{"schema_version":1,"generated_at":"2026-09-21T08:06:05+00:00","edge_direction":"Prerequisite to consumer for teaching/import/order edges; parent to child for contains edges.","evidence_boundary":"This is a generated navigation graph over source containment, module imports, reviewed teaching dependencies, milestone evidence, and textbook mappings. It is not a kernel trace or the frozen exact environment graph.","root":"library:banditrl","views":{"overview":["library:banditrl","group:book-map","group:textbook-spine","group:milestones","group:proof-laboratory"],"book":["library:banditrl","group:book-map","chapter:foundations","chapter:probability","chapter:etc","chapter:ucb","chapter:oful","chapter:thompson","chapter:exp3","chapter:tsallis","chapter:finite-horizon-rl","chapter:frontier"],"spine":["library:banditrl","group:textbook-spine","spine:13","spine:14","spine:15","spine:16","spine:17"],"milestones":["library:banditrl","group:milestones"]},"nodes":[{"id":"library:banditrl","label":"BanditRLlib","kind":"library","status":"compiled","subtitle":"Lean 4 library for bandit and reinforcement-learning theory","description":"The public, searchable library produced by ABRL. Expand one branch at a time to move from teaching chapters to modules, exact declarations, and reviewed routes.","url":"../index.html","parent":"","order":0,"meta":[["Lean modules","796"],["Indexed declarations","10,434"],["Mapped milestones","97"]],"statement":"","missing":[],"search":"banditrllib lean 4 library for bandit and reinforcement-learning theory the public, searchable library produced by abrl. expand one branch at a time to move from teaching chapters to modules, exact declarations, and reviewed routes. library compiled","shard":"overview.json"},{"id":"group:book-map","label":"Bandit Book · Teaching routes","kind":"curriculum","status":"partial","subtitle":"10 teaching chapters","description":"The student-facing route from foundations and concentration to bandit algorithms and finite-horizon RL.","url":"../learning/index.html","parent":"library:banditrl","order":1,"meta":[["Chapters","10"],["Mapped declarations","10,434"]],"statement":"","missing":[],"search":"bandit book · teaching routes 10 teaching chapters the student-facing route from foundations and concentration to bandit algorithms and finite-horizon rl. curriculum partial","shard":"views/book.json"},{"id":"group:textbook-spine","label":"Bandit Book · Source Ch.13–17","kind":"textbook spine","status":"partial","subtitle":"Chapters 13–17 · 5 source chapters","description":"The source-numbered group of the Bandit Book. Chapters 13–17 have merged required main-text contracts; optional exercises are not all complete and Chapter 17 uses explicit source corrections. Exact chapter and build status remain separate.","url":"../textbook-spine/index.html","parent":"library:banditrl","order":2,"meta":[["Source chapters","5"]],"statement":"","missing":[],"search":"bandit book · source ch.13–17 chapters 13–17 · 5 source chapters the source-numbered group of the bandit book. chapters 13–17 have merged required main-text contracts; optional exercises are not all complete and chapter 17 uses explicit source corrections. exact chapter and build status remain separate. textbook spine partial","shard":"views/spine.json"},{"id":"group:milestones","label":"Implementation milestones","kind":"milestone map","status":"partial","subtitle":"97 exact route contracts","description":"Reviewed mathematical routes with independent compiled, partial, planned, and blocked status.","url":"../implementation-map/index.html","parent":"library:banditrl","order":3,"meta":[["Compiled","86"],["Partial","8"],["Blocked","2"],["Planned","1"]],"statement":"","missing":[],"search":"implementation milestones 97 exact route contracts reviewed mathematical routes with independent compiled, partial, planned, and blocked status. milestone map partial","shard":"views/milestones.json"},{"id":"group:proof-laboratory","label":"Proof Graph Laboratory","kind":"research laboratory","status":"prototype","subtitle":"Frozen environment graph and proof-structure prototypes","description":"A separate evidence page for compiled-environment dependency counts, fixed benchmark supports, ZDD/hypergraph prototypes, and the partial Curvature–Noise–Gap candidate.","url":"../proof-graph-laboratory/index.html","parent":"library:banditrl","order":4,"meta":[["Frozen project nodes","13,512"],["Frozen direct edges","664,837"],["Benchmark roots","3"],["CNG status","partial"]],"statement":"","missing":[],"search":"proof graph laboratory frozen environment graph and proof-structure prototypes a separate evidence page for compiled-environment dependency counts, fixed benchmark supports, zdd/hypergraph prototypes, and the partial curvature–noise–gap candidate. research laboratory prototype","shard":"overview.json"},{"id":"chapter:foundations","label":"01 · Foundations","kind":"book chapter","status":"compiled","subtitle":"1. Finite bandits, traces, and regret","description":"The deterministic language shared by the entire project: finite models, action and reward traces, pull counts, gaps, reward sums, pseudo-regret, and count-to-regret decompositions.","url":"../chapters/foundations/index.html","parent":"group:book-map","order":1,"meta":[["Lean modules","223"],["Declarations","2,579"],["Milestones","14"]],"statement":"","missing":["The deterministic identities are reusable; each probabilistic algorithm still needs its own measurable generated law and integrability contracts.","A local model-facing endpoint need not match every textbook or upstream theorem signature."],"search":"01 · foundations 1. finite bandits, traces, and regret the deterministic language shared by the entire project: finite models, action and reward traces, pull counts, gaps, reward sums, pseudo-regret, and count-to-regret decompositions. book chapter compiled","shard":"views/book.json"},{"id":"chapter:probability","label":"02 · Probability layer","kind":"book chapter","status":"compiled","subtitle":"2. Probability, kernels, filtrations, and concentration","description":"Measure-theoretic infrastructure for generated histories, conditional reward laws, martingale differences, posterior kernels, stopping times, and concentration.","url":"../chapters/probability/index.html","parent":"group:book-map","order":2,"meta":[["Lean modules","46"],["Declarations","807"],["Milestones","7"]],"statement":"","missing":["The telescoping all-time empirical-mean producer now has an explicit same-source fixed-policy UCB consumer; the geometric producer remains confidence infrastructure rather than the logarithmic scheduled-regret route.","There is intentionally no single producer for every adaptive environment or every fixed-policy anytime algorithm.","Self-normalized and RL-specific concentration remain explicit downstream theorem routes rather than one universal wrapper."],"search":"02 · probability layer 2. probability, kernels, filtrations, and concentration measure-theoretic infrastructure for generated histories, conditional reward laws, martingale differences, posterior kernels, stopping times, and concentration. book chapter compiled","shard":"views/book.json"},{"id":"chapter:etc","label":"03 · ETC","kind":"book chapter","status":"compiled","subtitle":"3. Explore-Then-Commit","description":"Round-robin exploration, empirical means, measurable commit choices, wrong-commit tails, and expected-regret assemblies under bounded or sub-Gaussian arm laws.","url":"../chapters/etc/index.html","parent":"group:book-map","order":3,"meta":[["Lean modules","37"],["Declarations","339"],["Milestones","3"]],"statement":"","missing":["Direct LML symbol integration remains optional cross-toolchain work even though the local canonical chapter route compiles.","The website never upgrades a theorem card into a local proof certificate."],"search":"03 · etc 3. explore-then-commit round-robin exploration, empirical means, measurable commit choices, wrong-commit tails, and expected-regret assemblies under bounded or sub-gaussian arm laws. book chapter compiled","shard":"views/book.json"},{"id":"chapter:ucb","label":"04 · UCB","kind":"book chapter","status":"compiled","subtitle":"4. UCB: confidence events to regret","description":"History-based ordinary-UCB and KL-UCB scores, count thresholds, generated policies, arm streams, same-trajectory confidence, finite-arm reward kernels, and expected regret.","url":"../chapters/ucb/index.html","parent":"group:book-map","order":4,"meta":[["Lean modules","35"],["Declarations","586"],["Milestones","19"]],"statement":"","missing":["The fixed-policy telescoping anytime confidence/count/finite-time expected-regret route is compiled. Its bound retains the explicit T times delta failure term, so fixed-delta expected-average consistency is not claimed.","The geometric all-time radius is retained as a valid confidence producer but its exponentially shrinking share is not the route to logarithmic fixed-policy UCB regret.","A distinct generated KL-UCB extension now compiles Bernoulli-KL endpoint semantics, confidence-set supremum, a horizon-free policy, same-source all-time confidence, all-horizon counts, and conservative finite-time expected pseudo-regret under AE unit support and a common interior-mean margin.","Sharp KL-Chernoff concentration, the Garivier-Cappe leading constant, KL-UCB limsup asymptotic optimality, and exact pinned-LML compatibility remain separate."],"search":"04 · ucb 4. ucb: confidence events to regret history-based ordinary-ucb and kl-ucb scores, count thresholds, generated policies, arm streams, same-trajectory confidence, finite-arm reward kernels, and expected regret. book chapter compiled","shard":"views/book.json"},{"id":"chapter:oful","label":"05 · OFUL","kind":"book chapter","status":"compiled","subtitle":"5. OFUL, self-normalized confidence, and stopping times","description":"The scoped canonical finite-action linear-bandit route compiles from elliptical potential and self-normalized ridge confidence through one horizon-free generated OFUL policy with all-horizon regret and stopping consumers, plus a separately identified horizon-indexed expected-consistency family.","url":"../chapters/oful/index.html","parent":"group:book-map","order":5,"meta":[["Lean modules","65"],["Declarations","706"],["Milestones","8"]],"statement":"","missing":["Contextual or time-varying action sets, dynamic linear bandits, paper-sharp/minimax constants, uniform-over-parameter guarantees, and infinite-dimensional Hilbert-space OFUL remain extensions.","Arbitrary history environments without a centered conditional-MGF producer, pathwise/almost-sure/universal optional-stopping consistency, and full primal-dual Bandits-with-Knapsacks remain outside the completed scope.","Budget-forced schedules compile under explicit contracts, but they are not labeled as a complete BwK theorem."],"search":"05 · oful 5. oful, self-normalized confidence, and stopping times the scoped canonical finite-action linear-bandit route compiles from elliptical potential and self-normalized ridge confidence through one horizon-free generated oful policy with all-horizon regret and stopping consumers, plus a separately identified horizon-indexed expected-consistency family. book chapter compiled","shard":"views/book.json"},{"id":"chapter:thompson","label":"06 · Thompson sampling","kind":"book chapter","status":"compiled","subtitle":"6. Thompson sampling and Bayesian regret","description":"The scoped stationary finite-arm Thompson route compiles from posterior kernels and probability matching on the actual recursive generated history through clipped-UCB decomposition, latent-stream confidence, and an explicit Bayesian regret terminal.","url":"../chapters/thompson/index.html","parent":"group:book-map","order":6,"meta":[["Lean modules","11"],["Declarations","297"],["Milestones","7"]],"statement":"","missing":["Arbitrary nonstationary posterior models, contextual or linear Thompson sampling, posterior-sampling RL, and user-supplied posteriors without a law producer remain extensions.","Sharp problem-dependent or asymptotically optimal constants are not claimed by the compiled stationary terminal.","Exact LeanMachineLearning declaration/toolchain identity remains independently blocked; LML cards are retrieval evidence, not imported proof terms."],"search":"06 · thompson sampling 6. thompson sampling and bayesian regret the scoped stationary finite-arm thompson route compiles from posterior kernels and probability matching on the actual recursive generated history through clipped-ucb decomposition, latent-stream confidence, and an explicit bayesian regret terminal. book chapter compiled","shard":"views/book.json"},{"id":"chapter:exp3","label":"07 · EXP3","kind":"book chapter","status":"compiled","subtitle":"7. EXP3 and adversarial concentration","description":"The scoped canonical generated EXP3 route compiles from exponential-weight potentials and importance-weighted conditional moments through horizon-tuned expected and best-arm high-probability endpoints, plus a distinct fixed-process all-positive-prefix realized-regret event and a sparse-loss extension.","url":"../chapters/exp3/index.html","parent":"group:book-map","order":7,"meta":[["Lean modules","86"],["Declarations","841"],["Milestones","7"]],"statement":"","missing":["The tuned expected and fixed-window best-arm theorems rebuild eta, gamma, and the generated law from the queried horizon; they are not a single horizon-free policy or simultaneous anytime theorem.","The geometric all-prefix theorem instead fixes eta, gamma, prior, arms, loss, and one supported comparator; its scheduled radius is not a tuned sublinear confidence sequence or a best-arm minimum.","Ville/Doob or mixture boundaries, optional stopping, ideal EXP3.P, contextual/delayed EXP3, and one universal theorem subsuming every hypothesis regime remain extensions. The sparse theorem retains its supplied sparsity-failure probability."],"search":"07 · exp3 7. exp3 and adversarial concentration the scoped canonical generated exp3 route compiles from exponential-weight potentials and importance-weighted conditional moments through horizon-tuned expected and best-arm high-probability endpoints, plus a distinct fixed-process all-positive-prefix realized-regret event and a sparse-loss extension. book chapter compiled","shard":"views/book.json"},{"id":"chapter:tsallis","label":"08 · Tsallis-FTRL","kind":"book chapter","status":"compiled","subtitle":"8. Tsallis-FTRL, corruption, and nonstationarity","description":"The scoped canonical half-Tsallis FTRL route compiles from finite-simplex minimizers and one-step stability through a measurable scheduled generated trajectory, score alignment, expected self-bounding, and a finite-arm IID bounded reward-law logarithmic regret terminal; corruption and nonstationary routes remain labelled extensions.","url":"../chapters/tsallis/index.html","parent":"group:book-map","order":8,"meta":[["Lean modules","86"],["Declarations","896"],["Milestones","4"]],"statement":"","missing":["Paper-sharp or minimax-optimal Tsallis-INF constants, a complete best-of-both-worlds theorem, and high-probability or realized-regret guarantees are not claimed by the canonical expected-regret terminal.","History-adaptive corruption, drifting laws, dynamic comparators, and the population-mean oracle restart compile as separately labelled extensions; an observed-reward detector still needs random history-dependent scheduling plus delay/false-alarm concentration.","The strict Fin 2 refined-averaged-stability counterexample remains an explicit obstruction; contextual/linear Tsallis-INF and broader adaptive-learning-rate paper routes remain separate."],"search":"08 · tsallis-ftrl 8. tsallis-ftrl, corruption, and nonstationarity the scoped canonical half-tsallis ftrl route compiles from finite-simplex minimizers and one-step stability through a measurable scheduled generated trajectory, score alignment, expected self-bounding, and a finite-arm iid bounded reward-law logarithmic regret terminal; corruption and nonstationary routes remain labelled extensions. book chapter compiled","shard":"views/book.json"},{"id":"chapter:finite-horizon-rl","label":"09 · Finite-horizon RL","kind":"book chapter","status":"compiled","subtitle":"9. Finite-horizon reinforcement learning","description":"Finite MDPs, Bellman optimality, generated trajectories and occupancy regret, plus a canonical known-reward Hoeffding UCBVI-CH route whose recurrent planner, same-source confidence, optimism, raw cumulative episode pseudo-regret, high-probability terminal, and failure-aware expectation consumer share one adaptive generated process.","url":"../chapters/finite-horizon-rl/index.html","parent":"group:book-map","order":9,"meta":[["Lean modules","166"],["Declarations","2,697"],["Milestones","7"]],"statement":"","missing":["Bernstein/variance-aware minimax UCB-VI and its paper-leading rate remain a separate second milestone.","Stochastic-reward UCBVI, adversarial or nonstationary initial-state sequences, realized sampled-return high-probability UCBVI, posterior-sampling RL, model-free Q-learning, and infinite or continuous state-action spaces remain extensions.","Natural-causal consistency and stopping-time RL results are independent extension branches; they are not consequences or aliases of the compiled raw cumulative UCBVI-CH terminal.","The joint confidence event includes singleton Bernstein coordinates and a same-generated-law optimal-tail scalar probe. It does not claim the sharp scalar bound follows from singleton envelopes alone."],"search":"09 · finite-horizon rl 9. finite-horizon reinforcement learning finite mdps, bellman optimality, generated trajectories and occupancy regret, plus a canonical known-reward hoeffding ucbvi-ch route whose recurrent planner, same-source confidence, optimism, raw cumulative episode pseudo-regret, high-probability terminal, and failure-aware expectation consumer share one adaptive generated process. book chapter compiled","shard":"views/book.json"},{"id":"chapter:frontier","label":"10 · Frontier","kind":"book chapter","status":"planned","subtitle":"10. Automation, resources, and open routes","description":"The proof harness, task vocabulary, resource stopping leaves, literature registry, partial source-frozen delayed-feedback, succinct-lower-bound, and stochastic-gradient-bandit audits, and planned BwK, preference, robust, federated, neural-bandit, and sharp KL-asymptotic work.","url":"../chapters/frontier/index.html","parent":"group:book-map","order":10,"meta":[["Lean modules","41"],["Declarations","686"],["Milestones","21"]],"statement":"","missing":["The conservative generated KL-UCB finite-time route has named compiled declarations. The source-frozen delayed-bandit audit now compiles accounting, causal-view, new-arrival processing, probability allocation, the deterministic optimal-arm-survival core, and a causal one-round action measure. EAP/BSC state preservation, a measurable recursive trajectory, and stochastic/adversarial regret endpoints remain blocked. The source-frozen succinct-lower-bound audit compiles 54 declarations covering Definitions 3.1–3.3 and Lemmas 3.1–3.4, including the finite-Bessel strict-support route, plus an explicit global-R boundedness diagnostic; the global Lemmas 3.5–3.6 and Theorem 3.8 remain open. The stochastic-gradient-bandit audit retains the exact counted 361 = 223 + 23 + 25 + 26 + 7 + 8 + 13 + 28 + 8 audit-slice inventory through deterministic-time selected-reward freshness, terminal-count events, nth-pull-to-count bridges, and a generic low-count regret consumer. A separate ten-declaration module proves equality of the complete visible/native trajectory measures. The selected-block module now contains eight declarations for missing-pull-aware block transport, fourteen for the exact finite Appendix-C `S0/S1` event, ten for the exact disjoint all-present/missing-pull probability split, and four for missing-pull inclusion, visible-marginal probability transport, and the finite-horizon expected-regret charge. None is a selected-IID theorem or a positive-probability producer. Corollary 1 remains a direct Theorem-1 consumer, not Theorem-2 evidence. The frozen K = 2 Theorem-2 terminal remains blocked: the generated all-present phase still needs a fixed-chronological-cutoff trigger, the stopped-prefix future-cylinder law must yield conditional no-return probability at least one half, and the Rademacher/ballot probability plus asymptotic assembly remain uncompiled. The frozen terminal twoArmRademacherDirac_theoremTwo_polynomialRegret is not claimed. The Theorem-4 terminal also remains blocked because the general-K generated process, uniform buffer/survival producer, stopped-process argument, and regret assembly are absent. Full BwK/primal-dual regret, dueling, robust, federated, and neural-bandit routes remain planned or partial.","Direct LeanMachineLearning declaration identity remains blocked on a deliberate cross-toolchain migration and real upstream-symbol import; local theorem-card-shaped ETC/UCB proofs do not satisfy that gate.","The harness records completion gates but does not replace Lean elaboration or mathematical review."],"search":"10 · frontier 10. automation, resources, and open routes the proof harness, task vocabulary, resource stopping leaves, literature registry, partial source-frozen delayed-feedback, succinct-lower-bound, and stochastic-gradient-bandit audits, and planned bwk, preference, robust, federated, neural-bandit, and sharp kl-asymptotic work. book chapter planned","shard":"views/book.json"},{"id":"module:BanditRLProof","label":"BanditRLProof","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof","description":"Generated source map for BanditRLProof.lean.","url":"../modules/banditrlproof/index.html","parent":"chapter:foundations","order":0,"meta":[["Source","BanditRLProof.lean"],["Declarations","0"],["Project imports","700"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"banditrlproof banditrlproof generated source map for banditrlproof.lean. lean module compiled","shard":"modules/a703ba0de6e1bf2d.json"},{"id":"module:BanditRLProof.Algorithms.ArmStreamPolicy","label":"Algorithms.ArmStreamPolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ArmStreamPolicy","description":"A measurable history selector driven by the existing latent reward streams. This generalizes the UCB recursion without changing its reward/count semantics.","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html","parent":"chapter:foundations","order":1,"meta":[["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.armstreampolicy banditrlproof.algorithms.armstreampolicy a measurable history selector driven by the existing latent reward streams. this generalizes the ucb recursion without changing its reward/count semantics. lean module compiled","shard":"modules/8ebdcfdac06a056e.json"},{"id":"module:BanditRLProof.Algorithms.CUCBActualReward","label":"Algorithms.CUCBActualReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBActualReward","description":"Actual reward and true-score expectations on every horizon of the same CUCB trajectory. The joint action/feedback distribution is retained.","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html","parent":"chapter:foundations","order":2,"meta":[["Source","BanditRLProof/Algorithms/CUCBActualReward.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbactualreward banditrlproof.algorithms.cucbactualreward actual reward and true-score expectations on every horizon of the same cucb trajectory. the joint action/feedback distribution is retained. lean module compiled","shard":"modules/2c93a6e6b0be6d62.json"},{"id":"module:BanditRLProof.Algorithms.CUCBCharge","label":"Algorithms.CUCBCharge","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBCharge","description":"Integer analysis counters for the repaired source charging rule. These counters use only actions and fixed instance data, never current feedback. They are analysis objects, not an additional learner input.","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html","parent":"chapter:foundations","order":3,"meta":[["Source","BanditRLProof/Algorithms/CUCBCharge.lean"],["Declarations","20"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbcharge banditrlproof.algorithms.cucbcharge integer analysis counters for the repaired source charging rule. these counters use only actions and fixed instance data, never current feedback. they are analysis objects, not an additional learner input. lean module compiled","shard":"modules/5d47e3a0553cec38.json"},{"id":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","label":"Algorithms.CUCBChargedConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBChargedConcentration","description":"Accumulating charged-trigger exponential bounds on the actual CUCB path.","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html","parent":"chapter:foundations","order":4,"meta":[["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbchargedconcentration banditrlproof.algorithms.cucbchargedconcentration accumulating charged-trigger exponential bounds on the actual cucb path. lean module compiled","shard":"modules/076a5861e0ca3f9f.json"},{"id":"module:BanditRLProof.Algorithms.CUCBChargedConditional","label":"Algorithms.CUCBChargedConditional","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBChargedConditional","description":"Conditional trigger factor along the actual history-dependent charging recursion and actual randomized CUCB trajectory.","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html","parent":"chapter:foundations","order":5,"meta":[["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbchargedconditional banditrlproof.algorithms.cucbchargedconditional conditional trigger factor along the actual history-dependent charging recursion and actual randomized cucb trajectory. lean module compiled","shard":"modules/7f9b7b51d7849e54.json"},{"id":"module:BanditRLProof.Algorithms.CUCBChargedMGF","label":"Algorithms.CUCBChargedMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBChargedMGF","description":"The actual oracle mixture preserves the trigger bound for the normalized charge, which is chosen before the environment draws the current feedback.","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html","parent":"chapter:foundations","order":6,"meta":[["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbchargedmgf banditrlproof.algorithms.cucbchargedmgf the actual oracle mixture preserves the trigger bound for the normalized charge, which is chosen before the environment draws the current feedback. lean module compiled","shard":"modules/68465bba9f63c3af.json"},{"id":"module:BanditRLProof.Algorithms.CUCBConcentration","label":"Algorithms.CUCBConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBConcentration","description":"Count-compensated concentration on the actual randomized CUCB path. Both initial and conditional MGF premises are produced from its environment.","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html","parent":"chapter:foundations","order":7,"meta":[["Source","BanditRLProof/Algorithms/CUCBConcentration.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbconcentration banditrlproof.algorithms.cucbconcentration count-compensated concentration on the actual randomized cucb path. both initial and conditional mgf premises are produced from its environment. lean module compiled","shard":"modules/2d562cf58d444a24.json"},{"id":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","label":"Algorithms.CUCBConditionalMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBConditionalMGF","description":"Conditional compensated MGF for the actual CUCB trajectory.","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html","parent":"chapter:foundations","order":8,"meta":[["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbconditionalmgf banditrlproof.algorithms.cucbconditionalmgf conditional compensated mgf for the actual cucb trajectory. lean module compiled","shard":"modules/d415b5a7946463ea.json"},{"id":"module:BanditRLProof.Algorithms.CUCBConfidence","label":"Algorithms.CUCBConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBConfidence","description":"Peeling over the actual random number of observed outcomes.","url":"../modules/banditrlproof-algorithms-cucbconfidence/index.html","parent":"chapter:foundations","order":9,"meta":[["Source","BanditRLProof/Algorithms/CUCBConfidence.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbconfidence banditrlproof.algorithms.cucbconfidence peeling over the actual random number of observed outcomes. lean module compiled","shard":"modules/29afa36b1f71f0fe.json"},{"id":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","label":"Algorithms.CUCBDeterministicTrigger","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBDeterministicTrigger","description":"Deterministic triggering on the actual CUCB law, including all horizons. No claim is made about feedback records in null sets.","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html","parent":"chapter:foundations","order":10,"meta":[["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbdeterministictrigger banditrlproof.algorithms.cucbdeterministictrigger deterministic triggering on the actual cucb law, including all horizons. no claim is made about feedback records in null sets. lean module compiled","shard":"modules/aa338b4ac58b9cc2.json"},{"id":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","label":"Algorithms.CUCBFeedbackModel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBFeedbackModel","description":"Finite feasible superarms and primitive triggered-feedback laws. Trigger minima are computed from actual environment probabilities.","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html","parent":"chapter:foundations","order":11,"meta":[["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbfeedbackmodel banditrlproof.algorithms.cucbfeedbackmodel finite feasible superarms and primitive triggered-feedback laws. trigger minima are computed from actual environment probabilities. lean module compiled","shard":"modules/3c74c73b1fcf73e7.json"},{"id":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","label":"Algorithms.CUCBFiniteConcavity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBFiniteConcavity","description":"The source finite-concavity obligation, instantiated with actual under-sampled charge counts, including zero counts and zero horizon.","url":"../modules/banditrlproof-algorithms-cucbfiniteconcavity/index.html","parent":"chapter:foundations","order":12,"meta":[["Source","BanditRLProof/Algorithms/CUCBFiniteConcavity.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbfiniteconcavity banditrlproof.algorithms.cucbfiniteconcavity the source finite-concavity obligation, instantiated with actual under-sampled charge counts, including zero counts and zero horizon. lean module compiled","shard":"modules/a022e5888d0985fd.json"},{"id":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","label":"Algorithms.CUCBFiniteDeterministicExample","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","description":"Deterministic-trigger recovery of the same noisy three-arm example.","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html","parent":"chapter:foundations","order":13,"meta":[["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbfinitedeterministicexample banditrlproof.algorithms.cucbfinitedeterministicexample deterministic-trigger recovery of the same noisy three-arm example. lean module compiled","shard":"modules/d2e3d092c7ae8876.json"},{"id":"module:BanditRLProof.Algorithms.CUCBFiniteExample","label":"Algorithms.CUCBFiniteExample","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBFiniteExample","description":"A finite noisy triggered-feedback witness. Independent primitive coordinates produce three Bernoulli arms and a fresh extra-trigger coin.","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html","parent":"chapter:foundations","order":14,"meta":[["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean"],["Declarations","21"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbfiniteexample banditrlproof.algorithms.cucbfiniteexample a finite noisy triggered-feedback witness. independent primitive coordinates produce three bernoulli arms and a fresh extra-trigger coin. lean module compiled","shard":"modules/a7217debf832ac74.json"},{"id":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","label":"Algorithms.CUCBFiniteSourceExample","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBFiniteSourceExample","description":"Concrete nonlinear score and input-dependent maximizing oracle for the finite noisy feedback model. Ties choose the inferior true-reward action.","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html","parent":"chapter:foundations","order":15,"meta":[["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean"],["Declarations","23"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbfinitesourceexample banditrlproof.algorithms.cucbfinitesourceexample concrete nonlinear score and input-dependent maximizing oracle for the finite noisy feedback model. ties choose the inferior true-reward action. lean module compiled","shard":"modules/f04f3933b1c8ff08.json"},{"id":"module:BanditRLProof.Algorithms.CUCBGapCutoff","label":"Algorithms.CUCBGapCutoff","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBGapCutoff","description":"Actual gap cutoff bound used by the distribution-independent CUCB proof. The baseline is paid once per round, not once per arm threshold crossing.","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html","parent":"chapter:foundations","order":16,"meta":[["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbgapcutoff banditrlproof.algorithms.cucbgapcutoff actual gap cutoff bound used by the distribution-independent cucb proof. the baseline is paid once per round, not once per arm threshold crossing. lean module compiled","shard":"modules/ca6b8cc77d816ca7.json"},{"id":"module:BanditRLProof.Algorithms.CUCBGapInverse","label":"Algorithms.CUCBGapInverse","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBGapInverse","description":"Scalar inverse on the full frozen positive-gap interval, and the exact integrable sampling threshold used in the refined source regret integral.","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html","parent":"chapter:foundations","order":17,"meta":[["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbgapinverse banditrlproof.algorithms.cucbgapinverse scalar inverse on the full frozen positive-gap interval, and the exact integrable sampling threshold used in the refined source regret integral. lean module compiled","shard":"modules/b3d4b0f63cd17de0.json"},{"id":"module:BanditRLProof.Algorithms.CUCBHistory","label":"Algorithms.CUCBHistory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBHistory","description":"The source CUCB statistics computed from chronological triggered feedback. Index n uses exactly rounds 0,...,n-1, so its source round number is n+1. The reward coordinate is not used to estimate individual arm means.","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html","parent":"chapter:foundations","order":18,"meta":[["Source","BanditRLProof/Algorithms/CUCBHistory.lean"],["Declarations","28"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbhistory banditrlproof.algorithms.cucbhistory the source cucb statistics computed from chronological triggered feedback. index n uses exactly rounds 0,...,n-1, so its source round number is n+1. the reward coordinate is not used to estimate individual arm means. lean module compiled","shard":"modules/49fc793c1f05d1cd.json"},{"id":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","label":"Algorithms.CUCBImpossibleCase","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBImpossibleCase","description":"The source impossible-case argument with actual CUCB indices, true score gaps, finite possible-trigger sets and the source sampling constant six.","url":"../modules/banditrlproof-algorithms-cucbimpossiblecase/index.html","parent":"chapter:foundations","order":19,"meta":[["Source","BanditRLProof/Algorithms/CUCBImpossibleCase.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbimpossiblecase banditrlproof.algorithms.cucbimpossiblecase the source impossible-case argument with actual cucb indices, true score gaps, finite possible-trigger sets and the source sampling constant six. lean module compiled","shard":"modules/e948bf39a752d49c.json"},{"id":"module:BanditRLProof.Algorithms.CUCBNiceEvent","label":"Algorithms.CUCBNiceEvent","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBNiceEvent","description":"The clipped empirical-mean confidence event on the actual CUCB path. The history length is n; the source decision round is n+1.","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html","parent":"chapter:foundations","order":20,"meta":[["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbniceevent banditrlproof.algorithms.cucbniceevent the clipped empirical-mean confidence event on the actual cucb path. the history length is n; the source decision round is n+1. lean module compiled","shard":"modules/9a587b170484b277.json"},{"id":"module:BanditRLProof.Algorithms.CUCBObservationMGF","label":"Algorithms.CUCBObservationMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBObservationMGF","description":"Bounded-outcome MGF produced from the primitive uncensored marginal law. The current random observation mask remains inside the exponent.","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html","parent":"chapter:foundations","order":21,"meta":[["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbobservationmgf banditrlproof.algorithms.cucbobservationmgf bounded-outcome mgf produced from the primitive uncensored marginal law. the current random observation mask remains inside the exponent. lean module compiled","shard":"modules/559cf5080721c22b.json"},{"id":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","label":"Algorithms.CUCBOracleMeasurable","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBOracleMeasurable","description":"Source bounded smoothness produces score continuity and measurable oracle events; no extra score-measurability assumption is added to the model.","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html","parent":"chapter:foundations","order":22,"meta":[["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucboraclemeasurable banditrlproof.algorithms.cucboraclemeasurable source bounded smoothness produces score continuity and measurable oracle events; no extra score-measurability assumption is added to the model. lean module compiled","shard":"modules/5aaf77eb214b613d.json"},{"id":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","label":"Algorithms.CUCBOracleSuccess","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBOracleSuccess","description":"Per-input approximate oracle success lifted to the actual initial and conditional successor laws, with the actual history-dependent oracle input.","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html","parent":"chapter:foundations","order":23,"meta":[["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucboraclesuccess banditrlproof.algorithms.cucboraclesuccess per-input approximate oracle success lifted to the actual initial and conditional successor laws, with the actual history-dependent oracle input. lean module compiled","shard":"modules/9720d8ff492c808f.json"},{"id":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","label":"Algorithms.CUCBPolynomialIntegral","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBPolynomialIntegral","description":"Integration of the actual source thresholds under a polynomial modulus.","url":"../modules/banditrlproof-algorithms-cucbpolynomialintegral/index.html","parent":"chapter:foundations","order":24,"meta":[["Source","BanditRLProof/Algorithms/CUCBPolynomialIntegral.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbpolynomialintegral banditrlproof.algorithms.cucbpolynomialintegral integration of the actual source thresholds under a polynomial modulus. lean module compiled","shard":"modules/6d2a3d7dc92f5f06.json"},{"id":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","label":"Algorithms.CUCBPolynomialRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBPolynomialRegret","description":"Exact source polynomial-smoothness endpoints, including small horizons.","url":"../modules/banditrlproof-algorithms-cucbpolynomialregret/index.html","parent":"chapter:foundations","order":25,"meta":[["Source","BanditRLProof/Algorithms/CUCBPolynomialRegret.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbpolynomialregret banditrlproof.algorithms.cucbpolynomialregret exact source polynomial-smoothness endpoints, including small horizons. lean module compiled","shard":"modules/3e4ea64c4acc8035.json"},{"id":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","label":"Algorithms.CUCBPolynomialThreshold","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBPolynomialThreshold","description":"Exact polynomial-modulus inverse and source threshold expressions.","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html","parent":"chapter:foundations","order":26,"meta":[["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbpolynomialthreshold banditrlproof.algorithms.cucbpolynomialthreshold exact polynomial-modulus inverse and source threshold expressions. lean module compiled","shard":"modules/d24a7d0182819f29.json"},{"id":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","label":"Algorithms.CUCBRefinedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBRefinedRegret","description":"The full refined integral approximation-regret endpoint, with the disclosed normalized analysis-counter repair and the unchanged source CUCB learner.","url":"../modules/banditrlproof-algorithms-cucbrefinedregret/index.html","parent":"chapter:foundations","order":27,"meta":[["Source","BanditRLProof/Algorithms/CUCBRefinedRegret.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbrefinedregret banditrlproof.algorithms.cucbrefinedregret the full refined integral approximation-regret endpoint, with the disclosed normalized analysis-counter repair and the unchanged source cucb learner. lean module compiled","shard":"modules/0afac4c1b20b09ab.json"},{"id":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","label":"Algorithms.CUCBRegretDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBRegretDecomposition","description":"Actual charged-gap decomposition preserving the signed oracle-failure credit.","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html","parent":"chapter:foundations","order":28,"meta":[["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbregretdecomposition banditrlproof.algorithms.cucbregretdecomposition actual charged-gap decomposition preserving the signed oracle-failure credit. lean module compiled","shard":"modules/25775f8ae73dfb18.json"},{"id":"module:BanditRLProof.Algorithms.CUCBRegretTail","label":"Algorithms.CUCBRegretTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBRegretTail","description":"Exact cumulative source tail and signed regret reduction to the actual under-sampled charge weight. The refined gap integral remains separate.","url":"../modules/banditrlproof-algorithms-cucbregrettail/index.html","parent":"chapter:foundations","order":29,"meta":[["Source","BanditRLProof/Algorithms/CUCBRegretTail.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbregrettail banditrlproof.algorithms.cucbregrettail exact cumulative source tail and signed regret reduction to the actual under-sampled charge weight. the refined gap integral remains separate. lean module compiled","shard":"modules/ff9f71f1783d3b37.json"},{"id":"module:BanditRLProof.Algorithms.CUCBRewardKernel","label":"Algorithms.CUCBRewardKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBRewardKernel","description":"Actual reward integrability and round expectations from the primitive nonnegative L1 reward law. No bound on realized rewards is imposed.","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html","parent":"chapter:foundations","order":30,"meta":[["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbrewardkernel banditrlproof.algorithms.cucbrewardkernel actual reward integrability and round expectations from the primitive nonnegative l1 reward law. no bound on realized rewards is imposed. lean module compiled","shard":"modules/9f7508adc9e5df3b.json"},{"id":"module:BanditRLProof.Algorithms.CUCBRoundMGF","label":"Algorithms.CUCBRoundMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBRoundMGF","description":"Integrating the primitive masked MGF over the actual oracle action. The action and its triggered feedback retain their joint round kernel.","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html","parent":"chapter:foundations","order":31,"meta":[["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbroundmgf banditrlproof.algorithms.cucbroundmgf integrating the primitive masked mgf over the actual oracle action. the action and its triggered feedback retain their joint round kernel. lean module compiled","shard":"modules/6d3d75cbcb1c4711.json"},{"id":"module:BanditRLProof.Algorithms.CUCBSourceModel","label":"Algorithms.CUCBSourceModel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBSourceModel","description":"The frozen full triggered-CUCB source model: arbitrary finite feasible superarms, nonlinear smooth reward scores and randomized approximation oracle. No confidence, counting or regret premise is part of this structure.","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html","parent":"chapter:foundations","order":32,"meta":[["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbsourcemodel banditrlproof.algorithms.cucbsourcemodel the frozen full triggered-cucb source model: arbitrary finite feasible superarms, nonlinear smooth reward scores and randomized approximation oracle. no confidence, counting or regret premise is part of this structure. lean module compiled","shard":"modules/9daa49c94010d214.json"},{"id":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","label":"Algorithms.CUCBSufficientSampling","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBSufficientSampling","description":"The actual successful-oracle bad-action event after the normalized charged counter crosses its source threshold. The nice-event and trigger-tail constants are retained, including the probability-one observation branch.","url":"../modules/banditrlproof-algorithms-cucbsufficientsampling/index.html","parent":"chapter:foundations","order":33,"meta":[["Source","BanditRLProof/Algorithms/CUCBSufficientSampling.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbsufficientsampling banditrlproof.algorithms.cucbsufficientsampling the actual successful-oracle bad-action event after the normalized charged counter crosses its source threshold. the nice-event and trigger-tail constants are retained, including the probability-one observation branch. lean module compiled","shard":"modules/d6ed59ac96fd3dd0.json"},{"id":"module:BanditRLProof.Algorithms.CUCBThreshold","label":"Algorithms.CUCBThreshold","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBThreshold","description":"The exact piecewise sampling threshold of Chen et al. (JMLR 2016). Normalized analysis counters repair the mixed triggering-probability step. This file does not assert a CUCB trajectory or regret theorem.","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html","parent":"chapter:foundations","order":34,"meta":[["Source","BanditRLProof/Algorithms/CUCBThreshold.lean"],["Declarations","10"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbthreshold banditrlproof.algorithms.cucbthreshold the exact piecewise sampling threshold of chen et al. (jmlr 2016). normalized analysis counters repair the mixed triggering-probability step. this file does not assert a cucb trajectory or regret theorem. lean module compiled","shard":"modules/486f0f6275048d07.json"},{"id":"module:BanditRLProof.Algorithms.CUCBThresholdTail","label":"Algorithms.CUCBThresholdTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBThresholdTail","description":"Exact source threshold crossing converted into an actual fixed-count observation-shortfall probability, with decision time n+1 and horizon H.","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html","parent":"chapter:foundations","order":35,"meta":[["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbthresholdtail banditrlproof.algorithms.cucbthresholdtail exact source threshold crossing converted into an actual fixed-count observation-shortfall probability, with decision time n+1 and horizon h. lean module compiled","shard":"modules/88212ea011b4dcd4.json"},{"id":"module:BanditRLProof.Algorithms.CUCBTrajectory","label":"Algorithms.CUCBTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBTrajectory","description":"The actual CUCB trajectory first samples a feasible oracle action, then its fresh triggered feedback. Neither concentration nor oracle success is assumed by this construction; those properties must be derived separately.","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html","parent":"chapter:foundations","order":36,"meta":[["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbtrajectory banditrlproof.algorithms.cucbtrajectory the actual cucb trajectory first samples a feasible oracle action, then its fresh triggered feedback. neither concentration nor oracle success is assumed by this construction; those properties must be derived separately. lean module compiled","shard":"modules/98705dd65b9428c0.json"},{"id":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","label":"Algorithms.CUCBTriggerMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBTriggerMGF","description":"Primitive Bernoulli exponential bound for actual triggered observations. No stopping-time or adaptive-count tail is assumed here.","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html","parent":"chapter:foundations","order":37,"meta":[["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbtriggermgf banditrlproof.algorithms.cucbtriggermgf primitive bernoulli exponential bound for actual triggered observations. no stopping-time or adaptive-count tail is assumed here. lean module compiled","shard":"modules/e593aad3b3b259c2.json"},{"id":"module:BanditRLProof.Algorithms.CUCBUnderCount","label":"Algorithms.CUCBUnderCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBUnderCount","description":"Actual distinct charged counters and the refined under-sampling gap-tail cardinality. No count bound is supplied as a source-model hypothesis.","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html","parent":"chapter:foundations","order":38,"meta":[["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbundercount banditrlproof.algorithms.cucbundercount actual distinct charged counters and the refined under-sampling gap-tail cardinality. no count bound is supplied as a source-model hypothesis. lean module compiled","shard":"modules/859dcab10a79ce60.json"},{"id":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","label":"Algorithms.CUCBUnderCountIntegral","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CUCBUnderCountIntegral","description":"Refined integral bound for the actual under-sampled charge weights.","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html","parent":"chapter:foundations","order":39,"meta":[["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.cucbundercountintegral banditrlproof.algorithms.cucbundercountintegral refined integral bound for the actual under-sampled charge weights. lean module compiled","shard":"modules/f90a78b0865d09b2.json"},{"id":"module:BanditRLProof.Algorithms.CausalAllocation","label":"Algorithms.CausalAllocation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalAllocation","description":"Allocation bounds for the actual finite mixture design objective.","url":"../modules/banditrlproof-algorithms-causalallocation/index.html","parent":"chapter:foundations","order":40,"meta":[["Source","BanditRLProof/Algorithms/CausalAllocation.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalallocation banditrlproof.algorithms.causalallocation allocation bounds for the actual finite mixture design objective. lean module compiled","shard":"modules/17b3f0149e2d074b.json"},{"id":"module:BanditRLProof.Algorithms.CausalAllocationRegret","label":"Algorithms.CausalAllocationRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalAllocationRegret","description":"The actual learner under uniform and attained optimal covered allocations.","url":"../modules/banditrlproof-algorithms-causalallocationregret/index.html","parent":"chapter:foundations","order":41,"meta":[["Source","BanditRLProof/Algorithms/CausalAllocationRegret.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalallocationregret banditrlproof.algorithms.causalallocationregret the actual learner under uniform and attained optimal covered allocations. lean module compiled","shard":"modules/437c067495ee2b84.json"},{"id":"module:BanditRLProof.Algorithms.CausalConfidence","label":"Algorithms.CausalConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalConfidence","description":"Source-tuned confidence for the actual fixed-budget intervention estimator.","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html","parent":"chapter:foundations","order":42,"meta":[["Source","BanditRLProof/Algorithms/CausalConfidence.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalconfidence banditrlproof.algorithms.causalconfidence source-tuned confidence for the actual fixed-budget intervention estimator. lean module compiled","shard":"modules/660aa4f2c75caaef.json"},{"id":"module:BanditRLProof.Algorithms.CausalExpectedRegret","label":"Algorithms.CausalExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalExpectedRegret","description":"Expected simple regret of the actual intervention sampler and recommendation. The finite-budget residual is retained explicitly.","url":"../modules/banditrlproof-algorithms-causalexpectedregret/index.html","parent":"chapter:foundations","order":43,"meta":[["Source","BanditRLProof/Algorithms/CausalExpectedRegret.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalexpectedregret banditrlproof.algorithms.causalexpectedregret expected simple regret of the actual intervention sampler and recommendation. the finite-budget residual is retained explicitly. lean module compiled","shard":"modules/f9be698157d645d3.json"},{"id":"module:BanditRLProof.Algorithms.CausalHeterogeneous","label":"Algorithms.CausalHeterogeneous","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalHeterogeneous","description":"Genuinely dependent node laws and their common-alphabet encoding.","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html","parent":"chapter:foundations","order":44,"meta":[["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean"],["Declarations","30"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalheterogeneous banditrlproof.algorithms.causalheterogeneous genuinely dependent node laws and their common-alphabet encoding. lean module compiled","shard":"modules/804f24a949b91951.json"},{"id":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","label":"Algorithms.CausalHeterogeneousLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalHeterogeneousLaw","description":"Native dependent joint factorization, independent of the common-alphabet target.","url":"../modules/banditrlproof-algorithms-causalheterogeneouslaw/index.html","parent":"chapter:foundations","order":45,"meta":[["Source","BanditRLProof/Algorithms/CausalHeterogeneousLaw.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalheterogeneouslaw banditrlproof.algorithms.causalheterogeneouslaw native dependent joint factorization, independent of the common-alphabet target. lean module compiled","shard":"modules/b3fd8fd5daadd708.json"},{"id":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","label":"Algorithms.CausalHeterogeneousRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalHeterogeneousRegret","description":"Actual expected simple regret on native heterogeneous finite DAGs.","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html","parent":"chapter:foundations","order":46,"meta":[["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalheterogeneousregret banditrlproof.algorithms.causalheterogeneousregret actual expected simple regret on native heterogeneous finite dags. lean module compiled","shard":"modules/d06124f51676b9e5.json"},{"id":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","label":"Algorithms.CausalHeterogeneousSampling","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalHeterogeneousSampling","description":"Native heterogeneous observations, estimates, and their actual sampling law.","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html","parent":"chapter:foundations","order":47,"meta":[["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean"],["Declarations","12"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalheterogeneoussampling banditrlproof.algorithms.causalheterogeneoussampling native heterogeneous observations, estimates, and their actual sampling law. lean module compiled","shard":"modules/21256850cfc44718.json"},{"id":"module:BanditRLProof.Algorithms.CausalImportance","label":"Algorithms.CausalImportance","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalImportance","description":"Covered finite mixtures and exact importance-weight identities.","url":"../modules/banditrlproof-algorithms-causalimportance/index.html","parent":"chapter:foundations","order":48,"meta":[["Source","BanditRLProof/Algorithms/CausalImportance.lean"],["Declarations","27"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalimportance banditrlproof.algorithms.causalimportance covered finite mixtures and exact importance-weight identities. lean module compiled","shard":"modules/9b0bfe2a017b5287.json"},{"id":"module:BanditRLProof.Algorithms.CausalImportanceTransport","label":"Algorithms.CausalImportanceTransport","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalImportanceTransport","description":"Importance ratios and design cost under injective finite-state encoding.","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html","parent":"chapter:foundations","order":49,"meta":[["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalimportancetransport banditrlproof.algorithms.causalimportancetransport importance ratios and design cost under injective finite-state encoding. lean module compiled","shard":"modules/452583592f22cd6a.json"},{"id":"module:BanditRLProof.Algorithms.CausalMarginalLaw","label":"Algorithms.CausalMarginalLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalMarginalLaw","description":"Marginal laws derived from topologically ordered sampling.","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html","parent":"chapter:foundations","order":50,"meta":[["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalmarginallaw banditrlproof.algorithms.causalmarginallaw marginal laws derived from topologically ordered sampling. lean module compiled","shard":"modules/82d9a180187b8ad8.json"},{"id":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","label":"Algorithms.CausalOptimalAllocation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalOptimalAllocation","description":"Attainment of the covered finite causal allocation objective.","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html","parent":"chapter:foundations","order":51,"meta":[["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean"],["Declarations","21"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causaloptimalallocation banditrlproof.algorithms.causaloptimalallocation attainment of the covered finite causal allocation objective. lean module compiled","shard":"modules/1e9c6c116515e2db.json"},{"id":"module:BanditRLProof.Algorithms.CausalOrderedLaw","label":"Algorithms.CausalOrderedLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalOrderedLaw","description":"Topologically ordered conditional tables and their actual joint PMF. Interventions replace node tables before sampling the joint law.","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html","parent":"chapter:foundations","order":52,"meta":[["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean"],["Declarations","14"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalorderedlaw banditrlproof.algorithms.causalorderedlaw topologically ordered conditional tables and their actual joint pmf. interventions replace node tables before sampling the joint law. lean module compiled","shard":"modules/2bfec485485647c1.json"},{"id":"module:BanditRLProof.Algorithms.CausalParallelDesign","label":"Algorithms.CausalParallelDesign","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalParallelDesign","description":"Constructed rarity parameter and normalized finite intervention design.","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html","parent":"chapter:foundations","order":53,"meta":[["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalparalleldesign banditrlproof.algorithms.causalparalleldesign constructed rarity parameter and normalized finite intervention design. lean module compiled","shard":"modules/33beb34818f8020b.json"},{"id":"module:BanditRLProof.Algorithms.CausalParallelLaw","label":"Algorithms.CausalParallelLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalParallelLaw","description":"Actual independent-root intervention laws and their covered allocation cost.","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html","parent":"chapter:foundations","order":54,"meta":[["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalparallellaw banditrlproof.algorithms.causalparallellaw actual independent-root intervention laws and their covered allocation cost. lean module compiled","shard":"modules/cecc5e09203a86a4.json"},{"id":"module:BanditRLProof.Algorithms.CausalParallelRegret","label":"Algorithms.CausalParallelRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalParallelRegret","description":"Parallel DAG adapter and actual allocation-dependent expected simple regret.","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html","parent":"chapter:foundations","order":55,"meta":[["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean"],["Declarations","11"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalparallelregret banditrlproof.algorithms.causalparallelregret parallel dag adapter and actual allocation-dependent expected simple regret. lean module compiled","shard":"modules/e0ecb1b3c65f64c3.json"},{"id":"module:BanditRLProof.Algorithms.CausalRecommendation","label":"Algorithms.CausalRecommendation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalRecommendation","description":"A fixed-order recommendation from observed estimates and its pathwise regret bound.","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html","parent":"chapter:foundations","order":56,"meta":[["Source","BanditRLProof/Algorithms/CausalRecommendation.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalrecommendation banditrlproof.algorithms.causalrecommendation a fixed-order recommendation from observed estimates and its pathwise regret bound. lean module compiled","shard":"modules/d51c6dd1d7db3a28.json"},{"id":"module:BanditRLProof.Algorithms.CausalSampleMGF","label":"Algorithms.CausalSampleMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalSampleMGF","description":"Centered MGF and signed sum tails on actual causal intervention samples.","url":"../modules/banditrlproof-algorithms-causalsamplemgf/index.html","parent":"chapter:foundations","order":57,"meta":[["Source","BanditRLProof/Algorithms/CausalSampleMGF.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalsamplemgf banditrlproof.algorithms.causalsamplemgf centered mgf and signed sum tails on actual causal intervention samples. lean module compiled","shard":"modules/b5ecf0464d8d2ccd.json"},{"id":"module:BanditRLProof.Algorithms.CausalSampling","label":"Algorithms.CausalSampling","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalSampling","description":"Actual intervention/assignment rounds and their fixed-budget product law.","url":"../modules/banditrlproof-algorithms-causalsampling/index.html","parent":"chapter:foundations","order":58,"meta":[["Source","BanditRLProof/Algorithms/CausalSampling.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causalsampling banditrlproof.algorithms.causalsampling actual intervention/assignment rounds and their fixed-budget product law. lean module compiled","shard":"modules/2d6837cf2bd017ea.json"},{"id":"module:BanditRLProof.Algorithms.CausalTuning","label":"Algorithms.CausalTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.CausalTuning","description":"The fixed source tuning used for the causal importance estimator.","url":"../modules/banditrlproof-algorithms-causaltuning/index.html","parent":"chapter:foundations","order":59,"meta":[["Source","BanditRLProof/Algorithms/CausalTuning.lean"],["Declarations","15"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.causaltuning banditrlproof.algorithms.causaltuning the fixed source tuning used for the causal importance estimator. lean module compiled","shard":"modules/25bf2ab2c7d635bc.json"},{"id":"module:BanditRLProof.Algorithms.ETC","label":"Algorithms.ETC","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETC","description":"Explore-Then-Commit surfaces","url":"../modules/banditrlproof-algorithms-etc/index.html","parent":"chapter:etc","order":60,"meta":[["Source","BanditRLProof/Algorithms/ETC.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etc banditrlproof.algorithms.etc explore-then-commit surfaces lean module compiled","shard":"modules/f261e9d5d3185843.json"},{"id":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","label":"Algorithms.ETCArgmaxOracle","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCArgmaxOracle","description":"This file promotes the concrete finite-argmax ETC commit-oracle route into a compiled deterministic leaf. It only constructs a score-maximizing oracle over Fin K -> Rat and proves the maximality certificate consumed by the existing wrong-commit event wrappers.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html","parent":"chapter:etc","order":61,"meta":[["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcargmaxoracle banditrlproof.algorithms.etcargmaxoracle this file promotes the concrete finite-argmax etc commit-oracle route into a compiled deterministic leaf. it only constructs a score-maximizing oracle over fin k -> rat and proves the maximality certificate consumed by the existing wrong-commit event wrappers. lean module compiled","shard":"modules/4479fa60fbcd9782.json"},{"id":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","label":"Algorithms.ETCBoundedRewardInfinitePiSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","description":"This module instantiates the action-matched bounded reward source contract for a fixed product-coordinate reward trace law. It is still a fixed-commit ETC source leaf: it introduces no adaptive filtration, random commit arm, or final expected-regret result.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html","parent":"chapter:etc","order":62,"meta":[["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean"],["Declarations","7"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcboundedrewardinfinitepisource banditrlproof.algorithms.etcboundedrewardinfinitepisource this module instantiates the action-matched bounded reward source contract for a fixed product-coordinate reward trace law. it is still a fixed-commit etc source leaf: it introduces no adaptive filtration, random commit arm, or final expected-regret result. lean module compiled","shard":"modules/d5d4f17da7e21a6e.json"},{"id":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","label":"Algorithms.ETCBoundedRewardSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCBoundedRewardSource","description":"This module packages the stochastic source facts consumed by the bounded-reward ETC wrong-commit route. It does not construct those facts from a concrete environment, product space, kernel, or filtration.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html","parent":"chapter:etc","order":63,"meta":[["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean"],["Declarations","17"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcboundedrewardsource banditrlproof.algorithms.etcboundedrewardsource this module packages the stochastic source facts consumed by the bounded-reward etc wrong-commit route. it does not construct those facts from a concrete environment, product space, kernel, or filtration. lean module compiled","shard":"modules/f2f7901b95f27b0c.json"},{"id":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","label":"Algorithms.ETCBoundedRewardSubGaussian","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","description":"This module connects Mathlib's bounded-variable Hoeffding lemma to the ETC reward-coordinate wrong-commit route. It still leaves the stochastic source of reward-coordinate independence, boundedness, and mean identities explicit.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html","parent":"chapter:etc","order":64,"meta":[["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean"],["Declarations","8"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcboundedrewardsubgaussian banditrlproof.algorithms.etcboundedrewardsubgaussian this module connects mathlib's bounded-variable hoeffding lemma to the etc reward-coordinate wrong-commit route. it still leaves the stochastic source of reward-coordinate independence, boundedness, and mean identities explicit. lean module compiled","shard":"modules/ecfd9ad299323fdc.json"},{"id":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","label":"Algorithms.ETCCenteredDiffCanonicalTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","description":"This module fixes the canonical exponential tail budget produced by the centered reward-difference independent sub-Gaussian route. It removes the need for downstream users to provide a separate tail-domination hypothesis when they choose the exact Mathlib sub-Gaussian bound as their tail function.","url":"../modules/banditrlproof-algorithms-etccentereddiffcanonicaltail/index.html","parent":"chapter:etc","order":65,"meta":[["Source","BanditRLProof/Algorithms/ETCCenteredDiffCanonicalTail.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etccentereddiffcanonicaltail banditrlproof.algorithms.etccentereddiffcanonicaltail this module fixes the canonical exponential tail budget produced by the centered reward-difference independent sub-gaussian route. it removes the need for downstream users to provide a separate tail-domination hypothesis when they choose the exact mathlib sub-gaussian bound as their tail function. lean module compiled","shard":"modules/631636627679e640.json"},{"id":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","label":"Algorithms.ETCCenteredDiffRewardIndependence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","description":"This module transfers time-coordinate independence of a stochastic reward trace through the deterministic centered pairwise reward-difference transform used by the ETC wrong-commit tail route.","url":"../modules/banditrlproof-algorithms-etccentereddiffrewardindependence/index.html","parent":"chapter:etc","order":66,"meta":[["Source","BanditRLProof/Algorithms/ETCCenteredDiffRewardIndependence.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etccentereddiffrewardindependence banditrlproof.algorithms.etccentereddiffrewardindependence this module transfers time-coordinate independence of a stochastic reward trace through the deterministic centered pairwise reward-difference transform used by the etc wrong-commit tail route. lean module compiled","shard":"modules/4547db0ccc8c596c.json"},{"id":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","label":"Algorithms.ETCCenteredDiffRewardSubGaussian","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","description":"This module transfers per-time centered reward sub-Gaussian witnesses through the deterministic centered pairwise reward-difference transform used by the ETC wrong-commit probability route.","url":"../modules/banditrlproof-algorithms-etccentereddiffrewardsubgaussian/index.html","parent":"chapter:etc","order":67,"meta":[["Source","BanditRLProof/Algorithms/ETCCenteredDiffRewardSubGaussian.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etccentereddiffrewardsubgaussian banditrlproof.algorithms.etccentereddiffrewardsubgaussian this module transfers per-time centered reward sub-gaussian witnesses through the deterministic centered pairwise reward-difference transform used by the etc wrong-commit probability route. lean module compiled","shard":"modules/40506a24111ce02e.json"},{"id":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","label":"Algorithms.ETCCenteredDiffSubGaussianWitnesses","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","description":"This module packages the concrete reward-law witnesses needed by the centered reward-difference ETC tail producer. It does not prove those witnesses from a reward distribution, filtration, or kernel. Instead, it gives downstream work one exact Lean-facing contract to target.","url":"../modules/banditrlproof-algorithms-etccentereddiffsubgaussianwitnesses/index.html","parent":"chapter:etc","order":68,"meta":[["Source","BanditRLProof/Algorithms/ETCCenteredDiffSubGaussianWitnesses.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etccentereddiffsubgaussianwitnesses banditrlproof.algorithms.etccentereddiffsubgaussianwitnesses this module packages the concrete reward-law witnesses needed by the centered reward-difference etc tail producer. it does not prove those witnesses from a reward distribution, filtration, or kernel. instead, it gives downstream work one exact lean-facing contract to target. lean module compiled","shard":"modules/e21481b8a0246af2.json"},{"id":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","label":"Algorithms.ETCCondSubGaussianWitnesses","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","description":"This module packages the conditional reward-law witnesses needed to use the compiled Mathlib conditional sub-Gaussian concentration wrapper on ETC centered reward differences. It does not prove those witnesses from a concrete reward kernel or conditional expectation identity.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html","parent":"chapter:etc","order":69,"meta":[["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean"],["Declarations","24"],["Project imports","4"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etccondsubgaussianwitnesses banditrlproof.algorithms.etccondsubgaussianwitnesses this module packages the conditional reward-law witnesses needed to use the compiled mathlib conditional sub-gaussian concentration wrapper on etc centered reward differences. it does not prove those witnesses from a concrete reward kernel or conditional expectation identity. lean module compiled","shard":"modules/6afd18aa9d81068c.json"},{"id":"module:BanditRLProof.Algorithms.ETCCountLemmas","label":"Algorithms.ETCCountLemmas","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCCountLemmas","description":"This module records small Explore-Then-Commit count facts over the existing round-robin exploration primitive. It deliberately stays below full ETC traces, commit behavior, probability, concentration, and regret bounds.","url":"../modules/banditrlproof-algorithms-etccountlemmas/index.html","parent":"chapter:etc","order":70,"meta":[["Source","BanditRLProof/Algorithms/ETCCountLemmas.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etccountlemmas banditrlproof.algorithms.etccountlemmas this module records small explore-then-commit count facts over the existing round-robin exploration primitive. it deliberately stays below full etc traces, commit behavior, probability, concentration, and regret bounds. lean module compiled","shard":"modules/ab29fc9835e88284.json"},{"id":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","label":"Algorithms.ETCEmpiricalMean","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCEmpiricalMean","description":"This module defines the deterministic empirical mean for a fixed-commit ETC trace at the configured exploration horizon. It deliberately stays below probability, measurability, and concentration assumptions.","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html","parent":"chapter:etc","order":71,"meta":[["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean"],["Declarations","6"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcempiricalmean banditrlproof.algorithms.etcempiricalmean this module defines the deterministic empirical mean for a fixed-commit etc trace at the configured exploration horizon. it deliberately stays below probability, measurability, and concentration assumptions. lean module compiled","shard":"modules/584d56d84b375281.json"},{"id":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","label":"Algorithms.ETCEmpiricalMeanMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","description":"This module starts wiring the deterministic fixed-commit ETC empirical-mean surface to stochastic reward traces. It proves numerator measurability and the first full empirical-mean measurability wrapper under an explicit division-by-constant measurability contract, then discharges that contract using the local Rat measurable-singleton wrapper, plus a coordinate-shaped wrapper for downstream event measurability. Argm…","url":"../modules/banditrlproof-algorithms-etcempiricalmeanmeasurability/index.html","parent":"chapter:etc","order":72,"meta":[["Source","BanditRLProof/Algorithms/ETCEmpiricalMeanMeasurability.lean"],["Declarations","4"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcempiricalmeanmeasurability banditrlproof.algorithms.etcempiricalmeanmeasurability this module starts wiring the deterministic fixed-commit etc empirical-mean surface to stochastic reward traces. it proves numerator measurability and the first full empirical-mean measurability wrapper under an explicit division-by-constant measurability contract, then discharges that contract using the local rat measurable-singleton wrapper, plus a coordinate-shaped wrapper for downstream event measurability. argmax wiring, concentration, and filtration remain separate leaves. lean module compiled","shard":"modules/1d18399676b48c7d.json"},{"id":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","label":"Algorithms.ETCExactSubGaussianTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCExactSubGaussianTail","description":"This module normalizes the canonical common-proxy ETC pairwise tail to the closed exponential used by the exact LML theorem route. It remains over the project's existing Rat reward-law model, with all random variables embedded in Real; transport to native Real reward kernels is downstream work.","url":"../modules/banditrlproof-algorithms-etcexactsubgaussiantail/index.html","parent":"chapter:etc","order":73,"meta":[["Source","BanditRLProof/Algorithms/ETCExactSubGaussianTail.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcexactsubgaussiantail banditrlproof.algorithms.etcexactsubgaussiantail this module normalizes the canonical common-proxy etc pairwise tail to the closed exponential used by the exact lml theorem route. it remains over the project's existing rat reward-law model, with all random variables embedded in real; transport to native real reward kernels is downstream work. lean module compiled","shard":"modules/62a6d8bad834f56d.json"},{"id":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","label":"Algorithms.ETCExpectedPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCExpectedPullCount","description":"This module integrates the deterministic pull-count formula for an Omega-indexed ETC commit selector. It isolates the exact interface between ETC counting and concentration: a later tail argument only has to bound the commit-fiber probability mu.real {omega | commit omega = a}.","url":"../modules/banditrlproof-algorithms-etcexpectedpullcount/index.html","parent":"chapter:etc","order":74,"meta":[["Source","BanditRLProof/Algorithms/ETCExpectedPullCount.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcexpectedpullcount banditrlproof.algorithms.etcexpectedpullcount this module integrates the deterministic pull-count formula for an omega-indexed etc commit selector. it isolates the exact interface between etc counting and concentration: a later tail argument only has to bound the commit-fiber probability mu.real {omega | commit omega = a}. lean module compiled","shard":"modules/09433af36a9049c5.json"},{"id":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","label":"Algorithms.ETCExpectedRegretAssembly","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCExpectedRegretAssembly","description":"This module lifts the pointwise wrong-commit regret bridge to the project's expectation surfaces: an ENNReal.ofReal lower-integral surrogate and an ordinary Real-valued Bochner integral wrapper. It does not introduce concentration, filtrations, or a final ETC theorem.","url":"../modules/banditrlproof-algorithms-etcexpectedregretassembly/index.html","parent":"chapter:etc","order":75,"meta":[["Source","BanditRLProof/Algorithms/ETCExpectedRegretAssembly.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcexpectedregretassembly banditrlproof.algorithms.etcexpectedregretassembly this module lifts the pointwise wrong-commit regret bridge to the project's expectation surfaces: an ennreal.ofreal lower-integral surrogate and an ordinary real-valued bochner integral wrapper. it does not introduce concentration, filtrations, or a final etc theorem. lean module compiled","shard":"modules/c64c4638920377a0.json"},{"id":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","label":"Algorithms.ETCFiniteArmRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","description":"This module turns bounded arm-indexed probability laws with the finite-bandit model means into the centered reward-kernel contract consumed by the canonical ETC trajectory theorem. It then exposes the resulting conditional sub-Gaussian MGF directly, without an abstract centered-law or variance-ceiling argument.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html","parent":"chapter:etc","order":76,"meta":[["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean"],["Declarations","51"],["Project imports","5"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcfinitearmrewardlaw banditrlproof.algorithms.etcfinitearmrewardlaw this module turns bounded arm-indexed probability laws with the finite-bandit model means into the centered reward-kernel contract consumed by the canonical etc trajectory theorem. it then exposes the resulting conditional sub-gaussian mgf directly, without an abstract centered-law or variance-ceiling argument. lean module compiled","shard":"modules/298234f9f4270afd.json"},{"id":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","label":"Algorithms.ETCGeneratedHistoryPolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","description":"This module constructs the canonical ETC action trace as a policy generated from finite reward histories. It closes the deterministic action-alignment layer needed before an adaptive reward law can be transported through the existing generated-action conditional-expectation surfaces.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html","parent":"chapter:etc","order":77,"meta":[["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean"],["Declarations","8"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcgeneratedhistorypolicy banditrlproof.algorithms.etcgeneratedhistorypolicy this module constructs the canonical etc action trace as a policy generated from finite reward histories. it closes the deterministic action-alignment layer needed before an adaptive reward law can be transported through the existing generated-action conditional-expectation surfaces. lean module compiled","shard":"modules/057f6f9dc38b5ea3.json"},{"id":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","label":"Algorithms.ETCInfinitePiExpectedRegretAssembly","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","description":"This module instantiates the abstract lower-integral wrong-commit regret assembly with the concrete finite-argmax oracle and the infinite-product bounded-reward wrong-commit probability bound. The Real/Bochner wrapper still remains fixed-product and fixed-exploration, not the final adaptive ETC theorem.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html","parent":"chapter:etc","order":78,"meta":[["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean"],["Declarations","27"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcinfinitepiexpectedregretassembly banditrlproof.algorithms.etcinfinitepiexpectedregretassembly this module instantiates the abstract lower-integral wrong-commit regret assembly with the concrete finite-argmax oracle and the infinite-product bounded-reward wrong-commit probability bound. the real/bochner wrapper still remains fixed-product and fixed-exploration, not the final adaptive etc theorem. lean module compiled","shard":"modules/69436358672ab94d.json"},{"id":"module:BanditRLProof.Algorithms.ETCMeasurability","label":"Algorithms.ETCMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCMeasurability","description":"This module starts the probability-facing ETC layer with event measurability. It deliberately avoids measures, probability inequalities, empirical means, filtrations, concentration, and final regret theorems.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html","parent":"chapter:etc","order":79,"meta":[["Source","BanditRLProof/Algorithms/ETCMeasurability.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcmeasurability banditrlproof.algorithms.etcmeasurability this module starts the probability-facing etc layer with event measurability. it deliberately avoids measures, probability inequalities, empirical means, filtrations, concentration, and final regret theorems. lean module compiled","shard":"modules/26c597b04de5215d.json"},{"id":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","label":"Algorithms.ETCPairwiseCenteredSubGaussianTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","description":"This module specializes the abstract ETC pairwise sub-Gaussian producer to the compiled centered reward-difference finite-sum event. It keeps the actual probabilistic reward-law work explicit: independence and sub-Gaussian witnesses for the concrete summands remain hypotheses.","url":"../modules/banditrlproof-algorithms-etcpairwisecenteredsubgaussiantail/index.html","parent":"chapter:etc","order":80,"meta":[["Source","BanditRLProof/Algorithms/ETCPairwiseCenteredSubGaussianTail.lean"],["Declarations","1"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcpairwisecenteredsubgaussiantail banditrlproof.algorithms.etcpairwisecenteredsubgaussiantail this module specializes the abstract etc pairwise sub-gaussian producer to the compiled centered reward-difference finite-sum event. it keeps the actual probabilistic reward-law work explicit: independence and sub-gaussian witnesses for the concrete summands remain hypotheses. lean module compiled","shard":"modules/44db1ed1ab6221d5.json"},{"id":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","label":"Algorithms.ETCPairwiseSubGaussianTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","description":"This module connects the reusable independent sub-Gaussian finite-sum tail wrapper to the fixed-commit ETC pairwise empirical-mean tail contract. It does not instantiate a reward law, prove independence for ETC rewards, add filtration, or prove final ETC regret.","url":"../modules/banditrlproof-algorithms-etcpairwisesubgaussiantail/index.html","parent":"chapter:etc","order":81,"meta":[["Source","BanditRLProof/Algorithms/ETCPairwiseSubGaussianTail.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcpairwisesubgaussiantail banditrlproof.algorithms.etcpairwisesubgaussiantail this module connects the reusable independent sub-gaussian finite-sum tail wrapper to the fixed-commit etc pairwise empirical-mean tail contract. it does not instantiate a reward law, prove independence for etc rewards, add filtration, or prove final etc regret. lean module compiled","shard":"modules/5f7b0c14eb92f5a2.json"},{"id":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","label":"Algorithms.ETCPairwiseTailContract","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCPairwiseTailContract","description":"This module introduces the narrow contract surface needed between concrete ETC empirical means and the already compiled concrete argmax wrong-commit probability wrapper. It deliberately packages the abstract pairwise-tail hypothesis without proving any concentration theorem, adding filtration, or proving final ETC regret.","url":"../modules/banditrlproof-algorithms-etcpairwisetailcontract/index.html","parent":"chapter:etc","order":82,"meta":[["Source","BanditRLProof/Algorithms/ETCPairwiseTailContract.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcpairwisetailcontract banditrlproof.algorithms.etcpairwisetailcontract this module introduces the narrow contract surface needed between concrete etc empirical means and the already compiled concrete argmax wrong-commit probability wrapper. it deliberately packages the abstract pairwise-tail hypothesis without proving any concentration theorem, adding filtration, or proving final etc regret. lean module compiled","shard":"modules/d038bef37a47bcd6.json"},{"id":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","label":"Algorithms.ETCRatArmLawRealKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRatArmLawRealKernel","description":"This module pushes the existing Rat arm laws forward along the cast to Real, identifies the resulting identity-integral kernel means and gaps, and then assembles the canonical exact per-arm ETC pull-count bounds into the LML-shaped finite sum for Real kernel regret.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html","parent":"chapter:etc","order":83,"meta":[["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean"],["Declarations","8"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcratarmlawrealkernel banditrlproof.algorithms.etcratarmlawrealkernel this module pushes the existing rat arm laws forward along the cast to real, identifies the resulting identity-integral kernel means and gaps, and then assembles the canonical exact per-arm etc pull-count bounds into the lml-shaped finite sum for real kernel regret. lean module compiled","shard":"modules/62af8cc17d32d317.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","label":"Algorithms.ETCRealArgmaxTie","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealArgmaxTie","description":"This module identifies the strict-update fold used by the native Real ETC route with the least-encoded maximizing arm selected by the Nat.find scheme used in LML's measurable argmax. It then assembles round-robin exploration, the commit action, and post-commit persistence into the action equality consumed by the exact native Real source-law theorem.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html","parent":"chapter:etc","order":84,"meta":[["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcrealargmaxtie banditrlproof.algorithms.etcrealargmaxtie this module identifies the strict-update fold used by the native real etc route with the least-encoded maximizing arm selected by the nat.find scheme used in lml's measurable argmax. it then assembles round-robin exploration, the commit action, and post-commit persistence into the action equality consumed by the exact native real source-law theorem. lean module compiled","shard":"modules/fb441a84537c03e3.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","label":"Algorithms.ETCRealEmpiricalMean","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealEmpiricalMean","description":"This module supplies the Real-valued empirical-mean and measurable finite argmax surface needed before transporting ETC to an arbitrary native Real environment. It stays below reward-law identification, concentration, and the final IsAlgEnvSeq regret theorem.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html","parent":"chapter:etc","order":85,"meta":[["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcrealempiricalmean banditrlproof.algorithms.etcrealempiricalmean this module supplies the real-valued empirical-mean and measurable finite argmax surface needed before transporting etc to an arbitrary native real environment. it stays below reward-law identification, concentration, and the final isalgenvseq regret theorem. lean module compiled","shard":"modules/9ae2ff24a48ff0e8.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","label":"Algorithms.ETCRealHistoryScore","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealHistoryScore","description":"This module mirrors the finite-pair-history pullCount', sumRewards', and empMean' score surface used by the pinned LML ETC source. It identifies that history score with the existing native Real exploration score, then feeds the history-shaped commit law into the exact source adapter.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html","parent":"chapter:etc","order":86,"meta":[["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcrealhistoryscore banditrlproof.algorithms.etcrealhistoryscore this module mirrors the finite-pair-history pullcount', sumrewards', and empmean' score surface used by the pinned lml etc source. it identifies that history score with the existing native real exploration score, then feeds the history-shaped commit law into the exact source adapter. lean module compiled","shard":"modules/e5b1389df6ad46a2.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","label":"Algorithms.ETCRealInfinitePiTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealInfinitePiTail","description":"This module proves the exact single-arm ETC wrong-commit tail for a native Real reward kernel under the canonical independent-coordinate product law. It then consumes that tail in the expected pull-count and finite-sum kernel regret identities. Transport from an arbitrary external algorithm/environment sequence remains a downstream law-identification obligation.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html","parent":"chapter:etc","order":87,"meta":[["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean"],["Declarations","17"],["Project imports","4"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcrealinfinitepitail banditrlproof.algorithms.etcrealinfinitepitail this module proves the exact single-arm etc wrong-commit tail for a native real reward kernel under the canonical independent-coordinate product law. it then consumes that tail in the expected pull-count and finite-sum kernel regret identities. transport from an arbitrary external algorithm/environment sequence remains a downstream law-identification obligation. lean module compiled","shard":"modules/8da6ef3ba368d237.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","label":"Algorithms.ETCRealLMLCompat","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealLMLCompat","description":"The pinned LML source currently uses a newer Lean/mathlib toolchain, so ABRL cannot import its IsAlgEnvSeq declaration directly. This module packages the exact measurable-action, measurable-feedback, action-behavior, and stationary feedback-law consequences consumed by the local ETC theorem. It is a local compatibility structure, not an imported LML proof.","url":"../modules/banditrlproof-algorithms-etcreallmlcompat/index.html","parent":"chapter:etc","order":88,"meta":[["Source","BanditRLProof/Algorithms/ETCRealLMLCompat.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcreallmlcompat banditrlproof.algorithms.etcreallmlcompat the pinned lml source currently uses a newer lean/mathlib toolchain, so abrl cannot import its isalgenvseq declaration directly. this module packages the exact measurable-action, measurable-feedback, action-behavior, and stationary feedback-law consequences consumed by the local etc theorem. it is a local compatibility structure, not an imported lml proof. lean module compiled","shard":"modules/6d2389e9aefd14aa.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","label":"Algorithms.ETCRealPrefixLawTransport","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealPrefixLawTransport","description":"This module factors the native Real ETC commit, action, and finite-horizon kernel regret through the finite exploration reward prefix. It then transports the canonical infinite-product regret bound to an arbitrary probability space from equality of that finite-prefix pushforward law. A final adapter permits an external action process that agrees almost surely with the local ETC action.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html","parent":"chapter:etc","order":89,"meta":[["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean"],["Declarations","23"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcrealprefixlawtransport banditrlproof.algorithms.etcrealprefixlawtransport this module factors the native real etc commit, action, and finite-horizon kernel regret through the finite exploration reward prefix. it then transports the canonical infinite-product regret bound to an arbitrary probability space from equality of that finite-prefix pushforward law. a final adapter permits an external action process that agrees almost surely with the local etc action. lean module compiled","shard":"modules/6ea3d46cfbf19c2d.json"},{"id":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","label":"Algorithms.ETCRealSourceAdapter","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRealSourceAdapter","description":"This module converts the action-selected feedback-law fields exposed by an algorithm/environment sequence into the reward-only scheduled laws consumed by the native Real exact ETC regret theorem. It does not import LML: the theorem statement mirrors the relevant IsAlgEnvSeq fields so the remaining upstream wrapper is limited to source-name and action/tie alignment.","url":"../modules/banditrlproof-algorithms-etcrealsourceadapter/index.html","parent":"chapter:etc","order":90,"meta":[["Source","BanditRLProof/Algorithms/ETCRealSourceAdapter.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcrealsourceadapter banditrlproof.algorithms.etcrealsourceadapter this module converts the action-selected feedback-law fields exposed by an algorithm/environment sequence into the reward-only scheduled laws consumed by the native real exact etc regret theorem. it does not import lml: the theorem statement mirrors the relevant isalgenvseq fields so the remaining upstream wrapper is limited to source-name and action/tie alignment. lean module compiled","shard":"modules/3f97a6d8cd07d916.json"},{"id":"module:BanditRLProof.Algorithms.ETCRegretLemmas","label":"Algorithms.ETCRegretLemmas","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCRegretLemmas","description":"This module contains ETC-specific deterministic regret scaffolds. It consumes the round-robin exploration count layer and deliberately stays below commit behavior, empirical means, probability, concentration, and final ETC regret theorems.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html","parent":"chapter:etc","order":91,"meta":[["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean"],["Declarations","8"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcregretlemmas banditrlproof.algorithms.etcregretlemmas this module contains etc-specific deterministic regret scaffolds. it consumes the round-robin exploration count layer and deliberately stays below commit behavior, empirical means, probability, concentration, and final etc regret theorems. lean module compiled","shard":"modules/c0416491b1744f12.json"},{"id":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","label":"Algorithms.ETCSumRewardsDiff","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCSumRewardsDiff","description":"This module bridges the fixed-horizon sumRewards comparison produced by the ETC empirical-mean algebra layer to a centered pairwise finite-sum event. It stays deterministic: no probability measure, independence, sub-Gaussianity, filtration, conditional expectation, or final ETC result is introduced here.","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html","parent":"chapter:etc","order":92,"meta":[["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcsumrewardsdiff banditrlproof.algorithms.etcsumrewardsdiff this module bridges the fixed-horizon sumrewards comparison produced by the etc empirical-mean algebra layer to a centered pairwise finite-sum event. it stays deterministic: no probability measure, independence, sub-gaussianity, filtration, conditional expectation, or final etc result is introduced here. lean module compiled","shard":"modules/869b43c0a8211722.json"},{"id":"module:BanditRLProof.Algorithms.ETCTrace","label":"Algorithms.ETCTrace","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCTrace","description":"This module introduces the first deterministic boundary for phase-switching Explore-Then-Commit traces. The commit arm is supplied explicitly; empirical mean selection, probability, concentration, and regret facts live in later leaves.","url":"../modules/banditrlproof-algorithms-etctrace/index.html","parent":"chapter:etc","order":93,"meta":[["Source","BanditRLProof/Algorithms/ETCTrace.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etctrace banditrlproof.algorithms.etctrace this module introduces the first deterministic boundary for phase-switching explore-then-commit traces. the commit arm is supplied explicitly; empirical mean selection, probability, concentration, and regret facts live in later leaves. lean module compiled","shard":"modules/383dc9cfc419125b.json"},{"id":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","label":"Algorithms.ETCTraceCountLemmas","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCTraceCountLemmas","description":"This module contains deterministic pull-count facts for the fixed-commit ETC trace. It stays below regret, empirical commit selection, probability, and concentration.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html","parent":"chapter:etc","order":94,"meta":[["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean"],["Declarations","9"],["Project imports","3"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etctracecountlemmas banditrlproof.algorithms.etctracecountlemmas this module contains deterministic pull-count facts for the fixed-commit etc trace. it stays below regret, empirical commit selection, probability, and concentration. lean module compiled","shard":"modules/fdc005960b7c90ed.json"},{"id":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","label":"Algorithms.ETCWrongCommitCanonicalTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","description":"This module closes the local composition from the canonical centered reward difference independent sub-Gaussian route to the concrete argmax-oracle wrong-commit probability bound. It still leaves the actual reward-law independence and sub-Gaussian witnesses as explicit hypotheses.","url":"../modules/banditrlproof-algorithms-etcwrongcommitcanonicaltail/index.html","parent":"chapter:etc","order":95,"meta":[["Source","BanditRLProof/Algorithms/ETCWrongCommitCanonicalTail.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcwrongcommitcanonicaltail banditrlproof.algorithms.etcwrongcommitcanonicaltail this module closes the local composition from the canonical centered reward difference independent sub-gaussian route to the concrete argmax-oracle wrong-commit probability bound. it still leaves the actual reward-law independence and sub-gaussian witnesses as explicit hypotheses. lean module compiled","shard":"modules/693f5ab3d775f0ba.json"},{"id":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","label":"Algorithms.ETCWrongCommitRegretAssembly","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","description":"This module gives a pointwise bridge from the deterministic fixed-commit ETC regret scaffolds to a wrong-commit-shaped suffix penalty. It deliberately stays below integration, probability bounds, concentration, filtrations, and a final ETC expected-regret theorem.","url":"../modules/banditrlproof-algorithms-etcwrongcommitregretassembly/index.html","parent":"chapter:etc","order":96,"meta":[["Source","BanditRLProof/Algorithms/ETCWrongCommitRegretAssembly.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","ETC"]],"statement":"","missing":[],"search":"algorithms.etcwrongcommitregretassembly banditrlproof.algorithms.etcwrongcommitregretassembly this module gives a pointwise bridge from the deterministic fixed-commit etc regret scaffolds to a wrong-commit-shaped suffix penalty. it deliberately stays below integration, probability bounds, concentration, filtrations, and a final etc expected-regret theorem. lean module compiled","shard":"modules/575319f5853bded4.json"},{"id":"module:BanditRLProof.Algorithms.HOOActualRegret","label":"Algorithms.HOOActualRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOActualRegret","description":"Expected actual/cumulative regret equals expected pseudo-regret for the constructed HOO trajectory, with bounded integrability proved.","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html","parent":"chapter:foundations","order":97,"meta":[["Source","BanditRLProof/Algorithms/HOOActualRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooactualregret banditrlproof.algorithms.hooactualregret expected actual/cumulative regret equals expected pseudo-regret for the constructed hoo trajectory, with bounded integrability proved. lean module compiled","shard":"modules/ff36e5a52ef33eea.json"},{"id":"module:BanditRLProof.Algorithms.HOOConcentration","label":"Algorithms.HOOConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOConcentration","description":"Count-dependent concentration for the actual chronological HOO trajectory, including its first reward.","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html","parent":"chapter:foundations","order":98,"meta":[["Source","BanditRLProof/Algorithms/HOOConcentration.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooconcentration banditrlproof.algorithms.hooconcentration count-dependent concentration for the actual chronological hoo trajectory, including its first reward. lean module compiled","shard":"modules/2f1e7407e86f21fa.json"},{"id":"module:BanditRLProof.Algorithms.HOOConditionalMGF","label":"Algorithms.HOOConditionalMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOConditionalMGF","description":"Bounded-reward conditional MGF producers for the actual HOO step kernel. Region selection depends on the observed prefix, never on the next reward.","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html","parent":"chapter:foundations","order":99,"meta":[["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean"],["Declarations","16"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooconditionalmgf banditrlproof.algorithms.hooconditionalmgf bounded-reward conditional mgf producers for the actual hoo step kernel. region selection depends on the observed prefix, never on the next reward. lean module compiled","shard":"modules/fb89660da334fb1b.json"},{"id":"module:BanditRLProof.Algorithms.HOOConfidence","label":"Algorithms.HOOConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOConfidence","description":"Finite visit-count peeling for HOO. The time parameter is the number of already observed chronological rewards; the bound applies to the next decision.","url":"../modules/banditrlproof-algorithms-hooconfidence/index.html","parent":"chapter:foundations","order":100,"meta":[["Source","BanditRLProof/Algorithms/HOOConfidence.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooconfidence banditrlproof.algorithms.hooconfidence finite visit-count peeling for hoo. the time parameter is the number of already observed chronological rewards; the bound applies to the next decision. lean module compiled","shard":"modules/03ba75ef32eb5fb3.json"},{"id":"module:BanditRLProof.Algorithms.HOODepthOptimization","label":"Algorithms.HOODepthOptimization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOODepthOptimization","description":"Integer depth optimization with positive logarithms, including horizon one.","url":"../modules/banditrlproof-algorithms-hoodepthoptimization/index.html","parent":"chapter:foundations","order":101,"meta":[["Source","BanditRLProof/Algorithms/HOODepthOptimization.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hoodepthoptimization banditrlproof.algorithms.hoodepthoptimization integer depth optimization with positive logarithms, including horizon one. lean module compiled","shard":"modules/e10b13a303147a7f.json"},{"id":"module:BanditRLProof.Algorithms.HOOExpectedRegret","label":"Algorithms.HOOExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOExpectedRegret","description":"Integrability and the unoptimized expected HOO regret bound, on the actual trajectory. Mean boundedness is a native source-model hypothesis.","url":"../modules/banditrlproof-algorithms-hooexpectedregret/index.html","parent":"chapter:foundations","order":102,"meta":[["Source","BanditRLProof/Algorithms/HOOExpectedRegret.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooexpectedregret banditrlproof.algorithms.hooexpectedregret integrability and the unoptimized expected hoo regret bound, on the actual trajectory. mean boundedness is a native source-model hypothesis. lean module compiled","shard":"modules/9b924db43bd01fc1.json"},{"id":"module:BanditRLProof.Algorithms.HOOExpectedVisits","label":"Algorithms.HOOExpectedVisits","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOExpectedVisits","description":"Expected HOO regional visits through the shared threshold-count argument. The binary trace below records visits to one fixed region; it does not replace the infinite HOO action space or its reward trajectory by a finite-arm model.","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html","parent":"chapter:foundations","order":103,"meta":[["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean"],["Declarations","10"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooexpectedvisits banditrlproof.algorithms.hooexpectedvisits expected hoo regional visits through the shared threshold-count argument. the binary trace below records visits to one fixed region; it does not replace the infinite hoo action space or its reward trajectory by a finite-arm model. lean module compiled","shard":"modules/4d0f92fff34a01ff.json"},{"id":"module:BanditRLProof.Algorithms.HOOHistory","label":"Algorithms.HOOHistory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOHistory","description":"Causal HOO history recursion. Input rewards are chronological observations; the sequential reward law is a separate mandatory probability construction.","url":"../modules/banditrlproof-algorithms-hoohistory/index.html","parent":"chapter:foundations","order":104,"meta":[["Source","BanditRLProof/Algorithms/HOOHistory.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hoohistory banditrlproof.algorithms.hoohistory causal hoo history recursion. input rewards are chronological observations; the sequential reward law is a separate mandatory probability construction. lean module compiled","shard":"modules/31aaa32a23098242.json"},{"id":"module:BanditRLProof.Algorithms.HOOIndexConfidence","label":"Algorithms.HOOIndexConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOIndexConfidence","description":"Transfer from actual centered regional observations to HOO U indices.","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html","parent":"chapter:foundations","order":105,"meta":[["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooindexconfidence banditrlproof.algorithms.hooindexconfidence transfer from actual centered regional observations to hoo u indices. lean module compiled","shard":"modules/ddb172e710ed8869.json"},{"id":"module:BanditRLProof.Algorithms.HOOMeasurable","label":"Algorithms.HOOMeasurable","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOMeasurable","description":"Measurability of the actual finite tree computations. Node labels are countable and discrete; the confidence comparisons remain real-valued.","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html","parent":"chapter:foundations","order":106,"meta":[["Source","BanditRLProof/Algorithms/HOOMeasurable.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hoomeasurable banditrlproof.algorithms.hoomeasurable measurability of the actual finite tree computations. node labels are countable and discrete; the confidence comparisons remain real-valued. lean module compiled","shard":"modules/2a8467a1ee82da16.json"},{"id":"module:BanditRLProof.Algorithms.HOOPathComparison","label":"Algorithms.HOOPathComparison","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOPathComparison","description":"Actual HOO search-path comparison with a supremum-preserving branch.","url":"../modules/banditrlproof-algorithms-hoopathcomparison/index.html","parent":"chapter:foundations","order":107,"meta":[["Source","BanditRLProof/Algorithms/HOOPathComparison.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hoopathcomparison banditrlproof.algorithms.hoopathcomparison actual hoo search-path comparison with a supremum-preserving branch. lean module compiled","shard":"modules/b7f75ec18030ba8c.json"},{"id":"module:BanditRLProof.Algorithms.HOOPrefix","label":"Algorithms.HOOPrefix","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOPrefix","description":"Actual selected-path and prefix-closure producers for source Lemma 14.","url":"../modules/banditrlproof-algorithms-hooprefix/index.html","parent":"chapter:foundations","order":108,"meta":[["Source","BanditRLProof/Algorithms/HOOPrefix.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooprefix banditrlproof.algorithms.hooprefix actual selected-path and prefix-closure producers for source lemma 14. lean module compiled","shard":"modules/9c2a059eeb853017.json"},{"id":"module:BanditRLProof.Algorithms.HOORate","label":"Algorithms.HOORate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOORate","description":"Repaired source Theorem 6: actual HOO expected pseudo-regret at all positive horizons, with a horizon-independent constant.","url":"../modules/banditrlproof-algorithms-hoorate/index.html","parent":"chapter:foundations","order":109,"meta":[["Source","BanditRLProof/Algorithms/HOORate.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hoorate banditrlproof.algorithms.hoorate repaired source theorem 6: actual hoo expected pseudo-regret at all positive horizons, with a horizon-independent constant. lean module compiled","shard":"modules/2218ec96e6a2ed8f.json"},{"id":"module:BanditRLProof.Algorithms.HOORegretAlgebra","label":"Algorithms.HOORegretAlgebra","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOORegretAlgebra","description":"Explicit geometric reduction of the source's three regret sums.","url":"../modules/banditrlproof-algorithms-hooregretalgebra/index.html","parent":"chapter:foundations","order":110,"meta":[["Source","BanditRLProof/Algorithms/HOORegretAlgebra.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooregretalgebra banditrlproof.algorithms.hooregretalgebra explicit geometric reduction of the source's three regret sums. lean module compiled","shard":"modules/f456b6b4e4f64e38.json"},{"id":"module:BanditRLProof.Algorithms.HOORegretPartition","label":"Algorithms.HOORegretPartition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOORegretPartition","description":"Pathwise regret contributions for the actual fresh-node HOO trace.","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html","parent":"chapter:foundations","order":111,"meta":[["Source","BanditRLProof/Algorithms/HOORegretPartition.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooregretpartition banditrlproof.algorithms.hooregretpartition pathwise regret contributions for the actual fresh-node hoo trace. lean module compiled","shard":"modules/f0009a9444670c3b.json"},{"id":"module:BanditRLProof.Algorithms.HOORewardFamily","label":"Algorithms.HOORewardFamily","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOORewardFamily","description":"Arbitrary reward-family interface for the source-repaired HOO rate. Only the countable representative-node law must be measurable. The internal discrete-domain transport preserves the actual action and trajectory.","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html","parent":"chapter:foundations","order":112,"meta":[["Source","BanditRLProof/Algorithms/HOORewardFamily.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hoorewardfamily banditrlproof.algorithms.hoorewardfamily arbitrary reward-family interface for the source-repaired hoo rate. only the countable representative-node law must be measurable. the internal discrete-domain transport preserves the actual action and trajectory. lean module compiled","shard":"modules/b29d5fe13bc16a56.json"},{"id":"module:BanditRLProof.Algorithms.HOOSelectionTail","label":"Algorithms.HOOSelectionTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOSelectionTail","description":"Selection of a sufficiently visited poor region forces a confidence failure for that region or for a finite prefix of a supremum-optimal branch.","url":"../modules/banditrlproof-algorithms-hooselectiontail/index.html","parent":"chapter:foundations","order":113,"meta":[["Source","BanditRLProof/Algorithms/HOOSelectionTail.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hooselectiontail banditrlproof.algorithms.hooselectiontail selection of a sufficiently visited poor region forces a confidence failure for that region or for a finite prefix of a supremum-optimal branch. lean module compiled","shard":"modules/315d6f0fd319eb3f.json"},{"id":"module:BanditRLProof.Algorithms.HOOTrajectory","label":"Algorithms.HOOTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOTrajectory","description":"The actual HOO reward law on one infinite chronological trajectory. Every successor kernel selects its region using only the observed prefix.","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html","parent":"chapter:foundations","order":114,"meta":[["Source","BanditRLProof/Algorithms/HOOTrajectory.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hootrajectory banditrlproof.algorithms.hootrajectory the actual hoo reward law on one infinite chronological trajectory. every successor kernel selects its region using only the observed prefix. lean module compiled","shard":"modules/a41329574d895035.json"},{"id":"module:BanditRLProof.Algorithms.HOOTree","label":"Algorithms.HOOTree","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HOOTree","description":"Finite implementation of the infinite HOO binary-tree search. Fuel is an implementation device; the depth proofs show it cannot truncate the search.","url":"../modules/banditrlproof-algorithms-hootree/index.html","parent":"chapter:foundations","order":115,"meta":[["Source","BanditRLProof/Algorithms/HOOTree.lean"],["Declarations","23"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.hootree banditrlproof.algorithms.hootree finite implementation of the infinite hoo binary-tree search. fuel is an implementation device; the depth proofs show it cannot truncate the search. lean module compiled","shard":"modules/07448909ff0903c4.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","label":"Algorithms.HeavyTailAdaptive","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailAdaptive","description":"Actual causal observations and their adaptive-count confidence.","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html","parent":"chapter:foundations","order":116,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailadaptive banditrlproof.algorithms.heavytailadaptive actual causal observations and their adaptive-count confidence. lean module compiled","shard":"modules/1900adfcc28d5961.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","label":"Algorithms.HeavyTailExpectedCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailExpectedCount","description":"Expected pull counts for the actual robust policy under raw-moment reward laws.","url":"../modules/banditrlproof-algorithms-heavytailexpectedcount/index.html","parent":"chapter:foundations","order":117,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailExpectedCount.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailexpectedcount banditrlproof.algorithms.heavytailexpectedcount expected pull counts for the actual robust policy under raw-moment reward laws. lean module compiled","shard":"modules/952a4e1f3873a4d8.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailHistory","label":"Algorithms.HeavyTailHistory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailHistory","description":"Sample-index truncation computed solely from the finite observed history.","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html","parent":"chapter:foundations","order":118,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailhistory banditrlproof.algorithms.heavytailhistory sample-index truncation computed solely from the finite observed history. lean module compiled","shard":"modules/66bef50c33586d09.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailRegret","label":"Algorithms.HeavyTailRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailRegret","description":"Conservative robust-UCB expected pseudo-regret endpoint. This is the separately documented repair/adaptation, not the unchanged printed BCL constant theorem. The policy and measure are fixed across horizons. Zero-gap arms contribute zero.","url":"../modules/banditrlproof-algorithms-heavytailregret/index.html","parent":"chapter:foundations","order":119,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailRegret.lean"],["Declarations","3"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailregret banditrlproof.algorithms.heavytailregret conservative robust-ucb expected pseudo-regret endpoint. this is the separately documented repair/adaptation, not the unchanged printed bcl constant theorem. the policy and measure are fixed across horizons. zero-gap arms contribute zero. lean module compiled","shard":"modules/539c7b0648284a1d.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","label":"Algorithms.HeavyTailRegretCap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailRegretCap","description":"Finite raw-moment regret cap and extended-real supremum obstruction. Source: Genalti et al., COLT2024, Theorem2 Eq5. This enlarged trace-law class supplies an upper obstruction, not a fixed-algorithm lower bound or an asymptotic impossibility theorem. See the paired reader page and source deltas.","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html","parent":"chapter:foundations","order":120,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailregretcap banditrlproof.algorithms.heavytailregretcap finite raw-moment regret cap and extended-real supremum obstruction. source: genalti et al., colt2024, theorem2 eq5. this enlarged trace-law class supplies an upper obstruction, not a fixed-algorithm lower bound or an asymptotic impossibility theorem. see the paired reader page and source deltas. lean module compiled","shard":"modules/9aa14106ccd22260.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","label":"Algorithms.HeavyTailSourceAdaptive","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailSourceAdaptive","description":"Strict good-event index comparison and actual selected-large-count events.","url":"../modules/banditrlproof-algorithms-heavytailsourceadaptive/index.html","parent":"chapter:foundations","order":121,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailSourceAdaptive.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailsourceadaptive banditrlproof.algorithms.heavytailsourceadaptive strict good-event index comparison and actual selected-large-count events. lean module compiled","shard":"modules/233681694f444278.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","label":"Algorithms.HeavyTailSourceCounterexample","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailSourceCounterexample","description":"Finite counterexample to the literal BCL13 printed regret coefficient for the unchanged radius-four policy. Deterministic arms0,-1, raw second moment<=1, horizon2^50. The corrected upper bound remains a separate theorem.","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html","parent":"chapter:foundations","order":122,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean"],["Declarations","34"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailsourcecounterexample banditrlproof.algorithms.heavytailsourcecounterexample finite counterexample to the literal bcl13 printed regret coefficient for the unchanged radius-four policy. deterministic arms0,-1, raw second moment<=1, horizon2^50. the corrected upper bound remains a separate theorem. lean module compiled","shard":"modules/9b3ed05d349ab7f7.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","label":"Algorithms.HeavyTailSourceExpectedCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","description":"Corrected expected counts for the unchanged source policy. The shared threshold-count integral producer is reused; confidence is derived from raw moments.","url":"../modules/banditrlproof-algorithms-heavytailsourceexpectedcount/index.html","parent":"chapter:foundations","order":123,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailSourceExpectedCount.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailsourceexpectedcount banditrlproof.algorithms.heavytailsourceexpectedcount corrected expected counts for the unchanged source policy. the shared threshold-count integral producer is reused; confidence is derived from raw moments. lean module compiled","shard":"modules/bdee61a9893b6589.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","label":"Algorithms.HeavyTailSourcePolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailSourcePolicy","description":"Source-parameter robust UCB: round-robin initialization is a tie convention for unpulled arms; later choices maximize the source radius-four index using only the observed history. Paper round is zero-based decision time plus one.","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html","parent":"chapter:foundations","order":124,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean"],["Declarations","21"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailsourcepolicy banditrlproof.algorithms.heavytailsourcepolicy source-parameter robust ucb: round-robin initialization is a tie convention for unpulled arms; later choices maximize the source radius-four index using only the observed history. paper round is zero-based decision time plus one. lean module compiled","shard":"modules/ae4833da8f0ce0f7.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","label":"Algorithms.HeavyTailSourceRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailSourceRegret","description":"Complete corrected expected-regret bound for the unchanged source policy. The rejected printed coefficient remains a separate finite-counterexample obligation.","url":"../modules/banditrlproof-algorithms-heavytailsourceregret/index.html","parent":"chapter:foundations","order":125,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailSourceRegret.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailsourceregret banditrlproof.algorithms.heavytailsourceregret complete corrected expected-regret bound for the unchanged source policy. the rejected printed coefficient remains a separate finite-counterexample obligation. lean module compiled","shard":"modules/c0b77d752e163bbc.json"},{"id":"module:BanditRLProof.Algorithms.HeavyTailUCB","label":"Algorithms.HeavyTailUCB","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.HeavyTailUCB","description":"Causal robust UCB with the audited conservative confidence schedule. Time is zero-based; the logarithm uses max(t,2), and failure probability t^-4 is the explicit repair recorded in DERIVATION.md. No regret endpoint is claimed here.","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html","parent":"chapter:foundations","order":126,"meta":[["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.heavytailucb banditrlproof.algorithms.heavytailucb causal robust ucb with the audited conservative confidence schedule. time is zero-based; the logarithm uses max(t,2), and failure probability t^-4 is the explicit repair recorded in derivation.md. no regret endpoint is claimed here. lean module compiled","shard":"modules/5e16c61dc9dcdbd7.json"},{"id":"module:BanditRLProof.Algorithms.KLUCBBernoulli","label":"Algorithms.KLUCBBernoulli","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.KLUCBBernoulli","description":"This file owns the project-local binary relative entropy used by KL-UCB. The codomain is ENNReal: singular Bernoulli comparisons are genuinely top, not the accidental finite value obtained from Mathlib's totalized Real.log 0.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html","parent":"chapter:foundations","order":127,"meta":[["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean"],["Declarations","41"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.klucbbernoulli banditrlproof.algorithms.klucbbernoulli this file owns the project-local binary relative entropy used by kl-ucb. the codomain is ennreal: singular bernoulli comparisons are genuinely top, not the accidental finite value obtained from mathlib's totalized real.log 0. lean module compiled","shard":"modules/b1dad9e08fe974cd.json"},{"id":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","label":"Algorithms.KLUCBGeneratedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.KLUCBGeneratedRegret","description":"This module defines one horizon-free KL-UCB policy on the canonical generated Rat action/reward trajectory. The confidence budget is calibrated from the accepted telescoping empirical-mean radius and an explicit common interior margin. The selected score is nevertheless the supremum of the genuine Bernoulli-KL confidence set; it is not the ordinary additive UCB score.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html","parent":"chapter:foundations","order":128,"meta":[["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean"],["Declarations","47"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.klucbgeneratedregret banditrlproof.algorithms.klucbgeneratedregret this module defines one horizon-free kl-ucb policy on the canonical generated rat action/reward trajectory. the confidence budget is calibrated from the accepted telescoping empirical-mean radius and an explicit common interior margin. the selected score is nevertheless the supremum of the genuine bernoulli-kl confidence set; it is not the ordinary additive ucb score. lean module compiled","shard":"modules/978c2a806ed308d8.json"},{"id":"module:BanditRLProof.Algorithms.MOSS","label":"Algorithms.MOSS","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSS","description":"Algorithm 7 in Lattimore--Szepesvari, *Bandit Algorithms*, uses a horizon dependent, realized-pull-count index. These definitions keep its factor four and log-plus truncation. They do not yet construct a stochastic history law or prove Theorem 9.1's regret bound.","url":"../modules/banditrlproof-algorithms-moss/index.html","parent":"chapter:foundations","order":129,"meta":[["Source","BanditRLProof/Algorithms/MOSS.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.moss banditrlproof.algorithms.moss algorithm 7 in lattimore--szepesvari, *bandit algorithms*, uses a horizon dependent, realized-pull-count index. these definitions keep its factor four and log-plus truncation. they do not yet construct a stochastic history law or prove theorem 9.1's regret bound. lean module compiled","shard":"modules/9770649f4b4075ea.json"},{"id":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","label":"Algorithms.MOSSCanonicalHistory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSCanonicalHistory","description":"Generated source map for BanditRLProof/Algorithms/MOSSCanonicalHistory.lean.","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html","parent":"chapter:foundations","order":130,"meta":[["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mosscanonicalhistory banditrlproof.algorithms.mosscanonicalhistory generated source map for banditrlproof/algorithms/mosscanonicalhistory.lean. lean module compiled","shard":"modules/1893505bcdd62921.json"},{"id":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","label":"Algorithms.MOSSCanonicalReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSCanonicalReward","description":"Generated source map for BanditRLProof/Algorithms/MOSSCanonicalReward.lean.","url":"../modules/banditrlproof-algorithms-mosscanonicalreward/index.html","parent":"chapter:foundations","order":131,"meta":[["Source","BanditRLProof/Algorithms/MOSSCanonicalReward.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mosscanonicalreward banditrlproof.algorithms.mosscanonicalreward generated source map for banditrlproof/algorithms/mosscanonicalreward.lean. lean module compiled","shard":"modules/6321412dacffae06.json"},{"id":"module:BanditRLProof.Algorithms.MOSSConditionalReward","label":"Algorithms.MOSSConditionalReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSConditionalReward","description":"Generated source map for BanditRLProof/Algorithms/MOSSConditionalReward.lean.","url":"../modules/banditrlproof-algorithms-mossconditionalreward/index.html","parent":"chapter:foundations","order":132,"meta":[["Source","BanditRLProof/Algorithms/MOSSConditionalReward.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossconditionalreward banditrlproof.algorithms.mossconditionalreward generated source map for banditrlproof/algorithms/mossconditionalreward.lean. lean module compiled","shard":"modules/56fa1bb6573cdf39.json"},{"id":"module:BanditRLProof.Algorithms.MOSSConstants","label":"Algorithms.MOSSConstants","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSConstants","description":"Generated source map for BanditRLProof/Algorithms/MOSSConstants.lean.","url":"../modules/banditrlproof-algorithms-mossconstants/index.html","parent":"chapter:foundations","order":133,"meta":[["Source","BanditRLProof/Algorithms/MOSSConstants.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossconstants banditrlproof.algorithms.mossconstants generated source map for banditrlproof/algorithms/mossconstants.lean. lean module compiled","shard":"modules/7201346a5b40b9e9.json"},{"id":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","label":"Algorithms.MOSSExpectedOccupancy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSExpectedOccupancy","description":"Generated source map for BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean.","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html","parent":"chapter:foundations","order":134,"meta":[["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean"],["Declarations","7"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossexpectedoccupancy banditrlproof.algorithms.mossexpectedoccupancy generated source map for banditrlproof/algorithms/mossexpectedoccupancy.lean. lean module compiled","shard":"modules/ad66c789e303c560.json"},{"id":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","label":"Algorithms.MOSSExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSExpectedRegret","description":"Generated source map for BanditRLProof/Algorithms/MOSSExpectedRegret.lean.","url":"../modules/banditrlproof-algorithms-mossexpectedregret/index.html","parent":"chapter:foundations","order":135,"meta":[["Source","BanditRLProof/Algorithms/MOSSExpectedRegret.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossexpectedregret banditrlproof.algorithms.mossexpectedregret generated source map for banditrlproof/algorithms/mossexpectedregret.lean. lean module compiled","shard":"modules/a2b8202084a39122.json"},{"id":"module:BanditRLProof.Algorithms.MOSSHistory","label":"Algorithms.MOSSHistory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSHistory","description":"An inclusive history at t contains t+1 observations; the successor selector therefore calls the source action at t+1. No reward law or concentration certificate is required to construct this deterministic Markov policy.","url":"../modules/banditrlproof-algorithms-mosshistory/index.html","parent":"chapter:foundations","order":136,"meta":[["Source","BanditRLProof/Algorithms/MOSSHistory.lean"],["Declarations","6"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mosshistory banditrlproof.algorithms.mosshistory an inclusive history at t contains t+1 observations; the successor selector therefore calls the source action at t+1. no reward law or concentration certificate is required to construct this deterministic markov policy. lean module compiled","shard":"modules/69098d034e7a2950.json"},{"id":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","label":"Algorithms.MOSSHistoryLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSHistoryLaw","description":"Generated source map for BanditRLProof/Algorithms/MOSSHistoryLaw.lean.","url":"../modules/banditrlproof-algorithms-mosshistorylaw/index.html","parent":"chapter:foundations","order":137,"meta":[["Source","BanditRLProof/Algorithms/MOSSHistoryLaw.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mosshistorylaw banditrlproof.algorithms.mosshistorylaw generated source map for banditrlproof/algorithms/mosshistorylaw.lean. lean module compiled","shard":"modules/ccacf09fe51dcb36.json"},{"id":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","label":"Algorithms.MOSSHistoryRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSHistoryRegret","description":"Generated source map for BanditRLProof/Algorithms/MOSSHistoryRegret.lean.","url":"../modules/banditrlproof-algorithms-mosshistoryregret/index.html","parent":"chapter:foundations","order":138,"meta":[["Source","BanditRLProof/Algorithms/MOSSHistoryRegret.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mosshistoryregret banditrlproof.algorithms.mosshistoryregret generated source map for banditrlproof/algorithms/mosshistoryregret.lean. lean module compiled","shard":"modules/c8da34ff2e770ab4.json"},{"id":"module:BanditRLProof.Algorithms.MOSSOccupancy","label":"Algorithms.MOSSOccupancy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSOccupancy","description":"Generated source map for BanditRLProof/Algorithms/MOSSOccupancy.lean.","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html","parent":"chapter:foundations","order":139,"meta":[["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossoccupancy banditrlproof.algorithms.mossoccupancy generated source map for banditrlproof/algorithms/mossoccupancy.lean. lean module compiled","shard":"modules/01b1e451bc8da9d8.json"},{"id":"module:BanditRLProof.Algorithms.MOSSOptimism","label":"Algorithms.MOSSOptimism","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSOptimism","description":"Generated source map for BanditRLProof/Algorithms/MOSSOptimism.lean.","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html","parent":"chapter:foundations","order":140,"meta":[["Source","BanditRLProof/Algorithms/MOSSOptimism.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossoptimism banditrlproof.algorithms.mossoptimism generated source map for banditrlproof/algorithms/mossoptimism.lean. lean module compiled","shard":"modules/dd62fe8c9d9c274e.json"},{"id":"module:BanditRLProof.Algorithms.MOSSPeeling","label":"Algorithms.MOSSPeeling","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSPeeling","description":"The source barrier and explicit dyadic maximum-event bridge for Lemma 9.3.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html","parent":"chapter:foundations","order":141,"meta":[["Source","BanditRLProof/Algorithms/MOSSPeeling.lean"],["Declarations","15"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mosspeeling banditrlproof.algorithms.mosspeeling the source barrier and explicit dyadic maximum-event bridge for lemma 9.3. lean module compiled","shard":"modules/0b6768b766971b05.json"},{"id":"module:BanditRLProof.Algorithms.MOSSRegret","label":"Algorithms.MOSSRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSRegret","description":"Generated source map for BanditRLProof/Algorithms/MOSSRegret.lean.","url":"../modules/banditrlproof-algorithms-mossregret/index.html","parent":"chapter:foundations","order":142,"meta":[["Source","BanditRLProof/Algorithms/MOSSRegret.lean"],["Declarations","3"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossregret banditrlproof.algorithms.mossregret generated source map for banditrlproof/algorithms/mossregret.lean. lean module compiled","shard":"modules/68217195bbceb365.json"},{"id":"module:BanditRLProof.Algorithms.MOSSRewardBranch","label":"Algorithms.MOSSRewardBranch","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSRewardBranch","description":"Generated source map for BanditRLProof/Algorithms/MOSSRewardBranch.lean.","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html","parent":"chapter:foundations","order":143,"meta":[["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossrewardbranch banditrlproof.algorithms.mossrewardbranch generated source map for banditrlproof/algorithms/mossrewardbranch.lean. lean module compiled","shard":"modules/14a9b50ee60dc408.json"},{"id":"module:BanditRLProof.Algorithms.MOSSStream","label":"Algorithms.MOSSStream","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSStream","description":"Generated source map for BanditRLProof/Algorithms/MOSSStream.lean.","url":"../modules/banditrlproof-algorithms-mossstream/index.html","parent":"chapter:foundations","order":144,"meta":[["Source","BanditRLProof/Algorithms/MOSSStream.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossstream banditrlproof.algorithms.mossstream generated source map for banditrlproof/algorithms/mossstream.lean. lean module compiled","shard":"modules/1a1e43c6a99ac364.json"},{"id":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","label":"Algorithms.MOSSStreamMeasurable","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSStreamMeasurable","description":"Generated source map for BanditRLProof/Algorithms/MOSSStreamMeasurable.lean.","url":"../modules/banditrlproof-algorithms-mossstreammeasurable/index.html","parent":"chapter:foundations","order":145,"meta":[["Source","BanditRLProof/Algorithms/MOSSStreamMeasurable.lean"],["Declarations","5"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossstreammeasurable banditrlproof.algorithms.mossstreammeasurable generated source map for banditrlproof/algorithms/mossstreammeasurable.lean. lean module compiled","shard":"modules/49bb04ef7e29c487.json"},{"id":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","label":"Algorithms.MOSSUnusedCoordinate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MOSSUnusedCoordinate","description":"Generated source map for BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean.","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html","parent":"chapter:foundations","order":146,"meta":[["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.mossunusedcoordinate banditrlproof.algorithms.mossunusedcoordinate generated source map for banditrlproof/algorithms/mossunusedcoordinate.lean. lean module compiled","shard":"modules/bb7fe370b32d2783.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsCollision","label":"Algorithms.MusicalChairsCollision","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsCollision","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsCollision.lean.","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html","parent":"chapter:foundations","order":147,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean"],["Declarations","34"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairscollision banditrlproof.algorithms.musicalchairscollision generated source map for banditrlproof/algorithms/musicalchairscollision.lean. lean module compiled","shard":"modules/48120dc14d0559d2.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","label":"Algorithms.MusicalChairsCoordination","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsCoordination","description":"Static Musical Chairs coordination component (Rosenski--Shamir--Szlak, ICML 2016, Algorithm 2 / supplement A.1 Lemma 4). Candidate sets are explicit outputs required from the still-separate exploration phase. This module proves the actual local transition, product-event law and statewise fixation hazard; it is not the full unknown-N algorithm or a regret endpoint.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html","parent":"chapter:foundations","order":148,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean"],["Declarations","38"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairscoordination banditrlproof.algorithms.musicalchairscoordination static musical chairs coordination component (rosenski--shamir--szlak, icml 2016, algorithm 2 / supplement a.1 lemma 4). candidate sets are explicit outputs required from the still-separate exploration phase. this module proves the actual local transition, product-event law and statewise fixation hazard; it is not the full unknown-n algorithm or a regret endpoint. lean module compiled","shard":"modules/1e0585c01fb5e6bb.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","label":"Algorithms.MusicalChairsCoordinationRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean.","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html","parent":"chapter:foundations","order":149,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairscoordinationregret banditrlproof.algorithms.musicalchairscoordinationregret generated source map for banditrlproof/algorithms/musicalchairscoordinationregret.lean. lean module compiled","shard":"modules/9950e6d728c0a33c.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","label":"Algorithms.MusicalChairsCoordinationTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsCoordinationTime","description":"Quantitative continuation of the actual static coordination kernel. This module does not provide the exploration-good event or the full learner.","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html","parent":"chapter:foundations","order":150,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairscoordinationtime banditrlproof.algorithms.musicalchairscoordinationtime quantitative continuation of the actual static coordination kernel. this module does not provide the exploration-good event or the full learner. lean module compiled","shard":"modules/ab2f68cc70feb0e8.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsExploration","label":"Algorithms.MusicalChairsExploration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsExploration","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsExploration.lean.","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html","parent":"chapter:foundations","order":151,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean"],["Declarations","30"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairsexploration banditrlproof.algorithms.musicalchairsexploration generated source map for banditrlproof/algorithms/musicalchairsexploration.lean. lean module compiled","shard":"modules/f7c1f4654d145886.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","label":"Algorithms.MusicalChairsHandoff","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsHandoff","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsHandoff.lean.","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html","parent":"chapter:foundations","order":152,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean"],["Declarations","33"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairshandoff banditrlproof.algorithms.musicalchairshandoff generated source map for banditrlproof/algorithms/musicalchairshandoff.lean. lean module compiled","shard":"modules/b0ed41d04a4a0fad.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","label":"Algorithms.MusicalChairsLearnerRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsLearnerRegret","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean.","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html","parent":"chapter:foundations","order":153,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean"],["Declarations","32"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairslearnerregret banditrlproof.algorithms.musicalchairslearnerregret generated source map for banditrlproof/algorithms/musicalchairslearnerregret.lean. lean module compiled","shard":"modules/2e493f6267affb53.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","label":"Algorithms.MusicalChairsMarginal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsMarginal","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsMarginal.lean.","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html","parent":"chapter:foundations","order":154,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean"],["Declarations","24"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairsmarginal banditrlproof.algorithms.musicalchairsmarginal generated source map for banditrlproof/algorithms/musicalchairsmarginal.lean. lean module compiled","shard":"modules/1cef4aee8da8ddd2.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","label":"Algorithms.MusicalChairsPopulation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsPopulation","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsPopulation.lean.","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html","parent":"chapter:foundations","order":155,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean"],["Declarations","30"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairspopulation banditrlproof.algorithms.musicalchairspopulation generated source map for banditrlproof/algorithms/musicalchairspopulation.lean. lean module compiled","shard":"modules/7c20d341d7664101.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsRanking","label":"Algorithms.MusicalChairsRanking","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsRanking","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsRanking.lean.","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html","parent":"chapter:foundations","order":156,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean"],["Declarations","27"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairsranking banditrlproof.algorithms.musicalchairsranking generated source map for banditrlproof/algorithms/musicalchairsranking.lean. lean module compiled","shard":"modules/ebfe70a38ba5b933.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsRealized","label":"Algorithms.MusicalChairsRealized","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsRealized","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsRealized.lean.","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html","parent":"chapter:foundations","order":157,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean"],["Declarations","50"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairsrealized banditrlproof.algorithms.musicalchairsrealized generated source map for banditrlproof/algorithms/musicalchairsrealized.lean. lean module compiled","shard":"modules/8a53c5ceb6f3f643.json"},{"id":"module:BanditRLProof.Algorithms.MusicalChairsReward","label":"Algorithms.MusicalChairsReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.MusicalChairsReward","description":"Generated source map for BanditRLProof/Algorithms/MusicalChairsReward.lean.","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html","parent":"chapter:foundations","order":158,"meta":[["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean"],["Declarations","38"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"algorithms.musicalchairsreward banditrlproof.algorithms.musicalchairsreward generated source map for banditrlproof/algorithms/musicalchairsreward.lean. lean module compiled","shard":"modules/b606c5c3a46ea860.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","label":"Algorithms.StochasticGradientBanditAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditAudit","description":"This module formalizes the finite-action algebra in Algorithm 1 and Equations (3)--(7) of Baudry--Johnson--Vary--Pike-Burke--Rebeschini (NeurIPS 2025).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html","parent":"chapter:frontier","order":159,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean"],["Declarations","26"],["Project imports","0"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbanditaudit banditrlproof.algorithms.stochasticgradientbanditaudit this module formalizes the finite-action algebra in algorithm 1 and equations (3)--(7) of baudry--johnson--vary--pike-burke--rebeschini (neurips 2025). lean module compiled","shard":"modules/ea45735030a545a9.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","label":"Algorithms.StochasticGradientBanditConditionalExponentialAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","description":"This module composes the source-exact bounded-reward Equation (8) with the actual action/reward kernel generated by the SGB history policy. Its main contract permits an action-dependent exponent q; this is the form needed by the two-arm Theorem-1 proof, where the selected-arm branches use different coefficients.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditconditionalexponentialaudit/index.html","parent":"chapter:frontier","order":160,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditConditionalExponentialAudit.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbanditconditionalexponentialaudit banditrlproof.algorithms.stochasticgradientbanditconditionalexponentialaudit this module composes the source-exact bounded-reward equation (8) with the actual action/reward kernel generated by the sgb history policy. its main contract permits an action-dependent exponent q; this is the form needed by the two-arm theorem-1 proof, where the selected-arm branches use different coefficients. lean module compiled","shard":"modules/fc97f5393dd37dd1.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","label":"Algorithms.StochasticGradientBanditCorollaryOne","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","description":"This module closes the bounded Corollary-1 companion from Baudry, Johnson, Vary, Pike-Burke, and Rebeschini, *Does Stochastic Gradient really succeed for Bandits?* (NeurIPS 2025). It remains on the generated Algorithm-1 trajectory and uses a separate fixed learning rate sqrt (log T / T) for each source horizon T.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html","parent":"chapter:frontier","order":161,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean"],["Declarations","23"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbanditcorollaryone banditrlproof.algorithms.stochasticgradientbanditcorollaryone this module closes the bounded corollary-1 companion from baudry, johnson, vary, pike-burke, and rebeschini, *does stochastic gradient really succeed for bandits?* (neurips 2025). it remains on the generated algorithm-1 trajectory and uses a separate fixed learning rate sqrt (log t / t) for each source horizon t. lean module compiled","shard":"modules/a7d847c979a636f6.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","label":"Algorithms.StochasticGradientBanditExponentialAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","description":"This module formalizes the source constant C_eta and Equation (8) from the two-arm proof of Baudry--Johnson--Vary--Pike-Burke--Rebeschini (NeurIPS 2025). For an almost-everywhere measurable reward supported on [-1, 1], it derives the exact second-order moment-generating-function inequality used by the paper. It also proves the source comparison C_eta <= exp (2 * eta) for nonnegative eta.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html","parent":"chapter:frontier","order":162,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbanditexponentialaudit banditrlproof.algorithms.stochasticgradientbanditexponentialaudit this module formalizes the source constant c_eta and equation (8) from the two-arm proof of baudry--johnson--vary--pike-burke--rebeschini (neurips 2025). for an almost-everywhere measurable reward supported on [-1, 1], it derives the exact second-order moment-generating-function inequality used by the paper. it also proves the source comparison c_eta <= exp (2 * eta) for nonnegative eta. lean module compiled","shard":"modules/bb4f78b4a499d0ac.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","label":"Algorithms.StochasticGradientBanditTheoremFourContractAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","description":"This module isolates finite scalar obligations from Appendix E, Steps 1, 3, and 4, of Baudry, Johnson, Vary, Pike-Burke, and Rebeschini, Does Stochastic Gradient really succeed for Bandits?* (NeurIPS 2025).","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html","parent":"chapter:frontier","order":163,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremfourcontractaudit banditrlproof.algorithms.stochasticgradientbandittheoremfourcontractaudit this module isolates finite scalar obligations from appendix e, steps 1, 3, and 4, of baudry, johnson, vary, pike-burke, and rebeschini, does stochastic gradient really succeed for bandits?* (neurips 2025). lean module compiled","shard":"modules/691fed6ec367d723.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","label":"Algorithms.StochasticGradientBanditTheoremTwoLatentReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","description":"Appendix C of Baudry--Johnson--Vary--Pike-Burke--Rebeschini treats the rewards collected from the optimal arm in pull order. This module exposes the safe part of that reindexing through the existing latent arm-stream coupling.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html","parent":"chapter:frontier","order":164,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean"],["Declarations","7"],["Project imports","3"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremtwolatentreward banditrlproof.algorithms.stochasticgradientbandittheoremtwolatentreward appendix c of baudry--johnson--vary--pike-burke--rebeschini treats the rewards collected from the optimal arm in pull order. this module exposes the safe part of that reindexing through the existing latent arm-stream coupling. lean module compiled","shard":"modules/5dca51c8946f1a6f.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","label":"Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","description":"The latent arm-stream coupling draws every reward coordinate before the algorithm runs, whereas the native fixed-IID process draws only the reward of the arm actually selected in each round. The compiled deterministic-time one-step selected-reward laws describe the coupling; this module states the native process against which those laws must be compared.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html","parent":"chapter:frontier","order":165,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremtwonativeprefix banditrlproof.algorithms.stochasticgradientbandittheoremtwonativeprefix the latent arm-stream coupling draws every reward coordinate before the algorithm runs, whereas the native fixed-iid process draws only the reward of the arm actually selected in each round. the compiled deterministic-time one-step selected-reward laws describe the coupling; this module states the native process against which those laws must be compared. lean module compiled","shard":"modules/91b486b71f97e146.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","label":"Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","description":"The latent arm-stream coupling samples every reward coordinate before the algorithm runs. The native fixed-IID process samples only the reward selected at each round. Connecting those constructions requires a deferred-decisions argument, not merely equality of one-coordinate marginals.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html","parent":"chapter:frontier","order":166,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean"],["Declarations","55"],["Project imports","5"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremtwonativetrajectory banditrlproof.algorithms.stochasticgradientbandittheoremtwonativetrajectory the latent arm-stream coupling samples every reward coordinate before the algorithm runs. the native fixed-iid process samples only the reward selected at each round. connecting those constructions requires a deferred-decisions argument, not merely equality of one-coordinate marginals. lean module compiled","shard":"modules/fbd8bc8e834f8afc.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","label":"Algorithms.StochasticGradientBanditTheoremTwoNthPull","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","description":"Appendix C of Baudry--Johnson--Vary--Pike-Burke--Rebeschini reindexes the optimal-arm dynamics by the number of times that arm has been selected. The generated Lean trajectory is instead indexed by chronological time. This module supplies the missing bridge without assuming that the adaptively selected reward subsequence is IID.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html","parent":"chapter:frontier","order":167,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean"],["Declarations","26"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremtwonthpull banditrlproof.algorithms.stochasticgradientbandittheoremtwonthpull appendix c of baudry--johnson--vary--pike-burke--rebeschini reindexes the optimal-arm dynamics by the number of times that arm has been selected. the generated lean trajectory is instead indexed by chronological time. this module supplies the missing bridge without assuming that the adaptively selected reward subsequence is iid. lean module compiled","shard":"modules/e4dc006c54eb66c5.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","label":"Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","description":"Appendix C of Baudry--Johnson--Vary--Pike-Burke--Rebeschini works with rewards in optimal-arm pull order. Adaptive selection makes a naive stopped-value IID statement false as a proof interface: a requested pull can be absent, and the event that a block of pulls occurs can itself depend on earlier rewards.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html","parent":"chapter:frontier","order":168,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean"],["Declarations","43"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremtwoselectediid banditrlproof.algorithms.stochasticgradientbandittheoremtwoselectediid appendix c of baudry--johnson--vary--pike-burke--rebeschini works with rewards in optimal-arm pull order. adaptive selection makes a naive stopped-value iid statement false as a proof interface: a requested pull can be absent, and the event that a block of pulls occurs can itself depend on earlier rewards. lean module compiled","shard":"modules/0bcc6ac641d11093.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","label":"Algorithms.StochasticGradientBanditTheoremTwoStarvation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","description":"This module isolates the deterministic consumer in Step 1 of Appendix C of Baudry--Johnson--Vary--Pike-Burke--Rebeschini (NeurIPS 2025). On the actual generated SGB action/reward trace, once exactly n optimal-arm pulls have occurred, a path with no later optimal-arm pull has exactly Delta * (T - n) sampled pseudo-regret.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html","parent":"chapter:frontier","order":169,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean"],["Declarations","26"],["Project imports","4"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittheoremtwostarvation banditrlproof.algorithms.stochasticgradientbandittheoremtwostarvation this module isolates the deterministic consumer in step 1 of appendix c of baudry--johnson--vary--pike-burke--rebeschini (neurips 2025). on the actual generated sgb action/reward trace, once exactly n optimal-arm pulls have occurred, a path with no later optimal-arm pull has exactly delta * (t - n) sampled pseudo-regret. lean module compiled","shard":"modules/0a3751859cc7e6fb.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","label":"Algorithms.StochasticGradientBanditTrajectoryAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","description":"This module lifts the finite Algorithm-1 / Equation-(5) algebra in StochasticGradientBanditAudit to the repository's canonical measurable action/reward trajectory. It constructs the recursive parameter vector, packages its softmax law as a history algorithm, and identifies the generated successor action and pair conditional laws.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html","parent":"chapter:frontier","order":170,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean"],["Declarations","18"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittrajectoryaudit banditrlproof.algorithms.stochasticgradientbandittrajectoryaudit this module lifts the finite algorithm-1 / equation-(5) algebra in stochasticgradientbanditaudit to the repository's canonical measurable action/reward trajectory. it constructs the recursive parameter vector, packages its softmax law as a history algorithm, and identifies the generated successor action and pair conditional laws. lean module compiled","shard":"modules/d5a41e0521dc0920.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","label":"Algorithms.StochasticGradientBanditTwoArmFixedIID","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","description":"This module realizes the reward model used by the two-arm source theorem from an arm-indexed family of fixed probability laws. The environment is the stationary history environment over Unit; hence its initial and successor reward fibers are exactly the selected arm law and cannot reveal a latent, changing environment.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html","parent":"chapter:frontier","order":171,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean"],["Declarations","9"],["Project imports","3"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmfixediid banditrlproof.algorithms.stochasticgradientbandittwoarmfixediid this module realizes the reward model used by the two-arm source theorem from an arm-indexed family of fixed probability laws. the environment is the stationary history environment over unit; hence its initial and successor reward fibers are exactly the selected arm law and cannot reveal a latent, changing environment. lean module compiled","shard":"modules/6afc2c2832aaebe3.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","label":"Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","description":"This module closes the source-round t = 1 base case for the two exponential recurrences used in Theorem 1 of Baudry--Johnson--Vary--Pike-Burke--Rebeschini (NeurIPS 2025). The generated initial pair kernel samples from the untouched source parameter theta_1 = 0, hence its action law is uniform on Fin 2; consuming that pair produces theta_2.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarminitialrecurrence/index.html","parent":"chapter:frontier","order":172,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmInitialRecurrence.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarminitialrecurrence banditrlproof.algorithms.stochasticgradientbandittwoarminitialrecurrence this module closes the source-round t = 1 base case for the two exponential recurrences used in theorem 1 of baudry--johnson--vary--pike-burke--rebeschini (neurips 2025). the generated initial pair kernel samples from the untouched source parameter theta_1 = 0, hence its action law is uniform on fin 2; consuming that pair produces theta_2. lean module compiled","shard":"modules/50a1a42d6bbf9223.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","label":"Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","description":"This module lifts the fixed-history recurrences to the jointly measurable environment kernel used by the canonical SGB trajectory. It also records one uniform, source-faithful environment contract: every initial and successor reward fiber is supported in [-1,1] almost everywhere and has the same fixed arm mean, uniformly over environment values, times, and finite histories.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html","parent":"chapter:frontier","order":173,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmmeasurablerecurrence banditrlproof.algorithms.stochasticgradientbandittwoarmmeasurablerecurrence this module lifts the fixed-history recurrences to the jointly measurable environment kernel used by the canonical sgb trajectory. it also records one uniform, source-faithful environment contract: every initial and successor reward fiber is supported in [-1,1] almost everywhere and has the same fixed arm mean, uniformly over environment values, times, and finite histories. lean module compiled","shard":"modules/f4d40b144252fac4.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","label":"Algorithms.StochasticGradientBanditTwoArmPathIntegrability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","description":"This module closes the finite-time regularity boundary needed before the conditional-distribution recurrences can be used in a tower argument.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html","parent":"chapter:frontier","order":174,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmpathintegrability banditrlproof.algorithms.stochasticgradientbandittwoarmpathintegrability this module closes the finite-time regularity boundary needed before the conditional-distribution recurrences can be used in a tower argument. lean module compiled","shard":"modules/3bf2a8d9e511a7c8.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","label":"Algorithms.StochasticGradientBanditTwoArmRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","description":"This module starts the source-facing rate layer for Theorem 1 of Baudry--Johnson--Vary--Pike-Burke--Rebeschini (NeurIPS 2025). It proves that Algorithm 1 preserves the zero sum of its parameter vector on every generated finite history, then specializes the two-arm softmax law to the exact odds identities used as Equation (11) in the source proof.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html","parent":"chapter:frontier","order":175,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmrate banditrlproof.algorithms.stochasticgradientbandittwoarmrate this module starts the source-facing rate layer for theorem 1 of baudry--johnson--vary--pike-burke--rebeschini (neurips 2025). it proves that algorithm 1 preserves the zero sum of its parameter vector on every generated finite history, then specializes the two-arm softmax law to the exact odds identities used as equation (11) in the source proof. lean module compiled","shard":"modules/e1923f1f1d395ac4.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","label":"Algorithms.StochasticGradientBanditTwoArmRecurrence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","description":"This module instantiates the generated conditional Equation (8) at the two action-dependent coefficients used in the proof of Theorem 1 of Baudry--Johnson--Vary--Pike-Burke--Rebeschini (NeurIPS 2025). It proves both one-step exponential recurrences over the actual generated history-step kernel, under the source initialization theta = 0.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html","parent":"chapter:frontier","order":176,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmrecurrence banditrlproof.algorithms.stochasticgradientbandittwoarmrecurrence this module instantiates the generated conditional equation (8) at the two action-dependent coefficients used in the proof of theorem 1 of baudry--johnson--vary--pike-burke--rebeschini (neurips 2025). it proves both one-step exponential recurrences over the actual generated history-step kernel, under the source initialization theta = 0. lean module compiled","shard":"modules/2f9129badec0ede0.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","label":"Algorithms.StochasticGradientBanditTwoArmTheoremOne","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","description":"This module closes Theorem 1 of Baudry, Johnson, Vary, Pike-Burke, and Rebeschini, *Does Stochastic Gradient really succeed for Bandits?* (NeurIPS 2025), for the paper's fixed two-arm IID reward model.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html","parent":"chapter:frontier","order":177,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean"],["Declarations","32"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmtheoremone banditrlproof.algorithms.stochasticgradientbandittwoarmtheoremone this module closes theorem 1 of baudry, johnson, vary, pike-burke, and rebeschini, *does stochastic gradient really succeed for bandits?* (neurips 2025), for the paper's fixed two-arm iid reward model. lean module compiled","shard":"modules/166646e93f03ff43.json"},{"id":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","label":"Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","description":"This module integrates the tower-ready conditional-expectation recurrences on the canonical trajectory and performs the finite scalar iterations used by the two-arm Theorem-1 proof.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html","parent":"chapter:frontier","order":178,"meta":[["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean"],["Declarations","37"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"algorithms.stochasticgradientbandittwoarmunconditionalrecurrence banditrlproof.algorithms.stochasticgradientbandittwoarmunconditionalrecurrence this module integrates the tower-ready conditional-expectation recurrences on the canonical trajectory and performs the finite scalar iterations used by the two-arm theorem-1 proof. lean module compiled","shard":"modules/4b10006a2c474128.json"},{"id":"module:BanditRLProof.Algorithms.Thompson","label":"Algorithms.Thompson","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.Thompson","description":"Thompson sampling and Bayesian regret surfaces","url":"../modules/banditrlproof-algorithms-thompson/index.html","parent":"chapter:thompson","order":179,"meta":[["Source","BanditRLProof/Algorithms/Thompson.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompson banditrlproof.algorithms.thompson thompson sampling and bayesian regret surfaces lean module compiled","shard":"modules/b2be9e5cabd45d88.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","label":"Algorithms.ThompsonAlgorithmDensity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonAlgorithmDensity","description":"This module isolates the measure-theoretic core of LML's algorithm-density posterior transport. If the actual history law and the actual history/environment joint law are obtained from their reference counterparts by the same density depending only on history, then the reference and actual environment posteriors agree at the actual history law.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html","parent":"chapter:thompson","order":180,"meta":[["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonalgorithmdensity banditrlproof.algorithms.thompsonalgorithmdensity this module isolates the measure-theoretic core of lml's algorithm-density posterior transport. if the actual history law and the actual history/environment joint law are obtained from their reference counterparts by the same density depending only on history, then the reference and actual environment posteriors agree at the actual history law. lean module compiled","shard":"modules/e097ae87eb6a9118.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","label":"Algorithms.ThompsonAlgorithmDensityProcess","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","description":"This module ports the process-facing core of LML's algorithm-density theorem. Two stochastic history policies interact with the same feedback environment. If every actual action law is absolutely continuous with respect to the reference action law, the actual finite action/reward history law is the reference history law weighted by the recursive product of policy Radon-Nikodym derivatives.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html","parent":"chapter:thompson","order":181,"meta":[["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean"],["Declarations","31"],["Project imports","2"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonalgorithmdensityprocess banditrlproof.algorithms.thompsonalgorithmdensityprocess this module ports the process-facing core of lml's algorithm-density theorem. two stochastic history policies interact with the same feedback environment. if every actual action law is absolutely continuous with respect to the reference action law, the actual finite action/reward history law is the reference history law weighted by the recursive product of policy radon-nikodym derivatives. lean module compiled","shard":"modules/0f531e9851a10d66.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","label":"Algorithms.ThompsonBayesRegretDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","description":"This module ports the probability-matching algebra behind LML's TS.integral_regret_eq_add to the locally generated Thompson trajectory. The confidence score is abstract but must depend only on the visible history and the candidate action. Clipped-UCB concentration can therefore be attached downstream without restoring an assumed sampler or posterior law.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html","parent":"chapter:thompson","order":182,"meta":[["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean"],["Declarations","16"],["Project imports","2"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonbayesregretdecomposition banditrlproof.algorithms.thompsonbayesregretdecomposition this module ports the probability-matching algebra behind lml's ts.integral_regret_eq_add to the locally generated thompson trajectory. the confidence score is abstract but must depend only on the visible history and the candidate action. clipped-ucb concentration can therefore be attached downstream without restoring an assumed sampler or posterior law. lean module compiled","shard":"modules/c7c8af48cb268eac.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","label":"Algorithms.ThompsonCanonicalSampler","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonCanonicalSampler","description":"This module constructs the one-step joint law obtained by sampling an environment/history pair from a Bayesian prior-likelihood model and then sampling an action from the canonical posterior pushed through a measurable best-action selector. The resulting probability-matching theorem has no separate pair-law or action-law premise.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html","parent":"chapter:thompson","order":183,"meta":[["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsoncanonicalsampler banditrlproof.algorithms.thompsoncanonicalsampler this module constructs the one-step joint law obtained by sampling an environment/history pair from a bayesian prior-likelihood model and then sampling an action from the canonical posterior pushed through a measurable best-action selector. the resulting probability-matching theorem has no separate pair-law or action-law premise. lean module compiled","shard":"modules/7a9bfcad3c6d2758.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","label":"Algorithms.ThompsonCanonicalTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","description":"This module realizes a HistoryAlgorithm interacting with one fixed HistoryEnvironment on Mathlib's Ionescu-Tulcea trajMeasure. It proves both the combined action/reward process contract and the four split conditional-law fields used by the environment-indexed Thompson density route.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html","parent":"chapter:thompson","order":184,"meta":[["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean"],["Declarations","31"],["Project imports","1"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsoncanonicaltrajectory banditrlproof.algorithms.thompsoncanonicaltrajectory this module realizes a historyalgorithm interacting with one fixed historyenvironment on mathlib's ionescu-tulcea trajmeasure. it proves both the combined action/reward process contract and the four split conditional-law fields used by the environment-indexed thompson density route. lean module compiled","shard":"modules/4d6cbf6acc86ef2a.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","label":"Algorithms.ThompsonClippedUCBScore","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonClippedUCBScore","description":"This module instantiates HistoryActionScore with the clipped upper-confidence score used by the pinned LML Thompson regret proof. The finite-history score is measurable and lies in [l, u]; those bounds discharge every score integrability premise in the compiled Bayesian-regret decomposition.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html","parent":"chapter:thompson","order":185,"meta":[["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean"],["Declarations","20"],["Project imports","3"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonclippeducbscore banditrlproof.algorithms.thompsonclippeducbscore this module instantiates historyactionscore with the clipped upper-confidence score used by the pinned lml thompson regret proof. the finite-history score is measurable and lies in [l, u]; those bounds discharge every score integrability premise in the compiled bayesian-regret decomposition. lean module compiled","shard":"modules/b76d90399c48b5dd.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","label":"Algorithms.ThompsonMeasurableTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","description":"The pointwise HistoryEnvironment API does not by itself say that feedback laws vary measurably with the environment. This module records that missing joint regularity and uses Mathlib's kernel-valued Ionescu-Tulcea theorem to construct the complete environment-indexed pair-trajectory kernel.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html","parent":"chapter:thompson","order":186,"meta":[["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean"],["Declarations","31"],["Project imports","2"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonmeasurabletrajectory banditrlproof.algorithms.thompsonmeasurabletrajectory the pointwise historyenvironment api does not by itself say that feedback laws vary measurably with the environment. this module records that missing joint regularity and uses mathlib's kernel-valued ionescu-tulcea theorem to construct the complete environment-indexed pair-trajectory kernel. lean module compiled","shard":"modules/a712a405dcde03e2.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","label":"Algorithms.ThompsonRecursiveSampler","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonRecursiveSampler","description":"This module closes the gap between a separately adjoined one-step posterior sampler and the action coordinate of one recursive trajectory. The Thompson policy is defined non-circularly from a fixed reference trajectory, following the uniform-reference design of LML's TS.policy.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html","parent":"chapter:thompson","order":187,"meta":[["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonrecursivesampler banditrlproof.algorithms.thompsonrecursivesampler this module closes the gap between a separately adjoined one-step posterior sampler and the action coordinate of one recursive trajectory. the thompson policy is defined non-circularly from a fixed reference trajectory, following the uniform-reference design of lml's ts.policy. lean module compiled","shard":"modules/c3b4d82ee57cc6a8.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","label":"Algorithms.ThompsonReferencePolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonReferencePolicy","description":"LML defines Thompson sampling from the posterior under a fixed reference algorithm and then transports that posterior to the actual process by an algorithm-density theorem. This module isolates the corresponding local Mathlib boundary.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html","parent":"chapter:thompson","order":188,"meta":[["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean"],["Declarations","17"],["Project imports","2"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonreferencepolicy banditrlproof.algorithms.thompsonreferencepolicy lml defines thompson sampling from the posterior under a fixed reference algorithm and then transports that posterior to the actual process by an algorithm-density theorem. this module isolates the corresponding local mathlib boundary. lean module compiled","shard":"modules/5b47603f52531259.json"},{"id":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","label":"Algorithms.ThompsonStationaryReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.ThompsonStationaryReward","description":"This module isolates the stationary feedback adapter and the algorithm-independent latent-arm-stream tail used by the Thompson Bayesian-regret route. The tail theorems deliberately quantify over an arbitrary action trace: only the next-unused-coordinate reward rule and the product arm-stream law matter.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html","parent":"chapter:thompson","order":189,"meta":[["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean"],["Declarations","97"],["Project imports","2"],["Teaching chapter","Thompson sampling"]],"statement":"","missing":[],"search":"algorithms.thompsonstationaryreward banditrlproof.algorithms.thompsonstationaryreward this module isolates the stationary feedback adapter and the algorithm-independent latent-arm-stream tail used by the thompson bayesian-regret route. the tail theorems deliberately quantify over an arbitrary action trace: only the next-unused-coordinate reward rule and the product arm-stream law matter. lean module compiled","shard":"modules/2357638fc93777eb.json"},{"id":"module:BanditRLProof.Algorithms.UCB","label":"Algorithms.UCB","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCB","description":"UCB surfaces","url":"../modules/banditrlproof-algorithms-ucb/index.html","parent":"chapter:ucb","order":190,"meta":[["Source","BanditRLProof/Algorithms/UCB.lean"],["Declarations","142"],["Project imports","8"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucb banditrlproof.algorithms.ucb ucb surfaces lean module compiled","shard":"modules/bb0657d69d348d5d.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","label":"Algorithms.UCBArmStreamAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamAsymptotics","description":"This module keeps one recursive armStreamAction and one armStreamMeasure fixed across all horizons. At exploration scale c = 4, the finite tail term in the exact LML-shaped regret bound is uniformly bounded by a convergent NNReal p-series, yielding logarithmic expected regret and vanishing expected average regret for the same policy and measure.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html","parent":"chapter:ucb","order":191,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamasymptotics banditrlproof.algorithms.ucbarmstreamasymptotics this module keeps one recursive armstreamaction and one armstreammeasure fixed across all horizons. at exploration scale c = 4, the finite tail term in the exact lml-shaped regret bound is uniformly bounded by a convergent nnreal p-series, yielding logarithmic expected regret and vanishing expected average regret for the same policy and measure. lean module compiled","shard":"modules/cf57712367db9821.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","label":"Algorithms.UCBArmStreamConditionalReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamConditionalReward","description":"This module isolates the product-measure part of the selected-reward route. Every (pull index, arm) coordinate is independent of the function containing all other coordinates, so its conditional distribution given that complement is the prescribed stationary arm law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html","parent":"chapter:ucb","order":192,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean"],["Declarations","46"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamconditionalreward banditrlproof.algorithms.ucbarmstreamconditionalreward this module isolates the product-measure part of the selected-reward route. every (pull index, arm) coordinate is independent of the function containing all other coordinates, so its conditional distribution given that complement is the prescribed stationary arm law. lean module compiled","shard":"modules/e3cce453d6a2762e.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","label":"Algorithms.UCBArmStreamExpectedPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","description":"This module connects the source-faithful one-sided index tails to the local selected-small/selected-large pull-count decomposition. It stays in ENNReal until a downstream Bochner-regret wrapper requests a Real expectation.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html","parent":"chapter:ucb","order":193,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean"],["Declarations","39"],["Project imports","7"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamexpectedpullcount banditrlproof.algorithms.ucbarmstreamexpectedpullcount this module connects the source-faithful one-sided index tails to the local selected-small/selected-large pull-count decomposition. it stays in ennreal until a downstream bochner-regret wrapper requests a real expectation. lean module compiled","shard":"modules/bf6ede8be12dac50.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","label":"Algorithms.UCBArmStreamFiniteArmRewardLaws","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","description":"This module packages stationary Real-valued arm laws as a Mathlib kernel and instantiates the canonical arm-stream expected-consistency theorem. The final bounded-law endpoint keeps one recursive policy and one product measure fixed across all horizons.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html","parent":"chapter:ucb","order":194,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamfinitearmrewardlaws banditrlproof.algorithms.ucbarmstreamfinitearmrewardlaws this module packages stationary real-valued arm laws as a mathlib kernel and instantiates the canonical arm-stream expected-consistency theorem. the final bounded-law endpoint keeps one recursive policy and one product measure fixed across all horizons. lean module compiled","shard":"modules/1b9fd22e2fff11f6.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","label":"Algorithms.UCBArmStreamProcess","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamProcess","description":"This module mirrors the deterministic part of the pinned LML array model. It recursively builds the inclusive action/reward history, uses round-robin initialization followed by the native Real history-index selector, and reads the selected arm's next unused latent reward coordinate.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html","parent":"chapter:ucb","order":195,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamprocess banditrlproof.algorithms.ucbarmstreamprocess this module mirrors the deterministic part of the pinned lml array model. it recursively builds the inclusive action/reward history, uses round-robin initialization followed by the native real history-index selector, and reads the selected arm's next unused latent reward coordinate. lean module compiled","shard":"modules/efbe8812ce2ea351.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamSource","label":"Algorithms.UCBArmStreamSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamSource","description":"This module implements the pathwise reward-consumption rule used by the pinned LML UCB route. Whenever an action is selected, the observed reward is the next unused coordinate of that arm's latent stream. Consequently, the selected reward sum is exactly the stream prefix whose length is the realized pull count, so it supplies FixedArmPrefixSource without an extra pathwise hypothesis.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html","parent":"chapter:ucb","order":196,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamsource banditrlproof.algorithms.ucbarmstreamsource this module implements the pathwise reward-consumption rule used by the pinned lml ucb route. whenever an action is selected, the observed reward is the next unused coordinate of that arm's latent stream. consequently, the selected reward sum is exactly the stream prefix whose length is the realized pull count, so it supplies fixedarmprefixsource without an extra pathwise hypothesis. lean module compiled","shard":"modules/2783add4e3846770.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmStreamTail","label":"Algorithms.UCBArmStreamTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmStreamTail","description":"The stationary product arm-stream measure makes each fixed arm an independent reward trace with its prescribed kernel law. This module transports centered sub-Gaussian witnesses to that stream space and combines the resulting fixed prefix tails with the compiled UCB fixed-count peeling theorem.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html","parent":"chapter:ucb","order":197,"meta":[["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean"],["Declarations","38"],["Project imports","3"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmstreamtail banditrlproof.algorithms.ucbarmstreamtail the stationary product arm-stream measure makes each fixed arm an independent reward trace with its prescribed kernel law. this module transports centered sub-gaussian witnesses to that stream space and combines the resulting fixed prefix tails with the compiled ucb fixed-count peeling theorem. lean module compiled","shard":"modules/1cbc2bb47b2d036a.json"},{"id":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","label":"Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","description":"This module derives the direct centered sub-Gaussian contracts required by the stationary finite-arm sampled-pair consistency theorem from arm-dependent almost-sure reward intervals and exact arm means.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html","parent":"chapter:ucb","order":198,"meta":[["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbarmwiseboundedfinitearmsampledasymptotics banditrlproof.algorithms.ucbarmwiseboundedfinitearmsampledasymptotics this module derives the direct centered sub-gaussian contracts required by the stationary finite-arm sampled-pair consistency theorem from arm-dependent almost-sure reward intervals and exact arm means. lean module compiled","shard":"modules/d5057ea7d0211b1c.json"},{"id":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","label":"Algorithms.UCBBoundedFiniteArmRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","description":"This module instantiates the centered-kernel canonical Real theorem with stationary action-indexed reward laws. It supports both one common interval and arm-dependent nondegenerate intervals.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmrewardlaw/index.html","parent":"chapter:ucb","order":199,"meta":[["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmRewardLaw.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbboundedfinitearmrewardlaw banditrlproof.algorithms.ucbboundedfinitearmrewardlaw this module instantiates the centered-kernel canonical real theorem with stationary action-indexed reward laws. it supports both one common interval and arm-dependent nondegenerate intervals. lean module compiled","shard":"modules/e306ae44ca5c44b7.json"},{"id":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","label":"Algorithms.UCBBoundedFiniteArmSampledAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","description":"This module derives the direct centered sub-Gaussian contracts required by the stationary finite-arm sampled-pair consistency theorem from a common almost- sure reward interval and exact arm means.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmsampledasymptotics/index.html","parent":"chapter:ucb","order":200,"meta":[["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmSampledAsymptotics.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbboundedfinitearmsampledasymptotics banditrlproof.algorithms.ucbboundedfinitearmsampledasymptotics this module derives the direct centered sub-gaussian contracts required by the stationary finite-arm sampled-pair consistency theorem from a common almost- sure reward interval and exact arm means. lean module compiled","shard":"modules/62c0a91378211d0f.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","label":"Algorithms.UCBConditionalRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardLaw","description":"This module connects the selected-policy simultaneous empirical-mean event to the deterministic UCB score algebra. Its confidence width depends on the realized pull count, so it deliberately does not use the older UCB.finiteHorizonConfidenceBadEvent, whose radius is deterministic in the sample point.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html","parent":"chapter:ucb","order":201,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardlaw banditrlproof.algorithms.ucbconditionalrewardlaw this module connects the selected-policy simultaneous empirical-mean event to the deterministic ucb score algebra. its confidence width depends on the realized pull count, so it deliberately does not use the older ucb.finitehorizonconfidencebadevent, whose radius is deterministic in the sample point. lean module compiled","shard":"modules/35442cfe57b139df.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","label":"Algorithms.UCBConditionalRewardLawCenteredKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","description":"Generated source map for BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html","parent":"chapter:ucb","order":202,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardlawcenteredkernel banditrlproof.algorithms.ucbconditionalrewardlawcenteredkernel generated source map for banditrlproof/algorithms/ucbconditionalrewardlawcenteredkernel.lean. lean module compiled","shard":"modules/b4dbdbef832594d1.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","label":"Algorithms.UCBConditionalRewardLawCenteredKernelReal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","description":"Generated source map for BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernelReal.lean.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernelreal/index.html","parent":"chapter:ucb","order":203,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernelReal.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardlawcenteredkernelreal banditrlproof.algorithms.ucbconditionalrewardlawcenteredkernelreal generated source map for banditrlproof/algorithms/ucbconditionalrewardlawcenteredkernelreal.lean. lean module compiled","shard":"modules/db412e11f4d68c1c.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","label":"Algorithms.UCBConditionalRewardLawPolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","description":"This module constructs a reward-history-generated finite-arm UCB policy whose score is exactly the realized-count score from UCBConditionalRewardLaw. Successor actions 1, ..., K initialize every arm once. Later actions maximize the practical random-width index computed from the inclusive pair history.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html","parent":"chapter:ucb","order":204,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean"],["Declarations","46"],["Project imports","3"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardlawpolicy banditrlproof.algorithms.ucbconditionalrewardlawpolicy this module constructs a reward-history-generated finite-arm ucb policy whose score is exactly the realized-count score from ucbconditionalrewardlaw. successor actions 1, ..., k initialize every arm once. later actions maximize the practical random-width index computed from the inclusive pair history. lean module compiled","shard":"modules/7a549d50bf316db6.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","label":"Algorithms.UCBConditionalRewardLawRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","description":"This module connects the explicit expected successor pull-count theorem to the local finite-bandit pseudo-regret surface. The regret action shifts generated coordinates 1, ..., T to the standard pull-count coordinates 0, ..., T-1.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html","parent":"chapter:ucb","order":205,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean"],["Declarations","12"],["Project imports","3"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardlawregret banditrlproof.algorithms.ucbconditionalrewardlawregret this module connects the explicit expected successor pull-count theorem to the local finite-bandit pseudo-regret surface. the regret action shifts generated coordinates 1, ..., t to the standard pull-count coordinates 0, ..., t-1. lean module compiled","shard":"modules/1f1de711bc498fa3.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","label":"Algorithms.UCBConditionalRewardLawTrajMeasure","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","description":"This module specializes the canonical reward-only Ionescu-Tulcea trajectory law to selectedPolicySuccessorHistoryPolicy. It closes the selected-reward condExpKernel.map premise used by the practical UCB regret route, while leaving reward-range regularity as a separate consumer obligation.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html","parent":"chapter:ucb","order":206,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardlawtrajmeasure banditrlproof.algorithms.ucbconditionalrewardlawtrajmeasure this module specializes the canonical reward-only ionescu-tulcea trajectory law to selectedpolicysuccessorhistorypolicy. it closes the selected-reward condexpkernel.map premise used by the practical ucb regret route, while leaving reward-range regularity as a separate consumer obligation. lean module compiled","shard":"modules/df390a781c1ccf9d.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","label":"Algorithms.UCBConditionalRewardPairTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","description":"This module consumes the canonical action/reward trajectory simultaneous empirical-mean event in the random-width UCB score algebra. It proves a fixed finite-arm/time large-gap event bound, a positive-gap chosen-arm explicit-threshold tail/ENNReal expected pull-count bound, and the resulting finite-arm explicit-threshold and textbook positive-gap ENNReal pseudo-regret sums. It requires a CenteredRewardKernelLaw, but…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html","parent":"chapter:ucb","order":207,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean"],["Declarations","6"],["Project imports","4"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardpairtrajectory banditrlproof.algorithms.ucbconditionalrewardpairtrajectory this module consumes the canonical action/reward trajectory simultaneous empirical-mean event in the random-width ucb score algebra. it proves a fixed finite-arm/time large-gap event bound, a positive-gap chosen-arm explicit-threshold tail/ennreal expected pull-count bound, and the resulting finite-arm explicit-threshold and textbook positive-gap ennreal pseudo-regret sums. it requires a centeredrewardkernellaw, but no caller selected-reward trajectory law or reward-range premise. it does not prove anytime confidence, an asymptotic normalization, or a real/bochner expectation endpoint. lean module compiled","shard":"modules/defbb2ae1e1d9c48.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","label":"Algorithms.UCBConditionalRewardPairTrajectoryReal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","description":"Generated source map for BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectoryReal.lean.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectoryreal/index.html","parent":"chapter:ucb","order":208,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectoryReal.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardpairtrajectoryreal banditrlproof.algorithms.ucbconditionalrewardpairtrajectoryreal generated source map for banditrlproof/algorithms/ucbconditionalrewardpairtrajectoryreal.lean. lean module compiled","shard":"modules/e1e962c51532b4c6.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","label":"Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","description":"This module runs the canonical pair-trajectory sampled-successor UCB theorem with confidence budget 1 / (T + 1), proves logarithmic expected pseudo-regret for fixed model data, and derives vanishing expected average pseudo-regret for the resulting horizon-indexed policy family.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html","parent":"chapter:ucb","order":209,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardpairtrajectorysampledasymptotics banditrlproof.algorithms.ucbconditionalrewardpairtrajectorysampledasymptotics this module runs the canonical pair-trajectory sampled-successor ucb theorem with confidence budget 1 / (t + 1), proves logarithmic expected pseudo-regret for fixed model data, and derives vanishing expected average pseudo-regret for the resulting horizon-indexed policy family. lean module compiled","shard":"modules/5605cd1a2cef9a33.json"},{"id":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","label":"Algorithms.UCBConditionalRewardPairTrajectorySampledReal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","description":"Generated source map for BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledReal.lean.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledreal/index.html","parent":"chapter:ucb","order":210,"meta":[["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledReal.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbconditionalrewardpairtrajectorysampledreal banditrlproof.algorithms.ucbconditionalrewardpairtrajectorysampledreal generated source map for banditrlproof/algorithms/ucbconditionalrewardpairtrajectorysampledreal.lean. lean module compiled","shard":"modules/594aa16a9d71c0c7.json"},{"id":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","label":"Algorithms.UCBContextDependentBoundedRewardKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","description":"The reward distribution may vary with context and action. Arm means remain stationary and all selected one-step laws share one nondegenerate interval.","url":"../modules/banditrlproof-algorithms-ucbcontextdependentboundedrewardkernel/index.html","parent":"chapter:ucb","order":211,"meta":[["Source","BanditRLProof/Algorithms/UCBContextDependentBoundedRewardKernel.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbcontextdependentboundedrewardkernel banditrlproof.algorithms.ucbcontextdependentboundedrewardkernel the reward distribution may vary with context and action. arm means remain stationary and all selected one-step laws share one nondegenerate interval. lean module compiled","shard":"modules/45ce4fd8138cd9f4.json"},{"id":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","label":"Algorithms.UCBContextDependentSubGaussianRewardKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","description":"The reward distribution and its pointwise sub-Gaussian proxy may vary with context and action. Arm means remain stationary, and a positive uniform proxy ceiling is supplied for the UCB confidence width.","url":"../modules/banditrlproof-algorithms-ucbcontextdependentsubgaussianrewardkernel/index.html","parent":"chapter:ucb","order":212,"meta":[["Source","BanditRLProof/Algorithms/UCBContextDependentSubGaussianRewardKernel.lean"],["Declarations","3"],["Project imports","3"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbcontextdependentsubgaussianrewardkernel banditrlproof.algorithms.ucbcontextdependentsubgaussianrewardkernel the reward distribution and its pointwise sub-gaussian proxy may vary with context and action. arm means remain stationary, and a positive uniform proxy ceiling is supplied for the ucb confidence width. lean module compiled","shard":"modules/584e0b20292298cb.json"},{"id":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","label":"Algorithms.UCBFiniteArmSubGaussianRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","description":"This module instantiates the centered-kernel canonical Real theorem directly from stationary action-indexed sub-Gaussian reward laws, without bounded-support assumptions.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussianrewardlaw/index.html","parent":"chapter:ucb","order":213,"meta":[["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianRewardLaw.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbfinitearmsubgaussianrewardlaw banditrlproof.algorithms.ucbfinitearmsubgaussianrewardlaw this module instantiates the centered-kernel canonical real theorem directly from stationary action-indexed sub-gaussian reward laws, without bounded-support assumptions. lean module compiled","shard":"modules/04cdf84537509fed.json"},{"id":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","label":"Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","description":"This module instantiates the canonical sampled pair-trajectory asymptotic UCB route from stationary action-indexed reward laws. The fixed initial action is paired with a reward sampled from its arm law, while all successor laws use the context-independent Markov reward kernel.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html","parent":"chapter:ucb","order":214,"meta":[["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean"],["Declarations","8"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbfinitearmsubgaussiansampledasymptotics banditrlproof.algorithms.ucbfinitearmsubgaussiansampledasymptotics this module instantiates the canonical sampled pair-trajectory asymptotic ucb route from stationary action-indexed reward laws. the fixed initial action is paired with a reward sampled from its arm law, while all successor laws use the context-independent markov reward kernel. lean module compiled","shard":"modules/c6553666b0c493f1.json"},{"id":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","label":"Algorithms.UCBFixedCountPeeling","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBFixedCountPeeling","description":"This module isolates the law transport used by the pinned LML UCB proof. A FixedArmPrefixSource records the pathwise fact that rewards selected from one arm are the prefix of an arm-indexed reward stream. The main theorems peel the random pull count into finitely many fixed counts and transport every fixed prefix event through an IdentDistrib stream law.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html","parent":"chapter:ucb","order":215,"meta":[["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean"],["Declarations","8"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbfixedcountpeeling banditrlproof.algorithms.ucbfixedcountpeeling this module isolates the law transport used by the pinned lml ucb proof. a fixedarmprefixsource records the pathwise fact that rewards selected from one arm are the prefix of an arm-indexed reward stream. the main theorems peel the random pull count into finitely many fixed counts and transport every fixed prefix event through an identdistrib stream law. lean module compiled","shard":"modules/e5071123e12c7774.json"},{"id":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","label":"Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","description":"This module defines one finite-arm UCB policy. Its score at history index t uses the summable confidence share telescopingConfidenceShare delta t / K; no terminal horizon occurs in the policy, state, score, or generated-action declarations. Finite horizons occur only in downstream count and regret consumers.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html","parent":"chapter:ucb","order":216,"meta":[["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean"],["Declarations","43"],["Project imports","3"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbfixedpolicytelescopinganytimeregret banditrlproof.algorithms.ucbfixedpolicytelescopinganytimeregret this module defines one finite-arm ucb policy. its score at history index t uses the summable confidence share telescopingconfidenceshare delta t / k; no terminal horizon occurs in the policy, state, score, or generated-action declarations. finite horizons occur only in downstream count and regret consumers. lean module compiled","shard":"modules/0fa477c4629fa244.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","label":"Algorithms.UCBRealHistoryIndex","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealHistoryIndex","description":"This module mirrors the path-dependent score used by the pinned LML UCB route. In particular, the confidence width divides by the realized pull count. It is therefore distinct from the earlier UCB surface whose proxy is deterministic in the sample point.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html","parent":"chapter:ucb","order":217,"meta":[["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean"],["Declarations","22"],["Project imports","4"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbrealhistoryindex banditrlproof.algorithms.ucbrealhistoryindex this module mirrors the path-dependent score used by the pinned lml ucb route. in particular, the confidence width divides by the realized pull count. it is therefore distinct from the earlier ucb surface whose proxy is deterministic in the sample point. lean module compiled","shard":"modules/1bf735e22a8440f6.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","label":"Algorithms.UCBRealLMLCompat","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealLMLCompat","description":"The pinned LML source currently uses a newer Lean/mathlib toolchain, so ABRL cannot import its IsAlgEnvSeq declaration directly. This module packages the exact measurable action/feedback and split conditional-law consequences used by the local UCB trajectory and regret route. It is a local compatibility structure, not an imported LML proof.","url":"../modules/banditrlproof-algorithms-ucbreallmlcompat/index.html","parent":"chapter:ucb","order":218,"meta":[["Source","BanditRLProof/Algorithms/UCBRealLMLCompat.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbreallmlcompat banditrlproof.algorithms.ucbreallmlcompat the pinned lml source currently uses a newer lean/mathlib toolchain, so abrl cannot import its isalgenvseq declaration directly. this module packages the exact measurable action/feedback and split conditional-law consequences used by the local ucb trajectory and regret route. it is a local compatibility structure, not an imported lml proof. lean module compiled","shard":"modules/56a897e873d445d1.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","label":"Algorithms.UCBRealStationaryCanonicalKernelTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","description":"This module packages the canonical arm-stream UCB split conditional laws as a history algorithm and environment, then independently regenerates their observable action/reward pair process with Mathlib's Ionescu-Tulcea Kernel.trajMeasure. The resulting coordinate process supplies every field of RealStationaryUCBSequence without a caller-provided sample space or law.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html","parent":"chapter:ucb","order":219,"meta":[["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbrealstationarycanonicalkerneltrajectory banditrlproof.algorithms.ucbrealstationarycanonicalkerneltrajectory this module packages the canonical arm-stream ucb split conditional laws as a history algorithm and environment, then independently regenerates their observable action/reward pair process with mathlib's ionescu-tulcea kernel.trajmeasure. the resulting coordinate process supplies every field of realstationaryucbsequence without a caller-provided sample space or law. lean module compiled","shard":"modules/c6a3921acc084883.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","label":"Algorithms.UCBRealStationaryExplicitPolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","description":"This module identifies the canonical arm-stream successor action kernel with the deterministic realHistoryNextArm kernel. It then transports the complete explicit-policy graph to the independently generated canonical pair trajectory and pairs that graph with the existing expected-average consistency theorem.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html","parent":"chapter:ucb","order":220,"meta":[["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbrealstationaryexplicitpolicy banditrlproof.algorithms.ucbrealstationaryexplicitpolicy this module identifies the canonical arm-stream successor action kernel with the deterministic realhistorynextarm kernel. it then transports the complete explicit-policy graph to the independently generated canonical pair trajectory and pairs that graph with the existing expected-average consistency theorem. lean module compiled","shard":"modules/09bfaa40c7f3aaa0.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","label":"Algorithms.UCBRealStationaryFiniteArmRewardLaws","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","description":"This module transports the canonical one-policy arm-stream asymptotics through the complete observable law supplied by RealStationaryUCBSequence. The final endpoint instantiates the route from armwise-bounded Real reward laws.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html","parent":"chapter:ucb","order":221,"meta":[["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbrealstationaryfinitearmrewardlaws banditrlproof.algorithms.ucbrealstationaryfinitearmrewardlaws this module transports the canonical one-policy arm-stream asymptotics through the complete observable law supplied by realstationaryucbsequence. the final endpoint instantiates the route from armwise-bounded real reward laws. lean module compiled","shard":"modules/9271d25e3a91d1bd.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","label":"Algorithms.UCBRealStationaryMeasurePreservingSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","description":"This module constructs RealStationaryUCBSequence by pulling the canonical arm-stream process through a measure-preserving source map. A product-space specialization permits arbitrary independent nuisance randomness and closes the armwise-bounded expected-average consistency route without caller-supplied split conditional laws.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html","parent":"chapter:ucb","order":222,"meta":[["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbrealstationarymeasurepreservingsource banditrlproof.algorithms.ucbrealstationarymeasurepreservingsource this module constructs realstationaryucbsequence by pulling the canonical arm-stream process through a measure-preserving source map. a product-space specialization permits arbitrary independent nuisance randomness and closes the armwise-bounded expected-average consistency route without caller-supplied split conditional laws. lean module compiled","shard":"modules/8383a63eb5f3bdc1.json"},{"id":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","label":"Algorithms.UCBRealStationarySelectedRewardConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","description":"This module transports the stationary selected-reward laws from the latent arm-stream process to the independently regenerated canonical kernel trajectory, then pairs those laws with the compiled explicit-policy expected-average result.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryselectedrewardconsistency/index.html","parent":"chapter:ucb","order":223,"meta":[["Source","BanditRLProof/Algorithms/UCBRealStationarySelectedRewardConsistency.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"algorithms.ucbrealstationaryselectedrewardconsistency banditrlproof.algorithms.ucbrealstationaryselectedrewardconsistency this module transports the stationary selected-reward laws from the latent arm-stream process to the independently regenerated canonical kernel trajectory, then pairs those laws with the compiled explicit-policy expected-average result. lean module compiled","shard":"modules/bc73ac9a07423c53.json"},{"id":"module:BanditRLProof.Automation","label":"Automation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Automation","description":"This file makes the harness roles and artifacts part of the compiled Lean project. It does not run agents; it records the protocol that external agents must satisfy.","url":"../modules/banditrlproof-automation/index.html","parent":"chapter:frontier","order":224,"meta":[["Source","BanditRLProof/Automation.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"automation banditrlproof.automation this file makes the harness roles and artifacts part of the compiled lean project. it does not run agents; it records the protocol that external agents must satisfy. lean module compiled","shard":"modules/b9c9548157412714.json"},{"id":"module:BanditRLProof.BoundedRewardKernelLaw","label":"BoundedRewardKernelLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.BoundedRewardKernelLaw","description":"This module constructs the one-step centered reward-kernel contract directly from pointwise MGF witnesses or common almost-sure bounds. It is independent of any bandit algorithm or trajectory construction.","url":"../modules/banditrlproof-boundedrewardkernellaw/index.html","parent":"chapter:probability","order":225,"meta":[["Source","BanditRLProof/BoundedRewardKernelLaw.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"boundedrewardkernellaw banditrlproof.boundedrewardkernellaw this module constructs the one-step centered reward-kernel contract directly from pointwise mgf witnesses or common almost-sure bounds. it is independent of any bandit algorithm or trajectory construction. lean module compiled","shard":"modules/6a357237165f37b3.json"},{"id":"module:BanditRLProof.BudgetStoppingTime","label":"BudgetStoppingTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.BudgetStoppingTime","description":"This module exposes a narrow Mathlib-backed stopping-time surface for resource/budget processes. It deliberately stays at the filtration foundation layer: no knapsack model, policy construction, optional stopping theorem, or regret theorem is introduced here.","url":"../modules/banditrlproof-budgetstoppingtime/index.html","parent":"chapter:frontier","order":226,"meta":[["Source","BanditRLProof/BudgetStoppingTime.lean"],["Declarations","3"],["Project imports","0"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"budgetstoppingtime banditrlproof.budgetstoppingtime this module exposes a narrow mathlib-backed stopping-time surface for resource/budget processes. it deliberately stays at the filtration foundation layer: no knapsack model, policy construction, optional stopping theorem, or regret theorem is introduced here. lean module compiled","shard":"modules/9e24be65f9e8509f.json"},{"id":"module:BanditRLProof.ConcentrationCappedOccupancy","label":"ConcentrationCappedOccupancy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationCappedOccupancy","description":"Generated source map for BanditRLProof/ConcentrationCappedOccupancy.lean.","url":"../modules/banditrlproof-concentrationcappedoccupancy/index.html","parent":"chapter:probability","order":227,"meta":[["Source","BanditRLProof/ConcentrationCappedOccupancy.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationcappedoccupancy banditrlproof.concentrationcappedoccupancy generated source map for banditrlproof/concentrationcappedoccupancy.lean. lean module compiled","shard":"modules/06349c79cd6a2f4d.json"},{"id":"module:BanditRLProof.ConcentrationConditionalMGF","label":"ConcentrationConditionalMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationConditionalMGF","description":"Constructors connecting conditional expectation bounds to the shared fixed-tilt conditional MGF interface.","url":"../modules/banditrlproof-concentrationconditionalmgf/index.html","parent":"chapter:probability","order":228,"meta":[["Source","BanditRLProof/ConcentrationConditionalMGF.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationconditionalmgf banditrlproof.concentrationconditionalmgf constructors connecting conditional expectation bounds to the shared fixed-tilt conditional mgf interface. lean module compiled","shard":"modules/77b5898376537822.json"},{"id":"module:BanditRLProof.ConcentrationConfidenceSchedule","label":"ConcentrationConfidenceSchedule","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationConfidenceSchedule","description":"This module records reusable deterministic schedules for countable confidence budgets. It contains no stochastic-process or algorithm assumptions.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html","parent":"chapter:probability","order":229,"meta":[["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean"],["Declarations","12"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationconfidenceschedule banditrlproof.concentrationconfidenceschedule this module records reusable deterministic schedules for countable confidence budgets. it contains no stochastic-process or algorithm assumptions. lean module compiled","shard":"modules/2f4ee0c9ee5b49e6.json"},{"id":"module:BanditRLProof.ConcentrationDyadicExponential","label":"ConcentrationDyadicExponential","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationDyadicExponential","description":"A convex-exponential comparison gives a dyadic sum estimate strong enough to imply the constant 15 used in source Lemma 9.3.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html","parent":"chapter:probability","order":230,"meta":[["Source","BanditRLProof/ConcentrationDyadicExponential.lean"],["Declarations","6"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationdyadicexponential banditrlproof.concentrationdyadicexponential a convex-exponential comparison gives a dyadic sum estimate strong enough to imply the constant 15 used in source lemma 9.3. lean module compiled","shard":"modules/0ed469de9bc35975.json"},{"id":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","label":"ConcentrationFintypeGeometricAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationFintypeGeometricAllTime","description":"This module composes two already-compiled outer-measure adapters: a finite equal-share union at each time and a countable geometric schedule over time. It is deliberately only a confidence-budget composition theorem. It does not produce the per-index tails and is not a Ville/Doob, mixture, optional-stopping, self-normalized, or general Freedman inequality.","url":"../modules/banditrlproof-concentrationfintypegeometricalltime/index.html","parent":"chapter:probability","order":231,"meta":[["Source","BanditRLProof/ConcentrationFintypeGeometricAllTime.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationfintypegeometricalltime banditrlproof.concentrationfintypegeometricalltime this module composes two already-compiled outer-measure adapters: a finite equal-share union at each time and a countable geometric schedule over time. it is deliberately only a confidence-budget composition theorem. it does not produce the per-index tails and is not a ville/doob, mixture, optional-stopping, self-normalized, or general freedman inequality. lean module compiled","shard":"modules/a86ff861bccc9795.json"},{"id":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","label":"ConcentrationFintypeTelescopingAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationFintypeTelescopingAllTime","description":"This module specializes the reusable finite-index/countable-time outer-measure union bound to the confidence schedule delta / ((n+1)(n+2)). Its reciprocal grows polynomially, so logarithmic confidence radii retain logarithmic time growth. The theorem only composes supplied event bounds; it is not a stochastic-process or UCB result.","url":"../modules/banditrlproof-concentrationfintypetelescopingalltime/index.html","parent":"chapter:probability","order":232,"meta":[["Source","BanditRLProof/ConcentrationFintypeTelescopingAllTime.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationfintypetelescopingalltime banditrlproof.concentrationfintypetelescopingalltime this module specializes the reusable finite-index/countable-time outer-measure union bound to the confidence schedule delta / ((n+1)(n+2)). its reciprocal grows polynomially, so logarithmic confidence radii retain logarithmic time growth. the theorem only composes supplied event bounds; it is not a stochastic-process or ucb result. lean module compiled","shard":"modules/cde39c638d622b2e.json"},{"id":"module:BanditRLProof.ConcentrationFixedMGF","label":"ConcentrationFixedMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationFixedMGF","description":"This module isolates the part of a martingale Bernstein/Freedman route that does not depend on the particular one-step exponential inequality. Unlike HasSubgaussianMGF, the upper bound is required at one fixed tilt only. Exponential integrability at every real multiple is retained because it is the regularity needed to compose kernel laws.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html","parent":"chapter:probability","order":233,"meta":[["Source","BanditRLProof/ConcentrationFixedMGF.lean"],["Declarations","24"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationfixedmgf banditrlproof.concentrationfixedmgf this module isolates the part of a martingale bernstein/freedman route that does not depend on the particular one-step exponential inequality. unlike hassubgaussianmgf, the upper bound is required at one fixed tilt only. exponential integrability at every real multiple is retained because it is the regularity needed to compose kernel laws. lean module compiled","shard":"modules/8e19fc3b9fb87664.json"},{"id":"module:BanditRLProof.ConcentrationGaussianOccupancy","label":"ConcentrationGaussianOccupancy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationGaussianOccupancy","description":"Generated source map for BanditRLProof/ConcentrationGaussianOccupancy.lean.","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html","parent":"chapter:probability","order":234,"meta":[["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean"],["Declarations","12"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationgaussianoccupancy banditrlproof.concentrationgaussianoccupancy generated source map for banditrlproof/concentrationgaussianoccupancy.lean. lean module compiled","shard":"modules/68489066b720b681.json"},{"id":"module:BanditRLProof.ConcentrationIndexOccupancy","label":"ConcentrationIndexOccupancy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationIndexOccupancy","description":"Generated source map for BanditRLProof/ConcentrationIndexOccupancy.lean.","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html","parent":"chapter:probability","order":235,"meta":[["Source","BanditRLProof/ConcentrationIndexOccupancy.lean"],["Declarations","8"],["Project imports","3"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationindexoccupancy banditrlproof.concentrationindexoccupancy generated source map for banditrlproof/concentrationindexoccupancy.lean. lean module compiled","shard":"modules/1c8fe1693a2e0a2c.json"},{"id":"module:BanditRLProof.ConcentrationMartingaleMaximal","label":"ConcentrationMartingaleMaximal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationMartingaleMaximal","description":"This module supplies the martingale analytic dependency for source Theorem 9.2. Identifying an independent reward partial sum as this martingale is separate.","url":"../modules/banditrlproof-concentrationmartingalemaximal/index.html","parent":"chapter:probability","order":236,"meta":[["Source","BanditRLProof/ConcentrationMartingaleMaximal.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationmartingalemaximal banditrlproof.concentrationmartingalemaximal this module supplies the martingale analytic dependency for source theorem 9.2. identifying an independent reward partial sum as this martingale is separate. lean module compiled","shard":"modules/0edaacebcfcfa795.json"},{"id":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","label":"ConcentrationQuadraticFixedMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationQuadraticFixedMGF","description":"This module turns a family of fixed-tilt quadratic exponential tails into a delta-shaped bound. The probabilistic construction of each fixed-tilt tail remains separate, so model-specific consumers only need to expose the common quadratic exponent and admissible tilt cap.","url":"../modules/banditrlproof-concentrationquadraticfixedmgf/index.html","parent":"chapter:probability","order":237,"meta":[["Source","BanditRLProof/ConcentrationQuadraticFixedMGF.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationquadraticfixedmgf banditrlproof.concentrationquadraticfixedmgf this module turns a family of fixed-tilt quadratic exponential tails into a delta-shaped bound. the probabilistic construction of each fixed-tilt tail remains separate, so model-specific consumers only need to expose the common quadratic exponent and admissible tilt cap. lean module compiled","shard":"modules/d7a6971ac1bd292e.json"},{"id":"module:BanditRLProof.ConcentrationQuadraticMaximal","label":"ConcentrationQuadraticMaximal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationQuadraticMaximal","description":"This module adds a finite-index maximal surface to the quadratic fixed-MGF route. It uses equal confidence shares and the finite outer-measure union bound; it is not a Ville, Doob, or infinite-horizon maximal inequality.","url":"../modules/banditrlproof-concentrationquadraticmaximal/index.html","parent":"chapter:probability","order":238,"meta":[["Source","BanditRLProof/ConcentrationQuadraticMaximal.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationquadraticmaximal banditrlproof.concentrationquadraticmaximal this module adds a finite-index maximal surface to the quadratic fixed-mgf route. it uses equal confidence shares and the finite outer-measure union bound; it is not a ville, doob, or infinite-horizon maximal inequality. lean module compiled","shard":"modules/24ac619ce203f767.json"},{"id":"module:BanditRLProof.ConcentrationQuadraticScheduled","label":"ConcentrationQuadraticScheduled","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationQuadraticScheduled","description":"This module lifts the one-event quadratic fixed-MGF delta theorem to a countable family with time-varying parameters and confidence shares. The result uses countable outer-measure subadditivity; it is not a Ville, Doob, mixture, optional-stopping, or general Freedman theorem.","url":"../modules/banditrlproof-concentrationquadraticscheduled/index.html","parent":"chapter:probability","order":239,"meta":[["Source","BanditRLProof/ConcentrationQuadraticScheduled.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationquadraticscheduled banditrlproof.concentrationquadraticscheduled this module lifts the one-event quadratic fixed-mgf delta theorem to a countable family with time-varying parameters and confidence shares. the result uses countable outer-measure subadditivity; it is not a ville, doob, mixture, optional-stopping, or general freedman theorem. lean module compiled","shard":"modules/010cd7d40ab36a07.json"},{"id":"module:BanditRLProof.ConcentrationSubGaussian","label":"ConcentrationSubGaussian","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationSubGaussian","description":"This module exposes small Mathlib-backed concentration imports under the project namespace. It deliberately stays at the reusable concentration layer: no ETC reward model, empirical-mean construction, or final regret result is introduced here.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html","parent":"chapter:probability","order":240,"meta":[["Source","BanditRLProof/ConcentrationSubGaussian.lean"],["Declarations","31"],["Project imports","3"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationsubgaussian banditrlproof.concentrationsubgaussian this module exposes small mathlib-backed concentration imports under the project namespace. it deliberately stays at the reusable concentration layer: no etc reward model, empirical-mean construction, or final regret result is introduced here. lean module compiled","shard":"modules/86d9c3e7f592c6de.json"},{"id":"module:BanditRLProof.ConcentrationTailIntegration","label":"ConcentrationTailIntegration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationTailIntegration","description":"Generated source map for BanditRLProof/ConcentrationTailIntegration.lean.","url":"../modules/banditrlproof-concentrationtailintegration/index.html","parent":"chapter:probability","order":241,"meta":[["Source","BanditRLProof/ConcentrationTailIntegration.lean"],["Declarations","1"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationtailintegration banditrlproof.concentrationtailintegration generated source map for banditrlproof/concentrationtailintegration.lean. lean module compiled","shard":"modules/ffaf3f81fc32fd9d.json"},{"id":"module:BanditRLProof.ConcentrationVariance","label":"ConcentrationVariance","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConcentrationVariance","description":"This module exposes small Mathlib-backed variance/Chebyshev imports under the project namespace. It is only the reusable finite-variance tail layer: no robust mean estimator, bandit reward law, or final regret theorem is introduced here.","url":"../modules/banditrlproof-concentrationvariance/index.html","parent":"chapter:probability","order":242,"meta":[["Source","BanditRLProof/ConcentrationVariance.lean"],["Declarations","3"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"concentrationvariance banditrlproof.concentrationvariance this module exposes small mathlib-backed variance/chebyshev imports under the project namespace. it is only the reusable finite-variance tail layer: no robust mean estimator, bandit reward law, or final regret theorem is introduced here. lean module compiled","shard":"modules/29402fc04ca866ee.json"},{"id":"module:BanditRLProof.ConditionalExpectationReward","label":"ConditionalExpectationReward","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward","description":"This module exposes a narrow COND-EXPECT-REWARD support leaf: if the conditional-expectation kernel already identifies the next centered reward law and its conditional integral is zero, then the ordinary conditional expectation of that centered reward is zero.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html","parent":"chapter:probability","order":243,"meta":[["Source","BanditRLProof/ConditionalExpectationReward.lean"],["Declarations","89"],["Project imports","3"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalexpectationreward banditrlproof.conditionalexpectationreward this module exposes a narrow cond-expect-reward support leaf: if the conditional-expectation kernel already identifies the next centered reward law and its conditional integral is zero, then the ordinary conditional expectation of that centered reward is zero. lean module compiled","shard":"modules/c84d405641caee32.json"},{"id":"module:BanditRLProof.ConditionalRewardFoundation","label":"ConditionalRewardFoundation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalRewardFoundation","description":"This module closes the canonical COND-EXPECT-REWARD route for reward-only trajMeasure processes. A centered reward-kernel law and deterministic historywise proxy ceilings yield, on the generated history filtration:","url":"../modules/banditrlproof-conditionalrewardfoundation/index.html","parent":"chapter:probability","order":244,"meta":[["Source","BanditRLProof/ConditionalRewardFoundation.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalrewardfoundation banditrlproof.conditionalrewardfoundation this module closes the canonical cond-expect-reward route for reward-only trajmeasure processes. a centered reward-kernel law and deterministic historywise proxy ceilings yield, on the generated history filtration: lean module compiled","shard":"modules/cac8bcf91b5e47d4.json"},{"id":"module:BanditRLProof.ConditionalRewardLawSource","label":"ConditionalRewardLawSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalRewardLawSource","description":"This module records a narrow COND-EXPECT-REWARD support leaf. It packages the remaining generated-policy conditional next-pair law assumption as a reusable contract, then consumes the existing ConditionalExpectationReward route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html","parent":"chapter:probability","order":245,"meta":[["Source","BanditRLProof/ConditionalRewardLawSource.lean"],["Declarations","347"],["Project imports","3"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalrewardlawsource banditrlproof.conditionalrewardlawsource this module records a narrow cond-expect-reward support leaf. it packages the remaining generated-policy conditional next-pair law assumption as a reusable contract, then consumes the existing conditionalexpectationreward route. lean module compiled","shard":"modules/3bcc27636785d725.json"},{"id":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","label":"ConditionalRewardPartialTrajectoryGeometricAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","description":"This module instantiates the canonical action/reward trajectory's fixed-arm, fixed-successor-horizon random-pull-count empirical-mean tail at one geometric confidence share per time and arm. The compiled finite-index geometric all-time adapter then controls the union over every positive successor horizon and every arm in a nonempty finite action type.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorygeometricalltime/index.html","parent":"chapter:probability","order":246,"meta":[["Source","BanditRLProof/ConditionalRewardPartialTrajectoryGeometricAllTime.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalrewardpartialtrajectorygeometricalltime banditrlproof.conditionalrewardpartialtrajectorygeometricalltime this module instantiates the canonical action/reward trajectory's fixed-arm, fixed-successor-horizon random-pull-count empirical-mean tail at one geometric confidence share per time and arm. the compiled finite-index geometric all-time adapter then controls the union over every positive successor horizon and every arm in a nonempty finite action type. lean module compiled","shard":"modules/d795266350ff4336.json"},{"id":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","label":"ConditionalRewardPartialTrajectoryLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalRewardPartialTrajectoryLaw","description":"Generated source map for BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html","parent":"chapter:probability","order":247,"meta":[["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalrewardpartialtrajectorylaw banditrlproof.conditionalrewardpartialtrajectorylaw generated source map for banditrlproof/conditionalrewardpartialtrajectorylaw.lean. lean module compiled","shard":"modules/60d0c5c2fec7b756.json"},{"id":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","label":"ConditionalRewardPartialTrajectoryMaskedLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","description":"This module specializes the canonical action/reward trajectory law to one history-policy-selected arm and the generic predictable-variance concentration interface. It identifies the policy mask with the sampled successor action almost everywhere on the canonical trajectory and rewrites the masked constant proxy as an actual successor pull count. It also normalizes positive exact count fibers into empirical means, pe…","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html","parent":"chapter:probability","order":248,"meta":[["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalrewardpartialtrajectorymaskedlaw banditrlproof.conditionalrewardpartialtrajectorymaskedlaw this module specializes the canonical action/reward trajectory law to one history-policy-selected arm and the generic predictable-variance concentration interface. it identifies the policy mask with the sampled successor action almost everywhere on the canonical trajectory and rewrites the masked constant proxy as an actual successor pull count. it also normalizes positive exact count fibers into empirical means, performs fixed-horizon count peeling, and closes the finite arm/time union on the canonical trajectory. it does not prove uniform-time confidence, ucb, or regret. lean module compiled","shard":"modules/1d074e4c5894e81f.json"},{"id":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","label":"ConditionalRewardPartialTrajectoryTelescopingAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","description":"This module instantiates the canonical action/reward trajectory's fixed-arm, fixed-successor-horizon random-pull-count empirical-mean tail at the telescoping confidence share delta / ((n+1)(n+2)) per time, divided equally across arms. The resulting countable event has outer measure at most delta on one generated trajectory measure.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorytelescopingalltime/index.html","parent":"chapter:probability","order":249,"meta":[["Source","BanditRLProof/ConditionalRewardPartialTrajectoryTelescopingAllTime.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"conditionalrewardpartialtrajectorytelescopingalltime banditrlproof.conditionalrewardpartialtrajectorytelescopingalltime this module instantiates the canonical action/reward trajectory's fixed-arm, fixed-successor-horizon random-pull-count empirical-mean tail at the telescoping confidence share delta / ((n+1)(n+2)) per time, divided equally across arms. the resulting countable event has outer measure at most delta on one generated trajectory measure. lean module compiled","shard":"modules/3386baaf829e7267.json"},{"id":"module:BanditRLProof.Core","label":"Core","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Core","description":"This file keeps the first project layer dependency-light. It provides a small executable language for finite action traces, pull counts, reward sums, and finite-arm mean models. Strong probabilistic statements can later import Mathlib or external libraries without changing this public surface.","url":"../modules/banditrlproof-core/index.html","parent":"chapter:foundations","order":250,"meta":[["Source","BanditRLProof/Core.lean"],["Declarations","15"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"core banditrlproof.core this file keeps the first project layer dependency-light. it provides a small executable language for finite action traces, pull counts, reward sums, and finite-arm mean models. strong probabilistic statements can later import mathlib or external libraries without changing this public surface. lean module compiled","shard":"modules/384912856a4c897c.json"},{"id":"module:BanditRLProof.CurvatureNoiseGapGeometry","label":"CurvatureNoiseGapGeometry","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGapGeometry","description":"This file contains only route-independent finite-dimensional algebra. It does not formalize the full Curvature--Noise--Gap calculus, Tsallis-INF, or a new bandit theorem. The intended later use is to test whether these abstractions replace repeated route-specific proof subgraphs and transfer to held-out proof families.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html","parent":"chapter:foundations","order":251,"meta":[["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean"],["Declarations","12"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"curvaturenoisegapgeometry banditrlproof.curvaturenoisegapgeometry this file contains only route-independent finite-dimensional algebra. it does not formalize the full curvature--noise--gap calculus, tsallis-inf, or a new bandit theorem. the intended later use is to test whether these abstractions replace repeated route-specific proof subgraphs and transfer to held-out proof families. lean module compiled","shard":"modules/1762b537fe4635c9.json"},{"id":"module:BanditRLProof.DelayedFeedback.Accounting","label":"DelayedFeedback.Accounting","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.Accounting","description":"Generated source map for BanditRLProof/DelayedFeedback/Accounting.lean.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html","parent":"chapter:frontier","order":252,"meta":[["Source","BanditRLProof/DelayedFeedback/Accounting.lean"],["Declarations","17"],["Project imports","0"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.accounting banditrlproof.delayedfeedback.accounting generated source map for banditrlproof/delayedfeedback/accounting.lean. lean module compiled","shard":"modules/8ce75c939e50ce73.json"},{"id":"module:BanditRLProof.DelayedFeedback.ActionLaw","label":"DelayedFeedback.ActionLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.ActionLaw","description":"Generated source map for BanditRLProof/DelayedFeedback/ActionLaw.lean.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html","parent":"chapter:frontier","order":253,"meta":[["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.actionlaw banditrlproof.delayedfeedback.actionlaw generated source map for banditrlproof/delayedfeedback/actionlaw.lean. lean module compiled","shard":"modules/96574241bd97d88f.json"},{"id":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","label":"DelayedFeedback.ActiveAllocation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.ActiveAllocation","description":"Generated source map for BanditRLProof/DelayedFeedback/ActiveAllocation.lean.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html","parent":"chapter:frontier","order":254,"meta":[["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.activeallocation banditrlproof.delayedfeedback.activeallocation generated source map for banditrlproof/delayedfeedback/activeallocation.lean. lean module compiled","shard":"modules/f0f22e9a6d0b345e.json"},{"id":"module:BanditRLProof.DelayedFeedback.CausalView","label":"DelayedFeedback.CausalView","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.CausalView","description":"Generated source map for BanditRLProof/DelayedFeedback/CausalView.lean.","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html","parent":"chapter:frontier","order":255,"meta":[["Source","BanditRLProof/DelayedFeedback/CausalView.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.causalview banditrlproof.delayedfeedback.causalview generated source map for banditrlproof/delayedfeedback/causalview.lean. lean module compiled","shard":"modules/e941e09c0a7ab1e9.json"},{"id":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","label":"DelayedFeedback.EliminatedArmInitialization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.EliminatedArmInitialization","description":"This module records the numerical state created for each arm in the line-7 elimination set after line 8 of the frozen Delayed-SAPO algorithm. It follows physical PDF page 22, Algorithm 5 lines 9--10: the elimination round and processed prefix are frozen, the initial inactive-arm probability is formed from the processed pull count, the surrogate gap is eight empirical widths, and the first EAP phase is initialized.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html","parent":"chapter:frontier","order":256,"meta":[["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean"],["Declarations","31"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.eliminatedarminitialization banditrlproof.delayedfeedback.eliminatedarminitialization this module records the numerical state created for each arm in the line-7 elimination set after line 8 of the frozen delayed-sapo algorithm. it follows physical pdf page 22, algorithm 5 lines 9--10: the elimination round and processed prefix are frozen, the initial inactive-arm probability is formed from the processed pull count, the surrogate gap is eight empirical widths, and the first eap phase is initialized. lean module compiled","shard":"modules/93a392d1ccd0bdca.json"},{"id":"module:BanditRLProof.DelayedFeedback.Elimination","label":"DelayedFeedback.Elimination","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.Elimination","description":"Generated source map for BanditRLProof/DelayedFeedback/Elimination.lean.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html","parent":"chapter:frontier","order":257,"meta":[["Source","BanditRLProof/DelayedFeedback/Elimination.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.elimination banditrlproof.delayedfeedback.elimination generated source map for banditrlproof/delayedfeedback/elimination.lean. lean module compiled","shard":"modules/ee44238470f31349.json"},{"id":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","label":"DelayedFeedback.MultiRegimeContract","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.MultiRegimeContract","description":"Generated source map for BanditRLProof/DelayedFeedback/MultiRegimeContract.lean.","url":"../modules/banditrlproof-delayedfeedback-multiregimecontract/index.html","parent":"chapter:frontier","order":258,"meta":[["Source","BanditRLProof/DelayedFeedback/MultiRegimeContract.lean"],["Declarations","5"],["Project imports","0"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.multiregimecontract banditrlproof.delayedfeedback.multiregimecontract generated source map for banditrlproof/delayedfeedback/multiregimecontract.lean. lean module compiled","shard":"modules/a0c7dad176accccb.json"},{"id":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","label":"DelayedFeedback.OrderedNoSwitchTrace","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","description":"This module composes the source-faithful one-item transition from OrderedProcessingTransition.lean across an arbitrary finite trace. A trace may either process one member of B(t) \\ S through the no-switch structural projection of Algorithm 5 lines 3--4 and 7--8, or close an exhausted inner loop and advance to the next action round.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html","parent":"chapter:frontier","order":259,"meta":[["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.orderednoswitchtrace banditrlproof.delayedfeedback.orderednoswitchtrace this module composes the source-faithful one-item transition from orderedprocessingtransition.lean across an arbitrary finite trace. a trace may either process one member of b(t) \\ s through the no-switch structural projection of algorithm 5 lines 3--4 and 7--8, or close an exhausted inner loop and advance to the next action round. lean module compiled","shard":"modules/e439675c7b7680f1.json"},{"id":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","label":"DelayedFeedback.OrderedProcessingTransition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.OrderedProcessingTransition","description":"This module formalizes the deterministic structural content of one iteration of Algorithm 5 lines 3--4 and 7--8. A newly observed source round is appended to the paper sequence before the line-7 confidence snapshot is formed. The line-8 successor then removes exactly the arms selected by that snapshot.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html","parent":"chapter:frontier","order":260,"meta":[["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.orderedprocessingtransition banditrlproof.delayedfeedback.orderedprocessingtransition this module formalizes the deterministic structural content of one iteration of algorithm 5 lines 3--4 and 7--8. a newly observed source round is appended to the paper sequence before the line-7 confidence snapshot is formed. the line-8 successor then removes exactly the arms selected by that snapshot. lean module compiled","shard":"modules/df0e9f28799a2ec2.json"},{"id":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","label":"DelayedFeedback.ProcessedPrefixCounts","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","description":"Generated source map for BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html","parent":"chapter:frontier","order":261,"meta":[["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.processedprefixcounts banditrlproof.delayedfeedback.processedprefixcounts generated source map for banditrlproof/delayedfeedback/processedprefixcounts.lean. lean module compiled","shard":"modules/02feff4ea12028f0.json"},{"id":"module:BanditRLProof.DelayedFeedback.Processing","label":"DelayedFeedback.Processing","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.Processing","description":"Generated source map for BanditRLProof/DelayedFeedback/Processing.lean.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html","parent":"chapter:frontier","order":262,"meta":[["Source","BanditRLProof/DelayedFeedback/Processing.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.processing banditrlproof.delayedfeedback.processing generated source map for banditrlproof/delayedfeedback/processing.lean. lean module compiled","shard":"modules/be65cd8268927f5e.json"},{"id":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","label":"DelayedFeedback.RecursiveProcessedState","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.RecursiveProcessedState","description":"This module makes one deterministic interface preceding the probabilistic count clause in source Definition D.1 / Lemma D.4 explicit. A processed trace summary stores the source round of every processed item and reads the allocation that was present at that source round. It never reconstructs a source-time probability from the later processing-time state. Source indices are unique and carry the strict availability w…","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html","parent":"chapter:frontier","order":263,"meta":[["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.recursiveprocessedstate banditrlproof.delayedfeedback.recursiveprocessedstate this module makes one deterministic interface preceding the probabilistic count clause in source definition d.1 / lemma d.4 explicit. a processed trace summary stores the source round of every processed item and reads the allocation that was present at that source round. it never reconstructs a source-time probability from the later processing-time state. source indices are unique and carry the strict availability witness s + d_s < t used by algorithm 5. lean module compiled","shard":"modules/03ff1d65e572fce8.json"},{"id":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","label":"DelayedFeedback.StochasticGapHalfSet","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.StochasticGapHalfSet","description":"The frozen Delayed SAPO source uses a Markov-style count for stochastic loss gaps: at most half of the gaps are greater than 2 * mu. This formalization promotes exactly the nonnegative domain used by that application and handles the zero-average case separately.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html","parent":"chapter:frontier","order":264,"meta":[["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean"],["Declarations","6"],["Project imports","0"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.stochasticgaphalfset banditrlproof.delayedfeedback.stochasticgaphalfset the frozen delayed sapo source uses a markov-style count for stochastic loss gaps: at most half of the gaps are greater than 2 * mu. this formalization promotes exactly the nonnegative domain used by that application and handles the zero-average case separately. lean module compiled","shard":"modules/fc10501d3a23a119.json"},{"id":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","label":"DelayedFeedback.StochasticGapOrderingAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","description":"Generated source map for BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html","parent":"chapter:frontier","order":265,"meta":[["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.stochasticgaporderingaudit banditrlproof.delayedfeedback.stochasticgaporderingaudit generated source map for banditrlproof/delayedfeedback/stochasticgaporderingaudit.lean. lean module compiled","shard":"modules/a204c8cacf49e908.json"},{"id":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","label":"DelayedFeedback.StochasticGoodEvent","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.StochasticGoodEvent","description":"Generated source map for BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html","parent":"chapter:frontier","order":266,"meta":[["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.stochasticgoodevent banditrlproof.delayedfeedback.stochasticgoodevent generated source map for banditrlproof/delayedfeedback/stochasticgoodevent.lean. lean module compiled","shard":"modules/4a04fac997a67e65.json"},{"id":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","label":"DelayedFeedback.StochasticGoodEventAssembly","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","description":"Generated source map for BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html","parent":"chapter:frontier","order":267,"meta":[["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"delayedfeedback.stochasticgoodeventassembly banditrlproof.delayedfeedback.stochasticgoodeventassembly generated source map for banditrlproof/delayedfeedback/stochasticgoodeventassembly.lean. lean module compiled","shard":"modules/9ae235d6048c94ed.json"},{"id":"module:BanditRLProof.Exp3ActionProcess","label":"Exp3ActionProcess","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ActionProcess","description":"This module constructs a history-adaptive finite-action Markov kernel from measurable probability coordinates. Its canonical sample measure is the history law composed with that kernel, so the sampled action has the requested conditional distribution by Mathlib's condDistrib/compProd uniqueness theorem. The final wrappers discharge the law premises of the one-round EXP3 importance-weighted moment transport.","url":"../modules/banditrlproof-exp3actionprocess/index.html","parent":"chapter:exp3","order":268,"meta":[["Source","BanditRLProof/Exp3ActionProcess.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3actionprocess banditrlproof.exp3actionprocess this module constructs a history-adaptive finite-action markov kernel from measurable probability coordinates. its canonical sample measure is the history law composed with that kernel, so the sampled action has the requested conditional distribution by mathlib's conddistrib/compprod uniqueness theorem. the final wrappers discharge the law premises of the one-round exp3 importance-weighted moment transport. lean module compiled","shard":"modules/cf33b341e5d9121d.json"},{"id":"module:BanditRLProof.Exp3BernsteinAllHorizon","label":"Exp3BernsteinAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3BernsteinAllHorizon","description":"The explicit Bernstein schedule gives its 11 * gamma * T threshold when the three large-horizon inequalities make clipping inactive. This module closes the complementary branch honestly: generated realized losses are at most one almost surely, comparator predictable losses are nonnegative pointwise, and therefore realized regret is at most T almost surely. A branch threshold uses T + 1 outside the Bernstein regime,…","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html","parent":"chapter:exp3","order":269,"meta":[["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3bernsteinallhorizon banditrlproof.exp3bernsteinallhorizon the explicit bernstein schedule gives its 11 * gamma * t threshold when the three large-horizon inequalities make clipping inactive. this module closes the complementary branch honestly: generated realized losses are at most one almost surely, comparator predictable losses are nonnegative pointwise, and therefore realized regret is at most t almost surely. a branch threshold uses t + 1 outside the bernstein regime, yielding one theorem for every positive horizon without pretending that the clipped branch satisfies cubic dominance. lean module compiled","shard":"modules/5b5bbb53d9e8c6d3.json"},{"id":"module:BanditRLProof.Exp3BernsteinExplicitTuning","label":"Exp3BernsteinExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3BernsteinExplicitTuning","description":"This module discharges the three dominance premises of the characterized 11 * gamma * T theorem with an explicit maximum of two cube-root scales and one square-root scale. The schedule is clipped at 1 / 2; transparent large-horizon premises ensure that the clip is inactive.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html","parent":"chapter:exp3","order":270,"meta":[["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3bernsteinexplicittuning banditrlproof.exp3bernsteinexplicittuning this module discharges the three dominance premises of the characterized 11 * gamma * t theorem with an explicit maximum of two cube-root scales and one square-root scale. the schedule is clipped at 1 / 2; transparent large-horizon premises ensure that the clip is inactive. lean module compiled","shard":"modules/254fb3c4a0a0b351.json"},{"id":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","label":"Exp3BernsteinHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3BernsteinHighProbabilityRegret","description":"This module reassembles the generated predictable EXP3 regret theorem with the variance-sensitive pure-cross and fixed-comparator confidence radii. The random Hedge-square contribution retains its existing pathwise reciprocal-floor bound; no general Freedman or predictable-variance theorem is claimed.","url":"../modules/banditrlproof-exp3bernsteinhighprobabilityregret/index.html","parent":"chapter:exp3","order":271,"meta":[["Source","BanditRLProof/Exp3BernsteinHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3bernsteinhighprobabilityregret banditrlproof.exp3bernsteinhighprobabilityregret this module reassembles the generated predictable exp3 regret theorem with the variance-sensitive pure-cross and fixed-comparator confidence radii. the random hedge-square contribution retains its existing pathwise reciprocal-floor bound; no general freedman or predictable-variance theorem is claimed. lean module compiled","shard":"modules/c7ceb5537c8e4d02.json"},{"id":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","label":"Exp3BernsteinRealizedHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","description":"This module composes the generated predictable Bernstein-radius EXP3 theorem with the one-sided realized-minus-exploration deviation tail. The two importance-weighted confidence events use the variance-sensitive fixed-tilt route; the realized-deviation event retains its bounded-loss Hoeffding/Azuma radius. The deterministic Hedge-square contribution also remains unchanged.","url":"../modules/banditrlproof-exp3bernsteinrealizedhighprobabilityregret/index.html","parent":"chapter:exp3","order":272,"meta":[["Source","BanditRLProof/Exp3BernsteinRealizedHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3bernsteinrealizedhighprobabilityregret banditrlproof.exp3bernsteinrealizedhighprobabilityregret this module composes the generated predictable bernstein-radius exp3 theorem with the one-sided realized-minus-exploration deviation tail. the two importance-weighted confidence events use the variance-sensitive fixed-tilt route; the realized-deviation event retains its bounded-loss hoeffding/azuma radius. the deterministic hedge-square contribution also remains unchanged. lean module compiled","shard":"modules/51275df4f4321a21.json"},{"id":"module:BanditRLProof.Exp3BernsteinTuning","label":"Exp3BernsteinTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3BernsteinTuning","description":"The compiled high-probability theorem retains a pathwise estimator-square term. Consequently, the expected-regret choice eta = gamma / K leaves a linear term. This module instead balances the Hedge terms with eta = sqrt (log K * gamma / (T * K)) and records the cubic exploration conditions required by the current Bernstein confidence radii.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html","parent":"chapter:exp3","order":273,"meta":[["Source","BanditRLProof/Exp3BernsteinTuning.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3bernsteintuning banditrlproof.exp3bernsteintuning the compiled high-probability theorem retains a pathwise estimator-square term. consequently, the expected-regret choice eta = gamma / k leaves a linear term. this module instead balances the hedge terms with eta = sqrt (log k * gamma / (t * k)) and records the cubic exploration conditions required by the current bernstein confidence radii. lean module compiled","shard":"modules/232ae82a12ecc628.json"},{"id":"module:BanditRLProof.Exp3BestArm","label":"Exp3BestArm","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3BestArm","description":"This module packages the comparator-independent order step used by finite-arm best-in-hindsight regret wrappers. It contains no probability bound: downstream routes identify the best-arm event with a finite union of fixed-comparator events and supply their own confidence schedules.","url":"../modules/banditrlproof-exp3bestarm/index.html","parent":"chapter:exp3","order":274,"meta":[["Source","BanditRLProof/Exp3BestArm.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3bestarm banditrlproof.exp3bestarm this module packages the comparator-independent order step used by finite-arm best-in-hindsight regret wrappers. it contains no probability bound: downstream routes identify the best-arm event with a finite union of fixed-comparator events and supply their own confidence schedules. lean module compiled","shard":"modules/0d63c9d1d6b80df1.json"},{"id":"module:BanditRLProof.Exp3ComparatorBernstein","label":"Exp3ComparatorBernstein","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ComparatorBernstein","description":"This module replaces the range-squared Hoeffding proxy for one fixed comparator estimator by a fixed-tilt second-moment bound. The scalar input is the quadratic exponential remainder on [-1, 1]; the probabilistic input is the exact second moment of the importance-weighted estimator under its finite sampling law.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html","parent":"chapter:exp3","order":275,"meta":[["Source","BanditRLProof/Exp3ComparatorBernstein.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3comparatorbernstein banditrlproof.exp3comparatorbernstein this module replaces the range-squared hoeffding proxy for one fixed comparator estimator by a fixed-tilt second-moment bound. the scalar input is the quadratic exponential remainder on [-1, 1]; the probabilistic input is the exact second moment of the importance-weighted estimator under its finite sampling law. lean module compiled","shard":"modules/96ce3115ea0c36d6.json"},{"id":"module:BanditRLProof.Exp3ComparatorConfidence","label":"Exp3ComparatorConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ComparatorConfidence","description":"Generated source map for BanditRLProof/Exp3ComparatorConfidence.lean.","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html","parent":"chapter:exp3","order":276,"meta":[["Source","BanditRLProof/Exp3ComparatorConfidence.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3comparatorconfidence banditrlproof.exp3comparatorconfidence generated source map for banditrlproof/exp3comparatorconfidence.lean. lean module compiled","shard":"modules/1cfd6644c7b64780.json"},{"id":"module:BanditRLProof.Exp3ConditionalMoments","label":"Exp3ConditionalMoments","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ConditionalMoments","description":"This module turns the finite weighted-sum identities for the EXP3 estimator into Bochner-integral identities under an actual history-conditional action law. The conditional distribution is represented by a Markov kernel whose values are explicit finite sums of Dirac measures.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html","parent":"chapter:exp3","order":277,"meta":[["Source","BanditRLProof/Exp3ConditionalMoments.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3conditionalmoments banditrlproof.exp3conditionalmoments this module turns the finite weighted-sum identities for the exp3 estimator into bochner-integral identities under an actual history-conditional action law. the conditional distribution is represented by a markov kernel whose values are explicit finite sums of dirac measures. lean module compiled","shard":"modules/b9c971db79a12e84.json"},{"id":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","label":"Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","description":"For fixed model parameters, the four deterministic large-horizon inequalities used by the exact double-predictable-variance sparse EXP3 schedule eventually hold automatically. This module therefore removes the coarse T + 1 fallback eventually, identifies the best-arm threshold with 16 * gamma_T * T, and reuses the existing off-sparsity, residual, and practical tails under the same horizon-indexed generated trajector…","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html","parent":"chapter:exp3","order":278,"meta":[["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3doublevariancesparsebestarmeventualrefinedregret banditrlproof.exp3doublevariancesparsebestarmeventualrefinedregret for fixed model parameters, the four deterministic large-horizon inequalities used by the exact double-predictable-variance sparse exp3 schedule eventually hold automatically. this module therefore removes the coarse t + 1 fallback eventually, identifies the best-arm threshold with 16 * gamma_t * t, and reuses the existing off-sparsity, residual, and practical tails under the same horizon-indexed generated trajectory measures. lean module compiled","shard":"modules/4d795fab0daae3b8.json"},{"id":"module:BanditRLProof.Exp3ExpectedRegret","label":"Exp3ExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ExpectedRegret","description":"This module keeps parameter algebra separate from the generated trajectory and conditional-law proof. The first theorem simplifies the unoptimized budget when eta = gamma / |A|; the second applies that deterministic fact to the compiled predictable EXP3 endpoint.","url":"../modules/banditrlproof-exp3expectedregret/index.html","parent":"chapter:exp3","order":279,"meta":[["Source","BanditRLProof/Exp3ExpectedRegret.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3expectedregret banditrlproof.exp3expectedregret this module keeps parameter algebra separate from the generated trajectory and conditional-law proof. the first theorem simplifies the unoptimized budget when eta = gamma / |a|; the second applies that deterministic fact to the compiled predictable exp3 endpoint. lean module compiled","shard":"modules/5875be88110a5986.json"},{"id":"module:BanditRLProof.Exp3ExplorationBias","label":"Exp3ExplorationBias","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ExplorationBias","description":"This module compares the pure exponential-weights distribution q_t used by the Hedge potential with the exploration-mixed sampling distribution p_t = (1 - gamma) q_t + gamma / |arms|.","url":"../modules/banditrlproof-exp3explorationbias/index.html","parent":"chapter:exp3","order":280,"meta":[["Source","BanditRLProof/Exp3ExplorationBias.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3explorationbias banditrlproof.exp3explorationbias this module compares the pure exponential-weights distribution q_t used by the hedge potential with the exploration-mixed sampling distribution p_t = (1 - gamma) q_t + gamma / |arms|. lean module compiled","shard":"modules/45e8ae535c89af6c.json"},{"id":"module:BanditRLProof.Exp3HedgeRegret","label":"Exp3HedgeRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3HedgeRegret","description":"This module closes the full-information exponential-weights/Hedge potential argument that sits immediately below EXP3. It defines weights from cumulative losses, normalizes them on an explicit nonempty finite action set, proves the one-step logarithmic potential inequality, and telescopes it into second-order and [0,1] finite-horizon regret bounds.","url":"../modules/banditrlproof-exp3hedgeregret/index.html","parent":"chapter:exp3","order":281,"meta":[["Source","BanditRLProof/Exp3HedgeRegret.lean"],["Declarations","26"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3hedgeregret banditrlproof.exp3hedgeregret this module closes the full-information exponential-weights/hedge potential argument that sits immediately below exp3. it defines weights from cumulative losses, normalizes them on an explicit nonempty finite action set, proves the one-step logarithmic potential inequality, and telescopes it into second-order and [0,1] finite-horizon regret bounds. lean module compiled","shard":"modules/ac9368b3cb1556bb.json"},{"id":"module:BanditRLProof.Exp3HighProbabilityRegret","label":"Exp3HighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3HighProbabilityRegret","description":"This module completes the finite-horizon high-probability pseudo-regret route for the generated predictable EXP3 trajectory. A pathwise reciprocal-floor bound controls the random estimator-square sum almost surely. The final theorem then combines the sampled Hedge inequality, exploration bias, the pure-Hedge predictable-minus-observed confidence event, and the comparator-estimator confidence event by a two-event uni…","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html","parent":"chapter:exp3","order":282,"meta":[["Source","BanditRLProof/Exp3HighProbabilityRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3highprobabilityregret banditrlproof.exp3highprobabilityregret this module completes the finite-horizon high-probability pseudo-regret route for the generated predictable exp3 trajectory. a pathwise reciprocal-floor bound controls the random estimator-square sum almost surely. the final theorem then combines the sampled hedge inequality, exploration bias, the pure-hedge predictable-minus-observed confidence event, and the comparator-estimator confidence event by a two-event union bound. lean module compiled","shard":"modules/0823a558d097504b.json"},{"id":"module:BanditRLProof.Exp3ImportanceWeighted","label":"Exp3ImportanceWeighted","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ImportanceWeighted","description":"This module supplies the finite-distribution algebra immediately above the deterministic Hedge theorem and below the probabilistic EXP3 process. It proves the estimator's armwise weighted-sum cancellation and exact mixed-square identity on an explicit finite action set.","url":"../modules/banditrlproof-exp3importanceweighted/index.html","parent":"chapter:exp3","order":283,"meta":[["Source","BanditRLProof/Exp3ImportanceWeighted.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3importanceweighted banditrlproof.exp3importanceweighted this module supplies the finite-distribution algebra immediately above the deterministic hedge theorem and below the probabilistic exp3 process. it proves the estimator's armwise weighted-sum cancellation and exact mixed-square identity on an explicit finite action set. lean module compiled","shard":"modules/436ffdc4a38efa23.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernstein","label":"Exp3MixedSquareBernstein","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernstein","description":"The interval-sub-Gaussian route treats the mixed estimator square as an arbitrary variable in [0, 1 / epsilon], producing a variance proxy quadratic in the reciprocal exploration floor. Here the exact finite sampling law gives a centered second moment at most K / epsilon. A fixed-tilt conditional MGF argument then yields a generated finite-horizon Bernstein tail whose square-root term is linear, rather than quadrati…","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html","parent":"chapter:exp3","order":284,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernstein.lean"],["Declarations","11"],["Project imports","3"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernstein banditrlproof.exp3mixedsquarebernstein the interval-sub-gaussian route treats the mixed estimator square as an arbitrary variable in [0, 1 / epsilon], producing a variance proxy quadratic in the reciprocal exploration floor. here the exact finite sampling law gives a centered second moment at most k / epsilon. a fixed-tilt conditional mgf argument then yields a generated finite-horizon bernstein tail whose square-root term is linear, rather than quadratic, in that reciprocal floor. lean module compiled","shard":"modules/4f2dbc86b81509b4.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","label":"Exp3MixedSquareBernsteinHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","description":"This module consumes the generated mixed-square Bernstein tail in the existing three-event predictable-regret assembly. The square-event radius now uses the second-moment coefficient K / epsilon; the pure-cross and fixed-comparator events retain their compiled Bernstein radii.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinhighprobabilityregret/index.html","parent":"chapter:exp3","order":285,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernsteinHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernsteinhighprobabilityregret banditrlproof.exp3mixedsquarebernsteinhighprobabilityregret this module consumes the generated mixed-square bernstein tail in the existing three-event predictable-regret assembly. the square-event radius now uses the second-moment coefficient k / epsilon; the pure-cross and fixed-comparator events retain their compiled bernstein radii. lean module compiled","shard":"modules/e567f4449dfa2df4.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","label":"Exp3MixedSquareBernsteinRealizedAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","description":"The explicit variance-sensitive schedule has its refined 14 * gamma * T threshold when four horizon inequalities make clipping inactive. Outside that regime this module reuses the compiled almost-sure horizon bound and the strict T + 1 zero-probability threshold. The fallback covers every positive horizon without claiming a sharp active-clipping, Freedman, or EXP3.P rate.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedallhorizon/index.html","parent":"chapter:exp3","order":286,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedAllHorizon.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernsteinrealizedallhorizon banditrlproof.exp3mixedsquarebernsteinrealizedallhorizon the explicit variance-sensitive schedule has its refined 14 * gamma * t threshold when four horizon inequalities make clipping inactive. outside that regime this module reuses the compiled almost-sure horizon bound and the strict t + 1 zero-probability threshold. the fallback covers every positive horizon without claiming a sharp active-clipping, freedman, or exp3.p rate. lean module compiled","shard":"modules/c9017662babd9653.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","label":"Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","description":"This module upgrades the fixed-comparator all-horizon Bernstein mixed-square tail to the best supported arm in hindsight. The fixed-comparator schedule is calibrated at delta / K; all comparators share the same generated trajectory measure, so a finite union gives total failure probability at most delta.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedbestarmallhorizon/index.html","parent":"chapter:exp3","order":287,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedBestArmAllHorizon.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernsteinrealizedbestarmallhorizon banditrlproof.exp3mixedsquarebernsteinrealizedbestarmallhorizon this module upgrades the fixed-comparator all-horizon bernstein mixed-square tail to the best supported arm in hindsight. the fixed-comparator schedule is calibrated at delta / k; all comparators share the same generated trajectory measure, so a finite union gives total failure probability at most delta. lean module compiled","shard":"modules/b6cad411093bcbaa.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","label":"Exp3MixedSquareBernsteinRealizedExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","description":"The existing four-scale clipped exploration schedule is strong enough for the variance-sensitive mixed-square radius. Its sixth-power contract controls the new square-root term when gamma <= 1/2, while its arm and confidence contracts jointly control the linear log_+ / epsilon term. Thus the complete tuned threshold remains bounded by 14 * gamma * T.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html","parent":"chapter:exp3","order":288,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernsteinrealizedexplicittuning banditrlproof.exp3mixedsquarebernsteinrealizedexplicittuning the existing four-scale clipped exploration schedule is strong enough for the variance-sensitive mixed-square radius. its sixth-power contract controls the new square-root term when gamma <= 1/2, while its arm and confidence contracts jointly control the linear log_+ / epsilon term. thus the complete tuned threshold remains bounded by 14 * gamma * t. lean module compiled","shard":"modules/a4e6b128fc438c5d.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","label":"Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","description":"This module composes the generated predictable Bernstein-square regret theorem with the one-sided realized-minus-exploration deviation tail. The resulting four-event route controls generated selected scalar loss while using the mixed-square second-moment coefficient K / epsilon.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedhighprobabilityregret/index.html","parent":"chapter:exp3","order":289,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernsteinrealizedhighprobabilityregret banditrlproof.exp3mixedsquarebernsteinrealizedhighprobabilityregret this module composes the generated predictable bernstein-square regret theorem with the one-sided realized-minus-exploration deviation tail. the resulting four-event route controls generated selected scalar loss while using the mixed-square second-moment coefficient k / epsilon. lean module compiled","shard":"modules/6024f47d9033dc0e.json"},{"id":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","label":"Exp3MixedSquareBernsteinRealizedTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","description":"The variance-sensitive mixed-square route uses","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html","parent":"chapter:exp3","order":290,"meta":[["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarebernsteinrealizedtuning banditrlproof.exp3mixedsquarebernsteinrealizedtuning the variance-sensitive mixed-square route uses lean module compiled","shard":"modules/f519189cf91b7ea6.json"},{"id":"module:BanditRLProof.Exp3MixedSquareConfidence","label":"Exp3MixedSquareConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareConfidence","description":"This module replaces the Markov-only confidence step for the observed mixed importance-weighted estimator-square sum by a finite-action conditional sub-Gaussian route. The raw score lies in [0, 1 / epsilon] and its conditional mean is the armwise predictable loss-square sum. Consequently, the centered generated process has an exponential finite-horizon tail.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html","parent":"chapter:exp3","order":291,"meta":[["Source","BanditRLProof/Exp3MixedSquareConfidence.lean"],["Declarations","24"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquareconfidence banditrlproof.exp3mixedsquareconfidence this module replaces the markov-only confidence step for the observed mixed importance-weighted estimator-square sum by a finite-action conditional sub-gaussian route. the raw score lies in [0, 1 / epsilon] and its conditional mean is the armwise predictable loss-square sum. consequently, the centered generated process has an exponential finite-horizon tail. lean module compiled","shard":"modules/d70f7e49468531f1.json"},{"id":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","label":"Exp3MixedSquareExponentialHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","description":"This module replaces the Markov estimator-square event in the generated predictable EXP3 regret assembly by the compiled exponential mixed-square confidence theorem. The pure-cross and fixed-comparator confidence terms remain the existing variance-sensitive Bernstein radii.","url":"../modules/banditrlproof-exp3mixedsquareexponentialhighprobabilityregret/index.html","parent":"chapter:exp3","order":292,"meta":[["Source","BanditRLProof/Exp3MixedSquareExponentialHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquareexponentialhighprobabilityregret banditrlproof.exp3mixedsquareexponentialhighprobabilityregret this module replaces the markov estimator-square event in the generated predictable exp3 regret assembly by the compiled exponential mixed-square confidence theorem. the pure-cross and fixed-comparator confidence terms remain the existing variance-sensitive bernstein radii. lean module compiled","shard":"modules/8e28484e2139b0f5.json"},{"id":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","label":"Exp3MixedSquareExponentialRealizedAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","description":"The explicit exponential-square schedule has its refined 14 * gamma * T threshold when four horizon inequalities make clipping inactive. Outside that regime this module reuses the compiled almost-sure horizon bound and the strict T + 1 zero-probability threshold. The fallback closes the active-clipping presentation without claiming a sharp short-horizon, Freedman, or EXP3.P rate.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedallhorizon/index.html","parent":"chapter:exp3","order":293,"meta":[["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedAllHorizon.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquareexponentialrealizedallhorizon banditrlproof.exp3mixedsquareexponentialrealizedallhorizon the explicit exponential-square schedule has its refined 14 * gamma * t threshold when four horizon inequalities make clipping inactive. outside that regime this module reuses the compiled almost-sure horizon bound and the strict t + 1 zero-probability threshold. the fallback closes the active-clipping presentation without claiming a sharp short-horizon, freedman, or exp3.p rate. lean module compiled","shard":"modules/5bb3146d79843e5b.json"},{"id":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","label":"Exp3MixedSquareExponentialRealizedExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","description":"This module closes the remaining exploration-parameter leaf in the generated exponential mixed-square route. The interval sub-Gaussian proxy contributes a sixth-root scale because its range is |arms| / gamma; the two Bernstein radii contribute a cube-root scale, and the realized deviation contributes a square-root scale. The resulting rate is deliberately recorded as the output of the current Hoeffding-proxy route,…","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html","parent":"chapter:exp3","order":294,"meta":[["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquareexponentialrealizedexplicittuning banditrlproof.exp3mixedsquareexponentialrealizedexplicittuning this module closes the remaining exploration-parameter leaf in the generated exponential mixed-square route. the interval sub-gaussian proxy contributes a sixth-root scale because its range is |arms| / gamma; the two bernstein radii contribute a cube-root scale, and the realized deviation contributes a square-root scale. the resulting rate is deliberately recorded as the output of the current hoeffding-proxy route, not as a freedman or exp3.p rate. lean module compiled","shard":"modules/6d24f3b068b5394d.json"},{"id":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","label":"Exp3MixedSquareExponentialRealizedHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","description":"This module composes the generated predictable exponential-square Bernstein result with the one-sided realized-minus-exploration deviation tail. The resulting four-event route controls generated selected scalar loss while replacing the Markov estimator-square threshold by logarithmic confidence.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedhighprobabilityregret/index.html","parent":"chapter:exp3","order":295,"meta":[["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquareexponentialrealizedhighprobabilityregret banditrlproof.exp3mixedsquareexponentialrealizedhighprobabilityregret this module composes the generated predictable exponential-square bernstein result with the one-sided realized-minus-exploration deviation tail. the resulting four-event route controls generated selected scalar loss while replacing the markov estimator-square threshold by logarithmic confidence. lean module compiled","shard":"modules/6bf101da91579426.json"},{"id":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","label":"Exp3MixedSquareExponentialRealizedTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","description":"The exponential mixed-square route replaces the Markov square threshold by K * T + sampledMixedSquaredConfidenceRadius. This module chooses the exact learning rate","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html","parent":"chapter:exp3","order":296,"meta":[["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquareexponentialrealizedtuning banditrlproof.exp3mixedsquareexponentialrealizedtuning the exponential mixed-square route replaces the markov square threshold by k * t + sampledmixedsquaredconfidenceradius. this module chooses the exact learning rate lean module compiled","shard":"modules/23c5a73a7619a967.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","label":"Exp3MixedSquarePredictableVariance","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVariance","description":"This module promotes the exact finite-action centered second moment used by the fixed-tilt Bernstein proof to an explicit generated predictable process. It supplies the variance-process input needed by a future local Freedman iteration without claiming that such a tail theorem already exists in Mathlib.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html","parent":"chapter:exp3","order":297,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariance banditrlproof.exp3mixedsquarepredictablevariance this module promotes the exact finite-action centered second moment used by the fixed-tilt bernstein proof to an explicit generated predictable process. it supplies the variance-process input needed by a future local freedman iteration without claiming that such a tail theorem already exists in mathlib. lean module compiled","shard":"modules/cd8ef79537333a57.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","description":"This module transports the fixed-horizon predictable-variance tail to the observed mixed estimator-square sum used by the sampled Hedge inequality. It then exposes a predictable-regret theorem whose only uncontrolled probability is the overflow event for the cumulative predictable variance.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html","parent":"chapter:exp3","order":298,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancehighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancehighprobabilityregret this module transports the fixed-horizon predictable-variance tail to the observed mixed estimator-square sum used by the sampled hedge inequality. it then exposes a predictable-regret theorem whose only uncontrolled probability is the overflow event for the cumulative predictable variance. lean module compiled","shard":"modules/8179b50f04b867f0.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","description":"This module bounds the predictable mixed-square variance by the armwise predictable loss-square energy. It discharges the Markov route's variance lintegral contract from a pathwise cumulative loss-energy budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html","parent":"chapter:exp3","order":299,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret this module bounds the predictable mixed-square variance by the armwise predictable loss-square energy. it discharges the markov route's variance lintegral contract from a pathwise cumulative loss-energy budget. lean module compiled","shard":"modules/92b6fc3e407e4601.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","description":"The predictable-regret component retains the mixed-square estimator variance, while the realized-minus-predictable selected-loss component retains its own exact predictable variance. This replaces the fixed Hoeffding proxy in the realized component without changing the existing predictable-regret route.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizeddoublepredictablevariancehi-73ffccd3bfbc/index.html","parent":"chapter:exp3","order":300,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancerealizeddoublepredictablevariancehighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancerealizeddoublepredictablevariancehighprobabilityregret the predictable-regret component retains the mixed-square estimator variance, while the realized-minus-predictable selected-loss component retains its own exact predictable variance. this replaces the fixed hoeffding proxy in the realized component without changing the existing predictable-regret route. lean module compiled","shard":"modules/6275faacb1d275a2.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","description":"This module adds the generated realized-minus-predictable deviation to the random predictable-variance regret route. The resulting selected-loss regret bound preserves the cumulative variance overflow event explicitly.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedhighprobabilityregret/index.html","parent":"chapter:exp3","order":301,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancerealizedhighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancerealizedhighprobabilityregret this module adds the generated realized-minus-predictable deviation to the random predictable-variance regret route. the resulting selected-loss regret bound preserves the cumulative variance overflow event explicitly. lean module compiled","shard":"modules/022f2888bf1bc5ee.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","description":"This module discharges the explicit predictable-variance overflow residual by Markov's inequality. The resulting theorem requires a caller-supplied lintegral bound for the cumulative predictable mixed-square variance; no such algorithm-specific expectation bound is inferred from predictability.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html","parent":"chapter:exp3","order":302,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret this module discharges the explicit predictable-variance overflow residual by markov's inequality. the resulting theorem requires a caller-supplied lintegral bound for the cumulative predictable mixed-square variance; no such algorithm-specific expectation bound is inferred from predictability. lean module compiled","shard":"modules/176571e6347c5e89.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","description":"The predictable-regret component uses the armwise loss-mass budget and the mixed-square predictable variance. The realized-minus-predictable component uses its exact selected-loss predictable variance. An explicit bad set is retained so probabilistic sparsity can discharge both pathwise budgets without charging the same failure event more than once.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizeddoublepredictablev-6b1be70bfcd6/index.html","parent":"chapter:exp3","order":303,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesmalllossrealizeddoublepredictablevariancehighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancesmalllossrealizeddoublepredictablevariancehighprobabilityregret the predictable-regret component uses the armwise loss-mass budget and the mixed-square predictable variance. the realized-minus-predictable component uses its exact selected-loss predictable variance. an explicit bad set is retained so probabilistic sparsity can discharge both pathwise budgets without charging the same failure event more than once. lean module compiled","shard":"modules/e1c5c3939b6c0e15.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","description":"This module bounds predictable loss-square energy by armwise predictable loss mass. It turns an almost-everywhere small-loss budget under the exact generated trajectory measure into the variance lintegral contract used by the Markov-closed realized regret route. Universal pathwise budgets remain valid as a special case via Filter.Eventually.of_forall.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html","parent":"chapter:exp3","order":304,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret this module bounds predictable loss-square energy by armwise predictable loss mass. it turns an almost-everywhere small-loss budget under the exact generated trajectory measure into the variance lintegral contract used by the markov-closed realized regret route. universal pathwise budgets remain valid as a special case via filter.eventually.of_forall. lean module compiled","shard":"modules/fcce993f0a19b8c3.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","description":"Outside the explicit sparsity-failure event, the mixed-square predictable variance is bounded by (1 / (gamma / K)) * (S * T) and the exact selected-loss predictable variance is bounded by S * T. Four confidence events receive delta / 4, and the common sparsity-failure set is charged exactly once.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html","parent":"chapter:exp3","order":305,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsity banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsity outside the explicit sparsity-failure event, the mixed-square predictable variance is bounded by (1 / (gamma / k)) * (s * t) and the exact selected-loss predictable variance is bounded by s * t. four confidence events receive delta / 4, and the common sparsity-failure set is charged exactly once. lean module compiled","shard":"modules/659a04d8d9ce0435.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","description":"The exact double-variance schedule has the refined 16 * gamma * T threshold when four horizon inequalities make clipping inactive. Outside that regime this module uses the strict T + 1 zero-probability threshold under the identical internal eta, gamma, and generated trajectory measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f7cd8884d456/index.html","parent":"chapter:exp3","order":306,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsityallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsityallhorizon the exact double-variance schedule has the refined 16 * gamma * t threshold when four horizon inequalities make clipping inactive. outside that regime this module uses the strict t + 1 zero-probability threshold under the identical internal eta, gamma, and generated trajectory measure. lean module compiled","shard":"modules/877b96a01f3788a5.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","description":"This module upgrades the exact double-variance fixed-comparator all-horizon tail to the best supported arm in hindsight. Confidence is calibrated at delta / K; the common support-sparsity failure event is removed before the finite comparator union and added only once afterward.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f70e76a4b1d2/index.html","parent":"chapter:exp3","order":307,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsitybestarmallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsitybestarmallhorizon this module upgrades the exact double-variance fixed-comparator all-horizon tail to the best supported arm in hindsight. confidence is calibrated at delta / k; the common support-sparsity failure event is removed before the finite comparator union and added only once afterward. lean module compiled","shard":"modules/529aad5bcb19c46d.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","description":"This module closes gamma tuning for the sparse generated-regret theorem that uses both the mixed-square predictable variance and the exact selected-loss predictable variance. The first three exploration scales are shared with the single-variance pathwise route. The additional selected-loss scale is","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html","parent":"chapter:exp3","order":308,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsityexplicittuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsityexplicittuning this module closes gamma tuning for the sparse generated-regret theorem that uses both the mixed-square predictable variance and the exact selected-loss predictable variance. the first three exploration scales are shared with the single-variance pathwise route. the additional selected-loss scale is lean module compiled","shard":"modules/2410943acc92ee4e.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","description":"This module reuses the pathwise-sparsity learning rate that balances entropy against the sparse loss-mass and mixed-square radius. The realized-loss term is replaced by the exact selected-loss predictable-variance radius. Gamma remains caller-selected, and the common sparsity-failure event is still charged once.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-86f08ff5ad1a/index.html","parent":"chapter:exp3","order":309,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsitytuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevarianceprobabilisticsparsitytuning this module reuses the pathwise-sparsity learning rate that balances entropy against the sparse loss-mass and mixed-square radius. the realized-loss term is replaced by the exact selected-loss predictable-variance radius. gamma remains caller-selected, and the common sparsity-failure event is still charged once. lean module compiled","shard":"modules/507ceac1b8ba411e.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","description":"This module replaces the universal pathwise sparse-support contract of the existing all-horizon route by an almost-everywhere contract under the exact generated trajectory measure. The large-horizon branch combines the compiled raw sparse-loss tail with the tuned and explicit budget comparisons. The complementary branch keeps the strict T + 1 zero-probability fallback.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovaesparsityallhorizon/index.html","parent":"chapter:exp3","order":310,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovaesparsityallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovaesparsityallhorizon this module replaces the universal pathwise sparse-support contract of the existing all-horizon route by an almost-everywhere contract under the exact generated trajectory measure. the large-horizon branch combines the compiled raw sparse-loss tail with the tuned and explicit budget comparisons. the complementary branch keeps the strict t + 1 zero-probability fallback. lean module compiled","shard":"modules/fae895a5c7b5f65c.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","description":"The explicit sparse-loss schedule has its refined 14 * gamma * T threshold when four horizon inequalities make clipping inactive. Outside that regime this module reuses the compiled almost-sure horizon bound and the strict T + 1 zero-probability threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovallhorizon/index.html","parent":"chapter:exp3","order":311,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovallhorizon the explicit sparse-loss schedule has its refined 14 * gamma * t threshold when four horizon inequalities make clipping inactive. outside that regime this module reuses the compiled almost-sure horizon bound and the strict t + 1 zero-probability threshold. lean module compiled","shard":"modules/95cf2e92e47865e3.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","description":"This module closes the exploration-parameter leaf for the sparse-loss predictable-variance Markov route. Besides the usual sparse arm, Bernstein confidence, and realized-deviation scales, the Markov variance threshold introduces the fifth-root scale","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html","parent":"chapter:exp3","order":312,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean"],["Declarations","22"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning this module closes the exploration-parameter leaf for the sparse-loss predictable-variance markov route. besides the usual sparse arm, bernstein confidence, and realized-deviation scales, the markov variance threshold introduces the fifth-root scale lean module compiled","shard":"modules/a2135a1b34dce681.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","description":"This module discharges the pathwise armwise loss-mass premise of the small-loss route from a per-round support-cardinality contract. The local proof route is:","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html","parent":"chapter:exp3","order":313,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret this module discharges the pathwise armwise loss-mass premise of the small-loss route from a per-round support-cardinality contract. the local proof route is: lean module compiled","shard":"modules/9f2f32b15d579ccb.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","description":"This module allows the per-round support-cardinality contract to fail on an explicit generated-trajectory event. On paths outside that event, sparsity still supplies the sparsity * horizon loss-mass budget used by the observed mixed-square and Hedge terms. The Markov variance closure instead uses the global arms.card * horizon loss-mass envelope, so the exceptional event costs its actual generated-measure probabilit…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html","parent":"chapter:exp3","order":314,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity this module allows the per-round support-cardinality contract to fail on an explicit generated-trajectory event. on paths outside that event, sparsity still supplies the sparsity * horizon loss-mass budget used by the observed mixed-square and hedge terms. the markov variance closure instead uses the global arms.card * horizon loss-mass envelope, so the exceptional event costs its actual generated-measure probability without assuming it is null. lean module compiled","shard":"modules/f101f1a70bd37105.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","description":"The explicit probabilistic-sparsity schedule has its refined 14 * gamma * T threshold when four horizon inequalities make clipping inactive. Outside that regime this module uses the strict T + 1 zero-probability threshold under exactly the same internal eta, gamma, and generated trajectory measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-306e3ec30e1f/index.html","parent":"chapter:exp3","order":315,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsityallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsityallhorizon the explicit probabilistic-sparsity schedule has its refined 14 * gamma * t threshold when four horizon inequalities make clipping inactive. outside that regime this module uses the strict t + 1 zero-probability threshold under exactly the same internal eta, gamma, and generated trajectory measure. lean module compiled","shard":"modules/62f8e0413cc173dd.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","description":"This module closes the exploration-rate leaf for the generated realized predictable-variance EXP3 route whose support-sparsity condition may fail with positive probability. The global K * T loss-mass envelope makes the Markov component","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html","parent":"chapter:exp3","order":316,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean"],["Declarations","18"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsityexplicittuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsityexplicittuning this module closes the exploration-rate leaf for the generated realized predictable-variance exp3 route whose support-sparsity condition may fail with positive probability. the global k * t loss-mass envelope makes the markov component lean module compiled","shard":"modules/71993992e89b3475.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","description":"This module tunes eta for the generated realized predictable-variance EXP3 route whose support-sparsity contract may fail with positive probability. The exact learning-rate scale is","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html","parent":"chapter:exp3","order":317,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsitytuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsitytuning this module tunes eta for the generated realized predictable-variance exp3 route whose support-sparsity contract may fail with positive probability. the exact learning-rate scale is lean module compiled","shard":"modules/f5373e7fc179aad6.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","description":"This module chooses the learning rate for the compiled sparse-loss realized Markov route. With L = sparsity * horizon, the exact Hedge stability scale is","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html","parent":"chapter:exp3","order":318,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning this module chooses the learning rate for the compiled sparse-loss realized markov route. with l = sparsity * horizon, the exact hedge stability scale is lean module compiled","shard":"modules/24c4d202a35479be.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","description":"This module removes the global arms.card * horizon Markov envelope from the positive-probability sparsity route. Outside the explicit sparsity-failure event, the armwise loss mass is at most sparsity * horizon; the existing pointwise mixed-square variance inequality therefore gives the deterministic variance budget","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html","parent":"chapter:exp3","order":319,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsity banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsity this module removes the global arms.card * horizon markov envelope from the positive-probability sparsity route. outside the explicit sparsity-failure event, the armwise loss mass is at most sparsity * horizon; the existing pointwise mixed-square variance inequality therefore gives the deterministic variance budget lean module compiled","shard":"modules/626461cb8ae51bf1.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","description":"The explicit pathwise-variance schedule has its refined 14 * gamma * T threshold when four horizon inequalities make clipping inactive. Outside that regime this module uses the strict T + 1 zero-probability threshold under exactly the same internal eta, gamma, and generated trajectory measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-40d17aa0b72d/index.html","parent":"chapter:exp3","order":320,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsityallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsityallhorizon the explicit pathwise-variance schedule has its refined 14 * gamma * t threshold when four horizon inequalities make clipping inactive. outside that regime this module uses the strict t + 1 zero-probability threshold under exactly the same internal eta, gamma, and generated trajectory measure. lean module compiled","shard":"modules/0aa582151645075c.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","description":"This module upgrades the fixed supported-comparator all-horizon tail to the best supported arm in hindsight. The confidence budget is calibrated armwise as delta / K. The fixed-comparator off-sparsityFailure tail is unioned over the arms, so the common sparsity-failure event is added only once afterward. The strengthened residual is delta + mu(sparsityFailure), and its practical consumer needs only mu(sparsityFailur…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html","parent":"chapter:exp3","order":321,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsitybestarmallhorizon banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsitybestarmallhorizon this module upgrades the fixed supported-comparator all-horizon tail to the best supported arm in hindsight. the confidence budget is calibrated armwise as delta / k. the fixed-comparator off-sparsityfailure tail is unioned over the arms, so the common sparsity-failure event is added only once afterward. the strengthened residual is delta + mu(sparsityfailure), and its practical consumer needs only mu(sparsityfailure) <= ofreal epsilon. lean module compiled","shard":"modules/ecf5383114beb843.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","description":"This module closes the exploration-rate route for the four-event generated realized-regret theorem under probabilistic sparsity. On the good event, the predictable-variance budget is","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html","parent":"chapter:exp3","order":322,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean"],["Declarations","20"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsityexplicittuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsityexplicittuning this module closes the exploration-rate route for the four-event generated realized-regret theorem under probabilistic sparsity. on the good event, the predictable-variance budget is lean module compiled","shard":"modules/1a76e426f30074c2.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","label":"Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","description":"This module tunes eta for the four-event probabilistic-sparsity route. Its variance budget is the deterministic good-path bound","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html","parent":"chapter:exp3","order":323,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsitytuning banditrlproof.exp3mixedsquarepredictablevariancesparselossrealizedpathwisevarianceprobabilisticsparsitytuning this module tunes eta for the four-event probabilistic-sparsity route. its variance budget is the deterministic good-path bound lean module compiled","shard":"modules/26a04dea238c920a.json"},{"id":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","label":"Exp3MixedSquarePredictableVarianceTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3MixedSquarePredictableVarianceTail","description":"This module keeps the exact conditional second moment random. It first compensates each centered mixed-square increment by its finite-law variance, then iterates the resulting zero-budget conditional MGF. The main endpoint is a fixed-tilt tail on the event that the cumulative predictable variance is at most a caller-supplied budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html","parent":"chapter:exp3","order":324,"meta":[["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3mixedsquarepredictablevariancetail banditrlproof.exp3mixedsquarepredictablevariancetail this module keeps the exact conditional second moment random. it first compensates each centered mixed-square increment by its finite-law variance, then iterates the resulting zero-budget conditional mgf. the main endpoint is a fixed-tilt tail on the event that the cumulative predictable variance is at most a caller-supplied budget. lean module compiled","shard":"modules/0e68ce7abd9ca78a.json"},{"id":"module:BanditRLProof.Exp3Potential","label":"Exp3Potential","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3Potential","description":"This module records the deterministic finite-action potential surface used by exponential-weights/EXP3 routes. It deliberately stops before importance weighted estimators, logarithmic inequalities, learning-rate optimization, or a regret theorem.","url":"../modules/banditrlproof-exp3potential/index.html","parent":"chapter:exp3","order":325,"meta":[["Source","BanditRLProof/Exp3Potential.lean"],["Declarations","10"],["Project imports","0"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3potential banditrlproof.exp3potential this module records the deterministic finite-action potential surface used by exponential-weights/exp3 routes. it deliberately stops before importance weighted estimators, logarithmic inequalities, learning-rate optimization, or a regret theorem. lean module compiled","shard":"modules/a90d8336912dfc27.json"},{"id":"module:BanditRLProof.Exp3PredictableAdversary","label":"Exp3PredictableAdversary","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PredictableAdversary","description":"This module packages history-dependent loss vectors that are fixed before the current action is sampled. It realizes their chosen coordinate as a deterministic feedback environment and keeps the EXP3 action law valid after conditioning on both the latent environment and the visible finite history.","url":"../modules/banditrlproof-exp3predictableadversary/index.html","parent":"chapter:exp3","order":326,"meta":[["Source","BanditRLProof/Exp3PredictableAdversary.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3predictableadversary banditrlproof.exp3predictableadversary this module packages history-dependent loss vectors that are fixed before the current action is sampled. it realizes their chosen coordinate as a deterministic feedback environment and keeps the exp3 action law valid after conditioning on both the latent environment and the visible finite history. lean module compiled","shard":"modules/71f59f5ffeeaf39a.json"},{"id":"module:BanditRLProof.Exp3PredictableHedge","label":"Exp3PredictableHedge","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PredictableHedge","description":"This module discharges the pathwise nonnegative-feedback premise of Exp3SampledHedge from the generated predictable [0,1] reward law. It first aggregates the time-zero and successor reward identifications into one finite-horizon almost-sure event, then applies the concrete pathwise Hedge bound on that event.","url":"../modules/banditrlproof-exp3predictablehedge/index.html","parent":"chapter:exp3","order":327,"meta":[["Source","BanditRLProof/Exp3PredictableHedge.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3predictablehedge banditrlproof.exp3predictablehedge this module discharges the pathwise nonnegative-feedback premise of exp3sampledhedge from the generated predictable [0,1] reward law. it first aggregates the time-zero and successor reward identifications into one finite-horizon almost-sure event, then applies the concrete pathwise hedge bound on that event. lean module compiled","shard":"modules/5f867901da88b643.json"},{"id":"module:BanditRLProof.Exp3PredictableIntegration","label":"Exp3PredictableIntegration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PredictableIntegration","description":"This module connects the pathwise sampled-Hedge inequality to the generated predictable trajectory moments. The key law transport uses the exploration distribution p_t to sample an action while a distinct predictable pure-Hedge distribution q_t weights the importance-weighted estimator.","url":"../modules/banditrlproof-exp3predictableintegration/index.html","parent":"chapter:exp3","order":328,"meta":[["Source","BanditRLProof/Exp3PredictableIntegration.lean"],["Declarations","22"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3predictableintegration banditrlproof.exp3predictableintegration this module connects the pathwise sampled-hedge inequality to the generated predictable trajectory moments. the key law transport uses the exploration distribution p_t to sample an action while a distinct predictable pure-hedge distribution q_t weights the importance-weighted estimator. lean module compiled","shard":"modules/da27cf4f49cd00c0.json"},{"id":"module:BanditRLProof.Exp3PredictableMoments","label":"Exp3PredictableMoments","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PredictableMoments","description":"This file transports the fixed-environment canonical trajectory law through an environment prior. The resulting global joint law retains the latent environment in the conditioning history, which is the law surface needed to identify predictable feedback coordinates and their roundwise moments.","url":"../modules/banditrlproof-exp3predictablemoments/index.html","parent":"chapter:exp3","order":329,"meta":[["Source","BanditRLProof/Exp3PredictableMoments.lean"],["Declarations","33"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3predictablemoments banditrlproof.exp3predictablemoments this file transports the fixed-environment canonical trajectory law through an environment prior. the resulting global joint law retains the latent environment in the conditioning history, which is the law surface needed to identify predictable feedback coordinates and their roundwise moments. lean module compiled","shard":"modules/43b280a37cc70ee8.json"},{"id":"module:BanditRLProof.Exp3PredictableRegretAllTime","label":"Exp3PredictableRegretAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PredictableRegretAllTime","description":"This module gives every positive prefix of one fixed generated EXP3 process a geometric confidence share and reuses the compiled fixed-horizon pathwise potential, exploration, and comparator assembly. The result is countable outer-measure subadditivity, not a Ville/Doob, mixture, optional-stopping, self-normalized, general Freedman, or tuned horizon-free EXP3 theorem.","url":"../modules/banditrlproof-exp3predictableregretalltime/index.html","parent":"chapter:exp3","order":330,"meta":[["Source","BanditRLProof/Exp3PredictableRegretAllTime.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3predictableregretalltime banditrlproof.exp3predictableregretalltime this module gives every positive prefix of one fixed generated exp3 process a geometric confidence share and reuses the compiled fixed-horizon pathwise potential, exploration, and comparator assembly. the result is countable outer-measure subadditivity, not a ville/doob, mixture, optional-stopping, self-normalized, general freedman, or tuned horizon-free exp3 theorem. lean module compiled","shard":"modules/092d1623420c79bb.json"},{"id":"module:BanditRLProof.Exp3PureBernstein","label":"Exp3PureBernstein","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PureBernstein","description":"This module replaces the range-squared Hoeffding proxy for the pure-Hedge cross-weighted estimator by a fixed-tilt second-moment budget. The sign is predictable pure loss - observed cross-weighted loss, as consumed by the sampled Hedge regret decomposition.","url":"../modules/banditrlproof-exp3purebernstein/index.html","parent":"chapter:exp3","order":331,"meta":[["Source","BanditRLProof/Exp3PureBernstein.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3purebernstein banditrlproof.exp3purebernstein this module replaces the range-squared hoeffding proxy for the pure-hedge cross-weighted estimator by a fixed-tilt second-moment budget. the sign is predictable pure loss - observed cross-weighted loss, as consumed by the sampled hedge regret decomposition. lean module compiled","shard":"modules/10b14b64d48b6652.json"},{"id":"module:BanditRLProof.Exp3PureConfidence","label":"Exp3PureConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3PureConfidence","description":"This module proves finite-horizon confidence bounds for the pure-Hedge weighted importance estimator, including the predictable-minus-observed direction needed by the regret decomposition. It first identifies the conditional mean under the exploration action law, transports the latent predictable score to the observed trajectory score, proves adaptedness, and applies the local conditional sub-Gaussian finite-sum tai…","url":"../modules/banditrlproof-exp3pureconfidence/index.html","parent":"chapter:exp3","order":332,"meta":[["Source","BanditRLProof/Exp3PureConfidence.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3pureconfidence banditrlproof.exp3pureconfidence this module proves finite-horizon confidence bounds for the pure-hedge weighted importance estimator, including the predictable-minus-observed direction needed by the regret decomposition. it first identifies the conditional mean under the exploration action law, transports the latent predictable score to the observed trajectory score, proves adaptedness, and applies the local conditional sub-gaussian finite-sum tail theorem. lean module compiled","shard":"modules/895c11c7bda6720a.json"},{"id":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","label":"Exp3RandomSquareBernsteinRealizedAllHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","description":"The explicit random-square schedule has its refined threshold when two large-horizon inequalities make clipping inactive. Outside that regime this module reuses the compiled almost-sure horizon bound and a strict T + 1 threshold. The result covers every positive horizon without asserting cubic or quadratic dominance on the active-clipping branch.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedallhorizon/index.html","parent":"chapter:exp3","order":333,"meta":[["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedAllHorizon.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3randomsquarebernsteinrealizedallhorizon banditrlproof.exp3randomsquarebernsteinrealizedallhorizon the explicit random-square schedule has its refined threshold when two large-horizon inequalities make clipping inactive. outside that regime this module reuses the compiled almost-sure horizon bound and a strict t + 1 threshold. the result covers every positive horizon without asserting cubic or quadratic dominance on the active-clipping branch. lean module compiled","shard":"modules/52284718dbde5c44.json"},{"id":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","label":"Exp3RandomSquareBernsteinRealizedExplicitTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","description":"The random-square theorem already tunes the learning rate independently of the exploration parameter. This module chooses the remaining exploration parameter from the two confidence scales at failure budget delta / 4 and obtains a fully explicit generated realized-regret threshold.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html","parent":"chapter:exp3","order":334,"meta":[["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3randomsquarebernsteinrealizedexplicittuning banditrlproof.exp3randomsquarebernsteinrealizedexplicittuning the random-square theorem already tunes the learning rate independently of the exploration parameter. this module chooses the remaining exploration parameter from the two confidence scales at failure budget delta / 4 and obtains a fully explicit generated realized-regret threshold. lean module compiled","shard":"modules/263eb29d89a3e112.json"},{"id":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","label":"Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","description":"This module composes the generated predictable random-square Bernstein theorem with the one-sided realized-minus-exploration deviation tail. The resulting four-event route controls generated selected scalar loss while preserving the random |arms| * T / deltaSquare Hedge-square budget.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedhighprobabilityregret/index.html","parent":"chapter:exp3","order":335,"meta":[["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3randomsquarebernsteinrealizedhighprobabilityregret banditrlproof.exp3randomsquarebernsteinrealizedhighprobabilityregret this module composes the generated predictable random-square bernstein theorem with the one-sided realized-minus-exploration deviation tail. the resulting four-event route controls generated selected scalar loss while preserving the random |arms| * t / deltasquare hedge-square budget. lean module compiled","shard":"modules/4004bf5482c4c440.json"},{"id":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","label":"Exp3RandomSquareBernsteinRealizedTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","description":"The random estimator-square route replaces the pathwise K * T / gamma budget by K * T / deltaSquare. At the public four-event allocation, deltaSquare = delta / 4. This module chooses","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html","parent":"chapter:exp3","order":336,"meta":[["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3randomsquarebernsteinrealizedtuning banditrlproof.exp3randomsquarebernsteinrealizedtuning the random estimator-square route replaces the pathwise k * t / gamma budget by k * t / deltasquare. at the public four-event allocation, deltasquare = delta / 4. this module chooses lean module compiled","shard":"modules/84268fd8bd9855c7.json"},{"id":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","label":"Exp3RandomSquareHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RandomSquareHighProbabilityRegret","description":"The pathwise EXP3 assembly bounds the mixed importance-weighted estimator-square sum by |arms| * T / gamma. Its expectation is at most |arms| * T. This module turns that expectation bound into a Markov tail and includes the square event beside the two existing Bernstein confidence events. The resulting regret theorem removes the reciprocal exploration factor from the Hedge-square budget, at the honest cost of a 1 /…","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html","parent":"chapter:exp3","order":337,"meta":[["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3randomsquarehighprobabilityregret banditrlproof.exp3randomsquarehighprobabilityregret the pathwise exp3 assembly bounds the mixed importance-weighted estimator-square sum by |arms| * t / gamma. its expectation is at most |arms| * t. this module turns that expectation bound into a markov tail and includes the square event beside the two existing bernstein confidence events. the resulting regret theorem removes the reciprocal exploration factor from the hedge-square budget, at the honest cost of a 1 / deltasquare failure allocation. lean module compiled","shard":"modules/d605b90fd453a37a.json"},{"id":"module:BanditRLProof.Exp3RealizedConcentration","label":"Exp3RealizedConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedConcentration","description":"This module identifies the successor action law inside condExpKernel without requiring a countable ambient action type. It then freezes the predictable environment/history coordinates and applies the bounded centered Hoeffding MGF bound to the selected and realized one-step loss deviations.","url":"../modules/banditrlproof-exp3realizedconcentration/index.html","parent":"chapter:exp3","order":338,"meta":[["Source","BanditRLProof/Exp3RealizedConcentration.lean"],["Declarations","6"],["Project imports","3"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedconcentration banditrlproof.exp3realizedconcentration this module identifies the successor action law inside condexpkernel without requiring a countable ambient action type. it then freezes the predictable environment/history coordinates and applies the bounded centered hoeffding mgf bound to the selected and realized one-step loss deviations. lean module compiled","shard":"modules/de2dee354111cae8.json"},{"id":"module:BanditRLProof.Exp3RealizedConfidence","label":"Exp3RealizedConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedConfidence","description":"This module converts the finite-horizon ENNReal Azuma bound into an explicit square-root confidence radius. It closes the one-sided confidence theorem for realized loss minus the exploration-mixed predictable conditional mean; it does not identify the estimator-valued Hedge comparator with true comparator loss.","url":"../modules/banditrlproof-exp3realizedconfidence/index.html","parent":"chapter:exp3","order":339,"meta":[["Source","BanditRLProof/Exp3RealizedConfidence.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedconfidence banditrlproof.exp3realizedconfidence this module converts the finite-horizon ennreal azuma bound into an explicit square-root confidence radius. it closes the one-sided confidence theorem for realized loss minus the exploration-mixed predictable conditional mean; it does not identify the estimator-valued hedge comparator with true comparator loss. lean module compiled","shard":"modules/52f01b4c40547c85.json"},{"id":"module:BanditRLProof.Exp3RealizedDeviationAllTime","label":"Exp3RealizedDeviationAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedDeviationAllTime","description":"This module discharges the variance-good conjunct in the geometric all-time tail with the deterministic unit bound on each exact selected-loss predictable variance. The process parameters remain fixed outside the countable index.","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html","parent":"chapter:exp3","order":340,"meta":[["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizeddeviationalltime banditrlproof.exp3realizeddeviationalltime this module discharges the variance-good conjunct in the geometric all-time tail with the deterministic unit bound on each exact selected-loss predictable variance. the process parameters remain fixed outside the countable index. lean module compiled","shard":"modules/274f260682a6b672.json"},{"id":"module:BanditRLProof.Exp3RealizedDeviationTail","label":"Exp3RealizedDeviationTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedDeviationTail","description":"This module shifts the generated predictable realized-deviation process by one time step so that its deterministic zero initial value and every actual round fit Mathlib's conditional sub-Gaussian sum theorem. The resulting public theorem bounds the full finite-horizon realized-minus-exploration-mixed loss sum under a probability environment prior.","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html","parent":"chapter:exp3","order":341,"meta":[["Source","BanditRLProof/Exp3RealizedDeviationTail.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizeddeviationtail banditrlproof.exp3realizeddeviationtail this module shifts the generated predictable realized-deviation process by one time step so that its deterministic zero initial value and every actual round fit mathlib's conditional sub-gaussian sum theorem. the resulting public theorem bounds the full finite-horizon realized-minus-exploration-mixed loss sum under a probability environment prior. lean module compiled","shard":"modules/ab3091012b05494b.json"},{"id":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","label":"Exp3RealizedHighProbabilityRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedHighProbabilityRegret","description":"This module composes the generated predictable EXP3 high-probability theorem with the one-sided realized-minus-exploration-mixed deviation tail. The primary endpoint controls the scalar loss stored in the generated trajectory, with the requested total failure probability split equally across the pure-q, comparator-estimator, and realized-deviation events.","url":"../modules/banditrlproof-exp3realizedhighprobabilityregret/index.html","parent":"chapter:exp3","order":342,"meta":[["Source","BanditRLProof/Exp3RealizedHighProbabilityRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedhighprobabilityregret banditrlproof.exp3realizedhighprobabilityregret this module composes the generated predictable exp3 high-probability theorem with the one-sided realized-minus-exploration-mixed deviation tail. the primary endpoint controls the scalar loss stored in the generated trajectory, with the requested total failure probability split equally across the pure-q, comparator-estimator, and realized-deviation events. lean module compiled","shard":"modules/469a5d2c5a4fca15.json"},{"id":"module:BanditRLProof.Exp3RealizedPredictableVariance","label":"Exp3RealizedPredictableVariance","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedPredictableVariance","description":"This module replaces the fixed interval proxy for the selected-loss deviation by its exact finite-action centered second moment. It constructs the generated predictable variance process and the zero-budget conditional MGF of the variance-compensated realized deviation.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html","parent":"chapter:exp3","order":343,"meta":[["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean"],["Declarations","23"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedpredictablevariance banditrlproof.exp3realizedpredictablevariance this module replaces the fixed interval proxy for the selected-loss deviation by its exact finite-action centered second moment. it constructs the generated predictable variance process and the zero-budget conditional mgf of the variance-compensated realized deviation. lean module compiled","shard":"modules/6d2efe210f9cede2.json"},{"id":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","label":"Exp3RealizedPredictableVarianceAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedPredictableVarianceAllTime","description":"This module gives every positive prefix of one fixed generated EXP3 process a geometric confidence share. The resulting countable union is controlled by one outer confidence budget. This is countable outer-measure subadditivity, not a Ville/Doob, mixture, optional-stopping, self-normalized, or general Freedman theorem.","url":"../modules/banditrlproof-exp3realizedpredictablevariancealltime/index.html","parent":"chapter:exp3","order":344,"meta":[["Source","BanditRLProof/Exp3RealizedPredictableVarianceAllTime.lean"],["Declarations","4"],["Project imports","3"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedpredictablevariancealltime banditrlproof.exp3realizedpredictablevariancealltime this module gives every positive prefix of one fixed generated exp3 process a geometric confidence share. the resulting countable union is controlled by one outer confidence budget. this is countable outer-measure subadditivity, not a ville/doob, mixture, optional-stopping, self-normalized, or general freedman theorem. lean module compiled","shard":"modules/3dec3ce7e621a217.json"},{"id":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","label":"Exp3RealizedPredictableVarianceMaximal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedPredictableVarianceMaximal","description":"This module applies the finite maximal quadratic fixed-MGF route to every positive prefix of the generated realized-loss deviation process.","url":"../modules/banditrlproof-exp3realizedpredictablevariancemaximal/index.html","parent":"chapter:exp3","order":345,"meta":[["Source","BanditRLProof/Exp3RealizedPredictableVarianceMaximal.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedpredictablevariancemaximal banditrlproof.exp3realizedpredictablevariancemaximal this module applies the finite maximal quadratic fixed-mgf route to every positive prefix of the generated realized-loss deviation process. lean module compiled","shard":"modules/f84a454b7f2d26d3.json"},{"id":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","label":"Exp3RealizedPredictableVarianceTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedPredictableVarianceTail","description":"This module iterates the exact selected-loss variance-compensated conditional MGF along the generated trajectory. It yields a Bernstein-shaped upper tail for realized-minus-predictable selected loss jointly with a pathwise budget on the cumulative selected-loss predictable variance.","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html","parent":"chapter:exp3","order":346,"meta":[["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedpredictablevariancetail banditrlproof.exp3realizedpredictablevariancetail this module iterates the exact selected-loss variance-compensated conditional mgf along the generated trajectory. it yields a bernstein-shaped upper tail for realized-minus-predictable selected loss jointly with a pathwise budget on the cumulative selected-loss predictable variance. lean module compiled","shard":"modules/5a031350159eb85e.json"},{"id":"module:BanditRLProof.Exp3RealizedRegret","label":"Exp3RealizedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedRegret","description":"This module transports the compiled predictable p_t-mixed expected-regret bound to the scalar loss actually observed on the generated trajectory. The only probabilistic step is the existing conditional action law: conditionally on the pre-action history, the sampled action has finite distribution p_t.","url":"../modules/banditrlproof-exp3realizedregret/index.html","parent":"chapter:exp3","order":347,"meta":[["Source","BanditRLProof/Exp3RealizedRegret.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedregret banditrlproof.exp3realizedregret this module transports the compiled predictable p_t-mixed expected-regret bound to the scalar loss actually observed on the generated trajectory. the only probabilistic step is the existing conditional action law: conditionally on the pre-action history, the sampled action has finite distribution p_t. lean module compiled","shard":"modules/33851dc85a03b8c2.json"},{"id":"module:BanditRLProof.Exp3RealizedRegretAllTime","label":"Exp3RealizedRegretAllTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RealizedRegretAllTime","description":"This module combines the accepted same-process predictable-regret and pure realized-deviation all-time events. The total confidence budget is split equally between those two event families. The result is an outer-measure bound for realized selected-loss regret at every positive prefix, with fixed process parameters and one fixed supported comparator.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html","parent":"chapter:exp3","order":348,"meta":[["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3realizedregretalltime banditrlproof.exp3realizedregretalltime this module combines the accepted same-process predictable-regret and pure realized-deviation all-time events. the total confidence budget is split equally between those two event families. the result is an outer-measure bound for realized selected-loss regret at every positive prefix, with fixed process parameters and one fixed supported comparator. lean module compiled","shard":"modules/b2fdeaca9a90732f.json"},{"id":"module:BanditRLProof.Exp3RecursiveTrajectory","label":"Exp3RecursiveTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3RecursiveTrajectory","description":"This module turns a measurable finite-history score into the exploration-mixed exponential policy used by EXP3. It proves normalization, coordinate measurability, and a uniform exploration floor, packages the policy as the project's stochastic finite-history algorithm, and invokes the Mathlib-backed Ionescu--Tulcea trajectory kernel. The resulting theorem identifies every successor action's conditional law with the…","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html","parent":"chapter:exp3","order":349,"meta":[["Source","BanditRLProof/Exp3RecursiveTrajectory.lean"],["Declarations","25"],["Project imports","2"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3recursivetrajectory banditrlproof.exp3recursivetrajectory this module turns a measurable finite-history score into the exploration-mixed exponential policy used by exp3. it proves normalization, coordinate measurability, and a uniform exploration floor, packages the policy as the project's stochastic finite-history algorithm, and invokes the mathlib-backed ionescu--tulcea trajectory kernel. the resulting theorem identifies every successor action's conditional law with the explicit finite-action policy. lean module compiled","shard":"modules/7b2b37874141e198.json"},{"id":"module:BanditRLProof.Exp3SampledHedge","label":"Exp3SampledHedge","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3SampledHedge","description":"This module identifies the recursively accumulated sampledHistoryScore with the deterministic cumulative-loss surface used by Exp3HedgeRegret. It also identifies the corresponding pure exponential-weights distribution and the exploration-mixed trajectory probability. The final theorem specializes the deterministic second-order Hedge bound to one concrete sampled trajectory.","url":"../modules/banditrlproof-exp3sampledhedge/index.html","parent":"chapter:exp3","order":350,"meta":[["Source","BanditRLProof/Exp3SampledHedge.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3sampledhedge banditrlproof.exp3sampledhedge this module identifies the recursively accumulated sampledhistoryscore with the deterministic cumulative-loss surface used by exp3hedgeregret. it also identifies the corresponding pure exponential-weights distribution and the exploration-mixed trajectory probability. the final theorem specializes the deterministic second-order hedge bound to one concrete sampled trajectory. lean module compiled","shard":"modules/3a1d59db36248686.json"},{"id":"module:BanditRLProof.Exp3SampledHistoryScore","label":"Exp3SampledHistoryScore","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3SampledHistoryScore","description":"This module closes the input-score boundary of Exp3RecursiveTrajectory for real-valued observed losses. The score at an inclusive finite history is the previous score plus the importance-weighted loss of the newly observed pair. The probability in that increment is exactly the exploration-mixed policy computed from the preceding score and history prefix.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html","parent":"chapter:exp3","order":351,"meta":[["Source","BanditRLProof/Exp3SampledHistoryScore.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3sampledhistoryscore banditrlproof.exp3sampledhistoryscore this module closes the input-score boundary of exp3recursivetrajectory for real-valued observed losses. the score at an inclusive finite history is the previous score plus the importance-weighted loss of the newly observed pair. the probability in that increment is exactly the exploration-mixed policy computed from the preceding score and history prefix. lean module compiled","shard":"modules/5c770646523f2639.json"},{"id":"module:BanditRLProof.Exp3ScoreRegularity","label":"Exp3ScoreRegularity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3ScoreRegularity","description":"This module discharges the measurable-score and integrability premises of the generated one-round EXP3 action process. A uniform positive probability floor and measurable losses in [0, 1] give explicit pointwise bounds for the armwise, mixed first-moment, and mixed second-moment scores.","url":"../modules/banditrlproof-exp3scoreregularity/index.html","parent":"chapter:exp3","order":352,"meta":[["Source","BanditRLProof/Exp3ScoreRegularity.lean"],["Declarations","20"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3scoreregularity banditrlproof.exp3scoreregularity this module discharges the measurable-score and integrability premises of the generated one-round exp3 action process. a uniform positive probability floor and measurable losses in [0, 1] give explicit pointwise bounds for the armwise, mixed first-moment, and mixed second-moment scores. lean module compiled","shard":"modules/48999943f50a0b47.json"},{"id":"module:BanditRLProof.Exp3UniformRegret","label":"Exp3UniformRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Exp3UniformRegret","description":"This module closes the small-horizon branch left by the tuned square-root theorem. It clips the exploration rate at 1/2, uses the tuned theorem when 4 K log K <= T, and otherwise uses the pathwise [0,1] loss budget.","url":"../modules/banditrlproof-exp3uniformregret/index.html","parent":"chapter:exp3","order":353,"meta":[["Source","BanditRLProof/Exp3UniformRegret.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","EXP3"]],"statement":"","missing":[],"search":"exp3uniformregret banditrlproof.exp3uniformregret this module closes the small-horizon branch left by the tuned square-root theorem. it clips the exploration rate at 1/2, uses the tuned theorem when 4 k log k <= t, and otherwise uses the pathwise [0,1] loss budget. lean module compiled","shard":"modules/12e447357fda0ba1.json"},{"id":"module:BanditRLProof.ExpectationBochnerSums","label":"ExpectationBochnerSums","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationBochnerSums","description":"Thin Mathlib-backed wrappers for finite-sum linearity of the Bochner integral. These are the expectation-level companion to IntegrabilitySums: each summand must be integrable, and then the integral of the finite sum is the finite sum of the integrals.","url":"../modules/banditrlproof-expectationbochnersums/index.html","parent":"chapter:foundations","order":354,"meta":[["Source","BanditRLProof/ExpectationBochnerSums.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationbochnersums banditrlproof.expectationbochnersums thin mathlib-backed wrappers for finite-sum linearity of the bochner integral. these are the expectation-level companion to integrabilitysums: each summand must be integrable, and then the integral of the finite sum is the finite sum of the integrals. lean module compiled","shard":"modules/a7c5ae2e45e6eff4.json"},{"id":"module:BanditRLProof.ExpectationFiniteBanditBounds","label":"ExpectationFiniteBanditBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationFiniteBanditBounds","description":"This module specializes the generic finite-action weighted pull-count budget bound to the canonical finite action type Fin K and Finset.univ.","url":"../modules/banditrlproof-expectationfinitebanditbounds/index.html","parent":"chapter:foundations","order":355,"meta":[["Source","BanditRLProof/ExpectationFiniteBanditBounds.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationfinitebanditbounds banditrlproof.expectationfinitebanditbounds this module specializes the generic finite-action weighted pull-count budget bound to the canonical finite action type fin k and finset.univ. lean module compiled","shard":"modules/69a6cc9ead26a426.json"},{"id":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","label":"ExpectationFiniteBanditModelBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationFiniteBanditModelBounds","description":"This module connects the finite-action ENNReal pull-count budget bound to the current FiniteBanditModel.gap : Fin K -> Rat surface through ENNReal.ofReal. This is a scalar-conversion canary only: because ENNReal.ofReal clamps negative real values to zero, this file does not prove faithfulness for Rat-valued pseudo-regret.","url":"../modules/banditrlproof-expectationfinitebanditmodelbounds/index.html","parent":"chapter:foundations","order":356,"meta":[["Source","BanditRLProof/ExpectationFiniteBanditModelBounds.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationfinitebanditmodelbounds banditrlproof.expectationfinitebanditmodelbounds this module connects the finite-action ennreal pull-count budget bound to the current finitebanditmodel.gap : fin k -> rat surface through ennreal.ofreal. this is a scalar-conversion canary only: because ennreal.ofreal clamps negative real values to zero, this file does not prove faithfulness for rat-valued pseudo-regret. lean module compiled","shard":"modules/2070dac73dddce67.json"},{"id":"module:BanditRLProof.ExpectationFoundation","label":"ExpectationFoundation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationFoundation","description":"This module introduces the first integration canary through the lower Lebesgue integral of a pull-event indicator. It deliberately avoids Bochner expectation, probability measures, conditional expectation, filtrations, kernels, and concentration assumptions.","url":"../modules/banditrlproof-expectationfoundation/index.html","parent":"chapter:foundations","order":357,"meta":[["Source","BanditRLProof/ExpectationFoundation.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationfoundation banditrlproof.expectationfoundation this module introduces the first integration canary through the lower lebesgue integral of a pull-event indicator. it deliberately avoids bochner expectation, probability measures, conditional expectation, filtrations, kernels, and concentration assumptions. lean module compiled","shard":"modules/6fbaa91da864ae05.json"},{"id":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","label":"ExpectationPseudoRegretOfRealBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationPseudoRegretOfRealBounds","description":"This module lifts the pointwise scalar/model pseudo-regret faithfulness bridge to the existing ENNReal lower-integral model-gap budget bound. It remains an ENNReal.ofReal lower-integral theorem, not a Rat-valued or Bochner expected regret statement.","url":"../modules/banditrlproof-expectationpseudoregretofrealbounds/index.html","parent":"chapter:foundations","order":358,"meta":[["Source","BanditRLProof/ExpectationPseudoRegretOfRealBounds.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationpseudoregretofrealbounds banditrlproof.expectationpseudoregretofrealbounds this module lifts the pointwise scalar/model pseudo-regret faithfulness bridge to the existing ennreal lower-integral model-gap budget bound. it remains an ennreal.ofreal lower-integral theorem, not a rat-valued or bochner expected regret statement. lean module compiled","shard":"modules/4aea156c48d03fe0.json"},{"id":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","label":"ExpectationPseudoRegretRatBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationPseudoRegretRatBounds","description":"This module adapts the lower-integral ENNReal.ofReal pseudo-regret bound to a more natural Rat-valued model-gap nonnegativity contract, then discharges that contract from the local finite-bandit model invariant.","url":"../modules/banditrlproof-expectationpseudoregretratbounds/index.html","parent":"chapter:foundations","order":359,"meta":[["Source","BanditRLProof/ExpectationPseudoRegretRatBounds.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationpseudoregretratbounds banditrlproof.expectationpseudoregretratbounds this module adapts the lower-integral ennreal.ofreal pseudo-regret bound to a more natural rat-valued model-gap nonnegativity contract, then discharges that contract from the local finite-bandit model invariant. lean module compiled","shard":"modules/0f71a05fd49771da.json"},{"id":"module:BanditRLProof.ExpectationPullCount","label":"ExpectationPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationPullCount","description":"This module connects the ENNReal finite-sum lower-integral bridge back to the local recursive pullCount quantity. It remains before Bochner expectation, expected regret, filtrations, kernels, or concentration.","url":"../modules/banditrlproof-expectationpullcount/index.html","parent":"chapter:foundations","order":360,"meta":[["Source","BanditRLProof/ExpectationPullCount.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationpullcount banditrlproof.expectationpullcount this module connects the ennreal finite-sum lower-integral bridge back to the local recursive pullcount quantity. it remains before bochner expectation, expected regret, filtrations, kernels, or concentration. lean module compiled","shard":"modules/766c259f05a14bce.json"},{"id":"module:BanditRLProof.ExpectationPullCountBounds","label":"ExpectationPullCountBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationPullCountBounds","description":"This module proves the first probability-measure corollary of the local pull-count lower-integral identity. It stays in ENNReal and only uses the probability mass bound for measurable sets.","url":"../modules/banditrlproof-expectationpullcountbounds/index.html","parent":"chapter:foundations","order":361,"meta":[["Source","BanditRLProof/ExpectationPullCountBounds.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationpullcountbounds banditrlproof.expectationpullcountbounds this module proves the first probability-measure corollary of the local pull-count lower-integral identity. it stays in ennreal and only uses the probability mass bound for measurable sets. lean module compiled","shard":"modules/610130ea3000fa86.json"},{"id":"module:BanditRLProof.ExpectationRegretPullCount","label":"ExpectationRegretPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationRegretPullCount","description":"This module lifts the deterministic REGRET-PULLCOUNT equality to a Real-valued Bochner expectation statement. It stays at the bookkeeping layer: the only probabilistic regularity assumption is integrability of each finite horizon pull-count random variable after casting to Real.","url":"../modules/banditrlproof-expectationregretpullcount/index.html","parent":"chapter:foundations","order":362,"meta":[["Source","BanditRLProof/ExpectationRegretPullCount.lean"],["Declarations","4"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationregretpullcount banditrlproof.expectationregretpullcount this module lifts the deterministic regret-pullcount equality to a real-valued bochner expectation statement. it stays at the bookkeeping layer: the only probabilistic regularity assumption is integrability of each finite horizon pull-count random variable after casting to real. lean module compiled","shard":"modules/97cd0850e41034cf.json"},{"id":"module:BanditRLProof.ExpectationSums","label":"ExpectationSums","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationSums","description":"This module proves the finite-sum lower-integral bridge for pull-event indicators. It stays in ENNReal and arbitrary measures, before Bochner expectation, pull-count identities, filtrations, kernels, or concentration.","url":"../modules/banditrlproof-expectationsums/index.html","parent":"chapter:foundations","order":363,"meta":[["Source","BanditRLProof/ExpectationSums.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationsums banditrlproof.expectationsums this module proves the finite-sum lower-integral bridge for pull-event indicators. it stays in ennreal and arbitrary measures, before bochner expectation, pull-count identities, filtrations, kernels, or concentration. lean module compiled","shard":"modules/9048db3abdecc673.json"},{"id":"module:BanditRLProof.ExpectationWeightedPullCount","label":"ExpectationWeightedPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationWeightedPullCount","description":"This module proves the nonnegative weighted-count lower-integral bridge. It remains in ENNReal, with an arbitrary finite action set and arbitrary nonnegative gap weights, before any Rat/Real or Bochner-expectation route.","url":"../modules/banditrlproof-expectationweightedpullcount/index.html","parent":"chapter:foundations","order":364,"meta":[["Source","BanditRLProof/ExpectationWeightedPullCount.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationweightedpullcount banditrlproof.expectationweightedpullcount this module proves the nonnegative weighted-count lower-integral bridge. it remains in ennreal, with an arbitrary finite action set and arbitrary nonnegative gap weights, before any rat/real or bochner-expectation route. lean module compiled","shard":"modules/2b7426220aad76b5.json"},{"id":"module:BanditRLProof.ExpectationWeightedPullCountBounds","label":"ExpectationWeightedPullCountBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ExpectationWeightedPullCountBounds","description":"This module proves the finite weighted-count budget bound under a probability measure. It stays in ENNReal; no Rat/Real, Bochner expectation, filtration, kernel, or concentration interface is selected here.","url":"../modules/banditrlproof-expectationweightedpullcountbounds/index.html","parent":"chapter:foundations","order":365,"meta":[["Source","BanditRLProof/ExpectationWeightedPullCountBounds.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"expectationweightedpullcountbounds banditrlproof.expectationweightedpullcountbounds this module proves the finite weighted-count budget bound under a probability measure. it stays in ennreal; no rat/real, bochner expectation, filtration, kernel, or concentration interface is selected here. lean module compiled","shard":"modules/9c04fa97f43b1e61.json"},{"id":"module:BanditRLProof.FTRLOneStep","label":"FTRLOneStep","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FTRLOneStep","description":"This module records the deterministic optimization primitive used by FTRL/OMD routes. It consumes an explicit minimizer certificate for the regularized finite-action objective and returns the one-step linear-loss inequality against any feasible comparator.","url":"../modules/banditrlproof-ftrlonestep/index.html","parent":"chapter:tsallis","order":366,"meta":[["Source","BanditRLProof/FTRLOneStep.lean"],["Declarations","6"],["Project imports","0"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"ftrlonestep banditrlproof.ftrlonestep this module records the deterministic optimization primitive used by ftrl/omd routes. it consumes an explicit minimizer certificate for the regularized finite-action objective and returns the one-step linear-loss inequality against any feasible comparator. lean module compiled","shard":"modules/6ac6b62dc49dfa98.json"},{"id":"module:BanditRLProof.FiniteArmRewardKernelLaw","label":"FiniteArmRewardKernelLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FiniteArmRewardKernelLaw","description":"This module packages action-indexed probability laws into the centered reward kernel contract shared by bandit algorithms. It is deliberately independent of ETC, UCB, or any trajectory construction.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html","parent":"chapter:probability","order":367,"meta":[["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"finitearmrewardkernellaw banditrlproof.finitearmrewardkernellaw this module packages action-indexed probability laws into the centered reward kernel contract shared by bandit algorithms. it is deliberately independent of etc, ucb, or any trajectory construction. lean module compiled","shard":"modules/a4afe9b44c4fbe74.json"},{"id":"module:BanditRLProof.FiniteBanditModelInvariants","label":"FiniteBanditModelInvariants","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModelInvariants","description":"This module contains model-semantic facts about the local FiniteBanditModel.bestArm selector. It stays below regret, expectation, filtrations, kernels, and concentration.","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html","parent":"chapter:foundations","order":368,"meta":[["Source","BanditRLProof/FiniteBanditModelInvariants.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"finitebanditmodelinvariants banditrlproof.finitebanditmodelinvariants this module contains model-semantic facts about the local finitebanditmodel.bestarm selector. it stays below regret, expectation, filtrations, kernels, and concentration. lean module compiled","shard":"modules/f277ccb6e7cb9000.json"},{"id":"module:BanditRLProof.FiniteContextVarianceProxy","label":"FiniteContextVarianceProxy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FiniteContextVarianceProxy","description":"This module computes one common NNReal proxy for a finite context space and finite arm set. It is independent of any particular bandit algorithm.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html","parent":"chapter:probability","order":369,"meta":[["Source","BanditRLProof/FiniteContextVarianceProxy.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"finitecontextvarianceproxy banditrlproof.finitecontextvarianceproxy this module computes one common nnreal proxy for a finite context space and finite arm set. it is independent of any particular bandit algorithm. lean module compiled","shard":"modules/6fb7932810277bae.json"},{"id":"module:BanditRLProof.FiniteGapCutoff","label":"FiniteGapCutoff","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FiniteGapCutoff","description":"A cutoff layer-cake envelope that pays the baseline once per observation.","url":"../modules/banditrlproof-finitegapcutoff/index.html","parent":"chapter:foundations","order":370,"meta":[["Source","BanditRLProof/FiniteGapCutoff.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"finitegapcutoff banditrlproof.finitegapcutoff a cutoff layer-cake envelope that pays the baseline once per observation. lean module compiled","shard":"modules/d903dd5bd78971a1.json"},{"id":"module:BanditRLProof.FiniteGapLayerCake","label":"FiniteGapLayerCake","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake","description":"Finite gap layer-cake identities and a refined integral envelope.","url":"../modules/banditrlproof-finitegaplayercake/index.html","parent":"chapter:foundations","order":371,"meta":[["Source","BanditRLProof/FiniteGapLayerCake.lean"],["Declarations","5"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"finitegaplayercake banditrlproof.finitegaplayercake finite gap layer-cake identities and a refined integral envelope. lean module compiled","shard":"modules/147f959a8571008f.json"},{"id":"module:BanditRLProof.FiniteRealArgmax","label":"FiniteRealArgmax","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax","description":"This module provides a fixed-enumeration maximizer for Real-valued scores on a nonempty finite type. Unlike a bare Classical.choose over existence of a maximum, the explicit fold remains measurable when every score coordinate is measurable in an external parameter.","url":"../modules/banditrlproof-finiterealargmax/index.html","parent":"chapter:foundations","order":372,"meta":[["Source","BanditRLProof/FiniteRealArgmax.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"finiterealargmax banditrlproof.finiterealargmax this module provides a fixed-enumeration maximizer for real-valued scores on a nonempty finite type. unlike a bare classical.choose over existence of a maximum, the explicit fold remains measurable when every score coordinate is measurable in an external parameter. lean module compiled","shard":"modules/aa256c808ed3b441.json"},{"id":"module:BanditRLProof.HOOCantorModel","label":"HOOCantorModel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOCantorModel","description":"An infinite-arm, noisy HOO model on binary sequences.","url":"../modules/banditrlproof-hoocantormodel/index.html","parent":"chapter:foundations","order":373,"meta":[["Source","BanditRLProof/HOOCantorModel.lean"],["Declarations","29"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoocantormodel banditrlproof.hoocantormodel an infinite-arm, noisy hoo model on binary sequences. lean module compiled","shard":"modules/5ed75fc881ffdeb7.json"},{"id":"module:BanditRLProof.HOOCantorRate","label":"HOOCantorRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOCantorRate","description":"A conservative finite dimension certificate and full-rate instantiation for the infinite-arm noisy Cantor model. The certificate is an upper bound, not a claim that its exact near-optimality dimension equals two.","url":"../modules/banditrlproof-hoocantorrate/index.html","parent":"chapter:foundations","order":374,"meta":[["Source","BanditRLProof/HOOCantorRate.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoocantorrate banditrlproof.hoocantorrate a conservative finite dimension certificate and full-rate instantiation for the infinite-arm noisy cantor model. the certificate is an upper bound, not a claim that its exact near-optimality dimension equals two. lean module compiled","shard":"modules/74a1f773f2dc1a6f.json"},{"id":"module:BanditRLProof.HOODimension","label":"HOODimension","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOODimension","description":"Definition 5 with the explicit extended-real log(0)=-infinity convention frozen in LIPSCHITZ-HOO-CONTRACT.md before this implementation.","url":"../modules/banditrlproof-hoodimension/index.html","parent":"chapter:foundations","order":375,"meta":[["Source","BanditRLProof/HOODimension.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoodimension banditrlproof.hoodimension definition 5 with the explicit extended-real log(0)=-infinity convention frozen in lipschitz-hoo-contract.md before this implementation. lean module compiled","shard":"modules/ce32616e42a5f606.json"},{"id":"module:BanditRLProof.HOOGeometry","label":"HOOGeometry","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOGeometry","description":"Source A2 and Lemma 3 for general dissimilarities. No metric triangle inequality, maximizer, or attainment of a regional supremum is assumed.","url":"../modules/banditrlproof-hoogeometry/index.html","parent":"chapter:foundations","order":376,"meta":[["Source","BanditRLProof/HOOGeometry.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoogeometry banditrlproof.hoogeometry source a2 and lemma 3 for general dissimilarities. no metric triangle inequality, maximizer, or attainment of a regional supremum is assumed. lean module compiled","shard":"modules/6b3741ed62ab0f12.json"},{"id":"module:BanditRLProof.HOOLevels","label":"HOOLevels","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOLevels","description":"Finite complete levels of the infinite HOO covering tree.","url":"../modules/banditrlproof-hoolevels/index.html","parent":"chapter:foundations","order":377,"meta":[["Source","BanditRLProof/HOOLevels.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoolevels banditrlproof.hoolevels finite complete levels of the infinite hoo covering tree. lean module compiled","shard":"modules/794e45f5a86bd3d2.json"},{"id":"module:BanditRLProof.HOOModel","label":"HOOModel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOModel","description":"Source tree-of-coverings assumptions and the actual arm/reward realization. The infinite arm space is not replaced by a finite discretization.","url":"../modules/banditrlproof-hoomodel/index.html","parent":"chapter:foundations","order":378,"meta":[["Source","BanditRLProof/HOOModel.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoomodel banditrlproof.hoomodel source tree-of-coverings assumptions and the actual arm/reward realization. the infinite arm space is not replaced by a finite discretization. lean module compiled","shard":"modules/1509ac4b769586f6.json"},{"id":"module:BanditRLProof.HOOOptimalBranch","label":"HOOOptimalBranch","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOOptimalBranch","description":"A deterministic branch preserving regional suprema. No maximizer or attainment of any regional supremum is assumed.","url":"../modules/banditrlproof-hoooptimalbranch/index.html","parent":"chapter:foundations","order":379,"meta":[["Source","BanditRLProof/HOOOptimalBranch.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoooptimalbranch banditrlproof.hoooptimalbranch a deterministic branch preserving regional suprema. no maximizer or attainment of any regional supremum is assumed. lean module compiled","shard":"modules/18b477f5f279c715.json"},{"id":"module:BanditRLProof.HOOPacking","label":"HOOPacking","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOPacking","description":"Source Definition 4: whole open balls, not merely their centers, are contained in the target set. Under A1 every positive-radius packing is finite.","url":"../modules/banditrlproof-hoopacking/index.html","parent":"chapter:foundations","order":380,"meta":[["Source","BanditRLProof/HOOPacking.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoopacking banditrlproof.hoopacking source definition 4: whole open balls, not merely their centers, are contained in the target set. under a1 every positive-radius packing is finite. lean module compiled","shard":"modules/e619d0a70af0d280.json"},{"id":"module:BanditRLProof.HOOPartition","label":"HOOPartition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOPartition","description":"Source Theorem 6's actual covering-tree partition. Boundary nodes are children of near-optimal parents which fail the next-level near-optimal test.","url":"../modules/banditrlproof-hoopartition/index.html","parent":"chapter:foundations","order":381,"meta":[["Source","BanditRLProof/HOOPartition.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hoopartition banditrlproof.hoopartition source theorem 6's actual covering-tree partition. boundary nodes are children of near-optimal parents which fail the next-level near-optimal test. lean module compiled","shard":"modules/29186340eeacbfed.json"},{"id":"module:BanditRLProof.HOOTailSum","label":"HOOTailSum","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HOOTailSum","description":"Summable confidence failures after the HOO path/depth union.","url":"../modules/banditrlproof-hootailsum/index.html","parent":"chapter:foundations","order":382,"meta":[["Source","BanditRLProof/HOOTailSum.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"hootailsum banditrlproof.hootailsum summable confidence failures after the hoo path/depth union. lean module compiled","shard":"modules/122daf1ff2aaebcb.json"},{"id":"module:BanditRLProof.HeavyTailArmLaw","label":"HeavyTailArmLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailArmLaw","description":"Stationary raw-moment reward laws supply every fixed-coordinate hypothesis.","url":"../modules/banditrlproof-heavytailarmlaw/index.html","parent":"chapter:foundations","order":383,"meta":[["Source","BanditRLProof/HeavyTailArmLaw.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailarmlaw banditrlproof.heavytailarmlaw stationary raw-moment reward laws supply every fixed-coordinate hypothesis. lean module compiled","shard":"modules/905d14b915b1c1b2.json"},{"id":"module:BanditRLProof.HeavyTailClippedConfidence","label":"HeavyTailClippedConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailClippedConfidence","description":"Clean clipped-mean confidence from raw moments, through shared MGF interfaces.","url":"../modules/banditrlproof-heavytailclippedconfidence/index.html","parent":"chapter:foundations","order":384,"meta":[["Source","BanditRLProof/HeavyTailClippedConfidence.lean"],["Declarations","3"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailclippedconfidence banditrlproof.heavytailclippedconfidence clean clipped-mean confidence from raw moments, through shared mgf interfaces. lean module compiled","shard":"modules/51d6d2fcfa2e68d3.json"},{"id":"module:BanditRLProof.HeavyTailClippedMoments","label":"HeavyTailClippedMoments","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailClippedMoments","description":"Clean clipping bias and second moments, produced from raw moments.","url":"../modules/banditrlproof-heavytailclippedmoments/index.html","parent":"chapter:foundations","order":385,"meta":[["Source","BanditRLProof/HeavyTailClippedMoments.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailclippedmoments banditrlproof.heavytailclippedmoments clean clipping bias and second moments, produced from raw moments. lean module compiled","shard":"modules/e297979e39992423.json"},{"id":"module:BanditRLProof.HeavyTailClippedScheduled","label":"HeavyTailClippedScheduled","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailClippedScheduled","description":"The actual algorithm's scheduled radius, produced from raw moments.","url":"../modules/banditrlproof-heavytailclippedscheduled/index.html","parent":"chapter:foundations","order":386,"meta":[["Source","BanditRLProof/HeavyTailClippedScheduled.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailclippedscheduled banditrlproof.heavytailclippedscheduled the actual algorithm's scheduled radius, produced from raw moments. lean module compiled","shard":"modules/36c1a5551903455f.json"},{"id":"module:BanditRLProof.HeavyTailClippedTransfer","label":"HeavyTailClippedTransfer","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailClippedTransfer","description":"Raw-moment confidence under a pathwise corruption budget on the consumed prefix. The count and corruption may depend on the entire outcome.","url":"../modules/banditrlproof-heavytailclippedtransfer/index.html","parent":"chapter:foundations","order":387,"meta":[["Source","BanditRLProof/HeavyTailClippedTransfer.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailclippedtransfer banditrlproof.heavytailclippedtransfer raw-moment confidence under a pathwise corruption budget on the consumed prefix. the count and corruption may depend on the entire outcome. lean module compiled","shard":"modules/e380497bdb656e0d.json"},{"id":"module:BanditRLProof.HeavyTailClipping","label":"HeavyTailClipping","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailClipping","description":"Reserved transfer: winsorized estimators under an L1 corruption budget. The budget is on the consumed prefix. No corruption-robust regret claim.","url":"../modules/banditrlproof-heavytailclipping/index.html","parent":"chapter:foundations","order":388,"meta":[["Source","BanditRLProof/HeavyTailClipping.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailclipping banditrlproof.heavytailclipping reserved transfer: winsorized estimators under an l1 corruption budget. the budget is on the consumed prefix. no corruption-robust regret claim. lean module compiled","shard":"modules/403e67ada2edd0be.json"},{"id":"module:BanditRLProof.HeavyTailConfidence","label":"HeavyTailConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailConfidence","description":"Two-sided, tuned heavy-tail confidence. This closes the fixed-prefix probability producer before the causal policy and adaptive-count assembly.","url":"../modules/banditrlproof-heavytailconfidence/index.html","parent":"chapter:foundations","order":389,"meta":[["Source","BanditRLProof/HeavyTailConfidence.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailconfidence banditrlproof.heavytailconfidence two-sided, tuned heavy-tail confidence. this closes the fixed-prefix probability producer before the causal policy and adaptive-count assembly. lean module compiled","shard":"modules/0f956effc6634cf3.json"},{"id":"module:BanditRLProof.HeavyTailFixedTilt","label":"HeavyTailFixedTilt","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailFixedTilt","description":"A reusable bounded, centered, second-moment MGF producer. The existing EXP3 exponential-remainder leaf is genuinely reused for a new probability law. This is not yet the independent-sum or adaptive-policy concentration theorem.","url":"../modules/banditrlproof-heavytailfixedtilt/index.html","parent":"chapter:foundations","order":390,"meta":[["Source","BanditRLProof/HeavyTailFixedTilt.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailfixedtilt banditrlproof.heavytailfixedtilt a reusable bounded, centered, second-moment mgf producer. the existing exp3 exponential-remainder leaf is genuinely reused for a new probability law. this is not yet the independent-sum or adaptive-policy concentration theorem. lean module compiled","shard":"modules/d2e8745d536fa2e6.json"},{"id":"module:BanditRLProof.HeavyTailGapThreshold","label":"HeavyTailGapThreshold","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailGapThreshold","description":"An explicit integer sample budget makes twice the chosen radius smaller than the gap.","url":"../modules/banditrlproof-heavytailgapthreshold/index.html","parent":"chapter:foundations","order":391,"meta":[["Source","BanditRLProof/HeavyTailGapThreshold.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailgapthreshold banditrlproof.heavytailgapthreshold an explicit integer sample budget makes twice the chosen radius smaller than the gap. lean module compiled","shard":"modules/5cdaeedbc5ec6757.json"},{"id":"module:BanditRLProof.HeavyTailPowerSum","label":"HeavyTailPowerSum","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailPowerSum","description":"Finite fractional-power sum needed for the sample-index truncation bias.","url":"../modules/banditrlproof-heavytailpowersum/index.html","parent":"chapter:foundations","order":392,"meta":[["Source","BanditRLProof/HeavyTailPowerSum.lean"],["Declarations","2"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailpowersum banditrlproof.heavytailpowersum finite fractional-power sum needed for the sample-index truncation bias. lean module compiled","shard":"modules/f7af9e80a1514c97.json"},{"id":"module:BanditRLProof.HeavyTailScheduledConfidence","label":"HeavyTailScheduledConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailScheduledConfidence","description":"The actual algorithm's scheduled radius, produced from raw moments.","url":"../modules/banditrlproof-heavytailscheduledconfidence/index.html","parent":"chapter:foundations","order":393,"meta":[["Source","BanditRLProof/HeavyTailScheduledConfidence.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailscheduledconfidence banditrlproof.heavytailscheduledconfidence the actual algorithm's scheduled radius, produced from raw moments. lean module compiled","shard":"modules/7e10c23dd66ba6aa.json"},{"id":"module:BanditRLProof.HeavyTailSourceConfidence","label":"HeavyTailSourceConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailSourceConfidence","description":"Arbitrary-log-confidence sample-index truncation. The target is the constant four confidence radius of BCL 2013 Lemma 1; algorithm regret remains separate.","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html","parent":"chapter:foundations","order":394,"meta":[["Source","BanditRLProof/HeavyTailSourceConfidence.lean"],["Declarations","13"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailsourceconfidence banditrlproof.heavytailsourceconfidence arbitrary-log-confidence sample-index truncation. the target is the constant four confidence radius of bcl 2013 lemma 1; algorithm regret remains separate. lean module compiled","shard":"modules/2ec6f853c774d792.json"},{"id":"module:BanditRLProof.HeavyTailSourceGap","label":"HeavyTailSourceGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailSourceGap","description":"Explicit corrected gap cutoff for the unchanged source-parameter policy.","url":"../modules/banditrlproof-heavytailsourcegap/index.html","parent":"chapter:foundations","order":395,"meta":[["Source","BanditRLProof/HeavyTailSourceGap.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailsourcegap banditrlproof.heavytailsourcegap explicit corrected gap cutoff for the unchanged source-parameter policy. lean module compiled","shard":"modules/0b05d49bb6e92bb8.json"},{"id":"module:BanditRLProof.HeavyTailSourceSchedule","label":"HeavyTailSourceSchedule","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailSourceSchedule","description":"Original source log schedule and signed adaptive-prefix confidence. Counts are handled by a finite union, never by asserting selected samples IID.","url":"../modules/banditrlproof-heavytailsourceschedule/index.html","parent":"chapter:foundations","order":396,"meta":[["Source","BanditRLProof/HeavyTailSourceSchedule.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailsourceschedule banditrlproof.heavytailsourceschedule original source log schedule and signed adaptive-prefix confidence. counts are handled by a finite union, never by asserting selected samples iid. lean module compiled","shard":"modules/548e0673851b3aee.json"},{"id":"module:BanditRLProof.HeavyTailTailSum","label":"HeavyTailTailSum","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailTailSum","description":"Finite uniform budget for the actual two-arm, two-sided confidence union.","url":"../modules/banditrlproof-heavytailtailsum/index.html","parent":"chapter:foundations","order":397,"meta":[["Source","BanditRLProof/HeavyTailTailSum.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailtailsum banditrlproof.heavytailtailsum finite uniform budget for the actual two-arm, two-sided confidence union. lean module compiled","shard":"modules/3f46588cba47b9f6.json"},{"id":"module:BanditRLProof.HeavyTailTruncation","label":"HeavyTailTruncation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailTruncation","description":"Raw absolute moments; no sub-Gaussian assumption. These are producer leaves, not a complete robust-UCB regret theorem. Thresholds may depend on sample index and on the evaluation round (by choosing a different transform each round).","url":"../modules/banditrlproof-heavytailtruncation/index.html","parent":"chapter:foundations","order":398,"meta":[["Source","BanditRLProof/HeavyTailTruncation.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailtruncation banditrlproof.heavytailtruncation raw absolute moments; no sub-gaussian assumption. these are producer leaves, not a complete robust-ucb regret theorem. thresholds may depend on sample index and on the evaluation round (by choosing a different transform each round). lean module compiled","shard":"modules/90be129d6eb38816.json"},{"id":"module:BanditRLProof.HeavyTailTuning","label":"HeavyTailTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailTuning","description":"Algebraic tuning for the sample-index threshold; all exponents remain real.","url":"../modules/banditrlproof-heavytailtuning/index.html","parent":"chapter:foundations","order":399,"meta":[["Source","BanditRLProof/HeavyTailTuning.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailtuning banditrlproof.heavytailtuning algebraic tuning for the sample-index threshold; all exponents remain real. lean module compiled","shard":"modules/52d27245a640134c.json"},{"id":"module:BanditRLProof.HeavyTailUnshiftedMGF","label":"HeavyTailUnshiftedMGF","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HeavyTailUnshiftedMGF","description":"Centering after the exponential bound preserves the raw-variable tilt range. This uses a raw second moment, not the variance of the centered variable.","url":"../modules/banditrlproof-heavytailunshiftedmgf/index.html","parent":"chapter:foundations","order":400,"meta":[["Source","BanditRLProof/HeavyTailUnshiftedMGF.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"heavytailunshiftedmgf banditrlproof.heavytailunshiftedmgf centering after the exponential bound preserves the raw-variable tilt range. this uses a raw second moment, not the variance of the centered variable. lean module compiled","shard":"modules/3d2a5ac4eb840c7a.json"},{"id":"module:BanditRLProof.HistoryFiltration","label":"HistoryFiltration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.HistoryFiltration","description":"This module gives a narrow project-local filtration canary: the sigma-algebras generated by past action/reward singleton events form a Mathlib filtration. It does not construct kernels, policies, conditional expectations, conditional MGF witnesses, or adaptive regret theorems.","url":"../modules/banditrlproof-historyfiltration/index.html","parent":"chapter:probability","order":401,"meta":[["Source","BanditRLProof/HistoryFiltration.lean"],["Declarations","49"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"historyfiltration banditrlproof.historyfiltration this module gives a narrow project-local filtration canary: the sigma-algebras generated by past action/reward singleton events form a mathlib filtration. it does not construct kernels, policies, conditional expectations, conditional mgf witnesses, or adaptive regret theorems. lean module compiled","shard":"modules/024bfdf00d4ef14a.json"},{"id":"module:BanditRLProof.IndependenceFoundation","label":"IndependenceFoundation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.IndependenceFoundation","description":"This module exposes small Mathlib-backed independence imports under the project namespace. It stays at the product-coordinate source layer: no bandit policy, filtration, conditional expectation, or regret theorem is introduced here.","url":"../modules/banditrlproof-independencefoundation/index.html","parent":"chapter:probability","order":402,"meta":[["Source","BanditRLProof/IndependenceFoundation.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"independencefoundation banditrlproof.independencefoundation this module exposes small mathlib-backed independence imports under the project namespace. it stays at the product-coordinate source layer: no bandit policy, filtration, conditional expectation, or regret theorem is introduced here. lean module compiled","shard":"modules/4fba937e92935ca6.json"},{"id":"module:BanditRLProof.IntegrabilitySums","label":"IntegrabilitySums","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.IntegrabilitySums","description":"Thin Mathlib-backed wrappers for the reusable integrability fact needed by finite regret decompositions: a finite sum of integrable terms is integrable. This module does not state Bochner expectation linearity; that is a separate leaf.","url":"../modules/banditrlproof-integrabilitysums/index.html","parent":"chapter:probability","order":403,"meta":[["Source","BanditRLProof/IntegrabilitySums.lean"],["Declarations","2"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"integrabilitysums banditrlproof.integrabilitysums thin mathlib-backed wrappers for the reusable integrability fact needed by finite regret decompositions: a finite sum of integrable terms is integrable. this module does not state bochner expectation linearity; that is a separate leaf. lean module compiled","shard":"modules/5ad915b30d58d24a.json"},{"id":"module:BanditRLProof.KernelIndependentExtension","label":"KernelIndependentExtension","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.KernelIndependentExtension","description":"This module records the semidirect-product transport needed by stochastic bandit trajectory laws. Sampling from a Markov kernel that only sees a past summary cannot create dependence between its output and a random variable already independent of that summary.","url":"../modules/banditrlproof-kernelindependentextension/index.html","parent":"chapter:probability","order":404,"meta":[["Source","BanditRLProof/KernelIndependentExtension.lean"],["Declarations","6"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"kernelindependentextension banditrlproof.kernelindependentextension this module records the semidirect-product transport needed by stochastic bandit trajectory laws. sampling from a markov kernel that only sees a past summary cannot create dependence between its output and a random variable already independent of that summary. lean module compiled","shard":"modules/cf893df6918904a0.json"},{"id":"module:BanditRLProof.KernelTrajectoryPrefix","label":"KernelTrajectoryPrefix","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.KernelTrajectoryPrefix","description":"The infinite Kernel.traj construction has a finite marginal determined only by the initial law and the step kernels used before the marginal endpoint. These wrappers expose that fact in the form needed by environment-prefix factorizations.","url":"../modules/banditrlproof-kerneltrajectoryprefix/index.html","parent":"chapter:probability","order":405,"meta":[["Source","BanditRLProof/KernelTrajectoryPrefix.lean"],["Declarations","2"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"kerneltrajectoryprefix banditrlproof.kerneltrajectoryprefix the infinite kernel.traj construction has a finite marginal determined only by the initial law and the step kernels used before the marginal endpoint. these wrappers expose that fact in the form needed by environment-prefix factorizations. lean module compiled","shard":"modules/3a808c9e36eff4fd.json"},{"id":"module:BanditRLProof.LeafLemmas","label":"LeafLemmas","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LeafLemmas","description":"These lemmas are the first compiled ABRL leaf library. They deliberately avoid Mathlib imports while exposing stable theorem names that later Mathlib-backed tasks can replace, generalize, or upstream.","url":"../modules/banditrlproof-leaflemmas/index.html","parent":"chapter:foundations","order":406,"meta":[["Source","BanditRLProof/LeafLemmas.lean"],["Declarations","38"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"leaflemmas banditrlproof.leaflemmas these lemmas are the first compiled abrl leaf library. they deliberately avoid mathlib imports while exposing stable theorem names that later mathlib-backed tasks can replace, generalize, or upstream. lean module compiled","shard":"modules/5ecd1e8ad608dabc.json"},{"id":"module:BanditRLProof.Literature","label":"Literature","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Literature","description":"Literature and upstream theorem registry","url":"../modules/banditrlproof-literature/index.html","parent":"chapter:frontier","order":407,"meta":[["Source","BanditRLProof/Literature.lean"],["Declarations","3"],["Project imports","3"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"literature banditrlproof.literature literature and upstream theorem registry lean module compiled","shard":"modules/1da684ee192cb448.json"},{"id":"module:BanditRLProof.LowerBounds.AffinityKL","label":"LowerBounds.AffinityKL","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.AffinityKL","description":"The measure-level Jensen step in the Chapter 14 overlap proof.","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html","parent":"chapter:foundations","order":408,"meta":[["Source","BanditRLProof/LowerBounds/AffinityKL.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.affinitykl banditrlproof.lowerbounds.affinitykl the measure-level jensen step in the chapter 14 overlap proof. lean module compiled","shard":"modules/00f13da6645c930b.json"},{"id":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","label":"LowerBounds.ArithmeticBlockCoding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ArithmeticBlockCoding","description":"Generated source map for BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean.","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html","parent":"chapter:foundations","order":409,"meta":[["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.arithmeticblockcoding banditrlproof.lowerbounds.arithmeticblockcoding generated source map for banditrlproof/lowerbounds/arithmeticblockcoding.lean. lean module compiled","shard":"modules/c863a1e7129c5a32.json"},{"id":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","label":"LowerBounds.ArithmeticIntervals","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ArithmeticIntervals","description":"Generated source map for BanditRLProof/LowerBounds/ArithmeticIntervals.lean.","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html","parent":"chapter:foundations","order":410,"meta":[["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.arithmeticintervals banditrlproof.lowerbounds.arithmeticintervals generated source map for banditrlproof/lowerbounds/arithmeticintervals.lean. lean module compiled","shard":"modules/baa92f181aa335e8.json"},{"id":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","label":"LowerBounds.ArithmeticPrefixCode","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ArithmeticPrefixCode","description":"Generated source map for BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean.","url":"../modules/banditrlproof-lowerbounds-arithmeticprefixcode/index.html","parent":"chapter:foundations","order":411,"meta":[["Source","BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.arithmeticprefixcode banditrlproof.lowerbounds.arithmeticprefixcode generated source map for banditrlproof/lowerbounds/arithmeticprefixcode.lean. lean module compiled","shard":"modules/529bb9fb74a6dab5.json"},{"id":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","label":"LowerBounds.ArithmeticZeroExtension","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ArithmeticZeroExtension","description":"Generated source map for BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean.","url":"../modules/banditrlproof-lowerbounds-arithmeticzeroextension/index.html","parent":"chapter:foundations","order":412,"meta":[["Source","BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.arithmeticzeroextension banditrlproof.lowerbounds.arithmeticzeroextension generated source map for banditrlproof/lowerbounds/arithmeticzeroextension.lean. lean module compiled","shard":"modules/36b23875b3139ebf.json"},{"id":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","label":"LowerBounds.BanditHistoryDataProcessing","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BanditHistoryDataProcessing","description":"This module proves the measurable-observation data-processing leaf needed by the stopping-time extension in Lattimore--Szepesvari, *Bandit Algorithms*, Exercise 15.7. It also specializes the leaf to the compiled deterministic finite-history divergence decomposition of Lemma 15.1.","url":"../modules/banditrlproof-lowerbounds-bandithistorydataprocessing/index.html","parent":"chapter:foundations","order":413,"meta":[["Source","BanditRLProof/LowerBounds/BanditHistoryDataProcessing.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.bandithistorydataprocessing banditrlproof.lowerbounds.bandithistorydataprocessing this module proves the measurable-observation data-processing leaf needed by the stopping-time extension in lattimore--szepesvari, *bandit algorithms*, exercise 15.7. it also specializes the leaf to the compiled deterministic finite-history divergence decomposition of lemma 15.1. lean module compiled","shard":"modules/281da1b93079ff83.json"},{"id":"module:BanditRLProof.LowerBounds.BanditHistoryKL","label":"LowerBounds.BanditHistoryKL","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BanditHistoryKL","description":"This module connects the conditional-kernel KL integral to the repository's kernel-valued HistoryAlgorithm and canonical Ionescu--Tulcea trajectory. The observable history includes every sampled action and reward, so the same possibly randomized nonanticipating policy is shared by both environments.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html","parent":"chapter:foundations","order":414,"meta":[["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean"],["Declarations","32"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.bandithistorykl banditrlproof.lowerbounds.bandithistorykl this module connects the conditional-kernel kl integral to the repository's kernel-valued historyalgorithm and canonical ionescu--tulcea trajectory. the observable history includes every sampled action and reward, so the same possibly randomized nonanticipating policy is shared by both environments. lean module compiled","shard":"modules/4185177a88ce70a9.json"},{"id":"module:BanditRLProof.LowerBounds.BasicIdeas","label":"LowerBounds.BasicIdeas","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BasicIdeas","description":"This module formalizes the semantic and deterministic interfaces developed in Lattimore--Szepesvári, *Bandit Algorithms* (2020), Part IV, Chapter 13.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html","parent":"chapter:foundations","order":415,"meta":[["Source","BanditRLProof/LowerBounds/BasicIdeas.lean"],["Declarations","15"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.basicideas banditrlproof.lowerbounds.basicideas this module formalizes the semantic and deterministic interfaces developed in lattimore--szepesvári, *bandit algorithms* (2020), part iv, chapter 13. lean module compiled","shard":"modules/7efce33e480c5631.json"},{"id":"module:BanditRLProof.LowerBounds.BlockEntropy","label":"LowerBounds.BlockEntropy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BlockEntropy","description":"Generated source map for BanditRLProof/LowerBounds/BlockEntropy.lean.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html","parent":"chapter:foundations","order":416,"meta":[["Source","BanditRLProof/LowerBounds/BlockEntropy.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.blockentropy banditrlproof.lowerbounds.blockentropy generated source map for banditrlproof/lowerbounds/blockentropy.lean. lean module compiled","shard":"modules/766f791ee94508ad.json"},{"id":"module:BanditRLProof.LowerBounds.CodingEntropyBound","label":"LowerBounds.CodingEntropyBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.CodingEntropyBound","description":"Generated source map for BanditRLProof/LowerBounds/CodingEntropyBound.lean.","url":"../modules/banditrlproof-lowerbounds-codingentropybound/index.html","parent":"chapter:foundations","order":417,"meta":[["Source","BanditRLProof/LowerBounds/CodingEntropyBound.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.codingentropybound banditrlproof.lowerbounds.codingentropybound generated source map for banditrlproof/lowerbounds/codingentropybound.lean. lean module compiled","shard":"modules/70bffeb024e0eca8.json"},{"id":"module:BanditRLProof.LowerBounds.CommonDensityKL","label":"LowerBounds.CommonDensityKL","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.CommonDensityKL","description":"Chapter 14 Eq. (14.6), with exact common-density and infinite branches.","url":"../modules/banditrlproof-lowerbounds-commondensitykl/index.html","parent":"chapter:foundations","order":418,"meta":[["Source","BanditRLProof/LowerBounds/CommonDensityKL.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.commondensitykl banditrlproof.lowerbounds.commondensitykl chapter 14 eq. (14.6), with exact common-density and infinite branches. lean module compiled","shard":"modules/63200a032e0cec96.json"},{"id":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","label":"LowerBounds.CommonDensityOverlap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.CommonDensityOverlap","description":"Measure-level overlap in Chapter 14, including the optimal testing event.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html","parent":"chapter:foundations","order":419,"meta":[["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.commondensityoverlap banditrlproof.lowerbounds.commondensityoverlap measure-level overlap in chapter 14, including the optimal testing event. lean module compiled","shard":"modules/a9cb2127ecf6e8de.json"},{"id":"module:BanditRLProof.LowerBounds.CommonDomination","label":"LowerBounds.CommonDomination","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.CommonDomination","description":"Generated source map for BanditRLProof/LowerBounds/CommonDomination.lean.","url":"../modules/banditrlproof-lowerbounds-commondomination/index.html","parent":"chapter:foundations","order":420,"meta":[["Source","BanditRLProof/LowerBounds/CommonDomination.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.commondomination banditrlproof.lowerbounds.commondomination generated source map for banditrlproof/lowerbounds/commondomination.lean. lean module compiled","shard":"modules/fd48243cca1b4852.json"},{"id":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","label":"LowerBounds.ConditionalKernelKL","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ConditionalKernelKL","description":"Mathlib's composition-product chain rule deliberately leaves the conditional term as another measure-level KL divergence because measurability of x \\mapsto klDiv (kappa x) (eta x) is not automatic. For countably generated target spaces, kernel Radon--Nikodym derivatives provide a measurable replacement. This module proves the resulting iterated-lintegral identity.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html","parent":"chapter:foundations","order":421,"meta":[["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean"],["Declarations","10"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.conditionalkernelkl banditrlproof.lowerbounds.conditionalkernelkl mathlib's composition-product chain rule deliberately leaves the conditional term as another measure-level kl divergence because measurability of x \\mapsto kldiv (kappa x) (eta x) is not automatic. for countably generated target spaces, kernel radon--nikodym derivatives provide a measurable replacement. this module proves the resulting iterated-lintegral identity. lean module compiled","shard":"modules/d9b7b3d2d0ef672c.json"},{"id":"module:BanditRLProof.LowerBounds.CrossEntropy","label":"LowerBounds.CrossEntropy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.CrossEntropy","description":"Generated source map for BanditRLProof/LowerBounds/CrossEntropy.lean.","url":"../modules/banditrlproof-lowerbounds-crossentropy/index.html","parent":"chapter:foundations","order":422,"meta":[["Source","BanditRLProof/LowerBounds/CrossEntropy.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.crossentropy banditrlproof.lowerbounds.crossentropy generated source map for banditrlproof/lowerbounds/crossentropy.lean. lean module compiled","shard":"modules/0bfae14292396ff9.json"},{"id":"module:BanditRLProof.LowerBounds.DyadicAddresses","label":"LowerBounds.DyadicAddresses","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.DyadicAddresses","description":"Generated source map for BanditRLProof/LowerBounds/DyadicAddresses.lean.","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html","parent":"chapter:foundations","order":423,"meta":[["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.dyadicaddresses banditrlproof.lowerbounds.dyadicaddresses generated source map for banditrlproof/lowerbounds/dyadicaddresses.lean. lean module compiled","shard":"modules/8f71f7f78637e4d3.json"},{"id":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","label":"LowerBounds.FiniteDiscreteKL","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteDiscreteKL","description":"Generated source map for BanditRLProof/LowerBounds/FiniteDiscreteKL.lean.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html","parent":"chapter:foundations","order":424,"meta":[["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.finitediscretekl banditrlproof.lowerbounds.finitediscretekl generated source map for banditrlproof/lowerbounds/finitediscretekl.lean. lean module compiled","shard":"modules/724de932ea7680fc.json"},{"id":"module:BanditRLProof.LowerBounds.FinitePartitionKL","label":"LowerBounds.FinitePartitionKL","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FinitePartitionKL","description":"The supremum in textbook Eq. (14.5) ranges over every finite measurable partition, represented here by measurable maps to Fin n. Empty cells are harmless. The singular branch below is only one part of the required equivalence with the Radon--Nikodym definition.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html","parent":"chapter:foundations","order":425,"meta":[["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.finitepartitionkl banditrlproof.lowerbounds.finitepartitionkl the supremum in textbook eq. (14.5) ranges over every finite measurable partition, represented here by measurable maps to fin n. empty cells are harmless. the singular branch below is only one part of the required equivalence with the radon--nikodym definition. lean module compiled","shard":"modules/6dea14d3460ba942.json"},{"id":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","label":"LowerBounds.FinitePartitionKLRecovery","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FinitePartitionKLRecovery","description":"The finite-discretisation definition of KL equals the RN definition.","url":"../modules/banditrlproof-lowerbounds-finitepartitionklrecovery/index.html","parent":"chapter:foundations","order":426,"meta":[["Source","BanditRLProof/LowerBounds/FinitePartitionKLRecovery.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.finitepartitionklrecovery banditrlproof.lowerbounds.finitepartitionklrecovery the finite-discretisation definition of kl equals the rn definition. lean module compiled","shard":"modules/096ca9a547e4399c.json"},{"id":"module:BanditRLProof.LowerBounds.FixedLengthCoding","label":"LowerBounds.FixedLengthCoding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FixedLengthCoding","description":"Generated source map for BanditRLProof/LowerBounds/FixedLengthCoding.lean.","url":"../modules/banditrlproof-lowerbounds-fixedlengthcoding/index.html","parent":"chapter:foundations","order":427,"meta":[["Source","BanditRLProof/LowerBounds/FixedLengthCoding.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.fixedlengthcoding banditrlproof.lowerbounds.fixedlengthcoding generated source map for banditrlproof/lowerbounds/fixedlengthcoding.lean. lean module compiled","shard":"modules/83fdd270b00e8a08.json"},{"id":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","label":"LowerBounds.GaussianHypothesisTesting","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.GaussianHypothesisTesting","description":"This module formalizes the distribution-level threshold test used at the start of Lattimore--Szepesvári, Chapter 13.1. The source observes that a sample mean from n independent unit-variance Gaussian observations has variance 1 / n. We expose that Gaussian mean-observation law, identify the two error events, and prove the standard Chernoff upper bound exp (-n * gap^2 / 8) for both hypotheses and therefore for their…","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html","parent":"chapter:foundations","order":428,"meta":[["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean"],["Declarations","22"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.gaussianhypothesistesting banditrlproof.lowerbounds.gaussianhypothesistesting this module formalizes the distribution-level threshold test used at the start of lattimore--szepesvári, chapter 13.1. the source observes that a sample mean from n independent unit-variance gaussian observations has variance 1 / n. we expose that gaussian mean-observation law, identify the two error events, and prove the standard chernoff upper bound exp (-n * gap^2 / 8) for both hypotheses and therefore for their maximum error probability. lean module compiled","shard":"modules/6d573e942ee2905b.json"},{"id":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","label":"LowerBounds.GaussianMillsRatio","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.GaussianMillsRatio","description":"Analytic comparison functions for Chapter 13 Eq. (13.4). Both exact improper-integral bounds compile. Probability rescaling is separate.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html","parent":"chapter:foundations","order":429,"meta":[["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean"],["Declarations","26"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.gaussianmillsratio banditrlproof.lowerbounds.gaussianmillsratio analytic comparison functions for chapter 13 eq. (13.4). both exact improper-integral bounds compile. probability rescaling is separate. lean module compiled","shard":"modules/578fb1a5c76d9667.json"},{"id":"module:BanditRLProof.LowerBounds.GaussianMinimax","label":"LowerBounds.GaussianMinimax","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.GaussianMinimax","description":"This module formalizes the proof of Lattimore--Szepesvari, Theorem 15.2, on the repository's canonical finite-history law. The local history index is inclusive: lastRound represents exactly lastRound + 1 observations.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html","parent":"chapter:foundations","order":430,"meta":[["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean"],["Declarations","51"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.gaussianminimax banditrlproof.lowerbounds.gaussianminimax this module formalizes the proof of lattimore--szepesvari, theorem 15.2, on the repository's canonical finite-history law. the local history index is inclusive: lastround represents exactly lastround + 1 observations. lean module compiled","shard":"modules/585f4c0b4877ae79.json"},{"id":"module:BanditRLProof.LowerBounds.GaussianTesting","label":"LowerBounds.GaussianTesting","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.GaussianTesting","description":"Generated source map for BanditRLProof/LowerBounds/GaussianTesting.lean.","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html","parent":"chapter:foundations","order":431,"meta":[["Source","BanditRLProof/LowerBounds/GaussianTesting.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.gaussiantesting banditrlproof.lowerbounds.gaussiantesting generated source map for banditrlproof/lowerbounds/gaussiantesting.lean. lean module compiled","shard":"modules/b880d6ac0ebab727.json"},{"id":"module:BanditRLProof.LowerBounds.HighProbability","label":"LowerBounds.HighProbability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.HighProbability","description":"This module formalizes the source-faithful threshold surfaces and reusable probability/algebra leaves from Lattimore--Szepesvari, *Bandit Algorithms* (2020), Part IV, Chapter 17.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html","parent":"chapter:foundations","order":432,"meta":[["Source","BanditRLProof/LowerBounds/HighProbability.lean"],["Declarations","164"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.highprobability banditrlproof.lowerbounds.highprobability this module formalizes the source-faithful threshold surfaces and reusable probability/algebra leaves from lattimore--szepesvari, *bandit algorithms* (2020), part iv, chapter 17. lean module compiled","shard":"modules/c48a84f74677737d.json"},{"id":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","label":"LowerBounds.HuffmanAlphabet","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.HuffmanAlphabet","description":"Generated source map for BanditRLProof/LowerBounds/HuffmanAlphabet.lean.","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html","parent":"chapter:foundations","order":433,"meta":[["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.huffmanalphabet banditrlproof.lowerbounds.huffmanalphabet generated source map for banditrlproof/lowerbounds/huffmanalphabet.lean. lean module compiled","shard":"modules/b32960fded537975.json"},{"id":"module:BanditRLProof.LowerBounds.HuffmanConstruction","label":"LowerBounds.HuffmanConstruction","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.HuffmanConstruction","description":"Generated source map for BanditRLProof/LowerBounds/HuffmanConstruction.lean.","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html","parent":"chapter:foundations","order":434,"meta":[["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.huffmanconstruction banditrlproof.lowerbounds.huffmanconstruction generated source map for banditrlproof/lowerbounds/huffmanconstruction.lean. lean module compiled","shard":"modules/6f010ee120079082.json"},{"id":"module:BanditRLProof.LowerBounds.HuffmanStep","label":"LowerBounds.HuffmanStep","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.HuffmanStep","description":"Generated source map for BanditRLProof/LowerBounds/HuffmanStep.lean.","url":"../modules/banditrlproof-lowerbounds-huffmanstep/index.html","parent":"chapter:foundations","order":435,"meta":[["Source","BanditRLProof/LowerBounds/HuffmanStep.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.huffmanstep banditrlproof.lowerbounds.huffmanstep generated source map for banditrlproof/lowerbounds/huffmanstep.lean. lean module compiled","shard":"modules/02427867ffc46e50.json"},{"id":"module:BanditRLProof.LowerBounds.InformationTheory","label":"LowerBounds.InformationTheory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.InformationTheory","description":"This file formalizes the finite prefix-code/entropy definitions, a Kraft adapter, the measure-KL and data-processing surfaces, and event testing used in Part IV, Chapter 14 of Lattimore--Szepesvári, *Bandit Algorithms*. Measure-level relative entropy is Mathlib's extended-real InformationTheory.klDiv. The project-local work keeps codeword regularity, absolute continuity, integrability, KL direction, Bernoulli endpoi…","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html","parent":"chapter:foundations","order":436,"meta":[["Source","BanditRLProof/LowerBounds/InformationTheory.lean"],["Declarations","32"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.informationtheory banditrlproof.lowerbounds.informationtheory this file formalizes the finite prefix-code/entropy definitions, a kraft adapter, the measure-kl and data-processing surfaces, and event testing used in part iv, chapter 14 of lattimore--szepesvári, *bandit algorithms*. measure-level relative entropy is mathlib's extended-real informationtheory.kldiv. the project-local work keeps codeword regularity, absolute continuity, integrability, kl direction, bernoulli endpoints, and the infinite-divergence branch explicit. it does not claim huffman optimality or source coding. lean module compiled","shard":"modules/71eb4f20c1620f5f.json"},{"id":"module:BanditRLProof.LowerBounds.InstanceDependent","label":"LowerBounds.InstanceDependent","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.InstanceDependent","description":"This file starts the source-faithful Chapter 16 spine for Lattimore--Szepesvári, Bandit Algorithms* (2020). The compiled surface freezes the exact subpolynomial consistency quantifier, the distribution-class d_inf definition, the exact unit-Gaussian row of Table 16.1, the one-arm change-of-measure/event-information layer built on the compiled Chapter 15 history-KL identity, and the canonical gap-times-pull-count eve…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html","parent":"chapter:foundations","order":437,"meta":[["Source","BanditRLProof/LowerBounds/InstanceDependent.lean"],["Declarations","90"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.instancedependent banditrlproof.lowerbounds.instancedependent this file starts the source-faithful chapter 16 spine for lattimore--szepesvári, bandit algorithms* (2020). the compiled surface freezes the exact subpolynomial consistency quantifier, the distribution-class d_inf definition, the exact unit-gaussian row of table 16.1, the one-arm change-of-measure/event-information layer built on the compiled chapter 15 history-kl identity, and the canonical gap-times-pull-count event-to-regret producers. it also isolates the elementary asymptotic and scalar logarithmic steps used by the source proof. lean module compiled","shard":"modules/361365b9df4da239.json"},{"id":"module:BanditRLProof.LowerBounds.Minimax","label":"LowerBounds.Minimax","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Minimax","description":"This file provides the unit-Gaussian analytic leaves for the source-faithful Chapter 15 spine of Lattimore--Szepesvári, *Bandit Algorithms* (2020). The adaptive-history decomposition is proved in BanditHistoryKL; the complete finite-armed Gaussian minimax endpoint is assembled in GaussianMinimax.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html","parent":"chapter:foundations","order":438,"meta":[["Source","BanditRLProof/LowerBounds/Minimax.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.minimax banditrlproof.lowerbounds.minimax this file provides the unit-gaussian analytic leaves for the source-faithful chapter 15 spine of lattimore--szepesvári, *bandit algorithms* (2020). the adaptive-history decomposition is proved in bandithistorykl; the complete finite-armed gaussian minimax endpoint is assembled in gaussianminimax. lean module compiled","shard":"modules/6f019e70c12d35c3.json"},{"id":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","label":"LowerBounds.PrefixCodeConstruction","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.PrefixCodeConstruction","description":"Generated source map for BanditRLProof/LowerBounds/PrefixCodeConstruction.lean.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html","parent":"chapter:foundations","order":439,"meta":[["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.prefixcodeconstruction banditrlproof.lowerbounds.prefixcodeconstruction generated source map for banditrlproof/lowerbounds/prefixcodeconstruction.lean. lean module compiled","shard":"modules/468516b3f46bed90.json"},{"id":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","label":"LowerBounds.PrefixCodeExchange","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.PrefixCodeExchange","description":"Generated source map for BanditRLProof/LowerBounds/PrefixCodeExchange.lean.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html","parent":"chapter:foundations","order":440,"meta":[["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.prefixcodeexchange banditrlproof.lowerbounds.prefixcodeexchange generated source map for banditrlproof/lowerbounds/prefixcodeexchange.lean. lean module compiled","shard":"modules/b426b006e5df32f0.json"},{"id":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","label":"LowerBounds.PrefixCodeGreedy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.PrefixCodeGreedy","description":"Generated source map for BanditRLProof/LowerBounds/PrefixCodeGreedy.lean.","url":"../modules/banditrlproof-lowerbounds-prefixcodegreedy/index.html","parent":"chapter:foundations","order":441,"meta":[["Source","BanditRLProof/LowerBounds/PrefixCodeGreedy.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.prefixcodegreedy banditrlproof.lowerbounds.prefixcodegreedy generated source map for banditrlproof/lowerbounds/prefixcodegreedy.lean. lean module compiled","shard":"modules/5d36df5c68c900cb.json"},{"id":"module:BanditRLProof.LowerBounds.PrefixCodePruning","label":"LowerBounds.PrefixCodePruning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.PrefixCodePruning","description":"Generated source map for BanditRLProof/LowerBounds/PrefixCodePruning.lean.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html","parent":"chapter:foundations","order":442,"meta":[["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.prefixcodepruning banditrlproof.lowerbounds.prefixcodepruning generated source map for banditrlproof/lowerbounds/prefixcodepruning.lean. lean module compiled","shard":"modules/0ebf053c16ab922f.json"},{"id":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","label":"LowerBounds.PrefixCodeSiblings","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.PrefixCodeSiblings","description":"Generated source map for BanditRLProof/LowerBounds/PrefixCodeSiblings.lean.","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html","parent":"chapter:foundations","order":443,"meta":[["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.prefixcodesiblings banditrlproof.lowerbounds.prefixcodesiblings generated source map for banditrlproof/lowerbounds/prefixcodesiblings.lean. lean module compiled","shard":"modules/e00e04f22275603a.json"},{"id":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","label":"LowerBounds.RelativeEntropyFiltration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.RelativeEntropyFiltration","description":"Recovery of relative entropy from a filtration resolving the RN density.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html","parent":"chapter:foundations","order":444,"meta":[["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.relativeentropyfiltration banditrlproof.lowerbounds.relativeentropyfiltration recovery of relative entropy from a filtration resolving the rn density. lean module compiled","shard":"modules/a00cc609302d8339.json"},{"id":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","label":"LowerBounds.RelativeEntropyNonMetric","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.RelativeEntropyNonMetric","description":"Generated source map for BanditRLProof/LowerBounds/RelativeEntropyNonMetric.lean.","url":"../modules/banditrlproof-lowerbounds-relativeentropynonmetric/index.html","parent":"chapter:foundations","order":445,"meta":[["Source","BanditRLProof/LowerBounds/RelativeEntropyNonMetric.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.relativeentropynonmetric banditrlproof.lowerbounds.relativeentropynonmetric generated source map for banditrlproof/lowerbounds/relativeentropynonmetric.lean. lean module compiled","shard":"modules/d2261a7f5cc626f8.json"},{"id":"module:BanditRLProof.LowerBounds.ShannonLengths","label":"LowerBounds.ShannonLengths","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ShannonLengths","description":"Generated source map for BanditRLProof/LowerBounds/ShannonLengths.lean.","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html","parent":"chapter:foundations","order":446,"meta":[["Source","BanditRLProof/LowerBounds/ShannonLengths.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.shannonlengths banditrlproof.lowerbounds.shannonlengths generated source map for banditrlproof/lowerbounds/shannonlengths.lean. lean module compiled","shard":"modules/6911224bad3322b1.json"},{"id":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","label":"LowerBounds.SubgaussianMinimax","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.SubgaussianMinimax","description":"Generated source map for BanditRLProof/LowerBounds/SubgaussianMinimax.lean.","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html","parent":"chapter:foundations","order":447,"meta":[["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.subgaussianminimax banditrlproof.lowerbounds.subgaussianminimax generated source map for banditrlproof/lowerbounds/subgaussianminimax.lean. lean module compiled","shard":"modules/a86ee01fe165bae1.json"},{"id":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","label":"LowerBounds.SuccinctGeometryAudit","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.SuccinctGeometryAudit","description":"This module formalizes the first geometric layer of Zeng--Honorio (NeurIPS 2025), Definitions 3.1--3.3 and Lemmas 3.1--3.4. It keeps the atom set possibly infinite and makes every boundedness premise for real sSup explicit.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html","parent":"chapter:frontier","order":448,"meta":[["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean"],["Declarations","54"],["Project imports","0"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"lowerbounds.succinctgeometryaudit banditrlproof.lowerbounds.succinctgeometryaudit this module formalizes the first geometric layer of zeng--honorio (neurips 2025), definitions 3.1--3.3 and lemmas 3.1--3.4. it keeps the atom set possibly infinite and makes every boundedness premise for real ssup explicit. lean module compiled","shard":"modules/c766d95e3c82faab.json"},{"id":"module:BanditRLProof.LowerBounds.UniformCoding","label":"LowerBounds.UniformCoding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UniformCoding","description":"Generated source map for BanditRLProof/LowerBounds/UniformCoding.lean.","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html","parent":"chapter:foundations","order":449,"meta":[["Source","BanditRLProof/LowerBounds/UniformCoding.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"lowerbounds.uniformcoding banditrlproof.lowerbounds.uniformcoding generated source map for banditrlproof/lowerbounds/uniformcoding.lean. lean module compiled","shard":"modules/6e5ad5e5e5d2a9a5.json"},{"id":"module:BanditRLProof.MartingaleDifference","label":"MartingaleDifference","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MartingaleDifference","description":"This module exposes a small finite-prefix contract for martingale-difference increments. It deliberately stays below optional stopping or final regret theorems: the witness records integrability, adaptedness, and succ-indexed conditional mean-zero facts in the shape used by the local bandit filtration/concentration route.","url":"../modules/banditrlproof-martingaledifference/index.html","parent":"chapter:probability","order":450,"meta":[["Source","BanditRLProof/MartingaleDifference.lean"],["Declarations","15"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"martingaledifference banditrlproof.martingaledifference this module exposes a small finite-prefix contract for martingale-difference increments. it deliberately stays below optional stopping or final regret theorems: the witness records integrability, adaptedness, and succ-indexed conditional mean-zero facts in the shape used by the local bandit filtration/concentration route. lean module compiled","shard":"modules/a0be58a134567394.json"},{"id":"module:BanditRLProof.MathlibWrappers","label":"MathlibWrappers","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MathlibWrappers","description":"This module is the first intentional Mathlib interop layer. The dependency-light lemmas remain in BanditRLProof.LeafLemmas; this file only bridges those local recursive definitions to Mathlib-facing finite containers.","url":"../modules/banditrlproof-mathlibwrappers/index.html","parent":"chapter:foundations","order":451,"meta":[["Source","BanditRLProof/MathlibWrappers.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"mathlibwrappers banditrlproof.mathlibwrappers this module is the first intentional mathlib interop layer. the dependency-light lemmas remain in banditrlproof.leaflemmas; this file only bridges those local recursive definitions to mathlib-facing finite containers. lean module compiled","shard":"modules/bf55ef0c3ad0f92f.json"},{"id":"module:BanditRLProof.MeasurableLocalQuantities","label":"MeasurableLocalQuantities","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasurableLocalQuantities","description":"This module connects generic measurable finite-sum bridges back to the local recursive quantities used by the bandit vocabulary.","url":"../modules/banditrlproof-measurablelocalquantities/index.html","parent":"chapter:probability","order":452,"meta":[["Source","BanditRLProof/MeasurableLocalQuantities.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurablelocalquantities banditrlproof.measurablelocalquantities this module connects generic measurable finite-sum bridges back to the local recursive quantities used by the bandit vocabulary. lean module compiled","shard":"modules/5382c55ae34cc823.json"},{"id":"module:BanditRLProof.MeasurablePullCount","label":"MeasurablePullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasurablePullCount","description":"This module proves that the recursive pull-count process is measurable when the action trace is timewise measurable. It stays before expectation and before scalar-cast pull-count identities.","url":"../modules/banditrlproof-measurablepullcount/index.html","parent":"chapter:probability","order":453,"meta":[["Source","BanditRLProof/MeasurablePullCount.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurablepullcount banditrlproof.measurablepullcount this module proves that the recursive pull-count process is measurable when the action trace is timewise measurable. it stays before expectation and before scalar-cast pull-count identities. lean module compiled","shard":"modules/ebd93cbb17611ff1.json"},{"id":"module:BanditRLProof.MeasurablePullCountCast","label":"MeasurablePullCountCast","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasurablePullCountCast","description":"This module proves scalar-valued measurability of local pull counts. It stays before expectation, while matching the scalar form used by regret decompositions.","url":"../modules/banditrlproof-measurablepullcountcast/index.html","parent":"chapter:probability","order":454,"meta":[["Source","BanditRLProof/MeasurablePullCountCast.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurablepullcountcast banditrlproof.measurablepullcountcast this module proves scalar-valued measurability of local pull counts. it stays before expectation, while matching the scalar form used by regret decompositions. lean module compiled","shard":"modules/67083779814e664f.json"},{"id":"module:BanditRLProof.MeasurableRegret","label":"MeasurableRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasurableRegret","description":"This module keeps regret measurability before expectation. It only proves that the local deterministic pseudo-regret quantity becomes a measurable random variable when the action trace is timewise measurable.","url":"../modules/banditrlproof-measurableregret/index.html","parent":"chapter:probability","order":455,"meta":[["Source","BanditRLProof/MeasurableRegret.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurableregret banditrlproof.measurableregret this module keeps regret measurability before expectation. it only proves that the local deterministic pseudo-regret quantity becomes a measurable random variable when the action trace is timewise measurable. lean module compiled","shard":"modules/5385f9f8f4af17ff.json"},{"id":"module:BanditRLProof.MeasurableSums","label":"MeasurableSums","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasurableSums","description":"This module keeps the probability-facing layer before integration. It only proves measurability of finite sums built from the selected-reward indicator bridge in MeasureFoundation.","url":"../modules/banditrlproof-measurablesums/index.html","parent":"chapter:probability","order":456,"meta":[["Source","BanditRLProof/MeasurableSums.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurablesums banditrlproof.measurablesums this module keeps the probability-facing layer before integration. it only proves measurability of finite sums built from the selected-reward indicator bridge in measurefoundation. lean module compiled","shard":"modules/b5c7f53727a4fe1c.json"},{"id":"module:BanditRLProof.MeasureFoundation","label":"MeasureFoundation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasureFoundation","description":"This module starts the probability-facing layer with measurable events only. It deliberately avoids measure, integration, probability, filtration, and concentration imports.","url":"../modules/banditrlproof-measurefoundation/index.html","parent":"chapter:probability","order":457,"meta":[["Source","BanditRLProof/MeasureFoundation.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurefoundation banditrlproof.measurefoundation this module starts the probability-facing layer with measurable events only. it deliberately avoids measure, integration, probability, filtration, and concentration imports. lean module compiled","shard":"modules/d5dcabb5695737f5.json"},{"id":"module:BanditRLProof.MeasureL2Indicator","label":"MeasureL2Indicator","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.MeasureL2Indicator","description":"This is the nonnegative 2,2 Holder specialization used to turn an event probability bound and an exact second moment into an expected overflow bound.","url":"../modules/banditrlproof-measurel2indicator/index.html","parent":"chapter:probability","order":458,"meta":[["Source","BanditRLProof/MeasureL2Indicator.lean"],["Declarations","1"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"measurel2indicator banditrlproof.measurel2indicator this is the nonnegative 2,2 holder specialization used to turn an event probability bound and an exact second moment into an expected overflow bound. lean module compiled","shard":"modules/735ac88bb6fc3ec6.json"},{"id":"module:BanditRLProof.OFULAllTimeConfidence","label":"OFULAllTimeConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULAllTimeConfidence","description":"This module upgrades the fixed-time scalar-ridge confidence theorem to one countable union over every deterministic horizon on the same probability space. A telescoping confidence schedule allocates the total failure budget exactly:","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html","parent":"chapter:oful","order":459,"meta":[["Source","BanditRLProof/OFULAllTimeConfidence.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulalltimeconfidence banditrlproof.ofulalltimeconfidence this module upgrades the fixed-time scalar-ridge confidence theorem to one countable union over every deterministic horizon on the same probability space. a telescoping confidence schedule allocates the total failure budget exactly: lean module compiled","shard":"modules/acf5d661f53afac0.json"},{"id":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","label":"OFULConcreteHistoryRidgeSelection","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULConcreteHistoryRidgeSelection","description":"This module reconstructs the standard scalar-ridge OFUL state from an inclusive finite action/Real-reward history. It proves the complete finite-dimensional measurability chain and instantiates the measurable recursive selector without a caller-supplied score-measurability premise.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html","parent":"chapter:oful","order":460,"meta":[["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean"],["Declarations","26"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulconcretehistoryridgeselection banditrlproof.ofulconcretehistoryridgeselection this module reconstructs the standard scalar-ridge oful state from an inclusive finite action/real-reward history. it proves the complete finite-dimensional measurability chain and instantiates the measurable recursive selector without a caller-supplied score-measurability premise. lean module compiled","shard":"modules/9372d8d027df328b.json"},{"id":"module:BanditRLProof.OFULConfidenceEllipsoid","label":"OFULConfidenceEllipsoid","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULConfidenceEllipsoid","description":"This module consumes the common-R vector self-normalized tail. It defines the finite-horizon ridge estimator from scalar observations, proves its exact error decomposition into the martingale score and regularization bias, and transports the score tail to a parameter confidence ellipsoid.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html","parent":"chapter:oful","order":461,"meta":[["Source","BanditRLProof/OFULConfidenceEllipsoid.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulconfidenceellipsoid banditrlproof.ofulconfidenceellipsoid this module consumes the common-r vector self-normalized tail. it defines the finite-horizon ridge estimator from scalar observations, proves its exact error decomposition into the martingale score and regularization bias, and transports the score tail to a parameter confidence ellipsoid. lean module compiled","shard":"modules/2e6fae67537e0551.json"},{"id":"module:BanditRLProof.OFULEllipticalPotential","label":"OFULEllipticalPotential","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULEllipticalPotential","description":"This module records the deterministic finite-dimensional linear-algebra surface needed by OFUL/LinUCB routes. It packages rank-one and finite-history feature Gram matrices, positive-definiteness and inverse facts, determinant updates, log-determinant telescoping, trace/eigenvalue determinant bounds, and the standard logarithmic clipped elliptical-potential endpoint.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html","parent":"chapter:oful","order":462,"meta":[["Source","BanditRLProof/OFULEllipticalPotential.lean"],["Declarations","115"],["Project imports","0"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulellipticalpotential banditrlproof.ofulellipticalpotential this module records the deterministic finite-dimensional linear-algebra surface needed by oful/linucb routes. it packages rank-one and finite-history feature gram matrices, positive-definiteness and inverse facts, determinant updates, log-determinant telescoping, trace/eigenvalue determinant bounds, and the standard logarithmic clipped elliptical-potential endpoint. lean module compiled","shard":"modules/889e011214c8f29f.json"},{"id":"module:BanditRLProof.OFULEllipticalPotentialFoundation","label":"OFULEllipticalPotentialFoundation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULEllipticalPotentialFoundation","description":"This module packages the determinant-growth and clipped inverse-quadratic conclusions of the deterministic OFUL linear-algebra route in one theorem-facing statement. It is the handoff expected by a later self-normalized concentration and confidence-ellipsoid proof.","url":"../modules/banditrlproof-ofulellipticalpotentialfoundation/index.html","parent":"chapter:oful","order":463,"meta":[["Source","BanditRLProof/OFULEllipticalPotentialFoundation.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulellipticalpotentialfoundation banditrlproof.ofulellipticalpotentialfoundation this module packages the determinant-growth and clipped inverse-quadratic conclusions of the deterministic oful linear-algebra route in one theorem-facing statement. it is the handoff expected by a later self-normalized concentration and confidence-ellipsoid proof. lean module compiled","shard":"modules/28ca18f59b716403.json"},{"id":"module:BanditRLProof.OFULExpectedRegret","label":"OFULExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULExpectedRegret","description":"This module turns the compiled all-round high-probability cumulative-gap tail into a Bochner expected-gap bound. The bad event is charged by a deterministic finite-window envelope obtained from the same parameter and arm norm bounds.","url":"../modules/banditrlproof-ofulexpectedregret/index.html","parent":"chapter:oful","order":464,"meta":[["Source","BanditRLProof/OFULExpectedRegret.lean"],["Declarations","20"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulexpectedregret banditrlproof.ofulexpectedregret this module turns the compiled all-round high-probability cumulative-gap tail into a bochner expected-gap bound. the bad event is charged by a deterministic finite-window envelope obtained from the same parameter and arm norm bounds. lean module compiled","shard":"modules/b138104eaf9568de.json"},{"id":"module:BanditRLProof.OFULExpectedRegretAsymptotics","label":"OFULExpectedRegretAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULExpectedRegretAsymptotics","description":"This module upgrades the explicit finite-window canonical OFUL expected pseudo-regret theorem to an IsBigO statement at Filter.atTop. The feature dimension and model parameters are fixed while the horizon varies.","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html","parent":"chapter:oful","order":465,"meta":[["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulexpectedregretasymptotics banditrlproof.ofulexpectedregretasymptotics this module upgrades the explicit finite-window canonical oful expected pseudo-regret theorem to an isbigo statement at filter.attop. the feature dimension and model parameters are fixed while the horizon varies. lean module compiled","shard":"modules/2291c64b9b778162.json"},{"id":"module:BanditRLProof.OFULExpectedRegretConsistency","label":"OFULExpectedRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULExpectedRegretConsistency","description":"This module turns the fixed-model asymptotic expected pseudo-regret bound into convergence of expected pseudo-regret per played round. The generated algorithm still uses the horizon-dependent parameter 1 / (T + 1)^2.","url":"../modules/banditrlproof-ofulexpectedregretconsistency/index.html","parent":"chapter:oful","order":466,"meta":[["Source","BanditRLProof/OFULExpectedRegretConsistency.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulexpectedregretconsistency banditrlproof.ofulexpectedregretconsistency this module turns the fixed-model asymptotic expected pseudo-regret bound into convergence of expected pseudo-regret per played round. the generated algorithm still uses the horizon-dependent parameter 1 / (t + 1)^2. lean module compiled","shard":"modules/c08bbca52c3ee0df.json"},{"id":"module:BanditRLProof.OFULExpectedRegretRate","label":"OFULExpectedRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULExpectedRegretRate","description":"This module normalizes the horizon-tuned confidence parameter used by the canonical OFUL expected-regret theorem and exposes its logarithmic square-root bound without hiding the rate inside the radius-width definitions.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html","parent":"chapter:oful","order":467,"meta":[["Source","BanditRLProof/OFULExpectedRegretRate.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulexpectedregretrate banditrlproof.ofulexpectedregretrate this module normalizes the horizon-tuned confidence parameter used by the canonical oful expected-regret theorem and exposes its logarithmic square-root bound without hiding the rate inside the radius-width definitions. lean module compiled","shard":"modules/2dbf1f19a2a258b4.json"},{"id":"module:BanditRLProof.OFULFiniteActionOptimism","label":"OFULFiniteActionOptimism","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULFiniteActionOptimism","description":"This module turns the compiled scalar-ridge confidence ellipsoid into a finite-action optimistic selector and a one-step gap certificate.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html","parent":"chapter:oful","order":468,"meta":[["Source","BanditRLProof/OFULFiniteActionOptimism.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulfiniteactionoptimism banditrlproof.ofulfiniteactionoptimism this module turns the compiled scalar-ridge confidence ellipsoid into a finite-action optimistic selector and a one-step gap certificate. lean module compiled","shard":"modules/1e1baacfe4aed984.json"},{"id":"module:BanditRLProof.OFULFiniteHorizonScoreGram","label":"OFULFiniteHorizonScoreGram","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULFiniteHorizonScoreGram","description":"This module identifies the scalar compensated process from the fixed-direction conditional-MGF theorem with the random score/random-Gram quadratic exponential used by the Gaussian method of mixtures.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html","parent":"chapter:oful","order":469,"meta":[["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulfinitehorizonscoregram banditrlproof.ofulfinitehorizonscoregram this module identifies the scalar compensated process from the fixed-direction conditional-mgf theorem with the random score/random-gram quadratic exponential used by the gaussian method of mixtures. lean module compiled","shard":"modules/9111c1c9811fcb37.json"},{"id":"module:BanditRLProof.OFULGaussianCovarianceMixture","label":"OFULGaussianCovarianceMixture","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGaussianCovarianceMixture","description":"This file transports the standard-Gaussian quadratic-exponential identity through the square root of V0⁻¹. For a positive-definite prior precision V0 and a positive-semidefinite Gram matrix G, it collects the transformed determinant and quadratic form into the OFUL-facing matrices V0 + G and (V0 + G)⁻¹.","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html","parent":"chapter:oful","order":470,"meta":[["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgaussiancovariancemixture banditrlproof.ofulgaussiancovariancemixture this file transports the standard-gaussian quadratic-exponential identity through the square root of v0⁻¹. for a positive-definite prior precision v0 and a positive-semidefinite gram matrix g, it collects the transformed determinant and quadratic form into the oful-facing matrices v0 + g and (v0 + g)⁻¹. lean module compiled","shard":"modules/183124b07255e581.json"},{"id":"module:BanditRLProof.OFULGaussianEvaluatedMixture","label":"OFULGaussianEvaluatedMixture","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGaussianEvaluatedMixture","description":"This module evaluates the inner Gaussian direction integral in the finite-horizon score/variance-Gram mixture and transports the product-space bound to the determinant-ratio inverse-Gram exponential.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html","parent":"chapter:oful","order":471,"meta":[["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgaussianevaluatedmixture banditrlproof.ofulgaussianevaluatedmixture this module evaluates the inner gaussian direction integral in the finite-horizon score/variance-gram mixture and transports the product-space bound to the determinant-ratio inverse-gram exponential. lean module compiled","shard":"modules/4479e55dbdf79eaf.json"},{"id":"module:BanditRLProof.OFULGaussianMixture","label":"OFULGaussianMixture","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGaussianMixture","description":"This module develops the exact Gaussian quadratic-exponential integrals used after the fixed-direction exponential-supermartingale bound. The scalar identity is the normalization step needed before the finite-dimensional spectral/product assembly.","url":"../modules/banditrlproof-ofulgaussianmixture/index.html","parent":"chapter:oful","order":472,"meta":[["Source","BanditRLProof/OFULGaussianMixture.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgaussianmixture banditrlproof.ofulgaussianmixture this module develops the exact gaussian quadratic-exponential integrals used after the fixed-direction exponential-supermartingale bound. the scalar identity is the normalization step needed before the finite-dimensional spectral/product assembly. lean module compiled","shard":"modules/d5becfe86da00ed6.json"},{"id":"module:BanditRLProof.OFULGaussianMixtureMeasurability","label":"OFULGaussianMixtureMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGaussianMixtureMeasurability","description":"This file packages the random-score/random-Gram quadratic exponential on the product of a sample space and a finite-dimensional Gaussian parameter space. It proves the joint measurable surface needed by Tonelli and exposes both a generic product-measure identity and the N(0, V0⁻¹) specialization.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html","parent":"chapter:oful","order":473,"meta":[["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgaussianmixturemeasurability banditrlproof.ofulgaussianmixturemeasurability this file packages the random-score/random-gram quadratic exponential on the product of a sample space and a finite-dimensional gaussian parameter space. it proves the joint measurable surface needed by tonelli and exposes both a generic product-measure identity and the n(0, v0⁻¹) specialization. lean module compiled","shard":"modules/4559641663483d22.json"},{"id":"module:BanditRLProof.OFULGaussianSpectralMixture","label":"OFULGaussianSpectralMixture","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGaussianSpectralMixture","description":"This module transports the compiled independent-coordinate Gaussian quadratic-exponential identity through the orthonormal eigenbasis of a real positive-semidefinite matrix. It then collects the spectral product and coordinate quadratic form as det (1 + A) and (1 + A)⁻¹.","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html","parent":"chapter:oful","order":474,"meta":[["Source","BanditRLProof/OFULGaussianSpectralMixture.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgaussianspectralmixture banditrlproof.ofulgaussianspectralmixture this module transports the compiled independent-coordinate gaussian quadratic-exponential identity through the orthonormal eigenbasis of a real positive-semidefinite matrix. it then collects the spectral product and coordinate quadratic form as det (1 + a) and (1 + a)⁻¹. lean module compiled","shard":"modules/5d04c0c7389d56b5.json"},{"id":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","label":"OFULGeneratedTrajectoryConfidenceGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","description":"This module identifies the scalar-ridge state reconstructed from an inclusive finite pair history with the generic finite-horizon state on the underlying trajectory. It then transports the measurable strict-fold selector's score maximality to one-step and finite successor-window optimism-gap bounds.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html","parent":"chapter:oful","order":475,"meta":[["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgeneratedtrajectoryconfidencegap banditrlproof.ofulgeneratedtrajectoryconfidencegap this module identifies the scalar-ridge state reconstructed from an inclusive finite pair history with the generic finite-horizon state on the underlying trajectory. it then transports the measurable strict-fold selector's score maximality to one-step and finite successor-window optimism-gap bounds. lean module compiled","shard":"modules/69bb82127102175c.json"},{"id":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","label":"OFULGeneratedTrajectoryPredictableConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","description":"The actual canonical action coordinate is only almost surely equal to the deterministic history selector. This module constructs the pointwise predictable selector feature under a strict-past filtration, proves its almost-everywhere alignment with the actual selected feature, and transports the compiled uniform confidence and successor-gap tails back to the actual trajectory.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html","parent":"chapter:oful","order":476,"meta":[["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean"],["Declarations","20"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgeneratedtrajectorypredictableconfidence banditrlproof.ofulgeneratedtrajectorypredictableconfidence the actual canonical action coordinate is only almost surely equal to the deterministic history selector. this module constructs the pointwise predictable selector feature under a strict-past filtration, proves its almost-everywhere alignment with the actual selected feature, and transports the compiled uniform confidence and successor-gap tails back to the actual trajectory. lean module compiled","shard":"modules/8f536a33f6c96662.json"},{"id":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","label":"OFULGeneratedTrajectoryRadiusWidth","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","description":"This module connects the canonical successor-gap tail to the deterministic selected-width theorem. It first bounds every finite-horizon scalar confidence radius by one standard log-determinant radius at the terminal horizon.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html","parent":"chapter:oful","order":477,"meta":[["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgeneratedtrajectoryradiuswidth banditrlproof.ofulgeneratedtrajectoryradiuswidth this module connects the canonical successor-gap tail to the deterministic selected-width theorem. it first bounds every finite-horizon scalar confidence radius by one standard log-determinant radius at the terminal horizon. lean module compiled","shard":"modules/e3fcc58bd11cf3fd.json"},{"id":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","label":"OFULGeneratedTrajectoryUniformConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","description":"This module packages the precise filtration and conditional-law contracts needed to apply the compiled finite-window scalar-ridge confidence theorem to one canonical history-algorithm trajectory. It then combines that probability bound with the compiled successor-window good-event gap transport.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryuniformconfidence/index.html","parent":"chapter:oful","order":478,"meta":[["Source","BanditRLProof/OFULGeneratedTrajectoryUniformConfidence.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulgeneratedtrajectoryuniformconfidence banditrlproof.ofulgeneratedtrajectoryuniformconfidence this module packages the precise filtration and conditional-law contracts needed to apply the compiled finite-window scalar-ridge confidence theorem to one canonical history-algorithm trajectory. it then combines that probability bound with the compiled successor-window good-event gap transport. lean module compiled","shard":"modules/79c711e35b34d2bf.json"},{"id":"module:BanditRLProof.OFULHighProbabilityRegretRate","label":"OFULHighProbabilityRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULHighProbabilityRegretRate","description":"This module expands the named all-round OFUL gap budget into an explicit logarithmic radius-width expression and specializes the compiled cumulative-gap tail to a certified optimal fixed arm.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html","parent":"chapter:oful","order":479,"meta":[["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulhighprobabilityregretrate banditrlproof.ofulhighprobabilityregretrate this module expands the named all-round oful gap budget into an explicit logarithmic radius-width expression and specializes the compiled cumulative-gap tail to a certified optimal fixed arm. lean module compiled","shard":"modules/f8cc3317a82480ed.json"},{"id":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","label":"OFULHistoryEnvironmentRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULHistoryEnvironmentRewardLaw","description":"This module connects the Markov kernels stored in Thompson.HistoryEnvironment to the strict-past predictable residual law consumed by the compiled OFUL confidence route.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html","parent":"chapter:oful","order":480,"meta":[["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulhistoryenvironmentrewardlaw banditrlproof.ofulhistoryenvironmentrewardlaw this module connects the markov kernels stored in thompson.historyenvironment to the strict-past predictable residual law consumed by the compiled oful confidence route. lean module compiled","shard":"modules/be65fb476a795048.json"},{"id":"module:BanditRLProof.OFULInitialRoundGap","label":"OFULInitialRoundGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULInitialRoundGap","description":"This module bounds the fixed time-zero linear gap by the parameter and feature norm envelopes, then adds that deterministic charge to the compiled normalized successor-gap tail.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html","parent":"chapter:oful","order":481,"meta":[["Source","BanditRLProof/OFULInitialRoundGap.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulinitialroundgap banditrlproof.ofulinitialroundgap this module bounds the fixed time-zero linear gap by the parameter and feature norm envelopes, then adds that deterministic charge to the compiled normalized successor-gap tail. lean module compiled","shard":"modules/ad567d08ec5c40f2.json"},{"id":"module:BanditRLProof.OFULMeasurableRecursiveSelection","label":"OFULMeasurableRecursiveSelection","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULMeasurableRecursiveSelection","description":"This module replaces the nonconstructive finite OFUL argmax by the existing strict finite fold on Fin K. Under measurable score coordinates, the fold defines a deterministic history policy. The canonical history trajectory then realizes that selector at every successor time.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html","parent":"chapter:oful","order":482,"meta":[["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean"],["Declarations","9"],["Project imports","3"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulmeasurablerecursiveselection banditrlproof.ofulmeasurablerecursiveselection this module replaces the nonconstructive finite oful argmax by the existing strict finite fold on fin k. under measurable score coordinates, the fold defines a deterministic history policy. the canonical history trajectory then realizes that selector at every successor time. lean module compiled","shard":"modules/f8e33688f91d2e87.json"},{"id":"module:BanditRLProof.OFULNormalizedRadiusWidth","label":"OFULNormalizedRadiusWidth","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULNormalizedRadiusWidth","description":"This module derives the raw confidenceWidth <= 1 contract from the explicit normalization dotProduct x x <= L2 <= lambda, then exposes the canonical standard successor-gap tail without a caller-supplied width premise.","url":"../modules/banditrlproof-ofulnormalizedradiuswidth/index.html","parent":"chapter:oful","order":483,"meta":[["Source","BanditRLProof/OFULNormalizedRadiusWidth.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulnormalizedradiuswidth banditrlproof.ofulnormalizedradiuswidth this module derives the raw confidencewidth <= 1 contract from the explicit normalization dotproduct x x <= l2 <= lambda, then exposes the canonical standard successor-gap tail without a caller-supplied width premise. lean module compiled","shard":"modules/a72da139f18910bb.json"},{"id":"module:BanditRLProof.OFULScalarRegularizationBias","label":"OFULScalarRegularizationBias","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScalarRegularizationBias","description":"This module discharges the deterministic ridge-bias contract in the compiled finite-horizon confidence ellipsoid when the base matrix is lambda I.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html","parent":"chapter:oful","order":484,"meta":[["Source","BanditRLProof/OFULScalarRegularizationBias.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscalarregularizationbias banditrlproof.ofulscalarregularizationbias this module discharges the deterministic ridge-bias contract in the compiled finite-horizon confidence ellipsoid when the base matrix is lambda i. lean module compiled","shard":"modules/b7ab38cd1a862f90.json"},{"id":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","label":"OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","description":"This module replaces the earlier pointwise reach contract on every theoretical trajectory by the measure-theoretically sufficient almost-everywhere contract under the canonical generated trajectory law. It then specializes the route to deterministic Nat-valued action costs.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":485,"meta":[["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret banditrlproof.ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret this module replaces the earlier pointwise reach contract on every theoretical trajectory by the measure-theoretically sufficient almost-everywhere contract under the canonical generated trajectory law. it then specializes the route to deterministic nat-valued action costs. lean module compiled","shard":"modules/3afa54f7197ac783.json"},{"id":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","label":"OFULScheduledAllHorizonAllRoundGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledAllHorizonAllRoundGap","description":"This module adds the fixed canonical initial round to the compiled one-policy all-horizon successor-gap tail. The resulting event still uses the same telescoping-schedule policy and the same all-time confidence failure event.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html","parent":"chapter:oful","order":486,"meta":[["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledallhorizonallroundgap banditrlproof.ofulscheduledallhorizonallroundgap this module adds the fixed canonical initial round to the compiled one-policy all-horizon successor-gap tail. the resulting event still uses the same telescoping-schedule policy and the same all-time confidence failure event. lean module compiled","shard":"modules/8fb09f760fbe9d13.json"},{"id":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","label":"OFULScheduledAllHorizonCumulativeGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledAllHorizonCumulativeGap","description":"This module combines the one-policy all-time scheduled confidence event with a deterministic varying-budget radius-width envelope. The terminal event quantifies over every finite horizon but is absorbed by the same all-time confidence failure event.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html","parent":"chapter:oful","order":487,"meta":[["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledallhorizoncumulativegap banditrlproof.ofulscheduledallhorizoncumulativegap this module combines the one-policy all-time scheduled confidence event with a deterministic varying-budget radius-width envelope. the terminal event quantifies over every finite horizon but is absorbed by the same all-time confidence failure event. lean module compiled","shard":"modules/46980bd208ebf135.json"},{"id":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","label":"OFULScheduledAllHorizonHighProbabilityRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","description":"This module eliminates the telescoping confidence schedule from the displayed budget of the one-policy all-horizon pseudo-regret theorem.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html","parent":"chapter:oful","order":488,"meta":[["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledallhorizonhighprobabilityregretrate banditrlproof.ofulscheduledallhorizonhighprobabilityregretrate this module eliminates the telescoping confidence schedule from the displayed budget of the one-policy all-horizon pseudo-regret theorem. lean module compiled","shard":"modules/7a9a85d762279023.json"},{"id":"module:BanditRLProof.OFULScheduledAllTimeConfidence","label":"OFULScheduledAllTimeConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledAllTimeConfidence","description":"This module defines one history algorithm whose successor selector at history index n uses the confidence budget assigned to horizon n + 1. It then constructs the corresponding strict-past predictable feature and residual process and transports the all-time scalar-ridge confidence event to the actual canonical trajectory.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html","parent":"chapter:oful","order":489,"meta":[["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean"],["Declarations","33"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledalltimeconfidence banditrlproof.ofulscheduledalltimeconfidence this module defines one history algorithm whose successor selector at history index n uses the confidence budget assigned to horizon n + 1. it then constructs the corresponding strict-past predictable feature and residual process and transports the all-time scalar-ridge confidence event to the actual canonical trajectory. lean module compiled","shard":"modules/1d83be5691f4b15d.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","label":"OFULScheduledBlockStartForcedActionChargeBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","description":"Generated source map for BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html","parent":"chapter:oful","order":490,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedactionchargebound banditrlproof.ofulscheduledblockstartforcedactionchargebound generated source map for banditrlproof/ofulscheduledblockstartforcedactionchargebound.lean. lean module compiled","shard":"modules/2d0b08cb2da05bd2.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","label":"OFULScheduledBlockStartForcedAllTimeConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","description":"This module parameterizes the predictable-feature concentration layer over an arbitrary measurable deterministic finite-history selector. It then specializes that layer to the block-start forced telescoping OFUL policy, so the final confidence tail is proved under the modified policy's own canonical trajectory measure.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html","parent":"chapter:oful","order":491,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean"],["Declarations","21"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedalltimeconfidence banditrlproof.ofulscheduledblockstartforcedalltimeconfidence this module parameterizes the predictable-feature concentration layer over an arbitrary measurable deterministic finite-history selector. it then specializes that layer to the block-start forced telescoping oful policy, so the final confidence tail is proved under the modified policy's own canonical trajectory measure. lean module compiled","shard":"modules/2a3e8e8910a8b322.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","label":"OFULScheduledBlockStartForcedHistoryAlgorithm","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","description":"This module packages measurable finite-history selectors as deterministic history algorithms and proves their canonical successor-action graph. It then defines a modified telescoping OFUL selector that forces one prescribed action after every block-start history while retaining the ordinary optimistic selector at all other history indices.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html","parent":"chapter:oful","order":492,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedhistoryalgorithm banditrlproof.ofulscheduledblockstartforcedhistoryalgorithm this module packages measurable finite-history selectors as deterministic history algorithms and proves their canonical successor-action graph. it then defines a modified telescoping oful selector that forces one prescribed action after every block-start history while retaining the ordinary optimistic selector at all other history indices. lean module compiled","shard":"modules/d489a2ed3f54e800.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","label":"OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","description":"Generated source map for BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate/index.html","parent":"chapter:oful","order":493,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate banditrlproof.ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate generated source map for banditrlproof/ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate.lean. lean module compiled","shard":"modules/49e027c27d113924.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","label":"OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","description":"Generated source map for BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html","parent":"chapter:oful","order":494,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedhorizonwindowfinitehorizontail banditrlproof.ofulscheduledblockstartforcedhorizonwindowfinitehorizontail generated source map for banditrlproof/ofulscheduledblockstartforcedhorizonwindowfinitehorizontail.lean. lean module compiled","shard":"modules/fde36a91ea377d79.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","label":"OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","description":"The telescoping-confidence OFUL algorithm is deterministic conditional on its finite pair history. This module exposes its selected action as a named history function and transports the canonical successor-action law to an aligned-window positive-cost contract.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":495,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret banditrlproof.ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret the telescoping-confidence oful algorithm is deterministic conditional on its finite pair history. this module exposes its selected action as a named history function and transports the canonical successor-action law to an aligned-window positive-cost contract. lean module compiled","shard":"modules/984cbb88235ee835.json"},{"id":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","label":"OFULScheduledBlockStartForcedPseudoRegretDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","description":"This module separates the complete finite-horizon pseudo-regret of the block-start forced policy into the initial action, forced successor actions, and ordinary optimistic successor actions. The successor index n denotes the history used to select action n + 1.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html","parent":"chapter:oful","order":496,"meta":[["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledblockstartforcedpseudoregretdecomposition banditrlproof.ofulscheduledblockstartforcedpseudoregretdecomposition this module separates the complete finite-horizon pseudo-regret of the block-start forced policy into the initial action, forced successor actions, and ordinary optimistic successor actions. the successor index n denotes the history used to select action n + 1. lean module compiled","shard":"modules/1355bf4b8c8863f8.json"},{"id":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","label":"OFULScheduledBoundedStoppingTimeExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","description":"This module integrates the compiled bounded-stopping-time high-probability pseudo-regret theorem. Horizon monotonicity moves both the explicit scheduled budget and the deterministic gap envelope to the deterministic stopping-time bound. The expectation proof is an indicator decomposition, not optional stopping.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html","parent":"chapter:oful","order":497,"meta":[["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledboundedstoppingtimeexpectedregret banditrlproof.ofulscheduledboundedstoppingtimeexpectedregret this module integrates the compiled bounded-stopping-time high-probability pseudo-regret theorem. horizon monotonicity moves both the explicit scheduled budget and the deterministic gap envelope to the deterministic stopping-time bound. the expectation proof is an indicator decomposition, not optional stopping. lean module compiled","shard":"modules/e464efabbb613273.json"},{"id":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","label":"OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","description":"This module upgrades the explicit pointwise expected pseudo-regret bound for the horizon-indexed telescoping-schedule policy and bounded stopping-time family to an IsBigO statement at Filter.atTop.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretasymptotics/index.html","parent":"chapter:oful","order":498,"meta":[["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledboundedstoppingtimeexpectedregretasymptotics banditrlproof.ofulscheduledboundedstoppingtimeexpectedregretasymptotics this module upgrades the explicit pointwise expected pseudo-regret bound for the horizon-indexed telescoping-schedule policy and bounded stopping-time family to an isbigo statement at filter.attop. lean module compiled","shard":"modules/2e58e522b3f30601.json"},{"id":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","label":"OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","description":"This module normalizes the fixed-model horizon-indexed expected stopped pseudo-regret family by the number of available rounds. The policy at horizon T retains the telescoping schedule with outer budget 1 / (T + 1).","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretconsistency/index.html","parent":"chapter:oful","order":499,"meta":[["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretConsistency.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledboundedstoppingtimeexpectedregretconsistency banditrlproof.ofulscheduledboundedstoppingtimeexpectedregretconsistency this module normalizes the fixed-model horizon-indexed expected stopped pseudo-regret family by the number of available rounds. the policy at horizon t retains the telescoping schedule with outer budget 1 / (t + 1). lean module compiled","shard":"modules/40002839ba7154da.json"},{"id":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","label":"OFULScheduledBoundedStoppingTimeExpectedRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","description":"This module tunes the outer confidence budget of the scheduled bounded stopping-time expectation theorem to 1 / (T + 1). The stopped bad-event charge becomes one initial-gap envelope, while the scheduled confidence logarithm remains explicit.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html","parent":"chapter:oful","order":500,"meta":[["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledboundedstoppingtimeexpectedregretrate banditrlproof.ofulscheduledboundedstoppingtimeexpectedregretrate this module tunes the outer confidence budget of the scheduled bounded stopping-time expectation theorem to 1 / (t + 1). the stopped bad-event charge becomes one initial-gap envelope, while the scheduled confidence logarithm remains explicit. lean module compiled","shard":"modules/3886869feb38791d.json"},{"id":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","label":"OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","description":"This module evaluates the explicit one-policy all-horizon OFUL rate at a bounded stopping time. The probability argument is pathwise event domination, not optional stopping and not a new union bound.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html","parent":"chapter:oful","order":501,"meta":[["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledboundedstoppingtimehighprobabilityregretrate banditrlproof.ofulscheduledboundedstoppingtimehighprobabilityregretrate this module evaluates the explicit one-policy all-horizon oful rate at a bounded stopping time. the probability argument is pathwise event domination, not optional stopping and not a new union bound. lean module compiled","shard":"modules/db54f0cbc6436ea5.json"},{"id":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","label":"OFULScheduledBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","description":"This module connects the Mathlib-backed budget hitting time to the single telescoping-schedule OFUL policy. A pathwise reach-by-budget premise supplies a deterministic stopping bound, hence the square-integrable stopping contract and an explicit round-count second-moment bound.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":502,"meta":[["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledbudgetexhaustionexpectedregret banditrlproof.ofulscheduledbudgetexhaustionexpectedregret this module connects the mathlib-backed budget hitting time to the single telescoping-schedule oful policy. a pathwise reach-by-budget premise supplies a deterministic stopping bound, hence the square-integrable stopping contract and an explicit round-count second-moment bound. lean module compiled","shard":"modules/be274911e9fa198a.json"},{"id":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","label":"OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","description":"This module separates a resource threshold from the deterministic horizon by which it is reached. It then allows individual rounds to have zero cost while requiring every aligned block of a fixed length to spend at least one unit.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":503,"meta":[["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret banditrlproof.ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret this module separates a resource threshold from the deterministic horizon by which it is reached. it then allows individual rounds to have zero cost while requiring every aligned block of a fixed length to spend at least one unit. lean module compiled","shard":"modules/813bbf66cd6c8851.json"},{"id":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","label":"OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","description":"This module constructs an adapted Nat-valued resource process by summing per-round costs over completed rounds. Pointwise positive costs give the unit-growth contract required by the compiled budget-exhaustion OFUL theorem.","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":504,"meta":[["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret banditrlproof.ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret this module constructs an adapted nat-valued resource process by summing per-round costs over completed rounds. pointwise positive costs give the unit-growth contract required by the compiled budget-exhaustion oful theorem. lean module compiled","shard":"modules/4ac1df2117868a14.json"},{"id":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","label":"OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","description":"This module specializes the cumulative positive-cost budget-exhaustion route to a deterministic Nat-valued cost on the finite action set. At round t, the cost process reads the current action from the canonical generated trajectory.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":505,"meta":[["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret banditrlproof.ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret this module specializes the cumulative positive-cost budget-exhaustion route to a deterministic nat-valued cost on the finite action set. at round t, the cost process reads the current action from the canonical generated trajectory. lean module compiled","shard":"modules/4fb104d9366fa1a9.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","label":"OFULScheduledPowerOfTwoForcedAllTimeConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","description":"This module specializes the generic deterministic-history-selector confidence machinery to the horizon-independent power-of-two forced policy. Its terminal all-horizon pseudo-regret tail keeps the deterministic forced-action charge explicit; a logarithmic scalar bound for that charge is a separate downstream obligation.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html","parent":"chapter:oful","order":506,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedalltimeconfidence banditrlproof.ofulscheduledpoweroftwoforcedalltimeconfidence this module specializes the generic deterministic-history-selector confidence machinery to the horizon-independent power-of-two forced policy. its terminal all-horizon pseudo-regret tail keeps the deterministic forced-action charge explicit; a logarithmic scalar bound for that charge is a separate downstream obligation. lean module compiled","shard":"modules/4f4deace0566166d.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","label":"OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","description":"This module squeezes complete pseudo-regret per round to zero on every trajectory outside the existing all-horizon violation event. It then bounds the set of trajectories where this limit fails by the same fixed outer confidence budget, under the same policy and canonical measure.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency/index.html","parent":"chapter:oful","order":507,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency banditrlproof.ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency this module squeezes complete pseudo-regret per round to zero on every trajectory outside the existing all-horizon violation event. it then bounds the set of trajectories where this limit fails by the same fixed outer confidence budget, under the same policy and canonical measure. lean module compiled","shard":"modules/5040ce7e3935c6d7.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","label":"OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","description":"This module divides the exact fixed-model scalar all-horizon budget by the number of available rounds. It preserves the same power-of-two forced policy, canonical measure, violation event, and fixed outer confidence budget.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html","parent":"chapter:oful","order":508,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedhighprobabilityaverageregret banditrlproof.ofulscheduledpoweroftwoforcedhighprobabilityaverageregret this module divides the exact fixed-model scalar all-horizon budget by the number of available rounds. it preserves the same power-of-two forced policy, canonical measure, violation event, and fixed outer confidence budget. lean module compiled","shard":"modules/78456b0d1df0a2bd.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","label":"OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","description":"This module packages the exact scalar all-horizon budget of the fixed power-of-two forced policy and proves its fixed-model asymptotic growth. The probability event, policy, canonical measure, and confidence budget are inherited unchanged from the finite-horizon scalar theorem.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html","parent":"chapter:oful","order":509,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedhighprobabilityregretrate banditrlproof.ofulscheduledpoweroftwoforcedhighprobabilityregretrate this module packages the exact scalar all-horizon budget of the fixed power-of-two forced policy and proves its fixed-model asymptotic growth. the probability event, policy, canonical measure, and confidence budget are inherited unchanged from the finite-horizon scalar theorem. lean module compiled","shard":"modules/d1e2a899a845ad30.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","label":"OFULScheduledPowerOfTwoForcedHistoryAlgorithm","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","description":"This module packages one horizon-independent deterministic history algorithm. At history indices one below a power of two it selects the arm indexed by that power's exponent; at every other index it uses the ordinary telescoping OFUL selector.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html","parent":"chapter:oful","order":510,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedhistoryalgorithm banditrlproof.ofulscheduledpoweroftwoforcedhistoryalgorithm this module packages one horizon-independent deterministic history algorithm. at history indices one below a power of two it selects the arm indexed by that power's exponent; at every other index it uses the ordinary telescoping oful selector. lean module compiled","shard":"modules/b055bad5ea7bf136.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","label":"OFULScheduledPowerOfTwoForcedIndexCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","description":"Generated source map for BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html","parent":"chapter:oful","order":511,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean"],["Declarations","7"],["Project imports","0"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedindexcount banditrlproof.ofulscheduledpoweroftwoforcedindexcount generated source map for banditrlproof/ofulscheduledpoweroftwoforcedindexcount.lean. lean module compiled","shard":"modules/0c6215825c60c5b3.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","label":"OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","description":"This module separates the complete finite-horizon pseudo-regret of the horizon-independent power-of-two forced policy into the initial action, power-of-two forced successor actions, and ordinary optimistic successor actions. The successor index n denotes the history used to select action n + 1.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html","parent":"chapter:oful","order":512,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedpseudoregretdecomposition banditrlproof.ofulscheduledpoweroftwoforcedpseudoregretdecomposition this module separates the complete finite-horizon pseudo-regret of the horizon-independent power-of-two forced policy into the initial action, power-of-two forced successor actions, and ordinary optimistic successor actions. the successor index n denotes the history used to select action n + 1. lean module compiled","shard":"modules/138f6a849a8fc2c9.json"},{"id":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","label":"OFULScheduledPowerOfTwoForcedScalarChargeBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","description":"This module replaces the explicit prescribed-action charge in the all-time power-of-two forced OFUL theorem by a logarithmic scalar budget. The probability space, policy, and all-time confidence event are inherited unchanged from the upstream leaf.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html","parent":"chapter:oful","order":513,"meta":[["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledpoweroftwoforcedscalarchargebound banditrlproof.ofulscheduledpoweroftwoforcedscalarchargebound this module replaces the explicit prescribed-action charge in the all-time power-of-two forced oful theorem by a logarithmic scalar budget. the probability space, policy, and all-time confidence event are inherited unchanged from the upstream leaf. lean module compiled","shard":"modules/af164904f3022bdb.json"},{"id":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","label":"OFULScheduledUnboundedStoppingTimeExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","description":"This module removes the deterministic stopping-horizon bound from the expected pseudo-regret interface. It retains the exact bad-event random-envelope integral: controlling that term by delta requires an additional moment or tail contract and is not a consequence of first-moment stopping-time integrability alone. This is an event decomposition, not optional stopping.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html","parent":"chapter:oful","order":514,"meta":[["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledunboundedstoppingtimeexpectedregret banditrlproof.ofulscheduledunboundedstoppingtimeexpectedregret this module removes the deterministic stopping-horizon bound from the expected pseudo-regret interface. it retains the exact bad-event random-envelope integral: controlling that term by delta requires an additional moment or tail contract and is not a consequence of first-moment stopping-time integrability alone. this is an event decomposition, not optional stopping. lean module compiled","shard":"modules/0174dd6d0f187983.json"},{"id":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","label":"OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","description":"This module proves that the explicit telescoping OFUL budget grows at most quadratically in the round count. The existing L2 stopping-time contract therefore supplies stopped-budget integrability automatically.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html","parent":"chapter:oful","order":515,"meta":[["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledunboundedstoppingtimeexpectedregretclosed banditrlproof.ofulscheduledunboundedstoppingtimeexpectedregretclosed this module proves that the explicit telescoping oful budget grows at most quadratically in the round count. the existing l2 stopping-time contract therefore supplies stopped-budget integrability automatically. lean module compiled","shard":"modules/f9942c9d2c8fd217.json"},{"id":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","label":"OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","description":"This module names the actual second moment of the stopping-time round count and uses it directly in the canonical terminal theorem. Concrete stopping-rule analyses can subsequently bound this named quantity without rebuilding the stopped-regret argument.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretexactmoment/index.html","parent":"chapter:oful","order":516,"meta":[["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledunboundedstoppingtimeexpectedregretexactmoment banditrlproof.ofulscheduledunboundedstoppingtimeexpectedregretexactmoment this module names the actual second moment of the stopping-time round count and uses it directly in the canonical terminal theorem. concrete stopping-rule analyses can subsequently bound this named quantity without rebuilding the stopped-regret argument. lean module compiled","shard":"modules/1175000d7c838a28.json"},{"id":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","label":"OFULScheduledUnboundedStoppingTimeExpectedRegretRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","description":"This module controls the random bad-event overflow from the exact unbounded stopping-time decomposition by an L2 Cauchy-Schwarz bound. It yields a valid sqrt delta overflow rate from the compiled stopped-event tail. This remains an event decomposition and does not invoke optional stopping.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretrate/index.html","parent":"chapter:oful","order":517,"meta":[["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretRate.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledunboundedstoppingtimeexpectedregretrate banditrlproof.ofulscheduledunboundedstoppingtimeexpectedregretrate this module controls the random bad-event overflow from the exact unbounded stopping-time decomposition by an l2 cauchy-schwarz bound. it yields a valid sqrt delta overflow rate from the compiled stopped-event tail. this remains an event decomposition and does not invoke optional stopping. lean module compiled","shard":"modules/59628a6743778641.json"},{"id":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","label":"OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","description":"This module integrates the quadratic stopped-budget envelope. The canonical terminal theorem therefore depends only on the supplied round-count second moment, not on an unevaluated expected stopped-budget term.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretsecondmoment/index.html","parent":"chapter:oful","order":518,"meta":[["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledunboundedstoppingtimeexpectedregretsecondmoment banditrlproof.ofulscheduledunboundedstoppingtimeexpectedregretsecondmoment this module integrates the quadratic stopped-budget envelope. the canonical terminal theorem therefore depends only on the supplied round-count second moment, not on an unevaluated expected stopped-budget term. lean module compiled","shard":"modules/d44da43c79de5426.json"},{"id":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","label":"OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","description":"This module derives the pathwise reach-by-budget premise from the more primitive contract that the accumulated Nat-valued resource grows by at least one at every round, then reuses the compiled budget-exhaustion OFUL theorem.","url":"../modules/banditrlproof-ofulscheduledunitgrowthbudgetexhaustionexpectedregret/index.html","parent":"chapter:oful","order":519,"meta":[["Source","BanditRLProof/OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulscheduledunitgrowthbudgetexhaustionexpectedregret banditrlproof.ofulscheduledunitgrowthbudgetexhaustionexpectedregret this module derives the pathwise reach-by-budget premise from the more primitive contract that the accumulated nat-valued resource grows by at least one at every round, then reuses the compiled budget-exhaustion oful theorem. lean module compiled","shard":"modules/0283d78b22bb68d3.json"},{"id":"module:BanditRLProof.OFULSelectedWidthSummation","label":"OFULSelectedWidthSummation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULSelectedWidthSummation","description":"This module converts the compiled logarithmic elliptical-potential endpoint into cumulative clipped confidence-width bounds for arbitrary feature and selected-action sequences. The route is deterministic and finite-horizon.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html","parent":"chapter:oful","order":520,"meta":[["Source","BanditRLProof/OFULSelectedWidthSummation.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulselectedwidthsummation banditrlproof.ofulselectedwidthsummation this module converts the compiled logarithmic elliptical-potential endpoint into cumulative clipped confidence-width bounds for arbitrary feature and selected-action sequences. the route is deterministic and finite-horizon. lean module compiled","shard":"modules/d258cb96e21daf12.json"},{"id":"module:BanditRLProof.OFULSelfNormalizedConfidence","label":"OFULSelfNormalizedConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULSelfNormalizedConfidence","description":"This module starts the probabilistic OFUL route after the deterministic elliptical-potential theorem. It first formalizes the predictable-projection conditional exponential inequality used by the method of mixtures. The multivariate Gaussian mixture identity and final self-normalized event bound remain separate until they are compiled locally.","url":"../modules/banditrlproof-ofulselfnormalizedconfidence/index.html","parent":"chapter:oful","order":521,"meta":[["Source","BanditRLProof/OFULSelfNormalizedConfidence.lean"],["Declarations","4"],["Project imports","3"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulselfnormalizedconfidence banditrlproof.ofulselfnormalizedconfidence this module starts the probabilistic oful route after the deterministic elliptical-potential theorem. it first formalizes the predictable-projection conditional exponential inequality used by the method of mixtures. the multivariate gaussian mixture identity and final self-normalized event bound remain separate until they are compiled locally. lean module compiled","shard":"modules/ec572ee7cb9ffe63.json"},{"id":"module:BanditRLProof.OFULSelfNormalizedMarkov","label":"OFULSelfNormalizedMarkov","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULSelfNormalizedMarkov","description":"This module consumes the evaluated Gaussian-mixture lintegral <= 1 and converts it into a probability bound. It then transports the paper-facing inverse-Gram/log-determinant bad event into the Markov threshold event.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html","parent":"chapter:oful","order":522,"meta":[["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofulselfnormalizedmarkov banditrlproof.ofulselfnormalizedmarkov this module consumes the evaluated gaussian-mixture lintegral <= 1 and converts it into a probability bound. it then transports the paper-facing inverse-gram/log-determinant bad event into the markov threshold event. lean module compiled","shard":"modules/61b98a7b71d70c75.json"},{"id":"module:BanditRLProof.OFULUniformTimeConfidence","label":"OFULUniformTimeConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OFULUniformTimeConfidence","description":"This module assembles the compiled deterministic-horizon scalar-ridge confidence theorem over every horizon in a finite inclusive window. It supports arbitrary confidence schedules with values in (0, 1] and an equal allocation whose budgets sum exactly to delta, giving failure probability at most delta.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html","parent":"chapter:oful","order":523,"meta":[["Source","BanditRLProof/OFULUniformTimeConfidence.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","OFUL"]],"statement":"","missing":[],"search":"ofuluniformtimeconfidence banditrlproof.ofuluniformtimeconfidence this module assembles the compiled deterministic-horizon scalar-ridge confidence theorem over every horizon in a finite inclusive window. it supports arbitrary confidence schedules with values in (0, 1] and an equal allocation whose budgets sum exactly to delta, giving failure probability at most delta. lean module compiled","shard":"modules/50b065f3471e8cdd.json"},{"id":"module:BanditRLProof.OpenProblems","label":"OpenProblems","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.OpenProblems","description":"Open problem registry","url":"../modules/banditrlproof-openproblems/index.html","parent":"chapter:frontier","order":524,"meta":[["Source","BanditRLProof/OpenProblems.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Frontier"]],"statement":"","missing":[],"search":"openproblems banditrlproof.openproblems open problem registry lean module compiled","shard":"modules/173ca7392389ffa8.json"},{"id":"module:BanditRLProof.PolicyMeasurability","label":"PolicyMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.PolicyMeasurability","description":"This module records a narrow MEAS-POLICY leaf: a measurable policy applied to a measurable history/context state yields a measurable action. It is only a measurability and predictability surface; it does not construct policy kernels, trajectory laws, reward laws, or adaptive regret theorems.","url":"../modules/banditrlproof-policymeasurability/index.html","parent":"chapter:probability","order":525,"meta":[["Source","BanditRLProof/PolicyMeasurability.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"policymeasurability banditrlproof.policymeasurability this module records a narrow meas-policy leaf: a measurable policy applied to a measurable history/context state yields a measurable action. it is only a measurability and predictability surface; it does not construct policy kernels, trajectory laws, reward laws, or adaptive regret theorems. lean module compiled","shard":"modules/bfb5aa112ca61654.json"},{"id":"module:BanditRLProof.PosteriorKernel","label":"PosteriorKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel","description":"This module records the narrow POSTERIOR-KERNEL leaf: a posterior over environments, indexed by the observed history, is represented as a Mathlib Markov kernel from histories to environments. It deliberately does not prove a Bayes formula, a regular-conditional-distribution existence theorem, Thompson probability matching, or Bayesian regret.","url":"../modules/banditrlproof-posteriorkernel/index.html","parent":"chapter:probability","order":526,"meta":[["Source","BanditRLProof/PosteriorKernel.lean"],["Declarations","18"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"posteriorkernel banditrlproof.posteriorkernel this module records the narrow posterior-kernel leaf: a posterior over environments, indexed by the observed history, is represented as a mathlib markov kernel from histories to environments. it deliberately does not prove a bayes formula, a regular-conditional-distribution existence theorem, thompson probability matching, or bayesian regret. lean module compiled","shard":"modules/a39773eeb1d9043d.json"},{"id":"module:BanditRLProof.PowerCutoffNormalization","label":"PowerCutoffNormalization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.PowerCutoffNormalization","description":"Exact normalization of the balanced polynomial cutoff into source powers.","url":"../modules/banditrlproof-powercutoffnormalization/index.html","parent":"chapter:foundations","order":527,"meta":[["Source","BanditRLProof/PowerCutoffNormalization.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"powercutoffnormalization banditrlproof.powercutoffnormalization exact normalization of the balanced polynomial cutoff into source powers. lean module compiled","shard":"modules/fb515e09d31724d7.json"},{"id":"module:BanditRLProof.PowerTailIntegral","label":"PowerTailIntegral","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.PowerTailIntegral","description":"Exact power-tail integration and balancing identities.","url":"../modules/banditrlproof-powertailintegral/index.html","parent":"chapter:foundations","order":528,"meta":[["Source","BanditRLProof/PowerTailIntegral.lean"],["Declarations","3"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"powertailintegral banditrlproof.powertailintegral exact power-tail integration and balancing identities. lean module compiled","shard":"modules/8dd7d79ea72bf61c.json"},{"id":"module:BanditRLProof.ProbabilityUnionBound","label":"ProbabilityUnionBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ProbabilityUnionBound","description":"Thin Mathlib-backed finite-union wrappers. These are outer-measure bounds: no event measurability or probability-measure assumption is required.","url":"../modules/banditrlproof-probabilityunionbound/index.html","parent":"chapter:probability","order":529,"meta":[["Source","BanditRLProof/ProbabilityUnionBound.lean"],["Declarations","3"],["Project imports","0"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"probabilityunionbound banditrlproof.probabilityunionbound thin mathlib-backed finite-union wrappers. these are outer-measure bounds: no event measurability or probability-measure assumption is required. lean module compiled","shard":"modules/807bc11a7f36bc43.json"},{"id":"module:BanditRLProof.PullCountDecomposition","label":"PullCountDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.PullCountDecomposition","description":"This module contains deterministic count identities that consume the Mathlib-backed Finset.range wrappers. It stays below the probability and algorithm-specific layers.","url":"../modules/banditrlproof-pullcountdecomposition/index.html","parent":"chapter:foundations","order":530,"meta":[["Source","BanditRLProof/PullCountDecomposition.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"pullcountdecomposition banditrlproof.pullcountdecomposition this module contains deterministic count identities that consume the mathlib-backed finset.range wrappers. it stays below the probability and algorithm-specific layers. lean module compiled","shard":"modules/0d00fbec8502d108.json"},{"id":"module:BanditRLProof.PullCountReindex","label":"PullCountReindex","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.PullCountReindex","description":"Generated source map for BanditRLProof/PullCountReindex.lean.","url":"../modules/banditrlproof-pullcountreindex/index.html","parent":"chapter:foundations","order":531,"meta":[["Source","BanditRLProof/PullCountReindex.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"pullcountreindex banditrlproof.pullcountreindex generated source map for banditrlproof/pullcountreindex.lean. lean module compiled","shard":"modules/709c9ec5e35c3a52.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","label":"RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","description":"This module supplies the statistical producer required by the cumulative count-radius planner. Each raw episode-batch count is centered by its history-kernel integral. The exact adaptive iid batch law then gives a conditionally sub-Gaussian increment with the within-batch Bernoulli proxy, rather than the weaker whole-batch bounded-range proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html","parent":"chapter:finite-horizon-rl","order":532,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean"],["Declarations","46"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativecountmartingaleconfidence banditrlproof.rl.finitehorizonadaptivecumulativecountmartingaleconfidence this module supplies the statistical producer required by the cumulative count-radius planner. each raw episode-batch count is centered by its history-kernel integral. the exact adaptive iid batch law then gives a conditionally sub-gaussian increment with the within-batch bernoulli proxy, rather than the weaker whole-batch bounded-range proxy. lean module compiled","shard":"modules/d537a32c5a68b945.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","description":"This module closes the fixed-exploration residual charge in the exploratory behavior-regret route. It starts from one path-support visit floor at full exploration, proves the exact stagewise power scaling at a smaller exploration rate, and chooses","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html","parent":"chapter:finite-horizon-rl","order":533,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean"],["Declarations","26"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency banditrlproof.rl.finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency this module closes the fixed-exploration residual charge in the exploratory behavior-regret route. it starts from one path-support visit floor at full exploration, proves the exact stagewise power scaling at a smaller exploration rate, and chooses lean module compiled","shard":"modules/c2c0349d2431cbfc.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","description":"This module closes the scalar asymptotic boundary left by the finite-window realized-regret transport. The coarse whole-batch return proxy simplifies exactly, so the scheduled normalized return radius is bounded by 2 * horizon / (n + 2) and tends to zero independently of the scheduled batch size. Adding this radius to the compiled exploratory-behavior expected-regret certificate yields a realized certificate tending…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html","parent":"chapter:finite-horizon-rl","order":534,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency banditrlproof.rl.finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency this module closes the scalar asymptotic boundary left by the finite-window realized-regret transport. the coarse whole-batch return proxy simplifies exactly, so the scheduled normalized return radius is bounded by 2 * horizon / (n + 2) and tends to zero independently of the scheduled batch size. adding this radius to the compiled exploratory-behavior expected-regret certificate yields a realized certificate tending to zero, while the union of the count and return events has a two-share failure budget tending to zero. lean module compiled","shard":"modules/277247cf8ab0a5f8.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","label":"RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","description":"This module changes the adaptive empirical planner from a latest-batch summary to the sum of every transition count in the observed finite prefix. A nonnegative antitone count-radius object makes the plan's transition radius a function of the cumulative state-action visit count. The cumulative selector is measurable, so it also defines a concrete exploratory adaptive batch source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html","parent":"chapter:finite-horizon-rl","order":535,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean"],["Declarations","23"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeempiricaloptimisticregret banditrlproof.rl.finitehorizonadaptivecumulativeempiricaloptimisticregret this module changes the adaptive empirical planner from a latest-batch summary to the sum of every transition count in the observed finite prefix. a nonnegative antitone count-radius object makes the plan's transition radius a function of the cumulative state-action visit count. the cumulative selector is measurable, so it also defines a concrete exploratory adaptive batch source. lean module compiled","shard":"modules/b000410b6d687ac3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","description":"The preceding episodewise route gives one sharp certificate for every scheduled finite window, but those windows have different trajectory types. This module places the finite-window laws on one dependent infinite product. Coordinate n therefore has exactly the compiled adaptive trajectory law for schedule n, and the scheduled realized-regret coordinates form one random process.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html","parent":"chapter:finite-horizon-rl","order":536,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency banditrlproof.rl.finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency the preceding episodewise route gives one sharp certificate for every scheduled finite window, but those windows have different trajectory types. this module places the finite-window laws on one dependent infinite product. coordinate n therefore has exactly the compiled adaptive trajectory law for schedule n, and the scheduled realized-regret coordinates form one random process. lean module compiled","shard":"modules/b1dfbc5b216b0574.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","description":"This module strengthens the independent-coordinate common-space convergence in probability theorem to convergence of expected absolute realized-behavior regret. The additional input is an almost-everywhere deterministic envelope: generated adaptive batches are reward-consistent, so every scheduled realized average regret has absolute value at most twice the horizon.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html","parent":"chapter:finite-horizon-rl","order":537,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency banditrlproof.rl.finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency this module strengthens the independent-coordinate common-space convergence in probability theorem to convergence of expected absolute realized-behavior regret. the additional input is an almost-everywhere deterministic envelope: generated adaptive batches are reward-consistent, so every scheduled realized average regret has absolute value at most twice the horizon. lean module compiled","shard":"modules/0a573959ce6556ae.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","label":"RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","description":"This module packages the compiled expected absolute realized-regret convergence in Mathlib's native MemLp, eLpNorm, and Lp interfaces. The underlying probability space is still the independent product of complete scheduled finite-window experiments; no nested causal coupling is constructed here.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html","parent":"chapter:finite-horizon-rl","order":538,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency banditrlproof.rl.finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency this module packages the compiled expected absolute realized-regret convergence in mathlib's native memlp, elpnorm, and lp interfaces. the underlying probability space is still the independent product of complete scheduled finite-window experiments; no nested causal coupling is constructed here. lean module compiled","shard":"modules/2f23217ea804e4e1.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","description":"This module sharpens the finite-window realized-return route by preserving the iid structure inside each generated batch. Complete episodes are independent; stages inside one episode are not. Centering one bounded full-episode return at a time gives the batch proxy episodes * horizon^2, replacing the coarse whole-batch proxy (episodes * horizon)^2.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html","parent":"chapter:finite-horizon-rl","order":539,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean"],["Declarations","33"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency banditrlproof.rl.finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency this module sharpens the finite-window realized-return route by preserving the iid structure inside each generated batch. complete episodes are independent; stages inside one episode are not. centering one bounded full-episode return at a time gives the batch proxy episodes * horizon^2, replacing the coarse whole-batch proxy (episodes * horizon)^2. lean module compiled","shard":"modules/a8dc325512317251.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","label":"RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","description":"This module transports cumulative expected-regret certificates from deterministic recommended policies to the exploratory behavior policies actually used by the adaptive source. The explicit charge is linear in the exploration rate and quadratic in the horizon. For a fixed exploration rate, the vanishing statistical certificate therefore converges to that charge, not to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html","parent":"chapter:finite-horizon-rl","order":540,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean"],["Declarations","20"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeexploratorybehaviorregret banditrlproof.rl.finitehorizonadaptivecumulativeexploratorybehaviorregret this module transports cumulative expected-regret certificates from deterministic recommended policies to the exploratory behavior policies actually used by the adaptive source. the explicit charge is linear in the exploration rate and quadratic in the horizon. for a fixed exploration rate, the vanishing statistical certificate therefore converges to that charge, not to zero. lean module compiled","shard":"modules/539fd9fbc8cd3b09.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","label":"RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","description":"This module records the first generated-process leaves for the Hoeffding UCB-VI route. It adds the reward-sum component that was absent from the existing cumulative transition-count state, proves exact prefix updates and measurability, and defines a one-episode-at-a-time adaptive source whose next policy is the cumulative empirical optimistic policy with a clipped inverse-square-root Hoeffding radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html","parent":"chapter:finite-horizon-rl","order":541,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean"],["Declarations","36"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativehoeffdingucbvi banditrlproof.rl.finitehorizonadaptivecumulativehoeffdingucbvi this module records the first generated-process leaves for the hoeffding ucb-vi route. it adds the reward-sum component that was absent from the existing cumulative transition-count state, proves exact prefix updates and measurability, and defines a one-episode-at-a-time adaptive source whose next policy is the cumulative empirical optimistic policy with a clipped inverse-square-root hoeffding radius. lean module compiled","shard":"modules/56041dd13028c84c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","description":"This module chooses an explicit natural number of exploratory episodes per batch. The schedule simultaneously clears the normalized calibration threshold and places the logarithmic factor below one batch's visit mass. The resulting scalar average recommendation-regret bound is at most a fixed constant times 1 / sqrt(rounds) and therefore tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html","parent":"chapter:finite-horizon-rl","order":542,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeinversesqrtaverageconsistency banditrlproof.rl.finitehorizonadaptivecumulativeinversesqrtaverageconsistency this module chooses an explicit natural number of exploratory episodes per batch. the schedule simultaneously clears the normalized calibration threshold and places the logarithmic factor below one batch's visit mass. the resulting scalar average recommendation-regret bound is at most a fixed constant times 1 / sqrt(rounds) and therefore tends to zero. lean module compiled","shard":"modules/57c1674abe5c367c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","label":"RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","description":"This module divides the normalized cumulative recommendation-regret endpoint by a positive number of recommendation rounds. It rewrites the statistical term using the total number of exploratory episodes across all batches while preserving the parent event, probability tail, and optimism statement.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaveragerate/index.html","parent":"chapter:finite-horizon-rl","order":543,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeinversesqrtaveragerate banditrlproof.rl.finitehorizonadaptivecumulativeinversesqrtaveragerate this module divides the normalized cumulative recommendation-regret endpoint by a positive number of recommendation rounds. it rewrites the statistical term using the total number of exploratory episodes across all batches while preserving the parent event, probability tail, and optimism statement. lean module compiled","shard":"modules/68d37f57cb8b326f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","label":"RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","description":"This module calibrates the cumulative count-martingale confidence producer to a concrete count-dependent optimistic planner. Path-support exploration gives every adaptive batch a common predictable visit floor. Outside the compiled global count event, the accumulated realized visits exceed that predictable floor minus the cumulative confidence radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html","parent":"chapter:finite-horizon-rl","order":544,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean"],["Declarations","16"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeinversesqrtcalibration banditrlproof.rl.finitehorizonadaptivecumulativeinversesqrtcalibration this module calibrates the cumulative count-martingale confidence producer to a concrete count-dependent optimistic planner. path-support exploration gives every adaptive batch a common predictable visit floor. outside the compiled global count event, the accumulated realized visits exceed that predictable floor minus the cumulative confidence radius. lean module compiled","shard":"modules/5729df991a0898dd.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","label":"RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","description":"This module constructs the two-scale path calibration from deterministic episode and scale inequalities. It also sums the resulting capped inverse-square-root round envelopes and feeds the closed form into the existing optimism/recommended-policy expected-regret terminal.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html","parent":"chapter:finite-horizon-rl","order":545,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeinversesqrtexplicitrate banditrlproof.rl.finitehorizonadaptivecumulativeinversesqrtexplicitrate this module constructs the two-scale path calibration from deterministic episode and scale inequalities. it also sums the resulting capped inverse-square-root round envelopes and feeds the closed form into the existing optimism/recommended-policy expected-regret terminal. lean module compiled","shard":"modules/ea9784a14729365e.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","label":"RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","description":"This module specializes the scheduled average recommendation-regret route to the confidence budget delta_n = 1 / (n + 2). Both the failure budget and the deterministic average-regret certificate tend to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html","parent":"chapter:finite-horizon-rl","order":546,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency banditrlproof.rl.finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency this module specializes the scheduled average recommendation-regret route to the confidence budget delta_n = 1 / (n + 2). both the failure budget and the deterministic average-regret certificate tend to zero. lean module compiled","shard":"modules/a45d7177ee4135c4.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","label":"RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","description":"This module specializes the explicit two-scale calibration to rewards bounded in absolute value by one. It fixes the zero-count budget to one, chooses the inverse-square-root scale from the logarithmic factor and visit floor, and replaces both scalar calibration premises by one episode threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html","parent":"chapter:finite-horizon-rl","order":547,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeinversesqrtnormalizedrate banditrlproof.rl.finitehorizonadaptivecumulativeinversesqrtnormalizedrate this module specializes the explicit two-scale calibration to rewards bounded in absolute value by one. it fixes the zero-count budget to one, chooses the inverse-square-root scale from the logarithmic factor and visit floor, and replaces both scalar calibration premises by one episode threshold. lean module compiled","shard":"modules/31c20cbc996b71d9.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","label":"RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","description":"This module transports the compiled adaptive exploratory-policy expected-regret route to realized rewards on the same infinite episode-batch trajectory law. Successor coordinates 1 through rounds are charged; coordinate zero is the uncontrolled initial batch and is excluded.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html","parent":"chapter:finite-horizon-rl","order":548,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean"],["Declarations","46"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativerealizedbehaviorregret banditrlproof.rl.finitehorizonadaptivecumulativerealizedbehaviorregret this module transports the compiled adaptive exploratory-policy expected-regret route to realized rewards on the same infinite episode-batch trajectory law. successor coordinates 1 through rounds are charged; coordinate zero is the uncontrolled initial batch and is excluded. lean module compiled","shard":"modules/bece1f714ca529d9.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","description":"Predictable episode-level Bellman innovations on the recurrent source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html","parent":"chapter:finite-horizon-rl","order":549,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean"],["Declarations","31"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale banditrlproof.rl.finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale predictable episode-level bellman innovations on the recurrent source. lean module compiled","shard":"modules/2489d370681ce83c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","description":"UCBVI-CH pools every visit of a state-action pair across stages. The earlier foundation exposed the aggregate denominator N_k(x,a) but retained only the stage-indexed transition numerator. This module closes that structural gap: it defines N_k(x,a,y), proves that its row sum is exactly the aggregate visit count, proves exact generated-prefix and successor identities, and normalizes the row into a probability kernel…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html","parent":"chapter:finite-horizon-rl","order":550,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviaggregatetransition banditrlproof.rl.finitehorizonadaptivecumulativeucbviaggregatetransition ucbvi-ch pools every visit of a state-action pair across stages. the earlier foundation exposed the aggregate denominator n_k(x,a) but retained only the stage-indexed transition numerator. this module closes that structural gap: it defines n_k(x,a,y), proves that its row sum is exactly the aggregate visit count, proves exact generated-prefix and successor identities, and normalizes the row into a probability kernel with an explicit dirac fallback at zero. lean module compiled","shard":"modules/92a9beb680429b6f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","description":"This module proves, rather than assumes, that each successor batch on the adaptive Kernel.trajMeasure is the literal image of a trajectory generated by the deterministic recurrent policy computed from its strict prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvialignment/index.html","parent":"chapter:finite-horizon-rl","order":551,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAlignment.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvialignment banditrlproof.rl.finitehorizonadaptivecumulativeucbvialignment this module proves, rather than assumes, that each successor batch on the adaptive kernel.trajmeasure is the literal image of a trajectory generated by the deterministic recurrent policy computed from its strict prefix. lean module compiled","shard":"modules/c24ee449f56274f6.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","description":"This file proves the within-episode martingale bound for an arbitrary fixed chronological table of continuation gaps in [0,H]. The proof recurses through the actual policy trajectory kernel one transition at a time, so its variance budget is H * H^2 / 4, rather than the invalid whole-episode range-square budget H^4 / 4.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html","parent":"chapter:finite-horizon-rl","order":552,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean"],["Declarations","24"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvibellmaninnovation banditrlproof.rl.finitehorizonadaptivecumulativeucbvibellmaninnovation this file proves the within-episode martingale bound for an arbitrary fixed chronological table of continuation gaps in [0,h]. the proof recurses through the actual policy trajectory kernel one transition at a time, so its variance budget is h * h^2 / 4, rather than the invalid whole-episode range-square budget h^4 / 4. lean module compiled","shard":"modules/9e286f533f4df7e7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","description":"The Hoeffding producer records the exact generated visit compensator. The UCBVI-CH analysis additionally needs the Bernoulli variance of each next-state coordinate. This module proves that stronger fixed-tilt leaf from the same transition kernel. It is not an independent sample model.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html","parent":"chapter:finite-horizon-rl","order":553,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvibernsteinconfidence banditrlproof.rl.finitehorizonadaptivecumulativeucbvibernsteinconfidence the hoeffding producer records the exact generated visit compensator. the ucbvi-ch analysis additionally needs the bernoulli variance of each next-state coordinate. this module proves that stronger fixed-tilt leaf from the same transition kernel. it is not an independent sample model. lean module compiled","shard":"modules/ecc80b678ea37274.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","description":"Canonical generated-record alignment and UCBVI-CH charge summation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html","parent":"chapter:finite-horizon-rl","order":554,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean"],["Declarations","31"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvichargesummation banditrlproof.rl.finitehorizonadaptivecumulativeucbvichargesummation canonical generated-record alignment and ucbvi-ch charge summation. lean module compiled","shard":"modules/4ac7987599d9953d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","description":"This module defines the actual recurrent planner used by the Chapter 9 route. Coordinate zero uses the all-H initial Q table and its fixed finite argmax. After observing coordinates 0,...,n, the successor planner folds those exact transition summaries, normalizes their cross-stage aggregate row, and applies the backward recurrence","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html","parent":"chapter:finite-horizon-rl","order":555,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean"],["Declarations","21"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviclippedplanner banditrlproof.rl.finitehorizonadaptivecumulativeucbviclippedplanner this module defines the actual recurrent planner used by the chapter 9 route. coordinate zero uses the all-h initial q table and its fixed finite argmax. after observing coordinates 0,...,n, the successor planner folds those exact transition summaries, normalizes their cross-stage aggregate row, and applies the backward recurrence lean module compiled","shard":"modules/255177b2865b2c62.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","description":"Fixed-tilt arithmetic used by the finite UCBVI confidence union.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html","parent":"chapter:finite-horizon-rl","order":556,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviconfidencetuning banditrlproof.rl.finitehorizonadaptivecumulativeucbviconfidencetuning fixed-tilt arithmetic used by the finite ucbvi confidence union. lean module compiled","shard":"modules/7af76fc4fbb4e66a.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","description":"This file turns the variance-sensitive residual stored by the simultaneous same-source event into the literal empirical transition mass used by the planner. The denominator is the actual pooled generated visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicoordinatealignment/index.html","parent":"chapter:finite-horizon-rl","order":557,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvicoordinatealignment banditrlproof.rl.finitehorizonadaptivecumulativeucbvicoordinatealignment this file turns the variance-sensitive residual stored by the simultaneous same-source event into the literal empirical transition mass used by the planner. the denominator is the actual pooled generated visit count. lean module compiled","shard":"modules/7baff7bd1aabc720.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","description":"The policy is frozen within each length-H episode, so every visit in that episode uses the same strict-prefix count. These lemmas retain that batching exactly: low-count overshoot costs at most one batch, while positive-count terms telescope through square-root and logarithmic potentials.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html","parent":"chapter:finite-horizon-rl","order":558,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvicounting banditrlproof.rl.finitehorizonadaptivecumulativeucbvicounting the policy is frozen within each length-h episode, so every visit in that episode uses the same strict-prefix count. these lemmas retain that batching exactly: low-count overshoot costs at most one batch, while positive-count terms telescope through square-root and logarithmic potentials. lean module compiled","shard":"modules/9d51a190447b4fc3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","description":"Generated episode pseudo-regret for the canonical recurrent source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html","parent":"chapter:finite-horizon-rl","order":559,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviepisoderegret banditrlproof.rl.finitehorizonadaptivecumulativeucbviepisoderegret generated episode pseudo-regret for the canonical recurrent source. lean module compiled","shard":"modules/772b4376bbc41857.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","description":"This file first closes the small measurability bridge from the recurrent finite table to policy-value pseudo-regret. It then integrates the compiled high-probability terminal, charging the deterministic K H envelope only on the measurable hull of the proved failure event. The resulting corollary therefore retains the required K H delta term.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviexpectedregret/index.html","parent":"chapter:finite-horizon-rl","order":560,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviexpectedregret banditrlproof.rl.finitehorizonadaptivecumulativeucbviexpectedregret this file first closes the small measurability bridge from the recurrent finite table to policy-value pseudo-regret. it then integrates the compiled high-probability terminal, charging the deterministic k h envelope only on the measurable hull of the proved failure event. the resulting corollary therefore retains the required k h delta term. lean module compiled","shard":"modules/6941bfb05ef21651.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","description":"The declarations in this file are deterministic. The following generated source layer discharges their two model-error inputs from the compiled joint transition event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html","parent":"chapter:finite-horizon-rl","order":561,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvilocalbellman banditrlproof.rl.finitehorizonadaptivecumulativeucbvilocalbellman the declarations in this file are deterministic. the following generated source layer discharges their two model-error inputs from the compiled joint transition event. lean module compiled","shard":"modules/6b56a3a95d9a8545.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","description":"Tuned Bellman-innovation tail for canonical recurrent UCBVI-CH.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html","parent":"chapter:finite-horizon-rl","order":562,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvimartingaletuning banditrlproof.rl.finitehorizonadaptivecumulativeucbvimartingaletuning tuned bellman-innovation tail for canonical recurrent ucbvi-ch. lean module compiled","shard":"modules/dc864b477258d0b5.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","description":"Exact affine transport from the normalized generated probe to V*.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimaltailalignment/index.html","parent":"chapter:finite-horizon-rl","order":563,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvioptimaltailalignment banditrlproof.rl.finitehorizonadaptivecumulativeucbvioptimaltailalignment exact affine transport from the normalized generated probe to v*. lean module compiled","shard":"modules/b50caa3959474ba5.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","description":"Bellman optimism of one clipped recurrent UCBVI-CH update.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html","parent":"chapter:finite-horizon-rl","order":564,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvioptimism banditrlproof.rl.finitehorizonadaptivecumulativeucbvioptimism bellman optimism of one clipped recurrent ucbvi-ch update. lean module compiled","shard":"modules/4749049da345191f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","description":"Paper-scale probability arithmetic for canonical recurrent UCBVI-CH.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html","parent":"chapter:finite-horizon-rl","order":565,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviprobabilitybudget banditrlproof.rl.finitehorizonadaptivecumulativeucbviprobabilitybudget paper-scale probability arithmetic for canonical recurrent ucbvi-ch. lean module compiled","shard":"modules/aaf9cea6994cf22b.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","description":"Bellman optimism for every queried policy of the one recurrent source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvirecurrentoptimism/index.html","parent":"chapter:finite-horizon-rl","order":566,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvirecurrentoptimism banditrlproof.rl.finitehorizonadaptivecumulativeucbvirecurrentoptimism bellman optimism for every queried policy of the one recurrent source. lean module compiled","shard":"modules/d7010abdf823bb53.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","description":"Pathwise recurrent UCBVI-CH episode-regret decomposition.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html","parent":"chapter:finite-horizon-rl","order":567,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean"],["Declarations","23"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviregretdecomposition banditrlproof.rl.finitehorizonadaptivecumulativeucbviregretdecomposition pathwise recurrent ucbvi-ch episode-regret decomposition. lean module compiled","shard":"modules/2bc011b11da9fa3a.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","description":"This module starts the statistical producer on the actual finite episode law. For a fixed state/action/next-state coordinate, the residual at a generated stage is","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html","parent":"chapter:finite-horizon-rl","order":568,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean"],["Declarations","42"],["Project imports","4"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvisamesourceconfidence banditrlproof.rl.finitehorizonadaptivecumulativeucbvisamesourceconfidence this module starts the statistical producer on the actual finite episode law. for a fixed state/action/next-state coordinate, the residual at a generated stage is lean module compiled","shard":"modules/02eed265a80dad25.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","description":"The family contains two genuinely generated transition coordinates:","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html","parent":"chapter:finite-horizon-rl","order":569,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvisimultaneousconfidence banditrlproof.rl.finitehorizonadaptivecumulativeucbvisimultaneousconfidence the family contains two genuinely generated transition coordinates: lean module compiled","shard":"modules/211bebb9426c7476.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","description":"The terminal below uses one recurrent source, its exact Kernel.trajMeasure, the confidence family proved on that law, the policy-value regret generated by the same strict-prefix planner, and no caller-supplied confidence premise.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html","parent":"chapter:finite-horizon-rl","order":570,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbviterminal banditrlproof.rl.finitehorizonadaptivecumulativeucbviterminal the terminal below uses one recurrent source, its exact kernel.trajmeasure, the confidence family proved on that law, the policy-value regret generated by the same strict-prefix planner, and no caller-supplied confidence premise. lean module compiled","shard":"modules/774753115e42f47c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","label":"RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","description":"Singleton transition coordinates do not by themselves imply a sharp bound for (P_hat - P) V* without a sqrt |State| loss. This module therefore proves the bounded scalar projection as another coordinate of the same generated transition residual family. It never introduces an offline or independent sample law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html","parent":"chapter:finite-horizon-rl","order":571,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean"],["Declarations","29"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivecumulativeucbvitransitionvalueconfidence banditrlproof.rl.finitehorizonadaptivecumulativeucbvitransitionvalueconfidence singleton transition coordinates do not by themselves imply a sharp bound for (p_hat - p) v* without a sqrt |state| loss. this module therefore proves the bounded scalar projection as another coordinate of the same generated transition residual family. it never introduces an offline or independent sample law. lean module compiled","shard":"modules/7a25aa1e1906511c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","label":"RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","description":"This module calibrates a latest-batch empirical-transition source whose behavior policy uniformly explores around the latest optimistic deterministic table. For every policy that generates a batch, a local contract records genuine expected-visit margins and a finite-state coordinate-radius cover for the fixed transition bonus. Outside the existing adaptive simultaneous-count event, those contracts produce coordinate…","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html","parent":"chapter:finite-horizon-rl","order":572,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptiveempiricaloptimisticconfidence banditrlproof.rl.finitehorizonadaptiveempiricaloptimisticconfidence this module calibrates a latest-batch empirical-transition source whose behavior policy uniformly explores around the latest optimistic deterministic table. for every policy that generates a batch, a local contract records genuine expected-visit margins and a finite-state coordinate-radius cover for the fixed transition bonus. outside the existing adaptive simultaneous-count event, those contracts produce coordinate confidence for every known-reward empirical plan. lean module compiled","shard":"modules/1b73d33366a0cecb.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","label":"RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","description":"The current known-reward empirical optimistic plan uses zero reward radius and one fixed transition bonus at every coordinate. This module evaluates its recursive occupancy-radius sum exactly and attaches that explicit envelope to the compiled path-support episode-threshold event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticoccupancyenvelope/index.html","parent":"chapter:finite-horizon-rl","order":573,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptiveempiricaloptimisticoccupancyenvelope banditrlproof.rl.finitehorizonadaptiveempiricaloptimisticoccupancyenvelope the current known-reward empirical optimistic plan uses zero reward radius and one fixed transition bonus at every coordinate. this module evaluates its recursive occupancy-radius sum exactly and attaches that explicit envelope to the compiled path-support episode-threshold event. lean module compiled","shard":"modules/b84f7effecf44232.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","label":"RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","description":"This module constructs the measurable source left abstract by FiniteHorizonAdaptiveEpisodeBatchLaw. Every observed batch is compressed to its finite family of transition counts. Those counts define a normalized empirical transition kernel, which is combined with the known deterministic MDP reward and a fixed transition bonus. The resulting optimistic action table is a measurable function of the batch because the cou…","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html","parent":"chapter:finite-horizon-rl","order":574,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean"],["Declarations","26"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptiveempiricaloptimisticsource banditrlproof.rl.finitehorizonadaptiveempiricaloptimisticsource this module constructs the measurable source left abstract by finitehorizonadaptiveepisodebatchlaw. every observed batch is compressed to its finite family of transition counts. those counts define a normalized empirical transition kernel, which is combined with the known deterministic mdp reward and a fixed transition bonus. the resulting optimistic action table is a measurable function of the batch because the count-summary space is countable with measurable singletons. lean module compiled","shard":"modules/8b1ce42ec8462187.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","label":"RL.FiniteHorizonAdaptiveEpisodeBatchLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","description":"This module replaces the independent finite product of episode batches by an Ionescu--Tulcea trajectory whose successor batch kernel may depend on the full finite batch history. A source records the selected Markov policy and the exact equality between its generated iid batch law and the configured history kernel. The resulting trajectory exposes both the regular conditional law of the next batch and a finite-horizo…","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html","parent":"chapter:finite-horizon-rl","order":575,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean"],["Declarations","27"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptiveepisodebatchlaw banditrlproof.rl.finitehorizonadaptiveepisodebatchlaw this module replaces the independent finite product of episode batches by an ionescu--tulcea trajectory whose successor batch kernel may depend on the full finite batch history. a source records the selected markov policy and the exact equality between its generated iid batch law and the configured history kernel. the resulting trajectory exposes both the regular conditional law of the next batch and a finite-horizon union bound for arbitrary measurable adapted bad events. lean module compiled","shard":"modules/2913f0f8ee26ce69.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","description":"The compiled stochastic cumulative route gives a sharp certificate on each scheduled finite-window trajectory space. This module places those complete finite-window experiments on one dependent infinite product and proves that the scheduled stochastic realized-regret process converges to zero in measure.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html","parent":"chapter:finite-horizon-rl","order":576,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationcommonspaceconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationcommonspaceconsistency the compiled stochastic cumulative route gives a sharp certificate on each scheduled finite-window trajectory space. this module places those complete finite-window experiments on one dependent infinite product and proves that the scheduled stochastic realized-regret process converges to zero in measure. lean module compiled","shard":"modules/306fe948ddbfcdcc.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","description":"This module strengthens the scheduled stochastic-reward common-space theorem from convergence in probability to convergence of the expected absolute realized-behavior regret and then to convergence in L1.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html","parent":"chapter:finite-horizon-rl","order":577,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean"],["Declarations","30"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationcommonspacel1consistency banditrlproof.rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationcommonspacel1consistency this module strengthens the scheduled stochastic-reward common-space theorem from convergence in probability to convergence of the expected absolute realized-behavior regret and then to convergence in l1. lean module compiled","shard":"modules/9a92f03a55ac7941.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","description":"This module lifts the cumulative empirical-optimistic exploratory source to complete stochastic-reward episode batches. Policy selection reads only the known-reward projection of the sampled prefix. The first layer proves the exact complete trajectory pushforward to the deterministic cumulative source; later layers consume the deterministic decaying-exploration count certificate and the globally centered stochastic…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html","parent":"chapter:finite-horizon-rl","order":578,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean"],["Declarations","32"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency this module lifts the cumulative empirical-optimistic exploratory source to complete stochastic-reward episode batches. policy selection reads only the known-reward projection of the sampled prefix. the first layer proves the exact complete trajectory pushforward to the deterministic cumulative source; later layers consume the deterministic decaying-exploration count certificate and the globally centered stochastic return tail. lean module compiled","shard":"modules/2a78859d916b17e8.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","description":"This module discharges the composite Standard Borel premises exposed by the finite-window and all-window realized-behavior consistency endpoints. The state and action spaces remain the only caller-facing Standard Borel contract; the finite batch and countable trajectory instances are inferred locally.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-54b2a9d89968/index.html","parent":"chapter:finite-horizon-rl","order":579,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationregularityclosedconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardcumulativedecayingexplorationregularityclosedconsistency this module discharges the composite standard borel premises exposed by the finite-window and all-window realized-behavior consistency endpoints. the state and action spaces remain the only caller-facing standard borel contract; the finite batch and countable trajectory instances are inferred locally. lean module compiled","shard":"modules/f37d4f8f41f1fe22.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","label":"RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","description":"This module lifts the existing known-reward exploratory empirical-transition source to stochastic rewards. Its policy selector reads only the complete known-reward projection of the observed stochastic prefix. The resulting adaptive stochastic trajectory therefore maps exactly to the deterministic source trajectory, so the compiled count-confidence, optimism, and recommended expected-regret terminal can be pulled ba…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html","parent":"chapter:finite-horizon-rl","order":580,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean"],["Declarations","21"],["Project imports","4"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardempiricaloptimisticprojection banditrlproof.rl.finitehorizonadaptivestochasticrewardempiricaloptimisticprojection this module lifts the existing known-reward exploratory empirical-transition source to stochastic rewards. its policy selector reads only the complete known-reward projection of the observed stochastic prefix. the resulting adaptive stochastic trajectory therefore maps exactly to the deterministic source trajectory, so the compiled count-confidence, optimism, and recommended expected-regret terminal can be pulled back without estimating sampled rewards. lean module compiled","shard":"modules/3fb09df8644b7887.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","description":"This module combines the concrete stochastic known-mean empirical-transition source with the global sampled-return transport. The count/optimism event and the sampled-return event retain separate confidence budgets. The expected regret of the exploratory behavior is charged explicitly against the projected recommended policy before the realized-return deviation is added.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html","parent":"chapter:finite-horizon-rl","order":581,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean"],["Declarations","9"],["Project imports","4"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret this module combines the concrete stochastic known-mean empirical-transition source with the global sampled-return transport. the count/optimism event and the sampled-return event retain separate confidence budgets. the expected regret of the exploratory behavior is charged explicitly against the projected recommended policy before the realized-return deviation is added. lean module compiled","shard":"modules/732d9425e49f22fc.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","description":"This module changes the stochastic episode return from centering at each sampled initial state's policy value to centering at the selected policy's global initial-law value. The missing term is the policy-value fluctuation of the sampled initial state. Complete episodes remain the independent units inside one batch; no independence is assumed between the two terms inside an episode or across adaptive rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html","parent":"chapter:finite-horizon-rl","order":582,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean"],["Declarations","53"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardrealizedbehaviorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardrealizedbehaviorregret this module changes the stochastic episode return from centering at each sampled initial state's policy value to centering at the selected policy's global initial-law value. the missing term is the policy-value fluctuation of the sampled initial state. complete episodes remain the independent units inside one batch; no independence is assumed between the two terms inside an episode or across adaptive rounds. lean module compiled","shard":"modules/7eb45ee6543142db.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","description":"This module transports the fixed-policy sampled count-and-reward empirical model event through the exact history-fiber laws of the adaptive stochastic source. The resulting finite-round event controls models computed from the actual sampled rewards, not their known-mean projection.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html","parent":"chapter:finite-horizon-rl","order":583,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean"],["Declarations","24"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence this module transports the fixed-policy sampled count-and-reward empirical model event through the exact history-fiber laws of the adaptive stochastic source. the resulting finite-round event controls models computed from the actual sampled rewards, not their known-mean projection. lean module compiled","shard":"modules/e52a060cdefebe8d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","description":"This module charges uniform exploration around the actual sampled-model recommendations. The policy built from coordinate round is the adaptive source policy for successor batch round + 1; the initial policy and realized sampled returns remain outside this theorem route.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html","parent":"chapter:finite-horizon-rl","order":584,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcumulativeexploratorybehaviorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcumulativeexploratorybehaviorregret this module charges uniform exploration around the actual sampled-model recommendations. the policy built from coordinate round is the adaptive source policy for successor batch round + 1; the initial policy and realized sampled returns remain outside this theorem route. lean module compiled","shard":"modules/3d1c93c4686a0400.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","description":"This module sums the finite-round, actual-sampled-model recommendation bounds from the adaptive stochastic confidence route. It does not add exploratory behavior or realized-return costs.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-48200eab5c41/index.html","parent":"chapter:finite-horizon-rl","order":585,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcumulativerecommendedregret banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcumulativerecommendedregret this module sums the finite-round, actual-sampled-model recommendation bounds from the adaptive stochastic confidence route. it does not add exploratory behavior or realized-return costs. lean module compiled","shard":"modules/9ae36733e5c8ca42.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","description":"This module evaluates the selected-radius occupancy term of the actual sampled-reward empirical optimistic plans. The resulting three-share terminal has a deterministic planning envelope, while retaining the globally centered sampled-return radius and excluding the initial batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html","parent":"chapter:finite-horizon-rl","order":586,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexplicitbudgetrealizedbehaviorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexplicitbudgetrealizedbehaviorregret this module evaluates the selected-radius occupancy term of the actual sampled-reward empirical optimistic plans. the resulting three-share terminal has a deterministic planning envelope, while retaining the globally centered sampled-return radius and excluding the initial batch. lean module compiled","shard":"modules/a75a4e61326820fd.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","description":"This module combines the actual sampled-model confidence and exploratory successor-policy route with the existing globally centered stochastic return tail. Model count, model reward, and realized-return failures retain three separate confidence shares. The result covers successor batches 1..rounds; the initial batch and explicit occupancy-radius rates remain downstream.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticrea-28c98ba81ea7/index.html","parent":"chapter:finite-horizon-rl","order":587,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticrealizedbehaviorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticrealizedbehaviorregret this module combines the actual sampled-model confidence and exploratory successor-policy route with the existing globally centered stochastic return tail. model count, model reward, and realized-return failures retain three separate confidence shares. the result covers successor batches 1..rounds; the initial batch and explicit occupancy-radius rates remain downstream. lean module compiled","shard":"modules/49b43e9e188d456d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","description":"This module proves that the actual exploratory successor policy's expected regret tends to zero along almost every trajectory of the one heterogeneous causal source. The key project-local leaf extracts a pointwise behavior- policy rate from the existing coordinate confidence proof. First Borel-Cantelli supplies eventual model-goodness, and the deterministic causal planning rate then squeezes the nonnegative behavior…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-20fa414d9371/index.html","parent":"chapter:finite-horizon-rl","order":588,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacealmostsurebehaviorexpectedregretconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacealmostsurebehaviorexpectedregretconsistency this module proves that the actual exploratory successor policy's expected regret tends to zero along almost every trajectory of the one heterogeneous causal source. the key project-local leaf extracts a pointwise behavior- policy rate from the existing coordinate confidence proof. first borel-cantelli supplies eventual model-goodness, and the deterministic causal planning rate then squeezes the nonnegative behavior regret to zero. lean module compiled","shard":"modules/675fd8e299ce3279.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","description":"This module upgrades the heterogeneous natural-causal realized-regret process from L1 convergence to almost-sure convergence. The proof keeps the same dependent trajectory measure. It applies the first Borel-Cantelli lemma to the summable coordinate model events and to the mass-adapted successor-return events, then consumes the existing fixed-burn-in absolute-regret envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html","parent":"chapter:finite-horizon-rl","order":589,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacealmostsureconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacealmostsureconsistency this module upgrades the heterogeneous natural-causal realized-regret process from l1 convergence to almost-sure convergence. the proof keeps the same dependent trajectory measure. it applies the first borel-cantelli lemma to the summable coordinate model events and to the mass-adapted successor-return events, then consumes the existing fixed-burn-in absolute-regret envelope. lean module compiled","shard":"modules/14c1bcee766d279d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","description":"This module turns the actual exploratory source.successorPolicyAt pointwise planning certificate into a finite-coordinate expectation bound on the same genuine heterogeneous dependent causal trajectory measure. A generic event split integrates a local bound off one measurable bad event and a global bound on it. For the natural causal source, the local term is the compiled planning rate, the global term is 2 * horizo…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html","parent":"chapter:finite-horizon-rl","order":590,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretexplicitintegratedrate banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretexplicitintegratedrate this module turns the actual exploratory source.successorpolicyat pointwise planning certificate into a finite-coordinate expectation bound on the same genuine heterogeneous dependent causal trajectory measure. a generic event split integrates a local bound off one measurable bad event and a global bound on it. for the natural causal source, the local term is the compiled planning rate, the global term is 2 * horizon, and the event has the compiled two-share coordinate model-confidence budget. lean module compiled","shard":"modules/c4ecfc2d77946946.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","description":"This module sums the actual source.successorPolicyAt expected-regret process over Finset.range rounds on the one genuine heterogeneous dependent causal trajectory measure. ExpectationBochnerSums.integral_finset_sum identifies the integral of that pathwise finite sum with the sum of coordinate expectations; coordinate nonnegativity then identifies those terms with the compiled expected absolute regrets. Summing the e…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html","parent":"chapter:finite-horizon-rl","order":591,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean"],["Declarations","17"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretfiniteprefixcumulativeaveragerate banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretfiniteprefixcumulativeaveragerate this module sums the actual source.successorpolicyat expected-regret process over finset.range rounds on the one genuine heterogeneous dependent causal trajectory measure. expectationbochnersums.integral_finset_sum identifies the integral of that pathwise finite sum with the sum of coordinate expectations; coordinate nonnegativity then identifies those terms with the compiled expected absolute regrets. summing the explicit integrated coordinate bounds gives a finite-prefix cumulative rate, and division by rounds gives its cesaro rate. the existing tendsto_natweightedaverage_zero theorem at unit natural weights pr lean module compiled","shard":"modules/248397c259af1e3c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","description":"This module closes the regularity boundary left by the a.e. behavior route. For a measurable finite policy-table selector, expected regret of the selected exploratory policy is measurable by a finite indicator-sum representation. The heterogeneous trajectory coordinate selector is already measurable, so every actual successor-policy expected-regret coordinate is measurable. Mathlib then transports the compiled a.e.…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-9d6602db783e/index.html","parent":"chapter:finite-horizon-rl","order":592,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretinmeasureconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretinmeasureconsistency this module closes the regularity boundary left by the a.e. behavior route. for a measurable finite policy-table selector, expected regret of the selected exploratory policy is measurable by a finite indicator-sum representation. the heterogeneous trajectory coordinate selector is already measurable, so every actual successor-policy expected-regret coordinate is measurable. mathlib then transports the compiled a.e. limit to tendstoinmeasure on the same genuine dependent causal trajectory measure. lean module compiled","shard":"modules/f62a2dca35fd9200.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","description":"This module upgrades the actual exploratory source.successorPolicyAt expected-regret process from a.e./in-measure convergence to L1 on the same genuine heterogeneous dependent causal trajectory measure. Every coordinate is measurable, nonnegative, and bounded by the deterministic 2 * horizon policy envelope. Mathlib dominated convergence therefore gives expected absolute convergence directly; no realized-return MGF,…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html","parent":"chapter:finite-horizon-rl","order":593,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretl1consistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacebehaviorexpectedregretl1consistency this module upgrades the actual exploratory source.successorpolicyat expected-regret process from a.e./in-measure convergence to l1 on the same genuine heterogeneous dependent causal trajectory measure. every coordinate is measurable, nonnegative, and bounded by the deterministic 2 * horizon policy envelope. mathlib dominated convergence therefore gives expected absolute convergence directly; no realized-return mgf, independent-window coupling, or extra uniform-integrability assumption is used. lean module compiled","shard":"modules/3121aa04f88d38e9.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","description":"Target theorem route: prove convergence in probability of realized successor- average regret on the single infinite causal sampled trajectory, rather than coupling separate finite-window experiments.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html","parent":"chapter:finite-horizon-rl","order":594,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean"],["Declarations","38"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspaceconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspaceconsistency target theorem route: prove convergence in probability of realized successor- average regret on the single infinite causal sampled trajectory, rather than coupling separate finite-window experiments. lean module compiled","shard":"modules/f810818bfcf103b6.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","description":"This route strengthens the compiled convergence-in-probability theorem on the single heterogeneous causal trajectory measure. It does not use the independent product of complete finite-window experiments.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html","parent":"chapter:finite-horizon-rl","order":595,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean"],["Declarations","27"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacel1consistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalcommonspacel1consistency this route strengthens the compiled convergence-in-probability theorem on the single heterogeneous causal trajectory measure. it does not use the independent product of complete finite-window experiments. lean module compiled","shard":"modules/8de342c11ea19258.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","description":"The causal source uses a genuinely round-varying batch size. Its regret rate therefore cannot reuse the constant-window episodes * rounds algebra. This module keeps the exact successor mass and proves that the corresponding positive-weight average of the coordinatewise scheduled rate tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html","parent":"chapter:finite-horizon-rl","order":596,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean"],["Declarations","20"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalexplicitrate banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalexplicitrate the causal source uses a genuinely round-varying batch size. its regret rate therefore cannot reuse the constant-window episodes * rounds algebra. this module keeps the exact successor mass and proves that the corresponding positive-weight average of the coordinatewise scheduled rate tends to zero. lean module compiled","shard":"modules/08096d6e6305c654.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","description":"This module transports actual sampled count-and-reward empirical-model events to the dependent source whose coordinate t contains episodes t complete stochastic episodes. Unlike the constant-window route, local count and reward shares are functions of the coordinate. The fixed-prefix failure budget is therefore their genuine finite ENNReal sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html","parent":"chapter:finite-horizon-rl","order":597,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean"],["Declarations","23"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalmodelconfidence banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalmodelconfidence this module transports actual sampled count-and-reward empirical-model events to the dependent source whose coordinate t contains episodes t complete stochastic episodes. unlike the constant-window route, local count and reward shares are functions of the coordinate. the fixed-prefix failure budget is therefore their genuine finite ennreal sum. lean module compiled","shard":"modules/3e6e674281b0bfa7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","description":"This module combines the actual-sampled empirical-model event with the successor-only globally centered return event on the genuinely dependent round-varying source. Coordinate n + 1 is generated by the exploratory policy selected from the prefix through n and contains episodes (n + 1) complete episodes. Expected and realized successor regret are therefore weighted by the actual coordinate batch sizes and normalized…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html","parent":"chapter:finite-horizon-rl","order":598,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean"],["Declarations","26"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalrealizedsuccessorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalrealizedsuccessorregret this module combines the actual-sampled empirical-model event with the successor-only globally centered return event on the genuinely dependent round-varying source. coordinate n + 1 is generated by the exploratory policy selected from the prefix through n and contains episodes (n + 1) complete episodes. expected and realized successor regret are therefore weighted by the actual coordinate batch sizes and normalized by their finite sum, never by a constant-window episodes * rounds denominator. lean module compiled","shard":"modules/f4d04e51161ffd03.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","description":"This module transports two adaptive sampled-return concentration processes to the dependent source whose coordinate n contains episodes n complete stochastic episodes. The supporting process uses the initial policy and episodes 0 at coordinate zero, then the prefix-selected policy and episodes (n + 1) at coordinate n + 1; each batch is centered at its own sampled initial-state value. The regret-facing global process…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html","parent":"chapter:finite-horizon-rl","order":599,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean"],["Declarations","38"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalreturnconcentration banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalreturnconcentration this module transports two adaptive sampled-return concentration processes to the dependent source whose coordinate n contains episodes n complete stochastic episodes. the supporting process uses the initial policy and episodes 0 at coordinate zero, then the prefix-selected policy and episodes (n + 1) at coordinate n + 1; each batch is centered at its own sampled initial-state value. the regret-facing global process is zero at coordinate zero and globally centers successor coordinates 1..rounds by the initial-law expected value of the selected policy. its measurable-selector boundary is packaged by globalreturnme lean module compiled","shard":"modules/fc32965a64f21720.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","description":"The compiled self-consistent finite-window theorems vary the batch size and algorithm parameters with the outer window index. Their laws therefore are not prefixes of one fixed adaptive source. This module defines the distinct causal algorithm in which those parameters vary at each trajectory coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html","parent":"chapter:finite-horizon-rl","order":600,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean"],["Declarations","15"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalsource banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcausalsource the compiled self-consistent finite-window theorems vary the batch size and algorithm parameters with the outer window index. their laws therefore are not prefixes of one fixed adaptive source. this module defines the distinct causal algorithm in which those parameters vary at each trajectory coordinate. lean module compiled","shard":"modules/d1a06f4cd913d098.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","description":"Target theorem route: place the complete actual-sampled scheduled experiments on one exact-marginal dependent product space and prove that realized successor- average regret converges to zero in Mathlib TendstoInMeasure.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html","parent":"chapter:finite-horizon-rl","order":601,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcommonspaceconsistency banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentcommonspaceconsistency target theorem route: place the complete actual-sampled scheduled experiments on one exact-marginal dependent product space and prove that realized successor- average regret converges to zero in mathlib tendstoinmeasure. lean module compiled","shard":"modules/2c234f58b610d060.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","description":"This consumer keeps the scheduled source and all its regularity contracts unchanged. It combines the compiled bounds","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html","parent":"chapter:finite-horizon-rl","order":602,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentexplicitrate banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentexplicitrate this consumer keeps the scheduled source and all its regularity contracts unchanged. it combines the compiled bounds lean module compiled","shard":"modules/872311f9b2035ad5.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","description":"This module replaces the coarse transition budget rewardBound + 2 * rewardBudget by the exact fixed point generated by a contraction coefficient q < 1. It retains actual sampled rewards, adaptive policy selection, three independent confidence shares, and the globally centered successor-return transport.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a072b4ceb8e1/index.html","parent":"chapter:finite-horizon-rl","order":603,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentrealizedbehaviorregret banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentrealizedbehaviorregret this module replaces the coarse transition budget rewardbound + 2 * rewardbudget by the exact fixed point generated by a contraction coefficient q < 1. it retains actual sampled rewards, adaptive policy selection, three independent confidence shares, and the globally centered successor-return transport. lean module compiled","shard":"modules/62f926b73e242815.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","description":"This module gives the self-consistent actual-sampled route an explicit horizon-indexed batch schedule. It reuses the decaying exploration, round, visit-floor, and confidence schedules, and chooses one ceil + 1 episode count above three explicit thresholds: the existing count-calibration threshold, a shrinking count-ratio threshold, and a shrinking sampled-reward threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html","parent":"chapter:finite-horizon-rl","order":604,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean"],["Declarations","39"],["Project imports","4"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentschedule banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticselfconsistentschedule this module gives the self-consistent actual-sampled route an explicit horizon-indexed batch schedule. it reuses the decaying exploration, round, visit-floor, and confidence schedules, and chooses one ceil + 1 episode count above three explicit thresholds: the existing count-calibration threshold, a shrinking count-ratio threshold, and a shrinking sampled-reward threshold. lean module compiled","shard":"modules/9aab168693a59424.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","label":"RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","description":"This module constructs a genuinely sampled-reward adaptive optimistic source. The successor table is computed from the latest complete stochastic batch, including its observed rewards, rather than from the known-mean projection.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html","parent":"chapter:finite-horizon-rl","order":605,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource banditrlproof.rl.finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource this module constructs a genuinely sampled-reward adaptive optimistic source. the successor table is computed from the latest complete stochastic batch, including its observed rewards, rather than from the known-mean projection. lean module compiled","shard":"modules/13af6ace591b20c3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","label":"RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","description":"This module generates complete reward-bearing episode batches with a policy selected from the preceding batch history. It retains that history while mapping the next-batch kernel, identifies the resulting dynamic sampled-return law through condDistrib and condExpKernel, and applies the conditional sub-Gaussian sum theorem across a fixed finite number of adaptive rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html","parent":"chapter:finite-horizon-rl","order":606,"meta":[["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean"],["Declarations","29"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonadaptivestochasticrewardtotalreturnconcentration banditrlproof.rl.finitehorizonadaptivestochasticrewardtotalreturnconcentration this module generates complete reward-bearing episode batches with a policy selected from the preceding batch history. it retains that history while mapping the next-batch kernel, identifies the resulting dynamic sampled-return law through conddistrib and condexpkernel, and applies the conditional sub-gaussian sum theorem across a fixed finite number of adaptive rounds. lean module compiled","shard":"modules/84480fd4aa05d31d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","label":"RL.FiniteHorizonCoordinateModelConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","description":"This module turns finite-state singleton transition-mass errors into the transition-expectation confidence consumed by the estimated-model optimistic regret route. The only value regularity required is an absolute envelope for the recursively generated tail upper value.","url":"../modules/banditrlproof-rl-finitehorizoncoordinatemodelconfidence/index.html","parent":"chapter:finite-horizon-rl","order":607,"meta":[["Source","BanditRLProof/RL/FiniteHorizonCoordinateModelConfidence.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoncoordinatemodelconfidence banditrlproof.rl.finitehorizoncoordinatemodelconfidence this module turns finite-state singleton transition-mass errors into the transition-expectation confidence consumed by the estimated-model optimistic regret route. the only value regularity required is an absolute envelope for the recursively generated tail upper value. lean module compiled","shard":"modules/e4a4d6678bda70c7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","label":"RL.FiniteHorizonEmpiricalModel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonEmpiricalModel","description":"This module builds a genuine finite-state empirical transition kernel from a finite family of recorded episode steps. Positive visit counts are normalized into a finite PMF; zero visit counts use an explicit default-state Dirac PMF. The resulting empirical reward and transition model is then connected to the compiled coordinate-confidence optimistic-regret route.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html","parent":"chapter:finite-horizon-rl","order":608,"meta":[["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonempiricalmodel banditrlproof.rl.finitehorizonempiricalmodel this module builds a genuine finite-state empirical transition kernel from a finite family of recorded episode steps. positive visit counts are normalized into a finite pmf; zero visit counts use an explicit default-state dirac pmf. the resulting empirical reward and transition model is then connected to the compiled coordinate-confidence optimistic-regret route. lean module compiled","shard":"modules/7783b9dec15003e8.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","label":"RL.FiniteHorizonEpisodeBatchStandardBorel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","description":"EpisodeStep uses an explicit product-coordinate MeasurableSpace.comap, so the generic product instance is not visible to typeclass search. This module transports a Polish topology through the coordinate equivalence and then installs the corresponding Standard Borel instance. The deterministic and stochastic infinite batch trajectories are countable products of their finite batch spaces.","url":"../modules/banditrlproof-rl-finitehorizonepisodebatchstandardborel/index.html","parent":"chapter:finite-horizon-rl","order":609,"meta":[["Source","BanditRLProof/RL/FiniteHorizonEpisodeBatchStandardBorel.lean"],["Declarations","1"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonepisodebatchstandardborel banditrlproof.rl.finitehorizonepisodebatchstandardborel episodestep uses an explicit product-coordinate measurablespace.comap, so the generic product instance is not visible to typeclass search. this module transports a polish topology through the coordinate equivalence and then installs the corresponding standard borel instance. the deterministic and stochastic infinite batch trajectories are countable products of their finite batch spaces. lean module compiled","shard":"modules/2c82e1f72ed73243.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","label":"RL.FiniteHorizonEstimatedModelCertificate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","description":"This module connects a stage-indexed estimated reward/transition model to the compiled deterministic optimistic-certificate route. Separate two-sided reward and transition-expectation radii are required only on the recursively generated tail upper values. They produce a true Bellman certificate. The deterministic policy greedy for the estimated optimistic backup then has true Bellman residual at most twice its selec…","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html","parent":"chapter:finite-horizon-rl","order":610,"meta":[["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean"],["Declarations","30"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonestimatedmodelcertificate banditrlproof.rl.finitehorizonestimatedmodelcertificate this module connects a stage-indexed estimated reward/transition model to the compiled deterministic optimistic-certificate route. separate two-sided reward and transition-expectation radii are required only on the recursively generated tail upper values. they produce a true bellman certificate. the deterministic policy greedy for the estimated optimistic backup then has true bellman residual at most twice its selected reward-plus-transition radius, so the existing occupancy theorem gives a single-episode expected-regret bound. lean module compiled","shard":"modules/e0b34e8febac75ca.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","label":"RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","description":"This module solves the scalar square-root/logarithm obligations left by the explicit path-support calibration route. A closed-form lower bound on the number of episodes implies both the strict count margin and the finite-state, finite-horizon half contraction, then recovers the same adaptive confidence, optimism, and recommended-policy expected-regret endpoint.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html","parent":"chapter:finite-horizon-rl","order":611,"meta":[["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonexploratorypathsupportepisodethreshold banditrlproof.rl.finitehorizonexploratorypathsupportepisodethreshold this module solves the scalar square-root/logarithm obligations left by the explicit path-support calibration route. a closed-form lower bound on the number of episodes implies both the strict count margin and the finite-state, finite-horizon half contraction, then recovers the same adaptive confidence, optimism, and recommended-policy expected-regret endpoint. lean module compiled","shard":"modules/b2cf52ffdd58f3de.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","label":"RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","description":"This module removes the two remaining abstract calibration inputs from the path-support endpoint. A common state-action visit floor controls every expected-count denominator. If the resulting finite-state, finite-horizon transition coefficient is at most one half, the deterministic reward bound itself is a sufficient transition bonus.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html","parent":"chapter:finite-horizon-rl","order":612,"meta":[["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonexploratorypathsupportexplicitcalibration banditrlproof.rl.finitehorizonexploratorypathsupportexplicitcalibration this module removes the two remaining abstract calibration inputs from the path-support endpoint. a common state-action visit floor controls every expected-count denominator. if the resulting finite-state, finite-horizon transition coefficient is at most one half, the deterministic reward bound itself is a sufficient transition bonus. lean module compiled","shard":"modules/eb62366eb7df2f7f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","label":"RL.FiniteHorizonExploratoryPathSupportReachability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","description":"This module derives a policy-independent state-probability lower envelope from explicit initial singleton floors and one chosen predecessor transition for each successor-stage state. Uniform exploration supplies the chosen action's probability. The resulting envelope applies to every exploratory policy table and therefore to every behavior policy in the adaptive empirical optimistic source.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html","parent":"chapter:finite-horizon-rl","order":613,"meta":[["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonexploratorypathsupportreachability banditrlproof.rl.finitehorizonexploratorypathsupportreachability this module derives a policy-independent state-probability lower envelope from explicit initial singleton floors and one chosen predecessor transition for each successor-stage state. uniform exploration supplies the chosen action's probability. the resulting envelope applies to every exploratory policy table and therefore to every behavior policy in the adaptive empirical optimistic source. lean module compiled","shard":"modules/cb6c725009331d24.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","label":"RL.FiniteHorizonExploratoryReachabilityCalibration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","description":"This module turns a generated stage-state probability lower envelope into the state/action expected-count margins needed by the adaptive empirical optimistic confidence route. Uniform exploration supplies only the action factor; state reachability and the transition-bonus cover remain explicit contracts.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html","parent":"chapter:finite-horizon-rl","order":614,"meta":[["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean"],["Declarations","14"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonexploratoryreachabilitycalibration banditrlproof.rl.finitehorizonexploratoryreachabilitycalibration this module turns a generated stage-state probability lower envelope into the state/action expected-count margins needed by the adaptive empirical optimistic confidence route. uniform exploration supplies only the action factor; state reachability and the transition-bonus cover remain explicit contracts. lean module compiled","shard":"modules/305f505d9aa83d08.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","label":"RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","description":"This module turns the compiled generated reward and transition laws into an actual MDP.FiniteBatchModel.Confidence producer. The construction is noncircular: reward radius is zero, transition radius is one fixed external budget, transition coordinate radii use genuine expected-count lower margins, and the recursive value envelope is the explicit linear function remaining * (rewardBound + transitionBudget).","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html","parent":"chapter:finite-horizon-rl","order":615,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniidallcoordinatefinitebatchconfidence banditrlproof.rl.finitehorizoniidallcoordinatefinitebatchconfidence this module turns the compiled generated reward and transition laws into an actual mdp.finitebatchmodel.confidence producer. the construction is noncircular: reward radius is zero, transition radius is one fixed external budget, transition coordinate radii use genuine expected-count lower margins, and the recursive value envelope is the explicit linear function remaining * (rewardbound + transitionbudget). lean module compiled","shard":"modules/b60cbc42af79d2ee.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","label":"RL.FiniteHorizonIIDCountConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDCountConcentration","description":"This module applies the Mathlib-backed bounded-variable Hoeffding route to the fixed-policy iid episode-batch law. It proves two-sided confidence tails for one visit coordinate and one transition coordinate. Simultaneous finite coordinate events, random-denominator ratios, and adaptive episode policies remain downstream.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html","parent":"chapter:finite-horizon-rl","order":616,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean"],["Declarations","37"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniidcountconcentration banditrlproof.rl.finitehorizoniidcountconcentration this module applies the mathlib-backed bounded-variable hoeffding route to the fixed-policy iid episode-batch law. it proves two-sided confidence tails for one visit coordinate and one transition coordinate. simultaneous finite coordinate events, random-denominator ratios, and adaptive episode policies remain downstream. lean module compiled","shard":"modules/f6528463e67d2f76.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","label":"RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","description":"This module combines the compiled simultaneous visit/joint-count event, eligible positive denominators, and generated population-law factorization. For every eligible state-action-stage coordinate, the empirical singleton transition mass is within 2 * countRadius / visitCount of the true transition kernel singleton mass. The bundled endpoint reuses the existing global-delta event; it does not spend another failure b…","url":"../modules/banditrlproof-rl-finitehorizoniideligibleempiricaltransitionconfidence/index.html","parent":"chapter:finite-horizon-rl","order":617,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleEmpiricalTransitionConfidence.lean"],["Declarations","2"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniideligibleempiricaltransitionconfidence banditrlproof.rl.finitehorizoniideligibleempiricaltransitionconfidence this module combines the compiled simultaneous visit/joint-count event, eligible positive denominators, and generated population-law factorization. for every eligible state-action-stage coordinate, the empirical singleton transition mass is within 2 * countradius / visitcount of the true transition kernel singleton mass. the bundled endpoint reuses the existing global-delta event; it does not spend another failure budget or add reward confidence. lean module compiled","shard":"modules/cc8ccf653922594c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","label":"RL.FiniteHorizonIIDEligibleVisitCountPositivity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","description":"This module turns the compiled simultaneous visit-count deviation into a positive-denominator guarantee on an arbitrary finite set of eligible visit coordinates. Eligibility carries the necessary strict expected-count margin; unreachable coordinates are not assigned a false positivity conclusion.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html","parent":"chapter:finite-horizon-rl","order":618,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniideligiblevisitcountpositivity banditrlproof.rl.finitehorizoniideligiblevisitcountpositivity this module turns the compiled simultaneous visit-count deviation into a positive-denominator guarantee on an arbitrary finite set of eligible visit coordinates. eligibility carries the necessary strict expected-count margin; unreachable coordinates are not assigned a false positivity conclusion. lean module compiled","shard":"modules/d0660da0b0b6fa3b.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","label":"RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","description":"The finite-horizon MDP surface has a deterministic reward function. Consequently, records extracted from genuine trajectories have exact empirical rewards at every visited coordinate; no reward concentration or additional failure budget is needed. This module records that structural fact and combines it with the existing eligible empirical-transition confidence event.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html","parent":"chapter:finite-horizon-rl","order":619,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniidgeneratedempiricalrewardexactness banditrlproof.rl.finitehorizoniidgeneratedempiricalrewardexactness the finite-horizon mdp surface has a deterministic reward function. consequently, records extracted from genuine trajectories have exact empirical rewards at every visited coordinate; no reward concentration or additional failure budget is needed. this module records that structural fact and combines it with the existing eligible empirical-transition confidence event. lean module compiled","shard":"modules/8fb2f0caabd20af3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","label":"RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","description":"This module takes a finite Mathlib product of the compiled fixed-policy iid episode-batch law. Each product coordinate receives an equal confidence share. Outside the finite union of pulled-back count events, every batch-specific empirical model has a confidence witness, and the resulting one-episode expected-regret bounds sum over the finite product index.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html","parent":"chapter:finite-horizon-rl","order":620,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean"],["Declarations","16"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniidmultibatchcumulativeconfidenceregret banditrlproof.rl.finitehorizoniidmultibatchcumulativeconfidenceregret this module takes a finite mathlib product of the compiled fixed-policy iid episode-batch law. each product coordinate receives an equal confidence share. outside the finite union of pulled-back count events, every batch-specific empirical model has a confidence witness, and the resulting one-episode expected-regret bounds sum over the finite product index. lean module compiled","shard":"modules/7b98124166f88ccc.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","label":"RL.FiniteHorizonIIDSimultaneousCountConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","description":"This module puts every finite visit and joint-transition count coordinate into one index type and applies an equal-share finite union bound to the compiled fixed-coordinate tails. The route remains fixed-policy iid. It does not form visit-conditioned ratios or claim adaptive, anytime, or cumulative regret.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html","parent":"chapter:finite-horizon-rl","order":621,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean"],["Declarations","20"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniidsimultaneouscountconfidence banditrlproof.rl.finitehorizoniidsimultaneouscountconfidence this module puts every finite visit and joint-transition count coordinate into one index type and applies an equal-share finite union bound to the compiled fixed-coordinate tails. the route remains fixed-policy iid. it does not form visit-conditioned ratios or claim adaptive, anytime, or cumulative regret. lean module compiled","shard":"modules/0115ce4b57aba6bb.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","label":"RL.FiniteHorizonIIDTrajectoryBatch","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","description":"This module maps a finite product of genuine policy trajectory laws into the finite-batch empirical-model surface. It exposes both the episode/stage marginal law and independence across episode coordinates. The policy is fixed across episodes; adaptive cross-episode policy updates and concentration of visit-conditioned empirical ratios remain downstream.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html","parent":"chapter:finite-horizon-rl","order":622,"meta":[["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean"],["Declarations","27"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizoniidtrajectorybatch banditrlproof.rl.finitehorizoniidtrajectorybatch this module maps a finite product of genuine policy trajectory laws into the finite-batch empirical-model surface. it exposes both the episode/stage marginal law and independence across episode coordinates. the policy is fixed across episodes; adaptive cross-episode policy updates and concentration of visit-conditioned empirical ratios remain downstream. lean module compiled","shard":"modules/67ee00221d433aae.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonMDP","label":"RL.FiniteHorizonMDP","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonMDP","description":"This module fixes the dependency layer for the first finite-horizon RL leaf. A finite MDP uses Mathlib Markov kernels for transitions and a measurable Real reward. The only derived objects here are the one-step continuation value and Bellman action value. Policies, trajectories, value recursion, optimality, and regret remain downstream.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html","parent":"chapter:finite-horizon-rl","order":623,"meta":[["Source","BanditRLProof/RL/FiniteHorizonMDP.lean"],["Declarations","6"],["Project imports","0"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonmdp banditrlproof.rl.finitehorizonmdp this module fixes the dependency layer for the first finite-horizon rl leaf. a finite mdp uses mathlib markov kernels for transitions and a measurable real reward. the only derived objects here are the one-step continuation value and bellman action value. policies, trajectories, value recursion, optimality, and regret remain downstream. lean module compiled","shard":"modules/1b81b75c6f8471fa.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","label":"RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","description":"This module strengthens the fourth-power-prefix almost-sure theorem to every deterministic natural prefix for the same per-batch-normalized, equal-round-weighted process on the one heterogeneous causal trajectory measure.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html","parent":"chapter:finite-horizon-rl","order":624,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean"],["Declarations","21"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency banditrlproof.rl.finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency this module strengthens the fourth-power-prefix almost-sure theorem to every deterministic natural prefix for the same per-batch-normalized, equal-round-weighted process on the one heterogeneous causal trajectory measure. lean module compiled","shard":"modules/39bb2d69667e182e.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","label":"RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","description":"This module keeps the exact process which divides every successor batch by its own positive episode count and then weights rounds equally. Along the deterministic fourth-power prefixes (n + 1)^4, its compiled all-prefix L1 envelope is dominated by a sum of shifted exponent-three and exponent-two p-series. Markov's inequality and the first Borel-Cantelli lemma then give almost-everywhere convergence on this subsequen…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html","parent":"chapter:finite-horizon-rl","order":625,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean"],["Declarations","21"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureexplicitschedule banditrlproof.rl.finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureexplicitschedule this module keeps the exact process which divides every successor batch by its own positive episode count and then weights rounds equally. along the deterministic fourth-power prefixes (n + 1)^4, its compiled all-prefix l1 envelope is dominated by a sum of shifted exponent-three and exponent-two p-series. markov's inequality and the first borel-cantelli lemma then give almost-everywhere convergence on this subsequence. lean module compiled","shard":"modules/c727419ba9ecf20b.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","label":"RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","description":"This module controls the exact natural process which first normalizes every successor batch by its own positive episode count and then weights rounds equally. It is not the total-episode-mass-weighted process from the earlier L1 route. The argument exposes the cumulative normalized-return MGF, obtains a first-moment inverse-square-root bound, and combines it with the compiled logarithmic behavior expected-regret rat…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html","parent":"chapter:finite-horizon-rl","order":626,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean"],["Declarations","26"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency banditrlproof.rl.finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency this module controls the exact natural process which first normalizes every successor batch by its own positive episode count and then weights rounds equally. it is not the total-episode-mass-weighted process from the earlier l1 route. the argument exposes the cumulative normalized-return mgf, obtains a first-moment inverse-square-root bound, and combines it with the compiled logarithmic behavior expected-regret rate. lean module compiled","shard":"modules/b977209f4618234f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","label":"RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","description":"This module upgrades the natural-prefix pathwise process to a fixed-prefix high-probability statement on the genuine heterogeneous dependent causal source. The probability is inherited from the existing finite union of selected count-and-reward empirical-model events; no new independence, MGF, or concentration claim is introduced here.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html","parent":"chapter:finite-horizon-rl","order":627,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte banditrlproof.rl.finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte this module upgrades the natural-prefix pathwise process to a fixed-prefix high-probability statement on the genuine heterogeneous dependent causal source. the probability is inherited from the existing finite union of selected count-and-reward empirical-model events; no new independence, mgf, or concentration claim is introduced here. lean module compiled","shard":"modules/58bf0ff46185918e.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","label":"RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","description":"This module closes the symbolic finite-prefix envelope for the actual source.successorPolicyAt behavior on the one genuine heterogeneous dependent causal trajectory measure. The expanded coordinate rate has an inverse-square term, the scheduled exploration harmonic term, and a confidence term with exponent mdp.horizon + 5.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html","parent":"chapter:finite-horizon-rl","order":628,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean"],["Declarations","30"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalbehaviorexpectedregretlograte banditrlproof.rl.finitehorizonnaturalcausalbehaviorexpectedregretlograte this module closes the symbolic finite-prefix envelope for the actual source.successorpolicyat behavior on the one genuine heterogeneous dependent causal trajectory measure. the expanded coordinate rate has an inverse-square term, the scheduled exploration harmonic term, and a confidence term with exponent mdp.horizon + 5. lean module compiled","shard":"modules/2c1d7f021910a352.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","label":"RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","description":"This module turns the exact stopped second moment from the previous expected terminal into a finite deterministic budget. It uses finite-coordinate selection and MGF moment control, not optional stopping.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html","parent":"chapter:finite-horizon-rl","order":629,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean"],["Declarations","8"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministicmomentexpectedaveragerealizedbehaviorregret banditrlproof.rl.finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministicmomentexpectedaveragerealizedbehaviorregret this module turns the exact stopped second moment from the previous expected terminal into a finite deterministic budget. it uses finite-coordinate selection and mgf moment control, not optional stopping. lean module compiled","shard":"modules/ad54d36e4a57dcfe.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","label":"RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","description":"The fixed-confidence three-quarter event is integrated using an exact L2 second moment. The expectation proof is a pointwise event decomposition and a 2,2 Holder bound, not optional stopping.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html","parent":"chapter:finite-horizon-rl","order":630,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean"],["Declarations","10"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedaveragerealizedbehaviorregret banditrlproof.rl.finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedaveragerealizedbehaviorregret the fixed-confidence three-quarter event is integrated using an exact l2 second moment. the expectation proof is a pointwise event decomposition and a 2,2 holder bound, not optional stopping. lean module compiled","shard":"modules/0dccafeeffcf5dbe.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","label":"RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","description":"The existing self-consistent schedule assigns two shifted high-power model confidence shares to every coordinate. This module proves that their complete finite-prefix budget is at most 1/8, fixes the global return budget to 1/8, and turns the single-model bounded-stopping theorem into a joint bad event of probability at most 1/4. Its measurable complement therefore has real probability at least 3/4 and carries the s…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequarterg-2d26031403a4/index.html","parent":"chapter:finite-horizon-rl","order":631,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequartergoodeventaveragerealizedbehaviorregret banditrlproof.rl.finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequartergoodeventaveragerealizedbehaviorregret the existing self-consistent schedule assigns two shifted high-power model confidence shares to every coordinate. this module proves that their complete finite-prefix budget is at most 1/8, fixes the global return budget to 1/8, and turns the single-model bounded-stopping theorem into a joint bad event of probability at most 1/4. its measurable complement therefore has real probability at least 3/4 and carries the stopped logarithmic rate. lean module compiled","shard":"modules/058aae79cf868369.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","label":"RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","description":"The exact natural average realized behavior-regret process and its scheduled fixed-prefix logarithmic rate are evaluated at one positive bounded Mathlib stopping time. The stopped violation is measurable at the deterministic bound and is covered pathwise by the finite union of fixed-prefix violations.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html","parent":"chapter:finite-horizon-rl","order":632,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean"],["Declarations","14"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaveragerealizedbehaviorregret banditrlproof.rl.finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaveragerealizedbehaviorregret the exact natural average realized behavior-regret process and its scheduled fixed-prefix logarithmic rate are evaluated at one positive bounded mathlib stopping time. the stopped violation is measurable at the deterministic bound and is covered pathwise by the finite union of fixed-prefix violations. lean module compiled","shard":"modules/0f03556de4324f19.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","label":"RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","description":"The finite stopped-prefix route is sharpened by charging model-confidence failures only once at the deterministic horizon. Return deviations still use an equal-share finite union over the possible positive stopped prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html","parent":"chapter:finite-horizon-rl","order":633,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighprobabilityaveragerealizedbehaviorregret banditrlproof.rl.finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighprobabilityaveragerealizedbehaviorregret the finite stopped-prefix route is sharpened by charging model-confidence failures only once at the deterministic horizon. return deviations still use an equal-share finite union over the possible positive stopped prefixes. lean module compiled","shard":"modules/a6eaf19b50e052d0.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","label":"RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","description":"This module transports the compiled all-prefix L1 consistency theorem through stopping times constrained to a fixed deterministic window around each natural prefix. The selector is charged to a finite sum of shifted coordinate L1 envelopes. No optional-stopping identity or independence assumption is used.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html","parent":"chapter:finite-horizon-rl","order":634,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean"],["Declarations","12"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealizedbehaviorregretconsistency banditrlproof.rl.finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealizedbehaviorregretconsistency this module transports the compiled all-prefix l1 consistency theorem through stopping times constrained to a fixed deterministic window around each natural prefix. the selector is charged to a finite sum of shifted coordinate l1 envelopes. no optional-stopping identity or independence assumption is used. lean module compiled","shard":"modules/39e5c563a3541cad.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","label":"RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","description":"At schedule index n, this module scans the exact natural average realized behavior-regret process from the fourth-power prefix (n+1)^4 through the right endpoint (n+1)^4+(2*n+1). It stops at the first prefix where the process is at most a deterministic threshold, or at the right endpoint if no earlier prefix crosses the threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html","parent":"chapter:finite-horizon-rl","order":635,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingtimel1averagerealizedbehaviorregretconsistency banditrlproof.rl.finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingtimel1averagerealizedbehaviorregretconsistency at schedule index n, this module scans the exact natural average realized behavior-regret process from the fourth-power prefix (n+1)^4 through the right endpoint (n+1)^4+(2*n+1). it stops at the first prefix where the process is at most a deterministic threshold, or at the right endpoint if no earlier prefix crosses the threshold. lean module compiled","shard":"modules/d01efa4903b1d7de.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","label":"RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","description":"This module transports the compiled natural-causal L1 process through stopping families whose values lie in a growing finite window of the fourth-power prefix grid. Every finite candidate sum is dominated by the infinite tail of the compiled summable grid envelope, so the window width may grow arbitrarily.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html","parent":"chapter:finite-horizon-rl","order":636,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean"],["Declarations","18"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagerealizedbehaviorregretconsistency banditrlproof.rl.finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagerealizedbehaviorregretconsistency this module transports the compiled natural-causal l1 process through stopping families whose values lie in a growing finite window of the fourth-power prefix grid. every finite candidate sum is dominated by the infinite tail of the compiled summable grid envelope, so the window width may grow arbitrarily. lean module compiled","shard":"modules/5d2fc239ecf72c08.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","description":"This theorem route slows the moving first-passage threshold from 1/(n+1) to 1/sqrt(n+1). Dividing the compiled inverse-cubic plus inverse-square L1 envelope by that threshold gives shifted p-series with exponents 5/2 and 3/2. Their summability permits first Borel-Cantelli and an almost-sure eventual exact base-stop conclusion. No independence or optional-stopping argument is used.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html","parent":"chapter:finite-horizon-rl","order":637,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean"],["Declarations","22"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagesummabledelayandeventualimmediatestoppingl1consistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagesummabledelayandeventualimmediatestoppingl1consistency this theorem route slows the moving first-passage threshold from 1/(n+1) to 1/sqrt(n+1). dividing the compiled inverse-cubic plus inverse-square l1 envelope by that threshold gives shifted p-series with exponents 5/2 and 3/2. their summability permits first borel-cantelli and an almost-sure eventual exact base-stop conclusion. no independence or optional-stopping argument is used. lean module compiled","shard":"modules/141cb8fa34505117.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","description":"This module transports the accepted capped/uncapped L1 truncation equivalence through the Bochner integral. The result compares signed expected stopped average realized behavior regret on the same generated trajectory law. It is not an optional-stopping or policy-value identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html","parent":"chapter:finite-horizon-rl","order":638,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterexpectedregrettruncationreplacement banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterexpectedregrettruncationreplacement this module transports the accepted capped/uncapped l1 truncation equivalence through the bochner integral. the result compares signed expected stopped average realized behavior regret on the same generated trajectory law. it is not an optional-stopping or policy-value identity. lean module compiled","shard":"modules/8ced8c9611756ef1.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","description":"The capped double-linear scan and the genuine uncapped hittingAfter rule use the same deterministic base. Outside the existing capped delayed set, both rules stop exactly at that base and their stopped regret coordinates agree. Their difference also converges to zero in L1, in the named Lp Real 1 space, and in measure.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html","parent":"chapter:finite-horizon-rl","order":639,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterl1truncationequivalence banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterl1truncationequivalence the capped double-linear scan and the genuine uncapped hittingafter rule use the same deterministic base. outside the existing capped delayed set, both rules stop exactly at that base and their stopped regret coordinates agree. their difference also converges to zero in l1, in the named lp real 1 space, and in measure. lean module compiled","shard":"modules/eb8cdeca410f3171.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","description":"This module exposes the observed successor-batch sample means whose complement from the optimal initial value is the natural average realized-regret process. The empty prefix is assigned the optimal value, so the complement identity is total and remains valid at the WithTop.untopA fallback. The identity is then transported through the capped and genuine uncapped stopping prefixes and through Bochner integration. No…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html","parent":"chapter:finite-horizon-rl","order":640,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean"],["Declarations","22"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragesampledreturnexpectedoptimality banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragesampledreturnexpectedoptimality this module exposes the observed successor-batch sample means whose complement from the optimal initial value is the natural average realized-regret process. the empty prefix is assigned the optimal value, so the complement identity is total and remains valid at the withtop.untopa fallback. the identity is then transported through the capped and genuine uncapped stopping prefixes and through bochner integration. no expectation/stopping-index interchange or optional-stopping theorem is used. lean module compiled","shard":"modules/5983195049ff0445.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","description":"This module transports the accepted componentwise L1 truncation theorem through the Bochner integral. It compares capped and genuine uncapped successor-policy value gaps and return deviations on the same generated trajectory law, and preserves the exact expected realized-regret decomposition. It does not exchange expectation with a stopping index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html","parent":"chapter:finite-horizon-rl","order":641,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretandreturndeviationexpectedtruncationreplacement banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretandreturndeviationexpectedtruncationreplacement this module transports the accepted componentwise l1 truncation theorem through the bochner integral. it compares capped and genuine uncapped successor-policy value gaps and return deviations on the same generated trajectory law, and preserves the exact expected realized-regret decomposition. it does not exchange expectation with a stopping index. lean module compiled","shard":"modules/a25d67f6ee687193.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","description":"This module transports the genuine uncapped hittingAfter policy-value and return-deviation semantics to the capped double-linear first-passage prefix. It proves exponent-one norm replacement for both semantic components on the exact generated trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html","parent":"chapter:finite-horizon-rl","order":642,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean"],["Declarations","18"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretandreturndeviationl1truncationequivalence banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretandreturndeviationl1truncationequivalence this module transports the genuine uncapped hittingafter policy-value and return-deviation semantics to the capped double-linear first-passage prefix. it proves exponent-one norm replacement for both semantic components on the exact generated trajectory law. lean module compiled","shard":"modules/b9d5504243a1df3c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","description":"This module identifies the expected gap between stopped average realized behavior regret and its successor-policy value-gap coordinate with the negative expected return deviation. It proves that gap vanishes for both the capped first-passage approximation and the genuine uncapped hittingAfter prefix, then combines those vertical limits with the accepted horizontal truncation limits. It does not exchange expectation…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html","parent":"chapter:finite-horizon-rl","order":643,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedrealizedbehaviorregretandpolicyvalueexpectedconsistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedrealizedbehaviorregretandpolicyvalueexpectedconsistency this module identifies the expected gap between stopped average realized behavior regret and its successor-policy value-gap coordinate with the negative expected return deviation. it proves that gap vanishes for both the capped first-passage approximation and the genuine uncapped hittingafter prefix, then combines those vertical limits with the accepted horizontal truncation limits. it does not exchange expectation with a stopping index. lean module compiled","shard":"modules/f1d931ec59766ce7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","description":"This module extracts a deterministic tail start from the qualitative simultaneous stopped-return probability limit. At every later schedule index, the complement of the existing six-way violation event has real probability strictly greater than 1 - delta and is exactly the event on which all six literal errors are strictly below epsilon.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html","parent":"chapter:finite-horizon-rl","order":644,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturndeterministictailhighprobabilityoptimality banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturndeterministictailhighprobabilityoptimality this module extracts a deterministic tail start from the qualitative simultaneous stopped-return probability limit. at every later schedule index, the complement of the existing six-way violation event has real probability strictly greater than 1 - delta and is exactly the event on which all six literal errors are strictly below epsilon. lean module compiled","shard":"modules/c3ed91af080dc8bd.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","description":"This module upgrades the exact capped and uncapped stopped-return L1 route to literal convergence in measure and almost-sure convergence. The sampled return and the trajectory-law expected return of the actually selected successor policies converge to the optimal initial expected return, while their same-prefix difference converges to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html","parent":"chapter:finite-horizon-rl","order":645,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturninmeasurealmostsureoptimality banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturninmeasurealmostsureoptimality this module upgrades the exact capped and uncapped stopped-return l1 route to literal convergence in measure and almost-sure convergence. the sampled return and the trajectory-law expected return of the actually selected successor policies converge to the optimal initial expected return, while their same-prefix difference converges to zero. lean module compiled","shard":"modules/8b5632841b59a829.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","description":"This module upgrades the stopped expectation-level consistency theorem to actual L1 convergence. The sampled-return optimality error is exactly the negative realized behavior regret, the literal successor-policy return error is exactly the negative behavior expected regret, and their same-prefix gap is exactly the return deviation. The already compiled capped and genuine uncapped hittingAfter L1 results therefore tr…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html","parent":"chapter:finite-horizon-rl","order":646,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean"],["Declarations","20"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturnl1optimality banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturnl1optimality this module upgrades the stopped expectation-level consistency theorem to actual l1 convergence. the sampled-return optimality error is exactly the negative realized behavior regret, the literal successor-policy return error is exactly the negative behavior expected regret, and their same-prefix gap is exactly the return deviation. the already compiled capped and genuine uncapped hittingafter l1 results therefore transport to all three return processes. lean module compiled","shard":"modules/071a18b7f72511e3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","description":"This module combines the six literal capped/uncapped stopped-return convergence-in-measure endpoints into one measurable violation event at a common schedule index. Its probability tends to zero, so every positive accuracy and confidence budget is eventually satisfied simultaneously.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html","parent":"chapter:finite-horizon-rl","order":647,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturnsimultaneoushighprobabilityoptimality banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledandsuccessorpolicyreturnsimultaneoushighprobabilityoptimality this module combines the six literal capped/uncapped stopped-return convergence-in-measure endpoints into one measurable violation event at a common schedule index. its probability tends to zero, so every positive accuracy and confidence budget is eventually satisfied simultaneously. lean module compiled","shard":"modules/2427fcc33fef519d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","description":"This module gives a literal policy-value interpretation to the stopped sampled-return theorem. At coordinate t, the successor policy is the actual exploratory policy selected from the dependent prefix through t; its expected return is the integral of cumulative reward under its generated trajectory law. These literal policy returns are averaged over the same natural prefix as the sampled-return and regret processes,…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html","parent":"chapter:finite-horizon-rl","order":648,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean"],["Declarations","29"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledreturnandsuccessorpolicyexpectedreturnconsistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedsampledreturnandsuccessorpolicyexpectedreturnconsistency this module gives a literal policy-value interpretation to the stopped sampled-return theorem. at coordinate t, the successor policy is the actual exploratory policy selected from the dependent prefix through t; its expected return is the integral of cumulative reward under its generated trajectory law. these literal policy returns are averaged over the same natural prefix as the sampled-return and regret processes, with the optimal initial expected return at the empty prefix. lean module compiled","shard":"modules/a7011fb34dc315d6.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","description":"This route replaces the finite hittingBtwn scan by Mathlib's genuine uncapped hittingAfter. All-prefix almost-sure convergence proves that every fixed schedule-indexed hitting time is finite almost surely. The compiled summable-delay route separately proves that almost every trajectory eventually hits immediately at the fourth-power base. These facts yield a diverging stopped subsequence and hence stopped-process co…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html","parent":"chapter:finite-horizon-rl","order":649,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean"],["Declarations","12"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafteraefiniteeventualimmediatestoppingandinmeasureconsistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafteraefiniteeventualimmediatestoppingandinmeasureconsistency this route replaces the finite hittingbtwn scan by mathlib's genuine uncapped hittingafter. all-prefix almost-sure convergence proves that every fixed schedule-indexed hitting time is finite almost surely. the compiled summable-delay route separately proves that almost every trajectory eventually hits immediately at the fourth-power base. these facts yield a diverging stopped subsequence and hence stopped-process convergence almost everywhere and in measure. no expected-delay, uniform-integrability, l1, or optional- stopping claim is made. lean module compiled","shard":"modules/dccd43c981ca7bb6.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","description":"This module replaces the Young-inequality absolute-first-moment envelope by a Cauchy--Schwarz estimate on the stopping fibers. The actual stopping-round second moment still uses the accepted degree-eight polynomial envelope, while its square root yields a degree-four expected-absolute growth bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html","parent":"chapter:finite-horizon-rl","order":650,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean"],["Declarations","10"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzexpectedabsoluteasymptotics banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzexpectedabsoluteasymptotics this module replaces the young-inequality absolute-first-moment envelope by a cauchy--schwarz estimate on the stopping fibers. the actual stopping-round second moment still uses the accepted degree-eight polynomial envelope, while its square root yields a degree-four expected-absolute growth bound. lean module compiled","shard":"modules/5256cac1c538aadf.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","description":"This module combines the accepted uniform-integrability interface for the exact uncapped stopped-regret process with the compiled vanishing probability of the capped first-passage delayed event. The absolute and signed expected contributions on that concrete rare event both vanish.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-b0637306eb5b/index.html","parent":"chapter:finite-horizon-rl","order":651,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterdelayedeventexpectedcontribution banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterdelayedeventexpectedcontribution this module combines the accepted uniform-integrability interface for the exact uncapped stopped-regret process with the compiled vanishing probability of the capped first-passage delayed event. the absolute and signed expected contributions on that concrete rare event both vanish. lean module compiled","shard":"modules/ca4515a29f24abbf.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","description":"The exact stopped average realized behavior regret can be negative. This file isolates its one-sided excess, proves that its expectation is bounded by the hit threshold, and sends that threshold to zero. The argument uses the exact at-hit inequality and finite-measure integration, not optional stopping.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-d9d3109ce9f7/index.html","parent":"chapter:finite-horizon-rl","order":652,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterexpectedpositivepartconsistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterexpectedpositivepartconsistency the exact stopped average realized behavior regret can be negative. this file isolates its one-sided excess, proves that its expectation is bounded by the hit threshold, and sends that threshold to zero. the argument uses the exact at-hit inequality and finite-measure integration, not optional stopping. lean module compiled","shard":"modules/e656c5924aebfa74.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","description":"This module replaces the convergence-selected tail start in the fixed-index second-moment route by a concrete ceiling expression. It then transports that explicit witness into the deterministic stopped-regret absolute-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html","parent":"chapter:finite-horizon-rl","order":653,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterexplicittailstartexpectedabsolutebound banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterexplicittailstartexpectedabsolutebound this module replaces the convergence-selected tail start in the fixed-index second-moment route by a concrete ceiling expression. it then transports that explicit witness into the deterministic stopped-regret absolute-moment budget. lean module compiled","shard":"modules/077479f4659a43a5.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","description":"For each fixed threshold index, the genuine Mathlib hittingAfter has an L2 round count. Uniform deterministic-coordinate second moments therefore make the exact stopped average realized behavior regret integrable. Finite-hit membership then gives the expected threshold upper bound. This is a stopping- fiber argument, not optional stopping, and no lower expectation bound is claimed.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html","parent":"chapter:finite-horizon-rl","order":654,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean"],["Declarations","10"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterintegrableexpectedupperbound banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterintegrableexpectedupperbound for each fixed threshold index, the genuine mathlib hittingafter has an l2 round count. uniform deterministic-coordinate second moments therefore make the exact stopped average realized behavior regret integrable. finite-hit membership then gives the expected threshold upper bound. this is a stopping- fiber argument, not optional stopping, and no lower expectation bound is claimed. lean module compiled","shard":"modules/e5f6b467a20c8791.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","description":"This module upgrades each fixed-index genuine hittingAfter first passage from almost-sure finiteness to first-moment integrability. The proof uses the compiled fourth-power burn-in checkpoints. A cubic block-width envelope is summable against both the infinite model-tail budget and the exponentially small return share. Once the deterministic checkpoint regret rate lies below the fixed positive threshold, delayed che…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html","parent":"chapter:finite-horizon-rl","order":655,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean"],["Declarations","21"],["Project imports","4"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterintegrablefinitestoppingtime banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterintegrablefinitestoppingtime this module upgrades each fixed-index genuine hittingafter first passage from almost-sure finiteness to first-moment integrability. the proof uses the compiled fourth-power burn-in checkpoints. a cubic block-width envelope is summable against both the infinite model-tail budget and the exponentially small return share. once the deterministic checkpoint regret rate lies below the fixed positive threshold, delayed checkpoints are contained in the compiled violation events. lean module compiled","shard":"modules/f0bba0fa4b54c7a3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","description":"The positive part of the stopped average is already controlled by the hit threshold. For a delayed finite hit, first-hit minimality makes the previous average positive, so the negative overshoot can only come from the final successor-batch realized-regret coordinate divided by the hit index. A square-summable reciprocal weight and uniform coordinate L2 control then make that overshoot vanish. This is not optional st…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html","parent":"chapter:finite-horizon-rl","order":656,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean"],["Declarations","21"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterl1consistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterl1consistency the positive part of the stopped average is already controlled by the hit threshold. for a delayed finite hit, first-hit minimality makes the previous average positive, so the negative overshoot can only come from the final successor-batch realized-regret coordinate divided by the hit index. a square-summable reciprocal weight and uniform coordinate l2 control then make that overshoot vanish. this is not optional stopping. lean module compiled","shard":"modules/060c2dc5d2b73eb3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","description":"This module packages the expected-absolute convergence theorem for the exact uncapped hittingAfter stopped average realized behavior-regret process into Mathlib's MemLp 1, eLpNorm 1, and Lp Real 1 interfaces. It does not use optional stopping or add a uniform-integrability claim.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html","parent":"chapter:finite-horizon-rl","order":657,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterlpconsistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterlpconsistency this module packages the expected-absolute convergence theorem for the exact uncapped hittingafter stopped average realized behavior-regret process into mathlib's memlp 1, elpnorm 1, and lp real 1 interfaces. it does not use optional stopping or add a uniform-integrability claim. lean module compiled","shard":"modules/f6ab9e44e83da294.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","description":"This module transports the explicit polynomial moment envelope to the actual stopping-round second moment and stopped-regret expected absolute value. All model and source parameters are fixed while the threshold index varies.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html","parent":"chapter:finite-horizon-rl","order":658,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialsecondmomentexpectedabsoluteasymptotics banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialsecondmomentexpectedabsoluteasymptotics this module transports the explicit polynomial moment envelope to the actual stopping-round second moment and stopped-regret expected absolute value. all model and source parameters are fixed while the threshold index varies. lean module compiled","shard":"modules/0b1376067a4a4888.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","description":"This module bounds the explicit ceiling tail start by a model-dependent natural coefficient times the schedule scale. The resulting fourth-power checkpoint square is an explicit degree-eight polynomial in the fixed threshold index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html","parent":"chapter:finite-horizon-rl","order":659,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialsecondmomentexpectedabsolutebound banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialsecondmomentexpectedabsolutebound this module bounds the explicit ceiling tail start by a model-dependent natural coefficient times the schedule scale. the resulting fourth-power checkpoint square is an explicit degree-eight polynomial in the fixed threshold index. lean module compiled","shard":"modules/f2b75c0aecc1158c.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","description":"This module upgrades each fixed-index genuine hittingAfter first passage from a first moment to OFUL.SquareIntegrableFiniteStoppingTime when 4 < mdp.horizon. Squaring the fourth-power checkpoint values produces a seventh-degree block weight. The stronger horizon contract supplies an inverse-tenth local confidence share, leaving a summable inverse-square diagonal after shifted-tail reindexing.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html","parent":"chapter:finite-horizon-rl","order":660,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean"],["Declarations","27"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingaftersquareintegrablefinitestoppingtime banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingaftersquareintegrablefinitestoppingtime this module upgrades each fixed-index genuine hittingafter first passage from a first moment to oful.squareintegrablefinitestoppingtime when 4 < mdp.horizon. squaring the fourth-power checkpoint values produces a seventh-degree block weight. the stronger horizon contract supplies an inverse-tenth local confidence share, leaving a summable inverse-square diagonal after shifted-tail reindexing. lean module compiled","shard":"modules/5c53f54a74e9eb34.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","description":"This module evaluates the pathwise average successor-policy value gap and the normalized return deviation at the same genuine uncapped hittingAfter prefix as the accepted stopped realized-regret process. A deterministic 2H envelope and almost-everywhere random-prefix composition give behavior expected-regret L1 consistency. The exact realized/behavior/return decomposition then gives return-deviation L1 consistency.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html","parent":"chapter:finite-horizon-rl","order":661,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean"],["Declarations","23"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregretandreturndeviationl1consistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregretandreturndeviationl1consistency this module evaluates the pathwise average successor-policy value gap and the normalized return deviation at the same genuine uncapped hittingafter prefix as the accepted stopped realized-regret process. a deterministic 2h envelope and almost-everywhere random-prefix composition give behavior expected-regret l1 consistency. the exact realized/behavior/return decomposition then gives return-deviation l1 consistency. lean module compiled","shard":"modules/327c3e55ce790481.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","description":"This module consumes the accepted probability-theory uniform-integrability interface. Uniformly over the exact stopped-process schedule index, sufficiently small measurable trajectory events have small absolute stopped-regret integrals. This is an epsilon-delta consequence of uniform integrability, not an optional-stopping or quantitative tail theorem.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-f72757b422c4/index.html","parent":"chapter:finite-horizon-rl","order":662,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafteruniformabsolutecontinuity banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafteruniformabsolutecontinuity this module consumes the accepted probability-theory uniform-integrability interface. uniformly over the exact stopped-process schedule index, sufficiently small measurable trajectory events have small absolute stopped-regret integrals. this is an epsilon-delta consequence of uniform integrability, not an optional-stopping or quantitative tail theorem. lean module compiled","shard":"modules/96497046a61986e1.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","label":"RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","description":"This module upgrades the accepted exact hittingAfter L1/Lp convergence package to Mathlib probability-theory uniform integrability and signed expectation convergence. It does not use or prove an optional-stopping identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-22fe344b7372/index.html","parent":"chapter:finite-horizon-rl","order":663,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafteruniformintegrabilityexpectedconsistency banditrlproof.rl.finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafteruniformintegrabilityexpectedconsistency this module upgrades the accepted exact hittingafter l1/lp convergence package to mathlib probability-theory uniform integrability and signed expectation convergence. it does not use or prove an optional-stopping identity. lean module compiled","shard":"modules/34fff1c8285532b7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","label":"RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","description":"This module strengthens the fourth-power grid stopping theorem to a contiguous window of raw natural prefixes. At schedule index n, the stopping prefix may select any integer in [(n + 1)^4, (n + 1)^4 + n]. The all-prefix L1 envelope is bounded by one inverse square root, so every candidate costs at most D / (n + 1)^2; summing the n + 1 raw candidates gives D / (n + 1).","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html","parent":"chapter:finite-horizon-rl","order":664,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean"],["Declarations","22"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingtimel1averagerealizedbehaviorregretconsistency banditrlproof.rl.finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingtimel1averagerealizedbehaviorregretconsistency this module strengthens the fourth-power grid stopping theorem to a contiguous window of raw natural prefixes. at schedule index n, the stopping prefix may select any integer in [(n + 1)^4, (n + 1)^4 + n]. the all-prefix l1 envelope is bounded by one inverse square root, so every candidate costs at most d / (n + 1)^2; summing the n + 1 raw candidates gives d / (n + 1). lean module compiled","shard":"modules/3da26b2e8ecfe6ac.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","label":"RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","description":"This module consumes the compiled deterministic all-prefix almost-sure theorem for the exact per-batch-normalized, equal-round-weighted natural average realized behavior-regret process. A countable random-index measurability wrapper makes every random-prefix evaluation measurable, while pathwise composition transports the limit through any measurable random-prefix schedule which diverges almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html","parent":"chapter:finite-horizon-rl","order":665,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregretalmostsureconsistency banditrlproof.rl.finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregretalmostsureconsistency this module consumes the compiled deterministic all-prefix almost-sure theorem for the exact per-batch-normalized, equal-round-weighted natural average realized behavior-regret process. a countable random-index measurability wrapper makes every random-prefix evaluation measurable, while pathwise composition transports the limit through any measurable random-prefix schedule which diverges almost everywhere. lean module compiled","shard":"modules/f878fac0ec6429cf.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","label":"RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","description":"This module replaces the fixed fourth-power base and width n by deterministic functions baseRounds and windowWidth. The exact regularity contract is","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html","parent":"chapter:finite-horizon-rl","order":666,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1averagerealizedbehaviorregretconsistency banditrlproof.rl.finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1averagerealizedbehaviorregretconsistency this module replaces the fixed fourth-power base and width n by deterministic functions baserounds and windowwidth. the exact regularity contract is lean module compiled","shard":"modules/9f887de0e0743805.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","label":"RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","description":"This module repairs the nonvanishing finite-prefix model budget in the fixed-prefix logarithmic realized-regret route. Model failures before a deterministic burnin are paid for by the uniform 2 * horizon behavior regret bound. Only the infinite model tail from burnin onward enters the probability event. The return component remains the fixed-prefix sum of successor-batch-average deviations, with each batch divided b…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html","parent":"chapter:finite-horizon-rl","order":667,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte banditrlproof.rl.finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte this module repairs the nonvanishing finite-prefix model budget in the fixed-prefix logarithmic realized-regret route. model failures before a deterministic burnin are paid for by the uniform 2 * horizon behavior regret bound. only the infinite model tail from burnin onward enters the probability event. the return component remains the fixed-prefix sum of successor-batch-average deviations, with each batch divided by its own positive scheduled episode count. lean module compiled","shard":"modules/a63cc67926ea20ec.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","label":"RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","description":"This module chooses the concrete schedule scale n = n + 1, burnin n = scale n, rounds n = scale n ^ 4, and returnDelta n = exp (-scale n) for the compiled natural-causal burn-in terminal. The fourth-power prefix simultaneously absorbs the linear burn-in charge and the fixed-prefix normalized-return confidence radius.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html","parent":"chapter:finite-horizon-rl","order":668,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean"],["Declarations","25"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule banditrlproof.rl.finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule this module chooses the concrete schedule scale n = n + 1, burnin n = scale n, rounds n = scale n ^ 4, and returndelta n = exp (-scale n) for the compiled natural-causal burn-in terminal. the fourth-power prefix simultaneously absorbs the linear burn-in charge and the fixed-prefix normalized-return confidence radius. lean module compiled","shard":"modules/bfa07784d0928c6f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","label":"RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","description":"This module converts the compiled fixed-prefix logarithmic behavior-expected regret route into a realized route on the same heterogeneous dependent causal source. Natural round t uses the sample average of the actual successor batch at coordinate t + 1, generated by the exploratory policy selected from the prefix through t.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html","parent":"chapter:finite-horizon-rl","order":669,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean"],["Declarations","43"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte banditrlproof.rl.finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte this module converts the compiled fixed-prefix logarithmic behavior-expected regret route into a realized route on the same heterogeneous dependent causal source. natural round t uses the sample average of the actual successor batch at coordinate t + 1, generated by the exploratory policy selected from the prefix through t. lean module compiled","shard":"modules/4bec63ff105ca405.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","label":"RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","description":"This module upgrades the one-sided scheduled certificate to Mathlib TendstoInMeasure for the same equal-round-weighted natural realized behavior-regret process. Outside the compiled model-tail/return event, the parent route supplies the upper bound. The exact expected-minus-deviation identity, expected-regret nonnegativity, and the return-event complement supply the missing lower bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html","parent":"chapter:finite-horizon-rl","order":670,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule banditrlproof.rl.finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule this module upgrades the one-sided scheduled certificate to mathlib tendstoinmeasure for the same equal-round-weighted natural realized behavior-regret process. outside the compiled model-tail/return event, the parent route supplies the upper bound. the exact expected-minus-deviation identity, expected-regret nonnegativity, and the return-event complement supply the missing lower bound. lean module compiled","shard":"modules/5bf9e4aedf8353f7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","label":"RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","description":"This module consumes the explicit fourth-power prefix high-probability terminal. For every fixed positive threshold, the deterministic regret envelope is eventually below that threshold, so the threshold violation is contained in the compiled envelope violation and inherits its vanishing exact failure budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html","parent":"chapter:finite-horizon-rl","order":671,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability banditrlproof.rl.finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability this module consumes the explicit fourth-power prefix high-probability terminal. for every fixed positive threshold, the deterministic regret envelope is eventually below that threshold, so the threshold violation is contained in the compiled envelope violation and inherits its vanishing exact failure budget. lean module compiled","shard":"modules/f4bd7cceaf19f542.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","label":"RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","description":"At schedule index n, use threshold 1/(n+1) in the compiled capped first-passage scan from (n+1)^4 through (n+1)^4+(2*n+1).","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html","parent":"chapter:finite-horizon-rl","order":672,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean"],["Declarations","16"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagevanishingdelayprobabilityandl1consistency banditrlproof.rl.finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagevanishingdelayprobabilityandl1consistency at schedule index n, use threshold 1/(n+1) in the compiled capped first-passage scan from (n+1)^4 through (n+1)^4+(2*n+1). lean module compiled","shard":"modules/bceb71c80a576ce6.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","label":"RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","description":"This module equips the heterogeneous causal batch trajectory with its dependent Filtration.piLE natural filtration. The exact per-batch-normalized, equal-round-weighted average realized behavior-regret process is strongly adapted to that filtration: its value at prefix r uses only successor batches with coordinates at most r.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html","parent":"chapter:finite-horizon-rl","order":673,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregretalmostsureconsistency banditrlproof.rl.finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregretalmostsureconsistency this module equips the heterogeneous causal batch trajectory with its dependent filtration.pile natural filtration. the exact per-batch-normalized, equal-round-weighted average realized behavior-regret process is strongly adapted to that filtration: its value at prefix r uses only successor batches with coordinates at most r. lean module compiled","shard":"modules/c0efccc76ef87329.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","label":"RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","description":"At schedule index n, this module observes the compiled natural average realized behavior-regret process at the fourth-power prefix (n+1)^4. It stops at that prefix when the observation is at most a deterministic threshold and otherwise waits exactly 2*n+1 additional raw prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html","parent":"chapter:finite-horizon-rl","order":674,"meta":[["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindowstoppingtimel1averagerealizedbehaviorregretconsistency banditrlproof.rl.finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindowstoppingtimel1averagerealizedbehaviorregretconsistency at schedule index n, this module observes the compiled natural average realized behavior-regret process at the fourth-power prefix (n+1)^4. it stops at that prefix when the observation is at most a deterministic threshold and otherwise waits exactly 2*n+1 additional raw prefixes. lean module compiled","shard":"modules/51d238693033e960.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","label":"RL.FiniteHorizonOccupancyRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonOccupancyRegret","description":"This module connects finite-horizon Bellman optimality to expected trajectory regret. It constructs chronological state occupancies, records the expected one-step Bellman optimality gap under a policy, and recursively sums those gaps along the policy-induced state laws. The resulting occupancy functional is exactly the difference between optimal and policy value, hence exactly the expected trajectory regret. The mea…","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html","parent":"chapter:finite-horizon-rl","order":675,"meta":[["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonoccupancyregret banditrlproof.rl.finitehorizonoccupancyregret this module connects finite-horizon bellman optimality to expected trajectory regret. it constructs chronological state occupancies, records the expected one-step bellman optimality gap under a policy, and recursively sums those gaps along the policy-induced state laws. the resulting occupancy functional is exactly the difference between optimal and policy value, hence exactly the expected trajectory regret. the measurable greedy policy has zero regret. lean module compiled","shard":"modules/682a8686065935b9.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonOptimality","label":"RL.FiniteHorizonOptimality","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonOptimality","description":"This module adds finite-action maximization to the compiled policy-evaluation route. On a finite discrete state space, the pointwise finite argmax is a measurable deterministic selector. The resulting greedy Markov policy attains the backward optimal value, which dominates every Markov policy value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html","parent":"chapter:finite-horizon-rl","order":676,"meta":[["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean"],["Declarations","22"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonoptimality banditrlproof.rl.finitehorizonoptimality this module adds finite-action maximization to the compiled policy-evaluation route. on a finite discrete state space, the pointwise finite argmax is a measurable deterministic selector. the resulting greedy markov policy attains the backward optimal value, which dominates every markov policy value. lean module compiled","shard":"modules/08a6cdf0c7e8de8d.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","label":"RL.FiniteHorizonOptimisticCertificate","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonOptimisticCertificate","description":"This module isolates the deterministic dynamic-programming interface needed by optimistic finite-horizon RL. A certificate supplies upper values with zero terminal value and a one-step upper Bellman inequality. Backward induction turns that local contract into global optimism. For any Markov policy, the upper-value minus policy-value difference is then exactly a recursive sum of Bellman residuals under the true poli…","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html","parent":"chapter:finite-horizon-rl","order":677,"meta":[["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonoptimisticcertificate banditrlproof.rl.finitehorizonoptimisticcertificate this module isolates the deterministic dynamic-programming interface needed by optimistic finite-horizon rl. a certificate supplies upper values with zero terminal value and a one-step upper bellman inequality. backward induction turns that local contract into global optimism. for any markov policy, the upper-value minus policy-value difference is then exactly a recursive sum of bellman residuals under the true policy-induced state laws. a pointwise bonus bound on those residuals yields a finite-episode expected-regret bound. lean module compiled","shard":"modules/25f18bdd51c05545.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonPolicy","label":"RL.FiniteHorizonPolicy","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonPolicy","description":"This module defines a Markov action kernel at every valid decision stage of a finite-horizon MDP. It constructs the induced next-state kernel, the policy Bellman operator, and the backward finite-horizon policy value. The terminal and Bellman recursion theorems are policy-evaluation facts; no maximization, optimality, occupancy, or regret statement is made here.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html","parent":"chapter:finite-horizon-rl","order":678,"meta":[["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonpolicy banditrlproof.rl.finitehorizonpolicy this module defines a markov action kernel at every valid decision stage of a finite-horizon mdp. it constructs the induced next-state kernel, the policy bellman operator, and the backward finite-horizon policy value. the terminal and bellman recursion theorems are policy-evaluation facts; no maximization, optimality, occupancy, or regret statement is made here. lean module compiled","shard":"modules/273e54deb76050f3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","label":"RL.FiniteHorizonStageTransitionJointFactorization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","description":"This module proves that a fixed generated trajectory-stage joint transition mass factors into its state-action visit mass and the true MDP transition kernel singleton mass. The proof follows the recursive finite trajectory kernel; it does not divide by visit probabilities or introduce empirical confidence.","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html","parent":"chapter:finite-horizon-rl","order":679,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean"],["Declarations","18"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstagetransitionjointfactorization banditrlproof.rl.finitehorizonstagetransitionjointfactorization this module proves that a fixed generated trajectory-stage joint transition mass factors into its state-action visit mass and the true mdp transition kernel singleton mass. the proof follows the recursive finite trajectory kernel; it does not divide by visit probabilities or introduce empirical confidence. lean module compiled","shard":"modules/b8553f8c07d7c892.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","label":"RL.FiniteHorizonStageVisitFactorization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStageVisitFactorization","description":"This module factors a generated trajectory's stage/state/action visit mass into the corresponding stage-state mass and the policy action-kernel singleton mass. It is a population-law identity only: no reachability, action support, episode count, concentration, or regret premise is introduced.","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html","parent":"chapter:finite-horizon-rl","order":680,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstagevisitfactorization banditrlproof.rl.finitehorizonstagevisitfactorization this module factors a generated trajectory's stage/state/action visit mass into the corresponding stage-state mass and the policy action-kernel singleton mass. it is a population-law identity only: no reachability, action support, episode count, concentration, or regret premise is introduced. lean module compiled","shard":"modules/6bc2deb294183cbd.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","label":"RL.FiniteHorizonStochasticRewardBellman","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","description":"This module extends the finite-horizon MDP planning surface with a separate Real reward kernel. The compatibility contract says that every selected reward is integrable and has the deterministic mdp.reward field as its mean. The product with the transition kernel models conditional independence of the reward and next state given the current state-action pair.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html","parent":"chapter:finite-horizon-rl","order":681,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardbellman banditrlproof.rl.finitehorizonstochasticrewardbellman this module extends the finite-horizon mdp planning surface with a separate real reward kernel. the compatibility contract says that every selected reward is integrable and has the deterministic mdp.reward field as its mean. the product with the transition kernel models conditional independence of the reward and next state given the current state-action pair. lean module compiled","shard":"modules/30774244479759d1.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","label":"RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","description":"This module controls the policy and transition randomness in the mean Bellman return r(s, a) + V(s'). Sampled reward noise is deliberately left to the separate cumulative reward-deviation theorem.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html","parent":"chapter:finite-horizon-rl","order":682,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardbellmaninnovationconcentration banditrlproof.rl.finitehorizonstochasticrewardbellmaninnovationconcentration this module controls the policy and transition randomness in the mean bellman return r(s, a) + v(s'). sampled reward noise is deliberately left to the separate cumulative reward-deviation theorem. lean module compiled","shard":"modules/7e9a1796f28c42cb.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","label":"RL.FiniteHorizonStochasticRewardConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","description":"This module transports a selected reward-kernel sub-Gaussian law through the compiled generated-head conditional distribution. It exposes the centered head reward as a conditional and unconditional sub-Gaussian random variable, then specializes the existing finite-sum concentration route to a one-step two-sided tail.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html","parent":"chapter:finite-horizon-rl","order":683,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardconcentration banditrlproof.rl.finitehorizonstochasticrewardconcentration this module transports a selected reward-kernel sub-gaussian law through the compiled generated-head conditional distribution. it exposes the centered head reward as a conditional and unconditional sub-gaussian random variable, then specializes the existing finite-sum concentration route to a one-step two-sided tail. lean module compiled","shard":"modules/b5c0fe62bbf0bf28.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","label":"RL.FiniteHorizonStochasticRewardConditionalLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","description":"This module identifies the selected reward kernel as the conditional law of the first sampled reward given the first sampled action. The result is first proved for the one-step action/reward marginal, then transported to every positive generated stochastic trajectory and to the corresponding trimmed condExpKernel.map surface.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html","parent":"chapter:finite-horizon-rl","order":684,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardconditionallaw banditrlproof.rl.finitehorizonstochasticrewardconditionallaw this module identifies the selected reward kernel as the conditional law of the first sampled reward given the first sampled action. the result is first proved for the one-step action/reward marginal, then transported to every positive generated stochastic trajectory and to the corresponding trimmed condexpkernel.map surface. lean module compiled","shard":"modules/2520e3477eb2999f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","label":"RL.FiniteHorizonStochasticRewardCumulativeConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","description":"This module centers every sampled reward at the mean of its actual pre-step state and sampled action. A recursive kernel composition proof combines the common one-step sub-Gaussian proxy additively, yielding a total proxy linear in the remaining horizon and a fixed-horizon two-sided delta tail.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html","parent":"chapter:finite-horizon-rl","order":685,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardcumulativeconcentration banditrlproof.rl.finitehorizonstochasticrewardcumulativeconcentration this module centers every sampled reward at the mean of its actual pre-step state and sampled action. a recursive kernel composition proof combines the common one-step sub-gaussian proxy additively, yielding a total proxy linear in the remaining horizon and a fixed-horizon two-sided delta tail. lean module compiled","shard":"modules/da549a6056d78ebd.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","label":"RL.FiniteHorizonStochasticRewardErasureLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","description":"This module discards sampled Real rewards from generated stochastic-reward trajectories while retaining every action and next state. The resulting law is exactly the ordinary finite-horizon policy trajectory law. The equality is then lifted through the initial-state mixture, a finite iid episode family, and the existing known-reward EpisodeBatch conversion.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html","parent":"chapter:finite-horizon-rl","order":686,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean"],["Declarations","14"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewarderasurelaw banditrlproof.rl.finitehorizonstochasticrewarderasurelaw this module discards sampled real rewards from generated stochastic-reward trajectories while retaining every action and next state. the resulting law is exactly the ordinary finite-horizon policy trajectory law. the equality is then lifted through the initial-state mixture, a finite iid episode family, and the existing known-reward episodebatch conversion. lean module compiled","shard":"modules/92f38ad159165af3.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","label":"RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","description":"This module combines the compiled sampled-reward coordinate tail with the existing simultaneous count/transition event. It keeps the complete reward-bearing iid trajectory family as the probability space and constructs the empirical model from the actual sampled rewards.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html","parent":"chapter:finite-horizon-rl","order":687,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean"],["Declarations","32"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence banditrlproof.rl.finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence this module combines the compiled sampled-reward coordinate tail with the existing simultaneous count/transition event. it keeps the complete reward-bearing iid trajectory family as the probability space and constructs the empirical model from the actual sampled rewards. lean module compiled","shard":"modules/e4faea21c4e46cf7.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","label":"RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","description":"This module retains the sampled reward in every generated episode record. It proves that a fixed stage/state/action reward sum, centered by its stored MDP mean and masked by the visit event, is sub-Gaussian across iid complete trajectories. The total proxy is the conservative episode-linear proxy episodes * varianceProxy; no bounded sampled-reward or exact-count claim is used.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html","parent":"chapter:finite-horizon-rl","order":688,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean"],["Declarations","22"],["Project imports","3"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardiidempiricalrewardconfidence banditrlproof.rl.finitehorizonstochasticrewardiidempiricalrewardconfidence this module retains the sampled reward in every generated episode record. it proves that a fixed stage/state/action reward sum, centered by its stored mdp mean and masked by the visit event, is sub-gaussian across iid complete trajectories. the total proxy is the conservative episode-linear proxy episodes * varianceproxy; no bounded sampled-reward or exact-count claim is used. lean module compiled","shard":"modules/4e005ff9f4704783.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","label":"RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","description":"This module replaces the coordinatewise margin and cover inputs of the stochastic-reward iid empirical-model terminal by one common expected-count floor and one scalar half-contraction condition. The reward and transition budgets are explicit functions of that floor.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html","parent":"chapter:finite-horizon-rl","order":689,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardiidexplicitcalibration banditrlproof.rl.finitehorizonstochasticrewardiidexplicitcalibration this module replaces the coordinatewise margin and cover inputs of the stochastic-reward iid empirical-model terminal by one common expected-count floor and one scalar half-contraction condition. the reward and transition budgets are explicit functions of that floor. lean module compiled","shard":"modules/86e110363a4d0d40.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","label":"RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","description":"The half-contraction calibration uses the coarse fixed point transitionBudget = rewardBound + 2 * rewardBudget. Here the actual contraction factor q < 1 is retained and the fixed point is solved exactly, so the transition budget shrinks with q.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html","parent":"chapter:finite-horizon-rl","order":690,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardiidselfconsistentcalibration banditrlproof.rl.finitehorizonstochasticrewardiidselfconsistentcalibration the half-contraction calibration uses the coarse fixed point transitionbudget = rewardbound + 2 * rewardbudget. here the actual contraction factor q < 1 is retained and the fixed point is solved exactly, so the transition budget shrinks with q. lean module compiled","shard":"modules/a2dd38ed46a3af40.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","label":"RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","description":"This module takes finite products of complete reward-bearing trajectories. Independence is only asserted across episodes. Each episode deviation remains centered by the policy value at that trajectory's own sampled initial state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html","parent":"chapter:finite-horizon-rl","order":691,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardiidtotalreturnconcentration banditrlproof.rl.finitehorizonstochasticrewardiidtotalreturnconcentration this module takes finite products of complete reward-bearing trajectories. independence is only asserted across episodes. each episode deviation remains centered by the policy value at that trajectory's own sampled initial state. lean module compiled","shard":"modules/4e222b16e8fb2ba2.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","label":"RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","description":"This module lifts the statewise sampled-return concentration theorem to the full stochastic trajectory measure generated from a finite initial-state law. The random variable remains centered by the recursive value of its own initial state; no concentration claim is made for mixing those state-dependent values.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardinitiallawtotalreturnconcentration/index.html","parent":"chapter:finite-horizon-rl","order":692,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardinitiallawtotalreturnconcentration banditrlproof.rl.finitehorizonstochasticrewardinitiallawtotalreturnconcentration this module lifts the statewise sampled-return concentration theorem to the full stochastic trajectory measure generated from a finite initial-state law. the random variable remains centered by the recursive value of its own initial state; no concentration claim is made for mixing those state-dependent values. lean module compiled","shard":"modules/340d176746b0eace.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","label":"RL.FiniteHorizonStochasticRewardMarginal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","description":"This module identifies the first generated stochastic-reward coordinate at every positive recursive horizon. It exposes the exact action/reward joint law and reward-only policy mixture needed before adding conditional-law or concentration assumptions.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html","parent":"chapter:finite-horizon-rl","order":693,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardmarginal banditrlproof.rl.finitehorizonstochasticrewardmarginal this module identifies the first generated stochastic-reward coordinate at every positive recursive horizon. it exposes the exact action/reward joint law and reward-only policy mixture needed before adding conditional-law or concentration assumptions. lean module compiled","shard":"modules/6ecdd848334c317a.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","label":"RL.FiniteHorizonStochasticRewardTotalReturnConcentration","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","description":"This module combines selected-reward noise with policy-action and transition noise. The target random variable is the actual sampled cumulative return minus the recursive policy value. The additive proxy relies on the product structure of reward and next-state sampling given the current state and action.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html","parent":"chapter:finite-horizon-rl","order":694,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean"],["Declarations","28"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardtotalreturnconcentration banditrlproof.rl.finitehorizonstochasticrewardtotalreturnconcentration this module combines selected-reward noise with policy-action and transition noise. the target random variable is the actual sampled cumulative return minus the recursive policy value. the additive proxy relies on the product structure of reward and next-state sampling given the current state and action. lean module compiled","shard":"modules/c1517f0e64986a8f.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","label":"RL.FiniteHorizonStochasticRewardTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","description":"This module generates finite policy trajectories whose coordinates retain the sampled action, sampled Real reward, and next state. Selected rewards need only the L1 mean-compatibility contract from the stochastic Bellman layer. The cumulative sampled reward is proved integrable recursively and its expectation is identified with both stochastic and mean policy evaluation.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html","parent":"chapter:finite-horizon-rl","order":695,"meta":[["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizonstochasticrewardtrajectory banditrlproof.rl.finitehorizonstochasticrewardtrajectory this module generates finite policy trajectories whose coordinates retain the sampled action, sampled real reward, and next state. selected rewards need only the l1 mean-compatibility contract from the stochastic bellman layer. the cumulative sampled reward is proved integrable recursively and its expectation is identified with both stochastic and mean policy evaluation. lean module compiled","shard":"modules/876b9bd0cd8904e2.json"},{"id":"module:BanditRLProof.RL.FiniteHorizonTrajectory","label":"RL.FiniteHorizonTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.FiniteHorizonTrajectory","description":"This module constructs the genuine finite trajectory generated by a Markov policy. A trace with n decisions records the sampled action and resulting next state at each coordinate. The recursive trajectory kernel is then used to identify expected cumulative reward with the backward policy value.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html","parent":"chapter:finite-horizon-rl","order":696,"meta":[["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean"],["Declarations","15"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.finitehorizontrajectory banditrlproof.rl.finitehorizontrajectory this module constructs the genuine finite trajectory generated by a markov policy. a trace with n decisions records the sampled action and resulting next state at each coordinate. the recursive trajectory kernel is then used to identify expected cumulative reward with the backward policy value. lean module compiled","shard":"modules/1a1bc34fd7cc0065.json"},{"id":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","label":"RL.StoppedReturnJointErrorDeterministicTailHighProbability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","description":"This module replaces the six-coordinate conjunction-only confidence surface by one measurable scalar random variable: the maximum of the capped/uncapped sampled-return, actual successor-policy-return, and same-prefix-gap errors. The accepted deterministic-tail certificate then transports to direct strict sublevel and weak superlevel events for this joint error.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html","parent":"chapter:finite-horizon-rl","order":697,"meta":[["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Finite-horizon RL"]],"statement":"","missing":[],"search":"rl.stoppedreturnjointerrordeterministictailhighprobability banditrlproof.rl.stoppedreturnjointerrordeterministictailhighprobability this module replaces the six-coordinate conjunction-only confidence surface by one measurable scalar random variable: the maximum of the capped/uncapped sampled-return, actual successor-policy-return, and same-prefix-gap errors. the accepted deterministic-tail certificate then transports to direct strict sublevel and weak superlevel events for this joint error. lean module compiled","shard":"modules/34c63cc214810d21.json"},{"id":"module:BanditRLProof.RatMeasurability","label":"RatMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RatMeasurability","description":"This module contains narrow measurability wrappers for the project's Rat-valued bandit quantities. It deliberately does not choose probability, filtration, or concentration assumptions.","url":"../modules/banditrlproof-ratmeasurability/index.html","parent":"chapter:foundations","order":698,"meta":[["Source","BanditRLProof/RatMeasurability.lean"],["Declarations","1"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"ratmeasurability banditrlproof.ratmeasurability this module contains narrow measurability wrappers for the project's rat-valued bandit quantities. it deliberately does not choose probability, filtration, or concentration assumptions. lean module compiled","shard":"modules/ec5dd601bec2a62a.json"},{"id":"module:BanditRLProof.RealKernelRegretPullCount","label":"RealKernelRegretPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RealKernelRegretPullCount","description":"This module specializes the Real mean-regret bookkeeping surface to an arm-indexed Mathlib kernel. The mean of arm a is the Bochner integral of the identity under the measure nu a, exactly as in the LML finite-bandit regret definitions. Algorithm laws and concentration assumptions remain downstream.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html","parent":"chapter:foundations","order":699,"meta":[["Source","BanditRLProof/RealKernelRegretPullCount.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"realkernelregretpullcount banditrlproof.realkernelregretpullcount this module specializes the real mean-regret bookkeeping surface to an arm-indexed mathlib kernel. the mean of arm a is the bochner integral of the identity under the measure nu a, exactly as in the lml finite-bandit regret definitions. algorithm laws and concentration assumptions remain downstream. lean module compiled","shard":"modules/bd78f0f47c67d1d3.json"},{"id":"module:BanditRLProof.RealMeanRegretPullCount","label":"RealMeanRegretPullCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RealMeanRegretPullCount","description":"This module provides the Real-valued finite-arm bookkeeping surface needed by the exact LML ETC route. It is parameterized by an arm-mean function, so a later kernel bridge can instantiate mean a with the integral of the identity under arm a without changing the deterministic or Bochner proofs here.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html","parent":"chapter:foundations","order":700,"meta":[["Source","BanditRLProof/RealMeanRegretPullCount.lean"],["Declarations","6"],["Project imports","3"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"realmeanregretpullcount banditrlproof.realmeanregretpullcount this module provides the real-valued finite-arm bookkeeping surface needed by the exact lml etc route. it is parameterized by an arm-mean function, so a later kernel bridge can instantiate mean a with the integral of the identity under arm a without changing the deterministic or bochner proofs here. lean module compiled","shard":"modules/a8b2b232507c2380.json"},{"id":"module:BanditRLProof.Regret","label":"Regret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.Regret","description":"The first layer records pseudo-regret as an executable recursive object over finite arms and rational means. Later Mathlib-heavy files can connect this to expectations, conditional distributions, martingales, or concentration.","url":"../modules/banditrlproof-regret/index.html","parent":"chapter:foundations","order":701,"meta":[["Source","BanditRLProof/Regret.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"regret banditrlproof.regret the first layer records pseudo-regret as an executable recursive object over finite arms and rational means. later mathlib-heavy files can connect this to expectations, conditional distributions, martingales, or concentration. lean module compiled","shard":"modules/094980085c9c35a8.json"},{"id":"module:BanditRLProof.RegretCountBounds","label":"RegretCountBounds","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RegretCountBounds","description":"This module contains algorithm-neutral deterministic scaffolds that convert per-arm pull-count upper bounds into pseudo-regret upper bounds. It stays below probability, expectation, filtrations, concentration, and algorithm final theorem work.","url":"../modules/banditrlproof-regretcountbounds/index.html","parent":"chapter:foundations","order":702,"meta":[["Source","BanditRLProof/RegretCountBounds.lean"],["Declarations","3"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"regretcountbounds banditrlproof.regretcountbounds this module contains algorithm-neutral deterministic scaffolds that convert per-arm pull-count upper bounds into pseudo-regret upper bounds. it stays below probability, expectation, filtrations, concentration, and algorithm final theorem work. lean module compiled","shard":"modules/a21ec45143e9825b.json"},{"id":"module:BanditRLProof.RegretDecomposition","label":"RegretDecomposition","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RegretDecomposition","description":"This module consumes the Mathlib-backed finite bookkeeping wrappers. It should stay deterministic: probability, measurability, and concentration imports belong in later layers.","url":"../modules/banditrlproof-regretdecomposition/index.html","parent":"chapter:foundations","order":703,"meta":[["Source","BanditRLProof/RegretDecomposition.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"regretdecomposition banditrlproof.regretdecomposition this module consumes the mathlib-backed finite bookkeeping wrappers. it should stay deterministic: probability, measurability, and concentration imports belong in later layers. lean module compiled","shard":"modules/b064adda19239008.json"},{"id":"module:BanditRLProof.RewardKernel","label":"RewardKernel","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RewardKernel","description":"This module records a narrow KERNEL-REWARD leaf: a reward law is represented as a Mathlib Markov kernel indexed by an arm/context object. It exposes the kernel and the measurability/probability regularity facts needed by later policy-kernel and trajectory-law leaves, but it does not bind kernels or build a trajectory measure.","url":"../modules/banditrlproof-rewardkernel/index.html","parent":"chapter:probability","order":704,"meta":[["Source","BanditRLProof/RewardKernel.lean"],["Declarations","74"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"rewardkernel banditrlproof.rewardkernel this module records a narrow kernel-reward leaf: a reward law is represented as a mathlib markov kernel indexed by an arm/context object. it exposes the kernel and the measurability/probability regularity facts needed by later policy-kernel and trajectory-law leaves, but it does not bind kernels or build a trajectory measure. lean module compiled","shard":"modules/6baeb7b115f3b6c1.json"},{"id":"module:BanditRLProof.RewardTraceLaw","label":"RewardTraceLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.RewardTraceLaw","description":"This module contains the process-law foundation shared by adaptive bandit routes. Initial reward marginals and successor regular conditional distributions determine every finite prefix and hence the complete reward-trace law. The results are independent of any ETC/UCB/Thompson algorithm layer.","url":"../modules/banditrlproof-rewardtracelaw/index.html","parent":"chapter:probability","order":705,"meta":[["Source","BanditRLProof/RewardTraceLaw.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"rewardtracelaw banditrlproof.rewardtracelaw this module contains the process-law foundation shared by adaptive bandit routes. initial reward marginals and successor regular conditional distributions determine every finite prefix and hence the complete reward-trace law. the results are independent of any etc/ucb/thompson algorithm layer. lean module compiled","shard":"modules/e81e7eb551d58a10.json"},{"id":"module:BanditRLProof.ScalarENNReal","label":"ScalarENNReal","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ScalarENNReal","description":"This module contains scalar conversion leaves used before any probability or bandit-specific expectation statement. The lemmas here are deliberately independent of FiniteBanditModel, traces, integrals, and filtrations.","url":"../modules/banditrlproof-scalarennreal/index.html","parent":"chapter:foundations","order":706,"meta":[["Source","BanditRLProof/ScalarENNReal.lean"],["Declarations","1"],["Project imports","0"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"scalarennreal banditrlproof.scalarennreal this module contains scalar conversion leaves used before any probability or bandit-specific expectation statement. the lemmas here are deliberately independent of finitebanditmodel, traces, integrals, and filtrations. lean module compiled","shard":"modules/ffe67e47b1342d9b.json"},{"id":"module:BanditRLProof.ScalarPseudoRegret","label":"ScalarPseudoRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.ScalarPseudoRegret","description":"This module is still pointwise scalar algebra. It connects the deterministic pull-count regret decomposition to the scalar ENNReal.ofReal finite-sum faithfulness lemma under an explicit nonnegativity contract on model gaps.","url":"../modules/banditrlproof-scalarpseudoregret/index.html","parent":"chapter:foundations","order":707,"meta":[["Source","BanditRLProof/ScalarPseudoRegret.lean"],["Declarations","2"],["Project imports","2"],["Teaching chapter","Foundations"]],"statement":"","missing":[],"search":"scalarpseudoregret banditrlproof.scalarpseudoregret this module is still pointwise scalar algebra. it connects the deterministic pull-count regret decomposition to the scalar ennreal.ofreal finite-sum faithfulness lemma under an explicit nonnegativity contract on model gaps. lean module compiled","shard":"modules/04f2599e28c848fe.json"},{"id":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","label":"TsallisConjugatePotentialFiniteHorizon","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisConjugatePotentialFiniteHorizon","description":"This module transports the compiled ordinary-importance-weighted conjugate-potential bound through identified finite conditional action laws and sums it over a finite horizon. It also records the exact deterministic potential telescope. The potential includes the paper-normalizing 1 / eta, so the process is ready for later cross-learning-rate comparisons.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html","parent":"chapter:tsallis","order":708,"meta":[["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean"],["Declarations","21"],["Project imports","4"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisconjugatepotentialfinitehorizon banditrlproof.tsallisconjugatepotentialfinitehorizon this module transports the compiled ordinary-importance-weighted conjugate-potential bound through identified finite conditional action laws and sums it over a finite horizon. it also records the exact deterministic potential telescope. the potential includes the paper-normalizing 1 / eta, so the process is ready for later cross-learning-rate comparisons. lean module compiled","shard":"modules/95b2d9c1b7c75723.json"},{"id":"module:BanditRLProof.TsallisConjugatePotentialStability","label":"TsallisConjugatePotentialStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisConjugatePotentialStability","description":"This module formalizes the deterministic potential quantity used in the paper-faithful refined Tsallis-INF stability route. The local learning-rate normalization is half the paper normalization, so the translated Lemma 19 coefficients are eta for the quadratic term and 2 * eta^2 for the positive cubic remainder.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html","parent":"chapter:tsallis","order":709,"meta":[["Source","BanditRLProof/TsallisConjugatePotentialStability.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisconjugatepotentialstability banditrlproof.tsallisconjugatepotentialstability this module formalizes the deterministic potential quantity used in the paper-faithful refined tsallis-inf stability route. the local learning-rate normalization is half the paper normalization, so the translated lemma 19 coefficients are eta for the quadratic term and 2 * eta^2 for the positive cubic remainder. lean module compiled","shard":"modules/36d26052d5725552.json"},{"id":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","label":"TsallisConstrainedQuadraticOptimization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisConstrainedQuadraticOptimization","description":"This module gives a rigorous finite-sum version of the one-round optimization used after lambda interpolation. In contrast to the informal paper lemma, the quadratic coefficients are required to be strictly positive and the simplex-derived square-root mass constraint is explicit.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html","parent":"chapter:tsallis","order":710,"meta":[["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisconstrainedquadraticoptimization banditrlproof.tsallisconstrainedquadraticoptimization this module gives a rigorous finite-sum version of the one-round optimization used after lambda interpolation. in contrast to the informal paper lemma, the quadratic coefficients are required to be strictly positive and the simplex-derived square-root mass constraint is explicit. lean module compiled","shard":"modules/0092d0a097e173b2.json"},{"id":"module:BanditRLProof.TsallisFTRLConditionalStability","label":"TsallisFTRLConditionalStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLConditionalStability","description":"This module transports the compiled finite sampling-law half-Tsallis stability bound through an identified conditional action law. It also exposes a canonical generated-action consumer for the fixed half-Tsallis minimizer.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html","parent":"chapter:tsallis","order":711,"meta":[["Source","BanditRLProof/TsallisFTRLConditionalStability.lean"],["Declarations","8"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlconditionalstability banditrlproof.tsallisftrlconditionalstability this module transports the compiled finite sampling-law half-tsallis stability bound through an identified conditional action law. it also exposes a canonical generated-action consumer for the fixed half-tsallis minimizer. lean module compiled","shard":"modules/6f0056731c8550fd.json"},{"id":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","label":"TsallisFTRLEstimatedEnvironmentRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","description":"This module joins the deterministic half-Tsallis FTRL decomposition to the generated predictable trajectory. Time zero is kept separate: the recursive history score at level n already contains observations through time n, so the generated successor-stability theorem covers times 1, ..., horizon.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html","parent":"chapter:tsallis","order":712,"meta":[["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean"],["Declarations","39"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlestimatedenvironmentregret banditrlproof.tsallisftrlestimatedenvironmentregret this module joins the deterministic half-tsallis ftrl decomposition to the generated predictable trajectory. time zero is kept separate: the recursive history score at level n already contains observations through time n, so the generated successor-stability theorem covers times 1, ..., horizon. lean module compiled","shard":"modules/4184d01947728c69.json"},{"id":"module:BanditRLProof.TsallisFTRLExpectedStability","label":"TsallisFTRLExpectedStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLExpectedStability","description":"This module sums the one-round conditional half-Tsallis stability theorem over a finite horizon under a common ambient trajectory measure. Conditional-law identification is used both for the one-round inequality and for transporting product-law integrability back to each realized history/action score.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html","parent":"chapter:tsallis","order":713,"meta":[["Source","BanditRLProof/TsallisFTRLExpectedStability.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlexpectedstability banditrlproof.tsallisftrlexpectedstability this module sums the one-round conditional half-tsallis stability theorem over a finite horizon under a common ambient trajectory measure. conditional-law identification is used both for the one-round inequality and for transporting product-law integrability back to each realized history/action score. lean module compiled","shard":"modules/80ea1cc48bfe0748.json"},{"id":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","label":"TsallisFTRLFiniteHorizonSelection","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLFiniteHorizonSelection","description":"This module applies the fixed half-Tsallis minimizer choice to cumulative loss vectors. It removes the caller-supplied minimizer certificates from the deterministic finite-horizon FTRL decomposition and records the successor indexing needed by importance-weighted updates.","url":"../modules/banditrlproof-tsallisftrlfinitehorizonselection/index.html","parent":"chapter:tsallis","order":714,"meta":[["Source","BanditRLProof/TsallisFTRLFiniteHorizonSelection.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlfinitehorizonselection banditrlproof.tsallisftrlfinitehorizonselection this module applies the fixed half-tsallis minimizer choice to cumulative loss vectors. it removes the caller-supplied minimizer certificates from the deterministic finite-horizon ftrl decomposition and records the successor indexing needed by importance-weighted updates. lean module compiled","shard":"modules/2a195d5c65ce1238.json"},{"id":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","label":"TsallisFTRLGeneratedMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLGeneratedMeasurability","description":"This module derives measurability of the generated one-round stability score from coordinate measurability of the current and updated half-Tsallis selectors. The updated selector remains noncomputable, so its coordinate measurability is kept as an explicit selector contract rather than inferred from Classical.choose.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html","parent":"chapter:tsallis","order":715,"meta":[["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlgeneratedmeasurability banditrlproof.tsallisftrlgeneratedmeasurability this module derives measurability of the generated one-round stability score from coordinate measurability of the current and updated half-tsallis selectors. the updated selector remains noncomputable, so its coordinate measurability is kept as an explicit selector contract rather than inferred from classical.choose. lean module compiled","shard":"modules/f8c727f70ad58870.json"},{"id":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","label":"TsallisFTRLGeneratedRegularity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLGeneratedRegularity","description":"This module discharges integrability assumptions for the generated pure half-Tsallis stability route. The finite sampling law cancels the importance-weight denominator in the conditional absolute moment, while the simplex contracts uniformly bound the remaining finite sums.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html","parent":"chapter:tsallis","order":716,"meta":[["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlgeneratedregularity banditrlproof.tsallisftrlgeneratedregularity this module discharges integrability assumptions for the generated pure half-tsallis stability route. the finite sampling law cancels the importance-weight denominator in the conditional absolute moment, while the simplex contracts uniformly bound the remaining finite sums. lean module compiled","shard":"modules/fd9ae68b9a1257ea.json"},{"id":"module:BanditRLProof.TsallisFTRLInteriority","label":"TsallisFTRLInteriority","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLInteriority","description":"The square-root regularizer has an infinite inward slope at a zero coordinate. This module makes that boundary argument finite and algebraic: transfer a sufficiently small positive mass from any positive donor coordinate to the zero coordinate and contradict global minimality.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html","parent":"chapter:tsallis","order":717,"meta":[["Source","BanditRLProof/TsallisFTRLInteriority.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlinteriority banditrlproof.tsallisftrlinteriority the square-root regularizer has an infinite inward slope at a zero coordinate. this module makes that boundary argument finite and algebraic: transfer a sufficiently small positive mass from any positive donor coordinate to the zero coordinate and contradict global minimality. lean module compiled","shard":"modules/0cc655f0d890710a.json"},{"id":"module:BanditRLProof.TsallisFTRLMinimizerExistence","label":"TsallisFTRLMinimizerExistence","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLMinimizerExistence","description":"The project simplex only constrains coordinates in an explicit Finset, so it is not compact as a subset of the full function space when the ambient action type is infinite. We minimize instead on Mathlib's compact standard simplex over the finite subtype ↥arms, then extend the minimizer by zero.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html","parent":"chapter:tsallis","order":718,"meta":[["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean"],["Declarations","20"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlminimizerexistence banditrlproof.tsallisftrlminimizerexistence the project simplex only constrains coordinates in an explicit finset, so it is not compact as a subset of the full function space when the ambient action type is infinite. we minimize instead on mathlib's compact standard simplex over the finite subtype ↥arms, then extend the minimizer by zero. lean module compiled","shard":"modules/c7437c4d4a66ce50.json"},{"id":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","label":"TsallisFTRLMinimizerMeasurability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLMinimizerMeasurability","description":"The project minimizer is chosen noncomputably on the ambient action space, but the objective only sees the explicit finite arm set. We restrict the selected minimizer to that finite subtype and prove continuity in the restricted score vector. Compactness supplies cluster points and strict convexity identifies every cluster point with the unique minimizer.","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html","parent":"chapter:tsallis","order":719,"meta":[["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlminimizermeasurability banditrlproof.tsallisftrlminimizermeasurability the project minimizer is chosen noncomputably on the ambient action space, but the objective only sees the explicit finite arm set. we restrict the selected minimizer to that finite subtype and prove continuity in the restricted score vector. compactness supplies cluster points and strict convexity identifies every cluster point with the unique minimizer. lean module compiled","shard":"modules/0a891c6c0a977227.json"},{"id":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","label":"TsallisFTRLMinimizerUniqueness","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLMinimizerUniqueness","description":"The square-root sum is strictly concave on a nonempty standard simplex, so the half-Tsallis regularized objective is strictly convex there. The project-level finite simplex leaves coordinates outside arms unconstrained; accordingly, the public uniqueness theorem identifies exactly the supported coordinates.","url":"../modules/banditrlproof-tsallisftrlminimizeruniqueness/index.html","parent":"chapter:tsallis","order":720,"meta":[["Source","BanditRLProof/TsallisFTRLMinimizerUniqueness.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlminimizeruniqueness banditrlproof.tsallisftrlminimizeruniqueness the square-root sum is strictly concave on a nonempty standard simplex, so the half-tsallis regularized objective is strictly convex there. the project-level finite simplex leaves coordinates outside arms unconstrained; accordingly, the public uniqueness theorem identifies exactly the supported coordinates. lean module compiled","shard":"modules/b19721492d6eb203.json"},{"id":"module:BanditRLProof.TsallisFTRLOneStepStability","label":"TsallisFTRLOneStepStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLOneStepStability","description":"This module proves the deterministic one-step stability estimate used by the alpha = 1 / 2 Tsallis-INF route. It consumes explicit interior stationarity certificates for the current and importance-weighted updated distributions. The certificates expose exactly the KKT equation for the local objective eta * <p, score> + negEntropyRegularizer arms (1 / 2) p.","url":"../modules/banditrlproof-tsallisftrlonestepstability/index.html","parent":"chapter:tsallis","order":721,"meta":[["Source","BanditRLProof/TsallisFTRLOneStepStability.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlonestepstability banditrlproof.tsallisftrlonestepstability this module proves the deterministic one-step stability estimate used by the alpha = 1 / 2 tsallis-inf route. it consumes explicit interior stationarity certificates for the current and importance-weighted updated distributions. the certificates expose exactly the kkt equation for the local objective eta * <p, score> + negentropyregularizer arms (1 / 2) p. lean module compiled","shard":"modules/8302fed7a3be334d.json"},{"id":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","label":"TsallisFTRLRecursiveTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLRecursiveTrajectory","description":"This module builds the pure half-Tsallis finite-history policy and its Ionescu--Tulcea trajectory. The only selector-specific input is coordinate measurability of the canonical noncomputable minimizer for every measurable finite-history score. Simplex feasibility, score recursion, policy kernels, and conditional action laws are constructed locally.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html","parent":"chapter:tsallis","order":722,"meta":[["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean"],["Declarations","27"],["Project imports","4"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlrecursivetrajectory banditrlproof.tsallisftrlrecursivetrajectory this module builds the pure half-tsallis finite-history policy and its ionescu--tulcea trajectory. the only selector-specific input is coordinate measurability of the canonical noncomputable minimizer for every measurable finite-history score. simplex feasibility, score recursion, policy kernels, and conditional action laws are constructed locally. lean module compiled","shard":"modules/eca159a1335906e7.json"},{"id":"module:BanditRLProof.TsallisFTRLRegret","label":"TsallisFTRLRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLRegret","description":"This module lifts the one-step explicit-minimizer API to a finite-horizon regularized be-the-leader theorem and the standard FTRL stability/penalty regret decomposition. It then specializes the regularizer to negative Tsallis entropy and exposes the penalty as a difference of finite power sums.","url":"../modules/banditrlproof-tsallisftrlregret/index.html","parent":"chapter:tsallis","order":723,"meta":[["Source","BanditRLProof/TsallisFTRLRegret.lean"],["Declarations","12"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlregret banditrlproof.tsallisftrlregret this module lifts the one-step explicit-minimizer api to a finite-horizon regularized be-the-leader theorem and the standard ftrl stability/penalty regret decomposition. it then specializes the regularizer to negative tsallis entropy and exposes the penalty as a difference of finite power sums. lean module compiled","shard":"modules/2ffee97809724358.json"},{"id":"module:BanditRLProof.TsallisFTRLStationarity","label":"TsallisFTRLStationarity","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFTRLStationarity","description":"This module connects the finite-simplex minimizer certificate used by the generic FTRL route to the explicit interior stationarity certificate consumed by the half-Tsallis one-step stability theorem. The proof perturbs two simplex coordinates in opposite directions, differentiates the resulting scalar objective at an interior minimizer, and uses the vanishing derivative to identify a common multiplier.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html","parent":"chapter:tsallis","order":724,"meta":[["Source","BanditRLProof/TsallisFTRLStationarity.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisftrlstationarity banditrlproof.tsallisftrlstationarity this module connects the finite-simplex minimizer certificate used by the generic ftrl route to the explicit interior stationarity certificate consumed by the half-tsallis one-step stability theorem. the proof perturbs two simplex coordinates in opposite directions, differentiates the resulting scalar objective at an interior minimizer, and uses the vanishing derivative to identify a common multiplier. lean module compiled","shard":"modules/d144522ae1500cb1.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","label":"TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html","parent":"chapter:tsallis","order":725,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidarmdependentsuboptimalboostregret banditrlproof.tsallisfinitearmiidarmdependentsuboptimalboostregret generated source map for banditrlproof/tsallisfinitearmiidarmdependentsuboptimalboostregret.lean. lean module compiled","shard":"modules/68d103113a899997.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","label":"TsallisFiniteArmIIDCorruptedRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","description":"This module replaces the free nonnegative corruption parameter of the abstract self-bounding endpoint by a quantity derived from a concrete process. Each arm receives a fixed real reward shift, the shifted reward is projected back to [0, 1], and fresh finite reward vectors remain IID across rounds.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html","parent":"chapter:tsallis","order":726,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean"],["Declarations","20"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidcorruptedrewardlaw banditrlproof.tsallisfinitearmiidcorruptedrewardlaw this module replaces the free nonnegative corruption parameter of the abstract self-bounding endpoint by a quantity derived from a concrete process. each arm receives a fixed real reward shift, the shifted reward is projected back to [0, 1], and fresh finite reward vectors remain iid across rounds. lean module compiled","shard":"modules/a8370caa8eaa971c.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","label":"TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","description":"The base reward vector is fresh IID at every round. Before the current action is sampled, a measurable arm-dependent reward shift may be selected from the preceding finite pair history. Shifted rewards are clipped to [0, 1]. A deterministic time-and-arm envelope controls the resulting baseline-gap perturbation and therefore supplies the additive regret budget.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html","parent":"chapter:tsallis","order":727,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean"],["Declarations","15"],["Project imports","3"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw banditrlproof.tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw the base reward vector is fresh iid at every round. before the current action is sampled, a measurable arm-dependent reward shift may be selected from the preceding finite pair history. shifted rewards are clipped to [0, 1]. a deterministic time-and-arm envelope controls the resulting baseline-gap perturbation and therefore supplies the additive regret budget. lean module compiled","shard":"modules/41dd857f9f56104b.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","label":"TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html","parent":"chapter:tsallis","order":728,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean"],["Declarations","18"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw banditrlproof.tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw generated source map for banditrlproof/tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw.lean. lean module compiled","shard":"modules/0ea12c1f9e104637.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","label":"TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw.lean.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw/index.html","parent":"chapter:tsallis","order":729,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw.lean"],["Declarations","4"],["Project imports","3"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw banditrlproof.tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw generated source map for banditrlproof/tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw.lean. lean module compiled","shard":"modules/1e549752339dbd02.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","label":"TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","description":"This module packages only the predictable reward-shift data used through a fixed finite horizon. The source is extended by zero after that horizon and then consumed by the existing all-time expected-corruption theorem route.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html","parent":"chapter:tsallis","order":730,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean"],["Declarations","19"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw banditrlproof.tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw this module packages only the predictable reward-shift data used through a fixed finite horizon. the source is extended by zero after that horizon and then consumed by the existing all-time expected-corruption theorem route. lean module compiled","shard":"modules/3a1cebcc93d6b123.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","label":"TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret.lean.","url":"../modules/banditrlproof-tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret/index.html","parent":"chapter:tsallis","order":731,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret banditrlproof.tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret generated source map for banditrlproof/tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret.lean. lean module compiled","shard":"modules/b5361afee43a7f6c.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","label":"TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret.lean.","url":"../modules/banditrlproof-tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret/index.html","parent":"chapter:tsallis","order":732,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret banditrlproof.tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret generated source map for banditrlproof/tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret.lean. lean module compiled","shard":"modules/a79a1b34f8c9e3b5.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","label":"TsallisFiniteArmIIDRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDRewardLaw","description":"This module turns one probability law per rational-valued arm into a finite product reward-vector law. Rewards are clipped pointwise before conversion to losses so that the abstract IID loss-state API has global [0, 1] bounds; an almost-sure [0, 1] arm-law contract then proves that clipping does not change the coordinate means or the finite-bandit gaps.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html","parent":"chapter:tsallis","order":733,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidrewardlaw banditrlproof.tsallisfinitearmiidrewardlaw this module turns one probability law per rational-valued arm into a finite product reward-vector law. rewards are clipped pointwise before conversion to losses so that the abstract iid loss-state api has global [0, 1] bounds; an almost-sure [0, 1] arm-law contract then proves that clipping does not change the coordinate means or the finite-bandit gaps. lean module compiled","shard":"modules/ffb1c4fe34ae905a.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","label":"TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","description":"The base reward vector is fresh IID at every round. Before play begins, a deterministic arm-dependent reward shift is fixed for every round; shifted rewards are clipped back to [0,1]. Such a schedule is predictable and oblivious. It need not be stationary, but it does not depend on the realized history.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html","parent":"chapter:tsallis","order":734,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean"],["Declarations","11"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidtimevaryingcorruptedrewardlaw banditrlproof.tsallisfinitearmiidtimevaryingcorruptedrewardlaw the base reward vector is fresh iid at every round. before play begins, a deterministic arm-dependent reward shift is fixed for every round; shifted rewards are clipped back to [0,1]. such a schedule is predictable and oblivious. it need not be stationary, but it does not depend on the realized history. lean module compiled","shard":"modules/d6d5939c9663cbce.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","label":"TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html","parent":"chapter:tsallis","order":735,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiidtimevaryingsuboptimalboostregret banditrlproof.tsallisfinitearmiidtimevaryingsuboptimalboostregret generated source map for banditrlproof/tsallisfinitearmiidtimevaryingsuboptimalboostregret.lean. lean module compiled","shard":"modules/cc17252d1756c662.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","label":"TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","description":"Generated source map for BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html","parent":"chapter:tsallis","order":736,"meta":[["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmiiduniformsuboptimalboostrefinedregret banditrlproof.tsallisfinitearmiiduniformsuboptimalboostrefinedregret generated source map for banditrlproof/tsallisfinitearmiiduniformsuboptimalboostrefinedregret.lean. lean module compiled","shard":"modules/f6fddcb46ffd96bb.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","label":"TsallisFiniteArmIndependentDriftingMeanAllRegimes","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","description":"This module combines the compact-window refined endpoint with the logarithmic fallback for the same independent, nonidentical generated reward law. The comparator remains the fixed baseline arm model.bestArm.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanallregimes/index.html","parent":"chapter:tsallis","order":737,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanAllRegimes.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentdriftingmeanallregimes banditrlproof.tsallisfinitearmindependentdriftingmeanallregimes this module combines the compact-window refined endpoint with the logarithmic fallback for the same independent, nonidentical generated reward law. the comparator remains the fixed baseline arm model.bestarm. lean module compiled","shard":"modules/51c9062786bf444c.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","label":"TsallisFiniteArmIndependentDriftingMeanDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","description":"This module upgrades the fixed-baseline all-regimes theorem to the comparator that maximizes the actual reward mean at every round. The dynamic regret is decomposed into fixed-model.bestArm regret and the actual mean advantage of the moving comparator. The supplied armwise mean-deviation envelope controls that second term.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html","parent":"chapter:tsallis","order":738,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentdriftingmeandynamicregret banditrlproof.tsallisfinitearmindependentdriftingmeandynamicregret this module upgrades the fixed-baseline all-regimes theorem to the comparator that maximizes the actual reward mean at every round. the dynamic regret is decomposed into fixed-model.bestarm regret and the actual mean advantage of the moving comparator. the supplied armwise mean-deviation envelope controls that second term. lean module compiled","shard":"modules/28ce4068a99827b0.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","label":"TsallisFiniteArmIndependentDriftingMeanRefinedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","description":"This module connects the explicit mean-deviation budget of independent, nonidentical finite-arm reward laws to the compiled refined local self-bounding optimizer. The comparator remains the fixed baseline arm model.bestArm.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrefinedregret/index.html","parent":"chapter:tsallis","order":739,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRefinedRegret.lean"],["Declarations","2"],["Project imports","3"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentdriftingmeanrefinedregret banditrlproof.tsallisfinitearmindependentdriftingmeanrefinedregret this module connects the explicit mean-deviation budget of independent, nonidentical finite-arm reward laws to the compiled refined local self-bounding optimizer. the comparator remains the fixed baseline arm model.bestarm. lean module compiled","shard":"modules/55162eabfb656bf3.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","label":"TsallisFiniteArmIndependentDriftingMeanRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","description":"Each round may use a different probability law for every arm. The roundwise arm means may drift from a fixed finite-bandit model, with a deterministic coordinatewise deviation envelope. The resulting regret theorem charges the induced two-arm gap deviation explicitly.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html","parent":"chapter:tsallis","order":740,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentdriftingmeanrewardlaw banditrlproof.tsallisfinitearmindependentdriftingmeanrewardlaw each round may use a different probability law for every arm. the roundwise arm means may drift from a fixed finite-bandit model, with a deterministic coordinatewise deviation envelope. the resulting regret theorem charges the induced two-arm gap deviation explicitly. lean module compiled","shard":"modules/3309325c46f3735e.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","label":"TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","description":"This module replaces every time-indexed global prefix count in the logarithmic dynamic-regret route by the single terminal count at horizon. The resulting bound is explicit and finite-sum free in its nonstationarity term, but remains linear rather than minimax-sharp.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html","parent":"chapter:tsallis","order":741,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret banditrlproof.tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret this module replaces every time-indexed global prefix count in the logarithmic dynamic-regret route by the single terminal count at horizon. the resulting bound is explicit and finite-sum free in its nonstationarity term, but remains linear rather than minimax-sharp. lean module compiled","shard":"modules/d1f98a4d6221487f.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","label":"TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","description":"This module replaces the arm-indexed population-mean switch envelope by one global prefix count. A round contributes exactly when at least one arm's population mean changes.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html","parent":"chapter:tsallis","order":742,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentglobalmeanswitchcountdynamicregret banditrlproof.tsallisfinitearmindependentglobalmeanswitchcountdynamicregret this module replaces the arm-indexed population-mean switch envelope by one global prefix count. a round contributes exactly when at least one arm's population mean changes. lean module compiled","shard":"modules/ece891261c1097c5.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","label":"TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","description":"This module bounds each arm's cumulative population-mean path variation by the number of its nonzero consecutive mean changes. The generated dynamic regret theorem therefore needs no caller-supplied variation or switch-count budget.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html","parent":"chapter:tsallis","order":743,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean"],["Declarations","9"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentmeanswitchcountdynamicregret banditrlproof.tsallisfinitearmindependentmeanswitchcountdynamicregret this module bounds each arm's cumulative population-mean path variation by the number of its nonzero consecutive mean changes. the generated dynamic regret theorem therefore needs no caller-supplied variation or switch-count budget. lean module compiled","shard":"modules/35ce12786adfe8b9.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","label":"TsallisFiniteArmIndependentPathVariationDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","description":"This module derives the mean-deviation envelope used by the compiled drifting-mean dynamic-regret theorem from the actual reward-mean path. The only model-alignment premise is equality of the actual and baseline means at round zero.","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html","parent":"chapter:tsallis","order":744,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentpathvariationdynamicregret banditrlproof.tsallisfinitearmindependentpathvariationdynamicregret this module derives the mean-deviation envelope used by the compiled drifting-mean dynamic-regret theorem from the actual reward-mean path. the only model-alignment premise is equality of the actual and baseline means at round zero. lean module compiled","shard":"modules/5aa0257adad200c0.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","label":"TsallisFiniteArmIndependentRewardLaw","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentRewardLaw","description":"Each round may use a different probability law for every arm, while all roundwise arm laws retain the means of one fixed finite-bandit model. The roundwise product reward vectors are independent across time but need not be identically distributed.","url":"../modules/banditrlproof-tsallisfinitearmindependentrewardlaw/index.html","parent":"chapter:tsallis","order":745,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentRewardLaw.lean"],["Declarations","4"],["Project imports","3"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentrewardlaw banditrlproof.tsallisfinitearmindependentrewardlaw each round may use a different probability law for every arm, while all roundwise arm laws retain the means of one fixed finite-bandit model. the roundwise product reward vectors are independent across time but need not be identically distributed. lean module compiled","shard":"modules/4555088178dcc4f4.json"},{"id":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","label":"TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","description":"This module gives an exact two-arm, one-switch calculation. The global population-mean switch count is one, while the moving-comparator mean advantage charged by the current fixed-plus-advantage decomposition grows linearly with the horizon. This is a proof-route obstruction, not a dynamic regret lower bound.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html","parent":"chapter:tsallis","order":746,"meta":[["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean"],["Declarations","21"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitearmindependentsingleswitchcomparatorobstruction banditrlproof.tsallisfinitearmindependentsingleswitchcomparatorobstruction this module gives an exact two-arm, one-switch calculation. the global population-mean switch count is one, while the moving-comparator mean advantage charged by the current fixed-plus-advantage decomposition grows linearly with the horizon. this is a proof-route obstruction, not a dynamic regret lower bound. lean module compiled","shard":"modules/a6a975af03ac1544.json"},{"id":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","label":"TsallisFiniteBanditMeanLoss","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisFiniteBanditMeanLoss","description":"This module turns a bounded finite-bandit mean-reward model into the history-independent predictable loss family 1 - mean. Its loss differences are exactly the model gaps, so the compiled square-root-schedule fixed-gap theorem applies without a caller-supplied predictable gap law.","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html","parent":"chapter:tsallis","order":747,"meta":[["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean"],["Declarations","6"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisfinitebanditmeanloss banditrlproof.tsallisfinitebanditmeanloss this module turns a bounded finite-bandit mean-reward model into the history-independent predictable loss family 1 - mean. its loss differences are exactly the model gaps, so the compiled square-root-schedule fixed-gap theorem applies without a caller-supplied predictable gap law. lean module compiled","shard":"modules/d5a6aca3bc377c2a.json"},{"id":"module:BanditRLProof.TsallisImportanceWeightedMoment","label":"TsallisImportanceWeightedMoment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisImportanceWeightedMoment","description":"This module isolates the finite-sum power-moment calculation used by the Tsallis-INF stability analysis. The sampled importance-weighted loss is weighted by the inverse Hessian scale p^(2 - alpha). Taking the finite sum weighted by the sampling masses gives exactly the power-weighted second moment sum_i loss_i^2 * p_i^(1 - alpha).","url":"../modules/banditrlproof-tsallisimportanceweightedmoment/index.html","parent":"chapter:tsallis","order":748,"meta":[["Source","BanditRLProof/TsallisImportanceWeightedMoment.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisimportanceweightedmoment banditrlproof.tsallisimportanceweightedmoment this module isolates the finite-sum power-moment calculation used by the tsallis-inf stability analysis. the sampled importance-weighted loss is weighted by the inverse hessian scale p^(2 - alpha). taking the finite sum weighted by the sampling masses gives exactly the power-weighted second moment sum_i loss_i^2 * p_i^(1 - alpha). lean module compiled","shard":"modules/f8ebe049226fec43.json"},{"id":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","label":"TsallisOracleRestartDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartDynamicRegret","description":"This module proves the finite-sum assembly behind a switch-aligned restart route. It does not construct a restarted selector or trajectory kernel: epoch-local fixed-comparator regret certificates remain explicit inputs.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html","parent":"chapter:tsallis","order":749,"meta":[["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartdynamicregret banditrlproof.tsallisoraclerestartdynamicregret this module proves the finite-sum assembly behind a switch-aligned restart route. it does not construct a restarted selector or trajectory kernel: epoch-local fixed-comparator regret certificates remain explicit inputs. lean module compiled","shard":"modules/05d4c2590f7ed16f.json"},{"id":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","label":"TsallisOracleRestartExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartExpectedRegret","description":"This module transports predictable importance-weighted first moments under the generated restart trajectory. The transport is performed on the actual global law and therefore does not assume independent fresh epoch runs.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html","parent":"chapter:tsallis","order":750,"meta":[["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean"],["Declarations","25"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartexpectedregret banditrlproof.tsallisoraclerestartexpectedregret this module transports predictable importance-weighted first moments under the generated restart trajectory. the transport is performed on the actual global law and therefore does not assume independent fresh epoch runs. lean module compiled","shard":"modules/f2bb346c779f5a7f.json"},{"id":"module:BanditRLProof.TsallisOracleRestartExpectedStability","label":"TsallisOracleRestartExpectedStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartExpectedStability","description":"This module transports one restart-local half-Tsallis potential-stability term through the conditional action law of the single generated restart trajectory. The shifted trajectory is used only for pathwise reindexing; no fresh independent epoch law is introduced.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html","parent":"chapter:tsallis","order":751,"meta":[["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean"],["Declarations","26"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartexpectedstability banditrlproof.tsallisoraclerestartexpectedstability this module transports one restart-local half-tsallis potential-stability term through the conditional action law of the single generated restart trajectory. the shifted trajectory is used only for pathwise reindexing; no fresh independent epoch law is introduced. lean module compiled","shard":"modules/5fe8711e9fdf9945.json"},{"id":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","label":"TsallisOracleRestartGeneratedDynamicRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","description":"This module feeds the generated, observed epoch certificate for the square-root half-Tsallis schedule into the predictable moving-comparator assembly. The switch-count-facing theorem keeps schedule cardinality as an explicit contract; deriving that contract from a concrete restart law is a separate obligation.","url":"../modules/banditrlproof-tsallisoraclerestartgenerateddynamicregret/index.html","parent":"chapter:tsallis","order":752,"meta":[["Source","BanditRLProof/TsallisOracleRestartGeneratedDynamicRegret.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartgenerateddynamicregret banditrlproof.tsallisoraclerestartgenerateddynamicregret this module feeds the generated, observed epoch certificate for the square-root half-tsallis schedule into the predictable moving-comparator assembly. the switch-count-facing theorem keeps schedule cardinality as an explicit contract; deriving that contract from a concrete restart law is a separate obligation. lean module compiled","shard":"modules/163d2cea7bd3fa4a.json"},{"id":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","label":"TsallisOracleRestartGeneratedTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartGeneratedTrajectory","description":"This module constructs the restarted selector and trajectory kernel needed by the oracle-restart dynamic-regret route. It proves the conditional action law of that generated process. Epoch-local regret transport remains downstream.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html","parent":"chapter:tsallis","order":753,"meta":[["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean"],["Declarations","22"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartgeneratedtrajectory banditrlproof.tsallisoraclerestartgeneratedtrajectory this module constructs the restarted selector and trajectory kernel needed by the oracle-restart dynamic-regret route. it proves the conditional action law of that generated process. epoch-local regret transport remains downstream. lean module compiled","shard":"modules/eb8cb529c6ce11ba.json"},{"id":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","label":"TsallisOracleRestartGlobalMeanSwitchCount","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","description":"This module builds a deterministic oracle schedule that restarts exactly after each supplied change point. It then specializes the change predicate to population-mean changes of an independent finite-arm reward law.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html","parent":"chapter:tsallis","order":754,"meta":[["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean"],["Declarations","16"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartglobalmeanswitchcount banditrlproof.tsallisoraclerestartglobalmeanswitchcount this module builds a deterministic oracle schedule that restarts exactly after each supplied change point. it then specializes the change predicate to population-mean changes of an independent finite-arm reward law. lean module compiled","shard":"modules/e59ee124d04434ff.json"},{"id":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","label":"TsallisOracleRestartPredictableRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartPredictableRegret","description":"This module puts the generated restart probability into the pointwise regret summand and aligns the deterministic epoch assembly with schedule.start. Epoch-local regret certificates remain explicit downstream inputs.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html","parent":"chapter:tsallis","order":755,"meta":[["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartpredictableregret banditrlproof.tsallisoraclerestartpredictableregret this module puts the generated restart probability into the pointwise regret summand and aligns the deterministic epoch assembly with schedule.start. epoch-local regret certificates remain explicit downstream inputs. lean module compiled","shard":"modules/8feacfdbb9c9dfb6.json"},{"id":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","label":"TsallisOracleRestartRefinedStabilityTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","description":"This module tunes the refined half-Tsallis one-round stability certificate on an actual contiguous oracle-restart epoch. All expectations remain under the single global generated restart law; the shifted trajectory is used only for pathwise local-time indexing.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html","parent":"chapter:tsallis","order":756,"meta":[["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean"],["Declarations","12"],["Project imports","5"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartrefinedstabilitytuning banditrlproof.tsallisoraclerestartrefinedstabilitytuning this module tunes the refined half-tsallis one-round stability certificate on an actual contiguous oracle-restart epoch. all expectations remain under the single global generated restart law; the shifted trajectory is used only for pathwise local-time indexing. lean module compiled","shard":"modules/accbe1b237996ad8.json"},{"id":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","label":"TsallisOracleRestartScoreAlignment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisOracleRestartScoreAlignment","description":"This module identifies every actual restart probability and stored-reward importance-weighted loss with the corresponding scheduled surface after shifting to the current epoch's local time. It then exposes the pathwise FTRL certificate on any explicit contiguous prefix of one restart epoch.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html","parent":"chapter:tsallis","order":757,"meta":[["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisoraclerestartscorealignment banditrlproof.tsallisoraclerestartscorealignment this module identifies every actual restart probability and stored-reward importance-weighted loss with the corresponding scheduled surface after shifting to the current epoch's local time. it then exposes the pathwise ftrl certificate on any explicit contiguous prefix of one restart epoch. lean module compiled","shard":"modules/109822b415d3fc83.json"},{"id":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","label":"TsallisRefinedAveragedStabilityObstruction","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisRefinedAveragedStabilityObstruction","description":"The shifted ordinary-importance-weighted moments have the paper-shaped sampled-action average, but the current local stability quantity <p - p_next, hatLoss> is not the conjugate-potential stability used in the Tsallis-INF proof. This module gives a fully rational two-arm counterexample: the current and sampled-update distributions are strict simplex minimizers and the losses lie in [0,1], yet the proposed refined a…","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html","parent":"chapter:tsallis","order":758,"meta":[["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean"],["Declarations","30"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisrefinedaveragedstabilityobstruction banditrlproof.tsallisrefinedaveragedstabilityobstruction the shifted ordinary-importance-weighted moments have the paper-shaped sampled-action average, but the current local stability quantity <p - p_next, hatloss> is not the conjugate-potential stability used in the tsallis-inf proof. this module gives a fully rational two-arm counterexample: the current and sampled-update distributions are strict simplex minimizers and the losses lie in [0,1], yet the proposed refined averaged upper bound fails. lean module compiled","shard":"modules/0d8cb34db93ae3ba.json"},{"id":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","label":"TsallisRefinedImportanceWeightedMoment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisRefinedImportanceWeightedMoment","description":"This module isolates the finite-sum calculation behind the ordinary importance-weighted part of the refined half-Tsallis stability route. The baseline is the sampled raw loss. The quadratic moment contracts to sum_a sqrt (p a) * (1 - p a), while the positive cubic remainder is at most one.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html","parent":"chapter:tsallis","order":759,"meta":[["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisrefinedimportanceweightedmoment banditrlproof.tsallisrefinedimportanceweightedmoment this module isolates the finite-sum calculation behind the ordinary importance-weighted part of the refined half-tsallis stability route. the baseline is the sampled raw loss. the quadratic moment contracts to sum_a sqrt (p a) * (1 - p a), while the positive cubic remainder is at most one. lean module compiled","shard":"modules/277aeb8ba81dc22e.json"},{"id":"module:BanditRLProof.TsallisRefinedSuboptimalStability","label":"TsallisRefinedSuboptimalStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisRefinedSuboptimalStability","description":"The paper's refined stability term contains sum_a sqrt (p a) * (1 - p a). This module proves the finite-simplex elimination of one optimal arm and connects that bound to the compiled self-bounding completion-of-squares consumer.","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html","parent":"chapter:tsallis","order":760,"meta":[["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisrefinedsuboptimalstability banditrlproof.tsallisrefinedsuboptimalstability the paper's refined stability term contains sum_a sqrt (p a) * (1 - p a). this module proves the finite-simplex elimination of one optimal arm and connects that bound to the compiled self-bounding completion-of-squares consumer. lean module compiled","shard":"modules/08461fc43d8decb4.json"},{"id":"module:BanditRLProof.TsallisRegularizer","label":"TsallisRegularizer","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisRegularizer","description":"This module records the first deterministic Real.rpow surface for Tsallis-INF/FTRL routes. It defines the finite-action Tsallis power sum, Tsallis entropy, and the negative-entropy regularizer used as an FTRL regularizer, then packages the small well-definedness side conditions over the finite simplex.","url":"../modules/banditrlproof-tsallisregularizer/index.html","parent":"chapter:tsallis","order":761,"meta":[["Source","BanditRLProof/TsallisRegularizer.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisregularizer banditrlproof.tsallisregularizer this module records the first deterministic real.rpow surface for tsallis-inf/ftrl routes. it defines the finite-action tsallis power sum, tsallis entropy, and the negative-entropy regularizer used as an ftrl regularizer, then packages the small well-definedness side conditions over the finite simplex. lean module compiled","shard":"modules/bef5147b4a2f36c0.json"},{"id":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","label":"TsallisScheduledAllRateExpectedStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledAllRateExpectedStability","description":"This module supplies the coarse branch omitted by the refined scheduled stability theorem. Ordinary importance weighting has one-round expected conjugate-potential stability at most one for every positive local rate. The final generated-trajectory theorem combines that fallback with the refined bound available when the local ABRL rate is at most 1 / 2.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html","parent":"chapter:tsallis","order":762,"meta":[["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean"],["Declarations","16"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledallrateexpectedstability banditrlproof.tsallisscheduledallrateexpectedstability this module supplies the coarse branch omitted by the refined scheduled stability theorem. ordinary importance weighting has one-round expected conjugate-potential stability at most one for every positive local rate. the final generated-trajectory theorem combines that fallback with the refined bound available when the local abrl rate is at most 1 / 2. lean module compiled","shard":"modules/6f27333bec391692.json"},{"id":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","label":"TsallisScheduledAllTimesExpectedStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledAllTimesExpectedStability","description":"This module joins the distinct time-zero and successor conditioning surfaces. Its final left side is exactly the full Finset.range (horizon + 1) stability sum consumed by the scheduled pathwise regret decomposition.","url":"../modules/banditrlproof-tsallisscheduledalltimesexpectedstability/index.html","parent":"chapter:tsallis","order":763,"meta":[["Source","BanditRLProof/TsallisScheduledAllTimesExpectedStability.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledalltimesexpectedstability banditrlproof.tsallisscheduledalltimesexpectedstability this module joins the distinct time-zero and successor conditioning surfaces. its final left side is exactly the full finset.range (horizon + 1) stability sum consumed by the scheduled pathwise regret decomposition. lean module compiled","shard":"modules/6066298328b3c736.json"},{"id":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","label":"TsallisScheduledConditionalMeanGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledConditionalMeanGap","description":"This module exposes the sampling-before-action sigma-algebra at each scheduled time. The scheduled action probability is measurable with respect to that sigma-algebra, so Mathlib's conditional-expectation pull-out theorem turns a constant conditional loss-gap law into HasScheduledExpectedGapLaw.","url":"../modules/banditrlproof-tsallisscheduledconditionalmeangap/index.html","parent":"chapter:tsallis","order":764,"meta":[["Source","BanditRLProof/TsallisScheduledConditionalMeanGap.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledconditionalmeangap banditrlproof.tsallisscheduledconditionalmeangap this module exposes the sampling-before-action sigma-algebra at each scheduled time. the scheduled action probability is measurable with respect to that sigma-algebra, so mathlib's conditional-expectation pull-out theorem turns a constant conditional loss-gap law into hasscheduledexpectedgaplaw. lean module compiled","shard":"modules/8bdf146310c99785.json"},{"id":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","label":"TsallisScheduledExpectedGapSelfBounding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledExpectedGapSelfBounding","description":"This module replaces the samplewise fixed-gap premise by the stochastic first-moment law actually needed by self-bounding. For every time and suboptimal arm, the probability-weighted predictable loss difference has expectation gap action * E[p_t(action)].","url":"../modules/banditrlproof-tsallisscheduledexpectedgapselfbounding/index.html","parent":"chapter:tsallis","order":765,"meta":[["Source","BanditRLProof/TsallisScheduledExpectedGapSelfBounding.lean"],["Declarations","5"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledexpectedgapselfbounding banditrlproof.tsallisscheduledexpectedgapselfbounding this module replaces the samplewise fixed-gap premise by the stochastic first-moment law actually needed by self-bounding. for every time and suboptimal arm, the probability-weighted predictable loss difference has expectation gap action * e[p_t(action)]. lean module compiled","shard":"modules/7b012c9da4988fc4.json"},{"id":"module:BanditRLProof.TsallisScheduledExpectedRegret","label":"TsallisScheduledExpectedRegret","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledExpectedRegret","description":"This module transports the observed importance-weighted regret of the scheduled half-Tsallis trajectory to predictable environment regret. It then combines the pathwise time-varying penalty theorem with the all-rate expected stability theorem.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html","parent":"chapter:tsallis","order":766,"meta":[["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean"],["Declarations","13"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledexpectedregret banditrlproof.tsallisscheduledexpectedregret this module transports the observed importance-weighted regret of the scheduled half-tsallis trajectory to predictable environment regret. it then combines the pathwise time-varying penalty theorem with the all-rate expected stability theorem. lean module compiled","shard":"modules/afeded321c420efe.json"},{"id":"module:BanditRLProof.TsallisScheduledExpectedStability","label":"TsallisScheduledExpectedStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledExpectedStability","description":"This module transports the one-round ordinary-importance-weighted conjugate-potential bound through the generated scheduled action law. Round n + 1 uses eta (n + 1), so the finite sum can vary the learning rate without weakening the deterministic stability theorem.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html","parent":"chapter:tsallis","order":767,"meta":[["Source","BanditRLProof/TsallisScheduledExpectedStability.lean"],["Declarations","19"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledexpectedstability banditrlproof.tsallisscheduledexpectedstability this module transports the one-round ordinary-importance-weighted conjugate-potential bound through the generated scheduled action law. round n + 1 uses eta (n + 1), so the finite sum can vary the learning rate without weakening the deterministic stability theorem. lean module compiled","shard":"modules/3bfc0e062b30519e.json"},{"id":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","label":"TsallisScheduledFixedGapSelfBounding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledFixedGapSelfBounding","description":"This module identifies generated scheduled predictable environment regret with expected suboptimal-arm gap mass under an exact predictable fixed-gap law. It discharges the explicit self-bounding premise of the scheduled completion-of-squares theorem.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html","parent":"chapter:tsallis","order":768,"meta":[["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledfixedgapselfbounding banditrlproof.tsallisscheduledfixedgapselfbounding this module identifies generated scheduled predictable environment regret with expected suboptimal-arm gap mass under an exact predictable fixed-gap law. it discharges the explicit self-bounding premise of the scheduled completion-of-squares theorem. lean module compiled","shard":"modules/2efdc5dff2ff34e7.json"},{"id":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","label":"TsallisScheduledIIDHistoryAdaptive","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledIIDHistoryAdaptive","description":"The loss chosen at a successor round may inspect the complete pre-action pair history. It may inspect the fresh IID state only through the current state coordinate. This coordinate-locality contract is exactly what is needed for the canonical visible trajectory to factor through every finite state prefix.","url":"../modules/banditrlproof-tsallisschedulediidhistoryadaptive/index.html","parent":"chapter:tsallis","order":769,"meta":[["Source","BanditRLProof/TsallisScheduledIIDHistoryAdaptive.lean"],["Declarations","4"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisschedulediidhistoryadaptive banditrlproof.tsallisschedulediidhistoryadaptive the loss chosen at a successor round may inspect the complete pre-action pair history. it may inspect the fresh iid state only through the current state coordinate. this coordinate-locality contract is exactly what is needed for the canonical visible trajectory to factor through every finite state prefix. lean module compiled","shard":"modules/373736b152e0a15b.json"},{"id":"module:BanditRLProof.TsallisScheduledIIDMeanGap","label":"TsallisScheduledIIDMeanGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledIIDMeanGap","description":"This module instantiates the abstract scheduled independent-mean contract with an infinite product of loss states. The only trajectory-side input is an explicit finite-prefix kernel factorization: the visible prefix through n must depend on the loss-state stream only through its coordinates through n.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html","parent":"chapter:tsallis","order":770,"meta":[["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean"],["Declarations","12"],["Project imports","4"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisschedulediidmeangap banditrlproof.tsallisschedulediidmeangap this module instantiates the abstract scheduled independent-mean contract with an infinite product of loss states. the only trajectory-side input is an explicit finite-prefix kernel factorization: the visible prefix through n must depend on the loss-state stream only through its coordinates through n. lean module compiled","shard":"modules/20709c4534b72170.json"},{"id":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","label":"TsallisScheduledIIDTimeVaryingMeanGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","description":"Fresh states remain IID, but the deterministic evaluator may depend on the round. This is the process-law layer for an oblivious predictable corruption schedule. The generated trajectory still factors through each finite state prefix, so the current coordinate is independent of the pre-action trace.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html","parent":"chapter:tsallis","order":771,"meta":[["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisschedulediidtimevaryingmeangap banditrlproof.tsallisschedulediidtimevaryingmeangap fresh states remain iid, but the deterministic evaluator may depend on the round. this is the process-law layer for an oblivious predictable corruption schedule. the generated trajectory still factors through each finite state prefix, so the current coordinate is independent of the pre-action trace. lean module compiled","shard":"modules/fbe3fe5a6b58802d.json"},{"id":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","label":"TsallisScheduledIndependentMeanGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledIndependentMeanGap","description":"This module turns an independence-plus-global-mean contract for every predictable loss difference into the conditional-mean law used by scheduled self-bounding. It also exposes the resulting logarithmic square-root-schedule regret theorem on the generated trajectory law.","url":"../modules/banditrlproof-tsallisscheduledindependentmeangap/index.html","parent":"chapter:tsallis","order":772,"meta":[["Source","BanditRLProof/TsallisScheduledIndependentMeanGap.lean"],["Declarations","4"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledindependentmeangap banditrlproof.tsallisscheduledindependentmeangap this module turns an independence-plus-global-mean contract for every predictable loss difference into the conditional-mean law used by scheduled self-bounding. it also exposes the resulting logarithmic square-root-schedule regret theorem on the generated trajectory law. lean module compiled","shard":"modules/203d1924167df739.json"},{"id":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","label":"TsallisScheduledInitialExpectedStability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledInitialExpectedStability","description":"This module closes the time-zero conditioning surface omitted from the scheduled successor stability theorem. The left side is the actual time-zero conjugate-potential term from the scheduled pathwise decomposition.","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html","parent":"chapter:tsallis","order":773,"meta":[["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledinitialexpectedstability banditrlproof.tsallisscheduledinitialexpectedstability this module closes the time-zero conditioning surface omitted from the scheduled successor stability theorem. the left side is the actual time-zero conjugate-potential term from the scheduled pathwise decomposition. lean module compiled","shard":"modules/a01c4cca68684c12.json"},{"id":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","label":"TsallisScheduledRecursiveTrajectory","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledRecursiveTrajectory","description":"This module lifts the recursive pure half-Tsallis trajectory from a fixed learning rate to a deterministic schedule. The initial action uses eta 0; after the visible prefix through round n, the successor policy uses eta (n + 1). This indexing matches the scheduled FTRL penalty route.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html","parent":"chapter:tsallis","order":774,"meta":[["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean"],["Declarations","13"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledrecursivetrajectory banditrlproof.tsallisscheduledrecursivetrajectory this module lifts the recursive pure half-tsallis trajectory from a fixed learning rate to a deterministic schedule. the initial action uses eta 0; after the visible prefix through round n, the successor policy uses eta (n + 1). this indexing matches the scheduled ftrl penalty route. lean module compiled","shard":"modules/8634a6ad55dda625.json"},{"id":"module:BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","label":"TsallisScheduledReferenceGapExpectedDeviationSelfBounding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","description":"This variant retains the scheduled action probability inside the corruption allowance. A sample-dependent predictable deviation therefore contributes its actual probability-weighted expectation instead of a deterministic pointwise envelope.","url":"../modules/banditrlproof-tsallisscheduledreferencegapexpecteddeviationselfbounding/index.html","parent":"chapter:tsallis","order":775,"meta":[["Source","BanditRLProof/TsallisScheduledReferenceGapExpectedDeviationSelfBounding.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledreferencegapexpecteddeviationselfbounding banditrlproof.tsallisscheduledreferencegapexpecteddeviationselfbounding this variant retains the scheduled action probability inside the corruption allowance. a sample-dependent predictable deviation therefore contributes its actual probability-weighted expectation instead of a deterministic pointwise envelope. lean module compiled","shard":"modules/eca4fc0ad7196908.json"},{"id":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","label":"TsallisScheduledReferenceGapSelfBounding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledReferenceGapSelfBounding","description":"The action probabilities and regret are generated by actualLoss, while a second predictable loss vector on the same trajectory supplies the stochastic baseline gap law. A pointwise bound between the two loss differences then converts the reference expected-gap law into a self-bound for the actual regret. This is the generic bridge needed by history-adaptive corruption.","url":"../modules/banditrlproof-tsallisscheduledreferencegapselfbounding/index.html","parent":"chapter:tsallis","order":776,"meta":[["Source","BanditRLProof/TsallisScheduledReferenceGapSelfBounding.lean"],["Declarations","1"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledreferencegapselfbounding banditrlproof.tsallisscheduledreferencegapselfbounding the action probabilities and regret are generated by actualloss, while a second predictable loss vector on the same trajectory supplies the stochastic baseline gap law. a pointwise bound between the two loss differences then converts the reference expected-gap law into a self-bound for the actual regret. this is the generic bridge needed by history-adaptive corruption. lean module compiled","shard":"modules/963e29109b7e5e4c.json"},{"id":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","label":"TsallisScheduledRefinedExpectedPenalty","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledRefinedExpectedPenalty","description":"This module keeps the reciprocal-rate increments from the deterministic time-varying penalty theorem instead of collapsing them into a terminal potential mass. The mass above the point-mass baseline is eliminated in favor of suboptimal-arm terms, then transported through expectation by the compiled square-root Jensen bridge.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html","parent":"chapter:tsallis","order":777,"meta":[["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean"],["Declarations","8"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledrefinedexpectedpenalty banditrlproof.tsallisscheduledrefinedexpectedpenalty this module keeps the reciprocal-rate increments from the deterministic time-varying penalty theorem instead of collapsing them into a terminal potential mass. the mass above the point-mass baseline is eliminated in favor of suboptimal-arm terms, then transported through expectation by the compiled square-root jensen bridge. lean module compiled","shard":"modules/f4fa09eb862ed7a2.json"},{"id":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","label":"TsallisScheduledRefinedStabilityPenalty","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledRefinedStabilityPenalty","description":"This module combines the generated small-rate stability budget with the uncollapsed refined expected penalty. It exposes one coefficient per actual time and feeds the resulting square-root upper bound to the exact fixed-gap self-bounding consumer.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html","parent":"chapter:tsallis","order":778,"meta":[["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledrefinedstabilitypenalty banditrlproof.tsallisscheduledrefinedstabilitypenalty this module combines the generated small-rate stability budget with the uncollapsed refined expected penalty. it exposes one coefficient per actual time and feeds the resulting square-root upper bound to the exact fixed-gap self-bounding consumer. lean module compiled","shard":"modules/46556ada78c750ef.json"},{"id":"module:BanditRLProof.TsallisScheduledScoreAlignment","label":"TsallisScheduledScoreAlignment","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledScoreAlignment","description":"This module identifies the recursively accumulated scheduled half-Tsallis history score with FTRL.cumulativeLoss of the actual observed importance-weighted loss vectors. It then rewrites the deterministic time-varying penalty theorem onto one generated trajectory sample.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html","parent":"chapter:tsallis","order":779,"meta":[["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean"],["Declarations","11"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledscorealignment banditrlproof.tsallisscheduledscorealignment this module identifies the recursively accumulated scheduled half-tsallis history score with ftrl.cumulativeloss of the actual observed importance-weighted loss vectors. it then rewrites the deterministic time-varying penalty theorem onto one generated trajectory sample. lean module compiled","shard":"modules/b2f112affca17175.json"},{"id":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","label":"TsallisScheduledSelfBoundingInterpolation","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledSelfBoundingInterpolation","description":"This module formalizes the lambda interpolation step used before the joint probability/simplex optimization in self-bounding Tsallis-INF analyses. It combines an upper regret estimate with a terminal self-bounding inequality; no prefix self-bounding assumption is required.","url":"../modules/banditrlproof-tsallisscheduledselfboundinginterpolation/index.html","parent":"chapter:tsallis","order":780,"meta":[["Source","BanditRLProof/TsallisScheduledSelfBoundingInterpolation.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledselfboundinginterpolation banditrlproof.tsallisscheduledselfboundinginterpolation this module formalizes the lambda interpolation step used before the joint probability/simplex optimization in self-bounding tsallis-inf analyses. it combines an upper regret estimate with a terminal self-bounding inequality; no prefix self-bounding assumption is required. lean module compiled","shard":"modules/80ebd323c80dd90b.json"},{"id":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","label":"TsallisScheduledSelfBoundingOptimization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledSelfBoundingOptimization","description":"This module consumes the terminal lambda interpolation and the two one-round quadratic branches. It partitions a finite time set by the exact active-mass threshold, then exposes the resulting deterministic scalar sums to the remaining schedule and lambda optimization.","url":"../modules/banditrlproof-tsallisscheduledselfboundingoptimization/index.html","parent":"chapter:tsallis","order":781,"meta":[["Source","BanditRLProof/TsallisScheduledSelfBoundingOptimization.lean"],["Declarations","5"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledselfboundingoptimization banditrlproof.tsallisscheduledselfboundingoptimization this module consumes the terminal lambda interpolation and the two one-round quadratic branches. it partitions a finite time set by the exact active-mass threshold, then exposes the resulting deterministic scalar sums to the remaining schedule and lambda optimization. lean module compiled","shard":"modules/348c2c11d8f79600.json"},{"id":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","label":"TsallisScheduledSuboptimalExpectedBound","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledSuboptimalExpectedBound","description":"This module converts the generated pathwise refined half-power budget into a deterministic expression involving square roots of expected action probabilities. It is the Jensen bridge between the compiled scheduled expected-regret theorem and the self-bounding completion-of-squares consumer.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html","parent":"chapter:tsallis","order":782,"meta":[["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledsuboptimalexpectedbound banditrlproof.tsallisscheduledsuboptimalexpectedbound this module converts the generated pathwise refined half-power budget into a deterministic expression involving square roots of expected action probabilities. it is the jensen bridge between the compiled scheduled expected-regret theorem and the self-bounding completion-of-squares consumer. lean module compiled","shard":"modules/e53088856a6c064e.json"},{"id":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","label":"TsallisScheduledTimeVaryingExpectedGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","description":"This module lets the conditional or independent mean loss gap vary with the round. It is the law surface needed by deterministic predictable corruption schedules: the baseline stochastic gap remains fixed, while clipping a round-dependent reward shift perturbs the actual gap by a known per-round amount.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html","parent":"chapter:tsallis","order":783,"meta":[["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean"],["Declarations","9"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisscheduledtimevaryingexpectedgap banditrlproof.tsallisscheduledtimevaryingexpectedgap this module lets the conditional or independent mean loss gap vary with the round. it is the law surface needed by deterministic predictable corruption schedules: the baseline stochastic gap remains fixed, while clipping a round-dependent reward shift perturbs the actual gap by a known per-round amount. lean module compiled","shard":"modules/4962101e2a47ba19.json"},{"id":"module:BanditRLProof.TsallisSelfBounding","label":"TsallisSelfBounding","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSelfBounding","description":"This module records the paper-facing (Delta, C, T) self-bounding condition, identifies it on the generated predictable trajectory under an exact gap law, and proves the finite completion-of-squares conversion used after a refined suboptimal-arm stability bound is available.","url":"../modules/banditrlproof-tsallisselfbounding/index.html","parent":"chapter:tsallis","order":784,"meta":[["Source","BanditRLProof/TsallisSelfBounding.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisselfbounding banditrlproof.tsallisselfbounding this module records the paper-facing (delta, c, t) self-bounding condition, identifies it on the generated predictable trajectory under an exact gap law, and proves the finite completion-of-squares conversion used after a refined suboptimal-arm stability bound is available. lean module compiled","shard":"modules/d1fde7e2599ad19b.json"},{"id":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","label":"TsallisSelfBoundingBetaRoot","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSelfBoundingBetaRoot","description":"This module isolates the scalar intermediate-value step used by the refined self-bounding tuning route. It proves existence of a root in the paper's admissible beta interval without assuming a Lambert-W API. Quantitative bounds on that root remain downstream obligations.","url":"../modules/banditrlproof-tsallisselfboundingbetaroot/index.html","parent":"chapter:tsallis","order":785,"meta":[["Source","BanditRLProof/TsallisSelfBoundingBetaRoot.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallisselfboundingbetaroot banditrlproof.tsallisselfboundingbetaroot this module isolates the scalar intermediate-value step used by the refined self-bounding tuning route. it proves existence of a root in the paper's admissible beta interval without assuming a lambert-w api. quantitative bounds on that root remain downstream obligations. lean module compiled","shard":"modules/40098b95879e0116.json"},{"id":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","label":"TsallisSqrtScheduleFixedGap","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSqrtScheduleFixedGap","description":"This module instantiates the scheduled refined stability-penalty theorem with eta t = 1 / (2 * sqrt (t + 1)). The schedule contracts and its unified coefficient are bounded explicitly before the fixed-gap theorem is consumed.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html","parent":"chapter:tsallis","order":786,"meta":[["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean"],["Declarations","17"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallissqrtschedulefixedgap banditrlproof.tsallissqrtschedulefixedgap this module instantiates the scheduled refined stability-penalty theorem with eta t = 1 / (2 * sqrt (t + 1)). the schedule contracts and its unified coefficient are bounded explicitly before the fixed-gap theorem is consumed. lean module compiled","shard":"modules/ce68acf4a8761837.json"},{"id":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","label":"TsallisSqrtScheduleSelfBoundingOptimization","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","description":"This module specializes the refined generated quadratic split to the local schedule eta_t = 1 / (2 * sqrt (t + 1)). The refined coefficient is replaced by its compiled 5 / sqrt (t + 1) envelope and the deterministic rate-square base is identified with one half of the harmonic budget.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html","parent":"chapter:tsallis","order":787,"meta":[["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean"],["Declarations","7"],["Project imports","2"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallissqrtscheduleselfboundingoptimization banditrlproof.tsallissqrtscheduleselfboundingoptimization this module specializes the refined generated quadratic split to the local schedule eta_t = 1 / (2 * sqrt (t + 1)). the refined coefficient is replaced by its compiled 5 / sqrt (t + 1) envelope and the deterministic rate-square base is identified with one half of the harmonic budget. lean module compiled","shard":"modules/011d9f7321b4e3dd.json"},{"id":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","label":"TsallisSqrtScheduleSelfBoundingRefinedScalar","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","description":"The local refined stability envelope has coefficient 5. After the change of variables used by the square-root schedule, its scalar objective therefore has a beta equation with offset 2, rather than the paper's idealized offset 1. This module records the corrected root interval and an elementary root estimate that does not require Lambert W.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html","parent":"chapter:tsallis","order":788,"meta":[["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean"],["Declarations","14"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallissqrtscheduleselfboundingrefinedscalar banditrlproof.tsallissqrtscheduleselfboundingrefinedscalar the local refined stability envelope has coefficient 5. after the change of variables used by the square-root schedule, its scalar objective therefore has a beta equation with offset 2, rather than the paper's idealized offset 1. this module records the corrected root interval and an elementary root estimate that does not require lambert w. lean module compiled","shard":"modules/fb68ef46c66a1727.json"},{"id":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","label":"TsallisSqrtScheduleSelfBoundingRefinedTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","description":"This module transports the coefficient-aware scalar variables into the actual finite-arm square-root schedule. In particular, it identifies the continuous floor threshold exactly as 2 * (T + 1) / beta.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html","parent":"chapter:tsallis","order":789,"meta":[["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean"],["Declarations","6"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallissqrtscheduleselfboundingrefinedtuning banditrlproof.tsallissqrtscheduleselfboundingrefinedtuning this module transports the coefficient-aware scalar variables into the actual finite-arm square-root schedule. in particular, it identifies the continuous floor threshold exactly as 2 * (t + 1) / beta. lean module compiled","shard":"modules/84246877a8b9e43a.json"},{"id":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","label":"TsallisSqrtScheduleSelfBoundingRefinedWindow","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","description":"Generated source map for BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedWindow.lean.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedwindow/index.html","parent":"chapter:tsallis","order":790,"meta":[["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedWindow.lean"],["Declarations","2"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallissqrtscheduleselfboundingrefinedwindow banditrlproof.tsallissqrtscheduleselfboundingrefinedwindow generated source map for banditrlproof/tsallissqrtscheduleselfboundingrefinedwindow.lean. lean module compiled","shard":"modules/962b76b0d7bfa6f7.json"},{"id":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","label":"TsallisSqrtScheduleSelfBoundingTuning","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","description":"This module removes the caller-supplied cutoff from the refined generated self-bounding bound. It floors the continuous active-branch threshold and records the large-horizon conditions needed to keep that cutoff positive and inside the finite horizon. Joint optimization of lambda and corruption is left to a downstream scalar leaf.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html","parent":"chapter:tsallis","order":791,"meta":[["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean"],["Declarations","7"],["Project imports","1"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallissqrtscheduleselfboundingtuning banditrlproof.tsallissqrtscheduleselfboundingtuning this module removes the caller-supplied cutoff from the refined generated self-bounding bound. it floors the continuous active-branch threshold and records the large-horizon conditions needed to keep that cutoff positive and inside the finite horizon. joint optimization of lambda and corruption is left to a downstream scalar leaf. lean module compiled","shard":"modules/5744f46d69559543.json"},{"id":"module:BanditRLProof.TsallisTimeVaryingPenalty","label":"TsallisTimeVaryingPenalty","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.TsallisTimeVaryingPenalty","description":"This module formalizes the deterministic learning-rate-change term used by the Tsallis-INF penalty route. It telescopes paper-normalized half-Tsallis potential values across a positive, nonincreasing schedule and retains the negative terminal comparator contribution.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html","parent":"chapter:tsallis","order":792,"meta":[["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean"],["Declarations","14"],["Project imports","4"],["Teaching chapter","Tsallis-FTRL"]],"statement":"","missing":[],"search":"tsallistimevaryingpenalty banditrlproof.tsallistimevaryingpenalty this module formalizes the deterministic learning-rate-change term used by the tsallis-inf penalty route. it telescopes paper-normalized half-tsallis potential values across a positive, nonincreasing schedule and retains the negative terminal comparator contribution. lean module compiled","shard":"modules/c0d31421123d3a7b.json"},{"id":"module:BanditRLProof.UCBSummability","label":"UCBSummability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.UCBSummability","description":"This file records the finite-union/summation layer used by UCB-style bad events: once each arm-time event has a tail bound, the bad-event union over finite arms and a finite horizon is bounded by the corresponding double sum.","url":"../modules/banditrlproof-ucbsummability/index.html","parent":"chapter:ucb","order":793,"meta":[["Source","BanditRLProof/UCBSummability.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","UCB"]],"statement":"","missing":[],"search":"ucbsummability banditrlproof.ucbsummability this file records the finite-union/summation layer used by ucb-style bad events: once each arm-time event has a tail bound, the bad-event union over finite arms and a finite horizon is bounded by the corresponding double sum. lean module compiled","shard":"modules/dde1574221a3812c.json"},{"id":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","label":"UnboundedStoppingTimeL2CoordinateIntegrability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","description":"This module provides the countable-fiber transport needed when an unbounded WithTop Nat stopping time has an L2 round count and the deterministic-time coordinates have a uniform L2 bound. The argument is a measurable equality- fiber decomposition plus Holder; it is not optional stopping.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html","parent":"chapter:probability","order":794,"meta":[["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean"],["Declarations","10"],["Project imports","2"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"unboundedstoppingtimel2coordinateintegrability banditrlproof.unboundedstoppingtimel2coordinateintegrability this module provides the countable-fiber transport needed when an unbounded withtop nat stopping time has an l2 round count and the deterministic-time coordinates have a uniform l2 bound. the argument is a measurable equality- fiber decomposition plus holder; it is not optional stopping. lean module compiled","shard":"modules/b61978be8cbce3c1.json"},{"id":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","label":"UnboundedStoppingTimeWeightedL2CoordinateIntegrability","kind":"Lean module","status":"compiled","subtitle":"BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","description":"This module replaces a moment assumption on an a.e.-finite stopping time by a square-summable deterministic weight. Uniform L2 control of the deterministic coordinates and Cauchy--Schwarz over the measurable stopping fibers then give an integrable weighted stopped value. This is a countable-fiber argument, not optional stopping.","url":"../modules/banditrlproof-unboundedstoppingtimeweightedl2coordinateintegrability/index.html","parent":"chapter:probability","order":795,"meta":[["Source","BanditRLProof/UnboundedStoppingTimeWeightedL2CoordinateIntegrability.lean"],["Declarations","3"],["Project imports","1"],["Teaching chapter","Probability layer"]],"statement":"","missing":[],"search":"unboundedstoppingtimeweightedl2coordinateintegrability banditrlproof.unboundedstoppingtimeweightedl2coordinateintegrability this module replaces a moment assumption on an a.e.-finite stopping time by a square-summable deterministic weight. uniform l2 control of the deterministic coordinates and cauchy--schwarz over the measurable stopping fibers then give an integrable weighted stopped value. this is a countable-fiber argument, not optional stopping. lean module compiled","shard":"modules/ddbe6f306a28ebff.json"},{"id":"declaration:BanditRLProof.ArmStreamPolicy.history","label":"history","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.history","description":"noncomputable def history (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K) (stream : UCB.ArmRewardStream K) : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n | 0 => fun _ => (initial, stream 0 initial) | n + 1 => let h := history initial select stream n let a := select n h History.extendPairHistorySucc h (a, stream (ETC.realHistoryPullCount n h a) a) noncomputable def action (init…","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-23b075b051da","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":0,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def history (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K) (stream : UCB.ArmRewardStream K) : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n | 0 => fun _ => (initial, stream 0 initial) | n + 1 => let h := history initial select stream n let a := select n h History.extendPairHistorySucc h (a, stream (ETC.realHistoryPullCount n h a) a) noncomputable def action (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K) (stream : UCB.ArmRewardStream K) : ActionTrace (Fin K)","missing":[],"search":"history banditrlproof.armstreampolicy.history noncomputable def history (initial : fin k) (select : (n : ℕ) → history.finitepairhistory (fin k) ℝ n → fin k) (stream : ucb.armrewardstream k) : (n : ℕ) → history.finitepairhistory (fin k) ℝ n | 0 => fun _ => (initial, stream 0 initial) | n + 1 => let h := history initial select stream n let a := select n h history.extendpairhistorysucc h (a, stream (etc.realhistorypullcount n h a) a) noncomputable def action (initial : fin k) (select : (n : ℕ) → history.finitepairhistory (fin k) ℝ n → fin k) (stream : ucb.armrewardstream k) : actiontrace (fin k) definition compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.action","label":"action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.action","description":"noncomputable def action (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K) (stream : UCB.ArmRewardStream K) : ActionTrace (Fin K)","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-22d5efc471b7","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":1,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def action (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K) (stream : UCB.ArmRewardStream K) : ActionTrace (Fin K)","missing":[],"search":"action banditrlproof.armstreampolicy.action noncomputable def action (initial : fin k) (select : (n : ℕ) → history.finitepairhistory (fin k) ℝ n → fin k) (stream : ucb.armrewardstream k) : actiontrace (fin k) definition compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.reward","label":"reward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.reward","description":"noncomputable def reward (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K)","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-74823eb90596","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":2,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def reward (initial : Fin K) (select : (n : ℕ) → History.FinitePairHistory (Fin K) ℝ n → Fin K)","missing":[],"search":"reward banditrlproof.armstreampolicy.reward noncomputable def reward (initial : fin k) (select : (n : ℕ) → history.finitepairhistory (fin k) ℝ n → fin k) definition compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.action_zero","label":"action_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.action_zero","description":"@[simp] theorem action_zero (initial : Fin K) (select) (stream) : action initial select stream 0 = initial","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-1731b33c0ebe","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":3,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem action_zero (initial : Fin K) (select) (stream) : action initial select stream 0 = initial","missing":[],"search":"action_zero banditrlproof.armstreampolicy.action_zero @[simp] theorem action_zero (initial : fin k) (select) (stream) : action initial select stream 0 = initial theorem compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.action_succ","label":"action_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.action_succ","description":"@[simp] theorem action_succ (initial : Fin K) (select) (stream) (n : ℕ) : action initial select stream (n+1) = select n (history initial select stream n)","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-2c6cad9ec8f0","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":4,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem action_succ (initial : Fin K) (select) (stream) (n : ℕ) : action initial select stream (n+1) = select n (history initial select stream n)","missing":[],"search":"action_succ banditrlproof.armstreampolicy.action_succ @[simp] theorem action_succ (initial : fin k) (select) (stream) (n : ℕ) : action initial select stream (n+1) = select n (history initial select stream n) theorem compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.history_eq_trace","label":"history_eq_trace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.history_eq_trace","description":"theorem history_eq_trace (initial : Fin K) (select) (stream) (n : ℕ) : history initial select stream n = History.finitePairHistoryOfTrace (action initial select stream) (reward initial select stream) n","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-253f9be62b2e","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":5,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem history_eq_trace (initial : Fin K) (select) (stream) (n : ℕ) : history initial select stream n = History.finitePairHistoryOfTrace (action initial select stream) (reward initial select stream) n","missing":[],"search":"history_eq_trace banditrlproof.armstreampolicy.history_eq_trace theorem history_eq_trace (initial : fin k) (select) (stream) (n : ℕ) : history initial select stream n = history.finitepairhistoryoftrace (action initial select stream) (reward initial select stream) n theorem compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.measurable_history","label":"measurable_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.measurable_history","description":"theorem measurable_history (initial : Fin K) (select) (hm : ∀ n, Measurable (select n)) (n : ℕ) : Measurable (fun stream => history initial select stream n)","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-0f9e40576f67","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":6,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_history (initial : Fin K) (select) (hm : ∀ n, Measurable (select n)) (n : ℕ) : Measurable (fun stream => history initial select stream n)","missing":[],"search":"measurable_history banditrlproof.armstreampolicy.measurable_history theorem measurable_history (initial : fin k) (select) (hm : ∀ n, measurable (select n)) (n : ℕ) : measurable (fun stream => history initial select stream n) theorem compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.measurable_action","label":"measurable_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.measurable_action","description":"theorem measurable_action (initial : Fin K) (select) (hm : ∀ n, Measurable (select n)) (t : ℕ) : Measurable (fun stream => action initial select stream t)","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-1a02a7b92468","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":7,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_action (initial : Fin K) (select) (hm : ∀ n, Measurable (select n)) (t : ℕ) : Measurable (fun stream => action initial select stream t)","missing":[],"search":"measurable_action banditrlproof.armstreampolicy.measurable_action theorem measurable_action (initial : fin k) (select) (hm : ∀ n, measurable (select n)) (t : ℕ) : measurable (fun stream => action initial select stream t) theorem compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ArmStreamPolicy.history_ucb","label":"history_ucb","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ArmStreamPolicy.history_ucb","description":"Compatibility is equality of the actual recursively generated histories.","url":"../modules/banditrlproof-algorithms-armstreampolicy/index.html#decl-e703eacba06e","parent":"module:BanditRLProof.Algorithms.ArmStreamPolicy","order":8,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ArmStreamPolicy"],["Source","BanditRLProof/Algorithms/ArmStreamPolicy.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem history_ucb (hK : 0 < K) (c : ℝ) (stream) (n : ℕ) : history (UCB.initializationArm hK 0) (UCB.realHistoryNextArm hK c) stream n = UCB.armStreamHistory hK c stream n","missing":[],"search":"history_ucb banditrlproof.armstreampolicy.history_ucb compatibility is equality of the actual recursively generated histories. theorem compiled","shard":"modules/8ebdcfdac06a056e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_actual_reward","label":"integrable_actual_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_actual_reward","description":"theorem integrable_actual_reward (n : ℕ) : Integrable (fun Y : ℕ → Round A m => (Y n).2.2.2) (cucbTrajectory S.oracle M.environment)","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-a506dfd630ad","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":9,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_actual_reward (n : ℕ) : Integrable (fun Y : ℕ → Round A m => (Y n).2.2.2) (cucbTrajectory S.oracle M.environment)","missing":[],"search":"integrable_actual_reward banditrlproof.cucb.sourcemodel.integrable_actual_reward theorem integrable_actual_reward (n : ℕ) : integrable (fun y : ℕ → round a m => (y n).2.2.2) (cucbtrajectory s.oracle m.environment) theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.actual_reward_expectation","label":"actual_reward_expectation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.actual_reward_expectation","description":"theorem actual_reward_expectation (n : ℕ) : (∫Y : ℕ → Round A m, (Y n).2.2.2 ∂cucbTrajectory S.oracle M.environment)= ∫Y : ℕ → Round A m, S.score M.trueInput (Y n).1 ∂cucbTrajectory S.oracle M.environment","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-41e7a3da0297","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":10,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem actual_reward_expectation (n : ℕ) : (∫Y : ℕ → Round A m, (Y n).2.2.2 ∂cucbTrajectory S.oracle M.environment)= ∫Y : ℕ → Round A m, S.score M.trueInput (Y n).1 ∂cucbTrajectory S.oracle M.environment","missing":[],"search":"actual_reward_expectation banditrlproof.cucb.sourcemodel.actual_reward_expectation theorem actual_reward_expectation (n : ℕ) : (∫y : ℕ → round a m, (y n).2.2.2 ∂cucbtrajectory s.oracle m.environment)= ∫y : ℕ → round a m, s.score m.trueinput (y n).1 ∂cucbtrajectory s.oracle m.environment theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.cumulative_actual_reward_expectation","label":"cumulative_actual_reward_expectation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.cumulative_actual_reward_expectation","description":"theorem cumulative_actual_reward_expectation (n : ℕ) : (∫Y : ℕ → Round A m, ∑t∈Finset.range n, (Y t).2.2.2 ∂cucbTrajectory S.oracle M.environment)= ∫Y : ℕ → Round A m, ∑t∈Finset.range n, S.score M.trueInput (Y t).1 ∂cucbTrajectory S.oracle M.environment","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-23a623b903f4","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":11,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cumulative_actual_reward_expectation (n : ℕ) : (∫Y : ℕ → Round A m, ∑t∈Finset.range n, (Y t).2.2.2 ∂cucbTrajectory S.oracle M.environment)= ∫Y : ℕ → Round A m, ∑t∈Finset.range n, S.score M.trueInput (Y t).1 ∂cucbTrajectory S.oracle M.environment","missing":[],"search":"cumulative_actual_reward_expectation banditrlproof.cucb.sourcemodel.cumulative_actual_reward_expectation theorem cumulative_actual_reward_expectation (n : ℕ) : (∫y : ℕ → round a m, ∑t∈finset.range n, (y t).2.2.2 ∂cucbtrajectory s.oracle m.environment)= ∫y : ℕ → round a m, ∑t∈finset.range n, s.score m.trueinput (y t).1 ∂cucbtrajectory s.oracle m.environment theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret","label":"approximationRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret","description":"noncomputable def approximationRegret (n : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-8c64c804091b","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":12,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def approximationRegret (n : ℕ) : ℝ","missing":[],"search":"approximationregret banditrlproof.cucb.sourcemodel.approximationregret noncomputable def approximationregret (n : ℕ) : ℝ definition compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_zero","label":"approximationRegret_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_zero","description":"theorem approximationRegret_zero : S.approximationRegret 0=0","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-46aaed59ba1b","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":13,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_zero : S.approximationRegret 0=0","missing":[],"search":"approximationregret_zero banditrlproof.cucb.sourcemodel.approximationregret_zero theorem approximationregret_zero : s.approximationregret 0=0 theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_eq_mean","label":"approximationRegret_eq_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_eq_mean","description":"theorem approximationRegret_eq_mean (n : ℕ) : S.approximationRegret n=(n:ℝ)*S.alpha*S.beta*scoreOptimum S.score M.trueInput- ∫Y : ℕ → Round A m, ∑t∈Finset.range n, S.score M.trueInput (Y t).1 ∂cucbTrajectory S.oracle M.environment","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-fd59f1f2b94e","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":14,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_eq_mean (n : ℕ) : S.approximationRegret n=(n:ℝ)*S.alpha*S.beta*scoreOptimum S.score M.trueInput- ∫Y : ℕ → Round A m, ∑t∈Finset.range n, S.score M.trueInput (Y t).1 ∂cucbTrajectory S.oracle M.environment","missing":[],"search":"approximationregret_eq_mean banditrlproof.cucb.sourcemodel.approximationregret_eq_mean theorem approximationregret_eq_mean (n : ℕ) : s.approximationregret n=(n:ℝ)*s.alpha*s.beta*scoreoptimum s.score m.trueinput- ∫y : ℕ → round a m, ∑t∈finset.range n, s.score m.trueinput (y t).1 ∂cucbtrajectory s.oracle m.environment theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_actual_gap","label":"integrable_actual_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_actual_gap","description":"theorem integrable_actual_gap (t : ℕ) : Integrable (fun Y : ℕ → Round A m => S.gap (Y t).1) (cucbTrajectory S.oracle M.environment)","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-6b845ab95e09","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":15,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_actual_gap (t : ℕ) : Integrable (fun Y : ℕ → Round A m => S.gap (Y t).1) (cucbTrajectory S.oracle M.environment)","missing":[],"search":"integrable_actual_gap banditrlproof.cucb.sourcemodel.integrable_actual_gap theorem integrable_actual_gap (t : ℕ) : integrable (fun y : ℕ → round a m => s.gap (y t).1) (cucbtrajectory s.oracle m.environment) theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_eq_gap_sum","label":"approximationRegret_eq_gap_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_eq_gap_sum","description":"theorem approximationRegret_eq_gap_sum (n : ℕ) : S.approximationRegret n = (∫Y : ℕ → Round A m, ∑t∈Finset.range n, S.gap (Y t).1 ∂cucbTrajectory S.oracle M.environment)- (n:ℝ)*S.alpha*(1-S.beta)*scoreOptimum S.score M.trueInput","url":"../modules/banditrlproof-algorithms-cucbactualreward/index.html#decl-d819a7b4527f","parent":"module:BanditRLProof.Algorithms.CUCBActualReward","order":16,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBActualReward"],["Source","BanditRLProof/Algorithms/CUCBActualReward.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem approximationRegret_eq_gap_sum (n : ℕ) : S.approximationRegret n = (∫Y : ℕ → Round A m, ∑t∈Finset.range n, S.gap (Y t).1 ∂cucbTrajectory S.oracle M.environment)- (n:ℝ)*S.alpha*(1-S.beta)*scoreOptimum S.score M.trueInput","missing":[],"search":"approximationregret_eq_gap_sum banditrlproof.cucb.sourcemodel.approximationregret_eq_gap_sum theorem approximationregret_eq_gap_sum (n : ℕ) : s.approximationregret n = (∫y : ℕ → round a m, ∑t∈finset.range n, s.gap (y t).1 ∂cucbtrajectory s.oracle m.environment)- (n:ℝ)*s.alpha*(1-s.beta)*scoreoptimum s.score m.trueinput theorem compiled","shard":"modules/2c93a6e6b0be6d62.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.ChargeData","label":"ChargeData","kind":"structure","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData","description":"Fixed inputs to the analysis rule. Full source model obligations are separate.","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-35964c6bddea","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":17,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure ChargeData (A : Type*) (m : ℕ) where","missing":[],"search":"chargedata banditrlproof.cucb.chargedata fixed inputs to the analysis rule. full source model obligations are separate. structure compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.choose","label":"choose","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.choose","description":"noncomputable def choose (N : Fin m → ℕ) (a : A) : Option (Fin m)","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-bdadd10c5df8","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":18,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def choose (N : Fin m → ℕ) (a : A) : Option (Fin m)","missing":[],"search":"choose banditrlproof.cucb.chargedata.choose noncomputable def choose (n : fin m → ℕ) (a : a) : option (fin m) definition compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters","label":"counters","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters","description":"noncomputable def counters (actions : ℕ → A) : ℕ → Fin m → ℕ | 0 => fun _ => 0 | n+1 => fun i => counters actions n i + if C.choose (counters actions n) (actions n) = some i then 1 else 0 theorem counters_zero (actions : ℕ → A) (i : Fin m) : C.counters actions 0 i=0","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-d5540dd231a9","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":19,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def counters (actions : ℕ → A) : ℕ → Fin m → ℕ | 0 => fun _ => 0 | n+1 => fun i => counters actions n i + if C.choose (counters actions n) (actions n) = some i then 1 else 0 theorem counters_zero (actions : ℕ → A) (i : Fin m) : C.counters actions 0 i=0","missing":[],"search":"counters banditrlproof.cucb.chargedata.counters noncomputable def counters (actions : ℕ → a) : ℕ → fin m → ℕ | 0 => fun _ => 0 | n+1 => fun i => counters actions n i + if c.choose (counters actions n) (actions n) = some i then 1 else 0 theorem counters_zero (actions : ℕ → a) (i : fin m) : c.counters actions 0 i=0 definition compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_zero","label":"counters_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_zero","description":"theorem counters_zero (actions : ℕ → A) (i : Fin m) : C.counters actions 0 i=0","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-4838aa06e43b","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":20,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_zero (actions : ℕ → A) (i : Fin m) : C.counters actions 0 i=0","missing":[],"search":"counters_zero banditrlproof.cucb.chargedata.counters_zero theorem counters_zero (actions : ℕ → a) (i : fin m) : c.counters actions 0 i=0 theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_succ","label":"counters_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_succ","description":"theorem counters_succ (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions (n+1) i=C.counters actions n i+ if C.choose (C.counters actions n) (actions n)=some i then 1 else 0","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-4c29be8323cd","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":21,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_succ (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions (n+1) i=C.counters actions n i+ if C.choose (C.counters actions n) (actions n)=some i then 1 else 0","missing":[],"search":"counters_succ banditrlproof.cucb.chargedata.counters_succ theorem counters_succ (actions : ℕ → a) (n : ℕ) (i : fin m) : c.counters actions (n+1) i=c.counters actions n i+ if c.choose (c.counters actions n) (actions n)=some i then 1 else 0 theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.choose_mem","label":"choose_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.choose_mem","description":"theorem choose_mem (N : Fin m → ℕ) (a : A) (i : Fin m) (h : C.choose N a=some i) : C.bad a=true ∧ i∈C.triggers a","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-b64166ed4109","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":22,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem choose_mem (N : Fin m → ℕ) (a : A) (i : Fin m) (h : C.choose N a=some i) : C.bad a=true ∧ i∈C.triggers a","missing":[],"search":"choose_mem banditrlproof.cucb.chargedata.choose_mem theorem choose_mem (n : fin m → ℕ) (a : a) (i : fin m) (h : c.choose n a=some i) : c.bad a=true ∧ i∈c.triggers a theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.choose_sufficient","label":"choose_sufficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.choose_sufficient","description":"theorem choose_sufficient (N : Fin m → ℕ) (a : A) (i : Fin m) (h : C.choose N a=some i) (hu : 0<C.inverseGap a) (hp : ∀j∈C.triggers a, 0<C.triggerLower j) (n : ℕ) (hi : samplingThreshold n (C.inverseGap a) (C.triggerLower i)<N i) : ∀j∈C.triggers a, samplingThreshold n (C.inverseGap a) (C.triggerLower j)<N j","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-44d2329c0f00","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":23,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem choose_sufficient (N : Fin m → ℕ) (a : A) (i : Fin m) (h : C.choose N a=some i) (hu : 0<C.inverseGap a) (hp : ∀j∈C.triggers a, 0<C.triggerLower j) (n : ℕ) (hi : samplingThreshold n (C.inverseGap a) (C.triggerLower i)<N i) : ∀j∈C.triggers a, samplingThreshold n (C.inverseGap a) (C.triggerLower j)<N j","missing":[],"search":"choose_sufficient banditrlproof.cucb.chargedata.choose_sufficient theorem choose_sufficient (n : fin m → ℕ) (a : a) (i : fin m) (h : c.choose n a=some i) (hu : 0<c.inversegap a) (hp : ∀j∈c.triggers a, 0<c.triggerlower j) (n : ℕ) (hi : samplingthreshold n (c.inversegap a) (c.triggerlower i)<n i) : ∀j∈c.triggers a, samplingthreshold n (c.inversegap a) (c.triggerlower j)<n j theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_le_time","label":"counters_le_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_le_time","description":"theorem counters_le_time (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions n i≤n","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-22fa1f004ce9","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":24,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_le_time (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions n i≤n","missing":[],"search":"counters_le_time banditrlproof.cucb.chargedata.counters_le_time theorem counters_le_time (actions : ℕ → a) (n : ℕ) (i : fin m) : c.counters actions n i≤n theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_mono_step","label":"counters_mono_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_mono_step","description":"theorem counters_mono_step (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions n i≤C.counters actions (n+1) i","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-b7db2db06928","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":25,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_mono_step (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions n i≤C.counters actions (n+1) i","missing":[],"search":"counters_mono_step banditrlproof.cucb.chargedata.counters_mono_step theorem counters_mono_step (actions : ℕ → a) (n : ℕ) (i : fin m) : c.counters actions n i≤c.counters actions (n+1) i theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_causal","label":"counters_causal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_causal","description":"theorem counters_causal (actions actions' : ℕ → A) (n : ℕ) (h : ∀t<n, actions t=actions' t) : C.counters actions n=C.counters actions' n","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-5b7e2679a186","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":26,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_causal (actions actions' : ℕ → A) (n : ℕ) (h : ∀t<n, actions t=actions' t) : C.counters actions n=C.counters actions' n","missing":[],"search":"counters_causal banditrlproof.cucb.chargedata.counters_causal theorem counters_causal (actions actions' : ℕ → a) (n : ℕ) (h : ∀t<n, actions t=actions' t) : c.counters actions n=c.counters actions' n theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.charge_before_feedback","label":"charge_before_feedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.charge_before_feedback","description":"The charge may use the current action but not its sampled feedback.","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-0947a20b9fd3","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":27,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem charge_before_feedback (Y Z : ℕ → Round A m) (n : ℕ) (hpast : ∀t<n, (Y t).1=(Z t).1) (hcurrent : (Y n).1=(Z n).1) : C.choose (C.counters (fun t => (Y t).1) n) (Y n).1 = C.choose (C.counters (fun t => (Z t).1) n) (Z n).1","missing":[],"search":"charge_before_feedback banditrlproof.cucb.chargedata.charge_before_feedback the charge may use the current action but not its sampled feedback. theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_sum","label":"counters_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_sum","description":"theorem counters_sum (actions : ℕ → A) (n : ℕ) : ∑i:Fin m, C.counters actions n i = ∑t∈Finset.range n, if C.bad (actions t) then 1 else 0","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-989fbfd03ec1","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":28,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:87"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_sum (actions : ℕ → A) (n : ℕ) : ∑i:Fin m, C.counters actions n i = ∑t∈Finset.range n, if C.bad (actions t) then 1 else 0","missing":[],"search":"counters_sum banditrlproof.cucb.chargedata.counters_sum theorem counters_sum (actions : ℕ → a) (n : ℕ) : ∑i:fin m, c.counters actions n i = ∑t∈finset.range n, if c.bad (actions t) then 1 else 0 theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_eq_sum","label":"counters_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_eq_sum","description":"theorem counters_eq_sum (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions n i = ∑t∈Finset.range n, if C.choose (C.counters actions t) (actions t)=some i then 1 else 0","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-eb1d669359cd","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":29,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_eq_sum (actions : ℕ → A) (n : ℕ) (i : Fin m) : C.counters actions n i = ∑t∈Finset.range n, if C.choose (C.counters actions t) (actions t)=some i then 1 else 0","missing":[],"search":"counters_eq_sum banditrlproof.cucb.chargedata.counters_eq_sum theorem counters_eq_sum (actions : ℕ → a) (n : ℕ) (i : fin m) : c.counters actions n i = ∑t∈finset.range n, if c.choose (c.counters actions t) (actions t)=some i then 1 else 0 theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargedObservations","label":"chargedObservations","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargedObservations","description":"noncomputable def chargedObservations (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : ℕ","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-c3a0f453e89f","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":30,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:106"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def chargedObservations (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : ℕ","missing":[],"search":"chargedobservations banditrlproof.cucb.chargedata.chargedobservations noncomputable def chargedobservations (y : ℕ → round a m) (n : ℕ) (i : fin m) : ℕ definition compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargedObservations_le_observationCount","label":"chargedObservations_le_observationCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargedObservations_le_observationCount","description":"theorem chargedObservations_le_observationCount (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : C.chargedObservations Y n i≤observationCount (fun t => (Y t).2) n i","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-8c94bb3e639b","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":31,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chargedObservations_le_observationCount (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : C.chargedObservations Y n i≤observationCount (fun t => (Y t).2) n i","missing":[],"search":"chargedobservations_le_observationcount banditrlproof.cucb.chargedata.chargedobservations_le_observationcount theorem chargedobservations_le_observationcount (y : ℕ → round a m) (n : ℕ) (i : fin m) : c.chargedobservations y n i≤observationcount (fun t => (y t).2) n i theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargedObservations_le_counters","label":"chargedObservations_le_counters","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargedObservations_le_counters","description":"theorem chargedObservations_le_counters (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : C.chargedObservations Y n i≤C.counters (fun t => (Y t).1) n i","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-a0a8496c6b70","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":32,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:116"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chargedObservations_le_counters (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : C.chargedObservations Y n i≤C.counters (fun t => (Y t).1) n i","missing":[],"search":"chargedobservations_le_counters banditrlproof.cucb.chargedata.chargedobservations_le_counters theorem chargedobservations_le_counters (y : ℕ → round a m) (n : ℕ) (i : fin m) : c.chargedobservations y n i≤c.counters (fun t => (y t).1) n i theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_le_observations_of_always_triggered","label":"counters_le_observations_of_always_triggered","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_le_observations_of_always_triggered","description":"Deterministic triggering is recovered without a probabilistic tail bound.","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-1862b866206e","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":33,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_le_observations_of_always_triggered (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) (h : ∀t<n, C.choose (C.counters (fun s => (Y s).1) t) (Y t).1=some i → (Y t).2.1 i=true) : C.counters (fun t => (Y t).1) n i≤observationCount (fun t => (Y t).2) n i","missing":[],"search":"counters_le_observations_of_always_triggered banditrlproof.cucb.chargedata.counters_le_observations_of_always_triggered deterministic triggering is recovered without a probabilistic tail bound. theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_choose","label":"measurable_choose","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_choose","description":"theorem measurable_choose : Measurable (fun p : (Fin m → ℕ) × A => C.choose p.1 p.2)","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-71f1c9ed1fb0","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":34,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:137"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_choose : Measurable (fun p : (Fin m → ℕ) × A => C.choose p.1 p.2)","missing":[],"search":"measurable_choose banditrlproof.cucb.chargedata.measurable_choose theorem measurable_choose : measurable (fun p : (fin m → ℕ) × a => c.choose p.1 p.2) theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_counters","label":"measurable_counters","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_counters","description":"theorem measurable_counters (n : ℕ) : Measurable (fun actions : ℕ → A => C.counters actions n)","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-65f0a320982e","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":35,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:140"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_counters (n : ℕ) : Measurable (fun actions : ℕ → A => C.counters actions n)","missing":[],"search":"measurable_counters banditrlproof.cucb.chargedata.measurable_counters theorem measurable_counters (n : ℕ) : measurable (fun actions : ℕ → a => c.counters actions n) theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_path_charge","label":"measurable_path_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_path_charge","description":"theorem measurable_path_charge (n : ℕ) : Measurable (fun Y : ℕ → Round A m => C.choose (C.counters (fun t => (Y t).1) n) (Y n).1)","url":"../modules/banditrlproof-algorithms-cucbcharge/index.html#decl-b6be7b135537","parent":"module:BanditRLProof.Algorithms.CUCBCharge","order":36,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBCharge"],["Source","BanditRLProof/Algorithms/CUCBCharge.lean:150"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_path_charge (n : ℕ) : Measurable (fun Y : ℕ → Round A m => C.choose (C.counters (fun t => (Y t).1) n) (Y n).1)","missing":[],"search":"measurable_path_charge banditrlproof.cucb.chargedata.measurable_path_charge theorem measurable_path_charge (n : ℕ) : measurable (fun y : ℕ → round a m => c.choose (c.counters (fun t => (y t).1) n) (y n).1) theorem compiled","shard":"modules/5d47e3a0553cec38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargeValue","label":"chargeValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargeValue","description":"noncomputable def chargeValue (i : Fin m) (N : Fin m → ℕ) (z : Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-8b090d9f9b76","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":37,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def chargeValue (i : Fin m) (N : Fin m → ℕ) (z : Round A m) : ℝ","missing":[],"search":"chargevalue banditrlproof.cucb.chargedata.chargevalue noncomputable def chargevalue (i : fin m) (n : fin m → ℕ) (z : round a m) : ℝ definition compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.successValue","label":"successValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.successValue","description":"noncomputable def successValue (i : Fin m) (N : Fin m → ℕ) (z : Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-b934fcbfa884","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":38,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def successValue (i : Fin m) (N : Fin m → ℕ) (z : Round A m) : ℝ","missing":[],"search":"successvalue banditrlproof.cucb.chargedata.successvalue noncomputable def successvalue (i : fin m) (n : fin m → ℕ) (z : round a m) : ℝ definition compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.compensation","label":"compensation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.compensation","description":"noncomputable def compensation (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-f59fb63dc727","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":39,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def compensation (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : ℝ","missing":[],"search":"compensation banditrlproof.cucb.chargedata.compensation noncomputable def compensation (i : fin m) (tilt : ℝ) (n : fin m → ℕ) (z : round a m) : ℝ definition compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.exp_compensation","label":"exp_compensation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.exp_compensation","description":"theorem exp_compensation (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : Real.exp (C.compensation i tilt N z)=C.chargedTriggerFactor i tilt N z","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-cb2ed15184fa","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":40,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_compensation (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : Real.exp (C.compensation i tilt N z)=C.chargedTriggerFactor i tilt N z","missing":[],"search":"exp_compensation banditrlproof.cucb.chargedata.exp_compensation theorem exp_compensation (i : fin m) (tilt : ℝ) (n : fin m → ℕ) (z : round a m) : real.exp (c.compensation i tilt n z)=c.chargedtriggerfactor i tilt n z theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.compensation_abs_le","label":"compensation_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.compensation_abs_le","description":"theorem compensation_abs_le (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : |C.compensation i tilt N z|≤|(1-Real.exp (-tilt))*C.triggerLower i|+|tilt|","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-40ccf431e6c4","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":41,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem compensation_abs_le (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : |C.compensation i tilt N z|≤|(1-Real.exp (-tilt))*C.triggerLower i|+|tilt|","missing":[],"search":"compensation_abs_le banditrlproof.cucb.chargedata.compensation_abs_le theorem compensation_abs_le (i : fin m) (tilt : ℝ) (n : fin m → ℕ) (z : round a m) : |c.compensation i tilt n z|≤|(1-real.exp (-tilt))*c.triggerlower i|+|tilt| theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.sum_chargeValue","label":"sum_chargeValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.sum_chargeValue","description":"theorem sum_chargeValue (actions : ℕ → Round A m) (n : ℕ) (i : Fin m) : (∑t∈Finset.range n, C.chargeValue i (C.counters (fun s => (actions s).1) t) (actions t)) = (C.counters (fun t => (actions t).1) n i : ℝ)","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-8e3e2faa344b","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":42,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_chargeValue (actions : ℕ → Round A m) (n : ℕ) (i : Fin m) : (∑t∈Finset.range n, C.chargeValue i (C.counters (fun s => (actions s).1) t) (actions t)) = (C.counters (fun t => (actions t).1) n i : ℝ)","missing":[],"search":"sum_chargevalue banditrlproof.cucb.chargedata.sum_chargevalue theorem sum_chargevalue (actions : ℕ → round a m) (n : ℕ) (i : fin m) : (∑t∈finset.range n, c.chargevalue i (c.counters (fun s => (actions s).1) t) (actions t)) = (c.counters (fun t => (actions t).1) n i : ℝ) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.sum_successValue","label":"sum_successValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.sum_successValue","description":"theorem sum_successValue (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : (∑t∈Finset.range n, C.successValue i (C.counters (fun s => (Y s).1) t) (Y t)) = (C.chargedObservations Y n i : ℝ)","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-19dc067343a7","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":43,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_successValue (Y : ℕ → Round A m) (n : ℕ) (i : Fin m) : (∑t∈Finset.range n, C.successValue i (C.counters (fun s => (Y s).1) t) (Y t)) = (C.chargedObservations Y n i : ℝ)","missing":[],"search":"sum_successvalue banditrlproof.cucb.chargedata.sum_successvalue theorem sum_successvalue (y : ℕ → round a m) (n : ℕ) (i : fin m) : (∑t∈finset.range n, c.successvalue i (c.counters (fun s => (y s).1) t) (y t)) = (c.chargedobservations y n i : ℝ) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_compensation","label":"measurable_compensation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_compensation","description":"theorem measurable_compensation (i : Fin m) (tilt : ℝ) : Measurable (fun p : (Fin m → ℕ) × Round A m => C.compensation i tilt p.1 p.2)","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-0833b14eea11","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":44,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_compensation (i : Fin m) (tilt : ℝ) : Measurable (fun p : (Fin m → ℕ) × Round A m => C.compensation i tilt p.1 p.2)","missing":[],"search":"measurable_compensation banditrlproof.cucb.chargedata.measurable_compensation theorem measurable_compensation (i : fin m) (tilt : ℝ) : measurable (fun p : (fin m → ℕ) × round a m => c.compensation i tilt p.1 p.2) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.integrable_exp_compensation","label":"integrable_exp_compensation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.integrable_exp_compensation","description":"theorem integrable_exp_compensation {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (i : Fin m) (tilt scale : ℝ) (g : Ω → (Fin m → ℕ) × Round A m) (hg : Measurable g) : Integrable (fun ω => Real.exp (scale*C.compensation i tilt (g ω).1 (g ω).2)) ν","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-40287f43ab09","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":45,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_compensation {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (i : Fin m) (tilt scale : ℝ) (g : Ω → (Fin m → ℕ) × Round A m) (hg : Measurable g) : Integrable (fun ω => Real.exp (scale*C.compensation i tilt (g ω).1 (g ω).2)) ν","missing":[],"search":"integrable_exp_compensation banditrlproof.cucb.chargedata.integrable_exp_compensation theorem integrable_exp_compensation {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] (i : fin m) (tilt scale : ℝ) (g : ω → (fin m → ℕ) × round a m) (hg : measurable g) : integrable (fun ω => real.exp (scale*c.compensation i tilt (g ω).1 (g ω).2)) ν theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_counters_piLE","label":"measurable_counters_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_counters_piLE","description":"theorem measurable_counters_piLE (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → Round A m => C.counters (fun t => (Y t).1) n)","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-aae5b5b4661d","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":46,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_counters_piLE (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → Round A m => C.counters (fun t => (Y t).1) n)","missing":[],"search":"measurable_counters_pile banditrlproof.cucb.chargedata.measurable_counters_pile theorem measurable_counters_pile (n : ℕ) : measurable[filtration.pile n] (fun y : ℕ → round a m => c.counters (fun t => (y t).1) n) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.compensation_adapted","label":"compensation_adapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.compensation_adapted","description":"theorem compensation_adapted (i : Fin m) (tilt : ℝ) : StronglyAdapted Filtration.piLE (fun n (Y : ℕ → Round A m) => C.compensation i tilt (C.counters (fun t => (Y t).1) n) (Y n))","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-ea1870e680ee","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":47,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:100"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem compensation_adapted (i : Fin m) (tilt : ℝ) : StronglyAdapted Filtration.piLE (fun n (Y : ℕ → Round A m) => C.compensation i tilt (C.counters (fun t => (Y t).1) n) (Y n))","missing":[],"search":"compensation_adapted banditrlproof.cucb.chargedata.compensation_adapted theorem compensation_adapted (i : fin m) (tilt : ℝ) : stronglyadapted filtration.pile (fun n (y : ℕ → round a m) => c.compensation i tilt (c.counters (fun t => (y t).1) n) (y n)) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.charged_successor_condMGF","label":"charged_successor_condMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.charged_successor_condMGF","description":"theorem charged_successor_condMGF (n : ℕ) (tilt : ℝ) (htilt : 0≤tilt) : Concentration.HasCondMGFUpperBoundAt (MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance) (Preorder.measurable_frestrictLe n).comap_le (fun Y => C.compensation i tilt (C.counters (fun t => (Y t).1) (n+1)) (Y (n+1))) 1 0 (cucbTrajectory oracle environment)","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-fdfcb45b3d9b","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":48,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:114"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem charged_successor_condMGF (n : ℕ) (tilt : ℝ) (htilt : 0≤tilt) : Concentration.HasCondMGFUpperBoundAt (MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance) (Preorder.measurable_frestrictLe n).comap_le (fun Y => C.compensation i tilt (C.counters (fun t => (Y t).1) (n+1)) (Y (n+1))) 1 0 (cucbTrajectory oracle environment)","missing":[],"search":"charged_successor_condmgf banditrlproof.cucb.chargedata.charged_successor_condmgf theorem charged_successor_condmgf (n : ℕ) (tilt : ℝ) (htilt : 0≤tilt) : concentration.hascondmgfupperboundat (measurablespace.comap (preorder.frestrictle n) inferinstance) (preorder.measurable_frestrictle n).comap_le (fun y => c.compensation i tilt (c.counters (fun t => (y t).1) (n+1)) (y (n+1))) 1 0 (cucbtrajectory oracle environment) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.charged_initial_MGF","label":"charged_initial_MGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.charged_initial_MGF","description":"theorem charged_initial_MGF (tilt : ℝ) (htilt : 0≤tilt) : Concentration.HasMGFUpperBoundAt (fun Y => C.compensation i tilt (C.counters (fun t => (Y t).1) 0) (Y 0)) 1 0 (cucbTrajectory oracle environment)","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-43e46edc7d24","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":49,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem charged_initial_MGF (tilt : ℝ) (htilt : 0≤tilt) : Concentration.HasMGFUpperBoundAt (fun Y => C.compensation i tilt (C.counters (fun t => (Y t).1) 0) (Y 0)) 1 0 (cucbTrajectory oracle environment)","missing":[],"search":"charged_initial_mgf banditrlproof.cucb.chargedata.charged_initial_mgf theorem charged_initial_mgf (tilt : ℝ) (htilt : 0≤tilt) : concentration.hasmgfupperboundat (fun y => c.compensation i tilt (c.counters (fun t => (y t).1) 0) (y 0)) 1 0 (cucbtrajectory oracle environment) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.charged_count_tail","label":"charged_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.charged_count_tail","description":"theorem charged_count_tail (n : ℕ) (tilt k budget : ℝ) (htilt : 0≤tilt) (hp : 0≤C.triggerLower i) : (cucbTrajectory oracle environment) {Y | k≤(C.counters (fun t => (Y t).1) n i : ℝ) ∧ (C.chargedObservations Y n i : ℝ)≤budget} ≤ ENNReal.ofReal (Real.exp (-((1-Real.exp (-tilt))*C.triggerLower i)*k+tilt*budget))","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-a1cf7a919056","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":50,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem charged_count_tail (n : ℕ) (tilt k budget : ℝ) (htilt : 0≤tilt) (hp : 0≤C.triggerLower i) : (cucbTrajectory oracle environment) {Y | k≤(C.counters (fun t => (Y t).1) n i : ℝ) ∧ (C.chargedObservations Y n i : ℝ)≤budget} ≤ ENNReal.ofReal (Real.exp (-((1-Real.exp (-tilt))*C.triggerLower i)*k+tilt*budget))","missing":[],"search":"charged_count_tail banditrlproof.cucb.chargedata.charged_count_tail theorem charged_count_tail (n : ℕ) (tilt k budget : ℝ) (htilt : 0≤tilt) (hp : 0≤c.triggerlower i) : (cucbtrajectory oracle environment) {y | k≤(c.counters (fun t => (y t).1) n i : ℝ) ∧ (c.chargedobservations y n i : ℝ)≤budget} ≤ ennreal.ofreal (real.exp (-((1-real.exp (-tilt))*c.triggerlower i)*k+tilt*budget)) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.observation_below_charged_half","label":"observation_below_charged_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.observation_below_charged_half","description":"theorem observation_below_charged_half (n : ℕ) (k : ℝ) (hk : 0≤k) (hp : 0≤C.triggerLower i) : (cucbTrajectory oracle environment) {Y | k≤(C.counters (fun t => (Y t).1) n i : ℝ) ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤k*C.triggerLower i/2} ≤ ENNReal.ofReal (Real.exp (-k*C.triggerLower i/8))","url":"../modules/banditrlproof-algorithms-cucbchargedconcentration/index.html#decl-a6b1ed1b9b18","parent":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","order":51,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConcentration"],["Source","BanditRLProof/Algorithms/CUCBChargedConcentration.lean:166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observation_below_charged_half (n : ℕ) (k : ℝ) (hk : 0≤k) (hp : 0≤C.triggerLower i) : (cucbTrajectory oracle environment) {Y | k≤(C.counters (fun t => (Y t).1) n i : ℝ) ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤k*C.triggerLower i/2} ≤ ENNReal.ofReal (Real.exp (-k*C.triggerLower i/8))","missing":[],"search":"observation_below_charged_half banditrlproof.cucb.chargedata.observation_below_charged_half theorem observation_below_charged_half (n : ℕ) (k : ℝ) (hk : 0≤k) (hp : 0≤c.triggerlower i) : (cucbtrajectory oracle environment) {y | k≤(c.counters (fun t => (y t).1) n i : ℝ) ∧ (observationcount (fun t => (y t).2) n i : ℝ)≤k*c.triggerlower i/2} ≤ ennreal.ofreal (real.exp (-k*c.triggerlower i/8)) theorem compiled","shard":"modules/076a5861e0ca3f9f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.prefixActions","label":"prefixActions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.prefixActions","description":"noncomputable def prefixActions (n : ℕ) (h : (i : Finset.Iic n) → Round A m) : ℕ → A","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-1e7c39ed4461","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":52,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def prefixActions (n : ℕ) (h : (i : Finset.Iic n) → Round A m) : ℕ → A","missing":[],"search":"prefixactions banditrlproof.cucb.chargedata.prefixactions noncomputable def prefixactions (n : ℕ) (h : (i : finset.iic n) → round a m) : ℕ → a definition compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.prefixCounters","label":"prefixCounters","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.prefixCounters","description":"noncomputable def prefixCounters (n : ℕ) (h : (i : Finset.Iic n) → Round A m) : Fin m → ℕ","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-73b09ad7ce5b","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":53,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def prefixCounters (n : ℕ) (h : (i : Finset.Iic n) → Round A m) : Fin m → ℕ","missing":[],"search":"prefixcounters banditrlproof.cucb.chargedata.prefixcounters noncomputable def prefixcounters (n : ℕ) (h : (i : finset.iic n) → round a m) : fin m → ℕ definition compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.prefixCounters_eq","label":"prefixCounters_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.prefixCounters_eq","description":"theorem prefixCounters_eq (n : ℕ) (Y : ℕ → Round A m) : C.prefixCounters n (Preorder.frestrictLe n Y) = C.counters (fun t => (Y t).1) (n+1)","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-e6b30b85aef9","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":54,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem prefixCounters_eq (n : ℕ) (Y : ℕ → Round A m) : C.prefixCounters n (Preorder.frestrictLe n Y) = C.counters (fun t => (Y t).1) (n+1)","missing":[],"search":"prefixcounters_eq banditrlproof.cucb.chargedata.prefixcounters_eq theorem prefixcounters_eq (n : ℕ) (y : ℕ → round a m) : c.prefixcounters n (preorder.frestrictle n y) = c.counters (fun t => (y t).1) (n+1) theorem compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_prefixActions","label":"measurable_prefixActions","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_prefixActions","description":"theorem measurable_prefixActions (n : ℕ) : Measurable (prefixActions (A:=A) (m:=m) n)","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-75181f7ac81a","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":55,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_prefixActions (n : ℕ) : Measurable (prefixActions (A:=A) (m:=m) n)","missing":[],"search":"measurable_prefixactions banditrlproof.cucb.chargedata.measurable_prefixactions theorem measurable_prefixactions (n : ℕ) : measurable (prefixactions (a:=a) (m:=m) n) theorem compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_prefixCounters","label":"measurable_prefixCounters","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_prefixCounters","description":"theorem measurable_prefixCounters (n : ℕ) : Measurable (C.prefixCounters n)","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-43f4c3f72358","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":56,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_prefixCounters (n : ℕ) : Measurable (C.prefixCounters n)","missing":[],"search":"measurable_prefixcounters banditrlproof.cucb.chargedata.measurable_prefixcounters theorem measurable_prefixcounters (n : ℕ) : measurable (c.prefixcounters n) theorem compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.cucb_condExp_chargedTriggerFactor","label":"cucb_condExp_chargedTriggerFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.cucb_condExp_chargedTriggerFactor","description":"theorem cucb_condExp_chargedTriggerFactor (n : ℕ) (tilt : ℝ) (htilt : 0≤tilt) : (cucbTrajectory oracle environment)[fun Y => C.chargedTriggerFactor i tilt (C.counters (fun t => (Y t).1) (n+1)) (Y (n+1)) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance] ≤ᵐ[cucbTrajectory oracle environment] fun _ => 1","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-6e7298688bdf","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":57,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucb_condExp_chargedTriggerFactor (n : ℕ) (tilt : ℝ) (htilt : 0≤tilt) : (cucbTrajectory oracle environment)[fun Y => C.chargedTriggerFactor i tilt (C.counters (fun t => (Y t).1) (n+1)) (Y (n+1)) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance] ≤ᵐ[cucbTrajectory oracle environment] fun _ => 1","missing":[],"search":"cucb_condexp_chargedtriggerfactor banditrlproof.cucb.chargedata.cucb_condexp_chargedtriggerfactor theorem cucb_condexp_chargedtriggerfactor (n : ℕ) (tilt : ℝ) (htilt : 0≤tilt) : (cucbtrajectory oracle environment)[fun y => c.chargedtriggerfactor i tilt (c.counters (fun t => (y t).1) (n+1)) (y (n+1)) | measurablespace.comap (preorder.frestrictle n) inferinstance] ≤ᵐ[cucbtrajectory oracle environment] fun _ => 1 theorem compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.cucb_initial_chargedTriggerFactor","label":"cucb_initial_chargedTriggerFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.cucb_initial_chargedTriggerFactor","description":"theorem cucb_initial_chargedTriggerFactor (tilt : ℝ) (htilt : 0≤tilt) : (∫Y, C.chargedTriggerFactor i tilt (C.counters (fun t => (Y t).1) 0) (Y 0) ∂cucbTrajectory oracle environment)≤1","url":"../modules/banditrlproof-algorithms-cucbchargedconditional/index.html#decl-a0b4c7e11969","parent":"module:BanditRLProof.Algorithms.CUCBChargedConditional","order":58,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedConditional"],["Source","BanditRLProof/Algorithms/CUCBChargedConditional.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucb_initial_chargedTriggerFactor (tilt : ℝ) (htilt : 0≤tilt) : (∫Y, C.chargedTriggerFactor i tilt (C.counters (fun t => (Y t).1) 0) (Y 0) ∂cucbTrajectory oracle environment)≤1","missing":[],"search":"cucb_initial_chargedtriggerfactor banditrlproof.cucb.chargedata.cucb_initial_chargedtriggerfactor theorem cucb_initial_chargedtriggerfactor (tilt : ℝ) (htilt : 0≤tilt) : (∫y, c.chargedtriggerfactor i tilt (c.counters (fun t => (y t).1) 0) (y 0) ∂cucbtrajectory oracle environment)≤1 theorem compiled","shard":"modules/7f9b7b51d7849e54.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargedTriggerFactor","label":"chargedTriggerFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargedTriggerFactor","description":"noncomputable def chargedTriggerFactor (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html#decl-684cbfea34df","parent":"module:BanditRLProof.Algorithms.CUCBChargedMGF","order":59,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBChargedMGF"],["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def chargedTriggerFactor (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : ℝ","missing":[],"search":"chargedtriggerfactor banditrlproof.cucb.chargedata.chargedtriggerfactor noncomputable def chargedtriggerfactor (i : fin m) (tilt : ℝ) (n : fin m → ℕ) (z : round a m) : ℝ definition compiled","shard":"modules/68465bba9f63c3af.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargedTriggerFactor_nonneg","label":"chargedTriggerFactor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargedTriggerFactor_nonneg","description":"theorem chargedTriggerFactor_nonneg (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : 0≤C.chargedTriggerFactor i tilt N z","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html#decl-9ca8d29f3218","parent":"module:BanditRLProof.Algorithms.CUCBChargedMGF","order":60,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedMGF"],["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chargedTriggerFactor_nonneg (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : 0≤C.chargedTriggerFactor i tilt N z","missing":[],"search":"chargedtriggerfactor_nonneg banditrlproof.cucb.chargedata.chargedtriggerfactor_nonneg theorem chargedtriggerfactor_nonneg (i : fin m) (tilt : ℝ) (n : fin m → ℕ) (z : round a m) : 0≤c.chargedtriggerfactor i tilt n z theorem compiled","shard":"modules/68465bba9f63c3af.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.chargedTriggerFactor_le","label":"chargedTriggerFactor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.chargedTriggerFactor_le","description":"theorem chargedTriggerFactor_le (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : C.chargedTriggerFactor i tilt N z ≤ Real.exp (|tilt|+|(1-Real.exp (-tilt))*C.triggerLower i|)","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html#decl-520575a6b0bc","parent":"module:BanditRLProof.Algorithms.CUCBChargedMGF","order":61,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedMGF"],["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chargedTriggerFactor_le (i : Fin m) (tilt : ℝ) (N : Fin m → ℕ) (z : Round A m) : C.chargedTriggerFactor i tilt N z ≤ Real.exp (|tilt|+|(1-Real.exp (-tilt))*C.triggerLower i|)","missing":[],"search":"chargedtriggerfactor_le banditrlproof.cucb.chargedata.chargedtriggerfactor_le theorem chargedtriggerfactor_le (i : fin m) (tilt : ℝ) (n : fin m → ℕ) (z : round a m) : c.chargedtriggerfactor i tilt n z ≤ real.exp (|tilt|+|(1-real.exp (-tilt))*c.triggerlower i|) theorem compiled","shard":"modules/68465bba9f63c3af.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurable_chargedTriggerFactor","label":"measurable_chargedTriggerFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurable_chargedTriggerFactor","description":"theorem measurable_chargedTriggerFactor (i : Fin m) (tilt : ℝ) : Measurable (fun p : (Fin m → ℕ) × Round A m => C.chargedTriggerFactor i tilt p.1 p.2)","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html#decl-9d40406520d8","parent":"module:BanditRLProof.Algorithms.CUCBChargedMGF","order":62,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedMGF"],["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_chargedTriggerFactor (i : Fin m) (tilt : ℝ) : Measurable (fun p : (Fin m → ℕ) × Round A m => C.chargedTriggerFactor i tilt p.1 p.2)","missing":[],"search":"measurable_chargedtriggerfactor banditrlproof.cucb.chargedata.measurable_chargedtriggerfactor theorem measurable_chargedtriggerfactor (i : fin m) (tilt : ℝ) : measurable (fun p : (fin m → ℕ) × round a m => c.chargedtriggerfactor i tilt p.1 p.2) theorem compiled","shard":"modules/68465bba9f63c3af.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.integrable_chargedTriggerFactor_comp","label":"integrable_chargedTriggerFactor_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.integrable_chargedTriggerFactor_comp","description":"theorem integrable_chargedTriggerFactor_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (i : Fin m) (tilt : ℝ) (g : Ω → (Fin m → ℕ) × Round A m) (hg : Measurable g) : Integrable (fun ω => C.chargedTriggerFactor i tilt (g ω).1 (g ω).2) ν","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html#decl-0079c4aceb6c","parent":"module:BanditRLProof.Algorithms.CUCBChargedMGF","order":63,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedMGF"],["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_chargedTriggerFactor_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (i : Fin m) (tilt : ℝ) (g : Ω → (Fin m → ℕ) × Round A m) (hg : Measurable g) : Integrable (fun ω => C.chargedTriggerFactor i tilt (g ω).1 (g ω).2) ν","missing":[],"search":"integrable_chargedtriggerfactor_comp banditrlproof.cucb.chargedata.integrable_chargedtriggerfactor_comp theorem integrable_chargedtriggerfactor_comp {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] (i : fin m) (tilt : ℝ) (g : ω → (fin m → ℕ) × round a m) (hg : measurable g) : integrable (fun ω => c.chargedtriggerfactor i tilt (g ω).1 (g ω).2) ν theorem compiled","shard":"modules/68465bba9f63c3af.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.roundKernel_chargedTriggerFactor_le_one","label":"roundKernel_chargedTriggerFactor_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.roundKernel_chargedTriggerFactor_le_one","description":"theorem roundKernel_chargedTriggerFactor_le_one (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (v : Input m) (N : Fin m → ℕ) (tilt : ℝ) (htilt : 0≤tilt) : (∫z, C.chargedTriggerFactor i tilt N z ∂roundKernel oracle environment v)≤1","url":"../modules/banditrlproof-algorithms-cucbchargedmgf/index.html#decl-f85162ccc31d","parent":"module:BanditRLProof.Algorithms.CUCBChargedMGF","order":64,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBChargedMGF"],["Source","BanditRLProof/Algorithms/CUCBChargedMGF.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem roundKernel_chargedTriggerFactor_le_one (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (v : Input m) (N : Fin m → ℕ) (tilt : ℝ) (htilt : 0≤tilt) : (∫z, C.chargedTriggerFactor i tilt N z ∂roundKernel oracle environment v)≤1","missing":[],"search":"roundkernel_chargedtriggerfactor_le_one banditrlproof.cucb.chargedata.roundkernel_chargedtriggerfactor_le_one theorem roundkernel_chargedtriggerfactor_le_one (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (i : fin m) (htrigger : ∀a, i∈c.triggers a → c.triggerlower i≤(environment a (observedset i)).toreal) (v : input m) (n : fin m → ℕ) (tilt : ℝ) (htilt : 0≤tilt) : (∫z, c.chargedtriggerfactor i tilt n z ∂roundkernel oracle environment v)≤1 theorem compiled","shard":"modules/68465bba9f63c3af.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.pathNoise","label":"pathNoise","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.pathNoise","description":"noncomputable def pathNoise (D : Measure UnitOutcome) (i : Fin m) (t : ℕ) (Y : ℕ → Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-26c8aae13c0e","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":65,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def pathNoise (D : Measure UnitOutcome) (i : Fin m) (t : ℕ) (Y : ℕ → Round A m) : ℝ","missing":[],"search":"pathnoise banditrlproof.cucb.pathnoise noncomputable def pathnoise (d : measure unitoutcome) (i : fin m) (t : ℕ) (y : ℕ → round a m) : ℝ definition compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.pathCount","label":"pathCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.pathCount","description":"def pathCount (i : Fin m) (t : ℕ) (Y : ℕ → Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-68c37b85a4aa","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":66,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def pathCount (i : Fin m) (t : ℕ) (Y : ℕ → Round A m) : ℝ","missing":[],"search":"pathcount banditrlproof.cucb.pathcount def pathcount (i : fin m) (t : ℕ) (y : ℕ → round a m) : ℝ definition compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_round_piLE","label":"measurable_round_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_round_piLE","description":"theorem measurable_round_piLE (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → Round A m => Y n)","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-c95e8fbaa00d","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":67,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_round_piLE (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → Round A m => Y n)","missing":[],"search":"measurable_round_pile banditrlproof.cucb.measurable_round_pile theorem measurable_round_pile (n : ℕ) : measurable[filtration.pile n] (fun y : ℕ → round a m => y n) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_compensated_adapted","label":"path_compensated_adapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_compensated_adapted","description":"theorem path_compensated_adapted (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) : StronglyAdapted Filtration.piLE (fun t (Y : ℕ → Round A m) => tilt*pathNoise D i t Y-tilt^2/8*pathCount i t Y)","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-adc65e7c05ea","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":68,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_compensated_adapted (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) : StronglyAdapted Filtration.piLE (fun t (Y : ℕ → Round A m) => tilt*pathNoise D i t Y-tilt^2/8*pathCount i t Y)","missing":[],"search":"path_compensated_adapted banditrlproof.cucb.path_compensated_adapted theorem path_compensated_adapted (d : measure unitoutcome) (i : fin m) (tilt : ℝ) : stronglyadapted filtration.pile (fun t (y : ℕ → round a m) => tilt*pathnoise d i t y-tilt^2/8*pathcount i t y) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_noise_count_tail","label":"path_noise_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_noise_count_tail","description":"theorem path_noise_count_tail (n : ℕ) (tilt threshold budget : ℝ) (htilt : 0≤tilt) : (cucbTrajectory oracle environment) {Y | threshold≤∑t∈Finset.range n, pathNoise D i t Y ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-tilt*threshold+tilt^2/8*budget))","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-2bcdaf8addbf","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":69,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_noise_count_tail (n : ℕ) (tilt threshold budget : ℝ) (htilt : 0≤tilt) : (cucbTrajectory oracle environment) {Y | threshold≤∑t∈Finset.range n, pathNoise D i t Y ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-tilt*threshold+tilt^2/8*budget))","missing":[],"search":"path_noise_count_tail banditrlproof.cucb.path_noise_count_tail theorem path_noise_count_tail (n : ℕ) (tilt threshold budget : ℝ) (htilt : 0≤tilt) : (cucbtrajectory oracle environment) {y | threshold≤∑t∈finset.range n, pathnoise d i t y ∧ (∑t∈finset.range n, pathcount i t y)≤budget} ≤ ennreal.ofreal (real.exp (-tilt*threshold+tilt^2/8*budget)) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_noise_count_tail_optimized","label":"path_noise_count_tail_optimized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_noise_count_tail_optimized","description":"theorem path_noise_count_tail_optimized (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤∑t∈Finset.range n, pathNoise D i t Y ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-40affec3e489","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":70,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_noise_count_tail_optimized (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤∑t∈Finset.range n, pathNoise D i t Y ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","missing":[],"search":"path_noise_count_tail_optimized banditrlproof.cucb.path_noise_count_tail_optimized theorem path_noise_count_tail_optimized (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbtrajectory oracle environment) {y | threshold≤∑t∈finset.range n, pathnoise d i t y ∧ (∑t∈finset.range n, pathcount i t y)≤budget} ≤ ennreal.ofreal (real.exp (-2*threshold^2/budget)) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_negative_noise_count_tail","label":"path_negative_noise_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_negative_noise_count_tail","description":"theorem path_negative_noise_count_tail (n : ℕ) (tilt threshold budget : ℝ) (htilt : 0≤tilt) : (cucbTrajectory oracle environment) {Y | threshold≤-(∑t∈Finset.range n, pathNoise D i t Y) ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-tilt*threshold+tilt^2/8*budget))","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-8c4c0b3bfcb0","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":71,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_negative_noise_count_tail (n : ℕ) (tilt threshold budget : ℝ) (htilt : 0≤tilt) : (cucbTrajectory oracle environment) {Y | threshold≤-(∑t∈Finset.range n, pathNoise D i t Y) ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-tilt*threshold+tilt^2/8*budget))","missing":[],"search":"path_negative_noise_count_tail banditrlproof.cucb.path_negative_noise_count_tail theorem path_negative_noise_count_tail (n : ℕ) (tilt threshold budget : ℝ) (htilt : 0≤tilt) : (cucbtrajectory oracle environment) {y | threshold≤-(∑t∈finset.range n, pathnoise d i t y) ∧ (∑t∈finset.range n, pathcount i t y)≤budget} ≤ ennreal.ofreal (real.exp (-tilt*threshold+tilt^2/8*budget)) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_negative_noise_count_tail_optimized","label":"path_negative_noise_count_tail_optimized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_negative_noise_count_tail_optimized","description":"theorem path_negative_noise_count_tail_optimized (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤-(∑t∈Finset.range n, pathNoise D i t Y) ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-034c010c9750","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":72,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_negative_noise_count_tail_optimized (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤-(∑t∈Finset.range n, pathNoise D i t Y) ∧ (∑t∈Finset.range n, pathCount i t Y)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","missing":[],"search":"path_negative_noise_count_tail_optimized banditrlproof.cucb.path_negative_noise_count_tail_optimized theorem path_negative_noise_count_tail_optimized (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbtrajectory oracle environment) {y | threshold≤-(∑t∈finset.range n, pathnoise d i t y) ∧ (∑t∈finset.range n, pathcount i t y)≤budget} ≤ ennreal.ofreal (real.exp (-2*threshold^2/budget)) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.sum_pathCount","label":"sum_pathCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.sum_pathCount","description":"theorem sum_pathCount (n : ℕ) (Y : ℕ → Round A m) : (∑t∈Finset.range n, pathCount i t Y) = (observationCount (fun t => (Y t).2) n i : ℝ)","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-6964a69b6579","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":73,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_pathCount (n : ℕ) (Y : ℕ → Round A m) : (∑t∈Finset.range n, pathCount i t Y) = (observationCount (fun t => (Y t).2) n i : ℝ)","missing":[],"search":"sum_pathcount banditrlproof.cucb.sum_pathcount theorem sum_pathcount (n : ℕ) (y : ℕ → round a m) : (∑t∈finset.range n, pathcount i t y) = (observationcount (fun t => (y t).2) n i : ℝ) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.sum_pathNoise","label":"sum_pathNoise","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.sum_pathNoise","description":"theorem sum_pathNoise (n : ℕ) (Y : ℕ → Round A m) : (∑t∈Finset.range n, pathNoise D i t Y) = observationSum (fun t => (Y t).2) n i - (observationCount (fun t => (Y t).2) n i : ℝ)*marginalMean D","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-091c5e77cae5","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":74,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_pathNoise (n : ℕ) (Y : ℕ → Round A m) : (∑t∈Finset.range n, pathNoise D i t Y) = observationSum (fun t => (Y t).2) n i - (observationCount (fun t => (Y t).2) n i : ℝ)*marginalMean D","missing":[],"search":"sum_pathnoise banditrlproof.cucb.sum_pathnoise theorem sum_pathnoise (n : ℕ) (y : ℕ → round a m) : (∑t∈finset.range n, pathnoise d i t y) = observationsum (fun t => (y t).2) n i - (observationcount (fun t => (y t).2) n i : ℝ)*marginalmean d theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observed_sum_upper_tail","label":"observed_sum_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observed_sum_upper_tail","description":"theorem observed_sum_upper_tail (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤observationSum (fun t => (Y t).2) n i - (observationCount (fun t => (Y t).2) n i : ℝ)*marginalMean D ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-024bc1761cc9","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":75,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observed_sum_upper_tail (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤observationSum (fun t => (Y t).2) n i - (observationCount (fun t => (Y t).2) n i : ℝ)*marginalMean D ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","missing":[],"search":"observed_sum_upper_tail banditrlproof.cucb.observed_sum_upper_tail theorem observed_sum_upper_tail (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbtrajectory oracle environment) {y | threshold≤observationsum (fun t => (y t).2) n i - (observationcount (fun t => (y t).2) n i : ℝ)*marginalmean d ∧ (observationcount (fun t => (y t).2) n i : ℝ)≤budget} ≤ ennreal.ofreal (real.exp (-2*threshold^2/budget)) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observed_sum_lower_tail","label":"observed_sum_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observed_sum_lower_tail","description":"theorem observed_sum_lower_tail (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤(observationCount (fun t => (Y t).2) n i : ℝ)*marginalMean D - observationSum (fun t => (Y t).2) n i ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","url":"../modules/banditrlproof-algorithms-cucbconcentration/index.html#decl-2af19d63d0e7","parent":"module:BanditRLProof.Algorithms.CUCBConcentration","order":76,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConcentration"],["Source","BanditRLProof/Algorithms/CUCBConcentration.lean:135"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observed_sum_lower_tail (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbTrajectory oracle environment) {Y | threshold≤(observationCount (fun t => (Y t).2) n i : ℝ)*marginalMean D - observationSum (fun t => (Y t).2) n i ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤budget} ≤ ENNReal.ofReal (Real.exp (-2*threshold^2/budget))","missing":[],"search":"observed_sum_lower_tail banditrlproof.cucb.observed_sum_lower_tail theorem observed_sum_lower_tail (n : ℕ) (threshold budget : ℝ) (hx : 0≤threshold) (hb : 0<budget) : (cucbtrajectory oracle environment) {y | threshold≤(observationcount (fun t => (y t).2) n i : ℝ)*marginalmean d - observationsum (fun t => (y t).2) n i ∧ (observationcount (fun t => (y t).2) n i : ℝ)≤budget} ≤ ennreal.ofreal (real.exp (-2*threshold^2/budget)) theorem compiled","shard":"modules/2d562cf58d444a24.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedNoise","label":"observedNoise","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observedNoise","description":"noncomputable def observedNoise {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (z : Feedback m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-69a7f5ad0492","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":77,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def observedNoise {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (z : Feedback m) : ℝ","missing":[],"search":"observednoise banditrlproof.cucb.observednoise noncomputable def observednoise {m : ℕ} (d : measure unitoutcome) (i : fin m) (z : feedback m) : ℝ definition compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedIndicator","label":"observedIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observedIndicator","description":"def observedIndicator {m : ℕ} (i : Fin m) (z : Feedback m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-8b623b7680b5","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":78,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def observedIndicator {m : ℕ} (i : Fin m) (z : Feedback m) : ℝ","missing":[],"search":"observedindicator banditrlproof.cucb.observedindicator def observedindicator {m : ℕ} (i : fin m) (z : feedback m) : ℝ definition compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedCompensated","label":"observedCompensated","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observedCompensated","description":"noncomputable def observedCompensated {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-1c526b1e52b4","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":79,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def observedCompensated {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : ℝ","missing":[],"search":"observedcompensated banditrlproof.cucb.observedcompensated noncomputable def observedcompensated {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) (z : feedback m) : ℝ definition compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_observedCompensated","label":"measurable_observedCompensated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_observedCompensated","description":"theorem measurable_observedCompensated {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) : Measurable (observedCompensated D i tilt)","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-c2575eb229fb","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":80,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_observedCompensated {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) : Measurable (observedCompensated D i tilt)","missing":[],"search":"measurable_observedcompensated banditrlproof.cucb.measurable_observedcompensated theorem measurable_observedcompensated {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) : measurable (observedcompensated d i tilt) theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedCompensated_abs_le","label":"observedCompensated_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observedCompensated_abs_le","description":"theorem observedCompensated_abs_le {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt : ℝ) (z : Feedback m) : |observedCompensated D i tilt z|≤|tilt|+tilt^2/8","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-b5db47e2897f","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":81,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observedCompensated_abs_le {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt : ℝ) (z : Feedback m) : |observedCompensated D i tilt z|≤|tilt|+tilt^2/8","missing":[],"search":"observedcompensated_abs_le banditrlproof.cucb.observedcompensated_abs_le theorem observedcompensated_abs_le {m : ℕ} (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (tilt : ℝ) (z : feedback m) : |observedcompensated d i tilt z|≤|tilt|+tilt^2/8 theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integrable_exp_observedCompensated","label":"integrable_exp_observedCompensated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integrable_exp_observedCompensated","description":"theorem integrable_exp_observedCompensated {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt scale : ℝ) (g : Ω → Feedback m) (hg : Measurable g) : Integrable (fun ω => Real.exp (scale*observedCompensated D i tilt (g ω))) ν","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-1d838a8c71d6","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":82,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_observedCompensated {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt scale : ℝ) (g : Ω → Feedback m) (hg : Measurable g) : Integrable (fun ω => Real.exp (scale*observedCompensated D i tilt (g ω))) ν","missing":[],"search":"integrable_exp_observedcompensated banditrlproof.cucb.integrable_exp_observedcompensated theorem integrable_exp_observedcompensated {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] {m : ℕ} (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (tilt scale : ℝ) (g : ω → feedback m) (hg : measurable g) : integrable (fun ω => real.exp (scale*observedcompensated d i tilt (g ω))) ν theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.exp_observedCompensated","label":"exp_observedCompensated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.exp_observedCompensated","description":"theorem exp_observedCompensated {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : Real.exp (observedCompensated D i tilt z)=observedFactor D i tilt z","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-8484224ed682","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":83,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_observedCompensated {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : Real.exp (observedCompensated D i tilt z)=observedFactor D i tilt z","missing":[],"search":"exp_observedcompensated banditrlproof.cucb.exp_observedcompensated theorem exp_observedcompensated {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) (z : feedback m) : real.exp (observedcompensated d i tilt z)=observedfactor d i tilt z theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucb_condExp_observedCompensated","label":"cucb_condExp_observedCompensated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucb_condExp_observedCompensated","description":"theorem cucb_condExp_observedCompensated (n : ℕ) (tilt : ℝ) : (cucbTrajectory oracle environment)[fun Y => Real.exp (observedCompensated D i tilt (Y (n+1)).2) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance] ≤ᵐ[cucbTrajectory oracle environment] fun _ => 1","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-37bab3017e8c","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":84,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucb_condExp_observedCompensated (n : ℕ) (tilt : ℝ) : (cucbTrajectory oracle environment)[fun Y => Real.exp (observedCompensated D i tilt (Y (n+1)).2) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance] ≤ᵐ[cucbTrajectory oracle environment] fun _ => 1","missing":[],"search":"cucb_condexp_observedcompensated banditrlproof.cucb.cucb_condexp_observedcompensated theorem cucb_condexp_observedcompensated (n : ℕ) (tilt : ℝ) : (cucbtrajectory oracle environment)[fun y => real.exp (observedcompensated d i tilt (y (n+1)).2) | measurablespace.comap (preorder.frestrictle n) inferinstance] ≤ᵐ[cucbtrajectory oracle environment] fun _ => 1 theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucb_successor_condMGF","label":"cucb_successor_condMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucb_successor_condMGF","description":"theorem cucb_successor_condMGF (n : ℕ) (tilt : ℝ) : Concentration.HasCondMGFUpperBoundAt (MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance) (Preorder.measurable_frestrictLe n).comap_le (fun Y => observedCompensated D i tilt (Y (n+1)).2) 1 0 (cucbTrajectory oracle environment)","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-63afbe292883","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":85,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucb_successor_condMGF (n : ℕ) (tilt : ℝ) : Concentration.HasCondMGFUpperBoundAt (MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance) (Preorder.measurable_frestrictLe n).comap_le (fun Y => observedCompensated D i tilt (Y (n+1)).2) 1 0 (cucbTrajectory oracle environment)","missing":[],"search":"cucb_successor_condmgf banditrlproof.cucb.cucb_successor_condmgf theorem cucb_successor_condmgf (n : ℕ) (tilt : ℝ) : concentration.hascondmgfupperboundat (measurablespace.comap (preorder.frestrictle n) inferinstance) (preorder.measurable_frestrictle n).comap_le (fun y => observedcompensated d i tilt (y (n+1)).2) 1 0 (cucbtrajectory oracle environment) theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucb_initial_MGF","label":"cucb_initial_MGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucb_initial_MGF","description":"theorem cucb_initial_MGF (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (fun Y : ℕ → Round A m => observedCompensated D i tilt (Y 0).2) 1 0 (cucbTrajectory oracle environment)","url":"../modules/banditrlproof-algorithms-cucbconditionalmgf/index.html#decl-51c19be63a14","parent":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","order":86,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConditionalMGF"],["Source","BanditRLProof/Algorithms/CUCBConditionalMGF.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucb_initial_MGF (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (fun Y : ℕ → Round A m => observedCompensated D i tilt (Y 0).2) 1 0 (cucbTrajectory oracle environment)","missing":[],"search":"cucb_initial_mgf banditrlproof.cucb.cucb_initial_mgf theorem cucb_initial_mgf (tilt : ℝ) : concentration.hasmgfupperboundat (fun y : ℕ → round a m => observedcompensated d i tilt (y 0).2) 1 0 (cucbtrajectory oracle environment) theorem compiled","shard":"modules/d415b5a7946463ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationCount_le","label":"observationCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observationCount_le","description":"theorem observationCount_le {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationCount Y n i≤n","url":"../modules/banditrlproof-algorithms-cucbconfidence/index.html#decl-6a825506f6e8","parent":"module:BanditRLProof.Algorithms.CUCBConfidence","order":87,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConfidence"],["Source","BanditRLProof/Algorithms/CUCBConfidence.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observationCount_le {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationCount Y n i≤n","missing":[],"search":"observationcount_le banditrlproof.cucb.observationcount_le theorem observationcount_le {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : observationcount y n i≤n theorem compiled","shard":"modules/29afa36b1f71f0fe.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.pathDeviation","label":"pathDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.pathDeviation","description":"noncomputable def pathDeviation (D : Measure UnitOutcome) (i : Fin m) (n : ℕ) (lower : Bool) (Y : ℕ → Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbconfidence/index.html#decl-4224f147dfbc","parent":"module:BanditRLProof.Algorithms.CUCBConfidence","order":88,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBConfidence"],["Source","BanditRLProof/Algorithms/CUCBConfidence.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def pathDeviation (D : Measure UnitOutcome) (i : Fin m) (n : ℕ) (lower : Bool) (Y : ℕ → Round A m) : ℝ","missing":[],"search":"pathdeviation banditrlproof.cucb.pathdeviation noncomputable def pathdeviation (d : measure unitoutcome) (i : fin m) (n : ℕ) (lower : bool) (y : ℕ → round a m) : ℝ definition compiled","shard":"modules/29afa36b1f71f0fe.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_deviation_slice","label":"path_deviation_slice","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_deviation_slice","description":"theorem path_deviation_slice (n k : ℕ) (lower : Bool) (L : ℝ) (hL : 0≤L) (hk : 0<k) : (cucbTrajectory oracle environment) {Y | Real.sqrt ((k:ℝ)*L/2)≤pathDeviation D i n lower Y ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤k} ≤ ENNReal.ofReal (Real.exp (-L))","url":"../modules/banditrlproof-algorithms-cucbconfidence/index.html#decl-58e6818d04bd","parent":"module:BanditRLProof.Algorithms.CUCBConfidence","order":89,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConfidence"],["Source","BanditRLProof/Algorithms/CUCBConfidence.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_deviation_slice (n k : ℕ) (lower : Bool) (L : ℝ) (hL : 0≤L) (hk : 0<k) : (cucbTrajectory oracle environment) {Y | Real.sqrt ((k:ℝ)*L/2)≤pathDeviation D i n lower Y ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤k} ≤ ENNReal.ofReal (Real.exp (-L))","missing":[],"search":"path_deviation_slice banditrlproof.cucb.path_deviation_slice theorem path_deviation_slice (n k : ℕ) (lower : bool) (l : ℝ) (hl : 0≤l) (hk : 0<k) : (cucbtrajectory oracle environment) {y | real.sqrt ((k:ℝ)*l/2)≤pathdeviation d i n lower y ∧ (observationcount (fun t => (y t).2) n i : ℝ)≤k} ≤ ennreal.ofreal (real.exp (-l)) theorem compiled","shard":"modules/29afa36b1f71f0fe.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.path_deviation_confidence","label":"path_deviation_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.path_deviation_confidence","description":"theorem path_deviation_confidence (n : ℕ) (lower : Bool) (L : ℝ) (hL : 0≤L) : (cucbTrajectory oracle environment) {Y | 0<observationCount (fun t => (Y t).2) n i ∧ Real.sqrt ((observationCount (fun t => (Y t).2) n i : ℝ)*L/2)≤ pathDeviation D i n lower Y} ≤ (n:ENNReal)*ENNReal.ofReal (Real.exp (-L))","url":"../modules/banditrlproof-algorithms-cucbconfidence/index.html#decl-46feb896267b","parent":"module:BanditRLProof.Algorithms.CUCBConfidence","order":90,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBConfidence"],["Source","BanditRLProof/Algorithms/CUCBConfidence.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem path_deviation_confidence (n : ℕ) (lower : Bool) (L : ℝ) (hL : 0≤L) : (cucbTrajectory oracle environment) {Y | 0<observationCount (fun t => (Y t).2) n i ∧ Real.sqrt ((observationCount (fun t => (Y t).2) n i : ℝ)*L/2)≤ pathDeviation D i n lower Y} ≤ (n:ENNReal)*ENNReal.ofReal (Real.exp (-L))","missing":[],"search":"path_deviation_confidence banditrlproof.cucb.path_deviation_confidence theorem path_deviation_confidence (n : ℕ) (lower : bool) (l : ℝ) (hl : 0≤l) : (cucbtrajectory oracle environment) {y | 0<observationcount (fun t => (y t).2) n i ∧ real.sqrt ((observationcount (fun t => (y t).2) n i : ℝ)*l/2)≤ pathdeviation d i n lower y} ≤ (n:ennreal)*ennreal.ofreal (real.exp (-l)) theorem compiled","shard":"modules/29afa36b1f71f0fe.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucbTrajectory_ae_round_property","label":"cucbTrajectory_ae_round_property","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbTrajectory_ae_round_property","description":"theorem cucbTrajectory_ae_round_property (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (s : Set (Round A m)) (hs : MeasurableSet s) (hround : ∀v, ∀ᵐ z ∂roundKernel oracle environment v, z∈s) : ∀ᵐ Y ∂cucbTrajectory oracle environment, ∀n, Y n∈s","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html#decl-0f4f0df00a66","parent":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","order":91,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBDeterministicTrigger"],["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucbTrajectory_ae_round_property (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (s : Set (Round A m)) (hs : MeasurableSet s) (hround : ∀v, ∀ᵐ z ∂roundKernel oracle environment v, z∈s) : ∀ᵐ Y ∂cucbTrajectory oracle environment, ∀n, Y n∈s","missing":[],"search":"cucbtrajectory_ae_round_property banditrlproof.cucb.cucbtrajectory_ae_round_property theorem cucbtrajectory_ae_round_property (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (s : set (round a m)) (hs : measurableset s) (hround : ∀v, ∀ᵐ z ∂roundkernel oracle environment v, z∈s) : ∀ᵐ y ∂cucbtrajectory oracle environment, ∀n, y n∈s theorem compiled","shard":"modules/aa338b4ac58b9cc2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.triggerSupport","label":"triggerSupport","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.triggerSupport","description":"def triggerSupport (i : Fin m) : Set (Round A m)","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html#decl-a4531bcca3b4","parent":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","order":92,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBDeterministicTrigger"],["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def triggerSupport (i : Fin m) : Set (Round A m)","missing":[],"search":"triggersupport banditrlproof.cucb.chargedata.triggersupport def triggersupport (i : fin m) : set (round a m) definition compiled","shard":"modules/aa338b4ac58b9cc2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.measurableSet_triggerSupport","label":"measurableSet_triggerSupport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.measurableSet_triggerSupport","description":"theorem measurableSet_triggerSupport (i : Fin m) : MeasurableSet (C.triggerSupport i)","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html#decl-dff8fe6573d3","parent":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","order":93,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBDeterministicTrigger"],["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_triggerSupport (i : Fin m) : MeasurableSet (C.triggerSupport i)","missing":[],"search":"measurableset_triggersupport banditrlproof.cucb.chargedata.measurableset_triggersupport theorem measurableset_triggersupport (i : fin m) : measurableset (c.triggersupport i) theorem compiled","shard":"modules/aa338b4ac58b9cc2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.environment_ae_observed_of_one","label":"environment_ae_observed_of_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.environment_ae_observed_of_one","description":"theorem environment_ae_observed_of_one (environment : Kernel A (Feedback m)) [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (hp : C.triggerLower i=1) (a : A) (ha : i∈C.triggers a) : ∀ᵐ z ∂environment a, z.1 i=true","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html#decl-f150010efd66","parent":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","order":94,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBDeterministicTrigger"],["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem environment_ae_observed_of_one (environment : Kernel A (Feedback m)) [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (hp : C.triggerLower i=1) (a : A) (ha : i∈C.triggers a) : ∀ᵐ z ∂environment a, z.1 i=true","missing":[],"search":"environment_ae_observed_of_one banditrlproof.cucb.chargedata.environment_ae_observed_of_one theorem environment_ae_observed_of_one (environment : kernel a (feedback m)) [ismarkovkernel environment] (i : fin m) (htrigger : ∀a, i∈c.triggers a → c.triggerlower i≤(environment a (observedset i)).toreal) (hp : c.triggerlower i=1) (a : a) (ha : i∈c.triggers a) : ∀ᵐ z ∂environment a, z.1 i=true theorem compiled","shard":"modules/aa338b4ac58b9cc2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.roundKernel_ae_triggerSupport_of_one","label":"roundKernel_ae_triggerSupport_of_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.roundKernel_ae_triggerSupport_of_one","description":"theorem roundKernel_ae_triggerSupport_of_one (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (hp : C.triggerLower i=1) (v : Input m) : ∀ᵐ z ∂roundKernel oracle environment v, z∈C.triggerSupport i","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html#decl-c3e49d3e1f59","parent":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","order":95,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBDeterministicTrigger"],["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem roundKernel_ae_triggerSupport_of_one (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (hp : C.triggerLower i=1) (v : Input m) : ∀ᵐ z ∂roundKernel oracle environment v, z∈C.triggerSupport i","missing":[],"search":"roundkernel_ae_triggersupport_of_one banditrlproof.cucb.chargedata.roundkernel_ae_triggersupport_of_one theorem roundkernel_ae_triggersupport_of_one (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (i : fin m) (htrigger : ∀a, i∈c.triggers a → c.triggerlower i≤(environment a (observedset i)).toreal) (hp : c.triggerlower i=1) (v : input m) : ∀ᵐ z ∂roundkernel oracle environment v, z∈c.triggersupport i theorem compiled","shard":"modules/aa338b4ac58b9cc2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_le_observations_ae_of_one","label":"counters_le_observations_ae_of_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_le_observations_ae_of_one","description":"theorem counters_le_observations_ae_of_one (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (hp : C.triggerLower i=1) : ∀ᵐ Y ∂cucbTrajectory oracle environment, ∀n, C.counters (fun t => (Y t).1) n i≤observationCount (fun t => (Y t).2) n i","url":"../modules/banditrlproof-algorithms-cucbdeterministictrigger/index.html#decl-289abdeb73ec","parent":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","order":96,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBDeterministicTrigger"],["Source","BanditRLProof/Algorithms/CUCBDeterministicTrigger.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_le_observations_ae_of_one (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (i : Fin m) (htrigger : ∀a, i∈C.triggers a → C.triggerLower i≤(environment a (observedSet i)).toReal) (hp : C.triggerLower i=1) : ∀ᵐ Y ∂cucbTrajectory oracle environment, ∀n, C.counters (fun t => (Y t).1) n i≤observationCount (fun t => (Y t).2) n i","missing":[],"search":"counters_le_observations_ae_of_one banditrlproof.cucb.chargedata.counters_le_observations_ae_of_one theorem counters_le_observations_ae_of_one (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (i : fin m) (htrigger : ∀a, i∈c.triggers a → c.triggerlower i≤(environment a (observedset i)).toreal) (hp : c.triggerlower i=1) : ∀ᵐ y ∂cucbtrajectory oracle environment, ∀n, c.counters (fun t => (y t).1) n i≤observationcount (fun t => (y t).2) n i theorem compiled","shard":"modules/aa338b4ac58b9cc2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel","label":"FeedbackModel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel","description":"structure FeedbackModel (A : Type*) [MeasurableSpace A] (m : ℕ) where","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-1595ad5845db","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":97,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"structure FeedbackModel (A : Type*) [MeasurableSpace A] (m : ℕ) where","missing":[],"search":"feedbackmodel banditrlproof.cucb.feedbackmodel structure feedbackmodel (a : type*) [measurablespace a] (m : ℕ) where structure compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.trueInput","label":"trueInput","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.trueInput","description":"noncomputable def trueInput : Input m","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-c0f4332d0062","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":98,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def trueInput : Input m","missing":[],"search":"trueinput banditrlproof.cucb.feedbackmodel.trueinput noncomputable def trueinput : input m definition compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.expectedReward","label":"expectedReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.expectedReward","description":"noncomputable def expectedReward (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-486345476eb5","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":99,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedReward (a : A) : ℝ","missing":[],"search":"expectedreward banditrlproof.cucb.feedbackmodel.expectedreward noncomputable def expectedreward (a : a) : ℝ definition compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerProbability","label":"triggerProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.triggerProbability","description":"noncomputable def triggerProbability (a : A) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-d2e7a01f7515","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":100,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def triggerProbability (a : A) (i : Fin m) : ℝ","missing":[],"search":"triggerprobability banditrlproof.cucb.feedbackmodel.triggerprobability noncomputable def triggerprobability (a : a) (i : fin m) : ℝ definition compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.expectedReward_nonneg","label":"expectedReward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.expectedReward_nonneg","description":"theorem expectedReward_nonneg (a : A) : 0≤M.expectedReward a","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-c20dd3a2cd5e","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":101,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedReward_nonneg (a : A) : 0≤M.expectedReward a","missing":[],"search":"expectedreward_nonneg banditrlproof.cucb.feedbackmodel.expectedreward_nonneg theorem expectedreward_nonneg (a : a) : 0≤m.expectedreward a theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerProbability_le_one","label":"triggerProbability_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.triggerProbability_le_one","description":"theorem triggerProbability_le_one (a : A) (i : Fin m) : M.triggerProbability a i≤1","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-1175a9867cbd","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":102,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem triggerProbability_le_one (a : A) (i : Fin m) : M.triggerProbability a i≤1","missing":[],"search":"triggerprobability_le_one banditrlproof.cucb.feedbackmodel.triggerprobability_le_one theorem triggerprobability_le_one (a : a) (i : fin m) : m.triggerprobability a i≤1 theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerProbability_pos","label":"triggerProbability_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.triggerProbability_pos","description":"theorem triggerProbability_pos (a : A) (i : Fin m) (hi : i∈M.possible a) : 0<M.triggerProbability a i","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-5f79e1e12dd9","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":103,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem triggerProbability_pos (a : A) (i : Fin m) (hi : i∈M.possible a) : 0<M.triggerProbability a i","missing":[],"search":"triggerprobability_pos banditrlproof.cucb.feedbackmodel.triggerprobability_pos theorem triggerprobability_pos (a : a) (i : fin m) (hi : i∈m.possible a) : 0<m.triggerprobability a i theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerActions","label":"triggerActions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.triggerActions","description":"noncomputable def triggerActions (i : Fin m) : Finset A","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-b173f6bb5b97","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":104,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def triggerActions (i : Fin m) : Finset A","missing":[],"search":"triggeractions banditrlproof.cucb.feedbackmodel.triggeractions noncomputable def triggeractions (i : fin m) : finset a definition compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.mem_triggerActions","label":"mem_triggerActions","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.mem_triggerActions","description":"theorem mem_triggerActions (a : A) (i : Fin m) : a∈M.triggerActions i ↔ i∈M.possible a","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-9afd052cb0ad","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":105,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_triggerActions (a : A) (i : Fin m) : a∈M.triggerActions i ↔ i∈M.possible a","missing":[],"search":"mem_triggeractions banditrlproof.cucb.feedbackmodel.mem_triggeractions theorem mem_triggeractions (a : a) (i : fin m) : a∈m.triggeractions i ↔ i∈m.possible a theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerActions_nonempty","label":"triggerActions_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.triggerActions_nonempty","description":"theorem triggerActions_nonempty (i : Fin m) : (M.triggerActions i).Nonempty","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-e3f1bc2770ee","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":106,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem triggerActions_nonempty (i : Fin m) : (M.triggerActions i).Nonempty","missing":[],"search":"triggeractions_nonempty banditrlproof.cucb.feedbackmodel.triggeractions_nonempty theorem triggeractions_nonempty (i : fin m) : (m.triggeractions i).nonempty theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger","label":"minTrigger","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.minTrigger","description":"noncomputable def minTrigger (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-9f2e4c239da4","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":107,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def minTrigger (i : Fin m) : ℝ","missing":[],"search":"mintrigger banditrlproof.cucb.feedbackmodel.mintrigger noncomputable def mintrigger (i : fin m) : ℝ definition compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger_pos","label":"minTrigger_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.minTrigger_pos","description":"theorem minTrigger_pos (i : Fin m) : 0<M.minTrigger i","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-2e9e1009ea15","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":108,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem minTrigger_pos (i : Fin m) : 0<M.minTrigger i","missing":[],"search":"mintrigger_pos banditrlproof.cucb.feedbackmodel.mintrigger_pos theorem mintrigger_pos (i : fin m) : 0<m.mintrigger i theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger_le","label":"minTrigger_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.minTrigger_le","description":"theorem minTrigger_le (a : A) (i : Fin m) (hi : i∈M.possible a) : M.minTrigger i≤M.triggerProbability a i","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-91860f59414d","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":109,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem minTrigger_le (a : A) (i : Fin m) (hi : i∈M.possible a) : M.minTrigger i≤M.triggerProbability a i","missing":[],"search":"mintrigger_le banditrlproof.cucb.feedbackmodel.mintrigger_le theorem mintrigger_le (a : a) (i : fin m) (hi : i∈m.possible a) : m.mintrigger i≤m.triggerprobability a i theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger_le_one","label":"minTrigger_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.minTrigger_le_one","description":"theorem minTrigger_le_one (i : Fin m) : M.minTrigger i≤1","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-086858d5f19b","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":110,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem minTrigger_le_one (i : Fin m) : M.minTrigger i≤1","missing":[],"search":"mintrigger_le_one banditrlproof.cucb.feedbackmodel.mintrigger_le_one theorem mintrigger_le_one (i : fin m) : m.mintrigger i≤1 theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.chargeData","label":"chargeData","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.chargeData","description":"Source gaps will supply the two analysis fields; the trigger lower bound is already fixed to the actual environment minimum, with its proof below.","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-6abaeab1247b","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":111,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def chargeData (bad : A → Bool) (inverseGap : A → ℝ) : ChargeData A m where","missing":[],"search":"chargedata banditrlproof.cucb.feedbackmodel.chargedata source gaps will supply the two analysis fields; the trigger lower bound is already fixed to the actual environment minimum, with its proof below. definition compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.chargeData_trigger_bound","label":"chargeData_trigger_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.chargeData_trigger_bound","description":"theorem chargeData_trigger_bound (bad : A → Bool) (inverseGap : A → ℝ) (i : Fin m) : ∀a, i∈(M.chargeData bad inverseGap).triggers a → (M.chargeData bad inverseGap).triggerLower i≤(M.environment a (observedSet i)).toReal","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-9c7ab00be34e","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":112,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chargeData_trigger_bound (bad : A → Bool) (inverseGap : A → ℝ) (i : Fin m) : ∀a, i∈(M.chargeData bad inverseGap).triggers a → (M.chargeData bad inverseGap).triggerLower i≤(M.environment a (observedSet i)).toReal","missing":[],"search":"chargedata_trigger_bound banditrlproof.cucb.feedbackmodel.chargedata_trigger_bound theorem chargedata_trigger_bound (bad : a → bool) (inversegap : a → ℝ) (i : fin m) : ∀a, i∈(m.chargedata bad inversegap).triggers a → (m.chargedata bad inversegap).triggerlower i≤(m.environment a (observedset i)).toreal theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.deterministic_counter_bound","label":"deterministic_counter_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.deterministic_counter_bound","description":"theorem deterministic_counter_bound (oracle : Kernel (Input m) A) [IsMarkovKernel oracle] (bad : A → Bool) (inverseGap : A → ℝ) (i : Fin m) (hp : M.minTrigger i=1) : ∀ᵐ Y ∂cucbTrajectory oracle M.environment, ∀n, (M.chargeData bad inverseGap).counters (fun t => (Y t).1) n i≤ observationCount (fun t => (Y t).2) n i","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-acaf85784ba1","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":113,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:92"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem deterministic_counter_bound (oracle : Kernel (Input m) A) [IsMarkovKernel oracle] (bad : A → Bool) (inverseGap : A → ℝ) (i : Fin m) (hp : M.minTrigger i=1) : ∀ᵐ Y ∂cucbTrajectory oracle M.environment, ∀n, (M.chargeData bad inverseGap).counters (fun t => (Y t).1) n i≤ observationCount (fun t => (Y t).2) n i","missing":[],"search":"deterministic_counter_bound banditrlproof.cucb.feedbackmodel.deterministic_counter_bound theorem deterministic_counter_bound (oracle : kernel (input m) a) [ismarkovkernel oracle] (bad : a → bool) (inversegap : a → ℝ) (i : fin m) (hp : m.mintrigger i=1) : ∀ᵐ y ∂cucbtrajectory oracle m.environment, ∀n, (m.chargedata bad inversegap).counters (fun t => (y t).1) n i≤ observationcount (fun t => (y t).2) n i theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.charged_observation_tail","label":"charged_observation_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.charged_observation_tail","description":"theorem charged_observation_tail (oracle : Kernel (Input m) A) [IsMarkovKernel oracle] (bad : A → Bool) (inverseGap : A → ℝ) (i : Fin m) (n : ℕ) (k : ℝ) (hk : 0≤k) : (cucbTrajectory oracle M.environment) {Y | k≤((M.chargeData bad inverseGap).counters (fun t => (Y t).1) n i : ℝ) ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤k*M.minTrigger i/2} ≤ ENNReal.ofReal (Real.exp (-k*M.minTrigger i/8))","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-8d804da361bc","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":114,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem charged_observation_tail (oracle : Kernel (Input m) A) [IsMarkovKernel oracle] (bad : A → Bool) (inverseGap : A → ℝ) (i : Fin m) (n : ℕ) (k : ℝ) (hk : 0≤k) : (cucbTrajectory oracle M.environment) {Y | k≤((M.chargeData bad inverseGap).counters (fun t => (Y t).1) n i : ℝ) ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤k*M.minTrigger i/2} ≤ ENNReal.ofReal (Real.exp (-k*M.minTrigger i/8))","missing":[],"search":"charged_observation_tail banditrlproof.cucb.feedbackmodel.charged_observation_tail theorem charged_observation_tail (oracle : kernel (input m) a) [ismarkovkernel oracle] (bad : a → bool) (inversegap : a → ℝ) (i : fin m) (n : ℕ) (k : ℝ) (hk : 0≤k) : (cucbtrajectory oracle m.environment) {y | k≤((m.chargedata bad inversegap).counters (fun t => (y t).1) n i : ℝ) ∧ (observationcount (fun t => (y t).2) n i : ℝ)≤k*m.mintrigger i/2} ≤ ennreal.ofreal (real.exp (-k*m.mintrigger i/8)) theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.nice_event_probability","label":"nice_event_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.nice_event_probability","description":"theorem nice_event_probability (oracle : Kernel (Input m) A) [IsMarkovKernel oracle] (n : ℕ) : (cucbTrajectory oracle M.environment) (NiceEvent (fun i => (M.trueInput i:ℝ)) n)ᶜ ≤ ENNReal.ofReal (2*(m:ℝ)/((n:ℝ)+1)^2)","url":"../modules/banditrlproof-algorithms-cucbfeedbackmodel/index.html#decl-f9c66ffc76b5","parent":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","order":115,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFeedbackModel"],["Source","BanditRLProof/Algorithms/CUCBFeedbackModel.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem nice_event_probability (oracle : Kernel (Input m) A) [IsMarkovKernel oracle] (n : ℕ) : (cucbTrajectory oracle M.environment) (NiceEvent (fun i => (M.trueInput i:ℝ)) n)ᶜ ≤ ENNReal.ofReal (2*(m:ℝ)/((n:ℝ)+1)^2)","missing":[],"search":"nice_event_probability banditrlproof.cucb.feedbackmodel.nice_event_probability theorem nice_event_probability (oracle : kernel (input m) a) [ismarkovkernel oracle] (n : ℕ) : (cucbtrajectory oracle m.environment) (niceevent (fun i => (m.trueinput i:ℝ)) n)ᶜ ≤ ennreal.ofreal (2*(m:ℝ)/((n:ℝ)+1)^2) theorem compiled","shard":"modules/3c74c73b1fcf73e7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.finite_power_sum_le","label":"finite_power_sum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.finite_power_sum_le","description":"theorem finite_power_sum_le {m : ℕ} (hm : 0<m) (z : Fin m → ℝ) (hz : ∀i, 0≤z i) (H p : ℝ) (hp0 : 0≤p) (hp1 : p≤1) (hH : (∑i:Fin m, z i)≤H) : (∑i:Fin m, (z i)^p)≤(m:ℝ)^(1-p)*H^p","url":"../modules/banditrlproof-algorithms-cucbfiniteconcavity/index.html#decl-aecf307f07e6","parent":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","order":116,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteConcavity"],["Source","BanditRLProof/Algorithms/CUCBFiniteConcavity.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finite_power_sum_le {m : ℕ} (hm : 0<m) (z : Fin m → ℝ) (hz : ∀i, 0≤z i) (H p : ℝ) (hp0 : 0≤p) (hp1 : p≤1) (hH : (∑i:Fin m, z i)≤H) : (∑i:Fin m, (z i)^p)≤(m:ℝ)^(1-p)*H^p","missing":[],"search":"finite_power_sum_le banditrlproof.cucb.sourcemodel.finite_power_sum_le theorem finite_power_sum_le {m : ℕ} (hm : 0<m) (z : fin m → ℝ) (hz : ∀i, 0≤z i) (h p : ℝ) (hp0 : 0≤p) (hp1 : p≤1) (hh : (∑i:fin m, z i)≤h) : (∑i:fin m, (z i)^p)≤(m:ℝ)^(1-p)*h^p theorem compiled","shard":"modules/a022e5888d0985fd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeCount_power_sum_le","label":"underChargeCount_power_sum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeCount_power_sum_le","description":"theorem underChargeCount_power_sum_le (H : ℕ) (actions : ℕ → A) (ω : ℝ) (hω : 0<ω) (hω1 : ω≤1) : (∑i:Fin m, ((S.underChargeTimes H actions i).card:ℝ)^(1-ω/2))≤ (m:ℝ)^(ω/2)*(H:ℝ)^(1-ω/2)","url":"../modules/banditrlproof-algorithms-cucbfiniteconcavity/index.html#decl-ba80e997c49f","parent":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","order":117,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteConcavity"],["Source","BanditRLProof/Algorithms/CUCBFiniteConcavity.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem underChargeCount_power_sum_le (H : ℕ) (actions : ℕ → A) (ω : ℝ) (hω : 0<ω) (hω1 : ω≤1) : (∑i:Fin m, ((S.underChargeTimes H actions i).card:ℝ)^(1-ω/2))≤ (m:ℝ)^(ω/2)*(H:ℝ)^(1-ω/2)","missing":[],"search":"underchargecount_power_sum_le banditrlproof.cucb.sourcemodel.underchargecount_power_sum_le theorem underchargecount_power_sum_le (h : ℕ) (actions : ℕ → a) (ω : ℝ) (hω : 0<ω) (hω1 : ω≤1) : (∑i:fin m, ((s.underchargetimes h actions i).card:ℝ)^(1-ω/2))≤ (m:ℝ)^(ω/2)*(h:ℝ)^(1-ω/2) theorem compiled","shard":"modules/a022e5888d0985fd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.fullFeedback","label":"fullFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.fullFeedback","description":"def fullFeedback (a : Bool) (x : Sample) : Feedback 3","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-1fa31858cb7c","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":118,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def fullFeedback (a : Bool) (x : Sample) : Feedback 3","missing":[],"search":"fullfeedback banditrlproof.cucb.finiteexample.fullfeedback def fullfeedback (a : bool) (x : sample) : feedback 3 definition compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.fullEnvironment","label":"fullEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.fullEnvironment","description":"noncomputable def fullEnvironment : Kernel Bool (Feedback 3) where","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-e9f3524f1bf7","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":119,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def fullEnvironment : Kernel Bool (Feedback 3) where","missing":[],"search":"fullenvironment banditrlproof.cucb.finiteexample.fullenvironment noncomputable def fullenvironment : kernel bool (feedback 3) where definition compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_observed","label":"full_observed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_observed","description":"theorem full_observed (a : Bool) (i : Fin 3) : fullEnvironment a (observedSet i)=1","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-1c94d0c29885","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":120,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_observed (a : Bool) (i : Fin 3) : fullEnvironment a (observedSet i)=1","missing":[],"search":"full_observed banditrlproof.cucb.finiteexample.full_observed theorem full_observed (a : bool) (i : fin 3) : fullenvironment a (observedset i)=1 theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_compatible","label":"full_compatible","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_compatible","description":"theorem full_compatible (a : Bool) (i : Fin 3) : ObservationCompatible (fullEnvironment a) (law i) i","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-09a92d6ee28e","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":121,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_compatible (a : Bool) (i : Fin 3) : ObservationCompatible (fullEnvironment a) (law i) i","missing":[],"search":"full_compatible banditrlproof.cucb.finiteexample.full_compatible theorem full_compatible (a : bool) (i : fin 3) : observationcompatible (fullenvironment a) (law i) i theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_reward_integrable","label":"full_reward_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_reward_integrable","description":"theorem full_reward_integrable (a : Bool) : Integrable (fun z => z.2.2) (fullEnvironment a)","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-5984fac8e5b8","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":122,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_reward_integrable (a : Bool) : Integrable (fun z => z.2.2) (fullEnvironment a)","missing":[],"search":"full_reward_integrable banditrlproof.cucb.finiteexample.full_reward_integrable theorem full_reward_integrable (a : bool) : integrable (fun z => z.2.2) (fullenvironment a) theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_reward_nonneg","label":"full_reward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_reward_nonneg","description":"theorem full_reward_nonneg (a : Bool) : ∀ᵐ z ∂fullEnvironment a, 0≤z.2.2","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-833d60efec41","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":123,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_reward_nonneg (a : Bool) : ∀ᵐ z ∂fullEnvironment a, 0≤z.2.2","missing":[],"search":"full_reward_nonneg banditrlproof.cucb.finiteexample.full_reward_nonneg theorem full_reward_nonneg (a : bool) : ∀ᵐ z ∂fullenvironment a, 0≤z.2.2 theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.fullModel","label":"fullModel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.fullModel","description":"noncomputable def fullModel : FeedbackModel Bool 3 where","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-664e3e54bcca","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":124,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def fullModel : FeedbackModel Bool 3 where","missing":[],"search":"fullmodel banditrlproof.cucb.finiteexample.fullmodel noncomputable def fullmodel : feedbackmodel bool 3 where definition compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_expected_reward","label":"full_expected_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_expected_reward","description":"theorem full_expected_reward (a : Bool) : fullModel.expectedReward a=model.expectedReward a","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-28e39b6cc867","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":125,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_expected_reward (a : Bool) : fullModel.expectedReward a=model.expectedReward a","missing":[],"search":"full_expected_reward banditrlproof.cucb.finiteexample.full_expected_reward theorem full_expected_reward (a : bool) : fullmodel.expectedreward a=model.expectedreward a theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.fullSource","label":"fullSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.fullSource","description":"noncomputable def fullSource : SourceModel fullModel where","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-fe852cf21d12","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":126,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def fullSource : SourceModel fullModel where","missing":[],"search":"fullsource banditrlproof.cucb.finiteexample.fullsource noncomputable def fullsource : sourcemodel fullmodel where definition compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_minTrigger","label":"full_minTrigger","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_minTrigger","description":"theorem full_minTrigger (i : Fin 3) : fullModel.minTrigger i=1","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-7df4901449d6","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":127,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_minTrigger (i : Fin 3) : fullModel.minTrigger i=1","missing":[],"search":"full_mintrigger banditrlproof.cucb.finiteexample.full_mintrigger theorem full_mintrigger (i : fin 3) : fullmodel.mintrigger i=1 theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.full_globalMinTrigger","label":"full_globalMinTrigger","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.full_globalMinTrigger","description":"theorem full_globalMinTrigger : fullModel.globalMinTrigger=1","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-e9139713afbd","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":128,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem full_globalMinTrigger : fullModel.globalMinTrigger=1","missing":[],"search":"full_globalmintrigger banditrlproof.cucb.finiteexample.full_globalmintrigger theorem full_globalmintrigger : fullmodel.globalmintrigger=1 theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.deterministic_regret","label":"deterministic_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.deterministic_regret","description":"theorem deterministic_regret (H : ℕ) (hH : 1≤H) : fullSource.approximationRegret H≤ 4*(18*Real.log (H:ℝ))^((1:ℝ)/2)*(H:ℝ)^((1:ℝ)/2)+(1+Real.pi^2/3)*(3/4)","url":"../modules/banditrlproof-algorithms-cucbfinitedeterministicexample/index.html#decl-80c9c6dcddd7","parent":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","order":129,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteDeterministicExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteDeterministicExample.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem deterministic_regret (H : ℕ) (hH : 1≤H) : fullSource.approximationRegret H≤ 4*(18*Real.log (H:ℝ))^((1:ℝ)/2)*(H:ℝ)^((1:ℝ)/2)+(1+Real.pi^2/3)*(3/4)","missing":[],"search":"deterministic_regret banditrlproof.cucb.finiteexample.deterministic_regret theorem deterministic_regret (h : ℕ) (hh : 1≤h) : fullsource.approximationregret h≤ 4*(18*real.log (h:ℝ))^((1:ℝ)/2)*(h:ℝ)^((1:ℝ)/2)+(1+real.pi^2/3)*(3/4) theorem compiled","shard":"modules/d2e3d092c7ae8876.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.Sample","label":"Sample","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.Sample","description":"abbrev Sample","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-15a70dad55a6","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":130,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Sample","missing":[],"search":"sample banditrlproof.cucb.finiteexample.sample abbrev sample abbreviation compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.sampleLaw","label":"sampleLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.sampleLaw","description":"noncomputable def sampleLaw : Measure Sample","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-5a0bb61c3a22","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":131,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sampleLaw : Measure Sample","missing":[],"search":"samplelaw banditrlproof.cucb.finiteexample.samplelaw noncomputable def samplelaw : measure sample definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.bit","label":"bit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.bit","description":"def bit (b : Bool) : UnitOutcome","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-0d7ac7ac83cc","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":132,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def bit (b : Bool) : UnitOutcome","missing":[],"search":"bit banditrlproof.cucb.finiteexample.bit def bit (b : bool) : unitoutcome definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.outcome","label":"outcome","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.outcome","description":"def outcome (x : Sample) (i : Fin 3) : UnitOutcome","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-752f764926be","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":133,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def outcome (x : Sample) (i : Fin 3) : UnitOutcome","missing":[],"search":"outcome banditrlproof.cucb.finiteexample.outcome def outcome (x : sample) (i : fin 3) : unitoutcome definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.selected","label":"selected","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.selected","description":"def selected (a : Bool) : Finset (Fin 3)","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-735d0ba41daf","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":134,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def selected (a : Bool) : Finset (Fin 3)","missing":[],"search":"selected banditrlproof.cucb.finiteexample.selected def selected (a : bool) : finset (fin 3) definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.feedback","label":"feedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.feedback","description":"def feedback (a : Bool) (x : Sample) : Feedback 3","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-980adead46cf","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":135,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def feedback (a : Bool) (x : Sample) : Feedback 3","missing":[],"search":"feedback banditrlproof.cucb.finiteexample.feedback def feedback (a : bool) (x : sample) : feedback 3 definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.environment","label":"environment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.environment","description":"noncomputable def environment : Kernel Bool (Feedback 3) where","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-a9475f518ef4","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":136,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def environment : Kernel Bool (Feedback 3) where","missing":[],"search":"environment banditrlproof.cucb.finiteexample.environment noncomputable def environment : kernel bool (feedback 3) where definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.law","label":"law","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.law","description":"noncomputable def law (i : Fin 3) : Measure UnitOutcome","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-0e3947ff076b","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":137,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def law (i : Fin 3) : Measure UnitOutcome","missing":[],"search":"law banditrlproof.cucb.finiteexample.law noncomputable def law (i : fin 3) : measure unitoutcome definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.sampleLaw_apply","label":"sampleLaw_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.sampleLaw_apply","description":"theorem sampleLaw_apply (s : Set Sample) : sampleLaw s=∑x:Sample, if x∈s then (1/64:ENNReal) else 0","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-0e03a499e5b8","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":138,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleLaw_apply (s : Set Sample) : sampleLaw s=∑x:Sample, if x∈s then (1/64:ENNReal) else 0","missing":[],"search":"samplelaw_apply banditrlproof.cucb.finiteexample.samplelaw_apply theorem samplelaw_apply (s : set sample) : samplelaw s=∑x:sample, if x∈s then (1/64:ennreal) else 0 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.environment_apply","label":"environment_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.environment_apply","description":"theorem environment_apply (a : Bool) (s : Set (Feedback 3)) (hs : MeasurableSet s) : environment a s=∑x:Sample, if feedback a x∈s then (1/64:ENNReal) else 0","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-40e1257bcfb3","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":139,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem environment_apply (a : Bool) (s : Set (Feedback 3)) (hs : MeasurableSet s) : environment a s=∑x:Sample, if feedback a x∈s then (1/64:ENNReal) else 0","missing":[],"search":"environment_apply banditrlproof.cucb.finiteexample.environment_apply theorem environment_apply (a : bool) (s : set (feedback 3)) (hs : measurableset s) : environment a s=∑x:sample, if feedback a x∈s then (1/64:ennreal) else 0 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.law_apply","label":"law_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.law_apply","description":"theorem law_apply (i : Fin 3) (s : Set UnitOutcome) (hs : MeasurableSet s) : law i s=∑x:Sample, if outcome x i∈s then (1/64:ENNReal) else 0","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-cb80ad211719","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":140,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem law_apply (i : Fin 3) (s : Set UnitOutcome) (hs : MeasurableSet s) : law i s=∑x:Sample, if outcome x i∈s then (1/64:ENNReal) else 0","missing":[],"search":"law_apply banditrlproof.cucb.finiteexample.law_apply theorem law_apply (i : fin 3) (s : set unitoutcome) (hs : measurableset s) : law i s=∑x:sample, if outcome x i∈s then (1/64:ennreal) else 0 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.environment_observed","label":"environment_observed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.environment_observed","description":"theorem environment_observed (a : Bool) (i : Fin 3) : environment a (observedSet i)=if i∈selected a then 1 else 1/2","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-3369370a0479","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":141,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem environment_observed (a : Bool) (i : Fin 3) : environment a (observedSet i)=if i∈selected a then 1 else 1/2","missing":[],"search":"environment_observed banditrlproof.cucb.finiteexample.environment_observed theorem environment_observed (a : bool) (i : fin 3) : environment a (observedset i)=if i∈selected a then 1 else 1/2 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.observation_compatible","label":"observation_compatible","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.observation_compatible","description":"theorem observation_compatible (a : Bool) (i : Fin 3) : ObservationCompatible (environment a) (law i) i","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-60af4c676177","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":142,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observation_compatible (a : Bool) (i : Fin 3) : ObservationCompatible (environment a) (law i) i","missing":[],"search":"observation_compatible banditrlproof.cucb.finiteexample.observation_compatible theorem observation_compatible (a : bool) (i : fin 3) : observationcompatible (environment a) (law i) i theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.reward_integrable","label":"reward_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.reward_integrable","description":"theorem reward_integrable (a : Bool) : Integrable (fun z => z.2.2) (environment a)","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-ea7bc6ec1fc1","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":143,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem reward_integrable (a : Bool) : Integrable (fun z => z.2.2) (environment a)","missing":[],"search":"reward_integrable banditrlproof.cucb.finiteexample.reward_integrable theorem reward_integrable (a : bool) : integrable (fun z => z.2.2) (environment a) theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.reward_nonneg","label":"reward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.reward_nonneg","description":"theorem reward_nonneg (a : Bool) : ∀ᵐ z ∂environment a, 0≤z.2.2","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-9dd4419295d4","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":144,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem reward_nonneg (a : Bool) : ∀ᵐ z ∂environment a, 0≤z.2.2","missing":[],"search":"reward_nonneg banditrlproof.cucb.finiteexample.reward_nonneg theorem reward_nonneg (a : bool) : ∀ᵐ z ∂environment a, 0≤z.2.2 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.model","label":"model","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.model","description":"noncomputable def model : FeedbackModel Bool 3 where","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-4e2fe8801c05","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":145,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def model : FeedbackModel Bool 3 where","missing":[],"search":"model banditrlproof.cucb.finiteexample.model noncomputable def model : feedbackmodel bool 3 where definition compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.mean_value","label":"mean_value","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.mean_value","description":"theorem mean_value (i : Fin 3) : marginalMean (law i)=if i=0 then 1/4 else if i=1 then 1/2 else 3/4","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-337e97dabc0b","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":146,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_value (i : Fin 3) : marginalMean (law i)=if i=0 then 1/4 else if i=1 then 1/2 else 3/4","missing":[],"search":"mean_value banditrlproof.cucb.finiteexample.mean_value theorem mean_value (i : fin 3) : marginalmean (law i)=if i=0 then 1/4 else if i=1 then 1/2 else 3/4 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.expected_reward","label":"expected_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.expected_reward","description":"theorem expected_reward (a : Bool) : model.expectedReward a=if a then 3/8 else 1/8","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-6ae140342f71","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":147,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_reward (a : Bool) : model.expectedReward a=if a then 3/8 else 1/8","missing":[],"search":"expected_reward banditrlproof.cucb.finiteexample.expected_reward theorem expected_reward (a : bool) : model.expectedreward a=if a then 3/8 else 1/8 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.law_one_mass","label":"law_one_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.law_one_mass","description":"theorem law_one_mass (i : Fin 3) : (law i {bit true}).toReal=if i=0 then 1/4 else if i=1 then 1/2 else 3/4","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-9b3d40eb9c5a","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":148,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:140"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem law_one_mass (i : Fin 3) : (law i {bit true}).toReal=if i=0 then 1/4 else if i=1 then 1/2 else 3/4","missing":[],"search":"law_one_mass banditrlproof.cucb.finiteexample.law_one_mass theorem law_one_mass (i : fin 3) : (law i {bit true}).toreal=if i=0 then 1/4 else if i=1 then 1/2 else 3/4 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.noisy_each_arm","label":"noisy_each_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.noisy_each_arm","description":"theorem noisy_each_arm (i : Fin 3) : 0<(law i {bit true}).toReal ∧ (law i {bit true}).toReal<1","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-edd14709751f","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":149,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:149"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem noisy_each_arm (i : Fin 3) : 0<(law i {bit true}).toReal ∧ (law i {bit true}).toReal<1","missing":[],"search":"noisy_each_arm banditrlproof.cucb.finiteexample.noisy_each_arm theorem noisy_each_arm (i : fin 3) : 0<(law i {bit true}).toreal ∧ (law i {bit true}).toreal<1 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.selected_size","label":"selected_size","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.selected_size","description":"theorem selected_size (a : Bool) : (selected a).card=2","url":"../modules/banditrlproof-algorithms-cucbfiniteexample/index.html#decl-9e1b58e898ca","parent":"module:BanditRLProof.Algorithms.CUCBFiniteExample","order":150,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteExample.lean:153"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem selected_size (a : Bool) : (selected a).card=2","missing":[],"search":"selected_size banditrlproof.cucb.finiteexample.selected_size theorem selected_size (a : bool) : (selected a).card=2 theorem compiled","shard":"modules/a7217debf832ac74.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.score","label":"score","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.score","description":"def score (v : Input 3) (a : Bool) : ℝ","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-e7bf1e4213f5","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":151,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def score (v : Input 3) (a : Bool) : ℝ","missing":[],"search":"score banditrlproof.cucb.finiteexample.score def score (v : input 3) (a : bool) : ℝ definition compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.choose","label":"choose","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.choose","description":"noncomputable def choose (v : Input 3) : Bool","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-a9291f6099c7","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":152,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def choose (v : Input 3) : Bool","missing":[],"search":"choose banditrlproof.cucb.finiteexample.choose noncomputable def choose (v : input 3) : bool definition compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.measurable_choose","label":"measurable_choose","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.measurable_choose","description":"theorem measurable_choose : Measurable choose","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-9d9affe2e0f4","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":153,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_choose : Measurable choose","missing":[],"search":"measurable_choose banditrlproof.cucb.finiteexample.measurable_choose theorem measurable_choose : measurable choose theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.choose_max","label":"choose_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.choose_max","description":"theorem choose_max (v : Input 3) : scoreOptimum score v≤score v (choose v)","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-8ded261c64d2","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":154,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem choose_max (v : Input 3) : scoreOptimum score v≤score v (choose v)","missing":[],"search":"choose_max banditrlproof.cucb.finiteexample.choose_max theorem choose_max (v : input 3) : scoreoptimum score v≤score v (choose v) theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.product_smooth","label":"product_smooth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.product_smooth","description":"theorem product_smooth (x y u v : UnitOutcome) (L : ℝ) (hL : 0≤L) (hx : |(x:ℝ)-u|≤L) (hy : |(y:ℝ)-v|≤L) : |(x:ℝ)*y-(u:ℝ)*v|≤2*L","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-7f83483b5ae4","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":155,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem product_smooth (x y u v : UnitOutcome) (L : ℝ) (hL : 0≤L) (hx : |(x:ℝ)-u|≤L) (hy : |(y:ℝ)-v|≤L) : |(x:ℝ)*y-(u:ℝ)*v|≤2*L","missing":[],"search":"product_smooth banditrlproof.cucb.finiteexample.product_smooth theorem product_smooth (x y u v : unitoutcome) (l : ℝ) (hl : 0≤l) (hx : |(x:ℝ)-u|≤l) (hy : |(y:ℝ)-v|≤l) : |(x:ℝ)*y-(u:ℝ)*v|≤2*l theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.source","label":"source","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.source","description":"noncomputable def source : SourceModel model where","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-a27192cb9432","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":156,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"noncomputable def source : SourceModel model where","missing":[],"search":"source banditrlproof.cucb.finiteexample.source noncomputable def source : sourcemodel model where definition compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.trigger_value","label":"trigger_value","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.trigger_value","description":"theorem trigger_value (a : Bool) (i : Fin 3) : model.triggerProbability a i=if i∈selected a then 1 else 1/2","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-542d0054e315","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":157,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_value (a : Bool) (i : Fin 3) : model.triggerProbability a i=if i∈selected a then 1 else 1/2","missing":[],"search":"trigger_value banditrlproof.cucb.finiteexample.trigger_value theorem trigger_value (a : bool) (i : fin 3) : model.triggerprobability a i=if i∈selected a then 1 else 1/2 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.minTrigger_value","label":"minTrigger_value","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.minTrigger_value","description":"theorem minTrigger_value (i : Fin 3) : model.minTrigger i=if i=1 then 1 else 1/2","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-14c4bb19002e","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":158,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem minTrigger_value (i : Fin 3) : model.minTrigger i=if i=1 then 1 else 1/2","missing":[],"search":"mintrigger_value banditrlproof.cucb.finiteexample.mintrigger_value theorem mintrigger_value (i : fin 3) : model.mintrigger i=if i=1 then 1 else 1/2 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.globalMinTrigger_value","label":"globalMinTrigger_value","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.globalMinTrigger_value","description":"theorem globalMinTrigger_value : model.globalMinTrigger=1/2","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-fb304a19b082","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":159,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:107"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem globalMinTrigger_value : model.globalMinTrigger=1/2","missing":[],"search":"globalmintrigger_value banditrlproof.cucb.finiteexample.globalmintrigger_value theorem globalmintrigger_value : model.globalmintrigger=1/2 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.true_score","label":"true_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.true_score","description":"theorem true_score (a : Bool) : score model.trueInput a=if a then 3/8 else 1/8","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-f47f2885727b","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":160,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:116"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem true_score (a : Bool) : score model.trueInput a=if a then 3/8 else 1/8","missing":[],"search":"true_score banditrlproof.cucb.finiteexample.true_score theorem true_score (a : bool) : score model.trueinput a=if a then 3/8 else 1/8 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.true_optimum","label":"true_optimum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.true_optimum","description":"theorem true_optimum : scoreOptimum score model.trueInput=3/8","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-1a20fe94a271","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":161,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:119"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem true_optimum : scoreOptimum score model.trueInput=3/8","missing":[],"search":"true_optimum banditrlproof.cucb.finiteexample.true_optimum theorem true_optimum : scoreoptimum score model.trueinput=3/8 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.max_gap","label":"max_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.max_gap","description":"theorem max_gap : maxPositiveGap source.score model.trueInput source.alpha=1/4","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-ad590e7b047e","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":162,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem max_gap : maxPositiveGap source.score model.trueInput source.alpha=1/4","missing":[],"search":"max_gap banditrlproof.cucb.finiteexample.max_gap theorem max_gap : maxpositivegap source.score model.trueinput source.alpha=1/4 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.choose_initial","label":"choose_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.choose_initial","description":"theorem choose_initial : choose (initialInput 3)=false","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-1424e48d7b22","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":163,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem choose_initial : choose (initialInput 3)=false","missing":[],"search":"choose_initial banditrlproof.cucb.finiteexample.choose_initial theorem choose_initial : choose (initialinput 3)=false theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.initial_action_gap","label":"initial_action_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.initial_action_gap","description":"theorem initial_action_gap : source.gap (choose (initialInput 3))=1/4","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-7901dd128022","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":164,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem initial_action_gap : source.gap (choose (initialInput 3))=1/4","missing":[],"search":"initial_action_gap banditrlproof.cucb.finiteexample.initial_action_gap theorem initial_action_gap : source.gap (choose (initialinput 3))=1/4 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.probabilistic_regret","label":"probabilistic_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.probabilistic_regret","description":"theorem probabilistic_regret (H : ℕ) (hH : 1≤H) : source.approximationRegret H≤ 4*(72*Real.log (H:ℝ))^((1:ℝ)/2)*(H:ℝ)^((1:ℝ)/2)+ (1+Real.pi^2/2)*(3/4)+30*Real.log (H:ℝ)","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-541cc855fcd4","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":165,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:147"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem probabilistic_regret (H : ℕ) (hH : 1≤H) : source.approximationRegret H≤ 4*(72*Real.log (H:ℝ))^((1:ℝ)/2)*(H:ℝ)^((1:ℝ)/2)+ (1+Real.pi^2/2)*(3/4)+30*Real.log (H:ℝ)","missing":[],"search":"probabilistic_regret banditrlproof.cucb.finiteexample.probabilistic_regret theorem probabilistic_regret (h : ℕ) (hh : 1≤h) : source.approximationregret h≤ 4*(72*real.log (h:ℝ))^((1:ℝ)/2)*(h:ℝ)^((1:ℝ)/2)+ (1+real.pi^2/2)*(3/4)+30*real.log (h:ℝ) theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.randomizedSource","label":"randomizedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.randomizedSource","description":"noncomputable def randomizedSource : SourceModel model","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-a451a181f5b8","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":166,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"noncomputable def randomizedSource : SourceModel model","missing":[],"search":"randomizedsource banditrlproof.cucb.finiteexample.randomizedsource noncomputable def randomizedsource : sourcemodel model definition compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.randomized_action_mass","label":"randomized_action_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.randomized_action_mass","description":"theorem randomized_action_mass (v : Input 3) (a : Bool) : randomizedSource.oracle v {a}=(1/2:ENNReal)","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-f5b5eea3a5c7","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":167,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:177"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem randomized_action_mass (v : Input 3) (a : Bool) : randomizedSource.oracle v {a}=(1/2:ENNReal)","missing":[],"search":"randomized_action_mass banditrlproof.cucb.finiteexample.randomized_action_mass theorem randomized_action_mass (v : input 3) (a : bool) : randomizedsource.oracle v {a}=(1/2:ennreal) theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.randomized_probabilistic_regret","label":"randomized_probabilistic_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.randomized_probabilistic_regret","description":"theorem randomized_probabilistic_regret (H : ℕ) (hH : 1≤H) : randomizedSource.approximationRegret H≤ 4*(72*Real.log (H:ℝ))^((1:ℝ)/2)*(H:ℝ)^((1:ℝ)/2)+ (1+Real.pi^2/2)*(3/4)+30*Real.log (H:ℝ)","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-dfa20517b854","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":168,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:183"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem randomized_probabilistic_regret (H : ℕ) (hH : 1≤H) : randomizedSource.approximationRegret H≤ 4*(72*Real.log (H:ℝ))^((1:ℝ)/2)*(H:ℝ)^((1:ℝ)/2)+ (1+Real.pi^2/2)*(3/4)+30*Real.log (H:ℝ)","missing":[],"search":"randomized_probabilistic_regret banditrlproof.cucb.finiteexample.randomized_probabilistic_regret theorem randomized_probabilistic_regret (h : ℕ) (hh : 1≤h) : randomizedsource.approximationregret h≤ 4*(72*real.log (h:ℝ))^((1:ℝ)/2)*(h:ℝ)^((1:ℝ)/2)+ (1+real.pi^2/2)*(3/4)+30*real.log (h:ℝ) theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.noBadSource","label":"noBadSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.noBadSource","description":"noncomputable def noBadSource : SourceModel model","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-3f38b4fafc93","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":169,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:197"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def noBadSource : SourceModel model","missing":[],"search":"nobadsource banditrlproof.cucb.finiteexample.nobadsource noncomputable def nobadsource : sourcemodel model definition compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.no_bad_regret","label":"no_bad_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.no_bad_regret","description":"theorem no_bad_regret (H : ℕ) : noBadSource.approximationRegret H≤0","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-cd8b7537d74e","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":170,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:211"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem no_bad_regret (H : ℕ) : noBadSource.approximationRegret H≤0","missing":[],"search":"no_bad_regret banditrlproof.cucb.finiteexample.no_bad_regret theorem no_bad_regret (h : ℕ) : nobadsource.approximationregret h≤0 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.first_action_law","label":"first_action_law","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.first_action_law","description":"theorem first_action_law : (cucbTrajectory source.oracle model.environment).map (fun Y => (Y 0).1)=Measure.dirac false","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-2e781816e5b0","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":171,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:218"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem first_action_law : (cucbTrajectory source.oracle model.environment).map (fun Y => (Y 0).1)=Measure.dirac false","missing":[],"search":"first_action_law banditrlproof.cucb.finiteexample.first_action_law theorem first_action_law : (cucbtrajectory source.oracle model.environment).map (fun y => (y 0).1)=measure.dirac false theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.regret_one_positive","label":"regret_one_positive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.regret_one_positive","description":"theorem regret_one_positive : source.approximationRegret 1=1/4","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-8853f0a7189b","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":172,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:227"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem regret_one_positive : source.approximationRegret 1=1/4","missing":[],"search":"regret_one_positive banditrlproof.cucb.finiteexample.regret_one_positive theorem regret_one_positive : source.approximationregret 1=1/4 theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.FiniteExample.refined_regret","label":"refined_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FiniteExample.refined_regret","description":"theorem refined_regret (H : ℕ) (hH : 1≤H) : source.approximationRegret H≤(∑i:Fin 3,source.armRefinedTerm H i)+(1+Real.pi^2/2)*(3/4)","url":"../modules/banditrlproof-algorithms-cucbfinitesourceexample/index.html#decl-1b40692cbf47","parent":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","order":173,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBFiniteSourceExample"],["Source","BanditRLProof/Algorithms/CUCBFiniteSourceExample.lean:244"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem refined_regret (H : ℕ) (hH : 1≤H) : source.approximationRegret H≤(∑i:Fin 3,source.armRefinedTerm H i)+(1+Real.pi^2/2)*(3/4)","missing":[],"search":"refined_regret banditrlproof.cucb.finiteexample.refined_regret theorem refined_regret (h : ℕ) (hh : 1≤h) : source.approximationregret h≤(∑i:fin 3,source.armrefinedterm h i)+(1+real.pi^2/2)*(3/4) theorem compiled","shard":"modules/f04f3933b1c8ff08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.card_underChargeTimes_le","label":"card_underChargeTimes_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.card_underChargeTimes_le","description":"theorem card_underChargeTimes_le (H : ℕ) (actions : ℕ → A) (i : Fin m) : (S.underChargeTimes H actions i).card≤S.chargeData.counters actions H i","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html#decl-68e02a0e8218","parent":"module:BanditRLProof.Algorithms.CUCBGapCutoff","order":174,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapCutoff"],["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem card_underChargeTimes_le (H : ℕ) (actions : ℕ → A) (i : Fin m) : (S.underChargeTimes H actions i).card≤S.chargeData.counters actions H i","missing":[],"search":"card_underchargetimes_le banditrlproof.cucb.sourcemodel.card_underchargetimes_le theorem card_underchargetimes_le (h : ℕ) (actions : ℕ → a) (i : fin m) : (s.underchargetimes h actions i).card≤s.chargedata.counters actions h i theorem compiled","shard":"modules/ca6b8cc77d816ca7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sum_card_underChargeTimes_le","label":"sum_card_underChargeTimes_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sum_card_underChargeTimes_le","description":"theorem sum_card_underChargeTimes_le (H : ℕ) (actions : ℕ → A) : (∑i:Fin m, ((S.underChargeTimes H actions i).card:ℝ))≤H","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html#decl-53bcb0f554a9","parent":"module:BanditRLProof.Algorithms.CUCBGapCutoff","order":175,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapCutoff"],["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_card_underChargeTimes_le (H : ℕ) (actions : ℕ → A) : (∑i:Fin m, ((S.underChargeTimes H actions i).card:ℝ))≤H","missing":[],"search":"sum_card_underchargetimes_le banditrlproof.cucb.sourcemodel.sum_card_underchargetimes_le theorem sum_card_underchargetimes_le (h : ℕ) (actions : ℕ → a) : (∑i:fin m, ((s.underchargetimes h actions i).card:ℝ))≤h theorem compiled","shard":"modules/ca6b8cc77d816ca7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight_le_cutoff","label":"underChargeWeight_le_cutoff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeWeight_le_cutoff","description":"theorem underChargeWeight_le_cutoff (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) (a : ℝ) (ha : a∈S.gapDomain) : S.underChargeWeight H actions i≤ a*((S.underChargeTimes H actions i).card:ℝ)+ (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)+ (maxPositiveGap S.score M.trueInput S.alpha-a)","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html#decl-b75c11ed0813","parent":"module:BanditRLProof.Algorithms.CUCBGapCutoff","order":176,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapCutoff"],["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underChargeWeight_le_cutoff (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) (a : ℝ) (ha : a∈S.gapDomain) : S.underChargeWeight H actions i≤ a*((S.underChargeTimes H actions i).card:ℝ)+ (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)+ (maxPositiveGap S.score M.trueInput S.alpha-a)","missing":[],"search":"underchargeweight_le_cutoff banditrlproof.cucb.sourcemodel.underchargeweight_le_cutoff theorem underchargeweight_le_cutoff (h : ℕ) (hh : 1≤h) (actions : ℕ → a) (i : fin m) (a : ℝ) (ha : a∈s.gapdomain) : s.underchargeweight h actions i≤ a*((s.underchargetimes h actions i).card:ℝ)+ (∫x in a..maxpositivegap s.score m.trueinput s.alpha, s.gapthreshold h (m.mintrigger i) x)+ (maxpositivegap s.score m.trueinput s.alpha-a) theorem compiled","shard":"modules/ca6b8cc77d816ca7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sum_underSampledGap_le_cutoff","label":"sum_underSampledGap_le_cutoff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sum_underSampledGap_le_cutoff","description":"theorem sum_underSampledGap_le_cutoff (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (a : ℝ) (ha : a∈S.gapDomain) : (∑t∈Finset.range H, S.underSampledGap H (S.chargeData.counters actions t) (actions t))≤ (H:ℝ)*a+(∑i:Fin m, ∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)+(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html#decl-a9b58d2ad3eb","parent":"module:BanditRLProof.Algorithms.CUCBGapCutoff","order":177,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapCutoff"],["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_underSampledGap_le_cutoff (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (a : ℝ) (ha : a∈S.gapDomain) : (∑t∈Finset.range H, S.underSampledGap H (S.chargeData.counters actions t) (actions t))≤ (H:ℝ)*a+(∑i:Fin m, ∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)+(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"sum_undersampledgap_le_cutoff banditrlproof.cucb.sourcemodel.sum_undersampledgap_le_cutoff theorem sum_undersampledgap_le_cutoff (h : ℕ) (hh : 1≤h) (actions : ℕ → a) (a : ℝ) (ha : a∈s.gapdomain) : (∑t∈finset.range h, s.undersampledgap h (s.chargedata.counters actions t) (actions t))≤ (h:ℝ)*a+(∑i:fin m, ∫x in a..maxpositivegap s.score m.trueinput s.alpha, s.gapthreshold h (m.mintrigger i) x)+(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/ca6b8cc77d816ca7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_gap_cutoff","label":"approximationRegret_le_gap_cutoff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_le_gap_cutoff","description":"theorem approximationRegret_le_gap_cutoff (H : ℕ) (hH : 1≤H) (a : ℝ) (ha : a∈S.gapDomain) : S.approximationRegret H≤(H:ℝ)*a+ (∑i:Fin m, ∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)+ (1+(2+(if M.globalMinTrigger<1 then 1 else 0))*Real.pi^2/6)* (m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html#decl-3fb8e782cbbc","parent":"module:BanditRLProof.Algorithms.CUCBGapCutoff","order":178,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapCutoff"],["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_le_gap_cutoff (H : ℕ) (hH : 1≤H) (a : ℝ) (ha : a∈S.gapDomain) : S.approximationRegret H≤(H:ℝ)*a+ (∑i:Fin m, ∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)+ (1+(2+(if M.globalMinTrigger<1 then 1 else 0))*Real.pi^2/6)* (m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"approximationregret_le_gap_cutoff banditrlproof.cucb.sourcemodel.approximationregret_le_gap_cutoff theorem approximationregret_le_gap_cutoff (h : ℕ) (hh : 1≤h) (a : ℝ) (ha : a∈s.gapdomain) : s.approximationregret h≤(h:ℝ)*a+ (∑i:fin m, ∫x in a..maxpositivegap s.score m.trueinput s.alpha, s.gapthreshold h (m.mintrigger i) x)+ (1+(2+(if m.globalmintrigger<1 then 1 else 0))*real.pi^2/6)* (m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/ca6b8cc77d816ca7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_large_cutoff","label":"approximationRegret_le_large_cutoff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_le_large_cutoff","description":"theorem approximationRegret_le_large_cutoff (H : ℕ) (a : ℝ) (ha : maxPositiveGap S.score M.trueInput S.alpha≤a) : S.approximationRegret H≤(H:ℝ)*a+ (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)* maxPositiveGap S.score M.trueInput S.alpha*Real.pi^2/6","url":"../modules/banditrlproof-algorithms-cucbgapcutoff/index.html#decl-2e80717ab378","parent":"module:BanditRLProof.Algorithms.CUCBGapCutoff","order":179,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapCutoff"],["Source","BanditRLProof/Algorithms/CUCBGapCutoff.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_le_large_cutoff (H : ℕ) (a : ℝ) (ha : maxPositiveGap S.score M.trueInput S.alpha≤a) : S.approximationRegret H≤(H:ℝ)*a+ (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)* maxPositiveGap S.score M.trueInput S.alpha*Real.pi^2/6","missing":[],"search":"approximationregret_le_large_cutoff banditrlproof.cucb.sourcemodel.approximationregret_le_large_cutoff theorem approximationregret_le_large_cutoff (h : ℕ) (a : ℝ) (ha : maxpositivegap s.score m.trueinput s.alpha≤a) : s.approximationregret h≤(h:ℝ)*a+ (2+(if m.globalmintrigger<1 then 1 else 0))*(m:ℝ)* maxpositivegap s.score m.trueinput s.alpha*real.pi^2/6 theorem compiled","shard":"modules/ca6b8cc77d816ca7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.thresholdCoefficient_antitone","label":"thresholdCoefficient_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.thresholdCoefficient_antitone","description":"theorem thresholdCoefficient_antitone {u v p : ℝ} (hu : 0<u) (huv : u≤v) (hp : 0<p) : thresholdCoefficient v p≤thresholdCoefficient u p","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-56f0d3970777","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":180,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem thresholdCoefficient_antitone {u v p : ℝ} (hu : 0<u) (huv : u≤v) (hp : 0<p) : thresholdCoefficient v p≤thresholdCoefficient u p","missing":[],"search":"thresholdcoefficient_antitone banditrlproof.cucb.thresholdcoefficient_antitone theorem thresholdcoefficient_antitone {u v p : ℝ} (hu : 0<u) (huv : u≤v) (hp : 0<p) : thresholdcoefficient v p≤thresholdcoefficient u p theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapDomain","label":"gapDomain","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapDomain","description":"noncomputable def gapDomain : Set ℝ","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-c0635890a2d0","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":181,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapDomain : Set ℝ","missing":[],"search":"gapdomain banditrlproof.cucb.sourcemodel.gapdomain noncomputable def gapdomain : set ℝ definition compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt","label":"inverseAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt","description":"The value outside the source inverse domain is zero by convention; all inverse and integral claims below explicitly stay inside the domain.","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-67214352ab5d","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":182,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseAt (d : ℝ) : ℝ","missing":[],"search":"inverseat banditrlproof.cucb.sourcemodel.inverseat the value outside the source inverse domain is zero by convention; all inverse and integral claims below explicitly stay inside the domain. definition compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_spec","label":"inverseAt_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_spec","description":"theorem inverseAt_spec {d : ℝ} (hd : d∈S.gapDomain) : 0<S.inverseAt d ∧ S.modulus (S.inverseAt d)=d","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-24271ab166df","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":183,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_spec {d : ℝ} (hd : d∈S.gapDomain) : 0<S.inverseAt d ∧ S.modulus (S.inverseAt d)=d","missing":[],"search":"inverseat_spec banditrlproof.cucb.sourcemodel.inverseat_spec theorem inverseat_spec {d : ℝ} (hd : d∈s.gapdomain) : 0<s.inverseat d ∧ s.modulus (s.inverseat d)=d theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_unique","label":"inverseAt_unique","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_unique","description":"theorem inverseAt_unique {d u : ℝ} (hd : d∈S.gapDomain) (hu : 0≤u) (he : S.modulus u=d) : S.inverseAt d=u","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-091a39cf0cae","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":184,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_unique {d u : ℝ} (hd : d∈S.gapDomain) (hu : 0≤u) (he : S.modulus u=d) : S.inverseAt d=u","missing":[],"search":"inverseat_unique banditrlproof.cucb.sourcemodel.inverseat_unique theorem inverseat_unique {d u : ℝ} (hd : d∈s.gapdomain) (hu : 0≤u) (he : s.modulus u=d) : s.inverseat d=u theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_gap","label":"inverseAt_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_gap","description":"theorem inverseAt_gap (a : A) (ha : 0<S.gap a) : S.inverseAt (S.gap a)=S.inverseGap a","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-831ce4eea82a","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":185,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_gap (a : A) (ha : 0<S.gap a) : S.inverseAt (S.gap a)=S.inverseGap a","missing":[],"search":"inverseat_gap banditrlproof.cucb.sourcemodel.inverseat_gap theorem inverseat_gap (a : a) (ha : 0<s.gap a) : s.inverseat (s.gap a)=s.inversegap a theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_strictMono","label":"inverseAt_strictMono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_strictMono","description":"theorem inverseAt_strictMono : StrictMonoOn S.inverseAt S.gapDomain","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-fb4ea99493ad","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":186,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_strictMono : StrictMonoOn S.inverseAt S.gapDomain","missing":[],"search":"inverseat_strictmono banditrlproof.cucb.sourcemodel.inverseat_strictmono theorem inverseat_strictmono : strictmonoon s.inverseat s.gapdomain theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold","label":"gapThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold","description":"noncomputable def gapThreshold (n : ℕ) (p d : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-9ff2340d1b1f","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":187,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapThreshold (n : ℕ) (p d : ℝ) : ℝ","missing":[],"search":"gapthreshold banditrlproof.cucb.sourcemodel.gapthreshold noncomputable def gapthreshold (n : ℕ) (p d : ℝ) : ℝ definition compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_antitone","label":"gapThreshold_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold_antitone","description":"theorem gapThreshold_antitone (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) : AntitoneOn (S.gapThreshold n p) S.gapDomain","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-67c8b5b65ca7","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":188,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_antitone (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) : AntitoneOn (S.gapThreshold n p) S.gapDomain","missing":[],"search":"gapthreshold_antitone banditrlproof.cucb.sourcemodel.gapthreshold_antitone theorem gapthreshold_antitone (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) : antitoneon (s.gapthreshold n p) s.gapdomain theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_nonneg","label":"gapThreshold_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold_nonneg","description":"theorem gapThreshold_nonneg (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) {d : ℝ} (hd : d∈S.gapDomain) : 0≤S.gapThreshold n p d","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-2ab17a017ea3","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":189,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_nonneg (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) {d : ℝ} (hd : d∈S.gapDomain) : 0≤S.gapThreshold n p d","missing":[],"search":"gapthreshold_nonneg banditrlproof.cucb.sourcemodel.gapthreshold_nonneg theorem gapthreshold_nonneg (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) {d : ℝ} (hd : d∈s.gapdomain) : 0≤s.gapthreshold n p d theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_intervalIntegrable","label":"gapThreshold_intervalIntegrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold_intervalIntegrable","description":"theorem gapThreshold_intervalIntegrable (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) {a b : ℝ} (ha : a∈S.gapDomain) (hb : b∈S.gapDomain) : IntervalIntegrable (S.gapThreshold n p) volume a b","url":"../modules/banditrlproof-algorithms-cucbgapinverse/index.html#decl-eeae49b6b2b4","parent":"module:BanditRLProof.Algorithms.CUCBGapInverse","order":190,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBGapInverse"],["Source","BanditRLProof/Algorithms/CUCBGapInverse.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_intervalIntegrable (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) {a b : ℝ} (ha : a∈S.gapDomain) (hb : b∈S.gapDomain) : IntervalIntegrable (S.gapThreshold n p) volume a b","missing":[],"search":"gapthreshold_intervalintegrable banditrlproof.cucb.sourcemodel.gapthreshold_intervalintegrable theorem gapthreshold_intervalintegrable (n : ℕ) (hn : 1≤n) (p : ℝ) (hp : 0<p) {a b : ℝ} (ha : a∈s.gapdomain) (hb : b∈s.gapdomain) : intervalintegrable (s.gapthreshold n p) volume a b theorem compiled","shard":"modules/b3d4b0f63cd17de0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.UnitOutcome","label":"UnitOutcome","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.CUCB.UnitOutcome","description":"abbrev UnitOutcome","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-0053c8711f1e","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":191,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev UnitOutcome","missing":[],"search":"unitoutcome banditrlproof.cucb.unitoutcome abbrev unitoutcome abbreviation compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.Feedback","label":"Feedback","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.CUCB.Feedback","description":"abbrev Feedback (m : ℕ)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-bb92bec781bd","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":192,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Feedback (m : ℕ)","missing":[],"search":"feedback banditrlproof.cucb.feedback abbrev feedback (m : ℕ) abbreviation compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observation","label":"observation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observation","description":"def observation {m : ℕ} (z : Feedback m) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-46506ae3b0c2","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":193,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def observation {m : ℕ} (z : Feedback m) (i : Fin m) : ℝ","missing":[],"search":"observation banditrlproof.cucb.observation def observation {m : ℕ} (z : feedback m) (i : fin m) : ℝ definition compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationCount","label":"observationCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observationCount","description":"def observationCount {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℕ","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-f7970b6450d5","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":194,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def observationCount {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℕ","missing":[],"search":"observationcount banditrlproof.cucb.observationcount def observationcount {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : ℕ definition compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationSum","label":"observationSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observationSum","description":"noncomputable def observationSum {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-81569169b050","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":195,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def observationSum {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℝ","missing":[],"search":"observationsum banditrlproof.cucb.observationsum noncomputable def observationsum {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : ℝ definition compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empiricalMean","label":"empiricalMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.empiricalMean","description":"noncomputable def empiricalMean {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-427efd4fba2e","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":196,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalMean {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℝ","missing":[],"search":"empiricalmean banditrlproof.cucb.empiricalmean noncomputable def empiricalmean {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : ℝ definition compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.upperIndex","label":"upperIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.upperIndex","description":"noncomputable def upperIndex {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-2122984c2036","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":197,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def upperIndex {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : ℝ","missing":[],"search":"upperindex banditrlproof.cucb.upperindex noncomputable def upperindex {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : ℝ definition compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observation_nonneg","label":"observation_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observation_nonneg","description":"theorem observation_nonneg {m : ℕ} (z : Feedback m) (i : Fin m) : 0≤observation z i","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-4d36d8da2a10","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":198,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observation_nonneg {m : ℕ} (z : Feedback m) (i : Fin m) : 0≤observation z i","missing":[],"search":"observation_nonneg banditrlproof.cucb.observation_nonneg theorem observation_nonneg {m : ℕ} (z : feedback m) (i : fin m) : 0≤observation z i theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationCount_succ","label":"observationCount_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observationCount_succ","description":"theorem observationCount_succ {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationCount Y (n+1) i = observationCount Y n i + if (Y n).1 i then 1 else 0","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-748ca2b84041","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":199,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observationCount_succ {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationCount Y (n+1) i = observationCount Y n i + if (Y n).1 i then 1 else 0","missing":[],"search":"observationcount_succ banditrlproof.cucb.observationcount_succ theorem observationcount_succ {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : observationcount y (n+1) i = observationcount y n i + if (y n).1 i then 1 else 0 theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationSum_succ","label":"observationSum_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observationSum_succ","description":"theorem observationSum_succ {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationSum Y (n+1) i = observationSum Y n i + observation (Y n) i","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-1ac53908158b","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":200,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observationSum_succ {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationSum Y (n+1) i = observationSum Y n i + observation (Y n) i","missing":[],"search":"observationsum_succ banditrlproof.cucb.observationsum_succ theorem observationsum_succ {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : observationsum y (n+1) i = observationsum y n i + observation (y n) i theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationSum_nonneg","label":"observationSum_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observationSum_nonneg","description":"theorem observationSum_nonneg {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : 0≤observationSum Y n i","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-2a7a27ad9ab3","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":201,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observationSum_nonneg {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : 0≤observationSum Y n i","missing":[],"search":"observationsum_nonneg banditrlproof.cucb.observationsum_nonneg theorem observationsum_nonneg {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : 0≤observationsum y n i theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observationSum_le_count","label":"observationSum_le_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observationSum_le_count","description":"theorem observationSum_le_count {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationSum Y n i ≤ (observationCount Y n i : ℝ)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-abc12067bea8","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":202,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observationSum_le_count {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : observationSum Y n i ≤ (observationCount Y n i : ℝ)","missing":[],"search":"observationsum_le_count banditrlproof.cucb.observationsum_le_count theorem observationsum_le_count {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : observationsum y n i ≤ (observationcount y n i : ℝ) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empiricalMean_le_one","label":"empiricalMean_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empiricalMean_le_one","description":"theorem empiricalMean_le_one {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : empiricalMean Y n i ≤ 1","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-d2bb16fb5658","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":203,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empiricalMean_le_one {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : empiricalMean Y n i ≤ 1","missing":[],"search":"empiricalmean_le_one banditrlproof.cucb.empiricalmean_le_one theorem empiricalmean_le_one {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : empiricalmean y n i ≤ 1 theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empiricalMean_unobserved","label":"empiricalMean_unobserved","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empiricalMean_unobserved","description":"theorem empiricalMean_unobserved {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (h : (Y n).1 i=false) : empiricalMean Y (n+1) i=empiricalMean Y n i","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-d5d0d6979d12","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":204,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empiricalMean_unobserved {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (h : (Y n).1 i=false) : empiricalMean Y (n+1) i=empiricalMean Y n i","missing":[],"search":"empiricalmean_unobserved banditrlproof.cucb.empiricalmean_unobserved theorem empiricalmean_unobserved {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) (h : (y n).1 i=false) : empiricalmean y (n+1) i=empiricalmean y n i theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empiricalMean_observed","label":"empiricalMean_observed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empiricalMean_observed","description":"theorem empiricalMean_observed {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (h : (Y n).1 i=true) : empiricalMean Y (n+1) i = (observationSum Y n i + ((Y n).2.1 i : ℝ)) / ((observationCount Y n i : ℝ)+1)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-83eeee334bbc","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":205,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empiricalMean_observed {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (h : (Y n).1 i=true) : empiricalMean Y (n+1) i = (observationSum Y n i + ((Y n).2.1 i : ℝ)) / ((observationCount Y n i : ℝ)+1)","missing":[],"search":"empiricalmean_observed banditrlproof.cucb.empiricalmean_observed theorem empiricalmean_observed {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) (h : (y n).1 i=true) : empiricalmean y (n+1) i = (observationsum y n i + ((y n).2.1 i : ℝ)) / ((observationcount y n i : ℝ)+1) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empiricalMean_nonneg","label":"empiricalMean_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empiricalMean_nonneg","description":"theorem empiricalMean_nonneg {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : 0≤empiricalMean Y n i","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-b75ec24cbdb3","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":206,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empiricalMean_nonneg {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : 0≤empiricalMean Y n i","missing":[],"search":"empiricalmean_nonneg banditrlproof.cucb.empiricalmean_nonneg theorem empiricalmean_nonneg {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : 0≤empiricalmean y n i theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.upperIndex_mem","label":"upperIndex_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.upperIndex_mem","description":"theorem upperIndex_mem {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : upperIndex Y n i ∈ Set.Icc (0:ℝ) 1","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-bc22d87d1024","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":207,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem upperIndex_mem {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) : upperIndex Y n i ∈ Set.Icc (0:ℝ) 1","missing":[],"search":"upperindex_mem banditrlproof.cucb.upperindex_mem theorem upperindex_mem {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) : upperindex y n i ∈ set.icc (0:ℝ) 1 theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.oracleInput","label":"oracleInput","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.oracleInput","description":"noncomputable def oracleInput {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) : Fin m → UnitOutcome","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-3fd84b00b705","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":208,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def oracleInput {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) : Fin m → UnitOutcome","missing":[],"search":"oracleinput banditrlproof.cucb.oracleinput noncomputable def oracleinput {m : ℕ} (y : ℕ → feedback m) (n : ℕ) : fin m → unitoutcome definition compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.statistics_causal","label":"statistics_causal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.statistics_causal","description":"theorem statistics_causal {m : ℕ} (Y Z : ℕ → Feedback m) (n : ℕ) (h : ∀t<n, Y t=Z t) (i : Fin m) : observationCount Y n i=observationCount Z n i ∧ observationSum Y n i=observationSum Z n i","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-629d84218e4b","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":209,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem statistics_causal {m : ℕ} (Y Z : ℕ → Feedback m) (n : ℕ) (h : ∀t<n, Y t=Z t) (i : Fin m) : observationCount Y n i=observationCount Z n i ∧ observationSum Y n i=observationSum Z n i","missing":[],"search":"statistics_causal banditrlproof.cucb.statistics_causal theorem statistics_causal {m : ℕ} (y z : ℕ → feedback m) (n : ℕ) (h : ∀t<n, y t=z t) (i : fin m) : observationcount y n i=observationcount z n i ∧ observationsum y n i=observationsum z n i theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.oracleInput_causal","label":"oracleInput_causal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.oracleInput_causal","description":"theorem oracleInput_causal {m : ℕ} (Y Z : ℕ → Feedback m) (n : ℕ) (h : ∀t<n, Y t=Z t) : oracleInput Y n=oracleInput Z n","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-c447500a3ca0","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":210,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oracleInput_causal {m : ℕ} (Y Z : ℕ → Feedback m) (n : ℕ) (h : ∀t<n, Y t=Z t) : oracleInput Y n=oracleInput Z n","missing":[],"search":"oracleinput_causal banditrlproof.cucb.oracleinput_causal theorem oracleinput_causal {m : ℕ} (y z : ℕ → feedback m) (n : ℕ) (h : ∀t<n, y t=z t) : oracleinput y n=oracleinput z n theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.oracleInput_visible","label":"oracleInput_visible","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.oracleInput_visible","description":"Masked latent values and the aggregate reward cannot leak into CUCB's choice: only matching masks and matching observed arm values are needed.","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-81980013282d","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":211,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:120"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oracleInput_visible {m : ℕ} (Y Z : ℕ → Feedback m) (n : ℕ) (hm : ∀t<n, ∀i, (Y t).1 i=(Z t).1 i) (hx : ∀t<n, ∀i, (Y t).1 i=true → (Y t).2.1 i=(Z t).2.1 i) : oracleInput Y n=oracleInput Z n","missing":[],"search":"oracleinput_visible banditrlproof.cucb.oracleinput_visible masked latent values and the aggregate reward cannot leak into cucb's choice: only matching masks and matching observed arm values are needed. theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_observation","label":"measurable_observation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_observation","description":"theorem measurable_observation {m : ℕ} (i : Fin m) : Measurable (fun z : Feedback m => observation z i)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-07b1c15474fd","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":212,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_observation {m : ℕ} (i : Fin m) : Measurable (fun z : Feedback m => observation z i)","missing":[],"search":"measurable_observation banditrlproof.cucb.measurable_observation theorem measurable_observation {m : ℕ} (i : fin m) : measurable (fun z : feedback m => observation z i) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_observationCount","label":"measurable_observationCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_observationCount","description":"theorem measurable_observationCount {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => observationCount Y n i)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-641578bc63ab","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":213,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:146"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_observationCount {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => observationCount Y n i)","missing":[],"search":"measurable_observationcount banditrlproof.cucb.measurable_observationcount theorem measurable_observationcount {m : ℕ} (n : ℕ) (i : fin m) : measurable (fun y : ℕ → feedback m => observationcount y n i) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_observationCount_real","label":"measurable_observationCount_real","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_observationCount_real","description":"theorem measurable_observationCount_real {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => (observationCount Y n i : ℝ))","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-6e3a59760400","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":214,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:154"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_observationCount_real {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => (observationCount Y n i : ℝ))","missing":[],"search":"measurable_observationcount_real banditrlproof.cucb.measurable_observationcount_real theorem measurable_observationcount_real {m : ℕ} (n : ℕ) (i : fin m) : measurable (fun y : ℕ → feedback m => (observationcount y n i : ℝ)) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_observationSum","label":"measurable_observationSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_observationSum","description":"theorem measurable_observationSum {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => observationSum Y n i)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-e85be8dca7a3","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":215,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:158"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_observationSum {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => observationSum Y n i)","missing":[],"search":"measurable_observationsum banditrlproof.cucb.measurable_observationsum theorem measurable_observationsum {m : ℕ} (n : ℕ) (i : fin m) : measurable (fun y : ℕ → feedback m => observationsum y n i) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_empiricalMean","label":"measurable_empiricalMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_empiricalMean","description":"theorem measurable_empiricalMean {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => empiricalMean Y n i)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-db5ba9f43925","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":216,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:165"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_empiricalMean {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => empiricalMean Y n i)","missing":[],"search":"measurable_empiricalmean banditrlproof.cucb.measurable_empiricalmean theorem measurable_empiricalmean {m : ℕ} (n : ℕ) (i : fin m) : measurable (fun y : ℕ → feedback m => empiricalmean y n i) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_upperIndex","label":"measurable_upperIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_upperIndex","description":"theorem measurable_upperIndex {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => upperIndex Y n i)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-4f51b33f0cf3","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":217,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:171"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_upperIndex {m : ℕ} (n : ℕ) (i : Fin m) : Measurable (fun Y : ℕ → Feedback m => upperIndex Y n i)","missing":[],"search":"measurable_upperindex banditrlproof.cucb.measurable_upperindex theorem measurable_upperindex {m : ℕ} (n : ℕ) (i : fin m) : measurable (fun y : ℕ → feedback m => upperindex y n i) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_oracleInput","label":"measurable_oracleInput","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_oracleInput","description":"theorem measurable_oracleInput {m : ℕ} (n : ℕ) : Measurable (fun Y : ℕ → Feedback m => oracleInput Y n)","url":"../modules/banditrlproof-algorithms-cucbhistory/index.html#decl-d6b7a81ac8e2","parent":"module:BanditRLProof.Algorithms.CUCBHistory","order":218,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBHistory"],["Source","BanditRLProof/Algorithms/CUCBHistory.lean:179"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_oracleInput {m : ℕ} (n : ℕ) : Measurable (fun Y : ℕ → Feedback m => oracleInput Y n)","missing":[],"search":"measurable_oracleinput banditrlproof.cucb.measurable_oracleinput theorem measurable_oracleinput {m : ℕ} (n : ℕ) : measurable (fun y : ℕ → feedback m => oracleinput y n) theorem compiled","shard":"modules/49fc793c1f05d1cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.confidenceRadius_lt_half","label":"confidenceRadius_lt_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.confidenceRadius_lt_half","description":"theorem confidenceRadius_lt_half {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (u : ℝ) (hu : 0<u) (hc : 6*Real.log ((n:ℝ)+1)/u^2 < (observationCount Y n i : ℝ)) : confidenceRadius Y n i (3*Real.log ((n:ℝ)+1)) < u/2","url":"../modules/banditrlproof-algorithms-cucbimpossiblecase/index.html#decl-fbc7b90da764","parent":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","order":219,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBImpossibleCase"],["Source","BanditRLProof/Algorithms/CUCBImpossibleCase.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem confidenceRadius_lt_half {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (u : ℝ) (hu : 0<u) (hc : 6*Real.log ((n:ℝ)+1)/u^2 < (observationCount Y n i : ℝ)) : confidenceRadius Y n i (3*Real.log ((n:ℝ)+1)) < u/2","missing":[],"search":"confidenceradius_lt_half banditrlproof.cucb.confidenceradius_lt_half theorem confidenceradius_lt_half {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) (u : ℝ) (hu : 0<u) (hc : 6*real.log ((n:ℝ)+1)/u^2 < (observationcount y n i : ℝ)) : confidenceradius y n i (3*real.log ((n:ℝ)+1)) < u/2 theorem compiled","shard":"modules/e948bf39a752d49c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.not_bad_of_nice_and_sufficient_observations","label":"not_bad_of_nice_and_sufficient_observations","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.not_bad_of_nice_and_sufficient_observations","description":"theorem not_bad_of_nice_and_sufficient_observations (Y : ℕ → Feedback m) (n : ℕ) (a : A) (hnice : ∀i, |empiricalMean Y n i-(M.trueInput i:ℝ)|≤ confidenceRadius Y n i (3*Real.log ((n:ℝ)+1))) (horacle : S.alpha*scoreOptimum S.score (oracleInput Y n)≤S.score (oracleInput Y n) a) (hcount : ∀i∈M.possible a, 6*Real.log ((n:ℝ)+1)/(S.inverseGap a)^2 < (observationCount Y n i : ℝ)) : ¬0<S.gap a","url":"../modules/banditrlproof-algorithms-cucbimpossiblecase/index.html#decl-3ba1d5ab9524","parent":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","order":220,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBImpossibleCase"],["Source","BanditRLProof/Algorithms/CUCBImpossibleCase.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem not_bad_of_nice_and_sufficient_observations (Y : ℕ → Feedback m) (n : ℕ) (a : A) (hnice : ∀i, |empiricalMean Y n i-(M.trueInput i:ℝ)|≤ confidenceRadius Y n i (3*Real.log ((n:ℝ)+1))) (horacle : S.alpha*scoreOptimum S.score (oracleInput Y n)≤S.score (oracleInput Y n) a) (hcount : ∀i∈M.possible a, 6*Real.log ((n:ℝ)+1)/(S.inverseGap a)^2 < (observationCount Y n i : ℝ)) : ¬0<S.gap a","missing":[],"search":"not_bad_of_nice_and_sufficient_observations banditrlproof.cucb.sourcemodel.not_bad_of_nice_and_sufficient_observations theorem not_bad_of_nice_and_sufficient_observations (y : ℕ → feedback m) (n : ℕ) (a : a) (hnice : ∀i, |empiricalmean y n i-(m.trueinput i:ℝ)|≤ confidenceradius y n i (3*real.log ((n:ℝ)+1))) (horacle : s.alpha*scoreoptimum s.score (oracleinput y n)≤s.score (oracleinput y n) a) (hcount : ∀i∈m.possible a, 6*real.log ((n:ℝ)+1)/(s.inversegap a)^2 < (observationcount y n i : ℝ)) : ¬0<s.gap a theorem compiled","shard":"modules/e948bf39a752d49c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.confidenceRadius","label":"confidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.confidenceRadius","description":"noncomputable def confidenceRadius {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (L : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-645df0c076fd","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":221,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceRadius {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (L : ℝ) : ℝ","missing":[],"search":"confidenceradius banditrlproof.cucb.confidenceradius noncomputable def confidenceradius {m : ℕ} (y : ℕ → feedback m) (n : ℕ) (i : fin m) (l : ℝ) : ℝ definition compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_confidenceRadius","label":"measurable_confidenceRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_confidenceRadius","description":"theorem measurable_confidenceRadius {m : ℕ} (n : ℕ) (i : Fin m) (L : ℝ) : Measurable (fun Y : ℕ → Feedback m => confidenceRadius Y n i L)","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-6343b221e85e","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":222,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_confidenceRadius {m : ℕ} (n : ℕ) (i : Fin m) (L : ℝ) : Measurable (fun Y : ℕ → Feedback m => confidenceRadius Y n i L)","missing":[],"search":"measurable_confidenceradius banditrlproof.cucb.measurable_confidenceradius theorem measurable_confidenceradius {m : ℕ} (n : ℕ) (i : fin m) (l : ℝ) : measurable (fun y : ℕ → feedback m => confidenceradius y n i l) theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.count_mul_radius","label":"count_mul_radius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.count_mul_radius","description":"theorem count_mul_radius (T L : ℝ) (hT : 0<T) (hL : 0≤L) : T*Real.sqrt (L/(2*T)) = Real.sqrt (T*L/2)","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-3566b1b0280b","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":223,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem count_mul_radius (T L : ℝ) (hT : 0<T) (hL : 0≤L) : T*Real.sqrt (L/(2*T)) = Real.sqrt (T*L/2)","missing":[],"search":"count_mul_radius banditrlproof.cucb.count_mul_radius theorem count_mul_radius (t l : ℝ) (ht : 0<t) (hl : 0≤l) : t*real.sqrt (l/(2*t)) = real.sqrt (t*l/2) theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empirical_bad_implies_deviation","label":"empirical_bad_implies_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empirical_bad_implies_deviation","description":"theorem empirical_bad_implies_deviation {A : Type*} {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (n : ℕ) (L : ℝ) (hL : 0≤L) (Y : ℕ → Round A m) (hbad : confidenceRadius (fun t => (Y t).2) n i L < |empiricalMean (fun t => (Y t).2) n i-marginalMean D|) : ∃ lower : Bool, 0<observationCount (fun t => (Y t).2) n i ∧ Real.sqrt ((observationCount (fun t => (Y t).2) n i : ℝ)*L/2) ≤ pathDeviation D…","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-4edd7b33c6c0","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":224,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empirical_bad_implies_deviation {A : Type*} {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (n : ℕ) (L : ℝ) (hL : 0≤L) (Y : ℕ → Round A m) (hbad : confidenceRadius (fun t => (Y t).2) n i L < |empiricalMean (fun t => (Y t).2) n i-marginalMean D|) : ∃ lower : Bool, 0<observationCount (fun t => (Y t).2) n i ∧ Real.sqrt ((observationCount (fun t => (Y t).2) n i : ℝ)*L/2) ≤ pathDeviation D i n lower Y","missing":[],"search":"empirical_bad_implies_deviation banditrlproof.cucb.empirical_bad_implies_deviation theorem empirical_bad_implies_deviation {a : type*} {m : ℕ} (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (n : ℕ) (l : ℝ) (hl : 0≤l) (y : ℕ → round a m) (hbad : confidenceradius (fun t => (y t).2) n i l < |empiricalmean (fun t => (y t).2) n i-marginalmean d|) : ∃ lower : bool, 0<observationcount (fun t => (y t).2) n i ∧ real.sqrt ((observationcount (fun t => (y t).2) n i : ℝ)*l/2) ≤ pathdeviation d i n lower y theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.upperIndex_of_confidence","label":"upperIndex_of_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.upperIndex_of_confidence","description":"On the source confidence event the clipped index is optimistic, and its excess above the true mean is at most twice the clipped confidence radius.","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-2a7277009c89","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":225,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem upperIndex_of_confidence {m : ℕ} (Y : ℕ → Feedback m) (n : ℕ) (i : Fin m) (μ : ℝ) (hμ : μ ∈ Set.Icc (0:ℝ) 1) (hgood : |empiricalMean Y n i-μ| ≤ confidenceRadius Y n i (3*Real.log ((n:ℝ)+1))) : μ ≤ upperIndex Y n i ∧ upperIndex Y n i ≤ μ+2*confidenceRadius Y n i (3*Real.log ((n:ℝ)+1))","missing":[],"search":"upperindex_of_confidence banditrlproof.cucb.upperindex_of_confidence on the source confidence event the clipped index is optimistic, and its excess above the true mean is at most twice the clipped confidence radius. theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empirical_bad_probability","label":"empirical_bad_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empirical_bad_probability","description":"theorem empirical_bad_probability (n : ℕ) (L : ℝ) (hL : 0≤L) : (cucbTrajectory oracle environment) {Y | confidenceRadius (fun t => (Y t).2) n i L < |empiricalMean (fun t => (Y t).2) n i-marginalMean D|} ≤ 2*(n:ENNReal)*ENNReal.ofReal (Real.exp (-L))","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-418ddf64dfad","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":226,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empirical_bad_probability (n : ℕ) (L : ℝ) (hL : 0≤L) : (cucbTrajectory oracle environment) {Y | confidenceRadius (fun t => (Y t).2) n i L < |empiricalMean (fun t => (Y t).2) n i-marginalMean D|} ≤ 2*(n:ENNReal)*ENNReal.ofReal (Real.exp (-L))","missing":[],"search":"empirical_bad_probability banditrlproof.cucb.empirical_bad_probability theorem empirical_bad_probability (n : ℕ) (l : ℝ) (hl : 0≤l) : (cucbtrajectory oracle environment) {y | confidenceradius (fun t => (y t).2) n i l < |empiricalmean (fun t => (y t).2) n i-marginalmean d|} ≤ 2*(n:ennreal)*ennreal.ofreal (real.exp (-l)) theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.round_confidence_arithmetic","label":"round_confidence_arithmetic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.round_confidence_arithmetic","description":"theorem round_confidence_arithmetic (n : ℕ) : 2*(n:ℝ)*Real.exp (-(3*Real.log ((n:ℝ)+1))) ≤ 2/((n:ℝ)+1)^2","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-3ec869dcca2f","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":227,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem round_confidence_arithmetic (n : ℕ) : 2*(n:ℝ)*Real.exp (-(3*Real.log ((n:ℝ)+1))) ≤ 2/((n:ℝ)+1)^2","missing":[],"search":"round_confidence_arithmetic banditrlproof.cucb.round_confidence_arithmetic theorem round_confidence_arithmetic (n : ℕ) : 2*(n:ℝ)*real.exp (-(3*real.log ((n:ℝ)+1))) ≤ 2/((n:ℝ)+1)^2 theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.empirical_bad_probability_source","label":"empirical_bad_probability_source","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.empirical_bad_probability_source","description":"theorem empirical_bad_probability_source (n : ℕ) : (cucbTrajectory oracle environment) {Y | confidenceRadius (fun t => (Y t).2) n i (3*Real.log ((n:ℝ)+1)) < |empiricalMean (fun t => (Y t).2) n i-marginalMean D|} ≤ ENNReal.ofReal (2/((n:ℝ)+1)^2)","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-10a01f502d63","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":228,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem empirical_bad_probability_source (n : ℕ) : (cucbTrajectory oracle environment) {Y | confidenceRadius (fun t => (Y t).2) n i (3*Real.log ((n:ℝ)+1)) < |empiricalMean (fun t => (Y t).2) n i-marginalMean D|} ≤ ENNReal.ofReal (2/((n:ℝ)+1)^2)","missing":[],"search":"empirical_bad_probability_source banditrlproof.cucb.empirical_bad_probability_source theorem empirical_bad_probability_source (n : ℕ) : (cucbtrajectory oracle environment) {y | confidenceradius (fun t => (y t).2) n i (3*real.log ((n:ℝ)+1)) < |empiricalmean (fun t => (y t).2) n i-marginalmean d|} ≤ ennreal.ofreal (2/((n:ℝ)+1)^2) theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.NiceEvent","label":"NiceEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.NiceEvent","description":"Source nice event, before round `n+1`, simultaneously over all arms.","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-2f4b48f6703b","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":229,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:159"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NiceEvent (means : Fin m → ℝ) (n : ℕ) : Set (ℕ → Round A m)","missing":[],"search":"niceevent banditrlproof.cucb.niceevent source nice event, before round `n+1`, simultaneously over all arms. definition compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurableSet_niceEvent","label":"measurableSet_niceEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurableSet_niceEvent","description":"theorem measurableSet_niceEvent (means : Fin m → ℝ) (n : ℕ) : MeasurableSet (NiceEvent (A:=A) means n)","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-2a320564a926","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":230,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:164"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_niceEvent (means : Fin m → ℝ) (n : ℕ) : MeasurableSet (NiceEvent (A:=A) means n)","missing":[],"search":"measurableset_niceevent banditrlproof.cucb.measurableset_niceevent theorem measurableset_niceevent (means : fin m → ℝ) (n : ℕ) : measurableset (niceevent (a:=a) means n) theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.niceEvent_complement_probability","label":"niceEvent_complement_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.niceEvent_complement_probability","description":"theorem niceEvent_complement_probability (laws : Fin m → Measure UnitOutcome) [∀i, IsProbabilityMeasure (laws i)] (hcompatAll : ∀a i, ObservationCompatible (environment a) (laws i) i) (n : ℕ) : (cucbTrajectory oracle environment) (NiceEvent (fun i => marginalMean (laws i)) n)ᶜ ≤ ENNReal.ofReal (2*(m:ℝ)/((n:ℝ)+1)^2)","url":"../modules/banditrlproof-algorithms-cucbniceevent/index.html#decl-2cdbb260c4f4","parent":"module:BanditRLProof.Algorithms.CUCBNiceEvent","order":231,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBNiceEvent"],["Source","BanditRLProof/Algorithms/CUCBNiceEvent.lean:175"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem niceEvent_complement_probability (laws : Fin m → Measure UnitOutcome) [∀i, IsProbabilityMeasure (laws i)] (hcompatAll : ∀a i, ObservationCompatible (environment a) (laws i) i) (n : ℕ) : (cucbTrajectory oracle environment) (NiceEvent (fun i => marginalMean (laws i)) n)ᶜ ≤ ENNReal.ofReal (2*(m:ℝ)/((n:ℝ)+1)^2)","missing":[],"search":"niceevent_complement_probability banditrlproof.cucb.niceevent_complement_probability theorem niceevent_complement_probability (laws : fin m → measure unitoutcome) [∀i, isprobabilitymeasure (laws i)] (hcompatall : ∀a i, observationcompatible (environment a) (laws i) i) (n : ℕ) : (cucbtrajectory oracle environment) (niceevent (fun i => marginalmean (laws i)) n)ᶜ ≤ ennreal.ofreal (2*(m:ℝ)/((n:ℝ)+1)^2) theorem compiled","shard":"modules/9a587b170484b277.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.marginalMean","label":"marginalMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.marginalMean","description":"noncomputable def marginalMean (D : Measure UnitOutcome) : ℝ","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-dea35d7aeeb1","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":232,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def marginalMean (D : Measure UnitOutcome) : ℝ","missing":[],"search":"marginalmean banditrlproof.cucb.marginalmean noncomputable def marginalmean (d : measure unitoutcome) : ℝ definition compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.centeredFactor","label":"centeredFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.centeredFactor","description":"noncomputable def centeredFactor (D : Measure UnitOutcome) (tilt : ℝ) (x : UnitOutcome) : ℝ","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-609538020f1c","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":233,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredFactor (D : Measure UnitOutcome) (tilt : ℝ) (x : UnitOutcome) : ℝ","missing":[],"search":"centeredfactor banditrlproof.cucb.centeredfactor noncomputable def centeredfactor (d : measure unitoutcome) (tilt : ℝ) (x : unitoutcome) : ℝ definition compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_centeredFactor","label":"measurable_centeredFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_centeredFactor","description":"theorem measurable_centeredFactor (D : Measure UnitOutcome) (tilt : ℝ) : Measurable (centeredFactor D tilt)","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-1166997fbb33","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":234,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_centeredFactor (D : Measure UnitOutcome) (tilt : ℝ) : Measurable (centeredFactor D tilt)","missing":[],"search":"measurable_centeredfactor banditrlproof.cucb.measurable_centeredfactor theorem measurable_centeredfactor (d : measure unitoutcome) (tilt : ℝ) : measurable (centeredfactor d tilt) theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.marginal_subgaussian","label":"marginal_subgaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.marginal_subgaussian","description":"theorem marginal_subgaussian (D : Measure UnitOutcome) [IsProbabilityMeasure D] : HasSubgaussianMGF (fun x : UnitOutcome => (x:ℝ)-marginalMean D) (1/4:ℝ≥0) D","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-5994561d6cad","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":235,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem marginal_subgaussian (D : Measure UnitOutcome) [IsProbabilityMeasure D] : HasSubgaussianMGF (fun x : UnitOutcome => (x:ℝ)-marginalMean D) (1/4:ℝ≥0) D","missing":[],"search":"marginal_subgaussian banditrlproof.cucb.marginal_subgaussian theorem marginal_subgaussian (d : measure unitoutcome) [isprobabilitymeasure d] : hassubgaussianmgf (fun x : unitoutcome => (x:ℝ)-marginalmean d) (1/4:ℝ≥0) d theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integrable_centeredFactor","label":"integrable_centeredFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integrable_centeredFactor","description":"theorem integrable_centeredFactor (D : Measure UnitOutcome) [IsProbabilityMeasure D] (tilt : ℝ) : Integrable (centeredFactor D tilt) D","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-0177b69abbb3","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":236,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_centeredFactor (D : Measure UnitOutcome) [IsProbabilityMeasure D] (tilt : ℝ) : Integrable (centeredFactor D tilt) D","missing":[],"search":"integrable_centeredfactor banditrlproof.cucb.integrable_centeredfactor theorem integrable_centeredfactor (d : measure unitoutcome) [isprobabilitymeasure d] (tilt : ℝ) : integrable (centeredfactor d tilt) d theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integral_centeredFactor_le_one","label":"integral_centeredFactor_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integral_centeredFactor_le_one","description":"theorem integral_centeredFactor_le_one (D : Measure UnitOutcome) [IsProbabilityMeasure D] (tilt : ℝ) : (∫ x, centeredFactor D tilt x ∂D)≤1","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-b3e1a488da3d","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":237,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_centeredFactor_le_one (D : Measure UnitOutcome) [IsProbabilityMeasure D] (tilt : ℝ) : (∫ x, centeredFactor D tilt x ∂D)≤1","missing":[],"search":"integral_centeredfactor_le_one banditrlproof.cucb.integral_centeredfactor_le_one theorem integral_centeredfactor_le_one (d : measure unitoutcome) [isprobabilitymeasure d] (tilt : ℝ) : (∫ x, centeredfactor d tilt x ∂d)≤1 theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedSet","label":"observedSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observedSet","description":"def observedSet {m : ℕ} (i : Fin m) : Set (Feedback m)","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-c8cce2f564f9","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":238,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def observedSet {m : ℕ} (i : Fin m) : Set (Feedback m)","missing":[],"search":"observedset banditrlproof.cucb.observedset def observedset {m : ℕ} (i : fin m) : set (feedback m) definition compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurableSet_observedSet","label":"measurableSet_observedSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurableSet_observedSet","description":"theorem measurableSet_observedSet {m : ℕ} (i : Fin m) : MeasurableSet (observedSet i)","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-5bf93ece9b54","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":239,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_observedSet {m : ℕ} (i : Fin m) : MeasurableSet (observedSet i)","missing":[],"search":"measurableset_observedset banditrlproof.cucb.measurableset_observedset theorem measurableset_observedset {m : ℕ} (i : fin m) : measurableset (observedset i) theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ObservationCompatible","label":"ObservationCompatible","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.ObservationCompatible","description":"Primitive equality of measures, expressing the unchanged marginal law upon observation. It is not a concentration or MGF assumption.","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-0477054fb77d","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":240,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def ObservationCompatible {m : ℕ} (ν : Measure (Feedback m)) (D : Measure UnitOutcome) (i : Fin m) : Prop","missing":[],"search":"observationcompatible banditrlproof.cucb.observationcompatible primitive equality of measures, expressing the unchanged marginal law upon observation. it is not a concentration or mgf assumption. definition compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observed_centeredFactor_integrable","label":"observed_centeredFactor_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observed_centeredFactor_integrable","description":"theorem observed_centeredFactor_integrable {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : IntegrableOn (fun z => centeredFactor D tilt (z.2.1 i)) (observedSet i) ν","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-075d7cd4dd1f","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":241,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observed_centeredFactor_integrable {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : IntegrableOn (fun z => centeredFactor D tilt (z.2.1 i)) (observedSet i) ν","missing":[],"search":"observed_centeredfactor_integrable banditrlproof.cucb.observed_centeredfactor_integrable theorem observed_centeredfactor_integrable {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (h : observationcompatible ν d i) (tilt : ℝ) : integrableon (fun z => centeredfactor d tilt (z.2.1 i)) (observedset i) ν theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observed_centeredFactor_integral","label":"observed_centeredFactor_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observed_centeredFactor_integral","description":"theorem observed_centeredFactor_integral {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : (∫ z in observedSet i, centeredFactor D tilt (z.2.1 i) ∂ν) = (ν (observedSet i)).toReal * ∫ x, centeredFactor D tilt x ∂D","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-a54ab18fc615","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":242,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observed_centeredFactor_integral {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : (∫ z in observedSet i, centeredFactor D tilt (z.2.1 i) ∂ν) = (ν (observedSet i)).toReal * ∫ x, centeredFactor D tilt x ∂D","missing":[],"search":"observed_centeredfactor_integral banditrlproof.cucb.observed_centeredfactor_integral theorem observed_centeredfactor_integral {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (h : observationcompatible ν d i) (tilt : ℝ) : (∫ z in observedset i, centeredfactor d tilt (z.2.1 i) ∂ν) = (ν (observedset i)).toreal * ∫ x, centeredfactor d tilt x ∂d theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedFactor","label":"observedFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.observedFactor","description":"noncomputable def observedFactor {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-566f17287b53","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":243,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def observedFactor {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : ℝ","missing":[],"search":"observedfactor banditrlproof.cucb.observedfactor noncomputable def observedfactor {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) (z : feedback m) : ℝ definition compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedFactor_eq_piecewise","label":"observedFactor_eq_piecewise","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observedFactor_eq_piecewise","description":"theorem observedFactor_eq_piecewise {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) [DecidablePred (· ∈ observedSet i)] : observedFactor D i tilt = (observedSet i).piecewise (fun z => centeredFactor D tilt (z.2.1 i)) (fun _ => 1)","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-1b30faddd570","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":244,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observedFactor_eq_piecewise {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) [DecidablePred (· ∈ observedSet i)] : observedFactor D i tilt = (observedSet i).piecewise (fun z => centeredFactor D tilt (z.2.1 i)) (fun _ => 1)","missing":[],"search":"observedfactor_eq_piecewise banditrlproof.cucb.observedfactor_eq_piecewise theorem observedfactor_eq_piecewise {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) [decidablepred (· ∈ observedset i)] : observedfactor d i tilt = (observedset i).piecewise (fun z => centeredfactor d tilt (z.2.1 i)) (fun _ => 1) theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedFactor_eq_exp","label":"observedFactor_eq_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observedFactor_eq_exp","description":"theorem observedFactor_eq_exp {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : observedFactor D i tilt z = Real.exp (tilt * (if z.1 i then (z.2.1 i:ℝ)-marginalMean D else 0) - tilt^2/8 * (if z.1 i then 1 else 0))","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-f55db120a028","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":245,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observedFactor_eq_exp {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : observedFactor D i tilt z = Real.exp (tilt * (if z.1 i then (z.2.1 i:ℝ)-marginalMean D else 0) - tilt^2/8 * (if z.1 i then 1 else 0))","missing":[],"search":"observedfactor_eq_exp banditrlproof.cucb.observedfactor_eq_exp theorem observedfactor_eq_exp {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) (z : feedback m) : observedfactor d i tilt z = real.exp (tilt * (if z.1 i then (z.2.1 i:ℝ)-marginalmean d else 0) - tilt^2/8 * (if z.1 i then 1 else 0)) theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integrable_observedFactor","label":"integrable_observedFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integrable_observedFactor","description":"theorem integrable_observedFactor {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : Integrable (observedFactor D i tilt) ν","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-00ef7f1d6e73","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":246,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:92"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_observedFactor {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : Integrable (observedFactor D i tilt) ν","missing":[],"search":"integrable_observedfactor banditrlproof.cucb.integrable_observedfactor theorem integrable_observedfactor {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (h : observationcompatible ν d i) (tilt : ℝ) : integrable (observedfactor d i tilt) ν theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integral_observedFactor_le_one","label":"integral_observedFactor_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integral_observedFactor_le_one","description":"theorem integral_observedFactor_le_one {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : (∫ z, observedFactor D i tilt z ∂ν)≤1","url":"../modules/banditrlproof-algorithms-cucbobservationmgf/index.html#decl-37ceae339d4e","parent":"module:BanditRLProof.Algorithms.CUCBObservationMGF","order":247,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBObservationMGF"],["Source","BanditRLProof/Algorithms/CUCBObservationMGF.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_observedFactor_le_one {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (h : ObservationCompatible ν D i) (tilt : ℝ) : (∫ z, observedFactor D i tilt z ∂ν)≤1","missing":[],"search":"integral_observedfactor_le_one banditrlproof.cucb.integral_observedfactor_le_one theorem integral_observedfactor_le_one {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (h : observationcompatible ν d i) (tilt : ℝ) : (∫ z, observedfactor d i tilt z ∂ν)≤1 theorem compiled","shard":"modules/559cf5080721c22b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.continuous_score","label":"continuous_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.continuous_score","description":"theorem continuous_score (a : A) : Continuous (fun v => S.score v a)","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html#decl-d510c938ac1e","parent":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","order":248,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleMeasurable"],["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem continuous_score (a : A) : Continuous (fun v => S.score v a)","missing":[],"search":"continuous_score banditrlproof.cucb.sourcemodel.continuous_score theorem continuous_score (a : a) : continuous (fun v => s.score v a) theorem compiled","shard":"modules/5aaf77eb214b613d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.continuous_optimum","label":"continuous_optimum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.continuous_optimum","description":"theorem continuous_optimum : Continuous (scoreOptimum S.score)","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html#decl-ee7c73bf00fb","parent":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","order":249,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleMeasurable"],["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem continuous_optimum : Continuous (scoreOptimum S.score)","missing":[],"search":"continuous_optimum banditrlproof.cucb.sourcemodel.continuous_optimum theorem continuous_optimum : continuous (scoreoptimum s.score) theorem compiled","shard":"modules/5aaf77eb214b613d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.measurable_joint_score","label":"measurable_joint_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.measurable_joint_score","description":"theorem measurable_joint_score : Measurable (fun p : Input m × A => S.score p.1 p.2)","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html#decl-6536ee5b41c8","parent":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","order":250,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleMeasurable"],["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_joint_score : Measurable (fun p : Input m × A => S.score p.1 p.2)","missing":[],"search":"measurable_joint_score banditrlproof.cucb.sourcemodel.measurable_joint_score theorem measurable_joint_score : measurable (fun p : input m × a => s.score p.1 p.2) theorem compiled","shard":"modules/5aaf77eb214b613d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.oracleSuccess","label":"oracleSuccess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.oracleSuccess","description":"def oracleSuccess : Set (Input m × A)","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html#decl-2b36ffa08d27","parent":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","order":251,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBOracleMeasurable"],["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def oracleSuccess : Set (Input m × A)","missing":[],"search":"oraclesuccess banditrlproof.cucb.sourcemodel.oraclesuccess def oraclesuccess : set (input m × a) definition compiled","shard":"modules/5aaf77eb214b613d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.measurableSet_oracleSuccess","label":"measurableSet_oracleSuccess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.measurableSet_oracleSuccess","description":"theorem measurableSet_oracleSuccess : MeasurableSet S.oracleSuccess","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html#decl-b4724e91ad2a","parent":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","order":252,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleMeasurable"],["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_oracleSuccess : MeasurableSet S.oracleSuccess","missing":[],"search":"measurableset_oraclesuccess banditrlproof.cucb.sourcemodel.measurableset_oraclesuccess theorem measurableset_oraclesuccess : measurableset s.oraclesuccess theorem compiled","shard":"modules/5aaf77eb214b613d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.measurableSet_path_oracleSuccess","label":"measurableSet_path_oracleSuccess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.measurableSet_path_oracleSuccess","description":"theorem measurableSet_path_oracleSuccess (n : ℕ) : MeasurableSet {Y : ℕ → Round A m | (oracleInput (fun t => (Y t).2) n, (Y n).1)∈S.oracleSuccess}","url":"../modules/banditrlproof-algorithms-cucboraclemeasurable/index.html#decl-1324649abddc","parent":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","order":253,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleMeasurable"],["Source","BanditRLProof/Algorithms/CUCBOracleMeasurable.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_path_oracleSuccess (n : ℕ) : MeasurableSet {Y : ℕ → Round A m | (oracleInput (fun t => (Y t).2) n, (Y n).1)∈S.oracleSuccess}","missing":[],"search":"measurableset_path_oraclesuccess banditrlproof.cucb.sourcemodel.measurableset_path_oraclesuccess theorem measurableset_path_oraclesuccess (n : ℕ) : measurableset {y : ℕ → round a m | (oracleinput (fun t => (y t).2) n, (y n).1)∈s.oraclesuccess} theorem compiled","shard":"modules/5aaf77eb214b613d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.successIndicator","label":"successIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.successIndicator","description":"noncomputable def successIndicator (v : Input m) (z : Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-687ea271154e","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":254,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def successIndicator (v : Input m) (z : Round A m) : ℝ","missing":[],"search":"successindicator banditrlproof.cucb.sourcemodel.successindicator noncomputable def successindicator (v : input m) (z : round a m) : ℝ definition compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.measurable_successIndicator","label":"measurable_successIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.measurable_successIndicator","description":"theorem measurable_successIndicator : Measurable (fun p : Input m × Round A m => S.successIndicator p.1 p.2)","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-066f0cfe496f","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":255,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_successIndicator : Measurable (fun p : Input m × Round A m => S.successIndicator p.1 p.2)","missing":[],"search":"measurable_successindicator banditrlproof.cucb.sourcemodel.measurable_successindicator theorem measurable_successindicator : measurable (fun p : input m × round a m => s.successindicator p.1 p.2) theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.successIndicator_mem","label":"successIndicator_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.successIndicator_mem","description":"theorem successIndicator_mem (v : Input m) (z : Round A m) : S.successIndicator v z∈Set.Icc (0:ℝ) 1","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-640d9ac8cc86","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":256,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem successIndicator_mem (v : Input m) (z : Round A m) : S.successIndicator v z∈Set.Icc (0:ℝ) 1","missing":[],"search":"successindicator_mem banditrlproof.cucb.sourcemodel.successindicator_mem theorem successindicator_mem (v : input m) (z : round a m) : s.successindicator v z∈set.icc (0:ℝ) 1 theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_successIndicator_comp","label":"integrable_successIndicator_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_successIndicator_comp","description":"theorem integrable_successIndicator_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → Input m × Round A m) (hg : Measurable g) : Integrable (fun ω => S.successIndicator (g ω).1 (g ω).2) ν","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-f743cfe15ec2","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":257,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_successIndicator_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → Input m × Round A m) (hg : Measurable g) : Integrable (fun ω => S.successIndicator (g ω).1 (g ω).2) ν","missing":[],"search":"integrable_successindicator_comp banditrlproof.cucb.sourcemodel.integrable_successindicator_comp theorem integrable_successindicator_comp {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] (g : ω → input m × round a m) (hg : measurable g) : integrable (fun ω => s.successindicator (g ω).1 (g ω).2) ν theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.roundKernel_success_lower","label":"roundKernel_success_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.roundKernel_success_lower","description":"theorem roundKernel_success_lower (v : Input m) : S.beta≤∫z, S.successIndicator v z ∂roundKernel S.oracle M.environment v","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-4a60c86d7569","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":258,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem roundKernel_success_lower (v : Input m) : S.beta≤∫z, S.successIndicator v z ∂roundKernel S.oracle M.environment v","missing":[],"search":"roundkernel_success_lower banditrlproof.cucb.sourcemodel.roundkernel_success_lower theorem roundkernel_success_lower (v : input m) : s.beta≤∫z, s.successindicator v z ∂roundkernel s.oracle m.environment v theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.condExp_oracle_success","label":"condExp_oracle_success","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.condExp_oracle_success","description":"theorem condExp_oracle_success (n : ℕ) : (fun _ => S.beta) ≤ᵐ[cucbTrajectory S.oracle M.environment] (cucbTrajectory S.oracle M.environment)[fun Y => S.successIndicator (oracleInput (fun t => (Y t).2) (n+1)) (Y (n+1)) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance]","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-b0d75f0127a6","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":259,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem condExp_oracle_success (n : ℕ) : (fun _ => S.beta) ≤ᵐ[cucbTrajectory S.oracle M.environment] (cucbTrajectory S.oracle M.environment)[fun Y => S.successIndicator (oracleInput (fun t => (Y t).2) (n+1)) (Y (n+1)) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance]","missing":[],"search":"condexp_oracle_success banditrlproof.cucb.sourcemodel.condexp_oracle_success theorem condexp_oracle_success (n : ℕ) : (fun _ => s.beta) ≤ᵐ[cucbtrajectory s.oracle m.environment] (cucbtrajectory s.oracle m.environment)[fun y => s.successindicator (oracleinput (fun t => (y t).2) (n+1)) (y (n+1)) | measurablespace.comap (preorder.frestrictle n) inferinstance] theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.initial_oracle_success","label":"initial_oracle_success","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.initial_oracle_success","description":"theorem initial_oracle_success : S.beta≤∫Y, S.successIndicator (oracleInput (fun t => (Y t).2) 0) (Y 0) ∂cucbTrajectory S.oracle M.environment","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-69835b18d5cd","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":260,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem initial_oracle_success : S.beta≤∫Y, S.successIndicator (oracleInput (fun t => (Y t).2) 0) (Y 0) ∂cucbTrajectory S.oracle M.environment","missing":[],"search":"initial_oracle_success banditrlproof.cucb.sourcemodel.initial_oracle_success theorem initial_oracle_success : s.beta≤∫y, s.successindicator (oracleinput (fun t => (y t).2) 0) (y 0) ∂cucbtrajectory s.oracle m.environment theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.expected_oracle_success","label":"expected_oracle_success","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.expected_oracle_success","description":"theorem expected_oracle_success (n : ℕ) : S.beta≤∫Y, S.successIndicator (oracleInput (fun t => (Y t).2) n) (Y n) ∂cucbTrajectory S.oracle M.environment","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-c051deb8aa7c","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":261,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_oracle_success (n : ℕ) : S.beta≤∫Y, S.successIndicator (oracleInput (fun t => (Y t).2) n) (Y n) ∂cucbTrajectory S.oracle M.environment","missing":[],"search":"expected_oracle_success banditrlproof.cucb.sourcemodel.expected_oracle_success theorem expected_oracle_success (n : ℕ) : s.beta≤∫y, s.successindicator (oracleinput (fun t => (y t).2) n) (y n) ∂cucbtrajectory s.oracle m.environment theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.oracle_failure_probability","label":"oracle_failure_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.oracle_failure_probability","description":"theorem oracle_failure_probability (n : ℕ) : (cucbTrajectory S.oracle M.environment) {Y | (oracleInput (fun t => (Y t).2) n,(Y n).1)∈S.oracleSuccess}ᶜ ≤ ENNReal.ofReal (1-S.beta)","url":"../modules/banditrlproof-algorithms-cucboraclesuccess/index.html#decl-16eef553145b","parent":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","order":262,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBOracleSuccess"],["Source","BanditRLProof/Algorithms/CUCBOracleSuccess.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oracle_failure_probability (n : ℕ) : (cucbTrajectory S.oracle M.environment) {Y | (oracleInput (fun t => (Y t).2) n,(Y n).1)∈S.oracleSuccess}ᶜ ≤ ENNReal.ofReal (1-S.beta)","missing":[],"search":"oracle_failure_probability banditrlproof.cucb.sourcemodel.oracle_failure_probability theorem oracle_failure_probability (n : ℕ) : (cucbtrajectory s.oracle m.environment) {y | (oracleinput (fun t => (y t).2) n,(y n).1)∈s.oraclesuccess}ᶜ ≤ ennreal.ofreal (1-s.beta) theorem compiled","shard":"modules/9720d8ff492c808f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.threshold_integral_le_power","label":"threshold_integral_le_power","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.threshold_integral_le_power","description":"theorem threshold_integral_le_power (H : ℕ) (hH : 1≤H) (i : Fin m) (a : ℝ) (ha : a∈S.gapDomain) (q B c : ℝ) (hq : 1<q) (hB : 0≤B) (hc : 0≤c) (he : ∀x∈Set.Icc a (maxPositiveGap S.score M.trueInput S.alpha), S.gapThreshold H (M.minTrigger i) x≤B*x^(-q)+c) : (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)≤ B*a^(1-q)/(q-1)+c*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbpolynomialintegral/index.html#decl-b79dac9cd5cc","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","order":263,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialIntegral"],["Source","BanditRLProof/Algorithms/CUCBPolynomialIntegral.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem threshold_integral_le_power (H : ℕ) (hH : 1≤H) (i : Fin m) (a : ℝ) (ha : a∈S.gapDomain) (q B c : ℝ) (hq : 1<q) (hB : 0≤B) (hc : 0≤c) (he : ∀x∈Set.Icc a (maxPositiveGap S.score M.trueInput S.alpha), S.gapThreshold H (M.minTrigger i) x≤B*x^(-q)+c) : (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)≤ B*a^(1-q)/(q-1)+c*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"threshold_integral_le_power banditrlproof.cucb.sourcemodel.threshold_integral_le_power theorem threshold_integral_le_power (h : ℕ) (hh : 1≤h) (i : fin m) (a : ℝ) (ha : a∈s.gapdomain) (q b c : ℝ) (hq : 1<q) (hb : 0≤b) (hc : 0≤c) (he : ∀x∈set.icc a (maxpositivegap s.score m.trueinput s.alpha), s.gapthreshold h (m.mintrigger i) x≤b*x^(-q)+c) : (∫x in a..maxpositivegap s.score m.trueinput s.alpha, s.gapthreshold h (m.mintrigger i) x)≤ b*a^(1-q)/(q-1)+c*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/6d2a3d7dc92f5f06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_threshold_integral_deterministic","label":"polynomial_threshold_integral_deterministic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.polynomial_threshold_integral_deterministic","description":"theorem polynomial_threshold_integral_deterministic (H : ℕ) (hH : 1≤H) (i : Fin m) (hp : M.minTrigger i=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)≤ (6*Real.log (H:ℝ)*γ^(2/ω))*a^(1-2/ω)/(2/ω-1)","url":"../modules/banditrlproof-algorithms-cucbpolynomialintegral/index.html#decl-a60d88217877","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","order":264,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialIntegral"],["Source","BanditRLProof/Algorithms/CUCBPolynomialIntegral.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem polynomial_threshold_integral_deterministic (H : ℕ) (hH : 1≤H) (i : Fin m) (hp : M.minTrigger i=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)≤ (6*Real.log (H:ℝ)*γ^(2/ω))*a^(1-2/ω)/(2/ω-1)","missing":[],"search":"polynomial_threshold_integral_deterministic banditrlproof.cucb.sourcemodel.polynomial_threshold_integral_deterministic theorem polynomial_threshold_integral_deterministic (h : ℕ) (hh : 1≤h) (i : fin m) (hp : m.mintrigger i=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) (a : ℝ) (ha : a∈s.gapdomain) : (∫x in a..maxpositivegap s.score m.trueinput s.alpha, s.gapthreshold h (m.mintrigger i) x)≤ (6*real.log (h:ℝ)*γ^(2/ω))*a^(1-2/ω)/(2/ω-1) theorem compiled","shard":"modules/6d2a3d7dc92f5f06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_threshold_integral_probabilistic","label":"polynomial_threshold_integral_probabilistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.polynomial_threshold_integral_probabilistic","description":"theorem polynomial_threshold_integral_probabilistic (H : ℕ) (hH : 1≤H) (i : Fin m) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)≤ (12*Real.log (H:ℝ)/M.globalMinTrigger*γ^(2/ω))*a^(1-2/ω)/(2/ω-1)+ (24*Real.log (H:ℝ)/M.minTrigger i)*maxPositiveGap S.score M.trueInpu…","url":"../modules/banditrlproof-algorithms-cucbpolynomialintegral/index.html#decl-42f56398ee67","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","order":265,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialIntegral"],["Source","BanditRLProof/Algorithms/CUCBPolynomialIntegral.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem polynomial_threshold_integral_probabilistic (H : ℕ) (hH : 1≤H) (i : Fin m) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : (∫x in a..maxPositiveGap S.score M.trueInput S.alpha, S.gapThreshold H (M.minTrigger i) x)≤ (12*Real.log (H:ℝ)/M.globalMinTrigger*γ^(2/ω))*a^(1-2/ω)/(2/ω-1)+ (24*Real.log (H:ℝ)/M.minTrigger i)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"polynomial_threshold_integral_probabilistic banditrlproof.cucb.sourcemodel.polynomial_threshold_integral_probabilistic theorem polynomial_threshold_integral_probabilistic (h : ℕ) (hh : 1≤h) (i : fin m) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) (a : ℝ) (ha : a∈s.gapdomain) : (∫x in a..maxpositivegap s.score m.trueinput s.alpha, s.gapthreshold h (m.mintrigger i) x)≤ (12*real.log (h:ℝ)/m.globalmintrigger*γ^(2/ω))*a^(1-2/ω)/(2/ω-1)+ (24*real.log (h:ℝ)/m.mintrigger i)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/6d2a3d7dc92f5f06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_cutoff_regret_deterministic","label":"polynomial_cutoff_regret_deterministic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.polynomial_cutoff_regret_deterministic","description":"theorem polynomial_cutoff_regret_deterministic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : S.approximationRegret H≤(H:ℝ)*a+ ((m:ℝ)*(6*Real.log (H:ℝ)*γ^(2/ω)))*a^(1-2/ω)/(2/ω-1)+ (1+Real.pi^2/3)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbpolynomialintegral/index.html#decl-722ecdf1519e","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","order":266,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialIntegral"],["Source","BanditRLProof/Algorithms/CUCBPolynomialIntegral.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem polynomial_cutoff_regret_deterministic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : S.approximationRegret H≤(H:ℝ)*a+ ((m:ℝ)*(6*Real.log (H:ℝ)*γ^(2/ω)))*a^(1-2/ω)/(2/ω-1)+ (1+Real.pi^2/3)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"polynomial_cutoff_regret_deterministic banditrlproof.cucb.sourcemodel.polynomial_cutoff_regret_deterministic theorem polynomial_cutoff_regret_deterministic (h : ℕ) (hh : 1≤h) (hp : m.globalmintrigger=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) (a : ℝ) (ha : a∈s.gapdomain) : s.approximationregret h≤(h:ℝ)*a+ ((m:ℝ)*(6*real.log (h:ℝ)*γ^(2/ω)))*a^(1-2/ω)/(2/ω-1)+ (1+real.pi^2/3)*(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/6d2a3d7dc92f5f06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_cutoff_regret_probabilistic","label":"polynomial_cutoff_regret_probabilistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.polynomial_cutoff_regret_probabilistic","description":"theorem polynomial_cutoff_regret_probabilistic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger<1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : S.approximationRegret H≤(H:ℝ)*a+ ((m:ℝ)*(12*Real.log (H:ℝ)/M.globalMinTrigger*γ^(2/ω)))*a^(1-2/ω)/(2/ω-1)+ (1+Real.pi^2/2)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha+ ∑i:Fin m, (24*Real.log (H:ℝ)/M.minTrigger…","url":"../modules/banditrlproof-algorithms-cucbpolynomialintegral/index.html#decl-0fdb3bc9864d","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","order":267,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialIntegral"],["Source","BanditRLProof/Algorithms/CUCBPolynomialIntegral.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem polynomial_cutoff_regret_probabilistic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger<1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) (a : ℝ) (ha : a∈S.gapDomain) : S.approximationRegret H≤(H:ℝ)*a+ ((m:ℝ)*(12*Real.log (H:ℝ)/M.globalMinTrigger*γ^(2/ω)))*a^(1-2/ω)/(2/ω-1)+ (1+Real.pi^2/2)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha+ ∑i:Fin m, (24*Real.log (H:ℝ)/M.minTrigger i)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"polynomial_cutoff_regret_probabilistic banditrlproof.cucb.sourcemodel.polynomial_cutoff_regret_probabilistic theorem polynomial_cutoff_regret_probabilistic (h : ℕ) (hh : 1≤h) (hp : m.globalmintrigger<1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) (a : ℝ) (ha : a∈s.gapdomain) : s.approximationregret h≤(h:ℝ)*a+ ((m:ℝ)*(12*real.log (h:ℝ)/m.globalmintrigger*γ^(2/ω)))*a^(1-2/ω)/(2/ω-1)+ (1+real.pi^2/2)*(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha+ ∑i:fin m, (24*real.log (h:ℝ)/m.mintrigger i)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/6d2a3d7dc92f5f06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.base_arm_count_pos","label":"base_arm_count_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.base_arm_count_pos","description":"theorem base_arm_count_pos : 0<m","url":"../modules/banditrlproof-algorithms-cucbpolynomialregret/index.html#decl-5f9ae0a386ae","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","order":268,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialRegret"],["Source","BanditRLProof/Algorithms/CUCBPolynomialRegret.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem base_arm_count_pos : 0<m","missing":[],"search":"base_arm_count_pos banditrlproof.cucb.sourcemodel.base_arm_count_pos theorem base_arm_count_pos : 0<m theorem compiled","shard":"modules/3e4ea64c4acc8035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_linear_gap","label":"approximationRegret_le_linear_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_le_linear_gap","description":"theorem approximationRegret_le_linear_gap (H : ℕ) : S.approximationRegret H≤(H:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbpolynomialregret/index.html#decl-59a06c88ba3e","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","order":269,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialRegret"],["Source","BanditRLProof/Algorithms/CUCBPolynomialRegret.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_le_linear_gap (H : ℕ) : S.approximationRegret H≤(H:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"approximationregret_le_linear_gap banditrlproof.cucb.sourcemodel.approximationregret_le_linear_gap theorem approximationregret_le_linear_gap (h : ℕ) : s.approximationregret h≤(h:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/3e4ea64c4acc8035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_one_le","label":"approximationRegret_one_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_one_le","description":"theorem approximationRegret_one_le (c : ℝ) (hc : 1≤c) : S.approximationRegret 1≤c*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbpolynomialregret/index.html#decl-75b13edf1150","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","order":270,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialRegret"],["Source","BanditRLProof/Algorithms/CUCBPolynomialRegret.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_one_le (c : ℝ) (hc : 1≤c) : S.approximationRegret 1≤c*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"approximationregret_one_le banditrlproof.cucb.sourcemodel.approximationregret_one_le theorem approximationregret_one_le (c : ℝ) (hc : 1≤c) : s.approximationregret 1≤c*(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/3e4ea64c4acc8035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.theorem_two_deterministic","label":"theorem_two_deterministic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.theorem_two_deterministic","description":"theorem theorem_two_deterministic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) : S.approximationRegret H≤(2*γ/(2-ω))*(6*(m:ℝ)*Real.log (H:ℝ))^(ω/2)*(H:ℝ)^(1-ω/2)+ (1+Real.pi^2/3)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbpolynomialregret/index.html#decl-67a98522513e","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","order":271,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialRegret"],["Source","BanditRLProof/Algorithms/CUCBPolynomialRegret.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem theorem_two_deterministic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) : S.approximationRegret H≤(2*γ/(2-ω))*(6*(m:ℝ)*Real.log (H:ℝ))^(ω/2)*(H:ℝ)^(1-ω/2)+ (1+Real.pi^2/3)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"theorem_two_deterministic banditrlproof.cucb.sourcemodel.theorem_two_deterministic theorem theorem_two_deterministic (h : ℕ) (hh : 1≤h) (hp : m.globalmintrigger=1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) : s.approximationregret h≤(2*γ/(2-ω))*(6*(m:ℝ)*real.log (h:ℝ))^(ω/2)*(h:ℝ)^(1-ω/2)+ (1+real.pi^2/3)*(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/3e4ea64c4acc8035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.theorem_two_probabilistic","label":"theorem_two_probabilistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.theorem_two_probabilistic","description":"theorem theorem_two_probabilistic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger<1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) : S.approximationRegret H≤(2*γ/(2-ω))*(12*(m:ℝ)*Real.log (H:ℝ)/M.globalMinTrigger)^(ω/2)* (H:ℝ)^(1-ω/2)+(1+Real.pi^2/2)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha+ ∑i:Fin m, (24*Real.log (H:ℝ)/M.minTrigger i)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbpolynomialregret/index.html#decl-aa0343d755b7","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","order":272,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialRegret"],["Source","BanditRLProof/Algorithms/CUCBPolynomialRegret.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem theorem_two_probabilistic (H : ℕ) (hH : 1≤H) (hp : M.globalMinTrigger<1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) : S.approximationRegret H≤(2*γ/(2-ω))*(12*(m:ℝ)*Real.log (H:ℝ)/M.globalMinTrigger)^(ω/2)* (H:ℝ)^(1-ω/2)+(1+Real.pi^2/2)*(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha+ ∑i:Fin m, (24*Real.log (H:ℝ)/M.minTrigger i)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"theorem_two_probabilistic banditrlproof.cucb.sourcemodel.theorem_two_probabilistic theorem theorem_two_probabilistic (h : ℕ) (hh : 1≤h) (hp : m.globalmintrigger<1) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) : s.approximationregret h≤(2*γ/(2-ω))*(12*(m:ℝ)*real.log (h:ℝ)/m.globalmintrigger)^(ω/2)* (h:ℝ)^(1-ω/2)+(1+real.pi^2/2)*(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha+ ∑i:fin m, (24*real.log (h:ℝ)/m.mintrigger i)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/3e4ea64c4acc8035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_polynomial","label":"inverseAt_polynomial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_polynomial","description":"theorem inverseAt_polynomial (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.inverseAt d=(d/γ)^(1/ω)","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html#decl-5f01e789cc5f","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","order":273,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialThreshold"],["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_polynomial (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.inverseAt d=(d/γ)^(1/ω)","missing":[],"search":"inverseat_polynomial banditrlproof.cucb.sourcemodel.inverseat_polynomial theorem inverseat_polynomial (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) {d : ℝ} (hd : d∈s.gapdomain) : s.inverseat d=(d/γ)^(1/ω) theorem compiled","shard":"modules/d24a7d0182819f29.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_polynomial_square","label":"inverseAt_polynomial_square","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_polynomial_square","description":"theorem inverseAt_polynomial_square (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : (S.inverseAt d)^2=(d/γ)^(2/ω)","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html#decl-5b9e6633b413","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","order":274,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialThreshold"],["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_polynomial_square (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : (S.inverseAt d)^2=(d/γ)^(2/ω)","missing":[],"search":"inverseat_polynomial_square banditrlproof.cucb.sourcemodel.inverseat_polynomial_square theorem inverseat_polynomial_square (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) {d : ℝ} (hd : d∈s.gapdomain) : (s.inverseat d)^2=(d/γ)^(2/ω) theorem compiled","shard":"modules/d24a7d0182819f29.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_polynomial_reciprocal","label":"inverseAt_polynomial_reciprocal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseAt_polynomial_reciprocal","description":"theorem inverseAt_polynomial_reciprocal (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : ((S.inverseAt d)^2)⁻¹=γ^(2/ω)*d^(-(2/ω))","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html#decl-18b7e661841a","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","order":275,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialThreshold"],["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseAt_polynomial_reciprocal (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : ((S.inverseAt d)^2)⁻¹=γ^(2/ω)*d^(-(2/ω))","missing":[],"search":"inverseat_polynomial_reciprocal banditrlproof.cucb.sourcemodel.inverseat_polynomial_reciprocal theorem inverseat_polynomial_reciprocal (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) {d : ℝ} (hd : d∈s.gapdomain) : ((s.inverseat d)^2)⁻¹=γ^(2/ω)*d^(-(2/ω)) theorem compiled","shard":"modules/d24a7d0182819f29.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_deterministic","label":"gapThreshold_polynomial_deterministic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_deterministic","description":"theorem gapThreshold_polynomial_deterministic (H : ℕ) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.gapThreshold H 1 d=6*Real.log (H:ℝ)*γ^(2/ω)*d^(-(2/ω))","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html#decl-203cc67c9296","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","order":276,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialThreshold"],["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_polynomial_deterministic (H : ℕ) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.gapThreshold H 1 d=6*Real.log (H:ℝ)*γ^(2/ω)*d^(-(2/ω))","missing":[],"search":"gapthreshold_polynomial_deterministic banditrlproof.cucb.sourcemodel.gapthreshold_polynomial_deterministic theorem gapthreshold_polynomial_deterministic (h : ℕ) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) {d : ℝ} (hd : d∈s.gapdomain) : s.gapthreshold h 1 d=6*real.log (h:ℝ)*γ^(2/ω)*d^(-(2/ω)) theorem compiled","shard":"modules/d24a7d0182819f29.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_probabilistic","label":"gapThreshold_polynomial_probabilistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_probabilistic","description":"theorem gapThreshold_polynomial_probabilistic (H : ℕ) (hH : 1≤H) (γ ω p : ℝ) (hγ : 0<γ) (hω : 0<ω) (hp : p≠1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.gapThreshold H p d=max (12*Real.log (H:ℝ)/p*γ^(2/ω)*d^(-(2/ω))) (24*Real.log (H:ℝ)/p)","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html#decl-f9ed7b0afaef","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","order":277,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialThreshold"],["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_polynomial_probabilistic (H : ℕ) (hH : 1≤H) (γ ω p : ℝ) (hγ : 0<γ) (hω : 0<ω) (hp : p≠1) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.gapThreshold H p d=max (12*Real.log (H:ℝ)/p*γ^(2/ω)*d^(-(2/ω))) (24*Real.log (H:ℝ)/p)","missing":[],"search":"gapthreshold_polynomial_probabilistic banditrlproof.cucb.sourcemodel.gapthreshold_polynomial_probabilistic theorem gapthreshold_polynomial_probabilistic (h : ℕ) (hh : 1≤h) (γ ω p : ℝ) (hγ : 0<γ) (hω : 0<ω) (hp : p≠1) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) {d : ℝ} (hd : d∈s.gapdomain) : s.gapthreshold h p d=max (12*real.log (h:ℝ)/p*γ^(2/ω)*d^(-(2/ω))) (24*real.log (h:ℝ)/p) theorem compiled","shard":"modules/d24a7d0182819f29.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_upper","label":"gapThreshold_polynomial_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_upper","description":"theorem gapThreshold_polynomial_upper (H : ℕ) (hH : 1≤H) (i : Fin m) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.gapThreshold H (M.minTrigger i) d≤ 12*Real.log (H:ℝ)/M.globalMinTrigger*γ^(2/ω)*d^(-(2/ω))+ 24*Real.log (H:ℝ)/M.minTrigger i","url":"../modules/banditrlproof-algorithms-cucbpolynomialthreshold/index.html#decl-00699e10f10a","parent":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","order":278,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBPolynomialThreshold"],["Source","BanditRLProof/Algorithms/CUCBPolynomialThreshold.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_polynomial_upper (H : ℕ) (hH : 1≤H) (i : Fin m) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → S.modulus u=γ*u^ω) {d : ℝ} (hd : d∈S.gapDomain) : S.gapThreshold H (M.minTrigger i) d≤ 12*Real.log (H:ℝ)/M.globalMinTrigger*γ^(2/ω)*d^(-(2/ω))+ 24*Real.log (H:ℝ)/M.minTrigger i","missing":[],"search":"gapthreshold_polynomial_upper banditrlproof.cucb.sourcemodel.gapthreshold_polynomial_upper theorem gapthreshold_polynomial_upper (h : ℕ) (hh : 1≤h) (i : fin m) (γ ω : ℝ) (hγ : 0<γ) (hω : 0<ω) (hf : ∀u, 0≤u → s.modulus u=γ*u^ω) {d : ℝ} (hd : d∈s.gapdomain) : s.gapthreshold h (m.mintrigger i) d≤ 12*real.log (h:ℝ)/m.globalmintrigger*γ^(2/ω)*d^(-(2/ω))+ 24*real.log (h:ℝ)/m.mintrigger i theorem compiled","shard":"modules/d24a7d0182819f29.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.expected_underSampledGap_le_refined","label":"expected_underSampledGap_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.expected_underSampledGap_le_refined","description":"theorem expected_underSampledGap_le_refined (H : ℕ) (hH : 1≤H) : (∑n∈Finset.range H, ∫Y : ℕ → Round A m, S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 ∂cucbTrajectory S.oracle M.environment) ≤ (∑i:Fin m, S.armRefinedTerm H i)+(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbrefinedregret/index.html#decl-5e7c3903d527","parent":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","order":279,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRefinedRegret"],["Source","BanditRLProof/Algorithms/CUCBRefinedRegret.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_underSampledGap_le_refined (H : ℕ) (hH : 1≤H) : (∑n∈Finset.range H, ∫Y : ℕ → Round A m, S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 ∂cucbTrajectory S.oracle M.environment) ≤ (∑i:Fin m, S.armRefinedTerm H i)+(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"expected_undersampledgap_le_refined banditrlproof.cucb.sourcemodel.expected_undersampledgap_le_refined theorem expected_undersampledgap_le_refined (h : ℕ) (hh : 1≤h) : (∑n∈finset.range h, ∫y : ℕ → round a m, s.undersampledgap h (s.chargedata.counters (fun t => (y t).1) n) (y n).1 ∂cucbtrajectory s.oracle m.environment) ≤ (∑i:fin m, s.armrefinedterm h i)+(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/0afac4c1b20b09ab.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.theorem_one_refined_regret","label":"theorem_one_refined_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.theorem_one_refined_regret","description":"Chen et al. JMLR 2016 Theorem 1, with the documented normalized charge repair. The per-arm gap family is derived from actual bad actions and possible triggers; an empty family contributes zero. No performance/count premise is assumed.","url":"../modules/banditrlproof-algorithms-cucbrefinedregret/index.html#decl-409f59495c3c","parent":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","order":280,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRefinedRegret"],["Source","BanditRLProof/Algorithms/CUCBRefinedRegret.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem theorem_one_refined_regret (H : ℕ) (hH : 1≤H) : S.approximationRegret H ≤ (∑i:Fin m, S.armRefinedTerm H i)+ (1+(2+(if M.globalMinTrigger<1 then 1 else 0))*Real.pi^2/6)* (m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"theorem_one_refined_regret banditrlproof.cucb.sourcemodel.theorem_one_refined_regret chen et al. jmlr 2016 theorem 1, with the documented normalized charge repair. the per-arm gap family is derived from actual bad actions and possible triggers; an empty family contributes zero. no performance/count premise is assumed. theorem compiled","shard":"modules/0afac4c1b20b09ab.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.armRefinedTerm_eq_zero_of_no_bad","label":"armRefinedTerm_eq_zero_of_no_bad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.armRefinedTerm_eq_zero_of_no_bad","description":"theorem armRefinedTerm_eq_zero_of_no_bad (H : ℕ) (h : ∀a, S.gap a≤0) (i : Fin m) : S.armRefinedTerm H i=0","url":"../modules/banditrlproof-algorithms-cucbrefinedregret/index.html#decl-38278f33dda6","parent":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","order":281,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRefinedRegret"],["Source","BanditRLProof/Algorithms/CUCBRefinedRegret.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem armRefinedTerm_eq_zero_of_no_bad (H : ℕ) (h : ∀a, S.gap a≤0) (i : Fin m) : S.armRefinedTerm H i=0","missing":[],"search":"armrefinedterm_eq_zero_of_no_bad banditrlproof.cucb.sourcemodel.armrefinedterm_eq_zero_of_no_bad theorem armrefinedterm_eq_zero_of_no_bad (h : ℕ) (h : ∀a, s.gap a≤0) (i : fin m) : s.armrefinedterm h i=0 theorem compiled","shard":"modules/0afac4c1b20b09ab.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_nonpos_of_no_bad","label":"approximationRegret_nonpos_of_no_bad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_nonpos_of_no_bad","description":"theorem approximationRegret_nonpos_of_no_bad (H : ℕ) (h : ∀a, S.gap a≤0) : S.approximationRegret H≤0","url":"../modules/banditrlproof-algorithms-cucbrefinedregret/index.html#decl-a482e8054afb","parent":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","order":282,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRefinedRegret"],["Source","BanditRLProof/Algorithms/CUCBRefinedRegret.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_nonpos_of_no_bad (H : ℕ) (h : ∀a, S.gap a≤0) : S.approximationRegret H≤0","missing":[],"search":"approximationregret_nonpos_of_no_bad banditrlproof.cucb.sourcemodel.approximationregret_nonpos_of_no_bad theorem approximationregret_nonpos_of_no_bad (h : ℕ) (h : ∀a, s.gap a≤0) : s.approximationregret h≤0 theorem compiled","shard":"modules/0afac4c1b20b09ab.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underSampledGap","label":"underSampledGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underSampledGap","description":"noncomputable def underSampledGap (H : ℕ) (N : Fin m → ℕ) (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-1568f029f18e","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":283,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def underSampledGap (H : ℕ) (N : Fin m → ℕ) (a : A) : ℝ","missing":[],"search":"undersampledgap banditrlproof.cucb.sourcemodel.undersampledgap noncomputable def undersampledgap (h : ℕ) (n : fin m → ℕ) (a : a) : ℝ definition compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sufficientIndicator","label":"sufficientIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sufficientIndicator","description":"noncomputable def sufficientIndicator (H n : ℕ) (Y : ℕ → Round A m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-fe25ab01a7ed","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":284,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sufficientIndicator (H n : ℕ) (Y : ℕ → Round A m) : ℝ","missing":[],"search":"sufficientindicator banditrlproof.cucb.sourcemodel.sufficientindicator noncomputable def sufficientindicator (h n : ℕ) (y : ℕ → round a m) : ℝ definition compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.maxPositiveGap_nonneg","label":"maxPositiveGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.maxPositiveGap_nonneg","description":"theorem maxPositiveGap_nonneg : 0≤maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-562bab30280f","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":285,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem maxPositiveGap_nonneg : 0≤maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"maxpositivegap_nonneg banditrlproof.cucb.sourcemodel.maxpositivegap_nonneg theorem maxpositivegap_nonneg : 0≤maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underSampledGap_mem","label":"underSampledGap_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underSampledGap_mem","description":"theorem underSampledGap_mem (H : ℕ) (N : Fin m → ℕ) (a : A) : S.underSampledGap H N a∈Set.Icc 0 (maxPositiveGap S.score M.trueInput S.alpha)","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-20dd26d982c3","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":286,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underSampledGap_mem (H : ℕ) (N : Fin m → ℕ) (a : A) : S.underSampledGap H N a∈Set.Icc 0 (maxPositiveGap S.score M.trueInput S.alpha)","missing":[],"search":"undersampledgap_mem banditrlproof.cucb.sourcemodel.undersampledgap_mem theorem undersampledgap_mem (h : ℕ) (n : fin m → ℕ) (a : a) : s.undersampledgap h n a∈set.icc 0 (maxpositivegap s.score m.trueinput s.alpha) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sufficientIndicator_mem","label":"sufficientIndicator_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sufficientIndicator_mem","description":"theorem sufficientIndicator_mem (H n : ℕ) (Y : ℕ → Round A m) : S.sufficientIndicator H n Y∈Set.Icc (0:ℝ) 1","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-641acb35a2c1","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":287,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sufficientIndicator_mem (H n : ℕ) (Y : ℕ → Round A m) : S.sufficientIndicator H n Y∈Set.Icc (0:ℝ) 1","missing":[],"search":"sufficientindicator_mem banditrlproof.cucb.sourcemodel.sufficientindicator_mem theorem sufficientindicator_mem (h n : ℕ) (y : ℕ → round a m) : s.sufficientindicator h n y∈set.icc (0:ℝ) 1 theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gap_decomposition","label":"gap_decomposition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gap_decomposition","description":"theorem gap_decomposition (H n : ℕ) (Y : ℕ → Round A m) : S.gap (Y n).1≤S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 + maxPositiveGap S.score M.trueInput S.alpha*S.sufficientIndicator H n Y + S.alpha*scoreOptimum S.score M.trueInput* (1-S.successIndicator (oracleInput (fun t => (Y t).2) n) (Y n))","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-cd33555d7a27","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":288,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_decomposition (H n : ℕ) (Y : ℕ → Round A m) : S.gap (Y n).1≤S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 + maxPositiveGap S.score M.trueInput S.alpha*S.sufficientIndicator H n Y + S.alpha*scoreOptimum S.score M.trueInput* (1-S.successIndicator (oracleInput (fun t => (Y t).2) n) (Y n))","missing":[],"search":"gap_decomposition banditrlproof.cucb.sourcemodel.gap_decomposition theorem gap_decomposition (h n : ℕ) (y : ℕ → round a m) : s.gap (y n).1≤s.undersampledgap h (s.chargedata.counters (fun t => (y t).1) n) (y n).1 + maxpositivegap s.score m.trueinput s.alpha*s.sufficientindicator h n y + s.alpha*scoreoptimum s.score m.trueinput* (1-s.successindicator (oracleinput (fun t => (y t).2) n) (y n)) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.measurable_actual_underSampledGap","label":"measurable_actual_underSampledGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.measurable_actual_underSampledGap","description":"theorem measurable_actual_underSampledGap (H n : ℕ) : Measurable (fun Y : ℕ → Round A m => S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1)","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-e42d53dc64ca","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":289,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_actual_underSampledGap (H n : ℕ) : Measurable (fun Y : ℕ → Round A m => S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1)","missing":[],"search":"measurable_actual_undersampledgap banditrlproof.cucb.sourcemodel.measurable_actual_undersampledgap theorem measurable_actual_undersampledgap (h n : ℕ) : measurable (fun y : ℕ → round a m => s.undersampledgap h (s.chargedata.counters (fun t => (y t).1) n) (y n).1) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.measurableSet_sufficientSuccessfulCharge","label":"measurableSet_sufficientSuccessfulCharge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.measurableSet_sufficientSuccessfulCharge","description":"theorem measurableSet_sufficientSuccessfulCharge (H n : ℕ) : MeasurableSet (S.SufficientSuccessfulCharge H n)","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-69f2ea648ad3","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":290,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_sufficientSuccessfulCharge (H n : ℕ) : MeasurableSet (S.SufficientSuccessfulCharge H n)","missing":[],"search":"measurableset_sufficientsuccessfulcharge banditrlproof.cucb.sourcemodel.measurableset_sufficientsuccessfulcharge theorem measurableset_sufficientsuccessfulcharge (h n : ℕ) : measurableset (s.sufficientsuccessfulcharge h n) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_actual_underSampledGap","label":"integrable_actual_underSampledGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_actual_underSampledGap","description":"theorem integrable_actual_underSampledGap (H n : ℕ) : Integrable (fun Y : ℕ → Round A m => S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1) (cucbTrajectory S.oracle M.environment)","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-d1e386f08b55","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":291,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_actual_underSampledGap (H n : ℕ) : Integrable (fun Y : ℕ → Round A m => S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1) (cucbTrajectory S.oracle M.environment)","missing":[],"search":"integrable_actual_undersampledgap banditrlproof.cucb.sourcemodel.integrable_actual_undersampledgap theorem integrable_actual_undersampledgap (h n : ℕ) : integrable (fun y : ℕ → round a m => s.undersampledgap h (s.chargedata.counters (fun t => (y t).1) n) (y n).1) (cucbtrajectory s.oracle m.environment) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_sufficientIndicator","label":"integrable_sufficientIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_sufficientIndicator","description":"theorem integrable_sufficientIndicator (H n : ℕ) : Integrable (S.sufficientIndicator H n) (cucbTrajectory S.oracle M.environment)","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-15599857c4ff","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":292,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_sufficientIndicator (H n : ℕ) : Integrable (S.sufficientIndicator H n) (cucbTrajectory S.oracle M.environment)","missing":[],"search":"integrable_sufficientindicator banditrlproof.cucb.sourcemodel.integrable_sufficientindicator theorem integrable_sufficientindicator (h n : ℕ) : integrable (s.sufficientindicator h n) (cucbtrajectory s.oracle m.environment) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.expected_gap_decomposition","label":"expected_gap_decomposition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.expected_gap_decomposition","description":"theorem expected_gap_decomposition (H n : ℕ) : (∫Y : ℕ → Round A m, S.gap (Y n).1 ∂cucbTrajectory S.oracle M.environment) ≤ (∫Y : ℕ → Round A m, S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 ∂cucbTrajectory S.oracle M.environment) + maxPositiveGap S.score M.trueInput S.alpha* (∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment) + S.alpha*scoreOptimum S.score M.trueInput…","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-fcf22dc63361","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":293,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_gap_decomposition (H n : ℕ) : (∫Y : ℕ → Round A m, S.gap (Y n).1 ∂cucbTrajectory S.oracle M.environment) ≤ (∫Y : ℕ → Round A m, S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 ∂cucbTrajectory S.oracle M.environment) + maxPositiveGap S.score M.trueInput S.alpha* (∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment) + S.alpha*scoreOptimum S.score M.trueInput*(1-S.beta)","missing":[],"search":"expected_gap_decomposition banditrlproof.cucb.sourcemodel.expected_gap_decomposition theorem expected_gap_decomposition (h n : ℕ) : (∫y : ℕ → round a m, s.gap (y n).1 ∂cucbtrajectory s.oracle m.environment) ≤ (∫y : ℕ → round a m, s.undersampledgap h (s.chargedata.counters (fun t => (y t).1) n) (y n).1 ∂cucbtrajectory s.oracle m.environment) + maxpositivegap s.score m.trueinput s.alpha* (∫y, s.sufficientindicator h n y ∂cucbtrajectory s.oracle m.environment) + s.alpha*scoreoptimum s.score m.trueinput*(1-s.beta) theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_underSampled_add_sufficient","label":"approximationRegret_le_underSampled_add_sufficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_le_underSampled_add_sufficient","description":"Exact cancellation against the negative credit in the original signed approximation regret. The under-count term is still the actual random weight.","url":"../modules/banditrlproof-algorithms-cucbregretdecomposition/index.html#decl-bba23727c477","parent":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","order":294,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretDecomposition"],["Source","BanditRLProof/Algorithms/CUCBRegretDecomposition.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem approximationRegret_le_underSampled_add_sufficient (H : ℕ) : S.approximationRegret H ≤ (∑n∈Finset.range H, ∫Y : ℕ → Round A m, S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 ∂cucbTrajectory S.oracle M.environment) + maxPositiveGap S.score M.trueInput S.alpha* ∑n∈Finset.range H, ∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment","missing":[],"search":"approximationregret_le_undersampled_add_sufficient banditrlproof.cucb.sourcemodel.approximationregret_le_undersampled_add_sufficient exact cancellation against the negative credit in the original signed approximation regret. the under-count term is still the actual random weight. theorem compiled","shard":"modules/25775f8ae73dfb18.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.expected_sufficientIndicator_le","label":"expected_sufficientIndicator_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.expected_sufficientIndicator_le","description":"theorem expected_sufficientIndicator_le (H n : ℕ) (hH : n+1≤H) : (∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment) ≤ (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)/((n:ℝ)+1)^2","url":"../modules/banditrlproof-algorithms-cucbregrettail/index.html#decl-2d058aaa7fe1","parent":"module:BanditRLProof.Algorithms.CUCBRegretTail","order":295,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretTail"],["Source","BanditRLProof/Algorithms/CUCBRegretTail.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_sufficientIndicator_le (H n : ℕ) (hH : n+1≤H) : (∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment) ≤ (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)/((n:ℝ)+1)^2","missing":[],"search":"expected_sufficientindicator_le banditrlproof.cucb.sourcemodel.expected_sufficientindicator_le theorem expected_sufficientindicator_le (h n : ℕ) (hh : n+1≤h) : (∫y, s.sufficientindicator h n y ∂cucbtrajectory s.oracle m.environment) ≤ (2+(if m.globalmintrigger<1 then 1 else 0))*(m:ℝ)/((n:ℝ)+1)^2 theorem compiled","shard":"modules/ff9f71f1783d3b37.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sum_inverse_square_le","label":"sum_inverse_square_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sum_inverse_square_le","description":"theorem sum_inverse_square_le (H : ℕ) : (∑n∈Finset.range H, (((n:ℝ)+1)^2)⁻¹)≤Real.pi^2/6","url":"../modules/banditrlproof-algorithms-cucbregrettail/index.html#decl-70bfc1005076","parent":"module:BanditRLProof.Algorithms.CUCBRegretTail","order":296,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretTail"],["Source","BanditRLProof/Algorithms/CUCBRegretTail.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_inverse_square_le (H : ℕ) : (∑n∈Finset.range H, (((n:ℝ)+1)^2)⁻¹)≤Real.pi^2/6","missing":[],"search":"sum_inverse_square_le banditrlproof.cucb.sourcemodel.sum_inverse_square_le theorem sum_inverse_square_le (h : ℕ) : (∑n∈finset.range h, (((n:ℝ)+1)^2)⁻¹)≤real.pi^2/6 theorem compiled","shard":"modules/ff9f71f1783d3b37.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sum_expected_sufficientIndicator_le","label":"sum_expected_sufficientIndicator_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sum_expected_sufficientIndicator_le","description":"theorem sum_expected_sufficientIndicator_le (H : ℕ) : (∑n∈Finset.range H, ∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment) ≤ (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)*(Real.pi^2/6)","url":"../modules/banditrlproof-algorithms-cucbregrettail/index.html#decl-1214999b5295","parent":"module:BanditRLProof.Algorithms.CUCBRegretTail","order":297,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretTail"],["Source","BanditRLProof/Algorithms/CUCBRegretTail.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_expected_sufficientIndicator_le (H : ℕ) : (∑n∈Finset.range H, ∫Y, S.sufficientIndicator H n Y ∂cucbTrajectory S.oracle M.environment) ≤ (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)*(Real.pi^2/6)","missing":[],"search":"sum_expected_sufficientindicator_le banditrlproof.cucb.sourcemodel.sum_expected_sufficientindicator_le theorem sum_expected_sufficientindicator_le (h : ℕ) : (∑n∈finset.range h, ∫y, s.sufficientindicator h n y ∂cucbtrajectory s.oracle m.environment) ≤ (2+(if m.globalmintrigger<1 then 1 else 0))*(m:ℝ)*(real.pi^2/6) theorem compiled","shard":"modules/ff9f71f1783d3b37.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_underSampled_add_source_tail","label":"approximationRegret_le_underSampled_add_source_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.approximationRegret_le_underSampled_add_source_tail","description":"The original signed regret, with the oracle-failure credit cancelled and all sufficiently sampled rounds summed with the exact source constant.","url":"../modules/banditrlproof-algorithms-cucbregrettail/index.html#decl-c91bb4a445c6","parent":"module:BanditRLProof.Algorithms.CUCBRegretTail","order":298,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRegretTail"],["Source","BanditRLProof/Algorithms/CUCBRegretTail.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem approximationRegret_le_underSampled_add_source_tail (H : ℕ) : S.approximationRegret H ≤ (∑n∈Finset.range H, ∫Y : ℕ → Round A m, S.underSampledGap H (S.chargeData.counters (fun t => (Y t).1) n) (Y n).1 ∂cucbTrajectory S.oracle M.environment) + (2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)* maxPositiveGap S.score M.trueInput S.alpha*Real.pi^2/6","missing":[],"search":"approximationregret_le_undersampled_add_source_tail banditrlproof.cucb.sourcemodel.approximationregret_le_undersampled_add_source_tail the original signed regret, with the oracle-failure credit cancelled and all sufficiently sampled rounds summed with the exact source constant. theorem compiled","shard":"modules/ff9f71f1783d3b37.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_trueScore_comp","label":"integrable_trueScore_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_trueScore_comp","description":"theorem integrable_trueScore_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → A) (hg : Measurable g) : Integrable (fun ω => S.score M.trueInput (g ω)) ν","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-85db879bbec2","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":299,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_trueScore_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → A) (hg : Measurable g) : Integrable (fun ω => S.score M.trueInput (g ω)) ν","missing":[],"search":"integrable_truescore_comp banditrlproof.cucb.sourcemodel.integrable_truescore_comp theorem integrable_truescore_comp {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] (g : ω → a) (hg : measurable g) : integrable (fun ω => s.score m.trueinput (g ω)) ν theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.environment_norm_reward","label":"environment_norm_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.environment_norm_reward","description":"theorem environment_norm_reward (a : A) : (∫z, ‖z.2.2‖ ∂M.environment a)=S.score M.trueInput a","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-a68c1e808ccf","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":300,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem environment_norm_reward (a : A) : (∫z, ‖z.2.2‖ ∂M.environment a)=S.score M.trueInput a","missing":[],"search":"environment_norm_reward banditrlproof.cucb.sourcemodel.environment_norm_reward theorem environment_norm_reward (a : a) : (∫z, ‖z.2.2‖ ∂m.environment a)=s.score m.trueinput a theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_round_reward","label":"integrable_round_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_round_reward","description":"theorem integrable_round_reward (v : Input m) : Integrable (fun z : Round A m => z.2.2.2) (roundKernel S.oracle M.environment v)","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-7320cf3f71cd","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":301,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_round_reward (v : Input m) : Integrable (fun z : Round A m => z.2.2.2) (roundKernel S.oracle M.environment v)","missing":[],"search":"integrable_round_reward banditrlproof.cucb.sourcemodel.integrable_round_reward theorem integrable_round_reward (v : input m) : integrable (fun z : round a m => z.2.2.2) (roundkernel s.oracle m.environment v) theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.round_reward_nonneg","label":"round_reward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.round_reward_nonneg","description":"theorem round_reward_nonneg (v : Input m) : ∀ᵐ z ∂roundKernel S.oracle M.environment v, 0≤z.2.2.2","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-a6818a1536ff","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":302,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem round_reward_nonneg (v : Input m) : ∀ᵐ z ∂roundKernel S.oracle M.environment v, 0≤z.2.2.2","missing":[],"search":"round_reward_nonneg banditrlproof.cucb.sourcemodel.round_reward_nonneg theorem round_reward_nonneg (v : input m) : ∀ᵐ z ∂roundkernel s.oracle m.environment v, 0≤z.2.2.2 theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integral_round_reward","label":"integral_round_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integral_round_reward","description":"theorem integral_round_reward (v : Input m) : (∫z : Round A m, z.2.2.2 ∂roundKernel S.oracle M.environment v)= ∫a, S.score M.trueInput a ∂S.oracle v","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-585d7a235797","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":303,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_round_reward (v : Input m) : (∫z : Round A m, z.2.2.2 ∂roundKernel S.oracle M.environment v)= ∫a, S.score M.trueInput a ∂S.oracle v","missing":[],"search":"integral_round_reward banditrlproof.cucb.sourcemodel.integral_round_reward theorem integral_round_reward (v : input m) : (∫z : round a m, z.2.2.2 ∂roundkernel s.oracle m.environment v)= ∫a, s.score m.trueinput a ∂s.oracle v theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integral_round_norm_reward_le","label":"integral_round_norm_reward_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integral_round_norm_reward_le","description":"theorem integral_round_norm_reward_le (v : Input m) : (∫z : Round A m, ‖z.2.2.2‖ ∂roundKernel S.oracle M.environment v)≤ scoreOptimum S.score M.trueInput","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-9463f145acd0","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":304,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_round_norm_reward_le (v : Input m) : (∫z : Round A m, ‖z.2.2.2‖ ∂roundKernel S.oracle M.environment v)≤ scoreOptimum S.score M.trueInput","missing":[],"search":"integral_round_norm_reward_le banditrlproof.cucb.sourcemodel.integral_round_norm_reward_le theorem integral_round_norm_reward_le (v : input m) : (∫z : round a m, ‖z.2.2.2‖ ∂roundkernel s.oracle m.environment v)≤ scoreoptimum s.score m.trueinput theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integral_round_reward_eq_score","label":"integral_round_reward_eq_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integral_round_reward_eq_score","description":"theorem integral_round_reward_eq_score (v : Input m) : (∫z : Round A m, z.2.2.2 ∂roundKernel S.oracle M.environment v)= ∫z : Round A m, S.score M.trueInput z.1 ∂roundKernel S.oracle M.environment v","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-33709d114a08","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":305,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_round_reward_eq_score (v : Input m) : (∫z : Round A m, z.2.2.2 ∂roundKernel S.oracle M.environment v)= ∫z : Round A m, S.score M.trueInput z.1 ∂roundKernel S.oracle M.environment v","missing":[],"search":"integral_round_reward_eq_score banditrlproof.cucb.sourcemodel.integral_round_reward_eq_score theorem integral_round_reward_eq_score (v : input m) : (∫z : round a m, z.2.2.2 ∂roundkernel s.oracle m.environment v)= ∫z : round a m, s.score m.trueinput z.1 ∂roundkernel s.oracle m.environment v theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integrable_joint_reward","label":"integrable_joint_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integrable_joint_reward","description":"theorem integrable_joint_reward {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → Input m) (hg : Measurable g) : Integrable (fun p : Ω × Round A m => p.2.2.2.2) (ν ⊗ₘ (roundKernel S.oracle M.environment).comap g hg)","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-93484a45249d","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":306,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_joint_reward {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → Input m) (hg : Measurable g) : Integrable (fun p : Ω × Round A m => p.2.2.2.2) (ν ⊗ₘ (roundKernel S.oracle M.environment).comap g hg)","missing":[],"search":"integrable_joint_reward banditrlproof.cucb.sourcemodel.integrable_joint_reward theorem integrable_joint_reward {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] (g : ω → input m) (hg : measurable g) : integrable (fun p : ω × round a m => p.2.2.2.2) (ν ⊗ₘ (roundkernel s.oracle m.environment).comap g hg) theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.integral_joint_reward_eq_score","label":"integral_joint_reward_eq_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.integral_joint_reward_eq_score","description":"theorem integral_joint_reward_eq_score {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → Input m) (hg : Measurable g) : (∫p : Ω × Round A m, p.2.2.2.2 ∂(ν ⊗ₘ (roundKernel S.oracle M.environment).comap g hg))= ∫p : Ω × Round A m, S.score M.trueInput p.2.1 ∂(ν ⊗ₘ (roundKernel S.oracle M.environment).comap g hg)","url":"../modules/banditrlproof-algorithms-cucbrewardkernel/index.html#decl-8c5a51d937ef","parent":"module:BanditRLProof.Algorithms.CUCBRewardKernel","order":307,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRewardKernel"],["Source","BanditRLProof/Algorithms/CUCBRewardKernel.lean:87"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_joint_reward_eq_score {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] (g : Ω → Input m) (hg : Measurable g) : (∫p : Ω × Round A m, p.2.2.2.2 ∂(ν ⊗ₘ (roundKernel S.oracle M.environment).comap g hg))= ∫p : Ω × Round A m, S.score M.trueInput p.2.1 ∂(ν ⊗ₘ (roundKernel S.oracle M.environment).comap g hg)","missing":[],"search":"integral_joint_reward_eq_score banditrlproof.cucb.sourcemodel.integral_joint_reward_eq_score theorem integral_joint_reward_eq_score {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] (g : ω → input m) (hg : measurable g) : (∫p : ω × round a m, p.2.2.2.2 ∂(ν ⊗ₘ (roundkernel s.oracle m.environment).comap g hg))= ∫p : ω × round a m, s.score m.trueinput p.2.1 ∂(ν ⊗ₘ (roundkernel s.oracle m.environment).comap g hg) theorem compiled","shard":"modules/9f7508adc9e5df3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.marginalMean_mem","label":"marginalMean_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.marginalMean_mem","description":"theorem marginalMean_mem (D : Measure UnitOutcome) [IsProbabilityMeasure D] : marginalMean D ∈ Set.Icc (0:ℝ) 1","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-4dd90ec825d6","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":308,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem marginalMean_mem (D : Measure UnitOutcome) [IsProbabilityMeasure D] : marginalMean D ∈ Set.Icc (0:ℝ) 1","missing":[],"search":"marginalmean_mem banditrlproof.cucb.marginalmean_mem theorem marginalmean_mem (d : measure unitoutcome) [isprobabilitymeasure d] : marginalmean d ∈ set.icc (0:ℝ) 1 theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.centeredFactor_nonneg","label":"centeredFactor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.centeredFactor_nonneg","description":"theorem centeredFactor_nonneg (D : Measure UnitOutcome) (tilt : ℝ) (x : UnitOutcome) : 0≤centeredFactor D tilt x","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-a466fd340d0a","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":309,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem centeredFactor_nonneg (D : Measure UnitOutcome) (tilt : ℝ) (x : UnitOutcome) : 0≤centeredFactor D tilt x","missing":[],"search":"centeredfactor_nonneg banditrlproof.cucb.centeredfactor_nonneg theorem centeredfactor_nonneg (d : measure unitoutcome) (tilt : ℝ) (x : unitoutcome) : 0≤centeredfactor d tilt x theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.centeredFactor_le","label":"centeredFactor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.centeredFactor_le","description":"theorem centeredFactor_le (D : Measure UnitOutcome) [IsProbabilityMeasure D] (tilt : ℝ) (x : UnitOutcome) : centeredFactor D tilt x ≤ Real.exp |tilt|","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-18eb2872c047","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":310,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem centeredFactor_le (D : Measure UnitOutcome) [IsProbabilityMeasure D] (tilt : ℝ) (x : UnitOutcome) : centeredFactor D tilt x ≤ Real.exp |tilt|","missing":[],"search":"centeredfactor_le banditrlproof.cucb.centeredfactor_le theorem centeredfactor_le (d : measure unitoutcome) [isprobabilitymeasure d] (tilt : ℝ) (x : unitoutcome) : centeredfactor d tilt x ≤ real.exp |tilt| theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedFactor_nonneg","label":"observedFactor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observedFactor_nonneg","description":"theorem observedFactor_nonneg {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : 0≤observedFactor D i tilt z","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-d53ffebd6111","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":311,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observedFactor_nonneg {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) (z : Feedback m) : 0≤observedFactor D i tilt z","missing":[],"search":"observedfactor_nonneg banditrlproof.cucb.observedfactor_nonneg theorem observedfactor_nonneg {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) (z : feedback m) : 0≤observedfactor d i tilt z theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.observedFactor_le","label":"observedFactor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.observedFactor_le","description":"theorem observedFactor_le {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt : ℝ) (z : Feedback m) : observedFactor D i tilt z≤Real.exp |tilt|","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-33689a6054a2","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":312,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem observedFactor_le {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt : ℝ) (z : Feedback m) : observedFactor D i tilt z≤Real.exp |tilt|","missing":[],"search":"observedfactor_le banditrlproof.cucb.observedfactor_le theorem observedfactor_le {m : ℕ} (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (tilt : ℝ) (z : feedback m) : observedfactor d i tilt z≤real.exp |tilt| theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_observedFactor","label":"measurable_observedFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_observedFactor","description":"theorem measurable_observedFactor {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) : Measurable (observedFactor D i tilt)","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-8d4683f39e52","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":313,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_observedFactor {m : ℕ} (D : Measure UnitOutcome) (i : Fin m) (tilt : ℝ) : Measurable (observedFactor D i tilt)","missing":[],"search":"measurable_observedfactor banditrlproof.cucb.measurable_observedfactor theorem measurable_observedfactor {m : ℕ} (d : measure unitoutcome) (i : fin m) (tilt : ℝ) : measurable (observedfactor d i tilt) theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integrable_observedFactor_comp","label":"integrable_observedFactor_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integrable_observedFactor_comp","description":"theorem integrable_observedFactor_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt : ℝ) (g : Ω → Feedback m) (hg : Measurable g) : Integrable (fun ω => observedFactor D i tilt (g ω)) ν","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-a4d14e68851c","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":314,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_observedFactor_comp {Ω : Type*} [MeasurableSpace Ω] (ν : Measure Ω) [IsProbabilityMeasure ν] {m : ℕ} (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (tilt : ℝ) (g : Ω → Feedback m) (hg : Measurable g) : Integrable (fun ω => observedFactor D i tilt (g ω)) ν","missing":[],"search":"integrable_observedfactor_comp banditrlproof.cucb.integrable_observedfactor_comp theorem integrable_observedfactor_comp {ω : type*} [measurablespace ω] (ν : measure ω) [isprobabilitymeasure ν] {m : ℕ} (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (tilt : ℝ) (g : ω → feedback m) (hg : measurable g) : integrable (fun ω => observedfactor d i tilt (g ω)) ν theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.roundKernel_observedFactor_le_one","label":"roundKernel_observedFactor_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.roundKernel_observedFactor_le_one","description":"theorem roundKernel_observedFactor_le_one {A : Type*} [MeasurableSpace A] {m : ℕ} (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (hcompat : ∀a, ObservationCompatible (environment a) D i) (v : Input m) (tilt : ℝ) : (∫ z, observedFactor D i tilt z.2 ∂roundKernel oracle environment v)…","url":"../modules/banditrlproof-algorithms-cucbroundmgf/index.html#decl-5c469f1dd832","parent":"module:BanditRLProof.Algorithms.CUCBRoundMGF","order":315,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBRoundMGF"],["Source","BanditRLProof/Algorithms/CUCBRoundMGF.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem roundKernel_observedFactor_le_one {A : Type*} [MeasurableSpace A] {m : ℕ} (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (D : Measure UnitOutcome) [IsProbabilityMeasure D] (i : Fin m) (hcompat : ∀a, ObservationCompatible (environment a) D i) (v : Input m) (tilt : ℝ) : (∫ z, observedFactor D i tilt z.2 ∂roundKernel oracle environment v)≤1","missing":[],"search":"roundkernel_observedfactor_le_one banditrlproof.cucb.roundkernel_observedfactor_le_one theorem roundkernel_observedfactor_le_one {a : type*} [measurablespace a] {m : ℕ} (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (d : measure unitoutcome) [isprobabilitymeasure d] (i : fin m) (hcompat : ∀a, observationcompatible (environment a) d i) (v : input m) (tilt : ℝ) : (∫ z, observedfactor d i tilt z.2 ∂roundkernel oracle environment v)≤1 theorem compiled","shard":"modules/6d3d75cbcb1c4711.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.scoreOptimum","label":"scoreOptimum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.scoreOptimum","description":"noncomputable def scoreOptimum (score : Input m → A → ℝ) (v : Input m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-074f7cf6be48","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":316,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def scoreOptimum (score : Input m → A → ℝ) (v : Input m) : ℝ","missing":[],"search":"scoreoptimum banditrlproof.cucb.scoreoptimum noncomputable def scoreoptimum (score : input m → a → ℝ) (v : input m) : ℝ definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.score_le_optimum","label":"score_le_optimum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.score_le_optimum","description":"theorem score_le_optimum (score : Input m → A → ℝ) (v : Input m) (a : A) : score v a≤scoreOptimum score v","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-323a182f8a53","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":317,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem score_le_optimum (score : Input m → A → ℝ) (v : Input m) (a : A) : score v a≤scoreOptimum score v","missing":[],"search":"score_le_optimum banditrlproof.cucb.score_le_optimum theorem score_le_optimum (score : input m → a → ℝ) (v : input m) (a : a) : score v a≤scoreoptimum score v theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.sourceGap","label":"sourceGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.sourceGap","description":"noncomputable def sourceGap (score : Input m → A → ℝ) (v : Input m) (α : ℝ) (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-ad69042ae72a","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":318,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceGap (score : Input m → A → ℝ) (v : Input m) (α : ℝ) (a : A) : ℝ","missing":[],"search":"sourcegap banditrlproof.cucb.sourcegap noncomputable def sourcegap (score : input m → a → ℝ) (v : input m) (α : ℝ) (a : a) : ℝ definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.maxPositiveGap","label":"maxPositiveGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.maxPositiveGap","description":"noncomputable def maxPositiveGap (score : Input m → A → ℝ) (v : Input m) (α : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-9e59883112b2","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":319,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def maxPositiveGap (score : Input m → A → ℝ) (v : Input m) (α : ℝ) : ℝ","missing":[],"search":"maxpositivegap banditrlproof.cucb.maxpositivegap noncomputable def maxpositivegap (score : input m → a → ℝ) (v : input m) (α : ℝ) : ℝ definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.gap_le_maxPositiveGap","label":"gap_le_maxPositiveGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.gap_le_maxPositiveGap","description":"theorem gap_le_maxPositiveGap (score : Input m → A → ℝ) (v : Input m) (α : ℝ) (a : A) : sourceGap score v α a≤maxPositiveGap score v α","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-4a0900d68d6c","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":320,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_le_maxPositiveGap (score : Input m → A → ℝ) (v : Input m) (α : ℝ) (a : A) : sourceGap score v α a≤maxPositiveGap score v α","missing":[],"search":"gap_le_maxpositivegap banditrlproof.cucb.gap_le_maxpositivegap theorem gap_le_maxpositivegap (score : input m → a → ℝ) (v : input m) (α : ℝ) (a : a) : sourcegap score v α a≤maxpositivegap score v α theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel","label":"SourceModel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel","description":"structure SourceModel (M : FeedbackModel A m) where","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-b47cc0b5db9a","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":321,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"structure SourceModel (M : FeedbackModel A m) where","missing":[],"search":"sourcemodel banditrlproof.cucb.sourcemodel structure sourcemodel (m : feedbackmodel a m) where structure compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.gap","label":"gap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.gap","description":"noncomputable def gap (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-30380ef08608","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":322,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gap (a : A) : ℝ","missing":[],"search":"gap banditrlproof.cucb.sourcemodel.gap noncomputable def gap (a : a) : ℝ definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.bad","label":"bad","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.bad","description":"noncomputable def bad (a : A) : Bool","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-43597e545d25","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":323,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bad (a : A) : Bool","missing":[],"search":"bad banditrlproof.cucb.sourcemodel.bad noncomputable def bad (a : a) : bool definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseGap","label":"inverseGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseGap","description":"noncomputable def inverseGap (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-19fed9a120c2","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":324,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseGap (a : A) : ℝ","missing":[],"search":"inversegap banditrlproof.cucb.sourcemodel.inversegap noncomputable def inversegap (a : a) : ℝ definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.inverseGap_spec","label":"inverseGap_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.inverseGap_spec","description":"theorem inverseGap_spec (a : A) (ha : 0<S.gap a) : 0<S.inverseGap a ∧ S.modulus (S.inverseGap a)=S.gap a","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-6639d129e794","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":325,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverseGap_spec (a : A) (ha : 0<S.gap a) : 0<S.inverseGap a ∧ S.modulus (S.inverseGap a)=S.gap a","missing":[],"search":"inversegap_spec banditrlproof.cucb.sourcemodel.inversegap_spec theorem inversegap_spec (a : a) (ha : 0<s.gap a) : 0<s.inversegap a ∧ s.modulus (s.inversegap a)=s.gap a theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.chargeData","label":"chargeData","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.chargeData","description":"noncomputable def chargeData : ChargeData A m","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-7124384166c2","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":326,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def chargeData : ChargeData A m","missing":[],"search":"chargedata banditrlproof.cucb.sourcemodel.chargedata noncomputable def chargedata : chargedata a m definition compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.chargeData_sufficient","label":"chargeData_sufficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.chargeData_sufficient","description":"theorem chargeData_sufficient (N : Fin m → ℕ) (a : A) (i : Fin m) (h : S.chargeData.choose N a=some i) (n : ℕ) (hi : samplingThreshold n (S.inverseGap a) (M.minTrigger i)<N i) : ∀j∈M.possible a, samplingThreshold n (S.inverseGap a) (M.minTrigger j)<N j","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-75103273aa97","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":327,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"theorem chargeData_sufficient (N : Fin m → ℕ) (a : A) (i : Fin m) (h : S.chargeData.choose N a=some i) (n : ℕ) (hi : samplingThreshold n (S.inverseGap a) (M.minTrigger i)<N i) : ∀j∈M.possible a, samplingThreshold n (S.inverseGap a) (M.minTrigger j)<N j","missing":[],"search":"chargedata_sufficient banditrlproof.cucb.sourcemodel.chargedata_sufficient theorem chargedata_sufficient (n : fin m → ℕ) (a : a) (i : fin m) (h : s.chargedata.choose n a=some i) (n : ℕ) (hi : samplingthreshold n (s.inversegap a) (m.mintrigger i)<n i) : ∀j∈m.possible a, samplingthreshold n (s.inversegap a) (m.mintrigger j)<n j theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.optimum_monotone","label":"optimum_monotone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.optimum_monotone","description":"theorem optimum_monotone (v w : Input m) (h : ∀i, (v i:ℝ)≤(w i:ℝ)) : scoreOptimum S.score v≤scoreOptimum S.score w","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-c864487fa955","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":328,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem optimum_monotone (v w : Input m) (h : ∀i, (v i:ℝ)≤(w i:ℝ)) : scoreOptimum S.score v≤scoreOptimum S.score w","missing":[],"search":"optimum_monotone banditrlproof.cucb.sourcemodel.optimum_monotone theorem optimum_monotone (v w : input m) (h : ∀i, (v i:ℝ)≤(w i:ℝ)) : scoreoptimum s.score v≤scoreoptimum s.score w theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.maxPositiveGap_eq_zero_of_no_bad","label":"maxPositiveGap_eq_zero_of_no_bad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.maxPositiveGap_eq_zero_of_no_bad","description":"theorem maxPositiveGap_eq_zero_of_no_bad (h : ∀a, S.gap a≤0) : maxPositiveGap S.score M.trueInput S.alpha=0","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-2478e4b6ea5e","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":329,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem maxPositiveGap_eq_zero_of_no_bad (h : ∀a, S.gap a≤0) : maxPositiveGap S.score M.trueInput S.alpha=0","missing":[],"search":"maxpositivegap_eq_zero_of_no_bad banditrlproof.cucb.sourcemodel.maxpositivegap_eq_zero_of_no_bad theorem maxpositivegap_eq_zero_of_no_bad (h : ∀a, s.gap a≤0) : maxpositivegap s.score m.trueinput s.alpha=0 theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.counters_zero_of_no_bad","label":"counters_zero_of_no_bad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.counters_zero_of_no_bad","description":"theorem counters_zero_of_no_bad (h : ∀a, S.gap a≤0) (actions : ℕ → A) (n : ℕ) (i : Fin m) : S.chargeData.counters actions n i=0","url":"../modules/banditrlproof-algorithms-cucbsourcemodel/index.html#decl-b098458f01e2","parent":"module:BanditRLProof.Algorithms.CUCBSourceModel","order":330,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSourceModel"],["Source","BanditRLProof/Algorithms/CUCBSourceModel.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_zero_of_no_bad (h : ∀a, S.gap a≤0) (actions : ℕ → A) (n : ℕ) (i : Fin m) : S.chargeData.counters actions n i=0","missing":[],"search":"counters_zero_of_no_bad banditrlproof.cucb.sourcemodel.counters_zero_of_no_bad theorem counters_zero_of_no_bad (h : ∀a, s.gap a≤0) (actions : ℕ → a) (n : ℕ) (i : fin m) : s.chargedata.counters actions n i=0 theorem compiled","shard":"modules/9daa49c94010d214.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.SufficientSuccessfulCharge","label":"SufficientSuccessfulCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.SufficientSuccessfulCharge","description":"def SufficientSuccessfulCharge (H n : ℕ) : Set (ℕ → Round A m)","url":"../modules/banditrlproof-algorithms-cucbsufficientsampling/index.html#decl-ccedf0a26a8e","parent":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","order":331,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBSufficientSampling"],["Source","BanditRLProof/Algorithms/CUCBSufficientSampling.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def SufficientSuccessfulCharge (H n : ℕ) : Set (ℕ → Round A m)","missing":[],"search":"sufficientsuccessfulcharge banditrlproof.cucb.sourcemodel.sufficientsuccessfulcharge def sufficientsuccessfulcharge (h n : ℕ) : set (ℕ → round a m) definition compiled","shard":"modules/d6ed59ac96fd3dd0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_subset","label":"sufficient_successful_charge_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_subset","description":"theorem sufficient_successful_charge_subset (H n : ℕ) : S.SufficientSuccessfulCharge H n ⊆ (NiceEvent (fun i => (M.trueInput i:ℝ)) n)ᶜ ∪ ⋃i, S.TriggerShortfall H n i","url":"../modules/banditrlproof-algorithms-cucbsufficientsampling/index.html#decl-06396f89b1aa","parent":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","order":332,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSufficientSampling"],["Source","BanditRLProof/Algorithms/CUCBSufficientSampling.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sufficient_successful_charge_subset (H n : ℕ) : S.SufficientSuccessfulCharge H n ⊆ (NiceEvent (fun i => (M.trueInput i:ℝ)) n)ᶜ ∪ ⋃i, S.TriggerShortfall H n i","missing":[],"search":"sufficient_successful_charge_subset banditrlproof.cucb.sourcemodel.sufficient_successful_charge_subset theorem sufficient_successful_charge_subset (h n : ℕ) : s.sufficientsuccessfulcharge h n ⊆ (niceevent (fun i => (m.trueinput i:ℝ)) n)ᶜ ∪ ⋃i, s.triggershortfall h n i theorem compiled","shard":"modules/d6ed59ac96fd3dd0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability","label":"sufficient_successful_charge_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability","description":"theorem sufficient_successful_charge_probability (H n : ℕ) (hH : n+1≤H) : (cucbTrajectory S.oracle M.environment) (S.SufficientSuccessfulCharge H n) ≤ ENNReal.ofReal (3*(m:ℝ)/((n:ℝ)+1)^2)","url":"../modules/banditrlproof-algorithms-cucbsufficientsampling/index.html#decl-c3336a58a9ec","parent":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","order":333,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSufficientSampling"],["Source","BanditRLProof/Algorithms/CUCBSufficientSampling.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sufficient_successful_charge_probability (H n : ℕ) (hH : n+1≤H) : (cucbTrajectory S.oracle M.environment) (S.SufficientSuccessfulCharge H n) ≤ ENNReal.ofReal (3*(m:ℝ)/((n:ℝ)+1)^2)","missing":[],"search":"sufficient_successful_charge_probability banditrlproof.cucb.sourcemodel.sufficient_successful_charge_probability theorem sufficient_successful_charge_probability (h n : ℕ) (hh : n+1≤h) : (cucbtrajectory s.oracle m.environment) (s.sufficientsuccessfulcharge h n) ≤ ennreal.ofreal (3*(m:ℝ)/((n:ℝ)+1)^2) theorem compiled","shard":"modules/d6ed59ac96fd3dd0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability_of_all_one","label":"sufficient_successful_charge_probability_of_all_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability_of_all_one","description":"theorem sufficient_successful_charge_probability_of_all_one (H n : ℕ) (hH : n+1≤H) (hp : ∀i, M.minTrigger i=1) : (cucbTrajectory S.oracle M.environment) (S.SufficientSuccessfulCharge H n) ≤ ENNReal.ofReal (2*(m:ℝ)/((n:ℝ)+1)^2)","url":"../modules/banditrlproof-algorithms-cucbsufficientsampling/index.html#decl-a50100e07295","parent":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","order":334,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSufficientSampling"],["Source","BanditRLProof/Algorithms/CUCBSufficientSampling.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sufficient_successful_charge_probability_of_all_one (H n : ℕ) (hH : n+1≤H) (hp : ∀i, M.minTrigger i=1) : (cucbTrajectory S.oracle M.environment) (S.SufficientSuccessfulCharge H n) ≤ ENNReal.ofReal (2*(m:ℝ)/((n:ℝ)+1)^2)","missing":[],"search":"sufficient_successful_charge_probability_of_all_one banditrlproof.cucb.sourcemodel.sufficient_successful_charge_probability_of_all_one theorem sufficient_successful_charge_probability_of_all_one (h n : ℕ) (hh : n+1≤h) (hp : ∀i, m.mintrigger i=1) : (cucbtrajectory s.oracle m.environment) (s.sufficientsuccessfulcharge h n) ≤ ennreal.ofreal (2*(m:ℝ)/((n:ℝ)+1)^2) theorem compiled","shard":"modules/d6ed59ac96fd3dd0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability_source","label":"sufficient_successful_charge_probability_source","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability_source","description":"The source indicator uses the minimum of actual trigger probabilities, not an independent parameter. Nonempty base arms follow from the feasible family.","url":"../modules/banditrlproof-algorithms-cucbsufficientsampling/index.html#decl-2b930f0125b5","parent":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","order":335,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBSufficientSampling"],["Source","BanditRLProof/Algorithms/CUCBSufficientSampling.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sufficient_successful_charge_probability_source (H n : ℕ) (hH : n+1≤H) : (cucbTrajectory S.oracle M.environment) (S.SufficientSuccessfulCharge H n) ≤ ENNReal.ofReal ((2+(if M.globalMinTrigger<1 then 1 else 0))*(m:ℝ)/((n:ℝ)+1)^2)","missing":[],"search":"sufficient_successful_charge_probability_source banditrlproof.cucb.sourcemodel.sufficient_successful_charge_probability_source the source indicator uses the minimum of actual trigger probabilities, not an independent parameter. nonempty base arms follow from the feasible family. theorem compiled","shard":"modules/d6ed59ac96fd3dd0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.thresholdCoefficient","label":"thresholdCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.thresholdCoefficient","description":"`u` is the positive inverse-smoothness value at the current gap.","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-6a7526d19e46","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":336,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def thresholdCoefficient (u p : ℝ) : ℝ","missing":[],"search":"thresholdcoefficient banditrlproof.cucb.thresholdcoefficient `u` is the positive inverse-smoothness value at the current gap. definition compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.samplingThreshold","label":"samplingThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.samplingThreshold","description":"noncomputable def samplingThreshold (n : ℕ) (u p : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-f135d1be1f18","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":337,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def samplingThreshold (n : ℕ) (u p : ℝ) : ℝ","missing":[],"search":"samplingthreshold banditrlproof.cucb.samplingthreshold noncomputable def samplingthreshold (n : ℕ) (u p : ℝ) : ℝ definition compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.thresholdCoefficient_pos","label":"thresholdCoefficient_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.thresholdCoefficient_pos","description":"theorem thresholdCoefficient_pos {u p : ℝ} (hu : 0<u) (hp : 0<p) : 0<thresholdCoefficient u p","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-55a7b60a8f37","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":338,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem thresholdCoefficient_pos {u p : ℝ} (hu : 0<u) (hp : 0<p) : 0<thresholdCoefficient u p","missing":[],"search":"thresholdcoefficient_pos banditrlproof.cucb.thresholdcoefficient_pos theorem thresholdcoefficient_pos {u p : ℝ} (hu : 0<u) (hp : 0<p) : 0<thresholdcoefficient u p theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.samplingThreshold_deterministic","label":"samplingThreshold_deterministic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.samplingThreshold_deterministic","description":"theorem samplingThreshold_deterministic (n : ℕ) (u : ℝ) : samplingThreshold n u 1 = 6 * Real.log n / u^2","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-2edfff137fb6","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":339,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem samplingThreshold_deterministic (n : ℕ) (u : ℝ) : samplingThreshold n u 1 = 6 * Real.log n / u^2","missing":[],"search":"samplingthreshold_deterministic banditrlproof.cucb.samplingthreshold_deterministic theorem samplingthreshold_deterministic (n : ℕ) (u : ℝ) : samplingthreshold n u 1 = 6 * real.log n / u^2 theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.samplingThreshold_probabilistic","label":"samplingThreshold_probabilistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.samplingThreshold_probabilistic","description":"theorem samplingThreshold_probabilistic {n : ℕ} {u p : ℝ} (hn : 1≤n) (hp : p≠1) : samplingThreshold n u p = max (12*Real.log n/(u^2*p)) (24*Real.log n/p)","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-fdded7e44079","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":340,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem samplingThreshold_probabilistic {n : ℕ} {u p : ℝ} (hn : 1≤n) (hp : p≠1) : samplingThreshold n u p = max (12*Real.log n/(u^2*p)) (24*Real.log n/p)","missing":[],"search":"samplingthreshold_probabilistic banditrlproof.cucb.samplingthreshold_probabilistic theorem samplingthreshold_probabilistic {n : ℕ} {u p : ℝ} (hn : 1≤n) (hp : p≠1) : samplingthreshold n u p = max (12*real.log n/(u^2*p)) (24*real.log n/p) theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.normalizedCharge","label":"normalizedCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.normalizedCharge","description":"A definite, tie-fixed analysis choice. Its arguments depend only on the past counters, the current action and fixed instance parameters.","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-65bf7be18024","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":341,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCharge {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N c : ι → ℝ) : ι","missing":[],"search":"normalizedcharge banditrlproof.cucb.normalizedcharge a definite, tie-fixed analysis choice. its arguments depend only on the past counters, the current action and fixed instance parameters. definition compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.normalizedCharge_spec","label":"normalizedCharge_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.normalizedCharge_spec","description":"theorem normalizedCharge_spec {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N c : ι → ℝ) : normalizedCharge s hs N c ∈ s ∧ ∀j∈s, N (normalizedCharge s hs N c) / c (normalizedCharge s hs N c) ≤ N j / c j","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-2ce9c40ddeb6","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":342,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem normalizedCharge_spec {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N c : ι → ℝ) : normalizedCharge s hs N c ∈ s ∧ ∀j∈s, N (normalizedCharge s hs N c) / c (normalizedCharge s hs N c) ≤ N j / c j","missing":[],"search":"normalizedcharge_spec banditrlproof.cucb.normalizedcharge_spec theorem normalizedcharge_spec {ι : type*} (s : finset ι) (hs : s.nonempty) (n c : ι → ℝ) : normalizedcharge s hs n c ∈ s ∧ ∀j∈s, n (normalizedcharge s hs n c) / c (normalizedcharge s hs n c) ≤ n j / c j theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.normalizedCharge_sufficient","label":"normalizedCharge_sufficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.normalizedCharge_sufficient","description":"theorem normalizedCharge_sufficient {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N c : ι → ℝ) (hc : ∀i∈s, 0<c i) (L : ℝ) (hL : L*c (normalizedCharge s hs N c)<N (normalizedCharge s hs N c)) : ∀j∈s, L*c j<N j","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-be88ab8d6998","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":343,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem normalizedCharge_sufficient {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N c : ι → ℝ) (hc : ∀i∈s, 0<c i) (L : ℝ) (hL : L*c (normalizedCharge s hs N c)<N (normalizedCharge s hs N c)) : ∀j∈s, L*c j<N j","missing":[],"search":"normalizedcharge_sufficient banditrlproof.cucb.normalizedcharge_sufficient theorem normalizedcharge_sufficient {ι : type*} (s : finset ι) (hs : s.nonempty) (n c : ι → ℝ) (hc : ∀i∈s, 0<c i) (l : ℝ) (hl : l*c (normalizedcharge s hs n c)<n (normalizedcharge s hs n c)) : ∀j∈s, l*c j<n j theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.exists_normalized_charge","label":"exists_normalized_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.exists_normalized_charge","description":"For any finite nonempty possible-trigger set, an analysis charge exists whose sufficient sampling forces sufficient sampling of every member. The threshold coefficients may differ across arms, including p=1 vs p<1. No claim of predictability is made here; that requires the actual history.","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-8e32ba80dce8","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":344,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_normalized_charge {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N c : ι → ℝ) (hc : ∀i∈s, 0<c i) : ∃i∈s, ∀j∈s, ∀L : ℝ, L*c i<N i → L*c j<N j","missing":[],"search":"exists_normalized_charge banditrlproof.cucb.exists_normalized_charge for any finite nonempty possible-trigger set, an analysis charge exists whose sufficient sampling forces sufficient sampling of every member. the threshold coefficients may differ across arms, including p=1 vs p<1. no claim of predictability is made here; that requires the actual history. theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.exists_threshold_charge","label":"exists_threshold_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.exists_threshold_charge","description":"theorem exists_threshold_charge {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N p : ι → ℝ) (u : ℝ) (hu : 0<u) (hp : ∀i∈s, 0<p i) : ∃i∈s, ∀j∈s, ∀n : ℕ, samplingThreshold n u (p i)<N i → samplingThreshold n u (p j)<N j","url":"../modules/banditrlproof-algorithms-cucbthreshold/index.html#decl-ffd749c16dfd","parent":"module:BanditRLProof.Algorithms.CUCBThreshold","order":345,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThreshold"],["Source","BanditRLProof/Algorithms/CUCBThreshold.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_threshold_charge {ι : Type*} (s : Finset ι) (hs : s.Nonempty) (N p : ι → ℝ) (u : ℝ) (hu : 0<u) (hp : ∀i∈s, 0<p i) : ∃i∈s, ∀j∈s, ∀n : ℕ, samplingThreshold n u (p i)<N i → samplingThreshold n u (p j)<N j","missing":[],"search":"exists_threshold_charge banditrlproof.cucb.exists_threshold_charge theorem exists_threshold_charge {ι : type*} (s : finset ι) (hs : s.nonempty) (n p : ι → ℝ) (u : ℝ) (hu : 0<u) (hp : ∀i∈s, 0<p i) : ∃i∈s, ∀j∈s, ∀n : ℕ, samplingthreshold n u (p i)<n i → samplingthreshold n u (p j)<n j theorem compiled","shard":"modules/486f0f6275048d07.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.probabilistic_threshold_crossing","label":"probabilistic_threshold_crossing","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.probabilistic_threshold_crossing","description":"theorem probabilistic_threshold_crossing (H n : ℕ) (hH : n+1≤H) (u p k : ℝ) (hu : 0<u) (hp : 0<p) (hp1 : p≠1) (hk : samplingThreshold H u p<k) : 0<k ∧ 6*Real.log ((n:ℝ)+1)/u^2<k*p/2 ∧ -k*p/8≤-(3*Real.log ((n:ℝ)+1))","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-e3577e8bae0a","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":346,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem probabilistic_threshold_crossing (H n : ℕ) (hH : n+1≤H) (u p k : ℝ) (hu : 0<u) (hp : 0<p) (hp1 : p≠1) (hk : samplingThreshold H u p<k) : 0<k ∧ 6*Real.log ((n:ℝ)+1)/u^2<k*p/2 ∧ -k*p/8≤-(3*Real.log ((n:ℝ)+1))","missing":[],"search":"probabilistic_threshold_crossing banditrlproof.cucb.probabilistic_threshold_crossing theorem probabilistic_threshold_crossing (h n : ℕ) (hh : n+1≤h) (u p k : ℝ) (hu : 0<u) (hp : 0<p) (hp1 : p≠1) (hk : samplingthreshold h u p<k) : 0<k ∧ 6*real.log ((n:ℝ)+1)/u^2<k*p/2 ∧ -k*p/8≤-(3*real.log ((n:ℝ)+1)) theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.arms_nonempty","label":"arms_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.arms_nonempty","description":"theorem arms_nonempty : (Finset.univ : Finset (Fin m)).Nonempty","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-18973742e04c","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":347,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arms_nonempty : (Finset.univ : Finset (Fin m)).Nonempty","missing":[],"search":"arms_nonempty banditrlproof.cucb.feedbackmodel.arms_nonempty theorem arms_nonempty : (finset.univ : finset (fin m)).nonempty theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger","label":"globalMinTrigger","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.globalMinTrigger","description":"noncomputable def globalMinTrigger : ℝ","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-a4011b677fc9","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":348,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def globalMinTrigger : ℝ","missing":[],"search":"globalmintrigger banditrlproof.cucb.feedbackmodel.globalmintrigger noncomputable def globalmintrigger : ℝ definition compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_pos","label":"globalMinTrigger_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_pos","description":"theorem globalMinTrigger_pos : 0<M.globalMinTrigger","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-d217c3b17f3b","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":349,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem globalMinTrigger_pos : 0<M.globalMinTrigger","missing":[],"search":"globalmintrigger_pos banditrlproof.cucb.feedbackmodel.globalmintrigger_pos theorem globalmintrigger_pos : 0<m.globalmintrigger theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_le","label":"globalMinTrigger_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_le","description":"theorem globalMinTrigger_le (i : Fin m) : M.globalMinTrigger≤M.minTrigger i","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-8795da9b2bf3","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":350,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem globalMinTrigger_le (i : Fin m) : M.globalMinTrigger≤M.minTrigger i","missing":[],"search":"globalmintrigger_le banditrlproof.cucb.feedbackmodel.globalmintrigger_le theorem globalmintrigger_le (i : fin m) : m.globalmintrigger≤m.mintrigger i theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_le_one","label":"globalMinTrigger_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_le_one","description":"theorem globalMinTrigger_le_one : M.globalMinTrigger≤1","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-aa45edf0a2ec","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":351,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem globalMinTrigger_le_one : M.globalMinTrigger≤1","missing":[],"search":"globalmintrigger_le_one banditrlproof.cucb.feedbackmodel.globalmintrigger_le_one theorem globalmintrigger_le_one : m.globalmintrigger≤1 theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_eq_one_iff","label":"globalMinTrigger_eq_one_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_eq_one_iff","description":"theorem globalMinTrigger_eq_one_iff : M.globalMinTrigger=1 ↔ ∀i, M.minTrigger i=1","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-4c03c51d7d47","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":352,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem globalMinTrigger_eq_one_iff : M.globalMinTrigger=1 ↔ ∀i, M.minTrigger i=1","missing":[],"search":"globalmintrigger_eq_one_iff banditrlproof.cucb.feedbackmodel.globalmintrigger_eq_one_iff theorem globalmintrigger_eq_one_iff : m.globalmintrigger=1 ↔ ∀i, m.mintrigger i=1 theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_fixed_count","label":"trigger_shortfall_fixed_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.trigger_shortfall_fixed_count","description":"theorem trigger_shortfall_fixed_count (H n : ℕ) (hH : n+1≤H) (i : Fin m) (u k : ℝ) (hu : 0<u) (hp1 : M.minTrigger i≠1) (hk : samplingThreshold H u (M.minTrigger i)<k) : (cucbTrajectory S.oracle M.environment) {Y | k≤(S.chargeData.counters (fun t => (Y t).1) n i : ℝ) ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤6*Real.log ((n:ℝ)+1)/u^2} ≤ ENNReal.ofReal (((n:ℝ)+1)^3)⁻¹","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-f6a393e2bad6","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":353,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_shortfall_fixed_count (H n : ℕ) (hH : n+1≤H) (i : Fin m) (u k : ℝ) (hu : 0<u) (hp1 : M.minTrigger i≠1) (hk : samplingThreshold H u (M.minTrigger i)<k) : (cucbTrajectory S.oracle M.environment) {Y | k≤(S.chargeData.counters (fun t => (Y t).1) n i : ℝ) ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤6*Real.log ((n:ℝ)+1)/u^2} ≤ ENNReal.ofReal (((n:ℝ)+1)^3)⁻¹","missing":[],"search":"trigger_shortfall_fixed_count banditrlproof.cucb.sourcemodel.trigger_shortfall_fixed_count theorem trigger_shortfall_fixed_count (h n : ℕ) (hh : n+1≤h) (i : fin m) (u k : ℝ) (hu : 0<u) (hp1 : m.mintrigger i≠1) (hk : samplingthreshold h u (m.mintrigger i)<k) : (cucbtrajectory s.oracle m.environment) {y | k≤(s.chargedata.counters (fun t => (y t).1) n i : ℝ) ∧ (observationcount (fun t => (y t).2) n i : ℝ)≤6*real.log ((n:ℝ)+1)/u^2} ≤ ennreal.ofreal (((n:ℝ)+1)^3)⁻¹ theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_slice","label":"trigger_shortfall_slice","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.trigger_shortfall_slice","description":"A single fixed-count tail covers every eligible action: minimize its positive inverse gap before taking probabilities, instead of unioning over actions.","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-c6568dde3464","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":354,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_shortfall_slice (H n : ℕ) (hH : n+1≤H) (i : Fin m) (hp1 : M.minTrigger i≠1) (k : ℕ) : (cucbTrajectory S.oracle M.environment) {Y | S.chargeData.counters (fun t => (Y t).1) n i=k ∧ ∃a, 0<S.gap a ∧ i∈M.possible a ∧ samplingThreshold H (S.inverseGap a) (M.minTrigger i)<k ∧ (observationCount (fun t => (Y t).2) n i : ℝ)≤ 6*Real.log ((n:ℝ)+1)/(S.inverseGap a)^2} ≤ ENNReal.ofReal (((n:ℝ)+1)^3)⁻¹","missing":[],"search":"trigger_shortfall_slice banditrlproof.cucb.sourcemodel.trigger_shortfall_slice a single fixed-count tail covers every eligible action: minimize its positive inverse gap before taking probabilities, instead of unioning over actions. theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.TriggerShortfall","label":"TriggerShortfall","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.TriggerShortfall","description":"The actual observation shortfall after crossing the source horizon threshold.","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-43b1e0622474","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":355,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def TriggerShortfall (H n : ℕ) (i : Fin m) : Set (ℕ → Round A m)","missing":[],"search":"triggershortfall banditrlproof.cucb.sourcemodel.triggershortfall the actual observation shortfall after crossing the source horizon threshold. definition compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_probability","label":"trigger_shortfall_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.trigger_shortfall_probability","description":"theorem trigger_shortfall_probability (H n : ℕ) (hH : n+1≤H) (i : Fin m) (hp1 : M.minTrigger i≠1) : (cucbTrajectory S.oracle M.environment) (S.TriggerShortfall H n i) ≤ ENNReal.ofReal ((((n:ℝ)+1)^2)⁻¹)","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-394aafa4be86","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":356,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_shortfall_probability (H n : ℕ) (hH : n+1≤H) (i : Fin m) (hp1 : M.minTrigger i≠1) : (cucbTrajectory S.oracle M.environment) (S.TriggerShortfall H n i) ≤ ENNReal.ofReal ((((n:ℝ)+1)^2)⁻¹)","missing":[],"search":"trigger_shortfall_probability banditrlproof.cucb.sourcemodel.trigger_shortfall_probability theorem trigger_shortfall_probability (h n : ℕ) (hh : n+1≤h) (i : fin m) (hp1 : m.mintrigger i≠1) : (cucbtrajectory s.oracle m.environment) (s.triggershortfall h n i) ≤ ennreal.ofreal ((((n:ℝ)+1)^2)⁻¹) theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_probability_of_one","label":"trigger_shortfall_probability_of_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.trigger_shortfall_probability_of_one","description":"theorem trigger_shortfall_probability_of_one (H n : ℕ) (hH : n+1≤H) (i : Fin m) (hp : M.minTrigger i=1) : (cucbTrajectory S.oracle M.environment) (S.TriggerShortfall H n i)=0","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-7e05159776e7","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":357,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:175"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_shortfall_probability_of_one (H n : ℕ) (hH : n+1≤H) (i : Fin m) (hp : M.minTrigger i=1) : (cucbTrajectory S.oracle M.environment) (S.TriggerShortfall H n i)=0","missing":[],"search":"trigger_shortfall_probability_of_one banditrlproof.cucb.sourcemodel.trigger_shortfall_probability_of_one theorem trigger_shortfall_probability_of_one (h n : ℕ) (hh : n+1≤h) (i : fin m) (hp : m.mintrigger i=1) : (cucbtrajectory s.oracle m.environment) (s.triggershortfall h n i)=0 theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_union_probability","label":"trigger_shortfall_union_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.trigger_shortfall_union_probability","description":"Sum only over base arms; the action family has already been absorbed in each fixed-count tail. Deterministically observed arms contribute zero.","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-d6fa408184ce","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":358,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:195"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_shortfall_union_probability (H n : ℕ) (hH : n+1≤H) : (cucbTrajectory S.oracle M.environment) (⋃i, S.TriggerShortfall H n i) ≤ ENNReal.ofReal ((m:ℝ)/((n:ℝ)+1)^2)","missing":[],"search":"trigger_shortfall_union_probability banditrlproof.cucb.sourcemodel.trigger_shortfall_union_probability sum only over base arms; the action family has already been absorbed in each fixed-count tail. deterministically observed arms contribute zero. theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_union_probability_of_all_one","label":"trigger_shortfall_union_probability_of_all_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.trigger_shortfall_union_probability_of_all_one","description":"theorem trigger_shortfall_union_probability_of_all_one (H n : ℕ) (hH : n+1≤H) (hp : ∀i, M.minTrigger i=1) : (cucbTrajectory S.oracle M.environment) (⋃i, S.TriggerShortfall H n i)=0","url":"../modules/banditrlproof-algorithms-cucbthresholdtail/index.html#decl-06b196768eff","parent":"module:BanditRLProof.Algorithms.CUCBThresholdTail","order":359,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBThresholdTail"],["Source","BanditRLProof/Algorithms/CUCBThresholdTail.lean:215"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trigger_shortfall_union_probability_of_all_one (H n : ℕ) (hH : n+1≤H) (hp : ∀i, M.minTrigger i=1) : (cucbTrajectory S.oracle M.environment) (⋃i, S.TriggerShortfall H n i)=0","missing":[],"search":"trigger_shortfall_union_probability_of_all_one banditrlproof.cucb.sourcemodel.trigger_shortfall_union_probability_of_all_one theorem trigger_shortfall_union_probability_of_all_one (h n : ℕ) (hh : n+1≤h) (hp : ∀i, m.mintrigger i=1) : (cucbtrajectory s.oracle m.environment) (⋃i, s.triggershortfall h n i)=0 theorem compiled","shard":"modules/88212ea011b4dcd4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.Round","label":"Round","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.CUCB.Round","description":"abbrev Round (A : Type*) (m : ℕ)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-24e599b62b05","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":360,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Round (A : Type*) (m : ℕ)","missing":[],"search":"round banditrlproof.cucb.round abbrev round (a : type*) (m : ℕ) abbreviation compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.Input","label":"Input","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.CUCB.Input","description":"abbrev Input (m : ℕ)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-329b2329e0fb","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":361,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Input (m : ℕ)","missing":[],"search":"input banditrlproof.cucb.input abbrev input (m : ℕ) abbreviation compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.emptyFeedback","label":"emptyFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.emptyFeedback","description":"def emptyFeedback (m : ℕ) : Feedback m","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-7d4b7a5b0941","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":362,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def emptyFeedback (m : ℕ) : Feedback m","missing":[],"search":"emptyfeedback banditrlproof.cucb.emptyfeedback def emptyfeedback (m : ℕ) : feedback m definition compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.initialInput","label":"initialInput","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.initialInput","description":"def initialInput (m : ℕ) : Input m","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-39fa33b5f7dd","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":363,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def initialInput (m : ℕ) : Input m","missing":[],"search":"initialinput banditrlproof.cucb.initialinput def initialinput (m : ℕ) : input m definition compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.oracleInput_zero","label":"oracleInput_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.oracleInput_zero","description":"theorem oracleInput_zero {m : ℕ} (Y : ℕ → Feedback m) : oracleInput Y 0 = initialInput m","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-4dd88f17e7fe","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":364,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oracleInput_zero {m : ℕ} (Y : ℕ → Feedback m) : oracleInput Y 0 = initialInput m","missing":[],"search":"oracleinput_zero banditrlproof.cucb.oracleinput_zero theorem oracleinput_zero {m : ℕ} (y : ℕ → feedback m) : oracleinput y 0 = initialinput m theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.roundKernel","label":"roundKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.roundKernel","description":"noncomputable def roundKernel (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) : Kernel (Input m) (Round A m)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-3d574052617d","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":365,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def roundKernel (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) : Kernel (Input m) (Round A m)","missing":[],"search":"roundkernel banditrlproof.cucb.roundkernel noncomputable def roundkernel (oracle : kernel (input m) a) (environment : kernel a (feedback m)) : kernel (input m) (round a m) definition compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.roundKernel_rectangle","label":"roundKernel_rectangle","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.roundKernel_rectangle","description":"The oracle draw precedes environment feedback; the latter depends on the realized action, not on an independently redrawn action.","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-70fda255e7ae","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":366,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem roundKernel_rectangle (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (v : Input m) (s : Set A) (t : Set (Feedback m)) (hs : MeasurableSet s) (ht : MeasurableSet t) : roundKernel oracle environment v (s ×ˢ t) = ∫⁻ a in s, environment a t ∂oracle v","missing":[],"search":"roundkernel_rectangle banditrlproof.cucb.roundkernel_rectangle the oracle draw precedes environment feedback; the latter depends on the realized action, not on an independently redrawn action. theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.roundKernel_action_law","label":"roundKernel_action_law","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.roundKernel_action_law","description":"theorem roundKernel_action_law (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (v : Input m) : (roundKernel oracle environment v).map Prod.fst = oracle v","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-ba1398a51a8a","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":367,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem roundKernel_action_law (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (v : Input m) : (roundKernel oracle environment v).map Prod.fst = oracle v","missing":[],"search":"roundkernel_action_law banditrlproof.cucb.roundkernel_action_law theorem roundkernel_action_law (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (v : input m) : (roundkernel oracle environment v).map prod.fst = oracle v theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.feedbackExtension","label":"feedbackExtension","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.feedbackExtension","description":"def feedbackExtension (n : ℕ) (h : (i : Finset.Iic n) → Round A m) : ℕ → Feedback m","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-e93eefd63ce6","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":368,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def feedbackExtension (n : ℕ) (h : (i : Finset.Iic n) → Round A m) : ℕ → Feedback m","missing":[],"search":"feedbackextension banditrlproof.cucb.feedbackextension def feedbackextension (n : ℕ) (h : (i : finset.iic n) → round a m) : ℕ → feedback m definition compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_feedbackExtension","label":"measurable_feedbackExtension","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_feedbackExtension","description":"theorem measurable_feedbackExtension (n : ℕ) : Measurable (feedbackExtension (A:=A) (m:=m) n)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-91af47b8ef95","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":369,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_feedbackExtension (n : ℕ) : Measurable (feedbackExtension (A:=A) (m:=m) n)","missing":[],"search":"measurable_feedbackextension banditrlproof.cucb.measurable_feedbackextension theorem measurable_feedbackextension (n : ℕ) : measurable (feedbackextension (a:=a) (m:=m) n) theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.oracleInput_feedbackExtension","label":"oracleInput_feedbackExtension","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.oracleInput_feedbackExtension","description":"theorem oracleInput_feedbackExtension (Y : ℕ → Round A m) (n : ℕ) : oracleInput (feedbackExtension n (Preorder.frestrictLe n Y)) (n+1) = oracleInput (fun t => (Y t).2) (n+1)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-e2cbfbc3c256","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":370,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oracleInput_feedbackExtension (Y : ℕ → Round A m) (n : ℕ) : oracleInput (feedbackExtension n (Preorder.frestrictLe n Y)) (n+1) = oracleInput (fun t => (Y t).2) (n+1)","missing":[],"search":"oracleinput_feedbackextension banditrlproof.cucb.oracleinput_feedbackextension theorem oracleinput_feedbackextension (y : ℕ → round a m) (n : ℕ) : oracleinput (feedbackextension n (preorder.frestrictle n y)) (n+1) = oracleinput (fun t => (y t).2) (n+1) theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucbStepKernel","label":"cucbStepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbStepKernel","description":"noncomputable def cucbStepKernel (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) (n : ℕ) : Kernel ((i : Finset.Iic n) → Round A m) (Round A m)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-0f872bcc7794","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":371,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def cucbStepKernel (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) (n : ℕ) : Kernel ((i : Finset.Iic n) → Round A m) (Round A m)","missing":[],"search":"cucbstepkernel banditrlproof.cucb.cucbstepkernel noncomputable def cucbstepkernel (oracle : kernel (input m) a) (environment : kernel a (feedback m)) (n : ℕ) : kernel ((i : finset.iic n) → round a m) (round a m) definition compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucbTrajectory","label":"cucbTrajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbTrajectory","description":"noncomputable def cucbTrajectory (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] : Measure (ℕ → Round A m)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-8a6f0e2afe07","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":372,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:92"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","combinatorial"]],"statement":"noncomputable def cucbTrajectory (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] : Measure (ℕ → Round A m)","missing":[],"search":"cucbtrajectory banditrlproof.cucb.cucbtrajectory noncomputable def cucbtrajectory (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] : measure (ℕ → round a m) definition compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.CUCB.cucbStepKernel_apply_prefix","label":"cucbStepKernel_apply_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbStepKernel_apply_prefix","description":"theorem cucbStepKernel_apply_prefix (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) (Y : ℕ → Round A m) (n : ℕ) : cucbStepKernel oracle environment n (Preorder.frestrictLe n Y) = roundKernel oracle environment (oracleInput (fun t => (Y t).2) (n+1))","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-86ea0b1725ae","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":373,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucbStepKernel_apply_prefix (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) (Y : ℕ → Round A m) (n : ℕ) : cucbStepKernel oracle environment n (Preorder.frestrictLe n Y) = roundKernel oracle environment (oracleInput (fun t => (Y t).2) (n+1))","missing":[],"search":"cucbstepkernel_apply_prefix banditrlproof.cucb.cucbstepkernel_apply_prefix theorem cucbstepkernel_apply_prefix (oracle : kernel (input m) a) (environment : kernel a (feedback m)) (y : ℕ → round a m) (n : ℕ) : cucbstepkernel oracle environment n (preorder.frestrictle n y) = roundkernel oracle environment (oracleinput (fun t => (y t).2) (n+1)) theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucbTrajectory_prefix_compProd","label":"cucbTrajectory_prefix_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbTrajectory_prefix_compProd","description":"theorem cucbTrajectory_prefix_compProd (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (n : ℕ) : (cucbTrajectory oracle environment).map (Preorder.frestrictLe n) ⊗ₘ cucbStepKernel oracle environment n = (cucbTrajectory oracle environment).map (fun Y => (Preorder.frestrictLe n Y, Y (n+1)))","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-99e7ad2f0d78","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":374,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucbTrajectory_prefix_compProd (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (n : ℕ) : (cucbTrajectory oracle environment).map (Preorder.frestrictLe n) ⊗ₘ cucbStepKernel oracle environment n = (cucbTrajectory oracle environment).map (fun Y => (Preorder.frestrictLe n Y, Y (n+1)))","missing":[],"search":"cucbtrajectory_prefix_compprod banditrlproof.cucb.cucbtrajectory_prefix_compprod theorem cucbtrajectory_prefix_compprod (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (n : ℕ) : (cucbtrajectory oracle environment).map (preorder.frestrictle n) ⊗ₘ cucbstepkernel oracle environment n = (cucbtrajectory oracle environment).map (fun y => (preorder.frestrictle n y, y (n+1))) theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucbTrajectory_condDistrib","label":"cucbTrajectory_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbTrajectory_condDistrib","description":"theorem cucbTrajectory_condDistrib [StandardBorelSpace A] [Nonempty A] (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (n : ℕ) : condDistrib (fun Y : ℕ → Round A m => Y (n+1)) (Preorder.frestrictLe n) (cucbTrajectory oracle environment) =ᵐ[ (cucbTrajectory oracle environment).map (Preorder.frestrictLe n)] cucbStepKernel oracle environment n","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-a971e9df24da","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":375,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucbTrajectory_condDistrib [StandardBorelSpace A] [Nonempty A] (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] (n : ℕ) : condDistrib (fun Y : ℕ → Round A m => Y (n+1)) (Preorder.frestrictLe n) (cucbTrajectory oracle environment) =ᵐ[ (cucbTrajectory oracle environment).map (Preorder.frestrictLe n)] cucbStepKernel oracle environment n","missing":[],"search":"cucbtrajectory_conddistrib banditrlproof.cucb.cucbtrajectory_conddistrib theorem cucbtrajectory_conddistrib [standardborelspace a] [nonempty a] (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] (n : ℕ) : conddistrib (fun y : ℕ → round a m => y (n+1)) (preorder.frestrictle n) (cucbtrajectory oracle environment) =ᵐ[ (cucbtrajectory oracle environment).map (preorder.frestrictle n)] cucbstepkernel oracle environment n theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.cucbTrajectory_initial_law","label":"cucbTrajectory_initial_law","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.cucbTrajectory_initial_law","description":"theorem cucbTrajectory_initial_law (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] : (cucbTrajectory oracle environment).map (fun Y => Y 0) = roundKernel oracle environment (initialInput m)","url":"../modules/banditrlproof-algorithms-cucbtrajectory/index.html#decl-e54098d3588a","parent":"module:BanditRLProof.Algorithms.CUCBTrajectory","order":376,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTrajectory"],["Source","BanditRLProof/Algorithms/CUCBTrajectory.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cucbTrajectory_initial_law (oracle : Kernel (Input m) A) (environment : Kernel A (Feedback m)) [IsMarkovKernel oracle] [IsMarkovKernel environment] : (cucbTrajectory oracle environment).map (fun Y => Y 0) = roundKernel oracle environment (initialInput m)","missing":[],"search":"cucbtrajectory_initial_law banditrlproof.cucb.cucbtrajectory_initial_law theorem cucbtrajectory_initial_law (oracle : kernel (input m) a) (environment : kernel a (feedback m)) [ismarkovkernel oracle] [ismarkovkernel environment] : (cucbtrajectory oracle environment).map (fun y => y 0) = roundkernel oracle environment (initialinput m) theorem compiled","shard":"modules/98705dd65b9428c0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.triggerFactor","label":"triggerFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.triggerFactor","description":"noncomputable def triggerFactor {m : ℕ} (i : Fin m) (p tilt : ℝ) (z : Feedback m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-55d1dfb9a95e","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":377,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def triggerFactor {m : ℕ} (i : Fin m) (p tilt : ℝ) (z : Feedback m) : ℝ","missing":[],"search":"triggerfactor banditrlproof.cucb.triggerfactor noncomputable def triggerfactor {m : ℕ} (i : fin m) (p tilt : ℝ) (z : feedback m) : ℝ definition compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.measurable_triggerFactor","label":"measurable_triggerFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.measurable_triggerFactor","description":"theorem measurable_triggerFactor {m : ℕ} (i : Fin m) (p tilt : ℝ) : Measurable (triggerFactor i p tilt)","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-3d280ca6fedb","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":378,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_triggerFactor {m : ℕ} (i : Fin m) (p tilt : ℝ) : Measurable (triggerFactor i p tilt)","missing":[],"search":"measurable_triggerfactor banditrlproof.cucb.measurable_triggerfactor theorem measurable_triggerfactor {m : ℕ} (i : fin m) (p tilt : ℝ) : measurable (triggerfactor i p tilt) theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.triggerFactor_nonneg","label":"triggerFactor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.triggerFactor_nonneg","description":"theorem triggerFactor_nonneg {m : ℕ} (i : Fin m) (p tilt : ℝ) (z : Feedback m) : 0≤triggerFactor i p tilt z","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-b2d920f0f1bc","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":379,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem triggerFactor_nonneg {m : ℕ} (i : Fin m) (p tilt : ℝ) (z : Feedback m) : 0≤triggerFactor i p tilt z","missing":[],"search":"triggerfactor_nonneg banditrlproof.cucb.triggerfactor_nonneg theorem triggerfactor_nonneg {m : ℕ} (i : fin m) (p tilt : ℝ) (z : feedback m) : 0≤triggerfactor i p tilt z theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.triggerFactor_le","label":"triggerFactor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.triggerFactor_le","description":"theorem triggerFactor_le {m : ℕ} (i : Fin m) (p tilt : ℝ) (z : Feedback m) : triggerFactor i p tilt z≤Real.exp (|tilt|+|(1-Real.exp (-tilt))*p|)","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-c69c28b2f74e","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":380,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem triggerFactor_le {m : ℕ} (i : Fin m) (p tilt : ℝ) (z : Feedback m) : triggerFactor i p tilt z≤Real.exp (|tilt|+|(1-Real.exp (-tilt))*p|)","missing":[],"search":"triggerfactor_le banditrlproof.cucb.triggerfactor_le theorem triggerfactor_le {m : ℕ} (i : fin m) (p tilt : ℝ) (z : feedback m) : triggerfactor i p tilt z≤real.exp (|tilt|+|(1-real.exp (-tilt))*p|) theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.triggerFactor_piecewise","label":"triggerFactor_piecewise","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.triggerFactor_piecewise","description":"theorem triggerFactor_piecewise {m : ℕ} (i : Fin m) (p tilt : ℝ) [DecidablePred (· ∈ observedSet i)] : triggerFactor i p tilt = (observedSet i).piecewise (fun _ => Real.exp (-tilt+(1-Real.exp (-tilt))*p)) (fun _ => Real.exp ((1-Real.exp (-tilt))*p))","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-65cfee42d8f8","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":381,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem triggerFactor_piecewise {m : ℕ} (i : Fin m) (p tilt : ℝ) [DecidablePred (· ∈ observedSet i)] : triggerFactor i p tilt = (observedSet i).piecewise (fun _ => Real.exp (-tilt+(1-Real.exp (-tilt))*p)) (fun _ => Real.exp ((1-Real.exp (-tilt))*p))","missing":[],"search":"triggerfactor_piecewise banditrlproof.cucb.triggerfactor_piecewise theorem triggerfactor_piecewise {m : ℕ} (i : fin m) (p tilt : ℝ) [decidablepred (· ∈ observedset i)] : triggerfactor i p tilt = (observedset i).piecewise (fun _ => real.exp (-tilt+(1-real.exp (-tilt))*p)) (fun _ => real.exp ((1-real.exp (-tilt))*p)) theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integrable_triggerFactor","label":"integrable_triggerFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integrable_triggerFactor","description":"theorem integrable_triggerFactor {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (i : Fin m) (p tilt : ℝ) : Integrable (triggerFactor i p tilt) ν","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-17606f3515bb","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":382,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_triggerFactor {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (i : Fin m) (p tilt : ℝ) : Integrable (triggerFactor i p tilt) ν","missing":[],"search":"integrable_triggerfactor banditrlproof.cucb.integrable_triggerfactor theorem integrable_triggerfactor {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (i : fin m) (p tilt : ℝ) : integrable (triggerfactor i p tilt) ν theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integral_triggerFactor","label":"integral_triggerFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integral_triggerFactor","description":"theorem integral_triggerFactor {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (i : Fin m) (p tilt : ℝ) : (∫z, triggerFactor i p tilt z ∂ν) = (1-(ν (observedSet i)).toReal*(1-Real.exp (-tilt)))* Real.exp ((1-Real.exp (-tilt))*p)","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-4265724e6c32","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":383,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_triggerFactor {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (i : Fin m) (p tilt : ℝ) : (∫z, triggerFactor i p tilt z ∂ν) = (1-(ν (observedSet i)).toReal*(1-Real.exp (-tilt)))* Real.exp ((1-Real.exp (-tilt))*p)","missing":[],"search":"integral_triggerfactor banditrlproof.cucb.integral_triggerfactor theorem integral_triggerfactor {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (i : fin m) (p tilt : ℝ) : (∫z, triggerfactor i p tilt z ∂ν) = (1-(ν (observedset i)).toreal*(1-real.exp (-tilt)))* real.exp ((1-real.exp (-tilt))*p) theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.integral_triggerFactor_le_one","label":"integral_triggerFactor_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.integral_triggerFactor_le_one","description":"theorem integral_triggerFactor_le_one {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (i : Fin m) (p tilt : ℝ) (hp : p≤(ν (observedSet i)).toReal) (htilt : 0≤tilt) : (∫z, triggerFactor i p tilt z ∂ν)≤1","url":"../modules/banditrlproof-algorithms-cucbtriggermgf/index.html#decl-808fcf318cc4","parent":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","order":384,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBTriggerMGF"],["Source","BanditRLProof/Algorithms/CUCBTriggerMGF.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_triggerFactor_le_one {m : ℕ} (ν : Measure (Feedback m)) [IsProbabilityMeasure ν] (i : Fin m) (p tilt : ℝ) (hp : p≤(ν (observedSet i)).toReal) (htilt : 0≤tilt) : (∫z, triggerFactor i p tilt z ∂ν)≤1","missing":[],"search":"integral_triggerfactor_le_one banditrlproof.cucb.integral_triggerfactor_le_one theorem integral_triggerfactor_le_one {m : ℕ} (ν : measure (feedback m)) [isprobabilitymeasure ν] (i : fin m) (p tilt : ℝ) (hp : p≤(ν (observedset i)).toreal) (htilt : 0≤tilt) : (∫z, triggerfactor i p tilt z ∂ν)≤1 theorem compiled","shard":"modules/e593aad3b3b259c2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_monotone","label":"counters_monotone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_monotone","description":"theorem counters_monotone (actions : ℕ → A) (i : Fin m) : Monotone (fun n => C.counters actions n i)","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-0a96e2ed9277","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":385,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_monotone (actions : ℕ → A) (i : Fin m) : Monotone (fun n => C.counters actions n i)","missing":[],"search":"counters_monotone banditrlproof.cucb.chargedata.counters_monotone theorem counters_monotone (actions : ℕ → a) (i : fin m) : monotone (fun n => c.counters actions n i) theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_strict_of_charge","label":"counters_strict_of_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_strict_of_charge","description":"theorem counters_strict_of_charge (actions : ℕ → A) (i : Fin m) {t s : ℕ} (hts : t<s) (ht : C.choose (C.counters actions t) (actions t)=some i) : C.counters actions t i<C.counters actions s i","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-920a0ce03ad9","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":386,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_strict_of_charge (actions : ℕ → A) (i : Fin m) {t s : ℕ} (hts : t<s) (ht : C.choose (C.counters actions t) (actions t)=some i) : C.counters actions t i<C.counters actions s i","missing":[],"search":"counters_strict_of_charge banditrlproof.cucb.chargedata.counters_strict_of_charge theorem counters_strict_of_charge (actions : ℕ → a) (i : fin m) {t s : ℕ} (hts : t<s) (ht : c.choose (c.counters actions t) (actions t)=some i) : c.counters actions t i<c.counters actions s i theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.counters_injOn_charges","label":"counters_injOn_charges","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.counters_injOn_charges","description":"theorem counters_injOn_charges (actions : ℕ → A) (i : Fin m) : Set.InjOn (fun t => C.counters actions t i) {t | C.choose (C.counters actions t) (actions t)=some i}","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-ae9bd002d678","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":387,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem counters_injOn_charges (actions : ℕ → A) (i : Fin m) : Set.InjOn (fun t => C.counters actions t i) {t | C.choose (C.counters actions t) (actions t)=some i}","missing":[],"search":"counters_injon_charges banditrlproof.cucb.chargedata.counters_injon_charges theorem counters_injon_charges (actions : ℕ → a) (i : fin m) : set.injon (fun t => c.counters actions t i) {t | c.choose (c.counters actions t) (actions t)=some i} theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.ChargeData.card_charges_le","label":"card_charges_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.ChargeData.card_charges_le","description":"theorem card_charges_le (actions : ℕ → A) (i : Fin m) (times : Finset ℕ) (B : ℝ) (hB : 0≤B) (hc : ∀t∈times, C.choose (C.counters actions t) (actions t)=some i) (hb : ∀t∈times, (C.counters actions t i : ℝ)≤B) : (times.card:ℝ)≤B+1","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-c558ea918fc1","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":388,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem card_charges_le (actions : ℕ → A) (i : Fin m) (times : Finset ℕ) (B : ℝ) (hB : 0≤B) (hc : ∀t∈times, C.choose (C.counters actions t) (actions t)=some i) (hb : ∀t∈times, (C.counters actions t i : ℝ)≤B) : (times.card:ℝ)≤B+1","missing":[],"search":"card_charges_le banditrlproof.cucb.chargedata.card_charges_le theorem card_charges_le (actions : ℕ → a) (i : fin m) (times : finset ℕ) (b : ℝ) (hb : 0≤b) (hc : ∀t∈times, c.choose (c.counters actions t) (actions t)=some i) (hb : ∀t∈times, (c.counters actions t i : ℝ)≤b) : (times.card:ℝ)≤b+1 theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes","label":"underChargeTimes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeTimes","description":"noncomputable def underChargeTimes (H : ℕ) (actions : ℕ → A) (i : Fin m) : Finset ℕ","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-acf8088b6484","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":389,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def underChargeTimes (H : ℕ) (actions : ℕ → A) (i : Fin m) : Finset ℕ","missing":[],"search":"underchargetimes banditrlproof.cucb.sourcemodel.underchargetimes noncomputable def underchargetimes (h : ℕ) (actions : ℕ → a) (i : fin m) : finset ℕ definition compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes_spec","label":"underChargeTimes_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeTimes_spec","description":"theorem underChargeTimes_spec (H : ℕ) (actions : ℕ → A) (i : Fin m) {t : ℕ} (ht : t∈S.underChargeTimes H actions i) : t<H ∧ S.chargeData.choose (S.chargeData.counters actions t) (actions t)=some i ∧ 0<S.gap (actions t) ∧ i∈M.possible (actions t) ∧ (S.chargeData.counters actions t i:ℝ)≤S.gapThreshold H (M.minTrigger i) (S.gap (actions t))","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-9e5c22f0e07a","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":390,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underChargeTimes_spec (H : ℕ) (actions : ℕ → A) (i : Fin m) {t : ℕ} (ht : t∈S.underChargeTimes H actions i) : t<H ∧ S.chargeData.choose (S.chargeData.counters actions t) (actions t)=some i ∧ 0<S.gap (actions t) ∧ i∈M.possible (actions t) ∧ (S.chargeData.counters actions t i:ℝ)≤S.gapThreshold H (M.minTrigger i) (S.gap (actions t))","missing":[],"search":"underchargetimes_spec banditrlproof.cucb.sourcemodel.underchargetimes_spec theorem underchargetimes_spec (h : ℕ) (actions : ℕ → a) (i : fin m) {t : ℕ} (ht : t∈s.underchargetimes h actions i) : t<h ∧ s.chargedata.choose (s.chargedata.counters actions t) (actions t)=some i ∧ 0<s.gap (actions t) ∧ i∈m.possible (actions t) ∧ (s.chargedata.counters actions t i:ℝ)≤s.gapthreshold h (m.mintrigger i) (s.gap (actions t)) theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeGapTail","label":"underChargeGapTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeGapTail","description":"noncomputable def underChargeGapTail (H : ℕ) (actions : ℕ → A) (i : Fin m) (x : ℝ) : Finset ℕ","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-13d39b2a26c6","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":391,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def underChargeGapTail (H : ℕ) (actions : ℕ → A) (i : Fin m) (x : ℝ) : Finset ℕ","missing":[],"search":"underchargegaptail banditrlproof.cucb.sourcemodel.underchargegaptail noncomputable def underchargegaptail (h : ℕ) (actions : ℕ → a) (i : fin m) (x : ℝ) : finset ℕ definition compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.card_underChargeGapTail_le","label":"card_underChargeGapTail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.card_underChargeGapTail_le","description":"theorem card_underChargeGapTail_le (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) (x : ℝ) (hx : x∈S.gapDomain) : ((S.underChargeGapTail H actions i x).card:ℝ)≤S.gapThreshold H (M.minTrigger i) x+1","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-0e5f6b72350f","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":392,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:80"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem card_underChargeGapTail_le (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) (x : ℝ) (hx : x∈S.gapDomain) : ((S.underChargeGapTail H actions i x).card:ℝ)≤S.gapThreshold H (M.minTrigger i) x+1","missing":[],"search":"card_underchargegaptail_le banditrlproof.cucb.sourcemodel.card_underchargegaptail_le theorem card_underchargegaptail_le (h : ℕ) (hh : 1≤h) (actions : ℕ → a) (i : fin m) (x : ℝ) (hx : x∈s.gapdomain) : ((s.underchargegaptail h actions i x).card:ℝ)≤s.gapthreshold h (m.mintrigger i) x+1 theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.badActions","label":"badActions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.badActions","description":"noncomputable def badActions (i : Fin m) : Finset A","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-05e03d7e3748","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":393,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def badActions (i : Fin m) : Finset A","missing":[],"search":"badactions banditrlproof.cucb.sourcemodel.badactions noncomputable def badactions (i : fin m) : finset a definition compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes_mem_badActions","label":"underChargeTimes_mem_badActions","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeTimes_mem_badActions","description":"theorem underChargeTimes_mem_badActions (H : ℕ) (actions : ℕ → A) (i : Fin m) {t : ℕ} (ht : t∈S.underChargeTimes H actions i) : actions t∈S.badActions i","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-4a95f1712148","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":394,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underChargeTimes_mem_badActions (H : ℕ) (actions : ℕ → A) (i : Fin m) {t : ℕ} (ht : t∈S.underChargeTimes H actions i) : actions t∈S.badActions i","missing":[],"search":"underchargetimes_mem_badactions banditrlproof.cucb.sourcemodel.underchargetimes_mem_badactions theorem underchargetimes_mem_badactions (h : ℕ) (actions : ℕ → a) (i : fin m) {t : ℕ} (ht : t∈s.underchargetimes h actions i) : actions t∈s.badactions i theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes_empty_of_no_bad","label":"underChargeTimes_empty_of_no_bad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeTimes_empty_of_no_bad","description":"theorem underChargeTimes_empty_of_no_bad (H : ℕ) (actions : ℕ → A) (i : Fin m) (hi : S.badActions i=∅) : S.underChargeTimes H actions i=∅","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-a918ec2c3d4d","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":395,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underChargeTimes_empty_of_no_bad (H : ℕ) (actions : ℕ → A) (i : Fin m) (hi : S.badActions i=∅) : S.underChargeTimes H actions i=∅","missing":[],"search":"underchargetimes_empty_of_no_bad banditrlproof.cucb.sourcemodel.underchargetimes_empty_of_no_bad theorem underchargetimes_empty_of_no_bad (h : ℕ) (actions : ℕ → a) (i : fin m) (hi : s.badactions i=∅) : s.underchargetimes h actions i=∅ theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.minBadGap","label":"minBadGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.minBadGap","description":"noncomputable def minBadGap (i : Fin m) (hi : (S.badActions i).Nonempty) : ℝ","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-6a642ba69c96","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":396,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def minBadGap (i : Fin m) (hi : (S.badActions i).Nonempty) : ℝ","missing":[],"search":"minbadgap banditrlproof.cucb.sourcemodel.minbadgap noncomputable def minbadgap (i : fin m) (hi : (s.badactions i).nonempty) : ℝ definition compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.maxBadGap","label":"maxBadGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.maxBadGap","description":"noncomputable def maxBadGap (i : Fin m) (hi : (S.badActions i).Nonempty) : ℝ","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-e2aa06906cb2","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":397,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def maxBadGap (i : Fin m) (hi : (S.badActions i).Nonempty) : ℝ","missing":[],"search":"maxbadgap banditrlproof.cucb.sourcemodel.maxbadgap noncomputable def maxbadgap (i : fin m) (hi : (s.badactions i).nonempty) : ℝ definition compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.badGap_bounds","label":"badGap_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.badGap_bounds","description":"theorem badGap_bounds (i : Fin m) (hi : (S.badActions i).Nonempty) : S.minBadGap i hi∈S.gapDomain ∧ S.maxBadGap i hi∈S.gapDomain ∧ S.minBadGap i hi≤S.maxBadGap i hi","url":"../modules/banditrlproof-algorithms-cucbundercount/index.html#decl-a585d6fd72df","parent":"module:BanditRLProof.Algorithms.CUCBUnderCount","order":398,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCount"],["Source","BanditRLProof/Algorithms/CUCBUnderCount.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem badGap_bounds (i : Fin m) (hi : (S.badActions i).Nonempty) : S.minBadGap i hi∈S.gapDomain ∧ S.maxBadGap i hi∈S.gapDomain ∧ S.minBadGap i hi≤S.maxBadGap i hi","missing":[],"search":"badgap_bounds banditrlproof.cucb.sourcemodel.badgap_bounds theorem badgap_bounds (i : fin m) (hi : (s.badactions i).nonempty) : s.minbadgap i hi∈s.gapdomain ∧ s.maxbadgap i hi∈s.gapdomain ∧ s.minbadgap i hi≤s.maxbadgap i hi theorem compiled","shard":"modules/859dcab10a79ce60.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight","label":"underChargeWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeWeight","description":"noncomputable def underChargeWeight (H : ℕ) (actions : ℕ → A) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-b63fd97df243","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":399,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def underChargeWeight (H : ℕ) (actions : ℕ → A) (i : Fin m) : ℝ","missing":[],"search":"underchargeweight banditrlproof.cucb.sourcemodel.underchargeweight noncomputable def underchargeweight (h : ℕ) (actions : ℕ → a) (i : fin m) : ℝ definition compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight_le_refined","label":"underChargeWeight_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeWeight_le_refined","description":"theorem underChargeWeight_le_refined (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) (hi : (S.badActions i).Nonempty) : S.underChargeWeight H actions i ≤ S.minBadGap i hi*S.gapThreshold H (M.minTrigger i) (S.minBadGap i hi)+ (∫x in S.minBadGap i hi..S.maxBadGap i hi, S.gapThreshold H (M.minTrigger i) x)+ S.maxBadGap i hi","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-7daeb9944e41","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":400,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underChargeWeight_le_refined (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) (hi : (S.badActions i).Nonempty) : S.underChargeWeight H actions i ≤ S.minBadGap i hi*S.gapThreshold H (M.minTrigger i) (S.minBadGap i hi)+ (∫x in S.minBadGap i hi..S.maxBadGap i hi, S.gapThreshold H (M.minTrigger i) x)+ S.maxBadGap i hi","missing":[],"search":"underchargeweight_le_refined banditrlproof.cucb.sourcemodel.underchargeweight_le_refined theorem underchargeweight_le_refined (h : ℕ) (hh : 1≤h) (actions : ℕ → a) (i : fin m) (hi : (s.badactions i).nonempty) : s.underchargeweight h actions i ≤ s.minbadgap i hi*s.gapthreshold h (m.mintrigger i) (s.minbadgap i hi)+ (∫x in s.minbadgap i hi..s.maxbadgap i hi, s.gapthreshold h (m.mintrigger i) x)+ s.maxbadgap i hi theorem compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.armRefinedTerm","label":"armRefinedTerm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.armRefinedTerm","description":"noncomputable def armRefinedTerm (H : ℕ) (i : Fin m) : ℝ","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-fe00a1190ea9","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":401,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def armRefinedTerm (H : ℕ) (i : Fin m) : ℝ","missing":[],"search":"armrefinedterm banditrlproof.cucb.sourcemodel.armrefinedterm noncomputable def armrefinedterm (h : ℕ) (i : fin m) : ℝ definition compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight_le_armRefinedTerm","label":"underChargeWeight_le_armRefinedTerm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underChargeWeight_le_armRefinedTerm","description":"theorem underChargeWeight_le_armRefinedTerm (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) : S.underChargeWeight H actions i≤ S.armRefinedTerm H i+maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-5fe140fccb20","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":402,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underChargeWeight_le_armRefinedTerm (H : ℕ) (hH : 1≤H) (actions : ℕ → A) (i : Fin m) : S.underChargeWeight H actions i≤ S.armRefinedTerm H i+maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"underchargeweight_le_armrefinedterm banditrlproof.cucb.sourcemodel.underchargeweight_le_armrefinedterm theorem underchargeweight_le_armrefinedterm (h : ℕ) (hh : 1≤h) (actions : ℕ → a) (i : fin m) : s.underchargeweight h actions i≤ s.armrefinedterm h i+maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.underSampledGap_eq_sum_arms","label":"underSampledGap_eq_sum_arms","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.underSampledGap_eq_sum_arms","description":"theorem underSampledGap_eq_sum_arms (H : ℕ) (N : Fin m → ℕ) (a : A) : S.underSampledGap H N a=∑i:Fin m, if S.chargeData.choose N a=some i ∧ (N i:ℝ)≤samplingThreshold H (S.inverseGap a) (M.minTrigger i) then S.gap a else 0","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-0c1c6befc745","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":403,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem underSampledGap_eq_sum_arms (H : ℕ) (N : Fin m → ℕ) (a : A) : S.underSampledGap H N a=∑i:Fin m, if S.chargeData.choose N a=some i ∧ (N i:ℝ)≤samplingThreshold H (S.inverseGap a) (M.minTrigger i) then S.gap a else 0","missing":[],"search":"undersampledgap_eq_sum_arms banditrlproof.cucb.sourcemodel.undersampledgap_eq_sum_arms theorem undersampledgap_eq_sum_arms (h : ℕ) (n : fin m → ℕ) (a : a) : s.undersampledgap h n a=∑i:fin m, if s.chargedata.choose n a=some i ∧ (n i:ℝ)≤samplingthreshold h (s.inversegap a) (m.mintrigger i) then s.gap a else 0 theorem compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sum_underSampledGap_eq_weights","label":"sum_underSampledGap_eq_weights","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sum_underSampledGap_eq_weights","description":"theorem sum_underSampledGap_eq_weights (H : ℕ) (actions : ℕ → A) : (∑t∈Finset.range H, S.underSampledGap H (S.chargeData.counters actions t) (actions t))= ∑i:Fin m, S.underChargeWeight H actions i","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-09235268b064","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":404,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_underSampledGap_eq_weights (H : ℕ) (actions : ℕ → A) : (∑t∈Finset.range H, S.underSampledGap H (S.chargeData.counters actions t) (actions t))= ∑i:Fin m, S.underChargeWeight H actions i","missing":[],"search":"sum_undersampledgap_eq_weights banditrlproof.cucb.sourcemodel.sum_undersampledgap_eq_weights theorem sum_undersampledgap_eq_weights (h : ℕ) (actions : ℕ → a) : (∑t∈finset.range h, s.undersampledgap h (s.chargedata.counters actions t) (actions t))= ∑i:fin m, s.underchargeweight h actions i theorem compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CUCB.SourceModel.sum_underSampledGap_le_refined","label":"sum_underSampledGap_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CUCB.SourceModel.sum_underSampledGap_le_refined","description":"theorem sum_underSampledGap_le_refined (H : ℕ) (hH : 1≤H) (actions : ℕ → A) : (∑t∈Finset.range H, S.underSampledGap H (S.chargeData.counters actions t) (actions t))≤ (∑i:Fin m, S.armRefinedTerm H i)+(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","url":"../modules/banditrlproof-algorithms-cucbundercountintegral/index.html#decl-34409d8c17ba","parent":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","order":405,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CUCBUnderCountIntegral"],["Source","BanditRLProof/Algorithms/CUCBUnderCountIntegral.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_underSampledGap_le_refined (H : ℕ) (hH : 1≤H) (actions : ℕ → A) : (∑t∈Finset.range H, S.underSampledGap H (S.chargeData.counters actions t) (actions t))≤ (∑i:Fin m, S.armRefinedTerm H i)+(m:ℝ)*maxPositiveGap S.score M.trueInput S.alpha","missing":[],"search":"sum_undersampledgap_le_refined banditrlproof.cucb.sourcemodel.sum_undersampledgap_le_refined theorem sum_undersampledgap_le_refined (h : ℕ) (hh : 1≤h) (actions : ℕ → a) : (∑t∈finset.range h, s.undersampledgap h (s.chargedata.counters actions t) (actions t))≤ (∑i:fin m, s.armrefinedterm h i)+(m:ℝ)*maxpositivegap s.score m.trueinput s.alpha theorem compiled","shard":"modules/f90a78b0865d09b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment_ge_one","label":"secondMoment_ge_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment_ge_one","description":"theorem secondMoment_ge_one {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) : 1 ≤ secondMoment (p a) q","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-91ad4596f0d3","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":406,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem secondMoment_ge_one {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) : 1 ≤ secondMoment (p a) q","missing":[],"search":"secondmoment_ge_one banditrlproof.causal.secondmoment_ge_one theorem secondmoment_ge_one {a z : type*} [fintype z] (p : a → pmf z) (q : pmf z) (hc : covers p q) (a : a) : 1 ≤ secondmoment (p a) q theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.uniform_mass","label":"uniform_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.uniform_mass","description":"theorem uniform_mass {A : Type*} [Fintype A] [Nonempty A] (a : A) : mass (PMF.uniformOfFintype A) a = (Fintype.card A : ℝ)⁻¹","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-26f4e2407c97","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":407,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniform_mass {A : Type*} [Fintype A] [Nonempty A] (a : A) : mass (PMF.uniformOfFintype A) a = (Fintype.card A : ℝ)⁻¹","missing":[],"search":"uniform_mass banditrlproof.causal.uniform_mass theorem uniform_mass {a : type*} [fintype a] [nonempty a] (a : a) : mass (pmf.uniformoffintype a) a = (fintype.card a : ℝ)⁻¹ theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.uniform_covers","label":"uniform_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.uniform_covers","description":"theorem uniform_covers {A Z : Type*} [Fintype A] [Nonempty A] (p : A → PMF Z) : Covers p (mixture (PMF.uniformOfFintype A) p)","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-afb48ab5cb07","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":408,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniform_covers {A Z : Type*} [Fintype A] [Nonempty A] (p : A → PMF Z) : Covers p (mixture (PMF.uniformOfFintype A) p)","missing":[],"search":"uniform_covers banditrlproof.causal.uniform_covers theorem uniform_covers {a z : type*} [fintype a] [nonempty a] (p : a → pmf z) : covers p (mixture (pmf.uniformoffintype a) p) theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.uniform_ratio_le_card","label":"uniform_ratio_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.uniform_ratio_le_card","description":"theorem uniform_ratio_le_card {A Z : Type*} [Fintype A] [Nonempty A] (p : A → PMF Z) (a : A) (z : Z) : ratio (p a) (mixture (PMF.uniformOfFintype A) p) z ≤ Fintype.card A","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-0336681e38bd","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":409,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniform_ratio_le_card {A Z : Type*} [Fintype A] [Nonempty A] (p : A → PMF Z) (a : A) (z : Z) : ratio (p a) (mixture (PMF.uniformOfFintype A) p) z ≤ Fintype.card A","missing":[],"search":"uniform_ratio_le_card banditrlproof.causal.uniform_ratio_le_card theorem uniform_ratio_le_card {a z : type*} [fintype a] [nonempty a] (p : a → pmf z) (a : a) (z : z) : ratio (p a) (mixture (pmf.uniformoffintype a) p) z ≤ fintype.card a theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.uniform_secondMoment_le_card","label":"uniform_secondMoment_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.uniform_secondMoment_le_card","description":"theorem uniform_secondMoment_le_card {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (a : A) : secondMoment (p a) (mixture (PMF.uniformOfFintype A) p) ≤ Fintype.card A","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-ce3d106025e3","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":410,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniform_secondMoment_le_card {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (a : A) : secondMoment (p a) (mixture (PMF.uniformOfFintype A) p) ≤ Fintype.card A","missing":[],"search":"uniform_secondmoment_le_card banditrlproof.causal.uniform_secondmoment_le_card theorem uniform_secondmoment_le_card {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) (a : a) : secondmoment (p a) (mixture (pmf.uniformoffintype a) p) ≤ fintype.card a theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.designCost","label":"designCost","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.designCost","description":"noncomputable def designCost {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) : ℝ","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-299c94815a98","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":411,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def designCost {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) : ℝ","missing":[],"search":"designcost banditrlproof.causal.designcost noncomputable def designcost {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) (eta : pmf a) : ℝ definition compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment_le_designCost","label":"secondMoment_le_designCost","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment_le_designCost","description":"theorem secondMoment_le_designCost {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) (a : A) : secondMoment (p a) (mixture eta p) ≤ designCost p eta","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-0ecbdfbc6d9c","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":412,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem secondMoment_le_designCost {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) (a : A) : secondMoment (p a) (mixture eta p) ≤ designCost p eta","missing":[],"search":"secondmoment_le_designcost banditrlproof.causal.secondmoment_le_designcost theorem secondmoment_le_designcost {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) (eta : pmf a) (a : a) : secondmoment (p a) (mixture eta p) ≤ designcost p eta theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.designCost_ge_one","label":"designCost_ge_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.designCost_ge_one","description":"theorem designCost_ge_one {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) : 1 ≤ designCost p eta","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-a3fd3e3c29ea","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":413,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem designCost_ge_one {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) : 1 ≤ designCost p eta","missing":[],"search":"designcost_ge_one banditrlproof.causal.designcost_ge_one theorem designcost_ge_one {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) (eta : pmf a) (hc : covers p (mixture eta p)) : 1 ≤ designcost p eta theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.uniform_designCost_le_card","label":"uniform_designCost_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.uniform_designCost_le_card","description":"theorem uniform_designCost_le_card {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) : designCost p (PMF.uniformOfFintype A) ≤ Fintype.card A","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-4d0afcf87ea9","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":414,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniform_designCost_le_card {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) : designCost p (PMF.uniformOfFintype A) ≤ Fintype.card A","missing":[],"search":"uniform_designcost_le_card banditrlproof.causal.uniform_designcost_le_card theorem uniform_designcost_le_card {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) : designcost p (pmf.uniformoffintype a) ≤ fintype.card a theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.convexLaw","label":"convexLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.convexLaw","description":"noncomputable def convexLaw {Z : Type*} (t : ℝ≥0) (ht : t ≤ 1) (p q : PMF Z) : PMF Z","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-673d1e338aef","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":415,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def convexLaw {Z : Type*} (t : ℝ≥0) (ht : t ≤ 1) (p q : PMF Z) : PMF Z","missing":[],"search":"convexlaw banditrlproof.causal.convexlaw noncomputable def convexlaw {z : type*} (t : ℝ≥0) (ht : t ≤ 1) (p q : pmf z) : pmf z definition compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.convexLaw_mass","label":"convexLaw_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.convexLaw_mass","description":"theorem convexLaw_mass {Z : Type*} (t : ℝ≥0) (ht : t ≤ 1) (p q : PMF Z) (z : Z) : mass (convexLaw t ht p q) z = (t : ℝ) * mass p z + (1-(t : ℝ)) * mass q z","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-11d3d5d729e7","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":416,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem convexLaw_mass {Z : Type*} (t : ℝ≥0) (ht : t ≤ 1) (p q : PMF Z) (z : Z) : mass (convexLaw t ht p q) z = (t : ℝ) * mass p z + (1-(t : ℝ)) * mass q z","missing":[],"search":"convexlaw_mass banditrlproof.causal.convexlaw_mass theorem convexlaw_mass {z : type*} (t : ℝ≥0) (ht : t ≤ 1) (p q : pmf z) (z : z) : mass (convexlaw t ht p q) z = (t : ℝ) * mass p z + (1-(t : ℝ)) * mass q z theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.inverse_convex","label":"inverse_convex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.inverse_convex","description":"theorem inverse_convex (x y t : ℝ) (hx : 0 < x) (hy : 0 < y) (ht : 0 ≤ t) (ht1 : t ≤ 1) : (t*x+(1-t)*y)⁻¹ ≤ t*x⁻¹+(1-t)*y⁻¹","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-737a42f0a660","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":417,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverse_convex (x y t : ℝ) (hx : 0 < x) (hy : 0 < y) (ht : 0 ≤ t) (ht1 : t ≤ 1) : (t*x+(1-t)*y)⁻¹ ≤ t*x⁻¹+(1-t)*y⁻¹","missing":[],"search":"inverse_convex banditrlproof.causal.inverse_convex theorem inverse_convex (x y t : ℝ) (hx : 0 < x) (hy : 0 < y) (ht : 0 ≤ t) (ht1 : t ≤ 1) : (t*x+(1-t)*y)⁻¹ ≤ t*x⁻¹+(1-t)*y⁻¹ theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.convexLaw_covers","label":"convexLaw_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.convexLaw_covers","description":"theorem convexLaw_covers {A Z : Type*} (p : A → PMF Z) (q₀ q₁ : PMF Z) (hc₀ : Covers p q₀) (hc₁ : Covers p q₁) (t : ℝ≥0) (ht : t ≤ 1) : Covers p (convexLaw t ht q₀ q₁)","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-fdc5f3baf8b7","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":418,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem convexLaw_covers {A Z : Type*} (p : A → PMF Z) (q₀ q₁ : PMF Z) (hc₀ : Covers p q₀) (hc₁ : Covers p q₁) (t : ℝ≥0) (ht : t ≤ 1) : Covers p (convexLaw t ht q₀ q₁)","missing":[],"search":"convexlaw_covers banditrlproof.causal.convexlaw_covers theorem convexlaw_covers {a z : type*} (p : a → pmf z) (q₀ q₁ : pmf z) (hc₀ : covers p q₀) (hc₁ : covers p q₁) (t : ℝ≥0) (ht : t ≤ 1) : covers p (convexlaw t ht q₀ q₁) theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment_convex","label":"secondMoment_convex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment_convex","description":"theorem secondMoment_convex {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q₀ q₁ : PMF Z) (hc₀ : Covers p q₀) (hc₁ : Covers p q₁) (a : A) (t : ℝ≥0) (ht : t ≤ 1) : secondMoment (p a) (convexLaw t ht q₀ q₁) ≤ (t : ℝ)*secondMoment (p a) q₀+(1-(t : ℝ))*secondMoment (p a) q₁","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-287b2a033df8","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":419,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem secondMoment_convex {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q₀ q₁ : PMF Z) (hc₀ : Covers p q₀) (hc₁ : Covers p q₁) (a : A) (t : ℝ≥0) (ht : t ≤ 1) : secondMoment (p a) (convexLaw t ht q₀ q₁) ≤ (t : ℝ)*secondMoment (p a) q₀+(1-(t : ℝ))*secondMoment (p a) q₁","missing":[],"search":"secondmoment_convex banditrlproof.causal.secondmoment_convex theorem secondmoment_convex {a z : type*} [fintype z] (p : a → pmf z) (q₀ q₁ : pmf z) (hc₀ : covers p q₀) (hc₁ : covers p q₁) (a : a) (t : ℝ≥0) (ht : t ≤ 1) : secondmoment (p a) (convexlaw t ht q₀ q₁) ≤ (t : ℝ)*secondmoment (p a) q₀+(1-(t : ℝ))*secondmoment (p a) q₁ theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixture_convexLaw","label":"mixture_convexLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mixture_convexLaw","description":"theorem mixture_convexLaw {A Z : Type*} (p : A → PMF Z) (eta₀ eta₁ : PMF A) (t : ℝ≥0) (ht : t ≤ 1) : mixture (convexLaw t ht eta₀ eta₁) p = convexLaw t ht (mixture eta₀ p) (mixture eta₁ p)","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-7b7d04653a70","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":420,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mixture_convexLaw {A Z : Type*} (p : A → PMF Z) (eta₀ eta₁ : PMF A) (t : ℝ≥0) (ht : t ≤ 1) : mixture (convexLaw t ht eta₀ eta₁) p = convexLaw t ht (mixture eta₀ p) (mixture eta₁ p)","missing":[],"search":"mixture_convexlaw banditrlproof.causal.mixture_convexlaw theorem mixture_convexlaw {a z : type*} (p : a → pmf z) (eta₀ eta₁ : pmf a) (t : ℝ≥0) (ht : t ≤ 1) : mixture (convexlaw t ht eta₀ eta₁) p = convexlaw t ht (mixture eta₀ p) (mixture eta₁ p) theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.designCost_convex","label":"designCost_convex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.designCost_convex","description":"theorem designCost_convex {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta₀ eta₁ : PMF A) (hc₀ : Covers p (mixture eta₀ p)) (hc₁ : Covers p (mixture eta₁ p)) (t : ℝ≥0) (ht : t ≤ 1) : designCost p (convexLaw t ht eta₀ eta₁) ≤ (t : ℝ)*designCost p eta₀+(1-(t : ℝ))*designCost p eta₁","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-db790d6141f5","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":421,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem designCost_convex {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta₀ eta₁ : PMF A) (hc₀ : Covers p (mixture eta₀ p)) (hc₁ : Covers p (mixture eta₁ p)) (t : ℝ≥0) (ht : t ≤ 1) : designCost p (convexLaw t ht eta₀ eta₁) ≤ (t : ℝ)*designCost p eta₀+(1-(t : ℝ))*designCost p eta₁","missing":[],"search":"designcost_convex banditrlproof.causal.designcost_convex theorem designcost_convex {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) (eta₀ eta₁ : pmf a) (hc₀ : covers p (mixture eta₀ p)) (hc₁ : covers p (mixture eta₁ p)) (t : ℝ≥0) (ht : t ≤ 1) : designcost p (convexlaw t ht eta₀ eta₁) ≤ (t : ℝ)*designcost p eta₀+(1-(t : ℝ))*designcost p eta₁ theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.design_sublevel_mass_lower","label":"design_sublevel_mass_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.design_sublevel_mass_lower","description":"theorem design_sublevel_mass_lower {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) (K : ℝ) (hK : 0 < K) (hcost : designCost p eta ≤ K) (a : A) (z : Z) : mass (p a) z ^ 2 / K ≤ mass (mixture eta p) z","url":"../modules/banditrlproof-algorithms-causalallocation/index.html#decl-1e73e0304b19","parent":"module:BanditRLProof.Algorithms.CausalAllocation","order":422,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocation"],["Source","BanditRLProof/Algorithms/CausalAllocation.lean:166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem design_sublevel_mass_lower {A Z : Type*} [Fintype A] [Nonempty A] [Fintype Z] (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) (K : ℝ) (hK : 0 < K) (hcost : designCost p eta ≤ K) (a : A) (z : Z) : mass (p a) z ^ 2 / K ≤ mass (mixture eta p) z","missing":[],"search":"design_sublevel_mass_lower banditrlproof.causal.design_sublevel_mass_lower theorem design_sublevel_mass_lower {a z : type*} [fintype a] [nonempty a] [fintype z] (p : a → pmf z) (eta : pmf a) (hc : covers p (mixture eta p)) (k : ℝ) (hk : 0 < k) (hcost : designcost p eta ≤ k) (a : a) (z : z) : mass (p a) z ^ 2 / k ≤ mass (mixture eta p) z theorem compiled","shard":"modules/17b3f0149e2d074b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_uniform","label":"expected_simpleRegret_uniform","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.expected_simpleRegret_uniform","description":"theorem GraphModel.expected_simpleRegret_uniform (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let eta := PMF.uniformOfFintype A let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw a…","url":"../modules/banditrlproof-algorithms-causalallocationregret/index.html#decl-b8dde97459d4","parent":"module:BanditRLProof.Algorithms.CausalAllocationRegret","order":423,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocationRegret"],["Source","BanditRLProof/Algorithms/CausalAllocationRegret.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.expected_simpleRegret_uniform (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let eta := PMF.uniformOfFintype A let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt ((Fintype.card A:ℝ)*L/T) + 1/(T:ℝ)","missing":[],"search":"expected_simpleregret_uniform banditrlproof.causal.graphmodel.expected_simpleregret_uniform theorem graphmodel.expected_simpleregret_uniform (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (i : fin n) (hi : ∀ a, actions a i = none) (t : ℕ) (ht : 0 < t) : let eta := pmf.uniformoffintype a let m := designcost (fun a => g.parentlaw (actions a) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ≤ (2*real.sqrt 2+7)*real.sqrt ((fintype.card a:ℝ)*l/t) + 1/(t:ℝ) theorem compiled","shard":"modules/437c067495ee2b84.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_optimal","label":"expected_simpleRegret_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.expected_simpleRegret_optimal","description":"theorem GraphModel.expected_simpleRegret_optimal (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let p := fun a => g.parentLaw (actions a) i let eta := optimalAllocation p let m := designCost p eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampl…","url":"../modules/banditrlproof-algorithms-causalallocationregret/index.html#decl-6f061b632f11","parent":"module:BanditRLProof.Algorithms.CausalAllocationRegret","order":424,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalAllocationRegret"],["Source","BanditRLProof/Algorithms/CausalAllocationRegret.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.expected_simpleRegret_optimal (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let p := fun a => g.parentLaw (actions a) i let eta := optimalAllocation p let m := designCost p eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt (m*L/T) + 1/(T:ℝ)","missing":[],"search":"expected_simpleregret_optimal banditrlproof.causal.graphmodel.expected_simpleregret_optimal theorem graphmodel.expected_simpleregret_optimal (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (i : fin n) (hi : ∀ a, actions a i = none) (t : ℕ) (ht : 0 < t) : let p := fun a => g.parentlaw (actions a) i let eta := optimalallocation p let m := designcost p eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ≤ (2*real.sqrt 2+7)*real.sqrt (m*l/t) + 1/(t:ℝ) theorem compiled","shard":"modules/437c067495ee2b84.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate","label":"sampleEstimate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleEstimate","description":"noncomputable def GraphModel.sampleEstimate (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : ℝ","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-16ec1a98e062","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":425,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def GraphModel.sampleEstimate (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : ℝ","missing":[],"search":"sampleestimate banditrlproof.causal.graphmodel.sampleestimate noncomputable def graphmodel.sampleestimate (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (a : a) (b : ℝ) {t : ℕ} (w : fin t → a × (fin n → v)) : ℝ definition compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.estimateCenter","label":"estimateCenter","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.estimateCenter","description":"noncomputable def GraphModel.estimateCenter (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-40396598bcc6","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":426,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def GraphModel.estimateCenter (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) : ℝ","missing":[],"search":"estimatecenter banditrlproof.causal.graphmodel.estimatecenter noncomputable def graphmodel.estimatecenter (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (a : a) (b : ℝ) : ℝ definition compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_centered_sum","label":"sampleEstimate_centered_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleEstimate_centered_sum","description":"theorem GraphModel.sampleEstimate_centered_sum (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B c sign : ℝ) (T : ℕ) (hT : 0 < T) (w : Fin T → A × (Fin n → V)) : (∑ t : Fin T, sign*(g.sampleWeightedBit rewardBit actions eta i a B t w-c)) = (T:ℝ) * (sign*(g.sampleEstimate rewardBit actions eta i a B w-c))","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-6ed50e048cb6","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":427,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleEstimate_centered_sum (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B c sign : ℝ) (T : ℕ) (hT : 0 < T) (w : Fin T → A × (Fin n → V)) : (∑ t : Fin T, sign*(g.sampleWeightedBit rewardBit actions eta i a B t w-c)) = (T:ℝ) * (sign*(g.sampleEstimate rewardBit actions eta i a B w-c))","missing":[],"search":"sampleestimate_centered_sum banditrlproof.causal.graphmodel.sampleestimate_centered_sum theorem graphmodel.sampleestimate_centered_sum (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (a : a) (b c sign : ℝ) (t : ℕ) (ht : 0 < t) (w : fin t → a × (fin n → v)) : (∑ t : fin t, sign*(g.sampleweightedbit rewardbit actions eta i a b t w-c)) = (t:ℝ) * (sign*(g.sampleestimate rewardbit actions eta i a b w-c)) theorem compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_signed_tail","label":"sampleEstimate_signed_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleEstimate_signed_tail","description":"theorem GraphModel.sampleEstimate_signed_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (T : ℕ) (hT : 0 < T) (L : ℝ) (hL : 0 < L) (sign : ℝ) (hs : |sign| = 1) : let m := designCost (fun b => g.parentLaw (actions b) i)…","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-a144f4fecf2e","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":428,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleEstimate_signed_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (T : ℕ) (hT : 0 < T) (L : ℝ) (hL : 0 < L) (sign : ℝ) (hs : |sign| = 1) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let B := sourceThreshold m T L (g.sampleLaw actions eta T).real {w | sourceRadius m T L ≤ sign*(g.sampleEstimate rewardBit actions eta i a B w-g.estimateCenter rewardBit actions eta i a B)} ≤ Real.exp (-L)","missing":[],"search":"sampleestimate_signed_tail banditrlproof.causal.graphmodel.sampleestimate_signed_tail theorem graphmodel.sampleestimate_signed_tail (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (a : a) (t : ℕ) (ht : 0 < t) (l : ℝ) (hl : 0 < l) (sign : ℝ) (hs : |sign| = 1) : let m := designcost (fun b => g.parentlaw (actions b) i) eta let b := sourcethreshold m t l (g.samplelaw actions eta t).real {w | sourceradius m t l ≤ sign*(g.sampleestimate rewardbit actions eta i a b w-g.estimatecenter rewardbit actions eta i a b)} ≤ real.exp (-l) theorem compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_abs_tail","label":"sampleEstimate_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleEstimate_abs_tail","description":"theorem GraphModel.sampleEstimate_abs_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (T : ℕ) (hT : 0 < T) (L : ℝ) (hL : 0 < L) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let B := sourceThreshold m…","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-dccf772ad549","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":429,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleEstimate_abs_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (T : ℕ) (hT : 0 < T) (L : ℝ) (hL : 0 < L) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let B := sourceThreshold m T L (g.sampleLaw actions eta T).real {w | sourceRadius m T L ≤ |g.sampleEstimate rewardBit actions eta i a B w-g.estimateCenter rewardBit actions eta i a B|} ≤ 2*Real.exp (-L)","missing":[],"search":"sampleestimate_abs_tail banditrlproof.causal.graphmodel.sampleestimate_abs_tail theorem graphmodel.sampleestimate_abs_tail (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (a : a) (t : ℕ) (ht : 0 < t) (l : ℝ) (hl : 0 < l) : let m := designcost (fun b => g.parentlaw (actions b) i) eta let b := sourcethreshold m t l (g.samplelaw actions eta t).real {w | sourceradius m t l ≤ |g.sampleestimate rewardbit actions eta i a b w-g.estimatecenter rewardbit actions eta i a b|} ≤ 2*real.exp (-l) theorem compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_simultaneous_tail","label":"sampleEstimate_simultaneous_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleEstimate_simultaneous_tail","description":"theorem GraphModel.sampleEstimate_simultaneous_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) (L : ℝ) (hL : 0 < L) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let B := sourceThreshold m…","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-329db863d94b","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":430,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:107"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleEstimate_simultaneous_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) (L : ℝ) (hL : 0 < L) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let B := sourceThreshold m T L (g.sampleLaw actions eta T).real {w | ∃ a : A, sourceRadius m T L ≤ |g.sampleEstimate rewardBit actions eta i a B w-g.estimateCenter rewardBit actions eta i a B|} ≤ (Fintype.card A:ℝ) * (2*Real.exp (-L))","missing":[],"search":"sampleestimate_simultaneous_tail banditrlproof.causal.graphmodel.sampleestimate_simultaneous_tail theorem graphmodel.sampleestimate_simultaneous_tail (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (t : ℕ) (ht : 0 < t) (l : ℝ) (hl : 0 < l) : let m := designcost (fun b => g.parentlaw (actions b) i) eta let b := sourcethreshold m t l (g.samplelaw actions eta t).real {w | ∃ a : a, sourceradius m t l ≤ |g.sampleestimate rewardbit actions eta i a b w-g.estimatecenter rewardbit actions eta i a b|} ≤ (fintype.card a:ℝ) * (2*real.exp (-l)) theorem compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_source_confidence","label":"sampleEstimate_source_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleEstimate_source_confidence","description":"theorem GraphModel.sampleEstimate_source_confidence (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let L := sourceLog T (Fintype.card A) let B :=…","url":"../modules/banditrlproof-algorithms-causalconfidence/index.html#decl-9a3bae04b761","parent":"module:BanditRLProof.Algorithms.CausalConfidence","order":431,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalConfidence"],["Source","BanditRLProof/Algorithms/CausalConfidence.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.sampleEstimate_source_confidence (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun b => g.parentLaw (actions b) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (g.sampleLaw actions eta T).real {w | ∃ a : A, sourceRadius m T L ≤ |g.sampleEstimate rewardBit actions eta i a B w-g.estimateCenter rewardBit actions eta i a B|} ≤ 1/(T:ℝ)","missing":[],"search":"sampleestimate_source_confidence banditrlproof.causal.graphmodel.sampleestimate_source_confidence theorem graphmodel.sampleestimate_source_confidence (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (t : ℕ) (ht : 0 < t) : let m := designcost (fun b => g.parentlaw (actions b) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (g.samplelaw actions eta t).real {w | ∃ a : a, sourceradius m t l ≤ |g.sampleestimate rewardbit actions eta i a b w-g.estimatecenter rewardbit actions eta i a b|} ≤ 1/(t:ℝ) theorem compiled","shard":"modules/660aa4f2c75caaef.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_le_tuned","label":"expected_simpleRegret_le_tuned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.expected_simpleRegret_le_tuned","description":"theorem GraphModel.expected_simpleRegret_le_tuned (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := so…","url":"../modules/banditrlproof-algorithms-causalexpectedregret/index.html#decl-0296c9be3159","parent":"module:BanditRLProof.Algorithms.CausalExpectedRegret","order":432,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalExpectedRegret"],["Source","BanditRLProof/Algorithms/CausalExpectedRegret.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.expected_simpleRegret_le_tuned (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ 2*sourceRadius m T L + m/B + 1/(T:ℝ)","missing":[],"search":"expected_simpleregret_le_tuned banditrlproof.causal.graphmodel.expected_simpleregret_le_tuned theorem graphmodel.expected_simpleregret_le_tuned (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (t : ℕ) (ht : 0 < t) : let m := designcost (fun a => g.parentlaw (actions a) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ≤ 2*sourceradius m t l + m/b + 1/(t:ℝ) theorem compiled","shard":"modules/f9be698157d645d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_bounds","label":"expected_simpleRegret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.expected_simpleRegret_bounds","description":"theorem GraphModel.expected_simpleRegret_bounds (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) (T : ℕ) : 0 ≤ (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ∧ (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ 1","url":"../modules/banditrlproof-algorithms-causalexpectedregret/index.html#decl-96e1f139d261","parent":"module:BanditRLProof.Algorithms.CausalExpectedRegret","order":433,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalExpectedRegret"],["Source","BanditRLProof/Algorithms/CausalExpectedRegret.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.expected_simpleRegret_bounds (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) (T : ℕ) : 0 ≤ (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ∧ (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ 1","missing":[],"search":"expected_simpleregret_bounds banditrlproof.causal.graphmodel.expected_simpleregret_bounds theorem graphmodel.expected_simpleregret_bounds (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (b : ℝ) (t : ℕ) : 0 ≤ (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ∧ (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ≤ 1 theorem compiled","shard":"modules/f9be698157d645d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_source_bound","label":"expected_simpleRegret_source_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.expected_simpleRegret_source_bound","description":"theorem GraphModel.expected_simpleRegret_source_bound (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B :…","url":"../modules/banditrlproof-algorithms-causalexpectedregret/index.html#decl-43ac7c7a5ffe","parent":"module:BanditRLProof.Algorithms.CausalExpectedRegret","order":434,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalExpectedRegret"],["Source","BanditRLProof/Algorithms/CausalExpectedRegret.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.expected_simpleRegret_source_bound (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt (m*L/T) + 1/(T:ℝ)","missing":[],"search":"expected_simpleregret_source_bound banditrlproof.causal.graphmodel.expected_simpleregret_source_bound theorem graphmodel.expected_simpleregret_source_bound (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (t : ℕ) (ht : 0 < t) : let m := designcost (fun a => g.parentlaw (actions a) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ≤ (2*real.sqrt 2+7)*real.sqrt (m*l/t) + 1/(t:ℝ) theorem compiled","shard":"modules/f9be698157d645d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_explicit_rate","label":"expected_simpleRegret_explicit_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.expected_simpleRegret_explicit_rate","description":"theorem GraphModel.expected_simpleRegret_explicit_rate (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B…","url":"../modules/banditrlproof-algorithms-causalexpectedregret/index.html#decl-7a5014ebdad6","parent":"module:BanditRLProof.Algorithms.CausalExpectedRegret","order":435,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalExpectedRegret"],["Source","BanditRLProof/Algorithms/CausalExpectedRegret.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.expected_simpleRegret_explicit_rate (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret rewardBit actions eta i B w ∂g.sampleLaw actions eta T) ≤ (3*Real.sqrt 2+7)*Real.sqrt (m*L/T)","missing":[],"search":"expected_simpleregret_explicit_rate banditrlproof.causal.graphmodel.expected_simpleregret_explicit_rate theorem graphmodel.expected_simpleregret_explicit_rate (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (t : ℕ) (ht : 0 < t) : let m := designcost (fun a => g.parentlaw (actions a) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret rewardbit actions eta i b w ∂g.samplelaw actions eta t) ≤ (3*real.sqrt 2+7)*real.sqrt (m*l/t) theorem compiled","shard":"modules/f9be698157d645d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeHistory","label":"NodeHistory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Causal.NodeHistory","description":"abbrev NodeHistory {n : ℕ} (V : Fin n → Type*) (i : Fin n)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-4f988a5d3d35","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":436,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev NodeHistory {n : ℕ} (V : Fin n → Type*) (i : Fin n)","missing":[],"search":"nodehistory banditrlproof.causal.nodehistory abbrev nodehistory {n : ℕ} (v : fin n → type*) (i : fin n) abbreviation compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeTables","label":"NodeTables","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Causal.NodeTables","description":"abbrev NodeTables {n : ℕ} (V : Fin n → Type*)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-7ac6a79cf34e","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":437,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev NodeTables {n : ℕ} (V : Fin n → Type*)","missing":[],"search":"nodetables banditrlproof.causal.nodetables abbrev nodetables {n : ℕ} (v : fin n → type*) abbreviation compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeTables.prefix","label":"prefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeTables.prefix","description":"def NodeTables.prefix {n : ℕ} {V : Fin (n+1) → Type*} (p : NodeTables V) : NodeTables (fun i : Fin n => V i.castSucc)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-1df31fac807d","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":438,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeTables.prefix {n : ℕ} {V : Fin (n+1) → Type*} (p : NodeTables V) : NodeTables (fun i : Fin n => V i.castSucc)","missing":[],"search":"prefix banditrlproof.causal.nodetables.prefix def nodetables.prefix {n : ℕ} {v : fin (n+1) → type*} (p : nodetables v) : nodetables (fun i : fin n => v i.castsucc) definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.nodeJoint","label":"nodeJoint","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.nodeJoint","description":"noncomputable def nodeJoint : {n : ℕ} → {V : Fin n → Type*} → NodeTables V → PMF ((i : Fin n) → V i) | 0, _, _ => PMF.pure (fun i => Fin.elim0 i) | n+1, _, p => (nodeJoint p.prefix).bind fun h => (p (Fin.last n) h).bind fun x => PMF.pure (Fin.snoc h x) structure NodeCodec {n : ℕ} (V : Fin n → Type*) (W : Type*) where","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-3abab6a01320","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":439,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def nodeJoint : {n : ℕ} → {V : Fin n → Type*} → NodeTables V → PMF ((i : Fin n) → V i) | 0, _, _ => PMF.pure (fun i => Fin.elim0 i) | n+1, _, p => (nodeJoint p.prefix).bind fun h => (p (Fin.last n) h).bind fun x => PMF.pure (Fin.snoc h x) structure NodeCodec {n : ℕ} (V : Fin n → Type*) (W : Type*) where","missing":[],"search":"nodejoint banditrlproof.causal.nodejoint noncomputable def nodejoint : {n : ℕ} → {v : fin n → type*} → nodetables v → pmf ((i : fin n) → v i) | 0, _, _ => pmf.pure (fun i => fin.elim0 i) | n+1, _, p => (nodejoint p.prefix).bind fun h => (p (fin.last n) h).bind fun x => pmf.pure (fin.snoc h x) structure nodecodec {n : ℕ} (v : fin n → type*) (w : type*) where definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeCodec","label":"NodeCodec","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec","description":"structure NodeCodec {n : ℕ} (V : Fin n → Type*) (W : Type*) where","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-c2aaaa005030","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":440,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure NodeCodec {n : ℕ} (V : Fin n → Type*) (W : Type*) where","missing":[],"search":"nodecodec banditrlproof.causal.nodecodec structure nodecodec {n : ℕ} (v : fin n → type*) (w : type*) where structure compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.prefix","label":"prefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.prefix","description":"def NodeCodec.prefix {n : ℕ} {V : Fin (n+1) → Type*} {W : Type*} (c : NodeCodec V W) : NodeCodec (fun i : Fin n => V i.castSucc) W where","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-ee4446ee8229","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":441,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeCodec.prefix {n : ℕ} {V : Fin (n+1) → Type*} {W : Type*} (c : NodeCodec V W) : NodeCodec (fun i : Fin n => V i.castSucc) W where","missing":[],"search":"prefix banditrlproof.causal.nodecodec.prefix def nodecodec.prefix {n : ℕ} {v : fin (n+1) → type*} {w : type*} (c : nodecodec v w) : nodecodec (fun i : fin n => v i.castsucc) w where definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeAssignment","label":"encodeAssignment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeAssignment","description":"def NodeCodec.encodeAssignment {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (x : (i : Fin n) → V i) : Fin n → W","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-c6571eea1ff4","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":442,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeCodec.encodeAssignment {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (x : (i : Fin n) → V i) : Fin n → W","missing":[],"search":"encodeassignment banditrlproof.causal.nodecodec.encodeassignment def nodecodec.encodeassignment {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (x : (i : fin n) → v i) : fin n → w definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.decodeAssignment","label":"decodeAssignment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.decodeAssignment","description":"def NodeCodec.decodeAssignment {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (x : Fin n → W) : (i : Fin n) → V i","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-829e0bf797cb","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":443,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeCodec.decodeAssignment {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (x : Fin n → W) : (i : Fin n) → V i","missing":[],"search":"decodeassignment banditrlproof.causal.nodecodec.decodeassignment def nodecodec.decodeassignment {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (x : fin n → w) : (i : fin n) → v i definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.decode_encodeAssignment","label":"decode_encodeAssignment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.decode_encodeAssignment","description":"theorem NodeCodec.decode_encodeAssignment {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (x : (i : Fin n) → V i) : c.decodeAssignment (c.encodeAssignment x) = x","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-140e6da98819","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":444,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeCodec.decode_encodeAssignment {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (x : (i : Fin n) → V i) : c.decodeAssignment (c.encodeAssignment x) = x","missing":[],"search":"decode_encodeassignment banditrlproof.causal.nodecodec.decode_encodeassignment theorem nodecodec.decode_encodeassignment {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (x : (i : fin n) → v i) : c.decodeassignment (c.encodeassignment x) = x theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeTables","label":"encodeTables","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeTables","description":"noncomputable def NodeCodec.encodeTables {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) : Tables W n","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-2328fe9714cf","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":445,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def NodeCodec.encodeTables {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) : Tables W n","missing":[],"search":"encodetables banditrlproof.causal.nodecodec.encodetables noncomputable def nodecodec.encodetables {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (p : nodetables v) : tables w n definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeAssignment_snoc","label":"encodeAssignment_snoc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeAssignment_snoc","description":"theorem NodeCodec.encodeAssignment_snoc {n : ℕ} {V : Fin (n+1) → Type*} {W : Type*} (c : NodeCodec V W) (h : (i : Fin n) → V i.castSucc) (x : V (Fin.last n)) : c.encodeAssignment (Fin.snoc h x) = Fin.snoc (c.prefix.encodeAssignment h) (c.encode (Fin.last n) x)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-4bd2a1e4b444","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":446,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeCodec.encodeAssignment_snoc {n : ℕ} {V : Fin (n+1) → Type*} {W : Type*} (c : NodeCodec V W) (h : (i : Fin n) → V i.castSucc) (x : V (Fin.last n)) : c.encodeAssignment (Fin.snoc h x) = Fin.snoc (c.prefix.encodeAssignment h) (c.encode (Fin.last n) x)","missing":[],"search":"encodeassignment_snoc banditrlproof.causal.nodecodec.encodeassignment_snoc theorem nodecodec.encodeassignment_snoc {n : ℕ} {v : fin (n+1) → type*} {w : type*} (c : nodecodec v w) (h : (i : fin n) → v i.castsucc) (x : v (fin.last n)) : c.encodeassignment (fin.snoc h x) = fin.snoc (c.prefix.encodeassignment h) (c.encode (fin.last n) x) theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.joint_encodeTables","label":"joint_encodeTables","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.joint_encodeTables","description":"theorem NodeCodec.joint_encodeTables {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) : joint (c.encodeTables p) = (nodeJoint p).map c.encodeAssignment","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-4543e14fe847","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":447,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeCodec.joint_encodeTables {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) : joint (c.encodeTables p) = (nodeJoint p).map c.encodeAssignment","missing":[],"search":"joint_encodetables banditrlproof.causal.nodecodec.joint_encodetables theorem nodecodec.joint_encodetables {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (p : nodetables v) : joint (c.encodetables p) = (nodejoint p).map c.encodeassignment theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.decode_joint_encodeTables","label":"decode_joint_encodeTables","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.decode_joint_encodeTables","description":"theorem NodeCodec.decode_joint_encodeTables {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) : (joint (c.encodeTables p)).map c.decodeAssignment = nodeJoint p","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-040aec1b98db","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":448,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeCodec.decode_joint_encodeTables {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) : (joint (c.encodeTables p)).map c.decodeAssignment = nodeJoint p","missing":[],"search":"decode_joint_encodetables banditrlproof.causal.nodecodec.decode_joint_encodetables theorem nodecodec.decode_joint_encodetables {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (p : nodetables v) : (joint (c.encodetables p)).map c.decodeassignment = nodejoint p theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.productNodeCodec","label":"productNodeCodec","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.productNodeCodec","description":"A finite common alphabet is constructed, not assumed: use the finite product. Each node stores its value at its own coordinate and defaults elsewhere.","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-1242ef3d239e","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":449,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def productNodeCodec {n : ℕ} (V : Fin n → Type*) [∀ i, Inhabited (V i)] : NodeCodec V ((i : Fin n) → V i) where","missing":[],"search":"productnodecodec banditrlproof.causal.productnodecodec a finite common alphabet is constructed, not assumed: use the finite product. each node stores its value at its own coordinate and defaults elsewhere. definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.nodeIntervene","label":"nodeIntervene","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.nodeIntervene","description":"noncomputable def nodeIntervene {n : ℕ} {V : Fin n → Type*} (p : NodeTables V) (a : (i : Fin n) → Option (V i)) : NodeTables V","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-8d83bc8cb1a8","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":450,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def nodeIntervene {n : ℕ} {V : Fin n → Type*} (p : NodeTables V) (a : (i : Fin n) → Option (V i)) : NodeTables V","missing":[],"search":"nodeintervene banditrlproof.causal.nodeintervene noncomputable def nodeintervene {n : ℕ} {v : fin n → type*} (p : nodetables v) (a : (i : fin n) → option (v i)) : nodetables v definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeAction","label":"encodeAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeAction","description":"def NodeCodec.encodeAction {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (a : (i : Fin n) → Option (V i)) : Fin n → Option W","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-3d774736338f","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":451,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeCodec.encodeAction {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (a : (i : Fin n) → Option (V i)) : Fin n → Option W","missing":[],"search":"encodeaction banditrlproof.causal.nodecodec.encodeaction def nodecodec.encodeaction {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (a : (i : fin n) → option (v i)) : fin n → option w definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeTables_intervene","label":"encodeTables_intervene","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeTables_intervene","description":"theorem NodeCodec.encodeTables_intervene {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) (a : (i : Fin n) → Option (V i)) : c.encodeTables (nodeIntervene p a) = intervene (c.encodeTables p) (c.encodeAction a)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-cd163f8ac9ba","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":452,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeCodec.encodeTables_intervene {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) (a : (i : Fin n) → Option (V i)) : c.encodeTables (nodeIntervene p a) = intervene (c.encodeTables p) (c.encodeAction a)","missing":[],"search":"encodetables_intervene banditrlproof.causal.nodecodec.encodetables_intervene theorem nodecodec.encodetables_intervene {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (p : nodetables v) (a : (i : fin n) → option (v i)) : c.encodetables (nodeintervene p a) = intervene (c.encodetables p) (c.encodeaction a) theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.joint_intervene","label":"joint_intervene","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.joint_intervene","description":"theorem NodeCodec.joint_intervene {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) (a : (i : Fin n) → Option (V i)) : joint (intervene (c.encodeTables p) (c.encodeAction a)) = (nodeJoint (nodeIntervene p a)).map c.encodeAssignment","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-bd35608eefeb","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":453,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:119"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeCodec.joint_intervene {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) (a : (i : Fin n) → Option (V i)) : joint (intervene (c.encodeTables p) (c.encodeAction a)) = (nodeJoint (nodeIntervene p a)).map c.encodeAssignment","missing":[],"search":"joint_intervene banditrlproof.causal.nodecodec.joint_intervene theorem nodecodec.joint_intervene {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (p : nodetables v) (a : (i : fin n) → option (v i)) : joint (intervene (c.encodetables p) (c.encodeaction a)) = (nodejoint (nodeintervene p a)).map c.encodeassignment theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel","label":"NodeGraphModel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel","description":"structure NodeGraphModel {n : ℕ} (V : Fin n → Type*) where","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-b5993f15a63a","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":454,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure NodeGraphModel {n : ℕ} (V : Fin n → Type*) where","missing":[],"search":"nodegraphmodel banditrlproof.causal.nodegraphmodel structure nodegraphmodel {n : ℕ} (v : fin n → type*) where structure compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.encodeGraph","label":"encodeGraph","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.encodeGraph","description":"noncomputable def NodeGraphModel.encodeGraph {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) : GraphModel W n where","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-5acd2b5ddf34","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":455,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def NodeGraphModel.encodeGraph {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) : GraphModel W n where","missing":[],"search":"encodegraph banditrlproof.causal.nodegraphmodel.encodegraph noncomputable def nodegraphmodel.encodegraph {n : ℕ} {v : fin n → type*} {w : type*} (g : nodegraphmodel v) (c : nodecodec v w) : graphmodel w n where definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.intervention_joint_encoded","label":"intervention_joint_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.intervention_joint_encoded","description":"theorem NodeGraphModel.intervention_joint_encoded {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (a : (i : Fin n) → Option (V i)) : joint ((g.encodeGraph c).doModel (c.encodeAction a)).table = (nodeJoint (nodeIntervene g.table a)).map c.encodeAssignment","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-dd179b7942da","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":456,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:141"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.intervention_joint_encoded {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (a : (i : Fin n) → Option (V i)) : joint ((g.encodeGraph c).doModel (c.encodeAction a)).table = (nodeJoint (nodeIntervene g.table a)).map c.encodeAssignment","missing":[],"search":"intervention_joint_encoded banditrlproof.causal.nodegraphmodel.intervention_joint_encoded theorem nodegraphmodel.intervention_joint_encoded {n : ℕ} {v : fin n → type*} {w : type*} (g : nodegraphmodel v) (c : nodecodec v w) (a : (i : fin n) → option (v i)) : joint ((g.encodegraph c).domodel (c.encodeaction a)).table = (nodejoint (nodeintervene g.table a)).map c.encodeassignment theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encoded_joint_valid","label":"encoded_joint_valid","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encoded_joint_valid","description":"theorem NodeCodec.encoded_joint_valid {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) (w : Fin n → W) (hw : w ∈ (joint (c.encodeTables p)).support) : c.encodeAssignment (c.decodeAssignment w) = w","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-b49e1bc7ecf5","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":457,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:147"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeCodec.encoded_joint_valid {n : ℕ} {V : Fin n → Type*} {W : Type*} (c : NodeCodec V W) (p : NodeTables V) (w : Fin n → W) (hw : w ∈ (joint (c.encodeTables p)).support) : c.encodeAssignment (c.decodeAssignment w) = w","missing":[],"search":"encoded_joint_valid banditrlproof.causal.nodecodec.encoded_joint_valid theorem nodecodec.encoded_joint_valid {n : ℕ} {v : fin n → type*} {w : type*} (c : nodecodec v w) (p : nodetables v) (w : fin n → w) (hw : w ∈ (joint (c.encodetables p)).support) : c.encodeassignment (c.decodeassignment w) = w theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.nodeHistory","label":"nodeHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.nodeHistory","description":"def nodeHistory {n : ℕ} {V : Fin n → Type*} (x : (i : Fin n) → V i) (i : Fin n) : NodeHistory V i","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-f2b6dc2641e6","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":458,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:155"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def nodeHistory {n : ℕ} {V : Fin n → Type*} (x : (i : Fin n) → V i) (i : Fin n) : NodeHistory V i","missing":[],"search":"nodehistory banditrlproof.causal.nodehistory def nodehistory {n : ℕ} {v : fin n → type*} (x : (i : fin n) → v i) (i : fin n) : nodehistory v i definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.ParentConfig","label":"ParentConfig","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.ParentConfig","description":"abbrev NodeGraphModel.ParentConfig {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (i : Fin n)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-fae086880e7e","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":459,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:158"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev NodeGraphModel.ParentConfig {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (i : Fin n)","missing":[],"search":"parentconfig banditrlproof.causal.nodegraphmodel.parentconfig abbrev nodegraphmodel.parentconfig {n : ℕ} {v : fin n → type*} (g : nodegraphmodel v) (i : fin n) abbreviation compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.parentConfig","label":"parentConfig","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.parentConfig","description":"def NodeGraphModel.parentConfig {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (i : Fin n) (h : NodeHistory V i) : g.ParentConfig i","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-01b584463211","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":460,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:162"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeGraphModel.parentConfig {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (i : Fin n) (h : NodeHistory V i) : g.ParentConfig i","missing":[],"search":"parentconfig banditrlproof.causal.nodegraphmodel.parentconfig def nodegraphmodel.parentconfig {n : ℕ} {v : fin n → type*} (g : nodegraphmodel v) (i : fin n) (h : nodehistory v i) : g.parentconfig i definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.encodeParent","label":"encodeParent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.encodeParent","description":"def NodeGraphModel.encodeParent {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (i : Fin n) (z : g.ParentConfig i) : (g.encodeGraph c).ParentConfig i","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-83aa4e1e1e49","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":461,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeGraphModel.encodeParent {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (i : Fin n) (z : g.ParentConfig i) : (g.encodeGraph c).ParentConfig i","missing":[],"search":"encodeparent banditrlproof.causal.nodegraphmodel.encodeparent def nodegraphmodel.encodeparent {n : ℕ} {v : fin n → type*} {w : type*} (g : nodegraphmodel v) (c : nodecodec v w) (i : fin n) (z : g.parentconfig i) : (g.encodegraph c).parentconfig i definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.decodeParent","label":"decodeParent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.decodeParent","description":"def NodeGraphModel.decodeParent {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (i : Fin n) (z : (g.encodeGraph c).ParentConfig i) : g.ParentConfig i","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-14ef7f100b7e","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":462,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:171"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeGraphModel.decodeParent {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (i : Fin n) (z : (g.encodeGraph c).ParentConfig i) : g.ParentConfig i","missing":[],"search":"decodeparent banditrlproof.causal.nodegraphmodel.decodeparent def nodegraphmodel.decodeparent {n : ℕ} {v : fin n → type*} {w : type*} (g : nodegraphmodel v) (c : nodecodec v w) (i : fin n) (z : (g.encodegraph c).parentconfig i) : g.parentconfig i definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.decode_encodeParent","label":"decode_encodeParent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.decode_encodeParent","description":"theorem NodeGraphModel.decode_encodeParent {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (i : Fin n) (z : g.ParentConfig i) : g.decodeParent c i (g.encodeParent c i z) = z","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-42610ceeac79","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":463,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:176"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.decode_encodeParent {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (i : Fin n) (z : g.ParentConfig i) : g.decodeParent c i (g.encodeParent c i z) = z","missing":[],"search":"decode_encodeparent banditrlproof.causal.nodegraphmodel.decode_encodeparent theorem nodegraphmodel.decode_encodeparent {n : ℕ} {v : fin n → type*} {w : type*} (g : nodegraphmodel v) (c : nodecodec v w) (i : fin n) (z : g.parentconfig i) : g.decodeparent c i (g.encodeparent c i z) = z theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.parentLaw","label":"parentLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.parentLaw","description":"noncomputable def NodeGraphModel.parentLaw {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (a : (i : Fin n) → Option (V i)) (i : Fin n) : PMF (g.ParentConfig i)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-67702c67c6a0","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":464,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:182"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def NodeGraphModel.parentLaw {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (a : (i : Fin n) → Option (V i)) (i : Fin n) : PMF (g.ParentConfig i)","missing":[],"search":"parentlaw banditrlproof.causal.nodegraphmodel.parentlaw noncomputable def nodegraphmodel.parentlaw {n : ℕ} {v : fin n → type*} (g : nodegraphmodel v) (a : (i : fin n) → option (v i)) (i : fin n) : pmf (g.parentconfig i) definition compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.parentLaw_encoded","label":"parentLaw_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.parentLaw_encoded","description":"theorem NodeGraphModel.parentLaw_encoded {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (a : (i : Fin n) → Option (V i)) (i : Fin n) : (g.encodeGraph c).parentLaw (c.encodeAction a) i = (g.parentLaw a i).map (g.encodeParent c i)","url":"../modules/banditrlproof-algorithms-causalheterogeneous/index.html#decl-82191af35079","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneous","order":465,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneous"],["Source","BanditRLProof/Algorithms/CausalHeterogeneous.lean:187"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.parentLaw_encoded {n : ℕ} {V : Fin n → Type*} {W : Type*} (g : NodeGraphModel V) (c : NodeCodec V W) (a : (i : Fin n) → Option (V i)) (i : Fin n) : (g.encodeGraph c).parentLaw (c.encodeAction a) i = (g.parentLaw a i).map (g.encodeParent c i)","missing":[],"search":"parentlaw_encoded banditrlproof.causal.nodegraphmodel.parentlaw_encoded theorem nodegraphmodel.parentlaw_encoded {n : ℕ} {v : fin n → type*} {w : type*} (g : nodegraphmodel v) (c : nodecodec v w) (a : (i : fin n) → option (v i)) (i : fin n) : (g.encodegraph c).parentlaw (c.encodeaction a) i = (g.parentlaw a i).map (g.encodeparent c i) theorem compiled","shard":"modules/804f24a949b91951.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.nodeJoint_snoc","label":"nodeJoint_snoc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.nodeJoint_snoc","description":"theorem nodeJoint_snoc {n : ℕ} {V : Fin (n+1) → Type*} (p : NodeTables V) (h : (i : Fin n) → V i.castSucc) (x : V (Fin.last n)) : nodeJoint p (Fin.snoc h x) = nodeJoint p.prefix h * p (Fin.last n) h x","url":"../modules/banditrlproof-algorithms-causalheterogeneouslaw/index.html#decl-a7b29993bff8","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","order":466,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousLaw"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousLaw.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem nodeJoint_snoc {n : ℕ} {V : Fin (n+1) → Type*} (p : NodeTables V) (h : (i : Fin n) → V i.castSucc) (x : V (Fin.last n)) : nodeJoint p (Fin.snoc h x) = nodeJoint p.prefix h * p (Fin.last n) h x","missing":[],"search":"nodejoint_snoc banditrlproof.causal.nodejoint_snoc theorem nodejoint_snoc {n : ℕ} {v : fin (n+1) → type*} (p : nodetables v) (h : (i : fin n) → v i.castsucc) (x : v (fin.last n)) : nodejoint p (fin.snoc h x) = nodejoint p.prefix h * p (fin.last n) h x theorem compiled","shard":"modules/b3fd8fd5daadd708.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.nodeJoint_factorization","label":"nodeJoint_factorization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.nodeJoint_factorization","description":"theorem nodeJoint_factorization {n : ℕ} {V : Fin n → Type*} (p : NodeTables V) (x : (i : Fin n) → V i) : nodeJoint p x = ∏ i : Fin n, p i (nodeHistory x i) (x i)","url":"../modules/banditrlproof-algorithms-causalheterogeneouslaw/index.html#decl-903f2ff48d59","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","order":467,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousLaw"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousLaw.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem nodeJoint_factorization {n : ℕ} {V : Fin n → Type*} (p : NodeTables V) (x : (i : Fin n) → V i) : nodeJoint p x = ∏ i : Fin n, p i (nodeHistory x i) (x i)","missing":[],"search":"nodejoint_factorization banditrlproof.causal.nodejoint_factorization theorem nodejoint_factorization {n : ℕ} {v : fin n → type*} (p : nodetables v) (x : (i : fin n) → v i) : nodejoint p x = ∏ i : fin n, p i (nodehistory x i) (x i) theorem compiled","shard":"modules/b3fd8fd5daadd708.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.nodeJoint_normalized","label":"nodeJoint_normalized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.nodeJoint_normalized","description":"theorem nodeJoint_normalized {n : ℕ} {V : Fin n → Type*} (p : NodeTables V) : ∑' x, nodeJoint p x = 1","url":"../modules/banditrlproof-algorithms-causalheterogeneouslaw/index.html#decl-ad44b9eb0c8b","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","order":468,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousLaw"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousLaw.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem nodeJoint_normalized {n : ℕ} {V : Fin n → Type*} (p : NodeTables V) : ∑' x, nodeJoint p x = 1","missing":[],"search":"nodejoint_normalized banditrlproof.causal.nodejoint_normalized theorem nodejoint_normalized {n : ℕ} {v : fin n → type*} (p : nodetables v) : ∑' x, nodejoint p x = 1 theorem compiled","shard":"modules/b3fd8fd5daadd708.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.nodeIntervention_factorization","label":"nodeIntervention_factorization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.nodeIntervention_factorization","description":"theorem nodeIntervention_factorization {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (a : (i : Fin n) → Option (V i)) (x : (i : Fin n) → V i) : nodeJoint (nodeIntervene g.table a) x = ∏ i : Fin n, (match a i with | none => g.table i (nodeHistory x i) (x i) | some v => if x i = v then 1 else 0)","url":"../modules/banditrlproof-algorithms-causalheterogeneouslaw/index.html#decl-33d9be72b90a","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","order":469,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousLaw"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousLaw.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem nodeIntervention_factorization {n : ℕ} {V : Fin n → Type*} (g : NodeGraphModel V) (a : (i : Fin n) → Option (V i)) (x : (i : Fin n) → V i) : nodeJoint (nodeIntervene g.table a) x = ∏ i : Fin n, (match a i with | none => g.table i (nodeHistory x i) (x i) | some v => if x i = v then 1 else 0)","missing":[],"search":"nodeintervention_factorization banditrlproof.causal.nodeintervention_factorization theorem nodeintervention_factorization {n : ℕ} {v : fin n → type*} (g : nodegraphmodel v) (a : (i : fin n) → option (v i)) (x : (i : fin n) → v i) : nodejoint (nodeintervene g.table a) x = ∏ i : fin n, (match a i with | none => g.table i (nodehistory x i) (x i) | some v => if x i = v then 1 else 0) theorem compiled","shard":"modules/b3fd8fd5daadd708.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.rewardMean","label":"rewardMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.rewardMean","description":"noncomputable def NodeGraphModel.rewardMean (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-ed9c48a76867","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":470,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def NodeGraphModel.rewardMean (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (a : A) : ℝ","missing":[],"search":"rewardmean banditrlproof.causal.nodegraphmodel.rewardmean noncomputable def nodegraphmodel.rewardmean (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (i : fin n) (rewardbit : v i → bool) (a : a) : ℝ definition compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.rewardMean_encoded","label":"rewardMean_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.rewardMean_encoded","description":"theorem NodeGraphModel.rewardMean_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (a : A) (hi : actions a i = none) : (g.encodeGraph c).rewardMean (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) i a = g.rewardMean actions i rewardBit a","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-e202ad6bdb88","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":471,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.rewardMean_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (a : A) (hi : actions a i = none) : (g.encodeGraph c).rewardMean (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) i a = g.rewardMean actions i rewardBit a","missing":[],"search":"rewardmean_encoded banditrlproof.causal.nodegraphmodel.rewardmean_encoded theorem nodegraphmodel.rewardmean_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (j : fin n) → option (v j)) (i : fin n) (rewardbit : v i → bool) (a : a) (hi : actions a i = none) : (g.encodegraph c).rewardmean (fun v => rewardbit (c.decode i v)) (fun a => c.encodeaction (actions a)) i a = g.rewardmean actions i rewardbit a theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleRecommendation","label":"sampleRecommendation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleRecommendation","description":"noncomputable def NodeGraphModel.sampleRecommendation (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : A","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-28473a2fb7b6","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":472,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def NodeGraphModel.sampleRecommendation (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : A","missing":[],"search":"samplerecommendation banditrlproof.causal.nodegraphmodel.samplerecommendation noncomputable def nodegraphmodel.samplerecommendation (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (b : ℝ) {t : ℕ} (w : fin t → a × ((j : fin n) → v j)) : a definition compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.simpleRegret","label":"simpleRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.simpleRegret","description":"noncomputable def NodeGraphModel.simpleRegret (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : ℝ","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-f0dfe3054cce","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":473,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def NodeGraphModel.simpleRegret (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : ℝ","missing":[],"search":"simpleregret banditrlproof.causal.nodegraphmodel.simpleregret noncomputable def nodegraphmodel.simpleregret (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (b : ℝ) {t : ℕ} (w : fin t → a × ((j : fin n) → v j)) : ℝ definition compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleRecommendation_encoded","label":"sampleRecommendation_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleRecommendation_encoded","description":"theorem NodeGraphModel.sampleRecommendation_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : (g.encodeGraph c).sampleRecommendation (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i B (c.encodeSamples w) = g.sampleRecommendation actions…","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-41709b9415aa","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":474,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.sampleRecommendation_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : (g.encodeGraph c).sampleRecommendation (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i B (c.encodeSamples w) = g.sampleRecommendation actions eta i rewardBit B w","missing":[],"search":"samplerecommendation_encoded banditrlproof.causal.nodegraphmodel.samplerecommendation_encoded theorem nodegraphmodel.samplerecommendation_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (b : ℝ) {t : ℕ} (w : fin t → a × ((j : fin n) → v j)) : (g.encodegraph c).samplerecommendation (fun v => rewardbit (c.decode i v)) (fun a => c.encodeaction (actions a)) eta i b (c.encodesamples w) = g.samplerecommendation actions eta i rewardbit b w theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.simpleRegret_encoded","label":"simpleRegret_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.simpleRegret_encoded","description":"theorem NodeGraphModel.simpleRegret_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : (g.encodeGraph c).simpleRegret (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i B (c.encodeSamples w) = g.simpleRegret a…","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-b3e255966633","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":475,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.simpleRegret_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (B : ℝ) {T : ℕ} (w : Fin T → A × ((j : Fin n) → V j)) : (g.encodeGraph c).simpleRegret (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i B (c.encodeSamples w) = g.simpleRegret actions eta i rewardBit B w","missing":[],"search":"simpleregret_encoded banditrlproof.causal.nodegraphmodel.simpleregret_encoded theorem nodegraphmodel.simpleregret_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (hi : ∀ a, actions a i = none) (b : ℝ) {t : ℕ} (w : fin t → a × ((j : fin n) → v j)) : (g.encodegraph c).simpleregret (fun v => rewardbit (c.decode i v)) (fun a => c.encodeaction (actions a)) eta i b (c.encodesamples w) = g.simpleregret actions eta i rewardbit b w theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_encoded","label":"expected_simpleRegret_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_encoded","description":"theorem NodeGraphModel.expected_simpleRegret_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (B : ℝ) (T : ℕ) : (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) = ∫ w, (g.encodeGraph c).simpleRegret (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (act…","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-5c5139d653f5","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":476,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.expected_simpleRegret_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (B : ℝ) (T : ℕ) : (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) = ∫ w, (g.encodeGraph c).simpleRegret (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i B w ∂(g.encodeGraph c).sampleLaw (fun a => c.encodeAction (actions a)) eta T","missing":[],"search":"expected_simpleregret_encoded banditrlproof.causal.nodegraphmodel.expected_simpleregret_encoded theorem nodegraphmodel.expected_simpleregret_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (hi : ∀ a, actions a i = none) (b : ℝ) (t : ℕ) : (∫ w, g.simpleregret actions eta i rewardbit b w ∂g.samplelaw actions eta t) = ∫ w, (g.encodegraph c).simpleregret (fun v => rewardbit (c.decode i v)) (fun a => c.encodeaction (actions a)) eta i b w ∂(g.encodegraph c).samplelaw (fun a => c.encodeaction (actions a)) eta t theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.covers_encoded","label":"covers_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.covers_encoded","description":"theorem NodeGraphModel.covers_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) : Covers (fun a => (g.encodeGraph c).parentLaw (c.encodeAction (actions a)) i) (mixture eta (fun a => (g.encodeGraph c).parentLaw (c.encodeAction (actions a)) i))","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-97b3a9a4c3bb","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":477,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.covers_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) : Covers (fun a => (g.encodeGraph c).parentLaw (c.encodeAction (actions a)) i) (mixture eta (fun a => (g.encodeGraph c).parentLaw (c.encodeAction (actions a)) i))","missing":[],"search":"covers_encoded banditrlproof.causal.nodegraphmodel.covers_encoded theorem nodegraphmodel.covers_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) : covers (fun a => (g.encodegraph c).parentlaw (c.encodeaction (actions a)) i) (mixture eta (fun a => (g.encodegraph c).parentlaw (c.encodeaction (actions a)) i)) theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_source_bound","label":"expected_simpleRegret_source_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_source_bound","description":"Native finite-node theorem: no codec or common-alphabet premise is required.","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-b984145bab41","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":478,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.expected_simpleRegret_source_bound (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt (m*L/T)+1/(T:ℝ)","missing":[],"search":"expected_simpleregret_source_bound banditrlproof.causal.nodegraphmodel.expected_simpleregret_source_bound native finite-node theorem: no codec or common-alphabet premise is required. theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_bounds","label":"expected_simpleRegret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_bounds","description":"theorem NodeGraphModel.expected_simpleRegret_bounds (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (B : ℝ) (T : ℕ) : 0 ≤ (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ∧ (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ≤ 1","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-4179c128f316","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":479,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.expected_simpleRegret_bounds (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (B : ℝ) (T : ℕ) : 0 ≤ (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ∧ (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ≤ 1","missing":[],"search":"expected_simpleregret_bounds banditrlproof.causal.nodegraphmodel.expected_simpleregret_bounds theorem nodegraphmodel.expected_simpleregret_bounds (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (hi : ∀ a, actions a i = none) (b : ℝ) (t : ℕ) : 0 ≤ (∫ w, g.simpleregret actions eta i rewardbit b w ∂g.samplelaw actions eta t) ∧ (∫ w, g.simpleregret actions eta i rewardbit b w ∂g.samplelaw actions eta t) ≤ 1 theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_explicit_rate","label":"expected_simpleRegret_explicit_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_explicit_rate","description":"theorem NodeGraphModel.expected_simpleRegret_explicit_rate (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fint…","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-d98edb775ce4","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":480,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:119"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.expected_simpleRegret_explicit_rate (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (T : ℕ) (hT : 0 < T) : let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ≤ (3*Real.sqrt 2+7)*Real.sqrt (m*L/T)","missing":[],"search":"expected_simpleregret_explicit_rate banditrlproof.causal.nodegraphmodel.expected_simpleregret_explicit_rate theorem nodegraphmodel.expected_simpleregret_explicit_rate (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (t : ℕ) (ht : 0 < t) : let m := designcost (fun a => g.parentlaw (actions a) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret actions eta i rewardbit b w ∂g.samplelaw actions eta t) ≤ (3*real.sqrt 2+7)*real.sqrt (m*l/t) theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_uniform","label":"expected_simpleRegret_uniform","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_uniform","description":"theorem NodeGraphModel.expected_simpleRegret_uniform (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let eta := PMF.uniformOfFintype A let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret actions eta i rewardBit…","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-2a33bd312097","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":481,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:135"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.expected_simpleRegret_uniform (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let eta := PMF.uniformOfFintype A let m := designCost (fun a => g.parentLaw (actions a) i) eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt ((Fintype.card A:ℝ)*L/T)+1/(T:ℝ)","missing":[],"search":"expected_simpleregret_uniform banditrlproof.causal.nodegraphmodel.expected_simpleregret_uniform theorem nodegraphmodel.expected_simpleregret_uniform (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (i : fin n) (rewardbit : v i → bool) (hi : ∀ a, actions a i = none) (t : ℕ) (ht : 0 < t) : let eta := pmf.uniformoffintype a let m := designcost (fun a => g.parentlaw (actions a) i) eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret actions eta i rewardbit b w ∂g.samplelaw actions eta t) ≤ (2*real.sqrt 2+7)*real.sqrt ((fintype.card a:ℝ)*l/t)+1/(t:ℝ) theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_optimal","label":"expected_simpleRegret_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_optimal","description":"theorem NodeGraphModel.expected_simpleRegret_optimal (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let p := fun a => g.parentLaw (actions a) i let eta := optimalAllocation p let m := designCost p eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret actions eta i rewa…","url":"../modules/banditrlproof-algorithms-causalheterogeneousregret/index.html#decl-2dd9ec40e44e","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","order":482,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousRegret"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousRegret.lean:154"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.expected_simpleRegret_optimal (g : NodeGraphModel V) (actions : A → (j : Fin n) → Option (V j)) (i : Fin n) (rewardBit : V i → Bool) (hi : ∀ a, actions a i = none) (T : ℕ) (hT : 0 < T) : let p := fun a => g.parentLaw (actions a) i let eta := optimalAllocation p let m := designCost p eta let L := sourceLog T (Fintype.card A) let B := sourceThreshold m T L (∫ w, g.simpleRegret actions eta i rewardBit B w ∂g.sampleLaw actions eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt (m*L/T)+1/(T:ℝ)","missing":[],"search":"expected_simpleregret_optimal banditrlproof.causal.nodegraphmodel.expected_simpleregret_optimal theorem nodegraphmodel.expected_simpleregret_optimal (g : nodegraphmodel v) (actions : a → (j : fin n) → option (v j)) (i : fin n) (rewardbit : v i → bool) (hi : ∀ a, actions a i = none) (t : ℕ) (ht : 0 < t) : let p := fun a => g.parentlaw (actions a) i let eta := optimalallocation p let m := designcost p eta let l := sourcelog t (fintype.card a) let b := sourcethreshold m t l (∫ w, g.simpleregret actions eta i rewardbit b w ∂g.samplelaw actions eta t) ≤ (2*real.sqrt 2+7)*real.sqrt (m*l/t)+1/(t:ℝ) theorem compiled","shard":"modules/d06124f51676b9e5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.roundLaw","label":"roundLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.roundLaw","description":"noncomputable def NodeGraphModel.roundLaw (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) : PMF (A × ((i : Fin n) → V i))","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-355b71adea81","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":483,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def NodeGraphModel.roundLaw (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) : PMF (A × ((i : Fin n) → V i))","missing":[],"search":"roundlaw banditrlproof.causal.nodegraphmodel.roundlaw noncomputable def nodegraphmodel.roundlaw (g : nodegraphmodel v) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) : pmf (a × ((i : fin n) → v i)) definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleLaw","label":"sampleLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleLaw","description":"noncomputable def NodeGraphModel.sampleLaw (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (T : ℕ) : Measure (Fin T → A × ((i : Fin n) → V i))","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-2c41ed694653","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":484,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def NodeGraphModel.sampleLaw (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (T : ℕ) : Measure (Fin T → A × ((i : Fin n) → V i))","missing":[],"search":"samplelaw banditrlproof.causal.nodegraphmodel.samplelaw noncomputable def nodegraphmodel.samplelaw (g : nodegraphmodel v) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (t : ℕ) : measure (fin t → a × ((i : fin n) → v i)) definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeRound","label":"encodeRound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeRound","description":"def NodeCodec.encodeRound (c : NodeCodec V W) (ax : A × ((i : Fin n) → V i)) : A × (Fin n → W)","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-7b8c19d0b4c4","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":485,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeCodec.encodeRound (c : NodeCodec V W) (ax : A × ((i : Fin n) → V i)) : A × (Fin n → W)","missing":[],"search":"encoderound banditrlproof.causal.nodecodec.encoderound def nodecodec.encoderound (c : nodecodec v w) (ax : a × ((i : fin n) → v i)) : a × (fin n → w) definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeCodec.encodeSamples","label":"encodeSamples","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeCodec.encodeSamples","description":"def NodeCodec.encodeSamples (c : NodeCodec V W) {T : ℕ} (w : Fin T → A × ((i : Fin n) → V i)) : Fin T → A × (Fin n → W)","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-307f595ef8bf","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":486,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeCodec.encodeSamples (c : NodeCodec V W) {T : ℕ} (w : Fin T → A × ((i : Fin n) → V i)) : Fin T → A × (Fin n → W)","missing":[],"search":"encodesamples banditrlproof.causal.nodecodec.encodesamples def nodecodec.encodesamples (c : nodecodec v w) {t : ℕ} (w : fin t → a × ((i : fin n) → v i)) : fin t → a × (fin n → w) definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.roundLaw_encoded","label":"roundLaw_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.roundLaw_encoded","description":"theorem NodeGraphModel.roundLaw_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) : (g.roundLaw actions eta).map c.encodeRound = (g.encodeGraph c).roundLaw (fun a => c.encodeAction (actions a)) eta","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-b9e0993519c8","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":487,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.roundLaw_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) : (g.roundLaw actions eta).map c.encodeRound = (g.encodeGraph c).roundLaw (fun a => c.encodeAction (actions a)) eta","missing":[],"search":"roundlaw_encoded banditrlproof.causal.nodegraphmodel.roundlaw_encoded theorem nodegraphmodel.roundlaw_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) : (g.roundlaw actions eta).map c.encoderound = (g.encodegraph c).roundlaw (fun a => c.encodeaction (actions a)) eta theorem compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleLaw_encoded","label":"sampleLaw_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleLaw_encoded","description":"theorem NodeGraphModel.sampleLaw_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (T : ℕ) : (g.sampleLaw actions eta T).map c.encodeSamples = (g.encodeGraph c).sampleLaw (fun a => c.encodeAction (actions a)) eta T","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-441b13c7dad2","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":488,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.sampleLaw_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (T : ℕ) : (g.sampleLaw actions eta T).map c.encodeSamples = (g.encodeGraph c).sampleLaw (fun a => c.encodeAction (actions a)) eta T","missing":[],"search":"samplelaw_encoded banditrlproof.causal.nodegraphmodel.samplelaw_encoded theorem nodegraphmodel.samplelaw_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (t : ℕ) : (g.samplelaw actions eta t).map c.encodesamples = (g.encodegraph c).samplelaw (fun a => c.encodeaction (actions a)) eta t theorem compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.observation","label":"observation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.observation","description":"def NodeGraphModel.observation (g : NodeGraphModel V) (i : Fin n) (rewardBit : V i → Bool) (ax : A × ((i : Fin n) → V i)) : g.ParentConfig i × Bool","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-1d30080e74bb","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":489,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def NodeGraphModel.observation (g : NodeGraphModel V) (i : Fin n) (rewardBit : V i → Bool) (ax : A × ((i : Fin n) → V i)) : g.ParentConfig i × Bool","missing":[],"search":"observation banditrlproof.causal.nodegraphmodel.observation def nodegraphmodel.observation (g : nodegraphmodel v) (i : fin n) (rewardbit : v i → bool) (ax : a × ((i : fin n) → v i)) : g.parentconfig i × bool definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleWeightedBit","label":"sampleWeightedBit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleWeightedBit","description":"noncomputable def NodeGraphModel.sampleWeightedBit (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (t : Fin T) (w : Fin T → A × ((i : Fin n) → V i)) : ℝ","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-e28ec10ddddf","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":490,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def NodeGraphModel.sampleWeightedBit (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (t : Fin T) (w : Fin T → A × ((i : Fin n) → V i)) : ℝ","missing":[],"search":"sampleweightedbit banditrlproof.causal.nodegraphmodel.sampleweightedbit noncomputable def nodegraphmodel.sampleweightedbit (g : nodegraphmodel v) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (a : a) (b : ℝ) {t : ℕ} (t : fin t) (w : fin t → a × ((i : fin n) → v i)) : ℝ definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleEstimate","label":"sampleEstimate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleEstimate","description":"noncomputable def NodeGraphModel.sampleEstimate (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (w : Fin T → A × ((i : Fin n) → V i)) : ℝ","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-1622765ca9b3","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":491,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def NodeGraphModel.sampleEstimate (g : NodeGraphModel V) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (w : Fin T → A × ((i : Fin n) → V i)) : ℝ","missing":[],"search":"sampleestimate banditrlproof.causal.nodegraphmodel.sampleestimate noncomputable def nodegraphmodel.sampleestimate (g : nodegraphmodel v) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (a : a) (b : ℝ) {t : ℕ} (w : fin t → a × ((i : fin n) → v i)) : ℝ definition compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleWeightedBit_encoded","label":"sampleWeightedBit_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleWeightedBit_encoded","description":"theorem NodeGraphModel.sampleWeightedBit_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (t : Fin T) (w : Fin T → A × ((i : Fin n) → V i)) : (g.encodeGraph c).sampleWeightedBit (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i a B t (c.encodeSamples w) = g.sampleWeigh…","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-26e2054bd8ff","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":492,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.sampleWeightedBit_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (t : Fin T) (w : Fin T → A × ((i : Fin n) → V i)) : (g.encodeGraph c).sampleWeightedBit (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i a B t (c.encodeSamples w) = g.sampleWeightedBit actions eta i rewardBit a B t w","missing":[],"search":"sampleweightedbit_encoded banditrlproof.causal.nodegraphmodel.sampleweightedbit_encoded theorem nodegraphmodel.sampleweightedbit_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (a : a) (b : ℝ) {t : ℕ} (t : fin t) (w : fin t → a × ((i : fin n) → v i)) : (g.encodegraph c).sampleweightedbit (fun v => rewardbit (c.decode i v)) (fun a => c.encodeaction (actions a)) eta i a b t (c.encodesamples w) = g.sampleweightedbit actions eta i rewardbit a b t w theorem compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleEstimate_encoded","label":"sampleEstimate_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.sampleEstimate_encoded","description":"theorem NodeGraphModel.sampleEstimate_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (w : Fin T → A × ((i : Fin n) → V i)) : (g.encodeGraph c).sampleEstimate (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i a B (c.encodeSamples w) = g.sampleEstimate actions eta i re…","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-4eacd9822f6d","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":493,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem NodeGraphModel.sampleEstimate_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) (rewardBit : V i → Bool) (a : A) (B : ℝ) {T : ℕ} (w : Fin T → A × ((i : Fin n) → V i)) : (g.encodeGraph c).sampleEstimate (fun v => rewardBit (c.decode i v)) (fun a => c.encodeAction (actions a)) eta i a B (c.encodeSamples w) = g.sampleEstimate actions eta i rewardBit a B w","missing":[],"search":"sampleestimate_encoded banditrlproof.causal.nodegraphmodel.sampleestimate_encoded theorem nodegraphmodel.sampleestimate_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (i : fin n) (rewardbit : v i → bool) (a : a) (b : ℝ) {t : ℕ} (w : fin t → a × ((i : fin n) → v i)) : (g.encodegraph c).sampleestimate (fun v => rewardbit (c.decode i v)) (fun a => c.encodeaction (actions a)) eta i a b (c.encodesamples w) = g.sampleestimate actions eta i rewardbit a b w theorem compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.NodeGraphModel.designCost_encoded","label":"designCost_encoded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.NodeGraphModel.designCost_encoded","description":"theorem NodeGraphModel.designCost_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) : designCost (fun a => (g.encodeGraph c).parentLaw (c.encodeAction (actions a)) i) eta = designCost (fun a => g.parentLaw (actions a) i) eta","url":"../modules/banditrlproof-algorithms-causalheterogeneoussampling/index.html#decl-f3d69317c698","parent":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","order":494,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalHeterogeneousSampling"],["Source","BanditRLProof/Algorithms/CausalHeterogeneousSampling.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem NodeGraphModel.designCost_encoded (g : NodeGraphModel V) (c : NodeCodec V W) (actions : A → (i : Fin n) → Option (V i)) (eta : PMF A) (i : Fin n) : designCost (fun a => (g.encodeGraph c).parentLaw (c.encodeAction (actions a)) i) eta = designCost (fun a => g.parentLaw (actions a) i) eta","missing":[],"search":"designcost_encoded banditrlproof.causal.nodegraphmodel.designcost_encoded theorem nodegraphmodel.designcost_encoded (g : nodegraphmodel v) (c : nodecodec v w) (actions : a → (i : fin n) → option (v i)) (eta : pmf a) (i : fin n) : designcost (fun a => (g.encodegraph c).parentlaw (c.encodeaction (actions a)) i) eta = designcost (fun a => g.parentlaw (actions a) i) eta theorem compiled","shard":"modules/21256850cfc44718.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.mass","label":"mass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.mass","description":"noncomputable def mass {Z : Type*} (p : PMF Z) (z : Z) : ℝ","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-f36b3634d749","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":495,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def mass {Z : Type*} (p : PMF Z) (z : Z) : ℝ","missing":[],"search":"mass banditrlproof.causal.mass noncomputable def mass {z : type*} (p : pmf z) (z : z) : ℝ definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mass_nonneg","label":"mass_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mass_nonneg","description":"theorem mass_nonneg {Z : Type*} (p : PMF Z) (z : Z) : 0 ≤ mass p z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-b099d1ed16c9","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":496,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mass_nonneg {Z : Type*} (p : PMF Z) (z : Z) : 0 ≤ mass p z","missing":[],"search":"mass_nonneg banditrlproof.causal.mass_nonneg theorem mass_nonneg {z : type*} (p : pmf z) (z : z) : 0 ≤ mass p z theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mass_le_one","label":"mass_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mass_le_one","description":"theorem mass_le_one {Z : Type*} (p : PMF Z) (z : Z) : mass p z ≤ 1","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-d270c087c30c","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":497,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mass_le_one {Z : Type*} (p : PMF Z) (z : Z) : mass p z ≤ 1","missing":[],"search":"mass_le_one banditrlproof.causal.mass_le_one theorem mass_le_one {z : type*} (p : pmf z) (z : z) : mass p z ≤ 1 theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sum_mass","label":"sum_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sum_mass","description":"theorem sum_mass {Z : Type*} [Fintype Z] (p : PMF Z) : ∑ z, mass p z = 1","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-48d728f8afc4","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":498,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_mass {Z : Type*} [Fintype Z] (p : PMF Z) : ∑ z, mass p z = 1","missing":[],"search":"sum_mass banditrlproof.causal.sum_mass theorem sum_mass {z : type*} [fintype z] (p : pmf z) : ∑ z, mass p z = 1 theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixture","label":"mixture","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.mixture","description":"noncomputable def mixture {A Z : Type*} (eta : PMF A) (p : A → PMF Z) : PMF Z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-d657608753dc","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":499,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def mixture {A Z : Type*} (eta : PMF A) (p : A → PMF Z) : PMF Z","missing":[],"search":"mixture banditrlproof.causal.mixture noncomputable def mixture {a z : type*} (eta : pmf a) (p : a → pmf z) : pmf z definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixture_mass","label":"mixture_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mixture_mass","description":"theorem mixture_mass {A Z : Type*} [Fintype A] (eta : PMF A) (p : A → PMF Z) (z : Z) : mass (mixture eta p) z = ∑ a, mass eta a * mass (p a) z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-8a567e4047d1","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":500,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mixture_mass {A Z : Type*} [Fintype A] (eta : PMF A) (p : A → PMF Z) (z : Z) : mass (mixture eta p) z = ∑ a, mass eta a * mass (p a) z","missing":[],"search":"mixture_mass banditrlproof.causal.mixture_mass theorem mixture_mass {a z : type*} [fintype a] (eta : pmf a) (p : a → pmf z) (z : z) : mass (mixture eta p) z = ∑ a, mass eta a * mass (p a) z theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.Covers","label":"Covers","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.Covers","description":"def Covers {A Z : Type*} (p : A → PMF Z) (q : PMF Z) : Prop","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-95e5cf1d26a4","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":501,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def Covers {A Z : Type*} (p : A → PMF Z) (q : PMF Z) : Prop","missing":[],"search":"covers banditrlproof.causal.covers def covers {a z : type*} (p : a → pmf z) (q : pmf z) : prop definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ratio","label":"ratio","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ratio","description":"noncomputable def ratio {Z : Type*} (p q : PMF Z) (z : Z) : ℝ","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-9b206c497009","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":502,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def ratio {Z : Type*} (p q : PMF Z) (z : Z) : ℝ","missing":[],"search":"ratio banditrlproof.causal.ratio noncomputable def ratio {z : type*} (p q : pmf z) (z : z) : ℝ definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.covered_cancel","label":"covered_cancel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.covered_cancel","description":"theorem covered_cancel {A Z : Type*} (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (z : Z) : mass q z * ratio (p a) q z = mass (p a) z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-123e73c0d9d6","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":503,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem covered_cancel {A Z : Type*} (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (z : Z) : mass q z * ratio (p a) q z = mass (p a) z","missing":[],"search":"covered_cancel banditrlproof.causal.covered_cancel theorem covered_cancel {a z : type*} (p : a → pmf z) (q : pmf z) (hc : covers p q) (a : a) (z : z) : mass q z * ratio (p a) q z = mass (p a) z theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.importance_identity","label":"importance_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.importance_identity","description":"theorem importance_identity {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (f : Z → ℝ) : ∑ z, mass q z * (ratio (p a) q z * f z) = ∑ z, mass (p a) z * f z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-dddd97cdee33","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":504,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem importance_identity {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (f : Z → ℝ) : ∑ z, mass q z * (ratio (p a) q z * f z) = ∑ z, mass (p a) z * f z","missing":[],"search":"importance_identity banditrlproof.causal.importance_identity theorem importance_identity {a z : type*} [fintype z] (p : a → pmf z) (q : pmf z) (hc : covers p q) (a : a) (f : z → ℝ) : ∑ z, mass q z * (ratio (p a) q z * f z) = ∑ z, mass (p a) z * f z theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.positive_allocation_covers","label":"positive_allocation_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.positive_allocation_covers","description":"theorem positive_allocation_covers {A Z : Type*} [Fintype A] (eta : PMF A) (p : A → PMF Z) (he : ∀ a, 0 < mass eta a) : Covers p (mixture eta p)","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-d93b2e0deb28","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":505,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem positive_allocation_covers {A Z : Type*} [Fintype A] (eta : PMF A) (p : A → PMF Z) (he : ∀ a, 0 < mass eta a) : Covers p (mixture eta p)","missing":[],"search":"positive_allocation_covers banditrlproof.causal.positive_allocation_covers theorem positive_allocation_covers {a z : type*} [fintype a] (eta : pmf a) (p : a → pmf z) (he : ∀ a, 0 < mass eta a) : covers p (mixture eta p) theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ratio_nonneg","label":"ratio_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ratio_nonneg","description":"theorem ratio_nonneg {Z : Type*} (p q : PMF Z) (z : Z) : 0 ≤ ratio p q z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-aa2f3236e5ee","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":506,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ratio_nonneg {Z : Type*} (p q : PMF Z) (z : Z) : 0 ≤ ratio p q z","missing":[],"search":"ratio_nonneg banditrlproof.causal.ratio_nonneg theorem ratio_nonneg {z : type*} (p q : pmf z) (z : z) : 0 ≤ ratio p q z theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment","label":"secondMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment","description":"noncomputable def secondMoment {Z : Type*} [Fintype Z] (p q : PMF Z) : ℝ","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-c5fbd8b63792","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":507,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def secondMoment {Z : Type*} [Fintype Z] (p q : PMF Z) : ℝ","missing":[],"search":"secondmoment banditrlproof.causal.secondmoment noncomputable def secondmoment {z : type*} [fintype z] (p q : pmf z) : ℝ definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ratio_second_moment","label":"ratio_second_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ratio_second_moment","description":"theorem ratio_second_moment {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) : ∑ z, mass q z * ratio (p a) q z ^ 2 = secondMoment (p a) q","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-ee20f3c4aedd","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":508,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ratio_second_moment {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) : ∑ z, mass q z * ratio (p a) q z ^ 2 = secondMoment (p a) q","missing":[],"search":"ratio_second_moment banditrlproof.causal.ratio_second_moment theorem ratio_second_moment {a z : type*} [fintype z] (p : a → pmf z) (q : pmf z) (hc : covers p q) (a : a) : ∑ z, mass q z * ratio (p a) q z ^ 2 = secondmoment (p a) q theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.truncationBias","label":"truncationBias","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.truncationBias","description":"noncomputable def truncationBias {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (B : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-009cded78361","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":509,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def truncationBias {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (B : ℝ) : ℝ","missing":[],"search":"truncationbias banditrlproof.causal.truncationbias noncomputable def truncationbias {z : type*} [fintype z] (p q : pmf z) (r : z → ℝ) (b : ℝ) : ℝ definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.truncatedMean","label":"truncatedMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.truncatedMean","description":"noncomputable def truncatedMean {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (B : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-542b8abd373a","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":510,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def truncatedMean {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (B : ℝ) : ℝ","missing":[],"search":"truncatedmean banditrlproof.causal.truncatedmean noncomputable def truncatedmean {z : type*} [fintype z] (p q : pmf z) (r : z → ℝ) (b : ℝ) : ℝ definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.truncatedMean_add_bias","label":"truncatedMean_add_bias","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.truncatedMean_add_bias","description":"theorem truncatedMean_add_bias {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (r : Z → ℝ) (B : ℝ) : truncatedMean (p a) q r B + truncationBias (p a) q r B = ∑ z, mass (p a) z * r z","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-1d30c7f714a4","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":511,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncatedMean_add_bias {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (r : Z → ℝ) (B : ℝ) : truncatedMean (p a) q r B + truncationBias (p a) q r B = ∑ z, mass (p a) z * r z","missing":[],"search":"truncatedmean_add_bias banditrlproof.causal.truncatedmean_add_bias theorem truncatedmean_add_bias {a z : type*} [fintype z] (p : a → pmf z) (q : pmf z) (hc : covers p q) (a : a) (r : z → ℝ) (b : ℝ) : truncatedmean (p a) q r b + truncationbias (p a) q r b = ∑ z, mass (p a) z * r z theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.truncationBias_nonneg","label":"truncationBias_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.truncationBias_nonneg","description":"theorem truncationBias_nonneg {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (hr : ∀ z, 0 ≤ r z) (B : ℝ) : 0 ≤ truncationBias p q r B","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-8bef899d514a","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":512,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncationBias_nonneg {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (hr : ∀ z, 0 ≤ r z) (B : ℝ) : 0 ≤ truncationBias p q r B","missing":[],"search":"truncationbias_nonneg banditrlproof.causal.truncationbias_nonneg theorem truncationbias_nonneg {z : type*} [fintype z] (p q : pmf z) (r : z → ℝ) (hr : ∀ z, 0 ≤ r z) (b : ℝ) : 0 ≤ truncationbias p q r b theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.truncationBias_le","label":"truncationBias_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.truncationBias_le","description":"theorem truncationBias_le {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (hr : ∀ z, r z ≤ 1) (B : ℝ) (hB : 0 < B) : truncationBias p q r B ≤ secondMoment p q / B","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-525c696acde0","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":513,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:106"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncationBias_le {Z : Type*} [Fintype Z] (p q : PMF Z) (r : Z → ℝ) (hr : ∀ z, r z ≤ 1) (B : ℝ) (hB : 0 < B) : truncationBias p q r B ≤ secondMoment p q / B","missing":[],"search":"truncationbias_le banditrlproof.causal.truncationbias_le theorem truncationbias_le {z : type*} [fintype z] (p q : pmf z) (r : z → ℝ) (hr : ∀ z, r z ≤ 1) (b : ℝ) (hb : 0 < b) : truncationbias p q r b ≤ secondmoment p q / b theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.pairedLaw","label":"pairedLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.pairedLaw","description":"noncomputable def pairedLaw {Z V : Type*} (p : PMF Z) (k : Z → PMF V) : PMF (Z × V)","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-34e0fc005fbf","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":514,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def pairedLaw {Z V : Type*} (p : PMF Z) (k : Z → PMF V) : PMF (Z × V)","missing":[],"search":"pairedlaw banditrlproof.causal.pairedlaw noncomputable def pairedlaw {z v : type*} (p : pmf z) (k : z → pmf v) : pmf (z × v) definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixture_pairedLaw","label":"mixture_pairedLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mixture_pairedLaw","description":"theorem mixture_pairedLaw {A Z V : Type*} (eta : PMF A) (p : A → PMF Z) (k : Z → PMF V) : mixture eta (fun a => pairedLaw (p a) k) = pairedLaw (mixture eta p) k","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-abb7eada21d1","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":515,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mixture_pairedLaw {A Z V : Type*} (eta : PMF A) (p : A → PMF Z) (k : Z → PMF V) : mixture eta (fun a => pairedLaw (p a) k) = pairedLaw (mixture eta p) k","missing":[],"search":"mixture_pairedlaw banditrlproof.causal.mixture_pairedlaw theorem mixture_pairedlaw {a z v : type*} (eta : pmf a) (p : a → pmf z) (k : z → pmf v) : mixture eta (fun a => pairedlaw (p a) k) = pairedlaw (mixture eta p) k theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.pairedLaw_mass","label":"pairedLaw_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.pairedLaw_mass","description":"theorem pairedLaw_mass {Z V : Type*} (p : PMF Z) (k : Z → PMF V) (z : Z) (y : V) : mass (pairedLaw p k) (z,y) = mass p z * mass (k z) y","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-58543308b56f","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":516,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pairedLaw_mass {Z V : Type*} (p : PMF Z) (k : Z → PMF V) (z : Z) (y : V) : mass (pairedLaw p k) (z,y) = mass p z * mass (k z) y","missing":[],"search":"pairedlaw_mass banditrlproof.causal.pairedlaw_mass theorem pairedlaw_mass {z v : type*} (p : pmf z) (k : z → pmf v) (z : z) (y : v) : mass (pairedlaw p k) (z,y) = mass p z * mass (k z) y theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.weightedBit","label":"weightedBit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.weightedBit","description":"noncomputable def weightedBit {Z : Type*} (p q : PMF Z) (B : ℝ) (zy : Z × Bool) : ℝ","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-0d99d93d1e39","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":517,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def weightedBit {Z : Type*} (p q : PMF Z) (B : ℝ) (zy : Z × Bool) : ℝ","missing":[],"search":"weightedbit banditrlproof.causal.weightedbit noncomputable def weightedbit {z : type*} (p q : pmf z) (b : ℝ) (zy : z × bool) : ℝ definition compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.weightedBit_mean","label":"weightedBit_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.weightedBit_mean","description":"theorem weightedBit_mean {Z : Type*} [Fintype Z] (p q : PMF Z) (k : Z → PMF Bool) (B : ℝ) : (∑ z, ∑ y : Bool, mass (pairedLaw q k) (z,y) * weightedBit p q B (z,y)) = truncatedMean p q (fun z => mass (k z) true) B","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-fe956d51b251","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":518,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedBit_mean {Z : Type*} [Fintype Z] (p q : PMF Z) (k : Z → PMF Bool) (B : ℝ) : (∑ z, ∑ y : Bool, mass (pairedLaw q k) (z,y) * weightedBit p q B (z,y)) = truncatedMean p q (fun z => mass (k z) true) B","missing":[],"search":"weightedbit_mean banditrlproof.causal.weightedbit_mean theorem weightedbit_mean {z : type*} [fintype z] (p q : pmf z) (k : z → pmf bool) (b : ℝ) : (∑ z, ∑ y : bool, mass (pairedlaw q k) (z,y) * weightedbit p q b (z,y)) = truncatedmean p q (fun z => mass (k z) true) b theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.weightedBit_bounds","label":"weightedBit_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.weightedBit_bounds","description":"theorem weightedBit_bounds {Z : Type*} (p q : PMF Z) (B : ℝ) (hB : 0 ≤ B) (zy : Z × Bool) : 0 ≤ weightedBit p q B zy ∧ weightedBit p q B zy ≤ B","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-f57f6511b244","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":519,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:149"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedBit_bounds {Z : Type*} (p q : PMF Z) (B : ℝ) (hB : 0 ≤ B) (zy : Z × Bool) : 0 ≤ weightedBit p q B zy ∧ weightedBit p q B zy ≤ B","missing":[],"search":"weightedbit_bounds banditrlproof.causal.weightedbit_bounds theorem weightedbit_bounds {z : type*} (p q : pmf z) (b : ℝ) (hb : 0 ≤ b) (zy : z × bool) : 0 ≤ weightedbit p q b zy ∧ weightedbit p q b zy ≤ b theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.weightedBit_second_le","label":"weightedBit_second_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.weightedBit_second_le","description":"theorem weightedBit_second_le {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (k : Z → PMF Bool) (B : ℝ) : (∑ z, ∑ y : Bool, mass (pairedLaw q k) (z,y) * weightedBit (p a) q B (z,y)^2) ≤ secondMoment (p a) q","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-723fc798e718","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":520,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedBit_second_le {A Z : Type*} [Fintype Z] (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (a : A) (k : Z → PMF Bool) (B : ℝ) : (∑ z, ∑ y : Bool, mass (pairedLaw q k) (z,y) * weightedBit (p a) q B (z,y)^2) ≤ secondMoment (p a) q","missing":[],"search":"weightedbit_second_le banditrlproof.causal.weightedbit_second_le theorem weightedbit_second_le {a z : type*} [fintype z] (p : a → pmf z) (q : pmf z) (hc : covers p q) (a : a) (k : z → pmf bool) (b : ℝ) : (∑ z, ∑ y : bool, mass (pairedlaw q k) (z,y) * weightedbit (p a) q b (z,y)^2) ≤ secondmoment (p a) q theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.mixture_parent_joint","label":"mixture_parent_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.mixture_parent_joint","description":"theorem GraphModel.mixture_parent_joint {A V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) : mixture eta (fun a => (joint (g.doModel (actions a)).table).map (fun x => (g.parentConfig i (history x i),x i))) = pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (g.parentTable i)","url":"../modules/banditrlproof-algorithms-causalimportance/index.html#decl-152b79f470d5","parent":"module:BanditRLProof.Algorithms.CausalImportance","order":521,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportance"],["Source","BanditRLProof/Algorithms/CausalImportance.lean:176"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.mixture_parent_joint {A V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) : mixture eta (fun a => (joint (g.doModel (actions a)).table).map (fun x => (g.parentConfig i (history x i),x i))) = pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (g.parentTable i)","missing":[],"search":"mixture_parent_joint banditrlproof.causal.graphmodel.mixture_parent_joint theorem graphmodel.mixture_parent_joint {a v : type*} [inhabited v] {n : ℕ} (g : graphmodel v n) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) : mixture eta (fun a => (joint (g.domodel (actions a)).table).map (fun x => (g.parentconfig i (history x i),x i))) = pairedlaw (mixture eta (fun a => g.parentlaw (actions a) i)) (g.parenttable i) theorem compiled","shard":"modules/9b0bfe2a017b5287.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mass_map_injective","label":"mass_map_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mass_map_injective","description":"theorem mass_map_injective (p : PMF Z) (e : Z → U) (he : Function.Injective e) (z : Z) : mass (p.map e) (e z) = mass p z","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-b56603f36cd1","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":522,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mass_map_injective (p : PMF Z) (e : Z → U) (he : Function.Injective e) (z : Z) : mass (p.map e) (e z) = mass p z","missing":[],"search":"mass_map_injective banditrlproof.causal.mass_map_injective theorem mass_map_injective (p : pmf z) (e : z → u) (he : function.injective e) (z : z) : mass (p.map e) (e z) = mass p z theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ratio_map_injective","label":"ratio_map_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ratio_map_injective","description":"theorem ratio_map_injective (p q : PMF Z) (e : Z → U) (he : Function.Injective e) (z : Z) : ratio (p.map e) (q.map e) (e z) = ratio p q z","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-f47858469906","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":523,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ratio_map_injective (p q : PMF Z) (e : Z → U) (he : Function.Injective e) (z : Z) : ratio (p.map e) (q.map e) (e z) = ratio p q z","missing":[],"search":"ratio_map_injective banditrlproof.causal.ratio_map_injective theorem ratio_map_injective (p q : pmf z) (e : z → u) (he : function.injective e) (z : z) : ratio (p.map e) (q.map e) (e z) = ratio p q z theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixture_map","label":"mixture_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mixture_map","description":"theorem mixture_map (eta : PMF A) (p : A → PMF Z) (e : Z → U) : mixture eta (fun a => (p a).map e) = (mixture eta p).map e","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-7b1ae810c175","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":524,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mixture_map (eta : PMF A) (p : A → PMF Z) (e : Z → U) : mixture eta (fun a => (p a).map e) = (mixture eta p).map e","missing":[],"search":"mixture_map banditrlproof.causal.mixture_map theorem mixture_map (eta : pmf a) (p : a → pmf z) (e : z → u) : mixture eta (fun a => (p a).map e) = (mixture eta p).map e theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.covers_map_injective","label":"covers_map_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.covers_map_injective","description":"theorem covers_map_injective (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (e : Z → U) (he : Function.Injective e) : Covers (fun a => (p a).map e) (q.map e)","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-13ff59ae3db3","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":525,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem covers_map_injective (p : A → PMF Z) (q : PMF Z) (hc : Covers p q) (e : Z → U) (he : Function.Injective e) : Covers (fun a => (p a).map e) (q.map e)","missing":[],"search":"covers_map_injective banditrlproof.causal.covers_map_injective theorem covers_map_injective (p : a → pmf z) (q : pmf z) (hc : covers p q) (e : z → u) (he : function.injective e) : covers (fun a => (p a).map e) (q.map e) theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.weightedBit_map_injective","label":"weightedBit_map_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.weightedBit_map_injective","description":"theorem weightedBit_map_injective (p q : PMF Z) (e : Z → U) (he : Function.Injective e) (B : ℝ) (z : Z) (y : Bool) : weightedBit (p.map e) (q.map e) B (e z,y) = weightedBit p q B (z,y)","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-5193479d599d","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":526,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedBit_map_injective (p q : PMF Z) (e : Z → U) (he : Function.Injective e) (B : ℝ) (z : Z) (y : Bool) : weightedBit (p.map e) (q.map e) B (e z,y) = weightedBit p q B (z,y)","missing":[],"search":"weightedbit_map_injective banditrlproof.causal.weightedbit_map_injective theorem weightedbit_map_injective (p q : pmf z) (e : z → u) (he : function.injective e) (b : ℝ) (z : z) (y : bool) : weightedbit (p.map e) (q.map e) b (e z,y) = weightedbit p q b (z,y) theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment_map_injective","label":"secondMoment_map_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment_map_injective","description":"theorem secondMoment_map_injective (p q : PMF Z) (e : Z → U) (he : Function.Injective e) : secondMoment (p.map e) (q.map e) = secondMoment p q","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-5a1df40b2538","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":527,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem secondMoment_map_injective (p q : PMF Z) (e : Z → U) (he : Function.Injective e) : secondMoment (p.map e) (q.map e) = secondMoment p q","missing":[],"search":"secondmoment_map_injective banditrlproof.causal.secondmoment_map_injective theorem secondmoment_map_injective (p q : pmf z) (e : z → u) (he : function.injective e) : secondmoment (p.map e) (q.map e) = secondmoment p q theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.designCost_map_injective","label":"designCost_map_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.designCost_map_injective","description":"theorem designCost_map_injective (p : A → PMF Z) (eta : PMF A) (e : Z → U) (he : Function.Injective e) : designCost (fun a => (p a).map e) eta = designCost p eta","url":"../modules/banditrlproof-algorithms-causalimportancetransport/index.html#decl-7fd1d01b3a50","parent":"module:BanditRLProof.Algorithms.CausalImportanceTransport","order":528,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalImportanceTransport"],["Source","BanditRLProof/Algorithms/CausalImportanceTransport.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem designCost_map_injective (p : A → PMF Z) (eta : PMF A) (e : Z → U) (he : Function.Injective e) : designCost (fun a => (p a).map e) eta = designCost p eta","missing":[],"search":"designcost_map_injective banditrlproof.causal.designcost_map_injective theorem designcost_map_injective (p : a → pmf z) (eta : pmf a) (e : z → u) (he : function.injective e) : designcost (fun a => (p a).map e) eta = designcost p eta theorem compiled","shard":"modules/452583592f22cd6a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_map_init","label":"joint_map_init","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_map_init","description":"theorem joint_map_init {V : Type*} {n : ℕ} (p : Tables V (n+1)) : (joint p).map Fin.init = joint p.prefix","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-f7d3c549cc2c","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":529,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_map_init {V : Type*} {n : ℕ} (p : Tables V (n+1)) : (joint p).map Fin.init = joint p.prefix","missing":[],"search":"joint_map_init banditrlproof.causal.joint_map_init theorem joint_map_init {v : type*} {n : ℕ} (p : tables v (n+1)) : (joint p).map fin.init = joint p.prefix theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.take","label":"take","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.take","description":"def take {V : Type*} {m n : ℕ} (hm : m ≤ n) (x : Fin n → V) : Fin m → V","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-37b814220584","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":530,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def take {V : Type*} {m n : ℕ} (hm : m ≤ n) (x : Fin n → V) : Fin m → V","missing":[],"search":"take banditrlproof.causal.take def take {v : type*} {m n : ℕ} (hm : m ≤ n) (x : fin n → v) : fin m → v definition compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.Tables.restrict","label":"restrict","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.Tables.restrict","description":"def Tables.restrict {V : Type*} {m n : ℕ} (p : Tables V n) (hm : m ≤ n) : Tables V m","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-6ed5219a4fb3","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":531,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def Tables.restrict {V : Type*} {m n : ℕ} (p : Tables V n) (hm : m ≤ n) : Tables V m","missing":[],"search":"restrict banditrlproof.causal.tables.restrict def tables.restrict {v : type*} {m n : ℕ} (p : tables v n) (hm : m ≤ n) : tables v m definition compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_map_take","label":"joint_map_take","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_map_take","description":"theorem joint_map_take {V : Type*} {m n : ℕ} (p : Tables V n) (hm : m ≤ n) : (joint p).map (take hm) = joint (p.restrict hm)","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-5a63297d8152","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":532,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_map_take {V : Type*} {m n : ℕ} (p : Tables V n) (hm : m ≤ n) : (joint p).map (take hm) = joint (p.restrict hm)","missing":[],"search":"joint_map_take banditrlproof.causal.joint_map_take theorem joint_map_take {v : type*} {m n : ℕ} (p : tables v n) (hm : m ≤ n) : (joint p).map (take hm) = joint (p.restrict hm) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.ParentConfig","label":"ParentConfig","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.ParentConfig","description":"abbrev GraphModel.ParentConfig {V : Type*} {n : ℕ} (g : GraphModel V n) (i : Fin n)","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-756f2e1fabf2","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":533,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev GraphModel.ParentConfig {V : Type*} {n : ℕ} (g : GraphModel V n) (i : Fin n)","missing":[],"search":"parentconfig banditrlproof.causal.graphmodel.parentconfig abbrev graphmodel.parentconfig {v : type*} {n : ℕ} (g : graphmodel v n) (i : fin n) abbreviation compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.parentConfig","label":"parentConfig","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.parentConfig","description":"def GraphModel.parentConfig {V : Type*} {n : ℕ} (g : GraphModel V n) (i : Fin n) (h : Fin i.val → V) : g.ParentConfig i","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-bf710e2666fa","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":534,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def GraphModel.parentConfig {V : Type*} {n : ℕ} (g : GraphModel V n) (i : Fin n) (h : Fin i.val → V) : g.ParentConfig i","missing":[],"search":"parentconfig banditrlproof.causal.graphmodel.parentconfig def graphmodel.parentconfig {v : type*} {n : ℕ} (g : graphmodel v n) (i : fin n) (h : fin i.val → v) : g.parentconfig i definition compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.parentTable","label":"parentTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.parentTable","description":"noncomputable def GraphModel.parentTable {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (i : Fin n) (z : g.ParentConfig i) : PMF V","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-bf5b546f8f07","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":535,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def GraphModel.parentTable {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (i : Fin n) (z : g.ParentConfig i) : PMF V","missing":[],"search":"parenttable banditrlproof.causal.graphmodel.parenttable noncomputable def graphmodel.parenttable {v : type*} [inhabited v] {n : ℕ} (g : graphmodel v n) (i : fin n) (z : g.parentconfig i) : pmf v definition compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.table_eq_parentTable","label":"table_eq_parentTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.table_eq_parentTable","description":"theorem GraphModel.table_eq_parentTable {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (i : Fin n) (h : Fin i.val → V) : g.table i h = g.parentTable i (g.parentConfig i h)","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-332efa7510e1","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":536,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.table_eq_parentTable {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (i : Fin n) (h : Fin i.val → V) : g.table i h = g.parentTable i (g.parentConfig i h)","missing":[],"search":"table_eq_parenttable banditrlproof.causal.graphmodel.table_eq_parenttable theorem graphmodel.table_eq_parenttable {v : type*} [inhabited v] {n : ℕ} (g : graphmodel v n) (i : fin n) (h : fin i.val → v) : g.table i h = g.parenttable i (g.parentconfig i h) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_map_last_pair","label":"joint_map_last_pair","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_map_last_pair","description":"theorem joint_map_last_pair {V Z : Type*} {n : ℕ} (p : Tables V (n+1)) (f : (Fin n → V) → Z) : (joint p).map (fun x => (f (Fin.init x), x (Fin.last n))) = (joint p.prefix).bind (fun h => (p (Fin.last n) h).map (fun y => (f h, y)))","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-e5e60f3ca3b0","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":537,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_map_last_pair {V Z : Type*} {n : ℕ} (p : Tables V (n+1)) (f : (Fin n → V) → Z) : (joint p).map (fun x => (f (Fin.init x), x (Fin.last n))) = (joint p.prefix).bind (fun h => (p (Fin.last n) h).map (fun y => (f h, y)))","missing":[],"search":"joint_map_last_pair banditrlproof.causal.joint_map_last_pair theorem joint_map_last_pair {v z : type*} {n : ℕ} (p : tables v (n+1)) (f : (fin n → v) → z) : (joint p).map (fun x => (f (fin.init x), x (fin.last n))) = (joint p.prefix).bind (fun h => (p (fin.last n) h).map (fun y => (f h, y))) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.last_parent_joint","label":"last_parent_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.last_parent_joint","description":"theorem GraphModel.last_parent_joint {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V (n+1)) : (joint g.table).map (fun x => (g.parentConfig (Fin.last n) (Fin.init x), x (Fin.last n))) = ((joint g.table.prefix).map (g.parentConfig (Fin.last n))).bind (fun z => (g.parentTable (Fin.last n) z).map (fun y => (z,y)))","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-d65a1a020b5b","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":538,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.last_parent_joint {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V (n+1)) : (joint g.table).map (fun x => (g.parentConfig (Fin.last n) (Fin.init x), x (Fin.last n))) = ((joint g.table.prefix).map (g.parentConfig (Fin.last n))).bind (fun z => (g.parentTable (Fin.last n) z).map (fun y => (z,y)))","missing":[],"search":"last_parent_joint banditrlproof.causal.graphmodel.last_parent_joint theorem graphmodel.last_parent_joint {v : type*} [inhabited v] {n : ℕ} (g : graphmodel v (n+1)) : (joint g.table).map (fun x => (g.parentconfig (fin.last n) (fin.init x), x (fin.last n))) = ((joint g.table.prefix).map (g.parentconfig (fin.last n))).bind (fun z => (g.parenttable (fin.last n) z).map (fun y => (z,y))) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_map_node_pair","label":"joint_map_node_pair","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_map_node_pair","description":"theorem joint_map_node_pair {V Z : Type*} {n : ℕ} (p : Tables V n) (i : Fin n) (f : (Fin i.val → V) → Z) : (joint p).map (fun x => (f (history x i), x i)) = (joint (p.restrict (Nat.le_of_lt i.isLt))).bind (fun h => (p i h).map (fun y => (f h,y)))","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-d5bc7d712dd4","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":539,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_map_node_pair {V Z : Type*} {n : ℕ} (p : Tables V n) (i : Fin n) (f : (Fin i.val → V) → Z) : (joint p).map (fun x => (f (history x i), x i)) = (joint (p.restrict (Nat.le_of_lt i.isLt))).bind (fun h => (p i h).map (fun y => (f h,y)))","missing":[],"search":"joint_map_node_pair banditrlproof.causal.joint_map_node_pair theorem joint_map_node_pair {v z : type*} {n : ℕ} (p : tables v n) (i : fin n) (f : (fin i.val → v) → z) : (joint p).map (fun x => (f (history x i), x i)) = (joint (p.restrict (nat.le_of_lt i.islt))).bind (fun h => (p i h).map (fun y => (f h,y))) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_map_history","label":"joint_map_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_map_history","description":"theorem joint_map_history {V : Type*} {n : ℕ} (p : Tables V n) (i : Fin n) : (joint p).map (fun x => history x i) = joint (p.restrict (Nat.le_of_lt i.isLt))","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-5b985a5892d1","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":540,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:80"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_map_history {V : Type*} {n : ℕ} (p : Tables V n) (i : Fin n) : (joint p).map (fun x => history x i) = joint (p.restrict (Nat.le_of_lt i.isLt))","missing":[],"search":"joint_map_history banditrlproof.causal.joint_map_history theorem joint_map_history {v : type*} {n : ℕ} (p : tables v n) (i : fin n) : (joint p).map (fun x => history x i) = joint (p.restrict (nat.le_of_lt i.islt)) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.parentLaw","label":"parentLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.parentLaw","description":"noncomputable def GraphModel.parentLaw {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) : PMF (g.ParentConfig i)","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-c6b17a22f50a","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":541,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def GraphModel.parentLaw {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) : PMF (g.ParentConfig i)","missing":[],"search":"parentlaw banditrlproof.causal.graphmodel.parentlaw noncomputable def graphmodel.parentlaw {v : type*} {n : ℕ} (g : graphmodel v n) (a : fin n → option v) (i : fin n) : pmf (g.parentconfig i) definition compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.parentLaw_prefix","label":"parentLaw_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.parentLaw_prefix","description":"theorem GraphModel.parentLaw_prefix {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) : g.parentLaw a i = (joint ((g.doModel a).table.restrict (Nat.le_of_lt i.isLt))).map (g.parentConfig i)","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-437f308cc243","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":542,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.parentLaw_prefix {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) : g.parentLaw a i = (joint ((g.doModel a).table.restrict (Nat.le_of_lt i.isLt))).map (g.parentConfig i)","missing":[],"search":"parentlaw_prefix banditrlproof.causal.graphmodel.parentlaw_prefix theorem graphmodel.parentlaw_prefix {v : type*} {n : ℕ} (g : graphmodel v n) (a : fin n → option v) (i : fin n) : g.parentlaw a i = (joint ((g.domodel a).table.restrict (nat.le_of_lt i.islt))).map (g.parentconfig i) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.intervention_parent_joint","label":"intervention_parent_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.intervention_parent_joint","description":"theorem GraphModel.intervention_parent_joint {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) (hi : a i = none) : (joint (g.doModel a).table).map (fun x => (g.parentConfig i (history x i),x i)) = (g.parentLaw a i).bind (fun z => (g.parentTable i z).map (fun y => (z,y)))","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-404f2e3ccb0a","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":543,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.intervention_parent_joint {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) (hi : a i = none) : (joint (g.doModel a).table).map (fun x => (g.parentConfig i (history x i),x i)) = (g.parentLaw a i).bind (fun z => (g.parentTable i z).map (fun y => (z,y)))","missing":[],"search":"intervention_parent_joint banditrlproof.causal.graphmodel.intervention_parent_joint theorem graphmodel.intervention_parent_joint {v : type*} [inhabited v] {n : ℕ} (g : graphmodel v n) (a : fin n → option v) (i : fin n) (hi : a i = none) : (joint (g.domodel a).table).map (fun x => (g.parentconfig i (history x i),x i)) = (g.parentlaw a i).bind (fun z => (g.parenttable i z).map (fun y => (z,y))) theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.paired_mass","label":"paired_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.paired_mass","description":"theorem paired_mass {Z V : Type*} (p : PMF Z) (k : Z → PMF V) (z : Z) (y : V) : (p.bind (fun w => (k w).map (fun v => (w,v)))) (z,y) = p z * k z y","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-df0d6488a040","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":544,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:107"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem paired_mass {Z V : Type*} (p : PMF Z) (k : Z → PMF V) (z : Z) (y : V) : (p.bind (fun w => (k w).map (fun v => (w,v)))) (z,y) = p z * k z y","missing":[],"search":"paired_mass banditrlproof.causal.paired_mass theorem paired_mass {z v : type*} (p : pmf z) (k : z → pmf v) (z : z) (y : v) : (p.bind (fun w => (k w).map (fun v => (w,v)))) (z,y) = p z * k z y theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.intervention_parent_mass","label":"intervention_parent_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.intervention_parent_mass","description":"theorem GraphModel.intervention_parent_mass {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) (hi : a i = none) (z : g.ParentConfig i) (y : V) : ((joint (g.doModel a).table).map (fun x => (g.parentConfig i (history x i),x i))) (z,y) = g.parentLaw a i z * g.parentTable i z y","url":"../modules/banditrlproof-algorithms-causalmarginallaw/index.html#decl-4584477f1d17","parent":"module:BanditRLProof.Algorithms.CausalMarginalLaw","order":545,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalMarginalLaw"],["Source","BanditRLProof/Algorithms/CausalMarginalLaw.lean:128"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.intervention_parent_mass {V : Type*} [Inhabited V] {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (i : Fin n) (hi : a i = none) (z : g.ParentConfig i) (y : V) : ((joint (g.doModel a).table).map (fun x => (g.parentConfig i (history x i),x i))) (z,y) = g.parentLaw a i z * g.parentTable i z y","missing":[],"search":"intervention_parent_mass banditrlproof.causal.graphmodel.intervention_parent_mass theorem graphmodel.intervention_parent_mass {v : type*} [inhabited v] {n : ℕ} (g : graphmodel v n) (a : fin n → option v) (i : fin n) (hi : a i = none) (z : g.parentconfig i) (y : v) : ((joint (g.domodel a).table).map (fun x => (g.parentconfig i (history x i),x i))) (z,y) = g.parentlaw a i z * g.parenttable i z y theorem compiled","shard":"modules/82d9a180187b8ad8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.allocationOfWeights","label":"allocationOfWeights","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.allocationOfWeights","description":"noncomputable def allocationOfWeights (w : A → ℝ) (hw : w ∈ stdSimplex ℝ A) : PMF A","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-a0c296201257","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":546,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def allocationOfWeights (w : A → ℝ) (hw : w ∈ stdSimplex ℝ A) : PMF A","missing":[],"search":"allocationofweights banditrlproof.causal.allocationofweights noncomputable def allocationofweights (w : a → ℝ) (hw : w ∈ stdsimplex ℝ a) : pmf a definition compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.allocationOfWeights_mass","label":"allocationOfWeights_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.allocationOfWeights_mass","description":"theorem allocationOfWeights_mass (w : A → ℝ) (hw : w ∈ stdSimplex ℝ A) (a : A) : mass (allocationOfWeights w hw) a = w a","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-f7f81cb93f3a","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":547,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem allocationOfWeights_mass (w : A → ℝ) (hw : w ∈ stdSimplex ℝ A) (a : A) : mass (allocationOfWeights w hw) a = w a","missing":[],"search":"allocationofweights_mass banditrlproof.causal.allocationofweights_mass theorem allocationofweights_mass (w : a → ℝ) (hw : w ∈ stdsimplex ℝ a) (a : a) : mass (allocationofweights w hw) a = w a theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mass_mem_simplex","label":"mass_mem_simplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mass_mem_simplex","description":"theorem mass_mem_simplex (eta : PMF A) : mass eta ∈ stdSimplex ℝ A","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-0431fde9e273","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":548,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mass_mem_simplex (eta : PMF A) : mass eta ∈ stdSimplex ℝ A","missing":[],"search":"mass_mem_simplex banditrlproof.causal.mass_mem_simplex theorem mass_mem_simplex (eta : pmf a) : mass eta ∈ stdsimplex ℝ a theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixtureWeight","label":"mixtureWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.mixtureWeight","description":"noncomputable def mixtureWeight (p : A → PMF Z) (w : A → ℝ) (z : Z) : ℝ","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-97e1dfae9a9a","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":549,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def mixtureWeight (p : A → PMF Z) (w : A → ℝ) (z : Z) : ℝ","missing":[],"search":"mixtureweight banditrlproof.causal.mixtureweight noncomputable def mixtureweight (p : a → pmf z) (w : a → ℝ) (z : z) : ℝ definition compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.coordinateCost","label":"coordinateCost","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.coordinateCost","description":"noncomputable def coordinateCost (p : A → PMF Z) (w : A → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-9fd69392ca1f","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":550,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def coordinateCost (p : A → PMF Z) (w : A → ℝ) : ℝ","missing":[],"search":"coordinatecost banditrlproof.causal.coordinatecost noncomputable def coordinatecost (p : a → pmf z) (w : a → ℝ) : ℝ definition compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.coordinateCost_mass","label":"coordinateCost_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.coordinateCost_mass","description":"theorem coordinateCost_mass (p : A → PMF Z) (eta : PMF A) : coordinateCost p (mass eta) = designCost p eta","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-485c8b7c9d31","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":551,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem coordinateCost_mass (p : A → PMF Z) (eta : PMF A) : coordinateCost p (mass eta) = designCost p eta","missing":[],"search":"coordinatecost_mass banditrlproof.causal.coordinatecost_mass theorem coordinatecost_mass (p : a → pmf z) (eta : pmf a) : coordinatecost p (mass eta) = designcost p eta theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.safeAllocations","label":"safeAllocations","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.safeAllocations","description":"def safeAllocations (p : A → PMF Z) : Set (A → ℝ)","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-2c5dd96446bf","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":552,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def safeAllocations (p : A → PMF Z) : Set (A → ℝ)","missing":[],"search":"safeallocations banditrlproof.causal.safeallocations def safeallocations (p : a → pmf z) : set (a → ℝ) definition compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.continuous_mixtureWeight","label":"continuous_mixtureWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.continuous_mixtureWeight","description":"theorem continuous_mixtureWeight (p : A → PMF Z) (z : Z) : Continuous (fun w => mixtureWeight p w z)","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-552a724483e7","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":553,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem continuous_mixtureWeight (p : A → PMF Z) (z : Z) : Continuous (fun w => mixtureWeight p w z)","missing":[],"search":"continuous_mixtureweight banditrlproof.causal.continuous_mixtureweight theorem continuous_mixtureweight (p : a → pmf z) (z : z) : continuous (fun w => mixtureweight p w z) theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.safeAllocations_compact","label":"safeAllocations_compact","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.safeAllocations_compact","description":"theorem safeAllocations_compact (p : A → PMF Z) : IsCompact (safeAllocations p)","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-c3b2d010e48e","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":554,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem safeAllocations_compact (p : A → PMF Z) : IsCompact (safeAllocations p)","missing":[],"search":"safeallocations_compact banditrlproof.causal.safeallocations_compact theorem safeallocations_compact (p : a → pmf z) : iscompact (safeallocations p) theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sublevel_mem_safe","label":"sublevel_mem_safe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sublevel_mem_safe","description":"theorem sublevel_mem_safe (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) (hcost : designCost p eta ≤ Fintype.card A) : mass eta ∈ safeAllocations p","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-c09973a6a857","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":555,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sublevel_mem_safe (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) (hcost : designCost p eta ≤ Fintype.card A) : mass eta ∈ safeAllocations p","missing":[],"search":"sublevel_mem_safe banditrlproof.causal.sublevel_mem_safe theorem sublevel_mem_safe (p : a → pmf z) (eta : pmf a) (hc : covers p (mixture eta p)) (hcost : designcost p eta ≤ fintype.card a) : mass eta ∈ safeallocations p theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.uniform_mem_safe","label":"uniform_mem_safe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.uniform_mem_safe","description":"theorem uniform_mem_safe (p : A → PMF Z) : mass (PMF.uniformOfFintype A) ∈ safeAllocations p","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-bb901c35c961","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":556,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniform_mem_safe (p : A → PMF Z) : mass (PMF.uniformOfFintype A) ∈ safeAllocations p","missing":[],"search":"uniform_mem_safe banditrlproof.causal.uniform_mem_safe theorem uniform_mem_safe (p : a → pmf z) : mass (pmf.uniformoffintype a) ∈ safeallocations p theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.safe_denominator_pos","label":"safe_denominator_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.safe_denominator_pos","description":"theorem safe_denominator_pos (p : A → PMF Z) (w : A → ℝ) (hw : w ∈ safeAllocations p) (a : A) (z : Z) (hp : mass (p a) z ≠ 0) : 0 < mixtureWeight p w z","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-ffdb1d328d55","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":557,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem safe_denominator_pos (p : A → PMF Z) (w : A → ℝ) (hw : w ∈ safeAllocations p) (a : A) (z : Z) (hp : mass (p a) z ≠ 0) : 0 < mixtureWeight p w z","missing":[],"search":"safe_denominator_pos banditrlproof.causal.safe_denominator_pos theorem safe_denominator_pos (p : a → pmf z) (w : a → ℝ) (hw : w ∈ safeallocations p) (a : a) (z : z) (hp : mass (p a) z ≠ 0) : 0 < mixtureweight p w z theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.coordinateCost_continuousOn","label":"coordinateCost_continuousOn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.coordinateCost_continuousOn","description":"theorem coordinateCost_continuousOn (p : A → PMF Z) : ContinuousOn (coordinateCost p) (safeAllocations p)","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-b00335ee56fb","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":558,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem coordinateCost_continuousOn (p : A → PMF Z) : ContinuousOn (coordinateCost p) (safeAllocations p)","missing":[],"search":"coordinatecost_continuouson banditrlproof.causal.coordinatecost_continuouson theorem coordinatecost_continuouson (p : a → pmf z) : continuouson (coordinatecost p) (safeallocations p) theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.safe_allocation_covers","label":"safe_allocation_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.safe_allocation_covers","description":"theorem safe_allocation_covers (p : A → PMF Z) (w : A → ℝ) (hw : w ∈ safeAllocations p) : Covers p (mixture (allocationOfWeights w hw.1) p)","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-350eb1f96293","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":559,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem safe_allocation_covers (p : A → PMF Z) (w : A → ℝ) (hw : w ∈ safeAllocations p) : Covers p (mixture (allocationOfWeights w hw.1) p)","missing":[],"search":"safe_allocation_covers banditrlproof.causal.safe_allocation_covers theorem safe_allocation_covers (p : a → pmf z) (w : a → ℝ) (hw : w ∈ safeallocations p) : covers p (mixture (allocationofweights w hw.1) p) theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.exists_optimal_allocation","label":"exists_optimal_allocation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.exists_optimal_allocation","description":"theorem exists_optimal_allocation (p : A → PMF Z) : ∃ eta : PMF A, Covers p (mixture eta p) ∧ designCost p eta ≤ Fintype.card A ∧ ∀ eta' : PMF A, Covers p (mixture eta' p) → designCost p eta ≤ designCost p eta'","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-a51b9ceca487","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":560,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_optimal_allocation (p : A → PMF Z) : ∃ eta : PMF A, Covers p (mixture eta p) ∧ designCost p eta ≤ Fintype.card A ∧ ∀ eta' : PMF A, Covers p (mixture eta' p) → designCost p eta ≤ designCost p eta'","missing":[],"search":"exists_optimal_allocation banditrlproof.causal.exists_optimal_allocation theorem exists_optimal_allocation (p : a → pmf z) : ∃ eta : pmf a, covers p (mixture eta p) ∧ designcost p eta ≤ fintype.card a ∧ ∀ eta' : pmf a, covers p (mixture eta' p) → designcost p eta ≤ designcost p eta' theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.optimalAllocation","label":"optimalAllocation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.optimalAllocation","description":"noncomputable def optimalAllocation (p : A → PMF Z) : PMF A","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-6479250e792a","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":561,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:119"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalAllocation (p : A → PMF Z) : PMF A","missing":[],"search":"optimalallocation banditrlproof.causal.optimalallocation noncomputable def optimalallocation (p : a → pmf z) : pmf a definition compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.optimalAllocation_covers","label":"optimalAllocation_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.optimalAllocation_covers","description":"theorem optimalAllocation_covers (p : A → PMF Z) : Covers p (mixture (optimalAllocation p) p)","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-95209745615f","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":562,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:122"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem optimalAllocation_covers (p : A → PMF Z) : Covers p (mixture (optimalAllocation p) p)","missing":[],"search":"optimalallocation_covers banditrlproof.causal.optimalallocation_covers theorem optimalallocation_covers (p : a → pmf z) : covers p (mixture (optimalallocation p) p) theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.optimalAllocation_cost_le_card","label":"optimalAllocation_cost_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.optimalAllocation_cost_le_card","description":"theorem optimalAllocation_cost_le_card (p : A → PMF Z) : designCost p (optimalAllocation p) ≤ Fintype.card A","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-5a60fd0f493c","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":563,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem optimalAllocation_cost_le_card (p : A → PMF Z) : designCost p (optimalAllocation p) ≤ Fintype.card A","missing":[],"search":"optimalallocation_cost_le_card banditrlproof.causal.optimalallocation_cost_le_card theorem optimalallocation_cost_le_card (p : a → pmf z) : designcost p (optimalallocation p) ≤ fintype.card a theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.optimalAllocation_minimizes","label":"optimalAllocation_minimizes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.optimalAllocation_minimizes","description":"theorem optimalAllocation_minimizes (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) : designCost p (optimalAllocation p) ≤ designCost p eta","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-062d45dba87f","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":564,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:129"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem optimalAllocation_minimizes (p : A → PMF Z) (eta : PMF A) (hc : Covers p (mixture eta p)) : designCost p (optimalAllocation p) ≤ designCost p eta","missing":[],"search":"optimalallocation_minimizes banditrlproof.causal.optimalallocation_minimizes theorem optimalallocation_minimizes (p : a → pmf z) (eta : pmf a) (hc : covers p (mixture eta p)) : designcost p (optimalallocation p) ≤ designcost p eta theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment_self","label":"secondMoment_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment_self","description":"theorem secondMoment_self (q : PMF Z) : secondMoment q q = 1","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-cf55107d36f1","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":565,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem secondMoment_self (q : PMF Z) : secondMoment q q = 1","missing":[],"search":"secondmoment_self banditrlproof.causal.secondmoment_self theorem secondmoment_self (q : pmf z) : secondmoment q q = 1 theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.constant_designCost","label":"constant_designCost","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.constant_designCost","description":"theorem constant_designCost (q : PMF Z) (eta : PMF A) : designCost (fun _ : A => q) eta = 1","url":"../modules/banditrlproof-algorithms-causaloptimalallocation/index.html#decl-12b6b465c9a1","parent":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","order":566,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOptimalAllocation"],["Source","BanditRLProof/Algorithms/CausalOptimalAllocation.lean:138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem constant_designCost (q : PMF Z) (eta : PMF A) : designCost (fun _ : A => q) eta = 1","missing":[],"search":"constant_designcost banditrlproof.causal.constant_designcost theorem constant_designcost (q : pmf z) (eta : pmf a) : designcost (fun _ : a => q) eta = 1 theorem compiled","shard":"modules/1e9c6c116515e2db.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.Tables","label":"Tables","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Causal.Tables","description":"abbrev Tables (V : Type*) (n : ℕ)","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-8b430688dcd6","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":567,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Tables (V : Type*) (n : ℕ)","missing":[],"search":"tables banditrlproof.causal.tables abbrev tables (v : type*) (n : ℕ) abbreviation compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.Tables.prefix","label":"prefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.Tables.prefix","description":"def Tables.prefix {V : Type*} {n : ℕ} (p : Tables V (n+1)) : Tables V n","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-d79c49c8deb9","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":568,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def Tables.prefix {V : Type*} {n : ℕ} (p : Tables V (n+1)) : Tables V n","missing":[],"search":"prefix banditrlproof.causal.tables.prefix def tables.prefix {v : type*} {n : ℕ} (p : tables v (n+1)) : tables v n definition compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint","label":"joint","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.joint","description":"noncomputable def joint {V : Type*} : {n : ℕ} → Tables V n → PMF (Fin n → V) | 0, _ => PMF.pure Fin.elim0 | n+1, p => (joint p.prefix).bind fun h => (p (Fin.last n) h).bind fun x => PMF.pure (Fin.snoc h x) noncomputable def intervene {V : Type*} {n : ℕ} (p : Tables V n) (a : Fin n → Option V) : Tables V n","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-4a62e02cdd99","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":569,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def joint {V : Type*} : {n : ℕ} → Tables V n → PMF (Fin n → V) | 0, _ => PMF.pure Fin.elim0 | n+1, p => (joint p.prefix).bind fun h => (p (Fin.last n) h).bind fun x => PMF.pure (Fin.snoc h x) noncomputable def intervene {V : Type*} {n : ℕ} (p : Tables V n) (a : Fin n → Option V) : Tables V n","missing":[],"search":"joint banditrlproof.causal.joint noncomputable def joint {v : type*} : {n : ℕ} → tables v n → pmf (fin n → v) | 0, _ => pmf.pure fin.elim0 | n+1, p => (joint p.prefix).bind fun h => (p (fin.last n) h).bind fun x => pmf.pure (fin.snoc h x) noncomputable def intervene {v : type*} {n : ℕ} (p : tables v n) (a : fin n → option v) : tables v n definition compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.intervene","label":"intervene","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.intervene","description":"noncomputable def intervene {V : Type*} {n : ℕ} (p : Tables V n) (a : Fin n → Option V) : Tables V n","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-55df7b6f91ac","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":570,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def intervene {V : Type*} {n : ℕ} (p : Tables V n) (a : Fin n → Option V) : Tables V n","missing":[],"search":"intervene banditrlproof.causal.intervene noncomputable def intervene {v : type*} {n : ℕ} (p : tables v n) (a : fin n → option v) : tables v n definition compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.intervene_none","label":"intervene_none","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.intervene_none","description":"theorem intervene_none {V : Type*} {n : ℕ} (p : Tables V n) : intervene p (fun _ => none) = p","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-b9422b50d96d","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":571,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem intervene_none {V : Type*} {n : ℕ} (p : Tables V n) : intervene p (fun _ => none) = p","missing":[],"search":"intervene_none banditrlproof.causal.intervene_none theorem intervene_none {v : type*} {n : ℕ} (p : tables v n) : intervene p (fun _ => none) = p theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.intervene_at","label":"intervene_at","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.intervene_at","description":"theorem intervene_at {V : Type*} {n : ℕ} (p : Tables V n) (a : Fin n → Option V) (i : Fin n) (x : V) (ha : a i = some x) (h : Fin i.val → V) : intervene p a i h = PMF.pure x","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-4a08932bade9","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":572,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem intervene_at {V : Type*} {n : ℕ} (p : Tables V n) (a : Fin n → Option V) (i : Fin n) (x : V) (ha : a i = some x) (h : Fin i.val → V) : intervene p a i h = PMF.pure x","missing":[],"search":"intervene_at banditrlproof.causal.intervene_at theorem intervene_at {v : type*} {n : ℕ} (p : tables v n) (a : fin n → option v) (i : fin n) (x : v) (ha : a i = some x) (h : fin i.val → v) : intervene p a i h = pmf.pure x theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_snoc","label":"joint_snoc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_snoc","description":"theorem joint_snoc {V : Type*} {n : ℕ} (p : Tables V (n+1)) (h : Fin n → V) (x : V) : joint p (Fin.snoc h x) = joint p.prefix h * p (Fin.last n) h x","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-01afd7f2d31f","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":573,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_snoc {V : Type*} {n : ℕ} (p : Tables V (n+1)) (h : Fin n → V) (x : V) : joint p (Fin.snoc h x) = joint p.prefix h * p (Fin.last n) h x","missing":[],"search":"joint_snoc banditrlproof.causal.joint_snoc theorem joint_snoc {v : type*} {n : ℕ} (p : tables v (n+1)) (h : fin n → v) (x : v) : joint p (fin.snoc h x) = joint p.prefix h * p (fin.last n) h x theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.history","label":"history","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.history","description":"def history {V : Type*} {n : ℕ} (x : Fin n → V) (i : Fin n) : Fin i.val → V","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-f1ffd497a77d","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":574,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def history {V : Type*} {n : ℕ} (x : Fin n → V) (i : Fin n) : Fin i.val → V","missing":[],"search":"history banditrlproof.causal.history def history {v : type*} {n : ℕ} (x : fin n → v) (i : fin n) : fin i.val → v definition compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_factorization","label":"joint_factorization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_factorization","description":"theorem joint_factorization {V : Type*} {n : ℕ} (p : Tables V n) (x : Fin n → V) : joint p x = ∏ i : Fin n, p i (history x i) (x i)","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-1036cf42e35f","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":575,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_factorization {V : Type*} {n : ℕ} (p : Tables V n) (x : Fin n → V) : joint p x = ∏ i : Fin n, p i (history x i) (x i)","missing":[],"search":"joint_factorization banditrlproof.causal.joint_factorization theorem joint_factorization {v : type*} {n : ℕ} (p : tables v n) (x : fin n → v) : joint p x = ∏ i : fin n, p i (history x i) (x i) theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.joint_normalized","label":"joint_normalized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.joint_normalized","description":"theorem joint_normalized {V : Type*} {n : ℕ} (p : Tables V n) : ∑' x, joint p x = 1","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-dcee074b75f7","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":576,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem joint_normalized {V : Type*} {n : ℕ} (p : Tables V n) : ∑' x, joint p x = 1","missing":[],"search":"joint_normalized banditrlproof.causal.joint_normalized theorem joint_normalized {v : type*} {n : ℕ} (p : tables v n) : ∑' x, joint p x = 1 theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel","label":"GraphModel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel","description":"Parent indices are strictly earlier in the topological order.","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-65bb1cd58399","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":577,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure GraphModel (V : Type*) (n : ℕ) where","missing":[],"search":"graphmodel banditrlproof.causal.graphmodel parent indices are strictly earlier in the topological order. structure compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.doModel","label":"doModel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.doModel","description":"noncomputable def GraphModel.doModel {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) : GraphModel V n where","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-5f765bcd9e6f","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":578,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:80"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def GraphModel.doModel {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) : GraphModel V n where","missing":[],"search":"domodel banditrlproof.causal.graphmodel.domodel noncomputable def graphmodel.domodel {v : type*} {n : ℕ} (g : graphmodel v n) (a : fin n → option v) : graphmodel v n where definition compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.doModel_factorization","label":"doModel_factorization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.doModel_factorization","description":"theorem doModel_factorization {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (x : Fin n → V) : joint (g.doModel a).table x = ∏ i : Fin n, (match a i with | none => g.table i (history x i) (x i) | some v => if x i = v then 1 else 0)","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-8e4174bbedab","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":579,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem doModel_factorization {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (x : Fin n → V) : joint (g.doModel a).table x = ∏ i : Fin n, (match a i with | none => g.table i (history x i) (x i) | some v => if x i = v then 1 else 0)","missing":[],"search":"domodel_factorization banditrlproof.causal.domodel_factorization theorem domodel_factorization {v : type*} {n : ℕ} (g : graphmodel v n) (a : fin n → option v) (x : fin n → v) : joint (g.domodel a).table x = ∏ i : fin n, (match a i with | none => g.table i (history x i) (x i) | some v => if x i = v then 1 else 0) theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.intervention_incompatible_zero","label":"intervention_incompatible_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.intervention_incompatible_zero","description":"theorem intervention_incompatible_zero {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (x : Fin n → V) (i : Fin n) (v : V) (ha : a i = some v) (hx : x i ≠ v) : joint (g.doModel a).table x = 0","url":"../modules/banditrlproof-algorithms-causalorderedlaw/index.html#decl-326b72be655d","parent":"module:BanditRLProof.Algorithms.CausalOrderedLaw","order":580,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalOrderedLaw"],["Source","BanditRLProof/Algorithms/CausalOrderedLaw.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem intervention_incompatible_zero {V : Type*} {n : ℕ} (g : GraphModel V n) (a : Fin n → Option V) (x : Fin n → V) (i : Fin n) (v : V) (ha : a i = some v) (hx : x i ≠ v) : joint (g.doModel a).table x = 0","missing":[],"search":"intervention_incompatible_zero banditrlproof.causal.intervention_incompatible_zero theorem intervention_incompatible_zero {v : type*} {n : ℕ} (g : graphmodel v n) (a : fin n → option v) (x : fin n → v) (i : fin n) (v : v) (ha : a i = some v) (hx : x i ≠ v) : joint (g.domodel a).table x = 0 theorem compiled","shard":"modules/2bfec485485647c1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters","label":"ParallelParameters","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters","description":"structure ParallelParameters (N : ℕ) where","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-773fc0e9f388","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":581,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure ParallelParameters (N : ℕ) where","missing":[],"search":"parallelparameters banditrlproof.causal.parallelparameters structure parallelparameters (n : ℕ) where structure compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rareIndices","label":"rareIndices","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rareIndices","description":"noncomputable def rareIndices (tau : ℕ) : Finset (Fin N)","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-59ac72a257f6","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":582,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def rareIndices (tau : ℕ) : Finset (Fin N)","missing":[],"search":"rareindices banditrlproof.causal.parallelparameters.rareindices noncomputable def rareindices (tau : ℕ) : finset (fin n) definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.exists_rarity","label":"exists_rarity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.exists_rarity","description":"theorem exists_rarity : ∃ m : ℕ, 2 ≤ m ∧ m ≤ N ∧ (p.rareIndices m).card ≤ m","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-4fa588724f0e","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":583,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_rarity : ∃ m : ℕ, 2 ≤ m ∧ m ≤ N ∧ (p.rareIndices m).card ≤ m","missing":[],"search":"exists_rarity banditrlproof.causal.parallelparameters.exists_rarity theorem exists_rarity : ∃ m : ℕ, 2 ≤ m ∧ m ≤ n ∧ (p.rareindices m).card ≤ m theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rarity","label":"rarity","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rarity","description":"noncomputable def rarity : ℕ","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-fc780b1faf31","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":584,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def rarity : ℕ","missing":[],"search":"rarity banditrlproof.causal.parallelparameters.rarity noncomputable def rarity : ℕ definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_spec","label":"rarity_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rarity_spec","description":"theorem rarity_spec : 2 ≤ p.rarity ∧ p.rarity ≤ N ∧ (p.rareIndices p.rarity).card ≤ p.rarity","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-f2edfe8281bf","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":585,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem rarity_spec : 2 ≤ p.rarity ∧ p.rarity ≤ N ∧ (p.rareIndices p.rarity).card ≤ p.rarity","missing":[],"search":"rarity_spec banditrlproof.causal.parallelparameters.rarity_spec theorem rarity_spec : 2 ≤ p.rarity ∧ p.rarity ≤ n ∧ (p.rareindices p.rarity).card ≤ p.rarity theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_minimal","label":"rarity_minimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rarity_minimal","description":"theorem rarity_minimal (tau : ℕ) (h2 : 2 ≤ tau) (hN : tau ≤ N) (hc : (p.rareIndices tau).card ≤ tau) : p.rarity ≤ tau","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-ee0850baaaa2","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":586,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem rarity_minimal (tau : ℕ) (h2 : 2 ≤ tau) (hN : tau ≤ N) (hc : (p.rareIndices tau).card ≤ tau) : p.rarity ≤ tau","missing":[],"search":"rarity_minimal banditrlproof.causal.parallelparameters.rarity_minimal theorem rarity_minimal (tau : ℕ) (h2 : 2 ≤ tau) (hn : tau ≤ n) (hc : (p.rareindices tau).card ≤ tau) : p.rarity ≤ tau theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.valueProbability","label":"valueProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.valueProbability","description":"def valueProbability (a : Fin N × Bool) : ℝ","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-cb429fc03daf","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":587,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def valueProbability (a : Fin N × Bool) : ℝ","missing":[],"search":"valueprobability banditrlproof.causal.parallelparameters.valueprobability def valueprobability (a : fin n × bool) : ℝ definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.valueProbability_nonneg","label":"valueProbability_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.valueProbability_nonneg","description":"theorem valueProbability_nonneg (a : Fin N × Bool) : 0 ≤ p.valueProbability a","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-06d6171ce02d","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":588,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem valueProbability_nonneg (a : Fin N × Bool) : 0 ≤ p.valueProbability a","missing":[],"search":"valueprobability_nonneg banditrlproof.causal.parallelparameters.valueprobability_nonneg theorem valueprobability_nonneg (a : fin n × bool) : 0 ≤ p.valueprobability a theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_pos","label":"rarity_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rarity_pos","description":"theorem rarity_pos : (0 : ℝ) < p.rarity","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-a823d30b4f14","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":589,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rarity_pos : (0 : ℝ) < p.rarity","missing":[],"search":"rarity_pos banditrlproof.causal.parallelparameters.rarity_pos theorem rarity_pos : (0 : ℝ) < p.rarity theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_reciprocal_le_half","label":"rarity_reciprocal_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rarity_reciprocal_le_half","description":"theorem rarity_reciprocal_le_half : 1/(p.rarity:ℝ) ≤ 1/2","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-72e805837aba","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":590,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rarity_reciprocal_le_half : 1/(p.rarity:ℝ) ≤ 1/2","missing":[],"search":"rarity_reciprocal_le_half banditrlproof.causal.parallelparameters.rarity_reciprocal_le_half theorem rarity_reciprocal_le_half : 1/(p.rarity:ℝ) ≤ 1/2 theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rareActions","label":"rareActions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rareActions","description":"noncomputable def rareActions : Finset (Fin N × Bool)","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-271c4ced1910","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":591,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def rareActions : Finset (Fin N × Bool)","missing":[],"search":"rareactions banditrlproof.causal.parallelparameters.rareactions noncomputable def rareactions : finset (fin n × bool) definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rareActions_card_le","label":"rareActions_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rareActions_card_le","description":"theorem rareActions_card_le : p.rareActions.card ≤ p.rarity","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-8b35dd316c50","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":592,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem rareActions_card_le : p.rareActions.card ≤ p.rarity","missing":[],"search":"rareactions_card_le banditrlproof.causal.parallelparameters.rareactions_card_le theorem rareactions_card_le : p.rareactions.card ≤ p.rarity theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rareWeight","label":"rareWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rareWeight","description":"noncomputable def rareWeight : ℝ","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-855c18df8f78","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":593,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def rareWeight : ℝ","missing":[],"search":"rareweight banditrlproof.causal.parallelparameters.rareweight noncomputable def rareweight : ℝ definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.atomicTotal","label":"atomicTotal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.atomicTotal","description":"noncomputable def atomicTotal : ℝ","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-50f3989327e8","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":594,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def atomicTotal : ℝ","missing":[],"search":"atomictotal banditrlproof.causal.parallelparameters.atomictotal noncomputable def atomictotal : ℝ definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rareWeight_pos","label":"rareWeight_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rareWeight_pos","description":"theorem rareWeight_pos : 0 < p.rareWeight","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-05bfb56c856b","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":595,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rareWeight_pos : 0 < p.rareWeight","missing":[],"search":"rareweight_pos banditrlproof.causal.parallelparameters.rareweight_pos theorem rareweight_pos : 0 < p.rareweight theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.atomicTotal_bounds","label":"atomicTotal_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.atomicTotal_bounds","description":"theorem atomicTotal_bounds : 0 ≤ p.atomicTotal ∧ p.atomicTotal ≤ 1/2","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-9d79f084f5ab","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":596,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem atomicTotal_bounds : 0 ≤ p.atomicTotal ∧ p.atomicTotal ≤ 1/2","missing":[],"search":"atomictotal_bounds banditrlproof.causal.parallelparameters.atomictotal_bounds theorem atomictotal_bounds : 0 ≤ p.atomictotal ∧ p.atomictotal ≤ 1/2 theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocationWeight","label":"allocationWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocationWeight","description":"noncomputable def allocationWeight : Option (Fin N × Bool) → ℝ | none => 1-p.atomicTotal | some a => if a ∈ p.rareActions then p.rareWeight else 0 theorem allocationWeight_nonneg (a : Option (Fin N × Bool)) : 0 ≤ p.allocationWeight a","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-662762171939","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":597,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def allocationWeight : Option (Fin N × Bool) → ℝ | none => 1-p.atomicTotal | some a => if a ∈ p.rareActions then p.rareWeight else 0 theorem allocationWeight_nonneg (a : Option (Fin N × Bool)) : 0 ≤ p.allocationWeight a","missing":[],"search":"allocationweight banditrlproof.causal.parallelparameters.allocationweight noncomputable def allocationweight : option (fin n × bool) → ℝ | none => 1-p.atomictotal | some a => if a ∈ p.rareactions then p.rareweight else 0 theorem allocationweight_nonneg (a : option (fin n × bool)) : 0 ≤ p.allocationweight a definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocationWeight_nonneg","label":"allocationWeight_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocationWeight_nonneg","description":"theorem allocationWeight_nonneg (a : Option (Fin N × Bool)) : 0 ≤ p.allocationWeight a","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-2b9fe14af3b4","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":598,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem allocationWeight_nonneg (a : Option (Fin N × Bool)) : 0 ≤ p.allocationWeight a","missing":[],"search":"allocationweight_nonneg banditrlproof.causal.parallelparameters.allocationweight_nonneg theorem allocationweight_nonneg (a : option (fin n × bool)) : 0 ≤ p.allocationweight a theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocationWeight_sum","label":"allocationWeight_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocationWeight_sum","description":"theorem allocationWeight_sum : ∑ a, p.allocationWeight a = 1","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-f138236c061b","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":599,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:100"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem allocationWeight_sum : ∑ a, p.allocationWeight a = 1","missing":[],"search":"allocationweight_sum banditrlproof.causal.parallelparameters.allocationweight_sum theorem allocationweight_sum : ∑ a, p.allocationweight a = 1 theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocation","label":"allocation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocation","description":"noncomputable def allocation : PMF (Option (Fin N × Bool))","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-df1d2125c875","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":600,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def allocation : PMF (Option (Fin N × Bool))","missing":[],"search":"allocation banditrlproof.causal.parallelparameters.allocation noncomputable def allocation : pmf (option (fin n × bool)) definition compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_mass","label":"allocation_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocation_mass","description":"theorem allocation_mass (a : Option (Fin N × Bool)) : mass p.allocation a = p.allocationWeight a","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-de12954d5b26","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":601,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem allocation_mass (a : Option (Fin N × Bool)) : mass p.allocation a = p.allocationWeight a","missing":[],"search":"allocation_mass banditrlproof.causal.parallelparameters.allocation_mass theorem allocation_mass (a : option (fin n × bool)) : mass p.allocation a = p.allocationweight a theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_empty_ge_half","label":"allocation_empty_ge_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocation_empty_ge_half","description":"theorem allocation_empty_ge_half : 1/2 ≤ mass p.allocation none","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-175806c33e47","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":602,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:114"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem allocation_empty_ge_half : 1/2 ≤ mass p.allocation none","missing":[],"search":"allocation_empty_ge_half banditrlproof.causal.parallelparameters.allocation_empty_ge_half theorem allocation_empty_ge_half : 1/2 ≤ mass p.allocation none theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.covers_of_mass_domination","label":"covers_of_mass_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.covers_of_mass_domination","description":"theorem covers_of_mass_domination {A Z : Type*} (p : A → PMF Z) (q : PMF Z) (C : ℝ) (h : ∀ a z, mass (p a) z ≤ C*mass q z) : Covers p q","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-11861747775c","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":603,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem covers_of_mass_domination {A Z : Type*} (p : A → PMF Z) (q : PMF Z) (C : ℝ) (h : ∀ a z, mass (p a) z ≤ C*mass q z) : Covers p q","missing":[],"search":"covers_of_mass_domination banditrlproof.causal.covers_of_mass_domination theorem covers_of_mass_domination {a z : type*} (p : a → pmf z) (q : pmf z) (c : ℝ) (h : ∀ a z, mass (p a) z ≤ c*mass q z) : covers p q theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ratio_le_of_mass_domination","label":"ratio_le_of_mass_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ratio_le_of_mass_domination","description":"theorem ratio_le_of_mass_domination {Z : Type*} (p q : PMF Z) (C : ℝ) (hC : 0 ≤ C) (h : ∀ z, mass p z ≤ C*mass q z) (z : Z) : ratio p q z ≤ C","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-97f33edbad70","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":604,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:128"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ratio_le_of_mass_domination {Z : Type*} (p q : PMF Z) (C : ℝ) (hC : 0 ≤ C) (h : ∀ z, mass p z ≤ C*mass q z) (z : Z) : ratio p q z ≤ C","missing":[],"search":"ratio_le_of_mass_domination banditrlproof.causal.ratio_le_of_mass_domination theorem ratio_le_of_mass_domination {z : type*} (p q : pmf z) (c : ℝ) (hc : 0 ≤ c) (h : ∀ z, mass p z ≤ c*mass q z) (z : z) : ratio p q z ≤ c theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.secondMoment_le_of_mass_domination","label":"secondMoment_le_of_mass_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.secondMoment_le_of_mass_domination","description":"theorem secondMoment_le_of_mass_domination {Z : Type*} [Fintype Z] (p q : PMF Z) (C : ℝ) (hC : 0 ≤ C) (h : ∀ z, mass p z ≤ C*mass q z) : secondMoment p q ≤ C","url":"../modules/banditrlproof-algorithms-causalparalleldesign/index.html#decl-55d03b6f0a98","parent":"module:BanditRLProof.Algorithms.CausalParallelDesign","order":605,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelDesign"],["Source","BanditRLProof/Algorithms/CausalParallelDesign.lean:134"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem secondMoment_le_of_mass_domination {Z : Type*} [Fintype Z] (p q : PMF Z) (C : ℝ) (hC : 0 ≤ C) (h : ∀ z, mass p z ≤ C*mass q z) : secondMoment p q ≤ C","missing":[],"search":"secondmoment_le_of_mass_domination banditrlproof.causal.secondmoment_le_of_mass_domination theorem secondmoment_le_of_mass_domination {z : type*} [fintype z] (p q : pmf z) (c : ℝ) (hc : 0 ≤ c) (h : ∀ z, mass p z ≤ c*mass q z) : secondmoment p q ≤ c theorem compiled","shard":"modules/33beb34818f8020b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.mixture_mass_ge_component","label":"mixture_mass_ge_component","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.mixture_mass_ge_component","description":"theorem mixture_mass_ge_component {A Z : Type*} [Fintype A] (eta : PMF A) (p : A → PMF Z) (a : A) (z : Z) : mass eta a * mass (p a) z ≤ mass (mixture eta p) z","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-4d2cf0f92190","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":606,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mixture_mass_ge_component {A Z : Type*} [Fintype A] (eta : PMF A) (p : A → PMF Z) (a : A) (z : Z) : mass eta a * mass (p a) z ≤ mass (mixture eta p) z","missing":[],"search":"mixture_mass_ge_component banditrlproof.causal.mixture_mass_ge_component theorem mixture_mass_ge_component {a z : type*} [fintype a] (eta : pmf a) (p : a → pmf z) (a : a) (z : z) : mass eta a * mass (p a) z ≤ mass (mixture eta p) z theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rootTable","label":"rootTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rootTable","description":"noncomputable def rootTable (i : Fin N) : PMF Bool","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-2c7a3f7cd9c6","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":607,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def rootTable (i : Fin N) : PMF Bool","missing":[],"search":"roottable banditrlproof.causal.parallelparameters.roottable noncomputable def roottable (i : fin n) : pmf bool definition compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rootTable_mass","label":"rootTable_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rootTable_mass","description":"theorem rootTable_mass (i : Fin N) (b : Bool) : mass (p.rootTable i) b = p.valueProbability (i,b)","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-159a047f82a7","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":608,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rootTable_mass (i : Fin N) (b : Bool) : mass (p.rootTable i) b = p.valueProbability (i,b)","missing":[],"search":"roottable_mass banditrlproof.causal.parallelparameters.roottable_mass theorem roottable_mass (i : fin n) (b : bool) : mass (p.roottable i) b = p.valueprobability (i,b) theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rootAction","label":"rootAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rootAction","description":"def rootAction (a : Option (Fin N × Bool)) (i : Fin N) : Option Bool","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-c3d3d676a2b4","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":609,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def rootAction (a : Option (Fin N × Bool)) (i : Fin N) : Option Bool","missing":[],"search":"rootaction banditrlproof.causal.parallelparameters.rootaction def rootaction (a : option (fin n × bool)) (i : fin n) : option bool definition compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rootLaw","label":"rootLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rootLaw","description":"noncomputable def rootLaw (a : Option (Fin N × Bool)) : PMF (Fin N → Bool)","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-c9aedbb912a0","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":610,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def rootLaw (a : Option (Fin N × Bool)) : PMF (Fin N → Bool)","missing":[],"search":"rootlaw banditrlproof.causal.parallelparameters.rootlaw noncomputable def rootlaw (a : option (fin n × bool)) : pmf (fin n → bool) definition compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rootLaw_mass","label":"rootLaw_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rootLaw_mass","description":"theorem rootLaw_mass (a : Option (Fin N × Bool)) (x : Fin N → Bool) : mass (p.rootLaw a) x = ∏ i, match a with | none => p.valueProbability (i,x i) | some b => if i = b.1 then (if x i = b.2 then 1 else 0) else p.valueProbability (i,x i)","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-2e30a95acee4","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":611,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem rootLaw_mass (a : Option (Fin N × Bool)) (x : Fin N → Bool) : mass (p.rootLaw a) x = ∏ i, match a with | none => p.valueProbability (i,x i) | some b => if i = b.1 then (if x i = b.2 then 1 else 0) else p.valueProbability (i,x i)","missing":[],"search":"rootlaw_mass banditrlproof.causal.parallelparameters.rootlaw_mass theorem rootlaw_mass (a : option (fin n × bool)) (x : fin n → bool) : mass (p.rootlaw a) x = ∏ i, match a with | none => p.valueprobability (i,x i) | some b => if i = b.1 then (if x i = b.2 then 1 else 0) else p.valueprobability (i,x i) theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.rootLaw_atom_mass","label":"rootLaw_atom_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.rootLaw_atom_mass","description":"theorem rootLaw_atom_mass (i : Fin N) (b : Bool) (x : Fin N → Bool) : mass (p.rootLaw (some (i,b))) x = (if x i = b then 1 else 0) * ∏ j ∈ Finset.univ.erase i, p.valueProbability (j,x j)","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-ee134c10c1e6","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":612,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rootLaw_atom_mass (i : Fin N) (b : Bool) (x : Fin N → Bool) : mass (p.rootLaw (some (i,b))) x = (if x i = b then 1 else 0) * ∏ j ∈ Finset.univ.erase i, p.valueProbability (j,x j)","missing":[],"search":"rootlaw_atom_mass banditrlproof.causal.parallelparameters.rootlaw_atom_mass theorem rootlaw_atom_mass (i : fin n) (b : bool) (x : fin n → bool) : mass (p.rootlaw (some (i,b))) x = (if x i = b then 1 else 0) * ∏ j ∈ finset.univ.erase i, p.valueprobability (j,x j) theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.atomic_observational_domination","label":"atomic_observational_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.atomic_observational_domination","description":"theorem atomic_observational_domination (a : Fin N × Bool) (x : Fin N → Bool) : p.valueProbability a * mass (p.rootLaw (some a)) x ≤ mass (p.rootLaw none) x","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-02d6c6101544","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":613,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem atomic_observational_domination (a : Fin N × Bool) (x : Fin N → Bool) : p.valueProbability a * mass (p.rootLaw (some a)) x ≤ mass (p.rootLaw none) x","missing":[],"search":"atomic_observational_domination banditrlproof.causal.parallelparameters.atomic_observational_domination theorem atomic_observational_domination (a : fin n × bool) (x : fin n → bool) : p.valueprobability a * mass (p.rootlaw (some a)) x ≤ mass (p.rootlaw none) x theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocated_mass_domination","label":"allocated_mass_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocated_mass_domination","description":"theorem allocated_mass_domination (a : Option (Fin N × Bool)) (x : Fin N → Bool) : mass (p.rootLaw a) x ≤ (2*p.rarity)*mass (mixture p.allocation p.rootLaw) x","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-e006b79d0a5f","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":614,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem allocated_mass_domination (a : Option (Fin N × Bool)) (x : Fin N → Bool) : mass (p.rootLaw a) x ≤ (2*p.rarity)*mass (mixture p.allocation p.rootLaw) x","missing":[],"search":"allocated_mass_domination banditrlproof.causal.parallelparameters.allocated_mass_domination theorem allocated_mass_domination (a : option (fin n × bool)) (x : fin n → bool) : mass (p.rootlaw a) x ≤ (2*p.rarity)*mass (mixture p.allocation p.rootlaw) x theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_covers","label":"allocation_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocation_covers","description":"theorem allocation_covers : Covers p.rootLaw (mixture p.allocation p.rootLaw)","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-6ce88db930ae","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":615,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem allocation_covers : Covers p.rootLaw (mixture p.allocation p.rootLaw)","missing":[],"search":"allocation_covers banditrlproof.causal.parallelparameters.allocation_covers theorem allocation_covers : covers p.rootlaw (mixture p.allocation p.rootlaw) theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_cost_le","label":"allocation_cost_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.allocation_cost_le","description":"theorem allocation_cost_le : designCost p.rootLaw p.allocation ≤ 2*p.rarity","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-8a16a99c9248","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":616,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem allocation_cost_le : designCost p.rootLaw p.allocation ≤ 2*p.rarity","missing":[],"search":"allocation_cost_le banditrlproof.causal.parallelparameters.allocation_cost_le theorem allocation_cost_le : designcost p.rootlaw p.allocation ≤ 2*p.rarity theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.optimal_cost_le","label":"optimal_cost_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.optimal_cost_le","description":"theorem optimal_cost_le : designCost p.rootLaw (optimalAllocation p.rootLaw) ≤ 2*p.rarity","url":"../modules/banditrlproof-algorithms-causalparallellaw/index.html#decl-2452f31955fe","parent":"module:BanditRLProof.Algorithms.CausalParallelLaw","order":617,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelLaw"],["Source","BanditRLProof/Algorithms/CausalParallelLaw.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem optimal_cost_le : designCost p.rootLaw (optimalAllocation p.rootLaw) ≤ 2*p.rarity","missing":[],"search":"optimal_cost_le banditrlproof.causal.parallelparameters.optimal_cost_le theorem optimal_cost_le : designcost p.rootlaw (optimalallocation p.rootlaw) ≤ 2*p.rarity theorem compiled","shard":"modules/cecc5e09203a86a4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph","label":"graph","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph","description":"noncomputable def graph : GraphModel Bool (N+1) where","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-1a0d5f74f619","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":618,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def graph : GraphModel Bool (N+1) where","missing":[],"search":"graph banditrlproof.causal.parallelparameters.graph noncomputable def graph : graphmodel bool (n+1) where definition compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graphAction","label":"graphAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graphAction","description":"def graphAction (a : Option (Fin N × Bool)) : Fin (N+1) → Option Bool","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-68779bd4a0fc","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":619,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def graphAction (a : Option (Fin N × Bool)) : Fin (N+1) → Option Bool","missing":[],"search":"graphaction banditrlproof.causal.parallelparameters.graphaction def graphaction (a : option (fin n × bool)) : fin (n+1) → option bool definition compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graphAction_reward","label":"graphAction_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graphAction_reward","description":"theorem graphAction_reward (a : Option (Fin N × Bool)) : graphAction a (Fin.last N) = none","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-d097dfc90976","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":620,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem graphAction_reward (a : Option (Fin N × Bool)) : graphAction a (Fin.last N) = none","missing":[],"search":"graphaction_reward banditrlproof.causal.parallelparameters.graphaction_reward theorem graphaction_reward (a : option (fin n × bool)) : graphaction a (fin.last n) = none theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph_parentLaw","label":"graph_parentLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph_parentLaw","description":"theorem graph_parentLaw (a : Option (Fin N × Bool)) : (p.graph reward).parentLaw (graphAction a) (Fin.last N) = (p.rootLaw a).map ((p.graph reward).parentConfig (Fin.last N))","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-50f691395db3","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":621,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem graph_parentLaw (a : Option (Fin N × Bool)) : (p.graph reward).parentLaw (graphAction a) (Fin.last N) = (p.rootLaw a).map ((p.graph reward).parentConfig (Fin.last N))","missing":[],"search":"graph_parentlaw banditrlproof.causal.parallelparameters.graph_parentlaw theorem graph_parentlaw (a : option (fin n × bool)) : (p.graph reward).parentlaw (graphaction a) (fin.last n) = (p.rootlaw a).map ((p.graph reward).parentconfig (fin.last n)) theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph_parentConfig_injective","label":"graph_parentConfig_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph_parentConfig_injective","description":"theorem graph_parentConfig_injective : Function.Injective ((p.graph reward).parentConfig (Fin.last N))","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-1a133ce1c5f0","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":622,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem graph_parentConfig_injective : Function.Injective ((p.graph reward).parentConfig (Fin.last N))","missing":[],"search":"graph_parentconfig_injective banditrlproof.causal.parallelparameters.graph_parentconfig_injective theorem graph_parentconfig_injective : function.injective ((p.graph reward).parentconfig (fin.last n)) theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph_allocation_covers","label":"graph_allocation_covers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph_allocation_covers","description":"theorem graph_allocation_covers : Covers (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)) (mixture p.allocation (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)))","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-03e3efca0627","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":623,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem graph_allocation_covers : Covers (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)) (mixture p.allocation (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)))","missing":[],"search":"graph_allocation_covers banditrlproof.causal.parallelparameters.graph_allocation_covers theorem graph_allocation_covers : covers (fun a => (p.graph reward).parentlaw (graphaction a) (fin.last n)) (mixture p.allocation (fun a => (p.graph reward).parentlaw (graphaction a) (fin.last n))) theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph_designCost","label":"graph_designCost","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph_designCost","description":"theorem graph_designCost (eta : PMF (Option (Fin N × Bool))) : designCost (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)) eta = designCost p.rootLaw eta","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-fe7447c9f03b","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":624,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem graph_designCost (eta : PMF (Option (Fin N × Bool))) : designCost (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)) eta = designCost p.rootLaw eta","missing":[],"search":"graph_designcost banditrlproof.causal.parallelparameters.graph_designcost theorem graph_designcost (eta : pmf (option (fin n × bool))) : designcost (fun a => (p.graph reward).parentlaw (graphaction a) (fin.last n)) eta = designcost p.rootlaw eta theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph_allocation_cost_le","label":"graph_allocation_cost_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph_allocation_cost_le","description":"theorem graph_allocation_cost_le : designCost (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)) p.allocation ≤ 2*p.rarity","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-d878c521e139","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":625,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem graph_allocation_cost_le : designCost (fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N)) p.allocation ≤ 2*p.rarity","missing":[],"search":"graph_allocation_cost_le banditrlproof.causal.parallelparameters.graph_allocation_cost_le theorem graph_allocation_cost_le : designcost (fun a => (p.graph reward).parentlaw (graphaction a) (fin.last n)) p.allocation ≤ 2*p.rarity theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.graph_optimal_cost_le","label":"graph_optimal_cost_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.graph_optimal_cost_le","description":"theorem graph_optimal_cost_le : let laws := fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N) designCost laws (optimalAllocation laws) ≤ 2*p.rarity","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-6a657ddf4b31","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":626,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem graph_optimal_cost_le : let laws := fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N) designCost laws (optimalAllocation laws) ≤ 2*p.rarity","missing":[],"search":"graph_optimal_cost_le banditrlproof.causal.parallelparameters.graph_optimal_cost_le theorem graph_optimal_cost_le : let laws := fun a => (p.graph reward).parentlaw (graphaction a) (fin.last n) designcost laws (optimalallocation laws) ≤ 2*p.rarity theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel","label":"expected_simpleRegret_parallel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel","description":"theorem expected_simpleRegret_parallel (T : ℕ) (hT : 0 < T) : let m := designCost p.rootLaw p.allocation let L := sourceLog T (Fintype.card (Option (Fin N × Bool))) let B := sourceThreshold m T L (∫ w, (p.graph reward).simpleRegret id graphAction p.allocation (Fin.last N) B w ∂(p.graph reward).sampleLaw graphAction p.allocation T) ≤ (2*Real.sqrt 2+7)*Real.sqrt ((2*p.rarity)*L/T)+1/(T:ℝ)","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-9da98aa43fdb","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":627,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem expected_simpleRegret_parallel (T : ℕ) (hT : 0 < T) : let m := designCost p.rootLaw p.allocation let L := sourceLog T (Fintype.card (Option (Fin N × Bool))) let B := sourceThreshold m T L (∫ w, (p.graph reward).simpleRegret id graphAction p.allocation (Fin.last N) B w ∂(p.graph reward).sampleLaw graphAction p.allocation T) ≤ (2*Real.sqrt 2+7)*Real.sqrt ((2*p.rarity)*L/T)+1/(T:ℝ)","missing":[],"search":"expected_simpleregret_parallel banditrlproof.causal.parallelparameters.expected_simpleregret_parallel theorem expected_simpleregret_parallel (t : ℕ) (ht : 0 < t) : let m := designcost p.rootlaw p.allocation let l := sourcelog t (fintype.card (option (fin n × bool))) let b := sourcethreshold m t l (∫ w, (p.graph reward).simpleregret id graphaction p.allocation (fin.last n) b w ∂(p.graph reward).samplelaw graphaction p.allocation t) ≤ (2*real.sqrt 2+7)*real.sqrt ((2*p.rarity)*l/t)+1/(t:ℝ) theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel_optimal","label":"expected_simpleRegret_parallel_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel_optimal","description":"theorem expected_simpleRegret_parallel_optimal (T : ℕ) (hT : 0 < T) : let laws := fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N) let eta := optimalAllocation laws let m := designCost laws eta let L := sourceLog T (Fintype.card (Option (Fin N × Bool))) let B := sourceThreshold m T L (∫ w, (p.graph reward).simpleRegret id graphAction eta (Fin.last N) B w ∂(p.graph reward).sampleLaw graphAction eta T)…","url":"../modules/banditrlproof-algorithms-causalparallelregret/index.html#decl-47220ce503c2","parent":"module:BanditRLProof.Algorithms.CausalParallelRegret","order":628,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalParallelRegret"],["Source","BanditRLProof/Algorithms/CausalParallelRegret.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem expected_simpleRegret_parallel_optimal (T : ℕ) (hT : 0 < T) : let laws := fun a => (p.graph reward).parentLaw (graphAction a) (Fin.last N) let eta := optimalAllocation laws let m := designCost laws eta let L := sourceLog T (Fintype.card (Option (Fin N × Bool))) let B := sourceThreshold m T L (∫ w, (p.graph reward).simpleRegret id graphAction eta (Fin.last N) B w ∂(p.graph reward).sampleLaw graphAction eta T) ≤ (2*Real.sqrt 2+7)*Real.sqrt ((2*p.rarity)*L/T)+1/(T:ℝ)","missing":[],"search":"expected_simpleregret_parallel_optimal banditrlproof.causal.parallelparameters.expected_simpleregret_parallel_optimal theorem expected_simpleregret_parallel_optimal (t : ℕ) (ht : 0 < t) : let laws := fun a => (p.graph reward).parentlaw (graphaction a) (fin.last n) let eta := optimalallocation laws let m := designcost laws eta let l := sourcelog t (fintype.card (option (fin n × bool))) let b := sourcethreshold m t l (∫ w, (p.graph reward).simpleregret id graphaction eta (fin.last n) b w ∂(p.graph reward).samplelaw graphaction eta t) ≤ (2*real.sqrt 2+7)*real.sqrt ((2*p.rarity)*l/t)+1/(t:ℝ) theorem compiled","shard":"modules/e0ecb1b3c65f64c3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.maximizingActions","label":"maximizingActions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.maximizingActions","description":"noncomputable def maximizingActions (score : A → ℝ) : Finset A","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-1392f728b15f","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":629,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def maximizingActions (score : A → ℝ) : Finset A","missing":[],"search":"maximizingactions banditrlproof.causal.maximizingactions noncomputable def maximizingactions (score : a → ℝ) : finset a definition compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.maximizingActions_nonempty","label":"maximizingActions_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.maximizingActions_nonempty","description":"theorem maximizingActions_nonempty (score : A → ℝ) : (maximizingActions score).Nonempty","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-1c3a7d57b8d3","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":630,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem maximizingActions_nonempty (score : A → ℝ) : (maximizingActions score).Nonempty","missing":[],"search":"maximizingactions_nonempty banditrlproof.causal.maximizingactions_nonempty theorem maximizingactions_nonempty (score : a → ℝ) : (maximizingactions score).nonempty theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.orderedArgmax","label":"orderedArgmax","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.orderedArgmax","description":"noncomputable def orderedArgmax (score : A → ℝ) : A","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-260b79a9c545","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":631,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def orderedArgmax (score : A → ℝ) : A","missing":[],"search":"orderedargmax banditrlproof.causal.orderedargmax noncomputable def orderedargmax (score : a → ℝ) : a definition compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.score_le_orderedArgmax","label":"score_le_orderedArgmax","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.score_le_orderedArgmax","description":"theorem score_le_orderedArgmax (score : A → ℝ) (a : A) : score a ≤ score (orderedArgmax score)","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-c11422bb7b47","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":632,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem score_le_orderedArgmax (score : A → ℝ) (a : A) : score a ≤ score (orderedArgmax score)","missing":[],"search":"score_le_orderedargmax banditrlproof.causal.score_le_orderedargmax theorem score_le_orderedargmax (score : a → ℝ) (a : a) : score a ≤ score (orderedargmax score) theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.orderedArgmax_le_of_maximal","label":"orderedArgmax_le_of_maximal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.orderedArgmax_le_of_maximal","description":"theorem orderedArgmax_le_of_maximal (score : A → ℝ) (a : A) (ha : ∀ b, score b ≤ score a) : orderedArgmax score ≤ a","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-f060a1a0d210","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":633,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem orderedArgmax_le_of_maximal (score : A → ℝ) (a : A) (ha : ∀ b, score b ≤ score a) : orderedArgmax score ≤ a","missing":[],"search":"orderedargmax_le_of_maximal banditrlproof.causal.orderedargmax_le_of_maximal theorem orderedargmax_le_of_maximal (score : a → ℝ) (a : a) (ha : ∀ b, score b ≤ score a) : orderedargmax score ≤ a theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.rewardMean","label":"rewardMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.rewardMean","description":"noncomputable def GraphModel.rewardMean (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (a : A) : ℝ","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-7a084e86398f","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":634,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def GraphModel.rewardMean (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (a : A) : ℝ","missing":[],"search":"rewardmean banditrlproof.causal.graphmodel.rewardmean noncomputable def graphmodel.rewardmean (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (i : fin n) (a : a) : ℝ definition compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.rewardMean_eq_integral","label":"rewardMean_eq_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.rewardMean_eq_integral","description":"theorem GraphModel.rewardMean_eq_integral (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (a : A) (hi : actions a i = none) : g.rewardMean rewardBit actions i a = ∫ x, (if rewardBit (x i) then (1:ℝ) else 0) ∂(joint (g.doModel (actions a)).table).toMeasure","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-cf36241fb041","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":635,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"theorem GraphModel.rewardMean_eq_integral (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (a : A) (hi : actions a i = none) : g.rewardMean rewardBit actions i a = ∫ x, (if rewardBit (x i) then (1:ℝ) else 0) ∂(joint (g.doModel (actions a)).table).toMeasure","missing":[],"search":"rewardmean_eq_integral banditrlproof.causal.graphmodel.rewardmean_eq_integral theorem graphmodel.rewardmean_eq_integral (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (i : fin n) (a : a) (hi : actions a i = none) : g.rewardmean rewardbit actions i a = ∫ x, (if rewardbit (x i) then (1:ℝ) else 0) ∂(joint (g.domodel (actions a)).table).tomeasure theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.rewardMean_bounds","label":"rewardMean_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.rewardMean_bounds","description":"theorem GraphModel.rewardMean_bounds (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (a : A) : 0 ≤ g.rewardMean rewardBit actions i a ∧ g.rewardMean rewardBit actions i a ≤ 1","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-d985e9d9bbd3","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":636,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.rewardMean_bounds (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (i : Fin n) (a : A) : 0 ≤ g.rewardMean rewardBit actions i a ∧ g.rewardMean rewardBit actions i a ≤ 1","missing":[],"search":"rewardmean_bounds banditrlproof.causal.graphmodel.rewardmean_bounds theorem graphmodel.rewardmean_bounds (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (i : fin n) (a : a) : 0 ≤ g.rewardmean rewardbit actions i a ∧ g.rewardmean rewardbit actions i a ≤ 1 theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleRecommendation","label":"sampleRecommendation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleRecommendation","description":"noncomputable def GraphModel.sampleRecommendation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : A","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-0953f8a60a15","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":637,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def GraphModel.sampleRecommendation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : A","missing":[],"search":"samplerecommendation banditrlproof.causal.graphmodel.samplerecommendation noncomputable def graphmodel.samplerecommendation (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (b : ℝ) {t : ℕ} (w : fin t → a × (fin n → v)) : a definition compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.simpleRegret","label":"simpleRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.simpleRegret","description":"noncomputable def GraphModel.simpleRegret (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : ℝ","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-d64253e09612","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":638,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def GraphModel.simpleRegret (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : ℝ","missing":[],"search":"simpleregret banditrlproof.causal.graphmodel.simpleregret noncomputable def graphmodel.simpleregret (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (b : ℝ) {t : ℕ} (w : fin t → a × (fin n → v)) : ℝ definition compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.simpleRegret_bounds","label":"simpleRegret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.simpleRegret_bounds","description":"theorem GraphModel.simpleRegret_bounds (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : 0 ≤ g.simpleRegret rewardBit actions eta i B w ∧ g.simpleRegret rewardBit actions eta i B w ≤ 1","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-a3b466c1ca0c","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":639,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.simpleRegret_bounds (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (B : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) : 0 ≤ g.simpleRegret rewardBit actions eta i B w ∧ g.simpleRegret rewardBit actions eta i B w ≤ 1","missing":[],"search":"simpleregret_bounds banditrlproof.causal.graphmodel.simpleregret_bounds theorem graphmodel.simpleregret_bounds (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (b : ℝ) {t : ℕ} (w : fin t → a × (fin n → v)) : 0 ≤ g.simpleregret rewardbit actions eta i b w ∧ g.simpleregret rewardbit actions eta i b w ≤ 1 theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.simpleRegret_le_on_confidence","label":"simpleRegret_le_on_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.simpleRegret_le_on_confidence","description":"theorem GraphModel.simpleRegret_le_on_confidence (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (B : ℝ) (hB : 0 < B) (epsilon : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) (hw : ∀ a, |g.sampleEstimate rewardBit actions eta i a B w-g.estimateCenter rewardBit action…","url":"../modules/banditrlproof-algorithms-causalrecommendation/index.html#decl-26546ad78ee0","parent":"module:BanditRLProof.Algorithms.CausalRecommendation","order":640,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalRecommendation"],["Source","BanditRLProof/Algorithms/CausalRecommendation.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.simpleRegret_le_on_confidence (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (B : ℝ) (hB : 0 < B) (epsilon : ℝ) {T : ℕ} (w : Fin T → A × (Fin n → V)) (hw : ∀ a, |g.sampleEstimate rewardBit actions eta i a B w-g.estimateCenter rewardBit actions eta i a B| ≤ epsilon) : g.simpleRegret rewardBit actions eta i B w ≤ 2*epsilon + designCost (fun a => g.parentLaw (actions a) i) eta / B","missing":[],"search":"simpleregret_le_on_confidence banditrlproof.causal.graphmodel.simpleregret_le_on_confidence theorem graphmodel.simpleregret_le_on_confidence (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (b : ℝ) (hb : 0 < b) (epsilon : ℝ) {t : ℕ} (w : fin t → a × (fin n → v)) (hw : ∀ a, |g.sampleestimate rewardbit actions eta i a b w-g.estimatecenter rewardbit actions eta i a b| ≤ epsilon) : g.simpleregret rewardbit actions eta i b w ≤ 2*epsilon + designcost (fun a => g.parentlaw (actions a) i) eta / b theorem compiled","shard":"modules/d51c6dd1d7db3a28.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_mgf","label":"sampleWeightedBit_signed_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_mgf","description":"theorem GraphModel.sampleWeightedBit_signed_mgf (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (hB : 0 ≤ B) (T : ℕ) (t : Fin T) (sign tilt : ℝ) (hs : |sign| = 1) (hsmall : |tilt| * (2*B) ≤ 1) : Concentration.HasMGF…","url":"../modules/banditrlproof-algorithms-causalsamplemgf/index.html#decl-8a33a5e8ef26","parent":"module:BanditRLProof.Algorithms.CausalSampleMGF","order":641,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampleMGF"],["Source","BanditRLProof/Algorithms/CausalSampleMGF.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleWeightedBit_signed_mgf (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (hB : 0 ≤ B) (T : ℕ) (t : Fin T) (sign tilt : ℝ) (hs : |sign| = 1) (hsmall : |tilt| * (2*B) ≤ 1) : Concentration.HasMGFUpperBoundAt (fun w => sign * (g.sampleWeightedBit rewardBit actions eta i a B t w - truncatedMean (g.parentLaw (actions a) i) (mixture eta (fun b => g.parentLaw (actions b) i)) (fun z => mass ((g.parentTable i z).map rewardBit) true) B)) tilt (tilt^2 * designCost (fun b => g.parentLaw (actions b) i) eta) (g.sampleLaw actions eta T)","missing":[],"search":"sampleweightedbit_signed_mgf banditrlproof.causal.graphmodel.sampleweightedbit_signed_mgf theorem graphmodel.sampleweightedbit_signed_mgf (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (a : a) (b : ℝ) (hb : 0 ≤ b) (t : ℕ) (t : fin t) (sign tilt : ℝ) (hs : |sign| = 1) (hsmall : |tilt| * (2*b) ≤ 1) : concentration.hasmgfupperboundat (fun w => sign * (g.sampleweightedbit rewardbit actions eta i a b t w - truncatedmean (g.parentlaw (actions a) i) (mixture eta (fun b => g.parentlaw (actions b) i)) (fun z => mass ((g.parenttable i z).map rewardbit) true) b)) tilt (tilt^2 * designcost (fun b => g.parentlaw (actions b) i) eta) (g.samplelaw actions eta t) theorem compiled","shard":"modules/b5ecf0464d8d2ccd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_sum_mgf","label":"sampleWeightedBit_signed_sum_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_sum_mgf","description":"theorem GraphModel.sampleWeightedBit_signed_sum_mgf (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (hB : 0 ≤ B) (T : ℕ) (sign tilt : ℝ) (hs : |sign| = 1) (hsmall : |tilt| * (2*B) ≤ 1) : Concentration.HasMGFUpperBou…","url":"../modules/banditrlproof-algorithms-causalsamplemgf/index.html#decl-a76dbeec5f13","parent":"module:BanditRLProof.Algorithms.CausalSampleMGF","order":642,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampleMGF"],["Source","BanditRLProof/Algorithms/CausalSampleMGF.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleWeightedBit_signed_sum_mgf (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (hB : 0 ≤ B) (T : ℕ) (sign tilt : ℝ) (hs : |sign| = 1) (hsmall : |tilt| * (2*B) ≤ 1) : Concentration.HasMGFUpperBoundAt (fun w => ∑ t : Fin T, sign * (g.sampleWeightedBit rewardBit actions eta i a B t w - truncatedMean (g.parentLaw (actions a) i) (mixture eta (fun b => g.parentLaw (actions b) i)) (fun z => mass ((g.parentTable i z).map rewardBit) true) B)) tilt ((T:ℝ) * (tilt^2 * designCost (fun b => g.parentLaw (actions b) i) eta)) (g.sampleLaw actions eta T)","missing":[],"search":"sampleweightedbit_signed_sum_mgf banditrlproof.causal.graphmodel.sampleweightedbit_signed_sum_mgf theorem graphmodel.sampleweightedbit_signed_sum_mgf (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (a : a) (b : ℝ) (hb : 0 ≤ b) (t : ℕ) (sign tilt : ℝ) (hs : |sign| = 1) (hsmall : |tilt| * (2*b) ≤ 1) : concentration.hasmgfupperboundat (fun w => ∑ t : fin t, sign * (g.sampleweightedbit rewardbit actions eta i a b t w - truncatedmean (g.parentlaw (actions a) i) (mixture eta (fun b => g.parentlaw (actions b) i)) (fun z => mass ((g.parenttable i z).map rewardbit) true) b)) tilt ((t:ℝ) * (tilt^2 * designcost (fun b => g.parentlaw (actions b) i) eta)) (g.samplelaw actions eta t) theorem compiled","shard":"modules/b5ecf0464d8d2ccd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_sum_tail","label":"sampleWeightedBit_signed_sum_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_sum_tail","description":"theorem GraphModel.sampleWeightedBit_signed_sum_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (hB : 0 ≤ B) (T : ℕ) (sign tilt r : ℝ) (hs : |sign| = 1) (ht : 0 ≤ tilt) (hsmall : |tilt| * (2*B) ≤ 1) : (g.sample…","url":"../modules/banditrlproof-algorithms-causalsamplemgf/index.html#decl-933b40c2d31e","parent":"module:BanditRLProof.Algorithms.CausalSampleMGF","order":643,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampleMGF"],["Source","BanditRLProof/Algorithms/CausalSampleMGF.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleWeightedBit_signed_sum_tail (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (hB : 0 ≤ B) (T : ℕ) (sign tilt r : ℝ) (hs : |sign| = 1) (ht : 0 ≤ tilt) (hsmall : |tilt| * (2*B) ≤ 1) : (g.sampleLaw actions eta T).real {w | r ≤ ∑ t : Fin T, sign * (g.sampleWeightedBit rewardBit actions eta i a B t w - truncatedMean (g.parentLaw (actions a) i) (mixture eta (fun b => g.parentLaw (actions b) i)) (fun z => mass ((g.parentTable i z).map rewardBit) true) B)} ≤ Real.exp (-tilt*r + (T:ℝ) * (tilt^2 * designCost (fun b => g.parentLaw (actions b) i) eta))","missing":[],"search":"sampleweightedbit_signed_sum_tail banditrlproof.causal.graphmodel.sampleweightedbit_signed_sum_tail theorem graphmodel.sampleweightedbit_signed_sum_tail (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (a : a) (b : ℝ) (hb : 0 ≤ b) (t : ℕ) (sign tilt r : ℝ) (hs : |sign| = 1) (ht : 0 ≤ tilt) (hsmall : |tilt| * (2*b) ≤ 1) : (g.samplelaw actions eta t).real {w | r ≤ ∑ t : fin t, sign * (g.sampleweightedbit rewardbit actions eta i a b t w - truncatedmean (g.parentlaw (actions a) i) (mixture eta (fun b => g.parentlaw (actions b) i)) (fun z => mass ((g.parenttable i z).map rewardbit) true) b)} ≤ real.exp (-tilt*r + (t:ℝ) * (tilt^2 * designcost (fun b => g.parentlaw (actions b) i) eta)) theorem compiled","shard":"modules/b5ecf0464d8d2ccd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.roundLaw","label":"roundLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.roundLaw","description":"noncomputable def GraphModel.roundLaw (g : GraphModel V n) (actions : A → Fin n → Option V) (eta : PMF A) : PMF (A × (Fin n → V))","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-17efe2fd19c8","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":644,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def GraphModel.roundLaw (g : GraphModel V n) (actions : A → Fin n → Option V) (eta : PMF A) : PMF (A × (Fin n → V))","missing":[],"search":"roundlaw banditrlproof.causal.graphmodel.roundlaw noncomputable def graphmodel.roundlaw (g : graphmodel v n) (actions : a → fin n → option v) (eta : pmf a) : pmf (a × (fin n → v)) definition compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.observation","label":"observation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.observation","description":"def GraphModel.observation (g : GraphModel V n) (rewardBit : V → Bool) (i : Fin n) (ax : A × (Fin n → V)) : g.ParentConfig i × Bool","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-7efd013f0afc","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":645,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def GraphModel.observation (g : GraphModel V n) (rewardBit : V → Bool) (i : Fin n) (ax : A × (Fin n → V)) : g.ParentConfig i × Bool","missing":[],"search":"observation banditrlproof.causal.graphmodel.observation def graphmodel.observation (g : graphmodel v n) (rewardbit : v → bool) (i : fin n) (ax : a × (fin n → v)) : g.parentconfig i × bool definition compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.roundLaw_observation","label":"roundLaw_observation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.roundLaw_observation","description":"theorem GraphModel.roundLaw_observation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) : (g.roundLaw actions eta).map (g.observation rewardBit i) = pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (fun z => (g.parentTable i z).map rewardBit)","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-4c2f7cfb70ea","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":646,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.roundLaw_observation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) : (g.roundLaw actions eta).map (g.observation rewardBit i) = pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (fun z => (g.parentTable i z).map rewardBit)","missing":[],"search":"roundlaw_observation banditrlproof.causal.graphmodel.roundlaw_observation theorem graphmodel.roundlaw_observation (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) : (g.roundlaw actions eta).map (g.observation rewardbit i) = pairedlaw (mixture eta (fun a => g.parentlaw (actions a) i)) (fun z => (g.parenttable i z).map rewardbit) theorem compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleLaw","label":"sampleLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleLaw","description":"noncomputable def GraphModel.sampleLaw (g : GraphModel V n) (actions : A → Fin n → Option V) (eta : PMF A) (T : ℕ) : Measure (Fin T → A × (Fin n → V))","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-a187954e901d","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":647,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def GraphModel.sampleLaw (g : GraphModel V n) (actions : A → Fin n → Option V) (eta : PMF A) (T : ℕ) : Measure (Fin T → A × (Fin n → V))","missing":[],"search":"samplelaw banditrlproof.causal.graphmodel.samplelaw noncomputable def graphmodel.samplelaw (g : graphmodel v n) (actions : a → fin n → option v) (eta : pmf a) (t : ℕ) : measure (fin t → a × (fin n → v)) definition compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleLaw_observation","label":"sampleLaw_observation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleLaw_observation","description":"theorem GraphModel.sampleLaw_observation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (T : ℕ) (i : Fin n) (hi : ∀ a, actions a i = none) (t : Fin T) : (g.sampleLaw actions eta T).map (fun w => g.observation rewardBit i (w t)) = (pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (fun z => (g.parentTable i z).map rewardBit)).toMeasure","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-8898a80084b3","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":648,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleLaw_observation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (T : ℕ) (i : Fin n) (hi : ∀ a, actions a i = none) (t : Fin T) : (g.sampleLaw actions eta T).map (fun w => g.observation rewardBit i (w t)) = (pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (fun z => (g.parentTable i z).map rewardBit)).toMeasure","missing":[],"search":"samplelaw_observation banditrlproof.causal.graphmodel.samplelaw_observation theorem graphmodel.samplelaw_observation (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (t : ℕ) (i : fin n) (hi : ∀ a, actions a i = none) (t : fin t) : (g.samplelaw actions eta t).map (fun w => g.observation rewardbit i (w t)) = (pairedlaw (mixture eta (fun a => g.parentlaw (actions a) i)) (fun z => (g.parenttable i z).map rewardbit)).tomeasure theorem compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.integral_sample_observation","label":"integral_sample_observation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.integral_sample_observation","description":"theorem GraphModel.integral_sample_observation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (T : ℕ) (i : Fin n) (hi : ∀ a, actions a i = none) (t : Fin T) (f : g.ParentConfig i × Bool → ℝ) : (∫ w, f (g.observation rewardBit i (w t)) ∂g.sampleLaw actions eta T) = ∫ zy, f zy ∂(pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (fun z => (g.parentTable i z).map re…","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-1fd518804b81","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":649,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.integral_sample_observation (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (T : ℕ) (i : Fin n) (hi : ∀ a, actions a i = none) (t : Fin T) (f : g.ParentConfig i × Bool → ℝ) : (∫ w, f (g.observation rewardBit i (w t)) ∂g.sampleLaw actions eta T) = ∫ zy, f zy ∂(pairedLaw (mixture eta (fun a => g.parentLaw (actions a) i)) (fun z => (g.parentTable i z).map rewardBit)).toMeasure","missing":[],"search":"integral_sample_observation banditrlproof.causal.graphmodel.integral_sample_observation theorem graphmodel.integral_sample_observation (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (t : ℕ) (i : fin n) (hi : ∀ a, actions a i = none) (t : fin t) (f : g.parentconfig i × bool → ℝ) : (∫ w, f (g.observation rewardbit i (w t)) ∂g.samplelaw actions eta t) = ∫ zy, f zy ∂(pairedlaw (mixture eta (fun a => g.parentlaw (actions a) i)) (fun z => (g.parenttable i z).map rewardbit)).tomeasure theorem compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit","label":"sampleWeightedBit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit","description":"noncomputable def GraphModel.sampleWeightedBit (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) {T : ℕ} (t : Fin T) (w : Fin T → A × (Fin n → V)) : ℝ","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-53cd35af6338","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":650,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","causal"]],"statement":"noncomputable def GraphModel.sampleWeightedBit (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) {T : ℕ} (t : Fin T) (w : Fin T → A × (Fin n → V)) : ℝ","missing":[],"search":"sampleweightedbit banditrlproof.causal.graphmodel.sampleweightedbit noncomputable def graphmodel.sampleweightedbit (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (a : a) (b : ℝ) {t : ℕ} (t : fin t) (w : fin t → a × (fin n → v)) : ℝ definition compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["causal"]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_independent","label":"sampleWeightedBit_independent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit_independent","description":"theorem GraphModel.sampleWeightedBit_independent (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) (T : ℕ) : iIndepFun (g.sampleWeightedBit rewardBit actions eta i a B (T := T)) (g.sampleLaw actions eta T)","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-a73469cfccc7","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":651,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleWeightedBit_independent (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (a : A) (B : ℝ) (T : ℕ) : iIndepFun (g.sampleWeightedBit rewardBit actions eta i a B (T := T)) (g.sampleLaw actions eta T)","missing":[],"search":"sampleweightedbit_independent banditrlproof.causal.graphmodel.sampleweightedbit_independent theorem graphmodel.sampleweightedbit_independent (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (a : a) (b : ℝ) (t : ℕ) : iindepfun (g.sampleweightedbit rewardbit actions eta i a b (t := t)) (g.samplelaw actions eta t) theorem compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_mean","label":"sampleWeightedBit_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit_mean","description":"theorem GraphModel.sampleWeightedBit_mean (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (a : A) (B : ℝ) (T : ℕ) (t : Fin T) : (∫ w, g.sampleWeightedBit rewardBit actions eta i a B t w ∂g.sampleLaw actions eta T) = truncatedMean (g.parentLaw (actions a) i) (mixture eta (fun b => g.parentLaw (actions b) i)) (fun z => mass ((g.paren…","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-293d9a118bf9","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":652,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleWeightedBit_mean (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (a : A) (B : ℝ) (T : ℕ) (t : Fin T) : (∫ w, g.sampleWeightedBit rewardBit actions eta i a B t w ∂g.sampleLaw actions eta T) = truncatedMean (g.parentLaw (actions a) i) (mixture eta (fun b => g.parentLaw (actions b) i)) (fun z => mass ((g.parentTable i z).map rewardBit) true) B","missing":[],"search":"sampleweightedbit_mean banditrlproof.causal.graphmodel.sampleweightedbit_mean theorem graphmodel.sampleweightedbit_mean (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (a : a) (b : ℝ) (t : ℕ) (t : fin t) : (∫ w, g.sampleweightedbit rewardbit actions eta i a b t w ∂g.samplelaw actions eta t) = truncatedmean (g.parentlaw (actions a) i) (mixture eta (fun b => g.parentlaw (actions b) i)) (fun z => mass ((g.parenttable i z).map rewardbit) true) b theorem compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_second_le","label":"sampleWeightedBit_second_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.GraphModel.sampleWeightedBit_second_le","description":"theorem GraphModel.sampleWeightedBit_second_le (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (T : ℕ) (t : Fin T) : (∫ w, (g.sampleWeightedBit rewardBit actions eta i a B t w)^2 ∂g.sampleLaw actions eta T) ≤ second…","url":"../modules/banditrlproof-algorithms-causalsampling/index.html#decl-d4dedbf3d33c","parent":"module:BanditRLProof.Algorithms.CausalSampling","order":653,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalSampling"],["Source","BanditRLProof/Algorithms/CausalSampling.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem GraphModel.sampleWeightedBit_second_le (g : GraphModel V n) (rewardBit : V → Bool) (actions : A → Fin n → Option V) (eta : PMF A) (i : Fin n) (hi : ∀ a, actions a i = none) (hc : Covers (fun a => g.parentLaw (actions a) i) (mixture eta (fun a => g.parentLaw (actions a) i))) (a : A) (B : ℝ) (T : ℕ) (t : Fin T) : (∫ w, (g.sampleWeightedBit rewardBit actions eta i a B t w)^2 ∂g.sampleLaw actions eta T) ≤ secondMoment (g.parentLaw (actions a) i) (mixture eta (fun b => g.parentLaw (actions b) i))","missing":[],"search":"sampleweightedbit_second_le banditrlproof.causal.graphmodel.sampleweightedbit_second_le theorem graphmodel.sampleweightedbit_second_le (g : graphmodel v n) (rewardbit : v → bool) (actions : a → fin n → option v) (eta : pmf a) (i : fin n) (hi : ∀ a, actions a i = none) (hc : covers (fun a => g.parentlaw (actions a) i) (mixture eta (fun a => g.parentlaw (actions a) i))) (a : a) (b : ℝ) (t : ℕ) (t : fin t) : (∫ w, (g.sampleweightedbit rewardbit actions eta i a b t w)^2 ∂g.samplelaw actions eta t) ≤ secondmoment (g.parentlaw (actions a) i) (mixture eta (fun b => g.parentlaw (actions b) i)) theorem compiled","shard":"modules/2d6837cf2bd017ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceThreshold","label":"sourceThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.sourceThreshold","description":"noncomputable def sourceThreshold (m T L : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-ff6e5a1d0153","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":654,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceThreshold (m T L : ℝ) : ℝ","missing":[],"search":"sourcethreshold banditrlproof.causal.sourcethreshold noncomputable def sourcethreshold (m t l : ℝ) : ℝ definition compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceRadius","label":"sourceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.sourceRadius","description":"noncomputable def sourceRadius (m T L : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-fdf17073a979","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":655,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceRadius (m T L : ℝ) : ℝ","missing":[],"search":"sourceradius banditrlproof.causal.sourceradius noncomputable def sourceradius (m t l : ℝ) : ℝ definition compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceLog","label":"sourceLog","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Causal.sourceLog","description":"noncomputable def sourceLog (T K : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-2d33859d1d05","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":656,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceLog (T K : ℕ) : ℝ","missing":[],"search":"sourcelog banditrlproof.causal.sourcelog noncomputable def sourcelog (t k : ℕ) : ℝ definition compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceLog_pos","label":"sourceLog_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceLog_pos","description":"theorem sourceLog_pos (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : 0 < sourceLog T K","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-5298d896f5b6","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":657,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceLog_pos (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : 0 < sourceLog T K","missing":[],"search":"sourcelog_pos banditrlproof.causal.sourcelog_pos theorem sourcelog_pos (t k : ℕ) (ht : 0 < t) (hk : 0 < k) : 0 < sourcelog t k theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceLog_union_budget","label":"sourceLog_union_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceLog_union_budget","description":"theorem sourceLog_union_budget (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : (K:ℝ)*(2*Real.exp (-sourceLog T K)) = 1/(T:ℝ)","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-f8f584cc9cd5","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":658,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceLog_union_budget (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : (K:ℝ)*(2*Real.exp (-sourceLog T K)) = 1/(T:ℝ)","missing":[],"search":"sourcelog_union_budget banditrlproof.causal.sourcelog_union_budget theorem sourcelog_union_budget (t k : ℕ) (ht : 0 < t) (hk : 0 < k) : (k:ℝ)*(2*real.exp (-sourcelog t k)) = 1/(t:ℝ) theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceThreshold_pos","label":"sourceThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceThreshold_pos","description":"theorem sourceThreshold_pos (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : 0 < sourceThreshold m T L","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-3de55d13f81d","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":659,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_pos (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : 0 < sourceThreshold m T L","missing":[],"search":"sourcethreshold_pos banditrlproof.causal.sourcethreshold_pos theorem sourcethreshold_pos (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : 0 < sourcethreshold m t l theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceThreshold_sq","label":"sourceThreshold_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceThreshold_sq","description":"theorem sourceThreshold_sq (m T L : ℝ) (hm : 0 ≤ m) (hT : 0 ≤ T) (hL : 0 ≤ L) : (sourceThreshold m T L)^2 = m*T/L","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-64fe33e28e35","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":660,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_sq (m T L : ℝ) (hm : 0 ≤ m) (hT : 0 ≤ T) (hL : 0 ≤ L) : (sourceThreshold m T L)^2 = m*T/L","missing":[],"search":"sourcethreshold_sq banditrlproof.causal.sourcethreshold_sq theorem sourcethreshold_sq (m t l : ℝ) (hm : 0 ≤ m) (ht : 0 ≤ t) (hl : 0 ≤ l) : (sourcethreshold m t l)^2 = m*t/l theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceTilt_admissible","label":"sourceTilt_admissible","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceTilt_admissible","description":"theorem sourceTilt_admissible (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : |1/(2*sourceThreshold m T L)| * (2*sourceThreshold m T L) ≤ 1","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-0ad45591193a","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":661,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceTilt_admissible (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : |1/(2*sourceThreshold m T L)| * (2*sourceThreshold m T L) ≤ 1","missing":[],"search":"sourcetilt_admissible banditrlproof.causal.sourcetilt_admissible theorem sourcetilt_admissible (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : |1/(2*sourcethreshold m t l)| * (2*sourcethreshold m t l) ≤ 1 theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceTilt_budget","label":"sourceTilt_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceTilt_budget","description":"theorem sourceTilt_budget (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : T * ((1/(2*sourceThreshold m T L))^2*m) = L/4","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-b82499d0b8b7","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":662,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceTilt_budget (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : T * ((1/(2*sourceThreshold m T L))^2*m) = L/4","missing":[],"search":"sourcetilt_budget banditrlproof.causal.sourcetilt_budget theorem sourcetilt_budget (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : t * ((1/(2*sourcethreshold m t l))^2*m) = l/4 theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceTilt_exponent_le","label":"sourceTilt_exponent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceTilt_exponent_le","description":"theorem sourceTilt_exponent_le (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : -(1/(2*sourceThreshold m T L)) * (T*sourceRadius m T L) + T * ((1/(2*sourceThreshold m T L))^2*m) ≤ -L","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-95bb8e14d451","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":663,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceTilt_exponent_le (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : -(1/(2*sourceThreshold m T L)) * (T*sourceRadius m T L) + T * ((1/(2*sourceThreshold m T L))^2*m) ≤ -L","missing":[],"search":"sourcetilt_exponent_le banditrlproof.causal.sourcetilt_exponent_le theorem sourcetilt_exponent_le (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : -(1/(2*sourcethreshold m t l)) * (t*sourceradius m t l) + t * ((1/(2*sourcethreshold m t l))^2*m) ≤ -l theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceThreshold_scale","label":"sourceThreshold_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceThreshold_scale","description":"theorem sourceThreshold_scale (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : sourceThreshold m T L * L / T = Real.sqrt (m*L/T)","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-d7f3bf6e887c","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":664,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_scale (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : sourceThreshold m T L * L / T = Real.sqrt (m*L/T)","missing":[],"search":"sourcethreshold_scale banditrlproof.causal.sourcethreshold_scale theorem sourcethreshold_scale (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : sourcethreshold m t l * l / t = real.sqrt (m*l/t) theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceThreshold_bias_scale","label":"sourceThreshold_bias_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceThreshold_bias_scale","description":"theorem sourceThreshold_bias_scale (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : m/sourceThreshold m T L = Real.sqrt (m*L/T)","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-75d46a06ec8a","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":665,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_bias_scale (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : m/sourceThreshold m T L = Real.sqrt (m*L/T)","missing":[],"search":"sourcethreshold_bias_scale banditrlproof.causal.sourcethreshold_bias_scale theorem sourcethreshold_bias_scale (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : m/sourcethreshold m t l = real.sqrt (m*l/t) theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceRegret_scale","label":"sourceRegret_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceRegret_scale","description":"theorem sourceRegret_scale (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : 2*sourceRadius m T L + m/sourceThreshold m T L = (2*Real.sqrt 2+7)*Real.sqrt (m*L/T)","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-92fb731028c8","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":666,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceRegret_scale (m T L : ℝ) (hm : 0 < m) (hT : 0 < T) (hL : 0 < L) : 2*sourceRadius m T L + m/sourceThreshold m T L = (2*Real.sqrt 2+7)*Real.sqrt (m*L/T)","missing":[],"search":"sourceregret_scale banditrlproof.causal.sourceregret_scale theorem sourceregret_scale (m t l : ℝ) (hm : 0 < m) (ht : 0 < t) (hl : 0 < l) : 2*sourceradius m t l + m/sourcethreshold m t l = (2*real.sqrt 2+7)*real.sqrt (m*l/t) theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceLog_half_le","label":"sourceLog_half_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceLog_half_le","description":"theorem sourceLog_half_le (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : (1:ℝ)/2 ≤ sourceLog T K","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-2ca40f825d14","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":667,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceLog_half_le (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : (1:ℝ)/2 ≤ sourceLog T K","missing":[],"search":"sourcelog_half_le banditrlproof.causal.sourcelog_half_le theorem sourcelog_half_le (t k : ℕ) (ht : 0 < t) (hk : 0 < k) : (1:ℝ)/2 ≤ sourcelog t k theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Causal.sourceResidual_le_scale","label":"sourceResidual_le_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Causal.sourceResidual_le_scale","description":"theorem sourceResidual_le_scale (m : ℝ) (hm : 1 ≤ m) (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : 1/(T:ℝ) ≤ Real.sqrt 2 * Real.sqrt (m*sourceLog T K/T)","url":"../modules/banditrlproof-algorithms-causaltuning/index.html#decl-c8d163c594c0","parent":"module:BanditRLProof.Algorithms.CausalTuning","order":668,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.CausalTuning"],["Source","BanditRLProof/Algorithms/CausalTuning.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceResidual_le_scale (m : ℝ) (hm : 1 ≤ m) (T K : ℕ) (hT : 0 < T) (hK : 0 < K) : 1/(T:ℝ) ≤ Real.sqrt 2 * Real.sqrt (m*sourceLog T K/T)","missing":[],"search":"sourceresidual_le_scale banditrlproof.causal.sourceresidual_le_scale theorem sourceresidual_le_scale (m : ℝ) (hm : 1 ≤ m) (t k : ℕ) (ht : 0 < t) (hk : 0 < k) : 1/(t:ℝ) ≤ real.sqrt 2 * real.sqrt (m*sourcelog t k/t) theorem compiled","shard":"modules/25bf2ab2c7d635bc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.Spec","label":"Spec","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.Spec","description":"Parameters for a finite-arm Explore-Then-Commit run.","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-1aa5c7354686","parent":"module:BanditRLProof.Algorithms.ETC","order":669,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:11"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure Spec (K : Nat) where","missing":[],"search":"spec banditrlproof.etc.spec parameters for a finite-arm explore-then-commit run. structure compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.exploreArm","label":"exploreArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.exploreArm","description":"Round-robin exploration arm at time `t`.","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-67e758967579","parent":"module:BanditRLProof.Algorithms.ETC","order":670,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:16"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def exploreArm (spec : Spec K) (t : Nat) : Fin K","missing":[],"search":"explorearm banditrlproof.etc.explorearm round-robin exploration arm at time `t`. definition compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.exploreArm_val","label":"exploreArm_val","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.exploreArm_val","description":"@[simp] theorem exploreArm_val (spec : Spec K) (t : Nat) : (exploreArm spec t).val = t % K","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-38c1d52ac98b","parent":"module:BanditRLProof.Algorithms.ETC","order":671,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:19"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"@[simp] theorem exploreArm_val (spec : Spec K) (t : Nat) : (exploreArm spec t).val = t % K","missing":[],"search":"explorearm_val banditrlproof.etc.explorearm_val @[simp] theorem explorearm_val (spec : spec k) (t : nat) : (explorearm spec t).val = t % k theorem compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.exploreArm_eq_of_mod_eq","label":"exploreArm_eq_of_mod_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.exploreArm_eq_of_mod_eq","description":"theorem exploreArm_eq_of_mod_eq (spec : Spec K) {s t : Nat} (h : s % K = t % K) : exploreArm spec s = exploreArm spec t","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-04674a2e2f0d","parent":"module:BanditRLProof.Algorithms.ETC","order":672,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:22"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem exploreArm_eq_of_mod_eq (spec : Spec K) {s t : Nat} (h : s % K = t % K) : exploreArm spec s = exploreArm spec t","missing":[],"search":"explorearm_eq_of_mod_eq banditrlproof.etc.explorearm_eq_of_mod_eq theorem explorearm_eq_of_mod_eq (spec : spec k) {s t : nat} (h : s % k = t % k) : explorearm spec s = explorearm spec t theorem compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.exploreArm_eq_iff_mod_eq_val","label":"exploreArm_eq_iff_mod_eq_val","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.exploreArm_eq_iff_mod_eq_val","description":"theorem exploreArm_eq_iff_mod_eq_val (spec : Spec K) (t : Nat) (a : Fin K) : exploreArm spec t = a ↔ t % K = a.val","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-ac8656095ef9","parent":"module:BanditRLProof.Algorithms.ETC","order":673,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:28"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem exploreArm_eq_iff_mod_eq_val (spec : Spec K) (t : Nat) (a : Fin K) : exploreArm spec t = a ↔ t % K = a.val","missing":[],"search":"explorearm_eq_iff_mod_eq_val banditrlproof.etc.explorearm_eq_iff_mod_eq_val theorem explorearm_eq_iff_mod_eq_val (spec : spec k) (t : nat) (a : fin k) : explorearm spec t = a ↔ t % k = a.val theorem compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.exploreArm_add_K","label":"exploreArm_add_K","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.exploreArm_add_K","description":"theorem exploreArm_add_K (spec : Spec K) (t : Nat) : exploreArm spec (t + K) = exploreArm spec t","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-a09aafa0f78f","parent":"module:BanditRLProof.Algorithms.ETC","order":674,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:37"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem exploreArm_add_K (spec : Spec K) (t : Nat) : exploreArm spec (t + K) = exploreArm spec t","missing":[],"search":"explorearm_add_k banditrlproof.etc.explorearm_add_k theorem explorearm_add_k (spec : spec k) (t : nat) : explorearm spec (t + k) = explorearm spec t theorem compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.CommitOracle","label":"CommitOracle","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.CommitOracle","description":"Commit-phase selector. A concrete theorem should replace this by argmax.","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-ed86dbd7a3c5","parent":"module:BanditRLProof.Algorithms.ETC","order":675,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:43"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure CommitOracle (K : Nat) where","missing":[],"search":"commitoracle banditrlproof.etc.commitoracle commit-phase selector. a concrete theorem should replace this by argmax. structure compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.obligationNames","label":"obligationNames","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.obligationNames","description":"The proof-DAG leaves usually needed for ETC regret formalization.","url":"../modules/banditrlproof-algorithms-etc/index.html#decl-01146b636aae","parent":"module:BanditRLProof.Algorithms.ETC","order":676,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETC"],["Source","BanditRLProof/Algorithms/ETC.lean:48"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def obligationNames : List String","missing":[],"search":"obligationnames banditrlproof.etc.obligationnames the proof-dag leaves usually needed for etc regret formalization. definition compiled","shard":"modules/f261e9d5d3185843.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.score_le_foldl_select","label":"score_le_foldl_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.score_le_foldl_select","description":"private theorem score_le_foldl_select {K : Nat} (scores : Fin K -> Rat) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha cases h…","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-e68de33153cc","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":677,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:18"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem score_le_foldl_select {K : Nat} (scores : Fin K -> Rat) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha cases ha) (by simp) | arm :: rest => by let select := fun best arm : Fin K => if scores best < scores arm then arm else best let next := select init arm have ih","missing":[],"search":"score_le_foldl_select banditrlproof.etc.score_le_foldl_select private theorem score_le_foldl_select {k : nat} (scores : fin k -> rat) (init : fin k) : forall l : list (fin k), (forall a : fin k, list.mem a l -> scores a <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init) | [] => by exact and.intro (by intro _ ha cases ha) (by simp) | arm :: rest => by let select := fun best arm : fin k => if scores best < scores arm then arm else best let next := select init arm have ih theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmax_cons_eq_some_foldl_rat_select","label":"argmax_cons_eq_some_foldl_rat_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.argmax_cons_eq_some_foldl_rat_select","description":"private theorem argmax_cons_eq_some_foldl_rat_select {K : Nat} (scores : Fin K -> Rat) (init : Fin K) (l : List (Fin K)) : List.argmax scores (init :: l) = some (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-53229257f66f","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":678,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:72"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem argmax_cons_eq_some_foldl_rat_select {K : Nat} (scores : Fin K -> Rat) (init : Fin K) (l : List (Fin K)) : List.argmax scores (init :: l) = some (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)","missing":[],"search":"argmax_cons_eq_some_foldl_rat_select banditrlproof.etc.argmax_cons_eq_some_foldl_rat_select private theorem argmax_cons_eq_some_foldl_rat_select {k : nat} (scores : fin k -> rat) (init : fin k) (l : list (fin k)) : list.argmax scores (init :: l) = some (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init) theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmaxCommitOracle","label":"argmaxCommitOracle","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.argmaxCommitOracle","description":"A concrete ETC commit oracle that selects a score-maximizing arm from `Fin K`. The selector scans `List.finRange K` and keeps the previous arm on ties, giving a deterministic total oracle whenever `hK : 0 < K` supplies the initial arm. This is the compiled `ETC-COMMIT-ORACLE-CONCRETE-ARGMAX` leaf.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-968848846153","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":679,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:100"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def argmaxCommitOracle {K : Nat} (hK : 0 < K) : ETC.CommitOracle K where","missing":[],"search":"argmaxcommitoracle banditrlproof.etc.argmaxcommitoracle a concrete etc commit oracle that selects a score-maximizing arm from `fin k`. the selector scans `list.finrange k` and keeps the previous arm on ties, giving a deterministic total oracle whenever `hk : 0 < k` supplies the initial arm. this is the compiled `etc-commit-oracle-concrete-argmax` leaf. definition compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmaxCommitOracle_argmax_finRange","label":"argmaxCommitOracle_argmax_finRange","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.argmaxCommitOracle_argmax_finRange","description":"The concrete Rat commit oracle is Mathlib's first-occurrence list argmax. Because `List.finRange K` is ordered by the canonical `Fin` encoding, this identity records the implementation-level tie rule used by the generated ETC policy: strict score improvements replace the current arm and equal scores do not.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-2cb4750a14c2","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":680,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:117"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem argmaxCommitOracle_argmax_finRange {K : Nat} (hK : 0 < K) (scores : Fin K -> Rat) : List.argmax scores (List.finRange K) = some ((ETC.argmaxCommitOracle hK).choose scores)","missing":[],"search":"argmaxcommitoracle_argmax_finrange banditrlproof.etc.argmaxcommitoracle_argmax_finrange the concrete rat commit oracle is mathlib's first-occurrence list argmax. because `list.finrange k` is ordered by the canonical `fin` encoding, this identity records the implementation-level tie rule used by the generated etc policy: strict score improvements replace the current arm and equal scores do not. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmaxCommitOracle_encode_le_of_score_le","label":"argmaxCommitOracle_encode_le_of_score_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.argmaxCommitOracle_encode_le_of_score_le","description":"Among arms tying the concrete Rat commit oracle's maximal score, the oracle chooses the least encoded arm. This is the public tie-semantics certificate for the canonical generated ETC route. It is deterministic and introduces no probability or concentration assumption.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-88d661dde186","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":681,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:136"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem argmaxCommitOracle_encode_le_of_score_le {K : Nat} (hK : 0 < K) (scores : Fin K -> Rat) (a : Fin K) (hscore : scores ((ETC.argmaxCommitOracle hK).choose scores) <= scores a) : Encodable.encode ((ETC.argmaxCommitOracle hK).choose scores) <= Encodable.encode a","missing":[],"search":"argmaxcommitoracle_encode_le_of_score_le banditrlproof.etc.argmaxcommitoracle_encode_le_of_score_le among arms tying the concrete rat commit oracle's maximal score, the oracle chooses the least encoded arm. this is the public tie-semantics certificate for the canonical generated etc route. it is deterministic and introduces no probability or concentration assumption. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmaxCommitOracle_choose_spec","label":"argmaxCommitOracle_choose_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.argmaxCommitOracle_choose_spec","description":"The concrete ETC argmax oracle returns an arm whose score dominates every arm. This is the maximality certificate needed by the already compiled abstract commit-oracle wrong-event and probability consumers. It does not introduce measures, empirical-mean construction, concentration, filtration, or final ETC regret.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-7ed4d9d1dbf0","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":682,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:157"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem argmaxCommitOracle_choose_spec {K : Nat} (hK : 0 < K) (scores : Fin K -> Rat) (a : Fin K) : scores a <= scores ((ETC.argmaxCommitOracle hK).choose scores)","missing":[],"search":"argmaxcommitoracle_choose_spec banditrlproof.etc.argmaxcommitoracle_choose_spec the concrete etc argmax oracle returns an arm whose score dominates every arm. this is the maximality certificate needed by the already compiled abstract commit-oracle wrong-event and probability consumers. it does not introduce measures, empirical-mean construction, concentration, filtration, or final etc regret. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmaxCommitOracle_eq_arm_subset_empMean_ge_bestArm","label":"argmaxCommitOracle_eq_arm_subset_empMean_ge_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.argmaxCommitOracle_eq_arm_subset_empMean_ge_bestArm","description":"Selecting a particular arm with the concrete ETC argmax implies that arm's score is at least the selected model best arm's score. This is a deterministic single-fiber refinement of the existing wrong-commit union reduction. It preserves the candidate arm instead of existentially or union bounding it.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-ce9391fc96ff","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":683,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:177"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem argmaxCommitOracle_eq_arm_subset_empMean_ge_bestArm {Omega : Type u} {K : Nat} (hK : 0 < K) (model : FiniteBanditModel K) (empMean : Omega -> Fin K -> Rat) (a : Fin K) : Set.Subset {omega : Omega | (ETC.argmaxCommitOracle hK).choose (empMean omega) = a} {omega : Omega | empMean omega a >= empMean omega model.bestArm}","missing":[],"search":"argmaxcommitoracle_eq_arm_subset_empmean_ge_bestarm banditrlproof.etc.argmaxcommitoracle_eq_arm_subset_empmean_ge_bestarm selecting a particular arm with the concrete etc argmax implies that arm's score is at least the selected model best arm's score. this is a deterministic single-fiber refinement of the existing wrong-commit union reduction. it preserves the candidate arm instead of existentially or union bounding it. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_eq_arm_le_pairwise_tail","label":"prob_argmaxCommitOracle_eq_arm_le_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_eq_arm_le_pairwise_tail","description":"The probability of committing to one concrete arm is bounded by any supplied tail bound for that arm's empirical mean exceeding the model best arm's mean. No union bound is taken. The result uses only measure monotonicity and the deterministic concrete-argmax fiber inclusion above.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-004ff947055f","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":684,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:201"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_eq_arm_le_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) (model : FiniteBanditModel K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (a : Fin K) (hpair_tail : mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (empMean omega) = a} <= tail a","missing":[],"search":"prob_argmaxcommitoracle_eq_arm_le_pairwise_tail banditrlproof.etc.prob_argmaxcommitoracle_eq_arm_le_pairwise_tail the probability of committing to one concrete arm is bounded by any supplied tail bound for that arm's empirical mean exceeding the model best arm's mean. no union bound is taken. the result uses only measure monotonicity and the deterministic concrete-argmax fiber inclusion above. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm_of_argmaxOracle","label":"wrong_commit_subset_exists_empMean_ge_bestArm_of_argmaxOracle","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm_of_argmaxOracle","description":"The concrete argmax oracle instantiates the existing deterministic wrong-commit event reduction. This wrapper records that the new concrete oracle feeds directly into the previous abstract `CommitOracle` consumer. It remains a deterministic set inclusion, with no probability or concentration assumptions.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-93b84c375f37","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":685,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:230"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem wrong_commit_subset_exists_empMean_ge_bestArm_of_argmaxOracle {Omega : Type u} {K : Nat} (hK : 0 < K) (model : FiniteBanditModel K) (empMean : Omega -> Fin K -> Rat) : Set.Subset {omega : Omega | (ETC.argmaxCommitOracle hK).choose (empMean omega) = model.bestArm -> False} {omega : Omega | exists a : Fin K, (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm}","missing":[],"search":"wrong_commit_subset_exists_empmean_ge_bestarm_of_argmaxoracle banditrlproof.etc.wrong_commit_subset_exists_empmean_ge_bestarm_of_argmaxoracle the concrete argmax oracle instantiates the existing deterministic wrong-commit event reduction. this wrapper records that the new concrete oracle feeds directly into the previous abstract `commitoracle` consumer. it remains a deterministic set inclusion, with no probability or concentration assumptions. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","description":"The concrete argmax oracle-selected wrong-commit event is bounded by the filtered finite sum of abstract non-best pairwise tail bounds. This is the `ETC-COMMIT-ORACLE-CONCRETE-FILTERED-SUM-PAIRWISE-TAIL` leaf. It only specializes the already compiled abstract oracle filtered-sum consumer to `ETC.argmaxCommitOracle`; it does not prove pairwise tails, add concentration, introduce filtration, or prove final ETC regret.","url":"../modules/banditrlproof-algorithms-etcargmaxoracle/index.html#decl-98038b72e602","parent":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","order":686,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCArgmaxOracle"],["Source","BanditRLProof/Algorithms/ETCArgmaxOracle.lean:260"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) (model : FiniteBanditModel K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (empMean omega) = model.bestArm -> False} <= ((Finset.univ : Finset (Fin K)).filter (fun a : Fin K => a = model.bestArm -> False)).sum tail","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail the concrete argmax oracle-selected wrong-commit event is bounded by the filtered finite sum of abstract non-best pairwise tail bounds. this is the `etc-commit-oracle-concrete-filtered-sum-pairwise-tail` leaf. it only specializes the already compiled abstract oracle filtered-sum consumer to `etc.argmaxcommitoracle`; it does not prove pairwise tails, add concentration, introduce filtration, or prove final etc regret. theorem compiled","shard":"modules/4479fa60fbcd9782.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_pastReward_iSup_infinitePi","label":"indep_centeredReward_succ_pastReward_iSup_infinitePi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.indep_centeredReward_succ_pastReward_iSup_infinitePi","description":"Under an infinite product reward-coordinate law, the centered reward at time `i + 1` is independent of the reward-only past sigma-algebra generated by coordinates `j <= i`. This is the concrete product-law specialization of `indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward`. It does not add the deterministic action generators from `History.historyFiltrationSucc`; use `indep_centeredReward_succ_historyFi…","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-e6658207208e","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":687,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:28"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem indep_centeredReward_succ_pastReward_iSup_infinitePi {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (model : FiniteBanditModel K) (b : Fin K) (i : Nat) : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : RewardTrace Rat => (((omega (i + 1) - model.mean b : Rat) : Real))) inferInstance) (iSup fun j : Nat => iSup fun _h : j <= i => MeasurableSpace.comap (fun omega : RewardTrace Rat => omega j) inferInstance) (MeasureTheory.Measure.infinitePi coordLaw)","missing":[],"search":"indep_centeredreward_succ_pastreward_isup_infinitepi banditrlproof.etc.indep_centeredreward_succ_pastreward_isup_infinitepi under an infinite product reward-coordinate law, the centered reward at time `i + 1` is independent of the reward-only past sigma-algebra generated by coordinates `j <= i`. this is the concrete product-law specialization of `indep_centeredreward_succ_pastreward_isup_of_iindepfun_reward`. it does not add the deterministic action generators from `history.historyfiltrationsucc`; use `indep_centeredreward_succ_historyfiltrationsucc_infinitepi` for that. theorem compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_historyFiltrationSucc_infinitePi","label":"indep_centeredReward_succ_historyFiltrationSucc_infinitePi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.indep_centeredReward_succ_historyFiltrationSucc_infinitePi","description":"Under an infinite product reward-coordinate law, the centered reward at time `i + 1` is independent of the full fixed-commit shifted history filtration. This specializes the deterministic action-history inclusion and reward-coordinate `iIndepFun` bridge to `Measure.infinitePi`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-8a05746467e4","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":688,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:60"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem indep_centeredReward_succ_historyFiltrationSucc_infinitePi {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (b : Fin K) (i : Nat) : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : RewardTrace Rat => (((omega (i + 1) - model.mean b : Rat) : Real))) inferInstance) (History.historyFiltrationSucc (fun _omega : RewardTrace Rat => ETC.actionWithCommit spec commitArm) (fun omega : RewardTrace Rat => omega) (fun _t : Nat => measurable_const) (fun t : Nat => measurable_pi_apply t) i) (MeasureTheory.Measure.infinitePi coordLaw)","missing":[],"search":"indep_centeredreward_succ_historyfiltrationsucc_infinitepi banditrlproof.etc.indep_centeredreward_succ_historyfiltrationsucc_infinitepi under an infinite product reward-coordinate law, the centered reward at time `i + 1` is independent of the full fixed-commit shifted history filtration. this specializes the deterministic action-history inclusion and reward-coordinate `iindepfun` bridge to `measure.infinitepi`. theorem compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.boundedRewardTraceSource_infinitePi_actionWithCommit","label":"boundedRewardTraceSource_infinitePi_actionWithCommit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.boundedRewardTraceSource_infinitePi_actionWithCommit","description":"The infinite product law over reward coordinates satisfies the action-matched bounded reward source contract when each coordinate law has the corresponding a.s. bound and mean for the arm pulled by the fixed-commit ETC trace.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-20cd9d5959b0","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":689,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:96"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem boundedRewardTraceSource_infinitePi_actionWithCommit {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (lo hi : Fin K -> Nat -> Real) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun r : Rat => Set.Icc (lo (ETC.actionWithCommit spec commitArm t) t) (hi (ETC.actionWithCommit spec commitArm t) t) (((r : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun r : Rat => (((r : Rat) : Real))) = (((model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) : ETC.BoundedRewardTraceSource (MeasureTheory.Measure.infinitePi coordLaw) spec model commitArm (fun omega : RewardTrace Rat => omega) lo hi where","missing":[],"search":"boundedrewardtracesource_infinitepi_actionwithcommit banditrlproof.etc.boundedrewardtracesource_infinitepi_actionwithcommit the infinite product law over reward coordinates satisfies the action-matched bounded reward source contract when each coordinate law has the corresponding a.s. bound and mean for the arm pulled by the fixed-commit etc trace. theorem compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_infinitePi_bounded_actionMean_canonicalTail","label":"centeredRewardCondSubGaussianWitnesses_of_infinitePi_bounded_actionMean_canonicalTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_infinitePi_bounded_actionMean_canonicalTail","description":"Under an infinite product reward-coordinate law, construct the reward-level conditional sub-Gaussian witness package for the fixed `actionWithCommit` route with the canonical centered-diff exponential tail.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-d2f58b23e25e","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":690,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:179"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredRewardCondSubGaussianWitnesses_of_infinitePi_bounded_actionMean_canonicalTail {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (lo hi : Fin K -> Nat -> Real) (horizon_pos : 0 < spec.explorationPulls * K) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun r : Rat => Set.Icc (lo (ETC.actionWithCommit spec commitArm t) t) (hi (ETC.actionWithCommit spec commitArm t) t) (((r : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun r : Rat => (((r : Rat) : Real))) = (((model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) : ETC.CenteredRewardCondSubGaussianWitnesses (MeasureTheory.Measure.i…","missing":[],"search":"centeredrewardcondsubgaussianwitnesses_of_infinitepi_bounded_actionmean_canonicaltail banditrlproof.etc.centeredrewardcondsubgaussianwitnesses_of_infinitepi_bounded_actionmean_canonicaltail under an infinite product reward-coordinate law, construct the reward-level conditional sub-gaussian witness package for the fixed `actionwithcommit` route with the canonical centered-diff exponential tail. definition compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_infinitePi_bounded_actionMean_condSubGaussian_canonicalTail","label":"pairwiseEmpMeanTailContract_of_infinitePi_bounded_actionMean_condSubGaussian_canonicalTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_infinitePi_bounded_actionMean_condSubGaussian_canonicalTail","description":"Under an infinite product reward-coordinate law, produce the fixed-commit pairwise empirical-mean tail contract through the conditional sub-Gaussian route with the canonical centered-diff exponential tail.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-2b75c0959121","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":691,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:223"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_infinitePi_bounded_actionMean_condSubGaussian_canonicalTail {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (lo hi : Fin K -> Nat -> Real) (horizon_pos : 0 < spec.explorationPulls * K) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun r : Rat => Set.Icc (lo (ETC.actionWithCommit spec commitArm t) t) (hi (ETC.actionWithCommit spec commitArm t) t) (((r : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun r : Rat => (((r : Rat) : Real))) = (((model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) : ETC.PairwiseEmpMeanT…","missing":[],"search":"pairwiseempmeantailcontract_of_infinitepi_bounded_actionmean_condsubgaussian_canonicaltail banditrlproof.etc.pairwiseempmeantailcontract_of_infinitepi_bounded_actionmean_condsubgaussian_canonicaltail under an infinite product reward-coordinate law, produce the fixed-commit pairwise empirical-mean tail contract through the conditional sub-gaussian route with the canonical centered-diff exponential tail. theorem compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean","description":"Concrete fixed-commit ETC wrong-commit probability bound under an infinite product reward-coordinate law with action-matched boundedness and means.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-92fbc4d7172b","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":692,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:267"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean {K : Nat} (hK : 0 < K) (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun r : Rat => Set.Icc (lo (ETC.actionWithCommit spec commitArm t) t) (hi (ETC.actionWithCommit spec commitArm t) t) (((r : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun r : Rat => (((r : Rat) : Real))) = (((model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) : MeasureTheory.Measure.infinitePi…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_infinitepi_bounded_actionmean banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_infinitepi_bounded_actionmean concrete fixed-commit etc wrong-commit probability bound under an infinite product reward-coordinate law with action-matched boundedness and means. theorem compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean_condSubGaussian","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean_condSubGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean_condSubGaussian","description":"Concrete fixed-commit ETC wrong-commit probability bound under an infinite product reward-coordinate law, routed through the conditional sub-Gaussian canonical-tail source package.","url":"../modules/banditrlproof-algorithms-etcboundedrewardinfinitepisource/index.html#decl-9c8800b72aff","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","order":693,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardInfinitePiSource.lean:314"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean_condSubGaussian {K : Nat} (hK : 0 < K) (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun r : Rat => Set.Icc (lo (ETC.actionWithCommit spec commitArm t) t) (hi (ETC.actionWithCommit spec commitArm t) t) (((r : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun r : Rat => (((r : Rat) : Real))) = (((model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) : MeasureTheory.Me…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_infinitepi_bounded_actionmean_condsubgaussian banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_infinitepi_bounded_actionmean_condsubgaussian concrete fixed-commit etc wrong-commit probability bound under an infinite product reward-coordinate law, routed through the conditional sub-gaussian canonical-tail source package. theorem compiled","shard":"modules/d5d4f17da7e21a6e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.BoundedRewardTraceSource","label":"BoundedRewardTraceSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.BoundedRewardTraceSource","description":"Source contract for the bounded-reward ETC wrong-commit route. This is intentionally a thin structure: it records trace-level reward-coordinate independence, a.e. coordinate measurability, a.s. interval bounds, and exact mean identities over the ETC exploration horizon.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-1147f58f18a0","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":694,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:23"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure BoundedRewardTraceSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) where","missing":[],"search":"boundedrewardtracesource banditrlproof.etc.boundedrewardtracesource source contract for the bounded-reward etc wrong-commit route. this is intentionally a thin structure: it records trace-level reward-coordinate independence, a.e. coordinate measurability, a.s. interval bounds, and exact mean identities over the etc exploration horizon. structure compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_integrable_of_boundedRewardTraceSource","label":"centeredReward_integrable_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_integrable_of_boundedRewardTraceSource","description":"A bounded reward source supplies raw reward integrability for the action actually pulled at time `t`. This packages the `meas` and `bound` fields into the bounded-to-integrable side condition consumed by the conditional mean-zero route.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-21515060d33c","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":695,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:60"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_integrable_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (t : Nat) (ht : t < spec.explorationPulls * K) : MeasureTheory.Integrable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu","missing":[],"search":"centeredreward_integrable_of_boundedrewardtracesource banditrlproof.etc.centeredreward_integrable_of_boundedrewardtracesource a bounded reward source supplies raw reward integrability for the action actually pulled at time `t`. this packages the `meas` and `bound` fields into the bounded-to-integrable side condition consumed by the conditional mean-zero route. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_integral_eq_zero_of_boundedRewardTraceSource_mean","label":"centeredReward_integral_eq_zero_of_boundedRewardTraceSource_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_integral_eq_zero_of_boundedRewardTraceSource_mean","description":"A bounded reward source supplies the action-matched zero-integral fact for the centered reward coordinate. This packages the `meas`, `bound`, and `mean` fields into the zero-integral side condition consumed by the conditional mean-zero route.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-7608e80fc3dc","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":696,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:86"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_integral_eq_zero_of_boundedRewardTraceSource_mean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (t : Nat) (ht : t < spec.explorationPulls * K) : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) = 0","missing":[],"search":"centeredreward_integral_eq_zero_of_boundedrewardtracesource_mean banditrlproof.etc.centeredreward_integral_eq_zero_of_boundedrewardtracesource_mean a bounded reward source supplies the action-matched zero-integral fact for the centered reward coordinate. this packages the `meas`, `bound`, and `mean` fields into the zero-integral side condition consumed by the conditional mean-zero route. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_centeredReward_actionWithCommit_historyFiltrationSucc","label":"measurable_centeredReward_actionWithCommit_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_centeredReward_actionWithCommit_historyFiltrationSucc","description":"The action-matched centered reward at time `t` is measurable with respect to the shifted generated history filtration at time `t`. This is the adaptedness side of the fixed-commit martingale-difference route.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-e12041efa8ea","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":697,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:115"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_centeredReward_actionWithCommit_historyFiltrationSucc {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (t : Nat) : @Measurable Omega Real (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) inferInstance (fun omega : Omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real)))","missing":[],"search":"measurable_centeredreward_actionwithcommit_historyfiltrationsucc banditrlproof.etc.measurable_centeredreward_actionwithcommit_historyfiltrationsucc the action-matched centered reward at time `t` is measurable with respect to the shifted generated history filtration at time `t`. this is the adaptedness side of the fixed-commit martingale-difference route. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.stronglyAdapted_centeredReward_actionWithCommit_historyFiltrationSucc","label":"stronglyAdapted_centeredReward_actionWithCommit_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.stronglyAdapted_centeredReward_actionWithCommit_historyFiltrationSucc","description":"The fixed-commit action-matched centered reward process is strongly adapted to the shifted generated history filtration.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-490870e5a49d","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":698,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:163"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem stronglyAdapted_centeredReward_actionWithCommit_historyFiltrationSucc {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) : MeasureTheory.StronglyAdapted (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward) (fun t omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real)))","missing":[],"search":"stronglyadapted_centeredreward_actionwithcommit_historyfiltrationsucc banditrlproof.etc.stronglyadapted_centeredreward_actionwithcommit_historyfiltrationsucc the fixed-commit action-matched centered reward process is strongly adapted to the shifted generated history filtration. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_actionWithCommit_integrable_of_boundedRewardTraceSource","label":"centeredReward_actionWithCommit_integrable_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_actionWithCommit_integrable_of_boundedRewardTraceSource","description":"A bounded reward source supplies integrability of the action-matched centered reward coordinate.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-44f3e82bcb09","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":699,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:192"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_actionWithCommit_integrable_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (t : Nat) (ht : t < spec.explorationPulls * K) : MeasureTheory.Integrable (fun omega : Omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) mu","missing":[],"search":"centeredreward_actionwithcommit_integrable_of_boundedrewardtracesource banditrlproof.etc.centeredreward_actionwithcommit_integrable_of_boundedrewardtracesource a bounded reward source supplies integrability of the action-matched centered reward coordinate. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_boundedRewardTraceSource","label":"centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_boundedRewardTraceSource","description":"A bounded reward source supplies the succ-indexed conditional mean-zero witness for the fixed-commit shifted history filtration. The source gives reward-coordinate independence and the action-matched zero-integral identity. The action equality rewrites that identity into the sampled-arm shape consumed by the reward-level conditional mean-zero wrapper.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-1a2c38cf6da2","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":700,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:241"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (b : Fin K) (i : Nat) (hact : ETC.actionWithCommit spec commitArm (i + 1) = b) (ht : i + 1 < spec.explorationPulls * K) : Filter.EventuallyEq (MeasureTheory.ae mu) (MeasureTheory.condExp (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i) mu (fun omega :…","missing":[],"search":"centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_boundedrewardtracesource banditrlproof.etc.centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_boundedrewardtracesource a bounded reward source supplies the succ-indexed conditional mean-zero witness for the fixed-commit shifted history filtration. the source gives reward-coordinate independence and the action-matched zero-integral identity. the action equality rewrites that identity into the sampled-arm shape consumed by the reward-level conditional mean-zero wrapper. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_actionWithCommit_succMartingaleDifferencePrefix_of_boundedRewardTraceSource","label":"centeredReward_actionWithCommit_succMartingaleDifferencePrefix_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_actionWithCommit_succMartingaleDifferencePrefix_of_boundedRewardTraceSource","description":"Bounded fixed-commit rewards form a finite-prefix martingale-difference witness after centering by the action-matched arm mean. This packages the adaptedness, integrability, and succ-indexed conditional mean-zero fields. It remains a witness surface, not a final martingale stopping or regret theorem.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-31e71fecfb2b","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":701,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:294"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_actionWithCommit_succMartingaleDifferencePrefix_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (n : Nat) (hn : n <= spec.explorationPulls * K) : MartingaleDiff.SuccMartingaleDifferencePrefix mu (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward) (fun t omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real)…","missing":[],"search":"centeredreward_actionwithcommit_succmartingaledifferenceprefix_of_boundedrewardtracesource banditrlproof.etc.centeredreward_actionwithcommit_succmartingaledifferenceprefix_of_boundedrewardtracesource bounded fixed-commit rewards form a finite-prefix martingale-difference witness after centering by the action-matched arm mean. this packages the adaptedness, integrability, and succ-indexed conditional mean-zero fields. it remains a witness surface, not a final martingale stopping or regret theorem. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_hasSubgaussianMGF_of_boundedRewardTraceSource","label":"centeredReward_hasSubgaussianMGF_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_hasSubgaussianMGF_of_boundedRewardTraceSource","description":"A bounded reward source contract yields the per-coordinate centered reward sub-Gaussian witness used by the reward-law route.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-56b3090dba17","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":702,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:342"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_hasSubgaussianMGF_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (t : Nat) (ht : t < spec.explorationPulls * K) : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) (ETC.centeredRewardBoundVarianceProxy lo hi (ETC.actionWithCommit spec commitArm t) t) mu","missing":[],"search":"centeredreward_hassubgaussianmgf_of_boundedrewardtracesource banditrlproof.etc.centeredreward_hassubgaussianmgf_of_boundedrewardtracesource a bounded reward source contract yields the per-coordinate centered reward sub-gaussian witness used by the reward-law route. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_boundedRewardTraceSource","label":"centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_boundedRewardTraceSource","description":"A bounded reward source supplies the succ-indexed conditional MGF witness for the fixed-commit shifted history filtration. The source only records action-matched boundedness and exact means, so this wrapper keeps the concrete reward measurability contract explicit. The action equality rewrites the action-matched unconditional sub-Gaussian witness into the sampled-arm shape consumed by the reward-level conditional wi…","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-fea79416fb8d","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":703,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:376"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (b : Fin K) (i : Nat) (hact : ETC.actionWithCommit spec commitArm (i + 1) = b) (ht : i + 1 < spec.explorationPulls * K) : ProbabilityTheory.HasCondSubgaussianMGF (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i) ((Histor…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_historyfiltrationsucc_of_boundedrewardtracesource banditrlproof.etc.centeredreward_succ_hascondsubgaussianmgf_historyfiltrationsucc_of_boundedrewardtracesource a bounded reward source supplies the succ-indexed conditional mgf witness for the fixed-commit shifted history filtration. the source only records action-matched boundedness and exact means, so this wrapper keeps the concrete reward measurability contract explicit. the action equality rewrites the action-matched unconditional sub-gaussian witness into the sampled-arm shape consumed by the reward-level conditional witness package. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource","label":"centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource","description":"Construct the reward-level conditional sub-Gaussian witness package from a bounded reward source. This is the bounded/source assembly leaf: bounded exact-mean rewards provide the zeroth unconditional MGF and the later conditional MGF witnesses via the independence bridge, while tail domination remains an explicit contract for the consumer theorem.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-54f3eecea511","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":704,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:437"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (tail : Fin K -> ENNReal) (horizon_pos : 0 < spec.explorationPulls * K) (htail : forall a : Fin K, (a = model.bestArm -> False) -> ENNReal.ofReal (Real.exp (-(ETC.centeredPairwiseGapThreshold spec model a) ^ 2 / (2 * (((Finset.range (spec.explorationPulls * K)).sum (fun t => ETC.centeredPairwiseRewardDiffVarianceProxy spec mode…","missing":[],"search":"centeredrewardcondsubgaussianwitnesses_of_boundedrewardtracesource banditrlproof.etc.centeredrewardcondsubgaussianwitnesses_of_boundedrewardtracesource construct the reward-level conditional sub-gaussian witness package from a bounded reward source. this is the bounded/source assembly leaf: bounded exact-mean rewards provide the zeroth unconditional mgf and the later conditional mgf witnesses via the independence bridge, while tail domination remains an explicit contract for the consumer theorem. definition compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian","label":"pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian","description":"Consume a bounded reward source contract through the conditional sub-Gaussian route to obtain the fixed-commit pairwise empirical-mean tail contract. This theorem closes the local source-to-tail-contract chain for fixed `actionWithCommit`; it still keeps the final tail domination contract explicit and does not address arbitrary policy predictability.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-d74a4876d571","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":705,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:501"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (tail : Fin K -> ENNReal) (horizon_pos : 0 < spec.explorationPulls * K) (hexplorationPulls_pos : 0 < spec.explorationPulls) (htail : forall a : Fin K, (a = model.bestArm -> False) -> ENNReal.ofReal (Real.exp (-(ETC.centeredPairwiseGapThreshold spec model a) ^ 2 / (2 * (((Finset.range (spec.explorationPulls * K)).sum (fun t => ETC.ce…","missing":[],"search":"pairwiseempmeantailcontract_of_boundedrewardtracesource_condsubgaussian banditrlproof.etc.pairwiseempmeantailcontract_of_boundedrewardtracesource_condsubgaussian consume a bounded reward source contract through the conditional sub-gaussian route to obtain the fixed-commit pairwise empirical-mean tail contract. this theorem closes the local source-to-tail-contract chain for fixed `actionwithcommit`; it still keeps the final tail domination contract explicit and does not address arbitrary policy predictability. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_boundedRewardTraceSource_condSubGaussian","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_boundedRewardTraceSource_condSubGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_boundedRewardTraceSource_condSubGaussian","description":"Consume a bounded reward source contract through the conditional sub-Gaussian route to obtain the fixed-commit argmax-oracle wrong-commit probability wrapper. This is the probability-facing wrapper for the bounded/source conditional route; the numerical tail budget is still supplied explicitly by `htail`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-caec757824cf","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":706,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:550"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_boundedRewardTraceSource_condSubGaussian {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (tail : Fin K -> ENNReal) (horizon_pos : 0 < spec.explorationPulls * K) (hexplorationPulls_pos : 0 < spec.explorationPulls) (htail : forall a : Fin K, (a = model.bestArm -> False) -> ENNReal.ofReal (Real.exp (-(ETC.centeredPairwiseGapThreshold spec model a) ^ 2 / (2 * (((Finset.range…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail_of_boundedrewardtracesource_condsubgaussian banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail_of_boundedrewardtracesource_condsubgaussian consume a bounded reward source contract through the conditional sub-gaussian route to obtain the fixed-commit argmax-oracle wrong-commit probability wrapper. this is the probability-facing wrapper for the bounded/source conditional route; the numerical tail budget is still supplied explicitly by `htail`. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource_canonicalTail","label":"centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource_canonicalTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource_canonicalTail","description":"Construct the reward-level conditional witness package from a bounded reward source when the caller chooses the canonical centered-diff exponential tail. This removes the explicit tail-domination assumption from `centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource` for the fixed `actionWithCommit` route.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-195a4d2d7a4f","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":707,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:603"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource_canonicalTail {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (horizon_pos : 0 < spec.explorationPulls * K) : ETC.CenteredRewardCondSubGaussianWitnesses mu spec model commitArm reward (ETC.centeredDiffSubGaussianTail spec model (ETC.centeredPairwiseRewardDiffVarianceProxy spec model commitArm (ETC.centeredRewardBoundVarianceProxy lo hi)))","missing":[],"search":"centeredrewardcondsubgaussianwitnesses_of_boundedrewardtracesource_canonicaltail banditrlproof.etc.centeredrewardcondsubgaussianwitnesses_of_boundedrewardtracesource_canonicaltail construct the reward-level conditional witness package from a bounded reward source when the caller chooses the canonical centered-diff exponential tail. this removes the explicit tail-domination assumption from `centeredrewardcondsubgaussianwitnesses_of_boundedrewardtracesource` for the fixed `actionwithcommit` route. definition compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian_canonicalTail","label":"pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian_canonicalTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian_canonicalTail","description":"Produce the fixed-commit pairwise empirical-mean tail contract from a bounded reward source through the conditional sub-Gaussian route, using the canonical centered-diff exponential tail.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-0c04443514a1","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":708,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:639"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian_canonicalTail {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (horizon_pos : 0 < spec.explorationPulls * K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : ETC.PairwiseEmpMeanTailContract mu spec model commitArm reward (ETC.centeredDiffSubGaussianTail spec model (ETC.centeredPairwiseRewardDiffVarianceProxy spec model commitArm (ETC.centeredRewardBoundVarianceProxy lo hi)))","missing":[],"search":"pairwiseempmeantailcontract_of_boundedrewardtracesource_condsubgaussian_canonicaltail banditrlproof.etc.pairwiseempmeantailcontract_of_boundedrewardtracesource_condsubgaussian_canonicaltail produce the fixed-commit pairwise empirical-mean tail contract from a bounded reward source through the conditional sub-gaussian route, using the canonical centered-diff exponential tail. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource_condSubGaussian","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource_condSubGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource_condSubGaussian","description":"Probability-facing fixed-commit bounded-source conditional sub-Gaussian route with the canonical centered-diff exponential tail budget. This theorem removes the explicit `htail` contract from `prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_boundedRewardTraceSource_condSubGaussian`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-559966f45a5a","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":709,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:682"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource_condSubGaussian {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (horizon_pos : 0 < spec.explorationPulls * K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (fun a : Fin K => ETC.empMeanAtExploration spec commitArm (reward omega) a) = model.bestArm -> False} <= ((Finset.u…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_boundedrewardtracesource_condsubgaussian banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_boundedrewardtracesource_condsubgaussian probability-facing fixed-commit bounded-source conditional sub-gaussian route with the canonical centered-diff exponential tail budget. this theorem removes the explicit `htail` contract from `prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail_of_boundedrewardtracesource_condsubgaussian`. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource","description":"Consume a bounded reward source contract to obtain the concrete argmax-oracle wrong-commit probability bound.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsource/index.html#decl-2cb7bdc90d51","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","order":710,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSource"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSource.lean:723"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (source : ETC.BoundedRewardTraceSource mu spec model commitArm reward lo hi) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (fun a : Fin K => ETC.empMeanAtExploration spec commitArm (reward omega) a) = model.bestArm -> False} <= ((Finset.univ : Finset (Fin K)).filter (fun a : Fin K => a = model.bestArm -> False)).sum (ETC.centeredDiffSubGaussianTail spec model (ETC.centeredPairwiseRewardDiffVarianceProxy spec model commitArm (ETC.centeredRewa…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_boundedrewardtracesource banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_boundedrewardtracesource consume a bounded reward source contract to obtain the concrete argmax-oracle wrong-commit probability bound. theorem compiled","shard":"modules/f2f7901b95f27b0c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredRewardBoundVarianceProxy","label":"centeredRewardBoundVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredRewardBoundVarianceProxy","description":"Sub-Gaussian variance proxy induced by an almost-sure interval bound on one raw reward coordinate.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-bbe490c07c1f","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":711,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:19"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredRewardBoundVarianceProxy {K : Nat} (lo hi : Fin K -> Nat -> Real) (a : Fin K) (t : Nat) : NNReal","missing":[],"search":"centeredrewardboundvarianceproxy banditrlproof.etc.centeredrewardboundvarianceproxy sub-gaussian variance proxy induced by an almost-sure interval bound on one raw reward coordinate. definition compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_integrable_of_mem_Icc","label":"centeredReward_integrable_of_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_integrable_of_mem_Icc","description":"An a.e. bounded raw reward coordinate is integrable. This is the bounded-to-integrable source needed by the conditional mean-zero route. It is a thin ETC-shaped wrapper around Mathlib's `MeasureTheory.Integrable.of_mem_Icc`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-e27e9062581c","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":712,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:32"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_integrable_of_mem_Icc {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (a : Fin K) (t : Nat) (hmeas : AEMeasurable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu) (hbound : Filter.Eventually (fun omega : Omega => Set.Icc (lo a t) (hi a t) (((reward omega t : Rat) : Real))) (MeasureTheory.ae mu)) : MeasureTheory.Integrable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu","missing":[],"search":"centeredreward_integrable_of_mem_icc banditrlproof.etc.centeredreward_integrable_of_mem_icc an a.e. bounded raw reward coordinate is integrable. this is the bounded-to-integrable source needed by the conditional mean-zero route. it is a thin etc-shaped wrapper around mathlib's `measuretheory.integrable.of_mem_icc`. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_integral_eq_zero_of_integral_eq_mean","label":"centeredReward_integral_eq_zero_of_integral_eq_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_integral_eq_zero_of_integral_eq_mean","description":"An exact raw-reward mean identity gives zero integral for the centered reward. This is the zero-integral source needed by the conditional mean-zero wrapper. It still keeps raw reward integrability explicit; boundedness or a concrete reward law can discharge that assumption in later leaves.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-1498f459f9c2","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":713,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:56"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_integral_eq_zero_of_integral_eq_mean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (t : Nat) (hint : MeasureTheory.Integrable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu) (hmean : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t : Rat) : Real))) = (((model.mean a : Rat) : Real))) : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t - model.mean a : Rat) : Real))) = 0","missing":[],"search":"centeredreward_integral_eq_zero_of_integral_eq_mean banditrlproof.etc.centeredreward_integral_eq_zero_of_integral_eq_mean an exact raw-reward mean identity gives zero integral for the centered reward. this is the zero-integral source needed by the conditional mean-zero wrapper. it still keeps raw reward integrability explicit; boundedness or a concrete reward law can discharge that assumption in later leaves. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_integral_eq_zero_of_mem_Icc_integral_eq_mean","label":"centeredReward_integral_eq_zero_of_mem_Icc_integral_eq_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_integral_eq_zero_of_mem_Icc_integral_eq_mean","description":"Bounded reward plus the correct raw-reward mean identity gives zero integral for the centered reward. This combines the bounded-to-integrable source with the exact-mean zero-integral wrapper.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-9468b9f90114","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":714,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:103"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_integral_eq_zero_of_mem_Icc_integral_eq_mean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (a : Fin K) (t : Nat) (hmeas : AEMeasurable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu) (hbound : Filter.Eventually (fun omega : Omega => Set.Icc (lo a t) (hi a t) (((reward omega t : Rat) : Real))) (MeasureTheory.ae mu)) (hmean : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t : Rat) : Real))) = (((model.mean a : Rat) : Real))) : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t - model.mean a : Rat) : Real))) = 0","missing":[],"search":"centeredreward_integral_eq_zero_of_mem_icc_integral_eq_mean banditrlproof.etc.centeredreward_integral_eq_zero_of_mem_icc_integral_eq_mean bounded reward plus the correct raw-reward mean identity gives zero integral for the centered reward. this combines the bounded-to-integrable source with the exact-mean zero-integral wrapper. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_hasSubgaussianMGF_of_mem_Icc_integral_eq_mean","label":"centeredReward_hasSubgaussianMGF_of_mem_Icc_integral_eq_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_hasSubgaussianMGF_of_mem_Icc_integral_eq_mean","description":"Bounded reward plus the correct mean identity gives a centered reward sub-Gaussian witness. This is the `ETC-CENTERED-REWARD-BOUNDED-SUBGAUSSIAN-SOURCE` leaf. It is a thin ETC-shaped wrapper around Mathlib's `ProbabilityTheory.hasSubgaussianMGF_of_mem_Icc`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-148daefd99e2","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":715,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:138"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_hasSubgaussianMGF_of_mem_Icc_integral_eq_mean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (a : Fin K) (t : Nat) (hmeas : AEMeasurable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu) (hbound : Filter.Eventually (fun omega : Omega => Set.Icc (lo a t) (hi a t) (((reward omega t : Rat) : Real))) (MeasureTheory.ae mu)) (hmean : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t : Rat) : Real))) = (((model.mean a : Rat) : Real))) : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => (((reward omega t - model.mean a : Rat) : Real))) (ETC.centeredRewardBoundVarianceProxy lo hi a t) mu","missing":[],"search":"centeredreward_hassubgaussianmgf_of_mem_icc_integral_eq_mean banditrlproof.etc.centeredreward_hassubgaussianmgf_of_mem_icc_integral_eq_mean bounded reward plus the correct mean identity gives a centered reward sub-gaussian witness. this is the `etc-centered-reward-bounded-subgaussian-source` leaf. it is a thin etc-shaped wrapper around mathlib's `probabilitytheory.hassubgaussianmgf_of_mem_icc`. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_bounded_centered","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_bounded_centered","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_bounded_centered","description":"Concrete argmax-oracle wrong-commit probability bound from reward-coordinate independence, almost-sure bounded rewards, and exact per-coordinate mean identities. This is the `ETC-WRONG-COMMIT-BOUNDED-REWARD-SUBGAUSSIAN-BOUND` leaf. It uses Mathlib's bounded-variable Hoeffding lemma only to produce the per-time centered reward `HasSubgaussianMGF` witnesses consumed by the already compiled reward-coordinate-law wrong-…","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-c108481a315d","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":716,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:177"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_bounded_centered {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (h_reward_meas : forall _b : Fin K, forall t, t < spec.explorationPulls * K -> AEMeasurable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu) (h_reward_bound : forall b : Fin K, forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun omega : Omega => Set.Icc (lo b t) (hi b t) (((reward omega t : Rat) : Real))) (MeasureTheory.ae mu)) (h_reward…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_bounded_centered banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_bounded_centered concrete argmax-oracle wrong-commit probability bound from reward-coordinate independence, almost-sure bounded rewards, and exact per-coordinate mean identities. this is the `etc-wrong-commit-bounded-reward-subgaussian-bound` leaf. it uses mathlib's bounded-variable hoeffding lemma only to produce the per-time centered reward `hassubgaussianmgf` witnesses consumed by the already compiled reward-coordinate-law wrong-commit theorem. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_centeredReward_subG","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_centeredReward_subG","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_centeredReward_subG","description":"Concrete argmax-oracle wrong-commit probability bound from reward-coordinate independence plus action-matched centered reward sub-Gaussian witnesses. Unlike `prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_centeredReward_subG`, this theorem only requires the centered witness for the arm actually pulled at time `t`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-64ae517553bb","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":717,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:236"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_centeredReward_subG {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (cReward : Fin K -> Nat -> NNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (h_reward_subG : forall t, t < spec.explorationPulls * K -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => (((reward omega t - model.mean (ETC.actionWithCommit spec commitArm t) : Rat) : Real))) (cReward (ETC.actionWithCommit spec commitArm t) t) mu) : mu {omega : Omega | (ETC.argmaxCommitOracle hK…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_action_centeredreward_subg banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_action_centeredreward_subg concrete argmax-oracle wrong-commit probability bound from reward-coordinate independence plus action-matched centered reward sub-gaussian witnesses. unlike `prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_centeredreward_subg`, this theorem only requires the centered witness for the arm actually pulled at time `t`. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_bounded_centered","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_bounded_centered","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_bounded_centered","description":"Concrete argmax-oracle wrong-commit probability bound from reward-coordinate independence, action-matched almost-sure reward bounds, and action-matched exact mean identities. This is the practical fixed-commit ETC source boundary for a single observed reward trace: `reward omega t` is centered at the mean of the arm actually pulled at time `t`.","url":"../modules/banditrlproof-algorithms-etcboundedrewardsubgaussian/index.html#decl-d53c079382c2","parent":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","order":718,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCBoundedRewardSubGaussian.lean:290"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_bounded_centered {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (h_reward_meas : forall t, t < spec.explorationPulls * K -> AEMeasurable (fun omega : Omega => (((reward omega t : Rat) : Real))) mu) (h_reward_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun omega : Omega => Set.Icc (lo (ETC.actionWithCommit spec commitArm t) t) (hi (ETC.actionWithCommit spec commitArm t) t) (((reward omega t : R…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_action_bounded_centered banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_action_bounded_centered concrete argmax-oracle wrong-commit probability bound from reward-coordinate independence, action-matched almost-sure reward bounds, and action-matched exact mean identities. this is the practical fixed-commit etc source boundary for a single observed reward trace: `reward omega t` is centered at the mean of the arm actually pulled at time `t`. theorem compiled","shard":"modules/ecfd9ad299323fdc.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredDiffSubGaussianTail","label":"centeredDiffSubGaussianTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredDiffSubGaussianTail","description":"Canonical pairwise tail budget for the centered reward-difference independent sub-Gaussian ETC route. The expression is exactly the `ENNReal.ofReal (Real.exp ...)` right-hand side used by the centered-diff producer over the ETC exploration horizon.","url":"../modules/banditrlproof-algorithms-etccentereddiffcanonicaltail/index.html#decl-9f73c6e27a88","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","order":719,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffCanonicalTail.lean:22"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredDiffSubGaussianTail {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (c : Fin K -> Nat -> NNReal) (a : Fin K) : ENNReal","missing":[],"search":"centereddiffsubgaussiantail banditrlproof.etc.centereddiffsubgaussiantail canonical pairwise tail budget for the centered reward-difference independent sub-gaussian etc route. the expression is exactly the `ennreal.ofreal (real.exp ...)` right-hand side used by the centered-diff producer over the etc exploration horizon. definition compiled","shard":"modules/631636627679e640.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredDiffSubGaussianWitnesses_of_indep_subG","label":"centeredDiffSubGaussianWitnesses_of_indep_subG","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredDiffSubGaussianWitnesses_of_indep_subG","description":"Build the centered-diff witness package when the caller chooses the canonical sub-Gaussian exponential tail budget. This is the `ETC-CENTERED-DIFF-SUBGAUSSIAN-CANONICAL-TAIL` leaf. It still requires the actual reward-law independence and sub-Gaussian witnesses, but it discharges the tail-domination field definitionally.","url":"../modules/banditrlproof-algorithms-etccentereddiffcanonicaltail/index.html#decl-5cd77b6084cf","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","order":720,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffCanonicalTail.lean:43"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredDiffSubGaussianWitnesses_of_indep_subG {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (c : Fin K -> Nat -> NNReal) (h_indep : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) mu) (h_subG : forall a : Fin K, (a = model.bestArm -> False) -> forall t, t ∈ Finset.range (spec.explorationPulls * K) -> ProbabilityTheory.HasSubgaussianMGF (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) (c a t) mu) : ETC.CenteredDiffSubGaussianWitnesses mu spec model commitArm reward (ETC.centeredDiffSubGaussianTail spec model c)","missing":[],"search":"centereddiffsubgaussianwitnesses_of_indep_subg banditrlproof.etc.centereddiffsubgaussianwitnesses_of_indep_subg build the centered-diff witness package when the caller chooses the canonical sub-gaussian exponential tail budget. this is the `etc-centered-diff-subgaussian-canonical-tail` leaf. it still requires the actual reward-law independence and sub-gaussian witnesses, but it discharges the tail-domination field definitionally. definition compiled","shard":"modules/631636627679e640.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiff_indep_subG","label":"pairwiseEmpMeanTailContract_of_centeredDiff_indep_subG","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiff_indep_subG","description":"Produce the fixed-commit ETC pairwise empirical-mean tail contract directly from concrete centered-diff independence and sub-Gaussian witnesses, using the canonical exponential tail budget. This theorem is the no-extra-tail-domination consumer for the independent sub-Gaussian centered-diff route.","url":"../modules/banditrlproof-algorithms-etccentereddiffcanonicaltail/index.html#decl-1bed57e7d8b0","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","order":721,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffCanonicalTail.lean:83"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_centeredDiff_indep_subG {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (c : Fin K -> Nat -> NNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_indep : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) mu) (h_subG : forall a : Fin K, (a = model.bestArm -> False) -> forall t, t ∈ Finset.range (spec.explorationPulls * K) -> ProbabilityTheory.HasSubgaussianMGF (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) (c a t) mu) : ETC.PairwiseEmpMeanTailContract mu spec model commitArm reward (ETC.centeredDiffSubGaussianTail s…","missing":[],"search":"pairwiseempmeantailcontract_of_centereddiff_indep_subg banditrlproof.etc.pairwiseempmeantailcontract_of_centereddiff_indep_subg produce the fixed-commit etc pairwise empirical-mean tail contract directly from concrete centered-diff independence and sub-gaussian witnesses, using the canonical exponential tail budget. this theorem is the no-extra-tail-domination consumer for the independent sub-gaussian centered-diff route. theorem compiled","shard":"modules/631636627679e640.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward","label":"iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward","description":"If the reward trace coordinates are independent across time, then the centered pairwise reward-difference summands are independent across time for every non-best arm. This is the `ETC-CENTERED-DIFF-INDEPENDENCE-WITNESS` leaf. It only proves the deterministic-transform part of the reward-law independence obligation. It does not prove the reward trace coordinates are independent from a kernel or environment model, and…","url":"../modules/banditrlproof-algorithms-etccentereddiffrewardindependence/index.html#decl-09e97bbef8d6","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","order":722,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffRewardIndependence.lean:26"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) mu","missing":[],"search":"iindepfun_centeredpairwiserewarddiff_of_iindepfun_reward banditrlproof.etc.iindepfun_centeredpairwiserewarddiff_of_iindepfun_reward if the reward trace coordinates are independent across time, then the centered pairwise reward-difference summands are independent across time for every non-best arm. this is the `etc-centered-diff-independence-witness` leaf. it only proves the deterministic-transform part of the reward-law independence obligation. it does not prove the reward trace coordinates are independent from a kernel or environment model, and it does not prove any sub-gaussian witness. theorem compiled","shard":"modules/4547db0ccc8c596c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiffVarianceProxy","label":"centeredPairwiseRewardDiffVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiffVarianceProxy","description":"Variance proxy induced on the centered pairwise reward-difference summand by per-arm centered reward variance proxies. Only the arm pulled by `ETC.actionWithCommit spec commitArm t` contributes: arm `a`, the model best arm, or zero if neither is pulled at time `t`.","url":"../modules/banditrlproof-algorithms-etccentereddiffrewardsubgaussian/index.html#decl-d27e61a4783b","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","order":723,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffRewardSubGaussian.lean:22"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredPairwiseRewardDiffVarianceProxy {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (cReward : Fin K -> Nat -> NNReal) (a : Fin K) (t : Nat) : NNReal","missing":[],"search":"centeredpairwiserewarddiffvarianceproxy banditrlproof.etc.centeredpairwiserewarddiffvarianceproxy variance proxy induced on the centered pairwise reward-difference summand by per-arm centered reward variance proxies. only the arm pulled by `etc.actionwithcommit spec commitarm t` contributes: arm `a`, the model best arm, or zero if neither is pulled at time `t`. definition compiled","shard":"modules/40506a24111ce02e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","label":"centeredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","description":"Transfer a centered reward sub-Gaussian witness at one time index to the corresponding centered pairwise reward-difference summand. This is the `ETC-CENTERED-DIFF-SUBGAUSSIAN-REWARD-WITNESS` leaf. It handles the deterministic action cases and the sign flip for the best-arm centered reward. It does not prove the raw centered reward is sub-Gaussian from a distributional model.","url":"../modules/banditrlproof-algorithms-etccentereddiffrewardsubgaussian/index.html#decl-1481efe92967","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","order":724,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffRewardSubGaussian.lean:43"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (cReward : Fin K -> Nat -> NNReal) (a : Fin K) (t : Nat) (hne : a = model.bestArm -> False) (h_subG : forall b : Fin K, ETC.actionWithCommit spec commitArm t = b -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) (cReward b t) mu) : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) (ETC.centeredPairwiseRewardDiffVarianceProxy spec model commitArm cReward a t) mu","missing":[],"search":"centeredpairwiserewarddiff_hassubgaussianmgf_of_centeredreward banditrlproof.etc.centeredpairwiserewarddiff_hassubgaussianmgf_of_centeredreward transfer a centered reward sub-gaussian witness at one time index to the corresponding centered pairwise reward-difference summand. this is the `etc-centered-diff-subgaussian-reward-witness` leaf. it handles the deterministic action cases and the sign flip for the best-arm centered reward. it does not prove the raw centered reward is sub-gaussian from a distributional model. theorem compiled","shard":"modules/40506a24111ce02e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_centeredReward_subG","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_centeredReward_subG","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_centeredReward_subG","description":"Canonical concrete argmax-oracle wrong-commit probability bound from trace-level reward independence plus per-time centered reward sub-Gaussian witnesses. This is the `ETC-WRONG-COMMIT-REWARD-LAW-SUBGAUSSIAN-BOUND` leaf. It composes the reward-coordinate independence transfer, the centered reward sub-Gaussian transfer, and the canonical wrong-commit probability consumer. It still leaves the source of trace-level ind…","url":"../modules/banditrlproof-algorithms-etccentereddiffrewardsubgaussian/index.html#decl-e2849a236c5e","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","order":725,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffRewardSubGaussian.lean:108"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_centeredReward_subG {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (cReward : Fin K -> Nat -> NNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (h_reward_subG : forall b : Fin K, forall t, t < spec.explorationPulls * K -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) (cReward b t) mu) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (fun a : Fin K => ETC.empMeanAtExploration spec commitAr…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_centeredreward_subg banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail_of_reward_iindepfun_centeredreward_subg canonical concrete argmax-oracle wrong-commit probability bound from trace-level reward independence plus per-time centered reward sub-gaussian witnesses. this is the `etc-wrong-commit-reward-law-subgaussian-bound` leaf. it composes the reward-coordinate independence transfer, the centered reward sub-gaussian transfer, and the canonical wrong-commit probability consumer. it still leaves the source of trace-level independence and centered reward sub-gaussianity as explicit assumptions. theorem compiled","shard":"modules/40506a24111ce02e.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.CenteredDiffSubGaussianWitnesses","label":"CenteredDiffSubGaussianWitnesses","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.CenteredDiffSubGaussianWitnesses","description":"Witness package for the concrete centered reward-difference sub-Gaussian ETC tail route. This is the `ETC-CENTERED-DIFF-SUBGAUSSIAN-WITNESS-CONTRACT` leaf. The fields are exactly the reward-law facts still missing after the compiled centered-diff producer specialization: a sub-Gaussian variance proxy, independence of the centered summands, per-index sub-Gaussian MGF witnesses on the exploration horizon, and dominati…","url":"../modules/banditrlproof-algorithms-etccentereddiffsubgaussianwitnesses/index.html#decl-de928cb93ed7","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","order":726,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffSubGaussianWitnesses.lean:25"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure CenteredDiffSubGaussianWitnesses {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) where","missing":[],"search":"centereddiffsubgaussianwitnesses banditrlproof.etc.centereddiffsubgaussianwitnesses witness package for the concrete centered reward-difference sub-gaussian etc tail route. this is the `etc-centered-diff-subgaussian-witness-contract` leaf. the fields are exactly the reward-law facts still missing after the compiled centered-diff producer specialization: a sub-gaussian variance proxy, independence of the centered summands, per-index sub-gaussian mgf witnesses on the exploration horizon, and domination by the chosen pairwise tail budget. structure compiled","shard":"modules/e21481b8a0246af2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiffSubGaussianWitnesses","label":"pairwiseEmpMeanTailContract_of_centeredDiffSubGaussianWitnesses","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiffSubGaussianWitnesses","description":"Consume a centered reward-difference witness package to build the fixed-commit ETC pairwise empirical-mean tail contract. This theorem is intentionally thin. It fixes the exact API boundary for the next reward-law leaf while reusing the already compiled centered-diff producer.","url":"../modules/banditrlproof-algorithms-etccentereddiffsubgaussianwitnesses/index.html#decl-22d2b58e4ad0","parent":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","order":727,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCenteredDiffSubGaussianWitnesses.lean:66"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_centeredDiffSubGaussianWitnesses {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (w : ETC.CenteredDiffSubGaussianWitnesses mu spec model commitArm reward tail) : ETC.PairwiseEmpMeanTailContract mu spec model commitArm reward tail","missing":[],"search":"pairwiseempmeantailcontract_of_centereddiffsubgaussianwitnesses banditrlproof.etc.pairwiseempmeantailcontract_of_centereddiffsubgaussianwitnesses consume a centered reward-difference witness package to build the fixed-commit etc pairwise empirical-mean tail contract. this theorem is intentionally thin. it fixes the exact api boundary for the next reward-law leaf while reusing the already compiled centered-diff producer. theorem compiled","shard":"modules/e21481b8a0246af2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_centeredPairwiseRewardDiff_historyFiltrationSucc","label":"measurable_centeredPairwiseRewardDiff_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_centeredPairwiseRewardDiff_historyFiltrationSucc","description":"Centered pairwise reward differences are measurable at time `t` with respect to the shifted generated history filtration. The shifted filtration contains action/reward observations with index `< t+1`, so the reward coordinate at time `t` is available. This proves only the adaptedness/measurability part of the conditional route; it does not prove any conditional MGF or conditional expectation identity.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-2e2c7c3932b7","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":728,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:28"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_centeredPairwiseRewardDiff_historyFiltrationSucc {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) (t : Nat) : @Measurable Omega Real (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) inferInstance (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega)","missing":[],"search":"measurable_centeredpairwiserewarddiff_historyfiltrationsucc banditrlproof.etc.measurable_centeredpairwiserewarddiff_historyfiltrationsucc centered pairwise reward differences are measurable at time `t` with respect to the shifted generated history filtration. the shifted filtration contains action/reward observations with index `< t+1`, so the reward coordinate at time `t` is available. this proves only the adaptedness/measurability part of the conditional route; it does not prove any conditional mgf or conditional expectation identity. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.stronglyAdapted_centeredPairwiseRewardDiff_historyFiltrationSucc","label":"stronglyAdapted_centeredPairwiseRewardDiff_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.stronglyAdapted_centeredPairwiseRewardDiff_historyFiltrationSucc","description":"The fixed-commit ETC centered reward-difference process is strongly adapted to the shifted generated history filtration. This discharges the `StronglyAdapted` witness field for the conditional centered-diff package under the local discrete history model. Conditional MGF witnesses remain separate assumptions.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-9d8aa2d183e7","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":729,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:82"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem stronglyAdapted_centeredPairwiseRewardDiff_historyFiltrationSucc {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) : MeasureTheory.StronglyAdapted (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward) (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega)","missing":[],"search":"stronglyadapted_centeredpairwiserewarddiff_historyfiltrationsucc banditrlproof.etc.stronglyadapted_centeredpairwiserewarddiff_historyfiltrationsucc the fixed-commit etc centered reward-difference process is strongly adapted to the shifted generated history filtration. this discharges the `stronglyadapted` witness field for the conditional centered-diff package under the local discrete history model. conditional mgf witnesses remain separate assumptions. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_condExp_eq_zero_of_indep","label":"centeredReward_condExp_eq_zero_of_indep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_condExp_eq_zero_of_indep","description":"Independence-based conditional mean-zero source for one centered reward coordinate. If the sigma-algebra generated by the centered reward coordinate is independent of the conditioning sigma-algebra, and the centered reward has global integral zero, then its conditional expectation against that sigma-algebra is zero. This is a thin ETC-shaped wrapper around Mathlib's `condExp_indep_eq`; it does not derive the require…","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-492f87b775dd","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":730,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:116"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_condExp_eq_zero_of_indep {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (b : Fin K) (t : Nat) (hmeas : @Measurable Omega Real mOmega inferInstance (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real)))) (h_indep : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) inferInstance) mcond mu) (h_integral : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) = 0) : Filter.EventuallyEq (MeasureTheory.ae mu) (MeasureTheory.condExp mcond mu (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real)))) (fun _omega : Omega => (0 : Rea…","missing":[],"search":"centeredreward_condexp_eq_zero_of_indep banditrlproof.etc.centeredreward_condexp_eq_zero_of_indep independence-based conditional mean-zero source for one centered reward coordinate. if the sigma-algebra generated by the centered reward coordinate is independent of the conditioning sigma-algebra, and the centered reward has global integral zero, then its conditional expectation against that sigma-algebra is zero. this is a thin etc-shaped wrapper around mathlib's `condexp_indep_eq`; it does not derive the required independence or integral identity from a reward kernel. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_condExp_historyFiltrationSucc_eq_zero_of_indep","label":"centeredReward_condExp_historyFiltrationSucc_eq_zero_of_indep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_condExp_historyFiltrationSucc_eq_zero_of_indep","description":"Shifted-history specialization of `centeredReward_condExp_eq_zero_of_indep`. This is the first compiled conditional mean-zero source for the fixed-commit ETC history filtration. It still assumes the centered reward coordinate is independent of the shifted generated history sigma-algebra and has integral zero; later product-law work should discharge those two assumptions.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-ed7ee391ffbc","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":731,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:172"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_condExp_historyFiltrationSucc_eq_zero_of_indep {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (b : Fin K) (t : Nat) (h_indep : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) inferInstance) (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) mu) (h_integral : MeasureTheory.integral mu (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) = 0) : Filter.EventuallyEq (MeasureTheory.ae mu) (Measure…","missing":[],"search":"centeredreward_condexp_historyfiltrationsucc_eq_zero_of_indep banditrlproof.etc.centeredreward_condexp_historyfiltrationsucc_eq_zero_of_indep shifted-history specialization of `centeredreward_condexp_eq_zero_of_indep`. this is the first compiled conditional mean-zero source for the fixed-commit etc history filtration. it still assumes the centered reward coordinate is independent of the shifted generated history sigma-algebra and has integral zero; later product-law work should discharge those two assumptions. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward","label":"indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward","description":"Future centered reward coordinates are independent of the reward-only past sigma-algebra generated by earlier coordinates, provided the reward trace coordinates are mutually independent. This is a product-law/history-independence bridge for the conditional mean-zero route. It deliberately targets only the reward-coordinate past `iSup`; adding the deterministic action generators from the local history filtration is a…","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-72cbf26ecd38","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":732,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:243"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (b : Fin K) (i : Nat) : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) inferInstance) (iSup fun j : Nat => iSup fun _h : j <= i => MeasurableSpace.comap (fun omega : Omega => reward omega j) inferInstance) mu","missing":[],"search":"indep_centeredreward_succ_pastreward_isup_of_iindepfun_reward banditrlproof.etc.indep_centeredreward_succ_pastreward_isup_of_iindepfun_reward future centered reward coordinates are independent of the reward-only past sigma-algebra generated by earlier coordinates, provided the reward trace coordinates are mutually independent. this is a product-law/history-independence bridge for the conditional mean-zero route. it deliberately targets only the reward-coordinate past `isup`; adding the deterministic action generators from the local history filtration is a separate wrapper. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.historyFiltrationSucc_actionWithCommit_le_pastReward_iSup","label":"historyFiltrationSucc_actionWithCommit_le_pastReward_iSup","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.historyFiltrationSucc_actionWithCommit_le_pastReward_iSup","description":"For a fixed-commit ETC action trace, the shifted generated history filtration is contained in the reward-only past sigma-algebra. The action generators are deterministic singleton preimages, hence either `univ` or `empty`; the reward generators at filtration level `i` are exactly coordinates `j <= i`. This is the missing deterministic-action bridge from the reward-only past independence leaf to the local `History.hi…","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-2ca4bd844fcf","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":733,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:327"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem historyFiltrationSucc_actionWithCommit_le_pastReward_iSup {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (spec : ETC.Spec K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (i : Nat) : (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i : MeasurableSpace Omega) <= (iSup fun j : Nat => iSup fun _h : j <= i => MeasurableSpace.comap (fun omega : Omega => reward omega j) inferInstance)","missing":[],"search":"historyfiltrationsucc_actionwithcommit_le_pastreward_isup banditrlproof.etc.historyfiltrationsucc_actionwithcommit_le_pastreward_isup for a fixed-commit etc action trace, the shifted generated history filtration is contained in the reward-only past sigma-algebra. the action generators are deterministic singleton preimages, hence either `univ` or `empty`; the reward generators at filtration level `i` are exactly coordinates `j <= i`. this is the missing deterministic-action bridge from the reward-only past independence leaf to the local `history.historyfiltrationsucc`. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_historyFiltrationSucc_of_iIndepFun_reward","label":"indep_centeredReward_succ_historyFiltrationSucc_of_iIndepFun_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.indep_centeredReward_succ_historyFiltrationSucc_of_iIndepFun_reward","description":"Future centered reward coordinates are independent of the full fixed-commit shifted history filtration when reward coordinates are mutually independent. This extends `indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward` from the reward-only past to `History.historyFiltrationSucc` by the deterministic action-generator inclusion above.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-12a5e2e5c27a","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":734,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:433"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem indep_centeredReward_succ_historyFiltrationSucc_of_iIndepFun_reward {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (b : Fin K) (i : Nat) : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) inferInstance) (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i) mu","missing":[],"search":"indep_centeredreward_succ_historyfiltrationsucc_of_iindepfun_reward banditrlproof.etc.indep_centeredreward_succ_historyfiltrationsucc_of_iindepfun_reward future centered reward coordinates are independent of the full fixed-commit shifted history filtration when reward coordinates are mutually independent. this extends `indep_centeredreward_succ_pastreward_isup_of_iindepfun_reward` from the reward-only past to `history.historyfiltrationsucc` by the deterministic action-generator inclusion above. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_indep","label":"centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_indep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_indep","description":"One-step conditional mean-zero wrapper for Mathlib's conditional tail API. Mathlib asks for the summand at `i + 1` to be conditionally controlled against filtration level `i`. This theorem specializes the generic independence-based conditional expectation wrapper to that index shape.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-df8345cf6638","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":735,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:475"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_indep {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (b : Fin K) (i : Nat) (h_indep : ProbabilityTheory.Indep (MeasurableSpace.comap (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) inferInstance) (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i) mu) (h_integral : MeasureTheory.integral mu (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) = 0) : Filter.EventuallyEq (MeasureTheor…","missing":[],"search":"centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_indep banditrlproof.etc.centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_indep one-step conditional mean-zero wrapper for mathlib's conditional tail api. mathlib asks for the summand at `i + 1` to be conditionally controlled against filtration level `i`. this theorem specializes the generic independence-based conditional expectation wrapper to that index shape. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_iIndepFun_reward","label":"centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_iIndepFun_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_iIndepFun_reward","description":"One-step conditional mean-zero source from reward-coordinate independence. This combines the full fixed-commit history independence theorem with the existing `condExp_indep_eq` wrapper. It still leaves the zero-integral identity as an explicit side condition, supplied separately by exact-mean or bounded source leaves.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-b952f8f67c1b","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":736,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:544"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_iIndepFun_reward {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (b : Fin K) (i : Nat) (h_integral : MeasureTheory.integral mu (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) = 0) : Filter.EventuallyEq (MeasureTheory.ae mu) (MeasureTheory.condExp (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i) mu (fun omega : Omega => (((re…","missing":[],"search":"centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_iindepfun_reward banditrlproof.etc.centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_iindepfun_reward one-step conditional mean-zero source from reward-coordinate independence. this combines the full fixed-commit history independence theorem with the existing `condexp_indep_eq` wrapper. it still leaves the zero-integral identity as an explicit side condition, supplied separately by exact-mean or bounded source leaves. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.hasCondSubgaussianMGF_of_indep_comap","label":"hasCondSubgaussianMGF_of_indep_comap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.hasCondSubgaussianMGF_of_indep_comap","description":"Unconditional sub-Gaussianity plus independence from the conditioning sigma-algebra gives conditional sub-Gaussianity. This is a Mathlib-shaped source for later sampled reward-law witnesses. It uses `Kernel.HasSubgaussianMGF.of_rat`: for each rational exponent the conditional MGF is the unconditional MGF by `condExp_indep_eq`, and Mathlib extends the bound to all real exponents by continuity/density.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-9b065252cebf","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":737,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:592"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem hasCondSubgaussianMGF_of_indep_comap {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (X : Omega -> Real) (c : NNReal) (hmeasX : @Measurable Omega Real mOmega inferInstance X) (h_subG : ProbabilityTheory.HasSubgaussianMGF X c mu) (h_indep : ProbabilityTheory.Indep (MeasurableSpace.comap X inferInstance) mcond mu) : ProbabilityTheory.HasCondSubgaussianMGF mcond hm X c mu","missing":[],"search":"hascondsubgaussianmgf_of_indep_comap banditrlproof.etc.hascondsubgaussianmgf_of_indep_comap unconditional sub-gaussianity plus independence from the conditioning sigma-algebra gives conditional sub-gaussianity. this is a mathlib-shaped source for later sampled reward-law witnesses. it uses `kernel.hassubgaussianmgf.of_rat`: for each rational exponent the conditional mgf is the unconditional mgf by `condexp_indep_eq`, and mathlib extends the bound to all real exponents by continuity/density. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_iIndepFun_reward","label":"centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_iIndepFun_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_iIndepFun_reward","description":"One-step conditional sub-Gaussian source from reward-coordinate independence. This is the sampled centered-reward MGF analogue of `centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_iIndepFun_reward`: unconditional sub-Gaussianity of the future centered reward, together with `iIndepFun` reward coordinates, gives the conditional MGF witness against the fixed-commit shifted history filtration.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-121efcf32cd3","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":738,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:723"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_iIndepFun_reward {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega t)) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (b : Fin K) (i : Nat) (c : NNReal) (h_subG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) c mu) : ProbabilityTheory.HasCondSubgaussianMGF (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward i) ((…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_historyfiltrationsucc_of_iindepfun_reward banditrlproof.etc.centeredreward_succ_hascondsubgaussianmgf_historyfiltrationsucc_of_iindepfun_reward one-step conditional sub-gaussian source from reward-coordinate independence. this is the sampled centered-reward mgf analogue of `centeredreward_succ_condexp_historyfiltrationsucc_eq_zero_of_iindepfun_reward`: unconditional sub-gaussianity of the future centered reward, together with `iindepfun` reward coordinates, gives the conditional mgf witness against the fixed-commit shifted history filtration. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_action_miss","label":"centeredPairwiseRewardDiff_hasSubgaussianMGF_of_action_miss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_action_miss","description":"If the fixed ETC trace at time `t` pulls neither arm `a` nor the selected best arm, the centered pairwise reward-difference summand is identically zero and therefore has a zero-variance sub-Gaussian MGF. This is a small reward-law source for the independent/zeroth MGF side only; it does not cover the sampled-arm times where the reward coordinate appears.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-a18a4726dcbe","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":739,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:793"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasSubgaussianMGF_of_action_miss {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (t : Nat) (h_ne_a : ETC.actionWithCommit spec commitArm t ≠ a) (h_ne_best : ETC.actionWithCommit spec commitArm t ≠ model.bestArm) : ProbabilityTheory.HasSubgaussianMGF (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) 0 mu","missing":[],"search":"centeredpairwiserewarddiff_hassubgaussianmgf_of_action_miss banditrlproof.etc.centeredpairwiserewarddiff_hassubgaussianmgf_of_action_miss if the fixed etc trace at time `t` pulls neither arm `a` nor the selected best arm, the centered pairwise reward-difference summand is identically zero and therefore has a zero-variance sub-gaussian mgf. this is a small reward-law source for the independent/zeroth mgf side only; it does not cover the sampled-arm times where the reward coordinate appears. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_miss","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_miss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_miss","description":"If the fixed ETC trace at time `t` pulls neither arm `a` nor the selected best arm, the centered pairwise reward-difference summand is conditionally sub-Gaussian with zero variance proxy with respect to any sub-sigma-algebra. This discharges only the zero summand cases of the conditional MGF field.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-f64327a213ff","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":740,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:820"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_miss {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (m : MeasurableSpace Omega) (hm : m ≤ mOmega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (t : Nat) (h_ne_a : ETC.actionWithCommit spec commitArm t ≠ a) (h_ne_best : ETC.actionWithCommit spec commitArm t ≠ model.bestArm) : ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) 0 mu","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_of_action_miss banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_of_action_miss if the fixed etc trace at time `t` pulls neither arm `a` nor the selected best arm, the centered pairwise reward-difference summand is conditionally sub-gaussian with zero variance proxy with respect to any sub-sigma-algebra. this discharges only the zero summand cases of the conditional mgf field. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_miss","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_miss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_miss","description":"History-filtration specialization of the zero-summand conditional MGF source. For conditional concentration over the shifted generated history filtration, the summand at process time `t` is conditionally sub-Gaussian with zero variance whenever `actionWithCommit t` is neither the comparison arm nor the best arm.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-6f4f5bcf3be7","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":741,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:849"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_miss {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) (t : Nat) (h_ne_a : ETC.actionWithCommit spec commitArm t ≠ a) (h_ne_best : ETC.actionWithCommit spec commitArm t ≠ model.bestArm) : ProbabilityTheory.HasCondSubgaussianMGF (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) ((History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hr…","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_action_miss banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_action_miss history-filtration specialization of the zero-summand conditional mgf source. for conditional concentration over the shifted generated history filtration, the summand at process time `t` is conditionally sub-gaussian with zero variance whenever `actionwithcommit t` is neither the comparison arm nor the best arm. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_arm","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_arm","description":"Sampled comparison-arm conditional MGF transfer. When the fixed ETC trace pulls the comparison arm `a` at time `t`, the centered pairwise reward-difference summand is exactly the centered reward `reward t - mean a`, so any conditional MGF witness for that centered reward transfers to the pairwise summand.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-e9d2307ac8f8","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":742,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:902"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_arm {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (m : MeasurableSpace Omega) (hm : m ≤ mOmega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (t : Nat) (c : NNReal) (hne : a = model.bestArm -> False) (h_action : ETC.actionWithCommit spec commitArm t = a) (h_subG : ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega : Omega => (((reward omega t - model.mean a : Rat) : Real))) c mu) : ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) c mu","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_of_action_eq_arm banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_of_action_eq_arm sampled comparison-arm conditional mgf transfer. when the fixed etc trace pulls the comparison arm `a` at time `t`, the centered pairwise reward-difference summand is exactly the centered reward `reward t - mean a`, so any conditional mgf witness for that centered reward transfers to the pairwise summand. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_bestArm","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_bestArm","description":"Sampled best-arm conditional MGF transfer. When the fixed ETC trace pulls `model.bestArm` at time `t`, the centered pairwise reward-difference summand is the negative centered best-arm reward `mean bestArm - reward t`. This theorem consumes that negative-direction conditional MGF witness explicitly.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-fb09790fe817","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":743,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:943"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_bestArm {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (m : MeasurableSpace Omega) (hm : m ≤ mOmega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (t : Nat) (c : NNReal) (hne : a = model.bestArm -> False) (h_action : ETC.actionWithCommit spec commitArm t = model.bestArm) (h_subG : ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega : Omega => (((model.mean model.bestArm - reward omega t : Rat) : Real))) c mu) : ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) c mu","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_of_action_eq_bestarm banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_of_action_eq_bestarm sampled best-arm conditional mgf transfer. when the fixed etc trace pulls `model.bestarm` at time `t`, the centered pairwise reward-difference summand is the negative centered best-arm reward `mean bestarm - reward t`. this theorem consumes that negative-direction conditional mgf witness explicitly. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_arm","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_arm","description":"Shifted-history specialization of the sampled comparison-arm conditional MGF transfer.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-37e504f296f4","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":744,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:983"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_arm {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) (t : Nat) (c : NNReal) (hne : a = model.bestArm -> False) (h_action : ETC.actionWithCommit spec commitArm t = a) (h_subG : ProbabilityTheory.HasCondSubgaussianMGF (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) ((History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward).l…","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_action_eq_arm banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_action_eq_arm shifted-history specialization of the sampled comparison-arm conditional mgf transfer. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_bestArm","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_bestArm","description":"Shifted-history specialization of the sampled best-arm conditional MGF transfer.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-5679520035b4","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":745,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1046"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_bestArm {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) (t : Nat) (c : NNReal) (hne : a = model.bestArm -> False) (h_action : ETC.actionWithCommit spec commitArm t = model.bestArm) (h_subG : ProbabilityTheory.HasCondSubgaussianMGF (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) ((History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_c…","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_action_eq_bestarm banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_action_eq_bestarm shifted-history specialization of the sampled best-arm conditional mgf transfer. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_centeredReward","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_centeredReward","description":"Action-case assembly for sampled centered-reward conditional MGF witnesses. This is the conditional analogue of `ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward`: if the reward coordinate pulled by the fixed ETC trace has the appropriate centered conditional MGF witness, then the concrete centered pairwise reward-difference summand has the corresponding action-selected variance proxy. The miss cas…","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-75e14e95716f","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":746,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1117"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_centeredReward {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (m : MeasurableSpace Omega) (hm : m ≤ mOmega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (cReward : Fin K -> Nat -> NNReal) (a : Fin K) (t : Nat) (hne : a = model.bestArm -> False) (h_subG : forall b : Fin K, ETC.actionWithCommit spec commitArm t = b -> ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega : Omega => (((reward omega t - model.mean b : Rat) : Real))) (cReward b t) mu) : ProbabilityTheory.HasCondSubgaussianMGF m hm (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) (ETC.centeredPairwiseRewardDiffVarianceProxy spec model commitArm cReward a t) mu","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_of_centeredreward banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_of_centeredreward action-case assembly for sampled centered-reward conditional mgf witnesses. this is the conditional analogue of `etc.centeredpairwiserewarddiff_hassubgaussianmgf_of_centeredreward`: if the reward coordinate pulled by the fixed etc trace has the appropriate centered conditional mgf witness, then the concrete centered pairwise reward-difference summand has the corresponding action-selected variance proxy. the miss case is closed by the zero-summand theorem, and the best-arm case uses the sign-flipped conditional mgf witness. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_centeredReward","label":"centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_centeredReward","description":"Shifted-history specialization of the action-case conditional MGF assembly from sampled centered-reward witnesses.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-ee1bb0db0057","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":747,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1194"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_centeredReward {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (cReward : Fin K -> Nat -> NNReal) (a : Fin K) (t : Nat) (hne : a = model.bestArm -> False) (h_subG : forall b : Fin K, ETC.actionWithCommit spec commitArm t = b -> ProbabilityTheory.HasCondSubgaussianMGF (History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat => measurable_const) hreward t) ((History.historyFiltrationSucc (fun _omega : Omega => ETC.actionWithCommit spec commitArm) reward (fun _t : Nat…","missing":[],"search":"centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_centeredreward banditrlproof.etc.centeredpairwiserewarddiff_hascondsubgaussianmgf_historyfiltrationsucc_of_centeredreward shifted-history specialization of the action-case conditional mgf assembly from sampled centered-reward witnesses. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.CenteredRewardCondSubGaussianWitnesses","label":"CenteredRewardCondSubGaussianWitnesses","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.CenteredRewardCondSubGaussianWitnesses","description":"Reward-level conditional sub-Gaussian source contract for the fixed-commit ETC centered-diff route. This contract is one layer closer to a concrete reward law than `CenteredDiffCondSubGaussianWitnesses`: it asks for conditional MGF witnesses for the sampled reward coordinate centered at the arm actually pulled by `actionWithCommit`, plus the zeroth unconditional witness and the final tail budget. The constructor bel…","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-8a38748ddc00","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":748,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1266"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure CenteredRewardCondSubGaussianWitnesses {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) where","missing":[],"search":"centeredrewardcondsubgaussianwitnesses banditrlproof.etc.centeredrewardcondsubgaussianwitnesses reward-level conditional sub-gaussian source contract for the fixed-commit etc centered-diff route. this contract is one layer closer to a concrete reward law than `centereddiffcondsubgaussianwitnesses`: it asks for conditional mgf witnesses for the sampled reward coordinate centered at the arm actually pulled by `actionwithcommit`, plus the zeroth unconditional witness and the final tail budget. the constructor below turns it into the centered-diff witness package. structure compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.CenteredDiffCondSubGaussianWitnesses","label":"CenteredDiffCondSubGaussianWitnesses","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.CenteredDiffCondSubGaussianWitnesses","description":"Witness package for the concrete centered reward-difference conditional sub-Gaussian ETC tail route. This is a project-local conditional reward-law surface: it records the filtration, strong adaptedness, zeroth unconditional MGF witness, later conditional MGF witnesses, and tail domination needed to produce the fixed ETC pairwise empirical-mean tail contract.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-c06a5edd5709","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":749,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1325"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure CenteredDiffCondSubGaussianWitnesses {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) where","missing":[],"search":"centereddiffcondsubgaussianwitnesses banditrlproof.etc.centereddiffcondsubgaussianwitnesses witness package for the concrete centered reward-difference conditional sub-gaussian etc tail route. this is a project-local conditional reward-law surface: it records the filtration, strong adaptedness, zeroth unconditional mgf witness, later conditional mgf witnesses, and tail domination needed to produce the fixed etc pairwise empirical-mean tail contract. structure compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredDiffCondSubGaussianWitnesses_of_centeredRewardCondSubGaussianWitnesses","label":"centeredDiffCondSubGaussianWitnesses_of_centeredRewardCondSubGaussianWitnesses","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredDiffCondSubGaussianWitnesses_of_centeredRewardCondSubGaussianWitnesses","description":"Build the centered-diff conditional witness package from reward-level sampled conditional MGF witnesses. This bridges the remaining action-case gap between raw reward-law conditional MGF facts and the already compiled centered-diff conditional tail consumer.","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-9f9b87f237c2","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":750,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1377"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredDiffCondSubGaussianWitnesses_of_centeredRewardCondSubGaussianWitnesses {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) (w : ETC.CenteredRewardCondSubGaussianWitnesses mu spec model commitArm reward tail) : ETC.CenteredDiffCondSubGaussianWitnesses mu spec model commitArm reward tail where","missing":[],"search":"centereddiffcondsubgaussianwitnesses_of_centeredrewardcondsubgaussianwitnesses banditrlproof.etc.centereddiffcondsubgaussianwitnesses_of_centeredrewardcondsubgaussianwitnesses build the centered-diff conditional witness package from reward-level sampled conditional mgf witnesses. this bridges the remaining action-case gap between raw reward-law conditional mgf facts and the already compiled centered-diff conditional tail consumer. definition compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiffCondSubGaussianWitnesses","label":"pairwiseEmpMeanTailContract_of_centeredDiffCondSubGaussianWitnesses","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiffCondSubGaussianWitnesses","description":"Consume a centered reward-difference conditional witness package to build the fixed-commit ETC pairwise empirical-mean tail contract. This is the first compiled conditional reward-law bridge for the ETC route. It uses the generated centered-diff event inclusion plus the Mathlib-backed conditional sub-Gaussian finite-prefix wrapper, while keeping the actual conditional MGF and adaptedness facts explicit in the witnes…","url":"../modules/banditrlproof-algorithms-etccondsubgaussianwitnesses/index.html#decl-e767092c2f16","parent":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","order":751,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses"],["Source","BanditRLProof/Algorithms/ETCCondSubGaussianWitnesses.lean:1442"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_centeredDiffCondSubGaussianWitnesses {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] [MeasureTheory.IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (w : ETC.CenteredDiffCondSubGaussianWitnesses mu spec model commitArm reward tail) : ETC.PairwiseEmpMeanTailContract mu spec model commitArm reward tail","missing":[],"search":"pairwiseempmeantailcontract_of_centereddiffcondsubgaussianwitnesses banditrlproof.etc.pairwiseempmeantailcontract_of_centereddiffcondsubgaussianwitnesses consume a centered reward-difference conditional witness package to build the fixed-commit etc pairwise empirical-mean tail contract. this is the first compiled conditional reward-law bridge for the etc route. it uses the generated centered-diff event inclusion plus the mathlib-backed conditional sub-gaussian finite-prefix wrapper, while keeping the actual conditional mgf and adaptedness facts explicit in the witness package. theorem compiled","shard":"modules/6afd18aa9d81068c.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_exploreArm_K_eq_one","label":"pullCount_exploreArm_K_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_exploreArm_K_eq_one","description":"During the first full round-robin exploration cycle, every arm is pulled exactly once. This is the `ETC-ROUND-ROBIN-FIRST-CYCLE-COUNT` project-local deterministic count scaffold.","url":"../modules/banditrlproof-algorithms-etccountlemmas/index.html#decl-0e49e4f13519","parent":"module:BanditRLProof.Algorithms.ETCCountLemmas","order":752,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCountLemmas"],["Source","BanditRLProof/Algorithms/ETCCountLemmas.lean:22"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_exploreArm_K_eq_one {K : Nat} (spec : ETC.Spec K) (a : Fin K) : pullCount (ETC.exploreArm spec) a K = 1","missing":[],"search":"pullcount_explorearm_k_eq_one banditrlproof.etc.pullcount_explorearm_k_eq_one during the first full round-robin exploration cycle, every arm is pulled exactly once. this is the `etc-round-robin-first-cycle-count` project-local deterministic count scaffold. theorem compiled","shard":"modules/ab29fc9835e88284.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_exploreArm_add_K_eq_add_one","label":"pullCount_exploreArm_add_K_eq_add_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_exploreArm_add_K_eq_add_one","description":"Extending a round-robin ETC exploration prefix by one full cycle adds exactly one pull of every arm. This is the `ETC-ROUND-ROBIN-ADD-K-COUNT` project-local deterministic count scaffold.","url":"../modules/banditrlproof-algorithms-etccountlemmas/index.html#decl-6f9456f9172d","parent":"module:BanditRLProof.Algorithms.ETCCountLemmas","order":753,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCountLemmas"],["Source","BanditRLProof/Algorithms/ETCCountLemmas.lean:58"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_exploreArm_add_K_eq_add_one {K : Nat} (spec : ETC.Spec K) (a : Fin K) (t : Nat) : pullCount (ETC.exploreArm spec) a (t + K) = pullCount (ETC.exploreArm spec) a t + 1","missing":[],"search":"pullcount_explorearm_add_k_eq_add_one banditrlproof.etc.pullcount_explorearm_add_k_eq_add_one extending a round-robin etc exploration prefix by one full cycle adds exactly one pull of every arm. this is the `etc-round-robin-add-k-count` project-local deterministic count scaffold. theorem compiled","shard":"modules/ab29fc9835e88284.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_exploreArm_mul_K_eq","label":"pullCount_exploreArm_mul_K_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_exploreArm_mul_K_eq","description":"Across `m` full round-robin ETC exploration cycles, every arm is pulled exactly `m` times. This is the `ETC-ROUND-ROBIN-MUL-K-COUNT` project-local deterministic count scaffold.","url":"../modules/banditrlproof-algorithms-etccountlemmas/index.html#decl-b0daf6ffae83","parent":"module:BanditRLProof.Algorithms.ETCCountLemmas","order":754,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCountLemmas"],["Source","BanditRLProof/Algorithms/ETCCountLemmas.lean:86"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_exploreArm_mul_K_eq {K : Nat} (spec : ETC.Spec K) (a : Fin K) (m : Nat) : pullCount (ETC.exploreArm spec) a (m * K) = m","missing":[],"search":"pullcount_explorearm_mul_k_eq banditrlproof.etc.pullcount_explorearm_mul_k_eq across `m` full round-robin etc exploration cycles, every arm is pulled exactly `m` times. this is the `etc-round-robin-mul-k-count` project-local deterministic count scaffold. theorem compiled","shard":"modules/ab29fc9835e88284.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_exploreArm_explorationPulls_mul_K_eq","label":"pullCount_exploreArm_explorationPulls_mul_K_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_exploreArm_explorationPulls_mul_K_eq","description":"At the configured ETC exploration horizon, every arm has been pulled exactly `spec.explorationPulls` times. This is the `ETC-ROUND-ROBIN-EXPLORATION-PULLS-COUNT` project-local deterministic count adapter.","url":"../modules/banditrlproof-algorithms-etccountlemmas/index.html#decl-f14b16218242","parent":"module:BanditRLProof.Algorithms.ETCCountLemmas","order":755,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCCountLemmas"],["Source","BanditRLProof/Algorithms/ETCCountLemmas.lean:105"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_exploreArm_explorationPulls_mul_K_eq {K : Nat} (spec : ETC.Spec K) (a : Fin K) : pullCount (ETC.exploreArm spec) a (spec.explorationPulls * K) = spec.explorationPulls","missing":[],"search":"pullcount_explorearm_explorationpulls_mul_k_eq banditrlproof.etc.pullcount_explorearm_explorationpulls_mul_k_eq at the configured etc exploration horizon, every arm has been pulled exactly `spec.explorationpulls` times. this is the `etc-round-robin-exploration-pulls-count` project-local deterministic count adapter. theorem compiled","shard":"modules/ab29fc9835e88284.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration","label":"empMeanAtExploration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration","description":"Empirical mean of arm `a` at the configured ETC exploration horizon for a fixed-commit trace. This is the `ETC-EMP-MEAN-ACTION-WITH-COMMIT-EXPLORATION` project-local leaf.","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html#decl-e0ffc1467dd7","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","order":756,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean:24"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def empMeanAtExploration {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) (reward : RewardTrace Rat) (a : Fin K) : Rat","missing":[],"search":"empmeanatexploration banditrlproof.etc.empmeanatexploration empirical mean of arm `a` at the configured etc exploration horizon for a fixed-commit trace. this is the `etc-emp-mean-action-with-commit-exploration` project-local leaf. definition compiled","shard":"modules/584d56d84b375281.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration_eq_of_eq_on_prefix","label":"empMeanAtExploration_eq_of_eq_on_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration_eq_of_eq_on_prefix","description":"Exploration empirical means depend only on the reward coordinates observed before the configured exploration horizon. This is the prefix-congruence bridge needed to reconstruct the ETC commit score from a finite reward history after exploration, rather than from an ambient reward trace with future coordinates.","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html#decl-cb7f58eefb4e","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","order":757,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean:39"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem empMeanAtExploration_eq_of_eq_on_prefix {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) (reward0 reward1 : RewardTrace Rat) (hprefix : forall t, t < spec.explorationPulls * K -> reward0 t = reward1 t) (a : Fin K) : ETC.empMeanAtExploration spec commitArm reward0 a = ETC.empMeanAtExploration spec commitArm reward1 a","missing":[],"search":"empmeanatexploration_eq_of_eq_on_prefix banditrlproof.etc.empmeanatexploration_eq_of_eq_on_prefix exploration empirical means depend only on the reward coordinates observed before the configured exploration horizon. this is the prefix-congruence bridge needed to reconstruct the etc commit score from a finite reward history after exploration, rather than from an ambient reward trace with future coordinates. theorem compiled","shard":"modules/584d56d84b375281.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration_completeRewardTrace_eq_of_explorationHorizon_le","label":"empMeanAtExploration_completeRewardTrace_eq_of_explorationHorizon_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration_completeRewardTrace_eq_of_explorationHorizon_le","description":"Once a finite reward history reaches the configured exploration horizon, its default-completed trace gives the same ETC exploration empirical mean as the ambient reward trace. This is the history-reconstruction bridge for a later generated commit policy: at generated action time `t + 1`, the state reads a history through `t`, so the explicit `spec.explorationPulls * K <= t + 1` contract supplies every score coordina…","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html#decl-7919ff2eb5fd","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","order":758,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean:69"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem empMeanAtExploration_completeRewardTrace_eq_of_explorationHorizon_le {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) (reward : RewardTrace Rat) (t : Nat) (horizon_le : spec.explorationPulls * K <= t + 1) (a : Fin K) : ETC.empMeanAtExploration spec commitArm (History.completeRewardTrace t (History.finiteRewardHistoryOfTrace reward t) (0 : Rat)) a = ETC.empMeanAtExploration spec commitArm reward a","missing":[],"search":"empmeanatexploration_completerewardtrace_eq_of_explorationhorizon_le banditrlproof.etc.empmeanatexploration_completerewardtrace_eq_of_explorationhorizon_le once a finite reward history reaches the configured exploration horizon, its default-completed trace gives the same etc exploration empirical mean as the ambient reward trace. this is the history-reconstruction bridge for a later generated commit policy: at generated action time `t + 1`, the state reads a history through `t`, so the explicit `spec.explorationpulls * k <= t + 1` contract supplies every score coordinate used by the commit rule. theorem compiled","shard":"modules/584d56d84b375281.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration_eq_sumRewards_div_explorationPulls","label":"empMeanAtExploration_eq_sumRewards_div_explorationPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration_eq_sumRewards_div_explorationPulls","description":"The empirical-mean denominator at the ETC exploration horizon rewrites to the configured number of exploration pulls per arm.","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html#decl-7975269dabf7","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","order":759,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean:88"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem empMeanAtExploration_eq_sumRewards_div_explorationPulls {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) (reward : RewardTrace Rat) (a : Fin K) : ETC.empMeanAtExploration spec commitArm reward a = sumRewards (ETC.actionWithCommit spec commitArm) reward a (spec.explorationPulls * K) / ((spec.explorationPulls : Nat) : Rat)","missing":[],"search":"empmeanatexploration_eq_sumrewards_div_explorationpulls banditrlproof.etc.empmeanatexploration_eq_sumrewards_div_explorationpulls the empirical-mean denominator at the etc exploration horizon rewrites to the configured number of exploration pulls per arm. theorem compiled","shard":"modules/584d56d84b375281.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration_le_iff_sumRewards_le_of_explorationPulls_pos","label":"empMeanAtExploration_le_iff_sumRewards_le_of_explorationPulls_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration_le_iff_sumRewards_le_of_explorationPulls_pos","description":"At a positive exploration count, comparing two fixed-commit ETC empirical means is equivalent to comparing their fixed-horizon reward sums. This is the `ETC-EMP-MEAN-COMPARISON-AS-FINITE-SUM` deterministic algebra leaf. It only removes the common positive denominator from two empirical means; it does not introduce probability, concentration, filtration, or final ETC regret.","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html#decl-a1e9162d7295","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","order":760,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean:107"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem empMeanAtExploration_le_iff_sumRewards_le_of_explorationPulls_pos {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) (reward : RewardTrace Rat) (a b : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : ETC.empMeanAtExploration spec commitArm reward b <= ETC.empMeanAtExploration spec commitArm reward a ↔ sumRewards (ETC.actionWithCommit spec commitArm) reward b (spec.explorationPulls * K) <= sumRewards (ETC.actionWithCommit spec commitArm) reward a (spec.explorationPulls * K)","missing":[],"search":"empmeanatexploration_le_iff_sumrewards_le_of_explorationpulls_pos banditrlproof.etc.empmeanatexploration_le_iff_sumrewards_le_of_explorationpulls_pos at a positive exploration count, comparing two fixed-commit etc empirical means is equivalent to comparing their fixed-horizon reward sums. this is the `etc-emp-mean-comparison-as-finite-sum` deterministic algebra leaf. it only removes the common positive denominator from two empirical means; it does not introduce probability, concentration, filtration, or final etc regret. theorem compiled","shard":"modules/584d56d84b375281.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration_ge_best_event_subset_sumRewards_tail_event_of_imp","label":"empMeanAtExploration_ge_best_event_subset_sumRewards_tail_event_of_imp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration_ge_best_event_subset_sumRewards_tail_event_of_imp","description":"At a positive exploration count, a pointwise implication from the fixed-horizon reward-sum comparison into a real finite-sum tail event yields the matching event inclusion from the non-best empirical-mean comparison event. This is the `ETC-EMPMEAN-EVENT-SUBSET-SUMREWARDS-TAIL-EVENT` bridge. It is only an event-shape adapter; it does not instantiate centered reward differences, prove sub-Gaussianity, introduce filtra…","url":"../modules/banditrlproof-algorithms-etcempiricalmean/index.html#decl-0aeb89d28671","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","order":761,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMean.lean:133"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem empMeanAtExploration_ge_best_event_subset_sumRewards_tail_event_of_imp {Omega : Type u} {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) {Idx : Type v} (idx : Finset Idx) (X : Idx -> Omega -> Real) (eps : Real) (himp : forall omega : Omega, sumRewards (ETC.actionWithCommit spec commitArm) (reward omega) model.bestArm (spec.explorationPulls * K) <= sumRewards (ETC.actionWithCommit spec commitArm) (reward omega) a (spec.explorationPulls * K) -> eps <= idx.sum (fun i => X i omega)) : Set.Subset {omega : Omega | ETC.empMeanAtExploration spec commitArm (reward omega) a >= ETC.empMeanAtExploration spec commitArm (reward omega) model.bestArm} {omega : Omega | eps <= idx.sum (fun i => X i omega)}","missing":[],"search":"empmeanatexploration_ge_best_event_subset_sumrewards_tail_event_of_imp banditrlproof.etc.empmeanatexploration_ge_best_event_subset_sumrewards_tail_event_of_imp at a positive exploration count, a pointwise implication from the fixed-horizon reward-sum comparison into a real finite-sum tail event yields the matching event inclusion from the non-best empirical-mean comparison event. this is the `etc-empmean-event-subset-sumrewards-tail-event` bridge. it is only an event-shape adapter; it does not instantiate centered reward differences, prove sub-gaussianity, introduce filtrations, or prove final etc regret. theorem compiled","shard":"modules/584d56d84b375281.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_sumRewards_actionWithCommit_exploration","label":"measurable_sumRewards_actionWithCommit_exploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_sumRewards_actionWithCommit_exploration","description":"The selected-reward numerator of the fixed-commit ETC empirical mean is measurable for stochastic reward traces with timewise measurable coordinates. This is the `ETC-MEASURABLE-SUMREWARDS-ACTION-WITH-COMMIT-EXPLORATION` project-local leaf.","url":"../modules/banditrlproof-algorithms-etcempiricalmeanmeasurability/index.html#decl-328a27e78a50","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","order":762,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMeanMeasurability.lean:27"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_sumRewards_actionWithCommit_exploration {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Rat] [MeasurableAdd₂ Rat] (spec : ETC.Spec K) (commitArm a : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Measurable (fun omega : Omega => sumRewards (ETC.actionWithCommit spec commitArm) (reward omega) a (spec.explorationPulls * K))","missing":[],"search":"measurable_sumrewards_actionwithcommit_exploration banditrlproof.etc.measurable_sumrewards_actionwithcommit_exploration the selected-reward numerator of the fixed-commit etc empirical mean is measurable for stochastic reward traces with timewise measurable coordinates. this is the `etc-measurable-sumrewards-action-with-commit-exploration` project-local leaf. theorem compiled","shard":"modules/1d18399676b48c7d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_empMeanAtExploration_of_measurable_div_const","label":"measurable_empMeanAtExploration_of_measurable_div_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_empMeanAtExploration_of_measurable_div_const","description":"The fixed-commit ETC empirical mean is measurable under stochastic reward traces once measurability of division by a constant Rat is supplied. This is the `ETC-MEASURABLE-EMPMEAN-ACTION-WITH-COMMIT-EXPLORATION-OF-DIV-CONST` project-local leaf. It deliberately leaves the Mathlib import/wrapper decision for Rat division measurability to a later leaf.","url":"../modules/banditrlproof-algorithms-etcempiricalmeanmeasurability/index.html#decl-bbf9a98472b1","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","order":763,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMeanMeasurability.lean:59"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_empMeanAtExploration_of_measurable_div_const {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Rat] [MeasurableAdd₂ Rat] (spec : ETC.Spec K) (commitArm a : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hdiv_const : forall c : Rat, Measurable (fun x : Rat => x / c)) : Measurable (fun omega : Omega => ETC.empMeanAtExploration spec commitArm (reward omega) a)","missing":[],"search":"measurable_empmeanatexploration_of_measurable_div_const banditrlproof.etc.measurable_empmeanatexploration_of_measurable_div_const the fixed-commit etc empirical mean is measurable under stochastic reward traces once measurability of division by a constant rat is supplied. this is the `etc-measurable-empmean-action-with-commit-exploration-of-div-const` project-local leaf. it deliberately leaves the mathlib import/wrapper decision for rat division measurability to a later leaf. theorem compiled","shard":"modules/1d18399676b48c7d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_empMeanAtExploration","label":"measurable_empMeanAtExploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_empMeanAtExploration","description":"The fixed-commit ETC empirical mean is measurable under stochastic reward traces, using the local Rat division-by-constant measurability wrapper. This is the `ETC-MEASURABLE-EMPMEAN-ACTION-WITH-COMMIT-EXPLORATION` project-local leaf. It removes the explicit `hdiv_const` argument from `measurable_empMeanAtExploration_of_measurable_div_const` by requiring measurable singletons on `Rat`.","url":"../modules/banditrlproof-algorithms-etcempiricalmeanmeasurability/index.html#decl-d0dc26dd6db4","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","order":764,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMeanMeasurability.lean:93"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_empMeanAtExploration {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Rat] [MeasurableSingletonClass Rat] [MeasurableAdd₂ Rat] (spec : ETC.Spec K) (commitArm a : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Measurable (fun omega : Omega => ETC.empMeanAtExploration spec commitArm (reward omega) a)","missing":[],"search":"measurable_empmeanatexploration banditrlproof.etc.measurable_empmeanatexploration the fixed-commit etc empirical mean is measurable under stochastic reward traces, using the local rat division-by-constant measurability wrapper. this is the `etc-measurable-empmean-action-with-commit-exploration` project-local leaf. it removes the explicit `hdiv_const` argument from `measurable_empmeanatexploration_of_measurable_div_const` by requiring measurable singletons on `rat`. theorem compiled","shard":"modules/1d18399676b48c7d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_empMeanAtExploration_coordinates","label":"measurable_empMeanAtExploration_coordinates","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_empMeanAtExploration_coordinates","description":"The fixed-commit ETC empirical-mean coordinates are measurable under stochastic reward traces. This is the `ETC-MEASURABLE-EMPMEAN-AT-EXPLORATION-COORDINATES` project-local leaf. It packages `measurable_empMeanAtExploration` in the `forall a : Fin K, Measurable ...` shape used by empirical-mean event measurability lemmas; it does not add argmax, concentration, or filtration contracts.","url":"../modules/banditrlproof-algorithms-etcempiricalmeanmeasurability/index.html#decl-fd7640b060de","parent":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","order":765,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability"],["Source","BanditRLProof/Algorithms/ETCEmpiricalMeanMeasurability.lean:118"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_empMeanAtExploration_coordinates {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Rat] [MeasurableSingletonClass Rat] [MeasurableAdd₂ Rat] (spec : ETC.Spec K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : forall a : Fin K, Measurable (fun omega : Omega => (fun b : Fin K => ETC.empMeanAtExploration spec commitArm (reward omega) b) a)","missing":[],"search":"measurable_empmeanatexploration_coordinates banditrlproof.etc.measurable_empmeanatexploration_coordinates the fixed-commit etc empirical-mean coordinates are measurable under stochastic reward traces. this is the `etc-measurable-empmean-at-exploration-coordinates` project-local leaf. it packages `measurable_empmeanatexploration` in the `forall a : fin k, measurable ...` shape used by empirical-mean event measurability lemmas; it does not add argmax, concentration, or filtration contracts. theorem compiled","shard":"modules/1d18399676b48c7d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.sum_centeredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","label":"sum_centeredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.sum_centeredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","description":"Over the round-robin exploration horizon, the constant common-proxy pairwise process charges exactly `m` pulls of the candidate arm and `m` pulls of the selected best arm.","url":"../modules/banditrlproof-algorithms-etcexactsubgaussiantail/index.html#decl-b53b5a072039","parent":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","order":766,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExactSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCExactSubGaussianTail.lean:23"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem sum_centeredPairwiseRewardDiffVarianceProxy_const_eq_two_mul {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (sigma2 : NNReal) (a : Fin K) (hne : a = model.bestArm -> False) : (Finset.range (spec.explorationPulls * K)).sum (fun t => ETC.centeredPairwiseRewardDiffVarianceProxy spec model model.bestArm (fun _ _ => sigma2) a t) = (2 : NNReal) * (spec.explorationPulls : NNReal) * sigma2","missing":[],"search":"sum_centeredpairwiserewarddiffvarianceproxy_const_eq_two_mul banditrlproof.etc.sum_centeredpairwiserewarddiffvarianceproxy_const_eq_two_mul over the round-robin exploration horizon, the constant common-proxy pairwise process charges exactly `m` pulls of the candidate arm and `m` pulls of the selected best arm. theorem compiled","shard":"modules/62a6d8bad834f56d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseGapThreshold_eq_explorationPulls_mul_gap","label":"centeredPairwiseGapThreshold_eq_explorationPulls_mul_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseGapThreshold_eq_explorationPulls_mul_gap","description":"The deterministic centered pairwise threshold is `m` times the model gap.","url":"../modules/banditrlproof-algorithms-etcexactsubgaussiantail/index.html#decl-7e55bb9a024f","parent":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","order":767,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExactSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCExactSubGaussianTail.lean:79"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem centeredPairwiseGapThreshold_eq_explorationPulls_mul_gap {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (a : Fin K) (hne : a = model.bestArm -> False) : ETC.centeredPairwiseGapThreshold spec model a = (spec.explorationPulls : Real) * ((model.gap a : Rat) : Real)","missing":[],"search":"centeredpairwisegapthreshold_eq_explorationpulls_mul_gap banditrlproof.etc.centeredpairwisegapthreshold_eq_explorationpulls_mul_gap the deterministic centered pairwise threshold is `m` times the model gap. theorem compiled","shard":"modules/62a6d8bad834f56d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalSubGaussianArmPairwiseTailReal_eq_exp_neg_explorationPulls_mul_gap_sq_div_four_mul","label":"canonicalSubGaussianArmPairwiseTailReal_eq_exp_neg_explorationPulls_mul_gap_sq_div_four_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalSubGaussianArmPairwiseTailReal_eq_exp_neg_explorationPulls_mul_gap_sq_div_four_mul","description":"The canonical direct-MGF pairwise tail equals the exact LML exponential `exp (-m * gap^2 / (4 * sigma2))` for every non-best arm. Only `m > 0` is needed to cancel the exploration multiplicity. The `sigma2 = 0` boundary is handled separately and remains a valid equality in Lean's total division.","url":"../modules/banditrlproof-algorithms-etcexactsubgaussiantail/index.html#decl-0363b99b8141","parent":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","order":768,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExactSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCExactSubGaussianTail.lean:99"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem canonicalSubGaussianArmPairwiseTailReal_eq_exp_neg_explorationPulls_mul_gap_sq_div_four_mul {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (sigma2 : NNReal) (a : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) (hne : a = model.bestArm -> False) : ETC.canonicalSubGaussianArmPairwiseTailReal spec model sigma2 a = Real.exp (-(spec.explorationPulls : Real) * ((model.gap a : Rat) : Real) ^ 2 / (4 * (sigma2 : Real)))","missing":[],"search":"canonicalsubgaussianarmpairwisetailreal_eq_exp_neg_explorationpulls_mul_gap_sq_div_four_mul banditrlproof.etc.canonicalsubgaussianarmpairwisetailreal_eq_exp_neg_explorationpulls_mul_gap_sq_div_four_mul the canonical direct-mgf pairwise tail equals the exact lml exponential `exp (-m * gap^2 / (4 * sigma2))` for every non-best arm. only `m > 0` is needed to cancel the exploration multiplicity. the `sigma2 = 0` boundary is handled separately and remains a valid equality in lean's total division. theorem compiled","shard":"modules/62a6d8bad834f56d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_exp_neg_explorationPulls_mul_gap_sq_div_four_mul_of_armLaws","label":"real_measure_explorationArgmaxCommit_eq_arm_le_exp_neg_explorationPulls_mul_gap_sq_div_four_mul_of_armLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_exp_neg_explorationPulls_mul_gap_sq_div_four_mul_of_armLaws","description":"Exact-LML-constant form of the canonical single-arm commit-fiber probability bound under the existing common-sub-Gaussian `Rat` arm laws. This is a genuine concentration producer, but not yet the native `Real` reward-kernel theorem: the arm laws are measures on `Rat`, their centered casts have a common `NNReal` proxy, and the generated trajectory is the local canonical history process.","url":"../modules/banditrlproof-algorithms-etcexactsubgaussiantail/index.html#decl-8e2893ef5925","parent":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","order":769,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExactSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCExactSubGaussianTail.lean:152"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_measure_explorationArgmaxCommit_eq_arm_le_exp_neg_explorationPulls_mul_gap_sq_div_four_mul_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> MeasureTheory.Measure Rat) (hprob : forall arm, MeasureTheory.IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, MeasureTheory.integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (a : Fin K) (hne : a = model.bestArm -> False) : let defaultAction := ETC.exp…","missing":[],"search":"real_measure_explorationargmaxcommit_eq_arm_le_exp_neg_explorationpulls_mul_gap_sq_div_four_mul_of_armlaws banditrlproof.etc.real_measure_explorationargmaxcommit_eq_arm_le_exp_neg_explorationpulls_mul_gap_sq_div_four_mul_of_armlaws exact-lml-constant form of the canonical single-arm commit-fiber probability bound under the existing common-sub-gaussian `rat` arm laws. this is a genuine concentration producer, but not yet the native `real` reward-kernel theorem: the arm laws are measures on `rat`, their centered casts have a common `nnreal` proxy, and the generated trajectory is the local canonical history process. theorem compiled","shard":"modules/62a6d8bad834f56d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_explorationArgmaxAction_le_exploration_add_remaining_mul_exp_of_armLaws","label":"integral_real_pullCount_explorationArgmaxAction_le_exploration_add_remaining_mul_exp_of_armLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_explorationArgmaxAction_le_exploration_add_remaining_mul_exp_of_armLaws","description":"Canonical generated-history per-arm expected pull-count bound with the exact LML exponential constant, under the existing common-sub-Gaussian `Rat` arm laws. This theorem composes the concrete commit-fiber concentration producer with the generic Real/Bochner `actionWithCommit` expected-count consumer. It is the first local theorem with the full `m + (n - K*m) * exp (...)` per-arm shape. Native `Real` reward-kernel t…","url":"../modules/banditrlproof-algorithms-etcexactsubgaussiantail/index.html#decl-e92c28595638","parent":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","order":770,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExactSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCExactSubGaussianTail.lean:207"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_explorationArgmaxAction_le_exploration_add_remaining_mul_exp_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> MeasureTheory.Measure Rat) (hprob : forall arm, MeasureTheory.IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, MeasureTheory.integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (a : Fin K) (n : Nat) (hn : K * spec.explorationPulls <= n) (hne : a = model.bestArm ->…","missing":[],"search":"integral_real_pullcount_explorationargmaxaction_le_exploration_add_remaining_mul_exp_of_armlaws banditrlproof.etc.integral_real_pullcount_explorationargmaxaction_le_exploration_add_remaining_mul_exp_of_armlaws canonical generated-history per-arm expected pull-count bound with the exact lml exponential constant, under the existing common-sub-gaussian `rat` arm laws. this theorem composes the concrete commit-fiber concentration producer with the generic real/bochner `actionwithcommit` expected-count consumer. it is the first local theorem with the full `m + (n - k*m) * exp (...)` per-arm shape. native `real` reward-kernel transport and external algorithm/environment-law alignment remain downstream. theorem compiled","shard":"modules/62a6d8bad834f56d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integrable_real_pullCount_actionWithCommit_choice_of_measurable_commit","label":"integrable_real_pullCount_actionWithCommit_choice_of_measurable_commit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integrable_real_pullCount_actionWithCommit_choice_of_measurable_commit","description":"The Real cast of a finite-horizon ETC pull count is integrable when the commit selector is measurable and the ambient measure is finite. This regularity adapter is independent of reward laws and concentration. Its proof uses timewise measurability of the finite-valued ETC action and the deterministic bound `pullCount <= n`.","url":"../modules/banditrlproof-algorithms-etcexpectedpullcount/index.html#decl-eeec125dcdc8","parent":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","order":771,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedPullCount"],["Source","BanditRLProof/Algorithms/ETCExpectedPullCount.lean:30"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integrable_real_pullCount_actionWithCommit_choice_of_measurable_commit {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (commit : Omega -> Fin K) (a : Fin K) (n : Nat) (hmeas_commit : Measurable commit) : Integrable (fun omega : Omega => ((pullCount (ETC.actionWithCommit spec (commit omega)) a n : Nat) : Real)) mu","missing":[],"search":"integrable_real_pullcount_actionwithcommit_choice_of_measurable_commit banditrlproof.etc.integrable_real_pullcount_actionwithcommit_choice_of_measurable_commit the real cast of a finite-horizon etc pull count is integrable when the commit selector is measurable and the ambient measure is finite. this regularity adapter is independent of reward laws and concentration. its proof uses timewise measurability of the finite-valued etc action and the deterministic bound `pullcount <= n`. theorem compiled","shard":"modules/09433af36a9049c5.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_suffix_mul_commit_prob","label":"integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_suffix_mul_commit_prob","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_suffix_mul_commit_prob","description":"At horizon `spec.explorationPulls * K + r`, the expected Real pull count of arm `a` is exactly the exploration count plus `r` times the probability of committing to `a`. This is the direct Bochner/indicator integration of `ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq`.","url":"../modules/banditrlproof-algorithms-etcexpectedpullcount/index.html#decl-5b128c9db5e1","parent":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","order":772,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedPullCount"],["Source","BanditRLProof/Algorithms/ETCExpectedPullCount.lean:69"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_suffix_mul_commit_prob {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (commit : Omega -> Fin K) (a : Fin K) (r : Nat) (hmeas_commit : Measurable commit) : MeasureTheory.integral mu (fun omega : Omega => ((pullCount (ETC.actionWithCommit spec (commit omega)) a (spec.explorationPulls * K + r) : Nat) : Real)) = (spec.explorationPulls : Real) + (r : Real) * mu.real {omega : Omega | commit omega = a}","missing":[],"search":"integral_real_pullcount_actionwithcommit_choice_eq_exploration_add_suffix_mul_commit_prob banditrlproof.etc.integral_real_pullcount_actionwithcommit_choice_eq_exploration_add_suffix_mul_commit_prob at horizon `spec.explorationpulls * k + r`, the expected real pull count of arm `a` is exactly the exploration count plus `r` times the probability of committing to `a`. this is the direct bochner/indicator integration of `etc.pullcount_actionwithcommit_explorationpulls_mul_k_add_eq`. theorem compiled","shard":"modules/09433af36a9049c5.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_remaining_mul_commit_prob","label":"integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_remaining_mul_commit_prob","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_remaining_mul_commit_prob","description":"LML-shaped horizon form of the exact ETC expected pull-count identity. The exploration condition is written as `K * explorationPulls <= n`, and the suffix is `n - K * explorationPulls`, matching the exact ETC theorem-card surface. No concentration or reward-law assumption is used.","url":"../modules/banditrlproof-algorithms-etcexpectedpullcount/index.html#decl-8401876e6e47","parent":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","order":773,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedPullCount"],["Source","BanditRLProof/Algorithms/ETCExpectedPullCount.lean:116"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_remaining_mul_commit_prob {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (commit : Omega -> Fin K) (a : Fin K) (n : Nat) (hn : K * spec.explorationPulls <= n) (hmeas_commit : Measurable commit) : MeasureTheory.integral mu (fun omega : Omega => ((pullCount (ETC.actionWithCommit spec (commit omega)) a n : Nat) : Real)) = (spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * mu.real {omega : Omega | commit omega = a}","missing":[],"search":"integral_real_pullcount_actionwithcommit_choice_eq_exploration_add_remaining_mul_commit_prob banditrlproof.etc.integral_real_pullcount_actionwithcommit_choice_eq_exploration_add_remaining_mul_commit_prob lml-shaped horizon form of the exact etc expected pull-count identity. the exploration condition is written as `k * explorationpulls <= n`, and the suffix is `n - k * explorationpulls`, matching the exact etc theorem-card surface. no concentration or reward-law assumption is used. theorem compiled","shard":"modules/09433af36a9049c5.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le","label":"integral_real_pullCount_actionWithCommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le","description":"Per-arm ETC expected pull-count bound from a commit-fiber probability bound. This is the concentration consumer needed by the exact LML route: a later empirical-mean tail theorem supplies only `mu.real {omega | commit omega = a} <= p`; the counting and integration steps are discharged here.","url":"../modules/banditrlproof-algorithms-etcexpectedpullcount/index.html#decl-e3b77a21bf97","parent":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","order":774,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedPullCount"],["Source","BanditRLProof/Algorithms/ETCExpectedPullCount.lean:163"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_actionWithCommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (commit : Omega -> Fin K) (a : Fin K) (n : Nat) (p : Real) (hn : K * spec.explorationPulls <= n) (hmeas_commit : Measurable commit) (hprob : mu.real {omega : Omega | commit omega = a} <= p) : MeasureTheory.integral mu (fun omega : Omega => ((pullCount (ETC.actionWithCommit spec (commit omega)) a n : Nat) : Real)) <= (spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * p","missing":[],"search":"integral_real_pullcount_actionwithcommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le banditrlproof.etc.integral_real_pullcount_actionwithcommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le per-arm etc expected pull-count bound from a commit-fiber probability bound. this is the concentration consumer needed by the exact lml route: a later empirical-mean tail theorem supplies only `mu.real {omega | commit omega = a} <= p`; the counting and integration steps are discharged here. theorem compiled","shard":"modules/09433af36a9049c5.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","label":"lintegral_ofReal_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","description":"Lower-integral assembly for an `Omega`-indexed ETC commit selector. The theorem consumes a pointwise non-best gap bound and an abstract upper bound `pWrong` on the wrong-commit event probability. It is intentionally still an `ENNReal.ofReal` lower-integral statement, matching the existing expectation surface in the project.","url":"../modules/banditrlproof-algorithms-etcexpectedregretassembly/index.html#decl-b2bd8fc82470","parent":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","order":775,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCExpectedRegretAssembly.lean:33"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commit : Omega -> Fin K) (r : Nat) (badGapBound : Rat) (pWrong : ENNReal) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (hmeas_wrong : MeasurableSet {omega : Omega | commit omega = model.bestArm -> False}) (hprob_wrong : mu {omega : Omega | commit omega = model.bestArm -> False} <= pWrong) : MeasureTheory.lintegral mu (fun omega : Omega => ENNReal.ofReal (((pseudoRegret model (ETC.actionWithCommit spec (commit omega)) (spec.explorationPulls * K + r) : Rat) : Real))) <= ENNReal.ofReal (((((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat)…","missing":[],"search":"lintegral_ofreal_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_badgap_prob banditrlproof.etc.lintegral_ofreal_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_badgap_prob lower-integral assembly for an `omega`-indexed etc commit selector. the theorem consumes a pointwise non-best gap bound and an abstract upper bound `pwrong` on the wrong-commit event probability. it is intentionally still an `ennreal.ofreal` lower-integral statement, matching the existing expectation surface in the project. theorem compiled","shard":"modules/c64c4638920377a0.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integrable_real_pseudoRegret_actionWithCommit_choice_of_measurable_commit","label":"integrable_real_pseudoRegret_actionWithCommit_choice_of_measurable_commit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integrable_real_pseudoRegret_actionWithCommit_choice_of_measurable_commit","description":"Finite-arm `actionWithCommit` pseudo-regret has an integrable Real cast when the selected commit arm is measurable and the ambient measure is finite. The proof uses only the finite range of `commit : Omega -> Fin K`: the integrand is a measurable finite-valued function, hence bounded by the finite sum of the absolute values of its arm-indexed constants.","url":"../modules/banditrlproof-algorithms-etcexpectedregretassembly/index.html#decl-00ced228627a","parent":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","order":776,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCExpectedRegretAssembly.lean:162"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integrable_real_pseudoRegret_actionWithCommit_choice_of_measurable_commit {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commit : Omega -> Fin K) (r : Nat) (hmeas_commit : Measurable commit) : Integrable (fun omega : Omega => (((pseudoRegret model (ETC.actionWithCommit spec (commit omega)) (spec.explorationPulls * K + r) : Rat) : Real))) mu","missing":[],"search":"integrable_real_pseudoregret_actionwithcommit_choice_of_measurable_commit banditrlproof.etc.integrable_real_pseudoregret_actionwithcommit_choice_of_measurable_commit finite-arm `actionwithcommit` pseudo-regret has an integrable real cast when the selected commit arm is measurable and the ambient measure is finite. the proof uses only the finite range of `commit : omega -> fin k`: the integrand is a measurable finite-valued function, hence bounded by the finite sum of the absolute values of its arm-indexed constants. theorem compiled","shard":"modules/c64c4638920377a0.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob","label":"integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob","description":"Bochner/Real expected-regret assembly with a separate probability charge for each possible commit arm. The suffix term is decomposed into the finite family of measurable events `{omega | commit omega = a}`. Unlike the coarser wrong-commit wrapper below, this theorem preserves every arm gap and therefore exposes the per-arm RHS needed by the LML ETC route. It does not supply the armwise probability bounds themselves.","url":"../modules/banditrlproof-algorithms-etcexpectedregretassembly/index.html#decl-a2a67b0e6f05","parent":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","order":777,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCExpectedRegretAssembly.lean:208"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commit : Omega -> Fin K) (r : Nat) (hmeas_commit : Measurable commit) : MeasureTheory.integral mu (fun omega : Omega => (((pseudoRegret model (ETC.actionWithCommit spec (commit omega)) (spec.explorationPulls * K + r) : Rat) : Real))) <= (((((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat) : Rat)) : Rat) : Real)) + (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ((((((r : Nat) : Rat) * model.gap a : Rat) : Real))) * mu.real {omega : Omega | commit omega = a})","missing":[],"search":"integral_real_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob banditrlproof.etc.integral_real_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob bochner/real expected-regret assembly with a separate probability charge for each possible commit arm. the suffix term is decomposed into the finite family of measurable events `{omega | commit omega = a}`. unlike the coarser wrong-commit wrapper below, this theorem preserves every arm gap and therefore exposes the per-arm rhs needed by the lml etc route. it does not supply the armwise probability bounds themselves. theorem compiled","shard":"modules/c64c4638920377a0.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","label":"integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","description":"Bochner/Real expected-regret assembly for an `Omega`-indexed ETC commit selector. This is the Real-valued analogue of `lintegral_ofReal_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob`. It still consumes an abstract wrong-commit probability bound, but the conclusion is an ordinary Bochner integral of the Real-cast pseudo-regret.","url":"../modules/banditrlproof-algorithms-etcexpectedregretassembly/index.html#decl-341c5281c0d5","parent":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","order":778,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCExpectedRegretAssembly.lean:339"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commit : Omega -> Fin K) (r : Nat) (badGapBound : Rat) (pWrong : Real) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (hbadGap_nonneg : (0 : Rat) <= badGapBound) (hmeas_wrong : MeasurableSet {omega : Omega | commit omega = model.bestArm -> False}) (hprob_wrong : mu.real {omega : Omega | commit omega = model.bestArm -> False} <= pWrong) (hinteg : Integrable (fun omega : Omega => (((pseudoRegret model (ETC.actionWithCommit spec (commit omega)) (spec.explorationPulls * K + r) : Rat) : Real))) mu) : MeasureTheory.integral mu (fun omega : Omega => (((pseudoRegret model (ETC.actionWithCommit spec…","missing":[],"search":"integral_real_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_badgap_prob banditrlproof.etc.integral_real_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_badgap_prob bochner/real expected-regret assembly for an `omega`-indexed etc commit selector. this is the real-valued analogue of `lintegral_ofreal_pseudoregret_actionwithcommit_choice_le_exploration_add_suffix_badgap_prob`. it still consumes an abstract wrong-commit probability bound, but the conclusion is an ordinary bochner integral of the real-cast pseudo-regret. theorem compiled","shard":"modules/c64c4638920377a0.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.pair_map_eq_compProd_of_map_eq_of_condDistrib","label":"pair_map_eq_compProd_of_map_eq_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.pair_map_eq_compProd_of_map_eq_of_condDistrib","description":"An action marginal and the conditional law of its feedback determine the joint action/feedback law. This is the measure-level composition used by the zeroth step of an `IsAlgEnvSeq`-shaped process contract.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-c1b3c1933d73","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":779,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:43"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pair_map_eq_compProd_of_map_eq_of_condDistrib {Omega Action Feedback : Type*} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Feedback] [StandardBorelSpace Feedback] [Nonempty Feedback] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> Action) (feedback : Omega -> Feedback) (hfeedback : Measurable feedback) (p0 : Measure Action) (feedbackKernel : ProbabilityTheory.Kernel Action Feedback) [ProbabilityTheory.IsFiniteKernel feedbackKernel] (haction : Measure.map action mu = p0) (hcond : ProbabilityTheory.condDistrib feedback action mu =ᵐ[mu.map action] feedbackKernel) : Measure.map (fun omega : Omega => (action omega, feedback omega)) mu = p0 ⊗ₘ feedbackKernel","missing":[],"search":"pair_map_eq_compprod_of_map_eq_of_conddistrib banditrlproof.rewardkernel.pair_map_eq_compprod_of_map_eq_of_conddistrib an action marginal and the conditional law of its feedback determine the joint action/feedback law. this is the measure-level composition used by the zeroth step of an `isalgenvseq`-shaped process contract. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.condDistrib_pair_ae_eq_compProd_of_split","label":"condDistrib_pair_ae_eq_compProd_of_split","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.condDistrib_pair_ae_eq_compProd_of_split","description":"The successor action policy and the feedback law conditional on `(history, action)` combine into the conditional law of the observable `(action, feedback)` pair given history. The proof is the joint-law characterization of `condDistrib` followed by associativity of `compProd`.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-5945b466b328","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":780,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:72"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_pair_ae_eq_compProd_of_split {Omega History Action Feedback : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Feedback] [StandardBorelSpace Feedback] [Nonempty Feedback] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (feedback : Omega -> Feedback) (hfeedback : Measurable feedback) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] (feedbackKernel : ProbabilityTheory.Kernel (History × Action) Feedback) [ProbabilityTheory.IsMarkovKernel feedbackKernel] (hactionCond : ProbabilityTheory.condDistrib action history mu =ᵐ[mu.map history] policy) (hfeedbackCond : ProbabilityTheory.condDistrib feedback (fun omega : Omega => (his…","missing":[],"search":"conddistrib_pair_ae_eq_compprod_of_split banditrlproof.rewardkernel.conddistrib_pair_ae_eq_compprod_of_split the successor action policy and the feedback law conditional on `(history, action)` combine into the conditional law of the observable `(action, feedback)` pair given history. the proof is the joint-law characterization of `conddistrib` followed by associativity of `compprod`. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.condDistrib_ae_eq_const_of_comp","label":"condDistrib_ae_eq_const_of_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.condDistrib_ae_eq_const_of_comp","description":"A constant conditional distribution given a fine conditioning variable stays constant after projecting to any measurable coarser conditioning variable. The proof uses the defining joint-law identity for `condDistrib`: the fine joint law is a product because the kernel is constant, and mapping that product by `(project, id)` gives the corresponding coarse joint law.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-94a53e86f496","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":781,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:143"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_ae_eq_const_of_comp {Omega Fine Coarse Target : Type*} [MeasurableSpace Omega] [MeasurableSpace Fine] [MeasurableSpace Coarse] [MeasurableSpace Target] [StandardBorelSpace Target] [Nonempty Target] (mu : Measure Omega) [IsFiniteMeasure mu] (fine : Omega -> Fine) (hfine : Measurable fine) (coarse : Omega -> Coarse) (target : Omega -> Target) (htarget : Measurable target) (project : Fine -> Coarse) (hproject : Measurable project) (hcomp : coarse = project ∘ fine) (Q : Measure Target) [IsProbabilityMeasure Q] (hcond : ProbabilityTheory.condDistrib target fine mu =ᵐ[mu.map fine] ProbabilityTheory.Kernel.const Fine Q) : ProbabilityTheory.condDistrib target coarse mu =ᵐ[mu.map coarse] ProbabilityTheory.Kernel.const Coarse Q","missing":[],"search":"conddistrib_ae_eq_const_of_comp banditrlproof.rewardkernel.conddistrib_ae_eq_const_of_comp a constant conditional distribution given a fine conditioning variable stays constant after projecting to any measurable coarser conditioning variable. the proof uses the defining joint-law identity for `conddistrib`: the fine joint law is a product because the kernel is constant, and mapping that product by `(project, id)` gives the corresponding coarse joint law. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.map_eq_of_condDistrib_ae_eq_const","label":"map_eq_of_condDistrib_ae_eq_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.map_eq_of_condDistrib_ae_eq_const","description":"A constant conditional distribution determines the target marginal.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-e74632078479","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":782,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:195"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem map_eq_of_condDistrib_ae_eq_const {Omega Condition Target : Type*} [MeasurableSpace Omega] [MeasurableSpace Condition] [MeasurableSpace Target] [StandardBorelSpace Target] [Nonempty Target] (mu : Measure Omega) [IsProbabilityMeasure mu] (condition : Omega -> Condition) (hcondition : Measurable condition) (target : Omega -> Target) (htarget : Measurable target) (Q : Measure Target) [IsProbabilityMeasure Q] (hcond : ProbabilityTheory.condDistrib target condition mu =ᵐ[mu.map condition] ProbabilityTheory.Kernel.const Condition Q) : Measure.map target mu = Q","missing":[],"search":"map_eq_of_conddistrib_ae_eq_const banditrlproof.rewardkernel.map_eq_of_conddistrib_ae_eq_const a constant conditional distribution determines the target marginal. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.condDistrib_ae_eq_const_of_ae_eq_selected","label":"condDistrib_ae_eq_const_of_ae_eq_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.condDistrib_ae_eq_const_of_ae_eq_selected","description":"An action-selected conditional kernel becomes constant when the selected value is almost surely constant under the source process. The selector equality is pushed to the conditioning-variable law with `ae_map_iff`; the supplied pointwise kernel-selection identity then rewrites the conditional kernel almost everywhere.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-16ed120bb645","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":783,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:236"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_ae_eq_const_of_ae_eq_selected {Omega Fine Action Target : Type*} [MeasurableSpace Omega] [MeasurableSpace Fine] [MeasurableSpace Action] [MeasurableEq Action] [MeasurableSpace Target] [StandardBorelSpace Target] [Nonempty Target] (mu : Measure Omega) [IsFiniteMeasure mu] (fine : Omega -> Fine) (hfine : Measurable fine) (target : Omega -> Target) (selected : Fine -> Action) (hselected : Measurable selected) (kernel : ProbabilityTheory.Kernel Fine Target) (actionLaw : Action -> Measure Target) (selectedValue : Action) (hselectedValue : (fun omega : Omega => selected (fine omega)) =ᵐ[mu] fun _omega => selectedValue) (hkernel : forall value : Fine, kernel value = actionLaw (selected value)) (hcond : ProbabilityTheory.condDistrib target fine mu =ᵐ[mu.map fine] kernel) : ProbabilityTheory.condDistrib target fine mu =ᵐ[mu.map fine] ProbabilityTheory.Kernel.const Fine (actio…","missing":[],"search":"conddistrib_ae_eq_const_of_ae_eq_selected banditrlproof.rewardkernel.conddistrib_ae_eq_const_of_ae_eq_selected an action-selected conditional kernel becomes constant when the selected value is almost surely constant under the source process. the selector equality is pushed to the conditioning-variable law with `ae_map_iff`; the supplied pointwise kernel-selection identity then rewrites the conditional kernel almost everywhere. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.finiteArmCenteredRewardKernelLaw_of_hasSubgaussianMGF","label":"finiteArmCenteredRewardKernelLaw_of_hasSubgaussianMGF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.finiteArmCenteredRewardKernelLaw_of_hasSubgaussianMGF","description":"Exact model means and direct per-arm sub-Gaussian witnesses turn finite-arm probability laws into a context-independent centered reward-kernel law. Unlike `finiteArmBoundedCenteredRewardKernelLaw`, this constructor does not derive the MGF from common bounded support. The supplied proxy is shared by all arms, matching the concentration contract of the LML ETC theorem while retaining ABRL's current `Rat` reward trace.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-d02576c6e14e","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":784,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:280"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmCenteredRewardKernelLaw_of_hasSubgaussianMGF {K : Nat} {Context : Type} [MeasurableSpace Context] (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) : RewardKernel.CenteredRewardKernelLaw (RewardKernel.contextIndependentOfActionLaws (Context := Context) armLaw hprob) (fun _ arm => model.mean arm) (fun _ _ => sigma2) where","missing":[],"search":"finitearmcenteredrewardkernellaw_of_hassubgaussianmgf banditrlproof.etc.finitearmcenteredrewardkernellaw_of_hassubgaussianmgf exact model means and direct per-arm sub-gaussian witnesses turn finite-arm probability laws into a context-independent centered reward-kernel law. unlike `finitearmboundedcenteredrewardkernellaw`, this constructor does not derive the mgf from common bounded support. the supplied proxy is shared by all arms, matching the concentration contract of the lml etc theorem while retaining abrl's current `rat` reward trace. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.finiteArmBoundedCenteredRewardKernelLaw","label":"finiteArmBoundedCenteredRewardKernelLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.finiteArmBoundedCenteredRewardKernelLaw","description":"Common boundedness and exact model means turn finite-arm probability laws into a context-independent centered reward-kernel law. The variance proxy is the common Hoeffding proxy for `[lo, hi]`. Raw reward integrability follows from `Integrable.of_mem_Icc`; the centered MGF is the Mathlib-backed bounded-variable wrapper in `ConcentrationSubGaussian`.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-d16d731a7d2f","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":785,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:332"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmBoundedCenteredRewardKernelLaw {K : Nat} {Context : Type} [MeasurableSpace Context] (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) : RewardKernel.CenteredRewardKernelLaw (RewardKernel.contextIndependentOfActionLaws (Context := Context) armLaw hprob) (fun _ arm => model.mean arm) (fun _ _ => Concentration.intervalVarianceProxy lo hi) where","missing":[],"search":"finitearmboundedcenteredrewardkernellaw banditrlproof.etc.finitearmboundedcenteredrewardkernellaw common boundedness and exact model means turn finite-arm probability laws into a context-independent centered reward-kernel law. the variance proxy is the common hoeffding proxy for `[lo, hi]`. raw reward integrability follows from `integrable.of_mem_icc`; the centered mgf is the mathlib-backed bounded-variable wrapper in `concentrationsubgaussian`. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_boundedArmLaws","label":"explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_boundedArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_boundedArmLaws","description":"Canonical generated-history ETC conditional MGF from bounded finite-arm laws. Because every arm uses the same interval `[lo, hi]`, the kernel variance proxy is constant. The selected-history variance ceiling required by the generic trajectory theorem is therefore discharged by reflexivity.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-7e25c88c2b17","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":786,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:393"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_boundedArmLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (i : Nat) : let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Cont…","missing":[],"search":"explorationargmaxhistory_centeredreward_succ_hascondsubgaussianmgf_of_boundedarmlaws banditrlproof.etc.explorationargmaxhistory_centeredreward_succ_hascondsubgaussianmgf_of_boundedarmlaws canonical generated-history etc conditional mgf from bounded finite-arm laws. because every arm uses the same interval `[lo, hi]`, the kernel variance proxy is constant. the selected-history variance ceiling required by the generic trajectory theorem is therefore discharged by reflexivity. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_armLaws","label":"explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_armLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_armLaws","description":"Canonical generated-history ETC successor conditional MGF from direct finite-arm sub-Gaussian laws with a common variance proxy. The theorem removes common bounded support from the canonical kernel route. Its remaining contracts are exact model means, a shared per-arm MGF proxy, and measurable history context; the reward trace is still `Rat`-valued.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-9a03a48ad0f4","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":787,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:472"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (i : Nat) : let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Context) armLaw hprob let policy := fun t => ETC.explorationArgma…","missing":[],"search":"explorationargmaxhistory_centeredreward_succ_hascondsubgaussianmgf_of_armlaws banditrlproof.etc.explorationargmaxhistory_centeredreward_succ_hascondsubgaussianmgf_of_armlaws canonical generated-history etc successor conditional mgf from direct finite-arm sub-gaussian laws with a common variance proxy. the theorem removes common bounded support from the canonical kernel route. its remaining contracts are exact model means, a shared per-arm mgf proxy, and measurable history context; the reward trace is still `rat`-valued. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardProcess_sum_tail_ennreal_of_boundedArmLaws","label":"explorationArgmaxHistory_centeredRewardProcess_sum_tail_ennreal_of_boundedArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardProcess_sum_tail_ennreal_of_boundedArmLaws","description":"Canonical ETC Azuma-Hoeffding bound for the full centered reward sum, including the actual reward at time zero. The initial trajectory law is fixed to the law of `ETC.exploreArm spec 0`. `RewardKernel.trajMeasure_map_eval_zero` transfers its bounded centered MGF to the zeroth trajectory coordinate, while the generated-history kernel law gives the successor conditional MGF witnesses. Unlike the earlier zero-initializ…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-93b92e9b2a26","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":788,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:550"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_centeredRewardProcess_sum_tail_ennreal_of_boundedArmLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (n : Nat) {eps : Real} (heps : 0 <= eps) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfAct…","missing":[],"search":"explorationargmaxhistory_centeredrewardprocess_sum_tail_ennreal_of_boundedarmlaws banditrlproof.etc.explorationargmaxhistory_centeredrewardprocess_sum_tail_ennreal_of_boundedarmlaws canonical etc azuma-hoeffding bound for the full centered reward sum, including the actual reward at time zero. the initial trajectory law is fixed to the law of `etc.explorearm spec 0`. `rewardkernel.trajmeasure_map_eval_zero` transfers its bounded centered mgf to the zeroth trajectory coordinate, while the generated-history kernel law gives the successor conditional mgf witnesses. unlike the earlier zero-initialized sum-tail surface, this sum contains rewards at every index in `finset.range n`. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_boundedArmLaws","label":"explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_boundedArmLaws","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_boundedArmLaws","description":"Bounded finite-arm laws provide the reward-level conditional sub-Gaussian witness package used by the existing fixed-commit pairwise ETC tail route. The ambient measure is the canonical generated-history trajectory. During the exploration horizon its generated action prefix agrees with `actionWithCommit spec model.bestArm`; the corresponding shifted history filtrations therefore agree. This transports the canonical…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-6da11ff37a59","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":789,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:736"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_boundedArmLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.co…","missing":[],"search":"explorationargmaxhistory_centeredrewardcondsubgaussianwitnesses_of_boundedarmlaws banditrlproof.etc.explorationargmaxhistory_centeredrewardcondsubgaussianwitnesses_of_boundedarmlaws bounded finite-arm laws provide the reward-level conditional sub-gaussian witness package used by the existing fixed-commit pairwise etc tail route. the ambient measure is the canonical generated-history trajectory. during the exploration horizon its generated action prefix agrees with `actionwithcommit spec model.bestarm`; the corresponding shifted history filtrations therefore agree. this transports the canonical selected-reward conditional mgf to the fixed-commit witness interface without assuming coordinate independence. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_armLaws","label":"explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_armLaws","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_armLaws","description":"Direct common-sub-Gaussian arm laws provide the reward-level conditional witness package used by the fixed-commit pairwise ETC tail route. The initial witness is transported from the initial arm law. Successor witnesses come from the generated-history kernel law and are moved to the fixed `actionWithCommit` filtration using exploration-prefix action equality. No bounded-support assumption or arm union is used.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-3b732c8e7e5d","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":790,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:905"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Context) armLaw hpro…","missing":[],"search":"explorationargmaxhistory_centeredrewardcondsubgaussianwitnesses_of_armlaws banditrlproof.etc.explorationargmaxhistory_centeredrewardcondsubgaussianwitnesses_of_armlaws direct common-sub-gaussian arm laws provide the reward-level conditional witness package used by the fixed-commit pairwise etc tail route. the initial witness is transported from the initial arm law. successor witnesses come from the generated-history kernel law and are moved to the fixed `actionwithcommit` filtration using exploration-prefix action equality. no bounded-support assumption or arm union is used. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_boundedArmLaws","label":"explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_boundedArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_boundedArmLaws","description":"Canonical pairwise empirical-mean tail contract for the generated-history ETC trajectory under bounded finite-arm reward laws. This consumes the reward-level witness above through the existing centered pairwise conditional sub-Gaussian adapter. The contract matches the exact empirical means used by `explorationArgmaxCommit`.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-b6fb1e0668a6","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":791,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1058"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_boundedArmLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfAct…","missing":[],"search":"explorationargmaxhistory_pairwiseempmeantailcontract_of_boundedarmlaws banditrlproof.etc.explorationargmaxhistory_pairwiseempmeantailcontract_of_boundedarmlaws canonical pairwise empirical-mean tail contract for the generated-history etc trajectory under bounded finite-arm reward laws. this consumes the reward-level witness above through the existing centered pairwise conditional sub-gaussian adapter. the contract matches the exact empirical means used by `explorationargmaxcommit`. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_armLaws","label":"explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_armLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_armLaws","description":"Canonical pairwise empirical-mean tail contract for generated-history ETC under direct common-sub-Gaussian finite-arm laws. This is the concentration endpoint of the direct-MGF leaf. It retains the exact empirical means and fixed-commit one-sided pairwise process used by the bounded route, but replaces common support by the caller's shared `sigma2`.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-09402bb85f86","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":792,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1130"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Context) armLaw hprob let policy := fun t…","missing":[],"search":"explorationargmaxhistory_pairwiseempmeantailcontract_of_armlaws banditrlproof.etc.explorationargmaxhistory_pairwiseempmeantailcontract_of_armlaws canonical pairwise empirical-mean tail contract for generated-history etc under direct common-sub-gaussian finite-arm laws. this is the concentration endpoint of the direct-mgf leaf. it retains the exact empirical means and fixed-commit one-sided pairwise process used by the bounded route, but replaces common support by the caller's shared `sigma2`. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_armLaws","label":"explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_armLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_armLaws","description":"Canonical generated-history ETC probability bound for committing to one fixed non-best arm under direct common-sub-Gaussian finite-arm laws. The theorem consumes the direct-MGF pairwise empirical-mean contract. It keeps the single concrete commit fiber and therefore introduces no arm union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-07a354fbe5eb","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":793,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1198"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (a : Fin K) (hne : a = model.bestArm -> False) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfActionLaws…","missing":[],"search":"explorationargmaxhistory_prob_commit_eq_arm_le_pairwisetail_of_armlaws banditrlproof.etc.explorationargmaxhistory_prob_commit_eq_arm_le_pairwisetail_of_armlaws canonical generated-history etc probability bound for committing to one fixed non-best arm under direct common-sub-gaussian finite-arm laws. the theorem consumes the direct-mgf pairwise empirical-mean contract. it keeps the single concrete commit fiber and therefore introduces no arm union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_boundedArmLaws","label":"explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_boundedArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_boundedArmLaws","description":"Canonical generated-history ETC probability bound for committing to one fixed non-best arm under bounded finite-arm reward laws. The event is the actual concrete empirical-mean argmax fiber. Its probability is bounded by the matching one-sided centered pairwise tail, with no union over the other arms. This is the armwise probability source required by the gap-weighted Bochner assembly.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-dcfc0e34d57c","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":794,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1271"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_boundedArmLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (a : Fin K) (hne : a = model.bestArm -> False) : let defaultAction := ETC.exploreArm spec 0 le…","missing":[],"search":"explorationargmaxhistory_prob_commit_eq_arm_le_pairwisetail_of_boundedarmlaws banditrlproof.etc.explorationargmaxhistory_prob_commit_eq_arm_le_pairwisetail_of_boundedarmlaws canonical generated-history etc probability bound for committing to one fixed non-best arm under bounded finite-arm reward laws. the event is the actual concrete empirical-mean argmax fiber. its probability is bounded by the matching one-sided centered pairwise tail, with no union over the other arms. this is the armwise probability source required by the gap-weighted bochner assembly. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_prob_wrongCommit_le_pairwiseTailSum_of_boundedArmLaws","label":"explorationArgmaxHistory_prob_wrongCommit_le_pairwiseTailSum_of_boundedArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_prob_wrongCommit_le_pairwiseTailSum_of_boundedArmLaws","description":"Canonical generated-history ETC wrong-commit probability bound from bounded finite-arm laws. The event is the actual empirical-mean argmax commit used by the generated ETC policy. Its probability is bounded by the finite union of the canonical centered pairwise sub-Gaussian tails. This is a one-horizon, one-sided, union-bounded result; it does not yet assemble expected regret or transport to an externally supplied e…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-6f2dad93a97c","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":795,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1348"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_prob_wrongCommit_le_pairwiseTailSum_of_boundedArmLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndepend…","missing":[],"search":"explorationargmaxhistory_prob_wrongcommit_le_pairwisetailsum_of_boundedarmlaws banditrlproof.etc.explorationargmaxhistory_prob_wrongcommit_le_pairwisetailsum_of_boundedarmlaws canonical generated-history etc wrong-commit probability bound from bounded finite-arm laws. the event is the actual empirical-mean argmax commit used by the generated etc policy. its probability is bounded by the finite union of the canonical centered pairwise sub-gaussian tails. this is a one-horizon, one-sided, union-bounded result; it does not yet assemble expected regret or transport to an externally supplied environment law. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalBoundedArmWrongCommitTailBudget","label":"canonicalBoundedArmWrongCommitTailBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalBoundedArmWrongCommitTailBudget","description":"Canonical finite-union wrong-commit tail budget for common-bounded arm laws. Every reward-level variance proxy is the common interval proxy. The pairwise mask retains that proxy only when the exploration trace samples the candidate arm or the selected best arm.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-59aa4d91a786","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":796,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1421"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBoundedArmWrongCommitTailBudget {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (lo hi : Real) : ENNReal","missing":[],"search":"canonicalboundedarmwrongcommittailbudget banditrlproof.etc.canonicalboundedarmwrongcommittailbudget canonical finite-union wrong-commit tail budget for common-bounded arm laws. every reward-level variance proxy is the common interval proxy. the pairwise mask retains that proxy only when the exploration trace samples the candidate arm or the selected best arm. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalBoundedArmWrongCommitTailBudgetReal","label":"canonicalBoundedArmWrongCommitTailBudgetReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalBoundedArmWrongCommitTailBudgetReal","description":"Real view of the canonical bounded-arm wrong-commit tail budget.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-26b1a34087bd","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":797,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1433"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBoundedArmWrongCommitTailBudgetReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (lo hi : Real) : Real","missing":[],"search":"canonicalboundedarmwrongcommittailbudgetreal banditrlproof.etc.canonicalboundedarmwrongcommittailbudgetreal real view of the canonical bounded-arm wrong-commit tail budget. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalBoundedArmMaxGapIntegralRegretBoundReal","label":"canonicalBoundedArmMaxGapIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalBoundedArmMaxGapIntegralRegretBoundReal","description":"Named Real regret budget for canonical bounded-arm generated ETC. The first term pays the round-robin exploration cost. The second charges the remaining `r` rounds by `model.maxGap` times the canonical wrong-commit budget.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-ca2866a42f13","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":798,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1444"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBoundedArmMaxGapIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (r : Nat) (lo hi : Real) : Real","missing":[],"search":"canonicalboundedarmmaxgapintegralregretboundreal banditrlproof.etc.canonicalboundedarmmaxgapintegralregretboundreal named real regret budget for canonical bounded-arm generated etc. the first term pays the round-robin exploration cost. the second charges the remaining `r` rounds by `model.maxgap` times the canonical wrong-commit budget. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalBoundedArmPairwiseTailReal","label":"canonicalBoundedArmPairwiseTailReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalBoundedArmPairwiseTailReal","description":"Real view of one canonical bounded-arm centered pairwise tail.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-77299867ea91","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":799,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1457"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBoundedArmPairwiseTailReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (lo hi : Real) (a : Fin K) : Real","missing":[],"search":"canonicalboundedarmpairwisetailreal banditrlproof.etc.canonicalboundedarmpairwisetailreal real view of one canonical bounded-arm centered pairwise tail. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalBoundedArmPerArmIntegralRegretBoundReal","label":"canonicalBoundedArmPerArmIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalBoundedArmPerArmIntegralRegretBoundReal","description":"Named Real regret budget for the canonical bounded-arm per-arm ETC route. The exploration term is unchanged. The suffix keeps every arm gap paired with that arm's own canonical pairwise tail instead of collapsing the finite family to `model.maxGap` times a union probability.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-6ac240ed2e52","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":800,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1473"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBoundedArmPerArmIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (r : Nat) (lo hi : Real) : Real","missing":[],"search":"canonicalboundedarmperarmintegralregretboundreal banditrlproof.etc.canonicalboundedarmperarmintegralregretboundreal named real regret budget for the canonical bounded-arm per-arm etc route. the exploration term is unchanged. the suffix keeps every arm gap paired with that arm's own canonical pairwise tail instead of collapsing the finite family to `model.maxgap` times a union probability. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalSubGaussianArmPairwiseTailReal","label":"canonicalSubGaussianArmPairwiseTailReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalSubGaussianArmPairwiseTailReal","description":"Real view of one canonical common-sub-Gaussian arm pairwise tail.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-30b58acab470","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":801,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1486"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalSubGaussianArmPairwiseTailReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (sigma2 : NNReal) (a : Fin K) : Real","missing":[],"search":"canonicalsubgaussianarmpairwisetailreal banditrlproof.etc.canonicalsubgaussianarmpairwisetailreal real view of one canonical common-sub-gaussian arm pairwise tail. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.canonicalSubGaussianArmPerArmIntegralRegretBoundReal","label":"canonicalSubGaussianArmPerArmIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.canonicalSubGaussianArmPerArmIntegralRegretBoundReal","description":"Named Real regret budget for the canonical common-sub-Gaussian per-arm ETC route over `Rat` rewards. The suffix preserves one gap-weighted direct-MGF pairwise tail per arm and takes no finite-arm union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-e1826e4177fa","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":802,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1501"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalSubGaussianArmPerArmIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (r : Nat) (sigma2 : NNReal) : Real","missing":[],"search":"canonicalsubgaussianarmperarmintegralregretboundreal banditrlproof.etc.canonicalsubgaussianarmperarmintegralregretboundreal named real regret budget for the canonical common-sub-gaussian per-arm etc route over `rat` rewards. the suffix preserves one gap-weighted direct-mgf pairwise tail per arm and takes no finite-arm union. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_canonicalSubGaussianArmPairwiseTailReal","label":"real_measure_explorationArgmaxCommit_eq_arm_le_canonicalSubGaussianArmPairwiseTailReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_canonicalSubGaussianArmPairwiseTailReal","description":"Real-valued canonical probability bound for one non-best commit fiber under direct common-sub-Gaussian finite-arm laws. The ENNReal pairwise tail is finite because it is an `ofReal` exponential, so the conversion preserves the concrete event and introduces no arm union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-e83ae394f131","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":803,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1520"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_measure_explorationArgmaxCommit_eq_arm_le_canonicalSubGaussianArmPairwiseTailReal {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (a : Fin K) (hne : a = model.bestArm -> False) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndepen…","missing":[],"search":"real_measure_explorationargmaxcommit_eq_arm_le_canonicalsubgaussianarmpairwisetailreal banditrlproof.etc.real_measure_explorationargmaxcommit_eq_arm_le_canonicalsubgaussianarmpairwisetailreal real-valued canonical probability bound for one non-best commit fiber under direct common-sub-gaussian finite-arm laws. the ennreal pairwise tail is finite because it is an `ofreal` exponential, so the conversion preserves the concrete event and introduces no arm union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_canonicalBoundedArmPairwiseTailReal","label":"real_measure_explorationArgmaxCommit_eq_arm_le_canonicalBoundedArmPairwiseTailReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_canonicalBoundedArmPairwiseTailReal","description":"Real-valued canonical probability bound for committing to one non-best arm. The corresponding ENNReal tail is finite because it is `ENNReal.ofReal` of a real exponential, so `ENNReal.toReal_mono` transports the compiled armwise probability theorem without changing the event or taking a union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-71686b289d31","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":804,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1591"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_measure_explorationArgmaxCommit_eq_arm_le_canonicalBoundedArmPairwiseTailReal {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (a : Fin K) (hne : a = model.bestArm -> False) : let defaultAction := ETC.exploreArm spec…","missing":[],"search":"real_measure_explorationargmaxcommit_eq_arm_le_canonicalboundedarmpairwisetailreal banditrlproof.etc.real_measure_explorationargmaxcommit_eq_arm_le_canonicalboundedarmpairwisetailreal real-valued canonical probability bound for committing to one non-best arm. the corresponding ennreal tail is finite because it is `ennreal.ofreal` of a real exponential, so `ennreal.toreal_mono` transports the compiled armwise probability theorem without changing the event or taking a union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_ne_bestArm_le_canonicalBoundedArmWrongCommitTailBudgetReal","label":"real_measure_explorationArgmaxCommit_ne_bestArm_le_canonicalBoundedArmWrongCommitTailBudgetReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_measure_explorationArgmaxCommit_ne_bestArm_le_canonicalBoundedArmWrongCommitTailBudgetReal","description":"Real-valued canonical wrong-commit probability bound. The ENNReal finite tail sum is never infinite because each summand is an `ENNReal.ofReal` exponential. This wrapper converts the compiled canonical pairwise probability theorem with `ENNReal.toReal_mono`.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-347af5d5f5af","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":805,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1664"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_measure_explorationArgmaxCommit_ne_bestArm_le_canonicalBoundedArmWrongCommitTailBudgetReal {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKerne…","missing":[],"search":"real_measure_explorationargmaxcommit_ne_bestarm_le_canonicalboundedarmwrongcommittailbudgetreal banditrlproof.etc.real_measure_explorationargmaxcommit_ne_bestarm_le_canonicalboundedarmwrongcommittailbudgetreal real-valued canonical wrong-commit probability bound. the ennreal finite tail sum is never infinite because each summand is an `ennreal.ofreal` exponential. this wrapper converts the compiled canonical pairwise probability theorem with `ennreal.toreal_mono`. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal","description":"Bochner expected pseudo-regret bound for the generated finite-history ETC action under its canonical bounded-arm trajectory law. The measurable empirical-mean argmax supplies both wrong-event measurability and pseudo-regret integrability. The preceding Real probability wrapper feeds the generic pointwise-to-expectation ETC assembly with `model.maxGap` as the suffix charge. This is the canonical action-dependent kern…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-f0a108603632","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":806,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1736"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (r : Nat) : let defaultAction := ETC.exploreArm spec 0 let r…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmmaxgapintegralregretboundreal banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmmaxgapintegralregretboundreal bochner expected pseudo-regret bound for the generated finite-history etc action under its canonical bounded-arm trajectory law. the measurable empirical-mean argmax supplies both wrong-event measurability and pseudo-regret integrability. the preceding real probability wrapper feeds the generic pointwise-to-expectation etc assembly with `model.maxgap` as the suffix charge. this is the canonical action-dependent kernel theorem; transport to an externally supplied environment law remains separate. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal","description":"Canonical bounded-arm generated-history ETC expected-regret theorem with a per-arm gap-weighted suffix budget. Each non-best commit fiber uses its own canonical centered pairwise tail. The best-arm term vanishes by `gap_bestArm`, so no artificial tail hypothesis is needed there. This endpoint removes the max-gap finite-union collapse from the canonical bounded-Rat theorem while retaining the same exploration term.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-16f83701976c","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":807,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1855"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (r : Nat) : let defaultAction := ETC.exploreArm spec 0 let r…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmperarmintegralregretboundreal banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmperarmintegralregretboundreal canonical bounded-arm generated-history etc expected-regret theorem with a per-arm gap-weighted suffix budget. each non-best commit fiber uses its own canonical centered pairwise tail. the best-arm term vanishes by `gap_bestarm`, so no artificial tail hypothesis is needed there. this endpoint removes the max-gap finite-union collapse from the canonical bounded-rat theorem while retaining the same exploration term. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","description":"Canonical generated-history ETC expected-regret theorem with direct common- sub-Gaussian finite-arm laws and a per-arm gap-weighted suffix budget. Each non-best commit fiber is charged by its own direct-MGF pairwise tail. The best-arm term vanishes by `gap_bestArm`; no bounded-support assumption, max-gap collapse, or arm union is used. Rewards and model means remain `Rat`-valued.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-1fe39cf602e2","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":808,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:1980"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (r : Nat) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfAc…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalsubgaussianarmperarmintegralregretboundreal banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalsubgaussianarmperarmintegralregretboundreal canonical generated-history etc expected-regret theorem with direct common- sub-gaussian finite-arm laws and a per-arm gap-weighted suffix budget. each non-best commit fiber is charged by its own direct-mgf pairwise tail. the best-arm term vanishes by `gap_bestarm`; no bounded-support assumption, max-gap collapse, or arm union is used. rewards and model means remain `rat`-valued. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxPrefixRegretReal","label":"explorationArgmaxPrefixRegretReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxPrefixRegretReal","description":"Finite-prefix realization of the generated ETC pseudo-regret integrand. The completed trace is only an implementation device. Once the supplied history contains the complete exploration prefix, the value agrees with the original generated-history action by the factorization theorem below.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-895d96217413","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":809,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2104"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxPrefixRegretReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (t r : Nat) (history : History.FiniteRewardHistory Rat t) : Real","missing":[],"search":"explorationargmaxprefixregretreal banditrlproof.etc.explorationargmaxprefixregretreal finite-prefix realization of the generated etc pseudo-regret integrand. the completed trace is only an implementation device. once the supplied history contains the complete exploration prefix, the value agrees with the original generated-history action by the factorization theorem below. definition compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_explorationArgmaxPrefixRegretReal","label":"measurable_explorationArgmaxPrefixRegretReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_explorationArgmaxPrefixRegretReal","description":"The finite-prefix ETC pseudo-regret realization is measurable.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-fe1128862852","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":810,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2113"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_explorationArgmaxPrefixRegretReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (t r : Nat) : Measurable (ETC.explorationArgmaxPrefixRegretReal spec model t r)","missing":[],"search":"measurable_explorationargmaxprefixregretreal banditrlproof.etc.measurable_explorationargmaxprefixregretreal the finite-prefix etc pseudo-regret realization is measurable. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace","label":"explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace","description":"The finite-prefix realization agrees with canonical ETC whenever the retained history covers every exploration reward coordinate.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-630836923213","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":811,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2169"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (reward : RewardTrace Rat) (t r : Nat) (horizon_le : spec.explorationPulls * K <= t + 1) : ETC.explorationArgmaxPrefixRegretReal spec model t r (History.finiteRewardHistoryOfTrace reward t) = (((pseudoRegret model (ETC.explorationArgmaxAction spec model reward) (spec.explorationPulls * K + r) : Rat) : Real))","missing":[],"search":"explorationargmaxprefixregretreal_finiterewardhistoryoftrace banditrlproof.etc.explorationargmaxprefixregretreal_finiterewardhistoryoftrace the finite-prefix realization agrees with canonical etc whenever the retained history covers every exploration reward coordinate. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace_generated","label":"explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace_generated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace_generated","description":"Generated-history form of the finite exploration-prefix regret factorization.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-73006d385b5d","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":812,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2196"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace_generated {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (reward : RewardTrace Rat) (t r : Nat) (hexplorationPulls_pos : 0 < spec.explorationPulls) (horizon_le : spec.explorationPulls * K <= t + 1) : ETC.explorationArgmaxPrefixRegretReal spec model t r (History.finiteRewardHistoryOfTrace reward t) = (((pseudoRegret model (ETC.explorationArgmaxGeneratedAction spec model reward) (spec.explorationPulls * K + r) : Rat) : Real))","missing":[],"search":"explorationargmaxprefixregretreal_finiterewardhistoryoftrace_generated banditrlproof.etc.explorationargmaxprefixregretreal_finiterewardhistoryoftrace_generated generated-history form of the finite exploration-prefix regret factorization. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_eq_of_explorationPrefix_map_eq","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_eq_of_explorationPrefix_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_eq_of_explorationPrefix_map_eq","description":"The generated-history ETC pseudo-regret integral depends only on the finite exploration-prefix pushforward. This law-transport equality is independent of the reward-law or concentration route used later to bound either integral. It packages the measurable prefix factorization once for both max-gap and per-arm canonical endpoints.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-677ac14e0936","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":813,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2219"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_eq_of_explorationPrefix_map_eq {K : Nat} (mu nu : Measure (RewardTrace Rat)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (hexplorationPulls_pos : 0 < spec.explorationPulls) (r : Nat) (hprefix : let explorationLast := spec.explorationPulls * K - 1 Measure.map (fun trajectory : RewardTrace Rat => History.finiteRewardHistoryOfTrace trajectory explorationLast) mu = Measure.map (fun trajectory : RewardTrace Rat => History.finiteRewardHistoryOfTrace trajectory explorationLast) nu) : integral mu (fun trajectory : RewardTrace Rat => (((pseudoRegret model (ETC.explorationArgmaxGeneratedAction spec model trajectory) (spec.explorationPulls * K + r) : Rat) : Real))) = integral nu (fun trajectory : RewardTrace Rat => (((pseudoRegret model (ETC.explorationArgmaxGeneratedAction spec model trajectory) (spec.explorationPulls * K +…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_eq_of_explorationprefix_map_eq banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_eq_of_explorationprefix_map_eq the generated-history etc pseudo-regret integral depends only on the finite exploration-prefix pushforward. this law-transport equality is independent of the reward-law or concentration route used later to bound either integral. it packages the measurable prefix factorization once for both max-gap and per-arm canonical endpoints. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_explorationPrefix_map_eq","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_explorationPrefix_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_explorationPrefix_map_eq","description":"External-law ETC expected-regret theorem from exploration-prefix law equality. The external reward law need not equal the canonical Ionescu-Tulcea law on the full infinite trajectory. It is enough that their pushforwards to the `m * K` exploration rewards agree: the generated ETC action and its finite horizon pseudo-regret factor through exactly that prefix. This is the law transport surface needed by external stati…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-5f137965c1be","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":814,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2301"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_explorationPrefix_map_eq {K : Nat} {Context : Type} [MeasurableSpace Context] (mu : Measure (RewardTrace Rat)) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos :…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmmaxgapintegralregretboundreal_of_explorationprefix_map_eq banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmmaxgapintegralregretboundreal_of_explorationprefix_map_eq external-law etc expected-regret theorem from exploration-prefix law equality. the external reward law need not equal the canonical ionescu-tulcea law on the full infinite trajectory. it is enough that their pushforwards to the `m * k` exploration rewards agree: the generated etc action and its finite horizon pseudo-regret factor through exactly that prefix. this is the law transport surface needed by external stationary-bandit environment models. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","description":"External-law canonical per-arm ETC expected-regret theorem from exploration- prefix law equality. The external reward law only needs the same `m * K` exploration-prefix pushforward as the canonical generated-history trajectory. The resulting bound preserves each arm's own gap-weighted pairwise tail and takes no union over commit arms.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-bf7af72acabb","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":815,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2422"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq {K : Nat} {Context : Type} [MeasurableSpace Context] (mu : Measure (RewardTrace Rat)) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos :…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmperarmintegralregretboundreal_of_explorationprefix_map_eq banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalboundedarmperarmintegralregretboundreal_of_explorationprefix_map_eq external-law canonical per-arm etc expected-regret theorem from exploration- prefix law equality. the external reward law only needs the same `m * k` exploration-prefix pushforward as the canonical generated-history trajectory. the resulting bound preserves each arm's own gap-weighted pairwise tail and takes no union over commit arms. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","description":"External-law common-sub-Gaussian per-arm ETC expected regret from exploration- prefix law equality. Only the `m * K` exploration reward-prefix pushforward must match the canonical generated-history trajectory. The direct-MGF gap-weighted armwise budget is transported without a full trajectory law, suffix law, or arm union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-e12ef6dcfdc9","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":816,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2509"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq {K : Nat} {Context : Type} [MeasurableSpace Context] (mu : Measure (RewardTrace Rat)) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (r : Nat) (hprefix : le…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_explorationprefix_map_eq banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_explorationprefix_map_eq external-law common-sub-gaussian per-arm etc expected regret from exploration- prefix law equality. only the `m * k` exploration reward-prefix pushforward must match the canonical generated-history trajectory. the direct-mgf gap-weighted armwise budget is transported without a full trajectory law, suffix law, or arm union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_condDistrib","description":"External-process ETC expected regret from initial and successor conditional reward laws. The contract is local to the exploration prefix: coordinate measurability, the zeroth reward marginal, and the conditional distribution of reward `i + 1` given rewards through `i` for `i < m * K - 1`. The generic finite-prefix law result identifies the required pushforward with the Ionescu-Tulcea process; the preceding prefix-tr…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-b6ad6691f221","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":817,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2597"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_condDistrib {Omega : Type*} {K : Nat} {Context : Type} [MeasurableSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) :…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_initial_map_eq_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_initial_map_eq_conddistrib external-process etc expected regret from initial and successor conditional reward laws. the contract is local to the exploration prefix: coordinate measurability, the zeroth reward marginal, and the conditional distribution of reward `i + 1` given rewards through `i` for `i < m * k - 1`. the generic finite-prefix law result identifies the required pushforward with the ionescu-tulcea process; the preceding prefix-transport theorem then supplies the regret bound. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","description":"External-process canonical per-arm ETC expected regret from initial and successor conditional reward laws. The conditional laws are needed only through the exploration prefix. They identify the reward-trace prefix pushforward with the canonical trajectory; the per-arm prefix transport then preserves each gap-weighted pairwise tail without a wrong-event union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-89495cfa6d70","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":818,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2722"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib {Omega : Type*} {K : Nat} {Context : Type} [MeasurableSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) :…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_initial_map_eq_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_initial_map_eq_conddistrib external-process canonical per-arm etc expected regret from initial and successor conditional reward laws. the conditional laws are needed only through the exploration prefix. they identify the reward-trace prefix pushforward with the canonical trajectory; the per-arm prefix transport then preserves each gap-weighted pairwise tail without a wrong-event union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","description":"External-process common-sub-Gaussian per-arm ETC expected regret from an initial reward marginal and successor conditional reward laws. The conditional laws are required only through exploration. They identify the external reward-prefix law with the canonical trajectory and then invoke the direct-MGF prefix transport theorem.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-864abb8b8f87","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":819,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2846"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib {Omega : Type*} {K : Nat} {Context : Type} [MeasurableSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_initial_map_eq_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_initial_map_eq_conddistrib external-process common-sub-gaussian per-arm etc expected regret from an initial reward marginal and successor conditional reward laws. the conditional laws are required only through exploration. they identify the external reward-prefix law with the canonical trajectory and then invoke the direct-mgf prefix transport theorem. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_stepKernel_apply_eq_exploreArmLaw_of_lt","label":"explorationArgmaxHistory_stepKernel_apply_eq_exploreArmLaw_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_stepKernel_apply_eq_exploreArmLaw_of_lt","description":"During exploration, the generated-history ETC step kernel is exactly the law of the arm scheduled at time `i + 1`. The context disappears because `contextIndependentOfActionLaws` selects only the policy action, and the history state disappears because the exploration branch of the policy is deterministic.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-a4e5c758fbbd","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":820,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:2969"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_stepKernel_apply_eq_exploreArmLaw_of_lt {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (i : Nat) (hi : i + 1 < spec.explorationPulls * K) (history : (j : Finset.Iic i) -> Rat) : RewardKernel.historyStepKernelFamily (RewardKernel.contextIndependentOfActionLaws (Context := Context) armLaw hprob) (fun t => ETC.explorationArgmaxHistoryPolicy spec model t) context (fun t history => ETC.explorationArgmaxHistoryState t history) hcontext (fun t => ETC.measurable_explorationArgmaxHistoryState t) i history = armLaw (ETC.exploreArm spec (i + 1))","missing":[],"search":"explorationargmaxhistory_stepkernel_apply_eq_explorearmlaw_of_lt banditrlproof.etc.explorationargmaxhistory_stepkernel_apply_eq_explorearmlaw_of_lt during exploration, the generated-history etc step kernel is exactly the law of the arm scheduled at time `i + 1`. the context disappears because `contextindependentofactionlaws` selects only the policy action, and the history state disappears because the exploration branch of the policy is deterministic. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","description":"Practical external-process ETC expected regret from the conditional laws of the scheduled exploration arms. Unlike the preceding theorem, callers do not mention the local `historyStepKernelFamily`. They provide the initial arm law and, before the end of exploration, the conditional law of reward `i + 1` as the stationary law of `exploreArm spec (i + 1)`. The exploration step-kernel equality above turns that environm…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-62f771b4de1e","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":821,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3001"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hexplorationPulls_…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_initial_map_eq_explorationarm_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_initial_map_eq_explorationarm_conddistrib practical external-process etc expected regret from the conditional laws of the scheduled exploration arms. unlike the preceding theorem, callers do not mention the local `historystepkernelfamily`. they provide the initial arm law and, before the end of exploration, the conditional law of reward `i + 1` as the stationary law of `explorearm spec (i + 1)`. the exploration step-kernel equality above turns that environment-facing contract into the canonical conditional-law contract. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","description":"Practical external-process per-arm ETC expected regret from the conditional laws of the scheduled exploration arms. The step-kernel equality above converts the environment-facing scheduled-arm laws into the canonical conditional-law contract. The per-arm consumer then keeps the gap-weighted pairwise tails separate, without a wrong-event union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-d3cae6a24fce","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":822,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3064"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hexplorationPulls_…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_initial_map_eq_explorationarm_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_initial_map_eq_explorationarm_conddistrib practical external-process per-arm etc expected regret from the conditional laws of the scheduled exploration arms. the step-kernel equality above converts the environment-facing scheduled-arm laws into the canonical conditional-law contract. the per-arm consumer then keeps the gap-weighted pairwise tails separate, without a wrong-event union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","description":"Practical external-process common-sub-Gaussian per-arm ETC expected regret from the conditional laws of the scheduled exploration arms. The public contract exposes no local context, policy state, reward kernel, or trajectory measure. It preserves the canonical direct-MGF gap-weighted armwise budget and requires no bounded support or arm union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-543d644c8ed1","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":823,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3127"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((model.mean arm : Rat) : Real))) sigma2 (armLaw arm)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (r : Nat) (hzero…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_initial_map_eq_explorationarm_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_initial_map_eq_explorationarm_conddistrib practical external-process common-sub-gaussian per-arm etc expected regret from the conditional laws of the scheduled exploration arms. the public contract exposes no local context, policy state, reward kernel, or trajectory measure. it preserves the canonical direct-mgf gap-weighted armwise budget and requires no bounded support or arm union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","description":"External ETC expected regret from LML-shaped action/reward-history conditional laws during exploration. The initial reward law is supplied conditionally on the first action. Each successor reward has the scheduled exploration-arm law conditionally on the complete action/reward prefix together with the next action. Since those laws are constant, `RewardKernel.condDistrib_ae_eq_const_of_comp` projects them to the rewa…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-f7d73e28f6bf","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":824,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3191"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, inte…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_actionrewardhistory_explorationarm_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_actionrewardhistory_explorationarm_conddistrib external etc expected regret from lml-shaped action/reward-history conditional laws during exploration. the initial reward law is supplied conditionally on the first action. each successor reward has the scheduled exploration-arm law conditionally on the complete action/reward prefix together with the next action. since those laws are constant, `rewardkernel.conddistrib_ae_eq_const_of_comp` projects them to the reward-only prefixes consumed by the preceding environment-facing theorem. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","description":"External per-arm ETC expected regret from LML-shaped action/reward-history conditional laws during exploration. Constant scheduled-arm laws conditioned on the complete action/reward prefix and next action coarsen to reward-only prefixes. The scheduled-arm per-arm consumer then preserves the gap-weighted armwise tails without a wrong-event union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-ab9d1fe063f2","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":825,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3284"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, inte…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_actionrewardhistory_explorationarm_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_actionrewardhistory_explorationarm_conddistrib external per-arm etc expected regret from lml-shaped action/reward-history conditional laws during exploration. constant scheduled-arm laws conditioned on the complete action/reward prefix and next action coarsen to reward-only prefixes. the scheduled-arm per-arm consumer then preserves the gap-weighted armwise tails without a wrong-event union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","description":"External direct-MGF per-arm ETC expected regret from LML-shaped action/reward-history conditional laws during exploration. The initial constant conditional reward law yields the time-zero marginal. Each successor constant scheduled-arm law is coarsened from the complete action/reward prefix and next action to the reward-only prefix, then the external scheduled-arm direct-MGF theorem preserves the gap-weighted armwis…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-d9049cee32be","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":826,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3378"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - (((…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_actionrewardhistory_explorationarm_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_actionrewardhistory_explorationarm_conddistrib external direct-mgf per-arm etc expected regret from lml-shaped action/reward-history conditional laws during exploration. the initial constant conditional reward law yields the time-zero marginal. each successor constant scheduled-arm law is coarsened from the complete action/reward prefix and next action to the reward-only prefix, then the external scheduled-arm direct-mgf theorem preserves the gap-weighted armwise budget without bounded support or an arm union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","description":"External bounded ETC regret from action-dependent stationary feedback kernels and almost-sure exploration action identities. This is the dependency-light local analogue of the law transport exposed by the exact-seed LML `IsAlgEnvSeq` fields: the initial feedback kernel is indexed by action zero, each later kernel is indexed by the next action in the complete history condition, and the exploration action identities m…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-c064f5529368","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":827,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3471"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, int…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_actiondependent_actionrewardhistory_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmmaxgapintegralregretboundreal_of_actiondependent_actionrewardhistory_conddistrib external bounded etc regret from action-dependent stationary feedback kernels and almost-sure exploration action identities. this is the dependency-light local analogue of the law transport exposed by the exact-seed lml `isalgenvseq` fields: the initial feedback kernel is indexed by action zero, each later kernel is indexed by the next action in the complete history condition, and the exploration action identities make those selectors almost surely equal to the scheduled round-robin arms. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","description":"External per-arm ETC regret from action-dependent stationary feedback kernels and almost-sure exploration action identities. The selector transport turns each action-indexed feedback kernel into the constant law of the scheduled exploration arm. The full-history per-arm consumer then preserves the gap-weighted armwise tails without a wrong-event union.","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-c745119e2ebd","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":828,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3566"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => (((reward : Rat) : Real))) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi (((reward : Rat) : Real))) (ae (armLaw arm))) (hmean : forall arm, int…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_actiondependent_actionrewardhistory_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalboundedarmperarmintegralregretboundreal_of_actiondependent_actionrewardhistory_conddistrib external per-arm etc regret from action-dependent stationary feedback kernels and almost-sure exploration action identities. the selector transport turns each action-indexed feedback kernel into the constant law of the scheduled exploration arm. the full-history per-arm consumer then preserves the gap-weighted armwise tails without a wrong-event union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","label":"integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","description":"External direct-MGF per-arm ETC regret from action-dependent stationary feedback kernels and almost-sure exploration action identities. The selector transport converts the raw action-indexed initial and successor feedback kernels into the constant laws of the scheduled exploration arms. The full action/reward-history direct-MGF consumer then returns the same gap-weighted armwise budget without bounded support or an…","url":"../modules/banditrlproof-algorithms-etcfinitearmrewardlaw/index.html#decl-ea23cc603a58","parent":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","order":829,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/ETCFiniteArmRewardLaw.lean:3661"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib {Omega : Type*} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => (((reward : Rat) : Real))) = (((model.mean arm : Rat) : Real))) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - ((…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_actiondependent_actionrewardhistory_conddistrib banditrlproof.etc.integral_real_pseudoregret_explorationargmaxgeneratedaction_reward_le_canonicalsubgaussianarmperarmintegralregretboundreal_of_actiondependent_actionrewardhistory_conddistrib external direct-mgf per-arm etc regret from action-dependent stationary feedback kernels and almost-sure exploration action identities. the selector transport converts the raw action-indexed initial and successor feedback kernels into the constant laws of the scheduled exploration arms. the full action/reward-history direct-mgf consumer then returns the same gap-weighted armwise budget without bounded support or an arm union. theorem compiled","shard":"modules/298234f9f4270afd.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistoryState","label":"explorationArgmaxHistoryState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistoryState","description":"The state reconstructed from a finite reward history for the ETC commit rule. The prefix through `t` is retained exactly and the unobserved suffix is zero. The defaulted suffix is only an implementation device: score reconstruction requires the exploration horizon to be contained in the retained prefix.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-b596490e8e1d","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":830,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:27"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def explorationArgmaxHistoryState (t : Nat) (history : History.FiniteRewardHistory Rat t) : RewardTrace Rat","missing":[],"search":"explorationargmaxhistorystate banditrlproof.etc.explorationargmaxhistorystate the state reconstructed from a finite reward history for the etc commit rule. the prefix through `t` is retained exactly and the unobserved suffix is zero. the defaulted suffix is only an implementation device: score reconstruction requires the exploration horizon to be contained in the retained prefix. definition compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_explorationArgmaxHistoryState","label":"measurable_explorationArgmaxHistoryState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_explorationArgmaxHistoryState","description":"The finite-history ETC state reconstruction is measurable as a trace.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-4ea8e1daa1ef","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":831,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:33"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_explorationArgmaxHistoryState (t : Nat) : Measurable (explorationArgmaxHistoryState t)","missing":[],"search":"measurable_explorationargmaxhistorystate banditrlproof.etc.measurable_explorationargmaxhistorystate the finite-history etc state reconstruction is measurable as a trace. theorem compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistoryPolicy","label":"explorationArgmaxHistoryPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistoryPolicy","description":"At policy time `t`, explore at action time `t + 1` while that action lies in the exploration prefix; otherwise choose the empirical-mean argmax computed from the completed finite-history state.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-cdde50b4df60","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":832,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:51"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxHistoryPolicy {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (t : Nat) : Policy.MeasurablePolicy (RewardTrace Rat) (Fin K) where","missing":[],"search":"explorationargmaxhistorypolicy banditrlproof.etc.explorationargmaxhistorypolicy at policy time `t`, explore at action time `t + 1` while that action lies in the exploration prefix; otherwise choose the empirical-mean argmax computed from the completed finite-history state. definition compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedAction","label":"explorationArgmaxGeneratedAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxGeneratedAction","description":"The canonical ETC action generator obtained by running the finite-history policy against the identity reward trace.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-8c8778149f75","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":833,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:77"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxGeneratedAction {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) : RewardTrace Rat -> ActionTrace (Fin K)","missing":[],"search":"explorationargmaxgeneratedaction banditrlproof.etc.explorationargmaxgeneratedaction the canonical etc action generator obtained by running the finite-history policy against the identity reward trace. definition compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedAction_eq_explorationArgmaxAction","label":"explorationArgmaxGeneratedAction_eq_explorationArgmaxAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxGeneratedAction_eq_explorationArgmaxAction","description":"The canonical empirical-mean ETC trace is generated by the measurable policy over finite reward histories. Positive exploration pulls are essential at time zero: the shifted generated trace starts with `ETC.exploreArm spec 0`, which matches the ETC trace only when time zero belongs to the exploration prefix. At later commit times, the finite-history score reconstruction theorem supplies the argmax equality.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-64c9b8720715","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":834,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:95"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxGeneratedAction_eq_explorationArgmaxAction {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : ETC.explorationArgmaxGeneratedAction spec model = ETC.explorationArgmaxAction spec model","missing":[],"search":"explorationargmaxgeneratedaction_eq_explorationargmaxaction banditrlproof.etc.explorationargmaxgeneratedaction_eq_explorationargmaxaction the canonical empirical-mean etc trace is generated by the measurable policy over finite reward histories. positive exploration pulls are essential at time zero: the shifted generated trace starts with `etc.explorearm spec 0`, which matches the etc trace only when time zero belongs to the exploration prefix. at later commit times, the finite-history score reconstruction theorem supplies the argmax equality. theorem compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedAction_eq_actionWithCommit_of_lt","label":"explorationArgmaxGeneratedAction_eq_actionWithCommit_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxGeneratedAction_eq_actionWithCommit_of_lt","description":"During the exploration prefix, the generated finite-history ETC policy agrees with every fixed-commit ETC trace because both choose the same round-robin arm. The post-exploration commit arm is deliberately arbitrary in this statement; only the strict exploration-horizon hypothesis is used.","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-7d1bec4fa8c9","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":835,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:143"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxGeneratedAction_eq_actionWithCommit_of_lt {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : RewardTrace Rat) {t : Nat} (hexplorationPulls_pos : 0 < spec.explorationPulls) (ht : t < spec.explorationPulls * K) : ETC.explorationArgmaxGeneratedAction spec model reward t = ETC.actionWithCommit spec commitArm t","missing":[],"search":"explorationargmaxgeneratedaction_eq_actionwithcommit_of_lt banditrlproof.etc.explorationargmaxgeneratedaction_eq_actionwithcommit_of_lt during the exploration prefix, the generated finite-history etc policy agrees with every fixed-commit etc trace because both choose the same round-robin arm. the post-exploration commit arm is deliberately arbitrary in this statement; only the strict exploration-horizon hypothesis is used. theorem compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedActionPartialTrajectoryPairLawSource_trajMeasure","label":"explorationArgmaxGeneratedActionPartialTrajectoryPairLawSource_trajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxGeneratedActionPartialTrajectoryPairLawSource_trajMeasure","description":"Canonical kernel-trajectory law source for the finite-history ETC policy. The reward process is the identity coordinate process under the `RewardKernel.historyStepKernelFamily` trajectory measure. The existing canonical `trajMeasure` construction proves the full finite-pair `partialTraj` law for the generated policy action without assuming an ambient selected-reward law. Combined with `explorationArgmaxGeneratedActi…","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-73931ed3247e","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":836,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:180"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxGeneratedActionPartialTrajectoryPairLawSource_trajMeasure {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (mu0 : MeasureTheory.Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) : ConditionalExpectationReward.GeneratedActionPartialTrajectoryPairLawSource (ProbabilityTheory.Kernel.trajMeasure (X := fun _ : Nat => Rat) mu0 (RewardKernel.historyStepKernelFamily rewardKernel (fun t => ETC.explorationArgmaxHistoryPolicy spec model t) context (fun t history => ETC.explorationArgmaxHistoryState t history) hcontext (fun t => ETC.measurable_explorationArgmaxHistoryState t))) rewardKernel (fun t => ETC.explorationArg…","missing":[],"search":"explorationargmaxgeneratedactionpartialtrajectorypairlawsource_trajmeasure banditrlproof.etc.explorationargmaxgeneratedactionpartialtrajectorypairlawsource_trajmeasure canonical kernel-trajectory law source for the finite-history etc policy. the reward process is the identity coordinate process under the `rewardkernel.historystepkernelfamily` trajectory measure. the existing canonical `trajmeasure` construction proves the full finite-pair `partialtraj` law for the generated policy action without assuming an ambient selected-reward law. combined with `explorationargmaxgeneratedaction_eq_explorationargmaxaction`, this law is the action-dependent probability foundation for the canonical etc action trace. it remains a kernel-generated trajectory law. transport to an arbitrary adaptive environment and identification with the fixed product-coordinate source remain separate obligations. definition compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","label":"explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","description":"Canonical trajectory conditional sub-Gaussian MGF for the successor reward centered at the finite-bandit model mean selected by the ETC policy. The mean surface is context-independent, `fun _ action => model.mean action`. The kernel centered-reward law supplies conditional mean-zero and the one-step MGF; the remaining regularity contract is the selected finite-history variance ceiling at the requested successor time…","url":"../modules/banditrlproof-algorithms-etcgeneratedhistorypolicy/index.html#decl-d01e12467438","parent":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","order":837,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy"],["Source","BanditRLProof/Algorithms/ETCGeneratedHistoryPolicy.lean:227"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (mu0 : MeasureTheory.Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (varianceProxy : Context -> Fin K -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel (fun _ action => model.mean action) varianceProxy) (i : Nat) (c : NNReal) (hvariance : forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((ETC.explorationArgmaxHistoryPolicy spec model i).action (ETC.explorationArgmaxHistoryState i history)) <= c) : let policy := fun t => ETC.explorationArgmaxHistoryPolicy sp…","missing":[],"search":"explorationargmaxhistory_centeredreward_succ_hascondsubgaussianmgf_trajmeasure banditrlproof.etc.explorationargmaxhistory_centeredreward_succ_hascondsubgaussianmgf_trajmeasure canonical trajectory conditional sub-gaussian mgf for the successor reward centered at the finite-bandit model mean selected by the etc policy. the mean surface is context-independent, `fun _ action => model.mean action`. the kernel centered-reward law supplies conditional mean-zero and the one-step mgf; the remaining regularity contract is the selected finite-history variance ceiling at the requested successor time. this theorem is therefore a concentration foundation for the canonical kernel trajectory, not a transport to the fixed product-coordinate reward source or an arbitrary adaptive bandit environment. theorem compiled","shard":"modules/057f6f9dc38b5ea3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductArgmaxCommit","label":"fixedProductArgmaxCommit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductArgmaxCommit","description":"The finite argmax commit arm selected from the fixed-commit exploration sample. This is only a naming wrapper around the existing argmax commit oracle and `ETC.empMeanAtExploration`; it does not add a new probability or filtration assumption.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-a5809b5fd862","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":838,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:26"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductArgmaxCommit {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (omega : RewardTrace Rat) : Fin K","missing":[],"search":"fixedproductargmaxcommit banditrlproof.etc.fixedproductargmaxcommit the finite argmax commit arm selected from the fixed-commit exploration sample. this is only a naming wrapper around the existing argmax commit oracle and `etc.empmeanatexploration`; it does not add a new probability or filtration assumption. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductArgmaxAction","label":"fixedProductArgmaxAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductArgmaxAction","description":"The ETC action trace that explores with `baseCommitArm` and then commits to the finite argmax empirical-mean arm selected from that exploration sample.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-2cd1b0a3518c","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":839,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:40"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductArgmaxAction {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (omega : RewardTrace Rat) : ActionTrace (Fin K)","missing":[],"search":"fixedproductargmaxaction banditrlproof.etc.fixedproductargmaxaction the etc action trace that explores with `basecommitarm` and then commits to the finite argmax empirical-mean arm selected from that exploration sample. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductMaxGapLintegralRegretBound","label":"fixedProductMaxGapLintegralRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductMaxGapLintegralRegretBound","description":"Named RHS budget for the fixed product-coordinate max-gap lower-integral ETC regret wrapper. This is still an `ENNReal.ofReal` surrogate budget. It does not claim Bochner or Rat-valued expected regret.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-bed57ee90657","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":840,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:56"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductMaxGapLintegralRegretBound {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) : ENNReal","missing":[],"search":"fixedproductmaxgaplintegralregretbound banditrlproof.etc.fixedproductmaxgaplintegralregretbound named rhs budget for the fixed product-coordinate max-gap lower-integral etc regret wrapper. this is still an `ennreal.ofreal` surrogate budget. it does not claim bochner or rat-valued expected regret. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductWrongCommitTailBudget","label":"fixedProductWrongCommitTailBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductWrongCommitTailBudget","description":"Named ENNReal wrong-commit tail budget used by the fixed product-coordinate ETC argmax route.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-e25b88c1c57a","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":841,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:79"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductWrongCommitTailBudget {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (lo hi : Fin K -> Nat -> Real) : ENNReal","missing":[],"search":"fixedproductwrongcommittailbudget banditrlproof.etc.fixedproductwrongcommittailbudget named ennreal wrong-commit tail budget used by the fixed product-coordinate etc argmax route. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductBadGapLintegralRegretBound","label":"fixedProductBadGapLintegralRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductBadGapLintegralRegretBound","description":"Named RHS budget for the fixed product-coordinate bad-gap lower-integral ETC regret wrapper. This is still an `ENNReal.ofReal` surrogate budget. It packages the explicit bad-gap suffix contract behind the same fixed-product tail budget used by the Bochner/Real bad-gap wrapper.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-adbf56c9cca6","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":842,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:99"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductBadGapLintegralRegretBound {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (badGapBound : Rat) (lo hi : Fin K -> Nat -> Real) : ENNReal","missing":[],"search":"fixedproductbadgaplintegralregretbound banditrlproof.etc.fixedproductbadgaplintegralregretbound named rhs budget for the fixed product-coordinate bad-gap lower-integral etc regret wrapper. this is still an `ennreal.ofreal` surrogate budget. it packages the explicit bad-gap suffix contract behind the same fixed-product tail budget used by the bochner/real bad-gap wrapper. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductSumGapLintegralRegretBound","label":"fixedProductSumGapLintegralRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductSumGapLintegralRegretBound","description":"Named RHS budget for the conservative fixed product-coordinate sum-gap lower-integral ETC regret wrapper. This is the sum-gap specialization of `fixedProductBadGapLintegralRegretBound`; it removes the explicit `badGapBound` parameter while remaining on the `ENNReal.ofReal` lower-integral surface.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-0402ebfe7226","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":843,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:124"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductSumGapLintegralRegretBound {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) : ENNReal","missing":[],"search":"fixedproductsumgaplintegralregretbound banditrlproof.etc.fixedproductsumgaplintegralregretbound named rhs budget for the conservative fixed product-coordinate sum-gap lower-integral etc regret wrapper. this is the sum-gap specialization of `fixedproductbadgaplintegralregretbound`; it removes the explicit `badgapbound` parameter while remaining on the `ennreal.ofreal` lower-integral surface. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductWrongCommitTailBudgetReal","label":"fixedProductWrongCommitTailBudgetReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductWrongCommitTailBudgetReal","description":"Real-valued view of the fixed product-coordinate wrong-commit tail budget.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-642d30ceadde","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":844,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:140"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductWrongCommitTailBudgetReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (lo hi : Fin K -> Nat -> Real) : Real","missing":[],"search":"fixedproductwrongcommittailbudgetreal banditrlproof.etc.fixedproductwrongcommittailbudgetreal real-valued view of the fixed product-coordinate wrong-commit tail budget. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductBadGapIntegralRegretBoundReal","label":"fixedProductBadGapIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductBadGapIntegralRegretBoundReal","description":"Named Real RHS for the fixed product-coordinate bad-gap Bochner ETC regret assembly.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-ba5c6bd3a6a8","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":845,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:152"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductBadGapIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (badGapBound : Rat) (lo hi : Fin K -> Nat -> Real) : Real","missing":[],"search":"fixedproductbadgapintegralregretboundreal banditrlproof.etc.fixedproductbadgapintegralregretboundreal named real rhs for the fixed product-coordinate bad-gap bochner etc regret assembly. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductSumGapIntegralRegretBoundReal","label":"fixedProductSumGapIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductSumGapIntegralRegretBoundReal","description":"Named Real RHS for the conservative fixed product-coordinate sum-gap Bochner ETC regret wrapper. This is the sum-gap specialization of `fixedProductBadGapIntegralRegretBoundReal`; it removes the explicit `badGapBound` parameter by using the total finite sum of model gaps as a conservative non-best suffix gap bound.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-9fc110571d6f","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":846,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:179"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductSumGapIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) : Real","missing":[],"search":"fixedproductsumgapintegralregretboundreal banditrlproof.etc.fixedproductsumgapintegralregretboundreal named real rhs for the conservative fixed product-coordinate sum-gap bochner etc regret wrapper. this is the sum-gap specialization of `fixedproductbadgapintegralregretboundreal`; it removes the explicit `badgapbound` parameter by using the total finite sum of model gaps as a conservative non-best suffix gap bound. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.fixedProductMaxGapIntegralRegretBoundReal","label":"fixedProductMaxGapIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.fixedProductMaxGapIntegralRegretBoundReal","description":"Named Real RHS for the fixed product-coordinate max-gap Bochner ETC regret wrapper. This is the max-gap specialization of `fixedProductBadGapIntegralRegretBoundReal`; it removes the explicit `badGapBound` parameter while keeping the same finite-product wrong-commit tail budget.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-c31c7a0304eb","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":847,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:201"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedProductMaxGapIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) : Real","missing":[],"search":"fixedproductmaxgapintegralregretboundreal banditrlproof.etc.fixedproductmaxgapintegralregretboundreal named real rhs for the fixed product-coordinate max-gap bochner etc regret wrapper. this is the max-gap specialization of `fixedproductbadgapintegralregretboundreal`; it removes the explicit `badgapbound` parameter while keeping the same finite-product wrong-commit tail budget. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxCommit","label":"explorationArgmaxCommit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxCommit","description":"The canonical fixed-product ETC commit arm uses the round-robin exploration prefix only. Internally it fixes the already-selected best arm as a harmless post-exploration seed; every public source contract below is stated directly against `ETC.exploreArm`.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-795e8e46895b","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":848,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:217"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxCommit {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (omega : RewardTrace Rat) : Fin K","missing":[],"search":"explorationargmaxcommit banditrlproof.etc.explorationargmaxcommit the canonical fixed-product etc commit arm uses the round-robin exploration prefix only. internally it fixes the already-selected best arm as a harmless post-exploration seed; every public source contract below is stated directly against `etc.explorearm`. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationArgmaxAction","label":"explorationArgmaxAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationArgmaxAction","description":"The canonical fixed-product ETC trace: round-robin exploration followed by the argmax commit computed from that exploration prefix.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-7d082db2b121","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":849,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:228"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationArgmaxAction {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (omega : RewardTrace Rat) : ActionTrace (Fin K)","missing":[],"search":"explorationargmaxaction banditrlproof.etc.explorationargmaxaction the canonical fixed-product etc trace: round-robin exploration followed by the argmax commit computed from that exploration prefix. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.explorationMaxGapIntegralRegretBoundReal","label":"explorationMaxGapIntegralRegretBoundReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.explorationMaxGapIntegralRegretBoundReal","description":"Real-valued max-gap regret budget for the canonical round-robin exploration fixed-product ETC endpoint.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-9a3e6a84f273","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":850,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:239"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def explorationMaxGapIntegralRegretBoundReal {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (r : Nat) (lo hi : Fin K -> Nat -> Real) : Real","missing":[],"search":"explorationmaxgapintegralregretboundreal banditrlproof.etc.explorationmaxgapintegralregretboundreal real-valued max-gap regret budget for the canonical round-robin exploration fixed-product etc endpoint. definition compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","label":"lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","description":"Concrete lower-integral ETC regret assembly for the finite argmax commit oracle under an infinite product reward-coordinate law. The empirical means are computed from the fixed exploration trace determined by `baseCommitArm`; the post-exploration action trace commits to the argmax oracle choice computed from those empirical means. The suffix term is charged by an explicit non-best gap bound and the compiled infinite…","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-9b897d6c7809","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":851,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:258"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (badGapBound : Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.inte…","missing":[],"search":"lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_badgap_prob_of_infinitepi_bounded_actionmean banditrlproof.etc.lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_badgap_prob_of_infinitepi_bounded_actionmean concrete lower-integral etc regret assembly for the finite argmax commit oracle under an infinite product reward-coordinate law. the empirical means are computed from the fixed exploration trace determined by `basecommitarm`; the post-exploration action trace commits to the argmax oracle choice computed from those empirical means. the suffix term is charged by an explicit non-best gap bound and the compiled infinite-product wrong-commit probability bound. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapLintegralRegretBound_of_infinitePi_bounded_actionMean","label":"lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapLintegralRegretBound_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapLintegralRegretBound_of_infinitePi_bounded_actionMean","description":"Polished fixed-product bad-gap lower-integral ETC regret wrapper. The statement names both the fixed-product argmax action trace and the bad-gap lower-integral RHS budget while remaining on the same fixed product-coordinate, fixed-exploration `ENNReal.ofReal` surface as the concrete assembly theorem.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-352c115ebbe0","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":852,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:378"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapLintegralRegretBound_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (badGapBound : Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (c…","missing":[],"search":"lintegral_ofreal_pseudoregret_fixedproductargmaxaction_le_fixedproductbadgaplintegralregretbound_of_infinitepi_bounded_actionmean banditrlproof.etc.lintegral_ofreal_pseudoregret_fixedproductargmaxaction_le_fixedproductbadgaplintegralregretbound_of_infinitepi_bounded_actionmean polished fixed-product bad-gap lower-integral etc regret wrapper. the statement names both the fixed-product argmax action trace and the bad-gap lower-integral rhs budget while remaining on the same fixed product-coordinate, fixed-exploration `ennreal.ofreal` surface as the concrete assembly theorem. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapLintegralRegretBound_of_infinitePi_bounded_actionMean","label":"lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapLintegralRegretBound_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapLintegralRegretBound_of_infinitePi_bounded_actionMean","description":"Polished fixed-product sum-gap lower-integral ETC regret wrapper. The statement names both the fixed-product argmax action trace and the conservative sum-gap lower-integral RHS budget, while remaining on the fixed product-coordinate, fixed-exploration `ENNReal.ofReal` surface.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-81658c59f585","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":853,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:438"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapLintegralRegretBound_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommit sp…","missing":[],"search":"lintegral_ofreal_pseudoregret_fixedproductargmaxaction_le_fixedproductsumgaplintegralregretbound_of_infinitepi_bounded_actionmean banditrlproof.etc.lintegral_ofreal_pseudoregret_fixedproductargmaxaction_le_fixedproductsumgaplintegralregretbound_of_infinitepi_bounded_actionmean polished fixed-product sum-gap lower-integral etc regret wrapper. the statement names both the fixed-product argmax action trace and the conservative sum-gap lower-integral rhs budget, while remaining on the fixed product-coordinate, fixed-exploration `ennreal.ofreal` surface. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_measure_fixedProductArgmaxCommit_ne_bestArm_le_fixedProductWrongCommitTailBudgetReal_of_infinitePi_bounded_actionMean","label":"real_measure_fixedProductArgmaxCommit_ne_bestArm_le_fixedProductWrongCommitTailBudgetReal_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_measure_fixedProductArgmaxCommit_ne_bestArm_le_fixedProductWrongCommitTailBudgetReal_of_infinitePi_bounded_actionMean","description":"Real-valued fixed-product wrong-commit probability bound. This exposes the `ENNReal.toReal` conversion used by the Bochner expected-regret assembly as a standalone reusable surface. It remains tied to the concrete infinite product-coordinate reward source and fixed exploration trace.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-594c3fe81aee","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":854,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:501"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_measure_fixedProductArgmaxCommit_ne_bestArm_le_fixedProductWrongCommitTailBudgetReal_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommit spec baseCommitArm…","missing":[],"search":"real_measure_fixedproductargmaxcommit_ne_bestarm_le_fixedproductwrongcommittailbudgetreal_of_infinitepi_bounded_actionmean banditrlproof.etc.real_measure_fixedproductargmaxcommit_ne_bestarm_le_fixedproductwrongcommittailbudgetreal_of_infinitepi_bounded_actionmean real-valued fixed-product wrong-commit probability bound. this exposes the `ennreal.toreal` conversion used by the bochner expected-regret assembly as a standalone reusable surface. it remains tied to the concrete infinite product-coordinate reward source and fixed exploration trace. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","label":"integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","description":"Concrete Bochner/Real ETC regret assembly for the finite argmax commit oracle under an infinite product reward-coordinate law. This is the Real-valued counterpart of `lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean`. It uses the same infinite-product wrong-commit probability source, converts that probability budget with `ENNReal…","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-04c75cfa9435","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":855,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:567"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (badGapBound : Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (hbadGap_nonneg : (0 : Rat) <= badGapBound) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.e…","missing":[],"search":"integral_real_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_badgap_prob_of_infinitepi_bounded_actionmean banditrlproof.etc.integral_real_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_badgap_prob_of_infinitepi_bounded_actionmean concrete bochner/real etc regret assembly for the finite argmax commit oracle under an infinite product reward-coordinate law. this is the real-valued counterpart of `lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_badgap_prob_of_infinitepi_bounded_actionmean`. it uses the same infinite-product wrong-commit probability source, converts that probability budget with `ennreal.toreal`, and discharges the abstract integrability side condition from the measurable finite-valued commit selector. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","label":"integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","description":"Polished fixed-product bad-gap Bochner/Real ETC regret wrapper. The underlying concrete theorem already uses the named fixed-product argmax action in its conclusion; this wrapper gives that endpoint the same API shape as the sum-gap and max-gap Bochner wrappers.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-3b65c599886b","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":856,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:705"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (badGapBound : Rat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (hbadGap_nonneg : (0 : Rat) <= badGapBound) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explo…","missing":[],"search":"integral_real_pseudoregret_fixedproductargmaxaction_le_fixedproductbadgapintegralregretboundreal_of_infinitepi_bounded_actionmean banditrlproof.etc.integral_real_pseudoregret_fixedproductargmaxaction_le_fixedproductbadgapintegralregretboundreal_of_infinitepi_bounded_actionmean polished fixed-product bad-gap bochner/real etc regret wrapper. the underlying concrete theorem already uses the named fixed-product argmax action in its conclusion; this wrapper gives that endpoint the same api shape as the sum-gap and max-gap bochner wrappers. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","label":"integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","description":"Conservative concrete Bochner/Real ETC regret assembly using the total sum of model gaps as the suffix bad-gap bound. This is the Real-valued counterpart of the existing sum-gap `ENNReal.ofReal` lower-integral adapter. It removes the explicit `badGapBound` and `hbadGap` contracts from the concrete fixed-product expected-regret theorem by using gap nonnegativity over the finite arm set.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-3f5e03bcc9e2","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":857,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:766"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommi…","missing":[],"search":"integral_real_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_sumgap_prob_of_infinitepi_bounded_actionmean banditrlproof.etc.integral_real_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_sumgap_prob_of_infinitepi_bounded_actionmean conservative concrete bochner/real etc regret assembly using the total sum of model gaps as the suffix bad-gap bound. this is the real-valued counterpart of the existing sum-gap `ennreal.ofreal` lower-integral adapter. it removes the explicit `badgapbound` and `hbadgap` contracts from the concrete fixed-product expected-regret theorem by using gap nonnegativity over the finite arm set. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","label":"integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","description":"Polished fixed-product sum-gap Bochner/Real ETC regret wrapper. The statement names both the argmax-commit action trace and the conservative Real-valued sum-gap RHS budget, while remaining on the fixed product-coordinate, fixed-exploration surface.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-29fc32de1d69","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":858,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:832"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommit sp…","missing":[],"search":"integral_real_pseudoregret_fixedproductargmaxaction_le_fixedproductsumgapintegralregretboundreal_of_infinitepi_bounded_actionmean banditrlproof.etc.integral_real_pseudoregret_fixedproductargmaxaction_le_fixedproductsumgapintegralregretboundreal_of_infinitepi_bounded_actionmean polished fixed-product sum-gap bochner/real etc regret wrapper. the statement names both the argmax-commit action trace and the conservative real-valued sum-gap rhs budget, while remaining on the fixed product-coordinate, fixed-exploration surface. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","label":"integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","description":"Sharper concrete Bochner/Real ETC regret assembly using `FiniteBanditModel.maxGap` as the suffix bad-gap bound. This is the Real-valued counterpart of the existing max-gap `ENNReal.ofReal` lower-integral adapter. It removes the explicit `badGapBound` and `hbadGap` contracts from the concrete fixed-product expected-regret theorem by using the compiled finite-model max-gap invariants.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-9099c9e42fa9","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":859,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:885"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommi…","missing":[],"search":"integral_real_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_maxgap_prob_of_infinitepi_bounded_actionmean banditrlproof.etc.integral_real_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_maxgap_prob_of_infinitepi_bounded_actionmean sharper concrete bochner/real etc regret assembly using `finitebanditmodel.maxgap` as the suffix bad-gap bound. this is the real-valued counterpart of the existing max-gap `ennreal.ofreal` lower-integral adapter. it removes the explicit `badgapbound` and `hbadgap` contracts from the concrete fixed-product expected-regret theorem by using the compiled finite-model max-gap invariants. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","label":"integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","description":"Polished fixed-product max-gap Bochner/Real ETC regret wrapper. The statement names both the argmax-commit action trace and the Real-valued RHS budget, while remaining on the fixed product-coordinate, fixed-exploration surface.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-eae2c6694772","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":860,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:945"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommit sp…","missing":[],"search":"integral_real_pseudoregret_fixedproductargmaxaction_le_fixedproductmaxgapintegralregretboundreal_of_infinitepi_bounded_actionmean banditrlproof.etc.integral_real_pseudoregret_fixedproductargmaxaction_le_fixedproductmaxgapintegralregretboundreal_of_infinitepi_bounded_actionmean polished fixed-product max-gap bochner/real etc regret wrapper. the statement names both the argmax-commit action trace and the real-valued rhs budget, while remaining on the fixed product-coordinate, fixed-exploration surface. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxAction_le_explorationMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_exploreMean","label":"integral_real_pseudoRegret_explorationArgmaxAction_le_explorationMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_exploreMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxAction_le_explorationMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_exploreMean","description":"Canonical fixed-product Bochner/Real ETC expected-regret theorem stated only with round-robin exploration coordinates. Unlike the lower-level fixed-product wrapper, callers do not supply an arbitrary `baseCommitArm`: every coordinate bound and mean identity is indexed by `ETC.exploreArm`. The result remains a fixed-product/fixed-exploration endpoint and does not construct an adaptive environment or policy law.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-fe38949cb421","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":861,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:998"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_explorationArgmaxAction_le_explorationMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_exploreMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.exploreArm spec t) t) (hi (ETC.exploreArm spec t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.exploreArm spec t) : Rat) : Real))) : MeasureTheory.integral (MeasureTheory.Measure.…","missing":[],"search":"integral_real_pseudoregret_explorationargmaxaction_le_explorationmaxgapintegralregretboundreal_of_infinitepi_bounded_exploremean banditrlproof.etc.integral_real_pseudoregret_explorationargmaxaction_le_explorationmaxgapintegralregretboundreal_of_infinitepi_bounded_exploremean canonical fixed-product bochner/real etc expected-regret theorem stated only with round-robin exploration coordinates. unlike the lower-level fixed-product wrapper, callers do not supply an arbitrary `basecommitarm`: every coordinate bound and mean identity is indexed by `etc.explorearm`. the result remains a fixed-product/fixed-exploration endpoint and does not construct an adaptive environment or policy law. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","label":"lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","description":"Conservative concrete lower-integral ETC regret assembly using the total sum of model gaps as the suffix bad-gap bound. This removes the explicit `badGapBound`/`hbadGap` contract from the concrete infinite-product wrapper. The price is a looser suffix constant: `sum_a model.gap a` bounds every non-best gap by nonnegativity.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-a642012cc9b8","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":862,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:1061"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCo…","missing":[],"search":"lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_sumgap_prob_of_infinitepi_bounded_actionmean banditrlproof.etc.lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_sumgap_prob_of_infinitepi_bounded_actionmean conservative concrete lower-integral etc regret assembly using the total sum of model gaps as the suffix bad-gap bound. this removes the explicit `badgapbound`/`hbadgap` contract from the concrete infinite-product wrapper. the price is a looser suffix constant: `sum_a model.gap a` bounds every non-best gap by nonnegativity. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","label":"lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","description":"Sharper concrete lower-integral ETC regret assembly using `FiniteBanditModel.maxGap` as the suffix bad-gap bound. This keeps the same infinite-product wrong-commit probability term as the sum-gap wrapper, but charges each wrong suffix pull by the maximum local gap instead of the total sum of all gaps.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-44a09bbf52e0","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":863,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:1136"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCo…","missing":[],"search":"lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_maxgap_prob_of_infinitepi_bounded_actionmean banditrlproof.etc.lintegral_ofreal_pseudoregret_argmaxcommitoracle_actionwithcommit_le_exploration_add_suffix_maxgap_prob_of_infinitepi_bounded_actionmean sharper concrete lower-integral etc regret assembly using `finitebanditmodel.maxgap` as the suffix bad-gap bound. this keeps the same infinite-product wrong-commit probability term as the sum-gap wrapper, but charges each wrong suffix pull by the maximum local gap instead of the total sum of all gaps. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapLintegralRegretBound_of_infinitePi_bounded_actionMean","label":"lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapLintegralRegretBound_of_infinitePi_bounded_actionMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapLintegralRegretBound_of_infinitePi_bounded_actionMean","description":"Polished fixed-product max-gap lower-integral ETC regret wrapper. The statement names both the argmax-commit action trace and the RHS max-gap budget, while remaining exactly on the same fixed product-coordinate, fixed-exploration, `ENNReal.ofReal` surface as the concrete max-gap assembly.","url":"../modules/banditrlproof-algorithms-etcinfinitepiexpectedregretassembly/index.html#decl-c5e479522149","parent":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","order":864,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCInfinitePiExpectedRegretAssembly.lean:1204"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapLintegralRegretBound_of_infinitePi_bounded_actionMean {K : Nat} (coordLaw : Nat -> MeasureTheory.Measure Rat) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] (spec : ETC.Spec K) (model : FiniteBanditModel K) (baseCommitArm : Fin K) (r : Nat) (lo hi : Fin K -> Nat -> Real) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_coord_bound : forall t, t < spec.explorationPulls * K -> Filter.Eventually (fun rewardValue : Rat => Set.Icc (lo (ETC.actionWithCommit spec baseCommitArm t) t) (hi (ETC.actionWithCommit spec baseCommitArm t) t) (((rewardValue : Rat) : Real))) (MeasureTheory.ae (coordLaw t))) (h_coord_mean : forall t, t < spec.explorationPulls * K -> MeasureTheory.integral (coordLaw t) (fun rewardValue : Rat => (((rewardValue : Rat) : Real))) = (((model.mean (ETC.actionWithCommit sp…","missing":[],"search":"lintegral_ofreal_pseudoregret_fixedproductargmaxaction_le_fixedproductmaxgaplintegralregretbound_of_infinitepi_bounded_actionmean banditrlproof.etc.lintegral_ofreal_pseudoregret_fixedproductargmaxaction_le_fixedproductmaxgaplintegralregretbound_of_infinitepi_bounded_actionmean polished fixed-product max-gap lower-integral etc regret wrapper. the statement names both the argmax-commit action trace and the rhs max-gap budget, while remaining exactly on the same fixed product-coordinate, fixed-exploration, `ennreal.ofreal` surface as the concrete max-gap assembly. theorem compiled","shard":"modules/69436358672ab94d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurableSet_commitArm_ne_bestArm","label":"measurableSet_commitArm_ne_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurableSet_commitArm_ne_bestArm","description":"If the commit arm is measurable, then the event that it is not the selected best arm is measurable. This is the first wrong-commit event canary after the deterministic fixed-commit ETC layer. It is intentionally only an event measurability fact.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-57addc76fede","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":865,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:31"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_commitArm_ne_bestArm {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (hmeas_commit : Measurable commitArm) : MeasurableSet {omega : Omega | commitArm omega = model.bestArm -> False}","missing":[],"search":"measurableset_commitarm_ne_bestarm banditrlproof.etc.measurableset_commitarm_ne_bestarm if the commit arm is measurable, then the event that it is not the selected best arm is measurable. this is the first wrong-commit event canary after the deterministic fixed-commit etc layer. it is intentionally only an event measurability fact. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_empMeanVector_of_forall_measurable","label":"measurable_empMeanVector_of_forall_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_empMeanVector_of_forall_measurable","description":"Coordinatewise empirical-mean measurability packages into measurability of the empirical-mean score vector. This is the `ETC-EMPMEAN-VECTOR-MEASURABILITY-BRIDGE` project-local leaf selected by the Extended Pro review after `ETC-COMMIT-ORACLE-CHOICE-MEASURABILITY-BRIDGE`. It is a direct producer for `measurable_commitOracle_choose_of_measurable_empMeanVector`; it does not construct an oracle, prove argmax correctness…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-2070504e0652","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":866,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:55"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_empMeanVector_of_forall_measurable {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace Rat] (empMean : Omega -> Fin K -> Rat) (hmeas_coord : forall a : Fin K, Measurable (fun omega : Omega => empMean omega a)) : Measurable (fun omega : Omega => (empMean omega : Fin K -> Rat))","missing":[],"search":"measurable_empmeanvector_of_forall_measurable banditrlproof.etc.measurable_empmeanvector_of_forall_measurable coordinatewise empirical-mean measurability packages into measurability of the empirical-mean score vector. this is the `etc-empmean-vector-measurability-bridge` project-local leaf selected by the extended pro review after `etc-commit-oracle-choice-measurability-bridge`. it is a direct producer for `measurable_commitoracle_choose_of_measurable_empmeanvector`; it does not construct an oracle, prove argmax correctness, add concentration, or introduce filtration. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_commitOracle_choose_of_measurable_empMeanVector","label":"measurable_commitOracle_choose_of_measurable_empMeanVector","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_commitOracle_choose_of_measurable_empMeanVector","description":"An abstract commit oracle has a measurable composed choice map whenever the empirical-mean score vector is measurable and the finite-score-vector domain is countable with measurable singletons. This is the compiled candidate identified by the `ETC-COMMIT-ORACLE-CHOICE-MEASURABILITY-ROUTE-CARD`. It discharges the direct `hmeas_choose` contract used by `measurableSet_commitOracle_ne_bestArm` for the current `Rat` scor…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-98c8a8cb05cd","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":867,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:76"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_commitOracle_choose_of_measurable_empMeanVector {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSpace (Fin K -> Rat)] [MeasurableSingletonClass (Fin K -> Rat)] [Countable (Fin K -> Rat)] (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (hmeas_emp : Measurable (fun omega : Omega => (empMean omega : Fin K -> Rat))) : Measurable (fun omega : Omega => oracle.choose (empMean omega))","missing":[],"search":"measurable_commitoracle_choose_of_measurable_empmeanvector banditrlproof.etc.measurable_commitoracle_choose_of_measurable_empmeanvector an abstract commit oracle has a measurable composed choice map whenever the empirical-mean score vector is measurable and the finite-score-vector domain is countable with measurable singletons. this is the compiled candidate identified by the `etc-commit-oracle-choice-measurability-route-card`. it discharges the direct `hmeas_choose` contract used by `measurableset_commitoracle_ne_bestarm` for the current `rat` score-vector surface. it does not construct a concrete argmax oracle, prove argmax correctness, introduce concentration, or use filtration. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_commitOracle_choose_of_forall_measurable_empMean","label":"measurable_commitOracle_choose_of_forall_measurable_empMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_commitOracle_choose_of_forall_measurable_empMean","description":"Coordinatewise empirical-mean measurability is enough to make an abstract commit oracle's composed choice map measurable. This is the `ETC-COMMIT-ORACLE-CHOICE-MEASURABILITY-OF-COORDINATES` project-local leaf selected after `ETC-EMPMEAN-VECTOR-MEASURABILITY-BRIDGE`. It composes the Mathlib Pi-space empirical-mean vector bridge with the countable score-vector oracle-choice bridge. It does not construct a concrete ora…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-4bbd51b6053f","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":868,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:104"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_commitOracle_choose_of_forall_measurable_empMean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace Rat] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K -> Rat)] [Countable (Fin K -> Rat)] (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (hmeas_coord : forall a : Fin K, Measurable (fun omega : Omega => empMean omega a)) : Measurable (fun omega : Omega => oracle.choose (empMean omega))","missing":[],"search":"measurable_commitoracle_choose_of_forall_measurable_empmean banditrlproof.etc.measurable_commitoracle_choose_of_forall_measurable_empmean coordinatewise empirical-mean measurability is enough to make an abstract commit oracle's composed choice map measurable. this is the `etc-commit-oracle-choice-measurability-of-coordinates` project-local leaf selected after `etc-empmean-vector-measurability-bridge`. it composes the mathlib pi-space empirical-mean vector bridge with the countable score-vector oracle-choice bridge. it does not construct a concrete oracle, prove argmax correctness, introduce probability, or use concentration/filtration facts. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurableSet_commitOracle_ne_bestArm","label":"measurableSet_commitOracle_ne_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurableSet_commitOracle_ne_bestArm","description":"If an abstract commit oracle's composed choice map is measurable, then the event that the oracle-selected arm is not the selected best arm is measurable. This is the `ETC-COMMIT-ORACLE-WRONG-EVENT-MEASURABILITY` project-local leaf. It deliberately assumes measurability of the composed oracle choice directly, so it does not construct a concrete argmax oracle, prove oracle measurability from empirical means, add proba…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-32182a044742","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":869,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:136"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_commitOracle_ne_bestArm {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (model : FiniteBanditModel K) (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (hmeas_choose : Measurable (fun omega : Omega => oracle.choose (empMean omega))) : MeasurableSet {omega : Omega | oracle.choose (empMean omega) = model.bestArm -> False}","missing":[],"search":"measurableset_commitoracle_ne_bestarm banditrlproof.etc.measurableset_commitoracle_ne_bestarm if an abstract commit oracle's composed choice map is measurable, then the event that the oracle-selected arm is not the selected best arm is measurable. this is the `etc-commit-oracle-wrong-event-measurability` project-local leaf. it deliberately assumes measurability of the composed oracle choice directly, so it does not construct a concrete argmax oracle, prove oracle measurability from empirical means, add probability assumptions, or introduce concentration and filtration obligations. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurableSet_commitOracle_ne_bestArm_of_forall_measurable_empMean","label":"measurableSet_commitOracle_ne_bestArm_of_forall_measurable_empMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurableSet_commitOracle_ne_bestArm_of_forall_measurable_empMean","description":"Coordinatewise empirical-mean measurability is enough to make the oracle-selected wrong-commit event measurable. This is the `ETC-COMMIT-ORACLE-WRONG-EVENT-MEASURABILITY-OF-COORDINATES` project-local leaf. It composes the coordinatewise oracle-choice measurability wrapper with the oracle-selected wrong-event measurability wrapper. It does not construct a concrete oracle, prove argmax correctness, introduce a probabi…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-4780de10be9c","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":870,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:165"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_commitOracle_ne_bestArm_of_forall_measurable_empMean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace Rat] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSingletonClass (Fin K -> Rat)] [Countable (Fin K -> Rat)] (model : FiniteBanditModel K) (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (hmeas_coord : forall a : Fin K, Measurable (fun omega : Omega => empMean omega a)) : MeasurableSet {omega : Omega | oracle.choose (empMean omega) = model.bestArm -> False}","missing":[],"search":"measurableset_commitoracle_ne_bestarm_of_forall_measurable_empmean banditrlproof.etc.measurableset_commitoracle_ne_bestarm_of_forall_measurable_empmean coordinatewise empirical-mean measurability is enough to make the oracle-selected wrong-commit event measurable. this is the `etc-commit-oracle-wrong-event-measurability-of-coordinates` project-local leaf. it composes the coordinatewise oracle-choice measurability wrapper with the oracle-selected wrong-event measurability wrapper. it does not construct a concrete oracle, prove argmax correctness, introduce a probability measure, or use concentration/filtration facts. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurableSet_empMean_ge_empMean","label":"measurableSet_empMean_ge_empMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurableSet_empMean_ge_empMean","description":"Pairwise empirical-mean comparison events are measurable when each empirical mean coordinate is measurable. This is an ordered-event regularity canary for the wrong-mean event. It does not use measures, probability, commit-arm argmax, finite unions, concentration, or filtration assumptions.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-f7503d04a706","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":871,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:202"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_empMean_ge_empMean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (empMean : Omega -> Fin K -> Rat) (hmeas_empMean : forall a : Fin K, Measurable (fun omega : Omega => empMean omega a)) (a b : Fin K) : MeasurableSet {omega : Omega | empMean omega a >= empMean omega b}","missing":[],"search":"measurableset_empmean_ge_empmean banditrlproof.etc.measurableset_empmean_ge_empmean pairwise empirical-mean comparison events are measurable when each empirical mean coordinate is measurable. this is an ordered-event regularity canary for the wrong-mean event. it does not use measures, probability, commit-arm argmax, finite unions, concentration, or filtration assumptions. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurableSet_exists_ne_bestArm_empMean_ge_bestArm","label":"measurableSet_exists_ne_bestArm_empMean_ge_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurableSet_exists_ne_bestArm_empMean_ge_bestArm","description":"The finite existential wrong-mean event is measurable when each empirical mean coordinate is measurable. This packages the pairwise empirical-mean comparison canary into the exact finite event shape used by the wrong-commit probability bridge. It does not use measures, probability, commit-arm argmax, concentration, filtration, or an empirical-mean construction.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-fedda5f3a6ee","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":872,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:222"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_exists_ne_bestArm_empMean_ge_bestArm {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (model : FiniteBanditModel K) (empMean : Omega -> Fin K -> Rat) (hmeas_empMean : forall a : Fin K, Measurable (fun omega : Omega => empMean omega a)) : MeasurableSet {omega : Omega | exists a : Fin K, (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm}","missing":[],"search":"measurableset_exists_ne_bestarm_empmean_ge_bestarm banditrlproof.etc.measurableset_exists_ne_bestarm_empmean_ge_bestarm the finite existential wrong-mean event is measurable when each empirical mean coordinate is measurable. this packages the pairwise empirical-mean comparison canary into the exact finite event shape used by the wrong-commit probability bridge. it does not use measures, probability, commit-arm argmax, concentration, filtration, or an empirical-mean construction. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm","label":"wrong_commit_subset_exists_empMean_ge_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm","description":"The wrong-commit event is contained in the event that some non-best arm's empirical mean beats or ties the selected best arm's empirical mean, assuming the commit arm is an empirical-mean argmax. This is a pure event-reduction leaf: no measurability, measure, probability, or concentration assumptions are used.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-9f3adc20cdb1","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":873,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:270"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem wrong_commit_subset_exists_empMean_ge_bestArm {Omega : Type u} {K : Nat} (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (empMean : Omega -> Fin K -> Rat) (hcommit_argmax : forall omega : Omega, forall a : Fin K, empMean omega a <= empMean omega (commitArm omega)) : Set.Subset {omega : Omega | commitArm omega = model.bestArm -> False} {omega : Omega | exists a : Fin K, (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm}","missing":[],"search":"wrong_commit_subset_exists_empmean_ge_bestarm banditrlproof.etc.wrong_commit_subset_exists_empmean_ge_bestarm the wrong-commit event is contained in the event that some non-best arm's empirical mean beats or ties the selected best arm's empirical mean, assuming the commit arm is an empirical-mean argmax. this is a pure event-reduction leaf: no measurability, measure, probability, or concentration assumptions are used. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm_of_commitOracle","label":"wrong_commit_subset_exists_empMean_ge_bestArm_of_commitOracle","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm_of_commitOracle","description":"A commit oracle with an explicit argmax certificate satisfies the deterministic wrong-commit event reduction when used as the commit-arm selector. This is the `ETC-COMMIT-ORACLE-ARGMAX-CONSUMER` project-local leaf. It consumes only the oracle's abstract argmax contract and the compiled set-inclusion event reduction; it does not prove oracle optimality, oracle measurability, concentration, filtration, or final ETC re…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-6828cafe5bd5","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":874,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:298"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem wrong_commit_subset_exists_empMean_ge_bestArm_of_commitOracle {Omega : Type u} {K : Nat} (model : FiniteBanditModel K) (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (hchoose_argmax : forall scores : Fin K -> Rat, forall a : Fin K, scores a <= scores (oracle.choose scores)) : Set.Subset {omega : Omega | oracle.choose (empMean omega) = model.bestArm -> False} {omega : Omega | exists a : Fin K, (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm}","missing":[],"search":"wrong_commit_subset_exists_empmean_ge_bestarm_of_commitoracle banditrlproof.etc.wrong_commit_subset_exists_empmean_ge_bestarm_of_commitoracle a commit oracle with an explicit argmax certificate satisfies the deterministic wrong-commit event reduction when used as the commit-arm selector. this is the `etc-commit-oracle-argmax-consumer` project-local leaf. it consumes only the oracle's abstract argmax contract and the compiled set-inclusion event reduction; it does not prove oracle optimality, oracle measurability, concentration, filtration, or final etc regret. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_wrong_mean_events_of_subset","label":"prob_commitArm_ne_bestArm_le_wrong_mean_events_of_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_wrong_mean_events_of_subset","description":"The measure of the wrong-commit event is bounded by the measure of the empirical wrong-mean event, using only the compiled set inclusion and measure monotonicity. This is a probability-facing wrapper leaf, but it does not require a probability measure, event measurability, empirical-mean measurability, concentration, or filtration assumptions.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-8bce62abbcee","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":875,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:330"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitArm_ne_bestArm_le_wrong_mean_events_of_subset {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (empMean : Omega -> Fin K -> Rat) (hcommit_argmax : forall omega : Omega, forall a : Fin K, empMean omega a <= empMean omega (commitArm omega)) : mu {omega : Omega | commitArm omega = model.bestArm -> False} <= mu {omega : Omega | exists a : Fin K, (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm}","missing":[],"search":"prob_commitarm_ne_bestarm_le_wrong_mean_events_of_subset banditrlproof.etc.prob_commitarm_ne_bestarm_le_wrong_mean_events_of_subset the measure of the wrong-commit event is bounded by the measure of the empirical wrong-mean event, using only the compiled set inclusion and measure monotonicity. this is a probability-facing wrapper leaf, but it does not require a probability measure, event measurability, empirical-mean measurability, concentration, or filtration assumptions. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_exists_ne_bestArm_empMean_ge_bestArm_le_sum","label":"prob_exists_ne_bestArm_empMean_ge_bestArm_le_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_exists_ne_bestArm_empMean_ge_bestArm_le_sum","description":"The finite existential wrong-mean event is bounded by the finite sum of its guarded pairwise arm events. This is an outer-measure finite-union wrapper. It intentionally does not require event measurability, a probability measure, empirical-mean measurability, commit-arm argmax, concentration, filtration, or an empirical-mean construction.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-d28c82ad1ee5","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":876,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:357"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_exists_ne_bestArm_empMean_ge_bestArm_le_sum {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (empMean : Omega -> Fin K -> Rat) : mu {omega : Omega | exists a : Fin K, (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm} <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => mu {omega : Omega | (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm})","missing":[],"search":"prob_exists_ne_bestarm_empmean_ge_bestarm_le_sum banditrlproof.etc.prob_exists_ne_bestarm_empmean_ge_bestarm_le_sum the finite existential wrong-mean event is bounded by the finite sum of its guarded pairwise arm events. this is an outer-measure finite-union wrapper. it intentionally does not require event measurability, a probability measure, empirical-mean measurability, commit-arm argmax, concentration, filtration, or an empirical-mean construction. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_wrong_mean_events","label":"prob_commitArm_ne_bestArm_le_sum_wrong_mean_events","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_wrong_mean_events","description":"The wrong-commit event is bounded by the finite sum of guarded pairwise wrong-mean event measures. This is the terminal elementary probability assembly for the wrong-commit event reduction. It composes the compiled set-inclusion measure wrapper with the compiled finite-union probability wrapper, without adding empirical-mean construction, concentration, filtration, or final regret assumptions.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-c0d520e026a6","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":877,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:398"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitArm_ne_bestArm_le_sum_wrong_mean_events {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (empMean : Omega -> Fin K -> Rat) (hcommit_argmax : forall omega : Omega, forall a : Fin K, empMean omega a <= empMean omega (commitArm omega)) : mu {omega : Omega | commitArm omega = model.bestArm -> False} <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => mu {omega : Omega | (a = model.bestArm -> False) /\\ empMean omega a >= empMean omega model.bestArm})","missing":[],"search":"prob_commitarm_ne_bestarm_le_sum_wrong_mean_events banditrlproof.etc.prob_commitarm_ne_bestarm_le_sum_wrong_mean_events the wrong-commit event is bounded by the finite sum of guarded pairwise wrong-mean event measures. this is the terminal elementary probability assembly for the wrong-commit event reduction. it composes the compiled set-inclusion measure wrapper with the compiled finite-union probability wrapper, without adding empirical-mean construction, concentration, filtration, or final regret assumptions. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_pairwise_tail","label":"prob_commitArm_ne_bestArm_le_sum_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_pairwise_tail","description":"The wrong-commit event is bounded by a finite sum of abstract pairwise tail bounds. This keeps the current ETC probability layer free of concentration, filtration, independence, and empirical-mean construction assumptions while exposing the interface that future Hoeffding-style leaves can discharge.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-2ce5f6f6db61","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":878,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:427"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitArm_ne_bestArm_le_sum_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hcommit_argmax : forall omega : Omega, forall a : Fin K, empMean omega a <= empMean omega (commitArm omega)) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | commitArm omega = model.bestArm -> False} <= (Finset.univ : Finset (Fin K)).sum tail","missing":[],"search":"prob_commitarm_ne_bestarm_le_sum_pairwise_tail banditrlproof.etc.prob_commitarm_ne_bestarm_le_sum_pairwise_tail the wrong-commit event is bounded by a finite sum of abstract pairwise tail bounds. this keeps the current etc probability layer free of concentration, filtration, independence, and empirical-mean construction assumptions while exposing the interface that future hoeffding-style leaves can discharge. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_sum_pairwise_tail","label":"prob_commitOracle_ne_bestArm_le_sum_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_sum_pairwise_tail","description":"The oracle-selected wrong-commit event is bounded by a finite sum of abstract pairwise tail bounds. This is the `ETC-COMMIT-ORACLE-PROB-WRAPPER` project-local leaf. It specializes the arbitrary-commit-arm pairwise-tail consumer to `oracle.choose (empMean omega)` and derives the needed commit-arm argmax contract from the abstract oracle certificate. It does not prove oracle measurability, construct a concrete argmax…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-ead22e13c601","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":879,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:487"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitOracle_ne_bestArm_le_sum_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hchoose_argmax : forall scores : Fin K -> Rat, forall a : Fin K, scores a <= scores (oracle.choose scores)) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | oracle.choose (empMean omega) = model.bestArm -> False} <= (Finset.univ : Finset (Fin K)).sum tail","missing":[],"search":"prob_commitoracle_ne_bestarm_le_sum_pairwise_tail banditrlproof.etc.prob_commitoracle_ne_bestarm_le_sum_pairwise_tail the oracle-selected wrong-commit event is bounded by a finite sum of abstract pairwise tail bounds. this is the `etc-commit-oracle-prob-wrapper` project-local leaf. it specializes the arbitrary-commit-arm pairwise-tail consumer to `oracle.choose (empmean omega)` and derives the needed commit-arm argmax contract from the abstract oracle certificate. it does not prove oracle measurability, construct a concrete argmax oracle, prove concentration, add filtration, or instantiate final etc regret. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_nonbest_pairwise_tail","label":"prob_commitArm_ne_bestArm_le_sum_nonbest_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_nonbest_pairwise_tail","description":"The wrong-commit event is bounded by a finite sum of abstract pairwise tail bounds with the selected best-arm summand forced to zero. This is a sharper tail-consumer wrapper than `prob_commitArm_ne_bestArm_le_sum_pairwise_tail`; it still avoids filtered-sum normalization, empirical-mean construction, and concentration assumptions.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-f3ff3b489128","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":880,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:525"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitArm_ne_bestArm_le_sum_nonbest_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hcommit_argmax : forall omega : Omega, forall a : Fin K, empMean omega a <= empMean omega (commitArm omega)) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | commitArm omega = model.bestArm -> False} <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => if a = model.bestArm then 0 else tail a)","missing":[],"search":"prob_commitarm_ne_bestarm_le_sum_nonbest_pairwise_tail banditrlproof.etc.prob_commitarm_ne_bestarm_le_sum_nonbest_pairwise_tail the wrong-commit event is bounded by a finite sum of abstract pairwise tail bounds with the selected best-arm summand forced to zero. this is a sharper tail-consumer wrapper than `prob_commitarm_ne_bestarm_le_sum_pairwise_tail`; it still avoids filtered-sum normalization, empirical-mean construction, and concentration assumptions. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_filtered_sum_pairwise_tail","label":"prob_commitArm_ne_bestArm_le_filtered_sum_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_filtered_sum_pairwise_tail","description":"The wrong-commit event is bounded by the filtered finite sum of abstract non-best pairwise tail bounds. This is a presentation-normalization wrapper around `prob_commitArm_ne_bestArm_le_sum_nonbest_pairwise_tail`: it replaces the if-zeroed `Finset.univ` sum by an explicit filtered sum over non-best arms.","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-e9198a1f8838","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":881,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:589"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitArm_ne_bestArm_le_filtered_sum_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (commitArm : Omega -> Fin K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hcommit_argmax : forall omega : Omega, forall a : Fin K, empMean omega a <= empMean omega (commitArm omega)) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | commitArm omega = model.bestArm -> False} <= ((Finset.univ : Finset (Fin K)).filter (fun a : Fin K => a = model.bestArm -> False)).sum tail","missing":[],"search":"prob_commitarm_ne_bestarm_le_filtered_sum_pairwise_tail banditrlproof.etc.prob_commitarm_ne_bestarm_le_filtered_sum_pairwise_tail the wrong-commit event is bounded by the filtered finite sum of abstract non-best pairwise tail bounds. this is a presentation-normalization wrapper around `prob_commitarm_ne_bestarm_le_sum_nonbest_pairwise_tail`: it replaces the if-zeroed `finset.univ` sum by an explicit filtered sum over non-best arms. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","label":"prob_commitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","description":"The oracle-selected wrong-commit event is bounded by the filtered finite sum of abstract non-best pairwise tail bounds. This is the `ETC-COMMIT-ORACLE-FILTERED-SUM-PAIRWISE-TAIL` project-local leaf. It specializes the arbitrary-commit-arm filtered probability consumer to `oracle.choose (empMean omega)` and derives the needed commit-arm argmax contract from the abstract oracle certificate. It does not add oracle meas…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-a85976c6c12b","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":882,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:639"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitOracle_ne_bestArm_le_filtered_sum_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hchoose_argmax : forall scores : Fin K -> Rat, forall a : Fin K, scores a <= scores (oracle.choose scores)) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | oracle.choose (empMean omega) = model.bestArm -> False} <= ((Finset.univ : Finset (Fin K)).filter (fun a : Fin K => a = model.bestArm -> False)).sum tail","missing":[],"search":"prob_commitoracle_ne_bestarm_le_filtered_sum_pairwise_tail banditrlproof.etc.prob_commitoracle_ne_bestarm_le_filtered_sum_pairwise_tail the oracle-selected wrong-commit event is bounded by the filtered finite sum of abstract non-best pairwise tail bounds. this is the `etc-commit-oracle-filtered-sum-pairwise-tail` project-local leaf. it specializes the arbitrary-commit-arm filtered probability consumer to `oracle.choose (empmean omega)` and derives the needed commit-arm argmax contract from the abstract oracle certificate. it does not add oracle measurability, a concrete argmax oracle, concentration, filtration, or final etc regret. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_sum_nonbest_pairwise_tail","label":"prob_commitOracle_ne_bestArm_le_sum_nonbest_pairwise_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_sum_nonbest_pairwise_tail","description":"The oracle-selected wrong-commit event is bounded by the if-zeroed finite sum of abstract non-best pairwise tail bounds. This is the `ETC-COMMIT-ORACLE-NONBEST-PAIRWISE-TAIL` project-local leaf. It specializes the arbitrary-commit-arm if-zeroed probability consumer to `oracle.choose (empMean omega)` and derives the needed commit-arm argmax contract from the abstract oracle certificate. It does not add oracle measura…","url":"../modules/banditrlproof-algorithms-etcmeasurability/index.html#decl-f2c7819c4242","parent":"module:BanditRLProof.Algorithms.ETCMeasurability","order":883,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCMeasurability"],["Source","BanditRLProof/Algorithms/ETCMeasurability.lean:681"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_commitOracle_ne_bestArm_le_sum_nonbest_pairwise_tail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (oracle : ETC.CommitOracle K) (empMean : Omega -> Fin K -> Rat) (tail : Fin K -> ENNReal) (hchoose_argmax : forall scores : Fin K -> Rat, forall a : Fin K, scores a <= scores (oracle.choose scores)) (hpair_tail : forall a : Fin K, (a = model.bestArm -> False) -> mu {omega : Omega | empMean omega a >= empMean omega model.bestArm} <= tail a) : mu {omega : Omega | oracle.choose (empMean omega) = model.bestArm -> False} <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => if a = model.bestArm then 0 else tail a)","missing":[],"search":"prob_commitoracle_ne_bestarm_le_sum_nonbest_pairwise_tail banditrlproof.etc.prob_commitoracle_ne_bestarm_le_sum_nonbest_pairwise_tail the oracle-selected wrong-commit event is bounded by the if-zeroed finite sum of abstract non-best pairwise tail bounds. this is the `etc-commit-oracle-nonbest-pairwise-tail` project-local leaf. it specializes the arbitrary-commit-arm if-zeroed probability consumer to `oracle.choose (empmean omega)` and derives the needed commit-arm argmax contract from the abstract oracle certificate. it does not add oracle measurability, a concrete argmax oracle, concentration, filtration, or final etc regret. theorem compiled","shard":"modules/26c597b04de5215d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centered_subGaussian_event_bounds","label":"pairwiseEmpMeanTailContract_of_centered_subGaussian_event_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centered_subGaussian_event_bounds","description":"Build the fixed-commit ETC pairwise empirical-mean tail contract from concrete centered pairwise reward-difference sub-Gaussian witnesses over the ETC exploration horizon. This is the `ETC-PAIRWISE-TAIL-PRODUCER-CENTERED-DIFF` bridge. It instantiates the abstract producer with `idx := Finset.range (spec.explorationPulls * K)`, `X := ETC.centeredPairwiseRewardDiff`, and `eps := ETC.centeredPairwiseGapThreshold`. It s…","url":"../modules/banditrlproof-algorithms-etcpairwisecenteredsubgaussiantail/index.html#decl-c9e5fd8687a4","parent":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","order":884,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCPairwiseCenteredSubGaussianTail.lean:30"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_centered_subGaussian_event_bounds {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) (c : Fin K -> Nat -> NNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_indep : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) mu) (h_subG : forall a : Fin K, (a = model.bestArm -> False) -> forall t, t ∈ Finset.range (spec.explorationPulls * K) -> ProbabilityTheory.HasSubgaussianMGF (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) (c a t) mu) (htail : forall a : Fin K, (a = model.bestArm -> False) -> ENNR…","missing":[],"search":"pairwiseempmeantailcontract_of_centered_subgaussian_event_bounds banditrlproof.etc.pairwiseempmeantailcontract_of_centered_subgaussian_event_bounds build the fixed-commit etc pairwise empirical-mean tail contract from concrete centered pairwise reward-difference sub-gaussian witnesses over the etc exploration horizon. this is the `etc-pairwise-tail-producer-centered-diff` bridge. it instantiates the abstract producer with `idx := finset.range (spec.explorationpulls * k)`, `x := etc.centeredpairwiserewarddiff`, and `eps := etc.centeredpairwisegapthreshold`. it still does not prove independence, sub-gaussianity of reward differences, filtration, or final etc regret. theorem compiled","shard":"modules/44db1ed1ab6221d5.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_subGaussian_event_bounds","label":"pairwiseEmpMeanTailContract_of_subGaussian_event_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_subGaussian_event_bounds","description":"Build the fixed-commit ETC pairwise empirical-mean tail contract from abstract independent sub-Gaussian sum witnesses for every non-best arm. This is the `ETC-PAIRWISE-TAIL-PRODUCER-SUBGAUSS` bridge. The event inclusion field is intentionally explicit: later leaves must prove that the ETC empirical-mean comparison event is contained in the corresponding centered reward-difference sum event.","url":"../modules/banditrlproof-algorithms-etcpairwisesubgaussiantail/index.html#decl-79c199369e01","parent":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","order":885,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail"],["Source","BanditRLProof/Algorithms/ETCPairwiseSubGaussianTail.lean:25"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pairwiseEmpMeanTailContract_of_subGaussian_event_bounds {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) {Idx : Type v} (idx : Finset Idx) (X : Fin K -> Idx -> Omega -> Real) (c : Fin K -> Idx -> NNReal) (eps : Fin K -> Real) (h_indep : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (X a) mu) (h_subG : forall a : Fin K, (a = model.bestArm -> False) -> forall i, i ∈ idx -> ProbabilityTheory.HasSubgaussianMGF ((X a) i) ((c a) i) mu) (heps : forall a : Fin K, (a = model.bestArm -> False) -> 0 <= eps a) (hsubset : forall a : Fin K, (a = model.bestArm -> False) -> Set.Subset {omega : Omega | ETC.empMeanAtExploration spec commitArm (reward omega) a >= ET…","missing":[],"search":"pairwiseempmeantailcontract_of_subgaussian_event_bounds banditrlproof.etc.pairwiseempmeantailcontract_of_subgaussian_event_bounds build the fixed-commit etc pairwise empirical-mean tail contract from abstract independent sub-gaussian sum witnesses for every non-best arm. this is the `etc-pairwise-tail-producer-subgauss` bridge. the event inclusion field is intentionally explicit: later leaves must prove that the etc empirical-mean comparison event is contained in the corresponding centered reward-difference sum event. theorem compiled","shard":"modules/5f7b0c14eb92f5a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.PairwiseEmpMeanTailContract","label":"PairwiseEmpMeanTailContract","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.PairwiseEmpMeanTailContract","description":"Abstract non-best pairwise tail contract for fixed-commit ETC empirical means. This is the `ETC-PAIRWISE-TAIL-CONTRACT-SURFACE` leaf. It records exactly the `hpair_tail` shape needed by the concrete argmax filtered-sum probability wrapper after instantiating `empMean` with `ETC.empMeanAtExploration`.","url":"../modules/banditrlproof-algorithms-etcpairwisetailcontract/index.html#decl-1c734d9e4821","parent":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","order":886,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETCPairwiseTailContract"],["Source","BanditRLProof/Algorithms/ETCPairwiseTailContract.lean:24"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure PairwiseEmpMeanTailContract {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) : Prop where","missing":[],"search":"pairwiseempmeantailcontract banditrlproof.etc.pairwiseempmeantailcontract abstract non-best pairwise tail contract for fixed-commit etc empirical means. this is the `etc-pairwise-tail-contract-surface` leaf. it records exactly the `hpair_tail` shape needed by the concrete argmax filtered-sum probability wrapper after instantiating `empmean` with `etc.empmeanatexploration`. structure compiled","shard":"modules/d038bef37a47bcd6.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_eq_arm_le_pairwise_tail_of_contract","label":"prob_argmaxCommitOracle_eq_arm_le_pairwise_tail_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_eq_arm_le_pairwise_tail_of_contract","description":"The concrete argmax commit probability for one non-best arm is bounded by the matching arm entry of a pairwise empirical-mean tail contract. Unlike the filtered-sum consumer below, this theorem performs no finite union and keeps the candidate arm visible for a later gap-weighted expected-regret sum.","url":"../modules/banditrlproof-algorithms-etcpairwisetailcontract/index.html#decl-830adab5f786","parent":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","order":887,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCPairwiseTailContract"],["Source","BanditRLProof/Algorithms/ETCPairwiseTailContract.lean:48"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_eq_arm_le_pairwise_tail_of_contract {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) (hcontract : ETC.PairwiseEmpMeanTailContract mu spec model commitArm reward tail) (a : Fin K) (hne : a = model.bestArm -> False) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (fun b : Fin K => ETC.empMeanAtExploration spec commitArm (reward omega) b) = a} <= tail a","missing":[],"search":"prob_argmaxcommitoracle_eq_arm_le_pairwise_tail_of_contract banditrlproof.etc.prob_argmaxcommitoracle_eq_arm_le_pairwise_tail_of_contract the concrete argmax commit probability for one non-best arm is bounded by the matching arm entry of a pairwise empirical-mean tail contract. unlike the filtered-sum consumer below, this theorem performs no finite union and keeps the candidate arm visible for a later gap-weighted expected-regret sum. theorem compiled","shard":"modules/d038bef37a47bcd6.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_contract","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_contract","description":"The concrete argmax wrong-commit probability wrapper consumes the fixed-commit ETC empirical-mean pairwise-tail contract directly. This leaf only connects the contract surface to the compiled probability consumer. It does not prove the contract, import Hoeffding/sub-Gaussian tails, introduce filtration, or prove final ETC regret.","url":"../modules/banditrlproof-algorithms-etcpairwisetailcontract/index.html#decl-5c278f53aeab","parent":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","order":888,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCPairwiseTailContract"],["Source","BanditRLProof/Algorithms/ETCPairwiseTailContract.lean:83"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_contract {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (tail : Fin K -> ENNReal) (hcontract : ETC.PairwiseEmpMeanTailContract mu spec model commitArm reward tail) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (fun a : Fin K => ETC.empMeanAtExploration spec commitArm (reward omega) a) = model.bestArm -> False} <= ((Finset.univ : Finset (Fin K)).filter (fun a : Fin K => a = model.bestArm -> False)).sum tail","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail_of_contract banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_pairwise_tail_of_contract the concrete argmax wrong-commit probability wrapper consumes the fixed-commit etc empirical-mean pairwise-tail contract directly. this leaf only connects the contract surface to the compiled probability consumer. it does not prove the contract, import hoeffding/sub-gaussian tails, introduce filtration, or prove final etc regret. theorem compiled","shard":"modules/d038bef37a47bcd6.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.ratArmLawRealKernel","label":"ratArmLawRealKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.ratArmLawRealKernel","description":"Push each countable `Rat` arm law forward to a Real-valued arm kernel.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-2aee81811993","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":889,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:24"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def ratArmLawRealKernel {K : Nat} (armLaw : Fin K -> Measure Rat) : ProbabilityTheory.Kernel (Fin K) Real","missing":[],"search":"ratarmlawrealkernel banditrlproof.etc.ratarmlawrealkernel push each countable `rat` arm law forward to a real-valued arm kernel. definition compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.ratArmLawRealKernel_apply","label":"ratArmLawRealKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.ratArmLawRealKernel_apply","description":"theorem ratArmLawRealKernel_apply {K : Nat} (armLaw : Fin K -> Measure Rat) (arm : Fin K) : ratArmLawRealKernel armLaw arm = Measure.map (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-ee67345af3e0","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":890,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:32"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ratArmLawRealKernel_apply {K : Nat} (armLaw : Fin K -> Measure Rat) (arm : Fin K) : ratArmLawRealKernel armLaw arm = Measure.map (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)","missing":[],"search":"ratarmlawrealkernel_apply banditrlproof.etc.ratarmlawrealkernel_apply theorem ratarmlawrealkernel_apply {k : nat} (armlaw : fin k -> measure rat) (arm : fin k) : ratarmlawrealkernel armlaw arm = measure.map (fun reward : rat => ((reward : rat) : real)) (armlaw arm) theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.isMarkovKernel_ratArmLawRealKernel","label":"isMarkovKernel_ratArmLawRealKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.isMarkovKernel_ratArmLawRealKernel","description":"Probability arm laws give a Markov Real pushforward kernel.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-2fa30129e940","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":891,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:40"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_ratArmLawRealKernel {K : Nat} (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) : ProbabilityTheory.IsMarkovKernel (ratArmLawRealKernel armLaw)","missing":[],"search":"ismarkovkernel_ratarmlawrealkernel banditrlproof.etc.ismarkovkernel_ratarmlawrealkernel probability arm laws give a markov real pushforward kernel. theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelMean_ratArmLawRealKernel_eq_integral_cast","label":"realKernelMean_ratArmLawRealKernel_eq_integral_cast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelMean_ratArmLawRealKernel_eq_integral_cast","description":"The pushforward kernel identity integral is the original casted mean.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-7e07a013a315","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":892,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:53"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelMean_ratArmLawRealKernel_eq_integral_cast {K : Nat} (armLaw : Fin K -> Measure Rat) (arm : Fin K) : realKernelMean (ratArmLawRealKernel armLaw) arm = integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real))","missing":[],"search":"realkernelmean_ratarmlawrealkernel_eq_integral_cast banditrlproof.etc.realkernelmean_ratarmlawrealkernel_eq_integral_cast the pushforward kernel identity integral is the original casted mean. theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelMean_ratArmLawRealKernel_eq_modelMean","label":"realKernelMean_ratArmLawRealKernel_eq_modelMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelMean_ratArmLawRealKernel_eq_modelMean","description":"Exact arm means identify the pushforward Real kernel mean with the model mean.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-1df44b07e1b3","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":893,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:65"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelMean_ratArmLawRealKernel_eq_modelMean {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (arm : Fin K) : realKernelMean (ratArmLawRealKernel armLaw) arm = ((model.mean arm : Rat) : Real)","missing":[],"search":"realkernelmean_ratarmlawrealkernel_eq_modelmean banditrlproof.etc.realkernelmean_ratarmlawrealkernel_eq_modelmean exact arm means identify the pushforward real kernel mean with the model mean. theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.ciSup_modelMean_cast_eq_bestArm","label":"ciSup_modelMean_cast_eq_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.ciSup_modelMean_cast_eq_bestArm","description":"The supremum of the cast model means is attained at the local best arm.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-63800097bdb5","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":894,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:78"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ciSup_modelMean_cast_eq_bestArm {K : Nat} (model : FiniteBanditModel K) : (⨆ arm : Fin K, ((model.mean arm : Rat) : Real)) = ((model.mean model.bestArm : Rat) : Real)","missing":[],"search":"cisup_modelmean_cast_eq_bestarm banditrlproof.etc.cisup_modelmean_cast_eq_bestarm the supremum of the cast model means is attained at the local best arm. theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelGap_ratArmLawRealKernel_eq_modelGap","label":"realKernelGap_ratArmLawRealKernel_eq_modelGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelGap_ratArmLawRealKernel_eq_modelGap","description":"The pushforward Real kernel gap is exactly the cast local model gap.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-d313867a24c6","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":895,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:90"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelGap_ratArmLawRealKernel_eq_modelGap {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (arm : Fin K) : realKernelGap (ratArmLawRealKernel armLaw) arm = ((model.gap arm : Rat) : Real)","missing":[],"search":"realkernelgap_ratarmlawrealkernel_eq_modelgap banditrlproof.etc.realkernelgap_ratarmlawrealkernel_eq_modelgap the pushforward real kernel gap is exactly the cast local model gap. theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_explorationArgmaxAction_le_exact_sum_of_armLaws","label":"integral_realKernelRegret_explorationArgmaxAction_le_exact_sum_of_armLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_explorationArgmaxAction_le_exact_sum_of_armLaws","description":"Canonical exact ETC expected regret against the Real pushforward arm kernel. This is the finite-arm assembly of the exact per-arm expected pull-count leaf: the best-arm summand vanishes, while each non-best summand uses the exact `exp (-m * gap^2 / (4 * sigma2))` count bound.","url":"../modules/banditrlproof-algorithms-etcratarmlawrealkernel/index.html#decl-3b20b2d741c7","parent":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","order":896,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRatArmLawRealKernel"],["Source","BanditRLProof/Algorithms/ETCRatArmLawRealKernel.lean:115"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_explorationArgmaxAction_le_exact_sum_of_armLaws {K : Nat} {Context : Type} [MeasurableSpace Context] (spec : ETC.Spec K) (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (sigma2 : NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => ((reward : Rat) : Real) - ((model.mean arm : Rat) : Real)) sigma2 (armLaw arm)) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (hexplorationPulls_pos : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) : let defaultAction := ETC.exploreArm spec 0 let rewardKernel := RewardKernel.contextIndependentOfActionLaws (…","missing":[],"search":"integral_realkernelregret_explorationargmaxaction_le_exact_sum_of_armlaws banditrlproof.etc.integral_realkernelregret_explorationargmaxaction_le_exact_sum_of_armlaws canonical exact etc expected regret against the real pushforward arm kernel. this is the finite-arm assembly of the exact per-arm expected pull-count leaf: the best-arm summand vanishes, while each non-best summand uses the exact `exp (-m * gap^2 / (4 * sigma2))` count bound. theorem compiled","shard":"modules/62af8cc17d32d317.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.argmax_cons_eq_some_foldl_real_select","label":"argmax_cons_eq_some_foldl_real_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.argmax_cons_eq_some_foldl_real_select","description":"private theorem argmax_cons_eq_some_foldl_real_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) (l : List (Fin K)) : List.argmax scores (init :: l) = some (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-f5d15c7e85ab","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":897,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:19"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem argmax_cons_eq_some_foldl_real_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) (l : List (Fin K)) : List.argmax scores (init :: l) = some (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)","missing":[],"search":"argmax_cons_eq_some_foldl_real_select banditrlproof.etc.argmax_cons_eq_some_foldl_real_select private theorem argmax_cons_eq_some_foldl_real_select {k : nat} (scores : fin k -> real) (init : fin k) (l : list (fin k)) : list.argmax scores (init :: l) = some (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init) theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realArgmaxCommit_argmax_finRange","label":"realArgmaxCommit_argmax_finRange","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realArgmaxCommit_argmax_finRange","description":"The native strict-update fold is Mathlib's first-occurrence list argmax.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-05d95e2546cc","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":898,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:41"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realArgmaxCommit_argmax_finRange {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : List.argmax scores (List.finRange K) = some (ETC.realArgmaxCommit hK scores)","missing":[],"search":"realargmaxcommit_argmax_finrange banditrlproof.etc.realargmaxcommit_argmax_finrange the native strict-update fold is mathlib's first-occurrence list argmax. theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realArgmaxCommit_encode_le_of_score_le","label":"realArgmaxCommit_encode_le_of_score_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realArgmaxCommit_encode_le_of_score_le","description":"Among score maximizers, the native fold chooses the least encoded arm.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-37bb5d19e415","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":899,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:53"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realArgmaxCommit_encode_le_of_score_le {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) (hscore : scores (ETC.realArgmaxCommit hK scores) <= scores a) : Encodable.encode (ETC.realArgmaxCommit hK scores) <= Encodable.encode a","missing":[],"search":"realargmaxcommit_encode_le_of_score_le banditrlproof.etc.realargmaxcommit_encode_le_of_score_le among score maximizers, the native fold chooses the least encoded arm. theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.RealEncodedArgmaxCandidate","label":"RealEncodedArgmaxCandidate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.RealEncodedArgmaxCandidate","description":"A natural number encodes a maximizing arm for the supplied score vector.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-8f98c883596b","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":900,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:65"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def RealEncodedArgmaxCandidate {K : Nat} (scores : Fin K -> Real) (n : Nat) : Prop","missing":[],"search":"realencodedargmaxcandidate banditrlproof.etc.realencodedargmaxcandidate a natural number encodes a maximizing arm for the supplied score vector. definition compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.exists_realEncodedArgmaxCandidate","label":"exists_realEncodedArgmaxCandidate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.exists_realEncodedArgmaxCandidate","description":"theorem exists_realEncodedArgmaxCandidate {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Exists fun n : Nat => ETC.RealEncodedArgmaxCandidate scores n","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-96d1c55ab0b9","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":901,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:71"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem exists_realEncodedArgmaxCandidate {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Exists fun n : Nat => ETC.RealEncodedArgmaxCandidate scores n","missing":[],"search":"exists_realencodedargmaxcandidate banditrlproof.etc.exists_realencodedargmaxcandidate theorem exists_realencodedargmaxcandidate {k : nat} (hk : 0 < k) (scores : fin k -> real) : exists fun n : nat => etc.realencodedargmaxcandidate scores n theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmaxIndex","label":"realLeastEncodedArgmaxIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmaxIndex","description":"The least encoded natural-number witness of a maximizing arm.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-8c846031832c","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":902,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:79"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realLeastEncodedArgmaxIndex {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Nat","missing":[],"search":"realleastencodedargmaxindex banditrlproof.etc.realleastencodedargmaxindex the least encoded natural-number witness of a maximizing arm. definition compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmaxIndex_candidate","label":"realLeastEncodedArgmaxIndex_candidate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmaxIndex_candidate","description":"theorem realLeastEncodedArgmaxIndex_candidate {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : ETC.RealEncodedArgmaxCandidate scores (ETC.realLeastEncodedArgmaxIndex hK scores)","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-0c9158cfc9d1","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":903,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:84"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realLeastEncodedArgmaxIndex_candidate {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : ETC.RealEncodedArgmaxCandidate scores (ETC.realLeastEncodedArgmaxIndex hK scores)","missing":[],"search":"realleastencodedargmaxindex_candidate banditrlproof.etc.realleastencodedargmaxindex_candidate theorem realleastencodedargmaxindex_candidate {k : nat} (hk : 0 < k) (scores : fin k -> real) : etc.realencodedargmaxcandidate scores (etc.realleastencodedargmaxindex hk scores) theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax","label":"realLeastEncodedArgmax","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmax","description":"The maximizing arm decoded from the least `Nat.find` witness.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-47ff574fc63b","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":904,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:93"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realLeastEncodedArgmax {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Fin K","missing":[],"search":"realleastencodedargmax banditrlproof.etc.realleastencodedargmax the maximizing arm decoded from the least `nat.find` witness. definition compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_encode_eq_index","label":"realLeastEncodedArgmax_encode_eq_index","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmax_encode_eq_index","description":"theorem realLeastEncodedArgmax_encode_eq_index {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Encodable.encode (ETC.realLeastEncodedArgmax hK scores) = ETC.realLeastEncodedArgmaxIndex hK scores","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-9a37686e4d23","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":905,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:97"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realLeastEncodedArgmax_encode_eq_index {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Encodable.encode (ETC.realLeastEncodedArgmax hK scores) = ETC.realLeastEncodedArgmaxIndex hK scores","missing":[],"search":"realleastencodedargmax_encode_eq_index banditrlproof.etc.realleastencodedargmax_encode_eq_index theorem realleastencodedargmax_encode_eq_index {k : nat} (hk : 0 < k) (scores : fin k -> real) : encodable.encode (etc.realleastencodedargmax hk scores) = etc.realleastencodedargmaxindex hk scores theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_spec","label":"realLeastEncodedArgmax_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmax_spec","description":"theorem realLeastEncodedArgmax_spec {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) : scores a <= scores (ETC.realLeastEncodedArgmax hK scores)","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-0c903d412f9a","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":906,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:104"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realLeastEncodedArgmax_spec {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) : scores a <= scores (ETC.realLeastEncodedArgmax hK scores)","missing":[],"search":"realleastencodedargmax_spec banditrlproof.etc.realleastencodedargmax_spec theorem realleastencodedargmax_spec {k : nat} (hk : 0 < k) (scores : fin k -> real) (a : fin k) : scores a <= scores (etc.realleastencodedargmax hk scores) theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_encode_le_of_isMax","label":"realLeastEncodedArgmax_encode_le_of_isMax","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmax_encode_le_of_isMax","description":"theorem realLeastEncodedArgmax_encode_le_of_isMax {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) (ha : forall z : Fin K, scores z <= scores a) : Encodable.encode (ETC.realLeastEncodedArgmax hK scores) <= Encodable.encode a","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-0bc4bcf58328","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":907,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:110"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realLeastEncodedArgmax_encode_le_of_isMax {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) (ha : forall z : Fin K, scores z <= scores a) : Encodable.encode (ETC.realLeastEncodedArgmax hK scores) <= Encodable.encode a","missing":[],"search":"realleastencodedargmax_encode_le_of_ismax banditrlproof.etc.realleastencodedargmax_encode_le_of_ismax theorem realleastencodedargmax_encode_le_of_ismax {k : nat} (hk : 0 < k) (scores : fin k -> real) (a : fin k) (ha : forall z : fin k, scores z <= scores a) : encodable.encode (etc.realleastencodedargmax hk scores) <= encodable.encode a theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_eq_realArgmaxCommit","label":"realLeastEncodedArgmax_eq_realArgmaxCommit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realLeastEncodedArgmax_eq_realArgmaxCommit","description":"The LML-shaped least-encoded selector equals the native strict fold.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-a01d45299d59","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":908,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:123"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realLeastEncodedArgmax_eq_realArgmaxCommit {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : ETC.realLeastEncodedArgmax hK scores = ETC.realArgmaxCommit hK scores","missing":[],"search":"realleastencodedargmax_eq_realargmaxcommit banditrlproof.etc.realleastencodedargmax_eq_realargmaxcommit the lml-shaped least-encoded selector equals the native strict fold. theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.eventually_realExplorationArgmaxAction_eq_of_roundRobin_leastEncodedCommit_persist","label":"eventually_realExplorationArgmaxAction_eq_of_roundRobin_leastEncodedCommit_persist","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.eventually_realExplorationArgmaxAction_eq_of_roundRobin_leastEncodedCommit_persist","description":"Round-robin exploration, least-encoded commit, and persistence determine the native Real ETC action trace. The boundary is written as `K * m` to match the upstream ETC behavior lemmas; the local trace uses the definitionally equivalent `m * K` boundary.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-97b5648ceb32","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":909,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:147"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem eventually_realExplorationArgmaxAction_eq_of_roundRobin_leastEncodedCommit_persist {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (spec : ETC.Spec K) (baseCommitArm : Fin K) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (hactionExplore : forall t, t < K * spec.explorationPulls -> Filter.EventuallyEq (ae mu) (fun omega => action omega t) (fun _omega => ETC.exploreArm spec t)) (hactionCommit : Filter.EventuallyEq (ae mu) (fun omega => action omega (K * spec.explorationPulls)) (fun omega => ETC.realLeastEncodedArgmax spec.hK (fun arm => ETC.realEmpMeanAtExploration spec baseCommitArm (reward omega) arm))) (hactionPersist : forall t, K * spec.explorationPulls <= t -> Filter.EventuallyEq (ae mu) (fun omega => action omega t) (fun omega => action omega (K * spec.explorationPulls))) : Filter.Eventually (fun omega => forall t, acti…","missing":[],"search":"eventually_realexplorationargmaxaction_eq_of_roundrobin_leastencodedcommit_persist banditrlproof.etc.eventually_realexplorationargmaxaction_eq_of_roundrobin_leastencodedcommit_persist round-robin exploration, least-encoded commit, and persistence determine the native real etc action trace. the boundary is written as `k * m` to match the upstream etc behavior lemmas; the local trace uses the definitionally equivalent `m * k` boundary. theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_leastEncodedCommit_persist","label":"integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_leastEncodedCommit_persist","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_leastEncodedCommit_persist","description":"Exact native Real ETC regret from upstream-shaped feedback laws and the three ETC action behavior fields. Unlike the lower source adapter, callers do not provide a preassembled horizon action equality: exploration, least-encoded commit, and persistence construct it here.","url":"../modules/banditrlproof-algorithms-etcrealargmaxtie/index.html#decl-3f1273fbd68f","parent":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","order":910,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealArgmaxTie"],["Source","BanditRLProof/Algorithms/ETCRealArgmaxTie.lean:211"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_leastEncodedCommit_persist {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (hactionExplore : forall t, t < K * spec.explorationPulls -> Filter.EventuallyEq (ae mu) (fun omega => action omega t) (fun _omega =…","missing":[],"search":"integral_realkernelregret_externalaction_le_exact_sum_of_actiondependent_actionrewardhistory_conddistrib_of_leastencodedcommit_persist banditrlproof.etc.integral_realkernelregret_externalaction_le_exact_sum_of_actiondependent_actionrewardhistory_conddistrib_of_leastencodedcommit_persist exact native real etc regret from upstream-shaped feedback laws and the three etc action behavior fields. unlike the lower source adapter, callers do not provide a preassembled horizon action equality: exploration, least-encoded commit, and persistence construct it here. theorem compiled","shard":"modules/fb441a84537c03e3.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realEmpMeanAtExploration","label":"realEmpMeanAtExploration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realEmpMeanAtExploration","description":"The exploration empirical mean formed directly from a Real reward trace.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-8105063bbe5d","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":911,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:22"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realEmpMeanAtExploration {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : RewardTrace Real) (a : Fin K) : Real","missing":[],"search":"realempmeanatexploration banditrlproof.etc.realempmeanatexploration the exploration empirical mean formed directly from a real reward trace. definition compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realEmpMeanAtExploration_eq_sumRewards_div_explorationPulls","label":"realEmpMeanAtExploration_eq_sumRewards_div_explorationPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realEmpMeanAtExploration_eq_sumRewards_div_explorationPulls","description":"The deterministic exploration count removes the pull-count denominator.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-3eae163be0dc","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":912,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:30"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realEmpMeanAtExploration_eq_sumRewards_div_explorationPulls {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : RewardTrace Real) (a : Fin K) : ETC.realEmpMeanAtExploration spec baseCommitArm reward a = sumRewards (ETC.actionWithCommit spec baseCommitArm) reward a (spec.explorationPulls * K) / (spec.explorationPulls : Real)","missing":[],"search":"realempmeanatexploration_eq_sumrewards_div_explorationpulls banditrlproof.etc.realempmeanatexploration_eq_sumrewards_div_explorationpulls the deterministic exploration count removes the pull-count denominator. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realEmpMeanAtExploration","label":"measurable_realEmpMeanAtExploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realEmpMeanAtExploration","description":"Real exploration empirical means are measurable from measurable rewards.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-6be0555a182a","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":913,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:41"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realEmpMeanAtExploration {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (spec : ETC.Spec K) (baseCommitArm a : Fin K) (reward : Omega -> RewardTrace Real) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Measurable (fun omega : Omega => ETC.realEmpMeanAtExploration spec baseCommitArm (reward omega) a)","missing":[],"search":"measurable_realempmeanatexploration banditrlproof.etc.measurable_realempmeanatexploration real exploration empirical means are measurable from measurable rewards. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_score_le_foldl_select","label":"real_score_le_foldl_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_score_le_foldl_select","description":"private theorem real_score_le_foldl_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha;…","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-d72eb9ba161b","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":914,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:64"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem real_score_le_foldl_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha; cases ha) (by simp) | arm :: rest => by let select := fun best arm : Fin K => if scores best < scores arm then arm else best let next := select init arm have ih","missing":[],"search":"real_score_le_foldl_select banditrlproof.etc.real_score_le_foldl_select private theorem real_score_le_foldl_select {k : nat} (scores : fin k -> real) (init : fin k) : forall l : list (fin k), (forall a : fin k, list.mem a l -> scores a <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init) | [] => by exact and.intro (by intro _ ha; cases ha) (by simp) | arm :: rest => by let select := fun best arm : fin k => if scores best < scores arm then arm else best let next := select init arm have ih theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realArgmaxCommit","label":"realArgmaxCommit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realArgmaxCommit","description":"A deterministic finite argmax for Real-valued arm scores.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-03798b606aae","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":915,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:107"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realArgmaxCommit {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Fin K","missing":[],"search":"realargmaxcommit banditrlproof.etc.realargmaxcommit a deterministic finite argmax for real-valued arm scores. definition compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realArgmaxCommit_spec","label":"realArgmaxCommit_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realArgmaxCommit_spec","description":"The Real finite argmax dominates every arm score.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-3b24f5489610","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":916,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:115"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realArgmaxCommit_spec {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) : scores a <= scores (ETC.realArgmaxCommit hK scores)","missing":[],"search":"realargmaxcommit_spec banditrlproof.etc.realargmaxcommit_spec the real finite argmax dominates every arm score. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realArgmaxCommit_const","label":"realArgmaxCommit_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realArgmaxCommit_const","description":"On a constant score vector, the tie rule keeps the initial arm `0`.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-3fa97d4a150b","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":917,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:124"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"@[simp] theorem realArgmaxCommit_const {K : Nat} (hK : 0 < K) (c : Real) : ETC.realArgmaxCommit hK (fun _a : Fin K => c) = Fin.mk 0 hK","missing":[],"search":"realargmaxcommit_const banditrlproof.etc.realargmaxcommit_const on a constant score vector, the tie rule keeps the initial arm `0`. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_selected_real_score","label":"measurable_selected_real_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_selected_real_score","description":"private theorem measurable_selected_real_score {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : Measurable (fun omega : Omega => scores omega (best omega))","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-279a88b5bd7b","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":918,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:128"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem measurable_selected_real_score {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : Measurable (fun omega : Omega => scores omega (best omega))","missing":[],"search":"measurable_selected_real_score banditrlproof.etc.measurable_selected_real_score private theorem measurable_selected_real_score {omega : type u} {k : nat} [measurablespace omega] (scores : omega -> fin k -> real) (hscores : forall a : fin k, measurable (fun omega : omega => scores omega a)) (best : omega -> fin k) (hbest : measurable best) : measurable (fun omega : omega => scores omega (best omega)) theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_foldl_real_select","label":"measurable_foldl_real_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_foldl_real_select","description":"private theorem measurable_foldl_real_select {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : forall l : List (Fin K), Measurable (fun omega : Omega => l.foldl (fun best arm : Fin K => if scores omega best < scores omega arm then arm else best) (best o…","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-15d371ddc3db","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":919,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:149"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem measurable_foldl_real_select {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : forall l : List (Fin K), Measurable (fun omega : Omega => l.foldl (fun best arm : Fin K => if scores omega best < scores omega arm then arm else best) (best omega)) | [] => by simpa using hbest | arm :: rest => by have hselected : Measurable (fun omega : Omega => scores omega (best omega))","missing":[],"search":"measurable_foldl_real_select banditrlproof.etc.measurable_foldl_real_select private theorem measurable_foldl_real_select {omega : type u} {k : nat} [measurablespace omega] (scores : omega -> fin k -> real) (hscores : forall a : fin k, measurable (fun omega : omega => scores omega a)) (best : omega -> fin k) (hbest : measurable best) : forall l : list (fin k), measurable (fun omega : omega => l.foldl (fun best arm : fin k => if scores omega best < scores omega arm then arm else best) (best omega)) | [] => by simpa using hbest | arm :: rest => by have hselected : measurable (fun omega : omega => scores omega (best omega)) theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realArgmaxCommit_of_forall_measurable","label":"measurable_realArgmaxCommit_of_forall_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realArgmaxCommit_of_forall_measurable","description":"The finite Real argmax is measurable when every score coordinate is.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-6e811691f839","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":920,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:185"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realArgmaxCommit_of_forall_measurable {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) : Measurable (fun omega : Omega => ETC.realArgmaxCommit hK (scores omega))","missing":[],"search":"measurable_realargmaxcommit_of_forall_measurable banditrlproof.etc.measurable_realargmaxcommit_of_forall_measurable the finite real argmax is measurable when every score coordinate is. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit","label":"realExplorationArgmaxCommit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationArgmaxCommit","description":"Commit to the arm maximizing the native Real exploration means.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-04ff2487b554","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":921,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:197"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realExplorationArgmaxCommit {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : RewardTrace Real) : Fin K","missing":[],"search":"realexplorationargmaxcommit banditrlproof.etc.realexplorationargmaxcommit commit to the arm maximizing the native real exploration means. definition compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationArgmaxAction","label":"realExplorationArgmaxAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationArgmaxAction","description":"Native Real ETC trace: round-robin exploration, then Real empirical argmax.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-5b1815caaf20","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":922,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:204"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realExplorationArgmaxAction {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : RewardTrace Real) : ActionTrace (Fin K)","missing":[],"search":"realexplorationargmaxaction banditrlproof.etc.realexplorationargmaxaction native real etc trace: round-robin exploration, then real empirical argmax. definition compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realExplorationArgmaxCommit","label":"measurable_realExplorationArgmaxCommit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realExplorationArgmaxCommit","description":"The reward-dependent native Real ETC commit arm is measurable.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-faf3965206c5","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":923,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:211"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realExplorationArgmaxCommit {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : Omega -> RewardTrace Real) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Measurable (fun omega : Omega => ETC.realExplorationArgmaxCommit spec baseCommitArm (reward omega))","missing":[],"search":"measurable_realexplorationargmaxcommit banditrlproof.etc.measurable_realexplorationargmaxcommit the reward-dependent native real etc commit arm is measurable. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_eq_exploration_add_remaining_mul_commit_prob","label":"integral_real_pullCount_realExplorationArgmaxAction_eq_exploration_add_remaining_mul_commit_prob","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_eq_exploration_add_remaining_mul_commit_prob","description":"Exact expected pull count for the native Real empirical-argmax ETC trace.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-0816c7009fbc","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":924,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:226"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_realExplorationArgmaxAction_eq_exploration_add_remaining_mul_commit_prob {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : Omega -> RewardTrace Real) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) (n : Nat) (hn : K * spec.explorationPulls <= n) : integral mu (fun omega : Omega => ((pullCount (ETC.realExplorationArgmaxAction spec baseCommitArm (reward omega)) a n : Nat) : Real)) = (spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * mu.real {omega : Omega | ETC.realExplorationArgmaxCommit spec baseCommitArm (reward omega) = a}","missing":[],"search":"integral_real_pullcount_realexplorationargmaxaction_eq_exploration_add_remaining_mul_commit_prob banditrlproof.etc.integral_real_pullcount_realexplorationargmaxaction_eq_exploration_add_remaining_mul_commit_prob exact expected pull count for the native real empirical-argmax etc trace. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_le_exploration_add_remaining_mul_of_commit_prob_le","label":"integral_real_pullCount_realExplorationArgmaxAction_le_exploration_add_remaining_mul_of_commit_prob_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_le_exploration_add_remaining_mul_of_commit_prob_le","description":"A commit-fiber bound immediately yields the native Real expected count bound.","url":"../modules/banditrlproof-algorithms-etcrealempiricalmean/index.html#decl-927a8e9bd2f6","parent":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","order":925,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealEmpiricalMean"],["Source","BanditRLProof/Algorithms/ETCRealEmpiricalMean.lean:255"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_realExplorationArgmaxAction_le_exploration_add_remaining_mul_of_commit_prob_le {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward : Omega -> RewardTrace Real) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Fin K) (n : Nat) (p : Real) (hn : K * spec.explorationPulls <= n) (hprob : mu.real {omega : Omega | ETC.realExplorationArgmaxCommit spec baseCommitArm (reward omega) = a} <= p) : integral mu (fun omega : Omega => ((pullCount (ETC.realExplorationArgmaxAction spec baseCommitArm (reward omega)) a n : Nat) : Real)) <= (spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * p","missing":[],"search":"integral_real_pullcount_realexplorationargmaxaction_le_exploration_add_remaining_mul_of_commit_prob_le banditrlproof.etc.integral_real_pullcount_realexplorationargmaxaction_le_exploration_add_remaining_mul_of_commit_prob_le a commit-fiber bound immediately yields the native real expected count bound. theorem compiled","shard":"modules/9ae2ff24a48ff0e8.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistoryPullCount","label":"realHistoryPullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realHistoryPullCount","description":"Number of occurrences of an arm in an inclusive finite pair history.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-664422ef202e","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":926,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:18"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistoryPullCount {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Nat","missing":[],"search":"realhistorypullcount banditrlproof.etc.realhistorypullcount number of occurrences of an arm in an inclusive finite pair history. definition compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistorySumRewards","label":"realHistorySumRewards","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realHistorySumRewards","description":"Sum of rewards of an arm in an inclusive finite pair history.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-f74e6d88fe8e","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":927,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:24"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistorySumRewards {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Real","missing":[],"search":"realhistorysumrewards banditrlproof.etc.realhistorysumrewards sum of rewards of an arm in an inclusive finite pair history. definition compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistoryEmpMean","label":"realHistoryEmpMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realHistoryEmpMean","description":"Empirical mean of an arm in an inclusive finite pair history.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-cc6864b66d91","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":928,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:30"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistoryEmpMean {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Real","missing":[],"search":"realhistoryempmean banditrlproof.etc.realhistoryempmean empirical mean of an arm in an inclusive finite pair history. definition compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistoryPullCount_finitePairHistoryOfTrace","label":"realHistoryPullCount_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realHistoryPullCount_finitePairHistoryOfTrace","description":"Inclusive history pull counts are exclusive trace pull counts at `n + 1`.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-ac2bc291b331","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":929,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:37"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realHistoryPullCount_finitePairHistoryOfTrace {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n : Nat) (arm : Fin K) : ETC.realHistoryPullCount n (History.finitePairHistoryOfTrace action reward n) arm = pullCount action arm (n + 1)","missing":[],"search":"realhistorypullcount_finitepairhistoryoftrace banditrlproof.etc.realhistorypullcount_finitepairhistoryoftrace inclusive history pull counts are exclusive trace pull counts at `n + 1`. theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistorySumRewards_finitePairHistoryOfTrace","label":"realHistorySumRewards_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realHistorySumRewards_finitePairHistoryOfTrace","description":"Inclusive history reward sums are exclusive trace sums at `n + 1`.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-702977536ea1","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":930,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:68"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realHistorySumRewards_finitePairHistoryOfTrace {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n : Nat) (arm : Fin K) : ETC.realHistorySumRewards n (History.finitePairHistoryOfTrace action reward n) arm = sumRewards action reward arm (n + 1)","missing":[],"search":"realhistorysumrewards_finitepairhistoryoftrace banditrlproof.etc.realhistorysumrewards_finitepairhistoryoftrace inclusive history reward sums are exclusive trace sums at `n + 1`. theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistoryEmpMean_finitePairHistoryOfTrace","label":"realHistoryEmpMean_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realHistoryEmpMean_finitePairHistoryOfTrace","description":"The source-shaped history mean is the trace empirical mean at `n + 1`.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-03a8bed7b801","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":931,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:87"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realHistoryEmpMean_finitePairHistoryOfTrace {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n : Nat) (arm : Fin K) : ETC.realHistoryEmpMean n (History.finitePairHistoryOfTrace action reward n) arm = sumRewards action reward arm (n + 1) / (pullCount action arm (n + 1) : Real)","missing":[],"search":"realhistoryempmean_finitepairhistoryoftrace banditrlproof.etc.realhistoryempmean_finitepairhistoryoftrace the source-shaped history mean is the trace empirical mean at `n + 1`. theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_eq_of_eq_on_lt","label":"pullCount_eq_of_eq_on_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_eq_of_eq_on_lt","description":"private theorem pullCount_eq_of_eq_on_lt {Action : Type} [DecidableEq Action] (action action' : ActionTrace Action) (arm : Action) (n : Nat) (haction : forall t, t < n -> action t = action' t) : pullCount action arm n = pullCount action' arm n","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-745dbf5df390","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":932,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:98"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem pullCount_eq_of_eq_on_lt {Action : Type} [DecidableEq Action] (action action' : ActionTrace Action) (arm : Action) (n : Nat) (haction : forall t, t < n -> action t = action' t) : pullCount action arm n = pullCount action' arm n","missing":[],"search":"pullcount_eq_of_eq_on_lt banditrlproof.etc.pullcount_eq_of_eq_on_lt private theorem pullcount_eq_of_eq_on_lt {action : type} [decidableeq action] (action action' : actiontrace action) (arm : action) (n : nat) (haction : forall t, t < n -> action t = action' t) : pullcount action arm n = pullcount action' arm n theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.sumRewards_eq_of_action_eq_on_lt","label":"sumRewards_eq_of_action_eq_on_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.sumRewards_eq_of_action_eq_on_lt","description":"private theorem sumRewards_eq_of_action_eq_on_lt {Action Reward : Type} [DecidableEq Action] [AddCommMonoid Reward] (action action' : ActionTrace Action) (reward : RewardTrace Reward) (arm : Action) (n : Nat) (haction : forall t, t < n -> action t = action' t) : sumRewards action reward arm n = sumRewards action' reward arm n","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-c3fad0b0ae23","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":933,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:111"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"private theorem sumRewards_eq_of_action_eq_on_lt {Action Reward : Type} [DecidableEq Action] [AddCommMonoid Reward] (action action' : ActionTrace Action) (reward : RewardTrace Reward) (arm : Action) (n : Nat) (haction : forall t, t < n -> action t = action' t) : sumRewards action reward arm n = sumRewards action' reward arm n","missing":[],"search":"sumrewards_eq_of_action_eq_on_lt banditrlproof.etc.sumrewards_eq_of_action_eq_on_lt private theorem sumrewards_eq_of_action_eq_on_lt {action reward : type} [decidableeq action] [addcommmonoid reward] (action action' : actiontrace action) (reward : rewardtrace reward) (arm : action) (n : nat) (haction : forall t, t < n -> action t = action' t) : sumrewards action reward arm n = sumrewards action' reward arm n theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realHistoryEmpMean_exploration_eq_realEmpMeanAtExploration","label":"realHistoryEmpMean_exploration_eq_realEmpMeanAtExploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realHistoryEmpMean_exploration_eq_realEmpMeanAtExploration","description":"At the exploration boundary, the pinned-source finite-history score equals the native Real exploration score whenever the observed actions are round robin.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-75f8c915c90d","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":934,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:131"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realHistoryEmpMean_exploration_eq_realEmpMeanAtExploration {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (hm : 0 < spec.explorationPulls) (hactionExplore : forall t, t < K * spec.explorationPulls -> action t = ETC.exploreArm spec t) (arm : Fin K) : ETC.realHistoryEmpMean (K * spec.explorationPulls - 1) (History.finitePairHistoryOfTrace action reward (K * spec.explorationPulls - 1)) arm = ETC.realEmpMeanAtExploration spec baseCommitArm reward arm","missing":[],"search":"realhistoryempmean_exploration_eq_realempmeanatexploration banditrlproof.etc.realhistoryempmean_exploration_eq_realempmeanatexploration at the exploration boundary, the pinned-source finite-history score equals the native real exploration score whenever the observed actions are round robin. theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_historyLeastEncodedCommit_persist","label":"integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_historyLeastEncodedCommit_persist","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_historyLeastEncodedCommit_persist","description":"Exact native Real ETC regret from a source-shaped finite-history commit score. The history score is rewritten locally; callers no longer provide a commit law already phrased with `realEmpMeanAtExploration`.","url":"../modules/banditrlproof-algorithms-etcrealhistoryscore/index.html#decl-25b8a6a44f53","parent":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","order":935,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealHistoryScore"],["Source","BanditRLProof/Algorithms/ETCRealHistoryScore.lean:181"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_historyLeastEncodedCommit_persist {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (hactionExplore : forall t, t < K * spec.explorationPulls -> Filter.EventuallyEq (ae mu) (fun omega => action omega t) (fun _…","missing":[],"search":"integral_realkernelregret_externalaction_le_exact_sum_of_actiondependent_actionrewardhistory_conddistrib_of_historyleastencodedcommit_persist banditrlproof.etc.integral_realkernelregret_externalaction_le_exact_sum_of_actiondependent_actionrewardhistory_conddistrib_of_historyleastencodedcommit_persist exact native real etc regret from a source-shaped finite-history commit score. the history score is rewritten locally; callers no longer provide a commit law already phrased with `realempmeanatexploration`. theorem compiled","shard":"modules/e5b1389df6ad46a2.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelBestArm","label":"realKernelBestArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelBestArm","description":"A finite maximizer of the identity-integral means of a Real reward kernel.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-b6fbf673e5f5","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":936,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:27"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realKernelBestArm {K : Nat} (hK : 0 < K) (nu : ProbabilityTheory.Kernel (Fin K) Real) : Fin K","missing":[],"search":"realkernelbestarm banditrlproof.etc.realkernelbestarm a finite maximizer of the identity-integral means of a real reward kernel. definition compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelMean_le_realKernelBestArm","label":"realKernelMean_le_realKernelBestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelMean_le_realKernelBestArm","description":"The selected kernel best arm dominates every arm mean.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-14c48a190d06","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":937,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:32"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelMean_le_realKernelBestArm {K : Nat} (hK : 0 < K) (nu : ProbabilityTheory.Kernel (Fin K) Real) (a : Fin K) : realKernelMean nu a <= realKernelMean nu (ETC.realKernelBestArm hK nu)","missing":[],"search":"realkernelmean_le_realkernelbestarm banditrlproof.etc.realkernelmean_le_realkernelbestarm the selected kernel best arm dominates every arm mean. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.ciSup_realKernelMean_eq_realKernelBestArm","label":"ciSup_realKernelMean_eq_realKernelBestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.ciSup_realKernelMean_eq_realKernelBestArm","description":"The finite supremum of kernel means is attained at `realKernelBestArm`.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-22a669c6213c","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":938,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:39"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ciSup_realKernelMean_eq_realKernelBestArm {K : Nat} (hK : 0 < K) (nu : ProbabilityTheory.Kernel (Fin K) Real) : (⨆ a : Fin K, realKernelMean nu a) = realKernelMean nu (ETC.realKernelBestArm hK nu)","missing":[],"search":"cisup_realkernelmean_eq_realkernelbestarm banditrlproof.etc.cisup_realkernelmean_eq_realkernelbestarm the finite supremum of kernel means is attained at `realkernelbestarm`. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelGap_eq_realKernelBestArm_sub","label":"realKernelGap_eq_realKernelBestArm_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelGap_eq_realKernelBestArm_sub","description":"Kernel gap is the selected best mean minus the queried arm mean.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-ee76936bbfa9","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":939,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:50"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelGap_eq_realKernelBestArm_sub {K : Nat} (hK : 0 < K) (nu : ProbabilityTheory.Kernel (Fin K) Real) (a : Fin K) : realKernelGap nu a = realKernelMean nu (ETC.realKernelBestArm hK nu) - realKernelMean nu a","missing":[],"search":"realkernelgap_eq_realkernelbestarm_sub banditrlproof.etc.realkernelgap_eq_realkernelbestarm_sub kernel gap is the selected best mean minus the queried arm mean. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realCenteredPairwiseRewardDiff","label":"realCenteredPairwiseRewardDiff","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realCenteredPairwiseRewardDiff","description":"Native Real centered candidate-minus-best reward difference.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-e67a0790d211","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":940,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:59"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realCenteredPairwiseRewardDiff {Omega : Type u} {K : Nat} (spec : ETC.Spec K) (mean : Fin K -> Real) (best commitArm : Fin K) (reward : Omega -> RewardTrace Real) (a : Fin K) (t : Nat) (omega : Omega) : Real","missing":[],"search":"realcenteredpairwiserewarddiff banditrlproof.etc.realcenteredpairwiserewarddiff native real centered candidate-minus-best reward difference. definition compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realCenteredPairwiseGapThreshold","label":"realCenteredPairwiseGapThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realCenteredPairwiseGapThreshold","description":"Native Real threshold of the centered pairwise exploration event.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-d1280fe0eefb","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":941,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:70"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realCenteredPairwiseGapThreshold {K : Nat} (spec : ETC.Spec K) (mean : Fin K -> Real) (best a : Fin K) : Real","missing":[],"search":"realcenteredpairwisegapthreshold banditrlproof.etc.realcenteredpairwisegapthreshold native real threshold of the centered pairwise exploration event. definition compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realCenteredPairwiseRewardDiffVarianceProxy","label":"realCenteredPairwiseRewardDiffVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realCenteredPairwiseRewardDiffVarianceProxy","description":"Variance proxy charged at candidate and best-arm exploration coordinates.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-7cbc5a7d0497","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":942,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:76"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realCenteredPairwiseRewardDiffVarianceProxy {K : Nat} (spec : ETC.Spec K) (best commitArm : Fin K) (cReward : Fin K -> Nat -> NNReal) (a : Fin K) (t : Nat) : NNReal","missing":[],"search":"realcenteredpairwiserewarddiffvarianceproxy banditrlproof.etc.realcenteredpairwiserewarddiffvarianceproxy variance proxy charged at candidate and best-arm exploration coordinates. definition compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","label":"real_selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","description":"Selected centered Real rewards equal reward sum minus count times mean.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-dd5cd45dc323","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":943,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:84"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Real) (a : Action) (n : Nat) (mu : Real) : (Finset.range n).sum (fun t => if action t = a then reward t - mu else 0) = sumRewards action reward a n - (pullCount action a n : Real) * mu","missing":[],"search":"real_selectedsubmean_sum_eq_sumrewards_sub_pullcount_mul banditrlproof.etc.real_selectedsubmean_sum_eq_sumrewards_sub_pullcount_mul selected centered real rewards equal reward sum minus count times mean. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","label":"real_meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","description":"Selected negative centered Real rewards equal count times mean minus sum.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-9486c2fc5d69","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":944,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:109"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Real) (a : Action) (n : Nat) (mu : Real) : (Finset.range n).sum (fun t => if action t = a then mu - reward t else 0) = (pullCount action a n : Real) * mu - sumRewards action reward a n","missing":[],"search":"real_meansubselected_sum_eq_pullcount_mul_sub_sumrewards banditrlproof.etc.real_meansubselected_sum_eq_pullcount_mul_sub_sumrewards selected negative centered real rewards equal count times mean minus sum. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_sumRewards_le_imp_centered_pairwise_sum_ge","label":"real_sumRewards_le_imp_centered_pairwise_sum_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_sumRewards_le_imp_centered_pairwise_sum_ge","description":"Equal exploration counts turn a raw reward comparison into a centered sum.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-baa6330473d7","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":945,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:141"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_sumRewards_le_imp_centered_pairwise_sum_ge {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Real) (a b : Action) (n m : Nat) (muA muB : Real) (hcount_a : pullCount action a n = m) (hcount_b : pullCount action b n = m) (hraw : sumRewards action reward b n <= sumRewards action reward a n) : (m : Real) * (muB - muA) <= (Finset.range n).sum (fun t => (if action t = a then reward t - muA else 0) + (if action t = b then muB - reward t else 0))","missing":[],"search":"real_sumrewards_le_imp_centered_pairwise_sum_ge banditrlproof.etc.real_sumrewards_le_imp_centered_pairwise_sum_ge equal exploration counts turn a raw reward comparison into a centered sum. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit_eq_arm_event_subset_centeredPairwise_sum_event","label":"realExplorationArgmaxCommit_eq_arm_event_subset_centeredPairwise_sum_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationArgmaxCommit_eq_arm_event_subset_centeredPairwise_sum_event","description":"A native Real commit fiber is contained in its centered pairwise tail event.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-12f431c34292","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":946,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:168"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realExplorationArgmaxCommit_eq_arm_event_subset_centeredPairwise_sum_event {Omega : Type u} {K : Nat} (spec : ETC.Spec K) (mean : Fin K -> Real) (best baseCommitArm : Fin K) (reward : Omega -> RewardTrace Real) (a : Fin K) (hm : 0 < spec.explorationPulls) : Set.Subset {omega | ETC.realExplorationArgmaxCommit spec baseCommitArm (reward omega) = a} {omega | ETC.realCenteredPairwiseGapThreshold spec mean best a <= (Finset.range (spec.explorationPulls * K)).sum (fun t => ETC.realCenteredPairwiseRewardDiff spec mean best baseCommitArm reward a t omega)}","missing":[],"search":"realexplorationargmaxcommit_eq_arm_event_subset_centeredpairwise_sum_event banditrlproof.etc.realexplorationargmaxcommit_eq_arm_event_subset_centeredpairwise_sum_event a native real commit fiber is contained in its centered pairwise tail event. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.iIndepFun_realCenteredPairwiseRewardDiff_of_iIndepFun_reward","label":"iIndepFun_realCenteredPairwiseRewardDiff_of_iIndepFun_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.iIndepFun_realCenteredPairwiseRewardDiff_of_iIndepFun_reward","description":"Coordinate independence survives the native Real pairwise transform.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-a84215038f8a","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":947,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:206"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_realCenteredPairwiseRewardDiff_of_iIndepFun_reward {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (spec : ETC.Spec K) (mean : Fin K -> Real) (best commitArm : Fin K) (reward : Omega -> RewardTrace Real) (h_reward_indep : ProbabilityTheory.iIndepFun (fun t omega => reward omega t) mu) (a : Fin K) : ProbabilityTheory.iIndepFun (fun t omega => ETC.realCenteredPairwiseRewardDiff spec mean best commitArm reward a t omega) mu","missing":[],"search":"iindepfun_realcenteredpairwiserewarddiff_of_iindepfun_reward banditrlproof.etc.iindepfun_realcenteredpairwiserewarddiff_of_iindepfun_reward coordinate independence survives the native real pairwise transform. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realCenteredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","label":"realCenteredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realCenteredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","description":"Per-coordinate centered sub-Gaussianity transfers to the pairwise summand.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-12173d7354b0","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":948,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:226"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realCenteredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsZeroOrProbabilityMeasure mu] (spec : ETC.Spec K) (mean : Fin K -> Real) (best commitArm : Fin K) (reward : Omega -> RewardTrace Real) (cReward : Fin K -> Nat -> NNReal) (a : Fin K) (t : Nat) (hne : a ≠ best) (h_subG : forall b, ETC.actionWithCommit spec commitArm t = b -> ProbabilityTheory.HasSubgaussianMGF (fun omega => reward omega t - mean b) (cReward b t) mu) : ProbabilityTheory.HasSubgaussianMGF (fun omega => ETC.realCenteredPairwiseRewardDiff spec mean best commitArm reward a t omega) (ETC.realCenteredPairwiseRewardDiffVarianceProxy spec best commitArm cReward a t) mu","missing":[],"search":"realcenteredpairwiserewarddiff_hassubgaussianmgf_of_centeredreward banditrlproof.etc.realcenteredpairwiserewarddiff_hassubgaussianmgf_of_centeredreward per-coordinate centered sub-gaussianity transfers to the pairwise summand. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.sum_realCenteredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","label":"sum_realCenteredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.sum_realCenteredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","description":"The constant pairwise proxy sums to exactly `2 * m * sigma2`.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-787bf55002c4","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":949,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:266"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem sum_realCenteredPairwiseRewardDiffVarianceProxy_const_eq_two_mul {K : Nat} (spec : ETC.Spec K) (best : Fin K) (sigma2 : NNReal) (a : Fin K) (hne : a ≠ best) : (Finset.range (spec.explorationPulls * K)).sum (fun t => ETC.realCenteredPairwiseRewardDiffVarianceProxy spec best best (fun _ _ => sigma2) a t) = (2 : NNReal) * (spec.explorationPulls : NNReal) * sigma2","missing":[],"search":"sum_realcenteredpairwiserewarddiffvarianceproxy_const_eq_two_mul banditrlproof.etc.sum_realcenteredpairwiserewarddiffvarianceproxy_const_eq_two_mul the constant pairwise proxy sums to exactly `2 * m * sigma2`. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_measure_realExplorationArgmaxCommit_eq_arm_le_exp_of_infinitePi_kernel","label":"real_measure_realExplorationArgmaxCommit_eq_arm_le_exp_of_infinitePi_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_measure_realExplorationArgmaxCommit_eq_arm_le_exp_of_infinitePi_kernel","description":"Exact native Real single-arm wrong-commit tail under action-matched independent kernel coordinates.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-6ebf57d3aba2","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":950,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:313"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_measure_realExplorationArgmaxCommit_eq_arm_le_exp_of_infinitePi_kernel {K : Nat} (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (a : Fin K) (hne : a ≠ ETC.realKernelBestArm spec.hK nu) : let best := ETC.realKernelBestArm spec.hK nu let coordLaw := fun t : Nat => nu (ETC.exploreArm spec t) (Measure.infinitePi coordLaw).real {trajectory : RewardTrace Real | ETC.realExplorationArgmaxCommit spec best trajectory = a} <= Real.exp (-(spec.explorationPulls : Real) * (realKernelGap nu a) ^ 2 / (4 * (sigma2 : Real)))","missing":[],"search":"real_measure_realexplorationargmaxcommit_eq_arm_le_exp_of_infinitepi_kernel banditrlproof.etc.real_measure_realexplorationargmaxcommit_eq_arm_le_exp_of_infinitepi_kernel exact native real single-arm wrong-commit tail under action-matched independent kernel coordinates. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_le_exp_of_infinitePi_kernel","label":"integral_real_pullCount_realExplorationArgmaxAction_le_exp_of_infinitePi_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_le_exp_of_infinitePi_kernel","description":"Exact native Real per-arm expected pull-count bound under `infinitePi`.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-ed6b1748f72d","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":951,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:439"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_realExplorationArgmaxAction_le_exp_of_infinitePi_kernel {K : Nat} (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (a : Fin K) (n : Nat) (hn : K * spec.explorationPulls <= n) (hne : a ≠ ETC.realKernelBestArm spec.hK nu) : let best := ETC.realKernelBestArm spec.hK nu let coordLaw := fun t : Nat => nu (ETC.exploreArm spec t) integral (Measure.infinitePi coordLaw) (fun trajectory : RewardTrace Real => (pullCount (ETC.realExplorationArgmaxAction spec best trajectory) a n : Real)) <= (spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * Real.exp (-(spec.explorationPulls : Real) * (realKernelGap nu a) ^ 2 / (4 *…","missing":[],"search":"integral_real_pullcount_realexplorationargmaxaction_le_exp_of_infinitepi_kernel banditrlproof.etc.integral_real_pullcount_realexplorationargmaxaction_le_exp_of_infinitepi_kernel exact native real per-arm expected pull-count bound under `infinitepi`. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_infinitePi_kernel","label":"integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_infinitePi_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_infinitePi_kernel","description":"Exact LML-shaped finite-sum native Real kernel regret bound under the canonical independent exploration-coordinate law.","url":"../modules/banditrlproof-algorithms-etcrealinfinitepitail/index.html#decl-0e7560ae5236","parent":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","order":952,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealInfinitePiTail"],["Source","BanditRLProof/Algorithms/ETCRealInfinitePiTail.lean:479"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_infinitePi_kernel {K : Nat} (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) : let best := ETC.realKernelBestArm spec.hK nu let coordLaw := fun t : Nat => nu (ETC.exploreArm spec t) integral (Measure.infinitePi coordLaw) (fun trajectory : RewardTrace Real => realKernelRegret nu (ETC.realExplorationArgmaxAction spec best trajectory) n) <= (Finset.univ : Finset (Fin K)).sum (fun arm => realKernelGap nu arm * ((spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * Real.exp (-(spec.explorationPulls : Real) * (realKerne…","missing":[],"search":"integral_realkernelregret_realexplorationargmaxaction_le_exact_sum_of_infinitepi_kernel banditrlproof.etc.integral_realkernelregret_realexplorationargmaxaction_le_exact_sum_of_infinitepi_kernel exact lml-shaped finite-sum native real kernel regret bound under the canonical independent exploration-coordinate law. theorem compiled","shard":"modules/8da6ef3ba368d237.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.RealStationaryETCSequence","label":"RealStationaryETCSequence","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ETC.RealStationaryETCSequence","description":"The exact consequences of a stationary Real ETC algorithm-environment sequence used by the local regret route. The fields correspond to the pinned source's `IsAlgEnvSeq` measurability and feedback fields together with `ETC.arm_of_lt`, `ETC.arm_mul`, and `ETC.arm_of_ge`. Conditional laws are stated as `condDistrib` equalities, which is the Mathlib-facing form consumed by ABRL.","url":"../modules/banditrlproof-algorithms-etcreallmlcompat/index.html#decl-29ad904116d7","parent":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","order":953,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ETCRealLMLCompat"],["Source","BanditRLProof/Algorithms/ETCRealLMLCompat.lean:27"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"structure RealStationaryETCSequence {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) : Prop where","missing":[],"search":"realstationaryetcsequence banditrlproof.etc.realstationaryetcsequence the exact consequences of a stationary real etc algorithm-environment sequence used by the local regret route. the fields correspond to the pinned source's `isalgenvseq` measurability and feedback fields together with `etc.arm_of_lt`, `etc.arm_mul`, and `etc.arm_of_ge`. conditional laws are stated as `conddistrib` equalities, which is the mathlib-facing form consumed by abrl. structure compiled","shard":"modules/6d2389e9aefd14aa.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.regret_le_of_realStationaryETCSequence","label":"regret_le_of_realStationaryETCSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.regret_le_of_realStationaryETCSequence","description":"Exact native Real ETC regret from the bundled stationary sequence fields. This is the local theorem corresponding to the mathematical statement of the pinned LML `Bandits.ETC.regret_le`. A direct theorem about the imported LML `IsAlgEnvSeq` symbol still requires a common Lean/mathlib toolchain.","url":"../modules/banditrlproof-algorithms-etcreallmlcompat/index.html#decl-c386b5c0d9b8","parent":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","order":954,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealLMLCompat"],["Source","BanditRLProof/Algorithms/ETCRealLMLCompat.lean:75"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem regret_le_of_realStationaryETCSequence {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (h : ETC.RealStationaryETCSequence mu spec nu action reward) (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun x => x - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) : integral mu (fun omega => realKernelRegret nu (action omega) n) <= (Finset.univ : Finset (Fin K)).sum (fun arm => realKernelGap nu arm * ((spec.explorationPulls : Real) + ((n - K * spec.explorationPulls : Nat) : Real) * Real.exp (-(spec.explorationPulls : Real) * (realKernelGap nu arm) ^ 2 / (4 * (si…","missing":[],"search":"regret_le_of_realstationaryetcsequence banditrlproof.etc.regret_le_of_realstationaryetcsequence exact native real etc regret from the bundled stationary sequence fields. this is the local theorem corresponding to the mathematical statement of the pinned lml `bandits.etc.regret_le`. a direct theorem about the imported lml `isalgenvseq` symbol still requires a common lean/mathlib toolchain. theorem compiled","shard":"modules/6d2389e9aefd14aa.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationRewardPrefix","label":"realExplorationRewardPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationRewardPrefix","description":"The reward coordinates read during ETC exploration.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-eda4cc65af0f","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":955,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:21"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def realExplorationRewardPrefix {K : Nat} (spec : ETC.Spec K) (trajectory : RewardTrace Real) : Fin (spec.explorationPulls * K) -> Real","missing":[],"search":"realexplorationrewardprefix banditrlproof.etc.realexplorationrewardprefix the reward coordinates read during etc exploration. definition compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realRewardTraceOfExplorationPrefix","label":"realRewardTraceOfExplorationPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realRewardTraceOfExplorationPrefix","description":"Extend a finite exploration prefix by zero outside the exploration phase.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-1703c304e63b","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":956,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:26"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def realRewardTraceOfExplorationPrefix {K : Nat} (spec : ETC.Spec K) (rewardPrefix : Fin (spec.explorationPulls * K) -> Real) : RewardTrace Real","missing":[],"search":"realrewardtraceofexplorationprefix banditrlproof.etc.realrewardtraceofexplorationprefix extend a finite exploration prefix by zero outside the exploration phase. definition compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationRewardPrefix_realRewardTraceOfExplorationPrefix","label":"realExplorationRewardPrefix_realRewardTraceOfExplorationPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationRewardPrefix_realRewardTraceOfExplorationPrefix","description":"Taking the exploration prefix after extension recovers the original prefix.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-1a5eecc26dc2","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":957,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:31"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"@[simp] theorem realExplorationRewardPrefix_realRewardTraceOfExplorationPrefix {K : Nat} (spec : ETC.Spec K) (rewardPrefix : Fin (spec.explorationPulls * K) -> Real) : ETC.realExplorationRewardPrefix spec (ETC.realRewardTraceOfExplorationPrefix spec rewardPrefix) = rewardPrefix","missing":[],"search":"realexplorationrewardprefix_realrewardtraceofexplorationprefix banditrlproof.etc.realexplorationrewardprefix_realrewardtraceofexplorationprefix taking the exploration prefix after extension recovers the original prefix. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realExplorationRewardPrefix","label":"measurable_realExplorationRewardPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realExplorationRewardPrefix","description":"Prefix extraction from a Real reward trace is measurable.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-ebb10f14363b","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":958,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:42"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realExplorationRewardPrefix {K : Nat} (spec : ETC.Spec K) : Measurable (ETC.realExplorationRewardPrefix spec)","missing":[],"search":"measurable_realexplorationrewardprefix banditrlproof.etc.measurable_realexplorationrewardprefix prefix extraction from a real reward trace is measurable. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realRewardTraceOfExplorationPrefix","label":"measurable_realRewardTraceOfExplorationPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realRewardTraceOfExplorationPrefix","description":"Zero extension of a finite Real reward prefix is measurable.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-309f82b2cf53","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":959,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:47"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realRewardTraceOfExplorationPrefix {K : Nat} (spec : ETC.Spec K) : Measurable (ETC.realRewardTraceOfExplorationPrefix spec)","missing":[],"search":"measurable_realrewardtraceofexplorationprefix banditrlproof.etc.measurable_realrewardtraceofexplorationprefix zero extension of a finite real reward prefix is measurable. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realExplorationRewardPrefix_comp","label":"measurable_realExplorationRewardPrefix_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realExplorationRewardPrefix_comp","description":"Timewise measurable rewards give a measurable finite exploration prefix.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-ae93d41b4fd0","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":960,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:58"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realExplorationRewardPrefix_comp {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (spec : ETC.Spec K) (reward : Omega -> RewardTrace Real) (hreward : forall t, Measurable (fun omega => reward omega t)) : Measurable (fun omega => ETC.realExplorationRewardPrefix spec (reward omega))","missing":[],"search":"measurable_realexplorationrewardprefix_comp banditrlproof.etc.measurable_realexplorationrewardprefix_comp timewise measurable rewards give a measurable finite exploration prefix. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.sumRewards_eq_of_eq_on_lt","label":"sumRewards_eq_of_eq_on_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.sumRewards_eq_of_eq_on_lt","description":"Reward sums agree when the reward traces agree before the horizon.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-beb9c87079ac","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":961,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:67"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_eq_of_eq_on_lt {Action Reward : Type} [DecidableEq Action] [AddCommMonoid Reward] (action : ActionTrace Action) (reward reward' : RewardTrace Reward) (a : Action) (n : Nat) (hreward : forall t, t < n -> reward t = reward' t) : sumRewards action reward a n = sumRewards action reward' a n","missing":[],"search":"sumrewards_eq_of_eq_on_lt banditrlproof.etc.sumrewards_eq_of_eq_on_lt reward sums agree when the reward traces agree before the horizon. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realEmpMeanAtExploration_eq_of_eq_on_exploration","label":"realEmpMeanAtExploration_eq_of_eq_on_exploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realEmpMeanAtExploration_eq_of_eq_on_exploration","description":"Native Real exploration means depend only on the exploration reward prefix.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-8d83d877fff7","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":962,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:79"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realEmpMeanAtExploration_eq_of_eq_on_exploration {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward reward' : RewardTrace Real) (hreward : forall t, t < spec.explorationPulls * K -> reward t = reward' t) (a : Fin K) : ETC.realEmpMeanAtExploration spec baseCommitArm reward a = ETC.realEmpMeanAtExploration spec baseCommitArm reward' a","missing":[],"search":"realempmeanatexploration_eq_of_eq_on_exploration banditrlproof.etc.realempmeanatexploration_eq_of_eq_on_exploration native real exploration means depend only on the exploration reward prefix. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit_eq_of_eq_on_exploration","label":"realExplorationArgmaxCommit_eq_of_eq_on_exploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationArgmaxCommit_eq_of_eq_on_exploration","description":"Native Real empirical argmax commit depends only on exploration rewards.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-1ec6e8080098","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":963,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:94"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realExplorationArgmaxCommit_eq_of_eq_on_exploration {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (reward reward' : RewardTrace Real) (hreward : forall t, t < spec.explorationPulls * K -> reward t = reward' t) : ETC.realExplorationArgmaxCommit spec baseCommitArm reward = ETC.realExplorationArgmaxCommit spec baseCommitArm reward'","missing":[],"search":"realexplorationargmaxcommit_eq_of_eq_on_exploration banditrlproof.etc.realexplorationargmaxcommit_eq_of_eq_on_exploration native real empirical argmax commit depends only on exploration rewards. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit_realRewardTraceOf_prefix_eq","label":"realExplorationArgmaxCommit_realRewardTraceOf_prefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationArgmaxCommit_realRewardTraceOf_prefix_eq","description":"Extending the extracted prefix does not change the ETC commit arm.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-134072dca94d","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":964,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:108"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realExplorationArgmaxCommit_realRewardTraceOf_prefix_eq {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (trajectory : RewardTrace Real) : ETC.realExplorationArgmaxCommit spec baseCommitArm trajectory = ETC.realExplorationArgmaxCommit spec baseCommitArm (ETC.realRewardTraceOfExplorationPrefix spec (ETC.realExplorationRewardPrefix spec trajectory))","missing":[],"search":"realexplorationargmaxcommit_realrewardtraceof_prefix_eq banditrlproof.etc.realexplorationargmaxcommit_realrewardtraceof_prefix_eq extending the extracted prefix does not change the etc commit arm. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationArgmaxAction_realRewardTraceOf_prefix_eq","label":"realExplorationArgmaxAction_realRewardTraceOf_prefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationArgmaxAction_realRewardTraceOf_prefix_eq","description":"Extending the extracted prefix does not change the native Real ETC action.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-2c35cedfd940","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":965,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:121"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realExplorationArgmaxAction_realRewardTraceOf_prefix_eq {K : Nat} (spec : ETC.Spec K) (baseCommitArm : Fin K) (trajectory : RewardTrace Real) : ETC.realExplorationArgmaxAction spec baseCommitArm trajectory = ETC.realExplorationArgmaxAction spec baseCommitArm (ETC.realRewardTraceOfExplorationPrefix spec (ETC.realExplorationRewardPrefix spec trajectory))","missing":[],"search":"realexplorationargmaxaction_realrewardtraceof_prefix_eq banditrlproof.etc.realexplorationargmaxaction_realrewardtraceof_prefix_eq extending the extracted prefix does not change the native real etc action. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realKernelRegret_of_forall_measurable_action","label":"measurable_realKernelRegret_of_forall_measurable_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realKernelRegret_of_forall_measurable_action","description":"Kernel regret is measurable from a timewise measurable finite-arm action.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-06d973a2fbf8","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":966,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:132"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realKernelRegret_of_forall_measurable_action {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (nu : ProbabilityTheory.Kernel (Fin K) Real) (action : Omega -> ActionTrace (Fin K)) (haction : forall t, Measurable (fun omega => action omega t)) (n : Nat) : Measurable (fun omega => realKernelRegret nu (action omega) n)","missing":[],"search":"measurable_realkernelregret_of_forall_measurable_action banditrlproof.etc.measurable_realkernelregret_of_forall_measurable_action kernel regret is measurable from a timewise measurable finite-arm action. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelRegretOfExplorationPrefix","label":"realKernelRegretOfExplorationPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelRegretOfExplorationPrefix","description":"The finite-prefix functional whose value is native Real ETC kernel regret.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-cbef35961e50","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":967,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:149"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def realKernelRegretOfExplorationPrefix {K : Nat} (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) (baseCommitArm : Fin K) (n : Nat) (rewardPrefix : Fin (spec.explorationPulls * K) -> Real) : Real","missing":[],"search":"realkernelregretofexplorationprefix banditrlproof.etc.realkernelregretofexplorationprefix the finite-prefix functional whose value is native real etc kernel regret. definition compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realKernelRegretOfExplorationPrefix","label":"measurable_realKernelRegretOfExplorationPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realKernelRegretOfExplorationPrefix","description":"The finite-prefix kernel-regret functional is measurable.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-d506dc171eac","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":968,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:158"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realKernelRegretOfExplorationPrefix {K : Nat} (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) (baseCommitArm : Fin K) (n : Nat) : Measurable (ETC.realKernelRegretOfExplorationPrefix spec nu baseCommitArm n)","missing":[],"search":"measurable_realkernelregretofexplorationprefix banditrlproof.etc.measurable_realkernelregretofexplorationprefix the finite-prefix kernel-regret functional is measurable. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelRegret_realExplorationArgmaxAction_eq_prefixFunctional","label":"realKernelRegret_realExplorationArgmaxAction_eq_prefixFunctional","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelRegret_realExplorationArgmaxAction_eq_prefixFunctional","description":"Native Real ETC kernel regret factors through the exploration prefix.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-08d0e078a10d","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":969,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:184"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelRegret_realExplorationArgmaxAction_eq_prefixFunctional {K : Nat} (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) (baseCommitArm : Fin K) (trajectory : RewardTrace Real) (n : Nat) : realKernelRegret nu (ETC.realExplorationArgmaxAction spec baseCommitArm trajectory) n = ETC.realKernelRegretOfExplorationPrefix spec nu baseCommitArm n (ETC.realExplorationRewardPrefix spec trajectory)","missing":[],"search":"realkernelregret_realexplorationargmaxaction_eq_prefixfunctional banditrlproof.etc.realkernelregret_realexplorationargmaxaction_eq_prefixfunctional native real etc kernel regret factors through the exploration prefix. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realKernelRegret_eq_of_action_eq_on_lt","label":"realKernelRegret_eq_of_action_eq_on_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realKernelRegret_eq_of_action_eq_on_lt","description":"Kernel regret agrees when two action traces agree before the horizon.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-d312bb4a8e7e","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":970,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:196"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realKernelRegret_eq_of_action_eq_on_lt {K : Nat} (nu : ProbabilityTheory.Kernel (Fin K) Real) (action action' : ActionTrace (Fin K)) (n : Nat) (haction : forall t, t < n -> action t = action' t) : realKernelRegret nu action n = realKernelRegret nu action' n","missing":[],"search":"realkernelregret_eq_of_action_eq_on_lt banditrlproof.etc.realkernelregret_eq_of_action_eq_on_lt kernel regret agrees when two action traces agree before the horizon. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.real_trajMeasure_const_eq_infinitePi","label":"real_trajMeasure_const_eq_infinitePi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.real_trajMeasure_const_eq_infinitePi","description":"A constant-kernel Ionescu-Tulcea trajectory is the infinite product law.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-d993423d76b0","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":971,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:208"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem real_trajMeasure_const_eq_infinitePi (coordLaw : Nat -> Measure Real) [forall t, IsProbabilityMeasure (coordLaw t)] : ProbabilityTheory.Kernel.trajMeasure (coordLaw 0) (fun i => ProbabilityTheory.Kernel.const ((j : Finset.Iic i) -> Real) (coordLaw (i + 1))) = Measure.infinitePi coordLaw","missing":[],"search":"real_trajmeasure_const_eq_infinitepi banditrlproof.etc.real_trajmeasure_const_eq_infinitepi a constant-kernel ionescu-tulcea trajectory is the infinite product law. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationPrefixOfFiniteRewardHistory","label":"realExplorationPrefixOfFiniteRewardHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationPrefixOfFiniteRewardHistory","description":"Convert the inclusive history through `horizon - 1` to a `Fin horizon` prefix.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-eeba14fb8643","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":972,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:231"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def realExplorationPrefixOfFiniteRewardHistory {K : Nat} (spec : ETC.Spec K) (history : History.FiniteRewardHistory Real (spec.explorationPulls * K - 1)) : Fin (spec.explorationPulls * K) -> Real","missing":[],"search":"realexplorationprefixoffiniterewardhistory banditrlproof.etc.realexplorationprefixoffiniterewardhistory convert the inclusive history through `horizon - 1` to a `fin horizon` prefix. definition compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.measurable_realExplorationPrefixOfFiniteRewardHistory","label":"measurable_realExplorationPrefixOfFiniteRewardHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.measurable_realExplorationPrefixOfFiniteRewardHistory","description":"The inclusive-history to `Fin` exploration-prefix conversion is measurable.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-100e9ff6d7f0","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":973,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:240"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem measurable_realExplorationPrefixOfFiniteRewardHistory {K : Nat} (spec : ETC.Spec K) : Measurable (ETC.realExplorationPrefixOfFiniteRewardHistory spec)","missing":[],"search":"measurable_realexplorationprefixoffiniterewardhistory banditrlproof.etc.measurable_realexplorationprefixoffiniterewardhistory the inclusive-history to `fin` exploration-prefix conversion is measurable. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.realExplorationPrefixOfFiniteRewardHistory_of_trace","label":"realExplorationPrefixOfFiniteRewardHistory_of_trace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.realExplorationPrefixOfFiniteRewardHistory_of_trace","description":"Conversion of a trace's inclusive history is its `Fin` exploration prefix.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-5143e0c2b02b","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":974,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:248"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem realExplorationPrefixOfFiniteRewardHistory_of_trace {K : Nat} (spec : ETC.Spec K) (trajectory : RewardTrace Real) : ETC.realExplorationPrefixOfFiniteRewardHistory spec (History.finiteRewardHistoryOfTrace trajectory (spec.explorationPulls * K - 1)) = ETC.realExplorationRewardPrefix spec trajectory","missing":[],"search":"realexplorationprefixoffiniterewardhistory_of_trace banditrlproof.etc.realexplorationprefixoffiniterewardhistory_of_trace conversion of a trace's inclusive history is its `fin` exploration prefix. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_prefixLaw_eq_infinitePi","label":"integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_prefixLaw_eq_infinitePi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_prefixLaw_eq_infinitePi","description":"Equality of finite exploration-prefix laws transports the canonical native Real ETC exact regret bound to an arbitrary reward process.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-61afb6d3d6c3","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":975,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:260"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_prefixLaw_eq_infinitePi {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) (reward : Omega -> RewardTrace Real) (hreward : forall t, Measurable (fun omega => reward omega t)) (hprefixLaw : Measure.map (fun omega => ETC.realExplorationRewardPrefix spec (reward omega)) mu = Measure.map (ETC.realExplorationRewardPrefix spec) (Measure.infinitePi (fun t : Nat => nu (ETC.exploreArm spec t)))) : let best := ETC.realKernelBestArm spec.hK nu integral mu (f…","missing":[],"search":"integral_realkernelregret_realexplorationargmaxaction_le_exact_sum_of_prefixlaw_eq_infinitepi banditrlproof.etc.integral_realkernelregret_realexplorationargmaxaction_le_exact_sum_of_prefixlaw_eq_infinitepi equality of finite exploration-prefix laws transports the canonical native real etc exact regret bound to an arbitrary reward process. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_prefixLaw_eq_infinitePi","label":"integral_realKernelRegret_externalAction_le_exact_sum_of_prefixLaw_eq_infinitePi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_prefixLaw_eq_infinitePi","description":"External-action version of the finite-prefix law transport theorem. The two remaining upstream obligations are explicit: finite-prefix reward law equality and almost-sure equality with the local native Real ETC action.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-4823a1f92a42","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":976,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:364"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_exact_sum_of_prefixLaw_eq_infinitePi {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (hreward : forall t, Measurable (fun omega => reward omega t)) (hprefixLaw : Measure.map (fun omega => ETC.realExplorationRewardPrefix spec (reward omega)) mu = Measure.map (ETC.realExplorationRewardPrefix spec) (Measure.infinitePi (fun t : Nat => nu (ETC.exploreArm spec t)))) (haction : Filter.Eventually (fun…","missing":[],"search":"integral_realkernelregret_externalaction_le_exact_sum_of_prefixlaw_eq_infinitepi banditrlproof.etc.integral_realkernelregret_externalaction_le_exact_sum_of_prefixlaw_eq_infinitepi external-action version of the finite-prefix law transport theorem. the two remaining upstream obligations are explicit: finite-prefix reward law equality and almost-sure equality with the local native real etc action. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_initial_map_eq_condDistrib","label":"integral_realKernelRegret_externalAction_le_exact_sum_of_initial_map_eq_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_initial_map_eq_condDistrib","description":"External native Real exact ETC regret from the scheduled exploration-arm initial marginal and successor conditional reward laws. This is the direct law surface needed before mapping an upstream `IsAlgEnvSeq` witness.","url":"../modules/banditrlproof-algorithms-etcrealprefixlawtransport/index.html#decl-3d8aa33a6ea7","parent":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","order":977,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealPrefixLawTransport"],["Source","BanditRLProof/Algorithms/ETCRealPrefixLawTransport.lean:426"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_exact_sum_of_initial_map_eq_condDistrib {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (hreward : forall t, Measurable (fun omega => reward omega t)) (hzero : Measure.map (fun omega => reward omega 0) mu = nu (ETC.exploreArm spec 0)) (hcond : forall i, i < spec.explorationPulls * K - 1 -> ProbabilityTheory.condDistrib (fun omega => reward omega (i + 1)) (fun omega => History.finiteRewardHistor…","missing":[],"search":"integral_realkernelregret_externalaction_le_exact_sum_of_initial_map_eq_conddistrib banditrlproof.etc.integral_realkernelregret_externalaction_le_exact_sum_of_initial_map_eq_conddistrib external native real exact etc regret from the scheduled exploration-arm initial marginal and successor conditional reward laws. this is the direct law surface needed before mapping an upstream `isalgenvseq` witness. theorem compiled","shard":"modules/6ea3d46cfbf19c2d.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib","label":"integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib","description":"External native Real exact ETC regret from action-selected initial and full-history successor feedback laws. The exploration action identities turn each action-selected kernel into the constant law of the scheduled round-robin arm. The complete action/reward history is then projected to the reward-only prefix expected by the compiled finite-prefix uniqueness theorem.","url":"../modules/banditrlproof-algorithms-etcrealsourceadapter/index.html#decl-75f4dc56c2ff","parent":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","order":978,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRealSourceAdapter"],["Source","BanditRLProof/Algorithms/ETCRealSourceAdapter.lean:27"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (spec : ETC.Spec K) (nu : ProbabilityTheory.Kernel (Fin K) Real) [ProbabilityTheory.IsMarkovKernel nu] (sigma2 : NNReal) (hsubG : forall arm, ProbabilityTheory.HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (hm : 0 < spec.explorationPulls) (n : Nat) (hn : K * spec.explorationPulls <= n) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (hactionExplore : forall t, t < spec.explorationPulls * K -> (fun omega => action omega t) =ᵐ[mu] fun _omega => ETC.exploreArm spec t) (hzero : ProbabilityTheory.…","missing":[],"search":"integral_realkernelregret_externalaction_le_exact_sum_of_actiondependent_actionrewardhistory_conddistrib banditrlproof.etc.integral_realkernelregret_externalaction_le_exact_sum_of_actiondependent_actionrewardhistory_conddistrib external native real exact etc regret from action-selected initial and full-history successor feedback laws. the exploration action identities turn each action-selected kernel into the constant law of the scheduled round-robin arm. the complete action/reward history is then projected to the reward-only prefix expected by the compiled finite-prefix uniqueness theorem. theorem compiled","shard":"modules/3f97a6d8cd07d916.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_exploreArm_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","label":"pseudoRegret_exploreArm_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_exploreArm_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","description":"The pseudo-regret of the pure round-robin ETC exploration prefix is bounded by the sum of arm gaps times the configured number of exploration pulls per arm. This is the `ETC-EXPLORATION-REGRET-BOUND` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-0dc65ad3eae2","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":979,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:22"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_exploreArm_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) : pseudoRegret model (ETC.exploreArm spec) (spec.explorationPulls * K) <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat) : Rat))","missing":[],"search":"pseudoregret_explorearm_explorationpulls_mul_k_le_sum_gap_mul_explorationpulls banditrlproof.etc.pseudoregret_explorearm_explorationpulls_mul_k_le_sum_gap_mul_explorationpulls the pseudo-regret of the pure round-robin etc exploration prefix is bounded by the sum of arm gaps times the configured number of exploration pulls per arm. this is the `etc-exploration-regret-bound` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","description":"The pseudo-regret of the fixed-commit ETC trace at the exploration horizon is bounded by the sum of arm gaps times the configured number of exploration pulls per arm. This is the `ETC-ACTION-WITH-COMMIT-EXPLORATION-HORIZON-REGRET-BOUND` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-5d039af34d0a","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":980,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:48"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K) <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat) : Rat))","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_le_sum_gap_mul_explorationpulls banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_le_sum_gap_mul_explorationpulls the pseudo-regret of the fixed-commit etc trace at the exploration horizon is bounded by the sum of arm gaps times the configured number of exploration pulls per arm. this is the `etc-action-with-commit-exploration-horizon-regret-bound` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_suffix_count_budget","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_suffix_count_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_suffix_count_budget","description":"The fixed-commit ETC trace at an exploration horizon plus suffix `r` is bounded by the gap-weighted count budget that allocates `explorationPulls` to every arm and the suffix budget only to the committed arm. This is the `ETC-ACTION-WITH-COMMIT-SUFFIX-COUNT-BUDGET-REGRET` project-local deterministic regret scaffold. It deliberately keeps the right-hand side in the unsimplified per-arm budget form.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-e340439dbfa7","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":981,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:79"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_suffix_count_budget {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (r : Nat) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K + r) <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a * (((spec.explorationPulls + (if commitArm = a then r else 0) : Nat) : Rat)))","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_suffix_count_budget banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_suffix_count_budget the fixed-commit etc trace at an exploration horizon plus suffix `r` is bounded by the gap-weighted count budget that allocates `explorationpulls` to every arm and the suffix budget only to the committed arm. this is the `etc-action-with-commit-suffix-count-budget-regret` project-local deterministic regret scaffold. it deliberately keeps the right-hand side in the unsimplified per-arm budget form. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix","description":"The fixed-commit ETC trace after an exploration horizon plus suffix `r` also satisfies the coarser uniform count-budget regret bound where every arm is allowed `explorationPulls + r` pulls. This is the `ETC-ACTION-WITH-COMMIT-COARSE-SUFFIX-REGRET-BOUND` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-9baa6d979677","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":982,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:114"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (r : Nat) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K + r) <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * ((((spec.explorationPulls + r : Nat) : Rat)))","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_explorationpulls_add_suffix banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_explorationpulls_add_suffix the fixed-commit etc trace after an exploration horizon plus suffix `r` also satisfies the coarser uniform count-budget regret bound where every arm is allowed `explorationpulls + r` pulls. this is the `etc-action-with-commit-coarse-suffix-regret-bound` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_add_suffix_gap","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_add_suffix_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_add_suffix_gap","description":"The fixed-commit ETC trace splits pseudo-regret after the exploration horizon into the exploration-horizon pseudo-regret plus one committed-arm gap for each suffix pull. This is the `ETC-ACTION-WITH-COMMIT-PHASE-SPLIT-REGRET` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-9f9f4d082245","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":983,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:149"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_add_suffix_gap {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (r : Nat) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K + r) = pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K) + (((r : Nat) : Rat) * model.gap commitArm)","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_add_eq_add_suffix_gap banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_add_eq_add_suffix_gap the fixed-commit etc trace splits pseudo-regret after the exploration horizon into the exploration-horizon pseudo-regret plus one committed-arm gap for each suffix pull. this is the `etc-action-with-commit-phase-split-regret` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_of_commitArm_eq_bestArm","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_of_commitArm_eq_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_of_commitArm_eq_bestArm","description":"If the fixed commit arm is the model's selected best arm, extending the ETC trace past the exploration horizon adds no pseudo-regret. This is the `ETC-ACTION-WITH-COMMIT-BESTARM-SUFFIX-NO-REGRET` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-ff56221ab662","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":984,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:184"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_of_commitArm_eq_bestArm {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (r : Nat) (hcommit : commitArm = model.bestArm) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K + r) = pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K)","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_add_eq_of_commitarm_eq_bestarm banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_add_eq_of_commitarm_eq_bestarm if the fixed commit arm is the model's selected best arm, extending the etc trace past the exploration horizon adds no pseudo-regret. this is the `etc-action-with-commit-bestarm-suffix-no-regret` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_of_commitArm_eq_bestArm","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_of_commitArm_eq_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_of_commitArm_eq_bestArm","description":"If the fixed commit arm is the model's selected best arm, the ETC regret after any post-exploration suffix is bounded by the exploration-horizon regret budget. This is the `ETC-ACTION-WITH-COMMIT-BESTARM-SUFFIX-REGRET-BOUND` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-32642bd0d869","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":985,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:209"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_of_commitArm_eq_bestArm {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (r : Nat) (hcommit : commitArm = model.bestArm) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K + r) <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat) : Rat))","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_explorationpulls_of_commitarm_eq_bestarm banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_explorationpulls_of_commitarm_eq_bestarm if the fixed commit arm is the model's selected best arm, the etc regret after any post-exploration suffix is bounded by the exploration-horizon regret budget. this is the `etc-action-with-commit-bestarm-suffix-regret-bound` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix_gap","label":"pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix_gap","description":"The phase-split equality and the exploration-horizon regret bound combine into a deterministic fixed-commit ETC regret bound with an explicit committed-arm suffix term. This is the `ETC-ACTION-WITH-COMMIT-PHASE-SPLIT-REGRET-BOUND` project-local deterministic regret scaffold.","url":"../modules/banditrlproof-algorithms-etcregretlemmas/index.html#decl-78f929f037e4","parent":"module:BanditRLProof.Algorithms.ETCRegretLemmas","order":986,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCRegretLemmas"],["Source","BanditRLProof/Algorithms/ETCRegretLemmas.lean:239"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix_gap {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (r : Nat) : pseudoRegret model (ETC.actionWithCommit spec commitArm) (spec.explorationPulls * K + r) <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat) : Rat)) + (((r : Nat) : Rat) * model.gap commitArm)","missing":[],"search":"pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_explorationpulls_add_suffix_gap banditrlproof.etc.pseudoregret_actionwithcommit_explorationpulls_mul_k_add_le_sum_gap_mul_explorationpulls_add_suffix_gap the phase-split equality and the exploration-horizon regret bound combine into a deterministic fixed-commit etc regret bound with an explicit committed-arm suffix term. this is the `etc-action-with-commit-phase-split-regret-bound` project-local deterministic regret scaffold. theorem compiled","shard":"modules/c0416491b1744f12.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff","label":"centeredPairwiseRewardDiff","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseRewardDiff","description":"Real-valued centered pairwise reward-difference summand for arm `a` against the model's selected best arm at one ETC exploration-horizon time index. The expression remains deterministic and pointwise. Future probability leaves may impose independence or sub-Gaussian contracts on this function.","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html#decl-f083d8347351","parent":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","order":987,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCSumRewardsDiff"],["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean:28"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredPairwiseRewardDiff {Omega : Type u} {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (t : Nat) (omega : Omega) : Real","missing":[],"search":"centeredpairwiserewarddiff banditrlproof.etc.centeredpairwiserewarddiff real-valued centered pairwise reward-difference summand for arm `a` against the model's selected best arm at one etc exploration-horizon time index. the expression remains deterministic and pointwise. future probability leaves may impose independence or sub-gaussian contracts on this function. definition compiled","shard":"modules/869b43c0a8211722.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.centeredPairwiseGapThreshold","label":"centeredPairwiseGapThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.centeredPairwiseGapThreshold","description":"Real-valued threshold corresponding to `explorationPulls` copies of the mean gap between the selected best arm and arm `a`.","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html#decl-6d4a15ae1758","parent":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","order":988,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCSumRewardsDiff"],["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean:44"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredPairwiseGapThreshold {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (a : Fin K) : Real","missing":[],"search":"centeredpairwisegapthreshold banditrlproof.etc.centeredpairwisegapthreshold real-valued threshold corresponding to `explorationpulls` copies of the mean gap between the selected best arm and arm `a`. definition compiled","shard":"modules/869b43c0a8211722.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","label":"selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","description":"The selected centered reward sum for one arm equals its selected reward total minus its pull count times the supplied mean. This helper is part of the `ETC-SUMREWARDS-PAIRWISE-DIFF-FINSET` bridge. It uses the existing Mathlib-backed `sumRewards` and `pullCount` Finset wrappers.","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html#decl-2dee71e21c67","parent":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","order":989,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCSumRewardsDiff"],["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean:59"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (a : Action) (n : Nat) (mu : Rat) : (Finset.range n).sum (fun t : Nat => if action t = a then reward t - mu else 0) = sumRewards action reward a n - (pullCount action a n : Rat) * mu","missing":[],"search":"selectedsubmean_sum_eq_sumrewards_sub_pullcount_mul banditrlproof.etc.selectedsubmean_sum_eq_sumrewards_sub_pullcount_mul the selected centered reward sum for one arm equals its selected reward total minus its pull count times the supplied mean. this helper is part of the `etc-sumrewards-pairwise-diff-finset` bridge. it uses the existing mathlib-backed `sumrewards` and `pullcount` finset wrappers. theorem compiled","shard":"modules/869b43c0a8211722.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","label":"meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","description":"The selected negative centered reward sum for one arm equals its pull count times the supplied mean minus its selected reward total. This is the best-arm companion to `selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul`.","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html#decl-fb239d0be81b","parent":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","order":990,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCSumRewardsDiff"],["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean:106"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (a : Action) (n : Nat) (mu : Rat) : (Finset.range n).sum (fun t : Nat => if action t = a then mu - reward t else 0) = (pullCount action a n : Rat) * mu - sumRewards action reward a n","missing":[],"search":"meansubselected_sum_eq_pullcount_mul_sub_sumrewards banditrlproof.etc.meansubselected_sum_eq_pullcount_mul_sub_sumrewards the selected negative centered reward sum for one arm equals its pull count times the supplied mean minus its selected reward total. this is the best-arm companion to `selectedsubmean_sum_eq_sumrewards_sub_pullcount_mul`. theorem compiled","shard":"modules/869b43c0a8211722.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.sumRewards_le_imp_centered_pairwise_sum_ge","label":"sumRewards_le_imp_centered_pairwise_sum_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.sumRewards_le_imp_centered_pairwise_sum_ge","description":"If two arms have the same pull count by a horizon and the fixed-horizon reward sum of `b` is at most that of `a`, then the centered pairwise reward-difference finite sum is at least `m * (muB - muA)`. This is the deterministic algebra core of `ETC-SUMREWARDS-PAIRWISE-DIFF-FINSET`. The equal-count assumptions are kept explicit so ETC can discharge them with its exploration-horizon count theorem.","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html#decl-578c287a8ba2","parent":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","order":991,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCSumRewardsDiff"],["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean:156"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_le_imp_centered_pairwise_sum_ge {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (a b : Action) (n m : Nat) (muA muB : Rat) (hcount_a : pullCount action a n = m) (hcount_b : pullCount action b n = m) (hraw : sumRewards action reward b n <= sumRewards action reward a n) : (m : Rat) * (muB - muA) <= (Finset.range n).sum (fun t : Nat => (if action t = a then reward t - muA else 0) + (if action t = b then muB - reward t else 0))","missing":[],"search":"sumrewards_le_imp_centered_pairwise_sum_ge banditrlproof.etc.sumrewards_le_imp_centered_pairwise_sum_ge if two arms have the same pull count by a horizon and the fixed-horizon reward sum of `b` is at most that of `a`, then the centered pairwise reward-difference finite sum is at least `m * (mub - mua)`. this is the deterministic algebra core of `etc-sumrewards-pairwise-diff-finset`. the equal-count assumptions are kept explicit so etc can discharge them with its exploration-horizon count theorem. theorem compiled","shard":"modules/869b43c0a8211722.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.empMeanAtExploration_ge_best_event_subset_centered_pairwise_sum_event","label":"empMeanAtExploration_ge_best_event_subset_centered_pairwise_sum_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.empMeanAtExploration_ge_best_event_subset_centered_pairwise_sum_event","description":"Concrete ETC event inclusion from the empirical-mean comparison event to the centered pairwise reward-difference finite-sum event over the configured exploration horizon. This leaf instantiates the earlier abstract event-shape adapter with `idx := Finset.range (spec.explorationPulls * K)` and the centered non-best-minus-best reward-difference summands. It still does not prove independence, sub-Gaussianity, filtratio…","url":"../modules/banditrlproof-algorithms-etcsumrewardsdiff/index.html#decl-664b07fd9261","parent":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","order":992,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCSumRewardsDiff"],["Source","BanditRLProof/Algorithms/ETCSumRewardsDiff.lean:200"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem empMeanAtExploration_ge_best_event_subset_centered_pairwise_sum_event {Omega : Type u} {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (a : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : Set.Subset {omega : Omega | ETC.empMeanAtExploration spec commitArm (reward omega) a >= ETC.empMeanAtExploration spec commitArm (reward omega) model.bestArm} {omega : Omega | ETC.centeredPairwiseGapThreshold spec model a <= (Finset.range (spec.explorationPulls * K)).sum (fun t : Nat => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega)}","missing":[],"search":"empmeanatexploration_ge_best_event_subset_centered_pairwise_sum_event banditrlproof.etc.empmeanatexploration_ge_best_event_subset_centered_pairwise_sum_event concrete etc event inclusion from the empirical-mean comparison event to the centered pairwise reward-difference finite-sum event over the configured exploration horizon. this leaf instantiates the earlier abstract event-shape adapter with `idx := finset.range (spec.explorationpulls * k)` and the centered non-best-minus-best reward-difference summands. it still does not prove independence, sub-gaussianity, filtration, or final etc regret. theorem compiled","shard":"modules/869b43c0a8211722.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.actionWithCommit","label":"actionWithCommit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ETC.actionWithCommit","description":"Explore by round-robin until the configured horizon, then play `commitArm`.","url":"../modules/banditrlproof-algorithms-etctrace/index.html#decl-10b99c35b847","parent":"module:BanditRLProof.Algorithms.ETCTrace","order":993,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ETCTrace"],["Source","BanditRLProof/Algorithms/ETCTrace.lean:17"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"def actionWithCommit {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) : ActionTrace (Fin K)","missing":[],"search":"actionwithcommit banditrlproof.etc.actionwithcommit explore by round-robin until the configured horizon, then play `commitarm`. definition compiled","shard":"modules/383dc9cfc419125b.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.actionWithCommit_eq_exploreArm_of_lt","label":"actionWithCommit_eq_exploreArm_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.actionWithCommit_eq_exploreArm_of_lt","description":"During the configured exploration prefix, the phase-switching ETC trace agrees with the pure round-robin exploration trace. This is the `ETC-ACTION-WITH-COMMIT-EXPLORE-PHASE` project-local trace-boundary leaf.","url":"../modules/banditrlproof-algorithms-etctrace/index.html#decl-81d136a91cf0","parent":"module:BanditRLProof.Algorithms.ETCTrace","order":994,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTrace"],["Source","BanditRLProof/Algorithms/ETCTrace.lean:32"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"@[simp] theorem actionWithCommit_eq_exploreArm_of_lt {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) {t : Nat} (h : t < spec.explorationPulls * K) : ETC.actionWithCommit spec commitArm t = ETC.exploreArm spec t","missing":[],"search":"actionwithcommit_eq_explorearm_of_lt banditrlproof.etc.actionwithcommit_eq_explorearm_of_lt during the configured exploration prefix, the phase-switching etc trace agrees with the pure round-robin exploration trace. this is the `etc-action-with-commit-explore-phase` project-local trace-boundary leaf. theorem compiled","shard":"modules/383dc9cfc419125b.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.actionWithCommit_eq_commitArm_of_ge","label":"actionWithCommit_eq_commitArm_of_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.actionWithCommit_eq_commitArm_of_ge","description":"After the configured exploration prefix, the phase-switching ETC trace plays the supplied commit arm. This is the `ETC-ACTION-WITH-COMMIT-COMMIT-PHASE` project-local trace-boundary leaf.","url":"../modules/banditrlproof-algorithms-etctrace/index.html#decl-3a0589bbc171","parent":"module:BanditRLProof.Algorithms.ETCTrace","order":995,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTrace"],["Source","BanditRLProof/Algorithms/ETCTrace.lean:45"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"@[simp] theorem actionWithCommit_eq_commitArm_of_ge {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) {t : Nat} (h : spec.explorationPulls * K <= t) : ETC.actionWithCommit spec commitArm t = commitArm","missing":[],"search":"actionwithcommit_eq_commitarm_of_ge banditrlproof.etc.actionwithcommit_eq_commitarm_of_ge after the configured exploration prefix, the phase-switching etc trace plays the supplied commit arm. this is the `etc-action-with-commit-commit-phase` project-local trace-boundary leaf. theorem compiled","shard":"modules/383dc9cfc419125b.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.actionWithCommit_eq_bestArm_of_commitArm_eq_bestArm_of_explorationPulls_mul_K_le","label":"actionWithCommit_eq_bestArm_of_commitArm_eq_bestArm_of_explorationPulls_mul_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.actionWithCommit_eq_bestArm_of_commitArm_eq_bestArm_of_explorationPulls_mul_K_le","description":"After the configured exploration prefix, if the supplied commit arm is the model's selected best arm, the phase-switching ETC trace plays that best arm. This is the `ETC-ACTION-WITH-COMMIT-BESTARM-COMMIT-PHASE` project-local trace-boundary leaf.","url":"../modules/banditrlproof-algorithms-etctrace/index.html#decl-4c409ae777eb","parent":"module:BanditRLProof.Algorithms.ETCTrace","order":996,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTrace"],["Source","BanditRLProof/Algorithms/ETCTrace.lean:59"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem actionWithCommit_eq_bestArm_of_commitArm_eq_bestArm_of_explorationPulls_mul_K_le {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (t : Nat) (hcommit : commitArm = model.bestArm) (ht : spec.explorationPulls * K <= t) : ETC.actionWithCommit spec commitArm t = model.bestArm","missing":[],"search":"actionwithcommit_eq_bestarm_of_commitarm_eq_bestarm_of_explorationpulls_mul_k_le banditrlproof.etc.actionwithcommit_eq_bestarm_of_commitarm_eq_bestarm_of_explorationpulls_mul_k_le after the configured exploration prefix, if the supplied commit arm is the model's selected best arm, the phase-switching etc trace plays that best arm. this is the `etc-action-with-commit-bestarm-commit-phase` project-local trace-boundary leaf. theorem compiled","shard":"modules/383dc9cfc419125b.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_eq_pullCount_exploreArm_of_le","label":"pullCount_actionWithCommit_eq_pullCount_exploreArm_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_eq_pullCount_exploreArm_of_le","description":"On any prefix contained inside the configured exploration horizon, the fixed-commit ETC trace has the same pull counts as the pure round-robin exploration trace. This is the `ETC-ACTION-WITH-COMMIT-EXPLORE-PREFIX-PULLCOUNT` project-local trace/count transfer leaf.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-a4948727dc73","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":997,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:23"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_eq_pullCount_exploreArm_of_le {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) (n : Nat) (hn : n <= spec.explorationPulls * K) : pullCount (ETC.actionWithCommit spec commitArm) a n = pullCount (ETC.exploreArm spec) a n","missing":[],"search":"pullcount_actionwithcommit_eq_pullcount_explorearm_of_le banditrlproof.etc.pullcount_actionwithcommit_eq_pullcount_explorearm_of_le on any prefix contained inside the configured exploration horizon, the fixed-commit etc trace has the same pull counts as the pure round-robin exploration trace. this is the `etc-action-with-commit-explore-prefix-pullcount` project-local trace/count transfer leaf. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_eq","label":"pullCount_actionWithCommit_explorationPulls_mul_K_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_eq","description":"At the configured exploration horizon, the fixed-commit ETC trace has pulled each arm exactly `spec.explorationPulls` times. This is the `ETC-ACTION-WITH-COMMIT-EXPLORATION-HORIZON-COUNT` project-local trace/count adapter.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-764fe58563e9","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":998,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:53"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_explorationPulls_mul_K_eq {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) : pullCount (ETC.actionWithCommit spec commitArm) a (spec.explorationPulls * K) = spec.explorationPulls","missing":[],"search":"pullcount_actionwithcommit_explorationpulls_mul_k_eq banditrlproof.etc.pullcount_actionwithcommit_explorationpulls_mul_k_eq at the configured exploration horizon, the fixed-commit etc trace has pulled each arm exactly `spec.explorationpulls` times. this is the `etc-action-with-commit-exploration-horizon-count` project-local trace/count adapter. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_pos","label":"pullCount_actionWithCommit_explorationPulls_mul_K_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_pos","description":"At the configured exploration horizon, every arm in the fixed-commit ETC trace has a positive pull count whenever the configured number of exploration pulls is positive. This is the first denominator-positivity leaf for future empirical-mean construction. It stays purely deterministic and Nat-valued.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-2dc3331246a1","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":999,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:76"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_explorationPulls_mul_K_pos {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : 0 < pullCount (ETC.actionWithCommit spec commitArm) a (spec.explorationPulls * K)","missing":[],"search":"pullcount_actionwithcommit_explorationpulls_mul_k_pos banditrlproof.etc.pullcount_actionwithcommit_explorationpulls_mul_k_pos at the configured exploration horizon, every arm in the fixed-commit etc trace has a positive pull count whenever the configured number of exploration pulls is positive. this is the first denominator-positivity leaf for future empirical-mean construction. it stays purely deterministic and nat-valued. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_pos","label":"ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_pos","description":"Rat-cast form of the fixed-commit ETC exploration pull-count positivity leaf. This is the first denominator adapter for future Rat-valued empirical means. It only transports the compiled Nat positivity theorem across the Nat-to-Rat cast.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-4e3dde3becbd","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":1000,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:96"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_pos {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : (0 : Rat) < (pullCount (ETC.actionWithCommit spec commitArm) a (spec.explorationPulls * K) : Rat)","missing":[],"search":"ratcast_pullcount_actionwithcommit_explorationpulls_mul_k_pos banditrlproof.etc.ratcast_pullcount_actionwithcommit_explorationpulls_mul_k_pos rat-cast form of the fixed-commit etc exploration pull-count positivity leaf. this is the first denominator adapter for future rat-valued empirical means. it only transports the compiled nat positivity theorem across the nat-to-rat cast. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_ne_zero","label":"ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_ne_zero","description":"Nonzero Rat-denominator form of the fixed-commit ETC exploration pull-count positivity leaf. This is still only a deterministic denominator adapter. It does not define empirical means or introduce probability assumptions.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-855a9b0bee1a","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":1001,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:115"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_ne_zero {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) (hexplorationPulls_pos : 0 < spec.explorationPulls) : Not ((pullCount (ETC.actionWithCommit spec commitArm) a (spec.explorationPulls * K) : Rat) = 0)","missing":[],"search":"ratcast_pullcount_actionwithcommit_explorationpulls_mul_k_ne_zero banditrlproof.etc.ratcast_pullcount_actionwithcommit_explorationpulls_mul_k_ne_zero nonzero rat-denominator form of the fixed-commit etc exploration pull-count positivity leaf. this is still only a deterministic denominator adapter. it does not define empirical means or introduce probability assumptions. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_succ_eq_add_if_commitArm_of_ge","label":"pullCount_actionWithCommit_succ_eq_add_if_commitArm_of_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_succ_eq_add_if_commitArm_of_ge","description":"After the configured exploration horizon, one step of the fixed-commit ETC trace updates pull counts according to whether the queried arm is the commit arm. This is the `ETC-ACTION-WITH-COMMIT-POST-COMMIT-SUCC-COUNT` project-local trace/count update leaf.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-856a13659f6f","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":1002,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:135"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_succ_eq_add_if_commitArm_of_ge {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) {t : Nat} (ht : spec.explorationPulls * K <= t) : pullCount (ETC.actionWithCommit spec commitArm) a (Nat.succ t) = pullCount (ETC.actionWithCommit spec commitArm) a t + if commitArm = a then 1 else 0","missing":[],"search":"pullcount_actionwithcommit_succ_eq_add_if_commitarm_of_ge banditrlproof.etc.pullcount_actionwithcommit_succ_eq_add_if_commitarm_of_ge after the configured exploration horizon, one step of the fixed-commit etc trace updates pull counts according to whether the queried arm is the commit arm. this is the `etc-action-with-commit-post-commit-succ-count` project-local trace/count update leaf. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq","label":"pullCount_actionWithCommit_explorationPulls_mul_K_add_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq","description":"After the configured exploration horizon, the fixed-commit ETC trace has a closed-form pull count: the commit arm receives every suffix pull, while all other arms keep their exploration-horizon count. This is the `ETC-ACTION-WITH-COMMIT-SUFFIX-COUNT` project-local trace/count closed-form leaf.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-3b7c4357654a","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":1003,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:155"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq {K : Nat} (spec : ETC.Spec K) (commitArm a : Fin K) (r : Nat) : pullCount (ETC.actionWithCommit spec commitArm) a (spec.explorationPulls * K + r) = spec.explorationPulls + (if commitArm = a then r else 0)","missing":[],"search":"pullcount_actionwithcommit_explorationpulls_mul_k_add_eq banditrlproof.etc.pullcount_actionwithcommit_explorationpulls_mul_k_add_eq after the configured exploration horizon, the fixed-commit etc trace has a closed-form pull count: the commit arm receives every suffix pull, while all other arms keep their exploration-horizon count. this is the `etc-action-with-commit-suffix-count` project-local trace/count closed-form leaf. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_of_ne","label":"pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_of_ne","description":"After the configured exploration horizon, every non-commit arm keeps its exploration-horizon pull count. This is the `ETC-ACTION-WITH-COMMIT-NONCOMMIT-SUFFIX-COUNT` project-local trace/count corollary.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-0d7f1fc4653b","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":1004,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:193"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_of_ne {K : Nat} (spec : ETC.Spec K) {commitArm a : Fin K} (hne : commitArm ≠ a) (r : Nat) : pullCount (ETC.actionWithCommit spec commitArm) a (spec.explorationPulls * K + r) = spec.explorationPulls","missing":[],"search":"pullcount_actionwithcommit_explorationpulls_mul_k_add_eq_of_ne banditrlproof.etc.pullcount_actionwithcommit_explorationpulls_mul_k_add_eq_of_ne after the configured exploration horizon, every non-commit arm keeps its exploration-horizon pull count. this is the `etc-action-with-commit-noncommit-suffix-count` project-local trace/count corollary. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_commitArm","label":"pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_commitArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_commitArm","description":"After the configured exploration horizon, the commit arm has the exploration count plus every suffix pull. This is the `ETC-ACTION-WITH-COMMIT-COMMITARM-SUFFIX-COUNT` project-local trace/count corollary.","url":"../modules/banditrlproof-algorithms-etctracecountlemmas/index.html#decl-17d47d86efb5","parent":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","order":1005,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCTraceCountLemmas"],["Source","BanditRLProof/Algorithms/ETCTraceCountLemmas.lean:210"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_commitArm {K : Nat} (spec : ETC.Spec K) (commitArm : Fin K) (r : Nat) : pullCount (ETC.actionWithCommit spec commitArm) commitArm (spec.explorationPulls * K + r) = spec.explorationPulls + r","missing":[],"search":"pullcount_actionwithcommit_explorationpulls_mul_k_add_eq_commitarm banditrlproof.etc.pullcount_actionwithcommit_explorationpulls_mul_k_add_eq_commitarm after the configured exploration horizon, the commit arm has the exploration count plus every suffix pull. this is the `etc-action-with-commit-commitarm-suffix-count` project-local trace/count corollary. theorem compiled","shard":"modules/fdc005960b7c90ed.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail","label":"prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail","description":"Concrete argmax-oracle wrong-commit probability bound using the canonical centered reward-difference sub-Gaussian tail budget. This is the `ETC-WRONG-COMMIT-CANONICAL-SUBGAUSSIAN-BOUND` leaf. It composes the compiled canonical centered-diff producer with the compiled filtered-sum wrong-commit probability consumer. It does not prove the reward-law independence or `HasSubgaussianMGF` witnesses.","url":"../modules/banditrlproof-algorithms-etcwrongcommitcanonicaltail/index.html#decl-033a8593bec5","parent":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","order":1006,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail"],["Source","BanditRLProof/Algorithms/ETCWrongCommitCanonicalTail.lean:24"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (spec : ETC.Spec K) (model : FiniteBanditModel K) (commitArm : Fin K) (reward : Omega -> RewardTrace Rat) (c : Fin K -> Nat -> NNReal) (hexplorationPulls_pos : 0 < spec.explorationPulls) (h_indep : forall a : Fin K, (a = model.bestArm -> False) -> ProbabilityTheory.iIndepFun (fun t omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) mu) (h_subG : forall a : Fin K, (a = model.bestArm -> False) -> forall t, t ∈ Finset.range (spec.explorationPulls * K) -> ProbabilityTheory.HasSubgaussianMGF (fun omega => ETC.centeredPairwiseRewardDiff spec model commitArm reward a t omega) (c a t) mu) : mu {omega : Omega | (ETC.argmaxCommitOracle hK).choose (fun…","missing":[],"search":"prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail banditrlproof.etc.prob_argmaxcommitoracle_ne_bestarm_le_filtered_sum_centereddiffsubgaussiantail concrete argmax-oracle wrong-commit probability bound using the canonical centered reward-difference sub-gaussian tail budget. this is the `etc-wrong-commit-canonical-subgaussian-bound` leaf. it composes the compiled canonical centered-diff producer with the compiled filtered-sum wrong-commit probability consumer. it does not prove the reward-law independence or `hassubgaussianmgf` witnesses. theorem compiled","shard":"modules/693f5ab3d775f0ba.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_choice_le_sum_gap_mul_explorationPulls_add_suffix_badGap","label":"pseudoRegret_actionWithCommit_choice_le_sum_gap_mul_explorationPulls_add_suffix_badGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ETC.pseudoRegret_actionWithCommit_choice_le_sum_gap_mul_explorationPulls_add_suffix_badGap","description":"For an `Omega`-indexed commit selector, fixed-commit ETC regret after a suffix is bounded by the exploration budget plus a suffix penalty that vanishes when the selected commit arm is the model's `bestArm`. The explicit `badGapBound` is the local bridge to a later probability layer: the suffix cost is charged only on the wrong-commit branch.","url":"../modules/banditrlproof-algorithms-etcwrongcommitregretassembly/index.html#decl-db5db2074df7","parent":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","order":1007,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly"],["Source","BanditRLProof/Algorithms/ETCWrongCommitRegretAssembly.lean:26"],["Chapter","ETC"],["Used in books","bandit"],["Reading references","teaching:etc"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_actionWithCommit_choice_le_sum_gap_mul_explorationPulls_add_suffix_badGap {Omega : Type u} {K : Nat} (spec : ETC.Spec K) (model : FiniteBanditModel K) (commit : Omega -> Fin K) (r : Nat) (badGapBound : Rat) (hbadGap : forall a : Fin K, (a = model.bestArm -> False) -> model.gap a <= badGapBound) (omega : Omega) : pseudoRegret model (ETC.actionWithCommit spec (commit omega)) (spec.explorationPulls * K + r) <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((spec.explorationPulls : Nat) : Rat)) + (((r : Nat) : Rat) * (if commit omega = model.bestArm then 0 else badGapBound))","missing":[],"search":"pseudoregret_actionwithcommit_choice_le_sum_gap_mul_explorationpulls_add_suffix_badgap banditrlproof.etc.pseudoregret_actionwithcommit_choice_le_sum_gap_mul_explorationpulls_add_suffix_badgap for an `omega`-indexed commit selector, fixed-commit etc regret after a suffix is bounded by the exploration budget plus a suffix penalty that vanishes when the selected commit arm is the model's `bestarm`. the explicit `badgapbound` is the local bridge to a later probability layer: the suffix cost is charged only on the wrong-commit branch. theorem compiled","shard":"modules/575319f5853bded4.json","books":["bandit"],"chapters":["teaching:etc"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory_reward_bounded","label":"trajectory_reward_bounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory_reward_bounded","description":"theorem trajectory_reward_bounded (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0:ℝ) 1) (n : ℕ) : ∀ᵐ Y ∂trajectory ν ρ law, Y n ∈ Set.Icc (0:ℝ) 1","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html#decl-951ffe233fe3","parent":"module:BanditRLProof.Algorithms.HOOActualRegret","order":1008,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOActualRegret"],["Source","BanditRLProof/Algorithms/HOOActualRegret.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trajectory_reward_bounded (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0:ℝ) 1) (n : ℕ) : ∀ᵐ Y ∂trajectory ν ρ law, Y n ∈ Set.Icc (0:ℝ) 1","missing":[],"search":"trajectory_reward_bounded banditrlproof.hoo.trajectory_reward_bounded theorem trajectory_reward_bounded (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ set.icc (0:ℝ) 1) (n : ℕ) : ∀ᵐ y ∂trajectory ν ρ law, y n ∈ set.icc (0:ℝ) 1 theorem compiled","shard":"modules/ff36e5a52ef33eea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.integrable_trajectory_reward","label":"integrable_trajectory_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.integrable_trajectory_reward","description":"theorem integrable_trajectory_reward (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0:ℝ) 1) (n : ℕ) : Integrable (fun Y => Y n) (trajectory ν ρ law)","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html#decl-27aa0c0d05ff","parent":"module:BanditRLProof.Algorithms.HOOActualRegret","order":1009,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOActualRegret"],["Source","BanditRLProof/Algorithms/HOOActualRegret.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_trajectory_reward (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0:ℝ) 1) (n : ℕ) : Integrable (fun Y => Y n) (trajectory ν ρ law)","missing":[],"search":"integrable_trajectory_reward banditrlproof.hoo.integrable_trajectory_reward theorem integrable_trajectory_reward (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ set.icc (0:ℝ) 1) (n : ℕ) : integrable (fun y => y n) (trajectory ν ρ law) theorem compiled","shard":"modules/ff36e5a52ef33eea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.integral_trajectory_reward","label":"integral_trajectory_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.integral_trajectory_reward","description":"The reward mean at each actual round equals the mean of its causally selected node. Both the initial round and the successor joint law are used.","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html#decl-96b4ffacb9d5","parent":"module:BanditRLProof.Algorithms.HOOActualRegret","order":1010,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOActualRegret"],["Source","BanditRLProof/Algorithms/HOOActualRegret.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_trajectory_reward (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0:ℝ) 1) (n : ℕ) : (∫ Y, Y n ∂trajectory ν ρ law) = ∫ Y, nodeMean law (action ν ρ Y n) ∂trajectory ν ρ law","missing":[],"search":"integral_trajectory_reward banditrlproof.hoo.integral_trajectory_reward the reward mean at each actual round equals the mean of its causally selected node. both the initial round and the successor joint law are used. theorem compiled","shard":"modules/ff36e5a52ef33eea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.integral_actual_reward","label":"integral_actual_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.integral_actual_reward","description":"Source model version of the one-round expectation identity.","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html#decl-eed2fbd01241","parent":"module:BanditRLProof.Algorithms.HOOActualRegret","order":1011,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOActualRegret"],["Source","BanditRLProof/Algorithms/HOOActualRegret.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.integral_actual_reward {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) [IsMarkovKernel law] (ν ρ : ℝ) (f : X → ℝ) (hmean : ∀x, (∫ y, y ∂law x)=f x) (hbound : ∀x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (n : ℕ) : (∫ Y, Y n ∂trajectory ν ρ (C.nodeLaw law)) = ∫ Y, f (C.arm ν ρ Y n) ∂trajectory ν ρ (C.nodeLaw law)","missing":[],"search":"integral_actual_reward banditrlproof.hoo.covering.integral_actual_reward source model version of the one-round expectation identity. theorem compiled","shard":"modules/ff36e5a52ef33eea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret","label":"expected_actual_eq_pseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret","description":"Expected cumulative realized regret and pseudo-regret coincide at every horizon for one fixed policy and compatible infinite reward trajectory.","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html#decl-223951ac8429","parent":"module:BanditRLProof.Algorithms.HOOActualRegret","order":1012,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOActualRegret"],["Source","BanditRLProof/Algorithms/HOOActualRegret.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.expected_actual_eq_pseudoRegret {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀x, (∫ y, y ∂law x)=f x) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (hbound : ∀x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (N : ℕ) : (∫ Y, (∑ n ∈ Finset.range N, (best-Y n)) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) = (∫ Y, (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law))","missing":[],"search":"expected_actual_eq_pseudoregret banditrlproof.hoo.regularcovering.expected_actual_eq_pseudoregret expected cumulative realized regret and pseudo-regret coincide at every horizon for one fixed policy and compatible infinite reward trajectory. theorem compiled","shard":"modules/ff36e5a52ef33eea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate","label":"expected_actualRegret_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate","description":"The same repaired source rate for expected realized cumulative regret.","url":"../modules/banditrlproof-algorithms-hooactualregret/index.html#decl-e067ec4b9408","parent":"module:BanditRLProof.Algorithms.HOOActualRegret","order":1013,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOActualRegret"],["Source","BanditRLProof/Algorithms/HOOActualRegret.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.expected_actualRegret_rate {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best d : ℝ) (hmean : ∀x, (∫ y, y ∂law x)=f x) (hf : ∀x, f x≤best) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd : C.nearOptimalityDimension f best (4*C.nu1/C.nu2) < (d:EReal)) : ∃ γ : ℝ, 0<γ ∧ ∀ N : ℕ, 1≤N → (∫ Y, (∑ n ∈ Finset.range N, (best-Y n)) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ γ*(N:ℝ)^((d+1)/(d+2))*(Real.log (max (N:ℝ) 2))^(1/(d+2))","missing":[],"search":"expected_actualregret_rate banditrlproof.hoo.regularcovering.expected_actualregret_rate the same repaired source rate for expected realized cumulative regret. theorem compiled","shard":"modules/ff36e5a52ef33eea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionNoise","label":"regionNoise","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionNoise","description":"noncomputable def regionNoise (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (i : ℕ) (Y : ℕ → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-7c2bdd08354e","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1014,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionNoise (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (i : ℕ) (Y : ℕ → ℝ) : ℝ","missing":[],"search":"regionnoise banditrlproof.hoo.regionnoise noncomputable def regionnoise (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (i : ℕ) (y : ℕ → ℝ) : ℝ definition compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionCount","label":"regionCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionCount","description":"noncomputable def regionCount (ν ρ : ℝ) (v : Node) (i : ℕ) (Y : ℕ → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-b14c35d6723f","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1015,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionCount (ν ρ : ℝ) (v : Node) (i : ℕ) (Y : ℕ → ℝ) : ℝ","missing":[],"search":"regioncount banditrlproof.hoo.regioncount noncomputable def regioncount (ν ρ : ℝ) (v : node) (i : ℕ) (y : ℕ → ℝ) : ℝ definition compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_prefixExtension_at","label":"action_prefixExtension_at","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_prefixExtension_at","description":"theorem action_prefixExtension_at (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : action ν ρ (prefixExtension n (Preorder.frestrictLe n Y)) n = action ν ρ Y n","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-086612460b07","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1016,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_prefixExtension_at (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : action ν ρ (prefixExtension n (Preorder.frestrictLe n Y)) n = action ν ρ Y n","missing":[],"search":"action_prefixextension_at banditrlproof.hoo.action_prefixextension_at theorem action_prefixextension_at (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : action ν ρ (prefixextension n (preorder.frestrictle n y)) n = action ν ρ y n theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_action_piLE","label":"measurable_action_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_action_piLE","description":"theorem measurable_action_piLE (ν ρ : ℝ) (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → ℝ => action ν ρ Y n)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-40ac357dd34b","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1017,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_action_piLE (ν ρ : ℝ) (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → ℝ => action ν ρ Y n)","missing":[],"search":"measurable_action_pile banditrlproof.hoo.measurable_action_pile theorem measurable_action_pile (ν ρ : ℝ) (n : ℕ) : measurable[filtration.pile n] (fun y : ℕ → ℝ => action ν ρ y n) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_coordinate_piLE","label":"measurable_coordinate_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_coordinate_piLE","description":"theorem measurable_coordinate_piLE (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → ℝ => Y n)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-62c242bda013","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1018,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_coordinate_piLE (n : ℕ) : Measurable[Filtration.piLE n] (fun Y : ℕ → ℝ => Y n)","missing":[],"search":"measurable_coordinate_pile banditrlproof.hoo.measurable_coordinate_pile theorem measurable_coordinate_pile (n : ℕ) : measurable[filtration.pile n] (fun y : ℕ → ℝ => y n) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_compensated_adapted","label":"region_compensated_adapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_compensated_adapted","description":"theorem region_compensated_adapted (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (tilt : ℝ) : StronglyAdapted Filtration.piLE (fun i Y => tilt * regionNoise ν ρ law v i Y - tilt^2/8 * regionCount ν ρ v i Y)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-9eb21def941f","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1019,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_compensated_adapted (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (tilt : ℝ) : StronglyAdapted Filtration.piLE (fun i Y => tilt * regionNoise ν ρ law v i Y - tilt^2/8 * regionCount ν ρ v i Y)","missing":[],"search":"region_compensated_adapted banditrlproof.hoo.region_compensated_adapted theorem region_compensated_adapted (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (tilt : ℝ) : stronglyadapted filtration.pile (fun i y => tilt * regionnoise ν ρ law v i y - tilt^2/8 * regioncount ν ρ v i y) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_compensated_successor","label":"region_compensated_successor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_compensated_successor","description":"theorem region_compensated_successor (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt : ℝ) : Concentration.HasCondMGFUpperBoundAt (Filtration.piLE n) (Filtration.piLE.le n) (fun Y => tilt * regionNoise ν ρ law v (n+1) Y - tilt^2/8 * regionCount ν ρ v (n+1) Y) 1 0 (trajectory ν ρ law)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-a2a7c3a2cbf4","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1020,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_compensated_successor (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt : ℝ) : Concentration.HasCondMGFUpperBoundAt (Filtration.piLE n) (Filtration.piLE.le n) (fun Y => tilt * regionNoise ν ρ law v (n+1) Y - tilt^2/8 * regionCount ν ρ v (n+1) Y) 1 0 (trajectory ν ρ law)","missing":[],"search":"region_compensated_successor banditrlproof.hoo.region_compensated_successor theorem region_compensated_successor (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (n : ℕ) (tilt : ℝ) : concentration.hascondmgfupperboundat (filtration.pile n) (filtration.pile.le n) (fun y => tilt * regionnoise ν ρ law v (n+1) y - tilt^2/8 * regioncount ν ρ v (n+1) y) 1 0 (trajectory ν ρ law) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_compensated_initial","label":"region_compensated_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_compensated_initial","description":"theorem region_compensated_initial (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (fun Y => tilt * regionNoise ν ρ law v 0 Y - tilt^2/8 * regionCount ν ρ v 0 Y) 1 0 (trajectory ν ρ law)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-150bf83b73f9","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1021,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_compensated_initial (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (fun Y => tilt * regionNoise ν ρ law v 0 Y - tilt^2/8 * regionCount ν ρ v 0 Y) 1 0 (trajectory ν ρ law)","missing":[],"search":"region_compensated_initial banditrlproof.hoo.region_compensated_initial theorem region_compensated_initial (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (tilt : ℝ) : concentration.hasmgfupperboundat (fun y => tilt * regionnoise ν ρ law v 0 y - tilt^2/8 * regioncount ν ρ v 0 y) 1 0 (trajectory ν ρ law) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_noise_count_tail","label":"region_noise_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_noise_count_tail","description":"A regional noise/count tail for the real HOO process. Both the initial and conditional MGF obligations are derived from its bounded reward kernels.","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-398b173b9234","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1022,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_noise_count_tail (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt threshold countBudget : ℝ) (htilt : 0 ≤ tilt) : (trajectory ν ρ law) {Y | threshold ≤ ∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y ∧ (∑ i ∈ Finset.range n, regionCount ν ρ v i Y) ≤ countBudget} ≤ ENNReal.ofReal (Real.exp (-tilt * threshold + tilt^2/8 * countBudget))","missing":[],"search":"region_noise_count_tail banditrlproof.hoo.region_noise_count_tail a regional noise/count tail for the real hoo process. both the initial and conditional mgf obligations are derived from its bounded reward kernels. theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.sum_regionCount_eq_visits","label":"sum_regionCount_eq_visits","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.sum_regionCount_eq_visits","description":"theorem sum_regionCount_eq_visits (ν ρ : ℝ) (v : Node) (n : ℕ) (Y : ℕ → ℝ) : (∑ i ∈ Finset.range n, regionCount ν ρ v i Y) = (visits (history ν ρ Y n) v : ℝ)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-cf26f74ceb2f","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1023,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:100"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_regionCount_eq_visits (ν ρ : ℝ) (v : Node) (n : ℕ) (Y : ℕ → ℝ) : (∑ i ∈ Finset.range n, regionCount ν ρ v i Y) = (visits (history ν ρ Y n) v : ℝ)","missing":[],"search":"sum_regioncount_eq_visits banditrlproof.hoo.sum_regioncount_eq_visits theorem sum_regioncount_eq_visits (ν ρ : ℝ) (v : node) (n : ℕ) (y : ℕ → ℝ) : (∑ i ∈ finset.range n, regioncount ν ρ v i y) = (visits (history ν ρ y n) v : ℝ) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.sum_regionNoise_eq_rewardSum_sub_means","label":"sum_regionNoise_eq_rewardSum_sub_means","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.sum_regionNoise_eq_rewardSum_sub_means","description":"theorem sum_regionNoise_eq_rewardSum_sub_means (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (Y : ℕ → ℝ) : (∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y) = rewardSum (history ν ρ Y n) v - ∑ i ∈ Finset.range n, regionCount ν ρ v i Y * nodeMean law (action ν ρ Y i)","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-226f0837ac57","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1024,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_regionNoise_eq_rewardSum_sub_means (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (Y : ℕ → ℝ) : (∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y) = rewardSum (history ν ρ Y n) v - ∑ i ∈ Finset.range n, regionCount ν ρ v i Y * nodeMean law (action ν ρ Y i)","missing":[],"search":"sum_regionnoise_eq_rewardsum_sub_means banditrlproof.hoo.sum_regionnoise_eq_rewardsum_sub_means theorem sum_regionnoise_eq_rewardsum_sub_means (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (n : ℕ) (y : ℕ → ℝ) : (∑ i ∈ finset.range n, regionnoise ν ρ law v i y) = rewardsum (history ν ρ y n) v - ∑ i ∈ finset.range n, regioncount ν ρ v i y * nodemean law (action ν ρ y i) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_negative_noise_count_tail","label":"region_negative_noise_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_negative_noise_count_tail","description":"theorem region_negative_noise_count_tail (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt threshold countBudget : ℝ) (htilt : 0 ≤ tilt) : (trajectory ν ρ law) {Y | threshold ≤ -(∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y) ∧ (visits (history ν ρ Y n) v : ℝ) ≤ countBudget} ≤ ENNReal.ofReal (Real.exp (-tilt * threshold + tilt^2/8 * cou…","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-126edc23747b","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1025,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_negative_noise_count_tail (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt threshold countBudget : ℝ) (htilt : 0 ≤ tilt) : (trajectory ν ρ law) {Y | threshold ≤ -(∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y) ∧ (visits (history ν ρ Y n) v : ℝ) ≤ countBudget} ≤ ENNReal.ofReal (Real.exp (-tilt * threshold + tilt^2/8 * countBudget))","missing":[],"search":"region_negative_noise_count_tail banditrlproof.hoo.region_negative_noise_count_tail theorem region_negative_noise_count_tail (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (n : ℕ) (tilt threshold countbudget : ℝ) (htilt : 0 ≤ tilt) : (trajectory ν ρ law) {y | threshold ≤ -(∑ i ∈ finset.range n, regionnoise ν ρ law v i y) ∧ (visits (history ν ρ y n) v : ℝ) ≤ countbudget} ≤ ennreal.ofreal (real.exp (-tilt * threshold + tilt^2/8 * countbudget)) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_noise_visits_tail","label":"region_noise_visits_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_noise_visits_tail","description":"Optimized upper tail, retaining the actual regional visit count in the event.","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-7e30b7264be0","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1026,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:149"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_noise_visits_tail (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (threshold countBudget : ℝ) (hthreshold : 0 ≤ threshold) (hcount : 0 < countBudget) : (trajectory ν ρ law) {Y | threshold ≤ ∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y ∧ (visits (history ν ρ Y n) v : ℝ) ≤ countBudget} ≤ ENNReal.ofReal (Real.exp (-2 * threshold^2 / countBudget))","missing":[],"search":"region_noise_visits_tail banditrlproof.hoo.region_noise_visits_tail optimized upper tail, retaining the actual regional visit count in the event. theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_negative_noise_visits_tail","label":"region_negative_noise_visits_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_negative_noise_visits_tail","description":"theorem region_negative_noise_visits_tail (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (threshold countBudget : ℝ) (hthreshold : 0 ≤ threshold) (hcount : 0 < countBudget) : (trajectory ν ρ law) {Y | threshold ≤ -(∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y) ∧ (visits (history ν ρ Y n) v : ℝ) ≤ countBudget} ≤ ENNReal.ofReal (Real.exp (-…","url":"../modules/banditrlproof-algorithms-hooconcentration/index.html#decl-bf31357a6d87","parent":"module:BanditRLProof.Algorithms.HOOConcentration","order":1027,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConcentration"],["Source","BanditRLProof/Algorithms/HOOConcentration.lean:164"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_negative_noise_visits_tail (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (threshold countBudget : ℝ) (hthreshold : 0 ≤ threshold) (hcount : 0 < countBudget) : (trajectory ν ρ law) {Y | threshold ≤ -(∑ i ∈ Finset.range n, regionNoise ν ρ law v i Y) ∧ (visits (history ν ρ Y n) v : ℝ) ≤ countBudget} ≤ ENNReal.ofReal (Real.exp (-2 * threshold^2 / countBudget))","missing":[],"search":"region_negative_noise_visits_tail banditrlproof.hoo.region_negative_noise_visits_tail theorem region_negative_noise_visits_tail (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (n : ℕ) (threshold countbudget : ℝ) (hthreshold : 0 ≤ threshold) (hcount : 0 < countbudget) : (trajectory ν ρ law) {y | threshold ≤ -(∑ i ∈ finset.range n, regionnoise ν ρ law v i y) ∧ (visits (history ν ρ y n) v : ℝ) ≤ countbudget} ≤ ennreal.ofreal (real.exp (-2 * threshold^2 / countbudget)) theorem compiled","shard":"modules/2f1e7407e86f21fa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.nodeMean","label":"nodeMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.nodeMean","description":"noncomputable def nodeMean (law : Kernel Node ℝ) (v : Node) : ℝ","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-26e1f75fabc8","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1028,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def nodeMean (law : Kernel Node ℝ) (v : Node) : ℝ","missing":[],"search":"nodemean banditrlproof.hoo.nodemean noncomputable def nodemean (law : kernel node ℝ) (v : node) : ℝ definition compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bounded_node_subgaussian","label":"bounded_node_subgaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bounded_node_subgaussian","description":"theorem bounded_node_subgaussian (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) : HasSubgaussianMGF (fun y => y - nodeMean law v) (1/4 : ℝ≥0) (law v)","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-09e2c3441352","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1029,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bounded_node_subgaussian (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) : HasSubgaussianMGF (fun y => y - nodeMean law v) (1/4 : ℝ≥0) (law v)","missing":[],"search":"bounded_node_subgaussian banditrlproof.hoo.bounded_node_subgaussian theorem bounded_node_subgaussian (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ v, ∀ᵐ y ∂law v, y ∈ set.icc (0 : ℝ) 1) (v : node) : hassubgaussianmgf (fun y => y - nodemean law v) (1/4 : ℝ≥0) (law v) theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bounded_node_fixedMGF","label":"bounded_node_fixedMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bounded_node_fixedMGF","description":"The new setting supplies the existing fixed-MGF interface from bounded reward laws, rather than assuming a conditional confidence theorem.","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-26b184c573a0","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1030,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bounded_node_fixedMGF (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ v, ∀ᵐ y ∂law v, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (fun y => y - nodeMean law v) tilt (tilt^2/8) (law v)","missing":[],"search":"bounded_node_fixedmgf banditrlproof.hoo.bounded_node_fixedmgf the new setting supplies the existing fixed-mgf interface from bounded reward laws, rather than assuming a conditional confidence theorem. theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.selectedNode","label":"selectedNode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.selectedNode","description":"noncomputable def selectedNode (ν ρ : ℝ) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) : Node","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-6c34e84307fa","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1031,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedNode (ν ρ : ℝ) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) : Node","missing":[],"search":"selectednode banditrlproof.hoo.selectednode noncomputable def selectednode (ν ρ : ℝ) (n : ℕ) (h : (i : finset.iic n) → ℝ) : node definition compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionIncrement","label":"regionIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionIncrement","description":"noncomputable def regionIncrement (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) (y : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-8a6d71db00f3","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1032,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionIncrement (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) (y : ℝ) : ℝ","missing":[],"search":"regionincrement banditrlproof.hoo.regionincrement noncomputable def regionincrement (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (n : ℕ) (h : (i : finset.iic n) → ℝ) (y : ℝ) : ℝ definition compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionSelected","label":"regionSelected","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionSelected","description":"noncomputable def regionSelected (ν ρ : ℝ) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-5b5b4b2ec324","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1033,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionSelected (ν ρ : ℝ) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) : ℝ","missing":[],"search":"regionselected banditrlproof.hoo.regionselected noncomputable def regionselected (ν ρ : ℝ) (v : node) (n : ℕ) (h : (i : finset.iic n) → ℝ) : ℝ definition compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_step_fixedMGF","label":"region_step_fixedMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_step_fixedMGF","description":"theorem region_step_fixedMGF (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (regionIncrement ν ρ law v n h) tilt (tilt^2/8 * regionSelected ν ρ v n h) (stepKernel ν ρ law n h)","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-bc6fb4351a5b","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1034,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_step_fixedMGF (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) (tilt : ℝ) : Concentration.HasMGFUpperBoundAt (regionIncrement ν ρ law v n h) tilt (tilt^2/8 * regionSelected ν ρ v n h) (stepKernel ν ρ law n h)","missing":[],"search":"region_step_fixedmgf banditrlproof.hoo.region_step_fixedmgf theorem region_step_fixedmgf (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (n : ℕ) (h : (i : finset.iic n) → ℝ) (tilt : ℝ) : concentration.hasmgfupperboundat (regionincrement ν ρ law v n h) tilt (tilt^2/8 * regionselected ν ρ v n h) (stepkernel ν ρ law n h) theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_step_compensated_integral_le_one","label":"region_step_compensated_integral_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_step_compensated_integral_le_one","description":"Predictable variance is paid only when this region is selected. This is the one-step exponential process needed for count-dependent concentration.","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-8fd21bf009fb","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1035,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_step_compensated_integral_le_one (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (h : (i : Finset.Iic n) → ℝ) (tilt : ℝ) : (∫ y, Real.exp (tilt * regionIncrement ν ρ law v n h y - tilt^2/8 * regionSelected ν ρ v n h) ∂stepKernel ν ρ law n h) ≤ 1","missing":[],"search":"region_step_compensated_integral_le_one banditrlproof.hoo.region_step_compensated_integral_le_one predictable variance is paid only when this region is selected. this is the one-step exponential process needed for count-dependent concentration. theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_selectedNode","label":"measurable_selectedNode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_selectedNode","description":"theorem measurable_selectedNode (ν ρ : ℝ) (n : ℕ) : Measurable (selectedNode ν ρ n)","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-dc2e513e8374","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1036,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedNode (ν ρ : ℝ) (n : ℕ) : Measurable (selectedNode ν ρ n)","missing":[],"search":"measurable_selectednode banditrlproof.hoo.measurable_selectednode theorem measurable_selectednode (ν ρ : ℝ) (n : ℕ) : measurable (selectednode ν ρ n) theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_regionSelected","label":"measurable_regionSelected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_regionSelected","description":"theorem measurable_regionSelected (ν ρ : ℝ) (v : Node) (n : ℕ) : Measurable (regionSelected ν ρ v n)","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-f12233a4ab2b","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1037,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_regionSelected (ν ρ : ℝ) (v : Node) (n : ℕ) : Measurable (regionSelected ν ρ v n)","missing":[],"search":"measurable_regionselected banditrlproof.hoo.measurable_regionselected theorem measurable_regionselected (ν ρ : ℝ) (v : node) (n : ℕ) : measurable (regionselected ν ρ v n) theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_regionIncrement","label":"measurable_regionIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_regionIncrement","description":"theorem measurable_regionIncrement (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) : Measurable (fun p : ((i : Finset.Iic n) → ℝ) × ℝ => regionIncrement ν ρ law v n p.1 p.2)","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-21f709f5361d","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1038,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_regionIncrement (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) : Measurable (fun p : ((i : Finset.Iic n) → ℝ) × ℝ => regionIncrement ν ρ law v n p.1 p.2)","missing":[],"search":"measurable_regionincrement banditrlproof.hoo.measurable_regionincrement theorem measurable_regionincrement (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (n : ℕ) : measurable (fun p : ((i : finset.iic n) → ℝ) × ℝ => regionincrement ν ρ law v n p.1 p.2) theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionCompensated","label":"regionCompensated","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionCompensated","description":"noncomputable def regionCompensated (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (tilt : ℝ) (p : ((i : Finset.Iic n) → ℝ) × ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-917c319c1e51","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1039,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionCompensated (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (tilt : ℝ) (p : ((i : Finset.Iic n) → ℝ) × ℝ) : ℝ","missing":[],"search":"regioncompensated banditrlproof.hoo.regioncompensated noncomputable def regioncompensated (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (n : ℕ) (tilt : ℝ) (p : ((i : finset.iic n) → ℝ) × ℝ) : ℝ definition compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_regionCompensated","label":"measurable_regionCompensated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_regionCompensated","description":"theorem measurable_regionCompensated (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (tilt : ℝ) : Measurable (regionCompensated ν ρ law v n tilt)","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-d6fdc8a5b5cf","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1040,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_regionCompensated (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (tilt : ℝ) : Measurable (regionCompensated ν ρ law v n tilt)","missing":[],"search":"measurable_regioncompensated banditrlproof.hoo.measurable_regioncompensated theorem measurable_regioncompensated (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (n : ℕ) (tilt : ℝ) : measurable (regioncompensated ν ρ law v n tilt) theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.integrable_regionCompensated_exp","label":"integrable_regionCompensated_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.integrable_regionCompensated_exp","description":"All exponential moments are integrable under the joint prefix/reward law. This supplies the global integrability required by conditional MGF iteration.","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-7796435c8bd6","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1041,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_regionCompensated_exp (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt s : ℝ) (μ : Measure ((i : Finset.Iic n) → ℝ)) [IsProbabilityMeasure μ] : Integrable (fun p => Real.exp (s * regionCompensated ν ρ law v n tilt p)) (μ ⊗ₘ stepKernel ν ρ law n)","missing":[],"search":"integrable_regioncompensated_exp banditrlproof.hoo.integrable_regioncompensated_exp all exponential moments are integrable under the joint prefix/reward law. this supplies the global integrability required by conditional mgf iteration. theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory_region_condExp_le_one","label":"trajectory_region_condExp_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory_region_condExp_le_one","description":"The conditional exponential inequality holds for the actual constructed trajectory and its observed prefix, without a supplied concentration premise.","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-ba39d7f25bf9","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1042,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:154"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trajectory_region_condExp_le_one (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt : ℝ) : (trajectory ν ρ law)[fun Y => Real.exp (regionCompensated ν ρ law v n tilt (Preorder.frestrictLe n Y, Y (n+1))) | MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance] ≤ᵐ[trajectory ν ρ law] fun _ => 1","missing":[],"search":"trajectory_region_condexp_le_one banditrlproof.hoo.trajectory_region_condexp_le_one the conditional exponential inequality holds for the actual constructed trajectory and its observed prefix, without a supplied concentration premise. theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory_region_condMGF","label":"trajectory_region_condMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory_region_condMGF","description":"Actual HOO successor increments now inhabit the shared conditional MGF interface. The only probabilistic premise is bounded support of each reward law.","url":"../modules/banditrlproof-algorithms-hooconditionalmgf/index.html#decl-1b580dbbb479","parent":"module:BanditRLProof.Algorithms.HOOConditionalMGF","order":1043,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConditionalMGF"],["Source","BanditRLProof/Algorithms/HOOConditionalMGF.lean:181"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trajectory_region_condMGF (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (tilt : ℝ) : Concentration.HasCondMGFUpperBoundAt (MeasurableSpace.comap (Preorder.frestrictLe n) inferInstance) (Preorder.measurable_frestrictLe n).comap_le (fun Y => regionCompensated ν ρ law v n tilt (Preorder.frestrictLe n Y, Y (n+1))) 1 0 (trajectory ν ρ law)","missing":[],"search":"trajectory_region_condmgf banditrlproof.hoo.trajectory_region_condmgf actual hoo successor increments now inhabit the shared conditional mgf interface. the only probabilistic premise is bounded support of each reward law. theorem compiled","shard":"modules/fb89660da334fb1b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionDeviation","label":"regionDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionDeviation","description":"noncomputable def regionDeviation (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (lower : Bool) (Y : ℕ → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooconfidence/index.html#decl-4db7dcc2537a","parent":"module:BanditRLProof.Algorithms.HOOConfidence","order":1044,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOConfidence"],["Source","BanditRLProof/Algorithms/HOOConfidence.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionDeviation (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (n : ℕ) (lower : Bool) (Y : ℕ → ℝ) : ℝ","missing":[],"search":"regiondeviation banditrlproof.hoo.regiondeviation noncomputable def regiondeviation (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (n : ℕ) (lower : bool) (y : ℕ → ℝ) : ℝ definition compiled","shard":"modules/03ba75ef32eb5fb3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visits_history_le","label":"visits_history_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visits_history_le","description":"theorem visits_history_le (ν ρ : ℝ) (v : Node) (n : ℕ) (Y : ℕ → ℝ) : visits (history ν ρ Y n) v ≤ n","url":"../modules/banditrlproof-algorithms-hooconfidence/index.html#decl-bf2142da165d","parent":"module:BanditRLProof.Algorithms.HOOConfidence","order":1045,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConfidence"],["Source","BanditRLProof/Algorithms/HOOConfidence.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem visits_history_le (ν ρ : ℝ) (v : Node) (n : ℕ) (Y : ℕ → ℝ) : visits (history ν ρ Y n) v ≤ n","missing":[],"search":"visits_history_le banditrlproof.hoo.visits_history_le theorem visits_history_le (ν ρ : ℝ) (v : node) (n : ℕ) (y : ℕ → ℝ) : visits (history ν ρ y n) v ≤ n theorem compiled","shard":"modules/03ba75ef32eb5fb3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_deviation_slice","label":"region_deviation_slice","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_deviation_slice","description":"theorem region_deviation_slice (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n k : ℕ) (lower : Bool) (L : ℝ) (hL : 0 ≤ L) (hk : 0 < k) : (trajectory ν ρ law) {Y | Real.sqrt (2 * (k : ℝ) * L) ≤ regionDeviation ν ρ law v n lower Y ∧ (visits (history ν ρ Y n) v : ℝ) ≤ k} ≤ ENNReal.ofReal (Real.exp (-4*L))","url":"../modules/banditrlproof-algorithms-hooconfidence/index.html#decl-3cb8a0f904b6","parent":"module:BanditRLProof.Algorithms.HOOConfidence","order":1046,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConfidence"],["Source","BanditRLProof/Algorithms/HOOConfidence.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_deviation_slice (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n k : ℕ) (lower : Bool) (L : ℝ) (hL : 0 ≤ L) (hk : 0 < k) : (trajectory ν ρ law) {Y | Real.sqrt (2 * (k : ℝ) * L) ≤ regionDeviation ν ρ law v n lower Y ∧ (visits (history ν ρ Y n) v : ℝ) ≤ k} ≤ ENNReal.ofReal (Real.exp (-4*L))","missing":[],"search":"region_deviation_slice banditrlproof.hoo.region_deviation_slice theorem region_deviation_slice (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (n k : ℕ) (lower : bool) (l : ℝ) (hl : 0 ≤ l) (hk : 0 < k) : (trajectory ν ρ law) {y | real.sqrt (2 * (k : ℝ) * l) ≤ regiondeviation ν ρ law v n lower y ∧ (visits (history ν ρ y n) v : ℝ) ≤ k} ≤ ennreal.ofreal (real.exp (-4*l)) theorem compiled","shard":"modules/03ba75ef32eb5fb3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_deviation_confidence","label":"region_deviation_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_deviation_confidence","description":"One-sided confidence for the empirical regional noise, with its actual random visit count. The Boolean selects the upper or lower deviation.","url":"../modules/banditrlproof-algorithms-hooconfidence/index.html#decl-c310f92326c3","parent":"module:BanditRLProof.Algorithms.HOOConfidence","order":1047,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOConfidence"],["Source","BanditRLProof/Algorithms/HOOConfidence.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_deviation_confidence (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (n : ℕ) (lower : Bool) (L : ℝ) (hL : 0 ≤ L) : (trajectory ν ρ law) {Y | 0 < visits (history ν ρ Y n) v ∧ Real.sqrt (2 * (visits (history ν ρ Y n) v : ℝ) * L) ≤ regionDeviation ν ρ law v n lower Y} ≤ (n : ENNReal) * ENNReal.ofReal (Real.exp (-4*L))","missing":[],"search":"region_deviation_confidence banditrlproof.hoo.region_deviation_confidence one-sided confidence for the empirical regional noise, with its actual random visit count. the boolean selects the upper or lower deviation. theorem compiled","shard":"modules/03ba75ef32eb5fb3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.balance_identities","label":"balance_identities","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.balance_identities","description":"private theorem balance_identities {N L d : ℝ} (hN : 0<N) (hL : 0<L) (hd : 0<d) : N*(L/N)^(1/(d+2)) = N^((d+1)/(d+2))*L^(1/(d+2)) ∧ L*((L/N)^(1/(d+2)))^(-(1+d)) = N^((d+1)/(d+2))*L^(1/(d+2))","url":"../modules/banditrlproof-algorithms-hoodepthoptimization/index.html#decl-ea9d06582745","parent":"module:BanditRLProof.Algorithms.HOODepthOptimization","order":1048,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOODepthOptimization"],["Source","BanditRLProof/Algorithms/HOODepthOptimization.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem balance_identities {N L d : ℝ} (hN : 0<N) (hL : 0<L) (hd : 0<d) : N*(L/N)^(1/(d+2)) = N^((d+1)/(d+2))*L^(1/(d+2)) ∧ L*((L/N)^(1/(d+2)))^(-(1+d)) = N^((d+1)/(d+2))*L^(1/(d+2))","missing":[],"search":"balance_identities banditrlproof.hoo.balance_identities private theorem balance_identities {n l d : ℝ} (hn : 0<n) (hl : 0<l) (hd : 0<d) : n*(l/n)^(1/(d+2)) = n^((d+1)/(d+2))*l^(1/(d+2)) ∧ l*((l/n)^(1/(d+2)))^(-(1+d)) = n^((d+1)/(d+2))*l^(1/(d+2)) theorem compiled","shard":"modules/e10b13a303147a7f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.exists_regret_depth","label":"exists_regret_depth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.exists_regret_depth","description":"Chooses a genuine integer H>=1; no real-valued cutoff or asymptotic rounding assumption is used. The same estimate works when L=N.","url":"../modules/banditrlproof-algorithms-hoodepthoptimization/index.html#decl-64694633f7d2","parent":"module:BanditRLProof.Algorithms.HOODepthOptimization","order":1049,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOODepthOptimization"],["Source","BanditRLProof/Algorithms/HOODepthOptimization.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_regret_depth {ρ d A B N L : ℝ} (hr : 0<ρ) (hr1 : ρ<1) (hd : 0<d) (hA : 0≤A) (hB : 0≤B) (hN : 0<N) (hL : 0<L) (hLN : L≤N) : ∃ H : ℕ, 1≤H ∧ A*N*ρ^H+B*L*(ρ^H)^(-(1+d)) ≤ (A+B*ρ^(-(1+d)))*N^((d+1)/(d+2))*L^(1/(d+2))","missing":[],"search":"exists_regret_depth banditrlproof.hoo.exists_regret_depth chooses a genuine integer h>=1; no real-valued cutoff or asymptotic rounding assumption is used. the same estimate works when l=n. theorem compiled","shard":"modules/e10b13a303147a7f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.log_horizon_pos_le","label":"log_horizon_pos_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.log_horizon_pos_le","description":"theorem log_horizon_pos_le (N : ℕ) (hN : 1≤N) : 0<Real.log (max (N:ℝ) 2) ∧ Real.log (max (N:ℝ) 2)≤(N:ℝ)","url":"../modules/banditrlproof-algorithms-hoodepthoptimization/index.html#decl-ee6bea9cf542","parent":"module:BanditRLProof.Algorithms.HOODepthOptimization","order":1050,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOODepthOptimization"],["Source","BanditRLProof/Algorithms/HOODepthOptimization.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem log_horizon_pos_le (N : ℕ) (hN : 1≤N) : 0<Real.log (max (N:ℝ) 2) ∧ Real.log (max (N:ℝ) 2)≤(N:ℝ)","missing":[],"search":"log_horizon_pos_le banditrlproof.hoo.log_horizon_pos_le theorem log_horizon_pos_le (n : ℕ) (hn : 1≤n) : 0<real.log (max (n:ℝ) 2) ∧ real.log (max (n:ℝ) 2)≤(n:ℝ) theorem compiled","shard":"modules/e10b13a303147a7f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.integrable_actual_gap","label":"integrable_actual_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.integrable_actual_gap","description":"theorem RegularCovering.integrable_actual_gap {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (μ : Measure (ℕ → ℝ)) [IsProbabilityMeasure μ] (n : ℕ) : Integrable (fun Y => best-f (C.toCovering.arm C.nu1 C.rho Y n)) μ","url":"../modules/banditrlproof-algorithms-hooexpectedregret/index.html#decl-015c632a60af","parent":"module:BanditRLProof.Algorithms.HOOExpectedRegret","order":1051,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedRegret"],["Source","BanditRLProof/Algorithms/HOOExpectedRegret.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.integrable_actual_gap {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (μ : Measure (ℕ → ℝ)) [IsProbabilityMeasure μ] (n : ℕ) : Integrable (fun Y => best-f (C.toCovering.arm C.nu1 C.rho Y n)) μ","missing":[],"search":"integrable_actual_gap banditrlproof.hoo.regularcovering.integrable_actual_gap theorem regularcovering.integrable_actual_gap {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hfrange : ∀x, f x ∈ set.icc (0:ℝ) 1) (μ : measure (ℕ → ℝ)) [isprobabilitymeasure μ] (n : ℕ) : integrable (fun y => best-f (c.tocovering.arm c.nu1 c.rho y n)) μ theorem compiled","shard":"modules/9b924db43bd01fc1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_regret_partition_bound","label":"expected_regret_partition_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_regret_partition_bound","description":"Expected regret with an arbitrary integer cutoff H, retaining the exact three terms. All random count bounds come from the actual HOO law.","url":"../modules/banditrlproof-algorithms-hooexpectedregret/index.html#decl-5552123cdb1e","parent":"module:BanditRLProof.Algorithms.HOOExpectedRegret","order":1052,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedRegret"],["Source","BanditRLProof/Algorithms/HOOExpectedRegret.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.expected_regret_partition_bound {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀x, (∫ y, y ∂law x)=f x) (hf : ∀x, f x≤best) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (H N : ℕ) : (∫ Y, (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ 4*(C.nu1*C.rho^H)*N + (∑ h ∈ Finset.range H, 4*(C.nu1*C.rho^h)*(C.nearOptimalNodes f best h).card) + ∑ h ∈ Finset.range H, 4*(C.nu1*C.rho^h)*(C.boundaryNodes f best h).card * (8*Real.log (max (N:ℝ) 2)/(C.nu1*C.rho^(h+1))^2+4)","missing":[],"search":"expected_regret_partition_bound banditrlproof.hoo.regularcovering.expected_regret_partition_bound expected regret with an arbitrary integer cutoff h, retaining the exact three terms. all random count bounds come from the actual hoo law. theorem compiled","shard":"modules/9b924db43bd01fc1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_regret_dimension_sums","label":"expected_regret_dimension_sums","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_regret_dimension_sums","description":"Definition-5 dimension supplies one constant for every cutoff and horizon; the remaining finite geometric sums are explicit, not hidden in an O premise.","url":"../modules/banditrlproof-algorithms-hooexpectedregret/index.html#decl-1b21fa76e87a","parent":"module:BanditRLProof.Algorithms.HOOExpectedRegret","order":1053,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedRegret"],["Source","BanditRLProof/Algorithms/HOOExpectedRegret.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.expected_regret_dimension_sums {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best d : ℝ) (hmean : ∀x, (∫ y, y ∂law x)=f x) (hf : ∀x, f x≤best) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd : C.nearOptimalityDimension f best (4*C.nu1/C.nu2) < (d:EReal)) : ∃ K : ℝ, 0<K ∧ ∀ H N : ℕ, (∫ Y, (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ 4*(C.nu1*C.rho^H)*N + (∑ h ∈ Finset.range H, 4*(C.nu1*C.rho^h)*(K*(C.nu2*C.rho^h)^(-d))) + ∑ h ∈ Finset.range H, 8*(C.nu1*C.rho^h)*(K*(C.nu2*C.rho^h)^(-d)) * (8*Real.log (max (N:ℝ) 2)/(C.nu1*C.rho^(h+1))^2+4)","missing":[],"search":"expected_regret_dimension_sums banditrlproof.hoo.regularcovering.expected_regret_dimension_sums definition-5 dimension supplies one constant for every cutoff and horizon; the remaining finite geometric sums are explicit, not hidden in an o premise. theorem compiled","shard":"modules/9b924db43bd01fc1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visitTrace","label":"visitTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.visitTrace","description":"noncomputable def visitTrace (ν ρ : ℝ) (v : Node) (Y : ℕ → ℝ) : ActionTrace (Fin 2)","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-2f04ffdf4e75","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1054,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def visitTrace (ν ρ : ℝ) (v : Node) (Y : ℕ → ℝ) : ActionTrace (Fin 2)","missing":[],"search":"visittrace banditrlproof.hoo.visittrace noncomputable def visittrace (ν ρ : ℝ) (v : node) (y : ℕ → ℝ) : actiontrace (fin 2) definition compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_visitTrace","label":"measurable_visitTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_visitTrace","description":"theorem measurable_visitTrace (ν ρ : ℝ) (v : Node) (i : ℕ) : Measurable (fun Y => visitTrace ν ρ v Y i)","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-f0c9a57bb682","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1055,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_visitTrace (ν ρ : ℝ) (v : Node) (i : ℕ) : Measurable (fun Y => visitTrace ν ρ v Y i)","missing":[],"search":"measurable_visittrace banditrlproof.hoo.measurable_visittrace theorem measurable_visittrace (ν ρ : ℝ) (v : node) (i : ℕ) : measurable (fun y => visittrace ν ρ v y i) theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visitTrace_zero_iff","label":"visitTrace_zero_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visitTrace_zero_iff","description":"@[simp] theorem visitTrace_zero_iff (ν ρ : ℝ) (v : Node) (Y : ℕ → ℝ) (i : ℕ) : visitTrace ν ρ v Y i = 0 ↔ v <+: action ν ρ Y i","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-6a0236ae05ca","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1056,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem visitTrace_zero_iff (ν ρ : ℝ) (v : Node) (Y : ℕ → ℝ) (i : ℕ) : visitTrace ν ρ v Y i = 0 ↔ v <+: action ν ρ Y i","missing":[],"search":"visittrace_zero_iff banditrlproof.hoo.visittrace_zero_iff @[simp] theorem visittrace_zero_iff (ν ρ : ℝ) (v : node) (y : ℕ → ℝ) (i : ℕ) : visittrace ν ρ v y i = 0 ↔ v <+: action ν ρ y i theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visitTrace_pullCount","label":"visitTrace_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visitTrace_pullCount","description":"theorem visitTrace_pullCount (ν ρ : ℝ) (v : Node) (Y : ℕ → ℝ) (n : ℕ) : pullCount (visitTrace ν ρ v Y) 0 n = visits (history ν ρ Y n) v","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-219383fa9c53","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1057,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem visitTrace_pullCount (ν ρ : ℝ) (v : Node) (Y : ℕ → ℝ) (n : ℕ) : pullCount (visitTrace ν ρ v Y) 0 n = visits (history ν ρ Y n) v","missing":[],"search":"visittrace_pullcount banditrlproof.hoo.visittrace_pullcount theorem visittrace_pullcount (ν ρ : ℝ) (v : node) (y : ℕ → ℝ) (n : ℕ) : pullcount (visittrace ν ρ v y) 0 n = visits (history ν ρ y n) v theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.lintegral_visits_threshold","label":"lintegral_visits_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.lintegral_visits_threshold","description":"theorem lintegral_visits_threshold (ν ρ : ℝ) (v : Node) (μ : Measure (ℕ → ℝ)) [IsProbabilityMeasure μ] (N B : ℕ) : (∫⁻ Y, (visits (history ν ρ Y N) v : ENNReal) ∂μ) ≤ B + ∑ n ∈ Finset.range N, μ {Y | v <+: action ν ρ Y n ∧ B ≤ visits (history ν ρ Y n) v}","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-257cc3c32cfd","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1058,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_visits_threshold (ν ρ : ℝ) (v : Node) (μ : Measure (ℕ → ℝ)) [IsProbabilityMeasure μ] (N B : ℕ) : (∫⁻ Y, (visits (history ν ρ Y N) v : ENNReal) ∂μ) ≤ B + ∑ n ∈ Finset.range N, μ {Y | v <+: action ν ρ Y n ∧ B ≤ visits (history ν ρ Y n) v}","missing":[],"search":"lintegral_visits_threshold banditrlproof.hoo.lintegral_visits_threshold theorem lintegral_visits_threshold (ν ρ : ℝ) (v : node) (μ : measure (ℕ → ℝ)) [isprobabilitymeasure μ] (n b : ℕ) : (∫⁻ y, (visits (history ν ρ y n) v : ennreal) ∂μ) ≤ b + ∑ n ∈ finset.range n, μ {y | v <+: action ν ρ y n ∧ b ≤ visits (history ν ρ y n) v} theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visitThreshold","label":"visitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.visitThreshold","description":"noncomputable def visitThreshold (gap : ℝ) (N : ℕ) : ℕ","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-27e35903df0c","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1059,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def visitThreshold (gap : ℝ) (N : ℕ) : ℕ","missing":[],"search":"visitthreshold banditrlproof.hoo.visitthreshold noncomputable def visitthreshold (gap : ℝ) (n : ℕ) : ℕ definition compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visitThreshold_controls","label":"visitThreshold_controls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visitThreshold_controls","description":"theorem visitThreshold_controls (gap : ℝ) (N n : ℕ) (hn : n ≤ N) : 8*Real.log (max (n:ℝ) 2)/gap^2 ≤ (visitThreshold gap N : ℝ)","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-8b2a483c9d29","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1060,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem visitThreshold_controls (gap : ℝ) (N n : ℕ) (hn : n ≤ N) : 8*Real.log (max (n:ℝ) 2)/gap^2 ≤ (visitThreshold gap N : ℝ)","missing":[],"search":"visitthreshold_controls banditrlproof.hoo.visitthreshold_controls theorem visitthreshold_controls (gap : ℝ) (n n : ℕ) (hn : n ≤ n) : 8*real.log (max (n:ℝ) 2)/gap^2 ≤ (visitthreshold gap n : ℝ) theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.poor_region_lintegral_visits","label":"poor_region_lintegral_visits","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.poor_region_lintegral_visits","description":"theorem RegularCovering.poor_region_lintegral_visits {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionSup f Set.univ = best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.re…","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-0fabc4e03d95","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1061,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.poor_region_lintegral_visits {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionSup f Set.univ = best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.region v)) (N : ℕ) : (∫⁻ Y, (visits (history C.nu1 C.rho Y N) v : ENNReal) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ visitThreshold (best-regionSup f (C.region v)-C.nu1*C.rho^v.length) N + 3","missing":[],"search":"poor_region_lintegral_visits banditrlproof.hoo.regularcovering.poor_region_lintegral_visits theorem regularcovering.poor_region_lintegral_visits {x : type*} [measurablespace x] (c : regularcovering x) (law : kernel x ℝ) [ismarkovkernel law] (f : x → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionsup f set.univ = best) (hw : weaklylipschitz f c.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ set.icc (0 : ℝ) 1) (v : node) (hv : c.nu1*c.rho^v.length < best-regionsup f (c.region v)) (n : ℕ) : (∫⁻ y, (visits (history c.nu1 c.rho y n) v : ennreal) ∂trajectory c.nu1 c.rho (c.tocovering.nodelaw law)) ≤ visitthreshold (best-regionsup f (c.region v)-c.nu1*c.rho^v.length) n + 3 theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.integrable_visits","label":"integrable_visits","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.integrable_visits","description":"theorem integrable_visits (ν ρ : ℝ) (v : Node) (μ : Measure (ℕ → ℝ)) [IsProbabilityMeasure μ] (N : ℕ) : Integrable (fun Y => (visits (history ν ρ Y N) v : ℝ)) μ","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-1c8b3d56405a","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1062,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_visits (ν ρ : ℝ) (v : Node) (μ : Measure (ℕ → ℝ)) [IsProbabilityMeasure μ] (N : ℕ) : Integrable (fun Y => (visits (history ν ρ Y N) v : ℝ)) μ","missing":[],"search":"integrable_visits banditrlproof.hoo.integrable_visits theorem integrable_visits (ν ρ : ℝ) (v : node) (μ : measure (ℕ → ℝ)) [isprobabilitymeasure μ] (n : ℕ) : integrable (fun y => (visits (history ν ρ y n) v : ℝ)) μ theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.poor_region_expected_visits","label":"poor_region_expected_visits","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.poor_region_expected_visits","description":"Expected poor-region visits, with the source additive-four shape and an explicit logarithm repair at horizons zero and one. All probability and search premises are produced from the actual HOO algorithm and A1/A2 model.","url":"../modules/banditrlproof-algorithms-hooexpectedvisits/index.html#decl-1b2d830f0557","parent":"module:BanditRLProof.Algorithms.HOOExpectedVisits","order":1063,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOExpectedVisits"],["Source","BanditRLProof/Algorithms/HOOExpectedVisits.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.poor_region_expected_visits {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionSup f Set.univ = best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.region v)) (N : ℕ) : (∫ Y, (visits (history C.nu1 C.rho Y N) v : ℝ) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ 8*Real.log (max (N:ℝ) 2)/(best-regionSup f (C.region v)-C.nu1*C.rho^v.length)^2 + 4","missing":[],"search":"poor_region_expected_visits banditrlproof.hoo.regularcovering.poor_region_expected_visits expected poor-region visits, with the source additive-four shape and an explicit logarithm repair at horizons zero and one. all probability and search premises are produced from the actual hoo algorithm and a1/a2 model. theorem compiled","shard":"modules/4d0f92fff34a01ff.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Observations","label":"Observations","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.HOO.Observations","description":"abbrev Observations","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-9273d0cdbcb6","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1064,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Observations","missing":[],"search":"observations banditrlproof.hoo.observations abbrev observations abbreviation compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.expanded","label":"expanded","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.expanded","description":"def expanded (h : Observations) : Finset Node","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-1eee381f986d","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1065,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def expanded (h : Observations) : Finset Node","missing":[],"search":"expanded banditrlproof.hoo.expanded def expanded (h : observations) : finset node definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visits","label":"visits","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.visits","description":"def visits (h : Observations) (v : Node) : ℕ","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-988fc30a2bb0","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1066,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def visits (h : Observations) (v : Node) : ℕ","missing":[],"search":"visits banditrlproof.hoo.visits def visits (h : observations) (v : node) : ℕ definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.rewardSum","label":"rewardSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.rewardSum","description":"noncomputable def rewardSum (h : Observations) (v : Node) : ℝ","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-b85be9dc3308","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1067,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardSum (h : Observations) (v : Node) : ℝ","missing":[],"search":"rewardsum banditrlproof.hoo.rewardsum noncomputable def rewardsum (h : observations) (v : node) : ℝ definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.upper","label":"upper","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.upper","description":"noncomputable def upper (ν ρ : ℝ) (h : Observations) (v : Node) : WithTop ℝ","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-6036a7d194fb","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1068,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def upper (ν ρ : ℝ) (h : Observations) (v : Node) : WithTop ℝ","missing":[],"search":"upper banditrlproof.hoo.upper noncomputable def upper (ν ρ : ℝ) (h : observations) (v : node) : withtop ℝ definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.next","label":"next","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.next","description":"noncomputable def next (ν ρ : ℝ) (h : Observations) : Node","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-93e72951fe4c","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1069,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def next (ν ρ : ℝ) (h : Observations) : Node","missing":[],"search":"next banditrlproof.hoo.next noncomputable def next (ν ρ : ℝ) (h : observations) : node definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.step","label":"step","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.step","description":"noncomputable def step (ν ρ : ℝ) (h : Observations) (y : ℝ) : Observations","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-cefee2b38b71","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1070,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def step (ν ρ : ℝ) (h : Observations) (y : ℝ) : Observations","missing":[],"search":"step banditrlproof.hoo.step noncomputable def step (ν ρ : ℝ) (h : observations) (y : ℝ) : observations definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.history","label":"history","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.history","description":"noncomputable def history (ν ρ : ℝ) (Y : ℕ → ℝ) : ℕ → Observations | 0 => [] | n+1 => step ν ρ (history ν ρ Y n) (Y n) noncomputable def action (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : Node","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-7d8e0b8fdf04","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1071,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def history (ν ρ : ℝ) (Y : ℕ → ℝ) : ℕ → Observations | 0 => [] | n+1 => step ν ρ (history ν ρ Y n) (Y n) noncomputable def action (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : Node","missing":[],"search":"history banditrlproof.hoo.history noncomputable def history (ν ρ : ℝ) (y : ℕ → ℝ) : ℕ → observations | 0 => [] | n+1 => step ν ρ (history ν ρ y n) (y n) noncomputable def action (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : node definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action","label":"action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.action","description":"noncomputable def action (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : Node","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-f55808553574","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1072,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"noncomputable def action (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : Node","missing":[],"search":"action banditrlproof.hoo.action noncomputable def action (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : node definition compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.step_length","label":"step_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.step_length","description":"@[simp] theorem step_length (ν ρ : ℝ) (h : Observations) (y : ℝ) : (step ν ρ h y).length = h.length + 1","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-ae4d197c0517","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1073,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem step_length (ν ρ : ℝ) (h : Observations) (y : ℝ) : (step ν ρ h y).length = h.length + 1","missing":[],"search":"step_length banditrlproof.hoo.step_length @[simp] theorem step_length (ν ρ : ℝ) (h : observations) (y : ℝ) : (step ν ρ h y).length = h.length + 1 theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.history_length","label":"history_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.history_length","description":"@[simp] theorem history_length (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : (history ν ρ Y n).length = n","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-387aa5821545","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1074,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem history_length (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : (history ν ρ Y n).length = n","missing":[],"search":"history_length banditrlproof.hoo.history_length @[simp] theorem history_length (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : (history ν ρ y n).length = n theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.expanded_step","label":"expanded_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.expanded_step","description":"theorem expanded_step (ν ρ : ℝ) (h : Observations) (y : ℝ) : expanded (step ν ρ h y) = insert (next ν ρ h) (expanded h)","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-5d87a2f8bd1b","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1075,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expanded_step (ν ρ : ℝ) (h : Observations) (y : ℝ) : expanded (step ν ρ h y) = insert (next ν ρ h) (expanded h)","missing":[],"search":"expanded_step banditrlproof.hoo.expanded_step theorem expanded_step (ν ρ : ℝ) (h : observations) (y : ℝ) : expanded (step ν ρ h y) = insert (next ν ρ h) (expanded h) theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.next_not_expanded","label":"next_not_expanded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.next_not_expanded","description":"theorem next_not_expanded (ν ρ : ℝ) (h : Observations) : next ν ρ h ∉ expanded h","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-e79291e682a1","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1076,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem next_not_expanded (ν ρ : ℝ) (h : Observations) : next ν ρ h ∉ expanded h","missing":[],"search":"next_not_expanded banditrlproof.hoo.next_not_expanded theorem next_not_expanded (ν ρ : ℝ) (h : observations) : next ν ρ h ∉ expanded h theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.next_ne_root","label":"next_ne_root","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.next_ne_root","description":"theorem next_ne_root (ν ρ : ℝ) (h : Observations) : next ν ρ h ≠ []","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-049227eb2f92","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1077,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem next_ne_root (ν ρ : ℝ) (h : Observations) : next ν ρ h ≠ []","missing":[],"search":"next_ne_root banditrlproof.hoo.next_ne_root theorem next_ne_root (ν ρ : ℝ) (h : observations) : next ν ρ h ≠ [] theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.next_not_previously_played","label":"next_not_previously_played","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.next_not_previously_played","description":"theorem next_not_previously_played (ν ρ : ℝ) (h : Observations) : next ν ρ h ∉ h.map Prod.fst","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-ebd86bc16ff8","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1078,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem next_not_previously_played (ν ρ : ℝ) (h : Observations) : next ν ρ h ∉ h.map Prod.fst","missing":[],"search":"next_not_previously_played banditrlproof.hoo.next_not_previously_played theorem next_not_previously_played (ν ρ : ℝ) (h : observations) : next ν ρ h ∉ h.map prod.fst theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visits_step","label":"visits_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visits_step","description":"theorem visits_step (ν ρ : ℝ) (h : Observations) (y : ℝ) (v : Node) : visits (step ν ρ h y) v = visits h v + if v <+: next ν ρ h then 1 else 0","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-f1deff0ed965","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1079,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem visits_step (ν ρ : ℝ) (h : Observations) (y : ℝ) (v : Node) : visits (step ν ρ h y) v = visits h v + if v <+: next ν ρ h then 1 else 0","missing":[],"search":"visits_step banditrlproof.hoo.visits_step theorem visits_step (ν ρ : ℝ) (h : observations) (y : ℝ) (v : node) : visits (step ν ρ h y) v = visits h v + if v <+: next ν ρ h then 1 else 0 theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.rewardSum_step","label":"rewardSum_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.rewardSum_step","description":"theorem rewardSum_step (ν ρ : ℝ) (h : Observations) (y : ℝ) (v : Node) : rewardSum (step ν ρ h y) v = rewardSum h v + if v <+: next ν ρ h then y else 0","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-93fd40a2d704","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1080,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rewardSum_step (ν ρ : ℝ) (h : Observations) (y : ℝ) (v : Node) : rewardSum (step ν ρ h y) v = rewardSum h v + if v <+: next ν ρ h then y else 0","missing":[],"search":"rewardsum_step banditrlproof.hoo.rewardsum_step theorem rewardsum_step (ν ρ : ℝ) (h : observations) (y : ℝ) (v : node) : rewardsum (step ν ρ h y) v = rewardsum h v + if v <+: next ν ρ h then y else 0 theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.history_causal","label":"history_causal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.history_causal","description":"Generated state uses only rewards strictly before the current round.","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-0884c593357a","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1081,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem history_causal (ν ρ : ℝ) (Y Z : ℕ → ℝ) (n : ℕ) (hYZ : ∀ i < n, Y i = Z i) : history ν ρ Y n = history ν ρ Z n","missing":[],"search":"history_causal banditrlproof.hoo.history_causal generated state uses only rewards strictly before the current round. theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_causal","label":"action_causal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_causal","description":"theorem action_causal (ν ρ : ℝ) (Y Z : ℕ → ℝ) (n : ℕ) (hYZ : ∀ i < n, Y i = Z i) : action ν ρ Y n = action ν ρ Z n","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-93ee4b3a985a","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1082,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_causal (ν ρ : ℝ) (Y Z : ℕ → ℝ) (n : ℕ) (hYZ : ∀ i < n, Y i = Z i) : action ν ρ Y n = action ν ρ Z n","missing":[],"search":"action_causal banditrlproof.hoo.action_causal theorem action_causal (ν ρ : ℝ) (y z : ℕ → ℝ) (n : ℕ) (hyz : ∀ i < n, y i = z i) : action ν ρ y n = action ν ρ z n theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.expanded_card_history","label":"expanded_card_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.expanded_card_history","description":"theorem expanded_card_history (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : (expanded (history ν ρ Y n)).card = n+1","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-efe1079d5aad","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1083,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expanded_card_history (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : (expanded (history ν ρ Y n)).card = n+1","missing":[],"search":"expanded_card_history banditrlproof.hoo.expanded_card_history theorem expanded_card_history (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : (expanded (history ν ρ y n)).card = n+1 theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.depthBound_history","label":"depthBound_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.depthBound_history","description":"theorem depthBound_history (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : depthBound (expanded (history ν ρ Y n)) ≤ n","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-d70b719719c5","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1084,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem depthBound_history (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : depthBound (expanded (history ν ρ Y n)) ≤ n","missing":[],"search":"depthbound_history banditrlproof.hoo.depthbound_history theorem depthbound_history (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : depthbound (expanded (history ν ρ y n)) ≤ n theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_depth_le","label":"action_depth_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_depth_le","description":"theorem action_depth_le (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : (action ν ρ Y n).length ≤ n+1","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-97d977685edd","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1085,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:107"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_depth_le (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : (action ν ρ Y n).length ≤ n+1","missing":[],"search":"action_depth_le banditrlproof.hoo.action_depth_le theorem action_depth_le (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : (action ν ρ y n).length ≤ n+1 theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.history_eq_ofFn","label":"history_eq_ofFn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.history_eq_ofFn","description":"The stored history is precisely the generated action/reward trace.","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-3f5ff1217e07","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1086,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem history_eq_ofFn (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : history ν ρ Y n = List.ofFn (fun i : Fin n => (action ν ρ Y i, Y i))","missing":[],"search":"history_eq_offn banditrlproof.hoo.history_eq_offn the stored history is precisely the generated action/reward trace. theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_ne_of_lt","label":"action_ne_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_ne_of_lt","description":"theorem action_ne_of_lt (ν ρ : ℝ) (Y : ℕ → ℝ) {m n : ℕ} (hmn : m < n) : action ν ρ Y n ≠ action ν ρ Y m","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-d2eee8f7be6a","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1087,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_ne_of_lt (ν ρ : ℝ) (Y : ℕ → ℝ) {m n : ℕ} (hmn : m < n) : action ν ρ Y n ≠ action ν ρ Y m","missing":[],"search":"action_ne_of_lt banditrlproof.hoo.action_ne_of_lt theorem action_ne_of_lt (ν ρ : ℝ) (y : ℕ → ℝ) {m n : ℕ} (hmn : m < n) : action ν ρ y n ≠ action ν ρ y m theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_injective","label":"action_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_injective","description":"theorem action_injective (ν ρ : ℝ) (Y : ℕ → ℝ) : Function.Injective (action ν ρ Y)","url":"../modules/banditrlproof-algorithms-hoohistory/index.html#decl-1475bb61ee0f","parent":"module:BanditRLProof.Algorithms.HOOHistory","order":1088,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOHistory"],["Source","BanditRLProof/Algorithms/HOOHistory.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_injective (ν ρ : ℝ) (Y : ℕ → ℝ) : Function.Injective (action ν ρ Y)","missing":[],"search":"action_injective banditrlproof.hoo.action_injective theorem action_injective (ν ρ : ℝ) (y : ℕ → ℝ) : function.injective (action ν ρ y) theorem compiled","shard":"modules/31aaa32a23098242.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regional_means_lower","label":"regional_means_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.regional_means_lower","description":"theorem regional_means_lower (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (m : ℝ) (hm : ∀ a, v <+: a → m ≤ nodeMean law a) (n : ℕ) (Y : ℕ → ℝ) : m * (visits (history ν ρ Y n) v : ℝ) ≤ ∑ i ∈ Finset.range n, regionCount ν ρ v i Y * nodeMean law (action ν ρ Y i)","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-f7d8e7dab3c0","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1089,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem regional_means_lower (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (m : ℝ) (hm : ∀ a, v <+: a → m ≤ nodeMean law a) (n : ℕ) (Y : ℕ → ℝ) : m * (visits (history ν ρ Y n) v : ℝ) ≤ ∑ i ∈ Finset.range n, regionCount ν ρ v i Y * nodeMean law (action ν ρ Y i)","missing":[],"search":"regional_means_lower banditrlproof.hoo.regional_means_lower theorem regional_means_lower (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (m : ℝ) (hm : ∀ a, v <+: a → m ≤ nodemean law a) (n : ℕ) (y : ℕ → ℝ) : m * (visits (history ν ρ y n) v : ℝ) ≤ ∑ i ∈ finset.range n, regioncount ν ρ v i y * nodemean law (action ν ρ y i) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regional_means_upper","label":"regional_means_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.regional_means_upper","description":"theorem regional_means_upper (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (m : ℝ) (hm : ∀ a, v <+: a → nodeMean law a ≤ m) (n : ℕ) (Y : ℕ → ℝ) : (∑ i ∈ Finset.range n, regionCount ν ρ v i Y * nodeMean law (action ν ρ Y i)) ≤ m * (visits (history ν ρ Y n) v : ℝ)","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-ff48c1764864","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1090,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem regional_means_upper (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (m : ℝ) (hm : ∀ a, v <+: a → nodeMean law a ≤ m) (n : ℕ) (Y : ℕ → ℝ) : (∑ i ∈ Finset.range n, regionCount ν ρ v i Y * nodeMean law (action ν ρ Y i)) ≤ m * (visits (history ν ρ Y n) v : ℝ)","missing":[],"search":"regional_means_upper banditrlproof.hoo.regional_means_upper theorem regional_means_upper (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (m : ℝ) (hm : ∀ a, v <+: a → nodemean law a ≤ m) (n : ℕ) (y : ℕ → ℝ) : (∑ i ∈ finset.range n, regioncount ν ρ v i y * nodemean law (action ν ρ y i)) ≤ m * (visits (history ν ρ y n) v : ℝ) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.count_mul_width","label":"count_mul_width","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.count_mul_width","description":"theorem count_mul_width (T L : ℝ) (hT : 0 < T) : T * Real.sqrt (2*L/T) = Real.sqrt (2*T*L)","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-601326a99713","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1091,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem count_mul_width (T L : ℝ) (hT : 0 < T) : T * Real.sqrt (2*L/T) = Real.sqrt (2*T*L)","missing":[],"search":"count_mul_width banditrlproof.hoo.count_mul_width theorem count_mul_width (t l : ℝ) (ht : 0 < t) : t * real.sqrt (2*l/t) = real.sqrt (2*t*l) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.upper_le_implies_lower_deviation","label":"upper_le_implies_lower_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.upper_le_implies_lower_deviation","description":"A low U value in a region whose descendant means exceed `best-D` forces a lower centered-noise deviation, on the exact generated history.","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-03ff0eaaf936","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1092,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem upper_le_implies_lower_deviation (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (best : ℝ) (hm : ∀ a, v <+: a → best - ν*ρ^v.length ≤ nodeMean law a) (n : ℕ) (Y : ℕ → ℝ) (hu : upper ν ρ (history ν ρ Y n) v ≤ (best : WithTop ℝ)) : 0 < visits (history ν ρ Y n) v ∧ Real.sqrt (2*(visits (history ν ρ Y n) v : ℝ)*Real.log (max (n:ℝ) 2)) ≤ regionDeviation ν ρ law v n true Y","missing":[],"search":"upper_le_implies_lower_deviation banditrlproof.hoo.upper_le_implies_lower_deviation a low u value in a region whose descendant means exceed `best-d` forces a lower centered-noise deviation, on the exact generated history. theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.upper_underestimate_probability","label":"upper_underestimate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.upper_underestimate_probability","description":"theorem upper_underestimate_probability (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (best : ℝ) (hm : ∀ a, v <+: a → best - ν*ρ^v.length ≤ nodeMean law a) (n : ℕ) : (trajectory ν ρ law) {Y | upper ν ρ (history ν ρ Y n) v ≤ (best : WithTop ℝ)} ≤ (n : ENNReal) * ENNReal.ofReal (Real.exp (-4*Real.log (max (n:ℝ) 2)))","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-ba2102ca3603","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1093,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem upper_underestimate_probability (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (best : ℝ) (hm : ∀ a, v <+: a → best - ν*ρ^v.length ≤ nodeMean law a) (n : ℕ) : (trajectory ν ρ law) {Y | upper ν ρ (history ν ρ Y n) v ≤ (best : WithTop ℝ)} ≤ (n : ENNReal) * ENNReal.ofReal (Real.exp (-4*Real.log (max (n:ℝ) 2)))","missing":[],"search":"upper_underestimate_probability banditrlproof.hoo.upper_underestimate_probability theorem upper_underestimate_probability (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (best : ℝ) (hm : ∀ a, v <+: a → best - ν*ρ^v.length ≤ nodemean law a) (n : ℕ) : (trajectory ν ρ law) {y | upper ν ρ (history ν ρ y n) v ≤ (best : withtop ℝ)} ≤ (n : ennreal) * ennreal.ofreal (real.exp (-4*real.log (max (n:ℝ) 2))) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.width_le_half_gap","label":"width_le_half_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.width_le_half_gap","description":"theorem width_le_half_gap (T L gap : ℝ) (hT : 0 < T) (hgap : 0 < gap) (hcount : 8*L/gap^2 ≤ T) : Real.sqrt (2*L/T) ≤ gap/2","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-9ef3f024258a","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1094,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem width_le_half_gap (T L gap : ℝ) (hT : 0 < T) (hgap : 0 < gap) (hcount : 8*L/gap^2 ≤ T) : Real.sqrt (2*L/T) ≤ gap/2","missing":[],"search":"width_le_half_gap banditrlproof.hoo.width_le_half_gap theorem width_le_half_gap (t l gap : ℝ) (ht : 0 < t) (hgap : 0 < gap) (hcount : 8*l/gap^2 ≤ t) : real.sqrt (2*l/t) ≤ gap/2 theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.upper_ge_implies_upper_deviation","label":"upper_ge_implies_upper_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.upper_ge_implies_upper_deviation","description":"theorem upper_ge_implies_upper_deviation (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (best gap : ℝ) (hgap : 0 < gap - ν*ρ^v.length) (hm : ∀ a, v <+: a → nodeMean law a ≤ best-gap) (n : ℕ) (Y : ℕ → ℝ) (hcount : 8*Real.log (max (n:ℝ) 2)/(gap-ν*ρ^v.length)^2 ≤ (visits (history ν ρ Y n) v : ℝ)) (hu : (best : WithTop ℝ) ≤ upper ν ρ (history ν ρ Y n) v) : 0 < visits (history ν ρ Y n) v ∧ Real.sqrt (2*(visits (history ν ρ Y…","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-e5ec5160dc52","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1095,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem upper_ge_implies_upper_deviation (ν ρ : ℝ) (law : Kernel Node ℝ) (v : Node) (best gap : ℝ) (hgap : 0 < gap - ν*ρ^v.length) (hm : ∀ a, v <+: a → nodeMean law a ≤ best-gap) (n : ℕ) (Y : ℕ → ℝ) (hcount : 8*Real.log (max (n:ℝ) 2)/(gap-ν*ρ^v.length)^2 ≤ (visits (history ν ρ Y n) v : ℝ)) (hu : (best : WithTop ℝ) ≤ upper ν ρ (history ν ρ Y n) v) : 0 < visits (history ν ρ Y n) v ∧ Real.sqrt (2*(visits (history ν ρ Y n) v : ℝ)*Real.log (max (n:ℝ) 2)) ≤ regionDeviation ν ρ law v n false Y","missing":[],"search":"upper_ge_implies_upper_deviation banditrlproof.hoo.upper_ge_implies_upper_deviation theorem upper_ge_implies_upper_deviation (ν ρ : ℝ) (law : kernel node ℝ) (v : node) (best gap : ℝ) (hgap : 0 < gap - ν*ρ^v.length) (hm : ∀ a, v <+: a → nodemean law a ≤ best-gap) (n : ℕ) (y : ℕ → ℝ) (hcount : 8*real.log (max (n:ℝ) 2)/(gap-ν*ρ^v.length)^2 ≤ (visits (history ν ρ y n) v : ℝ)) (hu : (best : withtop ℝ) ≤ upper ν ρ (history ν ρ y n) v) : 0 < visits (history ν ρ y n) v ∧ real.sqrt (2*(visits (history ν ρ y n) v : ℝ)*real.log (max (n:ℝ) 2)) ≤ regiondeviation ν ρ law v n false y theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.upper_overestimate_probability","label":"upper_overestimate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.upper_overestimate_probability","description":"theorem upper_overestimate_probability (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (best gap : ℝ) (hgap : 0 < gap - ν*ρ^v.length) (hm : ∀ a, v <+: a → nodeMean law a ≤ best-gap) (n : ℕ) : (trajectory ν ρ law) {Y | 8*Real.log (max (n:ℝ) 2)/(gap-ν*ρ^v.length)^2 ≤ (visits (history ν ρ Y n) v : ℝ) ∧ (best : WithTop ℝ) ≤ upper ν ρ (history ν ρ Y n) v}…","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-e074cc7548d2","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1096,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem upper_overestimate_probability (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (best gap : ℝ) (hgap : 0 < gap - ν*ρ^v.length) (hm : ∀ a, v <+: a → nodeMean law a ≤ best-gap) (n : ℕ) : (trajectory ν ρ law) {Y | 8*Real.log (max (n:ℝ) 2)/(gap-ν*ρ^v.length)^2 ≤ (visits (history ν ρ Y n) v : ℝ) ∧ (best : WithTop ℝ) ≤ upper ν ρ (history ν ρ Y n) v} ≤ (n : ENNReal) * ENNReal.ofReal (Real.exp (-4*Real.log (max (n:ℝ) 2)))","missing":[],"search":"upper_overestimate_probability banditrlproof.hoo.upper_overestimate_probability theorem upper_overestimate_probability (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] (hbound : ∀ a, ∀ᵐ y ∂law a, y ∈ set.icc (0 : ℝ) 1) (v : node) (best gap : ℝ) (hgap : 0 < gap - ν*ρ^v.length) (hm : ∀ a, v <+: a → nodemean law a ≤ best-gap) (n : ℕ) : (trajectory ν ρ law) {y | 8*real.log (max (n:ℝ) 2)/(gap-ν*ρ^v.length)^2 ≤ (visits (history ν ρ y n) v : ℝ) ∧ (best : withtop ℝ) ≤ upper ν ρ (history ν ρ y n) v} ≤ (n : ennreal) * ennreal.ofreal (real.exp (-4*real.log (max (n:ℝ) 2))) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.nodeMean_eq","label":"nodeMean_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.nodeMean_eq","description":"theorem Covering.nodeMean_eq {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) (f : X → ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (a : Node) : nodeMean (C.nodeLaw law) a = f (C.representative a)","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-ec3c0349d615","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1097,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.nodeMean_eq {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) (f : X → ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (a : Node) : nodeMean (C.nodeLaw law) a = f (C.representative a)","missing":[],"search":"nodemean_eq banditrlproof.hoo.covering.nodemean_eq theorem covering.nodemean_eq {x : type*} [measurablespace x] (c : covering x) (law : kernel x ℝ) (f : x → ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (a : node) : nodemean (c.nodelaw law) a = f (c.representative a) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.optimal_descendant_mean_lower","label":"optimal_descendant_mean_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.optimal_descendant_mean_lower","description":"theorem RegularCovering.optimal_descendant_mean_lower {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hw : WeaklyLipschitz f C.ell best) (v : Node) (hv : regionSup f (C.region v) = best) : ∀ a, v <+: a → best - C.nu1*C.rho^v.length ≤ nodeMean (C.toCovering.nodeLaw law) a","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-b4bb5db5f700","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1098,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.optimal_descendant_mean_lower {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hw : WeaklyLipschitz f C.ell best) (v : Node) (hv : regionSup f (C.region v) = best) : ∀ a, v <+: a → best - C.nu1*C.rho^v.length ≤ nodeMean (C.toCovering.nodeLaw law) a","missing":[],"search":"optimal_descendant_mean_lower banditrlproof.hoo.regularcovering.optimal_descendant_mean_lower theorem regularcovering.optimal_descendant_mean_lower {x : type*} [measurablespace x] (c : regularcovering x) (law : kernel x ℝ) (f : x → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hw : weaklylipschitz f c.ell best) (v : node) (hv : regionsup f (c.region v) = best) : ∀ a, v <+: a → best - c.nu1*c.rho^v.length ≤ nodemean (c.tocovering.nodelaw law) a theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.descendant_mean_upper","label":"descendant_mean_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.descendant_mean_upper","description":"theorem Covering.descendant_mean_upper {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (v : Node) : ∀ a, v <+: a → nodeMean (C.nodeLaw law) a ≤ regionSup f (C.region v)","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-2a012c051f77","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1099,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.descendant_mean_upper {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (v : Node) : ∀ a, v <+: a → nodeMean (C.nodeLaw law) a ≤ regionSup f (C.region v)","missing":[],"search":"descendant_mean_upper banditrlproof.hoo.covering.descendant_mean_upper theorem covering.descendant_mean_upper {x : type*} [measurablespace x] (c : covering x) (law : kernel x ℝ) (f : x → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (v : node) : ∀ a, v <+: a → nodemean (c.nodelaw law) a ≤ regionsup f (c.region v) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.optimal_upper_underestimate_probability","label":"optimal_upper_underestimate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.optimal_upper_underestimate_probability","description":"Source optimal-region underestimation bound, now produced from A1/A2, the reward-mean identity and bounded reward support. No confidence premise.","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-8b8a024175b0","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1100,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:154"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.optimal_upper_underestimate_probability {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : regionSup f (C.region v) = best) (n : ℕ) : (trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) {Y | upper C.nu1 C.rho (history C.nu1 C.rho Y n) v ≤ (best : WithTop ℝ)} ≤ (n : ENNReal) * ENNReal.ofReal (Real.exp (-4*Real.log (max (n:ℝ) 2)))","missing":[],"search":"optimal_upper_underestimate_probability banditrlproof.hoo.regularcovering.optimal_upper_underestimate_probability source optimal-region underestimation bound, now produced from a1/a2, the reward-mean identity and bounded reward support. no confidence premise. theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.poor_upper_overestimate_probability","label":"poor_upper_overestimate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.poor_upper_overestimate_probability","description":"theorem RegularCovering.poor_upper_overestimate_probability {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.region v)) (n : ℕ) : (trajectory C.nu1 C.rho (C.toCovering.nodeLaw la…","url":"../modules/banditrlproof-algorithms-hooindexconfidence/index.html#decl-5e23e85749bf","parent":"module:BanditRLProof.Algorithms.HOOIndexConfidence","order":1101,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOIndexConfidence"],["Source","BanditRLProof/Algorithms/HOOIndexConfidence.lean:167"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.poor_upper_overestimate_probability {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.region v)) (n : ℕ) : (trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) {Y | 8*Real.log (max (n:ℝ) 2)/(best-regionSup f (C.region v)-C.nu1*C.rho^v.length)^2 ≤ (visits (history C.nu1 C.rho Y n) v : ℝ) ∧ (best : WithTop ℝ) ≤ upper C.nu1 C.rho (history C.nu1 C.rho Y n) v} ≤ (n : ENNReal) * ENNReal.ofReal (Real.exp (-4*Real.log (max (n:ℝ) 2)))","missing":[],"search":"poor_upper_overestimate_probability banditrlproof.hoo.regularcovering.poor_upper_overestimate_probability theorem regularcovering.poor_upper_overestimate_probability {x : type*} [measurablespace x] (c : regularcovering x) (law : kernel x ℝ) [ismarkovkernel law] (f : x → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ set.icc (0 : ℝ) 1) (v : node) (hv : c.nu1*c.rho^v.length < best-regionsup f (c.region v)) (n : ℕ) : (trajectory c.nu1 c.rho (c.tocovering.nodelaw law)) {y | 8*real.log (max (n:ℝ) 2)/(best-regionsup f (c.region v)-c.nu1*c.rho^v.length)^2 ≤ (visits (history c.nu1 c.rho y n) v : ℝ) ∧ (best : withtop ℝ) ≤ upper c.nu1 c.rho (history c.nu1 c.rho y n) v} ≤ (n : ennreal) * ennreal.ofreal (real.exp (-4*real.log (max (n:ℝ) 2))) theorem compiled","shard":"modules/ddb172e710ed8869.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_backward","label":"measurable_backward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_backward","description":"theorem measurable_backward {Ω : Type*} [MeasurableSpace Ω] (S : Finset Node) (U : Ω → Node → WithTop ℝ) (hU : ∀ v, Measurable (fun ω => U ω v)) (k : ℕ) (v : Node) : Measurable (fun ω => backward S (U ω) k v)","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-434fcd71e95c","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1102,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_backward {Ω : Type*} [MeasurableSpace Ω] (S : Finset Node) (U : Ω → Node → WithTop ℝ) (hU : ∀ v, Measurable (fun ω => U ω v)) (k : ℕ) (v : Node) : Measurable (fun ω => backward S (U ω) k v)","missing":[],"search":"measurable_backward banditrlproof.hoo.measurable_backward theorem measurable_backward {ω : type*} [measurablespace ω] (s : finset node) (u : ω → node → withtop ℝ) (hu : ∀ v, measurable (fun ω => u ω v)) (k : ℕ) (v : node) : measurable (fun ω => backward s (u ω) k v) theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_walk","label":"measurable_walk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_walk","description":"theorem measurable_walk {Ω : Type*} [MeasurableSpace Ω] (S : Finset Node) (B : Ω → Node → WithTop ℝ) (hB : ∀ v, Measurable (fun ω => B ω v)) (k : ℕ) (v : Node) : Measurable (fun ω => walk S (B ω) k v)","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-bd836a800b61","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1103,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_walk {Ω : Type*} [MeasurableSpace Ω] (S : Finset Node) (B : Ω → Node → WithTop ℝ) (hB : ∀ v, Measurable (fun ω => B ω v)) (k : ℕ) (v : Node) : Measurable (fun ω => walk S (B ω) k v)","missing":[],"search":"measurable_walk banditrlproof.hoo.measurable_walk theorem measurable_walk {ω : type*} [measurablespace ω] (s : finset node) (b : ω → node → withtop ℝ) (hb : ∀ v, measurable (fun ω => b ω v)) (k : ℕ) (v : node) : measurable (fun ω => walk s (b ω) k v) theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_select","label":"measurable_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_select","description":"theorem measurable_select {Ω : Type*} [MeasurableSpace Ω] (S : Finset Node) (U : Ω → Node → WithTop ℝ) (hU : ∀ v, Measurable (fun ω => U ω v)) : Measurable (fun ω => select S (U ω))","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-a5aa43e6c271","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1104,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_select {Ω : Type*} [MeasurableSpace Ω] (S : Finset Node) (U : Ω → Node → WithTop ℝ) (hU : ∀ v, Measurable (fun ω => U ω v)) : Measurable (fun ω => select S (U ω))","missing":[],"search":"measurable_select banditrlproof.hoo.measurable_select theorem measurable_select {ω : type*} [measurablespace ω] (s : finset node) (u : ω → node → withtop ℝ) (hu : ∀ v, measurable (fun ω => u ω v)) : measurable (fun ω => select s (u ω)) theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visits_ofFn","label":"visits_ofFn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visits_ofFn","description":"theorem visits_ofFn {n : ℕ} (a : Fin n → Node) (r : Fin n → ℝ) (v : Node) : visits (List.ofFn (fun i => (a i, r i))) v = (List.ofFn a).countP (fun w => decide (v <+: w))","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-8e6388f4e812","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1105,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem visits_ofFn {n : ℕ} (a : Fin n → Node) (r : Fin n → ℝ) (v : Node) : visits (List.ofFn (fun i => (a i, r i))) v = (List.ofFn a).countP (fun w => decide (v <+: w))","missing":[],"search":"visits_offn banditrlproof.hoo.visits_offn theorem visits_offn {n : ℕ} (a : fin n → node) (r : fin n → ℝ) (v : node) : visits (list.offn (fun i => (a i, r i))) v = (list.offn a).countp (fun w => decide (v <+: w)) theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.expanded_ofFn","label":"expanded_ofFn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.expanded_ofFn","description":"theorem expanded_ofFn {n : ℕ} (a : Fin n → Node) (r : Fin n → ℝ) : expanded (List.ofFn (fun i => (a i, r i))) = insert [] (List.ofFn a).toFinset","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-e978406f7528","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1106,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expanded_ofFn {n : ℕ} (a : Fin n → Node) (r : Fin n → ℝ) : expanded (List.ofFn (fun i => (a i, r i))) = insert [] (List.ofFn a).toFinset","missing":[],"search":"expanded_offn banditrlproof.hoo.expanded_offn theorem expanded_offn {n : ℕ} (a : fin n → node) (r : fin n → ℝ) : expanded (list.offn (fun i => (a i, r i))) = insert [] (list.offn a).tofinset theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_next_ofFn","label":"measurable_next_ofFn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_next_ofFn","description":"A genuine history-to-action map: discrete past nodes and real past rewards.","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-28274fcea2a0","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1107,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_next_ofFn (ν ρ : ℝ) (n : ℕ) : Measurable (fun p : (Fin n → Node) × (Fin n → ℝ) => next ν ρ (List.ofFn (fun i => (p.1 i, p.2 i))))","missing":[],"search":"measurable_next_offn banditrlproof.hoo.measurable_next_offn a genuine history-to-action map: discrete past nodes and real past rewards. theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_action","label":"measurable_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_action","description":"theorem measurable_action (ν ρ : ℝ) (n : ℕ) : Measurable (fun Y : ℕ → ℝ => action ν ρ Y n)","url":"../modules/banditrlproof-algorithms-hoomeasurable/index.html#decl-633ff21c90c8","parent":"module:BanditRLProof.Algorithms.HOOMeasurable","order":1108,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOMeasurable"],["Source","BanditRLProof/Algorithms/HOOMeasurable.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_action (ν ρ : ℝ) (n : ℕ) : Measurable (fun Y : ℕ → ℝ => action ν ρ Y n)","missing":[],"search":"measurable_action banditrlproof.hoo.measurable_action theorem measurable_action (ν ρ : ℝ) (n : ℕ) : measurable (fun y : ℕ → ℝ => action ν ρ y n) theorem compiled","shard":"modules/2a8467a1ee82da16.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.branch_root_optimistic","label":"branch_root_optimistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.branch_root_optimistic","description":"theorem Covering.branch_root_optimistic {X : Type*} (C : Covering X) (f : X → ℝ) (S : Finset Node) (U : Node → WithTop ℝ) (best : ℝ) (k : ℕ) (hk : C.optimalPath f k ∉ S) (hpre : ∀ j < k, C.optimalPath f j ∈ S) (hU : ∀ j < k, (best : WithTop ℝ) ≤ U (C.optimalPath f j)) : (best : WithTop ℝ) ≤ bValue S U []","url":"../modules/banditrlproof-algorithms-hoopathcomparison/index.html#decl-288837b07f4e","parent":"module:BanditRLProof.Algorithms.HOOPathComparison","order":1109,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPathComparison"],["Source","BanditRLProof/Algorithms/HOOPathComparison.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.branch_root_optimistic {X : Type*} (C : Covering X) (f : X → ℝ) (S : Finset Node) (U : Node → WithTop ℝ) (best : ℝ) (k : ℕ) (hk : C.optimalPath f k ∉ S) (hpre : ∀ j < k, C.optimalPath f j ∈ S) (hU : ∀ j < k, (best : WithTop ℝ) ≤ U (C.optimalPath f j)) : (best : WithTop ℝ) ≤ bValue S U []","missing":[],"search":"branch_root_optimistic banditrlproof.hoo.covering.branch_root_optimistic theorem covering.branch_root_optimistic {x : type*} (c : covering x) (f : x → ℝ) (s : finset node) (u : node → withtop ℝ) (best : ℝ) (k : ℕ) (hk : c.optimalpath f k ∉ s) (hpre : ∀ j < k, c.optimalpath f j ∈ s) (hu : ∀ j < k, (best : withtop ℝ) ≤ u (c.optimalpath f j)) : (best : withtop ℝ) ≤ bvalue s u [] theorem compiled","shard":"modules/b7f75ec18030ba8c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.selected_underestimate_implies_branch_underestimate","label":"selected_underestimate_implies_branch_underestimate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.selected_underestimate_implies_branch_underestimate","description":"If an expanded region lies on the actual selected path but has U below `best`, some node of the comparison branch has U below `best`. The branch need only be inspected through the finite tree's maximum depth.","url":"../modules/banditrlproof-algorithms-hoopathcomparison/index.html#decl-0c6918e8a82c","parent":"module:BanditRLProof.Algorithms.HOOPathComparison","order":1110,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPathComparison"],["Source","BanditRLProof/Algorithms/HOOPathComparison.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.selected_underestimate_implies_branch_underestimate {X : Type*} (C : Covering X) (f : X → ℝ) (S : Finset Node) (U : Node → WithTop ℝ) (best : ℝ) (v : Node) (hv : v ∈ S) (hpath : v <+: select S U) (hU : U v < (best : WithTop ℝ)) : ∃ j ≤ depthBound S, U (C.optimalPath f j) < (best : WithTop ℝ)","missing":[],"search":"selected_underestimate_implies_branch_underestimate banditrlproof.hoo.covering.selected_underestimate_implies_branch_underestimate if an expanded region lies on the actual selected path but has u below `best`, some node of the comparison branch has u below `best`. the branch need only be inspected through the finite tree's maximum depth. theorem compiled","shard":"modules/b7f75ec18030ba8c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.history_selected_underestimate","label":"history_selected_underestimate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.history_selected_underestimate","description":"The same comparison stated directly on the generated chronological state.","url":"../modules/banditrlproof-algorithms-hoopathcomparison/index.html#decl-3971c888f0dc","parent":"module:BanditRLProof.Algorithms.HOOPathComparison","order":1111,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPathComparison"],["Source","BanditRLProof/Algorithms/HOOPathComparison.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.history_selected_underestimate {X : Type*} (C : Covering X) (f : X → ℝ) (ν ρ best : ℝ) (Y : ℕ → ℝ) (n : ℕ) (v : Node) (hv : v ∈ expanded (history ν ρ Y n)) (hpath : v <+: action ν ρ Y n) (hu : upper ν ρ (history ν ρ Y n) v < (best : WithTop ℝ)) : ∃ j ≤ n, upper ν ρ (history ν ρ Y n) (C.optimalPath f j) < (best : WithTop ℝ)","missing":[],"search":"history_selected_underestimate banditrlproof.hoo.covering.history_selected_underestimate the same comparison stated directly on the generated chronological state. theorem compiled","shard":"modules/b7f75ec18030ba8c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.proper_prefix_walk_mem","label":"proper_prefix_walk_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.proper_prefix_walk_mem","description":"theorem proper_prefix_walk_mem (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v w : Node) (hvw : v <+: w) (hw : w <+: walk S B k v) (hne : w ≠ walk S B k v) : w ∈ S","url":"../modules/banditrlproof-algorithms-hooprefix/index.html#decl-7f9bdb91cb13","parent":"module:BanditRLProof.Algorithms.HOOPrefix","order":1112,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPrefix"],["Source","BanditRLProof/Algorithms/HOOPrefix.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem proper_prefix_walk_mem (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v w : Node) (hvw : v <+: w) (hw : w <+: walk S B k v) (hne : w ≠ walk S B k v) : w ∈ S","missing":[],"search":"proper_prefix_walk_mem banditrlproof.hoo.proper_prefix_walk_mem theorem proper_prefix_walk_mem (s : finset node) (b : node → withtop ℝ) (k : ℕ) (v w : node) (hvw : v <+: w) (hw : w <+: walk s b k v) (hne : w ≠ walk s b k v) : w ∈ s theorem compiled","shard":"modules/9c2a059eeb853017.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.proper_prefix_select_mem","label":"proper_prefix_select_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.proper_prefix_select_mem","description":"theorem proper_prefix_select_mem (S : Finset Node) (U : Node → WithTop ℝ) (w : Node) (hw : w <+: select S U) (hne : w ≠ select S U) : w ∈ S","url":"../modules/banditrlproof-algorithms-hooprefix/index.html#decl-dadc23a5b773","parent":"module:BanditRLProof.Algorithms.HOOPrefix","order":1113,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPrefix"],["Source","BanditRLProof/Algorithms/HOOPrefix.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem proper_prefix_select_mem (S : Finset Node) (U : Node → WithTop ℝ) (w : Node) (hw : w <+: select S U) (hne : w ≠ select S U) : w ∈ S","missing":[],"search":"proper_prefix_select_mem banditrlproof.hoo.proper_prefix_select_mem theorem proper_prefix_select_mem (s : finset node) (u : node → withtop ℝ) (w : node) (hw : w <+: select s u) (hne : w ≠ select s u) : w ∈ s theorem compiled","shard":"modules/9c2a059eeb853017.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.expanded_history_prefix_closed","label":"expanded_history_prefix_closed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.expanded_history_prefix_closed","description":"The generated tree is prefix closed; it is not assumed to be a valid tree.","url":"../modules/banditrlproof-algorithms-hooprefix/index.html#decl-cf2257817389","parent":"module:BanditRLProof.Algorithms.HOOPrefix","order":1114,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPrefix"],["Source","BanditRLProof/Algorithms/HOOPrefix.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expanded_history_prefix_closed (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) (w v : Node) (hw : w ∈ expanded (history ν ρ Y n)) (hv : v <+: w) : v ∈ expanded (history ν ρ Y n)","missing":[],"search":"expanded_history_prefix_closed banditrlproof.hoo.expanded_history_prefix_closed the generated tree is prefix closed; it is not assumed to be a valid tree. theorem compiled","shard":"modules/9c2a059eeb853017.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bValue_le_upper","label":"bValue_le_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bValue_le_upper","description":"theorem bValue_le_upper (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) (hv : v ∈ S) : bValue S U v ≤ U v","url":"../modules/banditrlproof-algorithms-hooprefix/index.html#decl-68dca34718cf","parent":"module:BanditRLProof.Algorithms.HOOPrefix","order":1115,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPrefix"],["Source","BanditRLProof/Algorithms/HOOPrefix.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bValue_le_upper (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) (hv : v ∈ S) : bValue S U v ≤ U v","missing":[],"search":"bvalue_le_upper banditrlproof.hoo.bvalue_le_upper theorem bvalue_le_upper (s : finset node) (u : node → withtop ℝ) (v : node) (hv : v ∈ s) : bvalue s u v ≤ u v theorem compiled","shard":"modules/9c2a059eeb853017.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bValue_le_prefix_walk","label":"bValue_le_prefix_walk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bValue_le_prefix_walk","description":"B is nondecreasing at every intermediate node on the actual chosen path.","url":"../modules/banditrlproof-algorithms-hooprefix/index.html#decl-2d0fbabb2471","parent":"module:BanditRLProof.Algorithms.HOOPrefix","order":1116,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPrefix"],["Source","BanditRLProof/Algorithms/HOOPrefix.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bValue_le_prefix_walk (S : Finset Node) (U : Node → WithTop ℝ) (k : ℕ) (v w : Node) (hvw : v <+: w) (hw : w <+: walk S (bValue S U) k v) : bValue S U v ≤ bValue S U w","missing":[],"search":"bvalue_le_prefix_walk banditrlproof.hoo.bvalue_le_prefix_walk b is nondecreasing at every intermediate node on the actual chosen path. theorem compiled","shard":"modules/9c2a059eeb853017.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.visits_pos_mem_expanded","label":"visits_pos_mem_expanded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.visits_pos_mem_expanded","description":"theorem visits_pos_mem_expanded (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) (v : Node) (ht : 0 < visits (history ν ρ Y n) v) : v ∈ expanded (history ν ρ Y n)","url":"../modules/banditrlproof-algorithms-hooprefix/index.html#decl-4eb24465bff5","parent":"module:BanditRLProof.Algorithms.HOOPrefix","order":1117,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOPrefix"],["Source","BanditRLProof/Algorithms/HOOPrefix.lean:80"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem visits_pos_mem_expanded (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) (v : Node) (ht : 0 < visits (history ν ρ Y n) v) : v ∈ expanded (history ν ρ Y n)","missing":[],"search":"visits_pos_mem_expanded banditrlproof.hoo.visits_pos_mem_expanded theorem visits_pos_mem_expanded (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) (v : node) (ht : 0 < visits (history ν ρ y n) v) : v ∈ expanded (history ν ρ y n) theorem compiled","shard":"modules/9c2a059eeb853017.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate","label":"expected_pseudoRegret_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate","description":"Every real exponent strictly above the actual near-optimality dimension admits the source rate for the constructed causal HOO process. The logarithm repair is log(max(N,2)); the confidence and algorithm use the same repair.","url":"../modules/banditrlproof-algorithms-hoorate/index.html#decl-9b9c02a004ea","parent":"module:BanditRLProof.Algorithms.HOORate","order":1118,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORate"],["Source","BanditRLProof/Algorithms/HOORate.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.expected_pseudoRegret_rate {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best d : ℝ) (hmean : ∀x, (∫ y, y ∂law x)=f x) (hf : ∀x, f x≤best) (hfrange : ∀x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd : C.nearOptimalityDimension f best (4*C.nu1/C.nu2) < (d:EReal)) : ∃ γ : ℝ, 0<γ ∧ ∀ N : ℕ, 1≤N → (∫ Y, (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ γ*(N:ℝ)^((d+1)/(d+2))*(Real.log (max (N:ℝ) 2))^(1/(d+2))","missing":[],"search":"expected_pseudoregret_rate banditrlproof.hoo.regularcovering.expected_pseudoregret_rate every real exponent strictly above the actual near-optimality dimension admits the source rate for the constructed causal hoo process. the logarithm repair is log(max(n,2)); the confidence and algorithm use the same repair. theorem compiled","shard":"modules/2218ec96e6a2ed8f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regretSumConstant","label":"regretSumConstant","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regretSumConstant","description":"noncomputable def regretSumConstant (ν₁ ν₂ ρ d K : ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-hooregretalgebra/index.html#decl-7811fe54a472","parent":"module:BanditRLProof.Algorithms.HOORegretAlgebra","order":1119,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOORegretAlgebra"],["Source","BanditRLProof/Algorithms/HOORegretAlgebra.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regretSumConstant (ν₁ ν₂ ρ d K : ℝ) : ℝ","missing":[],"search":"regretsumconstant banditrlproof.hoo.regretsumconstant noncomputable def regretsumconstant (ν₁ ν₂ ρ d k : ℝ) : ℝ definition compiled","shard":"modules/f456b6b4e4f64e38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regretSumConstant_pos","label":"regretSumConstant_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.regretSumConstant_pos","description":"theorem regretSumConstant_pos {ν₁ ν₂ ρ d K : ℝ} (h1 : 0<ν₁) (h2 : 0<ν₂) (hr : 0<ρ) (hK : 0<K) : 0<regretSumConstant ν₁ ν₂ ρ d K","url":"../modules/banditrlproof-algorithms-hooregretalgebra/index.html#decl-8554777dacf2","parent":"module:BanditRLProof.Algorithms.HOORegretAlgebra","order":1120,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretAlgebra"],["Source","BanditRLProof/Algorithms/HOORegretAlgebra.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem regretSumConstant_pos {ν₁ ν₂ ρ d K : ℝ} (h1 : 0<ν₁) (h2 : 0<ν₂) (hr : 0<ρ) (hK : 0<K) : 0<regretSumConstant ν₁ ν₂ ρ d K","missing":[],"search":"regretsumconstant_pos banditrlproof.hoo.regretsumconstant_pos theorem regretsumconstant_pos {ν₁ ν₂ ρ d k : ℝ} (h1 : 0<ν₁) (h2 : 0<ν₂) (hr : 0<ρ) (hk : 0<k) : 0<regretsumconstant ν₁ ν₂ ρ d k theorem compiled","shard":"modules/f456b6b4e4f64e38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.nat_pow_rpow","label":"nat_pow_rpow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.nat_pow_rpow","description":"private theorem nat_pow_rpow (r z : ℝ) (hr : 0≤r) (h : ℕ) : (r^h)^z=(r^z)^h","url":"../modules/banditrlproof-algorithms-hooregretalgebra/index.html#decl-8a460208a4e3","parent":"module:BanditRLProof.Algorithms.HOORegretAlgebra","order":1121,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretAlgebra"],["Source","BanditRLProof/Algorithms/HOORegretAlgebra.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem nat_pow_rpow (r z : ℝ) (hr : 0≤r) (h : ℕ) : (r^h)^z=(r^z)^h","missing":[],"search":"nat_pow_rpow banditrlproof.hoo.nat_pow_rpow private theorem nat_pow_rpow (r z : ℝ) (hr : 0≤r) (h : ℕ) : (r^h)^z=(r^z)^h theorem compiled","shard":"modules/f456b6b4e4f64e38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regret_level_le","label":"regret_level_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.regret_level_le","description":"The single-level algebra keeps the child's depth h+1 in the visit bound, and only then reduces both contributions to one geometric sequence.","url":"../modules/banditrlproof-algorithms-hooregretalgebra/index.html#decl-925a9be37040","parent":"module:BanditRLProof.Algorithms.HOORegretAlgebra","order":1122,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretAlgebra"],["Source","BanditRLProof/Algorithms/HOORegretAlgebra.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem regret_level_le {ν₁ ν₂ ρ d K L : ℝ} (h1 : 0<ν₁) (h2 : 0<ν₂) (hr : 0<ρ) (hr1 : ρ≤1) (hK : 0<K) (hL : Real.log 2≤L) (h : ℕ) : 4*(ν₁*ρ^h)*(K*(ν₂*ρ^h)^(-d)) + 8*(ν₁*ρ^h)*(K*(ν₂*ρ^h)^(-d))*(8*L/(ν₁*ρ^(h+1))^2+4) ≤ regretSumConstant ν₁ ν₂ ρ d K * L * (ρ^(-(1+d)))^h","missing":[],"search":"regret_level_le banditrlproof.hoo.regret_level_le the single-level algebra keeps the child's depth h+1 in the visit bound, and only then reduces both contributions to one geometric sequence. theorem compiled","shard":"modules/f456b6b4e4f64e38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regret_sums_le","label":"regret_sums_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.regret_sums_le","description":"Finite sums are reduced with an explicit environment-only constant.","url":"../modules/banditrlproof-algorithms-hooregretalgebra/index.html#decl-83eade9b7be4","parent":"module:BanditRLProof.Algorithms.HOORegretAlgebra","order":1123,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretAlgebra"],["Source","BanditRLProof/Algorithms/HOORegretAlgebra.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem regret_sums_le {ν₁ ν₂ ρ d K L : ℝ} (h1 : 0<ν₁) (h2 : 0<ν₂) (hr : 0<ρ) (hr1 : ρ<1) (hd : 0<d) (hK : 0<K) (hL : Real.log 2≤L) (H : ℕ) : (∑ h ∈ Finset.range H, 4*(ν₁*ρ^h)*(K*(ν₂*ρ^h)^(-d))) + (∑ h ∈ Finset.range H, 8*(ν₁*ρ^h)*(K*(ν₂*ρ^h)^(-d)) * (8*L/(ν₁*ρ^(h+1))^2+4)) ≤ (regretSumConstant ν₁ ν₂ ρ d K / (ρ^(-(1+d))-1))*L*(ρ^H)^(-(1+d))","missing":[],"search":"regret_sums_le banditrlproof.hoo.regret_sums_le finite sums are reduced with an explicit environment-only constant. theorem compiled","shard":"modules/f456b6b4e4f64e38.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_singleton_count_le_one","label":"action_singleton_count_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_singleton_count_le_one","description":"theorem action_singleton_count_le_one (ν ρ : ℝ) (Y : ℕ → ℝ) (N : ℕ) (p : Node) : (∑ n ∈ Finset.range N, if action ν ρ Y n=p then (1:ℝ) else 0) ≤ 1","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html#decl-daae67eff44c","parent":"module:BanditRLProof.Algorithms.HOORegretPartition","order":1124,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretPartition"],["Source","BanditRLProof/Algorithms/HOORegretPartition.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_singleton_count_le_one (ν ρ : ℝ) (Y : ℕ → ℝ) (N : ℕ) (p : Node) : (∑ n ∈ Finset.range N, if action ν ρ Y n=p then (1:ℝ) else 0) ≤ 1","missing":[],"search":"action_singleton_count_le_one banditrlproof.hoo.action_singleton_count_le_one theorem action_singleton_count_le_one (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) (p : node) : (∑ n ∈ finset.range n, if action ν ρ y n=p then (1:ℝ) else 0) ≤ 1 theorem compiled","shard":"modules/f0009a9444670c3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.deep_regret_le","label":"deep_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.deep_regret_le","description":"theorem RegularCovering.deep_regret_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.deepGoodNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) ≤ 4*(C.nu1*C.rho^H)*N","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html#decl-78551b7b342c","parent":"module:BanditRLProof.Algorithms.HOORegretPartition","order":1125,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretPartition"],["Source","BanditRLProof/Algorithms/HOORegretPartition.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.deep_regret_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.deepGoodNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) ≤ 4*(C.nu1*C.rho^H)*N","missing":[],"search":"deep_regret_le banditrlproof.hoo.regularcovering.deep_regret_le theorem regularcovering.deep_regret_le {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hw : weaklylipschitz f c.ell best) (h n : ℕ) (y : ℕ → ℝ) : (∑ n ∈ finset.range n, if action c.nu1 c.rho y n ∈ c.deepgoodnodes f best h then best-f (c.tocovering.arm c.nu1 c.rho y n) else 0) ≤ 4*(c.nu1*c.rho^h)*n theorem compiled","shard":"modules/f0009a9444670c3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.shallow_regret_le","label":"shallow_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.shallow_regret_le","description":"theorem RegularCovering.shallow_regret_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.shallowGoodNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) ≤ ∑ h ∈ Finset.range H, 4*(C.nu1*C.rho^h)*(C.nearOptimalNodes f best h).card","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html#decl-55541ebb5357","parent":"module:BanditRLProof.Algorithms.HOORegretPartition","order":1126,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretPartition"],["Source","BanditRLProof/Algorithms/HOORegretPartition.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.shallow_regret_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.shallowGoodNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) ≤ ∑ h ∈ Finset.range H, 4*(C.nu1*C.rho^h)*(C.nearOptimalNodes f best h).card","missing":[],"search":"shallow_regret_le banditrlproof.hoo.regularcovering.shallow_regret_le theorem regularcovering.shallow_regret_le {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hw : weaklylipschitz f c.ell best) (h n : ℕ) (y : ℕ → ℝ) : (∑ n ∈ finset.range n, if action c.nu1 c.rho y n ∈ c.shallowgoodnodes f best h then best-f (c.tocovering.arm c.nu1 c.rho y n) else 0) ≤ ∑ h ∈ finset.range h, 4*(c.nu1*c.rho^h)*(c.nearoptimalnodes f best h).card theorem compiled","shard":"modules/f0009a9444670c3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.bad_regret_le","label":"bad_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.bad_regret_le","description":"Poor-subtree contribution on each actual reward path. Its prefix indicators are exactly the history's visits, including unvisited nodes.","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html#decl-a7302d7cb8d2","parent":"module:BanditRLProof.Algorithms.HOORegretPartition","order":1127,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretPartition"],["Source","BanditRLProof/Algorithms/HOORegretPartition.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.bad_regret_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.badSubtreeNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) ≤ ∑ h ∈ Finset.range H, ∑ p ∈ C.boundaryNodes f best h, 4*(C.nu1*C.rho^h)*(visits (history C.nu1 C.rho Y N) p : ℝ)","missing":[],"search":"bad_regret_le banditrlproof.hoo.regularcovering.bad_regret_le poor-subtree contribution on each actual reward path. its prefix indicators are exactly the history's visits, including unvisited nodes. theorem compiled","shard":"modules/f0009a9444670c3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.pathwise_regret_le","label":"pathwise_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.pathwise_regret_le","description":"The actual pathwise three-term regret bound before expectation or depth optimization. No partition, visit count or one-play property is assumed.","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html#decl-ea41f9c9f79c","parent":"module:BanditRLProof.Algorithms.HOORegretPartition","order":1128,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretPartition"],["Source","BanditRLProof/Algorithms/HOORegretPartition.lean:129"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.pathwise_regret_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ≤ 4*(C.nu1*C.rho^H)*N + (∑ h ∈ Finset.range H, 4*(C.nu1*C.rho^h)*(C.nearOptimalNodes f best h).card) + ∑ h ∈ Finset.range H, ∑ p ∈ C.boundaryNodes f best h, 4*(C.nu1*C.rho^h)*(visits (history C.nu1 C.rho Y N) p : ℝ)","missing":[],"search":"pathwise_regret_le banditrlproof.hoo.regularcovering.pathwise_regret_le the actual pathwise three-term regret bound before expectation or depth optimization. no partition, visit count or one-play property is assumed. theorem compiled","shard":"modules/f0009a9444670c3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.boundary_expected_visits","label":"boundary_expected_visits","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.boundary_expected_visits","description":"Source first-step simplification for the actual boundary nodes.","url":"../modules/banditrlproof-algorithms-hooregretpartition/index.html#decl-09fbbb528d46","parent":"module:BanditRLProof.Algorithms.HOORegretPartition","order":1129,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORegretPartition"],["Source","BanditRLProof/Algorithms/HOORegretPartition.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.boundary_expected_visits {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) {h : ℕ} {v : Node} (hv : v ∈ C.boundaryNodes f best h) (N : ℕ) : (∫ Y, (visits (history C.nu1 C.rho Y N) v : ℝ) ∂trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) ≤ 8*Real.log (max (N:ℝ) 2)/(C.nu1*C.rho^(h+1))^2+4","missing":[],"search":"boundary_expected_visits banditrlproof.hoo.regularcovering.boundary_expected_visits source first-step simplification for the actual boundary nodes. theorem compiled","shard":"modules/f0009a9444670c3b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.familyNodeLaw","label":"familyNodeLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.familyNodeLaw","description":"noncomputable def Covering.familyNodeLaw {X : Type*} (C : Covering X) (law : X → Measure ℝ) : Kernel Node ℝ where","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-be09c0600a4d","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1130,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"noncomputable def Covering.familyNodeLaw {X : Type*} (C : Covering X) (law : X → Measure ℝ) : Kernel Node ℝ where","missing":[],"search":"familynodelaw banditrlproof.hoo.covering.familynodelaw noncomputable def covering.familynodelaw {x : type*} (c : covering x) (law : x → measure ℝ) : kernel node ℝ where definition compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.Covering.familyNodeLaw_apply","label":"familyNodeLaw_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.familyNodeLaw_apply","description":"theorem Covering.familyNodeLaw_apply {X : Type*} (C : Covering X) (law : X → Measure ℝ) (v : Node) : C.familyNodeLaw law v = law (C.representative v)","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-2ffd0b0163f9","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1131,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.familyNodeLaw_apply {X : Type*} (C : Covering X) (law : X → Measure ℝ) (v : Node) : C.familyNodeLaw law v = law (C.representative v)","missing":[],"search":"familynodelaw_apply banditrlproof.hoo.covering.familynodelaw_apply theorem covering.familynodelaw_apply {x : type*} (c : covering x) (law : x → measure ℝ) (v : node) : c.familynodelaw law v = law (c.representative v) theorem compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.familyNodeLaw_eq_nodeLaw","label":"familyNodeLaw_eq_nodeLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.familyNodeLaw_eq_nodeLaw","description":"theorem Covering.familyNodeLaw_eq_nodeLaw {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) : C.familyNodeLaw (fun x => law x) = C.nodeLaw law","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-1875a99b2c64","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1132,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.familyNodeLaw_eq_nodeLaw {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) : C.familyNodeLaw (fun x => law x) = C.nodeLaw law","missing":[],"search":"familynodelaw_eq_nodelaw banditrlproof.hoo.covering.familynodelaw_eq_nodelaw theorem covering.familynodelaw_eq_nodelaw {x : type*} [measurablespace x] (c : covering x) (law : kernel x ℝ) : c.familynodelaw (fun x => law x) = c.nodelaw law theorem compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.discrete","label":"discrete","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.discrete","description":"noncomputable def RegularCovering.discrete {X : Type*} [MeasurableSpace X] (C : RegularCovering X) : @RegularCovering X ⊤","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-f6e99d2dadc0","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1133,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def RegularCovering.discrete {X : Type*} [MeasurableSpace X] (C : RegularCovering X) : @RegularCovering X ⊤","missing":[],"search":"discrete banditrlproof.hoo.regularcovering.discrete noncomputable def regularcovering.discrete {x : type*} [measurablespace x] (c : regularcovering x) : @regularcovering x ⊤ definition compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.familyDiscreteKernel","label":"familyDiscreteKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.familyDiscreteKernel","description":"noncomputable def familyDiscreteKernel {X : Type*} (law : X → Measure ℝ) : @Kernel X ℝ ⊤ (inferInstance : MeasurableSpace ℝ)","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-be4f5ceb490f","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1134,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def familyDiscreteKernel {X : Type*} (law : X → Measure ℝ) : @Kernel X ℝ ⊤ (inferInstance : MeasurableSpace ℝ)","missing":[],"search":"familydiscretekernel banditrlproof.hoo.familydiscretekernel noncomputable def familydiscretekernel {x : type*} (law : x → measure ℝ) : @kernel x ℝ ⊤ (inferinstance : measurablespace ℝ) definition compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate_family","label":"expected_pseudoRegret_rate_family","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate_family","description":"theorem RegularCovering.expected_pseudoRegret_rate_family {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : X → Measure ℝ) [∀ x, IsProbabilityMeasure (law x)] (f : X → ℝ) (best d : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hfrange : ∀ x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd :…","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-6a1925709a15","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1135,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"theorem RegularCovering.expected_pseudoRegret_rate_family {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : X → Measure ℝ) [∀ x, IsProbabilityMeasure (law x)] (f : X → ℝ) (best d : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hfrange : ∀ x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd : C.nearOptimalityDimension f best (4*C.nu1/C.nu2) < (d:EReal)) : ∃ γ : ℝ, 0<γ ∧ ∀ N : ℕ, 1≤N → (∫ Y, (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ∂trajectory C.nu1 C.rho (C.toCovering.familyNodeLaw law)) ≤ γ*(N:ℝ)^((d+1)/(d+2))*(Real.log (max (N:ℝ) 2))^(1/(d+2))","missing":[],"search":"expected_pseudoregret_rate_family banditrlproof.hoo.regularcovering.expected_pseudoregret_rate_family theorem regularcovering.expected_pseudoregret_rate_family {x : type*} [measurablespace x] (c : regularcovering x) (law : x → measure ℝ) [∀ x, isprobabilitymeasure (law x)] (f : x → ℝ) (best d : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hfrange : ∀ x, f x ∈ set.icc (0:ℝ) 1) (hbest : regionsup f set.univ=best) (hw : weaklylipschitz f c.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ set.icc (0:ℝ) 1) (hd : c.nearoptimalitydimension f best (4*c.nu1/c.nu2) < (d:ereal)) : ∃ γ : ℝ, 0<γ ∧ ∀ n : ℕ, 1≤n → (∫ y, (∑ n ∈ finset.range n, (best-f (c.tocovering.arm c.nu1 c.rho y n))) ∂trajectory c.nu1 c.rho (c.tocovering.familynodelaw law)) ≤ γ*(n:ℝ)^((d+1)/(d+2))*(real.log (max (n:ℝ) 2))^(1/(d+2)) theorem compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret_family","label":"expected_actual_eq_pseudoRegret_family","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret_family","description":"theorem RegularCovering.expected_actual_eq_pseudoRegret_family {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : X → Measure ℝ) [∀ x, IsProbabilityMeasure (law x)] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hfrange : ∀ x, f x ∈ Set.Icc (0:ℝ) 1) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (N : ℕ) : (∫ Y, (∑ n ∈ Finset.range N, (best-Y n)) ∂trajectory C.nu1 C.rho (C.toCovering.familyN…","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-c8e88254c355","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1136,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"theorem RegularCovering.expected_actual_eq_pseudoRegret_family {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : X → Measure ℝ) [∀ x, IsProbabilityMeasure (law x)] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hfrange : ∀ x, f x ∈ Set.Icc (0:ℝ) 1) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (N : ℕ) : (∫ Y, (∑ n ∈ Finset.range N, (best-Y n)) ∂trajectory C.nu1 C.rho (C.toCovering.familyNodeLaw law)) = (∫ Y, (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) ∂trajectory C.nu1 C.rho (C.toCovering.familyNodeLaw law))","missing":[],"search":"expected_actual_eq_pseudoregret_family banditrlproof.hoo.regularcovering.expected_actual_eq_pseudoregret_family theorem regularcovering.expected_actual_eq_pseudoregret_family {x : type*} [measurablespace x] (c : regularcovering x) (law : x → measure ℝ) [∀ x, isprobabilitymeasure (law x)] (f : x → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hfrange : ∀ x, f x ∈ set.icc (0:ℝ) 1) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ set.icc (0:ℝ) 1) (n : ℕ) : (∫ y, (∑ n ∈ finset.range n, (best-y n)) ∂trajectory c.nu1 c.rho (c.tocovering.familynodelaw law)) = (∫ y, (∑ n ∈ finset.range n, (best-f (c.tocovering.arm c.nu1 c.rho y n))) ∂trajectory c.nu1 c.rho (c.tocovering.familynodelaw law)) theorem compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate_family","label":"expected_actualRegret_rate_family","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate_family","description":"theorem RegularCovering.expected_actualRegret_rate_family {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : X → Measure ℝ) [∀ x, IsProbabilityMeasure (law x)] (f : X → ℝ) (best d : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hfrange : ∀ x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd :…","url":"../modules/banditrlproof-algorithms-hoorewardfamily/index.html#decl-fd0bfbd50f5b","parent":"module:BanditRLProof.Algorithms.HOORewardFamily","order":1137,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOORewardFamily"],["Source","BanditRLProof/Algorithms/HOORewardFamily.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"theorem RegularCovering.expected_actualRegret_rate_family {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : X → Measure ℝ) [∀ x, IsProbabilityMeasure (law x)] (f : X → ℝ) (best d : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hfrange : ∀ x, f x ∈ Set.Icc (0:ℝ) 1) (hbest : regionSup f Set.univ=best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0:ℝ) 1) (hd : C.nearOptimalityDimension f best (4*C.nu1/C.nu2) < (d:EReal)) : ∃ γ : ℝ, 0<γ ∧ ∀ N : ℕ, 1≤N → (∫ Y, (∑ n ∈ Finset.range N, (best-Y n)) ∂trajectory C.nu1 C.rho (C.toCovering.familyNodeLaw law)) ≤ γ*(N:ℝ)^((d+1)/(d+2))*(Real.log (max (N:ℝ) 2))^(1/(d+2))","missing":[],"search":"expected_actualregret_rate_family banditrlproof.hoo.regularcovering.expected_actualregret_rate_family theorem regularcovering.expected_actualregret_rate_family {x : type*} [measurablespace x] (c : regularcovering x) (law : x → measure ℝ) [∀ x, isprobabilitymeasure (law x)] (f : x → ℝ) (best d : ℝ) (hmean : ∀ x, (∫ y, y ∂law x)=f x) (hf : ∀ x, f x≤best) (hfrange : ∀ x, f x ∈ set.icc (0:ℝ) 1) (hbest : regionsup f set.univ=best) (hw : weaklylipschitz f c.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ set.icc (0:ℝ) 1) (hd : c.nearoptimalitydimension f best (4*c.nu1/c.nu2) < (d:ereal)) : ∃ γ : ℝ, 0<γ ∧ ∀ n : ℕ, 1≤n → (∫ y, (∑ n ∈ finset.range n, (best-y n)) ∂trajectory c.nu1 c.rho (c.tocovering.familynodelaw law)) ≤ γ*(n:ℝ)^((d+1)/(d+2))*(real.log (max (n:ℝ) 2))^(1/(d+2)) theorem compiled","shard":"modules/b29d5fe13bc16a56.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.poor_region_selection_tail","label":"poor_region_selection_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.poor_region_selection_tail","description":"theorem RegularCovering.poor_region_selection_tail {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionSup f Set.univ = best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.regi…","url":"../modules/banditrlproof-algorithms-hooselectiontail/index.html#decl-c791c43bd20c","parent":"module:BanditRLProof.Algorithms.HOOSelectionTail","order":1138,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOSelectionTail"],["Source","BanditRLProof/Algorithms/HOOSelectionTail.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.poor_region_selection_tail {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (law : Kernel X ℝ) [IsMarkovKernel law] (f : X → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionSup f Set.univ = best) (hw : WeaklyLipschitz f C.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1) (v : Node) (hv : C.nu1*C.rho^v.length < best-regionSup f (C.region v)) (n : ℕ) : (trajectory C.nu1 C.rho (C.toCovering.nodeLaw law)) {Y | 8*Real.log (max (n:ℝ) 2)/(best-regionSup f (C.region v)-C.nu1*C.rho^v.length)^2 ≤ (visits (history C.nu1 C.rho Y n) v : ℝ) ∧ v <+: action C.nu1 C.rho Y n} ≤ ((n:ENNReal)+2) * ((n:ENNReal) * ENNReal.ofReal (Real.exp (-4*Real.log (max (n:ℝ) 2))))","missing":[],"search":"poor_region_selection_tail banditrlproof.hoo.regularcovering.poor_region_selection_tail theorem regularcovering.poor_region_selection_tail {x : type*} [measurablespace x] (c : regularcovering x) (law : kernel x ℝ) [ismarkovkernel law] (f : x → ℝ) (best : ℝ) (hmean : ∀ x, (∫ y, y ∂law x) = f x) (hf : ∀ x, f x ≤ best) (hbest : regionsup f set.univ = best) (hw : weaklylipschitz f c.ell best) (hbound : ∀ x, ∀ᵐ y ∂law x, y ∈ set.icc (0 : ℝ) 1) (v : node) (hv : c.nu1*c.rho^v.length < best-regionsup f (c.region v)) (n : ℕ) : (trajectory c.nu1 c.rho (c.tocovering.nodelaw law)) {y | 8*real.log (max (n:ℝ) 2)/(best-regionsup f (c.region v)-c.nu1*c.rho^v.length)^2 ≤ (visits (history c.nu1 c.rho y n) v : ℝ) ∧ v <+: action c.nu1 c.rho y n} ≤ ((n:ennreal)+2) * ((n:ennreal) * ennreal.ofreal (real.exp (-4*real.log (max (n:ℝ) 2)))) theorem compiled","shard":"modules/315d6f0fd319eb3f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.prefixExtension","label":"prefixExtension","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.prefixExtension","description":"def prefixExtension (n : ℕ) (h : (i : Finset.Iic n) → ℝ) : ℕ → ℝ","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-55bd168d4ce3","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1139,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def prefixExtension (n : ℕ) (h : (i : Finset.Iic n) → ℝ) : ℕ → ℝ","missing":[],"search":"prefixextension banditrlproof.hoo.prefixextension def prefixextension (n : ℕ) (h : (i : finset.iic n) → ℝ) : ℕ → ℝ definition compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.measurable_prefixExtension","label":"measurable_prefixExtension","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.measurable_prefixExtension","description":"theorem measurable_prefixExtension (n : ℕ) : Measurable (prefixExtension n)","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-308d22fc07ab","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1140,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_prefixExtension (n : ℕ) : Measurable (prefixExtension n)","missing":[],"search":"measurable_prefixextension banditrlproof.hoo.measurable_prefixextension theorem measurable_prefixextension (n : ℕ) : measurable (prefixextension n) theorem compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.action_prefixExtension","label":"action_prefixExtension","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.action_prefixExtension","description":"theorem action_prefixExtension (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : action ν ρ (prefixExtension n (Preorder.frestrictLe n Y)) (n+1) = action ν ρ Y (n+1)","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-7fc0acd1d1ac","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1141,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_prefixExtension (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : action ν ρ (prefixExtension n (Preorder.frestrictLe n Y)) (n+1) = action ν ρ Y (n+1)","missing":[],"search":"action_prefixextension banditrlproof.hoo.action_prefixextension theorem action_prefixextension (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : action ν ρ (prefixextension n (preorder.frestrictle n y)) (n+1) = action ν ρ y (n+1) theorem compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.stepKernel","label":"stepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.stepKernel","description":"noncomputable def stepKernel (ν ρ : ℝ) (law : Kernel Node ℝ) (n : ℕ) : Kernel ((i : Finset.Iic n) → ℝ) ℝ","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-0628129096d2","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1142,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def stepKernel (ν ρ : ℝ) (law : Kernel Node ℝ) (n : ℕ) : Kernel ((i : Finset.Iic n) → ℝ) ℝ","missing":[],"search":"stepkernel banditrlproof.hoo.stepkernel noncomputable def stepkernel (ν ρ : ℝ) (law : kernel node ℝ) (n : ℕ) : kernel ((i : finset.iic n) → ℝ) ℝ definition compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory","label":"trajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory","description":"noncomputable def trajectory (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] : Measure (ℕ → ℝ)","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-2b355e78be9c","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1143,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"noncomputable def trajectory (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] : Measure (ℕ → ℝ)","missing":[],"search":"trajectory banditrlproof.hoo.trajectory noncomputable def trajectory (ν ρ : ℝ) (law : kernel node ℝ) [ismarkovkernel law] : measure (ℕ → ℝ) definition compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.stepKernel_apply_prefix","label":"stepKernel_apply_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.stepKernel_apply_prefix","description":"theorem stepKernel_apply_prefix (ν ρ : ℝ) (law : Kernel Node ℝ) (Y : ℕ → ℝ) (n : ℕ) : stepKernel ν ρ law n (Preorder.frestrictLe n Y) = law (action ν ρ Y (n+1))","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-8e20f0fccb1b","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1144,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stepKernel_apply_prefix (ν ρ : ℝ) (law : Kernel Node ℝ) (Y : ℕ → ℝ) (n : ℕ) : stepKernel ν ρ law n (Preorder.frestrictLe n Y) = law (action ν ρ Y (n+1))","missing":[],"search":"stepkernel_apply_prefix banditrlproof.hoo.stepkernel_apply_prefix theorem stepkernel_apply_prefix (ν ρ : ℝ) (law : kernel node ℝ) (y : ℕ → ℝ) (n : ℕ) : stepkernel ν ρ law n (preorder.frestrictle n y) = law (action ν ρ y (n+1)) theorem compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory_condDistrib","label":"trajectory_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory_condDistrib","description":"Conditional reward distribution is produced by the constructed trajectory, not supplied as a confidence or stochastic-process oracle.","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-b5c6283a7083","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1145,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trajectory_condDistrib (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (n : ℕ) : condDistrib (fun Y : ℕ → ℝ => Y (n+1)) (Preorder.frestrictLe n) (trajectory ν ρ law) =ᵐ[(trajectory ν ρ law).map (Preorder.frestrictLe n)] stepKernel ν ρ law n","missing":[],"search":"trajectory_conddistrib banditrlproof.hoo.trajectory_conddistrib conditional reward distribution is produced by the constructed trajectory, not supplied as a confidence or stochastic-process oracle. theorem compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory_prefix_compProd","label":"trajectory_prefix_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory_prefix_compProd","description":"The joint prefix/next-reward law, useful without choosing a conditional expectation version. It keeps the actual history-dependent kernel.","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-b545cbe09cca","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1146,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trajectory_prefix_compProd (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] (n : ℕ) : (trajectory ν ρ law).map (Preorder.frestrictLe n) ⊗ₘ stepKernel ν ρ law n = (trajectory ν ρ law).map (fun Y => (Preorder.frestrictLe n Y, Y (n+1)))","missing":[],"search":"trajectory_prefix_compprod banditrlproof.hoo.trajectory_prefix_compprod the joint prefix/next-reward law, useful without choosing a conditional expectation version. it keeps the actual history-dependent kernel. theorem compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.trajectory_initial_law","label":"trajectory_initial_law","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.trajectory_initial_law","description":"The chronological trajectory starts with exactly the first selected node's reward law. No artificial reward is inserted at index zero.","url":"../modules/banditrlproof-algorithms-hootrajectory/index.html#decl-5ed539b61284","parent":"module:BanditRLProof.Algorithms.HOOTrajectory","order":1147,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTrajectory"],["Source","BanditRLProof/Algorithms/HOOTrajectory.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trajectory_initial_law (ν ρ : ℝ) (law : Kernel Node ℝ) [IsMarkovKernel law] : (trajectory ν ρ law).map (fun Y => Y 0) = law (action ν ρ (fun _ => 0) 0)","missing":[],"search":"trajectory_initial_law banditrlproof.hoo.trajectory_initial_law the chronological trajectory starts with exactly the first selected node's reward law. no artificial reward is inserted at index zero. theorem compiled","shard":"modules/a41329574d895035.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Node","label":"Node","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.HOO.Node","description":"abbrev Node","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-062df4619fd3","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1148,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Node","missing":[],"search":"node banditrlproof.hoo.node abbrev node abbreviation compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.child","label":"child","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.child","description":"def child (v : Node) (b : Bool) : Node","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-3a19a3077729","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1149,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def child (v : Node) (b : Bool) : Node","missing":[],"search":"child banditrlproof.hoo.child def child (v : node) (b : bool) : node definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.child_length","label":"child_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.child_length","description":"@[simp] theorem child_length (v : Node) (b : Bool) : (child v b).length = v.length + 1","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-6ca8b83ad8cd","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1150,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem child_length (v : Node) (b : Bool) : (child v b).length = v.length + 1","missing":[],"search":"child_length banditrlproof.hoo.child_length @[simp] theorem child_length (v : node) (b : bool) : (child v b).length = v.length + 1 theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.prefix_child","label":"prefix_child","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.prefix_child","description":"theorem prefix_child (v : Node) (b : Bool) : v <+: child v b","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-276c51b14faa","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1151,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem prefix_child (v : Node) (b : Bool) : v <+: child v b","missing":[],"search":"prefix_child banditrlproof.hoo.prefix_child theorem prefix_child (v : node) (b : bool) : v <+: child v b theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.depthBound","label":"depthBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.depthBound","description":"def depthBound (S : Finset Node) : ℕ","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-e183d11edded","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1152,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def depthBound (S : Finset Node) : ℕ","missing":[],"search":"depthbound banditrlproof.hoo.depthbound def depthbound (s : finset node) : ℕ definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.length_le_depthBound","label":"length_le_depthBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.length_le_depthBound","description":"theorem length_le_depthBound {S : Finset Node} {v : Node} (hv : v ∈ S) : v.length ≤ depthBound S","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-c4698179b49a","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1153,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem length_le_depthBound {S : Finset Node} {v : Node} (hv : v ∈ S) : v.length ≤ depthBound S","missing":[],"search":"length_le_depthbound banditrlproof.hoo.length_le_depthbound theorem length_le_depthbound {s : finset node} {v : node} (hv : v ∈ s) : v.length ≤ depthbound s theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.not_mem_of_depth_lt","label":"not_mem_of_depth_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.not_mem_of_depth_lt","description":"theorem not_mem_of_depth_lt {S : Finset Node} {v : Node} (hv : depthBound S < v.length) : v ∉ S","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-896577235175","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1154,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem not_mem_of_depth_lt {S : Finset Node} {v : Node} (hv : depthBound S < v.length) : v ∉ S","missing":[],"search":"not_mem_of_depth_lt banditrlproof.hoo.not_mem_of_depth_lt theorem not_mem_of_depth_lt {s : finset node} {v : node} (hv : depthbound s < v.length) : v ∉ s theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.backward","label":"backward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.backward","description":"Unexpanded nodes have infinity; expanded nodes use the source min/max rule.","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-0a2a9648c44d","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1155,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def backward (S : Finset Node) (U : Node → WithTop ℝ) : ℕ → Node → WithTop ℝ | 0, _ => ⊤ | k+1, v => if v ∈ S then min (U v) (max (backward S U k (child v false)) (backward S U k (child v true))) else ⊤ theorem backward_not_mem (S : Finset Node) (U : Node → WithTop ℝ) (k : ℕ) (v : Node) (hv : v ∉ S) : backward S U k v = ⊤","missing":[],"search":"backward banditrlproof.hoo.backward unexpanded nodes have infinity; expanded nodes use the source min/max rule. definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.backward_not_mem","label":"backward_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.backward_not_mem","description":"theorem backward_not_mem (S : Finset Node) (U : Node → WithTop ℝ) (k : ℕ) (v : Node) (hv : v ∉ S) : backward S U k v = ⊤","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-c37ba7f71d39","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1156,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem backward_not_mem (S : Finset Node) (U : Node → WithTop ℝ) (k : ℕ) (v : Node) (hv : v ∉ S) : backward S U k v = ⊤","missing":[],"search":"backward_not_mem banditrlproof.hoo.backward_not_mem theorem backward_not_mem (s : finset node) (u : node → withtop ℝ) (k : ℕ) (v : node) (hv : v ∉ s) : backward s u k v = ⊤ theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.backward_stable_succ","label":"backward_stable_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.backward_stable_succ","description":"Increasing already sufficient fuel has no effect on B.","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-1c019eebe359","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1157,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem backward_stable_succ (S : Finset Node) (U : Node → WithTop ℝ) (k : ℕ) (v : Node) (hk : depthBound S < v.length + k) : backward S U (k+1) v = backward S U k v","missing":[],"search":"backward_stable_succ banditrlproof.hoo.backward_stable_succ increasing already sufficient fuel has no effect on b. theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bValue","label":"bValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.bValue","description":"noncomputable def bValue (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) : WithTop ℝ","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-0c4c74e257bb","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1158,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bValue (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) : WithTop ℝ","missing":[],"search":"bvalue banditrlproof.hoo.bvalue noncomputable def bvalue (s : finset node) (u : node → withtop ℝ) (v : node) : withtop ℝ definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bValue_not_mem","label":"bValue_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bValue_not_mem","description":"theorem bValue_not_mem (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) (hv : v ∉ S) : bValue S U v = ⊤","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-e56d63073aef","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1159,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bValue_not_mem (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) (hv : v ∉ S) : bValue S U v = ⊤","missing":[],"search":"bvalue_not_mem banditrlproof.hoo.bvalue_not_mem theorem bvalue_not_mem (s : finset node) (u : node → withtop ℝ) (v : node) (hv : v ∉ s) : bvalue s u v = ⊤ theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bValue_eq","label":"bValue_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bValue_eq","description":"The actual finite computation satisfies the exact recursive B equation.","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-2f9800d368d9","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1160,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bValue_eq (S : Finset Node) (U : Node → WithTop ℝ) (v : Node) (hv : v ∈ S) : bValue S U v = min (U v) (max (bValue S U (child v false)) (bValue S U (child v true)))","missing":[],"search":"bvalue_eq banditrlproof.hoo.bvalue_eq the actual finite computation satisfies the exact recursive b equation. theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.preferred","label":"preferred","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.preferred","description":"noncomputable def preferred (B : Node → WithTop ℝ) (v : Node) : Bool","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-e2d7178c32bb","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1161,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def preferred (B : Node → WithTop ℝ) (v : Node) : Bool","missing":[],"search":"preferred banditrlproof.hoo.preferred noncomputable def preferred (b : node → withtop ℝ) (v : node) : bool definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.preferred_max","label":"preferred_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.preferred_max","description":"theorem preferred_max (B : Node → WithTop ℝ) (v : Node) : B (child v (preferred B v)) = max (B (child v false)) (B (child v true))","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-00dd1e263cfd","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1162,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem preferred_max (B : Node → WithTop ℝ) (v : Node) : B (child v (preferred B v)) = max (B (child v false)) (B (child v true))","missing":[],"search":"preferred_max banditrlproof.hoo.preferred_max theorem preferred_max (b : node → withtop ℝ) (v : node) : b (child v (preferred b v)) = max (b (child v false)) (b (child v true)) theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.walk","label":"walk","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.walk","description":"noncomputable def walk (S : Finset Node) (B : Node → WithTop ℝ) : ℕ → Node → Node | 0, v => v | k+1, v => if v ∈ S then walk S B k (child v (preferred B v)) else v theorem walk_not_mem (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) (hk : depthBound S < v.length + k) : walk S B k v ∉ S","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-a922920b10e4","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1163,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def walk (S : Finset Node) (B : Node → WithTop ℝ) : ℕ → Node → Node | 0, v => v | k+1, v => if v ∈ S then walk S B k (child v (preferred B v)) else v theorem walk_not_mem (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) (hk : depthBound S < v.length + k) : walk S B k v ∉ S","missing":[],"search":"walk banditrlproof.hoo.walk noncomputable def walk (s : finset node) (b : node → withtop ℝ) : ℕ → node → node | 0, v => v | k+1, v => if v ∈ s then walk s b k (child v (preferred b v)) else v theorem walk_not_mem (s : finset node) (b : node → withtop ℝ) (k : ℕ) (v : node) (hk : depthbound s < v.length + k) : walk s b k v ∉ s definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.walk_not_mem","label":"walk_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.walk_not_mem","description":"theorem walk_not_mem (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) (hk : depthBound S < v.length + k) : walk S B k v ∉ S","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-67c7ce506c93","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1164,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem walk_not_mem (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) (hk : depthBound S < v.length + k) : walk S B k v ∉ S","missing":[],"search":"walk_not_mem banditrlproof.hoo.walk_not_mem theorem walk_not_mem (s : finset node) (b : node → withtop ℝ) (k : ℕ) (v : node) (hk : depthbound s < v.length + k) : walk s b k v ∉ s theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.prefix_walk","label":"prefix_walk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.prefix_walk","description":"theorem prefix_walk (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) : v <+: walk S B k v","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-028f29d35383","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1165,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem prefix_walk (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) : v <+: walk S B k v","missing":[],"search":"prefix_walk banditrlproof.hoo.prefix_walk theorem prefix_walk (s : finset node) (b : node → withtop ℝ) (k : ℕ) (v : node) : v <+: walk s b k v theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.walk_length_le","label":"walk_length_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.walk_length_le","description":"theorem walk_length_le (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) : (walk S B k v).length ≤ v.length + k","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-47f653a762f0","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1166,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem walk_length_le (S : Finset Node) (B : Node → WithTop ℝ) (k : ℕ) (v : Node) : (walk S B k v).length ≤ v.length + k","missing":[],"search":"walk_length_le banditrlproof.hoo.walk_length_le theorem walk_length_le (s : finset node) (b : node → withtop ℝ) (k : ℕ) (v : node) : (walk s b k v).length ≤ v.length + k theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.select","label":"select","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.select","description":"noncomputable def select (S : Finset Node) (U : Node → WithTop ℝ) : Node","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-da2df362c303","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1167,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def select (S : Finset Node) (U : Node → WithTop ℝ) : Node","missing":[],"search":"select banditrlproof.hoo.select noncomputable def select (s : finset node) (u : node → withtop ℝ) : node definition compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.select_not_mem","label":"select_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.select_not_mem","description":"theorem select_not_mem (S : Finset Node) (U : Node → WithTop ℝ) : select S U ∉ S","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-55f0253c9763","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1168,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem select_not_mem (S : Finset Node) (U : Node → WithTop ℝ) : select S U ∉ S","missing":[],"search":"select_not_mem banditrlproof.hoo.select_not_mem theorem select_not_mem (s : finset node) (u : node → withtop ℝ) : select s u ∉ s theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.select_length_le","label":"select_length_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.select_length_le","description":"theorem select_length_le (S : Finset Node) (U : Node → WithTop ℝ) : (select S U).length ≤ depthBound S + 1","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-056010781fa0","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1169,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem select_length_le (S : Finset Node) (U : Node → WithTop ℝ) : (select S U).length ≤ depthBound S + 1","missing":[],"search":"select_length_le banditrlproof.hoo.select_length_le theorem select_length_le (s : finset node) (u : node → withtop ℝ) : (select s u).length ≤ depthbound s + 1 theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.bValue_le_walk","label":"bValue_le_walk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.bValue_le_walk","description":"The recursive optimistic value cannot decrease along the selected path.","url":"../modules/banditrlproof-algorithms-hootree/index.html#decl-6850af3e9dec","parent":"module:BanditRLProof.Algorithms.HOOTree","order":1170,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HOOTree"],["Source","BanditRLProof/Algorithms/HOOTree.lean:138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bValue_le_walk (S : Finset Node) (U : Node → WithTop ℝ) (k : ℕ) (v : Node) : bValue S U v ≤ bValue S U (walk S (bValue S U) k v)","missing":[],"search":"bvalue_le_walk banditrlproof.hoo.bvalue_le_walk the recursive optimistic value cannot decrease along the selected path. theorem compiled","shard":"modules/07448909ff0903c4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustMean","label":"robustMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustMean","description":"noncomputable def robustMean (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-3a75bded0237","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1171,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def robustMean (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : ℝ","missing":[],"search":"robustmean banditrlproof.heavytail.robustmean noncomputable def robustmean (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (t : ℕ) : ℝ definition compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustMean_latent","label":"robustMean_latent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustMean_latent","description":"theorem robustMean_latent (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : robustMean hK ε u stream arm t = (∑ s ∈ Finset.range (pullCount (robustAction hK ε u stream) arm t), truncate (sampleThreshold ε u t s) (stream s arm)) / pullCount (robustAction hK ε u stream) arm t","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-cd043a49bd4b","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1172,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_latent (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : robustMean hK ε u stream arm t = (∑ s ∈ Finset.range (pullCount (robustAction hK ε u stream) arm t), truncate (sampleThreshold ε u t s) (stream s arm)) / pullCount (robustAction hK ε u stream) arm t","missing":[],"search":"robustmean_latent banditrlproof.heavytail.robustmean_latent theorem robustmean_latent (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (t : ℕ) : robustmean hk ε u stream arm t = (∑ s ∈ finset.range (pullcount (robustaction hk ε u stream) arm t), truncate (samplethreshold ε u t s) (stream s arm)) / pullcount (robustaction hk ε u stream) arm t theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_pullCount_pos","label":"robust_pullCount_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_pullCount_pos","description":"theorem robust_pullCount_pos (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) (ht : K ≤ t) : 0 < pullCount (robustAction hK ε u stream) arm t","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-4cf55c88f35a","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1173,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_pullCount_pos (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) (ht : K ≤ t) : 0 < pullCount (robustAction hK ε u stream) arm t","missing":[],"search":"robust_pullcount_pos banditrlproof.heavytail.robust_pullcount_pos theorem robust_pullcount_pos (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (t : ℕ) (ht : k ≤ t) : 0 < pullcount (robustaction hk ε u stream) arm t theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustMean_history","label":"robustMean_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustMean_history","description":"theorem robustMean_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyTruncatedMean (UCB.initializationArm hK 0) (sampleThreshold ε u (n+1)) n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1)","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-8963d06b5d9a","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1174,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyTruncatedMean (UCB.initializationArm hK 0) (sampleThreshold ε u (n+1)) n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1)","missing":[],"search":"robustmean_history banditrlproof.heavytail.robustmean_history theorem robustmean_history (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (n : ℕ) : historytruncatedmean (ucb.initializationarm hk 0) (samplethreshold ε u (n+1)) n (armstreampolicy.history (ucb.initializationarm hk 0) (nextarm hk ε u) stream n) arm = robustmean hk ε u stream arm (n+1) theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustIndex_history","label":"robustIndex_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustIndex_history","description":"theorem robustIndex_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyIndex (UCB.initializationArm hK 0) ε u n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1) + confidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) arm (n+1))","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-a8e28622b45d","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1175,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustIndex_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyIndex (UCB.initializationArm hK 0) ε u n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1) + confidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) arm (n+1))","missing":[],"search":"robustindex_history banditrlproof.heavytail.robustindex_history theorem robustindex_history (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (n : ℕ) : historyindex (ucb.initializationarm hk 0) ε u n (armstreampolicy.history (ucb.initializationarm hk 0) (nextarm hk ε u) stream n) arm = robustmean hk ε u stream arm (n+1) + confidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) arm (n+1)) theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustMean_tail","label":"robustMean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustMean_tail","description":"theorem robustMean_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | confidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ |robustMean…","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-34673517f264","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1176,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | confidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ |robustMean hK ε u stream arm t - ∫ x, x ∂ν arm|} ≤ t * (2 * Real.exp (-confidenceLog t))","missing":[],"search":"robustmean_tail banditrlproof.heavytail.robustmean_tail theorem robustmean_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (ε u : ℝ) (t : ℕ) (ht : k ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : integrable (fun x : ℝ => x) (ν arm)) (hm : integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (ucb.armstreammeasure ν).real {stream | confidenceradius ε u t (pullcount (robustaction hk ε u stream) arm t) ≤ |robustmean hk ε u stream arm t - ∫ x, x ∂ν arm|} ≤ t * (2 * real.exp (-confidencelog t)) theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_selected_gap_le","label":"robust_selected_gap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_selected_gap_le","description":"theorem robust_selected_gap_le (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (mean : Fin K → ℝ) (best : Fin K) (n : ℕ) (hn : K ≤ n+1) (hbest : |robustMean hK ε u stream best (n+1) - mean best| ≤ confidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) best (n+1))) (hchosen : |robustMean hK ε u stream (robustAction hK ε u stream (n+1)) (n+1) - mean (robustAction hK ε u stream (n+1))| ≤ confidenceR…","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-c4f2976e03f7","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1177,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_selected_gap_le (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (mean : Fin K → ℝ) (best : Fin K) (n : ℕ) (hn : K ≤ n+1) (hbest : |robustMean hK ε u stream best (n+1) - mean best| ≤ confidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) best (n+1))) (hchosen : |robustMean hK ε u stream (robustAction hK ε u stream (n+1)) (n+1) - mean (robustAction hK ε u stream (n+1))| ≤ confidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) (robustAction hK ε u stream (n+1)) (n+1))) : mean best - mean (robustAction hK ε u stream (n+1)) ≤ 2 * confidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) (robustAction hK ε u stream (n+1)) (n+1))","missing":[],"search":"robust_selected_gap_le banditrlproof.heavytail.robust_selected_gap_le theorem robust_selected_gap_le (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (mean : fin k → ℝ) (best : fin k) (n : ℕ) (hn : k ≤ n+1) (hbest : |robustmean hk ε u stream best (n+1) - mean best| ≤ confidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) best (n+1))) (hchosen : |robustmean hk ε u stream (robustaction hk ε u stream (n+1)) (n+1) - mean (robustaction hk ε u stream (n+1))| ≤ confidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) (robustaction hk ε u stream (n+1)) (n+1))) : mean best - mean (robustaction hk ε u stream (n+1)) ≤ 2 * confidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) (robustaction hk ε u stream (n+1)) (n+1)) theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_selected_small_radius_tail","label":"robust_selected_small_radius_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_selected_small_radius_tail","description":"theorem robust_selected_small_radius_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction hK ε u stream t = arm ∧ 2 * confidence…","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-b785a8ab10ab","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1178,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_selected_small_radius_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction hK ε u stream t = arm ∧ 2 * confidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm} ≤ 4*t*Real.exp (-confidenceLog t)","missing":[],"search":"robust_selected_small_radius_tail banditrlproof.heavytail.robust_selected_small_radius_tail theorem robust_selected_small_radius_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t : ℕ) (ht : k ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (ucb.armstreammeasure ν).real {stream | robustaction hk ε u stream t = arm ∧ 2 * confidenceradius ε u t (pullcount (robustaction hk ε u stream) arm t) < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm} ≤ 4*t*real.exp (-confidencelog t) theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_initial_count_zero","label":"robust_initial_count_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_initial_count_zero","description":"theorem robust_initial_count_zero (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : pullCount (robustAction hK ε u stream) (robustAction hK ε u stream t) t = 0","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-f4b33e68a2e0","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1179,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:129"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_initial_count_zero (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : pullCount (robustAction hK ε u stream) (robustAction hK ε u stream t) t = 0","missing":[],"search":"robust_initial_count_zero banditrlproof.heavytail.robust_initial_count_zero theorem robust_initial_count_zero (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (t : ℕ) (ht : t < k) : pullcount (robustaction hk ε u stream) (robustaction hk ε u stream t) t = 0 theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_large_count_tail","label":"robust_large_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_large_count_tail","description":"theorem robust_large_count_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T t : ℕ) (ht : t ≤ T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction…","url":"../modules/banditrlproof-algorithms-heavytailadaptive/index.html#decl-1a02e0a80df1","parent":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","order":1180,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailAdaptive.lean:140"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_large_count_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T t : ℕ) (ht : t ≤ T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction hK ε u stream t = arm ∧ gapThreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T ≤ pullCount (robustAction hK ε u stream) arm t} ≤ 4*t*Real.exp (-confidenceLog t)","missing":[],"search":"robust_large_count_tail banditrlproof.heavytail.robust_large_count_tail theorem robust_large_count_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t t : ℕ) (ht : t ≤ t) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (ucb.armstreammeasure ν).real {stream | robustaction hk ε u stream t = arm ∧ gapthreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) t ≤ pullcount (robustaction hk ε u stream) arm t} ≤ 4*t*real.exp (-confidencelog t) theorem compiled","shard":"modules/1900adfcc28d5961.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.lintegral_pullCount_threshold","label":"lintegral_pullCount_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.lintegral_pullCount_threshold","description":"theorem lintegral_pullCount_threshold {Ω : Type} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (a : Ω → ActionTrace (Fin K)) (ha : ∀ t, Measurable (fun ω => a ω t)) (arm : Fin K) (T B : ℕ) : (∫⁻ ω, (pullCount (a ω) arm T : ℝ≥0∞) ∂μ) ≤ B + ∑ t ∈ Finset.range T, μ {ω | a ω t = arm ∧ B ≤ pullCount (a ω) arm t}","url":"../modules/banditrlproof-algorithms-heavytailexpectedcount/index.html#decl-5f4a368d6b72","parent":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","order":1181,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailExpectedCount.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_pullCount_threshold {Ω : Type} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (a : Ω → ActionTrace (Fin K)) (ha : ∀ t, Measurable (fun ω => a ω t)) (arm : Fin K) (T B : ℕ) : (∫⁻ ω, (pullCount (a ω) arm T : ℝ≥0∞) ∂μ) ≤ B + ∑ t ∈ Finset.range T, μ {ω | a ω t = arm ∧ B ≤ pullCount (a ω) arm t}","missing":[],"search":"lintegral_pullcount_threshold banditrlproof.heavytail.lintegral_pullcount_threshold theorem lintegral_pullcount_threshold {ω : type} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (a : ω → actiontrace (fin k)) (ha : ∀ t, measurable (fun ω => a ω t)) (arm : fin k) (t b : ℕ) : (∫⁻ ω, (pullcount (a ω) arm t : ℝ≥0∞) ∂μ) ≤ b + ∑ t ∈ finset.range t, μ {ω | a ω t = arm ∧ b ≤ pullcount (a ω) arm t} theorem compiled","shard":"modules/952a4e1f3873a4d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_lintegral_count_le","label":"robust_lintegral_count_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_lintegral_count_le","description":"theorem robust_lintegral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫⁻ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ≥0∞)…","url":"../modules/banditrlproof-algorithms-heavytailexpectedcount/index.html#decl-84a8f8dadf5c","parent":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","order":1182,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailExpectedCount.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_lintegral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫⁻ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ≥0∞) ∂UCB.armStreamMeasure ν) ≤ gapThreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T + 2","missing":[],"search":"robust_lintegral_count_le banditrlproof.heavytail.robust_lintegral_count_le theorem robust_lintegral_count_le (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫⁻ stream, (pullcount (robustaction hk ε u stream) arm t : ℝ≥0∞) ∂ucb.armstreammeasure ν) ≤ gapthreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) t + 2 theorem compiled","shard":"modules/952a4e1f3873a4d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_integrable_count","label":"robust_integrable_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_integrable_count","description":"theorem robust_integrable_count (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (T : ℕ) : Integrable (fun stream => (pullCount (robustAction hK ε u stream) arm T : ℝ)) (UCB.armStreamMeasure ν)","url":"../modules/banditrlproof-algorithms-heavytailexpectedcount/index.html#decl-5a1460d8e09d","parent":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","order":1183,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailExpectedCount.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_integrable_count (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (T : ℕ) : Integrable (fun stream => (pullCount (robustAction hK ε u stream) arm T : ℝ)) (UCB.armStreamMeasure ν)","missing":[],"search":"robust_integrable_count banditrlproof.heavytail.robust_integrable_count theorem robust_integrable_count (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (ε u : ℝ) (t : ℕ) : integrable (fun stream => (pullcount (robustaction hk ε u stream) arm t : ℝ)) (ucb.armstreammeasure ν) theorem compiled","shard":"modules/952a4e1f3873a4d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_integral_count_le","label":"robust_integral_count_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_integral_count_le","description":"theorem robust_integral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ) ∂UCB.…","url":"../modules/banditrlproof-algorithms-heavytailexpectedcount/index.html#decl-d7fe921d9e9a","parent":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","order":1184,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailExpectedCount.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_integral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ) ∂UCB.armStreamMeasure ν) ≤ gapThreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T + 2","missing":[],"search":"robust_integral_count_le banditrlproof.heavytail.robust_integral_count_le theorem robust_integral_count_le (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullcount (robustaction hk ε u stream) arm t : ℝ) ∂ucb.armstreammeasure ν) ≤ gapthreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) t + 2 theorem compiled","shard":"modules/952a4e1f3873a4d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.historyAction","label":"historyAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.historyAction","description":"def historyAction (initial : Fin K) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : ActionTrace (Fin K)","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-8f2c22854d53","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1185,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def historyAction (initial : Fin K) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : ActionTrace (Fin K)","missing":[],"search":"historyaction banditrlproof.heavytail.historyaction def historyaction (initial : fin k) (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) : actiontrace (fin k) definition compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.historyReward","label":"historyReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.historyReward","description":"def historyReward (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : RewardTrace ℝ","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-eb009b822124","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1186,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def historyReward (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : RewardTrace ℝ","missing":[],"search":"historyreward banditrlproof.heavytail.historyreward def historyreward (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) : rewardtrace ℝ definition compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_historyAction","label":"measurable_historyAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_historyAction","description":"theorem measurable_historyAction (initial : Fin K) (n t : ℕ) : Measurable (fun h => historyAction initial n h t)","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-a460112a2b14","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1187,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyAction (initial : Fin K) (n t : ℕ) : Measurable (fun h => historyAction initial n h t)","missing":[],"search":"measurable_historyaction banditrlproof.heavytail.measurable_historyaction theorem measurable_historyaction (initial : fin k) (n t : ℕ) : measurable (fun h => historyaction initial n h t) theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_historyReward","label":"measurable_historyReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_historyReward","description":"theorem measurable_historyReward (n t : ℕ) : Measurable (fun h : History.FinitePairHistory (Fin K) ℝ n => historyReward n h t)","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-fa220079fb42","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1188,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyReward (n t : ℕ) : Measurable (fun h : History.FinitePairHistory (Fin K) ℝ n => historyReward n h t)","missing":[],"search":"measurable_historyreward banditrlproof.heavytail.measurable_historyreward theorem measurable_historyreward (n t : ℕ) : measurable (fun h : history.finitepairhistory (fin k) ℝ n => historyreward n h t) theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.historyTruncatedMean","label":"historyTruncatedMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.historyTruncatedMean","description":"noncomputable def historyTruncatedMean (initial : Fin K) (B : ℕ → ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) (arm : Fin K) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-d703e0d7483f","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1189,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def historyTruncatedMean (initial : Fin K) (B : ℕ → ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) (arm : Fin K) : ℝ","missing":[],"search":"historytruncatedmean banditrlproof.heavytail.historytruncatedmean noncomputable def historytruncatedmean (initial : fin k) (b : ℕ → ℝ) (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) (arm : fin k) : ℝ definition compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_historyTruncatedMean","label":"measurable_historyTruncatedMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_historyTruncatedMean","description":"theorem measurable_historyTruncatedMean (initial : Fin K) (B : ℕ → ℝ) (n : ℕ) (arm : Fin K) : Measurable (fun h => historyTruncatedMean initial B n h arm)","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-01c093af7da4","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1190,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyTruncatedMean (initial : Fin K) (B : ℕ → ℝ) (n : ℕ) (arm : Fin K) : Measurable (fun h => historyTruncatedMean initial B n h arm)","missing":[],"search":"measurable_historytruncatedmean banditrlproof.heavytail.measurable_historytruncatedmean theorem measurable_historytruncatedmean (initial : fin k) (b : ℕ → ℝ) (n : ℕ) (arm : fin k) : measurable (fun h => historytruncatedmean initial b n h arm) theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.history_count_trace","label":"history_count_trace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.history_count_trace","description":"theorem history_count_trace (initial : Fin K) (a : ActionTrace (Fin K)) (r : RewardTrace ℝ) (arm : Fin K) (n t : ℕ) (ht : t ≤ n+1) : pullCount (historyAction initial n (History.finitePairHistoryOfTrace a r n)) arm t = pullCount a arm t","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-2e4afdbd3ce8","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1191,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem history_count_trace (initial : Fin K) (a : ActionTrace (Fin K)) (r : RewardTrace ℝ) (arm : Fin K) (n t : ℕ) (ht : t ≤ n+1) : pullCount (historyAction initial n (History.finitePairHistoryOfTrace a r n)) arm t = pullCount a arm t","missing":[],"search":"history_count_trace banditrlproof.heavytail.history_count_trace theorem history_count_trace (initial : fin k) (a : actiontrace (fin k)) (r : rewardtrace ℝ) (arm : fin k) (n t : ℕ) (ht : t ≤ n+1) : pullcount (historyaction initial n (history.finitepairhistoryoftrace a r n)) arm t = pullcount a arm t theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.historyTruncatedMean_trace","label":"historyTruncatedMean_trace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.historyTruncatedMean_trace","description":"theorem historyTruncatedMean_trace (initial : Fin K) (B : ℕ → ℝ) (a : ActionTrace (Fin K)) (r : RewardTrace ℝ) (arm : Fin K) (n : ℕ) : historyTruncatedMean initial B n (History.finitePairHistoryOfTrace a r n) arm = sumRewards a (fun t => truncate (B (pullCount a arm t)) (r t)) arm (n+1) / pullCount a arm (n+1)","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-e86b8f035ea9","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1192,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem historyTruncatedMean_trace (initial : Fin K) (B : ℕ → ℝ) (a : ActionTrace (Fin K)) (r : RewardTrace ℝ) (arm : Fin K) (n : ℕ) : historyTruncatedMean initial B n (History.finitePairHistoryOfTrace a r n) arm = sumRewards a (fun t => truncate (B (pullCount a arm t)) (r t)) arm (n+1) / pullCount a arm (n+1)","missing":[],"search":"historytruncatedmean_trace banditrlproof.heavytail.historytruncatedmean_trace theorem historytruncatedmean_trace (initial : fin k) (b : ℕ → ℝ) (a : actiontrace (fin k)) (r : rewardtrace ℝ) (arm : fin k) (n : ℕ) : historytruncatedmean initial b n (history.finitepairhistoryoftrace a r n) arm = sumrewards a (fun t => truncate (b (pullcount a arm t)) (r t)) arm (n+1) / pullcount a arm (n+1) theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncated_observed_sum","label":"truncated_observed_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncated_observed_sum","description":"On selected observations, the fixed arm count is the selected arm count. This is the pathwise connection to the latent fixed-prefix confidence theorem.","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-f0e5f20adcc4","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1193,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncated_observed_sum {Ω : Type*} (a : Ω → ActionTrace (Fin K)) (stream : Ω → UCB.ArmRewardStream K) (B : ℕ → ℝ) (ω : Ω) (arm : Fin K) (n : ℕ) : sumRewards (a ω) (fun t => truncate (B (pullCount (a ω) arm t)) (UCB.rewardFromArmStream a stream ω t)) arm n = ∑ s ∈ Finset.range (pullCount (a ω) arm n), truncate (B s) (stream ω s arm)","missing":[],"search":"truncated_observed_sum banditrlproof.heavytail.truncated_observed_sum on selected observations, the fixed arm count is the selected arm count. this is the pathwise connection to the latent fixed-prefix confidence theorem. theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.historyTruncatedMean_latent","label":"historyTruncatedMean_latent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.historyTruncatedMean_latent","description":"theorem historyTruncatedMean_latent (initial : Fin K) (select) (B : ℕ → ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyTruncatedMean initial B n (ArmStreamPolicy.history initial select stream n) arm = (∑ s ∈ Finset.range (pullCount (ArmStreamPolicy.action initial select stream) arm (n+1)), truncate (B s) (stream s arm)) / pullCount (ArmStreamPolicy.action initial select stream) arm (n+1)","url":"../modules/banditrlproof-algorithms-heavytailhistory/index.html#decl-91f9306604ab","parent":"module:BanditRLProof.Algorithms.HeavyTailHistory","order":1194,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailHistory"],["Source","BanditRLProof/Algorithms/HeavyTailHistory.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem historyTruncatedMean_latent (initial : Fin K) (select) (B : ℕ → ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyTruncatedMean initial B n (ArmStreamPolicy.history initial select stream n) arm = (∑ s ∈ Finset.range (pullCount (ArmStreamPolicy.action initial select stream) arm (n+1)), truncate (B s) (stream s arm)) / pullCount (ArmStreamPolicy.action initial select stream) arm (n+1)","missing":[],"search":"historytruncatedmean_latent banditrlproof.heavytail.historytruncatedmean_latent theorem historytruncatedmean_latent (initial : fin k) (select) (b : ℕ → ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (n : ℕ) : historytruncatedmean initial b n (armstreampolicy.history initial select stream n) arm = (∑ s ∈ finset.range (pullcount (armstreampolicy.action initial select stream) arm (n+1)), truncate (b s) (stream s arm)) / pullcount (armstreampolicy.action initial select stream) arm (n+1) theorem compiled","shard":"modules/66bef50c33586d09.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.integrable_id_of_raw_moment","label":"integrable_id_of_raw_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.integrable_id_of_raw_moment","description":"theorem integrable_id_of_raw_moment (μ : Measure ℝ) [IsProbabilityMeasure μ] (ε : ℝ) (hε : 0 ≤ ε) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) μ) : Integrable (fun x : ℝ => x) μ","url":"../modules/banditrlproof-algorithms-heavytailregret/index.html#decl-3a6b4e102a69","parent":"module:BanditRLProof.Algorithms.HeavyTailRegret","order":1195,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegret"],["Source","BanditRLProof/Algorithms/HeavyTailRegret.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_id_of_raw_moment (μ : Measure ℝ) [IsProbabilityMeasure μ] (ε : ℝ) (hε : 0 ≤ ε) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) μ) : Integrable (fun x : ℝ => x) μ","missing":[],"search":"integrable_id_of_raw_moment banditrlproof.heavytail.integrable_id_of_raw_moment theorem integrable_id_of_raw_moment (μ : measure ℝ) [isprobabilitymeasure μ] (ε : ℝ) (hε : 0 ≤ ε) (hm : integrable (fun x : ℝ => |x|^(1+ε)) μ) : integrable (fun x : ℝ => x) μ theorem compiled","shard":"modules/539c7b0648284a1d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_expected_regret","label":"robust_expected_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_expected_regret","description":"theorem robust_expected_regret (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T ∂UCB.armStreamMeasure ν) ≤ ∑ arm : Fin K, realMeanGap (realKernelMean ν) arm * (gapThreshold ε u (realMe…","url":"../modules/banditrlproof-algorithms-heavytailregret/index.html#decl-c1c75dd10a46","parent":"module:BanditRLProof.Algorithms.HeavyTailRegret","order":1196,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegret"],["Source","BanditRLProof/Algorithms/HeavyTailRegret.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_expected_regret (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T ∂UCB.armStreamMeasure ν) ≤ ∑ arm : Fin K, realMeanGap (realKernelMean ν) arm * (gapThreshold ε u (realMeanGap (realKernelMean ν) arm) T + 2)","missing":[],"search":"robust_expected_regret banditrlproof.heavytail.robust_expected_regret theorem robust_expected_regret (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (ε u : ℝ) (t : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, realmeanregret (realkernelmean ν) (robustaction hk ε u stream) t ∂ucb.armstreammeasure ν) ≤ ∑ arm : fin k, realmeangap (realkernelmean ν) arm * (gapthreshold ε u (realmeangap (realkernelmean ν) arm) t + 2) theorem compiled","shard":"modules/539c7b0648284a1d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robust_regret_integrable","label":"robust_regret_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robust_regret_integrable","description":"theorem robust_regret_integrable (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) : Integrable (fun stream => realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T) (UCB.armStreamMeasure ν)","url":"../modules/banditrlproof-algorithms-heavytailregret/index.html#decl-718a7320027e","parent":"module:BanditRLProof.Algorithms.HeavyTailRegret","order":1197,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegret"],["Source","BanditRLProof/Algorithms/HeavyTailRegret.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_regret_integrable (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) : Integrable (fun stream => realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T) (UCB.armStreamMeasure ν)","missing":[],"search":"robust_regret_integrable banditrlproof.heavytail.robust_regret_integrable theorem robust_regret_integrable (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (ε u : ℝ) (t : ℕ) : integrable (fun stream => realmeanregret (realkernelmean ν) (robustaction hk ε u stream) t) (ucb.armstreammeasure ν) theorem compiled","shard":"modules/539c7b0648284a1d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.mean_abs_le_raw_scale","label":"mean_abs_le_raw_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.mean_abs_le_raw_scale","description":"theorem mean_abs_le_raw_scale (μ : Measure ℝ) [IsProbabilityMeasure μ] (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) μ) (hbound : (∫ x, |x|^(1+ε) ∂μ) ≤ u) : |∫ x, x ∂μ| ≤ u^(1/(1+ε))","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-4dfa4faffe03","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1198,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_abs_le_raw_scale (μ : Measure ℝ) [IsProbabilityMeasure μ] (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) μ) (hbound : (∫ x, |x|^(1+ε) ∂μ) ≤ u) : |∫ x, x ∂μ| ≤ u^(1/(1+ε))","missing":[],"search":"mean_abs_le_raw_scale banditrlproof.heavytail.genaltiaudit.mean_abs_le_raw_scale theorem mean_abs_le_raw_scale (μ : measure ℝ) [isprobabilitymeasure μ] (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hm : integrable (fun x : ℝ => |x|^(1+ε)) μ) (hbound : (∫ x, |x|^(1+ε) ∂μ) ≤ u) : |∫ x, x ∂μ| ≤ u^(1/(1+ε)) theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.trace_regret_bounds","label":"trace_regret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.trace_regret_bounds","description":"theorem trace_regret_bounds {K : ℕ} (hK : 0 < K) (m : Fin K → ℝ) (s : ℝ) (hs : 0 ≤ s) (hm : ∀ i, |m i| ≤ s) (action : ActionTrace (Fin K)) (T : ℕ) : 0 ≤ realMeanRegret m action T ∧ realMeanRegret m action T ≤ 2*T*s","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-0c05a3255588","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1199,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem trace_regret_bounds {K : ℕ} (hK : 0 < K) (m : Fin K → ℝ) (s : ℝ) (hs : 0 ≤ s) (hm : ∀ i, |m i| ≤ s) (action : ActionTrace (Fin K)) (T : ℕ) : 0 ≤ realMeanRegret m action T ∧ realMeanRegret m action T ≤ 2*T*s","missing":[],"search":"trace_regret_bounds banditrlproof.heavytail.genaltiaudit.trace_regret_bounds theorem trace_regret_bounds {k : ℕ} (hk : 0 < k) (m : fin k → ℝ) (s : ℝ) (hs : 0 ≤ s) (hm : ∀ i, |m i| ≤ s) (action : actiontrace (fin k)) (t : ℕ) : 0 ≤ realmeanregret m action t ∧ realmeanregret m action t ≤ 2*t*s theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.expected_normalized_regret_le","label":"expected_normalized_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.expected_normalized_regret_le","description":"theorem expected_normalized_regret_le {K : ℕ} (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hm : ∀ i, Integrable (fun x : ℝ => |x|^(1+ε)) (ν i)) (hb : ∀ i, (∫ x, |x|^(1+ε) ∂ν i) ≤ u) {Ω : Type*} [MeasurableSpace Ω] (P : Measure Ω) [IsProbabilityMeasure P] (action : Ω → ActionTrace (Fin K)) (T : ℕ) (ha : ∀ t, Measurable (fun ω => action ω t)) : (∫ ω, realMeanRegret (realK…","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-100b68443537","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1200,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_normalized_regret_le {K : ℕ} (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hm : ∀ i, Integrable (fun x : ℝ => |x|^(1+ε)) (ν i)) (hb : ∀ i, (∫ x, |x|^(1+ε) ∂ν i) ≤ u) {Ω : Type*} [MeasurableSpace Ω] (P : Measure Ω) [IsProbabilityMeasure P] (action : Ω → ActionTrace (Fin K)) (T : ℕ) (ha : ∀ t, Measurable (fun ω => action ω t)) : (∫ ω, realMeanRegret (realKernelMean ν) (action ω) T ∂P) / u^(1/(1+ε)) ≤ 2*T","missing":[],"search":"expected_normalized_regret_le banditrlproof.heavytail.genaltiaudit.expected_normalized_regret_le theorem expected_normalized_regret_le {k : ℕ} (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hm : ∀ i, integrable (fun x : ℝ => |x|^(1+ε)) (ν i)) (hb : ∀ i, (∫ x, |x|^(1+ε) ∂ν i) ≤ u) {ω : type*} [measurablespace ω] (p : measure ω) [isprobabilitymeasure p] (action : ω → actiontrace (fin k)) (t : ℕ) (ha : ∀ t, measurable (fun ω => action ω t)) : (∫ ω, realmeanregret (realkernelmean ν) (action ω) t ∂p) / u^(1/(1+ε)) ≤ 2*t theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalizedValues","label":"normalizedValues","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.normalizedValues","description":"Actual values, with moment laws and trace laws as witnesses, not a bound premise.","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-17331a12afd6","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1201,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:90"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def normalizedValues (K : ℕ) (ε : ℝ) (T : ℕ) : Set EReal","missing":[],"search":"normalizedvalues banditrlproof.heavytail.genaltiaudit.normalizedvalues actual values, with moment laws and trace laws as witnesses, not a bound premise. definition compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalizedValues_le","label":"normalizedValues_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.normalizedValues_le","description":"theorem normalizedValues_le {K : ℕ} (hK : 0 < K) (ε : ℝ) (hε : 0 ≤ ε) (T : ℕ) {r : EReal} (hr : r ∈ normalizedValues K ε T) : r ≤ ((2*(T:ℝ) : ℝ) : EReal)","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-ee3e7e738b45","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1202,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem normalizedValues_le {K : ℕ} (hK : 0 < K) (ε : ℝ) (hε : 0 ≤ ε) (T : ℕ) {r : EReal} (hr : r ∈ normalizedValues K ε T) : r ≤ ((2*(T:ℝ) : ℝ) : EReal)","missing":[],"search":"normalizedvalues_le banditrlproof.heavytail.genaltiaudit.normalizedvalues_le theorem normalizedvalues_le {k : ℕ} (hk : 0 < k) (ε : ℝ) (hε : 0 ≤ ε) (t : ℕ) {r : ereal} (hr : r ∈ normalizedvalues k ε t) : r ≤ ((2*(t:ℝ) : ℝ) : ereal) theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_le","label":"normalized_sSup_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_le","description":"theorem normalized_sSup_le {K : ℕ} (hK : 0 < K) (ε : ℝ) (hε : 0 ≤ ε) (T : ℕ) : sSup (normalizedValues K ε T) ≤ ((2*(T:ℝ) : ℝ) : EReal)","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-7f685005cf3b","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1203,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem normalized_sSup_le {K : ℕ} (hK : 0 < K) (ε : ℝ) (hε : 0 ≤ ε) (T : ℕ) : sSup (normalizedValues K ε T) ≤ ((2*(T:ℝ) : ℝ) : EReal)","missing":[],"search":"normalized_ssup_le banditrlproof.heavytail.genaltiaudit.normalized_ssup_le theorem normalized_ssup_le {k : ℕ} (hk : 0 < k) (ε : ℝ) (hε : 0 ≤ ε) (t : ℕ) : ssup (normalizedvalues k ε t) ≤ ((2*(t:ℝ) : ℝ) : ereal) theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_ne_top","label":"normalized_sSup_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_ne_top","description":"theorem normalized_sSup_ne_top {K : ℕ} (hK : 0 < K) (ε : ℝ) (hε : 0 ≤ ε) (T : ℕ) : sSup (normalizedValues K ε T) ≠ ⊤","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-9a88623d68de","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1204,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem normalized_sSup_ne_top {K : ℕ} (hK : 0 < K) (ε : ℝ) (hε : 0 ≤ ε) (T : ℕ) : sSup (normalizedValues K ε T) ≠ ⊤","missing":[],"search":"normalized_ssup_ne_top banditrlproof.heavytail.genaltiaudit.normalized_ssup_ne_top theorem normalized_ssup_ne_top {k : ℕ} (hk : 0 < k) (ε : ℝ) (hε : 0 ≤ ε) (t : ℕ) : ssup (normalizedvalues k ε t) ≠ ⊤ theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.process_value_mem","label":"process_value_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.process_value_mem","description":"Arbitrary measurable action processes enter the set via their actual image law.","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-e534b03bfec9","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1205,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem process_value_mem {K : ℕ} (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (hu : 0 < u) (hm : ∀ i, Integrable (fun x : ℝ => |x|^(1+ε)) (ν i)) (hb : ∀ i, (∫ x, |x|^(1+ε) ∂ν i) ≤ u) {Ω : Type*} [MeasurableSpace Ω] (P : Measure Ω) [IsProbabilityMeasure P] (action : Ω → ActionTrace (Fin K)) (ha : ∀ t, Measurable (fun ω => action ω t)) (T : ℕ) : (((∫ ω, realMeanRegret (realKernelMean ν) (action ω) T ∂P) / u^(1/(1+ε)) : ℝ) : EReal) ∈ normalizedValues K ε T","missing":[],"search":"process_value_mem banditrlproof.heavytail.genaltiaudit.process_value_mem arbitrary measurable action processes enter the set via their actual image law. theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extremeKernel","label":"extremeKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.extremeKernel","description":"noncomputable def extremeKernel : Kernel (Fin 2) ℝ","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-242b171a047c","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1206,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def extremeKernel : Kernel (Fin 2) ℝ","missing":[],"search":"extremekernel banditrlproof.heavytail.genaltiaudit.extremekernel noncomputable def extremekernel : kernel (fin 2) ℝ definition compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_apply","label":"extreme_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.extreme_apply","description":"@[simp] theorem extreme_apply (i : Fin 2) : extremeKernel i = Measure.dirac (if i=0 then (1:ℝ) else -1)","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-8b2aa1e7531f","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1207,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem extreme_apply (i : Fin 2) : extremeKernel i = Measure.dirac (if i=0 then (1:ℝ) else -1)","missing":[],"search":"extreme_apply banditrlproof.heavytail.genaltiaudit.extreme_apply @[simp] theorem extreme_apply (i : fin 2) : extremekernel i = measure.dirac (if i=0 then (1:ℝ) else -1) theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_moment","label":"extreme_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.extreme_moment","description":"theorem extreme_moment (i : Fin 2) : Integrable (fun x : ℝ => |x|^(1+(1:ℝ))) (extremeKernel i) ∧ (∫ x, |x|^(1+(1:ℝ)) ∂extremeKernel i) ≤ 1","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-d8781c0427b6","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1208,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem extreme_moment (i : Fin 2) : Integrable (fun x : ℝ => |x|^(1+(1:ℝ))) (extremeKernel i) ∧ (∫ x, |x|^(1+(1:ℝ)) ∂extremeKernel i) ≤ 1","missing":[],"search":"extreme_moment banditrlproof.heavytail.genaltiaudit.extreme_moment theorem extreme_moment (i : fin 2) : integrable (fun x : ℝ => |x|^(1+(1:ℝ))) (extremekernel i) ∧ (∫ x, |x|^(1+(1:ℝ)) ∂extremekernel i) ≤ 1 theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_mean","label":"extreme_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.extreme_mean","description":"@[simp] theorem extreme_mean (i : Fin 2) : realKernelMean extremeKernel i = if i=0 then 1 else -1","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-a5a5c03d86f7","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1209,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem extreme_mean (i : Fin 2) : realKernelMean extremeKernel i = if i=0 then 1 else -1","missing":[],"search":"extreme_mean banditrlproof.heavytail.genaltiaudit.extreme_mean @[simp] theorem extreme_mean (i : fin 2) : realkernelmean extremekernel i = if i=0 then 1 else -1 theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_best","label":"extreme_best","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.extreme_best","description":"theorem extreme_best : (⨆ i : Fin 2, realKernelMean extremeKernel i) = 1","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-f45cf052ccd4","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1210,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:161"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem extreme_best : (⨆ i : Fin 2, realKernelMean extremeKernel i) = 1","missing":[],"search":"extreme_best banditrlproof.heavytail.genaltiaudit.extreme_best theorem extreme_best : (⨆ i : fin 2, realkernelmean extremekernel i) = 1 theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_trace_regret","label":"extreme_trace_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.extreme_trace_regret","description":"theorem extreme_trace_regret (T : ℕ) : realMeanRegret (realKernelMean extremeKernel) (fun _ => 1) T = 2*T","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-f42fba135c7a","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1211,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:171"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem extreme_trace_regret (T : ℕ) : realMeanRegret (realKernelMean extremeKernel) (fun _ => 1) T = 2*T","missing":[],"search":"extreme_trace_regret banditrlproof.heavytail.genaltiaudit.extreme_trace_regret theorem extreme_trace_regret (t : ℕ) : realmeanregret (realkernelmean extremekernel) (fun _ => 1) t = 2*t theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.cap_mem","label":"cap_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.cap_mem","description":"theorem cap_mem (T : ℕ) : ((2*(T:ℝ) : ℝ) : EReal) ∈ normalizedValues 2 1 T","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-150f57f37eb0","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1212,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:177"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cap_mem (T : ℕ) : ((2*(T:ℝ) : ℝ) : EReal) ∈ normalizedValues 2 1 T","missing":[],"search":"cap_mem banditrlproof.heavytail.genaltiaudit.cap_mem theorem cap_mem (t : ℕ) : ((2*(t:ℝ) : ℝ) : ereal) ∈ normalizedvalues 2 1 t theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_two_eq","label":"normalized_sSup_two_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_two_eq","description":"theorem normalized_sSup_two_eq (T : ℕ) : sSup (normalizedValues 2 1 T) = ((2*(T:ℝ) : ℝ) : EReal)","url":"../modules/banditrlproof-algorithms-heavytailregretcap/index.html#decl-7fe1d8afb420","parent":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","order":1213,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailRegretCap"],["Source","BanditRLProof/Algorithms/HeavyTailRegretCap.lean:183"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem normalized_sSup_two_eq (T : ℕ) : sSup (normalizedValues 2 1 T) = ((2*(T:ℝ) : ℝ) : EReal)","missing":[],"search":"normalized_ssup_two_eq banditrlproof.heavytail.genaltiaudit.normalized_ssup_two_eq theorem normalized_ssup_two_eq (t : ℕ) : ssup (normalizedvalues 2 1 t) = ((2*(t:ℝ) : ℝ) : ereal) theorem compiled","shard":"modules/9aa14106ccd22260.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.selected_gap_lt","label":"selected_gap_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.selected_gap_lt","description":"theorem selected_gap_lt (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (mean : Fin K → ℝ) (best : Fin K) (n : ℕ) (hn : K ≤ n+1) (hbest : mean best - robustMean hK ε u stream best (n+1) < sourceConfidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) best (n+1))) (hchosen : robustMean hK ε u stream (robustAction hK ε u stream (n+1)) (n+1) - mean (robustAction hK ε u stream (n+1)) < sourceConfidence…","url":"../modules/banditrlproof-algorithms-heavytailsourceadaptive/index.html#decl-ed117e14abe5","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","order":1214,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailSourceAdaptive.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem selected_gap_lt (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (mean : Fin K → ℝ) (best : Fin K) (n : ℕ) (hn : K ≤ n+1) (hbest : mean best - robustMean hK ε u stream best (n+1) < sourceConfidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) best (n+1))) (hchosen : robustMean hK ε u stream (robustAction hK ε u stream (n+1)) (n+1) - mean (robustAction hK ε u stream (n+1)) < sourceConfidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) (robustAction hK ε u stream (n+1)) (n+1))) : mean best - mean (robustAction hK ε u stream (n+1)) < 2*sourceConfidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) (robustAction hK ε u stream (n+1)) (n+1))","missing":[],"search":"selected_gap_lt banditrlproof.heavytail.sourcepolicy.selected_gap_lt theorem selected_gap_lt (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (mean : fin k → ℝ) (best : fin k) (n : ℕ) (hn : k ≤ n+1) (hbest : mean best - robustmean hk ε u stream best (n+1) < sourceconfidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) best (n+1))) (hchosen : robustmean hk ε u stream (robustaction hk ε u stream (n+1)) (n+1) - mean (robustaction hk ε u stream (n+1)) < sourceconfidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) (robustaction hk ε u stream (n+1)) (n+1))) : mean best - mean (robustaction hk ε u stream (n+1)) < 2*sourceconfidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) (robustaction hk ε u stream (n+1)) (n+1)) theorem compiled","shard":"modules/233681694f444278.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_selected_small_radius_tail","label":"robust_selected_small_radius_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_selected_small_radius_tail","description":"theorem robust_selected_small_radius_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction hK ε u stream t = arm ∧ 2*sourceConfid…","url":"../modules/banditrlproof-algorithms-heavytailsourceadaptive/index.html#decl-5f123ef41d9f","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","order":1215,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailSourceAdaptive.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_selected_small_radius_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction hK ε u stream t = arm ∧ 2*sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ (∫ x, x ∂ν best) - ∫ x, x ∂ν arm} ≤ 2*t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"robust_selected_small_radius_tail banditrlproof.heavytail.sourcepolicy.robust_selected_small_radius_tail theorem robust_selected_small_radius_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t : ℕ) (ht : k ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (ucb.armstreammeasure ν).real {stream | robustaction hk ε u stream t = arm ∧ 2*sourceconfidenceradius ε u t (pullcount (robustaction hk ε u stream) arm t) ≤ (∫ x, x ∂ν best) - ∫ x, x ∂ν arm} ≤ 2*t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) theorem compiled","shard":"modules/233681694f444278.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_initial_count_zero","label":"robust_initial_count_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_initial_count_zero","description":"theorem robust_initial_count_zero (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : pullCount (robustAction hK ε u stream) (robustAction hK ε u stream t) t = 0","url":"../modules/banditrlproof-algorithms-heavytailsourceadaptive/index.html#decl-7a3ca29371fd","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","order":1216,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailSourceAdaptive.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_initial_count_zero (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : pullCount (robustAction hK ε u stream) (robustAction hK ε u stream t) t = 0","missing":[],"search":"robust_initial_count_zero banditrlproof.heavytail.sourcepolicy.robust_initial_count_zero theorem robust_initial_count_zero (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (t : ℕ) (ht : t < k) : pullcount (robustaction hk ε u stream) (robustaction hk ε u stream t) t = 0 theorem compiled","shard":"modules/233681694f444278.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_large_count_tail","label":"robust_large_count_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_large_count_tail","description":"theorem robust_large_count_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T t : ℕ) (hT : 2 ≤ T) (ht : t < T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream |…","url":"../modules/banditrlproof-algorithms-heavytailsourceadaptive/index.html#decl-856d5544a0fa","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","order":1217,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceAdaptive"],["Source","BanditRLProof/Algorithms/HeavyTailSourceAdaptive.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_large_count_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T t : ℕ) (hT : 2 ≤ T) (ht : t < T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (UCB.armStreamMeasure ν).real {stream | robustAction hK ε u stream t = arm ∧ gapThreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T ≤ pullCount (robustAction hK ε u stream) arm t} ≤ 2*t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"robust_large_count_tail banditrlproof.heavytail.sourcepolicy.robust_large_count_tail theorem robust_large_count_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t t : ℕ) (ht : 2 ≤ t) (ht : t < t) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (ucb.armstreammeasure ν).real {stream | robustaction hk ε u stream t = arm ∧ gapthreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) t ≤ pullcount (robustaction hk ε u stream) arm t} ≤ 2*t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) theorem compiled","shard":"modules/233681694f444278.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.dropped","label":"dropped","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.dropped","description":"noncomputable def dropped (L : ℝ) (n : ℕ) : ℕ","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-92965ffad5d3","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1218,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def dropped (L : ℝ) (n : ℕ) : ℕ","missing":[],"search":"dropped banditrlproof.heavytail.sourcecounterexample.dropped noncomputable def dropped (l : ℝ) (n : ℕ) : ℕ definition compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.truncate_neg_one","label":"truncate_neg_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.truncate_neg_one","description":"theorem truncate_neg_one (L : ℝ) (hL : 0 < L) (s : ℕ) : truncate (sourceTruncationThreshold 1 1 L s) (-1) = if (s : ℝ)+1 < L then 0 else -1","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-d1eadcd56d71","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1219,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncate_neg_one (L : ℝ) (hL : 0 < L) (s : ℕ) : truncate (sourceTruncationThreshold 1 1 L s) (-1) = if (s : ℝ)+1 < L then 0 else -1","missing":[],"search":"truncate_neg_one banditrlproof.heavytail.sourcecounterexample.truncate_neg_one theorem truncate_neg_one (l : ℝ) (hl : 0 < l) (s : ℕ) : truncate (sourcetruncationthreshold 1 1 l s) (-1) = if (s : ℝ)+1 < l then 0 else -1 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.dropped_filter","label":"dropped_filter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.dropped_filter","description":"theorem dropped_filter (L : ℝ) (n : ℕ) : (Finset.range n).filter (fun s : ℕ => (s : ℝ)+1 < L) = Finset.range (dropped L n)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-4dec98592531","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1220,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dropped_filter (L : ℝ) (n : ℕ) : (Finset.range n).filter (fun s : ℕ => (s : ℝ)+1 < L) = Finset.range (dropped L n)","missing":[],"search":"dropped_filter banditrlproof.heavytail.sourcecounterexample.dropped_filter theorem dropped_filter (l : ℝ) (n : ℕ) : (finset.range n).filter (fun s : ℕ => (s : ℝ)+1 < l) = finset.range (dropped l n) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.sum_truncate_neg_one","label":"sum_truncate_neg_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.sum_truncate_neg_one","description":"theorem sum_truncate_neg_one (L : ℝ) (hL : 0 < L) (n : ℕ) : (∑ s ∈ Finset.range n, truncate (sourceTruncationThreshold 1 1 L s) (-1)) = -(n : ℝ) + dropped L n","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-42654b639c51","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1221,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_truncate_neg_one (L : ℝ) (hL : 0 < L) (n : ℕ) : (∑ s ∈ Finset.range n, truncate (sourceTruncationThreshold 1 1 L s) (-1)) = -(n : ℝ) + dropped L n","missing":[],"search":"sum_truncate_neg_one banditrlproof.heavytail.sourcecounterexample.sum_truncate_neg_one theorem sum_truncate_neg_one (l : ℝ) (hl : 0 < l) (n : ℕ) : (∑ s ∈ finset.range n, truncate (sourcetruncationthreshold 1 1 l s) (-1)) = -(n : ℝ) + dropped l n theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.dropped_ge","label":"dropped_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.dropped_ge","description":"theorem dropped_ge (L : ℝ) (hL : 0 < L) (n : ℕ) (hn : L ≤ n) : L-1 ≤ (dropped L n : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-ab32ee7cbeff","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1222,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dropped_ge (L : ℝ) (hL : 0 < L) (n : ℕ) (hn : L ≤ n) : L-1 ≤ (dropped L n : ℝ)","missing":[],"search":"dropped_ge banditrlproof.heavytail.sourcecounterexample.dropped_ge theorem dropped_ge (l : ℝ) (hl : 0 < l) (n : ℕ) (hn : l ≤ n) : l-1 ≤ (dropped l n : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.suboptimal_index_gt","label":"suboptimal_index_gt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.suboptimal_index_gt","description":"theorem suboptimal_index_gt (H L n d : ℝ) (hH : 30 < H) (hL : 2*H-2 ≤ L) (hn : 0 < n) (hnM : n ≤ 32*H+5) (hd : L-1 ≤ d) : (1/100 : ℝ) < -1+d/n+4*Real.sqrt (L/n)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-d9f65ec1d71d","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1223,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem suboptimal_index_gt (H L n d : ℝ) (hH : 30 < H) (hL : 2*H-2 ≤ L) (hn : 0 < n) (hnM : n ≤ 32*H+5) (hd : L-1 ≤ d) : (1/100 : ℝ) < -1+d/n+4*Real.sqrt (L/n)","missing":[],"search":"suboptimal_index_gt banditrlproof.heavytail.sourcecounterexample.suboptimal_index_gt theorem suboptimal_index_gt (h l n d : ℝ) (hh : 30 < h) (hl : 2*h-2 ≤ l) (hn : 0 < n) (hnm : n ≤ 32*h+5) (hd : l-1 ≤ d) : (1/100 : ℝ) < -1+d/n+4*real.sqrt (l/n) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.optimal_index_lt","label":"optimal_index_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.optimal_index_lt","description":"theorem optimal_index_lt (H L m : ℝ) (hH : 0 < H) (hL0 : 0 ≤ L) (hL : L ≤ 2*H) (hm : 320000*H < m) : 4*Real.sqrt (L/m) < (1/100 : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-aac8df5095a9","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1224,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem optimal_index_lt (H L m : ℝ) (hH : 0 < H) (hL0 : 0 ≤ L) (hL : L ≤ 2*H) (hm : 320000*H < m) : 4*Real.sqrt (L/m) < (1/100 : ℝ)","missing":[],"search":"optimal_index_lt banditrlproof.heavytail.sourcecounterexample.optimal_index_lt theorem optimal_index_lt (h l m : ℝ) (hh : 0 < h) (hl0 : 0 ≤ l) (hl : l ≤ 2*h) (hm : 320000*h < m) : 4*real.sqrt (l/m) < (1/100 : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.deterministicStream","label":"deterministicStream","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.deterministicStream","description":"noncomputable def deterministicStream : UCB.ArmRewardStream 2","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-801a4806ccd2","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1225,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:90"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def deterministicStream : UCB.ArmRewardStream 2","missing":[],"search":"deterministicstream banditrlproof.heavytail.sourcecounterexample.deterministicstream noncomputable def deterministicstream : ucb.armrewardstream 2 definition compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.trace","label":"trace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.trace","description":"noncomputable def trace : ActionTrace (Fin 2)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-d995f0b5e95b","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1226,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def trace : ActionTrace (Fin 2)","missing":[],"search":"trace banditrlproof.heavytail.sourcecounterexample.trace noncomputable def trace : actiontrace (fin 2) definition compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.mean_zero","label":"mean_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.mean_zero","description":"theorem mean_zero (t : ℕ) : SourcePolicy.robustMean (by decide) 1 1 deterministicStream 0 t = 0","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-3c3d14d2b1d9","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1227,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_zero (t : ℕ) : SourcePolicy.robustMean (by decide) 1 1 deterministicStream 0 t = 0","missing":[],"search":"mean_zero banditrlproof.heavytail.sourcecounterexample.mean_zero theorem mean_zero (t : ℕ) : sourcepolicy.robustmean (by decide) 1 1 deterministicstream 0 t = 0 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.mean_one","label":"mean_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.mean_one","description":"theorem mean_one (t : ℕ) (ht : 2 ≤ t) : SourcePolicy.robustMean (by decide) 1 1 deterministicStream 1 t = -1 + dropped (sourceConfidenceLog t) (pullCount trace 1 t) / pullCount trace 1 t","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-b9a3ddc70e55","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1228,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_one (t : ℕ) (ht : 2 ≤ t) : SourcePolicy.robustMean (by decide) 1 1 deterministicStream 1 t = -1 + dropped (sourceConfidenceLog t) (pullCount trace 1 t) / pullCount trace 1 t","missing":[],"search":"mean_one banditrlproof.heavytail.sourcecounterexample.mean_one theorem mean_one (t : ℕ) (ht : 2 ≤ t) : sourcepolicy.robustmean (by decide) 1 1 deterministicstream 1 t = -1 + dropped (sourceconfidencelog t) (pullcount trace 1 t) / pullcount trace 1 t theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.radius_sqrt","label":"radius_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.radius_sqrt","description":"theorem radius_sqrt (t n : ℕ) : sourceConfidenceRadius 1 1 t n = 4 * Real.sqrt (sourceConfidenceLog t/n)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-eefccc564b14","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1229,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem radius_sqrt (t n : ℕ) : sourceConfidenceRadius 1 1 t n = 4 * Real.sqrt (sourceConfidenceLog t/n)","missing":[],"search":"radius_sqrt banditrlproof.heavytail.sourcecounterexample.radius_sqrt theorem radius_sqrt (t n : ℕ) : sourceconfidenceradius 1 1 t n = 4 * real.sqrt (sourceconfidencelog t/n) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.source_index_sub_gt","label":"source_index_sub_gt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.source_index_sub_gt","description":"theorem source_index_sub_gt (H : ℝ) (t : ℕ) (ht : 2 ≤ t) (hH : 30 < H) (hL : 2*H-2 ≤ sourceConfidenceLog t) (hnM : (pullCount trace 1 t : ℝ) ≤ 32*H+5) : (1/100 : ℝ) < SourcePolicy.robustMean (by decide) 1 1 deterministicStream 1 t + sourceConfidenceRadius 1 1 t (pullCount trace 1 t)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-509d3c09ef75","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1230,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_index_sub_gt (H : ℝ) (t : ℕ) (ht : 2 ≤ t) (hH : 30 < H) (hL : 2*H-2 ≤ sourceConfidenceLog t) (hnM : (pullCount trace 1 t : ℝ) ≤ 32*H+5) : (1/100 : ℝ) < SourcePolicy.robustMean (by decide) 1 1 deterministicStream 1 t + sourceConfidenceRadius 1 1 t (pullCount trace 1 t)","missing":[],"search":"source_index_sub_gt banditrlproof.heavytail.sourcecounterexample.source_index_sub_gt theorem source_index_sub_gt (h : ℝ) (t : ℕ) (ht : 2 ≤ t) (hh : 30 < h) (hl : 2*h-2 ≤ sourceconfidencelog t) (hnm : (pullcount trace 1 t : ℝ) ≤ 32*h+5) : (1/100 : ℝ) < sourcepolicy.robustmean (by decide) 1 1 deterministicstream 1 t + sourceconfidenceradius 1 1 t (pullcount trace 1 t) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.source_index_best_lt","label":"source_index_best_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.source_index_best_lt","description":"theorem source_index_best_lt (H : ℝ) (t : ℕ) (ht : 2 ≤ t) (hH : 0 < H) (hL : sourceConfidenceLog t ≤ 2*H) (hm : 320000*H < (pullCount trace 0 t : ℝ)) : SourcePolicy.robustMean (by decide) 1 1 deterministicStream 0 t + sourceConfidenceRadius 1 1 t (pullCount trace 0 t) < (1/100 : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-691ebfc5e552","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1231,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_index_best_lt (H : ℝ) (t : ℕ) (ht : 2 ≤ t) (hH : 0 < H) (hL : sourceConfidenceLog t ≤ 2*H) (hm : 320000*H < (pullCount trace 0 t : ℝ)) : SourcePolicy.robustMean (by decide) 1 1 deterministicStream 0 t + sourceConfidenceRadius 1 1 t (pullCount trace 0 t) < (1/100 : ℝ)","missing":[],"search":"source_index_best_lt banditrlproof.heavytail.sourcecounterexample.source_index_best_lt theorem source_index_best_lt (h : ℝ) (t : ℕ) (ht : 2 ≤ t) (hh : 0 < h) (hl : sourceconfidencelog t ≤ 2*h) (hm : 320000*h < (pullcount trace 0 t : ℝ)) : sourcepolicy.robustmean (by decide) 1 1 deterministicstream 0 t + sourceconfidenceradius 1 1 t (pullcount trace 0 t) < (1/100 : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.forced_suboptimal","label":"forced_suboptimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.forced_suboptimal","description":"theorem forced_suboptimal (H : ℝ) (t : ℕ) (ht : 2 ≤ t) (hH : 30 < H) (hLlo : 2*H-2 ≤ sourceConfidenceLog t) (hLhi : sourceConfidenceLog t ≤ 2*H) (hnM : (pullCount trace 1 t : ℝ) ≤ 32*H+5) (hm : 320000*H < (pullCount trace 0 t : ℝ)) : trace t = 1","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-9d020768bbb0","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1232,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem forced_suboptimal (H : ℝ) (t : ℕ) (ht : 2 ≤ t) (hH : 30 < H) (hLlo : 2*H-2 ≤ sourceConfidenceLog t) (hLhi : sourceConfidenceLog t ≤ 2*H) (hnM : (pullCount trace 1 t : ℝ) ≤ 32*H+5) (hm : 320000*H < (pullCount trace 0 t : ℝ)) : trace t = 1","missing":[],"search":"forced_suboptimal banditrlproof.heavytail.sourcecounterexample.forced_suboptimal theorem forced_suboptimal (h : ℝ) (t : ℕ) (ht : 2 ≤ t) (hh : 30 < h) (hllo : 2*h-2 ≤ sourceconfidencelog t) (hlhi : sourceconfidencelog t ≤ 2*h) (hnm : (pullcount trace 1 t : ℝ) ≤ 32*h+5) (hm : 320000*h < (pullcount trace 0 t : ℝ)) : trace t = 1 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.horizon","label":"horizon","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.horizon","description":"def horizon : ℕ","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-55914b35a4e1","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1233,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:173"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def horizon : ℕ","missing":[],"search":"horizon banditrlproof.heavytail.sourcecounterexample.horizon def horizon : ℕ definition compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.horizonHeight","label":"horizonHeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.horizonHeight","description":"noncomputable def horizonHeight : ℝ","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-a1b51812e316","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1234,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def horizonHeight : ℝ","missing":[],"search":"horizonheight banditrlproof.heavytail.sourcecounterexample.horizonheight noncomputable def horizonheight : ℝ definition compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.log_two_bounds","label":"log_two_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.log_two_bounds","description":"theorem log_two_bounds : (3/5 : ℝ) < Real.log 2 ∧ Real.log 2 < 1","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-02659b100405","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1235,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:176"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem log_two_bounds : (3/5 : ℝ) < Real.log 2 ∧ Real.log 2 < 1","missing":[],"search":"log_two_bounds banditrlproof.heavytail.sourcecounterexample.log_two_bounds theorem log_two_bounds : (3/5 : ℝ) < real.log 2 ∧ real.log 2 < 1 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.height_bounds","label":"height_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.height_bounds","description":"theorem height_bounds : (30 : ℝ) < horizonHeight ∧ horizonHeight < 50","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-6027fcfbc5c3","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1236,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:189"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem height_bounds : (30 : ℝ) < horizonHeight ∧ horizonHeight < 50","missing":[],"search":"height_bounds banditrlproof.heavytail.sourcecounterexample.height_bounds theorem height_bounds : (30 : ℝ) < horizonheight ∧ horizonheight < 50 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.late_log_bounds","label":"late_log_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.late_log_bounds","description":"theorem late_log_bounds (t : ℕ) (htlo : horizon/2 ≤ t) (hthi : t < horizon) : 2*horizonHeight-2 ≤ sourceConfidenceLog t ∧ sourceConfidenceLog t ≤ 2*horizonHeight","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-455f584305db","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1237,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:197"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem late_log_bounds (t : ℕ) (htlo : horizon/2 ≤ t) (hthi : t < horizon) : 2*horizonHeight-2 ≤ sourceConfidenceLog t ∧ sourceConfidenceLog t ≤ 2*horizonHeight","missing":[],"search":"late_log_bounds banditrlproof.heavytail.sourcecounterexample.late_log_bounds theorem late_log_bounds (t : ℕ) (htlo : horizon/2 ≤ t) (hthi : t < horizon) : 2*horizonheight-2 ≤ sourceconfidencelog t ∧ sourceconfidencelog t ≤ 2*horizonheight theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.late_best_count","label":"late_best_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.late_best_count","description":"theorem late_best_count (t : ℕ) (ht : horizon/2 ≤ t) (hthi : t ≤ horizon) (hcount : (pullCount trace 1 horizon : ℝ) ≤ 32*horizonHeight+5) : 320000*horizonHeight < (pullCount trace 0 t : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-26a5eec88ba0","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1238,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:210"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem late_best_count (t : ℕ) (ht : horizon/2 ≤ t) (hthi : t ≤ horizon) (hcount : (pullCount trace 1 horizon : ℝ) ≤ 32*horizonHeight+5) : 320000*horizonHeight < (pullCount trace 0 t : ℝ)","missing":[],"search":"late_best_count banditrlproof.heavytail.sourcecounterexample.late_best_count theorem late_best_count (t : ℕ) (ht : horizon/2 ≤ t) (hthi : t ≤ horizon) (hcount : (pullcount trace 1 horizon : ℝ) ≤ 32*horizonheight+5) : 320000*horizonheight < (pullcount trace 0 t : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.finite_count_obstruction","label":"finite_count_obstruction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.finite_count_obstruction","description":"theorem finite_count_obstruction : 32*horizonHeight+5 < (pullCount trace 1 horizon : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-5fa338e5631d","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1239,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:223"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finite_count_obstruction : 32*horizonHeight+5 < (pullCount trace 1 horizon : ℝ)","missing":[],"search":"finite_count_obstruction banditrlproof.heavytail.sourcecounterexample.finite_count_obstruction theorem finite_count_obstruction : 32*horizonheight+5 < (pullcount trace 1 horizon : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel","label":"kernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.kernel","description":"noncomputable def kernel : Kernel (Fin 2) ℝ","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-d9d4d1f8b9e6","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1240,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:249"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def kernel : Kernel (Fin 2) ℝ","missing":[],"search":"kernel banditrlproof.heavytail.sourcecounterexample.kernel noncomputable def kernel : kernel (fin 2) ℝ definition compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_apply","label":"kernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.kernel_apply","description":"@[simp] theorem kernel_apply (a : Fin 2) : kernel a = Measure.dirac (if a = 0 then (0 : ℝ) else -1)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-b5355c577b42","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1241,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:252"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem kernel_apply (a : Fin 2) : kernel a = Measure.dirac (if a = 0 then (0 : ℝ) else -1)","missing":[],"search":"kernel_apply banditrlproof.heavytail.sourcecounterexample.kernel_apply @[simp] theorem kernel_apply (a : fin 2) : kernel a = measure.dirac (if a = 0 then (0 : ℝ) else -1) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_raw_moment","label":"kernel_raw_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.kernel_raw_moment","description":"theorem kernel_raw_moment (a : Fin 2) : Integrable (fun x : ℝ => |x|^(1+(1 : ℝ))) (kernel a) ∧ (∫ x : ℝ, |x|^(1+(1 : ℝ)) ∂kernel a) ≤ 1","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-e13e32d18f4b","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1242,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:261"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem kernel_raw_moment (a : Fin 2) : Integrable (fun x : ℝ => |x|^(1+(1 : ℝ))) (kernel a) ∧ (∫ x : ℝ, |x|^(1+(1 : ℝ)) ∂kernel a) ≤ 1","missing":[],"search":"kernel_raw_moment banditrlproof.heavytail.sourcecounterexample.kernel_raw_moment theorem kernel_raw_moment (a : fin 2) : integrable (fun x : ℝ => |x|^(1+(1 : ℝ))) (kernel a) ∧ (∫ x : ℝ, |x|^(1+(1 : ℝ)) ∂kernel a) ≤ 1 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_mean","label":"kernel_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.kernel_mean","description":"@[simp] theorem kernel_mean (a : Fin 2) : realKernelMean kernel a = if a = 0 then 0 else -1","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-80b11fdea804","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1243,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:270"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem kernel_mean (a : Fin 2) : realKernelMean kernel a = if a = 0 then 0 else -1","missing":[],"search":"kernel_mean banditrlproof.heavytail.sourcecounterexample.kernel_mean @[simp] theorem kernel_mean (a : fin 2) : realkernelmean kernel a = if a = 0 then 0 else -1 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.ae_deterministic","label":"ae_deterministic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.ae_deterministic","description":"theorem ae_deterministic : ∀ᵐ stream ∂UCB.armStreamMeasure kernel, stream = deterministicStream","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-b2edce84161a","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1244,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:274"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ae_deterministic : ∀ᵐ stream ∂UCB.armStreamMeasure kernel, stream = deterministicStream","missing":[],"search":"ae_deterministic banditrlproof.heavytail.sourcecounterexample.ae_deterministic theorem ae_deterministic : ∀ᵐ stream ∂ucb.armstreammeasure kernel, stream = deterministicstream theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_best_mean","label":"kernel_best_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.kernel_best_mean","description":"theorem kernel_best_mean : (⨆ a : Fin 2, realKernelMean kernel a) = 0","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-f3b5995eef7f","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1245,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:293"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem kernel_best_mean : (⨆ a : Fin 2, realKernelMean kernel a) = 0","missing":[],"search":"kernel_best_mean banditrlproof.heavytail.sourcecounterexample.kernel_best_mean theorem kernel_best_mean : (⨆ a : fin 2, realkernelmean kernel a) = 0 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_gap","label":"kernel_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.kernel_gap","description":"theorem kernel_gap (a : Fin 2) : realMeanGap (realKernelMean kernel) a = if a = 0 then 0 else 1","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-9d9dcec9c4a8","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1246,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:303"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem kernel_gap (a : Fin 2) : realMeanGap (realKernelMean kernel) a = if a = 0 then 0 else 1","missing":[],"search":"kernel_gap banditrlproof.heavytail.sourcecounterexample.kernel_gap theorem kernel_gap (a : fin 2) : realmeangap (realkernelmean kernel) a = if a = 0 then 0 else 1 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.regret_eq_count","label":"regret_eq_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.regret_eq_count","description":"theorem regret_eq_count (action : ActionTrace (Fin 2)) (T : ℕ) : realMeanRegret (realKernelMean kernel) action T = (pullCount action 1 T : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-75fadc12b6a9","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1247,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:308"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem regret_eq_count (action : ActionTrace (Fin 2)) (T : ℕ) : realMeanRegret (realKernelMean kernel) action T = (pullCount action 1 T : ℝ)","missing":[],"search":"regret_eq_count banditrlproof.heavytail.sourcecounterexample.regret_eq_count theorem regret_eq_count (action : actiontrace (fin 2)) (t : ℕ) : realmeanregret (realkernelmean kernel) action t = (pullcount action 1 t : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.expected_regret_eq_count","label":"expected_regret_eq_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.expected_regret_eq_count","description":"theorem expected_regret_eq_count : (∫ stream, realMeanRegret (realKernelMean kernel) (SourcePolicy.robustAction (by decide) 1 1 stream) horizon ∂UCB.armStreamMeasure kernel) = (pullCount trace 1 horizon : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-b245cfd1e364","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1248,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:313"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_regret_eq_count : (∫ stream, realMeanRegret (realKernelMean kernel) (SourcePolicy.robustAction (by decide) 1 1 stream) horizon ∂UCB.armStreamMeasure kernel) = (pullCount trace 1 horizon : ℝ)","missing":[],"search":"expected_regret_eq_count banditrlproof.heavytail.sourcecounterexample.expected_regret_eq_count theorem expected_regret_eq_count : (∫ stream, realmeanregret (realkernelmean kernel) (sourcepolicy.robustaction (by decide) 1 1 stream) horizon ∂ucb.armstreammeasure kernel) = (pullcount trace 1 horizon : ℝ) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.printed_coefficient_counterexample","label":"printed_coefficient_counterexample","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.printed_coefficient_counterexample","description":"The literal printed coefficient fails for the actual source-parameter policy under valid raw-second-moment stationary two-arm laws at a finite horizon.","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-3a23ad3a9109","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1249,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:329"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem printed_coefficient_counterexample : 32 * Real.log (horizon : ℝ) + 5 < ∫ stream, realMeanRegret (realKernelMean kernel) (SourcePolicy.robustAction (by decide) 1 1 stream) horizon ∂UCB.armStreamMeasure kernel","missing":[],"search":"printed_coefficient_counterexample banditrlproof.heavytail.sourcecounterexample.printed_coefficient_counterexample the literal printed coefficient fails for the actual source-parameter policy under valid raw-second-moment stationary two-arm laws at a finite horizon. theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.printed_gap_sum","label":"printed_gap_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.printed_gap_sum","description":"theorem printed_gap_sum : (∑ a ∈ Finset.univ.filter (fun a : Fin 2 => 0 < realMeanGap (realKernelMean kernel) a), (8 * (4 / realMeanGap (realKernelMean kernel) a) ^ (1 / (1 : ℝ)) * Real.log (horizon : ℝ) + 5 * realMeanGap (realKernelMean kernel) a)) = 32 * Real.log (horizon : ℝ) + 5","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-d45206b7694b","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1250,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:338"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem printed_gap_sum : (∑ a ∈ Finset.univ.filter (fun a : Fin 2 => 0 < realMeanGap (realKernelMean kernel) a), (8 * (4 / realMeanGap (realKernelMean kernel) a) ^ (1 / (1 : ℝ)) * Real.log (horizon : ℝ) + 5 * realMeanGap (realKernelMean kernel) a)) = 32 * Real.log (horizon : ℝ) + 5","missing":[],"search":"printed_gap_sum banditrlproof.heavytail.sourcecounterexample.printed_gap_sum theorem printed_gap_sum : (∑ a ∈ finset.univ.filter (fun a : fin 2 => 0 < realmeangap (realkernelmean kernel) a), (8 * (4 / realmeangap (realkernelmean kernel) a) ^ (1 / (1 : ℝ)) * real.log (horizon : ℝ) + 5 * realmeangap (realkernelmean kernel) a)) = 32 * real.log (horizon : ℝ) + 5 theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.literal_printed_bound_false","label":"literal_printed_bound_false","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourceCounterexample.literal_printed_bound_false","description":"theorem literal_printed_bound_false : ¬ (∫ stream, realMeanRegret (realKernelMean kernel) (SourcePolicy.robustAction (by decide) 1 1 stream) horizon ∂UCB.armStreamMeasure kernel) ≤ ∑ a ∈ Finset.univ.filter (fun a : Fin 2 => 0 < realMeanGap (realKernelMean kernel) a), (8 * (4 / realMeanGap (realKernelMean kernel) a) ^ (1 / (1 : ℝ)) * Real.log (horizon : ℝ) + 5 * realMeanGap (realKernelMean kernel) a)","url":"../modules/banditrlproof-algorithms-heavytailsourcecounterexample/index.html#decl-981f16c46ca9","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","order":1251,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceCounterexample"],["Source","BanditRLProof/Algorithms/HeavyTailSourceCounterexample.lean:347"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem literal_printed_bound_false : ¬ (∫ stream, realMeanRegret (realKernelMean kernel) (SourcePolicy.robustAction (by decide) 1 1 stream) horizon ∂UCB.armStreamMeasure kernel) ≤ ∑ a ∈ Finset.univ.filter (fun a : Fin 2 => 0 < realMeanGap (realKernelMean kernel) a), (8 * (4 / realMeanGap (realKernelMean kernel) a) ^ (1 / (1 : ℝ)) * Real.log (horizon : ℝ) + 5 * realMeanGap (realKernelMean kernel) a)","missing":[],"search":"literal_printed_bound_false banditrlproof.heavytail.sourcecounterexample.literal_printed_bound_false theorem literal_printed_bound_false : ¬ (∫ stream, realmeanregret (realkernelmean kernel) (sourcepolicy.robustaction (by decide) 1 1 stream) horizon ∂ucb.armstreammeasure kernel) ≤ ∑ a ∈ finset.univ.filter (fun a : fin 2 => 0 < realmeangap (realkernelmean kernel) a), (8 * (4 / realmeangap (realkernelmean kernel) a) ^ (1 / (1 : ℝ)) * real.log (horizon : ℝ) + 5 * realmeangap (realkernelmean kernel) a) theorem compiled","shard":"modules/9b3ed05d349ab7f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_lintegral_count_le","label":"robust_lintegral_count_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_lintegral_count_le","description":"theorem robust_lintegral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hT : 2 ≤ T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫⁻ stream, (pullCount (robustAction hK ε u stream) a…","url":"../modules/banditrlproof-algorithms-heavytailsourceexpectedcount/index.html#decl-29d87bb705f5","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","order":1252,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailSourceExpectedCount.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_lintegral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hT : 2 ≤ T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫⁻ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ≥0∞) ∂UCB.armStreamMeasure ν) ≤ gapThreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T + 4","missing":[],"search":"robust_lintegral_count_le banditrlproof.heavytail.sourcepolicy.robust_lintegral_count_le theorem robust_lintegral_count_le (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t : ℕ) (ht : 2 ≤ t) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫⁻ stream, (pullcount (robustaction hk ε u stream) arm t : ℝ≥0∞) ∂ucb.armstreammeasure ν) ≤ gapthreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) t + 4 theorem compiled","shard":"modules/bdee61a9893b6589.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integrable_count","label":"robust_integrable_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_integrable_count","description":"theorem robust_integrable_count (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (T : ℕ) : Integrable (fun stream => (pullCount (robustAction hK ε u stream) arm T : ℝ)) (UCB.armStreamMeasure ν)","url":"../modules/banditrlproof-algorithms-heavytailsourceexpectedcount/index.html#decl-101e44f56bc1","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","order":1253,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailSourceExpectedCount.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_integrable_count (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (T : ℕ) : Integrable (fun stream => (pullCount (robustAction hK ε u stream) arm T : ℝ)) (UCB.armStreamMeasure ν)","missing":[],"search":"robust_integrable_count banditrlproof.heavytail.sourcepolicy.robust_integrable_count theorem robust_integrable_count (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (ε u : ℝ) (t : ℕ) : integrable (fun stream => (pullcount (robustaction hk ε u stream) arm t : ℝ)) (ucb.armstreammeasure ν) theorem compiled","shard":"modules/bdee61a9893b6589.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le","label":"robust_integral_count_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le","description":"theorem robust_integral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hT : 2 ≤ T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullCount (robustAction hK ε u stream) arm…","url":"../modules/banditrlproof-algorithms-heavytailsourceexpectedcount/index.html#decl-ac8044c3f9dd","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","order":1254,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailSourceExpectedCount.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_integral_count_le (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hT : 2 ≤ T) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ) ∂UCB.armStreamMeasure ν) ≤ gapThreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T + 4","missing":[],"search":"robust_integral_count_le banditrlproof.heavytail.sourcepolicy.robust_integral_count_le theorem robust_integral_count_le (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (best arm : fin k) (ε u : ℝ) (t : ℕ) (ht : 2 ≤ t) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hx : ∀ a, integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullcount (robustaction hk ε u stream) arm t : ℝ) ∂ucb.armstreammeasure ν) ≤ gapthreshold ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) t + 4 theorem compiled","shard":"modules/bdee61a9893b6589.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le_budget","label":"robust_integral_count_le_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le_budget","description":"All horizons, retaining the additive five without a positive-cutoff assumption at T=0/1.","url":"../modules/banditrlproof-algorithms-heavytailsourceexpectedcount/index.html#decl-a370f303933f","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","order":1255,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceExpectedCount"],["Source","BanditRLProof/Algorithms/HeavyTailSourceExpectedCount.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_integral_count_le_budget (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (best arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hgap : 0 < (∫ x, x ∂ν best) - ∫ x, x ∂ν arm) (hX : ∀ a, Integrable (fun x : ℝ => x) (ν a)) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, (pullCount (robustAction hK ε u stream) arm T : ℝ) ∂UCB.armStreamMeasure ν) ≤ gapBudget ε u ((∫ x, x ∂ν best) - ∫ x, x ∂ν arm) T + 5","missing":[],"search":"robust_integral_count_le_budget banditrlproof.heavytail.sourcepolicy.robust_integral_count_le_budget all horizons, retaining the additive five without a positive-cutoff assumption at t=0/1. theorem compiled","shard":"modules/bdee61a9893b6589.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.sampleThreshold","label":"sampleThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.sampleThreshold","description":"noncomputable def sampleThreshold (ε u : ℝ) (t s : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-5096a85410ae","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1256,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sampleThreshold (ε u : ℝ) (t s : ℕ) : ℝ","missing":[],"search":"samplethreshold banditrlproof.heavytail.sourcepolicy.samplethreshold noncomputable def samplethreshold (ε u : ℝ) (t s : ℕ) : ℝ definition compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.historyIndex","label":"historyIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.historyIndex","description":"noncomputable def historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) (arm : Fin K) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-cf5e781486fa","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1257,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) (arm : Fin K) : ℝ","missing":[],"search":"historyindex banditrlproof.heavytail.sourcepolicy.historyindex noncomputable def historyindex (initial : fin k) (ε u : ℝ) (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) (arm : fin k) : ℝ definition compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.nextArm","label":"nextArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.nextArm","description":"noncomputable def nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : Fin K","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-087bbec19bcc","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1258,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : Fin K","missing":[],"search":"nextarm banditrlproof.heavytail.sourcepolicy.nextarm noncomputable def nextarm (hk : 0 < k) (ε u : ℝ) (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) : fin k definition compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.measurable_historyIndex","label":"measurable_historyIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.measurable_historyIndex","description":"theorem measurable_historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (arm : Fin K) : Measurable (fun h => historyIndex initial ε u n h arm)","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-76f1a361232d","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1259,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (arm : Fin K) : Measurable (fun h => historyIndex initial ε u n h arm)","missing":[],"search":"measurable_historyindex banditrlproof.heavytail.sourcepolicy.measurable_historyindex theorem measurable_historyindex (initial : fin k) (ε u : ℝ) (n : ℕ) (arm : fin k) : measurable (fun h => historyindex initial ε u n h arm) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.measurable_nextArm","label":"measurable_nextArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.measurable_nextArm","description":"theorem measurable_nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) : Measurable (nextArm hK ε u n)","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-546edb353f42","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1260,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) : Measurable (nextArm hK ε u n)","missing":[],"search":"measurable_nextarm banditrlproof.heavytail.sourcepolicy.measurable_nextarm theorem measurable_nextarm (hk : 0 < k) (ε u : ℝ) (n : ℕ) : measurable (nextarm hk ε u n) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustAction","label":"robustAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustAction","description":"noncomputable def robustAction (hK : 0 < K) (ε u : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-7eaebec39969","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1261,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"noncomputable def robustAction (hK : 0 < K) (ε u : ℝ)","missing":[],"search":"robustaction banditrlproof.heavytail.sourcepolicy.robustaction noncomputable def robustaction (hk : 0 < k) (ε u : ℝ) definition compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustReward","label":"robustReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustReward","description":"noncomputable def robustReward (hK : 0 < K) (ε u : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-b2ec38957f72","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1262,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def robustReward (hK : 0 < K) (ε u : ℝ)","missing":[],"search":"robustreward banditrlproof.heavytail.sourcepolicy.robustreward noncomputable def robustreward (hk : 0 < k) (ε u : ℝ) definition compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.measurable_robustAction","label":"measurable_robustAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.measurable_robustAction","description":"theorem measurable_robustAction (hK : 0 < K) (ε u : ℝ) (t : ℕ) : Measurable (fun stream => robustAction hK ε u stream t)","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-5c00fd365a02","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1263,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_robustAction (hK : 0 < K) (ε u : ℝ) (t : ℕ) : Measurable (fun stream => robustAction hK ε u stream t)","missing":[],"search":"measurable_robustaction banditrlproof.heavytail.sourcepolicy.measurable_robustaction theorem measurable_robustaction (hk : 0 < k) (ε u : ℝ) (t : ℕ) : measurable (fun stream => robustaction hk ε u stream t) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustAction_initialization","label":"robustAction_initialization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustAction_initialization","description":"theorem robustAction_initialization (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : robustAction hK ε u stream t = UCB.initializationArm hK t","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-308f5efadd40","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1264,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustAction_initialization (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : robustAction hK ε u stream t = UCB.initializationArm hK t","missing":[],"search":"robustaction_initialization banditrlproof.heavytail.sourcepolicy.robustaction_initialization theorem robustaction_initialization (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (t : ℕ) (ht : t < k) : robustaction hk ε u stream t = ucb.initializationarm hk t theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustAction_maximizes","label":"robustAction_maximizes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustAction_maximizes","description":"The selected index is maximal on the actual observed history.","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-95fbab75b3e2","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1265,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustAction_maximizes (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (n : ℕ) (hn : K ≤ n+1) (arm : Fin K) : let h := ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n historyIndex (UCB.initializationArm hK 0) ε u n h arm ≤ historyIndex (UCB.initializationArm hK 0) ε u n h (robustAction hK ε u stream (n+1))","missing":[],"search":"robustaction_maximizes banditrlproof.heavytail.sourcepolicy.robustaction_maximizes the selected index is maximal on the actual observed history. theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean","label":"robustMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean","description":"noncomputable def robustMean (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-894a593859c5","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1266,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def robustMean (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : ℝ","missing":[],"search":"robustmean banditrlproof.heavytail.sourcepolicy.robustmean noncomputable def robustmean (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (t : ℕ) : ℝ definition compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_latent","label":"robustMean_latent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean_latent","description":"theorem robustMean_latent (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : robustMean hK ε u stream arm t = (∑ s ∈ Finset.range (pullCount (robustAction hK ε u stream) arm t), truncate (sampleThreshold ε u t s) (stream s arm)) / pullCount (robustAction hK ε u stream) arm t","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-d95c2ecf78b2","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1267,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_latent (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) : robustMean hK ε u stream arm t = (∑ s ∈ Finset.range (pullCount (robustAction hK ε u stream) arm t), truncate (sampleThreshold ε u t s) (stream s arm)) / pullCount (robustAction hK ε u stream) arm t","missing":[],"search":"robustmean_latent banditrlproof.heavytail.sourcepolicy.robustmean_latent theorem robustmean_latent (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (t : ℕ) : robustmean hk ε u stream arm t = (∑ s ∈ finset.range (pullcount (robustaction hk ε u stream) arm t), truncate (samplethreshold ε u t s) (stream s arm)) / pullcount (robustaction hk ε u stream) arm t theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_pullCount_pos","label":"robust_pullCount_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_pullCount_pos","description":"theorem robust_pullCount_pos (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) (ht : K ≤ t) : 0 < pullCount (robustAction hK ε u stream) arm t","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-ef83cfdc134f","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1268,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_pullCount_pos (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (t : ℕ) (ht : K ≤ t) : 0 < pullCount (robustAction hK ε u stream) arm t","missing":[],"search":"robust_pullcount_pos banditrlproof.heavytail.sourcepolicy.robust_pullcount_pos theorem robust_pullcount_pos (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (t : ℕ) (ht : k ≤ t) : 0 < pullcount (robustaction hk ε u stream) arm t theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_history","label":"robustMean_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean_history","description":"theorem robustMean_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyTruncatedMean (UCB.initializationArm hK 0) (sampleThreshold ε u (n+1)) n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1)","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-6f2e941c313b","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1269,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyTruncatedMean (UCB.initializationArm hK 0) (sampleThreshold ε u (n+1)) n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1)","missing":[],"search":"robustmean_history banditrlproof.heavytail.sourcepolicy.robustmean_history theorem robustmean_history (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (n : ℕ) : historytruncatedmean (ucb.initializationarm hk 0) (samplethreshold ε u (n+1)) n (armstreampolicy.history (ucb.initializationarm hk 0) (nextarm hk ε u) stream n) arm = robustmean hk ε u stream arm (n+1) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustIndex_history","label":"robustIndex_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustIndex_history","description":"theorem robustIndex_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyIndex (UCB.initializationArm hK 0) ε u n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1) + sourceConfidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) arm (n+1))","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-34b387664157","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1270,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustIndex_history (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (arm : Fin K) (n : ℕ) : historyIndex (UCB.initializationArm hK 0) ε u n (ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n) arm = robustMean hK ε u stream arm (n+1) + sourceConfidenceRadius ε u (n+1) (pullCount (robustAction hK ε u stream) arm (n+1))","missing":[],"search":"robustindex_history banditrlproof.heavytail.sourcepolicy.robustindex_history theorem robustindex_history (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (arm : fin k) (n : ℕ) : historyindex (ucb.initializationarm hk 0) ε u n (armstreampolicy.history (ucb.initializationarm hk 0) (nextarm hk ε u) stream n) arm = robustmean hk ε u stream arm (n+1) + sourceconfidenceradius ε u (n+1) (pullcount (robustaction hk ε u stream) arm (n+1)) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.arm_adaptive_upper_tail","label":"arm_adaptive_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.arm_adaptive_upper_tail","description":"theorem arm_adaptive_upper_tail (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ sourceConfidenceRadius ε u…","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-e22591b89fda","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1271,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arm_adaptive_upper_tail (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ sourceConfidenceRadius ε u t (count stream) ≤ (∑ s ∈ Finset.range (count stream), truncate (sampleThreshold ε u t s) (stream s arm)) / count stream - ∫ x, x ∂ν arm} ≤ t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"arm_adaptive_upper_tail banditrlproof.heavytail.sourcepolicy.arm_adaptive_upper_tail theorem arm_adaptive_upper_tail (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (count : ucb.armrewardstream k → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : integrable (fun x : ℝ => x) (ν arm)) (hm : integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (ucb.armstreammeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ sourceconfidenceradius ε u t (count stream) ≤ (∑ s ∈ finset.range (count stream), truncate (samplethreshold ε u t s) (stream s arm)) / count stream - ∫ x, x ∂ν arm} ≤ t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail","label":"robustMean_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail","description":"theorem robustMean_upper_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤…","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-2efd0fd4577e","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1272,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_upper_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ robustMean hK ε u stream arm t - ∫ x, x ∂ν arm} ≤ t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"robustmean_upper_tail banditrlproof.heavytail.sourcepolicy.robustmean_upper_tail theorem robustmean_upper_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (ε u : ℝ) (t : ℕ) (ht : k ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : integrable (fun x : ℝ => x) (ν arm)) (hm : integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (ucb.armstreammeasure ν).real {stream | sourceconfidenceradius ε u t (pullcount (robustaction hk ε u stream) arm t) ≤ robustmean hk ε u stream arm t - ∫ x, x ∂ν arm} ≤ t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.arm_adaptive_lower_tail","label":"arm_adaptive_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.arm_adaptive_lower_tail","description":"theorem arm_adaptive_lower_tail (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ sourceConfidenceRadius ε u…","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-0bfbb7f474a2","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1273,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:153"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arm_adaptive_lower_tail (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ sourceConfidenceRadius ε u t (count stream) ≤ (∫ x, x ∂ν arm) - (∑ s ∈ Finset.range (count stream), truncate (sampleThreshold ε u t s) (stream s arm)) / count stream} ≤ t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"arm_adaptive_lower_tail banditrlproof.heavytail.sourcepolicy.arm_adaptive_lower_tail theorem arm_adaptive_lower_tail (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (count : ucb.armrewardstream k → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : integrable (fun x : ℝ => x) (ν arm)) (hm : integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (ucb.armstreammeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ sourceconfidenceradius ε u t (count stream) ≤ (∫ x, x ∂ν arm) - (∑ s ∈ finset.range (count stream), truncate (samplethreshold ε u t s) (stream s arm)) / count stream} ≤ t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail","label":"robustMean_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail","description":"theorem robustMean_lower_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤…","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-3891a15e4770","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1274,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:176"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustMean_lower_tail (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (t : ℕ) (ht : K ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ (∫ x, x ∂ν arm) - robustMean hK ε u stream arm t} ≤ t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"robustmean_lower_tail banditrlproof.heavytail.sourcepolicy.robustmean_lower_tail theorem robustmean_lower_tail (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (ε u : ℝ) (t : ℕ) (ht : k ≤ t) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : integrable (fun x : ℝ => x) (ν arm)) (hm : integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (ucb.armstreammeasure ν).real {stream | sourceconfidenceradius ε u t (pullcount (robustaction hk ε u stream) arm t) ≤ (∫ x, x ∂ν arm) - robustmean hk ε u stream arm t} ≤ t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail_sum","label":"robustMean_upper_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail_sum","description":"Finite time budget for one signed event of the actual source-parameter policy.","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-6dda3c47c5d0","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1275,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:193"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem robustMean_upper_tail_sum (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (∑ t ∈ Finset.range T, (UCB.armStreamMeasure ν).real {stream | K ≤ t ∧ sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ robustMean hK ε u stream arm t - ∫ x, x ∂ν arm}) ≤ 2","missing":[],"search":"robustmean_upper_tail_sum banditrlproof.heavytail.sourcepolicy.robustmean_upper_tail_sum finite time budget for one signed event of the actual source-parameter policy. theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail_sum","label":"robustMean_lower_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail_sum","description":"Finite time budget for one signed event of the actual source-parameter policy.","url":"../modules/banditrlproof-algorithms-heavytailsourcepolicy/index.html#decl-8dd0c72b3190","parent":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","order":1276,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourcePolicy"],["Source","BanditRLProof/Algorithms/HeavyTailSourcePolicy.lean:210"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem robustMean_lower_tail_sum (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (ε u : ℝ) (T : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (∑ t ∈ Finset.range T, (UCB.armStreamMeasure ν).real {stream | K ≤ t ∧ sourceConfidenceRadius ε u t (pullCount (robustAction hK ε u stream) arm t) ≤ (∫ x, x ∂ν arm) - robustMean hK ε u stream arm t}) ≤ 2","missing":[],"search":"robustmean_lower_tail_sum banditrlproof.heavytail.sourcepolicy.robustmean_lower_tail_sum finite time budget for one signed event of the actual source-parameter policy. theorem compiled","shard":"modules/ae4833da8f0ce0f7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_expected_regret","label":"robust_expected_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_expected_regret","description":"Raw moments produce the entire causal algorithm-to-expected-pseudo-regret chain.","url":"../modules/banditrlproof-algorithms-heavytailsourceregret/index.html#decl-90ab9f96460c","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","order":1277,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceRegret"],["Source","BanditRLProof/Algorithms/HeavyTailSourceRegret.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem robust_expected_regret (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hm : ∀ a, Integrable (fun x : ℝ => |x|^(1+ε)) (ν a)) (hu : ∀ a, (∫ x, |x|^(1+ε) ∂ν a) ≤ u) : (∫ stream, realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T ∂UCB.armStreamMeasure ν) ≤ ∑ arm ∈ Finset.univ.filter (fun arm : Fin K => 0 < realMeanGap (realKernelMean ν) arm), realMeanGap (realKernelMean ν) arm * (gapBudget ε u (realMeanGap (realKernelMean ν) arm) T + 5)","missing":[],"search":"robust_expected_regret banditrlproof.heavytail.sourcepolicy.robust_expected_regret raw moments produce the entire causal algorithm-to-expected-pseudo-regret chain. theorem compiled","shard":"modules/c0b77d752e163bbc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_regret_integrable","label":"robust_regret_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.robust_regret_integrable","description":"theorem robust_regret_integrable (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) : Integrable (fun stream => realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T) (UCB.armStreamMeasure ν)","url":"../modules/banditrlproof-algorithms-heavytailsourceregret/index.html#decl-770efd57be6d","parent":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","order":1278,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailSourceRegret"],["Source","BanditRLProof/Algorithms/HeavyTailSourceRegret.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robust_regret_integrable (hK : 0 < K) (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (ε u : ℝ) (T : ℕ) : Integrable (fun stream => realMeanRegret (realKernelMean ν) (robustAction hK ε u stream) T) (UCB.armStreamMeasure ν)","missing":[],"search":"robust_regret_integrable banditrlproof.heavytail.sourcepolicy.robust_regret_integrable theorem robust_regret_integrable (hk : 0 < k) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (ε u : ℝ) (t : ℕ) : integrable (fun stream => realmeanregret (realkernelmean ν) (robustaction hk ε u stream) t) (ucb.armstreammeasure ν) theorem compiled","shard":"modules/c0b77d752e163bbc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.confidenceLog","label":"confidenceLog","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.confidenceLog","description":"noncomputable def confidenceLog (t : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-801c021af1ad","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1279,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceLog (t : ℕ) : ℝ","missing":[],"search":"confidencelog banditrlproof.heavytail.confidencelog noncomputable def confidencelog (t : ℕ) : ℝ definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold","label":"sampleThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold","description":"noncomputable def sampleThreshold (ε u : ℝ) (t s : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-d78937e405fe","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1280,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sampleThreshold (ε u : ℝ) (t s : ℕ) : ℝ","missing":[],"search":"samplethreshold banditrlproof.heavytail.samplethreshold noncomputable def samplethreshold (ε u : ℝ) (t s : ℕ) : ℝ definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.confidenceRadius","label":"confidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.confidenceRadius","description":"noncomputable def confidenceRadius (ε u : ℝ) (t count : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-ae996d5a690a","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1281,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceRadius (ε u : ℝ) (t count : ℕ) : ℝ","missing":[],"search":"confidenceradius banditrlproof.heavytail.confidenceradius noncomputable def confidenceradius (ε u : ℝ) (t count : ℕ) : ℝ definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.historyIndex","label":"historyIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.historyIndex","description":"noncomputable def historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) (arm : Fin K) : ℝ","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-50c48dac1567","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1282,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) (arm : Fin K) : ℝ","missing":[],"search":"historyindex banditrlproof.heavytail.historyindex noncomputable def historyindex (initial : fin k) (ε u : ℝ) (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) (arm : fin k) : ℝ definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.nextArm","label":"nextArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.nextArm","description":"noncomputable def nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : Fin K","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-456ed246b79f","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1283,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) (h : History.FinitePairHistory (Fin K) ℝ n) : Fin K","missing":[],"search":"nextarm banditrlproof.heavytail.nextarm noncomputable def nextarm (hk : 0 < k) (ε u : ℝ) (n : ℕ) (h : history.finitepairhistory (fin k) ℝ n) : fin k definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_historyIndex","label":"measurable_historyIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_historyIndex","description":"theorem measurable_historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (arm : Fin K) : Measurable (fun h => historyIndex initial ε u n h arm)","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-a7dfdf937996","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1284,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyIndex (initial : Fin K) (ε u : ℝ) (n : ℕ) (arm : Fin K) : Measurable (fun h => historyIndex initial ε u n h arm)","missing":[],"search":"measurable_historyindex banditrlproof.heavytail.measurable_historyindex theorem measurable_historyindex (initial : fin k) (ε u : ℝ) (n : ℕ) (arm : fin k) : measurable (fun h => historyindex initial ε u n h arm) theorem compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_nextArm","label":"measurable_nextArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_nextArm","description":"theorem measurable_nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) : Measurable (nextArm hK ε u n)","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-716424fb3391","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1285,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_nextArm (hK : 0 < K) (ε u : ℝ) (n : ℕ) : Measurable (nextArm hK ε u n)","missing":[],"search":"measurable_nextarm banditrlproof.heavytail.measurable_nextarm theorem measurable_nextarm (hk : 0 < k) (ε u : ℝ) (n : ℕ) : measurable (nextarm hk ε u n) theorem compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustAction","label":"robustAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustAction","description":"noncomputable def robustAction (hK : 0 < K) (ε u : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-94ed60edd98f","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1286,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def robustAction (hK : 0 < K) (ε u : ℝ)","missing":[],"search":"robustaction banditrlproof.heavytail.robustaction noncomputable def robustaction (hk : 0 < k) (ε u : ℝ) definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustReward","label":"robustReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustReward","description":"noncomputable def robustReward (hK : 0 < K) (ε u : ℝ)","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-0949cfaa9546","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1287,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def robustReward (hK : 0 < K) (ε u : ℝ)","missing":[],"search":"robustreward banditrlproof.heavytail.robustreward noncomputable def robustreward (hk : 0 < k) (ε u : ℝ) definition compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_robustAction","label":"measurable_robustAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_robustAction","description":"theorem measurable_robustAction (hK : 0 < K) (ε u : ℝ) (t : ℕ) : Measurable (fun stream => robustAction hK ε u stream t)","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-3fcac8ac1868","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1288,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_robustAction (hK : 0 < K) (ε u : ℝ) (t : ℕ) : Measurable (fun stream => robustAction hK ε u stream t)","missing":[],"search":"measurable_robustaction banditrlproof.heavytail.measurable_robustaction theorem measurable_robustaction (hk : 0 < k) (ε u : ℝ) (t : ℕ) : measurable (fun stream => robustaction hk ε u stream t) theorem compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustAction_initialization","label":"robustAction_initialization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustAction_initialization","description":"theorem robustAction_initialization (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : robustAction hK ε u stream t = UCB.initializationArm hK t","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-aaaceaeaed57","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1289,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustAction_initialization (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (t : ℕ) (ht : t < K) : robustAction hK ε u stream t = UCB.initializationArm hK t","missing":[],"search":"robustaction_initialization banditrlproof.heavytail.robustaction_initialization theorem robustaction_initialization (hk : 0 < k) (ε u : ℝ) (stream : ucb.armrewardstream k) (t : ℕ) (ht : t < k) : robustaction hk ε u stream t = ucb.initializationarm hk t theorem compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.robustAction_maximizes","label":"robustAction_maximizes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.robustAction_maximizes","description":"The selected index is maximal on the actual observed history.","url":"../modules/banditrlproof-algorithms-heavytailucb/index.html#decl-c43df1488516","parent":"module:BanditRLProof.Algorithms.HeavyTailUCB","order":1290,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.HeavyTailUCB"],["Source","BanditRLProof/Algorithms/HeavyTailUCB.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem robustAction_maximizes (hK : 0 < K) (ε u : ℝ) (stream : UCB.ArmRewardStream K) (n : ℕ) (hn : K ≤ n+1) (arm : Fin K) : let h := ArmStreamPolicy.history (UCB.initializationArm hK 0) (nextArm hK ε u) stream n historyIndex (UCB.initializationArm hK 0) ε u n h arm ≤ historyIndex (UCB.initializationArm hK 0) ε u n h (robustAction hK ε u stream (n+1))","missing":[],"search":"robustaction_maximizes banditrlproof.heavytail.robustaction_maximizes the selected index is maximal on the actual observed history. theorem compiled","shard":"modules/5e16c61dc9dcdbd7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.IsBernoulliParameter","label":"IsBernoulliParameter","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.IsBernoulliParameter","description":"The closed unit interval predicate used by every KL-UCB parameter.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-9fc855c29a30","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1291,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def IsBernoulliParameter (p : Real) : Prop","missing":[],"search":"isbernoulliparameter banditrlproof.klucb.isbernoulliparameter the closed unit interval predicate used by every kl-ucb parameter. definition compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLCore","label":"bernoulliKLCore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLCore","description":"The finite analytic Bernoulli relative-entropy expression. It is used only when the second parameter is strictly between zero and one.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-ac4843545118","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1292,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bernoulliKLCore (p q : Real) : Real","missing":[],"search":"bernoulliklcore banditrlproof.klucb.bernoulliklcore the finite analytic bernoulli relative-entropy expression. it is used only when the second parameter is strictly between zero and one. definition compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL","label":"bernoulliKL","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL","description":"Bernoulli relative entropy with exact support/endpoint conventions. parameters outside `[0,1]` have value `top`; `d(0,0)=d(1,1)=0`; `d(p,0)=top` for `p>0` and `d(p,1)=top` for `p<1`; otherwise the ordinary finite logarithmic expression is used.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-8f99f3947abe","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1293,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","stochastic-finite"]],"statement":"noncomputable def bernoulliKL (p q : Real) : ENNReal","missing":[],"search":"bernoullikl banditrlproof.klucb.bernoullikl bernoulli relative entropy with exact support/endpoint conventions. parameters outside `[0,1]` have value `top`; `d(0,0)=d(1,1)=0`; `d(p,0)=top` for `p>0` and `d(p,1)=top` for `p<1`; otherwise the ordinary finite logarithmic expression is used. definition compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["stochastic-finite"]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_nonneg","label":"bernoulliKL_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_nonneg","description":"theorem bernoulliKL_nonneg (p q : Real) : 0 <= bernoulliKL p q","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-5b3b86acc12b","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1294,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_nonneg (p q : Real) : 0 <= bernoulliKL p q","missing":[],"search":"bernoullikl_nonneg banditrlproof.klucb.bernoullikl_nonneg theorem bernoullikl_nonneg (p q : real) : 0 <= bernoullikl p q theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_of_not_left","label":"bernoulliKL_eq_top_of_not_left","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_eq_top_of_not_left","description":"theorem bernoulliKL_eq_top_of_not_left {p q : Real} (hp : ¬ IsBernoulliParameter p) : bernoulliKL p q = ⊤","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-33a6e7778907","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1295,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_eq_top_of_not_left {p q : Real} (hp : ¬ IsBernoulliParameter p) : bernoulliKL p q = ⊤","missing":[],"search":"bernoullikl_eq_top_of_not_left banditrlproof.klucb.bernoullikl_eq_top_of_not_left theorem bernoullikl_eq_top_of_not_left {p q : real} (hp : ¬ isbernoulliparameter p) : bernoullikl p q = ⊤ theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_of_not_right","label":"bernoulliKL_eq_top_of_not_right","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_eq_top_of_not_right","description":"theorem bernoulliKL_eq_top_of_not_right {p q : Real} (hq : ¬ IsBernoulliParameter q) : bernoulliKL p q = ⊤","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-7a9048ebc6d9","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1296,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_eq_top_of_not_right {p q : Real} (hq : ¬ IsBernoulliParameter q) : bernoulliKL p q = ⊤","missing":[],"search":"bernoullikl_eq_top_of_not_right banditrlproof.klucb.bernoullikl_eq_top_of_not_right theorem bernoullikl_eq_top_of_not_right {p q : real} (hq : ¬ isbernoulliparameter q) : bernoullikl p q = ⊤ theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_zero_zero","label":"bernoulliKL_zero_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_zero_zero","description":"theorem bernoulliKL_zero_zero : bernoulliKL 0 0 = 0","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-701861f1cc9d","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1297,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_zero_zero : bernoulliKL 0 0 = 0","missing":[],"search":"bernoullikl_zero_zero banditrlproof.klucb.bernoullikl_zero_zero theorem bernoullikl_zero_zero : bernoullikl 0 0 = 0 theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_one_one","label":"bernoulliKL_one_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_one_one","description":"theorem bernoulliKL_one_one : bernoulliKL 1 1 = 0","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-d00cbca89d69","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1298,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_one_one : bernoulliKL 1 1 = 0","missing":[],"search":"bernoullikl_one_one banditrlproof.klucb.bernoullikl_one_one theorem bernoullikl_one_one : bernoullikl 1 1 = 0 theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_right_zero","label":"bernoulliKL_right_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_right_zero","description":"theorem bernoulliKL_right_zero {p : Real} (hp : IsBernoulliParameter p) : bernoulliKL p 0 = if p = 0 then 0 else ⊤","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-3c1913a07cd0","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1299,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_right_zero {p : Real} (hp : IsBernoulliParameter p) : bernoulliKL p 0 = if p = 0 then 0 else ⊤","missing":[],"search":"bernoullikl_right_zero banditrlproof.klucb.bernoullikl_right_zero theorem bernoullikl_right_zero {p : real} (hp : isbernoulliparameter p) : bernoullikl p 0 = if p = 0 then 0 else ⊤ theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_right_one","label":"bernoulliKL_right_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_right_one","description":"theorem bernoulliKL_right_one {p : Real} (hp : IsBernoulliParameter p) : bernoulliKL p 1 = if p = 1 then 0 else ⊤","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-8dce7cbd5522","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1300,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_right_one {p : Real} (hp : IsBernoulliParameter p) : bernoulliKL p 1 = if p = 1 then 0 else ⊤","missing":[],"search":"bernoullikl_right_one banditrlproof.klucb.bernoullikl_right_one theorem bernoullikl_right_one {p : real} (hp : isbernoulliparameter p) : bernoullikl p 1 = if p = 1 then 0 else ⊤ theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_right_zero","label":"bernoulliKL_eq_top_right_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_eq_top_right_zero","description":"theorem bernoulliKL_eq_top_right_zero {p : Real} (hp : IsBernoulliParameter p) (hp0 : p ≠ 0) : bernoulliKL p 0 = ⊤","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-2c6153573601","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1301,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_eq_top_right_zero {p : Real} (hp : IsBernoulliParameter p) (hp0 : p ≠ 0) : bernoulliKL p 0 = ⊤","missing":[],"search":"bernoullikl_eq_top_right_zero banditrlproof.klucb.bernoullikl_eq_top_right_zero theorem bernoullikl_eq_top_right_zero {p : real} (hp : isbernoulliparameter p) (hp0 : p ≠ 0) : bernoullikl p 0 = ⊤ theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_right_one","label":"bernoulliKL_eq_top_right_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_eq_top_right_one","description":"theorem bernoulliKL_eq_top_right_one {p : Real} (hp : IsBernoulliParameter p) (hp1 : p ≠ 1) : bernoulliKL p 1 = ⊤","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-cb1d03cdc301","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1302,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_eq_top_right_one {p : Real} (hp : IsBernoulliParameter p) (hp1 : p ≠ 1) : bernoulliKL p 1 = ⊤","missing":[],"search":"bernoullikl_eq_top_right_one banditrlproof.klucb.bernoullikl_eq_top_right_one theorem bernoullikl_eq_top_right_one {p : real} (hp : isbernoulliparameter p) (hp1 : p ≠ 1) : bernoullikl p 1 = ⊤ theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_zero_left_of_interior","label":"bernoulliKL_zero_left_of_interior","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_zero_left_of_interior","description":"theorem bernoulliKL_zero_left_of_interior {q : Real} (hq0 : 0 < q) (hq1 : q < 1) : bernoulliKL 0 q = ENNReal.ofReal (Real.log (1 / (1 - q)))","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-f66642f4706f","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1303,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_zero_left_of_interior {q : Real} (hq0 : 0 < q) (hq1 : q < 1) : bernoulliKL 0 q = ENNReal.ofReal (Real.log (1 / (1 - q)))","missing":[],"search":"bernoullikl_zero_left_of_interior banditrlproof.klucb.bernoullikl_zero_left_of_interior theorem bernoullikl_zero_left_of_interior {q : real} (hq0 : 0 < q) (hq1 : q < 1) : bernoullikl 0 q = ennreal.ofreal (real.log (1 / (1 - q))) theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_one_left_of_interior","label":"bernoulliKL_one_left_of_interior","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_one_left_of_interior","description":"theorem bernoulliKL_one_left_of_interior {q : Real} (hq0 : 0 < q) (hq1 : q < 1) : bernoulliKL 1 q = ENNReal.ofReal (Real.log (1 / q))","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-95c0d4b8eaeb","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1304,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:116"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_one_left_of_interior {q : Real} (hq0 : 0 < q) (hq1 : q < 1) : bernoulliKL 1 q = ENNReal.ofReal (Real.log (1 / q))","missing":[],"search":"bernoullikl_one_left_of_interior banditrlproof.klucb.bernoullikl_one_left_of_interior theorem bernoullikl_one_left_of_interior {q : real} (hq0 : 0 < q) (hq1 : q < 1) : bernoullikl 1 q = ennreal.ofreal (real.log (1 / q)) theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_self","label":"bernoulliKLCore_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLCore_self","description":"The finite Bernoulli expression vanishes on the diagonal away from the singular endpoints.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-7ddd835c8384","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1305,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:129"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKLCore_self {p : Real} (hp0 : p ≠ 0) (hp1 : p ≠ 1) : bernoulliKLCore p p = 0","missing":[],"search":"bernoulliklcore_self banditrlproof.klucb.bernoulliklcore_self the finite bernoulli expression vanishes on the diagonal away from the singular endpoints. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_eq_klFun","label":"bernoulliKLCore_eq_klFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLCore_eq_klFun","description":"The analytic binary KL expression is the two-atom `klFun` integral.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-e47976cdf5c6","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1306,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:135"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKLCore_eq_klFun {p q : Real} (hq0 : q ≠ 0) (hq1 : q ≠ 1) : bernoulliKLCore p q = q * InformationTheory.klFun (p / q) + (1 - q) * InformationTheory.klFun ((1 - p) / (1 - q))","missing":[],"search":"bernoulliklcore_eq_klfun banditrlproof.klucb.bernoulliklcore_eq_klfun the analytic binary kl expression is the two-atom `klfun` integral. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_nonneg","label":"bernoulliKLCore_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLCore_nonneg","description":"Nontrivial nonnegativity of the finite logarithmic expression. This is stronger than the order-theoretic nonnegativity of its `ENNReal` wrapper.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-e273bb31e167","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1307,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKLCore_nonneg {p q : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : 0 <= bernoulliKLCore p q","missing":[],"search":"bernoulliklcore_nonneg banditrlproof.klucb.bernoulliklcore_nonneg nontrivial nonnegativity of the finite logarithmic expression. this is stronger than the order-theoretic nonnegativity of its `ennreal` wrapper. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_of_interior","label":"bernoulliKL_eq_of_interior","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_eq_of_interior","description":"Interior parameters expose the finite analytic expression without truncation: its real nonnegativity has already been proved.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-1e35dc95378b","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1308,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:163"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_eq_of_interior {p q : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : bernoulliKL p q = ENNReal.ofReal (bernoulliKLCore p q)","missing":[],"search":"bernoullikl_eq_of_interior banditrlproof.klucb.bernoullikl_eq_of_interior interior parameters expose the finite analytic expression without truncation: its real nonnegativity has already been proved. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLExpanded","label":"bernoulliKLExpanded","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLExpanded","description":"Algebraically expanded finite KL, convenient for differentiation in the second parameter.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-658d9233fdbc","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1309,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bernoulliKLExpanded (p q : Real) : Real","missing":[],"search":"bernoulliklexpanded banditrlproof.klucb.bernoulliklexpanded algebraically expanded finite kl, convenient for differentiation in the second parameter. definition compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_eq_expanded","label":"bernoulliKLCore_eq_expanded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLCore_eq_expanded","description":"theorem bernoulliKLCore_eq_expanded {p q : Real} (hp0 : p ≠ 0) (hp1 : p ≠ 1) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : bernoulliKLCore p q = bernoulliKLExpanded p q","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-a5b1877b476d","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1310,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:179"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKLCore_eq_expanded {p q : Real} (hp0 : p ≠ 0) (hp1 : p ≠ 1) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : bernoulliKLCore p q = bernoulliKLExpanded p q","missing":[],"search":"bernoulliklcore_eq_expanded banditrlproof.klucb.bernoulliklcore_eq_expanded theorem bernoulliklcore_eq_expanded {p q : real} (hp0 : p ≠ 0) (hp1 : p ≠ 1) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : bernoulliklcore p q = bernoulliklexpanded p q theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.hasDerivAt_bernoulliKLExpanded_right","label":"hasDerivAt_bernoulliKLExpanded_right","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.hasDerivAt_bernoulliKLExpanded_right","description":"theorem hasDerivAt_bernoulliKLExpanded_right (p q : Real) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : HasDerivAt (fun r => bernoulliKLExpanded p r) ((q - p) / (q * (1 - q))) q","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-69f550988d98","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1311,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:189"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_bernoulliKLExpanded_right (p q : Real) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : HasDerivAt (fun r => bernoulliKLExpanded p r) ((q - p) / (q * (1 - q))) q","missing":[],"search":"hasderivat_bernoulliklexpanded_right banditrlproof.klucb.hasderivat_bernoulliklexpanded_right theorem hasderivat_bernoulliklexpanded_right (p q : real) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : hasderivat (fun r => bernoulliklexpanded p r) ((q - p) / (q * (1 - q))) q theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.half_sq_sub_le_bernoulliKLCore","label":"half_sq_sub_le_bernoulliKLCore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.half_sq_sub_le_bernoulliKLCore","description":"A conservative binary Pinsker inequality. The constant `1/2` is weaker than the sharp natural-log constant `2`, but is sufficient to invert every KL-UCB confidence set without changing the KL score itself.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-b9aa274a9d11","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1312,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:209"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem half_sq_sub_le_bernoulliKLCore {p q : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : (1 / 2 : Real) * (p - q) ^ 2 <= bernoulliKLCore p q","missing":[],"search":"half_sq_sub_le_bernoulliklcore banditrlproof.klucb.half_sq_sub_le_bernoulliklcore a conservative binary pinsker inequality. the constant `1/2` is weaker than the sharp natural-log constant `2`, but is sufficient to invert every kl-ucb confidence set without changing the kl score itself. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_le_sq_div","label":"bernoulliKLCore_le_sq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKLCore_le_sq_div","description":"On an interior reference mean, binary KL is controlled by the squared deviation divided by the Bernoulli variance denominator. This is the bridge from the repository's bounded-reward empirical-mean tails to a genuine KL confidence event.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-7a3451b30ec3","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1313,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:305"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKLCore_le_sq_div {p q : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : bernoulliKLCore p q <= (p - q) ^ 2 / (q * (1 - q))","missing":[],"search":"bernoulliklcore_le_sq_div banditrlproof.klucb.bernoulliklcore_le_sq_div on an interior reference mean, binary kl is controlled by the squared deviation divided by the bernoulli variance denominator. this is the bridge from the repository's bounded-reward empirical-mean tails to a genuine kl confidence event. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.ennnreal_half_sq_sub_le_bernoulliKL","label":"ennnreal_half_sq_sub_le_bernoulliKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.ennnreal_half_sq_sub_le_bernoulliKL","description":"theorem ennnreal_half_sq_sub_le_bernoulliKL {p q : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : ENNReal.ofReal ((1 / 2 : Real) * (p - q) ^ 2) <= bernoulliKL p q","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-844077b728e8","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1314,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:356"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ennnreal_half_sq_sub_le_bernoulliKL {p q : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : ENNReal.ofReal ((1 / 2 : Real) * (p - q) ^ 2) <= bernoulliKL p q","missing":[],"search":"ennnreal_half_sq_sub_le_bernoullikl banditrlproof.klucb.ennnreal_half_sq_sub_le_bernoullikl theorem ennnreal_half_sq_sub_le_bernoullikl {p q : real} (hp : isbernoulliparameter p) (hq0 : 0 < q) (hq1 : q < 1) : ennreal.ofreal ((1 / 2 : real) * (p - q) ^ 2) <= bernoullikl p q theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_le_of_sq_le","label":"bernoulliKL_le_of_sq_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_le_of_sq_le","description":"theorem bernoulliKL_le_of_sq_le {p q budget : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) (hsq : (p - q) ^ 2 / (q * (1 - q)) <= budget) : bernoulliKL p q <= ENNReal.ofReal budget","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-a29538094880","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1315,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:363"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_le_of_sq_le {p q budget : Real} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) (hsq : (p - q) ^ 2 / (q * (1 - q)) <= budget) : bernoulliKL p q <= ENNReal.ofReal budget","missing":[],"search":"bernoullikl_le_of_sq_le banditrlproof.klucb.bernoullikl_le_of_sq_le theorem bernoullikl_le_of_sq_le {p q budget : real} (hp : isbernoulliparameter p) (hq0 : 0 < q) (hq1 : q < 1) (hsq : (p - q) ^ 2 / (q * (1 - q)) <= budget) : bernoullikl p q <= ennreal.ofreal budget theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.bernoulliKL_self","label":"bernoulliKL_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.bernoulliKL_self","description":"theorem bernoulliKL_self {p : Real} (hp : IsBernoulliParameter p) : bernoulliKL p p = 0","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-e19577cf21e8","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1316,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:373"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKL_self {p : Real} (hp : IsBernoulliParameter p) : bernoulliKL p p = 0","missing":[],"search":"bernoullikl_self banditrlproof.klucb.bernoullikl_self theorem bernoullikl_self {p : real} (hp : isbernoulliparameter p) : bernoullikl p p = 0 theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.continuousAt_bernoulliKLCore_right","label":"continuousAt_bernoulliKLCore_right","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.continuousAt_bernoulliKLCore_right","description":"On the nonsingular right-parameter domain the finite expression is continuous. This is the local analytic regularity used by later inversion leaves; endpoint singularities remain in `ENNReal`.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-aa647f7a3f3e","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1317,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:387"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem continuousAt_bernoulliKLCore_right (p q : Real) (hq0 : q ≠ 0) (hq1 : q ≠ 1) : ContinuousAt (fun r => bernoulliKLCore p r) q","missing":[],"search":"continuousat_bernoulliklcore_right banditrlproof.klucb.continuousat_bernoulliklcore_right on the nonsingular right-parameter domain the finite expression is continuous. this is the local analytic regularity used by later inversion leaves; endpoint singularities remain in `ennreal`. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.confidenceSet","label":"confidenceSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.confidenceSet","description":"KL confidence set at one empirical mean, pull count, and exploration budget.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-43b831ebd183","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1318,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:408"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def confidenceSet (empiricalMean : Real) (count : Nat) (budget : Real) : Set Real","missing":[],"search":"confidenceset banditrlproof.klucb.confidenceset kl confidence set at one empirical mean, pull count, and exploration budget. definition compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.index","label":"index","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.index","description":"The KL-UCB index is the supremum of its confidence set.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-734b1693685d","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1319,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:415"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def index (empiricalMean : Real) (count : Nat) (budget : Real) : Real","missing":[],"search":"index banditrlproof.klucb.index the kl-ucb index is the supremum of its confidence set. definition compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.confidenceSet_bddAbove","label":"confidenceSet_bddAbove","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.confidenceSet_bddAbove","description":"theorem confidenceSet_bddAbove (empiricalMean : Real) (count : Nat) (budget : Real) : BddAbove (confidenceSet empiricalMean count budget)","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-5e7a3f63bd61","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1320,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:419"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem confidenceSet_bddAbove (empiricalMean : Real) (count : Nat) (budget : Real) : BddAbove (confidenceSet empiricalMean count budget)","missing":[],"search":"confidenceset_bddabove banditrlproof.klucb.confidenceset_bddabove theorem confidenceset_bddabove (empiricalmean : real) (count : nat) (budget : real) : bddabove (confidenceset empiricalmean count budget) theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.mem_confidenceSet_self","label":"mem_confidenceSet_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.mem_confidenceSet_self","description":"theorem mem_confidenceSet_self {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : p ∈ confidenceSet p count budget","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-92c12908fd3e","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1321,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:426"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_confidenceSet_self {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : p ∈ confidenceSet p count budget","missing":[],"search":"mem_confidenceset_self banditrlproof.klucb.mem_confidenceset_self theorem mem_confidenceset_self {p : real} (hp : isbernoulliparameter p) (count : nat) {budget : real} (hbudget : 0 <= budget) : p ∈ confidenceset p count budget theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.mem_confidenceSet_of_natCast_mul_core_le","label":"mem_confidenceSet_of_natCast_mul_core_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.mem_confidenceSet_of_natCast_mul_core_le","description":"Real finite-KL arithmetic is sufficient to establish exact membership in the `ENNReal` confidence set.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-7c8ba4ae22d5","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1322,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:436"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_confidenceSet_of_natCast_mul_core_le {p q budget : Real} {count : Nat} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) (hbudget : (count : Real) * bernoulliKLCore p q <= budget) : q ∈ confidenceSet p count budget","missing":[],"search":"mem_confidenceset_of_natcast_mul_core_le banditrlproof.klucb.mem_confidenceset_of_natcast_mul_core_le real finite-kl arithmetic is sufficient to establish exact membership in the `ennreal` confidence set. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.natCast_mul_half_sq_sub_le_budget_of_mem","label":"natCast_mul_half_sq_sub_le_budget_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.natCast_mul_half_sq_sub_le_budget_of_mem","description":"theorem natCast_mul_half_sq_sub_le_budget_of_mem {p q budget : Real} {count : Nat} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) (hbudget : 0 <= budget) (hmem : q ∈ confidenceSet p count budget) : (count : Real) * ((1 / 2 : Real) * (p - q) ^ 2) <= budget","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-23fe60141295","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1323,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:448"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem natCast_mul_half_sq_sub_le_budget_of_mem {p q budget : Real} {count : Nat} (hp : IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) (hbudget : 0 <= budget) (hmem : q ∈ confidenceSet p count budget) : (count : Real) * ((1 / 2 : Real) * (p - q) ^ 2) <= budget","missing":[],"search":"natcast_mul_half_sq_sub_le_budget_of_mem banditrlproof.klucb.natcast_mul_half_sq_sub_le_budget_of_mem theorem natcast_mul_half_sq_sub_le_budget_of_mem {p q budget : real} {count : nat} (hp : isbernoulliparameter p) (hq0 : 0 < q) (hq1 : q < 1) (hbudget : 0 <= budget) (hmem : q ∈ confidenceset p count budget) : (count : real) * ((1 / 2 : real) * (p - q) ^ 2) <= budget theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.confidenceSet_nonempty","label":"confidenceSet_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.confidenceSet_nonempty","description":"theorem confidenceSet_nonempty {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : (confidenceSet p count budget).Nonempty","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-8b4a02acf103","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1324,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:463"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem confidenceSet_nonempty {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : (confidenceSet p count budget).Nonempty","missing":[],"search":"confidenceset_nonempty banditrlproof.klucb.confidenceset_nonempty theorem confidenceset_nonempty {p : real} (hp : isbernoulliparameter p) (count : nat) {budget : real} (hbudget : 0 <= budget) : (confidenceset p count budget).nonempty theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.index_le_one","label":"index_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.index_le_one","description":"theorem index_le_one {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : index p count budget <= 1","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-218c1b42bac2","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1325,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:469"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem index_le_one {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : index p count budget <= 1","missing":[],"search":"index_le_one banditrlproof.klucb.index_le_one theorem index_le_one {p : real} (hp : isbernoulliparameter p) (count : nat) {budget : real} (hbudget : 0 <= budget) : index p count budget <= 1 theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.index_nonneg","label":"index_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.index_nonneg","description":"theorem index_nonneg {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : 0 <= index p count budget","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-eb97a6d06977","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1326,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:477"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem index_nonneg {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : 0 <= index p count budget","missing":[],"search":"index_nonneg banditrlproof.klucb.index_nonneg theorem index_nonneg {p : real} (hp : isbernoulliparameter p) (count : nat) {budget : real} (hbudget : 0 <= budget) : 0 <= index p count budget theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.index_mem_Icc","label":"index_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.index_mem_Icc","description":"theorem index_mem_Icc {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : index p count budget ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-7f8c12a7238f","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1327,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:486"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem index_mem_Icc {p : Real} (hp : IsBernoulliParameter p) (count : Nat) {budget : Real} (hbudget : 0 <= budget) : index p count budget ∈ Set.Icc (0 : Real) 1","missing":[],"search":"index_mem_icc banditrlproof.klucb.index_mem_icc theorem index_mem_icc {p : real} (hp : isbernoulliparameter p) (count : nat) {budget : real} (hbudget : 0 <= budget) : index p count budget ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.mem_confidenceSet_zero_iff","label":"mem_confidenceSet_zero_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.mem_confidenceSet_zero_iff","description":"Zero empirical count makes every unit-interval parameter feasible. This is the explicit KL-UCB zero-count convention.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-757d1df232f2","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1328,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:494"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_confidenceSet_zero_iff (p q budget : Real) (hbudget : 0 <= budget) : q ∈ confidenceSet p 0 budget ↔ IsBernoulliParameter q","missing":[],"search":"mem_confidenceset_zero_iff banditrlproof.klucb.mem_confidenceset_zero_iff zero empirical count makes every unit-interval parameter feasible. this is the explicit kl-ucb zero-count convention. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.index_zero_count","label":"index_zero_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.index_zero_count","description":"theorem index_zero_count {p : Real} (hp : IsBernoulliParameter p) {budget : Real} (hbudget : 0 <= budget) : index p 0 budget = 1","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-4198c0c45fcc","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1329,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:499"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem index_zero_count {p : Real} (hp : IsBernoulliParameter p) {budget : Real} (hbudget : 0 <= budget) : index p 0 budget = 1","missing":[],"search":"index_zero_count banditrlproof.klucb.index_zero_count theorem index_zero_count {p : real} (hp : isbernoulliparameter p) {budget : real} (hbudget : 0 <= budget) : index p 0 budget = 1 theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.le_index_of_mem_confidenceSet","label":"le_index_of_mem_confidenceSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.le_index_of_mem_confidenceSet","description":"Membership of the true mean implies KL optimism.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-20ac99555da6","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1330,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:511"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem le_index_of_mem_confidenceSet {p q : Real} {count : Nat} {budget : Real} (hq : q ∈ confidenceSet p count budget) : q <= index p count budget","missing":[],"search":"le_index_of_mem_confidenceset banditrlproof.klucb.le_index_of_mem_confidenceset membership of the true mean implies kl optimism. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.exists_mem_confidenceSet_of_lt_index","label":"exists_mem_confidenceSet_of_lt_index","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.exists_mem_confidenceSet_of_lt_index","description":"A strict lower level below the supremum has a genuine feasible witness. No unproved claim that the supremum itself belongs to the set is used.","url":"../modules/banditrlproof-algorithms-klucbbernoulli/index.html#decl-8e974a5e1d38","parent":"module:BanditRLProof.Algorithms.KLUCBBernoulli","order":1331,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBBernoulli"],["Source","BanditRLProof/Algorithms/KLUCBBernoulli.lean:520"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_mem_confidenceSet_of_lt_index {p level : Real} {count : Nat} {budget : Real} (hp : IsBernoulliParameter p) (hbudget : 0 <= budget) (hlevel : level < index p count budget) : ∃ q ∈ confidenceSet p count budget, level < q","missing":[],"search":"exists_mem_confidenceset_of_lt_index banditrlproof.klucb.exists_mem_confidenceset_of_lt_index a strict lower level below the supremum has a genuine feasible witness. no unproved claim that the supremum itself belongs to the set is used. theorem compiled","shard":"modules/b1dad9e08fe974cd.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedBudgetAt","label":"generatedBudgetAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedBudgetAt","description":"KL exploration budget on a realized generated prefix. The margin is a known regularity contract, while the empirical count and radius are computed from the actual action/reward trace.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-2de4a3f0c998","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1332,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def generatedBudgetAt {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (sigma2 : NNReal) (delta margin : Real) (omega : Omega) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"generatedbudgetat banditrlproof.klucb.generatedbudgetat kl exploration budget on a realized generated prefix. the margin is a known regularity contract, while the empirical count and radius are computed from the actual action/reward trace. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedBudgetAt_nonneg","label":"generatedBudgetAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedBudgetAt_nonneg","description":"theorem generatedBudgetAt_nonneg {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmargin1 : margin < 1) (omega : Omega) (t : Nat) (arm : Fin K) : 0 <= generatedBudgetAt action sigma2 delta margin omega t arm","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-4a01a14fa5c5","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1333,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedBudgetAt_nonneg {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmargin1 : margin < 1) (omega : Omega) (t : Nat) (arm : Fin K) : 0 <= generatedBudgetAt action sigma2 delta margin omega t arm","missing":[],"search":"generatedbudgetat_nonneg banditrlproof.klucb.generatedbudgetat_nonneg theorem generatedbudgetat_nonneg {omega : type} {k : nat} (action : omega -> actiontrace (fin k)) (sigma2 : nnreal) (delta margin : real) (hmargin0 : 0 < margin) (hmargin1 : margin < 1) (omega : omega) (t : nat) (arm : fin k) : 0 <= generatedbudgetat action sigma2 delta margin omega t arm theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedIndexAt","label":"generatedIndexAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedIndexAt","description":"Genuine Bernoulli-KL index on one generated prefix.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-a373474e0a80","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1334,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def generatedIndexAt {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta margin : Real) (omega : Omega) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"generatedindexat banditrlproof.klucb.generatedindexat genuine bernoulli-kl index on one generated prefix. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.historyIndex","label":"historyIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.historyIndex","description":"KL score reconstructed from a finite generated pair history.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-96542d879a06","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1335,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def historyIndex {K : Nat} (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (arm : Fin K) : Real","missing":[],"search":"historyindex banditrlproof.klucb.historyindex kl score reconstructed from a finite generated pair history. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measurable_historyIndex","label":"measurable_historyIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measurable_historyIndex","description":"theorem measurable_historyIndex {K : Nat} (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (t : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Rat t => historyIndex sigma2 delta margin defaultAction t history arm)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-556dc4d5ccee","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1336,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyIndex {K : Nat} (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (t : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Rat t => historyIndex sigma2 delta margin defaultAction t history arm)","missing":[],"search":"measurable_historyindex banditrlproof.klucb.measurable_historyindex theorem measurable_historyindex {k : nat} (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (t : nat) (arm : fin k) : measurable (fun history : history.finitepairhistory (fin k) rat t => historyindex sigma2 delta margin defaultaction t history arm) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.historyNextArm","label":"historyNextArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.historyNextArm","description":"Round-robin initialization followed by KL-index maximization.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-3e2e01da1983","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1337,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:87"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def historyNextArm {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) : Fin K","missing":[],"search":"historynextarm banditrlproof.klucb.historynextarm round-robin initialization followed by kl-index maximization. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.historyIndex_le_nextArm_of_K_le","label":"historyIndex_le_nextArm_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.historyIndex_le_nextArm_of_K_le","description":"theorem historyIndex_le_nextArm_of_K_le {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (ht : K <= t) (arm : Fin K) : historyIndex sigma2 delta margin defaultAction t history arm <= historyIndex sigma2 delta margin defaultAction t history (historyNextArm hK sigma2 delta margin defaultAction t history)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-5c2afd949491","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1338,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem historyIndex_le_nextArm_of_K_le {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (ht : K <= t) (arm : Fin K) : historyIndex sigma2 delta margin defaultAction t history arm <= historyIndex sigma2 delta margin defaultAction t history (historyNextArm hK sigma2 delta margin defaultAction t history)","missing":[],"search":"historyindex_le_nextarm_of_k_le banditrlproof.klucb.historyindex_le_nextarm_of_k_le theorem historyindex_le_nextarm_of_k_le {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (t : nat) (history : history.finitepairhistory (fin k) rat t) (ht : k <= t) (arm : fin k) : historyindex sigma2 delta margin defaultaction t history arm <= historyindex sigma2 delta margin defaultaction t history (historynextarm hk sigma2 delta margin defaultaction t history) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.pairHistory","label":"pairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.pairHistory","description":"Pair-history reconstruction for the KL-UCB policy.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-7a767d26f373","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1339,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def pairHistory {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) : (n : Nat) -> History.FiniteRewardHistory Rat n -> History.FinitePairHistory (Fin K) Rat n | 0, rewardHistory => fun i => (defaultAction, rewardHistory i) | n + 1, rewardHistory => let previousRewardHistory : History.FiniteRewardHistory Rat n := fun i => rewardHistory ⟨i.1, Finset.mem_Iic.mpr ((Finset.mem_Iic.mp i.2).trans (Nat.le_succ n))⟩ let previousHistory := pairHistory hK sigma2 delta margin defaultAction n previousRewardHistory let nextAction := historyNextArm hK sigma2 delta margin defaultAction n previousHistory History.extendPairHistorySucc previousHistory (nextAction, rewardHistory ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) def historyState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (n : Nat) (rewardHistory : History.FiniteRewardHistory…","missing":[],"search":"pairhistory banditrlproof.klucb.pairhistory pair-history reconstruction for the kl-ucb policy. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.historyState","label":"historyState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.historyState","description":"def historyState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (n : Nat) (rewardHistory : History.FiniteRewardHistory Rat n) : UCB.SelectedPolicySuccessorFiniteHistoryState K","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-f190874d316f","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1340,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:126"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def historyState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (n : Nat) (rewardHistory : History.FiniteRewardHistory Rat n) : UCB.SelectedPolicySuccessorFiniteHistoryState K","missing":[],"search":"historystate banditrlproof.klucb.historystate def historystate {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (n : nat) (rewardhistory : history.finiterewardhistory rat n) : ucb.selectedpolicysuccessorfinitehistorystate k definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measurable_historyState","label":"measurable_historyState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measurable_historyState","description":"theorem measurable_historyState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (n : Nat) : Measurable (historyState hK sigma2 delta margin defaultAction n)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-03ae58d06644","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1341,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (n : Nat) : Measurable (historyState hK sigma2 delta margin defaultAction n)","missing":[],"search":"measurable_historystate banditrlproof.klucb.measurable_historystate theorem measurable_historystate {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (n : nat) : measurable (historystate hk sigma2 delta margin defaultaction n) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.historyPolicy","label":"historyPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.historyPolicy","description":"The measurable KL-UCB policy. Its declaration has no terminal horizon.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-1911081ac60d","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1342,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:140"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def historyPolicy {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (_t : Nat) : Policy.MeasurablePolicy (UCB.SelectedPolicySuccessorFiniteHistoryState K) (Fin K) where","missing":[],"search":"historypolicy banditrlproof.klucb.historypolicy the measurable kl-ucb policy. its declaration has no terminal horizon. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedAction","label":"generatedAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedAction","description":"Canonical generated action trace of the KL-UCB policy.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-87af22d5d1a4","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1343,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:150"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def generatedAction {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace (Fin K)","missing":[],"search":"generatedaction banditrlproof.klucb.generatedaction canonical generated action trace of the kl-ucb policy. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedAction_succ","label":"generatedAction_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedAction_succ","description":"theorem generatedAction_succ {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) : generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1) = historyNextArm hK sigma2 delta margin defaultAction t (pairHistory hK sigma2 delta margin defaultAction t (History.finiteRewardHistoryOfTrace (rewar…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-1daaf2e5c8ed","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1344,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedAction_succ {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) : generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1) = historyNextArm hK sigma2 delta margin defaultAction t (pairHistory hK sigma2 delta margin defaultAction t (History.finiteRewardHistoryOfTrace (reward omega) t))","missing":[],"search":"generatedaction_succ banditrlproof.klucb.generatedaction_succ theorem generatedaction_succ {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (t : nat) : generatedaction hk sigma2 delta margin defaultaction reward omega (t + 1) = historynextarm hk sigma2 delta margin defaultaction t (pairhistory hk sigma2 delta margin defaultaction t (history.finiterewardhistoryoftrace (reward omega) t)) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.pairHistory_eq_finitePairHistoryOfTrace","label":"pairHistory_eq_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.pairHistory_eq_finitePairHistoryOfTrace","description":"theorem pairHistory_eq_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (n : Nat) : pairHistory hK sigma2 delta margin defaultAction n (History.finiteRewardHistoryOfTrace (reward omega) n) = History.finitePairHistoryOfTrace (generatedAction hK sigma2 delta margin defaultAction reward omeg…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-e28e5a6090a5","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1345,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:172"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pairHistory_eq_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (n : Nat) : pairHistory hK sigma2 delta margin defaultAction n (History.finiteRewardHistoryOfTrace (reward omega) n) = History.finitePairHistoryOfTrace (generatedAction hK sigma2 delta margin defaultAction reward omega) (reward omega) n","missing":[],"search":"pairhistory_eq_finitepairhistoryoftrace banditrlproof.klucb.pairhistory_eq_finitepairhistoryoftrace theorem pairhistory_eq_finitepairhistoryoftrace {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (n : nat) : pairhistory hk sigma2 delta margin defaultaction n (history.finiterewardhistoryoftrace (reward omega) n) = history.finitepairhistoryoftrace (generatedaction hk sigma2 delta margin defaultaction reward omega) (reward omega) n theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.historyIndex_finitePairHistoryOfTrace","label":"historyIndex_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.historyIndex_finitePairHistoryOfTrace","description":"The finite-history score, count, mean, and budget are definitionally the ones on the same generated trajectory.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-daa74196f9cb","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1346,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:214"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem historyIndex_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (omega : Omega) (t : Nat) (arm : Fin K) : historyIndex sigma2 delta margin defaultAction t (History.finitePairHistoryOfTrace (action omega) (reward omega) t) arm = generatedIndexAt action reward sigma2 delta margin omega t arm","missing":[],"search":"historyindex_finitepairhistoryoftrace banditrlproof.klucb.historyindex_finitepairhistoryoftrace the finite-history score, count, mean, and budget are definitionally the ones on the same generated trajectory. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedIndexAt_le_selected_of_K_le","label":"generatedIndexAt_le_selected_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedIndexAt_le_selected_of_K_le","description":"Selected KL index maximality on the actual generated action/reward trace.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-904c5fff61ca","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1347,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:232"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedIndexAt_le_selected_of_K_le {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (ht : K <= t) (arm : Fin K) : let action := generatedAction hK sigma2 delta margin defaultAction reward generatedIndexAt action reward sigma2 delta margin omega t arm <= generatedIndexAt action reward sigma2 delta margin omega t (action omega (t + 1))","missing":[],"search":"generatedindexat_le_selected_of_k_le banditrlproof.klucb.generatedindexat_le_selected_of_k_le selected kl index maximality on the actual generated action/reward trace. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedAction_succ_eq_initializationArm_of_lt","label":"generatedAction_succ_eq_initializationArm_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedAction_succ_eq_initializationArm_of_lt","description":"theorem generatedAction_succ_eq_initializationArm_of_lt {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (ht : t < K) : generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1) = UCB.initializationArm hK t","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-4eeefd5c4767","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1348,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:266"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedAction_succ_eq_initializationArm_of_lt {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (ht : t < K) : generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1) = UCB.initializationArm hK t","missing":[],"search":"generatedaction_succ_eq_initializationarm_of_lt banditrlproof.klucb.generatedaction_succ_eq_initializationarm_of_lt theorem generatedaction_succ_eq_initializationarm_of_lt {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (t : nat) (ht : t < k) : generatedaction hk sigma2 delta margin defaultaction reward omega (t + 1) = ucb.initializationarm hk t theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.successorArmPullCount_generatedAction_K_add_one_eq_one","label":"successorArmPullCount_generatedAction_K_add_one_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.successorArmPullCount_generatedAction_K_add_one_eq_one","description":"theorem successorArmPullCount_generatedAction_K_add_one_eq_one {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) : ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction reward omega) arm (K + 1) = 1","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-fb77fc28513a","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1349,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:275"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_generatedAction_K_add_one_eq_one {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) : ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction reward omega) arm (K + 1) = 1","missing":[],"search":"successorarmpullcount_generatedaction_k_add_one_eq_one banditrlproof.klucb.successorarmpullcount_generatedaction_k_add_one_eq_one theorem successorarmpullcount_generatedaction_k_add_one_eq_one {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (arm : fin k) : conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) arm (k + 1) = 1 theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.successorArmPullCount_generatedAction_pos_of_K_le","label":"successorArmPullCount_generatedAction_pos_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.successorArmPullCount_generatedAction_pos_of_K_le","description":"theorem successorArmPullCount_generatedAction_pos_of_K_le {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (ht : K <= t) : 0 < ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction reward omega) arm (t + 1)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-3ef23b85aaa5","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1350,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:298"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_generatedAction_pos_of_K_le {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (ht : K <= t) : 0 < ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction reward omega) arm (t + 1)","missing":[],"search":"successorarmpullcount_generatedaction_pos_of_k_le banditrlproof.klucb.successorarmpullcount_generatedaction_pos_of_k_le theorem successorarmpullcount_generatedaction_pos_of_k_le {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (arm : fin k) (t : nat) (ht : k <= t) : 0 < conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) arm (t + 1) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.K_le_of_generatedAction_selected_and_count_pos","label":"K_le_of_generatedAction_selected_and_count_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.K_le_of_generatedAction_selected_and_count_pos","description":"theorem K_le_of_generatedAction_selected_and_count_pos {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (hselected : generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1) = arm) (hcount : 0 < ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-1107f628ba3d","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1351,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:317"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem K_le_of_generatedAction_selected_and_count_pos {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (hselected : generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1) = arm) (hcount : 0 < ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction reward omega) arm (t + 1)) : K <= t","missing":[],"search":"k_le_of_generatedaction_selected_and_count_pos banditrlproof.klucb.k_le_of_generatedaction_selected_and_count_pos theorem k_le_of_generatedaction_selected_and_count_pos {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (arm : fin k) (t : nat) (hselected : generatedaction hk sigma2 delta margin defaultaction reward omega (t + 1) = arm) (hcount : 0 < conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) arm (t + 1)) : k <= t theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.armMean_mem_confidenceSet_of_abs_lt_radius","label":"armMean_mem_confidenceSet_of_abs_lt_radius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.armMean_mem_confidenceSet_of_abs_lt_radius","description":"A bounded empirical-mean deviation on the actual prefix makes the true interior arm mean feasible for the KL confidence set used by the policy.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-4a6c3d291bd5","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1352,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:347"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem armMean_mem_confidenceSet_of_abs_lt_radius {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (armMean : Fin K -> Rat) (omega : Omega) (t : Nat) (arm : Fin K) (hemp : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm)) (hmean : (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hdev : |UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm - (armMean arm : Real)| < UCB.selectedPolicySuccessorTelescopingRadiusAt action sigma2 delta omega t arm) : (armMean arm : Real) ∈ confidenceSet (UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm) (ConditionalExpectationReward.successorArmPullCount (action omega) arm (t + 1)) (generatedBudgetAt action sigma2 delta margin omeg…","missing":[],"search":"armmean_mem_confidenceset_of_abs_lt_radius banditrlproof.klucb.armmean_mem_confidenceset_of_abs_lt_radius a bounded empirical-mean deviation on the actual prefix makes the true interior arm mean feasible for the kl confidence set used by the policy. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.armMean_le_generatedIndexAt_of_abs_lt_radius","label":"armMean_le_generatedIndexAt_of_abs_lt_radius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.armMean_le_generatedIndexAt_of_abs_lt_radius","description":"theorem armMean_le_generatedIndexAt_of_abs_lt_radius {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (armMean : Fin K -> Rat) (omega : Omega) (t : Nat) (arm : Fin K) (hemp : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm)) (hm…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-290d0c4dfde2","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1353,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:413"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem armMean_le_generatedIndexAt_of_abs_lt_radius {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (armMean : Fin K -> Rat) (omega : Omega) (t : Nat) (arm : Fin K) (hemp : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm)) (hmean : (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hdev : |UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm - (armMean arm : Real)| < UCB.selectedPolicySuccessorTelescopingRadiusAt action sigma2 delta omega t arm) : (armMean arm : Real) <= generatedIndexAt action reward sigma2 delta margin omega t arm","missing":[],"search":"armmean_le_generatedindexat_of_abs_lt_radius banditrlproof.klucb.armmean_le_generatedindexat_of_abs_lt_radius theorem armmean_le_generatedindexat_of_abs_lt_radius {omega : type} {k : nat} (action : omega -> actiontrace (fin k)) (reward : omega -> rewardtrace rat) (sigma2 : nnreal) (delta margin : real) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (armmean : fin k -> rat) (omega : omega) (t : nat) (arm : fin k) (hemp : isbernoulliparameter (ucb.selectedpolicysuccessorempiricalmeanat action reward omega t arm)) (hmean : (armmean arm : real) ∈ set.icc margin (1 - margin)) (hdev : |ucb.selectedpolicysuccessorempiricalmeanat action reward omega t arm - (armmean arm : real)| < ucb.selectedpolicysuccessortelescopingradiusat action sigma2 delta omega t arm) : (armmean arm : real) <= generatedindexat action reward sigma2 delta margin omega t arm theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.margin_mul_gap_div_eight_le_radius_of_selected_of_not_badEvent","label":"margin_mul_gap_div_eight_le_radius_of_selected_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.margin_mul_gap_div_eight_le_radius_of_selected_of_not_badEvent","description":"On the common generated all-time good event, selection of a positive-gap arm forces its realized KL calibration radius to remain large. This is the single-arm, single-round KL pull-threshold leaf.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-98e444052886","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1354,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:438"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem margin_mul_gap_div_eight_le_radius_of_selected_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (defaultAction best : Fin K) (omega : Omega) (t : Nat) (ht : K <= t) (hemp : forall arm : Fin K, IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt (generatedAction hK sigma2 delta margin defaultAction reward) reward omega t arm)) (hbest : forall arm : Fin K, (armMean arm : Real) <= (armMean best : Real)) (hgap : 0 < UCB.meanGap (fun arm => (armMean arm : Real)) best (generatedAction hK sigma2 delta margin defaultAction reward omega (t + 1))) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAl…","missing":[],"search":"margin_mul_gap_div_eight_le_radius_of_selected_of_not_badevent banditrlproof.klucb.margin_mul_gap_div_eight_le_radius_of_selected_of_not_badevent on the common generated all-time good event, selection of a positive-gap arm forces its realized kl calibration radius to remain large. this is the single-arm, single-round kl pull-threshold leaf. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.pullThreshold","label":"pullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.pullThreshold","description":"Explicit KL-UCB pull threshold, obtained by inverting the accepted telescoping radius at the effective gap `margin * gap / 4`.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-fb18a67a225b","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1355,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:610"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def pullThreshold (K : Nat) (sigma2 : NNReal) (T : Nat) (delta margin gap : Real) : Nat","missing":[],"search":"pullthreshold banditrlproof.klucb.pullthreshold explicit kl-ucb pull threshold, obtained by inverting the accepted telescoping radius at the effective gap `margin * gap / 4`. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.pullCount_le_of_not_badEvent","label":"pullCount_le_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.pullCount_le_of_not_badEvent","description":"theorem pullCount_le_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (defaultAction best chosen : Fin K) (omega : Omega) (T : Nat) (hbest : forall arm : F…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-02b10ceab239","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1356,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:616"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (defaultAction best chosen : Fin K) (omega : Omega) (T : Nat) (hbest : forall arm : Fin K, (armMean arm : Real) <= (armMean best : Real)) (hemp : forall t : Nat, forall arm : Fin K, IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt (generatedAction hK sigma2 delta margin defaultAction reward) reward omega t arm)) (hgap : 0 < UCB.meanGap (fun arm => (armMean arm : Real)) best chosen) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (generatedAction hK sigma2 delta margin defaultAction rewar…","missing":[],"search":"pullcount_le_of_not_badevent banditrlproof.klucb.pullcount_le_of_not_badevent theorem pullcount_le_of_not_badevent {omega : type} {k : nat} (hk : 0 < k) (reward : omega -> rewardtrace rat) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (hmean : forall arm : fin k, (armmean arm : real) ∈ set.icc margin (1 - margin)) (defaultaction best chosen : fin k) (omega : omega) (t : nat) (hbest : forall arm : fin k, (armmean arm : real) <= (armmean best : real)) (hemp : forall t : nat, forall arm : fin k, isbernoulliparameter (ucb.selectedpolicysuccessorempiricalmeanat (generatedaction hk sigma2 delta margin defaultaction reward) reward omega t arm)) (hgap : 0 < ucb.meangap (fun arm => (armmean arm : real)) best chosen) (hgood : omega ∉ conditionalexpectationreward.successorarmempiricalmeanfintypetelescopingalltimebadevent (generatedaction hk sigma2 delta margin defaultaction reward) reward armmean sigma2 delta) : conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) chosen (t + 1) <= pullthreshold k sigma2 t delta margin (ucb.meangap (fun arm => (armmean arm : real)) best chosen) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.actionRewardHistoryStepKernelFamily_allTimeConfidence","label":"actionRewardHistoryStepKernelFamily_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.actionRewardHistoryStepKernelFamily_allTimeConfidence","description":"The accepted all-time telescoping confidence producer instantiated on the KL-UCB policy. The theorem explicitly transports the sampled pair action to the reward-reconstructed action used by `generatedIndexAt`.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-ccdac7dfaa3f","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1357,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:708"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_allTimeConfidence {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((historyPolicy hK sigma2 delta margin defaultAction i).action (historyState hK sigma2…","missing":[],"search":"actionrewardhistorystepkernelfamily_alltimeconfidence banditrlproof.klucb.actionrewardhistorystepkernelfamily_alltimeconfidence the accepted all-time telescoping confidence producer instantiated on the kl-ucb policy. the theorem explicitly transports the sampled pair action to the reward-reconstructed action used by `generatedindexat`. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedKLAllTimeBadEvent","label":"generatedKLAllTimeBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedKLAllTimeBadEvent","description":"Arms-by-times KL-confidence failure event on the exact generated trace.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-d9b395c57d7b","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1358,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:839"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def generatedKLAllTimeBadEvent {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) : Set Omega","missing":[],"search":"generatedklalltimebadevent banditrlproof.klucb.generatedklalltimebadevent arms-by-times kl-confidence failure event on the exact generated trace. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedKLAllTimeBadEvent_subset_absBadEvent","label":"generatedKLAllTimeBadEvent_subset_absBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedKLAllTimeBadEvent_subset_absBadEvent","description":"theorem generatedKLAllTimeBadEvent_subset_absBadEvent {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hemp : forall omega t arm, IsBernoulliParameter (UCB.selected…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-85acfd2b7db4","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1359,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:855"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedKLAllTimeBadEvent_subset_absBadEvent {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hemp : forall omega t arm, IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm)) : generatedKLAllTimeBadEvent action reward armMean sigma2 delta margin ⊆ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent action reward armMean sigma2 delta","missing":[],"search":"generatedklalltimebadevent_subset_absbadevent banditrlproof.klucb.generatedklalltimebadevent_subset_absbadevent theorem generatedklalltimebadevent_subset_absbadevent {omega : type} {k : nat} (action : omega -> actiontrace (fin k)) (reward : omega -> rewardtrace rat) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (hmean : forall arm : fin k, (armmean arm : real) ∈ set.icc margin (1 - margin)) (hemp : forall omega t arm, isbernoulliparameter (ucb.selectedpolicysuccessorempiricalmeanat action reward omega t arm)) : generatedklalltimebadevent action reward armmean sigma2 delta margin ⊆ conditionalexpectationreward.successorarmempiricalmeanfintypetelescopingalltimebadevent action reward armmean sigma2 delta theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le","label":"measure_generatedKLAllTimeBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le","description":"theorem measure_generatedKLAllTimeBadEvent_le {Omega : Type} [MeasurableSpace Omega] {K : Nat} (mu : Measure Omega) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hemp : ∀ᵐ omega ∂mu, for…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-41a2d6031676","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1360,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:890"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_generatedKLAllTimeBadEvent_le {Omega : Type} [MeasurableSpace Omega] {K : Nat} (mu : Measure Omega) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hemp : ∀ᵐ omega ∂mu, forall t arm, IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt action reward omega t arm)) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent action reward armMean sigma2 delta) <= ENNReal.ofReal delta) : mu (generatedKLAllTimeBadEvent action reward armMean sigma2 delta margin) <= ENNReal.ofReal delta","missing":[],"search":"measure_generatedklalltimebadevent_le banditrlproof.klucb.measure_generatedklalltimebadevent_le theorem measure_generatedklalltimebadevent_le {omega : type} [measurablespace omega] {k : nat} (mu : measure omega) (action : omega -> actiontrace (fin k)) (reward : omega -> rewardtrace rat) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (hmean : forall arm : fin k, (armmean arm : real) ∈ set.icc margin (1 - margin)) (hemp : ∀ᵐ omega ∂mu, forall t arm, isbernoulliparameter (ucb.selectedpolicysuccessorempiricalmeanat action reward omega t arm)) (hconfidence : mu (conditionalexpectationreward.successorarmempiricalmeanfintypetelescopingalltimebadevent action reward armmean sigma2 delta) <= ennreal.ofreal delta) : mu (generatedklalltimebadevent action reward armmean sigma2 delta margin) <= ennreal.ofreal delta theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.sumRewards_nonneg_of_mem_Icc_zero_one","label":"sumRewards_nonneg_of_mem_Icc_zero_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.sumRewards_nonneg_of_mem_Icc_zero_one","description":"Selected reward sums remain nonnegative under a pathwise `[0,1]` reward contract.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-b4424e1e5f73","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1361,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:932"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_nonneg_of_mem_Icc_zero_one {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Real) (arm : Action) (n : Nat) (hraw : forall i, i < n -> reward i ∈ Set.Icc (0 : Real) 1) : 0 <= sumRewards action reward arm n","missing":[],"search":"sumrewards_nonneg_of_mem_icc_zero_one banditrlproof.klucb.sumrewards_nonneg_of_mem_icc_zero_one selected reward sums remain nonnegative under a pathwise `[0,1]` reward contract. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.sumRewards_le_pullCount_of_mem_Icc_zero_one","label":"sumRewards_le_pullCount_of_mem_Icc_zero_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.sumRewards_le_pullCount_of_mem_Icc_zero_one","description":"theorem sumRewards_le_pullCount_of_mem_Icc_zero_one {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Real) (arm : Action) (n : Nat) (hraw : forall i, i < n -> reward i ∈ Set.Icc (0 : Real) 1) : sumRewards action reward arm n <= (pullCount action arm n : Real)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-14516f86189f","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1362,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:947"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_le_pullCount_of_mem_Icc_zero_one {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Real) (arm : Action) (n : Nat) (hraw : forall i, i < n -> reward i ∈ Set.Icc (0 : Real) 1) : sumRewards action reward arm n <= (pullCount action arm n : Real)","missing":[],"search":"sumrewards_le_pullcount_of_mem_icc_zero_one banditrlproof.klucb.sumrewards_le_pullcount_of_mem_icc_zero_one theorem sumrewards_le_pullcount_of_mem_icc_zero_one {action : type} [decidableeq action] (action : actiontrace action) (reward : rewardtrace real) (arm : action) (n : nat) (hraw : forall i, i < n -> reward i ∈ set.icc (0 : real) 1) : sumrewards action reward arm n <= (pullcount action arm n : real) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","label":"successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","description":"Hence every positive-count empirical mean on the canonical successor trace is a Bernoulli parameter.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-b2ba03d346a9","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1363,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:966"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (arm : Action) (n : Nat) (hcount : 0 < ConditionalExpectationReward.successorArmPullCount action arm n) (hraw : forall i : Nat, (((reward i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) : IsBernoulliParameter (ConditionalExpectationReward.successorArmEmpiricalMean action reward arm n)","missing":[],"search":"successorarmempiricalmean_isbernoulliparameter_of_rewards_mem_icc banditrlproof.klucb.successorarmempiricalmean_isbernoulliparameter_of_rewards_mem_icc hence every positive-count empirical mean on the canonical successor trace is a bernoulli parameter. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","label":"generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","description":"theorem generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (hraw : forall i : Nat, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (t : Nat) (ht : K <= t) (arm : Fin K) : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt (genera…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-3b80306d23ce","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1364,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:998"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (hraw : forall i : Nat, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (t : Nat) (ht : K <= t) (arm : Fin K) : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt (generatedAction hK sigma2 delta margin defaultAction reward) reward omega t arm)","missing":[],"search":"generatedempiricalmean_isbernoulliparameter_of_rewards_mem_icc banditrlproof.klucb.generatedempiricalmean_isbernoulliparameter_of_rewards_mem_icc theorem generatedempiricalmean_isbernoulliparameter_of_rewards_mem_icc {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (hraw : forall i : nat, (((reward omega i : rat) : real)) ∈ set.icc (0 : real) 1) (t : nat) (ht : k <= t) (arm : fin k) : isbernoulliparameter (ucb.selectedpolicysuccessorempiricalmeanat (generatedaction hk sigma2 delta margin defaultaction reward) reward omega t arm) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","label":"successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","description":"theorem successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc' {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (arm : Action) (n : Nat) (hraw : forall i : Nat, (((reward i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) : IsBernoulliParameter (ConditionalExpectationReward.successorArmEmpiricalMean action reward arm n)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-f5940aa0069e","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1365,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1013"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc' {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (arm : Action) (n : Nat) (hraw : forall i : Nat, (((reward i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) : IsBernoulliParameter (ConditionalExpectationReward.successorArmEmpiricalMean action reward arm n)","missing":[],"search":"successorarmempiricalmean_isbernoulliparameter_of_rewards_mem_icc' banditrlproof.klucb.successorarmempiricalmean_isbernoulliparameter_of_rewards_mem_icc' theorem successorarmempiricalmean_isbernoulliparameter_of_rewards_mem_icc' {action : type} [decidableeq action] (action : actiontrace action) (reward : rewardtrace rat) (arm : action) (n : nat) (hraw : forall i : nat, (((reward i : rat) : real)) ∈ set.icc (0 : real) 1) : isbernoulliparameter (conditionalexpectationreward.successorarmempiricalmean action reward arm n) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","label":"generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","description":"theorem generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc' {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (hraw : forall i : Nat, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (t : Nat) (arm : Fin K) : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt (generatedAction hK…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-56b57bf12eec","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1366,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1027"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc' {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (hraw : forall i : Nat, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (t : Nat) (arm : Fin K) : IsBernoulliParameter (UCB.selectedPolicySuccessorEmpiricalMeanAt (generatedAction hK sigma2 delta margin defaultAction reward) reward omega t arm)","missing":[],"search":"generatedempiricalmean_isbernoulliparameter_of_rewards_mem_icc' banditrlproof.klucb.generatedempiricalmean_isbernoulliparameter_of_rewards_mem_icc' theorem generatedempiricalmean_isbernoulliparameter_of_rewards_mem_icc' {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (hraw : forall i : nat, (((reward omega i : rat) : real)) ∈ set.icc (0 : real) 1) (t : nat) (arm : fin k) : isbernoulliparameter (ucb.selectedpolicysuccessorempiricalmeanat (generatedaction hk sigma2 delta margin defaultaction reward) reward omega t arm) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measure_pullCount_gt_threshold_le_of_allTimeConfidence","label":"measure_pullCount_gt_threshold_le_of_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measure_pullCount_gt_threshold_le_of_allTimeConfidence","description":"theorem measure_pullCount_gt_threshold_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hraw : ∀ᵐ ome…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-67068023b655","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1367,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1040"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_gt_threshold_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hraw : ∀ᵐ omega ∂mu, forall i, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (defaultAction best chosen : Fin K) (T : Nat) (hbest : forall arm : Fin K, (armMean arm : Real) <= (armMean best : Real)) (hgap : 0 < UCB.meanGap (fun arm => (armMean arm : Real)) best chosen) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (generatedAction hK sigma2 delta margin defaultAction reward) reward armMean sigma2 delta) <= ENNReal.of…","missing":[],"search":"measure_pullcount_gt_threshold_le_of_alltimeconfidence banditrlproof.klucb.measure_pullcount_gt_threshold_le_of_alltimeconfidence theorem measure_pullcount_gt_threshold_le_of_alltimeconfidence {omega : type} [measurablespace omega] {k : nat} (hk : 0 < k) (mu : measure omega) (reward : omega -> rewardtrace rat) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (hmean : forall arm : fin k, (armmean arm : real) ∈ set.icc margin (1 - margin)) (hraw : ∀ᵐ omega ∂mu, forall i, (((reward omega i : rat) : real)) ∈ set.icc (0 : real) 1) (defaultaction best chosen : fin k) (t : nat) (hbest : forall arm : fin k, (armmean arm : real) <= (armmean best : real)) (hgap : 0 < ucb.meangap (fun arm => (armmean arm : real)) best chosen) (hconfidence : mu (conditionalexpectationreward.successorarmempiricalmeanfintypetelescopingalltimebadevent (generatedaction hk sigma2 delta margin defaultaction reward) reward armmean sigma2 delta) <= ennreal.ofreal delta) : mu {omega | pullthreshold k sigma2 t delta margin (ucb.meangap (fun arm => (armmean arm : real)) best chosen) < conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) chosen (t + 1)} <= ennreal.ofreal delta theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.lintegral_pullCount_le_of_allTimeConfidence","label":"lintegral_pullCount_le_of_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.lintegral_pullCount_le_of_allTimeConfidence","description":"theorem lintegral_pullCount_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hm…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-114978c834a6","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1368,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1075"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_pullCount_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hraw : ∀ᵐ omega ∂mu, forall i, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (defaultAction best chosen : Fin K) (T : Nat) (hbest : forall arm : Fin K, (armMean arm : Real) <= (armMean best : Real)) (hgap : 0 < UCB.meanGap (fun arm => (armMean arm : Real)) best chosen) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (generatedAc…","missing":[],"search":"lintegral_pullcount_le_of_alltimeconfidence banditrlproof.klucb.lintegral_pullcount_le_of_alltimeconfidence theorem lintegral_pullcount_le_of_alltimeconfidence {omega : type} [measurablespace omega] {k : nat} (hk : 0 < k) (mu : measure omega) [isprobabilitymeasure mu] (reward : omega -> rewardtrace rat) (hreward : forall t : nat, measurable (fun omega : omega => reward omega t)) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (hmean : forall arm : fin k, (armmean arm : real) ∈ set.icc margin (1 - margin)) (hraw : ∀ᵐ omega ∂mu, forall i, (((reward omega i : rat) : real)) ∈ set.icc (0 : real) 1) (defaultaction best chosen : fin k) (t : nat) (hbest : forall arm : fin k, (armmean arm : real) <= (armmean best : real)) (hgap : 0 < ucb.meangap (fun arm => (armmean arm : real)) best chosen) (hconfidence : mu (conditionalexpectationreward.successorarmempiricalmeanfintypetelescopingalltimebadevent (generatedaction hk sigma2 delta margin defaultaction reward) reward armmean sigma2 delta) <= ennreal.ofreal delta) : ∫⁻ omega, (conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) chosen (t + 1) : ennreal) ∂mu <= (pullthreshold k sigma2 t delta margin (ucb.meangap (fun arm => (armmean arm : real)) best chosen) : ennreal) + (t : ennrea…","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.generatedRegretAction","label":"generatedRegretAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.generatedRegretAction","description":"Regret-time shift of the canonical generated KL-UCB action.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-4f53f4c206ba","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1369,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def generatedRegretAction {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace (Fin K)","missing":[],"search":"generatedregretaction banditrlproof.klucb.generatedregretaction regret-time shift of the canonical generated kl-ucb action. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measurable_generatedRegretAction","label":"measurable_generatedRegretAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measurable_generatedRegretAction","description":"theorem measurable_generatedRegretAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => generatedRegretAction hK sigma2 delta margin defaultAction reward omega t)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-b82f84f25808","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1370,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedRegretAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => generatedRegretAction hK sigma2 delta margin defaultAction reward omega t)","missing":[],"search":"measurable_generatedregretaction banditrlproof.klucb.measurable_generatedregretaction theorem measurable_generatedregretaction {omega : type} [measurablespace omega] {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (hreward : forall t : nat, measurable (fun omega : omega => reward omega t)) (t : nat) : measurable (fun omega => generatedregretaction hk sigma2 delta margin defaultaction reward omega t) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.pullCount_generatedRegretAction_eq","label":"pullCount_generatedRegretAction_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.pullCount_generatedRegretAction_eq","description":"theorem pullCount_generatedRegretAction_eq {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (T : Nat) : pullCount (generatedRegretAction hK sigma2 delta margin defaultAction reward omega) arm T = ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-9f06560e8120","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1371,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1153"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_generatedRegretAction_eq {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (T : Nat) : pullCount (generatedRegretAction hK sigma2 delta margin defaultAction reward omega) arm T = ConditionalExpectationReward.successorArmPullCount (generatedAction hK sigma2 delta margin defaultAction reward omega) arm (T + 1)","missing":[],"search":"pullcount_generatedregretaction_eq banditrlproof.klucb.pullcount_generatedregretaction_eq theorem pullcount_generatedregretaction_eq {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (arm : fin k) (t : nat) : pullcount (generatedregretaction hk sigma2 delta margin defaultaction reward omega) arm t = conditionalexpectationreward.successorarmpullcount (generatedaction hk sigma2 delta margin defaultaction reward omega) arm (t + 1) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le","label":"lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le","description":"Finite-time expected pseudo-regret of the actual generated KL-UCB policy. The `T * delta` failure-event term is explicit; this theorem makes no asymptotic-optimality or leading-constant claim.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-42334c4c9831","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1372,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1169"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le {Omega : Type} [MeasurableSpace Omega] {K : Nat} (mu : Measure Omega) [IsProbabilityMeasure mu] (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (model.mean arm : Real) ∈ Set.Icc margin (1 - margin)) (hraw : ∀ᵐ omega ∂mu, forall i, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (defaultAction : Fin K) (T : Nat) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (generatedAction model.hK sigma2 delta margin defaultAction reward) reward model.mean sigma2 delta) <= ENNReal.ofReal delta) : ∫⁻ omega, ENNReal.ofReal (((pseudoRegr…","missing":[],"search":"lintegral_ofreal_pseudoregret_generatedklucbbounded_le banditrlproof.klucb.lintegral_ofreal_pseudoregret_generatedklucbbounded_le finite-time expected pseudo-regret of the actual generated kl-ucb policy. the `t * delta` failure-event term is explicit; this theorem makes no asymptotic-optimality or leading-constant claim. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.allHorizonPullCount_of_not_badEvent","label":"allHorizonPullCount_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.allHorizonPullCount_of_not_badEvent","description":"One good generated sample controls every finite horizon and positive-gap arm for the same KL-UCB policy.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-ba40d323e44f","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1373,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1228"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem allHorizonPullCount_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (hmean : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (defaultAction best : Fin K) (omega : Omega) (hraw : forall i : Nat, (((reward omega i : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (hbest : forall arm : Fin K, (armMean arm : Real) <= (armMean best : Real)) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (generatedAction hK sigma2 delta margin defaultAction reward) reward armMean sigma2 delta) : forall T : Nat, forall chosen : Fin K, 0 < UCB.meanGap (fun arm => (armMean arm : Real)) best chosen -> ConditionalExpectationReward.successorArmPullCount (g…","missing":[],"search":"allhorizonpullcount_of_not_badevent banditrlproof.klucb.allhorizonpullcount_of_not_badevent one good generated sample controls every finite horizon and positive-gap arm for the same kl-ucb policy. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.actionRewardTrajMeasure","label":"actionRewardTrajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.KLUCB.actionRewardTrajMeasure","description":"Canonical action/reward trajectory measure of the KL-UCB policy.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-127d1df8afe3","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1374,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1258"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def actionRewardTrajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) : Measure (Nat -> Prod (Fin K) Rat)","missing":[],"search":"actionrewardtrajmeasure banditrlproof.klucb.actionrewardtrajmeasure canonical action/reward trajectory measure of the kl-ucb policy. definition compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measure_allTimeBadEvent_le_trajMeasure","label":"measure_allTimeBadEvent_le_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measure_allTimeBadEvent_le_trajMeasure","description":"theorem measure_allTimeBadEvent_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (a…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-44c379c5a59c","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1375,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1287"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_allTimeBadEvent_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((historyPolicy hK sigma2 delta margin defaultAction i).action (historyState hK sigma2 delta margin de…","missing":[],"search":"measure_alltimebadevent_le_trajmeasure banditrlproof.klucb.measure_alltimebadevent_le_trajmeasure theorem measure_alltimebadevent_le_trajmeasure {context : type} {k : nat} [measurablespace context] (hk : 0 < k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction : fin k) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((historypolicy hk sigma2 delta margin defaultaction i).action (historystate hk sigma2 delta margin defaultaction i history)) <= sigma2) (harmmean : forall i : nat, forall history : ((j : finset.iic i) -> rat), forall arm : fin k, mean (context i history) arm = armmean arm) (hsigma2 : 0 < (((sigma2 : nnreal) : real))) (hdelta : 0 < delta) : let mu := actionrewardtrajmeasure hk mu0 rewardkernel context hcontext sigma2 delta margin defaultaction let reward : (nat -> prod (fin k) rat) -> rewardtrace ra…","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le_trajMeasure","label":"measure_generatedKLAllTimeBadEvent_le_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le_trajMeasure","description":"theorem measure_generatedKLAllTimeBadEvent_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction…","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-262848501d20","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1376,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1321"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_generatedKLAllTimeBadEvent_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta margin : Real) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (harmMargin : forall arm : Fin K, (armMean arm : Real) ∈ Set.Icc margin (1 - margin)) (hcontext : forall n : Nat, Measurable (context n)) (hraw : ∀ᵐ trajectory ∂(actionRewardTrajMeasure hK mu0 rewardKernel context hcontext sigma2 delta margin defaultAction), forall i : Nat, ((((trajectory i).2 : Rat) : Real)) ∈ Set.Icc (0 : Real) 1) (hmean : Measur…","missing":[],"search":"measure_generatedklalltimebadevent_le_trajmeasure banditrlproof.klucb.measure_generatedklalltimebadevent_le_trajmeasure theorem measure_generatedklalltimebadevent_le_trajmeasure {context : type} {k : nat} [measurablespace context] (hk : 0 < k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction : fin k) (armmean : fin k -> rat) (sigma2 : nnreal) (delta margin : real) (hmargin0 : 0 < margin) (hmarginhalf : margin <= 1 / 2) (harmmargin : forall arm : fin k, (armmean arm : real) ∈ set.icc margin (1 - margin)) (hcontext : forall n : nat, measurable (context n)) (hraw : ∀ᵐ trajectory ∂(actionrewardtrajmeasure hk mu0 rewardkernel context hcontext sigma2 delta margin defaultaction), forall i : nat, ((((trajectory i).2 : rat) : real)) ∈ set.icc (0 : real) 1) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((historypolicy hk sigma2 delta margin defaultaction i).action (historystate hk sigma2 delta margin defaultaction i history)) <= sig…","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","label":"lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","description":"Canonical generated-policy KL-UCB regret theorem on the very trajectory measure built from the same measurable KL policy and reward kernel.","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-f0183508accc","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1377,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1380"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","stochastic-finite"]],"statement":"theorem lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (delta margin : Real) (T : Nat) (hdelta : 0 < delta) (hmargin0 : 0 < margin) (hmarginHalf : margin <= 1 / 2) (harmMargin : forall arm : Fin K, (model.mean arm : Real) ∈ Set.Icc margin (1 - margin)) (hcontext : forall n : Nat, Measurable (context n)) (hraw : ∀ᵐ trajectory ∂(actionRewardTrajMeasure model.hK mu0 rewardKernel context hcontext sigma2 delta margin defaultAction), forall i : Nat, ((((trajectory i).2 : Rat)…","missing":[],"search":"lintegral_ofreal_pseudoregret_generatedklucbbounded_le_trajmeasure banditrlproof.klucb.lintegral_ofreal_pseudoregret_generatedklucbbounded_le_trajmeasure canonical generated-policy kl-ucb regret theorem on the very trajectory measure built from the same measurable kl policy and reward kernel. theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["stochastic-finite"]},{"id":"declaration:BanditRLProof.KLUCB.measurable_generatedAction","label":"measurable_generatedAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KLUCB.measurable_generatedAction","description":"theorem measurable_generatedAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => generatedAction hK sigma2 delta margin defaultAction reward omega t)","url":"../modules/banditrlproof-algorithms-klucbgeneratedregret/index.html#decl-c0c8fb8f5e3f","parent":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","order":1378,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.KLUCBGeneratedRegret"],["Source","BanditRLProof/Algorithms/KLUCBGeneratedRegret.lean:1442"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta margin : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => generatedAction hK sigma2 delta margin defaultAction reward omega t)","missing":[],"search":"measurable_generatedaction banditrlproof.klucb.measurable_generatedaction theorem measurable_generatedaction {omega : type} [measurablespace omega] {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta margin : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (hreward : forall t : nat, measurable (fun omega : omega => reward omega t)) (t : nat) : measurable (fun omega => generatedaction hk sigma2 delta margin defaultaction reward omega t) theorem compiled","shard":"modules/978c2a806ed308d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.logPlus","label":"logPlus","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.logPlus","description":"Source convention `log max {1,x}`.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-f8bd6778050d","parent":"module:BanditRLProof.Algorithms.MOSS","order":1379,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def logPlus (x : ℝ) : ℝ","missing":[],"search":"logplus banditrlproof.moss.logplus source convention `log max {1,x}`. definition compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.radius","label":"radius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.radius","description":"Source confidence radius at sample count `s`. Real division totalizes `s=0`; the post-initialization algorithm must be used on positive counts.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-8734447bddf8","parent":"module:BanditRLProof.Algorithms.MOSS","order":1380,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def radius (n k s : ℕ) : ℝ","missing":[],"search":"radius banditrlproof.moss.radius source confidence radius at sample count `s`. real division totalizes `s=0`; the post-initialization algorithm must be used on positive counts. definition compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.index","label":"index","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.index","description":"Source index for arbitrary current empirical means and pull counts.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-fba9ce8ff27c","parent":"module:BanditRLProof.Algorithms.MOSS","order":1381,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def index {k : ℕ} (n : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (a : Fin k) : ℝ","missing":[],"search":"index banditrlproof.moss.index source index for arbitrary current empirical means and pull counts. definition compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.action","label":"action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.action","description":"Zero-based Algorithm 7 action: first each arm once, then a real argmax. This accepts a state; consistency of that state with past feedback is a separate history-level obligation.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-07cd4f80ab34","parent":"module:BanditRLProof.Algorithms.MOSS","order":1382,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def action {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) : Fin k","missing":[],"search":"action banditrlproof.moss.action zero-based algorithm 7 action: first each arm once, then a real argmax. this accepts a state; consistency of that state with past feedback is a separate history-level obligation. definition compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.logPlus_nonneg","label":"logPlus_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.logPlus_nonneg","description":"theorem logPlus_nonneg (x : ℝ) : 0 ≤ logPlus x","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-2cc07d4bb4b7","parent":"module:BanditRLProof.Algorithms.MOSS","order":1383,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem logPlus_nonneg (x : ℝ) : 0 ≤ logPlus x","missing":[],"search":"logplus_nonneg banditrlproof.moss.logplus_nonneg theorem logplus_nonneg (x : ℝ) : 0 ≤ logplus x theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.radius_nonneg","label":"radius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.radius_nonneg","description":"theorem radius_nonneg (n k s : ℕ) : 0 ≤ radius n k s","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-13268ebdc25a","parent":"module:BanditRLProof.Algorithms.MOSS","order":1384,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem radius_nonneg (n k s : ℕ) : 0 ≤ radius n k s","missing":[],"search":"radius_nonneg banditrlproof.moss.radius_nonneg theorem radius_nonneg (n k s : ℕ) : 0 ≤ radius n k s theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.radius_zero","label":"radius_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.radius_zero","description":"@[simp] theorem radius_zero (n k : ℕ) : radius n k 0 = 0","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-898261ca02bc","parent":"module:BanditRLProof.Algorithms.MOSS","order":1385,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem radius_zero (n k : ℕ) : radius n k 0 = 0","missing":[],"search":"radius_zero banditrlproof.moss.radius_zero @[simp] theorem radius_zero (n k : ℕ) : radius n k 0 = 0 theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.radius_sq","label":"radius_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.radius_sq","description":"Exact squared source radius, including the totalized zero-count branch.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-f8167365e15d","parent":"module:BanditRLProof.Algorithms.MOSS","order":1386,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem radius_sq (n k s : ℕ) : radius n k s ^ 2 = 4 / (s : ℝ) * logPlus ((n : ℝ) / ((k : ℝ) * (s : ℝ)))","missing":[],"search":"radius_sq banditrlproof.moss.radius_sq exact squared source radius, including the totalized zero-count branch. theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.action_of_lt","label":"action_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.action_of_lt","description":"@[simp] theorem action_of_lt {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (ht : t < k) : action hk n t empiricalMean pulls = ⟨t, ht⟩","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-8d584a428613","parent":"module:BanditRLProof.Algorithms.MOSS","order":1387,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem action_of_lt {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (ht : t < k) : action hk n t empiricalMean pulls = ⟨t, ht⟩","missing":[],"search":"action_of_lt banditrlproof.moss.action_of_lt @[simp] theorem action_of_lt {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalmean : fin k → ℝ) (pulls : fin k → ℕ) (ht : t < k) : action hk n t empiricalmean pulls = ⟨t, ht⟩ theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.action_initial_arm","label":"action_initial_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.action_initial_arm","description":"Initialization selects each source arm at its own zero-based time.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-6f4f2c6ebb50","parent":"module:BanditRLProof.Algorithms.MOSS","order":1388,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_initial_arm {k : ℕ} (hk : 0 < k) (n : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (a : Fin k) : action hk n a.val empiricalMean pulls = a","missing":[],"search":"action_initial_arm banditrlproof.moss.action_initial_arm initialization selects each source arm at its own zero-based time. theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.action_index_max","label":"action_index_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.action_index_max","description":"theorem action_index_max {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (ht : k ≤ t) (a : Fin k) : index n empiricalMean pulls a ≤ index n empiricalMean pulls (action hk n t empiricalMean pulls)","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-ac45c09fc984","parent":"module:BanditRLProof.Algorithms.MOSS","order":1389,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem action_index_max {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (ht : k ≤ t) (a : Fin k) : index n empiricalMean pulls a ≤ index n empiricalMean pulls (action hk n t empiricalMean pulls)","missing":[],"search":"action_index_max banditrlproof.moss.action_index_max theorem action_index_max {k : ℕ} (hk : 0 < k) (n t : ℕ) (empiricalmean : fin k → ℝ) (pulls : fin k → ℕ) (ht : k ≤ t) (a : fin k) : index n empiricalmean pulls a ≤ index n empiricalmean pulls (action hk n t empiricalmean pulls) theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.selected_index_gt_mean_add_half_gap","label":"selected_index_gt_mean_add_half_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.selected_index_gt_mean_add_half_gap","description":"The large-gap selection implication in the proof of Theorem 9.1. The optimism-deficit bound is explicit and is not a concentration theorem.","url":"../modules/banditrlproof-algorithms-moss/index.html#decl-87637561719b","parent":"module:BanditRLProof.Algorithms.MOSS","order":1390,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSS"],["Source","BanditRLProof/Algorithms/MOSS.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem selected_index_gt_mean_add_half_gap {k : ℕ} (hk : 0 < k) (n t : ℕ) (mean empiricalMean : Fin k → ℝ) (pulls : Fin k → ℕ) (best chosen : Fin k) (deficit : ℝ) (ht : k ≤ t) (hselected : action hk n t empiricalMean pulls = chosen) (hoptimism : mean best - deficit ≤ index n empiricalMean pulls best) (hgap : 2 * deficit < mean best - mean chosen) : mean chosen + (mean best - mean chosen) / 2 < index n empiricalMean pulls chosen","missing":[],"search":"selected_index_gt_mean_add_half_gap banditrlproof.moss.selected_index_gt_mean_add_half_gap the large-gap selection implication in the proof of theorem 9.1. the optimism-deficit bound is explicit and is not a concentration theorem. theorem compiled","shard":"modules/9770649f4b4075ea.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalAction","label":"canonicalAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalAction","description":"def canonicalAction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-7f9886ec28e5","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1391,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def canonicalAction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ)","missing":[],"search":"canonicalaction banditrlproof.moss.canonicalaction def canonicalaction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) definition compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalReward","label":"canonicalReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalReward","description":"Consume the next unused one-based raw reward of the chosen arm.","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-af4b2880167e","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1392,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def canonicalReward {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ)","missing":[],"search":"canonicalreward banditrlproof.moss.canonicalreward consume the next unused one-based raw reward of the chosen arm. definition compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory","label":"canonicalHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory","description":"def canonicalHistory {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-2de7868b3acb","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1393,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def canonicalHistory {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ)","missing":[],"search":"canonicalhistory banditrlproof.moss.canonicalhistory def canonicalhistory {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (table : ucb.armrewardstream k) (t : ℕ) definition compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory_pullCount","label":"canonicalHistory_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory_pullCount","description":"theorem canonicalHistory_pullCount {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) (a : Fin k) : ETC.realHistoryPullCount t (canonicalHistory hk n mean table t) a = pullCount (canonicalAction hk n mean table) a (t+1)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-d8f824e88f26","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1394,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistory_pullCount {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) (a : Fin k) : ETC.realHistoryPullCount t (canonicalHistory hk n mean table t) a = pullCount (canonicalAction hk n mean table) a (t+1)","missing":[],"search":"canonicalhistory_pullcount banditrlproof.moss.canonicalhistory_pullcount theorem canonicalhistory_pullcount {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (table : ucb.armrewardstream k) (t : ℕ) (a : fin k) : etc.realhistorypullcount t (canonicalhistory hk n mean table t) a = pullcount (canonicalaction hk n mean table) a (t+1) theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory_empiricalMean","label":"canonicalHistory_empiricalMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory_empiricalMean","description":"theorem canonicalHistory_empiricalMean {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) (a : Fin k) : ETC.realHistoryEmpMean t (canonicalHistory hk n mean table t) a = (∑ j ∈ Finset.range (pullCount (canonicalAction hk n mean table) a (t+1)), table (j+1) a) / (pullCount (canonicalAction hk n mean table) a (t+1) : ℝ)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-0db3e5eae8fe","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1395,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistory_empiricalMean {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) (a : Fin k) : ETC.realHistoryEmpMean t (canonicalHistory hk n mean table t) a = (∑ j ∈ Finset.range (pullCount (canonicalAction hk n mean table) a (t+1)), table (j+1) a) / (pullCount (canonicalAction hk n mean table) a (t+1) : ℝ)","missing":[],"search":"canonicalhistory_empiricalmean banditrlproof.moss.canonicalhistory_empiricalmean theorem canonicalhistory_empiricalmean {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (table : ucb.armrewardstream k) (t : ℕ) (a : fin k) : etc.realhistoryempmean t (canonicalhistory hk n mean table t) a = (∑ j ∈ finset.range (pullcount (canonicalaction hk n mean table) a (t+1)), table (j+1) a) / (pullcount (canonicalaction hk n mean table) a (t+1) : ℝ) theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalAction_succ_eq_historyAction","label":"canonicalAction_succ_eq_historyAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalAction_succ_eq_historyAction","description":"The realized next action is exactly the common finite-history MOSS selector.","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-dc1c49bc3901","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1396,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalAction_succ_eq_historyAction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) : canonicalAction hk n mean table (t+1) = historyAction hk n t (canonicalHistory hk n mean table t)","missing":[],"search":"canonicalaction_succ_eq_historyaction banditrlproof.moss.canonicalaction_succ_eq_historyaction the realized next action is exactly the common finite-history moss selector. theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalAction_zero","label":"canonicalAction_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalAction_zero","description":"theorem canonicalAction_zero {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) : canonicalAction hk n mean table 0 = ⟨0, hk⟩","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-03edc0b5629e","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1397,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalAction_zero {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) : canonicalAction hk n mean table 0 = ⟨0, hk⟩","missing":[],"search":"canonicalaction_zero banditrlproof.moss.canonicalaction_zero theorem canonicalaction_zero {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (table : ucb.armrewardstream k) : canonicalaction hk n mean table 0 = ⟨0, hk⟩ theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_canonicalAction","label":"measurable_canonicalAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_canonicalAction","description":"theorem measurable_canonicalAction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (fun table => canonicalAction hk n mean table t)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-a3ddb8554e38","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1398,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalAction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (fun table => canonicalAction hk n mean table t)","missing":[],"search":"measurable_canonicalaction banditrlproof.moss.measurable_canonicalaction theorem measurable_canonicalaction {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) : measurable (fun table => canonicalaction hk n mean table t) theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_canonicalReward","label":"measurable_canonicalReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_canonicalReward","description":"theorem measurable_canonicalReward {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (fun table => canonicalReward hk n mean table t)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-62faaee829eb","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1399,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalReward {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (fun table => canonicalReward hk n mean table t)","missing":[],"search":"measurable_canonicalreward banditrlproof.moss.measurable_canonicalreward theorem measurable_canonicalreward {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) : measurable (fun table => canonicalreward hk n mean table t) theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_canonicalHistory","label":"measurable_canonicalHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_canonicalHistory","description":"theorem measurable_canonicalHistory {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (fun table => canonicalHistory hk n mean table t)","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-cdbbd94501be","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1400,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalHistory {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (fun table => canonicalHistory hk n mean table t)","missing":[],"search":"measurable_canonicalhistory banditrlproof.moss.measurable_canonicalhistory theorem measurable_canonicalhistory {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) : measurable (fun table => canonicalhistory hk n mean table t) theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory_succ","label":"canonicalHistory_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory_succ","description":"theorem canonicalHistory_succ {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) : canonicalHistory hk n mean table (t+1) = History.extendPairHistorySucc (canonicalHistory hk n mean table t) (historyAction hk n t (canonicalHistory hk n mean table t), table (ETC.realHistoryPullCount t (canonicalHistory hk n mean table t) (historyAction hk n t (canonicalHistory hk n mean table t))+…","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-06ee024fb63a","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1401,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistory_succ {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (t : ℕ) : canonicalHistory hk n mean table (t+1) = History.extendPairHistorySucc (canonicalHistory hk n mean table t) (historyAction hk n t (canonicalHistory hk n mean table t), table (ETC.realHistoryPullCount t (canonicalHistory hk n mean table t) (historyAction hk n t (canonicalHistory hk n mean table t))+1) (historyAction hk n t (canonicalHistory hk n mean table t)))","missing":[],"search":"canonicalhistory_succ banditrlproof.moss.canonicalhistory_succ theorem canonicalhistory_succ {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (table : ucb.armrewardstream k) (t : ℕ) : canonicalhistory hk n mean table (t+1) = history.extendpairhistorysucc (canonicalhistory hk n mean table t) (historyaction hk n t (canonicalhistory hk n mean table t), table (etc.realhistorypullcount t (canonicalhistory hk n mean table t) (historyaction hk n t (canonicalhistory hk n mean table t))+1) (historyaction hk n t (canonicalhistory hk n mean table t))) theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory_eq_of_eq_consumed","label":"canonicalHistory_eq_of_eq_consumed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory_eq_of_eq_consumed","description":"Changing rewards that have not been consumed cannot change the observed history.","url":"../modules/banditrlproof-algorithms-mosscanonicalhistory/index.html#decl-d5a457036b67","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","order":1402,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalHistory"],["Source","BanditRLProof/Algorithms/MOSSCanonicalHistory.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistory_eq_of_eq_consumed {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (table table' : UCB.ArmRewardStream k) (t : ℕ) (hagrees : ∀ a j, j < pullCount (canonicalAction hk n mean table) a (t+1) → table (j+1) a = table' (j+1) a) : canonicalHistory hk n mean table t = canonicalHistory hk n mean table' t","missing":[],"search":"canonicalhistory_eq_of_eq_consumed banditrlproof.moss.canonicalhistory_eq_of_eq_consumed changing rewards that have not been consumed cannot change the observed history. theorem compiled","shard":"modules/1893505bcdd62921.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.centeredRewardTable","label":"centeredRewardTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.centeredRewardTable","description":"One-based centered reward table; sample zero is unused by MOSS.","url":"../modules/banditrlproof-algorithms-mosscanonicalreward/index.html#decl-ca69bd137043","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","order":1403,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSCanonicalReward"],["Source","BanditRLProof/Algorithms/MOSSCanonicalReward.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def centeredRewardTable {k : ℕ} (mean : Fin k → ℝ) : Fin k → ℕ → UCB.ArmRewardStream k → ℝ","missing":[],"search":"centeredrewardtable banditrlproof.moss.centeredrewardtable one-based centered reward table; sample zero is unused by moss. definition compiled","shard":"modules/6321412dacffae06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_canonicalReward_regret_le","label":"integral_canonicalReward_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_canonicalReward_regret_le","description":"Exact MOSS bound on the canonical product space of arbitrary arm laws.","url":"../modules/banditrlproof-algorithms-mosscanonicalreward/index.html#decl-3986cf4cbe2a","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","order":1404,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalReward"],["Source","BanditRLProof/Algorithms/MOSSCanonicalReward.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalReward_regret_le {k : ℕ} (hk : 0 < k) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) (hmean : ∀ a, ∫ r, r ∂ν a = mean a) (hsubG : ∀ a, HasSubgaussianMGF (fun r => r-mean a) 1 (ν a)) : (∫ table, realMeanRegret mean (streamTrace hk n mean (centeredRewardTable mean) table) n ∂UCB.armStreamMeasure ν) ≤ 39*Real.sqrt ((n : ℝ)*k) + ∑ a, (mean best-mean a)","missing":[],"search":"integral_canonicalreward_regret_le banditrlproof.moss.integral_canonicalreward_regret_le exact moss bound on the canonical product space of arbitrary arm laws. theorem compiled","shard":"modules/6321412dacffae06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.mean_add_centeredRewardTable_average","label":"mean_add_centeredRewardTable_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.mean_add_centeredRewardTable_average","description":"Centering is analytical only: positive-count empirical means use raw rewards.","url":"../modules/banditrlproof-algorithms-mosscanonicalreward/index.html#decl-ba83898fdb4a","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","order":1405,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalReward"],["Source","BanditRLProof/Algorithms/MOSSCanonicalReward.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_add_centeredRewardTable_average {k : ℕ} (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) (a : Fin k) (s : ℕ) (hs : 0 < s) : mean a + streamMean (centeredRewardTable mean a) table s = (∑ j ∈ Finset.range s, table (j+1) a)/(s : ℝ)","missing":[],"search":"mean_add_centeredrewardtable_average banditrlproof.moss.mean_add_centeredrewardtable_average centering is analytical only: positive-count empirical means use raw rewards. theorem compiled","shard":"modules/6321412dacffae06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamTrace_pullCount_pos","label":"streamTrace_pullCount_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.streamTrace_pullCount_pos","description":"After initialization each arm has a positive realized count.","url":"../modules/banditrlproof-algorithms-mosscanonicalreward/index.html#decl-03d55e6c3370","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","order":1406,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalReward"],["Source","BanditRLProof/Algorithms/MOSSCanonicalReward.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem streamTrace_pullCount_pos {Ω : Type*} {k : ℕ} (hk : 0 < k) (n t : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (a : Fin k) (ht : k ≤ t) : 0 < pullCount (streamTrace hk n mean X ω) a t","missing":[],"search":"streamtrace_pullcount_pos banditrlproof.moss.streamtrace_pullcount_pos after initialization each arm has a positive realized count. theorem compiled","shard":"modules/6321412dacffae06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalReward_action_eq_raw","label":"canonicalReward_action_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalReward_action_eq_raw","description":"The centered execution selects solely from raw empirical rewards and counts. The unknown means cancel; zero-count states are handled by initialization.","url":"../modules/banditrlproof-algorithms-mosscanonicalreward/index.html#decl-3cb0314906a1","parent":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","order":1407,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSCanonicalReward"],["Source","BanditRLProof/Algorithms/MOSSCanonicalReward.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalReward_action_eq_raw {k : ℕ} (hk : 0 < k) (n t : ℕ) (mean : Fin k → ℝ) (table : UCB.ArmRewardStream k) : streamTrace hk n mean (centeredRewardTable mean) table t = action hk n t (fun a => (∑ j ∈ Finset.range (pullCount (streamTrace hk n mean (centeredRewardTable mean) table) a t), table (j+1) a)/(pullCount (streamTrace hk n mean (centeredRewardTable mean) table) a t : ℝ)) (fun a => pullCount (streamTrace hk n mean (centeredRewardTable mean) table) a t)","missing":[],"search":"canonicalreward_action_eq_raw banditrlproof.moss.canonicalreward_action_eq_raw the centered execution selects solely from raw empirical rewards and counts. the unknown means cancel; zero-count states are handled by initialization. theorem compiled","shard":"modules/6321412dacffae06.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.map_condition_reward_eq_compProd","label":"map_condition_reward_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.map_condition_reward_eq_compProd","description":"theorem map_condition_reward_eq_compProd {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Measure.map (fun table => (canonicalCondition hk n mean t table, canonicalReward hk n mean table (t+1))) (UCB.armStreamMeasure ν) = (Measure.map (canonicalCondition hk n mean t) (UCB.armStreamMeasure ν)).compProd (UCB.armStreamSelectedRewardKernel t ν)","url":"../modules/banditrlproof-algorithms-mossconditionalreward/index.html#decl-2ab1c11b2c9d","parent":"module:BanditRLProof.Algorithms.MOSSConditionalReward","order":1408,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSConditionalReward"],["Source","BanditRLProof/Algorithms/MOSSConditionalReward.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem map_condition_reward_eq_compProd {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Measure.map (fun table => (canonicalCondition hk n mean t table, canonicalReward hk n mean table (t+1))) (UCB.armStreamMeasure ν) = (Measure.map (canonicalCondition hk n mean t) (UCB.armStreamMeasure ν)).compProd (UCB.armStreamSelectedRewardKernel t ν)","missing":[],"search":"map_condition_reward_eq_compprod banditrlproof.moss.map_condition_reward_eq_compprod theorem map_condition_reward_eq_compprod {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] : measure.map (fun table => (canonicalcondition hk n mean t table, canonicalreward hk n mean table (t+1))) (ucb.armstreammeasure ν) = (measure.map (canonicalcondition hk n mean t) (ucb.armstreammeasure ν)).compprod (ucb.armstreamselectedrewardkernel t ν) theorem compiled","shard":"modules/56fa1bb6573cdf39.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalReward_condDistrib","label":"canonicalReward_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalReward_condDistrib","description":"The actual successor reward has the selected arm law conditionally on history/action.","url":"../modules/banditrlproof-algorithms-mossconditionalreward/index.html#decl-7f8681fecc25","parent":"module:BanditRLProof.Algorithms.MOSSConditionalReward","order":1409,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSConditionalReward"],["Source","BanditRLProof/Algorithms/MOSSConditionalReward.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalReward_condDistrib {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Filter.EventuallyEq (ae ((UCB.armStreamMeasure ν).map (canonicalCondition hk n mean t))) (condDistrib (fun table => canonicalReward hk n mean table (t+1)) (canonicalCondition hk n mean t) (UCB.armStreamMeasure ν)) (UCB.armStreamSelectedRewardKernel t ν)","missing":[],"search":"canonicalreward_conddistrib banditrlproof.moss.canonicalreward_conddistrib the actual successor reward has the selected arm law conditionally on history/action. theorem compiled","shard":"modules/56fa1bb6573cdf39.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.log_sixtyFour_le","label":"log_sixtyFour_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.log_sixtyFour_le","description":"theorem log_sixtyFour_le : log (64 : ℝ) ≤ 17/4","url":"../modules/banditrlproof-algorithms-mossconstants/index.html#decl-7d701a2f3067","parent":"module:BanditRLProof.Algorithms.MOSSConstants","order":1410,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSConstants"],["Source","BanditRLProof/Algorithms/MOSSConstants.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem log_sixtyFour_le : log (64 : ℝ) ≤ 17/4","missing":[],"search":"log_sixtyfour_le banditrlproof.moss.log_sixtyfour_le theorem log_sixtyfour_le : log (64 : ℝ) ≤ 17/4 theorem compiled","shard":"modules/7201346a5b40b9e9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.largeGap_constant_fifteen","label":"largeGap_constant_fifteen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.largeGap_constant_fifteen","description":"Dimensionless large-gap numerical estimate used in source Theorem 9.1.","url":"../modules/banditrlproof-algorithms-mossconstants/index.html#decl-aa151f210ade","parent":"module:BanditRLProof.Algorithms.MOSSConstants","order":1411,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSConstants"],["Source","BanditRLProof/Algorithms/MOSSConstants.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem largeGap_constant_fifteen (q : ℝ) (hq : 64 ≤ q) : (1+8*(2*log q+sqrt (Real.pi*(2*log q))+1))/sqrt q ≤ 15","missing":[],"search":"largegap_constant_fifteen banditrlproof.moss.largegap_constant_fifteen dimensionless large-gap numerical estimate used in source theorem 9.1. theorem compiled","shard":"modules/7201346a5b40b9e9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.largeGap_scaled_constant_fifteen","label":"largeGap_scaled_constant_fifteen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.largeGap_scaled_constant_fifteen","description":"theorem largeGap_scaled_constant_fifteen (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) : gap*(1/gap^2+1+(8/gap^2)*(2*logPlus (gap^2/δ)+ sqrt (Real.pi*(2*logPlus (gap^2/δ)))+1)) ≤ gap+15/sqrt δ","url":"../modules/banditrlproof-algorithms-mossconstants/index.html#decl-53d9c743045a","parent":"module:BanditRLProof.Algorithms.MOSSConstants","order":1412,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSConstants"],["Source","BanditRLProof/Algorithms/MOSSConstants.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem largeGap_scaled_constant_fifteen (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) : gap*(1/gap^2+1+(8/gap^2)*(2*logPlus (gap^2/δ)+ sqrt (Real.pi*(2*logPlus (gap^2/δ)))+1)) ≤ gap+15/sqrt δ","missing":[],"search":"largegap_scaled_constant_fifteen banditrlproof.moss.largegap_scaled_constant_fifteen theorem largegap_scaled_constant_fifteen (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) : gap*(1/gap^2+1+(8/gap^2)*(2*logplus (gap^2/δ)+ sqrt (real.pi*(2*logplus (gap^2/δ)))+1)) ≤ gap+15/sqrt δ theorem compiled","shard":"modules/7201346a5b40b9e9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamMean","label":"streamMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.streamMean","description":"def streamMean (X : ℕ → Ω → ℝ) (ω : Ω) (s : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-d263f576227e","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1413,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def streamMean (X : ℕ → Ω → ℝ) (ω : Ω) (s : ℕ) : ℝ","missing":[],"search":"streammean banditrlproof.moss.streammean def streammean (x : ℕ → ω → ℝ) (ω : ω) (s : ℕ) : ℝ definition compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.fixedLogExceedanceCount_eq_fixedRadiusCount","label":"fixedLogExceedanceCount_eq_fixedRadiusCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.fixedLogExceedanceCount_eq_fixedRadiusCount","description":"theorem fixedLogExceedanceCount_eq_fixedRadiusCount (X : ℕ → Ω → ℝ) (δ gap : ℝ) (n : ℕ) (ω : Ω) : fixedLogExceedanceCount (streamMean X ω) δ gap n = Concentration.fixedRadiusCount X (2*logPlus (gap^2/δ)) (gap/2) n ω","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-33bf87ef0bd6","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1414,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem fixedLogExceedanceCount_eq_fixedRadiusCount (X : ℕ → Ω → ℝ) (δ gap : ℝ) (n : ℕ) (ω : Ω) : fixedLogExceedanceCount (streamMean X ω) δ gap n = Concentration.fixedRadiusCount X (2*logPlus (gap^2/δ)) (gap/2) n ω","missing":[],"search":"fixedlogexceedancecount_eq_fixedradiuscount banditrlproof.moss.fixedlogexceedancecount_eq_fixedradiuscount theorem fixedlogexceedancecount_eq_fixedradiuscount (x : ℕ → ω → ℝ) (δ gap : ℝ) (n : ℕ) (ω : ω) : fixedlogexceedancecount (streammean x ω) δ gap n = concentration.fixedradiuscount x (2*logplus (gap^2/δ)) (gap/2) n ω theorem compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integrable_indexExceedanceCount","label":"integrable_indexExceedanceCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integrable_indexExceedanceCount","description":"Finite measurable indicator counts are integrable, without tail assumptions.","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-b210663754f1","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1415,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_indexExceedanceCount (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (δ gap : ℝ) (n : ℕ) : Integrable (fun ω => indexExceedanceCount (streamMean X ω) δ gap n) μ","missing":[],"search":"integrable_indexexceedancecount banditrlproof.moss.integrable_indexexceedancecount finite measurable indicator counts are integrable, without tail assumptions. theorem compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_indexExceedanceCount_le_sharp","label":"integral_indexExceedanceCount_le_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_indexExceedanceCount_le_sharp","description":"theorem integral_indexExceedanceCount_le_sharp (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : δ < gap^2) (n : ℕ) : (∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ 1/gap^2 + (8/gap^2)*(2*logPlus (gap^2/δ)+sqrt (Real.pi*(2*logPlus (gap^2/δ)))+1)","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-871ed684fe0e","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1416,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_indexExceedanceCount_le_sharp (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : δ < gap^2) (n : ℕ) : (∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ 1/gap^2 + (8/gap^2)*(2*logPlus (gap^2/δ)+sqrt (Real.pi*(2*logPlus (gap^2/δ)))+1)","missing":[],"search":"integral_indexexceedancecount_le_sharp banditrlproof.moss.integral_indexexceedancecount_le_sharp theorem integral_indexexceedancecount_le_sharp (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : δ < gap^2) (n : ℕ) : (∫ ω, indexexceedancecount (streammean x ω) δ gap n ∂μ) ≤ 1/gap^2 + (8/gap^2)*(2*logplus (gap^2/δ)+sqrt (real.pi*(2*logplus (gap^2/δ)))+1) theorem compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_indexExceedanceCount_le","label":"integral_indexExceedanceCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_indexExceedanceCount_le","description":"theorem integral_indexExceedanceCount_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : δ < gap^2) (n : ℕ) : (∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ 1/gap^2 + 1 + (8/gap^2)*(2*logPlus (gap^2/δ)+sqrt (Real.pi*(2*logPlus (gap^2/δ)))+1)","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-31588315ec73","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1417,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_indexExceedanceCount_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : δ < gap^2) (n : ℕ) : (∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ 1/gap^2 + 1 + (8/gap^2)*(2*logPlus (gap^2/δ)+sqrt (Real.pi*(2*logPlus (gap^2/δ)))+1)","missing":[],"search":"integral_indexexceedancecount_le banditrlproof.moss.integral_indexexceedancecount_le theorem integral_indexexceedancecount_le (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : δ < gap^2) (n : ℕ) : (∫ ω, indexexceedancecount (streammean x ω) δ gap n ∂μ) ≤ 1/gap^2 + 1 + (8/gap^2)*(2*logplus (gap^2/δ)+sqrt (real.pi*(2*logplus (gap^2/δ)))+1) theorem compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.gap_mul_integral_indexExceedanceCount_le_sharp","label":"gap_mul_integral_indexExceedanceCount_le_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.gap_mul_integral_indexExceedanceCount_le_sharp","description":"theorem gap_mul_integral_indexExceedanceCount_le_sharp (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) (n : ℕ) : gap*(∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ 15/sqrt δ","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-8fa5df794872","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1418,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_mul_integral_indexExceedanceCount_le_sharp (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) (n : ℕ) : gap*(∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ 15/sqrt δ","missing":[],"search":"gap_mul_integral_indexexceedancecount_le_sharp banditrlproof.moss.gap_mul_integral_indexexceedancecount_le_sharp theorem gap_mul_integral_indexexceedancecount_le_sharp (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) (n : ℕ) : gap*(∫ ω, indexexceedancecount (streammean x ω) δ gap n ∂μ) ≤ 15/sqrt δ theorem compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.gap_mul_integral_indexExceedanceCount_le","label":"gap_mul_integral_indexExceedanceCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.gap_mul_integral_indexExceedanceCount_le","description":"Source large-gap weighted occupancy bound, before selected-count transport.","url":"../modules/banditrlproof-algorithms-mossexpectedoccupancy/index.html#decl-4f3ccb667d4b","parent":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","order":1419,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedOccupancy"],["Source","BanditRLProof/Algorithms/MOSSExpectedOccupancy.lean:90"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_mul_integral_indexExceedanceCount_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (hlarge : 8*sqrt δ ≤ gap) (n : ℕ) : gap*(∫ ω, indexExceedanceCount (streamMean X ω) δ gap n ∂μ) ≤ gap+15/sqrt δ","missing":[],"search":"gap_mul_integral_indexexceedancecount_le banditrlproof.moss.gap_mul_integral_indexexceedancecount_le source large-gap weighted occupancy bound, before selected-count transport. theorem compiled","shard":"modules/ad66c789e303c560.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_streamTrace_regret_le","label":"integral_streamTrace_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_streamTrace_regret_le","description":"Theorem 9.1's exact constant for the concrete centered reward-table execution. Identification with the common bandit history law is a separate theorem.","url":"../modules/banditrlproof-algorithms-mossexpectedregret/index.html#decl-3fcdc369f42a","parent":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","order":1420,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSExpectedRegret"],["Source","BanditRLProof/Algorithms/MOSSExpectedRegret.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_streamTrace_regret_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) (hXm : ∀ a i, StronglyMeasurable (X a i)) (hind : ∀ a, iIndepFun (X a) μ) (hmean : ∀ a i, ∫ ω, X a i ω ∂μ = 0) (hsubG : ∀ a i, HasSubgaussianMGF (X a i) 1 μ) : (∫ ω, realMeanRegret mean (streamTrace hk n mean X ω) n ∂μ) ≤ 39*sqrt ((n : ℝ)*k) + ∑ a, (mean best-mean a)","missing":[],"search":"integral_streamtrace_regret_le banditrlproof.moss.integral_streamtrace_regret_le theorem 9.1's exact constant for the concrete centered reward-table execution. identification with the common bandit history law is a separate theorem. theorem compiled","shard":"modules/a2b8202084a39122.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.historyAction","label":"historyAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.historyAction","description":"Source MOSS action after an inclusive finite action/reward history.","url":"../modules/banditrlproof-algorithms-mosshistory/index.html#decl-1b8b7af2a0a6","parent":"module:BanditRLProof.Algorithms.MOSSHistory","order":1421,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSHistory"],["Source","BanditRLProof/Algorithms/MOSSHistory.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def historyAction {k : ℕ} (hk : 0 < k) (n t : ℕ) (history : History.FinitePairHistory (Fin k) ℝ t) : Fin k","missing":[],"search":"historyaction banditrlproof.moss.historyaction source moss action after an inclusive finite action/reward history. definition compiled","shard":"modules/69098d034e7a2950.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_historyAction","label":"measurable_historyAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_historyAction","description":"theorem measurable_historyAction {k : ℕ} (hk : 0 < k) (n t : ℕ) : Measurable (historyAction hk n t)","url":"../modules/banditrlproof-algorithms-mosshistory/index.html#decl-869c9792ff88","parent":"module:BanditRLProof.Algorithms.MOSSHistory","order":1422,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistory"],["Source","BanditRLProof/Algorithms/MOSSHistory.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyAction {k : ℕ} (hk : 0 < k) (n t : ℕ) : Measurable (historyAction hk n t)","missing":[],"search":"measurable_historyaction banditrlproof.moss.measurable_historyaction theorem measurable_historyaction {k : ℕ} (hk : 0 < k) (n t : ℕ) : measurable (historyaction hk n t) theorem compiled","shard":"modules/69098d034e7a2950.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.historyAlgorithm","label":"historyAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.historyAlgorithm","description":"Concrete MOSS policy on the history interface shared by Chapter 13's minimax regret functional. Initial action is arm zero.","url":"../modules/banditrlproof-algorithms-mosshistory/index.html#decl-11f466add90d","parent":"module:BanditRLProof.Algorithms.MOSSHistory","order":1423,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSHistory"],["Source","BanditRLProof/Algorithms/MOSSHistory.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"noncomputable def historyAlgorithm {k : ℕ} (hk : 0 < k) (n : ℕ) : Thompson.HistoryAlgorithm (Fin k) ℝ where","missing":[],"search":"historyalgorithm banditrlproof.moss.historyalgorithm concrete moss policy on the history interface shared by chapter 13's minimax regret functional. initial action is arm zero. definition compiled","shard":"modules/69098d034e7a2950.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.historyAlgorithm_policy_apply","label":"historyAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.historyAlgorithm_policy_apply","description":"@[simp] theorem historyAlgorithm_policy_apply {k : ℕ} (hk : 0 < k) (n t : ℕ) (history : History.FinitePairHistory (Fin k) ℝ t) : (historyAlgorithm hk n).policy t history = Measure.dirac (historyAction hk n t history)","url":"../modules/banditrlproof-algorithms-mosshistory/index.html#decl-5589021a9735","parent":"module:BanditRLProof.Algorithms.MOSSHistory","order":1424,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistory"],["Source","BanditRLProof/Algorithms/MOSSHistory.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem historyAlgorithm_policy_apply {k : ℕ} (hk : 0 < k) (n t : ℕ) (history : History.FinitePairHistory (Fin k) ℝ t) : (historyAlgorithm hk n).policy t history = Measure.dirac (historyAction hk n t history)","missing":[],"search":"historyalgorithm_policy_apply banditrlproof.moss.historyalgorithm_policy_apply @[simp] theorem historyalgorithm_policy_apply {k : ℕ} (hk : 0 < k) (n t : ℕ) (history : history.finitepairhistory (fin k) ℝ t) : (historyalgorithm hk n).policy t history = measure.dirac (historyaction hk n t history) theorem compiled","shard":"modules/69098d034e7a2950.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.historyAction_initialization","label":"historyAction_initialization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.historyAction_initialization","description":"Subsequent initialization rounds select the corresponding arm.","url":"../modules/banditrlproof-algorithms-mosshistory/index.html#decl-6c51e9330891","parent":"module:BanditRLProof.Algorithms.MOSSHistory","order":1425,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistory"],["Source","BanditRLProof/Algorithms/MOSSHistory.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem historyAction_initialization {k : ℕ} (hk : 0 < k) (n t : ℕ) (history : History.FinitePairHistory (Fin k) ℝ t) (ht : t + 1 < k) : historyAction hk n t history = ⟨t + 1, ht⟩","missing":[],"search":"historyaction_initialization banditrlproof.moss.historyaction_initialization subsequent initialization rounds select the corresponding arm. theorem compiled","shard":"modules/69098d034e7a2950.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.historyAction_index_max","label":"historyAction_index_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.historyAction_index_max","description":"Exact post-initialization source index maximality on visible history.","url":"../modules/banditrlproof-algorithms-mosshistory/index.html#decl-272c4db505e5","parent":"module:BanditRLProof.Algorithms.MOSSHistory","order":1426,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistory"],["Source","BanditRLProof/Algorithms/MOSSHistory.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem historyAction_index_max {k : ℕ} (hk : 0 < k) (n t : ℕ) (history : History.FinitePairHistory (Fin k) ℝ t) (ht : k ≤ t + 1) (a : Fin k) : index n (ETC.realHistoryEmpMean t history) (ETC.realHistoryPullCount t history) a ≤ index n (ETC.realHistoryEmpMean t history) (ETC.realHistoryPullCount t history) (historyAction hk n t history)","missing":[],"search":"historyaction_index_max banditrlproof.moss.historyaction_index_max exact post-initialization source index maximality on visible history. theorem compiled","shard":"modules/69098d034e7a2950.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonical_initialPair_map","label":"canonical_initialPair_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonical_initialPair_map","description":"theorem canonical_initialPair_map {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : (UCB.armStreamMeasure ν).map (fun table => (canonicalAction hk n mean table 0, canonicalReward hk n mean table 0)) = (historyAlgorithm hk n).initialAction.compProd ν","url":"../modules/banditrlproof-algorithms-mosshistorylaw/index.html#decl-83c620d0103d","parent":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","order":1427,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryLaw"],["Source","BanditRLProof/Algorithms/MOSSHistoryLaw.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonical_initialPair_map {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : (UCB.armStreamMeasure ν).map (fun table => (canonicalAction hk n mean table 0, canonicalReward hk n mean table 0)) = (historyAlgorithm hk n).initialAction.compProd ν","missing":[],"search":"canonical_initialpair_map banditrlproof.moss.canonical_initialpair_map theorem canonical_initialpair_map {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] : (ucb.armstreammeasure ν).map (fun table => (canonicalaction hk n mean table 0, canonicalreward hk n mean table 0)) = (historyalgorithm hk n).initialaction.compprod ν theorem compiled","shard":"modules/ccacf09fe51dcb36.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonical_action_condDistrib","label":"canonical_action_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonical_action_condDistrib","description":"theorem canonical_action_condDistrib {k : ℕ} [NeZero k] (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Filter.EventuallyEq (ae ((UCB.armStreamMeasure ν).map (fun table => canonicalHistory hk n mean table t))) (condDistrib (fun table => canonicalAction hk n mean table (t+1)) (fun table => canonicalHistory hk n mean table t) (UCB.armStreamMeasure ν)) ((historyAlgorithm hk n…","url":"../modules/banditrlproof-algorithms-mosshistorylaw/index.html#decl-d4709a093ed3","parent":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","order":1428,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryLaw"],["Source","BanditRLProof/Algorithms/MOSSHistoryLaw.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonical_action_condDistrib {k : ℕ} [NeZero k] (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Filter.EventuallyEq (ae ((UCB.armStreamMeasure ν).map (fun table => canonicalHistory hk n mean table t))) (condDistrib (fun table => canonicalAction hk n mean table (t+1)) (fun table => canonicalHistory hk n mean table t) (UCB.armStreamMeasure ν)) ((historyAlgorithm hk n).policy t)","missing":[],"search":"canonical_action_conddistrib banditrlproof.moss.canonical_action_conddistrib theorem canonical_action_conddistrib {k : ℕ} [nezero k] (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] : filter.eventuallyeq (ae ((ucb.armstreammeasure ν).map (fun table => canonicalhistory hk n mean table t))) (conddistrib (fun table => canonicalaction hk n mean table (t+1)) (fun table => canonicalhistory hk n mean table t) (ucb.armstreammeasure ν)) ((historyalgorithm hk n).policy t) theorem compiled","shard":"modules/ccacf09fe51dcb36.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonical_historySequence","label":"canonical_historySequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonical_historySequence","description":"theorem canonical_historySequence {k : ℕ} [NeZero k] (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Thompson.IsHistoryAlgorithmEnvironmentSequence (UCB.armStreamMeasure ν) (canonicalAction hk n mean) (canonicalReward hk n mean) (historyAlgorithm hk n) (LowerBounds.stationaryBanditHistoryEnvironment ν) where","url":"../modules/banditrlproof-algorithms-mosshistorylaw/index.html#decl-231aaacac590","parent":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","order":1429,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryLaw"],["Source","BanditRLProof/Algorithms/MOSSHistoryLaw.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonical_historySequence {k : ℕ} [NeZero k] (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] : Thompson.IsHistoryAlgorithmEnvironmentSequence (UCB.armStreamMeasure ν) (canonicalAction hk n mean) (canonicalReward hk n mean) (historyAlgorithm hk n) (LowerBounds.stationaryBanditHistoryEnvironment ν) where","missing":[],"search":"canonical_historysequence banditrlproof.moss.canonical_historysequence theorem canonical_historysequence {k : ℕ} [nezero k] (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] : thompson.ishistoryalgorithmenvironmentsequence (ucb.armstreammeasure ν) (canonicalaction hk n mean) (canonicalreward hk n mean) (historyalgorithm hk n) (lowerbounds.stationarybandithistoryenvironment ν) where theorem compiled","shard":"modules/ccacf09fe51dcb36.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.map_canonicalHistory_eq","label":"map_canonicalHistory_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.map_canonicalHistory_eq","description":"Canonical reward-table histories have exactly the common bandit history law.","url":"../modules/banditrlproof-algorithms-mosshistorylaw/index.html#decl-0f70037fd09a","parent":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","order":1430,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryLaw"],["Source","BanditRLProof/Algorithms/MOSSHistoryLaw.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem map_canonicalHistory_eq {k : ℕ} [NeZero k] (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (t : ℕ) : (UCB.armStreamMeasure ν).map (fun table => canonicalHistory hk n mean table t) = LowerBounds.canonicalBanditHistoryMeasure (historyAlgorithm hk n) ν t","missing":[],"search":"map_canonicalhistory_eq banditrlproof.moss.map_canonicalhistory_eq canonical reward-table histories have exactly the common bandit history law. theorem compiled","shard":"modules/ccacf09fe51dcb36.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.finiteHistoryPullCountENNReal_trace","label":"finiteHistoryPullCountENNReal_trace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.finiteHistoryPullCountENNReal_trace","description":"Inclusive finite-history counts agree with trace counts through t+1.","url":"../modules/banditrlproof-algorithms-mosshistoryregret/index.html#decl-2888b0a8db53","parent":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","order":1431,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryRegret"],["Source","BanditRLProof/Algorithms/MOSSHistoryRegret.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPullCountENNReal_trace {k : ℕ} {Reward : Type*} (action : ActionTrace (Fin k)) (reward : RewardTrace Reward) (t : ℕ) (a : Fin k) : LowerBounds.finiteHistoryPullCountENNReal t (History.finitePairHistoryOfTrace action reward t) a = (pullCount action a (t+1) : ℝ≥0∞)","missing":[],"search":"finitehistorypullcountennreal_trace banditrlproof.finitehistorypullcountennreal_trace inclusive finite-history counts agree with trace counts through t+1. theorem compiled","shard":"modules/c8da34ff2e770ab4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory_gapRegret_toReal","label":"canonicalHistory_gapRegret_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory_gapRegret_toReal","description":"The history gap functional is exactly the executed real pseudo-regret.","url":"../modules/banditrlproof-algorithms-mosshistoryregret/index.html#decl-1c642661df40","parent":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","order":1432,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryRegret"],["Source","BanditRLProof/Algorithms/MOSSHistoryRegret.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistory_gapRegret_toReal {k : ℕ} (hk : 0 < k) (n t : ℕ) (mean : Fin k → ℝ) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) (table : UCB.ArmRewardStream k) : (LowerBounds.finiteHistoryGapPseudoRegret (fun a => mean best-mean a) t (canonicalHistory hk n mean table t)).toReal = realMeanRegret mean (canonicalAction hk n mean table) (t+1)","missing":[],"search":"canonicalhistory_gapregret_toreal banditrlproof.moss.canonicalhistory_gapregret_toreal the history gap functional is exactly the executed real pseudo-regret. theorem compiled","shard":"modules/c8da34ff2e770ab4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalGapExpectedRegret_eq_integral","label":"canonicalGapExpectedRegret_eq_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalGapExpectedRegret_eq_integral","description":"Exact transport of the common-history expected regret to the reward table.","url":"../modules/banditrlproof-algorithms-mosshistoryregret/index.html#decl-afc1435c4f68","parent":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","order":1433,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryRegret"],["Source","BanditRLProof/Algorithms/MOSSHistoryRegret.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalGapExpectedRegret_eq_integral {k : ℕ} [NeZero k] (hk : 0 < k) (n t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (mean : Fin k → ℝ) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) : LowerBounds.canonicalGapExpectedPseudoRegretReal (historyAlgorithm hk n) ν (fun a => mean best-mean a) t = ∫ table, realMeanRegret mean (canonicalAction hk n mean table) (t+1) ∂UCB.armStreamMeasure ν","missing":[],"search":"canonicalgapexpectedregret_eq_integral banditrlproof.moss.canonicalgapexpectedregret_eq_integral exact transport of the common-history expected regret to the reward table. theorem compiled","shard":"modules/c8da34ff2e770ab4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalGapExpectedRegret_le","label":"canonicalGapExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalGapExpectedRegret_le","description":"Source constant 39, now for the same history-law regret used by lower bounds.","url":"../modules/banditrlproof-algorithms-mosshistoryregret/index.html#decl-238d04640441","parent":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","order":1434,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSHistoryRegret"],["Source","BanditRLProof/Algorithms/MOSSHistoryRegret.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem canonicalGapExpectedRegret_le {k : ℕ} [NeZero k] (hk : 0 < k) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (t : ℕ) (hkt : k ≤ t+1) (mean : Fin k → ℝ) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) (hmean : ∀ a, ∫ r, r ∂ν a = mean a) (hsubG : ∀ a, HasSubgaussianMGF (fun r => r-mean a) 1 (ν a)) : LowerBounds.canonicalGapExpectedPseudoRegretReal (historyAlgorithm hk (t+1)) ν (fun a => mean best-mean a) t ≤ 39*Real.sqrt (((t+1 : ℕ) : ℝ)*k) + ∑ a, (mean best-mean a)","missing":[],"search":"canonicalgapexpectedregret_le banditrlproof.moss.canonicalgapexpectedregret_le source constant 39, now for the same history-law regret used by lower bounds. theorem compiled","shard":"modules/c8da34ff2e770ab4.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.fixedLogRadius","label":"fixedLogRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.fixedLogRadius","description":"def fixedLogRadius (δ gap : ℝ) (s : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-38f85360ff04","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1435,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def fixedLogRadius (δ gap : ℝ) (s : ℕ) : ℝ","missing":[],"search":"fixedlogradius banditrlproof.moss.fixedlogradius def fixedlogradius (δ gap : ℝ) (s : ℕ) : ℝ definition compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.sampleRadius_le_fixedLogRadius","label":"sampleRadius_le_fixedLogRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.sampleRadius_le_fixedLogRadius","description":"theorem sampleRadius_le_fixedLogRadius (δ gap : ℝ) (s : ℕ) (hδ : 0 < δ) (hg : 0 < gap) (hs : 0 < s) (hlarge : 1 ≤ (s : ℝ)*gap^2) : sqrt (4/(s : ℝ)*logPlus (1/((s : ℝ)*δ))) ≤ fixedLogRadius δ gap s","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-c72f8cd5a943","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1436,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleRadius_le_fixedLogRadius (δ gap : ℝ) (s : ℕ) (hδ : 0 < δ) (hg : 0 < gap) (hs : 0 < s) (hlarge : 1 ≤ (s : ℝ)*gap^2) : sqrt (4/(s : ℝ)*logPlus (1/((s : ℝ)*δ))) ≤ fixedLogRadius δ gap s","missing":[],"search":"sampleradius_le_fixedlogradius banditrlproof.moss.sampleradius_le_fixedlogradius theorem sampleradius_le_fixedlogradius (δ gap : ℝ) (s : ℕ) (hδ : 0 < δ) (hg : 0 < gap) (hs : 0 < s) (hlarge : 1 ≤ (s : ℝ)*gap^2) : sqrt (4/(s : ℝ)*logplus (1/((s : ℝ)*δ))) ≤ fixedlogradius δ gap s theorem compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.indexExceedanceCount","label":"indexExceedanceCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.indexExceedanceCount","description":"def indexExceedanceCount (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-5412b4ee1f5d","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1437,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def indexExceedanceCount (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) : ℝ","missing":[],"search":"indexexceedancecount banditrlproof.moss.indexexceedancecount def indexexceedancecount (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) : ℝ definition compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.pullCount_le_one_add_indexExceedanceCount","label":"pullCount_le_one_add_indexExceedanceCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.pullCount_le_one_add_indexExceedanceCount","description":"Deterministic count transport; the MOSS policy event is a separate obligation.","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-26107ed0566d","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1438,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_one_add_indexExceedanceCount {Action : Type*} [DecidableEq Action] (action : ActionTrace Action) (a : Action) (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) (hselected : ∀ t < n, action t = a → 0 < pullCount action a t → gap/2 ≤ mean (pullCount action a t) + sqrt (4/(pullCount action a t : ℝ)*logPlus (1/((pullCount action a t : ℝ)*δ)))) : (pullCount action a n : ℝ) ≤ 1 + indexExceedanceCount mean δ gap n","missing":[],"search":"pullcount_le_one_add_indexexceedancecount banditrlproof.moss.pullcount_le_one_add_indexexceedancecount deterministic count transport; the moss policy event is a separate obligation. theorem compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.fixedLogExceedanceCount","label":"fixedLogExceedanceCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.fixedLogExceedanceCount","description":"def fixedLogExceedanceCount (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-7508aebd2f44","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1439,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def fixedLogExceedanceCount (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) : ℝ","missing":[],"search":"fixedlogexceedancecount banditrlproof.moss.fixedlogexceedancecount def fixedlogexceedancecount (mean : ℕ → ℝ) (δ gap : ℝ) (n : ℕ) : ℝ definition compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.smallSampleCount","label":"smallSampleCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.smallSampleCount","description":"def smallSampleCount (gap : ℝ) (n : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-cd132859b4c4","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1440,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def smallSampleCount (gap : ℝ) (n : ℕ) : ℝ","missing":[],"search":"smallsamplecount banditrlproof.moss.smallsamplecount def smallsamplecount (gap : ℝ) (n : ℕ) : ℝ definition compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.indexExceedanceCount_le_small_add_fixed","label":"indexExceedanceCount_le_small_add_fixed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.indexExceedanceCount_le_small_add_fixed","description":"Source correction step before applying Lemma 8.2, with no stochastic premise.","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-656efb9cab44","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1441,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem indexExceedanceCount_le_small_add_fixed (mean : ℕ → ℝ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : indexExceedanceCount mean δ gap n ≤ smallSampleCount gap n + fixedLogExceedanceCount mean δ gap n","missing":[],"search":"indexexceedancecount_le_small_add_fixed banditrlproof.moss.indexexceedancecount_le_small_add_fixed source correction step before applying lemma 8.2, with no stochastic premise. theorem compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.smallSampleCount_le_horizon","label":"smallSampleCount_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.smallSampleCount_le_horizon","description":"theorem smallSampleCount_le_horizon (gap : ℝ) (n : ℕ) : smallSampleCount gap n ≤ (n : ℝ)","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-6d24322d0ebd","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1442,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem smallSampleCount_le_horizon (gap : ℝ) (n : ℕ) : smallSampleCount gap n ≤ (n : ℝ)","missing":[],"search":"smallsamplecount_le_horizon banditrlproof.moss.smallsamplecount_le_horizon theorem smallsamplecount_le_horizon (gap : ℝ) (n : ℕ) : smallsamplecount gap n ≤ (n : ℝ) theorem compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.smallSampleCount_le_inv_sq","label":"smallSampleCount_le_inv_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.smallSampleCount_le_inv_sq","description":"theorem smallSampleCount_le_inv_sq (gap : ℝ) (hg : 0 < gap) (n : ℕ) : smallSampleCount gap n ≤ 1/gap^2","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-d66574d56456","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1443,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem smallSampleCount_le_inv_sq (gap : ℝ) (hg : 0 < gap) (n : ℕ) : smallSampleCount gap n ≤ 1/gap^2","missing":[],"search":"smallsamplecount_le_inv_sq banditrlproof.moss.smallsamplecount_le_inv_sq theorem smallsamplecount_le_inv_sq (gap : ℝ) (hg : 0 < gap) (n : ℕ) : smallsamplecount gap n ≤ 1/gap^2 theorem compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.indexExceedanceCount_le_inv_sq_add_fixed","label":"indexExceedanceCount_le_inv_sq_add_fixed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.indexExceedanceCount_le_inv_sq_add_fixed","description":"theorem indexExceedanceCount_le_inv_sq_add_fixed (mean : ℕ → ℝ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : indexExceedanceCount mean δ gap n ≤ 1/gap^2 + fixedLogExceedanceCount mean δ gap n","url":"../modules/banditrlproof-algorithms-mossoccupancy/index.html#decl-78142538fcb6","parent":"module:BanditRLProof.Algorithms.MOSSOccupancy","order":1444,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOccupancy"],["Source","BanditRLProof/Algorithms/MOSSOccupancy.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem indexExceedanceCount_le_inv_sq_add_fixed (mean : ℕ → ℝ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : indexExceedanceCount mean δ gap n ≤ 1/gap^2 + fixedLogExceedanceCount mean δ gap n","missing":[],"search":"indexexceedancecount_le_inv_sq_add_fixed banditrlproof.moss.indexexceedancecount_le_inv_sq_add_fixed theorem indexexceedancecount_le_inv_sq_add_fixed (mean : ℕ → ℝ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : indexexceedancecount mean δ gap n ≤ 1/gap^2 + fixedlogexceedancecount mean δ gap n theorem compiled","shard":"modules/01b1e451bc8da9d8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.centeredIndex","label":"centeredIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.centeredIndex","description":"def centeredIndex (X : ℕ → Ω → ℝ) (δ : ℝ) (s : ℕ) (ω : Ω) : ℝ","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-36fe452c7393","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1445,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def centeredIndex (X : ℕ → Ω → ℝ) (δ : ℝ) (s : ℕ) (ω : Ω) : ℝ","missing":[],"search":"centeredindex banditrlproof.moss.centeredindex def centeredindex (x : ℕ → ω → ℝ) (δ : ℝ) (s : ℕ) (ω : ω) : ℝ definition compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.optimismDeficit","label":"optimismDeficit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.optimismDeficit","description":"Positive part of the worst centered index through sample count n.","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-d072dc66b9f5","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1446,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def optimismDeficit (X : ℕ → Ω → ℝ) (δ : ℝ) : ℕ → Ω → ℝ | 0, _ => 0 | n+1, ω => max (optimismDeficit X δ n ω) (-centeredIndex X δ (n+1) ω) theorem optimismDeficit_nonneg (X : ℕ → Ω → ℝ) (δ : ℝ) (n : ℕ) (ω : Ω) : 0 ≤ optimismDeficit X δ n ω","missing":[],"search":"optimismdeficit banditrlproof.moss.optimismdeficit positive part of the worst centered index through sample count n. definition compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.optimismDeficit_nonneg","label":"optimismDeficit_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.optimismDeficit_nonneg","description":"theorem optimismDeficit_nonneg (X : ℕ → Ω → ℝ) (δ : ℝ) (n : ℕ) (ω : Ω) : 0 ≤ optimismDeficit X δ n ω","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-c12a960d28ae","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1447,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem optimismDeficit_nonneg (X : ℕ → Ω → ℝ) (δ : ℝ) (n : ℕ) (ω : Ω) : 0 ≤ optimismDeficit X δ n ω","missing":[],"search":"optimismdeficit_nonneg banditrlproof.moss.optimismdeficit_nonneg theorem optimismdeficit_nonneg (x : ℕ → ω → ℝ) (δ : ℝ) (n : ℕ) (ω : ω) : 0 ≤ optimismdeficit x δ n ω theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.le_optimismDeficit_iff","label":"le_optimismDeficit_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.le_optimismDeficit_iff","description":"theorem le_optimismDeficit_iff (X : ℕ → Ω → ℝ) (δ gap : ℝ) (hg : 0 < gap) (n : ℕ) (ω : Ω) : gap ≤ optimismDeficit X δ n ω ↔ ∃ s : ℕ, 0 < s ∧ s ≤ n ∧ centeredIndex X δ s ω + gap ≤ 0","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-c9e7963ffc34","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1448,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem le_optimismDeficit_iff (X : ℕ → Ω → ℝ) (δ gap : ℝ) (hg : 0 < gap) (n : ℕ) (ω : Ω) : gap ≤ optimismDeficit X δ n ω ↔ ∃ s : ℕ, 0 < s ∧ s ≤ n ∧ centeredIndex X δ s ω + gap ≤ 0","missing":[],"search":"le_optimismdeficit_iff banditrlproof.moss.le_optimismdeficit_iff theorem le_optimismdeficit_iff (x : ℕ → ω → ℝ) (δ gap : ℝ) (hg : 0 < gap) (n : ℕ) (ω : ω) : gap ≤ optimismdeficit x δ n ω ↔ ∃ s : ℕ, 0 < s ∧ s ≤ n ∧ centeredindex x δ s ω + gap ≤ 0 theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.stronglyMeasurable_centeredIndex","label":"stronglyMeasurable_centeredIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.stronglyMeasurable_centeredIndex","description":"theorem stronglyMeasurable_centeredIndex (X : ℕ → Ω → ℝ) (hX : ∀ i, StronglyMeasurable (X i)) (δ : ℝ) (s : ℕ) : StronglyMeasurable (centeredIndex X δ s)","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-61cd1eb3b794","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1449,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stronglyMeasurable_centeredIndex (X : ℕ → Ω → ℝ) (hX : ∀ i, StronglyMeasurable (X i)) (δ : ℝ) (s : ℕ) : StronglyMeasurable (centeredIndex X δ s)","missing":[],"search":"stronglymeasurable_centeredindex banditrlproof.moss.stronglymeasurable_centeredindex theorem stronglymeasurable_centeredindex (x : ℕ → ω → ℝ) (hx : ∀ i, stronglymeasurable (x i)) (δ : ℝ) (s : ℕ) : stronglymeasurable (centeredindex x δ s) theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integrable_centeredIndex","label":"integrable_centeredIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integrable_centeredIndex","description":"theorem integrable_centeredIndex [IsFiniteMeasure μ] (X : ℕ → Ω → ℝ) (hX : ∀ i, Integrable (X i) μ) (δ : ℝ) (s : ℕ) : Integrable (centeredIndex X δ s) μ","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-7f6db6198231","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1450,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_centeredIndex [IsFiniteMeasure μ] (X : ℕ → Ω → ℝ) (hX : ∀ i, Integrable (X i) μ) (δ : ℝ) (s : ℕ) : Integrable (centeredIndex X δ s) μ","missing":[],"search":"integrable_centeredindex banditrlproof.moss.integrable_centeredindex theorem integrable_centeredindex [isfinitemeasure μ] (x : ℕ → ω → ℝ) (hx : ∀ i, integrable (x i) μ) (δ : ℝ) (s : ℕ) : integrable (centeredindex x δ s) μ theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.stronglyMeasurable_optimismDeficit","label":"stronglyMeasurable_optimismDeficit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.stronglyMeasurable_optimismDeficit","description":"theorem stronglyMeasurable_optimismDeficit (X : ℕ → Ω → ℝ) (hX : ∀ i, StronglyMeasurable (X i)) (δ : ℝ) (n : ℕ) : StronglyMeasurable (optimismDeficit X δ n)","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-59ce3eb80345","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1451,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stronglyMeasurable_optimismDeficit (X : ℕ → Ω → ℝ) (hX : ∀ i, StronglyMeasurable (X i)) (δ : ℝ) (n : ℕ) : StronglyMeasurable (optimismDeficit X δ n)","missing":[],"search":"stronglymeasurable_optimismdeficit banditrlproof.moss.stronglymeasurable_optimismdeficit theorem stronglymeasurable_optimismdeficit (x : ℕ → ω → ℝ) (hx : ∀ i, stronglymeasurable (x i)) (δ : ℝ) (n : ℕ) : stronglymeasurable (optimismdeficit x δ n) theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integrable_optimismDeficit","label":"integrable_optimismDeficit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integrable_optimismDeficit","description":"theorem integrable_optimismDeficit [IsFiniteMeasure μ] (X : ℕ → Ω → ℝ) (hX : ∀ i, Integrable (X i) μ) (δ : ℝ) (n : ℕ) : Integrable (optimismDeficit X δ n) μ","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-407f7b6ab5f4","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1452,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_optimismDeficit [IsFiniteMeasure μ] (X : ℕ → Ω → ℝ) (hX : ∀ i, Integrable (X i) μ) (δ : ℝ) (n : ℕ) : Integrable (optimismDeficit X δ n) μ","missing":[],"search":"integrable_optimismdeficit banditrlproof.moss.integrable_optimismdeficit theorem integrable_optimismdeficit [isfinitemeasure μ] (x : ℕ → ω → ℝ) (hx : ∀ i, integrable (x i) μ) (δ : ℝ) (n : ℕ) : integrable (optimismdeficit x δ n) μ theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measure_optimismDeficit_ge_le","label":"measure_optimismDeficit_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measure_optimismDeficit_ge_le","description":"theorem measure_optimismDeficit_ge_le [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : μ {ω | gap ≤ optimismDeficit X δ n ω} ≤ ENNReal.ofReal (15*δ/gap^2)","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-a3008c9c6aa1","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1453,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_optimismDeficit_ge_le [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : μ {ω | gap ≤ optimismDeficit X δ n ω} ≤ ENNReal.ofReal (15*δ/gap^2)","missing":[],"search":"measure_optimismdeficit_ge_le banditrlproof.moss.measure_optimismdeficit_ge_le theorem measure_optimismdeficit_ge_le [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (n : ℕ) : μ {ω | gap ≤ optimismdeficit x δ n ω} ≤ ennreal.ofreal (15*δ/gap^2) theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_optimismDeficit_eq_integral_tail","label":"integral_optimismDeficit_eq_integral_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_optimismDeficit_eq_integral_tail","description":"Layer-cake identity with integrability derived from the coordinate MGF contracts.","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-afc153f7aa1b","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1454,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_optimismDeficit_eq_integral_tail [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ : ℝ) (n : ℕ) : ∫ ω, optimismDeficit X δ n ω ∂μ = ∫ gap in Set.Ioi 0, μ.real {ω | gap ≤ optimismDeficit X δ n ω}","missing":[],"search":"integral_optimismdeficit_eq_integral_tail banditrlproof.moss.integral_optimismdeficit_eq_integral_tail layer-cake identity with integrability derived from the coordinate mgf contracts. theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_optimismDeficit_le_two_sqrt","label":"integral_optimismDeficit_le_two_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_optimismDeficit_le_two_sqrt","description":"Source expected optimism-deficit bound, derived from the uniform tail.","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-79ace4fb7faf","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1455,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_optimismDeficit_le_two_sqrt [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ : ℝ) (hδ : 0 < δ) (n : ℕ) : ∫ ω, optimismDeficit X δ n ω ∂μ ≤ 2*sqrt (15*δ)","missing":[],"search":"integral_optimismdeficit_le_two_sqrt banditrlproof.moss.integral_optimismdeficit_le_two_sqrt source expected optimism-deficit bound, derived from the uniform tail. theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.twice_horizon_mul_integral_optimismDeficit_le","label":"twice_horizon_mul_integral_optimismDeficit_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.twice_horizon_mul_integral_optimismDeficit_le","description":"The printed 16*sqrt(n*k) optimism contribution in Theorem 9.1.","url":"../modules/banditrlproof-algorithms-mossoptimism/index.html#decl-8aa07860bf9c","parent":"module:BanditRLProof.Algorithms.MOSSOptimism","order":1456,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSOptimism"],["Source","BanditRLProof/Algorithms/MOSSOptimism.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem twice_horizon_mul_integral_optimismDeficit_le [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (n k : ℕ) (hn : 0 < n) (hk : 0 < k) : 2*(n : ℝ)*(∫ ω, optimismDeficit X ((k : ℝ)/n) n ω ∂μ) ≤ 16*sqrt ((n : ℝ)*k)","missing":[],"search":"twice_horizon_mul_integral_optimismdeficit_le banditrlproof.moss.twice_horizon_mul_integral_optimismdeficit_le the printed 16*sqrt(n*k) optimism contribution in theorem 9.1. theorem compiled","shard":"modules/dd62fe8c9d9c274e.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.logPlus_mono","label":"logPlus_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.logPlus_mono","description":"theorem logPlus_mono {x y : ℝ} (hxy : x ≤ y) : logPlus x ≤ logPlus y","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-7208b16bc122","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1457,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem logPlus_mono {x y : ℝ} (hxy : x ≤ y) : logPlus x ≤ logPlus y","missing":[],"search":"logplus_mono banditrlproof.moss.logplus_mono theorem logplus_mono {x y : ℝ} (hxy : x ≤ y) : logplus x ≤ logplus y theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.exp_neg_logPlus_inv_le","label":"exp_neg_logPlus_inv_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.exp_neg_logPlus_inv_le","description":"theorem exp_neg_logPlus_inv_le (x : ℝ) (hx : 0 < x) : exp (-logPlus (1 / x)) ≤ x","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-e242fc9b7c24","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1458,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_logPlus_inv_le (x : ℝ) (hx : 0 < x) : exp (-logPlus (1 / x)) ≤ x","missing":[],"search":"exp_neg_logplus_inv_le banditrlproof.moss.exp_neg_logplus_inv_le theorem exp_neg_logplus_inv_le (x : ℝ) (hx : 0 < x) : exp (-logplus (1 / x)) ≤ x theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.peelingBarrier","label":"peelingBarrier","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.peelingBarrier","description":"Scaled source confidence barrier, avoiding division by the sample count.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-225ea1c2d440","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1459,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def peelingBarrier (δ gap s : ℝ) : ℝ","missing":[],"search":"peelingbarrier banditrlproof.moss.peelingbarrier scaled source confidence barrier, avoiding division by the sample count. definition compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.blockBarrier","label":"blockBarrier","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.blockBarrier","description":"Lower barrier common to one dyadic block m<=s<=2m.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-9fe77553dd7c","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1460,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def blockBarrier (δ gap m : ℝ) : ℝ","missing":[],"search":"blockbarrier banditrlproof.moss.blockbarrier lower barrier common to one dyadic block m<=s<=2m. definition compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.blockBarrier_pos","label":"blockBarrier_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.blockBarrier_pos","description":"theorem blockBarrier_pos (δ gap m : ℝ) (hm : 0 < m) (hg : 0 < gap) : 0 < blockBarrier δ gap m","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-cd561052bc4a","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1461,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem blockBarrier_pos (δ gap m : ℝ) (hm : 0 < m) (hg : 0 < gap) : 0 < blockBarrier δ gap m","missing":[],"search":"blockbarrier_pos banditrlproof.moss.blockbarrier_pos theorem blockbarrier_pos (δ gap m : ℝ) (hm : 0 < m) (hg : 0 < gap) : 0 < blockbarrier δ gap m theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.blockBarrier_le_peelingBarrier","label":"blockBarrier_le_peelingBarrier","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.blockBarrier_le_peelingBarrier","description":"theorem blockBarrier_le_peelingBarrier (δ gap m s : ℝ) (hδ : 0 < δ) (hg : 0 ≤ gap) (hm : 0 < m) (hms : m ≤ s) (hsm : s ≤ 2*m) : blockBarrier δ gap m ≤ peelingBarrier δ gap s","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-20fe980a735c","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1462,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem blockBarrier_le_peelingBarrier (δ gap m s : ℝ) (hδ : 0 < δ) (hg : 0 ≤ gap) (hm : 0 < m) (hms : m ≤ s) (hsm : s ≤ 2*m) : blockBarrier δ gap m ≤ peelingBarrier δ gap s","missing":[],"search":"blockbarrier_le_peelingbarrier banditrlproof.moss.blockbarrier_le_peelingbarrier theorem blockbarrier_le_peelingbarrier (δ gap m s : ℝ) (hδ : 0 < δ) (hg : 0 ≤ gap) (hm : 0 < m) (hms : m ≤ s) (hsm : s ≤ 2*m) : blockbarrier δ gap m ≤ peelingbarrier δ gap s theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.exp_neg_blockBarrier_sq_le","label":"exp_neg_blockBarrier_sq_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.exp_neg_blockBarrier_sq_le","description":"Scalar exponential bound used after Doob on one dyadic block.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-b48c1fcc4bb3","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1463,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_blockBarrier_sq_le (δ gap m : ℝ) (hδ : 0 < δ) (hg : 0 ≤ gap) (hm : 0 < m) : exp (-(blockBarrier δ gap m)^2 / (4*m)) ≤ (2*m*δ) * exp (-(m*gap^2/4))","missing":[],"search":"exp_neg_blockbarrier_sq_le banditrlproof.moss.exp_neg_blockbarrier_sq_le scalar exponential bound used after doob on one dyadic block. theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.peelingSum","label":"peelingSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.peelingSum","description":"Centered partial sum in the source's one-based sample convention.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-e9ead723c859","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1464,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def peelingSum {Ω : Type*} (X : ℕ → Ω → ℝ) (s : ℕ) (ω : Ω) : ℝ","missing":[],"search":"peelingsum banditrlproof.moss.peelingsum centered partial sum in the source's one-based sample convention. definition compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.blockBadEvent","label":"blockBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.blockBadEvent","description":"Maximal partial-sum event dominating one dyadic block.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-0143a6db87f0","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1465,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def blockBadEvent {Ω : Type*} (X : ℕ → Ω → ℝ) (δ gap : ℝ) (m : ℕ) : Set Ω","missing":[],"search":"blockbadevent banditrlproof.moss.blockbadevent maximal partial-sum event dominating one dyadic block. definition compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measure_blockBadEvent_le","label":"measure_blockBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measure_blockBadEvent_le","description":"theorem measure_blockBadEvent_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (m : ℕ) (hm : 0 < m) : μ (blockBadEvent X δ gap m) ≤ ENNReal.ofReal ((2*(m : ℝ)*δ) * exp (-((m : ℝ)*gap^2/4)))","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-212a4e9a3c61","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1466,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_blockBadEvent_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (m : ℕ) (hm : 0 < m) : μ (blockBadEvent X δ gap m) ≤ ENNReal.ofReal ((2*(m : ℝ)*δ) * exp (-((m : ℝ)*gap^2/4)))","missing":[],"search":"measure_blockbadevent_le banditrlproof.moss.measure_blockbadevent_le theorem measure_blockbadevent_le (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) (m : ℕ) (hm : 0 < m) : μ (blockbadevent x δ gap m) ≤ ennreal.ofreal ((2*(m : ℝ)*δ) * exp (-((m : ℝ)*gap^2/4))) theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.scaledBadEvent","label":"scaledBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.scaledBadEvent","description":"All positive-sample source bad events, in scaled partial-sum form.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-eb70d041d8db","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1467,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def scaledBadEvent (X : ℕ → Ω → ℝ) (δ gap : ℝ) : Set Ω","missing":[],"search":"scaledbadevent banditrlproof.moss.scaledbadevent all positive-sample source bad events, in scaled partial-sum form. definition compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measure_scaledBadEvent_le_fifteen","label":"measure_scaledBadEvent_le_fifteen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measure_scaledBadEvent_le_fifteen","description":"Countable dyadic peeling with the printed constant 15.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-5dba186ad148","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1468,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_scaledBadEvent_le_fifteen (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) : μ (scaledBadEvent X δ gap) ≤ ENNReal.ofReal (15*δ/gap^2)","missing":[],"search":"measure_scaledbadevent_le_fifteen banditrlproof.moss.measure_scaledbadevent_le_fifteen countable dyadic peeling with the printed constant 15. theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.meanBadEvent","label":"meanBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.meanBadEvent","description":"The actual empirical-mean bad event appearing in source Lemma 9.3.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-cbc64644556c","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1469,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def meanBadEvent (X : ℕ → Ω → ℝ) (δ gap : ℝ) : Set Ω","missing":[],"search":"meanbadevent banditrlproof.moss.meanbadevent the actual empirical-mean bad event appearing in source lemma 9.3. definition compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.meanBadEvent_subset_scaledBadEvent","label":"meanBadEvent_subset_scaledBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.meanBadEvent_subset_scaledBadEvent","description":"theorem meanBadEvent_subset_scaledBadEvent (X : ℕ → Ω → ℝ) (δ gap : ℝ) : meanBadEvent X δ gap ⊆ scaledBadEvent X δ gap","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-445b52aa0a44","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1470,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:146"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem meanBadEvent_subset_scaledBadEvent (X : ℕ → Ω → ℝ) (δ gap : ℝ) : meanBadEvent X δ gap ⊆ scaledBadEvent X δ gap","missing":[],"search":"meanbadevent_subset_scaledbadevent banditrlproof.moss.meanbadevent_subset_scaledbadevent theorem meanbadevent_subset_scaledbadevent (x : ℕ → ω → ℝ) (δ gap : ℝ) : meanbadevent x δ gap ⊆ scaledbadevent x δ gap theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measure_meanBadEvent_le_fifteen","label":"measure_meanBadEvent_le_fifteen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measure_meanBadEvent_le_fifteen","description":"Source Lemma 9.3, with the printed constant and actual mean/radius event. The proof works for every positive delta, hence in particular delta in (0,1). The explicit centered-coordinate contracts are later instantiated by MOSS.","url":"../modules/banditrlproof-algorithms-mosspeeling/index.html#decl-b6819413cfdb","parent":"module:BanditRLProof.Algorithms.MOSSPeeling","order":1471,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSPeeling"],["Source","BanditRLProof/Algorithms/MOSSPeeling.lean:167"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measure_meanBadEvent_le_fifteen (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (δ gap : ℝ) (hδ : 0 < δ) (hg : 0 < gap) : μ (meanBadEvent X δ gap) ≤ ENNReal.ofReal (15*δ/gap^2)","missing":[],"search":"measure_meanbadevent_le_fifteen banditrlproof.moss.measure_meanbadevent_le_fifteen source lemma 9.3, with the printed constant and actual mean/radius event. the proof works for every positive delta, hence in particular delta in (0,1). the explicit centered-coordinate contracts are later instantiated by moss. theorem compiled","shard":"modules/0b6768b766971b05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamTrace_gapSum_le","label":"streamTrace_gapSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.streamTrace_gapSum_le","description":"Pathwise regret split with a deterministic large-gap filter for integration.","url":"../modules/banditrlproof-algorithms-mossregret/index.html#decl-5d6675c2d3d1","parent":"module:BanditRLProof.Algorithms.MOSSRegret","order":1472,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRegret"],["Source","BanditRLProof/Algorithms/MOSSRegret.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem streamTrace_gapSum_le {Ω : Type*} [MeasurableSpace Ω] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) : (∑ a, (mean best - mean a) * (pullCount (streamTrace hk n mean X ω) a n : ℝ)) ≤ (8*sqrt ((k : ℝ)/(n : ℝ)) + 2*optimismDeficit (X best) ((k : ℝ)/(n : ℝ)) n ω)*(n : ℝ) + ∑ a, if 8*sqrt ((k : ℝ)/(n : ℝ)) ≤ mean best - mean a then (mean best - mean a) * (1 + indexExceedanceCount (streamMean (X a) ω) ((k : ℝ)/(n : ℝ)) (mean best - mean a) n) else 0","missing":[],"search":"streamtrace_gapsum_le banditrlproof.moss.streamtrace_gapsum_le pathwise regret split with a deterministic large-gap filter for integration. theorem compiled","shard":"modules/68217195bbceb365.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamTrace_realMeanRegret_le","label":"streamTrace_realMeanRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.streamTrace_realMeanRegret_le","description":"theorem streamTrace_realMeanRegret_le {Ω : Type*} [MeasurableSpace Ω] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) : realMeanRegret mean (streamTrace hk n mean X ω) n ≤ (8*sqrt ((k : ℝ)/(n : ℝ)) + 2*optimismDeficit (X best) ((k : ℝ)/(n : ℝ)) n ω)*(n : ℝ) + ∑ a, if 8*sqrt ((k : ℝ)/(n : ℝ)) ≤ mean best - mean a then (mean…","url":"../modules/banditrlproof-algorithms-mossregret/index.html#decl-726932dadab6","parent":"module:BanditRLProof.Algorithms.MOSSRegret","order":1473,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRegret"],["Source","BanditRLProof/Algorithms/MOSSRegret.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem streamTrace_realMeanRegret_le {Ω : Type*} [MeasurableSpace Ω] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) : realMeanRegret mean (streamTrace hk n mean X ω) n ≤ (8*sqrt ((k : ℝ)/(n : ℝ)) + 2*optimismDeficit (X best) ((k : ℝ)/(n : ℝ)) n ω)*(n : ℝ) + ∑ a, if 8*sqrt ((k : ℝ)/(n : ℝ)) ≤ mean best - mean a then (mean best - mean a) * (1 + indexExceedanceCount (streamMean (X a) ω) ((k : ℝ)/(n : ℝ)) (mean best - mean a) n) else 0","missing":[],"search":"streamtrace_realmeanregret_le banditrlproof.moss.streamtrace_realmeanregret_le theorem streamtrace_realmeanregret_le {ω : type*} [measurablespace ω] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (ω : ω) (best : fin k) (hbest : ∀ a, mean a ≤ mean best) : realmeanregret mean (streamtrace hk n mean x ω) n ≤ (8*sqrt ((k : ℝ)/(n : ℝ)) + 2*optimismdeficit (x best) ((k : ℝ)/(n : ℝ)) n ω)*(n : ℝ) + ∑ a, if 8*sqrt ((k : ℝ)/(n : ℝ)) ≤ mean best - mean a then (mean best - mean a) * (1 + indexexceedancecount (streammean (x a) ω) ((k : ℝ)/(n : ℝ)) (mean best - mean a) n) else 0 theorem compiled","shard":"modules/68217195bbceb365.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integral_largeGapCountSum_le","label":"integral_largeGapCountSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integral_largeGapCountSum_le","description":"Integrated large-gap contribution, with the single initialization gap term.","url":"../modules/banditrlproof-algorithms-mossregret/index.html#decl-753309986465","parent":"module:BanditRLProof.Algorithms.MOSSRegret","order":1474,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRegret"],["Source","BanditRLProof/Algorithms/MOSSRegret.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_largeGapCountSum_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] {k : ℕ} (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (best : Fin k) (hbest : ∀ a, mean a ≤ mean best) (hXm : ∀ a i, StronglyMeasurable (X a i)) (hind : ∀ a, iIndepFun (X a) μ) (hmean : ∀ a i, ∫ ω, X a i ω ∂μ = 0) (hsubG : ∀ a i, HasSubgaussianMGF (X a i) 1 μ) (δ : ℝ) (hδ : 0 < δ) (n : ℕ) : (∫ ω, ∑ a, if 8*sqrt δ ≤ mean best - mean a then (mean best - mean a)*(1+indexExceedanceCount (streamMean (X a) ω) δ (mean best-mean a) n) else 0 ∂μ) ≤ (∑ a, (mean best-mean a)) + (k : ℝ)*(15/sqrt δ)","missing":[],"search":"integral_largegapcountsum_le banditrlproof.moss.integral_largegapcountsum_le integrated large-gap contribution, with the single initialization gap term. theorem compiled","shard":"modules/68217195bbceb365.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.conditionCoordinate","label":"conditionCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.conditionCoordinate","description":"def conditionCoordinate {k : ℕ} (t : ℕ) (c : History.FinitePairHistory (Fin k) ℝ t × Fin k) : ℕ × Fin k","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-218bab536bae","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1475,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def conditionCoordinate {k : ℕ} (t : ℕ) (c : History.FinitePairHistory (Fin k) ℝ t × Fin k) : ℕ × Fin k","missing":[],"search":"conditioncoordinate banditrlproof.moss.conditioncoordinate def conditioncoordinate {k : ℕ} (t : ℕ) (c : history.finitepairhistory (fin k) ℝ t × fin k) : ℕ × fin k definition compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.conditionBranch","label":"conditionBranch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.conditionBranch","description":"def conditionBranch {k : ℕ} (t : ℕ) (target : ℕ × Fin k)","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-fc1d188e4d75","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1476,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def conditionBranch {k : ℕ} (t : ℕ) (target : ℕ × Fin k)","missing":[],"search":"conditionbranch banditrlproof.moss.conditionbranch def conditionbranch {k : ℕ} (t : ℕ) (target : ℕ × fin k) definition compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_conditionCoordinate","label":"measurable_conditionCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_conditionCoordinate","description":"theorem measurable_conditionCoordinate {k : ℕ} (t : ℕ) : Measurable (conditionCoordinate (k := k) t)","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-055731b4543c","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1477,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_conditionCoordinate {k : ℕ} (t : ℕ) : Measurable (conditionCoordinate (k := k) t)","missing":[],"search":"measurable_conditioncoordinate banditrlproof.moss.measurable_conditioncoordinate theorem measurable_conditioncoordinate {k : ℕ} (t : ℕ) : measurable (conditioncoordinate (k := k) t) theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurableSet_conditionBranch","label":"measurableSet_conditionBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurableSet_conditionBranch","description":"theorem measurableSet_conditionBranch {k : ℕ} (t : ℕ) (target : ℕ × Fin k) : MeasurableSet (conditionBranch t target)","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-4f9d26b74f0b","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1478,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_conditionBranch {k : ℕ} (t : ℕ) (target : ℕ × Fin k) : MeasurableSet (conditionBranch t target)","missing":[],"search":"measurableset_conditionbranch banditrlproof.moss.measurableset_conditionbranch theorem measurableset_conditionbranch {k : ℕ} (t : ℕ) (target : ℕ × fin k) : measurableset (conditionbranch t target) theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_canonicalNextCoordinate","label":"measurable_canonicalNextCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_canonicalNextCoordinate","description":"theorem measurable_canonicalNextCoordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (canonicalNextCoordinate hk n mean t)","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-64e60d73a755","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1479,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalNextCoordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (canonicalNextCoordinate hk n mean t)","missing":[],"search":"measurable_canonicalnextcoordinate banditrlproof.moss.measurable_canonicalnextcoordinate theorem measurable_canonicalnextcoordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) : measurable (canonicalnextcoordinate hk n mean t) theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalReward_succ_eq_coordinate","label":"canonicalReward_succ_eq_coordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalReward_succ_eq_coordinate","description":"theorem canonicalReward_succ_eq_coordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k) : canonicalReward hk n mean table (t+1) = UCB.armStreamCoordinate (canonicalNextCoordinate hk n mean t table) table","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-e3a7337dbfc9","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1480,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalReward_succ_eq_coordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k) : canonicalReward hk n mean table (t+1) = UCB.armStreamCoordinate (canonicalNextCoordinate hk n mean t table) table","missing":[],"search":"canonicalreward_succ_eq_coordinate banditrlproof.moss.canonicalreward_succ_eq_coordinate theorem canonicalreward_succ_eq_coordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (table : ucb.armrewardstream k) : canonicalreward hk n mean table (t+1) = ucb.armstreamcoordinate (canonicalnextcoordinate hk n mean t table) table theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.rebuilt_mem_conditionBranch_iff","label":"rebuilt_mem_conditionBranch_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.rebuilt_mem_conditionBranch_iff","description":"theorem rebuilt_mem_conditionBranch_iff {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (table : UCB.ArmRewardStream k) : canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table) ∈ conditionBranch t target ↔ canonicalNextCoordinate hk n mean t table = target","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-2b008a24f8d0","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1481,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rebuilt_mem_conditionBranch_iff {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (table : UCB.ArmRewardStream k) : canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table) ∈ conditionBranch t target ↔ canonicalNextCoordinate hk n mean t table = target","missing":[],"search":"rebuilt_mem_conditionbranch_iff banditrlproof.moss.rebuilt_mem_conditionbranch_iff theorem rebuilt_mem_conditionbranch_iff {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (v : ℝ) (table : ucb.armrewardstream k) : canonicalconditionwithout hk n mean t target v (ucb.armstreamwithoutcoordinate target table) ∈ conditionbranch t target ↔ canonicalnextcoordinate hk n mean t table = target theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.map_rebuilt_restrict_conditionBranch","label":"map_rebuilt_restrict_conditionBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.map_rebuilt_restrict_conditionBranch","description":"theorem map_rebuilt_restrict_conditionBranch {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (μ : Measure (UCB.ArmRewardStream k)) : (Measure.map (fun table => canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table)) μ).restrict (conditionBranch t target) = (Measure.map (canonicalCondition hk n mean t) μ).restrict (conditionBranch t target)","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-3ec26008acda","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1482,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem map_rebuilt_restrict_conditionBranch {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (μ : Measure (UCB.ArmRewardStream k)) : (Measure.map (fun table => canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table)) μ).restrict (conditionBranch t target) = (Measure.map (canonicalCondition hk n mean t) μ).restrict (conditionBranch t target)","missing":[],"search":"map_rebuilt_restrict_conditionbranch banditrlproof.moss.map_rebuilt_restrict_conditionbranch theorem map_rebuilt_restrict_conditionbranch {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (v : ℝ) (μ : measure (ucb.armrewardstream k)) : (measure.map (fun table => canonicalconditionwithout hk n mean t target v (ucb.armstreamwithoutcoordinate target table)) μ).restrict (conditionbranch t target) = (measure.map (canonicalcondition hk n mean t) μ).restrict (conditionbranch t target) theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.map_condition_reward_restrict_branch","label":"map_condition_reward_restrict_branch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.map_condition_reward_restrict_branch","description":"theorem map_condition_reward_restrict_branch {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (target : ℕ × Fin k) (v : ℝ) : Measure.map (fun table => (canonicalCondition hk n mean t table, canonicalReward hk n mean table (t+1))) ((UCB.armStreamMeasure ν).restrict {table | canonicalNextCoordinate hk n mean t table = target}) = ((Measure.map (canonicalCondition hk n me…","url":"../modules/banditrlproof-algorithms-mossrewardbranch/index.html#decl-3c1fef0d48eb","parent":"module:BanditRLProof.Algorithms.MOSSRewardBranch","order":1483,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSRewardBranch"],["Source","BanditRLProof/Algorithms/MOSSRewardBranch.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem map_condition_reward_restrict_branch {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (target : ℕ × Fin k) (v : ℝ) : Measure.map (fun table => (canonicalCondition hk n mean t table, canonicalReward hk n mean table (t+1))) ((UCB.armStreamMeasure ν).restrict {table | canonicalNextCoordinate hk n mean t table = target}) = ((Measure.map (canonicalCondition hk n mean t) (UCB.armStreamMeasure ν)).restrict (conditionBranch t target)).prod (ν target.2)","missing":[],"search":"map_condition_reward_restrict_branch banditrlproof.moss.map_condition_reward_restrict_branch theorem map_condition_reward_restrict_branch {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (target : ℕ × fin k) (v : ℝ) : measure.map (fun table => (canonicalcondition hk n mean t table, canonicalreward hk n mean table (t+1))) ((ucb.armstreammeasure ν).restrict {table | canonicalnextcoordinate hk n mean t table = target}) = ((measure.map (canonicalcondition hk n mean t) (ucb.armstreammeasure ν)).restrict (conditionbranch t target)).prod (ν target.2) theorem compiled","shard":"modules/14a9b50ee60dc408.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.neg_optimismDeficit_le_centeredIndex","label":"neg_optimismDeficit_le_centeredIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.neg_optimismDeficit_le_centeredIndex","description":"theorem neg_optimismDeficit_le_centeredIndex {Ω : Type*} [MeasurableSpace Ω] (X : ℕ → Ω → ℝ) (δ : ℝ) (n s : ℕ) (ω : Ω) (hs : 0 < s) (hsn : s ≤ n) : -optimismDeficit X δ n ω ≤ centeredIndex X δ s ω","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-ab371830055d","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1484,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem neg_optimismDeficit_le_centeredIndex {Ω : Type*} [MeasurableSpace Ω] (X : ℕ → Ω → ℝ) (δ : ℝ) (n s : ℕ) (ω : Ω) (hs : 0 < s) (hsn : s ≤ n) : -optimismDeficit X δ n ω ≤ centeredIndex X δ s ω","missing":[],"search":"neg_optimismdeficit_le_centeredindex banditrlproof.moss.neg_optimismdeficit_le_centeredindex theorem neg_optimismdeficit_le_centeredindex {ω : type*} [measurablespace ω] (x : ℕ → ω → ℝ) (δ : ℝ) (n s : ℕ) (ω : ω) (hs : 0 < s) (hsn : s ≤ n) : -optimismdeficit x δ n ω ≤ centeredindex x δ s ω theorem compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.radius_eq_streamRadius","label":"radius_eq_streamRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.radius_eq_streamRadius","description":"theorem radius_eq_streamRadius (n k s : ℕ) (hn : 0 < n) (hk : 0 < k) : radius n k s = sqrt (4/(s : ℝ)*logPlus (1/((s : ℝ)*((k : ℝ)/(n : ℝ)))))","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-870d1c97210f","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1485,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem radius_eq_streamRadius (n k s : ℕ) (hn : 0 < n) (hk : 0 < k) : radius n k s = sqrt (4/(s : ℝ)*logPlus (1/((s : ℝ)*((k : ℝ)/(n : ℝ)))))","missing":[],"search":"radius_eq_streamradius banditrlproof.moss.radius_eq_streamradius theorem radius_eq_streamradius (n k s : ℕ) (hn : 0 < n) (hk : 0 < k) : radius n k s = sqrt (4/(s : ℝ)*logplus (1/((s : ℝ)*((k : ℝ)/(n : ℝ))))) theorem compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamEmpirical","label":"streamEmpirical","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.streamEmpirical","description":"Empirical state indexed by actual arm pulls, using centered reward streams.","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-3c02aa2dcaef","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1486,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def streamEmpirical {Ω : Type*} {k : ℕ} (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (trace : ActionTrace (Fin k)) (t : ℕ) (a : Fin k) : ℝ","missing":[],"search":"streamempirical banditrlproof.moss.streamempirical empirical state indexed by actual arm pulls, using centered reward streams. definition compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.pullCount_le_of_stream_policy","label":"pullCount_le_of_stream_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.pullCount_le_of_stream_policy","description":"Pathwise large-gap count bound from the MOSS policy equation itself. The law identifying centered streams with observed rewards is not assumed here.","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-f0f92ef66639","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1487,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_of_stream_policy {Ω : Type*} [MeasurableSpace Ω] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (trace : ActionTrace (Fin k)) (hpolicy : ∀ t < n, trace t = action hk n t (streamEmpirical mean X ω trace t) (fun a => pullCount trace a t)) (best chosen : Fin k) (hgap : 2 * optimismDeficit (X best) ((k : ℝ)/(n : ℝ)) n ω < mean best - mean chosen) : (pullCount trace chosen n : ℝ) ≤ 1 + indexExceedanceCount (streamMean (X chosen) ω) ((k : ℝ)/(n : ℝ)) (mean best - mean chosen) n","missing":[],"search":"pullcount_le_of_stream_policy banditrlproof.moss.pullcount_le_of_stream_policy pathwise large-gap count bound from the moss policy equation itself. the law identifying centered streams with observed rewards is not assumed here. theorem compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamCounts","label":"streamCounts","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.streamCounts","description":"Concrete MOSS execution on a fixed reward table, tracking pre-pull counts.","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-75aae15906db","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1488,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:100"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def streamCounts {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) : ℕ → Fin k → ℕ | 0 => fun _ => 0 | t+1 => fun a => let counts := streamCounts hk n mean X ω t counts a + if action hk n t (fun b => mean b + streamMean (X b) ω (counts b)) counts = a then 1 else 0 def streamTrace {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) : ActionTrace (Fin k)","missing":[],"search":"streamcounts banditrlproof.moss.streamcounts concrete moss execution on a fixed reward table, tracking pre-pull counts. definition compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamTrace","label":"streamTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.streamTrace","description":"def streamTrace {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) : ActionTrace (Fin k)","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-5a9630a73b4c","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1489,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def streamTrace {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) : ActionTrace (Fin k)","missing":[],"search":"streamtrace banditrlproof.moss.streamtrace def streamtrace {ω : type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (ω : ω) : actiontrace (fin k) definition compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.pullCount_streamTrace","label":"pullCount_streamTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.pullCount_streamTrace","description":"theorem pullCount_streamTrace {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (t : ℕ) (a : Fin k) : pullCount (streamTrace hk n mean X ω) a t = streamCounts hk n mean X ω t a","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-0bfd829aaea0","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1490,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:114"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_streamTrace {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (t : ℕ) (a : Fin k) : pullCount (streamTrace hk n mean X ω) a t = streamCounts hk n mean X ω t a","missing":[],"search":"pullcount_streamtrace banditrlproof.moss.pullcount_streamtrace theorem pullcount_streamtrace {ω : type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (ω : ω) (t : ℕ) (a : fin k) : pullcount (streamtrace hk n mean x ω) a t = streamcounts hk n mean x ω t a theorem compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamTrace_policy","label":"streamTrace_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.streamTrace_policy","description":"theorem streamTrace_policy {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (t : ℕ) : streamTrace hk n mean X ω t = action hk n t (streamEmpirical mean X ω (streamTrace hk n mean X ω) t) (fun a => pullCount (streamTrace hk n mean X ω) a t)","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-1ee9a7621581","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1491,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem streamTrace_policy {Ω : Type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (t : ℕ) : streamTrace hk n mean X ω t = action hk n t (streamEmpirical mean X ω (streamTrace hk n mean X ω) t) (fun a => pullCount (streamTrace hk n mean X ω) a t)","missing":[],"search":"streamtrace_policy banditrlproof.moss.streamtrace_policy theorem streamtrace_policy {ω : type*} {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (ω : ω) (t : ℕ) : streamtrace hk n mean x ω t = action hk n t (streamempirical mean x ω (streamtrace hk n mean x ω) t) (fun a => pullcount (streamtrace hk n mean x ω) a t) theorem compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.streamTrace_pullCount_le","label":"streamTrace_pullCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.streamTrace_pullCount_le","description":"Concrete generated-trace bound: no selected-event or policy-equation oracle.","url":"../modules/banditrlproof-algorithms-mossstream/index.html#decl-d32403620fee","parent":"module:BanditRLProof.Algorithms.MOSSStream","order":1492,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStream"],["Source","BanditRLProof/Algorithms/MOSSStream.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem streamTrace_pullCount_le {Ω : Type*} [MeasurableSpace Ω] {k : ℕ} (hk : 0 < k) (n : ℕ) (hkn : k ≤ n) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (ω : Ω) (best chosen : Fin k) (hgap : 2 * optimismDeficit (X best) ((k : ℝ)/(n : ℝ)) n ω < mean best - mean chosen) : (pullCount (streamTrace hk n mean X ω) chosen n : ℝ) ≤ 1 + indexExceedanceCount (streamMean (X chosen) ω) ((k : ℝ)/(n : ℝ)) (mean best - mean chosen) n","missing":[],"search":"streamtrace_pullcount_le banditrlproof.moss.streamtrace_pullcount_le concrete generated-trace bound: no selected-event or policy-equation oracle. theorem compiled","shard":"modules/1a1e43c6a99ac364.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_streamMean_at_count","label":"measurable_streamMean_at_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_streamMean_at_count","description":"theorem measurable_streamMean_at_count (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (c : Ω → ℕ) (hc : Measurable c) : Measurable (fun ω => streamMean X ω (c ω))","url":"../modules/banditrlproof-algorithms-mossstreammeasurable/index.html#decl-000705329051","parent":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","order":1493,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStreamMeasurable"],["Source","BanditRLProof/Algorithms/MOSSStreamMeasurable.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_streamMean_at_count (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (c : Ω → ℕ) (hc : Measurable c) : Measurable (fun ω => streamMean X ω (c ω))","missing":[],"search":"measurable_streammean_at_count banditrlproof.moss.measurable_streammean_at_count theorem measurable_streammean_at_count (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (c : ω → ℕ) (hc : measurable c) : measurable (fun ω => streammean x ω (c ω)) theorem compiled","shard":"modules/49bb04ef7e29c487.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_action_of_state","label":"measurable_action_of_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_action_of_state","description":"theorem measurable_action_of_state {k : ℕ} (hk : 0 < k) (n t : ℕ) (emp : Ω → Fin k → ℝ) (counts : Ω → Fin k → ℕ) (he : ∀ a, Measurable (fun ω => emp ω a)) (hc : ∀ a, Measurable (fun ω => counts ω a)) : Measurable (fun ω => action hk n t (emp ω) (counts ω))","url":"../modules/banditrlproof-algorithms-mossstreammeasurable/index.html#decl-5dc0541a84eb","parent":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","order":1494,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStreamMeasurable"],["Source","BanditRLProof/Algorithms/MOSSStreamMeasurable.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_action_of_state {k : ℕ} (hk : 0 < k) (n t : ℕ) (emp : Ω → Fin k → ℝ) (counts : Ω → Fin k → ℕ) (he : ∀ a, Measurable (fun ω => emp ω a)) (hc : ∀ a, Measurable (fun ω => counts ω a)) : Measurable (fun ω => action hk n t (emp ω) (counts ω))","missing":[],"search":"measurable_action_of_state banditrlproof.moss.measurable_action_of_state theorem measurable_action_of_state {k : ℕ} (hk : 0 < k) (n t : ℕ) (emp : ω → fin k → ℝ) (counts : ω → fin k → ℕ) (he : ∀ a, measurable (fun ω => emp ω a)) (hc : ∀ a, measurable (fun ω => counts ω a)) : measurable (fun ω => action hk n t (emp ω) (counts ω)) theorem compiled","shard":"modules/49bb04ef7e29c487.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_streamCounts","label":"measurable_streamCounts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_streamCounts","description":"theorem measurable_streamCounts {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (hXm : ∀ a i, StronglyMeasurable (X a i)) (t : ℕ) (a : Fin k) : Measurable (fun ω => streamCounts hk n mean X ω t a)","url":"../modules/banditrlproof-algorithms-mossstreammeasurable/index.html#decl-ed3df4bea5a1","parent":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","order":1495,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStreamMeasurable"],["Source","BanditRLProof/Algorithms/MOSSStreamMeasurable.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_streamCounts {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (hXm : ∀ a i, StronglyMeasurable (X a i)) (t : ℕ) (a : Fin k) : Measurable (fun ω => streamCounts hk n mean X ω t a)","missing":[],"search":"measurable_streamcounts banditrlproof.moss.measurable_streamcounts theorem measurable_streamcounts {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (hxm : ∀ a i, stronglymeasurable (x a i)) (t : ℕ) (a : fin k) : measurable (fun ω => streamcounts hk n mean x ω t a) theorem compiled","shard":"modules/49bb04ef7e29c487.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_streamTrace","label":"measurable_streamTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_streamTrace","description":"theorem measurable_streamTrace {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (hXm : ∀ a i, StronglyMeasurable (X a i)) (t : ℕ) : Measurable (fun ω => streamTrace hk n mean X ω t)","url":"../modules/banditrlproof-algorithms-mossstreammeasurable/index.html#decl-6255978a2067","parent":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","order":1496,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStreamMeasurable"],["Source","BanditRLProof/Algorithms/MOSSStreamMeasurable.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_streamTrace {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (hXm : ∀ a i, StronglyMeasurable (X a i)) (t : ℕ) : Measurable (fun ω => streamTrace hk n mean X ω t)","missing":[],"search":"measurable_streamtrace banditrlproof.moss.measurable_streamtrace theorem measurable_streamtrace {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (hxm : ∀ a i, stronglymeasurable (x a i)) (t : ℕ) : measurable (fun ω => streamtrace hk n mean x ω t) theorem compiled","shard":"modules/49bb04ef7e29c487.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.integrable_streamTrace_regret","label":"integrable_streamTrace_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.integrable_streamTrace_regret","description":"theorem integrable_streamTrace_regret (μ : Measure Ω) [IsFiniteMeasure μ] {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (hXm : ∀ a i, StronglyMeasurable (X a i)) : Integrable (fun ω => realMeanRegret mean (streamTrace hk n mean X ω) n) μ","url":"../modules/banditrlproof-algorithms-mossstreammeasurable/index.html#decl-935086047855","parent":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","order":1497,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSStreamMeasurable"],["Source","BanditRLProof/Algorithms/MOSSStreamMeasurable.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_streamTrace_regret (μ : Measure Ω) [IsFiniteMeasure μ] {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (X : Fin k → ℕ → Ω → ℝ) (hXm : ∀ a i, StronglyMeasurable (X a i)) : Integrable (fun ω => realMeanRegret mean (streamTrace hk n mean X ω) n) μ","missing":[],"search":"integrable_streamtrace_regret banditrlproof.moss.integrable_streamtrace_regret theorem integrable_streamtrace_regret (μ : measure ω) [isfinitemeasure μ] {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (x : fin k → ℕ → ω → ℝ) (hxm : ∀ a i, stronglymeasurable (x a i)) : integrable (fun ω => realmeanregret mean (streamtrace hk n mean x ω) n) μ theorem compiled","shard":"modules/49bb04ef7e29c487.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalCondition","label":"canonicalCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalCondition","description":"def canonicalCondition {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k)","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-74a52ad424f1","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1498,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def canonicalCondition {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k)","missing":[],"search":"canonicalcondition banditrlproof.moss.canonicalcondition def canonicalcondition {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (table : ucb.armrewardstream k) definition compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalNextCoordinate","label":"canonicalNextCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalNextCoordinate","description":"def canonicalNextCoordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k) : ℕ × Fin k","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-a2f42a42c9b0","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1499,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def canonicalNextCoordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k) : ℕ × Fin k","missing":[],"search":"canonicalnextcoordinate banditrlproof.moss.canonicalnextcoordinate def canonicalnextcoordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (table : ucb.armrewardstream k) : ℕ × fin k definition compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalNextCoordinate_count","label":"canonicalNextCoordinate_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalNextCoordinate_count","description":"theorem canonicalNextCoordinate_count {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k) : (canonicalNextCoordinate hk n mean t table).1 = pullCount (canonicalAction hk n mean table) (canonicalNextCoordinate hk n mean t table).2 (t+1)+1","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-5b3fa01c388f","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1500,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalNextCoordinate_count {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (table : UCB.ArmRewardStream k) : (canonicalNextCoordinate hk n mean t table).1 = pullCount (canonicalAction hk n mean table) (canonicalNextCoordinate hk n mean t table).2 (t+1)+1","missing":[],"search":"canonicalnextcoordinate_count banditrlproof.moss.canonicalnextcoordinate_count theorem canonicalnextcoordinate_count {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (table : ucb.armrewardstream k) : (canonicalnextcoordinate hk n mean t table).1 = pullcount (canonicalaction hk n mean table) (canonicalnextcoordinate hk n mean t table).2 (t+1)+1 theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalHistory_eq_of_complement_eq","label":"canonicalHistory_eq_of_complement_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalHistory_eq_of_complement_eq","description":"theorem canonicalHistory_eq_of_complement_eq {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (table table' : UCB.ArmRewardStream k) (hc : UCB.armStreamWithoutCoordinate target table = UCB.armStreamWithoutCoordinate target table') (hf : pullCount (canonicalAction hk n mean table) target.2 (t+1) < target.1) : canonicalHistory hk n mean table t = canonicalHistory hk n mean table' t","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-6d25c629ade5","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1501,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistory_eq_of_complement_eq {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (table table' : UCB.ArmRewardStream k) (hc : UCB.armStreamWithoutCoordinate target table = UCB.armStreamWithoutCoordinate target table') (hf : pullCount (canonicalAction hk n mean table) target.2 (t+1) < target.1) : canonicalHistory hk n mean table t = canonicalHistory hk n mean table' t","missing":[],"search":"canonicalhistory_eq_of_complement_eq banditrlproof.moss.canonicalhistory_eq_of_complement_eq theorem canonicalhistory_eq_of_complement_eq {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (table table' : ucb.armrewardstream k) (hc : ucb.armstreamwithoutcoordinate target table = ucb.armstreamwithoutcoordinate target table') (hf : pullcount (canonicalaction hk n mean table) target.2 (t+1) < target.1) : canonicalhistory hk n mean table t = canonicalhistory hk n mean table' t theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalNextCoordinate_eq_iff_insert","label":"canonicalNextCoordinate_eq_iff_insert","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalNextCoordinate_eq_iff_insert","description":"theorem canonicalNextCoordinate_eq_iff_insert {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (table : UCB.ArmRewardStream k) : canonicalNextCoordinate hk n mean t table = target ↔ canonicalNextCoordinate hk n mean t (UCB.armStreamInsertCoordinate target v (UCB.armStreamWithoutCoordinate target table)) = target","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-572882006458","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1502,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalNextCoordinate_eq_iff_insert {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (table : UCB.ArmRewardStream k) : canonicalNextCoordinate hk n mean t table = target ↔ canonicalNextCoordinate hk n mean t (UCB.armStreamInsertCoordinate target v (UCB.armStreamWithoutCoordinate target table)) = target","missing":[],"search":"canonicalnextcoordinate_eq_iff_insert banditrlproof.moss.canonicalnextcoordinate_eq_iff_insert theorem canonicalnextcoordinate_eq_iff_insert {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (v : ℝ) (table : ucb.armrewardstream k) : canonicalnextcoordinate hk n mean t table = target ↔ canonicalnextcoordinate hk n mean t (ucb.armstreaminsertcoordinate target v (ucb.armstreamwithoutcoordinate target table)) = target theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalConditionWithout","label":"canonicalConditionWithout","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalConditionWithout","description":"def canonicalConditionWithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ)","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-67eac1ec64d7","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1503,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def canonicalConditionWithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ)","missing":[],"search":"canonicalconditionwithout banditrlproof.moss.canonicalconditionwithout def canonicalconditionwithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (v : ℝ) definition compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.canonicalCondition_eq_without","label":"canonicalCondition_eq_without","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.canonicalCondition_eq_without","description":"theorem canonicalCondition_eq_without {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (table : UCB.ArmRewardStream k) (hnext : canonicalNextCoordinate hk n mean t table = target) : canonicalCondition hk n mean t table = canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table)","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-fe54cc860674","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1504,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalCondition_eq_without {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) (table : UCB.ArmRewardStream k) (hnext : canonicalNextCoordinate hk n mean t table = target) : canonicalCondition hk n mean t table = canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table)","missing":[],"search":"canonicalcondition_eq_without banditrlproof.moss.canonicalcondition_eq_without theorem canonicalcondition_eq_without {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (v : ℝ) (table : ucb.armrewardstream k) (hnext : canonicalnextcoordinate hk n mean t table = target) : canonicalcondition hk n mean t table = canonicalconditionwithout hk n mean t target v (ucb.armstreamwithoutcoordinate target table) theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_canonicalCondition","label":"measurable_canonicalCondition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_canonicalCondition","description":"theorem measurable_canonicalCondition {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (canonicalCondition hk n mean t)","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-8f0c8c88c67c","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1505,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalCondition {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) : Measurable (canonicalCondition hk n mean t)","missing":[],"search":"measurable_canonicalcondition banditrlproof.moss.measurable_canonicalcondition theorem measurable_canonicalcondition {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) : measurable (canonicalcondition hk n mean t) theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.measurable_canonicalConditionWithout","label":"measurable_canonicalConditionWithout","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.measurable_canonicalConditionWithout","description":"theorem measurable_canonicalConditionWithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) : Measurable (canonicalConditionWithout hk n mean t target v)","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-9b01386038e7","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1506,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalConditionWithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (target : ℕ × Fin k) (v : ℝ) : Measurable (canonicalConditionWithout hk n mean t target v)","missing":[],"search":"measurable_canonicalconditionwithout banditrlproof.moss.measurable_canonicalconditionwithout theorem measurable_canonicalconditionwithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (target : ℕ × fin k) (v : ℝ) : measurable (canonicalconditionwithout hk n mean t target v) theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.indepFun_coordinate_canonicalConditionWithout","label":"indepFun_coordinate_canonicalConditionWithout","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.indepFun_coordinate_canonicalConditionWithout","description":"A target reward is independent of the condition reconstructed without it.","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-de7acd6ba322","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1507,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:90"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem indepFun_coordinate_canonicalConditionWithout {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (target : ℕ × Fin k) (v : ℝ) : IndepFun (UCB.armStreamCoordinate target) (fun table => canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table)) (UCB.armStreamMeasure ν)","missing":[],"search":"indepfun_coordinate_canonicalconditionwithout banditrlproof.moss.indepfun_coordinate_canonicalconditionwithout a target reward is independent of the condition reconstructed without it. theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.MOSS.map_canonicalConditionWithout_coordinate","label":"map_canonicalConditionWithout_coordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MOSS.map_canonicalConditionWithout_coordinate","description":"theorem map_canonicalConditionWithout_coordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (target : ℕ × Fin k) (v : ℝ) : Measure.map (fun table => (canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table), UCB.armStreamCoordinate target table)) (UCB.armStreamMeasure ν) = (Measure.map (fun table => canonicalConditionWithout h…","url":"../modules/banditrlproof-algorithms-mossunusedcoordinate/index.html#decl-e128595638f3","parent":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","order":1508,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MOSSUnusedCoordinate"],["Source","BanditRLProof/Algorithms/MOSSUnusedCoordinate.lean:100"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem map_canonicalConditionWithout_coordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : Fin k → ℝ) (t : ℕ) (ν : Kernel (Fin k) ℝ) [IsMarkovKernel ν] (target : ℕ × Fin k) (v : ℝ) : Measure.map (fun table => (canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table), UCB.armStreamCoordinate target table)) (UCB.armStreamMeasure ν) = (Measure.map (fun table => canonicalConditionWithout hk n mean t target v (UCB.armStreamWithoutCoordinate target table)) (UCB.armStreamMeasure ν)).prod (ν target.2)","missing":[],"search":"map_canonicalconditionwithout_coordinate banditrlproof.moss.map_canonicalconditionwithout_coordinate theorem map_canonicalconditionwithout_coordinate {k : ℕ} (hk : 0 < k) (n : ℕ) (mean : fin k → ℝ) (t : ℕ) (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (target : ℕ × fin k) (v : ℝ) : measure.map (fun table => (canonicalconditionwithout hk n mean t target v (ucb.armstreamwithoutcoordinate target table), ucb.armstreamcoordinate target table)) (ucb.armstreammeasure ν) = (measure.map (fun table => canonicalconditionwithout hk n mean t target v (ucb.armstreamwithoutcoordinate target table)) (ucb.armstreammeasure ν)).prod (ν target.2) theorem compiled","shard":"modules/bb7fe370b32d2783.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FinitePMF.iid_toMeasure_pi","label":"iid_toMeasure_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_toMeasure_pi","description":"theorem iid_toMeasure_pi {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (T : ℕ) : (iid p T).toMeasure = Measure.pi (fun _ : Fin T => p.toMeasure)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-41536a3710a3","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1509,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_toMeasure_pi {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (T : ℕ) : (iid p T).toMeasure = Measure.pi (fun _ : Fin T => p.toMeasure)","missing":[],"search":"iid_tomeasure_pi banditrlproof.finitepmf.iid_tomeasure_pi theorem iid_tomeasure_pi {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (t : ℕ) : (iid p t).tomeasure = measure.pi (fun _ : fin t => p.tomeasure) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.eventIndicator","label":"eventIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FinitePMF.eventIndicator","description":"noncomputable def eventIndicator {α : Type*} (E : Set α) (a : α) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-f22ac17bde28","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1510,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def eventIndicator {α : Type*} (E : Set α) (a : α) : ℝ","missing":[],"search":"eventindicator banditrlproof.finitepmf.eventindicator noncomputable def eventindicator {α : type*} (e : set α) (a : α) : ℝ definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.eventIndicator_mean","label":"eventIndicator_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.eventIndicator_mean","description":"theorem eventIndicator_mean {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) : (∫ a, eventIndicator E a ∂p.toMeasure) = p.toMeasure.real E","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-202bab580941","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1511,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem eventIndicator_mean {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) : (∫ a, eventIndicator E a ∂p.toMeasure) = p.toMeasure.real E","missing":[],"search":"eventindicator_mean banditrlproof.finitepmf.eventindicator_mean theorem eventindicator_mean {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) : (∫ a, eventindicator e a ∂p.tomeasure) = p.tomeasure.real e theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_indicator_mean","label":"iid_indicator_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_indicator_mean","description":"theorem iid_indicator_mean {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) {T : ℕ} (t : Fin T) : (∫ x, eventIndicator E (x t) ∂(iid p T).toMeasure) = p.toMeasure.real E","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-02419e519e65","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1512,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_indicator_mean {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) {T : ℕ} (t : Fin T) : (∫ x, eventIndicator E (x t) ∂(iid p T).toMeasure) = p.toMeasure.real E","missing":[],"search":"iid_indicator_mean banditrlproof.finitepmf.iid_indicator_mean theorem iid_indicator_mean {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) {t : ℕ} (t : fin t) : (∫ x, eventindicator e (x t) ∂(iid p t).tomeasure) = p.tomeasure.real e theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_indicator_subGaussian","label":"iid_indicator_subGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_indicator_subGaussian","description":"theorem iid_indicator_subGaussian {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) {T : ℕ} (t : Fin T) : HasSubgaussianMGF (fun x => eventIndicator E (x t) - p.toMeasure.real E) (1/4 : NNReal) (iid p T).toMeasure","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-62ccc04ffc79","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1513,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_indicator_subGaussian {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) {T : ℕ} (t : Fin T) : HasSubgaussianMGF (fun x => eventIndicator E (x t) - p.toMeasure.real E) (1/4 : NNReal) (iid p T).toMeasure","missing":[],"search":"iid_indicator_subgaussian banditrlproof.finitepmf.iid_indicator_subgaussian theorem iid_indicator_subgaussian {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) {t : ℕ} (t : fin t) : hassubgaussianmgf (fun x => eventindicator e (x t) - p.tomeasure.real e) (1/4 : nnreal) (iid p t).tomeasure theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_indicator_independent","label":"iid_indicator_independent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_indicator_independent","description":"theorem iid_indicator_independent {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) (T : ℕ) : iIndepFun (fun t (x : Fin T → α) => eventIndicator E (x t) - p.toMeasure.real E) (iid p T).toMeasure","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-02cace151ddd","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1514,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_indicator_independent {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) (T : ℕ) : iIndepFun (fun t (x : Fin T → α) => eventIndicator E (x t) - p.toMeasure.real E) (iid p T).toMeasure","missing":[],"search":"iid_indicator_independent banditrlproof.finitepmf.iid_indicator_independent theorem iid_indicator_independent {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) (t : ℕ) : iindepfun (fun t (x : fin t → α) => eventindicator e (x t) - p.tomeasure.real e) (iid p t).tomeasure theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.eventCount_eq_sum","label":"eventCount_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.eventCount_eq_sum","description":"theorem eventCount_eq_sum {α : Type*} {T : ℕ} (E : Set α) (x : Fin T → α) : (eventCount E x : ℝ) = ∑ t, eventIndicator E (x t)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-8886515c7d60","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1515,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem eventCount_eq_sum {α : Type*} {T : ℕ} (E : Set α) (x : Fin T → α) : (eventCount E x : ℝ) = ∑ t, eventIndicator E (x t)","missing":[],"search":"eventcount_eq_sum banditrlproof.finitepmf.eventcount_eq_sum theorem eventcount_eq_sum {α : type*} {t : ℕ} (e : set α) (x : fin t → α) : (eventcount e x : ℝ) = ∑ t, eventindicator e (x t) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_centered_sum_subGaussian","label":"iid_centered_sum_subGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_centered_sum_subGaussian","description":"theorem iid_centered_sum_subGaussian {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) (T : ℕ) : HasSubgaussianMGF (fun x : Fin T → α => ∑ t, (eventIndicator E (x t) - p.toMeasure.real E)) ((T : NNReal)/4) (iid p T).toMeasure","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-542aa52b14f9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1516,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_centered_sum_subGaussian {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) (T : ℕ) : HasSubgaussianMGF (fun x : Fin T → α => ∑ t, (eventIndicator E (x t) - p.toMeasure.real E)) ((T : NNReal)/4) (iid p T).toMeasure","missing":[],"search":"iid_centered_sum_subgaussian banditrlproof.finitepmf.iid_centered_sum_subgaussian theorem iid_centered_sum_subgaussian {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) (t : ℕ) : hassubgaussianmgf (fun x : fin t → α => ∑ t, (eventindicator e (x t) - p.tomeasure.real e)) ((t : nnreal)/4) (iid p t).tomeasure theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_centered_sum_abs_tail","label":"iid_centered_sum_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_centered_sum_abs_tail","description":"theorem iid_centered_sum_abs_tail {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) (T : ℕ) (u : ℝ) (hu : 0 ≤ u) : (iid p T).toMeasure.real {x | u ≤ |∑ t, (eventIndicator E (x t) - p.toMeasure.real E)|} ≤ 2 * Real.exp (-u^2 / (2 * ((T : ℝ)/4)))","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-1af306b9db2d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1517,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_centered_sum_abs_tail {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) (T : ℕ) (u : ℝ) (hu : 0 ≤ u) : (iid p T).toMeasure.real {x | u ≤ |∑ t, (eventIndicator E (x t) - p.toMeasure.real E)|} ≤ 2 * Real.exp (-u^2 / (2 * ((T : ℝ)/4)))","missing":[],"search":"iid_centered_sum_abs_tail banditrlproof.finitepmf.iid_centered_sum_abs_tail theorem iid_centered_sum_abs_tail {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) (t : ℕ) (u : ℝ) (hu : 0 ≤ u) : (iid p t).tomeasure.real {x | u ≤ |∑ t, (eventindicator e (x t) - p.tomeasure.real e)|} ≤ 2 * real.exp (-u^2 / (2 * ((t : ℝ)/4))) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_frequency_tail","label":"iid_frequency_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_frequency_tail","description":"theorem iid_frequency_tail {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) {T : ℕ} (hT : 0 < T) (eta : ℝ) (heta : 0 ≤ eta) : (iid p T).toMeasure.real {x | eta ≤ |(eventCount E x : ℝ)/T - p.toMeasure.real E|} ≤ 2 * Real.exp (-2 * T * eta^2)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-e010fca85acc","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1518,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_frequency_tail {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (p : PMF α) (E : Set α) {T : ℕ} (hT : 0 < T) (eta : ℝ) (heta : 0 ≤ eta) : (iid p T).toMeasure.real {x | eta ≤ |(eventCount E x : ℝ)/T - p.toMeasure.real E|} ≤ 2 * Real.exp (-2 * T * eta^2)","missing":[],"search":"iid_frequency_tail banditrlproof.finitepmf.iid_frequency_tail theorem iid_frequency_tail {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p : pmf α) (e : set α) {t : ℕ} (ht : 0 < t) (eta : ℝ) (heta : 0 ≤ eta) : (iid p t).tomeasure.real {x | eta ≤ |(eventcount e x : ℝ)/t - p.tomeasure.real e|} ≤ 2 * real.exp (-2 * t * eta^2) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionProbReal","label":"collisionProbReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionProbReal","description":"noncomputable def collisionProbReal (n k : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-c26d4d4094b8","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1519,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def collisionProbReal (n k : ℕ) : ℝ","missing":[],"search":"collisionprobreal banditrlproof.musicalchairs.collisionprobreal noncomputable def collisionprobreal (n k : ℕ) : ℝ definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collision_indicator_mean","label":"collision_indicator_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collision_indicator_mean","description":"theorem collision_indicator_mean {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toMeasure.real {draw | ¬ CollisionFree draw i} = collisionProbReal n k","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-6dec29db0a85","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1520,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collision_indicator_mean {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toMeasure.real {draw | ¬ CollisionFree draw i} = collisionProbReal n k","missing":[],"search":"collision_indicator_mean banditrlproof.musicalchairs.collision_indicator_mean theorem collision_indicator_mean {n k : ℕ} (hk : 0 < k) (i : fin n) : (explorationdraw n k hk).tomeasure.real {draw | ¬ collisionfree draw i} = collisionprobreal n k theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCollisionRate","label":"localCollisionRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCollisionRate","description":"noncomputable def localCollisionRate {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-a31b516a8f00","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1521,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localCollisionRate {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) : ℝ","missing":[],"search":"localcollisionrate banditrlproof.musicalchairs.localcollisionrate noncomputable def localcollisionrate {n k t : ℕ} (x : fin t → fin n → fin k) (r : fin t → fin k → ℝ) (i : fin n) : ℝ definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCollisionRate_eq","label":"localCollisionRate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCollisionRate_eq","description":"theorem localCollisionRate_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) : localCollisionRate x r i = (collisionCount i x : ℝ)/T","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-6621a44e2399","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1522,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:152"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCollisionRate_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) : localCollisionRate x r i = (collisionCount i x : ℝ)/T","missing":[],"search":"localcollisionrate_eq banditrlproof.musicalchairs.localcollisionrate_eq theorem localcollisionrate_eq {n k t : ℕ} (x : fin t → fin n → fin k) (r : fin t → fin k → ℝ) (i : fin n) : localcollisionrate x r i = (collisioncount i x : ℝ)/t theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionBadEvent","label":"collisionBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionBadEvent","description":"def collisionBadEvent {n k T : ℕ} (i : Fin n) : Set (Fin T → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-293dfea9f50f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1523,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def collisionBadEvent {n k T : ℕ} (i : Fin n) : Set (Fin T → Fin n → Fin k)","missing":[],"search":"collisionbadevent banditrlproof.musicalchairs.collisionbadevent def collisionbadevent {n k t : ℕ} (i : fin n) : set (fin t → fin n → fin k) definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionBadEvent_bound","label":"collisionBadEvent_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionBadEvent_bound","description":"theorem collisionBadEvent_bound {n k T : ℕ} (hk : 0 < k) (hT : 0 < T) (i : Fin n) : (explorationLaw n k T hk).toMeasure (collisionBadEvent (k := k) (T := T) i) ≤ ENNReal.ofReal (2 * Real.exp (-(T : ℝ)/(50*(k : ℝ)^2)))","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-4faa86894132","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1524,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collisionBadEvent_bound {n k T : ℕ} (hk : 0 < k) (hT : 0 < T) (i : Fin n) : (explorationLaw n k T hk).toMeasure (collisionBadEvent (k := k) (T := T) i) ≤ ENNReal.ofReal (2 * Real.exp (-(T : ℝ)/(50*(k : ℝ)^2)))","missing":[],"search":"collisionbadevent_bound banditrlproof.musicalchairs.collisionbadevent_bound theorem collisionbadevent_bound {n k t : ℕ} (hk : 0 < k) (ht : 0 < t) (i : fin n) : (explorationlaw n k t hk).tomeasure (collisionbadevent (k := k) (t := t) i) ≤ ennreal.ofreal (2 * real.exp (-(t : ℝ)/(50*(k : ℝ)^2))) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionBadEvent","label":"allCollisionBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionBadEvent","description":"def allCollisionBadEvent {n k T : ℕ} : Set (Fin T → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-725150da9694","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1525,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:172"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def allCollisionBadEvent {n k T : ℕ} : Set (Fin T → Fin n → Fin k)","missing":[],"search":"allcollisionbadevent banditrlproof.musicalchairs.allcollisionbadevent def allcollisionbadevent {n k t : ℕ} : set (fin t → fin n → fin k) definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionBadEvent_bound","label":"allCollisionBadEvent_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionBadEvent_bound","description":"theorem allCollisionBadEvent_bound {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (hT : 0 < T) : (explorationLaw n k T hk).toMeasure (allCollisionBadEvent (n := n) (k := k) (T := T)) ≤ ENNReal.ofReal (2*(k : ℝ)*Real.exp (-(T : ℝ)/(50*(k : ℝ)^2)))","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-878a59aef2f2","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1526,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:175"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allCollisionBadEvent_bound {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (hT : 0 < T) : (explorationLaw n k T hk).toMeasure (allCollisionBadEvent (n := n) (k := k) (T := T)) ≤ ENNReal.ofReal (2*(k : ℝ)*Real.exp (-(T : ℝ)/(50*(k : ℝ)^2)))","missing":[],"search":"allcollisionbadevent_bound banditrlproof.musicalchairs.allcollisionbadevent_bound theorem allcollisionbadevent_bound {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (ht : 0 < t) : (explorationlaw n k t hk).tomeasure (allcollisionbadevent (n := n) (k := k) (t := t)) ≤ ennreal.ofreal (2*(k : ℝ)*real.exp (-(t : ℝ)/(50*(k : ℝ)^2))) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collision_exploration_threshold","label":"collision_exploration_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collision_exploration_threshold","description":"theorem collision_exploration_threshold (k T : ℕ) (hk : 0 < k) (delta : ℝ) (hdelta : 0 < delta) (hT : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 2*(k : ℝ) * Real.exp (-(T : ℝ)/(50*(k : ℝ)^2)) ≤ delta/2","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-0de4f5111e1f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1527,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:196"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collision_exploration_threshold (k T : ℕ) (hk : 0 < k) (delta : ℝ) (hdelta : 0 < delta) (hT : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 2*(k : ℝ) * Real.exp (-(T : ℝ)/(50*(k : ℝ)^2)) ≤ delta/2","missing":[],"search":"collision_exploration_threshold banditrlproof.musicalchairs.collision_exploration_threshold theorem collision_exploration_threshold (k t : ℕ) (hk : 0 < k) (delta : ℝ) (hdelta : 0 < delta) (ht : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ t) : 2*(k : ℝ) * real.exp (-(t : ℝ)/(50*(k : ℝ)^2)) ≤ delta/2 theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionBadEvent_le_half_delta","label":"allCollisionBadEvent_le_half_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionBadEvent_le_half_delta","description":"theorem allCollisionBadEvent_le_half_delta {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (hT : 0 < T) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : (explorationLaw n k T hk).toMeasure (allCollisionBadEvent (n := n) (k := k) (T := T)) ≤ ENNReal.ofReal (delta/2)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-b72ab400c81e","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1528,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:214"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allCollisionBadEvent_le_half_delta {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (hT : 0 < T) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : (explorationLaw n k T hk).toMeasure (allCollisionBadEvent (n := n) (k := k) (T := T)) ≤ ENNReal.ofReal (delta/2)","missing":[],"search":"allcollisionbadevent_le_half_delta banditrlproof.musicalchairs.allcollisionbadevent_le_half_delta theorem allcollisionbadevent_le_half_delta {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (ht : 0 < t) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ t) : (explorationlaw n k t hk).tomeasure (allcollisionbadevent (n := n) (k := k) (t := t)) ≤ ennreal.ofreal (delta/2) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate","label":"allCollisionAccurate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionAccurate","description":"def allCollisionAccurate {n k T : ℕ} : Set (Fin T → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-40c60707e1eb","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1529,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:222"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def allCollisionAccurate {n k T : ℕ} : Set (Fin T → Fin n → Fin k)","missing":[],"search":"allcollisionaccurate banditrlproof.musicalchairs.allcollisionaccurate def allcollisionaccurate {n k t : ℕ} : set (fin t → fin n → fin k) definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate_eq_compl","label":"allCollisionAccurate_eq_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionAccurate_eq_compl","description":"theorem allCollisionAccurate_eq_compl {n k T : ℕ} : allCollisionAccurate (n := n) (k := k) (T := T) = (allCollisionBadEvent)ᶜ","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-b2be3202ea6d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1530,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:225"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allCollisionAccurate_eq_compl {n k T : ℕ} : allCollisionAccurate (n := n) (k := k) (T := T) = (allCollisionBadEvent)ᶜ","missing":[],"search":"allcollisionaccurate_eq_compl banditrlproof.musicalchairs.allcollisionaccurate_eq_compl theorem allcollisionaccurate_eq_compl {n k t : ℕ} : allcollisionaccurate (n := n) (k := k) (t := t) = (allcollisionbadevent)ᶜ theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate_probability","label":"allCollisionAccurate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionAccurate_probability","description":"theorem allCollisionAccurate_probability {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (hT : 0 < T) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 1 - ENNReal.ofReal (delta/2) ≤ (explorationLaw n k T hk).toMeasure (allCollisionAccurate (n := n) (k := k) (T := T))","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-60cab3b8ef50","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1531,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:230"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allCollisionAccurate_probability {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (hT : 0 < T) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 1 - ENNReal.ofReal (delta/2) ≤ (explorationLaw n k T hk).toMeasure (allCollisionAccurate (n := n) (k := k) (T := T))","missing":[],"search":"allcollisionaccurate_probability banditrlproof.musicalchairs.allcollisionaccurate_probability theorem allcollisionaccurate_probability {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (ht : 0 < t) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ t) : 1 - ennreal.ofreal (delta/2) ≤ (explorationlaw n k t hk).tomeasure (allcollisionaccurate (n := n) (k := k) (t := t)) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationRewardLaw_fst_event","label":"explorationRewardLaw_fst_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationRewardLaw_fst_event","description":"theorem explorationRewardLaw_fst_event {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (E : Set (Fin T → Fin n → Fin k)) : explorationRewardLaw hk nu (Prod.fst ⁻¹' E) = (explorationLaw n k T hk).toMeasure E","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-6d758d7d521a","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1532,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:244"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationRewardLaw_fst_event {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (E : Set (Fin T → Fin n → Fin k)) : explorationRewardLaw hk nu (Prod.fst ⁻¹' E) = (explorationLaw n k T hk).toMeasure E","missing":[],"search":"explorationrewardlaw_fst_event banditrlproof.musicalchairs.explorationrewardlaw_fst_event theorem explorationrewardlaw_fst_event {n k t : ℕ} (hk : 0 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (e : set (fin t → fin n → fin k)) : explorationrewardlaw hk nu (prod.fst ⁻¹' e) = (explorationlaw n k t hk).tomeasure e theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.statisticsBadEvent","label":"statisticsBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.statisticsBadEvent","description":"def statisticsBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-09670164c5a2","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1533,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:252"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def statisticsBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"statisticsbadevent banditrlproof.musicalchairs.statisticsbadevent def statisticsbadevent {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.statisticsBadEvent_le_delta","label":"statisticsBadEvent_le_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.statisticsBadEvent_le_delta","description":"theorem statisticsBadEvent_le_delta {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : 0 < T) (hmeans : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) (hcoll : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : explorationRewardLaw hk nu (…","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-b08ded5e5456","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1534,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:256"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem statisticsBadEvent_le_delta {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : 0 < T) (hmeans : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) (hcoll : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : explorationRewardLaw hk nu (statisticsBadEvent (n := n) (T := T) nu eps) ≤ ENNReal.ofReal delta","missing":[],"search":"statisticsbadevent_le_delta banditrlproof.musicalchairs.statisticsbadevent_le_delta theorem statisticsbadevent_le_delta {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (ht : 0 < t) (hmeans : (16*(k : ℝ)/eps^2) * real.log (4*(k : ℝ)^2/delta) ≤ t) (hcoll : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ t) : explorationrewardlaw hk nu (statisticsbadevent (n := n) (t := t) nu eps) ≤ ennreal.ofreal delta theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate","label":"explorationStatisticsAccurate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationStatisticsAccurate","description":"def explorationStatisticsAccurate {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-b4396220a5d2","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1535,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:276"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def explorationStatisticsAccurate {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"explorationstatisticsaccurate banditrlproof.musicalchairs.explorationstatisticsaccurate def explorationstatisticsaccurate {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_eq_compl","label":"explorationStatisticsAccurate_eq_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationStatisticsAccurate_eq_compl","description":"theorem explorationStatisticsAccurate_eq_compl {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : explorationStatisticsAccurate (n := n) (T := T) nu eps = (statisticsBadEvent nu eps)ᶜ","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-09335df3d8d9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1536,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:282"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationStatisticsAccurate_eq_compl {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : explorationStatisticsAccurate (n := n) (T := T) nu eps = (statisticsBadEvent nu eps)ᶜ","missing":[],"search":"explorationstatisticsaccurate_eq_compl banditrlproof.musicalchairs.explorationstatisticsaccurate_eq_compl theorem explorationstatisticsaccurate_eq_compl {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : explorationstatisticsaccurate (n := n) (t := t) nu eps = (statisticsbadevent nu eps)ᶜ theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_probability","label":"explorationStatisticsAccurate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationStatisticsAccurate_probability","description":"theorem explorationStatisticsAccurate_probability {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : 0 < T) (hmeans : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) (hcoll : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 1 - ENNReal.of…","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-2396f09b9d3f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1537,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:289"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationStatisticsAccurate_probability {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : 0 < T) (hmeans : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) (hcoll : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw hk nu (explorationStatisticsAccurate (n := n) (T := T) nu eps)","missing":[],"search":"explorationstatisticsaccurate_probability banditrlproof.musicalchairs.explorationstatisticsaccurate_probability theorem explorationstatisticsaccurate_probability {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (ht : 0 < t) (hmeans : (16*(k : ℝ)/eps^2) * real.log (4*(k : ℝ)^2/delta) ≤ t) (hcoll : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ t) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw hk nu (explorationstatisticsaccurate (n := n) (t := t) nu eps) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationLength","label":"explorationLength","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationLength","description":"noncomputable def explorationLength (k : ℕ) (eps delta : ℝ) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-33acedcf1a1d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1538,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:312"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationLength (k : ℕ) (eps delta : ℝ) : ℕ","missing":[],"search":"explorationlength banditrlproof.musicalchairs.explorationlength noncomputable def explorationlength (k : ℕ) (eps delta : ℝ) : ℕ definition compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationLength_mean_budget","label":"explorationLength_mean_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationLength_mean_budget","description":"theorem explorationLength_mean_budget (k : ℕ) (eps delta : ℝ) : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ explorationLength k eps delta","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-398c96bddb8d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1539,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:316"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationLength_mean_budget (k : ℕ) (eps delta : ℝ) : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ explorationLength k eps delta","missing":[],"search":"explorationlength_mean_budget banditrlproof.musicalchairs.explorationlength_mean_budget theorem explorationlength_mean_budget (k : ℕ) (eps delta : ℝ) : (16*(k : ℝ)/eps^2) * real.log (4*(k : ℝ)^2/delta) ≤ explorationlength k eps delta theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationLength_collision_budget","label":"explorationLength_collision_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationLength_collision_budget","description":"theorem explorationLength_collision_budget (k : ℕ) (eps delta : ℝ) : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ explorationLength k eps delta","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-22d31871b768","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1540,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:320"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationLength_collision_budget (k : ℕ) (eps delta : ℝ) : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ explorationLength k eps delta","missing":[],"search":"explorationlength_collision_budget banditrlproof.musicalchairs.explorationlength_collision_budget theorem explorationlength_collision_budget (k : ℕ) (eps delta : ℝ) : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ explorationlength k eps delta theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationLength_pos","label":"explorationLength_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationLength_pos","description":"theorem explorationLength_pos {k : ℕ} (hk : 0 < k) (eps delta : ℝ) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 0 < explorationLength k eps delta","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-784340b42ad3","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1541,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:324"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationLength_pos {k : ℕ} (hk : 0 < k) (eps delta : ℝ) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 0 < explorationLength k eps delta","missing":[],"search":"explorationlength_pos banditrlproof.musicalchairs.explorationlength_pos theorem explorationlength_pos {k : ℕ} (hk : 0 < k) (eps delta : ℝ) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 0 < explorationlength k eps delta theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_at_explorationLength","label":"explorationStatisticsAccurate_at_explorationLength","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationStatisticsAccurate_at_explorationLength","description":"theorem explorationStatisticsAccurate_at_explorationLength {n k : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw hk nu (explorationStatisticsAccurate (n := n) (T := explorationLength k e…","url":"../modules/banditrlproof-algorithms-musicalchairscollision/index.html#decl-513fc581dca7","parent":"module:BanditRLProof.Algorithms.MusicalChairsCollision","order":1542,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCollision"],["Source","BanditRLProof/Algorithms/MusicalChairsCollision.lean:335"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationStatisticsAccurate_at_explorationLength {n k : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw hk nu (explorationStatisticsAccurate (n := n) (T := explorationLength k eps delta) nu eps)","missing":[],"search":"explorationstatisticsaccurate_at_explorationlength banditrlproof.musicalchairs.explorationstatisticsaccurate_at_explorationlength theorem explorationstatisticsaccurate_at_explorationlength {n k : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw hk nu (explorationstatisticsaccurate (n := n) (t := explorationlength k eps delta) nu eps) theorem compiled","shard":"modules/48120dc14d0559d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.State","label":"State","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.State","description":"abbrev State (n k : ℕ)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-a86c155033bd","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1543,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"abbrev State (n k : ℕ)","missing":[],"search":"state banditrlproof.musicalchairs.state abbrev state (n k : ℕ) abbreviation compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.action","label":"action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.action","description":"def action {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) : Fin k","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-144f1d456dc9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1544,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def action {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) : Fin k","missing":[],"search":"action banditrlproof.musicalchairs.action def action {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) : fin k definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.CollisionFree","label":"CollisionFree","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.CollisionFree","description":"def CollisionFree {n k : ℕ} (a : Fin n → Fin k) (i : Fin n) : Prop","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-122dbcf62bc4","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1545,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def CollisionFree {n k : ℕ} (a : Fin n → Fin k) (i : Fin n) : Prop","missing":[],"search":"collisionfree banditrlproof.musicalchairs.collisionfree def collisionfree {n k : ℕ} (a : fin n → fin k) (i : fin n) : prop definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.step","label":"step","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.step","description":"noncomputable def step {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) : State n k","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-113f56a6252a","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1546,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def step {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) : State n k","missing":[],"search":"step banditrlproof.musicalchairs.step noncomputable def step {n k : ℕ} (s : state n k) (draw : fin n → fin k) : state n k definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.DistinctFixed","label":"DistinctFixed","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.DistinctFixed","description":"def DistinctFixed {n k : ℕ} (s : State n k) : Prop","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-3a5b9ffa3b52","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1547,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def DistinctFixed {n k : ℕ} (s : State n k) : Prop","missing":[],"search":"distinctfixed banditrlproof.musicalchairs.distinctfixed def distinctfixed {n k : ℕ} (s : state n k) : prop definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.action_of_fixed","label":"action_of_fixed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.action_of_fixed","description":"theorem action_of_fixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (h : s i = some a) : action s draw i = a","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-b62e484c779b","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1548,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem action_of_fixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (h : s i = some a) : action s draw i = a","missing":[],"search":"action_of_fixed banditrlproof.musicalchairs.action_of_fixed theorem action_of_fixed {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) (a : fin k) (h : s i = some a) : action s draw i = a theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.step_preserves_fixed","label":"step_preserves_fixed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.step_preserves_fixed","description":"theorem step_preserves_fixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (h : s i = some a) : step s draw i = some a","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-051a1fab6e4b","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1549,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem step_preserves_fixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (h : s i = some a) : step s draw i = some a","missing":[],"search":"step_preserves_fixed banditrlproof.musicalchairs.step_preserves_fixed theorem step_preserves_fixed {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) (a : fin k) (h : s i = some a) : step s draw i = some a theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.step_fixed_action","label":"step_fixed_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.step_fixed_action","description":"theorem step_fixed_action {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (h : step s draw i = some a) : action s draw i = a","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-ddc33934f2e9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1550,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem step_fixed_action {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (h : step s draw i = some a) : action s draw i = a","missing":[],"search":"step_fixed_action banditrlproof.musicalchairs.step_fixed_action theorem step_fixed_action {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) (a : fin k) (h : step s draw i = some a) : action s draw i = a theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.new_fixed_collisionFree","label":"new_fixed_collisionFree","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.new_fixed_collisionFree","description":"theorem new_fixed_collisionFree {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (hn : s i = none) (h : step s draw i = some a) : CollisionFree (action s draw) i","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-50bbc5da82c9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1551,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem new_fixed_collisionFree {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (hn : s i = none) (h : step s draw i = some a) : CollisionFree (action s draw) i","missing":[],"search":"new_fixed_collisionfree banditrlproof.musicalchairs.new_fixed_collisionfree theorem new_fixed_collisionfree {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) (a : fin k) (hn : s i = none) (h : step s draw i = some a) : collisionfree (action s draw) i theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.step_distinct","label":"step_distinct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.step_distinct","description":"theorem step_distinct {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : DistinctFixed (step s draw)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-5aad0bf9a5f0","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1552,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem step_distinct {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : DistinctFixed (step s draw)","missing":[],"search":"step_distinct banditrlproof.musicalchairs.step_distinct theorem step_distinct {n k : ℕ} (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) : distinctfixed (step s draw) theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.initial","label":"initial","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.initial","description":"def initial (n k : ℕ) : State n k","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-69e3188730e9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1553,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def initial (n k : ℕ) : State n k","missing":[],"search":"initial banditrlproof.musicalchairs.initial def initial (n k : ℕ) : state n k definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.initial_distinct","label":"initial_distinct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.initial_distinct","description":"theorem initial_distinct (n k : ℕ) : DistinctFixed (initial n k)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-dd4f3acfe43b","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1554,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem initial_distinct (n k : ℕ) : DistinctFixed (initial n k)","missing":[],"search":"initial_distinct banditrlproof.musicalchairs.initial_distinct theorem initial_distinct (n k : ℕ) : distinctfixed (initial n k) theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trajectory","label":"trajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trajectory","description":"noncomputable def trajectory {n k : ℕ} (draws : ℕ → Fin n → Fin k) : ℕ → State n k | 0 => initial n k | t+1 => step (trajectory draws t) (draws t) theorem trajectory_distinct {n k : ℕ} (draws : ℕ → Fin n → Fin k) (t : ℕ) : DistinctFixed (trajectory draws t)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-32c08e37da90","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1555,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def trajectory {n k : ℕ} (draws : ℕ → Fin n → Fin k) : ℕ → State n k | 0 => initial n k | t+1 => step (trajectory draws t) (draws t) theorem trajectory_distinct {n k : ℕ} (draws : ℕ → Fin n → Fin k) (t : ℕ) : DistinctFixed (trajectory draws t)","missing":[],"search":"trajectory banditrlproof.musicalchairs.trajectory noncomputable def trajectory {n k : ℕ} (draws : ℕ → fin n → fin k) : ℕ → state n k | 0 => initial n k | t+1 => step (trajectory draws t) (draws t) theorem trajectory_distinct {n k : ℕ} (draws : ℕ → fin n → fin k) (t : ℕ) : distinctfixed (trajectory draws t) definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trajectory_distinct","label":"trajectory_distinct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trajectory_distinct","description":"theorem trajectory_distinct {n k : ℕ} (draws : ℕ → Fin n → Fin k) (t : ℕ) : DistinctFixed (trajectory draws t)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-e2a9c3b143c0","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1556,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem trajectory_distinct {n k : ℕ} (draws : ℕ → Fin n → Fin k) (t : ℕ) : DistinctFixed (trajectory draws t)","missing":[],"search":"trajectory_distinct banditrlproof.musicalchairs.trajectory_distinct theorem trajectory_distinct {n k : ℕ} (draws : ℕ → fin n → fin k) (t : ℕ) : distinctfixed (trajectory draws t) theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.action_local","label":"action_local","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.action_local","description":"The action of player i uses only its own fixed arm and its own private draw.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-174195cbb7e0","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1557,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem action_local {n k : ℕ} (s s' : State n k) (draw draw' : Fin n → Fin k) (i : Fin n) (hs : s i = s' i) (hd : draw i = draw' i) : action s draw i = action s' draw' i","missing":[],"search":"action_local banditrlproof.musicalchairs.action_local the action of player i uses only its own fixed arm and its own private draw. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.jointDraw","label":"jointDraw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.jointDraw","description":"Uniform law on the finite Cartesian product of local candidate sets.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-920a840638f0","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1558,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def jointDraw {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) : PMF (Fin n → Fin k)","missing":[],"search":"jointdraw banditrlproof.musicalchairs.jointdraw uniform law on the finite cartesian product of local candidate sets. definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition","label":"transition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition","description":"noncomputable def transition {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (s : State n k) : PMF (State n k)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-7ad653f197c0","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1559,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def transition {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (s : State n k) : PMF (State n k)","missing":[],"search":"transition banditrlproof.musicalchairs.transition noncomputable def transition {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) (s : state n k) : pmf (state n k) definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.stateLaw","label":"stateLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.stateLaw","description":"noncomputable def stateLaw {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) : ℕ → PMF (State n k) | 0 => PMF.pure (initial n k) | t+1 => (stateLaw candidates hne t).bind (transition candidates hne) theorem transition_distinct {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (s s' : State n k) (hs : DistinctFixed s) (h : s' ∈ (transition candidat…","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-0dd4243c9591","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1560,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:106"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def stateLaw {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) : ℕ → PMF (State n k) | 0 => PMF.pure (initial n k) | t+1 => (stateLaw candidates hne t).bind (transition candidates hne) theorem transition_distinct {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (s s' : State n k) (hs : DistinctFixed s) (h : s' ∈ (transition candidates hne s).support) : DistinctFixed s'","missing":[],"search":"statelaw banditrlproof.musicalchairs.statelaw noncomputable def statelaw {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) : ℕ → pmf (state n k) | 0 => pmf.pure (initial n k) | t+1 => (statelaw candidates hne t).bind (transition candidates hne) theorem transition_distinct {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) (s s' : state n k) (hs : distinctfixed s) (h : s' ∈ (transition candidates hne s).support) : distinctfixed s' definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_distinct","label":"transition_distinct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_distinct","description":"theorem transition_distinct {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (s s' : State n k) (hs : DistinctFixed s) (h : s' ∈ (transition candidates hne s).support) : DistinctFixed s'","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-faa0f9a3ab61","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1561,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_distinct {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (s s' : State n k) (hs : DistinctFixed s) (h : s' ∈ (transition candidates hne s).support) : DistinctFixed s'","missing":[],"search":"transition_distinct banditrlproof.musicalchairs.transition_distinct theorem transition_distinct {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) (s s' : state n k) (hs : distinctfixed s) (h : s' ∈ (transition candidates hne s).support) : distinctfixed s' theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.stateLaw_distinct","label":"stateLaw_distinct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.stateLaw_distinct","description":"theorem stateLaw_distinct {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (t : ℕ) (s : State n k) (h : s ∈ (stateLaw candidates hne t).support) : DistinctFixed s","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-2baf150f0537","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1562,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem stateLaw_distinct {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (t : ℕ) (s : State n k) (h : s ∈ (stateLaw candidates hne t).support) : DistinctFixed s","missing":[],"search":"statelaw_distinct banditrlproof.musicalchairs.statelaw_distinct theorem statelaw_distinct {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) (t : ℕ) (s : state n k) (h : s ∈ (statelaw candidates hne t).support) : distinctfixed s theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localUpdate","label":"localUpdate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localUpdate","description":"A player's update consumes only its own old state, action, and collision bit.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-8085c94be6b8","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1563,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def localUpdate {k : ℕ} (old : Option (Fin k)) (played : Fin k) (collided : Bool) : Option (Fin k)","missing":[],"search":"localupdate banditrlproof.musicalchairs.localupdate a player's update consumes only its own old state, action, and collision bit. definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionBit","label":"collisionBit","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionBit","description":"noncomputable def collisionBit {n k : ℕ} (a : Fin n → Fin k) (i : Fin n) : Bool","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-b075e28d91bf","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1564,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def collisionBit {n k : ℕ} (a : Fin n → Fin k) (i : Fin n) : Bool","missing":[],"search":"collisionbit banditrlproof.musicalchairs.collisionbit noncomputable def collisionbit {n k : ℕ} (a : fin n → fin k) (i : fin n) : bool definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.step_localUpdate","label":"step_localUpdate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.step_localUpdate","description":"theorem step_localUpdate {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) : step s draw i = localUpdate (s i) (action s draw i) (collisionBit (action s draw) i)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-f0a2f90a78b2","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1565,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:141"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem step_localUpdate {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) : step s draw i = localUpdate (s i) (action s draw i) (collisionBit (action s draw) i)","missing":[],"search":"step_localupdate banditrlproof.musicalchairs.step_localupdate theorem step_localupdate {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) : step s draw i = localupdate (s i) (action s draw i) (collisionbit (action s draw) i) theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.jointDraw_mem","label":"jointDraw_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.jointDraw_mem","description":"theorem jointDraw_mem {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (draw : Fin n → Fin k) (h : draw ∈ (jointDraw candidates hne).support) (i : Fin n) : draw i ∈ candidates i","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-af9562d517f7","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1566,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem jointDraw_mem {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (draw : Fin n → Fin k) (h : draw ∈ (jointDraw candidates hne).support) (i : Fin n) : draw i ∈ candidates i","missing":[],"search":"jointdraw_mem banditrlproof.musicalchairs.jointdraw_mem theorem jointdraw_mem {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) (draw : fin n → fin k) (h : draw ∈ (jointdraw candidates hne).support) (i : fin n) : draw i ∈ candidates i theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.FixedWithin","label":"FixedWithin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.FixedWithin","description":"def FixedWithin {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (s : State n k) : Prop","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-270c0f487aab","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1567,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:159"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def FixedWithin {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (s : State n k) : Prop","missing":[],"search":"fixedwithin banditrlproof.musicalchairs.fixedwithin def fixedwithin {n k : ℕ} (candidates : fin n → finset (fin k)) (s : state n k) : prop definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.step_within","label":"step_within","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.step_within","description":"theorem step_within {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (s : State n k) (draw : Fin n → Fin k) (hs : FixedWithin candidates s) (hd : ∀ i, draw i ∈ candidates i) : FixedWithin candidates (step s draw)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-56ec8ba5586c","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1568,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:162"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem step_within {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (s : State n k) (draw : Fin n → Fin k) (hs : FixedWithin candidates s) (hd : ∀ i, draw i ∈ candidates i) : FixedWithin candidates (step s draw)","missing":[],"search":"step_within banditrlproof.musicalchairs.step_within theorem step_within {n k : ℕ} (candidates : fin n → finset (fin k)) (s : state n k) (draw : fin n → fin k) (hs : fixedwithin candidates s) (hd : ∀ i, draw i ∈ candidates i) : fixedwithin candidates (step s draw) theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.stateLaw_within","label":"stateLaw_within","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.stateLaw_within","description":"theorem stateLaw_within {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (t : ℕ) (s : State n k) (h : s ∈ (stateLaw candidates hne t).support) : FixedWithin candidates s","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-09b7f2ff597e","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1569,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:175"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem stateLaw_within {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) (t : ℕ) (s : State n k) (h : s ∈ (stateLaw candidates hne t).support) : FixedWithin candidates s","missing":[],"search":"statelaw_within banditrlproof.musicalchairs.statelaw_within theorem statelaw_within {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ i, (candidates i).nonempty) (t : ℕ) (s : state n k) (h : s ∈ (statelaw candidates hne t).support) : fixedwithin candidates s theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.jointDraw_rectangle","label":"jointDraw_rectangle","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.jointDraw_rectangle","description":"Exact rectangular-event law, including empty restrictions.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-44aab30baf3f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1570,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:191"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem jointDraw_rectangle {n k : ℕ} (candidates allowed : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) : (jointDraw candidates hne).toOuterMeasure {draw | ∀ i, draw i ∈ allowed i} = (∏ i, ((candidates i ∩ allowed i).card : ℝ≥0∞)) / (∏ i, ((candidates i).card : ℝ≥0∞))","missing":[],"search":"jointdraw_rectangle banditrlproof.musicalchairs.jointdraw_rectangle exact rectangular-event law, including empty restrictions. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.jointDraw_rectangle_product","label":"jointDraw_rectangle_product","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.jointDraw_rectangle_product","description":"Rectangle probabilities factor into their coordinate counting ratios.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-95246013520f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1571,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:213"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem jointDraw_rectangle_product {n k : ℕ} (candidates allowed : Fin n → Finset (Fin k)) (hne : ∀ i, (candidates i).Nonempty) : (jointDraw candidates hne).toOuterMeasure {draw | ∀ i, draw i ∈ allowed i} = ∏ i, ((candidates i ∩ allowed i).card : ℝ≥0∞) / ((candidates i).card : ℝ≥0∞)","missing":[],"search":"jointdraw_rectangle_product banditrlproof.musicalchairs.jointdraw_rectangle_product rectangle probabilities factor into their coordinate counting ratios. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.fixationWindow","label":"fixationWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.fixationWindow","description":"noncomputable def fixationWindow {n k : ℕ} (s : State n k) (i : Fin n) (a : Fin k) (j : Fin n) : Finset (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-0290681609ce","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1572,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:223"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def fixationWindow {n k : ℕ} (s : State n k) (i : Fin n) (a : Fin k) (j : Fin n) : Finset (Fin k)","missing":[],"search":"fixationwindow banditrlproof.musicalchairs.fixationwindow noncomputable def fixationwindow {n k : ℕ} (s : state n k) (i : fin n) (a : fin k) (j : fin n) : finset (fin k) definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.fixationWindow_fixes","label":"fixationWindow_fixes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.fixationWindow_fixes","description":"theorem fixationWindow_fixes {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (hi : s i = none) (ha : ∀ j b, s j = some b → b ≠ a) (hd : ∀ j, draw j ∈ fixationWindow s i a j) : step s draw i = some a","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-e43e807bda5a","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1573,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:227"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem fixationWindow_fixes {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (a : Fin k) (hi : s i = none) (ha : ∀ j b, s j = some b → b ≠ a) (hd : ∀ j, draw j ∈ fixationWindow s i a j) : step s draw i = some a","missing":[],"search":"fixationwindow_fixes banditrlproof.musicalchairs.fixationwindow_fixes theorem fixationwindow_fixes {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) (a : fin k) (hi : s i = none) (ha : ∀ j b, s j = some b → b ≠ a) (hd : ∀ j, draw j ∈ fixationwindow s i a j) : step s draw i = some a theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_fixation_lower","label":"transition_fixation_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_fixation_lower","description":"A genuine lower bound on the constructed next-state law, obtained from one explicit collision-free draw event. Candidate sets may differ.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-00deec52012c","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1574,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:245"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_fixation_lower {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ j, (candidates j).Nonempty) (s : State n k) (i : Fin n) (a : Fin k) (hi : s i = none) (ha : ∀ j b, s j = some b → b ≠ a) : (∏ j, ((candidates j ∩ fixationWindow s i a j).card : ℝ≥0∞) / ((candidates j).card : ℝ≥0∞)) ≤ (transition candidates hne s).toOuterMeasure {s' | s' i = some a}","missing":[],"search":"transition_fixation_lower banditrlproof.musicalchairs.transition_fixation_lower a genuine lower bound on the constructed next-state law, obtained from one explicit collision-free draw event. candidate sets may differ. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.isolationWindow","label":"isolationWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.isolationWindow","description":"noncomputable def isolationWindow {n k : ℕ} (i : Fin n) (a : Fin k) (j : Fin n) : Finset (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-fe62cb0eaf90","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1575,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:258"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def isolationWindow {n k : ℕ} (i : Fin n) (a : Fin k) (j : Fin n) : Finset (Fin k)","missing":[],"search":"isolationwindow banditrlproof.musicalchairs.isolationwindow noncomputable def isolationwindow {n k : ℕ} (i : fin n) (a : fin k) (j : fin n) : finset (fin k) definition compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.isolationWindow_subset","label":"isolationWindow_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.isolationWindow_subset","description":"theorem isolationWindow_subset {n k : ℕ} (s : State n k) (i : Fin n) (a : Fin k) (j : Fin n) : isolationWindow i a j ⊆ fixationWindow s i a j","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-db9369e87383","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1576,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:261"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem isolationWindow_subset {n k : ℕ} (s : State n k) (i : Fin n) (a : Fin k) (j : Fin n) : isolationWindow i a j ⊆ fixationWindow s i a j","missing":[],"search":"isolationwindow_subset banditrlproof.musicalchairs.isolationwindow_subset theorem isolationwindow_subset {n k : ℕ} (s : state n k) (i : fin n) (a : fin k) (j : fin n) : isolationwindow i a j ⊆ fixationwindow s i a j theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.jointDraw_isolation","label":"jointDraw_isolation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.jointDraw_isolation","description":"Exact probability of a tagged draw and avoidance by every other private coin.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-a2d04941fc3d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1577,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:268"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem jointDraw_isolation {n k : ℕ} (S : Finset (Fin k)) (i : Fin n) (a : Fin k) (ha : a ∈ S) (hcard : S.card = n) : (jointDraw (fun _ : Fin n => S) (fun _ => ⟨a, ha⟩)).toOuterMeasure {draw | ∀ j, draw j ∈ isolationWindow i a j} = (1 / (n : ℝ≥0∞)) * (((n-1 : ℕ) : ℝ≥0∞) / (n : ℝ≥0∞))^(n-1)","missing":[],"search":"jointdraw_isolation banditrlproof.musicalchairs.jointdraw_isolation exact probability of a tagged draw and avoidance by every other private coin. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_common_fixation_lower","label":"transition_common_fixation_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_common_fixation_lower","description":"Statewise fixation hazard from the actual joint law on a common N-arm set. Existence of an unused arm and the numerical 1/(4N) bound are separate obligations.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-af06b77a9ef7","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1578,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:295"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_common_fixation_lower {n k : ℕ} (S : Finset (Fin k)) (s : State n k) (i : Fin n) (a : Fin k) (hmem : a ∈ S) (hcard : S.card = n) (hi : s i = none) (ha : ∀ j b, s j = some b → b ≠ a) : (1 / (n : ℝ≥0∞)) * (((n-1 : ℕ) : ℝ≥0∞) / (n : ℝ≥0∞))^(n-1) ≤ (transition (fun _ : Fin n => S) (fun _ => ⟨a, hmem⟩) s).toOuterMeasure {s' | s' i = some a}","missing":[],"search":"transition_common_fixation_lower banditrlproof.musicalchairs.transition_common_fixation_lower statewise fixation hazard from the actual joint law on a common n-arm set. existence of an unused arm and the numerical 1/(4n) bound are separate obligations. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exists_unoccupied","label":"exists_unoccupied","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exists_unoccupied","description":"One unfixed player leaves at most N-1 occupied labels, so an N-label set contains an unused label. No successful coordination is assumed.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-6e95d2e944d7","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1579,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:311"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exists_unoccupied {n k : ℕ} (S : Finset (Fin k)) (s : State n k) (i : Fin n) (hi : s i = none) (hcard : S.card = n) : ∃ a ∈ S, ∀ j b, s j = some b → b ≠ a","missing":[],"search":"exists_unoccupied banditrlproof.musicalchairs.exists_unoccupied one unfixed player leaves at most n-1 occupied labels, so an n-label set contains an unused label. no successful coordination is assumed. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_hazard","label":"transition_unfixed_hazard","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_unfixed_hazard","description":"A statewise hazard for every unfixed player, produced without an externally supplied unused arm, settling event, or hazard premise.","url":"../modules/banditrlproof-algorithms-musicalchairscoordination/index.html#decl-d26077445615","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","order":1580,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordination"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordination.lean:341"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_unfixed_hazard {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (s : State n k) (i : Fin n) (hi : s i = none) (hcard : S.card = n) : (1 / (n : ℝ≥0∞)) * (((n-1 : ℕ) : ℝ≥0∞) / (n : ℝ≥0∞))^(n-1) ≤ (transition (fun _ : Fin n => S) (fun _ => hne) s).toOuterMeasure {s' | s' i ≠ none}","missing":[],"search":"transition_unfixed_hazard banditrlproof.musicalchairs.transition_unfixed_hazard a statewise hazard for every unfixed player, produced without an externally supplied unused arm, settling event, or hazard premise. theorem compiled","shard":"modules/1e0585c01fb5e6bb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.fixedPlayers","label":"fixedPlayers","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.fixedPlayers","description":"noncomputable def fixedPlayers {n k : ℕ} (s : State n k) : Finset (Fin n)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-9db7b33b9a65","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1581,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def fixedPlayers {n k : ℕ} (s : State n k) : Finset (Fin n)","missing":[],"search":"fixedplayers banditrlproof.musicalchairs.fixedplayers noncomputable def fixedplayers {n k : ℕ} (s : state n k) : finset (fin n) definition compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.hitFixed","label":"hitFixed","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.hitFixed","description":"noncomputable def hitFixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) : Finset (Fin n)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-1480db1a393d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1582,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def hitFixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) : Finset (Fin n)","missing":[],"search":"hitfixed banditrlproof.musicalchairs.hitfixed noncomputable def hitfixed {n k : ℕ} (s : state n k) (draw : fin n → fin k) : finset (fin n) definition compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.safeFixed","label":"safeFixed","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.safeFixed","description":"noncomputable def safeFixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) : Finset (Fin n)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-57dc136d65f2","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1583,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def safeFixed {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) : Finset (Fin n)","missing":[],"search":"safefixed banditrlproof.musicalchairs.safefixed noncomputable def safefixed {n k : ℕ} (s : state n k) (draw : fin n → fin k) : finset (fin n) definition compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.fixed_action_injective","label":"fixed_action_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.fixed_action_injective","description":"theorem fixed_action_injective {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : Set.InjOn (action s draw) (fixedPlayers s)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-6acb0ee1d094","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1584,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem fixed_action_injective {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : Set.InjOn (action s draw) (fixedPlayers s)","missing":[],"search":"fixed_action_injective banditrlproof.musicalchairs.fixed_action_injective theorem fixed_action_injective {n k : ℕ} (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) : set.injon (action s draw) (fixedplayers s) theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.hitFixed_partner","label":"hitFixed_partner","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.hitFixed_partner","description":"theorem hitFixed_partner {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) (i : Fin n) (hi : i ∈ hitFixed s draw) : ∃ j, s j = none ∧ action s draw j = action s draw i","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-3cc4be16a54b","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1585,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem hitFixed_partner {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) (i : Fin n) (hi : i ∈ hitFixed s draw) : ∃ j, s j = none ∧ action s draw j = action s draw i","missing":[],"search":"hitfixed_partner banditrlproof.musicalchairs.hitfixed_partner theorem hitfixed_partner {n k : ℕ} (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) (i : fin n) (hi : i ∈ hitfixed s draw) : ∃ j, s j = none ∧ action s draw j = action s draw i theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.hitFixed_card_le","label":"hitFixed_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.hitFixed_card_le","description":"theorem hitFixed_card_le {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : (hitFixed s draw).card ≤ unfixedCount s","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-f3bcca54c957","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1586,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem hitFixed_card_le {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : (hitFixed s draw).card ≤ unfixedCount s","missing":[],"search":"hitfixed_card_le banditrlproof.musicalchairs.hitfixed_card_le theorem hitfixed_card_le {n k : ℕ} (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) : (hitfixed s draw).card ≤ unfixedcount s theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.safeFixed_count","label":"safeFixed_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.safeFixed_count","description":"theorem safeFixed_count {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : n ≤ (safeFixed s draw).card + 2 * unfixedCount s","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-de5d84c7b791","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1587,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem safeFixed_count {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : n ≤ (safeFixed s draw).card + 2 * unfixedCount s","missing":[],"search":"safefixed_count banditrlproof.musicalchairs.safefixed_count theorem safefixed_count {n k : ℕ} (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) : n ≤ (safefixed s draw).card + 2 * unfixedcount s theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.roundMeanReward","label":"roundMeanReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.roundMeanReward","description":"noncomputable def roundMeanReward {n k : ℕ} (mu : Fin k → ℝ) (s : State n k) (draw : Fin n → Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-6b6480e0a39d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1588,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def roundMeanReward {n k : ℕ} (mu : Fin k → ℝ) (s : State n k) (draw : Fin n → Fin k) : ℝ","missing":[],"search":"roundmeanreward banditrlproof.musicalchairs.roundmeanreward noncomputable def roundmeanreward {n k : ℕ} (mu : fin k → ℝ) (s : state n k) (draw : fin n → fin k) : ℝ definition compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret","label":"roundPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.roundPseudoRegret","description":"noncomputable def roundPseudoRegret {n k : ℕ} (S : Finset (Fin k)) (mu : Fin k → ℝ) (s : State n k) (draw : Fin n → Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-042c6e4046bc","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1589,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def roundPseudoRegret {n k : ℕ} (S : Finset (Fin k)) (mu : Fin k → ℝ) (s : State n k) (draw : Fin n → Fin k) : ℝ","missing":[],"search":"roundpseudoregret banditrlproof.musicalchairs.roundpseudoregret noncomputable def roundpseudoregret {n k : ℕ} (s : finset (fin k)) (mu : fin k → ℝ) (s : state n k) (draw : fin n → fin k) : ℝ definition compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.safeFixed_mem","label":"safeFixed_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.safeFixed_mem","description":"theorem safeFixed_mem {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (hi : i ∈ safeFixed s draw) : i ∈ fixedPlayers s ∧ CollisionFree (action s draw) i","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-3a14bce5c85b","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1590,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem safeFixed_mem {n k : ℕ} (s : State n k) (draw : Fin n → Fin k) (i : Fin n) (hi : i ∈ safeFixed s draw) : i ∈ fixedPlayers s ∧ CollisionFree (action s draw) i","missing":[],"search":"safefixed_mem banditrlproof.musicalchairs.safefixed_mem theorem safefixed_mem {n k : ℕ} (s : state n k) (draw : fin n → fin k) (i : fin n) (hi : i ∈ safefixed s draw) : i ∈ fixedplayers s ∧ collisionfree (action s draw) i theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.safeFixed_arms_subset","label":"safeFixed_arms_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.safeFixed_arms_subset","description":"theorem safeFixed_arms_subset {n k : ℕ} (S : Finset (Fin k)) (s : State n k) (draw : Fin n → Fin k) (hw : FixedWithin (fun _ => S) s) : (safeFixed s draw).image (action s draw) ⊆ S","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-65bf170eae49","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1591,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem safeFixed_arms_subset {n k : ℕ} (S : Finset (Fin k)) (s : State n k) (draw : Fin n → Fin k) (hw : FixedWithin (fun _ => S) s) : (safeFixed s draw).image (action s draw) ⊆ S","missing":[],"search":"safefixed_arms_subset banditrlproof.musicalchairs.safefixed_arms_subset theorem safefixed_arms_subset {n k : ℕ} (s : finset (fin k)) (s : state n k) (draw : fin n → fin k) (hw : fixedwithin (fun _ => s) s) : (safefixed s draw).image (action s draw) ⊆ s theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.safeFixed_reward_le","label":"safeFixed_reward_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.safeFixed_reward_le","description":"theorem safeFixed_reward_le {n k : ℕ} (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : ∑ a ∈ (safeFixed s draw).image (action s draw), mu a ≤ roundMeanReward mu s draw","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-345c55f6276b","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1592,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem safeFixed_reward_le {n k : ℕ} (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) : ∑ a ∈ (safeFixed s draw).image (action s draw), mu a ≤ roundMeanReward mu s draw","missing":[],"search":"safefixed_reward_le banditrlproof.musicalchairs.safefixed_reward_le theorem safefixed_reward_le {n k : ℕ} (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) : ∑ a ∈ (safefixed s draw).image (action s draw), mu a ≤ roundmeanreward mu s draw theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_le_twice_unfixed","label":"roundPseudoRegret_le_twice_unfixed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.roundPseudoRegret_le_twice_unfixed","description":"theorem roundPseudoRegret_le_twice_unfixed {n k : ℕ} (S : Finset (Fin k)) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) (hw : FixedWithin (fun _ => S) s) : roundPseudoRegret S mu s draw ≤ 2 * (unfixedCount s : ℝ)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-cad400d831d9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1593,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem roundPseudoRegret_le_twice_unfixed {n k : ℕ} (S : Finset (Fin k)) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (s : State n k) (draw : Fin n → Fin k) (hs : DistinctFixed s) (hw : FixedWithin (fun _ => S) s) : roundPseudoRegret S mu s draw ≤ 2 * (unfixedCount s : ℝ)","missing":[],"search":"roundpseudoregret_le_twice_unfixed banditrlproof.musicalchairs.roundpseudoregret_le_twice_unfixed theorem roundpseudoregret_le_twice_unfixed {n k : ℕ} (s : finset (fin k)) (hcard : s.card = n) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (s : state n k) (draw : fin n → fin k) (hs : distinctfixed s) (hw : fixedwithin (fun _ => s) s) : roundpseudoregret s mu s draw ≤ 2 * (unfixedcount s : ℝ) theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_nonneg","label":"roundPseudoRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.roundPseudoRegret_nonneg","description":"theorem roundPseudoRegret_nonneg {n k : ℕ} (S : Finset (Fin k)) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (s : State n k) (draw : Fin n → Fin k) (hw : FixedWithin (fun _ => S) s) (hd : ∀ i, draw i ∈ S) : 0 ≤ roundPseudoRegret S mu s draw","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-75b2a904bdf2","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1594,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem roundPseudoRegret_nonneg {n k : ℕ} (S : Finset (Fin k)) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (s : State n k) (draw : Fin n → Fin k) (hw : FixedWithin (fun _ => S) s) (hd : ∀ i, draw i ∈ S) : 0 ≤ roundPseudoRegret S mu s draw","missing":[],"search":"roundpseudoregret_nonneg banditrlproof.musicalchairs.roundpseudoregret_nonneg theorem roundpseudoregret_nonneg {n k : ℕ} (s : finset (fin k)) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (s : state n k) (draw : fin n → fin k) (hw : fixedwithin (fun _ => s) s) (hd : ∀ i, draw i ∈ s) : 0 ≤ roundpseudoregret s mu s draw theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expected_round_charge","label":"expected_round_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expected_round_charge","description":"theorem expected_round_charge {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (s : State n k) (hs : DistinctFixed s) (hw : FixedWithin (fun _ => S) s) : ∑ draw, (jointDraw (fun _ : Fin n => S) (fun _ => hne)) draw * ENNReal.ofReal (roundPseudoRegret S mu s draw) ≤ 2 * (unfixedCount s : ℝ≥0∞)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-af89929936a1","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1595,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:173"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem expected_round_charge {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (s : State n k) (hs : DistinctFixed s) (hw : FixedWithin (fun _ => S) s) : ∑ draw, (jointDraw (fun _ : Fin n => S) (fun _ => hne)) draw * ENNReal.ofReal (roundPseudoRegret S mu s draw) ≤ 2 * (unfixedCount s : ℝ≥0∞)","missing":[],"search":"expected_round_charge banditrlproof.musicalchairs.expected_round_charge theorem expected_round_charge {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (hcard : s.card = n) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (s : state n k) (hs : distinctfixed s) (hw : fixedwithin (fun _ => s) s) : ∑ draw, (jointdraw (fun _ : fin n => s) (fun _ => hne)) draw * ennreal.ofreal (roundpseudoregret s mu s draw) ≤ 2 * (unfixedcount s : ℝ≥0∞) theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret","label":"expectedCoordinationRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expectedCoordinationRegret","description":"noncomputable def expectedCoordinationRegret {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) (T : ℕ) : ℝ≥0∞","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-1d07e3c69ba9","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1596,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:190"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def expectedCoordinationRegret {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) (T : ℕ) : ℝ≥0∞","missing":[],"search":"expectedcoordinationregret banditrlproof.musicalchairs.expectedcoordinationregret noncomputable def expectedcoordinationregret {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (mu : fin k → ℝ) (t : ℕ) : ℝ≥0∞ definition compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_le","label":"expectedCoordinationRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expectedCoordinationRegret_le","description":"theorem expectedCoordinationRegret_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (T : ℕ) : expectedCoordinationRegret (n := n) S hne mu T ≤ 8 * (n : ℝ≥0∞)^2","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-79d41b824dea","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1597,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:196"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem expectedCoordinationRegret_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (T : ℕ) : expectedCoordinationRegret (n := n) S hne mu T ≤ 8 * (n : ℝ≥0∞)^2","missing":[],"search":"expectedcoordinationregret_le banditrlproof.musicalchairs.expectedcoordinationregret_le theorem expectedcoordinationregret_le {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (hcard : s.card = n) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (t : ℕ) : expectedcoordinationregret (n := n) s hne mu t ≤ 8 * (n : ℝ≥0∞)^2 theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_toReal","label":"expectedCoordinationRegret_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expectedCoordinationRegret_toReal","description":"theorem expectedCoordinationRegret_toReal {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (T : ℕ) : (expectedCoordinationRegret (n := n) S hne mu T).toReal = ∑ t ∈ Finset.range T, ∑ s, ((stateLaw (fun _ : Fin n => S) (fun _ => hne) t) s).toReal * ∑ draw, ((jointDraw (fun _ : Fin n => S) (fun _ => hne)) draw).toReal * roundPseudoRegret S mu s draw","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-e793e165041e","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1598,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:228"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem expectedCoordinationRegret_toReal {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (T : ℕ) : (expectedCoordinationRegret (n := n) S hne mu T).toReal = ∑ t ∈ Finset.range T, ∑ s, ((stateLaw (fun _ : Fin n => S) (fun _ => hne) t) s).toReal * ∑ draw, ((jointDraw (fun _ : Fin n => S) (fun _ => hne)) draw).toReal * roundPseudoRegret S mu s draw","missing":[],"search":"expectedcoordinationregret_toreal banditrlproof.musicalchairs.expectedcoordinationregret_toreal theorem expectedcoordinationregret_toreal {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a) (t : ℕ) : (expectedcoordinationregret (n := n) s hne mu t).toreal = ∑ t ∈ finset.range t, ∑ s, ((statelaw (fun _ : fin n => s) (fun _ => hne) t) s).toreal * ∑ draw, ((jointdraw (fun _ : fin n => s) (fun _ => hne)) draw).toreal * roundpseudoregret s mu s draw theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_real_le","label":"expectedCoordinationRegret_real_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expectedCoordinationRegret_real_le","description":"theorem expectedCoordinationRegret_real_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (T : ℕ) : (∑ t ∈ Finset.range T, ∑ s, ((stateLaw (fun _ : Fin n => S) (fun _ => hne) t) s).toReal * ∑ draw, ((jointDraw (fun _ : Fin n => S) (fun _ => hne)) draw).toReal * roundPseudoRegret S mu s draw) ≤ 8 * (n : ℝ)^2","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationregret/index.html#decl-0646ecd4d98a","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","order":1599,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationRegret.lean:259"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem expectedCoordinationRegret_real_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (T : ℕ) : (∑ t ∈ Finset.range T, ∑ s, ((stateLaw (fun _ : Fin n => S) (fun _ => hne) t) s).toReal * ∑ draw, ((jointDraw (fun _ : Fin n => S) (fun _ => hne)) draw).toReal * roundPseudoRegret S mu s draw) ≤ 8 * (n : ℝ)^2","missing":[],"search":"expectedcoordinationregret_real_le banditrlproof.musicalchairs.expectedcoordinationregret_real_le theorem expectedcoordinationregret_real_le {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (hcard : s.card = n) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (t : ℕ) : (∑ t ∈ finset.range t, ∑ s, ((statelaw (fun _ : fin n => s) (fun _ => hne) t) s).toreal * ∑ draw, ((jointdraw (fun _ : fin n => s) (fun _ => hne)) draw).toreal * roundpseudoregret s mu s draw) ≤ 8 * (n : ℝ)^2 theorem compiled","shard":"modules/9950e6d728c0a33c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.quarter_le_avoidance_succ","label":"quarter_le_avoidance_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.quarter_le_avoidance_succ","description":"theorem quarter_le_avoidance_succ (m : ℕ) (hm : 0 < m) : (1/4 : ℝ) ≤ ((m : ℝ) / ((m : ℝ)+1))^m","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-a50430ef3acf","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1600,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem quarter_le_avoidance_succ (m : ℕ) (hm : 0 < m) : (1/4 : ℝ) ≤ ((m : ℝ) / ((m : ℝ)+1))^m","missing":[],"search":"quarter_le_avoidance_succ banditrlproof.musicalchairs.quarter_le_avoidance_succ theorem quarter_le_avoidance_succ (m : ℕ) (hm : 0 < m) : (1/4 : ℝ) ≤ ((m : ℝ) / ((m : ℝ)+1))^m theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.quarter_le_avoidance","label":"quarter_le_avoidance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.quarter_le_avoidance","description":"theorem quarter_le_avoidance (n : ℕ) (hn : 2 ≤ n) : (1/4 : ℝ) ≤ (((n-1 : ℕ) : ℝ) / (n : ℝ))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-c9b4816e34c8","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1601,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem quarter_le_avoidance (n : ℕ) (hn : 2 ≤ n) : (1/4 : ℝ) ≤ (((n-1 : ℕ) : ℝ) / (n : ℝ))^(n-1)","missing":[],"search":"quarter_le_avoidance banditrlproof.musicalchairs.quarter_le_avoidance theorem quarter_le_avoidance (n : ℕ) (hn : 2 ≤ n) : (1/4 : ℝ) ≤ (((n-1 : ℕ) : ℝ) / (n : ℝ))^(n-1) theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.real_uniform_hazard_lower","label":"real_uniform_hazard_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.real_uniform_hazard_lower","description":"theorem real_uniform_hazard_lower (n : ℕ) (hn : 0 < n) : (1 : ℝ) / (4 * n) ≤ (1 / (n : ℝ)) * (((n-1 : ℕ) : ℝ) / (n : ℝ))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-c7652c78952c","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1602,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem real_uniform_hazard_lower (n : ℕ) (hn : 0 < n) : (1 : ℝ) / (4 * n) ≤ (1 / (n : ℝ)) * (((n-1 : ℕ) : ℝ) / (n : ℝ))^(n-1)","missing":[],"search":"real_uniform_hazard_lower banditrlproof.musicalchairs.real_uniform_hazard_lower theorem real_uniform_hazard_lower (n : ℕ) (hn : 0 < n) : (1 : ℝ) / (4 * n) ≤ (1 / (n : ℝ)) * (((n-1 : ℕ) : ℝ) / (n : ℝ))^(n-1) theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.uniform_hazard_lower","label":"uniform_hazard_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.uniform_hazard_lower","description":"theorem uniform_hazard_lower (n : ℕ) (hn : 0 < n) : (1 : ℝ≥0∞) / (4 * n) ≤ (1 / (n : ℝ≥0∞)) * (((n-1 : ℕ) : ℝ≥0∞) / (n : ℝ≥0∞))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-15c2ddce2aae","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1603,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem uniform_hazard_lower (n : ℕ) (hn : 0 < n) : (1 : ℝ≥0∞) / (4 * n) ≤ (1 / (n : ℝ≥0∞)) * (((n-1 : ℕ) : ℝ≥0∞) / (n : ℝ≥0∞))^(n-1)","missing":[],"search":"uniform_hazard_lower banditrlproof.musicalchairs.uniform_hazard_lower theorem uniform_hazard_lower (n : ℕ) (hn : 0 < n) : (1 : ℝ≥0∞) / (4 * n) ≤ (1 / (n : ℝ≥0∞)) * (((n-1 : ℕ) : ℝ≥0∞) / (n : ℝ≥0∞))^(n-1) theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_hazard_quarter","label":"transition_unfixed_hazard_quarter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_unfixed_hazard_quarter","description":"theorem transition_unfixed_hazard_quarter {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (s : State n k) (i : Fin n) (hi : s i = none) (hcard : S.card = n) : (1 : ℝ≥0∞) / (4 * n) ≤ (transition (fun _ : Fin n => S) (fun _ => hne) s).toOuterMeasure {s' | s' i ≠ none}","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-a1a99c00edb7","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1604,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_unfixed_hazard_quarter {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (s : State n k) (i : Fin n) (hi : s i = none) (hcard : S.card = n) : (1 : ℝ≥0∞) / (4 * n) ≤ (transition (fun _ : Fin n => S) (fun _ => hne) s).toOuterMeasure {s' | s' i ≠ none}","missing":[],"search":"transition_unfixed_hazard_quarter banditrlproof.musicalchairs.transition_unfixed_hazard_quarter theorem transition_unfixed_hazard_quarter {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (s : state n k) (i : fin n) (hi : s i = none) (hcard : s.card = n) : (1 : ℝ≥0∞) / (4 * n) ≤ (transition (fun _ : fin n => s) (fun _ => hne) s).tooutermeasure {s' | s' i ≠ none} theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.event_add_compl","label":"event_add_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.event_add_compl","description":"theorem event_add_compl {α : Type*} (p : PMF α) (E : Set α) : p.toOuterMeasure E + p.toOuterMeasure Eᶜ = 1","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-8a3f6b04166e","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1605,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem event_add_compl {α : Type*} (p : PMF α) (E : Set α) : p.toOuterMeasure E + p.toOuterMeasure Eᶜ = 1","missing":[],"search":"event_add_compl banditrlproof.musicalchairs.event_add_compl theorem event_add_compl {α : type*} (p : pmf α) (e : set α) : p.tooutermeasure e + p.tooutermeasure eᶜ = 1 theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_fixed_no_return","label":"transition_fixed_no_return","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_fixed_no_return","description":"theorem transition_fixed_no_return {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ j, (candidates j).Nonempty) (s : State n k) (i : Fin n) (hi : s i ≠ none) : (transition candidates hne s).toOuterMeasure {s' | s' i = none} = 0","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-9931bab18799","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1606,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_fixed_no_return {n k : ℕ} (candidates : Fin n → Finset (Fin k)) (hne : ∀ j, (candidates j).Nonempty) (s : State n k) (i : Fin n) (hi : s i ≠ none) : (transition candidates hne s).toOuterMeasure {s' | s' i = none} = 0","missing":[],"search":"transition_fixed_no_return banditrlproof.musicalchairs.transition_fixed_no_return theorem transition_fixed_no_return {n k : ℕ} (candidates : fin n → finset (fin k)) (hne : ∀ j, (candidates j).nonempty) (s : state n k) (i : fin n) (hi : s i ≠ none) : (transition candidates hne s).tooutermeasure {s' | s' i = none} = 0 theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_le","label":"transition_unfixed_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.transition_unfixed_le","description":"theorem transition_unfixed_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (s : State n k) (i : Fin n) (hcard : S.card = n) : (transition (fun _ : Fin n => S) (fun _ => hne) s).toOuterMeasure {s' | s' i = none} ≤ if s i = none then 1 - (1 : ℝ≥0∞)/(4*n) else 0","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-4e23f9b3681f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1607,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem transition_unfixed_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (s : State n k) (i : Fin n) (hcard : S.card = n) : (transition (fun _ : Fin n => S) (fun _ => hne) s).toOuterMeasure {s' | s' i = none} ≤ if s i = none then 1 - (1 : ℝ≥0∞)/(4*n) else 0","missing":[],"search":"transition_unfixed_le banditrlproof.musicalchairs.transition_unfixed_le theorem transition_unfixed_le {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (s : state n k) (i : fin n) (hcard : s.card = n) : (transition (fun _ : fin n => s) (fun _ => hne) s).tooutermeasure {s' | s' i = none} ≤ if s i = none then 1 - (1 : ℝ≥0∞)/(4*n) else 0 theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.unfixed_survival","label":"unfixed_survival","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.unfixed_survival","description":"theorem unfixed_survival {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (i : Fin n) (hcard : S.card = n) (t : ℕ) : (stateLaw (fun _ : Fin n => S) (fun _ => hne) t).toOuterMeasure {s | s i = none} ≤ (1 - (1 : ℝ≥0∞)/(4*n))^t","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-d190f127d925","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1608,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem unfixed_survival {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (i : Fin n) (hcard : S.card = n) (t : ℕ) : (stateLaw (fun _ : Fin n => S) (fun _ => hne) t).toOuterMeasure {s | s i = none} ≤ (1 - (1 : ℝ≥0∞)/(4*n))^t","missing":[],"search":"unfixed_survival banditrlproof.musicalchairs.unfixed_survival theorem unfixed_survival {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (i : fin n) (hcard : s.card = n) (t : ℕ) : (statelaw (fun _ : fin n => s) (fun _ => hne) t).tooutermeasure {s | s i = none} ≤ (1 - (1 : ℝ≥0∞)/(4*n))^t theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.quarter_rate_le_one","label":"quarter_rate_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.quarter_rate_le_one","description":"theorem quarter_rate_le_one (n : ℕ) (hn : 0 < n) : (1 : ℝ≥0∞) / (4*n) ≤ 1","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-f3c39efc4a4d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1609,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:134"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem quarter_rate_le_one (n : ℕ) (hn : 0 < n) : (1 : ℝ≥0∞) / (4*n) ≤ 1","missing":[],"search":"quarter_rate_le_one banditrlproof.musicalchairs.quarter_rate_le_one theorem quarter_rate_le_one (n : ℕ) (hn : 0 < n) : (1 : ℝ≥0∞) / (4*n) ≤ 1 theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.unfixed_survival_sum","label":"unfixed_survival_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.unfixed_survival_sum","description":"theorem unfixed_survival_sum {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (i : Fin n) (hcard : S.card = n) (T : ℕ) : ∑ t ∈ Finset.range T, (stateLaw (fun _ : Fin n => S) (fun _ => hne) t).toOuterMeasure {s | s i = none} ≤ 4*n","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-d4b4553f503f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1610,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem unfixed_survival_sum {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (i : Fin n) (hcard : S.card = n) (T : ℕ) : ∑ t ∈ Finset.range T, (stateLaw (fun _ : Fin n => S) (fun _ => hne) t).toOuterMeasure {s | s i = none} ≤ 4*n","missing":[],"search":"unfixed_survival_sum banditrlproof.musicalchairs.unfixed_survival_sum theorem unfixed_survival_sum {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) (i : fin n) (hcard : s.card = n) (t : ℕ) : ∑ t ∈ finset.range t, (statelaw (fun _ : fin n => s) (fun _ => hne) t).tooutermeasure {s | s i = none} ≤ 4*n theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.total_unfixed_occupation_le","label":"total_unfixed_occupation_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.total_unfixed_occupation_le","description":"Expected total unfixed occupancy, expressed by actual finite-time marginals. The full pathwise regret adapter is a separate obligation.","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-6f725a73d63a","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1611,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:159"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem total_unfixed_occupation_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (T : ℕ) : ∑ t ∈ Finset.range T, ∑ i : Fin n, (stateLaw (fun _ : Fin n => S) (fun _ => hne) t).toOuterMeasure {s | s i = none} ≤ 4*(n : ℝ≥0∞)^2","missing":[],"search":"total_unfixed_occupation_le banditrlproof.musicalchairs.total_unfixed_occupation_le expected total unfixed occupancy, expressed by actual finite-time marginals. the full pathwise regret adapter is a separate obligation. theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.unfixedCount","label":"unfixedCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.unfixedCount","description":"noncomputable def unfixedCount {n k : ℕ} (s : State n k) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-ef08df3b866d","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1612,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:170"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def unfixedCount {n k : ℕ} (s : State n k) : ℕ","missing":[],"search":"unfixedcount banditrlproof.musicalchairs.unfixedcount noncomputable def unfixedcount {n k : ℕ} (s : state n k) : ℕ definition compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.unfixedCount_eq_sum","label":"unfixedCount_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.unfixedCount_eq_sum","description":"theorem unfixedCount_eq_sum {n k : ℕ} (s : State n k) : (unfixedCount s : ℝ≥0∞) = ∑ i : Fin n, if s i = none then 1 else 0","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-511c4c95941f","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1613,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:173"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem unfixedCount_eq_sum {n k : ℕ} (s : State n k) : (unfixedCount s : ℝ≥0∞) = ∑ i : Fin n, if s i = none then 1 else 0","missing":[],"search":"unfixedcount_eq_sum banditrlproof.musicalchairs.unfixedcount_eq_sum theorem unfixedcount_eq_sum {n k : ℕ} (s : state n k) : (unfixedcount s : ℝ≥0∞) = ∑ i : fin n, if s i = none then 1 else 0 theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expected_unfixed_eq_prob_sum","label":"expected_unfixed_eq_prob_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expected_unfixed_eq_prob_sum","description":"theorem expected_unfixed_eq_prob_sum {n k : ℕ} (p : PMF (State n k)) : ∑ s, p s * (unfixedCount s : ℝ≥0∞) = ∑ i : Fin n, p.toOuterMeasure {s | s i = none}","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-7f864007641c","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1614,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:178"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem expected_unfixed_eq_prob_sum {n k : ℕ} (p : PMF (State n k)) : ∑ s, p s * (unfixedCount s : ℝ≥0∞) = ∑ i : Fin n, p.toOuterMeasure {s | s i = none}","missing":[],"search":"expected_unfixed_eq_prob_sum banditrlproof.musicalchairs.expected_unfixed_eq_prob_sum theorem expected_unfixed_eq_prob_sum {n k : ℕ} (p : pmf (state n k)) : ∑ s, p s * (unfixedcount s : ℝ≥0∞) = ∑ i : fin n, p.tooutermeasure {s | s i = none} theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.expected_unfixed_occupation_le","label":"expected_unfixed_occupation_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.expected_unfixed_occupation_le","description":"Finite cumulative expected occupancy of the actual recursive state law. Counts are charged before each transition, including the initial all-unfixed state.","url":"../modules/banditrlproof-algorithms-musicalchairscoordinationtime/index.html#decl-0fc3d981fa71","parent":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","order":1615,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsCoordinationTime"],["Source","BanditRLProof/Algorithms/MusicalChairsCoordinationTime.lean:192"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem expected_unfixed_occupation_le {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (T : ℕ) : ∑ t ∈ Finset.range T, ∑ s, (stateLaw (fun _ : Fin n => S) (fun _ => hne) t) s * (unfixedCount s : ℝ≥0∞) ≤ 4*(n : ℝ≥0∞)^2","missing":[],"search":"expected_unfixed_occupation_le banditrlproof.musicalchairs.expected_unfixed_occupation_le finite cumulative expected occupancy of the actual recursive state law. counts are charged before each transition, including the initial all-unfixed state. theorem compiled","shard":"modules/ab2f68cc70feb0e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid","label":"iid","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid","description":"noncomputable def iid {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) : PMF (Fin T → α)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-4bff36460b9d","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1616,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def iid {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) : PMF (Fin T → α)","missing":[],"search":"iid banditrlproof.finitepmf.iid noncomputable def iid {α : type*} [fintype α] (p : pmf α) (t : ℕ) : pmf (fin t → α) definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_apply","label":"iid_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_apply","description":"theorem iid_apply {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (x : Fin T → α) : iid p T x = ∏ t, p (x t)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-a869244a0c07","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1617,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_apply {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (x : Fin T → α) : iid p T x = ∏ t, p (x t)","missing":[],"search":"iid_apply banditrlproof.finitepmf.iid_apply theorem iid_apply {α : type*} [fintype α] (p : pmf α) (t : ℕ) (x : fin t → α) : iid p t x = ∏ t, p (x t) theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_product_expectation","label":"iid_product_expectation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_product_expectation","description":"theorem iid_product_expectation {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (f : Fin T → α → ℝ≥0∞) : ∑ x, iid p T x * ∏ t, f t (x t) = ∏ t, ∑ a, p a * f t a","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-97669e159edb","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1618,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_product_expectation {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (f : Fin T → α → ℝ≥0∞) : ∑ x, iid p T x * ∏ t, f t (x t) = ∏ t, ∑ a, p a * f t a","missing":[],"search":"iid_product_expectation banditrlproof.finitepmf.iid_product_expectation theorem iid_product_expectation {α : type*} [fintype α] (p : pmf α) (t : ℕ) (f : fin t → α → ℝ≥0∞) : ∑ x, iid p t x * ∏ t, f t (x t) = ∏ t, ∑ a, p a * f t a theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.eventCount","label":"eventCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FinitePMF.eventCount","description":"noncomputable def eventCount {α : Type*} {T : ℕ} (E : Set α) (x : Fin T → α) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-c50f068a47f9","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1619,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def eventCount {α : Type*} {T : ℕ} (E : Set α) (x : Fin T → α) : ℕ","missing":[],"search":"eventcount banditrlproof.finitepmf.eventcount noncomputable def eventcount {α : type*} {t : ℕ} (e : set α) (x : fin t → α) : ℕ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.pow_eventCount","label":"pow_eventCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.pow_eventCount","description":"theorem pow_eventCount {α : Type*} {T : ℕ} (E : Set α) (x : Fin T → α) (z : ℝ≥0∞) : z ^ eventCount E x = ∏ t, if x t ∈ E then z else 1","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-0080325d7534","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1620,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem pow_eventCount {α : Type*} {T : ℕ} (E : Set α) (x : Fin T → α) (z : ℝ≥0∞) : z ^ eventCount E x = ∏ t, if x t ∈ E then z else 1","missing":[],"search":"pow_eventcount banditrlproof.finitepmf.pow_eventcount theorem pow_eventcount {α : type*} {t : ℕ} (e : set α) (x : fin t → α) (z : ℝ≥0∞) : z ^ eventcount e x = ∏ t, if x t ∈ e then z else 1 theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.indicator_weight_sum","label":"indicator_weight_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.indicator_weight_sum","description":"theorem indicator_weight_sum {α : Type*} [Fintype α] (p : PMF α) (E : Set α) (z : ℝ≥0∞) : ∑ a, p a * (if a ∈ E then z else 1) = p.toOuterMeasure E * z + p.toOuterMeasure Eᶜ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-ab65183e48d0","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1621,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem indicator_weight_sum {α : Type*} [Fintype α] (p : PMF α) (E : Set α) (z : ℝ≥0∞) : ∑ a, p a * (if a ∈ E then z else 1) = p.toOuterMeasure E * z + p.toOuterMeasure Eᶜ","missing":[],"search":"indicator_weight_sum banditrlproof.finitepmf.indicator_weight_sum theorem indicator_weight_sum {α : type*} [fintype α] (p : pmf α) (e : set α) (z : ℝ≥0∞) : ∑ a, p a * (if a ∈ e then z else 1) = p.tooutermeasure e * z + p.tooutermeasure eᶜ theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_count_pgf","label":"iid_count_pgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_count_pgf","description":"theorem iid_count_pgf {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (E : Set α) (z : ℝ≥0∞) : ∑ x, iid p T x * z ^ eventCount E x = (p.toOuterMeasure E * z + p.toOuterMeasure Eᶜ)^T","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-533e6b6db6a1","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1622,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_count_pgf {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (E : Set α) (z : ℝ≥0∞) : ∑ x, iid p T x * z ^ eventCount E x = (p.toOuterMeasure E * z + p.toOuterMeasure Eᶜ)^T","missing":[],"search":"iid_count_pgf banditrlproof.finitepmf.iid_count_pgf theorem iid_count_pgf {α : type*} [fintype α] (p : pmf α) (t : ℕ) (e : set α) (z : ℝ≥0∞) : ∑ x, iid p t x * z ^ eventcount e x = (p.tooutermeasure e * z + p.tooutermeasure eᶜ)^t theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationDraw","label":"explorationDraw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationDraw","description":"noncomputable def explorationDraw (n k : ℕ) (hk : 0 < k) : PMF (Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-226910a6e8bc","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1623,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationDraw (n k : ℕ) (hk : 0 < k) : PMF (Fin n → Fin k)","missing":[],"search":"explorationdraw banditrlproof.musicalchairs.explorationdraw noncomputable def explorationdraw (n k : ℕ) (hk : 0 < k) : pmf (fin n → fin k) definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationLaw","label":"explorationLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationLaw","description":"noncomputable def explorationLaw (n k T : ℕ) (hk : 0 < k) : PMF (Fin T → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-e4a38f245059","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1624,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationLaw (n k T : ℕ) (hk : 0 < k) : PMF (Fin T → Fin n → Fin k)","missing":[],"search":"explorationlaw banditrlproof.musicalchairs.explorationlaw noncomputable def explorationlaw (n k t : ℕ) (hk : 0 < k) : pmf (fin t → fin n → fin k) definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.observes","label":"observes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.observes","description":"def observes {n k : ℕ} (i : Fin n) (a : Fin k) : Set (Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-d54757e25ab2","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1625,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def observes {n k : ℕ} (i : Fin n) (a : Fin k) : Set (Fin n → Fin k)","missing":[],"search":"observes banditrlproof.musicalchairs.observes def observes {n k : ℕ} (i : fin n) (a : fin k) : set (fin n → fin k) definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.observes_rectangle","label":"observes_rectangle","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.observes_rectangle","description":"theorem observes_rectangle {n k : ℕ} (i : Fin n) (a : Fin k) : observes i a = {draw | ∀ j, draw j ∈ isolationWindow i a j}","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-3d4c0212b72d","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1626,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem observes_rectangle {n k : ℕ} (i : Fin n) (a : Fin k) : observes i a = {draw | ∀ j, draw j ∈ isolationWindow i a j}","missing":[],"search":"observes_rectangle banditrlproof.musicalchairs.observes_rectangle theorem observes_rectangle {n k : ℕ} (i : fin n) (a : fin k) : observes i a = {draw | ∀ j, draw j ∈ isolationwindow i a j} theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exploration_observes_probability","label":"exploration_observes_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exploration_observes_probability","description":"theorem exploration_observes_probability {n k : ℕ} (hk : 0 < k) (i : Fin n) (a : Fin k) : (explorationDraw n k hk).toOuterMeasure (observes i a) = (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-e1d55b4c3e67","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1627,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:80"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exploration_observes_probability {n k : ℕ} (hk : 0 < k) (i : Fin n) (a : Fin k) : (explorationDraw n k hk).toOuterMeasure (observes i a) = (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1)","missing":[],"search":"exploration_observes_probability banditrlproof.musicalchairs.exploration_observes_probability theorem exploration_observes_probability {n k : ℕ} (hk : 0 < k) (i : fin n) (a : fin k) : (explorationdraw n k hk).tooutermeasure (observes i a) = (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exploration_collisionFree_split","label":"exploration_collisionFree_split","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exploration_collisionFree_split","description":"theorem exploration_collisionFree_split {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toOuterMeasure {draw | CollisionFree draw i} = ∑ a : Fin k, (explorationDraw n k hk).toOuterMeasure (observes i a)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-6b64b4eff037","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1628,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exploration_collisionFree_split {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toOuterMeasure {draw | CollisionFree draw i} = ∑ a : Fin k, (explorationDraw n k hk).toOuterMeasure (observes i a)","missing":[],"search":"exploration_collisionfree_split banditrlproof.musicalchairs.exploration_collisionfree_split theorem exploration_collisionfree_split {n k : ℕ} (hk : 0 < k) (i : fin n) : (explorationdraw n k hk).tooutermeasure {draw | collisionfree draw i} = ∑ a : fin k, (explorationdraw n k hk).tooutermeasure (observes i a) theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exploration_collisionFree_probability","label":"exploration_collisionFree_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exploration_collisionFree_probability","description":"theorem exploration_collisionFree_probability {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toOuterMeasure {draw | CollisionFree draw i} = (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-8fa717340616","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1629,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exploration_collisionFree_probability {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toOuterMeasure {draw | CollisionFree draw i} = (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1)","missing":[],"search":"exploration_collisionfree_probability banditrlproof.musicalchairs.exploration_collisionfree_probability theorem exploration_collisionfree_probability {n k : ℕ} (hk : 0 < k) (i : fin n) : (explorationdraw n k hk).tooutermeasure {draw | collisionfree draw i} = (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exploration_collision_probability","label":"exploration_collision_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exploration_collision_probability","description":"theorem exploration_collision_probability {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toOuterMeasure {draw | ¬ CollisionFree draw i} = 1 - (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-13a660dceb13","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1630,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exploration_collision_probability {n k : ℕ} (hk : 0 < k) (i : Fin n) : (explorationDraw n k hk).toOuterMeasure {draw | ¬ CollisionFree draw i} = 1 - (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1)","missing":[],"search":"exploration_collision_probability banditrlproof.musicalchairs.exploration_collision_probability theorem exploration_collision_probability {n k : ℕ} (hk : 0 < k) (i : fin n) : (explorationdraw n k hk).tooutermeasure {draw | ¬ collisionfree draw i} = 1 - (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.observationCount","label":"observationCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.observationCount","description":"noncomputable def observationCount {n k T : ℕ} (i : Fin n) (a : Fin k) (x : Fin T → Fin n → Fin k) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-7d95fec0fb78","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1631,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:141"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def observationCount {n k T : ℕ} (i : Fin n) (a : Fin k) (x : Fin T → Fin n → Fin k) : ℕ","missing":[],"search":"observationcount banditrlproof.musicalchairs.observationcount noncomputable def observationcount {n k t : ℕ} (i : fin n) (a : fin k) (x : fin t → fin n → fin k) : ℕ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionCount","label":"collisionCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionCount","description":"noncomputable def collisionCount {n k T : ℕ} (i : Fin n) (x : Fin T → Fin n → Fin k) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-acfa8727b7ba","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1632,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:144"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def collisionCount {n k T : ℕ} (i : Fin n) (x : Fin T → Fin n → Fin k) : ℕ","missing":[],"search":"collisioncount banditrlproof.musicalchairs.collisioncount noncomputable def collisioncount {n k t : ℕ} (i : fin n) (x : fin t → fin n → fin k) : ℕ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.observationCount_pgf","label":"observationCount_pgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.observationCount_pgf","description":"theorem observationCount_pgf {n k : ℕ} (hk : 0 < k) (T : ℕ) (i : Fin n) (a : Fin k) (z : ℝ≥0∞) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationLaw n k T hk x * z ^ observationCount i a x = (q * z + (1-q))^T","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-6f5c160cec68","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1633,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem observationCount_pgf {n k : ℕ} (hk : 0 < k) (T : ℕ) (i : Fin n) (a : Fin k) (z : ℝ≥0∞) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationLaw n k T hk x * z ^ observationCount i a x = (q * z + (1-q))^T","missing":[],"search":"observationcount_pgf banditrlproof.musicalchairs.observationcount_pgf theorem observationcount_pgf {n k : ℕ} (hk : 0 < k) (t : ℕ) (i : fin n) (a : fin k) (z : ℝ≥0∞) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationlaw n k t hk x * z ^ observationcount i a x = (q * z + (1-q))^t theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionCount_pgf","label":"collisionCount_pgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionCount_pgf","description":"theorem collisionCount_pgf {n k : ℕ} (hk : 0 < k) (T : ℕ) (i : Fin n) (z : ℝ≥0∞) : let b := (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationLaw n k T hk x * z ^ collisionCount i x = ((1-b) * z + b)^T","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-8ec936600841","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1634,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:170"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collisionCount_pgf {n k : ℕ} (hk : 0 < k) (T : ℕ) (i : Fin n) (z : ℝ≥0∞) : let b := (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationLaw n k T hk x * z ^ collisionCount i x = ((1-b) * z + b)^T","missing":[],"search":"collisioncount_pgf banditrlproof.musicalchairs.collisioncount_pgf theorem collisioncount_pgf {n k : ℕ} (hk : 0 < k) (t : ℕ) (i : fin n) (z : ℝ≥0∞) : let b := (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationlaw n k t hk x * z ^ collisioncount i x = ((1-b) * z + b)^t theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exploration_observes_lower","label":"exploration_observes_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exploration_observes_lower","description":"theorem exploration_observes_lower {n k : ℕ} (hk : 0 < k) (hnk : n ≤ k) (i : Fin n) (a : Fin k) : (1 : ℝ≥0∞) / (4*k) ≤ (explorationDraw n k hk).toOuterMeasure (observes i a)","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-4cb7d30573ec","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1635,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:184"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exploration_observes_lower {n k : ℕ} (hk : 0 < k) (hnk : n ≤ k) (i : Fin n) (a : Fin k) : (1 : ℝ≥0∞) / (4*k) ≤ (explorationDraw n k hk).toOuterMeasure (observes i a)","missing":[],"search":"exploration_observes_lower banditrlproof.musicalchairs.exploration_observes_lower theorem exploration_observes_lower {n k : ℕ} (hk : 0 < k) (hnk : n ≤ k) (i : fin n) (a : fin k) : (1 : ℝ≥0∞) / (4*k) ≤ (explorationdraw n k hk).tooutermeasure (observes i a) theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.observationCount_laplace","label":"observationCount_laplace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.observationCount_laplace","description":"theorem observationCount_laplace {n k : ℕ} (hk : 0 < k) (T : ℕ) (i : Fin n) (a : Fin k) (eta : ℝ) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationLaw n k T hk x * ENNReal.ofReal (Real.exp (-eta * (observationCount i a x : ℝ))) = (q * ENNReal.ofReal (Real.exp (-eta)) + (1-q))^T","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-cfff81b3ba0b","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1636,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:202"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem observationCount_laplace {n k : ℕ} (hk : 0 < k) (T : ℕ) (i : Fin n) (a : Fin k) (eta : ℝ) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationLaw n k T hk x * ENNReal.ofReal (Real.exp (-eta * (observationCount i a x : ℝ))) = (q * ENNReal.ofReal (Real.exp (-eta)) + (1-q))^T","missing":[],"search":"observationcount_laplace banditrlproof.musicalchairs.observationcount_laplace theorem observationcount_laplace {n k : ℕ} (hk : 0 < k) (t : ℕ) (i : fin n) (a : fin k) (eta : ℝ) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) ∑ x, explorationlaw n k t hk x * ennreal.ofreal (real.exp (-eta * (observationcount i a x : ℝ))) = (q * ennreal.ofreal (real.exp (-eta)) + (1-q))^t theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.ExplorationFeedback","label":"ExplorationFeedback","kind":"structure","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.ExplorationFeedback","description":"structure ExplorationFeedback (k : ℕ) where","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-39a394e997f7","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1637,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:217"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"structure ExplorationFeedback (k : ℕ) where","missing":[],"search":"explorationfeedback banditrlproof.musicalchairs.explorationfeedback structure explorationfeedback (k : ℕ) where structure compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationFeedback","label":"explorationFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationFeedback","description":"noncomputable def explorationFeedback {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) : Fin T → ExplorationFeedback k","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-4aad91b8f3f0","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1638,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:222"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationFeedback {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) : Fin T → ExplorationFeedback k","missing":[],"search":"explorationfeedback banditrlproof.musicalchairs.explorationfeedback noncomputable def explorationfeedback {n k t : ℕ} (x : fin t → fin n → fin k) (rewards : fin t → fin k → ℝ) (i : fin n) : fin t → explorationfeedback k definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localObservationCount","label":"localObservationCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localObservationCount","description":"noncomputable def localObservationCount {k T : ℕ} (f : Fin T → ExplorationFeedback k) (a : Fin k) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-392f0ee6a093","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1639,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:227"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localObservationCount {k T : ℕ} (f : Fin T → ExplorationFeedback k) (a : Fin k) : ℕ","missing":[],"search":"localobservationcount banditrlproof.musicalchairs.localobservationcount noncomputable def localobservationcount {k t : ℕ} (f : fin t → explorationfeedback k) (a : fin k) : ℕ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCollisionCount","label":"localCollisionCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCollisionCount","description":"noncomputable def localCollisionCount {k T : ℕ} (f : Fin T → ExplorationFeedback k) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-4659214b4da9","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1640,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:231"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localCollisionCount {k T : ℕ} (f : Fin T → ExplorationFeedback k) : ℕ","missing":[],"search":"localcollisioncount banditrlproof.musicalchairs.localcollisioncount noncomputable def localcollisioncount {k t : ℕ} (f : fin t → explorationfeedback k) : ℕ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localRewardSum","label":"localRewardSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localRewardSum","description":"noncomputable def localRewardSum {k T : ℕ} (f : Fin T → ExplorationFeedback k) (a : Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-4fe74a50f58d","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1641,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:234"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localRewardSum {k T : ℕ} (f : Fin T → ExplorationFeedback k) (a : Fin k) : ℝ","missing":[],"search":"localrewardsum banditrlproof.musicalchairs.localrewardsum noncomputable def localrewardsum {k t : ℕ} (f : fin t → explorationfeedback k) (a : fin k) : ℝ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localEmpiricalMean","label":"localEmpiricalMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localEmpiricalMean","description":"noncomputable def localEmpiricalMean {k T : ℕ} (f : Fin T → ExplorationFeedback k) (a : Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-ea00190ecb9b","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1642,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:238"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localEmpiricalMean {k T : ℕ} (f : Fin T → ExplorationFeedback k) (a : Fin k) : ℝ","missing":[],"search":"localempiricalmean banditrlproof.musicalchairs.localempiricalmean noncomputable def localempiricalmean {k t : ℕ} (f : fin t → explorationfeedback k) (a : fin k) : ℝ definition compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localObservationCount_eq","label":"localObservationCount_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localObservationCount_eq","description":"theorem localObservationCount_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) (a : Fin k) : localObservationCount (explorationFeedback x rewards i) a = observationCount i a x","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-3f8bd21dab95","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1643,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:242"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localObservationCount_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) (a : Fin k) : localObservationCount (explorationFeedback x rewards i) a = observationCount i a x","missing":[],"search":"localobservationcount_eq banditrlproof.musicalchairs.localobservationcount_eq theorem localobservationcount_eq {n k t : ℕ} (x : fin t → fin n → fin k) (rewards : fin t → fin k → ℝ) (i : fin n) (a : fin k) : localobservationcount (explorationfeedback x rewards i) a = observationcount i a x theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCollisionCount_eq","label":"localCollisionCount_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCollisionCount_eq","description":"theorem localCollisionCount_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) : localCollisionCount (explorationFeedback x rewards i) = collisionCount i x","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-a2dcc1c091fd","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1644,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:250"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCollisionCount_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) : localCollisionCount (explorationFeedback x rewards i) = collisionCount i x","missing":[],"search":"localcollisioncount_eq banditrlproof.musicalchairs.localcollisioncount_eq theorem localcollisioncount_eq {n k t : ℕ} (x : fin t → fin n → fin k) (rewards : fin t → fin k → ℝ) (i : fin n) : localcollisioncount (explorationfeedback x rewards i) = collisioncount i x theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localRewardSum_eq","label":"localRewardSum_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localRewardSum_eq","description":"theorem localRewardSum_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) (a : Fin k) : localRewardSum (explorationFeedback x rewards i) a = ∑ t ∈ Finset.univ.filter (fun t => x t ∈ observes i a), rewards t a","url":"../modules/banditrlproof-algorithms-musicalchairsexploration/index.html#decl-657ebbb6bfca","parent":"module:BanditRLProof.Algorithms.MusicalChairsExploration","order":1645,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsExploration"],["Source","BanditRLProof/Algorithms/MusicalChairsExploration.lean:258"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localRewardSum_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (rewards : Fin T → Fin k → ℝ) (i : Fin n) (a : Fin k) : localRewardSum (explorationFeedback x rewards i) a = ∑ t ∈ Finset.univ.filter (fun t => x t ∈ observes i a), rewards t a","missing":[],"search":"localrewardsum_eq banditrlproof.musicalchairs.localrewardsum_eq theorem localrewardsum_eq {n k t : ℕ} (x : fin t → fin n → fin k) (rewards : fin t → fin k → ℝ) (i : fin n) (a : fin k) : localrewardsum (explorationfeedback x rewards i) a = ∑ t ∈ finset.univ.filter (fun t => x t ∈ observes i a), rewards t a theorem compiled","shard":"modules/f7c1f4654d145886.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.rankedArm","label":"rankedArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.rankedArm","description":"noncomputable def rankedArm {k : ℕ} (score : Fin k → ℝ) (j : Fin k) : Fin k","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-8c812f208f79","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1646,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def rankedArm {k : ℕ} (score : Fin k → ℝ) (j : Fin k) : Fin k","missing":[],"search":"rankedarm banditrlproof.musicalchairs.rankedarm noncomputable def rankedarm {k : ℕ} (score : fin k → ℝ) (j : fin k) : fin k definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.rankIndex","label":"rankIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.rankIndex","description":"noncomputable def rankIndex {k : ℕ} (score : Fin k → ℝ) (a : Fin k) : Fin k","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-e233fcb147e4","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1647,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def rankIndex {k : ℕ} (score : Fin k → ℝ) (a : Fin k) : Fin k","missing":[],"search":"rankindex banditrlproof.musicalchairs.rankindex noncomputable def rankindex {k : ℕ} (score : fin k → ℝ) (a : fin k) : fin k definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.rankedArm_rankIndex","label":"rankedArm_rankIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.rankedArm_rankIndex","description":"theorem rankedArm_rankIndex {k : ℕ} (score : Fin k → ℝ) (a : Fin k) : rankedArm score (rankIndex score a) = a","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-cd589b450c6d","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1648,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem rankedArm_rankIndex {k : ℕ} (score : Fin k → ℝ) (a : Fin k) : rankedArm score (rankIndex score a) = a","missing":[],"search":"rankedarm_rankindex banditrlproof.musicalchairs.rankedarm_rankindex theorem rankedarm_rankindex {k : ℕ} (score : fin k → ℝ) (a : fin k) : rankedarm score (rankindex score a) = a theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.rankIndex_rankedArm","label":"rankIndex_rankedArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.rankIndex_rankedArm","description":"theorem rankIndex_rankedArm {k : ℕ} (score : Fin k → ℝ) (j : Fin k) : rankIndex score (rankedArm score j) = j","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-e6c4c6ffb463","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1649,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem rankIndex_rankedArm {k : ℕ} (score : Fin k → ℝ) (j : Fin k) : rankIndex score (rankedArm score j) = j","missing":[],"search":"rankindex_rankedarm banditrlproof.musicalchairs.rankindex_rankedarm theorem rankindex_rankedarm {k : ℕ} (score : fin k → ℝ) (j : fin k) : rankindex score (rankedarm score j) = j theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.mem_topArms_iff_rankIndex","label":"mem_topArms_iff_rankIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.mem_topArms_iff_rankIndex","description":"theorem mem_topArms_iff_rankIndex {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (a : Fin k) : a ∈ topArms score n ↔ (rankIndex score a).val < n","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-e1466a16fd4d","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1650,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem mem_topArms_iff_rankIndex {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (a : Fin k) : a ∈ topArms score n ↔ (rankIndex score a).val < n","missing":[],"search":"mem_toparms_iff_rankindex banditrlproof.musicalchairs.mem_toparms_iff_rankindex theorem mem_toparms_iff_rankindex {k : ℕ} (score : fin k → ℝ) (n : ℕ) (a : fin k) : a ∈ toparms score n ↔ (rankindex score a).val < n theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.ranked_score_antitone","label":"ranked_score_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.ranked_score_antitone","description":"theorem ranked_score_antitone {k : ℕ} (score : Fin k → ℝ) {i j : Fin k} (hij : i ≤ j) : score (rankedArm score j) ≤ score (rankedArm score i)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-3de256b8b913","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1651,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem ranked_score_antitone {k : ℕ} (score : Fin k → ℝ) {i j : Fin k} (hij : i ≤ j) : score (rankedArm score j) ≤ score (rankedArm score i)","missing":[],"search":"ranked_score_antitone banditrlproof.musicalchairs.ranked_score_antitone theorem ranked_score_antitone {k : ℕ} (score : fin k → ℝ) {i j : fin k} (hij : i ≤ j) : score (rankedarm score j) ≤ score (rankedarm score i) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.boundaryGap","label":"boundaryGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.boundaryGap","description":"noncomputable def boundaryGap {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-9074b281eaaf","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1652,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def boundaryGap {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : ℝ","missing":[],"search":"boundarygap banditrlproof.musicalchairs.boundarygap noncomputable def boundarygap {k : ℕ} (score : fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : ℝ definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.boundaryGap_nonneg","label":"boundaryGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.boundaryGap_nonneg","description":"theorem boundaryGap_nonneg {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : 0 ≤ boundaryGap score n hn hnk","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-d695aaab27ff","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1653,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem boundaryGap_nonneg {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : 0 ≤ boundaryGap score n hn hnk","missing":[],"search":"boundarygap_nonneg banditrlproof.musicalchairs.boundarygap_nonneg theorem boundarygap_nonneg {k : ℕ} (score : fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : 0 ≤ boundarygap score n hn hnk theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.boundaryGap_separates","label":"boundaryGap_separates","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.boundaryGap_separates","description":"theorem boundaryGap_separates {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : ∀ a ∈ topArms score n, ∀ b ∉ topArms score n, boundaryGap score n hn hnk ≤ score a - score b","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-e97af1c1fd20","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1654,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem boundaryGap_separates {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : ∀ a ∈ topArms score n, ∀ b ∉ topArms score n, boundaryGap score n hn hnk ≤ score a - score b","missing":[],"search":"boundarygap_separates banditrlproof.musicalchairs.boundarygap_separates theorem boundarygap_separates {k : ℕ} (score : fin k → ℝ) (n : ℕ) (hn : 0 < n) (hnk : n < k) : ∀ a ∈ toparms score n, ∀ b ∉ toparms score n, boundarygap score n hn hnk ≤ score a - score b theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_orderStatistic_probability","label":"explorationGoodEvent_orderStatistic_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationGoodEvent_orderStatistic_probability","description":"theorem explorationGoodEvent_orderStatistic_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationGoodEvent (n :=…","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-e98f2e413915","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1655,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationGoodEvent_orderStatistic_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n))","missing":[],"search":"explorationgoodevent_orderstatistic_probability banditrlproof.musicalchairs.explorationgoodevent_orderstatistic_probability theorem explorationgoodevent_orderstatistic_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw (by omega) nu (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurableSet_actualCandidate_eq","label":"measurableSet_actualCandidate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurableSet_actualCandidate_eq","description":"theorem measurableSet_actualCandidate_eq {n k T : ℕ} (i : Fin n) (S : Finset (Fin k)) : MeasurableSet {z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ) | localCandidateSet (explorationFeedback z.1 z.2 i) = S}","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-477fa9788023","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1656,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurableSet_actualCandidate_eq {n k T : ℕ} (i : Fin n) (S : Finset (Fin k)) : MeasurableSet {z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ) | localCandidateSet (explorationFeedback z.1 z.2 i) = S}","missing":[],"search":"measurableset_actualcandidate_eq banditrlproof.musicalchairs.measurableset_actualcandidate_eq theorem measurableset_actualcandidate_eq {n k t : ℕ} (i : fin n) (s : finset (fin k)) : measurableset {z : (fin t → fin n → fin k) × (fin t → fin k → ℝ) | localcandidateset (explorationfeedback z.1 z.2 i) = s} theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.CandidateConfig","label":"CandidateConfig","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.CandidateConfig","description":"def CandidateConfig (n k : ℕ)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-6d0cfb707b55","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1657,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def CandidateConfig (n k : ℕ)","missing":[],"search":"candidateconfig banditrlproof.musicalchairs.candidateconfig def candidateconfig (n k : ℕ) definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedConfig","label":"learnedConfig","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedConfig","description":"noncomputable def learnedConfig {n k T : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : CandidateConfig n k","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-e44bd39bb184","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1658,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def learnedConfig {n k T : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : CandidateConfig n k","missing":[],"search":"learnedconfig banditrlproof.musicalchairs.learnedconfig noncomputable def learnedconfig {n k t : ℕ} (hk : 1 < k) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) : candidateconfig n k definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_learnedConfig","label":"measurable_learnedConfig","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_learnedConfig","description":"theorem measurable_learnedConfig {n k T : ℕ} (hk : 1 < k) : Measurable (learnedConfig (n := n) (T := T) hk)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-8c547ff647cc","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1659,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_learnedConfig {n k T : ℕ} (hk : 1 < k) : Measurable (learnedConfig (n := n) (T := T) hk)","missing":[],"search":"measurable_learnedconfig banditrlproof.musicalchairs.measurable_learnedconfig theorem measurable_learnedconfig {n k t : ℕ} (hk : 1 < k) : measurable (learnedconfig (n := n) (t := t) hk) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.configDrawPathKernel","label":"configDrawPathKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.configDrawPathKernel","description":"noncomputable def configDrawPathKernel (n k L : ℕ) : Kernel (CandidateConfig n k) (Fin L → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-1bd66fe2dcf5","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1660,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:154"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def configDrawPathKernel (n k L : ℕ) : Kernel (CandidateConfig n k) (Fin L → Fin n → Fin k)","missing":[],"search":"configdrawpathkernel banditrlproof.musicalchairs.configdrawpathkernel noncomputable def configdrawpathkernel (n k l : ℕ) : kernel (candidateconfig n k) (fin l → fin n → fin k) definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel","label":"learnedDrawPathKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel","description":"noncomputable def learnedDrawPathKernel {n k T : ℕ} (hk : 1 < k) (L : ℕ) : Kernel ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) (Fin L → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-851cbb828006","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1661,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:164"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def learnedDrawPathKernel {n k T : ℕ} (hk : 1 < k) (L : ℕ) : Kernel ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) (Fin L → Fin n → Fin k)","missing":[],"search":"learneddrawpathkernel banditrlproof.musicalchairs.learneddrawpathkernel noncomputable def learneddrawpathkernel {n k t : ℕ} (hk : 1 < k) (l : ℕ) : kernel ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) (fin l → fin n → fin k) definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_apply","label":"learnedDrawPathKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel_apply","description":"theorem learnedDrawPathKernel_apply {n k T L : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : learnedDrawPathKernel hk L z = (FinitePMF.iid (jointDraw (learnedConfig hk z).val (learnedConfig hk z).property) L).toMeasure","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-9ed0ff37e535","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1662,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:173"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedDrawPathKernel_apply {n k T L : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : learnedDrawPathKernel hk L z = (FinitePMF.iid (jointDraw (learnedConfig hk z).val (learnedConfig hk z).property) L).toMeasure","missing":[],"search":"learneddrawpathkernel_apply banditrlproof.musicalchairs.learneddrawpathkernel_apply theorem learneddrawpathkernel_apply {n k t l : ℕ} (hk : 1 < k) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) : learneddrawpathkernel hk l z = (finitepmf.iid (jointdraw (learnedconfig hk z).val (learnedconfig hk z).property) l).tomeasure theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_pi","label":"learnedDrawPathKernel_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel_pi","description":"theorem learnedDrawPathKernel_pi {n k T L : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : learnedDrawPathKernel hk L z = Measure.pi (fun _ : Fin L => (jointDraw (learnedConfig hk z).val (learnedConfig hk z).property).toMeasure)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-35b438a787fb","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1663,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:178"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedDrawPathKernel_pi {n k T L : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : learnedDrawPathKernel hk L z = Measure.pi (fun _ : Fin L => (jointDraw (learnedConfig hk z).val (learnedConfig hk z).property).toMeasure)","missing":[],"search":"learneddrawpathkernel_pi banditrlproof.musicalchairs.learneddrawpathkernel_pi theorem learneddrawpathkernel_pi {n k t l : ℕ} (hk : 1 < k) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) : learneddrawpathkernel hk l z = measure.pi (fun _ : fin l => (jointdraw (learnedconfig hk z).val (learnedconfig hk z).property).tomeasure) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_time_independent","label":"learnedDrawPathKernel_time_independent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel_time_independent","description":"theorem learnedDrawPathKernel_time_independent {n k T L : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : iIndepFun (fun t (draws : Fin L → Fin n → Fin k) => draws t) (learnedDrawPathKernel hk L z)","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-b5357abc1c6a","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1664,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:184"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedDrawPathKernel_time_independent {n k T L : ℕ} (hk : 1 < k) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : iIndepFun (fun t (draws : Fin L → Fin n → Fin k) => draws t) (learnedDrawPathKernel hk L z)","missing":[],"search":"learneddrawpathkernel_time_independent banditrlproof.musicalchairs.learneddrawpathkernel_time_independent theorem learneddrawpathkernel_time_independent {n k t l : ℕ} (hk : 1 < k) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) : iindepfun (fun t (draws : fin l → fin n → fin k) => draws t) (learneddrawpathkernel hk l z) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.commonConfig","label":"commonConfig","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.commonConfig","description":"noncomputable def commonConfig {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) : CandidateConfig n k","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-0e2f6c65eceb","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1665,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:195"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def commonConfig {n k : ℕ} (S : Finset (Fin k)) (hne : S.Nonempty) : CandidateConfig n k","missing":[],"search":"commonconfig banditrlproof.musicalchairs.commonconfig noncomputable def commonconfig {n k : ℕ} (s : finset (fin k)) (hne : s.nonempty) : candidateconfig n k definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedConfig_eq_common_on_good","label":"learnedConfig_eq_common_on_good","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedConfig_eq_common_on_good","description":"theorem learnedConfig_eq_common_on_good {n k T : ℕ} (hk : 1 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) (hz : z ∈ explorationGoodEvent (n := n) S) : learnedConfig hk z = commonConfig S hne","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-df70e688155b","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1666,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:198"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedConfig_eq_common_on_good {n k T : ℕ} (hk : 1 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) (hz : z ∈ explorationGoodEvent (n := n) S) : learnedConfig hk z = commonConfig S hne","missing":[],"search":"learnedconfig_eq_common_on_good banditrlproof.musicalchairs.learnedconfig_eq_common_on_good theorem learnedconfig_eq_common_on_good {n k t : ℕ} (hk : 1 < k) (s : finset (fin k)) (hne : s.nonempty) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) (hz : z ∈ explorationgoodevent (n := n) s) : learnedconfig hk z = commonconfig s hne theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_on_good","label":"learnedDrawPathKernel_on_good","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel_on_good","description":"theorem learnedDrawPathKernel_on_good {n k T L : ℕ} (hk : 1 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) (hz : z ∈ explorationGoodEvent (n := n) S) : learnedDrawPathKernel hk L z = (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L).toMeasure","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-01fe223cce79","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1667,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:206"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedDrawPathKernel_on_good {n k T L : ℕ} (hk : 1 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) (hz : z ∈ explorationGoodEvent (n := n) S) : learnedDrawPathKernel hk L z = (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L).toMeasure","missing":[],"search":"learneddrawpathkernel_on_good banditrlproof.musicalchairs.learneddrawpathkernel_on_good theorem learneddrawpathkernel_on_good {n k t l : ℕ} (hk : 1 < k) (s : finset (fin k)) (hne : s.nonempty) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) (hz : z ∈ explorationgoodevent (n := n) s) : learneddrawpathkernel hk l z = (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l).tomeasure theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw","label":"explorationContinuationLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw","description":"noncomputable def explorationContinuationLaw {n k T : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) (L : ℕ) : Measure (((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) × (Fin L → Fin n → Fin k))","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-118a5e7ae4e8","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1668,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:215"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationContinuationLaw {n k T : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) (L : ℕ) : Measure (((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) × (Fin L → Fin n → Fin k))","missing":[],"search":"explorationcontinuationlaw banditrlproof.musicalchairs.explorationcontinuationlaw noncomputable def explorationcontinuationlaw {n k t : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) (l : ℕ) : measure (((fin t → fin n → fin k) × (fin t → fin k → ℝ)) × (fin l → fin n → fin k)) definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_fst_event","label":"explorationContinuationLaw_fst_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_fst_event","description":"theorem explorationContinuationLaw_fst_event {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) : explorationContinuationLaw hk nu L (Prod.fst ⁻¹' E) = explorationRewardLaw (by omega) nu E","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-9c54f136e23a","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1669,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:231"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_fst_event {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) : explorationContinuationLaw hk nu L (Prod.fst ⁻¹' E) = explorationRewardLaw (by omega) nu E","missing":[],"search":"explorationcontinuationlaw_fst_event banditrlproof.musicalchairs.explorationcontinuationlaw_fst_event theorem explorationcontinuationlaw_fst_event {n k t l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (e : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ))) (he : measurableset e) : explorationcontinuationlaw hk nu l (prod.fst ⁻¹' e) = explorationrewardlaw (by omega) nu e theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_rectangle_on_good","label":"explorationContinuationLaw_rectangle_on_good","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_rectangle_on_good","description":"theorem explorationContinuationLaw_rectangle_on_good {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) (B : Set (Fin L → Fin n → Fin k)) : explorationContinuationLaw hk nu L (E ×ˢ B) = explorationRewardLaw (by omega)…","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-30e1d1e8b657","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1670,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:240"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_rectangle_on_good {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) (B : Set (Fin L → Fin n → Fin k)) : explorationContinuationLaw hk nu L (E ×ˢ B) = explorationRewardLaw (by omega) nu E * (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L).toMeasure B","missing":[],"search":"explorationcontinuationlaw_rectangle_on_good banditrlproof.musicalchairs.explorationcontinuationlaw_rectangle_on_good theorem explorationcontinuationlaw_rectangle_on_good {n k t l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (s : finset (fin k)) (hne : s.nonempty) (e : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) s) (b : set (fin l → fin n → fin k)) : explorationcontinuationlaw hk nu l (e ×ˢ b) = explorationrewardlaw (by omega) nu e * (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l).tomeasure b theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_good_probability","label":"explorationContinuationLaw_good_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_good_probability","description":"theorem explorationContinuationLaw_good_probability {n k L : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationContinuationLaw (by omega) nu L (Prod.fst ⁻¹' (explora…","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-11cbb593ce8e","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1671,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:256"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_good_probability {n k L : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationContinuationLaw (by omega) nu L (Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)))","missing":[],"search":"explorationcontinuationlaw_good_probability banditrlproof.musicalchairs.explorationcontinuationlaw_good_probability theorem explorationcontinuationlaw_good_probability {n k l : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationcontinuationlaw (by omega) nu l (prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n))) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trajectory_eq_of_prefix","label":"trajectory_eq_of_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trajectory_eq_of_prefix","description":"theorem trajectory_eq_of_prefix {n k : ℕ} (d e : ℕ → Fin n → Fin k) (t : ℕ) (h : ∀ u < t, d u = e u) : trajectory d t = trajectory e t","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-68c04c243f3f","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1672,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:268"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem trajectory_eq_of_prefix {n k : ℕ} (d e : ℕ → Fin n → Fin k) (t : ℕ) (h : ∀ u < t, d u = e u) : trajectory d t = trajectory e t","missing":[],"search":"trajectory_eq_of_prefix banditrlproof.musicalchairs.trajectory_eq_of_prefix theorem trajectory_eq_of_prefix {n k : ℕ} (d e : ℕ → fin n → fin k) (t : ℕ) (h : ∀ u < t, d u = e u) : trajectory d t = trajectory e t theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.extendedCoordinationDraws","label":"extendedCoordinationDraws","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.extendedCoordinationDraws","description":"noncomputable def extendedCoordinationDraws {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (t : ℕ) : Fin n → Fin k","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-03855c19106b","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1673,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:275"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def extendedCoordinationDraws {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (t : ℕ) : Fin n → Fin k","missing":[],"search":"extendedcoordinationdraws banditrlproof.musicalchairs.extendedcoordinationdraws noncomputable def extendedcoordinationdraws {n k l : ℕ} (hk : 0 < k) (d : fin l → fin n → fin k) (t : ℕ) : fin n → fin k definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationAction","label":"continuationAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationAction","description":"noncomputable def continuationAction {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (t : Fin L) : Fin n → Fin k","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-230157e63ed6","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1674,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:279"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def continuationAction {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (t : Fin L) : Fin n → Fin k","missing":[],"search":"continuationaction banditrlproof.musicalchairs.continuationaction noncomputable def continuationaction {n k l : ℕ} (hk : 0 < k) (d : fin l → fin n → fin k) (t : fin l) : fin n → fin k definition compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationAction_prefix","label":"continuationAction_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationAction_prefix","description":"theorem continuationAction_prefix {n k L : ℕ} (hk : 0 < k) (d e : Fin L → Fin n → Fin k) (t : Fin L) (h : ∀ u : Fin L, u ≤ t → d u = e u) : continuationAction hk d t = continuationAction hk e t","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-1bc8648ca5b0","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1675,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:283"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem continuationAction_prefix {n k L : ℕ} (hk : 0 < k) (d e : Fin L → Fin n → Fin k) (t : Fin L) (h : ∀ u : Fin L, u ≤ t → d u = e u) : continuationAction hk d t = continuationAction hk e t","missing":[],"search":"continuationaction_prefix banditrlproof.musicalchairs.continuationaction_prefix theorem continuationaction_prefix {n k l : ℕ} (hk : 0 < k) (d e : fin l → fin n → fin k) (t : fin l) (h : ∀ u : fin l, u ≤ t → d u = e u) : continuationaction hk d t = continuationaction hk e t theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationAction_fixed_stays","label":"continuationAction_fixed_stays","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationAction_fixed_stays","description":"theorem continuationAction_fixed_stays {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (t : Fin L) (i : Fin n) (a : Fin k) (hs : trajectory (extendedCoordinationDraws hk d) t.val i = some a) : continuationAction hk d t i = a","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-227ea764aa62","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1676,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:295"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem continuationAction_fixed_stays {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (t : Fin L) (i : Fin n) (a : Fin k) (hs : trajectory (extendedCoordinationDraws hk d) t.val i = some a) : continuationAction hk d t i = a","missing":[],"search":"continuationaction_fixed_stays banditrlproof.musicalchairs.continuationaction_fixed_stays theorem continuationaction_fixed_stays {n k l : ℕ} (hk : 0 < k) (d : fin l → fin n → fin k) (t : fin l) (i : fin n) (a : fin k) (hs : trajectory (extendedcoordinationdraws hk d) t.val i = some a) : continuationaction hk d t i = a theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_positive","label":"explorationGoodEvent_positive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationGoodEvent_positive","description":"theorem explorationGoodEvent_positive {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 0 < explorationRewardLaw (by omega) nu (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (…","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-52bc1d31c922","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1677,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:305"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationGoodEvent_positive {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 0 < explorationRewardLaw (by omega) nu (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n))","missing":[],"search":"explorationgoodevent_positive banditrlproof.musicalchairs.explorationgoodevent_positive theorem explorationgoodevent_positive {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 0 < explorationrewardlaw (by omega) nu (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)) theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_normalized_rectangle","label":"explorationContinuationLaw_normalized_rectangle","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_normalized_rectangle","description":"theorem explorationContinuationLaw_normalized_rectangle {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) (hpos : 0 < explorationRewardLaw (by omega) nu E) (B : Set (Fin L → Fin n → Fin k)) : explorationContinuationLa…","url":"../modules/banditrlproof-algorithms-musicalchairshandoff/index.html#decl-7306f256cfaa","parent":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","order":1678,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsHandoff"],["Source","BanditRLProof/Algorithms/MusicalChairsHandoff.lean:316"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_normalized_rectangle {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) (hpos : 0 < explorationRewardLaw (by omega) nu E) (B : Set (Fin L → Fin n → Fin k)) : explorationContinuationLaw hk nu L (E ×ˢ B) / explorationContinuationLaw hk nu L (Prod.fst ⁻¹' E) = (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L).toMeasure B","missing":[],"search":"explorationcontinuationlaw_normalized_rectangle banditrlproof.musicalchairs.explorationcontinuationlaw_normalized_rectangle theorem explorationcontinuationlaw_normalized_rectangle {n k t l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (s : finset (fin k)) (hne : s.nonempty) (e : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) s) (hpos : 0 < explorationrewardlaw (by omega) nu e) (b : set (fin l → fin n → fin k)) : explorationcontinuationlaw hk nu l (e ×ˢ b) / explorationcontinuationlaw hk nu l (prod.fst ⁻¹' e) = (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l).tomeasure b theorem compiled","shard":"modules/b0ed41d04a4a0fad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.ordered_set_sum_max","label":"ordered_set_sum_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.ordered_set_sum_max","description":"theorem ordered_set_sum_max {α : Type*} [DecidableEq α] (score : α → ℝ) (S B : Finset α) (hcard : B.card ≤ S.card) (hnonneg : ∀ a, 0 ≤ score a) (horder : ∀ a ∈ S, ∀ b ∉ S, score b ≤ score a) : ∑ b ∈ B, score b ≤ ∑ a ∈ S, score a","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-181114c04585","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1679,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem ordered_set_sum_max {α : Type*} [DecidableEq α] (score : α → ℝ) (S B : Finset α) (hcard : B.card ≤ S.card) (hnonneg : ∀ a, 0 ≤ score a) (horder : ∀ a ∈ S, ∀ b ∉ S, score b ≤ score a) : ∑ b ∈ B, score b ≤ ∑ a ∈ S, score a","missing":[],"search":"ordered_set_sum_max banditrlproof.musicalchairs.ordered_set_sum_max theorem ordered_set_sum_max {α : type*} [decidableeq α] (score : α → ℝ) (s b : finset α) (hcard : b.card ≤ s.card) (hnonneg : ∀ a, 0 ≤ score a) (horder : ∀ a ∈ s, ∀ b ∉ s, score b ≤ score a) : ∑ b ∈ b, score b ≤ ∑ a ∈ s, score a theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms_sum_max","label":"topArms_sum_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms_sum_max","description":"theorem topArms_sum_max {k n : ℕ} (score : Fin k → ℝ) (hnk : n ≤ k) (hnonneg : ∀ a, 0 ≤ score a) (B : Finset (Fin k)) (hcard : B.card ≤ n) : ∑ b ∈ B, score b ≤ ∑ a ∈ topArms score n, score a","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-794ea4170740","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1680,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem topArms_sum_max {k n : ℕ} (score : Fin k → ℝ) (hnk : n ≤ k) (hnonneg : ∀ a, 0 ≤ score a) (B : Finset (Fin k)) (hcard : B.card ≤ n) : ∑ b ∈ B, score b ≤ ∑ a ∈ topArms score n, score a","missing":[],"search":"toparms_sum_max banditrlproof.musicalchairs.toparms_sum_max theorem toparms_sum_max {k n : ℕ} (score : fin k → ℝ) (hnk : n ≤ k) (hnonneg : ∀ a, 0 ≤ score a) (b : finset (fin k)) (hcard : b.card ≤ n) : ∑ b ∈ b, score b ≤ ∑ a ∈ toparms score n, score a theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionFreePlayers","label":"collisionFreePlayers","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionFreePlayers","description":"noncomputable def collisionFreePlayers {n k : ℕ} (a : Fin n → Fin k) : Finset (Fin n)","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-dfad933e72c9","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1681,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def collisionFreePlayers {n k : ℕ} (a : Fin n → Fin k) : Finset (Fin n)","missing":[],"search":"collisionfreeplayers banditrlproof.musicalchairs.collisionfreeplayers noncomputable def collisionfreeplayers {n k : ℕ} (a : fin n → fin k) : finset (fin n) definition compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.earnedArms","label":"earnedArms","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.earnedArms","description":"noncomputable def earnedArms {n k : ℕ} (a : Fin n → Fin k) : Finset (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-951b3611db15","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1682,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def earnedArms {n k : ℕ} (a : Fin n → Fin k) : Finset (Fin k)","missing":[],"search":"earnedarms banditrlproof.musicalchairs.earnedarms noncomputable def earnedarms {n k : ℕ} (a : fin n → fin k) : finset (fin k) definition compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionFree_action_injective","label":"collisionFree_action_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionFree_action_injective","description":"theorem collisionFree_action_injective {n k : ℕ} (a : Fin n → Fin k) : Set.InjOn a (collisionFreePlayers a)","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-c508dfc761af","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1683,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collisionFree_action_injective {n k : ℕ} (a : Fin n → Fin k) : Set.InjOn a (collisionFreePlayers a)","missing":[],"search":"collisionfree_action_injective banditrlproof.musicalchairs.collisionfree_action_injective theorem collisionfree_action_injective {n k : ℕ} (a : fin n → fin k) : set.injon a (collisionfreeplayers a) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.earnedArms_card_le","label":"earnedArms_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.earnedArms_card_le","description":"theorem earnedArms_card_le {n k : ℕ} (a : Fin n → Fin k) : (earnedArms a).card ≤ n","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-162c50524ee8","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1684,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem earnedArms_card_le {n k : ℕ} (a : Fin n → Fin k) : (earnedArms a).card ≤ n","missing":[],"search":"earnedarms_card_le banditrlproof.musicalchairs.earnedarms_card_le theorem earnedarms_card_le {n k : ℕ} (a : fin n → fin k) : (earnedarms a).card ≤ n theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.earnedArms_sum","label":"earnedArms_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.earnedArms_sum","description":"theorem earnedArms_sum {n k : ℕ} (mu : Fin k → ℝ) (a : Fin n → Fin k) : ∑ b ∈ earnedArms a, mu b = ∑ i, if CollisionFree a i then mu (a i) else 0","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-cd63ff918d95","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1685,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem earnedArms_sum {n k : ℕ} (mu : Fin k → ℝ) (a : Fin n → Fin k) : ∑ b ∈ earnedArms a, mu b = ∑ i, if CollisionFree a i then mu (a i) else 0","missing":[],"search":"earnedarms_sum banditrlproof.musicalchairs.earnedarms_sum theorem earnedarms_sum {n k : ℕ} (mu : fin k → ℝ) (a : fin n → fin k) : ∑ b ∈ earnedarms a, mu b = ∑ i, if collisionfree a i then mu (a i) else 0 theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.globalRoundRegret","label":"globalRoundRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.globalRoundRegret","description":"noncomputable def globalRoundRegret {n k : ℕ} (mu : Fin k → ℝ) (a : Fin n → Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-d8202d2008d5","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1686,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def globalRoundRegret {n k : ℕ} (mu : Fin k → ℝ) (a : Fin n → Fin k) : ℝ","missing":[],"search":"globalroundregret banditrlproof.musicalchairs.globalroundregret noncomputable def globalroundregret {n k : ℕ} (mu : fin k → ℝ) (a : fin n → fin k) : ℝ definition compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.globalRoundRegret_nonneg","label":"globalRoundRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.globalRoundRegret_nonneg","description":"theorem globalRoundRegret_nonneg {n k : ℕ} (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ b, 0 ≤ mu b) (a : Fin n → Fin k) : 0 ≤ globalRoundRegret mu a","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-755554eff240","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1687,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem globalRoundRegret_nonneg {n k : ℕ} (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ b, 0 ≤ mu b) (a : Fin n → Fin k) : 0 ≤ globalRoundRegret mu a","missing":[],"search":"globalroundregret_nonneg banditrlproof.musicalchairs.globalroundregret_nonneg theorem globalroundregret_nonneg {n k : ℕ} (mu : fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ b, 0 ≤ mu b) (a : fin n → fin k) : 0 ≤ globalroundregret mu a theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.globalRoundRegret_le","label":"globalRoundRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.globalRoundRegret_le","description":"theorem globalRoundRegret_le {n k : ℕ} (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ b, 0 ≤ mu b ∧ mu b ≤ 1) (a : Fin n → Fin k) : globalRoundRegret mu a ≤ n","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-d2ff66c2842e","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1688,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem globalRoundRegret_le {n k : ℕ} (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ b, 0 ≤ mu b ∧ mu b ≤ 1) (a : Fin n → Fin k) : globalRoundRegret mu a ≤ n","missing":[],"search":"globalroundregret_le banditrlproof.musicalchairs.globalroundregret_le theorem globalroundregret_le {n k : ℕ} (mu : fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ b, 0 ≤ mu b ∧ mu b ≤ 1) (a : fin n → fin k) : globalroundregret mu a ≤ n theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_global","label":"roundPseudoRegret_global","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.roundPseudoRegret_global","description":"theorem roundPseudoRegret_global {n k : ℕ} (mu : Fin k → ℝ) (s : State n k) (d : Fin n → Fin k) : roundPseudoRegret (topArms mu n) mu s d = globalRoundRegret mu (action s d)","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-89119dcd97cb","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1689,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem roundPseudoRegret_global {n k : ℕ} (mu : Fin k → ℝ) (s : State n k) (d : Fin n → Fin k) : roundPseudoRegret (topArms mu n) mu s d = globalRoundRegret mu (action s d)","missing":[],"search":"roundpseudoregret_global banditrlproof.musicalchairs.roundpseudoregret_global theorem roundpseudoregret_global {n k : ℕ} (mu : fin k → ℝ) (s : state n k) (d : fin n → fin k) : roundpseudoregret (toparms mu n) mu s d = globalroundregret mu (action s d) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerAction","label":"learnerAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerAction","description":"noncomputable def learnerAction {n k S L : ℕ} (hk : 0 < k) (x : Fin S → Fin n → Fin k) (d : Fin L → Fin n → Fin k) (t : ℕ) : Fin n → Fin k","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-901b4cad4679","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1690,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def learnerAction {n k S L : ℕ} (hk : 0 < k) (x : Fin S → Fin n → Fin k) (d : Fin L → Fin n → Fin k) (t : ℕ) : Fin n → Fin k","missing":[],"search":"learneraction banditrlproof.musicalchairs.learneraction noncomputable def learneraction {n k s l : ℕ} (hk : 0 < k) (x : fin s → fin n → fin k) (d : fin l → fin n → fin k) (t : ℕ) : fin n → fin k definition compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerAction_explore","label":"learnerAction_explore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerAction_explore","description":"theorem learnerAction_explore {n k S L : ℕ} (hk : 0 < k) (x : Fin S → Fin n → Fin k) (d : Fin L → Fin n → Fin k) (t : ℕ) (ht : t < S) : learnerAction hk x d t = x ⟨t,ht⟩","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-ac29d2c5c1af","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1691,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerAction_explore {n k S L : ℕ} (hk : 0 < k) (x : Fin S → Fin n → Fin k) (d : Fin L → Fin n → Fin k) (t : ℕ) (ht : t < S) : learnerAction hk x d t = x ⟨t,ht⟩","missing":[],"search":"learneraction_explore banditrlproof.musicalchairs.learneraction_explore theorem learneraction_explore {n k s l : ℕ} (hk : 0 < k) (x : fin s → fin n → fin k) (d : fin l → fin n → fin k) (t : ℕ) (ht : t < s) : learneraction hk x d t = x ⟨t,ht⟩ theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerAction_coordinate","label":"learnerAction_coordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerAction_coordinate","description":"theorem learnerAction_coordinate {n k S L : ℕ} (hk : 0 < k) (x : Fin S → Fin n → Fin k) (d : Fin L → Fin n → Fin k) (t : Fin L) : learnerAction hk x d (S+t.val) = continuationAction hk d t","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-e76fe8760c8a","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1692,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerAction_coordinate {n k S L : ℕ} (hk : 0 < k) (x : Fin S → Fin n → Fin k) (d : Fin L → Fin n → Fin k) (t : Fin L) : learnerAction hk x d (S+t.val) = continuationAction hk d t","missing":[],"search":"learneraction_coordinate banditrlproof.musicalchairs.learneraction_coordinate theorem learneraction_coordinate {n k s l : ℕ} (hk : 0 < k) (x : fin s → fin n → fin k) (d : fin l → fin n → fin k) (t : fin l) : learneraction hk x d (s+t.val) = continuationaction hk d t theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationPrefixRegret","label":"explorationPrefixRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationPrefixRegret","description":"noncomputable def explorationPrefixRegret {n k S : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (H : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-5e91e19c17ae","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1693,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:126"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationPrefixRegret {n k S : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (H : ℕ) : ℝ","missing":[],"search":"explorationprefixregret banditrlproof.musicalchairs.explorationprefixregret noncomputable def explorationprefixregret {n k s : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (x : fin s → fin n → fin k) (h : ℕ) : ℝ definition compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerRegret","label":"learnerRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerRegret","description":"noncomputable def learnerRegret {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-aa96b2722832","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1694,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def learnerRegret {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : ℝ","missing":[],"search":"learnerregret banditrlproof.musicalchairs.learnerregret noncomputable def learnerregret {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (x : fin s → fin n → fin k) (d : fin (h-s) → fin n → fin k) : ℝ definition compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerRegret_split","label":"learnerRegret_split","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerRegret_split","description":"theorem learnerRegret_split {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : learnerRegret hk mu x d = explorationPrefixRegret hk mu x H + pathCoordinationRegret hk (topArms mu n) mu d","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-5bb2da3a77fd","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1695,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:134"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerRegret_split {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : learnerRegret hk mu x d = explorationPrefixRegret hk mu x H + pathCoordinationRegret hk (topArms mu n) mu d","missing":[],"search":"learnerregret_split banditrlproof.musicalchairs.learnerregret_split theorem learnerregret_split {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (x : fin s → fin n → fin k) (d : fin (h-s) → fin n → fin k) : learnerregret hk mu x d = explorationprefixregret hk mu x h + pathcoordinationregret hk (toparms mu n) mu d theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationPrefixRegret_bounds","label":"explorationPrefixRegret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationPrefixRegret_bounds","description":"theorem explorationPrefixRegret_bounds {n k S : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : Fin S → Fin n → Fin k) (H : ℕ) : 0 ≤ explorationPrefixRegret hk mu x H ∧ explorationPrefixRegret hk mu x H ≤ (n : ℝ) * min H S","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-e1d110460d0e","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1696,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:172"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationPrefixRegret_bounds {n k S : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : Fin S → Fin n → Fin k) (H : ℕ) : 0 ≤ explorationPrefixRegret hk mu x H ∧ explorationPrefixRegret hk mu x H ≤ (n : ℝ) * min H S","missing":[],"search":"explorationprefixregret_bounds banditrlproof.musicalchairs.explorationprefixregret_bounds theorem explorationprefixregret_bounds {n k s : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : fin s → fin n → fin k) (h : ℕ) : 0 ≤ explorationprefixregret hk mu x h ∧ explorationprefixregret hk mu x h ≤ (n : ℝ) * min h s theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerRegret_bounds","label":"learnerRegret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerRegret_bounds","description":"theorem learnerRegret_bounds {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : 0 ≤ learnerRegret hk mu x d ∧ learnerRegret hk mu x d ≤ (n : ℝ) * H","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-da5391babc9d","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1697,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:184"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerRegret_bounds {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : 0 ≤ learnerRegret hk mu x d ∧ learnerRegret hk mu x d ≤ (n : ℝ) * H","missing":[],"search":"learnerregret_bounds banditrlproof.musicalchairs.learnerregret_bounds theorem learnerregret_bounds {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : fin s → fin n → fin k) (d : fin (h-s) → fin n → fin k) : 0 ≤ learnerregret hk mu x d ∧ learnerregret hk mu x d ≤ (n : ℝ) * h theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerRegret_le_exploration_add","label":"learnerRegret_le_exploration_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerRegret_le_exploration_add","description":"theorem learnerRegret_le_exploration_add {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : learnerRegret hk mu x d ≤ (n : ℝ)*min H S + pathCoordinationRegret hk (topArms mu n) mu d","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-5a60c2819bee","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1698,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:195"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerRegret_le_exploration_add {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : Fin S → Fin n → Fin k) (d : Fin (H-S) → Fin n → Fin k) : learnerRegret hk mu x d ≤ (n : ℝ)*min H S + pathCoordinationRegret hk (topArms mu n) mu d","missing":[],"search":"learnerregret_le_exploration_add banditrlproof.musicalchairs.learnerregret_le_exploration_add theorem learnerregret_le_exploration_add {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (hnk : n ≤ k) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (x : fin s → fin n → fin k) (d : fin (h-s) → fin n → fin k) : learnerregret hk mu x d ≤ (n : ℝ)*min h s + pathcoordinationregret hk (toparms mu n) mu d theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerAction_prefix","label":"learnerAction_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerAction_prefix","description":"theorem learnerAction_prefix {n k S L : ℕ} (hk : 0 < k) (x y : Fin S → Fin n → Fin k) (d e : Fin L → Fin n → Fin k) (t : ℕ) (ht : t < S+L) (hx : ∀ u : Fin S, u.val ≤ t → x u = y u) (hd : ∀ u : Fin L, S+u.val ≤ t → d u = e u) : learnerAction hk x d t = learnerAction hk y e t","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-54e8ffc10835","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1699,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:207"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerAction_prefix {n k S L : ℕ} (hk : 0 < k) (x y : Fin S → Fin n → Fin k) (d e : Fin L → Fin n → Fin k) (t : ℕ) (ht : t < S+L) (hx : ∀ u : Fin S, u.val ≤ t → x u = y u) (hd : ∀ u : Fin L, S+u.val ≤ t → d u = e u) : learnerAction hk x d t = learnerAction hk y e t","missing":[],"search":"learneraction_prefix banditrlproof.musicalchairs.learneraction_prefix theorem learneraction_prefix {n k s l : ℕ} (hk : 0 < k) (x y : fin s → fin n → fin k) (d e : fin l → fin n → fin k) (t : ℕ) (ht : t < s+l) (hx : ∀ u : fin s, u.val ≤ t → x u = y u) (hd : ∀ u : fin l, s+u.val ≤ t → d u = e u) : learneraction hk x d t = learneraction hk y e t theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_learnerRegret","label":"integrable_learnerRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_learnerRegret","description":"theorem integrable_learnerRegret {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (P : Measure (((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ)) × (Fin (H-S) → Fin n → Fin k))) [IsFiniteMeasure P] : Integrable (fun z => learnerRegret hk mu z.1.1 z.2) P","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-1f75c2b4e22f","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1700,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:225"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_learnerRegret {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (P : Measure (((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ)) × (Fin (H-S) → Fin n → Fin k))) [IsFiniteMeasure P] : Integrable (fun z => learnerRegret hk mu z.1.1 z.2) P","missing":[],"search":"integrable_learnerregret banditrlproof.musicalchairs.integrable_learnerregret theorem integrable_learnerregret {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (p : measure (((fin s → fin n → fin k) × (fin s → fin k → ℝ)) × (fin (h-s) → fin n → fin k))) [isfinitemeasure p] : integrable (fun z => learnerregret hk mu z.1.1 z.2) p theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_pathRegret_projection","label":"integrable_pathRegret_projection","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_pathRegret_projection","description":"theorem integrable_pathRegret_projection {n k S L : ℕ} (hk : 0 < k) (A : Finset (Fin k)) (mu : Fin k → ℝ) (P : Measure (((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ)) × (Fin L → Fin n → Fin k))) [IsFiniteMeasure P] : Integrable (fun z => pathCoordinationRegret hk A mu z.2) P","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-7d506b556aa2","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1701,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:234"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_pathRegret_projection {n k S L : ℕ} (hk : 0 < k) (A : Finset (Fin k)) (mu : Fin k → ℝ) (P : Measure (((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ)) × (Fin L → Fin n → Fin k))) [IsFiniteMeasure P] : Integrable (fun z => pathCoordinationRegret hk A mu z.2) P","missing":[],"search":"integrable_pathregret_projection banditrlproof.musicalchairs.integrable_pathregret_projection theorem integrable_pathregret_projection {n k s l : ℕ} (hk : 0 < k) (a : finset (fin k)) (mu : fin k → ℝ) (p : measure (((fin s → fin n → fin k) × (fin s → fin k → ℝ)) × (fin l → fin n → fin k))) [isfinitemeasure p] : integrable (fun z => pathcoordinationregret hk a mu z.2) p theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerRegret_good_integral_le","label":"learnerRegret_good_integral_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerRegret_good_integral_le","description":"theorem learnerRegret_good_integral_le {n k S H : ℕ} (hk : 1 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (hne : (topArms mu n).Nonempty) (E : Set ((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) (topArms mu n)) : (∫ z in Prod.fst ⁻¹' E, learnerRegret (n := n) (by omega)…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-6897cf605baa","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1702,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:242"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerRegret_good_integral_le {n k S H : ℕ} (hk : 1 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (hne : (topArms mu n).Nonempty) (E : Set ((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) (topArms mu n)) : (∫ z in Prod.fst ⁻¹' E, learnerRegret (n := n) (by omega) mu z.1.1 z.2 ∂explorationContinuationLaw hk nu (H-S)) ≤ (explorationRewardLaw (by omega) nu E).toReal * ((n : ℝ)*min H S + 8*(n : ℝ)^2)","missing":[],"search":"learnerregret_good_integral_le banditrlproof.musicalchairs.learnerregret_good_integral_le theorem learnerregret_good_integral_le {n k s h : ℕ} (hk : 1 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (hne : (toparms mu n).nonempty) (e : set ((fin s → fin n → fin k) × (fin s → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) (toparms mu n)) : (∫ z in prod.fst ⁻¹' e, learnerregret (n := n) (by omega) mu z.1.1 z.2 ∂explorationcontinuationlaw hk nu (h-s)) ≤ (explorationrewardlaw (by omega) nu e).toreal * ((n : ℝ)*min h s + 8*(n : ℝ)^2) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerRegret_expected_le","label":"learnerRegret_expected_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerRegret_expected_le","description":"theorem learnerRegret_expected_le {n k S H : ℕ} (hk : 1 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (hne : (topArms mu n).Nonempty) (E : Set ((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) (topArms mu n)) (delta : ℝ) (hdelta : 0 ≤ delta) (hprob : 1-ENNReal.ofReal delta…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-0145cce35904","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1703,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:280"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerRegret_expected_le {n k S H : ℕ} (hk : 1 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (hne : (topArms mu n).Nonempty) (E : Set ((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) (topArms mu n)) (delta : ℝ) (hdelta : 0 ≤ delta) (hprob : 1-ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu E) : (∫ z, learnerRegret (n := n) (by omega) mu z.1.1 z.2 ∂explorationContinuationLaw hk nu (H-S)) ≤ min ((n : ℝ)*H) ((n : ℝ)*min H S + 8*(n : ℝ)^2 + delta*((n : ℝ)*H))","missing":[],"search":"learnerregret_expected_le banditrlproof.musicalchairs.learnerregret_expected_le theorem learnerregret_expected_le {n k s h : ℕ} (hk : 1 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (hne : (toparms mu n).nonempty) (e : set ((fin s → fin n → fin k) × (fin s → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) (toparms mu n)) (delta : ℝ) (hdelta : 0 ≤ delta) (hprob : 1-ennreal.ofreal delta ≤ explorationrewardlaw (by omega) nu e) : (∫ z, learnerregret (n := n) (by omega) mu z.1.1 z.2 ∂explorationcontinuationlaw hk nu (h-s)) ≤ min ((n : ℝ)*h) ((n : ℝ)*min h s + 8*(n : ℝ)^2 + delta*((n : ℝ)*h)) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_le","label":"source_expected_learnerRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_expected_learnerRegret_le","description":"theorem source_expected_learnerRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps delta) (H := H) (by omega) (armMean nu) z.1.…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-d04cb997d21b","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1704,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:328"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_expected_learnerRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps delta) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta)) ≤ min ((n : ℝ)*H) ((n : ℝ)*min H (explorationLength k eps delta) + 8*(n : ℝ)^2 + delta*((n : ℝ)*H))","missing":[],"search":"source_expected_learnerregret_le banditrlproof.musicalchairs.source_expected_learnerregret_le theorem source_expected_learnerregret_le {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerregret (n := n) (s := explorationlength k eps delta) (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta)) ≤ min ((n : ℝ)*h) ((n : ℝ)*min h (explorationlength k eps delta) + 8*(n : ℝ)^2 + delta*((n : ℝ)*h)) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_conditional_learnerRegret_le","label":"source_conditional_learnerRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_conditional_learnerRegret_le","description":"theorem source_conditional_learnerRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z in Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArm…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-88dde1b9d6b0","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1705,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:348"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_conditional_learnerRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z in Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)), learnerRegret (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta)) / (explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta) (Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)))).toReal ≤ (n : ℝ)*min H (explorationLength k eps delta) + 8*(n : ℝ)^2","missing":[],"search":"source_conditional_learnerregret_le banditrlproof.musicalchairs.source_conditional_learnerregret_le theorem source_conditional_learnerregret_le {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z in prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)), learnerregret (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta)) / (explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta) (prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)))).toreal ≤ (n : ℝ)*min h (explorationlength k eps delta) + 8*(n : ℝ)^2 theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_coarse","label":"source_expected_learnerRegret_coarse","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_expected_learnerRegret_coarse","description":"theorem source_expected_learnerRegret_coarse {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps delta) (H := H) (by omega) (armMean nu)…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-58866c6e3a1f","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1706,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:377"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_expected_learnerRegret_coarse {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps delta) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta)) ≤ min ((n : ℝ)*H) ((n : ℝ)*explorationLength k eps delta + 8*(n : ℝ)^2 + delta*((n : ℝ)*H))","missing":[],"search":"source_expected_learnerregret_coarse banditrlproof.musicalchairs.source_expected_learnerregret_coarse theorem source_expected_learnerregret_coarse {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerregret (n := n) (s := explorationlength k eps delta) (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta)) ≤ min ((n : ℝ)*h) ((n : ℝ)*explorationlength k eps delta + 8*(n : ℝ)^2 + delta*((n : ℝ)*h)) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.coordination_residual_source_constant","label":"coordination_residual_source_constant","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.coordination_residual_source_constant","description":"theorem coordination_residual_source_constant (n : ℕ) : 8*(n : ℝ)^2 ≤ 2*Real.exp 2*(n : ℝ)^2","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-d91bd0401efb","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1707,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:394"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem coordination_residual_source_constant (n : ℕ) : 8*(n : ℝ)^2 ≤ 2*Real.exp 2*(n : ℝ)^2","missing":[],"search":"coordination_residual_source_constant banditrlproof.musicalchairs.coordination_residual_source_constant theorem coordination_residual_source_constant (n : ℕ) : 8*(n : ℝ)^2 ≤ 2*real.exp 2*(n : ℝ)^2 theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_published_residual","label":"source_expected_learnerRegret_published_residual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_expected_learnerRegret_published_residual","description":"theorem source_expected_learnerRegret_published_residual {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps delta) (H := H) (by omega) (…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-4f7681b1711f","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1708,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:403"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_expected_learnerRegret_published_residual {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps delta) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta)) ≤ min ((n : ℝ)*H) ((n : ℝ)*explorationLength k eps delta + 2*Real.exp 2*(n : ℝ)^2 + delta*((n : ℝ)*H))","missing":[],"search":"source_expected_learnerregret_published_residual banditrlproof.musicalchairs.source_expected_learnerregret_published_residual theorem source_expected_learnerregret_published_residual {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z, learnerregret (n := n) (s := explorationlength k eps delta) (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta)) ≤ min ((n : ℝ)*h) ((n : ℝ)*explorationlength k eps delta + 2*real.exp 2*(n : ℝ)^2 + delta*((n : ℝ)*h)) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_inverse_horizon","label":"source_expected_learnerRegret_inverse_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_expected_learnerRegret_inverse_horizon","description":"theorem source_expected_learnerRegret_inverse_horizon {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (hH : 2 ≤ H) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps (1/(H : ℝ))) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂exploratio…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-eabe4d53d2af","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1709,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:418"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_expected_learnerRegret_inverse_horizon {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (hH : 2 ≤ H) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) : (∫ z, learnerRegret (n := n) (S := explorationLength k eps (1/(H : ℝ))) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (by omega) nu (H-explorationLength k eps (1/(H : ℝ)))) ≤ min ((n : ℝ)*H) ((n : ℝ)*explorationLength k eps (1/(H : ℝ)) + 8*(n : ℝ)^2 + n)","missing":[],"search":"source_expected_learnerregret_inverse_horizon banditrlproof.musicalchairs.source_expected_learnerregret_inverse_horizon theorem source_expected_learnerregret_inverse_horizon {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (hh : 2 ≤ h) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) : (∫ z, learnerregret (n := n) (s := explorationlength k eps (1/(h : ℝ))) (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (by omega) nu (h-explorationlength k eps (1/(h : ℝ)))) ≤ min ((n : ℝ)*h) ((n : ℝ)*explorationlength k eps (1/(h : ℝ)) + 8*(n : ℝ)^2 + n) theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_conditional_learnerRegret_published_residual","label":"source_conditional_learnerRegret_published_residual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_conditional_learnerRegret_published_residual","description":"theorem source_conditional_learnerRegret_published_residual {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) (hSH : explorationLength k eps delta ≤ H) : (∫ z in Prod.fst ⁻¹' (explorationGoodEvent…","url":"../modules/banditrlproof-algorithms-musicalchairslearnerregret/index.html#decl-d4408ff8ab69","parent":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","order":1710,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsLearnerRegret"],["Source","BanditRLProof/Algorithms/MusicalChairsLearnerRegret.lean:443"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_conditional_learnerRegret_published_residual {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) (hSH : explorationLength k eps delta ≤ H) : (∫ z in Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)), learnerRegret (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta)) / (explorationContinuationLaw (by omega) nu (H-explorationLength k eps delta) (Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)))).toReal ≤ (n : ℝ)*explorationLength k eps delta + 2*Real.exp 2*(n : ℝ)^2","missing":[],"search":"source_conditional_learnerregret_published_residual banditrlproof.musicalchairs.source_conditional_learnerregret_published_residual theorem source_conditional_learnerregret_published_residual {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) (hsh : explorationlength k eps delta ≤ h) : (∫ z in prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)), learnerregret (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta)) / (explorationcontinuationlaw (by omega) nu (h-explorationlength k eps delta) (prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)))).toreal ≤ (n : ℝ)*explorationlength k eps delta + 2*real.exp 2*(n : ℝ)^2 theorem compiled","shard":"modules/2e493f6267affb53.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_sum_snoc","label":"iid_sum_snoc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_sum_snoc","description":"theorem iid_sum_snoc {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (f : (Fin (T+1) → α) → ℝ≥0∞) : ∑ d, iid p (T+1) d * f d = ∑ x, iid p T x * ∑ a, p a * f (Fin.snoc x a)","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-5db396b5c4f7","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1711,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_sum_snoc {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) (f : (Fin (T+1) → α) → ℝ≥0∞) : ∑ d, iid p (T+1) d * f d = ∑ x, iid p T x * ∑ a, p a * f (Fin.snoc x a)","missing":[],"search":"iid_sum_snoc banditrlproof.finitepmf.iid_sum_snoc theorem iid_sum_snoc {α : type*} [fintype α] (p : pmf α) (t : ℕ) (f : (fin (t+1) → α) → ℝ≥0∞) : ∑ d, iid p (t+1) d * f d = ∑ x, iid p t x * ∑ a, p a * f (fin.snoc x a) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_snoc","label":"iid_snoc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_snoc","description":"theorem iid_snoc {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) : iid p (T+1) = (iid p T).bind (fun x => p.map (Fin.snoc x))","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-3db125111e25","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1712,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_snoc {α : Type*} [Fintype α] (p : PMF α) (T : ℕ) : iid p (T+1) = (iid p T).bind (fun x => p.map (Fin.snoc x))","missing":[],"search":"iid_snoc banditrlproof.finitepmf.iid_snoc theorem iid_snoc {α : type*} [fintype α] (p : pmf α) (t : ℕ) : iid p (t+1) = (iid p t).bind (fun x => p.map (fin.snoc x)) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trajectory_snoc_prefix","label":"trajectory_snoc_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trajectory_snoc_prefix","description":"theorem trajectory_snoc_prefix {n k T : ℕ} (hk : 0 < k) (d : Fin T → Fin n → Fin k) (a : Fin n → Fin k) : trajectory (extendedCoordinationDraws hk (Fin.snoc d a)) T = trajectory (extendedCoordinationDraws hk d) T","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-058e8edeef58","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1713,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem trajectory_snoc_prefix {n k T : ℕ} (hk : 0 < k) (d : Fin T → Fin n → Fin k) (a : Fin n → Fin k) : trajectory (extendedCoordinationDraws hk (Fin.snoc d a)) T = trajectory (extendedCoordinationDraws hk d) T","missing":[],"search":"trajectory_snoc_prefix banditrlproof.musicalchairs.trajectory_snoc_prefix theorem trajectory_snoc_prefix {n k t : ℕ} (hk : 0 < k) (d : fin t → fin n → fin k) (a : fin n → fin k) : trajectory (extendedcoordinationdraws hk (fin.snoc d a)) t = trajectory (extendedcoordinationdraws hk d) t theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trajectory_snoc_last","label":"trajectory_snoc_last","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trajectory_snoc_last","description":"theorem trajectory_snoc_last {n k T : ℕ} (hk : 0 < k) (d : Fin T → Fin n → Fin k) (a : Fin n → Fin k) : trajectory (extendedCoordinationDraws hk (Fin.snoc d a)) (T+1) = step (trajectory (extendedCoordinationDraws hk d) T) a","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-2890c640c940","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1714,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem trajectory_snoc_last {n k T : ℕ} (hk : 0 < k) (d : Fin T → Fin n → Fin k) (a : Fin n → Fin k) : trajectory (extendedCoordinationDraws hk (Fin.snoc d a)) (T+1) = step (trajectory (extendedCoordinationDraws hk d) T) a","missing":[],"search":"trajectory_snoc_last banditrlproof.musicalchairs.trajectory_snoc_last theorem trajectory_snoc_last {n k t : ℕ} (hk : 0 < k) (d : fin t → fin n → fin k) (a : fin n → fin k) : trajectory (extendedcoordinationdraws hk (fin.snoc d a)) (t+1) = step (trajectory (extendedcoordinationdraws hk d) t) a theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_trajectory_stateLaw","label":"iid_trajectory_stateLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_trajectory_stateLaw","description":"theorem iid_trajectory_stateLaw {n k : ℕ} (hk : 0 < k) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) (T : ℕ) : (FinitePMF.iid (jointDraw C hne) T).map (fun d => trajectory (extendedCoordinationDraws hk d) T) = stateLaw C hne T","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-cdc12e677bb5","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1715,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_trajectory_stateLaw {n k : ℕ} (hk : 0 < k) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) (T : ℕ) : (FinitePMF.iid (jointDraw C hne) T).map (fun d => trajectory (extendedCoordinationDraws hk d) T) = stateLaw C hne T","missing":[],"search":"iid_trajectory_statelaw banditrlproof.musicalchairs.iid_trajectory_statelaw theorem iid_trajectory_statelaw {n k : ℕ} (hk : 0 < k) (c : fin n → finset (fin k)) (hne : ∀ i, (c i).nonempty) (t : ℕ) : (finitepmf.iid (jointdraw c hne) t).map (fun d => trajectory (extendedcoordinationdraws hk d) t) = statelaw c hne t theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.iid_take","label":"iid_take","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.iid_take","description":"theorem iid_take {α : Type*} [Fintype α] (p : PMF α) {T L : ℕ} (h : T ≤ L) : (iid p L).map (fun d t => d (Fin.castLE h t)) = iid p T","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-9d6b960651b7","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1716,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_take {α : Type*} [Fintype α] (p : PMF α) {T L : ℕ} (h : T ≤ L) : (iid p L).map (fun d t => d (Fin.castLE h t)) = iid p T","missing":[],"search":"iid_take banditrlproof.finitepmf.iid_take theorem iid_take {α : type*} [fintype α] (p : pmf α) {t l : ℕ} (h : t ≤ l) : (iid p l).map (fun d t => d (fin.castle h t)) = iid p t theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.sum_map","label":"sum_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.sum_map","description":"theorem sum_map {α β : Type*} [Fintype α] [Fintype β] (p : PMF α) (g : α → β) (f : β → ℝ≥0∞) : ∑ b, p.map g b * f b = ∑ a, p a * f (g a)","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-4034167b7ffe","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1717,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem sum_map {α β : Type*} [Fintype α] [Fintype β] (p : PMF α) (g : α → β) (f : β → ℝ≥0∞) : ∑ b, p.map g b * f b = ∑ a, p a * f (g a)","missing":[],"search":"sum_map banditrlproof.finitepmf.sum_map theorem sum_map {α β : type*} [fintype α] [fintype β] (p : pmf α) (g : α → β) (f : β → ℝ≥0∞) : ∑ b, p.map g b * f b = ∑ a, p a * f (g a) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trajectory_take","label":"trajectory_take","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trajectory_take","description":"theorem trajectory_take {n k T L : ℕ} (hk : 0 < k) (h : T ≤ L) (d : Fin L → Fin n → Fin k) : trajectory (extendedCoordinationDraws hk d) T = trajectory (extendedCoordinationDraws hk (fun t => d (Fin.castLE h t))) T","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-fd6c9c4d0467","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1718,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:128"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem trajectory_take {n k T L : ℕ} (hk : 0 < k) (h : T ≤ L) (d : Fin L → Fin n → Fin k) : trajectory (extendedCoordinationDraws hk d) T = trajectory (extendedCoordinationDraws hk (fun t => d (Fin.castLE h t))) T","missing":[],"search":"trajectory_take banditrlproof.musicalchairs.trajectory_take theorem trajectory_take {n k t l : ℕ} (hk : 0 < k) (h : t ≤ l) (d : fin l → fin n → fin k) : trajectory (extendedcoordinationdraws hk d) t = trajectory (extendedcoordinationdraws hk (fun t => d (fin.castle h t))) t theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_trajectory_marginal","label":"iid_trajectory_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_trajectory_marginal","description":"theorem iid_trajectory_marginal {n k T L : ℕ} (hk : 0 < k) (h : T ≤ L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) L).map (fun d => trajectory (extendedCoordinationDraws hk d) T) = stateLaw C hne T","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-c5a5f8ce0494","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1719,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_trajectory_marginal {n k T L : ℕ} (hk : 0 < k) (h : T ≤ L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) L).map (fun d => trajectory (extendedCoordinationDraws hk d) T) = stateLaw C hne T","missing":[],"search":"iid_trajectory_marginal banditrlproof.musicalchairs.iid_trajectory_marginal theorem iid_trajectory_marginal {n k t l : ℕ} (hk : 0 < k) (h : t ≤ l) (c : fin n → finset (fin k)) (hne : ∀ i, (c i).nonempty) : (finitepmf.iid (jointdraw c hne) l).map (fun d => trajectory (extendedcoordinationdraws hk d) t) = statelaw c hne t theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_last_state_draw","label":"iid_last_state_draw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_last_state_draw","description":"theorem iid_last_state_draw {n k T : ℕ} (hk : 0 < k) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) (T+1)).map (fun d => (trajectory (extendedCoordinationDraws hk d) T, d (Fin.last T))) = (stateLaw C hne T).bind (fun s => (jointDraw C hne).map (fun a => (s,a)))","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-4d6cffa30be7","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1720,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:147"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_last_state_draw {n k T : ℕ} (hk : 0 < k) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) (T+1)).map (fun d => (trajectory (extendedCoordinationDraws hk d) T, d (Fin.last T))) = (stateLaw C hne T).bind (fun s => (jointDraw C hne).map (fun a => (s,a)))","missing":[],"search":"iid_last_state_draw banditrlproof.musicalchairs.iid_last_state_draw theorem iid_last_state_draw {n k t : ℕ} (hk : 0 < k) (c : fin n → finset (fin k)) (hne : ∀ i, (c i).nonempty) : (finitepmf.iid (jointdraw c hne) (t+1)).map (fun d => (trajectory (extendedcoordinationdraws hk d) t, d (fin.last t))) = (statelaw c hne t).bind (fun s => (jointdraw c hne).map (fun a => (s,a))) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_round_state_draw","label":"iid_round_state_draw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_round_state_draw","description":"theorem iid_round_state_draw {n k L : ℕ} (hk : 0 < k) (t : Fin L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) L).map (fun d => (trajectory (extendedCoordinationDraws hk d) t.val, d t)) = (stateLaw C hne t.val).bind (fun s => (jointDraw C hne).map (fun a => (s,a)))","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-f8fe95f7ed1e","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1721,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_round_state_draw {n k L : ℕ} (hk : 0 < k) (t : Fin L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) L).map (fun d => (trajectory (extendedCoordinationDraws hk d) t.val, d t)) = (stateLaw C hne t.val).bind (fun s => (jointDraw C hne).map (fun a => (s,a)))","missing":[],"search":"iid_round_state_draw banditrlproof.musicalchairs.iid_round_state_draw theorem iid_round_state_draw {n k l : ℕ} (hk : 0 < k) (t : fin l) (c : fin n → finset (fin k)) (hne : ∀ i, (c i).nonempty) : (finitepmf.iid (jointdraw c hne) l).map (fun d => (trajectory (extendedcoordinationdraws hk d) t.val, d t)) = (statelaw c hne t.val).bind (fun s => (jointdraw c hne).map (fun a => (s,a))) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.sum_map_real","label":"sum_map_real","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.sum_map_real","description":"theorem sum_map_real {α β : Type*} [Fintype α] [Fintype β] (p : PMF α) (g : α → β) (f : β → ℝ) : ∑ b, (p.map g b).toReal * f b = ∑ a, (p a).toReal * f (g a)","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-db680dfb8847","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1722,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:182"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem sum_map_real {α β : Type*} [Fintype α] [Fintype β] (p : PMF α) (g : α → β) (f : β → ℝ) : ∑ b, (p.map g b).toReal * f b = ∑ a, (p a).toReal * f (g a)","missing":[],"search":"sum_map_real banditrlproof.finitepmf.sum_map_real theorem sum_map_real {α β : type*} [fintype α] [fintype β] (p : pmf α) (g : α → β) (f : β → ℝ) : ∑ b, (p.map g b).toreal * f b = ∑ a, (p a).toreal * f (g a) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.FinitePMF.sum_bind_real","label":"sum_bind_real","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FinitePMF.sum_bind_real","description":"theorem sum_bind_real {α β : Type*} [Fintype α] [Fintype β] (p : PMF α) (q : α → PMF β) (f : β → ℝ) : ∑ b, (p.bind q b).toReal * f b = ∑ a, (p a).toReal * ∑ b, (q a b).toReal * f b","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-cbbccba349ae","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1723,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:198"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem sum_bind_real {α β : Type*} [Fintype α] [Fintype β] (p : PMF α) (q : α → PMF β) (f : β → ℝ) : ∑ b, (p.bind q b).toReal * f b = ∑ a, (p a).toReal * ∑ b, (q a b).toReal * f b","missing":[],"search":"sum_bind_real banditrlproof.finitepmf.sum_bind_real theorem sum_bind_real {α β : type*} [fintype α] [fintype β] (p : pmf α) (q : α → pmf β) (f : β → ℝ) : ∑ b, (p.bind q b).toreal * f b = ∑ a, (p a).toreal * ∑ b, (q a b).toreal * f b theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.pathCoordinationRegret","label":"pathCoordinationRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.pathCoordinationRegret","description":"noncomputable def pathCoordinationRegret {n k L : ℕ} (hk : 0 < k) (S : Finset (Fin k)) (mu : Fin k → ℝ) (d : Fin L → Fin n → Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-d4bde6f99df3","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1724,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:214"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def pathCoordinationRegret {n k L : ℕ} (hk : 0 < k) (S : Finset (Fin k)) (mu : Fin k → ℝ) (d : Fin L → Fin n → Fin k) : ℝ","missing":[],"search":"pathcoordinationregret banditrlproof.musicalchairs.pathcoordinationregret noncomputable def pathcoordinationregret {n k l : ℕ} (hk : 0 < k) (s : finset (fin k)) (mu : fin k → ℝ) (d : fin l → fin n → fin k) : ℝ definition compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_round_regret_expectation","label":"iid_round_regret_expectation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_round_regret_expectation","description":"theorem iid_round_regret_expectation {n k L : ℕ} (hk : 0 < k) (t : Fin L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) (S : Finset (Fin k)) (mu : Fin k → ℝ) : (∑ d, (FinitePMF.iid (jointDraw C hne) L d).toReal * roundPseudoRegret S mu (trajectory (extendedCoordinationDraws hk d) t.val) (d t)) = ∑ s, (stateLaw C hne t.val s).toReal * ∑ a, (jointDraw C hne a).toReal * roundPseudoRegret S mu s a","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-e6b8745628e6","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1725,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:219"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_round_regret_expectation {n k L : ℕ} (hk : 0 < k) (t : Fin L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) (S : Finset (Fin k)) (mu : Fin k → ℝ) : (∑ d, (FinitePMF.iid (jointDraw C hne) L d).toReal * roundPseudoRegret S mu (trajectory (extendedCoordinationDraws hk d) t.val) (d t)) = ∑ s, (stateLaw C hne t.val s).toReal * ∑ a, (jointDraw C hne a).toReal * roundPseudoRegret S mu s a","missing":[],"search":"iid_round_regret_expectation banditrlproof.musicalchairs.iid_round_regret_expectation theorem iid_round_regret_expectation {n k l : ℕ} (hk : 0 < k) (t : fin l) (c : fin n → finset (fin k)) (hne : ∀ i, (c i).nonempty) (s : finset (fin k)) (mu : fin k → ℝ) : (∑ d, (finitepmf.iid (jointdraw c hne) l d).toreal * roundpseudoregret s mu (trajectory (extendedcoordinationdraws hk d) t.val) (d t)) = ∑ s, (statelaw c hne t.val s).toreal * ∑ a, (jointdraw c hne a).toreal * roundpseudoregret s mu s a theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_path_regret_expectation","label":"iid_path_regret_expectation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_path_regret_expectation","description":"theorem iid_path_regret_expectation {n k L : ℕ} (hk : 0 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) : (∑ d, (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L d).toReal * pathCoordinationRegret hk S mu d) = ∑ t ∈ Finset.range L, ∑ s, (stateLaw (fun _ : Fin n => S) (fun _ => hne) t s).toReal * ∑ a, (jointDraw (fun _ : Fin n => S) (fun _ => hne) a).toReal * roundPseudoRegret S mu s a","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-6befb14093cf","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1726,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:232"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_path_regret_expectation {n k L : ℕ} (hk : 0 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) : (∑ d, (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L d).toReal * pathCoordinationRegret hk S mu d) = ∑ t ∈ Finset.range L, ∑ s, (stateLaw (fun _ : Fin n => S) (fun _ => hne) t s).toReal * ∑ a, (jointDraw (fun _ : Fin n => S) (fun _ => hne) a).toReal * roundPseudoRegret S mu s a","missing":[],"search":"iid_path_regret_expectation banditrlproof.musicalchairs.iid_path_regret_expectation theorem iid_path_regret_expectation {n k l : ℕ} (hk : 0 < k) (s : finset (fin k)) (hne : s.nonempty) (mu : fin k → ℝ) : (∑ d, (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l d).toreal * pathcoordinationregret hk s mu d) = ∑ t ∈ finset.range l, ∑ s, (statelaw (fun _ : fin n => s) (fun _ => hne) t s).toreal * ∑ a, (jointdraw (fun _ : fin n => s) (fun _ => hne) a).toreal * roundpseudoregret s mu s a theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_path_regret_le","label":"iid_path_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_path_regret_le","description":"theorem iid_path_regret_le {n k L : ℕ} (hk : 0 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) : (∑ d, (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L d).toReal * pathCoordinationRegret hk S mu d) ≤ 8 * (n : ℝ)^2","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-1c017fe23988","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1727,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:247"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_path_regret_le {n k L : ℕ} (hk : 0 < k) (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) : (∑ d, (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L d).toReal * pathCoordinationRegret hk S mu d) ≤ 8 * (n : ℝ)^2","missing":[],"search":"iid_path_regret_le banditrlproof.musicalchairs.iid_path_regret_le theorem iid_path_regret_le {n k l : ℕ} (hk : 0 < k) (s : finset (fin k)) (hne : s.nonempty) (hcard : s.card = n) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) : (∑ d, (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l d).toreal * pathcoordinationregret hk s mu d) ≤ 8 * (n : ℝ)^2 theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_state_marginal","label":"learnedDrawPathKernel_state_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel_state_marginal","description":"theorem learnedDrawPathKernel_state_marginal {n k T L t : ℕ} (hk : 1 < k) (ht : t ≤ L) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : (learnedDrawPathKernel hk L z).map (fun d => trajectory (extendedCoordinationDraws (by omega) d) t) = (stateLaw (learnedConfig hk z).val (learnedConfig hk z).property t).toMeasure","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-ee2a2953e4be","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1728,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:267"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedDrawPathKernel_state_marginal {n k T L t : ℕ} (hk : 1 < k) (ht : t ≤ L) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : (learnedDrawPathKernel hk L z).map (fun d => trajectory (extendedCoordinationDraws (by omega) d) t) = (stateLaw (learnedConfig hk z).val (learnedConfig hk z).property t).toMeasure","missing":[],"search":"learneddrawpathkernel_state_marginal banditrlproof.musicalchairs.learneddrawpathkernel_state_marginal theorem learneddrawpathkernel_state_marginal {n k t l t : ℕ} (hk : 1 < k) (ht : t ≤ l) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) : (learneddrawpathkernel hk l z).map (fun d => trajectory (extendedcoordinationdraws (by omega) d) t) = (statelaw (learnedconfig hk z).val (learnedconfig hk z).property t).tomeasure theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_round_marginal","label":"learnedDrawPathKernel_round_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedDrawPathKernel_round_marginal","description":"theorem learnedDrawPathKernel_round_marginal {n k T L : ℕ} (hk : 1 < k) (t : Fin L) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : (learnedDrawPathKernel hk L z).map (fun d => (trajectory (extendedCoordinationDraws (by omega) d) t.val, d t)) = ((stateLaw (learnedConfig hk z).val (learnedConfig hk z).property t.val).bind (fun s => (jointDraw (learnedConfig hk z).val (learnedConfig hk z).property).map (fun a =>…","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-83308203f990","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1729,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:275"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedDrawPathKernel_round_marginal {n k T L : ℕ} (hk : 1 < k) (t : Fin L) (z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ)) : (learnedDrawPathKernel hk L z).map (fun d => (trajectory (extendedCoordinationDraws (by omega) d) t.val, d t)) = ((stateLaw (learnedConfig hk z).val (learnedConfig hk z).property t.val).bind (fun s => (jointDraw (learnedConfig hk z).val (learnedConfig hk z).property).map (fun a => (s,a)))).toMeasure","missing":[],"search":"learneddrawpathkernel_round_marginal banditrlproof.musicalchairs.learneddrawpathkernel_round_marginal theorem learneddrawpathkernel_round_marginal {n k t l : ℕ} (hk : 1 < k) (t : fin l) (z : (fin t → fin n → fin k) × (fin t → fin k → ℝ)) : (learneddrawpathkernel hk l z).map (fun d => (trajectory (extendedcoordinationdraws (by omega) d) t.val, d t)) = ((statelaw (learnedconfig hk z).val (learnedconfig hk z).property t.val).bind (fun s => (jointdraw (learnedconfig hk z).val (learnedconfig hk z).property).map (fun a => (s,a)))).tomeasure theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_restricted_path","label":"explorationContinuationLaw_restricted_path","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_restricted_path","description":"theorem explorationContinuationLaw_restricted_path {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) : ((explorationContinuationLaw hk nu L).restrict (Prod.fst ⁻¹' E)).map Prod.snd = explorationRewardLaw (by omega) nu…","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-148b37d300f6","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1730,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:285"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_restricted_path {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) : ((explorationContinuationLaw hk nu L).restrict (Prod.fst ⁻¹' E)).map Prod.snd = explorationRewardLaw (by omega) nu E • (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L).toMeasure","missing":[],"search":"explorationcontinuationlaw_restricted_path banditrlproof.musicalchairs.explorationcontinuationlaw_restricted_path theorem explorationcontinuationlaw_restricted_path {n k t l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (s : finset (fin k)) (hne : s.nonempty) (e : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) s) : ((explorationcontinuationlaw hk nu l).restrict (prod.fst ⁻¹' e)).map prod.snd = explorationrewardlaw (by omega) nu e • (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l).tomeasure theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_good_regret_integral","label":"explorationContinuationLaw_good_regret_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_good_regret_integral","description":"theorem explorationContinuationLaw_good_regret_integral {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) : (∫ z in Prod.fst ⁻¹' E, pathCoordinationRegret (by omega) S mu z.2 ∂explorationContinuationL…","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-d108410d65a5","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1731,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:302"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_good_regret_integral {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (mu : Fin k → ℝ) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) : (∫ z in Prod.fst ⁻¹' E, pathCoordinationRegret (by omega) S mu z.2 ∂explorationContinuationLaw hk nu L) = (explorationRewardLaw (by omega) nu E).toReal * ∑ d, (FinitePMF.iid (jointDraw (fun _ : Fin n => S) (fun _ => hne)) L d).toReal * pathCoordinationRegret (by omega) S mu d","missing":[],"search":"explorationcontinuationlaw_good_regret_integral banditrlproof.musicalchairs.explorationcontinuationlaw_good_regret_integral theorem explorationcontinuationlaw_good_regret_integral {n k t l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (s : finset (fin k)) (hne : s.nonempty) (mu : fin k → ℝ) (e : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) s) : (∫ z in prod.fst ⁻¹' e, pathcoordinationregret (by omega) s mu z.2 ∂explorationcontinuationlaw hk nu l) = (explorationrewardlaw (by omega) nu e).toreal * ∑ d, (finitepmf.iid (jointdraw (fun _ : fin n => s) (fun _ => hne)) l d).toreal * pathcoordinationregret (by omega) s mu d theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_conditional_coordination_regret","label":"explorationContinuationLaw_conditional_coordination_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_conditional_coordination_regret","description":"theorem explorationContinuationLaw_conditional_coordination_regret {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) (hpos : 0 < explorationReward…","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-b80f5ab25456","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1732,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:318"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_conditional_coordination_regret {n k T L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (S : Finset (Fin k)) (hne : S.Nonempty) (hcard : S.card = n) (mu : Fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (E : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))) (hE : MeasurableSet E) (hEG : E ⊆ explorationGoodEvent (n := n) S) (hpos : 0 < explorationRewardLaw (by omega) nu E) : (∫ z in Prod.fst ⁻¹' E, pathCoordinationRegret (by omega) S mu z.2 ∂explorationContinuationLaw hk nu L) / (explorationContinuationLaw hk nu L (Prod.fst ⁻¹' E)).toReal ≤ 8 * (n : ℝ)^2","missing":[],"search":"explorationcontinuationlaw_conditional_coordination_regret banditrlproof.musicalchairs.explorationcontinuationlaw_conditional_coordination_regret theorem explorationcontinuationlaw_conditional_coordination_regret {n k t l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (s : finset (fin k)) (hne : s.nonempty) (hcard : s.card = n) (mu : fin k → ℝ) (hmu : ∀ a, 0 ≤ mu a ∧ mu a ≤ 1) (e : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ))) (he : measurableset e) (heg : e ⊆ explorationgoodevent (n := n) s) (hpos : 0 < explorationrewardlaw (by omega) nu e) : (∫ z in prod.fst ⁻¹' e, pathcoordinationregret (by omega) s mu z.2 ∂explorationcontinuationlaw hk nu l) / (explorationcontinuationlaw hk nu l (prod.fst ⁻¹' e)).toreal ≤ 8 * (n : ℝ)^2 theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.iid_continuationAction_marginal","label":"iid_continuationAction_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.iid_continuationAction_marginal","description":"theorem iid_continuationAction_marginal {n k L : ℕ} (hk : 0 < k) (t : Fin L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) L).map (fun d => continuationAction hk d t) = (stateLaw C hne t.val).bind (fun s => (jointDraw C hne).map (action s))","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-3b4fbe6164c2","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1733,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:347"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem iid_continuationAction_marginal {n k L : ℕ} (hk : 0 < k) (t : Fin L) (C : Fin n → Finset (Fin k)) (hne : ∀ i, (C i).Nonempty) : (FinitePMF.iid (jointDraw C hne) L).map (fun d => continuationAction hk d t) = (stateLaw C hne t.val).bind (fun s => (jointDraw C hne).map (action s))","missing":[],"search":"iid_continuationaction_marginal banditrlproof.musicalchairs.iid_continuationaction_marginal theorem iid_continuationaction_marginal {n k l : ℕ} (hk : 0 < k) (t : fin l) (c : fin n → finset (fin k)) (hne : ∀ i, (c i).nonempty) : (finitepmf.iid (jointdraw c hne) l).map (fun d => continuationaction hk d t) = (statelaw c hne t.val).bind (fun s => (jointdraw c hne).map (action s)) theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_conditional_coordination_regret","label":"source_conditional_coordination_regret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_conditional_coordination_regret","description":"theorem source_conditional_coordination_regret {n k L : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z in Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTop…","url":"../modules/banditrlproof-algorithms-musicalchairsmarginal/index.html#decl-a5485a4f0cbe","parent":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","order":1734,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsMarginal"],["Source","BanditRLProof/Algorithms/MusicalChairsMarginal.lean:358"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_conditional_coordination_regret {n k L : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z in Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)), pathCoordinationRegret (by omega) (trueTopArms nu n) (armMean nu) z.2 ∂explorationContinuationLaw (by omega) nu L) / (explorationContinuationLaw (by omega) nu L (Prod.fst ⁻¹' (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n)))).toReal ≤ 8 * (n : ℝ)^2","missing":[],"search":"source_conditional_coordination_regret banditrlproof.musicalchairs.source_conditional_coordination_regret theorem source_conditional_coordination_regret {n k l : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z in prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)), pathcoordinationregret (by omega) (truetoparms nu n) (armmean nu) z.2 ∂explorationcontinuationlaw (by omega) nu l) / (explorationcontinuationlaw (by omega) nu l (prod.fst ⁻¹' (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)))).toreal ≤ 8 * (n : ℝ)^2 theorem compiled","shard":"modules/1cef4aee8da8ddd2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.base_gamma_upper","label":"base_gamma_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.base_gamma_upper","description":"theorem base_gamma_upper {k : ℝ} (hk : 1 < k) : (1-1/k) ^ (2/5 : ℝ) ≤ 1 - (2/5 : ℝ)/k","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-e2a1f18d77c8","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1735,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem base_gamma_upper {k : ℝ} (hk : 1 < k) : (1-1/k) ^ (2/5 : ℝ) ≤ 1 - (2/5 : ℝ)/k","missing":[],"search":"base_gamma_upper banditrlproof.musicalchairs.base_gamma_upper theorem base_gamma_upper {k : ℝ} (hk : 1 < k) : (1-1/k) ^ (2/5 : ℝ) ≤ 1 - (2/5 : ℝ)/k theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.base_neg_gamma_lower","label":"base_neg_gamma_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.base_neg_gamma_lower","description":"theorem base_neg_gamma_lower {k : ℝ} (hk : 1 < k) : 1 + (2/5 : ℝ)/k ≤ (1-1/k) ^ (-(2/5 : ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-8346c863387a","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1736,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem base_neg_gamma_lower {k : ℝ} (hk : 1 < k) : 1 + (2/5 : ℝ)/k ≤ (1-1/k) ^ (-(2/5 : ℝ))","missing":[],"search":"base_neg_gamma_lower banditrlproof.musicalchairs.base_neg_gamma_lower theorem base_neg_gamma_lower {k : ℝ} (hk : 1 < k) : 1 + (2/5 : ℝ)/k ≤ (1-1/k) ^ (-(2/5 : ℝ)) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.avoidanceBase_eq","label":"avoidanceBase_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.avoidanceBase_eq","description":"theorem avoidanceBase_eq {k : ℕ} (hk : 0 < k) : (((k-1 : ℕ) : ℝ)/(k : ℝ)) = 1-1/(k : ℝ)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-bdc8c88c12f9","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1737,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem avoidanceBase_eq {k : ℕ} (hk : 0 < k) : (((k-1 : ℕ) : ℝ)/(k : ℝ)) = 1-1/(k : ℝ)","missing":[],"search":"avoidancebase_eq banditrlproof.musicalchairs.avoidancebase_eq theorem avoidancebase_eq {k : ℕ} (hk : 0 < k) : (((k-1 : ℕ) : ℝ)/(k : ℝ)) = 1-1/(k : ℝ) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.avoidance_mass_lower","label":"avoidance_mass_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.avoidance_mass_lower","description":"theorem avoidance_mass_lower {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) : (1/4 : ℝ) ≤ (1-1/(k : ℝ))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-dac652de2043","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1738,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem avoidance_mass_lower {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) : (1/4 : ℝ) ≤ (1-1/(k : ℝ))^(n-1)","missing":[],"search":"avoidance_mass_lower banditrlproof.musicalchairs.avoidance_mass_lower theorem avoidance_mass_lower {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) : (1/4 : ℝ) ≤ (1-1/(k : ℝ))^(n-1) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collision_accuracy_sandwich","label":"collision_accuracy_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collision_accuracy_sandwich","description":"theorem collision_accuracy_sandwich {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : (1-1/(k : ℝ))^(n-1) * (1-1/(k : ℝ))^(2/5 : ℝ) ≤ 1-pHat ∧ 1-pHat ≤ (1-1/(k : ℝ))^(n-1) * (1-1/(k : ℝ))^(-(2/5 : ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-4e83da5cc629","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1739,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collision_accuracy_sandwich {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : (1-1/(k : ℝ))^(n-1) * (1-1/(k : ℝ))^(2/5 : ℝ) ≤ 1-pHat ∧ 1-pHat ≤ (1-1/(k : ℝ))^(n-1) * (1-1/(k : ℝ))^(-(2/5 : ℝ))","missing":[],"search":"collision_accuracy_sandwich banditrlproof.musicalchairs.collision_accuracy_sandwich theorem collision_accuracy_sandwich {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) (phat : ℝ) (hacc : |phat-collisionprobreal n k| ≤ 1/(10*(k : ℝ))) : (1-1/(k : ℝ))^(n-1) * (1-1/(k : ℝ))^(2/5 : ℝ) ≤ 1-phat ∧ 1-phat ≤ (1-1/(k : ℝ))^(n-1) * (1-1/(k : ℝ))^(-(2/5 : ℝ)) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collision_accuracy_survival_pos","label":"collision_accuracy_survival_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collision_accuracy_survival_pos","description":"theorem collision_accuracy_survival_pos {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : 0 < 1-pHat","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-54921d4622e5","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1740,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collision_accuracy_survival_pos {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : 0 < 1-pHat","missing":[],"search":"collision_accuracy_survival_pos banditrlproof.musicalchairs.collision_accuracy_survival_pos theorem collision_accuracy_survival_pos {n k : ℕ} (hk : 1 < k) (hnk : n ≤ k) (phat : ℝ) (hacc : |phat-collisionprobreal n k| ≤ 1/(10*(k : ℝ))) : 0 < 1-phat theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.population_inverse_margin","label":"population_inverse_margin","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.population_inverse_margin","description":"theorem population_inverse_margin {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : (n : ℝ)-(2/5 : ℝ) ≤ 1+Real.log (1-pHat)/Real.log (1-1/(k : ℝ)) ∧ 1+Real.log (1-pHat)/Real.log (1-1/(k : ℝ)) ≤ (n : ℝ)+(2/5 : ℝ)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-ccd9e0ba80d7","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1741,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem population_inverse_margin {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : (n : ℝ)-(2/5 : ℝ) ≤ 1+Real.log (1-pHat)/Real.log (1-1/(k : ℝ)) ∧ 1+Real.log (1-pHat)/Real.log (1-1/(k : ℝ)) ≤ (n : ℝ)+(2/5 : ℝ)","missing":[],"search":"population_inverse_margin banditrlproof.musicalchairs.population_inverse_margin theorem population_inverse_margin {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (phat : ℝ) (hacc : |phat-collisionprobreal n k| ≤ 1/(10*(k : ℝ))) : (n : ℝ)-(2/5 : ℝ) ≤ 1+real.log (1-phat)/real.log (1-1/(k : ℝ)) ∧ 1+real.log (1-phat)/real.log (1-1/(k : ℝ)) ≤ (n : ℝ)+(2/5 : ℝ) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.population_inverse_round","label":"population_inverse_round","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.population_inverse_round","description":"theorem population_inverse_round {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : round (1+Real.log (1-pHat)/Real.log (1-1/(k : ℝ))) = (n : ℤ)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-51c9d224079a","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1742,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem population_inverse_round {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (pHat : ℝ) (hacc : |pHat-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : round (1+Real.log (1-pHat)/Real.log (1-1/(k : ℝ))) = (n : ℤ)","missing":[],"search":"population_inverse_round banditrlproof.musicalchairs.population_inverse_round theorem population_inverse_round {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (phat : ℝ) (hacc : |phat-collisionprobreal n k| ≤ 1/(10*(k : ℝ))) : round (1+real.log (1-phat)/real.log (1-1/(k : ℝ))) = (n : ℤ) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationInverse","label":"populationInverse","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationInverse","description":"noncomputable def populationInverse (k T C : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-979d78773aef","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1743,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:122"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def populationInverse (k T C : ℕ) : ℝ","missing":[],"search":"populationinverse banditrlproof.musicalchairs.populationinverse noncomputable def populationinverse (k t c : ℕ) : ℝ definition compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate","label":"populationEstimate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate","description":"noncomputable def populationEstimate (k T C : ℕ) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-840e7eac3fe3","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1744,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def populationEstimate (k T C : ℕ) : ℕ","missing":[],"search":"populationestimate banditrlproof.musicalchairs.populationestimate noncomputable def populationestimate (k t c : ℕ) : ℕ definition compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localPopulationEstimate","label":"localPopulationEstimate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localPopulationEstimate","description":"noncomputable def localPopulationEstimate {k T : ℕ} (f : Fin T → ExplorationFeedback k) : ℕ","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-c429fe0d01fc","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1745,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:128"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localPopulationEstimate {k T : ℕ} (f : Fin T → ExplorationFeedback k) : ℕ","missing":[],"search":"localpopulationestimate banditrlproof.musicalchairs.localpopulationestimate noncomputable def localpopulationestimate {k t : ℕ} (f : fin t → explorationfeedback k) : ℕ definition compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCollisionCount_le","label":"localCollisionCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCollisionCount_le","description":"theorem localCollisionCount_le {k T : ℕ} (f : Fin T → ExplorationFeedback k) : localCollisionCount f ≤ T","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-0a4888dce23c","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1746,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCollisionCount_le {k T : ℕ} (f : Fin T → ExplorationFeedback k) : localCollisionCount f ≤ T","missing":[],"search":"localcollisioncount_le banditrlproof.musicalchairs.localcollisioncount_le theorem localcollisioncount_le {k t : ℕ} (f : fin t → explorationfeedback k) : localcollisioncount f ≤ t theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.noncollision_fraction_eq","label":"noncollision_fraction_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.noncollision_fraction_eq","description":"theorem noncollision_fraction_eq {T C : ℕ} (hT : 0 < T) (hC : C ≤ T) : (((T-C : ℕ) : ℝ)/(T : ℝ)) = 1-(C : ℝ)/(T : ℝ)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-4488cd2ce135","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1747,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem noncollision_fraction_eq {T C : ℕ} (hT : 0 < T) (hC : C ≤ T) : (((T-C : ℕ) : ℝ)/(T : ℝ)) = 1-(C : ℝ)/(T : ℝ)","missing":[],"search":"noncollision_fraction_eq banditrlproof.musicalchairs.noncollision_fraction_eq theorem noncollision_fraction_eq {t c : ℕ} (ht : 0 < t) (hc : c ≤ t) : (((t-c : ℕ) : ℝ)/(t : ℝ)) = 1-(c : ℝ)/(t : ℝ) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationInverse_ge_one","label":"populationInverse_ge_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationInverse_ge_one","description":"theorem populationInverse_ge_one {k T C : ℕ} (hk : 1 < k) (hC : C < T) : 1 ≤ populationInverse k T C","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-522e95b0c09f","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1748,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationInverse_ge_one {k T C : ℕ} (hk : 1 < k) (hC : C < T) : 1 ≤ populationInverse k T C","missing":[],"search":"populationinverse_ge_one banditrlproof.musicalchairs.populationinverse_ge_one theorem populationinverse_ge_one {k t c : ℕ} (hk : 1 < k) (hc : c < t) : 1 ≤ populationinverse k t c theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate_le","label":"populationEstimate_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate_le","description":"theorem populationEstimate_le (k T C : ℕ) : populationEstimate k T C ≤ k","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-7bf9c718fa49","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1749,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:163"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationEstimate_le (k T C : ℕ) : populationEstimate k T C ≤ k","missing":[],"search":"populationestimate_le banditrlproof.musicalchairs.populationestimate_le theorem populationestimate_le (k t c : ℕ) : populationestimate k t c ≤ k theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate_all_collision","label":"populationEstimate_all_collision","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate_all_collision","description":"theorem populationEstimate_all_collision (k T : ℕ) : populationEstimate k T T = k","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-735c962d348b","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1750,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:169"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationEstimate_all_collision (k T : ℕ) : populationEstimate k T T = k","missing":[],"search":"populationestimate_all_collision banditrlproof.musicalchairs.populationestimate_all_collision theorem populationestimate_all_collision (k t : ℕ) : populationestimate k t t = k theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate_zero_collision","label":"populationEstimate_zero_collision","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate_zero_collision","description":"theorem populationEstimate_zero_collision {k T : ℕ} (hk : 0 < k) (hT : 0 < T) : populationEstimate k T 0 = 1","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-8e708a56df06","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1751,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:172"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationEstimate_zero_collision {k T : ℕ} (hk : 0 < k) (hT : 0 < T) : populationEstimate k T 0 = 1","missing":[],"search":"populationestimate_zero_collision banditrlproof.musicalchairs.populationestimate_zero_collision theorem populationestimate_zero_collision {k t : ℕ} (hk : 0 < k) (ht : 0 < t) : populationestimate k t 0 = 1 theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate_correct","label":"populationEstimate_correct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate_correct","description":"theorem populationEstimate_correct {n k T C : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (hC : C ≤ T) (hacc : |(C : ℝ)/T-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : populationEstimate k T C = n","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-ce61961c2f1f","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1752,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:178"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationEstimate_correct {n k T C : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (hC : C ≤ T) (hacc : |(C : ℝ)/T-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : populationEstimate k T C = n","missing":[],"search":"populationestimate_correct banditrlproof.musicalchairs.populationestimate_correct theorem populationestimate_correct {n k t c : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (ht : 0 < t) (hc : c ≤ t) (hacc : |(c : ℝ)/t-collisionprobreal n k| ≤ 1/(10*(k : ℝ))) : populationestimate k t c = n theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localPopulationEstimate_correct","label":"localPopulationEstimate_correct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localPopulationEstimate_correct","description":"theorem localPopulationEstimate_correct {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (f : Fin T → ExplorationFeedback k) (hacc : |(localCollisionCount f : ℝ)/T-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : localPopulationEstimate f = n","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-557e78e204f1","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1753,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:191"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localPopulationEstimate_correct {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (f : Fin T → ExplorationFeedback k) (hacc : |(localCollisionCount f : ℝ)/T-collisionProbReal n k| ≤ 1/(10*(k : ℝ))) : localPopulationEstimate f = n","missing":[],"search":"localpopulationestimate_correct banditrlproof.musicalchairs.localpopulationestimate_correct theorem localpopulationestimate_correct {n k t : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (ht : 0 < t) (f : fin t → explorationfeedback k) (hacc : |(localcollisioncount f : ℝ)/t-collisionprobreal n k| ≤ 1/(10*(k : ℝ))) : localpopulationestimate f = n theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate_regular_cast","label":"populationEstimate_regular_cast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate_regular_cast","description":"theorem populationEstimate_regular_cast {k T C : ℕ} (hk : 1 < k) (hC : C < T) : (populationEstimate k T C : ℤ) = min (k : ℤ) (round (populationInverse k T C))","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-e99392eac600","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1754,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:202"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationEstimate_regular_cast {k T C : ℕ} (hk : 1 < k) (hC : C < T) : (populationEstimate k T C : ℤ) = min (k : ℤ) (round (populationInverse k T C))","missing":[],"search":"populationestimate_regular_cast banditrlproof.musicalchairs.populationestimate_regular_cast theorem populationestimate_regular_cast {k t c : ℕ} (hk : 1 < k) (hc : c < t) : (populationestimate k t c : ℤ) = min (k : ℤ) (round (populationinverse k t c)) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.collisionCount_le","label":"collisionCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.collisionCount_le","description":"theorem collisionCount_le {n k T : ℕ} (x : Fin T → Fin n → Fin k) (i : Fin n) : collisionCount i x ≤ T","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-a821e28f1e9d","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1755,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:211"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem collisionCount_le {n k T : ℕ} (x : Fin T → Fin n → Fin k) (i : Fin n) : collisionCount i x ≤ T","missing":[],"search":"collisioncount_le banditrlproof.musicalchairs.collisioncount_le theorem collisioncount_le {n k t : ℕ} (x : fin t → fin n → fin k) (i : fin n) : collisioncount i x ≤ t theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localPopulationEstimate_eq","label":"localPopulationEstimate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localPopulationEstimate_eq","description":"theorem localPopulationEstimate_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) : localPopulationEstimate (explorationFeedback x r i) = populationEstimate k T (collisionCount i x)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-abe86afae73e","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1756,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:216"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localPopulationEstimate_eq {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) : localPopulationEstimate (explorationFeedback x r i) = populationEstimate k T (collisionCount i x)","missing":[],"search":"localpopulationestimate_eq banditrlproof.musicalchairs.localpopulationestimate_eq theorem localpopulationestimate_eq {n k t : ℕ} (x : fin t → fin n → fin k) (r : fin t → fin k → ℝ) (i : fin n) : localpopulationestimate (explorationfeedback x r i) = populationestimate k t (collisioncount i x) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationRecovered","label":"populationRecovered","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationRecovered","description":"def populationRecovered {n k T : ℕ} : Set (Fin T → Fin n → Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-abc4130a9c93","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1757,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:221"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def populationRecovered {n k T : ℕ} : Set (Fin T → Fin n → Fin k)","missing":[],"search":"populationrecovered banditrlproof.musicalchairs.populationrecovered def populationrecovered {n k t : ℕ} : set (fin t → fin n → fin k) definition compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate_subset_populationRecovered","label":"allCollisionAccurate_subset_populationRecovered","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allCollisionAccurate_subset_populationRecovered","description":"theorem allCollisionAccurate_subset_populationRecovered {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) : allCollisionAccurate (n := n) (k := k) (T := T) ⊆ populationRecovered","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-be63c1cb112c","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1758,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:224"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allCollisionAccurate_subset_populationRecovered {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) : allCollisionAccurate (n := n) (k := k) (T := T) ⊆ populationRecovered","missing":[],"search":"allcollisionaccurate_subset_populationrecovered banditrlproof.musicalchairs.allcollisionaccurate_subset_populationrecovered theorem allcollisionaccurate_subset_populationrecovered {n k t : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (ht : 0 < t) : allcollisionaccurate (n := n) (k := k) (t := t) ⊆ populationrecovered theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationRecovered_probability","label":"populationRecovered_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationRecovered_probability","description":"theorem populationRecovered_probability {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 1 - ENNReal.ofReal (delta/2) ≤ (explorationLaw n k T (by omega)).toMeasure (populationRecovered (n := n) (k := k) (T := T))","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-93e30d6fc6c8","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1759,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:230"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationRecovered_probability {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * Real.log (4*(k : ℝ)/delta) ≤ T) : 1 - ENNReal.ofReal (delta/2) ≤ (explorationLaw n k T (by omega)).toMeasure (populationRecovered (n := n) (k := k) (T := T))","missing":[],"search":"populationrecovered_probability banditrlproof.musicalchairs.populationrecovered_probability theorem populationrecovered_probability {n k t : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (ht : 0 < t) (delta : ℝ) (hdelta : 0 < delta) (hbudget : (50*(k : ℝ)^2) * real.log (4*(k : ℝ)/delta) ≤ t) : 1 - ennreal.ofreal (delta/2) ≤ (explorationlaw n k t (by omega)).tomeasure (populationrecovered (n := n) (k := k) (t := t)) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect","label":"explorationEstimatesCorrect","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationEstimatesCorrect","description":"def explorationEstimatesCorrect {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-b9f78c0dd5e4","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1760,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:238"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def explorationEstimatesCorrect {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"explorationestimatescorrect banditrlproof.musicalchairs.explorationestimatescorrect def explorationestimatescorrect {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect_eq_inter","label":"explorationEstimatesCorrect_eq_inter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationEstimatesCorrect_eq_inter","description":"theorem explorationEstimatesCorrect_eq_inter {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : explorationEstimatesCorrect (n := n) (T := T) nu eps = allMeanAccurate nu eps ∩ Prod.fst ⁻¹' populationRecovered","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-7392b5501df3","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1761,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:244"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationEstimatesCorrect_eq_inter {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : explorationEstimatesCorrect (n := n) (T := T) nu eps = allMeanAccurate nu eps ∩ Prod.fst ⁻¹' populationRecovered","missing":[],"search":"explorationestimatescorrect_eq_inter banditrlproof.musicalchairs.explorationestimatescorrect_eq_inter theorem explorationestimatescorrect_eq_inter {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : explorationestimatescorrect (n := n) (t := t) nu eps = allmeanaccurate nu eps ∩ prod.fst ⁻¹' populationrecovered theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurableSet_explorationEstimatesCorrect","label":"measurableSet_explorationEstimatesCorrect","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurableSet_explorationEstimatesCorrect","description":"theorem measurableSet_explorationEstimatesCorrect {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : MeasurableSet (explorationEstimatesCorrect (n := n) (T := T) nu eps)","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-fa4bb030ccb5","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1762,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:250"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurableSet_explorationEstimatesCorrect {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : MeasurableSet (explorationEstimatesCorrect (n := n) (T := T) nu eps)","missing":[],"search":"measurableset_explorationestimatescorrect banditrlproof.musicalchairs.measurableset_explorationestimatescorrect theorem measurableset_explorationestimatescorrect {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : measurableset (explorationestimatescorrect (n := n) (t := t) nu eps) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_subset_estimatesCorrect","label":"explorationStatisticsAccurate_subset_estimatesCorrect","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationStatisticsAccurate_subset_estimatesCorrect","description":"theorem explorationStatisticsAccurate_subset_estimatesCorrect {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (nu : Fin k → Measure ℝ) (eps : ℝ) : explorationStatisticsAccurate (n := n) (T := T) nu eps ⊆ explorationEstimatesCorrect nu eps","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-608a5c3fe386","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1763,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:258"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationStatisticsAccurate_subset_estimatesCorrect {n k T : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (hT : 0 < T) (nu : Fin k → Measure ℝ) (eps : ℝ) : explorationStatisticsAccurate (n := n) (T := T) nu eps ⊆ explorationEstimatesCorrect nu eps","missing":[],"search":"explorationstatisticsaccurate_subset_estimatescorrect banditrlproof.musicalchairs.explorationstatisticsaccurate_subset_estimatescorrect theorem explorationstatisticsaccurate_subset_estimatescorrect {n k t : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (ht : 0 < t) (nu : fin k → measure ℝ) (eps : ℝ) : explorationstatisticsaccurate (n := n) (t := t) nu eps ⊆ explorationestimatescorrect nu eps theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect_probability","label":"explorationEstimatesCorrect_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationEstimatesCorrect_probability","description":"theorem explorationEstimatesCorrect_probability {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationEstimatesCorrect (n := n) (T := explorationLe…","url":"../modules/banditrlproof-algorithms-musicalchairspopulation/index.html#decl-cec612d867af","parent":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","order":1764,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsPopulation"],["Source","BanditRLProof/Algorithms/MusicalChairsPopulation.lean:266"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationEstimatesCorrect_probability {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationEstimatesCorrect (n := n) (T := explorationLength k eps delta) nu eps)","missing":[],"search":"explorationestimatescorrect_probability banditrlproof.musicalchairs.explorationestimatescorrect_probability theorem explorationestimatescorrect_probability {n k : ℕ} (hn : 0 < n) (hk : 1 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw (by omega) nu (explorationestimatescorrect (n := n) (t := explorationlength k eps delta) nu eps) theorem compiled","shard":"modules/7c20d341d7664101.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.scoreOrder","label":"scoreOrder","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.scoreOrder","description":"def scoreOrder {k : ℕ} (score : Fin k → ℝ) (a b : Fin k) : Prop","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-3e6995200f70","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1765,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def scoreOrder {k : ℕ} (score : Fin k → ℝ) (a b : Fin k) : Prop","missing":[],"search":"scoreorder banditrlproof.musicalchairs.scoreorder def scoreorder {k : ℕ} (score : fin k → ℝ) (a b : fin k) : prop definition compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.rankedArms","label":"rankedArms","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.rankedArms","description":"noncomputable def rankedArms {k : ℕ} (score : Fin k → ℝ) : List (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-24361841baa8","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1766,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def rankedArms {k : ℕ} (score : Fin k → ℝ) : List (Fin k)","missing":[],"search":"rankedarms banditrlproof.musicalchairs.rankedarms noncomputable def rankedarms {k : ℕ} (score : fin k → ℝ) : list (fin k) definition compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms","label":"topArms","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms","description":"noncomputable def topArms {k : ℕ} (score : Fin k → ℝ) (n : ℕ) : Finset (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-95ff6e95cc4c","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1767,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def topArms {k : ℕ} (score : Fin k → ℝ) (n : ℕ) : Finset (Fin k)","missing":[],"search":"toparms banditrlproof.musicalchairs.toparms noncomputable def toparms {k : ℕ} (score : fin k → ℝ) (n : ℕ) : finset (fin k) definition compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms_card","label":"topArms_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms_card","description":"theorem topArms_card {k : ℕ} (score : Fin k → ℝ) (n : ℕ) : (topArms score n).card = min n k","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-5a1187dc516e","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1768,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem topArms_card {k : ℕ} (score : Fin k → ℝ) (n : ℕ) : (topArms score n).card = min n k","missing":[],"search":"toparms_card banditrlproof.musicalchairs.toparms_card theorem toparms_card {k : ℕ} (score : fin k → ℝ) (n : ℕ) : (toparms score n).card = min n k theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms_before_unselected","label":"topArms_before_unselected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms_before_unselected","description":"theorem topArms_before_unselected {k : ℕ} (score : Fin k → ℝ) (n : ℕ) {a b : Fin k} (ha : a ∈ topArms score n) (hb : b ∉ topArms score n) : scoreOrder score a b","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-c57a7d601af0","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1769,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem topArms_before_unselected {k : ℕ} (score : Fin k → ℝ) (n : ℕ) {a b : Fin k} (ha : a ∈ topArms score n) (hb : b ∉ topArms score n) : scoreOrder score a b","missing":[],"search":"toparms_before_unselected banditrlproof.musicalchairs.toparms_before_unselected theorem toparms_before_unselected {k : ℕ} (score : fin k → ℝ) (n : ℕ) {a b : fin k} (ha : a ∈ toparms score n) (hb : b ∉ toparms score n) : scoreorder score a b theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms_eq_of_order_separated","label":"topArms_eq_of_order_separated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms_eq_of_order_separated","description":"theorem topArms_eq_of_order_separated {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (hsep : ∀ a ∈ S, ∀ b ∉ S, scoreOrder score a b) : topArms score n = S","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-23e35cdce9b9","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1770,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem topArms_eq_of_order_separated {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (hsep : ∀ a ∈ S, ∀ b ∉ S, scoreOrder score a b) : topArms score n = S","missing":[],"search":"toparms_eq_of_order_separated banditrlproof.musicalchairs.toparms_eq_of_order_separated theorem toparms_eq_of_order_separated {k : ℕ} (score : fin k → ℝ) (n : ℕ) (s : finset (fin k)) (hcard : s.card = n) (hnk : n ≤ k) (hsep : ∀ a ∈ s, ∀ b ∉ s, scoreorder score a b) : toparms score n = s theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms_eq_of_strict_separation","label":"topArms_eq_of_strict_separation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms_eq_of_strict_separation","description":"theorem topArms_eq_of_strict_separation {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (hsep : ∀ a ∈ S, ∀ b ∉ S, score b < score a) : topArms score n = S","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-ad4ad1324f8d","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1771,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem topArms_eq_of_strict_separation {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (hsep : ∀ a ∈ S, ∀ b ∉ S, score b < score a) : topArms score n = S","missing":[],"search":"toparms_eq_of_strict_separation banditrlproof.musicalchairs.toparms_eq_of_strict_separation theorem toparms_eq_of_strict_separation {k : ℕ} (score : fin k → ℝ) (n : ℕ) (s : finset (fin k)) (hcard : s.card = n) (hnk : n ≤ k) (hsep : ∀ a ∈ s, ∀ b ∉ s, score b < score a) : toparms score n = s theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.topArms_eq_iff","label":"topArms_eq_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.topArms_eq_iff","description":"theorem topArms_eq_iff {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hnk : n ≤ k) (S : Finset (Fin k)) : topArms score n = S ↔ S.card = n ∧ ∀ a ∈ S, ∀ b ∉ S, scoreOrder score a b","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-afe8816a517f","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1772,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:100"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem topArms_eq_iff {k : ℕ} (score : Fin k → ℝ) (n : ℕ) (hnk : n ≤ k) (S : Finset (Fin k)) : topArms score n = S ↔ S.card = n ∧ ∀ a ∈ S, ∀ b ∉ S, scoreOrder score a b","missing":[],"search":"toparms_eq_iff banditrlproof.musicalchairs.toparms_eq_iff theorem toparms_eq_iff {k : ℕ} (score : fin k → ℝ) (n : ℕ) (hnk : n ≤ k) (s : finset (fin k)) : toparms score n = s ↔ s.card = n ∧ ∀ a ∈ s, ∀ b ∉ s, scoreorder score a b theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.populationEstimate_pos","label":"populationEstimate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.populationEstimate_pos","description":"theorem populationEstimate_pos {k T C : ℕ} (hk : 1 < k) (hC : C ≤ T) : 0 < populationEstimate k T C","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-b54864f8dd21","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1773,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem populationEstimate_pos {k T C : ℕ} (hk : 1 < k) (hC : C ≤ T) : 0 < populationEstimate k T C","missing":[],"search":"populationestimate_pos banditrlproof.musicalchairs.populationestimate_pos theorem populationestimate_pos {k t c : ℕ} (hk : 1 < k) (hc : c ≤ t) : 0 < populationestimate k t c theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCandidateSet","label":"localCandidateSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCandidateSet","description":"noncomputable def localCandidateSet {k T : ℕ} (f : Fin T → ExplorationFeedback k) : Finset (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-c1e953a9aa9c","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1774,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def localCandidateSet {k T : ℕ} (f : Fin T → ExplorationFeedback k) : Finset (Fin k)","missing":[],"search":"localcandidateset banditrlproof.musicalchairs.localcandidateset noncomputable def localcandidateset {k t : ℕ} (f : fin t → explorationfeedback k) : finset (fin k) definition compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_card","label":"localCandidateSet_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCandidateSet_card","description":"theorem localCandidateSet_card {k T : ℕ} (f : Fin T → ExplorationFeedback k) : (localCandidateSet f).card = localPopulationEstimate f","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-3c56565f163b","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1775,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCandidateSet_card {k T : ℕ} (f : Fin T → ExplorationFeedback k) : (localCandidateSet f).card = localPopulationEstimate f","missing":[],"search":"localcandidateset_card banditrlproof.musicalchairs.localcandidateset_card theorem localcandidateset_card {k t : ℕ} (f : fin t → explorationfeedback k) : (localcandidateset f).card = localpopulationestimate f theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_nonempty","label":"localCandidateSet_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCandidateSet_nonempty","description":"theorem localCandidateSet_nonempty {k T : ℕ} (hk : 1 < k) (f : Fin T → ExplorationFeedback k) : (localCandidateSet f).Nonempty","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-02ece134c789","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1776,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:135"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCandidateSet_nonempty {k T : ℕ} (hk : 1 < k) (f : Fin T → ExplorationFeedback k) : (localCandidateSet f).Nonempty","missing":[],"search":"localcandidateset_nonempty banditrlproof.musicalchairs.localcandidateset_nonempty theorem localcandidateset_nonempty {k t : ℕ} (hk : 1 < k) (f : fin t → explorationfeedback k) : (localcandidateset f).nonempty theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.empirical_strict_separation","label":"empirical_strict_separation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.empirical_strict_separation","description":"theorem empirical_strict_separation {k : ℕ} (mu score : Fin k → ℝ) (S : Finset (Fin k)) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ mu a - mu b) (hacc : ∀ a, |score a-mu a| < eps/2) : ∀ a ∈ S, ∀ b ∉ S, score b < score a","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-f6d666c1f9e8","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1777,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:140"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem empirical_strict_separation {k : ℕ} (mu score : Fin k → ℝ) (S : Finset (Fin k)) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ mu a - mu b) (hacc : ∀ a, |score a-mu a| < eps/2) : ∀ a ∈ S, ∀ b ∉ S, score b < score a","missing":[],"search":"empirical_strict_separation banditrlproof.musicalchairs.empirical_strict_separation theorem empirical_strict_separation {k : ℕ} (mu score : fin k → ℝ) (s : finset (fin k)) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ mu a - mu b) (hacc : ∀ a, |score a-mu a| < eps/2) : ∀ a ∈ s, ∀ b ∉ s, score b < score a theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_correct","label":"localCandidateSet_correct","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCandidateSet_correct","description":"theorem localCandidateSet_correct {n k T : ℕ} (f : Fin T → ExplorationFeedback k) (mu : Fin k → ℝ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ mu a - mu b) (hacc : ∀ a, |localEmpiricalMean f a-mu a| < eps/2) (hpop : localPopulationEstimate f = n) : localCandidateSet f = S","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-735ecd4c368f","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1778,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCandidateSet_correct {n k T : ℕ} (f : Fin T → ExplorationFeedback k) (mu : Fin k → ℝ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ mu a - mu b) (hacc : ∀ a, |localEmpiricalMean f a-mu a| < eps/2) (hpop : localPopulationEstimate f = n) : localCandidateSet f = S","missing":[],"search":"localcandidateset_correct banditrlproof.musicalchairs.localcandidateset_correct theorem localcandidateset_correct {n k t : ℕ} (f : fin t → explorationfeedback k) (mu : fin k → ℝ) (s : finset (fin k)) (hcard : s.card = n) (hnk : n ≤ k) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ mu a - mu b) (hacc : ∀ a, |localempiricalmean f a-mu a| < eps/2) (hpop : localpopulationestimate f = n) : localcandidateset f = s theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_eq_iff","label":"localCandidateSet_eq_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localCandidateSet_eq_iff","description":"theorem localCandidateSet_eq_iff {k T : ℕ} (f : Fin T → ExplorationFeedback k) (S : Finset (Fin k)) : localCandidateSet f = S ↔ S.card = localPopulationEstimate f ∧ ∀ a ∈ S, ∀ b ∉ S, scoreOrder (localEmpiricalMean f) a b","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-b65134c27131","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1779,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localCandidateSet_eq_iff {k T : ℕ} (f : Fin T → ExplorationFeedback k) (S : Finset (Fin k)) : localCandidateSet f = S ↔ S.card = localPopulationEstimate f ∧ ∀ a ∈ S, ∀ b ∉ S, scoreOrder (localEmpiricalMean f) a b","missing":[],"search":"localcandidateset_eq_iff banditrlproof.musicalchairs.localcandidateset_eq_iff theorem localcandidateset_eq_iff {k t : ℕ} (f : fin t → explorationfeedback k) (s : finset (fin k)) : localcandidateset f = s ↔ s.card = localpopulationestimate f ∧ ∀ a ∈ s, ∀ b ∉ s, scoreorder (localempiricalmean f) a b theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurableSet_localScoreOrder","label":"measurableSet_localScoreOrder","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurableSet_localScoreOrder","description":"theorem measurableSet_localScoreOrder {n k T : ℕ} (i : Fin n) (a b : Fin k) : MeasurableSet {z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ) | scoreOrder (localEmpiricalMean (explorationFeedback z.1 z.2 i)) a b}","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-a85eb539ddf3","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1780,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:171"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurableSet_localScoreOrder {n k T : ℕ} (i : Fin n) (a b : Fin k) : MeasurableSet {z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ) | scoreOrder (localEmpiricalMean (explorationFeedback z.1 z.2 i)) a b}","missing":[],"search":"measurableset_localscoreorder banditrlproof.musicalchairs.measurableset_localscoreorder theorem measurableset_localscoreorder {n k t : ℕ} (i : fin n) (a b : fin k) : measurableset {z : (fin t → fin n → fin k) × (fin t → fin k → ℝ) | scoreorder (localempiricalmean (explorationfeedback z.1 z.2 i)) a b} theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent","label":"explorationGoodEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationGoodEvent","description":"def explorationGoodEvent {n k T : ℕ} (S : Finset (Fin k)) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-20b8df8bf7cf","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1781,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:182"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def explorationGoodEvent {n k T : ℕ} (S : Finset (Fin k)) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"explorationgoodevent banditrlproof.musicalchairs.explorationgoodevent def explorationgoodevent {n k t : ℕ} (s : finset (fin k)) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_eq_comparisons","label":"explorationGoodEvent_eq_comparisons","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationGoodEvent_eq_comparisons","description":"theorem explorationGoodEvent_eq_comparisons {n k T : ℕ} (S : Finset (Fin k)) (hcard : S.card = n) : explorationGoodEvent (n := n) (T := T) S = (Prod.fst ⁻¹' populationRecovered) ∩ {z | ∀ i : Fin n, ∀ a ∈ S, ∀ b ∉ S, scoreOrder (localEmpiricalMean (explorationFeedback z.1 z.2 i)) a b}","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-6fca597d3226","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1782,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:187"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationGoodEvent_eq_comparisons {n k T : ℕ} (S : Finset (Fin k)) (hcard : S.card = n) : explorationGoodEvent (n := n) (T := T) S = (Prod.fst ⁻¹' populationRecovered) ∩ {z | ∀ i : Fin n, ∀ a ∈ S, ∀ b ∉ S, scoreOrder (localEmpiricalMean (explorationFeedback z.1 z.2 i)) a b}","missing":[],"search":"explorationgoodevent_eq_comparisons banditrlproof.musicalchairs.explorationgoodevent_eq_comparisons theorem explorationgoodevent_eq_comparisons {n k t : ℕ} (s : finset (fin k)) (hcard : s.card = n) : explorationgoodevent (n := n) (t := t) s = (prod.fst ⁻¹' populationrecovered) ∩ {z | ∀ i : fin n, ∀ a ∈ s, ∀ b ∉ s, scoreorder (localempiricalmean (explorationfeedback z.1 z.2 i)) a b} theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurableSet_explorationGoodEvent","label":"measurableSet_explorationGoodEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurableSet_explorationGoodEvent","description":"theorem measurableSet_explorationGoodEvent {n k T : ℕ} (S : Finset (Fin k)) (hcard : S.card = n) : MeasurableSet (explorationGoodEvent (n := n) (T := T) S)","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-a401d7a9182c","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1783,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:205"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurableSet_explorationGoodEvent {n k T : ℕ} (S : Finset (Fin k)) (hcard : S.card = n) : MeasurableSet (explorationGoodEvent (n := n) (T := T) S)","missing":[],"search":"measurableset_explorationgoodevent banditrlproof.musicalchairs.measurableset_explorationgoodevent theorem measurableset_explorationgoodevent {n k t : ℕ} (s : finset (fin k)) (hcard : s.card = n) : measurableset (explorationgoodevent (n := n) (t := t) s) theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect_subset_goodEvent","label":"explorationEstimatesCorrect_subset_goodEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationEstimatesCorrect_subset_goodEvent","description":"theorem explorationEstimatesCorrect_subset_goodEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) : explorationEstimatesCorrect (n := n) (T := T) nu eps ⊆ explorationGoodEvent S","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-7cfd06d6688d","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1784,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:214"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationEstimatesCorrect_subset_goodEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) : explorationEstimatesCorrect (n := n) (T := T) nu eps ⊆ explorationGoodEvent S","missing":[],"search":"explorationestimatescorrect_subset_goodevent banditrlproof.musicalchairs.explorationestimatescorrect_subset_goodevent theorem explorationestimatescorrect_subset_goodevent {n k t : ℕ} (nu : fin k → measure ℝ) (s : finset (fin k)) (hcard : s.card = n) (hnk : n ≤ k) (eps gap : ℝ) (hepsgap : eps < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ armmean nu a - armmean nu b) : explorationestimatescorrect (n := n) (t := t) nu eps ⊆ explorationgoodevent s theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trueTopArms","label":"trueTopArms","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trueTopArms","description":"noncomputable def trueTopArms {k : ℕ} (nu : Fin k → Measure ℝ) (n : ℕ) : Finset (Fin k)","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-186c663f0354","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1785,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:223"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def trueTopArms {k : ℕ} (nu : Fin k → Measure ℝ) (n : ℕ) : Finset (Fin k)","missing":[],"search":"truetoparms banditrlproof.musicalchairs.truetoparms noncomputable def truetoparms {k : ℕ} (nu : fin k → measure ℝ) (n : ℕ) : finset (fin k) definition compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.trueTopArms_eq_of_gap","label":"trueTopArms_eq_of_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.trueTopArms_eq_of_gap","description":"theorem trueTopArms_eq_of_gap {n k : ℕ} (nu : Fin k → Measure ℝ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (gap : ℝ) (hgap0 : 0 < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) : trueTopArms nu n = S","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-3968a0186e04","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1786,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:226"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem trueTopArms_eq_of_gap {n k : ℕ} (nu : Fin k → Measure ℝ) (S : Finset (Fin k)) (hcard : S.card = n) (hnk : n ≤ k) (gap : ℝ) (hgap0 : 0 < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) : trueTopArms nu n = S","missing":[],"search":"truetoparms_eq_of_gap banditrlproof.musicalchairs.truetoparms_eq_of_gap theorem truetoparms_eq_of_gap {n k : ℕ} (nu : fin k → measure ℝ) (s : finset (fin k)) (hcard : s.card = n) (hnk : n ≤ k) (gap : ℝ) (hgap0 : 0 < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ armmean nu a - armmean nu b) : truetoparms nu n = s theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_probability","label":"explorationGoodEvent_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationGoodEvent_probability","description":"theorem explorationGoodEvent_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNRea…","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-cc3208385792","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1787,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:234"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationGoodEvent_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationGoodEvent (n := n) (T := explorationLength k eps delta) S)","missing":[],"search":"explorationgoodevent_probability banditrlproof.musicalchairs.explorationgoodevent_probability theorem explorationgoodevent_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin k)) (hcard : s.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hepsgap : eps < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ armmean nu a - armmean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw (by omega) nu (explorationgoodevent (n := n) (t := explorationlength k eps delta) s) theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_trueTop_probability","label":"explorationGoodEvent_trueTop_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationGoodEvent_trueTop_probability","description":"theorem explorationGoodEvent_trueTop_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1…","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-1233275ca877","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1788,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:247"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationGoodEvent_trueTop_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n))","missing":[],"search":"explorationgoodevent_truetop_probability banditrlproof.musicalchairs.explorationgoodevent_truetop_probability theorem explorationgoodevent_truetop_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin k)) (hcard : s.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hepsgap : eps < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ armmean nu a - armmean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw (by omega) nu (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)) theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.armMean_mem_unitInterval","label":"armMean_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.armMean_mem_unitInterval","description":"theorem armMean_mem_unitInterval {k : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (a : Fin k) : armMean nu a ∈ Set.Icc (0 : ℝ) 1","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-595c3d8788d2","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1789,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:265"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem armMean_mem_unitInterval {k : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (a : Fin k) : armMean nu a ∈ Set.Icc (0 : ℝ) 1","missing":[],"search":"armmean_mem_unitinterval banditrlproof.musicalchairs.armmean_mem_unitinterval theorem armmean_mem_unitinterval {k : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (a : fin k) : armmean nu a ∈ set.icc (0 : ℝ) 1 theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.separating_gap_le_one","label":"separating_gap_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.separating_gap_le_one","description":"theorem separating_gap_le_one {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (gap : ℝ) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) : gap ≤ 1","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-09dd2acff060","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1790,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:277"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem separating_gap_le_one {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (gap : ℝ) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) : gap ≤ 1","missing":[],"search":"separating_gap_le_one banditrlproof.musicalchairs.separating_gap_le_one theorem separating_gap_le_one {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin k)) (hcard : s.card = n) (gap : ℝ) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ armmean nu a - armmean nu b) : gap ≤ 1 theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_exploration_good_probability","label":"source_exploration_good_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_exploration_good_probability","description":"theorem source_exploration_good_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta…","url":"../modules/banditrlproof-algorithms-musicalchairsranking/index.html#decl-c571c99988d1","parent":"module:BanditRLProof.Algorithms.MusicalChairsRanking","order":1791,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRanking"],["Source","BanditRLProof/Algorithms/MusicalChairsRanking.lean:290"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_exploration_good_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin k)) (hcard : S.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (hepsgap : eps < gap) (hgap : ∀ a ∈ S, ∀ b ∉ S, gap ≤ armMean nu a - armMean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ENNReal.ofReal delta ≤ explorationRewardLaw (by omega) nu (explorationGoodEvent (n := n) (T := explorationLength k eps delta) (trueTopArms nu n))","missing":[],"search":"source_exploration_good_probability banditrlproof.musicalchairs.source_exploration_good_probability theorem source_exploration_good_probability {n k : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin k)) (hcard : s.card = n) (eps gap delta : ℝ) (heps : 0 < eps) (hepsgap : eps < gap) (hgap : ∀ a ∈ s, ∀ b ∉ s, gap ≤ armmean nu a - armmean nu b) (hdelta : 0 < delta) (hdelta1 : delta < 1) : 1 - ennreal.ofreal delta ≤ explorationrewardlaw (by omega) nu (explorationgoodevent (n := n) (t := explorationlength k eps delta) (truetoparms nu n)) theorem compiled","shard":"modules/ebfe70a38ba5b933.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_integrable","label":"reward_coordinate_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.reward_coordinate_integrable","description":"theorem reward_coordinate_integrable {k D : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (t : Fin D) (a : Fin k) : Integrable (fun r : Fin D → Fin k → ℝ => r t a) (rewardLaw nu D)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-0d55c8aece91","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1792,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem reward_coordinate_integrable {k D : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (t : Fin D) (a : Fin k) : Integrable (fun r : Fin D → Fin k → ℝ => r t a) (rewardLaw nu D)","missing":[],"search":"reward_coordinate_integrable banditrlproof.musicalchairs.reward_coordinate_integrable theorem reward_coordinate_integrable {k d : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (t : fin d) (a : fin k) : integrable (fun r : fin d → fin k → ℝ => r t a) (rewardlaw nu d) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.scheduleRealizedRegret","label":"scheduleRealizedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.scheduleRealizedRegret","description":"noncomputable def scheduleRealizedRegret {n k D T : ℕ} (mu : Fin k → ℝ) (j : Fin T → Fin D) (a : Fin T → Fin n → Fin k) (r : Fin D → Fin k → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-880686997a9e","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1793,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def scheduleRealizedRegret {n k D T : ℕ} (mu : Fin k → ℝ) (j : Fin T → Fin D) (a : Fin T → Fin n → Fin k) (r : Fin D → Fin k → ℝ) : ℝ","missing":[],"search":"schedulerealizedregret banditrlproof.musicalchairs.schedulerealizedregret noncomputable def schedulerealizedregret {n k d t : ℕ} (mu : fin k → ℝ) (j : fin t → fin d) (a : fin t → fin n → fin k) (r : fin d → fin k → ℝ) : ℝ definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_scheduleRealizedRegret","label":"integrable_scheduleRealizedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_scheduleRealizedRegret","description":"theorem integrable_scheduleRealizedRegret {n k D T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) (j : Fin T → Fin D) (a : Fin T → Fin n → Fin k) : Integrable (scheduleRealizedRegret mu j a) (rewardLaw nu D)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-4c87eb1d272f","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1794,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_scheduleRealizedRegret {n k D T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) (j : Fin T → Fin D) (a : Fin T → Fin n → Fin k) : Integrable (scheduleRealizedRegret mu j a) (rewardLaw nu D)","missing":[],"search":"integrable_schedulerealizedregret banditrlproof.musicalchairs.integrable_schedulerealizedregret theorem integrable_schedulerealizedregret {n k d t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (mu : fin k → ℝ) (j : fin t → fin d) (a : fin t → fin n → fin k) : integrable (schedulerealizedregret mu j a) (rewardlaw nu d) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.scheduleRealizedRegret_mean","label":"scheduleRealizedRegret_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.scheduleRealizedRegret_mean","description":"theorem scheduleRealizedRegret_mean {n k D T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (j : Fin T → Fin D) (a : Fin T → Fin n → Fin k) : (∫ r, scheduleRealizedRegret (armMean nu) j a r ∂rewardLaw nu D) = ∑ t : Fin T, globalRoundRegret (armMean nu) (a t)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-2551181ea158","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1795,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem scheduleRealizedRegret_mean {n k D T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (j : Fin T → Fin D) (a : Fin T → Fin n → Fin k) : (∫ r, scheduleRealizedRegret (armMean nu) j a r ∂rewardLaw nu D) = ∑ t : Fin T, globalRoundRegret (armMean nu) (a t)","missing":[],"search":"schedulerealizedregret_mean banditrlproof.musicalchairs.schedulerealizedregret_mean theorem schedulerealizedregret_mean {n k d t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (j : fin t → fin d) (a : fin t → fin n → fin k) : (∫ r, schedulerealizedregret (armmean nu) j a r ∂rewardlaw nu d) = ∑ t : fin t, globalroundregret (armmean nu) (a t) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_scheduleRealizedRegret","label":"measurable_scheduleRealizedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_scheduleRealizedRegret","description":"theorem measurable_scheduleRealizedRegret {n k D T : ℕ} (mu : Fin k → ℝ) (j : Fin T → Fin D) : Measurable (fun z : (Fin T → Fin n → Fin k) × (Fin D → Fin k → ℝ) => scheduleRealizedRegret mu j z.1 z.2)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-35793a3bd489","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1796,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_scheduleRealizedRegret {n k D T : ℕ} (mu : Fin k → ℝ) (j : Fin T → Fin D) : Measurable (fun z : (Fin T → Fin n → Fin k) × (Fin D → Fin k → ℝ) => scheduleRealizedRegret mu j z.1 z.2)","missing":[],"search":"measurable_schedulerealizedregret banditrlproof.musicalchairs.measurable_schedulerealizedregret theorem measurable_schedulerealizedregret {n k d t : ℕ} (mu : fin k → ℝ) (j : fin t → fin d) : measurable (fun z : (fin t → fin n → fin k) × (fin d → fin k → ℝ) => schedulerealizedregret mu j z.1 z.2) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_randomScheduleRegret","label":"integrable_randomScheduleRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_randomScheduleRegret","description":"theorem integrable_randomScheduleRegret {α : Type*} [MeasurableSpace α] {n k D T : ℕ} (P : Measure α) [IsFiniteMeasure P] (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) (j : Fin T → Fin D) (a : α → Fin T → Fin n → Fin k) (ha : Measurable a) : Integrable (fun z : α × (Fin D → Fin k → ℝ) => scheduleRealizedRegret mu j (a z.1) z.2) (P.prod (rew…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-be56f75ccc5d","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1797,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_randomScheduleRegret {α : Type*} [MeasurableSpace α] {n k D T : ℕ} (P : Measure α) [IsFiniteMeasure P] (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) (j : Fin T → Fin D) (a : α → Fin T → Fin n → Fin k) (ha : Measurable a) : Integrable (fun z : α × (Fin D → Fin k → ℝ) => scheduleRealizedRegret mu j (a z.1) z.2) (P.prod (rewardLaw nu D))","missing":[],"search":"integrable_randomscheduleregret banditrlproof.musicalchairs.integrable_randomscheduleregret theorem integrable_randomscheduleregret {α : type*} [measurablespace α] {n k d t : ℕ} (p : measure α) [isfinitemeasure p] (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (mu : fin k → ℝ) (j : fin t → fin d) (a : α → fin t → fin n → fin k) (ha : measurable a) : integrable (fun z : α × (fin d → fin k → ℝ) => schedulerealizedregret mu j (a z.1) z.2) (p.prod (rewardlaw nu d)) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.randomScheduleRegret_mean","label":"randomScheduleRegret_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.randomScheduleRegret_mean","description":"theorem randomScheduleRegret_mean {α : Type*} [MeasurableSpace α] {n k D T : ℕ} (P : Measure α) [IsFiniteMeasure P] (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (j : Fin T → Fin D) (a : α → Fin T → Fin n → Fin k) (ha : Measurable a) : (∫ z, scheduleRealizedRegret (armMean nu) j (a z.1) z.2 ∂P.prod (rewardLaw nu D)) = ∫ x, ∑ t : Fin T, globalRoundRegret (ar…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-7a17936d46e4","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1798,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem randomScheduleRegret_mean {α : Type*} [MeasurableSpace α] {n k D T : ℕ} (P : Measure α) [IsFiniteMeasure P] (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (j : Fin T → Fin D) (a : α → Fin T → Fin n → Fin k) (ha : Measurable a) : (∫ z, scheduleRealizedRegret (armMean nu) j (a z.1) z.2 ∂P.prod (rewardLaw nu D)) = ∫ x, ∑ t : Fin T, globalRoundRegret (armMean nu) (a x t) ∂P","missing":[],"search":"randomscheduleregret_mean banditrlproof.musicalchairs.randomscheduleregret_mean theorem randomscheduleregret_mean {α : type*} [measurablespace α] {n k d t : ℕ} (p : measure α) [isfinitemeasure p] (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (j : fin t → fin d) (a : α → fin t → fin n → fin k) (ha : measurable a) : (∫ z, schedulerealizedregret (armmean nu) j (a z.1) z.2 ∂p.prod (rewardlaw nu d)) = ∫ x, ∑ t : fin t, globalroundregret (armmean nu) (a x t) ∂p theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationRealizedRegret","label":"explorationRealizedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationRealizedRegret","description":"noncomputable def explorationRealizedRegret {n k S : ℕ} (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (r : Fin S → Fin k → ℝ) (H : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-b0b2fa2db157","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1799,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationRealizedRegret {n k S : ℕ} (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (r : Fin S → Fin k → ℝ) (H : ℕ) : ℝ","missing":[],"search":"explorationrealizedregret banditrlproof.musicalchairs.explorationrealizedregret noncomputable def explorationrealizedregret {n k s : ℕ} (mu : fin k → ℝ) (x : fin s → fin n → fin k) (r : fin s → fin k → ℝ) (h : ℕ) : ℝ definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationRealizedRegret","label":"continuationRealizedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationRealizedRegret","description":"noncomputable def continuationRealizedRegret {n k L : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (d : Fin L → Fin n → Fin k) (r : Fin L → Fin k → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-9b99ec41ad14","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1800,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:129"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def continuationRealizedRegret {n k L : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (d : Fin L → Fin n → Fin k) (r : Fin L → Fin k → ℝ) : ℝ","missing":[],"search":"continuationrealizedregret banditrlproof.musicalchairs.continuationrealizedregret noncomputable def continuationrealizedregret {n k l : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (d : fin l → fin n → fin k) (r : fin l → fin k → ℝ) : ℝ definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.exploration_schedule_pseudo","label":"exploration_schedule_pseudo","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.exploration_schedule_pseudo","description":"theorem exploration_schedule_pseudo {n k S : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (H : ℕ) : (∑ t : Fin (min H S), globalRoundRegret mu (x (Fin.castLE (Nat.min_le_right H S) t))) = explorationPrefixRegret hk mu x H","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-42da57612fe5","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1801,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:133"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem exploration_schedule_pseudo {n k S : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (x : Fin S → Fin n → Fin k) (H : ℕ) : (∑ t : Fin (min H S), globalRoundRegret mu (x (Fin.castLE (Nat.min_le_right H S) t))) = explorationPrefixRegret hk mu x H","missing":[],"search":"exploration_schedule_pseudo banditrlproof.musicalchairs.exploration_schedule_pseudo theorem exploration_schedule_pseudo {n k s : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (x : fin s → fin n → fin k) (h : ℕ) : (∑ t : fin (min h s), globalroundregret mu (x (fin.castle (nat.min_le_right h s) t))) = explorationprefixregret hk mu x h theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuation_schedule_pseudo","label":"continuation_schedule_pseudo","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuation_schedule_pseudo","description":"theorem continuation_schedule_pseudo {n k L : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (d : Fin L → Fin n → Fin k) : (∑ t : Fin L, globalRoundRegret mu (continuationAction hk d t)) = pathCoordinationRegret hk (topArms mu n) mu d","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-e61fc47daded","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1802,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem continuation_schedule_pseudo {n k L : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (d : Fin L → Fin n → Fin k) : (∑ t : Fin L, globalRoundRegret mu (continuationAction hk d t)) = pathCoordinationRegret hk (topArms mu n) mu d","missing":[],"search":"continuation_schedule_pseudo banditrlproof.musicalchairs.continuation_schedule_pseudo theorem continuation_schedule_pseudo {n k l : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (d : fin l → fin n → fin k) : (∑ t : fin l, globalroundregret mu (continuationaction hk d t)) = pathcoordinationregret hk (toparms mu n) mu d theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_explorationRealizedRegret","label":"integrable_explorationRealizedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_explorationRealizedRegret","description":"theorem integrable_explorationRealizedRegret {n k S : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) (H : ℕ) : Integrable (fun z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ) => explorationRealizedRegret mu z.1 z.2 H) (explorationRewardLaw hk nu)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-5f9b9dac0105","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1803,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_explorationRealizedRegret {n k S : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) (H : ℕ) : Integrable (fun z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ) => explorationRealizedRegret mu z.1 z.2 H) (explorationRewardLaw hk nu)","missing":[],"search":"integrable_explorationrealizedregret banditrlproof.musicalchairs.integrable_explorationrealizedregret theorem integrable_explorationrealizedregret {n k s : ℕ} (hk : 0 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (mu : fin k → ℝ) (h : ℕ) : integrable (fun z : (fin s → fin n → fin k) × (fin s → fin k → ℝ) => explorationrealizedregret mu z.1 z.2 h) (explorationrewardlaw hk nu) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationRealizedRegret_mean","label":"explorationRealizedRegret_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationRealizedRegret_mean","description":"theorem explorationRealizedRegret_mean {n k S : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (H : ℕ) : (∫ z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ), explorationRealizedRegret (armMean nu) z.1 z.2 H ∂explorationRewardLaw hk nu) = ∫ z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ), explorationPrefixRegret hk (armMean nu) z.1 H ∂explo…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-724762df6f73","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1804,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationRealizedRegret_mean {n k S : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (H : ℕ) : (∫ z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ), explorationRealizedRegret (armMean nu) z.1 z.2 H ∂explorationRewardLaw hk nu) = ∫ z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ), explorationPrefixRegret hk (armMean nu) z.1 H ∂explorationRewardLaw hk nu","missing":[],"search":"explorationrealizedregret_mean banditrlproof.musicalchairs.explorationrealizedregret_mean theorem explorationrealizedregret_mean {n k s : ℕ} (hk : 0 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (h : ℕ) : (∫ z : (fin s → fin n → fin k) × (fin s → fin k → ℝ), explorationrealizedregret (armmean nu) z.1 z.2 h ∂explorationrewardlaw hk nu) = ∫ z : (fin s → fin n → fin k) × (fin s → fin k → ℝ), explorationprefixregret hk (armmean nu) z.1 h ∂explorationrewardlaw hk nu theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.FullSample","label":"FullSample","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.FullSample","description":"abbrev FullSample (n k S H : ℕ)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-56b24e39a077","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1805,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:179"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"abbrev FullSample (n k S H : ℕ)","missing":[],"search":"fullsample banditrlproof.musicalchairs.fullsample abbrev fullsample (n k s h : ℕ) abbreviation compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.completeLearnerLaw","label":"completeLearnerLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.completeLearnerLaw","description":"noncomputable def completeLearnerLaw {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) : Measure (FullSample n k S H)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-82d95fdeaf10","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1806,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:183"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def completeLearnerLaw {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) : Measure (FullSample n k S H)","missing":[],"search":"completelearnerlaw banditrlproof.musicalchairs.completelearnerlaw noncomputable def completelearnerlaw {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) : measure (fullsample n k s h) definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.completeLearnerLaw_first_preserving","label":"completeLearnerLaw_first_preserving","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.completeLearnerLaw_first_preserving","description":"theorem completeLearnerLaw_first_preserving {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] : MeasurePreserving (Prod.fst : FullSample n k S H → _) (completeLearnerLaw hk nu) (explorationContinuationLaw hk nu (H-S))","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-aeb14d1547d6","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1807,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:193"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem completeLearnerLaw_first_preserving {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] : MeasurePreserving (Prod.fst : FullSample n k S H → _) (completeLearnerLaw hk nu) (explorationContinuationLaw hk nu (H-S))","missing":[],"search":"completelearnerlaw_first_preserving banditrlproof.musicalchairs.completelearnerlaw_first_preserving theorem completelearnerlaw_first_preserving {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] : measurepreserving (prod.fst : fullsample n k s h → _) (completelearnerlaw hk nu) (explorationcontinuationlaw hk nu (h-s)) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_first_preserving","label":"explorationContinuationLaw_first_preserving","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationContinuationLaw_first_preserving","description":"theorem explorationContinuationLaw_first_preserving {n k S L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] : MeasurePreserving (Prod.fst : ((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ)) × (Fin L → Fin n → Fin k) → _) (explorationContinuationLaw hk nu L) (explorationRewardLaw (by omega) nu)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-a115e5672d05","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1808,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:198"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationContinuationLaw_first_preserving {n k S L : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] : MeasurePreserving (Prod.fst : ((Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ)) × (Fin L → Fin n → Fin k) → _) (explorationContinuationLaw hk nu L) (explorationRewardLaw (by omega) nu)","missing":[],"search":"explorationcontinuationlaw_first_preserving banditrlproof.musicalchairs.explorationcontinuationlaw_first_preserving theorem explorationcontinuationlaw_first_preserving {n k s l : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] : measurepreserving (prod.fst : ((fin s → fin n → fin k) × (fin s → fin k → ℝ)) × (fin l → fin n → fin k) → _) (explorationcontinuationlaw hk nu l) (explorationrewardlaw (by omega) nu) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.completeLearnerLaw_exploration_preserving","label":"completeLearnerLaw_exploration_preserving","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.completeLearnerLaw_exploration_preserving","description":"theorem completeLearnerLaw_exploration_preserving {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] : MeasurePreserving (fun z : FullSample n k S H => z.1.1) (completeLearnerLaw hk nu) (explorationRewardLaw (by omega) nu)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-1e29ce94d377","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1809,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:207"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem completeLearnerLaw_exploration_preserving {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] : MeasurePreserving (fun z : FullSample n k S H => z.1.1) (completeLearnerLaw hk nu) (explorationRewardLaw (by omega) nu)","missing":[],"search":"completelearnerlaw_exploration_preserving banditrlproof.musicalchairs.completelearnerlaw_exploration_preserving theorem completelearnerlaw_exploration_preserving {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] : measurepreserving (fun z : fullsample n k s h => z.1.1) (completelearnerlaw hk nu) (explorationrewardlaw (by omega) nu) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret","label":"realizedLearnerRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.realizedLearnerRegret","description":"noncomputable def realizedLearnerRegret {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (z : FullSample n k S H) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-a4c190eee13b","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1810,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:213"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def realizedLearnerRegret {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (z : FullSample n k S H) : ℝ","missing":[],"search":"realizedlearnerregret banditrlproof.musicalchairs.realizedlearnerregret noncomputable def realizedlearnerregret {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (z : fullsample n k s h) : ℝ definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integral_preserving_pullback","label":"integral_preserving_pullback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integral_preserving_pullback","description":"theorem integral_preserving_pullback {α β : Type*} [MeasurableSpace α] [MeasurableSpace β] {P : Measure α} {Q : Measure β} (g : α → β) (hg : MeasurePreserving g P Q) (f : β → ℝ) (hf : Integrable f Q) : (∫ x, f (g x) ∂P) = ∫ y, f y ∂Q","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-84cc093dc246","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1811,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:222"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integral_preserving_pullback {α β : Type*} [MeasurableSpace α] [MeasurableSpace β] {P : Measure α} {Q : Measure β} (g : α → β) (hg : MeasurePreserving g P Q) (f : β → ℝ) (hf : Integrable f Q) : (∫ x, f (g x) ∂P) = ∫ y, f y ∂Q","missing":[],"search":"integral_preserving_pullback banditrlproof.musicalchairs.integral_preserving_pullback theorem integral_preserving_pullback {α β : type*} [measurablespace α] [measurablespace β] {p : measure α} {q : measure β} (g : α → β) (hg : measurepreserving g p q) (f : β → ℝ) (hf : integrable f q) : (∫ x, f (g x) ∂p) = ∫ y, f y ∂q theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_continuationSchedule","label":"measurable_continuationSchedule","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_continuationSchedule","description":"theorem measurable_continuationSchedule {n k L : ℕ} (hk : 0 < k) : Measurable (fun d : Fin L → Fin n → Fin k => fun t : Fin L => continuationAction hk d t)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-d7ef4a10bb61","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1812,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:228"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_continuationSchedule {n k L : ℕ} (hk : 0 < k) : Measurable (fun d : Fin L → Fin n → Fin k => fun t : Fin L => continuationAction hk d t)","missing":[],"search":"measurable_continuationschedule banditrlproof.musicalchairs.measurable_continuationschedule theorem measurable_continuationschedule {n k l : ℕ} (hk : 0 < k) : measurable (fun d : fin l → fin n → fin k => fun t : fin l => continuationaction hk d t) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_realizedLearnerRegret","label":"integrable_realizedLearnerRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_realizedLearnerRegret","description":"theorem integrable_realizedLearnerRegret {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) : Integrable (realizedLearnerRegret (n := n) (S := S) (H := H) (by omega) mu) (completeLearnerLaw hk nu)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-4a66dcc279d3","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1813,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:232"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_realizedLearnerRegret {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) : Integrable (realizedLearnerRegret (n := n) (S := S) (H := H) (by omega) mu) (completeLearnerLaw hk nu)","missing":[],"search":"integrable_realizedlearnerregret banditrlproof.musicalchairs.integrable_realizedlearnerregret theorem integrable_realizedlearnerregret {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (mu : fin k → ℝ) : integrable (realizedlearnerregret (n := n) (s := s) (h := h) (by omega) mu) (completelearnerlaw hk nu) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationRealizedRegret_mean_joint","label":"continuationRealizedRegret_mean_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationRealizedRegret_mean_joint","description":"theorem continuationRealizedRegret_mean_joint {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) : (∫ z : FullSample n k S H, continuationRealizedRegret (by omega) (armMean nu) z.1.2 z.2 ∂completeLearnerLaw hk nu) = ∫ z, pathCoordinationRegret (by omega) (trueTopArms nu n) (armMean nu) z.2 ∂explorationContinuationLaw (n := n) (T := S)…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-03e83e4d9204","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1814,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:244"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem continuationRealizedRegret_mean_joint {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) : (∫ z : FullSample n k S H, continuationRealizedRegret (by omega) (armMean nu) z.1.2 z.2 ∂completeLearnerLaw hk nu) = ∫ z, pathCoordinationRegret (by omega) (trueTopArms nu n) (armMean nu) z.2 ∂explorationContinuationLaw (n := n) (T := S) hk nu (H-S)","missing":[],"search":"continuationrealizedregret_mean_joint banditrlproof.musicalchairs.continuationrealizedregret_mean_joint theorem continuationrealizedregret_mean_joint {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) : (∫ z : fullsample n k s h, continuationrealizedregret (by omega) (armmean nu) z.1.2 z.2 ∂completelearnerlaw hk nu) = ∫ z, pathcoordinationregret (by omega) (truetoparms nu n) (armmean nu) z.2 ∂explorationcontinuationlaw (n := n) (t := s) hk nu (h-s) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret_expected_eq_pseudo","label":"realizedLearnerRegret_expected_eq_pseudo","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.realizedLearnerRegret_expected_eq_pseudo","description":"theorem realizedLearnerRegret_expected_eq_pseudo {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) : (∫ z : FullSample n k S H, realizedLearnerRegret (by omega) (armMean nu) z ∂completeLearnerLaw hk nu) = ∫ z, learnerRegret (n := n) (S := S) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (n := n) (T := S) hk nu…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-9918c0bfa206","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1815,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:256"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem realizedLearnerRegret_expected_eq_pseudo {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) : (∫ z : FullSample n k S H, realizedLearnerRegret (by omega) (armMean nu) z ∂completeLearnerLaw hk nu) = ∫ z, learnerRegret (n := n) (S := S) (H := H) (by omega) (armMean nu) z.1.1 z.2 ∂explorationContinuationLaw (n := n) (T := S) hk nu (H-S)","missing":[],"search":"realizedlearnerregret_expected_eq_pseudo banditrlproof.musicalchairs.realizedlearnerregret_expected_eq_pseudo theorem realizedlearnerregret_expected_eq_pseudo {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) : (∫ z : fullsample n k s h, realizedlearnerregret (by omega) (armmean nu) z ∂completelearnerlaw hk nu) = ∫ z, learnerregret (n := n) (s := s) (h := h) (by omega) (armmean nu) z.1.1 z.2 ∂explorationcontinuationlaw (n := n) (t := s) hk nu (h-s) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_expected_realizedLearnerRegret_le","label":"source_expected_realizedLearnerRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_expected_realizedLearnerRegret_le","description":"theorem source_expected_realizedLearnerRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z : FullSample n k (explorationLength k eps delta) H, realizedLearnerRegret (by omega) (armM…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-71acb7a11866","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1816,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:314"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_expected_realizedLearnerRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z : FullSample n k (explorationLength k eps delta) H, realizedLearnerRegret (by omega) (armMean nu) z ∂completeLearnerLaw (by omega) nu) ≤ min ((n : ℝ)*H) ((n : ℝ)*explorationLength k eps delta + 8*(n : ℝ)^2 + delta*((n : ℝ)*H))","missing":[],"search":"source_expected_realizedlearnerregret_le banditrlproof.musicalchairs.source_expected_realizedlearnerregret_le theorem source_expected_realizedlearnerregret_le {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ z : fullsample n k (explorationlength k eps delta) h, realizedlearnerregret (by omega) (armmean nu) z ∂completelearnerlaw (by omega) nu) ≤ min ((n : ℝ)*h) ((n : ℝ)*explorationlength k eps delta + 8*(n : ℝ)^2 + delta*((n : ℝ)*h)) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationFeedback","label":"continuationFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationFeedback","description":"noncomputable def continuationFeedback {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (r : Fin L → Fin k → ℝ) (i : Fin n) (t : Fin L) : ExplorationFeedback k","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-005fdb900dcc","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1817,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:332"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def continuationFeedback {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (r : Fin L → Fin k → ℝ) (i : Fin n) (t : Fin L) : ExplorationFeedback k","missing":[],"search":"continuationfeedback banditrlproof.musicalchairs.continuationfeedback noncomputable def continuationfeedback {n k l : ℕ} (hk : 0 < k) (d : fin l → fin n → fin k) (r : fin l → fin k → ℝ) (i : fin n) (t : fin l) : explorationfeedback k definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback","label":"learnerFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback","description":"noncomputable def learnerFeedback {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) : ExplorationFeedback k","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-e37629ae456e","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1818,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:338"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def learnerFeedback {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) : ExplorationFeedback k","missing":[],"search":"learnerfeedback banditrlproof.musicalchairs.learnerfeedback noncomputable def learnerfeedback {n k s h : ℕ} (hk : 0 < k) (z : fullsample n k s h) (i : fin n) (t : fin h) : explorationfeedback k definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_explore","label":"learnerFeedback_explore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback_explore","description":"theorem learnerFeedback_explore {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) (ht : t.val < S) : learnerFeedback hk z i t = explorationFeedback z.1.1.1 z.1.1.2 i ⟨t.val,ht⟩","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-460507a52abb","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1819,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:343"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerFeedback_explore {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) (ht : t.val < S) : learnerFeedback hk z i t = explorationFeedback z.1.1.1 z.1.1.2 i ⟨t.val,ht⟩","missing":[],"search":"learnerfeedback_explore banditrlproof.musicalchairs.learnerfeedback_explore theorem learnerfeedback_explore {n k s h : ℕ} (hk : 0 < k) (z : fullsample n k s h) (i : fin n) (t : fin h) (ht : t.val < s) : learnerfeedback hk z i t = explorationfeedback z.1.1.1 z.1.1.2 i ⟨t.val,ht⟩ theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_coordinate","label":"learnerFeedback_coordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback_coordinate","description":"theorem learnerFeedback_coordinate {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (u : Fin (H-S)) : learnerFeedback hk z i ⟨S+u.val,by omega⟩ = continuationFeedback hk z.1.2 z.2 i u","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-31ab3cfd221f","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1820,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:348"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerFeedback_coordinate {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (u : Fin (H-S)) : learnerFeedback hk z i ⟨S+u.val,by omega⟩ = continuationFeedback hk z.1.2 z.2 i u","missing":[],"search":"learnerfeedback_coordinate banditrlproof.musicalchairs.learnerfeedback_coordinate theorem learnerfeedback_coordinate {n k s h : ℕ} (hk : 0 < k) (z : fullsample n k s h) (i : fin n) (u : fin (h-s)) : learnerfeedback hk z i ⟨s+u.val,by omega⟩ = continuationfeedback hk z.1.2 z.2 i u theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_arm","label":"learnerFeedback_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback_arm","description":"theorem learnerFeedback_arm {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) : (learnerFeedback hk z i t).arm = learnerAction hk z.1.1.1 z.1.2 t.val i","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-ae46e1b3eac6","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1821,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:353"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerFeedback_arm {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) : (learnerFeedback hk z i t).arm = learnerAction hk z.1.1.1 z.1.2 t.val i","missing":[],"search":"learnerfeedback_arm banditrlproof.musicalchairs.learnerfeedback_arm theorem learnerfeedback_arm {n k s h : ℕ} (hk : 0 < k) (z : fullsample n k s h) (i : fin n) (t : fin h) : (learnerfeedback hk z i t).arm = learneraction hk z.1.1.1 z.1.2 t.val i theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_collision","label":"learnerFeedback_collision","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback_collision","description":"theorem learnerFeedback_collision {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) : (learnerFeedback hk z i t).collided = collisionBit (learnerAction hk z.1.1.1 z.1.2 t.val) i","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-70a53b35a58c","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1822,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:363"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerFeedback_collision {n k S H : ℕ} (hk : 0 < k) (z : FullSample n k S H) (i : Fin n) (t : Fin H) : (learnerFeedback hk z i t).collided = collisionBit (learnerAction hk z.1.1.1 z.1.2 t.val) i","missing":[],"search":"learnerfeedback_collision banditrlproof.musicalchairs.learnerfeedback_collision theorem learnerfeedback_collision {n k s h : ℕ} (hk : 0 < k) (z : fullsample n k s h) (i : fin n) (t : fin h) : (learnerfeedback hk z i t).collided = collisionbit (learneraction hk z.1.1.1 z.1.2 t.val) i theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_completed_exploration","label":"learnerFeedback_completed_exploration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback_completed_exploration","description":"theorem learnerFeedback_completed_exploration {n k S H : ℕ} (hk : 0 < k) (hSH : S ≤ H) (z : FullSample n k S H) (i : Fin n) : (fun t : Fin S => learnerFeedback hk z i (Fin.castLE hSH t)) = explorationFeedback z.1.1.1 z.1.1.2 i","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-37920c0787f1","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1823,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:374"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerFeedback_completed_exploration {n k S H : ℕ} (hk : 0 < k) (hSH : S ≤ H) (z : FullSample n k S H) (i : Fin n) : (fun t : Fin S => learnerFeedback hk z i (Fin.castLE hSH t)) = explorationFeedback z.1.1.1 z.1.1.2 i","missing":[],"search":"learnerfeedback_completed_exploration banditrlproof.musicalchairs.learnerfeedback_completed_exploration theorem learnerfeedback_completed_exploration {n k s h : ℕ} (hk : 0 < k) (hsh : s ≤ h) (z : fullsample n k s h) (i : fin n) : (fun t : fin s => learnerfeedback hk z i (fin.castle hsh t)) = explorationfeedback z.1.1.1 z.1.1.2 i theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnedConfig_from_feedback","label":"learnedConfig_from_feedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnedConfig_from_feedback","description":"theorem learnedConfig_from_feedback {n k S H : ℕ} (hk : 1 < k) (hSH : S ≤ H) (z : FullSample n k S H) (i : Fin n) : (learnedConfig hk z.1.1).val i = localCandidateSet (fun t : Fin S => learnerFeedback (by omega) z i (Fin.castLE hSH t))","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-330cc0ec70ac","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1824,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:382"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnedConfig_from_feedback {n k S H : ℕ} (hk : 1 < k) (hSH : S ≤ H) (z : FullSample n k S H) (i : Fin n) : (learnedConfig hk z.1.1).val i = localCandidateSet (fun t : Fin S => learnerFeedback (by omega) z i (Fin.castLE hSH t))","missing":[],"search":"learnedconfig_from_feedback banditrlproof.musicalchairs.learnedconfig_from_feedback theorem learnedconfig_from_feedback {n k s h : ℕ} (hk : 1 < k) (hsh : s ≤ h) (z : fullsample n k s h) (i : fin n) : (learnedconfig hk z.1.1).val i = localcandidateset (fun t : fin s => learnerfeedback (by omega) z i (fin.castle hsh t)) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.continuationFeedback_local_update","label":"continuationFeedback_local_update","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.continuationFeedback_local_update","description":"theorem continuationFeedback_local_update {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (r : Fin L → Fin k → ℝ) (i : Fin n) (t : Fin L) : trajectory (extendedCoordinationDraws hk d) (t.val+1) i = localUpdate (trajectory (extendedCoordinationDraws hk d) t.val i) (continuationFeedback hk d r i t).arm (continuationFeedback hk d r i t).collided","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-fad79831ca8b","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1825,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:389"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem continuationFeedback_local_update {n k L : ℕ} (hk : 0 < k) (d : Fin L → Fin n → Fin k) (r : Fin L → Fin k → ℝ) (i : Fin n) (t : Fin L) : trajectory (extendedCoordinationDraws hk d) (t.val+1) i = localUpdate (trajectory (extendedCoordinationDraws hk d) t.val i) (continuationFeedback hk d r i t).arm (continuationFeedback hk d r i t).collided","missing":[],"search":"continuationfeedback_local_update banditrlproof.musicalchairs.continuationfeedback_local_update theorem continuationfeedback_local_update {n k l : ℕ} (hk : 0 < k) (d : fin l → fin n → fin k) (r : fin l → fin k → ℝ) (i : fin n) (t : fin l) : trajectory (extendedcoordinationdraws hk d) (t.val+1) i = localupdate (trajectory (extendedcoordinationdraws hk d) t.val i) (continuationfeedback hk d r i t).arm (continuationfeedback hk d r i t).collided theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_prefix","label":"learnerFeedback_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.learnerFeedback_prefix","description":"theorem learnerFeedback_prefix {n k S H : ℕ} (hk : 0 < k) (z w : FullSample n k S H) (i : Fin n) (t : Fin H) (hx : ∀ u : Fin S, u.val ≤ t.val → z.1.1.1 u = w.1.1.1 u) (hr : ∀ u : Fin S, u.val ≤ t.val → z.1.1.2 u = w.1.1.2 u) (hd : ∀ u : Fin (H-S), S+u.val ≤ t.val → z.1.2 u = w.1.2 u) (hy : ∀ u : Fin (H-S), S+u.val ≤ t.val → z.2 u = w.2 u) : learnerFeedback hk z i t = learnerFeedback hk w i t","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-f6f53273906c","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1826,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:402"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem learnerFeedback_prefix {n k S H : ℕ} (hk : 0 < k) (z w : FullSample n k S H) (i : Fin n) (t : Fin H) (hx : ∀ u : Fin S, u.val ≤ t.val → z.1.1.1 u = w.1.1.1 u) (hr : ∀ u : Fin S, u.val ≤ t.val → z.1.1.2 u = w.1.1.2 u) (hd : ∀ u : Fin (H-S), S+u.val ≤ t.val → z.1.2 u = w.1.2 u) (hy : ∀ u : Fin (H-S), S+u.val ≤ t.val → z.2 u = w.2 u) : learnerFeedback hk z i t = learnerFeedback hk w i t","missing":[],"search":"learnerfeedback_prefix banditrlproof.musicalchairs.learnerfeedback_prefix theorem learnerfeedback_prefix {n k s h : ℕ} (hk : 0 < k) (z w : fullsample n k s h) (i : fin n) (t : fin h) (hx : ∀ u : fin s, u.val ≤ t.val → z.1.1.1 u = w.1.1.1 u) (hr : ∀ u : fin s, u.val ≤ t.val → z.1.1.2 u = w.1.1.2 u) (hd : ∀ u : fin (h-s), s+u.val ≤ t.val → z.1.2 u = w.1.2 u) (hy : ∀ u : fin (h-s), s+u.val ≤ t.val → z.2 u = w.2 u) : learnerfeedback hk z i t = learnerfeedback hk w i t theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.finite_sum_split_at","label":"finite_sum_split_at","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.finite_sum_split_at","description":"theorem finite_sum_split_at {H S : ℕ} (hSH : S ≤ H) (f : Fin H → ℝ) : (∑ t : Fin H, f t) = (∑ t : Fin S, f (Fin.castLE hSH t)) + ∑ u : Fin (H-S), f ⟨S+u.val,by omega⟩","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-834ab4251fe7","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1827,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:424"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem finite_sum_split_at {H S : ℕ} (hSH : S ≤ H) (f : Fin H → ℝ) : (∑ t : Fin H, f t) = (∑ t : Fin S, f (Fin.castLE hSH t)) + ∑ u : Fin (H-S), f ⟨S+u.val,by omega⟩","missing":[],"search":"finite_sum_split_at banditrlproof.musicalchairs.finite_sum_split_at theorem finite_sum_split_at {h s : ℕ} (hsh : s ≤ h) (f : fin h → ℝ) : (∑ t : fin h, f t) = (∑ t : fin s, f (fin.castle hsh t)) + ∑ u : fin (h-s), f ⟨s+u.val,by omega⟩ theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret_eq_feedback","label":"realizedLearnerRegret_eq_feedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.realizedLearnerRegret_eq_feedback","description":"theorem realizedLearnerRegret_eq_feedback {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (z : FullSample n k S H) : realizedLearnerRegret hk mu z = ∑ t : Fin H, ((∑ a ∈ topArms mu n, mu a) - ∑ i : Fin n, (learnerFeedback hk z i t).reward)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-cdbb9f78273e","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1828,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:435"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem realizedLearnerRegret_eq_feedback {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (z : FullSample n k S H) : realizedLearnerRegret hk mu z = ∑ t : Fin H, ((∑ a ∈ topArms mu n, mu a) - ∑ i : Fin n, (learnerFeedback hk z i t).reward)","missing":[],"search":"realizedlearnerregret_eq_feedback banditrlproof.musicalchairs.realizedlearnerregret_eq_feedback theorem realizedlearnerregret_eq_feedback {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (z : fullsample n k s h) : realizedlearnerregret hk mu z = ∑ t : fin h, ((∑ a ∈ toparms mu n, mu a) - ∑ i : fin n, (learnerfeedback hk z i t).reward) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_feedback_mk","label":"measurable_feedback_mk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_feedback_mk","description":"theorem measurable_feedback_mk {α : Type*} [MeasurableSpace α] {k : ℕ} (a : α → Fin k) (c : α → Bool) (r : α → ℝ) (ha : Measurable a) (hc : Measurable c) (hr : Measurable r) : Measurable (fun x => (⟨a x,c x,r x⟩ : ExplorationFeedback k))","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-cc4dc32f0cb0","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1829,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:475"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_feedback_mk {α : Type*} [MeasurableSpace α] {k : ℕ} (a : α → Fin k) (c : α → Bool) (r : α → ℝ) (ha : Measurable a) (hc : Measurable c) (hr : Measurable r) : Measurable (fun x => (⟨a x,c x,r x⟩ : ExplorationFeedback k))","missing":[],"search":"measurable_feedback_mk banditrlproof.musicalchairs.measurable_feedback_mk theorem measurable_feedback_mk {α : type*} [measurablespace α] {k : ℕ} (a : α → fin k) (c : α → bool) (r : α → ℝ) (ha : measurable a) (hc : measurable c) (hr : measurable r) : measurable (fun x => (⟨a x,c x,r x⟩ : explorationfeedback k)) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_explorationFeedback","label":"measurable_explorationFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_explorationFeedback","description":"theorem measurable_explorationFeedback {n k S : ℕ} (i : Fin n) (t : Fin S) : Measurable (fun z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ) => explorationFeedback z.1 z.2 i t)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-a8af01efb5ad","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1830,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:481"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_explorationFeedback {n k S : ℕ} (i : Fin n) (t : Fin S) : Measurable (fun z : (Fin S → Fin n → Fin k) × (Fin S → Fin k → ℝ) => explorationFeedback z.1 z.2 i t)","missing":[],"search":"measurable_explorationfeedback banditrlproof.musicalchairs.measurable_explorationfeedback theorem measurable_explorationfeedback {n k s : ℕ} (i : fin n) (t : fin s) : measurable (fun z : (fin s → fin n → fin k) × (fin s → fin k → ℝ) => explorationfeedback z.1 z.2 i t) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_continuationFeedback","label":"measurable_continuationFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_continuationFeedback","description":"theorem measurable_continuationFeedback {n k L : ℕ} (hk : 0 < k) (i : Fin n) (t : Fin L) : Measurable (fun z : (Fin L → Fin n → Fin k) × (Fin L → Fin k → ℝ) => continuationFeedback hk z.1 z.2 i t)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-0d63bae70f66","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1831,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:497"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_continuationFeedback {n k L : ℕ} (hk : 0 < k) (i : Fin n) (t : Fin L) : Measurable (fun z : (Fin L → Fin n → Fin k) × (Fin L → Fin k → ℝ) => continuationFeedback hk z.1 z.2 i t)","missing":[],"search":"measurable_continuationfeedback banditrlproof.musicalchairs.measurable_continuationfeedback theorem measurable_continuationfeedback {n k l : ℕ} (hk : 0 < k) (i : fin n) (t : fin l) : measurable (fun z : (fin l → fin n → fin k) × (fin l → fin k → ℝ) => continuationfeedback hk z.1 z.2 i t) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_learnerFeedback","label":"measurable_learnerFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_learnerFeedback","description":"theorem measurable_learnerFeedback {n k S H : ℕ} (hk : 0 < k) (i : Fin n) (t : Fin H) : Measurable (fun z : FullSample n k S H => learnerFeedback hk z i t)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-cd936f224dbb","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1832,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:513"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_learnerFeedback {n k S H : ℕ} (hk : 0 < k) (i : Fin n) (t : Fin H) : Measurable (fun z : FullSample n k S H => learnerFeedback hk z i t)","missing":[],"search":"measurable_learnerfeedback banditrlproof.musicalchairs.measurable_learnerfeedback theorem measurable_learnerfeedback {n k s h : ℕ} (hk : 0 < k) (i : fin n) (t : fin h) : measurable (fun z : fullsample n k s h => learnerfeedback hk z i t) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_learnerTrace","label":"measurable_learnerTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_learnerTrace","description":"theorem measurable_learnerTrace {n k S H : ℕ} (hk : 0 < k) : Measurable (fun z : FullSample n k S H => fun i t => learnerFeedback hk z i t)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-944a9b0119ac","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1833,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:522"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_learnerTrace {n k S H : ℕ} (hk : 0 < k) : Measurable (fun z : FullSample n k S H => fun i t => learnerFeedback hk z i t)","missing":[],"search":"measurable_learnertrace banditrlproof.musicalchairs.measurable_learnertrace theorem measurable_learnertrace {n k s h : ℕ} (hk : 0 < k) : measurable (fun z : fullsample n k s h => fun i t => learnerfeedback hk z i t) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.visibleLearnerLaw","label":"visibleLearnerLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.visibleLearnerLaw","description":"noncomputable def visibleLearnerLaw {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) : Measure (Fin n → Fin H → ExplorationFeedback k)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-86571db4c4a7","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1834,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:526"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def visibleLearnerLaw {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) : Measure (Fin n → Fin H → ExplorationFeedback k)","missing":[],"search":"visiblelearnerlaw banditrlproof.musicalchairs.visiblelearnerlaw noncomputable def visiblelearnerlaw {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) : measure (fin n → fin h → explorationfeedback k) definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.visibleRegret","label":"visibleRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.visibleRegret","description":"noncomputable def visibleRegret {n k H : ℕ} (mu : Fin k → ℝ) (f : Fin n → Fin H → ExplorationFeedback k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-a0a60a713b5f","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1835,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:541"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def visibleRegret {n k H : ℕ} (mu : Fin k → ℝ) (f : Fin n → Fin H → ExplorationFeedback k) : ℝ","missing":[],"search":"visibleregret banditrlproof.musicalchairs.visibleregret noncomputable def visibleregret {n k h : ℕ} (mu : fin k → ℝ) (f : fin n → fin h → explorationfeedback k) : ℝ definition compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_feedback_reward","label":"measurable_feedback_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_feedback_reward","description":"theorem measurable_feedback_reward {k : ℕ} : Measurable (ExplorationFeedback.reward : ExplorationFeedback k → ℝ)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-bf47b90fb232","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1836,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:545"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_feedback_reward {k : ℕ} : Measurable (ExplorationFeedback.reward : ExplorationFeedback k → ℝ)","missing":[],"search":"measurable_feedback_reward banditrlproof.musicalchairs.measurable_feedback_reward theorem measurable_feedback_reward {k : ℕ} : measurable (explorationfeedback.reward : explorationfeedback k → ℝ) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_visibleRegret","label":"measurable_visibleRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_visibleRegret","description":"theorem measurable_visibleRegret {n k H : ℕ} (mu : Fin k → ℝ) : Measurable (visibleRegret (n := n) (H := H) mu)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-b1b69db1a875","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1837,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:551"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_visibleRegret {n k H : ℕ} (mu : Fin k → ℝ) : Measurable (visibleRegret (n := n) (H := H) mu)","missing":[],"search":"measurable_visibleregret banditrlproof.musicalchairs.measurable_visibleregret theorem measurable_visibleregret {n k h : ℕ} (mu : fin k → ℝ) : measurable (visibleregret (n := n) (h := h) mu) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.visibleRegret_pullback","label":"visibleRegret_pullback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.visibleRegret_pullback","description":"theorem visibleRegret_pullback {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (z : FullSample n k S H) : visibleRegret mu (fun i t => learnerFeedback hk z i t) = realizedLearnerRegret hk mu z","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-27600836c505","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1838,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:560"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem visibleRegret_pullback {n k S H : ℕ} (hk : 0 < k) (mu : Fin k → ℝ) (z : FullSample n k S H) : visibleRegret mu (fun i t => learnerFeedback hk z i t) = realizedLearnerRegret hk mu z","missing":[],"search":"visibleregret_pullback banditrlproof.musicalchairs.visibleregret_pullback theorem visibleregret_pullback {n k s h : ℕ} (hk : 0 < k) (mu : fin k → ℝ) (z : fullsample n k s h) : visibleregret mu (fun i t => learnerfeedback hk z i t) = realizedlearnerregret hk mu z theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.integrable_visibleRegret","label":"integrable_visibleRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.integrable_visibleRegret","description":"theorem integrable_visibleRegret {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) : Integrable (visibleRegret (n := n) (H := H) mu) (visibleLearnerLaw (S := S) hk nu)","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-0ed960c8ff2d","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1839,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:565"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem integrable_visibleRegret {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (mu : Fin k → ℝ) : Integrable (visibleRegret (n := n) (H := H) mu) (visibleLearnerLaw (S := S) hk nu)","missing":[],"search":"integrable_visibleregret banditrlproof.musicalchairs.integrable_visibleregret theorem integrable_visibleregret {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (mu : fin k → ℝ) : integrable (visibleregret (n := n) (h := h) mu) (visiblelearnerlaw (s := s) hk nu) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.visibleRegret_expected_eq_realized","label":"visibleRegret_expected_eq_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.visibleRegret_expected_eq_realized","description":"theorem visibleRegret_expected_eq_realized {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) (mu : Fin k → ℝ) : (∫ f, visibleRegret (n := n) (H := H) mu f ∂visibleLearnerLaw (S := S) hk nu) = ∫ z : FullSample n k S H, realizedLearnerRegret (by omega) mu z ∂completeLearnerLaw hk nu","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-989bc46e6114","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1840,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:574"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem visibleRegret_expected_eq_realized {n k S H : ℕ} (hk : 1 < k) (nu : Fin k → Measure ℝ) (mu : Fin k → ℝ) : (∫ f, visibleRegret (n := n) (H := H) mu f ∂visibleLearnerLaw (S := S) hk nu) = ∫ z : FullSample n k S H, realizedLearnerRegret (by omega) mu z ∂completeLearnerLaw hk nu","missing":[],"search":"visibleregret_expected_eq_realized banditrlproof.musicalchairs.visibleregret_expected_eq_realized theorem visibleregret_expected_eq_realized {n k s h : ℕ} (hk : 1 < k) (nu : fin k → measure ℝ) (mu : fin k → ℝ) : (∫ f, visibleregret (n := n) (h := h) mu f ∂visiblelearnerlaw (s := s) hk nu) = ∫ z : fullsample n k s h, realizedlearnerregret (by omega) mu z ∂completelearnerlaw hk nu theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.source_expected_visibleRegret_le","label":"source_expected_visibleRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.source_expected_visibleRegret_le","description":"theorem source_expected_visibleRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ f, visibleRegret (n := n) (H := H) (armMean nu) f ∂visibleLearnerLaw (S := explorationLength k eps d…","url":"../modules/banditrlproof-algorithms-musicalchairsrealized/index.html#decl-28d23b286a11","parent":"module:BanditRLProof.Algorithms.MusicalChairsRealized","order":1841,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsRealized"],["Source","BanditRLProof/Algorithms/MusicalChairsRealized.lean:583"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem source_expected_visibleRegret_le {n k H : ℕ} (hn : 0 < n) (hnk : n < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundaryGap (armMean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ f, visibleRegret (n := n) (H := H) (armMean nu) f ∂visibleLearnerLaw (S := explorationLength k eps delta) (by omega) nu) ≤ min ((n : ℝ)*H) ((n : ℝ)*explorationLength k eps delta + 8*(n : ℝ)^2 + delta*((n : ℝ)*H))","missing":[],"search":"source_expected_visibleregret_le banditrlproof.musicalchairs.source_expected_visibleregret_le theorem source_expected_visibleregret_le {n k h : ℕ} (hn : 0 < n) (hnk : n < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (hepsgap : eps < boundarygap (armmean nu) n hn hnk) (hdelta : 0 < delta) (hdelta1 : delta < 1) : (∫ f, visibleregret (n := n) (h := h) (armmean nu) f ∂visiblelearnerlaw (s := explorationlength k eps delta) (by omega) nu) ≤ min ((n : ℝ)*h) ((n : ℝ)*explorationlength k eps delta + 8*(n : ℝ)^2 + delta*((n : ℝ)*h)) theorem compiled","shard":"modules/8a53c5ceb6f3f643.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.rewardLaw","label":"rewardLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.rewardLaw","description":"noncomputable def rewardLaw {k : ℕ} (nu : Fin k → Measure ℝ) (T : ℕ) : Measure (Fin T → Fin k → ℝ)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-c000dbb07719","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1842,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def rewardLaw {k : ℕ} (nu : Fin k → Measure ℝ) (T : ℕ) : Measure (Fin T → Fin k → ℝ)","missing":[],"search":"rewardlaw banditrlproof.musicalchairs.rewardlaw noncomputable def rewardlaw {k : ℕ} (nu : fin k → measure ℝ) (t : ℕ) : measure (fin t → fin k → ℝ) definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.armMean","label":"armMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.armMean","description":"noncomputable def armMean {k : ℕ} (nu : Fin k → Measure ℝ) (a : Fin k) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-f3c259336c1a","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1843,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def armMean {k : ℕ} (nu : Fin k → Measure ℝ) (a : Fin k) : ℝ","missing":[],"search":"armmean banditrlproof.musicalchairs.armmean noncomputable def armmean {k : ℕ} (nu : fin k → measure ℝ) (a : fin k) : ℝ definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_preserving","label":"reward_coordinate_preserving","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.reward_coordinate_preserving","description":"theorem reward_coordinate_preserving {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (t : Fin T) (a : Fin k) : MeasurePreserving (fun r : Fin T → Fin k → ℝ => r t a) (rewardLaw nu T) (nu a)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-29c4148d295f","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1844,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem reward_coordinate_preserving {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (t : Fin T) (a : Fin k) : MeasurePreserving (fun r : Fin T → Fin k → ℝ => r t a) (rewardLaw nu T) (nu a)","missing":[],"search":"reward_coordinate_preserving banditrlproof.musicalchairs.reward_coordinate_preserving theorem reward_coordinate_preserving {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (t : fin t) (a : fin k) : measurepreserving (fun r : fin t → fin k → ℝ => r t a) (rewardlaw nu t) (nu a) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_mean","label":"reward_coordinate_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.reward_coordinate_mean","description":"theorem reward_coordinate_mean {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (t : Fin T) (a : Fin k) : (∫ r, r t a ∂rewardLaw nu T) = armMean nu a","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-a0806720c725","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1845,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem reward_coordinate_mean {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (t : Fin T) (a : Fin k) : (∫ r, r t a ∂rewardLaw nu T) = armMean nu a","missing":[],"search":"reward_coordinate_mean banditrlproof.musicalchairs.reward_coordinate_mean theorem reward_coordinate_mean {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (t : fin t) (a : fin k) : (∫ r, r t a ∂rewardlaw nu t) = armmean nu a theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_bounded","label":"reward_coordinate_bounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.reward_coordinate_bounded","description":"theorem reward_coordinate_bounded {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (t : Fin T) (a : Fin k) : ∀ᵐ r ∂rewardLaw nu T, r t a ∈ Set.Icc (0 : ℝ) 1","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-6c84f821126a","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1846,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem reward_coordinate_bounded {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (t : Fin T) (a : Fin k) : ∀ᵐ r ∂rewardLaw nu T, r t a ∈ Set.Icc (0 : ℝ) 1","missing":[],"search":"reward_coordinate_bounded banditrlproof.musicalchairs.reward_coordinate_bounded theorem reward_coordinate_bounded {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (t : fin t) (a : fin k) : ∀ᵐ r ∂rewardlaw nu t, r t a ∈ set.icc (0 : ℝ) 1 theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_subGaussian","label":"reward_coordinate_subGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.reward_coordinate_subGaussian","description":"theorem reward_coordinate_subGaussian {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (t : Fin T) (a : Fin k) : HasSubgaussianMGF (fun r : Fin T → Fin k → ℝ => r t a - armMean nu a) (1/4 : NNReal) (rewardLaw nu T)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-ad3e84a0a9ff","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1847,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem reward_coordinate_subGaussian {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (t : Fin T) (a : Fin k) : HasSubgaussianMGF (fun r : Fin T → Fin k → ℝ => r t a - armMean nu a) (1/4 : NNReal) (rewardLaw nu T)","missing":[],"search":"reward_coordinate_subgaussian banditrlproof.musicalchairs.reward_coordinate_subgaussian theorem reward_coordinate_subgaussian {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (t : fin t) (a : fin k) : hassubgaussianmgf (fun r : fin t → fin k → ℝ => r t a - armmean nu a) (1/4 : nnreal) (rewardlaw nu t) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.reward_time_independent","label":"reward_time_independent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.reward_time_independent","description":"theorem reward_time_independent {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (a : Fin k) : iIndepFun (fun t (r : Fin T → Fin k → ℝ) => r t a - armMean nu a) (rewardLaw nu T)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-b7c041f7dcc3","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1848,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem reward_time_independent {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (a : Fin k) : iIndepFun (fun t (r : Fin T → Fin k → ℝ) => r t a - armMean nu a) (rewardLaw nu T)","missing":[],"search":"reward_time_independent banditrlproof.musicalchairs.reward_time_independent theorem reward_time_independent {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (a : fin k) : iindepfun (fun t (r : fin t → fin k → ℝ) => r t a - armmean nu a) (rewardlaw nu t) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.selected_sum_subGaussian","label":"selected_sum_subGaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.selected_sum_subGaussian","description":"theorem selected_sum_subGaussian {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) : HasSubgaussianMGF (fun r : Fin T → Fin k → ℝ => ∑ t ∈ S, (r t a - armMean nu a)) ((S.card : NNReal)/4) (rewardLaw nu T)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-f74648ad013d","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1849,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem selected_sum_subGaussian {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) : HasSubgaussianMGF (fun r : Fin T → Fin k → ℝ => ∑ t ∈ S, (r t a - armMean nu a)) ((S.card : NNReal)/4) (rewardLaw nu T)","missing":[],"search":"selected_sum_subgaussian banditrlproof.musicalchairs.selected_sum_subgaussian theorem selected_sum_subgaussian {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin t)) (a : fin k) : hassubgaussianmgf (fun r : fin t → fin k → ℝ => ∑ t ∈ s, (r t a - armmean nu a)) ((s.card : nnreal)/4) (rewardlaw nu t) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.selected_sum_abs_tail","label":"selected_sum_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.selected_sum_abs_tail","description":"theorem selected_sum_abs_tail {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) (u : ℝ) (hu : 0 ≤ u) : (rewardLaw nu T).real {r | u ≤ |∑ t ∈ S, (r t a - armMean nu a)|} ≤ 2 * Real.exp (-u^2 / (2 * ((S.card : ℝ)/4)))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-437a7786966f","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1850,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem selected_sum_abs_tail {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) (u : ℝ) (hu : 0 ≤ u) : (rewardLaw nu T).real {r | u ≤ |∑ t ∈ S, (r t a - armMean nu a)|} ≤ 2 * Real.exp (-u^2 / (2 * ((S.card : ℝ)/4)))","missing":[],"search":"selected_sum_abs_tail banditrlproof.musicalchairs.selected_sum_abs_tail theorem selected_sum_abs_tail {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin t)) (a : fin k) (u : ℝ) (hu : 0 ≤ u) : (rewardlaw nu t).real {r | u ≤ |∑ t ∈ s, (r t a - armmean nu a)|} ≤ 2 * real.exp (-u^2 / (2 * ((s.card : ℝ)/4))) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.selectedMean","label":"selectedMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.selectedMean","description":"noncomputable def selectedMean {k T : ℕ} (S : Finset (Fin T)) (a : Fin k) (r : Fin T → Fin k → ℝ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-f6ae76d6b6b3","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1851,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def selectedMean {k T : ℕ} (S : Finset (Fin T)) (a : Fin k) (r : Fin T → Fin k → ℝ) : ℝ","missing":[],"search":"selectedmean banditrlproof.musicalchairs.selectedmean noncomputable def selectedmean {k t : ℕ} (s : finset (fin t)) (a : fin k) (r : fin t → fin k → ℝ) : ℝ definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.selected_centered_sum","label":"selected_centered_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.selected_centered_sum","description":"theorem selected_centered_sum {k T : ℕ} (nu : Fin k → Measure ℝ) (S : Finset (Fin T)) (a : Fin k) (hS : 0 < S.card) (r : Fin T → Fin k → ℝ) : (∑ t ∈ S, (r t a - armMean nu a)) = (S.card : ℝ) * (selectedMean S a r - armMean nu a)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-dba76c18ad58","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1852,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem selected_centered_sum {k T : ℕ} (nu : Fin k → Measure ℝ) (S : Finset (Fin T)) (a : Fin k) (hS : 0 < S.card) (r : Fin T → Fin k → ℝ) : (∑ t ∈ S, (r t a - armMean nu a)) = (S.card : ℝ) * (selectedMean S a r - armMean nu a)","missing":[],"search":"selected_centered_sum banditrlproof.musicalchairs.selected_centered_sum theorem selected_centered_sum {k t : ℕ} (nu : fin k → measure ℝ) (s : finset (fin t)) (a : fin k) (hs : 0 < s.card) (r : fin t → fin k → ℝ) : (∑ t ∈ s, (r t a - armmean nu a)) = (s.card : ℝ) * (selectedmean s a r - armmean nu a) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.selectedMean_tail_pos","label":"selectedMean_tail_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.selectedMean_tail_pos","description":"theorem selectedMean_tail_pos {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) (hS : 0 < S.card) (eps : ℝ) (heps : 0 ≤ eps) : (rewardLaw nu T).real {r | eps/2 ≤ |selectedMean S a r - armMean nu a|} ≤ 2 * Real.exp (-(S.card : ℝ) * eps^2 / 2)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-abc595646070","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1853,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem selectedMean_tail_pos {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) (hS : 0 < S.card) (eps : ℝ) (heps : 0 ≤ eps) : (rewardLaw nu T).real {r | eps/2 ≤ |selectedMean S a r - armMean nu a|} ≤ 2 * Real.exp (-(S.card : ℝ) * eps^2 / 2)","missing":[],"search":"selectedmean_tail_pos banditrlproof.musicalchairs.selectedmean_tail_pos theorem selectedmean_tail_pos {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin t)) (a : fin k) (hs : 0 < s.card) (eps : ℝ) (heps : 0 ≤ eps) : (rewardlaw nu t).real {r | eps/2 ≤ |selectedmean s a r - armmean nu a|} ≤ 2 * real.exp (-(s.card : ℝ) * eps^2 / 2) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.selectedMean_tail","label":"selectedMean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.selectedMean_tail","description":"theorem selectedMean_tail {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) (eps : ℝ) (heps : 0 ≤ eps) : (rewardLaw nu T).real {r | eps/2 ≤ |selectedMean S a r - armMean nu a|} ≤ 2 * Real.exp (-(S.card : ℝ) * eps^2 / 2)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-1ea51bbebdb8","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1854,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem selectedMean_tail {k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (S : Finset (Fin T)) (a : Fin k) (eps : ℝ) (heps : 0 ≤ eps) : (rewardLaw nu T).real {r | eps/2 ≤ |selectedMean S a r - armMean nu a|} ≤ 2 * Real.exp (-(S.card : ℝ) * eps^2 / 2)","missing":[],"search":"selectedmean_tail banditrlproof.musicalchairs.selectedmean_tail theorem selectedmean_tail {k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (s : finset (fin t)) (a : fin k) (eps : ℝ) (heps : 0 ≤ eps) : (rewardlaw nu t).real {r | eps/2 ≤ |selectedmean s a r - armmean nu a|} ≤ 2 * real.exp (-(s.card : ℝ) * eps^2 / 2) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.observedTimes","label":"observedTimes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.observedTimes","description":"noncomputable def observedTimes {n k T : ℕ} (x : Fin T → Fin n → Fin k) (i : Fin n) (a : Fin k) : Finset (Fin T)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-407e1ac83266","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1855,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def observedTimes {n k T : ℕ} (x : Fin T → Fin n → Fin k) (i : Fin n) (a : Fin k) : Finset (Fin T)","missing":[],"search":"observedtimes banditrlproof.musicalchairs.observedtimes noncomputable def observedtimes {n k t : ℕ} (x : fin t → fin n → fin k) (i : fin n) (a : fin k) : finset (fin t) definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localEmpiricalMean_eq_selectedMean","label":"localEmpiricalMean_eq_selectedMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localEmpiricalMean_eq_selectedMean","description":"theorem localEmpiricalMean_eq_selectedMean {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) (a : Fin k) : localEmpiricalMean (explorationFeedback x r i) a = selectedMean (observedTimes x i a) a r","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-a91caf77e4a0","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1856,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:146"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localEmpiricalMean_eq_selectedMean {n k T : ℕ} (x : Fin T → Fin n → Fin k) (r : Fin T → Fin k → ℝ) (i : Fin n) (a : Fin k) : localEmpiricalMean (explorationFeedback x r i) a = selectedMean (observedTimes x i a) a r","missing":[],"search":"localempiricalmean_eq_selectedmean banditrlproof.musicalchairs.localempiricalmean_eq_selectedmean theorem localempiricalmean_eq_selectedmean {n k t : ℕ} (x : fin t → fin n → fin k) (r : fin t → fin k → ℝ) (i : fin n) (a : fin k) : localempiricalmean (explorationfeedback x r i) a = selectedmean (observedtimes x i a) a r theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.localEmpiricalMean_fixed_tail","label":"localEmpiricalMean_fixed_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.localEmpiricalMean_fixed_tail","description":"theorem localEmpiricalMean_fixed_tail {n k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (x : Fin T → Fin n → Fin k) (i : Fin n) (a : Fin k) (eps : ℝ) (heps : 0 ≤ eps) : (rewardLaw nu T).real {r | eps/2 ≤ |localEmpiricalMean (explorationFeedback x r i) a - armMean nu a|} ≤ 2 * Real.exp (-(observationCount i a x : ℝ) * eps^2 / 2)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-af0e8901f39c","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1857,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:152"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem localEmpiricalMean_fixed_tail {n k T : ℕ} (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (x : Fin T → Fin n → Fin k) (i : Fin n) (a : Fin k) (eps : ℝ) (heps : 0 ≤ eps) : (rewardLaw nu T).real {r | eps/2 ≤ |localEmpiricalMean (explorationFeedback x r i) a - armMean nu a|} ≤ 2 * Real.exp (-(observationCount i a x : ℝ) * eps^2 / 2)","missing":[],"search":"localempiricalmean_fixed_tail banditrlproof.musicalchairs.localempiricalmean_fixed_tail theorem localempiricalmean_fixed_tail {n k t : ℕ} (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (x : fin t → fin n → fin k) (i : fin n) (a : fin k) (eps : ℝ) (heps : 0 ≤ eps) : (rewardlaw nu t).real {r | eps/2 ≤ |localempiricalmean (explorationfeedback x r i) a - armmean nu a|} ≤ 2 * real.exp (-(observationcount i a x : ℝ) * eps^2 / 2) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_selectedMean","label":"measurable_selectedMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_selectedMean","description":"theorem measurable_selectedMean {k T : ℕ} (S : Finset (Fin T)) (a : Fin k) : Measurable (selectedMean S a)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-fb3daf22c56d","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1858,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:161"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_selectedMean {k T : ℕ} (S : Finset (Fin T)) (a : Fin k) : Measurable (selectedMean S a)","missing":[],"search":"measurable_selectedmean banditrlproof.musicalchairs.measurable_selectedmean theorem measurable_selectedmean {k t : ℕ} (s : finset (fin t)) (a : fin k) : measurable (selectedmean s a) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurable_jointEmpiricalMean","label":"measurable_jointEmpiricalMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurable_jointEmpiricalMean","description":"theorem measurable_jointEmpiricalMean {n k T : ℕ} (i : Fin n) (a : Fin k) : Measurable (fun z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ) => localEmpiricalMean (explorationFeedback z.1 z.2 i) a)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-c0ac53176491","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1859,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurable_jointEmpiricalMean {n k T : ℕ} (i : Fin n) (a : Fin k) : Measurable (fun z : (Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ) => localEmpiricalMean (explorationFeedback z.1 z.2 i) a)","missing":[],"search":"measurable_jointempiricalmean banditrlproof.musicalchairs.measurable_jointempiricalmean theorem measurable_jointempiricalmean {n k t : ℕ} (i : fin n) (a : fin k) : measurable (fun z : (fin t → fin n → fin k) × (fin t → fin k → ℝ) => localempiricalmean (explorationfeedback z.1 z.2 i) a) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationRewardLaw","label":"explorationRewardLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationRewardLaw","description":"noncomputable def explorationRewardLaw {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) : Measure ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-5fa81959027f","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1860,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationRewardLaw {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) : Measure ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"explorationrewardlaw banditrlproof.musicalchairs.explorationrewardlaw noncomputable def explorationrewardlaw {n k t : ℕ} (hk : 0 < k) (nu : fin k → measure ℝ) : measure ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.meanBadEvent","label":"meanBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.meanBadEvent","description":"def meanBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (i : Fin n) (a : Fin k) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-3d23f33c30cc","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1861,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:185"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def meanBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (i : Fin n) (a : Fin k) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"meanbadevent banditrlproof.musicalchairs.meanbadevent def meanbadevent {n k t : ℕ} (nu : fin k → measure ℝ) (i : fin n) (a : fin k) (eps : ℝ) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.measurableSet_meanBadEvent","label":"measurableSet_meanBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.measurableSet_meanBadEvent","description":"theorem measurableSet_meanBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (i : Fin n) (a : Fin k) (eps : ℝ) : MeasurableSet (meanBadEvent (T := T) nu i a eps)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-70d1a908bd5b","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1862,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:189"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem measurableSet_meanBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (i : Fin n) (a : Fin k) (eps : ℝ) : MeasurableSet (meanBadEvent (T := T) nu i a eps)","missing":[],"search":"measurableset_meanbadevent banditrlproof.musicalchairs.measurableset_meanbadevent theorem measurableset_meanbadevent {n k t : ℕ} (nu : fin k → measure ℝ) (i : fin n) (a : fin k) (eps : ℝ) : measurableset (meanbadevent (t := t) nu i a eps) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.meanBadEvent_mixture","label":"meanBadEvent_mixture","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.meanBadEvent_mixture","description":"theorem meanBadEvent_mixture {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (i : Fin n) (a : Fin k) (eps : ℝ) : explorationRewardLaw hk nu (meanBadEvent (T := T) nu i a eps) = ∑ x, explorationLaw n k T hk x * rewardLaw nu T {r | eps/2 ≤ |localEmpiricalMean (explorationFeedback x r i) a - armMean nu a|}","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-9e6461234765","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1863,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:193"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem meanBadEvent_mixture {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (i : Fin n) (a : Fin k) (eps : ℝ) : explorationRewardLaw hk nu (meanBadEvent (T := T) nu i a eps) = ∑ x, explorationLaw n k T hk x * rewardLaw nu T {r | eps/2 ≤ |localEmpiricalMean (explorationFeedback x r i) a - armMean nu a|}","missing":[],"search":"meanbadevent_mixture banditrlproof.musicalchairs.meanbadevent_mixture theorem meanbadevent_mixture {n k t : ℕ} (hk : 0 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (i : fin n) (a : fin k) (eps : ℝ) : explorationrewardlaw hk nu (meanbadevent (t := t) nu i a eps) = ∑ x, explorationlaw n k t hk x * rewardlaw nu t {r | eps/2 ≤ |localempiricalmean (explorationfeedback x r i) a - armmean nu a|} theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.meanBadEvent_count_bound","label":"meanBadEvent_count_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.meanBadEvent_count_bound","description":"theorem meanBadEvent_count_bound {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (i : Fin n) (a : Fin k) (eps : ℝ) (heps : 0 ≤ eps) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) explorationRewardLaw hk nu (meanBadEvent (T := T) nu i a eps) ≤ 2 * (q * ENNReal.ofReal (Real.exp (-(eps^2/2))) + (1-q))^T","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-b5fd9ec54b85","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1864,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:205"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem meanBadEvent_count_bound {n k T : ℕ} (hk : 0 < k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (i : Fin n) (a : Fin k) (eps : ℝ) (heps : 0 ≤ eps) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) explorationRewardLaw hk nu (meanBadEvent (T := T) nu i a eps) ≤ 2 * (q * ENNReal.ofReal (Real.exp (-(eps^2/2))) + (1-q))^T","missing":[],"search":"meanbadevent_count_bound banditrlproof.musicalchairs.meanbadevent_count_bound theorem meanbadevent_count_bound {n k t : ℕ} (hk : 0 < k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (i : fin n) (a : fin k) (eps : ℝ) (heps : 0 ≤ eps) : let q := (1 / (k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞) / (k : ℝ≥0∞))^(n-1) explorationrewardlaw hk nu (meanbadevent (t := t) nu i a eps) ≤ 2 * (q * ennreal.ofreal (real.exp (-(eps^2/2))) + (1-q))^t theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.half_le_one_sub_exp_neg","label":"half_le_one_sub_exp_neg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.half_le_one_sub_exp_neg","description":"theorem half_le_one_sub_exp_neg (x : ℝ) (hx0 : 0 ≤ x) (hx1 : x ≤ 1) : x/2 ≤ 1 - Real.exp (-x)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-9845f5d7df31","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1865,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:235"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem half_le_one_sub_exp_neg (x : ℝ) (hx0 : 0 ≤ x) (hx1 : x ≤ 1) : x/2 ≤ 1 - Real.exp (-x)","missing":[],"search":"half_le_one_sub_exp_neg banditrlproof.musicalchairs.half_le_one_sub_exp_neg theorem half_le_one_sub_exp_neg (x : ℝ) (hx0 : 0 ≤ x) (hx1 : x ≤ 1) : x/2 ≤ 1 - real.exp (-x) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.count_mixture_exp_bound","label":"count_mixture_exp_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.count_mixture_exp_bound","description":"theorem count_mixture_exp_bound (q eps : ℝ) (T : ℕ) (hq0 : 0 ≤ q) (hq1 : q ≤ 1) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : (q * Real.exp (-(eps^2/2)) + (1-q))^T ≤ Real.exp (-(T : ℝ)*q*eps^2/4)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-d76030750ce5","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1866,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:250"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem count_mixture_exp_bound (q eps : ℝ) (T : ℕ) (hq0 : 0 ≤ q) (hq1 : q ≤ 1) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : (q * Real.exp (-(eps^2/2)) + (1-q))^T ≤ Real.exp (-(T : ℝ)*q*eps^2/4)","missing":[],"search":"count_mixture_exp_bound banditrlproof.musicalchairs.count_mixture_exp_bound theorem count_mixture_exp_bound (q eps : ℝ) (t : ℕ) (hq0 : 0 ≤ q) (hq1 : q ≤ 1) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : (q * real.exp (-(eps^2/2)) + (1-q))^t ≤ real.exp (-(t : ℝ)*q*eps^2/4) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationProbReal","label":"explorationProbReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationProbReal","description":"noncomputable def explorationProbReal (n k : ℕ) : ℝ","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-a921e7bc42fa","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1867,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:271"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"noncomputable def explorationProbReal (n k : ℕ) : ℝ","missing":[],"search":"explorationprobreal banditrlproof.musicalchairs.explorationprobreal noncomputable def explorationprobreal (n k : ℕ) : ℝ definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationProbReal_nonneg","label":"explorationProbReal_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationProbReal_nonneg","description":"theorem explorationProbReal_nonneg (n k : ℕ) : 0 ≤ explorationProbReal n k","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-20b3bee44f8b","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1868,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:274"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationProbReal_nonneg (n k : ℕ) : 0 ≤ explorationProbReal n k","missing":[],"search":"explorationprobreal_nonneg banditrlproof.musicalchairs.explorationprobreal_nonneg theorem explorationprobreal_nonneg (n k : ℕ) : 0 ≤ explorationprobreal n k theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationProbReal_le_one","label":"explorationProbReal_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationProbReal_le_one","description":"theorem explorationProbReal_le_one (n k : ℕ) (hk : 0 < k) : explorationProbReal n k ≤ 1","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-0405f8e276f1","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1869,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:278"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationProbReal_le_one (n k : ℕ) (hk : 0 < k) : explorationProbReal n k ≤ 1","missing":[],"search":"explorationprobreal_le_one banditrlproof.musicalchairs.explorationprobreal_le_one theorem explorationprobreal_le_one (n k : ℕ) (hk : 0 < k) : explorationprobreal n k ≤ 1 theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.explorationProbReal_lower","label":"explorationProbReal_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.explorationProbReal_lower","description":"theorem explorationProbReal_lower (n k : ℕ) (hk : 0 < k) (hnk : n ≤ k) : (1 : ℝ)/(4*k) ≤ explorationProbReal n k","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-4d2b92b0d019","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1870,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:292"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem explorationProbReal_lower (n k : ℕ) (hk : 0 < k) (hnk : n ≤ k) : (1 : ℝ)/(4*k) ≤ explorationProbReal n k","missing":[],"search":"explorationprobreal_lower banditrlproof.musicalchairs.explorationprobreal_lower theorem explorationprobreal_lower (n k : ℕ) (hk : 0 < k) (hnk : n ≤ k) : (1 : ℝ)/(4*k) ≤ explorationprobreal n k theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.ofReal_explorationProbReal","label":"ofReal_explorationProbReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.ofReal_explorationProbReal","description":"theorem ofReal_explorationProbReal (n k : ℕ) (hk : 0 < k) : ENNReal.ofReal (explorationProbReal n k) = (1/(k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞)/(k : ℝ≥0∞))^(n-1)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-f728f8932c94","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1871,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:303"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem ofReal_explorationProbReal (n k : ℕ) (hk : 0 < k) : ENNReal.ofReal (explorationProbReal n k) = (1/(k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞)/(k : ℝ≥0∞))^(n-1)","missing":[],"search":"ofreal_explorationprobreal banditrlproof.musicalchairs.ofreal_explorationprobreal theorem ofreal_explorationprobreal (n k : ℕ) (hk : 0 < k) : ennreal.ofreal (explorationprobreal n k) = (1/(k : ℝ≥0∞)) * (((k-1 : ℕ) : ℝ≥0∞)/(k : ℝ≥0∞))^(n-1) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.meanBadEvent_exponential_bound","label":"meanBadEvent_exponential_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.meanBadEvent_exponential_bound","description":"theorem meanBadEvent_exponential_bound {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (i : Fin n) (a : Fin k) (eps : ℝ) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : explorationRewardLaw hk nu (meanBadEvent (T := T) nu i a eps) ≤ ENNReal.ofReal (2 * Real.exp (-(T : ℝ)*eps^2/(16*k)))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-43711adfc71b","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1872,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:312"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem meanBadEvent_exponential_bound {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (i : Fin n) (a : Fin k) (eps : ℝ) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : explorationRewardLaw hk nu (meanBadEvent (T := T) nu i a eps) ≤ ENNReal.ofReal (2 * Real.exp (-(T : ℝ)*eps^2/(16*k)))","missing":[],"search":"meanbadevent_exponential_bound banditrlproof.musicalchairs.meanbadevent_exponential_bound theorem meanbadevent_exponential_bound {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (i : fin n) (a : fin k) (eps : ℝ) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : explorationrewardlaw hk nu (meanbadevent (t := t) nu i a eps) ≤ ennreal.ofreal (2 * real.exp (-(t : ℝ)*eps^2/(16*k))) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allMeanBadEvent","label":"allMeanBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allMeanBadEvent","description":"def allMeanBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-4003690d67bf","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1873,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:350"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def allMeanBadEvent {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"allmeanbadevent banditrlproof.musicalchairs.allmeanbadevent def allmeanbadevent {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allMeanBadEvent_exponential_bound","label":"allMeanBadEvent_exponential_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allMeanBadEvent_exponential_bound","description":"theorem allMeanBadEvent_exponential_bound {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps : ℝ) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : explorationRewardLaw hk nu (allMeanBadEvent (n := n) (T := T) nu eps) ≤ ENNReal.ofReal (2 * (k : ℝ)^2 * Real.exp (-(T : ℝ)*eps^2/(16*k)))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-997b108d2d0f","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1874,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:354"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allMeanBadEvent_exponential_bound {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps : ℝ) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : explorationRewardLaw hk nu (allMeanBadEvent (n := n) (T := T) nu eps) ≤ ENNReal.ofReal (2 * (k : ℝ)^2 * Real.exp (-(T : ℝ)*eps^2/(16*k)))","missing":[],"search":"allmeanbadevent_exponential_bound banditrlproof.musicalchairs.allmeanbadevent_exponential_bound theorem allmeanbadevent_exponential_bound {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps : ℝ) (heps0 : 0 ≤ eps) (heps1 : eps ≤ 1) : explorationrewardlaw hk nu (allmeanbadevent (n := n) (t := t) nu eps) ≤ ennreal.ofreal (2 * (k : ℝ)^2 * real.exp (-(t : ℝ)*eps^2/(16*k))) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.mean_exploration_threshold","label":"mean_exploration_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.mean_exploration_threshold","description":"theorem mean_exploration_threshold (k T : ℕ) (hk : 0 < k) (eps delta : ℝ) (heps : 0 < eps) (hdelta : 0 < delta) (hT : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) : 2*(k : ℝ)^2 * Real.exp (-(T : ℝ)*eps^2/(16*k)) ≤ delta/2","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-c4c2831eafca","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1875,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:385"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem mean_exploration_threshold (k T : ℕ) (hk : 0 < k) (eps delta : ℝ) (heps : 0 < eps) (hdelta : 0 < delta) (hT : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) : 2*(k : ℝ)^2 * Real.exp (-(T : ℝ)*eps^2/(16*k)) ≤ delta/2","missing":[],"search":"mean_exploration_threshold banditrlproof.musicalchairs.mean_exploration_threshold theorem mean_exploration_threshold (k t : ℕ) (hk : 0 < k) (eps delta : ℝ) (heps : 0 < eps) (hdelta : 0 < delta) (ht : (16*(k : ℝ)/eps^2) * real.log (4*(k : ℝ)^2/delta) ≤ t) : 2*(k : ℝ)^2 * real.exp (-(t : ℝ)*eps^2/(16*k)) ≤ delta/2 theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allMeanBadEvent_le_half_delta","label":"allMeanBadEvent_le_half_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allMeanBadEvent_le_half_delta","description":"theorem allMeanBadEvent_le_half_delta {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) : explorationRewardLaw hk nu (allMeanBadEvent (n := n) (T := T) nu eps) ≤ ENNReal.ofReal (delta/2)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-387f581d8cd7","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1876,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:406"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allMeanBadEvent_le_half_delta {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) : explorationRewardLaw hk nu (allMeanBadEvent (n := n) (T := T) nu eps) ≤ ENNReal.ofReal (delta/2)","missing":[],"search":"allmeanbadevent_le_half_delta banditrlproof.musicalchairs.allmeanbadevent_le_half_delta theorem allmeanbadevent_le_half_delta {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (ht : (16*(k : ℝ)/eps^2) * real.log (4*(k : ℝ)^2/delta) ≤ t) : explorationrewardlaw hk nu (allmeanbadevent (n := n) (t := t) nu eps) ≤ ennreal.ofreal (delta/2) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allMeanAccurate","label":"allMeanAccurate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allMeanAccurate","description":"def allMeanAccurate {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-6c6473b27913","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1877,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:418"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"def allMeanAccurate {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : Set ((Fin T → Fin n → Fin k) × (Fin T → Fin k → ℝ))","missing":[],"search":"allmeanaccurate banditrlproof.musicalchairs.allmeanaccurate def allmeanaccurate {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : set ((fin t → fin n → fin k) × (fin t → fin k → ℝ)) definition compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allMeanAccurate_eq_compl","label":"allMeanAccurate_eq_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allMeanAccurate_eq_compl","description":"theorem allMeanAccurate_eq_compl {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : allMeanAccurate (n := n) (T := T) nu eps = (allMeanBadEvent nu eps)ᶜ","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-8724b3f65142","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1878,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:423"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allMeanAccurate_eq_compl {n k T : ℕ} (nu : Fin k → Measure ℝ) (eps : ℝ) : allMeanAccurate (n := n) (T := T) nu eps = (allMeanBadEvent nu eps)ᶜ","missing":[],"search":"allmeanaccurate_eq_compl banditrlproof.musicalchairs.allmeanaccurate_eq_compl theorem allmeanaccurate_eq_compl {n k t : ℕ} (nu : fin k → measure ℝ) (eps : ℝ) : allmeanaccurate (n := n) (t := t) nu eps = (allmeanbadevent nu eps)ᶜ theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.MusicalChairs.allMeanAccurate_probability","label":"allMeanAccurate_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MusicalChairs.allMeanAccurate_probability","description":"theorem allMeanAccurate_probability {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) : 1 - ENNReal.ofReal (delta/2) ≤ explorationRewardLaw hk nu (allMeanAccurate (n := n) (T := T) nu eps)","url":"../modules/banditrlproof-algorithms-musicalchairsreward/index.html#decl-899e2b598b8c","parent":"module:BanditRLProof.Algorithms.MusicalChairsReward","order":1879,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.MusicalChairsReward"],["Source","BanditRLProof/Algorithms/MusicalChairsReward.lean:428"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","multi-agent"]],"statement":"theorem allMeanAccurate_probability {n k T : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : Fin k → Measure ℝ) [∀ a, IsProbabilityMeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ Set.Icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (hT : (16*(k : ℝ)/eps^2) * Real.log (4*(k : ℝ)^2/delta) ≤ T) : 1 - ENNReal.ofReal (delta/2) ≤ explorationRewardLaw hk nu (allMeanAccurate (n := n) (T := T) nu eps)","missing":[],"search":"allmeanaccurate_probability banditrlproof.musicalchairs.allmeanaccurate_probability theorem allmeanaccurate_probability {n k t : ℕ} (hk : 0 < k) (hnk : n ≤ k) (nu : fin k → measure ℝ) [∀ a, isprobabilitymeasure (nu a)] (hb : ∀ a, ∀ᵐ y ∂nu a, y ∈ set.icc (0 : ℝ) 1) (eps delta : ℝ) (heps : 0 < eps) (heps1 : eps ≤ 1) (hdelta : 0 < delta) (ht : (16*(k : ℝ)/eps^2) * real.log (4*(k : ℝ)^2/delta) ≤ t) : 1 - ennreal.ofreal (delta/2) ≤ explorationrewardlaw hk nu (allmeanaccurate (n := n) (t := t) nu eps) theorem compiled","shard":"modules/b606c5c3a46ea860.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["multi-agent"]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxDenominator","label":"softmaxDenominator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxDenominator","description":"The denominator in the source softmax rule, Equation (3).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-de7fcee89e25","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1880,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:27"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def softmaxDenominator (theta : Action -> Real) : Real","missing":[],"search":"softmaxdenominator banditrlproof.stochasticgradientbandit.softmaxdenominator the denominator in the source softmax rule, equation (3). definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability","label":"softmaxProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability","description":"The source softmax sampling probability, Equation (3).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-76040a968987","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1881,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:31"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def softmaxProbability (theta : Action -> Real) (a : Action) : Real","missing":[],"search":"softmaxprobability banditrlproof.stochasticgradientbandit.softmaxprobability the source softmax sampling probability, equation (3). definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxDenominator_pos","label":"softmaxDenominator_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxDenominator_pos","description":"theorem softmaxDenominator_pos [Nonempty Action] (theta : Action -> Real) : 0 < softmaxDenominator theta","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-a0efcdae77a1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1882,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:35"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxDenominator_pos [Nonempty Action] (theta : Action -> Real) : 0 < softmaxDenominator theta","missing":[],"search":"softmaxdenominator_pos banditrlproof.stochasticgradientbandit.softmaxdenominator_pos theorem softmaxdenominator_pos [nonempty action] (theta : action -> real) : 0 < softmaxdenominator theta theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_pos","label":"softmaxProbability_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_pos","description":"theorem softmaxProbability_pos [Nonempty Action] (theta : Action -> Real) (a : Action) : 0 < softmaxProbability theta a","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-038e4fe70818","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1883,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:43"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_pos [Nonempty Action] (theta : Action -> Real) (a : Action) : 0 < softmaxProbability theta a","missing":[],"search":"softmaxprobability_pos banditrlproof.stochasticgradientbandit.softmaxprobability_pos theorem softmaxprobability_pos [nonempty action] (theta : action -> real) (a : action) : 0 < softmaxprobability theta a theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_nonneg","label":"softmaxProbability_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_nonneg","description":"theorem softmaxProbability_nonneg [Nonempty Action] (theta : Action -> Real) (a : Action) : 0 <= softmaxProbability theta a","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-1df76028de84","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1884,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:48"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_nonneg [Nonempty Action] (theta : Action -> Real) (a : Action) : 0 <= softmaxProbability theta a","missing":[],"search":"softmaxprobability_nonneg banditrlproof.stochasticgradientbandit.softmaxprobability_nonneg theorem softmaxprobability_nonneg [nonempty action] (theta : action -> real) (a : action) : 0 <= softmaxprobability theta a theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_sum","label":"softmaxProbability_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_sum","description":"Equation (3) defines a normalized finite sampling law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-b8691929b8b6","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1885,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:54"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_sum [Nonempty Action] (theta : Action -> Real) : ∑ a, softmaxProbability theta a = 1","missing":[],"search":"softmaxprobability_sum banditrlproof.stochasticgradientbandit.softmaxprobability_sum equation (3) defines a normalized finite sampling law. theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_le_one","label":"softmaxProbability_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_le_one","description":"theorem softmaxProbability_le_one [Nonempty Action] (theta : Action -> Real) (a : Action) : softmaxProbability theta a <= 1","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-9fcaf7bff75c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1886,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:61"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_le_one [Nonempty Action] (theta : Action -> Real) (a : Action) : softmaxProbability theta a <= 1","missing":[],"search":"softmaxprobability_le_one banditrlproof.stochasticgradientbandit.softmaxprobability_le_one theorem softmaxprobability_le_one [nonempty action] (theta : action -> real) (a : action) : softmaxprobability theta a <= 1 theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceIncrement","label":"sourceIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceIncrement","description":"Algorithm 1 / Equation (4), before multiplication by the learning rate.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-8e1110123de8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1887,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:72"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def sourceIncrement (p : Action -> Real) (reward : Real) (selected k : Action) : Real","missing":[],"search":"sourceincrement banditrlproof.stochasticgradientbandit.sourceincrement algorithm 1 / equation (4), before multiplication by the learning rate. definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceIncrement_eq_indicator","label":"sourceIncrement_eq_indicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceIncrement_eq_indicator","description":"theorem sourceIncrement_eq_indicator (p : Action -> Real) (reward : Real) (selected k : Action) : sourceIncrement p reward selected k = reward * ((if selected = k then 1 else 0) - p k)","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-1db4de62c3a3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1888,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:77"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceIncrement_eq_indicator (p : Action -> Real) (reward : Real) (selected k : Action) : sourceIncrement p reward selected k = reward * ((if selected = k then 1 else 0) - p k)","missing":[],"search":"sourceincrement_eq_indicator banditrlproof.stochasticgradientbandit.sourceincrement_eq_indicator theorem sourceincrement_eq_indicator (p : action -> real) (reward : real) (selected k : action) : sourceincrement p reward selected k = reward * ((if selected = k then 1 else 0) - p k) theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sum_sourceIncrement","label":"sum_sourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sum_sourceIncrement","description":"Algorithm 1 preserves the zero sum of its parameter vector.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-3dc8cc43b551","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1889,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:84"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sum_sourceIncrement (p : Action -> Real) (reward : Real) (selected : Action) (hp : ∑ k, p k = 1) : ∑ k, sourceIncrement p reward selected k = 0","missing":[],"search":"sum_sourceincrement banditrlproof.stochasticgradientbandit.sum_sourceincrement algorithm 1 preserves the zero sum of its parameter vector. theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.policyValue","label":"policyValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.policyValue","description":"The policy value at a fixed pre-action history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-929e9a20b4dc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1890,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:98"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def policyValue (p mean : Action -> Real) : Real","missing":[],"search":"policyvalue banditrlproof.stochasticgradientbandit.policyvalue the policy value at a fixed pre-action history. definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement","label":"expectedSourceIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expectedSourceIncrement","description":"The finite conditional-mean version of the source expected update.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-494a9d9743be","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1891,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:102"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def expectedSourceIncrement (p mean : Action -> Real) (k : Action) : Real","missing":[],"search":"expectedsourceincrement banditrlproof.stochasticgradientbandit.expectedsourceincrement the finite conditional-mean version of the source expected update. definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gradientCoordinate","label":"expectedSourceIncrement_eq_gradientCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gradientCoordinate","description":"Equation (5), in policy-gradient-coordinate form.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-ba0a711df1cc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1892,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:106"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expectedSourceIncrement_eq_gradientCoordinate (p mean : Action -> Real) (k : Action) : expectedSourceIncrement p mean k = p k * (mean k - policyValue p mean)","missing":[],"search":"expectedsourceincrement_eq_gradientcoordinate banditrlproof.stochasticgradientbandit.expectedsourceincrement_eq_gradientcoordinate equation (5), in policy-gradient-coordinate form. theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap","label":"instantaneousGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.instantaneousGap","description":"The source instantaneous expected gap `E_t[Delta_{A_t}]`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-577c7bfb8007","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1893,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:129"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def instantaneousGap (p gap : Action -> Real) : Real","missing":[],"search":"instantaneousgap banditrlproof.stochasticgradientbandit.instantaneousgap the source instantaneous expected gap `e_t[delta_{a_t}]`. definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_eq_bestMean_sub_policyValue","label":"instantaneousGap_eq_bestMean_sub_policyValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.instantaneousGap_eq_bestMean_sub_policyValue","description":"theorem instantaneousGap_eq_bestMean_sub_policyValue (p mean gap : Action -> Real) (bestMean : Real) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestMean - mean a) : instantaneousGap p gap = bestMean - policyValue p mean","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-107b95e163fb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1894,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:133"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem instantaneousGap_eq_bestMean_sub_policyValue (p mean gap : Action -> Real) (bestMean : Real) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestMean - mean a) : instantaneousGap p gap = bestMean - policyValue p mean","missing":[],"search":"instantaneousgap_eq_bestmean_sub_policyvalue banditrlproof.stochasticgradientbandit.instantaneousgap_eq_bestmean_sub_policyvalue theorem instantaneousgap_eq_bestmean_sub_policyvalue (p mean gap : action -> real) (bestmean : real) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestmean - mean a) : instantaneousgap p gap = bestmean - policyvalue p mean theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","label":"expectedSourceIncrement_eq_gapCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","description":"Equation (5), in instantaneous-gap-coordinate form.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-3930a668bbf0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1895,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:155"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expectedSourceIncrement_eq_gapCoordinate (p mean gap : Action -> Real) (bestMean : Real) (k : Action) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestMean - mean a) : expectedSourceIncrement p mean k = p k * (instantaneousGap p gap - gap k)","missing":[],"search":"expectedsourceincrement_eq_gapcoordinate banditrlproof.stochasticgradientbandit.expectedsourceincrement_eq_gapcoordinate equation (5), in instantaneous-gap-coordinate form. theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.gapExpectedIncrement","label":"gapExpectedIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.gapExpectedIncrement","description":"The gap-coordinate update isolated from Equation (5).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-8eb99e4fed46","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1896,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:167"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def gapExpectedIncrement (p gap : Action -> Real) (k : Action) : Real","missing":[],"search":"gapexpectedincrement banditrlproof.stochasticgradientbandit.gapexpectedincrement the gap-coordinate update isolated from equation (5). definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapExpectedIncrement","label":"expectedSourceIncrement_eq_gapExpectedIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapExpectedIncrement","description":"theorem expectedSourceIncrement_eq_gapExpectedIncrement (p mean gap : Action -> Real) (bestMean : Real) (k : Action) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestMean - mean a) : expectedSourceIncrement p mean k = gapExpectedIncrement p gap k","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-0cb541dae244","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1897,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:170"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expectedSourceIncrement_eq_gapExpectedIncrement (p mean gap : Action -> Real) (bestMean : Real) (k : Action) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestMean - mean a) : expectedSourceIncrement p mean k = gapExpectedIncrement p gap k","missing":[],"search":"expectedsourceincrement_eq_gapexpectedincrement banditrlproof.stochasticgradientbandit.expectedsourceincrement_eq_gapexpectedincrement theorem expectedsourceincrement_eq_gapexpectedincrement (p mean gap : action -> real) (bestmean : real) (k : action) (hp : ∑ a, p a = 1) (hgap : ∀ a, gap a = bestmean - mean a) : expectedsourceincrement p mean k = gapexpectedincrement p gap k theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_ge_minGap_mul_failureMass","label":"instantaneousGap_ge_minGap_mul_failureMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.instantaneousGap_ge_minGap_mul_failureMass","description":"theorem instantaneousGap_ge_minGap_mul_failureMass (p gap : Action -> Real) (best : Action) (Delta : Real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_min : ∀ a, a ≠ best -> Delta <= gap a) : Delta * (1 - p best) <= instantaneousGap p gap","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-1b64ece9e702","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1898,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:177"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem instantaneousGap_ge_minGap_mul_failureMass (p gap : Action -> Real) (best : Action) (Delta : Real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_min : ∀ a, a ≠ best -> Delta <= gap a) : Delta * (1 - p best) <= instantaneousGap p gap","missing":[],"search":"instantaneousgap_ge_mingap_mul_failuremass banditrlproof.stochasticgradientbandit.instantaneousgap_ge_mingap_mul_failuremass theorem instantaneousgap_ge_mingap_mul_failuremass (p gap : action -> real) (best : action) (delta : real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_min : ∀ a, a ≠ best -> delta <= gap a) : delta * (1 - p best) <= instantaneousgap p gap theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.gapExpectedIncrement_best_ge","label":"gapExpectedIncrement_best_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.gapExpectedIncrement_best_ge","description":"Pointwise Equation (6): the best coordinate gains at least the positive gap times its success/failure probability product.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-f83fcab6a228","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1899,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:214"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gapExpectedIncrement_best_ge (p gap : Action -> Real) (best : Action) (Delta : Real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_min : ∀ a, a ≠ best -> Delta <= gap a) : Delta * (p best * (1 - p best)) <= gapExpectedIncrement p gap best","missing":[],"search":"gapexpectedincrement_best_ge banditrlproof.stochasticgradientbandit.gapexpectedincrement_best_ge pointwise equation (6): the best coordinate gains at least the positive gap times its success/failure probability product. theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum","label":"bestParameterIncrementSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum","description":"The finite-horizon best-parameter expectation represented by Equation (6), after conditioning at each round.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-38cb5e956354","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1900,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:229"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def bestParameterIncrementSum (eta : Real) (p : Nat -> Action -> Real) (gap : Action -> Real) (best : Action) (horizon : Nat) : Real","missing":[],"search":"bestparameterincrementsum banditrlproof.stochasticgradientbandit.bestparameterincrementsum the finite-horizon best-parameter expectation represented by equation (6), after conditioning at each round. definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum_ge","label":"bestParameterIncrementSum_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum_ge","description":"Finite-horizon Equation (6).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-9fcf611d3903","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1901,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:234"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem bestParameterIncrementSum_ge (eta Delta : Real) (p : Nat -> Action -> Real) (gap : Action -> Real) (best : Action) (horizon : Nat) (heta : 0 <= eta) (hp : ∀ t, ∑ a, p t a = 1) (hp_nonneg : ∀ t a, 0 <= p t a) (hgap_best : gap best = 0) (hgap_min : ∀ a, a ≠ best -> Delta <= gap a) : eta * Delta * (∑ t ∈ Finset.range horizon, p t best * (1 - p t best)) <= bestParameterIncrementSum eta p gap best horizon","missing":[],"search":"bestparameterincrementsum_ge banditrlproof.stochasticgradientbandit.bestparameterincrementsum_ge finite-horizon equation (6). theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceExpectedPseudoRegret","label":"sourceExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceExpectedPseudoRegret","description":"The gap-weighted finite-horizon expected pseudo-regret from Equation (2).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-dc81dbd38bdb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1902,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:257"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def sourceExpectedPseudoRegret (p : Nat -> Action -> Real) (gap : Action -> Real) (horizon : Nat) : Real","missing":[],"search":"sourceexpectedpseudoregret banditrlproof.stochasticgradientbandit.sourceexpectedpseudoregret the gap-weighted finite-horizon expected pseudo-regret from equation (2). definition compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_le_maxGap_mul_failureMass","label":"instantaneousGap_le_maxGap_mul_failureMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.instantaneousGap_le_maxGap_mul_failureMass","description":"theorem instantaneousGap_le_maxGap_mul_failureMass (p gap : Action -> Real) (best : Action) (DeltaMax : Real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_max : ∀ a, a ≠ best -> gap a <= DeltaMax) : instantaneousGap p gap <= DeltaMax * (1 - p best)","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-80aaaa219b2d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1903,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:261"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem instantaneousGap_le_maxGap_mul_failureMass (p gap : Action -> Real) (best : Action) (DeltaMax : Real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_max : ∀ a, a ≠ best -> gap a <= DeltaMax) : instantaneousGap p gap <= DeltaMax * (1 - p best)","missing":[],"search":"instantaneousgap_le_maxgap_mul_failuremass banditrlproof.stochasticgradientbandit.instantaneousgap_le_maxgap_mul_failuremass theorem instantaneousgap_le_maxgap_mul_failuremass (p gap : action -> real) (best : action) (deltamax : real) (hp : ∑ a, p a = 1) (hp_nonneg : ∀ a, 0 <= p a) (hgap_best : gap best = 0) (hgap_max : ∀ a, a ≠ best -> gap a <= deltamax) : instantaneousgap p gap <= deltamax * (1 - p best) theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.failureMass_eq_successFailure_add_sq","label":"failureMass_eq_successFailure_add_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.failureMass_eq_successFailure_add_sq","description":"theorem failureMass_eq_successFailure_add_sq (x : Real) : 1 - x = x * (1 - x) + (1 - x) ^ 2","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-4d5f776320db","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1904,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:297"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem failureMass_eq_successFailure_add_sq (x : Real) : 1 - x = x * (1 - x) + (1 - x) ^ 2","missing":[],"search":"failuremass_eq_successfailure_add_sq banditrlproof.stochasticgradientbandit.failuremass_eq_successfailure_add_sq theorem failuremass_eq_successfailure_add_sq (x : real) : 1 - x = x * (1 - x) + (1 - x) ^ 2 theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","label":"sourceRegretDecomposition_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","description":"Equation (7), with its positive learning-rate/minimum-gap denominator and maximum-gap envelope exposed explicitly.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditaudit/index.html#decl-3f8721f7f9cd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","order":1905,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditAudit.lean:302"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceRegretDecomposition_le (eta Delta DeltaMax : Real) (p : Nat -> Action -> Real) (gap : Action -> Real) (best : Action) (horizon : Nat) (heta : 0 < eta) (hDelta : 0 < Delta) (hDeltaMax : 0 <= DeltaMax) (hp : ∀ t, ∑ a, p t a = 1) (hp_nonneg : ∀ t a, 0 <= p t a) (hgap_best : gap best = 0) (hgap_min : ∀ a, a ≠ best -> Delta <= gap a) (hgap_max : ∀ a, a ≠ best -> gap a <= DeltaMax) : sourceExpectedPseudoRegret p gap horizon <= (DeltaMax / (eta * Delta)) * bestParameterIncrementSum eta p gap best horizon + DeltaMax * (∑ t ∈ Finset.range horizon, (1 - p t best) ^ 2)","missing":[],"search":"sourceregretdecomposition_le banditrlproof.stochasticgradientbandit.sourceregretdecomposition_le equation (7), with its positive learning-rate/minimum-gap denominator and maximum-gap envelope exposed explicitly. theorem compiled","shard":"modules/ea45735030a545a9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_exp_actionReward_le_sourceEqEight_of_mean","label":"integral_measurableEnvironmentInitialPairKernel_exp_actionReward_le_sourceEqEight_of_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_exp_actionReward_le_sourceEqEight_of_mean","description":"Action-dependent Equation (8) for the generated initial action/reward pair. Source-time fence: this kernel samples source round `t = 1` from the untouched parameter `initialTheta`; consuming pair zero then constructs source `theta_2`. In the source Theorem 1, `initialTheta` is the zero vector. This base bridge is therefore separate from the successor bridge below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditconditionalexponentialaudit/index.html#decl-f12b2ee78530","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","order":1906,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditConditionalExponentialAudit.lean:56"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableEnvironmentInitialPairKernel_exp_actionReward_le_sourceEqEight_of_mean {Env : Type v} [MeasurableSpace Env] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (env : Env) (q mean : Action -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.initialFeedback (env, selected), |reward| <= 1) (hmean : forall selected, integral (environment.initialFeedback (env, selected)) id = mean selected) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm initialTheta eta) environment env) (fun pair : Action × Real => Real.exp (q pair.1 * pair.2)) <= 1 + ∑ selected, softmaxProbability initialTheta selected * (q selected * mean selected + q selected ^ 2 / 2 * sourceC (|q selected| / 2))","missing":[],"search":"integral_measurableenvironmentinitialpairkernel_exp_actionreward_le_sourceeqeight_of_mean banditrlproof.stochasticgradientbandit.integral_measurableenvironmentinitialpairkernel_exp_actionreward_le_sourceeqeight_of_mean action-dependent equation (8) for the generated initial action/reward pair. source-time fence: this kernel samples source round `t = 1` from the untouched parameter `initialtheta`; consuming pair zero then constructs source `theta_2`. in the source theorem 1, `initialtheta` is the zero vector. this base bridge is therefore separate from the successor bridge below. theorem compiled","shard":"modules/fc97f5393dd37dd1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight","label":"integral_historyStepKernel_exp_actionReward_le_sourceEqEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight","description":"Action-dependent Equation (8) under the actual generated SGB history-step kernel. The real reward coordinate is measurable by construction, so the only reward-law premise needed beyond the Markov-kernel contract is its almost-sure source support in `[-1, 1]`. Source-time fence: `historyParameter initialTheta eta n history` has already consumed trace pairs `0, ..., n`, so it is source `theta_{n+2}`. This successor ke…","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditconditionalexponentialaudit/index.html#decl-419d05052794","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","order":1907,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditConditionalExponentialAudit.lean:197"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_historyStepKernel_exp_actionReward_le_sourceEqEight (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.HistoryEnvironment Action Real) (n : Nat) (history : History.FinitePairHistory Action Real n) (q : Action -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (history, selected), |reward| <= 1) : integral (Thompson.historyStepKernel (historyAlgorithm initialTheta eta) environment n history) (fun pair : Action × Real => Real.exp (q pair.1 * pair.2)) <= 1 + ∑ selected, softmaxProbability (historyParameter initialTheta eta n history) selected * (q selected * integral (environment.feedback n (history, selected)) id + q selected ^ 2 / 2 * sourceC (|q selected| / 2))","missing":[],"search":"integral_historystepkernel_exp_actionreward_le_sourceeqeight banditrlproof.stochasticgradientbandit.integral_historystepkernel_exp_actionreward_le_sourceeqeight action-dependent equation (8) under the actual generated sgb history-step kernel. the real reward coordinate is measurable by construction, so the only reward-law premise needed beyond the markov-kernel contract is its almost-sure source support in `[-1, 1]`. source-time fence: `historyparameter initialtheta eta n history` has already consumed trace pairs `0, ..., n`, so it is source `theta_{n+2}`. this successor kernel samples source round `n+2`; consuming its new pair produces `theta_{n+3}`. consequently it does not replace the initial-pair bridge above when assembling a recurrence that begins at source `t = 1`. for the later `fin 2` consumer with best-arm probability `p`, the positive recurrence instantiates this generic coefficient by `q 0 = 2 * eta * (1 - p)` and `q 1 = -2 * eta * p`; the inverse-odds recurrence uses the pointwise negation. those recurrence simplifications are deliberately left to the next layer. theorem compiled","shard":"modules/fc97f5393dd37dd1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight_of_mean","label":"integral_historyStepKernel_exp_actionReward_le_sourceEqEight_of_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight_of_mean","description":"Mean-surface form of the generated action-dependent Equation-(8) bound.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditconditionalexponentialaudit/index.html#decl-401c3a2b46d5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","order":1908,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditConditionalExponentialAudit.lean:313"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_historyStepKernel_exp_actionReward_le_sourceEqEight_of_mean (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.HistoryEnvironment Action Real) (n : Nat) (history : History.FinitePairHistory Action Real n) (q mean : Action -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (history, selected), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (history, selected)) id = mean selected) : integral (Thompson.historyStepKernel (historyAlgorithm initialTheta eta) environment n history) (fun pair : Action × Real => Real.exp (q pair.1 * pair.2)) <= 1 + ∑ selected, softmaxProbability (historyParameter initialTheta eta n history) selected * (q selected * mean selected + q selected ^ 2 / 2 * sourceC (|q selected| / 2))","missing":[],"search":"integral_historystepkernel_exp_actionreward_le_sourceeqeight_of_mean banditrlproof.stochasticgradientbandit.integral_historystepkernel_exp_actionreward_le_sourceeqeight_of_mean mean-surface form of the generated action-dependent equation-(8) bound. theorem compiled","shard":"modules/fc97f5393dd37dd1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_exp_actionReward_le_sourceEqEight_of_mean","label":"integral_measurableEnvironmentHistoryStepKernel_exp_actionReward_le_sourceEqEight_of_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_exp_actionReward_le_sourceEqEight_of_mean","description":"Environment-indexed wrapper over the jointly measurable feedback kernel used by the canonical generated trajectory.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditconditionalexponentialaudit/index.html#decl-46e9a89fd7df","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","order":1909,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditConditionalExponentialAudit.lean:361"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableEnvironmentHistoryStepKernel_exp_actionReward_le_sourceEqEight_of_mean {Env : Type v} [MeasurableSpace Env] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (q mean : Action -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (env, (history, selected)), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (env, (history, selected))) id = mean selected) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) environment n (env, history)) (fun pair : Action × Real => Real.exp (q pair.1 * pair.2)) <= 1 + ∑ selected, softmaxProbability (historyParameter initialTheta eta n history) selected * (q selected * mean selected + q selected ^ 2 / 2 * source…","missing":[],"search":"integral_measurableenvironmenthistorystepkernel_exp_actionreward_le_sourceeqeight_of_mean banditrlproof.stochasticgradientbandit.integral_measurableenvironmenthistorystepkernel_exp_actionreward_le_sourceeqeight_of_mean environment-indexed wrapper over the jointly measurable feedback kernel used by the canonical generated trajectory. theorem compiled","shard":"modules/fc97f5393dd37dd1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmActionGap_le_gap","label":"twoArmActionGap_le_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmActionGap_le_gap","description":"Every realized two-arm gap is at most `Delta` when `Delta` is nonnegative.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-052a653ea0ee","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1910,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:36"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmActionGap_le_gap (Delta : Real) (hDelta : 0 <= Delta) (action : Fin 2) : twoArmActionGap Delta action <= Delta","missing":[],"search":"twoarmactiongap_le_gap banditrlproof.stochasticgradientbandit.twoarmactiongap_le_gap every realized two-arm gap is at most `delta` when `delta` is nonnegative. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_le_gap_mul_horizon","label":"twoArmSampledPseudoRegret_le_gap_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_le_gap_mul_horizon","description":"Pathwise trivial regret bound for the actions actually sampled by the generated process.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-6b8430d14f5c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1911,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:45"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSampledPseudoRegret_le_gap_mul_horizon {Env : Type v} (Delta : Real) (hDelta : 0 <= Delta) (horizon : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmSampledPseudoRegret Delta horizon sample <= Delta * (horizon : Real)","missing":[],"search":"twoarmsampledpseudoregret_le_gap_mul_horizon banditrlproof.stochasticgradientbandit.twoarmsampledpseudoregret_le_gap_mul_horizon pathwise trivial regret bound for the actions actually sampled by the generated process. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSampledPseudoRegret","label":"measurable_twoArmSampledPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmSampledPseudoRegret","description":"theorem measurable_twoArmSampledPseudoRegret {Env : Type v} [MeasurableSpace Env] (Delta : Real) (horizon : Nat) : Measurable (twoArmSampledPseudoRegret (Env := Env) Delta horizon)","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-c6c148a44e6f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1912,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:60"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmSampledPseudoRegret {Env : Type v} [MeasurableSpace Env] (Delta : Real) (horizon : Nat) : Measurable (twoArmSampledPseudoRegret (Env := Env) Delta horizon)","missing":[],"search":"measurable_twoarmsampledpseudoregret banditrlproof.stochasticgradientbandit.measurable_twoarmsampledpseudoregret theorem measurable_twoarmsampledpseudoregret {env : type v} [measurablespace env] (delta : real) (horizon : nat) : measurable (twoarmsampledpseudoregret (env := env) delta horizon) theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret","label":"integrable_twoArmSampledPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret","description":"theorem integrable_twoArmSampledPseudoRegret {Env : Type v} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) -> Fin 2 × Real))) [IsFiniteMeasure mu] (Delta : Real) (horizon : Nat) : Integrable (twoArmSampledPseudoRegret (Env := Env) Delta horizon) mu","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-426cb9576e85","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1913,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:66"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmSampledPseudoRegret {Env : Type v} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) -> Fin 2 × Real))) [IsFiniteMeasure mu] (Delta : Real) (horizon : Nat) : Integrable (twoArmSampledPseudoRegret (Env := Env) Delta horizon) mu","missing":[],"search":"integrable_twoarmsampledpseudoregret banditrlproof.stochasticgradientbandit.integrable_twoarmsampledpseudoregret theorem integrable_twoarmsampledpseudoregret {env : type v} [measurablespace env] (mu : measure (env × ((k : nat) -> fin 2 × real))) [isfinitemeasure mu] (delta : real) (horizon : nat) : integrable (twoarmsampledpseudoregret (env := env) delta horizon) mu theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_gap_mul_horizon","label":"integral_twoArmSampledPseudoRegret_le_gap_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_gap_mul_horizon","description":"Integral form of the pathwise `Delta * T` bound. The measure is the actual generated trajectory measure in the Corollary-1 consumer below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-96249b29b185","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1914,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:92"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmSampledPseudoRegret_le_gap_mul_horizon {Env : Type v} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) -> Fin 2 × Real))) [IsProbabilityMeasure mu] (Delta : Real) (hDelta : 0 <= Delta) (horizon : Nat) : integral mu (twoArmSampledPseudoRegret (Env := Env) Delta horizon) <= Delta * (horizon : Real)","missing":[],"search":"integral_twoarmsampledpseudoregret_le_gap_mul_horizon banditrlproof.stochasticgradientbandit.integral_twoarmsampledpseudoregret_le_gap_mul_horizon integral form of the pathwise `delta * t` bound. the measure is the actual generated trajectory measure in the corollary-1 consumer below. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta","label":"corollaryOneEta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneEta","description":"The horizon-indexed fixed learning rate used in Corollary 1.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-90eb7cc910e8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1915,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:111"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def corollaryOneEta (horizon : Nat) : Real","missing":[],"search":"corollaryoneeta banditrlproof.stochasticgradientbandit.corollaryoneeta the horizon-indexed fixed learning rate used in corollary 1. definition compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceTheoremOne_margin_of_two_mul_eta_sourceC_le","label":"sourceTheoremOne_margin_of_two_mul_eta_sourceC_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceTheoremOne_margin_of_two_mul_eta_sourceC_le","description":"The source small-learning-rate branch implies the strict Theorem-1 margin.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-dbffadee6256","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1916,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:116"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceTheoremOne_margin_of_two_mul_eta_sourceC_le (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (hsmall : 2 * eta * sourceC eta <= Delta) : eta * sourceC eta < Delta","missing":[],"search":"sourcetheoremone_margin_of_two_mul_eta_sourcec_le banditrlproof.stochasticgradientbandit.sourcetheoremone_margin_of_two_mul_eta_sourcec_le the source small-learning-rate branch implies the strict theorem-1 margin. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceTheoremOne_constant_le_inv_eta","label":"sourceTheoremOne_constant_le_inv_eta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceTheoremOne_constant_le_inv_eta","description":"Under the Corollary-1 small-learning-rate branch, the constant term in Theorem 1 is at most `1 / eta`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-a4e9f0e3dc63","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1917,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:125"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceTheoremOne_constant_le_inv_eta (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (hsmall : 2 * eta * sourceC eta <= Delta) : Delta / (2 * eta * (Delta - eta * sourceC eta)) <= 1 / eta","missing":[],"search":"sourcetheoremone_constant_le_inv_eta banditrlproof.stochasticgradientbandit.sourcetheoremone_constant_le_inv_eta under the corollary-1 small-learning-rate branch, the constant term in theorem 1 is at most `1 / eta`. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_pos","label":"corollaryOneEta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneEta_pos","description":"theorem corollaryOneEta_pos (horizon : Nat) (hhorizon : 2 <= horizon) : 0 < corollaryOneEta horizon","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-bed88b477766","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1918,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:139"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOneEta_pos (horizon : Nat) (hhorizon : 2 <= horizon) : 0 < corollaryOneEta horizon","missing":[],"search":"corollaryoneeta_pos banditrlproof.stochasticgradientbandit.corollaryoneeta_pos theorem corollaryoneeta_pos (horizon : nat) (hhorizon : 2 <= horizon) : 0 < corollaryoneeta horizon theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_sq","label":"corollaryOneEta_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneEta_sq","description":"theorem corollaryOneEta_sq (horizon : Nat) (hhorizon : 2 <= horizon) : corollaryOneEta horizon ^ 2 = Real.log (horizon : Real) / (horizon : Real)","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-906b3c7a9cfd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1919,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:146"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOneEta_sq (horizon : Nat) (hhorizon : 2 <= horizon) : corollaryOneEta horizon ^ 2 = Real.log (horizon : Real) / (horizon : Real)","missing":[],"search":"corollaryoneeta_sq banditrlproof.stochasticgradientbandit.corollaryoneeta_sq theorem corollaryoneeta_sq (horizon : nat) (hhorizon : 2 <= horizon) : corollaryoneeta horizon ^ 2 = real.log (horizon : real) / (horizon : real) theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_le_one","label":"corollaryOneEta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneEta_le_one","description":"The Corollary-1 learning rate stays in the range where `C_eta <= exp 2` is available.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-8eaf32bc5de3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1920,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:158"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOneEta_le_one (horizon : Nat) (hhorizon : 2 <= horizon) : corollaryOneEta horizon <= 1","missing":[],"search":"corollaryoneeta_le_one banditrlproof.stochasticgradientbandit.corollaryoneeta_le_one the corollary-1 learning rate stays in the range where `c_eta <= exp 2` is available. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneRate","label":"corollaryOneRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneRate","description":"The square-root rate appearing in the explicit Corollary-1 endpoint.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-c57af2cf50dc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1921,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:171"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def corollaryOneRate (horizon : Nat) : Real","missing":[],"search":"corollaryonerate banditrlproof.stochasticgradientbandit.corollaryonerate the square-root rate appearing in the explicit corollary-1 endpoint. definition compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneRate_nonneg","label":"corollaryOneRate_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneRate_nonneg","description":"theorem corollaryOneRate_nonneg (horizon : Nat) : 0 <= corollaryOneRate horizon","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-bd2f47f5d815","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1922,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:174"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOneRate_nonneg (horizon : Nat) : 0 <= corollaryOneRate horizon","missing":[],"search":"corollaryonerate_nonneg banditrlproof.stochasticgradientbandit.corollaryonerate_nonneg theorem corollaryonerate_nonneg (horizon : nat) : 0 <= corollaryonerate horizon theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_horizon_eq_rate","label":"corollaryOneEta_mul_horizon_eq_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_horizon_eq_rate","description":"The horizon-indexed learning rate times the horizon is exactly the square-root rate.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-8f80a73ffabf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1923,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:179"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOneEta_mul_horizon_eq_rate (horizon : Nat) (hhorizon : 2 <= horizon) : corollaryOneEta horizon * (horizon : Real) = corollaryOneRate horizon","missing":[],"search":"corollaryoneeta_mul_horizon_eq_rate banditrlproof.stochasticgradientbandit.corollaryoneeta_mul_horizon_eq_rate the horizon-indexed learning rate times the horizon is exactly the square-root rate. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_rate_eq_log","label":"corollaryOneEta_mul_rate_eq_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_rate_eq_log","description":"theorem corollaryOneEta_mul_rate_eq_log (horizon : Nat) (hhorizon : 2 <= horizon) : corollaryOneEta horizon * corollaryOneRate horizon = Real.log (horizon : Real)","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-bb8c3ee7a8a7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1924,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:202"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOneEta_mul_rate_eq_log (horizon : Nat) (hhorizon : 2 <= horizon) : corollaryOneEta horizon * corollaryOneRate horizon = Real.log (horizon : Real)","missing":[],"search":"corollaryoneeta_mul_rate_eq_log banditrlproof.stochasticgradientbandit.corollaryoneeta_mul_rate_eq_log theorem corollaryoneeta_mul_rate_eq_log (horizon : nat) (hhorizon : 2 <= horizon) : corollaryoneeta horizon * corollaryonerate horizon = real.log (horizon : real) theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_inv_eta_le_inv_log_two_mul_rate","label":"corollaryOne_inv_eta_le_inv_log_two_mul_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOne_inv_eta_le_inv_log_two_mul_rate","description":"The inverse learning-rate term is an absolute-constant multiple of the square-root rate. We keep the exact constant `1 / log 2`, avoiding a hidden asymptotic threshold.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-002115eaf9ba","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1925,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:219"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOne_inv_eta_le_inv_log_two_mul_rate (horizon : Nat) (hhorizon : 2 <= horizon) : 1 / corollaryOneEta horizon <= (1 / Real.log 2) * corollaryOneRate horizon","missing":[],"search":"corollaryone_inv_eta_le_inv_log_two_mul_rate banditrlproof.stochasticgradientbandit.corollaryone_inv_eta_le_inv_log_two_mul_rate the inverse learning-rate term is an absolute-constant multiple of the square-root rate. we keep the exact constant `1 / log 2`, avoiding a hidden asymptotic threshold. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_log_argument_le_horizon_pow_four","label":"corollaryOne_log_argument_le_horizon_pow_four","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOne_log_argument_le_horizon_pow_four","description":"For `T >= 2`, `0 < Delta < 1`, and the Corollary-1 learning rate, the argument of the Theorem-1 logarithm is at most `T^4`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-cecdb84b198c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1926,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:240"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOne_log_argument_le_horizon_pow_four (horizon : Nat) (hhorizon : 2 <= horizon) (Delta : Real) (hDelta : 0 < Delta) (hDelta_lt_one : Delta < 1) : 1 + 4 * corollaryOneEta horizon * Delta * (horizon : Real) <= (horizon : Real) ^ 4","missing":[],"search":"corollaryone_log_argument_le_horizon_pow_four banditrlproof.stochasticgradientbandit.corollaryone_log_argument_le_horizon_pow_four for `t >= 2`, `0 < delta < 1`, and the corollary-1 learning rate, the argument of the theorem-1 logarithm is at most `t^4`. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_log_term_le_two_mul_rate","label":"corollaryOne_log_term_le_two_mul_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOne_log_term_le_two_mul_rate","description":"The logarithmic term in the Theorem-1 branch is at most twice the square-root rate.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-55505fe39825","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1927,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:282"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOne_log_term_le_two_mul_rate (horizon : Nat) (hhorizon : 2 <= horizon) (Delta : Real) (hDelta : 0 < Delta) (hDelta_lt_one : Delta < 1) : Real.log (1 + 4 * corollaryOneEta horizon * Delta * (horizon : Real)) / (2 * corollaryOneEta horizon) <= 2 * corollaryOneRate horizon","missing":[],"search":"corollaryone_log_term_le_two_mul_rate banditrlproof.stochasticgradientbandit.corollaryone_log_term_le_two_mul_rate the logarithmic term in the theorem-1 branch is at most twice the square-root rate. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_gap_mul_horizon_le_exp_constant_mul_rate","label":"corollaryOne_gap_mul_horizon_le_exp_constant_mul_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOne_gap_mul_horizon_le_exp_constant_mul_rate","description":"On the complementary branch, the pathwise `Delta * T` bound is still a square-root rate because failure of `2 * eta * C_eta <= Delta` forces the gap below the learning-rate scale.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-69bf0532a100","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1928,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:314"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOne_gap_mul_horizon_le_exp_constant_mul_rate (horizon : Nat) (hhorizon : 2 <= horizon) (Delta : Real) (hlarge : ¬ 2 * corollaryOneEta horizon * sourceC (corollaryOneEta horizon) <= Delta) : Delta * (horizon : Real) <= (2 * Real.exp 2) * corollaryOneRate horizon","missing":[],"search":"corollaryone_gap_mul_horizon_le_exp_constant_mul_rate banditrlproof.stochasticgradientbandit.corollaryone_gap_mul_horizon_le_exp_constant_mul_rate on the complementary branch, the pathwise `delta * t` bound is still a square-root rate because failure of `2 * eta * c_eta <= delta` forces the gap below the learning-rate scale. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneAbsoluteConstant","label":"corollaryOneAbsoluteConstant","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOneAbsoluteConstant","description":"An explicit horizon-independent constant for the finite Corollary-1 endpoint.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-902907405f62","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1929,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:350"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def corollaryOneAbsoluteConstant : Real","missing":[],"search":"corollaryoneabsoluteconstant banditrlproof.stochasticgradientbandit.corollaryoneabsoluteconstant an explicit horizon-independent constant for the finite corollary-1 endpoint. definition compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_piecewise_bound","label":"corollaryOne_piecewise_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.corollaryOne_piecewise_bound","description":"theorem corollaryOne_piecewise_bound (horizon : Nat) (hhorizon : 2 <= horizon) (Delta : Real) (hDelta : 0 < Delta) (hDelta_lt_one : Delta < 1) : (if 2 * corollaryOneEta horizon * sourceC (corollaryOneEta horizon) <= Delta then Real.log (1 + 4 * corollaryOneEta horizon * Delta * (horizon : Real)) / (2 * corollaryOneEta horizon) + 1 / corollaryOneEta horizon else Delta * (horizon : Real)) <= corollaryOneAbsoluteConsta…","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-2e80d83c34c2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1930,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:353"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem corollaryOne_piecewise_bound (horizon : Nat) (hhorizon : 2 <= horizon) (Delta : Real) (hDelta : 0 < Delta) (hDelta_lt_one : Delta < 1) : (if 2 * corollaryOneEta horizon * sourceC (corollaryOneEta horizon) <= Delta then Real.log (1 + 4 * corollaryOneEta horizon * Delta * (horizon : Real)) / (2 * corollaryOneEta horizon) + 1 / corollaryOneEta horizon else Delta * (horizon : Real)) <= corollaryOneAbsoluteConstant * corollaryOneRate horizon","missing":[],"search":"corollaryone_piecewise_bound banditrlproof.stochasticgradientbandit.corollaryone_piecewise_bound theorem corollaryone_piecewise_bound (horizon : nat) (hhorizon : 2 <= horizon) (delta : real) (hdelta : 0 < delta) (hdelta_lt_one : delta < 1) : (if 2 * corollaryoneeta horizon * sourcec (corollaryoneeta horizon) <= delta then real.log (1 + 4 * corollaryoneeta horizon * delta * (horizon : real)) / (2 * corollaryoneeta horizon) + 1 / corollaryoneeta horizon else delta * (horizon : real)) <= corollaryoneabsoluteconstant * corollaryonerate horizon theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne_piecewise","label":"twoArmFixedIIDDirac_corollaryOne_piecewise","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne_piecewise","description":"Exact finite two-branch version of source Corollary 1. A separate fixed rate is used for each source horizon `T = tailHorizon + 1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-0e43e2467b03","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1931,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:414"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDDirac_corollaryOne_piecewise (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (mean : Fin 2 -> Real) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, |reward| <= 1) (hmean : forall arm, integral (armLaw arm) id = mean arm) (Delta : Real) (hDelta : 0 < Delta) (hDelta_lt_one : Delta < 1) (hgap : mean 0 - mean 1 = Delta) (tailHorizon : Nat) (horizon_ge_two : 1 <= tailHorizon) : let eta := corollaryOneEta (tailHorizon + 1) integral (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmSampledPseudoRegret (Env := Unit) Delta (tailHorizon + 1)) <= if 2 * eta * sourceC eta <= Delta then Real.log (1 + 4 * eta * Delta * ((tailHorizon + 1 : Nat) : Real)) / (2 * eta) + 1 / eta else Delta * ((tailHorizon + 1 : Nat) : Real)","missing":[],"search":"twoarmfixediiddirac_corollaryone_piecewise banditrlproof.stochasticgradientbandit.twoarmfixediiddirac_corollaryone_piecewise exact finite two-branch version of source corollary 1. a separate fixed rate is used for each source horizon `t = tailhorizon + 1`. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","label":"twoArmFixedIIDDirac_corollaryOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","description":"Source Corollary 1 on the generated two-arm fixed-IID trajectory, with an explicit absolute constant and no asymptotic notation. The learning rate is fixed within each horizon and may vary across the horizon-indexed family.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditcorollaryone/index.html#decl-f135a6147c83","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","order":1932,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditCorollaryOne.lean:468"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDDirac_corollaryOne (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (mean : Fin 2 -> Real) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, |reward| <= 1) (hmean : forall arm, integral (armLaw arm) id = mean arm) (Delta : Real) (hDelta : 0 < Delta) (hDelta_lt_one : Delta < 1) (hgap : mean 0 - mean 1 = Delta) (tailHorizon : Nat) (horizon_ge_two : 1 <= tailHorizon) : let eta := corollaryOneEta (tailHorizon + 1) integral (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmSampledPseudoRegret (Env := Unit) Delta (tailHorizon + 1)) <= corollaryOneAbsoluteConstant * corollaryOneRate (tailHorizon + 1)","missing":[],"search":"twoarmfixediiddirac_corollaryone banditrlproof.stochasticgradientbandit.twoarmfixediiddirac_corollaryone source corollary 1 on the generated two-arm fixed-iid trajectory, with an explicit absolute constant and no asymptotic notation. the learning rate is fixed within each horizon and may vary across the horizon-indexed family. theorem compiled","shard":"modules/a7d847c979a636f6.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.two_mul_abs_pow_div_factorial_add_two_le","label":"two_mul_abs_pow_div_factorial_add_two_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.two_mul_abs_pow_div_factorial_add_two_le","description":"A factorial comparison used to dominate the shifted exponential series.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-894691d13651","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1933,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:26"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem two_mul_abs_pow_div_factorial_add_two_le (x : Real) (n : Nat) : 2 * |x| ^ n / ((n + 2).factorial : Real) <= |x| ^ n / (n.factorial : Real)","missing":[],"search":"two_mul_abs_pow_div_factorial_add_two_le banditrlproof.stochasticgradientbandit.two_mul_abs_pow_div_factorial_add_two_le a factorial comparison used to dominate the shifted exponential series. theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceC","label":"sourceC","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceC","description":"The source constant `C_eta = 2 * sum_{n >= 0} (2 * eta)^n / (n + 2)!`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-e05fb4154392","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1934,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:45"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceC (eta : Real) : Real","missing":[],"search":"sourcec banditrlproof.stochasticgradientbandit.sourcec the source constant `c_eta = 2 * sum_{n >= 0} (2 * eta)^n / (n + 2)!`. definition compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_terms_summable","label":"sourceC_terms_summable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceC_terms_summable","description":"theorem sourceC_terms_summable (eta : Real) : Summable (fun n : Nat => (2 * eta) ^ n / ((n + 2).factorial : Real))","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-0401bde33984","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1935,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:48"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceC_terms_summable (eta : Real) : Summable (fun n : Nat => (2 * eta) ^ n / ((n + 2).factorial : Real))","missing":[],"search":"sourcec_terms_summable banditrlproof.stochasticgradientbandit.sourcec_terms_summable theorem sourcec_terms_summable (eta : real) : summable (fun n : nat => (2 * eta) ^ n / ((n + 2).factorial : real)) theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_nonneg","label":"sourceC_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceC_nonneg","description":"theorem sourceC_nonneg (eta : Real) (heta : 0 <= eta) : 0 <= sourceC eta","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-ce7be71b66d1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1936,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:65"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceC_nonneg (eta : Real) (heta : 0 <= eta) : 0 <= sourceC eta","missing":[],"search":"sourcec_nonneg banditrlproof.stochasticgradientbandit.sourcec_nonneg theorem sourcec_nonneg (eta : real) (heta : 0 <= eta) : 0 <= sourcec eta theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_mono","label":"sourceC_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceC_mono","description":"Monotonicity needed to replace the time-varying source constants in the Theorem-1 recurrences by a common `C_eta`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-53d99167d071","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1937,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:72"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceC_mono {eta eta' : Real} (heta : 0 <= eta) (hle : eta <= eta') : sourceC eta <= sourceC eta'","missing":[],"search":"sourcec_mono banditrlproof.stochasticgradientbandit.sourcec_mono monotonicity needed to replace the time-varying source constants in the theorem-1 recurrences by a common `c_eta`. theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_le_exp_two_mul","label":"sourceC_le_exp_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sourceC_le_exp_two_mul","description":"The source comparison following Theorem 1: `C_eta <= exp (2 * eta)`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-d9c5c442fa23","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1938,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:84"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceC_le_exp_two_mul (eta : Real) (heta : 0 <= eta) : sourceC eta <= Real.exp (2 * eta)","missing":[],"search":"sourcec_le_exp_two_mul banditrlproof.stochasticgradientbandit.sourcec_le_exp_two_mul the source comparison following theorem 1: `c_eta <= exp (2 * eta)`. theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo","label":"expTailTwo","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expTailTwo","description":"The exponential-series tail beginning at degree two.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-2fb7a910f0e7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1939,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:102"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def expTailTwo (x : Real) : Real","missing":[],"search":"exptailtwo banditrlproof.stochasticgradientbandit.exptailtwo the exponential-series tail beginning at degree two. definition compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo_terms_summable","label":"expTailTwo_terms_summable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expTailTwo_terms_summable","description":"theorem expTailTwo_terms_summable (x : Real) : Summable (fun n : Nat => x ^ (n + 2) / ((n + 2).factorial : Real))","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-2e0d6ad105c3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1940,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:105"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expTailTwo_terms_summable (x : Real) : Summable (fun n : Nat => x ^ (n + 2) / ((n + 2).factorial : Real))","missing":[],"search":"exptailtwo_terms_summable banditrlproof.stochasticgradientbandit.exptailtwo_terms_summable theorem exptailtwo_terms_summable (x : real) : summable (fun n : nat => x ^ (n + 2) / ((n + 2).factorial : real)) theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.exp_eq_one_add_self_add_expTailTwo","label":"exp_eq_one_add_self_add_expTailTwo","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.exp_eq_one_add_self_add_expTailTwo","description":"theorem exp_eq_one_add_self_add_expTailTwo (x : Real) : Real.exp x = 1 + x + expTailTwo x","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-e9404b7f013e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1941,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:111"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem exp_eq_one_add_self_add_expTailTwo (x : Real) : Real.exp x = 1 + x + expTailTwo x","missing":[],"search":"exp_eq_one_add_self_add_exptailtwo banditrlproof.stochasticgradientbandit.exp_eq_one_add_self_add_exptailtwo theorem exp_eq_one_add_self_add_exptailtwo (x : real) : real.exp x = 1 + x + exptailtwo x theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo_le_of_abs_le","label":"expTailTwo_le_of_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.expTailTwo_le_of_abs_le","description":"theorem expTailTwo_le_of_abs_le {x y : Real} (hxy : |x| <= y) : expTailTwo x <= expTailTwo y","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-eb128bde22bf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1942,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:119"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expTailTwo_le_of_abs_le {x y : Real} (hxy : |x| <= y) : expTailTwo x <= expTailTwo y","missing":[],"search":"exptailtwo_le_of_abs_le banditrlproof.stochasticgradientbandit.exptailtwo_le_of_abs_le theorem exptailtwo_le_of_abs_le {x y : real} (hxy : |x| <= y) : exptailtwo x <= exptailtwo y theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.sq_div_two_mul_sourceC_abs_div_two","label":"sq_div_two_mul_sourceC_abs_div_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.sq_div_two_mul_sourceC_abs_div_two","description":"theorem sq_div_two_mul_sourceC_abs_div_two (q : Real) : q ^ 2 / 2 * sourceC (|q| / 2) = expTailTwo |q|","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-41b298ff5fe7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1943,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:132"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sq_div_two_mul_sourceC_abs_div_two (q : Real) : q ^ 2 / 2 * sourceC (|q| / 2) = expTailTwo |q|","missing":[],"search":"sq_div_two_mul_sourcec_abs_div_two banditrlproof.stochasticgradientbandit.sq_div_two_mul_sourcec_abs_div_two theorem sq_div_two_mul_sourcec_abs_div_two (q : real) : q ^ 2 / 2 * sourcec (|q| / 2) = exptailtwo |q| theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.exp_mul_le_sourceEqEight","label":"exp_mul_le_sourceEqEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.exp_mul_le_sourceEqEight","description":"Pointwise form of source Equation (8).","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-1423ab0e9a0d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1944,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:150"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem exp_mul_le_sourceEqEight (q reward : Real) (hreward : |reward| <= 1) : Real.exp (q * reward) <= 1 + q * reward + q ^ 2 / 2 * sourceC (|q| / 2)","missing":[],"search":"exp_mul_le_sourceeqeight banditrlproof.stochasticgradientbandit.exp_mul_le_sourceeqeight pointwise form of source equation (8). theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight","label":"integral_exp_mul_le_sourceEqEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight","description":"Expectation form of source Equation (8), with integrability explicit.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-073f28ba0b36","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1945,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:162"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_mul_le_sourceEqEight {Omega : Type*} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (q : Real) (reward : Omega -> Real) (hrewardIntegrable : MeasureTheory.Integrable reward mu) (hexpIntegrable : MeasureTheory.Integrable (fun omega => Real.exp (q * reward omega)) mu) (hreward : ∀ᵐ omega ∂mu, |reward omega| <= 1) : (∫ omega, Real.exp (q * reward omega) ∂mu) <= 1 + q * (∫ omega, reward omega ∂mu) + q ^ 2 / 2 * sourceC (|q| / 2)","missing":[],"search":"integral_exp_mul_le_sourceeqeight banditrlproof.stochasticgradientbandit.integral_exp_mul_le_sourceeqeight expectation form of source equation (8), with integrability explicit. theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","label":"integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","description":"Equation (8) from measurability and almost-sure support in `[-1, 1]`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbanditexponentialaudit/index.html#decl-ebab60a052ae","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","order":1946,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditExponentialAudit.lean:193"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one {Omega : Type*} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (q : Real) (reward : Omega -> Real) (hrewardMeasurable : MeasureTheory.AEStronglyMeasurable reward mu) (hreward : ∀ᵐ omega ∂mu, |reward omega| <= 1) : (∫ omega, Real.exp (q * reward omega) ∂mu) <= 1 + q * (∫ omega, reward omega ∂mu) + q ^ 2 / 2 * sourceC (|q| / 2)","missing":[],"search":"integral_exp_mul_le_sourceeqeight_of_ae_abs_le_one banditrlproof.stochasticgradientbandit.integral_exp_mul_le_sourceeqeight_of_ae_abs_le_one equation (8) from measurability and almost-sure support in `[-1, 1]`. theorem compiled","shard":"modules/bb4f78b4a499d0ac.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin","label":"theoremFourStepOneMargin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin","description":"The positive scalar margin used in Appendix E after Equation (22).","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-ea3978636c4e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1947,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:37"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def theoremFourStepOneMargin (K : Nat) (eta Delta : Real) : Real","missing":[],"search":"theoremfoursteponemargin banditrlproof.stochasticgradientbandit.theoremfoursteponemargin the positive scalar margin used in appendix e after equation (22). definition compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound","label":"theoremFourStepFourSurvivalLowerBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound","description":"The unconditional survival mass required by the audited Step-4 event composition.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-93921ffd45b0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1948,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:42"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def theoremFourStepFourSurvivalLowerBound (pPrime c : Real) : Real","missing":[],"search":"theoremfourstepfoursurvivallowerbound banditrlproof.stochasticgradientbandit.theoremfourstepfoursurvivallowerbound the unconditional survival mass required by the audited step-4 event composition. definition compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin_pos","label":"theoremFourStepOneMargin_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin_pos","description":"The source learning-rate condition implies that the Equation-(22) drift margin is strictly positive.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-86405e025293","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1949,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:47"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem theoremFourStepOneMargin_pos (K : Nat) (eta Delta : Real) (hmargin : eta * sourceC eta < 2 * Delta / ((K : Real) + 2)) : 0 < theoremFourStepOneMargin K eta Delta","missing":[],"search":"theoremfoursteponemargin_pos banditrlproof.stochasticgradientbandit.theoremfoursteponemargin_pos the source learning-rate condition implies that the equation-(22) drift margin is strictly positive. theorem compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound_pos","label":"theoremFourStepFourSurvivalLowerBound_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound_pos","description":"If `pPrime > 0` and `c < 1/2`, the audited Step-4 survival lower bound is positive.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-3f3e781f24b4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1950,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:58"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem theoremFourStepFourSurvivalLowerBound_pos (pPrime c : Real) (hpPrime : 0 < pPrime) (hc_half : c < 1 / 2) : 0 < theoremFourStepFourSurvivalLowerBound pPrime c","missing":[],"search":"theoremfourstepfoursurvivallowerbound_pos banditrlproof.stochasticgradientbandit.theoremfourstepfoursurvivallowerbound_pos if `pprime > 0` and `c < 1/2`, the audited step-4 survival lower bound is positive. theorem compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_ge","label":"theoremFourStepFour_survivalMass_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_ge","description":"Finite total-probability contract for Appendix E, Step 4. `bufferedMass` represents the probability of entering the strict buffer `q_{s+1} < c`; `jointSurvivalMass` represents the probability of both entering that buffer and not returning before the audited finite horizon. The two middle hypotheses are the multiplication-free form of `P(buffer) >= pPrime` and `P(survival | buffer) >= 1 - 2*c`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-4a23c61cbc9b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1951,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:75"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem theoremFourStepFour_survivalMass_ge (pPrime c bufferedMass jointSurvivalMass survivalMass : Real) (hc_half : c < 1 / 2) (hbuffer : pPrime <= bufferedMass) (hconditional : (1 - 2 * c) * bufferedMass <= jointSurvivalMass) (hsubset : jointSurvivalMass <= survivalMass) : theoremFourStepFourSurvivalLowerBound pPrime c <= survivalMass","missing":[],"search":"theoremfourstepfour_survivalmass_ge banditrlproof.stochasticgradientbandit.theoremfourstepfour_survivalmass_ge finite total-probability contract for appendix e, step 4. `bufferedmass` represents the probability of entering the strict buffer `q_{s+1} < c`; `jointsurvivalmass` represents the probability of both entering that buffer and not returning before the audited finite horizon. the two middle hypotheses are the multiplication-free form of `p(buffer) >= pprime` and `p(survival | buffer) >= 1 - 2*c`. theorem compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_pos","label":"theoremFourStepFour_survivalMass_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_pos","description":"The audited finite event contract yields a strictly positive survival mass when its buffered event has positive mass.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-8d760702a67b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1952,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:93"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem theoremFourStepFour_survivalMass_pos (pPrime c bufferedMass jointSurvivalMass survivalMass : Real) (hpPrime : 0 < pPrime) (hc_half : c < 1 / 2) (hbuffer : pPrime <= bufferedMass) (hconditional : (1 - 2 * c) * bufferedMass <= jointSurvivalMass) (hsubset : jointSurvivalMass <= survivalMass) : 0 < survivalMass","missing":[],"search":"theoremfourstepfour_survivalmass_pos banditrlproof.stochasticgradientbandit.theoremfourstepfour_survivalmass_pos the audited finite event contract yields a strictly positive survival mass when its buffered event has positive mass. theorem compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourFiniteGeometricPhaseMass_le_inv","label":"theoremFourFiniteGeometricPhaseMass_le_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourFiniteGeometricPhaseMass_le_inv","description":"For `0 < rho <= 1`, the finite geometric phase envelope is at most `1 / rho`. This is the finite statement needed before any infinite expected-phase claim.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-df58a09b6209","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1953,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:109"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem theoremFourFiniteGeometricPhaseMass_le_inv (rho : Real) (hrho_pos : 0 < rho) (hrho_le_one : rho <= 1) (phaseCount : Nat) : (Finset.range phaseCount).sum (fun phase => (1 - rho) ^ phase) <= 1 / rho","missing":[],"search":"theoremfourfinitegeometricphasemass_le_inv banditrlproof.stochasticgradientbandit.theoremfourfinitegeometricphasemass_le_inv for `0 < rho <= 1`, the finite geometric phase envelope is at most `1 / rho`. this is the finite statement needed before any infinite expected-phase claim. theorem compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourFiniteTransientMass_le_inv","label":"theoremFourFiniteTransientMass_le_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.theoremFourFiniteTransientMass_le_inv","description":"Any finite transient-phase mass dominated termwise by the geometric return envelope inherits the same `1 / rho` bound.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremfourcontractaudit/index.html#decl-ac36ab0f2be5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","order":1954,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremFourContractAudit.lean:124"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem theoremFourFiniteTransientMass_le_inv (rho : Real) (hrho_pos : 0 < rho) (hrho_le_one : rho <= 1) (phaseMass : Nat -> Real) (hphase : forall phase, phaseMass phase <= (1 - rho) ^ phase) (phaseCount : Nat) : (Finset.range phaseCount).sum phaseMass <= 1 / rho","missing":[],"search":"theoremfourfinitetransientmass_le_inv banditrlproof.stochasticgradientbandit.theoremfourfinitetransientmass_le_inv any finite transient-phase mass dominated termwise by the geometric return envelope inherits the same `1 / rho` bound. theorem compiled","shard":"modules/691fed6ec367d723.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_fixedArmFinitePrefix_eq_pi","label":"armStreamMeasure_map_fixedArmFinitePrefix_eq_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_fixedArmFinitePrefix_eq_pi","description":"A fixed arm's first `m` latent rewards have the finite IID product law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-650c104c6bfe","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1955,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:32"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_fixedArmFinitePrefix_eq_pi {K m : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) : Measure.map (fun stream : ArmRewardStream K => fun i : Fin m => stream (i : Nat) arm) (armStreamMeasure nu) = Measure.pi (fun _ : Fin m => nu arm)","missing":[],"search":"armstreammeasure_map_fixedarmfiniteprefix_eq_pi banditrlproof.ucb.armstreammeasure_map_fixedarmfiniteprefix_eq_pi a fixed arm's first `m` latent rewards have the finite iid product law. theorem compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi","label":"latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi","description":"The fixed-arm product law lifted through the exact stream marginal of the latent trajectory coupling.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-e034707a2716","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1956,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:81"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi {Env : Type u} {K m : Nat} [MeasurableSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) : Measure.map (fun sample : UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real) => fun i : Fin m => sample.1 (i : Nat) arm) (latentArmStreamTrajectoryMeasure algorithm env nu) = Measure.pi (fun _ : Fin m => nu arm)","missing":[],"search":"latentarmstreamtrajectorymeasure_map_fixedarmfiniteprefix_eq_pi banditrlproof.thompson.latentarmstreamtrajectorymeasure_map_fixedarmfiniteprefix_eq_pi the fixed-arm product law lifted through the exact stream marginal of the latent trajectory coupling. theorem compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure","label":"twoArmFixedIIDLatentTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure","description":"The latent-stream coupling specialized to the zero-initialized two-arm SGB policy and the fixed arm laws used by the source instance.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-fa40dbb32608","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1957,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:127"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def twoArmFixedIIDLatentTrajectoryMeasure (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) : Measure (UCB.ArmRewardStream 2 × ((n : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmfixediidlatenttrajectorymeasure banditrlproof.stochasticgradientbandit.twoarmfixediidlatenttrajectorymeasure the latent-stream coupling specialized to the zero-initialized two-arm sgb policy and the fixed arm laws used by the source instance. definition compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_dirac_eq_map_trajectoryKernel","label":"twoArmTrajectoryMeasure_dirac_eq_map_trajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_dirac_eq_map_trajectoryKernel","description":"The native `Unit`-environment law is its trajectory kernel with the trivial environment coordinate reattached. This is normalization for the still-open latent-to-native trajectory adapter.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-b5632ea23bb5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1958,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:152"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectoryMeasure_dirac_eq_map_trajectoryKernel (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) : twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob) = Measure.map (Prod.mk ()) (trajectoryKernel (fun _ : Fin 2 => 0) eta (twoArmFixedIIDEnvironment armLaw hprob) ())","missing":[],"search":"twoarmtrajectorymeasure_dirac_eq_map_trajectorykernel banditrlproof.stochasticgradientbandit.twoarmtrajectorymeasure_dirac_eq_map_trajectorykernel the native `unit`-environment law is its trajectory kernel with the trivial environment coordinate reattached. this is normalization for the still-open latent-to-native trajectory adapter. theorem compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.stationaryRewardKernelAt_twoArmFixedIIDRewardKernel_eq","label":"stationaryRewardKernelAt_twoArmFixedIIDRewardKernel_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.stationaryRewardKernelAt_twoArmFixedIIDRewardKernel_eq","description":"Freezing the direct fixed-IID reward kernel at `Unit` gives the same finite-arm kernel used by the latent-stream construction.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-1c63bcbb2256","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1959,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:168"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem stationaryRewardKernelAt_twoArmFixedIIDRewardKernel_eq (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) : letI : IsMarkovKernel (twoArmFixedIIDRewardKernel armLaw)","missing":[],"search":"stationaryrewardkernelat_twoarmfixediidrewardkernel_eq banditrlproof.stochasticgradientbandit.stationaryrewardkernelat_twoarmfixediidrewardkernel_eq freezing the direct fixed-iid reward kernel at `unit` gives the same finite-arm kernel used by the latent-stream construction. theorem compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","label":"twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","description":"On the specialized coupling, the first `m` latent optimal-arm rewards have exactly the product of the source arm law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-6710267e04b8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1960,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:183"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (m : Nat) : Measure.map (fun sample : UCB.ArmRewardStream 2 × ((n : Nat) -> Fin 2 × Real) => fun i : Fin m => sample.1 (i : Nat) 0) (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) = Measure.pi (fun _ : Fin m => armLaw 0)","missing":[],"search":"twoarmfixediidlatenttrajectorymeasure_map_optimalprefix_eq_pi banditrlproof.stochasticgradientbandit.twoarmfixediidlatenttrajectorymeasure_map_optimalprefix_eq_pi on the specialized coupling, the first `m` latent optimal-arm rewards have exactly the product of the source arm law. theorem compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_latentCoordinate_ae","label":"twoArmNthOptimalPullReward_eq_latentCoordinate_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_latentCoordinate_ae","description":"At every finite nth optimal-arm pull, the observed stopped reward is the corresponding latent arm-`0` coordinate almost surely. This is pathwise support, not an IID statement about totalized or occurrence-conditioned stopped rewards.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwolatentreward/index.html#decl-2273d9582ab2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","order":1961,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoLatentReward.lean:205"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullReward_eq_latentCoordinate_ae (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (pullIndex : Nat) : ∀ᵐ sample ∂twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta, ∀ t : Nat, twoArmNthOptimalPullTime (Env := Unit) pullIndex ((), sample.2) = (t : WithTop Nat) -> twoArmNthOptimalPullReward (Env := Unit) pullIndex ((), sample.2) = sample.1 pullIndex 0","missing":[],"search":"twoarmnthoptimalpullreward_eq_latentcoordinate_ae banditrlproof.stochasticgradientbandit.twoarmnthoptimalpullreward_eq_latentcoordinate_ae at every finite nth optimal-arm pull, the observed stopped reward is the corresponding latent arm-`0` coordinate almost surely. this is pathwise support, not an iid statement about totalized or occurrence-conditioned stopped rewards. theorem compiled","shard":"modules/5dca51c8946f1a6f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryRewardHistoryEnvironment","label":"stationaryRewardHistoryEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryRewardHistoryEnvironment","description":"The native stationary environment: feedback is the law of the selected arm, independent of the observed history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-71a4df4e15f0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1962,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:31"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def stationaryRewardHistoryEnvironment {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : HistoryEnvironment (Fin K) Real where","missing":[],"search":"stationaryrewardhistoryenvironment banditrlproof.thompson.stationaryrewardhistoryenvironment the native stationary environment: feedback is the law of the selected arm, independent of the observed history. definition compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel_stationaryRewardHistoryEnvironment","label":"historyStepKernel_stationaryRewardHistoryEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel_stationaryRewardHistoryEnvironment","description":"The native step kernel composes the policy with the selected-arm law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-9f42a9b18392","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1963,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:38"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernel_stationaryRewardHistoryEnvironment {K : Nat} (algorithm : HistoryAlgorithm (Fin K) Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : historyStepKernel algorithm (stationaryRewardHistoryEnvironment nu) n = algorithm.policy n ⊗ₖ UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"historystepkernel_stationaryrewardhistoryenvironment banditrlproof.thompson.historystepkernel_stationaryrewardhistoryenvironment the native step kernel composes the policy with the selected-arm law. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextPair_eq_compProd","label":"latentArmStreamVisiblePrefixNextPair_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextPair_eq_compProd","description":"One-step native extension of the visible prefix under the latent coupling. At every deterministic time `n`, the joint law of the visible prefix through `n` and the next observed action/reward pair is the prefix law composed with the native stationary step kernel. This is a single-step statement; it does not identify whole prefixes or the full native law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-b670be380a73","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1964,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:51"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextPair_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => (Preorder.frestrictLe n sample.2, sample.2 (n + 1))) (latentArmStreamTrajectoryMeasure algorithm env nu) = Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => Preorder.frestrictLe n sample.2) (latentArmStreamTrajectoryMeasure algorithm env nu) ⊗ₘ historyStepKernel algorithm (stationaryRewardHistoryEnvironment nu) n","missing":[],"search":"latentarmstreamvisibleprefixnextpair_eq_compprod banditrlproof.thompson.latentarmstreamvisibleprefixnextpair_eq_compprod one-step native extension of the visible prefix under the latent coupling. at every deterministic time `n`, the joint law of the visible prefix through `n` and the next observed action/reward pair is the prefix law composed with the native stationary step kernel. this is a single-step statement; it does not identify whole prefixes or the full native law. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleInitialPair_eq_compProd","label":"latentArmStreamVisibleInitialPair_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleInitialPair_eq_compProd","description":"Time-zero pair law of the latent coupling: the initial action follows the algorithm's initial law and the initial reward is a fresh draw from the selected arm law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-914abeffb6b1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1965,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:93"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleInitialPair_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => sample.2 0) (latentArmStreamTrajectoryMeasure algorithm env nu) = algorithm.initialAction ⊗ₘ nu","missing":[],"search":"latentarmstreamvisibleinitialpair_eq_compprod banditrlproof.thompson.latentarmstreamvisibleinitialpair_eq_compprod time-zero pair law of the latent coupling: the initial action follows the algorithm's initial law and the initial reward is a fresh draw from the selected arm law. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.trajMeasure_map_eval_zero","label":"trajMeasure_map_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.trajMeasure_map_eval_zero","description":"The Ionescu-Tulcea trajectory law reproduces its initial law at time zero.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-49acb36ad67b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1966,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:202"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem trajMeasure_map_eval_zero {X : Nat -> Type u} [forall n, MeasurableSpace (X n)] (mu0 : Measure (X 0)) [IsProbabilityMeasure mu0] (kappa : (n : Nat) -> Kernel ((i : Finset.Iic n) -> X i) (X (n + 1))) [forall n, IsMarkovKernel (kappa n)] : (Kernel.trajMeasure mu0 kappa).map (fun x => x 0) = mu0","missing":[],"search":"trajmeasure_map_eval_zero banditrlproof.thompson.trajmeasure_map_eval_zero the ionescu-tulcea trajectory law reproduces its initial law at time zero. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.frestrictLe_succ_eq_extendPairHistorySucc","label":"frestrictLe_succ_eq_extendPairHistorySucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.frestrictLe_succ_eq_extendPairHistorySucc","description":"The visible prefix at `n + 1` is the prefix at `n` extended by the next pair.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-6f90568c9f60","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1967,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:221"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem frestrictLe_succ_eq_extendPairHistorySucc {K : Nat} (n : Nat) (x : (t : Nat) -> Fin K × Real) : Preorder.frestrictLe (n + 1) x = History.extendPairHistorySucc (Preorder.frestrictLe n x) (x (n + 1))","missing":[],"search":"frestrictle_succ_eq_extendpairhistorysucc banditrlproof.thompson.frestrictle_succ_eq_extendpairhistorysucc the visible prefix at `n + 1` is the prefix at `n` extended by the next pair. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.nativeStationaryTrajectoryMeasure","label":"nativeStationaryTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.nativeStationaryTrajectoryMeasure","description":"The native fixed-i.i.d. trajectory law: the same algorithm run against the stationary environment whose feedback is the selected-arm law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-71e725258d13","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1968,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:238"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def nativeStationaryTrajectoryMeasure {K : Nat} (algorithm : HistoryAlgorithm (Fin K) Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Measure ((n : Nat) -> Fin K × Real)","missing":[],"search":"nativestationarytrajectorymeasure banditrlproof.thompson.nativestationarytrajectorymeasure the native fixed-i.i.d. trajectory law: the same algorithm run against the stationary environment whose feedback is the selected-arm law. definition compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_map_frestrictLe_eq_native","label":"latentArmStreamVisibleTrajectoryMeasure_map_frestrictLe_eq_native","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_map_frestrictLe_eq_native","description":"*Native-prefix identification.** At every finite horizon `n`, the visible trajectory marginal of the latent arm-stream coupling and the native fixed-i.i.d. process induce the same law on prefixes through `n`. This is a finite-prefix identity. It does not by itself give the full native visible law, selected- or stopped-reward i.i.d. statements, the stopped-prefix future/no-return law, or Theorem 2.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-fba23ffc4301","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1969,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:259"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleTrajectoryMeasure_map_frestrictLe_eq_native {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : ((latentArmStreamTrajectoryMeasure algorithm env nu).map Prod.snd).map (Preorder.frestrictLe n) = (nativeStationaryTrajectoryMeasure algorithm nu).map (Preorder.frestrictLe n)","missing":[],"search":"latentarmstreamvisibletrajectorymeasure_map_frestrictle_eq_native banditrlproof.thompson.latentarmstreamvisibletrajectorymeasure_map_frestrictle_eq_native *native-prefix identification.** at every finite horizon `n`, the visible trajectory marginal of the latent arm-stream coupling and the native fixed-i.i.d. process induce the same law on prefixes through `n`. this is a finite-prefix identity. it does not by itself give the full native visible law, selected- or stopped-reward i.i.d. statements, the stopped-prefix future/no-return law, or theorem 2. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","label":"latentArmStreamVisibleTrajectoryMeasure_eq_native","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","description":"*Full native visible-law identification.** Forgetting the latent reward stream from the arm-stream coupling gives exactly the native stationary fixed-i.i.d. SGB trajectory law. The proof promotes the compiled equality of every inclusive finite prefix to an equality of complete trajectory measures by projective-limit uniqueness. It does not assert that totalized stopped rewards are i.i.d., identify a random-time futu…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativeprefix/index.html#decl-248ea0e376e8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","order":1970,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativePrefix.lean:356"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleTrajectoryMeasure_eq_native {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : (latentArmStreamTrajectoryMeasure algorithm env nu).map Prod.snd = nativeStationaryTrajectoryMeasure algorithm nu","missing":[],"search":"latentarmstreamvisibletrajectorymeasure_eq_native banditrlproof.thompson.latentarmstreamvisibletrajectorymeasure_eq_native *full native visible-law identification.** forgetting the latent reward stream from the arm-stream coupling gives exactly the native stationary fixed-i.i.d. sgb trajectory law. the proof promotes the compiled equality of every inclusive finite prefix to an equality of complete trajectory measures by projective-limit uniqueness. it does not assert that totalized stopped rewards are i.i.d., identify a random-time future cylinder, or prove the source theorem 2 endpoint. theorem compiled","shard":"modules/91b486b71f97e146.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_frestrictLe_eq_pi","label":"armStreamMeasure_map_frestrictLe_eq_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_frestrictLe_eq_pi","description":"Restricting the latent arm stream through time `n` gives the exact finite product of the per-round arm-vector laws.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-37e5f6abda84","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1971,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:41"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_frestrictLe_eq_pi {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (Preorder.frestrictLe n) (armStreamMeasure nu) = Measure.pi (fun _ : Finset.Iic n => Measure.infinitePi fun arm : Fin K => nu arm)","missing":[],"search":"armstreammeasure_map_frestrictle_eq_pi banditrlproof.ucb.armstreammeasure_map_frestrictle_eq_pi restricting the latent arm stream through time `n` gives the exact finite product of the per-round arm-vector laws. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.extendArmStreamFinitePrefix","label":"extendArmStreamFinitePrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.extendArmStreamFinitePrefix","description":"Extend a finite arm-stream box by zero rows after its endpoint.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-9b8e704a8de1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1972,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:52"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def extendArmStreamFinitePrefix {K : Nat} (n : Nat) (streamBox : (i : Finset.Iic n) -> Fin K -> Real) : ArmRewardStream K","missing":[],"search":"extendarmstreamfiniteprefix banditrlproof.ucb.extendarmstreamfiniteprefix extend a finite arm-stream box by zero rows after its endpoint. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_extendArmStreamFinitePrefix","label":"measurable_extendArmStreamFinitePrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_extendArmStreamFinitePrefix","description":"theorem measurable_extendArmStreamFinitePrefix {K : Nat} (n : Nat) : Measurable (extendArmStreamFinitePrefix (K := K) n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-118c9622c40b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1973,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:57"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_extendArmStreamFinitePrefix {K : Nat} (n : Nat) : Measurable (extendArmStreamFinitePrefix (K := K) n)","missing":[],"search":"measurable_extendarmstreamfiniteprefix banditrlproof.ucb.measurable_extendarmstreamfiniteprefix theorem measurable_extendarmstreamfiniteprefix {k : nat} (n : nat) : measurable (extendarmstreamfiniteprefix (k := k) n) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.extendArmStreamFinitePrefix_apply_of_le","label":"extendArmStreamFinitePrefix_apply_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.extendArmStreamFinitePrefix_apply_of_le","description":"theorem extendArmStreamFinitePrefix_apply_of_le {K : Nat} (n t : Nat) (ht : t <= n) (streamBox : (i : Finset.Iic n) -> Fin K -> Real) : extendArmStreamFinitePrefix n streamBox t = streamBox ⟨t, Finset.mem_Iic.mpr ht⟩","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-6a0e21fcd6a0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1974,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:68"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem extendArmStreamFinitePrefix_apply_of_le {K : Nat} (n t : Nat) (ht : t <= n) (streamBox : (i : Finset.Iic n) -> Fin K -> Real) : extendArmStreamFinitePrefix n streamBox t = streamBox ⟨t, Finset.mem_Iic.mpr ht⟩","missing":[],"search":"extendarmstreamfiniteprefix_apply_of_le banditrlproof.ucb.extendarmstreamfiniteprefix_apply_of_le theorem extendarmstreamfiniteprefix_apply_of_le {k : nat} (n t : nat) (ht : t <= n) (streambox : (i : finset.iic n) -> fin k -> real) : extendarmstreamfiniteprefix n streambox t = streambox ⟨t, finset.mem_iic.mpr ht⟩ theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod","label":"armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod","description":"A finite kernel driven only by the complement of one arm-stream coordinate leaves that coordinate's prescribed marginal independent of the kernel output. The output kernel may be sub-Markov, as needed after branch restriction.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-a6cd1a78f02d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1975,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:79"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod {K : Nat} {Output : Type*} [MeasurableSpace Output] (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (target : Nat × Fin K) (kernel : Kernel ({index : Nat × Fin K // index ≠ target} -> Real) Output) [IsFiniteKernel kernel] : Measure.map (fun sample : ArmRewardStream K × Output => (sample.2, armStreamCoordinate target sample.1)) (armStreamMeasure nu ⊗ₘ kernel.comap (armStreamWithoutCoordinate target) (measurable_armStreamWithoutCoordinate target)) = (Measure.map Prod.snd (armStreamMeasure nu ⊗ₘ kernel.comap (armStreamWithoutCoordinate target) (measurable_armStreamWithoutCoordinate target))).prod (nu target.2)","missing":[],"search":"armstreammeasure_map_output_coordinate_compprod_comap_without_eq_prod banditrlproof.ucb.armstreammeasure_map_output_coordinate_compprod_comap_without_eq_prod a finite kernel driven only by the complement of one arm-stream coordinate leaves that coordinate's prescribed marginal independent of the kernel output. the output kernel may be sub-markov, as needed after branch restriction. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_prefix_next_eq_compProd","label":"latentArmStreamTrajectoryKernel_map_prefix_next_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_prefix_next_eq_compProd","description":"Fixed-stream finite-prefix/next-pair recursion for the latent arm-stream trajectory. This is the trajectory-level recurrence used by the count-capped locality induction below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-d4045c359afd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1976,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:119"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryKernel_map_prefix_next_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (stream : UCB.ArmRewardStream K) (n : Nat) : (latentArmStreamTrajectoryKernel algorithm env stream).map (fun trajectory => (Preorder.frestrictLe n trajectory, trajectory (n + 1))) = (latentArmStreamTrajectoryKernel algorithm env stream).map (Preorder.frestrictLe n) ⊗ₘ historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream)) n","missing":[],"search":"latentarmstreamtrajectorykernel_map_prefix_next_eq_compprod banditrlproof.thompson.latentarmstreamtrajectorykernel_map_prefix_next_eq_compprod fixed-stream finite-prefix/next-pair recursion for the latent arm-stream trajectory. this is the trajectory-level recurrence used by the count-capped locality induction below. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamFeedback_eq_of_withoutCoordinate_eq_of_selectedCoordinate_ne","label":"latentArmStreamFeedback_eq_of_withoutCoordinate_eq_of_selectedCoordinate_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamFeedback_eq_of_withoutCoordinate_eq_of_selectedCoordinate_ne","description":"If two latent reward streams agree away from one coordinate, then their feedback laws agree at every history/action pair whose next-unused coordinate is not the omitted one.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-1fedd6b70acf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1977,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:143"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamFeedback_eq_of_withoutCoordinate_eq_of_selectedCoordinate_ne {Env : Type u} {K : Nat} [MeasurableSpace Env] (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) (hne : (ETC.realHistoryPullCount n history arm, arm) ≠ target) : ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₁)).feedback n (history, arm) = ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₂)).feedback n (history, arm)","missing":[],"search":"latentarmstreamfeedback_eq_of_withoutcoordinate_eq_of_selectedcoordinate_ne banditrlproof.thompson.latentarmstreamfeedback_eq_of_withoutcoordinate_eq_of_selectedcoordinate_ne if two latent reward streams agree away from one coordinate, then their feedback laws agree at every history/action pair whose next-unused coordinate is not the omitted one. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_eq_of_withoutCoordinate_eq_of_target_count_lt","label":"historyStepKernel_apply_eq_of_withoutCoordinate_eq_of_target_count_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel_apply_eq_of_withoutCoordinate_eq_of_target_count_lt","description":"Before the target coordinate can be the next coordinate of its arm, the entire next-pair law agrees for streams with the same omitted-coordinate projection.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-8b8815d057bb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1978,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:169"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernel_apply_eq_of_withoutCoordinate_eq_of_target_count_lt {Env : Type u} {K : Nat} [MeasurableSpace Env] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) (hcount : ETC.realHistoryPullCount n history target.2 < target.1) : historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₁)) n history = historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₂)) n history","missing":[],"search":"historystepkernel_apply_eq_of_withoutcoordinate_eq_of_target_count_lt banditrlproof.thompson.historystepkernel_apply_eq_of_withoutcoordinate_eq_of_target_count_lt before the target coordinate can be the next coordinate of its arm, the entire next-pair law agrees for streams with the same omitted-coordinate projection. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamNextActionNeSet","label":"latentArmStreamNextActionNeSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamNextActionNeSet","description":"Next pairs whose action avoids a designated arm.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-bb6dd594bbf1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1979,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:200"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamNextActionNeSet {K : Nat} (arm : Fin K) : Set (Fin K × Real)","missing":[],"search":"latentarmstreamnextactionneset banditrlproof.thompson.latentarmstreamnextactionneset next pairs whose action avoids a designated arm. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamInitialSafeArmSet","label":"latentArmStreamInitialSafeArmSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamInitialSafeArmSet","description":"Initial actions whose time-zero reward coordinate is not the omitted target.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-710a1fc6cf0a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1980,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:206"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamInitialSafeArmSet {K : Nat} (target : Nat × Fin K) : Set (Fin K)","missing":[],"search":"latentarmstreaminitialsafearmset banditrlproof.thompson.latentarmstreaminitialsafearmset initial actions whose time-zero reward coordinate is not the omitted target. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamNextActionNeSet","label":"measurableSet_latentArmStreamNextActionNeSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableSet_latentArmStreamNextActionNeSet","description":"theorem measurableSet_latentArmStreamNextActionNeSet {K : Nat} (arm : Fin K) : MeasurableSet (latentArmStreamNextActionNeSet arm)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-90e5a9f73743","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1981,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:210"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_latentArmStreamNextActionNeSet {K : Nat} (arm : Fin K) : MeasurableSet (latentArmStreamNextActionNeSet arm)","missing":[],"search":"measurableset_latentarmstreamnextactionneset banditrlproof.thompson.measurableset_latentarmstreamnextactionneset theorem measurableset_latentarmstreamnextactionneset {k : nat} (arm : fin k) : measurableset (latentarmstreamnextactionneset arm) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_restrict_nextActionNe_eq_of_withoutCoordinate_eq","label":"historyStepKernel_apply_restrict_nextActionNe_eq_of_withoutCoordinate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel_apply_restrict_nextActionNe_eq_of_withoutCoordinate_eq","description":"Even when the target pull index may already be next, the part of the next-pair law selecting another arm remains independent of the target coordinate.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-fab6d5d2c85f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1982,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:218"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernel_apply_restrict_nextActionNe_eq_of_withoutCoordinate_eq {Env : Type u} {K : Nat} [MeasurableSpace Env] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) : (historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₁)) n history).restrict (latentArmStreamNextActionNeSet target.2) = (historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₂)) n history).restrict (latentArmStreamNextActionNeSet target.2)","missing":[],"search":"historystepkernel_apply_restrict_nextactionne_eq_of_withoutcoordinate_eq banditrlproof.thompson.historystepkernel_apply_restrict_nextactionne_eq_of_withoutcoordinate_eq even when the target pull index may already be next, the part of the next-pair law selecting another arm remains independent of the target coordinate. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_zero","label":"latentArmStreamTrajectoryKernel_map_frestrictLe_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_zero","description":"The time-zero visible prefix is the initial action/feedback pair pushed through the singleton-history constructor.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-16264ec4bd71","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1983,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:265"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryKernel_map_frestrictLe_zero {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (stream : UCB.ArmRewardStream K) : (latentArmStreamTrajectoryKernel algorithm env stream).map (Preorder.frestrictLe 0) = (algorithm.initialAction ⊗ₘ ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream)).initialFeedback).map singletonPairHistory","missing":[],"search":"latentarmstreamtrajectorykernel_map_frestrictle_zero banditrlproof.thompson.latentarmstreamtrajectorykernel_map_frestrictle_zero the time-zero visible prefix is the initial action/feedback pair pushed through the singleton-history constructor. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq","label":"latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq","description":"Two latent streams agreeing through `n` generate the same visible trajectory law through `n`. Action randomization remains inside the canonical trajectory kernel; the argument changes only deterministic reward fibers.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-088ec7eb059d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1984,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:319"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (stream₁ stream₂ : UCB.ArmRewardStream K) (n : Nat) (hstream : Preorder.frestrictLe n stream₁ = Preorder.frestrictLe n stream₂) : (latentArmStreamTrajectoryKernel algorithm env stream₁).map (Preorder.frestrictLe n) = (latentArmStreamTrajectoryKernel algorithm env stream₂).map (Preorder.frestrictLe n)","missing":[],"search":"latentarmstreamtrajectorykernel_map_frestrictle_eq_of_streamprefix_eq banditrlproof.thompson.latentarmstreamtrajectorykernel_map_frestrictle_eq_of_streamprefix_eq two latent streams agreeing through `n` generate the same visible trajectory law through `n`. action randomization remains inside the canonical trajectory kernel; the argument changes only deterministic reward fibers. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixKernel","label":"latentArmStreamVisiblePrefixKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixKernel","description":"Finite visible-prefix kernel after replacing the infinite reward stream by its zero-extended finite box.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-943e7054c566","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1985,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:385"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def latentArmStreamVisiblePrefixKernel {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (n : Nat) : Kernel ((i : Finset.Iic n) -> Fin K -> Real) (History.FinitePairHistory (Fin K) Real n)","missing":[],"search":"latentarmstreamvisibleprefixkernel banditrlproof.thompson.latentarmstreamvisibleprefixkernel finite visible-prefix kernel after replacing the infinite reward stream by its zero-extended finite box. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap","label":"latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap","description":"The latent visible-prefix kernel factors exactly through the finite stream box through the same endpoint.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-8a30601bc48c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1986,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:409"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (n : Nat) : (latentArmStreamTrajectoryKernel algorithm env).map (Preorder.frestrictLe n) = (latentArmStreamVisiblePrefixKernel algorithm env n).comap (Preorder.frestrictLe n) (Preorder.measurable_frestrictLe n)","missing":[],"search":"latentarmstreamtrajectorykernel_map_frestrictle_eq_prefixkernel_comap banditrlproof.thompson.latentarmstreamtrajectorykernel_map_frestrictle_eq_prefixkernel_comap the latent visible-prefix kernel factors exactly through the finite stream box through the same endpoint. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction","label":"latentArmStreamVisiblePrefixNextAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction","description":"Visible history through `n` paired with the action selected at `n + 1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-388cb0432568","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1987,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:438"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamVisiblePrefixNextAction {K : Nat} (n : Nat) : ((t : Nat) -> Fin K × Real) -> History.FinitePairHistory (Fin K) Real n × Fin K","missing":[],"search":"latentarmstreamvisibleprefixnextaction banditrlproof.thompson.latentarmstreamvisibleprefixnextaction visible history through `n` paired with the action selected at `n + 1`. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamVisiblePrefixNextAction","label":"measurable_latentArmStreamVisiblePrefixNextAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamVisiblePrefixNextAction","description":"theorem measurable_latentArmStreamVisiblePrefixNextAction {K : Nat} (n : Nat) : Measurable (latentArmStreamVisiblePrefixNextAction (K := K) n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-9ee04539f258","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1988,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:445"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamVisiblePrefixNextAction {K : Nat} (n : Nat) : Measurable (latentArmStreamVisiblePrefixNextAction (K := K) n)","missing":[],"search":"measurable_latentarmstreamvisibleprefixnextaction banditrlproof.thompson.measurable_latentarmstreamvisibleprefixnextaction theorem measurable_latentarmstreamvisibleprefixnextaction {k : nat} (n : nat) : measurable (latentarmstreamvisibleprefixnextaction (k := k) n) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward","label":"latentArmStreamVisibleNextReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleNextReward","description":"Reward observed at the shifted successor time `n + 1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-18b065325106","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1989,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:452"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamVisibleNextReward {K : Nat} (n : Nat) : ((t : Nat) -> Fin K × Real) -> Real","missing":[],"search":"latentarmstreamvisiblenextreward banditrlproof.thompson.latentarmstreamvisiblenextreward reward observed at the shifted successor time `n + 1`. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamVisibleNextReward","label":"measurable_latentArmStreamVisibleNextReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamVisibleNextReward","description":"theorem measurable_latentArmStreamVisibleNextReward {K : Nat} (n : Nat) : Measurable (latentArmStreamVisibleNextReward (K := K) n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-65288b130450","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1990,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:457"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamVisibleNextReward {K : Nat} (n : Nat) : Measurable (latentArmStreamVisibleNextReward (K := K) n)","missing":[],"search":"measurable_latentarmstreamvisiblenextreward banditrlproof.thompson.measurable_latentarmstreamvisiblenextreward theorem measurable_latentarmstreamvisiblenextreward {k : nat} (n : nat) : measurable (latentarmstreamvisiblenextreward (k := k) n) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchKernel","label":"latentArmStreamVisiblePrefixNextActionBranchKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchKernel","description":"Canonical candidate for the condition kernel on one next-coordinate branch, reconstructed after fixing the omitted coordinate to zero.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-f33f06374be7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1991,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:464"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def latentArmStreamVisiblePrefixNextActionBranchKernel {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (n : Nat) (target : Nat × Fin K) : Kernel ({index : Nat × Fin K // index ≠ target} -> Real) (History.FinitePairHistory (Fin K) Real n × Fin K)","missing":[],"search":"latentarmstreamvisibleprefixnextactionbranchkernel banditrlproof.thompson.latentarmstreamvisibleprefixnextactionbranchkernel canonical candidate for the condition kernel on one next-coordinate branch, reconstructed after fixing the omitted coordinate to zero. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCap","label":"latentArmStreamPrefixCountCap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamPrefixCountCap","description":"Histories through `n` that have not consumed the designated latent coordinate. The inclusive history count can equal the coordinate index: that coordinate is used only by the next reward from the same arm.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-72d78749f01b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1992,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:491"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamPrefixCountCap {K : Nat} (n : Nat) (target : Nat × Fin K) : Set (History.FinitePairHistory (Fin K) Real n)","missing":[],"search":"latentarmstreamprefixcountcap banditrlproof.thompson.latentarmstreamprefixcountcap histories through `n` that have not consumed the designated latent coordinate. the inclusive history count can equal the coordinate index: that coordinate is used only by the next reward from the same arm. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountCap","label":"measurableSet_latentArmStreamPrefixCountCap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountCap","description":"theorem measurableSet_latentArmStreamPrefixCountCap {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (latentArmStreamPrefixCountCap n target)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-22d2fced79de","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1993,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:497"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_latentArmStreamPrefixCountCap {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (latentArmStreamPrefixCountCap n target)","missing":[],"search":"measurableset_latentarmstreamprefixcountcap banditrlproof.thompson.measurableset_latentarmstreamprefixcountcap theorem measurableset_latentarmstreamprefixcountcap {k : nat} (n : nat) (target : nat × fin k) : measurableset (latentarmstreamprefixcountcap n target) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.singletonPairHistory_preimage_latentArmStreamPrefixCountCap_zero","label":"singletonPairHistory_preimage_latentArmStreamPrefixCountCap_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.singletonPairHistory_preimage_latentArmStreamPrefixCountCap_zero","description":"Pulling the time-zero count cap back through the singleton-history constructor leaves exactly the actions whose initial reward coordinate is not the omitted target.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-8cd9445aca1b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1994,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:506"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem singletonPairHistory_preimage_latentArmStreamPrefixCountCap_zero {K : Nat} (target : Nat × Fin K) : (@singletonPairHistory (Fin K) Real) ⁻¹' latentArmStreamPrefixCountCap 0 target = latentArmStreamInitialSafeArmSet target ×ˢ Set.univ","missing":[],"search":"singletonpairhistory_preimage_latentarmstreamprefixcountcap_zero banditrlproof.thompson.singletonpairhistory_preimage_latentarmstreamprefixcountcap_zero pulling the time-zero count cap back through the singleton-history constructor leaves exactly the actions whose initial reward coordinate is not the omitted target. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCapLocality_zero","label":"latentArmStreamPrefixCountCapLocality_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamPrefixCountCapLocality_zero","description":"Count-cap locality at the initial pair.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-b2c9a96360ca","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1995,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:526"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamPrefixCountCapLocality_zero {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) : ((latentArmStreamTrajectoryKernel algorithm env stream₁).map (Preorder.frestrictLe 0)).restrict (latentArmStreamPrefixCountCap 0 target) = ((latentArmStreamTrajectoryKernel algorithm env stream₂).map (Preorder.frestrictLe 0)).restrict (latentArmStreamPrefixCountCap 0 target)","missing":[],"search":"latentarmstreamprefixcountcaplocality_zero banditrlproof.thompson.latentarmstreamprefixcountcaplocality_zero count-cap locality at the initial pair. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountLt","label":"latentArmStreamPrefixCountLt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamPrefixCountLt","description":"Strict count region used by the first rectangle in the successor cap decomposition.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-9908f682afac","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1996,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:577"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamPrefixCountLt {K : Nat} (n : Nat) (target : Nat × Fin K) : Set (History.FinitePairHistory (Fin K) Real n)","missing":[],"search":"latentarmstreamprefixcountlt banditrlproof.thompson.latentarmstreamprefixcountlt strict count region used by the first rectangle in the successor cap decomposition. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountLt","label":"measurableSet_latentArmStreamPrefixCountLt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountLt","description":"theorem measurableSet_latentArmStreamPrefixCountLt {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (latentArmStreamPrefixCountLt n target)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-47b3b643d31a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1997,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:582"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_latentArmStreamPrefixCountLt {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (latentArmStreamPrefixCountLt n target)","missing":[],"search":"measurableset_latentarmstreamprefixcountlt banditrlproof.thompson.measurableset_latentarmstreamprefixcountlt theorem measurableset_latentarmstreamprefixcountlt {k : nat} (n : nat) (target : nat × fin k) : measurableset (latentarmstreamprefixcountlt n target) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountEq","label":"latentArmStreamPrefixCountEq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamPrefixCountEq","description":"Exact count region used by the second rectangle in the successor cap decomposition.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-733124a22f80","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1998,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:590"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamPrefixCountEq {K : Nat} (n : Nat) (target : Nat × Fin K) : Set (History.FinitePairHistory (Fin K) Real n)","missing":[],"search":"latentarmstreamprefixcounteq banditrlproof.thompson.latentarmstreamprefixcounteq exact count region used by the second rectangle in the successor cap decomposition. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountEq","label":"measurableSet_latentArmStreamPrefixCountEq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountEq","description":"theorem measurableSet_latentArmStreamPrefixCountEq {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (latentArmStreamPrefixCountEq n target)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-e77086cd1410","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":1999,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:595"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_latentArmStreamPrefixCountEq {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (latentArmStreamPrefixCountEq n target)","missing":[],"search":"measurableset_latentarmstreamprefixcounteq banditrlproof.thompson.measurableset_latentarmstreamprefixcounteq theorem measurableset_latentarmstreamprefixcounteq {k : nat} (n : nat) (target : nat × fin k) : measurableset (latentarmstreamprefixcounteq n target) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.realHistoryPullCount_extendPairHistorySucc","label":"realHistoryPullCount_extendPairHistorySucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.realHistoryPullCount_extendPairHistorySucc","description":"Extending an inclusive finite pair history increments exactly the count of the arm selected by the appended pair.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-60e7c147ab4c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2000,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:604"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem realHistoryPullCount_extendPairHistorySucc {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (next : Fin K × Real) (arm : Fin K) : ETC.realHistoryPullCount (n + 1) (History.extendPairHistorySucc history next) arm = ETC.realHistoryPullCount n history arm + if next.1 = arm then 1 else 0","missing":[],"search":"realhistorypullcount_extendpairhistorysucc banditrlproof.thompson.realhistorypullcount_extendpairhistorysucc extending an inclusive finite pair history increments exactly the count of the arm selected by the appended pair. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.mem_latentArmStreamPrefixCountCap_extendPairHistorySucc_iff","label":"mem_latentArmStreamPrefixCountCap_extendPairHistorySucc_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.mem_latentArmStreamPrefixCountCap_extendPairHistorySucc_iff","description":"The successor prefix remains below the target count cap exactly when the old prefix is capped and either has strict slack or avoids the target arm.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-0e7aedeae510","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2001,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:651"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_latentArmStreamPrefixCountCap_extendPairHistorySucc_iff {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (next : Fin K × Real) (target : Nat × Fin K) : History.extendPairHistorySucc history next ∈ latentArmStreamPrefixCountCap (n + 1) target ↔ history ∈ latentArmStreamPrefixCountCap n target ∧ (ETC.realHistoryPullCount n history target.2 < target.1 ∨ next.1 ≠ target.2)","missing":[],"search":"mem_latentarmstreamprefixcountcap_extendpairhistorysucc_iff banditrlproof.thompson.mem_latentarmstreamprefixcountcap_extendpairhistorysucc_iff the successor prefix remains below the target count cap exactly when the old prefix is capped and either has strict slack or avoids the target arm. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCap_of_extendPairHistorySucc_mem","label":"latentArmStreamPrefixCountCap_of_extendPairHistorySucc_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamPrefixCountCap_of_extendPairHistorySucc_mem","description":"A capped successor prefix was already capped before appending its last action/reward pair.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-84dff706da36","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2002,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:670"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamPrefixCountCap_of_extendPairHistorySucc_mem {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (next : Fin K × Real) (target : Nat × Fin K) (hcap : History.extendPairHistorySucc history next ∈ latentArmStreamPrefixCountCap (n + 1) target) : history ∈ latentArmStreamPrefixCountCap n target","missing":[],"search":"latentarmstreamprefixcountcap_of_extendpairhistorysucc_mem banditrlproof.thompson.latentarmstreamprefixcountcap_of_extendpairhistorysucc_mem a capped successor prefix was already capped before appending its last action/reward pair. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.selectedCoordinate_ne_of_extendPairHistorySucc_mem_prefixCountCap","label":"selectedCoordinate_ne_of_extendPairHistorySucc_mem_prefixCountCap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.selectedCoordinate_ne_of_extendPairHistorySucc_mem_prefixCountCap","description":"Every appended pair that remains in the successor cap reads a stream coordinate different from the omitted target.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-fded07654e58","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2003,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:683"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem selectedCoordinate_ne_of_extendPairHistorySucc_mem_prefixCountCap {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (next : Fin K × Real) (target : Nat × Fin K) (hcap : History.extendPairHistorySucc history next ∈ latentArmStreamPrefixCountCap (n + 1) target) : (ETC.realHistoryPullCount n history next.1, next.1) ≠ target","missing":[],"search":"selectedcoordinate_ne_of_extendpairhistorysucc_mem_prefixcountcap banditrlproof.thompson.selectedcoordinate_ne_of_extendpairhistorysucc_mem_prefixcountcap every appended pair that remains in the successor cap reads a stream coordinate different from the omitted target. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamSuccessorCountCap_preimage","label":"latentArmStreamSuccessorCountCap_preimage","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamSuccessorCountCap_preimage","description":"Pulling the successor count cap back through history extension gives two measurable rectangles: below the target count every action is safe, while at the target count only actions avoiding the target arm are safe.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-dc0c6a82cfe6","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2004,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:703"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamSuccessorCountCap_preimage {K : Nat} (n : Nat) (target : Nat × Fin K) : (fun sample : History.FinitePairHistory (Fin K) Real n × (Fin K × Real) => History.extendPairHistorySucc sample.1 sample.2) ⁻¹' latentArmStreamPrefixCountCap (n + 1) target = (latentArmStreamPrefixCountLt n target ×ˢ Set.univ) ∪ (latentArmStreamPrefixCountEq n target ×ˢ latentArmStreamNextActionNeSet target.2)","missing":[],"search":"latentarmstreamsuccessorcountcap_preimage banditrlproof.thompson.latentarmstreamsuccessorcountcap_preimage pulling the successor count cap back through history extension gives two measurable rectangles: below the target count every action is safe, while at the target count only actions avoiding the target arm are safe. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamSuccessorCountCapSection","label":"latentArmStreamSuccessorCountCapSection","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamSuccessorCountCapSection","description":"Next action/reward pairs that keep a fixed old prefix below the successor count cap.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-60ad06be7479","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2005,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:724"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def latentArmStreamSuccessorCountCapSection {K : Nat} (n : Nat) (target : Nat × Fin K) (history : History.FinitePairHistory (Fin K) Real n) : Set (Fin K × Real)","missing":[],"search":"latentarmstreamsuccessorcountcapsection banditrlproof.thompson.latentarmstreamsuccessorcountcapsection next action/reward pairs that keep a fixed old prefix below the successor count cap. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamSuccessorCountCapSection","label":"measurableSet_latentArmStreamSuccessorCountCapSection","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableSet_latentArmStreamSuccessorCountCapSection","description":"theorem measurableSet_latentArmStreamSuccessorCountCapSection {K : Nat} (n : Nat) (target : Nat × Fin K) (history : History.FinitePairHistory (Fin K) Real n) : MeasurableSet (latentArmStreamSuccessorCountCapSection n target history)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-cac4225c8436","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2006,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:732"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_latentArmStreamSuccessorCountCapSection {K : Nat} (n : Nat) (target : Nat × Fin K) (history : History.FinitePairHistory (Fin K) Real n) : MeasurableSet (latentArmStreamSuccessorCountCapSection n target history)","missing":[],"search":"measurableset_latentarmstreamsuccessorcountcapsection banditrlproof.thompson.measurableset_latentarmstreamsuccessorcountcapsection theorem measurableset_latentarmstreamsuccessorcountcapsection {k : nat} (n : nat) (target : nat × fin k) (history : history.finitepairhistory (fin k) real n) : measurableset (latentarmstreamsuccessorcountcapsection n target history) theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_restrict_successorCountCap_eq_of_withoutCoordinate_eq","label":"historyStepKernel_apply_restrict_successorCountCap_eq_of_withoutCoordinate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel_apply_restrict_successorCountCap_eq_of_withoutCoordinate_eq","description":"Over a previously capped prefix, the one-step law restricted to capped successors depends only on the complement of the omitted stream coordinate.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-f71bbb1a5ff2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2007,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:743"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernel_apply_restrict_successorCountCap_eq_of_withoutCoordinate_eq {Env : Type u} {K : Nat} [MeasurableSpace Env] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) (hcap : history ∈ latentArmStreamPrefixCountCap n target) : (historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₁)) n history).restrict (latentArmStreamSuccessorCountCapSection n target history) = (historyStepKernel algorithm ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream₂)) n history).restrict (latentArmStreamSuccessorCountCapSection n target history)","missing":[],"search":"historystepkernel_apply_restrict_successorcountcap_eq_of_withoutcoordinate_eq banditrlproof.thompson.historystepkernel_apply_restrict_successorcountcap_eq_of_withoutcoordinate_eq over a previously capped prefix, the one-step law restricted to capped successors depends only on the complement of the omitted stream coordinate. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_succ","label":"latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_succ","description":"Successor step for count-capped locality of the fixed-stream visible trajectory prefix law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-887672d948b0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2008,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:779"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_succ {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) (n : Nat) (hprefix : ((latentArmStreamTrajectoryKernel algorithm env stream₁).map (Preorder.frestrictLe n)).restrict (latentArmStreamPrefixCountCap n target) = ((latentArmStreamTrajectoryKernel algorithm env stream₂).map (Preorder.frestrictLe n)).restrict (latentArmStreamPrefixCountCap n target)) : ((latentArmStreamTrajectoryKernel algorithm env stream₁).map (Preorder.frestrictLe (n + 1))).restrict (latentArmStreamPrefixCountCap (n + 1) target) = ((latentArmStreamTrajectoryKernel algorithm env stream₂).m…","missing":[],"search":"latentarmstreamtrajectorykernel_map_frestrictle_restrict_countcap_succ banditrlproof.thompson.latentarmstreamtrajectorykernel_map_frestrictle_restrict_countcap_succ successor step for count-capped locality of the fixed-stream visible trajectory prefix law. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq","label":"latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq","description":"Count-capped fixed-stream prefix locality under equality of all latent coordinates except the omitted target.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-ae0301d5a15f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2009,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:861"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K) (hwithout : UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂) (n : Nat) : ((latentArmStreamTrajectoryKernel algorithm env stream₁).map (Preorder.frestrictLe n)).restrict (latentArmStreamPrefixCountCap n target) = ((latentArmStreamTrajectoryKernel algorithm env stream₂).map (Preorder.frestrictLe n)).restrict (latentArmStreamPrefixCountCap n target)","missing":[],"search":"latentarmstreamtrajectorykernel_map_frestrictle_restrict_countcap_eq_of_withoutcoordinate_eq banditrlproof.thompson.latentarmstreamtrajectorykernel_map_frestrictle_restrict_countcap_eq_of_withoutcoordinate_eq count-capped fixed-stream prefix locality under equality of all latent coordinates except the omitted target. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.LatentArmStreamVisiblePrefixNextActionBranchLocality","label":"LatentArmStreamVisiblePrefixNextActionBranchLocality","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.LatentArmStreamVisiblePrefixNextActionBranchLocality","description":"Exact contract for branchwise prefix/action locality. The count-capped trajectory induction above proves this contract below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-334d57e43247","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2010,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:886"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def LatentArmStreamVisiblePrefixNextActionBranchLocality {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (n : Nat) : Prop","missing":[],"search":"latentarmstreamvisibleprefixnextactionbranchlocality banditrlproof.thompson.latentarmstreamvisibleprefixnextactionbranchlocality exact contract for branchwise prefix/action locality. the count-capped trajectory induction above proves this contract below. definition compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality_of_prefixCountCapLocality","label":"latentArmStreamVisiblePrefixNextActionBranchLocality_of_prefixCountCapLocality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality_of_prefixCountCapLocality","description":"Count-cap locality of visible prefixes implies the exact history/action branch-locality contract.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-d397c0229889","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2011,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:902"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextActionBranchLocality_of_prefixCountCapLocality {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (n : Nat) (hcap : ∀ (target : Nat × Fin K) (stream₁ stream₂ : UCB.ArmRewardStream K), UCB.armStreamWithoutCoordinate target stream₁ = UCB.armStreamWithoutCoordinate target stream₂ → (((latentArmStreamTrajectoryKernel algorithm env stream₁).map (Preorder.frestrictLe n)).restrict (latentArmStreamPrefixCountCap n target)) = (((latentArmStreamTrajectoryKernel algorithm env stream₂).map (Preorder.frestrictLe n)).restrict (latentArmStreamPrefixCountCap n target))) : LatentArmStreamVisiblePrefixNextActionBranchLocality algorithm env n","missing":[],"search":"latentarmstreamvisibleprefixnextactionbranchlocality_of_prefixcountcaplocality banditrlproof.thompson.latentarmstreamvisibleprefixnextactionbranchlocality_of_prefixcountcaplocality count-cap locality of visible prefixes implies the exact history/action branch-locality contract. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality","label":"latentArmStreamVisiblePrefixNextActionBranchLocality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality","description":"Every latent arm-stream trajectory kernel satisfies branch locality: on the exact next-coordinate branch, the visible prefix and next action depend only on the complementary latent coordinates. This does not yet identify the latent visible law with the native fixed-IID process.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-250f22c0fa35","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2012,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:984"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextActionBranchLocality {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (n : Nat) : LatentArmStreamVisiblePrefixNextActionBranchLocality algorithm env n","missing":[],"search":"latentarmstreamvisibleprefixnextactionbranchlocality banditrlproof.thompson.latentarmstreamvisibleprefixnextactionbranchlocality every latent arm-stream trajectory kernel satisfies branch locality: on the exact next-coordinate branch, the visible prefix and next action depend only on the complementary latent coordinates. this does not yet identify the latent visible law with the native fixed-iid process. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod_of_locality","label":"latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod_of_locality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod_of_locality","description":"Once the branch-locality producer is available, coordinate independence immediately yields the exact branchwise condition/selected-coordinate product law. This consumer is valid for the sub-Markov branch kernel and therefore does not misstate branch restriction as ordinary probability independence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-12ad327c3f5a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2013,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1002"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod_of_locality {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) (hlocal : LatentArmStreamVisiblePrefixNextActionBranchLocality algorithm env n) (target : Nat × Fin K) : let branchKernel := ((latentArmStreamTrajectoryKernel algorithm env).map (latentArmStreamVisiblePrefixNextAction n)).restrict (UCB.measurableSet_armStreamHistoryActionCoordinateBranch n target) Measure.map (fun sample : UCB.ArmRewardStream K × (History.FinitePairHistory (Fin K) Real n × Fin K) => (sample.2, UCB.armStreamCoordinate target sample.1)) (UCB.armStreamMeasure nu ⊗ₘ branchKernel) = (Measure.map Prod.snd (UCB.armStreamMeasure nu ⊗ₘ branchKernel)).prod (nu target.2)","missing":[],"search":"latentarmstreamvisibleprefixnextaction_coordinate_branch_eq_prod_of_locality banditrlproof.thompson.latentarmstreamvisibleprefixnextaction_coordinate_branch_eq_prod_of_locality once the branch-locality producer is available, coordinate independence immediately yields the exact branchwise condition/selected-coordinate product law. this consumer is valid for the sub-markov branch kernel and therefore does not misstate branch restriction as ordinary probability independence. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","label":"latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","description":"Unconditional compiled branchwise product law obtained from the proved count-capped locality producer.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-d02c065ead41","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2014,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1033"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) (target : Nat × Fin K) : let branchKernel := ((latentArmStreamTrajectoryKernel algorithm env).map (latentArmStreamVisiblePrefixNextAction n)).restrict (UCB.measurableSet_armStreamHistoryActionCoordinateBranch n target) Measure.map (fun sample : UCB.ArmRewardStream K × (History.FinitePairHistory (Fin K) Real n × Fin K) => (sample.2, UCB.armStreamCoordinate target sample.1)) (UCB.armStreamMeasure nu ⊗ₘ branchKernel) = (Measure.map Prod.snd (UCB.armStreamMeasure nu ⊗ₘ branchKernel)).prod (nu target.2)","missing":[],"search":"latentarmstreamvisibleprefixnextaction_coordinate_branch_eq_prod banditrlproof.thompson.latentarmstreamvisibleprefixnextaction_coordinate_branch_eq_prod unconditional compiled branchwise product law obtained from the proved count-capped locality producer. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","label":"latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","description":"After the latent reward stream is mixed out, the next action still follows the algorithm's history policy. This isolates action randomization from the remaining selected-reward freshness obligation.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-ff488b617f94","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2015,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1060"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => latentArmStreamVisiblePrefixNextAction n sample.2) (latentArmStreamTrajectoryMeasure algorithm env nu) = Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => Preorder.frestrictLe n sample.2) (latentArmStreamTrajectoryMeasure algorithm env nu) ⊗ₘ algorithm.policy n","missing":[],"search":"latentarmstreamtrajectorymeasure_map_visibleprefix_nextaction_eq_compprod banditrlproof.thompson.latentarmstreamtrajectorymeasure_map_visibleprefix_nextaction_eq_compprod after the latent reward stream is mixed out, the next action still follows the algorithm's history policy. this isolates action randomization from the remaining selected-reward freshness obligation. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae","label":"latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae","description":"On the coupling, the actual successor reward is the latent coordinate encoded by the visible prefix and the sampled next action. This is pathwise support; conditional freshness still requires the branchwise product proof.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-80ee5b28044b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2016,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1096"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : ∀ᵐ sample ∂latentArmStreamTrajectoryMeasure algorithm env nu, latentArmStreamVisibleNextReward n sample.2 = UCB.armStreamCoordinate (UCB.armStreamCoordinateOfHistoryAction n (latentArmStreamVisiblePrefixNextAction n sample.2)) sample.1","missing":[],"search":"latentarmstreamvisiblenextreward_eq_selectedcoordinate_ae banditrlproof.thompson.latentarmstreamvisiblenextreward_eq_selectedcoordinate_ae on the coupling, the actual successor reward is the latent coordinate encoded by the visible prefix and the sampled next action. this is pathwise support; conditional freshness still requires the branchwise product proof. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamSelectedCoordinate","label":"measurable_latentArmStreamSelectedCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamSelectedCoordinate","description":"The reward coordinate selected by a visible history/action condition is a measurable function of the latent stream and that condition.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-0535da3f5c80","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2017,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1129"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamSelectedCoordinate {K : Nat} (n : Nat) : Measurable (fun sample : UCB.ArmRewardStream K × (History.FinitePairHistory (Fin K) Real n × Fin K) => UCB.armStreamCoordinate (UCB.armStreamCoordinateOfHistoryAction n sample.2) sample.1)","missing":[],"search":"measurable_latentarmstreamselectedcoordinate banditrlproof.thompson.measurable_latentarmstreamselectedcoordinate the reward coordinate selected by a visible history/action condition is a measurable function of the latent stream and that condition. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_branch_eq_prod","label":"latentArmStreamVisiblePrefixNextAction_selectedCoordinate_branch_eq_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_branch_eq_prod","description":"On a fixed pull-count/arm branch, the dynamically selected latent reward coordinate has the restricted condition marginal times the selected arm law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-2f9421b1ca8b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2018,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1148"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextAction_selectedCoordinate_branch_eq_prod {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) (target : Nat × Fin K) : let conditionKernel := (latentArmStreamTrajectoryKernel algorithm env).map (latentArmStreamVisiblePrefixNextAction n) let fullMixed := UCB.armStreamMeasure nu ⊗ₘ conditionKernel let branch := UCB.armStreamHistoryActionCoordinateBranch n target let branchKernel := conditionKernel.restrict (UCB.measurableSet_armStreamHistoryActionCoordinateBranch n target) let branchMixed := UCB.armStreamMeasure nu ⊗ₘ branchKernel Measure.map (fun sample : UCB.ArmRewardStream K × (History.FinitePairHistory (Fin K) Real n × Fin K) => (sample.2, UCB.armStreamCoordinate (UCB.armStreamCoordinateOfHistoryAction n…","missing":[],"search":"latentarmstreamvisibleprefixnextaction_selectedcoordinate_branch_eq_prod banditrlproof.thompson.latentarmstreamvisibleprefixnextaction_selectedcoordinate_branch_eq_prod on a fixed pull-count/arm branch, the dynamically selected latent reward coordinate has the restricted condition marginal times the selected arm law. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd","label":"latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd","description":"Summing the countable pull-count/arm partition gives the selected-coordinate joint law under the mixed latent-stream/visible-condition measure.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-48b5ae13bfed","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2019,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1232"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : let conditionKernel := (latentArmStreamTrajectoryKernel algorithm env).map (latentArmStreamVisiblePrefixNextAction n) let fullMixed := UCB.armStreamMeasure nu ⊗ₘ conditionKernel Measure.map (fun sample : UCB.ArmRewardStream K × (History.FinitePairHistory (Fin K) Real n × Fin K) => (sample.2, UCB.armStreamCoordinate (UCB.armStreamCoordinateOfHistoryAction n sample.2) sample.1)) fullMixed = Measure.map Prod.snd fullMixed ⊗ₘ UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"latentarmstreamvisibleprefixnextaction_selectedcoordinate_mixed_eq_compprod banditrlproof.thompson.latentarmstreamvisibleprefixnextaction_selectedcoordinate_mixed_eq_compprod summing the countable pull-count/arm partition gives the selected-coordinate joint law under the mixed latent-stream/visible-condition measure. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_eq_compProd","label":"latentArmStreamVisiblePrefixNextAction_selectedCoordinate_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_eq_compProd","description":"Transport the selected-coordinate product law from the mixed representation back to the original latent-stream trajectory coupling.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-60219bd630f4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2020,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1372"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisiblePrefixNextAction_selectedCoordinate_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => (latentArmStreamVisiblePrefixNextAction n sample.2, UCB.armStreamCoordinate (UCB.armStreamCoordinateOfHistoryAction n (latentArmStreamVisiblePrefixNextAction n sample.2)) sample.1)) (latentArmStreamTrajectoryMeasure algorithm env nu) = Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => latentArmStreamVisiblePrefixNextAction n sample.2) (latentArmStreamTrajectoryMeasure algorithm env nu) ⊗ₘ UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"latentarmstreamvisibleprefixnextaction_selectedcoordinate_eq_compprod banditrlproof.thompson.latentarmstreamvisibleprefixnextaction_selectedcoordinate_eq_compprod transport the selected-coordinate product law from the mixed representation back to the original latent-stream trajectory coupling. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_joint_eq_compProd","label":"latentArmStreamVisibleNextReward_joint_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleNextReward_joint_eq_compProd","description":"Given the visible prefix through `n` and the action at `n + 1`, the actual next reward in the latent-stream coupling has exactly the selected arm law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-d0bb7188eafe","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2021,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1464"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleNextReward_joint_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => (latentArmStreamVisiblePrefixNextAction n sample.2, latentArmStreamVisibleNextReward n sample.2)) (latentArmStreamTrajectoryMeasure algorithm env nu) = Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => latentArmStreamVisiblePrefixNextAction n sample.2) (latentArmStreamTrajectoryMeasure algorithm env nu) ⊗ₘ UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"latentarmstreamvisiblenextreward_joint_eq_compprod banditrlproof.thompson.latentarmstreamvisiblenextreward_joint_eq_compprod given the visible prefix through `n` and the action at `n + 1`, the actual next reward in the latent-stream coupling has exactly the selected arm law. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_condDistrib_ae_eq_nu","label":"latentArmStreamVisibleNextReward_condDistrib_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleNextReward_condDistrib_ae_eq_nu","description":"Conditional-law form of one-step selected-reward freshness on the full latent-stream trajectory coupling.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-6a972ec63060","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2022,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1515"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleNextReward_condDistrib_ae_eq_nu {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : condDistrib (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => latentArmStreamVisibleNextReward n sample.2) (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => latentArmStreamVisiblePrefixNextAction n sample.2) (latentArmStreamTrajectoryMeasure algorithm env nu) =ᵐ[ (latentArmStreamTrajectoryMeasure algorithm env nu).map (fun sample => latentArmStreamVisiblePrefixNextAction n sample.2)] UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"latentarmstreamvisiblenextreward_conddistrib_ae_eq_nu banditrlproof.thompson.latentarmstreamvisiblenextreward_conddistrib_ae_eq_nu conditional-law form of one-step selected-reward freshness on the full latent-stream trajectory coupling. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_joint_eq_compProd","label":"latentArmStreamVisibleTrajectoryMeasure_nextReward_joint_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_joint_eq_compProd","description":"Observable-marginal form of the one-step selected-reward joint law. This removes the latent arm stream from the theorem's source space without claiming that the entire visible trajectory already equals the native construction.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-02d6c3b1f4ba","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2023,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1545"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleTrajectoryMeasure_nextReward_joint_eq_compProd {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : let visibleMeasure := (latentArmStreamTrajectoryMeasure algorithm env nu).map Prod.snd Measure.map (fun trajectory : (t : Nat) -> Fin K × Real => (latentArmStreamVisiblePrefixNextAction n trajectory, latentArmStreamVisibleNextReward n trajectory)) visibleMeasure = Measure.map (latentArmStreamVisiblePrefixNextAction n) visibleMeasure ⊗ₘ UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"latentarmstreamvisibletrajectorymeasure_nextreward_joint_eq_compprod banditrlproof.thompson.latentarmstreamvisibletrajectorymeasure_nextreward_joint_eq_compprod observable-marginal form of the one-step selected-reward joint law. this removes the latent arm stream from the theorem's source space without claiming that the entire visible trajectory already equals the native construction. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","label":"latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","description":"Conditional-law form of one-step selected-reward freshness on the visible trajectory marginal.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-7e477b470105","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2024,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1604"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : let visibleMeasure := (latentArmStreamTrajectoryMeasure algorithm env nu).map Prod.snd condDistrib (latentArmStreamVisibleNextReward n) (latentArmStreamVisiblePrefixNextAction n) visibleMeasure =ᵐ[ visibleMeasure.map (latentArmStreamVisiblePrefixNextAction n)] UCB.armStreamSelectedRewardKernel n nu","missing":[],"search":"latentarmstreamvisibletrajectorymeasure_nextreward_conddistrib_ae_eq_nu banditrlproof.thompson.latentarmstreamvisibletrajectorymeasure_nextreward_conddistrib_ae_eq_nu conditional-law form of one-step selected-reward freshness on the visible trajectory marginal. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","label":"latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","description":"Exact finite mixture law for the latent stream box and the visible SGB trajectory prefix. This is the deferred-decisions representation that the remaining native-prefix comparison must consume.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonativetrajectory/index.html#decl-53c811e737f8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","order":2025,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNativeTrajectory.lean:1629"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun sample : UCB.ArmRewardStream K × ((t : Nat) -> Fin K × Real) => (Preorder.frestrictLe n sample.1, Preorder.frestrictLe n sample.2)) (latentArmStreamTrajectoryMeasure algorithm env nu) = Measure.pi (fun _ : Finset.Iic n => Measure.infinitePi fun arm : Fin K => nu arm) ⊗ₘ latentArmStreamVisiblePrefixKernel algorithm env n","missing":[],"search":"latentarmstreamtrajectorymeasure_map_stream_visibleprefix_eq banditrlproof.thompson.latentarmstreamtrajectorymeasure_map_stream_visibleprefix_eq exact finite mixture law for the latent stream box and the visible sgb trajectory prefix. this is the deferred-decisions representation that the remaining native-prefix comparison must consume. theorem compiled","shard":"modules/fbd8bc8e834f8afc.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixGeneratedAction","label":"twoArmPrefixGeneratedAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmPrefixGeneratedAction","description":"Complete an inclusive finite prefix to an action trace, using arm `0` outside the prefix. Only coordinates through `prefix` are consumed below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-5422052ef5bd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2026,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:36"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmPrefixGeneratedAction {Env : Type v} (chron : Nat) (context : Env × History.FinitePairHistory (Fin 2) Real chron) : ActionTrace (Fin 2)","missing":[],"search":"twoarmprefixgeneratedaction banditrlproof.stochasticgradientbandit.twoarmprefixgeneratedaction complete an inclusive finite prefix to an action trace, using arm `0` outside the prefix. only coordinates through `prefix` are consumed below. definition compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount","label":"twoArmPrefixOptimalPullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount","description":"Optimal-arm pulls visible in an inclusive finite prefix.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-2cf0453aa698","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2027,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:47"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmPrefixOptimalPullCount {Env : Type v} (chron : Nat) (context : Env × History.FinitePairHistory (Fin 2) Real chron) : Nat","missing":[],"search":"twoarmprefixoptimalpullcount banditrlproof.stochasticgradientbandit.twoarmprefixoptimalpullcount optimal-arm pulls visible in an inclusive finite prefix. definition compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixGeneratedAction","label":"measurable_twoArmPrefixGeneratedAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixGeneratedAction","description":"theorem measurable_twoArmPrefixGeneratedAction {Env : Type v} [MeasurableSpace Env] (chron t : Nat) : Measurable (fun context : Env × History.FinitePairHistory (Fin 2) Real chron => twoArmPrefixGeneratedAction chron context t)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-063cb0e351b2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2028,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:52"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmPrefixGeneratedAction {Env : Type v} [MeasurableSpace Env] (chron t : Nat) : Measurable (fun context : Env × History.FinitePairHistory (Fin 2) Real chron => twoArmPrefixGeneratedAction chron context t)","missing":[],"search":"measurable_twoarmprefixgeneratedaction banditrlproof.stochasticgradientbandit.measurable_twoarmprefixgeneratedaction theorem measurable_twoarmprefixgeneratedaction {env : type v} [measurablespace env] (chron t : nat) : measurable (fun context : env × history.finitepairhistory (fin 2) real chron => twoarmprefixgeneratedaction chron context t) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixOptimalPullCount","label":"measurable_twoArmPrefixOptimalPullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixOptimalPullCount","description":"theorem measurable_twoArmPrefixOptimalPullCount {Env : Type v} [MeasurableSpace Env] (chron : Nat) : Measurable (twoArmPrefixOptimalPullCount (Env := Env) chron)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-3ba8597144e3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2029,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:66"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmPrefixOptimalPullCount {Env : Type v} [MeasurableSpace Env] (chron : Nat) : Measurable (twoArmPrefixOptimalPullCount (Env := Env) chron)","missing":[],"search":"measurable_twoarmprefixoptimalpullcount banditrlproof.stochasticgradientbandit.measurable_twoarmprefixoptimalpullcount theorem measurable_twoarmprefixoptimalpullcount {env : type v} [measurablespace env] (chron : nat) : measurable (twoarmprefixoptimalpullcount (env := env) chron) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount_environmentPrefix_eq","label":"twoArmPrefixOptimalPullCount_environmentPrefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount_environmentPrefix_eq","description":"The finite-prefix count is exactly the chronological count on the ambient trajectory.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-0de9d5c5c7fd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2030,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:78"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"@[simp] theorem twoArmPrefixOptimalPullCount_environmentPrefix_eq {Env : Type v} [MeasurableSpace Env] (chron : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmPrefixOptimalPullCount chron (twoArmEnvironmentPrefix chron sample) = twoArmOptimalPullCount (chron + 1) sample","missing":[],"search":"twoarmprefixoptimalpullcount_environmentprefix_eq banditrlproof.stochasticgradientbandit.twoarmprefixoptimalpullcount_environmentprefix_eq the finite-prefix count is exactly the chronological count on the ambient trajectory. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInclusiveOptimalPullCountProcess","label":"twoArmInclusiveOptimalPullCountProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInclusiveOptimalPullCountProcess","description":"Inclusive optimal-arm pull count as a chronological stochastic process.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-9f222be85e33","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2031,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:92"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInclusiveOptimalPullCountProcess {Env : Type v} [MeasurableSpace Env] (chron : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Nat","missing":[],"search":"twoarminclusiveoptimalpullcountprocess banditrlproof.stochasticgradientbandit.twoarminclusiveoptimalpullcountprocess inclusive optimal-arm pull count as a chronological stochastic process. definition compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmInclusiveOptimalPullCountProcess","label":"adapted_twoArmInclusiveOptimalPullCountProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.adapted_twoArmInclusiveOptimalPullCountProcess","description":"The inclusive count process is adapted to the canonical environment and generated-prefix filtration.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-020be1ac6bcc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2032,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:99"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem adapted_twoArmInclusiveOptimalPullCountProcess {Env : Type v} [MeasurableSpace Env] : Adapted (twoArmPrefixFiltration (Env := Env)) (twoArmInclusiveOptimalPullCountProcess (Env := Env))","missing":[],"search":"adapted_twoarminclusiveoptimalpullcountprocess banditrlproof.stochasticgradientbandit.adapted_twoarminclusiveoptimalpullcountprocess the inclusive count process is adapted to the canonical environment and generated-prefix filtration. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime","label":"twoArmNthOptimalPullTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime","description":"Zero-based time of the requested optimal-arm pull. The value is `top` when that pull never occurs.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-44ab7ec28f43","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2033,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:128"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmNthOptimalPullTime {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : Env × ((k : Nat) -> Fin 2 × Real) -> WithTop Nat","missing":[],"search":"twoarmnthoptimalpulltime banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime zero-based time of the requested optimal-arm pull. the value is `top` when that pull never occurs. definition compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.isStoppingTime_twoArmNthOptimalPullTime","label":"isStoppingTime_twoArmNthOptimalPullTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.isStoppingTime_twoArmNthOptimalPullTime","description":"theorem isStoppingTime_twoArmNthOptimalPullTime {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : IsStoppingTime (twoArmPrefixFiltration (Env := Env)) (twoArmNthOptimalPullTime (Env := Env) pullIndex)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-1969cbc127a4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2034,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:135"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem isStoppingTime_twoArmNthOptimalPullTime {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : IsStoppingTime (twoArmPrefixFiltration (Env := Env)) (twoArmNthOptimalPullTime (Env := Env) pullIndex)","missing":[],"search":"isstoppingtime_twoarmnthoptimalpulltime banditrlproof.stochasticgradientbandit.isstoppingtime_twoarmnthoptimalpulltime theorem isstoppingtime_twoarmnthoptimalpulltime {env : type v} [measurablespace env] (pullindex : nat) : isstoppingtime (twoarmprefixfiltration (env := env)) (twoarmnthoptimalpulltime (env := env) pullindex) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullTime","label":"measurable_twoArmNthOptimalPullTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullTime","description":"theorem measurable_twoArmNthOptimalPullTime {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : Measurable (twoArmNthOptimalPullTime (Env := Env) pullIndex)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-6df0ba677bb5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2035,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:143"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmNthOptimalPullTime {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : Measurable (twoArmNthOptimalPullTime (Env := Env) pullIndex)","missing":[],"search":"measurable_twoarmnthoptimalpulltime banditrlproof.stochasticgradientbandit.measurable_twoarmnthoptimalpulltime theorem measurable_twoarmnthoptimalpulltime {env : type v} [measurablespace env] (pullindex : nat) : measurable (twoarmnthoptimalpulltime (env := env) pullindex) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_eq_top_iff","label":"twoArmNthOptimalPullTime_eq_top_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_eq_top_iff","description":"`top` is precisely the explicit not-yet-pulled case.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-2285d15e91a0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2036,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:150"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullTime_eq_top_iff {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmNthOptimalPullTime pullIndex sample = (⊤ : WithTop Nat) <-> forall chron : Nat, twoArmOptimalPullCount (chron + 1) sample ≠ pullIndex + 1","missing":[],"search":"twoarmnthoptimalpulltime_eq_top_iff banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime_eq_top_iff `top` is precisely the explicit not-yet-pulled case. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_succ_of_nthOptimalPullTime_eq_top","label":"twoArmOptimalPullCount_lt_succ_of_nthOptimalPullTime_eq_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_succ_of_nthOptimalPullTime_eq_top","description":"A missing zero-based requested pull keeps every finite-horizon count below the corresponding positive count level.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-e8e19f9753f6","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2037,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:161"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmOptimalPullCount_lt_succ_of_nthOptimalPullTime_eq_top {Env : Type v} [MeasurableSpace Env] (pullIndex horizon : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htop : twoArmNthOptimalPullTime pullIndex sample = (⊤ : WithTop Nat)) : twoArmOptimalPullCount horizon sample < pullIndex + 1","missing":[],"search":"twoarmoptimalpullcount_lt_succ_of_nthoptimalpulltime_eq_top banditrlproof.stochasticgradientbandit.twoarmoptimalpullcount_lt_succ_of_nthoptimalpulltime_eq_top a missing zero-based requested pull keeps every finite-horizon count below the corresponding positive count level. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top","label":"twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top","description":"If one of the first `m` requested optimal-arm pulls is missing, every finite-horizon optimal-arm count is strictly below `m`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-4d9731e703fd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2038,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:179"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top {Env : Type v} [MeasurableSpace Env] (m horizon : Nat) (i : Fin m) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htop : twoArmNthOptimalPullTime (i : Nat) sample = (⊤ : WithTop Nat)) : twoArmOptimalPullCount horizon sample < m","missing":[],"search":"twoarmoptimalpullcount_lt_of_fin_nthoptimalpulltime_eq_top banditrlproof.stochasticgradientbandit.twoarmoptimalpullcount_lt_of_fin_nthoptimalpulltime_eq_top if one of the first `m` requested optimal-arm pulls is missing, every finite-horizon optimal-arm count is strictly below `m`. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq","label":"twoArmNthOptimalPullTime_count_succ_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq","description":"At every finite nth-pull time, the inclusive count hits its target.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-b1a2f3051d1d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2039,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:190"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullTime_count_succ_eq {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (hfinite : twoArmNthOptimalPullTime pullIndex sample ≠ (⊤ : WithTop Nat)) : twoArmOptimalPullCount ((twoArmNthOptimalPullTime pullIndex sample).untopA + 1) sample = pullIndex + 1","missing":[],"search":"twoarmnthoptimalpulltime_count_succ_eq banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime_count_succ_eq at every finite nth-pull time, the inclusive count hits its target. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq_of_eq","label":"twoArmNthOptimalPullTime_count_succ_eq_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq_of_eq","description":"theorem twoArmNthOptimalPullTime_count_succ_eq_of_eq {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmOptimalPullCount (t + 1) sample = pullIndex + 1","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-a82471962706","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2040,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:204"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullTime_count_succ_eq_of_eq {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmOptimalPullCount (t + 1) sample = pullIndex + 1","missing":[],"search":"twoarmnthoptimalpulltime_count_succ_eq_of_eq banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime_count_succ_eq_of_eq theorem twoarmnthoptimalpulltime_count_succ_eq_of_eq {env : type v} [measurablespace env] (pullindex t : nat) (sample : env × ((k : nat) -> fin 2 × real)) (htime : twoarmnthoptimalpulltime pullindex sample = (t : withtop nat)) : twoarmoptimalpullcount (t + 1) sample = pullindex + 1 theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_action_eq_zero","label":"twoArmNthOptimalPullTime_action_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_action_eq_zero","description":"A finite nth-pull time is a genuine selection of the optimal arm.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-6b6173129895","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2041,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:218"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullTime_action_eq_zero {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmGeneratedAction sample t = 0","missing":[],"search":"twoarmnthoptimalpulltime_action_eq_zero banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime_action_eq_zero a finite nth-pull time is a genuine selection of the optimal arm. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_eq","label":"twoArmNthOptimalPullTime_count_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_eq","description":"Immediately before the finite nth-pull time there are exactly `pullIndex` optimal pulls.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-9ace4c16ddd6","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2042,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:247"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullTime_count_eq {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmOptimalPullCount t sample = pullIndex","missing":[],"search":"twoarmnthoptimalpulltime_count_eq banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime_count_eq immediately before the finite nth-pull time there are exactly `pullindex` optimal pulls. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","label":"twoArmNthOptimalPullTime_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","description":"The complete deterministic chronological-to-pull-index bridge.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-176029d2149a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2043,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:262"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullTime_spec {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmOptimalPullCount t sample = pullIndex /\\ twoArmGeneratedAction sample t = 0 /\\ twoArmOptimalPullCount (t + 1) sample = pullIndex + 1","missing":[],"search":"twoarmnthoptimalpulltime_spec banditrlproof.stochasticgradientbandit.twoarmnthoptimalpulltime_spec the complete deterministic chronological-to-pull-index bridge. theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward","label":"twoArmNthOptimalPullReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward","description":"Reward observed at the requested optimal-arm pull. The value at `top` is Mathlib's totalized stopped-value default and is never used without a finite time witness.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-994dc618a200","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2044,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:278"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmNthOptimalPullReward {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmnthoptimalpullreward banditrlproof.stochasticgradientbandit.twoarmnthoptimalpullreward reward observed at the requested optimal-arm pull. the value at `top` is mathlib's totalized stopped-value default and is never used without a finite time witness. definition compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmGeneratedReward","label":"adapted_twoArmGeneratedReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.adapted_twoArmGeneratedReward","description":"theorem adapted_twoArmGeneratedReward {Env : Type v} [MeasurableSpace Env] : Adapted (twoArmPrefixFiltration (Env := Env)) (fun (t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) => (sample.2 t).2)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-37c35dc78a68","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2045,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:285"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem adapted_twoArmGeneratedReward {Env : Type v} [MeasurableSpace Env] : Adapted (twoArmPrefixFiltration (Env := Env)) (fun (t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) => (sample.2 t).2)","missing":[],"search":"adapted_twoarmgeneratedreward banditrlproof.stochasticgradientbandit.adapted_twoarmgeneratedreward theorem adapted_twoarmgeneratedreward {env : type v} [measurablespace env] : adapted (twoarmprefixfiltration (env := env)) (fun (t : nat) (sample : env × ((k : nat) -> fin 2 × real)) => (sample.2 t).2) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullReward","label":"measurable_twoArmNthOptimalPullReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullReward","description":"theorem measurable_twoArmNthOptimalPullReward {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : Measurable (twoArmNthOptimalPullReward (Env := Env) pullIndex)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-3000203cf37b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2046,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:308"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmNthOptimalPullReward {Env : Type v} [MeasurableSpace Env] (pullIndex : Nat) : Measurable (twoArmNthOptimalPullReward (Env := Env) pullIndex)","missing":[],"search":"measurable_twoarmnthoptimalpullreward banditrlproof.stochasticgradientbandit.measurable_twoarmnthoptimalpullreward theorem measurable_twoarmnthoptimalpullreward {env : type v} [measurablespace env] (pullindex : nat) : measurable (twoarmnthoptimalpullreward (env := env) pullindex) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_of_time_eq","label":"twoArmNthOptimalPullReward_eq_of_time_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_of_time_eq","description":"@[simp] theorem twoArmNthOptimalPullReward_eq_of_time_eq {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmNthOptimalPullReward pullIndex sample = (sample.2 t).2","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-cb2ae96c34cf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2047,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:324"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"@[simp] theorem twoArmNthOptimalPullReward_eq_of_time_eq {Env : Type v} [MeasurableSpace Env] (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmNthOptimalPullReward pullIndex sample = (sample.2 t).2","missing":[],"search":"twoarmnthoptimalpullreward_eq_of_time_eq banditrlproof.stochasticgradientbandit.twoarmnthoptimalpullreward_eq_of_time_eq @[simp] theorem twoarmnthoptimalpullreward_eq_of_time_eq {env : type v} [measurablespace env] (pullindex t : nat) (sample : env × ((k : nat) -> fin 2 × real)) (htime : twoarmnthoptimalpulltime pullindex sample = (t : withtop nat)) : twoarmnthoptimalpullreward pullindex sample = (sample.2 t).2 theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability","label":"twoArmNthOptimalPullSuccessProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability","description":"Post-pull optimal-arm probability at the requested optimal-arm pull.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-169123050a71","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2048,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:336"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmNthOptimalPullSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (pullIndex : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmnthoptimalpullsuccessprobability banditrlproof.stochasticgradientbandit.twoarmnthoptimalpullsuccessprobability post-pull optimal-arm probability at the requested optimal-arm pull. definition compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmSuccessProbability","label":"adapted_twoArmSuccessProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.adapted_twoArmSuccessProbability","description":"theorem adapted_twoArmSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) : Adapted (twoArmPrefixFiltration (Env := Env)) (twoArmSuccessProbability (Env := Env) eta)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-d6b9e6e45b1f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2049,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:343"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem adapted_twoArmSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) : Adapted (twoArmPrefixFiltration (Env := Env)) (twoArmSuccessProbability (Env := Env) eta)","missing":[],"search":"adapted_twoarmsuccessprobability banditrlproof.stochasticgradientbandit.adapted_twoarmsuccessprobability theorem adapted_twoarmsuccessprobability {env : type v} [measurablespace env] (eta : real) : adapted (twoarmprefixfiltration (env := env)) (twoarmsuccessprobability (env := env) eta) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullSuccessProbability","label":"measurable_twoArmNthOptimalPullSuccessProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullSuccessProbability","description":"theorem measurable_twoArmNthOptimalPullSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (pullIndex : Nat) : Measurable (twoArmNthOptimalPullSuccessProbability (Env := Env) eta pullIndex)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-59d3691b4974","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2050,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:368"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmNthOptimalPullSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (pullIndex : Nat) : Measurable (twoArmNthOptimalPullSuccessProbability (Env := Env) eta pullIndex)","missing":[],"search":"measurable_twoarmnthoptimalpullsuccessprobability banditrlproof.stochasticgradientbandit.measurable_twoarmnthoptimalpullsuccessprobability theorem measurable_twoarmnthoptimalpullsuccessprobability {env : type v} [measurablespace env] (eta : real) (pullindex : nat) : measurable (twoarmnthoptimalpullsuccessprobability (env := env) eta pullindex) theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_eq_of_time_eq","label":"twoArmNthOptimalPullSuccessProbability_eq_of_time_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_eq_of_time_eq","description":"@[simp] theorem twoArmNthOptimalPullSuccessProbability_eq_of_time_eq {Env : Type v} [MeasurableSpace Env] (eta : Real) (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmNthOptimalPullSuccessProbability eta pullIndex sample = twoArmSuccessProbability eta t sample","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwonthpull/index.html#decl-11f06bde8c38","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","order":2051,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoNthPull.lean:384"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"@[simp] theorem twoArmNthOptimalPullSuccessProbability_eq_of_time_eq {Env : Type v} [MeasurableSpace Env] (eta : Real) (pullIndex t : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (t : WithTop Nat)) : twoArmNthOptimalPullSuccessProbability eta pullIndex sample = twoArmSuccessProbability eta t sample","missing":[],"search":"twoarmnthoptimalpullsuccessprobability_eq_of_time_eq banditrlproof.stochasticgradientbandit.twoarmnthoptimalpullsuccessprobability_eq_of_time_eq @[simp] theorem twoarmnthoptimalpullsuccessprobability_eq_of_time_eq {env : type v} [measurablespace env] (eta : real) (pullindex t : nat) (sample : env × ((k : nat) -> fin 2 × real)) (htime : twoarmnthoptimalpulltime pullindex sample = (t : withtop nat)) : twoarmnthoptimalpullsuccessprobability eta pullindex sample = twoarmsuccessprobability eta t sample theorem compiled","shard":"modules/e4dc006c54eb66c5.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmHistoryEnvironment_ext","label":"twoArmHistoryEnvironment_ext","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmHistoryEnvironment_ext","description":"private theorem twoArmHistoryEnvironment_ext (environment₁ environment₂ : Thompson.HistoryEnvironment (Fin 2) Real) (hfeedback : environment₁.feedback = environment₂.feedback) (hinitial : environment₁.initialFeedback = environment₂.initialFeedback) : environment₁ = environment₂","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-ff1aff9e73ce","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2052,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:35"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"private theorem twoArmHistoryEnvironment_ext (environment₁ environment₂ : Thompson.HistoryEnvironment (Fin 2) Real) (hfeedback : environment₁.feedback = environment₂.feedback) (hinitial : environment₁.initialFeedback = environment₂.initialFeedback) : environment₁ = environment₂","missing":[],"search":"twoarmhistoryenvironment_ext banditrlproof.stochasticgradientbandit.twoarmhistoryenvironment_ext private theorem twoarmhistoryenvironment_ext (environment₁ environment₂ : thompson.historyenvironment (fin 2) real) (hfeedback : environment₁.feedback = environment₂.feedback) (hinitial : environment₁.initialfeedback = environment₂.initialfeedback) : environment₁ = environment₂ theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullTimeRewardBlock","label":"twoArmOptimalPullTimeRewardBlock","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullTimeRewardBlock","description":"The first `m` optimal-arm pull records on an observable trajectory. The time coordinate is part of the value, so `top` remains visible. The reward coordinate has source semantics only when that time is finite.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-bcc853c96efb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2053,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:50"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmOptimalPullTimeRewardBlock {Env : Type u} [MeasurableSpace Env] (m : Nat) : Env × ((t : Nat) -> Fin 2 × Real) -> ((i : Fin m) -> WithTop Nat × Real)","missing":[],"search":"twoarmoptimalpulltimerewardblock banditrlproof.stochasticgradientbandit.twoarmoptimalpulltimerewardblock the first `m` optimal-arm pull records on an observable trajectory. the time coordinate is part of the value, so `top` remains visible. the reward coordinate has source semantics only when that time is finite. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullTimeRewardBlock","label":"measurable_twoArmOptimalPullTimeRewardBlock","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullTimeRewardBlock","description":"theorem measurable_twoArmOptimalPullTimeRewardBlock {Env : Type u} [MeasurableSpace Env] (m : Nat) : Measurable (twoArmOptimalPullTimeRewardBlock (Env := Env) m)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-3924d27d7135","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2054,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:58"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmOptimalPullTimeRewardBlock {Env : Type u} [MeasurableSpace Env] (m : Nat) : Measurable (twoArmOptimalPullTimeRewardBlock (Env := Env) m)","missing":[],"search":"measurable_twoarmoptimalpulltimerewardblock banditrlproof.stochasticgradientbandit.measurable_twoarmoptimalpulltimerewardblock theorem measurable_twoarmoptimalpulltimerewardblock {env : type u} [measurablespace env] (m : nat) : measurable (twoarmoptimalpulltimerewardblock (env := env) m) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock","label":"twoArmLatentMaskedOptimalPullBlock","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock","description":"The latent comparison block. A finite pull reads its corresponding arm-`0` stream coordinate. A missing pull retains the stopped-value fallback rather than silently turning an absent observation into an IID reward.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-429b04c9d3d4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2055,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:69"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmLatentMaskedOptimalPullBlock (m : Nat) : UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real) -> ((i : Fin m) -> WithTop Nat × Real)","missing":[],"search":"twoarmlatentmaskedoptimalpullblock banditrlproof.stochasticgradientbandit.twoarmlatentmaskedoptimalpullblock the latent comparison block. a finite pull reads its corresponding arm-`0` stream coordinate. a missing pull retains the stopped-value fallback rather than silently turning an absent observation into an iid reward. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmLatentMaskedOptimalPullBlock","label":"measurable_twoArmLatentMaskedOptimalPullBlock","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmLatentMaskedOptimalPullBlock","description":"theorem measurable_twoArmLatentMaskedOptimalPullBlock (m : Nat) : Measurable (twoArmLatentMaskedOptimalPullBlock m)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-2af477d6355e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2056,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:84"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmLatentMaskedOptimalPullBlock (m : Nat) : Measurable (twoArmLatentMaskedOptimalPullBlock m)","missing":[],"search":"measurable_twoarmlatentmaskedoptimalpullblock banditrlproof.stochasticgradientbandit.measurable_twoarmlatentmaskedoptimalpullblock theorem measurable_twoarmlatentmaskedoptimalpullblock (m : nat) : measurable (twoarmlatentmaskedoptimalpullblock m) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullTimeRewardBlock_eq_latentMasked_ae","label":"twoArmOptimalPullTimeRewardBlock_eq_latentMasked_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullTimeRewardBlock_eq_latentMasked_ae","description":"On the latent coupling, the observable finite pull block agrees almost surely with the missing-pull-aware latent block.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-fe68e51bf2fc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2057,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:116"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmOptimalPullTimeRewardBlock_eq_latentMasked_ae (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (m : Nat) : (fun sample : UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real) => twoArmOptimalPullTimeRewardBlock (Env := Unit) m ((), sample.2)) =ᵐ[ twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta] twoArmLatentMaskedOptimalPullBlock m","missing":[],"search":"twoarmoptimalpulltimerewardblock_eq_latentmasked_ae banditrlproof.stochasticgradientbandit.twoarmoptimalpulltimerewardblock_eq_latentmasked_ae on the latent coupling, the observable finite pull block agrees almost surely with the missing-pull-aware latent block. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNativeOptimalPullTimeRewardBlock_map_eq_latentMasked","label":"twoArmNativeOptimalPullTimeRewardBlock_map_eq_latentMasked","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNativeOptimalPullTimeRewardBlock_map_eq_latentMasked","description":"Exact finite selected-block law on the native stationary fixed-IID SGB process. The right side is a masked latent-coupling law, not a product law; this retained dependence is what makes the statement valid under adaptive selection and possible missing pulls.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-533621ed24d1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2058,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:163"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNativeOptimalPullTimeRewardBlock_map_eq_latentMasked (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (m : Nat) : letI : IsMarkovKernel (UCB.finiteArmRealRewardKernel armLaw)","missing":[],"search":"twoarmnativeoptimalpulltimerewardblock_map_eq_latentmasked banditrlproof.stochasticgradientbandit.twoarmnativeoptimalpulltimerewardblock_map_eq_latentmasked exact finite selected-block law on the native stationary fixed-iid sgb process. the right side is a masked latent-coupling law, not a product law; this retained dependence is what makes the statement valid under adaptive selection and possible missing pulls. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_snd_eq_nativeStationary","label":"twoArmFixedIIDTrajectoryMeasure_map_snd_eq_nativeStationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_snd_eq_nativeStationary","description":"The source-shaped `Unit`-environment trajectory measure has the same observable marginal as the native stationary history construction.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-93dcdb7008f0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2059,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:212"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDTrajectoryMeasure_map_snd_eq_nativeStationary (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) : letI : IsMarkovKernel (UCB.finiteArmRealRewardKernel armLaw)","missing":[],"search":"twoarmfixediidtrajectorymeasure_map_snd_eq_nativestationary banditrlproof.stochasticgradientbandit.twoarmfixediidtrajectorymeasure_map_snd_eq_nativestationary the source-shaped `unit`-environment trajectory measure has the same observable marginal as the native stationary history construction. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_visible_eq_generated","label":"twoArmFixedIIDLatentTrajectoryMeasure_map_visible_eq_generated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_visible_eq_generated","description":"The latent coupling and the source-generated fixed-IID trajectory have exactly the same visible `Unit`-environment marginal. This is a full trajectory-law transport; it does not identify selected rewards as IID or condition on occurrence of any pull.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-3713c2e1224f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2060,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:276"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDLatentTrajectoryMeasure_map_visible_eq_generated (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) : Measure.map (fun sample : UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real) => ((), sample.2)) (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) = twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)","missing":[],"search":"twoarmfixediidlatenttrajectorymeasure_map_visible_eq_generated banditrlproof.stochasticgradientbandit.twoarmfixediidlatenttrajectorymeasure_map_visible_eq_generated the latent coupling and the source-generated fixed-iid trajectory have exactly the same visible `unit`-environment marginal. this is a full trajectory-law transport; it does not identify selected rewards as iid or condition on occurrence of any pull. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","label":"twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","description":"Source-facing finite selected-block transport on the actual generated two-arm trajectory measure. Missing pulls remain visible in the `WithTop` time coordinates; the theorem does not promote the masked block to IID.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-05346404836e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2061,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:335"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (m : Nat) : Measure.map (twoArmOptimalPullTimeRewardBlock (Env := Unit) m) (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) = Measure.map (twoArmLatentMaskedOptimalPullBlock m) (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta)","missing":[],"search":"twoarmfixediidtrajectorymeasure_map_optimalpulltimerewardblock_eq_latentmasked banditrlproof.stochasticgradientbandit.twoarmfixediidtrajectorymeasure_map_optimalpulltimerewardblock_eq_latentmasked source-facing finite selected-block transport on the actual generated two-arm trajectory measure. missing pulls remain visible in the `withtop` time coordinates; the theorem does not promote the masked block to iid. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPhaseOnePrefixSum","label":"twoArmAppendixCPhaseOnePrefixSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCPhaseOnePrefixSum","description":"The running reward sum inside Appendix C's recovery phase. The index `k : Fin (n1 + 1)` permits every prefix length from `0` through `n1`. The ambient reward block contains the unlucky phase of length `n0` followed by the recovery phase of length `n1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-eee0dad40e2d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2062,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:383"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCPhaseOnePrefixSum (n0 n1 : Nat) (rewardBlock : Fin (n0 + n1) -> Real) (k : Fin (n1 + 1)) : Real","missing":[],"search":"twoarmappendixcphaseoneprefixsum banditrlproof.stochasticgradientbandit.twoarmappendixcphaseoneprefixsum the running reward sum inside appendix c's recovery phase. the index `k : fin (n1 + 1)` permits every prefix length from `0` through `n1`. the ambient reward block contains the unlucky phase of length `n0` followed by the recovery phase of length `n1`. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmAppendixCPhaseOnePrefixSum","label":"measurable_twoArmAppendixCPhaseOnePrefixSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmAppendixCPhaseOnePrefixSum","description":"theorem measurable_twoArmAppendixCPhaseOnePrefixSum (n0 n1 : Nat) (k : Fin (n1 + 1)) : Measurable (fun rewardBlock : Fin (n0 + n1) -> Real => twoArmAppendixCPhaseOnePrefixSum n0 n1 rewardBlock k)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-3f45ecac5696","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2063,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:388"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmAppendixCPhaseOnePrefixSum (n0 n1 : Nat) (k : Fin (n1 + 1)) : Measurable (fun rewardBlock : Fin (n0 + n1) -> Real => twoArmAppendixCPhaseOnePrefixSum n0 n1 rewardBlock k)","missing":[],"search":"measurable_twoarmappendixcphaseoneprefixsum banditrlproof.stochasticgradientbandit.measurable_twoarmappendixcphaseoneprefixsum theorem measurable_twoarmappendixcphaseoneprefixsum (n0 n1 : nat) (k : fin (n1 + 1)) : measurable (fun rewardblock : fin (n0 + n1) -> real => twoarmappendixcphaseoneprefixsum n0 n1 rewardblock k) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseEvent","label":"twoArmAppendixCRewardPhaseEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseEvent","description":"The exact finite reward event used by the two phases in Appendix C. Phase `S0` consists of `n0` rewards equal to `-1`. Phase `S1` consists only of `{-1, 1}` rewards, has the specified exact terminal sum, and has running sum at most zero at every prefix. The later arithmetic layer will instantiate `phaseOneTotal` with the rounded Rademacher count selected by the source; this definition itself contains no probability…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-4c40b30a1d7e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2064,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:403"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCRewardPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : Set (Fin (n0 + n1) -> Real)","missing":[],"search":"twoarmappendixcrewardphaseevent banditrlproof.stochasticgradientbandit.twoarmappendixcrewardphaseevent the exact finite reward event used by the two phases in appendix c. phase `s0` consists of `n0` rewards equal to `-1`. phase `s1` consists only of `{-1, 1}` rewards, has the specified exact terminal sum, and has running sum at most zero at every prefix. the later arithmetic layer will instantiate `phaseonetotal` with the rounded rademacher count selected by the source; this definition itself contains no probability or iid premise. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCRewardPhaseEvent","label":"measurableSet_twoArmAppendixCRewardPhaseEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCRewardPhaseEvent","description":"theorem measurableSet_twoArmAppendixCRewardPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCRewardPhaseEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-4d80268b7063","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2065,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:416"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCRewardPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCRewardPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"measurableset_twoarmappendixcrewardphaseevent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixcrewardphaseevent theorem measurableset_twoarmappendixcrewardphaseevent (n0 n1 : nat) (phaseonetotal : real) : measurableset (twoarmappendixcrewardphaseevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCAllPullsPresent","label":"twoArmAppendixCAllPullsPresent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCAllPullsPresent","description":"Every requested optimal-arm pull in a finite block has occurred. This set is kept separate from the reward pattern because occurrence depends on the adaptive trajectory.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-9e3efb5b73e1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2066,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:464"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCAllPullsPresent (m : Nat) : Set ((i : Fin m) -> WithTop Nat × Real)","missing":[],"search":"twoarmappendixcallpullspresent banditrlproof.stochasticgradientbandit.twoarmappendixcallpullspresent every requested optimal-arm pull in a finite block has occurred. this set is kept separate from the reward pattern because occurrence depends on the adaptive trajectory. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCAllPullsPresent","label":"measurableSet_twoArmAppendixCAllPullsPresent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCAllPullsPresent","description":"theorem measurableSet_twoArmAppendixCAllPullsPresent (m : Nat) : MeasurableSet (twoArmAppendixCAllPullsPresent m)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-dba5c71750d4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2067,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:468"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCAllPullsPresent (m : Nat) : MeasurableSet (twoArmAppendixCAllPullsPresent m)","missing":[],"search":"measurableset_twoarmappendixcallpullspresent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixcallpullspresent theorem measurableset_twoarmappendixcallpullspresent (m : nat) : measurableset (twoarmappendixcallpullspresent m) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCObservedPhaseEvent","label":"twoArmAppendixCObservedPhaseEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCObservedPhaseEvent","description":"Observable Appendix-C phase event on a pull-time/reward block. It requires the full block to occur and only then reads the phase reward pattern.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-21e4e1ad4393","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2068,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:483"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCObservedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : Set ((i : Fin (n0 + n1)) -> WithTop Nat × Real)","missing":[],"search":"twoarmappendixcobservedphaseevent banditrlproof.stochasticgradientbandit.twoarmappendixcobservedphaseevent observable appendix-c phase event on a pull-time/reward block. it requires the full block to occur and only then reads the phase reward pattern. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCObservedPhaseEvent","label":"measurableSet_twoArmAppendixCObservedPhaseEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCObservedPhaseEvent","description":"theorem measurableSet_twoArmAppendixCObservedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCObservedPhaseEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-9b5ada0e8ea2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2069,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:490"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCObservedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCObservedPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"measurableset_twoarmappendixcobservedphaseevent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixcobservedphaseevent theorem measurableset_twoarmappendixcobservedphaseevent (n0 n1 : nat) (phaseonetotal : real) : measurableset (twoarmappendixcobservedphaseevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCLatentPhaseEvent","label":"twoArmAppendixCLatentPhaseEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCLatentPhaseEvent","description":"The latent Appendix-C event without occurrence conditioning. The first conjunct still depends on the generated visible trajectory and says that all requested pulls occur. The second conjunct reads the unconditional latent arm-`0` stream. Keeping both in the same event avoids the invalid step of declaring the reward block IID after conditioning on occurrence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-78dd6975a552","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2070,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:506"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCLatentPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : Set (UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmappendixclatentphaseevent banditrlproof.stochasticgradientbandit.twoarmappendixclatentphaseevent the latent appendix-c event without occurrence conditioning. the first conjunct still depends on the generated visible trajectory and says that all requested pulls occur. the second conjunct reads the unconditional latent arm-`0` stream. keeping both in the same event avoids the invalid step of declaring the reward block iid after conditioning on occurrence. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCLatentPhaseEvent","label":"measurableSet_twoArmAppendixCLatentPhaseEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCLatentPhaseEvent","description":"theorem measurableSet_twoArmAppendixCLatentPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-90e745092757","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2071,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:516"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCLatentPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"measurableset_twoarmappendixclatentphaseevent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixclatentphaseevent theorem measurableset_twoarmappendixclatentphaseevent (n0 n1 : nat) (phaseonetotal : real) : measurableset (twoarmappendixclatentphaseevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock_preimage_appendixCObservedPhaseEvent","label":"twoArmLatentMaskedOptimalPullBlock_preimage_appendixCObservedPhaseEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock_preimage_appendixCObservedPhaseEvent","description":"On the all-pulls-present boundary, the masked block reads exactly the latent arm-`0` prefix.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-22d2d32ba847","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2072,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:531"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmLatentMaskedOptimalPullBlock_preimage_appendixCObservedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : (twoArmLatentMaskedOptimalPullBlock (n0 + n1)) ⁻¹' twoArmAppendixCObservedPhaseEvent n0 n1 phaseOneTotal = twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal","missing":[],"search":"twoarmlatentmaskedoptimalpullblock_preimage_appendixcobservedphaseevent banditrlproof.stochasticgradientbandit.twoarmlatentmaskedoptimalpullblock_preimage_appendixcobservedphaseevent on the all-pulls-present boundary, the masked block reads exactly the latent arm-`0` prefix. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCGeneratedPhaseEvent","label":"twoArmAppendixCGeneratedPhaseEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCGeneratedPhaseEvent","description":"Source-shaped generated-process event corresponding to the finite Appendix-C pull-ordered phase.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-4e7693a9613e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2073,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:589"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCGeneratedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : Set (Unit × ((t : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmappendixcgeneratedphaseevent banditrlproof.stochasticgradientbandit.twoarmappendixcgeneratedphaseevent source-shaped generated-process event corresponding to the finite appendix-c pull-ordered phase. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCGeneratedPhaseEvent","label":"measurableSet_twoArmAppendixCGeneratedPhaseEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCGeneratedPhaseEvent","description":"theorem measurableSet_twoArmAppendixCGeneratedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCGeneratedPhaseEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-9536ce065cbf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2074,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:595"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCGeneratedPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCGeneratedPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"measurableset_twoarmappendixcgeneratedphaseevent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixcgeneratedphaseevent theorem measurableset_twoarmappendixcgeneratedphaseevent (n0 n1 : nat) (phaseonetotal : real) : measurableset (twoarmappendixcgeneratedphaseevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","label":"twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","description":"Exact transport of the finite Appendix-C phase event to the source-shaped generated SGB trajectory. The right side is an intersection of the latent reward pattern with the adaptive all-pulls-present event. This theorem does not assert a product law, selected IID, a probability lower bound, future no-return, or Theorem 2.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-f48b673b0321","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2075,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:610"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (n0 n1 : Nat) (phaseOneTotal : Real) : (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmAppendixCGeneratedPhaseEvent n0 n1 phaseOneTotal) = (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) (twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"twoarmfixediidtrajectorymeasure_appendixcgeneratedphaseevent_eq_latent banditrlproof.stochasticgradientbandit.twoarmfixediidtrajectorymeasure_appendixcgeneratedphaseevent_eq_latent exact transport of the finite appendix-c phase event to the source-shaped generated sgb trajectory. the right side is an intersection of the latent reward pattern with the adaptive all-pulls-present event. this theorem does not assert a product law, selected iid, a probability lower bound, future no-return, or theorem 2. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent","label":"twoArmAppendixCPureLatentRewardEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent","description":"The pure latent Appendix-C reward pattern, before intersecting it with the adaptive event that all requested optimal-arm pulls occur.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-4cfa896d57a5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2076,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:652"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCPureLatentRewardEvent (n0 n1 : Nat) (phaseOneTotal : Real) : Set (UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmappendixcpurelatentrewardevent banditrlproof.stochasticgradientbandit.twoarmappendixcpurelatentrewardevent the pure latent appendix-c reward pattern, before intersecting it with the adaptive event that all requested optimal-arm pulls occur. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCPureLatentRewardEvent","label":"measurableSet_twoArmAppendixCPureLatentRewardEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCPureLatentRewardEvent","description":"theorem measurableSet_twoArmAppendixCPureLatentRewardEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCPureLatentRewardEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-7483c1098c6e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2077,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:660"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCPureLatentRewardEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCPureLatentRewardEvent n0 n1 phaseOneTotal)","missing":[],"search":"measurableset_twoarmappendixcpurelatentrewardevent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixcpurelatentrewardevent theorem measurableset_twoarmappendixcpurelatentrewardevent (n0 n1 : nat) (phaseonetotal : real) : measurableset (twoarmappendixcpurelatentrewardevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent","label":"twoArmAppendixCMissingPullLatentPhaseEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent","description":"The pure latent reward pattern together with failure of at least one requested optimal-arm pull to occur. This is the complementary branch to the existing all-present latent phase event, not yet a fixed-cutoff starvation event.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-f02a9e8c45bb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2078,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:674"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmAppendixCMissingPullLatentPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : Set (UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmappendixcmissingpulllatentphaseevent banditrlproof.stochasticgradientbandit.twoarmappendixcmissingpulllatentphaseevent the pure latent reward pattern together with failure of at least one requested optimal-arm pull to occur. this is the complementary branch to the existing all-present latent phase event, not yet a fixed-cutoff starvation event. definition compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCMissingPullLatentPhaseEvent","label":"measurableSet_twoArmAppendixCMissingPullLatentPhaseEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCMissingPullLatentPhaseEvent","description":"theorem measurableSet_twoArmAppendixCMissingPullLatentPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-13047c4e98f3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2079,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:681"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmAppendixCMissingPullLatentPhaseEvent (n0 n1 : Nat) (phaseOneTotal : Real) : MeasurableSet (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"measurableset_twoarmappendixcmissingpulllatentphaseevent banditrlproof.stochasticgradientbandit.measurableset_twoarmappendixcmissingpulllatentphaseevent theorem measurableset_twoarmappendixcmissingpulllatentphaseevent (n0 n1 : nat) (phaseonetotal : real) : measurableset (twoarmappendixcmissingpulllatentphaseevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.mem_twoArmAppendixCMissingPullLatentPhaseEvent_iff","label":"mem_twoArmAppendixCMissingPullLatentPhaseEvent_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.mem_twoArmAppendixCMissingPullLatentPhaseEvent_iff","description":"Membership in the missing branch exposes an actual `WithTop.top` pull-time coordinate while retaining the pure latent reward pattern.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-f48ff16d9404","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2080,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:695"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_twoArmAppendixCMissingPullLatentPhaseEvent_iff (n0 n1 : Nat) (phaseOneTotal : Real) (sample : UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real)) : sample ∈ twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal ↔ (fun i : Fin (n0 + n1) => sample.1 (i : Nat) 0) ∈ twoArmAppendixCRewardPhaseEvent n0 n1 phaseOneTotal ∧ ∃ i : Fin (n0 + n1), twoArmNthOptimalPullTime (Env := Unit) (i : Nat) ((), sample.2) = (⊤ : WithTop Nat)","missing":[],"search":"mem_twoarmappendixcmissingpulllatentphaseevent_iff banditrlproof.stochasticgradientbandit.mem_twoarmappendixcmissingpulllatentphaseevent_iff membership in the missing branch exposes an actual `withtop.top` pull-time coordinate while retaining the pure latent reward pattern. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","label":"twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","description":"Every missing-pull Appendix-C phase lies in the finite-horizon below-threshold count event, at every chosen finite horizon. This is the deterministic missing-pull-to-starvation bridge; it asserts neither the source trigger inequality nor a probability lower bound.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-8d8bced0077d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2081,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:716"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow (n0 n1 : Nat) (phaseOneTotal : Real) (horizon : Nat) : twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal ⊆ (fun sample : UCB.ArmRewardStream 2 × ((t : Nat) -> Fin 2 × Real) => ((), sample.2)) ⁻¹' twoArmOptimalPullCountBelowEvent (Env := Unit) (n0 + n1) horizon","missing":[],"search":"twoarmappendixcmissingpulllatentphaseevent_subset_terminalcountbelow banditrlproof.stochasticgradientbandit.twoarmappendixcmissingpulllatentphaseevent_subset_terminalcountbelow every missing-pull appendix-c phase lies in the finite-horizon below-threshold count event, at every chosen finite horizon. this is the deterministic missing-pull-to-starvation bridge; it asserts neither the source trigger inequality nor a probability lower bound. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_probability_le_countBelow","label":"twoArmFixedIIDMissingPullLatentPhase_probability_le_countBelow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_probability_le_countBelow","description":"The latent missing-pull phase mass is bounded by the probability of the visible generated trajectory having fewer than the requested block of optimal-arm pulls. The inequality transports existing mass only; it does not prove that the missing branch has positive probability.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-e45e6d945121","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2082,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:737"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDMissingPullLatentPhase_probability_le_countBelow (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (n0 n1 : Nat) (phaseOneTotal : Real) (horizon : Nat) : (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta).real (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal) ≤ (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)).real (twoArmOptimalPullCountBelowEvent (Env := Unit) (n0 + n1) horizon)","missing":[],"search":"twoarmfixediidmissingpulllatentphase_probability_le_countbelow banditrlproof.stochasticgradientbandit.twoarmfixediidmissingpulllatentphase_probability_le_countbelow the latent missing-pull phase mass is bounded by the probability of the visible generated trajectory having fewer than the requested block of optimal-arm pulls. the inequality transports existing mass only; it does not prove that the missing branch has positive probability. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","label":"twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","description":"Finite-horizon expected-regret consumer for the latent missing-pull branch. Nonnegative gap times the horizon-minus-block-size charge times the existing missing-branch probability is bounded by expected sampled pseudo-regret on the actual generated trajectory. This theorem supplies no lower bound on that probability and no source trigger, selected-IID, future/no-return, ballot, asymptotic, or Theorem-2 conclusion.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-2f32eec8ea0c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2083,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:791"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta Delta : Real) (hDelta : 0 ≤ Delta) (n0 n1 : Nat) (phaseOneTotal : Real) (horizon : Nat) : Delta * ((horizon - (n0 + n1) : Nat) : Real) * (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta).real (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal) ≤ integral (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmSampledPseudoRegret (Env := Unit) Delta horizon)","missing":[],"search":"twoarmfixediidmissingpulllatentphase_charge_mul_probability_le_integral banditrlproof.stochasticgradientbandit.twoarmfixediidmissingpulllatentphase_charge_mul_probability_le_integral finite-horizon expected-regret consumer for the latent missing-pull branch. nonnegative gap times the horizon-minus-block-size charge times the existing missing-branch probability is bounded by expected sampled pseudo-regret on the actual generated trajectory. this theorem supplies no lower bound on that probability and no source trigger, selected-iid, future/no-return, ballot, asymptotic, or theorem-2 conclusion. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent_eq_union_phase_missing","label":"twoArmAppendixCPureLatentRewardEvent_eq_union_phase_missing","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent_eq_union_phase_missing","description":"The unconditional latent reward event is exactly the union of the all-present phase and the missing-pull phase.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-b4179a3ca784","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2084,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:835"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmAppendixCPureLatentRewardEvent_eq_union_phase_missing (n0 n1 : Nat) (phaseOneTotal : Real) : twoArmAppendixCPureLatentRewardEvent n0 n1 phaseOneTotal = twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal ∪ twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal","missing":[],"search":"twoarmappendixcpurelatentrewardevent_eq_union_phase_missing banditrlproof.stochasticgradientbandit.twoarmappendixcpurelatentrewardevent_eq_union_phase_missing the unconditional latent reward event is exactly the union of the all-present phase and the missing-pull phase. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.disjoint_twoArmAppendixCLatentPhaseEvent_missing","label":"disjoint_twoArmAppendixCLatentPhaseEvent_missing","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.disjoint_twoArmAppendixCLatentPhaseEvent_missing","description":"theorem disjoint_twoArmAppendixCLatentPhaseEvent_missing (n0 n1 : Nat) (phaseOneTotal : Real) : Disjoint (twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal) (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-a0d714eb4a72","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2085,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:852"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem disjoint_twoArmAppendixCLatentPhaseEvent_missing (n0 n1 : Nat) (phaseOneTotal : Real) : Disjoint (twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal) (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"disjoint_twoarmappendixclatentphaseevent_missing banditrlproof.stochasticgradientbandit.disjoint_twoarmappendixclatentphaseevent_missing theorem disjoint_twoarmappendixclatentphaseevent_missing (n0 n1 : nat) (phaseonetotal : real) : disjoint (twoarmappendixclatentphaseevent n0 n1 phaseonetotal) (twoarmappendixcmissingpulllatentphaseevent n0 n1 phaseonetotal) theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_pi","label":"twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_pi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_pi","description":"The pure latent phase probability is evaluated under the already compiled finite arm-0 product law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-fbcb80c28eb3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2086,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:864"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_pi (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (n0 n1 : Nat) (phaseOneTotal : Real) : (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) (twoArmAppendixCPureLatentRewardEvent n0 n1 phaseOneTotal) = (Measure.pi (fun _ : Fin (n0 + n1) => armLaw 0) : Measure (Fin (n0 + n1) -> Real)) (twoArmAppendixCRewardPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"twoarmfixediidlatenttrajectorymeasure_purephaseevent_eq_pi banditrlproof.stochasticgradientbandit.twoarmfixediidlatenttrajectorymeasure_purephaseevent_eq_pi the pure latent phase probability is evaluated under the already compiled finite arm-0 product law. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_phase_add_missing","label":"twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_phase_add_missing","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_phase_add_missing","description":"Probability additivity for the disjoint all-present and missing-pull branches of the pure latent phase event.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-f0e3e09364d6","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2087,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:899"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_phase_add_missing (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (n0 n1 : Nat) (phaseOneTotal : Real) : (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) (twoArmAppendixCPureLatentRewardEvent n0 n1 phaseOneTotal) = (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) (twoArmAppendixCLatentPhaseEvent n0 n1 phaseOneTotal) + (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"twoarmfixediidlatenttrajectorymeasure_purephaseevent_eq_phase_add_missing banditrlproof.stochasticgradientbandit.twoarmfixediidlatenttrajectorymeasure_purephaseevent_eq_phase_add_missing probability additivity for the disjoint all-present and missing-pull branches of the pure latent phase event. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","label":"twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","description":"Source-facing missing-pull/all-present dichotomy. The unconditional product-law probability of the pure finite reward phase is the sum of the generated all-present phase probability and the latent missing-pull branch probability. This theorem does not identify the missing branch with a fixed-cutoff starvation event, condition rewards on occurrence, or prove a positive phase bound, future no-return, ballot asymptotic…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-e051a3f15441","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2088,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:925"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta : Real) (n0 n1 : Nat) (phaseOneTotal : Real) : (Measure.pi (fun _ : Fin (n0 + n1) => armLaw 0) : Measure (Fin (n0 + n1) -> Real)) (twoArmAppendixCRewardPhaseEvent n0 n1 phaseOneTotal) = (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmAppendixCGeneratedPhaseEvent n0 n1 phaseOneTotal) + (twoArmFixedIIDLatentTrajectoryMeasure armLaw hprob eta) (twoArmAppendixCMissingPullLatentPhaseEvent n0 n1 phaseOneTotal)","missing":[],"search":"twoarmappendixcrewardphaseprobability_eq_generated_add_missing banditrlproof.stochasticgradientbandit.twoarmappendixcrewardphaseprobability_eq_generated_add_missing source-facing missing-pull/all-present dichotomy. the unconditional product-law probability of the pure finite reward phase is the sum of the generated all-present phase probability and the latent missing-pull branch probability. this theorem does not identify the missing branch with a fixed-cutoff starvation event, condition rewards on occurrence, or prove a positive phase bound, future no-return, ballot asymptotics, or theorem 2. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le","label":"softmaxProbability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le","description":"A zero-sum two-arm softmax vector reaches the source Step-1 threshold once its exact odds are at most `1 / (2 * T - 1)`. This is the final algebraic implication in Appendix C's deterministic trigger. It does not supply the preceding phase-to-parameter or phase-to-odds bound.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-087995c034b8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2089,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:969"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le (theta : Fin 2 -> Real) (horizon : Nat) (hhorizon : 1 <= horizon) (hsum : ∑ coordinate, theta coordinate = 0) (hexp : Real.exp (2 * theta 0) <= 1 / (2 * (horizon : Real) - 1)) : softmaxProbability theta 0 <= 1 / (2 * (horizon : Real))","missing":[],"search":"softmaxprobability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le banditrlproof.stochasticgradientbandit.softmaxprobability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le a zero-sum two-arm softmax vector reaches the source step-1 threshold once its exact odds are at most `1 / (2 * t - 1)`. this is the final algebraic implication in appendix c's deterministic trigger. it does not supply the preceding phase-to-parameter or phase-to-odds bound. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_one_div_two_mul_nat_of_exp_parameter_le","label":"twoArmSuccessProbability_le_one_div_two_mul_nat_of_exp_parameter_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_one_div_two_mul_nat_of_exp_parameter_le","description":"Generated-trajectory specialization of the exact two-arm odds threshold.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-d9e8295d1c60","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2090,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:1001"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSuccessProbability_le_one_div_two_mul_nat_of_exp_parameter_le {Env : Type v} [MeasurableSpace Env] (eta : Real) (time horizon : Nat) (hhorizon : 1 <= horizon) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (hexp : Real.exp (2 * twoArmTrajectoryParameterZero eta time sample) <= 1 / (2 * (horizon : Real) - 1)) : twoArmSuccessProbability eta time sample <= 1 / (2 * (horizon : Real))","missing":[],"search":"twoarmsuccessprobability_le_one_div_two_mul_nat_of_exp_parameter_le banditrlproof.stochasticgradientbandit.twoarmsuccessprobability_le_one_div_two_mul_nat_of_exp_parameter_le generated-trajectory specialization of the exact two-arm odds threshold. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_le_one_div_two_mul_nat_of_time_eq","label":"twoArmNthOptimalPullSuccessProbability_le_one_div_two_mul_nat_of_time_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_le_one_div_two_mul_nat_of_time_eq","description":"At a finite requested optimal-arm pull, the generated parameter odds cap transfers to the stopped post-pull probability used by Appendix C.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-0a96fa2ac1cb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2091,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:1022"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmNthOptimalPullSuccessProbability_le_one_div_two_mul_nat_of_time_eq {Env : Type v} [MeasurableSpace Env] (eta : Real) (pullIndex time horizon : Nat) (hhorizon : 1 <= horizon) (sample : Env × ((k : Nat) -> Fin 2 × Real)) (htime : twoArmNthOptimalPullTime pullIndex sample = (time : WithTop Nat)) (hexp : Real.exp (2 * twoArmTrajectoryParameterZero eta time sample) <= 1 / (2 * (horizon : Real) - 1)) : twoArmNthOptimalPullSuccessProbability eta pullIndex sample <= 1 / (2 * (horizon : Real))","missing":[],"search":"twoarmnthoptimalpullsuccessprobability_le_one_div_two_mul_nat_of_time_eq banditrlproof.stochasticgradientbandit.twoarmnthoptimalpullsuccessprobability_le_one_div_two_mul_nat_of_time_eq at a finite requested optimal-arm pull, the generated parameter odds cap transfers to the stopped post-pull probability used by appendix c. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_exp_two_mul_parameter","label":"twoArmSuccessProbability_le_exp_two_mul_parameter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_exp_two_mul_parameter","description":"The two-arm optimal-arm probability is bounded by the exponential of twice its zero-sum parameter coordinate. This is the deterministic softmax terminal used in Appendix C. It does not derive a parameter bound from the reward phase.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-94798e8d11fd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2092,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:1045"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSuccessProbability_le_exp_two_mul_parameter {Env : Type u} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((t : Nat) -> Fin 2 × Real)) : twoArmSuccessProbability eta n sample <= Real.exp (2 * twoArmTrajectoryParameterZero eta n sample)","missing":[],"search":"twoarmsuccessprobability_le_exp_two_mul_parameter banditrlproof.stochasticgradientbandit.twoarmsuccessprobability_le_exp_two_mul_parameter the two-arm optimal-arm probability is bounded by the exponential of twice its zero-sum parameter coordinate. this is the deterministic softmax terminal used in appendix c. it does not derive a parameter bound from the reward phase. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_one_div_two_mul_horizon_of_parameter","label":"twoArmSuccessProbability_le_one_div_two_mul_horizon_of_parameter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_one_div_two_mul_horizon_of_parameter","description":"A sufficiently negative post-prefix parameter implies the exact `1/(2*T)` optimal-arm probability threshold used by source Lemma 9. The hard phase-recurrence obligation is deliberately a producer for `hparameter`; this theorem only closes the final softmax/exponential step.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-3dff1f70299f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2093,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:1075"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSuccessProbability_le_one_div_two_mul_horizon_of_parameter {Env : Type u} [MeasurableSpace Env] (eta : Real) (n T : Nat) (sample : Env × ((t : Nat) -> Fin 2 × Real)) (hT : 0 < T) (hparameter : 2 * twoArmTrajectoryParameterZero eta n sample <= -Real.log (2 * (T : Real))) : twoArmSuccessProbability eta n sample <= 1 / (2 * (T : Real))","missing":[],"search":"twoarmsuccessprobability_le_one_div_two_mul_horizon_of_parameter banditrlproof.stochasticgradientbandit.twoarmsuccessprobability_le_one_div_two_mul_horizon_of_parameter a sufficiently negative post-prefix parameter implies the exact `1/(2*t)` optimal-arm probability threshold used by source lemma 9. the hard phase-recurrence obligation is deliberately a producer for `hparameter`; this theorem only closes the final softmax/exponential step. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCGeneratedPhaseEvent_exists_lastPullTime","label":"twoArmAppendixCGeneratedPhaseEvent_exists_lastPullTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmAppendixCGeneratedPhaseEvent_exists_lastPullTime","description":"Membership in the generated all-present Appendix-C phase exposes the finite chronological time of its last requested optimal-arm pull. The zero-based pull index is `n0+n1-1`; the conclusion records the exact before/action/after count specification at the same cutoff. This theorem is only the occurrence bridge and makes no probability or recurrence claim.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwoselectediid/index.html#decl-04a2ff78f68c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","order":2094,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoSelectedIID.lean:1101"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmAppendixCGeneratedPhaseEvent_exists_lastPullTime (n0 n1 : Nat) (phaseOneTotal : Real) (hpositive : 0 < n0 + n1) (sample : Unit × ((t : Nat) -> Fin 2 × Real)) (hphase : sample ∈ twoArmAppendixCGeneratedPhaseEvent n0 n1 phaseOneTotal) : exists cutoff : Nat, twoArmNthOptimalPullTime (n0 + n1 - 1) sample = (cutoff : WithTop Nat) /\\ twoArmOptimalPullCount cutoff sample = n0 + n1 - 1 /\\ twoArmGeneratedAction sample cutoff = 0 /\\ twoArmOptimalPullCount (cutoff + 1) sample = n0 + n1","missing":[],"search":"twoarmappendixcgeneratedphaseevent_exists_lastpulltime banditrlproof.stochasticgradientbandit.twoarmappendixcgeneratedphaseevent_exists_lastpulltime membership in the generated all-present appendix-c phase exposes the finite chronological time of its last requested optimal-arm pull. the zero-based pull index is `n0+n1-1`; the conclusion records the exact before/action/after count specification at the same cutoff. this theorem is only the occurrence bridge and makes no probability or recurrence claim. theorem compiled","shard":"modules/0bcc6ac641d11093.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedAction","label":"twoArmGeneratedAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmGeneratedAction","description":"The action coordinate of the canonical SGB trajectory.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-cc551c3494c2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2095,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:33"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmGeneratedAction {Env : Type v} (sample : Env × ((k : Nat) → Fin 2 × Real)) : ActionTrace (Fin 2)","missing":[],"search":"twoarmgeneratedaction banditrlproof.stochasticgradientbandit.twoarmgeneratedaction the action coordinate of the canonical sgb trajectory. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount","label":"twoArmOptimalPullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount","description":"Number of optimal-arm (`0`) pulls in the first `horizon` generated rounds.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-666f9fa7eb21","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2096,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:39"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmOptimalPullCount {Env : Type v} (horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) : Nat","missing":[],"search":"twoarmoptimalpullcount banditrlproof.stochasticgradientbandit.twoarmoptimalpullcount number of optimal-arm (`0`) pulls in the first `horizon` generated rounds. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneThreshold","label":"twoArmStepOneThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmStepOneThreshold","description":"The source Step-1 threshold `1 / (2*T)`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-08ec00c7c4ef","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2097,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:45"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmStepOneThreshold (horizon : Nat) : Real","missing":[],"search":"twoarmsteponethreshold banditrlproof.stochasticgradientbandit.twoarmsteponethreshold the source step-1 threshold `1 / (2*t)`. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneTriggerEvent","label":"twoArmStepOneTriggerEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmStepOneTriggerEvent","description":"At a chronological prefix ending at `prefix`, the source Step-1 trigger says that the next optimal-arm probability is at most `1/(2*T)` and that exactly `n` optimal-arm pulls have occurred through that prefix. The distinction between chronological `prefix` and pull index `n` is intentional and source-critical.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-de2217800dda","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2098,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:56"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmStepOneTriggerEvent {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) : Set (Env × ((k : Nat) → Fin 2 × Real))","missing":[],"search":"twoarmsteponetriggerevent banditrlproof.stochasticgradientbandit.twoarmsteponetriggerevent at a chronological prefix ending at `prefix`, the source step-1 trigger says that the next optimal-arm probability is at most `1/(2*t)` and that exactly `n` optimal-arm pulls have occurred through that prefix. the distinction between chronological `prefix` and pull index `n` is intentional and source-critical. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent","label":"twoArmStepOneStarvationEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent","description":"The measurable part of the Appendix-C Step-1 starvation event: a trigger prefix occurs and the total number of optimal-arm pulls by `horizon` remains exactly `n`. Equality of the two pull counts is the finite-trace statement that there is no later optimal-arm selection.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-7029af20abcc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2099,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:71"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmStepOneStarvationEvent {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) : Set (Env × ((k : Nat) → Fin 2 × Real))","missing":[],"search":"twoarmsteponestarvationevent banditrlproof.stochasticgradientbandit.twoarmsteponestarvationevent the measurable part of the appendix-c step-1 starvation event: a trigger prefix occurs and the total number of optimal-arm pulls by `horizon` remains exactly `n`. equality of the two pull counts is the finite-trace statement that there is no later optimal-arm selection. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmGeneratedAction","label":"measurable_twoArmGeneratedAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmGeneratedAction","description":"theorem measurable_twoArmGeneratedAction {Env : Type v} [MeasurableSpace Env] (t : Nat) : Measurable (fun sample : Env × ((k : Nat) → Fin 2 × Real) => twoArmGeneratedAction sample t)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-5aca557f2cd4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2100,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:79"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmGeneratedAction {Env : Type v} [MeasurableSpace Env] (t : Nat) : Measurable (fun sample : Env × ((k : Nat) → Fin 2 × Real) => twoArmGeneratedAction sample t)","missing":[],"search":"measurable_twoarmgeneratedaction banditrlproof.stochasticgradientbandit.measurable_twoarmgeneratedaction theorem measurable_twoarmgeneratedaction {env : type v} [measurablespace env] (t : nat) : measurable (fun sample : env × ((k : nat) → fin 2 × real) => twoarmgeneratedaction sample t) theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullCount","label":"measurable_twoArmOptimalPullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullCount","description":"theorem measurable_twoArmOptimalPullCount {Env : Type v} [MeasurableSpace Env] (horizon : Nat) : Measurable (twoArmOptimalPullCount (Env := Env) horizon)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-97bc34b8147e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2101,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:86"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmOptimalPullCount {Env : Type v} [MeasurableSpace Env] (horizon : Nat) : Measurable (twoArmOptimalPullCount (Env := Env) horizon)","missing":[],"search":"measurable_twoarmoptimalpullcount banditrlproof.stochasticgradientbandit.measurable_twoarmoptimalpullcount theorem measurable_twoarmoptimalpullcount {env : type v} [measurablespace env] (horizon : nat) : measurable (twoarmoptimalpullcount (env := env) horizon) theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent","label":"twoArmTerminalOptimalPullCountEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent","description":"The exact terminal optimal-arm count fiber at a finite horizon.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-d6e4bc096668","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2102,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:111"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmTerminalOptimalPullCountEvent {Env : Type v} [MeasurableSpace Env] (n horizon : Nat) : Set (Env × ((k : Nat) → Fin 2 × Real))","missing":[],"search":"twoarmterminaloptimalpullcountevent banditrlproof.stochasticgradientbandit.twoarmterminaloptimalpullcountevent the exact terminal optimal-arm count fiber at a finite horizon. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmTerminalOptimalPullCountEvent","label":"measurableSet_twoArmTerminalOptimalPullCountEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmTerminalOptimalPullCountEvent","description":"theorem measurableSet_twoArmTerminalOptimalPullCountEvent {Env : Type v} [MeasurableSpace Env] (n horizon : Nat) : MeasurableSet (twoArmTerminalOptimalPullCountEvent (Env := Env) n horizon)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-43dbaeb279d7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2103,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:116"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmTerminalOptimalPullCountEvent {Env : Type v} [MeasurableSpace Env] (n horizon : Nat) : MeasurableSet (twoArmTerminalOptimalPullCountEvent (Env := Env) n horizon)","missing":[],"search":"measurableset_twoarmterminaloptimalpullcountevent banditrlproof.stochasticgradientbandit.measurableset_twoarmterminaloptimalpullcountevent theorem measurableset_twoarmterminaloptimalpullcountevent {env : type v} [measurablespace env] (n horizon : nat) : measurableset (twoarmterminaloptimalpullcountevent (env := env) n horizon) theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent","label":"twoArmOptimalPullCountBelowEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent","description":"The finite-horizon event that fewer than `m` optimal-arm pulls occurred.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-c04c2c01ddc8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2104,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:124"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmOptimalPullCountBelowEvent {Env : Type v} [MeasurableSpace Env] (m horizon : Nat) : Set (Env × ((k : Nat) → Fin 2 × Real))","missing":[],"search":"twoarmoptimalpullcountbelowevent banditrlproof.stochasticgradientbandit.twoarmoptimalpullcountbelowevent the finite-horizon event that fewer than `m` optimal-arm pulls occurred. definition compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmOptimalPullCountBelowEvent","label":"measurableSet_twoArmOptimalPullCountBelowEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmOptimalPullCountBelowEvent","description":"theorem measurableSet_twoArmOptimalPullCountBelowEvent {Env : Type v} [MeasurableSpace Env] (m horizon : Nat) : MeasurableSet (twoArmOptimalPullCountBelowEvent (Env := Env) m horizon)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-e4a394fe13b5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2105,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:129"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmOptimalPullCountBelowEvent {Env : Type v} [MeasurableSpace Env] (m horizon : Nat) : MeasurableSet (twoArmOptimalPullCountBelowEvent (Env := Env) m horizon)","missing":[],"search":"measurableset_twoarmoptimalpullcountbelowevent banditrlproof.stochasticgradientbandit.measurableset_twoarmoptimalpullcountbelowevent theorem measurableset_twoarmoptimalpullcountbelowevent {env : type v} [measurablespace env] (m horizon : nat) : measurableset (twoarmoptimalpullcountbelowevent (env := env) m horizon) theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_eq_iUnion_terminalCount","label":"twoArmOptimalPullCountBelowEvent_eq_iUnion_terminalCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_eq_iUnion_terminalCount","description":"The below-threshold count event is the finite disjoint-by-value partition into its exact terminal-count fibers.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-571af1c2029e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2106,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:138"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmOptimalPullCountBelowEvent_eq_iUnion_terminalCount {Env : Type v} [MeasurableSpace Env] (m horizon : Nat) : twoArmOptimalPullCountBelowEvent (Env := Env) m horizon = ⋃ n : Fin m, twoArmTerminalOptimalPullCountEvent (Env := Env) (n : Nat) horizon","missing":[],"search":"twoarmoptimalpullcountbelowevent_eq_iunion_terminalcount banditrlproof.stochasticgradientbandit.twoarmoptimalpullcountbelowevent_eq_iunion_terminalcount the below-threshold count event is the finite disjoint-by-value partition into its exact terminal-count fibers. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneTriggerEvent","label":"measurableSet_twoArmStepOneTriggerEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneTriggerEvent","description":"theorem measurableSet_twoArmStepOneTriggerEvent {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) : MeasurableSet (twoArmStepOneTriggerEvent (Env := Env) eta cutoff n horizon)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-0e3379346085","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2107,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:157"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmStepOneTriggerEvent {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) : MeasurableSet (twoArmStepOneTriggerEvent (Env := Env) eta cutoff n horizon)","missing":[],"search":"measurableset_twoarmsteponetriggerevent banditrlproof.stochasticgradientbandit.measurableset_twoarmsteponetriggerevent theorem measurableset_twoarmsteponetriggerevent {env : type v} [measurablespace env] (eta : real) (cutoff n horizon : nat) : measurableset (twoarmsteponetriggerevent (env := env) eta cutoff n horizon) theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneStarvationEvent","label":"measurableSet_twoArmStepOneStarvationEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneStarvationEvent","description":"theorem measurableSet_twoArmStepOneStarvationEvent {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) : MeasurableSet (twoArmStepOneStarvationEvent (Env := Env) eta cutoff n horizon)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-6bab895edb69","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2108,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:175"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_twoArmStepOneStarvationEvent {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) : MeasurableSet (twoArmStepOneStarvationEvent (Env := Env) eta cutoff n horizon)","missing":[],"search":"measurableset_twoarmsteponestarvationevent banditrlproof.stochasticgradientbandit.measurableset_twoarmsteponestarvationevent theorem measurableset_twoarmsteponestarvationevent {env : type v} [measurablespace env] (eta : real) (cutoff n horizon : nat) : measurableset (twoarmsteponestarvationevent (env := env) eta cutoff n horizon) theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.finTwo_eq_zero_or_one","label":"finTwo_eq_zero_or_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.finTwo_eq_zero_or_one","description":"private theorem finTwo_eq_zero_or_one (action : Fin 2) : action = 0 ∨ action = 1","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-a0e30f56a4f1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2109,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:186"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"private theorem finTwo_eq_zero_or_one (action : Fin 2) : action = 0 ∨ action = 1","missing":[],"search":"fintwo_eq_zero_or_one banditrlproof.stochasticgradientbandit.fintwo_eq_zero_or_one private theorem fintwo_eq_zero_or_one (action : fin 2) : action = 0 ∨ action = 1 theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_nonneg","label":"twoArmSampledPseudoRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_nonneg","description":"theorem twoArmSampledPseudoRegret_nonneg {Env : Type v} (Delta : Real) (hDelta : 0 ≤ Delta) (horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) : 0 ≤ twoArmSampledPseudoRegret Delta horizon sample","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-7c7917d6e85a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2110,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:190"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSampledPseudoRegret_nonneg {Env : Type v} (Delta : Real) (hDelta : 0 ≤ Delta) (horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) : 0 ≤ twoArmSampledPseudoRegret Delta horizon sample","missing":[],"search":"twoarmsampledpseudoregret_nonneg banditrlproof.stochasticgradientbandit.twoarmsampledpseudoregret_nonneg theorem twoarmsampledpseudoregret_nonneg {env : type v} (delta : real) (hdelta : 0 ≤ delta) (horizon : nat) (sample : env × ((k : nat) → fin 2 × real)) : 0 ≤ twoarmsampledpseudoregret delta horizon sample theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_suboptimalPullCount","label":"twoArmSampledPseudoRegret_eq_gap_mul_suboptimalPullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_suboptimalPullCount","description":"The two-arm sampled pseudo-regret is exactly the gap times arm-`1` pulls.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-9125632b1fb5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2111,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:205"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSampledPseudoRegret_eq_gap_mul_suboptimalPullCount {Env : Type v} (Delta : Real) (horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) : twoArmSampledPseudoRegret Delta horizon sample = Delta * (pullCount (twoArmGeneratedAction sample) 1 horizon : Real)","missing":[],"search":"twoarmsampledpseudoregret_eq_gap_mul_suboptimalpullcount banditrlproof.stochasticgradientbandit.twoarmsampledpseudoregret_eq_gap_mul_suboptimalpullcount the two-arm sampled pseudo-regret is exactly the gap times arm-`1` pulls. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon","label":"twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon","description":"theorem twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon {Env : Type v} (horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) : twoArmOptimalPullCount horizon sample + pullCount (twoArmGeneratedAction sample) 1 horizon = horizon","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-8752398c56cc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2112,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:228"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon {Env : Type v} (horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) : twoArmOptimalPullCount horizon sample + pullCount (twoArmGeneratedAction sample) 1 horizon = horizon","missing":[],"search":"twoarmoptimalpullcount_add_suboptimalpullcount_eq_horizon banditrlproof.stochasticgradientbandit.twoarmoptimalpullcount_add_suboptimalpullcount_eq_horizon theorem twoarmoptimalpullcount_add_suboptimalpullcount_eq_horizon {env : type v} (horizon : nat) (sample : env × ((k : nat) → fin 2 × real)) : twoarmoptimalpullcount horizon sample + pullcount (twoarmgeneratedaction sample) 1 horizon = horizon theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_horizon_sub_of_optimalPullCount_eq","label":"twoArmSampledPseudoRegret_eq_gap_mul_horizon_sub_of_optimalPullCount_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_horizon_sub_of_optimalPullCount_eq","description":"Exactly `n` optimal pulls force the exact source starvation charge.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-4dea7c9da852","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2113,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:238"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmSampledPseudoRegret_eq_gap_mul_horizon_sub_of_optimalPullCount_eq {Env : Type v} (Delta : Real) (n horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) (hcount : twoArmOptimalPullCount horizon sample = n) : twoArmSampledPseudoRegret Delta horizon sample = Delta * ((horizon - n : Nat) : Real)","missing":[],"search":"twoarmsampledpseudoregret_eq_gap_mul_horizon_sub_of_optimalpullcount_eq banditrlproof.stochasticgradientbandit.twoarmsampledpseudoregret_eq_gap_mul_horizon_sub_of_optimalpullcount_eq exactly `n` optimal pulls force the exact source starvation charge. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent_sampledPseudoRegret_eq","label":"twoArmTerminalOptimalPullCountEvent_sampledPseudoRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent_sampledPseudoRegret_eq","description":"On one terminal-count fiber, sampled pseudo-regret has the exact charge used by the Appendix-C starvation consumer.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-0412abc6249d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2114,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:254"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTerminalOptimalPullCountEvent_sampledPseudoRegret_eq {Env : Type v} [MeasurableSpace Env] (Delta : Real) (n horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) (hcount : sample ∈ twoArmTerminalOptimalPullCountEvent (Env := Env) n horizon) : twoArmSampledPseudoRegret Delta horizon sample = Delta * ((horizon - n : Nat) : Real)","missing":[],"search":"twoarmterminaloptimalpullcountevent_sampledpseudoregret_eq banditrlproof.stochasticgradientbandit.twoarmterminaloptimalpullcountevent_sampledpseudoregret_eq on one terminal-count fiber, sampled pseudo-regret has the exact charge used by the appendix-c starvation consumer. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.mem_twoArmStepOneStarvationEvent_of_lowProbability_noFurtherOptimalPull","label":"mem_twoArmStepOneStarvationEvent_of_lowProbability_noFurtherOptimalPull","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.mem_twoArmStepOneStarvationEvent_of_lowProbability_noFurtherOptimalPull","description":"The explicit chronological no-return premise constructs membership in the measurable starvation event. This is the pathwise half of Appendix-C Step 1; it does not assign a probability to the event.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-b9758c91a09e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2115,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:271"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_twoArmStepOneStarvationEvent_of_lowProbability_noFurtherOptimalPull {Env : Type v} [MeasurableSpace Env] (eta : Real) (cutoff n horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) (hcutoff : cutoff + 1 ≤ horizon) (hlow : twoArmSuccessProbability eta cutoff sample ≤ twoArmStepOneThreshold horizon) (hcount : twoArmOptimalPullCount (cutoff + 1) sample = n) (hnoFurther : ∀ t, cutoff + 1 ≤ t → t < horizon → twoArmGeneratedAction sample t ≠ 0) : sample ∈ twoArmStepOneStarvationEvent (Env := Env) eta cutoff n horizon","missing":[],"search":"mem_twoarmsteponestarvationevent_of_lowprobability_nofurtheroptimalpull banditrlproof.stochasticgradientbandit.mem_twoarmsteponestarvationevent_of_lowprobability_nofurtheroptimalpull the explicit chronological no-return premise constructs membership in the measurable starvation event. this is the pathwise half of appendix-c step 1; it does not assign a probability to the event. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_sampledPseudoRegret_eq","label":"twoArmStepOneStarvationEvent_sampledPseudoRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_sampledPseudoRegret_eq","description":"Every path in the measurable starvation event has the exact Step-1 charge.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-59f053e697e2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2116,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:296"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmStepOneStarvationEvent_sampledPseudoRegret_eq {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (cutoff n horizon : Nat) (sample : Env × ((k : Nat) → Fin 2 × Real)) (hstarve : sample ∈ twoArmStepOneStarvationEvent (Env := Env) eta cutoff n horizon) : twoArmSampledPseudoRegret Delta horizon sample = Delta * ((horizon - n : Nat) : Real)","missing":[],"search":"twoarmsteponestarvationevent_sampledpseudoregret_eq banditrlproof.stochasticgradientbandit.twoarmsteponestarvationevent_sampledpseudoregret_eq every path in the measurable starvation event has the exact step-1 charge. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret_of_finiteMeasure","label":"integrable_twoArmSampledPseudoRegret_of_finiteMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret_of_finiteMeasure","description":"Integrability of finite-horizon sampled regret under any finite trace law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-63793395030f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2117,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:309"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmSampledPseudoRegret_of_finiteMeasure {Env : Type v} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) → Fin 2 × Real))) [IsFiniteMeasure mu] (Delta : Real) (horizon : Nat) : Integrable (twoArmSampledPseudoRegret (Env := Env) Delta horizon) mu","missing":[],"search":"integrable_twoarmsampledpseudoregret_of_finitemeasure banditrlproof.stochasticgradientbandit.integrable_twoarmsampledpseudoregret_of_finitemeasure integrability of finite-horizon sampled regret under any finite trace law. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral","label":"twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral","description":"Fewer than `m` optimal-arm pulls at `horizon` force at least the uniform `Delta * (horizon - m)` sampled-pseudo-regret charge. This is the finite-horizon consumer for a below-count event; it does not provide any positive-mass lower bound for that event. The measure need only be finite; probability and expectation terminology is reserved for probability-measure wrappers such as the fixed-IID consumer below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-40bdca13e5f3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2118,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:333"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral {Env : Type v} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) → Fin 2 × Real))) [IsFiniteMeasure mu] (Delta : Real) (hDelta : 0 ≤ Delta) (m horizon : Nat) : Delta * ((horizon - m : Nat) : Real) * mu.real (twoArmOptimalPullCountBelowEvent (Env := Env) m horizon) ≤ integral mu (twoArmSampledPseudoRegret (Env := Env) Delta horizon)","missing":[],"search":"twoarmoptimalpullcountbelowevent_charge_mul_probability_le_integral banditrlproof.stochasticgradientbandit.twoarmoptimalpullcountbelowevent_charge_mul_probability_le_integral fewer than `m` optimal-arm pulls at `horizon` force at least the uniform `delta * (horizon - m)` sampled-pseudo-regret charge. this is the finite-horizon consumer for a below-count event; it does not provide any positive-mass lower bound for that event. the measure need only be finite; probability and expectation terminology is reserved for probability-measure wrappers such as the fixed-iid consumer below. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_charge_mul_probability_le_integral","label":"twoArmStepOneStarvationEvent_charge_mul_probability_le_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_charge_mul_probability_le_integral","description":"Expectation-level deterministic Step-1 consumer. It lower-bounds expected sampled regret by the exact starvation charge times the probability of the actual generated starvation event. The missing source producer is precisely the separate lower bound on this event probability from the trigger event.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-28a3ebf1d887","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2119,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:396"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmStepOneStarvationEvent_charge_mul_probability_le_integral {Env : Type v} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) → Fin 2 × Real))) [IsFiniteMeasure mu] (eta Delta : Real) (hDelta : 0 ≤ Delta) (cutoff n horizon : Nat) : Delta * ((horizon - n : Nat) : Real) * mu.real (twoArmStepOneStarvationEvent (Env := Env) eta cutoff n horizon) ≤ integral mu (twoArmSampledPseudoRegret (Env := Env) Delta horizon)","missing":[],"search":"twoarmsteponestarvationevent_charge_mul_probability_le_integral banditrlproof.stochasticgradientbandit.twoarmsteponestarvationevent_charge_mul_probability_le_integral expectation-level deterministic step-1 consumer. it lower-bounds expected sampled regret by the exact starvation charge times the probability of the actual generated starvation event. the missing source producer is precisely the separate lower bound on this event probability from the trigger event. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","label":"twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","description":"The same deterministic Step-1 consumer specialized to the canonical generated fixed-IID trajectory with a Dirac environment prior. This wrapper covers the paper's Rademacher/Dirac arm-law instance once that explicit arm-law adapter is supplied; it still does not lower-bound the starvation-event probability.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittheoremtwostarvation/index.html#decl-6a5b1b3dbaab","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","order":2120,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTheoremTwoStarvation.lean:451"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (eta Delta : Real) (hDelta : 0 <= Delta) (cutoff n horizon : Nat) : Delta * ((horizon - n : Nat) : Real) * (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)).real (twoArmStepOneStarvationEvent (Env := Unit) eta cutoff n horizon) <= integral (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmSampledPseudoRegret (Env := Unit) Delta horizon)","missing":[],"search":"twoarmfixediidsteponestarvationevent_charge_mul_probability_le_integral banditrlproof.stochasticgradientbandit.twoarmfixediidsteponestarvationevent_charge_mul_probability_le_integral the same deterministic step-1 consumer specialized to the canonical generated fixed-iid trajectory with a dirac environment prior. this wrapper covers the paper's rademacher/dirac arm-law instance once that explicit arm-law adapter is supplied; it still does not lower-bound the starvation-event probability. theorem compiled","shard":"modules/0a3751859cc7e6fb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_softmaxProbability","label":"measurable_softmaxProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_softmaxProbability","description":"Coordinate measurability of the source softmax transform.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-b5974de47b1a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2121,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:30"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_softmaxProbability [Fintype Action] {History : Type*} [MeasurableSpace History] (theta : History -> Action -> Real) (htheta : forall action, Measurable (fun history => theta history action)) (action : Action) : Measurable (fun history => softmaxProbability (theta history) action)","missing":[],"search":"measurable_softmaxprobability banditrlproof.stochasticgradientbandit.measurable_softmaxprobability coordinate measurability of the source softmax transform. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_sourceIncrement","label":"measurable_sourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_sourceIncrement","description":"Measurability of one Algorithm-1 coordinate update.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-fd77180b2c3a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2122,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:43"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_sourceIncrement [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] {History : Type*} [MeasurableSpace History] (prob : History -> Action -> Real) (reward : History -> Real) (selected : History -> Action) (coordinate : Action) (hprob : Measurable (fun history => prob history coordinate)) (hreward : Measurable reward) (hselected : Measurable selected) : Measurable (fun history => sourceIncrement (prob history) (reward history) (selected history) coordinate)","missing":[],"search":"measurable_sourceincrement banditrlproof.stochasticgradientbandit.measurable_sourceincrement measurability of one algorithm-1 coordinate update. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter","label":"historyParameter","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyParameter","description":"The Algorithm-1 parameter vector after an inclusive finite action/reward history. At round zero the first observed pair updates `initialTheta`; every successor uses the softmax vector generated by the preceding inclusive prefix.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-fdf04bfbddb1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2123,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:66"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def historyParameter [Fintype Action] [DecidableEq Action] (initialTheta : Action -> Real) (eta : Real) : (n : Nat) -> History.FinitePairHistory Action Real n -> Action -> Real | 0, history, coordinate => initialTheta coordinate + eta * sourceIncrement (softmaxProbability initialTheta) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2 (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 coordinate | n + 1, history, coordinate => let previous := Exp3.previousPairHistory history historyParameter initialTheta eta n previous coordinate + eta * sourceIncrement (softmaxProbability (historyParameter initialTheta eta n previous)) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2 (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 coordinate @[simp] theorem historyParameter_zero [Fintype Action] [DecidableEq Action] (initialTheta : Action -> Real) (eta : Real) (history : History.FinitePairHistory Action Re…","missing":[],"search":"historyparameter banditrlproof.stochasticgradientbandit.historyparameter the algorithm-1 parameter vector after an inclusive finite action/reward history. at round zero the first observed pair updates `initialtheta`; every successor uses the softmax vector generated by the preceding inclusive prefix. definition compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_zero","label":"historyParameter_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyParameter_zero","description":"theorem historyParameter_zero [Fintype Action] [DecidableEq Action] (initialTheta : Action -> Real) (eta : Real) (history : History.FinitePairHistory Action Real 0) (coordinate : Action) : historyParameter initialTheta eta 0 history coordinate = initialTheta coordinate + eta * sourceIncrement (softmaxProbability initialTheta) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2 (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 coord…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-14a8cc20f674","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2124,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:84"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyParameter_zero [Fintype Action] [DecidableEq Action] (initialTheta : Action -> Real) (eta : Real) (history : History.FinitePairHistory Action Real 0) (coordinate : Action) : historyParameter initialTheta eta 0 history coordinate = initialTheta coordinate + eta * sourceIncrement (softmaxProbability initialTheta) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2 (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 coordinate","missing":[],"search":"historyparameter_zero banditrlproof.stochasticgradientbandit.historyparameter_zero theorem historyparameter_zero [fintype action] [decidableeq action] (initialtheta : action -> real) (eta : real) (history : history.finitepairhistory action real 0) (coordinate : action) : historyparameter initialtheta eta 0 history coordinate = initialtheta coordinate + eta * sourceincrement (softmaxprobability initialtheta) (history ⟨0, finset.mem_iic.mpr le_rfl⟩).2 (history ⟨0, finset.mem_iic.mpr le_rfl⟩).1 coordinate theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_succ","label":"historyParameter_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyParameter_succ","description":"theorem historyParameter_succ [Fintype Action] [DecidableEq Action] (initialTheta : Action -> Real) (eta : Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (coordinate : Action) : historyParameter initialTheta eta (n + 1) history coordinate = historyParameter initialTheta eta n (Exp3.previousPairHistory history) coordinate + eta * sourceIncrement (softmaxProbability (historyParameter initial…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-c47721fd60c7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2125,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:97"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyParameter_succ [Fintype Action] [DecidableEq Action] (initialTheta : Action -> Real) (eta : Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (coordinate : Action) : historyParameter initialTheta eta (n + 1) history coordinate = historyParameter initialTheta eta n (Exp3.previousPairHistory history) coordinate + eta * sourceIncrement (softmaxProbability (historyParameter initialTheta eta n (Exp3.previousPairHistory history))) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2 (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 coordinate","missing":[],"search":"historyparameter_succ banditrlproof.stochasticgradientbandit.historyparameter_succ theorem historyparameter_succ [fintype action] [decidableeq action] (initialtheta : action -> real) (eta : real) (n : nat) (history : history.finitepairhistory action real (n + 1)) (coordinate : action) : historyparameter initialtheta eta (n + 1) history coordinate = historyparameter initialtheta eta n (exp3.previouspairhistory history) coordinate + eta * sourceincrement (softmaxprobability (historyparameter initialtheta eta n (exp3.previouspairhistory history))) (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).2 (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).1 coordinate theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_historyParameter","label":"measurable_historyParameter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_historyParameter","description":"Every recursive parameter coordinate is measurable in the finite history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-92b44bdbb88b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2126,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:117"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyParameter (initialTheta : Action -> Real) (eta : Real) : forall n coordinate, Measurable (fun history : History.FinitePairHistory Action Real n => historyParameter initialTheta eta n history coordinate)","missing":[],"search":"measurable_historyparameter banditrlproof.stochasticgradientbandit.measurable_historyparameter every recursive parameter coordinate is measurable in the finite history. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxFiniteActionDistribution","label":"softmaxFiniteActionDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxFiniteActionDistribution","description":"The source softmax vector as a finite probability distribution on all actions.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-fd53781d68af","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2127,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:184"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def softmaxFiniteActionDistribution (theta : Action -> Real) : Exp3.FiniteActionDistribution (Finset.univ : Finset Action) (softmaxProbability theta) where","missing":[],"search":"softmaxfiniteactiondistribution banditrlproof.stochasticgradientbandit.softmaxfiniteactiondistribution the source softmax vector as a finite probability distribution on all actions. definition compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historySoftmaxDistributionSource","label":"historySoftmaxDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historySoftmaxDistributionSource","description":"Measurable softmax policy generated from the recursive SGB state.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-6c819cbe1f68","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2128,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:191"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def historySoftmaxDistributionSource (initialTheta : Action -> Real) (eta : Real) (n : Nat) : Exp3.MeasurableFiniteActionDistribution (Finset.univ : Finset Action) (fun history : History.FinitePairHistory Action Real n => softmaxProbability (historyParameter initialTheta eta n history)) where","missing":[],"search":"historysoftmaxdistributionsource banditrlproof.stochasticgradientbandit.historysoftmaxdistributionsource measurable softmax policy generated from the recursive sgb state. definition compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyAlgorithm","label":"historyAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyAlgorithm","description":"The recursive stochastic-gradient-bandit history policy.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-27fd3dbbe27f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2129,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:206"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def historyAlgorithm (initialTheta : Action -> Real) (eta : Real) : Thompson.HistoryAlgorithm Action Real where","missing":[],"search":"historyalgorithm banditrlproof.stochasticgradientbandit.historyalgorithm the recursive stochastic-gradient-bandit history policy. definition compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyAlgorithm_policy","label":"historyAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyAlgorithm_policy","description":"theorem historyAlgorithm_policy (initialTheta : Action -> Real) (eta : Real) (n : Nat) : (historyAlgorithm initialTheta eta).policy n = Exp3.finiteActionKernel (Finset.univ : Finset Action) (fun history : History.FinitePairHistory Action Real n => softmaxProbability (historyParameter initialTheta eta n history)) (historySoftmaxDistributionSource initialTheta eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-df636e30a200","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2130,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:226"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyAlgorithm_policy (initialTheta : Action -> Real) (eta : Real) (n : Nat) : (historyAlgorithm initialTheta eta).policy n = Exp3.finiteActionKernel (Finset.univ : Finset Action) (fun history : History.FinitePairHistory Action Real n => softmaxProbability (historyParameter initialTheta eta n history)) (historySoftmaxDistributionSource initialTheta eta n)","missing":[],"search":"historyalgorithm_policy banditrlproof.stochasticgradientbandit.historyalgorithm_policy theorem historyalgorithm_policy (initialtheta : action -> real) (eta : real) (n : nat) : (historyalgorithm initialtheta eta).policy n = exp3.finiteactionkernel (finset.univ : finset action) (fun history : history.finitepairhistory action real n => softmaxprobability (historyparameter initialtheta eta n history)) (historysoftmaxdistributionsource initialtheta eta n) theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_sourceIncrement_eq_expectedSourceIncrement","label":"integral_measurableEnvironmentInitialPairKernel_sourceIncrement_eq_expectedSourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_sourceIncrement_eq_expectedSourceIncrement","description":"Equation (5) for the generated initial action/reward pair law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-69f4cc669ee1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2131,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:236"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableEnvironmentInitialPairKernel_sourceIncrement_eq_expectedSourceIncrement {Env : Type v} [MeasurableSpace Env] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (env : Env) (mean : Action -> Real) (coordinate : Action) (hIntegrable : Integrable (fun pair : Action × Real => sourceIncrement (softmaxProbability initialTheta) pair.2 pair.1 coordinate) (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm initialTheta eta) environment env)) (hmean : forall selected, integral (environment.initialFeedback (env, selected)) id = mean selected) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm initialTheta eta) environment env) (fun pair : Action × Real => sourceIncrement (softmaxProbability initialTheta) pair.2 pair.1 coordinate) = expectedSourceIncrement (softmaxPro…","missing":[],"search":"integral_measurableenvironmentinitialpairkernel_sourceincrement_eq_expectedsourceincrement banditrlproof.stochasticgradientbandit.integral_measurableenvironmentinitialpairkernel_sourceincrement_eq_expectedsourceincrement equation (5) for the generated initial action/reward pair law. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_sourceIncrement_eq_expectedSourceIncrement","label":"integral_historyStepKernel_sourceIncrement_eq_expectedSourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_sourceIncrement_eq_expectedSourceIncrement","description":"The exact Equation-(5) conditional-kernel calculation at a fixed generated history. The reward-law hypotheses expose precisely the two facts used here: integrability of the source increment and the selected arm's reward mean.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-9c3a7af95cd9","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2132,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:311"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_historyStepKernel_sourceIncrement_eq_expectedSourceIncrement (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.HistoryEnvironment Action Real) (n : Nat) (history : History.FinitePairHistory Action Real n) (mean : Action -> Real) (coordinate : Action) (hIntegrable : Integrable (fun pair : Action × Real => sourceIncrement (softmaxProbability (historyParameter initialTheta eta n history)) pair.2 pair.1 coordinate) (Thompson.historyStepKernel (historyAlgorithm initialTheta eta) environment n history)) (hmean : forall selected, integral (environment.feedback n (history, selected)) id = mean selected) : integral (Thompson.historyStepKernel (historyAlgorithm initialTheta eta) environment n history) (fun pair : Action × Real => sourceIncrement (softmaxProbability (historyParameter initialTheta eta n history)) pair.2 pair.1 coordinate) = expectedSourceIncremen…","missing":[],"search":"integral_historystepkernel_sourceincrement_eq_expectedsourceincrement banditrlproof.stochasticgradientbandit.integral_historystepkernel_sourceincrement_eq_expectedsourceincrement the exact equation-(5) conditional-kernel calculation at a fixed generated history. the reward-law hypotheses expose precisely the two facts used here: integrability of the source increment and the selected arm's reward mean. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_expectedSourceIncrement","label":"integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_expectedSourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_expectedSourceIncrement","description":"Environment-indexed form of the same conditional-kernel calculation. This is the exact kernel that appears on the right-hand side of the generated trajectory's successor-pair `condDistrib` theorem below.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-3af8485d05f0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2133,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:396"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_expectedSourceIncrement {Env : Type v} [MeasurableSpace Env] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (mean : Action -> Real) (coordinate : Action) (hIntegrable : Integrable (fun pair : Action × Real => sourceIncrement (softmaxProbability (historyParameter initialTheta eta n history)) pair.2 pair.1 coordinate) (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) environment n (env, history))) (hmean : forall selected, integral (environment.feedback n (env, (history, selected))) id = mean selected) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) environment n (env, history)) (fun pair : Act…","missing":[],"search":"integral_measurableenvironmenthistorystepkernel_sourceincrement_eq_expectedsourceincrement banditrlproof.stochasticgradientbandit.integral_measurableenvironmenthistorystepkernel_sourceincrement_eq_expectedsourceincrement environment-indexed form of the same conditional-kernel calculation. this is the exact kernel that appears on the right-hand side of the generated trajectory's successor-pair `conddistrib` theorem below. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","label":"integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","description":"Generated Equation (5), in the source instantaneous-gap coordinates.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-70387be35574","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2134,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:433"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate {Env : Type v} [MeasurableSpace Env] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (mean gap : Action -> Real) (bestMean : Real) (coordinate : Action) (hIntegrable : Integrable (fun pair : Action × Real => sourceIncrement (softmaxProbability (historyParameter initialTheta eta n history)) pair.2 pair.1 coordinate) (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) environment n (env, history))) (hmean : forall selected, integral (environment.feedback n (env, (history, selected))) id = mean selected) (hgap : forall action, gap action = bestMean - mean action) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyA…","missing":[],"search":"integral_measurableenvironmenthistorystepkernel_sourceincrement_eq_gapcoordinate banditrlproof.stochasticgradientbandit.integral_measurableenvironmenthistorystepkernel_sourceincrement_eq_gapcoordinate generated equation (5), in the source instantaneous-gap coordinates. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryKernel","label":"trajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.trajectoryKernel","description":"Canonical generated action/reward trajectory for the recursive SGB policy.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-7f172e6089d9","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2135,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:471"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def trajectoryKernel {Env : Type v} [MeasurableSpace Env] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) : Kernel Env ((n : Nat) -> Action × Real)","missing":[],"search":"trajectorykernel banditrlproof.stochasticgradientbandit.trajectorykernel canonical generated action/reward trajectory for the recursive sgb policy. definition compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action_zero_given_environment","label":"trajectoryMeasure_condDistrib_action_zero_given_environment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action_zero_given_environment","description":"The generated initial action follows the initial softmax law given the environment.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-3ca256a23991","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2136,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:488"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_action_zero_given_environment {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Action] (prior : Measure Env) [IsFiniteMeasure prior] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 0).1) (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1) (prior ⊗ₘ trajectoryKernel initialTheta eta environment) =ᵐ[ (prior ⊗ₘ trajectoryKernel initialTheta eta environment).map (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)] Kernel.const Env (Exp3.finiteActionMeasure (Finset.univ : Finset Action) (softmaxProbability initialTheta))","missing":[],"search":"trajectorymeasure_conddistrib_action_zero_given_environment banditrlproof.stochasticgradientbandit.trajectorymeasure_conddistrib_action_zero_given_environment the generated initial action follows the initial softmax law given the environment. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action","label":"trajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action","description":"Every generated successor action has the recursive softmax conditional law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-833841ad1a1f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2137,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:509"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_action {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [StandardBorelSpace Action] (prior : Measure Env) [IsFiniteMeasure prior] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ trajectoryKernel initialTheta eta environment) =ᵐ[ (prior ⊗ₘ trajectoryKernel initialTheta eta environment).map (fun sample => Preorder.frestrictLe n sample.2)] Exp3.finiteActionKernel (Finset.univ : Finset Action) (fun history : History.FinitePairHistory Action Real n => softmaxProbability (historyParameter initialTheta eta n history)) (historySoftmaxDistributionSource initialTheta eta n)","missing":[],"search":"trajectorymeasure_conddistrib_action banditrlproof.stochasticgradientbandit.trajectorymeasure_conddistrib_action every generated successor action has the recursive softmax conditional law. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","label":"trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","description":"Conditional on the retained environment and visible prefix, the next generated action/reward pair follows the canonical SGB history-step kernel.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittrajectoryaudit/index.html#decl-e80cc945bf98","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","order":2138,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTrajectoryAudit.lean:535"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_nextPair_given_environment_prefix {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [StandardBorelSpace Action] (prior : Measure Env) [IsFiniteMeasure prior] (initialTheta : Action -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => sample.2 (n + 1)) (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2)) (prior ⊗ₘ trajectoryKernel initialTheta eta environment) =ᵐ[ (prior ⊗ₘ trajectoryKernel initialTheta eta environment).map (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2))] Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) environment n","missing":[],"search":"trajectorymeasure_conddistrib_nextpair_given_environment_prefix banditrlproof.stochasticgradientbandit.trajectorymeasure_conddistrib_nextpair_given_environment_prefix conditional on the retained environment and visible prefix, the next generated action/reward pair follows the canonical sgb history-step kernel. theorem compiled","shard":"modules/d5a41e0521dc0920.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel","label":"twoArmFixedIIDRewardKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel","description":"The fixed two-arm reward kernel, with a trivial environment coordinate added for the measurable-history-environment interface.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-75a24a86584b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2139,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:30"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def twoArmFixedIIDRewardKernel (armLaw : Fin 2 -> Measure Real) : Kernel (Unit × Fin 2) Real","missing":[],"search":"twoarmfixediidrewardkernel banditrlproof.stochasticgradientbandit.twoarmfixediidrewardkernel the fixed two-arm reward kernel, with a trivial environment coordinate added for the measurable-history-environment interface. definition compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_apply","label":"twoArmFixedIIDRewardKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_apply","description":"theorem twoArmFixedIIDRewardKernel_apply (armLaw : Fin 2 -> Measure Real) (env : Unit) (arm : Fin 2) : twoArmFixedIIDRewardKernel armLaw (env, arm) = armLaw arm","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-735a3b215d30","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2140,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:35"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDRewardKernel_apply (armLaw : Fin 2 -> Measure Real) (env : Unit) (arm : Fin 2) : twoArmFixedIIDRewardKernel armLaw (env, arm) = armLaw arm","missing":[],"search":"twoarmfixediidrewardkernel_apply banditrlproof.stochasticgradientbandit.twoarmfixediidrewardkernel_apply theorem twoarmfixediidrewardkernel_apply (armlaw : fin 2 -> measure real) (env : unit) (arm : fin 2) : twoarmfixediidrewardkernel armlaw (env, arm) = armlaw arm theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_isMarkov","label":"twoArmFixedIIDRewardKernel_isMarkov","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_isMarkov","description":"Pointwise probability laws make the fixed two-arm kernel Markov.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-fed34419f353","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2141,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:42"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDRewardKernel_isMarkov (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) : IsMarkovKernel (twoArmFixedIIDRewardKernel armLaw)","missing":[],"search":"twoarmfixediidrewardkernel_ismarkov banditrlproof.stochasticgradientbandit.twoarmfixediidrewardkernel_ismarkov pointwise probability laws make the fixed two-arm kernel markov. theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment","label":"twoArmFixedIIDEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment","description":"The fixed-IID source environment. Conditional on the selected arm, every round uses the same arm law and ignores the observed history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-858c8210001a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2142,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:55"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def twoArmFixedIIDEnvironment (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) : Thompson.MeasurableHistoryEnvironment Unit (Fin 2) Real","missing":[],"search":"twoarmfixediidenvironment banditrlproof.stochasticgradientbandit.twoarmfixediidenvironment the fixed-iid source environment. conditional on the selected arm, every round uses the same arm law and ignores the observed history. definition compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_initialFeedback_apply","label":"twoArmFixedIIDEnvironment_initialFeedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_initialFeedback_apply","description":"theorem twoArmFixedIIDEnvironment_initialFeedback_apply (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (env : Unit) (arm : Fin 2) : (twoArmFixedIIDEnvironment armLaw hprob).initialFeedback (env, arm) = armLaw arm","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-435dd15f7079","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2143,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:65"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDEnvironment_initialFeedback_apply (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (env : Unit) (arm : Fin 2) : (twoArmFixedIIDEnvironment armLaw hprob).initialFeedback (env, arm) = armLaw arm","missing":[],"search":"twoarmfixediidenvironment_initialfeedback_apply banditrlproof.stochasticgradientbandit.twoarmfixediidenvironment_initialfeedback_apply theorem twoarmfixediidenvironment_initialfeedback_apply (armlaw : fin 2 -> measure real) (hprob : forall arm, isprobabilitymeasure (armlaw arm)) (env : unit) (arm : fin 2) : (twoarmfixediidenvironment armlaw hprob).initialfeedback (env, arm) = armlaw arm theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_feedback_apply","label":"twoArmFixedIIDEnvironment_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_feedback_apply","description":"theorem twoArmFixedIIDEnvironment_feedback_apply (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (n : Nat) (env : Unit) (history : History.FinitePairHistory (Fin 2) Real n) (arm : Fin 2) : (twoArmFixedIIDEnvironment armLaw hprob).feedback n (env, (history, arm)) = armLaw arm","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-83a075e9f786","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2144,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:74"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDEnvironment_feedback_apply (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (n : Nat) (env : Unit) (history : History.FinitePairHistory (Fin 2) Real n) (arm : Fin 2) : (twoArmFixedIIDEnvironment armLaw hprob).feedback n (env, (history, arm)) = armLaw arm","missing":[],"search":"twoarmfixediidenvironment_feedback_apply banditrlproof.stochasticgradientbandit.twoarmfixediidenvironment_feedback_apply theorem twoarmfixediidenvironment_feedback_apply (armlaw : fin 2 -> measure real) (hprob : forall arm, isprobabilitymeasure (armlaw arm)) (n : nat) (env : unit) (history : history.finitepairhistory (fin 2) real n) (arm : fin 2) : (twoarmfixediidenvironment armlaw hprob).feedback n (env, (history, arm)) = armlaw arm theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDReward_aestronglyMeasurable","label":"twoArmFixedIIDReward_aestronglyMeasurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDReward_aestronglyMeasurable","description":"Real-valued rewards need no extra measurability assumption beyond their law.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-0d2a71735955","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2145,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:85"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDReward_aestronglyMeasurable (armLaw : Fin 2 -> Measure Real) (arm : Fin 2) : AEStronglyMeasurable (fun reward : Real => reward) (armLaw arm)","missing":[],"search":"twoarmfixediidreward_aestronglymeasurable banditrlproof.stochasticgradientbandit.twoarmfixediidreward_aestronglymeasurable real-valued rewards need no extra measurability assumption beyond their law. theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_contract","label":"twoArmFixedIIDEnvironment_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_contract","description":"Fixed probability laws supported on `[-1,1]` with the stated arm means produce the uniform bounded fixed-mean contract used by the compiled two-arm recurrences.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-394e236b70de","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2146,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:95"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDEnvironment_contract (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (mean : Fin 2 -> Real) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, |reward| <= 1) (hmean : forall arm, integral (armLaw arm) id = mean arm) : TwoArmBoundedFixedMeanEnvironmentContract (twoArmFixedIIDEnvironment armLaw hprob) mean","missing":[],"search":"twoarmfixediidenvironment_contract banditrlproof.stochasticgradientbandit.twoarmfixediidenvironment_contract fixed probability laws supported on `[-1,1]` with the stated arm means produce the uniform bounded fixed-mean contract used by the compiled two-arm recurrences. theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","label":"integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","description":"Equation (5) on every generated fixed-IID successor history, with the source-increment integrability premise discharged from the fixed reward-law support. This remains a one-step conditional-kernel identity; it does not perform a global tower iteration.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmfixediid/index.html#decl-d0eff58a8cba","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","order":2147,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmFixedIID.lean:117"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (mean : Fin 2 -> Real) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, |reward| <= 1) (hmean : forall arm, integral (armLaw arm) id = mean arm) (initialTheta : Fin 2 -> Real) (eta : Real) (n : Nat) (history : History.FinitePairHistory (Fin 2) Real n) (gap : Fin 2 -> Real) (bestMean : Real) (coordinate : Fin 2) (hgap : forall action, gap action = bestMean - mean action) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) (twoArmFixedIIDEnvironment armLaw hprob) n ((), history)) (fun pair : Fin 2 × Real => sourceIncrement (softmaxProbability (historyParameter initialTheta eta n history)) pair.2 pair.1 coordinate) = softmaxProbability (historyParameter initialTheta eta n history) c…","missing":[],"search":"integral_twoarmfixediidhistorystepkernel_sourceincrement_eq_gapcoordinate banditrlproof.stochasticgradientbandit.integral_twoarmfixediidhistorystepkernel_sourceincrement_eq_gapcoordinate equation (5) on every generated fixed-iid successor history, with the source-increment integrability premise discharged from the fixed reward-law support. this remains a one-step conditional-kernel identity; it does not perform a global tower iteration. theorem compiled","shard":"modules/6afc2c2832aaebe3.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zeroInitialization_finTwo","label":"softmaxProbability_zeroInitialization_finTwo","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_zeroInitialization_finTwo","description":"Zero initialization gives the exact source probability `p_1 = 1 / 2` on both arms.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarminitialrecurrence/index.html#decl-c1c30e2b1069","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","order":2148,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmInitialRecurrence.lean:48"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_zeroInitialization_finTwo (selected : Fin 2) : softmaxProbability (fun _ : Fin 2 => 0) selected = (1 : Real) / 2","missing":[],"search":"softmaxprobability_zeroinitialization_fintwo banditrlproof.stochasticgradientbandit.softmaxprobability_zeroinitialization_fintwo zero initialization gives the exact source probability `p_1 = 1 / 2` on both arms. theorem compiled","shard":"modules/50a1a42d6bbf9223.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le","label":"integral_twoArmInitialPairKernel_exp_forwardIncrement_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le","description":"Source-round `t = 1` forward exponential base recurrence. The left side is `E[exp(2 * theta_{1,2})]`, because `theta_{1,1} = 0` and the initial pair contributes exactly `eta * sourceIncrement`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarminitialrecurrence/index.html#decl-f7ed0fbe7c4a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","order":2149,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmInitialRecurrence.lean:56"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInitialPairKernel_exp_forwardIncrement_le {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.initialFeedback (env, selected), |reward| <= 1) (hmean : forall selected, integral (environment.initialFeedback (env, selected)) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) (fun pair : Fin 2 × Real => Real.exp (2 * eta * sourceIncrement (fun _ : Fin 2 => (1 : Real) / 2) pair.2 pair.1 0)) <= 1 + (eta * Delta + eta ^ 2 * sourceC eta) / 2","missing":[],"search":"integral_twoarminitialpairkernel_exp_forwardincrement_le banditrlproof.stochasticgradientbandit.integral_twoarminitialpairkernel_exp_forwardincrement_le source-round `t = 1` forward exponential base recurrence. the left side is `e[exp(2 * theta_{1,2})]`, because `theta_{1,1} = 0` and the initial pair contributes exactly `eta * sourceincrement`. theorem compiled","shard":"modules/50a1a42d6bbf9223.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le","label":"integral_twoArmInitialPairKernel_exp_inverseIncrement_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le","description":"Source-round `t = 1` inverse exponential base recurrence. This is the initial value of the inverse-odds potential used to telescope the expected squared failure mass.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarminitialrecurrence/index.html#decl-981929470068","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","order":2150,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmInitialRecurrence.lean:146"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInitialPairKernel_exp_inverseIncrement_le {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.initialFeedback (env, selected), |reward| <= 1) (hmean : forall selected, integral (environment.initialFeedback (env, selected)) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) (fun pair : Fin 2 × Real => Real.exp (-2 * eta * sourceIncrement (fun _ : Fin 2 => (1 : Real) / 2) pair.2 pair.1 0)) <= 1 - eta / 2 * (Delta - eta * sourceC eta)","missing":[],"search":"integral_twoarminitialpairkernel_exp_inverseincrement_le banditrlproof.stochasticgradientbandit.integral_twoarminitialpairkernel_exp_inverseincrement_le source-round `t = 1` inverse exponential base recurrence. this is the initial value of the inverse-odds potential used to telescope the expected squared failure mass. theorem compiled","shard":"modules/50a1a42d6bbf9223.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessorPotential","label":"twoArmForwardSuccessorPotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessorPotential","description":"The forward successor exponential at a retained finite two-arm history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-91c0cdc7fc57","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2151,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:53"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmForwardSuccessorPotential (eta : Real) {n : Nat} (history : History.FinitePairHistory (Fin 2) Real n) (pair : Fin 2 × Real) : Real","missing":[],"search":"twoarmforwardsuccessorpotential banditrlproof.stochasticgradientbandit.twoarmforwardsuccessorpotential the forward successor exponential at a retained finite two-arm history. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessorPotential","label":"twoArmInverseSuccessorPotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessorPotential","description":"The inverse-odds successor exponential at the same time fence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-4b39dbf3442a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2152,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:65"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInverseSuccessorPotential (eta : Real) {n : Nat} (history : History.FinitePairHistory (Fin 2) Real n) (pair : Fin 2 × Real) : Real","missing":[],"search":"twoarminversesuccessorpotential banditrlproof.stochasticgradientbandit.twoarminversesuccessorpotential the inverse-odds successor exponential at the same time fence. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardRecurrenceBound","label":"twoArmForwardRecurrenceBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardRecurrenceBound","description":"The additive right side of the forward conditional recurrence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-52b9961a0c7d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2153,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:77"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmForwardRecurrenceBound (eta Delta : Real) {n : Nat} (history : History.FinitePairHistory (Fin 2) Real n) : Real","missing":[],"search":"twoarmforwardrecurrencebound banditrlproof.stochasticgradientbandit.twoarmforwardrecurrencebound the additive right side of the forward conditional recurrence. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseRecurrenceBound","label":"twoArmInverseRecurrenceBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseRecurrenceBound","description":"The additive right side of the inverse conditional recurrence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-b9ce910d7c2c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2154,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:87"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInverseRecurrenceBound (eta Delta : Real) {n : Nat} (history : History.FinitePairHistory (Fin 2) Real n) : Real","missing":[],"search":"twoarminverserecurrencebound banditrlproof.stochasticgradientbandit.twoarminverserecurrencebound the additive right side of the inverse conditional recurrence. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardSuccessorPotential","label":"measurable_twoArmForwardSuccessorPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardSuccessorPotential","description":"theorem measurable_twoArmForwardSuccessorPotential (eta : Real) (n : Nat) : Measurable (fun input : History.FinitePairHistory (Fin 2) Real n × (Fin 2 × Real) => twoArmForwardSuccessorPotential eta input.1 input.2)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-2a365faf8075","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2155,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:96"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmForwardSuccessorPotential (eta : Real) (n : Nat) : Measurable (fun input : History.FinitePairHistory (Fin 2) Real n × (Fin 2 × Real) => twoArmForwardSuccessorPotential eta input.1 input.2)","missing":[],"search":"measurable_twoarmforwardsuccessorpotential banditrlproof.stochasticgradientbandit.measurable_twoarmforwardsuccessorpotential theorem measurable_twoarmforwardsuccessorpotential (eta : real) (n : nat) : measurable (fun input : history.finitepairhistory (fin 2) real n × (fin 2 × real) => twoarmforwardsuccessorpotential eta input.1 input.2) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseSuccessorPotential","label":"measurable_twoArmInverseSuccessorPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseSuccessorPotential","description":"theorem measurable_twoArmInverseSuccessorPotential (eta : Real) (n : Nat) : Measurable (fun input : History.FinitePairHistory (Fin 2) Real n × (Fin 2 × Real) => twoArmInverseSuccessorPotential eta input.1 input.2)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-5acb84589a25","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2156,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:126"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInverseSuccessorPotential (eta : Real) (n : Nat) : Measurable (fun input : History.FinitePairHistory (Fin 2) Real n × (Fin 2 × Real) => twoArmInverseSuccessorPotential eta input.1 input.2)","missing":[],"search":"measurable_twoarminversesuccessorpotential banditrlproof.stochasticgradientbandit.measurable_twoarminversesuccessorpotential theorem measurable_twoarminversesuccessorpotential (eta : real) (n : nat) : measurable (fun input : history.finitepairhistory (fin 2) real n × (fin 2 × real) => twoarminversesuccessorpotential eta input.1 input.2) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardRecurrenceBound","label":"measurable_twoArmForwardRecurrenceBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardRecurrenceBound","description":"theorem measurable_twoArmForwardRecurrenceBound (eta Delta : Real) (n : Nat) : Measurable (twoArmForwardRecurrenceBound eta Delta : History.FinitePairHistory (Fin 2) Real n -> Real)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-92eb58968a9a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2157,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:156"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmForwardRecurrenceBound (eta Delta : Real) (n : Nat) : Measurable (twoArmForwardRecurrenceBound eta Delta : History.FinitePairHistory (Fin 2) Real n -> Real)","missing":[],"search":"measurable_twoarmforwardrecurrencebound banditrlproof.stochasticgradientbandit.measurable_twoarmforwardrecurrencebound theorem measurable_twoarmforwardrecurrencebound (eta delta : real) (n : nat) : measurable (twoarmforwardrecurrencebound eta delta : history.finitepairhistory (fin 2) real n -> real) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseRecurrenceBound","label":"measurable_twoArmInverseRecurrenceBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseRecurrenceBound","description":"theorem measurable_twoArmInverseRecurrenceBound (eta Delta : Real) (n : Nat) : Measurable (twoArmInverseRecurrenceBound eta Delta : History.FinitePairHistory (Fin 2) Real n -> Real)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-073926ec5b06","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2158,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:172"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInverseRecurrenceBound (eta Delta : Real) (n : Nat) : Measurable (twoArmInverseRecurrenceBound eta Delta : History.FinitePairHistory (Fin 2) Real n -> Real)","missing":[],"search":"measurable_twoarminverserecurrencebound banditrlproof.stochasticgradientbandit.measurable_twoarminverserecurrencebound theorem measurable_twoarminverserecurrencebound (eta delta : real) (n : nat) : measurable (twoarminverserecurrencebound eta delta : history.finitepairhistory (fin 2) real n -> real) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.TwoArmBoundedFixedMeanEnvironmentContract","label":"TwoArmBoundedFixedMeanEnvironmentContract","kind":"structure","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.TwoArmBoundedFixedMeanEnvironmentContract","description":"Uniform source model for the two-arm Theorem-1 route. The same `mean` is used at the initial pair and at every history-dependent successor reward law. This is a conditional fixed-mean contract, not an independence assertion.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-687251928fbf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2159,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:194"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure TwoArmBoundedFixedMeanEnvironmentContract {Env : Type v} [MeasurableSpace Env] (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) : Prop where","missing":[],"search":"twoarmboundedfixedmeanenvironmentcontract banditrlproof.stochasticgradientbandit.twoarmboundedfixedmeanenvironmentcontract uniform source model for the two-arm theorem-1 route. the same `mean` is used at the initial pair and at every history-dependent successor reward law. this is a conditional fixed-mean contract, not an independence assertion. structure compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le","label":"integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le","description":"Forward recurrence on the jointly measurable environment/history kernel.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-c4eef73f34d0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2160,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:212"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin 2) Real n) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (env, (history, selected)), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (env, (history, selected))) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n (env, history)) (twoArmForwardSuccessorPotential eta history) <= twoArmForwardRecurrenceBound eta Delta history","missing":[],"search":"integral_measurabletwoarmhistorystepkernel_forwardsuccessor_le banditrlproof.stochasticgradientbandit.integral_measurabletwoarmhistorystepkernel_forwardsuccessor_le forward recurrence on the jointly measurable environment/history kernel. theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le","label":"integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le","description":"Inverse recurrence on the jointly measurable environment/history kernel.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-beb4d1422d1d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2161,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:244"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin 2) Real n) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (env, (history, selected)), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (env, (history, selected))) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n (env, history)) (twoArmInverseSuccessorPotential eta history) <= twoArmInverseRecurrenceBound eta Delta history","missing":[],"search":"integral_measurabletwoarmhistorystepkernel_inversesuccessor_le banditrlproof.stochasticgradientbandit.integral_measurabletwoarmhistorystepkernel_inversesuccessor_le inverse recurrence on the jointly measurable environment/history kernel. theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract","label":"integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract","description":"The uniform environment contract supplies the forward recurrence at every environment value and retained history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-a23eacf91926","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2162,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:277"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin 2) Real n) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n (env, history)) (twoArmForwardSuccessorPotential eta history) <= twoArmForwardRecurrenceBound eta Delta history","missing":[],"search":"integral_measurabletwoarmhistorystepkernel_forwardsuccessor_le_of_contract banditrlproof.stochasticgradientbandit.integral_measurabletwoarmhistorystepkernel_forwardsuccessor_le_of_contract the uniform environment contract supplies the forward recurrence at every environment value and retained history. theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le_of_contract","label":"integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le_of_contract","description":"The same contract supplies the inverse recurrence at every successor.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-8540253e8cdb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2163,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:298"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le_of_contract {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin 2) Real n) : integral (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n (env, history)) (twoArmInverseSuccessorPotential eta history) <= twoArmInverseRecurrenceBound eta Delta history","missing":[],"search":"integral_measurabletwoarmhistorystepkernel_inversesuccessor_le_of_contract banditrlproof.stochasticgradientbandit.integral_measurabletwoarmhistorystepkernel_inversesuccessor_le_of_contract the same contract supplies the inverse recurrence at every successor. theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure","label":"twoArmTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure","description":"Canonical two-arm trajectory measure under zero initialization.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-07357c81b4b3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2164,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:319"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmTrajectoryMeasure {Env : Type v} [MeasurableSpace Env] (prior : Measure Env) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) : Measure (Env × ((k : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmtrajectorymeasure banditrlproof.stochasticgradientbandit.twoarmtrajectorymeasure canonical two-arm trajectory measure under zero initialization. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmEnvironmentPrefix","label":"twoArmEnvironmentPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmEnvironmentPrefix","description":"The environment together with the visible inclusive prefix through `n`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-6e4b5faf71ba","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2165,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:336"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmEnvironmentPrefix {Env : Type v} (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Env × History.FinitePairHistory (Fin 2) Real n","missing":[],"search":"twoarmenvironmentprefix banditrlproof.stochasticgradientbandit.twoarmenvironmentprefix the environment together with the visible inclusive prefix through `n`. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNextPair","label":"twoArmNextPair","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmNextPair","description":"The next generated action/reward pair after the inclusive prefix.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-1ba92aca7028","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2166,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:343"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmNextPair {Env : Type v} (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Fin 2 × Real","missing":[],"search":"twoarmnextpair banditrlproof.stochasticgradientbandit.twoarmnextpair the next generated action/reward pair after the inclusive prefix. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmEnvironmentPrefix","label":"measurable_twoArmEnvironmentPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmEnvironmentPrefix","description":"theorem measurable_twoArmEnvironmentPrefix {Env : Type v} [MeasurableSpace Env] (n : Nat) : Measurable (twoArmEnvironmentPrefix (Env := Env) n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-af1b5c12d72c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2167,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:348"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmEnvironmentPrefix {Env : Type v} [MeasurableSpace Env] (n : Nat) : Measurable (twoArmEnvironmentPrefix (Env := Env) n)","missing":[],"search":"measurable_twoarmenvironmentprefix banditrlproof.stochasticgradientbandit.measurable_twoarmenvironmentprefix theorem measurable_twoarmenvironmentprefix {env : type v} [measurablespace env] (n : nat) : measurable (twoarmenvironmentprefix (env := env) n) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNextPair","label":"measurable_twoArmNextPair","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmNextPair","description":"theorem measurable_twoArmNextPair {Env : Type v} [MeasurableSpace Env] (n : Nat) : Measurable (twoArmNextPair (Env := Env) n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-8838f2e197bd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2168,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:354"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmNextPair {Env : Type v} [MeasurableSpace Env] (n : Nat) : Measurable (twoArmNextPair (Env := Env) n)","missing":[],"search":"measurable_twoarmnextpair banditrlproof.stochasticgradientbandit.measurable_twoarmnextpair theorem measurable_twoarmnextpair {env : type v} [measurablespace env] (n : nat) : measurable (twoarmnextpair (env := env) n) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma","label":"twoArmPrefixSigma","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma","description":"The sigma-algebra generated by retaining the environment and the inclusive trace prefix through `n`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-24a42ede78b8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2169,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:361"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"@[reducible] def twoArmPrefixSigma {Env : Type v} [MeasurableSpace Env] (n : Nat) : MeasurableSpace (Env × ((k : Nat) -> Fin 2 × Real))","missing":[],"search":"twoarmprefixsigma banditrlproof.stochasticgradientbandit.twoarmprefixsigma the sigma-algebra generated by retaining the environment and the inclusive trace prefix through `n`. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma_mono","label":"twoArmPrefixSigma_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma_mono","description":"theorem twoArmPrefixSigma_mono {Env : Type v} [MeasurableSpace Env] {n m : Nat} (hnm : n <= m) : twoArmPrefixSigma (Env := Env) n <= twoArmPrefixSigma (Env := Env) m","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-276e2e8f6fac","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2170,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:368"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmPrefixSigma_mono {Env : Type v} [MeasurableSpace Env] {n m : Nat} (hnm : n <= m) : twoArmPrefixSigma (Env := Env) n <= twoArmPrefixSigma (Env := Env) m","missing":[],"search":"twoarmprefixsigma_mono banditrlproof.stochasticgradientbandit.twoarmprefixsigma_mono theorem twoarmprefixsigma_mono {env : type v} [measurablespace env] {n m : nat} (hnm : n <= m) : twoarmprefixsigma (env := env) n <= twoarmprefixsigma (env := env) m theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixFiltration","label":"twoArmPrefixFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmPrefixFiltration","description":"The retained environment/prefix sigma-algebras form the canonical discrete-time filtration needed by the later tower argument.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-cd3ee89dad2a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2171,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:396"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmPrefixFiltration {Env : Type v} [MeasurableSpace Env] : Filtration Nat (inferInstance : MeasurableSpace (Env × ((k : Nat) -> Fin 2 × Real))) where","missing":[],"search":"twoarmprefixfiltration banditrlproof.stochasticgradientbandit.twoarmprefixfiltration the retained environment/prefix sigma-algebras form the canonical discrete-time filtration needed by the later tower argument. definition compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardTrajectorySuccessorPotential","label":"measurable_twoArmForwardTrajectorySuccessorPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardTrajectorySuccessorPotential","description":"theorem measurable_twoArmForwardTrajectorySuccessorPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample))","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-6f162a7874fd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2172,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:405"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmForwardTrajectorySuccessorPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample))","missing":[],"search":"measurable_twoarmforwardtrajectorysuccessorpotential banditrlproof.stochasticgradientbandit.measurable_twoarmforwardtrajectorysuccessorpotential theorem measurable_twoarmforwardtrajectorysuccessorpotential {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (fun sample : env × ((k : nat) -> fin 2 × real) => twoarmforwardsuccessorpotential eta (twoarmenvironmentprefix n sample).2 (twoarmnextpair n sample)) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseTrajectorySuccessorPotential","label":"measurable_twoArmInverseTrajectorySuccessorPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseTrajectorySuccessorPotential","description":"theorem measurable_twoArmInverseTrajectorySuccessorPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample))","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-90878248af35","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2173,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:416"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInverseTrajectorySuccessorPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample))","missing":[],"search":"measurable_twoarminversetrajectorysuccessorpotential banditrlproof.stochasticgradientbandit.measurable_twoarminversetrajectorysuccessorpotential theorem measurable_twoarminversetrajectorysuccessorpotential {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (fun sample : env × ((k : nat) -> fin 2 × real) => twoarminversesuccessorpotential eta (twoarmenvironmentprefix n sample).2 (twoarmnextpair n sample)) theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","label":"trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","description":"On the actual canonical trajectory, the conditional-distribution integral of the forward successor potential obeys the source recurrence for almost every retained environment/prefix.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-c0c122057c1e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2174,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:430"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem trajectoryPrefix_condDistrib_integral_forwardSuccessor_le {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : ∀ᵐ context ∂(twoArmTrajectoryMeasure prior eta environment).map (twoArmEnvironmentPrefix n), integral (condDistrib (twoArmNextPair n) (twoArmEnvironmentPrefix n) (twoArmTrajectoryMeasure prior eta environment) context) (twoArmForwardSuccessorPotential eta context.2) <= twoArmForwardRecurrenceBound eta Delta context.2","missing":[],"search":"trajectoryprefix_conddistrib_integral_forwardsuccessor_le banditrlproof.stochasticgradientbandit.trajectoryprefix_conddistrib_integral_forwardsuccessor_le on the actual canonical trajectory, the conditional-distribution integral of the forward successor potential obeys the source recurrence for almost every retained environment/prefix. theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_inverseSuccessor_le","label":"trajectoryPrefix_condDistrib_integral_inverseSuccessor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_inverseSuccessor_le","description":"The inverse-potential recurrence transported to the same canonical trajectory conditional distribution.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmmeasurablerecurrence/index.html#decl-db42252f4752","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","order":2175,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmMeasurableRecurrence.lean:466"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem trajectoryPrefix_condDistrib_integral_inverseSuccessor_le {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : ∀ᵐ context ∂(twoArmTrajectoryMeasure prior eta environment).map (twoArmEnvironmentPrefix n), integral (condDistrib (twoArmNextPair n) (twoArmEnvironmentPrefix n) (twoArmTrajectoryMeasure prior eta environment) context) (twoArmInverseSuccessorPotential eta context.2) <= twoArmInverseRecurrenceBound eta Delta context.2","missing":[],"search":"trajectoryprefix_conddistrib_integral_inversesuccessor_le banditrlproof.stochasticgradientbandit.trajectoryprefix_conddistrib_integral_inversesuccessor_le the inverse-potential recurrence transported to the same canonical trajectory conditional distribution. theorem compiled","shard":"modules/f4d40b144252fac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le_of_contract","label":"integral_twoArmInitialPairKernel_exp_forwardIncrement_le_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le_of_contract","description":"The uniform contract supplies the source-round forward base recurrence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-b25eed3ca5b1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2176,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:49"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInitialPairKernel_exp_forwardIncrement_le_of_contract {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (env : Env) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) (fun pair : Fin 2 × Real => Real.exp (2 * eta * sourceIncrement (fun _ : Fin 2 => (1 : Real) / 2) pair.2 pair.1 0)) <= 1 + (eta * Delta + eta ^ 2 * sourceC eta) / 2","missing":[],"search":"integral_twoarminitialpairkernel_exp_forwardincrement_le_of_contract banditrlproof.stochasticgradientbandit.integral_twoarminitialpairkernel_exp_forwardincrement_le_of_contract the uniform contract supplies the source-round forward base recurrence. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le_of_contract","label":"integral_twoArmInitialPairKernel_exp_inverseIncrement_le_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le_of_contract","description":"The same contract supplies the source-round inverse base recurrence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-b23232daf047","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2177,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:71"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInitialPairKernel_exp_inverseIncrement_le_of_contract {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (env : Env) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) (fun pair : Fin 2 × Real => Real.exp (-2 * eta * sourceIncrement (fun _ : Fin 2 => (1 : Real) / 2) pair.2 pair.1 0)) <= 1 - eta / 2 * (Delta - eta * sourceC eta)","missing":[],"search":"integral_twoarminitialpairkernel_exp_inverseincrement_le_of_contract banditrlproof.stochasticgradientbandit.integral_twoarminitialpairkernel_exp_inverseincrement_le_of_contract the same contract supplies the source-round inverse base recurrence. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_zero_abs_le_one_ae","label":"twoArmTrajectoryMeasure_reward_zero_abs_le_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_zero_abs_le_one_ae","description":"The initial reward bound holds on the full prior-mixed canonical path.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-b139fd008706","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2178,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:95"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectoryMeasure_reward_zero_abs_le_one_ae {Env : Type v} [MeasurableSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) : ∀ᵐ sample ∂twoArmTrajectoryMeasure prior eta environment, |(sample.2 0).2| <= 1","missing":[],"search":"twoarmtrajectorymeasure_reward_zero_abs_le_one_ae banditrlproof.stochasticgradientbandit.twoarmtrajectorymeasure_reward_zero_abs_le_one_ae the initial reward bound holds on the full prior-mixed canonical path. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_succ_abs_le_one_ae","label":"twoArmTrajectoryMeasure_reward_succ_abs_le_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_succ_abs_le_one_ae","description":"Every fixed successor reward bound also holds on the full canonical path.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-753f62a7b0b8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2179,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:142"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectoryMeasure_reward_succ_abs_le_one_ae {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : ∀ᵐ sample ∂twoArmTrajectoryMeasure prior eta environment, |(sample.2 (n + 1)).2| <= 1","missing":[],"search":"twoarmtrajectorymeasure_reward_succ_abs_le_one_ae banditrlproof.stochasticgradientbandit.twoarmtrajectorymeasure_reward_succ_abs_le_one_ae every fixed successor reward bound also holds on the full canonical path. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_abs_le_one_ae","label":"twoArmTrajectoryMeasure_reward_abs_le_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_abs_le_one_ae","description":"Uniform fixed-coordinate wrapper, separating coordinate zero from shifted successor coordinates.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-04780a200a2a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2180,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:197"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectoryMeasure_reward_abs_le_one_ae {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (t : Nat) : ∀ᵐ sample ∂twoArmTrajectoryMeasure prior eta environment, |(sample.2 t).2| <= 1","missing":[],"search":"twoarmtrajectorymeasure_reward_abs_le_one_ae banditrlproof.stochasticgradientbandit.twoarmtrajectorymeasure_reward_abs_le_one_ae uniform fixed-coordinate wrapper, separating coordinate zero from shifted successor coordinates. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_prefix_rewards_abs_le_one_ae","label":"twoArmTrajectoryMeasure_prefix_rewards_abs_le_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_prefix_rewards_abs_le_one_ae","description":"For a fixed finite prefix, all reward coordinates are simultaneously in the source support almost everywhere.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-7e3dd3b140fe","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2181,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:217"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectoryMeasure_prefix_rewards_abs_le_one_ae {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : ∀ᵐ sample ∂twoArmTrajectoryMeasure prior eta environment, ∀ i : Finset.Iic n, |(sample.2 i.1).2| <= 1","missing":[],"search":"twoarmtrajectorymeasure_prefix_rewards_abs_le_one_ae banditrlproof.stochasticgradientbandit.twoarmtrajectorymeasure_prefix_rewards_abs_le_one_ae for a fixed finite prefix, all reward coordinates are simultaneously in the source support almost everywhere. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_le_abs_reward_of_mem_Icc","label":"abs_sourceIncrement_le_abs_reward_of_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_le_abs_reward_of_mem_Icc","description":"One Algorithm-1 coordinate update is no larger than the observed reward when the coordinate probability lies in `[0,1]`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-47fd2fed8cb2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2182,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:236"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_sourceIncrement_le_abs_reward_of_mem_Icc {Action : Type*} [DecidableEq Action] (p : Action -> Real) (reward : Real) (selected coordinate : Action) (hp_nonneg : 0 <= p coordinate) (hp_le_one : p coordinate <= 1) : |sourceIncrement p reward selected coordinate| <= |reward|","missing":[],"search":"abs_sourceincrement_le_abs_reward_of_mem_icc banditrlproof.stochasticgradientbandit.abs_sourceincrement_le_abs_reward_of_mem_icc one algorithm-1 coordinate update is no larger than the observed reward when the coordinate probability lies in `[0,1]`. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_softmax_le_abs_reward","label":"abs_sourceIncrement_softmax_le_abs_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_softmax_le_abs_reward","description":"Softmax coordinates satisfy the probability premises of the preceding generic update bound.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-ada6999c14a8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2183,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:252"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_sourceIncrement_softmax_le_abs_reward {Action : Type*} [Fintype Action] [DecidableEq Action] [Nonempty Action] (theta : Action -> Real) (reward : Real) (selected coordinate : Action) : |sourceIncrement (softmaxProbability theta) reward selected coordinate| <= |reward|","missing":[],"search":"abs_sourceincrement_softmax_le_abs_reward banditrlproof.stochasticgradientbandit.abs_sourceincrement_softmax_le_abs_reward softmax coordinates satisfy the probability premises of the preceding generic update bound. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmInitialPairKernel_sourceIncrement_of_contract","label":"integrable_measurableTwoArmInitialPairKernel_sourceIncrement_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmInitialPairKernel_sourceIncrement_of_contract","description":"The bounded fixed-mean contract discharges the initial-pair `Integrable sourceIncrement` premise used by Equation (5). Only the support field of the contract is needed: softmax updates are measurable and have absolute value at most the observed reward.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-a5eb5a55fe93","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2184,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:268"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_measurableTwoArmInitialPairKernel_sourceIncrement_of_contract {Env : Type v} [MeasurableSpace Env] (initialTheta : Fin 2 -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (env : Env) (coordinate : Fin 2) : Integrable (fun pair : Fin 2 × Real => sourceIncrement (softmaxProbability initialTheta) pair.2 pair.1 coordinate) (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm initialTheta eta) environment env)","missing":[],"search":"integrable_measurabletwoarminitialpairkernel_sourceincrement_of_contract banditrlproof.stochasticgradientbandit.integrable_measurabletwoarminitialpairkernel_sourceincrement_of_contract the bounded fixed-mean contract discharges the initial-pair `integrable sourceincrement` premise used by equation (5). only the support field of the contract is needed: softmax updates are measurable and have absolute value at most the observed reward. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract","label":"integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract","description":"The same contract discharges Equation-(5)'s source-increment integrability premise at every generated successor history. No independence, gap, or learning-rate assumption is introduced.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-a963c4695acf","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2185,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:312"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract {Env : Type v} [MeasurableSpace Env] (initialTheta : Fin 2 -> Real) (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin 2) Real n) (coordinate : Fin 2) : Integrable (fun pair : Fin 2 × Real => sourceIncrement (softmaxProbability (historyParameter initialTheta eta n history)) pair.2 pair.1 coordinate) (Thompson.measurableEnvironmentHistoryStepKernel (historyAlgorithm initialTheta eta) environment n (env, history))","missing":[],"search":"integrable_measurabletwoarmhistorystepkernel_sourceincrement_of_contract banditrlproof.stochasticgradientbandit.integrable_measurabletwoarmhistorystepkernel_sourceincrement_of_contract the same contract discharges equation-(5)'s source-increment integrability premise at every generated successor history. no independence, gap, or learning-rate assumption is introduced. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.abs_historyParameter_zeroInitialization_le","label":"abs_historyParameter_zeroInitialization_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.abs_historyParameter_zeroInitialization_le","description":"On a finite history supported in `[-1,1]`, every zero-initialized Algorithm-1 parameter coordinate grows by at most `|eta|` per consumed pair.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-83c38318a5d9","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2186,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:369"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_historyParameter_zeroInitialization_le (eta : Real) : forall n (history : History.FinitePairHistory (Fin 2) Real n) (coordinate : Fin 2), (forall i, |(history i).2| <= 1) -> |historyParameter (fun _ : Fin 2 => 0) eta n history coordinate| <= ((n + 1 : Nat) : Real) * |eta|","missing":[],"search":"abs_historyparameter_zeroinitialization_le banditrlproof.stochasticgradientbandit.abs_historyparameter_zeroinitialization_le on a finite history supported in `[-1,1]`, every zero-initialized algorithm-1 parameter coordinate grows by at most `|eta|` per consumed pair. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessorPotential_eq_exp_historyParameter","label":"twoArmForwardTrajectorySuccessorPotential_eq_exp_historyParameter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessorPotential_eq_exp_historyParameter","description":"The forward successor potential is exactly the exponential of the next finite-history parameter.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-616714c62be2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2187,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:456"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardTrajectorySuccessorPotential_eq_exp_historyParameter {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample) = Real.exp (2 * historyParameter (fun _ : Fin 2 => 0) eta (n + 1) (Preorder.frestrictLe (n + 1) sample.2) 0)","missing":[],"search":"twoarmforwardtrajectorysuccessorpotential_eq_exp_historyparameter banditrlproof.stochasticgradientbandit.twoarmforwardtrajectorysuccessorpotential_eq_exp_historyparameter the forward successor potential is exactly the exponential of the next finite-history parameter. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessorPotential_eq_exp_historyParameter","label":"twoArmInverseTrajectorySuccessorPotential_eq_exp_historyParameter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessorPotential_eq_exp_historyParameter","description":"The inverse successor potential has the analogous next-parameter form.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-f0fbbce595f1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2188,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:475"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseTrajectorySuccessorPotential_eq_exp_historyParameter {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample) = Real.exp (-2 * historyParameter (fun _ : Fin 2 => 0) eta (n + 1) (Preorder.frestrictLe (n + 1) sample.2) 0)","missing":[],"search":"twoarminversetrajectorysuccessorpotential_eq_exp_historyparameter banditrlproof.stochasticgradientbandit.twoarminversetrajectorysuccessorpotential_eq_exp_historyparameter the inverse successor potential has the analogous next-parameter form. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardTrajectorySuccessorPotential","label":"integrable_twoArmForwardTrajectorySuccessorPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardTrajectorySuccessorPotential","description":"The actual forward successor exponential is integrable on the canonical trajectory at every fixed successor index. A finite prior is sufficient; no normalization of the prior is claimed or used.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-15ac37a2e9aa","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2189,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:498"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmForwardTrajectorySuccessorPotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample)) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmforwardtrajectorysuccessorpotential banditrlproof.stochasticgradientbandit.integrable_twoarmforwardtrajectorysuccessorpotential the actual forward successor exponential is integrable on the canonical trajectory at every fixed successor index. a finite prior is sufficient; no normalization of the prior is claimed or used. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseTrajectorySuccessorPotential","label":"integrable_twoArmInverseTrajectorySuccessorPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseTrajectorySuccessorPotential","description":"The inverse-odds successor exponential has the same deterministic fixed-time envelope and is integrable on the same canonical trajectory.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-aeb2051f68b6","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2190,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:543"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmInverseTrajectorySuccessorPotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample)) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarminversetrajectorysuccessorpotential banditrlproof.stochasticgradientbandit.integrable_twoarminversetrajectorysuccessorpotential the inverse-odds successor exponential has the same deterministic fixed-time envelope and is integrable on the same canonical trajectory. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","label":"twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","description":"The fixed-time forward potential has the regular-conditional-distribution representation with respect to the environment/prefix sigma-algebra.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-1f1a66ff0797","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2191,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:593"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : (twoArmTrajectoryMeasure prior eta environment)[ fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample) | twoArmPrefixSigma (Env := Env) n] =ᵐ[ twoArmTrajectoryMeasure prior eta environment] fun sample : Env × ((k : Nat) -> Fin 2 × Real) => integral (condDistrib (twoArmNextPair n) (twoArmEnvironmentPrefix n) (twoArmTrajectoryMeasure prior eta environment) (twoArmEnvironmentPrefix n sample)) (twoArmForwardSuccessorPotential eta…","missing":[],"search":"twoarmforwardtrajectorysuccessor_condexp_ae_eq_integral_conddistrib banditrlproof.stochasticgradientbandit.twoarmforwardtrajectorysuccessor_condexp_ae_eq_integral_conddistrib the fixed-time forward potential has the regular-conditional-distribution representation with respect to the environment/prefix sigma-algebra. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","label":"twoArmInverseTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","description":"The inverse potential has the same conditional-distribution representation.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-17a007ef3836","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2192,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:635"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseTrajectorySuccessor_condExp_ae_eq_integral_condDistrib {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : (twoArmTrajectoryMeasure prior eta environment)[ fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample) | twoArmPrefixSigma (Env := Env) n] =ᵐ[ twoArmTrajectoryMeasure prior eta environment] fun sample : Env × ((k : Nat) -> Fin 2 × Real) => integral (condDistrib (twoArmNextPair n) (twoArmEnvironmentPrefix n) (twoArmTrajectoryMeasure prior eta environment) (twoArmEnvironmentPrefix n sample)) (twoArmInverseSuccessorPotential eta…","missing":[],"search":"twoarminversetrajectorysuccessor_condexp_ae_eq_integral_conddistrib banditrlproof.stochasticgradientbandit.twoarminversetrajectorysuccessor_condexp_ae_eq_integral_conddistrib the inverse potential has the same conditional-distribution representation. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","label":"twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","description":"Tower-ready forward conditional recurrence on the actual canonical path.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-8c7f1b339695","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2193,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:676"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : (twoArmTrajectoryMeasure prior eta environment)[ fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample) | twoArmPrefixSigma (Env := Env) n] ≤ᵐ[ twoArmTrajectoryMeasure prior eta environment] fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardRecurrenceBound eta Delta (twoArmEnvironmentPrefix n sample).2","missing":[],"search":"twoarmforwardtrajectorysuccessor_condexp_le_recurrencebound banditrlproof.stochasticgradientbandit.twoarmforwardtrajectorysuccessor_condexp_le_recurrencebound tower-ready forward conditional recurrence on the actual canonical path. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_le_recurrenceBound","label":"twoArmInverseTrajectorySuccessor_condExp_le_recurrenceBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_le_recurrenceBound","description":"Tower-ready inverse conditional recurrence on the same path.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmpathintegrability/index.html#decl-40e8045e765e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","order":2194,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmPathIntegrability.lean:706"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseTrajectorySuccessor_condExp_le_recurrenceBound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : (twoArmTrajectoryMeasure prior eta environment)[ fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample) | twoArmPrefixSigma (Env := Env) n] ≤ᵐ[ twoArmTrajectoryMeasure prior eta environment] fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseRecurrenceBound eta Delta (twoArmEnvironmentPrefix n sample).2","missing":[],"search":"twoarminversetrajectorysuccessor_condexp_le_recurrencebound banditrlproof.stochasticgradientbandit.twoarminversetrajectorysuccessor_condexp_le_recurrencebound tower-ready inverse conditional recurrence on the same path. theorem compiled","shard":"modules/3bf2a8d9e511a7c8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_sum_eq_initial","label":"historyParameter_sum_eq_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyParameter_sum_eq_initial","description":"Algorithm 1 preserves the sum of an arbitrary initial parameter vector on every inclusive finite action/reward history.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-036497fb4e93","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2195,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:28"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyParameter_sum_eq_initial {Action : Type u} [Fintype Action] [DecidableEq Action] [Nonempty Action] (initialTheta : Action -> Real) (eta : Real) : forall n (history : History.FinitePairHistory Action Real n), (∑ coordinate, historyParameter initialTheta eta n history coordinate) = ∑ coordinate, initialTheta coordinate","missing":[],"search":"historyparameter_sum_eq_initial banditrlproof.stochasticgradientbandit.historyparameter_sum_eq_initial algorithm 1 preserves the sum of an arbitrary initial parameter vector on every inclusive finite action/reward history. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_zeroInitialization_sum","label":"historyParameter_zeroInitialization_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyParameter_zeroInitialization_sum","description":"The source initialization `theta = 0` therefore stays zero-sum pathwise.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-49648f58f3bb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2196,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:77"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyParameter_zeroInitialization_sum {Action : Type u} [Fintype Action] [DecidableEq Action] [Nonempty Action] (eta : Real) (n : Nat) (history : History.FinitePairHistory Action Real n) : ∑ coordinate, historyParameter (fun _ : Action => 0) eta n history coordinate = 0","missing":[],"search":"historyparameter_zeroinitialization_sum banditrlproof.stochasticgradientbandit.historyparameter_zeroinitialization_sum the source initialization `theta = 0` therefore stays zero-sum pathwise. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt","label":"twoArmParameterAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmParameterAt","description":"Source-time parameter adapter for one infinite two-arm action/reward trace. Lean time `0` is the pre-action source parameter `theta_{.,1} = 0`; Lean time `n + 1` is the parameter after consuming trace pair `n`, namely source `theta_{.,n+2}`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-3a34833cb2c3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2197,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:90"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def twoArmParameterAt (eta : Real) (trace : Nat -> Fin 2 × Real) : Nat -> Fin 2 -> Real | 0 => fun _ => 0 | n + 1 => historyParameter (fun _ : Fin 2 => 0) eta n (Preorder.frestrictLe n trace) /-- The two-arm softmax law generated from the source-time parameter adapter. -/ def twoArmProbabilityAt (eta : Real) (trace : Nat -> Fin 2 × Real) (time : Nat) : Fin 2 -> Real","missing":[],"search":"twoarmparameterat banditrlproof.stochasticgradientbandit.twoarmparameterat source-time parameter adapter for one infinite two-arm action/reward trace. lean time `0` is the pre-action source parameter `theta_{.,1} = 0`; lean time `n + 1` is the parameter after consuming trace pair `n`, namely source `theta_{.,n+2}`. definition compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt","label":"twoArmProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt","description":"The two-arm softmax law generated from the source-time parameter adapter.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-1b6bf8f3bd1b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2198,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:98"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmProbabilityAt (eta : Real) (trace : Nat -> Fin 2 × Real) (time : Nat) : Fin 2 -> Real","missing":[],"search":"twoarmprobabilityat banditrlproof.stochasticgradientbandit.twoarmprobabilityat the two-arm softmax law generated from the source-time parameter adapter. definition compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_zero","label":"twoArmParameterAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmParameterAt_zero","description":"theorem twoArmParameterAt_zero (eta : Real) (trace : Nat -> Fin 2 × Real) (arm : Fin 2) : twoArmParameterAt eta trace 0 arm = 0","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-feacc9545ffa","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2199,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:103"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmParameterAt_zero (eta : Real) (trace : Nat -> Fin 2 × Real) (arm : Fin 2) : twoArmParameterAt eta trace 0 arm = 0","missing":[],"search":"twoarmparameterat_zero banditrlproof.stochasticgradientbandit.twoarmparameterat_zero theorem twoarmparameterat_zero (eta : real) (trace : nat -> fin 2 × real) (arm : fin 2) : twoarmparameterat eta trace 0 arm = 0 theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_succ","label":"twoArmParameterAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmParameterAt_succ","description":"theorem twoArmParameterAt_succ (eta : Real) (trace : Nat -> Fin 2 × Real) (n : Nat) (arm : Fin 2) : twoArmParameterAt eta trace (n + 1) arm = historyParameter (fun _ : Fin 2 => 0) eta n (Preorder.frestrictLe n trace) arm","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-eb100e464115","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2200,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:108"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmParameterAt_succ (eta : Real) (trace : Nat -> Fin 2 × Real) (n : Nat) (arm : Fin 2) : twoArmParameterAt eta trace (n + 1) arm = historyParameter (fun _ : Fin 2 => 0) eta n (Preorder.frestrictLe n trace) arm","missing":[],"search":"twoarmparameterat_succ banditrlproof.stochasticgradientbandit.twoarmparameterat_succ theorem twoarmparameterat_succ (eta : real) (trace : nat -> fin 2 × real) (n : nat) (arm : fin 2) : twoarmparameterat eta trace (n + 1) arm = historyparameter (fun _ : fin 2 => 0) eta n (preorder.frestrictle n trace) arm theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_sum_eq_zero","label":"twoArmParameterAt_sum_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmParameterAt_sum_eq_zero","description":"The source-time two-arm parameter stays zero-sum on every path.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-0a5081df693c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2201,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:115"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmParameterAt_sum_eq_zero (eta : Real) (trace : Nat -> Fin 2 × Real) (time : Nat) : ∑ arm, twoArmParameterAt eta trace time arm = 0","missing":[],"search":"twoarmparameterat_sum_eq_zero banditrlproof.stochasticgradientbandit.twoarmparameterat_sum_eq_zero the source-time two-arm parameter stays zero-sum on every path. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_one_eq_neg_zero","label":"twoArmParameterAt_one_eq_neg_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmParameterAt_one_eq_neg_zero","description":"Hence source arm `2` has the negative parameter of source arm `1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-00630218227d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2202,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:125"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmParameterAt_one_eq_neg_zero (eta : Real) (trace : Nat -> Fin 2 × Real) (time : Nat) : twoArmParameterAt eta trace time 1 = -twoArmParameterAt eta trace time 0","missing":[],"search":"twoarmparameterat_one_eq_neg_zero banditrlproof.stochasticgradientbandit.twoarmparameterat_one_eq_neg_zero hence source arm `2` has the negative parameter of source arm `1`. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero","label":"twoArmProbabilityAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero","description":"Algorithm 1's zero initialization gives the uniform two-arm law before the first action.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-5a8dbfb5de91","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2203,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:136"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmProbabilityAt_zero (eta : Real) (trace : Nat -> Fin 2 × Real) (arm : Fin 2) : twoArmProbabilityAt eta trace 0 arm = 1 / 2","missing":[],"search":"twoarmprobabilityat_zero banditrlproof.stochasticgradientbandit.twoarmprobabilityat_zero algorithm 1's zero initialization gives the uniform two-arm law before the first action. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_one_eq_one_sub_zero","label":"softmaxProbability_one_eq_one_sub_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_one_eq_one_sub_zero","description":"On two arms, normalization identifies the second softmax probability with the failure probability of the first arm.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-1bc13b7bfd6e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2204,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:145"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_one_eq_one_sub_zero (theta : Fin 2 -> Real) : softmaxProbability theta 1 = 1 - softmaxProbability theta 0","missing":[],"search":"softmaxprobability_one_eq_one_sub_zero banditrlproof.stochasticgradientbandit.softmaxprobability_one_eq_one_sub_zero on two arms, normalization identifies the second softmax probability with the failure probability of the first arm. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one","label":"softmaxProbability_zero_div_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one","description":"The exact two-arm softmax odds before using the zero-sum invariant.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-33e0ff05f870","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2205,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:152"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_zero_div_one (theta : Fin 2 -> Real) : softmaxProbability theta 0 / softmaxProbability theta 1 = Real.exp (theta 0 - theta 1)","missing":[],"search":"softmaxprobability_zero_div_one banditrlproof.stochasticgradientbandit.softmaxprobability_zero_div_one the exact two-arm softmax odds before using the zero-sum invariant. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.finTwo_one_eq_neg_zero_of_sum_eq_zero","label":"finTwo_one_eq_neg_zero_of_sum_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.finTwo_one_eq_neg_zero_of_sum_eq_zero","description":"A zero-sum two-arm parameter has opposite coordinates.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-3e4c07eab53a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2206,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:166"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem finTwo_one_eq_neg_zero_of_sum_eq_zero (theta : Fin 2 -> Real) (hsum : ∑ coordinate, theta coordinate = 0) : theta 1 = -theta 0","missing":[],"search":"fintwo_one_eq_neg_zero_of_sum_eq_zero banditrlproof.stochasticgradientbandit.fintwo_one_eq_neg_zero_of_sum_eq_zero a zero-sum two-arm parameter has opposite coordinates. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul","label":"softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul","description":"The printed division form of Equation (11). Lean arm `0` is source arm `1`, and strict softmax positivity makes the denominator nonzero.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-50d2bb0ada8f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2207,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:174"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul (theta : Fin 2 -> Real) (hsum : ∑ coordinate, theta coordinate = 0) : softmaxProbability theta 0 / (1 - softmaxProbability theta 0) = Real.exp (2 * theta 0)","missing":[],"search":"softmaxprobability_zero_div_one_sub_zero_eq_exp_two_mul banditrlproof.stochasticgradientbandit.softmaxprobability_zero_div_one_sub_zero_eq_exp_two_mul the printed division form of equation (11). lean arm `0` is source arm `1`, and strict softmax positivity makes the denominator nonzero. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.exp_two_mul_zero_mul_one_sub_softmaxProbability_zero","label":"exp_two_mul_zero_mul_one_sub_softmaxProbability_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.exp_two_mul_zero_mul_one_sub_softmaxProbability_zero","description":"The multiplication form of Equation (11) used in the source Theorem 1 proof: `exp(2 theta_1) (1 - p_1) = p_1`. Lean arm `0` is source arm `1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-f5ee71ec9b2a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2208,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:185"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem exp_two_mul_zero_mul_one_sub_softmaxProbability_zero (theta : Fin 2 -> Real) (hsum : ∑ coordinate, theta coordinate = 0) : Real.exp (2 * theta 0) * (1 - softmaxProbability theta 0) = softmaxProbability theta 0","missing":[],"search":"exp_two_mul_zero_mul_one_sub_softmaxprobability_zero banditrlproof.stochasticgradientbandit.exp_two_mul_zero_mul_one_sub_softmaxprobability_zero the multiplication form of equation (11) used in the source theorem 1 proof: `exp(2 theta_1) (1 - p_1) = p_1`. lean arm `0` is source arm `1`. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.exp_neg_two_mul_zero_mul_softmaxProbability_zero","label":"exp_neg_two_mul_zero_mul_softmaxProbability_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.exp_neg_two_mul_zero_mul_softmaxProbability_zero","description":"The inverse-odds form used for the failure-mass telescoping potential in the second half of the source Theorem 1 proof.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-13a8d02667e4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2209,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:199"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_two_mul_zero_mul_softmaxProbability_zero (theta : Fin 2 -> Real) (hsum : ∑ coordinate, theta coordinate = 0) : Real.exp (-2 * theta 0) * softmaxProbability theta 0 = 1 - softmaxProbability theta 0","missing":[],"search":"exp_neg_two_mul_zero_mul_softmaxprobability_zero banditrlproof.stochasticgradientbandit.exp_neg_two_mul_zero_mul_softmaxprobability_zero the inverse-odds form used for the failure-mass telescoping potential in the second half of the source theorem 1 proof. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_exp_two_mul_failure_eq_success","label":"twoArmProbabilityAt_exp_two_mul_failure_eq_success","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_exp_two_mul_failure_eq_success","description":"Source-time Equation (11) on every infinite action/reward trace.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-9b184439db46","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2210,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:212"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmProbabilityAt_exp_two_mul_failure_eq_success (eta : Real) (trace : Nat -> Fin 2 × Real) (time : Nat) : Real.exp (2 * twoArmParameterAt eta trace time 0) * (1 - twoArmProbabilityAt eta trace time 0) = twoArmProbabilityAt eta trace time 0","missing":[],"search":"twoarmprobabilityat_exp_two_mul_failure_eq_success banditrlproof.stochasticgradientbandit.twoarmprobabilityat_exp_two_mul_failure_eq_success source-time equation (11) on every infinite action/reward trace. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","label":"twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","description":"The printed Equation (11) on the explicitly fenced source-time trace.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-f538371ddba4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2211,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:222"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul (eta : Real) (trace : Nat -> Fin 2 × Real) (time : Nat) : twoArmProbabilityAt eta trace time 0 / (1 - twoArmProbabilityAt eta trace time 0) = Real.exp (2 * twoArmParameterAt eta trace time 0)","missing":[],"search":"twoarmprobabilityat_zero_div_failure_eq_exp_two_mul banditrlproof.stochasticgradientbandit.twoarmprobabilityat_zero_div_failure_eq_exp_two_mul the printed equation (11) on the explicitly fenced source-time trace. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_exp_two_mul_zero_eq_odds","label":"historyParameter_exp_two_mul_zero_eq_odds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.historyParameter_exp_two_mul_zero_eq_odds","description":"Finite-history multiplication form of Equation (11) under Algorithm 1's source initialization.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrate/index.html#decl-7583426f4266","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","order":2212,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRate.lean:233"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem historyParameter_exp_two_mul_zero_eq_odds (eta : Real) (n : Nat) (history : History.FinitePairHistory (Fin 2) Real n) : Real.exp (2 * historyParameter (fun _ : Fin 2 => 0) eta n history 0) * (1 - softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n history) 0) = softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n history) 0","missing":[],"search":"historyparameter_exp_two_mul_zero_eq_odds banditrlproof.stochasticgradientbandit.historyparameter_exp_two_mul_zero_eq_odds finite-history multiplication form of equation (11) under algorithm 1's source initialization. theorem compiled","shard":"modules/e1923f1f1d395ac4.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardQ","label":"twoArmForwardQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardQ","description":"The action-dependent coefficient in the forward exponential potential. For source arm `1` (Lean arm `0`) it is `2 eta (1-p)`; for source arm `2` (Lean arm `1`) it is `-2 eta p`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-8b2c46ddda17","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2213,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:49"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmForwardQ (eta : Real) (prob : Fin 2 -> Real) (selected : Fin 2) : Real","missing":[],"search":"twoarmforwardq banditrlproof.stochasticgradientbandit.twoarmforwardq the action-dependent coefficient in the forward exponential potential. for source arm `1` (lean arm `0`) it is `2 eta (1-p)`; for source arm `2` (lean arm `1`) it is `-2 eta p`. definition compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseQ","label":"twoArmInverseQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseQ","description":"The coefficient for the inverse-odds exponential potential.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-4fed5c72c253","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2214,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:54"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInverseQ (eta : Real) (prob : Fin 2 -> Real) (selected : Fin 2) : Real","missing":[],"search":"twoarminverseq banditrlproof.stochasticgradientbandit.twoarminverseq the coefficient for the inverse-odds exponential potential. definition compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardQ_mul_reward_eq_sourceIncrement","label":"twoArmForwardQ_mul_reward_eq_sourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardQ_mul_reward_eq_sourceIncrement","description":"`q_+ reward` is exactly twice the learning-rate-scaled best-coordinate source update.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-7a110cc4c2d7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2215,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:60"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardQ_mul_reward_eq_sourceIncrement (eta reward : Real) (prob : Fin 2 -> Real) (selected : Fin 2) : twoArmForwardQ eta prob selected * reward = 2 * eta * sourceIncrement prob reward selected 0","missing":[],"search":"twoarmforwardq_mul_reward_eq_sourceincrement banditrlproof.stochasticgradientbandit.twoarmforwardq_mul_reward_eq_sourceincrement `q_+ reward` is exactly twice the learning-rate-scaled best-coordinate source update. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseQ_mul_reward_eq_sourceIncrement","label":"twoArmInverseQ_mul_reward_eq_sourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseQ_mul_reward_eq_sourceIncrement","description":"`q_- reward` is the negative of twice the learning-rate-scaled best-coordinate source update.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-e91c0fd9c636","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2216,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:68"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseQ_mul_reward_eq_sourceIncrement (eta reward : Real) (prob : Fin 2 -> Real) (selected : Fin 2) : twoArmInverseQ eta prob selected * reward = -2 * eta * sourceIncrement prob reward selected 0","missing":[],"search":"twoarminverseq_mul_reward_eq_sourceincrement banditrlproof.stochasticgradientbandit.twoarminverseq_mul_reward_eq_sourceincrement `q_- reward` is the negative of twice the learning-rate-scaled best-coordinate source update. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardEqEightRemainder_le","label":"twoArmForwardEqEightRemainder_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardEqEightRemainder_le","description":"The two-branch Equation-(8) remainder for the forward potential, after replacing both branch constants by the common source constant `C_eta`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-ef3d870fa15b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2217,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:78"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardEqEightRemainder_le (eta p meanZero meanOne Delta : Real) (heta : 0 <= eta) (hp_nonneg : 0 <= p) (hp_le_one : p <= 1) (hgap : meanZero - meanOne = Delta) : p * ((2 * eta * (1 - p)) * meanZero + (2 * eta * (1 - p)) ^ 2 / 2 * sourceC (|2 * eta * (1 - p)| / 2)) + (1 - p) * ((-(2 * eta * p)) * meanOne + (-(2 * eta * p)) ^ 2 / 2 * sourceC (|-(2 * eta * p)| / 2)) <= 2 * p * (1 - p) * (eta * Delta + eta ^ 2 * sourceC eta)","missing":[],"search":"twoarmforwardeqeightremainder_le banditrlproof.stochasticgradientbandit.twoarmforwardeqeightremainder_le the two-branch equation-(8) remainder for the forward potential, after replacing both branch constants by the common source constant `c_eta`. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseEqEightRemainder_le","label":"twoArmInverseEqEightRemainder_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseEqEightRemainder_le","description":"The analogous Equation-(8) remainder for the inverse potential.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-a0512a70681e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2218,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:162"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseEqEightRemainder_le (eta p meanZero meanOne Delta : Real) (heta : 0 <= eta) (hp_nonneg : 0 <= p) (hp_le_one : p <= 1) (hgap : meanZero - meanOne = Delta) : p * ((-(2 * eta * (1 - p))) * meanZero + (-(2 * eta * (1 - p))) ^ 2 / 2 * sourceC (|-(2 * eta * (1 - p))| / 2)) + (1 - p) * ((2 * eta * p) * meanOne + (2 * eta * p) ^ 2 / 2 * sourceC (|2 * eta * p| / 2)) <= -2 * eta * p * (1 - p) * (Delta - eta * sourceC eta)","missing":[],"search":"twoarminverseeqeightremainder_le banditrlproof.stochasticgradientbandit.twoarminverseeqeightremainder_le the analogous equation-(8) remainder for the inverse potential. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le","label":"integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le","description":"Multiplicative forward one-step recurrence at a fixed generated history. The current parameter is source `theta_{.,n+2}` and the integrand is the source successor `theta_{1,n+3}`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-d45caea73efe","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2219,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:245"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.HistoryEnvironment (Fin 2) Real) (n : Nat) (history : History.FinitePairHistory (Fin 2) Real n) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (history, selected), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (history, selected)) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.historyStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n history) (fun pair : Fin 2 × Real => Real.exp (2 * (historyParameter (fun _ : Fin 2 => 0) eta n history 0 + eta * sourceIncrement (softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n history)) pair.2 pair.1 0))) <= Real.exp (2 * historyParameter (fun _ : Fin 2 => 0) eta n history 0) * (1 + 2 * softmaxProbability…","missing":[],"search":"integral_twoarmhistorystepkernel_exp_forwardsuccessor_le banditrlproof.stochasticgradientbandit.integral_twoarmhistorystepkernel_exp_forwardsuccessor_le multiplicative forward one-step recurrence at a fixed generated history. the current parameter is source `theta_{.,n+2}` and the integrand is the source successor `theta_{1,n+3}`. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le_add_success_sq","label":"integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le_add_success_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le_add_success_sq","description":"Forward recurrence in the additive source form `E_t exp(2 theta_{1,t+1}) <= exp(2 theta_{1,t}) + 2 p_t^2 (eta Delta + eta^2 C_eta)`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-b48a31268137","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2220,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:340"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le_add_success_sq (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.HistoryEnvironment (Fin 2) Real) (n : Nat) (history : History.FinitePairHistory (Fin 2) Real n) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (history, selected), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (history, selected)) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.historyStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n history) (fun pair : Fin 2 × Real => Real.exp (2 * (historyParameter (fun _ : Fin 2 => 0) eta n history 0 + eta * sourceIncrement (softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n history)) pair.2 pair.1 0))) <= Real.exp (2 * historyParameter (fun _ : Fin 2 => 0) eta n history 0) + 2 * softmaxPr…","missing":[],"search":"integral_twoarmhistorystepkernel_exp_forwardsuccessor_le_add_success_sq banditrlproof.stochasticgradientbandit.integral_twoarmhistorystepkernel_exp_forwardsuccessor_le_add_success_sq forward recurrence in the additive source form `e_t exp(2 theta_{1,t+1}) <= exp(2 theta_{1,t}) + 2 p_t^2 (eta delta + eta^2 c_eta)`. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le","label":"integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le","description":"Multiplicative inverse-potential one-step recurrence at the same time fence.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-be7cb65a4fd0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2221,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:394"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.HistoryEnvironment (Fin 2) Real) (n : Nat) (history : History.FinitePairHistory (Fin 2) Real n) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (history, selected), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (history, selected)) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.historyStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n history) (fun pair : Fin 2 × Real => Real.exp (-2 * (historyParameter (fun _ : Fin 2 => 0) eta n history 0 + eta * sourceIncrement (softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n history)) pair.2 pair.1 0))) <= Real.exp (-2 * historyParameter (fun _ : Fin 2 => 0) eta n history 0) * (1 - 2 * eta * softmaxProb…","missing":[],"search":"integral_twoarmhistorystepkernel_exp_inversesuccessor_le banditrlproof.stochasticgradientbandit.integral_twoarmhistorystepkernel_exp_inversesuccessor_le multiplicative inverse-potential one-step recurrence at the same time fence. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le_sub_failure_sq","label":"integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le_sub_failure_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le_sub_failure_sq","description":"Inverse recurrence in the source telescoping form `E_t exp(-2 theta_{1,t+1}) <= exp(-2 theta_{1,t}) - 2 eta (1-p_t)^2 (Delta-eta C_eta)`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmrecurrence/index.html#decl-5c21160d3bb5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","order":2222,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmRecurrence.lean:494"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le_sub_failure_sq (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.HistoryEnvironment (Fin 2) Real) (n : Nat) (history : History.FinitePairHistory (Fin 2) Real n) (mean : Fin 2 -> Real) (hreward : forall selected, ∀ᵐ reward ∂environment.feedback n (history, selected), |reward| <= 1) (hmean : forall selected, integral (environment.feedback n (history, selected)) id = mean selected) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.historyStepKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment n history) (fun pair : Fin 2 × Real => Real.exp (-2 * (historyParameter (fun _ : Fin 2 => 0) eta n history 0 + eta * sourceIncrement (softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n history)) pair.2 pair.1 0))) <= Real.exp (-2 * historyParameter (fun _ : Fin 2 => 0) eta n history 0) - 2 * eta * (…","missing":[],"search":"integral_twoarmhistorystepkernel_exp_inversesuccessor_le_sub_failure_sq banditrlproof.stochasticgradientbandit.integral_twoarmhistorystepkernel_exp_inversesuccessor_le_sub_failure_sq inverse recurrence in the source telescoping form `e_t exp(-2 theta_{1,t+1}) <= exp(-2 theta_{1,t}) - 2 eta (1-p_t)^2 (delta-eta c_eta)`. theorem compiled","shard":"modules/2f9129badec0ede0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement","label":"twoArmTrajectorySourceIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement","description":"def twoArmTrajectorySourceIncrement {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-7e34f36dc42f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2223,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:34"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmTrajectorySourceIncrement {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmtrajectorysourceincrement banditrlproof.stochasticgradientbandit.twoarmtrajectorysourceincrement def twoarmtrajectorysourceincrement {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectorySourceIncrement","label":"measurable_twoArmTrajectorySourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectorySourceIncrement","description":"theorem measurable_twoArmTrajectorySourceIncrement {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmTrajectorySourceIncrement (Env := Env) eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-079de65c2783","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2224,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:44"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmTrajectorySourceIncrement {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmTrajectorySourceIncrement (Env := Env) eta n)","missing":[],"search":"measurable_twoarmtrajectorysourceincrement banditrlproof.stochasticgradientbandit.measurable_twoarmtrajectorysourceincrement theorem measurable_twoarmtrajectorysourceincrement {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (twoarmtrajectorysourceincrement (env := env) eta n) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectorySourceIncrement","label":"integrable_twoArmTrajectorySourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectorySourceIncrement","description":"theorem integrable_twoArmTrajectorySourceIncrement {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmTrajectorySourceIncrement (Env := Env) eta n) (twoA…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-3863d6b38fc8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2225,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:65"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmTrajectorySourceIncrement {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmTrajectorySourceIncrement (Env := Env) eta n) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmtrajectorysourceincrement banditrlproof.stochasticgradientbandit.integrable_twoarmtrajectorysourceincrement theorem integrable_twoarmtrajectorysourceincrement {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : integrable (twoarmtrajectorysourceincrement (env := env) eta n) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib","label":"twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib","description":"theorem twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : (twoArmTrajectoryMeasure prior eta environmen…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-7ee470fdab59","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2226,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:87"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : (twoArmTrajectoryMeasure prior eta environment)[ twoArmTrajectorySourceIncrement (Env := Env) eta n | twoArmPrefixSigma (Env := Env) n] =ᵐ[ twoArmTrajectoryMeasure prior eta environment] fun sample : Env × ((k : Nat) -> Fin 2 × Real) => integral (condDistrib (twoArmNextPair n) (twoArmEnvironmentPrefix n) (twoArmTrajectoryMeasure prior eta environment) (twoArmEnvironmentPrefix n sample)) (fun pair : Fin 2 × Real => sourceIncrement (softmaxProbability (historyParameter (fun _ : Fin 2 => 0) eta n (twoArmEnvironmentPrefix n…","missing":[],"search":"twoarmtrajectorysourceincrement_condexp_ae_eq_integral_conddistrib banditrlproof.stochasticgradientbandit.twoarmtrajectorysourceincrement_condexp_ae_eq_integral_conddistrib theorem twoarmtrajectorysourceincrement_condexp_ae_eq_integral_conddistrib {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : (twoarmtrajectorymeasure prior eta environment)[ twoarmtrajectorysourceincrement (env := env) eta n | twoarmprefixsigma (env := env) n] =ᵐ[ twoarmtrajectorymeasure prior eta environment] fun sample : env × ((k : nat) -> fin 2 × real) => integral (conddistrib (twoarmnextpair n) (twoarmenvironmentprefix n) (twoarmtrajectorymeasure prior eta environment) (twoarmenvironmentprefix n sample)) (fun pair : fin 2 × real => sourceincrement (softmaxprobability (historyparameter (fun _ : fin 2 => 0) eta n (twoarmenvironmentprefix n sample).2)) pair.2 pair.1 0) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure","label":"twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure","description":"theorem twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : (twoArmTraje…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-b2d35a05e1d1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2227,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:162"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : (twoArmTrajectoryMeasure prior eta environment)[ twoArmTrajectorySourceIncrement (Env := Env) eta n | twoArmPrefixSigma (Env := Env) n] =ᵐ[ twoArmTrajectoryMeasure prior eta environment] fun sample => Delta * twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample","missing":[],"search":"twoarmtrajectorysourceincrement_condexp_ae_eq_successfailure banditrlproof.stochasticgradientbandit.twoarmtrajectorysourceincrement_condexp_ae_eq_successfailure theorem twoarmtrajectorysourceincrement_condexp_ae_eq_successfailure {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (n : nat) : (twoarmtrajectorymeasure prior eta environment)[ twoarmtrajectorysourceincrement (env := env) eta n | twoarmprefixsigma (env := env) n] =ᵐ[ twoarmtrajectorymeasure prior eta environment] fun sample => delta * twoarmsuccessprobability eta n sample * twoarmfailuremass eta n sample theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero_succ","label":"twoArmTrajectoryParameterZero_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero_succ","description":"theorem twoArmTrajectoryParameterZero_succ {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmTrajectoryParameterZero eta (n + 1) sample = twoArmTrajectoryParameterZero eta n sample + eta * twoArmTrajectorySourceIncrement eta n sample","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-32926ca08b17","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2228,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:235"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmTrajectoryParameterZero_succ {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmTrajectoryParameterZero eta (n + 1) sample = twoArmTrajectoryParameterZero eta n sample + eta * twoArmTrajectorySourceIncrement eta n sample","missing":[],"search":"twoarmtrajectoryparameterzero_succ banditrlproof.stochasticgradientbandit.twoarmtrajectoryparameterzero_succ theorem twoarmtrajectoryparameterzero_succ {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : twoarmtrajectoryparameterzero eta (n + 1) sample = twoarmtrajectoryparameterzero eta n sample + eta * twoarmtrajectorysourceincrement eta n sample theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectoryParameterZero","label":"integrable_twoArmTrajectoryParameterZero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectoryParameterZero","description":"theorem integrable_twoArmTrajectoryParameterZero {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmTrajectoryParameterZero (Env := Env) eta n) (twoArmTr…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-9d4e039079fc","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2229,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:252"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmTrajectoryParameterZero {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmTrajectoryParameterZero (Env := Env) eta n) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmtrajectoryparameterzero banditrlproof.stochasticgradientbandit.integrable_twoarmtrajectoryparameterzero theorem integrable_twoarmtrajectoryparameterzero {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : integrable (twoarmtrajectoryparameterzero (env := env) eta n) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectorySourceIncrement_eq_successFailure","label":"integral_twoArmTrajectorySourceIncrement_eq_successFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectorySourceIncrement_eq_successFailure","description":"theorem integral_twoArmTrajectorySourceIncrement_eq_successFailure {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoA…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-f1060148561f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2230,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:273"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmTrajectorySourceIncrement_eq_successFailure {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectorySourceIncrement (Env := Env) eta n) = integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => Delta * twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample)","missing":[],"search":"integral_twoarmtrajectorysourceincrement_eq_successfailure banditrlproof.stochasticgradientbandit.integral_twoarmtrajectorysourceincrement_eq_successfailure theorem integral_twoarmtrajectorysourceincrement_eq_successfailure {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectorysourceincrement (env := env) eta n) = integral (twoarmtrajectorymeasure prior eta environment) (fun sample => delta * twoarmsuccessprobability eta n sample * twoarmfailuremass eta n sample) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_succ","label":"integral_twoArmTrajectoryParameterZero_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_succ","description":"theorem integral_twoArmTrajectoryParameterZero_succ {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTrajectoryMea…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-25678c02942c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2231,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:305"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmTrajectoryParameterZero_succ {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta (n + 1)) = integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta n) + eta * integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => Delta * twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample)","missing":[],"search":"integral_twoarmtrajectoryparameterzero_succ banditrlproof.stochasticgradientbandit.integral_twoarmtrajectoryparameterzero_succ theorem integral_twoarmtrajectoryparameterzero_succ {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta (n + 1)) = integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta n) + eta * integral (twoarmtrajectorymeasure prior eta environment) (fun sample => delta * twoarmsuccessprobability eta n sample * twoarmfailuremass eta n sample) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialSourceIncrement","label":"twoArmInitialSourceIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInitialSourceIncrement","description":"def twoArmInitialSourceIncrement (pair : Fin 2 × Real) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-0315cc4a4ed0","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2232,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:345"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInitialSourceIncrement (pair : Fin 2 × Real) : Real","missing":[],"search":"twoarminitialsourceincrement banditrlproof.stochasticgradientbandit.twoarminitialsourceincrement def twoarminitialsourceincrement (pair : fin 2 × real) : real definition compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialSourceIncrement","label":"measurable_twoArmInitialSourceIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialSourceIncrement","description":"theorem measurable_twoArmInitialSourceIncrement : Measurable twoArmInitialSourceIncrement","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-59d0b7f66cbb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2233,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:350"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInitialSourceIncrement : Measurable twoArmInitialSourceIncrement","missing":[],"search":"measurable_twoarminitialsourceincrement banditrlproof.stochasticgradientbandit.measurable_twoarminitialsourceincrement theorem measurable_twoarminitialsourceincrement : measurable twoarminitialsourceincrement theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialSourceIncrement_eq_quarter_gap","label":"integral_twoArmInitialSourceIncrement_eq_quarter_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInitialSourceIncrement_eq_quarter_gap","description":"theorem integral_twoArmInitialSourceIncrement_eq_quarter_gap {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-c5da08299c65","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2234,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:356"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInitialSourceIncrement_eq_quarter_gap {Env : Type v} [MeasurableSpace Env] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) twoArmInitialSourceIncrement = Delta / 4","missing":[],"search":"integral_twoarminitialsourceincrement_eq_quarter_gap banditrlproof.stochasticgradientbandit.integral_twoarminitialsourceincrement_eq_quarter_gap theorem integral_twoarminitialsourceincrement_eq_quarter_gap {env : type v} [measurablespace env] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (env : env) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) : integral (thompson.measurableenvironmentinitialpairkernel (historyalgorithm (fun _ : fin 2 => 0) eta) environment env) twoarminitialsourceincrement = delta / 4 theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_zero","label":"integral_twoArmTrajectoryParameterZero_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_zero","description":"theorem integral_twoArmTrajectoryParameterZero_zero {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (twoArmTrajectoryMeasure…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-108f022c6d5a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2235,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:396"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmTrajectoryParameterZero_zero {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta 0) = eta * Delta / 4","missing":[],"search":"integral_twoarmtrajectoryparameterzero_zero banditrlproof.stochasticgradientbandit.integral_twoarmtrajectoryparameterzero_zero theorem integral_twoarmtrajectoryparameterzero_zero {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta 0) = eta * delta / 4 theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_eq_successFailureSum","label":"integral_twoArmTrajectoryParameterZero_eq_successFailureSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_eq_successFailureSum","description":"theorem integral_twoArmTrajectoryParameterZero_eq_successFailureSum {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (tailHorizon : Nat)…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-10fdbf81e69b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2236,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:486"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmTrajectoryParameterZero_eq_successFailureSum {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta tailHorizon) = eta * Delta * ((1 : Real) / 4 + (Finset.range tailHorizon).sum (fun n => integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample)))","missing":[],"search":"integral_twoarmtrajectoryparameterzero_eq_successfailuresum banditrlproof.stochasticgradientbandit.integral_twoarmtrajectoryparameterzero_eq_successfailuresum theorem integral_twoarmtrajectoryparameterzero_eq_successfailuresum {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta tailhorizon) = eta * delta * ((1 : real) / 4 + (finset.range tailhorizon).sum (fun n => integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmsuccessprobability eta n sample * twoarmfailuremass eta n sample))) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessFailureMass","label":"measurable_twoArmSuccessFailureMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessFailureMass","description":"theorem measurable_twoArmSuccessFailureMass {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-85be282ff5ff","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2237,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:532"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmSuccessFailureMass {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample)","missing":[],"search":"measurable_twoarmsuccessfailuremass banditrlproof.stochasticgradientbandit.measurable_twoarmsuccessfailuremass theorem measurable_twoarmsuccessfailuremass {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (fun sample : env × ((k : nat) -> fin 2 × real) => twoarmsuccessprobability eta n sample * twoarmfailuremass eta n sample) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessFailureMass","label":"integrable_twoArmSuccessFailureMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessFailureMass","description":"theorem integrable_twoArmSuccessFailureMass {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : Integrable (fun sample => twoArmSuccessProbability (Env := Env) eta n sample * twoArmFailureMass eta n sample) (twoArmTrajectoryMeasure prior eta environment)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-e6f67eaa0692","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2238,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:541"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmSuccessFailureMass {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : Integrable (fun sample => twoArmSuccessProbability (Env := Env) eta n sample * twoArmFailureMass eta n sample) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmsuccessfailuremass banditrlproof.stochasticgradientbandit.integrable_twoarmsuccessfailuremass theorem integrable_twoarmsuccessfailuremass {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (n : nat) : integrable (fun sample => twoarmsuccessprobability (env := env) eta n sample * twoarmfailuremass eta n sample) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFailureMass_eq_successFailure_add_sq","label":"integral_twoArmFailureMass_eq_successFailure_add_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmFailureMass_eq_successFailure_add_sq","description":"theorem integral_twoArmFailureMass_eq_successFailure_add_sq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample) = integral (twoArmTrajectoryMeasure pr…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-9472500f3c3f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2239,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:565"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmFailureMass_eq_successFailure_add_sq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample) = integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability eta n sample * twoArmFailureMass eta n sample) + integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass eta n sample ^ 2)","missing":[],"search":"integral_twoarmfailuremass_eq_successfailure_add_sq banditrlproof.stochasticgradientbandit.integral_twoarmfailuremass_eq_successfailure_add_sq theorem integral_twoarmfailuremass_eq_successfailure_add_sq {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass (env := env) eta n sample) = integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmsuccessprobability eta n sample * twoarmfailuremass eta n sample) + integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass eta n sample ^ 2) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret","label":"twoArmGeneratedExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret","description":"Expected pseudo-regret for source rounds `1, ..., tailHorizon + 1`. The first round is uniform, and Lean tail index `n` is source round `n + 2`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-1b23e9149fa1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2240,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:597"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmGeneratedExpectedPseudoRegret {Env : Type v} [MeasurableSpace Env] (prior : Measure Env) (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (tailHorizon : Nat) : Real","missing":[],"search":"twoarmgeneratedexpectedpseudoregret banditrlproof.stochasticgradientbandit.twoarmgeneratedexpectedpseudoregret expected pseudo-regret for source rounds `1, ..., tailhorizon + 1`. the first round is uniform, and lean tail index `n` is source round `n + 2`. definition compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq","label":"twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq","description":"theorem twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta)…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-0950628466a5","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2241,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:607"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (tailHorizon : Nat) : twoArmGeneratedExpectedPseudoRegret prior eta Delta environment tailHorizon = integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta tailHorizon) / eta + Delta * ((1 : Real) / 4 + (Finset.range tailHorizon).sum (fun n => integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass eta n sample ^ 2)))","missing":[],"search":"twoarmgeneratedexpectedpseudoregret_eq_parameter_add_failuresq banditrlproof.stochasticgradientbandit.twoarmgeneratedexpectedpseudoregret_eq_parameter_add_failuresq theorem twoarmgeneratedexpectedpseudoregret_eq_parameter_add_failuresq {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 < eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (tailhorizon : nat) : twoarmgeneratedexpectedpseudoregret prior eta delta environment tailhorizon = integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta tailhorizon) / eta + delta * ((1 : real) / 4 + (finset.range tailhorizon).sum (fun n => integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass eta n sample ^ 2))) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential","label":"integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential","description":"theorem integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (tailHorizon : Nat) : integral (twoArmTrajectoryMea…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-10ee15e5b28b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2242,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:646"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta tailHorizon) <= Real.log (integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta tailHorizon)) / 2","missing":[],"search":"integral_twoarmtrajectoryparameterzero_le_half_log_forwardpotential banditrlproof.stochasticgradientbandit.integral_twoarmtrajectoryparameterzero_le_half_log_forwardpotential theorem integral_twoarmtrajectoryparameterzero_le_half_log_forwardpotential {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta tailhorizon) <= real.log (integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta tailhorizon)) / 2 theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessProbability_sq_le_one","label":"integral_twoArmSuccessProbability_sq_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessProbability_sq_le_one","description":"theorem integral_twoArmSuccessProbability_sq_le_one {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) <= 1","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-b417c09e418a","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2243,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:694"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmSuccessProbability_sq_le_one {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) <= 1","missing":[],"search":"integral_twoarmsuccessprobability_sq_le_one banditrlproof.stochasticgradientbandit.integral_twoarmsuccessprobability_sq_le_one theorem integral_twoarmsuccessprobability_sq_le_one {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmsuccessprobability (env := env) eta n sample ^ 2) <= 1 theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_le_source_bound","label":"integral_twoArmForwardPotential_le_source_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_le_source_bound","description":"theorem integral_twoArmForwardPotential_le_source_bound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = D…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-3b2a9a38f0ac","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2244,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:725"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmForwardPotential_le_source_bound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta tailHorizon) <= 1 + 4 * eta * Delta * ((tailHorizon + 1 : Nat) : Real)","missing":[],"search":"integral_twoarmforwardpotential_le_source_bound banditrlproof.stochasticgradientbandit.integral_twoarmforwardpotential_le_source_bound theorem integral_twoarmforwardpotential_le_source_bound {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 < eta) (hdelta : 0 < delta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (hmargin : eta * sourcec eta < delta) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta tailhorizon) <= 1 + 4 * eta * delta * ((tailhorizon + 1 : nat) : real) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_source_log_bound","label":"integral_twoArmTrajectoryParameterZero_le_source_log_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_source_log_bound","description":"theorem integral_twoArmTrajectoryParameterZero_le_source_log_bound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 -…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-cb50279e11f4","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2245,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:790"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmTrajectoryParameterZero_le_source_log_bound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmTrajectoryParameterZero (Env := Env) eta tailHorizon) <= Real.log (1 + 4 * eta * Delta * ((tailHorizon + 1 : Nat) : Real)) / 2","missing":[],"search":"integral_twoarmtrajectoryparameterzero_le_source_log_bound banditrlproof.stochasticgradientbandit.integral_twoarmtrajectoryparameterzero_le_source_log_bound theorem integral_twoarmtrajectoryparameterzero_le_source_log_bound {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 < eta) (hdelta : 0 < delta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (hmargin : eta * sourcec eta < delta) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmtrajectoryparameterzero (env := env) eta tailhorizon) <= real.log (1 + 4 * eta * delta * ((tailhorizon + 1 : nat) : real)) / 2 theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmActionGap","label":"twoArmActionGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmActionGap","description":"def twoArmActionGap (Delta : Real) (action : Fin 2) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-aef4926772b2","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2246,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:844"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmActionGap (Delta : Real) (action : Fin 2) : Real","missing":[],"search":"twoarmactiongap banditrlproof.stochasticgradientbandit.twoarmactiongap def twoarmactiongap (delta : real) (action : fin 2) : real definition compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmActionGap","label":"measurable_twoArmActionGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmActionGap","description":"theorem measurable_twoArmActionGap (Delta : Real) : Measurable (twoArmActionGap Delta)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-f2d37592db40","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2247,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:847"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmActionGap (Delta : Real) : Measurable (twoArmActionGap Delta)","missing":[],"search":"measurable_twoarmactiongap banditrlproof.stochasticgradientbandit.measurable_twoarmactiongap theorem measurable_twoarmactiongap (delta : real) : measurable (twoarmactiongap delta) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialActionGap_eq_half","label":"integral_twoArmInitialActionGap_eq_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInitialActionGap_eq_half","description":"theorem integral_twoArmInitialActionGap_eq_half {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmActionGap Delta (sample.2 0).1) = Delta / 2","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-26436bbdcd1f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2248,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:851"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInitialActionGap_eq_half {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmActionGap Delta (sample.2 0).1) = Delta / 2","missing":[],"search":"integral_twoarminitialactiongap_eq_half banditrlproof.stochasticgradientbandit.integral_twoarminitialactiongap_eq_half theorem integral_twoarminitialactiongap_eq_half {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) : integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmactiongap delta (sample.2 0).1) = delta / 2 theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessorActionGap_eq_failureMass","label":"integral_twoArmSuccessorActionGap_eq_failureMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessorActionGap_eq_failureMass","description":"theorem integral_twoArmSuccessorActionGap_eq_failureMass {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmActionGap Delta (sample.2 (n + 1)).1) = Delta * integral (twoArmTraje…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-da04c63f5848","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2249,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:915"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmSuccessorActionGap_eq_failureMass {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmActionGap Delta (sample.2 (n + 1)).1) = Delta * integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample)","missing":[],"search":"integral_twoarmsuccessoractiongap_eq_failuremass banditrlproof.stochasticgradientbandit.integral_twoarmsuccessoractiongap_eq_failuremass theorem integral_twoarmsuccessoractiongap_eq_failuremass {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmactiongap delta (sample.2 (n + 1)).1) = delta * integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass (env := env) eta n sample) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret","label":"twoArmSampledPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret","description":"def twoArmSampledPseudoRegret {Env : Type v} (Delta : Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-66e26454a0b8","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2250,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:1007"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmSampledPseudoRegret {Env : Type v} (Delta : Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmsampledpseudoregret banditrlproof.stochasticgradientbandit.twoarmsampledpseudoregret def twoarmsampledpseudoregret {env : type v} (delta : real) (horizon : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_eq_generated","label":"integral_twoArmSampledPseudoRegret_eq_generated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_eq_generated","description":"theorem integral_twoArmSampledPseudoRegret_eq_generated {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmSampledPseudoRegret (Env := Env) Delta (tailHorizon + 1)) = twoArmGenerate…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-cc6e96d93725","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2251,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:1013"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmSampledPseudoRegret_eq_generated {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmSampledPseudoRegret (Env := Env) Delta (tailHorizon + 1)) = twoArmGeneratedExpectedPseudoRegret prior eta Delta environment tailHorizon","missing":[],"search":"integral_twoarmsampledpseudoregret_eq_generated banditrlproof.stochasticgradientbandit.integral_twoarmsampledpseudoregret_eq_generated theorem integral_twoarmsampledpseudoregret_eq_generated {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmsampledpseudoregret (env := env) delta (tailhorizon + 1)) = twoarmgeneratedexpectedpseudoregret prior eta delta environment tailhorizon theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne","label":"twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne","description":"theorem twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - me…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-ac1c5d499dcb","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2252,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:1057"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (tailHorizon : Nat) : twoArmGeneratedExpectedPseudoRegret prior eta Delta environment tailHorizon <= Real.log (1 + 4 * eta * Delta * ((tailHorizon + 1 : Nat) : Real)) / (2 * eta) + Delta / (2 * eta * (Delta - eta * sourceC eta))","missing":[],"search":"twoarmgeneratedexpectedpseudoregret_le_sourcetheoremone banditrlproof.stochasticgradientbandit.twoarmgeneratedexpectedpseudoregret_le_sourcetheoremone theorem twoarmgeneratedexpectedpseudoregret_le_sourcetheoremone {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 < eta) (hdelta : 0 < delta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (hmargin : eta * sourcec eta < delta) (tailhorizon : nat) : twoarmgeneratedexpectedpseudoregret prior eta delta environment tailhorizon <= real.log (1 + 4 * eta * delta * ((tailhorizon + 1 : nat) : real)) / (2 * eta) + delta / (2 * eta * (delta - eta * sourcec eta)) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_sourceTheoremOne","label":"integral_twoArmSampledPseudoRegret_le_sourceTheoremOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_sourceTheoremOne","description":"theorem integral_twoArmSampledPseudoRegret_le_sourceTheoremOne {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mea…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-234c90645f52","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2253,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:1106"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmSampledPseudoRegret_le_sourceTheoremOne {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmSampledPseudoRegret (Env := Env) Delta (tailHorizon + 1)) <= Real.log (1 + 4 * eta * Delta * ((tailHorizon + 1 : Nat) : Real)) / (2 * eta) + Delta / (2 * eta * (Delta - eta * sourceC eta))","missing":[],"search":"integral_twoarmsampledpseudoregret_le_sourcetheoremone banditrlproof.stochasticgradientbandit.integral_twoarmsampledpseudoregret_le_sourcetheoremone theorem integral_twoarmsampledpseudoregret_le_sourcetheoremone {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 < eta) (hdelta : 0 < delta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (hmargin : eta * sourcec eta < delta) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmsampledpseudoregret (env := env) delta (tailhorizon + 1)) <= real.log (1 + 4 * eta * delta * ((tailhorizon + 1 : nat) : real)) / (2 * eta) + delta / (2 * eta * (delta - eta * sourcec eta)) theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","label":"twoArmFixedIIDDirac_theoremOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","description":"Source-faithful Theorem 1 of Baudry--Johnson--Vary--Pike-Burke-- Rebeschini (NeurIPS 2025), for a fixed pair of IID arm laws. Lean horizon `tailHorizon + 1` is source `T`, including the uniform first action.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmtheoremone/index.html#decl-0c2e512d4027","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","order":2254,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmTheoremOne.lean:1131"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFixedIIDDirac_theoremOne (armLaw : Fin 2 -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (mean : Fin 2 -> Real) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, |reward| <= 1) (hmean : forall arm, integral (armLaw arm) id = mean arm) (eta Delta : Real) (heta : 0 < eta) (hDelta : 0 < Delta) (_hDelta_lt_one : Delta < 1) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure (Measure.dirac ()) eta (twoArmFixedIIDEnvironment armLaw hprob)) (twoArmSampledPseudoRegret (Env := Unit) Delta (tailHorizon + 1)) <= Real.log (1 + 4 * eta * Delta * ((tailHorizon + 1 : Nat) : Real)) / (2 * eta) + Delta / (2 * eta * (Delta - eta * sourceC eta))","missing":[],"search":"twoarmfixediiddirac_theoremone banditrlproof.stochasticgradientbandit.twoarmfixediiddirac_theoremone source-faithful theorem 1 of baudry--johnson--vary--pike-burke-- rebeschini (neurips 2025), for a fixed pair of iid arm laws. lean horizon `tailhorizon + 1` is source `t`, including the uniform first action. theorem compiled","shard":"modules/166646e93f03ff43.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero","label":"twoArmTrajectoryParameterZero","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero","description":"def twoArmTrajectoryParameterZero {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-41946dfdacda","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2255,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:39"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmTrajectoryParameterZero {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmtrajectoryparameterzero banditrlproof.stochasticgradientbandit.twoarmtrajectoryparameterzero def twoarmtrajectoryparameterzero {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardPotential","label":"twoArmForwardPotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardPotential","description":"def twoArmForwardPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-75cd0160bd6f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2256,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:46"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmForwardPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmforwardpotential banditrlproof.stochasticgradientbandit.twoarmforwardpotential def twoarmforwardpotential {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInversePotential","label":"twoArmInversePotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInversePotential","description":"def twoArmInversePotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-074115839675","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2257,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:52"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInversePotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarminversepotential banditrlproof.stochasticgradientbandit.twoarminversepotential def twoarminversepotential {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability","label":"twoArmSuccessProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability","description":"def twoArmSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-7f45b8c8ddae","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2258,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:58"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmsuccessprobability banditrlproof.stochasticgradientbandit.twoarmsuccessprobability def twoarmsuccessprobability {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFailureMass","label":"twoArmFailureMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFailureMass","description":"def twoArmFailureMass {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-bfb8bd2b7b08","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2259,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:66"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmFailureMass {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : Real","missing":[],"search":"twoarmfailuremass banditrlproof.stochasticgradientbandit.twoarmfailuremass def twoarmfailuremass {env : type v} [measurablespace env] (eta : real) (n : nat) (sample : env × ((k : nat) -> fin 2 × real)) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectoryParameterZero","label":"measurable_twoArmTrajectoryParameterZero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectoryParameterZero","description":"theorem measurable_twoArmTrajectoryParameterZero {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmTrajectoryParameterZero (Env := Env) eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-58b1753e3981","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2260,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:72"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmTrajectoryParameterZero {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmTrajectoryParameterZero (Env := Env) eta n)","missing":[],"search":"measurable_twoarmtrajectoryparameterzero banditrlproof.stochasticgradientbandit.measurable_twoarmtrajectoryparameterzero theorem measurable_twoarmtrajectoryparameterzero {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (twoarmtrajectoryparameterzero (env := env) eta n) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardPotential","label":"measurable_twoArmForwardPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardPotential","description":"theorem measurable_twoArmForwardPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmForwardPotential (Env := Env) eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-7fe208fbcf08","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2261,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:78"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmForwardPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmForwardPotential (Env := Env) eta n)","missing":[],"search":"measurable_twoarmforwardpotential banditrlproof.stochasticgradientbandit.measurable_twoarmforwardpotential theorem measurable_twoarmforwardpotential {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (twoarmforwardpotential (env := env) eta n) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInversePotential","label":"measurable_twoArmInversePotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInversePotential","description":"theorem measurable_twoArmInversePotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmInversePotential (Env := Env) eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-3b934af7ed6c","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2262,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:85"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInversePotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmInversePotential (Env := Env) eta n)","missing":[],"search":"measurable_twoarminversepotential banditrlproof.stochasticgradientbandit.measurable_twoarminversepotential theorem measurable_twoarminversepotential {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (twoarminversepotential (env := env) eta n) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessProbability","label":"measurable_twoArmSuccessProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessProbability","description":"theorem measurable_twoArmSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmSuccessProbability (Env := Env) eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-7c3718eb8ebd","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2263,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:92"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmSuccessProbability {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmSuccessProbability (Env := Env) eta n)","missing":[],"search":"measurable_twoarmsuccessprobability banditrlproof.stochasticgradientbandit.measurable_twoarmsuccessprobability theorem measurable_twoarmsuccessprobability {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (twoarmsuccessprobability (env := env) eta n) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmFailureMass","label":"measurable_twoArmFailureMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmFailureMass","description":"theorem measurable_twoArmFailureMass {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmFailureMass (Env := Env) eta n)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-75f27fdec931","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2264,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:104"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmFailureMass {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : Measurable (twoArmFailureMass (Env := Env) eta n)","missing":[],"search":"measurable_twoarmfailuremass banditrlproof.stochasticgradientbandit.measurable_twoarmfailuremass theorem measurable_twoarmfailuremass {env : type v} [measurablespace env] (eta : real) (n : nat) : measurable (twoarmfailuremass (env := env) eta n) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardPotential","label":"integrable_twoArmForwardPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardPotential","description":"theorem integrable_twoArmForwardPotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmForwardPotential (Env := Env) eta n) (twoArmTrajectoryMeasur…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-7c6b89e84f32","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2265,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:109"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmForwardPotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmForwardPotential (Env := Env) eta n) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmforwardpotential banditrlproof.stochasticgradientbandit.integrable_twoarmforwardpotential theorem integrable_twoarmforwardpotential {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : integrable (twoarmforwardpotential (env := env) eta n) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInversePotential","label":"integrable_twoArmInversePotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmInversePotential","description":"theorem integrable_twoArmInversePotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmInversePotential (Env := Env) eta n) (twoArmTrajectoryMeasur…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-46281205c4ce","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2266,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:139"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmInversePotential {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (twoArmInversePotential (Env := Env) eta n) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarminversepotential banditrlproof.stochasticgradientbandit.integrable_twoarminversepotential theorem integrable_twoarminversepotential {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : integrable (twoarminversepotential (env := env) eta n) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessProbability_sq","label":"integrable_twoArmSuccessProbability_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessProbability_sq","description":"theorem integrable_twoArmSuccessProbability_sq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : Integrable (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) (twoArmTrajectoryMeasure prior eta environment)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-dc42efb887fa","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2267,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:170"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmSuccessProbability_sq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : Integrable (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmsuccessprobability_sq banditrlproof.stochasticgradientbandit.integrable_twoarmsuccessprobability_sq theorem integrable_twoarmsuccessprobability_sq {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (n : nat) : integrable (fun sample => twoarmsuccessprobability (env := env) eta n sample ^ 2) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmFailureMass_sq","label":"integrable_twoArmFailureMass_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmFailureMass_sq","description":"theorem integrable_twoArmFailureMass_sq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : Integrable (fun sample => twoArmFailureMass (Env := Env) eta n sample ^ 2) (twoArmTrajectoryMeasure prior eta environment)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-849f8d422417","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2268,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:189"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmFailureMass_sq {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (n : Nat) : Integrable (fun sample => twoArmFailureMass (Env := Env) eta n sample ^ 2) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmfailuremass_sq banditrlproof.stochasticgradientbandit.integrable_twoarmfailuremass_sq theorem integrable_twoarmfailuremass_sq {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (n : nat) : integrable (fun sample => twoarmfailuremass (env := env) eta n sample ^ 2) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessor_eq_nextPotential","label":"twoArmForwardSuccessor_eq_nextPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessor_eq_nextPotential","description":"theorem twoArmForwardSuccessor_eq_nextPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample)) = twoArmForwardPotential (Env := Env) eta (n + 1)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-3b855ce5698b","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2269,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:209"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardSuccessor_eq_nextPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample)) = twoArmForwardPotential (Env := Env) eta (n + 1)","missing":[],"search":"twoarmforwardsuccessor_eq_nextpotential banditrlproof.stochasticgradientbandit.twoarmforwardsuccessor_eq_nextpotential theorem twoarmforwardsuccessor_eq_nextpotential {env : type v} [measurablespace env] (eta : real) (n : nat) : (fun sample : env × ((k : nat) -> fin 2 × real) => twoarmforwardsuccessorpotential eta (twoarmenvironmentprefix n sample).2 (twoarmnextpair n sample)) = twoarmforwardpotential (env := env) eta (n + 1) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessor_eq_nextPotential","label":"twoArmInverseSuccessor_eq_nextPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessor_eq_nextPotential","description":"theorem twoArmInverseSuccessor_eq_nextPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample)) = twoArmInversePotential (Env := Env) eta (n + 1)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-61e057371741","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2270,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:220"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseSuccessor_eq_nextPotential {Env : Type v} [MeasurableSpace Env] (eta : Real) (n : Nat) : (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseSuccessorPotential eta (twoArmEnvironmentPrefix n sample).2 (twoArmNextPair n sample)) = twoArmInversePotential (Env := Env) eta (n + 1)","missing":[],"search":"twoarminversesuccessor_eq_nextpotential banditrlproof.stochasticgradientbandit.twoarminversesuccessor_eq_nextpotential theorem twoarminversesuccessor_eq_nextpotential {env : type v} [measurablespace env] (eta : real) (n : nat) : (fun sample : env × ((k : nat) -> fin 2 × real) => twoarminversesuccessorpotential eta (twoarmenvironmentprefix n sample).2 (twoarmnextpair n sample)) = twoarminversepotential (env := env) eta (n + 1) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardRecurrenceBound","label":"integrable_twoArmForwardRecurrenceBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardRecurrenceBound","description":"theorem integrable_twoArmForwardRecurrenceBound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoA…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-347a8673c434","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2271,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:231"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmForwardRecurrenceBound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmForwardRecurrenceBound eta Delta (twoArmEnvironmentPrefix n sample).2) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarmforwardrecurrencebound banditrlproof.stochasticgradientbandit.integrable_twoarmforwardrecurrencebound theorem integrable_twoarmforwardrecurrencebound {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : integrable (fun sample : env × ((k : nat) -> fin 2 × real) => twoarmforwardrecurrencebound eta delta (twoarmenvironmentprefix n sample).2) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseRecurrenceBound","label":"integrable_twoArmInverseRecurrenceBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseRecurrenceBound","description":"theorem integrable_twoArmInverseRecurrenceBound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoA…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-27dc893b625f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2272,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:254"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integrable_twoArmInverseRecurrenceBound {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (n : Nat) : Integrable (fun sample : Env × ((k : Nat) -> Fin 2 × Real) => twoArmInverseRecurrenceBound eta Delta (twoArmEnvironmentPrefix n sample).2) (twoArmTrajectoryMeasure prior eta environment)","missing":[],"search":"integrable_twoarminverserecurrencebound banditrlproof.stochasticgradientbandit.integrable_twoarminverserecurrencebound theorem integrable_twoarminverserecurrencebound {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (n : nat) : integrable (fun sample : env × ((k : nat) -> fin 2 × real) => twoarminverserecurrencebound eta delta (twoarmenvironmentprefix n sample).2) (twoarmtrajectorymeasure prior eta environment) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardUnconditionalRecurrence","label":"twoArmForwardUnconditionalRecurrence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardUnconditionalRecurrence","description":"theorem twoArmForwardUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTr…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-36e73992e2ad","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2273,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:288"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta (n + 1)) <= integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta n) + 2 * integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) * (eta * Delta + eta ^ 2 * sourceC eta)","missing":[],"search":"twoarmforwardunconditionalrecurrence banditrlproof.stochasticgradientbandit.twoarmforwardunconditionalrecurrence theorem twoarmforwardunconditionalrecurrence {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta (n + 1)) <= integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta n) + 2 * integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmsuccessprobability (env := env) eta n sample ^ 2) * (eta * delta + eta ^ 2 * sourcec eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseUnconditionalRecurrence","label":"twoArmInverseUnconditionalRecurrence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseUnconditionalRecurrence","description":"theorem twoArmInverseUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTr…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-b812244f7848","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2274,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:339"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (n : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmInversePotential (Env := Env) eta (n + 1)) <= integral (twoArmTrajectoryMeasure prior eta environment) (twoArmInversePotential (Env := Env) eta n) - 2 * eta * integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample ^ 2) * (Delta - eta * sourceC eta)","missing":[],"search":"twoarminverseunconditionalrecurrence banditrlproof.stochasticgradientbandit.twoarminverseunconditionalrecurrence theorem twoarminverseunconditionalrecurrence {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (n : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarminversepotential (env := env) eta (n + 1)) <= integral (twoarmtrajectorymeasure prior eta environment) (twoarminversepotential (env := env) eta n) - 2 * eta * integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass (env := env) eta n sample ^ 2) * (delta - eta * sourcec eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmScalarForwardIterate","label":"twoArmScalarForwardIterate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmScalarForwardIterate","description":"theorem twoArmScalarForwardIterate (value increment : Nat -> Real) (hstep : forall n, value (n + 1) <= value n + increment n) : forall horizon, value horizon <= value 0 + (Finset.range horizon).sum increment","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-e9f1dc93a7e3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2275,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:392"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmScalarForwardIterate (value increment : Nat -> Real) (hstep : forall n, value (n + 1) <= value n + increment n) : forall horizon, value horizon <= value 0 + (Finset.range horizon).sum increment","missing":[],"search":"twoarmscalarforwarditerate banditrlproof.stochasticgradientbandit.twoarmscalarforwarditerate theorem twoarmscalarforwarditerate (value increment : nat -> real) (hstep : forall n, value (n + 1) <= value n + increment n) : forall horizon, value horizon <= value 0 + (finset.range horizon).sum increment theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmScalarInverseTelescope","label":"twoArmScalarInverseTelescope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmScalarInverseTelescope","description":"theorem twoArmScalarInverseTelescope (value failure : Nat -> Real) (coefficient : Real) (hstep : forall n, value (n + 1) <= value n - coefficient * failure n) : forall horizon, coefficient * (Finset.range horizon).sum failure <= value 0 - value horizon","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-b0425ed26289","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2276,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:409"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmScalarInverseTelescope (value failure : Nat -> Real) (coefficient : Real) (hstep : forall n, value (n + 1) <= value n - coefficient * failure n) : forall horizon, coefficient * (Finset.range horizon).sum failure <= value 0 - value horizon","missing":[],"search":"twoarmscalarinversetelescope banditrlproof.stochasticgradientbandit.twoarmscalarinversetelescope theorem twoarmscalarinversetelescope (value failure : nat -> real) (coefficient : real) (hstep : forall n, value (n + 1) <= value n - coefficient * failure n) : forall horizon, coefficient * (finset.range horizon).sum failure <= value 0 - value horizon theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration","label":"twoArmForwardFiniteIteration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration","description":"theorem twoArmForwardFiniteIteration {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (horizon : Nat) : integral (twoArmTraj…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-23fe8a8d74b3","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2277,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:431"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardFiniteIteration {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (horizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta horizon) <= integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta 0) + (Finset.range horizon).sum (fun n => 2 * integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) * (eta * Delta + eta ^ 2 * sourceC eta))","missing":[],"search":"twoarmforwardfiniteiteration banditrlproof.stochasticgradientbandit.twoarmforwardfiniteiteration theorem twoarmforwardfiniteiteration {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (horizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta horizon) <= integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta 0) + (finset.range horizon).sum (fun n => 2 * integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmsuccessprobability (env := env) eta n sample ^ 2) * (eta * delta + eta ^ 2 * sourcec eta)) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqTelescope","label":"twoArmInverseFailureMassSqTelescope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqTelescope","description":"theorem twoArmInverseFailureMassSqTelescope {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (horizon : Nat) : (2 * eta * (D…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-7d0bd275a5d1","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2278,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:460"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseFailureMassSqTelescope {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (horizon : Nat) : (2 * eta * (Delta - eta * sourceC eta)) * (Finset.range horizon).sum (fun n => integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample ^ 2)) <= integral (twoArmTrajectoryMeasure prior eta environment) (twoArmInversePotential (Env := Env) eta 0) - integral (twoArmTrajectoryMeasure prior eta environment) (twoArmInversePotential (Env := Env) eta horizon)","missing":[],"search":"twoarminversefailuremasssqtelescope banditrlproof.stochasticgradientbandit.twoarminversefailuremasssqtelescope theorem twoarminversefailuremasssqtelescope {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (horizon : nat) : (2 * eta * (delta - eta * sourcec eta)) * (finset.range horizon).sum (fun n => integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass (env := env) eta n sample ^ 2)) <= integral (twoarmtrajectorymeasure prior eta environment) (twoarminversepotential (env := env) eta 0) - integral (twoarmtrajectorymeasure prior eta environment) (twoarminversepotential (env := env) eta horizon) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqSum_le_initial_div","label":"twoArmInverseFailureMassSqSum_le_initial_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqSum_le_initial_div","description":"theorem twoArmInverseFailureMassSqSum_le_initial_div {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 < eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * source…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-2761dd76b53d","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2279,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:490"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseFailureMassSqSum_le_initial_div {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsFiniteMeasure prior] (eta Delta : Real) (heta : 0 < eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (horizon : Nat) : (Finset.range horizon).sum (fun n => integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample ^ 2)) <= integral (twoArmTrajectoryMeasure prior eta environment) (twoArmInversePotential (Env := Env) eta 0) / (2 * eta * (Delta - eta * sourceC eta))","missing":[],"search":"twoarminversefailuremasssqsum_le_initial_div banditrlproof.stochasticgradientbandit.twoarminversefailuremasssqsum_le_initial_div theorem twoarminversefailuremasssqsum_le_initial_div {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isfinitemeasure prior] (eta delta : real) (heta : 0 < eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (hmargin : eta * sourcec eta < delta) (horizon : nat) : (finset.range horizon).sum (fun n => integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmfailuremass (env := env) eta n sample ^ 2)) <= integral (twoarmtrajectorymeasure prior eta environment) (twoarminversepotential (env := env) eta 0) / (2 * eta * (delta - eta * sourcec eta)) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialForwardPotential","label":"twoArmInitialForwardPotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInitialForwardPotential","description":"def twoArmInitialForwardPotential (eta : Real) (pair : Fin 2 × Real) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-8a3a4565ac5f","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2280,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:522"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInitialForwardPotential (eta : Real) (pair : Fin 2 × Real) : Real","missing":[],"search":"twoarminitialforwardpotential banditrlproof.stochasticgradientbandit.twoarminitialforwardpotential def twoarminitialforwardpotential (eta : real) (pair : fin 2 × real) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialInversePotential","label":"twoArmInitialInversePotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInitialInversePotential","description":"def twoArmInitialInversePotential (eta : Real) (pair : Fin 2 × Real) : Real","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-5b3ef1e1594e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2281,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:528"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def twoArmInitialInversePotential (eta : Real) (pair : Fin 2 × Real) : Real","missing":[],"search":"twoarminitialinversepotential banditrlproof.stochasticgradientbandit.twoarminitialinversepotential def twoarminitialinversepotential (eta : real) (pair : fin 2 × real) : real definition compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialForwardPotential","label":"measurable_twoArmInitialForwardPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialForwardPotential","description":"theorem measurable_twoArmInitialForwardPotential (eta : Real) : Measurable (twoArmInitialForwardPotential eta)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-6c7cc57e1c62","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2282,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:534"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInitialForwardPotential (eta : Real) : Measurable (twoArmInitialForwardPotential eta)","missing":[],"search":"measurable_twoarminitialforwardpotential banditrlproof.stochasticgradientbandit.measurable_twoarminitialforwardpotential theorem measurable_twoarminitialforwardpotential (eta : real) : measurable (twoarminitialforwardpotential eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialInversePotential","label":"measurable_twoArmInitialInversePotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialInversePotential","description":"theorem measurable_twoArmInitialInversePotential (eta : Real) : Measurable (twoArmInitialInversePotential eta)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-3f1a59a13f39","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2283,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:548"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurable_twoArmInitialInversePotential (eta : Real) : Measurable (twoArmInitialInversePotential eta)","missing":[],"search":"measurable_twoarminitialinversepotential banditrlproof.stochasticgradientbandit.measurable_twoarminitialinversepotential theorem measurable_twoarminitialinversepotential (eta : real) : measurable (twoarminitialinversepotential eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardPotential_zero_eq_initial","label":"twoArmForwardPotential_zero_eq_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardPotential_zero_eq_initial","description":"theorem twoArmForwardPotential_zero_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmForwardPotential eta 0 sample = twoArmInitialForwardPotential eta (sample.2 0)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-3c447bcc4d86","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2284,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:562"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardPotential_zero_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmForwardPotential eta 0 sample = twoArmInitialForwardPotential eta (sample.2 0)","missing":[],"search":"twoarmforwardpotential_zero_eq_initial banditrlproof.stochasticgradientbandit.twoarmforwardpotential_zero_eq_initial theorem twoarmforwardpotential_zero_eq_initial {env : type v} [measurablespace env] (eta : real) (sample : env × ((k : nat) -> fin 2 × real)) : twoarmforwardpotential eta 0 sample = twoarminitialforwardpotential eta (sample.2 0) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInversePotential_zero_eq_initial","label":"twoArmInversePotential_zero_eq_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInversePotential_zero_eq_initial","description":"theorem twoArmInversePotential_zero_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmInversePotential eta 0 sample = twoArmInitialInversePotential eta (sample.2 0)","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-05ea8ba1f368","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2285,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:584"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInversePotential_zero_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (sample : Env × ((k : Nat) -> Fin 2 × Real)) : twoArmInversePotential eta 0 sample = twoArmInitialInversePotential eta (sample.2 0)","missing":[],"search":"twoarminversepotential_zero_eq_initial banditrlproof.stochasticgradientbandit.twoarminversepotential_zero_eq_initial theorem twoarminversepotential_zero_eq_initial {env : type v} [measurablespace env] (eta : real) (sample : env × ((k : nat) -> fin 2 × real)) : twoarminversepotential eta 0 sample = twoarminitialinversepotential eta (sample.2 0) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_zero_kernel_eq_initial","label":"integral_twoArmForwardPotential_zero_kernel_eq_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_zero_kernel_eq_initial","description":"theorem integral_twoArmForwardPotential_zero_kernel_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) : integral (trajectoryKernel (fun _ : Fin 2 => 0) eta environment env) (fun trajectory => twoArmForwardPotential eta 0 (env, trajectory)) = integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-8474d61a6ae7","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2286,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:606"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmForwardPotential_zero_kernel_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) : integral (trajectoryKernel (fun _ : Fin 2 => 0) eta environment env) (fun trajectory => twoArmForwardPotential eta 0 (env, trajectory)) = integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) (twoArmInitialForwardPotential eta)","missing":[],"search":"integral_twoarmforwardpotential_zero_kernel_eq_initial banditrlproof.stochasticgradientbandit.integral_twoarmforwardpotential_zero_kernel_eq_initial theorem integral_twoarmforwardpotential_zero_kernel_eq_initial {env : type v} [measurablespace env] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (env : env) : integral (trajectorykernel (fun _ : fin 2 => 0) eta environment env) (fun trajectory => twoarmforwardpotential eta 0 (env, trajectory)) = integral (thompson.measurableenvironmentinitialpairkernel (historyalgorithm (fun _ : fin 2 => 0) eta) environment env) (twoarminitialforwardpotential eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInversePotential_zero_kernel_eq_initial","label":"integral_twoArmInversePotential_zero_kernel_eq_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.integral_twoArmInversePotential_zero_kernel_eq_initial","description":"theorem integral_twoArmInversePotential_zero_kernel_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) : integral (trajectoryKernel (fun _ : Fin 2 => 0) eta environment env) (fun trajectory => twoArmInversePotential eta 0 (env, trajectory)) = integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-edc10e800d17","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2287,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:642"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem integral_twoArmInversePotential_zero_kernel_eq_initial {Env : Type v} [MeasurableSpace Env] (eta : Real) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (env : Env) : integral (trajectoryKernel (fun _ : Fin 2 => 0) eta environment env) (fun trajectory => twoArmInversePotential eta 0 (env, trajectory)) = integral (Thompson.measurableEnvironmentInitialPairKernel (historyAlgorithm (fun _ : Fin 2 => 0) eta) environment env) (twoArmInitialInversePotential eta)","missing":[],"search":"integral_twoarminversepotential_zero_kernel_eq_initial banditrlproof.stochasticgradientbandit.integral_twoarminversepotential_zero_kernel_eq_initial theorem integral_twoarminversepotential_zero_kernel_eq_initial {env : type v} [measurablespace env] (eta : real) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (env : env) : integral (trajectorykernel (fun _ : fin 2 => 0) eta environment env) (fun trajectory => twoarminversepotential eta 0 (env, trajectory)) = integral (thompson.measurableenvironmentinitialpairkernel (historyalgorithm (fun _ : fin 2 => 0) eta) environment env) (twoarminitialinversepotential eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardInitialUnconditionalRecurrence","label":"twoArmForwardInitialUnconditionalRecurrence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardInitialUnconditionalRecurrence","description":"theorem twoArmForwardInitialUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (twoArm…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-8cbd167f129e","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2288,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:678"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardInitialUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta 0) <= 1 + (eta * Delta + eta ^ 2 * sourceC eta) / 2","missing":[],"search":"twoarmforwardinitialunconditionalrecurrence banditrlproof.stochasticgradientbandit.twoarmforwardinitialunconditionalrecurrence theorem twoarmforwardinitialunconditionalrecurrence {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta 0) <= 1 + (eta * delta + eta ^ 2 * sourcec eta) / 2 theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseInitialUnconditionalRecurrence","label":"twoArmInverseInitialUnconditionalRecurrence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmInverseInitialUnconditionalRecurrence","description":"theorem twoArmInverseInitialUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (twoArm…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-8e922eb7a144","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2289,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:733"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmInverseInitialUnconditionalRecurrence {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmInversePotential (Env := Env) eta 0) <= 1 - eta / 2 * (Delta - eta * sourceC eta)","missing":[],"search":"twoarminverseinitialunconditionalrecurrence banditrlproof.stochasticgradientbandit.twoarminverseinitialunconditionalrecurrence theorem twoarminverseinitialunconditionalrecurrence {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) : integral (twoarmtrajectorymeasure prior eta environment) (twoarminversepotential (env := env) eta 0) <= 1 - eta / 2 * (delta - eta * sourcec eta) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration_from_source_initial","label":"twoArmForwardFiniteIteration_from_source_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration_from_source_initial","description":"theorem twoArmForwardFiniteIteration_from_source_initial {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (tailHorizon…","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-fa241b4f8fbe","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2290,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:788"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmForwardFiniteIteration_from_source_initial {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 <= eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (tailHorizon : Nat) : integral (twoArmTrajectoryMeasure prior eta environment) (twoArmForwardPotential (Env := Env) eta tailHorizon) <= 1 + (eta * Delta + eta ^ 2 * sourceC eta) / 2 + (Finset.range tailHorizon).sum (fun n => 2 * integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmSuccessProbability (Env := Env) eta n sample ^ 2) * (eta * Delta + eta ^ 2 * sourceC eta))","missing":[],"search":"twoarmforwardfiniteiteration_from_source_initial banditrlproof.stochasticgradientbandit.twoarmforwardfiniteiteration_from_source_initial theorem twoarmforwardfiniteiteration_from_source_initial {env : type v} [measurablespace env] [standardborelspace env] (prior : measure env) [isprobabilitymeasure prior] (eta delta : real) (heta : 0 <= eta) (environment : thompson.measurablehistoryenvironment env (fin 2) real) (mean : fin 2 -> real) (contract : twoarmboundedfixedmeanenvironmentcontract environment mean) (hgap : mean 0 - mean 1 = delta) (tailhorizon : nat) : integral (twoarmtrajectorymeasure prior eta environment) (twoarmforwardpotential (env := env) eta tailhorizon) <= 1 + (eta * delta + eta ^ 2 * sourcec eta) / 2 + (finset.range tailhorizon).sum (fun n => 2 * integral (twoarmtrajectorymeasure prior eta environment) (fun sample => twoarmsuccessprobability (env := env) eta n sample ^ 2) * (eta * delta + eta ^ 2 * sourcec eta)) theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","label":"twoArmFullFailureMassSqSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","description":"The source round `t = 1` contributes exactly `(1 - 1/2)^2 = 1/4`; the `tailHorizon` summands are source rounds `t = 2, ..., tailHorizon + 1`.","url":"../modules/banditrlproof-algorithms-stochasticgradientbandittwoarmunconditionalrecurrence/index.html#decl-7773b5ffa203","parent":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","order":2291,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence"],["Source","BanditRLProof/Algorithms/StochasticGradientBanditTwoArmUnconditionalRecurrence.lean:812"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem twoArmFullFailureMassSqSum_le {Env : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (eta Delta : Real) (heta : 0 < eta) (environment : Thompson.MeasurableHistoryEnvironment Env (Fin 2) Real) (mean : Fin 2 -> Real) (contract : TwoArmBoundedFixedMeanEnvironmentContract environment mean) (hgap : mean 0 - mean 1 = Delta) (hmargin : eta * sourceC eta < Delta) (tailHorizon : Nat) : (1 : Real) / 4 + (Finset.range tailHorizon).sum (fun n => integral (twoArmTrajectoryMeasure prior eta environment) (fun sample => twoArmFailureMass (Env := Env) eta n sample ^ 2)) <= 1 / (2 * eta * (Delta - eta * sourceC eta))","missing":[],"search":"twoarmfullfailuremasssqsum_le banditrlproof.stochasticgradientbandit.twoarmfullfailuremasssqsum_le the source round `t = 1` contributes exactly `(1 - 1/2)^2 = 1/4`; the `tailhorizon` summands are source rounds `t = 2, ..., tailhorizon + 1`. theorem compiled","shard":"modules/4b10006a2c474128.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PriorSketch","label":"PriorSketch","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.PriorSketch","description":"A lightweight descriptor for a Bayesian bandit parameter space.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-42760d09ca07","parent":"module:BanditRLProof.Algorithms.Thompson","order":2292,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:17"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure PriorSketch where","missing":[],"search":"priorsketch banditrlproof.thompson.priorsketch a lightweight descriptor for a bayesian bandit parameter space. structure compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.obligationNames","label":"obligationNames","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.obligationNames","description":"The proof-DAG leaves usually needed for Thompson sampling regret.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-8c6b3a17884c","parent":"module:BanditRLProof.Algorithms.Thompson","order":2293,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:24"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def obligationNames : List String","missing":[],"search":"obligationnames banditrlproof.thompson.obligationnames the proof-dag leaves usually needed for thompson sampling regret. definition compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger","label":"PosteriorActionIdentityLedger","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.PosteriorActionIdentityLedger","description":"Source contract for the Thompson probability-matching identity. The ledger records the exact law expected from a Thompson action sampler: its action kernel at a history agrees on every measurable action event with the posterior distribution pushed forward by the environment-to-best-action map. It is a contract surface, not a Bayes-rule proof or posterior-sampler construction.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-8794e6f5dd66","parent":"module:BanditRLProof.Algorithms.Thompson","order":2294,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:49"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure PosteriorActionIdentityLedger (History : Type u) (Env : Type v) (Action : Type w) [MeasurableSpace History] [MeasurableSpace Env] [MeasurableSpace Action] where","missing":[],"search":"posterioractionidentityledger banditrlproof.thompson.posterioractionidentityledger source contract for the thompson probability-matching identity. the ledger records the exact law expected from a thompson action sampler: its action kernel at a history agrees on every measurable action event with the posterior distribution pushed forward by the environment-to-best-action map. it is a contract surface, not a bayes-rule proof or posterior-sampler construction. structure compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.bestAction_measurable_of_countable_env","label":"bestAction_measurable_of_countable_env","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.bestAction_measurable_of_countable_env","description":"Any best-action selector out of a countable singleton-measurable environment space is measurable. This is the regularity wrapper needed by finite or countable posterior model spaces before constructing a Thompson posterior-action identity ledger.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-b1e22bb41eb9","parent":"module:BanditRLProof.Algorithms.Thompson","order":2295,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:71"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem bestAction_measurable_of_countable_env {Env : Type v} {Action : Type w} [MeasurableSpace Env] [MeasurableSingletonClass Env] [Countable Env] [MeasurableSpace Action] (bestAction : Env -> Action) : Measurable bestAction","missing":[],"search":"bestaction_measurable_of_countable_env banditrlproof.thompson.bestaction_measurable_of_countable_env any best-action selector out of a countable singleton-measurable environment space is measurable. this is the regularity wrapper needed by finite or countable posterior model spaces before constructing a thompson posterior-action identity ledger. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.ofCountableEnv","label":"ofCountableEnv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.PosteriorActionIdentityLedger.ofCountableEnv","description":"Build a posterior-action identity ledger over a countable environment space without separately supplying best-action measurability.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-7f3170a162cf","parent":"module:BanditRLProof.Algorithms.Thompson","order":2296,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:93"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def ofCountableEnv [MeasurableSingletonClass Env] [Countable Env] (posterior : PosteriorKernel.MarkovPosteriorKernel History Env) (actionKernel : ProbabilityTheory.Kernel History Action) (hactionKernel : ProbabilityTheory.IsMarkovKernel actionKernel) (bestAction : Env -> Action) (hmatch : forall (history : History) {event : Set Action}, MeasurableSet event -> actionKernel history event = Measure.map bestAction (posterior.kernel history) event) : PosteriorActionIdentityLedger History Env Action where","missing":[],"search":"ofcountableenv banditrlproof.thompson.posterioractionidentityledger.ofcountableenv build a posterior-action identity ledger over a countable environment space without separately supplying best-action measurability. definition compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.ofPosteriorMap","label":"ofPosteriorMap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.PosteriorActionIdentityLedger.ofPosteriorMap","description":"Build the Thompson action ledger directly by mapping a posterior kernel through a measurable best-action selector.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-b0bd3217e24c","parent":"module:BanditRLProof.Algorithms.Thompson","order":2297,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:116"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def ofPosteriorMap (posterior : PosteriorKernel.MarkovPosteriorKernel History Env) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : PosteriorActionIdentityLedger History Env Action where","missing":[],"search":"ofposteriormap banditrlproof.thompson.posterioractionidentityledger.ofposteriormap build the thompson action ledger directly by mapping a posterior kernel through a measurable best-action selector. definition compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_eq_posterior_map","label":"actionKernel_eq_posterior_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_eq_posterior_map","description":"The event-level ledger identity is equality of the two Markov kernels.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-789d4a8088f2","parent":"module:BanditRLProof.Algorithms.Thompson","order":2298,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:131"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem actionKernel_eq_posterior_map (ledger : PosteriorActionIdentityLedger History Env Action) : ledger.actionKernel = ledger.posterior.kernel.map ledger.bestAction","missing":[],"search":"actionkernel_eq_posterior_map banditrlproof.thompson.posterioractionidentityledger.actionkernel_eq_posterior_map the event-level ledger identity is equality of the two markov kernels. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_apply_eq_posteriorBest_map","label":"actionKernel_apply_eq_posteriorBest_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_apply_eq_posteriorBest_map","description":"Event-level Thompson probability matching from the packaged ledger.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-f92b18b65ede","parent":"module:BanditRLProof.Algorithms.Thompson","order":2299,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:141"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem actionKernel_apply_eq_posteriorBest_map (ledger : PosteriorActionIdentityLedger History Env Action) (history : History) {event : Set Action} (hevent : MeasurableSet event) : ledger.actionKernel history event = Measure.map ledger.bestAction (ledger.posterior.kernel history) event","missing":[],"search":"actionkernel_apply_eq_posteriorbest_map banditrlproof.thompson.posterioractionidentityledger.actionkernel_apply_eq_posteriorbest_map event-level thompson probability matching from the packaged ledger. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_apply_singleton_eq_posteriorBest_preimage","label":"actionKernel_apply_singleton_eq_posteriorBest_preimage","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_apply_singleton_eq_posteriorBest_preimage","description":"Singleton form of the posterior action identity. For discrete or singleton-measurable action spaces, the Thompson probability of choosing `action` is the posterior probability that `action` is the best action.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-5891bf1cea02","parent":"module:BanditRLProof.Algorithms.Thompson","order":2300,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:155"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem actionKernel_apply_singleton_eq_posteriorBest_preimage [MeasurableSingletonClass Action] (ledger : PosteriorActionIdentityLedger History Env Action) (history : History) (action : Action) : ledger.actionKernel history ({action} : Set Action) = ledger.posterior.kernel history {env : Env | ledger.bestAction env = action}","missing":[],"search":"actionkernel_apply_singleton_eq_posteriorbest_preimage banditrlproof.thompson.posterioractionidentityledger.actionkernel_apply_singleton_eq_posteriorbest_preimage singleton form of the posterior action identity. for discrete or singleton-measurable action spaces, the thompson probability of choosing `action` is the posterior probability that `action` is the best action. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.BayesianPosteriorActionSource","label":"BayesianPosteriorActionSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.BayesianPosteriorActionSource","description":"Source fields needed for the Thompson posterior-action conditional-law theorem. The first law says the process samples its next action from the ledger action kernel. The second identifies the ledger posterior with the conditional law of the latent environment given the observed history. These are the two law surfaces used by the pinned LML proof.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-ab0b1efc7364","parent":"module:BanditRLProof.Algorithms.Thompson","order":2301,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:177"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure BayesianPosteriorActionSource {Omega : Type u} (History : Type v) (Env : Type w) (Action : Type*) [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (nextAction : Omega -> Action) (ledger : PosteriorActionIdentityLedger History Env Action) : Prop where","missing":[],"search":"bayesianposterioractionsource banditrlproof.thompson.bayesianposterioractionsource source fields needed for the thompson posterior-action conditional-law theorem. the first law says the process samples its next action from the ledger action kernel. the second identifies the ledger posterior with the conditional law of the latent environment given the observed history. these are the two law surfaces used by the pinned lml proof. structure compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_bayesianPosteriorActionSource","label":"condDistrib_action_ae_eq_bestAction_of_bayesianPosteriorActionSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_bayesianPosteriorActionSource","description":"Thompson probability matching in Mathlib `condDistrib` form. Mapping the posterior-kernel equality through `bestAction` identifies the ledger action kernel with the mapped environment conditional law. Mathlib's `condDistrib_comp` then identifies that map with the conditional law of the random best action itself.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-32db512fc70a","parent":"module:BanditRLProof.Algorithms.Thompson","order":2302,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:204"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_action_ae_eq_bestAction_of_bayesianPosteriorActionSource {Omega : Type u} {History : Type v} {Env : Type w} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (nextAction : Omega -> Action) (ledger : PosteriorActionIdentityLedger History Env Action) (source : BayesianPosteriorActionSource History Env Action mu env history nextAction ledger) : ProbabilityTheory.condDistrib nextAction history mu =ᵐ[mu.map history] ProbabilityTheory.condDistrib (ledger.bestAction ∘ env) history mu","missing":[],"search":"conddistrib_action_ae_eq_bestaction_of_bayesianposterioractionsource banditrlproof.thompson.conddistrib_action_ae_eq_bestaction_of_bayesianposterioractionsource thompson probability matching in mathlib `conddistrib` form. mapping the posterior-kernel equality through `bestaction` identifies the ledger action kernel with the mapped environment conditional law. mathlib's `conddistrib_comp` then identifies that map with the conditional law of the random best action itself. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_posteriorMap","label":"condDistrib_action_ae_eq_bestAction_of_posteriorMap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_posteriorMap","description":"Direct posterior-map form of Thompson probability matching. This is the local Mathlib-facing counterpart of pinned LML `Bandits.TS.hasCondDistrib_action`: it avoids a local `HasCondDistrib` wrapper and states the resulting regular conditional-kernel equality directly.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-1c1ff119df73","parent":"module:BanditRLProof.Algorithms.Thompson","order":2303,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:241"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_action_ae_eq_bestAction_of_posteriorMap {Omega : Type u} {History : Type v} {Env : Type w} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (nextAction : Omega -> Action) (posterior : PosteriorKernel.MarkovPosteriorKernel History Env) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (henv : Measurable env) (hhistory : Measurable history) (hnextAction : Measurable nextAction) (haction : ProbabilityTheory.condDistrib nextAction history mu =ᵐ[mu.map history] posterior.kernel.map bestAction) (hposterior : posterior.kernel =ᵐ[mu.map history] ProbabilityTheory.condDistrib env history mu) : ProbabilityTheory.condDistr…","missing":[],"search":"conddistrib_action_ae_eq_bestaction_of_posteriormap banditrlproof.thompson.conddistrib_action_ae_eq_bestaction_of_posteriormap direct posterior-map form of thompson probability matching. this is the local mathlib-facing counterpart of pinned lml `bandits.ts.hasconddistrib_action`: it avoids a local `hasconddistrib` wrapper and states the resulting regular conditional-kernel equality directly. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_bayesianPairMap","label":"condDistrib_action_ae_eq_bestAction_of_bayesianPairMap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_bayesianPairMap","description":"Thompson probability matching from a Bayesian environment/history pair law. Unlike `condDistrib_action_ae_eq_bestAction_of_posteriorMap`, this theorem does not assume the posterior/environment conditional-law equality. It constructs that equality from the source pair law using Mathlib's canonical posterior.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-fa0fea92c585","parent":"module:BanditRLProof.Algorithms.Thompson","order":2304,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:283"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_action_ae_eq_bestAction_of_bayesianPairMap {Omega : Type u} {History : Type v} {Env : Type w} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (nextAction : Omega -> Action) (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (henv : Measurable env) (hhistory : Measurable history) (hnextAction : Measurable nextAction) (hpair : mu.map (fun omega => (env omega, history omega)) = PosteriorKernel.canonicalJointMeasure prior likelihood) (haction : ProbabilityTheory.condDis…","missing":[],"search":"conddistrib_action_ae_eq_bestaction_of_bayesianpairmap banditrlproof.thompson.conddistrib_action_ae_eq_bestaction_of_bayesianpairmap thompson probability matching from a bayesian environment/history pair law. unlike `conddistrib_action_ae_eq_bestaction_of_posteriormap`, this theorem does not assume the posterior/environment conditional-law equality. it constructs that equality from the source pair law using mathlib's canonical posterior. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_canonicalPriorLikelihood","label":"condDistrib_action_ae_eq_bestAction_of_canonicalPriorLikelihood","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_canonicalPriorLikelihood","description":"Canonical-product specialization of the Thompson posterior-action law. The source space is `Env × History` with law `prior ⊗ₘ likelihood`, so the posterior conditional-law premise is discharged entirely by the canonical Bayesian construction. The remaining law premise is exactly the Thompson action sampler's conditional law.","url":"../modules/banditrlproof-algorithms-thompson/index.html#decl-f0d5a86ff012","parent":"module:BanditRLProof.Algorithms.Thompson","order":2305,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.Thompson"],["Source","BanditRLProof/Algorithms/Thompson.lean:321"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_action_ae_eq_bestAction_of_canonicalPriorLikelihood {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (nextAction : Env × History -> Action) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (hnextAction : Measurable nextAction) (haction : ProbabilityTheory.condDistrib nextAction Prod.snd (PosteriorKernel.canonicalJointMeasure prior likelihood) =ᵐ[ (PosteriorKernel.canonicalJointMeasure prior likelihood).map Prod.snd] (PosteriorKernel.canonicalPosterior prior likelihood).kernel.map bestAction) : ProbabilityTheory.condDistrib nextAction Prod.snd (…","missing":[],"search":"conddistrib_action_ae_eq_bestaction_of_canonicalpriorlikelihood banditrlproof.thompson.conddistrib_action_ae_eq_bestaction_of_canonicalpriorlikelihood canonical-product specialization of the thompson posterior-action law. the source space is `env × history` with law `prior ⊗ₘ likelihood`, so the posterior conditional-law premise is discharged entirely by the canonical bayesian construction. the remaining law premise is exactly the thompson action sampler's conditional law. theorem compiled","shard":"modules/b2be9e5cabd45d88.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.compProd_withDensity_left","label":"compProd_withDensity_left","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.compProd_withDensity_left","description":"Weighting the base measure of a composition product is the same as weighting the product by the density pulled back through the first projection.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-24516159cb42","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2306,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:26"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem compProd_withDensity_left {History : Type u} {Env : Type v} [MeasurableSpace History] [MeasurableSpace Env] (historyLaw : Measure History) [SFinite historyLaw] (posterior : ProbabilityTheory.Kernel History Env) [ProbabilityTheory.IsSFiniteKernel posterior] (density : History -> ENNReal) (hdensity : Measurable density) : historyLaw.withDensity density ⊗ₘ posterior = (historyLaw ⊗ₘ posterior).withDensity (density ∘ Prod.fst)","missing":[],"search":"compprod_withdensity_left banditrlproof.thompson.compprod_withdensity_left weighting the base measure of a composition product is the same as weighting the product by the density pulled back through the first projection. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.comp_withDensity_history","label":"comp_withDensity_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.comp_withDensity_history","description":"Composing a kernel weighted by a density independent of its input is the same as weighting the composed output measure.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-165e4a235615","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2307,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:66"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem comp_withDensity_history {Env : Type u} {History : Type v} [MeasurableSpace Env] [MeasurableSpace History] (envLaw : Measure Env) (historyKernel : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsSFiniteKernel historyKernel] (density : History -> ENNReal) (hdensity : Measurable density) : (historyKernel.withDensity (fun _ history => density history)) ∘ₘ envLaw = (historyKernel ∘ₘ envLaw).withDensity density","missing":[],"search":"comp_withdensity_history banditrlproof.thompson.comp_withdensity_history composing a kernel weighted by a density independent of its input is the same as weighting the composed output measure. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.map_swap_withDensity_snd","label":"map_swap_withDensity_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.map_swap_withDensity_snd","description":"Swapping a joint law weighted by its second coordinate moves the density to the first coordinate of the swapped law.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-5f340aba8a5d","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2308,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:93"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem map_swap_withDensity_snd {Env : Type u} {History : Type v} [MeasurableSpace Env] [MeasurableSpace History] (joint : Measure (Env × History)) (density : History -> ENNReal) (hdensity : Measurable density) : (joint.withDensity (density ∘ Prod.snd)).map Prod.swap = (joint.map Prod.swap).withDensity (density ∘ Prod.fst)","missing":[],"search":"map_swap_withdensity_snd banditrlproof.thompson.map_swap_withdensity_snd swapping a joint law weighted by its second coordinate moves the density to the first coordinate of the swapped law. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.compProd_eq_compProd_withDensity_snd_of_ae_eq","label":"compProd_eq_compProd_withDensity_snd_of_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.compProd_eq_compProd_withDensity_snd_of_ae_eq","description":"Composition products transport an a.e. output-only kernel density directly, without requiring an `IsSFiniteKernel` instance for the weighted kernel.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-8638540e78ce","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2309,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:112"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem compProd_eq_compProd_withDensity_snd_of_ae_eq {Env : Type u} {History : Type v} [MeasurableSpace Env] [MeasurableSpace History] (envLaw : Measure Env) [SFinite envLaw] (actualHistoryKernel referenceHistoryKernel : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsSFiniteKernel actualHistoryKernel] [ProbabilityTheory.IsSFiniteKernel referenceHistoryKernel] (density : History -> ENNReal) (hdensity : Measurable density) (hkernel : actualHistoryKernel =ᵐ[envLaw] referenceHistoryKernel.withDensity (fun _ history => density history)) : envLaw ⊗ₘ actualHistoryKernel = (envLaw ⊗ₘ referenceHistoryKernel).withDensity (density ∘ Prod.snd)","missing":[],"search":"compprod_eq_compprod_withdensity_snd_of_ae_eq banditrlproof.thompson.compprod_eq_compprod_withdensity_snd_of_ae_eq composition products transport an a.e. output-only kernel density directly, without requiring an `issfinitekernel` instance for the weighted kernel. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.AlgorithmDensityPosteriorSource","label":"AlgorithmDensityPosteriorSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.AlgorithmDensityPosteriorSource","description":"The exact law interface produced by an algorithm-density/change-of-algorithm argument. Both the history marginal and the history/environment joint law are weighted by the same measurable density depending only on history.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-b1f710e7be55","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2310,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:164"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure AlgorithmDensityPosteriorSource {Omega : Type u} {OmegaRef : Type v} {History : Type w} {Env : Type x} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] (mu : Measure Omega) (env : Omega -> Env) (history : Omega -> History) (referenceMu : Measure OmegaRef) (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) where","missing":[],"search":"algorithmdensityposteriorsource banditrlproof.thompson.algorithmdensityposteriorsource the exact law interface produced by an algorithm-density/change-of-algorithm argument. both the history marginal and the history/environment joint law are weighted by the same measurable density depending only on history. structure compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.algorithmDensityPosteriorSource_of_condDistrib_history_withDensity","label":"algorithmDensityPosteriorSource_of_condDistrib_history_withDensity","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.algorithmDensityPosteriorSource_of_condDistrib_history_withDensity","description":"Construct the two algorithm-density pushforward laws from a closer-to-process interface: the actual and reference environment marginals agree, and the actual conditional history kernel is the reference conditional history kernel weighted by one history-only density.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-c296ed5364f5","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2311,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:190"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def algorithmDensityPosteriorSource_of_condDistrib_history_withDensity {Omega : Type u} {OmegaRef : Type v} {History : Type w} {Env : Type x} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace History] [StandardBorelSpace History] [Nonempty History] [MeasurableSpace Env] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (hhistory : Measurable history) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) (density : History -> ENNReal) (hdensity : Measurable density) (henvLaw : mu.map env = referenceMu.map referenceEnv) (hcond : ProbabilityTheory.condDistrib history env mu =ᵐ[mu.map env] (ProbabilityTheory.condDistrib re…","missing":[],"search":"algorithmdensityposteriorsource_of_conddistrib_history_withdensity banditrlproof.thompson.algorithmdensityposteriorsource_of_conddistrib_history_withdensity construct the two algorithm-density pushforward laws from a closer-to-process interface: the actual and reference environment marginals agree, and the actual conditional history kernel is the reference conditional history kernel weighted by one history-only density. definition compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosterior_ae_eq_condDistrib_of_algorithmDensitySource","label":"referencePosterior_ae_eq_condDistrib_of_algorithmDensitySource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosterior_ae_eq_condDistrib_of_algorithmDensitySource","description":"The reference environment posterior equals the actual posterior whenever the two algorithm-density pushforward laws use the same history density.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-ddff54a407f6","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2312,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:296"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePosterior_ae_eq_condDistrib_of_algorithmDensitySource {Omega : Type u} {OmegaRef : Type v} {History : Type w} {Env : Type x} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) (source : AlgorithmDensityPosteriorSource mu env history referenceMu referenceEnv referenceHistory) : (referencePosterior referenceMu referenceEnv referenceHistory hreferenceEnv hreferenceHistory).kernel =ᵐ[mu.map history] ProbabilityTheory.condDistrib env history mu","missing":[],"search":"referenceposterior_ae_eq_conddistrib_of_algorithmdensitysource banditrlproof.thompson.referenceposterior_ae_eq_conddistrib_of_algorithmdensitysource the reference environment posterior equals the actual posterior whenever the two algorithm-density pushforward laws use the same history density. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","label":"referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","description":"Reference-policy Thompson probability matching with algorithm-density laws as the only process-level input.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-edd1696e4118","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2313,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:357"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource {Omega : Type u} {OmegaRef : Type v} {History : Type w} {Env : Type x} {Action : Type y} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (hhistory : Measurable history) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (source : AlgorithmDensityPosteriorSource mu env history refere…","missing":[],"search":"referencepolicysampler_conddistrib_action_ae_eq_bestaction_of_algorithmdensitysource banditrlproof.thompson.referencepolicysampler_conddistrib_action_ae_eq_bestaction_of_algorithmdensitysource reference-policy thompson probability matching with algorithm-density laws as the only process-level input. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","label":"referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","description":"Reference-policy Thompson probability matching from an equal environment marginal and a conditional-history density law, without separately assuming the two algorithm-density pushforward laws.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-135ef950c0c4","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2314,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:396"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity {Omega : Type u} {OmegaRef : Type v} {History : Type w} {Env : Type x} {Action : Type y} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace History] [StandardBorelSpace History] [Nonempty History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (hhistory : Measurable history) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) (density : History -> ENNReal) (hdensity : Measurable density) (henvLaw :…","missing":[],"search":"referencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conddistrib_history_withdensity banditrlproof.thompson.referencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conddistrib_history_withdensity reference-policy thompson probability matching from an equal environment marginal and a conditional-history density law, without separately assuming the two algorithm-density pushforward laws. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","description":"Finite action/reward-prefix Thompson probability matching from a packaged pair of algorithm-density laws.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-8e22f9b84ced","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2315,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:442"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (henv : Measurable env) (haction : forall t : Nat, Measurable (fun omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_algorithmdensitysource banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_algorithmdensitysource finite action/reward-prefix thompson probability matching from a packaged pair of algorithm-density laws. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","description":"Finite action/reward-prefix Thompson probability matching from an equal environment marginal and a conditional finite-history density law.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensity/index.html#decl-0e788ec3e379","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","order":2316,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensity"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensity.lean:511"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (henv : Measurable env) (haction : forall t : Nat, Measurable (fun omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Ac…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conddistrib_history_withdensity banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conddistrib_history_withdensity finite action/reward-prefix thompson probability matching from an equal environment marginal and a conditional finite-history density law. theorem compiled","shard":"modules/e097ae87eb6a9118.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryAlgorithm","label":"HistoryAlgorithm","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryAlgorithm","description":"A stochastic policy indexed by inclusive finite action/reward histories.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-66f892cbe72e","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2317,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:25"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure HistoryAlgorithm (Action : Type u) (Reward : Type v) [MeasurableSpace Action] [MeasurableSpace Reward] where","missing":[],"search":"historyalgorithm banditrlproof.thompson.historyalgorithm a stochastic policy indexed by inclusive finite action/reward histories. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryEnvironment","label":"HistoryEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryEnvironment","description":"A stochastic feedback environment shared by the compared algorithms.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-10a52704cc34","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2318,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:49"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure HistoryEnvironment (Action : Type u) (Reward : Type v) [MeasurableSpace Action] [MeasurableSpace Reward] where","missing":[],"search":"historyenvironment banditrlproof.thompson.historyenvironment a stochastic feedback environment shared by the compared algorithms. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel","label":"historyStepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel","description":"Conditional law of the next action/reward pair after a finite history.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-a31a3feeacda","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2319,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:75"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def historyStepKernel {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (n : Nat) : ProbabilityTheory.Kernel (History.FinitePairHistory Action Reward n) (Action × Reward)","missing":[],"search":"historystepkernel banditrlproof.thompson.historystepkernel conditional law of the next action/reward pair after a finite history. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryAlgorithmAbsolutelyContinuous","label":"HistoryAlgorithmAbsolutelyContinuous","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryAlgorithmAbsolutelyContinuous","description":"Pointwise action-law absolute continuity between two algorithms.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-d336a4122ca7","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2320,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:95"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure HistoryAlgorithmAbsolutelyContinuous {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) : Prop where","missing":[],"search":"historyalgorithmabsolutelycontinuous banditrlproof.thompson.historyalgorithmabsolutelycontinuous pointwise action-law absolute continuity between two algorithms. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.singletonPairHistory","label":"singletonPairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.singletonPairHistory","description":"The unique history at time zero containing one action/reward pair.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-002fd66fdd00","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2321,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:104"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def singletonPairHistory {Action : Type u} {Reward : Type v} (pair : Action × Reward) : History.FinitePairHistory Action Reward 0","missing":[],"search":"singletonpairhistory banditrlproof.thompson.singletonpairhistory the unique history at time zero containing one action/reward pair. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.pairHistoryPrefix","label":"pairHistoryPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.pairHistoryPrefix","description":"Remove the last coordinate from an inclusive successor history.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-d10ee66e5872","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2322,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:110"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def pairHistoryPrefix {Action : Type u} {Reward : Type v} {n : Nat} (history : History.FinitePairHistory Action Reward (n + 1)) : History.FinitePairHistory Action Reward n","missing":[],"search":"pairhistoryprefix banditrlproof.thompson.pairhistoryprefix remove the last coordinate from an inclusive successor history. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.pairHistoryLast","label":"pairHistoryLast","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.pairHistoryLast","description":"Last pair in an inclusive successor history.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-ded988fc109b","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2323,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:118"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def pairHistoryLast {Action : Type u} {Reward : Type v} {n : Nat} (history : History.FinitePairHistory Action Reward (n + 1)) : Action × Reward","missing":[],"search":"pairhistorylast banditrlproof.thompson.pairhistorylast last pair in an inclusive successor history. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_singletonPairHistory","label":"measurable_singletonPairHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_singletonPairHistory","description":"theorem measurable_singletonPairHistory {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (@singletonPairHistory Action Reward)","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-dfe384a57f5d","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2324,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:124"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_singletonPairHistory {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (@singletonPairHistory Action Reward)","missing":[],"search":"measurable_singletonpairhistory banditrlproof.thompson.measurable_singletonpairhistory theorem measurable_singletonpairhistory {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] : measurable (@singletonpairhistory action reward) theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_pairHistoryPrefix","label":"measurable_pairHistoryPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_pairHistoryPrefix","description":"theorem measurable_pairHistoryPrefix {Action : Type u} {Reward : Type v} {n : Nat} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (@pairHistoryPrefix Action Reward n)","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-2f6e0399bcb2","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2325,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:130"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_pairHistoryPrefix {Action : Type u} {Reward : Type v} {n : Nat} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (@pairHistoryPrefix Action Reward n)","missing":[],"search":"measurable_pairhistoryprefix banditrlproof.thompson.measurable_pairhistoryprefix theorem measurable_pairhistoryprefix {action : type u} {reward : type v} {n : nat} [measurablespace action] [measurablespace reward] : measurable (@pairhistoryprefix action reward n) theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_pairHistoryLast","label":"measurable_pairHistoryLast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_pairHistoryLast","description":"theorem measurable_pairHistoryLast {Action : Type u} {Reward : Type v} {n : Nat} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (@pairHistoryLast Action Reward n)","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-4b95bf00b53f","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2326,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:136"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_pairHistoryLast {Action : Type u} {Reward : Type v} {n : Nat} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (@pairHistoryLast Action Reward n)","missing":[],"search":"measurable_pairhistorylast banditrlproof.thompson.measurable_pairhistorylast theorem measurable_pairhistorylast {action : type u} {reward : type v} {n : nat} [measurablespace action] [measurablespace reward] : measurable (@pairhistorylast action reward n) theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.pairHistoryPrefix_extendPairHistorySucc","label":"pairHistoryPrefix_extendPairHistorySucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.pairHistoryPrefix_extendPairHistorySucc","description":"theorem pairHistoryPrefix_extendPairHistorySucc {Action : Type u} {Reward : Type v} {n : Nat} (history : History.FinitePairHistory Action Reward n) (next : Action × Reward) : pairHistoryPrefix (History.extendPairHistorySucc history next) = history","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-63e6f6611d23","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2327,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:143"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem pairHistoryPrefix_extendPairHistorySucc {Action : Type u} {Reward : Type v} {n : Nat} (history : History.FinitePairHistory Action Reward n) (next : Action × Reward) : pairHistoryPrefix (History.extendPairHistorySucc history next) = history","missing":[],"search":"pairhistoryprefix_extendpairhistorysucc banditrlproof.thompson.pairhistoryprefix_extendpairhistorysucc theorem pairhistoryprefix_extendpairhistorysucc {action : type u} {reward : type v} {n : nat} (history : history.finitepairhistory action reward n) (next : action × reward) : pairhistoryprefix (history.extendpairhistorysucc history next) = history theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.pairHistoryLast_extendPairHistorySucc","label":"pairHistoryLast_extendPairHistorySucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.pairHistoryLast_extendPairHistorySucc","description":"theorem pairHistoryLast_extendPairHistorySucc {Action : Type u} {Reward : Type v} {n : Nat} (history : History.FinitePairHistory Action Reward n) (next : Action × Reward) : pairHistoryLast (History.extendPairHistorySucc history next) = next","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-b421bbc23362","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2328,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:153"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem pairHistoryLast_extendPairHistorySucc {Action : Type u} {Reward : Type v} {n : Nat} (history : History.FinitePairHistory Action Reward n) (next : Action × Reward) : pairHistoryLast (History.extendPairHistorySucc history next) = next","missing":[],"search":"pairhistorylast_extendpairhistorysucc banditrlproof.thompson.pairhistorylast_extendpairhistorysucc theorem pairhistorylast_extendpairhistorysucc {action : type u} {reward : type v} {n : nat} (history : history.finitepairhistory action reward n) (next : action × reward) : pairhistorylast (history.extendpairhistorysucc history next) = next theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyDensity","label":"historyDensity","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.historyDensity","description":"Recursive product of the initial and policy action likelihood ratios.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-b82a5ea6056c","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2329,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:161"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def historyDensity {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Action] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) : (n : Nat) -> History.FinitePairHistory Action Reward n -> ENNReal | 0, history => algorithm.initialAction.rnDeriv referenceAlgorithm.initialAction (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 | n + 1, history => historyDensity algorithm referenceAlgorithm n (pairHistoryPrefix history) * (algorithm.policy n).rnDeriv (referenceAlgorithm.policy n) (pairHistoryPrefix history) (pairHistoryLast history).1 theorem measurable_historyDensity {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Action] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (n : Nat) : Measurable (historyDensity al…","missing":[],"search":"historydensity banditrlproof.thompson.historydensity recursive product of the initial and policy action likelihood ratios. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_historyDensity","label":"measurable_historyDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_historyDensity","description":"theorem measurable_historyDensity {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Action] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (n : Nat) : Measurable (historyDensity algorithm referenceAlgorithm n)","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-641f8335eecc","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2330,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:175"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyDensity {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Action] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (n : Nat) : Measurable (historyDensity algorithm referenceAlgorithm n)","missing":[],"search":"measurable_historydensity banditrlproof.thompson.measurable_historydensity theorem measurable_historydensity {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] [measurablespace.countablygenerated action] (algorithm referencealgorithm : historyalgorithm action reward) (n : nat) : measurable (historydensity algorithm referencealgorithm n) theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.map_withDensity_comp","label":"map_withDensity_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.map_withDensity_comp","description":"Mapping a weighted measure transports a density pulled back by the map.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-c022f658ff0d","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2331,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:192"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem map_withDensity_comp {Source : Type u} {Target : Type v} [MeasurableSpace Source] [MeasurableSpace Target] (mu : Measure Source) (map : Source -> Target) (density : Target -> ENNReal) (hmap : Measurable map) (hdensity : Measurable density) : (mu.withDensity (density ∘ map)).map map = (mu.map map).withDensity density","missing":[],"search":"map_withdensity_comp banditrlproof.thompson.map_withdensity_comp mapping a weighted measure transports a density pulled back by the map. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.compProd_withDensity_withDensity","label":"compProd_withDensity_withDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.compProd_withDensity_withDensity","description":"Weighting both a composition-product base and kernel multiplies densities.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-bcbaecb36507","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2332,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:207"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem compProd_withDensity_withDensity {Base : Type u} {Target : Type v} [MeasurableSpace Base] [MeasurableSpace Target] (base : Measure Base) [SFinite base] (kernel : ProbabilityTheory.Kernel Base Target) [ProbabilityTheory.IsSFiniteKernel kernel] (baseDensity : Base -> ENNReal) (kernelDensity : Base -> Target -> ENNReal) (hbaseDensity : Measurable baseDensity) (hkernelDensity : Measurable (Function.uncurry kernelDensity)) [ProbabilityTheory.IsSFiniteKernel (kernel.withDensity kernelDensity)] : base.withDensity baseDensity ⊗ₘ kernel.withDensity kernelDensity = (base ⊗ₘ kernel).withDensity (fun pair => baseDensity pair.1 * kernelDensity pair.1 pair.2)","missing":[],"search":"compprod_withdensity_withdensity banditrlproof.thompson.compprod_withdensity_withdensity weighting both a composition-product base and kernel multiplies densities. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.kernel_withDensity_rnDeriv_eq_of_absolutelyContinuous","label":"kernel_withDensity_rnDeriv_eq_of_absolutelyContinuous","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.kernel_withDensity_rnDeriv_eq_of_absolutelyContinuous","description":"Kernel RN derivatives reconstruct a pointwise absolutely continuous kernel.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-17d871ce80ad","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2333,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:228"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem kernel_withDensity_rnDeriv_eq_of_absolutelyContinuous {Index : Type u} {Target : Type v} [MeasurableSpace Index] [MeasurableSpace Target] [MeasurableSpace.CountableOrCountablyGenerated Index Target] (kernel referenceKernel : ProbabilityTheory.Kernel Index Target) [ProbabilityTheory.IsFiniteKernel kernel] [ProbabilityTheory.IsFiniteKernel referenceKernel] (h : forall index, kernel index ≪ referenceKernel index) : referenceKernel.withDensity (kernel.rnDeriv referenceKernel) = kernel","missing":[],"search":"kernel_withdensity_rnderiv_eq_of_absolutelycontinuous banditrlproof.thompson.kernel_withdensity_rnderiv_eq_of_absolutelycontinuous kernel rn derivatives reconstruct a pointwise absolutely continuous kernel. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.kernel_compProd_withDensity_left","label":"kernel_compProd_withDensity_left","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.kernel_compProd_withDensity_left","description":"Weighting the left kernel of a kernel composition product.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-3044d07090fb","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2334,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:243"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem kernel_compProd_withDensity_left {Index : Type u} {Middle : Type v} {Target : Type w} [MeasurableSpace Index] [MeasurableSpace Middle] [MeasurableSpace Target] (kernel : ProbabilityTheory.Kernel Index Middle) (nextKernel : ProbabilityTheory.Kernel (Index × Middle) Target) [ProbabilityTheory.IsSFiniteKernel kernel] [ProbabilityTheory.IsSFiniteKernel nextKernel] (density : Index -> Middle -> ENNReal) (hdensity : Measurable (Function.uncurry density)) [ProbabilityTheory.IsSFiniteKernel (kernel.withDensity density)] : kernel.withDensity density ⊗ₖ nextKernel = (kernel ⊗ₖ nextKernel).withDensity (fun index pair => density index pair.1)","missing":[],"search":"kernel_compprod_withdensity_left banditrlproof.thompson.kernel_compprod_withdensity_left weighting the left kernel of a kernel composition product. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel_eq_withDensity","label":"historyStepKernel_eq_withDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel_eq_withDensity","description":"The actual one-step pair kernel is a density-weighted reference kernel.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-cd4595b83d7b","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2335,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:273"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernel_eq_withDensity {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) [MeasurableSpace.CountableOrCountablyGenerated (History.FinitePairHistory Action Reward n) Action] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (hcontinuous : HistoryAlgorithmAbsolutelyContinuous algorithm referenceAlgorithm) : historyStepKernel algorithm environment n = (historyStepKernel referenceAlgorithm environment n).withDensity (fun history pair => (algorithm.policy n).rnDeriv (referenceAlgorithm.policy n) history pair.1)","missing":[],"search":"historystepkernel_eq_withdensity banditrlproof.thompson.historystepkernel_eq_withdensity the actual one-step pair kernel is a density-weighted reference kernel. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.IsHistoryAlgorithmEnvironmentSequence","label":"IsHistoryAlgorithmEnvironmentSequence","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.IsHistoryAlgorithmEnvironmentSequence","description":"Process contract using the combined initial and successor pair laws. Split action/feedback conditional laws are assembled into this contract below.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-1a97895d43a4","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2336,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:312"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure IsHistoryAlgorithmEnvironmentSequence {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : Prop where","missing":[],"search":"ishistoryalgorithmenvironmentsequence banditrlproof.thompson.ishistoryalgorithmenvironmentsequence process contract using the combined initial and successor pair laws. split action/feedback conditional laws are assembled into this contract below. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.isHistoryAlgorithmEnvironmentSequence_of_split","label":"isHistoryAlgorithmEnvironmentSequence_of_split","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.isHistoryAlgorithmEnvironmentSequence_of_split","description":"Build the pair-law process contract from LML-shaped split fields.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-b84d167910b0","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2337,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:337"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def isHistoryAlgorithmEnvironmentSequence_of_split {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (haction : forall n, Measurable (fun omega => action omega n)) (hreward : forall n, Measurable (fun omega => reward omega n)) (hinitialAction : mu.map (fun omega => action omega 0) = algorithm.initialAction) (hinitialFeedback : ProbabilityTheory.condDistrib (fun omega => reward omega 0) (fun omega => action omega 0) mu =ᵐ[ mu.map (fun omega => action omega 0)] environment.initialFeedback) (hpolicy…","missing":[],"search":"ishistoryalgorithmenvironmentsequence_of_split banditrlproof.thompson.ishistoryalgorithmenvironmentsequence_of_split build the pair-law process contract from lml-shaped split fields. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.nextPairJointLaw_eq_compProd","label":"nextPairJointLaw_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.nextPairJointLaw_eq_compProd","description":"The process contract identifies the joint law of the current finite history and the next action/reward pair with the corresponding measure composition product.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-2ce1ae593004","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2338,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:396"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem nextPairJointLaw_eq_compProd {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) (n : Nat) : mu.map (fun omega => (History.finitePairHistoryOfTrace (action omega) (reward omega) n, (action omega (n + 1), reward omega (n + 1)))) = mu.map (fun omega => History.finitePairHistoryOfTrace (action omega) (reward omega) n) ⊗ₘ historyStepKernel algorithm environment n","missing":[],"search":"nextpairjointlaw_eq_compprod banditrlproof.thompson.nextpairjointlaw_eq_compprod the process contract identifies the joint law of the current finite history and the next action/reward pair with the corresponding measure composition product. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairHistory_map_eq_withDensity","label":"finitePairHistory_map_eq_withDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairHistory_map_eq_withDensity","description":"Finite-history algorithm-density transport. Two history-dependent stochastic policies use the same feedback environment. Pointwise absolute continuity of the actual policy with respect to the reference policy implies that every inclusive finite pair-history law is the reference law weighted by the recursive product of action likelihood ratios.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-02edfd356de7","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2339,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:433"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_map_eq_withDensity {Omega : Type w} {OmegaRef : Type x} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (muRef : Measure OmegaRef) [IsFiniteMeasure muRef] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (actionRef : OmegaRef -> ActionTrace Action) (rewardRef : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) (referenceSource : IsHistoryAlgorithmEnvironmentSequence muRef actionRef rewardRef referenceAlgorithm environment) (hcontinuou…","missing":[],"search":"finitepairhistory_map_eq_withdensity banditrlproof.thompson.finitepairhistory_map_eq_withdensity finite-history algorithm-density transport. two history-dependent stochastic policies use the same feedback environment. pointwise absolute continuity of the actual policy with respect to the reference policy implies that every inclusive finite pair-history law is the reference law weighted by the recursive product of action likelihood ratios. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.ConditionalHistoryAlgorithmDensitySource","label":"ConditionalHistoryAlgorithmDensitySource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.ConditionalHistoryAlgorithmDensitySource","description":"Environment-indexed process realization for algorithm-density transport. The regular conditional sample laws `condDistrib id env mu` and `condDistrib id referenceEnv referenceMu` must satisfy the actual and reference process contracts almost everywhere. The compared algorithms share the same environment-indexed feedback law.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-9c67f3e96281","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2340,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:732"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure ConditionalHistoryAlgorithmDensitySource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment…","missing":[],"search":"conditionalhistoryalgorithmdensitysource banditrlproof.thompson.conditionalhistoryalgorithmdensitysource environment-indexed process realization for algorithm-density transport. the regular conditional sample laws `conddistrib id env mu` and `conddistrib id referenceenv referencemu` must satisfy the actual and reference process contracts almost everywhere. the compared algorithms share the same environment-indexed feedback law. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_finitePairHistory_eq_withDensity_of_conditionalProcessSource","label":"condDistrib_finitePairHistory_eq_withDensity_of_conditionalProcessSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_finitePairHistory_eq_withDensity_of_conditionalProcessSource","description":"The environment-indexed process realization produces the conditional finite history density law required by the posterior-invariance source constructor.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-577b0e764216","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2341,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:776"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_finitePairHistory_eq_withDensity_of_conditionalProcessSource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironme…","missing":[],"search":"conddistrib_finitepairhistory_eq_withdensity_of_conditionalprocesssource banditrlproof.thompson.conddistrib_finitepairhistory_eq_withdensity_of_conditionalprocesssource the environment-indexed process realization produces the conditional finite history density law required by the posterior-invariance source constructor. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalProcessSource","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalProcessSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalProcessSource","description":"Finite-prefix Thompson probability matching produced directly from the environment-indexed recursive process contracts.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-ede21f609052","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2342,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:873"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalProcessSource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referen…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conditionalprocesssource banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conditionalprocesssource finite-prefix thompson probability matching produced directly from the environment-indexed recursive process contracts. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.ConditionalHistoryAlgorithmEnvironmentSplitSource","label":"ConditionalHistoryAlgorithmEnvironmentSplitSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.ConditionalHistoryAlgorithmEnvironmentSplitSource","description":"LML-shaped split laws for one algorithm/environment process under the regular conditional sample measure at almost every environment.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-3e139cf7b5a8","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2343,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:935"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure ConditionalHistoryAlgorithmEnvironmentSplitSource {Omega : Type u} {Env : Type v} {Action : Type w} {Reward : Type x} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) : Prop where","missing":[],"search":"conditionalhistoryalgorithmenvironmentsplitsource banditrlproof.thompson.conditionalhistoryalgorithmenvironmentsplitsource lml-shaped split laws for one algorithm/environment process under the regular conditional sample measure at almost every environment. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.ConditionalHistoryAlgorithmDensitySplitSource","label":"ConditionalHistoryAlgorithmDensitySplitSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.ConditionalHistoryAlgorithmDensitySplitSource","description":"Actual/reference split conditional-process laws together with the common environment marginal and policy absolute-continuity contract.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-e645a957d990","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2344,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:989"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure ConditionalHistoryAlgorithmDensitySplitSource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnviro…","missing":[],"search":"conditionalhistoryalgorithmdensitysplitsource banditrlproof.thompson.conditionalhistoryalgorithmdensitysplitsource actual/reference split conditional-process laws together with the common environment marginal and policy absolute-continuity contract. structure compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmDensitySource_of_split","label":"conditionalHistoryAlgorithmDensitySource_of_split","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.conditionalHistoryAlgorithmDensitySource_of_split","description":"Assemble the conditional process source from the four split law families for the actual and reference processes.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-5fa0c114453f","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2345,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:1021"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def conditionalHistoryAlgorithmDensitySource_of_split {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> H…","missing":[],"search":"conditionalhistoryalgorithmdensitysource_of_split banditrlproof.thompson.conditionalhistoryalgorithmdensitysource_of_split assemble the conditional process source from the four split law families for the actual and reference processes. definition compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_finitePairHistory_eq_withDensity_of_conditionalSplitSource","label":"condDistrib_finitePairHistory_eq_withDensity_of_conditionalSplitSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_finitePairHistory_eq_withDensity_of_conditionalSplitSource","description":"The split conditional laws directly produce the conditional finite-history density equality.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-e9d2b909708a","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2346,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:1141"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_finitePairHistory_eq_withDensity_of_conditionalSplitSource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment…","missing":[],"search":"conddistrib_finitepairhistory_eq_withdensity_of_conditionalsplitsource banditrlproof.thompson.conddistrib_finitepairhistory_eq_withdensity_of_conditionalsplitsource the split conditional laws directly produce the conditional finite-history density equality. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalSplitSource","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalSplitSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalSplitSource","description":"Finite-prefix Thompson probability matching from the concrete four-family split conditional-law interface.","url":"../modules/banditrlproof-algorithms-thompsonalgorithmdensityprocess/index.html#decl-621e65a44d62","parent":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","order":2347,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess"],["Source","BanditRLProof/Algorithms/ThompsonAlgorithmDensityProcess.lean:1187"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalSplitSource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm reference…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conditionalsplitsource banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_conditionalsplitsource finite-prefix thompson probability matching from the concrete four-family split conditional-law interface. theorem compiled","shard":"modules/0f531e9851a10d66.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryActionScore","label":"HistoryActionScore","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryActionScore","description":"A real score whose time-`n + 1` value sees exactly the history through `n`.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-e8761b5612e1","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2348,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:23"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure HistoryActionScore (Action : Type u) (Reward : Type v) [MeasurableSpace Action] [MeasurableSpace Reward] where","missing":[],"search":"historyactionscore banditrlproof.thompson.historyactionscore a real score whose time-`n + 1` value sees exactly the history through `n`. structure compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryActionScore.atTrace","label":"atTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryActionScore.atTrace","description":"Evaluate a history score on the action selected by one complete trace.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-f78a6f63c37b","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2349,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:37"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def atTrace {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (action : ActionTrace Action) (reward : RewardTrace Reward) : Nat -> Real | 0 => score.initial (action 0) | n + 1 => score.successor n (History.finitePairHistoryOfTrace action reward n) (action (n + 1)) /-- Evaluate the same visible-history score at a comparison action. -/ def atBestTrace {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (bestAction : Action) (action : ActionTrace Action) (reward : RewardTrace Reward) : Nat -> Real | 0 => score.initial bestAction | n + 1 => score.successor n (History.finitePairHistoryOfTrace action reward n) bestAction theorem measurable_atTrace {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace…","missing":[],"search":"attrace banditrlproof.thompson.historyactionscore.attrace evaluate a history score on the action selected by one complete trace. definition compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryActionScore.atBestTrace","label":"atBestTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryActionScore.atBestTrace","description":"Evaluate the same visible-history score at a comparison action.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-e97583e5630b","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2350,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:48"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def atBestTrace {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (bestAction : Action) (action : ActionTrace Action) (reward : RewardTrace Reward) : Nat -> Real | 0 => score.initial bestAction | n + 1 => score.successor n (History.finitePairHistoryOfTrace action reward n) bestAction theorem measurable_atTrace {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (t : Nat) : Measurable (fun omega => score.atTrace (action omega) (reward omega) t)","missing":[],"search":"atbesttrace banditrlproof.thompson.historyactionscore.atbesttrace evaluate the same visible-history score at a comparison action. definition compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryActionScore.measurable_atTrace","label":"measurable_atTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryActionScore.measurable_atTrace","description":"theorem measurable_atTrace {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (t : Nat) : Measur…","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-307dc6ff89f0","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2351,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:59"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_atTrace {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (t : Nat) : Measurable (fun omega => score.atTrace (action omega) (reward omega) t)","missing":[],"search":"measurable_attrace banditrlproof.thompson.historyactionscore.measurable_attrace theorem measurable_attrace {omega : type w} {action : type u} {reward : type v} [measurablespace omega] [measurablespace action] [measurablespace reward] (score : historyactionscore action reward) (action : omega -> actiontrace action) (reward : omega -> rewardtrace reward) (haction : forall t, measurable (fun omega => action omega t)) (hreward : forall t, measurable (fun omega => reward omega t)) (t : nat) : measurable (fun omega => score.attrace (action omega) (reward omega) t) theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryActionScore.measurable_atBestTrace","label":"measurable_atBestTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryActionScore.measurable_atBestTrace","description":"theorem measurable_atBestTrace {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (bestAction : Omega -> Action) (hbestAction : Measurable bestAction) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t, Measurable (fun omega => action omega t)) (hreward…","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-bec41512f3da","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2352,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:77"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_atBestTrace {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (bestAction : Omega -> Action) (hbestAction : Measurable bestAction) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (t : Nat) : Measurable (fun omega => score.atBestTrace (bestAction omega) (action omega) (reward omega) t)","missing":[],"search":"measurable_atbesttrace banditrlproof.thompson.historyactionscore.measurable_atbesttrace theorem measurable_atbesttrace {omega : type w} {action : type u} {reward : type v} [measurablespace omega] [measurablespace action] [measurablespace reward] (score : historyactionscore action reward) (bestaction : omega -> action) (hbestaction : measurable bestaction) (action : omega -> actiontrace action) (reward : omega -> rewardtrace reward) (haction : forall t, measurable (fun omega => action omega t)) (hreward : forall t, measurable (fun omega => reward omega t)) (t : nat) : measurable (fun omega => score.atbesttrace (bestaction omega) (action omega) (reward omega) t) theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integral_comp_eq_of_map_eq","label":"integral_comp_eq_of_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integral_comp_eq_of_map_eq","description":"Equal pushforwards give equal integrals of every measurable real score.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-393fd7dcb066","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2353,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:100"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integral_comp_eq_of_map_eq {Omega : Type u} {Target : Type v} [MeasurableSpace Omega] [MeasurableSpace Target] (mu : Measure Omega) (left right : Omega -> Target) (hleft : Measurable left) (hright : Measurable right) (hmap : mu.map left = mu.map right) (score : Target -> Real) (hscore : Measurable score) : integral mu (fun omega => score (left omega)) = integral mu (fun omega => score (right omega))","missing":[],"search":"integral_comp_eq_of_map_eq banditrlproof.thompson.integral_comp_eq_of_map_eq equal pushforwards give equal integrals of every measurable real score. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integral_historyAction_eq_of_condDistrib_ae_eq","label":"integral_historyAction_eq_of_condDistrib_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integral_historyAction_eq_of_condDistrib_ae_eq","description":"If two actions have the same conditional law given a history, every measurable history/action score has the same expectation under those actions.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-5928f7f62e78","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2354,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:121"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integral_historyAction_eq_of_condDistrib_ae_eq {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action bestAction : Omega -> Action) (haction : Measurable action) (hbestAction : Measurable bestAction) (hcond : condDistrib action history mu =ᵐ[mu.map history] condDistrib bestAction history mu) (score : History × Action -> Real) (hscore : Measurable score) : integral mu (fun omega => score (history omega, action omega)) = integral mu (fun omega => score (history omega, bestAction omega))","missing":[],"search":"integral_historyaction_eq_of_conddistrib_ae_eq banditrlproof.thompson.integral_historyaction_eq_of_conddistrib_ae_eq if two actions have the same conditional law given a history, every measurable history/action score has the same expectation under those actions. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_action_zero","label":"canonicalMeasurableEnvironmentTrajectoryKernel_map_action_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_action_zero","description":"The environment-indexed trajectory kernel has the algorithm's action law at time zero.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-b8defb80e2fb","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2355,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:146"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_map_action_zero {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) : (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun trajectory => (trajectory 0).1) = Kernel.const Env algorithm.initialAction","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_map_action_zero banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_map_action_zero the environment-indexed trajectory kernel has the algorithm's action law at time zero. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryMeasure_map_action_zero","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_map_action_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryMeasure_map_action_zero","description":"Mixing the trajectory kernel through a probability prior preserves its initial action law.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-e962cad7ea7a","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2356,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:168"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_map_action_zero {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) : (prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample => (sample.2 0).1) = algorithm.initialAction","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_map_action_zero banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorymeasure_map_action_zero mixing the trajectory kernel through a probability prior preserves its initial action law. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_map_action_zero_eq_bestAction","label":"uniformReferenceThompsonAlgorithm_map_action_zero_eq_bestAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_map_action_zero_eq_bestAction","description":"The initial action and latent best action have the same marginal on the actual TS trajectory.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-59e5ff2877ba","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2357,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:201"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem uniformReferenceThompsonAlgorithm_map_action_zero_eq_bestAction {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Fintype Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : let algorithm := uniformReferenceThompsonAlgorithm prior environment bestAction hbestAction let actualMeasure := prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment actualMeasure.map (fun sample => environmentTrajectoryAction sample 0) = actualMeasure.map (bestAction ∘ Prod.fst)","missing":[],"search":"uniformreferencethompsonalgorithm_map_action_zero_eq_bestaction banditrlproof.thompson.uniformreferencethompsonalgorithm_map_action_zero_eq_bestaction the initial action and latent best action have the same marginal on the actual ts trajectory. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.trajectoryHistoryScore","label":"trajectoryHistoryScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.trajectoryHistoryScore","description":"Evaluate the history score on the actual trajectory action.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-658fd648cdd5","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2358,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:245"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def trajectoryHistoryScore {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (sample : Env × ((n : Nat) -> Action × Reward)) (t : Nat) : Real","missing":[],"search":"trajectoryhistoryscore banditrlproof.thompson.trajectoryhistoryscore evaluate the history score on the actual trajectory action. definition compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.trajectoryBestHistoryScore","label":"trajectoryBestHistoryScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.trajectoryBestHistoryScore","description":"Evaluate the same score at the latent environment's best action.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-425cd323b130","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2359,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:254"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def trajectoryBestHistoryScore {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Action] [MeasurableSpace Reward] (score : HistoryActionScore Action Reward) (bestAction : Env -> Action) (sample : Env × ((n : Nat) -> Action × Reward)) (t : Nat) : Real","missing":[],"search":"trajectorybesthistoryscore banditrlproof.thompson.trajectorybesthistoryscore evaluate the same score at the latent environment's best action. definition compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_integral_historyScore_eq_bestAction","label":"uniformReferenceThompsonAlgorithm_integral_historyScore_eq_bestAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_integral_historyScore_eq_bestAction","description":"Probability matching on the actual recursive trajectory implies equality of every visible-history score at the selected and latent-best actions.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-af40a1afc295","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2360,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:268"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem uniformReferenceThompsonAlgorithm_integral_historyScore_eq_bestAction {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Fintype Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (score : HistoryActionScore Action Reward) (t : Nat) : let algorithm := uniformReferenceThompsonAlgorithm prior environment bestAction hbestAction let actualMeasure := prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment integral actualMeasure (fun sample => trajectoryHistoryScore score sample t) = integral actualMeasure (fun sample => traje…","missing":[],"search":"uniformreferencethompsonalgorithm_integral_historyscore_eq_bestaction banditrlproof.thompson.uniformreferencethompsonalgorithm_integral_historyscore_eq_bestaction probability matching on the actual recursive trajectory implies equality of every visible-history score at the selected and latent-best actions. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.IsOptimalMeanSelector","label":"IsOptimalMeanSelector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.IsOptimalMeanSelector","description":"A selector is mean-optimal when it maximizes the environment-dependent action mean pointwise. This contract separates a genuine best-action regret interpretation from the comparator-relative algebra below.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-0681eefab214","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2361,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:341"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def IsOptimalMeanSelector {Env : Type u} {Action : Type v} (mean : Env -> Action -> Real) (bestAction : Env -> Action) : Prop","missing":[],"search":"isoptimalmeanselector banditrlproof.thompson.isoptimalmeanselector a selector is mean-optimal when it maximizes the environment-dependent action mean pointwise. this contract separates a genuine best-action regret interpretation from the comparator-relative algebra below. definition compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.trajectoryBayesMeanRegret","label":"trajectoryBayesMeanRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.trajectoryBayesMeanRegret","description":"Finite-horizon comparator-relative mean regret. It is Bayesian pseudo-regret when `bestAction` satisfies `IsOptimalMeanSelector mean bestAction` and the environment is integrated against a prior.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-1dc7b773f4fb","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2362,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:349"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def trajectoryBayesMeanRegret {Env : Type u} {Action : Type v} {Reward : Type w} (mean : Env -> Action -> Real) (bestAction : Env -> Action) (sample : Env × ((n : Nat) -> Action × Reward)) (horizon : Nat) : Real","missing":[],"search":"trajectorybayesmeanregret banditrlproof.thompson.trajectorybayesmeanregret finite-horizon comparator-relative mean regret. it is bayesian pseudo-regret when `bestaction` satisfies `isoptimalmeanselector mean bestaction` and the environment is integrated against a prior. definition compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_historyScore","label":"integral_trajectoryBayesMeanRegret_eq_add_historyScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_historyScore","description":"LML-shaped Thompson Bayesian regret decomposition on the locally generated recursive trajectory. Integrability is explicit; bounded clipped scores and bounded action means can discharge these four contracts downstream.","url":"../modules/banditrlproof-algorithms-thompsonbayesregretdecomposition/index.html#decl-d3efce1c2be8","parent":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","order":2363,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition"],["Source","BanditRLProof/Algorithms/ThompsonBayesRegretDecomposition.lean:362"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integral_trajectoryBayesMeanRegret_eq_add_historyScore {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Fintype Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (mean : Env -> Action -> Real) (score : HistoryActionScore Action Reward) (horizon : Nat) (hmeanBest : Integrable (fun sample : Env × ((n : Nat) -> Action × Reward) => mean sample.1 (bestAction sample.1)) (prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel (uniformReferenceThompsonAlgorithm prior environment bestAction hbestAction) environment)) (hmeanAction : forall t : Nat,…","missing":[],"search":"integral_trajectorybayesmeanregret_eq_add_historyscore banditrlproof.thompson.integral_trajectorybayesmeanregret_eq_add_historyscore lml-shaped thompson bayesian regret decomposition on the locally generated recursive trajectory. integrability is explicit; bounded clipped scores and bounded action means can discharge these four contracts downstream. theorem compiled","shard":"modules/c7c8af48cb268eac.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalActionKernel","label":"canonicalActionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalActionKernel","description":"The Thompson action kernel induced by Mathlib's canonical posterior.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-54f87899443b","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2364,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:23"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalActionKernel {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (_hbestAction : Measurable bestAction) : ProbabilityTheory.Kernel History Action","missing":[],"search":"canonicalactionkernel banditrlproof.thompson.canonicalactionkernel the thompson action kernel induced by mathlib's canonical posterior. definition compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalActionKernelOnPair","label":"canonicalActionKernelOnPair","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalActionKernelOnPair","description":"Lift the history-indexed action kernel to environment/history pairs.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-299282318ae7","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2365,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:50"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalActionKernelOnPair {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : ProbabilityTheory.Kernel (Env × History) Action","missing":[],"search":"canonicalactionkernelonpair banditrlproof.thompson.canonicalactionkernelonpair lift the history-indexed action kernel to environment/history pairs. definition compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerMeasure","label":"canonicalSamplerMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerMeasure","description":"The canonical one-step Thompson law on `(Env × History) × Action`. The first component is sampled from `prior ⊗ₘ likelihood`; the action is then sampled from the canonical posterior mapped by `bestAction` at that history.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-4856c639e20d","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2366,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:83"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalSamplerMeasure {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : Measure ((Env × History) × Action)","missing":[],"search":"canonicalsamplermeasure banditrlproof.thompson.canonicalsamplermeasure the canonical one-step thompson law on `(env × history) × action`. the first component is sampled from `prior ⊗ₘ likelihood`; the action is then sampled from the canonical posterior mapped by `bestaction` at that history. definition compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerEnv","label":"canonicalSamplerEnv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerEnv","description":"Environment coordinate of the canonical sampler source.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-855bbeb60d5b","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2367,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:111"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def canonicalSamplerEnv {Env History Action : Type*} : (Env × History) × Action -> Env","missing":[],"search":"canonicalsamplerenv banditrlproof.thompson.canonicalsamplerenv environment coordinate of the canonical sampler source. definition compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerHistory","label":"canonicalSamplerHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerHistory","description":"History coordinate of the canonical sampler source.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-43aaeaff504d","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2368,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:116"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def canonicalSamplerHistory {Env History Action : Type*} : (Env × History) × Action -> History","missing":[],"search":"canonicalsamplerhistory banditrlproof.thompson.canonicalsamplerhistory history coordinate of the canonical sampler source. definition compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerAction","label":"canonicalSamplerAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerAction","description":"Action coordinate of the canonical sampler source.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-7aed5e3b8e76","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2369,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:121"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def canonicalSamplerAction {Env History Action : Type*} : (Env × History) × Action -> Action","missing":[],"search":"canonicalsampleraction banditrlproof.thompson.canonicalsampleraction action coordinate of the canonical sampler source. definition compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerEnv_measurable","label":"canonicalSamplerEnv_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerEnv_measurable","description":"theorem canonicalSamplerEnv_measurable {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@canonicalSamplerEnv Env History Action)","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-725f579fbef0","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2370,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:125"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSamplerEnv_measurable {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@canonicalSamplerEnv Env History Action)","missing":[],"search":"canonicalsamplerenv_measurable banditrlproof.thompson.canonicalsamplerenv_measurable theorem canonicalsamplerenv_measurable {env history action : type*} [measurablespace env] [measurablespace history] [measurablespace action] : measurable (@canonicalsamplerenv env history action) theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerHistory_measurable","label":"canonicalSamplerHistory_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerHistory_measurable","description":"theorem canonicalSamplerHistory_measurable {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@canonicalSamplerHistory Env History Action)","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-f28a3115baf1","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2371,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:131"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSamplerHistory_measurable {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@canonicalSamplerHistory Env History Action)","missing":[],"search":"canonicalsamplerhistory_measurable banditrlproof.thompson.canonicalsamplerhistory_measurable theorem canonicalsamplerhistory_measurable {env history action : type*} [measurablespace env] [measurablespace history] [measurablespace action] : measurable (@canonicalsamplerhistory env history action) theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSamplerAction_measurable","label":"canonicalSamplerAction_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSamplerAction_measurable","description":"theorem canonicalSamplerAction_measurable {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@canonicalSamplerAction Env History Action)","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-f8898debb790","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2372,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:137"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSamplerAction_measurable {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@canonicalSamplerAction Env History Action)","missing":[],"search":"canonicalsampleraction_measurable banditrlproof.thompson.canonicalsampleraction_measurable theorem canonicalsampleraction_measurable {env history action : type*} [measurablespace env] [measurablespace history] [measurablespace action] : measurable (@canonicalsampleraction env history action) theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.map_compProd_comap_snd","label":"map_compProd_comap_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.map_compProd_comap_snd","description":"Projecting a composition product whose kernel depends only on the second base coordinate gives the second-coordinate marginal composed with that kernel.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-88defae578db","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2373,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:147"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem map_compProd_comap_snd {Env History Action : Type*} [MeasurableSpace Env] [MeasurableSpace History] [MeasurableSpace Action] (mu : Measure (Env × History)) [IsFiniteMeasure mu] (actionKernel : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel actionKernel] : (mu ⊗ₘ actionKernel.comap Prod.snd measurable_snd).map (fun sample => (sample.1.2, sample.2)) = mu.map Prod.snd ⊗ₘ actionKernel","missing":[],"search":"map_compprod_comap_snd banditrlproof.thompson.map_compprod_comap_snd projecting a composition product whose kernel depends only on the second base coordinate gives the second-coordinate marginal composed with that kernel. theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSampler_env_history_map_eq","label":"canonicalSampler_env_history_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSampler_env_history_map_eq","description":"The canonical sampler preserves the prescribed environment/history law.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-1771ffb85257","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2374,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:181"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSampler_env_history_map_eq {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : (canonicalSamplerMeasure prior likelihood bestAction hbestAction).map (fun sample => (canonicalSamplerEnv sample, canonicalSamplerHistory sample)) = PosteriorKernel.canonicalJointMeasure prior likelihood","missing":[],"search":"canonicalsampler_env_history_map_eq banditrlproof.thompson.canonicalsampler_env_history_map_eq the canonical sampler preserves the prescribed environment/history law. theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSampler_history_action_map_eq","label":"canonicalSampler_history_action_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSampler_history_action_map_eq","description":"The history/action marginal is generated by the canonical action kernel.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-e49aa60a24d0","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2375,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:200"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSampler_history_action_map_eq {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : (canonicalSamplerMeasure prior likelihood bestAction hbestAction).map (fun sample => (canonicalSamplerHistory sample, canonicalSamplerAction sample)) = (canonicalSamplerMeasure prior likelihood bestAction hbestAction).map canonicalSamplerHistory ⊗ₘ canonicalActionKernel prior likelihood bestAction hbestAction","missing":[],"search":"canonicalsampler_history_action_map_eq banditrlproof.thompson.canonicalsampler_history_action_map_eq the history/action marginal is generated by the canonical action kernel. theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_actionKernel","label":"canonicalSampler_condDistrib_action_ae_eq_actionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_actionKernel","description":"The constructed sampler has the intended next-action conditional law.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-342a0ff7bbcd","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2376,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:242"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSampler_condDistrib_action_ae_eq_actionKernel {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : ProbabilityTheory.condDistrib canonicalSamplerAction canonicalSamplerHistory (canonicalSamplerMeasure prior likelihood bestAction hbestAction) =ᵐ[ (canonicalSamplerMeasure prior likelihood bestAction hbestAction).map canonicalSamplerHistory] canonicalActionKernel prior likelihood bestAction hbestAction","missing":[],"search":"canonicalsampler_conddistrib_action_ae_eq_actionkernel banditrlproof.thompson.canonicalsampler_conddistrib_action_ae_eq_actionkernel the constructed sampler has the intended next-action conditional law. theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_bestAction","label":"canonicalSampler_condDistrib_action_ae_eq_bestAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_bestAction","description":"Premise-free one-step Thompson probability matching for the canonical sampler. Both law premises of the generic theorem are discharged by the constructed composition-product measure: its environment/history marginal is the canonical Bayesian joint law, and its history/action marginal is generated by the mapped canonical posterior.","url":"../modules/banditrlproof-algorithms-thompsoncanonicalsampler/index.html#decl-946d27e10ae5","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","order":2377,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalSampler"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalSampler.lean:271"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalSampler_condDistrib_action_ae_eq_bestAction {History : Type u} {Env : Type v} {Action : Type w} [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : ProbabilityTheory.condDistrib canonicalSamplerAction canonicalSamplerHistory (canonicalSamplerMeasure prior likelihood bestAction hbestAction) =ᵐ[ (canonicalSamplerMeasure prior likelihood bestAction hbestAction).map canonicalSamplerHistory] ProbabilityTheory.condDistrib (bestAction ∘ canonicalSamplerEnv) canonicalSamplerHistory (canonicalSamplerMeasure prior likelihood bestAction hbestAction)","missing":[],"search":"canonicalsampler_conddistrib_action_ae_eq_bestaction banditrlproof.thompson.canonicalsampler_conddistrib_action_ae_eq_bestaction premise-free one-step thompson probability matching for the canonical sampler. both law premises of the generic theorem are discharged by the constructed composition-product measure: its environment/history marginal is the canonical bayesian joint law, and its history/action marginal is generated by the mapped canonical posterior. theorem compiled","shard":"modules/7a9bfcad3c6d2758.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectoryMeasure","label":"canonicalHistoryTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryTrajectoryMeasure","description":"The canonical Ionescu-Tulcea law of the observable action/reward pairs.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-ddfe855c839e","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2378,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:21"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryTrajectoryMeasure {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : Measure ((n : Nat) -> Action × Reward)","missing":[],"search":"canonicalhistorytrajectorymeasure banditrlproof.thompson.canonicalhistorytrajectorymeasure the canonical ionescu-tulcea law of the observable action/reward pairs. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectoryAction","label":"canonicalHistoryTrajectoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryTrajectoryAction","description":"Action trace projected from the canonical pair trajectory.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-0559bfad03b0","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2379,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:42"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectoryAction {Action : Type u} {Reward : Type v} : ((n : Nat) -> Action × Reward) -> ActionTrace Action","missing":[],"search":"canonicalhistorytrajectoryaction banditrlproof.thompson.canonicalhistorytrajectoryaction action trace projected from the canonical pair trajectory. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectoryReward","label":"canonicalHistoryTrajectoryReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryTrajectoryReward","description":"Reward trace projected from the canonical pair trajectory.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-d1201f3bd76a","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2380,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:48"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectoryReward {Action : Type u} {Reward : Type v} : ((n : Nat) -> Action × Reward) -> RewardTrace Reward","missing":[],"search":"canonicalhistorytrajectoryreward banditrlproof.thompson.canonicalhistorytrajectoryreward reward trace projected from the canonical pair trajectory. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_canonicalHistoryTrajectoryAction_apply","label":"measurable_canonicalHistoryTrajectoryAction_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_canonicalHistoryTrajectoryAction_apply","description":"theorem measurable_canonicalHistoryTrajectoryAction_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun trajectory : (k : Nat) -> Action × Reward => canonicalHistoryTrajectoryAction trajectory n)","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-dafb496d07f2","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2381,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:53"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalHistoryTrajectoryAction_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun trajectory : (k : Nat) -> Action × Reward => canonicalHistoryTrajectoryAction trajectory n)","missing":[],"search":"measurable_canonicalhistorytrajectoryaction_apply banditrlproof.thompson.measurable_canonicalhistorytrajectoryaction_apply theorem measurable_canonicalhistorytrajectoryaction_apply {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] (n : nat) : measurable (fun trajectory : (k : nat) -> action × reward => canonicalhistorytrajectoryaction trajectory n) theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_canonicalHistoryTrajectoryReward_apply","label":"measurable_canonicalHistoryTrajectoryReward_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_canonicalHistoryTrajectoryReward_apply","description":"theorem measurable_canonicalHistoryTrajectoryReward_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun trajectory : (k : Nat) -> Action × Reward => canonicalHistoryTrajectoryReward trajectory n)","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-85b7fa795fce","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2382,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:60"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalHistoryTrajectoryReward_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun trajectory : (k : Nat) -> Action × Reward => canonicalHistoryTrajectoryReward trajectory n)","missing":[],"search":"measurable_canonicalhistorytrajectoryreward_apply banditrlproof.thompson.measurable_canonicalhistorytrajectoryreward_apply theorem measurable_canonicalhistorytrajectoryreward_apply {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] (n : nat) : measurable (fun trajectory : (k : nat) -> action × reward => canonicalhistorytrajectoryreward trajectory n) theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectory_initialPair_map_eq","label":"canonicalHistoryTrajectory_initialPair_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryTrajectory_initialPair_map_eq","description":"theorem canonicalHistoryTrajectory_initialPair_map_eq {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : Measure.map (fun trajectory => (canonicalHistoryTrajectoryAction trajectory 0, canonicalHistoryTrajectoryReward trajectory 0)) (canonicalHistoryTrajectoryMeasure algorithm environment…","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-439a6ce45837","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2383,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:67"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_initialPair_map_eq {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : Measure.map (fun trajectory => (canonicalHistoryTrajectoryAction trajectory 0, canonicalHistoryTrajectoryReward trajectory 0)) (canonicalHistoryTrajectoryMeasure algorithm environment) = algorithm.initialAction ⊗ₘ environment.initialFeedback","missing":[],"search":"canonicalhistorytrajectory_initialpair_map_eq banditrlproof.thompson.canonicalhistorytrajectory_initialpair_map_eq theorem canonicalhistorytrajectory_initialpair_map_eq {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] (algorithm : historyalgorithm action reward) (environment : historyenvironment action reward) : measure.map (fun trajectory => (canonicalhistorytrajectoryaction trajectory 0, canonicalhistorytrajectoryreward trajectory 0)) (canonicalhistorytrajectorymeasure algorithm environment) = algorithm.initialaction ⊗ₘ environment.initialfeedback theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectory_step_condDistrib","label":"canonicalHistoryTrajectory_step_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryTrajectory_step_condDistrib","description":"theorem canonicalHistoryTrajectory_step_condDistrib {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory => (canonicalHistoryTrajectoryAction…","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-fd81fda3d409","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2384,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:92"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_step_condDistrib {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory => (canonicalHistoryTrajectoryAction trajectory (n + 1), canonicalHistoryTrajectoryReward trajectory (n + 1))) (fun trajectory => History.finitePairHistoryOfTrace (canonicalHistoryTrajectoryAction trajectory) (canonicalHistoryTrajectoryReward trajectory) n) (canonicalHistoryTrajectoryMeasure algorithm environment) =ᵐ[ (canonicalHistoryTrajectoryMeasure algorithm environment).map (fun trajectory => History.finitePairHistoryOfTrace (canonicalHistoryTrajectoryAction trajectory) (canonicalHistoryTrajectoryReward tra…","missing":[],"search":"canonicalhistorytrajectory_step_conddistrib banditrlproof.thompson.canonicalhistorytrajectory_step_conddistrib theorem canonicalhistorytrajectory_step_conddistrib {action : type u} {reward : type v} [measurablespace action] [standardborelspace action] [nonempty action] [measurablespace reward] [standardborelspace reward] [nonempty reward] (algorithm : historyalgorithm action reward) (environment : historyenvironment action reward) (n : nat) : probabilitytheory.conddistrib (fun trajectory => (canonicalhistorytrajectoryaction trajectory (n + 1), canonicalhistorytrajectoryreward trajectory (n + 1))) (fun trajectory => history.finitepairhistoryoftrace (canonicalhistorytrajectoryaction trajectory) (canonicalhistorytrajectoryreward trajectory) n) (canonicalhistorytrajectorymeasure algorithm environment) =ᵐ[ (canonicalhistorytrajectorymeasure algorithm environment).map (fun trajectory => history.finitepairhistoryoftrace (canonicalhistorytrajectoryaction trajectory) (canonicalhistorytrajectoryreward trajectory) n)] historystepkernel algorithm environment n theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSequence","label":"canonicalHistoryAlgorithmEnvironmentSequence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSequence","description":"The canonical pair trajectory satisfies the combined history-process contract without any externally supplied law premise.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-44c5faf410a9","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2385,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:136"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryAlgorithmEnvironmentSequence {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : IsHistoryAlgorithmEnvironmentSequence (canonicalHistoryTrajectoryMeasure algorithm environment) canonicalHistoryTrajectoryAction canonicalHistoryTrajectoryReward algorithm environment where","missing":[],"search":"canonicalhistoryalgorithmenvironmentsequence banditrlproof.thompson.canonicalhistoryalgorithmenvironmentsequence the canonical pair trajectory satisfies the combined history-process contract without any externally supplied law premise. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.initialAction_map_eq_of_historyAlgorithmEnvironmentSequence","label":"initialAction_map_eq_of_historyAlgorithmEnvironmentSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.initialAction_map_eq_of_historyAlgorithmEnvironmentSequence","description":"The combined initial pair law determines the initial action marginal.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-ac9f89580bdd","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2386,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:154"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem initialAction_map_eq_of_historyAlgorithmEnvironmentSequence {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) : mu.map (fun omega => action omega 0) = algorithm.initialAction","missing":[],"search":"initialaction_map_eq_of_historyalgorithmenvironmentsequence banditrlproof.thompson.initialaction_map_eq_of_historyalgorithmenvironmentsequence the combined initial pair law determines the initial action marginal. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.initialFeedback_condDistrib_of_historyAlgorithmEnvironmentSequence","label":"initialFeedback_condDistrib_of_historyAlgorithmEnvironmentSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.initialFeedback_condDistrib_of_historyAlgorithmEnvironmentSequence","description":"The combined initial pair law determines the initial feedback conditional law.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-38c40f0c7749","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2387,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:180"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem initialFeedback_condDistrib_of_historyAlgorithmEnvironmentSequence {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) : ProbabilityTheory.condDistrib (fun omega => reward omega 0) (fun omega => action omega 0) mu =ᵐ[ mu.map (fun omega => action omega 0)] environment.initialFeedback","missing":[],"search":"initialfeedback_conddistrib_of_historyalgorithmenvironmentsequence banditrlproof.thompson.initialfeedback_conddistrib_of_historyalgorithmenvironmentsequence the combined initial pair law determines the initial feedback conditional law. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyStepKernel_map_fst","label":"historyStepKernel_map_fst","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyStepKernel_map_fst","description":"The action marginal of a history step kernel is its policy kernel.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-8768c56178fa","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2388,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:202"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernel_map_fst {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (n : Nat) : (historyStepKernel algorithm environment n).map Prod.fst = algorithm.policy n","missing":[],"search":"historystepkernel_map_fst banditrlproof.thompson.historystepkernel_map_fst the action marginal of a history step kernel is its policy kernel. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policy_condDistrib_of_historyAlgorithmEnvironmentSequence","label":"policy_condDistrib_of_historyAlgorithmEnvironmentSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policy_condDistrib_of_historyAlgorithmEnvironmentSequence","description":"The combined successor pair law determines the successor action policy.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-e785f49c60c6","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2389,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:213"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policy_condDistrib_of_historyAlgorithmEnvironmentSequence {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) (n : Nat) : ProbabilityTheory.condDistrib (fun omega => action omega (n + 1)) (fun omega => History.finitePairHistoryOfTrace (action omega) (reward omega) n) mu =ᵐ[ mu.map (fun omega => History.finitePairHistoryOfTrace (action omega) (reward omega) n)] algorithm.policy n","missing":[],"search":"policy_conddistrib_of_historyalgorithmenvironmentsequence banditrlproof.thompson.policy_conddistrib_of_historyalgorithmenvironmentsequence the combined successor pair law determines the successor action policy. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.feedback_condDistrib_of_historyAlgorithmEnvironmentSequence","label":"feedback_condDistrib_of_historyAlgorithmEnvironmentSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.feedback_condDistrib_of_historyAlgorithmEnvironmentSequence","description":"The combined successor pair law also determines the feedback conditional law given the finite history and the newly sampled action.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-bc168284eada","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2390,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:262"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem feedback_condDistrib_of_historyAlgorithmEnvironmentSequence {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) (n : Nat) : ProbabilityTheory.condDistrib (fun omega => reward omega (n + 1)) (fun omega => (History.finitePairHistoryOfTrace (action omega) (reward omega) n, action omega (n + 1))) mu =ᵐ[ mu.map (fun omega => (History.finitePairHistoryOfTrace (action omega) (reward omega) n, action omega (n + 1)))] environme…","missing":[],"search":"feedback_conddistrib_of_historyalgorithmenvironmentsequence banditrlproof.thompson.feedback_conddistrib_of_historyalgorithmenvironmentsequence the combined successor pair law also determines the feedback conditional law given the finite history and the newly sampled action. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryAlgorithmEnvironmentSplitSource","label":"HistoryAlgorithmEnvironmentSplitSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryAlgorithmEnvironmentSplitSource","description":"Split action/feedback laws for one fixed history environment.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-95333108f2ee","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2391,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:339"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure HistoryAlgorithmEnvironmentSplitSource {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : Prop where","missing":[],"search":"historyalgorithmenvironmentsplitsource banditrlproof.thompson.historyalgorithmenvironmentsplitsource split action/feedback laws for one fixed history environment. structure compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.HistoryAlgorithmEnvironmentSplitSource.toSequence","label":"toSequence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.HistoryAlgorithmEnvironmentSplitSource.toSequence","description":"Assemble the combined process contract from a fixed-environment split source.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-e46097bc7d13","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2392,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:374"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def HistoryAlgorithmEnvironmentSplitSource.toSequence {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (source : HistoryAlgorithmEnvironmentSplitSource mu action reward algorithm environment) : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment","missing":[],"search":"tosequence banditrlproof.thompson.historyalgorithmenvironmentsplitsource.tosequence assemble the combined process contract from a fixed-environment split source. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSplitSource","label":"canonicalHistoryAlgorithmEnvironmentSplitSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSplitSource","description":"The canonical trajectory realizes all four fixed-environment split laws.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-9d794e2ac76a","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2393,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:395"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryAlgorithmEnvironmentSplitSource {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : HistoryAlgorithmEnvironmentSplitSource (canonicalHistoryTrajectoryMeasure algorithm environment) canonicalHistoryTrajectoryAction canonicalHistoryTrajectoryReward algorithm environment where","missing":[],"search":"canonicalhistoryalgorithmenvironmentsplitsource banditrlproof.thompson.canonicalhistoryalgorithmenvironmentsplitsource the canonical trajectory realizes all four fixed-environment split laws. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSequence_of_split","label":"canonicalHistoryAlgorithmEnvironmentSequence_of_split","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSequence_of_split","description":"The canonical combined process reconstructed specifically from its split laws.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-9ad0555ebc6f","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2394,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:433"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryAlgorithmEnvironmentSequence_of_split {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) : IsHistoryAlgorithmEnvironmentSequence (canonicalHistoryTrajectoryMeasure algorithm environment) canonicalHistoryTrajectoryAction canonicalHistoryTrajectoryReward algorithm environment","missing":[],"search":"canonicalhistoryalgorithmenvironmentsequence_of_split banditrlproof.thompson.canonicalhistoryalgorithmenvironmentsequence_of_split the canonical combined process reconstructed specifically from its split laws. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.kernelWithInput","label":"kernelWithInput","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.kernelWithInput","description":"Lift a kernel to samples that retain the kernel input as their first coordinate.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-15a6407641e6","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2395,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:449"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def kernelWithInput {Input : Type u} {Output : Type v} [MeasurableSpace Input] [MeasurableSpace Output] (kernel : ProbabilityTheory.Kernel Input Output) : ProbabilityTheory.Kernel Input (Input × Output)","missing":[],"search":"kernelwithinput banditrlproof.thompson.kernelwithinput lift a kernel to samples that retain the kernel input as their first coordinate. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.kernelWithInput_apply","label":"kernelWithInput_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.kernelWithInput_apply","description":"theorem kernelWithInput_apply {Input : Type u} {Output : Type v} [MeasurableSpace Input] [MeasurableSingletonClass Input] [MeasurableSpace Output] (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (input : Input) : kernelWithInput kernel input = (kernel input).map (Prod.mk input)","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-9a5b7b569979","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2396,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:466"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem kernelWithInput_apply {Input : Type u} {Output : Type v} [MeasurableSpace Input] [MeasurableSingletonClass Input] [MeasurableSpace Output] (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (input : Input) : kernelWithInput kernel input = (kernel input).map (Prod.mk input)","missing":[],"search":"kernelwithinput_apply banditrlproof.thompson.kernelwithinput_apply theorem kernelwithinput_apply {input : type u} {output : type v} [measurablespace input] [measurablesingletonclass input] [measurablespace output] (kernel : probabilitytheory.kernel input output) [probabilitytheory.ismarkovkernel kernel] (input : input) : kernelwithinput kernel input = (kernel input).map (prod.mk input) theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.condDistrib_id_fst_compProd_ae_eq_kernelWithInput","label":"condDistrib_id_fst_compProd_ae_eq_kernelWithInput","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.condDistrib_id_fst_compProd_ae_eq_kernelWithInput","description":"For a composition-product sample, conditioning the complete sample on its first coordinate keeps that coordinate and uses the supplied second-coordinate kernel.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-4d4f87b0ea0c","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2397,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:487"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_id_fst_compProd_ae_eq_kernelWithInput {Input : Type u} {Output : Type v} [MeasurableSpace Input] [StandardBorelSpace Input] [Nonempty Input] [MeasurableSpace Output] [StandardBorelSpace Output] [Nonempty Output] (prior : Measure Input) [IsFiniteMeasure prior] (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] : ProbabilityTheory.condDistrib id Prod.fst (prior ⊗ₘ kernel) =ᵐ[prior] kernelWithInput kernel","missing":[],"search":"conddistrib_id_fst_compprod_ae_eq_kernelwithinput banditrlproof.thompson.conddistrib_id_fst_compprod_ae_eq_kernelwithinput for a composition-product sample, conditioning the complete sample on its first coordinate keeps that coordinate and uses the supplied second-coordinate kernel. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.environmentTrajectoryAction","label":"environmentTrajectoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.environmentTrajectoryAction","description":"Action trace of an environment/trajectory sample.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-10a649701674","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2398,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:531"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def environmentTrajectoryAction {Env : Type u} {Action : Type v} {Reward : Type w} : (Env × ((n : Nat) -> Action × Reward)) -> ActionTrace Action","missing":[],"search":"environmenttrajectoryaction banditrlproof.thompson.environmenttrajectoryaction action trace of an environment/trajectory sample. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.environmentTrajectoryReward","label":"environmentTrajectoryReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.environmentTrajectoryReward","description":"Reward trace of an environment/trajectory sample.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-2ff66059dfab","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2399,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:537"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def environmentTrajectoryReward {Env : Type u} {Action : Type v} {Reward : Type w} : (Env × ((n : Nat) -> Action × Reward)) -> RewardTrace Reward","missing":[],"search":"environmenttrajectoryreward banditrlproof.thompson.environmenttrajectoryreward reward trace of an environment/trajectory sample. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_environmentTrajectoryAction_apply","label":"measurable_environmentTrajectoryAction_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_environmentTrajectoryAction_apply","description":"theorem measurable_environmentTrajectoryAction_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Action × Reward) => environmentTrajectoryAction sample n)","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-ba623ce18342","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2400,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:542"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_environmentTrajectoryAction_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Action × Reward) => environmentTrajectoryAction sample n)","missing":[],"search":"measurable_environmenttrajectoryaction_apply banditrlproof.thompson.measurable_environmenttrajectoryaction_apply theorem measurable_environmenttrajectoryaction_apply {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] (n : nat) : measurable (fun sample : env × ((k : nat) -> action × reward) => environmenttrajectoryaction sample n) theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_environmentTrajectoryReward_apply","label":"measurable_environmentTrajectoryReward_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_environmentTrajectoryReward_apply","description":"theorem measurable_environmentTrajectoryReward_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Action × Reward) => environmentTrajectoryReward sample n)","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-b265d2559afb","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2401,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:550"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_environmentTrajectoryReward_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Action × Reward) => environmentTrajectoryReward sample n)","missing":[],"search":"measurable_environmenttrajectoryreward_apply banditrlproof.thompson.measurable_environmenttrajectoryreward_apply theorem measurable_environmenttrajectoryreward_apply {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] (n : nat) : measurable (fun sample : env × ((k : nat) -> action × reward) => environmenttrajectoryreward sample n) theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyAlgorithmEnvironmentSequence_of_measure_eq","label":"historyAlgorithmEnvironmentSequence_of_measure_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.historyAlgorithmEnvironmentSequence_of_measure_eq","description":"Transport a history-process contract across an equality of source measures.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-34be140020d7","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2402,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:559"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def historyAlgorithmEnvironmentSequence_of_measure_eq {Omega : Type w} {Action : Type u} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu nu : Measure Omega) [IsFiniteMeasure mu] [IsFiniteMeasure nu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (algorithm : HistoryAlgorithm Action Reward) (environment : HistoryEnvironment Action Reward) (hmeasure : mu = nu) (source : IsHistoryAlgorithmEnvironmentSequence mu action reward algorithm environment) : IsHistoryAlgorithmEnvironmentSequence nu action reward algorithm environment","missing":[],"search":"historyalgorithmenvironmentsequence_of_measure_eq banditrlproof.thompson.historyalgorithmenvironmentsequence_of_measure_eq transport a history-process contract across an equality of source measures. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.mappedCanonicalHistoryAlgorithmEnvironmentSequence","label":"mappedCanonicalHistoryAlgorithmEnvironmentSequence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.mappedCanonicalHistoryAlgorithmEnvironmentSequence","description":"Retaining a fixed environment coordinate preserves the canonical process.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-6b367fb341bf","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2403,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:578"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def mappedCanonicalHistoryAlgorithmEnvironmentSequence {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) (environment : Env) : IsHistoryAlgorithmEnvironmentSequence ((canonicalHistoryTrajectoryMeasure algorithm (feedbackEnvironment environment)).map (Prod.mk environment)) environmentTrajectoryAction environmentTrajectoryReward algorithm (feedbackEnvironment environment) where","missing":[],"search":"mappedcanonicalhistoryalgorithmenvironmentsequence banditrlproof.thompson.mappedcanonicalhistoryalgorithmenvironmentsequence retaining a fixed environment coordinate preserves the canonical process. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.kernelWithInputHistoryAlgorithmEnvironmentSequence","label":"kernelWithInputHistoryAlgorithmEnvironmentSequence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.kernelWithInputHistoryAlgorithmEnvironmentSequence","description":"A trajectory kernel whose value is the canonical fixed-environment law yields the combined history-process contract after retaining its environment input.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-372ff3e3eb2a","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2404,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:686"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def kernelWithInputHistoryAlgorithmEnvironmentSequence {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) (trajectoryKernel : ProbabilityTheory.Kernel Env ((n : Nat) -> Action × Reward)) [ProbabilityTheory.IsMarkovKernel trajectoryKernel] (environment : Env) (hkernel : trajectoryKernel environment = canonicalHistoryTrajectoryMeasure algorithm (feedbackEnvironment environment)) : IsHistoryAlgorithmEnvironmentSequence (kernelWithInput trajectoryKernel environment) environmentTrajectoryAction environmentTrajectoryReward algorithm (feedbackEnvironment environment)","missing":[],"search":"kernelwithinputhistoryalgorithmenvironmentsequence banditrlproof.thompson.kernelwithinputhistoryalgorithmenvironmentsequence a trajectory kernel whose value is the canonical fixed-environment law yields the combined history-process contract after retaining its environment input. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmEnvironmentSequence_of_canonicalTrajectoryKernel","label":"conditionalHistoryAlgorithmEnvironmentSequence_of_canonicalTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.conditionalHistoryAlgorithmEnvironmentSequence_of_canonicalTrajectoryKernel","description":"The regular conditional complete-sample law of a canonical environment/ trajectory composition product satisfies the history-process contract.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-bfe85d409ae2","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2405,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:722"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem conditionalHistoryAlgorithmEnvironmentSequence_of_canonicalTrajectoryKernel {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) (trajectoryKernel : ProbabilityTheory.Kernel Env ((n : Nat) -> Action × Reward)) [ProbabilityTheory.IsMarkovKernel trajectoryKernel] (hkernel : forall environment, trajectoryKernel environment = canonicalHistoryTrajectoryMeasure algorithm (feedbackEnvironment environment)) : ∀ᵐ environment ∂(prior ⊗ₘ trajectoryKernel).map Prod.fst, IsHistoryAlgorithmEnvironmentSequence (ProbabilityTheory.condDistrib id…","missing":[],"search":"conditionalhistoryalgorithmenvironmentsequence_of_canonicaltrajectorykernel banditrlproof.thompson.conditionalhistoryalgorithmenvironmentsequence_of_canonicaltrajectorykernel the regular conditional complete-sample law of a canonical environment/ trajectory composition product satisfies the history-process contract. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmEnvironmentSplitSource_of_canonicalTrajectoryKernel","label":"conditionalHistoryAlgorithmEnvironmentSplitSource_of_canonicalTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.conditionalHistoryAlgorithmEnvironmentSplitSource_of_canonicalTrajectoryKernel","description":"Canonical environment-indexed trajectory kernels supply all four conditional split law families required by the Thompson density route.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-501a27433c07","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2406,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:764"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def conditionalHistoryAlgorithmEnvironmentSplitSource_of_canonicalTrajectoryKernel {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) (trajectoryKernel : ProbabilityTheory.Kernel Env ((n : Nat) -> Action × Reward)) [ProbabilityTheory.IsMarkovKernel trajectoryKernel] (hkernel : forall environment, trajectoryKernel environment = canonicalHistoryTrajectoryMeasure algorithm (feedbackEnvironment environment)) : ConditionalHistoryAlgorithmEnvironmentSplitSource (prior ⊗ₘ trajectoryKernel) Prod.fst environmentTrajectoryAction e…","missing":[],"search":"conditionalhistoryalgorithmenvironmentsplitsource_of_canonicaltrajectorykernel banditrlproof.thompson.conditionalhistoryalgorithmenvironmentsplitsource_of_canonicaltrajectorykernel canonical environment-indexed trajectory kernels supply all four conditional split law families required by the thompson density route. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmDensitySplitSource_of_canonicalTrajectoryKernels","label":"conditionalHistoryAlgorithmDensitySplitSource_of_canonicalTrajectoryKernels","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.conditionalHistoryAlgorithmDensitySplitSource_of_canonicalTrajectoryKernels","description":"Paired canonical actual/reference trajectory kernels construct the complete conditional split source consumed by algorithm-density transport.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-79c29b9ed7ba","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2407,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:831"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def conditionalHistoryAlgorithmDensitySplitSource_of_canonicalTrajectoryKernels {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) (trajectoryKernel referenceTrajectoryKernel : ProbabilityTheory.Kernel Env ((n : Nat) -> Action × Reward)) [ProbabilityTheory.IsMarkovKernel trajectoryKernel] [ProbabilityTheory.IsMarkovKernel referenceTrajectoryKernel] (htrajectoryKernel : forall environment, trajectoryKernel environment = canonicalHistoryTrajectoryMeasure algorithm (feedbackEnvironment environment)) (href…","missing":[],"search":"conditionalhistoryalgorithmdensitysplitsource_of_canonicaltrajectorykernels banditrlproof.thompson.conditionalhistoryalgorithmdensitysplitsource_of_canonicaltrajectorykernels paired canonical actual/reference trajectory kernels construct the complete conditional split source consumed by algorithm-density transport. definition compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_canonicalTrajectoryKernels","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_canonicalTrajectoryKernels","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_canonicalTrajectoryKernels","description":"Finite-prefix Thompson probability matching for paired canonical recursive trajectory kernels. The remaining producer obligation is exactly the measurable kernel family whose values are the fixed-environment canonical `trajMeasure`s.","url":"../modules/banditrlproof-algorithms-thompsoncanonicaltrajectory/index.html#decl-8324d4d2291e","parent":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","order":2408,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonCanonicalTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonCanonicalTrajectory.lean:878"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_canonicalTrajectoryKernels {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (feedbackEnvironment : Env -> HistoryEnvironment Action Reward) (trajectoryKernel referenceTrajectoryKernel : ProbabilityTheory.Kernel Env ((n : Nat) -> Action × Reward)) [ProbabilityTheory.IsMarkovKernel trajectoryKernel] [ProbabilityTheory.IsMarkovKernel referenceTrajectoryKernel] (htrajectoryKernel : forall environment, trajectoryKernel environment = canonicalHistoryTrajectoryMeasure algorithm (feedbackEnvironment enviro…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_canonicaltrajectorykernels banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_canonicaltrajectorykernels finite-prefix thompson probability matching for paired canonical recursive trajectory kernels. the remaining producer obligation is exactly the measurable kernel family whose values are the fixed-environment canonical `trajmeasure`s. theorem compiled","shard":"modules/4d6cbf6acc86ef2a.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCB","label":"clippedUCB","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCB","description":"The clipped UCB score on a complete action/reward trace.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-1923b536baef","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2409,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:24"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def clippedUCB {K : Nat} (l u sigma2 delta : Real) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (t : Nat) : Real","missing":[],"search":"clippeducb banditrlproof.thompson.clippeducb the clipped ucb score on a complete action/reward trace. definition compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCBHistory","label":"clippedUCBHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCBHistory","description":"The same clipped score on the inclusive history through time `n`.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-b2268a491820","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2410,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:37"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def clippedUCBHistory {K : Nat} (l u sigma2 delta : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Real","missing":[],"search":"clippeducbhistory banditrlproof.thompson.clippeducbhistory the same clipped score on the inclusive history through time `n`. definition compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCB_zero","label":"clippedUCB_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCB_zero","description":"theorem clippedUCB_zero {K : Nat} (l u sigma2 delta : Real) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) : clippedUCB l u sigma2 delta action reward arm 0 = u","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-8ec9a62e471e","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2411,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:50"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedUCB_zero {K : Nat} (l u sigma2 delta : Real) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) : clippedUCB l u sigma2 delta action reward arm 0 = u","missing":[],"search":"clippeducb_zero banditrlproof.thompson.clippeducb_zero theorem clippeducb_zero {k : nat} (l u sigma2 delta : real) (action : actiontrace (fin k)) (reward : rewardtrace real) (arm : fin k) : clippeducb l u sigma2 delta action reward arm 0 = u theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCB_mem_Icc","label":"clippedUCB_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCB_mem_Icc","description":"theorem clippedUCB_mem_Icc {K : Nat} (l u sigma2 delta : Real) (hlu : l <= u) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (t : Nat) : clippedUCB l u sigma2 delta action reward arm t ∈ Set.Icc l u","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-8c851e912230","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2412,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:57"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedUCB_mem_Icc {K : Nat} (l u sigma2 delta : Real) (hlu : l <= u) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (t : Nat) : clippedUCB l u sigma2 delta action reward arm t ∈ Set.Icc l u","missing":[],"search":"clippeducb_mem_icc banditrlproof.thompson.clippeducb_mem_icc theorem clippeducb_mem_icc {k : nat} (l u sigma2 delta : real) (hlu : l <= u) (action : actiontrace (fin k)) (reward : rewardtrace real) (arm : fin k) (t : nat) : clippeducb l u sigma2 delta action reward arm t ∈ set.icc l u theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finset_sum_sqrt_le","label":"finset_sum_sqrt_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finset_sum_sqrt_le","description":"private theorem finset_sum_sqrt_le {ι : Type*} (s : Finset ι) (c : ι -> Real) (hc : forall i, 0 <= c i) : ∑ i ∈ s, Real.sqrt (c i) <= Real.sqrt (s.card * ∑ i ∈ s, c i)","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-fc3fc350842e","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2413,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:68"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"private theorem finset_sum_sqrt_le {ι : Type*} (s : Finset ι) (c : ι -> Real) (hc : forall i, 0 <= c i) : ∑ i ∈ s, Real.sqrt (c i) <= Real.sqrt (s.card * ∑ i ∈ s, c i)","missing":[],"search":"finset_sum_sqrt_le banditrlproof.thompson.finset_sum_sqrt_le private theorem finset_sum_sqrt_le {ι : type*} (s : finset ι) (c : ι -> real) (hc : forall i, 0 <= c i) : ∑ i ∈ s, real.sqrt (c i) <= real.sqrt (s.card * ∑ i ∈ s, c i) theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finset_sum_one_div_sqrt_le","label":"finset_sum_one_div_sqrt_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finset_sum_one_div_sqrt_le","description":"private theorem finset_sum_one_div_sqrt_le {n : Nat} (hn : 0 < n) : ∑ k ∈ Finset.range (n + 1), 1 / Real.sqrt k <= 2 * Real.sqrt n - 1","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-694639d50535","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2414,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:77"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"private theorem finset_sum_one_div_sqrt_le {n : Nat} (hn : 0 < n) : ∑ k ∈ Finset.range (n + 1), 1 / Real.sqrt k <= 2 * Real.sqrt n - 1","missing":[],"search":"finset_sum_one_div_sqrt_le banditrlproof.thompson.finset_sum_one_div_sqrt_le private theorem finset_sum_one_div_sqrt_le {n : nat} (hn : 0 < n) : ∑ k ∈ finset.range (n + 1), 1 / real.sqrt k <= 2 * real.sqrt n - 1 theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.sum_clippedUCB_action_sub_mean_le","label":"sum_clippedUCB_action_sub_mean_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.sum_clippedUCB_action_sub_mean_le","description":"Pathwise selected-action clipped-UCB excess bound from the pinned LML Thompson route. The only hypothesis is that every positive-count empirical mean is strictly below its arm mean plus the confidence width.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-6cc39e90186e","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2415,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:106"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem sum_clippedUCB_action_sub_mean_le {K : Nat} [NeZero K] (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (mean : Fin K -> Real) (l u sigma2 delta : Real) (hmeanMem : forall arm, mean arm ∈ Set.Icc l u) (hlu : l <= u) (hgood : forall s, s < n -> pullCount action (action s) s ≠ 0 -> UCB.realEmpiricalMean action reward (action s) s - mean (action s) < Real.sqrt (2 * sigma2 * Real.log (1 / delta) / (pullCount action (action s) s : Real))) : ∑ s ∈ Finset.range n, (clippedUCB l u sigma2 delta action reward (action s) s - mean (action s)) <= (u - l) * K + 4 * Real.sqrt (2 * sigma2 * Real.log (1 / delta) * K * n)","missing":[],"search":"sum_clippeducb_action_sub_mean_le banditrlproof.thompson.sum_clippeducb_action_sub_mean_le pathwise selected-action clipped-ucb excess bound from the pinned lml thompson route. the only hypothesis is that every positive-count empirical mean is strictly below its arm mean plus the confidence width. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCBHistory_mem_Icc","label":"clippedUCBHistory_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCBHistory_mem_Icc","description":"theorem clippedUCBHistory_mem_Icc {K : Nat} (l u sigma2 delta : Real) (hlu : l <= u) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : clippedUCBHistory l u sigma2 delta n history arm ∈ Set.Icc l u","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-21cc6b3a0d28","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2416,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:237"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedUCBHistory_mem_Icc {K : Nat} (l u sigma2 delta : Real) (hlu : l <= u) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : clippedUCBHistory l u sigma2 delta n history arm ∈ Set.Icc l u","missing":[],"search":"clippeducbhistory_mem_icc banditrlproof.thompson.clippeducbhistory_mem_icc theorem clippeducbhistory_mem_icc {k : nat} (l u sigma2 delta : real) (hlu : l <= u) (n : nat) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : clippeducbhistory l u sigma2 delta n history arm ∈ set.icc l u theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_clippedUCBHistory","label":"measurable_clippedUCBHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_clippedUCBHistory","description":"A fixed-arm clipped score is measurable on inclusive pair histories.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-6670e85e609d","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2417,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:249"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_clippedUCBHistory {K : Nat} (l u sigma2 delta : Real) (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => clippedUCBHistory l u sigma2 delta n history arm)","missing":[],"search":"measurable_clippeducbhistory banditrlproof.thompson.measurable_clippeducbhistory a fixed-arm clipped score is measurable on inclusive pair histories. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_uncurry_clippedUCBHistory","label":"measurable_uncurry_clippedUCBHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_uncurry_clippedUCBHistory","description":"Joint measurability in the visible history and candidate action.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-96b2a68eb559","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2418,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:275"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_uncurry_clippedUCBHistory {K : Nat} (l u sigma2 delta : Real) (n : Nat) : Measurable (fun pair : History.FinitePairHistory (Fin K) Real n × Fin K => clippedUCBHistory l u sigma2 delta n pair.1 pair.2)","missing":[],"search":"measurable_uncurry_clippeducbhistory banditrlproof.thompson.measurable_uncurry_clippeducbhistory joint measurability in the visible history and candidate action. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCBHistoryScore","label":"clippedUCBHistoryScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCBHistoryScore","description":"`HistoryActionScore` instance for the pinned LML clipped-UCB formula.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-f83ac72246e2","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2419,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:285"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def clippedUCBHistoryScore {K : Nat} (l u sigma2 delta : Real) : HistoryActionScore (Fin K) Real where","missing":[],"search":"clippeducbhistoryscore banditrlproof.thompson.clippeducbhistoryscore `historyactionscore` instance for the pinned lml clipped-ucb formula. definition compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCBHistory_finitePairHistoryOfTrace","label":"clippedUCBHistory_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCBHistory_finitePairHistoryOfTrace","description":"Inclusive finite-history and exclusive trace versions agree at `n + 1`.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-a315a631c8eb","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2420,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:296"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedUCBHistory_finitePairHistoryOfTrace {K : Nat} (l u sigma2 delta : Real) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n : Nat) (arm : Fin K) : clippedUCBHistory l u sigma2 delta n (History.finitePairHistoryOfTrace action reward n) arm = clippedUCB l u sigma2 delta action reward arm (n + 1)","missing":[],"search":"clippeducbhistory_finitepairhistoryoftrace banditrlproof.thompson.clippeducbhistory_finitepairhistoryoftrace inclusive finite-history and exclusive trace versions agree at `n + 1`. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCBHistoryScore_atTrace","label":"clippedUCBHistoryScore_atTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCBHistoryScore_atTrace","description":"Evaluating the history score at the selected action recovers `clippedUCB`.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-fb5ac984899f","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2421,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:308"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedUCBHistoryScore_atTrace {K : Nat} (l u sigma2 delta : Real) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (t : Nat) : (clippedUCBHistoryScore l u sigma2 delta).atTrace action reward t = clippedUCB l u sigma2 delta action reward (action t) t","missing":[],"search":"clippeducbhistoryscore_attrace banditrlproof.thompson.clippeducbhistoryscore_attrace evaluating the history score at the selected action recovers `clippeducb`. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedUCBHistoryScore_atBestTrace","label":"clippedUCBHistoryScore_atBestTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedUCBHistoryScore_atBestTrace","description":"Evaluating at a comparison arm recovers the same trace score.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-976910f375de","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2422,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:320"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedUCBHistoryScore_atBestTrace {K : Nat} (l u sigma2 delta : Real) (bestArm : Fin K) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (t : Nat) : (clippedUCBHistoryScore l u sigma2 delta).atBestTrace bestArm action reward t = clippedUCB l u sigma2 delta action reward bestArm t","missing":[],"search":"clippeducbhistoryscore_atbesttrace banditrlproof.thompson.clippeducbhistoryscore_atbesttrace evaluating at a comparison arm recovers the same trace score. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integrable_of_measurable_mem_Icc","label":"integrable_of_measurable_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integrable_of_measurable_mem_Icc","description":"A measurable real function with pointwise range in `[l, u]` is integrable.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-67a42aad2928","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2423,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:334"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integrable_of_measurable_mem_Icc {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (f : Omega -> Real) (hf : Measurable f) {l u : Real} (hmem : forall omega, f omega ∈ Set.Icc l u) : Integrable f mu","missing":[],"search":"integrable_of_measurable_mem_icc banditrlproof.thompson.integrable_of_measurable_mem_icc a measurable real function with pointwise range in `[l, u]` is integrable. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integrable_trajectoryHistoryScore_clippedUCB","label":"integrable_trajectoryHistoryScore_clippedUCB","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integrable_trajectoryHistoryScore_clippedUCB","description":"theorem integrable_trajectoryHistoryScore_clippedUCB {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (l u sigma2 delta : Real) (hlu : l <= u) (t : Nat) : Integrable (fun sample => trajectoryHistoryScore (clippedUCBHistoryScore l u sigma2 delta) sample t) mu","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-b58693d758cc","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2424,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:347"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integrable_trajectoryHistoryScore_clippedUCB {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (l u sigma2 delta : Real) (hlu : l <= u) (t : Nat) : Integrable (fun sample => trajectoryHistoryScore (clippedUCBHistoryScore l u sigma2 delta) sample t) mu","missing":[],"search":"integrable_trajectoryhistoryscore_clippeducb banditrlproof.thompson.integrable_trajectoryhistoryscore_clippeducb theorem integrable_trajectoryhistoryscore_clippeducb {env : type u} {k : nat} [measurablespace env] (mu : measure (env × ((n : nat) -> fin k × real))) [isfinitemeasure mu] (l u sigma2 delta : real) (hlu : l <= u) (t : nat) : integrable (fun sample => trajectoryhistoryscore (clippeducbhistoryscore l u sigma2 delta) sample t) mu theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integrable_trajectoryBestHistoryScore_clippedUCB","label":"integrable_trajectoryBestHistoryScore_clippedUCB","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integrable_trajectoryBestHistoryScore_clippedUCB","description":"theorem integrable_trajectoryBestHistoryScore_clippedUCB {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) (l u sigma2 delta : Real) (hlu : l <= u) (t : Nat) : Integrable (fun sample => trajectoryBestHistoryScore (clippedUCBHistoryScore l u sigma2 delta) bestAction sample t) mu","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-d467c024853b","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2425,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:368"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integrable_trajectoryBestHistoryScore_clippedUCB {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) (l u sigma2 delta : Real) (hlu : l <= u) (t : Nat) : Integrable (fun sample => trajectoryBestHistoryScore (clippedUCBHistoryScore l u sigma2 delta) bestAction sample t) mu","missing":[],"search":"integrable_trajectorybesthistoryscore_clippeducb banditrlproof.thompson.integrable_trajectorybesthistoryscore_clippeducb theorem integrable_trajectorybesthistoryscore_clippeducb {env : type u} {k : nat} [measurablespace env] (mu : measure (env × ((n : nat) -> fin k × real))) [isfinitemeasure mu] (bestaction : env -> fin k) (hbestaction : measurable bestaction) (l u sigma2 delta : real) (hlu : l <= u) (t : nat) : integrable (fun sample => trajectorybesthistoryscore (clippeducbhistoryscore l u sigma2 delta) bestaction sample t) mu theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integrable_trajectoryMean_bestAction","label":"integrable_trajectoryMean_bestAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integrable_trajectoryMean_bestAction","description":"theorem integrable_trajectoryMean_bestAction {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (mean : Env -> Fin K -> Real) (hmean : Measurable (fun pair : Env × Fin K => mean pair.1 pair.2)) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) : Integrable (fun sample : Env × ((…","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-c1ed377822d1","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2426,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:391"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integrable_trajectoryMean_bestAction {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (mean : Env -> Fin K -> Real) (hmean : Measurable (fun pair : Env × Fin K => mean pair.1 pair.2)) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) : Integrable (fun sample : Env × ((n : Nat) -> Fin K × Real) => mean sample.1 (bestAction sample.1)) mu","missing":[],"search":"integrable_trajectorymean_bestaction banditrlproof.thompson.integrable_trajectorymean_bestaction theorem integrable_trajectorymean_bestaction {env : type u} {k : nat} [measurablespace env] (mu : measure (env × ((n : nat) -> fin k × real))) [isfinitemeasure mu] (mean : env -> fin k -> real) (hmean : measurable (fun pair : env × fin k => mean pair.1 pair.2)) (hmeanmem : forall env arm, mean env arm ∈ set.icc l u) (bestaction : env -> fin k) (hbestaction : measurable bestaction) : integrable (fun sample : env × ((n : nat) -> fin k × real) => mean sample.1 (bestaction sample.1)) mu theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integrable_trajectoryMean_action","label":"integrable_trajectoryMean_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integrable_trajectoryMean_action","description":"theorem integrable_trajectoryMean_action {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (mean : Env -> Fin K -> Real) (hmean : Measurable (fun pair : Env × Fin K => mean pair.1 pair.2)) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (t : Nat) : Integrable (fun sample : Env × ((n : Nat) -> Fin K × Real) => mean sample.1 (environmentTraje…","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-49b337dbf0a1","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2427,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:407"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integrable_trajectoryMean_action {Env : Type u} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((n : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (mean : Env -> Fin K -> Real) (hmean : Measurable (fun pair : Env × Fin K => mean pair.1 pair.2)) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (t : Nat) : Integrable (fun sample : Env × ((n : Nat) -> Fin K × Real) => mean sample.1 (environmentTrajectoryAction sample t)) mu","missing":[],"search":"integrable_trajectorymean_action banditrlproof.thompson.integrable_trajectorymean_action theorem integrable_trajectorymean_action {env : type u} {k : nat} [measurablespace env] (mu : measure (env × ((n : nat) -> fin k × real))) [isfinitemeasure mu] (mean : env -> fin k -> real) (hmean : measurable (fun pair : env × fin k => mean pair.1 pair.2)) (hmeanmem : forall env arm, mean env arm ∈ set.icc l u) (t : nat) : integrable (fun sample : env × ((n : nat) -> fin k × real) => mean sample.1 (environmenttrajectoryaction sample t)) mu theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","label":"integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","description":"Concrete clipped-UCB specialization of the actual-trajectory Thompson Bayesian-regret decomposition. Range and measurability assumptions discharge all four integrability families from `integral_trajectoryBayesMeanRegret_eq_add_historyScore`.","url":"../modules/banditrlproof-algorithms-thompsonclippeducbscore/index.html#decl-cdc4f419f1bd","parent":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","order":2428,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonClippedUCBScore"],["Source","BanditRLProof/Algorithms/ThompsonClippedUCBScore.lean:428"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem integral_trajectoryBayesMeanRegret_eq_add_clippedUCB {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [Nonempty (Fin K)] (prior : Measure Env) [IsProbabilityMeasure prior] (environment : MeasurableHistoryEnvironment Env (Fin K) Real) (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) (mean : Env -> Fin K -> Real) (hmean : Measurable (fun pair : Env × Fin K => mean pair.1 pair.2)) (l u sigma2 delta : Real) (hlu : l <= u) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (horizon : Nat) : let algorithm := uniformReferenceThompsonAlgorithm prior environment bestAction hbestAction let actualMeasure := prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment integral actualMeasure (fun sample => trajectoryBayesMeanRegret mean bestAction sample horizon) = integral actualMeasure (fun sample => ∑ t ∈ range h…","missing":[],"search":"integral_trajectorybayesmeanregret_eq_add_clippeducb banditrlproof.thompson.integral_trajectorybayesmeanregret_eq_add_clippeducb concrete clipped-ucb specialization of the actual-trajectory thompson bayesian-regret decomposition. range and measurability assumptions discharge all four integrability families from `integral_trajectorybayesmeanregret_eq_add_historyscore`. theorem compiled","shard":"modules/b76d90399c48b5dd.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.MeasurableHistoryEnvironment","label":"MeasurableHistoryEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Thompson.MeasurableHistoryEnvironment","description":"Jointly measurable feedback environment. Freezing the first kernel input recovers the pointwise `HistoryEnvironment` consumed by the density route.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-937614d24fb9","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2429,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:25"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"structure MeasurableHistoryEnvironment (Env : Type u) (Action : Type v) (Reward : Type w) [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] where","missing":[],"search":"measurablehistoryenvironment banditrlproof.thompson.measurablehistoryenvironment jointly measurable feedback environment. freezing the first kernel input recovers the pointwise `historyenvironment` consumed by the density route. structure compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.MeasurableHistoryEnvironment.at","label":"at","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.MeasurableHistoryEnvironment.at","description":"Freeze the measurable environment input to recover the pointwise API.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-db762f1462af","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2430,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:52"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def MeasurableHistoryEnvironment.at {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) : HistoryEnvironment Action Reward where","missing":[],"search":"at banditrlproof.thompson.measurablehistoryenvironment.at freeze the measurable environment input to recover the pointwise api. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableEnvironmentInitialPairKernel","label":"measurableEnvironmentInitialPairKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableEnvironmentInitialPairKernel","description":"Jointly measurable law of the initial action/reward pair.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-9fd5f63b7d15","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2431,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:63"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def measurableEnvironmentInitialPairKernel {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) : ProbabilityTheory.Kernel Env (Action × Reward)","missing":[],"search":"measurableenvironmentinitialpairkernel banditrlproof.thompson.measurableenvironmentinitialpairkernel jointly measurable law of the initial action/reward pair. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableEnvironmentHistoryStepKernel","label":"measurableEnvironmentHistoryStepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableEnvironmentHistoryStepKernel","description":"Jointly measurable successor pair kernel over environment and history.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-7d5aa17d13fe","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2432,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:83"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def measurableEnvironmentHistoryStepKernel {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (n : Nat) : ProbabilityTheory.Kernel (Env × History.FinitePairHistory Action Reward n) (Action × Reward)","missing":[],"search":"measurableenvironmenthistorystepkernel banditrlproof.thompson.measurableenvironmenthistorystepkernel jointly measurable successor pair kernel over environment and history. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableEnvironmentInitialPairKernel_apply","label":"measurableEnvironmentInitialPairKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableEnvironmentInitialPairKernel_apply","description":"theorem measurableEnvironmentInitialPairKernel_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) : measurableEnvironmentInitialPairKernel algorithm environment env = algorithm.initialAction ⊗ₘ (environment.at env).initia…","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-bdd8a2c70022","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2433,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:104"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurableEnvironmentInitialPairKernel_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) : measurableEnvironmentInitialPairKernel algorithm environment env = algorithm.initialAction ⊗ₘ (environment.at env).initialFeedback","missing":[],"search":"measurableenvironmentinitialpairkernel_apply banditrlproof.thompson.measurableenvironmentinitialpairkernel_apply theorem measurableenvironmentinitialpairkernel_apply {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] (algorithm : historyalgorithm action reward) (environment : measurablehistoryenvironment env action reward) (env : env) : measurableenvironmentinitialpairkernel algorithm environment env = algorithm.initialaction ⊗ₘ (environment.at env).initialfeedback theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableEnvironmentHistoryStepKernel_apply","label":"measurableEnvironmentHistoryStepKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableEnvironmentHistoryStepKernel_apply","description":"theorem measurableEnvironmentHistoryStepKernel_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Reward n) : measurableEnvironmentHistoryStepKernel algorithm environm…","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-8270c1751b15","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2434,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:117"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurableEnvironmentHistoryStepKernel_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Reward n) : measurableEnvironmentHistoryStepKernel algorithm environment n (env, history) = historyStepKernel algorithm (environment.at env) n history","missing":[],"search":"measurableenvironmenthistorystepkernel_apply banditrlproof.thompson.measurableenvironmenthistorystepkernel_apply theorem measurableenvironmenthistorystepkernel_apply {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] (algorithm : historyalgorithm action reward) (environment : measurablehistoryenvironment env action reward) (n : nat) (env : env) (history : history.finitepairhistory action reward n) : measurableenvironmenthistorystepkernel algorithm environment n (env, history) = historystepkernel algorithm (environment.at env) n history theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment","label":"measurableTrajectoryPrefixEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment","description":"Environment coordinate stored in a finite internal state prefix.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-963267078ec9","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2435,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:133"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def measurableTrajectoryPrefixEnvironment {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} (statePrefix : (i : Finset.Iic n) -> Env × (Action × Reward)) : Env","missing":[],"search":"measurabletrajectoryprefixenvironment banditrlproof.thompson.measurabletrajectoryprefixenvironment environment coordinate stored in a finite internal state prefix. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixHistory","label":"measurableTrajectoryPrefixHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableTrajectoryPrefixHistory","description":"Pair history stored after the dummy zeroth internal state.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-c55f2dcfde73","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2436,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:139"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def measurableTrajectoryPrefixHistory {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} (statePrefix : (i : Finset.Iic (n + 1)) -> Env × (Action × Reward)) : History.FinitePairHistory Action Reward n","missing":[],"search":"measurabletrajectoryprefixhistory banditrlproof.thompson.measurabletrajectoryprefixhistory pair history stored after the dummy zeroth internal state. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixEnvironment","label":"measurable_measurableTrajectoryPrefixEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixEnvironment","description":"theorem measurable_measurableTrajectoryPrefixEnvironment {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (measurableTrajectoryPrefixEnvironment (Env := Env) (Action := Action) (Reward := Reward) (n := n))","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-3f7818671e0a","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2437,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:148"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_measurableTrajectoryPrefixEnvironment {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (measurableTrajectoryPrefixEnvironment (Env := Env) (Action := Action) (Reward := Reward) (n := n))","missing":[],"search":"measurable_measurabletrajectoryprefixenvironment banditrlproof.thompson.measurable_measurabletrajectoryprefixenvironment theorem measurable_measurabletrajectoryprefixenvironment {env : type u} {action : type v} {reward : type w} {n : nat} [measurablespace env] [measurablespace action] [measurablespace reward] : measurable (measurabletrajectoryprefixenvironment (env := env) (action := action) (reward := reward) (n := n)) theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixHistory","label":"measurable_measurableTrajectoryPrefixHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixHistory","description":"theorem measurable_measurableTrajectoryPrefixHistory {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (measurableTrajectoryPrefixHistory (Env := Env) (Action := Action) (Reward := Reward) (n := n))","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-a8ecc7b6affe","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2438,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:158"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_measurableTrajectoryPrefixHistory {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (measurableTrajectoryPrefixHistory (Env := Env) (Action := Action) (Reward := Reward) (n := n))","missing":[],"search":"measurable_measurabletrajectoryprefixhistory banditrlproof.thompson.measurable_measurabletrajectoryprefixhistory theorem measurable_measurabletrajectoryprefixhistory {env : type u} {action : type v} {reward : type w} {n : nat} [measurablespace env] [measurablespace action] [measurablespace reward] : measurable (measurabletrajectoryprefixhistory (env := env) (action := action) (reward := reward) (n := n)) theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixEnvironmentHistory","label":"measurable_measurableTrajectoryPrefixEnvironmentHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixEnvironmentHistory","description":"theorem measurable_measurableTrajectoryPrefixEnvironmentHistory {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (fun statePrefix : (i : Finset.Iic (n + 1)) -> Env × (Action × Reward) => (measurableTrajectoryPrefixEnvironment (n := n + 1) statePrefix, measurableTrajectoryPrefixHistory (n := n) statePrefix))","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-0ca3455adf6b","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2439,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:170"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_measurableTrajectoryPrefixEnvironmentHistory {Env : Type u} {Action : Type v} {Reward : Type w} {n : Nat} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (fun statePrefix : (i : Finset.Iic (n + 1)) -> Env × (Action × Reward) => (measurableTrajectoryPrefixEnvironment (n := n + 1) statePrefix, measurableTrajectoryPrefixHistory (n := n) statePrefix))","missing":[],"search":"measurable_measurabletrajectoryprefixenvironmenthistory banditrlproof.thompson.measurable_measurabletrajectoryprefixenvironmenthistory theorem measurable_measurabletrajectoryprefixenvironmenthistory {env : type u} {action : type v} {reward : type w} {n : nat} [measurablespace env] [measurablespace action] [measurablespace reward] : measurable (fun stateprefix : (i : finset.iic (n + 1)) -> env × (action × reward) => (measurabletrajectoryprefixenvironment (n := n + 1) stateprefix, measurabletrajectoryprefixhistory (n := n) stateprefix)) theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.retainEnvironmentKernel","label":"retainEnvironmentKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.retainEnvironmentKernel","description":"Attach a kernel output to the environment already present in its input.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-4e0a8a6f9240","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2440,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:185"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def retainEnvironmentKernel {Input : Type u} {Env : Type v} {Output : Type w} [MeasurableSpace Input] [MeasurableSpace Env] [MeasurableSpace Output] (env : Input -> Env) (_henv : Measurable env) (kernel : ProbabilityTheory.Kernel Input Output) : ProbabilityTheory.Kernel Input (Env × Output)","missing":[],"search":"retainenvironmentkernel banditrlproof.thompson.retainenvironmentkernel attach a kernel output to the environment already present in its input. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.retainEnvironmentKernel_apply","label":"retainEnvironmentKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.retainEnvironmentKernel_apply","description":"theorem retainEnvironmentKernel_apply {Input : Type u} {Env : Type v} {Output : Type w} [MeasurableSpace Input] [MeasurableSpace Env] [MeasurableSpace Output] (env : Input -> Env) (henv : Measurable env) (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (input : Input) : retainEnvironmentKernel env henv kernel input = (kernel input).map (Prod.mk (env input))","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-1866b0bc610e","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2441,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:206"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem retainEnvironmentKernel_apply {Input : Type u} {Env : Type v} {Output : Type w} [MeasurableSpace Input] [MeasurableSpace Env] [MeasurableSpace Output] (env : Input -> Env) (henv : Measurable env) (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (input : Input) : retainEnvironmentKernel env henv kernel input = (kernel input).map (Prod.mk (env input))","missing":[],"search":"retainenvironmentkernel_apply banditrlproof.thompson.retainenvironmentkernel_apply theorem retainenvironmentkernel_apply {input : type u} {env : type v} {output : type w} [measurablespace input] [measurablespace env] [measurablespace output] (env : input -> env) (henv : measurable env) (kernel : probabilitytheory.kernel input output) [probabilitytheory.ismarkovkernel kernel] (input : input) : retainenvironmentkernel env henv kernel input = (kernel input).map (prod.mk (env input)) theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.retainEnvironmentKernel_map_snd","label":"retainEnvironmentKernel_map_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.retainEnvironmentKernel_map_snd","description":"theorem retainEnvironmentKernel_map_snd {Input : Type u} {Env : Type v} {Output : Type w} [MeasurableSpace Input] [MeasurableSpace Env] [MeasurableSpace Output] (env : Input -> Env) (henv : Measurable env) (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] : (retainEnvironmentKernel env henv kernel).map Prod.snd = kernel","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-663e705d32be","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2442,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:230"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem retainEnvironmentKernel_map_snd {Input : Type u} {Env : Type v} {Output : Type w} [MeasurableSpace Input] [MeasurableSpace Env] [MeasurableSpace Output] (env : Input -> Env) (henv : Measurable env) (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] : (retainEnvironmentKernel env henv kernel).map Prod.snd = kernel","missing":[],"search":"retainenvironmentkernel_map_snd banditrlproof.thompson.retainenvironmentkernel_map_snd theorem retainenvironmentkernel_map_snd {input : type u} {env : type v} {output : type w} [measurablespace input] [measurablespace env] [measurablespace output] (env : input -> env) (henv : measurable env) (kernel : probabilitytheory.kernel input output) [probabilitytheory.ismarkovkernel kernel] : (retainenvironmentkernel env henv kernel).map prod.snd = kernel theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentStepKernel","label":"canonicalMeasurableEnvironmentStepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentStepKernel","description":"Stable internal step family used by the measurable trajectory producer.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-e1eb07241810","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2443,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:249"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalMeasurableEnvironmentStepKernel {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> Env × (Action × Reward)) (Env × (Action × Reward)) | 0 => retainEnvironmentKernel measurableTrajectoryPrefixEnvironment measurable_measurableTrajectoryPrefixEnvironment ((measurableEnvironmentInitialPairKernel algorithm environment).comap measurableTrajectoryPrefixEnvironment measurable_measurableTrajectoryPrefixEnvironment) | n + 1 => retainEnvironmentKernel measurableTrajectoryPrefixEnvironment (measurable_measurableTrajectoryPrefixEnvironment (n := n + 1)) ((measurableEnvironmentHistoryStepKernel algorithm environment n).comap (fun state…","missing":[],"search":"canonicalmeasurableenvironmentstepkernel banditrlproof.thompson.canonicalmeasurableenvironmentstepkernel stable internal step family used by the measurable trajectory producer. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentStepKernel_succ_apply_map_snd","label":"canonicalMeasurableEnvironmentStepKernel_succ_apply_map_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentStepKernel_succ_apply_map_snd","description":"Dropping the retained environment recovers the visible pair step law.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-4a12df6619b9","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2444,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:287"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentStepKernel_succ_apply_map_snd {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (n : Nat) (statePrefix : (i : Finset.Iic (n + 1)) -> Env × (Action × Reward)) : (canonicalMeasurableEnvironmentStepKernel algorithm environment (n + 1) statePrefix).map Prod.snd = historyStepKernel algorithm (environment.at (measurableTrajectoryPrefixEnvironment statePrefix)) n (measurableTrajectoryPrefixHistory statePrefix)","missing":[],"search":"canonicalmeasurableenvironmentstepkernel_succ_apply_map_snd banditrlproof.thompson.canonicalmeasurableenvironmentstepkernel_succ_apply_map_snd dropping the retained environment recovers the visible pair step law. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableEnvironmentInitialStatePrefix","label":"measurableEnvironmentInitialStatePrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableEnvironmentInitialStatePrefix","description":"Dummy time-zero prefix used only to seed `Kernel.traj`.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-7d5976eb2c36","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2445,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:312"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def measurableEnvironmentInitialStatePrefix {Env : Type u} {Action : Type v} {Reward : Type w} [Nonempty Action] [Nonempty Reward] (env : Env) : (i : Finset.Iic 0) -> Env × (Action × Reward)","missing":[],"search":"measurableenvironmentinitialstateprefix banditrlproof.thompson.measurableenvironmentinitialstateprefix dummy time-zero prefix used only to seed `kernel.traj`. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_measurableEnvironmentInitialStatePrefix","label":"measurable_measurableEnvironmentInitialStatePrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_measurableEnvironmentInitialStatePrefix","description":"theorem measurable_measurableEnvironmentInitialStatePrefix {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] : Measurable (measurableEnvironmentInitialStatePrefix (Env := Env) (Action := Action) (Reward := Reward))","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-6e88a96a62fd","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2446,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:318"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_measurableEnvironmentInitialStatePrefix {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] : Measurable (measurableEnvironmentInitialStatePrefix (Env := Env) (Action := Action) (Reward := Reward))","missing":[],"search":"measurable_measurableenvironmentinitialstateprefix banditrlproof.thompson.measurable_measurableenvironmentinitialstateprefix theorem measurable_measurableenvironmentinitialstateprefix {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] [nonempty action] [nonempty reward] : measurable (measurableenvironmentinitialstateprefix (env := env) (action := action) (reward := reward)) theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableEnvironmentPairTrace","label":"measurableEnvironmentPairTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableEnvironmentPairTrace","description":"Pair trace obtained by dropping the dummy internal time-zero state.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-224b7333a373","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2447,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:329"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def measurableEnvironmentPairTrace {Env : Type u} {Action : Type v} {Reward : Type w} (trajectory : (n : Nat) -> Env × (Action × Reward)) : (n : Nat) -> Action × Reward","missing":[],"search":"measurableenvironmentpairtrace banditrlproof.thompson.measurableenvironmentpairtrace pair trace obtained by dropping the dummy internal time-zero state. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_measurableEnvironmentPairTrace","label":"measurable_measurableEnvironmentPairTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_measurableEnvironmentPairTrace","description":"theorem measurable_measurableEnvironmentPairTrace {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (measurableEnvironmentPairTrace (Env := Env) (Action := Action) (Reward := Reward))","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-7f4498fd5fb1","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2448,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:335"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_measurableEnvironmentPairTrace {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (measurableEnvironmentPairTrace (Env := Env) (Action := Action) (Reward := Reward))","missing":[],"search":"measurable_measurableenvironmentpairtrace banditrlproof.thompson.measurable_measurableenvironmentpairtrace theorem measurable_measurableenvironmentpairtrace {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] : measurable (measurableenvironmentpairtrace (env := env) (action := action) (reward := reward)) theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel","label":"canonicalMeasurableEnvironmentTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel","description":"Complete measurable pair-trajectory kernel generated from the joint feedback environment, with no externally supplied trajectory-kernel premise.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-0b4c32b43d39","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2449,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:348"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalMeasurableEnvironmentTrajectoryKernel {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) : ProbabilityTheory.Kernel Env ((n : Nat) -> Action × Reward)","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel complete measurable pair-trajectory kernel generated from the joint feedback environment, with no externally supplied trajectory-kernel premise. definition compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply","label":"canonicalMeasurableEnvironmentTrajectoryKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply","description":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) : canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env = (P…","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-7764dbe159da","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2450,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:373"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_apply {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) : canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env = (ProbabilityTheory.Kernel.traj (canonicalMeasurableEnvironmentStepKernel algorithm environment) 0 (measurableEnvironmentInitialStatePrefix (Action := Action) (Reward := Reward) env)).map measurableEnvironmentPairTrace","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_apply banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_apply theorem canonicalmeasurableenvironmenttrajectorykernel_apply {env : type u} {action : type v} {reward : type w} [measurablespace env] [measurablespace action] [measurablespace reward] [nonempty action] [nonempty reward] (algorithm : historyalgorithm action reward) (environment : measurablehistoryenvironment env action reward) (env : env) : canonicalmeasurableenvironmenttrajectorykernel algorithm environment env = (probabilitytheory.kernel.traj (canonicalmeasurableenvironmentstepkernel algorithm environment) 0 (measurableenvironmentinitialstateprefix (action := action) (reward := reward) env)).map measurableenvironmentpairtrace theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment_initialStatePrefix","label":"measurableTrajectoryPrefixEnvironment_initialStatePrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment_initialStatePrefix","description":"theorem measurableTrajectoryPrefixEnvironment_initialStatePrefix {Env : Type u} {Action : Type v} {Reward : Type w} [Nonempty Action] [Nonempty Reward] (env : Env) : measurableTrajectoryPrefixEnvironment (measurableEnvironmentInitialStatePrefix (Action := Action) (Reward := Reward) env) = env","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-304b31850470","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2451,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:390"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurableTrajectoryPrefixEnvironment_initialStatePrefix {Env : Type u} {Action : Type v} {Reward : Type w} [Nonempty Action] [Nonempty Reward] (env : Env) : measurableTrajectoryPrefixEnvironment (measurableEnvironmentInitialStatePrefix (Action := Action) (Reward := Reward) env) = env","missing":[],"search":"measurabletrajectoryprefixenvironment_initialstateprefix banditrlproof.thompson.measurabletrajectoryprefixenvironment_initialstateprefix theorem measurabletrajectoryprefixenvironment_initialstateprefix {env : type u} {action : type v} {reward : type w} [nonempty action] [nonempty reward] (env : env) : measurabletrajectoryprefixenvironment (measurableenvironmentinitialstateprefix (action := action) (reward := reward) env) = env theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment_ae_eq_of_traj","label":"measurableTrajectoryPrefixEnvironment_ae_eq_of_traj","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment_ae_eq_of_traj","description":"Under a fixed input environment, every finite internal prefix retains it.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-f3ffa10c2e84","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2452,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:399"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurableTrajectoryPrefixEnvironment_ae_eq_of_traj {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) (n : Nat) : (measurableTrajectoryPrefixEnvironment (Env := Env) (Action := Action) (Reward := Reward) (n := n + 1)) =ᵐ[ ((ProbabilityTheory.Kernel.traj (X := fun _ : Nat => Env × (Action × Reward)) (canonicalMeasurableEnvironmentStepKernel algorithm environment) 0) (measurableEnvironmentInitialStatePrefix (Action := Action) (Reward := Reward) env)).map (Preorder.frestrictLe (n + 1))] fun _ : ((i : Finset.Iic (n + 1)) -> Env × (Action × Reward)) => env","missing":[],"search":"measurabletrajectoryprefixenvironment_ae_eq_of_traj banditrlproof.thompson.measurabletrajectoryprefixenvironment_ae_eq_of_traj under a fixed input environment, every finite internal prefix retains it. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_eval_zero","label":"canonicalMeasurableEnvironmentTrajectoryKernel_map_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_eval_zero","description":"The generated trajectory kernel has the configured initial pair law.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-5cc0b691fed9","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2453,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:460"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_map_eval_zero {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) : (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun trajectory => trajectory 0) = measurableEnvironmentInitialPairKernel algorithm environment","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_map_eval_zero banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_map_eval_zero the generated trajectory kernel has the configured initial pair law. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_prefix_next_eq_compProd","label":"canonicalMeasurableEnvironmentTrajectoryKernel_map_prefix_next_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_prefix_next_eq_compProd","description":"Joint finite-prefix/next-pair law of the projected measurable trajectory.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-d493037beddd","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2454,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:533"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_map_prefix_next_eq_compProd {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) (n : Nat) : (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env).map (fun trajectory => (Preorder.frestrictLe n trajectory, trajectory (n + 1))) = (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env).map (Preorder.frestrictLe n) ⊗ₘ historyStepKernel algorithm (environment.at env) n","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_map_prefix_next_eq_compprod banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_map_prefix_next_eq_compprod joint finite-prefix/next-pair law of the projected measurable trajectory. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_condDistrib_succ","label":"canonicalMeasurableEnvironmentTrajectoryKernel_condDistrib_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_condDistrib_succ","description":"The projected trajectory has the configured shifted successor pair law.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-6ff9b9a67d37","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2455,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:660"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_condDistrib_succ {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : (k : Nat) -> Action × Reward => trajectory (n + 1)) (Preorder.frestrictLe n) (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env) =ᵐ[ (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env).map (Preorder.frestrictLe n)] historyStepKernel algorithm (environment.at env) n","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_conddistrib_succ banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_conddistrib_succ the projected trajectory has the configured shifted successor pair law. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical_of_step_condDistrib","label":"canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical_of_step_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical_of_step_condDistrib","description":"The generated measurable trajectory kernel has the canonical fixed-environment law once its shifted successor conditional laws are identified. The initial law is discharged internally by the preceding theorem.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-4c00cdbc4388","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2456,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:687"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical_of_step_condDistrib {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) (hstep : forall n, ProbabilityTheory.condDistrib (fun trajectory : (k : Nat) -> Action × Reward => trajectory (n + 1)) (Preorder.frestrictLe n) (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env) =ᵐ[ (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env).map (Preorder.frestrictLe n)] historyStepKernel algorithm (environment.at env) n) : canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env = canonicalHist…","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_apply_eq_canonical_of_step_conddistrib banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_apply_eq_canonical_of_step_conddistrib the generated measurable trajectory kernel has the canonical fixed-environment law once its shifted successor conditional laws are identified. the initial law is discharged internally by the preceding theorem. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical","label":"canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical","description":"The generated measurable trajectory kernel is pointwise canonical.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-a1ea23b0dba3","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2457,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:734"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) : canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env = canonicalHistoryTrajectoryMeasure algorithm (environment.at env)","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_apply_eq_canonical banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_apply_eq_canonical the generated measurable trajectory kernel is pointwise canonical. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment_stepCondDistrib","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment_stepCondDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment_stepCondDistrib","description":"Finite-prefix probability matching from jointly measurable actual/reference feedback environments. The only remaining process premise is the shifted successor conditional law of the two generated `Kernel.traj` kernels.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-e9a00d11f51a","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2458,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:754"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment_stepCondDistrib {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (absolutelyContinuous : HistoryAlgorithmAbsolutelyContinuous algorithm referenceAlgorithm) (hstep : forall env n, ProbabilityTheory.condDistrib (fun trajectory : (k : Nat) -> Action × Reward => trajectory (n + 1)) (Preorder.frestrictLe n) (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env) =ᵐ[ (canonicalMeasurableEnvironmentTraj…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_measurableenvironment_stepconddistrib banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_measurableenvironment_stepconddistrib finite-prefix probability matching from jointly measurable actual/reference feedback environments. the only remaining process premise is the shifted successor conditional law of the two generated `kernel.traj` kernels. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment","description":"Finite-prefix Thompson probability matching directly from a jointly measurable feedback environment, with both trajectory kernels and their process laws constructed internally.","url":"../modules/banditrlproof-algorithms-thompsonmeasurabletrajectory/index.html#decl-2d77f5f922a2","parent":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","order":2459,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonMeasurableTrajectory"],["Source","BanditRLProof/Algorithms/ThompsonMeasurableTrajectory.lean:832"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (absolutelyContinuous : HistoryAlgorithmAbsolutelyContinuous algorithm referenceAlgorithm) (n : Nat) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : let trajectoryKernel := canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment let referenceTrajectoryKernel := canonicalMeasurableEnvironmentTrajectoryKernel referenceAlgorithm environ…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_measurableenvironment banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_measurableenvironment finite-prefix thompson probability matching directly from a jointly measurable feedback environment, with both trajectory kernels and their process laws constructed internally. theorem compiled","shard":"modules/a712a405dcde03e2.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformActionMeasure","label":"uniformActionMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformActionMeasure","description":"Uniform probability measure on a nonempty finite action space.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-129222abc389","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2460,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:22"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","combinatorial"]],"statement":"noncomputable def uniformActionMeasure (Action : Type u) [Fintype Action] [Nonempty Action] [MeasurableSpace Action] : Measure Action","missing":[],"search":"uniformactionmeasure banditrlproof.thompson.uniformactionmeasure uniform probability measure on a nonempty finite action space. definition compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":["combinatorial"]},{"id":"declaration:BanditRLProof.Thompson.absolutelyContinuous_uniformActionMeasure","label":"absolutelyContinuous_uniformActionMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.absolutelyContinuous_uniformActionMeasure","description":"Every measure on a finite space is dominated by its uniform measure.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-37f4b1a1194a","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2461,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:35"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem absolutelyContinuous_uniformActionMeasure {Action : Type u} [Fintype Action] [Nonempty Action] [MeasurableSpace Action] (mu : Measure Action) : mu ≪ uniformActionMeasure Action","missing":[],"search":"absolutelycontinuous_uniformactionmeasure banditrlproof.thompson.absolutelycontinuous_uniformactionmeasure every measure on a finite space is dominated by its uniform measure. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformHistoryAlgorithm","label":"uniformHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformHistoryAlgorithm","description":"History-independent uniform reference algorithm.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-d312bf8087e3","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2462,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:57"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformHistoryAlgorithm (Action : Type u) (Reward : Type v) [Fintype Action] [Nonempty Action] [MeasurableSpace Action] [MeasurableSpace Reward] : HistoryAlgorithm Action Reward where","missing":[],"search":"uniformhistoryalgorithm banditrlproof.thompson.uniformhistoryalgorithm history-independent uniform reference algorithm. definition compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.historyAlgorithmAbsolutelyContinuous_uniform","label":"historyAlgorithmAbsolutelyContinuous_uniform","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.historyAlgorithmAbsolutelyContinuous_uniform","description":"Any history algorithm is absolutely continuous with respect to uniform.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-789163f39e9d","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2463,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:67"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem historyAlgorithmAbsolutelyContinuous_uniform {Action : Type u} {Reward : Type v} [Fintype Action] [Nonempty Action] [MeasurableSpace Action] [MeasurableSpace Reward] (algorithm : HistoryAlgorithm Action Reward) : HistoryAlgorithmAbsolutelyContinuous algorithm (uniformHistoryAlgorithm Action Reward) where","missing":[],"search":"historyalgorithmabsolutelycontinuous_uniform banditrlproof.thompson.historyalgorithmabsolutelycontinuous_uniform any history algorithm is absolutely continuous with respect to uniform. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.trajectoryMixture_map_history_action_eq_compProd","label":"trajectoryMixture_map_history_action_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.trajectoryMixture_map_history_action_eq_compProd","description":"Mixing environment-indexed trajectory laws preserves a common conditional action kernel. This is the generic measure transport needed to turn pointwise trajectory laws into one global recursive process law.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-d8a996a9d8a7","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2464,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:85"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMixture_map_history_action_eq_compProd {Env : Type u} {Omega : Type v} {History : Type w} {Action : Type x} [MeasurableSpace Env] [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] (prior : Measure Env) [IsFiniteMeasure prior] (trajectory : ProbabilityTheory.Kernel Env Omega) [ProbabilityTheory.IsMarkovKernel trajectory] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] (hlaw : forall env, (trajectory env).map (fun omega => (history omega, action omega)) = (trajectory env).map history ⊗ₘ policy) : (prior ⊗ₘ trajectory).map (fun sample => (history sample.2, action sample.2)) = (prior ⊗ₘ trajectory).map (history ∘ Prod.snd) ⊗ₘ policy","missing":[],"search":"trajectorymixture_map_history_action_eq_compprod banditrlproof.thompson.trajectorymixture_map_history_action_eq_compprod mixing environment-indexed trajectory laws preserves a common conditional action kernel. this is the generic measure transport needed to turn pointwise trajectory laws into one global recursive process law. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.trajectoryMixture_condDistrib_action","label":"trajectoryMixture_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.trajectoryMixture_condDistrib_action","description":"Conditional-law form of `trajectoryMixture_map_history_action_eq_compProd`.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-c84acf149e92","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2465,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:148"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMixture_condDistrib_action {Env : Type u} {Omega : Type v} {History : Type w} {Action : Type x} [MeasurableSpace Env] [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (prior : Measure Env) [IsFiniteMeasure prior] (trajectory : ProbabilityTheory.Kernel Env Omega) [ProbabilityTheory.IsMarkovKernel trajectory] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] (hlaw : forall env, (trajectory env).map (fun omega => (history omega, action omega)) = (trajectory env).map history ⊗ₘ policy) : ProbabilityTheory.condDistrib (action ∘ Prod.snd) (history ∘ Prod.snd) (prior ⊗ₘ trajectory) =ᵐ[(prior ⊗ₘ trajectory).map (history ∘ Prod.snd)] policy","missing":[],"search":"trajectorymixture_conddistrib_action banditrlproof.thompson.trajectorymixture_conddistrib_action conditional-law form of `trajectorymixture_map_history_action_eq_compprod`. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_history_action_eq_compProd","label":"canonicalMeasurableEnvironmentTrajectoryKernel_map_history_action_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_history_action_eq_compProd","description":"The visible action marginal of each fixed-environment successor law.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-835d6da86415","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2466,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:172"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryKernel_map_history_action_eq_compProd {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (env : Env) (n : Nat) : (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env).map (fun trajectory => (Preorder.frestrictLe n trajectory, (trajectory (n + 1)).1)) = (canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment env).map (Preorder.frestrictLe n) ⊗ₘ algorithm.policy n","missing":[],"search":"canonicalmeasurableenvironmenttrajectorykernel_map_history_action_eq_compprod banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorykernel_map_history_action_eq_compprod the visible action marginal of each fixed-environment successor law. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action","description":"The next action coordinate of the global prior/trajectory measure has the algorithm policy as its conditional law given the visible finite history.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-860136cc5d16","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2467,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:209"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (n : Nat) : ProbabilityTheory.condDistrib (fun sample : Env × ((k : Nat) -> Action × Reward) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment) =ᵐ[ (prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample => Preorder.frestrictLe n sample.2)] algorithm.policy n","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_conddistrib_action banditrlproof.thompson.canonicalmeasurableenvironmenttrajectorymeasure_conddistrib_action the next action coordinate of the global prior/trajectory measure has the algorithm policy as its conditional law given the visible finite history. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePosterior_ae_eq_condDistrib_of_conditionalProcessSource","label":"finitePairReferencePosterior_ae_eq_condDistrib_of_conditionalProcessSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePosterior_ae_eq_condDistrib_of_conditionalProcessSource","description":"Expose the posterior-invariance conclusion of the conditional process-density route without adjoining a fresh action sampler.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-13ef918442f3","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2468,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:241"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePosterior_ae_eq_condDistrib_of_conditionalProcessSource {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type*} [MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [MeasurableSpace OmegaRef] [StandardBorelSpace OmegaRef] [Nonempty OmegaRef] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward) (algorithm referenceAlgorithm : HistoryAl…","missing":[],"search":"finitepairreferenceposterior_ae_eq_conddistrib_of_conditionalprocesssource banditrlproof.thompson.finitepairreferenceposterior_ae_eq_conddistrib_of_conditionalprocesssource expose the posterior-invariance conclusion of the conditional process-density route without adjoining a fresh action sampler. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm","label":"referencePosteriorHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm","description":"Thompson's non-circular history algorithm: every policy is the posterior under one fixed reference trajectory, mapped through `bestAction`.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-62b409f22792","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2469,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:305"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def referencePosteriorHistoryAlgorithm {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : HistoryAlgorithm Action Reward where","missing":[],"search":"referenceposteriorhistoryalgorithm banditrlproof.thompson.referenceposteriorhistoryalgorithm thompson's non-circular history algorithm: every policy is the posterior under one fixed reference trajectory, mapped through `bestaction`. definition compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_initialAction","label":"referencePosteriorHistoryAlgorithm_initialAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_initialAction","description":"theorem referencePosteriorHistoryAlgorithm_initialAction {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Rew…","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-af7f8e5a3f57","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2470,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:334"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePosteriorHistoryAlgorithm_initialAction {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : (referencePosteriorHistoryAlgorithm prior referenceAlgorithm environment bestAction hbestAction).initialAction = prior.map bestAction","missing":[],"search":"referenceposteriorhistoryalgorithm_initialaction banditrlproof.thompson.referenceposteriorhistoryalgorithm_initialaction theorem referenceposteriorhistoryalgorithm_initialaction {env : type u} {action : type v} {reward : type w} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablespace reward] [nonempty action] [nonempty reward] (prior : measure env) [isprobabilitymeasure prior] (referencealgorithm : historyalgorithm action reward) (environment : measurablehistoryenvironment env action reward) (bestaction : env -> action) (hbestaction : measurable bestaction) : (referenceposteriorhistoryalgorithm prior referencealgorithm environment bestaction hbestaction).initialaction = prior.map bestaction theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_policy","label":"referencePosteriorHistoryAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_policy","description":"theorem referencePosteriorHistoryAlgorithm_policy {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (b…","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-adbcb226ae8d","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2471,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:347"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePosteriorHistoryAlgorithm_policy {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSpace Reward] [Nonempty Action] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (n : Nat) : (referencePosteriorHistoryAlgorithm prior referenceAlgorithm environment bestAction hbestAction).policy n = referenceActionKernel (prior ⊗ₘ canonicalMeasurableEnvironmentTrajectoryKernel referenceAlgorithm environment) Prod.fst (fun sample => History.finitePairHistoryOfTrace (environmentTrajectoryAction sample) (environmentTrajectoryReward sample) n) measurable_fst (History.measurable_finitePairHisto…","missing":[],"search":"referenceposteriorhistoryalgorithm_policy banditrlproof.thompson.referenceposteriorhistoryalgorithm_policy theorem referenceposteriorhistoryalgorithm_policy {env : type u} {action : type v} {reward : type w} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablespace reward] [nonempty action] [nonempty reward] (prior : measure env) [isprobabilitymeasure prior] (referencealgorithm : historyalgorithm action reward) (environment : measurablehistoryenvironment env action reward) (bestaction : env -> action) (hbestaction : measurable bestaction) (n : nat) : (referenceposteriorhistoryalgorithm prior referencealgorithm environment bestaction hbestaction).policy n = referenceactionkernel (prior ⊗ₘ canonicalmeasurableenvironmenttrajectorykernel referencealgorithm environment) prod.fst (fun sample => history.finitepairhistoryoftrace (environmenttrajectoryaction sample) (environmenttrajectoryreward sample) n) measurable_fst (history.measurable_finitepairhistoryoftrace environmenttrajectoryaction environmenttrajectoryreward measurable_environmenttrajectoryaction_apply measurable_environmenttrajectoryreward_apply n) bestaction hbestaction theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","label":"referencePosteriorHistoryAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","description":"Probability matching for the action coordinate of one globally generated Thompson trajectory. Unlike the earlier finite-prefix sampler endpoint, the next action here is the actual successor coordinate of the same recursive trajectory whose history appears in the conditioning variable. The remaining support contract is the standard algorithm-density condition against the fixed reference algorithm. A finite uniform re…","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-b67d98f65318","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2472,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:383"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePosteriorHistoryAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (referenceAlgorithm : HistoryAlgorithm Action Reward) (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (absolutelyContinuous : HistoryAlgorithmAbsolutelyContinuous (referencePosteriorHistoryAlgorithm prior referenceAlgorithm environment bestAction hbestAction) referenceAlgorithm) (n : Nat) : let algorithm := referencePosteriorHistoryAlgorithm prior referenceAlgorithm environment bestAction hbestAction let trajectoryKer…","missing":[],"search":"referenceposteriorhistoryalgorithm_trajectory_conddistrib_action_ae_eq_bestaction banditrlproof.thompson.referenceposteriorhistoryalgorithm_trajectory_conddistrib_action_ae_eq_bestaction probability matching for the action coordinate of one globally generated thompson trajectory. unlike the earlier finite-prefix sampler endpoint, the next action here is the actual successor coordinate of the same recursive trajectory whose history appears in the conditioning variable. the remaining support contract is the standard algorithm-density condition against the fixed reference algorithm. a finite uniform reference discharges that contract in the downstream specialization. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm","label":"uniformReferenceThompsonAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm","description":"Concrete finite-action Thompson algorithm using one uniform reference process. The definition is non-circular: its posterior policy is computed from the uniform algorithm's trajectory, not from the Thompson trajectory being built.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-7e7ffba57d7b","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2473,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:490"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformReferenceThompsonAlgorithm {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [Fintype Action] [Nonempty Action] [MeasurableSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) : HistoryAlgorithm Action Reward","missing":[],"search":"uniformreferencethompsonalgorithm banditrlproof.thompson.uniformreferencethompsonalgorithm concrete finite-action thompson algorithm using one uniform reference process. the definition is non-circular: its posterior policy is computed from the uniform algorithm's trajectory, not from the thompson trajectory being built. definition compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","label":"uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","description":"Premise-free finite-action probability matching on the actual globally recursive Thompson trajectory. Uniform full support discharges every algorithm-density absolute-continuity obligation internally.","url":"../modules/banditrlproof-algorithms-thompsonrecursivesampler/index.html#decl-103dd0175542","parent":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","order":2474,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonRecursiveSampler"],["Source","BanditRLProof/Algorithms/ThompsonRecursiveSampler.lean:507"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Fintype Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsProbabilityMeasure prior] (environment : MeasurableHistoryEnvironment Env Action Reward) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (n : Nat) : let algorithm := uniformReferenceThompsonAlgorithm prior environment bestAction hbestAction let trajectoryKernel := canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment let actualMeasure := prior ⊗ₘ trajectoryKernel let actualHistory := fun sample => History.finitePairHistoryOfTrace (environmentTrajectoryAction sample) (environ…","missing":[],"search":"uniformreferencethompsonalgorithm_trajectory_conddistrib_action_ae_eq_bestaction banditrlproof.thompson.uniformreferencethompsonalgorithm_trajectory_conddistrib_action_ae_eq_bestaction premise-free finite-action probability matching on the actual globally recursive thompson trajectory. uniform full support discharges every algorithm-density absolute-continuity obligation internally. theorem compiled","shard":"modules/c3b4d82ee57cc6a8.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosterior","label":"referencePosterior","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosterior","description":"A reference source's environment posterior given its history.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-15e79ea28b0b","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2475,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:27"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def referencePosterior {OmegaRef : Type u} {History : Type v} {Env : Type w} [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (_hreferenceEnv : Measurable referenceEnv) (_hreferenceHistory : Measurable referenceHistory) : PosteriorKernel.MarkovPosteriorKernel History Env","missing":[],"search":"referenceposterior banditrlproof.thompson.referenceposterior a reference source's environment posterior given its history. definition compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePosterior_kernel","label":"referencePosterior_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePosterior_kernel","description":"theorem referencePosterior_kernel {OmegaRef : Type u} {History : Type v} {Env : Type w} [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable refer…","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-de06988d73cc","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2476,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:42"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePosterior_kernel {OmegaRef : Type u} {History : Type v} {Env : Type w} [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) : (referencePosterior referenceMu referenceEnv referenceHistory hreferenceEnv hreferenceHistory).kernel = ProbabilityTheory.condDistrib referenceEnv referenceHistory referenceMu","missing":[],"search":"referenceposterior_kernel banditrlproof.thompson.referenceposterior_kernel theorem referenceposterior_kernel {omegaref : type u} {history : type v} {env : type w} [measurablespace omegaref] [measurablespace history] [measurablespace env] [standardborelspace env] [nonempty env] (referencemu : measure omegaref) [isfinitemeasure referencemu] (referenceenv : omegaref -> env) (referencehistory : omegaref -> history) (hreferenceenv : measurable referenceenv) (hreferencehistory : measurable referencehistory) : (referenceposterior referencemu referenceenv referencehistory hreferenceenv hreferencehistory).kernel = probabilitytheory.conddistrib referenceenv referencehistory referencemu theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referenceActionKernel","label":"referenceActionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.referenceActionKernel","description":"The Thompson action policy obtained by mapping a reference posterior.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-af9d16e44fa5","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2477,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:57"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def referenceActionKernel {OmegaRef : Type u} {History : Type v} {Env : Type w} {Action : Type x} [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) (bestAction : Env -> Action) (_hbestAction : Measurable bestAction) : ProbabilityTheory.Kernel History Action","missing":[],"search":"referenceactionkernel banditrlproof.thompson.referenceactionkernel the thompson action policy obtained by mapping a reference posterior. definition compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerMeasure","label":"policySamplerMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerMeasure","description":"Extend a base process by sampling an action from a history-indexed policy.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-040f005a3140","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2478,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:90"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def policySamplerMeasure {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] : Measure (Omega × Action)","missing":[],"search":"policysamplermeasure banditrlproof.thompson.policysamplermeasure extend a base process by sampling an action from a history-indexed policy. definition compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerEnv","label":"policySamplerEnv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerEnv","description":"Base environment coordinate after adjoining a sampled action.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-511d795339b2","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2479,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:112"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def policySamplerEnv {Omega Env Action : Type*} (env : Omega -> Env) : Omega × Action -> Env","missing":[],"search":"policysamplerenv banditrlproof.thompson.policysamplerenv base environment coordinate after adjoining a sampled action. definition compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerHistory","label":"policySamplerHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerHistory","description":"Base history coordinate after adjoining a sampled action.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-e825f90fc6b2","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2480,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:118"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def policySamplerHistory {Omega History Action : Type*} (history : Omega -> History) : Omega × Action -> History","missing":[],"search":"policysamplerhistory banditrlproof.thompson.policysamplerhistory base history coordinate after adjoining a sampled action. definition compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerAction","label":"policySamplerAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerAction","description":"Newly sampled action coordinate.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-c19e63a28623","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2481,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:124"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def policySamplerAction {Omega Action : Type*} : Omega × Action -> Action","missing":[],"search":"policysampleraction banditrlproof.thompson.policysampleraction newly sampled action coordinate. definition compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerEnv_measurable","label":"policySamplerEnv_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerEnv_measurable","description":"theorem policySamplerEnv_measurable {Omega Env Action : Type*} [MeasurableSpace Omega] [MeasurableSpace Env] [MeasurableSpace Action] (env : Omega -> Env) (henv : Measurable env) : Measurable (@policySamplerEnv Omega Env Action env)","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-a0d2e288fc28","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2482,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:127"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySamplerEnv_measurable {Omega Env Action : Type*} [MeasurableSpace Omega] [MeasurableSpace Env] [MeasurableSpace Action] (env : Omega -> Env) (henv : Measurable env) : Measurable (@policySamplerEnv Omega Env Action env)","missing":[],"search":"policysamplerenv_measurable banditrlproof.thompson.policysamplerenv_measurable theorem policysamplerenv_measurable {omega env action : type*} [measurablespace omega] [measurablespace env] [measurablespace action] (env : omega -> env) (henv : measurable env) : measurable (@policysamplerenv omega env action env) theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerHistory_measurable","label":"policySamplerHistory_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerHistory_measurable","description":"theorem policySamplerHistory_measurable {Omega History Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] (history : Omega -> History) (hhistory : Measurable history) : Measurable (@policySamplerHistory Omega History Action history)","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-63685fae1d1c","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2483,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:134"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySamplerHistory_measurable {Omega History Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] (history : Omega -> History) (hhistory : Measurable history) : Measurable (@policySamplerHistory Omega History Action history)","missing":[],"search":"policysamplerhistory_measurable banditrlproof.thompson.policysamplerhistory_measurable theorem policysamplerhistory_measurable {omega history action : type*} [measurablespace omega] [measurablespace history] [measurablespace action] (history : omega -> history) (hhistory : measurable history) : measurable (@policysamplerhistory omega history action history) theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySamplerAction_measurable","label":"policySamplerAction_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySamplerAction_measurable","description":"theorem policySamplerAction_measurable {Omega Action : Type*} [MeasurableSpace Omega] [MeasurableSpace Action] : Measurable (@policySamplerAction Omega Action)","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-5bee4e7bbea1","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2484,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:141"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySamplerAction_measurable {Omega Action : Type*} [MeasurableSpace Omega] [MeasurableSpace Action] : Measurable (@policySamplerAction Omega Action)","missing":[],"search":"policysampleraction_measurable banditrlproof.thompson.policysampleraction_measurable theorem policysampleraction_measurable {omega action : type*} [measurablespace omega] [measurablespace action] : measurable (@policysampleraction omega action) theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.map_compProd_comap_history","label":"map_compProd_comap_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.map_compProd_comap_history","description":"History/action projection of a policy sampler. This is the arbitrary-history-map version of `map_compProd_comap_snd` and is a generic Mathlib candidate.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-d0902bbc089f","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2485,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:153"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem map_compProd_comap_history {Omega History Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] : (mu ⊗ₘ policy.comap history hhistory).map (fun sample => (history sample.1, sample.2)) = mu.map history ⊗ₘ policy","missing":[],"search":"map_compprod_comap_history banditrlproof.thompson.map_compprod_comap_history history/action projection of a policy sampler. this is the arbitrary-history-map version of `map_compprod_comap_snd` and is a generic mathlib candidate. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySampler_base_map_eq","label":"policySampler_base_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySampler_base_map_eq","description":"Adjoining a Markov-policy sample preserves every measurable base map.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-19ec1b25039b","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2486,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:183"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySampler_base_map_eq {Omega History Action Target : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSpace Target] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] (f : Omega -> Target) (hf : Measurable f) : (policySamplerMeasure mu history hhistory policy).map (fun sample => f sample.1) = mu.map f","missing":[],"search":"policysampler_base_map_eq banditrlproof.thompson.policysampler_base_map_eq adjoining a markov-policy sample preserves every measurable base map. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySampler_history_action_map_eq","label":"policySampler_history_action_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySampler_history_action_map_eq","description":"The constructed sampler's history/action joint law is `historyLaw ⊗ policy`.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-985aa7199498","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2487,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:207"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySampler_history_action_map_eq {Omega History Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] : (policySamplerMeasure mu history hhistory policy).map (fun sample => (policySamplerHistory history sample, policySamplerAction sample)) = (policySamplerMeasure mu history hhistory policy).map (policySamplerHistory history) ⊗ₘ policy","missing":[],"search":"policysampler_history_action_map_eq banditrlproof.thompson.policysampler_history_action_map_eq the constructed sampler's history/action joint law is `historylaw ⊗ policy`. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySampler_condDistrib_action_ae_eq_policy","label":"policySampler_condDistrib_action_ae_eq_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySampler_condDistrib_action_ae_eq_policy","description":"The sampled action has the policy as its conditional law given history.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-b0cdd1fdd061","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2488,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:229"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySampler_condDistrib_action_ae_eq_policy {Omega History Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] : ProbabilityTheory.condDistrib policySamplerAction (policySamplerHistory history) (policySamplerMeasure mu history hhistory policy) =ᵐ[ (policySamplerMeasure mu history hhistory policy).map (policySamplerHistory history)] policy","missing":[],"search":"policysampler_conddistrib_action_ae_eq_policy banditrlproof.thompson.policysampler_conddistrib_action_ae_eq_policy the sampled action has the policy as its conditional law given history. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.policySampler_condDistrib_env_ae_eq_of_base","label":"policySampler_condDistrib_env_ae_eq_of_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.policySampler_condDistrib_env_ae_eq_of_base","description":"Adjoining a history-dependent action sample preserves an environment posterior.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-0dbdd6da4173","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2489,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:251"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem policySampler_condDistrib_env_ae_eq_of_base {Omega History Env Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (hhistory : Measurable history) (policy : ProbabilityTheory.Kernel History Action) [ProbabilityTheory.IsMarkovKernel policy] (posterior : ProbabilityTheory.Kernel History Env) [ProbabilityTheory.IsMarkovKernel posterior] (hbase : ProbabilityTheory.condDistrib env history mu =ᵐ[mu.map history] posterior) : ProbabilityTheory.condDistrib (policySamplerEnv env) (policySamplerHistory history) (policySamplerMeasure mu history hhistory policy) =ᵐ[ (policySamplerMeasure mu history hhistory policy).map (policySamplerHistory history)] posterior","missing":[],"search":"policysampler_conddistrib_env_ae_eq_of_base banditrlproof.thompson.policysampler_conddistrib_env_ae_eq_of_base adjoining a history-dependent action sample preserves an environment posterior. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","label":"referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","description":"Reference-posterior Thompson probability matching with a constructed sampler. The action law is generated by `policySamplerMeasure`. The sole law transport premise is that the reference posterior agrees with the actual environment posterior at the actual history law; this is the conclusion supplied by LML's algorithm-density route.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-03957617c6e4","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2490,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:307"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance {Omega : Type u} {OmegaRef : Type v} {History : Type w} {Env : Type x} {Action : Type y} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace History] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (hhistory : Measurable history) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceHistory : OmegaRef -> History) (hreferenceEnv : Measurable referenceEnv) (hreferenceHistory : Measurable referenceHistory) (bestAction : Env -> Action) (hbestAction : Measurable bestAction) (hposteriorInvariance : (referencePosterior referenceMu reference…","missing":[],"search":"referencepolicysampler_conddistrib_action_ae_eq_bestaction_of_posterior_invariance banditrlproof.thompson.referencepolicysampler_conddistrib_action_ae_eq_bestaction_of_posterior_invariance reference-posterior thompson probability matching with a constructed sampler. the action law is generated by `policysamplermeasure`. the sole law transport premise is that the reference posterior agrees with the actual environment posterior at the actual history law; this is the conclusion supplied by lml's algorithm-density route. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","label":"finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","description":"Finite action/reward-prefix specialization of the reference-policy theorem. This is the per-time interface used by the recursive bandit route. The remaining premise is exactly posterior invariance between the reference and actual finite-pair history laws.","url":"../modules/banditrlproof-algorithms-thompsonreferencepolicy/index.html#decl-2d16a8c5f700","parent":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","order":2491,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonReferencePolicy"],["Source","BanditRLProof/Algorithms/ThompsonReferencePolicy.lean:373"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance {Omega : Type u} {OmegaRef : Type v} {Env : Type w} {Action : Type x} {Reward : Type y} [MeasurableSpace Omega] [MeasurableSpace OmegaRef] [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (henv : Measurable env) (haction : forall t : Nat, Measurable (fun omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (referenceMu : Measure OmegaRef) [IsFiniteMeasure referenceMu] (referenceEnv : OmegaRef -> Env) (referenceAction : OmegaRef -> ActionTrace Action) (referenceReward : OmegaRef -> RewardTrace Reward)…","missing":[],"search":"finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_posterior_invariance banditrlproof.thompson.finitepairreferencepolicysampler_conddistrib_action_ae_eq_bestaction_of_posterior_invariance finite action/reward-prefix specialization of the reference-policy theorem. this is the per-time interface used by the recursive bandit route. the remaining premise is exactly posterior invariance between the reference and actual finite-pair history laws. theorem compiled","shard":"modules/5b47603f52531259.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.UnitArmStream","label":"UnitArmStream","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Thompson.UnitArmStream","description":"Uniform random table used to represent all stationary reward-kernel draws.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-8f9acf67648c","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2492,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:23"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"abbrev UnitArmStream (K : Nat)","missing":[],"search":"unitarmstream banditrlproof.thompson.unitarmstream uniform random table used to represent all stationary reward-kernel draws. abbreviation compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.uniformUnitArmStreamMeasure","label":"uniformUnitArmStreamMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.uniformUnitArmStreamMeasure","description":"Independent uniform coordinates indexed by pull number and arm.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-e2734bedd9a6","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2493,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:27"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformUnitArmStreamMeasure (K : Nat) : Measure (UnitArmStream K)","missing":[],"search":"uniformunitarmstreammeasure banditrlproof.thompson.uniformunitarmstreammeasure independent uniform coordinates indexed by pull number and arm. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryRewardSampler","label":"stationaryRewardSampler","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryRewardSampler","description":"A fixed measurable uniform-randomness representation of a reward kernel.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-f6530b069cb3","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2494,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:38"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryRewardSampler {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : (Env × Fin K) -> Set.Icc (0 : Real) 1 -> Real","missing":[],"search":"stationaryrewardsampler banditrlproof.thompson.stationaryrewardsampler a fixed measurable uniform-randomness representation of a reward kernel. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_uncurry_stationaryRewardSampler","label":"measurable_uncurry_stationaryRewardSampler","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_uncurry_stationaryRewardSampler","description":"theorem measurable_uncurry_stationaryRewardSampler {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Measurable (Function.uncurry (stationaryRewardSampler rewardKernel))","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-79399ec87555","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2495,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:44"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_uncurry_stationaryRewardSampler {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Measurable (Function.uncurry (stationaryRewardSampler rewardKernel))","missing":[],"search":"measurable_uncurry_stationaryrewardsampler banditrlproof.thompson.measurable_uncurry_stationaryrewardsampler theorem measurable_uncurry_stationaryrewardsampler {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] : measurable (function.uncurry (stationaryrewardsampler rewardkernel)) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryRewardSampler_map_volume","label":"stationaryRewardSampler_map_volume","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryRewardSampler_map_volume","description":"theorem stationaryRewardSampler_map_volume {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (input : Env × Fin K) : Measure.map (stationaryRewardSampler rewardKernel input) volume = rewardKernel input","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-169dbed4bb07","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2496,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:50"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryRewardSampler_map_volume {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (input : Env × Fin K) : Measure.map (stationaryRewardSampler rewardKernel input) volume = rewardKernel input","missing":[],"search":"stationaryrewardsampler_map_volume banditrlproof.thompson.stationaryrewardsampler_map_volume theorem stationaryrewardsampler_map_volume {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] (input : env × fin k) : measure.map (stationaryrewardsampler rewardkernel input) volume = rewardkernel input theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.rewardStreamOfUnitArmStream","label":"rewardStreamOfUnitArmStream","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.rewardStreamOfUnitArmStream","description":"Turn one environment and one uniform table into its latent reward table.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-59fc604a7fd5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2497,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:59"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardStreamOfUnitArmStream {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Env × UnitArmStream K -> UCB.ArmRewardStream K","missing":[],"search":"rewardstreamofunitarmstream banditrlproof.thompson.rewardstreamofunitarmstream turn one environment and one uniform table into its latent reward table. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_rewardStreamOfUnitArmStream","label":"measurable_rewardStreamOfUnitArmStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_rewardStreamOfUnitArmStream","description":"theorem measurable_rewardStreamOfUnitArmStream {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Measurable (rewardStreamOfUnitArmStream rewardKernel)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-3e8c61b7a2ef","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2498,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:66"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_rewardStreamOfUnitArmStream {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Measurable (rewardStreamOfUnitArmStream rewardKernel)","missing":[],"search":"measurable_rewardstreamofunitarmstream banditrlproof.thompson.measurable_rewardstreamofunitarmstream theorem measurable_rewardstreamofunitarmstream {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] : measurable (rewardstreamofunitarmstream rewardkernel) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryRewardKernelAt","label":"stationaryRewardKernelAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryRewardKernelAt","description":"The arm-indexed reward kernel obtained by freezing the environment.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-3c9d4caaa403","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2499,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:77"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryRewardKernelAt {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) : Kernel (Fin K) Real","missing":[],"search":"stationaryrewardkernelat banditrlproof.thompson.stationaryrewardkernelat the arm-indexed reward kernel obtained by freezing the environment. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryRewardKernelAt_apply","label":"stationaryRewardKernelAt_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryRewardKernelAt_apply","description":"theorem stationaryRewardKernelAt_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) (arm : Fin K) : stationaryRewardKernelAt rewardKernel env arm = rewardKernel (env, arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-6d01f9e8920a","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2500,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:91"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryRewardKernelAt_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) (arm : Fin K) : stationaryRewardKernelAt rewardKernel env arm = rewardKernel (env, arm)","missing":[],"search":"stationaryrewardkernelat_apply banditrlproof.thompson.stationaryrewardkernelat_apply theorem stationaryrewardkernelat_apply {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] (env : env) (arm : fin k) : stationaryrewardkernelat rewardkernel env arm = rewardkernel (env, arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryArmStreamKernel","label":"stationaryArmStreamKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryArmStreamKernel","description":"Markov kernel that samples an independent latent reward stream at each environment.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-900f10e7339b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2501,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:99"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryArmStreamKernel {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Kernel Env (UCB.ArmRewardStream K)","missing":[],"search":"stationaryarmstreamkernel banditrlproof.thompson.stationaryarmstreamkernel markov kernel that samples an independent latent reward stream at each environment. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryArmStreamKernel_apply","label":"stationaryArmStreamKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryArmStreamKernel_apply","description":"At a fixed environment, the sampled latent table has the canonical product arm-stream law.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-2230797d56a0","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2502,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:115"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryArmStreamKernel_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) : stationaryArmStreamKernel rewardKernel env = UCB.armStreamMeasure (stationaryRewardKernelAt rewardKernel env)","missing":[],"search":"stationaryarmstreamkernel_apply banditrlproof.thompson.stationaryarmstreamkernel_apply at a fixed environment, the sampled latent table has the canonical product arm-stream law. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment","label":"stationaryMeasurableHistoryEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment","description":"A stationary reward kernel is a measurable history environment that ignores history.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-0f9650178d4b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2503,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:176"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryMeasurableHistoryEnvironment {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : MeasurableHistoryEnvironment Env (Fin K) Real where","missing":[],"search":"stationarymeasurablehistoryenvironment banditrlproof.thompson.stationarymeasurablehistoryenvironment a stationary reward kernel is a measurable history environment that ignores history. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_initialFeedback","label":"stationaryMeasurableHistoryEnvironment_initialFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_initialFeedback","description":"theorem stationaryMeasurableHistoryEnvironment_initialFeedback {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : (stationaryMeasurableHistoryEnvironment rewardKernel).initialFeedback = rewardKernel","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-a3cc0f0d5d8f","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2504,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:185"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryMeasurableHistoryEnvironment_initialFeedback {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : (stationaryMeasurableHistoryEnvironment rewardKernel).initialFeedback = rewardKernel","missing":[],"search":"stationarymeasurablehistoryenvironment_initialfeedback banditrlproof.thompson.stationarymeasurablehistoryenvironment_initialfeedback theorem stationarymeasurablehistoryenvironment_initialfeedback {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] : (stationarymeasurablehistoryenvironment rewardkernel).initialfeedback = rewardkernel theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_feedback_apply","label":"stationaryMeasurableHistoryEnvironment_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_feedback_apply","description":"theorem stationaryMeasurableHistoryEnvironment_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : (stationaryMeasurableHistoryEnvironment rewardKernel).feedback n (env, (history, arm)) = rewardKernel (env, arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-41df03efeedd","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2505,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:192"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryMeasurableHistoryEnvironment_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : (stationaryMeasurableHistoryEnvironment rewardKernel).feedback n (env, (history, arm)) = rewardKernel (env, arm)","missing":[],"search":"stationarymeasurablehistoryenvironment_feedback_apply banditrlproof.thompson.stationarymeasurablehistoryenvironment_feedback_apply theorem stationarymeasurablehistoryenvironment_feedback_apply {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] (n : nat) (env : env) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : (stationarymeasurablehistoryenvironment rewardkernel).feedback n (env, (history, arm)) = rewardkernel (env, arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_at_initialFeedback_apply","label":"stationaryMeasurableHistoryEnvironment_at_initialFeedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_at_initialFeedback_apply","description":"theorem stationaryMeasurableHistoryEnvironment_at_initialFeedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) (arm : Fin K) : ((stationaryMeasurableHistoryEnvironment rewardKernel).at env).initialFeedback arm = rewardKernel (env, arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-56f849fb55ac","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2506,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:203"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryMeasurableHistoryEnvironment_at_initialFeedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) (arm : Fin K) : ((stationaryMeasurableHistoryEnvironment rewardKernel).at env).initialFeedback arm = rewardKernel (env, arm)","missing":[],"search":"stationarymeasurablehistoryenvironment_at_initialfeedback_apply banditrlproof.thompson.stationarymeasurablehistoryenvironment_at_initialfeedback_apply theorem stationarymeasurablehistoryenvironment_at_initialfeedback_apply {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] (env : env) (arm : fin k) : ((stationarymeasurablehistoryenvironment rewardkernel).at env).initialfeedback arm = rewardkernel (env, arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_at_feedback_apply","label":"stationaryMeasurableHistoryEnvironment_at_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_at_feedback_apply","description":"theorem stationaryMeasurableHistoryEnvironment_at_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : ((stationaryMeasurableHistoryEnvironment rewardKernel).at env).feedback n (history, arm) = rewardKernel (env, arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-97553bd99a0c","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2507,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:213"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryMeasurableHistoryEnvironment_at_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (env : Env) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : ((stationaryMeasurableHistoryEnvironment rewardKernel).at env).feedback n (history, arm) = rewardKernel (env, arm)","missing":[],"search":"stationarymeasurablehistoryenvironment_at_feedback_apply banditrlproof.thompson.stationarymeasurablehistoryenvironment_at_feedback_apply theorem stationarymeasurablehistoryenvironment_at_feedback_apply {env : type u} {k : nat} [measurablespace env] (rewardkernel : kernel (env × fin k) real) [ismarkovkernel rewardkernel] (env : env) (n : nat) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : ((stationarymeasurablehistoryenvironment rewardkernel).at env).feedback n (history, arm) = rewardkernel (env, arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamInitialReward","label":"latentArmStreamInitialReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamInitialReward","description":"Initial reward read from coordinate zero of the selected latent arm.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-61317951e056","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2508,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:225"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def latentArmStreamInitialReward {Env : Type u} {K : Nat} : ((Env × UCB.ArmRewardStream K) × Fin K) → Real","missing":[],"search":"latentarmstreaminitialreward banditrlproof.thompson.latentarmstreaminitialreward initial reward read from coordinate zero of the selected latent arm. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamInitialReward","label":"measurable_latentArmStreamInitialReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamInitialReward","description":"theorem measurable_latentArmStreamInitialReward {Env : Type u} {K : Nat} [MeasurableSpace Env] : Measurable (latentArmStreamInitialReward (Env := Env) (K := K))","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-b95a5cf73c1b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2509,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:230"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamInitialReward {Env : Type u} {K : Nat} [MeasurableSpace Env] : Measurable (latentArmStreamInitialReward (Env := Env) (K := K))","missing":[],"search":"measurable_latentarmstreaminitialreward banditrlproof.thompson.measurable_latentarmstreaminitialreward theorem measurable_latentarmstreaminitialreward {env : type u} {k : nat} [measurablespace env] : measurable (latentarmstreaminitialreward (env := env) (k := k)) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamNextReward","label":"latentArmStreamNextReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamNextReward","description":"Next unused reward coordinate determined by the visible finite history.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-fc517c9a5d3a","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2510,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:238"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def latentArmStreamNextReward {Env : Type u} {K : Nat} (n : Nat) : ((Env × UCB.ArmRewardStream K) × (History.FinitePairHistory (Fin K) Real n × Fin K)) → Real","missing":[],"search":"latentarmstreamnextreward banditrlproof.thompson.latentarmstreamnextreward next unused reward coordinate determined by the visible finite history. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamNextReward","label":"measurable_latentArmStreamNextReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamNextReward","description":"theorem measurable_latentArmStreamNextReward {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) : Measurable (latentArmStreamNextReward (Env := Env) (K := K) n)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-c513f444cc3e","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2511,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:245"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamNextReward {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) : Measurable (latentArmStreamNextReward (Env := Env) (K := K) n)","missing":[],"search":"measurable_latentarmstreamnextreward banditrlproof.thompson.measurable_latentarmstreamnextreward theorem measurable_latentarmstreamnextreward {env : type u} {k : nat} [measurablespace env] (n : nat) : measurable (latentarmstreamnextreward (env := env) (k := k) n) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment","label":"latentArmStreamMeasurableHistoryEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment","description":"Measurable deterministic feedback environment that exposes a latent reward table through the next-unused-coordinate rule. The first environment component is retained for the Bayesian model but does not affect feedback.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-00ed196a90f5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2512,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:267"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def latentArmStreamMeasurableHistoryEnvironment {Env : Type u} {K : Nat} [MeasurableSpace Env] : MeasurableHistoryEnvironment (Env × UCB.ArmRewardStream K) (Fin K) Real where","missing":[],"search":"latentarmstreammeasurablehistoryenvironment banditrlproof.thompson.latentarmstreammeasurablehistoryenvironment measurable deterministic feedback environment that exposes a latent reward table through the next-unused-coordinate rule. the first environment component is retained for the bayesian model but does not affect feedback. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_initialFeedback_apply","label":"latentArmStreamMeasurableHistoryEnvironment_initialFeedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_initialFeedback_apply","description":"theorem latentArmStreamMeasurableHistoryEnvironment_initialFeedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (env : Env) (stream : UCB.ArmRewardStream K) (arm : Fin K) : (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).initialFeedback ((env, stream), arm) = Measure.dirac (stream 0 arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-7467d204c022","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2513,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:277"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamMeasurableHistoryEnvironment_initialFeedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (env : Env) (stream : UCB.ArmRewardStream K) (arm : Fin K) : (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).initialFeedback ((env, stream), arm) = Measure.dirac (stream 0 arm)","missing":[],"search":"latentarmstreammeasurablehistoryenvironment_initialfeedback_apply banditrlproof.thompson.latentarmstreammeasurablehistoryenvironment_initialfeedback_apply theorem latentarmstreammeasurablehistoryenvironment_initialfeedback_apply {env : type u} {k : nat} [measurablespace env] (env : env) (stream : ucb.armrewardstream k) (arm : fin k) : (latentarmstreammeasurablehistoryenvironment (env := env) (k := k)).initialfeedback ((env, stream), arm) = measure.dirac (stream 0 arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_feedback_apply","label":"latentArmStreamMeasurableHistoryEnvironment_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_feedback_apply","description":"theorem latentArmStreamMeasurableHistoryEnvironment_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) (env : Env) (stream : UCB.ArmRewardStream K) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).feedback n ((env, stream), (history, arm)) = Measure.dirac (stream (ETC.realHistoryPullCount n history arm) arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-c15c585376e7","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2514,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:286"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamMeasurableHistoryEnvironment_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) (env : Env) (stream : UCB.ArmRewardStream K) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).feedback n ((env, stream), (history, arm)) = Measure.dirac (stream (ETC.realHistoryPullCount n history arm) arm)","missing":[],"search":"latentarmstreammeasurablehistoryenvironment_feedback_apply banditrlproof.thompson.latentarmstreammeasurablehistoryenvironment_feedback_apply theorem latentarmstreammeasurablehistoryenvironment_feedback_apply {env : type u} {k : nat} [measurablespace env] (n : nat) (env : env) (stream : ucb.armrewardstream k) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : (latentarmstreammeasurablehistoryenvironment (env := env) (k := k)).feedback n ((env, stream), (history, arm)) = measure.dirac (stream (etc.realhistorypullcount n history arm) arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamNextReward_fixed","label":"measurable_latentArmStreamNextReward_fixed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamNextReward_fixed","description":"theorem measurable_latentArmStreamNextReward_fixed {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) (env : Env) (stream : UCB.ArmRewardStream K) : Measurable (fun input : History.FinitePairHistory (Fin K) Real n × Fin K => stream (ETC.realHistoryPullCount n input.1 input.2) input.2)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-63e36d3dbde4","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2515,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:296"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamNextReward_fixed {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) (env : Env) (stream : UCB.ArmRewardStream K) : Measurable (fun input : History.FinitePairHistory (Fin K) Real n × Fin K => stream (ETC.realHistoryPullCount n input.1 input.2) input.2)","missing":[],"search":"measurable_latentarmstreamnextreward_fixed banditrlproof.thompson.measurable_latentarmstreamnextreward_fixed theorem measurable_latentarmstreamnextreward_fixed {env : type u} {k : nat} [measurablespace env] (n : nat) (env : env) (stream : ucb.armrewardstream k) : measurable (fun input : history.finitepairhistory (fin k) real n × fin k => stream (etc.realhistorypullcount n input.1 input.2) input.2) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_at_initialFeedback_apply","label":"latentArmStreamMeasurableHistoryEnvironment_at_initialFeedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_at_initialFeedback_apply","description":"theorem latentArmStreamMeasurableHistoryEnvironment_at_initialFeedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (env : Env) (stream : UCB.ArmRewardStream K) (arm : Fin K) : ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream)).initialFeedback arm = Measure.dirac (stream 0 arm)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-97685eb52105","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2516,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:311"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamMeasurableHistoryEnvironment_at_initialFeedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (env : Env) (stream : UCB.ArmRewardStream K) (arm : Fin K) : ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream)).initialFeedback arm = Measure.dirac (stream 0 arm)","missing":[],"search":"latentarmstreammeasurablehistoryenvironment_at_initialfeedback_apply banditrlproof.thompson.latentarmstreammeasurablehistoryenvironment_at_initialfeedback_apply theorem latentarmstreammeasurablehistoryenvironment_at_initialfeedback_apply {env : type u} {k : nat} [measurablespace env] (env : env) (stream : ucb.armrewardstream k) (arm : fin k) : ((latentarmstreammeasurablehistoryenvironment (env := env) (k := k)).at (env, stream)).initialfeedback arm = measure.dirac (stream 0 arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_at_feedback_apply","label":"latentArmStreamMeasurableHistoryEnvironment_at_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_at_feedback_apply","description":"theorem latentArmStreamMeasurableHistoryEnvironment_at_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) (env : Env) (stream : UCB.ArmRewardStream K) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream)).feedback n (history, arm) = Measure.dirac (stream (ETC.realHistoryPullCount n history arm)…","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-52cb8e080e1b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2517,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:321"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamMeasurableHistoryEnvironment_at_feedback_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (n : Nat) (env : Env) (stream : UCB.ArmRewardStream K) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : ((latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)).at (env, stream)).feedback n (history, arm) = Measure.dirac (stream (ETC.realHistoryPullCount n history arm) arm)","missing":[],"search":"latentarmstreammeasurablehistoryenvironment_at_feedback_apply banditrlproof.thompson.latentarmstreammeasurablehistoryenvironment_at_feedback_apply theorem latentarmstreammeasurablehistoryenvironment_at_feedback_apply {env : type u} {k : nat} [measurablespace env] (n : nat) (env : env) (stream : ucb.armrewardstream k) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : ((latentarmstreammeasurablehistoryenvironment (env := env) (k := k)).at (env, stream)).feedback n (history, arm) = measure.dirac (stream (etc.realhistorypullcount n history arm) arm) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_zero_ae","label":"canonicalLatentArmStreamTrajectory_reward_zero_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_zero_ae","description":"The initial canonical reward reads coordinate zero of the selected arm.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-5c35b166842d","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2518,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:332"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalLatentArmStreamTrajectory_reward_zero_ae {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (stream : UCB.ArmRewardStream K) : ∀ᵐ trajectory ∂ canonicalMeasurableEnvironmentTrajectoryKernel algorithm (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)) (env, stream), (trajectory 0).2 = stream 0 (trajectory 0).1","missing":[],"search":"canonicallatentarmstreamtrajectory_reward_zero_ae banditrlproof.thompson.canonicallatentarmstreamtrajectory_reward_zero_ae the initial canonical reward reads coordinate zero of the selected arm. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_succ_ae","label":"canonicalLatentArmStreamTrajectory_reward_succ_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_succ_ae","description":"Every shifted canonical reward reads the selected arm's next unused coordinate.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-1ea492a6d6a0","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2519,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:376"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalLatentArmStreamTrajectory_reward_succ_ae {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (stream : UCB.ArmRewardStream K) (n : Nat) : ∀ᵐ trajectory ∂ canonicalMeasurableEnvironmentTrajectoryKernel algorithm (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)) (env, stream), (trajectory (n + 1)).2 = stream (pullCount (canonicalHistoryTrajectoryAction (Reward := Real) trajectory) (trajectory (n + 1)).1 (n + 1)) (trajectory (n + 1)).1","missing":[],"search":"canonicallatentarmstreamtrajectory_reward_succ_ae banditrlproof.thompson.canonicallatentarmstreamtrajectory_reward_succ_ae every shifted canonical reward reads the selected arm's next unused coordinate. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","label":"canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","description":"The canonical trajectory under deterministic latent-stream feedback has the same reward trace as the pathwise next-unused-coordinate construction.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-031ca318d6e0","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2520,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:455"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (stream : UCB.ArmRewardStream K) : canonicalHistoryTrajectoryReward =ᵐ[ canonicalMeasurableEnvironmentTrajectoryKernel algorithm (latentArmStreamMeasurableHistoryEnvironment (Env := Env) (K := K)) (env, stream)] UCB.rewardFromArmStream canonicalHistoryTrajectoryAction (fun _ => stream)","missing":[],"search":"canonicallatentarmstreamtrajectory_reward_eq_rewardfromarmstream_ae banditrlproof.thompson.canonicallatentarmstreamtrajectory_reward_eq_rewardfromarmstream_ae the canonical trajectory under deterministic latent-stream feedback has the same reward trace as the pathwise next-unused-coordinate construction. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_rewardFromArmStream_apply","label":"measurable_rewardFromArmStream_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_rewardFromArmStream_apply","description":"A next-unused arm-stream reward coordinate is measurable at each time.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-f217968268cb","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2521,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:493"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_rewardFromArmStream_apply {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (action : Omega -> ActionTrace (Fin K)) (haction : forall t, Measurable (fun omega => action omega t)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) (t : Nat) : Measurable (fun omega => rewardFromArmStream action armStream omega t)","missing":[],"search":"measurable_rewardfromarmstream_apply banditrlproof.ucb.measurable_rewardfromarmstream_apply a next-unused arm-stream reward coordinate is measurable at each time. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","label":"measure_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","description":"Adaptive-count upper tail on any sample space whose latent-stream projection has the canonical stationary product law. The action trace may use additional algorithmic randomness carried by the sample space.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-93a6a3ab16b8","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2522,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:524"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) (hstreamLaw : IdentDistrib armStream id mu (armStreamMeasure nu)) (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : mu {omega : Omega | threshold (pullCount (action omega) arm n) <= sumRewards (action omega) (rewardFromArmStream action armStream omega) arm n - (pullCount (action omega) arm n : Real) * mean} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(thresho…","missing":[],"search":"measure_sumrewards_sub_pullcount_mul_ge_le_of_armstream_identdistrib banditrlproof.ucb.measure_sumrewards_sub_pullcount_mul_ge_le_of_armstream_identdistrib adaptive-count upper tail on any sample space whose latent-stream projection has the canonical stationary product law. the action trace may use additional algorithmic randomness carried by the sample space. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","label":"measure_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","description":"Lower-tail counterpart of the arbitrary-sample latent-stream transport.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-0922fda6c8cf","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2523,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:583"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) (hstreamLaw : IdentDistrib armStream id mu (armStreamMeasure nu)) (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : mu {omega : Omega | threshold (pullCount (action omega) arm n) <= (pullCount (action omega) arm n : Real) * mean - sumRewards (action omega) (rewardFromArmStream action armStream omega) arm n} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(thresho…","missing":[],"search":"measure_pullcount_mul_sub_sumrewards_ge_le_of_armstream_identdistrib banditrlproof.ucb.measure_pullcount_mul_sub_sumrewards_ge_le_of_armstream_identdistrib lower-tail counterpart of the arbitrary-sample latent-stream transport. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pos_and_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","label":"measure_pos_and_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pos_and_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","description":"Positive-count upper tail under an arbitrary algorithmic-randomness coupling.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-b0948aae3b09","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2524,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:642"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_pos_and_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) (hstreamLaw : IdentDistrib armStream id mu (armStreamMeasure nu)) (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : mu {omega : Omega | 0 < pullCount (action omega) arm n ∧ threshold (pullCount (action omega) arm n) <= sumRewards (action omega) (rewardFromArmStream action armStream omega) arm n - (pullCount (action omega) arm n : Real) * mean} <= ((Finset.range (n…","missing":[],"search":"measure_pos_and_sumrewards_sub_pullcount_mul_ge_le_of_armstream_identdistrib banditrlproof.ucb.measure_pos_and_sumrewards_sub_pullcount_mul_ge_le_of_armstream_identdistrib positive-count upper tail under an arbitrary algorithmic-randomness coupling. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pos_and_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","label":"measure_pos_and_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pos_and_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","description":"Positive-count lower tail under an arbitrary algorithmic-randomness coupling.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-21e5e7dec2ca","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2525,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:715"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_pos_and_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) (hstreamLaw : IdentDistrib armStream id mu (armStreamMeasure nu)) (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : mu {omega : Omega | 0 < pullCount (action omega) arm n ∧ threshold (pullCount (action omega) arm n) <= (pullCount (action omega) arm n : Real) * mean - sumRewards (action omega) (rewardFromArmStream action armStream omega) arm n} <= ((Finset.range (n…","missing":[],"search":"measure_pos_and_pullcount_mul_sub_sumrewards_ge_le_of_armstream_identdistrib banditrlproof.ucb.measure_pos_and_pullcount_mul_sub_sumrewards_ge_le_of_armstream_identdistrib positive-count lower tail under an arbitrary algorithmic-randomness coupling. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le_of_canonicalArmStream","label":"measure_sumRewards_sub_pullCount_mul_ge_le_of_canonicalArmStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le_of_canonicalArmStream","description":"Adaptive-count upper tail for any action trace driven by a canonical latent arm stream. No UCB action rule is used.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-801a91252751","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2526,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:790"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_sumRewards_sub_pullCount_mul_ge_le_of_canonicalArmStream {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : ArmRewardStream K -> ActionTrace (Fin K)) (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : armStreamMeasure nu {stream : ArmRewardStream K | threshold (pullCount (action stream) arm n) <= sumRewards (action stream) (rewardFromArmStream action id stream) arm n - (pullCount (action stream) arm n : Real) * mean} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_sumrewards_sub_pullcount_mul_ge_le_of_canonicalarmstream banditrlproof.ucb.measure_sumrewards_sub_pullcount_mul_ge_le_of_canonicalarmstream adaptive-count upper tail for any action trace driven by a canonical latent arm stream. no ucb action rule is used. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le_of_canonicalArmStream","label":"measure_pullCount_mul_sub_sumRewards_ge_le_of_canonicalArmStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le_of_canonicalArmStream","description":"Adaptive-count lower tail for any action trace driven by a canonical latent arm stream. No UCB action rule is used.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-afc1e30dcfbf","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2527,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:838"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_mul_sub_sumRewards_ge_le_of_canonicalArmStream {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : ArmRewardStream K -> ActionTrace (Fin K)) (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : armStreamMeasure nu {stream : ArmRewardStream K | threshold (pullCount (action stream) arm n) <= (pullCount (action stream) arm n : Real) * mean - sumRewards (action stream) (rewardFromArmStream action id stream) arm n} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_pullcount_mul_sub_sumrewards_ge_le_of_canonicalarmstream banditrlproof.ucb.measure_pullcount_mul_sub_sumrewards_ge_le_of_canonicalarmstream adaptive-count lower tail for any action trace driven by a canonical latent arm stream. no ucb action rule is used. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel","label":"latentArmStreamTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryKernel","description":"Canonical trajectory kernel after freezing the non-stream environment.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-b754880eb82d","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2528,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:887"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def latentArmStreamTrajectoryKernel {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) : Kernel (UCB.ArmRewardStream K) ((n : Nat) -> Fin K × Real)","missing":[],"search":"latentarmstreamtrajectorykernel banditrlproof.thompson.latentarmstreamtrajectorykernel canonical trajectory kernel after freezing the non-stream environment. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure","label":"latentArmStreamTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure","description":"Joint law of a stationary latent arm stream and its generated trajectory.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-8a8519c5e048","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2529,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:903"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def latentArmStreamTrajectoryMeasure {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Measure (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real))","missing":[],"search":"latentarmstreamtrajectorymeasure banditrlproof.thompson.latentarmstreamtrajectorymeasure joint law of a stationary latent arm stream and its generated trajectory. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryAction","label":"latentArmStreamTrajectoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryAction","description":"Action trace projected from the trajectory coordinate of the coupling.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-be8617bb8ded","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2530,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:919"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def latentArmStreamTrajectoryAction {K : Nat} : (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)) -> ActionTrace (Fin K)","missing":[],"search":"latentarmstreamtrajectoryaction banditrlproof.thompson.latentarmstreamtrajectoryaction action trace projected from the trajectory coordinate of the coupling. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryReward","label":"latentArmStreamTrajectoryReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryReward","description":"Reward trace projected from the trajectory coordinate of the coupling.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-cc0d0a9aeed5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2531,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:926"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def latentArmStreamTrajectoryReward {K : Nat} : (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)) -> RewardTrace Real","missing":[],"search":"latentarmstreamtrajectoryreward banditrlproof.thompson.latentarmstreamtrajectoryreward reward trace projected from the trajectory coordinate of the coupling. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamTrajectoryAction_apply","label":"measurable_latentArmStreamTrajectoryAction_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamTrajectoryAction_apply","description":"theorem measurable_latentArmStreamTrajectoryAction_apply {K : Nat} (t : Nat) : Measurable (fun sample : UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real) => latentArmStreamTrajectoryAction sample t)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-3ed5c8c7bb9b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2532,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:932"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamTrajectoryAction_apply {K : Nat} (t : Nat) : Measurable (fun sample : UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real) => latentArmStreamTrajectoryAction sample t)","missing":[],"search":"measurable_latentarmstreamtrajectoryaction_apply banditrlproof.thompson.measurable_latentarmstreamtrajectoryaction_apply theorem measurable_latentarmstreamtrajectoryaction_apply {k : nat} (t : nat) : measurable (fun sample : ucb.armrewardstream k × ((n : nat) -> fin k × real) => latentarmstreamtrajectoryaction sample t) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamTrajectoryReward_apply","label":"measurable_latentArmStreamTrajectoryReward_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_latentArmStreamTrajectoryReward_apply","description":"theorem measurable_latentArmStreamTrajectoryReward_apply {K : Nat} (t : Nat) : Measurable (fun sample : UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real) => latentArmStreamTrajectoryReward sample t)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-213136218a28","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2533,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:939"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_latentArmStreamTrajectoryReward_apply {K : Nat} (t : Nat) : Measurable (fun sample : UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real) => latentArmStreamTrajectoryReward sample t)","missing":[],"search":"measurable_latentarmstreamtrajectoryreward_apply banditrlproof.thompson.measurable_latentarmstreamtrajectoryreward_apply theorem measurable_latentarmstreamtrajectoryreward_apply {k : nat} (t : nat) : measurable (fun sample : ucb.armrewardstream k × ((n : nat) -> fin k × real) => latentarmstreamtrajectoryreward sample t) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_pullCount_selectedArm","label":"measurable_pullCount_selectedArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_pullCount_selectedArm","description":"Pull counts remain measurable when the queried arm is sample-dependent.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-94cb719c9693","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2534,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:947"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_pullCount_selectedArm {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (action : Omega -> ActionTrace (Fin K)) (haction : forall t, Measurable (fun omega => action omega t)) (arm : Omega -> Fin K) (harm : Measurable arm) (n : Nat) : Measurable (fun omega => pullCount (action omega) (arm omega) n)","missing":[],"search":"measurable_pullcount_selectedarm banditrlproof.thompson.measurable_pullcount_selectedarm pull counts remain measurable when the queried arm is sample-dependent. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_sumRewards_selectedArm","label":"measurable_sumRewards_selectedArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_sumRewards_selectedArm","description":"Selected-arm reward sums are measurable for a measurable arm selector.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-a18aeaa29585","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2535,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:964"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_sumRewards_selectedArm {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (arm : Omega -> Fin K) (harm : Measurable arm) (n : Nat) : Measurable (fun omega => sumRewards (action omega) (reward omega) (arm omega) n)","missing":[],"search":"measurable_sumrewards_selectedarm banditrlproof.thompson.measurable_sumrewards_selectedarm selected-arm reward sums are measurable for a measurable arm selector. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_realEmpiricalMean_selectedArm","label":"measurable_realEmpiricalMean_selectedArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_realEmpiricalMean_selectedArm","description":"The real empirical mean is measurable for a measurable arm selector.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-2a2c8c4345dd","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2536,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:984"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_realEmpiricalMean_selectedArm {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (arm : Omega -> Fin K) (harm : Measurable arm) (n : Nat) : Measurable (fun omega => UCB.realEmpiricalMean (action omega) (reward omega) (arm omega) n)","missing":[],"search":"measurable_realempiricalmean_selectedarm banditrlproof.thompson.measurable_realempiricalmean_selectedarm the real empirical mean is measurable for a measurable arm selector. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.identDistrib_fst_latentArmStreamTrajectoryMeasure","label":"identDistrib_fst_latentArmStreamTrajectoryMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.identDistrib_fst_latentArmStreamTrajectoryMeasure","description":"The stream coordinate of the coupling has exactly the canonical stream law.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-eb6c21dd8651","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2537,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:999"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_fst_latentArmStreamTrajectoryMeasure {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : IdentDistrib Prod.fst id (latentArmStreamTrajectoryMeasure algorithm env nu) (UCB.armStreamMeasure nu) where","missing":[],"search":"identdistrib_fst_latentarmstreamtrajectorymeasure banditrlproof.thompson.identdistrib_fst_latentarmstreamtrajectorymeasure the stream coordinate of the coupling has exactly the canonical stream law. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryReward_eq_rewardFromArmStream_ae","label":"latentArmStreamTrajectoryReward_eq_rewardFromArmStream_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.latentArmStreamTrajectoryReward_eq_rewardFromArmStream_ae","description":"The joint coupling reads its actual rewards from its own latent stream.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-849bafcbb4ef","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2538,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1015"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem latentArmStreamTrajectoryReward_eq_rewardFromArmStream_ae {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : latentArmStreamTrajectoryReward =ᵐ[ latentArmStreamTrajectoryMeasure algorithm env nu] UCB.rewardFromArmStream latentArmStreamTrajectoryAction Prod.fst","missing":[],"search":"latentarmstreamtrajectoryreward_eq_rewardfromarmstream_ae banditrlproof.thompson.latentarmstreamtrajectoryreward_eq_rewardfromarmstream_ae the joint coupling reads its actual rewards from its own latent stream. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_sumRewards_sub_pullCount_mul_ge_le","label":"measure_latentArmStreamTrajectory_sumRewards_sub_pullCount_mul_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_sumRewards_sub_pullCount_mul_ge_le","description":"Upper adaptive-count tail for rewards on the coupled Thompson trajectory.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-f63bdaa24415","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2539,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1055"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_sumRewards_sub_pullCount_mul_ge_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n - (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real…","missing":[],"search":"measure_latentarmstreamtrajectory_sumrewards_sub_pullcount_mul_ge_le banditrlproof.thompson.measure_latentarmstreamtrajectory_sumrewards_sub_pullcount_mul_ge_le upper adaptive-count tail for rewards on the coupled thompson trajectory. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pullCount_mul_sub_sumRewards_ge_le","label":"measure_latentArmStreamTrajectory_pullCount_mul_sub_sumRewards_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pullCount_mul_sub_sumRewards_ge_le","description":"Lower adaptive-count tail for rewards on the coupled Thompson trajectory.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-7da747c64696","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2540,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1119"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_pullCount_mul_sub_sumRewards_ge_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean - sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real…","missing":[],"search":"measure_latentarmstreamtrajectory_pullcount_mul_sub_sumrewards_ge_le banditrlproof.thompson.measure_latentarmstreamtrajectory_pullcount_mul_sub_sumrewards_ge_le lower adaptive-count tail for rewards on the coupled thompson trajectory. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pos_and_sumRewards_sub_pullCount_mul_ge_le","label":"measure_latentArmStreamTrajectory_pos_and_sumRewards_sub_pullCount_mul_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pos_and_sumRewards_sub_pullCount_mul_ge_le","description":"Positive-count upper tail for rewards on the coupled trajectory.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-dd64f024baeb","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2541,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1183"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_pos_and_sumRewards_sub_pullCount_mul_ge_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | 0 < pullCount (latentArmStreamTrajectoryAction sample) arm n ∧ threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n - (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean} <= ((Finset.range (n + 1)).filter (fun…","missing":[],"search":"measure_latentarmstreamtrajectory_pos_and_sumrewards_sub_pullcount_mul_ge_le banditrlproof.thompson.measure_latentarmstreamtrajectory_pos_and_sumrewards_sub_pullcount_mul_ge_le positive-count upper tail for rewards on the coupled trajectory. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pos_and_pullCount_mul_sub_sumRewards_ge_le","label":"measure_latentArmStreamTrajectory_pos_and_pullCount_mul_sub_sumRewards_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pos_and_pullCount_mul_sub_sumRewards_ge_le","description":"Positive-count lower tail for rewards on the coupled trajectory.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-8bc81422a227","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2542,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1249"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_pos_and_pullCount_mul_sub_sumRewards_ge_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | 0 < pullCount (latentArmStreamTrajectoryAction sample) arm n ∧ threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean - sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n} <= ((Finset.range (n + 1)).filter (fun…","missing":[],"search":"measure_latentarmstreamtrajectory_pos_and_pullcount_mul_sub_sumrewards_ge_le banditrlproof.thompson.measure_latentarmstreamtrajectory_pos_and_pullcount_mul_sub_sumrewards_ge_le positive-count lower tail for rewards on the coupled trajectory. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel","label":"stationaryLatentArmStreamTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel","description":"Environment-indexed kernel that samples the stationary latent stream and then runs the actual history algorithm against next-unused deterministic feedback.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-a56a2e0df39f","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2543,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1318"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryLatentArmStreamTrajectoryKernel {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) : Kernel Env (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real))","missing":[],"search":"stationarylatentarmstreamtrajectorykernel banditrlproof.thompson.stationarylatentarmstreamtrajectorykernel environment-indexed kernel that samples the stationary latent stream and then runs the actual history algorithm against next-unused deterministic feedback. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_apply","label":"stationaryLatentArmStreamTrajectoryKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_apply","description":"Its conditional law is exactly the fixed-environment coupling above.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-8092ad6d17ff","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2544,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1340"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryKernel_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) : stationaryLatentArmStreamTrajectoryKernel rewardKernel algorithm env = latentArmStreamTrajectoryMeasure algorithm env (stationaryRewardKernelAt rewardKernel env)","missing":[],"search":"stationarylatentarmstreamtrajectorykernel_apply banditrlproof.thompson.stationarylatentarmstreamtrajectorykernel_apply its conditional law is exactly the fixed-environment coupling above. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_sumRewards_upper_tail","label":"stationaryLatentArmStreamTrajectoryKernel_sumRewards_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_sumRewards_upper_tail","description":"Pointwise upper tail for the actual augmented Thompson trajectory kernel.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-c47e96db2977","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2545,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1355"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryKernel_sumRewards_upper_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (env : Env) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryKernel rewardKernel algorithm env {sample | threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n - (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean env arm} <= (Finset.ran…","missing":[],"search":"stationarylatentarmstreamtrajectorykernel_sumrewards_upper_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorykernel_sumrewards_upper_tail pointwise upper tail for the actual augmented thompson trajectory kernel. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_sumRewards_lower_tail","label":"stationaryLatentArmStreamTrajectoryKernel_sumRewards_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_sumRewards_lower_tail","description":"Pointwise lower tail for the actual augmented Thompson trajectory kernel.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-9bdc652be61e","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2546,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1388"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryKernel_sumRewards_lower_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (env : Env) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryKernel rewardKernel algorithm env {sample | threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean env arm - sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n} <= (Finset.ran…","missing":[],"search":"stationarylatentarmstreamtrajectorykernel_sumrewards_lower_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorykernel_sumrewards_lower_tail pointwise lower tail for the actual augmented thompson trajectory kernel. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_upper_tail","label":"stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_upper_tail","description":"Positive-count upper tail for the environment-indexed augmented kernel.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-7c4991c64430","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2547,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1421"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_upper_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (env : Env) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryKernel rewardKernel algorithm env {sample | 0 < pullCount (latentArmStreamTrajectoryAction sample) arm n ∧ threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= sumRewards (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm n - (pullCount (late…","missing":[],"search":"stationarylatentarmstreamtrajectorykernel_pos_and_sumrewards_upper_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorykernel_pos_and_sumrewards_upper_tail positive-count upper tail for the environment-indexed augmented kernel. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_lower_tail","label":"stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_lower_tail","description":"Positive-count lower tail for the environment-indexed augmented kernel.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-1c06ce90cdbd","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2548,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1455"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_lower_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (env : Env) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryKernel rewardKernel algorithm env {sample | 0 < pullCount (latentArmStreamTrajectoryAction sample) arm n ∧ threshold (pullCount (latentArmStreamTrajectoryAction sample) arm n) <= (pullCount (latentArmStreamTrajectoryAction sample) arm n : Real) * mean env arm - sumRewards (latentArmStreamTraject…","missing":[],"search":"stationarylatentarmstreamtrajectorykernel_pos_and_sumrewards_lower_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorykernel_pos_and_sumrewards_lower_tail positive-count lower tail for the environment-indexed augmented kernel. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryEnvironment","label":"stationaryLatentArmStreamTrajectoryEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryEnvironment","description":"Environment coordinate of the stationary augmented trajectory sample.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-2c0a691fa606","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2549,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1489"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def stationaryLatentArmStreamTrajectoryEnvironment {Env : Type u} {K : Nat} : (Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real))) -> Env","missing":[],"search":"stationarylatentarmstreamtrajectoryenvironment banditrlproof.thompson.stationarylatentarmstreamtrajectoryenvironment environment coordinate of the stationary augmented trajectory sample. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryAction","label":"stationaryLatentArmStreamTrajectoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryAction","description":"Action trace of the stationary augmented trajectory sample.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-a7a41ca5936a","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2550,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1495"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def stationaryLatentArmStreamTrajectoryAction {Env : Type u} {K : Nat} : (Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real))) -> ActionTrace (Fin K)","missing":[],"search":"stationarylatentarmstreamtrajectoryaction banditrlproof.thompson.stationarylatentarmstreamtrajectoryaction action trace of the stationary augmented trajectory sample. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryReward","label":"stationaryLatentArmStreamTrajectoryReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryReward","description":"Reward trace of the stationary augmented trajectory sample.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-c2f03fb1efab","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2551,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1502"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"def stationaryLatentArmStreamTrajectoryReward {Env : Type u} {K : Nat} : (Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real))) -> RewardTrace Real","missing":[],"search":"stationarylatentarmstreamtrajectoryreward banditrlproof.thompson.stationarylatentarmstreamtrajectoryreward reward trace of the stationary augmented trajectory sample. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryEnvironment","label":"measurable_stationaryLatentArmStreamTrajectoryEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryEnvironment","description":"theorem measurable_stationaryLatentArmStreamTrajectoryEnvironment {Env : Type u} {K : Nat} [MeasurableSpace Env] : Measurable (stationaryLatentArmStreamTrajectoryEnvironment (Env := Env) (K := K))","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-dea654e28cbd","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2552,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1508"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_stationaryLatentArmStreamTrajectoryEnvironment {Env : Type u} {K : Nat} [MeasurableSpace Env] : Measurable (stationaryLatentArmStreamTrajectoryEnvironment (Env := Env) (K := K))","missing":[],"search":"measurable_stationarylatentarmstreamtrajectoryenvironment banditrlproof.thompson.measurable_stationarylatentarmstreamtrajectoryenvironment theorem measurable_stationarylatentarmstreamtrajectoryenvironment {env : type u} {k : nat} [measurablespace env] : measurable (stationarylatentarmstreamtrajectoryenvironment (env := env) (k := k)) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryAction_apply","label":"measurable_stationaryLatentArmStreamTrajectoryAction_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryAction_apply","description":"theorem measurable_stationaryLatentArmStreamTrajectoryAction_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (t : Nat) : Measurable (fun sample : Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)) => stationaryLatentArmStreamTrajectoryAction sample t)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-b1d7cced5bf2","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2553,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1514"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_stationaryLatentArmStreamTrajectoryAction_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (t : Nat) : Measurable (fun sample : Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)) => stationaryLatentArmStreamTrajectoryAction sample t)","missing":[],"search":"measurable_stationarylatentarmstreamtrajectoryaction_apply banditrlproof.thompson.measurable_stationarylatentarmstreamtrajectoryaction_apply theorem measurable_stationarylatentarmstreamtrajectoryaction_apply {env : type u} {k : nat} [measurablespace env] (t : nat) : measurable (fun sample : env × (ucb.armrewardstream k × ((n : nat) -> fin k × real)) => stationarylatentarmstreamtrajectoryaction sample t) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryReward_apply","label":"measurable_stationaryLatentArmStreamTrajectoryReward_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryReward_apply","description":"theorem measurable_stationaryLatentArmStreamTrajectoryReward_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (t : Nat) : Measurable (fun sample : Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)) => stationaryLatentArmStreamTrajectoryReward sample t)","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-9b3c5804a0fa","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2554,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1521"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measurable_stationaryLatentArmStreamTrajectoryReward_apply {Env : Type u} {K : Nat} [MeasurableSpace Env] (t : Nat) : Measurable (fun sample : Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)) => stationaryLatentArmStreamTrajectoryReward sample t)","missing":[],"search":"measurable_stationarylatentarmstreamtrajectoryreward_apply banditrlproof.thompson.measurable_stationarylatentarmstreamtrajectoryreward_apply theorem measurable_stationarylatentarmstreamtrajectoryreward_apply {env : type u} {k : nat} [measurablespace env] (t : nat) : measurable (fun sample : env × (ucb.armrewardstream k × ((n : nat) -> fin k × real)) => stationarylatentarmstreamtrajectoryreward sample t) theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure","label":"stationaryLatentArmStreamTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure","description":"Prior mixture of the stationary latent stream and its actual trajectory.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-632159fbafb0","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2555,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1529"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryLatentArmStreamTrajectoryMeasure {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (prior : Measure Env) (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) : Measure (Env × (UCB.ArmRewardStream K × ((n : Nat) -> Fin K × Real)))","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure prior mixture of the stationary latent stream and its actual trajectory. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_sumRewards_upper_tail","label":"stationaryLatentArmStreamTrajectoryMeasure_sumRewards_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_sumRewards_upper_tail","description":"Upper fixed-arm adaptive-count tail after mixing the fixed-environment augmented trajectory through the environment prior.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-4b962f9bcbf5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2556,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1553"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_sumRewards_upper_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | threshold (pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n) <= sumRewards (stationaryLatentArmStreamTrajectoryAction sample) (stati…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_sumrewards_upper_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_sumrewards_upper_tail upper fixed-arm adaptive-count tail after mixing the fixed-environment augmented trajectory through the environment prior. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_sumRewards_lower_tail","label":"stationaryLatentArmStreamTrajectoryMeasure_sumRewards_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_sumRewards_lower_tail","description":"Lower-tail counterpart of the augmented-prior mixture theorem.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-bfc0ce85f117","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2557,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1646"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_sumRewards_lower_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | threshold (pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n) <= (pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_sumrewards_lower_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_sumrewards_lower_tail lower-tail counterpart of the augmented-prior mixture theorem. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_compProd_le_of_forall_kernel_apply_le","label":"measure_compProd_le_of_forall_kernel_apply_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_compProd_le_of_forall_kernel_apply_le","description":"Integrate a uniform pointwise kernel event bound through a probability prior.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-c9fc7b731834","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2558,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1738"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_compProd_le_of_forall_kernel_apply_le {Env : Type u} {Omega : Type v} [MeasurableSpace Env] [MeasurableSpace Omega] (prior : Measure Env) [IsProbabilityMeasure prior] (kernel : Kernel Env Omega) [IsMarkovKernel kernel] (event : Set (Env × Omega)) (hevent : MeasurableSet event) (bound : ENNReal) (hbound : forall env, kernel env (Prod.mk env ⁻¹' event) <= bound) : (prior ⊗ₘ kernel) event <= bound","missing":[],"search":"measure_compprod_le_of_forall_kernel_apply_le banditrlproof.thompson.measure_compprod_le_of_forall_kernel_apply_le integrate a uniform pointwise kernel event bound through a probability prior. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_upper_tail","label":"stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_upper_tail","description":"Positive-count upper tail after mixing through the environment prior.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-ab865ea0ab74","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2559,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1754"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_upper_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n ∧ threshold (pullCount (stationaryLatentArmStreamTrajectoryAct…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_pos_and_sumrewards_upper_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_pos_and_sumrewards_upper_tail positive-count upper tail after mixing through the environment prior. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_lower_tail","label":"stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_lower_tail","description":"Positive-count lower tail after mixing through the environment prior.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-1fee28d007c6","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2560,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1834"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_lower_tail {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (arm : Fin K) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n ∧ threshold (pullCount (stationaryLatentArmStreamTrajectoryAct…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_pos_and_sumrewards_lower_tail banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_pos_and_sumrewards_lower_tail positive-count lower tail after mixing through the environment prior. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold","label":"clippedCountWidthThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedCountWidthThreshold","description":"Pull-count-scaled radius used by the clipped Thompson score.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-81adb11f4eab","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2561,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1913"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedCountWidthThreshold (sigma2 : NNReal) (delta : Real) (k : Nat) : Real","missing":[],"search":"clippedcountwidththreshold banditrlproof.thompson.clippedcountwidththreshold pull-count-scaled radius used by the clipped thompson score. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_nonneg","label":"clippedCountWidthThreshold_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedCountWidthThreshold_nonneg","description":"theorem clippedCountWidthThreshold_nonneg (sigma2 : NNReal) (delta : Real) (k : Nat) : 0 <= clippedCountWidthThreshold sigma2 delta k","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-1106a0eeee9e","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2562,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1919"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedCountWidthThreshold_nonneg (sigma2 : NNReal) (delta : Real) (k : Nat) : 0 <= clippedCountWidthThreshold sigma2 delta k","missing":[],"search":"clippedcountwidththreshold_nonneg banditrlproof.thompson.clippedcountwidththreshold_nonneg theorem clippedcountwidththreshold_nonneg (sigma2 : nnreal) (delta : real) (k : nat) : 0 <= clippedcountwidththreshold sigma2 delta k theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_sq_div_eq","label":"clippedCountWidthThreshold_sq_div_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedCountWidthThreshold_sq_div_eq","description":"The clipped-score count threshold has the intended confidence exponent.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-829ddae6c821","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2563,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1925"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedCountWidthThreshold_sq_div_eq (sigma2 : NNReal) (delta : Real) (k : Nat) (hsigma2 : sigma2 ≠ 0) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (hk : 0 < k) : (clippedCountWidthThreshold sigma2 delta k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)) = Real.log (1 / delta)","missing":[],"search":"clippedcountwidththreshold_sq_div_eq banditrlproof.thompson.clippedcountwidththreshold_sq_div_eq the clipped-score count threshold has the intended confidence exponent. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.sum_clippedCountWidthThreshold_tail_eq","label":"sum_clippedCountWidthThreshold_tail_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.sum_clippedCountWidthThreshold_tail_eq","description":"The positive-count clipped-radius exponential sum is exactly `n * delta`.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-ec78b4c27b76","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2564,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1948"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem sum_clippedCountWidthThreshold_tail_eq (sigma2 : NNReal) (delta : Real) (n : Nat) (hsigma2 : sigma2 ≠ 0) (hdelta : 0 < delta) (hdelta_one : delta <= 1) : ((Finset.range (n + 1)).filter (fun k => 0 < k)).sum (fun k => ENNReal.ofReal (Real.exp (-(clippedCountWidthThreshold sigma2 delta k) ^ 2 / (2 * (k : Real) * (sigma2 : Real))))) = (n : ENNReal) * ENNReal.ofReal delta","missing":[],"search":"sum_clippedcountwidththreshold_tail_eq banditrlproof.thompson.sum_clippedcountwidththreshold_tail_eq the positive-count clipped-radius exponential sum is exactly `n * delta`. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_biUnion_clippedCountWidthThreshold_le_mul_sub_armPrefixSum_le","label":"measure_biUnion_clippedCountWidthThreshold_le_mul_sub_armPrefixSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_biUnion_clippedCountWidthThreshold_le_mul_sub_armPrefixSum_le","description":"A finite union of lower prefix deviations pays once per positive pull count.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-5bef246ef35b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2565,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:1976"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_biUnion_clippedCountWidthThreshold_le_mul_sub_armPrefixSum_le {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (m : Nat) : UCB.armStreamMeasure nu (⋃ k ∈ Finset.Icc 1 m, {stream : UCB.ArmRewardStream K | clippedCountWidthThreshold sigma2 delta k <= (k : Real) * mean - UCB.armPrefixSum arm k stream}) <= (m : ENNReal) * ENNReal.ofReal delta","missing":[],"search":"measure_biunion_clippedcountwidththreshold_le_mul_sub_armprefixsum_le banditrlproof.thompson.measure_biunion_clippedcountwidththreshold_le_mul_sub_armprefixsum_le a finite union of lower prefix deviations pays once per positive pull count. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_le_mul_mean_sub_sumRewards","label":"clippedCountWidthThreshold_le_mul_mean_sub_sumRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedCountWidthThreshold_le_mul_mean_sub_sumRewards","description":"A clipped-score lower-confidence failure implies a lower sum deviation.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-4435234d0d78","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2566,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2015"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedCountWidthThreshold_le_mul_mean_sub_sumRewards {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (n : Nat) (mean : Real) (sigma2 : NNReal) (delta : Real) (hcount : 0 < pullCount action arm n) (hindex : UCB.realEmpiricalMean action reward arm n + Real.sqrt (2 * (sigma2 : Real) * Real.log (1 / delta) / (pullCount action arm n : Real)) <= mean) : clippedCountWidthThreshold sigma2 delta (pullCount action arm n) <= (pullCount action arm n : Real) * mean - sumRewards action reward arm n","missing":[],"search":"clippedcountwidththreshold_le_mul_mean_sub_sumrewards banditrlproof.thompson.clippedcountwidththreshold_le_mul_mean_sub_sumrewards a clipped-score lower-confidence failure implies a lower sum deviation. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_le_sumRewards_sub_mul_mean","label":"clippedCountWidthThreshold_le_sumRewards_sub_mul_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.clippedCountWidthThreshold_le_sumRewards_sub_mul_mean","description":"A clipped-score upper-confidence failure implies an upper sum deviation.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-0f128b6aa94b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2567,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2044"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem clippedCountWidthThreshold_le_sumRewards_sub_mul_mean {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (n : Nat) (mean : Real) (sigma2 : NNReal) (delta : Real) (hcount : 0 < pullCount action arm n) (hindex : mean <= UCB.realEmpiricalMean action reward arm n - Real.sqrt (2 * (sigma2 : Real) * Real.log (1 / delta) / (pullCount action arm n : Real))) : clippedCountWidthThreshold sigma2 delta (pullCount action arm n) <= sumRewards action reward arm n - (pullCount action arm n : Real) * mean","missing":[],"search":"clippedcountwidththreshold_le_sumrewards_sub_mul_mean banditrlproof.thompson.clippedcountwidththreshold_le_sumrewards_sub_mul_mean a clipped-score upper-confidence failure implies an upper sum deviation. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_realEmpiricalMean_add_width_le_mean_le","label":"measure_latentArmStreamTrajectory_exists_realEmpiricalMean_add_width_le_mean_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_realEmpiricalMean_add_width_le_mean_le","description":"Finite-horizon lower-confidence failure for one arm on the coupled trajectory. All times with the same realized pull count reduce to the same latent-stream prefix event, so the bound pays for `1, ..., n - 1` once rather than unioning the fixed-time bounds.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-cc1acb0fad6b","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2568,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2078"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_exists_realEmpiricalMean_add_width_le_mean_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | ∃ t < n, 0 < pullCount (latentArmStreamTrajectoryAction sample) arm t ∧ UCB.realEmpiricalMean (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm t + Real.sqrt (2 * (sigma2 : Real) * Real.log (1 / delta) / (pullCount (latentArmStreamTrajectoryAction sample) arm t : Real)) <= mean} <= (n - 1 : Nat) * ENNReal.ofRea…","missing":[],"search":"measure_latentarmstreamtrajectory_exists_realempiricalmean_add_width_le_mean_le banditrlproof.thompson.measure_latentarmstreamtrajectory_exists_realempiricalmean_add_width_le_mean_le finite-horizon lower-confidence failure for one arm on the coupled trajectory. all times with the same realized pull count reduce to the same latent-stream prefix event, so the bound pays for `1, ..., n - 1` once rather than unioning the fixed-time bounds. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","label":"stationaryLatentArmStreamTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","description":"Prior-mixed finite-horizon lower-confidence failure for an environment-dependent measurable arm selector.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-856f455d65c0","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2569,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2180"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (selectedArm : Env -> Fin K) (hselectedArm : Measurable selectedArm) (n : Nat) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | ∃ t < n, 0 < pullCount (stationaryLatentArmStreamTra…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_exists_selectedarm_realempiricalmean_add_width_le_mean_le banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_exists_selectedarm_realempiricalmean_add_width_le_mean_le prior-mixed finite-horizon lower-confidence failure for an environment-dependent measurable arm selector. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean","label":"stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean","description":"Prior-mixed lower-confidence failure bound for the clipped score radius.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-d06350c673f9","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2570,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2301"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (delta : Real) (arm : Fin K) (n : Nat) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n ∧ UCB.realEmpiricalMean (stationaryLatentArmStreamTrajectoryAction sample) (stationaryLatentArmStreamTrajectoryReward sample) ar…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_realempiricalmean_add_width_le_mean banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_realempiricalmean_add_width_le_mean prior-mixed lower-confidence failure bound for the clipped score radius. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width","label":"stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width","description":"Prior-mixed upper-confidence failure bound for the clipped score radius.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-644711ccd344","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2571,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2379"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (delta : Real) (arm : Fin K) (n : Nat) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n ∧ mean (stationaryLatentArmStreamTrajectoryEnvironment sample) arm <= UCB.realEmpiricalMean (stationaryLatentArmStreamTrajectory…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_mean_le_realempiricalmean_sub_width banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_mean_le_realempiricalmean_sub_width prior-mixed upper-confidence failure bound for the clipped score radius. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","label":"stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","description":"Simplified `n * delta` lower-confidence failure bound.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-8e4662072989","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2572,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2457"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (arm : Fin K) (n : Nat) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n ∧ UCB.realEmpiricalMean (stationaryLatent…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_realempiricalmean_add_width_le_mean_le_nat_mul_delta banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_realempiricalmean_add_width_le_mean_le_nat_mul_delta simplified `n * delta` lower-confidence failure bound. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","label":"stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","description":"Simplified `n * delta` upper-confidence failure bound.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-5ef1728878ad","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2573,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2502"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (arm : Fin K) (n : Nat) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm n ∧ mean (stationaryLatentArmStreamTrajecto…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_mean_le_realempiricalmean_sub_width_le_nat_mul_delta banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_mean_le_realempiricalmean_sub_width_le_nat_mul_delta simplified `n * delta` upper-confidence failure bound. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamPrior","label":"stationaryLatentArmStreamPrior","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamPrior","description":"Bayesian prior augmented with the stationary latent arm stream.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-4ba79a4916f5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2574,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2547"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryLatentArmStreamPrior {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (prior : Measure Env) (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] : Measure (Env × UCB.ArmRewardStream K)","missing":[],"search":"stationarylatentarmstreamprior banditrlproof.thompson.stationarylatentarmstreamprior bayesian prior augmented with the stationary latent arm stream. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure","description":"Canonical trajectory measure in the left-associated sample shape consumed by the Thompson Bayesian-regret decomposition.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-9d84bf857ef9","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2575,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2567"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryLatentArmStreamCanonicalTrajectoryMeasure {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (prior : Measure Env) (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) : Measure ((Env × UCB.ArmRewardStream K) × ((n : Nat) -> Fin K × Real))","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure canonical trajectory measure in the left-associated sample shape consumed by the thompson bayesian-regret decomposition. definition compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_map_prodAssoc_symm","label":"stationaryLatentArmStreamTrajectoryMeasure_map_prodAssoc_symm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_map_prodAssoc_symm","description":"The concentration and decomposition sample shapes agree by product associativity.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-95ad4ed0021a","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2576,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2591"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_map_prodAssoc_symm {Env : Type u} {K : Nat} [MeasurableSpace Env] [NeZero K] (prior : Measure Env) (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) : (stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm).map MeasurableEquiv.prodAssoc.symm = stationaryLatentArmStreamCanonicalTrajectoryMeasure prior rewardKernel algorithm","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_map_prodassoc_symm banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_map_prodassoc_symm the concentration and decomposition sample shapes agree by product associativity. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","description":"Decomposition-facing lower-confidence failure bound on the canonical augmented trajectory measure.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-788687fa6f8d","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2577,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2611"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (arm : Fin K) (n : Nat) : stationaryLatentArmStreamCanonicalTrajectoryMeasure prior rewardKernel algorithm {sample : (Env × UCB.ArmRewardStream K) × ((n : Nat) → Fin K × Real) | 0 < pullCount (environmentTraject…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_realempiricalmean_add_width_le_mean_le_nat_mul_delta banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_realempiricalmean_add_width_le_mean_le_nat_mul_delta decomposition-facing lower-confidence failure bound on the canonical augmented trajectory measure. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","description":"Decomposition-facing upper-confidence failure bound on the canonical augmented trajectory measure.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-711def45b7f5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2578,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2691"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (arm : Fin K) (n : Nat) : stationaryLatentArmStreamCanonicalTrajectoryMeasure prior rewardKernel algorithm {sample : (Env × UCB.ArmRewardStream K) × ((n : Nat) → Fin K × Real) | 0 < pullCount (environmentTraject…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_mean_le_realempiricalmean_sub_width_le_nat_mul_delta banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_mean_le_realempiricalmean_sub_width_le_nat_mul_delta decomposition-facing upper-confidence failure bound on the canonical augmented trajectory measure. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","description":"Decomposition-facing horizon-uniform lower-confidence failure for a measurable environment-dependent arm, with the exact `(n - 1) * delta` cost.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-3a7862d78cf7","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2579,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2771"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (selectedArm : Env -> Fin K) (hselectedArm : Measurable selectedArm) (n : Nat) : stationaryLatentArmStreamCanonicalTrajectoryMeasure prior rewardKernel algorithm {sample : (Env × UCB.ArmRewardStream K) × ((…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_exists_selectedarm_realempiricalmean_add_width_le_mean_le banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_exists_selectedarm_realempiricalmean_add_width_le_mean_le decomposition-facing horizon-uniform lower-confidence failure for a measurable environment-dependent arm, with the exact `(n - 1) * delta` cost. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_mean_bestAction_sub_clippedUCB_le","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_mean_bestAction_sub_clippedUCB_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_mean_bestAction_sub_clippedUCB_le","description":"The first pinned-LML Thompson concentration expectation: the finite-horizon best-action mean minus clipped-UCB sum is controlled by the horizon-uniform lower-confidence failure event.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-33bf066dbee5","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2580,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:2892"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_mean_bestAction_sub_clippedUCB_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (l u : Real) (hlu : l <= u) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : integral (stationaryLatentArmStreamCanonicalTrajectoryM…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_integral_sum_mean_bestaction_sub_clippeducb_le banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_integral_sum_mean_bestaction_sub_clippeducb_le the first pinned-lml thompson concentration expectation: the finite-horizon best-action mean minus clipped-ucb sum is controlled by the horizon-uniform lower-confidence failure event. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_biUnion_clippedCountWidthThreshold_le_armPrefixSum_sub_mul_le","label":"measure_biUnion_clippedCountWidthThreshold_le_armPrefixSum_sub_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_biUnion_clippedCountWidthThreshold_le_armPrefixSum_sub_mul_le","description":"A finite union of upper prefix deviations pays once per positive pull count.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-77c2c7692a06","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2581,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3093"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_biUnion_clippedCountWidthThreshold_le_armPrefixSum_sub_mul_le {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (m : Nat) : UCB.armStreamMeasure nu (⋃ k ∈ Finset.Icc 1 m, {stream : UCB.ArmRewardStream K | clippedCountWidthThreshold sigma2 delta k <= UCB.armPrefixSum arm k stream - (k : Real) * mean}) <= (m : ENNReal) * ENNReal.ofReal delta","missing":[],"search":"measure_biunion_clippedcountwidththreshold_le_armprefixsum_sub_mul_le banditrlproof.thompson.measure_biunion_clippedcountwidththreshold_le_armprefixsum_sub_mul_le a finite union of upper prefix deviations pays once per positive pull count. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_mean_le_realEmpiricalMean_sub_width_le","label":"measure_latentArmStreamTrajectory_exists_mean_le_realEmpiricalMean_sub_width_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_mean_le_realEmpiricalMean_sub_width_le","description":"Finite-horizon upper-confidence failure for one arm on the coupled trajectory. Times with the same positive pull count collapse to one latent-stream prefix event, preserving the exact `(n - 1) * delta` cost.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-abf9a3628536","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2582,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3136"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_exists_mean_le_realEmpiricalMean_sub_width_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | ∃ t < n, 0 < pullCount (latentArmStreamTrajectoryAction sample) arm t ∧ mean <= UCB.realEmpiricalMean (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm t - Real.sqrt (2 * (sigma2 : Real) * Real.log (1 / delta) / (pullCount (latentArmStreamTrajectoryAction sample) arm t : Real))} <= (n - 1 : Nat) * ENNReal.ofRea…","missing":[],"search":"measure_latentarmstreamtrajectory_exists_mean_le_realempiricalmean_sub_width_le banditrlproof.thompson.measure_latentarmstreamtrajectory_exists_mean_le_realempiricalmean_sub_width_le finite-horizon upper-confidence failure for one arm on the coupled trajectory. times with the same positive pull count collapse to one latent-stream prefix event, preserving the exact `(n - 1) * delta` cost. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","label":"measure_latentArmStreamTrajectory_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","description":"Finite-arm union of the horizon upper-confidence failures at fixed environment.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-b03ed8e33ee0","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2583,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3236"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem measure_latentArmStreamTrajectory_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (algorithm : HistoryAlgorithm (Fin K) Real) (env : Env) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (mean : Fin K -> Real) (sigma2 : NNReal) (hsubG : forall arm, HasSubgaussianMGF (fun reward => reward - mean arm) sigma2 (nu arm)) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : latentArmStreamTrajectoryMeasure algorithm env nu {sample | ∃ arm : Fin K, ∃ t < n, 0 < pullCount (latentArmStreamTrajectoryAction sample) arm t ∧ mean arm <= UCB.realEmpiricalMean (latentArmStreamTrajectoryAction sample) (latentArmStreamTrajectoryReward sample) arm t - Real.sqrt (2 * (sigma2 : Real) * Real.log (1 / delta) / (pullCount (latentArmStreamTrajectoryAction sample) arm t :…","missing":[],"search":"measure_latentarmstreamtrajectory_exists_arm_exists_mean_le_realempiricalmean_sub_width_le banditrlproof.thompson.measure_latentarmstreamtrajectory_exists_arm_exists_mean_le_realempiricalmean_sub_width_le finite-arm union of the horizon upper-confidence failures at fixed environment. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","label":"stationaryLatentArmStreamTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","description":"Prior mixing preserves the finite-arm horizon upper-confidence budget.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-5b348b032fd2","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2584,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3299"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : stationaryLatentArmStreamTrajectoryMeasure prior rewardKernel algorithm {sample | ∃ arm : Fin K, ∃ t < n, 0 < pullCount (stationaryLatentArmStreamTrajectoryAction sample) arm t ∧ mean (stationaryLatentArm…","missing":[],"search":"stationarylatentarmstreamtrajectorymeasure_exists_arm_exists_mean_le_realempiricalmean_sub_width_le banditrlproof.thompson.stationarylatentarmstreamtrajectorymeasure_exists_arm_exists_mean_le_realempiricalmean_sub_width_le prior mixing preserves the finite-arm horizon upper-confidence budget. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","description":"Decomposition-facing finite-arm horizon upper-confidence failure with exact `K * (n - 1) * delta` cost.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-fbca9b0f1e7e","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2585,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3406"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : stationaryLatentArmStreamCanonicalTrajectoryMeasure prior rewardKernel algorithm {sample | ∃ arm : Fin K, ∃ t < n, 0 < pullCount (environmentTrajectoryAction sample) arm t ∧ mean sample.1.1 arm <…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_exists_arm_exists_mean_le_realempiricalmean_sub_width_le banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_exists_arm_exists_mean_le_realempiricalmean_sub_width_le decomposition-facing finite-arm horizon upper-confidence failure with exact `k * (n - 1) * delta` cost. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_clippedUCB_action_sub_mean_le","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_clippedUCB_action_sub_mean_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_clippedUCB_action_sub_mean_le","description":"The second pinned-LML Thompson concentration expectation: selected-action clipped-UCB excess is bounded by deterministic pull-count summation plus the finite-arm horizon upper-confidence failure budget.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-0f2074e7fd9a","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2586,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3512"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_clippedUCB_action_sub_mean_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (algorithm : HistoryAlgorithm (Fin K) Real) (mean : Env -> Fin K -> Real) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (l u : Real) (hlu : l <= u) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : integral (stationaryLatentArmStreamCanonicalTrajectoryMeasure prior rewardKernel algorithm) (fun sample => ∑ t ∈ Finset.range…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_integral_sum_clippeducb_action_sub_mean_le banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_integral_sum_clippeducb_action_sub_mean_le the second pinned-lml thompson concentration expectation: selected-action clipped-ucb excess is bounded by deterministic pull-count summation plus the finite-arm horizon upper-confidence failure budget. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le_of_delta","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le_of_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le_of_delta","description":"Stationary-reward Thompson Bayesian regret with an explicit confidence parameter. This is the decomposition-facing join of the two pinned-LML clipped-UCB expectation bounds. The analytic upper bound is in fact comparator-uniform, but the public Bayesian-regret endpoint deliberately retains `IsOptimalMeanSelector mean bestAction` as an interpretation contract.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-ded655f64179","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2587,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3723"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le_of_delta {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) (mean : Env -> Fin K -> Real) (hbest : IsOptimalMeanSelector mean bestAction) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (l u : Real) (hlu : l <= u) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : let augmentedPrior := stationaryLate…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_integral_trajectorybayesmeanregret_le_of_delta banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_integral_trajectorybayesmeanregret_le_of_delta stationary-reward thompson bayesian regret with an explicit confidence parameter. this is the decomposition-facing join of the two pinned-lml clipped-ucb expectation bounds. the analytic upper bound is in fact comparator-uniform, but the public bayesian-regret endpoint deliberately retains `isoptimalmeanselector mean bestaction` as an interpretation contract. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","label":"stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","description":"Pinned-LML stationary-reward Thompson Bayesian-regret bound, obtained from the explicit-confidence theorem with `delta = 1 / n ^ 2`. Unlike the comparator-relative decomposition, this endpoint explicitly requires the selector to maximize the declared mean surface pointwise.","url":"../modules/banditrlproof-algorithms-thompsonstationaryreward/index.html#decl-04335aae6150","parent":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","order":2588,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.ThompsonStationaryReward"],["Source","BanditRLProof/Algorithms/ThompsonStationaryReward.lean:3848"],["Chapter","Thompson sampling"],["Used in books","bandit"],["Reading references","teaching:thompson"],["Indexed settings","None registered"]],"statement":"theorem stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le {Env : Type u} {K : Nat} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [NeZero K] (prior : Measure Env) [IsProbabilityMeasure prior] (rewardKernel : Kernel (Env × Fin K) Real) [IsMarkovKernel rewardKernel] (bestAction : Env -> Fin K) (hbestAction : Measurable bestAction) (mean : Env -> Fin K -> Real) (hbest : IsOptimalMeanSelector mean bestAction) (hmeas_mean : Measurable (fun input : Env × Fin K => mean input.1 input.2)) (l u : Real) (hlu : l <= u) (hmeanMem : forall env arm, mean env arm ∈ Set.Icc l u) (sigma2 : NNReal) (hsubG : forall env arm, HasSubgaussianMGF (fun reward => reward - mean env arm) sigma2 (rewardKernel (env, arm))) (hsigma2 : sigma2 ≠ 0) (n : Nat) : let augmentedPrior := stationaryLatentArmStreamPrior prior rewardKernel let feedbackEnvironment := latentAr…","missing":[],"search":"stationarylatentarmstreamcanonicaltrajectorymeasure_integral_trajectorybayesmeanregret_le banditrlproof.thompson.stationarylatentarmstreamcanonicaltrajectorymeasure_integral_trajectorybayesmeanregret_le pinned-lml stationary-reward thompson bayesian-regret bound, obtained from the explicit-confidence theorem with `delta = 1 / n ^ 2`. unlike the comparator-relative decomposition, this endpoint explicitly requires the selector to maximize the declared mean surface pointwise. theorem compiled","shard":"modules/2357638fc93777eb.json","books":["bandit"],"chapters":["teaching:thompson"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.Spec","label":"Spec","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UCB.Spec","description":"Parameters for a finite-arm UCB run.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-480f7e39a0c1","parent":"module:BanditRLProof.Algorithms.UCB","order":2589,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:28"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"structure Spec (K : Nat) where","missing":[],"search":"spec banditrlproof.ucb.spec parameters for a finite-arm ucb run. structure compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.IndexState","label":"IndexState","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UCB.IndexState","description":"State visible to an index policy at one time step.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-71c3b8e4ee19","parent":"module:BanditRLProof.Algorithms.UCB","order":2590,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:33"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"structure IndexState (K : Nat) where","missing":[],"search":"indexstate banditrlproof.ucb.indexstate state visible to an index policy at one time step. structure compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.score","label":"score","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.score","description":"Placeholder score surface for the first dependency-light layer. The Mathlib/LML migration will refine this to `mean + sqrt (2 * c * log n / pulls)`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-dabeef1f8e2f","parent":"module:BanditRLProof.Algorithms.UCB","order":2591,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:43"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def score (_spec : Spec K) (state : IndexState K) (arm : Fin K) : Rat","missing":[],"search":"score banditrlproof.ucb.score placeholder score surface for the first dependency-light layer. the mathlib/lml migration will refine this to `mean + sqrt (2 * c * log n / pulls)`. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.score_eq_empiricalMean","label":"score_eq_empiricalMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.score_eq_empiricalMean","description":"@[simp] theorem score_eq_empiricalMean (spec : Spec K) (state : IndexState K) (arm : Fin K) : score spec state arm = state.empiricalMean arm","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-720ff520b5e7","parent":"module:BanditRLProof.Algorithms.UCB","order":2592,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:46"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem score_eq_empiricalMean (spec : Spec K) (state : IndexState K) (arm : Fin K) : score spec state arm = state.empiricalMean arm","missing":[],"search":"score_eq_empiricalmean banditrlproof.ucb.score_eq_empiricalmean @[simp] theorem score_eq_empiricalmean (spec : spec k) (state : indexstate k) (arm : fin k) : score spec state arm = state.empiricalmean arm theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScore","label":"confidenceScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScore","description":"Real-valued UCB confidence score `empirical mean + radius`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-f91da85455d4","parent":"module:BanditRLProof.Algorithms.UCB","order":2593,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:51"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def confidenceScore {Arm : Type} (empiricalMean radius : Arm -> Real) (arm : Arm) : Real","missing":[],"search":"confidencescore banditrlproof.ucb.confidencescore real-valued ucb confidence score `empirical mean + radius`. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScore_apply","label":"confidenceScore_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScore_apply","description":"@[simp] theorem confidenceScore_apply {Arm : Type} (empiricalMean radius : Arm -> Real) (arm : Arm) : confidenceScore empiricalMean radius arm = empiricalMean arm + radius arm","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-003b5216561e","parent":"module:BanditRLProof.Algorithms.UCB","order":2594,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:55"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem confidenceScore_apply {Arm : Type} (empiricalMean radius : Arm -> Real) (arm : Arm) : confidenceScore empiricalMean radius arm = empiricalMean arm + radius arm","missing":[],"search":"confidencescore_apply banditrlproof.ucb.confidencescore_apply @[simp] theorem confidencescore_apply {arm : type} (empiricalmean radius : arm -> real) (arm : arm) : confidencescore empiricalmean radius arm = empiricalmean arm + radius arm theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.score_le_foldl_select","label":"score_le_foldl_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.score_le_foldl_select","description":"private theorem score_le_foldl_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha cases…","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-e6c02099ee85","parent":"module:BanditRLProof.Algorithms.UCB","order":2595,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:60"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"private theorem score_le_foldl_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha cases ha) (by simp) | arm :: rest => by let select := fun best arm : Fin K => if scores best < scores arm then arm else best let next := select init arm have ih","missing":[],"search":"score_le_foldl_select banditrlproof.ucb.score_le_foldl_select private theorem score_le_foldl_select {k : nat} (scores : fin k -> real) (init : fin k) : forall l : list (fin k), (forall a : fin k, list.mem a l -> scores a <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init) | [] => by exact and.intro (by intro _ ha cases ha) (by simp) | arm :: rest => by let select := fun best arm : fin k => if scores best < scores arm then arm else best let next := select init arm have ih theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.scoreArgmax","label":"scoreArgmax","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.scoreArgmax","description":"Concrete finite-arm Real score argmax selector. The selector scans `List.finRange K`, keeps the previous arm on ties, and uses `hK` only to seed the nonempty finite arm set. This mirrors the ETC argmax oracle but targets the Real-valued UCB confidence-score surface.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a17b7e7aabaf","parent":"module:BanditRLProof.Algorithms.UCB","order":2596,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:121"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def scoreArgmax {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Fin K","missing":[],"search":"scoreargmax banditrlproof.ucb.scoreargmax concrete finite-arm real score argmax selector. the selector scans `list.finrange k`, keeps the previous arm on ties, and uses `hk` only to seed the nonempty finite arm set. this mirrors the etc argmax oracle but targets the real-valued ucb confidence-score surface. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.scoreArgmax_spec","label":"scoreArgmax_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.scoreArgmax_spec","description":"The concrete Real score argmax dominates every arm score.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-90cbc46c18f6","parent":"module:BanditRLProof.Algorithms.UCB","order":2597,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:129"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem scoreArgmax_spec {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) : scores a <= scores (scoreArgmax hK scores)","missing":[],"search":"scoreargmax_spec banditrlproof.ucb.scoreargmax_spec the concrete real score argmax dominates every arm score. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxAction","label":"confidenceScoreArgmaxAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScoreArgmaxAction","description":"Concrete UCB action that maximizes the current confidence score.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-041c25d5a8a9","parent":"module:BanditRLProof.Algorithms.UCB","order":2598,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:141"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceScoreArgmaxAction {Omega : Type} {K : Nat} (hK : 0 < K) (empiricalMean : Omega -> Nat -> Fin K -> Real) (radius : Nat -> Fin K -> Real) : Omega -> Nat -> Fin K","missing":[],"search":"confidencescoreargmaxaction banditrlproof.ucb.confidencescoreargmaxaction concrete ucb action that maximizes the current confidence score. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxAction_score_max","label":"confidenceScoreArgmaxAction_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScoreArgmaxAction_score_max","description":"The concrete confidence-score argmax action supplies score maximality against any comparison arm.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-77659d8a9546","parent":"module:BanditRLProof.Algorithms.UCB","order":2599,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:154"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem confidenceScoreArgmaxAction_score_max {Omega : Type} {K : Nat} (hK : 0 < K) (empiricalMean : Omega -> Nat -> Fin K -> Real) (radius : Nat -> Fin K -> Real) (omega : Omega) (t : Nat) (arm : Fin K) : confidenceScore (empiricalMean omega t) (radius t) arm <= confidenceScore (empiricalMean omega t) (radius t) (confidenceScoreArgmaxAction hK empiricalMean radius omega t)","missing":[],"search":"confidencescoreargmaxaction_score_max banditrlproof.ucb.confidencescoreargmaxaction_score_max the concrete confidence-score argmax action supplies score maximality against any comparison arm. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxAction_score_max_of_selected","label":"confidenceScoreArgmaxAction_score_max_of_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScoreArgmaxAction_score_max_of_selected","description":"Selected-arm form of `confidenceScoreArgmaxAction_score_max`, matching the abstract selected-action bridge contract.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-4d65af31d7fb","parent":"module:BanditRLProof.Algorithms.UCB","order":2600,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:172"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem confidenceScoreArgmaxAction_score_max_of_selected {Omega : Type} {K : Nat} (hK : 0 < K) (empiricalMean : Omega -> Nat -> Fin K -> Real) (radius : Nat -> Fin K -> Real) (omega : Omega) (t : Nat) (best chosen : Fin K) (hselected : confidenceScoreArgmaxAction hK empiricalMean radius omega t = chosen) : confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen","missing":[],"search":"confidencescoreargmaxaction_score_max_of_selected banditrlproof.ucb.confidencescoreargmaxaction_score_max_of_selected selected-arm form of `confidencescoreargmaxaction_score_max`, matching the abstract selected-action bridge contract. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.meanGap","label":"meanGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.meanGap","description":"Mean gap against a designated best arm for Real-valued UCB algebra.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-2e724ee6a6a9","parent":"module:BanditRLProof.Algorithms.UCB","order":2601,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:187"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def meanGap {Arm : Type} (trueMean : Arm -> Real) (best arm : Arm) : Real","missing":[],"search":"meangap banditrlproof.ucb.meangap mean gap against a designated best arm for real-valued ucb algebra. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.meanGap_apply","label":"meanGap_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.meanGap_apply","description":"@[simp] theorem meanGap_apply {Arm : Type} (trueMean : Arm -> Real) (best arm : Arm) : meanGap trueMean best arm = trueMean best - trueMean arm","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a6db585c49aa","parent":"module:BanditRLProof.Algorithms.UCB","order":2602,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:190"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem meanGap_apply {Arm : Type} (trueMean : Arm -> Real) (best arm : Arm) : meanGap trueMean best arm = trueMean best - trueMean arm","missing":[],"search":"meangap_apply banditrlproof.ucb.meangap_apply @[simp] theorem meangap_apply {arm : type} (truemean : arm -> real) (best arm : arm) : meangap truemean best arm = truemean best - truemean arm theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_confidenceScore_max","label":"meanGap_le_two_radius_of_confidenceScore_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.meanGap_le_two_radius_of_confidenceScore_max","description":"UCB good-event algebra: if the best arm's true mean is below its upper confidence score, the chosen arm's true mean is above its lower confidence score, and the chosen arm maximizes the confidence score against the best arm, then the chosen arm's mean gap is at most twice its radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-ab3e02b66556","parent":"module:BanditRLProof.Algorithms.UCB","order":2603,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:200"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem meanGap_le_two_radius_of_confidenceScore_max {Arm : Type} (trueMean empiricalMean radius : Arm -> Real) (best chosen : Arm) (hbest_upper : trueMean best <= confidenceScore empiricalMean radius best) (hchosen_lower : empiricalMean chosen - radius chosen <= trueMean chosen) (hscore : confidenceScore empiricalMean radius best <= confidenceScore empiricalMean radius chosen) : meanGap trueMean best chosen <= 2 * radius chosen","missing":[],"search":"meangap_le_two_radius_of_confidencescore_max banditrlproof.ucb.meangap_le_two_radius_of_confidencescore_max ucb good-event algebra: if the best arm's true mean is below its upper confidence score, the chosen arm's true mean is above its lower confidence score, and the chosen arm maximizes the confidence score against the best arm, then the chosen arm's mean gap is at most twice its radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.not_two_radius_lt_meanGap_of_confidenceScore_max","label":"not_two_radius_lt_meanGap_of_confidenceScore_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.not_two_radius_lt_meanGap_of_confidenceScore_max","description":"Contrapositive consumer for the UCB good-event algebra: under the same good event and score-maximality hypotheses, a strict `2 * radius < gap` certificate rules out choosing that arm.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-152d037aac45","parent":"module:BanditRLProof.Algorithms.UCB","order":2604,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:219"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem not_two_radius_lt_meanGap_of_confidenceScore_max {Arm : Type} (trueMean empiricalMean radius : Arm -> Real) (best chosen : Arm) (hbest_upper : trueMean best <= confidenceScore empiricalMean radius best) (hchosen_lower : empiricalMean chosen - radius chosen <= trueMean chosen) (hscore : confidenceScore empiricalMean radius best <= confidenceScore empiricalMean radius chosen) (hgap_large : 2 * radius chosen < meanGap trueMean best chosen) : False","missing":[],"search":"not_two_radius_lt_meangap_of_confidencescore_max banditrlproof.ucb.not_two_radius_lt_meangap_of_confidencescore_max contrapositive consumer for the ucb good-event algebra: under the same good event and score-maximality hypotheses, a strict `2 * radius < gap` certificate rules out choosing that arm. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.upperConfidenceBad","label":"upperConfidenceBad","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.upperConfidenceBad","description":"Upper-confidence failure for a random empirical-mean surface.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-30f267ce9455","parent":"module:BanditRLProof.Algorithms.UCB","order":2605,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:236"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def upperConfidenceBad {Omega Arm : Type} (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) : Set Omega","missing":[],"search":"upperconfidencebad banditrlproof.ucb.upperconfidencebad upper-confidence failure for a random empirical-mean surface. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lowerConfidenceBad","label":"lowerConfidenceBad","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.lowerConfidenceBad","description":"Lower-confidence failure for a random empirical-mean surface.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-1d9d2d6e3686","parent":"module:BanditRLProof.Algorithms.UCB","order":2606,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:242"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def lowerConfidenceBad {Omega Arm : Type} (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) : Set Omega","missing":[],"search":"lowerconfidencebad banditrlproof.ucb.lowerconfidencebad lower-confidence failure for a random empirical-mean surface. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_upperConfidenceBad","label":"measurableSet_upperConfidenceBad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_upperConfidenceBad","description":"Upper-confidence failure is measurable when the arm empirical mean is measurable.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-877ecbdc47d8","parent":"module:BanditRLProof.Algorithms.UCB","order":2607,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:248"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_upperConfidenceBad {Omega Arm : Type} [MeasurableSpace Omega] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) (hmeas : Measurable (fun omega : Omega => empiricalMean omega arm)) : MeasurableSet (upperConfidenceBad trueMean empiricalMean radius arm)","missing":[],"search":"measurableset_upperconfidencebad banditrlproof.ucb.measurableset_upperconfidencebad upper-confidence failure is measurable when the arm empirical mean is measurable. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_lowerConfidenceBad","label":"measurableSet_lowerConfidenceBad","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_lowerConfidenceBad","description":"Lower-confidence failure is measurable when the arm empirical mean is measurable.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-cbe522f1186e","parent":"module:BanditRLProof.Algorithms.UCB","order":2608,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:258"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_lowerConfidenceBad {Omega Arm : Type} [MeasurableSpace Omega] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) (hmeas : Measurable (fun omega : Omega => empiricalMean omega arm)) : MeasurableSet (lowerConfidenceBad trueMean empiricalMean radius arm)","missing":[],"search":"measurableset_lowerconfidencebad banditrlproof.ucb.measurableset_lowerconfidencebad lower-confidence failure is measurable when the arm empirical mean is measurable. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceBadEvent","label":"confidenceBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceBadEvent","description":"The finite-arm UCB confidence bad event, as a union of upper/lower failures.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d5186437f1c4","parent":"module:BanditRLProof.Algorithms.UCB","order":2609,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:268"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def confidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) : Set Omega","missing":[],"search":"confidencebadevent banditrlproof.ucb.confidencebadevent the finite-arm ucb confidence bad event, as a union of upper/lower failures. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_confidenceBadEvent","label":"measurableSet_confidenceBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_confidenceBadEvent","description":"The finite-arm UCB confidence bad event is measurable from per-arm empirical-mean measurability.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-2c4e200c076b","parent":"module:BanditRLProof.Algorithms.UCB","order":2610,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:275"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_confidenceBadEvent {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (hmeas : forall arm : Arm, Measurable (fun omega : Omega => empiricalMean omega arm)) : MeasurableSet (confidenceBadEvent trueMean empiricalMean radius)","missing":[],"search":"measurableset_confidencebadevent banditrlproof.ucb.measurableset_confidencebadevent the finite-arm ucb confidence bad event is measurable from per-arm empirical-mean measurability. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.not_upperConfidenceBad_of_not_confidenceBadEvent","label":"not_upperConfidenceBad_of_not_confidenceBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.not_upperConfidenceBad_of_not_confidenceBadEvent","description":"theorem not_upperConfidenceBad_of_not_confidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (omega : Omega) (arm : Arm) (hgood : omega ∉ confidenceBadEvent trueMean empiricalMean radius) : omega ∉ upperConfidenceBad trueMean empiricalMean radius arm","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-6d236a32d0de","parent":"module:BanditRLProof.Algorithms.UCB","order":2611,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:289"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem not_upperConfidenceBad_of_not_confidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (omega : Omega) (arm : Arm) (hgood : omega ∉ confidenceBadEvent trueMean empiricalMean radius) : omega ∉ upperConfidenceBad trueMean empiricalMean radius arm","missing":[],"search":"not_upperconfidencebad_of_not_confidencebadevent banditrlproof.ucb.not_upperconfidencebad_of_not_confidencebadevent theorem not_upperconfidencebad_of_not_confidencebadevent {omega arm : type} [fintype arm] (truemean : arm -> real) (empiricalmean : omega -> arm -> real) (radius : arm -> real) (omega : omega) (arm : arm) (hgood : omega ∉ confidencebadevent truemean empiricalmean radius) : omega ∉ upperconfidencebad truemean empiricalmean radius arm theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.not_lowerConfidenceBad_of_not_confidenceBadEvent","label":"not_lowerConfidenceBad_of_not_confidenceBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.not_lowerConfidenceBad_of_not_confidenceBadEvent","description":"theorem not_lowerConfidenceBad_of_not_confidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (omega : Omega) (arm : Arm) (hgood : omega ∉ confidenceBadEvent trueMean empiricalMean radius) : omega ∉ lowerConfidenceBad trueMean empiricalMean radius arm","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-c9ee6a1860c2","parent":"module:BanditRLProof.Algorithms.UCB","order":2612,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:300"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem not_lowerConfidenceBad_of_not_confidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (omega : Omega) (arm : Arm) (hgood : omega ∉ confidenceBadEvent trueMean empiricalMean radius) : omega ∉ lowerConfidenceBad trueMean empiricalMean radius arm","missing":[],"search":"not_lowerconfidencebad_of_not_confidencebadevent banditrlproof.ucb.not_lowerconfidencebad_of_not_confidencebadevent theorem not_lowerconfidencebad_of_not_confidencebadevent {omega arm : type} [fintype arm] (truemean : arm -> real) (empiricalmean : omega -> arm -> real) (radius : arm -> real) (omega : omega) (arm : arm) (hgood : omega ∉ confidencebadevent truemean empiricalmean radius) : omega ∉ lowerconfidencebad truemean empiricalmean radius arm theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_not_confidenceBadEvent","label":"meanGap_le_two_radius_of_not_confidenceBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.meanGap_le_two_radius_of_not_confidenceBadEvent","description":"Event-level UCB good-event consumer: outside the finite-arm confidence bad event, score maximality against the best arm implies the standard `gap <= 2 * chosenRadius` conclusion.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-28a90821bfbf","parent":"module:BanditRLProof.Algorithms.UCB","order":2613,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:316"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem meanGap_le_two_radius_of_not_confidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (omega : Omega) (best chosen : Arm) (hgood : omega ∉ confidenceBadEvent trueMean empiricalMean radius) (hscore : confidenceScore (empiricalMean omega) radius best <= confidenceScore (empiricalMean omega) radius chosen) : meanGap trueMean best chosen <= 2 * radius chosen","missing":[],"search":"meangap_le_two_radius_of_not_confidencebadevent banditrlproof.ucb.meangap_le_two_radius_of_not_confidencebadevent event-level ucb good-event consumer: outside the finite-arm confidence bad event, score maximality against the best arm implies the standard `gap <= 2 * chosenradius` conclusion. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_confidenceBadEvent_le_sum_upper_lower","label":"measure_confidenceBadEvent_le_sum_upper_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_confidenceBadEvent_le_sum_upper_lower","description":"Finite-arm union bound for the UCB confidence bad event. This is still an outer-measure bound: it does not require measurability of the upper/lower confidence failure events.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-478e74265442","parent":"module:BanditRLProof.Algorithms.UCB","order":2614,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:347"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_confidenceBadEvent_le_sum_upper_lower {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) : mu (confidenceBadEvent trueMean empiricalMean radius) <= (Finset.univ : Finset Arm).sum (fun arm => mu (upperConfidenceBad trueMean empiricalMean radius arm) + mu (lowerConfidenceBad trueMean empiricalMean radius arm))","missing":[],"search":"measure_confidencebadevent_le_sum_upper_lower banditrlproof.ucb.measure_confidencebadevent_le_sum_upper_lower finite-arm union bound for the ucb confidence bad event. this is still an outer-measure bound: it does not require measurability of the upper/lower confidence failure events. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceBadEventAt","label":"confidenceBadEventAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceBadEventAt","description":"Time-indexed UCB confidence bad event.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-c14b54e42f34","parent":"module:BanditRLProof.Algorithms.UCB","order":2615,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:379"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def confidenceBadEventAt {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (t : Nat) : Set Omega","missing":[],"search":"confidencebadeventat banditrlproof.ucb.confidencebadeventat time-indexed ucb confidence bad event. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_confidenceBadEventAt","label":"measurableSet_confidenceBadEventAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_confidenceBadEventAt","description":"The time-indexed UCB confidence bad event is measurable from per-arm empirical mean measurability at that time.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-6476b5c01e9e","parent":"module:BanditRLProof.Algorithms.UCB","order":2616,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:390"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_confidenceBadEventAt {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (t : Nat) (hmeas : forall arm : Arm, Measurable (fun omega : Omega => empiricalMean omega t arm)) : MeasurableSet (confidenceBadEventAt trueMean empiricalMean radius t)","missing":[],"search":"measurableset_confidencebadeventat banditrlproof.ucb.measurableset_confidencebadeventat the time-indexed ucb confidence bad event is measurable from per-arm empirical mean measurability at that time. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteHorizonConfidenceBadEvent","label":"finiteHorizonConfidenceBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.finiteHorizonConfidenceBadEvent","description":"Finite-horizon union of time-indexed UCB confidence bad events.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a28cfa70dcdc","parent":"module:BanditRLProof.Algorithms.UCB","order":2617,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:403"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def finiteHorizonConfidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) : Set Omega","missing":[],"search":"finitehorizonconfidencebadevent banditrlproof.ucb.finitehorizonconfidencebadevent finite-horizon union of time-indexed ucb confidence bad events. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.not_confidenceBadEventAt_of_not_finiteHorizonConfidenceBadEvent","label":"not_confidenceBadEventAt_of_not_finiteHorizonConfidenceBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.not_confidenceBadEventAt_of_not_finiteHorizonConfidenceBadEvent","description":"Outside the finite-horizon confidence bad event, every time-indexed bad event inside the horizon is absent.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-00283b2502c1","parent":"module:BanditRLProof.Algorithms.UCB","order":2618,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:414"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem not_confidenceBadEventAt_of_not_finiteHorizonConfidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) (omega : Omega) (t : Nat) (ht : t < T) (hgood : omega ∉ finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) : omega ∉ confidenceBadEventAt trueMean empiricalMean radius t","missing":[],"search":"not_confidencebadeventat_of_not_finitehorizonconfidencebadevent banditrlproof.ucb.not_confidencebadeventat_of_not_finitehorizonconfidencebadevent outside the finite-horizon confidence bad event, every time-indexed bad event inside the horizon is absent. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_not_finiteHorizonConfidenceBadEvent","label":"meanGap_le_two_radius_of_not_finiteHorizonConfidenceBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.meanGap_le_two_radius_of_not_finiteHorizonConfidenceBadEvent","description":"Finite-horizon good-event consumer: outside the finite-horizon confidence bad event, score maximality at any `t < T` gives the standard UCB gap-radius bound for the chosen arm.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-5a17cce1bdd6","parent":"module:BanditRLProof.Algorithms.UCB","order":2619,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:435"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem meanGap_le_two_radius_of_not_finiteHorizonConfidenceBadEvent {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) (omega : Omega) (t : Nat) (best chosen : Arm) (ht : t < T) (hgood : omega ∉ finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) (hscore : confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen) : meanGap trueMean best chosen <= 2 * radius t chosen","missing":[],"search":"meangap_le_two_radius_of_not_finitehorizonconfidencebadevent banditrlproof.ucb.meangap_le_two_radius_of_not_finitehorizonconfidencebadevent finite-horizon good-event consumer: outside the finite-horizon confidence bad event, score maximality at any `t < t` gives the standard ucb gap-radius bound for the chosen arm. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap_of_score_max","label":"mem_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap_of_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap_of_score_max","description":"Contrapositive finite-horizon good-event consumer: if a chosen arm's gap is larger than twice its current radius and it beats the best arm's UCB score, the finite-horizon confidence bad event must occur.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-170dabdcc759","parent":"module:BanditRLProof.Algorithms.UCB","order":2620,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:462"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem mem_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap_of_score_max {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) (omega : Omega) (t : Nat) (best chosen : Arm) (ht : t < T) (hscore : confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen) (hgap_large : 2 * radius t chosen < meanGap trueMean best chosen) : omega ∈ finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T","missing":[],"search":"mem_finitehorizonconfidencebadevent_of_two_radius_lt_meangap_of_score_max banditrlproof.ucb.mem_finitehorizonconfidencebadevent_of_two_radius_lt_meangap_of_score_max contrapositive finite-horizon good-event consumer: if a chosen arm's gap is larger than twice its current radius and it beats the best arm's ucb score, the finite-horizon confidence bad event must occur. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.scoreMaxEvent_subset_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap","label":"scoreMaxEvent_subset_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.scoreMaxEvent_subset_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap","description":"Event-level form of the finite-horizon large-gap consumer. This is the set inclusion shape needed before applying measure monotonicity in pull-count arguments.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-2a0414a8c62a","parent":"module:BanditRLProof.Algorithms.UCB","order":2621,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:487"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem scoreMaxEvent_subset_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) (t : Nat) (best chosen : Arm) (ht : t < T) (hgap_large : 2 * radius t chosen < meanGap trueMean best chosen) : {omega : Omega | confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen} ⊆ finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T","missing":[],"search":"scoremaxevent_subset_finitehorizonconfidencebadevent_of_two_radius_lt_meangap banditrlproof.ucb.scoremaxevent_subset_finitehorizonconfidencebadevent_of_two_radius_lt_meangap event-level form of the finite-horizon large-gap consumer. this is the set inclusion shape needed before applying measure monotonicity in pull-count arguments. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_sum_upper_lower","label":"measure_finiteHorizonConfidenceBadEvent_le_sum_upper_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_sum_upper_lower","description":"Finite-horizon union bound for UCB confidence bad events. This assembles the single-time upper/lower confidence-event union bound across `t < T`. It does not produce concentration tails or simplify the resulting double finite sum.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-821482b76634","parent":"module:BanditRLProof.Algorithms.UCB","order":2622,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:512"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_sum_upper_lower {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => mu (upperConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) + mu (lowerConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm)))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_sum_upper_lower banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_sum_upper_lower finite-horizon union bound for ucb confidence bad events. this assembles the single-time upper/lower confidence-event union bound across `t < t`. it does not produce concentration tails or simplify the resulting double finite sum. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_tail_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_tail_sum","description":"Finite-horizon tail-bound consumer for UCB confidence bad events. The hypotheses `hupper` and `hlower` are the per-time/per-arm concentration inputs for the upper and lower confidence failures. This wrapper only assembles those local tail budgets across `t < T` and finite arms.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-3fb74858d518","parent":"module:BanditRLProof.Algorithms.UCB","order":2623,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:565"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_tail_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (upperTail lowerTail : Nat -> Arm -> ENNReal) (T : Nat) (hupper : forall t arm, t < T -> mu (upperConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) <= upperTail t arm) (hlower : forall t arm, t < T -> mu (lowerConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) <= lowerTail t arm) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => upperTail t arm + lowerTail t arm))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_tail_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_tail_sum finite-horizon tail-bound consumer for ucb confidence bad events. the hypotheses `hupper` and `hlower` are the per-time/per-arm concentration inputs for the upper and lower confidence failures. this wrapper only assembles those local tail budgets across `t < t` and finite arms. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.upperConfidenceBad_subset_absDeviation","label":"upperConfidenceBad_subset_absDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.upperConfidenceBad_subset_absDeviation","description":"An upper-confidence failure implies an absolute empirical-mean deviation at least as large as the radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a9c1609a780b","parent":"module:BanditRLProof.Algorithms.UCB","order":2624,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:598"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem upperConfidenceBad_subset_absDeviation {Omega Arm : Type} (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) : upperConfidenceBad trueMean empiricalMean radius arm ⊆ {omega | radius arm <= |empiricalMean omega arm - trueMean arm|}","missing":[],"search":"upperconfidencebad_subset_absdeviation banditrlproof.ucb.upperconfidencebad_subset_absdeviation an upper-confidence failure implies an absolute empirical-mean deviation at least as large as the radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lowerConfidenceBad_subset_absDeviation","label":"lowerConfidenceBad_subset_absDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lowerConfidenceBad_subset_absDeviation","description":"A lower-confidence failure implies an absolute empirical-mean deviation at least as large as the radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-652104ec93ee","parent":"module:BanditRLProof.Algorithms.UCB","order":2625,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:620"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lowerConfidenceBad_subset_absDeviation {Omega Arm : Type} (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) : lowerConfidenceBad trueMean empiricalMean radius arm ⊆ {omega | radius arm <= |empiricalMean omega arm - trueMean arm|}","missing":[],"search":"lowerconfidencebad_subset_absdeviation banditrlproof.ucb.lowerconfidencebad_subset_absdeviation a lower-confidence failure implies an absolute empirical-mean deviation at least as large as the radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_absDeviation","label":"measure_upperConfidenceBad_le_absDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_upperConfidenceBad_le_absDeviation","description":"Measure monotonicity form of `upperConfidenceBad_subset_absDeviation`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-920a2f633d95","parent":"module:BanditRLProof.Algorithms.UCB","order":2626,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:634"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_upperConfidenceBad_le_absDeviation {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) : mu (upperConfidenceBad trueMean empiricalMean radius arm) <= mu {omega | radius arm <= |empiricalMean omega arm - trueMean arm|}","missing":[],"search":"measure_upperconfidencebad_le_absdeviation banditrlproof.ucb.measure_upperconfidencebad_le_absdeviation measure monotonicity form of `upperconfidencebad_subset_absdeviation`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_absDeviation","label":"measure_lowerConfidenceBad_le_absDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_lowerConfidenceBad_le_absDeviation","description":"Measure monotonicity form of `lowerConfidenceBad_subset_absDeviation`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-5d055da7513b","parent":"module:BanditRLProof.Algorithms.UCB","order":2627,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:646"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_lowerConfidenceBad_le_absDeviation {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) (trueMean : Arm -> Real) (empiricalMean : Omega -> Arm -> Real) (radius : Arm -> Real) (arm : Arm) : mu (lowerConfidenceBad trueMean empiricalMean radius arm) <= mu {omega | radius arm <= |empiricalMean omega arm - trueMean arm|}","missing":[],"search":"measure_lowerconfidencebad_le_absdeviation banditrlproof.ucb.measure_lowerconfidencebad_le_absdeviation measure monotonicity form of `lowerconfidencebad_subset_absdeviation`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_absDeviation_tail_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_absDeviation_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_absDeviation_tail_sum","description":"Finite-horizon UCB confidence bad-event bound from absolute-deviation tails. This is the UCB-facing adapter for concentration inequalities that bound `mu {omega | radius <= |empiricalMean - trueMean|}`. The same absolute deviation tail controls both the upper and lower confidence failures.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-322732f57e4f","parent":"module:BanditRLProof.Algorithms.UCB","order":2628,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:664"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_absDeviation_tail_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (tail : Nat -> Arm -> ENNReal) (T : Nat) (htail : forall t arm, t < T -> mu {omega | radius t arm <= |empiricalMean omega t arm - trueMean arm|} <= tail t arm) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => tail t arm + tail t arm))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_absdeviation_tail_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_absdeviation_tail_sum finite-horizon ucb confidence bad-event bound from absolute-deviation tails. this is the ucb-facing adapter for concentration inequalities that bound `mu {omega | radius <= |empiricalmean - truemean|}`. the same absolute deviation tail controls both the upper and lower confidence failures. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.chebyshevAbsDeviationTail","label":"chebyshevAbsDeviationTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.chebyshevAbsDeviationTail","description":"Chebyshev tail budget for a UCB empirical mean at time `t` and arm `arm`. This is intentionally an abstract finite-variance budget: it does not prove the variance rate of an empirical mean or choose a log/sqrt UCB radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-335997e9ef0b","parent":"module:BanditRLProof.Algorithms.UCB","order":2629,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:698"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def chebyshevAbsDeviationTail {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : ENNReal","missing":[],"search":"chebyshevabsdeviationtail banditrlproof.ucb.chebyshevabsdeviationtail chebyshev tail budget for a ucb empirical mean at time `t` and arm `arm`. this is intentionally an abstract finite-variance budget: it does not prove the variance rate of an empirical mean or choose a log/sqrt ucb radius. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_absDeviation_le_chebyshev_tail","label":"measure_absDeviation_le_chebyshev_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_absDeviation_le_chebyshev_tail","description":"Single-time Chebyshev tail for the UCB absolute-deviation event, under an explicit mean-identification contract.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-5459ce5154d1","parent":"module:BanditRLProof.Algorithms.UCB","order":2630,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:712"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_absDeviation_le_chebyshev_tail {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hmem : MemLp (fun omega : Omega => empiricalMean omega t arm) 2 mu) (hradius : 0 < radius t arm) (hmean : integral mu (fun omega : Omega => empiricalMean omega t arm) = trueMean arm) : mu {omega | radius t arm <= |empiricalMean omega t arm - trueMean arm|} <= chebyshevAbsDeviationTail mu empiricalMean radius t arm","missing":[],"search":"measure_absdeviation_le_chebyshev_tail banditrlproof.ucb.measure_absdeviation_le_chebyshev_tail single-time chebyshev tail for the ucb absolute-deviation event, under an explicit mean-identification contract. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_chebyshev_tail_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_chebyshev_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_chebyshev_tail_sum","description":"Finite-horizon UCB confidence bad-event bound from Chebyshev absolute-deviation tails. This is a concrete concentration producer for the abstract absolute-deviation tail adapter. It still leaves empirical-mean construction, variance-rate simplification, log/sqrt radius choice, pull-count bounds, and final regret to later leaves.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-aa91a57d9ffd","parent":"module:BanditRLProof.Algorithms.UCB","order":2631,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:743"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_chebyshev_tail_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (T : Nat) (hmem : forall t arm, t < T -> MemLp (fun omega : Omega => empiricalMean omega t arm) 2 mu) (hradius : forall t arm, t < T -> 0 < radius t arm) (hmean : forall t arm, t < T -> integral mu (fun omega : Omega => empiricalMean omega t arm) = trueMean arm) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => chebyshevAbsDeviationTail mu empiricalMean radius t arm + chebyshevAbsDeviationTail mu empiricalMean radius t arm))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_chebyshev_tail_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_chebyshev_tail_sum finite-horizon ucb confidence bad-event bound from chebyshev absolute-deviation tails. this is a concrete concentration producer for the abstract absolute-deviation tail adapter. it still leaves empirical-mean construction, variance-rate simplification, log/sqrt radius choice, pull-count bounds, and final regret to later leaves. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail","label":"subGaussianOneSidedDeviationTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianOneSidedDeviationTail","description":"One-sided sub-Gaussian tail budget for a UCB empirical mean at time `t` and arm `arm`. The proxy is for the centered variable `empiricalMean t arm - trueMean arm`. This one-sided budget is the sharper producer for the existing upper/lower confidence tail consumer; the absolute-deviation wrapper remains available for two-sided concentration statements.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-f2714670b554","parent":"module:BanditRLProof.Algorithms.UCB","order":2632,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:780"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianOneSidedDeviationTail {Arm : Type} (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (t : Nat) (arm : Arm) : ENNReal","missing":[],"search":"subgaussianonesideddeviationtail banditrlproof.ucb.subgaussianonesideddeviationtail one-sided sub-gaussian tail budget for a ucb empirical mean at time `t` and arm `arm`. the proxy is for the centered variable `empiricalmean t arm - truemean arm`. this one-sided budget is the sharper producer for the existing upper/lower confidence tail consumer; the absolute-deviation wrapper remains available for two-sided concentration statements. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail_le_exp_neg_budget","label":"subGaussianOneSidedDeviationTail_le_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianOneSidedDeviationTail_le_exp_neg_budget","description":"Radius-budget simplification for the one-sided UCB sub-Gaussian tail. If `radius^2` dominates `2 * proxy * budget`, the canonical exponential producer is bounded by `exp (-budget)`. This is the algebraic handoff between abstract sub-Gaussian tails and later log/sqrt UCB radius choices.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-5d484e44b3ac","parent":"module:BanditRLProof.Algorithms.UCB","order":2633,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:795"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianOneSidedDeviationTail_le_exp_neg_budget {Arm : Type} (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hradius_sq : 2 * ((proxy t arm : NNReal) : Real) * budget t arm <= (radius t arm) ^ 2) : subGaussianOneSidedDeviationTail radius proxy t arm <= ENNReal.ofReal (Real.exp (-(budget t arm)))","missing":[],"search":"subgaussianonesideddeviationtail_le_exp_neg_budget banditrlproof.ucb.subgaussianonesideddeviationtail_le_exp_neg_budget radius-budget simplification for the one-sided ucb sub-gaussian tail. if `radius^2` dominates `2 * proxy * budget`, the canonical exponential producer is bounded by `exp (-budget)`. this is the algebraic handoff between abstract sub-gaussian tails and later log/sqrt ucb radius choices. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianBudgetRadius","label":"subGaussianBudgetRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianBudgetRadius","description":"Concrete square-root radius associated with a one-sided sub-Gaussian budget. The budget is left abstract so later leaves can instantiate it with logarithmic schedules such as `log (T * |A| / delta)`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-6119d633e601","parent":"module:BanditRLProof.Algorithms.UCB","order":2634,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:825"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianBudgetRadius {Arm : Type} (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) : Nat -> Arm -> Real","missing":[],"search":"subgaussianbudgetradius banditrlproof.ucb.subgaussianbudgetradius concrete square-root radius associated with a one-sided sub-gaussian budget. the budget is left abstract so later leaves can instantiate it with logarithmic schedules such as `log (t * |a| / delta)`. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianBudgetRadius_nonneg","label":"subGaussianBudgetRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianBudgetRadius_nonneg","description":"The concrete square-root budget radius is nonnegative.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-558913b1282c","parent":"module:BanditRLProof.Algorithms.UCB","order":2635,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:833"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianBudgetRadius_nonneg {Arm : Type} (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : 0 <= subGaussianBudgetRadius proxy budget t arm","missing":[],"search":"subgaussianbudgetradius_nonneg banditrlproof.ucb.subgaussianbudgetradius_nonneg the concrete square-root budget radius is nonnegative. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianBudgetRadius_sq_domination","label":"subGaussianBudgetRadius_sq_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianBudgetRadius_sq_domination","description":"The concrete square-root budget radius satisfies the radius-square domination contract consumed by `subGaussianOneSidedDeviationTail_le_exp_neg_budget`. This uses `Real.sq_sqrt'`, so it does not need a separate nonnegativity assumption on `budget`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-b44e6d9edb83","parent":"module:BanditRLProof.Algorithms.UCB","order":2636,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:847"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianBudgetRadius_sq_domination {Arm : Type} (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : 2 * ((proxy t arm : NNReal) : Real) * budget t arm <= (subGaussianBudgetRadius proxy budget t arm) ^ 2","missing":[],"search":"subgaussianbudgetradius_sq_domination banditrlproof.ucb.subgaussianbudgetradius_sq_domination the concrete square-root budget radius satisfies the radius-square domination contract consumed by `subgaussianonesideddeviationtail_le_exp_neg_budget`. this uses `real.sq_sqrt'`, so it does not need a separate nonnegativity assumption on `budget`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail_budgetRadius_le_exp_neg_budget","label":"subGaussianOneSidedDeviationTail_budgetRadius_le_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianOneSidedDeviationTail_budgetRadius_le_exp_neg_budget","description":"One-sided sub-Gaussian tail bound specialized to the concrete square-root budget radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-efbe8f5ab659","parent":"module:BanditRLProof.Algorithms.UCB","order":2637,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:861"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianOneSidedDeviationTail_budgetRadius_le_exp_neg_budget {Arm : Type} (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) : subGaussianOneSidedDeviationTail (subGaussianBudgetRadius proxy budget) proxy t arm <= ENNReal.ofReal (Real.exp (-(budget t arm)))","missing":[],"search":"subgaussianonesideddeviationtail_budgetradius_le_exp_neg_budget banditrlproof.ucb.subgaussianonesideddeviationtail_budgetradius_le_exp_neg_budget one-sided sub-gaussian tail bound specialized to the concrete square-root budget radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_tail","label":"measure_upperConfidenceBad_le_subGaussian_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_tail","description":"Single-time one-sided sub-Gaussian tail for an upper-confidence failure.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-6cc517c0ded7","parent":"module:BanditRLProof.Algorithms.UCB","order":2638,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:874"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_upperConfidenceBad_le_subGaussian_tail {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (t : Nat) (arm : Arm) (hradius : 0 <= radius t arm) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (upperConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) <= subGaussianOneSidedDeviationTail radius proxy t arm","missing":[],"search":"measure_upperconfidencebad_le_subgaussian_tail banditrlproof.ucb.measure_upperconfidencebad_le_subgaussian_tail single-time one-sided sub-gaussian tail for an upper-confidence failure. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_exp_neg_budget","label":"measure_upperConfidenceBad_le_subGaussian_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_exp_neg_budget","description":"Single-time upper-confidence failure bound with an explicit exponential budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-59f575f2eeaf","parent":"module:BanditRLProof.Algorithms.UCB","order":2639,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:926"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_upperConfidenceBad_le_subGaussian_exp_neg_budget {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hradius : 0 <= radius t arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hradius_sq : 2 * ((proxy t arm : NNReal) : Real) * budget t arm <= (radius t arm) ^ 2) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (upperConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) <= ENNReal.ofReal (Real.exp (-(budget t arm)))","missing":[],"search":"measure_upperconfidencebad_le_subgaussian_exp_neg_budget banditrlproof.ucb.measure_upperconfidencebad_le_subgaussian_exp_neg_budget single-time upper-confidence failure bound with an explicit exponential budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_budgetRadius","label":"measure_upperConfidenceBad_le_subGaussian_budgetRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_budgetRadius","description":"Single-time upper-confidence failure bound for the concrete square-root budget radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-23edb389f4cf","parent":"module:BanditRLProof.Algorithms.UCB","order":2640,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:956"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_upperConfidenceBad_le_subGaussian_budgetRadius {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (upperConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (subGaussianBudgetRadius proxy budget t) arm) <= ENNReal.ofReal (Real.exp (-(budget t arm)))","missing":[],"search":"measure_upperconfidencebad_le_subgaussian_budgetradius banditrlproof.ucb.measure_upperconfidencebad_le_subgaussian_budgetradius single-time upper-confidence failure bound for the concrete square-root budget radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_tail","label":"measure_lowerConfidenceBad_le_subGaussian_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_tail","description":"Single-time one-sided sub-Gaussian tail for a lower-confidence failure.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-7b91e6abbdab","parent":"module:BanditRLProof.Algorithms.UCB","order":2641,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:979"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_lowerConfidenceBad_le_subGaussian_tail {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (t : Nat) (arm : Arm) (hradius : 0 <= radius t arm) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (lowerConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) <= subGaussianOneSidedDeviationTail radius proxy t arm","missing":[],"search":"measure_lowerconfidencebad_le_subgaussian_tail banditrlproof.ucb.measure_lowerconfidencebad_le_subgaussian_tail single-time one-sided sub-gaussian tail for a lower-confidence failure. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_exp_neg_budget","label":"measure_lowerConfidenceBad_le_subGaussian_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_exp_neg_budget","description":"Single-time lower-confidence failure bound with an explicit exponential budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d8107414d875","parent":"module:BanditRLProof.Algorithms.UCB","order":2642,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1029"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_lowerConfidenceBad_le_subGaussian_exp_neg_budget {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hradius : 0 <= radius t arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hradius_sq : 2 * ((proxy t arm : NNReal) : Real) * budget t arm <= (radius t arm) ^ 2) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (lowerConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (radius t) arm) <= ENNReal.ofReal (Real.exp (-(budget t arm)))","missing":[],"search":"measure_lowerconfidencebad_le_subgaussian_exp_neg_budget banditrlproof.ucb.measure_lowerconfidencebad_le_subgaussian_exp_neg_budget single-time lower-confidence failure bound with an explicit exponential budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_budgetRadius","label":"measure_lowerConfidenceBad_le_subGaussian_budgetRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_budgetRadius","description":"Single-time lower-confidence failure bound for the concrete square-root budget radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-bff338b50831","parent":"module:BanditRLProof.Algorithms.UCB","order":2643,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1059"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_lowerConfidenceBad_le_subGaussian_budgetRadius {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (lowerConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (subGaussianBudgetRadius proxy budget t) arm) <= ENNReal.ofReal (Real.exp (-(budget t arm)))","missing":[],"search":"measure_lowerconfidencebad_le_subgaussian_budgetradius banditrlproof.ucb.measure_lowerconfidencebad_le_subgaussian_budgetradius single-time lower-confidence failure bound for the concrete square-root budget radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_oneSided_tail_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_oneSided_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_oneSided_tail_sum","description":"Finite-horizon UCB confidence bad-event bound from one-sided sub-Gaussian upper/lower tails. This is the sharper UCB-facing sub-Gaussian producer for the existing upper/lower tail consumer. It still leaves empirical-mean construction, proxy/radius simplification to the textbook log/sqrt form, pull-count bounds, and final regret to later leaves.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-328673d5f98b","parent":"module:BanditRLProof.Algorithms.UCB","order":2644,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1090"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_oneSided_tail_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (T : Nat) (hradius : forall t arm, t < T -> 0 <= radius t arm) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => subGaussianOneSidedDeviationTail radius proxy t arm + subGaussianOneSidedDeviationTail radius proxy t arm))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_onesided_tail_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_onesided_tail_sum finite-horizon ucb confidence bad-event bound from one-sided sub-gaussian upper/lower tails. this is the sharper ucb-facing sub-gaussian producer for the existing upper/lower tail consumer. it still leaves empirical-mean construction, proxy/radius simplification to the textbook log/sqrt form, pull-count bounds, and final regret to later leaves. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_exp_neg_budget_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_exp_neg_budget_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_exp_neg_budget_sum","description":"Finite-horizon confidence bad-event bound with explicit one-sided exponential budgets. This is the UCB radius-budget handoff: later leaves can instantiate `budget` with a log schedule and prove the displayed radius-square domination.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-7af2503930b5","parent":"module:BanditRLProof.Algorithms.UCB","order":2645,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1130"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_exp_neg_budget_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (T : Nat) (hradius : forall t arm, t < T -> 0 <= radius t arm) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hradius_sq : forall t arm, t < T -> 2 * ((proxy t arm : NNReal) : Real) * budget t arm <= (radius t arm) ^ 2) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => ENNReal.ofReal…","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_exp_neg_budget_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_exp_neg_budget_sum finite-horizon confidence bad-event bound with explicit one-sided exponential budgets. this is the ucb radius-budget handoff: later leaves can instantiate `budget` with a log schedule and prove the displayed radius-square domination. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_budgetRadius_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_budgetRadius_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_budgetRadius_sum","description":"Finite-horizon confidence bad-event bound for the concrete square-root budget radius. This is the direct UCB-facing consumer for later logarithmic budget schedules.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-c92731f885cc","parent":"module:BanditRLProof.Algorithms.UCB","order":2646,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1176"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_budgetRadius_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (budget : Nat -> Arm -> Real) (T : Nat) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean (subGaussianBudgetRadius proxy budget) T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => ENNReal.ofReal (Real.exp (-(budget t arm))) + ENNReal.ofReal (Real.exp (-(budget t arm)))))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_budgetradius_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_budgetradius_sum finite-horizon confidence bad-event bound for the concrete square-root budget radius. this is the direct ucb-facing consumer for later logarithmic budget schedules. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.exp_neg_log_eq_inv","label":"exp_neg_log_eq_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.exp_neg_log_eq_inv","description":"Elementary log-budget simplification used by UCB tail producers. The positivity hypothesis is the regularity contract for later concrete schedules such as `scale = T * |A| / delta`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-6abf4dc57607","parent":"module:BanditRLProof.Algorithms.UCB","order":2647,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1212"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_log_eq_inv {x : Real} (hx : 0 < x) : Real.exp (-(Real.log x)) = x⁻¹","missing":[],"search":"exp_neg_log_eq_inv banditrlproof.ucb.exp_neg_log_eq_inv elementary log-budget simplification used by ucb tail producers. the positivity hypothesis is the regularity contract for later concrete schedules such as `scale = t * |a| / delta`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius","label":"subGaussianLogBudgetRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianLogBudgetRadius","description":"Concrete square-root radius with a logarithmic budget. This is still schedule-agnostic: `scale` is the positive quantity whose inverse will become the one-sided tail budget after simplifying `exp (-log scale)`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-2325448cc53f","parent":"module:BanditRLProof.Algorithms.UCB","order":2648,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1222"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianLogBudgetRadius {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) : Nat -> Arm -> Real","missing":[],"search":"subgaussianlogbudgetradius banditrlproof.ucb.subgaussianlogbudgetradius concrete square-root radius with a logarithmic budget. this is still schedule-agnostic: `scale` is the positive quantity whose inverse will become the one-sided tail budget after simplifying `exp (-log scale)`. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius_apply","label":"subGaussianLogBudgetRadius_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianLogBudgetRadius_apply","description":"@[simp] theorem subGaussianLogBudgetRadius_apply {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : subGaussianLogBudgetRadius proxy scale t arm = Real.sqrt (2 * ((proxy t arm : NNReal) : Real) * Real.log (scale t arm))","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-e8bd5d698c65","parent":"module:BanditRLProof.Algorithms.UCB","order":2649,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1228"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem subGaussianLogBudgetRadius_apply {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : subGaussianLogBudgetRadius proxy scale t arm = Real.sqrt (2 * ((proxy t arm : NNReal) : Real) * Real.log (scale t arm))","missing":[],"search":"subgaussianlogbudgetradius_apply banditrlproof.ucb.subgaussianlogbudgetradius_apply @[simp] theorem subgaussianlogbudgetradius_apply {arm : type} (proxy : nat -> arm -> nnreal) (scale : nat -> arm -> real) (t : nat) (arm : arm) : subgaussianlogbudgetradius proxy scale t arm = real.sqrt (2 * ((proxy t arm : nnreal) : real) * real.log (scale t arm)) theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius_nonneg","label":"subGaussianLogBudgetRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianLogBudgetRadius_nonneg","description":"The logarithmic square-root budget radius is nonnegative.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-4c143f8bde9a","parent":"module:BanditRLProof.Algorithms.UCB","order":2650,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1237"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianLogBudgetRadius_nonneg {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : 0 <= subGaussianLogBudgetRadius proxy scale t arm","missing":[],"search":"subgaussianlogbudgetradius_nonneg banditrlproof.ucb.subgaussianlogbudgetradius_nonneg the logarithmic square-root budget radius is nonnegative. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius_sq_domination","label":"subGaussianLogBudgetRadius_sq_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianLogBudgetRadius_sq_domination","description":"The logarithmic square-root budget radius satisfies the square-domination contract with budget `log scale`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-9170a338f290","parent":"module:BanditRLProof.Algorithms.UCB","order":2651,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1250"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianLogBudgetRadius_sq_domination {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) : 2 * ((proxy t arm : NNReal) : Real) * Real.log (scale t arm) <= (subGaussianLogBudgetRadius proxy scale t arm) ^ 2","missing":[],"search":"subgaussianlogbudgetradius_sq_domination banditrlproof.ucb.subgaussianlogbudgetradius_sq_domination the logarithmic square-root budget radius satisfies the square-domination contract with budget `log scale`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail_logBudgetRadius_le_inv_scale","label":"subGaussianOneSidedDeviationTail_logBudgetRadius_le_inv_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianOneSidedDeviationTail_logBudgetRadius_le_inv_scale","description":"One-sided sub-Gaussian tail bound specialized to a logarithmic square-root budget radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-f2d9c5558b35","parent":"module:BanditRLProof.Algorithms.UCB","order":2652,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1264"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianOneSidedDeviationTail_logBudgetRadius_le_inv_scale {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hscale : 0 < scale t arm) : subGaussianOneSidedDeviationTail (subGaussianLogBudgetRadius proxy scale) proxy t arm <= ENNReal.ofReal ((scale t arm)⁻¹)","missing":[],"search":"subgaussianonesideddeviationtail_logbudgetradius_le_inv_scale banditrlproof.ucb.subgaussianonesideddeviationtail_logbudgetradius_le_inv_scale one-sided sub-gaussian tail bound specialized to a logarithmic square-root budget radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_logBudgetRadius","label":"measure_upperConfidenceBad_le_subGaussian_logBudgetRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_logBudgetRadius","description":"Single-time upper-confidence failure bound for the logarithmic square-root budget radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d5c7e5723d03","parent":"module:BanditRLProof.Algorithms.UCB","order":2653,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1282"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_upperConfidenceBad_le_subGaussian_logBudgetRadius {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hscale : 0 < scale t arm) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (upperConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (subGaussianLogBudgetRadius proxy scale t) arm) <= ENNReal.ofReal ((scale t arm)⁻¹)","missing":[],"search":"measure_upperconfidencebad_le_subgaussian_logbudgetradius banditrlproof.ucb.measure_upperconfidencebad_le_subgaussian_logbudgetradius single-time upper-confidence failure bound for the logarithmic square-root budget radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_logBudgetRadius","label":"measure_lowerConfidenceBad_le_subGaussian_logBudgetRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_logBudgetRadius","description":"Single-time lower-confidence failure bound for the logarithmic square-root budget radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-b6e91277ee8d","parent":"module:BanditRLProof.Algorithms.UCB","order":2654,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1309"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_lowerConfidenceBad_le_subGaussian_logBudgetRadius {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (t : Nat) (arm : Arm) (hproxy : 0 < ((proxy t arm : NNReal) : Real)) (hscale : 0 < scale t arm) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (lowerConfidenceBad trueMean (fun omega arm => empiricalMean omega t arm) (subGaussianLogBudgetRadius proxy scale t) arm) <= ENNReal.ofReal ((scale t arm)⁻¹)","missing":[],"search":"measure_lowerconfidencebad_le_subgaussian_logbudgetradius banditrlproof.ucb.measure_lowerconfidencebad_le_subgaussian_logbudgetradius single-time lower-confidence failure bound for the logarithmic square-root budget radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_logBudgetRadius_inv_scale_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_logBudgetRadius_inv_scale_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_logBudgetRadius_inv_scale_sum","description":"Finite-horizon confidence bad-event bound for logarithmic square-root budget radii. This is the schedule-agnostic log-budget producer. Later UCB leaves can set `scale t arm` to a concrete positive expression and then simplify the double sum.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-3467fd74ef9f","parent":"module:BanditRLProof.Algorithms.UCB","order":2655,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1340"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_logBudgetRadius_inv_scale_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (scale : Nat -> Arm -> Real) (T : Nat) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hscale : forall t arm, t < T -> 0 < scale t arm) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean (subGaussianLogBudgetRadius proxy scale) T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => ENNReal.ofReal ((scale t arm)⁻¹) + ENNReal.ofReal ((scale t arm)⁻¹)))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_logbudgetradius_inv_scale_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_logbudgetradius_inv_scale_sum finite-horizon confidence bad-event bound for logarithmic square-root budget radii. this is the schedule-agnostic log-budget producer. later ucb leaves can set `scale t arm` to a concrete positive expression and then simplify the double sum. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius","label":"subGaussianConstantLogBudgetRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianConstantLogBudgetRadius","description":"Logarithmic square-root radius with a constant positive scale. This is the direct finite-horizon shape for later choices such as `scale = T * |A| / delta`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-53fc430a1202","parent":"module:BanditRLProof.Algorithms.UCB","order":2656,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1387"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianConstantLogBudgetRadius {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Real) : Nat -> Arm -> Real","missing":[],"search":"subgaussianconstantlogbudgetradius banditrlproof.ucb.subgaussianconstantlogbudgetradius logarithmic square-root radius with a constant positive scale. this is the direct finite-horizon shape for later choices such as `scale = t * |a| / delta`. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_apply","label":"subGaussianConstantLogBudgetRadius_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_apply","description":"@[simp] theorem subGaussianConstantLogBudgetRadius_apply {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Real) (t : Nat) (arm : Arm) : subGaussianConstantLogBudgetRadius proxy scale t arm = Real.sqrt (2 * ((proxy t arm : NNReal) : Real) * Real.log scale)","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d5ee40b3f080","parent":"module:BanditRLProof.Algorithms.UCB","order":2657,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1393"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem subGaussianConstantLogBudgetRadius_apply {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Real) (t : Nat) (arm : Arm) : subGaussianConstantLogBudgetRadius proxy scale t arm = Real.sqrt (2 * ((proxy t arm : NNReal) : Real) * Real.log scale)","missing":[],"search":"subgaussianconstantlogbudgetradius_apply banditrlproof.ucb.subgaussianconstantlogbudgetradius_apply @[simp] theorem subgaussianconstantlogbudgetradius_apply {arm : type} (proxy : nat -> arm -> nnreal) (scale : real) (t : nat) (arm : arm) : subgaussianconstantlogbudgetradius proxy scale t arm = real.sqrt (2 * ((proxy t arm : nnreal) : real) * real.log scale) theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_nonneg","label":"subGaussianConstantLogBudgetRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_nonneg","description":"Constant logarithmic square-root budget radii are nonnegative.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-979dd2ae3f75","parent":"module:BanditRLProof.Algorithms.UCB","order":2658,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1402"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianConstantLogBudgetRadius_nonneg {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Real) (t : Nat) (arm : Arm) : 0 <= subGaussianConstantLogBudgetRadius proxy scale t arm","missing":[],"search":"subgaussianconstantlogbudgetradius_nonneg banditrlproof.ucb.subgaussianconstantlogbudgetradius_nonneg constant logarithmic square-root budget radii are nonnegative. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_sq_domination","label":"subGaussianConstantLogBudgetRadius_sq_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_sq_domination","description":"Constant logarithmic square-root budget radii satisfy the square-domination contract with budget `log scale`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a095ac604cc7","parent":"module:BanditRLProof.Algorithms.UCB","order":2659,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1414"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianConstantLogBudgetRadius_sq_domination {Arm : Type} (proxy : Nat -> Arm -> NNReal) (scale : Real) (t : Nat) (arm : Arm) : 2 * ((proxy t arm : NNReal) : Real) * Real.log scale <= (subGaussianConstantLogBudgetRadius proxy scale t arm) ^ 2","missing":[],"search":"subgaussianconstantlogbudgetradius_sq_domination banditrlproof.ucb.subgaussianconstantlogbudgetradius_sq_domination constant logarithmic square-root budget radii satisfy the square-domination contract with budget `log scale`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constant_invScale_double_sum","label":"constant_invScale_double_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.constant_invScale_double_sum","description":"Double finite sum of a constant inverse-scale one-sided tail budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a529e601bdaf","parent":"module:BanditRLProof.Algorithms.UCB","order":2660,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1427"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem constant_invScale_double_sum {Arm : Type} [Fintype Arm] (T : Nat) (scale : Real) : (Finset.range T).sum (fun _ => (Finset.univ : Finset Arm).sum (fun _ => ENNReal.ofReal scale⁻¹ + ENNReal.ofReal scale⁻¹)) = HSMul.hSMul T (HSMul.hSMul (Fintype.card Arm) (ENNReal.ofReal scale⁻¹ + ENNReal.ofReal scale⁻¹))","missing":[],"search":"constant_invscale_double_sum banditrlproof.ucb.constant_invscale_double_sum double finite sum of a constant inverse-scale one-sided tail budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_constantLogBudgetRadius_card","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_constantLogBudgetRadius_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_constantLogBudgetRadius_card","description":"Finite-horizon confidence bad-event bound for a constant logarithmic scale, with the time/arm double sum folded into `T` and `Fintype.card Arm`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d716c2a758dc","parent":"module:BanditRLProof.Algorithms.UCB","order":2661,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1443"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_constantLogBudgetRadius_card {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (scale : Real) (T : Nat) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hscale : 0 < scale) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean (subGaussianConstantLogBudgetRadius proxy scale) T) <= HSMul.hSMul T (HSMul.hSMul (Fintype.card Arm) (ENNReal.ofReal scale⁻¹ + ENNReal.ofReal scale⁻¹))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_constantlogbudgetradius_card banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_constantlogbudgetradius_card finite-horizon confidence bad-event bound for a constant logarithmic scale, with the time/arm double sum folded into `t` and `fintype.card arm`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constant_invScale_double_sum_le_of_real","label":"constant_invScale_double_sum_le_of_real","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.constant_invScale_double_sum_le_of_real","description":"Convert a constant inverse-scale finite-horizon ENNReal tail budget back to an ordinary real budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-c6ad961f9c8f","parent":"module:BanditRLProof.Algorithms.UCB","order":2662,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1480"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem constant_invScale_double_sum_le_of_real {Arm : Type} [Fintype Arm] (T : Nat) (scale delta : Real) (hscale : 0 < scale) (hreal : (T : Real) * ((Fintype.card Arm : Real) * (scale⁻¹ + scale⁻¹)) <= delta) : HSMul.hSMul T (HSMul.hSMul (Fintype.card Arm) (ENNReal.ofReal scale⁻¹ + ENNReal.ofReal scale⁻¹)) <= ENNReal.ofReal delta","missing":[],"search":"constant_invscale_double_sum_le_of_real banditrlproof.ucb.constant_invscale_double_sum_le_of_real convert a constant inverse-scale finite-horizon ennreal tail budget back to an ordinary real budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.textbookDeltaScale","label":"textbookDeltaScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.textbookDeltaScale","description":"Textbook UCB finite-horizon scale for allocating two one-sided tails across `T` times and all arms. The factor `2` accounts for upper and lower confidence failures.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-2b6e7cd52e66","parent":"module:BanditRLProof.Algorithms.UCB","order":2663,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1508"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def textbookDeltaScale {Arm : Type} [Fintype Arm] (T : Nat) (delta : Real) : Real","missing":[],"search":"textbookdeltascale banditrlproof.ucb.textbookdeltascale textbook ucb finite-horizon scale for allocating two one-sided tails across `t` times and all arms. the factor `2` accounts for upper and lower confidence failures. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.textbookDeltaScale_pos","label":"textbookDeltaScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.textbookDeltaScale_pos","description":"The textbook delta scale is positive under the usual horizon/arm/delta contracts.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-b5e4a451ee00","parent":"module:BanditRLProof.Algorithms.UCB","order":2664,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1513"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem textbookDeltaScale_pos {Arm : Type} [Fintype Arm] [Nonempty Arm] (T : Nat) (delta : Real) (hT : 0 < T) (hdelta : 0 < delta) : 0 < textbookDeltaScale (Arm := Arm) T delta","missing":[],"search":"textbookdeltascale_pos banditrlproof.ucb.textbookdeltascale_pos the textbook delta scale is positive under the usual horizon/arm/delta contracts. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.textbookDeltaScale_total_inv_budget_eq_delta","label":"textbookDeltaScale_total_inv_budget_eq_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.textbookDeltaScale_total_inv_budget_eq_delta","description":"The textbook delta scale makes the folded constant inverse-scale tail budget equal to `delta` at the real-number level.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-ff249f6e2076","parent":"module:BanditRLProof.Algorithms.UCB","order":2665,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1525"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem textbookDeltaScale_total_inv_budget_eq_delta {Arm : Type} [Fintype Arm] [Nonempty Arm] (T : Nat) (delta : Real) (hT : 0 < T) (hdelta : 0 < delta) : (T : Real) * ((Fintype.card Arm : Real) * ((textbookDeltaScale (Arm := Arm) T delta)⁻¹ + (textbookDeltaScale (Arm := Arm) T delta)⁻¹)) = delta","missing":[],"search":"textbookdeltascale_total_inv_budget_eq_delta banditrlproof.ucb.textbookdeltascale_total_inv_budget_eq_delta the textbook delta scale makes the folded constant inverse-scale tail budget equal to `delta` at the real-number level. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constant_invScale_double_sum_textbookDeltaScale_le_delta","label":"constant_invScale_double_sum_textbookDeltaScale_le_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.constant_invScale_double_sum_textbookDeltaScale_le_delta","description":"The folded constant-scale UCB tail budget with textbook delta scale is bounded by `delta`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-31c43554c76f","parent":"module:BanditRLProof.Algorithms.UCB","order":2666,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1548"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem constant_invScale_double_sum_textbookDeltaScale_le_delta {Arm : Type} [Fintype Arm] [Nonempty Arm] (T : Nat) (delta : Real) (hT : 0 < T) (hdelta : 0 < delta) : HSMul.hSMul T (HSMul.hSMul (Fintype.card Arm) (ENNReal.ofReal (textbookDeltaScale (Arm := Arm) T delta)⁻¹ + ENNReal.ofReal (textbookDeltaScale (Arm := Arm) T delta)⁻¹)) <= ENNReal.ofReal delta","missing":[],"search":"constant_invscale_double_sum_textbookdeltascale_le_delta banditrlproof.ucb.constant_invscale_double_sum_textbookdeltascale_le_delta the folded constant-scale ucb tail budget with textbook delta scale is bounded by `delta`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius","label":"subGaussianTextbookDeltaRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius","description":"Textbook delta-scale logarithmic UCB radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-1bbcc345785f","parent":"module:BanditRLProof.Algorithms.UCB","order":2667,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1564"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianTextbookDeltaRadius {Arm : Type} [Fintype Arm] (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) : Nat -> Arm -> Real","missing":[],"search":"subgaussiantextbookdeltaradius banditrlproof.ucb.subgaussiantextbookdeltaradius textbook delta-scale logarithmic ucb radius. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_apply","label":"subGaussianTextbookDeltaRadius_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_apply","description":"@[simp] theorem subGaussianTextbookDeltaRadius_apply {Arm : Type} [Fintype Arm] (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) (t : Nat) (arm : Arm) : subGaussianTextbookDeltaRadius proxy T delta t arm = Real.sqrt (2 * ((proxy t arm : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Arm) T delta))","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-61dac685c4e0","parent":"module:BanditRLProof.Algorithms.UCB","order":2668,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1571"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem subGaussianTextbookDeltaRadius_apply {Arm : Type} [Fintype Arm] (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) (t : Nat) (arm : Arm) : subGaussianTextbookDeltaRadius proxy T delta t arm = Real.sqrt (2 * ((proxy t arm : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Arm) T delta))","missing":[],"search":"subgaussiantextbookdeltaradius_apply banditrlproof.ucb.subgaussiantextbookdeltaradius_apply @[simp] theorem subgaussiantextbookdeltaradius_apply {arm : type} [fintype arm] (proxy : nat -> arm -> nnreal) (t : nat) (delta : real) (t : nat) (arm : arm) : subgaussiantextbookdeltaradius proxy t delta t arm = real.sqrt (2 * ((proxy t arm : nnreal) : real) * real.log (textbookdeltascale (arm := arm) t delta)) theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_textbookDeltaRadius_delta","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_textbookDeltaRadius_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_textbookDeltaRadius_delta","description":"Finite-horizon confidence bad-event bound for the textbook delta-scale logarithmic UCB radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d8b59e999799","parent":"module:BanditRLProof.Algorithms.UCB","order":2669,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1584"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_textbookDeltaRadius_delta {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] [Nonempty Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) (hT : 0 < T) (hdelta : 0 < delta) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) T) <= ENNReal.ofReal delta","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_textbookdeltaradius_delta banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_textbookdeltaradius_delta finite-horizon confidence bad-event bound for the textbook delta-scale logarithmic ucb radius. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_scoreMaxEvent_le_subGaussian_textbookDeltaRadius_delta_of_gap","label":"measure_scoreMaxEvent_le_subGaussian_textbookDeltaRadius_delta_of_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_scoreMaxEvent_le_subGaussian_textbookDeltaRadius_delta_of_gap","description":"Large-gap score-max events under the textbook delta radius are controlled by the finite-horizon confidence budget. This is a probability-facing handoff for later pull-count arguments: once a chosen arm has gap larger than twice its current radius, selecting it by UCB score can only happen on the confidence bad event.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-4c3526c30fac","parent":"module:BanditRLProof.Algorithms.UCB","order":2670,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1617"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_scoreMaxEvent_le_subGaussian_textbookDeltaRadius_delta_of_gap {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] [Nonempty Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) (t : Nat) (best chosen : Arm) (hT : 0 < T) (hdelta : 0 < delta) (ht : t < T) (hgap_large : 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu {omega : Omega | confidenceScore (empiricalMean omega t) (subGaussianTextbookDeltaRadius proxy T delta t) best <= confidenceScore (empiricalMean omega t) (subGaussianTextbook…","missing":[],"search":"measure_scoremaxevent_le_subgaussian_textbookdeltaradius_delta_of_gap banditrlproof.ucb.measure_scoremaxevent_le_subgaussian_textbookdeltaradius_delta_of_gap large-gap score-max events under the textbook delta radius are controlled by the finite-horizon confidence budget. this is a probability-facing handoff for later pull-count arguments: once a chosen arm has gap larger than twice its current radius, selecting it by ucb score can only happen on the confidence bad event. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedEvent_subset_scoreMaxEvent_of_action_score_max","label":"selectedEvent_subset_scoreMaxEvent_of_action_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedEvent_subset_scoreMaxEvent_of_action_score_max","description":"Selecting `chosen` is contained in the corresponding UCB score-max event when the action trace exposes score maximality against `best`. This is intentionally abstract: the concrete argmax/tie-breaking policy can later discharge `hscore_of_selected`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-35a932f8b1f6","parent":"module:BanditRLProof.Algorithms.UCB","order":2671,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1662"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedEvent_subset_scoreMaxEvent_of_action_score_max {Omega Arm : Type} (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (action : Omega -> Nat -> Arm) (t : Nat) (best chosen : Arm) (hscore_of_selected : forall omega, action omega t = chosen -> confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen) : Set.Subset {omega : Omega | action omega t = chosen} {omega : Omega | confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen}","missing":[],"search":"selectedevent_subset_scoremaxevent_of_action_score_max banditrlproof.ucb.selectedevent_subset_scoremaxevent_of_action_score_max selecting `chosen` is contained in the corresponding ucb score-max event when the action trace exposes score maximality against `best`. this is intentionally abstract: the concrete argmax/tie-breaking policy can later discharge `hscore_of_selected`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","label":"measure_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","description":"Selected large-gap arms inherit the textbook delta probability budget once the selected-action trace certifies UCB score maximality. This is the action-trace-facing bridge before a concrete pull-count summation: selection plus a large-gap radius condition implies that the selected event is covered by the finite-horizon confidence bad event.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-0ed54318af8d","parent":"module:BanditRLProof.Algorithms.UCB","order":2672,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1688"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] [Nonempty Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (action : Omega -> Nat -> Arm) (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) (t : Nat) (best chosen : Arm) (hT : 0 < T) (hdelta : 0 < delta) (ht : t < T) (hgap_large : 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hscore_of_selected : forall omega, action omega t = chosen -> confidenceScore (empiricalMean omega t) (subGaussianTextbookDeltaRadius proxy T delta t) best <= confidenceScore (empiricalMean omega t) (subGaussianTextbookDeltaRadius proxy T delta t) chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> Probabil…","missing":[],"search":"measure_selectedlargegapevent_le_subgaussian_textbookdeltaradius_delta banditrlproof.ucb.measure_selectedlargegapevent_le_subgaussian_textbookdeltaradius_delta selected large-gap arms inherit the textbook delta probability budget once the selected-action trace certifies ucb score maximality. this is the action-trace-facing bridge before a concrete pull-count summation: selection plus a large-gap radius condition implies that the selected event is covered by the finite-horizon confidence bad event. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedEventOn_subset_finiteHorizonConfidenceBadEvent_of_action_score_max","label":"selectedEventOn_subset_finiteHorizonConfidenceBadEvent_of_action_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedEventOn_subset_finiteHorizonConfidenceBadEvent_of_action_score_max","description":"Finite-time selected-action events are covered by the finite-horizon confidence bad event when every selected time in the index set has a large enough gap and certifies UCB score maximality. This is the event-level bridge needed before turning selected-time collections into pull-count or suffix-time bounds.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-1098add6947a","parent":"module:BanditRLProof.Algorithms.UCB","order":2673,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1737"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedEventOn_subset_finiteHorizonConfidenceBadEvent_of_action_score_max {Omega Arm : Type} [Fintype Arm] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (action : Omega -> Nat -> Arm) (T : Nat) (times : Finset Nat) (best chosen : Arm) (htimes : forall t, t ∈ times -> t < T) (hscore_of_selected : forall omega t, t ∈ times -> action omega t = chosen -> confidenceScore (empiricalMean omega t) (radius t) best <= confidenceScore (empiricalMean omega t) (radius t) chosen) (hgap_large : forall t, t ∈ times -> 2 * radius t chosen < meanGap trueMean best chosen) : Set.Subset {omega : Omega | exists t, t ∈ times /\\ action omega t = chosen} (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T)","missing":[],"search":"selectedeventon_subset_finitehorizonconfidencebadevent_of_action_score_max banditrlproof.ucb.selectedeventon_subset_finitehorizonconfidencebadevent_of_action_score_max finite-time selected-action events are covered by the finite-horizon confidence bad event when every selected time in the index set has a large enough gap and certifies ucb score maximality. this is the event-level bridge needed before turning selected-time collections into pull-count or suffix-time bounds. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","label":"measure_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","description":"The textbook delta budget also controls the event that a fixed arm is selected at any time from a finite index set, provided each such time satisfies the large-gap radius condition and selected-action score maximality.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-8c6d8e1eb5e0","parent":"module:BanditRLProof.Algorithms.UCB","order":2674,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1769"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] [Nonempty Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (action : Omega -> Nat -> Arm) (proxy : Nat -> Arm -> NNReal) (T : Nat) (delta : Real) (times : Finset Nat) (best chosen : Arm) (hT : 0 < T) (hdelta : 0 < delta) (htimes : forall t, t ∈ times -> t < T) (hgap_large : forall t, t ∈ times -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hscore_of_selected : forall omega t, t ∈ times -> action omega t = chosen -> confidenceScore (empiricalMean omega t) (subGaussianTextbookDeltaRadius proxy T delta t) best <= confidenceScore (empiricalMean omega t) (subGaussianTextbookDeltaRadius proxy T delta t) chosen) (hproxy : forall t arm, t < T ->…","missing":[],"search":"measure_selectedlargegapeventon_le_subgaussian_textbookdeltaradius_delta banditrlproof.ucb.measure_selectedlargegapeventon_le_subgaussian_textbookdeltaradius_delta the textbook delta budget also controls the event that a fixed arm is selected at any time from a finite index set, provided each such time satisfies the large-gap radius condition and selected-action score maximality. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","label":"measure_confidenceScoreArgmax_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","description":"Concrete confidence-score argmax version of the selected large-gap delta bound. The score-maximality contract is discharged by `confidenceScoreArgmaxAction`, so the remaining assumptions are the textbook radius/concentration contracts and the large-gap condition for the selected arm.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a00f6ffac3b2","parent":"module:BanditRLProof.Algorithms.UCB","order":2675,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1815"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_confidenceScoreArgmax_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (t : Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (ht : t < T) (hgap_large : 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu {omega : Omega | confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t = chosen} <= ENNReal.ofReal delta","missing":[],"search":"measure_confidencescoreargmax_selectedlargegapevent_le_subgaussian_textbookdeltaradius_delta banditrlproof.ucb.measure_confidencescoreargmax_selectedlargegapevent_le_subgaussian_textbookdeltaradius_delta concrete confidence-score argmax version of the selected large-gap delta bound. the score-maximality contract is discharged by `confidencescoreargmaxaction`, so the remaining assumptions are the textbook radius/concentration contracts and the large-gap condition for the selected arm. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","label":"measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","description":"Finite-time-set concrete confidence-score argmax version of the selected large-gap delta bound.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-1a9cd4416008","parent":"module:BanditRLProof.Algorithms.UCB","order":2676,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1857"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (times : Finset Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (htimes : forall t, t ∈ times -> t < T) (hgap_large : forall t, t ∈ times -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu {omega : Omega | exists t, t ∈ times /\\ confidenceScoreArgmaxAction hK empiricalMean (subG…","missing":[],"search":"measure_confidencescoreargmax_selectedlargegapeventon_le_subgaussian_textbookdeltaradius_delta banditrlproof.ucb.measure_confidencescoreargmax_selectedlargegapeventon_le_subgaussian_textbookdeltaradius_delta finite-time-set concrete confidence-score argmax version of the selected large-gap delta bound. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sum_measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_card_mul_delta","label":"sum_measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_card_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sum_measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_card_mul_delta","description":"Summing the single-time concrete score-argmax selected large-gap bounds over a finite time set gives a finite-count probability budget. This is the first counting-facing UCB bridge: it keeps the statement as a sum of selected-action event probabilities before converting it to a lower integral of selected-time indicators.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-994adbbd45c0","parent":"module:BanditRLProof.Algorithms.UCB","order":2677,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1905"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sum_measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_card_mul_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (times : Finset Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (htimes : forall t, t ∈ times -> t < T) (hgap_large : forall t, t ∈ times -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : times.sum (fun t : Nat => mu {omega : Omega | confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookD…","missing":[],"search":"sum_measure_confidencescoreargmax_selectedlargegapeventon_le_card_mul_delta banditrlproof.ucb.sum_measure_confidencescoreargmax_selectedlargegapeventon_le_card_mul_delta summing the single-time concrete score-argmax selected large-gap bounds over a finite time set gives a finite-count probability budget. this is the first counting-facing ucb bridge: it keeps the statement as a sum of selected-action event probabilities before converting it to a lower integral of selected-time indicators. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_selectedLargeGapCountOn_le_card_mul_delta","label":"lintegral_confidenceScoreArgmax_selectedLargeGapCountOn_le_card_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_selectedLargeGapCountOn_le_card_mul_delta","description":"Lower-integral selected-count budget for concrete score-argmax UCB over an explicit finite time set. The integrand is the finite sum of selected-action indicators over `times`. This is not yet the recursive `pullCount`, but it is the expectation-facing finite-count surface needed to bridge the selected-event probability bounds into pull-count and regret arguments.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-7077760c5fe8","parent":"module:BanditRLProof.Algorithms.UCB","order":2678,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:1960"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_selectedLargeGapCountOn_le_card_mul_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (times : Finset Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (htimes : forall t, t ∈ times -> t < T) (hgap_large : forall t, t ∈ times -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> Probability…","missing":[],"search":"lintegral_confidencescoreargmax_selectedlargegapcounton_le_card_mul_delta banditrlproof.ucb.lintegral_confidencescoreargmax_selectedlargegapcounton_le_card_mul_delta lower-integral selected-count budget for concrete score-argmax ucb over an explicit finite time set. the integrand is the finite sum of selected-action indicators over `times`. this is not yet the recursive `pullcount`, but it is the expectation-facing finite-count surface needed to bridge the selected-event probability bounds into pull-count and regret arguments. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_horizon_mul_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_horizon_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_horizon_mul_delta","description":"Recursive pull-count lower-integral budget for a concrete score-argmax UCB arm whose large-gap condition holds throughout the horizon. This specializes the finite-time selected-count bridge to `Finset.range T` and then uses the existing project-local `pullCount` lower-integral identity. It does not yet split the horizon into small-radius and large-radius phases.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-b619e6140d34","parent":"module:BanditRLProof.Algorithms.UCB","order":2679,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2016"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_horizon_mul_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (hgap_large : forall t, t < T -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - t…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_horizon_mul_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_horizon_mul_delta recursive pull-count lower-integral budget for a concrete score-argmax ucb arm whose large-gap condition holds throughout the horizon. this specializes the finite-time selected-count bridge to `finset.range t` and then uses the existing project-local `pullcount` lower-integral identity. it does not yet split the horizon into small-radius and large-radius phases. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_free_or_delta_sum","label":"lintegral_confidenceScoreArgmax_pullCount_le_free_or_delta_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_free_or_delta_sum","description":"Threshold/suffix-shaped pull-count budget for concrete score-argmax UCB. Times in `freeTimes` are charged by the trivial probability bound `1`; every other horizon time must be listed in `chargedTimes` and satisfy the large-gap condition, so those selected events are charged by `delta`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-1330e3f30dd1","parent":"module:BanditRLProof.Algorithms.UCB","order":2680,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2074"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_free_or_delta_sum {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (freeTimes chargedTimes : Finset Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (hcharged_of_not_free : forall t, t < T -> t ∉ freeTimes -> t ∈ chargedTimes) (hgap_large : forall t, t ∈ chargedTimes -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0 < ((prox…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_free_or_delta_sum banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_free_or_delta_sum threshold/suffix-shaped pull-count budget for concrete score-argmax ucb. times in `freetimes` are charged by the trivial probability bound `1`; every other horizon time must be listed in `chargedtimes` and satisfy the large-gap condition, so those selected events are charged by `delta`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_freeBudget_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_freeBudget_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_freeBudget_add_horizon_delta","description":"Budgeted form of the threshold/suffix pull-count split. The only new input is a bound on the free-time indicator sum. Future radius-threshold leaves can discharge `hfree_budget` by proving a cardinality bound for the low-radius/small-sample times.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-5b1b07e0da82","parent":"module:BanditRLProof.Algorithms.UCB","order":2681,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2145"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_freeBudget_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (freeTimes chargedTimes : Finset Nat) (freeBudget : ENNReal) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (hfree_budget : (Finset.range T).sum (fun t : Nat => if t ∈ freeTimes then (1 : ENNReal) else 0) <= freeBudget) (hcharged_of_not_free : forall t, t < T -> t ∉ freeTimes -> t ∈ chargedTimes) (hgap_large : forall t, t ∈ chargedTimes -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (haction : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_freebudget_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_freebudget_add_horizon_delta budgeted form of the threshold/suffix pull-count split. the only new input is a bound on the free-time indicator sum. future radius-threshold leaves can discharge `hfree_budget` by proving a cardinality bound for the low-radius/small-sample times. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.freeTimes_indicator_sum_le_card","label":"freeTimes_indicator_sum_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.freeTimes_indicator_sum_le_card","description":"The ENNReal indicator count of a finite set of free horizon times is bounded by the total number of declared free times.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-6bec22f741d2","parent":"module:BanditRLProof.Algorithms.UCB","order":2682,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2226"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem freeTimes_indicator_sum_le_card (T : Nat) (freeTimes : Finset Nat) : (Finset.range T).sum (fun t : Nat => if t ∈ freeTimes then (1 : ENNReal) else 0) <= (freeTimes.card : ENNReal)","missing":[],"search":"freetimes_indicator_sum_le_card banditrlproof.ucb.freetimes_indicator_sum_le_card the ennreal indicator count of a finite set of free horizon times is bounded by the total number of declared free times. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedSmallPullCount_sum_eq_min_pullCount","label":"selectedSmallPullCount_sum_eq_min_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedSmallPullCount_sum_eq_min_pullCount","description":"Along one concrete action trace, the number of selected times whose previous pull count is still below threshold `B` is the minimum of the terminal pull count and `B`. This is the pathwise source of the usual UCB small-count budget: selected occurrences with `pullCount < B` can happen at most `B` times, regardless of the ambient horizon length.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-854db5886b8b","parent":"module:BanditRLProof.Algorithms.UCB","order":2683,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2256"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedSmallPullCount_sum_eq_min_pullCount {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T B : Nat) : (Finset.range T).sum (fun t : Nat => if action t = chosen ∧ pullCount action chosen t < B then (1 : Nat) else 0) = Nat.min (pullCount action chosen T) B","missing":[],"search":"selectedsmallpullcount_sum_eq_min_pullcount banditrlproof.ucb.selectedsmallpullcount_sum_eq_min_pullcount along one concrete action trace, the number of selected times whose previous pull count is still below threshold `b` is the minimum of the terminal pull count and `b`. this is the pathwise source of the usual ucb small-count budget: selected occurrences with `pullcount < b` can happen at most `b` times, regardless of the ambient horizon length. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedSmallPullCount_sum_le_threshold","label":"selectedSmallPullCount_sum_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedSmallPullCount_sum_le_threshold","description":"Pathwise UCB small-count budget in Nat form.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-ae25d234f433","parent":"module:BanditRLProof.Algorithms.UCB","order":2684,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2316"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedSmallPullCount_sum_le_threshold {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T B : Nat) : (Finset.range T).sum (fun t : Nat => if action t = chosen ∧ pullCount action chosen t < B then (1 : Nat) else 0) <= B","missing":[],"search":"selectedsmallpullcount_sum_le_threshold banditrlproof.ucb.selectedsmallpullcount_sum_le_threshold pathwise ucb small-count budget in nat form. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedSmallPullCount_indicator_sum_le_threshold","label":"selectedSmallPullCount_indicator_sum_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedSmallPullCount_indicator_sum_le_threshold","description":"ENNReal-facing pathwise UCB small-count budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-33a3a653db8b","parent":"module:BanditRLProof.Algorithms.UCB","order":2685,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2331"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedSmallPullCount_indicator_sum_le_threshold {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T B : Nat) : (Finset.range T).sum (fun t : Nat => if action t = chosen ∧ pullCount action chosen t < B then (1 : ENNReal) else 0) <= (B : ENNReal)","missing":[],"search":"selectedsmallpullcount_indicator_sum_le_threshold banditrlproof.ucb.selectedsmallpullcount_indicator_sum_le_threshold ennreal-facing pathwise ucb small-count budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_selectedSmallPullCount_indicator_sum_le_threshold","label":"lintegral_selectedSmallPullCount_indicator_sum_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_selectedSmallPullCount_indicator_sum_le_threshold","description":"Probability-facing version of the pathwise selected-small budget. No measurability assumption is needed for this upper bound: the lower integral is dominated pointwise by the constant `B`, and the measure is a probability measure.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-03bd699c57eb","parent":"module:BanditRLProof.Algorithms.UCB","order":2686,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2361"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_selectedSmallPullCount_indicator_sum_le_threshold {Omega Action : Type} [MeasurableSpace Omega] [DecidableEq Action] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (action : Omega -> ActionTrace Action) (chosen : Action) (T B : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => (Finset.range T).sum (fun t : Nat => if action omega t = chosen ∧ pullCount (action omega) chosen t < B then (1 : ENNReal) else 0)) <= (B : ENNReal)","missing":[],"search":"lintegral_selectedsmallpullcount_indicator_sum_le_threshold banditrlproof.ucb.lintegral_selectedsmallpullcount_indicator_sum_le_threshold probability-facing version of the pathwise selected-small budget. no measurability assumption is needed for this upper bound: the lower integral is dominated pointwise by the constant `b`, and the measure is a probability measure. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPullCount_sum_eq_pullCount","label":"selectedPullCount_sum_eq_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPullCount_sum_eq_pullCount","description":"The selected-time Nat indicator sum is exactly the recursive pull count.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-87fcd6d01e40","parent":"module:BanditRLProof.Algorithms.UCB","order":2687,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2398"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPullCount_sum_eq_pullCount {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T : Nat) : (Finset.range T).sum (fun t : Nat => if action t = chosen then (1 : Nat) else 0) = pullCount action chosen T","missing":[],"search":"selectedpullcount_sum_eq_pullcount banditrlproof.ucb.selectedpullcount_sum_eq_pullcount the selected-time nat indicator sum is exactly the recursive pull count. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPullCount_indicator_sum_eq_natCast_pullCount","label":"selectedPullCount_indicator_sum_eq_natCast_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPullCount_indicator_sum_eq_natCast_pullCount","description":"ENNReal-facing selected-time indicator identity for the recursive pull count.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-8ddbed632523","parent":"module:BanditRLProof.Algorithms.UCB","order":2688,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2422"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPullCount_indicator_sum_eq_natCast_pullCount {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T : Nat) : (Finset.range T).sum (fun t : Nat => if action t = chosen then (1 : ENNReal) else 0) = ((pullCount action chosen T : Nat) : ENNReal)","missing":[],"search":"selectedpullcount_indicator_sum_eq_natcast_pullcount banditrlproof.ucb.selectedpullcount_indicator_sum_eq_natcast_pullcount ennreal-facing selected-time indicator identity for the recursive pull count. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPullCount_indicator_sum_eq_selectedSmall_add_selectedLargePullCount","label":"selectedPullCount_indicator_sum_eq_selectedSmall_add_selectedLargePullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPullCount_indicator_sum_eq_selectedSmall_add_selectedLargePullCount","description":"Every selected time is either a selected-small time or a selected-large-count time, split by the threshold `B`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d713cf767f3f","parent":"module:BanditRLProof.Algorithms.UCB","order":2689,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2441"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPullCount_indicator_sum_eq_selectedSmall_add_selectedLargePullCount {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T B : Nat) : (Finset.range T).sum (fun t : Nat => if action t = chosen then (1 : ENNReal) else 0) = (Finset.range T).sum (fun t : Nat => if action t = chosen ∧ pullCount action chosen t < B then (1 : ENNReal) else 0) + (Finset.range T).sum (fun t : Nat => if action t = chosen ∧ B <= pullCount action chosen t then (1 : ENNReal) else 0)","missing":[],"search":"selectedpullcount_indicator_sum_eq_selectedsmall_add_selectedlargepullcount banditrlproof.ucb.selectedpullcount_indicator_sum_eq_selectedsmall_add_selectedlargepullcount every selected time is either a selected-small time or a selected-large-count time, split by the threshold `b`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.natCast_pullCount_le_threshold_add_selectedLargePullCount_indicator_sum","label":"natCast_pullCount_le_threshold_add_selectedLargePullCount_indicator_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.natCast_pullCount_le_threshold_add_selectedLargePullCount_indicator_sum","description":"Pointwise ENNReal UCB count budget after isolating selected-large-count times. The selected-small part is charged by `B`; only selected times whose prior pull count is at least `B` remain for a future large-gap tail bound.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-74a37ce84fa7","parent":"module:BanditRLProof.Algorithms.UCB","order":2690,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2483"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem natCast_pullCount_le_threshold_add_selectedLargePullCount_indicator_sum {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (chosen : Action) (T B : Nat) : ((pullCount action chosen T : Nat) : ENNReal) <= (B : ENNReal) + (Finset.range T).sum (fun t : Nat => if action t = chosen ∧ B <= pullCount action chosen t then (1 : ENNReal) else 0)","missing":[],"search":"natcast_pullcount_le_threshold_add_selectedlargepullcount_indicator_sum banditrlproof.ucb.natcast_pullcount_le_threshold_add_selectedlargepullcount_indicator_sum pointwise ennreal ucb count budget after isolating selected-large-count times. the selected-small part is charged by `b`; only selected times whose prior pull count is at least `b` remain for a future large-gap tail bound. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_selectedLargePullCount","label":"measurableSet_selectedLargePullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_selectedLargePullCount","description":"Measurability of a selected-large-count event for a fixed time.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-9896da46aa63","parent":"module:BanditRLProof.Algorithms.UCB","order":2691,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2533"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selectedLargePullCount {Omega Action : Type} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] [DecidableEq Action] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (chosen : Action) (t B : Nat) : MeasurableSet {omega : Omega | action omega t = chosen ∧ B <= pullCount (action omega) chosen t}","missing":[],"search":"measurableset_selectedlargepullcount banditrlproof.ucb.measurableset_selectedlargepullcount measurability of a selected-large-count event for a fixed time. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_selectedLargePullCount_indicator_sum_eq_sum_measure","label":"lintegral_selectedLargePullCount_indicator_sum_eq_sum_measure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_selectedLargePullCount_indicator_sum_eq_sum_measure","description":"The lower integral of selected-large-count indicators is the corresponding finite sum of event measures.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-f34173288846","parent":"module:BanditRLProof.Algorithms.UCB","order":2692,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2555"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_selectedLargePullCount_indicator_sum_eq_sum_measure {Omega Action : Type} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] [DecidableEq Action] (mu : Measure Omega) (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (chosen : Action) (T B : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => (Finset.range T).sum (fun t : Nat => if action omega t = chosen ∧ B <= pullCount (action omega) chosen t then (1 : ENNReal) else 0)) = (Finset.range T).sum (fun t : Nat => mu {omega : Omega | action omega t = chosen ∧ B <= pullCount (action omega) chosen t})","missing":[],"search":"lintegral_selectedlargepullcount_indicator_sum_eq_sum_measure banditrlproof.ucb.lintegral_selectedlargepullcount_indicator_sum_eq_sum_measure the lower integral of selected-large-count indicators is the corresponding finite sum of event measures. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_subGaussian_textbookDeltaRadius_delta","label":"measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_subGaussian_textbookDeltaRadius_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_subGaussian_textbookDeltaRadius_delta","description":"Single-time selected-large-count event budget for concrete score-argmax UCB. If the event is nonempty but the deterministic large-gap inequality fails, the pointwise large-count-to-large-gap contract gives a contradiction. Otherwise it reduces to the existing selected-event `delta` bound.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-f530a1ac3857","parent":"module:BanditRLProof.Algorithms.UCB","order":2693,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2667"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_subGaussian_textbookDeltaRadius_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (t B : Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (ht : t < T) (hlarge_count_gap : forall omega : Omega, confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t = chosen -> B <= pullCount ((confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta)) omega) chosen t -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T ->…","missing":[],"search":"measure_confidencescoreargmax_selectedlargepullcountevent_le_subgaussian_textbookdeltaradius_delta banditrlproof.ucb.measure_confidencescoreargmax_selectedlargepullcountevent_le_subgaussian_textbookdeltaradius_delta single-time selected-large-count event budget for concrete score-argmax ucb. if the event is nonempty but the deterministic large-gap inequality fails, the pointwise large-count-to-large-gap contract gives a contradiction. otherwise it reduces to the existing selected-event `delta` bound. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sum_measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_horizon_mul_delta","label":"sum_measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_horizon_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sum_measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_horizon_mul_delta","description":"Finite-horizon sum budget for selected-large-count concrete score-argmax events.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-be963beeee67","parent":"module:BanditRLProof.Algorithms.UCB","order":2694,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2739"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sum_measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_horizon_mul_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (hlarge_count_gap : forall omega t, t < T -> confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t = chosen -> B <= pullCount ((confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta)) omega) chosen t -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgau…","missing":[],"search":"sum_measure_confidencescoreargmax_selectedlargepullcountevent_le_horizon_mul_delta banditrlproof.ucb.sum_measure_confidencescoreargmax_selectedlargepullcountevent_le_horizon_mul_delta finite-horizon sum budget for selected-large-count concrete score-argmax events. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_selectedLargePullCount_indicator_sum_le_horizon_mul_delta","label":"lintegral_confidenceScoreArgmax_selectedLargePullCount_indicator_sum_le_horizon_mul_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_selectedLargePullCount_indicator_sum_le_horizon_mul_delta","description":"Lower-integral finite-sum budget for selected-large-count concrete score-argmax events.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-4467d5d82277","parent":"module:BanditRLProof.Algorithms.UCB","order":2695,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2807"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_selectedLargePullCount_indicator_sum_le_horizon_mul_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hlarge_count_gap : forall omega t, t < T -> confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t = chosen -> B <= pullCount ((confidenceScoreArgmaxAction hK…","missing":[],"search":"lintegral_confidencescoreargmax_selectedlargepullcount_indicator_sum_le_horizon_mul_delta banditrlproof.ucb.lintegral_confidencescoreargmax_selectedlargepullcount_indicator_sum_le_horizon_mul_delta lower-integral finite-sum budget for selected-large-count concrete score-argmax events. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_threshold_add_horizon_delta_of_selectedLargePullCount","label":"lintegral_confidenceScoreArgmax_pullCount_le_threshold_add_horizon_delta_of_selectedLargePullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_threshold_add_horizon_delta_of_selectedLargePullCount","description":"Integrated UCB pull-count budget from a selected-large-count large-gap source.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-e50a6599ecb9","parent":"module:BanditRLProof.Algorithms.UCB","order":2696,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2870"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_threshold_add_horizon_delta_of_selectedLargePullCount {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hlarge_count_gap : forall omega t, t < T -> confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t = chosen -> B <= pullCount ((con…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_threshold_add_horizon_delta_of_selectedlargepullcount banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_threshold_add_horizon_delta_of_selectedlargepullcount integrated ucb pull-count budget from a selected-large-count large-gap source. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_freeCard_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_freeCard_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_freeCard_add_horizon_delta","description":"Concrete cardinality-budget version of the threshold/suffix pull-count split. This discharges the abstract `freeBudget` input with `freeTimes.card`. A later radius-threshold leaf can instantiate `freeTimes` and prove its cardinality is the usual logarithmic/gap-dependent budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-e0e2efcffe06","parent":"module:BanditRLProof.Algorithms.UCB","order":2697,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:2960"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_freeCard_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (freeTimes chargedTimes : Finset Nat) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (hcharged_of_not_free : forall t, t < T -> t ∉ freeTimes -> t ∈ chargedTimes) (hgap_large : forall t, t ∈ chargedTimes -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_freecard_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_freecard_add_horizon_delta concrete cardinality-budget version of the threshold/suffix pull-count split. this discharges the abstract `freebudget` input with `freetimes.card`. a later radius-threshold leaf can instantiate `freetimes` and prove its cardinality is the usual logarithmic/gap-dependent budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes","label":"subGaussianTextbookDeltaRadiusChargedTimes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes","description":"Horizon times where the textbook delta radius is already small enough for the selected arm to satisfy the large-gap condition.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-92d9f274b0f8","parent":"module:BanditRLProof.Algorithms.UCB","order":2698,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3003"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianTextbookDeltaRadiusChargedTimes {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) : Finset Nat","missing":[],"search":"subgaussiantextbookdeltaradiuschargedtimes banditrlproof.ucb.subgaussiantextbookdeltaradiuschargedtimes horizon times where the textbook delta radius is already small enough for the selected arm to satisfy the large-gap condition. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes","label":"subGaussianTextbookDeltaRadiusFreeTimes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes","description":"Horizon times not yet discharged by the textbook large-gap radius condition. The next cardinality leaf can bound this concrete set by a closed-form gap/log/sample threshold.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-9188e8873df0","parent":"module:BanditRLProof.Algorithms.UCB","order":2699,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3018"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianTextbookDeltaRadiusFreeTimes {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) : Finset Nat","missing":[],"search":"subgaussiantextbookdeltaradiusfreetimes banditrlproof.ucb.subgaussiantextbookdeltaradiusfreetimes horizon times not yet discharged by the textbook large-gap radius condition. the next cardinality leaf can bound this concrete set by a closed-form gap/log/sample threshold. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_subGaussianTextbookDeltaRadiusChargedTimes_iff","label":"mem_subGaussianTextbookDeltaRadiusChargedTimes_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_subGaussianTextbookDeltaRadiusChargedTimes_iff","description":"@[simp] theorem mem_subGaussianTextbookDeltaRadiusChargedTimes_iff {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t ∈ subGaussianTextbookDeltaRadiusChargedTimes trueMean proxy T delta best chosen ↔ t < T ∧ 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-080c7cb4f404","parent":"module:BanditRLProof.Algorithms.UCB","order":2700,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3027"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem mem_subGaussianTextbookDeltaRadiusChargedTimes_iff {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t ∈ subGaussianTextbookDeltaRadiusChargedTimes trueMean proxy T delta best chosen ↔ t < T ∧ 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","missing":[],"search":"mem_subgaussiantextbookdeltaradiuschargedtimes_iff banditrlproof.ucb.mem_subgaussiantextbookdeltaradiuschargedtimes_iff @[simp] theorem mem_subgaussiantextbookdeltaradiuschargedtimes_iff {k : nat} (truemean : fin k -> real) (proxy : nat -> fin k -> nnreal) (t : nat) (delta : real) (best chosen : fin k) (t : nat) : t ∈ subgaussiantextbookdeltaradiuschargedtimes truemean proxy t delta best chosen ↔ t < t ∧ 2 * subgaussiantextbookdeltaradius proxy t delta t chosen < meangap truemean best chosen theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_subGaussianTextbookDeltaRadiusFreeTimes_iff","label":"mem_subGaussianTextbookDeltaRadiusFreeTimes_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_subGaussianTextbookDeltaRadiusFreeTimes_iff","description":"@[simp] theorem mem_subGaussianTextbookDeltaRadiusFreeTimes_iff {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t ∈ subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen ↔ t < T ∧ ¬ 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-ed2463dc4021","parent":"module:BanditRLProof.Algorithms.UCB","order":2701,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3038"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"@[simp] theorem mem_subGaussianTextbookDeltaRadiusFreeTimes_iff {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t ∈ subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen ↔ t < T ∧ ¬ 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","missing":[],"search":"mem_subgaussiantextbookdeltaradiusfreetimes_iff banditrlproof.ucb.mem_subgaussiantextbookdeltaradiusfreetimes_iff @[simp] theorem mem_subgaussiantextbookdeltaradiusfreetimes_iff {k : nat} (truemean : fin k -> real) (proxy : nat -> fin k -> nnreal) (t : nat) (delta : real) (best chosen : fin k) (t : nat) : t ∈ subgaussiantextbookdeltaradiusfreetimes truemean proxy t delta best chosen ↔ t < t ∧ ¬ 2 * subgaussiantextbookdeltaradius proxy t delta t chosen < meangap truemean best chosen theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes_of_not_free","label":"subGaussianTextbookDeltaRadiusChargedTimes_of_not_free","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes_of_not_free","description":"theorem subGaussianTextbookDeltaRadiusChargedTimes_of_not_free {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t < T -> t ∉ subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen -> t ∈ subGaussianTextbookDeltaRadiusChargedTimes trueMean proxy T delta best chosen","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-242d69d189bc","parent":"module:BanditRLProof.Algorithms.UCB","order":2702,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3049"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadiusChargedTimes_of_not_free {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t < T -> t ∉ subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen -> t ∈ subGaussianTextbookDeltaRadiusChargedTimes trueMean proxy T delta best chosen","missing":[],"search":"subgaussiantextbookdeltaradiuschargedtimes_of_not_free banditrlproof.ucb.subgaussiantextbookdeltaradiuschargedtimes_of_not_free theorem subgaussiantextbookdeltaradiuschargedtimes_of_not_free {k : nat} (truemean : fin k -> real) (proxy : nat -> fin k -> nnreal) (t : nat) (delta : real) (best chosen : fin k) (t : nat) : t < t -> t ∉ subgaussiantextbookdeltaradiusfreetimes truemean proxy t delta best chosen -> t ∈ subgaussiantextbookdeltaradiuschargedtimes truemean proxy t delta best chosen theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes_gap_large","label":"subGaussianTextbookDeltaRadiusChargedTimes_gap_large","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes_gap_large","description":"theorem subGaussianTextbookDeltaRadiusChargedTimes_gap_large {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t ∈ subGaussianTextbookDeltaRadiusChargedTimes trueMean proxy T delta best chosen -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-551d6796eb5a","parent":"module:BanditRLProof.Algorithms.UCB","order":2703,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3066"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadiusChargedTimes_gap_large {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) : t ∈ subGaussianTextbookDeltaRadiusChargedTimes trueMean proxy T delta best chosen -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","missing":[],"search":"subgaussiantextbookdeltaradiuschargedtimes_gap_large banditrlproof.ucb.subgaussiantextbookdeltaradiuschargedtimes_gap_large theorem subgaussiantextbookdeltaradiuschargedtimes_gap_large {k : nat} (truemean : fin k -> real) (proxy : nat -> fin k -> nnreal) (t : nat) (delta : real) (best chosen : fin k) (t : nat) : t ∈ subgaussiantextbookdeltaradiuschargedtimes truemean proxy t delta best chosen -> 2 * subgaussiantextbookdeltaradius proxy t delta t chosen < meangap truemean best chosen theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusFreeCard_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusFreeCard_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusFreeCard_add_horizon_delta","description":"Concrete radius-threshold split for the textbook delta UCB pull-count budget. This instantiates the abstract `freeTimes`/`chargedTimes` split with the large-gap predicate induced by `subGaussianTextbookDeltaRadius`. It leaves the closed-form cardinality bound for the concrete free-time set to the next leaf.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-df0df40c0c3c","parent":"module:BanditRLProof.Algorithms.UCB","order":2704,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3086"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusFreeCard_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (hT : 0 < T) (hdelta : 0 < delta) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : MeasureTheory.lintegral mu (fun omega : Ome…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusfreecard_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusfreecard_add_horizon_delta concrete radius-threshold split for the textbook delta ucb pull-count budget. this instantiates the abstract `freetimes`/`chargedtimes` split with the large-gap predicate induced by `subgaussiantextbookdeltaradius`. it leaves the closed-form cardinality bound for the concrete free-time set to the next leaf. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold","label":"subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold","description":"If every horizon time at or beyond threshold `B` satisfies the textbook large-gap radius condition, then the concrete free-time set has cardinality at most `B`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-772aadd2df4f","parent":"module:BanditRLProof.Algorithms.UCB","order":2705,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3141"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hlarge_after : forall t, t < T -> B <= t -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) : (subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen).card <= B","missing":[],"search":"subgaussiantextbookdeltaradiusfreetimes_card_le_threshold banditrlproof.ucb.subgaussiantextbookdeltaradiusfreetimes_card_le_threshold if every horizon time at or beyond threshold `b` satisfies the textbook large-gap radius condition, then the concrete free-time set has cardinality at most `b`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold_ennreal","label":"subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold_ennreal","description":"theorem subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold_ennreal {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hlarge_after : forall t, t < T -> B <= t -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) : ((subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen).car…","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-56e760930c7a","parent":"module:BanditRLProof.Algorithms.UCB","order":2706,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3168"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold_ennreal {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hlarge_after : forall t, t < T -> B <= t -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) : ((subGaussianTextbookDeltaRadiusFreeTimes trueMean proxy T delta best chosen).card : ENNReal) <= (B : ENNReal)","missing":[],"search":"subgaussiantextbookdeltaradiusfreetimes_card_le_threshold_ennreal banditrlproof.ucb.subgaussiantextbookdeltaradiusfreetimes_card_le_threshold_ennreal theorem subgaussiantextbookdeltaradiusfreetimes_card_le_threshold_ennreal {k : nat} (truemean : fin k -> real) (proxy : nat -> fin k -> nnreal) (t : nat) (delta : real) (best chosen : fin k) (b : nat) (hlarge_after : forall t, t < t -> b <= t -> 2 * subgaussiantextbookdeltaradius proxy t delta t chosen < meangap truemean best chosen) : ((subgaussiantextbookdeltaradiusfreetimes truemean proxy t delta best chosen).card : ennreal) <= (b : ennreal) theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusThreshold_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusThreshold_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusThreshold_add_horizon_delta","description":"Threshold-budget version of the concrete textbook-radius UCB pull-count split. The only new deterministic input is that all times `t >= B` in the horizon satisfy the large-gap radius condition. A later leaf can instantiate `B` with a closed-form logarithmic/gap-dependent expression.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-a1dfaa02fc4b","parent":"module:BanditRLProof.Algorithms.UCB","order":2707,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3189"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusThreshold_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hT : 0 < T) (hdelta : 0 < delta) (hlarge_after : forall t, t < T -> B <= t -> 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> ProbabilityTheory…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusthreshold_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusthreshold_add_horizon_delta threshold-budget version of the concrete textbook-radius ucb pull-count split. the only new deterministic input is that all times `t >= b` in the horizon satisfy the large-gap radius condition. a later leaf can instantiate `b` with a closed-form logarithmic/gap-dependent expression. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_large_gap_of_lt_half_meanGap","label":"subGaussianTextbookDeltaRadius_large_gap_of_lt_half_meanGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_large_gap_of_lt_half_meanGap","description":"Half-gap radius condition in the textbook form implies the large-gap condition consumed by the UCB selected-event and pull-count budget wrappers.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-77359224693c","parent":"module:BanditRLProof.Algorithms.UCB","order":2708,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3236"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadius_large_gap_of_lt_half_meanGap {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) (hhalf : subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen / 2) : 2 * subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen","missing":[],"search":"subgaussiantextbookdeltaradius_large_gap_of_lt_half_meangap banditrlproof.ucb.subgaussiantextbookdeltaradius_large_gap_of_lt_half_meangap half-gap radius condition in the textbook form implies the large-gap condition consumed by the ucb selected-event and pull-count budget wrappers. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusHalfGapThreshold_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusHalfGapThreshold_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusHalfGapThreshold_add_horizon_delta","description":"Half-gap threshold version of the concrete textbook-radius UCB pull-count budget. This is the surface normally targeted by the remaining logarithmic/gap algebra: prove that after threshold `B`, the textbook radius is below half the chosen arm's gap.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-9bb5f8eeb3e2","parent":"module:BanditRLProof.Algorithms.UCB","order":2709,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3255"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusHalfGapThreshold_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hT : 0 < T) (hdelta : 0 < delta) (hhalf_after : forall t, t < T -> B <= t -> subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen / 2) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t arm, t < T -> 0 < ((proxy t arm : NNReal) : Real)) (hsubG : forall t arm, t < T -> Probability…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiushalfgapthreshold_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiushalfgapthreshold_add_horizon_delta half-gap threshold version of the concrete textbook-radius ucb pull-count budget. this is the surface normally targeted by the remaining logarithmic/gap algebra: prove that after threshold `b`, the textbook radius is below half the chosen arm's gap. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_sq_lt","label":"subGaussianTextbookDeltaRadius_lt_half_meanGap_of_sq_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_sq_lt","description":"Square-form deterministic algebra for the textbook delta radius: if the quantity under the square root is below `(gap / 2)^2`, the radius is below half the gap.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-179938c7b72a","parent":"module:BanditRLProof.Algorithms.UCB","order":2710,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3301"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadius_lt_half_meanGap_of_sq_lt {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) (hgap_pos : 0 < meanGap trueMean best chosen) (hsq : 2 * ((proxy t chosen : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) < (meanGap trueMean best chosen / 2) ^ 2) : subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen / 2","missing":[],"search":"subgaussiantextbookdeltaradius_lt_half_meangap_of_sq_lt banditrlproof.ucb.subgaussiantextbookdeltaradius_lt_half_meangap_of_sq_lt square-form deterministic algebra for the textbook delta radius: if the quantity under the square root is below `(gap / 2)^2`, the radius is below half the gap. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_eight_mul_lt_sq","label":"subGaussianTextbookDeltaRadius_lt_half_meanGap_of_eight_mul_lt_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_eight_mul_lt_sq","description":"Common UCB algebra form for the textbook delta radius: the sufficient condition `8 * proxy * log(scale) < gap^2` implies `radius < gap / 2`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-8d65a4cd69fc","parent":"module:BanditRLProof.Algorithms.UCB","order":2711,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3321"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadius_lt_half_meanGap_of_eight_mul_lt_sq {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) (hgap_pos : 0 < meanGap trueMean best chosen) (height : 8 * ((proxy t chosen : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) < (meanGap trueMean best chosen) ^ 2) : subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen / 2","missing":[],"search":"subgaussiantextbookdeltaradius_lt_half_meangap_of_eight_mul_lt_sq banditrlproof.ucb.subgaussiantextbookdeltaradius_lt_half_meangap_of_eight_mul_lt_sq common ucb algebra form for the textbook delta radius: the sufficient condition `8 * proxy * log(scale) < gap^2` implies `radius < gap / 2`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusEightProxyLogThreshold_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusEightProxyLogThreshold_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusEightProxyLogThreshold_add_horizon_delta","description":"Eight-proxy-log threshold version of the concrete textbook-radius UCB pull-count budget. The remaining closed-form work is to prove the displayed square inequality from a concrete choice of `B` and whatever sample-count/proxy monotonicity the eventual empirical-mean construction provides.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-852178213fee","parent":"module:BanditRLProof.Algorithms.UCB","order":2712,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3350"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusEightProxyLogThreshold_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (height_after : forall t, t < T -> B <= t -> 8 * ((proxy t chosen : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) < (meanGap trueMean best chosen) ^ 2) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subGaussianTextbookDeltaRadius proxy T delta) omega t)) (hproxy : forall t…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiuseightproxylogthreshold_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiuseightproxylogthreshold_add_horizon_delta eight-proxy-log threshold version of the concrete textbook-radius ucb pull-count budget. the remaining closed-form work is to prove the displayed square inequality from a concrete choice of `b` and whatever sample-count/proxy monotonicity the eventual empirical-mean construction provides. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_lt_gap_sq_div","label":"subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_lt_gap_sq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_lt_gap_sq_div","description":"Proxy-small form of the textbook radius half-gap algebra. Under a positive logarithmic scale, bounding the selected arm's proxy by `gap^2 / (8 * log scale)` implies the usual eight-proxy-log condition and hence `radius < gap / 2`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-ed2c8ad2b79e","parent":"module:BanditRLProof.Algorithms.UCB","order":2713,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3400"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_lt_gap_sq_div {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hproxy_small : ((proxy t chosen : NNReal) : Real) < (meanGap trueMean best chosen) ^ 2 / (8 * Real.log (textbookDeltaScale (Arm := Fin K) T delta))) : subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen / 2","missing":[],"search":"subgaussiantextbookdeltaradius_lt_half_meangap_of_proxy_lt_gap_sq_div banditrlproof.ucb.subgaussiantextbookdeltaradius_lt_half_meangap_of_proxy_lt_gap_sq_div proxy-small form of the textbook radius half-gap algebra. under a positive logarithmic scale, bounding the selected arm's proxy by `gap^2 / (8 * log scale)` implies the usual eight-proxy-log condition and hence `radius < gap / 2`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusProxyThreshold_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusProxyThreshold_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusProxyThreshold_add_horizon_delta","description":"Proxy-small threshold version of the concrete textbook-radius UCB pull-count budget. This is the handoff expected from a later empirical-mean/sample-count leaf: after threshold `B`, prove the selected arm's sub-Gaussian proxy is below the displayed gap/log scale.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-831c988c39c7","parent":"module:BanditRLProof.Algorithms.UCB","order":2714,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3446"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusProxyThreshold_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hproxy_small_after : forall t, t < T -> B <= t -> ((proxy t chosen : NNReal) : Real) < (meanGap trueMean best chosen) ^ 2 / (8 * Real.log (textbookDeltaScale (Arm := Fin K) T delta))) (haction : forall t : Nat, Measurable (fun omega : Omega => confidenceScoreArgmaxAction hK empiricalMean (subG…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusproxythreshold_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusproxythreshold_add_horizon_delta proxy-small threshold version of the concrete textbook-radius ucb pull-count budget. this is the handoff expected from a later empirical-mean/sample-count leaf: after threshold `b`, prove the selected arm's sub-gaussian proxy is below the displayed gap/log scale. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_le_variance_div_count","label":"subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_le_variance_div_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_le_variance_div_count","description":"Sample-count proxy form of the textbook radius half-gap algebra. If the selected arm's proxy is bounded by `varianceProxy / count`, then a count threshold of the form `8 * varianceProxy * log(scale) < gap^2 * count` implies the proxy-small condition and hence `radius < gap / 2`.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-bcde9144be95","parent":"module:BanditRLProof.Algorithms.UCB","order":2715,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3498"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_le_variance_div_count {K : Nat} (trueMean : Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (t : Nat) (varianceProxy : NNReal) (count : Nat) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hcount_pos : 0 < count) (hproxy_le : ((proxy t chosen : NNReal) : Real) <= ((varianceProxy : NNReal) : Real) / (count : Real)) (hcount_large : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) < (meanGap trueMean best chosen) ^ 2 * (count : Real)) : subGaussianTextbookDeltaRadius proxy T delta t chosen < meanGap trueMean best chosen / 2","missing":[],"search":"subgaussiantextbookdeltaradius_lt_half_meangap_of_proxy_le_variance_div_count banditrlproof.ucb.subgaussiantextbookdeltaradius_lt_half_meangap_of_proxy_le_variance_div_count sample-count proxy form of the textbook radius half-gap algebra. if the selected arm's proxy is bounded by `varianceproxy / count`, then a count threshold of the form `8 * varianceproxy * log(scale) < gap^2 * count` implies the proxy-small condition and hence `radius < gap / 2`. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountThreshold_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountThreshold_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountThreshold_add_horizon_delta","description":"Sample-count threshold version of the concrete textbook-radius UCB pull-count budget. This keeps the probabilistic/concentration assumptions abstract, but turns the remaining radius-threshold algebra into explicit count and proxy contracts.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-cc8add36e0aa","parent":"module:BanditRLProof.Algorithms.UCB","order":2716,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3542"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountThreshold_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (varianceProxy : NNReal) (count : Nat -> Fin K -> Nat) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hcount_pos_after : forall t, t < T -> B <= t -> 0 < count t chosen) (hproxy_le_after : forall t, t < T -> B <= t -> ((proxy t chosen : NNReal) : Real) <= ((varianceProxy : NNReal) : Real) / (count t chosen : Real)) (hcount_large_afte…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiussamplecountthreshold_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiussamplecountthreshold_add_horizon_delta sample-count threshold version of the concrete textbook-radius ucb pull-count budget. this keeps the probabilistic/concentration assumptions abstract, but turns the remaining radius-threshold algebra into explicit count and proxy contracts. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_count_large_of_threshold_lt_bound","label":"subGaussianTextbookDeltaRadius_count_large_of_threshold_lt_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianTextbookDeltaRadius_count_large_of_threshold_lt_bound","description":"Closed threshold-to-count algebra: if the real threshold `8 * varianceProxy * log(scale) / gap^2` is below `B`, and `B <= count`, then the count is large enough for the sample-count UCB radius condition.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-e10a052e5503","parent":"module:BanditRLProof.Algorithms.UCB","order":2717,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3602"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem subGaussianTextbookDeltaRadius_count_large_of_threshold_lt_bound {K : Nat} (trueMean : Fin K -> Real) (T : Nat) (delta : Real) (best chosen : Fin K) (varianceProxy : NNReal) (B count : Nat) (hgap_pos : 0 < meanGap trueMean best chosen) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) / (meanGap trueMean best chosen) ^ 2 < (B : Real)) (hB_le_count : B <= count) : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) < (meanGap trueMean best chosen) ^ 2 * (count : Real)","missing":[],"search":"subgaussiantextbookdeltaradius_count_large_of_threshold_lt_bound banditrlproof.ucb.subgaussiantextbookdeltaradius_count_large_of_threshold_lt_bound closed threshold-to-count algebra: if the real threshold `8 * varianceproxy * log(scale) / gap^2` is below `b`, and `b <= count`, then the count is large enough for the sample-count ucb radius condition. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountLowerBound_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountLowerBound_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountLowerBound_add_horizon_delta","description":"Lower-bound-on-count version of the concrete textbook-radius UCB pull-count budget. A later adaptive trace leaf can aim to prove `B <= count t chosen` after the same threshold `B`; this wrapper then supplies the usual `B + T * delta` pull-count budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-b30222696cec","parent":"module:BanditRLProof.Algorithms.UCB","order":2718,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3637"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountLowerBound_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (best chosen : Fin K) (B : Nat) (varianceProxy : NNReal) (count : Nat -> Fin K -> Nat) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hB_pos : 0 < B) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) / (meanGap trueMean best chosen) ^ 2 < (B : Real)) (hcount_lower_after : forall t, t < T -> B <= t -> B…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiussamplecountlowerbound_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiussamplecountlowerbound_add_horizon_delta lower-bound-on-count version of the concrete textbook-radius ucb pull-count budget. a later adaptive trace leaf can aim to prove `b <= count t chosen` after the same threshold `b`; this wrapper then supplies the usual `b + t * delta` pull-count budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusRecursiveSampleCount_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusRecursiveSampleCount_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusRecursiveSampleCount_add_horizon_delta","description":"Recursive sample-count adapter for the selected-large-count UCB budget. For selected times whose previous recursive pull count is at least `B`, a variance-over-count proxy bound plus the closed real threshold certificate implies the textbook radius is below half the gap. The selected-large-count wrapper then yields the usual `B + T * delta` pull-count budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-4f8e4f4c5e2d","parent":"module:BanditRLProof.Algorithms.UCB","order":2719,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3704"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusRecursiveSampleCount_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (varianceProxy : NNReal) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hB_pos : 0 < B) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) / (meanGap trueMean best chosen) ^ 2 < (B : Real)) (haction : for…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusrecursivesamplecount_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiusrecursivesamplecount_add_horizon_delta recursive sample-count adapter for the selected-large-count ucb budget. for selected times whose previous recursive pull count is at least `b`, a variance-over-count proxy bound plus the closed real threshold certificate implies the textbook radius is below half the gap. the selected-large-count wrapper then yields the usual `b + t * delta` pull-count budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","label":"lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","description":"Source-count version of the recursive sample-count UCB budget. This wrapper is meant for later empirical-mean leaves: they can expose their own history-derived `sampleCount`, prove it agrees with recursive `pullCount` on selected-large events, and provide the usual variance-over-count proxy bound for that source count. The existing recursive sample-count adapter then gives the same `B + T * delta` pull-count budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d6a5a9cdd335","parent":"module:BanditRLProof.Algorithms.UCB","order":2720,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3801"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (varianceProxy : NNReal) (sampleCount : Omega -> Nat -> Fin K -> Nat) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hB_pos : 0 < B) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) / (meanGap trueMean bes…","missing":[],"search":"lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmax_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta source-count version of the recursive sample-count ucb budget. this wrapper is meant for later empirical-mean leaves: they can expose their own history-derived `samplecount`, prove it agrees with recursive `pullcount` on selected-large events, and provide the usual variance-over-count proxy bound for that source count. the existing recursive sample-count adapter then gives the same `b + t * delta` pull-count budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_historyAction_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","label":"lintegral_historyAction_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_historyAction_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","description":"History-action source-count version of the textbook-radius UCB pull-count budget. If an externally generated history trace agrees with the concrete score-argmax UCB trace throughout the horizon, and its own recursive pull count supplies the variance-over-count proxy contract on selected-large events, then that history trace inherits the same `B + T * delta` selected-arm count budget.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-0f9310a6dc36","parent":"module:BanditRLProof.Algorithms.UCB","order":2721,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:3888"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_historyAction_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (varianceProxy : NNReal) (historyAction : Omega -> ActionTrace (Fin K)) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hB_pos : 0 < B) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) / (meanGap trueMean best chos…","missing":[],"search":"lintegral_historyaction_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta banditrlproof.ucb.lintegral_historyaction_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta history-action source-count version of the textbook-radius ucb pull-count budget. if an externally generated history trace agrees with the concrete score-argmax ucb trace throughout the horizon, and its own recursive pull count supplies the variance-over-count proxy contract on selected-large events, then that history trace inherits the same `b + t * delta` selected-arm count budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_generatedActionTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","label":"lintegral_generatedActionTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_generatedActionTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","description":"Generated-policy source-count version of the textbook-radius UCB pull-count budget. This packages the previous history-action wrapper for a concrete `Policy.generatedActionTrace`. Pointwise equality with score argmax over all time coordinates transfers measurability from the generated policy trace to the score-argmax trace, and the existing history-action adapter supplies the `B + T * delta` selected-arm count budge…","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-edcb805d69e2","parent":"module:BanditRLProof.Algorithms.UCB","order":2722,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4009"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_generatedActionTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta {Omega State : Type} [MeasurableSpace Omega] [MeasurableSpace State] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (varianceProxy : NNReal) (policy : Policy.MeasurablePolicy State (Fin K)) (state : Nat -> Omega -> State) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hB_pos : 0 < B) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (t…","missing":[],"search":"lintegral_generatedactiontrace_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta banditrlproof.ucb.lintegral_generatedactiontrace_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta generated-policy source-count version of the textbook-radius ucb pull-count budget. this packages the previous history-action wrapper for a concrete `policy.generatedactiontrace`. pointwise equality with score argmax over all time coordinates transfers measurability from the generated policy trace to the score-argmax trace, and the existing history-action adapter supplies the `b + t * delta` selected-arm count budget. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.identityActionPolicy","label":"identityActionPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.identityActionPolicy","description":"Identity measurable policy on an action space.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-927d369d3341","parent":"module:BanditRLProof.Algorithms.UCB","order":2723,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4086"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def identityActionPolicy (Action : Type) [MeasurableSpace Action] : Policy.MeasurablePolicy Action Action where","missing":[],"search":"identityactionpolicy banditrlproof.ucb.identityactionpolicy identity measurable policy on an action space. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxGeneratedState","label":"confidenceScoreArgmaxGeneratedState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScoreArgmaxGeneratedState","description":"State process whose value is the current concrete UCB score-argmax action. This is a thin policy/state adapter: a generated trace using `identityActionPolicy` over this state is definitionally the concrete score-argmax action trace.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-e9faa569cc58","parent":"module:BanditRLProof.Algorithms.UCB","order":2724,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4099"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceScoreArgmaxGeneratedState {Omega : Type} {K : Nat} (hK : 0 < K) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) : Nat -> Omega -> Fin K","missing":[],"search":"confidencescoreargmaxgeneratedstate banditrlproof.ucb.confidencescoreargmaxgeneratedstate state process whose value is the current concrete ucb score-argmax action. this is a thin policy/state adapter: a generated trace using `identityactionpolicy` over this state is definitionally the concrete score-argmax action trace. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxGeneratedTrace","label":"confidenceScoreArgmaxGeneratedTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.confidenceScoreArgmaxGeneratedTrace","description":"Concrete UCB score-argmax action trace expressed as a generated policy trace.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-5435128e0c18","parent":"module:BanditRLProof.Algorithms.UCB","order":2725,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4109"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceScoreArgmaxGeneratedTrace {Omega : Type} {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) : Omega -> ActionTrace (Fin K)","missing":[],"search":"confidencescoreargmaxgeneratedtrace banditrlproof.ucb.confidencescoreargmaxgeneratedtrace concrete ucb score-argmax action trace expressed as a generated policy trace. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmaxGeneratedTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","label":"lintegral_confidenceScoreArgmaxGeneratedTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_confidenceScoreArgmaxGeneratedTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","description":"Generated-trace instantiation of the textbook-radius UCB pull-count budget. The generated trace is built from the identity action policy and the concrete score-argmax state, so the generated-policy equality contract from the previous wrapper is discharged definitionally.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-8c7cd47fa2a6","parent":"module:BanditRLProof.Algorithms.UCB","order":2726,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4126"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_confidenceScoreArgmaxGeneratedTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [OpensMeasurableSpace Nat] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (trueMean : Fin K -> Real) (empiricalMean : Omega -> Nat -> Fin K -> Real) (proxy : Nat -> Fin K -> NNReal) (T : Nat) (delta : Real) (B : Nat) (best chosen : Fin K) (varianceProxy : NNReal) (hT : 0 < T) (hdelta : 0 < delta) (hgap_pos : 0 < meanGap trueMean best chosen) (hlog_pos : 0 < Real.log (textbookDeltaScale (Arm := Fin K) T delta)) (hB_pos : 0 < B) (hthreshold_lt_B : 8 * ((varianceProxy : NNReal) : Real) * Real.log (textbookDeltaScale (Arm := Fin K) T delta) / (meanGap trueMean best chosen) ^ 2 < (B : Real)) (ha…","missing":[],"search":"lintegral_confidencescoreargmaxgeneratedtrace_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta banditrlproof.ucb.lintegral_confidencescoreargmaxgeneratedtrace_pullcount_le_textbookdeltaradiussamplecountsource_add_horizon_delta generated-trace instantiation of the textbook-radius ucb pull-count budget. the generated trace is built from the identity action policy and the concrete score-argmax state, so the generated-policy equality contract from the previous wrapper is discharged definitionally. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.subGaussianAbsDeviationTail","label":"subGaussianAbsDeviationTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.subGaussianAbsDeviationTail","description":"Two-sided sub-Gaussian tail budget for a UCB empirical mean at time `t` and arm `arm`. The proxy is for the centered variable `empiricalMean t arm - trueMean arm`. This is still abstract: a later empirical-mean construction must prove the sub-Gaussian proxy and choose the usual log/sqrt radius.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d9c0ecd12bae","parent":"module:BanditRLProof.Algorithms.UCB","order":2727,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4202"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianAbsDeviationTail {Arm : Type} (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (t : Nat) (arm : Arm) : ENNReal","missing":[],"search":"subgaussianabsdeviationtail banditrlproof.ucb.subgaussianabsdeviationtail two-sided sub-gaussian tail budget for a ucb empirical mean at time `t` and arm `arm`. the proxy is for the centered variable `empiricalmean t arm - truemean arm`. this is still abstract: a later empirical-mean construction must prove the sub-gaussian proxy and choose the usual log/sqrt radius. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_absDeviation_le_subGaussian_tail","label":"measure_absDeviation_le_subGaussian_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_absDeviation_le_subGaussian_tail","description":"Single-time two-sided sub-Gaussian tail for the UCB absolute-deviation event.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-d5bca77d26a6","parent":"module:BanditRLProof.Algorithms.UCB","order":2728,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4214"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_absDeviation_le_subGaussian_tail {Omega Arm : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (t : Nat) (arm : Arm) (hradius : 0 <= radius t arm) (hsubG : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu {omega | radius t arm <= |empiricalMean omega t arm - trueMean arm|} <= subGaussianAbsDeviationTail radius proxy t arm","missing":[],"search":"measure_absdeviation_le_subgaussian_tail banditrlproof.ucb.measure_absdeviation_le_subgaussian_tail single-time two-sided sub-gaussian tail for the ucb absolute-deviation event. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_tail_sum","label":"measure_finiteHorizonConfidenceBadEvent_le_subGaussian_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_tail_sum","description":"Finite-horizon UCB confidence bad-event bound from abstract sub-Gaussian absolute-deviation tails. This is the UCB-facing sub-Gaussian producer layer. It still leaves empirical mean construction, proxy simplification, log/sqrt radius choice, pull-count bounds, and final regret to later leaves.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-ae3075d39e5e","parent":"module:BanditRLProof.Algorithms.UCB","order":2729,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4298"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonConfidenceBadEvent_le_subGaussian_tail_sum {Omega Arm : Type} [MeasurableSpace Omega] [Fintype Arm] (mu : Measure Omega) [IsFiniteMeasure mu] (trueMean : Arm -> Real) (empiricalMean : Omega -> Nat -> Arm -> Real) (radius : Nat -> Arm -> Real) (proxy : Nat -> Arm -> NNReal) (T : Nat) (hradius : forall t arm, t < T -> 0 <= radius t arm) (hsubG : forall t arm, t < T -> ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => empiricalMean omega t arm - trueMean arm) (proxy t arm) mu) : mu (finiteHorizonConfidenceBadEvent trueMean empiricalMean radius T) <= (Finset.range T).sum (fun t => (Finset.univ : Finset Arm).sum (fun arm => subGaussianAbsDeviationTail radius proxy t arm + subGaussianAbsDeviationTail radius proxy t arm))","missing":[],"search":"measure_finitehorizonconfidencebadevent_le_subgaussian_tail_sum banditrlproof.ucb.measure_finitehorizonconfidencebadevent_le_subgaussian_tail_sum finite-horizon ucb confidence bad-event bound from abstract sub-gaussian absolute-deviation tails. this is the ucb-facing sub-gaussian producer layer. it still leaves empirical mean construction, proxy simplification, log/sqrt radius choice, pull-count bounds, and final regret to later leaves. theorem compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.obligationNames","label":"obligationNames","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.obligationNames","description":"The proof-DAG leaves usually needed for UCB regret formalization.","url":"../modules/banditrlproof-algorithms-ucb/index.html#decl-1e8c0c38728a","parent":"module:BanditRLProof.Algorithms.UCB","order":2730,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCB"],["Source","BanditRLProof/Algorithms/UCB.lean:4327"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def obligationNames : List String","missing":[],"search":"obligationnames banditrlproof.ucb.obligationnames the proof-dag leaves usually needed for ucb regret formalization. definition compiled","shard":"modules/bb0657d69d348d5d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamPSeriesTerm","label":"armStreamPSeriesTerm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamPSeriesTerm","description":"The summable cubic tail that controls `indexTail 4`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-585c7a3aff32","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2731,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:23"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamPSeriesTerm (t : Nat) : NNReal","missing":[],"search":"armstreampseriesterm banditrlproof.ucb.armstreampseriesterm the summable cubic tail that controls `indextail 4`. definition compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamPSeriesTailBound","label":"armStreamPSeriesTailBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamPSeriesTailBound","description":"A fixed finite upper bound for every `constSum 4 n`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-0881fd693d50","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2732,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:27"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamPSeriesTailBound : NNReal","missing":[],"search":"armstreampseriestailbound banditrlproof.ucb.armstreampseriestailbound a fixed finite upper bound for every `constsum 4 n`. definition compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamPSeriesTerm_summable","label":"armStreamPSeriesTerm_summable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamPSeriesTerm_summable","description":"theorem armStreamPSeriesTerm_summable : Summable armStreamPSeriesTerm","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-742f98e31153","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2733,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:30"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamPSeriesTerm_summable : Summable armStreamPSeriesTerm","missing":[],"search":"armstreampseriesterm_summable banditrlproof.ucb.armstreampseriesterm_summable theorem armstreampseriesterm_summable : summable armstreampseriesterm theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.indexTail_four_eq_coe_armStreamPSeriesTerm","label":"indexTail_four_eq_coe_armStreamPSeriesTerm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.indexTail_four_eq_coe_armStreamPSeriesTerm","description":"theorem indexTail_four_eq_coe_armStreamPSeriesTerm (t : Nat) : indexTail 4 t = (armStreamPSeriesTerm t : ENNReal)","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-8f92f1260786","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2734,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:35"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem indexTail_four_eq_coe_armStreamPSeriesTerm (t : Nat) : indexTail 4 t = (armStreamPSeriesTerm t : ENNReal)","missing":[],"search":"indextail_four_eq_coe_armstreampseriesterm banditrlproof.ucb.indextail_four_eq_coe_armstreampseriesterm theorem indextail_four_eq_coe_armstreampseriesterm (t : nat) : indextail 4 t = (armstreampseriesterm t : ennreal) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constSum_four_le_armStreamPSeriesTailBound","label":"constSum_four_le_armStreamPSeriesTailBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.constSum_four_le_armStreamPSeriesTailBound","description":"theorem constSum_four_le_armStreamPSeriesTailBound (n : Nat) : constSum 4 n <= (armStreamPSeriesTailBound : ENNReal)","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-f452ee745ce3","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2735,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:40"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem constSum_four_le_armStreamPSeriesTailBound (n : Nat) : constSum 4 n <= (armStreamPSeriesTailBound : ENNReal)","missing":[],"search":"constsum_four_le_armstreampseriestailbound banditrlproof.ucb.constsum_four_le_armstreampseriestailbound theorem constsum_four_le_armstreampseriestailbound (n : nat) : constsum 4 n <= (armstreampseriestailbound : ennreal) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constSum_four_toReal_le_armStreamPSeriesTailBound","label":"constSum_four_toReal_le_armStreamPSeriesTailBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.constSum_four_toReal_le_armStreamPSeriesTailBound","description":"theorem constSum_four_toReal_le_armStreamPSeriesTailBound (n : Nat) : (constSum 4 n).toReal <= (armStreamPSeriesTailBound : Real)","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-e208d6f976ef","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2736,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:52"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem constSum_four_toReal_le_armStreamPSeriesTailBound (n : Nat) : (constSum 4 n).toReal <= (armStreamPSeriesTailBound : Real)","missing":[],"search":"constsum_four_toreal_le_armstreampseriestailbound banditrlproof.ucb.constsum_four_toreal_le_armstreampseriestailbound theorem constsum_four_toreal_le_armstreampseriestailbound (n : nat) : (constsum 4 n).toreal <= (armstreampseriestailbound : real) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAsymptoticModelCoefficient","label":"armStreamAsymptoticModelCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAsymptoticModelCoefficient","description":"Fixed kernel-dependent coefficient for the one-policy logarithmic envelope.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-30a7d459c0d4","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2737,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:62"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamAsymptoticModelCoefficient {K : Nat} (nu : Kernel (Fin K) Real) (sigma2 : NNReal) : Real","missing":[],"search":"armstreamasymptoticmodelcoefficient banditrlproof.ucb.armstreamasymptoticmodelcoefficient fixed kernel-dependent coefficient for the one-policy logarithmic envelope. definition compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAsymptoticModelCoefficient_nonneg","label":"armStreamAsymptoticModelCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAsymptoticModelCoefficient_nonneg","description":"theorem armStreamAsymptoticModelCoefficient_nonneg {K : Nat} (hK : 0 < K) (nu : Kernel (Fin K) Real) (sigma2 : NNReal) : 0 <= armStreamAsymptoticModelCoefficient nu sigma2","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-e4cb07e5e75d","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2738,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:69"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAsymptoticModelCoefficient_nonneg {K : Nat} (hK : 0 < K) (nu : Kernel (Fin K) Real) (sigma2 : NNReal) : 0 <= armStreamAsymptoticModelCoefficient nu sigma2","missing":[],"search":"armstreamasymptoticmodelcoefficient_nonneg banditrlproof.ucb.armstreamasymptoticmodelcoefficient_nonneg theorem armstreamasymptoticmodelcoefficient_nonneg {k : nat} (hk : 0 < k) (nu : kernel (fin k) real) (sigma2 : nnreal) : 0 <= armstreamasymptoticmodelcoefficient nu sigma2 theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lml_sum_four_le_armStreamAsymptoticModelCoefficient","label":"lml_sum_four_le_armStreamAsymptoticModelCoefficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lml_sum_four_le_armStreamAsymptoticModelCoefficient","description":"theorem lml_sum_four_le_armStreamAsymptoticModelCoefficient {K : Nat} (hK : 0 < K) (nu : Kernel (Fin K) Real) (sigma2 : NNReal) (n : Nat) : (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * 4 * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm + realKernelGap nu arm * (2 + 2 * (constSum 4 n).toReal)) <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Real.log ((n + 1 : Nat) : Real))","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-a8468d7c9853","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2739,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:86"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lml_sum_four_le_armStreamAsymptoticModelCoefficient {K : Nat} (hK : 0 < K) (nu : Kernel (Fin K) Real) (sigma2 : NNReal) (n : Nat) : (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * 4 * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm + realKernelGap nu arm * (2 + 2 * (constSum 4 n).toReal)) <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Real.log ((n + 1 : Nat) : Real))","missing":[],"search":"lml_sum_four_le_armstreamasymptoticmodelcoefficient banditrlproof.ucb.lml_sum_four_le_armstreamasymptoticmodelcoefficient theorem lml_sum_four_le_armstreamasymptoticmodelcoefficient {k : nat} (hk : 0 < k) (nu : kernel (fin k) real) (sigma2 : nnreal) (n : nat) : (finset.univ : finset (fin k)).sum (fun arm => 8 * 4 * (sigma2 : real) * real.log ((n + 1 : nat) : real) / realkernelgap nu arm + realkernelgap nu arm * (2 + 2 * (constsum 4 n).toreal)) <= armstreamasymptoticmodelcoefficient nu sigma2 * (1 + real.log ((n + 1 : nat) : real)) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamExpectedRegret","label":"armStreamExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamExpectedRegret","description":"Expected regret of one fixed recursive arm-stream UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-1e94cf8aa1f9","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2740,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:149"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamExpectedRegret {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Real","missing":[],"search":"armstreamexpectedregret banditrlproof.ucb.armstreamexpectedregret expected regret of one fixed recursive arm-stream ucb process. definition compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamExpectedRegret_nonneg_and_le","label":"armStreamExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamExpectedRegret_nonneg_and_le","description":"theorem armStreamExpectedRegret_nonneg_and_le {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (n : Nat) : 0 <= armStreamExpectedRegret hK sigma2 nu n ∧ armStreamExpectedRegret hK sigma2 nu n <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Rea…","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-a1e4b1cce6e4","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2741,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:157"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamExpectedRegret_nonneg_and_le {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) (n : Nat) : 0 <= armStreamExpectedRegret hK sigma2 nu n ∧ armStreamExpectedRegret hK sigma2 nu n <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Real.log ((n + 1 : Nat) : Real))","missing":[],"search":"armstreamexpectedregret_nonneg_and_le banditrlproof.ucb.armstreamexpectedregret_nonneg_and_le theorem armstreamexpectedregret_nonneg_and_le {k : nat} (hk : 0 < k) (sigma2 : nnreal) (nu : kernel (fin k) real) [ismarkovkernel nu] (hsigma2 : sigma2 ≠ 0) (hsubg : ∀ arm : fin k, hassubgaussianmgf (fun reward => reward - realkernelmean nu arm) sigma2 (nu arm)) (n : nat) : 0 <= armstreamexpectedregret hk sigma2 nu n ∧ armstreamexpectedregret hk sigma2 nu n <= armstreamasymptoticmodelcoefficient nu sigma2 * (1 + real.log ((n + 1 : nat) : real)) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamExpectedRegret_isBigO_log","label":"armStreamExpectedRegret_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamExpectedRegret_isBigO_log","description":"theorem armStreamExpectedRegret_isBigO_log {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : (armStreamExpectedRegret hK sigma2 nu) =O[atTop] (fun n : Nat => Real.log ((n + 1 : Nat) : Real))","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-cd72b5e1b950","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2742,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:182"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamExpectedRegret_isBigO_log {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : (armStreamExpectedRegret hK sigma2 nu) =O[atTop] (fun n : Nat => Real.log ((n + 1 : Nat) : Real))","missing":[],"search":"armstreamexpectedregret_isbigo_log banditrlproof.ucb.armstreamexpectedregret_isbigo_log theorem armstreamexpectedregret_isbigo_log {k : nat} (hk : 0 < k) (sigma2 : nnreal) (nu : kernel (fin k) real) [ismarkovkernel nu] (hsigma2 : sigma2 ≠ 0) (hsubg : ∀ arm : fin k, hassubgaussianmgf (fun reward => reward - realkernelmean nu arm) sigma2 (nu arm)) : (armstreamexpectedregret hk sigma2 nu) =o[attop] (fun n : nat => real.log ((n + 1 : nat) : real)) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamExpectedRegret_isLittleO_natCast_succ","label":"armStreamExpectedRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamExpectedRegret_isLittleO_natCast_succ","description":"theorem armStreamExpectedRegret_isLittleO_natCast_succ {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : (armStreamExpectedRegret hK sigma2 nu) =o[atTop] (fun n : Nat => ((n + 1 : Nat) : Real))","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-bdb8cc45245e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2743,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:223"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamExpectedRegret_isLittleO_natCast_succ {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : (armStreamExpectedRegret hK sigma2 nu) =o[atTop] (fun n : Nat => ((n + 1 : Nat) : Real))","missing":[],"search":"armstreamexpectedregret_islittleo_natcast_succ banditrlproof.ucb.armstreamexpectedregret_islittleo_natcast_succ theorem armstreamexpectedregret_islittleo_natcast_succ {k : nat} (hk : 0 < k) (sigma2 : nnreal) (nu : kernel (fin k) real) [ismarkovkernel nu] (hsigma2 : sigma2 ≠ 0) (hsubg : ∀ arm : fin k, hassubgaussianmgf (fun reward => reward - realkernelmean nu arm) sigma2 (nu arm)) : (armstreamexpectedregret hk sigma2 nu) =o[attop] (fun n : nat => ((n + 1 : nat) : real)) theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamExpectedAverageRegret","label":"armStreamExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamExpectedAverageRegret","description":"Expected regret of the fixed arm-stream policy normalized by `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-7417fbdde3bf","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2744,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:236"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamExpectedAverageRegret {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Real","missing":[],"search":"armstreamexpectedaverageregret banditrlproof.ucb.armstreamexpectedaverageregret expected regret of the fixed arm-stream policy normalized by `n + 1`. definition compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamExpectedAverageRegret_tendsto_zero","label":"armStreamExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamExpectedAverageRegret_tendsto_zero","description":"One fixed canonical arm-stream UCB policy has vanishing expected average regret under its one fixed product measure.","url":"../modules/banditrlproof-algorithms-ucbarmstreamasymptotics/index.html#decl-2025f0cd6691","parent":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","order":2745,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmStreamAsymptotics.lean:245"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamExpectedAverageRegret_tendsto_zero {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (hsigma2 : sigma2 ≠ 0) (hsubG : ∀ arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : Tendsto (armStreamExpectedAverageRegret hK sigma2 nu) atTop (nhds 0)","missing":[],"search":"armstreamexpectedaverageregret_tendsto_zero banditrlproof.ucb.armstreamexpectedaverageregret_tendsto_zero one fixed canonical arm-stream ucb policy has vanishing expected average regret under its one fixed product measure. theorem compiled","shard":"modules/cf57712367db9821.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamCoordinate","label":"armStreamCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamCoordinate","description":"A flattened pull-index/arm coordinate of a latent arm-reward stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-dae33d6732c5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2746,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:21"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def armStreamCoordinate {K : Nat} (index : Nat × Fin K) : ArmRewardStream K -> Real","missing":[],"search":"armstreamcoordinate banditrlproof.ucb.armstreamcoordinate a flattened pull-index/arm coordinate of a latent arm-reward stream. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamWithoutCoordinate","label":"armStreamWithoutCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamWithoutCoordinate","description":"The latent arm stream with one specified coordinate omitted.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-6a1124d788f9","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2747,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:26"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def armStreamWithoutCoordinate {K : Nat} (target : Nat × Fin K) : ArmRewardStream K -> ({index : Nat × Fin K // index ≠ target} -> Real)","missing":[],"search":"armstreamwithoutcoordinate banditrlproof.ucb.armstreamwithoutcoordinate the latent arm stream with one specified coordinate omitted. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamInsertCoordinate","label":"armStreamInsertCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamInsertCoordinate","description":"Reconstruct a latent arm stream after supplying one omitted coordinate.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-65d9ac196028","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2748,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:31"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def armStreamInsertCoordinate {K : Nat} (target : Nat × Fin K) (value : Real) : ({index : Nat × Fin K // index ≠ target} -> Real) -> ArmRewardStream K","missing":[],"search":"armstreaminsertcoordinate banditrlproof.ucb.armstreaminsertcoordinate reconstruct a latent arm stream after supplying one omitted coordinate. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistoryAction","label":"armStreamHistoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistoryAction","description":"The history/action condition used by the successor reward kernel.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-35e447331cc8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2749,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:38"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamHistoryAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : ArmRewardStream K -> History.FinitePairHistory (Fin K) Real n × Fin K","missing":[],"search":"armstreamhistoryaction banditrlproof.ucb.armstreamhistoryaction the history/action condition used by the successor reward kernel. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamNextCoordinate","label":"armStreamNextCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamNextCoordinate","description":"The next unused pull-index/arm coordinate selected after history `n`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-173a48f34876","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2750,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:47"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamNextCoordinate {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : ArmRewardStream K -> Nat × Fin K","missing":[],"search":"armstreamnextcoordinate banditrlproof.ucb.armstreamnextcoordinate the next unused pull-index/arm coordinate selected after history `n`. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamCoordinateOfHistoryAction","label":"armStreamCoordinateOfHistoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamCoordinateOfHistoryAction","description":"The pull-index/arm coordinate encoded by a successor history/action condition.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-5b028b76f3c8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2751,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:55"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamCoordinateOfHistoryAction {K : Nat} (n : Nat) : History.FinitePairHistory (Fin K) Real n × Fin K -> Nat × Fin K","missing":[],"search":"armstreamcoordinateofhistoryaction banditrlproof.ucb.armstreamcoordinateofhistoryaction the pull-index/arm coordinate encoded by a successor history/action condition. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistoryActionCoordinateBranch","label":"armStreamHistoryActionCoordinateBranch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistoryActionCoordinateBranch","description":"Conditions that encode one fixed next pull-index/arm coordinate.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-9ce3929a4565","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2752,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:62"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) (target : Nat × Fin K) : Set (History.FinitePairHistory (Fin K) Real n × Fin K)","missing":[],"search":"armstreamhistoryactioncoordinatebranch banditrlproof.ucb.armstreamhistoryactioncoordinatebranch conditions that encode one fixed next pull-index/arm coordinate. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamSelectedRewardKernel","label":"armStreamSelectedRewardKernel","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamSelectedRewardKernel","description":"The stationary reward kernel selected by the arm component of a condition.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-3ed0cd2afb53","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2753,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:68"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable abbrev armStreamSelectedRewardKernel {K : Nat} (n : Nat) (nu : Kernel (Fin K) Real) : Kernel (History.FinitePairHistory (Fin K) Real n × Fin K) Real","missing":[],"search":"armstreamselectedrewardkernel banditrlproof.ucb.armstreamselectedrewardkernel the stationary reward kernel selected by the arm component of a condition. abbreviation compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamNextCoordinateBranch","label":"armStreamNextCoordinateBranch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamNextCoordinateBranch","description":"Latent streams for which one fixed pull-index/arm coordinate is selected next.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-8d4f7abe902e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2754,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:74"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) : Set (ArmRewardStream K)","missing":[],"search":"armstreamnextcoordinatebranch banditrlproof.ucb.armstreamnextcoordinatebranch latent streams for which one fixed pull-index/arm coordinate is selected next. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistoryActionFromWithout","label":"armStreamHistoryActionFromWithout","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistoryActionFromWithout","description":"Reconstructed successor condition using only a fixed coordinate's complement.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-b87872b30e4e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2755,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:80"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamHistoryActionFromWithout {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) : ({index : Nat × Fin K // index ≠ target} -> Real) -> History.FinitePairHistory (Fin K) Real n × Fin K","missing":[],"search":"armstreamhistoryactionfromwithout banditrlproof.ucb.armstreamhistoryactionfromwithout reconstructed successor condition using only a fixed coordinate's complement. definition compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamCoordinate","label":"measurable_armStreamCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamCoordinate","description":"theorem measurable_armStreamCoordinate {K : Nat} (index : Nat × Fin K) : Measurable (armStreamCoordinate index)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-41cc0a0dad18","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2756,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:87"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamCoordinate {K : Nat} (index : Nat × Fin K) : Measurable (armStreamCoordinate index)","missing":[],"search":"measurable_armstreamcoordinate banditrlproof.ucb.measurable_armstreamcoordinate theorem measurable_armstreamcoordinate {k : nat} (index : nat × fin k) : measurable (armstreamcoordinate index) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamWithoutCoordinate","label":"measurable_armStreamWithoutCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamWithoutCoordinate","description":"theorem measurable_armStreamWithoutCoordinate {K : Nat} (target : Nat × Fin K) : Measurable (armStreamWithoutCoordinate target)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-3dc518ca0235","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2757,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:92"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamWithoutCoordinate {K : Nat} (target : Nat × Fin K) : Measurable (armStreamWithoutCoordinate target)","missing":[],"search":"measurable_armstreamwithoutcoordinate banditrlproof.ucb.measurable_armstreamwithoutcoordinate theorem measurable_armstreamwithoutcoordinate {k : nat} (target : nat × fin k) : measurable (armstreamwithoutcoordinate target) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamInsertCoordinate","label":"measurable_armStreamInsertCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamInsertCoordinate","description":"theorem measurable_armStreamInsertCoordinate {K : Nat} (target : Nat × Fin K) (value : Real) : Measurable (armStreamInsertCoordinate target value)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-c449766a4bc2","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2758,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:98"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamInsertCoordinate {K : Nat} (target : Nat × Fin K) (value : Real) : Measurable (armStreamInsertCoordinate target value)","missing":[],"search":"measurable_armstreaminsertcoordinate banditrlproof.ucb.measurable_armstreaminsertcoordinate theorem measurable_armstreaminsertcoordinate {k : nat} (target : nat × fin k) (value : real) : measurable (armstreaminsertcoordinate target value) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamWithoutCoordinate_insertCoordinate","label":"armStreamWithoutCoordinate_insertCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamWithoutCoordinate_insertCoordinate","description":"theorem armStreamWithoutCoordinate_insertCoordinate {K : Nat} (target : Nat × Fin K) (value : Real) (rest : {index : Nat × Fin K // index ≠ target} -> Real) : armStreamWithoutCoordinate target (armStreamInsertCoordinate target value rest) = rest","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-fa6b50b55f6f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2759,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:113"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamWithoutCoordinate_insertCoordinate {K : Nat} (target : Nat × Fin K) (value : Real) (rest : {index : Nat × Fin K // index ≠ target} -> Real) : armStreamWithoutCoordinate target (armStreamInsertCoordinate target value rest) = rest","missing":[],"search":"armstreamwithoutcoordinate_insertcoordinate banditrlproof.ucb.armstreamwithoutcoordinate_insertcoordinate theorem armstreamwithoutcoordinate_insertcoordinate {k : nat} (target : nat × fin k) (value : real) (rest : {index : nat × fin k // index ≠ target} -> real) : armstreamwithoutcoordinate target (armstreaminsertcoordinate target value rest) = rest theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamHistoryAction","label":"measurable_armStreamHistoryAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamHistoryAction","description":"theorem measurable_armStreamHistoryAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (armStreamHistoryAction hK c n)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-b94f829bb6bc","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2760,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:122"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamHistoryAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (armStreamHistoryAction hK c n)","missing":[],"search":"measurable_armstreamhistoryaction banditrlproof.ucb.measurable_armstreamhistoryaction theorem measurable_armstreamhistoryaction {k : nat} (hk : 0 < k) (c : real) (n : nat) : measurable (armstreamhistoryaction hk c n) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamNextCoordinate","label":"measurable_armStreamNextCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamNextCoordinate","description":"theorem measurable_armStreamNextCoordinate {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (armStreamNextCoordinate hK c n)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-c56350eced97","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2761,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:129"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamNextCoordinate {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (armStreamNextCoordinate hK c n)","missing":[],"search":"measurable_armstreamnextcoordinate banditrlproof.ucb.measurable_armstreamnextcoordinate theorem measurable_armstreamnextcoordinate {k : nat} (hk : 0 < k) (c : real) (n : nat) : measurable (armstreamnextcoordinate hk c n) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamCoordinateOfHistoryAction","label":"measurable_armStreamCoordinateOfHistoryAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamCoordinateOfHistoryAction","description":"theorem measurable_armStreamCoordinateOfHistoryAction {K : Nat} (n : Nat) : Measurable (armStreamCoordinateOfHistoryAction (K := K) n)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-53a9bee4569f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2762,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:142"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamCoordinateOfHistoryAction {K : Nat} (n : Nat) : Measurable (armStreamCoordinateOfHistoryAction (K := K) n)","missing":[],"search":"measurable_armstreamcoordinateofhistoryaction banditrlproof.ucb.measurable_armstreamcoordinateofhistoryaction theorem measurable_armstreamcoordinateofhistoryaction {k : nat} (n : nat) : measurable (armstreamcoordinateofhistoryaction (k := k) n) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_armStreamHistoryActionCoordinateBranch","label":"measurableSet_armStreamHistoryActionCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_armStreamHistoryActionCoordinateBranch","description":"theorem measurableSet_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (armStreamHistoryActionCoordinateBranch n target)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-6d46ed34a7dd","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2763,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:154"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) (target : Nat × Fin K) : MeasurableSet (armStreamHistoryActionCoordinateBranch n target)","missing":[],"search":"measurableset_armstreamhistoryactioncoordinatebranch banditrlproof.ucb.measurableset_armstreamhistoryactioncoordinatebranch theorem measurableset_armstreamhistoryactioncoordinatebranch {k : nat} (n : nat) (target : nat × fin k) : measurableset (armstreamhistoryactioncoordinatebranch n target) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurableSet_armStreamNextCoordinateBranch","label":"measurableSet_armStreamNextCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurableSet_armStreamNextCoordinateBranch","description":"theorem measurableSet_armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) : MeasurableSet (armStreamNextCoordinateBranch hK c n target)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-495e58cdece5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2764,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:160"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) : MeasurableSet (armStreamNextCoordinateBranch hK c n target)","missing":[],"search":"measurableset_armstreamnextcoordinatebranch banditrlproof.ucb.measurableset_armstreamnextcoordinatebranch theorem measurableset_armstreamnextcoordinatebranch {k : nat} (hk : 0 < k) (c : real) (n : nat) (target : nat × fin k) : measurableset (armstreamnextcoordinatebranch hk c n target) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamNextCoordinate_eq_coordinateOfHistoryAction","label":"armStreamNextCoordinate_eq_coordinateOfHistoryAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamNextCoordinate_eq_coordinateOfHistoryAction","description":"theorem armStreamNextCoordinate_eq_coordinateOfHistoryAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (stream : ArmRewardStream K) : armStreamNextCoordinate hK c n stream = armStreamCoordinateOfHistoryAction n (armStreamHistoryAction hK c n stream)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-f085debf52db","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2765,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:167"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamNextCoordinate_eq_coordinateOfHistoryAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (stream : ArmRewardStream K) : armStreamNextCoordinate hK c n stream = armStreamCoordinateOfHistoryAction n (armStreamHistoryAction hK c n stream)","missing":[],"search":"armstreamnextcoordinate_eq_coordinateofhistoryaction banditrlproof.ucb.armstreamnextcoordinate_eq_coordinateofhistoryaction theorem armstreamnextcoordinate_eq_coordinateofhistoryaction {k : nat} (hk : 0 < k) (c : real) (n : nat) (stream : armrewardstream k) : armstreamnextcoordinate hk c n stream = armstreamcoordinateofhistoryaction n (armstreamhistoryaction hk c n stream) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pairwise_disjoint_armStreamHistoryActionCoordinateBranch","label":"pairwise_disjoint_armStreamHistoryActionCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pairwise_disjoint_armStreamHistoryActionCoordinateBranch","description":"theorem pairwise_disjoint_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) : Pairwise (Function.onFun (fun s t : Set (History.FinitePairHistory (Fin K) Real n × Fin K) => Disjoint s t) (fun target : Nat × Fin K => armStreamHistoryActionCoordinateBranch n target))","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-43c20d212bd2","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2766,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:175"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pairwise_disjoint_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) : Pairwise (Function.onFun (fun s t : Set (History.FinitePairHistory (Fin K) Real n × Fin K) => Disjoint s t) (fun target : Nat × Fin K => armStreamHistoryActionCoordinateBranch n target))","missing":[],"search":"pairwise_disjoint_armstreamhistoryactioncoordinatebranch banditrlproof.ucb.pairwise_disjoint_armstreamhistoryactioncoordinatebranch theorem pairwise_disjoint_armstreamhistoryactioncoordinatebranch {k : nat} (n : nat) : pairwise (function.onfun (fun s t : set (history.finitepairhistory (fin k) real n × fin k) => disjoint s t) (fun target : nat × fin k => armstreamhistoryactioncoordinatebranch n target)) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pairwise_disjoint_armStreamNextCoordinateBranch","label":"pairwise_disjoint_armStreamNextCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pairwise_disjoint_armStreamNextCoordinateBranch","description":"theorem pairwise_disjoint_armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Pairwise (Function.onFun (fun s t : Set (ArmRewardStream K) => Disjoint s t) (fun target : Nat × Fin K => armStreamNextCoordinateBranch hK c n target))","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-43a29c40628c","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2767,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:191"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pairwise_disjoint_armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Pairwise (Function.onFun (fun s t : Set (ArmRewardStream K) => Disjoint s t) (fun target : Nat × Fin K => armStreamNextCoordinateBranch hK c n target))","missing":[],"search":"pairwise_disjoint_armstreamnextcoordinatebranch banditrlproof.ucb.pairwise_disjoint_armstreamnextcoordinatebranch theorem pairwise_disjoint_armstreamnextcoordinatebranch {k : nat} (hk : 0 < k) (c : real) (n : nat) : pairwise (function.onfun (fun s t : set (armrewardstream k) => disjoint s t) (fun target : nat × fin k => armstreamnextcoordinatebranch hk c n target)) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.iUnion_armStreamHistoryActionCoordinateBranch","label":"iUnion_armStreamHistoryActionCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.iUnion_armStreamHistoryActionCoordinateBranch","description":"theorem iUnion_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) : (⋃ target : Nat × Fin K, armStreamHistoryActionCoordinateBranch n target) = Set.univ","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-0e3363a1f7d5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2768,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:206"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem iUnion_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) : (⋃ target : Nat × Fin K, armStreamHistoryActionCoordinateBranch n target) = Set.univ","missing":[],"search":"iunion_armstreamhistoryactioncoordinatebranch banditrlproof.ucb.iunion_armstreamhistoryactioncoordinatebranch theorem iunion_armstreamhistoryactioncoordinatebranch {k : nat} (n : nat) : (⋃ target : nat × fin k, armstreamhistoryactioncoordinatebranch n target) = set.univ theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.iUnion_armStreamNextCoordinateBranch","label":"iUnion_armStreamNextCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.iUnion_armStreamNextCoordinateBranch","description":"theorem iUnion_armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : (⋃ target : Nat × Fin K, armStreamNextCoordinateBranch hK c n target) = Set.univ","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-22d8da96eaf2","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2769,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:213"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem iUnion_armStreamNextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : (⋃ target : Nat × Fin K, armStreamNextCoordinateBranch hK c n target) = Set.univ","missing":[],"search":"iunion_armstreamnextcoordinatebranch banditrlproof.ucb.iunion_armstreamnextcoordinatebranch theorem iunion_armstreamnextcoordinatebranch {k : nat} (hk : 0 < k) (c : real) (n : nat) : (⋃ target : nat × fin k, armstreamnextcoordinatebranch hk c n target) = set.univ theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_eq_sum_restrict_armStreamHistoryActionCoordinateBranch","label":"measure_eq_sum_restrict_armStreamHistoryActionCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_eq_sum_restrict_armStreamHistoryActionCoordinateBranch","description":"theorem measure_eq_sum_restrict_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) (mu : Measure (History.FinitePairHistory (Fin K) Real n × Fin K)) : mu = Measure.sum fun target : Nat × Fin K => mu.restrict (armStreamHistoryActionCoordinateBranch n target)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-3f5af2b2c7cf","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2770,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:220"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_eq_sum_restrict_armStreamHistoryActionCoordinateBranch {K : Nat} (n : Nat) (mu : Measure (History.FinitePairHistory (Fin K) Real n × Fin K)) : mu = Measure.sum fun target : Nat × Fin K => mu.restrict (armStreamHistoryActionCoordinateBranch n target)","missing":[],"search":"measure_eq_sum_restrict_armstreamhistoryactioncoordinatebranch banditrlproof.ucb.measure_eq_sum_restrict_armstreamhistoryactioncoordinatebranch theorem measure_eq_sum_restrict_armstreamhistoryactioncoordinatebranch {k : nat} (n : nat) (mu : measure (history.finitepairhistory (fin k) real n × fin k)) : mu = measure.sum fun target : nat × fin k => mu.restrict (armstreamhistoryactioncoordinatebranch n target) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_eq_sum_restrict_nextCoordinateBranch","label":"armStreamMeasure_eq_sum_restrict_nextCoordinateBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_eq_sum_restrict_nextCoordinateBranch","description":"theorem armStreamMeasure_eq_sum_restrict_nextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : armStreamMeasure nu = Measure.sum fun target : Nat × Fin K => (armStreamMeasure nu).restrict (armStreamNextCoordinateBranch hK c n target)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-c05878669260","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2771,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:232"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_eq_sum_restrict_nextCoordinateBranch {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : armStreamMeasure nu = Measure.sum fun target : Nat × Fin K => (armStreamMeasure nu).restrict (armStreamNextCoordinateBranch hK c n target)","missing":[],"search":"armstreammeasure_eq_sum_restrict_nextcoordinatebranch banditrlproof.ucb.armstreammeasure_eq_sum_restrict_nextcoordinatebranch theorem armstreammeasure_eq_sum_restrict_nextcoordinatebranch {k : nat} (hk : 0 < k) (c : real) (n : nat) (nu : kernel (fin k) real) [ismarkovkernel nu] : armstreammeasure nu = measure.sum fun target : nat × fin k => (armstreammeasure nu).restrict (armstreamnextcoordinatebranch hk c n target) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamHistoryActionFromWithout","label":"measurable_armStreamHistoryActionFromWithout","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamHistoryActionFromWithout","description":"theorem measurable_armStreamHistoryActionFromWithout {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) : Measurable (armStreamHistoryActionFromWithout hK c n target value)","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-04ea7251325e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2772,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:244"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamHistoryActionFromWithout {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) : Measurable (armStreamHistoryActionFromWithout hK c n target value)","missing":[],"search":"measurable_armstreamhistoryactionfromwithout banditrlproof.ucb.measurable_armstreamhistoryactionfromwithout theorem measurable_armstreamhistoryactionfromwithout {k : nat} (hk : 0 < k) (c : real) (n : nat) (target : nat × fin k) (value : real) : measurable (armstreamhistoryactionfromwithout hk c n target value) theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamNextCoordinate_fst_eq_pullCount","label":"armStreamNextCoordinate_fst_eq_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamNextCoordinate_fst_eq_pullCount","description":"The first component of the next coordinate is the selected arm's current count.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-c23030a228e9","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2773,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:252"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamNextCoordinate_fst_eq_pullCount {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) : (armStreamNextCoordinate hK c n stream).1 = pullCount (armStreamAction hK c stream) (armStreamNextCoordinate hK c n stream).2 (n + 1)","missing":[],"search":"armstreamnextcoordinate_fst_eq_pullcount banditrlproof.ucb.armstreamnextcoordinate_fst_eq_pullcount the first component of the next coordinate is the selected arm's current count. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamReward_succ_eq_nextCoordinate","label":"armStreamReward_succ_eq_nextCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamReward_succ_eq_nextCoordinate","description":"The successor reward reads exactly the next coordinate selected after history `n`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-e9528e7cf96e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2774,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:263"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamReward_succ_eq_nextCoordinate {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) : armStreamReward hK c stream (n + 1) = armStreamCoordinate (armStreamNextCoordinate hK c n stream) stream","missing":[],"search":"armstreamreward_succ_eq_nextcoordinate banditrlproof.ucb.armstreamreward_succ_eq_nextcoordinate the successor reward reads exactly the next coordinate selected after history `n`. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistory_eq_of_eq_below_pullCount","label":"armStreamHistory_eq_of_eq_below_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistory_eq_of_eq_below_pullCount","description":"The recursive history through `n` only reads coordinates whose pull index is strictly below the corresponding arm count at time `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-1c9a48cb49fa","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2775,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:279"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamHistory_eq_of_eq_below_pullCount {K : Nat} (hK : 0 < K) (c : Real) (stream stream' : ArmRewardStream K) (n : Nat) (hagrees : ∀ arm index, index < pullCount (armStreamAction hK c stream) arm (n + 1) → stream index arm = stream' index arm) : armStreamHistory hK c stream n = armStreamHistory hK c stream' n","missing":[],"search":"armstreamhistory_eq_of_eq_below_pullcount banditrlproof.ucb.armstreamhistory_eq_of_eq_below_pullcount the recursive history through `n` only reads coordinates whose pull index is strictly below the corresponding arm count at time `n + 1`. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistory_eq_of_withoutCoordinate_eq_of_pullCount_le","label":"armStreamHistory_eq_of_withoutCoordinate_eq_of_pullCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistory_eq_of_withoutCoordinate_eq_of_pullCount_le","description":"Changing one not-yet-consumed coordinate cannot alter the history already generated from the latent stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-173c939a087f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2776,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:328"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamHistory_eq_of_withoutCoordinate_eq_of_pullCount_le {K : Nat} (hK : 0 < K) (c : Real) (target : Nat × Fin K) (stream stream' : ArmRewardStream K) (n : Nat) (hwithout : armStreamWithoutCoordinate target stream = armStreamWithoutCoordinate target stream') (hfuture : pullCount (armStreamAction hK c stream) target.2 (n + 1) ≤ target.1) : armStreamHistory hK c stream n = armStreamHistory hK c stream' n","missing":[],"search":"armstreamhistory_eq_of_withoutcoordinate_eq_of_pullcount_le banditrlproof.ucb.armstreamhistory_eq_of_withoutcoordinate_eq_of_pullcount_le changing one not-yet-consumed coordinate cannot alter the history already generated from the latent stream. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamNextCoordinate_eq_iff_insertCoordinate","label":"armStreamNextCoordinate_eq_iff_insertCoordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamNextCoordinate_eq_iff_insertCoordinate","description":"The event that a fixed coordinate is selected next factors through the stream with that coordinate omitted.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-eb6cdcb82935","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2777,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:350"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamNextCoordinate_eq_iff_insertCoordinate {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) (stream : ArmRewardStream K) : armStreamNextCoordinate hK c n stream = target ↔ armStreamNextCoordinate hK c n (armStreamInsertCoordinate target value (armStreamWithoutCoordinate target stream)) = target","missing":[],"search":"armstreamnextcoordinate_eq_iff_insertcoordinate banditrlproof.ucb.armstreamnextcoordinate_eq_iff_insertcoordinate the event that a fixed coordinate is selected next factors through the stream with that coordinate omitted. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistoryAction_eq_fromWithout_of_nextCoordinate_eq","label":"armStreamHistoryAction_eq_fromWithout_of_nextCoordinate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistoryAction_eq_fromWithout_of_nextCoordinate_eq","description":"On the branch selecting `target`, the actual successor condition equals its reconstruction from all coordinates except `target`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-04977b855e20","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2778,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:384"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamHistoryAction_eq_fromWithout_of_nextCoordinate_eq {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) (stream : ArmRewardStream K) (hnext : armStreamNextCoordinate hK c n stream = target) : armStreamHistoryAction hK c n stream = armStreamHistoryActionFromWithout hK c n target value (armStreamWithoutCoordinate target stream)","missing":[],"search":"armstreamhistoryaction_eq_fromwithout_of_nextcoordinate_eq banditrlproof.ucb.armstreamhistoryaction_eq_fromwithout_of_nextcoordinate_eq on the branch selecting `target`, the actual successor condition equals its reconstruction from all coordinates except `target`. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.iIndepFun_armStreamMeasure_coordinate","label":"iIndepFun_armStreamMeasure_coordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.iIndepFun_armStreamMeasure_coordinate","description":"All pull-index/arm coordinates are mutually independent.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-83a0f33b41a5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2779,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:405"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_armStreamMeasure_coordinate {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : iIndepFun (fun index : Nat × Fin K => armStreamCoordinate index) (armStreamMeasure nu)","missing":[],"search":"iindepfun_armstreammeasure_coordinate banditrlproof.ucb.iindepfun_armstreammeasure_coordinate all pull-index/arm coordinates are mutually independent. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.indepFun_armStreamMeasure_coordinate_without","label":"indepFun_armStreamMeasure_coordinate_without","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.indepFun_armStreamMeasure_coordinate_without","description":"A fixed coordinate is independent of the collection of all other coordinates.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-56c3eca87794","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2780,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:417"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem indepFun_armStreamMeasure_coordinate_without {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (target : Nat × Fin K) : IndepFun (armStreamCoordinate target) (armStreamWithoutCoordinate target) (armStreamMeasure nu)","missing":[],"search":"indepfun_armstreammeasure_coordinate_without banditrlproof.ucb.indepfun_armstreammeasure_coordinate_without a fixed coordinate is independent of the collection of all other coordinates. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.indepFun_armStreamMeasure_coordinate_historyActionFromWithout","label":"indepFun_armStreamMeasure_coordinate_historyActionFromWithout","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.indepFun_armStreamMeasure_coordinate_historyActionFromWithout","description":"A fixed reward coordinate is independent of the reconstructed successor condition.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-b2e5843fd07f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2781,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:459"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem indepFun_armStreamMeasure_coordinate_historyActionFromWithout {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (target : Nat × Fin K) (value : Real) : IndepFun (armStreamCoordinate target) (fun stream : ArmRewardStream K => armStreamHistoryActionFromWithout hK c n target value (armStreamWithoutCoordinate target stream)) (armStreamMeasure nu)","missing":[],"search":"indepfun_armstreammeasure_coordinate_historyactionfromwithout banditrlproof.ucb.indepfun_armstreammeasure_coordinate_historyactionfromwithout a fixed reward coordinate is independent of the reconstructed successor condition. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_historyActionFromWithout_coordinate_eq_prod","label":"armStreamMeasure_map_historyActionFromWithout_coordinate_eq_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_historyActionFromWithout_coordinate_eq_prod","description":"The reconstructed successor condition and a fixed omitted coordinate have a product joint law, with the prescribed arm marginal on the reward coordinate.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-3e3016df354f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2782,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:478"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_historyActionFromWithout_coordinate_eq_prod {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (target : Nat × Fin K) (value : Real) : Measure.map (fun stream : ArmRewardStream K => (armStreamHistoryActionFromWithout hK c n target value (armStreamWithoutCoordinate target stream), armStreamCoordinate target stream)) (armStreamMeasure nu) = (Measure.map (fun stream : ArmRewardStream K => armStreamHistoryActionFromWithout hK c n target value (armStreamWithoutCoordinate target stream)) (armStreamMeasure nu)).prod (nu target.2)","missing":[],"search":"armstreammeasure_map_historyactionfromwithout_coordinate_eq_prod banditrlproof.ucb.armstreammeasure_map_historyactionfromwithout_coordinate_eq_prod the reconstructed successor condition and a fixed omitted coordinate have a product joint law, with the prescribed arm marginal on the reward coordinate. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistoryActionFromWithout_mem_coordinateBranch_iff","label":"armStreamHistoryActionFromWithout_mem_coordinateBranch_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistoryActionFromWithout_mem_coordinateBranch_iff","description":"The reconstructed condition is in `target`'s branch exactly on the actual branch.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-a4933a943610","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2783,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:511"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamHistoryActionFromWithout_mem_coordinateBranch_iff {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) (stream : ArmRewardStream K) : armStreamHistoryActionFromWithout hK c n target value (armStreamWithoutCoordinate target stream) ∈ armStreamHistoryActionCoordinateBranch n target ↔ stream ∈ armStreamNextCoordinateBranch hK c n target","missing":[],"search":"armstreamhistoryactionfromwithout_mem_coordinatebranch_iff banditrlproof.ucb.armstreamhistoryactionfromwithout_mem_coordinatebranch_iff the reconstructed condition is in `target`'s branch exactly on the actual branch. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.map_historyActionFromWithout_restrict_coordinateBranch_eq_historyAction","label":"map_historyActionFromWithout_restrict_coordinateBranch_eq_historyAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.map_historyActionFromWithout_restrict_coordinateBranch_eq_historyAction","description":"Both actual and reconstructed condition maps induce the same branch restriction.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-1b5b7b82c81b","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2784,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:528"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem map_historyActionFromWithout_restrict_coordinateBranch_eq_historyAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (target : Nat × Fin K) (value : Real) (mu : Measure (ArmRewardStream K)) : (Measure.map (fun stream : ArmRewardStream K => armStreamHistoryActionFromWithout hK c n target value (armStreamWithoutCoordinate target stream)) mu).restrict (armStreamHistoryActionCoordinateBranch n target) = (Measure.map (armStreamHistoryAction hK c n) mu).restrict (armStreamHistoryActionCoordinateBranch n target)","missing":[],"search":"map_historyactionfromwithout_restrict_coordinatebranch_eq_historyaction banditrlproof.ucb.map_historyactionfromwithout_restrict_coordinatebranch_eq_historyaction both actual and reconstructed condition maps induce the same branch restriction. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_historyAction_reward_restrict_branch_eq_prod","label":"armStreamMeasure_map_historyAction_reward_restrict_branch_eq_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_historyAction_reward_restrict_branch_eq_prod","description":"On one fixed next-coordinate branch, the actual successor condition/reward pair has the restricted condition marginal times the prescribed arm law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-d8ad617a69a8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2785,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:577"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_historyAction_reward_restrict_branch_eq_prod {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (target : Nat × Fin K) (value : Real) : Measure.map (fun stream : ArmRewardStream K => (armStreamHistoryAction hK c n stream, armStreamReward hK c stream (n + 1))) ((armStreamMeasure nu).restrict (armStreamNextCoordinateBranch hK c n target)) = ((Measure.map (armStreamHistoryAction hK c n) (armStreamMeasure nu)).restrict (armStreamHistoryActionCoordinateBranch n target)).prod (nu target.2)","missing":[],"search":"armstreammeasure_map_historyaction_reward_restrict_branch_eq_prod banditrlproof.ucb.armstreammeasure_map_historyaction_reward_restrict_branch_eq_prod on one fixed next-coordinate branch, the actual successor condition/reward pair has the restricted condition marginal times the prescribed arm law. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_historyAction_reward_succ_eq_compProd","label":"armStreamMeasure_map_historyAction_reward_succ_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_historyAction_reward_succ_eq_compProd","description":"The full successor condition/reward joint law is the condition marginal followed by the stationary law of the arm selected in that condition.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-d16fcbeca916","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2786,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:657"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_historyAction_reward_succ_eq_compProd {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Measure.map (fun stream : ArmRewardStream K => (armStreamHistoryAction hK c n stream, armStreamReward hK c stream (n + 1))) (armStreamMeasure nu) = Measure.compProd (Measure.map (armStreamHistoryAction hK c n) (armStreamMeasure nu)) (armStreamSelectedRewardKernel n nu)","missing":[],"search":"armstreammeasure_map_historyaction_reward_succ_eq_compprod banditrlproof.ucb.armstreammeasure_map_historyaction_reward_succ_eq_compprod the full successor condition/reward joint law is the condition marginal followed by the stationary law of the arm selected in that condition. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamReward_succ_condDistrib_ae_eq_nu","label":"armStreamReward_succ_condDistrib_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamReward_succ_condDistrib_ae_eq_nu","description":"The successor arm-stream reward conditional law is the selected arm law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-a21f8fa2a131","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2787,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:756"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamReward_succ_condDistrib_ae_eq_nu {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Filter.EventuallyEq (ae ((armStreamMeasure nu).map (armStreamHistoryAction hK c n))) (condDistrib (fun stream : ArmRewardStream K => armStreamReward hK c stream (n + 1)) (armStreamHistoryAction hK c n) (armStreamMeasure nu)) (armStreamSelectedRewardKernel n nu)","missing":[],"search":"armstreamreward_succ_conddistrib_ae_eq_nu banditrlproof.ucb.armstreamreward_succ_conddistrib_ae_eq_nu the successor arm-stream reward conditional law is the selected arm law. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment_feedback_ae_eq_nu","label":"canonicalArmStreamHistoryEnvironment_feedback_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment_feedback_ae_eq_nu","description":"The packaged canonical successor-feedback kernel is the stationary selected law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-269876c40a07","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2788,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:777"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalArmStreamHistoryEnvironment_feedback_ae_eq_nu {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Filter.EventuallyEq (ae ((armStreamMeasure nu).map (armStreamHistoryAction hK (c * (sigma2 : Real)) n))) ((canonicalArmStreamHistoryEnvironment hK c sigma2 nu).feedback n) (armStreamSelectedRewardKernel n nu)","missing":[],"search":"canonicalarmstreamhistoryenvironment_feedback_ae_eq_nu banditrlproof.ucb.canonicalarmstreamhistoryenvironment_feedback_ae_eq_nu the packaged canonical successor-feedback kernel is the stationary selected law. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_condDistrib_coordinate_given_without","label":"armStreamMeasure_condDistrib_coordinate_given_without","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_condDistrib_coordinate_given_without","description":"Given every other latent coordinate, a fixed coordinate still has its prescribed stationary arm law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-7d5b7c0bb98d","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2789,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:805"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_condDistrib_coordinate_given_without {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (target : Nat × Fin K) : Filter.EventuallyEq (ae ((armStreamMeasure nu).map (armStreamWithoutCoordinate target))) (condDistrib (armStreamCoordinate target) (armStreamWithoutCoordinate target) (armStreamMeasure nu)) (Kernel.const _ (nu target.2))","missing":[],"search":"armstreammeasure_conddistrib_coordinate_given_without banditrlproof.ucb.armstreammeasure_conddistrib_coordinate_given_without given every other latent coordinate, a fixed coordinate still has its prescribed stationary arm law. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamReward_zero_condDistrib_ae_eq_nu","label":"armStreamReward_zero_condDistrib_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamReward_zero_condDistrib_ae_eq_nu","description":"The initial arm-stream reward conditional law is the selected arm law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-b91fe11b2a54","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2790,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:830"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamReward_zero_condDistrib_ae_eq_nu {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Filter.EventuallyEq (ae ((armStreamMeasure nu).map (fun stream : ArmRewardStream K => armStreamAction hK (c * (sigma2 : Real)) stream 0))) (condDistrib (fun stream : ArmRewardStream K => armStreamReward hK (c * (sigma2 : Real)) stream 0) (fun stream : ArmRewardStream K => armStreamAction hK (c * (sigma2 : Real)) stream 0) (armStreamMeasure nu)) nu","missing":[],"search":"armstreamreward_zero_conddistrib_ae_eq_nu banditrlproof.ucb.armstreamreward_zero_conddistrib_ae_eq_nu the initial arm-stream reward conditional law is the selected arm law. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment_initialFeedback_ae_eq_nu","label":"canonicalArmStreamHistoryEnvironment_initialFeedback_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment_initialFeedback_ae_eq_nu","description":"The packaged canonical initial-feedback kernel is the stationary selected law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamconditionalreward/index.html#decl-36621135302d","parent":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","order":2791,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamConditionalReward"],["Source","BanditRLProof/Algorithms/UCBArmStreamConditionalReward.lean:882"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalArmStreamHistoryEnvironment_initialFeedback_ae_eq_nu {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Filter.EventuallyEq (ae ((canonicalArmStreamHistoryAlgorithm hK c sigma2 nu).initialAction)) (canonicalArmStreamHistoryEnvironment hK c sigma2 nu).initialFeedback nu","missing":[],"search":"canonicalarmstreamhistoryenvironment_initialfeedback_ae_eq_nu banditrlproof.ucb.canonicalarmstreamhistoryenvironment_initialfeedback_ae_eq_nu the packaged canonical initial-feedback kernel is the stationary selected law. theorem compiled","shard":"modules/e3cce453d6a2762e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction_eq_initializationArm_of_lt","label":"armStreamAction_eq_initializationArm_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction_eq_initializationArm_of_lt","description":"During initialization the recursive process follows round robin exactly.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-7f3ebe2abce8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2792,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:26"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAction_eq_initializationArm_of_lt {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (t : Nat) (ht : t < K) : armStreamAction hK c stream t = initializationArm hK t","missing":[],"search":"armstreamaction_eq_initializationarm_of_lt banditrlproof.ucb.armstreamaction_eq_initializationarm_of_lt during initialization the recursive process follows round robin exactly. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullCount_armStreamAction_K_eq_one","label":"pullCount_armStreamAction_K_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pullCount_armStreamAction_K_eq_one","description":"Every arm is pulled exactly once in the first full initialization cycle.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-f01df755547b","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2793,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:38"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pullCount_armStreamAction_K_eq_one {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (arm : Fin K) : pullCount (armStreamAction hK c stream) arm K = 1","missing":[],"search":"pullcount_armstreamaction_k_eq_one banditrlproof.ucb.pullcount_armstreamaction_k_eq_one every arm is pulled exactly once in the first full initialization cycle. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullCount_armStreamAction_pos_of_K_le","label":"pullCount_armStreamAction_pos_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pullCount_armStreamAction_pos_of_K_le","description":"After initialization every arm has positive prior pull count.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-c9442c111dd8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2794,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:61"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pullCount_armStreamAction_pos_of_K_le {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (arm : Fin K) (t : Nat) (ht : K <= t) : 0 < pullCount (armStreamAction hK c stream) arm t","missing":[],"search":"pullcount_armstreamaction_pos_of_k_le banditrlproof.ucb.pullcount_armstreamaction_pos_of_k_le after initialization every arm has positive prior pull count. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.K_lt_of_one_lt_pullCount_armStreamAction","label":"K_lt_of_one_lt_pullCount_armStreamAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.K_lt_of_one_lt_pullCount_armStreamAction","description":"A second pull cannot occur before the round-robin cycle is complete.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-2f5a06ce3e10","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2795,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:71"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem K_lt_of_one_lt_pullCount_armStreamAction {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (arm : Fin K) (t : Nat) (hcount : 1 < pullCount (armStreamAction hK c stream) arm t) : K < t","missing":[],"search":"k_lt_of_one_lt_pullcount_armstreamaction banditrlproof.ucb.k_lt_of_one_lt_pullcount_armstreamaction a second pull cannot occur before the round-robin cycle is complete. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction_eq_realIndexAction_of_K_le","label":"armStreamAction_eq_realIndexAction_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction_eq_realIndexAction_of_K_le","description":"After initialization the recursive action is the native trace UCB argmax.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-0102698c8d8a","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2796,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:84"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAction_eq_realIndexAction_of_K_le {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (t : Nat) (ht : K <= t) : armStreamAction hK c stream t = realIndexAction hK (armStreamAction hK c stream) (armStreamReward hK c stream) c t","missing":[],"search":"armstreamaction_eq_realindexaction_of_k_le banditrlproof.ucb.armstreamaction_eq_realindexaction_of_k_le after initialization the recursive action is the native trace ucb argmax. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realIndex_le_realIndex_armStreamAction_of_K_le","label":"realIndex_le_realIndex_armStreamAction_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realIndex_le_realIndex_armStreamAction_of_K_le","description":"The selected recursive action maximizes the actual random-width index.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-101f19497129","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2797,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:107"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realIndex_le_realIndex_armStreamAction_of_K_le {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (t : Nat) (ht : K <= t) (arm : Fin K) : realIndex (armStreamAction hK c stream) (armStreamReward hK c stream) c arm t <= realIndex (armStreamAction hK c stream) (armStreamReward hK c stream) c (armStreamAction hK c stream t) t","missing":[],"search":"realindex_le_realindex_armstreamaction_of_k_le banditrlproof.ucb.realindex_le_realindex_armstreamaction_of_k_le the selected recursive action maximizes the actual random-width index. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.meanGap_le_two_realWidth_of_selected","label":"meanGap_le_two_realWidth_of_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.meanGap_le_two_realWidth_of_selected","description":"Good confidence inequalities force the selected arm gap below twice its width.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-08aaec316b42","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2798,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:120"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem meanGap_le_two_realWidth_of_selected {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (t : Nat) (ht : K <= t) (trueMean : Fin K -> Real) (best chosen : Fin K) (hselected : armStreamAction hK c stream t = chosen) (hbest : trueMean best <= realEmpiricalMean (armStreamAction hK c stream) (armStreamReward hK c stream) best t + realWidth (armStreamAction hK c stream) c best t) (hchosen : realEmpiricalMean (armStreamAction hK c stream) (armStreamReward hK c stream) chosen t - realWidth (armStreamAction hK c stream) c chosen t <= trueMean chosen) : meanGap trueMean best chosen <= 2 * realWidth (armStreamAction hK c stream) c chosen t","missing":[],"search":"meangap_le_two_realwidth_of_selected banditrlproof.ucb.meangap_le_two_realwidth_of_selected good confidence inequalities force the selected arm gap below twice its width. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullCount_le_eight_scale_log_div_gap_sq","label":"pullCount_le_eight_scale_log_div_gap_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pullCount_le_eight_scale_log_div_gap_sq","description":"Squaring the good-event gap inequality gives the standard UCB count bound.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-4b7176bd2c3f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2799,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:146"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_eight_scale_log_div_gap_sq {K : Nat} (action : ActionTrace (Fin K)) (arm : Fin K) (scale gap : Real) (t : Nat) (hscale : 0 <= scale) (hgap : 0 < gap) (hcount : 0 < pullCount action arm t) (hgap_le : gap <= 2 * realWidth action scale arm t) : (pullCount action arm t : Real) <= 8 * scale * Real.log ((t + 1 : Nat) : Real) / gap ^ 2","missing":[],"search":"pullcount_le_eight_scale_log_div_gap_sq banditrlproof.ucb.pullcount_le_eight_scale_log_div_gap_sq squaring the good-event gap inequality gives the standard ucb count bound. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realPullThreshold","label":"realPullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realPullThreshold","description":"Real threshold used for the horizon-wide pull-count split.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-1d971d292906","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2800,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:173"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realPullThreshold (c : Real) (sigma2 : NNReal) (gap : Real) (n : Nat) : Real","missing":[],"search":"realpullthreshold banditrlproof.ucb.realpullthreshold real threshold used for the horizon-wide pull-count split. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullThreshold","label":"pullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.pullThreshold","description":"Integer split point: one more than the ceiling of the real threshold.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-bd754e87887e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2801,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:178"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def pullThreshold (c : Real) (sigma2 : NNReal) (gap : Real) (n : Nat) : Nat","missing":[],"search":"pullthreshold banditrlproof.ucb.pullthreshold integer split point: one more than the ceiling of the real threshold. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.indexTail","label":"indexTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.indexTail","description":"The one-sided inverse-power tail budget at time `t`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-c196f599a787","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2802,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:183"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def indexTail (c : Real) (t : Nat) : ENNReal","missing":[],"search":"indextail banditrlproof.ucb.indextail the one-sided inverse-power tail budget at time `t`. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constSum","label":"constSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.constSum","description":"Finite-horizon sum of the source-faithful one-sided index tails.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-0f326be918a4","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2803,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:187"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def constSum (c : Real) (n : Nat) : ENNReal","missing":[],"search":"constsum banditrlproof.ucb.constsum finite-horizon sum of the source-faithful one-sided index tails. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedLargePullCountEvent","label":"selectedLargePullCountEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedLargePullCountEvent","description":"A selected arm whose prior count has crossed the horizon threshold.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-c7b972f8dd3c","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2804,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:191"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedLargePullCountEvent {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (gap : Real) (arm : Fin K) (n t : Nat) : Set (ArmRewardStream K)","missing":[],"search":"selectedlargepullcountevent banditrlproof.ucb.selectedlargepullcountevent a selected arm whose prior count has crossed the horizon threshold. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lowerIndexFailure","label":"lowerIndexFailure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.lowerIndexFailure","description":"Lower-index failure for a fixed arm of the recursive process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-5e1461e5d65e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2805,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:200"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def lowerIndexFailure {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (arm : Fin K) (mean : Real) (t : Nat) : Set (ArmRewardStream K)","missing":[],"search":"lowerindexfailure banditrlproof.ucb.lowerindexfailure lower-index failure for a fixed arm of the recursive process. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.upperIndexFailure","label":"upperIndexFailure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.upperIndexFailure","description":"Upper-index failure for a fixed arm of the recursive process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-569a59f9c1b5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2806,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:214"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def upperIndexFailure {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (arm : Fin K) (mean : Real) (t : Nat) : Set (ArmRewardStream K)","missing":[],"search":"upperindexfailure banditrlproof.ucb.upperindexfailure upper-index failure for a fixed arm of the recursive process. definition compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedLargePullCountEvent_subset_lower_union_upper","label":"selectedLargePullCountEvent_subset_lower_union_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedLargePullCountEvent_subset_lower_union_upper","description":"Crossing the horizon count threshold while selecting a positive-gap arm forces one of the two one-sided index failures used by the arm-stream tail module.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-69fff5afa877","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2807,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:232"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedLargePullCountEvent_subset_lower_union_upper {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) (arm : Fin K) (n t : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hgap : 0 < realKernelGap nu arm) (ht : t < n) : selectedLargePullCountEvent hK c sigma2 (realKernelGap nu arm) arm n t ⊆ lowerIndexFailure hK c sigma2 (ETC.realKernelBestArm hK nu) (realKernelMean nu (ETC.realKernelBestArm hK nu)) t ∪ upperIndexFailure hK c sigma2 arm (realKernelMean nu arm) t","missing":[],"search":"selectedlargepullcountevent_subset_lower_union_upper banditrlproof.ucb.selectedlargepullcountevent_subset_lower_union_upper crossing the horizon count threshold while selecting a positive-gap arm forces one of the two one-sided index failures used by the arm-stream tail module. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedLargePullCountEvent_le_two_mul_indexTail","label":"measure_selectedLargePullCountEvent_le_two_mul_indexTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedLargePullCountEvent_le_two_mul_indexTail","description":"The selected-large event has twice the one-sided inverse-power budget.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-f25578b6f831","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2808,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:375"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedLargePullCountEvent_le_two_mul_indexTail {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n t : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hgap : 0 < realKernelGap nu arm) (ht : t < n) (hsubGBest : HasSubgaussianMGF (fun reward => reward - realKernelMean nu (ETC.realKernelBestArm hK nu)) sigma2 (nu (ETC.realKernelBestArm hK nu))) (hsubGArm : HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : armStreamMeasure nu (selectedLargePullCountEvent hK c sigma2 (realKernelGap nu arm) arm n t) <= 2 * indexTail c t","missing":[],"search":"measure_selectedlargepullcountevent_le_two_mul_indextail banditrlproof.ucb.measure_selectedlargepullcountevent_le_two_mul_indextail the selected-large event has twice the one-sided inverse-power budget. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_selectedLargePullCount_indicator_sum_le_two_mul_constSum","label":"lintegral_selectedLargePullCount_indicator_sum_le_two_mul_constSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_selectedLargePullCount_indicator_sum_le_two_mul_constSum","description":"Finite-time selected-large indicators integrate to twice `constSum`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-0b884217c611","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2809,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:423"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_selectedLargePullCount_indicator_sum_le_two_mul_constSum {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hgap : 0 < realKernelGap nu arm) (hsubGBest : HasSubgaussianMGF (fun reward => reward - realKernelMean nu (ETC.realKernelBestArm hK nu)) sigma2 (nu (ETC.realKernelBestArm hK nu))) (hsubGArm : HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : ∫⁻ stream : ArmRewardStream K, (Finset.range n).sum (fun t : Nat => if armStreamAction hK (c * (sigma2 : Real)) stream t = arm ∧ pullThreshold c sigma2 (realKernelGap nu arm) n <= pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm t then (1 : ENNReal) else 0) ∂armStreamMeasure nu <= 2 * constSum c n","missing":[],"search":"lintegral_selectedlargepullcount_indicator_sum_le_two_mul_constsum banditrlproof.ucb.lintegral_selectedlargepullcount_indicator_sum_le_two_mul_constsum finite-time selected-large indicators integrate to twice `constsum`. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_natCast_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","label":"lintegral_natCast_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_natCast_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","description":"ENNReal expected pull-count bound for one positive-gap arm of the concrete recursive UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-d903b6fe7b54","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2810,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:473"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_natCast_pullCount_armStreamAction_le_threshold_add_two_mul_constSum {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hgap : 0 < realKernelGap nu arm) (hsubGBest : HasSubgaussianMGF (fun reward => reward - realKernelMean nu (ETC.realKernelBestArm hK nu)) sigma2 (nu (ETC.realKernelBestArm hK nu))) (hsubGArm : HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : ∫⁻ stream : ArmRewardStream K, (pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n : ENNReal) ∂armStreamMeasure nu <= (pullThreshold c sigma2 (realKernelGap nu arm) n : ENNReal) + 2 * constSum c n","missing":[],"search":"lintegral_natcast_pullcount_armstreamaction_le_threshold_add_two_mul_constsum banditrlproof.ucb.lintegral_natcast_pullcount_armstreamaction_le_threshold_add_two_mul_constsum ennreal expected pull-count bound for one positive-gap arm of the concrete recursive ucb process. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.indexTail_ne_top","label":"indexTail_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.indexTail_ne_top","description":"Every finite inverse-power tail budget is finite.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-cf4c3ef03706","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2811,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:538"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem indexTail_ne_top (c : Real) (t : Nat) : indexTail c t ≠ ∞","missing":[],"search":"indextail_ne_top banditrlproof.ucb.indextail_ne_top every finite inverse-power tail budget is finite. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.constSum_ne_top","label":"constSum_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.constSum_ne_top","description":"The finite-horizon inverse-power tail sum is finite.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-e39642f71194","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2812,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:543"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem constSum_ne_top (c : Real) (n : Nat) : constSum c n ≠ ∞","missing":[],"search":"constsum_ne_top banditrlproof.ucb.constsum_ne_top the finite-horizon inverse-power tail sum is finite. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integrable_real_pullCount_armStreamAction","label":"integrable_real_pullCount_armStreamAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integrable_real_pullCount_armStreamAction","description":"Real pull counts of the measurable recursive process are integrable.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-2780b290deb3","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2813,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:548"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integrable_real_pullCount_armStreamAction {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n : Nat) : Integrable (fun stream : ArmRewardStream K => (pullCount (armStreamAction hK c stream) arm n : Real)) (armStreamMeasure nu)","missing":[],"search":"integrable_real_pullcount_armstreamaction banditrlproof.ucb.integrable_real_pullcount_armstreamaction real pull counts of the measurable recursive process are integrable. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","label":"integral_real_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","description":"Real Bochner expected pull-count bound obtained from the ENNReal endpoint.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-05c53ab30cfc","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2814,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:567"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_armStreamAction_le_threshold_add_two_mul_constSum {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hgap : 0 < realKernelGap nu arm) (hsubGBest : HasSubgaussianMGF (fun reward => reward - realKernelMean nu (ETC.realKernelBestArm hK nu)) sigma2 (nu (ETC.realKernelBestArm hK nu))) (hsubGArm : HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : ∫ stream : ArmRewardStream K, (pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n : Real) ∂armStreamMeasure nu <= (pullThreshold c sigma2 (realKernelGap nu arm) n : Real) + 2 * (constSum c n).toReal","missing":[],"search":"integral_real_pullcount_armstreamaction_le_threshold_add_two_mul_constsum banditrlproof.ucb.integral_real_pullcount_armstreamaction_le_threshold_add_two_mul_constsum real bochner expected pull-count bound obtained from the ennreal endpoint. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullThreshold_cast_le_realPullThreshold_add_two","label":"pullThreshold_cast_le_realPullThreshold_add_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pullThreshold_cast_le_realPullThreshold_add_two","description":"The ceiling threshold is bounded by the exact LML real threshold plus two.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-ca931f0cfc33","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2815,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:635"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pullThreshold_cast_le_realPullThreshold_add_two (c : Real) (sigma2 : NNReal) (gap : Real) (n : Nat) (hthreshold : 0 <= realPullThreshold c sigma2 gap n) : (pullThreshold c sigma2 gap n : Real) <= realPullThreshold c sigma2 gap n + 2","missing":[],"search":"pullthreshold_cast_le_realpullthreshold_add_two banditrlproof.ucb.pullthreshold_cast_le_realpullthreshold_add_two the ceiling threshold is bounded by the exact lml real threshold plus two. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pullCount_armStreamAction_le_realThreshold_add_two_add_two_mul_constSum","label":"integral_real_pullCount_armStreamAction_le_realThreshold_add_two_add_two_mul_constSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pullCount_armStreamAction_le_realThreshold_add_two_add_two_mul_constSum","description":"LML-shaped Real expected pull-count bound without a ceiling term.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-083b83feba82","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2816,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:646"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pullCount_armStreamAction_le_realThreshold_add_two_add_two_mul_constSum {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hgap : 0 < realKernelGap nu arm) (hsubGBest : HasSubgaussianMGF (fun reward => reward - realKernelMean nu (ETC.realKernelBestArm hK nu)) sigma2 (nu (ETC.realKernelBestArm hK nu))) (hsubGArm : HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : ∫ stream : ArmRewardStream K, (pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n : Real) ∂armStreamMeasure nu <= realPullThreshold c sigma2 (realKernelGap nu arm) n + 2 + 2 * (constSum c n).toReal","missing":[],"search":"integral_real_pullcount_armstreamaction_le_realthreshold_add_two_add_two_mul_constsum banditrlproof.ucb.integral_real_pullcount_armstreamaction_le_realthreshold_add_two_add_two_mul_constsum lml-shaped real expected pull-count bound without a ceiling term. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_sum_gap_mul_realThreshold_add_two_add_two_mul_constSum","label":"integral_realKernelRegret_armStreamAction_le_sum_gap_mul_realThreshold_add_two_add_two_mul_constSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_sum_gap_mul_realThreshold_add_two_add_two_mul_constSum","description":"Finite-arm expected regret bound for the concrete recursive arm-stream UCB process, in the same gap-weighted shape as the pinned LML route.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-016f4d890653","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2817,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:675"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_armStreamAction_le_sum_gap_mul_realThreshold_add_two_add_two_mul_constSum {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : ∫ stream : ArmRewardStream K, realKernelRegret nu (armStreamAction hK (c * (sigma2 : Real)) stream) n ∂armStreamMeasure nu <= (Finset.univ : Finset (Fin K)).sum (fun arm => realKernelGap nu arm * (realPullThreshold c sigma2 (realKernelGap nu arm) n + 2 + 2 * (constSum c n).toReal))","missing":[],"search":"integral_realkernelregret_armstreamaction_le_sum_gap_mul_realthreshold_add_two_add_two_mul_constsum banditrlproof.ucb.integral_realkernelregret_armstreamaction_le_sum_gap_mul_realthreshold_add_two_add_two_mul_constsum finite-arm expected regret bound for the concrete recursive arm-stream ucb process, in the same gap-weighted shape as the pinned lml route. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_lml_sum","label":"integral_realKernelRegret_armStreamAction_le_lml_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_lml_sum","description":"Canonical arm-stream specialization of the pinned LML UCB regret theorem, with exactly the upstream gap-weighted finite-sum right-hand side.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-36887a320e3f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2818,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:718"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_armStreamAction_le_lml_sum {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : ∫ stream : ArmRewardStream K, realKernelRegret nu (armStreamAction hK (c * (sigma2 : Real)) stream) n ∂armStreamMeasure nu <= (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * c * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm + realKernelGap nu arm * (2 + 2 * (constSum c n).toReal))","missing":[],"search":"integral_realkernelregret_armstreamaction_le_lml_sum banditrlproof.ucb.integral_realkernelregret_armstreamaction_le_lml_sum canonical arm-stream specialization of the pinned lml ucb regret theorem, with exactly the upstream gap-weighted finite-sum right-hand side. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamActionTrace","label":"measurable_armStreamActionTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamActionTrace","description":"The complete recursive arm-stream action trace is measurable.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-aa2717b6c784","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2819,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:743"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamActionTrace {K : Nat} (hK : 0 < K) (c : Real) : Measurable (armStreamAction hK c : ArmRewardStream K -> ActionTrace (Fin K))","missing":[],"search":"measurable_armstreamactiontrace banditrlproof.ucb.measurable_armstreamactiontrace the complete recursive arm-stream action trace is measurable. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realKernelRegret_actionTrace","label":"measurable_realKernelRegret_actionTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realKernelRegret_actionTrace","description":"Kernel regret is a measurable functional of a finite-arm action trace.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-1db9b08490c4","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2820,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:750"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realKernelRegret_actionTrace {K : Nat} (nu : Kernel (Fin K) Real) (n : Nat) : Measurable (fun action : ActionTrace (Fin K) => realKernelRegret nu action n)","missing":[],"search":"measurable_realkernelregret_actiontrace banditrlproof.ucb.measurable_realkernelregret_actiontrace kernel regret is a measurable functional of a finite-arm action trace. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStreamAction","label":"integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStreamAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStreamAction","description":"Any external action process with the same complete action-trace law as the canonical arm-stream UCB process inherits its exact LML-shaped regret bound.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-a63bd2fc05c0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2821,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:769"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStreamAction {Omega : Type*} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (hident : IdentDistrib action (armStreamAction hK (c * (sigma2 : Real))) mu (armStreamMeasure nu)) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : integral mu (fun omega => realKernelRegret nu (action omega) n) <= (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * c * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm + realKernelGap nu arm * (2 + 2 * (constSum c n).toReal))","missing":[],"search":"integral_realkernelregret_externalaction_le_lml_sum_of_identdistrib_armstreamaction banditrlproof.ucb.integral_realkernelregret_externalaction_le_lml_sum_of_identdistrib_armstreamaction any external action process with the same complete action-trace law as the canonical arm-stream ucb process inherits its exact lml-shaped regret bound. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.identDistrib_action_armStreamAction_of_identDistrib_armStream","label":"identDistrib_action_armStreamAction_of_identDistrib_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.identDistrib_action_armStreamAction_of_identDistrib_armStream","description":"An external action generated almost surely by the recursive UCB map inherits the canonical action-trace law from an identically distributed latent arm stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-68e504c3cd0e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2822,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:801"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_action_armStreamAction_of_identDistrib_armStream {Omega : Type*} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (c : Real) (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (armStream : Omega -> ArmRewardStream K) (action : Omega -> ActionTrace (Fin K)) (hstreamLaw : IdentDistrib armStream (id : ArmRewardStream K -> ArmRewardStream K) mu (armStreamMeasure nu)) (haction : ∀ᵐ omega ∂mu, action omega = armStreamAction hK c (armStream omega)) : IdentDistrib action (armStreamAction hK c) mu (armStreamMeasure nu)","missing":[],"search":"identdistrib_action_armstreamaction_of_identdistrib_armstream banditrlproof.ucb.identdistrib_action_armstreamaction_of_identdistrib_armstream an external action generated almost surely by the recursive ucb map inherits the canonical action-trace law from an identically distributed latent arm stream. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStream","label":"integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStream","description":"Exact LML-shaped regret for an external action generated from a latent arm stream with the canonical complete stream law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-e6b8f9bb8437","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2823,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:829"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStream {Omega : Type*} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (armStream : Omega -> ArmRewardStream K) (action : Omega -> ActionTrace (Fin K)) (hstreamLaw : IdentDistrib armStream (id : ArmRewardStream K -> ArmRewardStream K) mu (armStreamMeasure nu)) (haction : ∀ᵐ omega ∂mu, action omega = armStreamAction hK (c * (sigma2 : Real)) (armStream omega)) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun reward => reward - realKernelMean nu arm) sigma2 (nu arm)) : integral mu (fun omega => realKernelRegret nu (action omega) n) <= (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * c * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm + realKer…","missing":[],"search":"integral_realkernelregret_externalaction_le_lml_sum_of_identdistrib_armstream banditrlproof.ucb.integral_realkernelregret_externalaction_le_lml_sum_of_identdistrib_armstream exact lml-shaped regret for an external action generated from a latent arm stream with the canonical complete stream law. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.identDistrib_action_of_identDistrib_actionRewardTrace","label":"identDistrib_action_of_identDistrib_actionRewardTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.identDistrib_action_of_identDistrib_actionRewardTrace","description":"Identical laws of complete action/reward trajectories imply identical laws of their action traces. This is the projection used by LML's `IsAlgEnvSeq.identDistrib_trajectory` route.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-e08cd6da8c13","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2824,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:862"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_action_of_identDistrib_actionRewardTrace {Omega Xi : Type*} [MeasurableSpace Omega] [MeasurableSpace Xi] {K : Nat} (mu : Measure Omega) (mu' : Measure Xi) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (action' : Xi -> ActionTrace (Fin K)) (reward' : Xi -> RewardTrace Real) (htrajectory : IdentDistrib (fun omega t => (action omega t, reward omega t)) (fun xi t => (action' xi t, reward' xi t)) mu mu') : IdentDistrib action action' mu mu'","missing":[],"search":"identdistrib_action_of_identdistrib_actionrewardtrace banditrlproof.ucb.identdistrib_action_of_identdistrib_actionrewardtrace identical laws of complete action/reward trajectories imply identical laws of their action traces. this is the projection used by lml's `isalgenvseq.identdistrib_trajectory` route. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_actionRewardTrace","label":"integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_actionRewardTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_actionRewardTrace","description":"Exact LML-shaped regret transported from an observable action/reward trajectory law matching the canonical recursive arm-stream UCB trajectory.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-4c968691ecfd","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2825,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:885"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_actionRewardTrace {Omega : Type*} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (htrajectory : IdentDistrib (fun omega t => (action omega t, reward omega t)) (fun stream t => (armStreamAction hK (c * (sigma2 : Real)) stream t, armStreamReward hK (c * (sigma2 : Real)) stream t)) mu (armStreamMeasure nu)) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun x => x - realKernelMean nu arm) sigma2 (nu arm)) : integral mu (fun omega => realKernelRegret nu (action omega) n) <= (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * c * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm…","missing":[],"search":"integral_realkernelregret_externalaction_le_lml_sum_of_identdistrib_actionrewardtrace banditrlproof.ucb.integral_realkernelregret_externalaction_le_lml_sum_of_identdistrib_actionrewardtrace exact lml-shaped regret transported from an observable action/reward trajectory law matching the canonical recursive arm-stream ucb trajectory. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_common_actionReward_condDistrib","label":"integral_realKernelRegret_externalAction_le_lml_sum_of_common_actionReward_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_common_actionReward_condDistrib","description":"Exact LML-shaped regret from upstream-style initial and successor action/reward conditional laws. The external process and the canonical arm-stream UCB process need only share the same initial pair marginal and the same history-indexed successor pair kernels; full trajectory `IdentDistrib` is then supplied by Ionescu-Tulcea/projective-limit uniqueness.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-166994aff9d0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2826,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:922"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_lml_sum_of_common_actionReward_condDistrib {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) [IsFiniteMeasure mu] (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (mu0 : Measure (Fin K × Real)) [IsProbabilityMeasure mu0] (pairKernel : (i : Nat) -> Kernel (History.FinitePairHistory (Fin K) Real i) (Fin K × Real)) [forall i, IsMarkovKernel (pairKernel i)] (hzero : Measure.map (fun omega : Omega => (action omega 0, reward omega 0)) mu = mu0) (hzeroCanonical : Measure.map (fun stream : ArmRewardStream K => (armStreamAction hK (c * (sigma2 : R…","missing":[],"search":"integral_realkernelregret_externalaction_le_lml_sum_of_common_actionreward_conddistrib banditrlproof.ucb.integral_realkernelregret_externalaction_le_lml_sum_of_common_actionreward_conddistrib exact lml-shaped regret from upstream-style initial and successor action/reward conditional laws. the external process and the canonical arm-stream ucb process need only share the same initial pair marginal and the same history-indexed successor pair kernels; full trajectory `identdistrib` is then supplied by ionescu-tulcea/projective-limit uniqueness. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.identDistrib_actionRewardTrace_of_condDistrib_eq_armStream","label":"identDistrib_actionRewardTrace_of_condDistrib_eq_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.identDistrib_actionRewardTrace_of_condDistrib_eq_armStream","description":"An external action/reward process has the same complete observable trajectory law as canonical arm-stream UCB when its initial pair marginal and every successor pair conditional distribution agree with the corresponding canonical ones. The canonical initial measure and kernel family are chosen internally.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-05eeb619042a","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2827,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:1011"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_actionRewardTrace_of_condDistrib_eq_armStream {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) [IsFiniteMeasure mu] (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hzero : Measure.map (fun omega : Omega => (action omega 0, reward omega 0)) mu = Measure.map (fun stream : ArmRewardStream K => (armStreamAction hK (c * (sigma2 : Real)) stream 0, armStreamReward hK (c * (sigma2 : Real)) stream 0)) (armStreamMeasure nu)) (hcond : forall i : Nat, condDistrib (fun omega : Omega => (action omega (i + 1), reward omega (i + 1))) (fun omega : Omega => History.finitePairHistoryO…","missing":[],"search":"identdistrib_actionrewardtrace_of_conddistrib_eq_armstream banditrlproof.ucb.identdistrib_actionrewardtrace_of_conddistrib_eq_armstream an external action/reward process has the same complete observable trajectory law as canonical arm-stream ucb when its initial pair marginal and every successor pair conditional distribution agree with the corresponding canonical ones. the canonical initial measure and kernel family are chosen internally. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.identDistrib_actionRewardTrace_of_split_condDistrib_eq_armStream","label":"identDistrib_actionRewardTrace_of_split_condDistrib_eq_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.identDistrib_actionRewardTrace_of_split_condDistrib_eq_armStream","description":"An external process has the canonical observable UCB trajectory law from the four split law surfaces used by LML's `IsAlgEnvSeq`: the initial action law, the initial feedback law given that action, the successor action law given finite pair history, and the successor feedback law given history and the next action. The canonical split kernels and their `compProd` pair kernels are chosen internally.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-e9033fe2c1ab","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2828,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:1097"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_actionRewardTrace_of_split_condDistrib_eq_armStream {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) [IsFiniteMeasure mu] (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hzeroAction : Measure.map (fun omega : Omega => action omega 0) mu = Measure.map (fun stream : ArmRewardStream K => armStreamAction hK (c * (sigma2 : Real)) stream 0) (armStreamMeasure nu)) (hzeroFeedback : condDistrib (fun omega : Omega => reward omega 0) (fun omega : Omega => action omega 0) mu =ᵐ[ mu.map (fun omega : Omega => action omega 0)] condDistrib (fun stream : ArmRewardStream K => armStre…","missing":[],"search":"identdistrib_actionrewardtrace_of_split_conddistrib_eq_armstream banditrlproof.ucb.identdistrib_actionrewardtrace_of_split_conddistrib_eq_armstream an external process has the canonical observable ucb trajectory law from the four split law surfaces used by lml's `isalgenvseq`: the initial action law, the initial feedback law given that action, the successor action law given finite pair history, and the successor feedback law given history and the next action. the canonical split kernels and their `compprod` pair kernels are chosen internally. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_split_condDistrib_eq_armStream","label":"integral_realKernelRegret_externalAction_le_lml_sum_of_split_condDistrib_eq_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_split_condDistrib_eq_armStream","description":"Exact canonical arm-stream UCB regret from the four split action/feedback law fields corresponding to LML's `IsAlgEnvSeq` interface.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-895a04ca4eac","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2829,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:1251"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_lml_sum_of_split_condDistrib_eq_armStream {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) [IsFiniteMeasure mu] (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hzeroAction : Measure.map (fun omega : Omega => action omega 0) mu = Measure.map (fun stream : ArmRewardStream K => armStreamAction hK (c * (sigma2 : Real)) stream 0) (armStreamMeasure nu)) (hzeroFeedback : condDistrib (fun omega : Omega => reward omega 0) (fun omega : Omega => action omega 0) mu =ᵐ[ mu.map (fun omega : Omega => action omega 0)] condDistrib (fun stream : ArmRewa…","missing":[],"search":"integral_realkernelregret_externalaction_le_lml_sum_of_split_conddistrib_eq_armstream banditrlproof.ucb.integral_realkernelregret_externalaction_le_lml_sum_of_split_conddistrib_eq_armstream exact canonical arm-stream ucb regret from the four split action/feedback law fields corresponding to lml's `isalgenvseq` interface. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_condDistrib_eq_armStream","label":"integral_realKernelRegret_externalAction_le_lml_sum_of_condDistrib_eq_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_condDistrib_eq_armStream","description":"Exact LML-shaped regret when the external initial and successor observable pair laws agree directly with canonical arm-stream UCB.","url":"../modules/banditrlproof-algorithms-ucbarmstreamexpectedpullcount/index.html#decl-76de36355e66","parent":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","order":2830,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount"],["Source","BanditRLProof/Algorithms/UCBArmStreamExpectedPullCount.lean:1332"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_externalAction_le_lml_sum_of_condDistrib_eq_armStream {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (mu : Measure Omega) [IsFiniteMeasure mu] (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hzero : Measure.map (fun omega : Omega => (action omega 0, reward omega 0)) mu = Measure.map (fun stream : ArmRewardStream K => (armStreamAction hK (c * (sigma2 : Real)) stream 0, armStreamReward hK (c * (sigma2 : Real)) stream 0)) (armStreamMeasure nu)) (hcond : forall i : Nat, condDistrib (fun omega : Omega => (action omega (i + 1), reward omega (i + 1))) (fun omega : Omega => Histo…","missing":[],"search":"integral_realkernelregret_externalaction_le_lml_sum_of_conddistrib_eq_armstream banditrlproof.ucb.integral_realkernelregret_externalaction_le_lml_sum_of_conddistrib_eq_armstream exact lml-shaped regret when the external initial and successor observable pair laws agree directly with canonical arm-stream ucb. theorem compiled","shard":"modules/bf6ede8be12dac50.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmRealRewardKernel","label":"finiteArmRealRewardKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmRealRewardKernel","description":"A finite family of Real reward laws, viewed as an arm-indexed kernel.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-8f26157bbc57","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2831,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:21"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmRealRewardKernel {K : Nat} (armLaw : Fin K -> Measure Real) : Kernel (Fin K) Real","missing":[],"search":"finitearmrealrewardkernel banditrlproof.ucb.finitearmrealrewardkernel a finite family of real reward laws, viewed as an arm-indexed kernel. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmRealRewardKernel_apply","label":"finiteArmRealRewardKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmRealRewardKernel_apply","description":"theorem finiteArmRealRewardKernel_apply {K : Nat} (armLaw : Fin K -> Measure Real) (arm : Fin K) : finiteArmRealRewardKernel armLaw arm = armLaw arm","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-1e8bf056d398","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2832,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:26"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem finiteArmRealRewardKernel_apply {K : Nat} (armLaw : Fin K -> Measure Real) (arm : Fin K) : finiteArmRealRewardKernel armLaw arm = armLaw arm","missing":[],"search":"finitearmrealrewardkernel_apply banditrlproof.ucb.finitearmrealrewardkernel_apply theorem finitearmrealrewardkernel_apply {k : nat} (armlaw : fin k -> measure real) (arm : fin k) : finitearmrealrewardkernel armlaw arm = armlaw arm theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmRealRewardKernel_isMarkov","label":"finiteArmRealRewardKernel_isMarkov","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmRealRewardKernel_isMarkov","description":"Pointwise probability laws make the finite-arm kernel Markov.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-8a9ef4c1aa2f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2833,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:32"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem finiteArmRealRewardKernel_isMarkov {K : Nat} (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) : IsMarkovKernel (finiteArmRealRewardKernel armLaw) where","missing":[],"search":"finitearmrealrewardkernel_ismarkov banditrlproof.ucb.finitearmrealrewardkernel_ismarkov pointwise probability laws make the finite-arm kernel markov. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedRegret","label":"armStreamFiniteArmSubgaussianExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedRegret","description":"Expected regret of the canonical one-policy arm-stream process built from direct finite-arm sub-Gaussian laws. The common tuning proxy is the padded finite maximum of the genuine armwise proxies.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-2a28cdbbd3f5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2834,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:44"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamFiniteArmSubgaussianExpectedRegret {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (n : Nat) : Real","missing":[],"search":"armstreamfinitearmsubgaussianexpectedregret banditrlproof.ucb.armstreamfinitearmsubgaussianexpectedregret expected regret of the canonical one-policy arm-stream process built from direct finite-arm sub-gaussian laws. the common tuning proxy is the padded finite maximum of the genuine armwise proxies. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedRegret_nonneg_and_le","label":"armStreamFiniteArmSubgaussianExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedRegret_nonneg_and_le","description":"Direct finite-arm sub-Gaussian laws satisfy the fixed logarithmic envelope.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-ed584444b9b9","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2835,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:55"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamFiniteArmSubgaussianExpectedRegret_nonneg_and_le {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Real => reward - integral (armLaw arm) id) (varianceProxy arm) (armLaw arm)) (n : Nat) : let nu := finiteArmRealRewardKernel armLaw let sigma2 := Concentration.finiteArmPositiveVarianceProxy varianceProxy 0 <= armStreamFiniteArmSubgaussianExpectedRegret hK armLaw hprob varianceProxy n /\\ armStreamFiniteArmSubgaussianExpectedRegret hK armLaw hprob varianceProxy n <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Real.log ((n + 1 : Nat) : Real))","missing":[],"search":"armstreamfinitearmsubgaussianexpectedregret_nonneg_and_le banditrlproof.ucb.armstreamfinitearmsubgaussianexpectedregret_nonneg_and_le direct finite-arm sub-gaussian laws satisfy the fixed logarithmic envelope. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedAverageRegret","label":"armStreamFiniteArmSubgaussianExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedAverageRegret","description":"Direct finite-arm sub-Gaussian expected regret normalized by `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-42a8e55b1649","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2836,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:95"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamFiniteArmSubgaussianExpectedAverageRegret {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (n : Nat) : Real","missing":[],"search":"armstreamfinitearmsubgaussianexpectedaverageregret banditrlproof.ucb.armstreamfinitearmsubgaussianexpectedaverageregret direct finite-arm sub-gaussian expected regret normalized by `n + 1`. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedAverageRegret_tendsto_zero","label":"armStreamFiniteArmSubgaussianExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedAverageRegret_tendsto_zero","description":"Direct stationary finite-arm sub-Gaussian laws instantiate one fixed canonical arm-stream policy with vanishing expected average regret.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-412748e7e3f0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2837,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:108"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamFiniteArmSubgaussianExpectedAverageRegret_tendsto_zero {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Real => reward - integral (armLaw arm) id) (varianceProxy arm) (armLaw arm)) : Tendsto (armStreamFiniteArmSubgaussianExpectedAverageRegret hK armLaw hprob varianceProxy) atTop (nhds 0)","missing":[],"search":"armstreamfinitearmsubgaussianexpectedaverageregret_tendsto_zero banditrlproof.ucb.armstreamfinitearmsubgaussianexpectedaverageregret_tendsto_zero direct stationary finite-arm sub-gaussian laws instantiate one fixed canonical arm-stream policy with vanishing expected average regret. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedRegret","label":"armStreamBoundedFiniteArmExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedRegret","description":"Expected regret of the one-policy process for common-bounded arm laws.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-6911c417815b","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2838,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:145"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamBoundedFiniteArmExpectedRegret {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (n : Nat) : Real","missing":[],"search":"armstreamboundedfinitearmexpectedregret banditrlproof.ucb.armstreamboundedfinitearmexpectedregret expected regret of the one-policy process for common-bounded arm laws. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedRegret_nonneg_and_le","label":"armStreamBoundedFiniteArmExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedRegret_nonneg_and_le","description":"Common-bounded finite-arm laws satisfy the fixed logarithmic envelope.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-2f6fe07074c1","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2839,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:154"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamBoundedFiniteArmExpectedRegret_nonneg_and_le {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hbound : forall arm, Filter.Eventually (fun reward : Real => Set.Icc lo hi reward) (ae (armLaw arm))) (n : Nat) : let nu := finiteArmRealRewardKernel armLaw let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun _ : Fin K => Concentration.intervalVarianceProxy lo hi) 0 <= armStreamBoundedFiniteArmExpectedRegret hK armLaw hprob lo hi n /\\ armStreamBoundedFiniteArmExpectedRegret hK armLaw hprob lo hi n <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Real.log ((n + 1 : Nat) : Real))","missing":[],"search":"armstreamboundedfinitearmexpectedregret_nonneg_and_le banditrlproof.ucb.armstreamboundedfinitearmexpectedregret_nonneg_and_le common-bounded finite-arm laws satisfy the fixed logarithmic envelope. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedAverageRegret","label":"armStreamBoundedFiniteArmExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedAverageRegret","description":"Expected common-bounded one-policy regret normalized by `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-658c55f35353","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2840,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:188"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamBoundedFiniteArmExpectedAverageRegret {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (n : Nat) : Real","missing":[],"search":"armstreamboundedfinitearmexpectedaverageregret banditrlproof.ucb.armstreamboundedfinitearmexpectedaverageregret expected common-bounded one-policy regret normalized by `n + 1`. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedAverageRegret_tendsto_zero","label":"armStreamBoundedFiniteArmExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedAverageRegret_tendsto_zero","description":"Stationary finite-arm Real reward laws bounded almost surely in one common interval induce one fixed canonical UCB process with vanishing expected average regret. The positive padded tuning proxy removes any `lo < hi` premise.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-ba94709ffe5a","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2841,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:201"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamBoundedFiniteArmExpectedAverageRegret_tendsto_zero {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hbound : forall arm, Filter.Eventually (fun reward : Real => Set.Icc lo hi reward) (ae (armLaw arm))) : Tendsto (armStreamBoundedFiniteArmExpectedAverageRegret hK armLaw hprob lo hi) atTop (nhds 0)","missing":[],"search":"armstreamboundedfinitearmexpectedaverageregret_tendsto_zero banditrlproof.ucb.armstreamboundedfinitearmexpectedaverageregret_tendsto_zero stationary finite-arm real reward laws bounded almost surely in one common interval induce one fixed canonical ucb process with vanishing expected average regret. the positive padded tuning proxy removes any `lo < hi` premise. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedRegret","label":"armStreamArmwiseBoundedFiniteArmExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedRegret","description":"Expected regret of the one-policy process for armwise-bounded laws.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-84d81f5de2ef","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2842,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:231"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamArmwiseBoundedFiniteArmExpectedRegret {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (n : Nat) : Real","missing":[],"search":"armstreamarmwiseboundedfinitearmexpectedregret banditrlproof.ucb.armstreamarmwiseboundedfinitearmexpectedregret expected regret of the one-policy process for armwise-bounded laws. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","label":"armStreamArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","description":"Armwise-bounded finite-arm laws satisfy the fixed logarithmic envelope.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-666e18f6265a","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2843,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:240"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hbound : forall arm, Filter.Eventually (fun reward : Real => Set.Icc (lo arm) (hi arm) reward) (ae (armLaw arm))) (n : Nat) : let nu := finiteArmRealRewardKernel armLaw let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) 0 <= armStreamArmwiseBoundedFiniteArmExpectedRegret hK armLaw hprob lo hi n /\\ armStreamArmwiseBoundedFiniteArmExpectedRegret hK armLaw hprob lo hi n <= armStreamAsymptoticModelCoefficient nu sigma2 * (1 + Real.log ((n + 1 : Nat) : Real))","missing":[],"search":"armstreamarmwiseboundedfinitearmexpectedregret_nonneg_and_le banditrlproof.ucb.armstreamarmwiseboundedfinitearmexpectedregret_nonneg_and_le armwise-bounded finite-arm laws satisfy the fixed logarithmic envelope. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedAverageRegret","label":"armStreamArmwiseBoundedFiniteArmExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedAverageRegret","description":"Expected armwise-bounded one-policy regret normalized by `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-6fa58f832b27","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2844,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:279"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamArmwiseBoundedFiniteArmExpectedAverageRegret {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (n : Nat) : Real","missing":[],"search":"armstreamarmwiseboundedfinitearmexpectedaverageregret banditrlproof.ucb.armstreamarmwiseboundedfinitearmexpectedaverageregret expected armwise-bounded one-policy regret normalized by `n + 1`. definition compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","label":"armStreamArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","description":"Stationary finite-arm Real reward laws with arm-dependent almost-sure interval bounds induce one fixed canonical UCB process with vanishing expected average regret. Positive padding removes every pointwise interval-order premise.","url":"../modules/banditrlproof-algorithms-ucbarmstreamfinitearmrewardlaws/index.html#decl-e89be8fe31a1","parent":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","order":2845,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBArmStreamFiniteArmRewardLaws.lean:293"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero {K : Nat} (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hbound : forall arm, Filter.Eventually (fun reward : Real => Set.Icc (lo arm) (hi arm) reward) (ae (armLaw arm))) : Tendsto (armStreamArmwiseBoundedFiniteArmExpectedAverageRegret hK armLaw hprob lo hi) atTop (nhds 0)","missing":[],"search":"armstreamarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero banditrlproof.ucb.armstreamarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero stationary finite-arm real reward laws with arm-dependent almost-sure interval bounds induce one fixed canonical ucb process with vanishing expected average regret. positive padding removes every pointwise interval-order premise. theorem compiled","shard":"modules/1b9fd22e2fff11f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.initializationArm","label":"initializationArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.initializationArm","description":"Round-robin arm at an absolute time coordinate.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-a1dbe2ee4d7f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2846,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:20"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def initializationArm {K : Nat} (hK : 0 < K) (t : Nat) : Fin K","missing":[],"search":"initializationarm banditrlproof.ucb.initializationarm round-robin arm at an absolute time coordinate. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.initializationArm_zero","label":"initializationArm_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.initializationArm_zero","description":"theorem initializationArm_zero {K : Nat} (hK : 0 < K) : initializationArm hK 0 = Fin.mk 0 hK","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-9d65ceb37a73","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2847,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:24"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem initializationArm_zero {K : Nat} (hK : 0 < K) : initializationArm hK 0 = Fin.mk 0 hK","missing":[],"search":"initializationarm_zero banditrlproof.ucb.initializationarm_zero theorem initializationarm_zero {k : nat} (hk : 0 < k) : initializationarm hk 0 = fin.mk 0 hk theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryNextArm","label":"realHistoryNextArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryNextArm","description":"LML-shaped next action after an inclusive history through time `n`. The first `K` actions are round robin. Once that prefix is complete, the native Real UCB history index selects a least-encoded maximizing arm.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-392d6c0306a3","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2848,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:34"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistoryNextArm {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"realhistorynextarm banditrlproof.ucb.realhistorynextarm lml-shaped next action after an inclusive history through time `n`. the first `k` actions are round robin. once that prefix is complete, the native real ucb history index selects a least-encoded maximizing arm. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistoryNextArm","label":"measurable_realHistoryNextArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistoryNextArm","description":"The next-arm selector is measurable on the finite history space.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-77ebef884249","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2849,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:43"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistoryNextArm {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => realHistoryNextArm hK c n history)","missing":[],"search":"measurable_realhistorynextarm banditrlproof.ucb.measurable_realhistorynextarm the next-arm selector is measurable on the finite history space. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armRewardStream_apply","label":"measurable_armRewardStream_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armRewardStream_apply","description":"Joint evaluation of an arm stream at countable random coordinates.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-0107a96941a6","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2850,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:53"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armRewardStream_apply {K : Nat} : Measurable (fun input : ArmRewardStream K × (Nat × Fin K) => input.1 input.2.1 input.2.2)","missing":[],"search":"measurable_armrewardstream_apply banditrlproof.ucb.measurable_armrewardstream_apply joint evaluation of an arm stream at countable random coordinates. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistory","label":"armStreamHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistory","description":"Inclusive pair history of the deterministic UCB process driven by a latent arm-reward stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-42499580a929","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2851,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:66"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamHistory {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) : (n : Nat) -> History.FinitePairHistory (Fin K) Real n | 0 => let arm := initializationArm hK 0 fun _ => (arm, stream 0 arm) | n + 1 => let history := armStreamHistory hK c stream n let arm := realHistoryNextArm hK c n history let count := ETC.realHistoryPullCount n history arm History.extendPairHistorySucc history (arm, stream count arm) /-- Action trace extracted from the recursive inclusive histories. -/ noncomputable def armStreamAction {K : Nat} (hK : 0 < K) (c : Real) : ArmRewardStream K -> ActionTrace (Fin K)","missing":[],"search":"armstreamhistory banditrlproof.ucb.armstreamhistory inclusive pair history of the deterministic ucb process driven by a latent arm-reward stream. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction","label":"armStreamAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction","description":"Action trace extracted from the recursive inclusive histories.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-9d1e2fe8a10f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2852,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:80"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamAction {K : Nat} (hK : 0 < K) (c : Real) : ArmRewardStream K -> ActionTrace (Fin K)","missing":[],"search":"armstreamaction banditrlproof.ucb.armstreamaction action trace extracted from the recursive inclusive histories. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamReward","label":"armStreamReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamReward","description":"Reward trace obtained by reading each selected arm's next unused value.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-b76a1204af2e","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2853,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:88"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamReward {K : Nat} (hK : 0 < K) (c : Real) : ArmRewardStream K -> RewardTrace Real","missing":[],"search":"armstreamreward banditrlproof.ucb.armstreamreward reward trace obtained by reading each selected arm's next unused value. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction_zero","label":"armStreamAction_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction_zero","description":"theorem armStreamAction_zero {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) : armStreamAction hK c stream 0 = initializationArm hK 0","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-da2148cfa15f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2854,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:94"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAction_zero {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) : armStreamAction hK c stream 0 = initializationArm hK 0","missing":[],"search":"armstreamaction_zero banditrlproof.ucb.armstreamaction_zero theorem armstreamaction_zero {k : nat} (hk : 0 < k) (c : real) (stream : armrewardstream k) : armstreamaction hk c stream 0 = initializationarm hk 0 theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction_succ","label":"armStreamAction_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction_succ","description":"theorem armStreamAction_succ {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) : armStreamAction hK c stream (n + 1) = realHistoryNextArm hK c n (armStreamHistory hK c stream n)","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-77db75e6916a","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2855,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:100"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAction_succ {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) : armStreamAction hK c stream (n + 1) = realHistoryNextArm hK c n (armStreamHistory hK c stream n)","missing":[],"search":"armstreamaction_succ banditrlproof.ucb.armstreamaction_succ theorem armstreamaction_succ {k : nat} (hk : 0 < k) (c : real) (stream : armrewardstream k) (n : nat) : armstreamaction hk c stream (n + 1) = realhistorynextarm hk c n (armstreamhistory hk c stream n) theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamHistory_eq_finitePairHistoryOfTrace","label":"armStreamHistory_eq_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamHistory_eq_finitePairHistoryOfTrace","description":"The recursively maintained history is exactly the finite pair history of the extracted action and next-unused-coordinate reward traces.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-7681beb6b0ab","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2856,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:111"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamHistory_eq_finitePairHistoryOfTrace {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) : armStreamHistory hK c stream n = History.finitePairHistoryOfTrace (armStreamAction hK c stream) (armStreamReward hK c stream) n","missing":[],"search":"armstreamhistory_eq_finitepairhistoryoftrace banditrlproof.ucb.armstreamhistory_eq_finitepairhistoryoftrace the recursively maintained history is exactly the finite pair history of the extracted action and next-unused-coordinate reward traces. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamHistory","label":"measurable_armStreamHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamHistory","description":"Every recursively generated inclusive history is measurable in the stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-aecd439b8572","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2857,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:141"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamHistory {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (fun stream : ArmRewardStream K => armStreamHistory hK c stream n)","missing":[],"search":"measurable_armstreamhistory banditrlproof.ucb.measurable_armstreamhistory every recursively generated inclusive history is measurable in the stream. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamAction","label":"measurable_armStreamAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamAction","description":"Every action coordinate of the recursive arm-stream process is measurable.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-072b653828b3","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2858,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:217"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamAction {K : Nat} (hK : 0 < K) (c : Real) (t : Nat) : Measurable (fun stream : ArmRewardStream K => armStreamAction hK c stream t)","missing":[],"search":"measurable_armstreamaction banditrlproof.ucb.measurable_armstreamaction every action coordinate of the recursive arm-stream process is measurable. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armStreamReward","label":"measurable_armStreamReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armStreamReward","description":"Every reward coordinate of the recursive arm-stream process is measurable.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-9a10ca83bbe7","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2859,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:227"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armStreamReward {K : Nat} (hK : 0 < K) (c : Real) (t : Nat) : Measurable (fun stream : ArmRewardStream K => armStreamReward hK c stream t)","missing":[],"search":"measurable_armstreamreward banditrlproof.ucb.measurable_armstreamreward every reward coordinate of the recursive arm-stream process is measurable. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction_succ_eq_realHistoryNextArm_actualHistory","label":"armStreamAction_succ_eq_realHistoryNextArm_actualHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction_succ_eq_realHistoryNextArm_actualHistory","description":"The action after time `n` is the LML-shaped selector on its actual history.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-6109e5c21d39","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2860,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:246"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAction_succ_eq_realHistoryNextArm_actualHistory {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) : armStreamAction hK c stream (n + 1) = realHistoryNextArm hK c n (History.finitePairHistoryOfTrace (armStreamAction hK c stream) (armStreamReward hK c stream) n)","missing":[],"search":"armstreamaction_succ_eq_realhistorynextarm_actualhistory banditrlproof.ucb.armstreamaction_succ_eq_realhistorynextarm_actualhistory the action after time `n` is the lml-shaped selector on its actual history. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamAction_succ_eq_realHistoryIndexAction_of_not_lt","label":"armStreamAction_succ_eq_realHistoryIndexAction_of_not_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamAction_succ_eq_realHistoryIndexAction_of_not_lt","description":"After initialization, the generated action is the native Real UCB index action.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-151e03a4216f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2861,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:257"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamAction_succ_eq_realHistoryIndexAction_of_not_lt {K : Nat} (hK : 0 < K) (c : Real) (stream : ArmRewardStream K) (n : Nat) (hn : ¬ n < K - 1) : armStreamAction hK c stream (n + 1) = realHistoryIndexAction hK c n (History.finitePairHistoryOfTrace (armStreamAction hK c stream) (armStreamReward hK c stream) n)","missing":[],"search":"armstreamaction_succ_eq_realhistoryindexaction_of_not_lt banditrlproof.ucb.armstreamaction_succ_eq_realhistoryindexaction_of_not_lt after initialization, the generated action is the native real ucb index action. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamUCBFixedArmPrefixSource","label":"armStreamUCBFixedArmPrefixSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamUCBFixedArmPrefixSource","description":"The recursive process satisfies the fixed-arm latent-prefix source.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-5fcb97d8e16b","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2862,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:269"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamUCBFixedArmPrefixSource {K : Nat} (hK : 0 < K) (c : Real) : FixedArmPrefixSource (armStreamAction hK c) (armStreamReward hK c)","missing":[],"search":"armstreamucbfixedarmprefixsource banditrlproof.ucb.armstreamucbfixedarmprefixsource the recursive process satisfies the fixed-arm latent-prefix source. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure","label":"armStreamMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure","description":"Stationary product law: one independent time stream for every arm law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-8fce645393be","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2863,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:276"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armStreamMeasure {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Measure (ArmRewardStream K)","missing":[],"search":"armstreammeasure banditrlproof.ucb.armstreammeasure stationary product law: one independent time stream for every arm law. definition compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_armStreamUCB_mem_le","label":"measure_pullCount_prod_sumRewards_armStreamUCB_mem_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_prod_sumRewards_armStreamUCB_mem_le","description":"Fixed-count peeling for the actual recursive UCB process under the stationary product arm-stream law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamprocess/index.html#decl-87f297d77d01","parent":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","order":2864,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamProcess"],["Source","BanditRLProof/Algorithms/UCBArmStreamProcess.lean:291"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_prod_sumRewards_armStreamUCB_mem_le {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (n : Nat) (s : Set (Nat × Real)) [DecidablePred (fun k : Nat => k ∈ Prod.fst '' s)] : armStreamMeasure nu {stream | (pullCount (armStreamAction hK c stream) arm n, sumRewards (armStreamAction hK c stream) (armStreamReward hK c stream) arm n) ∈ s} ≤ ((Finset.range (n + 1)).filter (fun k => k ∈ Prod.fst '' s)).sum (fun k => armStreamMeasure nu {stream | armPrefixSum arm k stream ∈ Prod.mk k ⁻¹' s})","missing":[],"search":"measure_pullcount_prod_sumrewards_armstreamucb_mem_le banditrlproof.ucb.measure_pullcount_prod_sumrewards_armstreamucb_mem_le fixed-count peeling for the actual recursive ucb process under the stationary product arm-stream law. theorem compiled","shard":"modules/efbe8812ce2ea351.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.rewardFromArmStream","label":"rewardFromArmStream","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.rewardFromArmStream","description":"Read the next unused latent reward of the arm selected at time `t`. The action trace may depend on the whole sample point. The coordinate index is the number of earlier selections of the currently selected arm.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html#decl-f80e9e674072","parent":"module:BanditRLProof.Algorithms.UCBArmStreamSource","order":2865,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamSource"],["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean:28"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def rewardFromArmStream {Omega : Type u} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) : Omega -> RewardTrace Real","missing":[],"search":"rewardfromarmstream banditrlproof.ucb.rewardfromarmstream read the next unused latent reward of the arm selected at time `t`. the action trace may depend on the whole sample point. the coordinate index is the number of earlier selections of the currently selected arm. definition compiled","shard":"modules/2783add4e3846770.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sumRewards_rewardFromArmStream_eq_armPrefixSum","label":"sumRewards_rewardFromArmStream_eq_armPrefixSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sumRewards_rewardFromArmStream_eq_armPrefixSum","description":"Rewards consumed from one arm are exactly the prefix of that arm's latent stream whose length is the realized pull count.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html#decl-84e6d4a0bc9c","parent":"module:BanditRLProof.Algorithms.UCBArmStreamSource","order":2866,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamSource"],["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean:42"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_rewardFromArmStream_eq_armPrefixSum {Omega : Type u} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (omega : Omega) (arm : Fin K) (n : Nat) : sumRewards (action omega) (rewardFromArmStream action armStream omega) arm n = armPrefixSum arm (pullCount (action omega) arm n) (armStream omega)","missing":[],"search":"sumrewards_rewardfromarmstream_eq_armprefixsum banditrlproof.ucb.sumrewards_rewardfromarmstream_eq_armprefixsum rewards consumed from one arm are exactly the prefix of that arm's latent stream whose length is the realized pull count. theorem compiled","shard":"modules/2783add4e3846770.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.fixedArmPrefixSourceOfArmStream","label":"fixedArmPrefixSourceOfArmStream","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.fixedArmPrefixSourceOfArmStream","description":"The next-unused-coordinate reward rule produces a fixed-arm prefix source as soon as the latent stream coordinates are measurable.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html#decl-85479513a1e0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamSource","order":2867,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamSource"],["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean:65"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def fixedArmPrefixSourceOfArmStream {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) : FixedArmPrefixSource action (rewardFromArmStream action armStream) where","missing":[],"search":"fixedarmprefixsourceofarmstream banditrlproof.ucb.fixedarmprefixsourceofarmstream the next-unused-coordinate reward rule produces a fixed-arm prefix source as soon as the latent stream coordinates are measurable. definition compiled","shard":"modules/2783add4e3846770.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_rewardFromArmStream_mem_le_identDistrib","label":"measure_pullCount_prod_sumRewards_rewardFromArmStream_mem_le_identDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_prod_sumRewards_rewardFromArmStream_mem_le_identDistrib","description":"Fixed-count peeling for rewards generated by the next-unused-coordinate rule, with complete-stream law transport to a canonical stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html#decl-317d6c0a1446","parent":"module:BanditRLProof.Algorithms.UCBArmStreamSource","order":2868,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamSource"],["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean:81"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_prod_sumRewards_rewardFromArmStream_mem_le_identDistrib {Omega : Type u} {Xi : Type v} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace Xi] (mu : Measure Omega) (nu : Measure Xi) (action : Omega -> ActionTrace (Fin K)) (armStream : Omega -> ArmRewardStream K) (hmeasurable : forall i arm, Measurable (fun omega => armStream omega i arm)) (canonicalStream : Xi -> ArmRewardStream K) (hstreamLaw : IdentDistrib armStream canonicalStream mu nu) (arm : Fin K) (n : Nat) (s : Set (Nat × Real)) [DecidablePred (fun k : Nat => k ∈ Prod.fst '' s)] (hs : MeasurableSet s) : mu {omega | (pullCount (action omega) arm n, sumRewards (action omega) (rewardFromArmStream action armStream omega) arm n) ∈ s} ≤ ((Finset.range (n + 1)).filter (fun k => k ∈ Prod.fst '' s)).sum (fun k => nu {xi | armPrefixSum arm k (canonicalStream xi) ∈ Prod.mk k ⁻¹' s})","missing":[],"search":"measure_pullcount_prod_sumrewards_rewardfromarmstream_mem_le_identdistrib banditrlproof.ucb.measure_pullcount_prod_sumrewards_rewardfromarmstream_mem_le_identdistrib fixed-count peeling for rewards generated by the next-unused-coordinate rule, with complete-stream law transport to a canonical stream. theorem compiled","shard":"modules/2783add4e3846770.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalFixedArmPrefixSource","label":"canonicalFixedArmPrefixSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalFixedArmPrefixSource","description":"The canonical stream space supplies its own measurable latent stream.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html#decl-7ec08358d921","parent":"module:BanditRLProof.Algorithms.UCBArmStreamSource","order":2869,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamSource"],["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean:109"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def canonicalFixedArmPrefixSource {K : Nat} (action : ArmRewardStream K -> ActionTrace (Fin K)) : FixedArmPrefixSource action (rewardFromArmStream action id)","missing":[],"search":"canonicalfixedarmprefixsource banditrlproof.ucb.canonicalfixedarmprefixsource the canonical stream space supplies its own measurable latent stream. definition compiled","shard":"modules/2783add4e3846770.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_rewardFromCanonicalArmStream_mem_le","label":"measure_pullCount_prod_sumRewards_rewardFromCanonicalArmStream_mem_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_prod_sumRewards_rewardFromCanonicalArmStream_mem_le","description":"Pathwise peeling directly on the canonical latent stream space. This theorem is ready for any recursively defined action trace on that space; probability or stationarity assumptions enter only in later fixed-prefix tail leaves.","url":"../modules/banditrlproof-algorithms-ucbarmstreamsource/index.html#decl-4a1302f7b2b1","parent":"module:BanditRLProof.Algorithms.UCBArmStreamSource","order":2870,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamSource"],["Source","BanditRLProof/Algorithms/UCBArmStreamSource.lean:121"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_prod_sumRewards_rewardFromCanonicalArmStream_mem_le {K : Nat} (mu : Measure (ArmRewardStream K)) (action : ArmRewardStream K -> ActionTrace (Fin K)) (arm : Fin K) (n : Nat) (s : Set (Nat × Real)) [DecidablePred (fun k : Nat => k ∈ Prod.fst '' s)] : mu {stream | (pullCount (action stream) arm n, sumRewards (action stream) (rewardFromArmStream action id stream) arm n) ∈ s} ≤ ((Finset.range (n + 1)).filter (fun k => k ∈ Prod.fst '' s)).sum (fun k => mu {stream | armPrefixSum arm k stream ∈ Prod.mk k ⁻¹' s})","missing":[],"search":"measure_pullcount_prod_sumrewards_rewardfromcanonicalarmstream_mem_le banditrlproof.ucb.measure_pullcount_prod_sumrewards_rewardfromcanonicalarmstream_mem_le pathwise peeling directly on the canonical latent stream space. this theorem is ready for any recursively defined action trace on that space; probability or stationarity assumptions enter only in later fixed-prefix tail leaves. theorem compiled","shard":"modules/2783add4e3846770.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armStreamMeasure_map_coord","label":"armStreamMeasure_map_coord","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.armStreamMeasure_map_coord","description":"Every time/arm coordinate has its prescribed stationary kernel law.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-85efc809aab8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2871,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:21"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armStreamMeasure_map_coord {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (i : Nat) (arm : Fin K) : Measure.map (fun stream : ArmRewardStream K => stream i arm) (armStreamMeasure nu) = nu arm","missing":[],"search":"armstreammeasure_map_coord banditrlproof.ucb.armstreammeasure_map_coord every time/arm coordinate has its prescribed stationary kernel law. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.iIndepFun_armStreamMeasure_coord_sub","label":"iIndepFun_armStreamMeasure_coord_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.iIndepFun_armStreamMeasure_coord_sub","description":"Centered coordinates of one fixed arm are independent across time.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-ad78b6592595","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2872,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:45"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_armStreamMeasure_coord_sub {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) : iIndepFun (fun i (stream : ArmRewardStream K) => stream i arm - mean) (armStreamMeasure nu)","missing":[],"search":"iindepfun_armstreammeasure_coord_sub banditrlproof.ucb.iindepfun_armstreammeasure_coord_sub centered coordinates of one fixed arm are independent across time. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.iIndepFun_armStreamMeasure_sub_coord","label":"iIndepFun_armStreamMeasure_sub_coord","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.iIndepFun_armStreamMeasure_sub_coord","description":"Lower-tail centered coordinates are also independent across time.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-39553adfe6fb","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2873,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:59"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_armStreamMeasure_sub_coord {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) : iIndepFun (fun i (stream : ArmRewardStream K) => mean - stream i arm) (armStreamMeasure nu)","missing":[],"search":"iindepfun_armstreammeasure_sub_coord banditrlproof.ucb.iindepfun_armstreammeasure_sub_coord lower-tail centered coordinates are also independent across time. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.hasSubgaussianMGF_armStreamMeasure_coord_sub","label":"hasSubgaussianMGF_armStreamMeasure_coord_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.hasSubgaussianMGF_armStreamMeasure_coord_sub","description":"A fixed arm's centered one-coordinate MGF transports to stream space.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-33b3e089a282","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2874,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:73"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_armStreamMeasure_coord_sub {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (i : Nat) : HasSubgaussianMGF (fun stream : ArmRewardStream K => stream i arm - mean) sigma2 (armStreamMeasure nu)","missing":[],"search":"hassubgaussianmgf_armstreammeasure_coord_sub banditrlproof.ucb.hassubgaussianmgf_armstreammeasure_coord_sub a fixed arm's centered one-coordinate mgf transports to stream space. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.hasSubgaussianMGF_armStreamMeasure_sub_coord","label":"hasSubgaussianMGF_armStreamMeasure_sub_coord","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.hasSubgaussianMGF_armStreamMeasure_sub_coord","description":"The corresponding lower-tail coordinate has the same variance proxy.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-5cfc1669dcf7","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2875,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:92"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_armStreamMeasure_sub_coord {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (i : Nat) : HasSubgaussianMGF (fun stream : ArmRewardStream K => mean - stream i arm) sigma2 (armStreamMeasure nu)","missing":[],"search":"hassubgaussianmgf_armstreammeasure_sub_coord banditrlproof.ucb.hassubgaussianmgf_armstreammeasure_sub_coord the corresponding lower-tail coordinate has the same variance proxy. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sum_coord_sub_eq_armPrefixSum_sub","label":"sum_coord_sub_eq_armPrefixSum_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sum_coord_sub_eq_armPrefixSum_sub","description":"theorem sum_coord_sub_eq_armPrefixSum_sub {K : Nat} (stream : ArmRewardStream K) (arm : Fin K) (mean : Real) (k : Nat) : (Finset.range k).sum (fun i => stream i arm - mean) = armPrefixSum arm k stream - (k : Real) * mean","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-29823418100d","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2876,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:108"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sum_coord_sub_eq_armPrefixSum_sub {K : Nat} (stream : ArmRewardStream K) (arm : Fin K) (mean : Real) (k : Nat) : (Finset.range k).sum (fun i => stream i arm - mean) = armPrefixSum arm k stream - (k : Real) * mean","missing":[],"search":"sum_coord_sub_eq_armprefixsum_sub banditrlproof.ucb.sum_coord_sub_eq_armprefixsum_sub theorem sum_coord_sub_eq_armprefixsum_sub {k : nat} (stream : armrewardstream k) (arm : fin k) (mean : real) (k : nat) : (finset.range k).sum (fun i => stream i arm - mean) = armprefixsum arm k stream - (k : real) * mean theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sum_sub_coord_eq_mul_sub_armPrefixSum","label":"sum_sub_coord_eq_mul_sub_armPrefixSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sum_sub_coord_eq_mul_sub_armPrefixSum","description":"theorem sum_sub_coord_eq_mul_sub_armPrefixSum {K : Nat} (stream : ArmRewardStream K) (arm : Fin K) (mean : Real) (k : Nat) : (Finset.range k).sum (fun i => mean - stream i arm) = (k : Real) * mean - armPrefixSum arm k stream","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-975fcfb5a6b2","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2877,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:115"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sum_sub_coord_eq_mul_sub_armPrefixSum {K : Nat} (stream : ArmRewardStream K) (arm : Fin K) (mean : Real) (k : Nat) : (Finset.range k).sum (fun i => mean - stream i arm) = (k : Real) * mean - armPrefixSum arm k stream","missing":[],"search":"sum_sub_coord_eq_mul_sub_armprefixsum banditrlproof.ucb.sum_sub_coord_eq_mul_sub_armprefixsum theorem sum_sub_coord_eq_mul_sub_armprefixsum {k : nat} (stream : armrewardstream k) (arm : fin k) (mean : real) (k : nat) : (finset.range k).sum (fun i => mean - stream i arm) = (k : real) * mean - armprefixsum arm k stream theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_armPrefixSum_sub_mul_ge_le","label":"measure_armPrefixSum_sub_mul_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_armPrefixSum_sub_mul_ge_le","description":"ENNReal upper tail for one fixed arm prefix.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-9c3f872d3e15","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2878,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:123"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_armPrefixSum_sub_mul_ge_le {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (k : Nat) {eps : Real} (heps : 0 <= eps) : armStreamMeasure nu {stream : ArmRewardStream K | eps <= armPrefixSum arm k stream - (k : Real) * mean} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * (k : Real) * (sigma2 : Real))))","missing":[],"search":"measure_armprefixsum_sub_mul_ge_le banditrlproof.ucb.measure_armprefixsum_sub_mul_ge_le ennreal upper tail for one fixed arm prefix. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_mul_sub_armPrefixSum_ge_le","label":"measure_mul_sub_armPrefixSum_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_mul_sub_armPrefixSum_ge_le","description":"ENNReal lower tail for one fixed arm prefix.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-456264fc59c0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2879,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:146"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_mul_sub_armPrefixSum_ge_le {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (k : Nat) {eps : Real} (heps : 0 <= eps) : armStreamMeasure nu {stream : ArmRewardStream K | eps <= (k : Real) * mean - armPrefixSum arm k stream} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * (k : Real) * (sigma2 : Real))))","missing":[],"search":"measure_mul_sub_armprefixsum_ge_le banditrlproof.ucb.measure_mul_sub_armprefixsum_ge_le ennreal lower tail for one fixed arm prefix. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armPrefixEmpiricalMean","label":"armPrefixEmpiricalMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armPrefixEmpiricalMean","description":"Empirical mean of the first `k` latent rewards of one fixed arm.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-ac5ac8e744a8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2880,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:169"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armPrefixEmpiricalMean {K : Nat} (arm : Fin K) (k : Nat) (stream : ArmRewardStream K) : Real","missing":[],"search":"armprefixempiricalmean banditrlproof.ucb.armprefixempiricalmean empirical mean of the first `k` latent rewards of one fixed arm. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armPrefixAverageConfidenceRadius","label":"armPrefixAverageConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armPrefixAverageConfidenceRadius","description":"Two-sided fixed-sample confidence radius for one arm with per-reward proxy variance `sigma2`.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-9e4d872747c5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2881,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:177"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def armPrefixAverageConfidenceRadius (sigma2 : NNReal) (k : Nat) (delta : Real) : Real","missing":[],"search":"armprefixaverageconfidenceradius banditrlproof.ucb.armprefixaverageconfidenceradius two-sided fixed-sample confidence radius for one arm with per-reward proxy variance `sigma2`. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_armPrefixAverageConfidenceRadius_le_abs_empiricalMean_sub","label":"measure_armPrefixAverageConfidenceRadius_le_abs_empiricalMean_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_armPrefixAverageConfidenceRadius_le_abs_empiricalMean_sub","description":"Two-sided `delta` confidence theorem for the empirical mean of exactly `k` latent rewards of one arm under the stationary product arm-stream law. The theorem derives positivity of the total proxy variance from `0 < k` and `sigma2 ≠ 0`; callers do not supply a separate aggregate-variance contract.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-7b30d9ae22b0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2882,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:188"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_armPrefixAverageConfidenceRadius_le_abs_empiricalMean_sub {K : Nat} (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (k : Nat) (hk : 0 < k) (hsigma2 : sigma2 ≠ 0) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : armStreamMeasure nu {stream : ArmRewardStream K | armPrefixAverageConfidenceRadius sigma2 k delta <= |armPrefixEmpiricalMean arm k stream - mean|} <= ENNReal.ofReal delta","missing":[],"search":"measure_armprefixaverageconfidenceradius_le_abs_empiricalmean_sub banditrlproof.ucb.measure_armprefixaverageconfidenceradius_le_abs_empiricalmean_sub two-sided `delta` confidence theorem for the empirical mean of exactly `k` latent rewards of one arm under the stationary product arm-stream law. the theorem derives positivity of the total proxy variance from `0 < k` and `sigma2 ≠ 0`; callers do not supply a separate aggregate-variance contract. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.upperDeviationPairs","label":"upperDeviationPairs","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.upperDeviationPairs","description":"Pair event used to peel an upper selected-reward deviation by pull count.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-818730ebf113","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2883,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:243"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def upperDeviationPairs (mean : Real) (threshold : Nat -> Real) : Set (Nat × Real)","missing":[],"search":"upperdeviationpairs banditrlproof.ucb.upperdeviationpairs pair event used to peel an upper selected-reward deviation by pull count. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lowerDeviationPairs","label":"lowerDeviationPairs","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.lowerDeviationPairs","description":"Pair event used to peel a lower selected-reward deviation by pull count.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-68c2f0cad597","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2884,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:248"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def lowerDeviationPairs (mean : Real) (threshold : Nat -> Real) : Set (Nat × Real)","missing":[],"search":"lowerdeviationpairs banditrlproof.ucb.lowerdeviationpairs pair event used to peel a lower selected-reward deviation by pull count. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.positiveUpperDeviationPairs","label":"positiveUpperDeviationPairs","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.positiveUpperDeviationPairs","description":"Positive-count upper deviation pair event used by UCB index tails.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-3a0aa528e5fe","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2885,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:253"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def positiveUpperDeviationPairs (mean : Real) (threshold : Nat -> Real) : Set (Nat × Real)","missing":[],"search":"positiveupperdeviationpairs banditrlproof.ucb.positiveupperdeviationpairs positive-count upper deviation pair event used by ucb index tails. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.positiveLowerDeviationPairs","label":"positiveLowerDeviationPairs","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.positiveLowerDeviationPairs","description":"Positive-count lower deviation pair event used by UCB index tails.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-ff2ddaa7cfe9","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2886,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:259"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def positiveLowerDeviationPairs (mean : Real) (threshold : Nat -> Real) : Set (Nat × Real)","missing":[],"search":"positivelowerdeviationpairs banditrlproof.ucb.positivelowerdeviationpairs positive-count lower deviation pair event used by ucb index tails. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_fst_image_upperDeviationPairs","label":"mem_fst_image_upperDeviationPairs","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_fst_image_upperDeviationPairs","description":"theorem mem_fst_image_upperDeviationPairs (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' upperDeviationPairs mean threshold","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-b77c8ed332de","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2887,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:265"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem mem_fst_image_upperDeviationPairs (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' upperDeviationPairs mean threshold","missing":[],"search":"mem_fst_image_upperdeviationpairs banditrlproof.ucb.mem_fst_image_upperdeviationpairs theorem mem_fst_image_upperdeviationpairs (mean : real) (threshold : nat -> real) (k : nat) : k ∈ prod.fst '' upperdeviationpairs mean threshold theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_fst_image_lowerDeviationPairs","label":"mem_fst_image_lowerDeviationPairs","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_fst_image_lowerDeviationPairs","description":"theorem mem_fst_image_lowerDeviationPairs (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' lowerDeviationPairs mean threshold","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-ec16707df6b4","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2888,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:274"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem mem_fst_image_lowerDeviationPairs (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' lowerDeviationPairs mean threshold","missing":[],"search":"mem_fst_image_lowerdeviationpairs banditrlproof.ucb.mem_fst_image_lowerdeviationpairs theorem mem_fst_image_lowerdeviationpairs (mean : real) (threshold : nat -> real) (k : nat) : k ∈ prod.fst '' lowerdeviationpairs mean threshold theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_fst_image_positiveUpperDeviationPairs_iff","label":"mem_fst_image_positiveUpperDeviationPairs_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_fst_image_positiveUpperDeviationPairs_iff","description":"theorem mem_fst_image_positiveUpperDeviationPairs_iff (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' positiveUpperDeviationPairs mean threshold ↔ 0 < k","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-e2b633750abc","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2889,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:283"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem mem_fst_image_positiveUpperDeviationPairs_iff (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' positiveUpperDeviationPairs mean threshold ↔ 0 < k","missing":[],"search":"mem_fst_image_positiveupperdeviationpairs_iff banditrlproof.ucb.mem_fst_image_positiveupperdeviationpairs_iff theorem mem_fst_image_positiveupperdeviationpairs_iff (mean : real) (threshold : nat -> real) (k : nat) : k ∈ prod.fst '' positiveupperdeviationpairs mean threshold ↔ 0 < k theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.mem_fst_image_positiveLowerDeviationPairs_iff","label":"mem_fst_image_positiveLowerDeviationPairs_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.mem_fst_image_positiveLowerDeviationPairs_iff","description":"theorem mem_fst_image_positiveLowerDeviationPairs_iff (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' positiveLowerDeviationPairs mean threshold ↔ 0 < k","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-99d66678a9b6","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2890,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:298"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem mem_fst_image_positiveLowerDeviationPairs_iff (mean : Real) (threshold : Nat -> Real) (k : Nat) : k ∈ Prod.fst '' positiveLowerDeviationPairs mean threshold ↔ 0 < k","missing":[],"search":"mem_fst_image_positivelowerdeviationpairs_iff banditrlproof.ucb.mem_fst_image_positivelowerdeviationpairs_iff theorem mem_fst_image_positivelowerdeviationpairs_iff (mean : real) (threshold : nat -> real) (k : nat) : k ∈ prod.fst '' positivelowerdeviationpairs mean threshold ↔ 0 < k theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le","label":"measure_sumRewards_sub_pullCount_mul_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le","description":"Adaptive-count upper selected-reward tail for the recursive UCB process. The threshold may depend on the realized pull count. Peeling turns the event into a finite sum over all fixed counts `k <= n`, and each term is discharged by the stationary fixed-arm sub-Gaussian prefix tail.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-aee262b98dfe","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2891,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:319"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_sumRewards_sub_pullCount_mul_ge_le {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : armStreamMeasure nu {stream : ArmRewardStream K | threshold (pullCount (armStreamAction hK c stream) arm n) <= sumRewards (armStreamAction hK c stream) (armStreamReward hK c stream) arm n - (pullCount (armStreamAction hK c stream) arm n : Real) * mean} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_sumrewards_sub_pullcount_mul_ge_le banditrlproof.ucb.measure_sumrewards_sub_pullcount_mul_ge_le adaptive-count upper selected-reward tail for the recursive ucb process. the threshold may depend on the realized pull count. peeling turns the event into a finite sum over all fixed counts `k <= n`, and each term is discharged by the stationary fixed-arm sub-gaussian prefix tail. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le","label":"measure_pullCount_mul_sub_sumRewards_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le","description":"Adaptive-count lower selected-reward tail for the recursive UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-1d8de38edc1f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2892,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:364"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_mul_sub_sumRewards_ge_le {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, k <= n -> 0 <= threshold k) : armStreamMeasure nu {stream : ArmRewardStream K | threshold (pullCount (armStreamAction hK c stream) arm n) <= (pullCount (armStreamAction hK c stream) arm n : Real) * mean - sumRewards (armStreamAction hK c stream) (armStreamReward hK c stream) arm n} <= (Finset.range (n + 1)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_pullcount_mul_sub_sumrewards_ge_le banditrlproof.ucb.measure_pullcount_mul_sub_sumrewards_ge_le adaptive-count lower selected-reward tail for the recursive ucb process. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pos_and_sumRewards_sub_pullCount_mul_ge_le","label":"measure_pos_and_sumRewards_sub_pullCount_mul_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pos_and_sumRewards_sub_pullCount_mul_ge_le","description":"Positive-pull-count adaptive upper tail, with the zero-count fiber removed.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-d67e1781dc46","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2893,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:409"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pos_and_sumRewards_sub_pullCount_mul_ge_le {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK c stream) arm n ∧ threshold (pullCount (armStreamAction hK c stream) arm n) <= sumRewards (armStreamAction hK c stream) (armStreamReward hK c stream) arm n - (pullCount (armStreamAction hK c stream) arm n : Real) * mean} <= ((Finset.range (n + 1)).filter (fun k => 0 < k)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_pos_and_sumrewards_sub_pullcount_mul_ge_le banditrlproof.ucb.measure_pos_and_sumrewards_sub_pullcount_mul_ge_le positive-pull-count adaptive upper tail, with the zero-count fiber removed. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pos_and_pullCount_mul_sub_sumRewards_ge_le","label":"measure_pos_and_pullCount_mul_sub_sumRewards_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pos_and_pullCount_mul_sub_sumRewards_ge_le","description":"Positive-pull-count adaptive lower tail, with the zero-count fiber removed.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-c44fc4797934","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2894,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:467"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pos_and_pullCount_mul_sub_sumRewards_ge_le {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) (threshold : Nat -> Real) (hthreshold : forall k, 0 < k -> k <= n -> 0 <= threshold k) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK c stream) arm n ∧ threshold (pullCount (armStreamAction hK c stream) arm n) <= (pullCount (armStreamAction hK c stream) arm n : Real) * mean - sumRewards (armStreamAction hK c stream) (armStreamReward hK c stream) arm n} <= ((Finset.range (n + 1)).filter (fun k => 0 < k)).sum (fun k => ENNReal.ofReal (Real.exp (-(threshold k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_pos_and_pullcount_mul_sub_sumrewards_ge_le banditrlproof.ucb.measure_pos_and_pullcount_mul_sub_sumrewards_ge_le positive-pull-count adaptive lower tail, with the zero-count fiber removed. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.countWidthThreshold","label":"countWidthThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.countWidthThreshold","description":"Pull-count-scaled confidence width used in the fixed-count tail fibers.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-5f8852dc3bd2","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2895,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:525"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def countWidthThreshold (c : Real) (sigma2 : NNReal) (n k : Nat) : Real","missing":[],"search":"countwidththreshold banditrlproof.ucb.countwidththreshold pull-count-scaled confidence width used in the fixed-count tail fibers. definition compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.countWidthThreshold_nonneg","label":"countWidthThreshold_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.countWidthThreshold_nonneg","description":"theorem countWidthThreshold_nonneg (c : Real) (sigma2 : NNReal) (n k : Nat) : 0 <= countWidthThreshold c sigma2 n k","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-1293e61ad984","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2896,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:532"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem countWidthThreshold_nonneg (c : Real) (sigma2 : NNReal) (n k : Nat) : 0 <= countWidthThreshold c sigma2 n k","missing":[],"search":"countwidththreshold_nonneg banditrlproof.ucb.countwidththreshold_nonneg theorem countwidththreshold_nonneg (c : real) (sigma2 : nnreal) (n k : nat) : 0 <= countwidththreshold c sigma2 n k theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.countWidthThreshold_le_mul_mean_sub_sumRewards_of_empiricalMean_add_width_le","label":"countWidthThreshold_le_mul_mean_sub_sumRewards_of_empiricalMean_add_width_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.countWidthThreshold_le_mul_mean_sub_sumRewards_of_empiricalMean_add_width_le","description":"Lower index failure implies the corresponding count-scaled lower sum deviation.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-b6c91495648b","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2897,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:538"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem countWidthThreshold_le_mul_mean_sub_sumRewards_of_empiricalMean_add_width_le {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (n : Nat) (mean c : Real) (sigma2 : NNReal) (hcount : 0 < pullCount action arm n) (hindex : realEmpiricalMean action reward arm n + realWidth action (c * (sigma2 : Real)) arm n <= mean) : countWidthThreshold c sigma2 n (pullCount action arm n) <= (pullCount action arm n : Real) * mean - sumRewards action reward arm n","missing":[],"search":"countwidththreshold_le_mul_mean_sub_sumrewards_of_empiricalmean_add_width_le banditrlproof.ucb.countwidththreshold_le_mul_mean_sub_sumrewards_of_empiricalmean_add_width_le lower index failure implies the corresponding count-scaled lower sum deviation. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.countWidthThreshold_le_sumRewards_sub_mul_mean_of_mean_le_empiricalMean_sub_width","label":"countWidthThreshold_le_sumRewards_sub_mul_mean_of_mean_le_empiricalMean_sub_width","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.countWidthThreshold_le_sumRewards_sub_mul_mean_of_mean_le_empiricalMean_sub_width","description":"Upper index failure implies the corresponding count-scaled upper sum deviation.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-9a2199a96c90","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2898,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:566"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem countWidthThreshold_le_sumRewards_sub_mul_mean_of_mean_le_empiricalMean_sub_width {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (n : Nat) (mean c : Real) (sigma2 : NNReal) (hcount : 0 < pullCount action arm n) (hindex : mean <= realEmpiricalMean action reward arm n - realWidth action (c * (sigma2 : Real)) arm n) : countWidthThreshold c sigma2 n (pullCount action arm n) <= sumRewards action reward arm n - (pullCount action arm n : Real) * mean","missing":[],"search":"countwidththreshold_le_sumrewards_sub_mul_mean_of_mean_le_empiricalmean_sub_width banditrlproof.ucb.countwidththreshold_le_sumrewards_sub_mul_mean_of_mean_le_empiricalmean_sub_width upper index failure implies the corresponding count-scaled upper sum deviation. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean","label":"measure_realEmpiricalMean_add_realWidth_le_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean","description":"LML-shaped lower UCB-index tail for the actual recursive arm-stream process, before simplifying the finite fixed-count exponential sum.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-bab47dac2504","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2899,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:597"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_realEmpiricalMean_add_realWidth_le_mean {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n ∧ realEmpiricalMean (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) arm n + realWidth (armStreamAction hK (c * (sigma2 : Real)) stream) (c * (sigma2 : Real)) arm n <= mean} <= ((Finset.range (n + 1)).filter (fun k => 0 < k)).sum (fun k => ENNReal.ofReal (Real.exp (-(countWidthThreshold c sigma2 n k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_realempiricalmean_add_realwidth_le_mean banditrlproof.ucb.measure_realempiricalmean_add_realwidth_le_mean lml-shaped lower ucb-index tail for the actual recursive arm-stream process, before simplifying the finite fixed-count exponential sum. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth","label":"measure_mean_le_realEmpiricalMean_sub_realWidth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth","description":"LML-shaped upper UCB-index tail for the actual recursive arm-stream process, before simplifying the finite fixed-count exponential sum.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-cfda7fc6147f","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2900,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:661"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_mean_le_realEmpiricalMean_sub_realWidth {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (n : Nat) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n ∧ mean <= realEmpiricalMean (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) arm n - realWidth (armStreamAction hK (c * (sigma2 : Real)) stream) (c * (sigma2 : Real)) arm n} <= ((Finset.range (n + 1)).filter (fun k => 0 < k)).sum (fun k => ENNReal.ofReal (Real.exp (-(countWidthThreshold c sigma2 n k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)))))","missing":[],"search":"measure_mean_le_realempiricalmean_sub_realwidth banditrlproof.ucb.measure_mean_le_realempiricalmean_sub_realwidth lml-shaped upper ucb-index tail for the actual recursive arm-stream process, before simplifying the finite fixed-count exponential sum. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.countWidthThreshold_sq_div_eq","label":"countWidthThreshold_sq_div_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.countWidthThreshold_sq_div_eq","description":"Every positive fixed-count fiber has the same logarithmic exponent.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-2c04ba0f42b9","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2901,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:723"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem countWidthThreshold_sq_div_eq (c : Real) (sigma2 : NNReal) (n k : Nat) (hc : 0 <= c) (hsigma2 : sigma2 ≠ 0) (hk : 0 < k) : (countWidthThreshold c sigma2 n k) ^ 2 / (2 * (k : Real) * (sigma2 : Real)) = c * Real.log ((n + 1 : Nat) : Real)","missing":[],"search":"countwidththreshold_sq_div_eq banditrlproof.ucb.countwidththreshold_sq_div_eq every positive fixed-count fiber has the same logarithmic exponent. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.positiveCountFilter_eq_Icc","label":"positiveCountFilter_eq_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.positiveCountFilter_eq_Icc","description":"theorem positiveCountFilter_eq_Icc (n : Nat) : (Finset.range (n + 1)).filter (fun k => 0 < k) = Finset.Icc 1 n","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-b6f0cb8233d0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2902,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:745"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem positiveCountFilter_eq_Icc (n : Nat) : (Finset.range (n + 1)).filter (fun k => 0 < k) = Finset.Icc 1 n","missing":[],"search":"positivecountfilter_eq_icc banditrlproof.ucb.positivecountfilter_eq_icc theorem positivecountfilter_eq_icc (n : nat) : (finset.range (n + 1)).filter (fun k => 0 < k) = finset.icc 1 n theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sum_countWidthThreshold_tail_eq","label":"sum_countWidthThreshold_tail_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sum_countWidthThreshold_tail_eq","description":"The fixed-count exponential sum collapses to `n` identical terms.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-c80f57cb0fc2","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2903,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:752"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sum_countWidthThreshold_tail_eq (c : Real) (sigma2 : NNReal) (n : Nat) (hc : 0 <= c) (hsigma2 : sigma2 ≠ 0) : ((Finset.range (n + 1)).filter (fun k => 0 < k)).sum (fun k => ENNReal.ofReal (Real.exp (-(countWidthThreshold c sigma2 n k) ^ 2 / (2 * (k : Real) * (sigma2 : Real))))) = (n : ENNReal) * ENNReal.ofReal (Real.exp (-c * Real.log ((n + 1 : Nat) : Real)))","missing":[],"search":"sum_countwidththreshold_tail_eq banditrlproof.ucb.sum_countwidththreshold_tail_eq the fixed-count exponential sum collapses to `n` identical terms. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.natCast_mul_exp_neg_log_le_inv_rpow_sub_one","label":"natCast_mul_exp_neg_log_le_inv_rpow_sub_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.natCast_mul_exp_neg_log_le_inv_rpow_sub_one","description":"Convert the peeled logarithmic tail into the inverse-power form used by LML.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-7734450c98bc","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2904,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:780"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem natCast_mul_exp_neg_log_le_inv_rpow_sub_one (c : Real) (n : Nat) : (n : ENNReal) * ENNReal.ofReal (Real.exp (-c * Real.log ((n + 1 : Nat) : Real))) <= (1 : ENNReal) / (((n + 1 : Nat) : ENNReal) ^ (c - 1))","missing":[],"search":"natcast_mul_exp_neg_log_le_inv_rpow_sub_one banditrlproof.ucb.natcast_mul_exp_neg_log_le_inv_rpow_sub_one convert the peeled logarithmic tail into the inverse-power form used by lml. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean_log_bound","label":"measure_realEmpiricalMean_add_realWidth_le_mean_log_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean_log_bound","description":"Simplified logarithmic lower-index tail for the recursive UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-acf2301119c0","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2905,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:826"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_realEmpiricalMean_add_realWidth_le_mean_log_bound {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hc : 0 <= c) (hsigma2 : sigma2 ≠ 0) (n : Nat) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n ∧ realEmpiricalMean (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) arm n + realWidth (armStreamAction hK (c * (sigma2 : Real)) stream) (c * (sigma2 : Real)) arm n <= mean} <= (n : ENNReal) * ENNReal.ofReal (Real.exp (-c * Real.log ((n + 1 : Nat) : Real)))","missing":[],"search":"measure_realempiricalmean_add_realwidth_le_mean_log_bound banditrlproof.ucb.measure_realempiricalmean_add_realwidth_le_mean_log_bound simplified logarithmic lower-index tail for the recursive ucb process. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth_log_bound","label":"measure_mean_le_realEmpiricalMean_sub_realWidth_log_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth_log_bound","description":"Simplified logarithmic upper-index tail for the recursive UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-4658f666b1b8","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2906,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:851"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_mean_le_realEmpiricalMean_sub_realWidth_log_bound {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hc : 0 <= c) (hsigma2 : sigma2 ≠ 0) (n : Nat) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n ∧ mean <= realEmpiricalMean (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) arm n - realWidth (armStreamAction hK (c * (sigma2 : Real)) stream) (c * (sigma2 : Real)) arm n} <= (n : ENNReal) * ENNReal.ofReal (Real.exp (-c * Real.log ((n + 1 : Nat) : Real)))","missing":[],"search":"measure_mean_le_realempiricalmean_sub_realwidth_log_bound banditrlproof.ucb.measure_mean_le_realempiricalmean_sub_realwidth_log_bound simplified logarithmic upper-index tail for the recursive ucb process. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean_rpow_bound","label":"measure_realEmpiricalMean_add_realWidth_le_mean_rpow_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean_rpow_bound","description":"LML-shaped inverse-power lower-index tail for the recursive UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-f0e558bc74ca","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2907,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:877"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_realEmpiricalMean_add_realWidth_le_mean_rpow_bound {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hc : 0 <= c) (hsigma2 : sigma2 ≠ 0) (n : Nat) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n /\\ realEmpiricalMean (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) arm n + realWidth (armStreamAction hK (c * (sigma2 : Real)) stream) (c * (sigma2 : Real)) arm n <= mean} <= (1 : ENNReal) / (((n + 1 : Nat) : ENNReal) ^ (c - 1))","missing":[],"search":"measure_realempiricalmean_add_realwidth_le_mean_rpow_bound banditrlproof.ucb.measure_realempiricalmean_add_realwidth_le_mean_rpow_bound lml-shaped inverse-power lower-index tail for the recursive ucb process. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth_rpow_bound","label":"measure_mean_le_realEmpiricalMean_sub_realWidth_rpow_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth_rpow_bound","description":"LML-shaped inverse-power upper-index tail for the recursive UCB process.","url":"../modules/banditrlproof-algorithms-ucbarmstreamtail/index.html#decl-5ea28623c4d5","parent":"module:BanditRLProof.Algorithms.UCBArmStreamTail","order":2908,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmStreamTail"],["Source","BanditRLProof/Algorithms/UCBArmStreamTail.lean:900"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_mean_le_realEmpiricalMean_sub_realWidth_rpow_bound {K : Nat} (hK : 0 < K) (c : Real) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (arm : Fin K) (mean : Real) (sigma2 : NNReal) (hsubG : HasSubgaussianMGF (fun reward => reward - mean) sigma2 (nu arm)) (hc : 0 <= c) (hsigma2 : sigma2 ≠ 0) (n : Nat) : armStreamMeasure nu {stream : ArmRewardStream K | 0 < pullCount (armStreamAction hK (c * (sigma2 : Real)) stream) arm n /\\ mean <= realEmpiricalMean (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) arm n - realWidth (armStreamAction hK (c * (sigma2 : Real)) stream) (c * (sigma2 : Real)) arm n} <= (1 : ENNReal) / (((n + 1 : Nat) : ENNReal) ^ (c - 1))","missing":[],"search":"measure_mean_le_realempiricalmean_sub_realwidth_rpow_bound banditrlproof.ucb.measure_mean_le_realempiricalmean_sub_realwidth_rpow_bound lml-shaped inverse-power upper-index tail for the recursive ucb process. theorem compiled","shard":"modules/1cbc2bb47b2d036a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret","label":"selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret","description":"Exact sampled-successor expected pseudo-regret for stationary finite-arm laws bounded almost surely in arm-dependent intervals. The parent practical route pads the finite maximum of the genuine armwise Hoeffding proxies.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html#decl-5833b5374624","parent":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","order":2909,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean:24"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (defaultAction : Fin K) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret banditrlproof.ucb.selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret exact sampled-successor expected pseudo-regret for stationary finite-arm laws bounded almost surely in arm-dependent intervals. the parent practical route pads the finite maximum of the genuine armwise hoeffding proxies. definition compiled","shard":"modules/d5057ea7d0211b1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_nonneg_and_le","label":"selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_nonneg_and_le","description":"The exact armwise-bounded sampled-successor expected pseudo-regret is nonnegative and satisfies the fixed-model logarithmic envelope at every large horizon.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html#decl-8eeb1581fac3","parent":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","order":2910,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean:42"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_nonneg_and_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc (lo arm) (hi arm) ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) (T : Nat) (hlarge : 2 * K <= T + 1) : 0 <= selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret model armLaw hprob lo hi defaultAction T /\\ selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret model armLaw hprob lo hi defaultAction T <= selectedPolicySucces…","missing":[],"search":"selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret_nonneg_and_le banditrlproof.ucb.selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret_nonneg_and_le the exact armwise-bounded sampled-successor expected pseudo-regret is nonnegative and satisfies the fixed-model logarithmic envelope at every large horizon. theorem compiled","shard":"modules/d5057ea7d0211b1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isBigO_log","label":"selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isBigO_log","description":"The exact armwise-bounded expected pseudo-regret family is logarithmic.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html#decl-8a043d5b5f9a","parent":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","order":2911,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean:87"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isBigO_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc (lo arm) (hi arm) ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) : (selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret model armLaw hprob lo hi defaultAction) =O[atTop] (fun T : Nat => Real.log (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret_isbigo_log banditrlproof.ucb.selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret_isbigo_log the exact armwise-bounded expected pseudo-regret family is logarithmic. theorem compiled","shard":"modules/d5057ea7d0211b1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","label":"selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","description":"The exact armwise-bounded expected pseudo-regret is little-o of `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html#decl-f8376e13948a","parent":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","order":2912,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean:126"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc (lo arm) (hi arm) ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) : (selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret model armLaw hprob lo hi defaultAction) =o[atTop] (fun T : Nat => (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret_islittleo_natcast_succ banditrlproof.ucb.selectedpolicysuccessorarmwiseboundedfinitearmexpectedpseudoregret_islittleo_natcast_succ the exact armwise-bounded expected pseudo-regret is little-o of `t + 1`. theorem compiled","shard":"modules/d5057ea7d0211b1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret","label":"selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret","description":"Expected armwise-bounded sampled-successor regret normalized by `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html#decl-0d5258c8b3f7","parent":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","order":2913,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean:153"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (defaultAction : Fin K) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorarmwiseboundedfinitearmexpectedaveragepseudoregret banditrlproof.ucb.selectedpolicysuccessorarmwiseboundedfinitearmexpectedaveragepseudoregret expected armwise-bounded sampled-successor regret normalized by `t + 1`. definition compiled","shard":"modules/d5057ea7d0211b1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","label":"selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","description":"For stationary finite-arm reward laws bounded almost surely in arm-dependent intervals, the expected pseudo-regret of the horizon-indexed canonical sampled-pair UCB family, normalized by `T + 1`, tends to zero. No pointwise nondegeneracy premise `lo arm < hi arm` is needed because the parent practical route pads the finite maximum of the genuine Hoeffding proxies before using it as the UCB parameter.","url":"../modules/banditrlproof-algorithms-ucbarmwiseboundedfinitearmsampledasymptotics/index.html#decl-7fe5baf3cbb2","parent":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","order":2914,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBArmwiseBoundedFiniteArmSampledAsymptotics.lean:175"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc (lo arm) (hi arm) ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) : Tendsto (selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret model armLaw hprob lo hi defaultAction) atTop (nhds 0)","missing":[],"search":"selectedpolicysuccessorarmwiseboundedfinitearmexpectedaveragepseudoregret_tendsto_zero banditrlproof.ucb.selectedpolicysuccessorarmwiseboundedfinitearmexpectedaveragepseudoregret_tendsto_zero for stationary finite-arm reward laws bounded almost surely in arm-dependent intervals, the expected pseudo-regret of the horizon-indexed canonical sampled-pair ucb family, normalized by `t + 1`, tends to zero. no pointwise nondegeneracy premise `lo arm < hi arm` is needed because the parent practical route pads the finite maximum of the genuine hoeffding proxies before using it as the ucb parameter. theorem compiled","shard":"modules/d5057ea7d0211b1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_boundedFiniteArmLaws","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_boundedFiniteArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_boundedFiniteArmLaws","description":"Canonical Real expected pseudo-regret bound for finite-arm stationary reward laws supported almost surely on one nondegenerate interval. The initial reward is sampled from the default arm law. Successor rewards use the context-independent action law kernel, while the context is `Unit`.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmrewardlaw/index.html#decl-8b8ce08f1572","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","order":2915,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmRewardLaw.lean:26"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_boundedFiniteArmLaws {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hlohi : lo < hi) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) : let sigma2 := Concentration.intervalVarianceProxy lo hi let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Unit) armLaw hprob let context : (n : Nat) -> History.FiniteRe…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_boundedfinitearmlaws banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_boundedfinitearmlaws canonical real expected pseudo-regret bound for finite-arm stationary reward laws supported almost surely on one nondegenerate interval. the initial reward is sampled from the default arm law. successor rewards use the context-independent action law kernel, while the context is `unit`. theorem compiled","shard":"modules/e306ae44ca5c44b7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_armwiseBoundedFiniteArmLaws","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_armwiseBoundedFiniteArmLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_armwiseBoundedFiniteArmLaws","description":"Canonical Real expected pseudo-regret bound for finite-arm stationary reward laws with arm-dependent nondegenerate support intervals. The UCB variance parameter is the maximum of the armwise Hoeffding proxies, computed internally with `Finset.sup` over all arms.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmrewardlaw/index.html#decl-3774c6cae787","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","order":2916,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmRewardLaw.lean:115"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_armwiseBoundedFiniteArmLaws {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hlohi : forall arm, lo arm < hi arm) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc (lo arm) (hi arm) ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) : let sigma2 := Concentration.finiteArmIntervalVarianceProxy lo hi let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Unit)…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_armwiseboundedfinitearmlaws banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_armwiseboundedfinitearmlaws canonical real expected pseudo-regret bound for finite-arm stationary reward laws with arm-dependent nondegenerate support intervals. the ucb variance parameter is the maximum of the armwise hoeffding proxies, computed internally with `finset.sup` over all arms. theorem compiled","shard":"modules/e306ae44ca5c44b7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret","label":"selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret","description":"Exact sampled-successor expected pseudo-regret for stationary finite-arm laws bounded almost surely in a common interval. The genuine armwise proxy is the common Hoeffding proxy; the parent practical route pads its finite maximum.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmsampledasymptotics/index.html#decl-7c1e1ca704be","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","order":2917,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmSampledAsymptotics.lean:24"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (defaultAction : Fin K) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorboundedfinitearmexpectedpseudoregret banditrlproof.ucb.selectedpolicysuccessorboundedfinitearmexpectedpseudoregret exact sampled-successor expected pseudo-regret for stationary finite-arm laws bounded almost surely in a common interval. the genuine armwise proxy is the common hoeffding proxy; the parent practical route pads its finite maximum. definition compiled","shard":"modules/62c0a91378211d0f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isBigO_log","label":"selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isBigO_log","description":"The exact common-bounded expected pseudo-regret family is logarithmic.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmsampledasymptotics/index.html#decl-aab3a082e569","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","order":2918,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmSampledAsymptotics.lean:38"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isBigO_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) : (selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret model armLaw hprob lo hi defaultAction) =O[atTop] (fun T : Nat => Real.log (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessorboundedfinitearmexpectedpseudoregret_isbigo_log banditrlproof.ucb.selectedpolicysuccessorboundedfinitearmexpectedpseudoregret_isbigo_log the exact common-bounded expected pseudo-regret family is logarithmic. theorem compiled","shard":"modules/62c0a91378211d0f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","label":"selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","description":"The exact common-bounded expected pseudo-regret is little-o of `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmsampledasymptotics/index.html#decl-307af1cfc1f0","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","order":2919,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmSampledAsymptotics.lean:79"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) : (selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret model armLaw hprob lo hi defaultAction) =o[atTop] (fun T : Nat => (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessorboundedfinitearmexpectedpseudoregret_islittleo_natcast_succ banditrlproof.ucb.selectedpolicysuccessorboundedfinitearmexpectedpseudoregret_islittleo_natcast_succ the exact common-bounded expected pseudo-regret is little-o of `t + 1`. theorem compiled","shard":"modules/62c0a91378211d0f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret","label":"selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret","description":"Expected common-bounded sampled-successor regret normalized by `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmsampledasymptotics/index.html#decl-a4cc01693132","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","order":2920,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmSampledAsymptotics.lean:105"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (defaultAction : Fin K) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorboundedfinitearmexpectedaveragepseudoregret banditrlproof.ucb.selectedpolicysuccessorboundedfinitearmexpectedaveragepseudoregret expected common-bounded sampled-successor regret normalized by `t + 1`. definition compiled","shard":"modules/62c0a91378211d0f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","label":"selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","description":"For stationary finite-arm reward laws bounded almost surely in a common interval, the expected pseudo-regret of the horizon-indexed canonical sampled-pair UCB family, normalized by `T + 1`, tends to zero. No nondegeneracy premise `lo < hi` is needed because the parent practical route pads the genuine Hoeffding proxy before using it as the UCB parameter.","url":"../modules/banditrlproof-algorithms-ucbboundedfinitearmsampledasymptotics/index.html#decl-b7b0288ea044","parent":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","order":2921,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBBoundedFiniteArmSampledAsymptotics.lean:126"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (defaultAction : Fin K) : Tendsto (selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret model armLaw hprob lo hi defaultAction) atTop (nhds 0)","missing":[],"search":"selectedpolicysuccessorboundedfinitearmexpectedaveragepseudoregret_tendsto_zero banditrlproof.ucb.selectedpolicysuccessorboundedfinitearmexpectedaveragepseudoregret_tendsto_zero for stationary finite-arm reward laws bounded almost surely in a common interval, the expected pseudo-regret of the horizon-indexed canonical sampled-pair ucb family, normalized by `t + 1`, tends to zero. no nondegeneracy premise `lo < hi` is needed because the parent practical route pads the genuine hoeffding proxy before using it as the ucb parameter. theorem compiled","shard":"modules/62c0a91378211d0f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorEmpiricalMeanAt","label":"selectedPolicySuccessorEmpiricalMeanAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorEmpiricalMeanAt","description":"Successor empirical mean at the positive horizon `t + 1`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-933a6b4c9149","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2922,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:22"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorEmpiricalMeanAt {Omega : Type u} {Action : Type} [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (arm : Action) : Real","missing":[],"search":"selectedpolicysuccessorempiricalmeanat banditrlproof.ucb.selectedpolicysuccessorempiricalmeanat successor empirical mean at the positive horizon `t + 1`. definition compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRadiusAt","label":"selectedPolicySuccessorRadiusAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorRadiusAt","description":"Realized-count confidence width used at one arm/time pair.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-b4442f0f693c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2923,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:31"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorRadiusAt {Omega : Type u} {Action : Type} [DecidableEq Action] (action : Omega -> ActionTrace Action) (sigma2 : NNReal) (arms : Finset Action) (T : Nat) (delta : Real) (omega : Omega) (t : Nat) (arm : Action) : Real","missing":[],"search":"selectedpolicysuccessorradiusat banditrlproof.ucb.selectedpolicysuccessorradiusat realized-count confidence width used at one arm/time pair. definition compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorIndexAt","label":"selectedPolicySuccessorIndexAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorIndexAt","description":"Practical UCB index with a sample-dependent realized-count width.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-ee82c891289c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2924,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:43"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorIndexAt {Omega : Type u} {Action : Type} [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (arms : Finset Action) (T : Nat) (delta : Real) (omega : Omega) (t : Nat) (arm : Action) : Real","missing":[],"search":"selectedpolicysuccessorindexat banditrlproof.ucb.selectedpolicysuccessorindexat practical ucb index with a sample-dependent realized-count width. definition compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorInitializedScoreMaxSource","label":"SelectedPolicySuccessorInitializedScoreMaxSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UCB.SelectedPolicySuccessorInitializedScoreMaxSource","description":"Initialization and score-maximality contract for the finite set of charged UCB times. The time set may omit initialization rounds; every retained time must lie below `T`, and both the designated best arm and the selected arm must have positive realized successor pull counts.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-b34ecc6c5fe1","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2925,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:63"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"structure SelectedPolicySuccessorInitializedScoreMaxSource {Omega : Type u} {Action : Type} [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (arms : Finset Action) (armMean : Action -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) where","missing":[],"search":"selectedpolicysuccessorinitializedscoremaxsource banditrlproof.ucb.selectedpolicysuccessorinitializedscoremaxsource initialization and score-maximality contract for the finite set of charged ucb times. the time set may omit initialization rounds; every retained time must lie below `t`, and both the designated best arm and the selected arm must have positive realized successor pull counts. structure compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorLargeGapEvent","label":"selectedPolicySuccessorLargeGapEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorLargeGapEvent","description":"Large-gap selected-time event charged by the practical confidence event.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-ccbd77560ee6","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2926,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:88"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorLargeGapEvent {Omega : Type u} {Action : Type} [DecidableEq Action] {action : Omega -> ActionTrace Action} {reward : Omega -> RewardTrace Rat} {arms : Finset Action} {armMean : Action -> Rat} {sigma2 : NNReal} {T : Nat} {delta : Real} (source : SelectedPolicySuccessorInitializedScoreMaxSource action reward arms armMean sigma2 T delta) : Set Omega","missing":[],"search":"selectedpolicysuccessorlargegapevent banditrlproof.ucb.selectedpolicysuccessorlargegapevent large-gap selected-time event charged by the practical confidence event. definition compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorInitializedScoreMaxSource.meanGap_le_two_radius_of_not_badEvent","label":"meanGap_le_two_radius_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.SelectedPolicySuccessorInitializedScoreMaxSource.meanGap_le_two_radius_of_not_badEvent","description":"Outside the practical simultaneous confidence event, score maximality implies the standard UCB gap bound at every initialized charged time.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-9eb4bc4d580c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2927,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:106"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem SelectedPolicySuccessorInitializedScoreMaxSource.meanGap_le_two_radius_of_not_badEvent {Omega : Type u} {Action : Type} [DecidableEq Action] {action : Omega -> ActionTrace Action} {reward : Omega -> RewardTrace Rat} {arms : Finset Action} {armMean : Action -> Rat} {sigma2 : NNReal} {T : Nat} {delta : Real} (source : SelectedPolicySuccessorInitializedScoreMaxSource action reward arms armMean sigma2 T delta) (omega : Omega) (t : Nat) (ht : t ∈ source.times) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeBadEvent action reward arms armMean sigma2 T delta) : meanGap (fun arm => (armMean arm : Real)) source.best (source.chosen omega t) <= 2 * selectedPolicySuccessorRadiusAt action sigma2 arms T delta omega t (source.chosen omega t)","missing":[],"search":"meangap_le_two_radius_of_not_badevent banditrlproof.ucb.selectedpolicysuccessorinitializedscoremaxsource.meangap_le_two_radius_of_not_badevent outside the practical simultaneous confidence event, score maximality implies the standard ucb gap bound at every initialized charged time. theorem compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Practical selected-policy UCB large-gap event bound on initialized times. The simultaneous empirical-mean theorem supplies the probability bound. The source contract turns any large-gap score-maximal selection outside that event into a contradiction via the deterministic UCB confidence algebra.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlaw/index.html#decl-2073f705ed21","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","order":2928,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLaw"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLaw.lean:189"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (arms : Finset Action) (harms : arms.Nonempty) (ar…","missing":[],"search":"measure_selectedpolicysuccessorlargegapevent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.ucb.measure_selectedpolicysuccessorlargegapevent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded practical selected-policy ucb large-gap event bound on initialized times. the simultaneous empirical-mean theorem supplies the probability bound. the source contract turns any large-gap score-maximal selection outside that event into a contradiction via the deterministic ucb confidence algebra. theorem compiled","shard":"modules/35442cfe57b139df.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_centeredKernel_of_variance_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_centeredKernel_of_variance_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_centeredKernel_of_variance_le","description":"Direct selected-law conditional-MGF bridge for a centered reward kernel. Unlike the older raw-range source route, this theorem consumes the `CenteredRewardKernelLaw` MGF field directly. No pointwise or almost-everywhere reward-range hypothesis is needed.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-3d09cdb57b6c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2929,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:17"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_centeredKernel_of_variance_le {Omega Context State Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (state : (n : Nat) -> History.FiniteRewardHistory Rat n -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (hcontext : forall n : Nat, Measurable (context…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_centeredkernel_of_variance_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_centeredkernel_of_variance_le direct selected-law conditional-mgf bridge for a centered reward kernel. unlike the older raw-range source route, this theorem consumes the `centeredrewardkernellaw` mgf field directly. no pointwise or almost-everywhere reward-range hypothesis is needed. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Predictable-variance two-sided tail for one arm, obtained directly from the centered-kernel conditional MGF. The selected arm is charged `sigma2` only at successor times when it is pulled.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-5b9770e1b8d9","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2930,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:183"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context State Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (state : (n : Nat) -> History.FiniteRewardHistory Rat n -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (sigma2 : NNR…","missing":[],"search":"armmaskedcenteredrewardsuccprocess_sum_abs_tail_predictablevariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.conditionalexpectationreward.armmaskedcenteredrewardsuccprocess_sum_abs_tail_predictablevariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel predictable-variance two-sided tail for one arm, obtained directly from the centered-kernel conditional mgf. the selected arm is charged `sigma2` only at successor times when it is pulled. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Exact positive pull-count confidence using the centered-kernel route.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-21cb72e48e82","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2931,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:367"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context State Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (state : (n : Nat) -> History.FiniteRewardHistory Rat n -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (armMean : Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward ome…","missing":[],"search":"successorarmempiricalmean_abs_tail_exact_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.conditionalexpectationreward.successorarmempiricalmean_abs_tail_exact_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel exact positive pull-count confidence using the centered-kernel route. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Positive random pull-count confidence via finite exact-count peeling.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-5e987cfdf9da","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2932,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:572"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context State Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (state : (n : Nat) -> History.FiniteRewardHistory Rat n -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (armMean : Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward om…","missing":[],"search":"successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.conditionalexpectationreward.successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel positive random pull-count confidence via finite exact-count peeling. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Finite-arm, finite-time empirical-mean confidence from the centered-kernel selected-law route. This is fixed-horizon and union-bounded, not anytime.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-802b1c4cd321","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2933,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:698"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context State Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (state : (n : Nat) -> History.FiniteRewardHistory Rat n -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (arms : Finset Action) (harms : arms.Nonempty) (armMean : Action -> Rat) (reward : Omega -> RewardTrace Rat…","missing":[],"search":"successorarmempiricalmean_simultaneous_finitearmtime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.conditionalexpectationreward.successorarmempiricalmean_simultaneous_finitearmtime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel finite-arm, finite-time empirical-mean confidence from the centered-kernel selected-law route. this is fixed-horizon and union-bounded, not anytime. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Practical selected-policy UCB large-gap event bound using the centered-kernel conditional-MGF route. No reward or mean range contract is exposed.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-cc5656f700db","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2934,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:841"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context State Action : Type} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (state : (n : Nat) -> History.FiniteRewardHistory Rat n -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (arms : Finset Action) (harms : arms.Nonempty) (armMean : Action -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : f…","missing":[],"search":"measure_selectedpolicysuccessorlargegapevent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.ucb.measure_selectedpolicysuccessorlargegapevent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel practical selected-policy ucb large-gap event bound using the centered-kernel conditional-mgf route. no reward or mean range contract is exposed. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorRewardMapLaw","label":"SelectedPolicySuccessorRewardMapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.SelectedPolicySuccessorRewardMapLaw","description":"Exact selected-reward conditional-law contract for the generated UCB policy.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-80df0995dfba","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2935,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:950"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def SelectedPolicySuccessorRewardMapLaw {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (sigma2 : NNReal) (T : Nat) (delta : Real) : Prop","missing":[],"search":"selectedpolicysuccessorrewardmaplaw banditrlproof.ucb.selectedpolicysuccessorrewardmaplaw exact selected-reward conditional-law contract for the generated ucb policy. definition compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Generated-UCB large-gap event bound with no range assumptions.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-0846a84e92e7","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2936,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:1009"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best : Fin K) (armMean : Fin K -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.CenteredRewardKernelLaw rewardKernel mean var…","missing":[],"search":"measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.ucb.measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredkernel generated-ucb large-gap event bound with no range assumptions. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredKernel","label":"lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredKernel","description":"Expected pull count at the explicit generated-UCB threshold, using only the centered-kernel law and the selected reward-map identity.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-23a8df16de25","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2937,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:1095"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredKernel {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best chosen : Fin K) (armMean : Fin K -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : Rewar…","missing":[],"search":"lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.ucb.lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredkernel expected pull count at the explicit generated-ucb threshold, using only the centered-kernel law and the selected reward-map identity. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy_centeredKernel","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy_centeredKernel","description":"Finite-arm ENNReal pseudo-regret assembly for the centered-kernel generated-UCB route. Positive-gap arms use their explicit pull threshold; zero gaps vanish.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-861f26d41640","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2938,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:1174"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy_centeredKernel {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (model : FiniteBanditModel K) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.CenteredRewardKernelLaw rewardK…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_of_reward_map_eq_selected_policy_centeredkernel finite-arm ennreal pseudo-regret assembly for the centered-kernel generated-ucb route. positive-gap arms use their explicit pull threshold; zero gaps vanish. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy_centeredKernel","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy_centeredKernel","description":"Textbook reciprocal-gap pseudo-regret sum for the centered-kernel route.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-6f42c4cf7d77","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2939,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:1280"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy_centeredKernel {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (model : FiniteBanditModel K) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega => reward omega t)) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.CenteredRewardKernelLaw rewardKernel…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_of_reward_map_eq_selected_policy_centeredkernel banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_of_reward_map_eq_selected_policy_centeredkernel textbook reciprocal-gap pseudo-regret sum for the centered-kernel route. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","description":"Canonical reward-only generated-UCB textbook pseudo-regret theorem. The canonical `trajMeasure` supplies the selected conditional reward law, and `CenteredRewardKernelLaw` supplies the analytic MGF/integrability contract. Consequently no raw-reward or mean-range premise remains.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernel/index.html#decl-4db7741199cc","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","order":2940,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernel.lean:1361"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel {Context : Type} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (T : Nat) (hT : 0 < T) (hsigma2 : 0 < (((sigma2 : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hvariance : forall i : Nat, forall history : History.FiniteRewardHi…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_centeredkernel banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_centeredkernel canonical reward-only generated-ucb textbook pseudo-regret theorem. the canonical `trajmeasure` supplies the selected conditional reward law, and `centeredrewardkernellaw` supplies the analytic mgf/integrability contract. consequently no raw-reward or mean-range premise remains. theorem compiled","shard":"modules/b4dbdbef832594d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget_nonneg","label":"selectedPolicySuccessorTextbookGapBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget_nonneg","description":"A positive-gap arm has a nonnegative textbook threshold contribution.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernelreal/index.html#decl-34959dd2457e","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","order":2941,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernelReal.lean:12"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTextbookGapBudget_nonneg (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) (hgap : 0 < gap) : 0 <= selectedPolicySuccessorTextbookGapBudget K sigma2 T delta gap","missing":[],"search":"selectedpolicysuccessortextbookgapbudget_nonneg banditrlproof.ucb.selectedpolicysuccessortextbookgapbudget_nonneg a positive-gap arm has a nonnegative textbook threshold contribution. theorem compiled","shard":"modules/db412e11f4d68c1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integrable_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction","label":"integrable_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integrable_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction","description":"Finite-horizon shifted UCB pseudo-regret is Bochner integrable whenever the underlying reward coordinates are measurable and the ambient measure is finite. No reward-law or concentration assumption is needed for this regularity statement.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernelreal/index.html#decl-65ff99f31998","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","order":2942,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernelReal.lean:28"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integrable_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsFiniteMeasure mu] (model : FiniteBanditModel K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Integrable (fun omega : Omega => ((pseudoRegret model (selectedPolicySuccessorGeneratedUCBRegretAction hK sigma2 T delta defaultAction reward omega) T : Rat) : Real)) mu","missing":[],"search":"integrable_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction banditrlproof.ucb.integrable_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction finite-horizon shifted ucb pseudo-regret is bochner integrable whenever the underlying reward coordinates are measurable and the ambient measure is finite. no reward-law or concentration assumption is needed for this regularity statement. theorem compiled","shard":"modules/db412e11f4d68c1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","description":"Canonical Real/Bochner expected pseudo-regret theorem for the generated UCB trajectory measure and a centered sub-Gaussian reward kernel. The probabilistic work is inherited from the ENNReal canonical theorem. This wrapper proves finite-horizon integrability, uses nonnegativity of model gaps, and converts the finite ENNReal textbook sum term by term to `Real`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawcenteredkernelreal/index.html#decl-c9c17464d0ce","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","order":2943,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawCenteredKernelReal.lean:62"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel {Context : Type} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (T : Nat) (hT : 0 < T) (hsigma2 : 0 < (((sigma2 : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hvariance : forall i : Nat, forall history : History.FiniteRewardHisto…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_centeredkernel banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_centeredkernel canonical real/bochner expected pseudo-regret theorem for the generated ucb trajectory measure and a centered sub-gaussian reward kernel. the probabilistic work is inherited from the ennreal canonical theorem. this wrapper proves finite-horizon integrability, uses nonnegativity of model gaps, and converts the finite ennreal textbook sum term by term to `real`. theorem compiled","shard":"modules/db412e11f4d68c1c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorFiniteHistoryState","label":"SelectedPolicySuccessorFiniteHistoryState","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UCB.SelectedPolicySuccessorFiniteHistoryState","description":"A finite pair history packaged with its dependent horizon.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-0134075f6987","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2944,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:20"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"structure SelectedPolicySuccessorFiniteHistoryState (K : Nat) where","missing":[],"search":"selectedpolicysuccessorfinitehistorystate banditrlproof.ucb.selectedpolicysuccessorfinitehistorystate a finite pair history packaged with its dependent horizon. structure compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.completeFinitePairHistory","label":"completeFinitePairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.completeFinitePairHistory","description":"Complete a finite pair history with a fixed pair outside its prefix.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-0b1d4a965162","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2945,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:29"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def completeFinitePairHistory {Action Reward : Type} (t : Nat) (history : History.FinitePairHistory Action Reward t) (defaultAction : Action) (defaultReward : Reward) : Nat -> Action × Reward","missing":[],"search":"completefinitepairhistory banditrlproof.ucb.completefinitepairhistory complete a finite pair history with a fixed pair outside its prefix. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.completeFinitePairHistoryAction","label":"completeFinitePairHistoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.completeFinitePairHistoryAction","description":"Action projection of a completed finite pair history.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-9a1e66754157","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2946,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:41"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def completeFinitePairHistoryAction {Action Reward : Type} (t : Nat) (history : History.FinitePairHistory Action Reward t) (defaultAction : Action) (defaultReward : Reward) : ActionTrace Action","missing":[],"search":"completefinitepairhistoryaction banditrlproof.ucb.completefinitepairhistoryaction action projection of a completed finite pair history. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.completeFinitePairHistoryReward","label":"completeFinitePairHistoryReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.completeFinitePairHistoryReward","description":"Reward projection of a completed finite pair history.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-a8e67b0b25c5","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2947,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:49"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def completeFinitePairHistoryReward {Action Reward : Type} (t : Nat) (history : History.FinitePairHistory Action Reward t) (defaultAction : Action) (defaultReward : Reward) : RewardTrace Reward","missing":[],"search":"completefinitepairhistoryreward banditrlproof.ucb.completefinitepairhistoryreward reward projection of a completed finite pair history. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.completeFinitePairHistory_finitePairHistoryOfTrace_apply_of_le","label":"completeFinitePairHistory_finitePairHistoryOfTrace_apply_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.completeFinitePairHistory_finitePairHistoryOfTrace_apply_of_le","description":"theorem completeFinitePairHistory_finitePairHistoryOfTrace_apply_of_le {Action Reward : Type} (action : ActionTrace Action) (reward : RewardTrace Reward) (t s : Nat) (defaultAction : Action) (defaultReward : Reward) (hs : s <= t) : completeFinitePairHistory t (History.finitePairHistoryOfTrace action reward t) defaultAction defaultReward s = (action s, reward s)","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-0a7c248a8ba6","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2948,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:57"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem completeFinitePairHistory_finitePairHistoryOfTrace_apply_of_le {Action Reward : Type} (action : ActionTrace Action) (reward : RewardTrace Reward) (t s : Nat) (defaultAction : Action) (defaultReward : Reward) (hs : s <= t) : completeFinitePairHistory t (History.finitePairHistoryOfTrace action reward t) defaultAction defaultReward s = (action s, reward s)","missing":[],"search":"completefinitepairhistory_finitepairhistoryoftrace_apply_of_le banditrlproof.ucb.completefinitepairhistory_finitepairhistoryoftrace_apply_of_le theorem completefinitepairhistory_finitepairhistoryoftrace_apply_of_le {action reward : type} (action : actiontrace action) (reward : rewardtrace reward) (t s : nat) (defaultaction : action) (defaultreward : reward) (hs : s <= t) : completefinitepairhistory t (history.finitepairhistoryoftrace action reward t) defaultaction defaultreward s = (action s, reward s) theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex","label":"selectedPolicySuccessorHistoryIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex","description":"Practical random-width UCB score reconstructed from one finite history.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-44cad01dbe03","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2949,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:69"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorHistoryIndex {K : Nat} (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (arm : Fin K) : Real","missing":[],"search":"selectedpolicysuccessorhistoryindex banditrlproof.ucb.selectedpolicysuccessorhistoryindex practical random-width ucb score reconstructed from one finite history. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryNextArm","label":"selectedPolicySuccessorHistoryNextArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorHistoryNextArm","description":"Round-robin initialization followed by finite-history score maximization.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-6f58567637c3","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2950,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:92"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorHistoryNextArm {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) : Fin K","missing":[],"search":"selectedpolicysuccessorhistorynextarm banditrlproof.ucb.selectedpolicysuccessorhistorynextarm round-robin initialization followed by finite-history score maximization. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex_le_nextArm_of_K_le","label":"selectedPolicySuccessorHistoryIndex_le_nextArm_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex_le_nextArm_of_K_le","description":"The post-initialization history selector maximizes every candidate score.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-f979533feab9","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2951,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:105"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorHistoryIndex_le_nextArm_of_K_le {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (ht : K <= t) (arm : Fin K) : selectedPolicySuccessorHistoryIndex sigma2 T delta defaultAction t history arm <= selectedPolicySuccessorHistoryIndex sigma2 T delta defaultAction t history (selectedPolicySuccessorHistoryNextArm hK sigma2 T delta defaultAction t history)","missing":[],"search":"selectedpolicysuccessorhistoryindex_le_nextarm_of_k_le banditrlproof.ucb.selectedpolicysuccessorhistoryindex_le_nextarm_of_k_le the post-initialization history selector maximizes every candidate score. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPairHistory","label":"selectedPolicySuccessorPairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorPairHistory","description":"Reconstruct the inclusive action/reward pair history from a finite reward history. The recursion uses only the preceding reconstructed pair prefix when selecting the next action.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-1beab219522c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2952,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:127"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorPairHistory {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) : (n : Nat) -> History.FiniteRewardHistory Rat n -> History.FinitePairHistory (Fin K) Rat n | 0, rewardHistory => fun i => (defaultAction, rewardHistory i) | n + 1, rewardHistory => let previousRewardHistory : History.FiniteRewardHistory Rat n := fun i => rewardHistory ⟨i.1, Finset.mem_Iic.mpr ((Finset.mem_Iic.mp i.2).trans (Nat.le_succ n))⟩ let previousHistory := selectedPolicySuccessorPairHistory hK sigma2 T delta defaultAction n previousRewardHistory let nextAction := selectedPolicySuccessorHistoryNextArm hK sigma2 T delta defaultAction n previousHistory History.extendPairHistorySucc previousHistory (nextAction, rewardHistory ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) /-- Package the reconstructed history as the fixed policy state type. -/ noncomputa…","missing":[],"search":"selectedpolicysuccessorpairhistory banditrlproof.ucb.selectedpolicysuccessorpairhistory reconstruct the inclusive action/reward pair history from a finite reward history. the recursion uses only the preceding reconstructed pair prefix when selecting the next action. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryState","label":"selectedPolicySuccessorHistoryState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorHistoryState","description":"Package the reconstructed history as the fixed policy state type.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-587e0061d4e9","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2953,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:151"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorHistoryState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (n : Nat) (rewardHistory : History.FiniteRewardHistory Rat n) : SelectedPolicySuccessorFiniteHistoryState K","missing":[],"search":"selectedpolicysuccessorhistorystate banditrlproof.ucb.selectedpolicysuccessorhistorystate package the reconstructed history as the fixed policy state type. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorHistoryState","label":"measurable_selectedPolicySuccessorHistoryState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorHistoryState","description":"Every finite reward-history state reconstruction is measurable.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-f9c98de96de6","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2954,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:162"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorHistoryState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (n : Nat) : Measurable (selectedPolicySuccessorHistoryState hK sigma2 T delta defaultAction n)","missing":[],"search":"measurable_selectedpolicysuccessorhistorystate banditrlproof.ucb.measurable_selectedpolicysuccessorhistorystate every finite reward-history state reconstruction is measurable. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryPolicy","label":"selectedPolicySuccessorHistoryPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorHistoryPolicy","description":"Measurable policy reading the packaged finite pair history.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-31b8944508ec","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2955,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:172"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorHistoryPolicy {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (_t : Nat) : Policy.MeasurablePolicy (SelectedPolicySuccessorFiniteHistoryState K) (Fin K) where","missing":[],"search":"selectedpolicysuccessorhistorypolicy banditrlproof.ucb.selectedpolicysuccessorhistorypolicy measurable policy reading the packaged finite pair history. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction","label":"selectedPolicySuccessorGeneratedUCBAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction","description":"Generated action trace for the concrete finite-history UCB policy.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-d935b48780c0","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2956,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:185"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorGeneratedUCBAction {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace (Fin K)","missing":[],"search":"selectedpolicysuccessorgenerateducbaction banditrlproof.ucb.selectedpolicysuccessorgenerateducbaction generated action trace for the concrete finite-history ucb policy. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction_succ","label":"selectedPolicySuccessorGeneratedUCBAction_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction_succ","description":"theorem selectedPolicySuccessorGeneratedUCBAction_succ {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) : selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega (t + 1) = selectedPolicySuccessorHistoryNextArm hK sigma2 T delta defaultAction t (selectedPolicySuccessorPa…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-0668ec52f16b","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2957,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:199"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorGeneratedUCBAction_succ {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) : selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega (t + 1) = selectedPolicySuccessorHistoryNextArm hK sigma2 T delta defaultAction t (selectedPolicySuccessorPairHistory hK sigma2 T delta defaultAction t (History.finiteRewardHistoryOfTrace (reward omega) t))","missing":[],"search":"selectedpolicysuccessorgenerateducbaction_succ banditrlproof.ucb.selectedpolicysuccessorgenerateducbaction_succ theorem selectedpolicysuccessorgenerateducbaction_succ {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (t : nat) (delta : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (t : nat) : selectedpolicysuccessorgenerateducbaction hk sigma2 t delta defaultaction reward omega (t + 1) = selectedpolicysuccessorhistorynextarm hk sigma2 t delta defaultaction t (selectedpolicysuccessorpairhistory hk sigma2 t delta defaultaction t (history.finiterewardhistoryoftrace (reward omega) t)) theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPairHistory_eq_finitePairHistoryOfTrace","label":"selectedPolicySuccessorPairHistory_eq_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorPairHistory_eq_finitePairHistoryOfTrace","description":"The reconstructed pair state is exactly the generated trace prefix.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-9398706aff9e","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2958,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:218"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorPairHistory_eq_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (n : Nat) : selectedPolicySuccessorPairHistory hK sigma2 T delta defaultAction n (History.finiteRewardHistoryOfTrace (reward omega) n) = History.finitePairHistoryOfTrace (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega) (reward omega) n","missing":[],"search":"selectedpolicysuccessorpairhistory_eq_finitepairhistoryoftrace banditrlproof.ucb.selectedpolicysuccessorpairhistory_eq_finitepairhistoryoftrace the reconstructed pair state is exactly the generated trace prefix. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sumRewards_eq_of_forall_lt","label":"sumRewards_eq_of_forall_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sumRewards_eq_of_forall_lt","description":"private theorem sumRewards_eq_of_forall_lt {Action : Type} [DecidableEq Action] (action action' : ActionTrace Action) (reward reward' : RewardTrace Real) (arm : Action) : forall n : Nat, (forall s, s < n -> action s = action' s) -> (forall s, s < n -> reward s = reward' s) -> sumRewards action reward arm n = sumRewards action' reward' arm n","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-f803c01fe4d5","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2959,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:269"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"private theorem sumRewards_eq_of_forall_lt {Action : Type} [DecidableEq Action] (action action' : ActionTrace Action) (reward reward' : RewardTrace Real) (arm : Action) : forall n : Nat, (forall s, s < n -> action s = action' s) -> (forall s, s < n -> reward s = reward' s) -> sumRewards action reward arm n = sumRewards action' reward' arm n","missing":[],"search":"sumrewards_eq_of_forall_lt banditrlproof.ucb.sumrewards_eq_of_forall_lt private theorem sumrewards_eq_of_forall_lt {action : type} [decidableeq action] (action action' : actiontrace action) (reward reward' : rewardtrace real) (arm : action) : forall n : nat, (forall s, s < n -> action s = action' s) -> (forall s, s < n -> reward s = reward' s) -> sumrewards action reward arm n = sumrewards action' reward' arm n theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmPullCount_completeFinitePairHistory","label":"successorArmPullCount_completeFinitePairHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmPullCount_completeFinitePairHistory","description":"Completed actual prefixes preserve every successor pull count.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-8e20c9cf2f1d","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2960,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:292"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_completeFinitePairHistory {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Rat) (defaultAction : Fin K) (t : Nat) (arm : Fin K) : ConditionalExpectationReward.successorArmPullCount (completeFinitePairHistoryAction t (History.finitePairHistoryOfTrace action reward t) defaultAction (0 : Rat)) arm (t + 1) = ConditionalExpectationReward.successorArmPullCount action arm (t + 1)","missing":[],"search":"successorarmpullcount_completefinitepairhistory banditrlproof.ucb.successorarmpullcount_completefinitepairhistory completed actual prefixes preserve every successor pull count. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmRewardSum_completeFinitePairHistory","label":"successorArmRewardSum_completeFinitePairHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmRewardSum_completeFinitePairHistory","description":"Completed actual prefixes preserve every successor selected reward sum.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-7a21433abc35","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2961,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:310"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmRewardSum_completeFinitePairHistory {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Rat) (defaultAction : Fin K) (t : Nat) (arm : Fin K) : ConditionalExpectationReward.successorArmRewardSum (completeFinitePairHistoryAction t (History.finitePairHistoryOfTrace action reward t) defaultAction (0 : Rat)) (completeFinitePairHistoryReward t (History.finitePairHistoryOfTrace action reward t) defaultAction (0 : Rat)) arm (t + 1) = ConditionalExpectationReward.successorArmRewardSum action reward arm (t + 1)","missing":[],"search":"successorarmrewardsum_completefinitepairhistory banditrlproof.ucb.successorarmrewardsum_completefinitepairhistory completed actual prefixes preserve every successor selected reward sum. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex_finitePairHistoryOfTrace","label":"selectedPolicySuccessorHistoryIndex_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex_finitePairHistoryOfTrace","description":"The finite-history score is exactly the score on the generated trace.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-49835a293014","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2962,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:334"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorHistoryIndex_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (omega : Omega) (t : Nat) (arm : Fin K) : selectedPolicySuccessorHistoryIndex sigma2 T delta defaultAction t (History.finitePairHistoryOfTrace (action omega) (reward omega) t) arm = selectedPolicySuccessorIndexAt action reward sigma2 (Finset.univ : Finset (Fin K)) T delta omega t arm","missing":[],"search":"selectedpolicysuccessorhistoryindex_finitepairhistoryoftrace banditrlproof.ucb.selectedpolicysuccessorhistoryindex_finitepairhistoryoftrace the finite-history score is exactly the score on the generated trace. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction_succ_eq_initializationArm_of_lt","label":"selectedPolicySuccessorGeneratedUCBAction_succ_eq_initializationArm_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction_succ_eq_initializationArm_of_lt","description":"During initialization, successor action `t + 1` follows round robin.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-c43219ed2816","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2963,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:355"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorGeneratedUCBAction_succ_eq_initializationArm_of_lt {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (ht : t < K) : selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega (t + 1) = initializationArm hK t","missing":[],"search":"selectedpolicysuccessorgenerateducbaction_succ_eq_initializationarm_of_lt banditrlproof.ucb.selectedpolicysuccessorgenerateducbaction_succ_eq_initializationarm_of_lt during initialization, successor action `t + 1` follows round robin. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_K_add_one_eq_one","label":"successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_K_add_one_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_K_add_one_eq_one","description":"Every arm appears once among successor actions `1, ..., K`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-75f564577779","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2964,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:367"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_K_add_one_eq_one {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) : ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega) arm (K + 1) = 1","missing":[],"search":"successorarmpullcount_selectedpolicysuccessorgenerateducbaction_k_add_one_eq_one banditrlproof.ucb.successorarmpullcount_selectedpolicysuccessorgenerateducbaction_k_add_one_eq_one every arm appears once among successor actions `1, ..., k`. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_pos_of_K_le","label":"successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_pos_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_pos_of_K_le","description":"After the successor initialization cycle, every arm count is positive.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-9a013aa87f4f","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2965,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:395"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_pos_of_K_le {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (ht : K <= t) : 0 < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega) arm (t + 1)","missing":[],"search":"successorarmpullcount_selectedpolicysuccessorgenerateducbaction_pos_of_k_le banditrlproof.ucb.successorarmpullcount_selectedpolicysuccessorgenerateducbaction_pos_of_k_le after the successor initialization cycle, every arm count is positive. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource","label":"selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource","description":"Concrete initialized score-max source for the practical selected-policy UCB route. Charged times are exactly the post-initialization times below `T`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-b6adb356d178","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2966,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:425"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction best : Fin K) : SelectedPolicySuccessorInitializedScoreMaxSource (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward) reward Finset.univ armMean sigma2 T delta where","missing":[],"search":"selectedpolicysuccessorgenerateducbinitializedscoremaxsource banditrlproof.ucb.selectedpolicysuccessorgenerateducbinitializedscoremaxsource concrete initialized score-max source for the practical selected-policy ucb route. charged times are exactly the post-initialization times below `t`. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.exists_selected_with_threshold_le_prior_pullCount","label":"exists_selected_with_threshold_le_prior_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.exists_selected_with_threshold_le_prior_pullCount","description":"If the final pull count exceeds `B`, some selected time has prior pull count at least `B`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-d5ca1b5d4d5c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2967,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:504"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem exists_selected_with_threshold_le_prior_pullCount {Action : Type} [DecidableEq Action] (action : ActionTrace Action) (arm : Action) (T B : Nat) (hcount : B < pullCount action arm T) : exists t, t < T ∧ action t = arm ∧ B <= pullCount action arm t","missing":[],"search":"exists_selected_with_threshold_le_prior_pullcount banditrlproof.ucb.exists_selected_with_threshold_le_prior_pullcount if the final pull count exceeds `b`, some selected time has prior pull count at least `b`. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget","label":"selectedPolicySuccessorFiniteArmTimeLogBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget","description":"Log budget hidden inside one finite-arm/time peeling radius.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-57749e10303a","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2968,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:524"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorFiniteArmTimeLogBudget (K T n : Nat) (delta : Real) : Real","missing":[],"search":"selectedpolicysuccessorfinitearmtimelogbudget banditrlproof.ucb.selectedpolicysuccessorfinitearmtimelogbudget log budget hidden inside one finite-arm/time peeling radius. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_le_horizon","label":"selectedPolicySuccessorFiniteArmTimeLogBudget_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_le_horizon","description":"The local log budget is maximized at the full positive horizon.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-98bc60387dce","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2969,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:532"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmTimeLogBudget_le_horizon (K T n : Nat) (delta : Real) (hK : 0 < K) (hT : 0 < T) (hnT : n <= T) (hdelta : 0 < delta) : selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta <= selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta","missing":[],"search":"selectedpolicysuccessorfinitearmtimelogbudget_le_horizon banditrlproof.ucb.selectedpolicysuccessorfinitearmtimelogbudget_le_horizon the local log budget is maximized at the full positive horizon. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRealPullThreshold","label":"selectedPolicySuccessorRealPullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorRealPullThreshold","description":"Real count threshold that simultaneously dominates the quadratic and linear parts of the practical random-width radius inversion.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-1b7391f4094a","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2970,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:576"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorRealPullThreshold (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) : Real","missing":[],"search":"selectedpolicysuccessorrealpullthreshold banditrlproof.ucb.selectedpolicysuccessorrealpullthreshold real count threshold that simultaneously dominates the quadratic and linear parts of the practical random-width radius inversion. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPullThreshold","label":"selectedPolicySuccessorPullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorPullThreshold","description":"One more than the ceiling supplies the strict margin needed by the radius.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-efd6c1d11b7d","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2971,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:584"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorPullThreshold (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) : Nat","missing":[],"search":"selectedpolicysuccessorpullthreshold banditrlproof.ucb.selectedpolicysuccessorpullthreshold one more than the ceiling supplies the strict margin needed by the radius. definition compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPullThreshold_contracts","label":"selectedPolicySuccessorPullThreshold_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorPullThreshold_contracts","description":"The explicit integer threshold is positive and satisfies both full-horizon strict radius-inversion inequalities.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-9d0d8fa72497","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2972,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:593"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorPullThreshold_contracts (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) (hgap : 0 < gap) : 0 < selectedPolicySuccessorPullThreshold K sigma2 T delta gap /\\ 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < gap ^ 2 * (selectedPolicySuccessorPullThreshold K sigma2 T delta gap : Real) /\\ 4 * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < gap * (selectedPolicySuccessorPullThreshold K sigma2 T delta gap : Real)","missing":[],"search":"selectedpolicysuccessorpullthreshold_contracts banditrlproof.ucb.selectedpolicysuccessorpullthreshold_contracts the explicit integer threshold is positive and satisfies both full-horizon strict radius-inversion inequalities. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_eq","label":"successorArmEmpiricalMeanFiniteArmTimePeelingRadius_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_eq","description":"Expanded algebraic form of the practical finite-arm/time radius.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-aeb3c75dda02","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2973,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:645"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMeanFiniteArmTimePeelingRadius_eq {K : Nat} (sigma2 : NNReal) (k n T : Nat) (delta : Real) : ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta = (2 * Real.sqrt ((1 / 2 : Real) * (((sigma2 : NNReal) : Real) * (k : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta) + selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta) / (k : Real)","missing":[],"search":"successorarmempiricalmeanfinitearmtimepeelingradius_eq banditrlproof.ucb.successorarmempiricalmeanfinitearmtimepeelingradius_eq expanded algebraic form of the practical finite-arm/time radius. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.two_mul_successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap","label":"two_mul_successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.two_mul_successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap","description":"Sufficient square-root and linear inequalities for inverting one realized-count peeling radius below half a positive arm gap.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-ba6d1a3a8312","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2974,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:666"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem two_mul_successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap {K : Nat} (sigma2 : NNReal) (k n T : Nat) (delta gap : Real) (hk : 0 < k) (hgap : 0 < gap) (hquadratic : 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta < gap ^ 2 * (k : Real)) (hlinear : 4 * selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta < gap * (k : Real)) : 2 * ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta < gap","missing":[],"search":"two_mul_successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap banditrlproof.ucb.two_mul_successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap sufficient square-root and linear inequalities for inverting one realized-count peeling radius below half a positive arm gap. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_threshold","label":"successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_threshold","description":"Uniform count-threshold inversion. It is enough to check the quadratic and linear log-budget inequalities at the threshold `B`; larger realized counts only improve the radius.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-5380ee2cf06c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2975,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:721"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_threshold {K : Nat} (sigma2 : NNReal) (T : Nat) (delta gap : Real) (B : Nat) (hB : 0 < B) (hgap : 0 < gap) (hquadratic : forall n : Nat, n <= T -> 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta < gap ^ 2 * (B : Real)) (hlinear : forall n : Nat, n <= T -> 4 * selectedPolicySuccessorFiniteArmTimeLogBudget K T n delta < gap * (B : Real)) : forall k n : Nat, B <= k -> n <= T -> 2 * ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta < gap","missing":[],"search":"successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap_of_threshold banditrlproof.ucb.successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap_of_threshold uniform count-threshold inversion. it is enough to check the quadratic and linear log-budget inequalities at the threshold `b`; larger realized counts only improve the radius. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_global_threshold","label":"successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_global_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_global_threshold","description":"Full-horizon sufficient condition for uniform radius inversion. The finite arm/time/count peeling log cost is summarized by the single deterministic budget at `n = T`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-a8a8c33a1046","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2976,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:759"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_global_threshold {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) (gap : Real) (hgap : 0 < gap) (B : Nat) (hB : 0 < B) (hquadratic : 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < gap ^ 2 * (B : Real)) (hlinear : 4 * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < gap * (B : Real)) : forall k n : Nat, B <= k -> n <= T -> 2 * ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta < gap","missing":[],"search":"successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap_of_global_threshold banditrlproof.ucb.successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap_of_global_threshold full-horizon sufficient condition for uniform radius inversion. the finite arm/time/count peeling log cost is summarized by the single deterministic budget at `n = t`. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_explicitPullThreshold","label":"successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_explicitPullThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_explicitPullThreshold","description":"Uniform radius inversion at the explicit one-more-than-ceiling pull threshold.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-bdabaa3d623d","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2977,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:794"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_explicitPullThreshold {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) (gap : Real) (hgap : 0 < gap) : forall k n : Nat, selectedPolicySuccessorPullThreshold K sigma2 T delta gap <= k -> n <= T -> 2 * ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta < gap","missing":[],"search":"successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap_of_explicitpullthreshold banditrlproof.ucb.successorarmempiricalmeanfinitearmtimepeelingradius_lt_gap_of_explicitpullthreshold uniform radius inversion at the explicit one-more-than-ceiling pull threshold. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.K_le_of_selectedPolicySuccessorGeneratedUCBAction_selected_and_count_pos","label":"K_le_of_selectedPolicySuccessorGeneratedUCBAction_selected_and_count_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.K_le_of_selectedPolicySuccessorGeneratedUCBAction_selected_and_count_pos","description":"A selected time with a positive prior count cannot lie inside the one-pass round-robin initialization prefix.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-fefe81b42899","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2978,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:818"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem K_le_of_selectedPolicySuccessorGeneratedUCBAction_selected_and_count_pos {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (hselected : selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega (t + 1) = arm) (hcount : 0 < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega) arm (t + 1)) : K <= t","missing":[],"search":"k_le_of_selectedpolicysuccessorgenerateducbaction_selected_and_count_pos banditrlproof.ucb.k_le_of_selectedpolicysuccessorgenerateducbaction_selected_and_count_pos a selected time with a positive prior count cannot lie inside the one-pass round-robin initialization prefix. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_largeGap","label":"measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_largeGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_largeGap","description":"Generated-policy high-probability pull-count consumer. The deterministic `hradius` contract is the exact remaining radius-inversion obligation: every count at least `B`, at every horizon at most `T`, must make twice the realized confidence radius strictly smaller than the chosen arm gap. Under that contract, exceeding `B` pulls forces the global large-gap event.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-2fb4379cac01","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2979,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:862"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_largeGap {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction best chosen : Fin K) (B : Nat) (hB : 0 < B) (hradius : forall k n : Nat, B <= k -> n <= T -> 2 * ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta < meanGap (fun arm => (armMean arm : Real)) best chosen) (hlargeGap : mu (selectedPolicySuccessorLargeGapEvent (selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource hK reward armMean sigma2 T delta defaultAction best)) <= ENNReal.ofReal delta) : mu {omega | B < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBActi…","missing":[],"search":"measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_threshold_le_of_largegap banditrlproof.ucb.measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_threshold_le_of_largegap generated-policy high-probability pull-count consumer. the deterministic `hradius` contract is the exact remaining radius-inversion obligation: every count at least `b`, at every horizon at most `t`, must make twice the realized confidence radius strictly smaller than the chosen arm gap. under that contract, exceeding `b` pulls forces the global large-gap event. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_global_threshold","label":"measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_global_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_global_threshold","description":"Generated-policy high-probability pull-count bound with the radius inversion discharged by the two full-horizon numeric threshold inequalities.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-90bf4af02bab","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2980,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:931"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_global_threshold {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (B : Nat) (hB : 0 < B) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hquadratic : 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < meanGap (fun arm => (armMean arm : Real)) best chosen ^ 2 * (B : Real)) (hlinear : 4 * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < meanGap (fun arm => (armMean arm : Real)) best chosen * (B : Real)) (hlargeGap : mu (selectedPolicySuccessorLargeGapEvent (selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSo…","missing":[],"search":"measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_threshold_le_of_global_threshold banditrlproof.ucb.measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_threshold_le_of_global_threshold generated-policy high-probability pull-count bound with the radius inversion discharged by the two full-horizon numeric threshold inequalities. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_of_largeGap","label":"measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_of_largeGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_of_largeGap","description":"Generated-policy pull-count tail at the explicit one-more-than-ceiling threshold; no caller-supplied radius or numeric threshold inequalities remain.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-724d8cde1577","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2981,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:974"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_of_largeGap {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hlargeGap : mu (selectedPolicySuccessorLargeGapEvent (selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource hK reward armMean sigma2 T delta defaultAction best)) <= ENNReal.ofReal delta) : mu {omega | selectedPolicySuccessorPullThreshold K sigma2 T delta (meanGap (fun arm => (armMean arm : Real)) best chosen) < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward o…","missing":[],"search":"measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_explicitpullthreshold_le_of_largegap banditrlproof.ucb.measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_explicitpullthreshold_le_of_largegap generated-policy pull-count tail at the explicit one-more-than-ceiling threshold; no caller-supplied radius or numeric threshold inequalities remain. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy","label":"measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy","description":"Practical selected-reward-law producer for the concrete generated UCB source's global random-width large-gap event.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-6ef6440437ea","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2982,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1012"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best : Fin K) (armMean : Fin K -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.Cent…","missing":[],"search":"measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_of_reward_map_eq_selected_policy banditrlproof.ucb.measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_of_reward_map_eq_selected_policy practical selected-reward-law producer for the concrete generated ucb source's global random-width large-gap event. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy","label":"measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy","description":"Practical conditional-reward-law endpoint for the concrete generated UCB policy. The only algorithmic numeric remainder is the explicit deterministic radius-inversion contract `hradius`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-e10bc613fcb2","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2983,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1152"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best chosen : Fin K) (armMean : Fin K -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pai…","missing":[],"search":"measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy banditrlproof.ucb.measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy practical conditional-reward-law endpoint for the concrete generated ucb policy. the only algorithmic numeric remainder is the explicit deterministic radius-inversion contract `hradius`. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorGeneratedUCBAction","label":"measurable_selectedPolicySuccessorGeneratedUCBAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorGeneratedUCBAction","description":"Timewise measurability of the concrete generated UCB action trace.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-4318d1451375","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2984,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1303"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorGeneratedUCBAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega t)","missing":[],"search":"measurable_selectedpolicysuccessorgenerateducbaction banditrlproof.ucb.measurable_selectedpolicysuccessorgenerateducbaction timewise measurability of the concrete generated ucb action trace. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_natCast_le_threshold_add_bound_mul_of_measure_gt","label":"lintegral_natCast_le_threshold_add_bound_mul_of_measure_gt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_natCast_le_threshold_add_bound_mul_of_measure_gt","description":"Integrate a bounded Nat-valued random variable from one upper-tail probability bound. This is the exact `B + horizon * delta` layer used by the generated UCB pull-count theorem below.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-5585ec979e42","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2985,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1332"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_natCast_le_threshold_add_bound_mul_of_measure_gt {Omega : Type} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (count : Omega -> Nat) (hcount : Measurable count) (threshold bound : Nat) (hbound : forall omega, count omega <= bound) (epsilon : ENNReal) (htail : mu {omega | threshold < count omega} <= epsilon) : ∫⁻ omega, (count omega : ENNReal) ∂mu <= (threshold : ENNReal) + (bound : ENNReal) * epsilon","missing":[],"search":"lintegral_natcast_le_threshold_add_bound_mul_of_measure_gt banditrlproof.ucb.lintegral_natcast_le_threshold_add_bound_mul_of_measure_gt integrate a bounded nat-valued random variable from one upper-tail probability bound. this is the exact `b + horizon * delta` layer used by the generated ucb pull-count theorem below. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_largeGap","label":"lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_largeGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_largeGap","description":"ENNReal expected pull-count bound for the concrete generated UCB process from the global large-gap probability theorem and the deterministic radius inversion contract.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-aab5b9287284","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2986,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1386"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_largeGap {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction best chosen : Fin K) (B : Nat) (hB : 0 < B) (hradius : forall k n : Nat, B <= k -> n <= T -> 2 * ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius sigma2 k n (Finset.univ : Finset (Fin K)) T delta < meanGap (fun arm => (armMean arm : Real)) best chosen) (hlargeGap : mu (selectedPolicySuccessorLargeGapEvent (selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource hK reward armMean sigma2 T delta defaultAction best)) <= ENNReal…","missing":[],"search":"lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_threshold_add_horizon_mul_delta_of_largegap banditrlproof.ucb.lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_threshold_add_horizon_mul_delta_of_largegap ennreal expected pull-count bound for the concrete generated ucb process from the global large-gap probability theorem and the deterministic radius inversion contract. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_global_threshold","label":"lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_global_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_global_threshold","description":"ENNReal expected pull-count bound whose algorithmic remainder is stated only through the two full-horizon numeric threshold inequalities.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-1c4a72d8bbc2","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2987,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1446"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_global_threshold {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (B : Nat) (hB : 0 < B) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hquadratic : 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < meanGap (fun arm => (armMean arm : Real)) best chosen ^ 2 * (B : Real)) (hlinear : 4 * selectedPolicySuccessorFiniteArmTimeLogBudget K T T delta < meanGap (fun arm => (armMean arm : Real)) best chosen *…","missing":[],"search":"lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_threshold_add_horizon_mul_delta_of_global_threshold banditrlproof.ucb.lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_threshold_add_horizon_mul_delta_of_global_threshold ennreal expected pull-count bound whose algorithmic remainder is stated only through the two full-horizon numeric threshold inequalities. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_largeGap","label":"lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_largeGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_largeGap","description":"ENNReal expected pull-count bound at the explicit one-more-than-ceiling threshold; the numeric radius inversion is fully discharged internally.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-fc1b36e10f62","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2988,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1491"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_largeGap {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hlargeGap : mu (selectedPolicySuccessorLargeGapEvent (selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource hK reward armMean sigma2 T delta defaultAction best)) <= ENNReal.ofReal delta) : ∫⁻ omega, (ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultActi…","missing":[],"search":"lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_of_largegap banditrlproof.ucb.lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_of_largegap ennreal expected pull-count bound at the explicit one-more-than-ceiling threshold; the numeric radius inversion is fully discharged internally. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy","label":"lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy","description":"End-to-end practical selected-reward-law expected pull-count theorem for the concrete generated UCB policy at its explicit integer threshold.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawpolicy/index.html#decl-443228c63644","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","order":2989,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawPolicy.lean:1531"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best chosen : Fin K) (armMean : Fin K -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K…","missing":[],"search":"lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy banditrlproof.ucb.lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy end-to-end practical selected-reward-law expected pull-count theorem for the concrete generated ucb policy at its explicit integer threshold. theorem compiled","shard":"modules/7a549d50bf316db6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.modelMeanGap_bestArm_eq_realGap","label":"modelMeanGap_bestArm_eq_realGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.modelMeanGap_bestArm_eq_realGap","description":"The UCB designated-best Real mean gap is the local model gap after casting.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-5a1cf8483c8c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2990,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:19"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem modelMeanGap_bestArm_eq_realGap {K : Nat} (model : FiniteBanditModel K) (arm : Fin K) : meanGap (fun a => ((model.mean a : Rat) : Real)) model.bestArm arm = ((model.gap arm : Rat) : Real)","missing":[],"search":"modelmeangap_bestarm_eq_realgap banditrlproof.ucb.modelmeangap_bestarm_eq_realgap the ucb designated-best real mean gap is the local model gap after casting. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget","label":"selectedPolicySuccessorTextbookGapBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget","description":"Textbook-style real contribution of one positive-gap arm's count threshold.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-9e86cb662b94","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2991,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:29"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTextbookGapBudget (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) : Real","missing":[],"search":"selectedpolicysuccessortextbookgapbudget banditrlproof.ucb.selectedpolicysuccessortextbookgapbudget textbook-style real contribution of one positive-gap arm's count threshold. definition compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPullThreshold_cast_le_realThreshold_add_two","label":"selectedPolicySuccessorPullThreshold_cast_le_realThreshold_add_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorPullThreshold_cast_le_realThreshold_add_two","description":"The explicit integer pull threshold is at most its real max plus two.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-ce6fbc08c2a2","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2992,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:37"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorPullThreshold_cast_le_realThreshold_add_two (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) (hgap : 0 < gap) : (selectedPolicySuccessorPullThreshold K sigma2 T delta gap : Real) <= selectedPolicySuccessorRealPullThreshold K sigma2 T delta gap + 2","missing":[],"search":"selectedpolicysuccessorpullthreshold_cast_le_realthreshold_add_two banditrlproof.ucb.selectedpolicysuccessorpullthreshold_cast_le_realthreshold_add_two the explicit integer pull threshold is at most its real max plus two. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.gap_mul_selectedPolicySuccessorPullThreshold_cast_le_textbookGapBudget","label":"gap_mul_selectedPolicySuccessorPullThreshold_cast_le_textbookGapBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.gap_mul_selectedPolicySuccessorPullThreshold_cast_le_textbookGapBudget","description":"Multiplying the explicit threshold by a positive gap removes one gap power.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-841a7473251a","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2993,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:55"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem gap_mul_selectedPolicySuccessorPullThreshold_cast_le_textbookGapBudget (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) (hgap : 0 < gap) : gap * (selectedPolicySuccessorPullThreshold K sigma2 T delta gap : Real) <= selectedPolicySuccessorTextbookGapBudget K sigma2 T delta gap","missing":[],"search":"gap_mul_selectedpolicysuccessorpullthreshold_cast_le_textbookgapbudget banditrlproof.ucb.gap_mul_selectedpolicysuccessorpullthreshold_cast_le_textbookgapbudget multiplying the explicit threshold by a positive gap removes one gap power. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.ofReal_gap_mul_selectedPolicySuccessorPullThreshold_le_textbookGapBudget","label":"ofReal_gap_mul_selectedPolicySuccessorPullThreshold_le_textbookGapBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.ofReal_gap_mul_selectedPolicySuccessorPullThreshold_le_textbookGapBudget","description":"ENNReal form of the one-arm textbook threshold contribution.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-108ca15e094e","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2994,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:93"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem ofReal_gap_mul_selectedPolicySuccessorPullThreshold_le_textbookGapBudget (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) (hgap : 0 < gap) : ENNReal.ofReal gap * (selectedPolicySuccessorPullThreshold K sigma2 T delta gap : ENNReal) <= ENNReal.ofReal (selectedPolicySuccessorTextbookGapBudget K sigma2 T delta gap)","missing":[],"search":"ofreal_gap_mul_selectedpolicysuccessorpullthreshold_le_textbookgapbudget banditrlproof.ucb.ofreal_gap_mul_selectedpolicysuccessorpullthreshold_le_textbookgapbudget ennreal form of the one-arm textbook threshold contribution. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.sum_gap_mul_explicitThreshold_add_failure_le_textbookGapSum","label":"sum_gap_mul_explicitThreshold_add_failure_le_textbookGapSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.sum_gap_mul_explicitThreshold_add_failure_le_textbookGapSum","description":"Finite-arm threshold simplification. Only positive model gaps remain in the textbook sum, and the confidence-failure contribution is preserved exactly.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-4dfc063e0d1f","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2995,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:111"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem sum_gap_mul_explicitThreshold_add_failure_le_textbookGapSum {K : Nat} (model : FiniteBanditModel K) (sigma2 : NNReal) (T : Nat) (delta : Real) : (Finset.univ : Finset (Fin K)).sum (fun arm => ENNReal.ofReal (((model.gap arm : Rat) : Real)) * (selectedPolicySuccessorPullThreshold K sigma2 T delta (((model.gap arm : Rat) : Real)) : ENNReal) + ENNReal.ofReal (((model.gap arm : Rat) : Real)) * ((T : ENNReal) * ENNReal.ofReal delta)) <= ((Finset.univ : Finset (Fin K)).filter (fun arm => 0 < (((model.gap arm : Rat) : Real)))).sum (fun arm => ENNReal.ofReal (selectedPolicySuccessorTextbookGapBudget K sigma2 T delta (((model.gap arm : Rat) : Real))) + ENNReal.ofReal (((model.gap arm : Rat) : Real)) * ((T : ENNReal) * ENNReal.ofReal delta))","missing":[],"search":"sum_gap_mul_explicitthreshold_add_failure_le_textbookgapsum banditrlproof.ucb.sum_gap_mul_explicitthreshold_add_failure_le_textbookgapsum finite-arm threshold simplification. only positive model gaps remain in the textbook sum, and the confidence-failure contribution is preserved exactly. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBRegretAction","label":"selectedPolicySuccessorGeneratedUCBRegretAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBRegretAction","description":"Shift successor generated actions `1, ..., T` to regret times `0, ..., T-1`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-ddc9ab3a8ba0","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2996,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:144"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorGeneratedUCBRegretAction {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace (Fin K)","missing":[],"search":"selectedpolicysuccessorgenerateducbregretaction banditrlproof.ucb.selectedpolicysuccessorgenerateducbregretaction shift successor generated actions `1, ..., t` to regret times `0, ..., t-1`. definition compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorGeneratedUCBRegretAction","label":"measurable_selectedPolicySuccessorGeneratedUCBRegretAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorGeneratedUCBRegretAction","description":"Every coordinate of the shifted generated UCB regret action is measurable.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-b9f1113243f6","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2997,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:154"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorGeneratedUCBRegretAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => selectedPolicySuccessorGeneratedUCBRegretAction hK sigma2 T delta defaultAction reward omega t)","missing":[],"search":"measurable_selectedpolicysuccessorgenerateducbregretaction banditrlproof.ucb.measurable_selectedpolicysuccessorgenerateducbregretaction every coordinate of the shifted generated ucb regret action is measurable. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullCount_selectedPolicySuccessorGeneratedUCBRegretAction_eq","label":"pullCount_selectedPolicySuccessorGeneratedUCBRegretAction_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pullCount_selectedPolicySuccessorGeneratedUCBRegretAction_eq","description":"The shifted regret pull count is exactly the existing successor pull count.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-cb07581cda16","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2998,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:170"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pullCount_selectedPolicySuccessorGeneratedUCBRegretAction_eq {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) : pullCount (selectedPolicySuccessorGeneratedUCBRegretAction hK sigma2 T delta defaultAction reward omega) arm T = ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorGeneratedUCBAction hK sigma2 T delta defaultAction reward omega) arm (T + 1)","missing":[],"search":"pullcount_selectedpolicysuccessorgenerateducbregretaction_eq banditrlproof.ucb.pullcount_selectedpolicysuccessorgenerateducbregretaction_eq the shifted regret pull count is exactly the existing successor pull count. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_le_sum_gap_mul_bound_of_positiveGap_pullCount","label":"lintegral_ofReal_pseudoRegret_le_sum_gap_mul_bound_of_positiveGap_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_le_sum_gap_mul_bound_of_positiveGap_pullCount","description":"Generic finite-arm ENNReal pseudo-regret assembly. Only positive-gap arms need a pull-count bound; zero-gap arms disappear after multiplication.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-beb216bef5e7","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":2999,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:191"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_le_sum_gap_mul_bound_of_positiveGap_pullCount {Omega : Type} [MeasurableSpace Omega] {K : Nat} (mu : Measure Omega) (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (n : Nat) (bound : Fin K -> ENNReal) (hcount : forall arm : Fin K, 0 < (((model.gap arm : Rat) : Real)) -> ∫⁻ omega, ((pullCount (action omega) arm n : Nat) : ENNReal) ∂mu <= bound arm) : ∫⁻ omega, ENNReal.ofReal (((pseudoRegret model (action omega) n : Rat) : Real)) ∂mu <= (Finset.univ : Finset (Fin K)).sum (fun arm => ENNReal.ofReal (((model.gap arm : Rat) : Real)) * bound arm)","missing":[],"search":"lintegral_ofreal_pseudoregret_le_sum_gap_mul_bound_of_positivegap_pullcount banditrlproof.ucb.lintegral_ofreal_pseudoregret_le_sum_gap_mul_bound_of_positivegap_pullcount generic finite-arm ennreal pseudo-regret assembly. only positive-gap arms need a pull-count bound; zero-gap arms disappear after multiplication. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy","description":"End-to-end practical selected-reward-law pseudo-regret bound for the concrete generated UCB policy. Every positive-gap arm uses its own explicit threshold; zero-gap arms vanish from the finite weighted sum.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-bfa751d89558","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":3000,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:262"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (model : FiniteBanditModel K) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : Rew…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_of_reward_map_eq_selected_policy banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_of_reward_map_eq_selected_policy end-to-end practical selected-reward-law pseudo-regret bound for the concrete generated ucb policy. every positive-gap arm uses its own explicit threshold; zero-gap arms vanish from the finite weighted sum. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy","description":"End-to-end practical selected-reward-law pseudo-regret bound with the integer threshold eliminated in favor of a textbook reciprocal-gap finite sum.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawregret/index.html#decl-6808da547841","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","order":3001,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawRegret"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawRegret.lean:427"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy {Omega Context : Type} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] (mu : Measure Omega) [IsProbabilityMeasure mu] (model : FiniteBanditModel K) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKer…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_of_reward_map_eq_selected_policy banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_of_reward_map_eq_selected_policy end-to-end practical selected-reward-law pseudo-regret bound with the integer threshold eliminated in favor of a textbook reciprocal-gap finite sum. theorem compiled","shard":"modules/1f1de711bc498fa3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRewardStepKernelFamily","label":"selectedPolicySuccessorRewardStepKernelFamily","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorRewardStepKernelFamily","description":"Reward-only history-step kernels for the practical generated UCB policy.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html#decl-859dc8005e00","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","order":3002,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean:18"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorRewardStepKernelFamily {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K)","missing":[],"search":"selectedpolicysuccessorrewardstepkernelfamily banditrlproof.ucb.selectedpolicysuccessorrewardstepkernelfamily reward-only history-step kernels for the practical generated ucb policy. definition compiled","shard":"modules/df390a781c1ccf9d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.isMarkovKernel_selectedPolicySuccessorRewardStepKernelFamily","label":"isMarkovKernel_selectedPolicySuccessorRewardStepKernelFamily","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.isMarkovKernel_selectedPolicySuccessorRewardStepKernelFamily","description":"Every UCB reward-only history-step kernel is Markov.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html#decl-a61334f3237e","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","order":3003,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean:37"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_selectedPolicySuccessorRewardStepKernelFamily {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) : forall n : Nat, ProbabilityTheory.IsMarkovKernel (selectedPolicySuccessorRewardStepKernelFamily hK rewardKernel context hcontext sigma2 T delta defaultAction n)","missing":[],"search":"ismarkovkernel_selectedpolicysuccessorrewardstepkernelfamily banditrlproof.ucb.ismarkovkernel_selectedpolicysuccessorrewardstepkernelfamily every ucb reward-only history-step kernel is markov. theorem compiled","shard":"modules/df390a781c1ccf9d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRewardTrajMeasure","label":"selectedPolicySuccessorRewardTrajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorRewardTrajMeasure","description":"Canonical reward-only trajectory measure for the practical generated UCB policy.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html#decl-36468b4c71c3","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","order":3004,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean:53"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorRewardTrajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) : Measure (RewardTrace Rat)","missing":[],"search":"selectedpolicysuccessorrewardtrajmeasure banditrlproof.ucb.selectedpolicysuccessorrewardtrajmeasure canonical reward-only trajectory measure for the practical generated ucb policy. definition compiled","shard":"modules/df390a781c1ccf9d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBSelectedRewardLawSource_trajMeasure","label":"selectedPolicySuccessorGeneratedUCBSelectedRewardLawSource_trajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBSelectedRewardLawSource_trajMeasure","description":"Canonical selected-reward law source for the practical generated UCB policy. The source constructor transports the canonical comap-trim law to the generated history filtration.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html#decl-73f244203aec","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","order":3005,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean:99"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorGeneratedUCBSelectedRewardLawSource_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) : ConditionalExpectationReward.GeneratedActionSelectedRewardFinitePairHistoryLawSource (selectedPolicySuccessorRewardTrajMeasure hK mu0 rewardKernel context hcontext sigma2 T delta defaultAction) rewardKernel (selectedPolicySuccessorHistoryPolicy hK sigma2 T delta defaultAction) context (selectedPolicySuccessorHistoryState hK sigma2 T delta defaultAction) defaultAction (fun trajectory : RewardTrace Rat => trajectory) (fun t => measur…","missing":[],"search":"selectedpolicysuccessorgenerateducbselectedrewardlawsource_trajmeasure banditrlproof.ucb.selectedpolicysuccessorgenerateducbselectedrewardlawsource_trajmeasure canonical selected-reward law source for the practical generated ucb policy. the source constructor transports the canonical comap-trim law to the generated history filtration. definition compiled","shard":"modules/df390a781c1ccf9d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCB_reward_map_eq_selected_policy_trajMeasure","label":"selectedPolicySuccessorGeneratedUCB_reward_map_eq_selected_policy_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCB_reward_map_eq_selected_policy_trajMeasure","description":"The canonical UCB reward-only trajectory measure satisfies the exact `historyFiltrationSucc` selected-reward law consumed by the practical regret theorem.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html#decl-2e74a231c402","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","order":3006,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean:157"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorGeneratedUCB_reward_map_eq_selected_policy_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (T : Nat) (delta : Real) (defaultAction : Fin K) (i : Nat) : let mu := selectedPolicySuccessorRewardTrajMeasure hK mu0 rewardKernel context hcontext sigma2 T delta defaultAction let reward : RewardTrace Rat -> RewardTrace Rat := fun trajectory => trajectory Filter.Eventually (fun trajectory : RewardTrace Rat => @Measure.map (RewardTrace Rat) Rat inferInstance inferInstance (fun y : RewardTrace Rat => reward y (i + 1)) (@ProbabilityTheory.condExpKernel (RewardTrace Rat) inferInstance _…","missing":[],"search":"selectedpolicysuccessorgenerateducb_reward_map_eq_selected_policy_trajmeasure banditrlproof.ucb.selectedpolicysuccessorgenerateducb_reward_map_eq_selected_policy_trajmeasure the canonical ucb reward-only trajectory measure satisfies the exact `historyfiltrationsucc` selected-reward law consumed by the practical regret theorem. theorem compiled","shard":"modules/df390a781c1ccf9d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure","description":"Canonical reward-only trajectory specialization of the practical textbook UCB pseudo-regret theorem. The selected-reward law is produced internally; the pointwise raw-range premise is retained explicitly.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardlawtrajmeasure/index.html#decl-fa027ded246f","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","order":3007,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardLawTrajMeasure.lean:215"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Context × Fin K => mean pair.1 pair.2)) (hkernel : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hraw : forall i : Nat, forall trajectory : RewardTrace Rat, Set.Icc (rewardLo i) (rewardHi i) (((trajectory (i + 1) : Rat) : Real))) (hmea…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure canonical reward-only trajectory specialization of the practical textbook ucb pseudo-regret theorem. the selected-reward law is produced internally; the pointwise raw-range premise is retained explicitly. theorem compiled","shard":"modules/df390a781c1ccf9d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_actionRewardHistoryStepKernelFamily_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_trajMeasure","label":"measure_actionRewardHistoryStepKernelFamily_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_actionRewardHistoryStepKernelFamily_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_trajMeasure","description":"theorem measure_actionRewardHistoryStepKernelFamily_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_trajMeasure {Context : Type u} {State : Type v} {Action : Type} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasur…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html#decl-a56d29eb2ee9","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","order":3008,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean:34"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_actionRewardHistoryStepKernelFamily_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_trajMeasure {Context : Type u} {State : Type v} {Action : Type} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKerne…","missing":[],"search":"measure_actionrewardhistorystepkernelfamily_selectedpolicysuccessorlargegapevent_le_ennreal_delta_trajmeasure banditrlproof.ucb.measure_actionrewardhistorystepkernelfamily_selectedpolicysuccessorlargegapevent_le_ennreal_delta_trajmeasure theorem measure_actionrewardhistorystepkernelfamily_selectedpolicysuccessorlargegapevent_le_ennreal_delta_trajmeasure {context : type u} {state : type v} {action : type} [measurablespace context] [measurablespace state] [measurablespace action] [standardborelspace action] [measurablesingletonclass action] [countable action] [nonempty action] [decidableeq action] (mu0 : measure (prod action rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context action) rat) (policy : nat -> policy.measurablepolicy state action) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (state : (n : nat) -> ((j : finset.iic n) -> rat) -> state) (hcontext : forall n : nat, measurable (context n)) (hstate : forall n : nat, measurable (state n)) (mean : context -> action -> rat) (varianceproxy : context -> action -> nnreal) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hmean : measurable (fun pair : prod context action => mean pair.1 pair.2)) (arms : finset action) (harms : arms.nonempty) (armmean : action -> rat) (sigma2 : nnreal) (t : nat) (hvariance : forall i : nat, i < t - 1 -> foral…","shard":"modules/defbb2ae1e1d9c48.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","label":"measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","description":"theorem measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K ->…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html#decl-1da98d103134","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","order":3009,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean:163"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((se…","missing":[],"search":"measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_actionrewardtrajmeasure_centeredkernel theorem measure_selectedpolicysuccessorlargegapevent_generateducb_le_ennreal_delta_actionrewardtrajmeasure_centeredkernel {context : type u} {k : nat} [measurablespace context] (hk : 0 < k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction best : fin k) (armmean : fin k -> rat) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, i < t - 1 -> forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((selectedpolicysuccessorhistorypolicy hk sigma2 t delta defaultaction i).action (selectedpolicysuccessorhistorystate hk sigma2 t delta defaultaction i history)) <= sigma2) (harmmean : forall i : nat, forall history : ((j : finset.iic i) -> rat), forall arm…","shard":"modules/defbb2ae1e1d9c48.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","label":"measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","description":"theorem measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat)…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html#decl-0dbc67085368","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","order":3010,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean:361"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best chosen : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.Iic i) -…","missing":[],"search":"measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_explicitpullthreshold_le_ennreal_delta_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_explicitpullthreshold_le_ennreal_delta_actionrewardtrajmeasure_centeredkernel theorem measure_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_gt_explicitpullthreshold_le_ennreal_delta_actionrewardtrajmeasure_centeredkernel {context : type u} {k : nat} [measurablespace context] (hk : 0 < k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction best chosen : fin k) (armmean : fin k -> rat) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, i < t - 1 -> forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((selectedpolicysuccessorhistorypolicy hk sigma2 t delta defaultaction i).action (selectedpolicysuccessorhistorystate hk sigma2 t delt…","shard":"modules/defbb2ae1e1d9c48.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_actionRewardTrajMeasure_centeredKernel","label":"lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_actionRewardTrajMeasure_centeredKernel","description":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html#decl-0a46beeb2133","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","order":3011,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean:485"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best chosen : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.I…","missing":[],"search":"lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_actionrewardtrajmeasure_centeredkernel theorem lintegral_successorarmpullcount_selectedpolicysuccessorgenerateducbaction_le_explicitpullthreshold_add_horizon_mul_delta_actionrewardtrajmeasure_centeredkernel {context : type u} {k : nat} [measurablespace context] (hk : 0 < k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction best chosen : fin k) (armmean : fin k -> rat) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, i < t - 1 -> forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((selectedpolicysuccessorhistorypolicy hk sigma2 t delta defaultaction i).action (selectedpolicysuccessorhistorys…","shard":"modules/defbb2ae1e1d9c48.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_actionRewardTrajMeasure_centeredKernel","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_actionRewardTrajMeasure_centeredKernel","description":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) ->…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html#decl-40fffbfcab9e","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","order":3012,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean:613"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_actionrewardtrajmeasure_centeredkernel theorem lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_explicitthresholdsum_actionrewardtrajmeasure_centeredkernel {context : type u} {k : nat} [measurablespace context] (model : finitebanditmodel k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction : fin k) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, i < t - 1 -> forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((selectedpolicysuccessorhistorypolicy model.hk sigma2 t delta defaultaction i).action (selectedpolicysuccessorhistorystate model.hk sigma2 t delta defaultaction i history)) <= sigma2) (harm…","shard":"modules/defbb2ae1e1d9c48.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","description":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectory/index.html#decl-8ed321d3e623","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","order":3013,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectory.lean:778"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i histo…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel theorem lintegral_ofreal_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel {context : type u} {k : nat} [measurablespace context] (model : finitebanditmodel k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction : fin k) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, i < t - 1 -> forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((selectedpolicysuccessorhistorypolicy model.hk sigma2 t delta defaultaction i).action (selectedpolicysuccessorhistorystate model.hk sigma2 t delta defaultaction i history)) <= sigma2) (harmmean : forall i :…","shard":"modules/defbb2ae1e1d9c48.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","description":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> C…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectoryreal/index.html#decl-60e1ea0c1325","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","order":3014,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectoryReal.lean:18"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history)…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel theorem integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel {context : type u} {k : nat} [measurablespace context] (model : finitebanditmodel k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction : fin k) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, i < t - 1 -> forall history : ((j : finset.iic i) -> rat), varianceproxy (context i history) ((selectedpolicysuccessorhistorypolicy model.hk sigma2 t delta defaultaction i).action (selectedpolicysuccessorhistorystate model.hk sigma2 t delta defaultaction i history)) <= sigma2) (harmmean : forall i : nat, fora…","shard":"modules/e1e962c51532b4c6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticDelta","label":"selectedPolicySuccessorAsymptoticDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorAsymptoticDelta","description":"Confidence schedule used by the fixed-model asymptotic UCB family.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-6dd8e5e0af88","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3015,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:22"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorAsymptoticDelta (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorasymptoticdelta banditrlproof.ucb.selectedpolicysuccessorasymptoticdelta confidence schedule used by the fixed-model asymptotic ucb family. definition compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticDelta_pos","label":"selectedPolicySuccessorAsymptoticDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorAsymptoticDelta_pos","description":"theorem selectedPolicySuccessorAsymptoticDelta_pos (T : Nat) : 0 < selectedPolicySuccessorAsymptoticDelta T","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-ebe66b5b40f0","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3016,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:25"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorAsymptoticDelta_pos (T : Nat) : 0 < selectedPolicySuccessorAsymptoticDelta T","missing":[],"search":"selectedpolicysuccessorasymptoticdelta_pos banditrlproof.ucb.selectedpolicysuccessorasymptoticdelta_pos theorem selectedpolicysuccessorasymptoticdelta_pos (t : nat) : 0 < selectedpolicysuccessorasymptoticdelta t theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.horizon_mul_selectedPolicySuccessorAsymptoticDelta_le_one","label":"horizon_mul_selectedPolicySuccessorAsymptoticDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.horizon_mul_selectedPolicySuccessorAsymptoticDelta_le_one","description":"theorem horizon_mul_selectedPolicySuccessorAsymptoticDelta_le_one (T : Nat) : (T : Real) * selectedPolicySuccessorAsymptoticDelta T <= 1","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-5ba93b65a221","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3017,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:30"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem horizon_mul_selectedPolicySuccessorAsymptoticDelta_le_one (T : Nat) : (T : Real) * selectedPolicySuccessorAsymptoticDelta T <= 1","missing":[],"search":"horizon_mul_selectedpolicysuccessorasymptoticdelta_le_one banditrlproof.ucb.horizon_mul_selectedpolicysuccessorasymptoticdelta_le_one theorem horizon_mul_selectedpolicysuccessorasymptoticdelta_le_one (t : nat) : (t : real) * selectedpolicysuccessorasymptoticdelta t <= 1 theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_eq","label":"selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_eq","description":"Under the asymptotic schedule, the full-horizon peeling argument is a polynomial.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-d8e35f416dec","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3018,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:39"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_eq (K T : Nat) (hK : 0 < K) (hT : 0 < T) : selectedPolicySuccessorFiniteArmTimeLogBudget K T T (selectedPolicySuccessorAsymptoticDelta T) = max (Real.log (2 * (K : Real) * (T : Real) * (T : Real) * (((T + 1 : Nat) : Real)))) 0","missing":[],"search":"selectedpolicysuccessorfinitearmtimelogbudget_asymptoticdelta_eq banditrlproof.ucb.selectedpolicysuccessorfinitearmtimelogbudget_asymptoticdelta_eq under the asymptotic schedule, the full-horizon peeling argument is a polynomial. theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_le","label":"selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_le","description":"Eventually, the scheduled finite-arm/time peeling budget is at most `4 log(T+1)`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-413376bf93b5","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3019,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:55"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_le (K T : Nat) (hK : 0 < K) (hT : 0 < T) (hlarge : 2 * K <= T + 1) : selectedPolicySuccessorFiniteArmTimeLogBudget K T T (selectedPolicySuccessorAsymptoticDelta T) <= 4 * Real.log (((T + 1 : Nat) : Real))","missing":[],"search":"selectedpolicysuccessorfinitearmtimelogbudget_asymptoticdelta_le banditrlproof.ucb.selectedpolicysuccessorfinitearmtimelogbudget_asymptoticdelta_le eventually, the scheduled finite-arm/time peeling budget is at most `4 log(t+1)`. theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticGapCoefficient","label":"selectedPolicySuccessorAsymptoticGapCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorAsymptoticGapCoefficient","description":"Fixed coefficient that absorbs one positive-gap arm's scheduled textbook budget.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-52f51fe2cfc6","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3020,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:95"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorAsymptoticGapCoefficient (sigma2 : NNReal) (gap : Real) : Real","missing":[],"search":"selectedpolicysuccessorasymptoticgapcoefficient banditrlproof.ucb.selectedpolicysuccessorasymptoticgapcoefficient fixed coefficient that absorbs one positive-gap arm's scheduled textbook budget. definition compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget_add_failure_asymptoticDelta_le","label":"selectedPolicySuccessorTextbookGapBudget_add_failure_asymptoticDelta_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget_add_failure_asymptoticDelta_le","description":"theorem selectedPolicySuccessorTextbookGapBudget_add_failure_asymptoticDelta_le (K : Nat) (sigma2 : NNReal) (T : Nat) (gap : Real) (hK : 0 < K) (hT : 0 < T) (hlarge : 2 * K <= T + 1) (hgap : 0 < gap) : selectedPolicySuccessorTextbookGapBudget K sigma2 T (selectedPolicySuccessorAsymptoticDelta T) gap + gap * ((T : Real) * selectedPolicySuccessorAsymptoticDelta T) <= selectedPolicySuccessorAsymptoticGapCoefficient sig…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-0838cceb6fd9","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3021,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:99"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTextbookGapBudget_add_failure_asymptoticDelta_le (K : Nat) (sigma2 : NNReal) (T : Nat) (gap : Real) (hK : 0 < K) (hT : 0 < T) (hlarge : 2 * K <= T + 1) (hgap : 0 < gap) : selectedPolicySuccessorTextbookGapBudget K sigma2 T (selectedPolicySuccessorAsymptoticDelta T) gap + gap * ((T : Real) * selectedPolicySuccessorAsymptoticDelta T) <= selectedPolicySuccessorAsymptoticGapCoefficient sigma2 gap * (1 + Real.log (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessortextbookgapbudget_add_failure_asymptoticdelta_le banditrlproof.ucb.selectedpolicysuccessortextbookgapbudget_add_failure_asymptoticdelta_le theorem selectedpolicysuccessortextbookgapbudget_add_failure_asymptoticdelta_le (k : nat) (sigma2 : nnreal) (t : nat) (gap : real) (hk : 0 < k) (ht : 0 < t) (hlarge : 2 * k <= t + 1) (hgap : 0 < gap) : selectedpolicysuccessortextbookgapbudget k sigma2 t (selectedpolicysuccessorasymptoticdelta t) gap + gap * ((t : real) * selectedpolicysuccessorasymptoticdelta t) <= selectedpolicysuccessorasymptoticgapcoefficient sigma2 gap * (1 + real.log (((t + 1 : nat) : real))) theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticModelCoefficient","label":"selectedPolicySuccessorAsymptoticModelCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorAsymptoticModelCoefficient","description":"Fixed finite-arm coefficient for the asymptotic textbook UCB envelope.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-006e72392a0f","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3022,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:175"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorAsymptoticModelCoefficient {K : Nat} (model : FiniteBanditModel K) (sigma2 : NNReal) : Real","missing":[],"search":"selectedpolicysuccessorasymptoticmodelcoefficient banditrlproof.ucb.selectedpolicysuccessorasymptoticmodelcoefficient fixed finite-arm coefficient for the asymptotic textbook ucb envelope. definition compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapSum_asymptoticDelta_le","label":"selectedPolicySuccessorTextbookGapSum_asymptoticDelta_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTextbookGapSum_asymptoticDelta_le","description":"theorem selectedPolicySuccessorTextbookGapSum_asymptoticDelta_le {K : Nat} (model : FiniteBanditModel K) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (hlarge : 2 * K <= T + 1) : ((Finset.univ : Finset (Fin K)).filter (fun arm => 0 < (((model.gap arm : Rat) : Real)))).sum (fun arm => selectedPolicySuccessorTextbookGapBudget K sigma2 T (selectedPolicySuccessorAsymptoticDelta T) (((model.gap arm : Rat) : Real)) + (((model.…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-ebb1a3bf57b9","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3023,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:182"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTextbookGapSum_asymptoticDelta_le {K : Nat} (model : FiniteBanditModel K) (sigma2 : NNReal) (T : Nat) (hT : 0 < T) (hlarge : 2 * K <= T + 1) : ((Finset.univ : Finset (Fin K)).filter (fun arm => 0 < (((model.gap arm : Rat) : Real)))).sum (fun arm => selectedPolicySuccessorTextbookGapBudget K sigma2 T (selectedPolicySuccessorAsymptoticDelta T) (((model.gap arm : Rat) : Real)) + (((model.gap arm : Rat) : Real)) * ((T : Real) * selectedPolicySuccessorAsymptoticDelta T)) <= selectedPolicySuccessorAsymptoticModelCoefficient model sigma2 * (1 + Real.log (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessortextbookgapsum_asymptoticdelta_le banditrlproof.ucb.selectedpolicysuccessortextbookgapsum_asymptoticdelta_le theorem selectedpolicysuccessortextbookgapsum_asymptoticdelta_le {k : nat} (model : finitebanditmodel k) (sigma2 : nnreal) (t : nat) (ht : 0 < t) (hlarge : 2 * k <= t + 1) : ((finset.univ : finset (fin k)).filter (fun arm => 0 < (((model.gap arm : rat) : real)))).sum (fun arm => selectedpolicysuccessortextbookgapbudget k sigma2 t (selectedpolicysuccessorasymptoticdelta t) (((model.gap arm : rat) : real)) + (((model.gap arm : rat) : real)) * ((t : real) * selectedpolicysuccessorasymptoticdelta t)) <= selectedpolicysuccessorasymptoticmodelcoefficient model sigma2 * (1 + real.log (((t + 1 : nat) : real))) theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret","label":"selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret","description":"Exact sampled-successor expected pseudo-regret for the canonical pair process at confidence budget `1 / (T + 1)`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-2f00a2ebeafc","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3024,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:210"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret banditrlproof.ucb.selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret exact sampled-successor expected pseudo-regret for the canonical pair process at confidence budget `1 / (t + 1)`. definition compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_nonneg_and_le","label":"selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_nonneg_and_le","description":"Pointwise fixed-horizon envelope used by the asymptotic route.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-10995de23df5","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3025,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:251"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_nonneg_and_le {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), forall arm : Fin K, varianceProxy (context i history) arm <= sigma2) (harmMean : forall i : Nat, forall history : ((j : Fi…","missing":[],"search":"selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_nonneg_and_le banditrlproof.ucb.selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_nonneg_and_le pointwise fixed-horizon envelope used by the asymptotic route. theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticModelCoefficient_nonneg","label":"selectedPolicySuccessorAsymptoticModelCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorAsymptoticModelCoefficient_nonneg","description":"theorem selectedPolicySuccessorAsymptoticModelCoefficient_nonneg {K : Nat} (model : FiniteBanditModel K) (sigma2 : NNReal) : 0 <= selectedPolicySuccessorAsymptoticModelCoefficient model sigma2","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-b852ff706448","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3026,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:330"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorAsymptoticModelCoefficient_nonneg {K : Nat} (model : FiniteBanditModel K) (sigma2 : NNReal) : 0 <= selectedPolicySuccessorAsymptoticModelCoefficient model sigma2","missing":[],"search":"selectedpolicysuccessorasymptoticmodelcoefficient_nonneg banditrlproof.ucb.selectedpolicysuccessorasymptoticmodelcoefficient_nonneg theorem selectedpolicysuccessorasymptoticmodelcoefficient_nonneg {k : nat} (model : finitebanditmodel k) (sigma2 : nnreal) : 0 <= selectedpolicysuccessorasymptoticmodelcoefficient model sigma2 theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.one_add_log_natCast_succ_isBigO_log_natCast_succ","label":"one_add_log_natCast_succ_isBigO_log_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.one_add_log_natCast_succ_isBigO_log_natCast_succ","description":"theorem one_add_log_natCast_succ_isBigO_log_natCast_succ : (fun T : Nat => 1 + Real.log (((T + 1 : Nat) : Real))) =O[atTop] (fun T : Nat => Real.log (((T + 1 : Nat) : Real)))","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-40084caca637","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3027,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:342"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem one_add_log_natCast_succ_isBigO_log_natCast_succ : (fun T : Nat => 1 + Real.log (((T + 1 : Nat) : Real))) =O[atTop] (fun T : Nat => Real.log (((T + 1 : Nat) : Real)))","missing":[],"search":"one_add_log_natcast_succ_isbigo_log_natcast_succ banditrlproof.ucb.one_add_log_natcast_succ_isbigo_log_natcast_succ theorem one_add_log_natcast_succ_isbigo_log_natcast_succ : (fun t : nat => 1 + real.log (((t + 1 : nat) : real))) =o[attop] (fun t : nat => real.log (((t + 1 : nat) : real))) theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isBigO_log","label":"selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isBigO_log","description":"For fixed model and reward-law data, the exact canonical sampled-successor expected pseudo-regret is logarithmic for the horizon-indexed confidence schedule `1 / (T + 1)`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-ee637cf035f6","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3028,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:363"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isBigO_log {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), forall arm : Fin K, varianceProxy (context i history) arm <= sigma2) (harmMean : forall i : Nat, forall history : ((j : Finse…","missing":[],"search":"selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_isbigo_log banditrlproof.ucb.selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_isbigo_log for fixed model and reward-law data, the exact canonical sampled-successor expected pseudo-regret is logarithmic for the horizon-indexed confidence schedule `1 / (t + 1)`. theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.log_natCast_succ_isLittleO_natCast_succ","label":"log_natCast_succ_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.log_natCast_succ_isLittleO_natCast_succ","description":"theorem log_natCast_succ_isLittleO_natCast_succ : (fun T : Nat => Real.log (((T + 1 : Nat) : Real))) =o[atTop] (fun T : Nat => (((T + 1 : Nat) : Real)))","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-17747dc8c4fc","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3029,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:425"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem log_natCast_succ_isLittleO_natCast_succ : (fun T : Nat => Real.log (((T + 1 : Nat) : Real))) =o[atTop] (fun T : Nat => (((T + 1 : Nat) : Real)))","missing":[],"search":"log_natcast_succ_islittleo_natcast_succ banditrlproof.ucb.log_natcast_succ_islittleo_natcast_succ theorem log_natcast_succ_islittleo_natcast_succ : (fun t : nat => real.log (((t + 1 : nat) : real))) =o[attop] (fun t : nat => (((t + 1 : nat) : real))) theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isLittleO_natCast_succ","label":"selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isLittleO_natCast_succ","description":"theorem selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isLittleO_natCast_succ {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-247ea1623d85","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3030,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:434"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isLittleO_natCast_succ {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), forall arm : Fin K, varianceProxy (context i history) arm <= sigma2) (harmMean : forall i : Nat, forall history :…","missing":[],"search":"selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_islittleo_natcast_succ banditrlproof.ucb.selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_islittleo_natcast_succ theorem selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret_islittleo_natcast_succ {context : type u} {k : nat} [measurablespace context] (model : finitebanditmodel k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (mean : context -> fin k -> rat) (varianceproxy : context -> fin k -> nnreal) (defaultaction : fin k) (sigma2 : nnreal) (hcontext : forall n : nat, measurable (context n)) (hmean : measurable (fun pair : prod context (fin k) => mean pair.1 pair.2)) (law : rewardkernel.centeredrewardkernellaw rewardkernel mean varianceproxy) (hvariance : forall i : nat, forall history : ((j : finset.iic i) -> rat), forall arm : fin k, varianceproxy (context i history) arm <= sigma2) (harmmean : forall i : nat, forall history : ((j : finset.iic i) -> rat), forall arm : fin k, mean (context i history) arm = model.mean arm) (hsigma2 : 0 < (((sigma2 : nnreal) : real))) : (selectedpolicysuccessoractionrewardtrajmeasureexpectedpseudoregret model mu0 rewardkernel context defaultaction sigma2 hcontext) =o[attop] (fun t : nat => ((…","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret","label":"selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret","description":"Expected sampled-successor pseudo-regret normalized by `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-f51557fb1428","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3031,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:465"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessoractionrewardtrajmeasureexpectedaveragepseudoregret banditrlproof.ucb.selectedpolicysuccessoractionrewardtrajmeasureexpectedaveragepseudoregret expected sampled-successor pseudo-regret normalized by `t + 1`. definition compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret_tendsto_zero","label":"selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret_tendsto_zero","description":"For the fixed-model horizon-indexed canonical UCB family, expected sampled-successor pseudo-regret normalized by `T + 1` tends to zero.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledasymptotics/index.html#decl-48b489b4b80c","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","order":3032,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledAsymptotics.lean:482"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret_tendsto_zero {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), forall arm : Fin K, varianceProxy (context i history) arm <= sigma2) (harmMean : forall i : Nat, forall history : ((…","missing":[],"search":"selectedpolicysuccessoractionrewardtrajmeasureexpectedaveragepseudoregret_tendsto_zero banditrlproof.ucb.selectedpolicysuccessoractionrewardtrajmeasureexpectedaveragepseudoregret_tendsto_zero for the fixed-model horizon-indexed canonical ucb family, expected sampled-successor pseudo-regret normalized by `t + 1` tends to zero. theorem compiled","shard":"modules/5605cd1a2cef9a33.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.actionRewardTrajectorySuccessorAction","label":"actionRewardTrajectorySuccessorAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.actionRewardTrajectorySuccessorAction","description":"Shift sampled pair-trajectory actions at coordinates `1, 2, ...` to times `0, 1, ...`.","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledreal/index.html#decl-b298b23f6335","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","order":3033,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledReal.lean:11"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def actionRewardTrajectorySuccessorAction {Action Reward : Type} (trajectory : Nat -> Prod Action Reward) : ActionTrace Action","missing":[],"search":"actionrewardtrajectorysuccessoraction banditrlproof.ucb.actionrewardtrajectorysuccessoraction shift sampled pair-trajectory actions at coordinates `1, 2, ...` to times `0, 1, ...`. definition compiled","shard":"modules/594aa16a9d71c0c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBRegretAction_ae_eq_actionRewardTrajectorySuccessorAction_trajMeasure","label":"selectedPolicySuccessorGeneratedUCBRegretAction_ae_eq_actionRewardTrajectorySuccessorAction_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBRegretAction_ae_eq_actionRewardTrajectorySuccessorAction_trajMeasure","description":"Shift sampled pair-trajectory actions at coordinates `1, 2, ...` to times `0, 1, ...`. -/ def actionRewardTrajectorySuccessorAction {Action Reward : Type} (trajectory : Nat -> Prod Action Reward) : ActionTrace Action := fun t => (trajectory (t + 1)).1 /- The shifted UCB action reconstructed from reward coordinates agrees almost everywhere with the sampled successor action trace on its canonical pair trajectory measu…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledreal/index.html#decl-886d5d0ad820","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","order":3034,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledReal.lean:21"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorGeneratedUCBRegretAction_ae_eq_actionRewardTrajectorySuccessorAction_trajMeasure {Context : Type u} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (defaultAction : Fin K) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) : let policy := selectedPolicySuccessorHistoryPolicy hK sigma2 T delta defaultAction let state := selectedPolicySuccessorHistoryState hK sigma2 T delta defaultAction let pairContext : (i : Nat) -> ((j : Finset.Iic i) -> Prod (Fin K) Rat) -> Context","missing":[],"search":"selectedpolicysuccessorgenerateducbregretaction_ae_eq_actionrewardtrajectorysuccessoraction_trajmeasure banditrlproof.ucb.selectedpolicysuccessorgenerateducbregretaction_ae_eq_actionrewardtrajectorysuccessoraction_trajmeasure shift sampled pair-trajectory actions at coordinates `1, 2, ...` to times `0, 1, ...`. -/ def actionrewardtrajectorysuccessoraction {action reward : type} (trajectory : nat -> prod action reward) : actiontrace action := fun t => (trajectory (t + 1)).1 /- the shifted ucb action reconstructed from reward coordinates agrees almost everywhere with the sampled successor action trace on its canonical pair trajectory measure. theorem compiled","shard":"modules/594aa16a9d71c0c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_actionRewardTrajectorySuccessorAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","label":"integral_real_pseudoRegret_actionRewardTrajectorySuccessorAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_actionRewardTrajectorySuccessorAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","description":"Shift sampled pair-trajectory actions at coordinates `1, 2, ...` to times `0, 1, ...`. -/ def actionRewardTrajectorySuccessorAction {Action Reward : Type} (trajectory : Nat -> Prod Action Reward) : ActionTrace Action := fun t => (trajectory (t + 1)).1 /- The shifted UCB action reconstructed from reward coordinates agrees almost everywhere with the sampled successor action trace on its canonical pair trajectory measu…","url":"../modules/banditrlproof-algorithms-ucbconditionalrewardpairtrajectorysampledreal/index.html#decl-7e6d70cd69d8","parent":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","order":3035,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal"],["Source","BanditRLProof/Algorithms/UCBConditionalRewardPairTrajectorySampledReal.lean:136"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_actionRewardTrajectorySuccessorAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel {Context : Type u} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (T : Nat) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, i < T - 1 -> forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((selecte…","missing":[],"search":"integral_real_pseudoregret_actionrewardtrajectorysuccessoraction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel banditrlproof.ucb.integral_real_pseudoregret_actionrewardtrajectorysuccessoraction_le_textbookgapsum_actionrewardtrajmeasure_centeredkernel shift sampled pair-trajectory actions at coordinates `1, 2, ...` to times `0, 1, ...`. -/ def actionrewardtrajectorysuccessoraction {action reward : type} (trajectory : nat -> prod action reward) : actiontrace action := fun t => (trajectory (t + 1)).1 /- the shifted ucb action reconstructed from reward coordinates agrees almost everywhere with the sampled successor action trace on its canonical pair trajectory measure. -/ theorem selectedpolicysuccessorgenerateducbregretaction_ae_eq_actionrewardtrajectorysuccessoraction_trajmeasure {context : type u} {k : nat} [measurablespace context] (hk : 0 < k) (mu0 : measure (prod (fin k) rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context (fin k)) rat) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (defaultaction : fin k) (sigma2 : nnreal) (t : nat) (delta : real) (hcontext : forall n : nat, measurable (context n)) : let policy := selectedpolicysuccessorhistorypolicy hk sigma2 t delta defaultaction let state := selectedpolicysuccessorhistorystate hk sigma2 t delta defaultaction let paircontext : (i : nat) -> ((j : finset.…","shard":"modules/594aa16a9d71c0c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentBoundedRewardKernel","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentBoundedRewardKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentBoundedRewardKernel","description":"Canonical Real expected pseudo-regret bound for a context-dependent Markov reward kernel with stationary arm means and common bounded support. The centered kernel law, constant Hoeffding proxy, selected reward law, trajectory law, and finite-horizon integrability are constructed internally.","url":"../modules/banditrlproof-algorithms-ucbcontextdependentboundedrewardkernel/index.html#decl-c000c82a6e82","parent":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","order":3036,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel"],["Source","BanditRLProof/Algorithms/UCBContextDependentBoundedRewardKernel.lean:25"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentBoundedRewardKernel {Context : Type} [MeasurableSpace Context] {K : Nat} (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n, Measurable (context n)) (lo hi : Real) (hlohi : lo < hi) (hmeas : forall ctx arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (RewardKernel.selectedMeasure rewardKernel ctx arm)) (hbound : forall ctx arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (RewardKernel.selectedMeasure rewardKernel ctx arm))) (hmean : forall ctx arm, integral (RewardKernel.selectedMeasure rewardKernel ctx arm) (fun reward : R…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_contextdependentboundedrewardkernel banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_contextdependentboundedrewardkernel canonical real expected pseudo-regret bound for a context-dependent markov reward kernel with stationary arm means and common bounded support. the centered kernel law, constant hoeffding proxy, selected reward law, trajectory law, and finite-horizon integrability are constructed internally. theorem compiled","shard":"modules/45ce4fd8138cd9f4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentSubgaussianRewardKernel","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentSubgaussianRewardKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentSubgaussianRewardKernel","description":"Canonical Real expected pseudo-regret bound for a context-dependent Markov reward kernel with stationary arm means and direct centered sub-Gaussian MGF witnesses. The centered kernel law, selected reward law, trajectory law, and finite-horizon integrability are constructed internally. The caller supplies a positive common ceiling because an arbitrary measurable context space has no finite maximum operation for the p…","url":"../modules/banditrlproof-algorithms-ucbcontextdependentsubgaussianrewardkernel/index.html#decl-e1f957b50d3d","parent":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","order":3037,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel"],["Source","BanditRLProof/Algorithms/UCBContextDependentSubGaussianRewardKernel.lean:30"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentSubgaussianRewardKernel {Context : Type} [MeasurableSpace Context] {K : Nat} (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n, Measurable (context n)) (varianceProxy : Context -> Fin K -> NNReal) (sigma2 : NNReal) (hsigma2 : 0 < ((sigma2 : NNReal) : Real)) (hvariance : forall ctx arm, varianceProxy ctx arm <= sigma2) (hmean : forall ctx arm, integral (RewardKernel.selectedMeasure rewardKernel ctx arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall ctx arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : R…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_contextdependentsubgaussianrewardkernel banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_contextdependentsubgaussianrewardkernel canonical real expected pseudo-regret bound for a context-dependent markov reward kernel with stationary arm means and direct centered sub-gaussian mgf witnesses. the centered kernel law, selected reward law, trajectory law, and finite-horizon integrability are constructed internally. the caller supplies a positive common ceiling because an arbitrary measurable context space has no finite maximum operation for the pointwise variance proxies. theorem compiled","shard":"modules/584e0b20292298cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel","description":"Canonical Real expected pseudo-regret bound for direct sub-Gaussian selected laws over a finite context space. The common UCB proxy is the finite maximum of all context-action proxies, so callers do not supply a separate ceiling or its pointwise domination proof.","url":"../modules/banditrlproof-algorithms-ucbcontextdependentsubgaussianrewardkernel/index.html#decl-ce3088b1cde0","parent":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","order":3038,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel"],["Source","BanditRLProof/Algorithms/UCBContextDependentSubGaussianRewardKernel.lean:105"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel {Context : Type} [MeasurableSpace Context] [Fintype Context] {K : Nat} (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n, Measurable (context n)) (varianceProxy : Context -> Fin K -> NNReal) (hvariancePos : exists ctx arm, 0 < ((varianceProxy ctx arm : NNReal) : Real)) (hmean : forall ctx arm, integral (RewardKernel.selectedMeasure rewardKernel ctx arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall ctx arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varia…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_finitecontextdependentsubgaussianrewardkernel banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_finitecontextdependentsubgaussianrewardkernel canonical real expected pseudo-regret bound for direct sub-gaussian selected laws over a finite context space. the common ucb proxy is the finite maximum of all context-action proxies, so callers do not supply a separate ceiling or its pointwise domination proof. theorem compiled","shard":"modules/584e0b20292298cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel_without_proxy_positivity","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel_without_proxy_positivity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel_without_proxy_positivity","description":"Canonical Real expected pseudo-regret bound for direct sub-Gaussian selected laws over a finite context space, including the all-zero proxy case. The UCB parameter is the finite context-action maximum padded by one, so no positivity or ceiling premise remains at the public theorem boundary.","url":"../modules/banditrlproof-algorithms-ucbcontextdependentsubgaussianrewardkernel/index.html#decl-a784a74dda07","parent":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","order":3039,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel"],["Source","BanditRLProof/Algorithms/UCBContextDependentSubGaussianRewardKernel.lean:176"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel_without_proxy_positivity {Context : Type} [MeasurableSpace Context] [Fintype Context] {K : Nat} (model : FiniteBanditModel K) (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Fin K) Rat) (context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Context) (hcontext : forall n, Measurable (context n)) (varianceProxy : Context -> Fin K -> NNReal) (hmean : forall ctx arm, integral (RewardKernel.selectedMeasure rewardKernel ctx arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall ctx arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy ctx arm) (RewardKernel.selectedMeasure reward…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_finitecontextdependentsubgaussianrewardkernel_without_proxy_positivity banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_trajmeasure_finitecontextdependentsubgaussianrewardkernel_without_proxy_positivity canonical real expected pseudo-regret bound for direct sub-gaussian selected laws over a finite context space, including the all-zero proxy case. the ucb parameter is the finite context-action maximum padded by one, so no positivity or ceiling premise remains at the public theorem boundary. theorem compiled","shard":"modules/584e0b20292298cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws","description":"Canonical Real expected pseudo-regret bound for stationary finite-arm reward laws with direct centered sub-Gaussian MGF witnesses. The UCB variance parameter is the maximum of the armwise proxies. At least one proxy must be positive because the existing canonical UCB route requires a strictly positive common proxy; zero proxies for other arms are allowed.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussianrewardlaw/index.html#decl-f52c5a74874f","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","order":3040,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianRewardLaw.lean:27"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hvariancePositive : exists arm, 0 < ((varianceProxy arm : NNReal) : Real)) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) (defaultAction : Fin K) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) : let sigma2 := Concentration.finiteArmVarianceProxy varianceProxy let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Unit) armLaw hprob let context : (n : Nat) -> H…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_finitearmsubgaussianlaws banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_finitearmsubgaussianlaws canonical real expected pseudo-regret bound for stationary finite-arm reward laws with direct centered sub-gaussian mgf witnesses. the ucb variance parameter is the maximum of the armwise proxies. at least one proxy must be positive because the existing canonical ucb route requires a strictly positive common proxy; zero proxies for other arms are allowed. theorem compiled","shard":"modules/04cdf84537509fed.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity","label":"integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity","description":"Canonical Real expected pseudo-regret bound for stationary finite-arm reward laws with direct centered sub-Gaussian MGF witnesses, including an all-zero proxy family. The common UCB tuning proxy is the finite maximum padded by one, so callers provide neither a positivity witness nor a separate ceiling.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussianrewardlaw/index.html#decl-a53c41f0d131","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","order":3041,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianRewardLaw.lean:119"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) (defaultAction : Fin K) (T : Nat) (hT : 0 < T) (delta : Real) (hdelta : 0 < delta) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy varianceProxy let rewardKernel := RewardKernel.contextIndependentOfActionLaws (Context := Unit) armLaw hprob let context : (n : Nat) -> History.FiniteRewardHistory Rat n -> Unit :=…","missing":[],"search":"integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_finitearmsubgaussianlaws_without_proxy_positivity banditrlproof.ucb.integral_real_pseudoregret_selectedpolicysuccessorgenerateducbregretaction_le_textbookgapsum_finitearmsubgaussianlaws_without_proxy_positivity canonical real expected pseudo-regret bound for stationary finite-arm reward laws with direct centered sub-gaussian mgf witnesses, including an all-zero proxy family. the common ucb tuning proxy is the finite maximum padded by one, so callers provide neither a positivity witness nor a separate ceiling. theorem compiled","shard":"modules/04cdf84537509fed.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmSubgaussianInitialActionRewardMeasure","label":"finiteArmSubgaussianInitialActionRewardMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmSubgaussianInitialActionRewardMeasure","description":"Initial pair law obtained by attaching a fixed action to its reward law.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-640f52e592cb","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3042,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:22"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmSubgaussianInitialActionRewardMeasure {K : Nat} (armLaw : Fin K -> Measure Rat) (defaultAction : Fin K) : Measure (Prod (Fin K) Rat)","missing":[],"search":"finitearmsubgaussianinitialactionrewardmeasure banditrlproof.ucb.finitearmsubgaussianinitialactionrewardmeasure initial pair law obtained by attaching a fixed action to its reward law. definition compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmSubgaussianInitialActionRewardMeasure_isProbabilityMeasure","label":"finiteArmSubgaussianInitialActionRewardMeasure_isProbabilityMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmSubgaussianInitialActionRewardMeasure_isProbabilityMeasure","description":"The fixed-action reward pushforward remains a probability measure.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-0c6f8e317b35","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3043,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:28"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem finiteArmSubgaussianInitialActionRewardMeasure_isProbabilityMeasure {K : Nat} (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (defaultAction : Fin K) : IsProbabilityMeasure (finiteArmSubgaussianInitialActionRewardMeasure armLaw defaultAction)","missing":[],"search":"finitearmsubgaussianinitialactionrewardmeasure_isprobabilitymeasure banditrlproof.ucb.finitearmsubgaussianinitialactionrewardmeasure_isprobabilitymeasure the fixed-action reward pushforward remains a probability measure. theorem compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret","label":"selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret","description":"Exact sampled-successor expected pseudo-regret for stationary finite-arm sub-Gaussian laws. The UCB proxy is the armwise maximum padded by one and the confidence schedule is `1 / (T + 1)`.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-a3d0ee40424d","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3044,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:43"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (defaultAction : Fin K) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret banditrlproof.ucb.selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret exact sampled-successor expected pseudo-regret for stationary finite-arm sub-gaussian laws. the ucb proxy is the armwise maximum padded by one and the confidence schedule is `1 / (t + 1)`. definition compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le","label":"selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le","description":"The exact practical sampled-successor expected pseudo-regret is nonnegative and satisfies the fixed-model logarithmic envelope at every large horizon.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-76c207dbaf15","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3045,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:69"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) (defaultAction : Fin K) (T : Nat) (hlarge : 2 * K <= T + 1) : 0 <= selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret model armLaw hprob varianceProxy defaultAction T /\\ selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret model armLaw hprob varianceProxy defaultAction T <= selectedPolicySuccessorAsymptoticModelCoefficient model (Concentration.finiteArmPositiveVa…","missing":[],"search":"selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret_nonneg_and_le banditrlproof.ucb.selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret_nonneg_and_le the exact practical sampled-successor expected pseudo-regret is nonnegative and satisfies the fixed-model logarithmic envelope at every large horizon. theorem compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isBigO_log","label":"selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isBigO_log","description":"The exact practical expected pseudo-regret family is logarithmic.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-16578b6c587a","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3046,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:142"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isBigO_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) (defaultAction : Fin K) : (selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret model armLaw hprob varianceProxy defaultAction) =O[atTop] (fun T : Nat => Real.log (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret_isbigo_log banditrlproof.ucb.selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret_isbigo_log the exact practical expected pseudo-regret family is logarithmic. theorem compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isLittleO_natCast_succ","label":"selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isLittleO_natCast_succ","description":"The exact practical expected pseudo-regret is little-o of `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-2891380599d4","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3047,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:204"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isLittleO_natCast_succ {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) (defaultAction : Fin K) : (selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret model armLaw hprob varianceProxy defaultAction) =o[atTop] (fun T : Nat => (((T + 1 : Nat) : Real)))","missing":[],"search":"selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret_islittleo_natcast_succ banditrlproof.ucb.selectedpolicysuccessorfinitearmsubgaussianexpectedpseudoregret_islittleo_natcast_succ the exact practical expected pseudo-regret is little-o of `t + 1`. theorem compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret","label":"selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret","description":"Expected practical sampled-successor pseudo-regret normalized by `T + 1`.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-d3dbb9ed5f3c","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3048,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:226"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (defaultAction : Fin K) (T : Nat) : Real","missing":[],"search":"selectedpolicysuccessorfinitearmsubgaussianexpectedaveragepseudoregret banditrlproof.ucb.selectedpolicysuccessorfinitearmsubgaussianexpectedaveragepseudoregret expected practical sampled-successor pseudo-regret normalized by `t + 1`. definition compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","label":"selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","description":"For stationary finite-arm sub-Gaussian reward laws, the expected pseudo-regret of the horizon-indexed canonical sampled-pair UCB family, normalized by `T + 1`, tends to zero.","url":"../modules/banditrlproof-algorithms-ucbfinitearmsubgaussiansampledasymptotics/index.html#decl-81ecd453ffb7","parent":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","order":3049,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics"],["Source","BanditRLProof/Algorithms/UCBFiniteArmSubGaussianSampledAsymptotics.lean:244"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (varianceProxy : Fin K -> NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - model.mean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) (defaultAction : Fin K) : Tendsto (selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret model armLaw hprob varianceProxy defaultAction) atTop (nhds 0)","missing":[],"search":"selectedpolicysuccessorfinitearmsubgaussianexpectedaveragepseudoregret_tendsto_zero banditrlproof.ucb.selectedpolicysuccessorfinitearmsubgaussianexpectedaveragepseudoregret_tendsto_zero for stationary finite-arm sub-gaussian reward laws, the expected pseudo-regret of the horizon-indexed canonical sampled-pair ucb family, normalized by `t + 1`, tends to zero. theorem compiled","shard":"modules/c6553666b0c493f1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.ArmRewardStream","label":"ArmRewardStream","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.UCB.ArmRewardStream","description":"A table containing one infinite reward stream for every finite arm.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-6d3f3c0dc783","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3050,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:24"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"abbrev ArmRewardStream (K : Nat)","missing":[],"search":"armrewardstream banditrlproof.ucb.armrewardstream a table containing one infinite reward stream for every finite arm. abbreviation compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.armPrefixSum","label":"armPrefixSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.armPrefixSum","description":"Sum of the first `k` rewards in one arm's stream.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-b69310d1583e","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3051,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:27"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def armPrefixSum {K : Nat} (arm : Fin K) (k : Nat) (stream : ArmRewardStream K) : Real","missing":[],"search":"armprefixsum banditrlproof.ucb.armprefixsum sum of the first `k` rewards in one arm's stream. definition compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_armPrefixSum","label":"measurable_armPrefixSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_armPrefixSum","description":"A fixed-arm prefix sum is measurable on the full stream space.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-088758fd2d70","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3052,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:32"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_armPrefixSum {K : Nat} (arm : Fin K) (k : Nat) : Measurable (armPrefixSum arm k)","missing":[],"search":"measurable_armprefixsum banditrlproof.ucb.measurable_armprefixsum a fixed-arm prefix sum is measurable on the full stream space. theorem compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.FixedArmPrefixSource","label":"FixedArmPrefixSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UCB.FixedArmPrefixSource","description":"Pathwise source contract behind fixed-count peeling. For each sample point, arm, and horizon, the selected reward sum must be the prefix sum of that arm's latent stream at the realized pull count.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-2d14b0dfcc87","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3053,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:45"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"structure FixedArmPrefixSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) where","missing":[],"search":"fixedarmprefixsource banditrlproof.ucb.fixedarmprefixsource pathwise source contract behind fixed-count peeling. for each sample point, arm, and horizon, the selected reward sum must be the prefix sum of that arm's latent stream at the realized pull count. structure compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.FixedArmPrefixSource.measurable_armStream","label":"measurable_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.FixedArmPrefixSource.measurable_armStream","description":"The complete latent arm stream supplied by a source is measurable.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-ef2963941c70","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3054,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:57"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem FixedArmPrefixSource.measurable_armStream {Omega : Type u} {K : Nat} [MeasurableSpace Omega] {action : Omega -> ActionTrace (Fin K)} {reward : Omega -> RewardTrace Real} (source : FixedArmPrefixSource action reward) : Measurable source.armStream","missing":[],"search":"measurable_armstream banditrlproof.ucb.fixedarmprefixsource.measurable_armstream the complete latent arm stream supplied by a source is measurable. theorem compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.FixedArmPrefixSource.measurable_armPrefixSum","label":"measurable_armPrefixSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.FixedArmPrefixSource.measurable_armPrefixSum","description":"Every fixed prefix sum read from a source is measurable.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-2d52ee66280e","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3055,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:68"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem FixedArmPrefixSource.measurable_armPrefixSum {Omega : Type u} {K : Nat} [MeasurableSpace Omega] {action : Omega -> ActionTrace (Fin K)} {reward : Omega -> RewardTrace Real} (source : FixedArmPrefixSource action reward) (arm : Fin K) (k : Nat) : Measurable (fun omega => UCB.armPrefixSum arm k (source.armStream omega))","missing":[],"search":"measurable_armprefixsum banditrlproof.ucb.fixedarmprefixsource.measurable_armprefixsum every fixed prefix sum read from a source is measurable. theorem compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource","label":"measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource","description":"Pathwise fixed-count peeling. The adaptive `(pullCount, sumRewards)` event is covered by the finite union of fixed-prefix events for counts at most `n`. This is an outer-measure bound, so the event set itself need not be measurable.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-b18cebae9f46","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3056,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:84"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (source : FixedArmPrefixSource action reward) (arm : Fin K) (n : Nat) (s : Set (Nat × Real)) [DecidablePred (fun k : Nat => k ∈ Prod.fst '' s)] : mu {omega | (pullCount (action omega) arm n, sumRewards (action omega) (reward omega) arm n) ∈ s} ≤ ((Finset.range (n + 1)).filter (fun k => k ∈ Prod.fst '' s)).sum (fun k => mu {omega | UCB.armPrefixSum arm k (source.armStream omega) ∈ Prod.mk k ⁻¹' s})","missing":[],"search":"measure_pullcount_prod_sumrewards_mem_le_of_fixedarmprefixsource banditrlproof.ucb.measure_pullcount_prod_sumrewards_mem_le_of_fixedarmprefixsource pathwise fixed-count peeling. the adaptive `(pullcount, sumrewards)` event is covered by the finite union of fixed-prefix events for counts at most `n`. this is an outer-measure bound, so the event set itself need not be measurable. theorem compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource_identDistrib","label":"measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource_identDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource_identDistrib","description":"Fixed-count peeling with law transport to a canonical arm-reward stream. One `IdentDistrib` hypothesis for the complete latent stream supplies every fixed-prefix law by measurable composition. This is the local counterpart of the law-transport step in LML `prob_pullCount_prod_sumRewards_mem_le`.","url":"../modules/banditrlproof-algorithms-ucbfixedcountpeeling/index.html#decl-7a0c1c8c5865","parent":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","order":3057,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedCountPeeling"],["Source","BanditRLProof/Algorithms/UCBFixedCountPeeling.lean:147"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource_identDistrib {Omega : Type u} {Xi : Type v} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace Xi] (mu : Measure Omega) (nu : Measure Xi) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (source : FixedArmPrefixSource action reward) (canonicalStream : Xi -> ArmRewardStream K) (hstreamLaw : IdentDistrib source.armStream canonicalStream mu nu) (arm : Fin K) (n : Nat) (s : Set (Nat × Real)) [DecidablePred (fun k : Nat => k ∈ Prod.fst '' s)] (hs : MeasurableSet s) : mu {omega | (pullCount (action omega) arm n, sumRewards (action omega) (reward omega) arm n) ∈ s} ≤ ((Finset.range (n + 1)).filter (fun k => k ∈ Prod.fst '' s)).sum (fun k => nu {xi | UCB.armPrefixSum arm k (canonicalStream xi) ∈ Prod.mk k ⁻¹' s})","missing":[],"search":"measure_pullcount_prod_sumrewards_mem_le_of_fixedarmprefixsource_identdistrib banditrlproof.ucb.measure_pullcount_prod_sumrewards_mem_le_of_fixedarmprefixsource_identdistrib fixed-count peeling with law transport to a canonical arm-reward stream. one `identdistrib` hypothesis for the complete latent stream supplies every fixed-prefix law by measurable composition. this is the local counterpart of the law-transport step in lml `prob_pullcount_prod_sumrewards_mem_le`. theorem compiled","shard":"modules/e5071123e12c7774.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingRadiusAt","label":"selectedPolicySuccessorTelescopingRadiusAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingRadiusAt","description":"Realized-count radius with the time-`t` telescoping share divided over the finite arm set.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-7a8ac8836695","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3058,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:22"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingRadiusAt {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (sigma2 : NNReal) (delta : Real) (omega : Omega) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"selectedpolicysuccessortelescopingradiusat banditrlproof.ucb.selectedpolicysuccessortelescopingradiusat realized-count radius with the time-`t` telescoping share divided over the finite arm set. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingRadiusAt_nonneg","label":"selectedPolicySuccessorTelescopingRadiusAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingRadiusAt_nonneg","description":"The scheduled radius is nonnegative, including its explicit zero-count convention inherited from division in `Real`.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-d847fd5581f2","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3059,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:35"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingRadiusAt_nonneg {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (sigma2 : NNReal) (delta : Real) (omega : Omega) (t : Nat) (arm : Fin K) : 0 <= selectedPolicySuccessorTelescopingRadiusAt action sigma2 delta omega t arm","missing":[],"search":"selectedpolicysuccessortelescopingradiusat_nonneg banditrlproof.ucb.selectedpolicysuccessortelescopingradiusat_nonneg the scheduled radius is nonnegative, including its explicit zero-count convention inherited from division in `real`. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingIndexAt","label":"selectedPolicySuccessorTelescopingIndexAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingIndexAt","description":"The horizon-free scheduled UCB index on an action/reward trace.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-79d23ce8e8e4","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3060,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:49"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingIndexAt {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta : Real) (omega : Omega) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"selectedpolicysuccessortelescopingindexat banditrlproof.ucb.selectedpolicysuccessortelescopingindexat the horizon-free scheduled ucb index on an action/reward trace. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex","label":"selectedPolicySuccessorTelescopingHistoryIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex","description":"Scheduled score reconstructed from one finite generated pair history.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-1f399a547028","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3061,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:64"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingHistoryIndex {K : Nat} (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (arm : Fin K) : Real","missing":[],"search":"selectedpolicysuccessortelescopinghistoryindex banditrlproof.ucb.selectedpolicysuccessortelescopinghistoryindex scheduled score reconstructed from one finite generated pair history. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingHistoryIndex","label":"measurable_selectedPolicySuccessorTelescopingHistoryIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingHistoryIndex","description":"Every fixed-arm scheduled history score is measurable.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-17ed82ef86e1","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3062,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:87"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorTelescopingHistoryIndex {K : Nat} (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (t : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Rat t => selectedPolicySuccessorTelescopingHistoryIndex sigma2 delta defaultAction t history arm)","missing":[],"search":"measurable_selectedpolicysuccessortelescopinghistoryindex banditrlproof.ucb.measurable_selectedpolicysuccessortelescopinghistoryindex every fixed-arm scheduled history score is measurable. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryNextArm","label":"selectedPolicySuccessorTelescopingHistoryNextArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryNextArm","description":"Round-robin initialization followed by the scheduled score argmax.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-2d95adb00db0","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3063,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:96"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingHistoryNextArm {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) : Fin K","missing":[],"search":"selectedpolicysuccessortelescopinghistorynextarm banditrlproof.ucb.selectedpolicysuccessortelescopinghistorynextarm round-robin initialization followed by the scheduled score argmax. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex_le_nextArm_of_K_le","label":"selectedPolicySuccessorTelescopingHistoryIndex_le_nextArm_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex_le_nextArm_of_K_le","description":"After initialization, the scheduled selector maximizes every arm score.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-350e06451f6a","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3064,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:108"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingHistoryIndex_le_nextArm_of_K_le {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (t : Nat) (history : History.FinitePairHistory (Fin K) Rat t) (ht : K <= t) (arm : Fin K) : selectedPolicySuccessorTelescopingHistoryIndex sigma2 delta defaultAction t history arm <= selectedPolicySuccessorTelescopingHistoryIndex sigma2 delta defaultAction t history (selectedPolicySuccessorTelescopingHistoryNextArm hK sigma2 delta defaultAction t history)","missing":[],"search":"selectedpolicysuccessortelescopinghistoryindex_le_nextarm_of_k_le banditrlproof.ucb.selectedpolicysuccessortelescopinghistoryindex_le_nextarm_of_k_le after initialization, the scheduled selector maximizes every arm score. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory","label":"selectedPolicySuccessorTelescopingPairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory","description":"Pair-history reconstruction for the single scheduled policy.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-33535fc45d10","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3065,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:126"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingPairHistory {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) : (n : Nat) -> History.FiniteRewardHistory Rat n -> History.FinitePairHistory (Fin K) Rat n | 0, rewardHistory => fun i => (defaultAction, rewardHistory i) | n + 1, rewardHistory => let previousRewardHistory : History.FiniteRewardHistory Rat n := fun i => rewardHistory ⟨i.1, Finset.mem_Iic.mpr ((Finset.mem_Iic.mp i.2).trans (Nat.le_succ n))⟩ let previousHistory := selectedPolicySuccessorTelescopingPairHistory hK sigma2 delta defaultAction n previousRewardHistory let nextAction := selectedPolicySuccessorTelescopingHistoryNextArm hK sigma2 delta defaultAction n previousHistory History.extendPairHistorySucc previousHistory (nextAction, rewardHistory ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) /-- Fixed state package reconstructed by the scheduled policy. -…","missing":[],"search":"selectedpolicysuccessortelescopingpairhistory banditrlproof.ucb.selectedpolicysuccessortelescopingpairhistory pair-history reconstruction for the single scheduled policy. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryState","label":"selectedPolicySuccessorTelescopingHistoryState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryState","description":"Fixed state package reconstructed by the scheduled policy.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-0ed6f1ceb2c2","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3066,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:149"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingHistoryState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (n : Nat) (rewardHistory : History.FiniteRewardHistory Rat n) : SelectedPolicySuccessorFiniteHistoryState K","missing":[],"search":"selectedpolicysuccessortelescopinghistorystate banditrlproof.ucb.selectedpolicysuccessortelescopinghistorystate fixed state package reconstructed by the scheduled policy. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingHistoryState","label":"measurable_selectedPolicySuccessorTelescopingHistoryState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingHistoryState","description":"Scheduled state reconstruction is measurable at every history index.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-0657fc249af9","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3067,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:159"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorTelescopingHistoryState {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (n : Nat) : Measurable (selectedPolicySuccessorTelescopingHistoryState hK sigma2 delta defaultAction n)","missing":[],"search":"measurable_selectedpolicysuccessortelescopinghistorystate banditrlproof.ucb.measurable_selectedpolicysuccessortelescopinghistorystate scheduled state reconstruction is measurable at every history index. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryPolicy","label":"selectedPolicySuccessorTelescopingHistoryPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryPolicy","description":"One measurable UCB policy. Its declaration has no terminal horizon.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-23dfee574caf","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3068,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:168"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingHistoryPolicy {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (_t : Nat) : Policy.MeasurablePolicy (SelectedPolicySuccessorFiniteHistoryState K) (Fin K) where","missing":[],"search":"selectedpolicysuccessortelescopinghistorypolicy banditrlproof.ucb.selectedpolicysuccessortelescopinghistorypolicy one measurable ucb policy. its declaration has no terminal horizon. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction","label":"selectedPolicySuccessorTelescopingGeneratedUCBAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction","description":"Generated action trace for the single scheduled UCB policy.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-004769e0a23f","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3069,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:180"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingGeneratedUCBAction {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace (Fin K)","missing":[],"search":"selectedpolicysuccessortelescopinggenerateducbaction banditrlproof.ucb.selectedpolicysuccessortelescopinggenerateducbaction generated action trace for the single scheduled ucb policy. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction_succ","label":"selectedPolicySuccessorTelescopingGeneratedUCBAction_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction_succ","description":"theorem selectedPolicySuccessorTelescopingGeneratedUCBAction_succ {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) : selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega (t + 1) = selectedPolicySuccessorTelescopingHistoryNextArm hK sigma2 delta defaultAction t (select…","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-63db48aa341d","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3070,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:192"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingGeneratedUCBAction_succ {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) : selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega (t + 1) = selectedPolicySuccessorTelescopingHistoryNextArm hK sigma2 delta defaultAction t (selectedPolicySuccessorTelescopingPairHistory hK sigma2 delta defaultAction t (History.finiteRewardHistoryOfTrace (reward omega) t))","missing":[],"search":"selectedpolicysuccessortelescopinggenerateducbaction_succ banditrlproof.ucb.selectedpolicysuccessortelescopinggenerateducbaction_succ theorem selectedpolicysuccessortelescopinggenerateducbaction_succ {omega : type} {k : nat} (hk : 0 < k) (sigma2 : nnreal) (delta : real) (defaultaction : fin k) (reward : omega -> rewardtrace rat) (omega : omega) (t : nat) : selectedpolicysuccessortelescopinggenerateducbaction hk sigma2 delta defaultaction reward omega (t + 1) = selectedpolicysuccessortelescopinghistorynextarm hk sigma2 delta defaultaction t (selectedpolicysuccessortelescopingpairhistory hk sigma2 delta defaultaction t (history.finiterewardhistoryoftrace (reward omega) t)) theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory_eq_finitePairHistoryOfTrace","label":"selectedPolicySuccessorTelescopingPairHistory_eq_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory_eq_finitePairHistoryOfTrace","description":"The reconstructed scheduled pair state is the generated trace prefix.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-25dc2de2fafd","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3071,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:210"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingPairHistory_eq_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (n : Nat) : selectedPolicySuccessorTelescopingPairHistory hK sigma2 delta defaultAction n (History.finiteRewardHistoryOfTrace (reward omega) n) = History.finitePairHistoryOfTrace (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) (reward omega) n","missing":[],"search":"selectedpolicysuccessortelescopingpairhistory_eq_finitepairhistoryoftrace banditrlproof.ucb.selectedpolicysuccessortelescopingpairhistory_eq_finitepairhistoryoftrace the reconstructed scheduled pair state is the generated trace prefix. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex_finitePairHistoryOfTrace","label":"selectedPolicySuccessorTelescopingHistoryIndex_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex_finitePairHistoryOfTrace","description":"Finite-history score, empirical count, and empirical mean are exactly the score, count, and mean on the generated scheduled trajectory.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-8d48d24d8557","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3072,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:261"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingHistoryIndex_finitePairHistoryOfTrace {Omega : Type} {K : Nat} (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Rat) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (omega : Omega) (t : Nat) (arm : Fin K) : selectedPolicySuccessorTelescopingHistoryIndex sigma2 delta defaultAction t (History.finitePairHistoryOfTrace (action omega) (reward omega) t) arm = selectedPolicySuccessorTelescopingIndexAt action reward sigma2 delta omega t arm","missing":[],"search":"selectedpolicysuccessortelescopinghistoryindex_finitepairhistoryoftrace banditrlproof.ucb.selectedpolicysuccessortelescopinghistoryindex_finitepairhistoryoftrace finite-history score, empirical count, and empirical mean are exactly the score, count, and mean on the generated scheduled trajectory. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingIndexAt_le_generatedAction_of_K_le","label":"selectedPolicySuccessorTelescopingIndexAt_le_generatedAction_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingIndexAt_le_generatedAction_of_K_le","description":"At every post-initialization generated round, the chosen scheduled score dominates every candidate arm score on the same generated trace.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-34f5c2445bea","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3073,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:283"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingIndexAt_le_generatedAction_of_K_le {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (ht : K <= t) (arm : Fin K) : let action := selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward selectedPolicySuccessorTelescopingIndexAt action reward sigma2 delta omega t arm <= selectedPolicySuccessorTelescopingIndexAt action reward sigma2 delta omega t (action omega (t + 1))","missing":[],"search":"selectedpolicysuccessortelescopingindexat_le_generatedaction_of_k_le banditrlproof.ucb.selectedpolicysuccessortelescopingindexat_le_generatedaction_of_k_le at every post-initialization generated round, the chosen scheduled score dominates every candidate arm score on the same generated trace. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction_succ_eq_initializationArm_of_lt","label":"selectedPolicySuccessorTelescopingGeneratedUCBAction_succ_eq_initializationArm_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction_succ_eq_initializationArm_of_lt","description":"Scheduled successor initialization remains the one-pass round robin.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-d519d332c601","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3074,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:330"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingGeneratedUCBAction_succ_eq_initializationArm_of_lt {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (t : Nat) (ht : t < K) : selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega (t + 1) = initializationArm hK t","missing":[],"search":"selectedpolicysuccessortelescopinggenerateducbaction_succ_eq_initializationarm_of_lt banditrlproof.ucb.selectedpolicysuccessortelescopinggenerateducbaction_succ_eq_initializationarm_of_lt scheduled successor initialization remains the one-pass round robin. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_K_add_one_eq_one","label":"successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_K_add_one_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_K_add_one_eq_one","description":"Every arm is pulled exactly once in the scheduled initialization cycle.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-81d527344102","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3075,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:341"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_K_add_one_eq_one {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) : ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) arm (K + 1) = 1","missing":[],"search":"successorarmpullcount_selectedpolicysuccessortelescopinggenerateducbaction_k_add_one_eq_one banditrlproof.ucb.successorarmpullcount_selectedpolicysuccessortelescopinggenerateducbaction_k_add_one_eq_one every arm is pulled exactly once in the scheduled initialization cycle. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_pos_of_K_le","label":"successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_pos_of_K_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_pos_of_K_le","description":"After initialization every scheduled successor arm count is positive.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-69d796a1dc79","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3076,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:369"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_pos_of_K_le {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (ht : K <= t) : 0 < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) arm (t + 1)","missing":[],"search":"successorarmpullcount_selectedpolicysuccessortelescopinggenerateducbaction_pos_of_k_le banditrlproof.ucb.successorarmpullcount_selectedpolicysuccessortelescopinggenerateducbaction_pos_of_k_le after initialization every scheduled successor arm count is positive. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.K_le_of_selectedPolicySuccessorTelescopingGeneratedUCBAction_selected_and_count_pos","label":"K_le_of_selectedPolicySuccessorTelescopingGeneratedUCBAction_selected_and_count_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.K_le_of_selectedPolicySuccessorTelescopingGeneratedUCBAction_selected_and_count_pos","description":"A selected scheduled round with positive prior count is post-initialization.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-f061f6761d42","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3077,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:392"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem K_le_of_selectedPolicySuccessorTelescopingGeneratedUCBAction_selected_and_count_pos {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (t : Nat) (hselected : selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega (t + 1) = arm) (hcount : 0 < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) arm (t + 1)) : K <= t","missing":[],"search":"k_le_of_selectedpolicysuccessortelescopinggenerateducbaction_selected_and_count_pos banditrlproof.ucb.k_le_of_selectedpolicysuccessortelescopinggenerateducbaction_selected_and_count_pos a selected scheduled round with positive prior count is post-initialization. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingGeneratedUCBAction","label":"measurable_selectedPolicySuccessorTelescopingGeneratedUCBAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingGeneratedUCBAction","description":"Timewise measurability of the fixed-policy generated action.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-9b03bcedbc97","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3078,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:426"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorTelescopingGeneratedUCBAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega t)","missing":[],"search":"measurable_selectedpolicysuccessortelescopinggenerateducbaction banditrlproof.ucb.measurable_selectedpolicysuccessortelescopinggenerateducbaction timewise measurability of the fixed-policy generated action. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget","label":"selectedPolicySuccessorTelescopingLogBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget","description":"The logarithmic budget inside the scheduled radius at history index `n`.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-55c93b8bf1b1","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3079,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:448"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingLogBudget (K n : Nat) (delta : Real) : Real","missing":[],"search":"selectedpolicysuccessortelescopinglogbudget banditrlproof.ucb.selectedpolicysuccessortelescopinglogbudget the logarithmic budget inside the scheduled radius at history index `n`. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget_eq","label":"selectedPolicySuccessorTelescopingLogBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget_eq","description":"Explicit polynomial reciprocal form of the scheduled log budget.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-5c0c01bf831a","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3080,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:454"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingLogBudget_eq (K n : Nat) (delta : Real) (hK : 0 < K) (hdelta : 0 < delta) : selectedPolicySuccessorTelescopingLogBudget K n delta = max (Real.log (2 * (K : Real) * ((n + 1 : Nat) : Real) ^ 2 * ((n + 2 : Nat) : Real) / delta)) 0","missing":[],"search":"selectedpolicysuccessortelescopinglogbudget_eq banditrlproof.ucb.selectedpolicysuccessortelescopinglogbudget_eq explicit polynomial reciprocal form of the scheduled log budget. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget_mono","label":"selectedPolicySuccessorTelescopingLogBudget_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget_mono","description":"The scheduled log budget is nondecreasing in the history index.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-efd4d138bf4f","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3081,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:468"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingLogBudget_mono (K n T : Nat) (delta : Real) (hK : 0 < K) (hdelta : 0 < delta) (hnT : n <= T) : selectedPolicySuccessorTelescopingLogBudget K n delta <= selectedPolicySuccessorTelescopingLogBudget K T delta","missing":[],"search":"selectedpolicysuccessortelescopinglogbudget_mono banditrlproof.ucb.selectedpolicysuccessortelescopinglogbudget_mono the scheduled log budget is nondecreasing in the history index. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanTelescopingPeelingRadius_eq","label":"successorArmEmpiricalMeanTelescopingPeelingRadius_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.successorArmEmpiricalMeanTelescopingPeelingRadius_eq","description":"Expanded algebraic form of one scheduled random-count radius.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-8f5f6a3b7d85","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3082,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:487"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMeanTelescopingPeelingRadius_eq {K : Nat} (sigma2 : NNReal) (k n : Nat) (delta : Real) : ConditionalExpectationReward.successorArmEmpiricalMeanPeelingRadius sigma2 k (n + 1) (Concentration.telescopingConfidenceShare delta n / (K : Real)) = (2 * Real.sqrt ((1 / 2 : Real) * (((sigma2 : NNReal) : Real) * (k : Real)) * selectedPolicySuccessorTelescopingLogBudget K n delta) + selectedPolicySuccessorTelescopingLogBudget K n delta) / (k : Real)","missing":[],"search":"successorarmempiricalmeantelescopingpeelingradius_eq banditrlproof.ucb.successorarmempiricalmeantelescopingpeelingradius_eq expanded algebraic form of one scheduled random-count radius. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingRealPullThreshold","label":"selectedPolicySuccessorTelescopingRealPullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingRealPullThreshold","description":"Real terminal envelope for inverting every scheduled radius up to index `T`.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-91a88037ddea","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3083,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:507"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingRealPullThreshold (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) : Real","missing":[],"search":"selectedpolicysuccessortelescopingrealpullthreshold banditrlproof.ucb.selectedpolicysuccessortelescopingrealpullthreshold real terminal envelope for inverting every scheduled radius up to index `t`. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPullThreshold","label":"selectedPolicySuccessorTelescopingPullThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingPullThreshold","description":"One more than the ceiling supplies the strict count-inversion margin.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-6dd4528b9ce2","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3084,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:515"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingPullThreshold (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) : Nat","missing":[],"search":"selectedpolicysuccessortelescopingpullthreshold banditrlproof.ucb.selectedpolicysuccessortelescopingpullthreshold one more than the ceiling supplies the strict count-inversion margin. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPullThreshold_contracts","label":"selectedPolicySuccessorTelescopingPullThreshold_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingPullThreshold_contracts","description":"The terminal scheduled threshold satisfies the quadratic and linear radius-inversion contracts.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-f93984a85d04","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3085,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:523"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescopingPullThreshold_contracts (K : Nat) (sigma2 : NNReal) (T : Nat) (delta gap : Real) (hgap : 0 < gap) : 0 < selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta gap /\\ 32 * (((sigma2 : NNReal) : Real)) * selectedPolicySuccessorTelescopingLogBudget K T delta < gap ^ 2 * (selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta gap : Real) /\\ 4 * selectedPolicySuccessorTelescopingLogBudget K T delta < gap * (selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta gap : Real)","missing":[],"search":"selectedpolicysuccessortelescopingpullthreshold_contracts banditrlproof.ucb.selectedpolicysuccessortelescopingpullthreshold_contracts the terminal scheduled threshold satisfies the quadratic and linear radius-inversion contracts. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.two_mul_successorArmEmpiricalMeanTelescopingPeelingRadius_lt_gap_of_threshold","label":"two_mul_successorArmEmpiricalMeanTelescopingPeelingRadius_lt_gap_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.two_mul_successorArmEmpiricalMeanTelescopingPeelingRadius_lt_gap_of_threshold","description":"Every count beyond the terminal envelope makes the scheduled radius at every earlier history index strictly smaller than half the positive gap.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-e2cd39aa9d9d","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3086,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:576"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem two_mul_successorArmEmpiricalMeanTelescopingPeelingRadius_lt_gap_of_threshold {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta gap : Real) (hdelta : 0 < delta) (hgap : 0 < gap) (k n T : Nat) (hk : selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta gap <= k) (hnT : n <= T) : 2 * ConditionalExpectationReward.successorArmEmpiricalMeanPeelingRadius sigma2 k (n + 1) (Concentration.telescopingConfidenceShare delta n / (K : Real)) < gap","missing":[],"search":"two_mul_successorarmempiricalmeantelescopingpeelingradius_lt_gap_of_threshold banditrlproof.ucb.two_mul_successorarmempiricalmeantelescopingpeelingradius_lt_gap_of_threshold every count beyond the terminal envelope makes the scheduled radius at every earlier history index strictly smaller than half the positive gap. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_meanGap_le_two_radius_of_not_badEvent","label":"selectedPolicySuccessorTelescoping_meanGap_le_two_radius_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescoping_meanGap_le_two_radius_of_not_badEvent","description":"Outside the one telescoping bad event, scheduled score maximality implies the standard UCB gap bound at every initialized generated round.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-ce0dc90db82b","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3087,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:635"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescoping_meanGap_le_two_radius_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (defaultAction best : Fin K) (omega : Omega) (t : Nat) (ht : K <= t) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward) reward armMean sigma2 delta) : let action := selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward meanGap (fun arm => (armMean arm : Real)) best (action omega (t + 1)) <= 2 * selectedPolicySuccessorTelescopingRadiusAt action sigma2 delta omega t (action omega (t + 1))","missing":[],"search":"selectedpolicysuccessortelescoping_meangap_le_two_radius_of_not_badevent banditrlproof.ucb.selectedpolicysuccessortelescoping_meangap_le_two_radius_of_not_badevent outside the one telescoping bad event, scheduled score maximality implies the standard ucb gap bound at every initialized generated round. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_pullCount_le_of_not_badEvent","label":"selectedPolicySuccessorTelescoping_pullCount_le_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescoping_pullCount_le_of_not_badEvent","description":"The single all-time good event controls one positive-gap arm at every finite horizon; the horizon appears only in the deterministic bound.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-df5f78585a7d","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3088,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:712"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescoping_pullCount_le_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (omega : Omega) (T : Nat) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward) reward armMean sigma2 delta) : ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) chosen (T + 1) <= selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta (meanGap (fun arm => (armMean arm : Real)) best chosen)","missing":[],"search":"selectedpolicysuccessortelescoping_pullcount_le_of_not_badevent banditrlproof.ucb.selectedpolicysuccessortelescoping_pullcount_le_of_not_badevent the single all-time good event controls one positive-gap arm at every finite horizon; the horizon appears only in the deterministic bound. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.actionRewardHistoryStepKernelFamily_selectedPolicySuccessorTelescoping_allTimeConfidence","label":"actionRewardHistoryStepKernelFamily_selectedPolicySuccessorTelescoping_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.actionRewardHistoryStepKernelFamily_selectedPolicySuccessorTelescoping_allTimeConfidence","description":"The accepted telescoping all-time confidence producer instantiated on the single scheduled policy. The sampled pair action and the reward-reconstructed policy action are transported on their explicit almost-everywhere alignment set, so the conclusion is about the action actually consumed by the UCB score and regret definitions.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-253ee28e0ef2","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3089,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:790"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedPolicySuccessorTelescoping_allTimeConfidence {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((selectedPolicySuccessorTelescopingHistoryPolicy hK sigma2…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedpolicysuccessortelescoping_alltimeconfidence banditrlproof.ucb.actionrewardhistorystepkernelfamily_selectedpolicysuccessortelescoping_alltimeconfidence the accepted telescoping all-time confidence producer instantiated on the single scheduled policy. the sampled pair action and the reward-reconstructed policy action are transported on their explicit almost-everywhere alignment set, so the conclusion is about the action actually consumed by the ucb score and regret definitions. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_of_allTimeConfidence","label":"measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_of_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_of_allTimeConfidence","description":"One global all-time confidence event yields the terminal pull-count tail for any requested finite horizon.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-a17d9ea35b64","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3090,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:956"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (T : Nat) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward) reward armMean sigma2 delta) <= ENNReal.ofReal delta) : mu {omega | selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta (meanGap (fun arm => (armMean arm : Real)) best chosen) < ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAc…","missing":[],"search":"measure_selectedpolicysuccessortelescoping_pullcount_gt_threshold_le_of_alltimeconfidence banditrlproof.ucb.measure_selectedpolicysuccessortelescoping_pullcount_gt_threshold_le_of_alltimeconfidence one global all-time confidence event yields the terminal pull-count tail for any requested finite horizon. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_selectedPolicySuccessorTelescoping_pullCount_le_of_allTimeConfidence","label":"lintegral_selectedPolicySuccessorTelescoping_pullCount_le_of_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_selectedPolicySuccessorTelescoping_pullCount_le_of_allTimeConfidence","description":"Expected scheduled pull count at any finite horizon. The explicit `T * delta` term is retained; no unconditional expectation claim is made from the good event alone.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-5fd3bdc3971c","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3091,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:991"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_selectedPolicySuccessorTelescoping_pullCount_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (mu : Measure Omega) [IsProbabilityMeasure mu] (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (hdelta : 0 < delta) (defaultAction best chosen : Fin K) (T : Nat) (hgap : 0 < meanGap (fun arm => (armMean arm : Real)) best chosen) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward) reward armMean sigma2 delta) <= ENNReal.ofReal delta) : ∫⁻ omega, (ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultActi…","missing":[],"search":"lintegral_selectedpolicysuccessortelescoping_pullcount_le_of_alltimeconfidence banditrlproof.ucb.lintegral_selectedpolicysuccessortelescoping_pullcount_le_of_alltimeconfidence expected scheduled pull count at any finite horizon. the explicit `t * delta` term is retained; no unconditional expectation claim is made from the good event alone. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","label":"selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","description":"Standard regret-time shift of the single scheduled generated policy.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-87e717f25ff1","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3092,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1050"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingGeneratedUCBRegretAction {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace (Fin K)","missing":[],"search":"selectedpolicysuccessortelescopinggenerateducbregretaction banditrlproof.ucb.selectedpolicysuccessortelescopinggenerateducbregretaction standard regret-time shift of the single scheduled generated policy. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","label":"measurable_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","description":"Timewise measurability of the shifted fixed-policy regret action.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-c891c49f6261","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3093,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1059"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction {Omega : Type} [MeasurableSpace Omega] {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega => selectedPolicySuccessorTelescopingGeneratedUCBRegretAction hK sigma2 delta defaultAction reward omega t)","missing":[],"search":"measurable_selectedpolicysuccessortelescopinggenerateducbregretaction banditrlproof.ucb.measurable_selectedpolicysuccessortelescopinggenerateducbregretaction timewise measurability of the shifted fixed-policy regret action. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.pullCount_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction_eq","label":"pullCount_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.pullCount_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction_eq","description":"Shifted fixed-policy pull counts are the existing successor counts.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-04d8828ff713","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3094,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1072"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem pullCount_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction_eq {Omega : Type} {K : Nat} (hK : 0 < K) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) (reward : Omega -> RewardTrace Rat) (omega : Omega) (arm : Fin K) (T : Nat) : pullCount (selectedPolicySuccessorTelescopingGeneratedUCBRegretAction hK sigma2 delta defaultAction reward omega) arm T = ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) arm (T + 1)","missing":[],"search":"pullcount_selectedpolicysuccessortelescopinggenerateducbregretaction_eq banditrlproof.ucb.pullcount_selectedpolicysuccessortelescopinggenerateducbregretaction_eq shifted fixed-policy pull counts are the existing successor counts. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_of_allTimeConfidence","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_of_allTimeConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_of_allTimeConfidence","description":"Finite-time expected pseudo-regret for the single scheduled policy from its one all-time confidence event.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-a6b2b232a3ed","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3095,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1091"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_of_allTimeConfidence {Omega : Type} [MeasurableSpace Omega] {K : Nat} (mu : Measure Omega) [IsProbabilityMeasure mu] (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (sigma2 : NNReal) (delta : Real) (hdelta : 0 < delta) (defaultAction : Fin K) (T : Nat) (hconfidence : mu (ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (selectedPolicySuccessorTelescopingGeneratedUCBAction model.hK sigma2 delta defaultAction reward) reward model.mean sigma2 delta) <= ENNReal.ofReal delta) : ∫⁻ omega, ENNReal.ofReal (((pseudoRegret model (selectedPolicySuccessorTelescopingGeneratedUCBRegretAction model.hK sigma2 delta defaultAction reward omega) T : Rat) : Real)) ∂mu <= (Finset.univ : Finset (Fin K)…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessortelescoping_le_of_alltimeconfidence banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessortelescoping_le_of_alltimeconfidence finite-time expected pseudo-regret for the single scheduled policy from its one all-time confidence event. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_allHorizonPullCount_of_not_badEvent","label":"selectedPolicySuccessorTelescoping_allHorizonPullCount_of_not_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescoping_allHorizonPullCount_of_not_badEvent","description":"A single good sample controls every finite horizon and every positive-gap arm simultaneously.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-9b5e2087d0ee","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3096,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1151"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicySuccessorTelescoping_allHorizonPullCount_of_not_badEvent {Omega : Type} {K : Nat} (hK : 0 < K) (reward : Omega -> RewardTrace Rat) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (hdelta : 0 < delta) (defaultAction best : Fin K) (omega : Omega) (hgood : omega ∉ ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward) reward armMean sigma2 delta) : forall T : Nat, forall chosen : Fin K, 0 < meanGap (fun arm => (armMean arm : Real)) best chosen -> ConditionalExpectationReward.successorArmPullCount (selectedPolicySuccessorTelescopingGeneratedUCBAction hK sigma2 delta defaultAction reward omega) chosen (T + 1) <= selectedPolicySuccessorTelescopingPullThreshold K sigma2 T delta (meanGap (fun arm => (armMean arm : Real)) best chosen)","missing":[],"search":"selectedpolicysuccessortelescoping_allhorizonpullcount_of_not_badevent banditrlproof.ucb.selectedpolicysuccessortelescoping_allhorizonpullcount_of_not_badevent a single good sample controls every finite horizon and every positive-gap arm simultaneously. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingActionRewardTrajMeasure","label":"selectedPolicySuccessorTelescopingActionRewardTrajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.selectedPolicySuccessorTelescopingActionRewardTrajMeasure","description":"The canonical action/reward trajectory measure of the single scheduled policy. This declaration itself contains no terminal horizon.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-8ebeb5679438","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3097,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1177"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedPolicySuccessorTelescopingActionRewardTrajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (hcontext : forall n : Nat, Measurable (context n)) (sigma2 : NNReal) (delta : Real) (defaultAction : Fin K) : Measure (Nat -> Prod (Fin K) Rat)","missing":[],"search":"selectedpolicysuccessortelescopingactionrewardtrajmeasure banditrlproof.ucb.selectedpolicysuccessortelescopingactionrewardtrajmeasure the canonical action/reward trajectory measure of the single scheduled policy. this declaration itself contains no terminal horizon. definition compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_allTimeBadEvent_le_trajMeasure","label":"measure_selectedPolicySuccessorTelescoping_allTimeBadEvent_le_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_allTimeBadEvent_le_trajMeasure","description":"Short canonical surface for the same-policy all-time confidence theorem.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-fbe42164523d","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3098,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1213"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorTelescoping_allTimeBadEvent_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((selectedPolicySuccessorTelescopingHistoryPolicy hK sigma2 delta defaultA…","missing":[],"search":"measure_selectedpolicysuccessortelescoping_alltimebadevent_le_trajmeasure banditrlproof.ucb.measure_selectedpolicysuccessortelescoping_alltimebadevent_le_trajmeasure short canonical surface for the same-policy all-time confidence theorem. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_trajMeasure","label":"measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_trajMeasure","description":"Canonical finite-horizon pull-count tail for the single scheduled policy and the same trajectory measure used by the all-time confidence theorem.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-9116314a18d2","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3099,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1257"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (hK : 0 < K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction best chosen : Fin K) (armMean : Fin K -> Rat) (sigma2 : NNReal) (delta : Real) (T : Nat) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((selectedPolicySuccessorTelescopingHistoryPo…","missing":[],"search":"measure_selectedpolicysuccessortelescoping_pullcount_gt_threshold_le_trajmeasure banditrlproof.ucb.measure_selectedpolicysuccessortelescoping_pullcount_gt_threshold_le_trajmeasure canonical finite-horizon pull-count tail for the single scheduled policy and the same trajectory measure used by the all-time confidence theorem. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","label":"lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","description":"Canonical finite-time expected pseudo-regret of the single scheduled policy. The same all-time event supplies every horizon, and the failure term is explicitly retained as `T * delta`.","url":"../modules/banditrlproof-algorithms-ucbfixedpolicytelescopinganytimeregret/index.html#decl-8c753c946182","parent":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","order":3100,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret"],["Source","BanditRLProof/Algorithms/UCBFixedPolicyTelescopingAnytimeRegret.lean:1314"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure {Context : Type} {K : Nat} [MeasurableSpace Context] (model : FiniteBanditModel K) (mu0 : Measure (Prod (Fin K) Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context (Fin K)) Rat) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (mean : Context -> Fin K -> Rat) (varianceProxy : Context -> Fin K -> NNReal) (defaultAction : Fin K) (sigma2 : NNReal) (delta : Real) (T : Nat) (hcontext : forall n : Nat, Measurable (context n)) (hmean : Measurable (fun pair : Prod Context (Fin K) => mean pair.1 pair.2)) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hvariance : forall i : Nat, forall history : ((j : Finset.Iic i) -> Rat), varianceProxy (context i history) ((selectedPolicySuccessorTelescopingHistoryPolicy model.hK sigma2…","missing":[],"search":"lintegral_ofreal_pseudoregret_selectedpolicysuccessortelescoping_le_trajmeasure banditrlproof.ucb.lintegral_ofreal_pseudoregret_selectedpolicysuccessortelescoping_le_trajmeasure canonical finite-time expected pseudo-regret of the single scheduled policy. the same all-time event supplies every horizon, and the failure term is explicitly retained as `t * delta`. theorem compiled","shard":"modules/0fa477c4629fa244.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realEmpiricalMean","label":"realEmpiricalMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realEmpiricalMean","description":"Real empirical mean of one arm before time `n`.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-fa311ca97c60","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3101,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:19"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realEmpiricalMean {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (arm : Fin K) (n : Nat) : Real","missing":[],"search":"realempiricalmean banditrlproof.ucb.realempiricalmean real empirical mean of one arm before time `n`. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realWidth","label":"realWidth","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realWidth","description":"LML-shaped path-dependent UCB width before time `n`.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-abd9331aca0f","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3102,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:25"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realWidth {K : Nat} (action : ActionTrace (Fin K)) (c : Real) (arm : Fin K) (n : Nat) : Real","missing":[],"search":"realwidth banditrlproof.ucb.realwidth lml-shaped path-dependent ucb width before time `n`. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realIndex","label":"realIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realIndex","description":"Real empirical mean plus the realized pull-count confidence width.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-79b6bfaae7ec","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3103,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:33"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realIndex {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (c : Real) (arm : Fin K) (n : Nat) : Real","missing":[],"search":"realindex banditrlproof.ucb.realindex real empirical mean plus the realized pull-count confidence width. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryWidth","label":"realHistoryWidth","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryWidth","description":"Inclusive finite-history version of the LML UCB width.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-045828132833","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3104,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:39"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistoryWidth {K : Nat} (c : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Real","missing":[],"search":"realhistorywidth banditrlproof.ucb.realhistorywidth inclusive finite-history version of the lml ucb width. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryIndex","label":"realHistoryIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryIndex","description":"Inclusive finite-history Real UCB score.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-ea0f8d883521","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3105,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:48"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistoryIndex {K : Nat} (c : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Real","missing":[],"search":"realhistoryindex banditrlproof.ucb.realhistoryindex inclusive finite-history real ucb score. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realIndexAction","label":"realIndexAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realIndexAction","description":"Least-encoded maximizer of the path-dependent Real UCB index.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-7602a43e65df","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3106,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:56"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realIndexAction {K : Nat} (hK : 0 < K) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (c : Real) (n : Nat) : Fin K","missing":[],"search":"realindexaction banditrlproof.ucb.realindexaction least-encoded maximizer of the path-dependent real ucb index. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryIndexAction","label":"realHistoryIndexAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryIndexAction","description":"Least-encoded maximizer on an inclusive finite pair history.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-c7036a0e1547","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3107,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:64"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realHistoryIndexAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"realhistoryindexaction banditrlproof.ucb.realhistoryindexaction least-encoded maximizer on an inclusive finite pair history. definition compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistoryPullCount","label":"measurable_realHistoryPullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistoryPullCount","description":"A fixed-arm pull count is measurable as a function of finite pair history.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-c16dddbd3cb2","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3108,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:71"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistoryPullCount {K : Nat} (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => ETC.realHistoryPullCount n history arm)","missing":[],"search":"measurable_realhistorypullcount banditrlproof.ucb.measurable_realhistorypullcount a fixed-arm pull count is measurable as a function of finite pair history. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistorySumRewards","label":"measurable_realHistorySumRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistorySumRewards","description":"A fixed-arm reward sum is measurable as a function of finite pair history.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-35584263f9c2","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3109,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:85"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistorySumRewards {K : Nat} (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => ETC.realHistorySumRewards n history arm)","missing":[],"search":"measurable_realhistorysumrewards banditrlproof.ucb.measurable_realhistorysumrewards a fixed-arm reward sum is measurable as a function of finite pair history. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistoryEmpMean","label":"measurable_realHistoryEmpMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistoryEmpMean","description":"A fixed-arm empirical mean is measurable on finite pair histories.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-27f2c3607e29","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3110,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:99"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistoryEmpMean {K : Nat} (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => ETC.realHistoryEmpMean n history arm)","missing":[],"search":"measurable_realhistoryempmean banditrlproof.ucb.measurable_realhistoryempmean a fixed-arm empirical mean is measurable on finite pair histories. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistoryWidth","label":"measurable_realHistoryWidth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistoryWidth","description":"The fixed-arm UCB width is measurable on finite pair histories.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-d8999cd3e774","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3111,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:109"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistoryWidth {K : Nat} (c : Real) (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => realHistoryWidth c n history arm)","missing":[],"search":"measurable_realhistorywidth banditrlproof.ucb.measurable_realhistorywidth the fixed-arm ucb width is measurable on finite pair histories. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistoryIndex","label":"measurable_realHistoryIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistoryIndex","description":"Every fixed-arm UCB score is measurable on finite pair histories.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-630161d6705b","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3112,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:119"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistoryIndex {K : Nat} (c : Real) (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => realHistoryIndex c n history arm)","missing":[],"search":"measurable_realhistoryindex banditrlproof.ucb.measurable_realhistoryindex every fixed-arm ucb score is measurable on finite pair histories. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realHistoryIndexAction","label":"measurable_realHistoryIndexAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realHistoryIndexAction","description":"The least-encoded finite-history UCB selector is measurable.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-b0f52e3c2628","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3113,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:127"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realHistoryIndexAction {K : Nat} (hK : 0 < K) (c : Real) (n : Nat) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => realHistoryIndexAction hK c n history)","missing":[],"search":"measurable_realhistoryindexaction banditrlproof.ucb.measurable_realhistoryindexaction the least-encoded finite-history ucb selector is measurable. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realEmpiricalMean","label":"measurable_realEmpiricalMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realEmpiricalMean","description":"The trace empirical mean is measurable under timewise measurable data.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-69a339176656","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3114,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:138"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realEmpiricalMean {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSingletonClass (Fin K)] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (arm : Fin K) (n : Nat) : Measurable (fun omega => realEmpiricalMean (action omega) (reward omega) arm n)","missing":[],"search":"measurable_realempiricalmean banditrlproof.ucb.measurable_realempiricalmean the trace empirical mean is measurable under timewise measurable data. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realWidth","label":"measurable_realWidth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realWidth","description":"The realized pull-count UCB width is measurable.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-7a0785296cfd","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3115,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:153"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realWidth {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSingletonClass (Fin K)] (action : Omega -> ActionTrace (Fin K)) (haction : forall t, Measurable (fun omega => action omega t)) (c : Real) (arm : Fin K) (n : Nat) : Measurable (fun omega => realWidth (action omega) c arm n)","missing":[],"search":"measurable_realwidth banditrlproof.ucb.measurable_realwidth the realized pull-count ucb width is measurable. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realIndex","label":"measurable_realIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realIndex","description":"Every coordinate of the realized UCB index is measurable.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-a896d404504f","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3116,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:166"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realIndex {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSingletonClass (Fin K)] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (c : Real) (arm : Fin K) (n : Nat) : Measurable (fun omega => realIndex (action omega) (reward omega) c arm n)","missing":[],"search":"measurable_realindex banditrlproof.ucb.measurable_realindex every coordinate of the realized ucb index is measurable. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realIndexAction_spec","label":"realIndexAction_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realIndexAction_spec","description":"The least-encoded realized UCB index action maximizes every arm score.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-b0e0a42afe43","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3117,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:180"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realIndexAction_spec {K : Nat} (hK : 0 < K) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (c : Real) (n : Nat) (arm : Fin K) : realIndex action reward c arm n <= realIndex action reward c (realIndexAction hK action reward c n) n","missing":[],"search":"realindexaction_spec banditrlproof.ucb.realindexaction_spec the least-encoded realized ucb index action maximizes every arm score. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_realIndexAction","label":"measurable_realIndexAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_realIndexAction","description":"The least-encoded realized UCB index action is measurable.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-318550d3cf2e","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3118,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:191"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_realIndexAction {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSingletonClass (Fin K)] (hK : 0 < K) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (haction : forall t, Measurable (fun omega => action omega t)) (hreward : forall t, Measurable (fun omega => reward omega t)) (c : Real) (n : Nat) : Measurable (fun omega => realIndexAction hK (action omega) (reward omega) c n)","missing":[],"search":"measurable_realindexaction banditrlproof.ucb.measurable_realindexaction the least-encoded realized ucb index action is measurable. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryEmpiricalMean_finitePairHistoryOfTrace","label":"realHistoryEmpiricalMean_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryEmpiricalMean_finitePairHistoryOfTrace","description":"Inclusive finite-history empirical means agree with the trace prefix.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-262920ad592c","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3119,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:207"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realHistoryEmpiricalMean_finitePairHistoryOfTrace {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n : Nat) (arm : Fin K) : ETC.realHistoryEmpMean n (History.finitePairHistoryOfTrace action reward n) arm = realEmpiricalMean action reward arm (n + 1)","missing":[],"search":"realhistoryempiricalmean_finitepairhistoryoftrace banditrlproof.ucb.realhistoryempiricalmean_finitepairhistoryoftrace inclusive finite-history empirical means agree with the trace prefix. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryWidth_finitePairHistoryOfTrace","label":"realHistoryWidth_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryWidth_finitePairHistoryOfTrace","description":"Inclusive finite-history widths agree with the trace width at `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-6e05e6ff58b5","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3120,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:216"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realHistoryWidth_finitePairHistoryOfTrace {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (c : Real) (n : Nat) (arm : Fin K) : realHistoryWidth c n (History.finitePairHistoryOfTrace action reward n) arm = realWidth action c arm (n + 1)","missing":[],"search":"realhistorywidth_finitepairhistoryoftrace banditrlproof.ucb.realhistorywidth_finitepairhistoryoftrace inclusive finite-history widths agree with the trace width at `n + 1`. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryIndex_finitePairHistoryOfTrace","label":"realHistoryIndex_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryIndex_finitePairHistoryOfTrace","description":"Inclusive finite-history UCB scores agree with the trace score.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-414c01c91969","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3121,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:226"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realHistoryIndex_finitePairHistoryOfTrace {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (c : Real) (n : Nat) (arm : Fin K) : realHistoryIndex c n (History.finitePairHistoryOfTrace action reward n) arm = realIndex action reward c arm (n + 1)","missing":[],"search":"realhistoryindex_finitepairhistoryoftrace banditrlproof.ucb.realhistoryindex_finitepairhistoryoftrace inclusive finite-history ucb scores agree with the trace score. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realHistoryIndexAction_finitePairHistoryOfTrace","label":"realHistoryIndexAction_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realHistoryIndexAction_finitePairHistoryOfTrace","description":"The least-encoded history selector is exactly the trace selector.","url":"../modules/banditrlproof-algorithms-ucbrealhistoryindex/index.html#decl-2ca71207df93","parent":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","order":3122,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealHistoryIndex"],["Source","BanditRLProof/Algorithms/UCBRealHistoryIndex.lean:237"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realHistoryIndexAction_finitePairHistoryOfTrace {K : Nat} (hK : 0 < K) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (c : Real) (n : Nat) : realHistoryIndexAction hK c n (History.finitePairHistoryOfTrace action reward n) = realIndexAction hK action reward c (n + 1)","missing":[],"search":"realhistoryindexaction_finitepairhistoryoftrace banditrlproof.ucb.realhistoryindexaction_finitepairhistoryoftrace the least-encoded history selector is exactly the trace selector. theorem compiled","shard":"modules/1bf735e22a8440f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.RealStationaryUCBSequence","label":"RealStationaryUCBSequence","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UCB.RealStationaryUCBSequence","description":"The exact `IsAlgEnvSeq`-shaped fields consumed by the local stationary Real UCB route. The law fields correspond to the pinned source's initial action law, initial feedback law, successor action law given finite observable history, and successor feedback law given history and the new action. They are stated as Mathlib `condDistrib` equalities because LML's `HasCondDistrib` symbol is not a local dependency.","url":"../modules/banditrlproof-algorithms-ucbreallmlcompat/index.html#decl-d01b3a524d8b","parent":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","order":3123,"meta":[["Kind","structure"],["Module","BanditRLProof.Algorithms.UCBRealLMLCompat"],["Source","BanditRLProof/Algorithms/UCBRealLMLCompat.lean:28"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"structure RealStationaryUCBSequence {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) : Prop where","missing":[],"search":"realstationaryucbsequence banditrlproof.ucb.realstationaryucbsequence the exact `isalgenvseq`-shaped fields consumed by the local stationary real ucb route. the law fields correspond to the pinned source's initial action law, initial feedback law, successor action law given finite observable history, and successor feedback law given history and the new action. they are stated as mathlib `conddistrib` equalities because lml's `hasconddistrib` symbol is not a local dependency. structure compiled","shard":"modules/56a897e873d445d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_armStream","label":"realStationaryUCBSequence_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryUCBSequence_armStream","description":"The canonical arm-stream process satisfies the local UCB field bundle.","url":"../modules/banditrlproof-algorithms-ucbreallmlcompat/index.html#decl-630c0191addd","parent":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","order":3124,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealLMLCompat"],["Source","BanditRLProof/Algorithms/UCBRealLMLCompat.lean:88"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryUCBSequence_armStream {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : RealStationaryUCBSequence (armStreamMeasure nu) hK c sigma2 nu (armStreamAction hK (c * (sigma2 : Real))) (armStreamReward hK (c * (sigma2 : Real)))","missing":[],"search":"realstationaryucbsequence_armstream banditrlproof.ucb.realstationaryucbsequence_armstream the canonical arm-stream process satisfies the local ucb field bundle. theorem compiled","shard":"modules/56a897e873d445d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.identDistrib_actionRewardTrace_of_realStationaryUCBSequence","label":"identDistrib_actionRewardTrace_of_realStationaryUCBSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.identDistrib_actionRewardTrace_of_realStationaryUCBSequence","description":"The complete external observable trajectory has the canonical arm-stream UCB law whenever the local stationary sequence field bundle holds.","url":"../modules/banditrlproof-algorithms-ucbreallmlcompat/index.html#decl-adfc7053e381","parent":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","order":3125,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealLMLCompat"],["Source","BanditRLProof/Algorithms/UCBRealLMLCompat.lean:109"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_actionRewardTrace_of_realStationaryUCBSequence {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (h : RealStationaryUCBSequence mu hK c sigma2 nu action reward) : IdentDistrib (fun omega t => (action omega t, reward omega t)) (fun stream t => (armStreamAction hK (c * (sigma2 : Real)) stream t, armStreamReward hK (c * (sigma2 : Real)) stream t)) mu (armStreamMeasure nu)","missing":[],"search":"identdistrib_actionrewardtrace_of_realstationaryucbsequence banditrlproof.ucb.identdistrib_actionrewardtrace_of_realstationaryucbsequence the complete external observable trajectory has the canonical arm-stream ucb law whenever the local stationary sequence field bundle holds. theorem compiled","shard":"modules/56a897e873d445d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.regret_le_of_realStationaryUCBSequence","label":"regret_le_of_realStationaryUCBSequence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.regret_le_of_realStationaryUCBSequence","description":"Exact native Real UCB regret from the bundled stationary sequence fields. This is the local theorem corresponding to the mathematical route of pinned LML `Bandits.UCB.regret_le`. A theorem about the imported LML `IsAlgEnvSeq` symbol still requires a common Lean/mathlib toolchain.","url":"../modules/banditrlproof-algorithms-ucbreallmlcompat/index.html#decl-6772de346993","parent":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","order":3126,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealLMLCompat"],["Source","BanditRLProof/Algorithms/UCBRealLMLCompat.lean:136"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem regret_le_of_realStationaryUCBSequence {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (h : RealStationaryUCBSequence mu hK c sigma2 nu action reward) (n : Nat) (hc : 0 < c) (hsigma2 : sigma2 ≠ 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun x => x - realKernelMean nu arm) sigma2 (nu arm)) : integral mu (fun omega => realKernelRegret nu (action omega) n) <= (Finset.univ : Finset (Fin K)).sum (fun arm => 8 * c * (sigma2 : Real) * Real.log ((n + 1 : Nat) : Real) / realKernelGap nu arm + realKernelGap nu arm * (2 + 2 * (constSum c n).toReal))","missing":[],"search":"regret_le_of_realstationaryucbsequence banditrlproof.ucb.regret_le_of_realstationaryucbsequence exact native real ucb regret from the bundled stationary sequence fields. this is the local theorem corresponding to the mathematical route of pinned lml `bandits.ucb.regret_le`. a theorem about the imported lml `isalgenvseq` symbol still requires a common lean/mathlib toolchain. theorem compiled","shard":"modules/56a897e873d445d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm","label":"canonicalArmStreamHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm","description":"Canonical arm-stream initial and successor action laws as a history algorithm.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-a2778b190227","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3127,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:22"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalArmStreamHistoryAlgorithm {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Thompson.HistoryAlgorithm (Fin K) Real where","missing":[],"search":"canonicalarmstreamhistoryalgorithm banditrlproof.ucb.canonicalarmstreamhistoryalgorithm canonical arm-stream initial and successor action laws as a history algorithm. definition compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment","label":"canonicalArmStreamHistoryEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment","description":"Canonical arm-stream initial and successor reward laws as a history environment.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-5e7e8812d6e9","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3128,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:43"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalArmStreamHistoryEnvironment {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Thompson.HistoryEnvironment (Fin K) Real where","missing":[],"search":"canonicalarmstreamhistoryenvironment banditrlproof.ucb.canonicalarmstreamhistoryenvironment canonical arm-stream initial and successor reward laws as a history environment. definition compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryMeasure","label":"canonicalKernelTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryMeasure","description":"Independently regenerated pair trajectory from the canonical UCB split kernels.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-058084528061","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3129,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:65"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalKernelTrajectoryMeasure {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Measure ((n : Nat) -> Fin K × Real)","missing":[],"search":"canonicalkerneltrajectorymeasure banditrlproof.ucb.canonicalkerneltrajectorymeasure independently regenerated pair trajectory from the canonical ucb split kernels. definition compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmCanonicalKernelTrajectoryMeasure","label":"finiteArmCanonicalKernelTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmCanonicalKernelTrajectoryMeasure","description":"Finite-arm-law specialization of the canonical-kernel trajectory measure.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-9b0d504fc1cf","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3130,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:84"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmCanonicalKernelTrajectoryMeasure {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) : Measure ((n : Nat) -> Fin K × Real)","missing":[],"search":"finitearmcanonicalkerneltrajectorymeasure banditrlproof.ucb.finitearmcanonicalkerneltrajectorymeasure finite-arm-law specialization of the canonical-kernel trajectory measure. definition compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction","label":"canonicalKernelTrajectoryAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryAction","description":"Action coordinates of the canonical-kernel pair trajectory.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-201a7a7fa770","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3131,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:109"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def canonicalKernelTrajectoryAction {K : Nat} : ((n : Nat) -> Fin K × Real) -> ActionTrace (Fin K)","missing":[],"search":"canonicalkerneltrajectoryaction banditrlproof.ucb.canonicalkerneltrajectoryaction action coordinates of the canonical-kernel pair trajectory. definition compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryReward","label":"canonicalKernelTrajectoryReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryReward","description":"Reward coordinates of the canonical-kernel pair trajectory.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-4c0f633efb62","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3132,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:114"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def canonicalKernelTrajectoryReward {K : Nat} : ((n : Nat) -> Fin K × Real) -> RewardTrace Real","missing":[],"search":"canonicalkerneltrajectoryreward banditrlproof.ucb.canonicalkerneltrajectoryreward reward coordinates of the canonical-kernel pair trajectory. definition compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_canonicalKernelTrajectory","label":"realStationaryUCBSequence_canonicalKernelTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryUCBSequence_canonicalKernelTrajectory","description":"The independently generated canonical-kernel trajectory satisfies all local stationary UCB split conditional-law fields.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-e349fa810b41","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3133,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:122"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryUCBSequence_canonicalKernelTrajectory {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : RealStationaryUCBSequence (canonicalKernelTrajectoryMeasure hK c sigma2 nu) hK c sigma2 nu canonicalKernelTrajectoryAction canonicalKernelTrajectoryReward","missing":[],"search":"realstationaryucbsequence_canonicalkerneltrajectory banditrlproof.ucb.realstationaryucbsequence_canonicalkerneltrajectory the independently generated canonical-kernel trajectory satisfies all local stationary ucb split conditional-law fields. theorem compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","label":"canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","description":"Armwise-bounded laws give the logarithmic expected-regret envelope on the independently regenerated canonical-kernel trajectory.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-5824be7774f1","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3134,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:168"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le {K : Nat} [NeZero K] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) (n : Nat) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) 0 <= realStationaryArmwiseBoundedFiniteArmExpectedRegret (finiteArmCanonicalKernelTrajectoryMeasure hK 4 sigma2 armLaw hprob) armLaw canonicalKernelTrajectoryAction n /\\ realStationaryArmwiseBoundedFiniteArmExpectedRegret (finiteArmCanonicalKernelTrajectoryMeasure hK 4 sigma2 armLaw hprob) armLaw canonicalKernelTrajectoryAction n <= realStationaryArmwiseBoundedFiniteArmModelCoefficient armLaw hprob lo hi * (1 + Real.log…","missing":[],"search":"canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedregret_nonneg_and_le banditrlproof.ucb.canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedregret_nonneg_and_le armwise-bounded laws give the logarithmic expected-regret envelope on the independently regenerated canonical-kernel trajectory. theorem compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","label":"canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","description":"The expected regret per round of the independently regenerated canonical-kernel UCB trajectory tends to zero.","url":"../modules/banditrlproof-algorithms-ucbrealstationarycanonicalkerneltrajectory/index.html#decl-98ca98b168a2","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","order":3135,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory"],["Source","BanditRLProof/Algorithms/UCBRealStationaryCanonicalKernelTrajectory.lean:218"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero {K : Nat} [NeZero K] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) Tendsto (realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret (finiteArmCanonicalKernelTrajectoryMeasure hK 4 sigma2 armLaw hprob) armLaw canonicalKernelTrajectoryAction) atTop (nhds 0)","missing":[],"search":"canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero banditrlproof.ucb.canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero the expected regret per round of the independently regenerated canonical-kernel ucb trajectory tends to zero. theorem compiled","shard":"modules/c6a3921acc084883.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalRealUCBHistorySelector","label":"canonicalRealUCBHistorySelector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalRealUCBHistorySelector","description":"The explicit finite-history selector encoded by the canonical arm-stream policy.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-172b4ece2a16","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3136,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:20"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalRealUCBHistorySelector {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (n : Nat) : History.FinitePairHistory (Fin K) Real n -> Fin K","missing":[],"search":"canonicalrealucbhistoryselector banditrlproof.ucb.canonicalrealucbhistoryselector the explicit finite-history selector encoded by the canonical arm-stream policy. definition compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalRealUCBPolicyKernel","label":"canonicalRealUCBPolicyKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalRealUCBPolicyKernel","description":"The explicit deterministic successor-action kernel for stationary Real UCB.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-da1a1e2eaf43","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3137,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:26"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalRealUCBPolicyKernel {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (n : Nat) : Kernel (History.FinitePairHistory (Fin K) Real n) (Fin K)","missing":[],"search":"canonicalrealucbpolicykernel banditrlproof.ucb.canonicalrealucbpolicykernel the explicit deterministic successor-action kernel for stationary real ucb. definition compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurable_canonicalRealUCBHistorySelector","label":"measurable_canonicalRealUCBHistorySelector","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.measurable_canonicalRealUCBHistorySelector","description":"The explicit canonical selector is measurable on every finite history.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-2f198723f204","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3138,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:34"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalRealUCBHistorySelector {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (n : Nat) : Measurable (canonicalRealUCBHistorySelector hK c sigma2 n)","missing":[],"search":"measurable_canonicalrealucbhistoryselector banditrlproof.ucb.measurable_canonicalrealucbhistoryselector the explicit canonical selector is measurable on every finite history. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalRealUCBPolicyKernel_apply","label":"canonicalRealUCBPolicyKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalRealUCBPolicyKernel_apply","description":"Every explicit policy-kernel section is the corresponding Dirac law.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-33ab3c6c1403","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3139,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:41"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalRealUCBPolicyKernel_apply {K : Nat} (hK : 0 < K) (c : Real) (sigma2 : NNReal) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : canonicalRealUCBPolicyKernel hK c sigma2 n history = Measure.dirac (canonicalRealUCBHistorySelector hK c sigma2 n history)","missing":[],"search":"canonicalrealucbpolicykernel_apply banditrlproof.ucb.canonicalrealucbpolicykernel_apply every explicit policy-kernel section is the corresponding dirac law. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm_initialAction_eq_dirac","label":"canonicalArmStreamHistoryAlgorithm_initialAction_eq_dirac","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm_initialAction_eq_dirac","description":"The canonical arm-stream initial-action package is the fixed initial arm.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-80d3da70ad18","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3140,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:49"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalArmStreamHistoryAlgorithm_initialAction_eq_dirac {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : (canonicalArmStreamHistoryAlgorithm hK c sigma2 nu).initialAction = Measure.dirac (initializationArm hK 0)","missing":[],"search":"canonicalarmstreamhistoryalgorithm_initialaction_eq_dirac banditrlproof.ucb.canonicalarmstreamhistoryalgorithm_initialaction_eq_dirac the canonical arm-stream initial-action package is the fixed initial arm. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm_policy_ae_eq_explicitPolicyKernel","label":"canonicalArmStreamHistoryAlgorithm_policy_ae_eq_explicitPolicyKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm_policy_ae_eq_explicitPolicyKernel","description":"The canonical arm-stream successor action conditional law is the explicit deterministic UCB policy kernel on its finite-history marginal.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-7c83f713dac6","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3141,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:61"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalArmStreamHistoryAlgorithm_policy_ae_eq_explicitPolicyKernel {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Filter.EventuallyEq (ae ((armStreamMeasure nu).map (fun stream : ArmRewardStream K => History.finitePairHistoryOfTrace (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) n))) ((canonicalArmStreamHistoryAlgorithm hK c sigma2 nu).policy n) (canonicalRealUCBPolicyKernel hK c sigma2 n)","missing":[],"search":"canonicalarmstreamhistoryalgorithm_policy_ae_eq_explicitpolicykernel banditrlproof.ucb.canonicalarmstreamhistoryalgorithm_policy_ae_eq_explicitpolicykernel the canonical arm-stream successor action conditional law is the explicit deterministic ucb policy kernel on its finite-history marginal. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectory_finitePairHistory_map_eq_armStream","label":"canonicalKernelTrajectory_finitePairHistory_map_eq_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectory_finitePairHistory_map_eq_armStream","description":"Every generated finite-pair-history marginal agrees with the corresponding canonical arm-stream history marginal.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-30ce2d5eefd4","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3142,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:95"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectory_finitePairHistory_map_eq_armStream {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun trajectory => History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n) (canonicalKernelTrajectoryMeasure hK c sigma2 nu) = Measure.map (fun stream : ArmRewardStream K => History.finitePairHistoryOfTrace (armStreamAction hK (c * (sigma2 : Real)) stream) (armStreamReward hK (c * (sigma2 : Real)) stream) n) (armStreamMeasure nu)","missing":[],"search":"canonicalkerneltrajectory_finitepairhistory_map_eq_armstream banditrlproof.ucb.canonicalkerneltrajectory_finitepairhistory_map_eq_armstream every generated finite-pair-history marginal agrees with the corresponding canonical arm-stream history marginal. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_condDistrib_ae_eq_explicitPolicyKernel","label":"canonicalKernelTrajectoryAction_condDistrib_ae_eq_explicitPolicyKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryAction_condDistrib_ae_eq_explicitPolicyKernel","description":"The generated successor action conditional law is the explicit deterministic UCB policy kernel on the generated finite-history marginal.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-806418a8d46d","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3143,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:137"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryAction_condDistrib_ae_eq_explicitPolicyKernel {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : condDistrib (fun trajectory => canonicalKernelTrajectoryAction trajectory (n + 1)) (fun trajectory => History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n) (canonicalKernelTrajectoryMeasure hK c sigma2 nu) =ᵐ[ (canonicalKernelTrajectoryMeasure hK c sigma2 nu).map (fun trajectory => History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n)] canonicalRealUCBPolicyKernel hK c sigma2 n","missing":[],"search":"canonicalkerneltrajectoryaction_conddistrib_ae_eq_explicitpolicykernel banditrlproof.ucb.canonicalkerneltrajectoryaction_conddistrib_ae_eq_explicitpolicykernel the generated successor action conditional law is the explicit deterministic ucb policy kernel on the generated finite-history marginal. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_zero_ae_eq_initializationArm","label":"canonicalKernelTrajectoryAction_zero_ae_eq_initializationArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryAction_zero_ae_eq_initializationArm","description":"The generated initial action is almost surely the canonical initial arm.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-53a44816bf3a","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3144,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:166"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryAction_zero_ae_eq_initializationArm {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Filter.EventuallyEq (ae (canonicalKernelTrajectoryMeasure hK c sigma2 nu)) (fun trajectory => canonicalKernelTrajectoryAction trajectory 0) (fun _trajectory => initializationArm hK 0)","missing":[],"search":"canonicalkerneltrajectoryaction_zero_ae_eq_initializationarm banditrlproof.ucb.canonicalkerneltrajectoryaction_zero_ae_eq_initializationarm the generated initial action is almost surely the canonical initial arm. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm","label":"canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm","description":"Every generated successor action follows the explicit UCB selector a.s.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-b193d53b1410","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3145,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:198"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Filter.Eventually (fun trajectory => canonicalKernelTrajectoryAction trajectory (n + 1) = canonicalRealUCBHistorySelector hK c sigma2 n (History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n)) (ae (canonicalKernelTrajectoryMeasure hK c sigma2 nu))","missing":[],"search":"canonicalkerneltrajectoryaction_succ_ae_eq_realhistorynextarm banditrlproof.ucb.canonicalkerneltrajectoryaction_succ_ae_eq_realhistorynextarm every generated successor action follows the explicit ucb selector a.s. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm_all","label":"canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm_all","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm_all","description":"One full-measure event carries every successor explicit-policy equality.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-9ebac5a2c396","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3146,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:255"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm_all {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Filter.Eventually (fun trajectory => forall n : Nat, canonicalKernelTrajectoryAction trajectory (n + 1) = canonicalRealUCBHistorySelector hK c sigma2 n (History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n)) (ae (canonicalKernelTrajectoryMeasure hK c sigma2 nu))","missing":[],"search":"canonicalkerneltrajectoryaction_succ_ae_eq_realhistorynextarm_all banditrlproof.ucb.canonicalkerneltrajectoryaction_succ_ae_eq_realhistorynextarm_all one full-measure event carries every successor explicit-policy equality. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_follows_realHistoryNextArm_ae","label":"canonicalKernelTrajectoryAction_follows_realHistoryNextArm_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryAction_follows_realHistoryNextArm_ae","description":"The complete generated action trace follows the explicit UCB policy a.s.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-49dc630c1c7d","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3147,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:272"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryAction_follows_realHistoryNextArm_ae {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : Filter.Eventually (fun trajectory => canonicalKernelTrajectoryAction trajectory 0 = initializationArm hK 0 /\\ forall n : Nat, canonicalKernelTrajectoryAction trajectory (n + 1) = canonicalRealUCBHistorySelector hK c sigma2 n (History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n)) (ae (canonicalKernelTrajectoryMeasure hK c sigma2 nu))","missing":[],"search":"canonicalkerneltrajectoryaction_follows_realhistorynextarm_ae banditrlproof.ucb.canonicalkerneltrajectoryaction_follows_realhistorynextarm_ae the complete generated action trace follows the explicit ucb policy a.s. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy","label":"canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy","description":"Armwise-bounded finite-arm laws give one fixed generated process whose actions follow explicit Real UCB almost surely and whose expected average pseudo-regret tends to zero.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryexplicitpolicy/index.html#decl-19a2b57d4e19","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","order":3148,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy"],["Source","BanditRLProof/Algorithms/UCBRealStationaryExplicitPolicy.lean:299"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy {K : Nat} [NeZero K] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) Tendsto (realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret (finiteArmCanonicalKernelTrajectoryMeasure hK 4 sigma2 armLaw hprob) armLaw canonicalKernelTrajectoryAction) atTop (nhds 0) /\\ Filter.Eventually (fun trajectory => canonicalKernelTrajectoryAction trajectory 0 = initializationArm hK 0 /\\ forall n : Nat, canonicalKernelTrajectoryAction trajectory (n + 1) = realHistoryNextArm hK (4 * (sigma2…","missing":[],"search":"canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero_and_explicitpolicy banditrlproof.ucb.canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero_and_explicitpolicy armwise-bounded finite-arm laws give one fixed generated process whose actions follow explicit real ucb almost surely and whose expected average pseudo-regret tends to zero. theorem compiled","shard":"modules/09bfaa40c7f3aaa0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryExpectedRegret","label":"realStationaryExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryExpectedRegret","description":"Expected Real pseudo-regret of one fixed external action process.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-4c9b55342251","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3149,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:20"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realStationaryExpectedRegret {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : Kernel (Fin K) Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) : Real","missing":[],"search":"realstationaryexpectedregret banditrlproof.ucb.realstationaryexpectedregret expected real pseudo-regret of one fixed external action process. definition compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryExpectedRegret_eq_armStreamExpectedRegret","label":"realStationaryExpectedRegret_eq_armStreamExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryExpectedRegret_eq_armStreamExpectedRegret","description":"The stationary UCB field bundle identifies each external expected-regret term exactly with the corresponding canonical arm-stream term at `c = 4`.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-d40fa0031895","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3150,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:30"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryExpectedRegret_eq_armStreamExpectedRegret {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (h : RealStationaryUCBSequence mu hK 4 sigma2 nu action reward) (n : Nat) : realStationaryExpectedRegret mu nu action n = armStreamExpectedRegret hK sigma2 nu n","missing":[],"search":"realstationaryexpectedregret_eq_armstreamexpectedregret banditrlproof.ucb.realstationaryexpectedregret_eq_armstreamexpectedregret the stationary ucb field bundle identifies each external expected-regret term exactly with the corresponding canonical arm-stream term at `c = 4`. theorem compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryExpectedAverageRegret","label":"realStationaryExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryExpectedAverageRegret","description":"External expected regret normalized by `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-02c197735496","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3151,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:54"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realStationaryExpectedAverageRegret {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : Kernel (Fin K) Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) : Real","missing":[],"search":"realstationaryexpectedaverageregret banditrlproof.ucb.realstationaryexpectedaverageregret external expected regret normalized by `n + 1`. definition compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryExpectedAverageRegret_tendsto_zero","label":"realStationaryExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryExpectedAverageRegret_tendsto_zero","description":"Any fixed external process satisfying the stationary UCB field bundle inherits the canonical sub-Gaussian expected-average consistency theorem.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-47ca831f6025","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3152,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:64"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryExpectedAverageRegret_tendsto_zero {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (h : RealStationaryUCBSequence mu hK 4 sigma2 nu action reward) (hsigma2 : Ne sigma2 0) (hsubG : forall arm : Fin K, HasSubgaussianMGF (fun x => x - realKernelMean nu arm) sigma2 (nu arm)) : Tendsto (realStationaryExpectedAverageRegret mu nu action) atTop (nhds 0)","missing":[],"search":"realstationaryexpectedaverageregret_tendsto_zero banditrlproof.ucb.realstationaryexpectedaverageregret_tendsto_zero any fixed external process satisfying the stationary ucb field bundle inherits the canonical sub-gaussian expected-average consistency theorem. theorem compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedRegret","label":"realStationaryArmwiseBoundedFiniteArmExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedRegret","description":"Expected regret of an external process over finite Real arm laws.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-0da583f10e59","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3153,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:89"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realStationaryArmwiseBoundedFiniteArmExpectedRegret {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (armLaw : Fin K -> Measure Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) : Real","missing":[],"search":"realstationaryarmwiseboundedfinitearmexpectedregret banditrlproof.ucb.realstationaryarmwiseboundedfinitearmexpectedregret expected regret of an external process over finite real arm laws. definition compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmModelCoefficient","label":"realStationaryArmwiseBoundedFiniteArmModelCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmModelCoefficient","description":"Canonical logarithmic-envelope coefficient for armwise-bounded laws.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-471a7c9dff0a","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3154,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:96"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realStationaryArmwiseBoundedFiniteArmModelCoefficient {K : Nat} (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) : Real","missing":[],"search":"realstationaryarmwiseboundedfinitearmmodelcoefficient banditrlproof.ucb.realstationaryarmwiseboundedfinitearmmodelcoefficient canonical logarithmic-envelope coefficient for armwise-bounded laws. definition compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","label":"realStationaryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","description":"External armwise-bounded laws inherit the canonical logarithmic envelope.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-94b945e47cd8","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3155,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:108"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) (h : @RealStationaryUCBSequence Omega K _ _ mu _ hK 4 (Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm))) (finiteArmRealRewardKernel armLaw) (finiteArmRealRewardKernel_isMarkov armLaw hprob) action reward) (n : Nat) : 0 <= realStationaryArmwiseBoundedFiniteArmExpectedRegret mu armLaw action n /\\ realStationaryArmwiseBoundedFiniteArmExpectedRegret mu…","missing":[],"search":"realstationaryarmwiseboundedfinitearmexpectedregret_nonneg_and_le banditrlproof.ucb.realstationaryarmwiseboundedfinitearmexpectedregret_nonneg_and_le external armwise-bounded laws inherit the canonical logarithmic envelope. theorem compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret","label":"realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret","description":"External finite-arm expected regret normalized by `n + 1`.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-248319dc624f","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3156,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:156"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (armLaw : Fin K -> Measure Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) : Real","missing":[],"search":"realstationaryarmwiseboundedfinitearmexpectedaverageregret banditrlproof.ucb.realstationaryarmwiseboundedfinitearmexpectedaverageregret external finite-arm expected regret normalized by `n + 1`. definition compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","label":"realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","description":"An external stationary UCB process driven by armwise-bounded finite Real laws has vanishing expected average regret. The same external measure, action trace, and reward trace are used at every horizon.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryfinitearmrewardlaws/index.html#decl-0f272dcee477","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","order":3157,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws"],["Source","BanditRLProof/Algorithms/UCBRealStationaryFiniteArmRewardLaws.lean:168"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (action : Omega -> ActionTrace (Fin K)) (reward : Omega -> RewardTrace Real) (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) (h : @RealStationaryUCBSequence Omega K _ _ mu _ hK 4 (Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm))) (finiteArmRealRewardKernel armLaw) (finiteArmRealRewardKernel_isMarkov armLaw hprob) action reward) : Tendsto (realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret mu armLaw action) atTop (nhds 0)","missing":[],"search":"realstationaryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero banditrlproof.ucb.realstationaryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero an external stationary ucb process driven by armwise-bounded finite real laws has vanishing expected average regret. the same external measure, action trace, and reward trace are used at every horizon. theorem compiled","shard":"modules/9271d25e3a91d1bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.condDistrib_comp_measurePreserving","label":"condDistrib_comp_measurePreserving","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.condDistrib_comp_measurePreserving","description":"Conditional distributions pull back through a measure-preserving map.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-e284a43633f1","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3158,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:27"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_comp_measurePreserving {Alpha : Type u} {Beta : Type v} {Output : Type w} {Source : Type x} [MeasurableSpace Alpha] [MeasurableSpace Beta] [MeasurableSpace Output] [StandardBorelSpace Output] [Nonempty Output] [MeasurableSpace Source] (mu : Measure Source) [IsFiniteMeasure mu] (nu : Measure Alpha) [IsFiniteMeasure nu] (source : Source -> Alpha) (hsource : MeasurePreserving source mu nu) (X : Alpha -> Beta) (Y : Alpha -> Output) (hX : Measurable X) (hY : Measurable Y) : condDistrib (Y ∘ source) (X ∘ source) mu =ᵐ[mu.map (X ∘ source)] condDistrib Y X nu","missing":[],"search":"conddistrib_comp_measurepreserving banditrlproof.ucb.conddistrib_comp_measurepreserving conditional distributions pull back through a measure-preserving map. theorem compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurePreservingArmStreamAction","label":"measurePreservingArmStreamAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.measurePreservingArmStreamAction","description":"Canonical UCB action trace composed with an external stream source.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-251c8645441f","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3159,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:54"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def measurePreservingArmStreamAction {Omega : Type u} {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (source : Omega -> ArmRewardStream K) : Omega -> ActionTrace (Fin K)","missing":[],"search":"measurepreservingarmstreamaction banditrlproof.ucb.measurepreservingarmstreamaction canonical ucb action trace composed with an external stream source. definition compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.measurePreservingArmStreamReward","label":"measurePreservingArmStreamReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.measurePreservingArmStreamReward","description":"Canonical selected-reward trace composed with an external stream source.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-f17be611157f","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3160,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:62"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def measurePreservingArmStreamReward {Omega : Type u} {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (source : Omega -> ArmRewardStream K) : Omega -> RewardTrace Real","missing":[],"search":"measurepreservingarmstreamreward banditrlproof.ucb.measurepreservingarmstreamreward canonical selected-reward trace composed with an external stream source. definition compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_comp_measurePreserving_armStream","label":"realStationaryUCBSequence_comp_measurePreserving_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryUCBSequence_comp_measurePreserving_armStream","description":"A measure-preserving source of canonical arm streams produces every field of the local stationary UCB compatibility bundle on the external sample space.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-4fd757bab78a","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3161,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:73"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryUCBSequence_comp_measurePreserving_armStream {Omega : Type u} {K : Nat} [NeZero K] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (source : Omega -> ArmRewardStream K) (hsource : MeasurePreserving source mu (armStreamMeasure nu)) : RealStationaryUCBSequence mu hK c sigma2 nu (measurePreservingArmStreamAction hK c sigma2 source) (measurePreservingArmStreamReward hK c sigma2 source)","missing":[],"search":"realstationaryucbsequence_comp_measurepreserving_armstream banditrlproof.ucb.realstationaryucbsequence_comp_measurepreserving_armstream a measure-preserving source of canonical arm streams produces every field of the local stationary ucb compatibility bundle on the external sample space. theorem compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.productNoiseArmStreamAction","label":"productNoiseArmStreamAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.productNoiseArmStreamAction","description":"Canonical UCB action on a stream product with auxiliary noise.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-032897ca19ed","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3162,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:149"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def productNoiseArmStreamAction {K : Nat} [NeZero K] {Aux : Type u} (hK : 0 < K) (c : Real) (sigma2 : NNReal) : ArmRewardStream K × Aux -> ActionTrace (Fin K)","missing":[],"search":"productnoisearmstreamaction banditrlproof.ucb.productnoisearmstreamaction canonical ucb action on a stream product with auxiliary noise. definition compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.productNoiseArmStreamReward","label":"productNoiseArmStreamReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.productNoiseArmStreamReward","description":"Canonical selected reward on a stream product with auxiliary noise.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-8ea59574a745","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3163,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:156"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def productNoiseArmStreamReward {K : Nat} [NeZero K] {Aux : Type u} (hK : 0 < K) (c : Real) (sigma2 : NNReal) : ArmRewardStream K × Aux -> RewardTrace Real","missing":[],"search":"productnoisearmstreamreward banditrlproof.ucb.productnoisearmstreamreward canonical selected reward on a stream product with auxiliary noise. definition compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_productNoise_armStream","label":"realStationaryUCBSequence_productNoise_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.realStationaryUCBSequence_productNoise_armStream","description":"Product-first projection is a concrete external stationary UCB source.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-687c5dad9c22","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3164,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:163"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem realStationaryUCBSequence_productNoise_armStream {K : Nat} [NeZero K] {Aux : Type u} [MeasurableSpace Aux] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (auxMu : Measure Aux) [IsProbabilityMeasure auxMu] : RealStationaryUCBSequence ((armStreamMeasure nu).prod auxMu) hK c sigma2 nu (productNoiseArmStreamAction hK c sigma2) (productNoiseArmStreamReward hK c sigma2)","missing":[],"search":"realstationaryucbsequence_productnoise_armstream banditrlproof.ucb.realstationaryucbsequence_productnoise_armstream product-first projection is a concrete external stationary ucb source. theorem compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.finiteArmProductNoiseMeasure","label":"finiteArmProductNoiseMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCB.finiteArmProductNoiseMeasure","description":"Product law for finite Real arm streams and independent auxiliary noise.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-bd665137d3bf","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3165,"meta":[["Kind","definition"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:177"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmProductNoiseMeasure {K : Nat} {Aux : Type u} [MeasurableSpace Aux] (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (auxMu : Measure Aux) : Measure (ArmRewardStream K × Aux)","missing":[],"search":"finitearmproductnoisemeasure banditrlproof.ucb.finitearmproductnoisemeasure product law for finite real arm streams and independent auxiliary noise. definition compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.productNoiseArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","label":"productNoiseArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.productNoiseArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","description":"Armwise-bounded laws give the canonical logarithmic expected-regret envelope on a product sample space carrying arbitrary independent auxiliary noise.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-232a5e3aeddd","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3166,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:191"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem productNoiseArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le {K : Nat} [NeZero K] {Aux : Type u} [MeasurableSpace Aux] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (auxMu : Measure Aux) [IsProbabilityMeasure auxMu] (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) (n : Nat) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) 0 <= realStationaryArmwiseBoundedFiniteArmExpectedRegret (finiteArmProductNoiseMeasure armLaw hprob auxMu) armLaw (productNoiseArmStreamAction hK 4 sigma2) n /\\ realStationaryArmwiseBoundedFiniteArmExpectedRegret (finiteArmProductNoiseMeasure armLaw hprob auxMu) armLaw (productNoiseArmStreamAction hK 4 sigma2) n <= realStationaryArmwiseBoundedFini…","missing":[],"search":"productnoisearmwiseboundedfinitearmexpectedregret_nonneg_and_le banditrlproof.ucb.productnoisearmwiseboundedfinitearmexpectedregret_nonneg_and_le armwise-bounded laws give the canonical logarithmic expected-regret envelope on a product sample space carrying arbitrary independent auxiliary noise. theorem compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.productNoiseArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","label":"productNoiseArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.productNoiseArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","description":"The expected regret per round of product-noise stationary UCB with armwise bounded finite Real reward laws tends to zero.","url":"../modules/banditrlproof-algorithms-ucbrealstationarymeasurepreservingsource/index.html#decl-b369679fd85c","parent":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","order":3167,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource"],["Source","BanditRLProof/Algorithms/UCBRealStationaryMeasurePreservingSource.lean:240"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem productNoiseArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero {K : Nat} [NeZero K] {Aux : Type u} [MeasurableSpace Aux] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (auxMu : Measure Aux) [IsProbabilityMeasure auxMu] (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) Tendsto (realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret (finiteArmProductNoiseMeasure armLaw hprob auxMu) armLaw (productNoiseArmStreamAction hK 4 sigma2)) atTop (nhds 0)","missing":[],"search":"productnoisearmwiseboundedfinitearmexpectedaverageregret_tendsto_zero banditrlproof.ucb.productnoisearmwiseboundedfinitearmexpectedaverageregret_tendsto_zero the expected regret per round of product-noise stationary ucb with armwise bounded finite real reward laws tends to zero. theorem compiled","shard":"modules/8383a63eb5f3bdc1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectory_historyAction_map_eq_armStream","label":"canonicalKernelTrajectory_historyAction_map_eq_armStream","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectory_historyAction_map_eq_armStream","description":"The generated history/action condition marginal agrees with its canonical arm-stream counterpart.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryselectedrewardconsistency/index.html#decl-441929633b15","parent":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","order":3168,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency"],["Source","BanditRLProof/Algorithms/UCBRealStationarySelectedRewardConsistency.lean:22"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectory_historyAction_map_eq_armStream {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : Measure.map (fun trajectory => (History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n, canonicalKernelTrajectoryAction trajectory (n + 1))) (canonicalKernelTrajectoryMeasure hK c sigma2 nu) = Measure.map (armStreamHistoryAction hK (c * (sigma2 : Real)) n) (armStreamMeasure nu)","missing":[],"search":"canonicalkerneltrajectory_historyaction_map_eq_armstream banditrlproof.ucb.canonicalkerneltrajectory_historyaction_map_eq_armstream the generated history/action condition marginal agrees with its canonical arm-stream counterpart. theorem compiled","shard":"modules/bc73ac9a07423c53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryReward_zero_condDistrib_ae_eq_nu","label":"canonicalKernelTrajectoryReward_zero_condDistrib_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryReward_zero_condDistrib_ae_eq_nu","description":"The generated initial reward has the stationary law of its selected arm.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryselectedrewardconsistency/index.html#decl-0fd1c70e07fe","parent":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","order":3169,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency"],["Source","BanditRLProof/Algorithms/UCBRealStationarySelectedRewardConsistency.lean:93"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryReward_zero_condDistrib_ae_eq_nu {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] : condDistrib (fun trajectory => canonicalKernelTrajectoryReward trajectory 0) (fun trajectory => canonicalKernelTrajectoryAction trajectory 0) (canonicalKernelTrajectoryMeasure hK c sigma2 nu) =ᵐ[ (canonicalKernelTrajectoryMeasure hK c sigma2 nu).map (fun trajectory => canonicalKernelTrajectoryAction trajectory 0)] nu","missing":[],"search":"canonicalkerneltrajectoryreward_zero_conddistrib_ae_eq_nu banditrlproof.ucb.canonicalkerneltrajectoryreward_zero_conddistrib_ae_eq_nu the generated initial reward has the stationary law of its selected arm. theorem compiled","shard":"modules/bc73ac9a07423c53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryReward_succ_condDistrib_ae_eq_nu","label":"canonicalKernelTrajectoryReward_succ_condDistrib_ae_eq_nu","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryReward_succ_condDistrib_ae_eq_nu","description":"Every generated successor reward has the stationary law of the selected arm.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryselectedrewardconsistency/index.html#decl-286afa1c1204","parent":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","order":3170,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency"],["Source","BanditRLProof/Algorithms/UCBRealStationarySelectedRewardConsistency.lean:126"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryReward_succ_condDistrib_ae_eq_nu {K : Nat} [NeZero K] (hK : 0 < K) (c : Real) (sigma2 : NNReal) (nu : Kernel (Fin K) Real) [IsMarkovKernel nu] (n : Nat) : condDistrib (fun trajectory => canonicalKernelTrajectoryReward trajectory (n + 1)) (fun trajectory => (History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n, canonicalKernelTrajectoryAction trajectory (n + 1))) (canonicalKernelTrajectoryMeasure hK c sigma2 nu) =ᵐ[ (canonicalKernelTrajectoryMeasure hK c sigma2 nu).map (fun trajectory => (History.finitePairHistoryOfTrace (canonicalKernelTrajectoryAction trajectory) (canonicalKernelTrajectoryReward trajectory) n, canonicalKernelTrajectoryAction trajectory (n + 1)))] armStreamSelectedRewardKernel n nu","missing":[],"search":"canonicalkerneltrajectoryreward_succ_conddistrib_ae_eq_nu banditrlproof.ucb.canonicalkerneltrajectoryreward_succ_conddistrib_ae_eq_nu every generated successor reward has the stationary law of the selected arm. theorem compiled","shard":"modules/bc73ac9a07423c53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy_and_selectedRewardLaws","label":"canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy_and_selectedRewardLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy_and_selectedRewardLaws","description":"The fresh canonical trajectory simultaneously exposes its stationary selected reward laws, explicit UCB policy graph, and expected-average consistency.","url":"../modules/banditrlproof-algorithms-ucbrealstationaryselectedrewardconsistency/index.html#decl-d266aed7d9e1","parent":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","order":3171,"meta":[["Kind","theorem"],["Module","BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency"],["Source","BanditRLProof/Algorithms/UCBRealStationarySelectedRewardConsistency.lean:161"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy_and_selectedRewardLaws {K : Nat} [NeZero K] (hK : 0 < K) (armLaw : Fin K -> Measure Real) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (lo hi : Fin K -> Real) (hbound : forall arm, Filter.Eventually (fun x : Real => Set.Icc (lo arm) (hi arm) x) (ae (armLaw arm))) : let sigma2 := Concentration.finiteArmPositiveVarianceProxy (fun arm => Concentration.intervalVarianceProxy (lo arm) (hi arm)) let nu := finiteArmRealRewardKernel armLaw Tendsto (realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret (finiteArmCanonicalKernelTrajectoryMeasure hK 4 sigma2 armLaw hprob) armLaw canonicalKernelTrajectoryAction) atTop (nhds 0) /\\ Filter.Eventually (fun trajectory => canonicalKernelTrajectoryAction trajectory 0 = initializationArm hK 0 /\\ forall n : Nat, canonicalKernelTraject…","missing":[],"search":"canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero_and_explicitpolicy_and_selectedrewardlaws banditrlproof.ucb.canonicalkerneltrajectoryarmwiseboundedfinitearmexpectedaverageregret_tendsto_zero_and_explicitpolicy_and_selectedrewardlaws the fresh canonical trajectory simultaneously exposes its stationary selected reward laws, explicit ucb policy graph, and expected-average consistency. theorem compiled","shard":"modules/bc73ac9a07423c53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.HarnessProfile","label":"HarnessProfile","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.HarnessProfile","description":"inductive HarnessProfile where","url":"../modules/banditrlproof-automation/index.html#decl-5dde55d4fa3e","parent":"module:BanditRLProof.Automation","order":3172,"meta":[["Kind","inductive type"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:13"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive HarnessProfile where","missing":[],"search":"harnessprofile banditrlproof.harnessprofile inductive harnessprofile where inductive type compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.AgentRole","label":"AgentRole","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.AgentRole","description":"inductive AgentRole where","url":"../modules/banditrlproof-automation/index.html#decl-1010178b18e3","parent":"module:BanditRLProof.Automation","order":3173,"meta":[["Kind","inductive type"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:17"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive AgentRole where","missing":[],"search":"agentrole banditrlproof.agentrole inductive agentrole where inductive type compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.TaskKind","label":"TaskKind","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.TaskKind","description":"inductive TaskKind where","url":"../modules/banditrlproof-automation/index.html#decl-1aba5b0b1297","parent":"module:BanditRLProof.Automation","order":3174,"meta":[["Kind","inductive type"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:24"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive TaskKind where","missing":[],"search":"taskkind banditrlproof.taskkind inductive taskkind where inductive type compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.TaskStatus","label":"TaskStatus","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.TaskStatus","description":"inductive TaskStatus where","url":"../modules/banditrlproof-automation/index.html#decl-380f5a6e2f56","parent":"module:BanditRLProof.Automation","order":3175,"meta":[["Kind","inductive type"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:32"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive TaskStatus where","missing":[],"search":"taskstatus banditrlproof.taskstatus inductive taskstatus where inductive type compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.ArtifactSpec","label":"ArtifactSpec","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ArtifactSpec","description":"structure ArtifactSpec where","url":"../modules/banditrlproof-automation/index.html#decl-16c561f3cffd","parent":"module:BanditRLProof.Automation","order":3176,"meta":[["Kind","structure"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:40"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure ArtifactSpec where","missing":[],"search":"artifactspec banditrlproof.artifactspec structure artifactspec where structure compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.AcceptanceGate","label":"AcceptanceGate","kind":"structure","status":"compiled","subtitle":"BanditRLProof.AcceptanceGate","description":"structure AcceptanceGate where","url":"../modules/banditrlproof-automation/index.html#decl-2ffe0641b9fd","parent":"module:BanditRLProof.Automation","order":3177,"meta":[["Kind","structure"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:46"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure AcceptanceGate where","missing":[],"search":"acceptancegate banditrlproof.acceptancegate structure acceptancegate where structure compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.HarnessTask","label":"HarnessTask","kind":"structure","status":"compiled","subtitle":"BanditRLProof.HarnessTask","description":"structure HarnessTask where","url":"../modules/banditrlproof-automation/index.html#decl-abd60dd00b8d","parent":"module:BanditRLProof.Automation","order":3178,"meta":[["Kind","structure"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:53"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure HarnessTask where","missing":[],"search":"harnesstask banditrlproof.harnesstask structure harnesstask where structure compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.defaultLeanGate","label":"defaultLeanGate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.defaultLeanGate","description":"def defaultLeanGate : AcceptanceGate where","url":"../modules/banditrlproof-automation/index.html#decl-090b6cc4b6b1","parent":"module:BanditRLProof.Automation","order":3179,"meta":[["Kind","definition"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:63"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def defaultLeanGate : AcceptanceGate where","missing":[],"search":"defaultleangate banditrlproof.defaultleangate def defaultleangate : acceptancegate where definition compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.defaultHarnessProfile","label":"defaultHarnessProfile","kind":"definition","status":"compiled","subtitle":"BanditRLProof.defaultHarnessProfile","description":"def defaultHarnessProfile : HarnessProfile","url":"../modules/banditrlproof-automation/index.html#decl-8939d548cbf7","parent":"module:BanditRLProof.Automation","order":3180,"meta":[["Kind","definition"],["Module","BanditRLProof.Automation"],["Source","BanditRLProof/Automation.lean:69"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def defaultHarnessProfile : HarnessProfile","missing":[],"search":"defaultharnessprofile banditrlproof.defaultharnessprofile def defaultharnessprofile : harnessprofile definition compiled","shard":"modules/b9c9548157412714.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.intervalVarianceProxy_pos_of_lt","label":"intervalVarianceProxy_pos_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.intervalVarianceProxy_pos_of_lt","description":"A nondegenerate interval has a strictly positive Hoeffding proxy.","url":"../modules/banditrlproof-boundedrewardkernellaw/index.html#decl-a854ff35ed06","parent":"module:BanditRLProof.BoundedRewardKernelLaw","order":3181,"meta":[["Kind","theorem"],["Module","BanditRLProof.BoundedRewardKernelLaw"],["Source","BanditRLProof/BoundedRewardKernelLaw.lean:19"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem intervalVarianceProxy_pos_of_lt {lo hi : Real} (hlohi : lo < hi) : 0 < ((intervalVarianceProxy lo hi : NNReal) : Real)","missing":[],"search":"intervalvarianceproxy_pos_of_lt banditrlproof.concentration.intervalvarianceproxy_pos_of_lt a nondegenerate interval has a strictly positive hoeffding proxy. theorem compiled","shard":"modules/6a357237165f37b3.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.centeredRewardKernelLaw_of_hasSubgaussianMGF","label":"centeredRewardKernelLaw_of_hasSubgaussianMGF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.centeredRewardKernelLaw_of_hasSubgaussianMGF","description":"Exact pointwise means and centered sub-Gaussian MGF witnesses form a centered law for an arbitrary context/action Markov reward kernel.","url":"../modules/banditrlproof-boundedrewardkernellaw/index.html#decl-2e499779c60b","parent":"module:BanditRLProof.BoundedRewardKernelLaw","order":3182,"meta":[["Kind","definition"],["Module","BanditRLProof.BoundedRewardKernelLaw"],["Source","BanditRLProof/BoundedRewardKernelLaw.lean:37"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredRewardKernelLaw_of_hasSubgaussianMGF {Context Action : Type} [MeasurableSpace Context] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hmean : forall context arm, integral (selectedMeasure rewardKernel context arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((mean context arm : Rat) : Real)) (hsubG : forall context arm, HasSubgaussianMGF (fun reward : Rat => (((reward - mean context arm : Rat) : Real))) (varianceProxy context arm) (selectedMeasure rewardKernel context arm)) : CenteredRewardKernelLaw rewardKernel mean varianceProxy where","missing":[],"search":"centeredrewardkernellaw_of_hassubgaussianmgf banditrlproof.rewardkernel.centeredrewardkernellaw_of_hassubgaussianmgf exact pointwise means and centered sub-gaussian mgf witnesses form a centered law for an arbitrary context/action markov reward kernel. definition compiled","shard":"modules/6a357237165f37b3.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.boundedCenteredRewardKernelLaw","label":"boundedCenteredRewardKernelLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.boundedCenteredRewardKernelLaw","description":"Common almost-sure reward bounds and exact pointwise means form a centered law for an arbitrary context/action Markov reward kernel.","url":"../modules/banditrlproof-boundedrewardkernellaw/index.html#decl-4b81ebf0713f","parent":"module:BanditRLProof.BoundedRewardKernelLaw","order":3183,"meta":[["Kind","definition"],["Module","BanditRLProof.BoundedRewardKernelLaw"],["Source","BanditRLProof/BoundedRewardKernelLaw.lean:82"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def boundedCenteredRewardKernelLaw {Context Action : Type} [MeasurableSpace Context] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (mean : Context -> Action -> Rat) (lo hi : Real) (hmeas : forall context arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (selectedMeasure rewardKernel context arm)) (hbound : forall context arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (selectedMeasure rewardKernel context arm))) (hmean : forall context arm, integral (selectedMeasure rewardKernel context arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((mean context arm : Rat) : Real)) : CenteredRewardKernelLaw rewardKernel mean (fun _ _ => Concentration.intervalVarianceProxy lo hi)","missing":[],"search":"boundedcenteredrewardkernellaw banditrlproof.rewardkernel.boundedcenteredrewardkernellaw common almost-sure reward bounds and exact pointwise means form a centered law for an arbitrary context/action markov reward kernel. definition compiled","shard":"modules/6a357237165f37b3.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budgetExhaustionTime","label":"budgetExhaustionTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Budget.budgetExhaustionTime","description":"First time an accumulated `Nat` resource process reaches a budget. This is a project-local name for Mathlib's `hittingAfter` specialized to the upper set `{spent >= budget}` and start time `0`.","url":"../modules/banditrlproof-budgetstoppingtime/index.html#decl-68ea54b14dfd","parent":"module:BanditRLProof.BudgetStoppingTime","order":3184,"meta":[["Kind","definition"],["Module","BanditRLProof.BudgetStoppingTime"],["Source","BanditRLProof/BudgetStoppingTime.lean:24"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def budgetExhaustionTime {Omega : Type u} (spent : Nat -> Omega -> Nat) (budget : Nat) : Omega -> WithTop Nat","missing":[],"search":"budgetexhaustiontime banditrlproof.budget.budgetexhaustiontime first time an accumulated `nat` resource process reaches a budget. this is a project-local name for mathlib's `hittingafter` specialized to the upper set `{spent >= budget}` and start time `0`. definition compiled","shard":"modules/9e24be65f9e8509f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.isStoppingTime_budgetExhaustionTime_of_adapted","label":"isStoppingTime_budgetExhaustionTime_of_adapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.isStoppingTime_budgetExhaustionTime_of_adapted","description":"An adapted accumulated-resource process has a budget-exhaustion stopping time. This is the `STOPPING-TIME-BUDGET` wrapper over `MeasureTheory.Adapted.isStoppingTime_hittingAfter`.","url":"../modules/banditrlproof-budgetstoppingtime/index.html#decl-5be9cb9dc135","parent":"module:BanditRLProof.BudgetStoppingTime","order":3185,"meta":[["Kind","theorem"],["Module","BanditRLProof.BudgetStoppingTime"],["Source","BanditRLProof/BudgetStoppingTime.lean:36"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem isStoppingTime_budgetExhaustionTime_of_adapted {Omega : Type u} [mOmega : MeasurableSpace Omega] {F : Filtration Nat mOmega} {spent : Nat -> Omega -> Nat} (budget : Nat) (hspent : Adapted F spent) : IsStoppingTime F (budgetExhaustionTime spent budget)","missing":[],"search":"isstoppingtime_budgetexhaustiontime_of_adapted banditrlproof.budget.isstoppingtime_budgetexhaustiontime_of_adapted an adapted accumulated-resource process has a budget-exhaustion stopping time. this is the `stopping-time-budget` wrapper over `measuretheory.adapted.isstoppingtime_hittingafter`. theorem compiled","shard":"modules/9e24be65f9e8509f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.measurableSet_budgetExhaustionTime_le_of_adapted","label":"measurableSet_budgetExhaustionTime_le_of_adapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.measurableSet_budgetExhaustionTime_le_of_adapted","description":"At each horizon `n`, the event that the budget has already been exhausted is measurable at filtration level `n`.","url":"../modules/banditrlproof-budgetstoppingtime/index.html#decl-1a1e2e28c619","parent":"module:BanditRLProof.BudgetStoppingTime","order":3186,"meta":[["Kind","theorem"],["Module","BanditRLProof.BudgetStoppingTime"],["Source","BanditRLProof/BudgetStoppingTime.lean:49"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_budgetExhaustionTime_le_of_adapted {Omega : Type u} [mOmega : MeasurableSpace Omega] {F : Filtration Nat mOmega} {spent : Nat -> Omega -> Nat} (budget n : Nat) (hspent : Adapted F spent) : MeasurableSet[F n] {omega | budgetExhaustionTime spent budget omega <= n}","missing":[],"search":"measurableset_budgetexhaustiontime_le_of_adapted banditrlproof.budget.measurableset_budgetexhaustiontime_le_of_adapted at each horizon `n`, the event that the budget has already been exhausted is measurable at filtration level `n`. theorem compiled","shard":"modules/9e24be65f9e8509f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.cappedOccupancyTail","label":"cappedOccupancyTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.cappedOccupancyTail","description":"def cappedOccupancyTail (a ε t : ℝ) : ℝ","url":"../modules/banditrlproof-concentrationcappedoccupancy/index.html#decl-66e537be39bf","parent":"module:BanditRLProof.ConcentrationCappedOccupancy","order":3187,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationCappedOccupancy"],["Source","BanditRLProof/ConcentrationCappedOccupancy.lean:7"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def cappedOccupancyTail (a ε t : ℝ) : ℝ","missing":[],"search":"cappedoccupancytail banditrlproof.concentration.cappedoccupancytail def cappedoccupancytail (a ε t : ℝ) : ℝ definition compiled","shard":"modules/06349c79cd6a2f4d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.cappedOccupancyTail_nonneg","label":"cappedOccupancyTail_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.cappedOccupancyTail_nonneg","description":"theorem cappedOccupancyTail_nonneg (a ε t : ℝ) : 0 ≤ cappedOccupancyTail a ε t","url":"../modules/banditrlproof-concentrationcappedoccupancy/index.html#decl-9a6f88c158aa","parent":"module:BanditRLProof.ConcentrationCappedOccupancy","order":3188,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationCappedOccupancy"],["Source","BanditRLProof/ConcentrationCappedOccupancy.lean:10"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem cappedOccupancyTail_nonneg (a ε t : ℝ) : 0 ≤ cappedOccupancyTail a ε t","missing":[],"search":"cappedoccupancytail_nonneg banditrlproof.concentration.cappedoccupancytail_nonneg theorem cappedoccupancytail_nonneg (a ε t : ℝ) : 0 ≤ cappedoccupancytail a ε t theorem compiled","shard":"modules/06349c79cd6a2f4d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.cappedOccupancyTail_antitone","label":"cappedOccupancyTail_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.cappedOccupancyTail_antitone","description":"theorem cappedOccupancyTail_antitone (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : Antitone (cappedOccupancyTail a ε)","url":"../modules/banditrlproof-concentrationcappedoccupancy/index.html#decl-6afe047b1034","parent":"module:BanditRLProof.ConcentrationCappedOccupancy","order":3189,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationCappedOccupancy"],["Source","BanditRLProof/ConcentrationCappedOccupancy.lean:14"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem cappedOccupancyTail_antitone (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : Antitone (cappedOccupancyTail a ε)","missing":[],"search":"cappedoccupancytail_antitone banditrlproof.concentration.cappedoccupancytail_antitone theorem cappedoccupancytail_antitone (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : antitone (cappedoccupancytail a ε) theorem compiled","shard":"modules/06349c79cd6a2f4d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_cappedOccupancyTail","label":"integral_cappedOccupancyTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_cappedOccupancyTail","description":"theorem integral_cappedOccupancyTail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : IntegrableOn (cappedOccupancyTail a ε) (Ioi 0) ∧ (∫ t in Ioi 0, cappedOccupancyTail a ε t) = (2/ε^2)*(a+sqrt (Real.pi*a)+1)","url":"../modules/banditrlproof-concentrationcappedoccupancy/index.html#decl-76ad3e7e7786","parent":"module:BanditRLProof.ConcentrationCappedOccupancy","order":3190,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationCappedOccupancy"],["Source","BanditRLProof/ConcentrationCappedOccupancy.lean:28"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_cappedOccupancyTail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : IntegrableOn (cappedOccupancyTail a ε) (Ioi 0) ∧ (∫ t in Ioi 0, cappedOccupancyTail a ε t) = (2/ε^2)*(a+sqrt (Real.pi*a)+1)","missing":[],"search":"integral_cappedoccupancytail banditrlproof.concentration.integral_cappedoccupancytail theorem integral_cappedoccupancytail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : integrableon (cappedoccupancytail a ε) (ioi 0) ∧ (∫ t in ioi 0, cappedoccupancytail a ε t) = (2/ε^2)*(a+sqrt (real.pi*a)+1) theorem compiled","shard":"modules/06349c79cd6a2f4d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_le_occupancy_bound_sharp","label":"sum_le_occupancy_bound_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_le_occupancy_bound_sharp","description":"theorem sum_le_occupancy_bound_sharp (p : ℕ → ℝ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (h1 : ∀ s, p s ≤ 1) (htail : ∀ s : ℕ, 2*a/ε^2 < (s : ℝ) → p s ≤ occupancyTail a ε s) (n : ℕ) : (∑ i ∈ Finset.range n, p (i+1)) ≤ (2/ε^2)*(a+sqrt (Real.pi*a)+1)","url":"../modules/banditrlproof-concentrationcappedoccupancy/index.html#decl-d9e10893dc08","parent":"module:BanditRLProof.ConcentrationCappedOccupancy","order":3191,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationCappedOccupancy"],["Source","BanditRLProof/ConcentrationCappedOccupancy.lean:54"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_le_occupancy_bound_sharp (p : ℕ → ℝ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (h1 : ∀ s, p s ≤ 1) (htail : ∀ s : ℕ, 2*a/ε^2 < (s : ℝ) → p s ≤ occupancyTail a ε s) (n : ℕ) : (∑ i ∈ Finset.range n, p (i+1)) ≤ (2/ε^2)*(a+sqrt (Real.pi*a)+1)","missing":[],"search":"sum_le_occupancy_bound_sharp banditrlproof.concentration.sum_le_occupancy_bound_sharp theorem sum_le_occupancy_bound_sharp (p : ℕ → ℝ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (h1 : ∀ s, p s ≤ 1) (htail : ∀ s : ℕ, 2*a/ε^2 < (s : ℝ) → p s ≤ occupancytail a ε s) (n : ℕ) : (∑ i ∈ finset.range n, p (i+1)) ≤ (2/ε^2)*(a+sqrt (real.pi*a)+1) theorem compiled","shard":"modules/06349c79cd6a2f4d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasCondMGFUpperBoundAt_of_condExp_le","label":"hasCondMGFUpperBoundAt_of_condExp_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasCondMGFUpperBoundAt_of_condExp_le","description":"theorem hasCondMGFUpperBoundAt_of_condExp_le {Ω : Type*} {mΩ : MeasurableSpace Ω} [StandardBorelSpace Ω] {μ : Measure Ω} [IsFiniteMeasure μ] (m : MeasurableSpace Ω) (hm : m ≤ mΩ) (X : Ω → ℝ) (t ψ : ℝ) (hi : ∀ s, Integrable (fun ω => Real.exp (s * X ω)) μ) (hc : μ[fun ω => Real.exp (t * X ω) | m] ≤ᵐ[μ] fun _ => Real.exp ψ) : HasCondMGFUpperBoundAt m hm X t ψ μ","url":"../modules/banditrlproof-concentrationconditionalmgf/index.html#decl-82e335e35987","parent":"module:BanditRLProof.ConcentrationConditionalMGF","order":3192,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConditionalMGF"],["Source","BanditRLProof/ConcentrationConditionalMGF.lean:8"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasCondMGFUpperBoundAt_of_condExp_le {Ω : Type*} {mΩ : MeasurableSpace Ω} [StandardBorelSpace Ω] {μ : Measure Ω} [IsFiniteMeasure μ] (m : MeasurableSpace Ω) (hm : m ≤ mΩ) (X : Ω → ℝ) (t ψ : ℝ) (hi : ∀ s, Integrable (fun ω => Real.exp (s * X ω)) μ) (hc : μ[fun ω => Real.exp (t * X ω) | m] ≤ᵐ[μ] fun _ => Real.exp ψ) : HasCondMGFUpperBoundAt m hm X t ψ μ","missing":[],"search":"hascondmgfupperboundat_of_condexp_le banditrlproof.concentration.hascondmgfupperboundat_of_condexp_le theorem hascondmgfupperboundat_of_condexp_le {ω : type*} {mω : measurablespace ω} [standardborelspace ω] {μ : measure ω} [isfinitemeasure μ] (m : measurablespace ω) (hm : m ≤ mω) (x : ω → ℝ) (t ψ : ℝ) (hi : ∀ s, integrable (fun ω => real.exp (s * x ω)) μ) (hc : μ[fun ω => real.exp (t * x ω) | m] ≤ᵐ[μ] fun _ => real.exp ψ) : hascondmgfupperboundat m hm x t ψ μ theorem compiled","shard":"modules/77b5898376537822.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.geometricConfidenceShare","label":"geometricConfidenceShare","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.geometricConfidenceShare","description":"The geometric confidence schedule `delta / 2 / 2^n`.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-8f211831a984","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3193,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:15"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def geometricConfidenceShare (delta : Real) (n : Nat) : Real","missing":[],"search":"geometricconfidenceshare banditrlproof.concentration.geometricconfidenceshare the geometric confidence schedule `delta / 2 / 2^n`. definition compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.geometricConfidenceShare_pos","label":"geometricConfidenceShare_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.geometricConfidenceShare_pos","description":"Every geometric confidence share is positive when its outer budget is.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-5b2612a49961","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3194,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:20"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem geometricConfidenceShare_pos {delta : Real} (hdelta : 0 < delta) (n : Nat) : 0 < geometricConfidenceShare delta n","missing":[],"search":"geometricconfidenceshare_pos banditrlproof.concentration.geometricconfidenceshare_pos every geometric confidence share is positive when its outer budget is. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.tsum_ofReal_geometricConfidenceShare","label":"tsum_ofReal_geometricConfidenceShare","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.tsum_ofReal_geometricConfidenceShare","description":"The ENNReal masses of the geometric confidence shares sum exactly to the outer budget.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-2e1ffa589537","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3195,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:28"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_ofReal_geometricConfidenceShare {delta : Real} (hdelta : 0 <= delta) : ∑' n, ENNReal.ofReal (geometricConfidenceShare delta n) = ENNReal.ofReal delta","missing":[],"search":"tsum_ofreal_geometricconfidenceshare banditrlproof.concentration.tsum_ofreal_geometricconfidenceshare the ennreal masses of the geometric confidence shares sum exactly to the outer budget. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.telescopingConfidenceWeight","label":"telescopingConfidenceWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.telescopingConfidenceWeight","description":"Positive telescoping weight `1 / ((n+1)(n+2))`. Unlike the geometric weight, its reciprocal grows only polynomially in time.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-b28213637f0e","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3196,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:57"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingConfidenceWeight (n : Nat) : Real","missing":[],"search":"telescopingconfidenceweight banditrlproof.concentration.telescopingconfidenceweight positive telescoping weight `1 / ((n+1)(n+2))`. unlike the geometric weight, its reciprocal grows only polynomially in time. definition compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.telescopingConfidenceWeight_eq_sub","label":"telescopingConfidenceWeight_eq_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.telescopingConfidenceWeight_eq_sub","description":"The telescoping weight is a difference of consecutive reciprocals.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-da9dcc8b883f","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3197,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:61"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem telescopingConfidenceWeight_eq_sub (n : Nat) : telescopingConfidenceWeight n = 1 / (((n + 1 : Nat) : Real)) - 1 / (((n + 2 : Nat) : Real))","missing":[],"search":"telescopingconfidenceweight_eq_sub banditrlproof.concentration.telescopingconfidenceweight_eq_sub the telescoping weight is a difference of consecutive reciprocals. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_range_telescopingConfidenceWeight","label":"sum_range_telescopingConfidenceWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_range_telescopingConfidenceWeight","description":"Exact finite partial sum of the telescoping weights.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-a6994a167ab2","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3198,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:70"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_range_telescopingConfidenceWeight (n : Nat) : (Finset.range n).sum telescopingConfidenceWeight = 1 - 1 / (((n + 1 : Nat) : Real))","missing":[],"search":"sum_range_telescopingconfidenceweight banditrlproof.concentration.sum_range_telescopingconfidenceweight exact finite partial sum of the telescoping weights. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.telescopingConfidenceWeight_nonneg","label":"telescopingConfidenceWeight_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.telescopingConfidenceWeight_nonneg","description":"Every telescoping confidence weight is nonnegative.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-09508307e7ff","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3199,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:79"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem telescopingConfidenceWeight_nonneg (n : Nat) : 0 <= telescopingConfidenceWeight n","missing":[],"search":"telescopingconfidenceweight_nonneg banditrlproof.concentration.telescopingconfidenceweight_nonneg every telescoping confidence weight is nonnegative. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasSum_telescopingConfidenceWeight","label":"hasSum_telescopingConfidenceWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasSum_telescopingConfidenceWeight","description":"The telescoping weights sum exactly to one.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-ba416cffa207","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3200,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:85"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasSum_telescopingConfidenceWeight : HasSum telescopingConfidenceWeight 1","missing":[],"search":"hassum_telescopingconfidenceweight banditrlproof.concentration.hassum_telescopingconfidenceweight the telescoping weights sum exactly to one. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.telescopingConfidenceShare","label":"telescopingConfidenceShare","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.telescopingConfidenceShare","description":"Time-`n` confidence share `delta / ((n+1)(n+2))`.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-76668d6daa79","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3201,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:103"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingConfidenceShare (delta : Real) (n : Nat) : Real","missing":[],"search":"telescopingconfidenceshare banditrlproof.concentration.telescopingconfidenceshare time-`n` confidence share `delta / ((n+1)(n+2))`. definition compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.telescopingConfidenceShare_eq_div","label":"telescopingConfidenceShare_eq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.telescopingConfidenceShare_eq_div","description":"Display the telescoping share as the intended quotient.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-3a94da06c288","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3202,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:108"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem telescopingConfidenceShare_eq_div (delta : Real) (n : Nat) : telescopingConfidenceShare delta n = delta / (((n + 1 : Nat) : Real) * ((n + 2 : Nat) : Real))","missing":[],"search":"telescopingconfidenceshare_eq_div banditrlproof.concentration.telescopingconfidenceshare_eq_div display the telescoping share as the intended quotient. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.telescopingConfidenceShare_pos","label":"telescopingConfidenceShare_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.telescopingConfidenceShare_pos","description":"Every telescoping confidence share is positive when its outer budget is.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-5edf51aa5567","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3203,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:116"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem telescopingConfidenceShare_pos {delta : Real} (hdelta : 0 < delta) (n : Nat) : 0 < telescopingConfidenceShare delta n","missing":[],"search":"telescopingconfidenceshare_pos banditrlproof.concentration.telescopingconfidenceshare_pos every telescoping confidence share is positive when its outer budget is. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.tsum_ofReal_telescopingConfidenceShare","label":"tsum_ofReal_telescopingConfidenceShare","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.tsum_ofReal_telescopingConfidenceShare","description":"The ENNReal masses of the telescoping confidence shares sum exactly to the outer budget.","url":"../modules/banditrlproof-concentrationconfidenceschedule/index.html#decl-22ecca312ba8","parent":"module:BanditRLProof.ConcentrationConfidenceSchedule","order":3204,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationConfidenceSchedule"],["Source","BanditRLProof/ConcentrationConfidenceSchedule.lean:124"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_ofReal_telescopingConfidenceShare {delta : Real} (hdelta : 0 <= delta) : ∑' n, ENNReal.ofReal (telescopingConfidenceShare delta n) = ENNReal.ofReal delta","missing":[],"search":"tsum_ofreal_telescopingconfidenceshare banditrlproof.concentration.tsum_ofreal_telescopingconfidenceshare the ennreal masses of the telescoping confidence shares sum exactly to the outer budget. theorem compiled","shard":"modules/2f4ee0c9ee5b49e6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.mul_exp_neg_le_exp_difference","label":"mul_exp_neg_le_exp_difference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.mul_exp_neg_le_exp_difference","description":"A telescoping majorant obtained from `x/3 ≤ sinh(x/3)`.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html#decl-93c2bbf75d30","parent":"module:BanditRLProof.ConcentrationDyadicExponential","order":3205,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationDyadicExponential"],["Source","BanditRLProof/ConcentrationDyadicExponential.lean:17"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem mul_exp_neg_le_exp_difference (x : ℝ) (hx : 0 ≤ x) : x * exp (-x) ≤ (3 / 2 : ℝ) * (exp (-(2 * x / 3)) - exp (-(4 * x / 3)))","missing":[],"search":"mul_exp_neg_le_exp_difference banditrlproof.concentration.mul_exp_neg_le_exp_difference a telescoping majorant obtained from `x/3 ≤ sinh(x/3)`. theorem compiled","shard":"modules/0ed469de9bc35975.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_dyadic_mul_exp_neg_le","label":"sum_dyadic_mul_exp_neg_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_dyadic_mul_exp_neg_le","description":"Finite dyadic exponential sum, with the remaining terminal mass retained.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html#decl-173ff941119e","parent":"module:BanditRLProof.ConcentrationDyadicExponential","order":3206,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationDyadicExponential"],["Source","BanditRLProof/ConcentrationDyadicExponential.lean:37"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_dyadic_mul_exp_neg_le (a : ℝ) (ha : 0 < a) (N : ℕ) : ∑ j ∈ range N, (2 : ℝ) ^ j * exp (-(a * 2 ^ j)) ≤ 3 / (2 * a) * (exp (-(2 * a / 3)) - exp (-(2 * (a * 2 ^ N) / 3)))","missing":[],"search":"sum_dyadic_mul_exp_neg_le banditrlproof.concentration.sum_dyadic_mul_exp_neg_le finite dyadic exponential sum, with the remaining terminal mass retained. theorem compiled","shard":"modules/0ed469de9bc35975.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_dyadic_mul_exp_neg_le_three_div_two","label":"sum_dyadic_mul_exp_neg_le_three_div_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_dyadic_mul_exp_neg_le_three_div_two","description":"Uniform finite-prefix bound; no logarithm or integral comparison loss.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html#decl-7292b593af1c","parent":"module:BanditRLProof.ConcentrationDyadicExponential","order":3207,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationDyadicExponential"],["Source","BanditRLProof/ConcentrationDyadicExponential.lean:58"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_dyadic_mul_exp_neg_le_three_div_two (a : ℝ) (ha : 0 < a) (N : ℕ) : ∑ j ∈ range N, (2 : ℝ) ^ j * exp (-(a * 2 ^ j)) ≤ 3 / (2 * a)","missing":[],"search":"sum_dyadic_mul_exp_neg_le_three_div_two banditrlproof.concentration.sum_dyadic_mul_exp_neg_le_three_div_two uniform finite-prefix bound; no logarithm or integral comparison loss. theorem compiled","shard":"modules/0ed469de9bc35975.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.tsum_dyadic_mul_exp_neg_le","label":"tsum_dyadic_mul_exp_neg_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.tsum_dyadic_mul_exp_neg_le","description":"Countable dyadic sum in the probability-friendly extended nonnegative reals.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html#decl-35c1476a122f","parent":"module:BanditRLProof.ConcentrationDyadicExponential","order":3208,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationDyadicExponential"],["Source","BanditRLProof/ConcentrationDyadicExponential.lean:67"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_dyadic_mul_exp_neg_le (a : ℝ) (ha : 0 < a) : (∑' j : ℕ, ENNReal.ofReal ((2 : ℝ) ^ j * exp (-(a * 2 ^ j)))) ≤ ENNReal.ofReal (3 / (2 * a))","missing":[],"search":"tsum_dyadic_mul_exp_neg_le banditrlproof.concentration.tsum_dyadic_mul_exp_neg_le countable dyadic sum in the probability-friendly extended nonnegative reals. theorem compiled","shard":"modules/0ed469de9bc35975.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_moss_peeling_exponential_le_twelve","label":"sum_moss_peeling_exponential_le_twelve","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_moss_peeling_exponential_le_twelve","description":"The geometric series in source Lemma 9.3 is at most 12 delta/gap^2.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html#decl-2b689d19aca8","parent":"module:BanditRLProof.ConcentrationDyadicExponential","order":3209,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationDyadicExponential"],["Source","BanditRLProof/ConcentrationDyadicExponential.lean:76"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_moss_peeling_exponential_le_twelve (δ gap : ℝ) (hδ : 0 ≤ δ) (hgap : 0 < gap) (N : ℕ) : ∑ j ∈ range N, δ * (2 : ℝ) ^ (j+1) * exp (-(gap ^ 2 / 4 * 2 ^ j)) ≤ 12 * δ / gap ^ 2","missing":[],"search":"sum_moss_peeling_exponential_le_twelve banditrlproof.concentration.sum_moss_peeling_exponential_le_twelve the geometric series in source lemma 9.3 is at most 12 delta/gap^2. theorem compiled","shard":"modules/0ed469de9bc35975.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.tsum_moss_peeling_exponential_le_fifteen","label":"tsum_moss_peeling_exponential_le_fifteen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.tsum_moss_peeling_exponential_le_fifteen","description":"Source-constant countable series bound, with no weakened constant.","url":"../modules/banditrlproof-concentrationdyadicexponential/index.html#decl-311f0f69b832","parent":"module:BanditRLProof.ConcentrationDyadicExponential","order":3210,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationDyadicExponential"],["Source","BanditRLProof/ConcentrationDyadicExponential.lean:95"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_moss_peeling_exponential_le_fifteen (δ gap : ℝ) (hδ : 0 ≤ δ) (hgap : 0 < gap) : (∑' j : ℕ, ENNReal.ofReal (δ * (2 : ℝ) ^ (j+1) * exp (-(gap ^ 2 / 4 * 2 ^ j)))) ≤ ENNReal.ofReal (15 * δ / gap ^ 2)","missing":[],"search":"tsum_moss_peeling_exponential_le_fifteen banditrlproof.concentration.tsum_moss_peeling_exponential_le_fifteen source-constant countable series bound, with no weakened constant. theorem compiled","shard":"modules/0ed469de9bc35975.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_tsum_of_uniform","label":"measure_iUnion_iUnion_fintype_le_tsum_of_uniform","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_tsum_of_uniform","description":"A countable family of finite-index bad events is controlled by the sum of its per-time budgets when every index receives an equal share. No event measurability or probability-measure assumption is required.","url":"../modules/banditrlproof-concentrationfintypegeometricalltime/index.html#decl-1479880bd324","parent":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","order":3211,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFintypeGeometricAllTime"],["Source","BanditRLProof/ConcentrationFintypeGeometricAllTime.lean:24"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_iUnion_iUnion_fintype_le_tsum_of_uniform {Omega : Type u} {Idx : Type v} [MeasurableSpace Omega] [Fintype Idx] [Nonempty Idx] (mu : Measure Omega) (bad : Nat -> Idx -> Set Omega) (deltaAt : Nat -> Real) (hbad : forall n i, mu (bad n i) <= ENNReal.ofReal (deltaAt n / (Fintype.card Idx : Real))) : mu (⋃ n, ⋃ i, bad n i) <= ∑' n, ENNReal.ofReal (deltaAt n)","missing":[],"search":"measure_iunion_iunion_fintype_le_tsum_of_uniform banditrlproof.concentration.measure_iunion_iunion_fintype_le_tsum_of_uniform a countable family of finite-index bad events is controlled by the sum of its per-time budgets when every index receives an equal share. no event measurability or probability-measure assumption is required. theorem compiled","shard":"modules/a86ff861bccc9795.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","label":"measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","description":"Equal sharing across a nonempty finite index type at every time, followed by the geometric schedule over time, gives one all-time outer confidence budget. No event measurability or probability-measure assumption is required.","url":"../modules/banditrlproof-concentrationfintypegeometricalltime/index.html#decl-761e6a46f5e6","parent":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","order":3212,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFintypeGeometricAllTime"],["Source","BanditRLProof/ConcentrationFintypeGeometricAllTime.lean:50"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare {Omega : Type u} {Idx : Type v} [MeasurableSpace Omega] [Fintype Idx] [Nonempty Idx] (mu : Measure Omega) (bad : Nat -> Idx -> Set Omega) (delta : Real) (hdelta : 0 <= delta) (hbad : forall n i, mu (bad n i) <= ENNReal.ofReal (geometricConfidenceShare delta n / (Fintype.card Idx : Real))) : mu (⋃ n, ⋃ i, bad n i) <= ENNReal.ofReal delta","missing":[],"search":"measure_iunion_iunion_fintype_le_delta_of_geometricconfidenceshare banditrlproof.concentration.measure_iunion_iunion_fintype_le_delta_of_geometricconfidenceshare equal sharing across a nonempty finite index type at every time, followed by the geometric schedule over time, gives one all-time outer confidence budget. no event measurability or probability-measure assumption is required. theorem compiled","shard":"modules/a86ff861bccc9795.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare","label":"measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare","description":"Equal per-index telescoping shares at every time compose to the original outer confidence budget.","url":"../modules/banditrlproof-concentrationfintypetelescopingalltime/index.html#decl-3316b8632e74","parent":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","order":3213,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFintypeTelescopingAllTime"],["Source","BanditRLProof/ConcentrationFintypeTelescopingAllTime.lean:22"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare {Omega : Type u} {Idx : Type v} [MeasurableSpace Omega] [Fintype Idx] [Nonempty Idx] (mu : Measure Omega) (bad : Nat -> Idx -> Set Omega) (delta : Real) (hdelta : 0 <= delta) (hbad : forall n i, mu (bad n i) <= ENNReal.ofReal (telescopingConfidenceShare delta n / (Fintype.card Idx : Real))) : mu (⋃ n, ⋃ i, bad n i) <= ENNReal.ofReal delta","missing":[],"search":"measure_iunion_iunion_fintype_le_delta_of_telescopingconfidenceshare banditrlproof.concentration.measure_iunion_iunion_fintype_le_delta_of_telescopingconfidenceshare equal per-index telescoping shares at every time compose to the original outer confidence budget. theorem compiled","shard":"modules/cde39c638d622b2e.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt","label":"HasMGFUpperBoundAt","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt","description":"A kernel-valued MGF upper bound at one fixed tilt. Integrability is required at every tilt so that successive conditional laws can be composed without adding boundedness assumptions.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-e1361ff5ee80","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3214,"meta":[["Kind","structure"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:25"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure HasMGFUpperBoundAt (X : Ω → ℝ) (t ψ : ℝ) (κ : ProbabilityTheory.Kernel Ω' Ω) (ν : Measure Ω' := by volume_tac) : Prop where","missing":[],"search":"hasmgfupperboundat banditrlproof.concentration.kernel.hasmgfupperboundat a kernel-valued mgf upper bound at one fixed tilt. integrability is required at every tilt so that successive conditional laws can be composed without adding boundedness assumptions. structure compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.ae_integrable_exp_mul","label":"ae_integrable_exp_mul","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.ae_integrable_exp_mul","description":"lemma ae_integrable_exp_mul (h : HasMGFUpperBoundAt X t ψ κ ν) (s : ℝ) : ∀ᵐ ω' ∂ν, Integrable (fun y ↦ exp (s * X y)) (κ ω')","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-f92432b5d5e1","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3215,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:32"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma ae_integrable_exp_mul (h : HasMGFUpperBoundAt X t ψ κ ν) (s : ℝ) : ∀ᵐ ω' ∂ν, Integrable (fun y ↦ exp (s * X y)) (κ ω')","missing":[],"search":"ae_integrable_exp_mul banditrlproof.concentration.kernel.hasmgfupperboundat.ae_integrable_exp_mul lemma ae_integrable_exp_mul (h : hasmgfupperboundat x t ψ κ ν) (s : ℝ) : ∀ᵐ ω' ∂ν, integrable (fun y ↦ exp (s * x y)) (κ ω') lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.ae_forall_integrable_exp_mul","label":"ae_forall_integrable_exp_mul","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.ae_forall_integrable_exp_mul","description":"lemma ae_forall_integrable_exp_mul (h : HasMGFUpperBoundAt X t ψ κ ν) : ∀ᵐ ω' ∂ν, ∀ s, Integrable (fun ω ↦ exp (s * X ω)) (κ ω')","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-57d742b0e56b","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3216,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:36"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma ae_forall_integrable_exp_mul (h : HasMGFUpperBoundAt X t ψ κ ν) : ∀ᵐ ω' ∂ν, ∀ s, Integrable (fun ω ↦ exp (s * X ω)) (κ ω')","missing":[],"search":"ae_forall_integrable_exp_mul banditrlproof.concentration.kernel.hasmgfupperboundat.ae_forall_integrable_exp_mul lemma ae_forall_integrable_exp_mul (h : hasmgfupperboundat x t ψ κ ν) : ∀ᵐ ω' ∂ν, ∀ s, integrable (fun ω ↦ exp (s * x ω)) (κ ω') lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.congr","label":"congr","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.congr","description":"lemma congr {Y : Ω -> Real} (h : HasMGFUpperBoundAt X t ψ κ ν) (hXY : X =ᵐ[κ ∘ₘ ν] Y) : HasMGFUpperBoundAt Y t ψ κ ν where","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-9709ee03098a","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3217,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:45"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma congr {Y : Ω -> Real} (h : HasMGFUpperBoundAt X t ψ κ ν) (hXY : X =ᵐ[κ ∘ₘ ν] Y) : HasMGFUpperBoundAt Y t ψ κ ν where","missing":[],"search":"congr banditrlproof.concentration.kernel.hasmgfupperboundat.congr lemma congr {y : ω -> real} (h : hasmgfupperboundat x t ψ κ ν) (hxy : x =ᵐ[κ ∘ₘ ν] y) : hasmgfupperboundat y t ψ κ ν where lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.congr_iff","label":"congr_iff","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.congr_iff","description":"lemma congr_iff {Y : Ω -> Real} (hXY : X =ᵐ[κ ∘ₘ ν] Y) : HasMGFUpperBoundAt X t ψ κ ν ↔ HasMGFUpperBoundAt Y t ψ κ ν","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-0932f10b4538","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3218,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:57"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma congr_iff {Y : Ω -> Real} (hXY : X =ᵐ[κ ∘ₘ ν] Y) : HasMGFUpperBoundAt X t ψ κ ν ↔ HasMGFUpperBoundAt Y t ψ κ ν","missing":[],"search":"congr_iff banditrlproof.concentration.kernel.hasmgfupperboundat.congr_iff lemma congr_iff {y : ω -> real} (hxy : x =ᵐ[κ ∘ₘ ν] y) : hasmgfupperboundat x t ψ κ ν ↔ hasmgfupperboundat y t ψ κ ν lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.memLp_exp_mul","label":"memLp_exp_mul","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.memLp_exp_mul","description":"lemma memLp_exp_mul (h : HasMGFUpperBoundAt X t ψ κ ν) (s : ℝ) (p : ℝ≥0) : MemLp (fun ω ↦ exp (s * X ω)) p (κ ∘ₘ ν)","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-bcf836687a2a","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3219,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:61"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma memLp_exp_mul (h : HasMGFUpperBoundAt X t ψ κ ν) (s : ℝ) (p : ℝ≥0) : MemLp (fun ω ↦ exp (s * X ω)) p (κ ∘ₘ ν)","missing":[],"search":"memlp_exp_mul banditrlproof.concentration.kernel.hasmgfupperboundat.memlp_exp_mul lemma memlp_exp_mul (h : hasmgfupperboundat x t ψ κ ν) (s : ℝ) (p : ℝ≥0) : memlp (fun ω ↦ exp (s * x ω)) p (κ ∘ₘ ν) lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.zero_kernel","label":"zero_kernel","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.zero_kernel","description":"lemma zero_kernel : HasMGFUpperBoundAt X t ψ (0 : ProbabilityTheory.Kernel Ω' Ω) ν","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-7dccbbe035e3","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3220,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:77"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma zero_kernel : HasMGFUpperBoundAt X t ψ (0 : ProbabilityTheory.Kernel Ω' Ω) ν","missing":[],"search":"zero_kernel banditrlproof.concentration.kernel.hasmgfupperboundat.zero_kernel lemma zero_kernel : hasmgfupperboundat x t ψ (0 : probabilitytheory.kernel ω' ω) ν lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.zero_measure","label":"zero_measure","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.zero_measure","description":"lemma zero_measure : HasMGFUpperBoundAt X t ψ κ (0 : Measure Ω')","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-9794643217bd","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3221,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:83"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma zero_measure : HasMGFUpperBoundAt X t ψ κ (0 : Measure Ω')","missing":[],"search":"zero_measure banditrlproof.concentration.kernel.hasmgfupperboundat.zero_measure lemma zero_measure : hasmgfupperboundat x t ψ κ (0 : measure ω') lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.prodMkLeft_compProd","label":"prodMkLeft_compProd","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.prodMkLeft_compProd","description":"lemma prodMkLeft_compProd {η : ProbabilityTheory.Kernel Ω Ω''} (h : HasMGFUpperBoundAt Y t ψY η (κ ∘ₘ ν)) : HasMGFUpperBoundAt Y t ψY (ProbabilityTheory.Kernel.prodMkLeft Ω' η) (ν ⊗ₘ κ)","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-6c7064c6aacf","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3222,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:90"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma prodMkLeft_compProd {η : ProbabilityTheory.Kernel Ω Ω''} (h : HasMGFUpperBoundAt Y t ψY η (κ ∘ₘ ν)) : HasMGFUpperBoundAt Y t ψY (ProbabilityTheory.Kernel.prodMkLeft Ω' η) (ν ⊗ₘ κ)","missing":[],"search":"prodmkleft_compprod banditrlproof.concentration.kernel.hasmgfupperboundat.prodmkleft_compprod lemma prodmkleft_compprod {η : probabilitytheory.kernel ω ω''} (h : hasmgfupperboundat y t ψy η (κ ∘ₘ ν)) : hasmgfupperboundat y t ψy (probabilitytheory.kernel.prodmkleft ω' η) (ν ⊗ₘ κ) lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.integrable_exp_add_compProd","label":"integrable_exp_add_compProd","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.integrable_exp_add_compProd","description":"lemma integrable_exp_add_compProd {η : ProbabilityTheory.Kernel (Ω' × Ω) Ω''} [ProbabilityTheory.IsZeroOrMarkovKernel η] (hX : HasMGFUpperBoundAt X t ψ κ ν) (hY : HasMGFUpperBoundAt Y t ψY η (ν ⊗ₘ κ)) (s : ℝ) : Integrable (fun ω ↦ exp (s * (X ω.1 + Y ω.2))) ((κ ⊗ₖ η) ∘ₘ ν)","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-7cafc8ee3d84","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3223,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:105"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma integrable_exp_add_compProd {η : ProbabilityTheory.Kernel (Ω' × Ω) Ω''} [ProbabilityTheory.IsZeroOrMarkovKernel η] (hX : HasMGFUpperBoundAt X t ψ κ ν) (hY : HasMGFUpperBoundAt Y t ψY η (ν ⊗ₘ κ)) (s : ℝ) : Integrable (fun ω ↦ exp (s * (X ω.1 + Y ω.2))) ((κ ⊗ₖ η) ∘ₘ ν)","missing":[],"search":"integrable_exp_add_compprod banditrlproof.concentration.kernel.hasmgfupperboundat.integrable_exp_add_compprod lemma integrable_exp_add_compprod {η : probabilitytheory.kernel (ω' × ω) ω''} [probabilitytheory.iszeroormarkovkernel η] (hx : hasmgfupperboundat x t ψ κ ν) (hy : hasmgfupperboundat y t ψy η (ν ⊗ₘ κ)) (s : ℝ) : integrable (fun ω ↦ exp (s * (x ω.1 + y ω.2))) ((κ ⊗ₖ η) ∘ₘ ν) lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.add_compProd","label":"add_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.add_compProd","description":"theorem add_compProd {η : ProbabilityTheory.Kernel (Ω' × Ω) Ω''} [ProbabilityTheory.IsZeroOrMarkovKernel η] (hX : HasMGFUpperBoundAt X t ψ κ ν) (hY : HasMGFUpperBoundAt Y t ψY η (ν ⊗ₘ κ)) : HasMGFUpperBoundAt (fun p ↦ X p.1 + Y p.2) t (ψ + ψY) (κ ⊗ₖ η) ν","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-02980c53466d","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3224,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:127"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem add_compProd {η : ProbabilityTheory.Kernel (Ω' × Ω) Ω''} [ProbabilityTheory.IsZeroOrMarkovKernel η] (hX : HasMGFUpperBoundAt X t ψ κ ν) (hY : HasMGFUpperBoundAt Y t ψY η (ν ⊗ₘ κ)) : HasMGFUpperBoundAt (fun p ↦ X p.1 + Y p.2) t (ψ + ψY) (κ ⊗ₖ η) ν","missing":[],"search":"add_compprod banditrlproof.concentration.kernel.hasmgfupperboundat.add_compprod theorem add_compprod {η : probabilitytheory.kernel (ω' × ω) ω''} [probabilitytheory.iszeroormarkovkernel η] (hx : hasmgfupperboundat x t ψ κ ν) (hy : hasmgfupperboundat y t ψy η (ν ⊗ₘ κ)) : hasmgfupperboundat (fun p ↦ x p.1 + y p.2) t (ψ + ψy) (κ ⊗ₖ η) ν theorem compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.add_comp","label":"add_comp","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.add_comp","description":"lemma add_comp {η : ProbabilityTheory.Kernel Ω Ω''} [ProbabilityTheory.IsZeroOrMarkovKernel η] (hX : HasMGFUpperBoundAt X t ψ κ ν) (hY : HasMGFUpperBoundAt Y t ψY η (κ ∘ₘ ν)) : HasMGFUpperBoundAt (fun p ↦ X p.1 + Y p.2) t (ψ + ψY) (κ ⊗ₖ ProbabilityTheory.Kernel.prodMkLeft Ω' η) ν","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-eb555e4c5ed6","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3225,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:156"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma add_comp {η : ProbabilityTheory.Kernel Ω Ω''} [ProbabilityTheory.IsZeroOrMarkovKernel η] (hX : HasMGFUpperBoundAt X t ψ κ ν) (hY : HasMGFUpperBoundAt Y t ψY η (κ ∘ₘ ν)) : HasMGFUpperBoundAt (fun p ↦ X p.1 + Y p.2) t (ψ + ψY) (κ ⊗ₖ ProbabilityTheory.Kernel.prodMkLeft Ω' η) ν","missing":[],"search":"add_comp banditrlproof.concentration.kernel.hasmgfupperboundat.add_comp lemma add_comp {η : probabilitytheory.kernel ω ω''} [probabilitytheory.iszeroormarkovkernel η] (hx : hasmgfupperboundat x t ψ κ ν) (hy : hasmgfupperboundat y t ψy η (κ ∘ₘ ν)) : hasmgfupperboundat (fun p ↦ x p.1 + y p.2) t (ψ + ψy) (κ ⊗ₖ probabilitytheory.kernel.prodmkleft ω' η) ν lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt","label":"HasMGFUpperBoundAt","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt","description":"A measure-level MGF upper bound at one fixed tilt.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-ef3054ce7cd2","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3226,"meta":[["Kind","structure"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:172"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure HasMGFUpperBoundAt (X : Ω → ℝ) (t ψ : ℝ) (μ : Measure Ω := by volume_tac) : Prop where","missing":[],"search":"hasmgfupperboundat banditrlproof.concentration.hasmgfupperboundat a measure-level mgf upper bound at one fixed tilt. structure compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasMGFUpperBoundAt_iff_kernel","label":"hasMGFUpperBoundAt_iff_kernel","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.hasMGFUpperBoundAt_iff_kernel","description":"lemma hasMGFUpperBoundAt_iff_kernel : HasMGFUpperBoundAt X t ψ μ ↔ Kernel.HasMGFUpperBoundAt X t ψ (ProbabilityTheory.Kernel.const Unit μ) (Measure.dirac ())","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-0de709a67cd6","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3227,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:177"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma hasMGFUpperBoundAt_iff_kernel : HasMGFUpperBoundAt X t ψ μ ↔ Kernel.HasMGFUpperBoundAt X t ψ (ProbabilityTheory.Kernel.const Unit μ) (Measure.dirac ())","missing":[],"search":"hasmgfupperboundat_iff_kernel banditrlproof.concentration.hasmgfupperboundat_iff_kernel lemma hasmgfupperboundat_iff_kernel : hasmgfupperboundat x t ψ μ ↔ kernel.hasmgfupperboundat x t ψ (probabilitytheory.kernel.const unit μ) (measure.dirac ()) lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.of_map","label":"of_map","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.of_map","description":"lemma of_map {Ω' : Type*} {mΩ' : MeasurableSpace Ω'} {μ' : Measure Ω'} {Z : Ω' → Ω} (hZ : AEMeasurable Z μ') (h : HasMGFUpperBoundAt X t ψ (μ'.map Z)) : HasMGFUpperBoundAt (X ∘ Z) t ψ μ' where","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-13f3fe5f8752","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3228,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:186"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma of_map {Ω' : Type*} {mΩ' : MeasurableSpace Ω'} {μ' : Measure Ω'} {Z : Ω' → Ω} (hZ : AEMeasurable Z μ') (h : HasMGFUpperBoundAt X t ψ (μ'.map Z)) : HasMGFUpperBoundAt (X ∘ Z) t ψ μ' where","missing":[],"search":"of_map banditrlproof.concentration.hasmgfupperboundat.of_map lemma of_map {ω' : type*} {mω' : measurablespace ω'} {μ' : measure ω'} {z : ω' → ω} (hz : aemeasurable z μ') (h : hasmgfupperboundat x t ψ (μ'.map z)) : hasmgfupperboundat (x ∘ z) t ψ μ' where lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.id_map_iff","label":"id_map_iff","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.id_map_iff","description":"lemma id_map_iff (hX : AEMeasurable X μ) : HasMGFUpperBoundAt id t ψ (μ.map X) ↔ HasMGFUpperBoundAt X t ψ μ","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-42ac11b98f39","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3229,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:197"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma id_map_iff (hX : AEMeasurable X μ) : HasMGFUpperBoundAt id t ψ (μ.map X) ↔ HasMGFUpperBoundAt X t ψ μ","missing":[],"search":"id_map_iff banditrlproof.concentration.hasmgfupperboundat.id_map_iff lemma id_map_iff (hx : aemeasurable x μ) : hasmgfupperboundat id t ψ (μ.map x) ↔ hasmgfupperboundat x t ψ μ lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.compensated","label":"compensated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.compensated","description":"Subtracting a deterministic log-MGF budget turns a fixed-tilt bound into a unit-tilt zero-budget bound. This is the algebraic step used to iterate exponential supermartingale increments.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-27d5a21754f9","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3230,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:210"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem compensated (h : HasMGFUpperBoundAt X t ψ μ) : HasMGFUpperBoundAt (fun ω => t * X ω - ψ) 1 0 μ","missing":[],"search":"compensated banditrlproof.concentration.hasmgfupperboundat.compensated subtracting a deterministic log-mgf budget turns a fixed-tilt bound into a unit-tilt zero-budget bound. this is the algebraic step used to iterate exponential supermartingale increments. theorem compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasCondMGFUpperBoundAt","label":"HasCondMGFUpperBoundAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.HasCondMGFUpperBoundAt","description":"Conditional fixed-tilt MGF bound, expressed through Mathlib's conditional expectation kernel.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-e85bf3f4a902","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3231,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:239"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def HasCondMGFUpperBoundAt (X : Ω → ℝ) (t ψ : ℝ) (μ : Measure Ω := by volume_tac) [IsFiniteMeasure μ] : Prop","missing":[],"search":"hascondmgfupperboundat banditrlproof.concentration.hascondmgfupperboundat conditional fixed-tilt mgf bound, expressed through mathlib's conditional expectation kernel. definition compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.trim","label":"trim","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.trim","description":"lemma trim (hm : m ≤ mΩ) (hXm : Measurable[m] X) (hX : HasMGFUpperBoundAt X t ψ μ) : HasMGFUpperBoundAt X t ψ (μ.trim hm) where","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-21618094ffd6","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3232,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:246"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma trim (hm : m ≤ mΩ) (hXm : Measurable[m] X) (hX : HasMGFUpperBoundAt X t ψ μ) : HasMGFUpperBoundAt X t ψ (μ.trim hm) where","missing":[],"search":"trim banditrlproof.concentration.hasmgfupperboundat.trim lemma trim (hm : m ≤ mω) (hxm : measurable[m] x) (hx : hasmgfupperboundat x t ψ μ) : hasmgfupperboundat x t ψ (μ.trim hm) where lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.add_of_hasCondMGFUpperBoundAt","label":"add_of_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.add_of_hasCondMGFUpperBoundAt","description":"theorem add_of_hasCondMGFUpperBoundAt {Y : Ω → ℝ} {ψX ψY : ℝ} (hm : m ≤ mΩ) (hX : HasMGFUpperBoundAt X t ψX (μ.trim hm)) (hY : HasCondMGFUpperBoundAt m hm Y t ψY μ) : HasMGFUpperBoundAt (X + Y) t (ψX + ψY) μ","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-487efa3aeacc","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3233,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:256"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem add_of_hasCondMGFUpperBoundAt {Y : Ω → ℝ} {ψX ψY : ℝ} (hm : m ≤ mΩ) (hX : HasMGFUpperBoundAt X t ψX (μ.trim hm)) (hY : HasCondMGFUpperBoundAt m hm Y t ψY μ) : HasMGFUpperBoundAt (X + Y) t (ψX + ψY) μ","missing":[],"search":"add_of_hascondmgfupperboundat banditrlproof.concentration.hasmgfupperboundat.add_of_hascondmgfupperboundat theorem add_of_hascondmgfupperboundat {y : ω → ℝ} {ψx ψy : ℝ} (hm : m ≤ mω) (hx : hasmgfupperboundat x t ψx (μ.trim hm)) (hy : hascondmgfupperboundat m hm y t ψy μ) : hasmgfupperboundat (x + y) t (ψx + ψy) μ theorem compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.measure_ge_le_exp_add","label":"measure_ge_le_exp_add","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.measure_ge_le_exp_add","description":"lemma measure_ge_le_exp_add (h : HasMGFUpperBoundAt X t ψ μ) (ε : ℝ) (ht : 0 ≤ t) : μ.real {ω | ε ≤ X ω} ≤ exp (-t * ε + ψ)","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-44e41ee71f8c","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3234,"meta":[["Kind","lemma"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:276"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"lemma measure_ge_le_exp_add (h : HasMGFUpperBoundAt X t ψ μ) (ε : ℝ) (ht : 0 ≤ t) : μ.real {ω | ε ≤ X ω} ≤ exp (-t * ε + ψ)","missing":[],"search":"measure_ge_le_exp_add banditrlproof.concentration.hasmgfupperboundat.measure_ge_le_exp_add lemma measure_ge_le_exp_add (h : hasmgfupperboundat x t ψ μ) (ε : ℝ) (ht : 0 ≤ t) : μ.real {ω | ε ≤ x ω} ≤ exp (-t * ε + ψ) lemma compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.sum_of_hasCondMGFUpperBoundAt","label":"sum_of_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.HasMGFUpperBoundAt.sum_of_hasCondMGFUpperBoundAt","description":"theorem HasMGFUpperBoundAt.sum_of_hasCondMGFUpperBoundAt [IsZeroOrProbabilityMeasure μ] (h_adapted : StronglyAdapted ℱ Y) (h0 : HasMGFUpperBoundAt (Y 0) t (ψY 0) μ) (n : ℕ) (h_mgf : ∀ i < n - 1, HasCondMGFUpperBoundAt (ℱ i) (ℱ.le i) (Y (i + 1)) t (ψY (i + 1)) μ) : HasMGFUpperBoundAt (fun ω ↦ ∑ i ∈ Finset.range n, Y i ω) t (∑ i ∈ Finset.range n, ψY i) μ","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-4b6b331b7b4c","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3235,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:290"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem HasMGFUpperBoundAt.sum_of_hasCondMGFUpperBoundAt [IsZeroOrProbabilityMeasure μ] (h_adapted : StronglyAdapted ℱ Y) (h0 : HasMGFUpperBoundAt (Y 0) t (ψY 0) μ) (n : ℕ) (h_mgf : ∀ i < n - 1, HasCondMGFUpperBoundAt (ℱ i) (ℱ.le i) (Y (i + 1)) t (ψY (i + 1)) μ) : HasMGFUpperBoundAt (fun ω ↦ ∑ i ∈ Finset.range n, Y i ω) t (∑ i ∈ Finset.range n, ψY i) μ","missing":[],"search":"sum_of_hascondmgfupperboundat banditrlproof.concentration.hasmgfupperboundat.sum_of_hascondmgfupperboundat theorem hasmgfupperboundat.sum_of_hascondmgfupperboundat [iszeroorprobabilitymeasure μ] (h_adapted : stronglyadapted ℱ y) (h0 : hasmgfupperboundat (y 0) t (ψy 0) μ) (n : ℕ) (h_mgf : ∀ i < n - 1, hascondmgfupperboundat (ℱ i) (ℱ.le i) (y (i + 1)) t (ψy (i + 1)) μ) : hasmgfupperboundat (fun ω ↦ ∑ i ∈ finset.range n, y i ω) t (∑ i ∈ finset.range n, ψy i) μ theorem compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_sum_ge_le_of_hasCondMGFUpperBoundAt","label":"measure_sum_ge_le_of_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_sum_ge_le_of_hasCondMGFUpperBoundAt","description":"Fixed-tilt Chernoff bound for a strongly adapted finite sum with conditional MGF budgets.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-8cd8d753cce5","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3236,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:316"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_sum_ge_le_of_hasCondMGFUpperBoundAt [IsZeroOrProbabilityMeasure μ] (h_adapted : StronglyAdapted ℱ Y) (h0 : HasMGFUpperBoundAt (Y 0) t (ψY 0) μ) (n : ℕ) (h_mgf : ∀ i < n - 1, HasCondMGFUpperBoundAt (ℱ i) (ℱ.le i) (Y (i + 1)) t (ψY (i + 1)) μ) (ε : ℝ) (ht : 0 ≤ t) : μ.real {ω | ε ≤ ∑ i ∈ Finset.range n, Y i ω} ≤ exp (-t * ε + ∑ i ∈ Finset.range n, ψY i)","missing":[],"search":"measure_sum_ge_le_of_hascondmgfupperboundat banditrlproof.concentration.measure_sum_ge_le_of_hascondmgfupperboundat fixed-tilt chernoff bound for a strongly adapted finite sum with conditional mgf budgets. theorem compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_sum_ge_inter_sum_le_of_compensated_hasCondMGFUpperBoundAt","label":"measure_sum_ge_inter_sum_le_of_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_sum_ge_inter_sum_le_of_compensated_hasCondMGFUpperBoundAt","description":"Fixed-tilt tail bound with a random predictable compensator retained in the event. The conditional MGF hypotheses are imposed on the compensated increments `tilt * Y i - varianceCoeff * V i`; on the event where the cumulative compensator is at most `varianceBudget`, their exponential tail controls the uncompensated sum.","url":"../modules/banditrlproof-concentrationfixedmgf/index.html#decl-a06c99e85a76","parent":"module:BanditRLProof.ConcentrationFixedMGF","order":3237,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationFixedMGF"],["Source","BanditRLProof/ConcentrationFixedMGF.lean:333"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_sum_ge_inter_sum_le_of_compensated_hasCondMGFUpperBoundAt [IsZeroOrProbabilityMeasure μ] (Y V : ℕ → Ω → ℝ) (n : ℕ) (tilt varianceCoeff threshold varianceBudget : ℝ) (h_adapted : StronglyAdapted ℱ (fun i ω => tilt * Y i ω - varianceCoeff * V i ω)) (h0 : HasMGFUpperBoundAt (fun ω => tilt * Y 0 ω - varianceCoeff * V 0 ω) 1 0 μ) (h_mgf : ∀ i < n - 1, HasCondMGFUpperBoundAt (ℱ i) (ℱ.le i) (fun ω => tilt * Y (i + 1) ω - varianceCoeff * V (i + 1) ω) 1 0 μ) (htilt : 0 ≤ tilt) (hvarianceCoeff : 0 ≤ varianceCoeff) : μ {ω | threshold ≤ ∑ i ∈ Finset.range n, Y i ω ∧ (∑ i ∈ Finset.range n, V i ω) ≤ varianceBudget} ≤ ENNReal.ofReal (Real.exp (-tilt * threshold + varianceCoeff * varianceBudget))","missing":[],"search":"measure_sum_ge_inter_sum_le_of_compensated_hascondmgfupperboundat banditrlproof.concentration.measure_sum_ge_inter_sum_le_of_compensated_hascondmgfupperboundat fixed-tilt tail bound with a random predictable compensator retained in the event. the conditional mgf hypotheses are imposed on the compensated increments `tilt * y i - variancecoeff * v i`; on the event where the cumulative compensator is at most `variancebudget`, their exponential tail controls the uncompensated sum. theorem compiled","shard":"modules/8e19fc3b9fb87664.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_mul_exp_neg_mul_sq_Ioi","label":"integral_mul_exp_neg_mul_sq_Ioi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_mul_exp_neg_mul_sq_Ioi","description":"theorem integral_mul_exp_neg_mul_sq_Ioi (b : ℝ) (hb : 0 < b) : ∫ x : ℝ in Ioi 0, x*exp (-b*x^2) = (2*b)⁻¹","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-4fffea825fe9","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3238,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:11"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_mul_exp_neg_mul_sq_Ioi (b : ℝ) (hb : 0 < b) : ∫ x : ℝ in Ioi 0, x*exp (-b*x^2) = (2*b)⁻¹","missing":[],"search":"integral_mul_exp_neg_mul_sq_ioi banditrlproof.concentration.integral_mul_exp_neg_mul_sq_ioi theorem integral_mul_exp_neg_mul_sq_ioi (b : ℝ) (hb : 0 < b) : ∫ x : ℝ in ioi 0, x*exp (-b*x^2) = (2*b)⁻¹ theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_transformed_occupancy_tail","label":"integral_transformed_occupancy_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_transformed_occupancy_tail","description":"Exact transformed tail integral in source Lemma 8.2.","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-52585d99c029","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3239,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:27"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_transformed_occupancy_tail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : ∫ z : ℝ in Ioi 0, (2/ε^2)*(z+sqrt (2*a))*exp (-(1/2 : ℝ)*z^2) = (2/ε^2)*(1+sqrt (Real.pi*a))","missing":[],"search":"integral_transformed_occupancy_tail banditrlproof.concentration.integral_transformed_occupancy_tail exact transformed tail integral in source lemma 8.2. theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.occupancyTail","label":"occupancyTail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.occupancyTail","description":"The shifted Gaussian kernel after the algebraic square-root rewrite.","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-10e6d3b96a57","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3240,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:47"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def occupancyTail (a ε t : ℝ) : ℝ","missing":[],"search":"occupancytail banditrlproof.concentration.occupancytail the shifted gaussian kernel after the algebraic square-root rewrite. definition compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.occupancyTail_antitoneOn","label":"occupancyTail_antitoneOn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.occupancyTail_antitoneOn","description":"theorem occupancyTail_antitoneOn (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : AntitoneOn (occupancyTail a ε) (Ici (2*a/ε^2))","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-ebd8467764e6","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3241,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:49"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem occupancyTail_antitoneOn (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : AntitoneOn (occupancyTail a ε) (Ici (2*a/ε^2))","missing":[],"search":"occupancytail_antitoneon banditrlproof.concentration.occupancytail_antitoneon theorem occupancytail_antitoneon (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : antitoneon (occupancytail a ε) (ici (2*a/ε^2)) theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.occupancySubstitution","label":"occupancySubstitution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.occupancySubstitution","description":"def occupancySubstitution (a ε z : ℝ) : ℝ","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-5ed5c490eefd","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3242,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:66"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def occupancySubstitution (a ε z : ℝ) : ℝ","missing":[],"search":"occupancysubstitution banditrlproof.concentration.occupancysubstitution def occupancysubstitution (a ε z : ℝ) : ℝ definition compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.occupancySubstitution_image","label":"occupancySubstitution_image","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.occupancySubstitution_image","description":"theorem occupancySubstitution_image (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : occupancySubstitution a ε '' Ioi 0 = Ioi (2*a/ε^2)","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-7f0354ab69ba","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3243,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:68"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem occupancySubstitution_image (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : occupancySubstitution a ε '' Ioi 0 = Ioi (2*a/ε^2)","missing":[],"search":"occupancysubstitution_image banditrlproof.concentration.occupancysubstitution_image theorem occupancysubstitution_image (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : occupancysubstitution a ε '' ioi 0 = ioi (2*a/ε^2) theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.occupancySubstitution_injOn","label":"occupancySubstitution_injOn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.occupancySubstitution_injOn","description":"theorem occupancySubstitution_injOn (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : InjOn (occupancySubstitution a ε) (Ioi 0)","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-7acaf212f32e","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3244,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:90"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem occupancySubstitution_injOn (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : InjOn (occupancySubstitution a ε) (Ioi 0)","missing":[],"search":"occupancysubstitution_injon banditrlproof.concentration.occupancysubstitution_injon theorem occupancysubstitution_injon (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : injon (occupancysubstitution a ε) (ioi 0) theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasDerivAt_occupancySubstitution","label":"hasDerivAt_occupancySubstitution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasDerivAt_occupancySubstitution","description":"theorem hasDerivAt_occupancySubstitution (a ε z : ℝ) : HasDerivAt (occupancySubstitution a ε) ((2/ε^2)*(z+sqrt (2*a))) z","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-c1765c8803f8","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3245,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:103"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_occupancySubstitution (a ε z : ℝ) : HasDerivAt (occupancySubstitution a ε) ((2/ε^2)*(z+sqrt (2*a))) z","missing":[],"search":"hasderivat_occupancysubstitution banditrlproof.concentration.hasderivat_occupancysubstitution theorem hasderivat_occupancysubstitution (a ε z : ℝ) : hasderivat (occupancysubstitution a ε) ((2/ε^2)*(z+sqrt (2*a))) z theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_occupancyTail","label":"integral_occupancyTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_occupancyTail","description":"theorem integral_occupancyTail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : ∫ t in Ioi (2*a/ε^2), occupancyTail a ε t = (2/ε^2)*(1+sqrt (Real.pi*a))","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-b6915282930b","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3246,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:109"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_occupancyTail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : ∫ t in Ioi (2*a/ε^2), occupancyTail a ε t = (2/ε^2)*(1+sqrt (Real.pi*a))","missing":[],"search":"integral_occupancytail banditrlproof.concentration.integral_occupancytail theorem integral_occupancytail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : ∫ t in ioi (2*a/ε^2), occupancytail a ε t = (2/ε^2)*(1+sqrt (real.pi*a)) theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integrableOn_occupancyTail","label":"integrableOn_occupancyTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integrableOn_occupancyTail","description":"theorem integrableOn_occupancyTail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : IntegrableOn (occupancyTail a ε) (Ioi (2*a/ε^2))","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-70fe5c982645","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3247,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:128"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrableOn_occupancyTail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : IntegrableOn (occupancyTail a ε) (Ioi (2*a/ε^2))","missing":[],"search":"integrableon_occupancytail banditrlproof.concentration.integrableon_occupancytail theorem integrableon_occupancytail (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) : integrableon (occupancytail a ε) (ioi (2*a/ε^2)) theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_occupancyTail_shift_le","label":"sum_occupancyTail_shift_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_occupancyTail_shift_le","description":"theorem sum_occupancyTail_shift_le (a ε r : ℝ) (ha : 0 < a) (hε : 0 < ε) (hr : 2*a/ε^2 ≤ r) (N : ℕ) : (∑ i ∈ Finset.range N, occupancyTail a ε (r+(i+1 : ℕ))) ≤ (2/ε^2)*(1+sqrt (Real.pi*a))","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-f6d61677a65f","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3248,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:149"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_occupancyTail_shift_le (a ε r : ℝ) (ha : 0 < a) (hε : 0 < ε) (hr : 2*a/ε^2 ≤ r) (N : ℕ) : (∑ i ∈ Finset.range N, occupancyTail a ε (r+(i+1 : ℕ))) ≤ (2/ε^2)*(1+sqrt (Real.pi*a))","missing":[],"search":"sum_occupancytail_shift_le banditrlproof.concentration.sum_occupancytail_shift_le theorem sum_occupancytail_shift_le (a ε r : ℝ) (ha : 0 < a) (hε : 0 < ε) (hr : 2*a/ε^2 ≤ r) (n : ℕ) : (∑ i ∈ finset.range n, occupancytail a ε (r+(i+1 : ℕ))) ≤ (2/ε^2)*(1+sqrt (real.pi*a)) theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_le_occupancy_bound","label":"sum_le_occupancy_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_le_occupancy_bound","description":"Integer cutoff aggregation with the exact source Lemma 8.2 constants.","url":"../modules/banditrlproof-concentrationgaussianoccupancy/index.html#decl-c796dd885dd1","parent":"module:BanditRLProof.ConcentrationGaussianOccupancy","order":3249,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationGaussianOccupancy"],["Source","BanditRLProof/ConcentrationGaussianOccupancy.lean:166"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_le_occupancy_bound (p : ℕ → ℝ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (h1 : ∀ s, p s ≤ 1) (htail : ∀ s : ℕ, 2*a/ε^2 < (s : ℝ) → p s ≤ occupancyTail a ε s) (n : ℕ) : (∑ i ∈ Finset.range n, p (i+1)) ≤ 1+(2/ε^2)*(a+sqrt (Real.pi*a)+1)","missing":[],"search":"sum_le_occupancy_bound banditrlproof.concentration.sum_le_occupancy_bound integer cutoff aggregation with the exact source lemma 8.2 constants. theorem compiled","shard":"modules/68489066b720b681.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.fixedRadiusMeanEvent","label":"fixedRadiusMeanEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.fixedRadiusMeanEvent","description":"def fixedRadiusMeanEvent (X : ℕ → Ω → ℝ) (a ε : ℝ) (s : ℕ) : Set Ω","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-a7e2549a51f9","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3250,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:11"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def fixedRadiusMeanEvent (X : ℕ → Ω → ℝ) (a ε : ℝ) (s : ℕ) : Set Ω","missing":[],"search":"fixedradiusmeanevent banditrlproof.concentration.fixedradiusmeanevent def fixedradiusmeanevent (x : ℕ → ω → ℝ) (a ε : ℝ) (s : ℕ) : set ω definition compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_fixedRadiusMeanEvent_le","label":"measure_fixedRadiusMeanEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_fixedRadiusMeanEvent_le","description":"theorem measure_fixedRadiusMeanEvent_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (s : ℕ) (hs : 2*a/ε^2 < (s : ℝ)) : μ (fixedRadiusMeanEvent X a ε s) ≤ ENNReal.ofReal (occupancyTail a ε s)","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-7b809d80d105","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3251,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:14"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_fixedRadiusMeanEvent_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (s : ℕ) (hs : 2*a/ε^2 < (s : ℝ)) : μ (fixedRadiusMeanEvent X a ε s) ≤ ENNReal.ofReal (occupancyTail a ε s)","missing":[],"search":"measure_fixedradiusmeanevent_le banditrlproof.concentration.measure_fixedradiusmeanevent_le theorem measure_fixedradiusmeanevent_le (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (s : ℕ) (hs : 2*a/ε^2 < (s : ℝ)) : μ (fixedradiusmeanevent x a ε s) ≤ ennreal.ofreal (occupancytail a ε s) theorem compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.sum_measureReal_fixedRadiusMeanEvent_le","label":"sum_measureReal_fixedRadiusMeanEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.sum_measureReal_fixedRadiusMeanEvent_le","description":"theorem sum_measureReal_fixedRadiusMeanEvent_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (n : ℕ) : (∑ i ∈ range n, μ.real (fixedRadiusMeanEvent X a ε (i+1))) ≤ 1+(2/ε^2)*(a+sqrt (Real.pi*a)+1)","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-44b419b12dbf","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3252,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:52"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem sum_measureReal_fixedRadiusMeanEvent_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (n : ℕ) : (∑ i ∈ range n, μ.real (fixedRadiusMeanEvent X a ε (i+1))) ≤ 1+(2/ε^2)*(a+sqrt (Real.pi*a)+1)","missing":[],"search":"sum_measurereal_fixedradiusmeanevent_le banditrlproof.concentration.sum_measurereal_fixedradiusmeanevent_le theorem sum_measurereal_fixedradiusmeanevent_le (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (hind : iindepfun x μ) (hmean : ∀ i, ∫ ω, x i ω ∂μ = 0) (hsubg : ∀ i, hassubgaussianmgf (x i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (n : ℕ) : (∑ i ∈ range n, μ.real (fixedradiusmeanevent x a ε (i+1))) ≤ 1+(2/ε^2)*(a+sqrt (real.pi*a)+1) theorem compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measurableSet_fixedRadiusMeanEvent","label":"measurableSet_fixedRadiusMeanEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measurableSet_fixedRadiusMeanEvent","description":"theorem measurableSet_fixedRadiusMeanEvent (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (a ε : ℝ) (s : ℕ) : MeasurableSet (fixedRadiusMeanEvent X a ε s)","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-64148c6f57cb","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3253,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:65"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_fixedRadiusMeanEvent (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (a ε : ℝ) (s : ℕ) : MeasurableSet (fixedRadiusMeanEvent X a ε s)","missing":[],"search":"measurableset_fixedradiusmeanevent banditrlproof.concentration.measurableset_fixedradiusmeanevent theorem measurableset_fixedradiusmeanevent (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (a ε : ℝ) (s : ℕ) : measurableset (fixedradiusmeanevent x a ε s) theorem compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.fixedRadiusCount","label":"fixedRadiusCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.fixedRadiusCount","description":"def fixedRadiusCount (X : ℕ → Ω → ℝ) (a ε : ℝ) (n : ℕ) (ω : Ω) : ℝ","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-58064c96e38b","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3254,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:73"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def fixedRadiusCount (X : ℕ → Ω → ℝ) (a ε : ℝ) (n : ℕ) (ω : Ω) : ℝ","missing":[],"search":"fixedradiuscount banditrlproof.concentration.fixedradiuscount def fixedradiuscount (x : ℕ → ω → ℝ) (a ε : ℝ) (n : ℕ) (ω : ω) : ℝ definition compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integrable_fixedRadiusCount","label":"integrable_fixedRadiusCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integrable_fixedRadiusCount","description":"theorem integrable_fixedRadiusCount (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (a ε : ℝ) (n : ℕ) : Integrable (fixedRadiusCount X a ε n) μ","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-f60137a6fc91","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3255,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:76"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_fixedRadiusCount (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (a ε : ℝ) (n : ℕ) : Integrable (fixedRadiusCount X a ε n) μ","missing":[],"search":"integrable_fixedradiuscount banditrlproof.concentration.integrable_fixedradiuscount theorem integrable_fixedradiuscount (x : ℕ → ω → ℝ) (hxm : ∀ i, stronglymeasurable (x i)) (a ε : ℝ) (n : ℕ) : integrable (fixedradiuscount x a ε n) μ theorem compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_fixedRadiusCount_le","label":"integral_fixedRadiusCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_fixedRadiusCount_le","description":"Source Lemma 8.2 expected-count conclusion for centered unit-subgaussian coordinates.","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-ee74fe8e2256","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3256,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:83"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_fixedRadiusCount_le (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (n : ℕ) : (∫ ω, fixedRadiusCount X a ε n ω ∂μ) ≤ 1+(2/ε^2)*(a+sqrt (Real.pi*a)+1)","missing":[],"search":"integral_fixedradiuscount_le banditrlproof.concentration.integral_fixedradiuscount_le source lemma 8.2 expected-count conclusion for centered unit-subgaussian coordinates. theorem compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_fixedRadiusCount_le_sharp","label":"integral_fixedRadiusCount_le_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_fixedRadiusCount_le_sharp","description":"Sharper count bound: the removed additive one pays for MOSS initialization.","url":"../modules/banditrlproof-concentrationindexoccupancy/index.html#decl-65643bd2e9d5","parent":"module:BanditRLProof.ConcentrationIndexOccupancy","order":3257,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationIndexOccupancy"],["Source","BanditRLProof/ConcentrationIndexOccupancy.lean:102"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_fixedRadiusCount_le_sharp (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (hsubG : ∀ i, HasSubgaussianMGF (X i) 1 μ) (a ε : ℝ) (ha : 0 < a) (hε : 0 < ε) (n : ℕ) : (∫ ω, fixedRadiusCount X a ε n ω ∂μ) ≤ (2/ε^2)*(a+sqrt (Real.pi*a)+1)","missing":[],"search":"integral_fixedradiuscount_le_sharp banditrlproof.concentration.integral_fixedradiuscount_le_sharp sharper count bound: the removed additive one pays for moss initialization. theorem compiled","shard":"modules/1c8fe1693a2e0a2c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.submartingale_exp_of_martingale","label":"submartingale_exp_of_martingale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.submartingale_exp_of_martingale","description":"Conditional Jensen turns an exponentially integrable real martingale into a nonnegative exponential submartingale.","url":"../modules/banditrlproof-concentrationmartingalemaximal/index.html#decl-bf7d0f50d062","parent":"module:BanditRLProof.ConcentrationMartingaleMaximal","order":3258,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationMartingaleMaximal"],["Source","BanditRLProof/ConcentrationMartingaleMaximal.lean:25"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem submartingale_exp_of_martingale (hS : Martingale S F μ) (hint : ∀ i, Integrable (fun ω => exp (S i ω)) μ) : Submartingale (fun i ω => exp (S i ω)) F μ","missing":[],"search":"submartingale_exp_of_martingale banditrlproof.concentration.submartingale_exp_of_martingale conditional jensen turns an exponentially integrable real martingale into a nonnegative exponential submartingale. theorem compiled","shard":"modules/0edaacebcfcfa795.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_exists_le_martingale_ge_le_exp","label":"measure_exists_le_martingale_ge_le_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_exists_le_martingale_ge_le_exp","description":"Finite-time maximal Chernoff bound with no union-bound cardinality factor. The terminal MGF supplies the variance budget; all-time exponential integrability is explicit for the Jensen producer.","url":"../modules/banditrlproof-concentrationmartingalemaximal/index.html#decl-4276294052a7","parent":"module:BanditRLProof.ConcentrationMartingaleMaximal","order":3259,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationMartingaleMaximal"],["Source","BanditRLProof/ConcentrationMartingaleMaximal.lean:38"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_exists_le_martingale_ge_le_exp (hS : Martingale S F μ) (hint : ∀ i t, Integrable (fun ω => exp (t * S i ω)) μ) (n : ℕ) (c : ℝ≥0) (hmgf : HasSubgaussianMGF (S n) c μ) (ε t : ℝ) (ht : 0 < t) : μ {ω | ∃ i, i ≤ n ∧ ε ≤ S i ω} ≤ ENNReal.ofReal (exp (-t * ε + (c : ℝ) * t ^ 2 / 2))","missing":[],"search":"measure_exists_le_martingale_ge_le_exp banditrlproof.concentration.measure_exists_le_martingale_ge_le_exp finite-time maximal chernoff bound with no union-bound cardinality factor. the terminal mgf supplies the variance budget; all-time exponential integrability is explicit for the jensen producer. theorem compiled","shard":"modules/0edaacebcfcfa795.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_exists_le_martingale_ge_le_subgaussian","label":"measure_exists_le_martingale_ge_le_subgaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_exists_le_martingale_ge_le_subgaussian","description":"Optimized finite maximal subgaussian bound.","url":"../modules/banditrlproof-concentrationmartingalemaximal/index.html#decl-77e207d059ad","parent":"module:BanditRLProof.ConcentrationMartingaleMaximal","order":3260,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationMartingaleMaximal"],["Source","BanditRLProof/ConcentrationMartingaleMaximal.lean:74"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_exists_le_martingale_ge_le_subgaussian (hS : Martingale S F μ) (hint : ∀ i t, Integrable (fun ω => exp (t * S i ω)) μ) (n : ℕ) (c : ℝ≥0) (hc : 0 < (c : ℝ)) (hmgf : HasSubgaussianMGF (S n) c μ) (ε : ℝ) (hε : 0 < ε) : μ {ω | ∃ i, i ≤ n ∧ ε ≤ S i ω} ≤ ENNReal.ofReal (exp (-(ε ^ 2) / (2 * (c : ℝ))))","missing":[],"search":"measure_exists_le_martingale_ge_le_subgaussian banditrlproof.concentration.measure_exists_le_martingale_ge_le_subgaussian optimized finite maximal subgaussian bound. theorem compiled","shard":"modules/0edaacebcfcfa795.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_exists_le_independent_partialSum_ge_le_subgaussian","label":"measure_exists_le_independent_partialSum_ge_le_subgaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_exists_le_independent_partialSum_ge_le_subgaussian","description":"Source Theorem 9.2 shape for independent centered subgaussian increments. The source variance is `c = σ²`; the partial sum uses X1 through Xn. Centering and coordinate measurability are explicit model contracts.","url":"../modules/banditrlproof-concentrationmartingalemaximal/index.html#decl-4be8a89f4f6b","parent":"module:BanditRLProof.ConcentrationMartingaleMaximal","order":3261,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationMartingaleMaximal"],["Source","BanditRLProof/ConcentrationMartingaleMaximal.lean:91"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem measure_exists_le_independent_partialSum_ge_le_subgaussian (X : ℕ → Ω → ℝ) (hXm : ∀ i, StronglyMeasurable (X i)) (hind : iIndepFun X μ) (hmean : ∀ i, ∫ ω, X i ω ∂μ = 0) (c : ℝ≥0) (hc : 0 < (c : ℝ)) (hsubG : ∀ i, HasSubgaussianMGF (X i) c μ) (n : ℕ) (hn : 0 < n) (ε : ℝ) (hε : 0 < ε) : μ {ω | ∃ i, i ≤ n ∧ ε ≤ ∑ j ∈ range i, X (j + 1) ω} ≤ ENNReal.ofReal (exp (-(ε ^ 2) / (2 * (n : ℝ) * (c : ℝ))))","missing":[],"search":"measure_exists_le_independent_partialsum_ge_le_subgaussian banditrlproof.concentration.measure_exists_le_independent_partialsum_ge_le_subgaussian source theorem 9.2 shape for independent centered subgaussian increments. the source variance is `c = σ²`; the partial sum uses x1 through xn. centering and coordinate measurability are explicit model contracts. theorem compiled","shard":"modules/0edaacebcfcfa795.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.exists_tilt_quadratic_fixedMGF_exponent_le_neg","label":"exists_tilt_quadratic_fixedMGF_exponent_le_neg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.exists_tilt_quadratic_fixedMGF_exponent_le_neg","description":"Optimize a quadratic fixed-tilt MGF budget when the quadratic coefficient and admissible tilt cap are separate parameters.","url":"../modules/banditrlproof-concentrationquadraticfixedmgf/index.html#decl-3041cf2c11c9","parent":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","order":3262,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationQuadraticFixedMGF"],["Source","BanditRLProof/ConcentrationQuadraticFixedMGF.lean:20"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem exists_tilt_quadratic_fixedMGF_exponent_le_neg (horizon variance cap budget : Real) (hhorizon : 0 <= horizon) (hvariance : 0 < variance) (hcap : 0 < cap) (hbudget : 0 <= budget) : exists tilt : Real, 0 <= tilt ∧ tilt <= cap ∧ -tilt * (2 * Real.sqrt (horizon * variance * budget) + budget / cap) + horizon * (tilt ^ 2 * variance) <= -budget","missing":[],"search":"exists_tilt_quadratic_fixedmgf_exponent_le_neg banditrlproof.concentration.exists_tilt_quadratic_fixedmgf_exponent_le_neg optimize a quadratic fixed-tilt mgf budget when the quadratic coefficient and admissible tilt cap are separate parameters. theorem compiled","shard":"modules/d7a6971ac1bd292e.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.quadraticFixedMGFRadius","label":"quadraticFixedMGFRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.quadraticFixedMGFRadius","description":"Radius obtained by optimizing a quadratic fixed-tilt exponent over `0 <= tilt <= tiltCap`.","url":"../modules/banditrlproof-concentrationquadraticfixedmgf/index.html#decl-beaa84da7e75","parent":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","order":3263,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationQuadraticFixedMGF"],["Source","BanditRLProof/ConcentrationQuadraticFixedMGF.lean:106"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def quadraticFixedMGFRadius (varianceScale varianceBudget tiltCap delta : Real) : Real","missing":[],"search":"quadraticfixedmgfradius banditrlproof.concentration.quadraticfixedmgfradius radius obtained by optimizing a quadratic fixed-tilt exponent over `0 <= tilt <= tiltcap`. definition compiled","shard":"modules/d7a6971ac1bd292e.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","label":"measure_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","description":"A family of fixed-tilt quadratic tails yields a delta-shaped joint deviation and variance-budget tail.","url":"../modules/banditrlproof-concentrationquadraticfixedmgf/index.html#decl-afdb4e478275","parent":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","order":3264,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationQuadraticFixedMGF"],["Source","BanditRLProof/ConcentrationQuadraticFixedMGF.lean:113"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) (deviation predictableVariance : Omega -> Real) (varianceScale varianceBudget tiltCap delta : Real) (hvarianceScale : 0 < varianceScale) (hvarianceBudget : 0 < varianceBudget) (htiltCap : 0 < tiltCap) (hdelta : 0 < delta) (hfixed : forall tilt, 0 <= tilt -> tilt <= tiltCap -> mu {omega | quadraticFixedMGFRadius varianceScale varianceBudget tiltCap delta <= deviation omega ∧ predictableVariance omega <= varianceBudget} <= ENNReal.ofReal (Real.exp (-tilt * quadraticFixedMGFRadius varianceScale varianceBudget tiltCap delta + varianceScale * (tilt ^ 2 * varianceBudget)))) : mu {omega | quadraticFixedMGFRadius varianceScale varianceBudget tiltCap delta <= deviation omega ∧ predictableVariance omega <= varianceBudget} <= ENNReal.ofReal delta","missing":[],"search":"measure_deviation_ge_inter_variance_le_delta_of_fixedtilt_quadratic_tail banditrlproof.concentration.measure_deviation_ge_inter_variance_le_delta_of_fixedtilt_quadratic_tail a family of fixed-tilt quadratic tails yields a delta-shaped joint deviation and variance-budget tail. theorem compiled","shard":"modules/d7a6971ac1bd292e.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.quadraticFixedMGFMaximalRadius","label":"quadraticFixedMGFMaximalRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.quadraticFixedMGFMaximalRadius","description":"Equal-share quadratic radius for a finite family of events.","url":"../modules/banditrlproof-concentrationquadraticmaximal/index.html#decl-a80ebfc5f3fc","parent":"module:BanditRLProof.ConcentrationQuadraticMaximal","order":3265,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationQuadraticMaximal"],["Source","BanditRLProof/ConcentrationQuadraticMaximal.lean:19"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def quadraticFixedMGFMaximalRadius {Idx : Type*} [DecidableEq Idx] (times : Finset Idx) (varianceScale varianceBudget tiltCap delta : Real) : Real","missing":[],"search":"quadraticfixedmgfmaximalradius banditrlproof.concentration.quadraticfixedmgfmaximalradius equal-share quadratic radius for a finite family of events. definition compiled","shard":"modules/24ac619ce203f767.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_biUnion_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","label":"measure_biUnion_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_biUnion_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","description":"Finite maximal delta tail obtained by optimizing every fixed-tilt event at confidence `delta / times.card` and taking the finite union.","url":"../modules/banditrlproof-concentrationquadraticmaximal/index.html#decl-807c0a275a54","parent":"module:BanditRLProof.ConcentrationQuadraticMaximal","order":3266,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationQuadraticMaximal"],["Source","BanditRLProof/ConcentrationQuadraticMaximal.lean:27"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_biUnion_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail {Omega Idx : Type*} [MeasurableSpace Omega] [DecidableEq Idx] (mu : Measure Omega) (times : Finset Idx) (htimes : times.Nonempty) (deviation predictableVariance : Idx -> Omega -> Real) (varianceScale varianceBudget tiltCap delta : Real) (hvarianceScale : 0 < varianceScale) (hvarianceBudget : 0 < varianceBudget) (htiltCap : 0 < tiltCap) (hdelta : 0 < delta) (hfixed : forall i, i ∈ times -> forall tilt, 0 <= tilt -> tilt <= tiltCap -> mu {omega | quadraticFixedMGFMaximalRadius times varianceScale varianceBudget tiltCap delta <= deviation i omega ∧ predictableVariance i omega <= varianceBudget} <= ENNReal.ofReal (Real.exp (-tilt * quadraticFixedMGFMaximalRadius times varianceScale varianceBudget tiltCap delta + varianceScale * (tilt ^ 2 * varianceBudget)))) : mu (⋃ i ∈ times, {omega | quadraticFixedMGF…","missing":[],"search":"measure_biunion_deviation_ge_inter_variance_le_delta_of_fixedtilt_quadratic_tail banditrlproof.concentration.measure_biunion_deviation_ge_inter_variance_le_delta_of_fixedtilt_quadratic_tail finite maximal delta tail obtained by optimizing every fixed-tilt event at confidence `delta / times.card` and taking the finite union. theorem compiled","shard":"modules/24ac619ce203f767.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.quadraticFixedMGFScheduledRadius","label":"quadraticFixedMGFScheduledRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.quadraticFixedMGFScheduledRadius","description":"The optimized quadratic radius evaluated at one index of a parameter schedule.","url":"../modules/banditrlproof-concentrationquadraticscheduled/index.html#decl-82ee3f53d453","parent":"module:BanditRLProof.ConcentrationQuadraticScheduled","order":3267,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationQuadraticScheduled"],["Source","BanditRLProof/ConcentrationQuadraticScheduled.lean:20"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def quadraticFixedMGFScheduledRadius (varianceScale varianceBudget tiltCap deltaAt : Nat -> Real) (n : Nat) : Real","missing":[],"search":"quadraticfixedmgfscheduledradius banditrlproof.concentration.quadraticfixedmgfscheduledradius the optimized quadratic radius evaluated at one index of a parameter schedule. definition compiled","shard":"modules/010cd7d40ab36a07.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedTilt_quadratic_tail","label":"measure_iUnion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedTilt_quadratic_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedTilt_quadratic_tail","description":"A countable family of quadratic fixed-MGF events is controlled by the sum of its scheduled confidence shares. Event measurability is not required because `Measure.measure_iUnion_le` is an outer-measure inequality.","url":"../modules/banditrlproof-concentrationquadraticscheduled/index.html#decl-0a79f448615c","parent":"module:BanditRLProof.ConcentrationQuadraticScheduled","order":3268,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationQuadraticScheduled"],["Source","BanditRLProof/ConcentrationQuadraticScheduled.lean:29"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_iUnion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedTilt_quadratic_tail {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) (deviation predictableVariance : Nat -> Omega -> Real) (varianceScale varianceBudget tiltCap deltaAt : Nat -> Real) (hvarianceScale : forall n, 0 < varianceScale n) (hvarianceBudget : forall n, 0 < varianceBudget n) (htiltCap : forall n, 0 < tiltCap n) (hdeltaAt : forall n, 0 < deltaAt n) (hfixed : forall n tilt, 0 <= tilt -> tilt <= tiltCap n -> mu {omega | quadraticFixedMGFScheduledRadius varianceScale varianceBudget tiltCap deltaAt n <= deviation n omega ∧ predictableVariance n omega <= varianceBudget n} <= ENNReal.ofReal (Real.exp (-tilt * quadraticFixedMGFScheduledRadius varianceScale varianceBudget tiltCap deltaAt n + varianceScale n * (tilt ^ 2 * varianceBudget n)))) : mu (⋃ n, {omega | quadraticFixedMGFScheduledRadius varia…","missing":[],"search":"measure_iunion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedtilt_quadratic_tail banditrlproof.concentration.measure_iunion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedtilt_quadratic_tail a countable family of quadratic fixed-mgf events is controlled by the sum of its scheduled confidence shares. event measurability is not required because `measure.measure_iunion_le` is an outer-measure inequality. theorem compiled","shard":"modules/010cd7d40ab36a07.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","label":"measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","description":"A caller-supplied total ENNReal budget turns the scheduled `tsum` bound into a direct outer confidence bound.","url":"../modules/banditrlproof-concentrationquadraticscheduled/index.html#decl-e61e72f328d0","parent":"module:BanditRLProof.ConcentrationQuadraticScheduled","order":3269,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationQuadraticScheduled"],["Source","BanditRLProof/ConcentrationQuadraticScheduled.lean:78"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) (deviation predictableVariance : Nat -> Omega -> Real) (varianceScale varianceBudget tiltCap deltaAt : Nat -> Real) (delta : Real) (hvarianceScale : forall n, 0 < varianceScale n) (hvarianceBudget : forall n, 0 < varianceBudget n) (htiltCap : forall n, 0 < tiltCap n) (hdeltaAt : forall n, 0 < deltaAt n) (hfixed : forall n tilt, 0 <= tilt -> tilt <= tiltCap n -> mu {omega | quadraticFixedMGFScheduledRadius varianceScale varianceBudget tiltCap deltaAt n <= deviation n omega ∧ predictableVariance n omega <= varianceBudget n} <= ENNReal.ofReal (Real.exp (-tilt * quadraticFixedMGFScheduledRadius varianceScale varianceBudget tiltCap deltaAt n + varianceScale n * (tilt ^ 2 * varianceBudget n)))) (hbudget : (∑' n, ENNReal.ofReal (deltaAt…","missing":[],"search":"measure_iunion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedtilt_quadratic_tail banditrlproof.concentration.measure_iunion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedtilt_quadratic_tail a caller-supplied total ennreal budget turns the scheduled `tsum` bound into a direct outer confidence bound. theorem compiled","shard":"modules/010cd7d40ab36a07.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.of_measurableSpace_eq","label":"of_measurableSpace_eq","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.of_measurableSpace_eq","description":"Transport a conditional sub-Gaussian witness across equality of the conditioning measurable spaces. The two sub-sigma-algebra proofs are propositionally irrelevant once the measurable spaces are identified. This is a general-purpose adapter for filtrations presented through different but extensionally equal histories.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-c3da72d1c1f2","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3270,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:17"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.of_measurableSpace_eq {Omega : Type u} {m0 m1 mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : MeasureTheory.Measure Omega} [MeasureTheory.IsFiniteMeasure mu] {X : Omega -> Real} {c : NNReal} (hm0 : m0 <= mOmega) (hm1 : m1 <= mOmega) (hm : m0 = m1) (hX : HasCondSubgaussianMGF m0 hm0 X c mu) : HasCondSubgaussianMGF m1 hm1 X c mu","missing":[],"search":"of_measurablespace_eq probabilitytheory.hascondsubgaussianmgf.of_measurablespace_eq transport a conditional sub-gaussian witness across equality of the conditioning measurable spaces. the two sub-sigma-algebra proofs are propositionally irrelevant once the measurable spaces are identified. this is a general-purpose adapter for filtrations presented through different but extensionally equal histories. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.integrable","label":"integrable","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.integrable","description":"A conditionally sub-Gaussian real random variable is integrable under the ambient measure. Mathlib's conditional MGF contract already includes exponential integrability for every real tilt. Applying it at `1` and `-1` gives integrability of the first absolute moment.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-4de8a6abba92","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3271,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:37"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.integrable {Omega : Type u} {m mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : MeasureTheory.Measure Omega} [MeasureTheory.IsFiniteMeasure mu] {X : Omega -> Real} {c : NNReal} (hm : m <= mOmega) (hX : HasCondSubgaussianMGF m hm X c mu) : MeasureTheory.Integrable X mu","missing":[],"search":"integrable probabilitytheory.hascondsubgaussianmgf.integrable a conditionally sub-gaussian real random variable is integrable under the ambient measure. mathlib's conditional mgf contract already includes exponential integrability for every real tilt. applying it at `1` and `-1` gives integrability of the first absolute moment. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.indicator","label":"indicator","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.indicator","description":"A conditionally sub-Gaussian variable remains conditionally sub-Gaussian with the same proxy after restriction to an event measurable in the conditioning sigma-algebra. The proof uses one common exceptional set for every exponential tilt: the conditional-expectation kernel is supported on the current side of a conditioning-measurable event, so the indicator-masked variable is kernel-a.e. equal either to the original…","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-a8a14d288e65","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3272,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:73"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.indicator {Omega : Type u} {m mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : MeasureTheory.Measure Omega} [MeasureTheory.IsProbabilityMeasure mu] {X : Omega -> Real} {c : NNReal} (hm : m <= mOmega) (hX : HasCondSubgaussianMGF m hm X c mu) {s : Set Omega} (hs : @MeasurableSet Omega m s) : HasCondSubgaussianMGF m hm (s.indicator X) c mu","missing":[],"search":"indicator probabilitytheory.hascondsubgaussianmgf.indicator a conditionally sub-gaussian variable remains conditionally sub-gaussian with the same proxy after restriction to an event measurable in the conditioning sigma-algebra. the proof uses one common exceptional set for every exponential tilt: the conditional-expectation kernel is supported on the current side of a conditioning-measurable event, so the indicator-masked variable is kernel-a.e. equal either to the original variable or to zero. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.indicator_compensated_hasCondMGFUpperBoundAt","label":"indicator_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.indicator_compensated_hasCondMGFUpperBoundAt","description":"At a fixed tilt, a conditioning-measurable mask pays the sub-Gaussian quadratic budget only on the masked event. This is the one-step predictable variance interface needed for count-sensitive adaptive-sampling tails.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-21c1076c402e","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3273,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:171"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.indicator_compensated_hasCondMGFUpperBoundAt {Omega : Type u} {m mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : MeasureTheory.Measure Omega} [MeasureTheory.IsProbabilityMeasure mu] {X : Omega -> Real} {c : NNReal} (hm : m <= mOmega) (hX : HasCondSubgaussianMGF m hm X c mu) {s : Set Omega} (hs : @MeasurableSet Omega m s) (tilt : Real) : BanditRLProof.Concentration.HasCondMGFUpperBoundAt m hm (fun omega => tilt * s.indicator X omega - (((c : NNReal) : Real) * tilt ^ 2 / 2) * s.indicator (fun _ : Omega => (1 : Real)) omega) 1 0 mu","missing":[],"search":"indicator_compensated_hascondmgfupperboundat probabilitytheory.hascondsubgaussianmgf.indicator_compensated_hascondmgfupperboundat at a fixed tilt, a conditioning-measurable mask pays the sub-gaussian quadratic budget only on the masked event. this is the one-step predictable variance interface needed for count-sensitive adaptive-sampling tails. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.intervalVarianceProxy","label":"intervalVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.intervalVarianceProxy","description":"Variance proxy induced by an almost-sure interval bound `[lo, hi]`. This is the reusable Hoeffding proxy `(hi - lo)^2 / 4`, represented in the same `NNReal` shape used by Mathlib's bounded-variable sub-Gaussian lemma.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-aa0d05030b22","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3274,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:332"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def intervalVarianceProxy (lo hi : Real) : NNReal","missing":[],"search":"intervalvarianceproxy banditrlproof.concentration.intervalvarianceproxy variance proxy induced by an almost-sure interval bound `[lo, hi]`. this is the reusable hoeffding proxy `(hi - lo)^2 / 4`, represented in the same `nnreal` shape used by mathlib's bounded-variable sub-gaussian lemma. definition compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.boundedCentered_hasSubgaussianMGF_of_mem_Icc_integral_eq","label":"boundedCentered_hasSubgaussianMGF_of_mem_Icc_integral_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.boundedCentered_hasSubgaussianMGF_of_mem_Icc_integral_eq","description":"Bounded variable plus an exact mean identity gives a centered sub-Gaussian witness. This is the generic `TAIL-HOEFFDING-BOUNDED` import wrapper. It keeps the Mathlib assumptions explicit: an a.e.-measurable real variable, an a.s. interval bound, and an equality between its integral and the supplied mean.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-3df8b97c8f84","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3275,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:343"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem boundedCentered_hasSubgaussianMGF_of_mem_Icc_integral_eq {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] {X : Omega -> Real} {lo hi mean : Real} (hmeas : AEMeasurable X mu) (hbound : Filter.Eventually (fun omega : Omega => Set.Icc lo hi (X omega)) (ae mu)) (hmean : integral mu X = mean) : ProbabilityTheory.HasSubgaussianMGF (fun omega : Omega => X omega - mean) (intervalVarianceProxy lo hi) mu","missing":[],"search":"boundedcentered_hassubgaussianmgf_of_mem_icc_integral_eq banditrlproof.concentration.boundedcentered_hassubgaussianmgf_of_mem_icc_integral_eq bounded variable plus an exact mean identity gives a centered sub-gaussian witness. this is the generic `tail-hoeffding-bounded` import wrapper. it keeps the mathlib assumptions explicit: an a.e.-measurable real variable, an a.s. interval bound, and an equality between its integral and the supplied mean. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_abs_le_two_mul_sqrt_mul_exp_half_of_hasSubgaussianMGF","label":"integral_abs_le_two_mul_sqrt_mul_exp_half_of_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_abs_le_two_mul_sqrt_mul_exp_half_of_hasSubgaussianMGF","description":"A sub-Gaussian MGF controls the first absolute moment at the natural square-root scale. For a positive proxy, evaluate the MGF at the dimensionless tilt `t = 1 / sqrt c` and use `exp |x| <= exp x + exp (-x)`. The zero-proxy case is Mathlib's a.e.-zero theorem. The constant is intentionally simple; the route needs the `sqrt c` scaling rather than the sharp Gaussian constant.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-fa55266825d3","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3276,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:372"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_le_two_mul_sqrt_mul_exp_half_of_hasSubgaussianMGF {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (X : Omega -> Real) (c : NNReal) (hX : ProbabilityTheory.HasSubgaussianMGF X c mu) : integral mu (fun omega => |X omega|) <= 2 * Real.sqrt (c : Real) * Real.exp (1 / 2 : Real)","missing":[],"search":"integral_abs_le_two_mul_sqrt_mul_exp_half_of_hassubgaussianmgf banditrlproof.concentration.integral_abs_le_two_mul_sqrt_mul_exp_half_of_hassubgaussianmgf a sub-gaussian mgf controls the first absolute moment at the natural square-root scale. for a positive proxy, evaluate the mgf at the dimensionless tilt `t = 1 / sqrt c` and use `exp |x| <= exp x + exp (-x)`. the zero-proxy case is mathlib's a.e.-zero theorem. the constant is intentionally simple; the route needs the `sqrt c` scaling rather than the sharp gaussian constant. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_sq_le_four_mul_proxy_mul_exp_half_of_hasSubgaussianMGF","label":"integral_sq_le_four_mul_proxy_mul_exp_half_of_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_sq_le_four_mul_proxy_mul_exp_half_of_hasSubgaussianMGF","description":"A global sub-Gaussian MGF gives a conservative explicit second-moment bound. For positive proxy `c`, evaluate the MGF at `t = 1 / sqrt c` and use the quadratic exponential inequality on `|t * X|`. The constant is intentionally non-sharp; it avoids assuming an unavailable derivative-to-variance bridge.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-180607ca45ea","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3277,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:456"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_le_four_mul_proxy_mul_exp_half_of_hasSubgaussianMGF {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (X : Omega -> Real) (c : NNReal) (hX : ProbabilityTheory.HasSubgaussianMGF X c mu) : integral mu (fun omega => X omega ^ 2) <= 4 * (c : Real) * Real.exp (1 / 2 : Real)","missing":[],"search":"integral_sq_le_four_mul_proxy_mul_exp_half_of_hassubgaussianmgf banditrlproof.concentration.integral_sq_le_four_mul_proxy_mul_exp_half_of_hassubgaussianmgf a global sub-gaussian mgf gives a conservative explicit second-moment bound. for positive proxy `c`, evaluate the mgf at `t = 1 / sqrt c` and use the quadratic exponential inequality on `|t * x|`. the constant is intentionally non-sharp; it avoids assuming an unavailable derivative-to-variance bridge. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussian_sum_tail_of_iIndepFun","label":"subGaussian_sum_tail_of_iIndepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussian_sum_tail_of_iIndepFun","description":"Mathlib-backed one-sided tail bound for a finite sum of independent sub-Gaussian real random variables. This is the `TAIL-SUBGAUSS-SUM` import wrapper. It is a thin project-local surface over `ProbabilityTheory.HasSubgaussianMGF.measure_sum_ge_le_of_iIndepFun`.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-903ad8ef6ae8","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3278,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:545"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussian_sum_tail_of_iIndepFun {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) {Idx : Type v} {X : Idx -> Omega -> Real} (h_indep : ProbabilityTheory.iIndepFun X mu) {c : Idx -> NNReal} {s : Finset Idx} (h_subG : forall i, i ∈ s -> ProbabilityTheory.HasSubgaussianMGF (X i) (c i) mu) {eps : Real} (heps : 0 <= eps) : mu.real {omega | eps <= s.sum (fun i => X i omega)} <= Real.exp (-eps ^ 2 / (2 * ((s.sum c : NNReal) : Real)))","missing":[],"search":"subgaussian_sum_tail_of_iindepfun banditrlproof.concentration.subgaussian_sum_tail_of_iindepfun mathlib-backed one-sided tail bound for a finite sum of independent sub-gaussian real random variables. this is the `tail-subgauss-sum` import wrapper. it is a thin project-local surface over `probabilitytheory.hassubgaussianmgf.measure_sum_ge_le_of_iindepfun`. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussian_sum_tail_ennreal_of_iIndepFun","label":"subGaussian_sum_tail_ennreal_of_iIndepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussian_sum_tail_ennreal_of_iIndepFun","description":"ENNReal-valued version of `subGaussian_sum_tail_of_iIndepFun`. This is the `TAIL-SUBGAUSS-DIFF-SUM-IMPORT` boundary adapter used before an ETC-specific reward-difference specialization exists. The summands `X i` stay abstract; later leaves may instantiate them with centered non-best-minus-best exploration reward differences.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-f78e91b1a726","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3279,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:567"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussian_sum_tail_ennreal_of_iIndepFun {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {Idx : Type v} {X : Idx -> Omega -> Real} (h_indep : ProbabilityTheory.iIndepFun X mu) {c : Idx -> NNReal} {s : Finset Idx} (h_subG : forall i, i ∈ s -> ProbabilityTheory.HasSubgaussianMGF (X i) (c i) mu) {eps : Real} (heps : 0 <= eps) : mu {omega | eps <= s.sum (fun i => X i omega)} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((s.sum c : NNReal) : Real))))","missing":[],"search":"subgaussian_sum_tail_ennreal_of_iindepfun banditrlproof.concentration.subgaussian_sum_tail_ennreal_of_iindepfun ennreal-valued version of `subgaussian_sum_tail_of_iindepfun`. this is the `tail-subgauss-diff-sum-import` boundary adapter used before an etc-specific reward-difference specialization exists. the summands `x i` stay abstract; later leaves may instantiate them with centered non-best-minus-best exploration reward differences. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_tail_of_stronglyAdapted","label":"condSubGaussian_sum_tail_of_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_sum_tail_of_stronglyAdapted","description":"Mathlib-backed Azuma-Hoeffding tail bound for a finite prefix of a strongly adapted conditionally sub-Gaussian process. This is the `TAIL-COND-SUBGAUSS` import wrapper. It keeps Mathlib's contract visible: the zeroth summand is unconditionally sub-Gaussian, later summands are conditionally sub-Gaussian with respect to the previous filtration level, and the process is strongly adapted.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-05708fc1bbc9","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3280,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:596"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_sum_tail_of_stronglyAdapted {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsZeroOrProbabilityMeasure mu] {Y : Nat -> Omega -> Real} {cY : Nat -> NNReal} {F : Filtration Nat mOmega} (h_adapted : StronglyAdapted F Y) (h0 : ProbabilityTheory.HasSubgaussianMGF (Y 0) (cY 0) mu) (n : Nat) (h_subG : forall i, i < n - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (Y (i + 1)) (cY (i + 1)) mu) {eps : Real} (heps : 0 <= eps) : mu.real {omega | eps <= (Finset.range n).sum (fun i => Y i omega)} <= Real.exp (-eps ^ 2 / (2 * (((Finset.range n).sum cY : NNReal) : Real)))","missing":[],"search":"condsubgaussian_sum_tail_of_stronglyadapted banditrlproof.concentration.condsubgaussian_sum_tail_of_stronglyadapted mathlib-backed azuma-hoeffding tail bound for a finite prefix of a strongly adapted conditionally sub-gaussian process. this is the `tail-cond-subgauss` import wrapper. it keeps mathlib's contract visible: the zeroth summand is unconditionally sub-gaussian, later summands are conditionally sub-gaussian with respect to the previous filtration level, and the process is strongly adapted. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_tail_ennreal_of_stronglyAdapted","label":"condSubGaussian_sum_tail_ennreal_of_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_sum_tail_ennreal_of_stronglyAdapted","description":"ENNReal-valued version of `condSubGaussian_sum_tail_of_stronglyAdapted`. This boundary adapter is shaped for later ETC conditional/filtration routes: it preserves the same Mathlib hypotheses but returns an ordinary measure bound against the canonical exponential RHS wrapped in `ENNReal.ofReal`.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-9385363b5cb0","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3281,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:623"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_sum_tail_ennreal_of_stronglyAdapted {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsFiniteMeasure mu] [IsZeroOrProbabilityMeasure mu] {Y : Nat -> Omega -> Real} {cY : Nat -> NNReal} {F : Filtration Nat mOmega} (h_adapted : StronglyAdapted F Y) (h0 : ProbabilityTheory.HasSubgaussianMGF (Y 0) (cY 0) mu) (n : Nat) (h_subG : forall i, i < n - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (Y (i + 1)) (cY (i + 1)) mu) {eps : Real} (heps : 0 <= eps) : mu {omega | eps <= (Finset.range n).sum (fun i => Y i omega)} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * (((Finset.range n).sum cY : NNReal) : Real))))","missing":[],"search":"condsubgaussian_sum_tail_ennreal_of_stronglyadapted banditrlproof.concentration.condsubgaussian_sum_tail_ennreal_of_stronglyadapted ennreal-valued version of `condsubgaussian_sum_tail_of_stronglyadapted`. this boundary adapter is shaped for later etc conditional/filtration routes: it preserves the same mathlib hypotheses but returns an ordinary measure bound against the canonical exponential rhs wrapped in `ennreal.ofreal`. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussian_sum_abs_tail_ennreal_of_iIndepFun","label":"subGaussian_sum_abs_tail_ennreal_of_iIndepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussian_sum_abs_tail_ennreal_of_iIndepFun","description":"Two-sided ENNReal Hoeffding tail for a finite sum of independent sub-Gaussian random variables. Mathlib's `HasSubgaussianMGF.sum_of_iIndepFun` supplies the global sum MGF. Upper and negated lower tails are then combined by an outer-measure union, so no event-measurability hypothesis is needed.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-21f6eb1e5ca6","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3282,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:655"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussian_sum_abs_tail_ennreal_of_iIndepFun {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {Idx : Type v} {X : Idx -> Omega -> Real} (h_indep : ProbabilityTheory.iIndepFun X mu) {c : Idx -> NNReal} {s : Finset Idx} (h_subG : forall i, i ∈ s -> ProbabilityTheory.HasSubgaussianMGF (X i) (c i) mu) {eps : Real} (heps : 0 <= eps) : mu {omega | eps <= |s.sum (fun i => X i omega)|} <= ENNReal.ofReal (2 * Real.exp (-eps ^ 2 / (2 * ((s.sum c : NNReal) : Real))))","missing":[],"search":"subgaussian_sum_abs_tail_ennreal_of_iindepfun banditrlproof.concentration.subgaussian_sum_abs_tail_ennreal_of_iindepfun two-sided ennreal hoeffding tail for a finite sum of independent sub-gaussian random variables. mathlib's `hassubgaussianmgf.sum_of_iindepfun` supplies the global sum mgf. upper and negated lower tails are then combined by an outer-measure union, so no event-measurability hypothesis is needed. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_abs_tail_ennreal_of_stronglyAdapted","label":"condSubGaussian_sum_abs_tail_ennreal_of_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_sum_abs_tail_ennreal_of_stronglyAdapted","description":"Two-sided ENNReal Azuma-Hoeffding tail for a finite prefix of a strongly adapted conditionally sub-Gaussian process. Mathlib first upgrades the conditional increment witnesses to a global `HasSubgaussianMGF` witness for the finite sum. Applying its one-sided tail to the sum and its negation, then taking an outer-measure union bound, gives the factor-two absolute-deviation estimate. No event-measurability hypothesis…","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-b6ffd6e8b8c5","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3283,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:729"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_sum_abs_tail_ennreal_of_stronglyAdapted {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsFiniteMeasure mu] [IsZeroOrProbabilityMeasure mu] {Y : Nat -> Omega -> Real} {cY : Nat -> NNReal} {F : Filtration Nat mOmega} (h_adapted : StronglyAdapted F Y) (h0 : ProbabilityTheory.HasSubgaussianMGF (Y 0) (cY 0) mu) (n : Nat) (h_subG : forall i, i < n - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (Y (i + 1)) (cY (i + 1)) mu) {eps : Real} (heps : 0 <= eps) : mu {omega | eps <= |(Finset.range n).sum (fun i => Y i omega)|} <= ENNReal.ofReal (2 * Real.exp (-eps ^ 2 / (2 * (((Finset.range n).sum cY : NNReal) : Real))))","missing":[],"search":"condsubgaussian_sum_abs_tail_ennreal_of_stronglyadapted banditrlproof.concentration.condsubgaussian_sum_abs_tail_ennreal_of_stronglyadapted two-sided ennreal azuma-hoeffding tail for a finite prefix of a strongly adapted conditionally sub-gaussian process. mathlib first upgrades the conditional increment witnesses to a global `hassubgaussianmgf` witness for the finite sum. applying its one-sided tail to the sum and its negation, then taking an outer-measure union bound, gives the factor-two absolute-deviation estimate. no event-measurability hypothesis is needed. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussianSumConfidenceRadius","label":"subGaussianSumConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussianSumConfidenceRadius","description":"Two-sided fixed-horizon sub-Gaussian confidence radius for total proxy variance `variance` and failure budget `delta`.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-b52f5661427e","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3284,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:804"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianSumConfidenceRadius (variance : NNReal) (delta : Real) : Real","missing":[],"search":"subgaussiansumconfidenceradius banditrlproof.concentration.subgaussiansumconfidenceradius two-sided fixed-horizon sub-gaussian confidence radius for total proxy variance `variance` and failure budget `delta`. definition compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussianSumConfidenceRadius_nonneg","label":"subGaussianSumConfidenceRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussianSumConfidenceRadius_nonneg","description":"theorem subGaussianSumConfidenceRadius_nonneg (variance : NNReal) (delta : Real) : 0 <= subGaussianSumConfidenceRadius variance delta","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-bb4965e58b70","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3285,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:808"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussianSumConfidenceRadius_nonneg (variance : NNReal) (delta : Real) : 0 <= subGaussianSumConfidenceRadius variance delta","missing":[],"search":"subgaussiansumconfidenceradius_nonneg banditrlproof.concentration.subgaussiansumconfidenceradius_nonneg theorem subgaussiansumconfidenceradius_nonneg (variance : nnreal) (delta : real) : 0 <= subgaussiansumconfidenceradius variance delta theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussianSumConfidenceRadius_sq","label":"subGaussianSumConfidenceRadius_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussianSumConfidenceRadius_sq","description":"theorem subGaussianSumConfidenceRadius_sq (variance : NNReal) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (subGaussianSumConfidenceRadius variance delta) ^ 2 = 2 * ((variance : NNReal) : Real) * Real.log (2 / delta)","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-170ec585e0ff","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3286,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:813"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussianSumConfidenceRadius_sq (variance : NNReal) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (subGaussianSumConfidenceRadius variance delta) ^ 2 = 2 * ((variance : NNReal) : Real) * Real.log (2 / delta)","missing":[],"search":"subgaussiansumconfidenceradius_sq banditrlproof.concentration.subgaussiansumconfidenceradius_sq theorem subgaussiansumconfidenceradius_sq (variance : nnreal) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (subgaussiansumconfidenceradius variance delta) ^ 2 = 2 * ((variance : nnreal) : real) * real.log (2 / delta) theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.two_mul_exp_neg_subGaussianSumConfidenceRadius_sq_div_eq_delta","label":"two_mul_exp_neg_subGaussianSumConfidenceRadius_sq_div_eq_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.two_mul_exp_neg_subGaussianSumConfidenceRadius_sq_div_eq_delta","description":"theorem two_mul_exp_neg_subGaussianSumConfidenceRadius_sq_div_eq_delta (variance : NNReal) (delta : Real) (hvariance : 0 < ((variance : NNReal) : Real)) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 2 * Real.exp (-(subGaussianSumConfidenceRadius variance delta) ^ 2 / (2 * ((variance : NNReal) : Real))) = delta","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-bcc6f58b4768","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3287,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:825"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem two_mul_exp_neg_subGaussianSumConfidenceRadius_sq_div_eq_delta (variance : NNReal) (delta : Real) (hvariance : 0 < ((variance : NNReal) : Real)) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 2 * Real.exp (-(subGaussianSumConfidenceRadius variance delta) ^ 2 / (2 * ((variance : NNReal) : Real))) = delta","missing":[],"search":"two_mul_exp_neg_subgaussiansumconfidenceradius_sq_div_eq_delta banditrlproof.concentration.two_mul_exp_neg_subgaussiansumconfidenceradius_sq_div_eq_delta theorem two_mul_exp_neg_subgaussiansumconfidenceradius_sq_div_eq_delta (variance : nnreal) (delta : real) (hvariance : 0 < ((variance : nnreal) : real)) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 2 * real.exp (-(subgaussiansumconfidenceradius variance delta) ^ 2 / (2 * ((variance : nnreal) : real))) = delta theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussian_sum_abs_tail_ennreal_delta_of_iIndepFun","label":"subGaussian_sum_abs_tail_ennreal_delta_of_iIndepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussian_sum_abs_tail_ennreal_delta_of_iIndepFun","description":"Delta-calibrated two-sided confidence bound for a finite sum of independent sub-Gaussian random variables.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-5d18422b0a29","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3288,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:849"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussian_sum_abs_tail_ennreal_delta_of_iIndepFun {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {Idx : Type v} {X : Idx -> Omega -> Real} (h_indep : ProbabilityTheory.iIndepFun X mu) {c : Idx -> NNReal} {s : Finset Idx} (h_subG : forall i, i ∈ s -> ProbabilityTheory.HasSubgaussianMGF (X i) (c i) mu) (hvariance : 0 < (((s.sum c : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : mu {omega | subGaussianSumConfidenceRadius (s.sum c) delta <= |s.sum (fun i => X i omega)|} <= ENNReal.ofReal delta","missing":[],"search":"subgaussian_sum_abs_tail_ennreal_delta_of_iindepfun banditrlproof.concentration.subgaussian_sum_abs_tail_ennreal_delta_of_iindepfun delta-calibrated two-sided confidence bound for a finite sum of independent sub-gaussian random variables. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_abs_tail_ennreal_delta_of_stronglyAdapted","label":"condSubGaussian_sum_abs_tail_ennreal_delta_of_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_sum_abs_tail_ennreal_delta_of_stronglyAdapted","description":"Delta-calibrated two-sided Azuma-Hoeffding confidence bound for a finite prefix of a strongly adapted conditionally sub-Gaussian process. The positive-total-variance contract is required because the bad event is written with non-strict `radius <= |sum|`; at zero variance the zero-radius event contains the almost-sure equality path.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-ae5641023079","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3289,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:880"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_sum_abs_tail_ennreal_delta_of_stronglyAdapted {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsFiniteMeasure mu] [IsZeroOrProbabilityMeasure mu] {Y : Nat -> Omega -> Real} {cY : Nat -> NNReal} {F : Filtration Nat mOmega} (h_adapted : StronglyAdapted F Y) (h0 : ProbabilityTheory.HasSubgaussianMGF (Y 0) (cY 0) mu) (n : Nat) (h_subG : forall i, i < n - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (Y (i + 1)) (cY (i + 1)) mu) (hvariance : 0 < ((((Finset.range n).sum cY : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : mu {omega | subGaussianSumConfidenceRadius ((Finset.range n).sum cY) delta <= |(Finset.range n).sum (fun i => Y i omega)|} <= ENNReal.ofReal delta","missing":[],"search":"condsubgaussian_sum_abs_tail_ennreal_delta_of_stronglyadapted banditrlproof.concentration.condsubgaussian_sum_abs_tail_ennreal_delta_of_stronglyadapted delta-calibrated two-sided azuma-hoeffding confidence bound for a finite prefix of a strongly adapted conditionally sub-gaussian process. the positive-total-variance contract is required because the bad event is written with non-strict `radius <= |sum|`; at zero variance the zero-radius event contains the almost-sure equality path. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussianAverageConfidenceRadius","label":"subGaussianAverageConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussianAverageConfidenceRadius","description":"Two-sided confidence radius for the average of `samples` centered increments. The total proxy variance belongs to the corresponding sum and the division by `samples` performs only the deterministic sum-to-average conversion.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-87ff058d10fa","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3290,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:915"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianAverageConfidenceRadius (variance : NNReal) (samples : Nat) (delta : Real) : Real","missing":[],"search":"subgaussianaverageconfidenceradius banditrlproof.concentration.subgaussianaverageconfidenceradius two-sided confidence radius for the average of `samples` centered increments. the total proxy variance belongs to the corresponding sum and the division by `samples` performs only the deterministic sum-to-average conversion. definition compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussianAverageConfidenceRadius_nonneg","label":"subGaussianAverageConfidenceRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussianAverageConfidenceRadius_nonneg","description":"theorem subGaussianAverageConfidenceRadius_nonneg (variance : NNReal) (samples : Nat) (delta : Real) : 0 <= subGaussianAverageConfidenceRadius variance samples delta","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-93a213d8419e","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3291,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:919"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussianAverageConfidenceRadius_nonneg (variance : NNReal) (samples : Nat) (delta : Real) : 0 <= subGaussianAverageConfidenceRadius variance samples delta","missing":[],"search":"subgaussianaverageconfidenceradius_nonneg banditrlproof.concentration.subgaussianaverageconfidenceradius_nonneg theorem subgaussianaverageconfidenceradius_nonneg (variance : nnreal) (samples : nat) (delta : real) : 0 <= subgaussianaverageconfidenceradius variance samples delta theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_average_abs_tail_le_of_measure_sum_abs_tail","label":"measure_average_abs_tail_le_of_measure_sum_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_average_abs_tail_le_of_measure_sum_abs_tail","description":"Deterministic positive-sample-count transport from a two-sided sum-confidence bound to the corresponding average-confidence bound.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-26ba8b7428c2","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3292,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:930"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_average_abs_tail_le_of_measure_sum_abs_tail {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) (X : Omega -> Real) (variance : NNReal) (m : Nat) (hm : 0 < m) (delta : Real) (htail : mu {omega | subGaussianSumConfidenceRadius variance delta <= |X omega|} <= ENNReal.ofReal delta) : mu {omega | subGaussianAverageConfidenceRadius variance m delta <= |X omega / (m : Real)|} <= ENNReal.ofReal delta","missing":[],"search":"measure_average_abs_tail_le_of_measure_sum_abs_tail banditrlproof.concentration.measure_average_abs_tail_le_of_measure_sum_abs_tail deterministic positive-sample-count transport from a two-sided sum-confidence bound to the corresponding average-confidence bound. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_randomCount_average_abs_tail_le_of_measure_sum_abs_tail","label":"measure_randomCount_average_abs_tail_le_of_measure_sum_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_randomCount_average_abs_tail_le_of_measure_sum_abs_tail","description":"Positive random-count transport from a two-sided sum-confidence bound to the corresponding average-confidence bound. No measurability of `count` is needed: Mathlib measures are outer measures on arbitrary sets, and the proof is the pointwise inclusion obtained by multiplying through by the positive realized count.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-f38304f93e07","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3293,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:962"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_randomCount_average_abs_tail_le_of_measure_sum_abs_tail {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) (X : Omega -> Real) (count : Omega -> Nat) (variance : NNReal) (delta : Real) (htail : mu {omega | subGaussianSumConfidenceRadius variance delta <= |X omega|} <= ENNReal.ofReal delta) : mu {omega | 0 < count omega ∧ subGaussianAverageConfidenceRadius variance (count omega) delta <= |X omega / (count omega : Real)|} <= ENNReal.ofReal delta","missing":[],"search":"measure_randomcount_average_abs_tail_le_of_measure_sum_abs_tail banditrlproof.concentration.measure_randomcount_average_abs_tail_le_of_measure_sum_abs_tail positive random-count transport from a two-sided sum-confidence bound to the corresponding average-confidence bound. no measurability of `count` is needed: mathlib measures are outer measures on arbitrary sets, and the proof is the pointwise inclusion obtained by multiplying through by the positive realized count. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_positive_randomCount_event_le_sum_exactCount","label":"measure_positive_randomCount_event_le_sum_exactCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_positive_randomCount_event_le_sum_exactCount","description":"A positive random-count event is covered by its exact-count fibers up to a deterministic count ceiling. This is an outer-measure statement, so neither the count nor the fiber events need to be measurable.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-205d3f87ac0e","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3294,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:989"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_positive_randomCount_event_le_sum_exactCount {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) (count : Omega -> Nat) (maxCount : Nat) (bad : Nat -> Set Omega) (hcount_le : forall omega, count omega <= maxCount) : mu {omega | 0 < count omega ∧ omega ∈ bad (count omega)} <= (Finset.range maxCount).sum (fun i => mu {omega | count omega = i + 1 ∧ omega ∈ bad (i + 1)})","missing":[],"search":"measure_positive_randomcount_event_le_sum_exactcount banditrlproof.concentration.measure_positive_randomcount_event_le_sum_exactcount a positive random-count event is covered by its exact-count fibers up to a deterministic count ceiling. this is an outer-measure statement, so neither the count nor the fiber events need to be measurable. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.measure_positive_randomCount_event_le_of_exactCount_uniform","label":"measure_positive_randomCount_event_le_of_exactCount_uniform","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.measure_positive_randomCount_event_le_of_exactCount_uniform","description":"Equal-share finite peeling over every positive exact-count fiber.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-d83632071cea","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3295,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:1026"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_positive_randomCount_event_le_of_exactCount_uniform {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) (count : Omega -> Nat) (maxCount : Nat) (bad : Nat -> Set Omega) (hcount_le : forall omega, count omega <= maxCount) (hmaxCount : 0 < maxCount) (delta : Real) (_hdelta : 0 < delta) (hfiber : forall k, 0 < k -> k <= maxCount -> mu {omega | count omega = k ∧ omega ∈ bad k} <= ENNReal.ofReal (delta / (maxCount : Real))) : mu {omega | 0 < count omega ∧ omega ∈ bad (count omega)} <= ENNReal.ofReal delta","missing":[],"search":"measure_positive_randomcount_event_le_of_exactcount_uniform banditrlproof.concentration.measure_positive_randomcount_event_le_of_exactcount_uniform equal-share finite peeling over every positive exact-count fiber. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussian_average_abs_tail_ennreal_delta_of_iIndepFun","label":"subGaussian_average_abs_tail_ennreal_delta_of_iIndepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussian_average_abs_tail_ennreal_delta_of_iIndepFun","description":"Delta-calibrated two-sided confidence bound for the average of exactly `m` independent sub-Gaussian random variables indexed by a finite set.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-9ba3ba80650c","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3296,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:1065"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem subGaussian_average_abs_tail_ennreal_delta_of_iIndepFun {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {X : Nat -> Omega -> Real} {c : Nat -> NNReal} (h_indep : ProbabilityTheory.iIndepFun X mu) (m : Nat) (hm : 0 < m) (h_subG : forall i, i < m -> ProbabilityTheory.HasSubgaussianMGF (X i) (c i) mu) (hvariance : 0 < ((((Finset.range m).sum c : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : mu {omega | subGaussianAverageConfidenceRadius ((Finset.range m).sum c) m delta <= |((Finset.range m).sum (fun i => X i omega)) / (m : Real)|} <= ENNReal.ofReal delta","missing":[],"search":"subgaussian_average_abs_tail_ennreal_delta_of_iindepfun banditrlproof.concentration.subgaussian_average_abs_tail_ennreal_delta_of_iindepfun delta-calibrated two-sided confidence bound for the average of exactly `m` independent sub-gaussian random variables indexed by a finite set. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_average_abs_tail_ennreal_delta_of_stronglyAdapted","label":"condSubGaussian_average_abs_tail_ennreal_delta_of_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_average_abs_tail_ennreal_delta_of_stronglyAdapted","description":"Delta-calibrated two-sided confidence bound for the average of exactly `m` successor increments in a zero-initialized conditional sub-Gaussian process. The process prefix is `Finset.range (m + 1)`: slot zero is the deterministic initial value and slots `1, ..., m` are the `m` averaged increments. The proof is a positive-denominator event transport from the compiled sum confidence theorem.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-9c756815c918","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3297,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:1101"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_average_abs_tail_ennreal_delta_of_stronglyAdapted {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsFiniteMeasure mu] [IsZeroOrProbabilityMeasure mu] {Y : Nat -> Omega -> Real} {cY : Nat -> NNReal} {F : Filtration Nat mOmega} (h_adapted : StronglyAdapted F Y) (h0 : ProbabilityTheory.HasSubgaussianMGF (Y 0) (cY 0) mu) (m : Nat) (hm : 0 < m) (h_subG : forall i, i < (m + 1) - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (Y (i + 1)) (cY (i + 1)) mu) (hvariance : 0 < ((((Finset.range (m + 1)).sum cY : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : mu {omega | subGaussianAverageConfidenceRadius ((Finset.range (m + 1)).sum cY) m delta <= |((Finset.range (m + 1)).sum (fun i => Y i omega)) / (m : Real)|} <= ENNReal.ofReal delta","missing":[],"search":"condsubgaussian_average_abs_tail_ennreal_delta_of_stronglyadapted banditrlproof.concentration.condsubgaussian_average_abs_tail_ennreal_delta_of_stronglyadapted delta-calibrated two-sided confidence bound for the average of exactly `m` successor increments in a zero-initialized conditional sub-gaussian process. the process prefix is `finset.range (m + 1)`: slot zero is the deterministic initial value and slots `1, ..., m` are the `m` averaged increments. the proof is a positive-denominator event transport from the compiled sum confidence theorem. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_indicator_sum_tail_predictableVariance_fixedTilt","label":"condSubGaussian_indicator_sum_tail_predictableVariance_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_indicator_sum_tail_predictableVariance_fixedTilt","description":"Fixed-tilt one-sided tail for a conditionally sub-Gaussian process masked by conditioning-measurable events. The event retains the random cumulative masked proxy instead of charging every time step.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-849397bcb44e","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3298,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:1137"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_indicator_sum_tail_predictableVariance_fixedTilt {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsProbabilityMeasure mu] (F : Filtration Nat mOmega) (X : Nat -> Omega -> Real) (c : Nat -> NNReal) (s : Nat -> Set Omega) (hY : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => (s i).indicator (X i) omega)) (hV : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => (s i).indicator (fun _ => (((c i : NNReal) : Real))) omega)) (hs : forall i, @MeasurableSet Omega (F i) (s i)) (n : Nat) (h_subG : forall i, i < n - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (X i) (c i) mu) (tilt : Real) (htilt : 0 <= tilt) (threshold varianceBudget : Real) : mu {omega | threshold <= (Finset.range n).sum (fun t => match t with | 0 => 0 | i + 1 => (s i).indicator (X i) omega) ∧ (Finset.range…","missing":[],"search":"condsubgaussian_indicator_sum_tail_predictablevariance_fixedtilt banditrlproof.concentration.condsubgaussian_indicator_sum_tail_predictablevariance_fixedtilt fixed-tilt one-sided tail for a conditionally sub-gaussian process masked by conditioning-measurable events. the event retains the random cumulative masked proxy instead of charging every time step. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.subGaussianPredictableVarianceRadius","label":"subGaussianPredictableVarianceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.subGaussianPredictableVarianceRadius","description":"Delta radius for a two-sided masked conditionally sub-Gaussian sum under a deterministic budget on its random cumulative predictable proxy.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-0dc5e2397e0c","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3299,"meta":[["Kind","definition"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:1219"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def subGaussianPredictableVarianceRadius (varianceBudget delta : Real) : Real","missing":[],"search":"subgaussianpredictablevarianceradius banditrlproof.concentration.subgaussianpredictablevarianceradius delta radius for a two-sided masked conditionally sub-gaussian sum under a deterministic budget on its random cumulative predictable proxy. definition compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.condSubGaussian_indicator_sum_abs_tail_predictableVariance_delta","label":"condSubGaussian_indicator_sum_abs_tail_predictableVariance_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.condSubGaussian_indicator_sum_abs_tail_predictableVariance_delta","description":"Two-sided delta tail for a conditionally sub-Gaussian process with a conditioning-measurable mask and a random cumulative predictable proxy.","url":"../modules/banditrlproof-concentrationsubgaussian/index.html#decl-25153e180683","parent":"module:BanditRLProof.ConcentrationSubGaussian","order":3300,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationSubGaussian"],["Source","BanditRLProof/ConcentrationSubGaussian.lean:1227"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condSubGaussian_indicator_sum_abs_tail_predictableVariance_delta {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] {mu : Measure Omega} [IsProbabilityMeasure mu] (F : Filtration Nat mOmega) (X : Nat -> Omega -> Real) (c : Nat -> NNReal) (s : Nat -> Set Omega) (hY : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => (s i).indicator (X i) omega)) (hV : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => (s i).indicator (fun _ => (((c i : NNReal) : Real))) omega)) (hs : forall i, @MeasurableSet Omega (F i) (s i)) (n : Nat) (h_subG : forall i, i < n - 1 -> ProbabilityTheory.HasCondSubgaussianMGF (F i) (F.le i) (X i) (c i) mu) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : mu {omega | subGaussianPredictableVarianceRadius varianceBudget delta <= |(Finset.range n).sum (fun t => match…","missing":[],"search":"condsubgaussian_indicator_sum_abs_tail_predictablevariance_delta banditrlproof.concentration.condsubgaussian_indicator_sum_abs_tail_predictablevariance_delta two-sided delta tail for a conditionally sub-gaussian process with a conditioning-measurable mask and a random cumulative predictable proxy. theorem compiled","shard":"modules/86d9c3e7f592c6de.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.integral_positive_tail_le_two_sqrt","label":"integral_positive_tail_le_two_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.integral_positive_tail_le_two_sqrt","description":"theorem integral_positive_tail_le_two_sqrt (f : ℝ → ℝ) (hf : Measurable f) (hn : ∀ t, 0 ≤ f t) (h1 : ∀ t, f t ≤ 1) (c : ℝ) (hc : 0 < c) (ht : ∀ t, 0 < t → f t ≤ c/t^2) : ∫ t in Ioi 0, f t ≤ 2*sqrt c","url":"../modules/banditrlproof-concentrationtailintegration/index.html#decl-0503dc15bec8","parent":"module:BanditRLProof.ConcentrationTailIntegration","order":3301,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationTailIntegration"],["Source","BanditRLProof/ConcentrationTailIntegration.lean:9"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_positive_tail_le_two_sqrt (f : ℝ → ℝ) (hf : Measurable f) (hn : ∀ t, 0 ≤ f t) (h1 : ∀ t, f t ≤ 1) (c : ℝ) (hc : 0 < c) (ht : ∀ t, 0 < t → f t ≤ c/t^2) : ∫ t in Ioi 0, f t ≤ 2*sqrt c","missing":[],"search":"integral_positive_tail_le_two_sqrt banditrlproof.concentration.integral_positive_tail_le_two_sqrt theorem integral_positive_tail_le_two_sqrt (f : ℝ → ℝ) (hf : measurable f) (hn : ∀ t, 0 ≤ f t) (h1 : ∀ t, f t ≤ 1) (c : ℝ) (hc : 0 < c) (ht : ∀ t, 0 < t → f t ≤ c/t^2) : ∫ t in ioi 0, f t ≤ 2*sqrt c theorem compiled","shard":"modules/ffaf3f81fc32fd9d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.variance_chebyshev_tail","label":"variance_chebyshev_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.variance_chebyshev_tail","description":"Mathlib-backed Chebyshev tail bound for a real random variable with finite second moment. This is the real-variance `TAIL-VARIANCE-ROBUST` import wrapper. It keeps the Mathlib contract explicit: a finite measure, a real variable in `L^2`, and a strictly positive deviation radius.","url":"../modules/banditrlproof-concentrationvariance/index.html#decl-f57d9909b124","parent":"module:BanditRLProof.ConcentrationVariance","order":3302,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationVariance"],["Source","BanditRLProof/ConcentrationVariance.lean:26"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem variance_chebyshev_tail {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {X : Omega -> Real} (hX : MemLp X 2 mu) {eps : Real} (heps : 0 < eps) : mu {omega | eps <= |X omega - integral mu X|} <= ENNReal.ofReal (ProbabilityTheory.variance X mu / eps ^ 2)","missing":[],"search":"variance_chebyshev_tail banditrlproof.concentration.variance_chebyshev_tail mathlib-backed chebyshev tail bound for a real random variable with finite second moment. this is the real-variance `tail-variance-robust` import wrapper. it keeps the mathlib contract explicit: a finite measure, a real variable in `l^2`, and a strictly positive deviation radius. theorem compiled","shard":"modules/29402fc04ca866ee.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.evariance_chebyshev_tail","label":"evariance_chebyshev_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.evariance_chebyshev_tail","description":"Extended-real Chebyshev tail bound using Mathlib's `evariance` formulation. This version only requires almost-everywhere strong measurability; if the extended variance is infinite the bound is correspondingly non-informative.","url":"../modules/banditrlproof-concentrationvariance/index.html#decl-7a5668438223","parent":"module:BanditRLProof.ConcentrationVariance","order":3303,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationVariance"],["Source","BanditRLProof/ConcentrationVariance.lean:42"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem evariance_chebyshev_tail {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) {X : Omega -> Real} (hX : AEStronglyMeasurable X mu) {eps : NNReal} (heps : Ne eps 0) : mu {omega | (eps : Real) <= |X omega - integral mu X|} <= ProbabilityTheory.evariance X mu / (eps : ENNReal) ^ 2","missing":[],"search":"evariance_chebyshev_tail banditrlproof.concentration.evariance_chebyshev_tail extended-real chebyshev tail bound using mathlib's `evariance` formulation. this version only requires almost-everywhere strong measurability; if the extended variance is infinite the bound is correspondingly non-informative. theorem compiled","shard":"modules/29402fc04ca866ee.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.variance_sum_of_pairwise_indep","label":"variance_sum_of_pairwise_indep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.variance_sum_of_pairwise_indep","description":"Mathlib-backed variance additivity for a finite sum of pairwise independent real random variables. This is a bookkeeping wrapper used by finite-variance routes before a bandit-specific empirical-mean specialization exists.","url":"../modules/banditrlproof-concentrationvariance/index.html#decl-8f94722764f0","parent":"module:BanditRLProof.ConcentrationVariance","order":3304,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConcentrationVariance"],["Source","BanditRLProof/ConcentrationVariance.lean:59"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem variance_sum_of_pairwise_indep {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {Idx : Type v} {X : Idx -> Omega -> Real} {s : Finset Idx} (h_mem : forall i : Idx, Membership.mem s i -> MemLp (X i) 2 mu) (h_pairwise : Set.Pairwise ((s : Finset Idx) : Set Idx) (fun i j => ProbabilityTheory.IndepFun (X i) (X j) mu)) : ProbabilityTheory.variance (Finset.sum s X) mu = Finset.sum s (fun i => ProbabilityTheory.variance (X i) mu)","missing":[],"search":"variance_sum_of_pairwise_indep banditrlproof.concentration.variance_sum_of_pairwise_indep mathlib-backed variance additivity for a finite sum of pairwise independent real random variables. this is a bookkeeping wrapper used by finite-variance routes before a bandit-specific empirical-mean specialization exists. theorem compiled","shard":"modules/29402fc04ca866ee.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_integral_eq_zero","label":"condExp_eq_zero_of_condExpKernel_integral_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_integral_eq_zero","description":"Turn a trimmed-a.e. zero conditional-kernel integral into a true conditional mean-zero statement. This is the generic kernel-facing bridge for `COND-EXPECT-REWARD`. The hard future work is to prove `h_kernel_zero` from a trajectory/kernel law; this wrapper only connects that law-shaped hypothesis to Mathlib's `condExp`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-c8e5f84870a0","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3305,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:36"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExp_eq_zero_of_condExpKernel_integral_eq_zero {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (X : Omega -> Real) (h_integrable : Integrable X mu) (h_kernel_zero : Filter.Eventually (fun omega : Omega => integral (ProbabilityTheory.condExpKernel (Ω := Omega) (mΩ := mOmega) mu mcond omega) X = 0) (ae (mu.trim hm))) : Filter.EventuallyEq (ae mu) (@condExp Omega Real mcond mOmega _ _ _ mu X) (fun _omega : Omega => (0 : Real))","missing":[],"search":"condexp_eq_zero_of_condexpkernel_integral_eq_zero banditrlproof.conditionalexpectationreward.condexp_eq_zero_of_condexpkernel_integral_eq_zero turn a trimmed-a.e. zero conditional-kernel integral into a true conditional mean-zero statement. this is the generic kernel-facing bridge for `cond-expect-reward`. the hard future work is to prove `h_kernel_zero` from a trajectory/kernel law; this wrapper only connects that law-shaped hypothesis to mathlib's `condexp`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.hasSubgaussianMGF_mono_varianceProxy","label":"hasSubgaussianMGF_mono_varianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.hasSubgaussianMGF_mono_varianceProxy","description":"Monotonicity of the variance proxy in Mathlib's unconditional sub-Gaussian MGF predicate. This small helper lets history-selected kernel witnesses with proxy `c` feed a conditional theorem stated with a deterministic upper proxy `d`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-81b89dbf9901","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3306,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:80"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_mono_varianceProxy {Omega : Type u} [MeasurableSpace Omega] {mu : Measure Omega} {X : Omega -> Real} {c d : NNReal} (hcd : c <= d) (h : ProbabilityTheory.HasSubgaussianMGF X c mu) : ProbabilityTheory.HasSubgaussianMGF X d mu where","missing":[],"search":"hassubgaussianmgf_mono_varianceproxy banditrlproof.conditionalexpectationreward.hassubgaussianmgf_mono_varianceproxy monotonicity of the variance proxy in mathlib's unconditional sub-gaussian mgf predicate. this small helper lets history-selected kernel witnesses with proxy `c` feed a conditional theorem stated with a deterministic upper proxy `d`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_deterministic_of_measurable","label":"condExpKernel_map_eq_deterministic_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_deterministic_of_measurable","description":"A random variable measurable in the conditioning sigma-algebra is frozen by the conditional-expectation kernel. The conclusion is a kernel equality on any countably generated target, rather than a singleton reconstruction on a countable target. The proof maps the diagonal composition-product identity for `condExpKernel` through `X` and then uses Mathlib's a.e. uniqueness theorem for finite kernels.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-caba9e92a54c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3307,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:109"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_eq_deterministic_of_measurable {Omega : Type u} {Target : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [mTarget : MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (X : Omega -> Target) (hX : @Measurable Omega Target mcond mTarget X) : Filter.EventuallyEq (ae (mu.trim hm)) (@ProbabilityTheory.Kernel.map Omega Omega mcond mOmega Target mTarget (@ProbabilityTheory.condExpKernel Omega mOmega inferInstance mu inferInstance mcond) X) (@ProbabilityTheory.Kernel.deterministic Omega Target mcond mTarget X hX)","missing":[],"search":"condexpkernel_map_eq_deterministic_of_measurable banditrlproof.conditionalexpectationreward.condexpkernel_map_eq_deterministic_of_measurable a random variable measurable in the conditioning sigma-algebra is frozen by the conditional-expectation kernel. the conclusion is a kernel equality on any countably generated target, rather than a singleton reconstruction on a countable target. the proof maps the diagonal composition-product identity for `condexpkernel` through `x` and then uses mathlib's a.e. uniqueness theorem for finite kernels. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_dirac_of_measurable","label":"condExpKernel_map_eq_dirac_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_dirac_of_measurable","description":"Pointwise measure form of `condExpKernel_map_eq_deterministic_of_measurable`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-edb4ea344d22","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3308,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:145"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_eq_dirac_of_measurable {Omega : Type u} {Target : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [mTarget : MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (X : Omega -> Target) (hX : @Measurable Omega Target mcond mTarget X) : Filter.Eventually (fun omega => @Measure.map Omega Target mOmega mTarget X (@ProbabilityTheory.condExpKernel Omega mOmega inferInstance mu inferInstance mcond omega) = @Measure.dirac Target mTarget (X omega)) (ae (mu.trim hm))","missing":[],"search":"condexpkernel_map_eq_dirac_of_measurable banditrlproof.conditionalexpectationreward.condexpkernel_map_eq_dirac_of_measurable pointwise measure form of `condexpkernel_map_eq_deterministic_of_measurable`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_countable","label":"condExpKernel_map_eq_of_condDistrib_ae_eq_countable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_countable","description":"Convert a `condDistrib` law into a `condExpKernel` pushforward law on a countable target. Mathlib supplies eventwise equality between regular conditional distributions and `condExpKernel` pushforwards. This wrapper packages those singleton equalities into a measure equality when the target type is countable. It is a local bridge from canonical `condDistrib` trajectory laws toward the `condExpKernel` map-law consumer…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-355591f9c432","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3309,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:177"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_eq_of_condDistrib_ae_eq_countable {Omega : Type u} {Target : Type v} {Condition : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mTarget : MeasurableSpace Target] [StandardBorelSpace Target] [Nonempty Target] [MeasurableSingletonClass Target] [Countable Target] [mCondition : MeasurableSpace Condition] (mu : Measure Omega) [IsFiniteMeasure mu] (X : Omega -> Target) (Y : Omega -> Condition) (hX : @Measurable Omega Target mOmega mTarget X) (hY : @Measurable Omega Condition mOmega mCondition Y) (kernel : ProbabilityTheory.Kernel Condition Target) (hcond : Filter.EventuallyEq (ae (mu.map Y)) (ProbabilityTheory.condDistrib X Y mu) kernel) : Filter.Eventually (fun omega : Omega => @Measure.map Omega Target mOmega mTarget X (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (mCondition.comap Y) omega) = kernel (Y omega)) (ae mu)","missing":[],"search":"condexpkernel_map_eq_of_conddistrib_ae_eq_countable banditrlproof.conditionalexpectationreward.condexpkernel_map_eq_of_conddistrib_ae_eq_countable convert a `conddistrib` law into a `condexpkernel` pushforward law on a countable target. mathlib supplies eventwise equality between regular conditional distributions and `condexpkernel` pushforwards. this wrapper packages those singleton equalities into a measure equality when the target type is countable. it is a local bridge from canonical `conddistrib` trajectory laws toward the `condexpkernel` map-law consumers below; it does not itself construct the trajectory law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_countable_trim","label":"condExpKernel_map_eq_of_condDistrib_ae_eq_countable_trim","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_countable_trim","description":"Trim-a.e. form of the countable-target `condDistrib` to `condExpKernel.map` bridge. The ordinary bridge gives equality almost everywhere for the ambient measure. For each target singleton, both event-probability functions are measurable in the conditioning sigma-algebra. Mathlib's `ae_eq_trim_of_measurable` therefore upgrades those scalar equalities to the trimmed measure, after which countability reconstructs equal…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-3d8b2e7d127c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3310,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:251"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_eq_of_condDistrib_ae_eq_countable_trim {Omega : Type u} {Target : Type v} {Condition : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mTarget : MeasurableSpace Target] [StandardBorelSpace Target] [Nonempty Target] [MeasurableSingletonClass Target] [Countable Target] [mCondition : MeasurableSpace Condition] (mu : Measure Omega) [IsFiniteMeasure mu] (X : Omega -> Target) (Y : Omega -> Condition) (hX : @Measurable Omega Target mOmega mTarget X) (hY : @Measurable Omega Condition mOmega mCondition Y) (kernel : ProbabilityTheory.Kernel Condition Target) (hcond : Filter.EventuallyEq (ae (mu.map Y)) (ProbabilityTheory.condDistrib X Y mu) kernel) : Filter.Eventually (fun omega : Omega => @Measure.map Omega Target mOmega mTarget X (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (mCondition.comap Y) omega) = kernel (Y omega)) (…","missing":[],"search":"condexpkernel_map_eq_of_conddistrib_ae_eq_countable_trim banditrlproof.conditionalexpectationreward.condexpkernel_map_eq_of_conddistrib_ae_eq_countable_trim trim-a.e. form of the countable-target `conddistrib` to `condexpkernel.map` bridge. the ordinary bridge gives equality almost everywhere for the ambient measure. for each target singleton, both event-probability functions are measurable in the conditioning sigma-algebra. mathlib's `ae_eq_trim_of_measurable` therefore upgrades those scalar equalities to the trimmed measure, after which countability reconstructs equality of the pushed-forward measures. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_real_trim","label":"condExpKernel_map_eq_of_condDistrib_ae_eq_real_trim","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_real_trim","description":"Convert a real-valued `condDistrib` law into a trimmed `condExpKernel.map` law without assuming `Countable Real`. The proof first obtains equality on every rational left ray `Iic q`, upgrades those scalar equalities to the trimmed measure using conditioning-space measurability, and then reconstructs the full Borel measure by Mathlib's countable rational-ray pi-system induction.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-691b269f8afa","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3311,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:330"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_eq_of_condDistrib_ae_eq_real_trim {Omega : Type u} {Condition : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mCondition : MeasurableSpace Condition] (mu : Measure Omega) [IsFiniteMeasure mu] (X : Omega -> Real) (Y : Omega -> Condition) (hX : @Measurable Omega Real mOmega inferInstance X) (hY : @Measurable Omega Condition mOmega mCondition Y) (kernel : ProbabilityTheory.Kernel Condition Real) [ProbabilityTheory.IsMarkovKernel kernel] (hcond : Filter.EventuallyEq (ae (mu.map Y)) (ProbabilityTheory.condDistrib X Y mu) kernel) : Filter.Eventually (fun omega : Omega => @Measure.map Omega Real mOmega inferInstance X (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (mCondition.comap Y) omega) = kernel (Y omega)) (ae (mu.trim hY.comap_le))","missing":[],"search":"condexpkernel_map_eq_of_conddistrib_ae_eq_real_trim banditrlproof.conditionalexpectationreward.condexpkernel_map_eq_of_conddistrib_ae_eq_real_trim convert a real-valued `conddistrib` law into a trimmed `condexpkernel.map` law without assuming `countable real`. the proof first obtains equality on every rational left ray `iic q`, upgrades those scalar equalities to the trimmed measure using conditioning-space measurability, and then reconstructs the full borel measure by mathlib's countable rational-ray pi-system induction. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.eventuallyEq_const_of_map_eq_dirac","label":"eventuallyEq_const_of_map_eq_dirac","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.eventuallyEq_const_of_map_eq_dirac","description":"If a measurable pushforward law is a Dirac measure, the original random variable is a.e. constant. This is the small Mathlib bridge used to turn selected-action `condExpKernel.map = dirac ...` laws into the conditional a.e. action equality needed by the next-pair split-law route.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-400d0c38d5fa","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3312,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:476"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem eventuallyEq_const_of_map_eq_dirac {Omega : Type u} {Target : Type v} [mOmega : MeasurableSpace Omega] [mTarget : MeasurableSpace Target] [MeasurableSingletonClass Target] (mu : Measure Omega) (X : Omega -> Target) (x : Target) (hX : @Measurable Omega Target mOmega mTarget X) (hmap : @Measure.map Omega Target mOmega mTarget X mu = Measure.dirac x) : Filter.EventuallyEq (ae mu) X (fun _omega : Omega => x)","missing":[],"search":"eventuallyeq_const_of_map_eq_dirac banditrlproof.conditionalexpectationreward.eventuallyeq_const_of_map_eq_dirac if a measurable pushforward law is a dirac measure, the original random variable is a.e. constant. this is the small mathlib bridge used to turn selected-action `condexpkernel.map = dirac ...` laws into the conditional a.e. action equality needed by the next-pair split-law route. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq","label":"pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq","description":"Build a pair pushforward law from a deterministic action side and a reward pushforward law. This is the measure-level split helper behind the next-pair route, stripped of filtration-specific hypotheses. It is useful both for canonical trajectory laws and for later ambient laws once the relevant `condExpKernel` has been identified.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d54621215b9a","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3313,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:510"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (mu : Measure Omega) (nextAction : Omega -> Action) (nextReward : Omega -> Reward) (selectedAction : Action) (selectedReward : Measure Reward) (h_nextReward : @Measurable Omega Reward mOmega inferInstance nextReward) (h_action_ae : Filter.EventuallyEq (ae mu) nextAction (fun _omega : Omega => selectedAction)) (h_reward_map_eq : @Measure.map Omega Reward mOmega inferInstance nextReward mu = selectedReward) : @Measure.map Omega (Prod Action Reward) mOmega inferInstance (fun omega : Omega => (nextAction omega, nextReward omega)) mu = Measure.map (Prod.mk selectedAction) selectedReward","missing":[],"search":"pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq banditrlproof.conditionalexpectationreward.pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq build a pair pushforward law from a deterministic action side and a reward pushforward law. this is the measure-level split helper behind the next-pair route, stripped of filtration-specific hypotheses. it is useful both for canonical trajectory laws and for later ambient laws once the relevant `condexpkernel` has been identified. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq","label":"condExpKernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq","description":"Build a conditional next-pair product law from split action and reward laws. For each conditioning point, the action side freezes a.e. under the `condExpKernel`, while the reward side supplies the selected reward measure. This is the ambient `condExpKernel` form of `pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-69bbdf43d22b","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3314,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:564"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (i : Nat) (nextAction : Omega -> Action) (nextReward : Omega -> Reward) (selectedAction : Omega -> Action) (selectedReward : Omega -> Measure Reward) (h_nextReward : @Measurable Omega Reward mOmega inferInstance nextReward) (h_action_ae : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (F i) omega)) nextAction (fun _y : Omega => selectedAction omega)) (ae (mu.trim (F.le i)))) (h_reward_map_eq : Filter.Eventually (fun omega : Omega => @Measure.map Omega Reward mOmega inferInstance nextReward (@ProbabilityTheory.condExp…","missing":[],"search":"condexpkernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq banditrlproof.conditionalexpectationreward.condexpkernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq build a conditional next-pair product law from split action and reward laws. for each conditioning point, the action side freezes a.e. under the `condexpkernel`, while the reward side supplies the selected reward measure. this is the ambient `condexpkernel` form of `pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","label":"historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","description":"Canonical reward-only `trajMeasure` selected-reward law in `condExpKernel.map` form. This applies Mathlib's `Kernel.condDistrib_trajMeasure` theorem directly to `RewardKernel.historyStepKernelFamily`, then uses the local countable-target `condDistrib`-to-`condExpKernel.map` bridge. Unlike the action/reward pair trajectory route below, the canonical process here has only reward coordinates, so its finite prefix is al…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-1a55e2f50a18","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3315,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:623"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Reward)] [Nonempty Reward] [Nonempty ((t : Nat) -> Reward)] [MeasurableSingletonClass Reward] [Countable Reward] (mu0 : Measure Reward) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : let stepKernel := RewardKernel.hist…","missing":[],"search":"historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure canonical reward-only `trajmeasure` selected-reward law in `condexpkernel.map` form. this applies mathlib's `kernel.conddistrib_trajmeasure` theorem directly to `rewardkernel.historystepkernelfamily`, then uses the local countable-target `conddistrib`-to-`condexpkernel.map` bridge. unlike the action/reward pair trajectory route below, the canonical process here has only reward coordinates, so its finite prefix is already the reward history consumed by the policy, context, and state maps. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_trim","label":"historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_trim","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_trim","description":"Trim-a.e. canonical reward-only `trajMeasure` selected-reward law. This is the source-facing strengthening of `historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure`. It uses the trim-aware countable-target bridge, so the conclusion is stated on the finite reward-prefix conditioning sigma-algebra itself.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-931270664f57","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3316,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:706"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_trim {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Reward)] [Nonempty Reward] [Nonempty ((t : Nat) -> Reward)] [MeasurableSingletonClass Reward] [Countable Reward] (mu0 : Measure Reward) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : let stepKernel := RewardKernel…","missing":[],"search":"historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure_trim banditrlproof.conditionalexpectationreward.historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure_trim trim-a.e. canonical reward-only `trajmeasure` selected-reward law. this is the source-facing strengthening of `historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure`. it uses the trim-aware countable-target bridge, so the conclusion is stated on the finite reward-prefix conditioning sigma-algebra itself. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_identDistrib_trajMeasure_of_condDistrib","label":"historyStepKernelFamily_identDistrib_trajMeasure_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_identDistrib_trajMeasure_of_condDistrib","description":"Identify an ambient reward trace with the canonical reward-only trajectory from its initial marginal and successor conditional distributions. The complete-law step is supplied by the foundation-level reward-trace uniqueness theorem. This wrapper specializes its kernel family to `RewardKernel.historyStepKernelFamily`, exposing the recursive process contract needed by the ambient selected-reward transport below.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-38ee1bd431e8","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3317,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:804"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_identDistrib_trajMeasure_of_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} {Reward : Type*} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Reward) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> ((t : Nat) -> Reward)) (hreward : forall t : Nat, Measurable (fun omega : Omega => rew…","missing":[],"search":"historystepkernelfamily_identdistrib_trajmeasure_of_conddistrib banditrlproof.conditionalexpectationreward.historystepkernelfamily_identdistrib_trajmeasure_of_conddistrib identify an ambient reward trace with the canonical reward-only trajectory from its initial marginal and successor conditional distributions. the complete-law step is supplied by the foundation-level reward-trace uniqueness theorem. this wrapper specializes its kernel family to `rewardkernel.historystepkernelfamily`, exposing the recursive process contract needed by the ambient selected-reward transport below. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_of_identDistrib_trajMeasure_trim","label":"historyStepKernelFamily_selectedMeasure_condExpKernel_map_of_identDistrib_trajMeasure_trim","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_of_identDistrib_trajMeasure_trim","description":"Transport the canonical reward-only selected-reward law to an ambient reward trace with the same complete distribution. The `IdentDistrib` contract is strictly upstream of the conditional law used by the source layer. Composing it with the finite-prefix/next-coordinate map transports the relevant joint law from Mathlib's canonical `trajMeasure`. The canonical `condDistrib_trajMeasure` factorization and disintegratio…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-78c14e033a54","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3318,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:865"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_selectedMeasure_condExpKernel_map_of_identDistrib_trajMeasure_trim {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} {Reward : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Reward)] [Nonempty Omega] [Nonempty Reward] [Nonempty ((t : Nat) -> Reward)] [MeasurableSingletonClass Reward] [Countable Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Reward) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcon…","missing":[],"search":"historystepkernelfamily_selectedmeasure_condexpkernel_map_of_identdistrib_trajmeasure_trim banditrlproof.conditionalexpectationreward.historystepkernelfamily_selectedmeasure_condexpkernel_map_of_identdistrib_trajmeasure_trim transport the canonical reward-only selected-reward law to an ambient reward trace with the same complete distribution. the `identdistrib` contract is strictly upstream of the conditional law used by the source layer. composing it with the finite-prefix/next-coordinate map transports the relevant joint law from mathlib's canonical `trajmeasure`. the canonical `conddistrib_trajmeasure` factorization and disintegration uniqueness then recover the ambient conditional distribution, after which the countable-target trim bridge yields the requested `condexpkernel.map` law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` next-pair law in `condExpKernel.map` form. This is the pair-coordinate analogue of `actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure`: on Mathlib's canonical action/reward trajectory measure, conditioning on the finite pair prefix and pushing `condExpKernel` forward by the next pair coordinate recovers the configured history-step kernel.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-f906758e6972","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3319,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1008"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass (Prod Action Reward)] [Countable (Prod Action Reward)] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat,…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_condexpkernel_map_trajmeasure canonical `trajmeasure` next-pair law in `condexpkernel.map` form. this is the pair-coordinate analogue of `actionrewardhistorystepkernelfamily_reward_condexpkernel_map_trajmeasure`: on mathlib's canonical action/reward trajectory measure, conditioning on the finite pair prefix and pushing `condexpkernel` forward by the next pair coordinate recovers the configured history-step kernel. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_action_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_action_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_action_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` next-action law in `condExpKernel.map` form. This is the action-coordinate projection of `actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure`: on Mathlib's canonical action/reward trajectory measure, conditioning on the finite pair prefix and pushing `condExpKernel` forward by the next action coordinate recovers the Dirac law at the policy-selected action.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8b3423663ee2","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3320,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1091"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_action_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass (Prod Action Reward)] [Countable (Prod Action Reward)] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat…","missing":[],"search":"actionrewardhistorystepkernelfamily_action_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_action_condexpkernel_map_trajmeasure canonical `trajmeasure` next-action law in `condexpkernel.map` form. this is the action-coordinate projection of `actionrewardhistorystepkernelfamily_pair_condexpkernel_map_trajmeasure`: on mathlib's canonical action/reward trajectory measure, conditioning on the finite pair prefix and pushing `condexpkernel` forward by the next action coordinate recovers the dirac law at the policy-selected action. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_actionMarginal_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_actionMarginal_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_actionMarginal_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` action-marginal law in `condExpKernel.map` form via the action-coordinate `condDistrib` route. This is the action-side analogue of `actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure`: on the canonical action/reward trajectory measure, conditioning on the finite pair prefix and pushing `condExpKernel` forward by the next action coordinate recovers the `Prod.fst` marginal…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-cd2ad6810ca2","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3321,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1189"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_actionMarginal_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Action] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Action] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State)…","missing":[],"search":"actionrewardhistorystepkernelfamily_actionmarginal_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_actionmarginal_condexpkernel_map_trajmeasure canonical `trajmeasure` action-marginal law in `condexpkernel.map` form via the action-coordinate `conddistrib` route. this is the action-side analogue of `actionrewardhistorystepkernelfamily_reward_condexpkernel_map_trajmeasure`: on the canonical action/reward trajectory measure, conditioning on the finite pair prefix and pushing `condexpkernel` forward by the next action coordinate recovers the `prod.fst` marginal of the configured history-step kernel. the target countability contract is on `action`, not on the whole `(action × reward)` pair. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` selected-action law in `condExpKernel.map` form via the action-coordinate `condDistrib` route. This has the same selected-action Dirac conclusion as `actionRewardHistoryStepKernelFamily_action_condExpKernel_map_trajMeasure`, but obtains it by applying the countable-target `condExpKernel_map_eq_of_condDistrib_ae_eq_countable` bridge directly to the next action coordinate. Thus the target count…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a9523d415b99","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3322,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1274"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Action] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Action] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State)…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedaction_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_selectedaction_condexpkernel_map_trajmeasure canonical `trajmeasure` selected-action law in `condexpkernel.map` form via the action-coordinate `conddistrib` route. this has the same selected-action dirac conclusion as `actionrewardhistorystepkernelfamily_action_condexpkernel_map_trajmeasure`, but obtains it by applying the countable-target `condexpkernel_map_eq_of_conddistrib_ae_eq_countable` bridge directly to the next action coordinate. thus the target countability contract is on `action`, not on the whole `(action × reward)` pair. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_ae_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_ae_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_ae_trajMeasure","description":"Canonical `trajMeasure` selected-action law in conditional a.e. form. The previous theorem identifies the pushforward of the conditional kernel by the next action coordinate as a Dirac law. This wrapper turns that Dirac pushforward equality into the `Filter.EventuallyEq` shape required by the next-pair split-law route.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-002243248968","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3323,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1357"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_ae_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Action] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Action] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedaction_condexpkernel_ae_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_selectedaction_condexpkernel_ae_trajmeasure canonical `trajmeasure` selected-action law in conditional a.e. form. the previous theorem identifies the pushforward of the conditional kernel by the next action coordinate as a dirac law. this wrapper turns that dirac pushforward equality into the `filter.eventuallyeq` shape required by the next-pair split-law route. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_condExpKernel_map_trajMeasure","label":"actionRewardPartialTrajectoryKernel_extend_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` extension-map law in `condExpKernel.map` form. This pushes `actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure` through `History.extendPairHistorySucc`, yielding the one-step `RewardKernel.actionRewardPartialTrajectoryKernel` surface used by the extension-map consumers. It remains a canonical `trajMeasure` theorem; it does not transport an arbitrary ambient process.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-9c8c147dc48a","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3324,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1440"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_extend_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass (Prod Action Reward)] [Countable (Prod Action Reward)] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat…","missing":[],"search":"actionrewardpartialtrajectorykernel_extend_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_extend_condexpkernel_map_trajmeasure canonical `trajmeasure` extension-map law in `condexpkernel.map` form. this pushes `actionrewardhistorystepkernelfamily_pair_condexpkernel_map_trajmeasure` through `history.extendpairhistorysucc`, yielding the one-step `rewardkernel.actionrewardpartialtrajectorykernel` surface used by the extension-map consumers. it remains a canonical `trajmeasure` theorem; it does not transport an arbitrary ambient process. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` reward law in `condExpKernel.map` form. This specializes `condExpKernel_map_eq_of_condDistrib_ae_eq_countable` to the Mathlib Ionescu-Tulcea trajectory measure generated by `RewardKernel.actionRewardHistoryStepKernelFamily`. It converts the canonical next-reward `condDistrib` law into the `condExpKernel` pushforward-map shape used by the project-local conditional reward consumers.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d95ce1277c9c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3325,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1582"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Reward] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Reward] [Countable Reward] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontex…","missing":[],"search":"actionrewardhistorystepkernelfamily_reward_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_reward_condexpkernel_map_trajmeasure canonical `trajmeasure` reward law in `condexpkernel.map` form. this specializes `condexpkernel_map_eq_of_conddistrib_ae_eq_countable` to the mathlib ionescu-tulcea trajectory measure generated by `rewardkernel.actionrewardhistorystepkernelfamily`. it converts the canonical next-reward `conddistrib` law into the `condexpkernel` pushforward-map shape used by the project-local conditional reward consumers. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` selected-reward law in `condExpKernel.map` form. This rewrites `actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure` through `RewardKernel.actionRewardHistoryStepKernelFamily_reward_map`, yielding the selected context/action reward measure directly at the finite pair prefix.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-c851970b05c7","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3326,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1664"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Reward] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Reward] [Countable Reward] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State)…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure canonical `trajmeasure` selected-reward law in `condexpkernel.map` form. this rewrites `actionrewardhistorystepkernelfamily_reward_condexpkernel_map_trajmeasure` through `rewardkernel.actionrewardhistorystepkernelfamily_reward_map`, yielding the selected context/action reward measure directly at the finite pair prefix. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedMeasure_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` selected-reward law in `History.finitePairHistoryOfTrace` notation. This is the same canonical Ionescu-Tulcea conditional reward law as `actionRewardHistoryStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure`; the theorem only exposes the finite pair history prefix in the notation used by the bandit-history API.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d091a7b06fdc","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3327,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1735"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedMeasure_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Reward] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Reward] [Countable Reward] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedmeasure_finitepairhistoryoftrace_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_selectedmeasure_finitepairhistoryoftrace_condexpkernel_map_trajmeasure canonical `trajmeasure` selected-reward law in `history.finitepairhistoryoftrace` notation. this is the same canonical ionescu-tulcea conditional reward law as `actionrewardhistorystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure`; the theorem only exposes the finite pair history prefix in the notation used by the bandit-history api. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_rewardHistoryOfTrace_condExpKernel_map_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedMeasure_rewardHistoryOfTrace_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_rewardHistoryOfTrace_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` selected-reward law with reward-history context/state extractors. This specializes the finite-pair-history notation wrapper to pair-context and pair-state maps obtained by projecting finite pair histories to reward histories. It is still a canonical Ionescu-Tulcea theorem, not an ambient `History.historyFiltrationSucc` transport result.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-cb1a4d002bc8","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3328,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1808"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedMeasure_rewardHistoryOfTrace_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace (Prod Action Rat)] [StandardBorelSpace Rat] [StandardBorelSpace ((t : Nat) -> Prod Action Rat)] [Nonempty (Prod Action Rat)] [Nonempty Rat] [Nonempty ((t : Nat) -> Prod Action Rat)] [MeasurableSingletonClass Rat] [Countable Rat] (mu0 : Measure (Prod Action Rat)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Mea…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedmeasure_rewardhistoryoftrace_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_selectedmeasure_rewardhistoryoftrace_condexpkernel_map_trajmeasure canonical `trajmeasure` selected-reward law with reward-history context/state extractors. this specializes the finite-pair-history notation wrapper to pair-context and pair-state maps obtained by projecting finite pair histories to reward histories. it is still a canonical ionescu-tulcea theorem, not an ambient `history.historyfiltrationsucc` transport result. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure_of_selectedAction_ae_selectedMeasure","label":"actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure_of_selectedAction_ae_selectedMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure_of_selectedAction_ae_selectedMeasure","description":"Canonical `trajMeasure` next-pair law rebuilt from split inputs. This theorem verifies the split route on Mathlib's canonical Ionescu-Tulcea trajectory measure: the selected-action conditional a.e. law and the selected-reward pushforward law combine to recover the full `RewardKernel.actionRewardHistoryStepKernelFamily` next-pair law. Unlike the direct pair-coordinate `condDistrib` bridge, this route only needs separ…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-b1513f2915b4","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3329,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:1919"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure_of_selectedAction_ae_selectedMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Action] [StandardBorelSpace Reward] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty Action] [Nonempty Reward] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass Action] [Countable Action] [MeasurableSingletonClass Reward] [Countable Reward] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_condexpkernel_map_trajmeasure_of_selectedaction_ae_selectedmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_condexpkernel_map_trajmeasure_of_selectedaction_ae_selectedmeasure canonical `trajmeasure` next-pair law rebuilt from split inputs. this theorem verifies the split route on mathlib's canonical ionescu-tulcea trajectory measure: the selected-action conditional a.e. law and the selected-reward pushforward law combine to recover the full `rewardkernel.actionrewardhistorystepkernelfamily` next-pair law. unlike the direct pair-coordinate `conddistrib` bridge, this route only needs separate countability of `action` and `reward`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_of_condExpKernel_map_eq","label":"hasCondSubgaussianMGF_of_condExpKernel_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_of_condExpKernel_map_eq","description":"Generic `condExpKernel` map-law consumer for conditional sub-Gaussianity. If the conditional kernel pushed forward by `X` is trim-a.e. a target measure whose identity random variable is sub-Gaussian with deterministic proxy `c`, then `X` is conditionally sub-Gaussian. The target MGF bound also supplies the global exponential-integrability field: `Measure.integrable_comp_iff` combines the target laws' pointwise integ…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d8857fbfe113","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3330,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2037"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasCondSubgaussianMGF_of_condExpKernel_map_eq {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (X : Omega -> Real) (c : NNReal) (hX : @Measurable Omega Real mOmega inferInstance X) (target : Omega -> Measure Real) (h_kernel_map_eq : Filter.Eventually (fun omega : Omega => @Measure.map Omega Real mOmega inferInstance X (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ mcond omega) = target omega) (ae (mu.trim hm))) (h_target_subG : Filter.Eventually (fun omega : Omega => ProbabilityTheory.HasSubgaussianMGF (fun z : Real => z) c (target omega)) (ae (mu.trim hm))) : ProbabilityTheory.HasCondSubgaussianMGF mcond hm X c mu","missing":[],"search":"hascondsubgaussianmgf_of_condexpkernel_map_eq banditrlproof.conditionalexpectationreward.hascondsubgaussianmgf_of_condexpkernel_map_eq generic `condexpkernel` map-law consumer for conditional sub-gaussianity. if the conditional kernel pushed forward by `x` is trim-a.e. a target measure whose identity random variable is sub-gaussian with deterministic proxy `c`, then `x` is conditionally sub-gaussian. the target mgf bound also supplies the global exponential-integrability field: `measure.integrable_comp_iff` combines the target laws' pointwise integrability with their common deterministic mgf bound over the finite trim measure. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_centered_of_condExpKernel_map_eq","label":"hasCondSubgaussianMGF_centered_of_condExpKernel_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_centered_of_condExpKernel_map_eq","description":"Center a real-valued conditional reward law by a quantity measurable in the conditioning sigma-algebra. The conditional kernel freezes the center at the conditioning point. The resulting residual pushforward is therefore the centered pushforward of the supplied reward law, which can be consumed by the generic conditional sub-Gaussian map-law theorem.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-86f759dc10a3","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3331,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2120"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasCondSubgaussianMGF_centered_of_condExpKernel_map_eq {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (X center : Omega -> Real) (c : NNReal) (hX : @Measurable Omega Real mOmega inferInstance X) (hcenter : @Measurable Omega Real mcond inferInstance center) (target : Omega -> Measure Real) (h_kernel_map_eq : Filter.Eventually (fun omega : Omega => @Measure.map Omega Real mOmega inferInstance X (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ mcond omega) = target omega) (ae (mu.trim hm))) (h_target_subG : Filter.Eventually (fun omega : Omega => ProbabilityTheory.HasSubgaussianMGF (fun z : Real => z - center omega) c (target omega)) (ae (mu.trim hm))) : ProbabilityTheory.HasCondSubgaussianMGF mcond hm (fun omega => X omega - center omega) c mu","missing":[],"search":"hascondsubgaussianmgf_centered_of_condexpkernel_map_eq banditrlproof.conditionalexpectationreward.hascondsubgaussianmgf_centered_of_condexpkernel_map_eq center a real-valued conditional reward law by a quantity measurable in the conditioning sigma-algebra. the conditional kernel freezes the center at the conditioning point. the resulting residual pushforward is therefore the centered pushforward of the supplied reward law, which can be consumed by the generic conditional sub-gaussian map-law theorem. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_congr_measurableSpace","label":"hasCondSubgaussianMGF_congr_measurableSpace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_congr_measurableSpace","description":"Transport conditional sub-Gaussianity across equal conditioning measurable spaces. The two inclusion proofs are propositionally irrelevant.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a2d3e179df7c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3332,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2197"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasCondSubgaussianMGF_congr_measurableSpace {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond mcond' : MeasurableSpace Omega) (hm : mcond <= mOmega) (hm' : mcond' <= mOmega) (hspaces : mcond = mcond') (X : Omega -> Real) (c : NNReal) (h : ProbabilityTheory.HasCondSubgaussianMGF (mΩ := mOmega) mcond hm X c mu) : ProbabilityTheory.HasCondSubgaussianMGF (mΩ := mOmega) mcond' hm' X c mu","missing":[],"search":"hascondsubgaussianmgf_congr_measurablespace banditrlproof.conditionalexpectationreward.hascondsubgaussianmgf_congr_measurablespace transport conditional sub-gaussianity across equal conditioning measurable spaces. the two inclusion proofs are propositionally irrelevant. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_condExpKernel_integral_eq_zero","label":"centeredReward_succ_condExp_eq_zero_of_condExpKernel_integral_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_condExpKernel_integral_eq_zero","description":"Centered-reward specialization of `condExp_eq_zero_of_condExpKernel_integral_eq_zero`. The statement matches the succ-indexed shape used by Mathlib's conditional tail API: the reward at `i + 1` is conditioned on filtration level `i`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-97041e39959c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3333,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2220"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_condExpKernel_integral_eq_zero {Omega : Type u} {K : Nat} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (model : FiniteBanditModel K) (reward : Omega -> RewardTrace Rat) (b : Fin K) (i : Nat) (h_integrable : Integrable (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real))) mu) (h_kernel_zero : Filter.Eventually (fun omega : Omega => integral (ProbabilityTheory.condExpKernel (Ω := Omega) (mΩ := mOmega) mu (F i) omega) (fun y : Omega => (((reward y (i + 1) - model.mean b : Rat) : Real))) = 0) (ae (mu.trim (F.le i)))) : Filter.EventuallyEq (ae mu) (@condExp Omega Real (F i) mOmega _ _ _ mu (fun omega : Omega => (((reward omega (i + 1) - model.mean b : Rat) : Real)))) (fun _omega : Omega => (0 : Real))","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_condexpkernel_integral_eq_zero banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_condexpkernel_integral_eq_zero centered-reward specialization of `condexp_eq_zero_of_condexpkernel_integral_eq_zero`. the statement matches the succ-indexed shape used by mathlib's conditional tail api: the reward at `i + 1` is conditioned on filtration level `i`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_integral_eq_historyStepKernel_centeredReward","label":"condExp_eq_zero_of_condExpKernel_integral_eq_historyStepKernel_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_integral_eq_historyStepKernel_centeredReward","description":"Consumer bridge from a trajectory-law-shaped conditional-kernel identification to ordinary conditional mean zero. The hypothesis `h_kernel_eq` is intentionally explicit: it is the future `condExpKernel`/history-step reward-law identification, already reduced to the centered integral shape needed here. The theorem then uses the compiled `RewardKernel.historyStepKernelFamily_centeredReward_integral_eq_zero` leaf and t…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-c0b16555e3af","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3334,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2266"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExp_eq_zero_of_condExpKernel_integral_eq_historyStepKernel_centeredReward {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (n…","missing":[],"search":"condexp_eq_zero_of_condexpkernel_integral_eq_historystepkernel_centeredreward banditrlproof.conditionalexpectationreward.condexp_eq_zero_of_condexpkernel_integral_eq_historystepkernel_centeredreward consumer bridge from a trajectory-law-shaped conditional-kernel identification to ordinary conditional mean zero. the hypothesis `h_kernel_eq` is intentionally explicit: it is the future `condexpkernel`/history-step reward-law identification, already reduced to the centered integral shape needed here. the theorem then uses the compiled `rewardkernel.historystepkernelfamily_centeredreward_integral_eq_zero` leaf and the generic `condexpkernel` bridge above. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_integral_eq","label":"centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_integral_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_integral_eq","description":"Succ-indexed selected-reward specialization of `condExp_eq_zero_of_condExpKernel_integral_eq_historyStepKernel_centeredReward`. The centered variable uses the history-selected context/action mean. The only law-identification input is `h_kernel_eq`; constructing that equality from a `partialTraj` trajectory measure is still a separate missing leaf.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-255323991c31","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3335,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2324"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_integral_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (reward : Omega ->…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_integral_eq banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_integral_eq succ-indexed selected-reward specialization of `condexp_eq_zero_of_condexpkernel_integral_eq_historystepkernel_centeredreward`. the centered variable uses the history-selected context/action mean. the only law-identification input is `h_kernel_eq`; constructing that equality from a `partialtraj` trajectory measure is still a separate missing leaf. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_map_eq_historyStepKernel_centeredReward","label":"condExp_eq_zero_of_condExpKernel_map_eq_historyStepKernel_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_map_eq_historyStepKernel_centeredReward","description":"Map-law consumer for the history-step conditional mean-zero route. Compared with the integral-equality consumer above, this theorem assumes a more structural reward-law identification: trim-a.e., the conditional kernel pushed forward by the next-reward coordinate is the corresponding `historyStepKernelFamily` reward law. The separate `h_kernel_X_eq` hypothesis records the usual \"past is frozen under conditioning\" ob…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-fbc0d9875703","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3336,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2417"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExp_eq_zero_of_condExpKernel_map_eq_historyStepKernel_centeredReward {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (n : Na…","missing":[],"search":"condexp_eq_zero_of_condexpkernel_map_eq_historystepkernel_centeredreward banditrlproof.conditionalexpectationreward.condexp_eq_zero_of_condexpkernel_map_eq_historystepkernel_centeredreward map-law consumer for the history-step conditional mean-zero route. compared with the integral-equality consumer above, this theorem assumes a more structural reward-law identification: trim-a.e., the conditional kernel pushed forward by the next-reward coordinate is the corresponding `historystepkernelfamily` reward law. the separate `h_kernel_x_eq` hypothesis records the usual \"past is frozen under conditioning\" obligation needed to replace the actual target variable by the centered function with the outer history fixed at `omega`. this still does not construct the `partialtraj`/`condexpkernel` identity; it only consumes a map-level version of that identity. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_of_condExpKernel_map_eq_historyStepKernel_centeredReward","label":"hasCondSubgaussianMGF_of_condExpKernel_map_eq_historyStepKernel_centeredReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_of_condExpKernel_map_eq_historyStepKernel_centeredReward","description":"Map-law consumer for the history-step conditional sub-Gaussian route. This is the MGF analogue of `condExp_eq_zero_of_condExpKernel_map_eq_historyStepKernel_centeredReward`. It keeps two contracts explicit: `h_kernel_X_eq` freezes the history-dependent centering under the conditional kernel, and `h_variance_le` bounds the history-selected variance proxy by the deterministic proxy `c` required by Mathlib's `HasCondSu…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a9fe0cacf06a","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3337,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2538"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem hasCondSubgaussianMGF_of_condExpKernel_map_eq_historyStepKernel_centeredReward {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy…","missing":[],"search":"hascondsubgaussianmgf_of_condexpkernel_map_eq_historystepkernel_centeredreward banditrlproof.conditionalexpectationreward.hascondsubgaussianmgf_of_condexpkernel_map_eq_historystepkernel_centeredreward map-law consumer for the history-step conditional sub-gaussian route. this is the mgf analogue of `condexp_eq_zero_of_condexpkernel_map_eq_historystepkernel_centeredreward`. it keeps two contracts explicit: `h_kernel_x_eq` freezes the history-dependent centering under the conditional kernel, and `h_variance_le` bounds the history-selected variance proxy by the deterministic proxy `c` required by mathlib's `hascondsubgaussianmgf`. exponential integrability is generated by the generic integrated target-law transfer above. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq","label":"centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq","description":"Succ-indexed selected-reward specialization of the conditional sub-Gaussian map-law consumer. This is the MGF analogue of `centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq`. It keeps measurability and deterministic variance domination explicit; the kernel-side sub-Gaussian law now also discharges exponential integrability via the integrated target-law transfer.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-7d1d4b95b335","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3338,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2708"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (reward : Omeg…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_historystepkernelfamily_condexpkernel_map_eq banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_historystepkernelfamily_condexpkernel_map_eq succ-indexed selected-reward specialization of the conditional sub-gaussian map-law consumer. this is the mgf analogue of `centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq`. it keeps measurability and deterministic variance domination explicit; the kernel-side sub-gaussian law now also discharges exponential integrability via the integrated target-law transfer. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_event_real_eq_indicator_of_measurableSet","label":"condExpKernel_event_real_eq_indicator_of_measurableSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_event_real_eq_indicator_of_measurableSet","description":"Event-level frozen-past canary for conditional-expectation kernels. If an event is measurable in the conditioning sigma-algebra, then trim-a.e. the conditional kernel assigns its real mass as the event indicator. This is the 0/1 event support fact used to freeze countable-valued past summaries below.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-33e23e39cc8e","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3339,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2815"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_event_real_eq_indicator_of_measurableSet {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (s : Set Omega) (hs : @MeasurableSet Omega mcond s) : Filter.Eventually (fun omega : Omega => (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ mcond omega).real s = Set.indicator s (fun _omega : Omega => (1 : Real)) omega) (ae (mu.trim hm))","missing":[],"search":"condexpkernel_event_real_eq_indicator_of_measurableset banditrlproof.conditionalexpectationreward.condexpkernel_event_real_eq_indicator_of_measurableset event-level frozen-past canary for conditional-expectation kernels. if an event is measurable in the conditioning sigma-algebra, then trim-a.e. the conditional kernel assigns its real mass as the event indicator. this is the 0/1 event support fact used to freeze countable-valued past summaries below. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_ae_eq_const_of_countable_measurable","label":"condExpKernel_ae_eq_const_of_countable_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.condExpKernel_ae_eq_const_of_countable_measurable","description":"Countable-valued frozen-past theorem for conditional-expectation kernels. Any countable-valued random variable measurable in the conditioning sigma-algebra is trim-a.e. constant under the corresponding conditional kernel. This packages the event-level 0/1 fact over all singleton fibers.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-b44246e66dab","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3340,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2853"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_ae_eq_const_of_countable_measurable {Omega : Type u} {A : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace A] [MeasurableSingletonClass A] [Countable A] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hm : mcond <= mOmega) (Y : Omega -> A) (hY : @Measurable Omega A mcond inferInstance Y) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ mcond omega)) Y (fun _y : Omega => Y omega)) (ae (mu.trim hm))","missing":[],"search":"condexpkernel_ae_eq_const_of_countable_measurable banditrlproof.conditionalexpectationreward.condexpkernel_ae_eq_const_of_countable_measurable countable-valued frozen-past theorem for conditional-expectation kernels. any countable-valued random variable measurable in the conditioning sigma-algebra is trim-a.e. constant under the corresponding conditional kernel. this packages the event-level 0/1 fact over all singleton fibers. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_prefix_condExpKernel_map_trajMeasure","label":"actionRewardPartialTrajectoryKernel_prefix_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_prefix_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` full finite-prefix law in `condExpKernel.map` form. The earlier canonical extension-map theorem states the law for the deterministic extension of the frozen prefix by the random next pair. This wrapper uses the standard `condExpKernel` frozen-prefix property for the conditioning map `Preorder.frestrictLe n`, then rewrites that extension map back to the full `Preorder.frestrictLe (n + 1)` pref…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8059ab31ff24","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3341,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:2929"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_prefix_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass (Prod Action Reward)] [Countable (Prod Action Reward)] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat…","missing":[],"search":"actionrewardpartialtrajectorykernel_prefix_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_prefix_condexpkernel_map_trajmeasure canonical `trajmeasure` full finite-prefix law in `condexpkernel.map` form. the earlier canonical extension-map theorem states the law for the deterministic extension of the frozen prefix by the random next pair. this wrapper uses the standard `condexpkernel` frozen-prefix property for the conditioning map `preorder.frestrictle n`, then rewrites that extension map back to the full `preorder.frestrictle (n + 1)` prefix under the conditional kernel. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","label":"actionRewardPartialTrajectoryKernel_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","description":"Canonical `trajMeasure` full finite-pair-history law in the project's concrete `History.finitePairHistoryOfTrace` notation. This is a notation-alignment wrapper around `actionRewardPartialTrajectoryKernel_extend_condExpKernel_map_trajMeasure`: the ambient theorem-card still needs a transport from canonical `trajMeasure` to an arbitrary generated process.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-b10dbbd9f317","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3342,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3112"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure {Context : Type x} {State : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace ((t : Nat) -> Prod Action Reward)] [Nonempty (Prod Action Reward)] [Nonempty ((t : Nat) -> Prod Action Reward)] [MeasurableSingletonClass (Prod Action Reward)] [Countable (Prod Action Reward)] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontex…","missing":[],"search":"actionrewardpartialtrajectorykernel_finitepairhistoryoftrace_condexpkernel_map_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_finitepairhistoryoftrace_condexpkernel_map_trajmeasure canonical `trajmeasure` full finite-pair-history law in the project's concrete `history.finitepairhistoryoftrace` notation. this is a notation-alignment wrapper around `actionrewardpartialtrajectorykernel_extend_condexpkernel_map_trajmeasure`: the ambient theorem-card still needs a transport from canonical `trajmeasure` to an arbitrary generated process. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_of_measurable_of_policy_eq","label":"action_condExpKernel_ae_eq_policy_of_measurable_of_policy_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_of_measurable_of_policy_eq","description":"Action-freezing hookup for the next-pair split-law route. If the next action is measurable at filtration level `F i`, the conditional kernel freezes it. If that frozen action is also trim-a.e. the policy-selected action for the finite pair history, this supplies exactly the conditional a.e. action equality consumed by `actionRewardHistoryStepKernelFamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq`. It does not…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-23fb34996580","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3343,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3210"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem action_condExpKernel_ae_eq_policy_of_measurable_of_policy_eq {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (action : Omega -> ActionTrace Action) (i : Nat) (pairHistory : Omega -> ((j : Finset.Iic i) -> Prod Action Rat)) (h_action_next_meas : @Measurable Omega Action (F i) inferInstance (fun omega : Omega => action omega (i + 1))) (h_action_policy_eq : Filter.Eventually (fun omega : Omega => action omega (i + 1) = (policy i).action (pairState i (pairHistory omega))) (ae (mu.trim (F.le i)))) : Filter.Eventually (fun o…","missing":[],"search":"action_condexpkernel_ae_eq_policy_of_measurable_of_policy_eq banditrlproof.conditionalexpectationreward.action_condexpkernel_ae_eq_policy_of_measurable_of_policy_eq action-freezing hookup for the next-pair split-law route. if the next action is measurable at filtration level `f i`, the conditional kernel freezes it. if that frozen action is also trim-a.e. the policy-selected action for the finite pair history, this supplies exactly the conditional a.e. action equality consumed by `actionrewardhistorystepkernelfamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq`. it does not prove the predictability or policy-generation equality hypotheses. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_of_pairHistory_measurable_of_action_eq","label":"action_condExpKernel_ae_eq_policy_of_pairHistory_measurable_of_action_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_of_pairHistory_measurable_of_action_eq","description":"Pair-history measurability hookup for the action side of the next-pair split. If the finite pair history is visible at `F i`, the pair-state extractor is measurable, and the next action is pointwise the policy action selected from that pair state, then the previous action-freezing theorem supplies the conditional a.e. action equality consumed by the split-law builder.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-9f5bbec7e3a4","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3344,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3267"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem action_condExpKernel_ae_eq_policy_of_pairHistory_measurable_of_action_eq {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairState : forall n : Nat, Measurable (pairState n)) (action : Omega -> ActionTrace Action) (i : Nat) (pairHistory : Omega -> ((j : Finset.Iic i) -> Prod Action Rat)) (h_pairHistory_meas : @Measurable Omega ((j : Finset.Iic i) -> Prod Action Rat) (F i) inferInstance pairHistory) (h_action_eq : (fun omega : Omega => action omega (i + 1)) = (fun omega : Omega => (policy i).action (pairState i (pairH…","missing":[],"search":"action_condexpkernel_ae_eq_policy_of_pairhistory_measurable_of_action_eq banditrlproof.conditionalexpectationreward.action_condexpkernel_ae_eq_policy_of_pairhistory_measurable_of_action_eq pair-history measurability hookup for the action side of the next-pair split. if the finite pair history is visible at `f i`, the pair-state extractor is measurable, and the next action is pointwise the policy action selected from that pair state, then the previous action-freezing theorem supplies the conditional a.e. action equality consumed by the split-law builder. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_action_eq","label":"action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_action_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_action_eq","description":"Generated-history specialization of the action side of the next-pair split. For `History.historyFiltrationSucc`, the finite action/reward pair prefix up to `i` is visible. Therefore a pointwise policy-generation equality for the next action supplies the conditional action a.e. side condition required by the split-law builder. The reward-coordinate law remains a separate input.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-7451398b2bad","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3345,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3337"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_action_eq {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (policy : Nat -> Policy.MeasurablePolicy State Action) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairState : forall n : Nat, Measurable (pairState n)) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) (h_action_eq : (fun omega : Omega => action omega (i + 1)) = (fun omega : Omega => (policy i).action (pairState i (…","missing":[],"search":"action_condexpkernel_ae_eq_policy_historyfiltrationsucc_finitepairhistoryoftrace_of_action_eq banditrlproof.conditionalexpectationreward.action_condexpkernel_ae_eq_policy_historyfiltrationsucc_finitepairhistoryoftrace_of_action_eq generated-history specialization of the action side of the next-pair split. for `history.historyfiltrationsucc`, the finite action/reward pair prefix up to `i` is visible. therefore a pointwise policy-generation equality for the next action supplies the conditional action a.e. side condition required by the split-law builder. the reward-coordinate law remains a separate input. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc","label":"action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc","description":"Generated-trace source for the action side of the next-pair split. If the action trace is the shifted policy-generated trace from finite action/reward pair histories, the pointwise policy-generation equality required by `..._finitePairHistoryOfTrace_of_action_eq` follows from the definition of `Policy.generatedActionTraceSucc`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-bcc1f9a2c730","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3346,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3431"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (policy : Nat -> Policy.MeasurablePolicy State Action) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairState : forall n : Nat, Measurable (pairState n)) (defaultAction : Action) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) (h_action_generated : action = Policy.generatedActionTraceSucc policy (fun…","missing":[],"search":"action_condexpkernel_ae_eq_policy_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc banditrlproof.conditionalexpectationreward.action_condexpkernel_ae_eq_policy_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc generated-trace source for the action side of the next-pair split. if the action trace is the shifted policy-generated trace from finite action/reward pair histories, the pointwise policy-generation equality required by `..._finitepairhistoryoftrace_of_action_eq` follows from the definition of `policy.generatedactiontracesucc`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_policy_of_action_eq","label":"reward_condExpKernel_map_eq_selected_policy_of_action_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_policy_of_action_eq","description":"Reward-law rewrite for the next-pair split route. Some future trajectory/source theorem may naturally identify the conditional reward law using the actual next action at the conditioning point. If that actual action is trim-a.e. equal to the policy-selected action, this adapter rewrites the reward-coordinate map law into the policy-selected shape consumed by `actionRewardHistoryStepKernelFamily_pair_map_eq_of_action…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-2ff9dfe79ae9","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3347,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3507"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_policy_of_action_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (i : Nat) (pairHistory : Omega -> ((j : Finset.Iic i) -> Prod Action Rat)) (h_action_policy_eq : Filter.Eventually (fun omega : Omega => action omega (i + 1) = (policy i).action (pairState i (pairH…","missing":[],"search":"reward_condexpkernel_map_eq_selected_policy_of_action_eq banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_policy_of_action_eq reward-law rewrite for the next-pair split route. some future trajectory/source theorem may naturally identify the conditional reward law using the actual next action at the conditioning point. if that actual action is trim-a.e. equal to the policy-selected action, this adapter rewrites the reward-coordinate map law into the policy-selected shape consumed by `actionrewardhistorystepkernelfamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_pair_map_eq","label":"reward_condExpKernel_map_eq_selected_actual_action_of_pair_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_pair_map_eq","description":"Actual-action pair-law marginalization for the reward-coordinate route. Future trajectory-law work may identify, under the conditional kernel, the pair `(actual next action, next reward)` as a fixed-action product of the selected reward law. Mapping that pair law through `Prod.snd` gives the actual-action reward-coordinate map law consumed by the generated-action route below.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-9a6a635a9d5c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3348,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3564"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_pair_map_eq {Omega : Type u} {Context : Type v} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (i : Nat) (pairHistory : Omega -> ((j : Finset.Iic i) -> Prod Action Rat)) (h_reward_next : @Measurable Omega Rat mOmega inferInstance (fun omega : Omega => reward omega (i + 1))) (h_pair_map_eq_actual_action : Filter.Eventually (fun omega : Omega => @Measure.map Omega (Prod Action Rat) mOmega inferInstance (fun y : Omega => (action omega (i + 1), reward y (…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_pair_map_eq banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_pair_map_eq actual-action pair-law marginalization for the reward-coordinate route. future trajectory-law work may identify, under the conditional kernel, the pair `(actual next action, next reward)` as a fixed-action product of the selected reward law. mapping that pair law through `prod.snd` gives the actual-action reward-coordinate map law consumed by the generated-action route below. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.random_pair_condExpKernel_map_eq_actual_action_of_generatedActionTraceSucc_reward_map_eq_actual_action","label":"random_pair_condExpKernel_map_eq_actual_action_of_generatedActionTraceSucc_reward_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.random_pair_condExpKernel_map_eq_actual_action_of_generatedActionTraceSucc_reward_map_eq_actual_action","description":"Generated-action reward-map source for the random next-pair product law. The shifted generated-action trace supplies the conditional action a.e. equality, while the hypothesis supplies the reward-coordinate selected-measure law. The conclusion is the fully random next-pair product law consumed by the trajectory-facing random-pair route.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-037948f3a836","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3349,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3651"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem random_pair_condExpKernel_map_eq_actual_action_of_generatedActionTraceSucc_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairState : forall n : Nat, Measurable (pairState n)) (defaultAction : Action) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun om…","missing":[],"search":"random_pair_condexpkernel_map_eq_actual_action_of_generatedactiontracesucc_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.random_pair_condexpkernel_map_eq_actual_action_of_generatedactiontracesucc_reward_map_eq_actual_action generated-action reward-map source for the random next-pair product law. the shifted generated-action trace supplies the conditional action a.e. equality, while the hypothesis supplies the reward-coordinate selected-measure law. the conclusion is the fully random next-pair product law consumed by the trajectory-facing random-pair route. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.pair_condExpKernel_map_eq_frozen_actual_action_of_generatedActionTraceSucc_random_pair_map_eq","label":"pair_condExpKernel_map_eq_frozen_actual_action_of_generatedActionTraceSucc_random_pair_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.pair_condExpKernel_map_eq_frozen_actual_action_of_generatedActionTraceSucc_random_pair_map_eq","description":"Freeze the action coordinate in a random next-pair law. A trajectory source may identify the conditional law of the fully random pair `(action y (i+1), reward y (i+1))`, while the actual-action marginal route expects the action coordinate frozen at the conditioning point `omega`. Under the shifted generated-action trace, both action coordinates are conditionally equal to the same policy-selected action, so Mathlib's…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-33319d280ef2","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3350,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3791"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem pair_condExpKernel_map_eq_frozen_actual_action_of_generatedActionTraceSucc_random_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairState : forall n : Nat, Measurable (pairState n)) (defaultAction : Action) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Ome…","missing":[],"search":"pair_condexpkernel_map_eq_frozen_actual_action_of_generatedactiontracesucc_random_pair_map_eq banditrlproof.conditionalexpectationreward.pair_condexpkernel_map_eq_frozen_actual_action_of_generatedactiontracesucc_random_pair_map_eq freeze the action coordinate in a random next-pair law. a trajectory source may identify the conditional law of the fully random pair `(action y (i+1), reward y (i+1))`, while the actual-action marginal route expects the action coordinate frozen at the conditioning point `omega`. under the shifted generated-action trace, both action coordinates are conditionally equal to the same policy-selected action, so mathlib's `measure.map_congr` transfers the random-pair map law to the frozen-action map law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","description":"Random next-pair source law in canonical history-step-kernel form. This rewrites a generated-action random next-pair law whose right side is stated as `Measure.map (Prod.mk actualAction) selectedMeasure` into the standard `RewardKernel.actionRewardHistoryStepKernelFamily` form consumed by the pair-law route. It is a law-shape adapter; the random next-pair law itself remains an explicit hypothesis.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-b37adff837c3","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3351,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:3942"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairContext : forall n : Nat, Measurable (pairContext n)) (hpairState : forall n : Nat, Measurable (pairState n)) (defaultAction : Action) (action…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_random_pair_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_random_pair_map_eq_actual_action random next-pair source law in canonical history-step-kernel form. this rewrites a generated-action random next-pair law whose right side is stated as `measure.map (prod.mk actualaction) selectedmeasure` into the standard `rewardkernel.actionrewardhistorystepkernelfamily` form consumed by the pair-law route. it is a law-shape adapter; the random next-pair law itself remains an explicit hypothesis. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_of_measurable","label":"finiteRewardHistory_condExpKernel_frozen_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_of_measurable","description":"Finite reward-history frozen-past specialization. Once the finite reward history at time `i` is measurable with respect to filtration level `F i`, the conditional-expectation kernel at that level freezes the whole finite history trim-a.e. This supplies the direct `h_history_frozen` hypothesis needed by `centeredReward_succ_frozenPast_ae_of_history_frozen`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-0db2b9ad3760","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3352,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4059"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteRewardHistory_condExpKernel_frozen_of_measurable {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (i : Nat) (history : Omega -> ((j : Finset.Iic i) -> Rat)) (hhistory : @Measurable Omega ((j : Finset.Iic i) -> Rat) (F i) inferInstance history) : Filter.Eventually (fun omega : Omega => history =ᵐ[ ProbabilityTheory.condExpKernel (Ω := Omega) (mΩ := mOmega) mu (F i) omega] (fun _y : Omega => history omega)) (ae (mu.trim (F.le i)))","missing":[],"search":"finiterewardhistory_condexpkernel_frozen_of_measurable banditrlproof.conditionalexpectationreward.finiterewardhistory_condexpkernel_frozen_of_measurable finite reward-history frozen-past specialization. once the finite reward history at time `i` is measurable with respect to filtration level `f i`, the conditional-expectation kernel at that level freezes the whole finite history trim-a.e. this supplies the direct `h_history_frozen` hypothesis needed by `centeredreward_succ_frozenpast_ae_of_history_frozen`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_of_coordinate_measurable","label":"finiteRewardHistory_condExpKernel_frozen_of_coordinate_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_of_coordinate_measurable","description":"Coordinate-measurability hookup for finite reward histories. If every coordinate in the finite reward prefix is measurable at filtration level `F i`, the whole `finiteRewardHistoryOfTrace` object is measurable and therefore frozen by the conditional-expectation kernel. This is still only a measurability bridge; it does not identify the conditional kernel with a trajectory/history-step reward law.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-04ebee20fee0","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3353,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4093"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteRewardHistory_condExpKernel_frozen_of_coordinate_measurable {Omega : Type u} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (reward : Omega -> RewardTrace Rat) (i : Nat) (hreward : forall j : Finset.Iic i, @Measurable Omega Rat (F i) inferInstance (fun omega : Omega => reward omega j.1)) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (F i) omega)) (fun y : Omega => History.finiteRewardHistoryOfTrace (reward y) i) (fun _y : Omega => History.finiteRewardHistoryOfTrace (reward omega) i)) (ae (mu.trim (F.le i)))","missing":[],"search":"finiterewardhistory_condexpkernel_frozen_of_coordinate_measurable banditrlproof.conditionalexpectationreward.finiterewardhistory_condexpkernel_frozen_of_coordinate_measurable coordinate-measurability hookup for finite reward histories. if every coordinate in the finite reward prefix is measurable at filtration level `f i`, the whole `finiterewardhistoryoftrace` object is measurable and therefore frozen by the conditional-expectation kernel. this is still only a measurability bridge; it does not identify the conditional kernel with a trajectory/history-step reward law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_historyFiltrationSucc","label":"finiteRewardHistory_condExpKernel_frozen_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_historyFiltrationSucc","description":"Generated-history-filtration specialization of finite reward-history freezing. For the shifted history filtration `historyFiltrationSucc`, all reward coordinates up to `i` are visible at level `i`, so the finite reward history prefix is frozen trim-a.e. under its conditional-expectation kernel. This closes the concrete finite-history measurability side of the frozen-past route; the remaining missing leaf is the rewa…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-b2452746c81d","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3354,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4134"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteRewardHistory_condExpKernel_frozen_historyFiltrationSucc {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ ((History.historyFiltrationSucc action reward haction hreward) i) omega)) (fun y : Omega => History.finiteRewardHistoryOfTrace (reward y) i) (fun _y : Omega => History.finiteRewardHistoryOfTrace (reward omega) i)) (ae (mu.trim ((History.historyFiltrationSucc action reward…","missing":[],"search":"finiterewardhistory_condexpkernel_frozen_historyfiltrationsucc banditrlproof.conditionalexpectationreward.finiterewardhistory_condexpkernel_frozen_historyfiltrationsucc generated-history-filtration specialization of finite reward-history freezing. for the shifted history filtration `historyfiltrationsucc`, all reward coordinates up to `i` are visible at level `i`, so the finite reward history prefix is frozen trim-a.e. under its conditional-expectation kernel. this closes the concrete finite-history measurability side of the frozen-past route; the remaining missing leaf is the reward-law/trajectory-law identification. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_of_measurable","label":"finitePairHistory_condExpKernel_frozen_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_of_measurable","description":"Finite action/reward pair-history frozen-past specialization. This is the pair-coordinate companion to `finiteRewardHistory_condExpKernel_frozen_of_measurable`. It needs a countable action space because the frozen object contains action coordinates.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8e7f38bfd76f","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3355,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4181"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_condExpKernel_frozen_of_measurable {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (i : Nat) (history : Omega -> ((j : Finset.Iic i) -> Prod Action Rat)) (hhistory : @Measurable Omega ((j : Finset.Iic i) -> Prod Action Rat) (F i) inferInstance history) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (F i) omega)) history (fun _y : Omega => history omega)) (ae (mu.trim (F.le i)))","missing":[],"search":"finitepairhistory_condexpkernel_frozen_of_measurable banditrlproof.conditionalexpectationreward.finitepairhistory_condexpkernel_frozen_of_measurable finite action/reward pair-history frozen-past specialization. this is the pair-coordinate companion to `finiterewardhistory_condexpkernel_frozen_of_measurable`. it needs a countable action space because the frozen object contains action coordinates. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_of_coordinate_measurable","label":"finitePairHistory_condExpKernel_frozen_of_coordinate_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_of_coordinate_measurable","description":"Coordinate-measurability hookup for finite action/reward pair histories. If every action and reward coordinate in the finite pair prefix is measurable at filtration level `F i`, the whole `finitePairHistoryOfTrace` object is measurable and therefore frozen under the conditional-expectation kernel.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-35977f682cdb","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3356,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4216"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_condExpKernel_frozen_of_coordinate_measurable {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (i : Nat) (haction : forall j : Finset.Iic i, @Measurable Omega Action (F i) inferInstance (fun omega : Omega => action omega j.1)) (hreward : forall j : Finset.Iic i, @Measurable Omega Rat (F i) inferInstance (fun omega : Omega => reward omega j.1)) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ (F i) omega)) (fun y : Omega => History.finitePairHistoryOfTrace (action y) (reward y) i) (fun _y : Omega => History.finitePairHistoryOfTr…","missing":[],"search":"finitepairhistory_condexpkernel_frozen_of_coordinate_measurable banditrlproof.conditionalexpectationreward.finitepairhistory_condexpkernel_frozen_of_coordinate_measurable coordinate-measurability hookup for finite action/reward pair histories. if every action and reward coordinate in the finite pair prefix is measurable at filtration level `f i`, the whole `finitepairhistoryoftrace` object is measurable and therefore frozen under the conditional-expectation kernel. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_historyFiltrationSucc","label":"finitePairHistory_condExpKernel_frozen_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_historyFiltrationSucc","description":"Generated-history-filtration specialization of finite pair-history freezing. For `History.historyFiltrationSucc`, all action and reward coordinates up to `i` are visible at level `i`, so the finite `(Action, Reward)` pair prefix is frozen trim-a.e. under its conditional-expectation kernel. This is a frozen past hook for future `partialTraj`/`condExpKernel` pair-law identification; it does not prove that law.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a0422b82fa3b","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3357,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4265"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_condExpKernel_frozen_historyFiltrationSucc {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ ((History.historyFiltrationSucc action reward haction hreward) i) omega)) (fun y : Omega => History.finitePairHistoryOfTrace (action y) (reward y) i) (fun _y : Omega => History.finitePairHistoryOfTrace (action omega) (reward omega) i)) (ae (mu.trim ((Histo…","missing":[],"search":"finitepairhistory_condexpkernel_frozen_historyfiltrationsucc banditrlproof.conditionalexpectationreward.finitepairhistory_condexpkernel_frozen_historyfiltrationsucc generated-history-filtration specialization of finite pair-history freezing. for `history.historyfiltrationsucc`, all action and reward coordinates up to `i` are visible at level `i`, so the finite `(action, reward)` pair prefix is frozen trim-a.e. under its conditional-expectation kernel. this is a frozen past hook for future `partialtraj`/`condexpkernel` pair-law identification; it does not prove that law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_ae_eq_extend_of_pairHistory_frozen","label":"finitePairHistory_succ_ae_eq_extend_of_pairHistory_frozen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_ae_eq_extend_of_pairHistory_frozen","description":"Successor finite-pair trace decomposition under any frozen-prefix hypothesis. If the old finite pair prefix is a.e. frozen at `omega`, then the full `i + 1` prefix is a.e. the deterministic successor extension of that frozen prefix by the random next `(Action, Reward)` pair.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a366b8bbd409","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3358,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4320"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_succ_ae_eq_extend_of_pairHistory_frozen {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] (nu : Measure Omega) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (i : Nat) (omega : Omega) (h_pair_history_frozen : Filter.EventuallyEq (ae nu) (fun y : Omega => History.finitePairHistoryOfTrace (action y) (reward y) i) (fun _y : Omega => History.finitePairHistoryOfTrace (action omega) (reward omega) i)) : Filter.EventuallyEq (ae nu) (fun y : Omega => History.finitePairHistoryOfTrace (action y) (reward y) (i + 1)) (fun y : Omega => History.extendPairHistorySucc (History.finitePairHistoryOfTrace (action omega) (reward omega) i) (action y (i + 1), reward y (i + 1)))","missing":[],"search":"finitepairhistory_succ_ae_eq_extend_of_pairhistory_frozen banditrlproof.conditionalexpectationreward.finitepairhistory_succ_ae_eq_extend_of_pairhistory_frozen successor finite-pair trace decomposition under any frozen-prefix hypothesis. if the old finite pair prefix is a.e. frozen at `omega`, then the full `i + 1` prefix is a.e. the deterministic successor extension of that frozen prefix by the random next `(action, reward)` pair. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_condExpKernel_ae_eq_extend_historyFiltrationSucc","label":"finitePairHistory_succ_condExpKernel_ae_eq_extend_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_condExpKernel_ae_eq_extend_historyFiltrationSucc","description":"Generated-history conditional-kernel successor decomposition for pair traces. This packages the previous frozen pair-prefix theorem into the concrete `History.historyFiltrationSucc` conditional kernel. It is not a joint law identification; it only rewrites the random `i + 1` pair trace into a frozen prefix plus random next pair under the conditional kernel.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-398847057fa4","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3359,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4361"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_succ_condExpKernel_ae_eq_extend_historyFiltrationSucc {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) : Filter.Eventually (fun omega : Omega => Filter.EventuallyEq (ae (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ ((History.historyFiltrationSucc action reward haction hreward) i) omega)) (fun y : Omega => History.finitePairHistoryOfTrace (action y) (reward y) (i + 1)) (fun y : Omega => History.extendPairHistorySucc (History.finitePairHistoryOfTrace (action…","missing":[],"search":"finitepairhistory_succ_condexpkernel_ae_eq_extend_historyfiltrationsucc banditrlproof.conditionalexpectationreward.finitepairhistory_succ_condexpkernel_ae_eq_extend_historyfiltrationsucc generated-history conditional-kernel successor decomposition for pair traces. this packages the previous frozen pair-prefix theorem into the concrete `history.historyfiltrationsucc` conditional kernel. it is not a joint law identification; it only rewrites the random `i + 1` pair trace into a frozen prefix plus random next pair under the conditional kernel. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_condExpKernel_map_eq_extend_historyFiltrationSucc","label":"finitePairHistory_succ_condExpKernel_map_eq_extend_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_condExpKernel_map_eq_extend_historyFiltrationSucc","description":"Pushforward form of the generated-history successor decomposition. The previous theorem is an a.e. equality under the conditional kernel. This lemma upgrades it via Mathlib's `Measure.map_congr`, so later route cards can state the remaining trajectory-law input against the deterministic `extendPairHistorySucc` map instead of the full `i + 1` trace restriction.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a08f3b59f096","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3360,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4420"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistory_succ_condExpKernel_map_eq_extend_historyFiltrationSucc {Omega : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) : Filter.Eventually (fun omega : Omega => @Measure.map Omega ((j : Finset.Iic (i + 1)) -> Prod Action Rat) mOmega inferInstance (fun y : Omega => History.finitePairHistoryOfTrace (action y) (reward y) (i + 1)) (@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ ((History.historyFiltrationSucc action reward haction hreward) i) omega) = @Measure.map Omega ((j :…","missing":[],"search":"finitepairhistory_succ_condexpkernel_map_eq_extend_historyfiltrationsucc banditrlproof.conditionalexpectationreward.finitepairhistory_succ_condexpkernel_map_eq_extend_historyfiltrationsucc pushforward form of the generated-history successor decomposition. the previous theorem is an a.e. equality under the conditional kernel. this lemma upgrades it via mathlib's `measure.map_congr`, so later route cards can state the remaining trajectory-law input against the deterministic `extendpairhistorysucc` map instead of the full `i + 1` trace restriction. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_frozenPast_ae_of_history_frozen","label":"centeredReward_succ_frozenPast_ae_of_history_frozen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_frozenPast_ae_of_history_frozen","description":"Frozen-history bridge for the succ-indexed map-law route. If the finite reward history is already frozen under the conditional kernel, then the history-selected context/action mean in the centered next-reward variable is frozen as well. This is the deterministic part of the `h_kernel_X_eq` side condition consumed by the map-law conditional-expectation bridge; it deliberately does not prove that the history itself is…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-0cd4d9e2df1f","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3361,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4476"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_frozenPast_ae_of_history_frozen {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (reward : Omega -> RewardTrace Rat) (i : Nat) (history : Omega -> ((j : Finset.Iic i) -> Rat)) (h_history_frozen : Filter.Eventually (fun omega : Omega => history =ᵐ[ ProbabilityTheory.condExpKernel (Ω := Omega) (mΩ := mOmega) mu (F i) omega] (fun _y : Omega => history omega)) (ae (mu.trim (F.le i)))) : Filter.Eventually (fun omega : Omega => (fun y : Omega => (((reward y (…","missing":[],"search":"centeredreward_succ_frozenpast_ae_of_history_frozen banditrlproof.conditionalexpectationreward.centeredreward_succ_frozenpast_ae_of_history_frozen frozen-history bridge for the succ-indexed map-law route. if the finite reward history is already frozen under the conditional kernel, then the history-selected context/action mean in the centered next-reward variable is frozen as well. this is the deterministic part of the `h_kernel_x_eq` side condition consumed by the map-law conditional-expectation bridge; it deliberately does not prove that the history itself is frozen. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","label":"centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","description":"Conditional sub-Gaussian map-law consumer with the frozen-past side condition discharged from prefix coordinate measurability. The remaining structural input is the reward-coordinate pushforward identity from `condExpKernel` to the history-step reward kernel. Ambient measurability of the centered reward and the deterministic variance-proxy upper bound remain explicit regularity contracts; exponential integrability i…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d2d66f0c03e0","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3362,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4525"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean vari…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_historystepkernelfamily_condexpkernel_map_eq_of_coordinate_measurable banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_historystepkernelfamily_condexpkernel_map_eq_of_coordinate_measurable conditional sub-gaussian map-law consumer with the frozen-past side condition discharged from prefix coordinate measurability. the remaining structural input is the reward-coordinate pushforward identity from `condexpkernel` to the history-step reward kernel. ambient measurability of the centered reward and the deterministic variance-proxy upper bound remain explicit regularity contracts; exponential integrability is derived. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","label":"centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","description":"Generated-history-filtration specialization of the conditional sub-Gaussian map-law consumer. At filtration level `History.historyFiltrationSucc ... i`, the reward prefix up to `i` is visible by construction, so the coordinate-measurable consumer applies directly. The reward-coordinate pushforward identity and analytic measurability/variance contracts remain explicit, while exponential integrability is derived from…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8c297c5d24a3","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3363,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4678"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.Cent…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_historystepkernelfamily_condexpkernel_map_eq_historyfiltrationsucc banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_historystepkernelfamily_condexpkernel_map_eq_historyfiltrationsucc generated-history-filtration specialization of the conditional sub-gaussian map-law consumer. at filtration level `history.historyfiltrationsucc ... i`, the reward prefix up to `i` is visible by construction, so the coordinate-measurable consumer applies directly. the reward-coordinate pushforward identity and analytic measurability/variance contracts remain explicit, while exponential integrability is derived from the selected target laws. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq","label":"centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq","description":"Succ-indexed selected-reward specialization of the map-law consumer. The `h_kernel_map_eq` hypothesis is the reward-coordinate pushforward form of the future trajectory-law identification. The `h_kernel_X_eq` hypothesis is the matching frozen-past condition for the centered variable.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-7729610af232","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3364,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4789"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (reward : Omega -> Rewa…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq succ-indexed selected-reward specialization of the map-law consumer. the `h_kernel_map_eq` hypothesis is the reward-coordinate pushforward form of the future trajectory-law identification. the `h_kernel_x_eq` hypothesis is the matching frozen-past condition for the centered variable. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","label":"centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","description":"Map-law consumer with the frozen-past side condition discharged from prefix coordinate measurability. The remaining structural input is still `h_kernel_map_eq`: the reward-coordinate pushforward identity from `condExpKernel` to the history-step reward kernel. This theorem only removes the separate `h_kernel_X_eq` obligation by proving finite-history frozen-past from coordinate measurability at `F i`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-cb16813b0ef0","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3365,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:4890"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean variancePr…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_of_coordinate_measurable banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_of_coordinate_measurable map-law consumer with the frozen-past side condition discharged from prefix coordinate measurability. the remaining structural input is still `h_kernel_map_eq`: the reward-coordinate pushforward identity from `condexpkernel` to the history-step reward kernel. this theorem only removes the separate `h_kernel_x_eq` obligation by proving finite-history frozen-past from coordinate measurability at `f i`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","label":"centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","description":"Generated-history-filtration specialization of the map-law consumer. At filtration level `History.historyFiltrationSucc ... i`, the reward prefix up to `i` is visible by construction, so the preceding coordinate-measurable consumer applies directly. The reward-coordinate pushforward identity remains an explicit assumption.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-39d7c9f40294","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3366,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5028"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRe…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_historyfiltrationsucc banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_historyfiltrationsucc generated-history-filtration specialization of the map-law consumer. at filtration level `history.historyfiltrationsucc ... i`, the reward prefix up to `i` is visible by construction, so the preceding coordinate-measurable consumer applies directly. the reward-coordinate pushforward identity remains an explicit assumption. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable","description":"Pair-law route into the map-law conditional mean-zero consumer. If the conditional kernel has the correct next `(Action × Reward)` law, then mapping both sides through `Prod.snd` gives the reward-coordinate law consumed by the map-law conditional-expectation bridge. This theorem still assumes the pair-law identity; it only packages the marginalization step through `RewardKernel.actionRewardHistoryStepKernelFamily_re…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-933308b8656b","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3367,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5125"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) ->…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_of_coordinate_measurable banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_of_coordinate_measurable pair-law route into the map-law conditional mean-zero consumer. if the conditional kernel has the correct next `(action × reward)` law, then mapping both sides through `prod.snd` gives the reward-coordinate law consumed by the map-law conditional-expectation bridge. this theorem still assumes the pair-law identity; it only packages the marginalization step through `rewardkernel.actionrewardhistorystepkernelfamily_reward_map`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc","description":"Generated-history-filtration specialization of the action/reward pair-law route. At `History.historyFiltrationSucc ... i`, the next action/reward coordinates are measurable in the ambient space and the reward prefix up to `i` is visible at filtration level `i`. Thus the coordinate-measurable pair-map consumer applies directly. The actual action/reward pair-law pushforward identity remains an explicit assumption.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-7afd441b4589","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3368,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5289"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> (…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc generated-history-filtration specialization of the action/reward pair-law route. at `history.historyfiltrationsucc ... i`, the next action/reward coordinates are measurable in the ambient space and the reward prefix up to `i` is visible at filtration level `i`. thus the coordinate-measurable pair-map consumer applies directly. the actual action/reward pair-law pushforward identity remains an explicit assumption. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected","description":"Generated-history pair-law route with the concrete trace pair-history. This specializes the previous generated-history consumer to the finite action/reward pair history obtained from the actual traces, and to context/state extractors that read only the reward projection of that pair history. It removes the separate pair-history compatibility hypotheses; the remaining structural assumption is exactly the generated-hi…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-060648932148","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3369,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5411"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (hpairContext : forall n : Nat, Measurable (fun history : (j : Finset.Iic n) -> Prod Action Rat =…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected generated-history pair-law route with the concrete trace pair-history. this specializes the previous generated-history consumer to the finite action/reward pair history obtained from the actual traces, and to context/state extractors that read only the reward projection of that pair history. it removes the separate pair-history compatibility hypotheses; the remaining structural assumption is exactly the generated-history `condexpkernel` pushforward equality for the next `(action, reward)` pair. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable","description":"Projected trace-pair route with projection measurability supplied locally. The remaining hypothesis is the concrete generated-history pair-law equality. The reward-projection context/state measurability proofs are derived from the original reward-history `context`/`state` measurability and `History.measurable_pairHistoryRewardProjection`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-c052d4be8246","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3370,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5531"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected_of_context_state_measurable banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected_of_context_state_measurable projected trace-pair route with projection measurability supplied locally. the remaining hypothesis is the concrete generated-history pair-law equality. the reward-projection context/state measurability proofs are derived from the original reward-history `context`/`state` measurability and `history.measurable_pairhistoryrewardprojection`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Projected trace-pair route using the named finite pair-history prefix. This is the same route as `..._projected_of_context_state_measurable`, but the remaining pair-law hypothesis is stated with `History.finitePairHistoryOfTrace`. That is the finite-prefix object aligned with `RewardKernel.actionRewardPartialTrajectoryKernel`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-f82689dbbc63","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3371,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5643"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (l…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace projected trace-pair route using the named finite pair-history prefix. this is the same route as `..._projected_of_context_state_measurable`, but the remaining pair-law hypothesis is stated with `history.finitepairhistoryoftrace`. that is the finite-prefix object aligned with `rewardkernel.actionrewardpartialtrajectorykernel`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Project a generated-history `partialTraj` trace law to the next-pair law. This names the reusable law-identification step that was previously only embedded inside the centered-reward consumer below. If the conditional kernel of the full finite pair trace at `i + 1` agrees with the one-step action/reward `partialTraj` kernel, then mapping both sides to the successor coordinate gives the concrete next `(Action, Reward…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-f0848c326e46","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3372,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5749"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace project a generated-history `partialtraj` trace law to the next-pair law. this names the reusable law-identification step that was previously only embedded inside the centered-reward consumer below. if the conditional kernel of the full finite pair trace at `i + 1` agrees with the one-step action/reward `partialtraj` kernel, then mapping both sides to the successor coordinate gives the concrete next `(action, reward)` pushforward law consumed by the pair-map route. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Project a history-step next-pair law to the actual-action reward law. This is the direct reward-coordinate adapter for a conditional kernel law already stated at the next `(Action, Reward)` level. Mapping both sides through `Prod.snd` reduces the target to `RewardKernel.actionRewardHistoryStepKernelFamily_reward_map`; the only remaining rewrite is the trim-a.e. successor action equality.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-1cca915fe07e","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3373,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:5958"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairContext : forall n : Nat, Measurable (pairContext n)) (hpairState : forall n : Nat, Measurable (pairState n)) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurabl…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardhistorystepkernelfamily_pair_map_eq project a history-step next-pair law to the actual-action reward law. this is the direct reward-coordinate adapter for a conditional kernel law already stated at the next `(action, reward)` level. mapping both sides through `prod.snd` reduces the target to `rewardkernel.actionrewardhistorystepkernelfamily_reward_map`; the only remaining rewrite is the trim-a.e. successor action equality. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Specialize the next-pair reward adapter to generated pair histories. The hypothesis is already the history-step next-pair `condExpKernel` law for `History.finitePairHistoryOfTrace`. The theorem only projects it through the reward coordinate and rewrites the selected policy action to the actual next action by the supplied successor equality.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-3deaaa8fc9af","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3374,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6066"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Meas…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace specialize the next-pair reward adapter to generated pair histories. the hypothesis is already the history-step next-pair `condexpkernel` law for `history.finitepairhistoryoftrace`. the theorem only projects it through the reward coordinate and rewrites the selected policy action to the actual next action by the supplied successor equality. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Generated-action wrapper for the finite-pair-history next-pair adapter. This removes the explicit successor action equality when the action trace is definitionally generated from the reward-history state.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a36fc6407b52","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3375,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6193"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_generatedactiontracesucc_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_generatedactiontracesucc_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace generated-action wrapper for the finite-pair-history next-pair adapter. this removes the explicit successor action equality when the action trace is definitionally generated from the reward-history state. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Project a full finite-pair `partialTraj` law to the actual-action reward law. The full trace law first gives the next `(Action, Reward)` marginal by the compiled `partialTraj` next-coordinate adapter. Mapping that marginal through `Prod.snd` and using `RewardKernel.actionRewardHistoryStepKernelFamily_reward_map` then gives the reward-coordinate law. The final rewrite only needs an `i + 1` action equality against the…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-91cb8eef2150","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3376,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6301"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurabl…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace project a full finite-pair `partialtraj` law to the actual-action reward law. the full trace law first gives the next `(action, reward)` marginal by the compiled `partialtraj` next-coordinate adapter. mapping that marginal through `prod.snd` and using `rewardkernel.actionrewardhistorystepkernelfamily_reward_map` then gives the reward-coordinate law. the final rewrite only needs an `i + 1` action equality against the policy-selected action. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Generated-action specialization of the full finite-pair `partialTraj` to actual-action reward-map adapter. `Policy.generatedActionTraceSucc` supplies the pointwise successor action equality required by `reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-c2aea271bf7e","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3377,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6489"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> Rew…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace generated-action specialization of the full finite-pair `partialtraj` to actual-action reward-map adapter. `policy.generatedactiontracesucc` supplies the pointwise successor action equality required by `reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Generated-history route from a one-step `partialTraj` law assumption. If the conditional kernel pushed forward to the extended finite pair trace agrees with the local action/reward `partialTraj` kernel from `i` to `i + 1`, then Mathlib's `partialTraj` next-coordinate marginal wrapper turns that into the concrete next `(Action, Reward)` pair-law consumed by `..._finitePairHistoryOfTrace`.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-c145d50fe469","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3378,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6599"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law :…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace generated-history route from a one-step `partialtraj` law assumption. if the conditional kernel pushed forward to the extended finite pair trace agrees with the local action/reward `partialtraj` kernel from `i` to `i + 1`, then mathlib's `partialtraj` next-coordinate marginal wrapper turns that into the concrete next `(action, reward)` pair-law consumed by `..._finitepairhistoryoftrace`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq","description":"Build the full next `(Action, Reward)` conditional-kernel law from split action and reward laws. The action side is deterministic/predictable under the conditional kernel, and the reward side supplies the selected reward law. Together they identify the next-pair pushforward with the local history-step action/reward kernel.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-1a0ed23191e4","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3379,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6856"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairContext : forall n : Nat, Measurable (pairContext n)) (hpairState : forall n : Nat, Measurable (pairState n)) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (i : Nat) (pairHistory : Omega -> ((j : Finset.Iic i)…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq build the full next `(action, reward)` conditional-kernel law from split action and reward laws. the action side is deterministic/predictable under the conditional kernel, and the reward side supplies the selected reward law. together they identify the next-pair pushforward with the local history-step action/reward kernel. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","description":"Generated-action source plus an actual-action reward law gives the full next-pair law. This is the generated-history specialization of the split-law builder. The action side is supplied by `Policy.generatedActionTraceSucc`; the remaining reward assumption may be stated with the actual next action at the conditioning point, and is rewritten to the policy-selected action before invoking the generic split-law theorem.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-de67f43eb372","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3380,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:6981"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairContext : forall n : Nat, Measurable (pairContext n)) (hpairState : forall n : Nat, Measurable (pairState n)) (defaultAction : Action) (action : Om…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_reward_map_eq_actual_action generated-action source plus an actual-action reward law gives the full next-pair law. this is the generated-history specialization of the split-law builder. the action side is supplied by `policy.generatedactiontracesucc`; the remaining reward assumption may be stated with the actual next action at the conditioning point, and is rewritten to the policy-selected action before invoking the generic split-law theorem. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Turn a next-pair `condExpKernel` law into the extension-map `partialTraj` law. This is the law-shaped bridge between the pair-map consumer and the extension-map `partialTraj` route: once the conditional kernel of the next `(Action, Reward)` pair is identified with the configured history-step kernel, pushing both sides through `History.extendPairHistorySucc` identifies the one-step finite-prefix trajectory kernel.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-e2a7f9c2bd26","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3381,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7134"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : Measure Omega) [IsFiniteMeasure mu] (F : Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> Context) (pairState : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> State) (hpairContext : forall n : Nat, Measurable (pairContext n)) (hpairState : forall n : Nat, Measurable (pairState n)) (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (i : Nat) (pairHistory : Omega -> ((j…","missing":[],"search":"actionrewardpartialtrajectorykernel_extend_map_eq_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_extend_map_eq_of_actionrewardhistorystepkernelfamily_pair_map_eq turn a next-pair `condexpkernel` law into the extension-map `partialtraj` law. this is the law-shaped bridge between the pair-map consumer and the extension-map `partialtraj` route: once the conditional kernel of the next `(action, reward)` pair is identified with the configured history-step kernel, pushing both sides through `history.extendpairhistorysucc` identifies the one-step finite-prefix trajectory kernel. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Generated-history finite-pair-trace specialization of the extension-map law builder. The remaining assumption is the concrete conditional next-pair law against `RewardKernel.actionRewardHistoryStepKernelFamily`; this theorem packages the pushforward through `History.extendPairHistorySucc` and the local `partialTraj` wrapper.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d3c0631ac6b6","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3382,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7256"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measu…","missing":[],"search":"actionrewardpartialtrajectorykernel_extend_map_eq_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_extend_map_eq_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace generated-history finite-pair-trace specialization of the extension-map law builder. the remaining assumption is the concrete conditional next-pair law against `rewardkernel.actionrewardhistorystepkernelfamily`; this theorem packages the pushforward through `history.extendpairhistorysucc` and the local `partialtraj` wrapper. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","label":"actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","description":"Generated-action and actual-action reward law source for the extension-map `partialTraj` route. This composes the generated-action next-pair split-law hookup with the extension-map `partialTraj` builder. The remaining law input is only the reward-coordinate conditional map law selected by the actual next action; the wrapper rewrites that action to the policy-selected action and pushes the resulting next-pair law thr…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8950fd3e86d9","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3383,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7372"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega ->…","missing":[],"search":"actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_reward_map_eq_actual_action generated-action and actual-action reward law source for the extension-map `partialtraj` route. this composes the generated-action next-pair split-law hookup with the extension-map `partialtraj` builder. the remaining law input is only the reward-coordinate conditional map law selected by the actual next action; the wrapper rewrites that action to the policy-selected action and pushes the resulting next-pair law through `history.extendpairhistorysucc`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_of_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"actionRewardPartialTrajectoryKernel_map_eq_of_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_of_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Turn an extension-map `partialTraj` law into the full finite-pair-trace law. The generated history filtration already freezes the old pair prefix under the conditional kernel, so the full `i + 1` trace pushforward agrees a.e. with the deterministic extension of the frozen `i` prefix by the random next pair. This adapter packages that successor decomposition separately from the centered reward consumer.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-db63209d8bf3","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3384,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7538"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_of_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Ome…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_of_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_of_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace turn an extension-map `partialtraj` law into the full finite-pair-trace law. the generated history filtration already freezes the old pair prefix under the conditional kernel, so the full `i + 1` trace pushforward agrees a.e. with the deterministic extension of the frozen `i` prefix by the random next pair. this adapter packages that successor decomposition separately from the centered reward consumer. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Project an extension-map `partialTraj` law to the actual-action reward law. This is the reward-coordinate counterpart of `actionRewardPartialTrajectoryKernel_map_eq_of_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace`: first lift the deterministic frozen-prefix extension law back to the full finite-pair trace law, then reuse the finite-pair trace reward-map adapter.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-dd6f8e989338","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3385,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7663"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> RewardTrace Rat) (haction :…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace project an extension-map `partialtraj` law to the actual-action reward law. this is the reward-coordinate counterpart of `actionrewardpartialtrajectorykernel_map_eq_of_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace`: first lift the deterministic frozen-prefix extension law back to the full finite-pair trace law, then reuse the finite-pair trace reward-map adapter. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Generated-action specialization of the extension-map reward-coordinate adapter. The generated shifted policy trace supplies the successor action equality; the remaining hypothesis is only the frozen-prefix extension-map `partialTraj` law.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8773b06fa94e","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3386,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7809"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Acti…","missing":[],"search":"reward_condexpkernel_map_eq_selected_actual_action_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_actual_action_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace generated-action specialization of the extension-map reward-coordinate adapter. the generated shifted policy trace supplies the successor action equality; the remaining hypothesis is only the frozen-prefix extension-map `partialtraj` law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","description":"Generated-action and actual-action reward law source for the full finite pair-trace `partialTraj` law. This exposes the trajectory-law part of the generated-history route without also consuming the centered-reward kernel law. The remaining external input is the actual-action reward-coordinate conditional map law.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-794bdba200ab","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3387,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:7922"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardT…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_reward_map_eq_actual_action generated-action and actual-action reward law source for the full finite pair-trace `partialtraj` law. this exposes the trajectory-law part of the generated-history route without also consuming the centered-reward kernel law. the remaining external input is the actual-action reward-coordinate conditional map law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","description":"Generated-history route from an extension-map `partialTraj` law assumption. This is the next narrowing of the `partialTraj`/`condExpKernel` gap. Instead of requiring a law identity for the full `i + 1` trace restriction, it only requires the conditional kernel pushed through the deterministic extension of the frozen old pair prefix by the random next `(Action, Reward)` pair to agree with the one-step action/reward `…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-a22ccc9ac4c9","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3388,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:8070"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context ->…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace generated-history route from an extension-map `partialtraj` law assumption. this is the next narrowing of the `partialtraj`/`condexpkernel` gap. instead of requiring a law identity for the full `i + 1` trace restriction, it only requires the conditional kernel pushed through the deterministic extension of the frozen old pair prefix by the random next `(action, reward)` pair to agree with the one-step action/reward `partialtraj` kernel. the compiled successor-decomposition bridge supplies the conversion back to the existing full-trace consumer. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action","description":"Generated-action and actual-action reward law source for succ-indexed conditional mean-zero. This is the current narrowest compiled consumer on the adaptive pair-law route: it combines the shifted generated-action trace, an actual-action reward-coordinate conditional map law, the extension-map `partialTraj` bridge, and the centered-reward law transfer to produce ordinary conditional mean-zero. It still assumes the a…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-d8ca01acaf26","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3389,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:8228"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.Cente…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_actual_action generated-action and actual-action reward law source for succ-indexed conditional mean-zero. this is the current narrowest compiled consumer on the adaptive pair-law route: it combines the shifted generated-action trace, an actual-action reward-coordinate conditional map law, the extension-map `partialtraj` bridge, and the centered-reward law transfer to produce ordinary conditional mean-zero. it still assumes the actual-action reward-coordinate law rather than constructing it from an ambient trajectory measure. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_pair_map_eq_actual_action","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_pair_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_pair_map_eq_actual_action","description":"Actual-action pair-product law source for the full finite-pair-trace `partialTraj` law. This exposes the trajectory-law layer that was previously only consumed inside the centered-reward theorem: an actual-action pair-product conditional law is first marginalized to the actual-action reward-coordinate law, then the generated-action route turns it into the full `i + 1` finite pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-2416ba2907ed","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3390,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:8376"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_pair_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTra…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_pair_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_pair_map_eq_actual_action actual-action pair-product law source for the full finite-pair-trace `partialtraj` law. this exposes the trajectory-law layer that was previously only consumed inside the centered-reward theorem: an actual-action pair-product conditional law is first marginalized to the actual-action reward-coordinate law, then the generated-action route turns it into the full `i + 1` finite pair-trace `partialtraj` law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action","description":"Actual-action pair-product law source for generated-action conditional mean-zero. This packages one more upstream law shape: if the conditional kernel identifies the pair `(actual next action, next reward)` with the selected reward law pushed through `Prod.mk` at the actual next action, then `Prod.snd` marginalization provides the actual-action reward-coordinate law required by `centeredReward_succ_condExp_eq_zero_o…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-8b0d930c29cd","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3391,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:8510"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.Centere…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_pair_map_eq_actual_action banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_pair_map_eq_actual_action actual-action pair-product law source for generated-action conditional mean-zero. this packages one more upstream law shape: if the conditional kernel identifies the pair `(actual next action, next reward)` with the selected reward law pushed through `prod.mk` at the actual next action, then `prod.snd` marginalization provides the actual-action reward-coordinate law required by `centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_actual_action`. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","description":"Fully-random next-pair product law source for the full finite-pair-trace `partialTraj` law. The hypothesis may state the conditional law of the sampled pair `(action y (i+1), reward y (i+1))`. The shifted generated-action trace freezes that action coordinate to the actual/policy-selected action under the conditional kernel, after which the actual-action pair-product adapter produces the full trace law.","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-44a745b68f1c","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3392,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:8660"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> Re…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_random_pair_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiontracesucc_random_pair_map_eq_actual_action fully-random next-pair product law source for the full finite-pair-trace `partialtraj` law. the hypothesis may state the conditional law of the sampled pair `(action y (i+1), reward y (i+1))`. the shifted generated-action trace freezes that action coordinate to the actual/policy-selected action under the conditional kernel, after which the actual-action pair-product adapter produces the full trace law. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","description":"Random next-pair product law source for generated-action conditional mean-zero. This is a slightly more trajectory-facing consumer than `..._pair_map_eq_actual_action`: the law hypothesis may state the conditional pushforward of the fully random next pair `(action y (i+1), reward y (i+1))`. The generated-action trace freezes the action coordinate, after which the existing actual-action pair-product consumer handles…","url":"../modules/banditrlproof-conditionalexpectationreward/index.html#decl-223f3d2f6cca","parent":"module:BanditRLProof.ConditionalExpectationReward","order":3393,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalExpectationReward"],["Source","BanditRLProof/ConditionalExpectationReward.lean:8837"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_random_pair_map_eq_actual_action banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_random_pair_map_eq_actual_action random next-pair product law source for generated-action conditional mean-zero. this is a slightly more trajectory-facing consumer than `..._pair_map_eq_actual_action`: the law hypothesis may state the conditional pushforward of the fully random next pair `(action y (i+1), reward y (i+1))`. the generated-action trace freezes the action coordinate, after which the existing actual-action pair-product consumer handles the centered-reward conditional expectation. theorem compiled","shard":"modules/c84d405641caee32.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure_of_condSubgaussian","label":"historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure_of_condSubgaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure_of_condSubgaussian","description":"Canonical successor conditional mean-zero with no caller integrability premise. The canonical conditional MGF theorem supplies exponential integrability, and `HasCondSubgaussianMGF.integrable` lowers it to the ordinary integrability needed by `condExp`.","url":"../modules/banditrlproof-conditionalrewardfoundation/index.html#decl-52092eeba913","parent":"module:BanditRLProof.ConditionalRewardFoundation","order":3394,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardFoundation"],["Source","BanditRLProof/ConditionalRewardFoundation.lean:28"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure_of_condSubgaussian {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hmean : Measurable (fun pair : Prod Context Action => mean…","missing":[],"search":"historystepkernelfamily_centeredreward_succ_condexp_eq_zero_trajmeasure_of_condsubgaussian banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredreward_succ_condexp_eq_zero_trajmeasure_of_condsubgaussian canonical successor conditional mean-zero with no caller integrability premise. the canonical conditional mgf theorem supplies exponential integrability, and `hascondsubgaussianmgf.integrable` lowers it to the ordinary integrability needed by `condexp`. theorem compiled","shard":"modules/cac8bcf91b5e47d4.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_conditionalRewardFoundation_trajMeasure","label":"historyStepKernelFamily_conditionalRewardFoundation_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_conditionalRewardFoundation_trajMeasure","description":"The canonical conditional reward foundation as one theorem-facing endpoint. For every successor time it exposes both conditional mean zero and the conditional MGF witness. The same assumptions also yield the finite-sum Azuma-Hoeffding upper tail for the zero-initialized centered process. Thus `Finset.range n` contains the deterministic slot `Y 0 = 0` and successor rewards `Y 1, ..., Y (n - 1)`. When the cumulative p…","url":"../modules/banditrlproof-conditionalrewardfoundation/index.html#decl-807e3fe453b3","parent":"module:BanditRLProof.ConditionalRewardFoundation","order":3395,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardFoundation"],["Source","BanditRLProof/ConditionalRewardFoundation.lean:139"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_conditionalRewardFoundation_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (defaultAct…","missing":[],"search":"historystepkernelfamily_conditionalrewardfoundation_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_conditionalrewardfoundation_trajmeasure the canonical conditional reward foundation as one theorem-facing endpoint. for every successor time it exposes both conditional mean zero and the conditional mgf witness. the same assumptions also yield the finite-sum azuma-hoeffding upper tail for the zero-initialized centered process. thus `finset.range n` contains the deterministic slot `y 0 = 0` and successor rewards `y 1, ..., y (n - 1)`. when the cumulative proxy is zero, lean's totalized division makes the displayed exponential bound equal to `1`; this endpoint does not claim a sharper degenerate-variance bound. theorem compiled","shard":"modules/cac8bcf91b5e47d4.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.eventually_ae_trim_of_eq_measurableSpace","label":"eventually_ae_trim_of_eq_measurableSpace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.eventually_ae_trim_of_eq_measurableSpace","description":"private theorem eventually_ae_trim_of_eq_measurableSpace {Omega : Type u} {m0 m1 m2 : MeasurableSpace Omega} {mu : @MeasureTheory.Measure Omega m0} (h : m1 = m2) (hm1 : m1 <= m0) {hm2 : m2 <= m0} {p : Omega -> Prop} (hp : Filter.Eventually p (ae (mu.trim hm2))) : Filter.Eventually p (ae (mu.trim hm1))","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4c83fb6aa9a5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3396,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"private theorem eventually_ae_trim_of_eq_measurableSpace {Omega : Type u} {m0 m1 m2 : MeasurableSpace Omega} {mu : @MeasureTheory.Measure Omega m0} (h : m1 = m2) (hm1 : m1 <= m0) {hm2 : m2 <= m0} {p : Omega -> Prop} (hp : Filter.Eventually p (ae (mu.trim hm2))) : Filter.Eventually p (ae (mu.trim hm1))","missing":[],"search":"eventually_ae_trim_of_eq_measurablespace banditrlproof.conditionalexpectationreward.eventually_ae_trim_of_eq_measurablespace private theorem eventually_ae_trim_of_eq_measurablespace {omega : type u} {m0 m1 m2 : measurablespace omega} {mu : @measuretheory.measure omega m0} (h : m1 = m2) (hm1 : m1 <= m0) {hm2 : m2 <= m0} {p : omega -> prop} (hp : filter.eventually p (ae (mu.trim hm2))) : filter.eventually p (ae (mu.trim hm1)) theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory","label":"generatedActionFromRewardHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory","description":"The shifted policy-generated action trace whose state reads the finite reward history of the ambient reward trace.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2cd7a59d6516","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3397,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:39"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionFromRewardHistory {Omega : Type u} {State : Type w} {Action : Type x} [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) : Omega -> ActionTrace Action","missing":[],"search":"generatedactionfromrewardhistory banditrlproof.conditionalexpectationreward.generatedactionfromrewardhistory the shifted policy-generated action trace whose state reads the finite reward history of the ambient reward trace. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_measurable","label":"generatedActionFromRewardHistory_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_measurable","description":"Timewise measurable reward traces plus measurable reward-history state extractors make the generated reward-history action trace timewise measurable.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-39c6be00870d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3398,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:56"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionFromRewardHistory_measurable {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hstate : forall n : Nat, Measurable (state n)) (t : Nat) : Measurable (fun omega : Omega => generatedActionFromRewardHistory policy state defaultAction reward omega t)","missing":[],"search":"generatedactionfromrewardhistory_measurable banditrlproof.conditionalexpectationreward.generatedactionfromrewardhistory_measurable timewise measurable reward traces plus measurable reward-history state extractors make the generated reward-history action trace timewise measurable. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.comap_finitePairHistoryOfTrace_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","label":"comap_finitePairHistoryOfTrace_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.comap_finitePairHistoryOfTrace_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","description":"For actions generated deterministically from finite reward histories, the finite pair prefix and the finite reward prefix generate the same measurable space. The nontrivial inclusion is that every generated action coordinate through time `n` is measurable from the reward prefix through time `n`; the reverse inclusion is the measurable reward projection from pair histories.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4a80609eb190","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3399,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:92"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem comap_finitePairHistoryOfTrace_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace {Omega : Type u} {State : Type w} {Action : Type x} [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (n : Nat) : ((inferInstance : MeasurableSpace ((j : Finset.Iic n) -> Prod Action Rat)).comap (fun omega : Omega => History.finitePairHistoryOfTrace (generatedActionFromRewardHistory policy state defaultAction reward omega) (reward omega) n)) = ((inferInstance : MeasurableSpace ((j : Finset.Iic n) -> Rat)).comap (fun omega : Omega => History.finiteRewardHistoryOfTrace (reward omega) n))","missing":[],"search":"comap_finitepairhistoryoftrace_generatedactionfromrewardhistory_eq_comap_finiterewardhistoryoftrace banditrlproof.conditionalexpectationreward.comap_finitepairhistoryoftrace_generatedactionfromrewardhistory_eq_comap_finiterewardhistoryoftrace for actions generated deterministically from finite reward histories, the finite pair prefix and the finite reward prefix generate the same measurable space. the nontrivial inclusion is that every generated action coordinate through time `n` is measurable from the reward prefix through time `n`; the reverse inclusion is the measurable reward projection from pair histories. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyFiltrationSucc_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","label":"historyFiltrationSucc_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyFiltrationSucc_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","description":"The shifted generated action/reward history filtration is therefore exactly the comap of the finite reward prefix. This removes the generated action coordinates from the conditioning surface without changing the filtration.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5257e608e65a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3400,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:241"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyFiltrationSucc_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (n : Nat) : ((History.historyFiltrationSucc (mOmega := mOmega) (generatedActionFromRewardHistory policy state defaultAction reward) reward (generatedActionFromRewardHistory_measurable (mOmega := mOmega) (policy := policy) (state := state) (defaultAction := defaultAction) (reward := reward) hreward hstate) hrewa…","missing":[],"search":"historyfiltrationsucc_generatedactionfromrewardhistory_eq_comap_finiterewardhistoryoftrace banditrlproof.conditionalexpectationreward.historyfiltrationsucc_generatedactionfromrewardhistory_eq_comap_finiterewardhistoryoftrace the shifted generated action/reward history filtration is therefore exactly the comap of the finite reward prefix. this removes the generated action coordinates from the conditioning surface without changing the filtration. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_succ_measurable_historyFiltrationSucc","label":"generatedActionFromRewardHistory_succ_measurable_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_succ_measurable_historyFiltrationSucc","description":"The generated action at time `i + 1` is measurable with respect to the generated history filtration at time `i`. This is the predictable-action surface needed to mask successor rewards by a fixed arm before applying conditional concentration.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a4282fbae942","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3401,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:286"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionFromRewardHistory_succ_measurable_historyFiltrationSucc {Omega : Type u} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) : @Measurable Omega Action ((History.historyFiltrationSucc (mOmega := mOmega) (generatedActionFromRewardHistory policy state defaultAction reward) reward (generatedActionFromRewardHistory_measurable (mOmega := mOmega) (policy := policy) (state := state) (defaultAction := defaultAction) (reward := reward) hreward hstate)…","missing":[],"search":"generatedactionfromrewardhistory_succ_measurable_historyfiltrationsucc banditrlproof.conditionalexpectationreward.generatedactionfromrewardhistory_succ_measurable_historyfiltrationsucc the generated action at time `i + 1` is measurable with respect to the generated history filtration at time `i`. this is the predictable-action surface needed to mask successor rewards by a fixed arm before applying conditional concentration. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace","label":"historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace","description":"Canonical reward-only `trajMeasure` selected-reward law on the generated finite-pair conditioning surface. The underlying trajectory contains only reward coordinates. The previous comap equality shows that adjoining the deterministically generated policy actions to each finite prefix does not change the conditioning measurable space, so the canonical reward-only law can be stated directly in the same finite-pair not…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-53fcb086af02","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3402,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:343"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (n : Nat) : let stepKernel := RewardKernel.historyStepKernelFamily rewardKernel policy context state hcontext hstate let trajMeasure := ProbabilityTheory.Kernel.trajMeasure (X := fun _ : Nat =>…","missing":[],"search":"historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure_generatedactionfromrewardhistory_finitepairhistoryoftrace banditrlproof.conditionalexpectationreward.historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure_generatedactionfromrewardhistory_finitepairhistoryoftrace canonical reward-only `trajmeasure` selected-reward law on the generated finite-pair conditioning surface. the underlying trajectory contains only reward coordinates. the previous comap equality shows that adjoining the deterministically generated policy actions to each finite prefix does not change the conditioning measurable space, so the canonical reward-only law can be stated directly in the same finite-pair notation used by the ambient source contracts. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace_trim","label":"historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace_trim","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace_trim","description":"Trim-a.e. canonical reward-only `trajMeasure` selected-reward law on the generated finite-pair conditioning surface. The trim-aware canonical reward law is first proved on the finite reward-prefix comap. Deterministically generated policy actions do not enlarge that comap, so the same law holds on the finite pair-prefix sigma-algebra used by the generated selected-reward source contract.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-98d64d93117e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3403,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:420"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace_trim {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (n : Nat) : let stepKernel := RewardKernel.historyStepKernelFamily rewardKernel policy context state hcontext hstate let trajMeasure := ProbabilityTheory.Kernel.trajMeasure (X := fun _ : Na…","missing":[],"search":"historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure_generatedactionfromrewardhistory_finitepairhistoryoftrace_trim banditrlproof.conditionalexpectationreward.historystepkernelfamily_selectedmeasure_condexpkernel_map_trajmeasure_generatedactionfromrewardhistory_finitepairhistoryoftrace_trim trim-a.e. canonical reward-only `trajmeasure` selected-reward law on the generated finite-pair conditioning surface. the trim-aware canonical reward law is first proved on the finite reward-prefix comap. deterministically generated policy actions do not enlarge that comap, so the same law holds on the finite pair-prefix sigma-algebra used by the generated selected-reward source contract. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionPartialTrajectoryPairLawSource","label":"GeneratedActionPartialTrajectoryPairLawSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionPartialTrajectoryPairLawSource","description":"Explicit source contract for the remaining generated-history `partialTraj`/`condExpKernel` pair-law gap. The field is exactly the full finite pair-trace law used by the downstream `partialTraj` consumers, specialized to the definitional generated action trace `generatedActionFromRewardHistory`. It does not prove that law from a global trajectory measure; it only gives future disintegration work a named target.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-16893259ca32","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3404,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:578"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionPartialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactionpartialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource explicit source contract for the remaining generated-history `partialtraj`/`condexpkernel` pair-law gap. the field is exactly the full finite pair-trace law used by the downstream `partialtraj` consumers, specialized to the definitional generated action trace `generatedactionfromrewardhistory`. it does not prove that law from a global trajectory measure; it only gives future disintegration work a named target. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_partialTrajectoryPairLawSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_partialTrajectoryPairLawSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_partialTrajectoryPairLawSource","description":"Expose the exact generated-history finite-pair `partialTraj` law stored in a `GeneratedActionPartialTrajectoryPairLawSource`. This is a source-projection wrapper for `COND-EXPECT-REWARD-PARTIALTRAJ-CONDEXPKERNEL-PAIR-LAW-CARD`: it does not construct the law from a global trajectory/disintegration argument, but it gives downstream consumers a named theorem with the card's full Lean-facing target.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-398e9468e22c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3405,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:657"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionPartialTrajectoryPair…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_partialtrajectorypairlawsource expose the exact generated-history finite-pair `partialtraj` law stored in a `generatedactionpartialtrajectorypairlawsource`. this is a source-projection wrapper for `cond-expect-reward-partialtraj-condexpkernel-pair-law-card`: it does not construct the law from a global trajectory/disintegration argument, but it gives downstream consumers a named theorem with the card's full lean-facing target. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_partialTrajectoryKernel_extend_map_eq","label":"generatedActionPartialTrajectoryPairLawSource_of_partialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_partialTrajectoryKernel_extend_map_eq","description":"Build the generated-history finite-pair `partialTraj` source from the narrower frozen-prefix extension-map law. This packages the existing extension-to-full-trace adapter at the source layer: future disintegration work may prove only the conditional law of extending the already frozen `i`-prefix by the random successor pair, and this constructor turns that into the full `GeneratedActionPartialTrajectoryPairLawSource…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8d04b73de2a7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3406,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:737"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_partialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => rewa…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_partialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_partialtrajectorykernel_extend_map_eq build the generated-history finite-pair `partialtraj` source from the narrower frozen-prefix extension-map law. this packages the existing extension-to-full-trace adapter at the source layer: future disintegration work may prove only the conditional law of extending the already frozen `i`-prefix by the random successor pair, and this constructor turns that into the full `generatedactionpartialtrajectorypairlawsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionPartialTrajectoryPairLawSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the generated-history finite-pair `partialTraj` source from a canonical history-step next-pair law. This lowers the future source obligation once more: a caller may identify the generated conditional law of the next `(action, reward)` pair with `RewardKernel.actionRewardHistoryStepKernelFamily`; the existing next-pair to extension-map adapter and the source constructor above then provide the full finite-pair `…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d2b941e6b4e9","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3407,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:847"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Ome…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the generated-history finite-pair `partialtraj` source from a canonical history-step next-pair law. this lowers the future source obligation once more: a caller may identify the generated conditional law of the next `(action, reward)` pair with `rewardkernel.actionrewardhistorystepkernelfamily`; the existing next-pair to extension-map adapter and the source constructor above then provide the full finite-pair `partialtraj` source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_action_ae_eq_policy_reward_map_eq","label":"generatedActionPartialTrajectoryPairLawSource_of_action_ae_eq_policy_reward_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_action_ae_eq_policy_reward_map_eq","description":"Build the generated-history finite-pair `partialTraj` source from split next-pair laws. This is the source-level version of the split next-pair route: a caller may separately supply the generated action conditional a.e. law and the policy-selected reward-coordinate map law, and the local split-law builder combines them into the history-step pair law consumed by `generatedActionPartialTrajectoryPairLawSource_of_actio…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8090f31be82f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3408,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:959"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_action_ae_eq_policy_reward_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward o…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_action_ae_eq_policy_reward_map_eq banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_action_ae_eq_policy_reward_map_eq build the generated-history finite-pair `partialtraj` source from split next-pair laws. this is the source-level version of the split next-pair route: a caller may separately supply the generated action conditional a.e. law and the policy-selected reward-coordinate map law, and the local split-law builder combines them into the history-step pair law consumed by `generatedactionpartialtrajectorypairlawsource_of_actionrewardhistorystepkernelfamily_pair_map_eq`. it still does not prove either split law from an ambient trajectory disintegration argument. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_selected_policy","label":"generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_selected_policy","description":"Build the generated-history finite-pair `partialTraj` source from only the policy-selected reward-coordinate law. For `generatedActionFromRewardHistory`, the action side of the split next-pair route is supplied by the shifted generated-trace action-freezing theorem. This constructor therefore leaves only the selected reward-coordinate `condExpKernel` map law as the explicit law input before reusing `generatedActionP…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b73b83b0ffe7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3409,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1155"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_reward_map_eq_selected_policy build the generated-history finite-pair `partialtraj` source from only the policy-selected reward-coordinate law. for `generatedactionfromrewardhistory`, the action side of the split next-pair route is supplied by the shifted generated-trace action-freezing theorem. this constructor therefore leaves only the selected reward-coordinate `condexpkernel` map law as the explicit law input before reusing `generatedactionpartialtrajectorypairlawsource_of_action_ae_eq_policy_reward_map_eq`. it still does not prove that reward-coordinate law from an ambient trajectory disintegration argument. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_selected_policy","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_selected_policy","description":"Consume a policy-selected reward-coordinate law stated at the generated reward history surface and expose the full generated finite-pair `partialTraj` law directly. This is the theorem-shaped wrapper around `generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_selected_policy`. It keeps the same reward-history selected-law contract and only projects the stored full finite-pair law from the generated partia…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3757a31a3cbf","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3410,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1273"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Na…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_reward_map_eq_selected_policy consume a policy-selected reward-coordinate law stated at the generated reward history surface and expose the full generated finite-pair `partialtraj` law directly. this is the theorem-shaped wrapper around `generatedactionpartialtrajectorypairlawsource_of_reward_map_eq_selected_policy`. it keeps the same reward-history selected-law contract and only projects the stored full finite-pair law from the generated partial-trajectory source; it does not prove the selected reward-coordinate law from an ambient trajectory construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_actual_action","label":"generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_actual_action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_actual_action","description":"Build the generated-history finite-pair `partialTraj` source from an actual-action reward-coordinate law stated directly for `generatedActionFromRewardHistory`. This removes the explicit action trace and generated-trace equality required by the generic generated-action theorem. The remaining law input is still the actual successor-action reward-coordinate `condExpKernel` map law; this wrapper does not prove that law…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5df32e7cf0ec","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3411,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1411"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_reward_map_eq_actual_action build the generated-history finite-pair `partialtraj` source from an actual-action reward-coordinate law stated directly for `generatedactionfromrewardhistory`. this removes the explicit action trace and generated-trace equality required by the generic generated-action theorem. the remaining law input is still the actual successor-action reward-coordinate `condexpkernel` map law; this wrapper does not prove that law from a global trajectory/disintegration argument. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_actual_action","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_actual_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_actual_action","description":"Consume an actual-action reward-coordinate law stated at the generated reward history surface and expose the full generated finite-pair `partialTraj` law directly. This is the theorem-shaped wrapper around `generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_actual_action`. It keeps the same reward-history actual-action law contract and only projects the stored full finite-pair law from the generated part…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-21eba1f7fdb4","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3412,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1500"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat,…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_reward_map_eq_actual_action consume an actual-action reward-coordinate law stated at the generated reward history surface and expose the full generated finite-pair `partialtraj` law directly. this is the theorem-shaped wrapper around `generatedactionpartialtrajectorypairlawsource_of_reward_map_eq_actual_action`. it keeps the same reward-history actual-action law contract and only projects the stored full finite-pair law from the generated partial-trajectory source. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_finitePairHistory_reward_map_eq_selected_policy","label":"generatedActionPartialTrajectoryPairLawSource_of_finitePairHistory_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_finitePairHistory_reward_map_eq_selected_policy","description":"Build the generated-history finite-pair `partialTraj` source from a policy-selected reward-coordinate law stated at the generated finite pair prefix. Future disintegration work often naturally phrases the selected reward law with `History.finitePairHistoryOfTrace`. This adapter rewrites that prefix through `History.pairHistoryRewardProjection_finitePairHistoryOfTrace` and reuses the existing reward-history source co…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-467b1436a5c3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3413,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1638"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_finitePairHistory_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Ome…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_finitepairhistory_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_finitepairhistory_reward_map_eq_selected_policy build the generated-history finite-pair `partialtraj` source from a policy-selected reward-coordinate law stated at the generated finite pair prefix. future disintegration work often naturally phrases the selected reward law with `history.finitepairhistoryoftrace`. this adapter rewrites that prefix through `history.pairhistoryrewardprojection_finitepairhistoryoftrace` and reuses the existing reward-history source constructor. it still assumes the ambient reward-coordinate law; it does not prove it. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionSelectedRewardFinitePairHistoryLawSource","label":"GeneratedActionSelectedRewardFinitePairHistoryLawSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionSelectedRewardFinitePairHistoryLawSource","description":"Explicit source contract for the generated ambient selected-reward law stated at the generated finite pair prefix. This is a narrower source surface than `GeneratedActionPartialTrajectoryPairLawSource`: it packages only the reward-coordinate `condExpKernel.map` law. The generated action side is later supplied by the definitional shifted-policy trace, so this source can be converted into the full finite-pair `partial…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d5e04c99cc37","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3414,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1725"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionSelectedRewardFinitePairHistoryLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource explicit source contract for the generated ambient selected-reward law stated at the generated finite pair prefix. this is a narrower source surface than `generatedactionpartialtrajectorypairlawsource`: it packages only the reward-coordinate `condexpkernel.map` law. the generated action side is later supplied by the definitional shifted-policy trace, so this source can be converted into the full finite-pair `partialtraj` law source without assuming a separate action law. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_reward_map_eq_selected_policy","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_reward_map_eq_selected_policy","description":"Build the selected-reward finite-pair-history source from the same `condExpKernel.map` law stated with the conditioning sigma-algebra as the comap of `History.finitePairHistoryOfTrace`. This is the source-layer consumer for `History.historyFiltrationSucc_eq_comap_finitePairHistoryOfTrace`: future trajectory/disintegration work can target the Mathlib-style finite-prefix comap conditioning surface and then enter the e…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9e04d9e962ec","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3415,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1796"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_comap_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_comap_reward_map_eq_selected_policy build the selected-reward finite-pair-history source from the same `condexpkernel.map` law stated with the conditioning sigma-algebra as the comap of `history.finitepairhistoryoftrace`. this is the source-layer consumer for `history.historyfiltrationsucc_eq_comap_finitepairhistoryoftrace`: future trajectory/disintegration work can target the mathlib-style finite-prefix comap conditioning surface and then enter the existing generated-history source route without restating the law at `history.historyfiltrationsucc`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_trim_reward_map_eq_selected_policy","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_trim_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_trim_reward_map_eq_selected_policy","description":"Build the selected-reward finite-pair-history source from the same selected-reward law, but with the a.e. filter stated directly on the comap finite-pair-prefix sigma-algebra. This is the Mathlib-facing source entry: a future disintegration proof can state the law entirely at the finite-prefix comap conditioning surface. The adapter only rewrites that surface through `History.historyFiltrationSucc_eq_comap_finitePai…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7e62733dad45","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3416,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:1970"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_trim_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_comap_trim_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_comap_trim_reward_map_eq_selected_policy build the selected-reward finite-pair-history source from the same selected-reward law, but with the a.e. filter stated directly on the comap finite-pair-prefix sigma-algebra. this is the mathlib-facing source entry: a future disintegration proof can state the law entirely at the finite-prefix comap conditioning surface. the adapter only rewrites that surface through `history.historyfiltrationsucc_eq_comap_finitepairhistoryoftrace`; it still consumes, rather than proves, the selected-reward law. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_trajMeasure","label":"historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_trajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_trajMeasure","description":"Construct the generated selected-reward finite-pair-history source for the canonical reward-only `trajMeasure`. Unlike the generic source adapters above, this constructor discharges the selected-reward law from Mathlib's `Kernel.trajMeasure` conditional-distribution result, routed through the trim-aware countable-target bridge. The remaining action assumptions are exactly those needed by the generated finite-pair fi…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b7d3ca2e0a85","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3417,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2144"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) : let stepKernel := RewardKernel.historyStepKernelFamily rewardKernel policy context state hcontext hstate let trajMeasure := ProbabilityTheory.Kernel.trajMeasure (X := fun _ : Nat…","missing":[],"search":"historystepkernelfamily_generatedactionselectedrewardfinitepairhistorylawsource_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_generatedactionselectedrewardfinitepairhistorylawsource_trajmeasure construct the generated selected-reward finite-pair-history source for the canonical reward-only `trajmeasure`. unlike the generic source adapters above, this constructor discharges the selected-reward law from mathlib's `kernel.trajmeasure` conditional-distribution result, routed through the trim-aware countable-target bridge. the remaining action assumptions are exactly those needed by the generated finite-pair filtration surface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_identDistrib_trajMeasure","label":"historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_identDistrib_trajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_identDistrib_trajMeasure","description":"Construct the generated selected-reward source on an ambient probability space whose complete reward trace is identically distributed with the canonical reward-only `trajMeasure`. The global `IdentDistrib` assumption is converted to the finite-prefix selected-reward law by the disintegration transport theorem. Deterministically generated actions do not enlarge the reward-prefix sigma-algebra, so the law then enters…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e13a10f79dce","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3418,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2211"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_identDistrib_trajMeasure {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hrewar…","missing":[],"search":"historystepkernelfamily_generatedactionselectedrewardfinitepairhistorylawsource_of_identdistrib_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_generatedactionselectedrewardfinitepairhistorylawsource_of_identdistrib_trajmeasure construct the generated selected-reward source on an ambient probability space whose complete reward trace is identically distributed with the canonical reward-only `trajmeasure`. the global `identdistrib` assumption is converted to the finite-prefix selected-reward law by the disintegration transport theorem. deterministically generated actions do not enlarge the reward-prefix sigma-algebra, so the law then enters the existing finite-pair source surface without an ambient `conddistrib` or `condexpkernel.map` assumption. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_condDistrib","label":"historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_condDistrib","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_condDistrib","description":"Construct the ambient generated selected-reward source from the recursive reward-process contract: the initial reward has law `mu0`, and each successor reward has the configured history-step conditional distribution. The recursive assumptions first identify the complete reward trace with the canonical `trajMeasure`; the existing distribution-transport constructor then supplies the selected-reward finite-pair-history…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-25b25c4972a6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3419,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2297"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t…","missing":[],"search":"historystepkernelfamily_generatedactionselectedrewardfinitepairhistorylawsource_of_conddistrib banditrlproof.conditionalexpectationreward.historystepkernelfamily_generatedactionselectedrewardfinitepairhistorylawsource_of_conddistrib construct the ambient generated selected-reward source from the recursive reward-process contract: the initial reward has law `mu0`, and each successor reward has the configured history-step conditional distribution. the recursive assumptions first identify the complete reward trace with the canonical `trajmeasure`; the existing distribution-transport constructor then supplies the selected-reward finite-pair-history law. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_selectedRewardFinitePairHistoryLawSource","label":"generatedActionPartialTrajectoryPairLawSource_of_selectedRewardFinitePairHistoryLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_selectedRewardFinitePairHistoryLawSource","description":"Convert the generated ambient selected-reward finite-pair-history source into the full generated finite-pair `partialTraj` source. This uses the existing generated-action split route; it still consumes the selected-reward source field rather than proving it.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-caea7473dea6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3420,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2353"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_selectedRewardFinitePairHistoryLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionSelectedRewardFinitePairHistoryLawSource mu rewardKernel po…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_selectedrewardfinitepairhistorylawsource banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_selectedrewardfinitepairhistorylawsource convert the generated ambient selected-reward finite-pair-history source into the full generated finite-pair `partialtraj` source. this uses the existing generated-action split route; it still consumes the selected-reward source field rather than proving it. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_of_condDistrib","label":"historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_of_condDistrib","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_of_condDistrib","description":"Construct the full ambient generated finite-pair `partialTraj` source from an initial reward law and successor conditional-distribution recursion. The selected-reward coordinate is obtained from the complete reward-trace law; the action coordinate is supplied by the existing deterministic generated policy split.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7c611a8d17c0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3421,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2394"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_of_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Mea…","missing":[],"search":"historystepkernelfamily_generatedactionpartialtrajectorypairlawsource_of_conddistrib banditrlproof.conditionalexpectationreward.historystepkernelfamily_generatedactionpartialtrajectorypairlawsource_of_conddistrib construct the full ambient generated finite-pair `partialtraj` source from an initial reward law and successor conditional-distribution recursion. the selected-reward coordinate is obtained from the complete reward-trace law; the action coordinate is supplied by the existing deterministic generated policy split. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_trajMeasure","label":"historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_trajMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_trajMeasure","description":"Construct the full generated finite-pair `partialTraj` source for the canonical reward-only `trajMeasure`. The selected-reward law is supplied by the canonical trim-aware source, while the action coordinate is the deterministic generated policy action. Thus this constructor reaches the full pair-history source without assuming an ambient selected-reward or random-pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0de8e72851a0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3422,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2447"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) : let stepKernel := RewardKernel.historyStepKernelFamily rewardKernel policy context state hcontext hstate let trajMeasure := ProbabilityTheory.Kernel.trajMeasure (X := fun _ : Nat => Rat) mu…","missing":[],"search":"historystepkernelfamily_generatedactionpartialtrajectorypairlawsource_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_generatedactionpartialtrajectorypairlawsource_trajmeasure construct the full generated finite-pair `partialtraj` source for the canonical reward-only `trajmeasure`. the selected-reward law is supplied by the canonical trim-aware source, while the action coordinate is the deterministic generated policy action. thus this constructor reaches the full pair-history source without assuming an ambient selected-reward or random-pair law. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_trajMeasure","label":"historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_trajMeasure","description":"Canonical reward-only `trajMeasure` full generated finite-pair `partialTraj` law. This is the theorem-shaped endpoint of the canonical selected-reward route: the next finite pair prefix under `condExpKernel`, conditioned on the generated finite pair history, has the configured one-step partial-trajectory kernel. No ambient selected-reward or random-pair source assumption remains.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-67398a4edb1d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3423,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2507"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (i : Nat) : let stepKernel := RewardKernel.historyStepKernelFamily rewardKernel policy context state hcontext hstate let trajMeasure := Probabi…","missing":[],"search":"historystepkernelfamily_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_trajmeasure canonical reward-only `trajmeasure` full generated finite-pair `partialtraj` law. this is the theorem-shaped endpoint of the canonical selected-reward route: the next finite pair prefix under `condexpkernel`, conditioned on the generated finite pair history, has the configured one-step partial-trajectory kernel. no ambient selected-reward or random-pair source assumption remains. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure","label":"historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure","description":"Canonical reward-only `trajMeasure` successor centered reward has conditional expectation zero under the generated finite-pair history. The canonical full `partialTraj` law discharges the probability-law identification. Ambient integrability of the centered successor reward remains an explicit regularity contract; this avoids imposing pointwise bounds on every trace in `Nat -> Rat`, including null trajectories outsi…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-24aeb62e199e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3424,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2599"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (defaultAction : Action) (i : Nat) (h_integrable : Integrable (fu…","missing":[],"search":"historystepkernelfamily_centeredreward_succ_condexp_eq_zero_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredreward_succ_condexp_eq_zero_trajmeasure canonical reward-only `trajmeasure` successor centered reward has conditional expectation zero under the generated finite-pair history. the canonical full `partialtraj` law discharges the probability-law identification. ambient integrability of the centered successor reward remains an explicit regularity contract; this avoids imposing pointwise bounds on every trace in `nat -> rat`, including null trajectories outside the generated law's support. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","label":"historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","description":"Canonical reward-only `trajMeasure` successor centered reward has a conditional sub-Gaussian MGF under the generated finite-pair history. The canonical full `partialTraj` law supplies the reward-coordinate conditional law. Measurability follows from the measurable mean surface, and the integrated target-law transfer derives exponential integrability from the selected kernel MGF bounds. A deterministic bound over fin…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1967507777d8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3425,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2719"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hmean : Measurable (fun pair : Prod Context Action => mean…","missing":[],"search":"historystepkernelfamily_centeredreward_succ_hascondsubgaussianmgf_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredreward_succ_hascondsubgaussianmgf_trajmeasure canonical reward-only `trajmeasure` successor centered reward has a conditional sub-gaussian mgf under the generated finite-pair history. the canonical full `partialtraj` law supplies the reward-coordinate conditional law. measurability follows from the measurable mean surface, and the integrated target-law transfer derives exponential integrability from the selected kernel mgf bounds. a deterministic bound over finite reward histories supplies the trim-a.e. variance domination. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource","description":"A generated full finite-pair `partialTraj` source directly yields the successor centered-reward conditional MGF witness. Unlike the older bounded-source adapters, exponential integrability is derived by the integrated target-law transfer from `CenteredRewardKernelLaw`; only measurability of the mean surface and a deterministic selected-history variance bound remain analytic side conditions.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9637815f6063","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3426,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:2956"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : Measure Omega) [IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (defaultAction : Action) (r…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource a generated full finite-pair `partialtraj` source directly yields the successor centered-reward conditional mgf witness. unlike the older bounded-source adapters, exponential integrability is derived by the integrated target-law transfer from `centeredrewardkernellaw`; only measurability of the mean surface and a deterministic selected-history variance bound remain analytic side conditions. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_of_condDistrib","label":"historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_of_condDistrib","description":"Recursive ambient reward-process laws directly yield the successor centered reward conditional MGF witness under generated finite-pair history. The initial marginal and successor `condDistrib` recursion construct the full `partialTraj` source; the source-level integrated transfer then consumes only a measurable mean, centered kernel law, and selected-history variance bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a446a056cd36","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3427,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3164"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_of_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : Rewa…","missing":[],"search":"historystepkernelfamily_centeredreward_succ_hascondsubgaussianmgf_of_conddistrib banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredreward_succ_hascondsubgaussianmgf_of_conddistrib recursive ambient reward-process laws directly yield the successor centered reward conditional mgf witness under generated finite-pair history. the initial marginal and successor `conddistrib` recursion construct the full `partialtraj` source; the source-level integrated transfer then consumes only a measurable mean, centered kernel law, and selected-history variance bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_centeredRewardSuccProcess_stronglyAdapted","label":"generatedActionFromRewardHistory_centeredRewardSuccProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_centeredRewardSuccProcess_stronglyAdapted","description":"The zero-initialized successor centered-reward process is strongly adapted to the generated shifted history filtration. Index zero is reserved for the deterministic value zero. At index `i + 1` the process uses reward `i + 1` centered by the context/action mean selected from the finite reward history through `i`. This indexing matches Mathlib's conditional sub-Gaussian finite-sum tail API.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-750732bb6686","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3428,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3265"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionFromRewardHistory_centeredRewardSuccProcess_stronglyAdapted {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : let action := generatedActionFromRewardHistory policy s…","missing":[],"search":"generatedactionfromrewardhistory_centeredrewardsuccprocess_stronglyadapted banditrlproof.conditionalexpectationreward.generatedactionfromrewardhistory_centeredrewardsuccprocess_stronglyadapted the zero-initialized successor centered-reward process is strongly adapted to the generated shifted history filtration. index zero is reserved for the deterministic value zero. at index `i + 1` the process uses reward `i + 1` centered by the context/action mean selected from the finite reward history through `i`. this indexing matches mathlib's conditional sub-gaussian finite-sum tail api. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_armMaskedCenteredRewardSuccProcess_stronglyAdapted","label":"generatedActionFromRewardHistory_armMaskedCenteredRewardSuccProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_armMaskedCenteredRewardSuccProcess_stronglyAdapted","description":"Masking the successor centered-reward process by selection of a fixed arm preserves strong adaptedness. The arm event at successor index `i + 1` is already measurable at filtration level `i`; monotonicity exposes it at level `i + 1`, where the centered reward it masks is measurable.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b99dc9756c15","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3429,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3428"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionFromRewardHistory_armMaskedCenteredRewardSuccProcess_stronglyAdapted {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (defaultAction arm : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : let action := generatedActionFromRewardHis…","missing":[],"search":"generatedactionfromrewardhistory_armmaskedcenteredrewardsuccprocess_stronglyadapted banditrlproof.conditionalexpectationreward.generatedactionfromrewardhistory_armmaskedcenteredrewardsuccprocess_stronglyadapted masking the successor centered-reward process by selection of a fixed arm preserves strong adaptedness. the arm event at successor index `i + 1` is already measurable at filtration level `i`; monotonicity exposes it at level `i + 1`, where the centered reward it masks is measurable. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_condDistrib","label":"historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_condDistrib","description":"Ambient Azuma-Hoeffding bound for a recursively specified generated reward process. The initial marginal identifies the ambient measure as a probability measure. Successor `condDistrib` laws construct the full `partialTraj` source and hence all conditional MGF witnesses, while generated-history measurability supplies strong adaptedness. The sum contains the zero slot followed by centered rewards at indices `1, ...,…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-209d419b4f92","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3430,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3541"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : Rew…","missing":[],"search":"historystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_of_conddistrib banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_of_conddistrib ambient azuma-hoeffding bound for a recursively specified generated reward process. the initial marginal identifies the ambient measure as a probability measure. successor `conddistrib` laws construct the full `partialtraj` source and hence all conditional mgf witnesses, while generated-history measurability supplies strong adaptedness. the sum contains the zero slot followed by centered rewards at indices `1, ..., n - 1`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure","label":"historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure","description":"Canonical reward-only `trajMeasure` Azuma-Hoeffding bound for the finite sum of successor policy-centered rewards. The zero initial term aligns the successor conditional-MGF witnesses with Mathlib's `Finset.range n` tail theorem. Thus the random sum contains centered rewards at indices `1, ..., n - 1`, and its proxy sum contains the corresponding history-selected deterministic ceilings.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6f068f7fd012","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3431,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3676"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.…","missing":[],"search":"historystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_trajmeasure canonical reward-only `trajmeasure` azuma-hoeffding bound for the finite sum of successor policy-centered rewards. the zero initial term aligns the successor conditional-mgf witnesses with mathlib's `finset.range n` tail theorem. thus the random sum contains centered rewards at indices `1, ..., n - 1`, and its proxy sum contains the corresponding history-selected deterministic ceilings. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_average_tail_ennreal_trajMeasure","label":"historyStepKernelFamily_centeredRewardSuccProcess_average_tail_ennreal_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_average_tail_ennreal_trajMeasure","description":"Canonical reward-only `trajMeasure` Azuma-Hoeffding bound for the empirical average of `m` successor policy-centered rewards. The `Finset.range (m + 1)` sum retains the deterministic zero slot required by the conditional-MGF tail API, so its remaining terms are exactly rewards `1, ..., m`. The strict positivity of `m` is the explicit denominator regularity contract used to rewrite the average-tail event as a sum-tai…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d2542e5f6b41","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3432,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3808"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredRewardSuccProcess_average_tail_ennreal_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 p…","missing":[],"search":"historystepkernelfamily_centeredrewardsuccprocess_average_tail_ennreal_trajmeasure banditrlproof.conditionalexpectationreward.historystepkernelfamily_centeredrewardsuccprocess_average_tail_ennreal_trajmeasure canonical reward-only `trajmeasure` azuma-hoeffding bound for the empirical average of `m` successor policy-centered rewards. the `finset.range (m + 1)` sum retains the deterministic zero slot required by the conditional-mgf tail api, so its remaining terms are exactly rewards `1, ..., m`. the strict positivity of `m` is the explicit denominator regularity contract used to rewrite the average-tail event as a sum-tail event with threshold `m * eps`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_comap_reward_map_eq_selected_policy","label":"generatedActionPartialTrajectoryPairLawSource_of_comap_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_comap_reward_map_eq_selected_policy","description":"Build the full generated finite-pair `partialTraj` source directly from a selected-reward law stated at the finite pair-prefix comap conditioning sigma-algebra. This composes the comap-to-selected-source adapter with the existing selected-source-to-`partialTraj` route, so future disintegration work can target the Mathlib-style comap conditioning surface and immediately obtain the full generated finite-pair source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-338ae768ef95","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3433,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:3924"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_comap_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_comap_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_comap_reward_map_eq_selected_policy build the full generated finite-pair `partialtraj` source directly from a selected-reward law stated at the finite pair-prefix comap conditioning sigma-algebra. this composes the comap-to-selected-source adapter with the existing selected-source-to-`partialtraj` route, so future disintegration work can target the mathlib-style comap conditioning surface and immediately obtain the full generated finite-pair source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_comap_trim_reward_map_eq_selected_policy","label":"generatedActionPartialTrajectoryPairLawSource_of_comap_trim_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_comap_trim_reward_map_eq_selected_policy","description":"Build the full generated finite-pair `partialTraj` source directly from a selected-reward law stated with both the conditioning sigma-algebra and the trim filter at the finite pair-prefix comap surface. This is the direct partialTraj entry for Mathlib-facing disintegration work: it composes the comap-trim selected-source adapter with the existing selected-source-to-`partialTraj` route, and still consumes rather than…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2bda65e5cb3c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3434,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4016"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_comap_trim_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => r…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_comap_trim_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_comap_trim_reward_map_eq_selected_policy build the full generated finite-pair `partialtraj` source directly from a selected-reward law stated with both the conditioning sigma-algebra and the trim filter at the finite pair-prefix comap surface. this is the direct partialtraj entry for mathlib-facing disintegration work: it composes the comap-trim selected-source adapter with the existing selected-source-to-`partialtraj` route, and still consumes rather than proves the selected-reward law. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_selectedRewardFinitePairHistoryLawSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_selectedRewardFinitePairHistoryLawSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_selectedRewardFinitePairHistoryLawSource","description":"Consume the generated selected-reward finite-pair-history source at the exact full finite-pair `partialTraj` law surface. This is the Lean-facing theorem-card target specialized to the generated reward-history action trace. It still consumes the selected-reward source field; the ambient disintegration proof of that field remains separate.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1368c5f04cf2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3435,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4160"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_selectedRewardFinitePairHistoryLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionSelectedRew…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_selectedrewardfinitepairhistorylawsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_selectedrewardfinitepairhistorylawsource consume the generated selected-reward finite-pair-history source at the exact full finite-pair `partialtraj` law surface. this is the lean-facing theorem-card target specialized to the generated reward-history action trace. it still consumes the selected-reward source field; the ambient disintegration proof of that field remains separate. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_reward_map_eq_selected_policy","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_reward_map_eq_selected_policy","description":"Consume a selected-reward law stated at the finite pair-prefix comap conditioning surface and expose the full generated finite-pair `partialTraj` law directly. This is a theorem-shaped wrapper around `generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_reward_map_eq_selected_policy` and `actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_selectedRewardFinitePair…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a742f648b5e8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3436,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4265"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_comap_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_comap_reward_map_eq_selected_policy consume a selected-reward law stated at the finite pair-prefix comap conditioning surface and expose the full generated finite-pair `partialtraj` law directly. this is a theorem-shaped wrapper around `generatedactionselectedrewardfinitepairhistorylawsource_of_comap_reward_map_eq_selected_policy` and `actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_selectedrewardfinitepairhistorylawsource`. it still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_trim_reward_map_eq_selected_policy","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_trim_reward_map_eq_selected_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_trim_reward_map_eq_selected_policy","description":"Consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, and expose the full generated finite-pair `partialTraj` law directly. This is the theorem-shaped companion to `generatedActionPartialTrajectoryPairLawSource_of_comap_trim_reward_map_eq_selected_policy`. It still consumes the selected-reward conditional law; it does not prove that law…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8442b0cbfdc0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3437,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4410"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_trim_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : fo…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_comap_trim_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_comap_trim_reward_map_eq_selected_policy consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, and expose the full generated finite-pair `partialtraj` law directly. this is the theorem-shaped companion to `generatedactionpartialtrajectorypairlawsource_of_comap_trim_reward_map_eq_selected_policy`. it still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairMapSource","label":"GeneratedActionRandomPairMapSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairMapSource","description":"A generated-policy conditional next-pair law source. The action trace is generated by the shifted policy from the reward-history state, and at each step the conditional law of the random next `(action, reward)` pair is the selected reward law pushed through `Prod.mk` at the actual next action. This is a contract surface: it packages the remaining law input, but does not prove it from a global trajectory measure.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a3ef0c25ee5b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3438,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4608"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactionrandompairmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource a generated-policy conditional next-pair law source. the action trace is generated by the shifted policy from the reward-history state, and at each step the conditional law of the random next `(action, reward)` pair is the selected reward law pushed through `prod.mk` at the actual next action. this is a contract surface: it packages the remaining law input, but does not prove it from a global trajectory measure. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionActualRewardMapSource","label":"GeneratedActionActualRewardMapSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionActualRewardMapSource","description":"Generated-policy conditional reward-coordinate law source. This is weaker than `GeneratedActionRandomPairMapSource`: it assumes only the conditional law of the next reward coordinate under the actual next action, not the full conditional law of the random `(action, reward)` pair. The action trace is still required to be the shifted policy-generated trace over finite reward histories.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-091abfaafbae","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3439,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4661"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactionactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource generated-policy conditional reward-coordinate law source. this is weaker than `generatedactionrandompairmapsource`: it assumes only the conditional law of the next reward coordinate under the actual next action, not the full conditional law of the random `(action, reward)` pair. the action trace is still required to be the shifted policy-generated trace over finite reward histories. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionDefinitionalActualRewardMapSource","label":"GeneratedActionDefinitionalActualRewardMapSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionDefinitionalActualRewardMapSource","description":"Definitional generated-policy conditional reward-coordinate law source. This is the definitional-action version of `GeneratedActionActualRewardMapSource`: the action trace is fixed to `generatedActionFromRewardHistory`, and its timewise measurability is derived from measurable reward-history state extractors plus timewise reward measurability. The reward-coordinate conditional law remains a source contract.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-18205e0f1e73","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3440,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4713"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource definitional generated-policy conditional reward-coordinate law source. this is the definitional-action version of `generatedactionactualrewardmapsource`: the action trace is fixed to `generatedactionfromrewardhistory`, and its timewise measurability is derived from measurable reward-history state extractors plus timewise reward measurability. the reward-coordinate conditional law remains a source contract. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_reward_map_eq_selected_policy","label":"generatedActionDefinitionalActualRewardMapSource_of_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_reward_map_eq_selected_policy","description":"Build the definitional actual reward-coordinate source from the same reward-coordinate law stated at the policy-selected action. For `generatedActionFromRewardHistory`, the successor coordinate `i + 1` is definitionally `(policy i).action` applied to the finite reward-history state. This adapter lets callers use that policy-facing law surface without first rewriting it to the generated successor action.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d5323bad9edd","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3441,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4772"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (h_reward_map_eq_policy : forall i : Nat, Fi…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_reward_map_eq_selected_policy build the definitional actual reward-coordinate source from the same reward-coordinate law stated at the policy-selected action. for `generatedactionfromrewardhistory`, the successor coordinate `i + 1` is definitionally `(policy i).action` applied to the finite reward-history state. this adapter lets callers use that policy-facing law surface without first rewriting it to the generated successor action. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_definitionalActualRewardMapSource","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_definitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_definitionalActualRewardMapSource","description":"Convert a definitional actual-action reward-coordinate source into the generated selected-reward finite-pair-history source. For `generatedActionFromRewardHistory`, the actual successor action is definitionally the policy-selected action at the finite reward history, and `History.pairHistoryRewardProjection_finitePairHistoryOfTrace` removes the irrelevant action coordinates from the generated finite pair prefix. Thi…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a0462f5a688a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3442,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4841"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_definitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionDefi…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_definitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_definitionalactualrewardmapsource convert a definitional actual-action reward-coordinate source into the generated selected-reward finite-pair-history source. for `generatedactionfromrewardhistory`, the actual successor action is definitionally the policy-selected action at the finite reward history, and `history.pairhistoryrewardprojection_finitepairhistoryoftrace` removes the irrelevant action coordinates from the generated finite pair prefix. this adapter exposes the same reward-coordinate law at the finite-pair-history source surface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq","label":"generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq","description":"Build the definitional actual reward-coordinate source from a full finite-pair `partialTraj` law over the generated action trace. This is a packaging step for the ambient trajectory-identification route: once callers identify the conditional kernel of the generated finite pair trace with the one-step action/reward `partialTraj` kernel, the compiled projection route supplies the reward-coordinate law required by `Gen…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-745f9c401228","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3443,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4881"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward o…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorykernel_map_eq build the definitional actual reward-coordinate source from a full finite-pair `partialtraj` law over the generated action trace. this is a packaging step for the ambient trajectory-identification route: once callers identify the conditional kernel of the generated finite pair trace with the one-step action/reward `partialtraj` kernel, the compiled projection route supplies the reward-coordinate law required by `generatedactiondefinitionalactualrewardmapsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryPairLawSource","label":"generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryPairLawSource","description":"Project a generated-history `partialTraj` pair-law source to the definitional actual-action reward-coordinate source. The `partialTraj` source already packages context/state measurability and the full finite-pair trace law. This wrapper exposes the weaker reward-coordinate source interface needed by mean-zero and conditional-MGF consumers without asking callers to unpack the source fields manually.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-98fc99392e34","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3444,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:4988"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionPartialTrajectoryPairLawSource mu rewardKernel policy context stat…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorypairlawsource project a generated-history `partialtraj` pair-law source to the definitional actual-action reward-coordinate source. the `partialtraj` source already packages context/state measurability and the full finite-pair trace law. this wrapper exposes the weaker reward-coordinate source interface needed by mean-zero and conditional-mgf consumers without asking callers to unpack the source fields manually. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_partialTrajectoryPairLawSource","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_partialTrajectoryPairLawSource","description":"Project a generated-history `partialTraj` pair-law source to the selected reward-coordinate source stated at the generated finite-pair-history surface. This is the source-level reverse of the selected-reward-to-`partialTraj` constructor: it first projects the full pair law to the definitional actual-action reward law, then rewrites the generated successor action and finite-pair reward projection into the policy-sele…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7e063acd29fa","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3445,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5032"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionPartialTrajectoryPairLawSource mu rewardKernel policy conte…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_partialtrajectorypairlawsource project a generated-history `partialtraj` pair-law source to the selected reward-coordinate source stated at the generated finite-pair-history surface. this is the source-level reverse of the selected-reward-to-`partialtraj` constructor: it first projects the full pair law to the definitional actual-action reward law, then rewrites the generated successor action and finite-pair reward projection into the policy-selected reward-law source. it does not construct the full pair law or transport canonical `trajmeasure` to an ambient process. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","label":"generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","description":"Build the definitional actual reward-coordinate source from the frozen-prefix extension-map form of the generated finite-pair `partialTraj` law. Compared with `generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq`, the input law only identifies the conditional kernel after appending the random successor pair to the already frozen `i`-prefix. The compiled extension-to-full adapter then s…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-103362869d8b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3446,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5085"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => r…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorykernel_extend_map_eq build the definitional actual reward-coordinate source from the frozen-prefix extension-map form of the generated finite-pair `partialtraj` law. compared with `generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorykernel_map_eq`, the input law only identifies the conditional kernel after appending the random successor pair to the already frozen `i`-prefix. the compiled extension-to-full adapter then supplies the full finite-pair trace law needed by the source constructor. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionDefinitionalActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the definitional actual reward-coordinate source directly from the canonical history-step next-pair law. This is the source-packaging version of the direct history-step reward-map adapter: once callers identify the generated conditional next-pair law with `RewardKernel.actionRewardHistoryStepKernelFamily`, projection through `Prod.snd` supplies the reward-coordinate law stored by `GeneratedActionDefinitionalAc…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-25794dc67bf7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3447,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5202"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the definitional actual reward-coordinate source directly from the canonical history-step next-pair law. this is the source-packaging version of the direct history-step reward-map adapter: once callers identify the generated conditional next-pair law with `rewardkernel.actionrewardhistorystepkernelfamily`, projection through `prod.snd` supplies the reward-coordinate law stored by `generatedactiondefinitionalactualrewardmapsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalMapSource","label":"GeneratedActionRandomPairDefinitionalMapSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalMapSource","description":"Definitional generated-policy conditional next-pair law source. Here the action trace is not a separate parameter with an equality witness: it is definitionally `generatedActionFromRewardHistory`. Timewise action measurability is derived from measurable state extractors and reward traces. The random next-pair law remains a source contract.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-796f1f65792c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3448,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5306"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactionrandompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource definitional generated-policy conditional next-pair law source. here the action trace is not a separate parameter with an equality witness: it is definitionally `generatedactionfromrewardhistory`. timewise action measurability is derived from measurable state extractors and reward traces. the random next-pair law remains a source contract. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","label":"generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","description":"Build the definitional generated-policy random next-pair source from a full finite-pair-trace `partialTraj` law. This is a source-side adapter for the remaining trajectory-law gap: once a future ambient construction identifies the conditional law of the whole `i + 1` finite pair trace with the Mathlib `partialTraj` kernel, the existing `partialTraj` next-coordinate projection and one-step action/reward kernel shape…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e98ab5d65156","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3449,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5372"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega =…","missing":[],"search":"generatedactionrandompairdefinitionalmapsource_of_actionrewardpartialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource_of_actionrewardpartialtrajectorykernel_map_eq build the definitional generated-policy random next-pair source from a full finite-pair-trace `partialtraj` law. this is a source-side adapter for the remaining trajectory-law gap: once a future ambient construction identifies the conditional law of the whole `i + 1` finite pair trace with the mathlib `partialtraj` kernel, the existing `partialtraj` next-coordinate projection and one-step action/reward kernel shape provide the `generatedactionrandompairdefinitionalmapsource` field. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_partialTrajectoryPairLawSource","label":"generatedActionRandomPairDefinitionalMapSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_partialTrajectoryPairLawSource","description":"Consume a generated-history `partialTraj` pair-law source into the existing definitional random next-pair source interface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-567c428bc758","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3450,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5581"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalMapSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionPartialTrajectoryPairLawSource mu rewardKernel policy context state…","missing":[],"search":"generatedactionrandompairdefinitionalmapsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource_of_partialtrajectorypairlawsource consume a generated-history `partialtraj` pair-law source into the existing definitional random next-pair source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","label":"generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","description":"Build the definitional generated-policy random next-pair source from the frozen-prefix extension-map form of the `partialTraj` law. This is one step narrower than `generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_map_eq`: the input law only needs to identify the conditional kernel after appending the random next pair to the already frozen prefix. The existing extension-to-full ad…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-32ae6cdc519c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3451,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5625"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"generatedactionrandompairdefinitionalmapsource_of_actionrewardpartialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource_of_actionrewardpartialtrajectorykernel_extend_map_eq build the definitional generated-policy random next-pair source from the frozen-prefix extension-map form of the `partialtraj` law. this is one step narrower than `generatedactionrandompairdefinitionalmapsource_of_actionrewardpartialtrajectorykernel_map_eq`: the input law only needs to identify the conditional kernel after appending the random next pair to the already frozen prefix. the existing extension-to-full adapter then supplies the full finite-pair-trace law needed by the source constructor above. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionRandomPairDefinitionalMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the definitional generated-policy random next-pair source directly from the canonical history-step next-pair law. This is the source-packaging version of the history-step consumer: once the conditional kernel is identified with `RewardKernel.actionRewardHistoryStepKernelFamily`, the kernel's one-step shape supplies the random-pair map law stored by `GeneratedActionRandomPairDefinitionalMapSource`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-415cc61564d1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3452,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5741"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Om…","missing":[],"search":"generatedactionrandompairdefinitionalmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the definitional generated-policy random next-pair source directly from the canonical history-step next-pair law. this is the source-packaging version of the history-step consumer: once the conditional kernel is identified with `rewardkernel.actionrewardhistorystepkernelfamily`, the kernel's one-step shape supplies the random-pair map law stored by `generatedactionrandompairdefinitionalmapsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalMapSource_of_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_reward_map_eq_selected_policy","description":"Build the definitional generated-policy random next-pair source from a policy-selected reward-coordinate law. The policy-facing reward law is first rewritten to the generated successor action. The existing generated-action reward-map route then supplies the frozen-prefix extension-map `partialTraj` law consumed by the bare `GeneratedActionRandomPairDefinitionalMapSource` constructor.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-881e59eddcee","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3453,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:5894"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalMapSource_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omeg…","missing":[],"search":"generatedactionrandompairdefinitionalmapsource_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource_of_reward_map_eq_selected_policy build the definitional generated-policy random next-pair source from a policy-selected reward-coordinate law. the policy-facing reward law is first rewritten to the generated successor action. the existing generated-action reward-map route then supplies the frozen-prefix extension-map `partialtraj` law consumed by the bare `generatedactionrandompairdefinitionalmapsource` constructor. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_definitionalMapSource","label":"generatedActionRandomPairMapSource_of_definitionalMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_definitionalMapSource","description":"Convert the definitional generated-action source into the existing generated random-pair map source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a04728f9a245","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3454,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6082"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_definitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionRandomPairDefinitionalMapSource mu rewardKernel policy context state defaultAction reward…","missing":[],"search":"generatedactionrandompairmapsource_of_definitionalmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_definitionalmapsource convert the definitional generated-action source into the existing generated random-pair map source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionRandomPairMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the explicit generated-action random next-pair source directly from the canonical history-step next-pair law. This is the explicit-action counterpart of `generatedActionRandomPairDefinitionalMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq`: callers keep an `action` trace and provide the shifted generated-action identity once. The one-step shape of `RewardKernel.actionRewardHistoryStepKernelFamily`…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b1667b83e5de","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3455,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6121"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat,…","missing":[],"search":"generatedactionrandompairmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the explicit generated-action random next-pair source directly from the canonical history-step next-pair law. this is the explicit-action counterpart of `generatedactionrandompairdefinitionalmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq`: callers keep an `action` trace and provide the shifted generated-action identity once. the one-step shape of `rewardkernel.actionrewardhistorystepkernelfamily` then supplies the packaged random-pair map law. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","label":"generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","description":"Build the explicit generated-action random next-pair source from a full finite-pair `partialTraj` law. The full trace law is first projected to the canonical history-step next-pair law; the history-step source constructor then packages the random-pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5140db7ec757","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3456,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6245"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Meas…","missing":[],"search":"generatedactionrandompairmapsource_of_actionrewardpartialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_actionrewardpartialtrajectorykernel_map_eq build the explicit generated-action random next-pair source from a full finite-pair `partialtraj` law. the full trace law is first projected to the canonical history-step next-pair law; the history-step source constructor then packages the random-pair law. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","label":"generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","description":"Build the explicit generated-action random next-pair source from the frozen-prefix extension-map form of the `partialTraj` law. This narrows the source construction surface to the deterministic old-prefix extension map, then reuses the compiled extension-to-full-trace adapter.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1a7790f6fad8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3457,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6344"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Na…","missing":[],"search":"generatedactionrandompairmapsource_of_actionrewardpartialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_actionrewardpartialtrajectorykernel_extend_map_eq build the explicit generated-action random next-pair source from the frozen-prefix extension-map form of the `partialtraj` law. this narrows the source construction surface to the deterministic old-prefix extension map, then reuses the compiled extension-to-full-trace adapter. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_definitionalActualRewardMapSource","label":"generatedActionActualRewardMapSource_of_definitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_definitionalActualRewardMapSource","description":"Convert the definitional generated-action reward-coordinate source into the existing explicit-action actual reward-coordinate source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e407bfeb964b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3458,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6442"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_definitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionDefinitionalActualRewardMapSource mu rewardKernel policy context state defa…","missing":[],"search":"generatedactionactualrewardmapsource_of_definitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_definitionalactualrewardmapsource convert the definitional generated-action reward-coordinate source into the existing explicit-action actual reward-coordinate source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryPairLawSource","label":"generatedActionActualRewardMapSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryPairLawSource","description":"Project a generated-history `partialTraj` pair-law source to the explicit generated-action actual reward-coordinate source. This is the explicit-source companion to `generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryPairLawSource`. It keeps the action trace at `generatedActionFromRewardHistory` and derives its measurability from the packaged state measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d395ba06c9b5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3459,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6479"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionPartialTrajectoryPairLawSource mu rewardKernel policy context state defaultAct…","missing":[],"search":"generatedactionactualrewardmapsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_partialtrajectorypairlawsource project a generated-history `partialtraj` pair-law source to the explicit generated-action actual reward-coordinate source. this is the explicit-source companion to `generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorypairlawsource`. it keeps the action trace at `generatedactionfromrewardhistory` and derives its measurability from the packaged state measurability. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryKernel_map_eq","label":"generatedActionActualRewardMapSource_of_partialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryKernel_map_eq","description":"Build the explicit generated-action actual reward-coordinate source from a full finite-pair `partialTraj` law. This is the non-definitional counterpart of `generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq`: callers keep their explicit `action` trace and supply the shifted generated trace equality once, then the compiled full-trace projection supplies the reward-coordinate law stored…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4d4744c702a5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3460,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6535"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_partialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fu…","missing":[],"search":"generatedactionactualrewardmapsource_of_partialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_partialtrajectorykernel_map_eq build the explicit generated-action actual reward-coordinate source from a full finite-pair `partialtraj` law. this is the non-definitional counterpart of `generatedactiondefinitionalactualrewardmapsource_of_partialtrajectorykernel_map_eq`: callers keep their explicit `action` trace and supply the shifted generated trace equality once, then the compiled full-trace projection supplies the reward-coordinate law stored in `generatedactionactualrewardmapsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","label":"generatedActionActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","description":"Build the explicit generated-action actual reward-coordinate source from the frozen-prefix extension-map form of the `partialTraj` law. The input law is narrower than the full finite-pair trace law: it only pushes forward by appending the random successor pair to the already frozen prefix. The existing extension-map reward-coordinate adapter discharges that projection before packaging the source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c7b6c41e9dd5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3461,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6627"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measura…","missing":[],"search":"generatedactionactualrewardmapsource_of_partialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_partialtrajectorykernel_extend_map_eq build the explicit generated-action actual reward-coordinate source from the frozen-prefix extension-map form of the `partialtraj` law. the input law is narrower than the full finite-pair trace law: it only pushes forward by appending the random successor pair to the already frozen prefix. the existing extension-map reward-coordinate adapter discharges that projection before packaging the source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the explicit generated-action actual reward-coordinate source directly from the canonical history-step next-pair law. This packages the history-step surface without first asking callers to project it manually through `Prod.snd` or to build a source record by hand.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6edd7202f0e8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3462,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6719"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Na…","missing":[],"search":"generatedactionactualrewardmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the explicit generated-action actual reward-coordinate source directly from the canonical history-step next-pair law. this packages the history-step surface without first asking callers to project it manually through `prod.snd` or to build a source record by hand. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairMapSource","label":"generatedActionActualRewardMapSource_of_randomPairMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairMapSource","description":"Convert a generated random next-pair source into the weaker generated actual-action reward-coordinate source. The conversion freezes the generated action coordinate under the conditional kernel, marginalizes the resulting actual-action pair-product law through `Prod.snd`, and keeps the original shifted generated-action equality. It still assumes the random next-pair source law; it only weakens the source surface exp…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6a0947820f56","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3463,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6808"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat,…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairmapsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairmapsource convert a generated random next-pair source into the weaker generated actual-action reward-coordinate source. the conversion freezes the generated action coordinate under the conditional kernel, marginalizes the resulting actual-action pair-product law through `prod.snd`, and keeps the original shifted generated-action equality. it still assumes the random next-pair source law; it only weakens the source surface exposed to later consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_generatedActionActualRewardMapSource","label":"generatedActionRandomPairMapSource_of_generatedActionActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_generatedActionActualRewardMapSource","description":"Upgrade a generated-action actual reward-coordinate source to the random next-pair map source. The source already fixes the action trace to the shifted generated policy and identifies the conditional reward-coordinate law at the actual successor action. The split-product condExpKernel law then recovers the full random `(action, reward)` next-pair pushforward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f37828d6c390","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3464,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:6937"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_generatedActionActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : f…","missing":[],"search":"generatedactionrandompairmapsource_of_generatedactionactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_generatedactionactualrewardmapsource upgrade a generated-action actual reward-coordinate source to the random next-pair map source. the source already fixes the action trace to the shifted generated policy and identifies the conditional reward-coordinate law at the actual successor action. the split-product condexpkernel law then recovers the full random `(action, reward)` next-pair pushforward. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_generatedActionDefinitionalActualRewardMapSource","label":"generatedActionRandomPairMapSource_of_generatedActionDefinitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_generatedActionDefinitionalActualRewardMapSource","description":"Definitional generated-action actual reward-coordinate sources also produce the explicit generated-action random next-pair map source. This is the direct explicit-action counterpart of the definitional random-pair source conversion: it first exposes the definitional actual source as an explicit actual reward-coordinate source, then applies the split-product source upgrade above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2c1e469e47eb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3465,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7049"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionDefinitionalActualRewardMapSource mu rewardKernel policy conte…","missing":[],"search":"generatedactionrandompairmapsource_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_generatedactiondefinitionalactualrewardmapsource definitional generated-action actual reward-coordinate sources also produce the explicit generated-action random next-pair map source. this is the direct explicit-action counterpart of the definitional random-pair source conversion: it first exposes the definitional actual source as an explicit actual reward-coordinate source, then applies the split-product source upgrade above. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalMapSource","label":"generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalMapSource","description":"Definitional generated-action version of `generatedActionActualRewardMapSource_of_randomPairMapSource`. This converts a definitional random next-pair source into the definitional actual-action reward-coordinate source, deriving the timewise action measurability from the source's state-measurability field.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0c8429cb0eba","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3466,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7109"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionRandomPairDefinitionalMapSource mu rewardKernel policy context st…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalmapsource definitional generated-action version of `generatedactionactualrewardmapsource_of_randompairmapsource`. this converts a definitional random next-pair source into the definitional actual-action reward-coordinate source, deriving the timewise action measurability from the source's state-measurability field. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_policy_of_generatedActionRandomPairDefinitionalMapSource","label":"reward_condExpKernel_map_eq_selected_policy_of_generatedActionRandomPairDefinitionalMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_policy_of_generatedActionRandomPairDefinitionalMapSource","description":"Project a definitional generated random next-pair source to the policy-selected reward-coordinate law. The definitional source first weakens to the actual generated-action reward-coordinate source; unfolding `generatedActionFromRewardHistory` rewrites that generated successor action to the policy-selected action at the visible finite reward history.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-81034cb382d9","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3467,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7179"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem reward_condExpKernel_map_eq_selected_policy_of_generatedActionRandomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionRandomPairDefinitionalMapSource mu rewardKernel pol…","missing":[],"search":"reward_condexpkernel_map_eq_selected_policy_of_generatedactionrandompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.reward_condexpkernel_map_eq_selected_policy_of_generatedactionrandompairdefinitionalmapsource project a definitional generated random next-pair source to the policy-selected reward-coordinate law. the definitional source first weakens to the actual generated-action reward-coordinate source; unfolding `generatedactionfromrewardhistory` rewrites that generated successor action to the policy-selected action at the visible finite reward history. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalMapSource","label":"generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalMapSource","description":"Convert a definitional generated random next-pair source into the generated finite-pair `partialTraj` source. The definitional random-pair source already contains a stronger next-pair law. Projecting it to the policy-selected reward-coordinate law and using the generated-trace action-freezing constructor above builds the `GeneratedActionPartialTrajectoryPairLawSource`. This is only a source-surface conversion; it do…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8e229edd193f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3468,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7255"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActionRandomPairDefini…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_randompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_randompairdefinitionalmapsource convert a definitional generated random next-pair source into the generated finite-pair `partialtraj` source. the definitional random-pair source already contains a stronger next-pair law. projecting it to the policy-selected reward-coordinate law and using the generated-trace action-freezing constructor above builds the `generatedactionpartialtrajectorypairlawsource`. this is only a source-surface conversion; it does not prove the random-pair source law itself. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairCenteredSource","label":"GeneratedActionRandomPairCenteredSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairCenteredSource","description":"Generated-policy random next-pair source plus centered-reward regularity. This packages the source contract together with the measurable context/state extractors, the centered reward-kernel law, and ambient integrability of the generated centered reward at every successor step. It is still a contract: the random next-pair law and integrability fields are supplied, not derived from a global trajectory construction.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-fc352024046f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3469,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7308"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward…","missing":[],"search":"generatedactionrandompaircenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompaircenteredsource generated-policy random next-pair source plus centered-reward regularity. this packages the source contract together with the measurable context/state extractors, the centered reward-kernel law, and ambient integrability of the generated centered reward at every successor step. it is still a contract: the random next-pair law and integrability fields are supplied, not derived from a global trajectory construction. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalCenteredSource","label":"GeneratedActionRandomPairDefinitionalCenteredSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalCenteredSource","description":"Definitional generated-policy variant of the centered random-pair source. The action trace is definitionally the shifted policy-generated trace over finite reward histories. This keeps the same centered-kernel and integrability contract as `GeneratedActionRandomPairCenteredSource`, while removing explicit `action` and `haction` inputs from callers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7173fb83bdd5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3470,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7356"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) where","missing":[],"search":"generatedactionrandompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalcenteredsource definitional generated-policy variant of the centered random-pair source. the action trace is definitionally the shifted policy-generated trace over finite reward histories. this keeps the same centered-kernel and integrability contract as `generatedactionrandompaircenteredsource`, while removing explicit `action` and `haction` inputs from callers. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalCenteredSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalCenteredSource","description":"A definitional centered source exposes ambient integrability of the generated centered successor reward directly.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-995dbcb1627c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3471,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7396"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (sou…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairdefinitionalcenteredsource a definitional centered source exposes ambient integrability of the generated centered successor reward directly. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_definitionalCenteredSource","label":"generatedActionRandomPairCenteredSource_of_definitionalCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_definitionalCenteredSource","description":"Turn a definitional centered source into the explicit centered source whose action trace is `generatedActionFromRewardHistory`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-071d2306d537","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3472,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7433"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairCenteredSource_of_definitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActio…","missing":[],"search":"generatedactionrandompaircenteredsource_of_definitionalcenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompaircenteredsource_of_definitionalcenteredsource turn a definitional centered source into the explicit centered source whose action trace is `generatedactionfromrewardhistory`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairDefinitionalCenteredSource","label":"generatedActionRandomPairMapSource_of_randomPairDefinitionalCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairDefinitionalCenteredSource","description":"Convert a definitional centered generated random-pair source into the explicit generated random-pair map source whose action trace is `generatedActionFromRewardHistory`. The centered source already packages the definitional random-pair map source; its centered-kernel and integrability fields are not needed for this weaker map-source interface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-53d1cb04d537","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3473,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7485"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : Generated…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairdefinitionalcenteredsource convert a definitional centered generated random-pair source into the explicit generated random-pair map source whose action trace is `generatedactionfromrewardhistory`. the centered source already packages the definitional random-pair map source; its centered-kernel and integrability fields are not needed for this weaker map-source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalCenteredSource","label":"generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalCenteredSource","description":"Convert a definitional centered generated random-pair source into the weaker definitional actual-action reward-coordinate source. The centered source already packages the definitional random-pair map source; its centered-kernel and integrability fields are not needed for this weaker reward-map interface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-33a550bd87d8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3474,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7531"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (sour…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalcenteredsource convert a definitional centered generated random-pair source into the weaker definitional actual-action reward-coordinate source. the centered source already packages the definitional random-pair map source; its centered-kernel and integrability fields are not needed for this weaker reward-map interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairDefinitionalCenteredSource","label":"generatedActionActualRewardMapSource_of_randomPairDefinitionalCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairDefinitionalCenteredSource","description":"Convert a definitional centered generated random-pair source into the explicit generated actual-action reward-coordinate source whose action trace is `generatedActionFromRewardHistory`. This is the explicit-action counterpart of `generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalCenteredSource`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f866ed1b9085","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3475,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7572"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : Generat…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairdefinitionalcenteredsource convert a definitional centered generated random-pair source into the explicit generated actual-action reward-coordinate source whose action trace is `generatedactionfromrewardhistory`. this is the explicit-action counterpart of `generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalcenteredsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairCenteredSource","label":"generatedActionRandomPairMapSource_of_randomPairCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairCenteredSource","description":"Project a centered generated random next-pair source directly to its packaged random-pair map source. The centered source already contains the random-pair source used by its history-step and actual-reward projections. This wrapper records that weaker interface under a stable name for downstream law consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-72c62a1997c4","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3476,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7630"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action om…","missing":[],"search":"generatedactionrandompairmapsource_of_randompaircenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompaircenteredsource project a centered generated random next-pair source directly to its packaged random-pair map source. the centered source already contains the random-pair source used by its history-step and actual-reward projections. this wrapper records that weaker interface under a stable name for downstream law consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairCenteredSource","label":"generatedActionActualRewardMapSource_of_randomPairCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairCenteredSource","description":"Convert a centered generated random next-pair source into the weaker generated actual-action reward-coordinate source. The centered source already contains the random-pair source and state measurability needed by `generatedActionActualRewardMapSource_of_randomPairMapSource`; this wrapper exposes that weaker interface for downstream consumers that do not need the centered-law or integrability fields.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d0e368d0c893","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3477,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7667"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompaircenteredsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompaircenteredsource convert a centered generated random next-pair source into the weaker generated actual-action reward-coordinate source. the centered source already contains the random-pair source and state measurability needed by `generatedactionactualrewardmapsource_of_randompairmapsource`; this wrapper exposes that weaker interface for downstream consumers that do not need the centered-law or integrability fields. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairBoundedCenteredSource","label":"GeneratedActionRandomPairBoundedCenteredSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairBoundedCenteredSource","description":"Generated-policy random next-pair source plus bounded centered-reward regularity. This is a bounded-input variant of `GeneratedActionRandomPairCenteredSource`: ambient integrability is derived from a.e. measurability and a per-step a.e. interval bound via `MeasureTheory.Integrable.of_mem_Icc`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5149f1d09809","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3478,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7714"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (h…","missing":[],"search":"generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompairboundedcenteredsource generated-policy random next-pair source plus bounded centered-reward regularity. this is a bounded-input variant of `generatedactionrandompaircenteredsource`: ambient integrability is derived from a.e. measurability and a per-step a.e. interval bound via `measuretheory.integrable.of_mem_icc`. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawMeanBoundedSource","label":"GeneratedActionRandomPairRawMeanBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawMeanBoundedSource","description":"Generated-policy random next-pair source plus raw-reward and selected-mean boundedness. This is a more primitive bounded-input variant: centered-reward a.e. measurability and bounds are derived from raw reward a.e. measurability/bounds and selected mean a.e. measurability/bounds.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-09b9a2c14e46","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3479,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7777"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hr…","missing":[],"search":"generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawmeanboundedsource generated-policy random next-pair source plus raw-reward and selected-mean boundedness. this is a more primitive bounded-input variant: centered-reward a.e. measurability and bounds are derived from raw reward a.e. measurability/bounds and selected mean a.e. measurability/bounds. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeanBoundedSource","label":"GeneratedActionRandomPairRawBoundMeanBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeanBoundedSource","description":"Generated-policy random next-pair source plus raw-reward bounds and selected-mean measurability/bounds. This variant derives raw reward a.e. measurability from the already available timewise measurable reward trace `hreward`; selected-mean measurability remains explicit.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0d82f2ca3767","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3480,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7849"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)…","missing":[],"search":"generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawboundmeanboundedsource generated-policy random next-pair source plus raw-reward bounds and selected-mean measurability/bounds. this variant derives raw reward a.e. measurability from the already available timewise measurable reward trace `hreward`; selected-mean measurability remains explicit. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"GeneratedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Generated-policy random next-pair source plus raw-reward bounds and a measurable selected mean. This variant derives selected-mean a.e. measurability by composing a measurable mean surface with the measurable finite reward history, context, state, and policy action. It still keeps selected-mean bounds as explicit source data.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-545380fba510","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3481,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7917"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => actio…","missing":[],"search":"generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawboundmeasurablemeanboundedsource generated-policy random next-pair source plus raw-reward bounds and a measurable selected mean. this variant derives selected-mean a.e. measurability by composing a measurable mean surface with the measurable finite reward history, context, state, and policy action. it still keeps selected-mean bounds as explicit source data. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"GeneratedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Generated-policy random next-pair source plus raw-reward bounds and a measurable mean with deterministic range bounds. This variant derives the selected-mean a.e. bound from a pointwise range contract for `mean`; it is still a source contract and does not derive the raw reward bounds or random next-pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8762c6a8375a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3482,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:7976"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawboundmeasurablemeanrangeboundedsource generated-policy random next-pair source plus raw-reward bounds and a measurable mean with deterministic range bounds. this variant derives the selected-mean a.e. bound from a pointwise range contract for `mean`; it is still a source contract and does not derive the raw reward bounds or random next-pair law. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"GeneratedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Generated-policy random next-pair source plus pointwise raw-reward range bounds and a measurable mean with deterministic range bounds. This variant derives both raw-reward and selected-mean a.e. bounds from pointwise range contracts; it is still a source contract and does not derive the random next-pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-cee10624e940","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3483,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8026"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawrangemeasurablemeanrangeboundedsource generated-policy random next-pair source plus pointwise raw-reward range bounds and a measurable mean with deterministic range bounds. this variant derives both raw-reward and selected-mean a.e. bounds from pointwise range contracts; it is still a source contract and does not derive the random next-pair law. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Definitional generated-policy variant of the most primitive raw-range source. The action trace is definitionally the shifted policy-generated trace over finite reward histories, so callers do not provide a separate action trace or timewise action measurability proof. The random next-pair law is still a source contract, inherited through `GeneratedActionRandomPairDefinitionalMapSource`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-61646b82c06e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3484,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8073"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource definitional generated-policy variant of the most primitive raw-range source. the action trace is definitionally the shifted policy-generated trace over finite reward histories, so callers do not provide a separate action trace or timewise action measurability proof. the random next-pair law is still a source contract, inherited through `generatedactionrandompairdefinitionalmapsource`. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","label":"GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","description":"Practical definitional raw-range source with a deterministic variance-proxy ceiling. The base source carries the generated-action random next-pair law, raw reward range, selected-mean range, and centered kernel law. This wrapper adds the uniform model-side upper bound needed to choose a deterministic `HasCondSubgaussianMGF` variance proxy.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-203b00904762","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3485,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8117"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource practical definitional raw-range source with a deterministic variance-proxy ceiling. the base source carries the generated-action random next-pair law, raw reward range, selected-mean range, and centered kernel law. this wrapper adds the uniform model-side upper bound needed to choose a deterministic `hascondsubgaussianmgf` variance proxy. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","label":"GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","description":"Practical definitional raw-range source with time-indexed selected-history variance-proxy ceilings. This is weaker than a global context/action ceiling: the bound only needs to hold on histories reachable by the reward-history context/state surface at the time being consumed.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-61f34534d67c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3486,"meta":[["Kind","structure"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8152"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource practical definitional raw-range source with time-indexed selected-history variance-proxy ceilings. this is weaker than a global context/action ceiling: the bound only needs to hold on histories reachable by the reward-history context/state surface at the time being consumed. structure compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniformVarianceBoundedSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniformVarianceBoundedSource","description":"Project a definitional raw-range/measurable-mean-range uniform-variance source to its packaged base raw-range/measurable-mean-range bounded source. The uniform source adds a global variance-proxy ceiling for MGF consumers. This wrapper records the weaker base-source interface for downstream consumers that only need the generated random-pair law and raw/mean range regularity.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bb0b6a076264","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3487,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8188"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => r…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_uniformvarianceboundedsource project a definitional raw-range/measurable-mean-range uniform-variance source to its packaged base raw-range/measurable-mean-range bounded source. the uniform source adds a global variance-proxy ceiling for mgf consumers. this wrapper records the weaker base-source interface for downstream consumers that only need the generated random-pair law and raw/mean range regularity. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_historyVarianceBoundedSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_historyVarianceBoundedSource","description":"Project a definitional raw-range/measurable-mean-range history-variance source to its packaged base raw-range/measurable-mean-range bounded source. The history-variance source adds selected-history variance-proxy ceilings for MGF consumers. This wrapper records the weaker base-source interface for downstream consumers that only need the generated random-pair law and raw/mean range regularity.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ac3d80a776d2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3488,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8225"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => r…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_historyvarianceboundedsource project a definitional raw-range/measurable-mean-range history-variance source to its packaged base raw-range/measurable-mean-range bounded source. the history-variance source adds selected-history variance-proxy ceilings for mgf consumers. this wrapper records the weaker base-source interface for downstream consumers that only need the generated random-pair law and raw/mean range regularity. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_uniformVarianceBoundedSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_uniformVarianceBoundedSource","description":"A uniform variance-proxy ceiling is a constant time-indexed selected-history variance ceiling. This adapter lets downstream consumers that are stated against the weaker history-specific source interface reuse a stronger global context/action source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f8075c0b2c2f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3489,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8260"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun ome…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_uniformvarianceboundedsource a uniform variance-proxy ceiling is a constant time-indexed selected-history variance ceiling. this adapter lets downstream consumers that are stated against the weaker history-specific source interface reuse a stronger global context/action source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","description":"Build the practical definitional generated-policy raw-range source from a full finite-pair-trace `partialTraj` law. This packages the new definitional map-source adapter together with the regularity fields needed by the top raw-range/measurable-mean-range source layer. The trajectory law is still an explicit hypothesis.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3c82cd275211","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3490,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8302"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omeg…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_map_eq build the practical definitional generated-policy raw-range source from a full finite-pair-trace `partialtraj` law. this packages the new definitional map-source adapter together with the regularity fields needed by the top raw-range/measurable-mean-range source layer. the trajectory law is still an explicit hypothesis. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_partialTrajectoryPairLawSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_partialTrajectoryPairLawSource","description":"Consume a generated-history `partialTraj` pair-law source into the practical definitional raw-range/measurable-mean-range source interface. The remaining trajectory-law input is still explicit; this wrapper only reuses the packaged source fields and adds the raw/mean regularity contracts needed by the top source layer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-13c982b9ce43","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3491,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8416"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_partialtrajectorypairlawsource consume a generated-history `partialtraj` pair-law source into the practical definitional raw-range/measurable-mean-range source interface. the remaining trajectory-law input is still explicit; this wrapper only reuses the packaged source fields and adds the raw/mean regularity contracts needed by the top source layer. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_reward_map_eq_selected_policy","description":"Build the practical raw-range/measurable-mean-range source directly from a finite-pair-history comap selected-reward law. This is the comap-selected-reward entry point for the base source-conversion leaf: the selected-reward law first constructs the generated full finite-pair `partialTraj` source, then the existing source-contract constructor packages the raw/mean range regularity.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f4aa49a7bd9f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3492,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8484"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_comap_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_comap_reward_map_eq_selected_policy build the practical raw-range/measurable-mean-range source directly from a finite-pair-history comap selected-reward law. this is the comap-selected-reward entry point for the base source-conversion leaf: the selected-reward law first constructs the generated full finite-pair `partialtraj` source, then the existing source-contract constructor packages the raw/mean range regularity. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_trim_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_trim_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_trim_reward_map_eq_selected_policy","description":"Build the practical raw-range/measurable-mean-range source directly from a finite-pair-history comap selected-reward law whose a.e. filter is also stated at the comap-trim conditioning surface. This is the direct comap-trim companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_reward_map_eq_selected_policy`. It changes only the selected-reward law entry surface: the pro…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-be1106b9a467","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3493,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8606"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_trim_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAc…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_comap_trim_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_comap_trim_reward_map_eq_selected_policy build the practical raw-range/measurable-mean-range source directly from a finite-pair-history comap selected-reward law whose a.e. filter is also stated at the comap-trim conditioning surface. this is the direct comap-trim companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_comap_reward_map_eq_selected_policy`. it changes only the selected-reward law entry surface: the proof first builds the generated full finite-pair `partialtraj` source through the comap-trim law constructor, then reuses the existing raw-range source-contract package. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","description":"Build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from a full finite-pair-trace `partialTraj` law. This is the uniform-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-409b7b959a7d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3494,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8779"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measu…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardpartialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardpartialtrajectorykernel_map_eq build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from a full finite-pair-trace `partialtraj` law. this is the uniform-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_map_eq`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_partialTrajectoryPairLawSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_partialTrajectoryPairLawSource","description":"Consume a generated-history `partialTraj` pair-law source into the packaged uniform-variance practical source interface. This is the source-contract version of `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq`: the trajectory law is still supplied by `source`, while raw/mean range and uniform variance regularity remain explici…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1d30c630721a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3495,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8904"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun o…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_partialtrajectorypairlawsource consume a generated-history `partialtraj` pair-law source into the packaged uniform-variance practical source interface. this is the source-contract version of `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardpartialtrajectorykernel_map_eq`: the trajectory law is still supplied by `source`, while raw/mean range and uniform variance regularity remain explicit. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","description":"Build the packaged uniform-variance practical source directly from a finite-pair-history comap selected-reward law. This is the comap-selected-reward entry point for the source-conversion leaf: the selected-reward law first constructs the generated full finite-pair `partialTraj` source, then the existing source-contract constructor packages the raw/mean range and global variance regularity.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-fbe5e2252558","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3496,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:8978"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n))…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_comap_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_comap_reward_map_eq_selected_policy build the packaged uniform-variance practical source directly from a finite-pair-history comap selected-reward law. this is the comap-selected-reward entry point for the source-conversion leaf: the selected-reward law first constructs the generated full finite-pair `partialtraj` source, then the existing source-contract constructor packages the raw/mean range and global variance regularity. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","description":"Build the packaged uniform-variance practical source directly from a finite-pair-history comap selected-reward law whose a.e. filter is also stated at the comap-trim conditioning surface. This is the direct comap-trim companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_reward_map_eq_selected_policy`. It changes only the selected-reward law entry surface…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-258fa099b3a6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3497,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9106"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_comap_trim_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_comap_trim_reward_map_eq_selected_policy build the packaged uniform-variance practical source directly from a finite-pair-history comap selected-reward law whose a.e. filter is also stated at the comap-trim conditioning surface. this is the direct comap-trim companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_comap_reward_map_eq_selected_policy`. it changes only the selected-reward law entry surface: the proof first builds the generated full finite-pair `partialtraj` source through the comap-trim law constructor, then reuses the existing source-contract uniform-variance package. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","description":"Build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from a full finite-pair-trace `partialTraj` law. This is the history-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c6f49c7ad440","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3498,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9285"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measu…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardpartialtrajectorykernel_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardpartialtrajectorykernel_map_eq build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from a full finite-pair-trace `partialtraj` law. this is the history-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_map_eq`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_partialTrajectoryPairLawSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_partialTrajectoryPairLawSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_partialTrajectoryPairLawSource","description":"Consume a generated-history `partialTraj` pair-law source into the packaged selected-history-variance practical source interface. This is the source-contract version of `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq`: the trajectory law is still supplied by `source`, while raw/mean range and selected-history variance regular…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2526d4fd7397","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3499,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9411"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_partialTrajectoryPairLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun o…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_partialtrajectorypairlawsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_partialtrajectorypairlawsource consume a generated-history `partialtraj` pair-law source into the packaged selected-history-variance practical source interface. this is the source-contract version of `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardpartialtrajectorykernel_map_eq`: the trajectory law is still supplied by `source`, while raw/mean range and selected-history variance regularity remain explicit. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","description":"Build the packaged selected-history-variance practical source directly from a finite-pair-history comap selected-reward law. This is the comap-selected-reward entry point for the selected-history source conversion leaf: the selected-reward law first constructs the generated full finite-pair `partialTraj` source, then the existing source-contract constructor packages the raw/mean range and selected-history variance r…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2624d86181dd","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3500,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9487"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n))…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_comap_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_comap_reward_map_eq_selected_policy build the packaged selected-history-variance practical source directly from a finite-pair-history comap selected-reward law. this is the comap-selected-reward entry point for the selected-history source conversion leaf: the selected-reward law first constructs the generated full finite-pair `partialtraj` source, then the existing source-contract constructor packages the raw/mean range and selected-history variance regularity. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","description":"Build the packaged selected-history-variance practical source directly from a finite-pair-history comap selected-reward law whose a.e. filter is also stated at the comap-trim conditioning surface. This is the direct comap-trim companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_reward_map_eq_selected_policy`. It changes only the selected-reward law entr…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f518ce3b9898","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3501,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9616"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_comap_trim_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_comap_trim_reward_map_eq_selected_policy build the packaged selected-history-variance practical source directly from a finite-pair-history comap selected-reward law whose a.e. filter is also stated at the comap-trim conditioning surface. this is the direct comap-trim companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_comap_reward_map_eq_selected_policy`. it changes only the selected-reward law entry surface: the proof first builds the generated full finite-pair `partialtraj` source through the comap-trim law constructor, then reuses the existing source-contract history-variance package. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","description":"Build the practical definitional generated-policy raw-range source from the narrower frozen-prefix extension-map `partialTraj` law. This is the current closest local surface to the final adaptive reward source: all regularity is packaged, while the remaining semantic gap is isolated to the explicit extension-map trajectory-law hypothesis.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f97f1b439abe","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3502,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9796"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (f…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq build the practical definitional generated-policy raw-range source from the narrower frozen-prefix extension-map `partialtraj` law. this is the current closest local surface to the final adaptive reward source: all regularity is packaged, while the remaining semantic gap is isolated to the explicit extension-map trajectory-law hypothesis. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","description":"Build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the narrower frozen-prefix extension-map `partialTraj` law. This is the uniform-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-94c5c1a2624c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3503,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:9914"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the narrower frozen-prefix extension-map `partialtraj` law. this is the uniform-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","description":"Build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the narrower frozen-prefix extension-map `partialTraj` law. This is the history-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2853b146b943","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3504,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10042"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the narrower frozen-prefix extension-map `partialtraj` law. this is the history-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardpartialtrajectorykernel_extend_map_eq`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the practical definitional generated-policy raw-range source from the canonical history-step next-pair law. This packages the same law shape consumed by `centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeBounded` as a reusable source, so downstream routes can keep the practical raw/mean regularity together with the definitional random-…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a4565ef6ba81","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3505,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10172"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the practical definitional generated-policy raw-range source from the canonical history-step next-pair law. this packages the same law shape consumed by `centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangebounded` as a reusable source, so downstream routes can keep the practical raw/mean regularity together with the definitional random-pair map source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the canonical history-step next-pair law. This is the uniform-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d6397d29aab0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3506,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10284"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat,…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the canonical history-step next-pair law. this is the uniform-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","description":"Build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the canonical history-step next-pair law. This is the history-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-55f19c1bfcf9","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3507,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10406"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat,…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the canonical history-step next-pair law. this is the history-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_actionrewardhistorystepkernelfamily_pair_map_eq`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.rawReward_succ_aemeasurable_of_measurable_reward","label":"rawReward_succ_aemeasurable_of_measurable_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.rawReward_succ_aemeasurable_of_measurable_reward","description":"Timewise measurable reward traces give raw reward a.e. measurability after casting `Rat` rewards to `Real`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4094f9a00829","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3508,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10525"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem rawReward_succ_aemeasurable_of_measurable_reward {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (i : Nat) : AEMeasurable (fun omega : Omega => (((reward omega (i + 1) : Rat) : Real))) mu","missing":[],"search":"rawreward_succ_aemeasurable_of_measurable_reward banditrlproof.conditionalexpectationreward.rawreward_succ_aemeasurable_of_measurable_reward timewise measurable reward traces give raw reward a.e. measurability after casting `rat` rewards to `real`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.selectedMean_succ_aemeasurable_of_measurable_mean","label":"selectedMean_succ_aemeasurable_of_measurable_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.selectedMean_succ_aemeasurable_of_measurable_mean","description":"Measurable finite reward history, context, state, policy action, and mean surface give selected-mean a.e. measurability after casting `Rat` to `Real`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1287ed76569b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3509,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10542"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem selectedMean_succ_aemeasurable_of_measurable_mean {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : MeasureTheory.Measure Omega) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (i : Nat) : AEMeasurable (fun omega : Omega => (((mean (context i (History.finiteRewardHistoryOfTrace (reward omega) i)) ((po…","missing":[],"search":"selectedmean_succ_aemeasurable_of_measurable_mean banditrlproof.conditionalexpectationreward.selectedmean_succ_aemeasurable_of_measurable_mean measurable finite reward history, context, state, policy action, and mean surface give selected-mean a.e. measurability after casting `rat` to `real`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.selectedMean_succ_bound_of_mean_range_bound","label":"selectedMean_succ_bound_of_mean_range_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.selectedMean_succ_bound_of_mean_range_bound","description":"A pointwise range bound on the mean surface gives the generated selected-mean a.e. bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6e80841c298b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3510,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10609"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem selectedMean_succ_bound_of_mean_range_bound {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : MeasureTheory.Measure Omega) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (reward : Omega -> RewardTrace Rat) (meanLo meanHi : Nat -> Real) (hmean_bound : forall i : Nat, forall context : Context, forall action : Action, Set.Icc (meanLo i) (meanHi i) (((mean context action : Rat) : Real))) (i : Nat) : Filter.Eventually (fun omega : Omega => Set.Icc (meanLo i) (meanHi i) (((mean (context i (History.finiteRewardHistoryOfTrace (reward omega) i)) ((policy i).action (state i (History.finiteRewa…","missing":[],"search":"selectedmean_succ_bound_of_mean_range_bound banditrlproof.conditionalexpectationreward.selectedmean_succ_bound_of_mean_range_bound a pointwise range bound on the mean surface gives the generated selected-mean a.e. bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.rawReward_succ_bound_of_reward_range_bound","label":"rawReward_succ_bound_of_reward_range_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.rawReward_succ_bound_of_reward_range_bound","description":"A pointwise range bound on the raw reward trace gives the raw successor reward a.e. bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-74bac1ebffed","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3511,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10649"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem rawReward_succ_bound_of_reward_range_bound {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) (reward : Omega -> RewardTrace Rat) (rewardLo rewardHi : Nat -> Real) (hreward_bound : forall i : Nat, forall omega : Omega, Set.Icc (rewardLo i) (rewardHi i) (((reward omega (i + 1) : Rat) : Real))) (i : Nat) : Filter.Eventually (fun omega : Omega => Set.Icc (rewardLo i) (rewardHi i) (((reward omega (i + 1) : Rat) : Real))) (ae mu)","missing":[],"search":"rawreward_succ_bound_of_reward_range_bound banditrlproof.conditionalexpectationreward.rawreward_succ_bound_of_reward_range_bound a pointwise range bound on the raw reward trace gives the raw successor reward a.e. bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_integrable_of_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_rawRangeMeasurableMeanRangeBounded","description":"Raw reward range bounds plus a measurable selected-mean surface with range bounds give ambient integrability for the centered successor reward. This helper intentionally does not require any conditional reward-law source: it is pure regularity, so weaker law sources can reuse the existing integrability-based conditional mean-zero route without assuming a random-pair map law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-086ac5f40305","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3512,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10677"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo meanHi : Nat -> Real) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (hraw : forall i : Nat,…","missing":[],"search":"centeredreward_succ_integrable_of_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_rawrangemeasurablemeanrangebounded raw reward range bounds plus a measurable selected-mean surface with range bounds give ambient integrability for the centered successor reward. this helper intentionally does not require any conditional reward-law source: it is pure regularity, so weaker law sources can reuse the existing integrability-based conditional mean-zero route without assuming a random-pair map law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","description":"Consume a reward-coordinate map law plus prefix-coordinate measurability and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the coordinate-measurable map-law consumer at an arbitrary filtration `F`: callers can provide the `RewardKernel.historyStepKernelFamily` pushforward law directly, with the finite reward prefix visib…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a2b372a529ed","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3513,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10847"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (F : MeasureTheory.Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNR…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_of_coordinate_measurable_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_of_coordinate_measurable_rawrangemeasurablemeanrangebounded consume a reward-coordinate map law plus prefix-coordinate measurability and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the coordinate-measurable map-law consumer at an arbitrary filtration `f`: callers can provide the `rewardkernel.historystepkernelfamily` pushforward law directly, with the finite reward prefix visible at `f i`, without separately proving centered-reward integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","description":"Consume a generated-history reward-coordinate map law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the earliest reward-coordinate map consumer: callers can provide the `RewardKernel.historyStepKernelFamily` pushforward law directly, without separately proving centered-reward integrability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-28612771a59c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3514,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:10965"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProx…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_historyfiltrationsucc_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_historystepkernelfamily_condexpkernel_map_eq_historyfiltrationsucc_rawrangemeasurablemeanrangebounded consume a generated-history reward-coordinate map law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the earliest reward-coordinate map consumer: callers can provide the `rewardkernel.historystepkernelfamily` pushforward law directly, without separately proving centered-reward integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","description":"Consume a direct action/reward pair map law plus prefix-coordinate measurability and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the coordinate-measurable pair-law consumer at an arbitrary filtration `F`: callers can provide the `RewardKernel.actionRewardHistoryStepKernelFamily` pushforward law directly, without separa…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-51747075f202","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3515,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11087"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (F : MeasureTheory.Filtration Nat mOmega) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (pairContext : (n : Nat) -> ((j : Finset.Iic n) -> Prod Action Rat) -> C…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_of_coordinate_measurable_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_of_coordinate_measurable_rawrangemeasurablemeanrangebounded consume a direct action/reward pair map law plus prefix-coordinate measurability and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the coordinate-measurable pair-law consumer at an arbitrary filtration `f`: callers can provide the `rewardkernel.actionrewardhistorystepkernelfamily` pushforward law directly, without separately proving centered-reward integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","description":"Generated-history-filtration specialization of the direct action/reward pair-map route with raw reward and selected-mean range regularity. This removes the separate centered-reward integrability hypothesis from the generated-history pair-law consumer. The actual `condExpKernel` next-pair law is still an explicit structural assumption.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1bb4adb780f8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3516,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11232"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (pairContext : (n : Nat) -> ((j : Finset.Iic…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_rawrangemeasurablemeanrangebounded generated-history-filtration specialization of the direct action/reward pair-map route with raw reward and selected-mean range regularity. this removes the separate centered-reward integrability hypothesis from the generated-history pair-law consumer. the actual `condexpkernel` next-pair law is still an explicit structural assumption. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_rawRangeMeasurableMeanRangeBounded","description":"Concrete trace-pair specialization of the generated-history pair-map route with raw reward and selected-mean range regularity. This fixes the pair history to the actual finite prefix `fun j => (action omega j, reward omega j)` and still derives the centered reward integrability from bounded raw reward and selected mean evidence.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bc10913ccd00","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3517,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11359"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (hpairContext : forall n : Nat, Me…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected_rawrangemeasurablemeanrangebounded concrete trace-pair specialization of the generated-history pair-map route with raw reward and selected-mean range regularity. this fixes the pair history to the actual finite prefix `fun j => (action omega j, reward omega j)` and still derives the centered reward integrability from bounded raw reward and selected mean evidence. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable_rawRangeMeasurableMeanRangeBounded","description":"Projected trace-pair route with projection measurability and raw reward/selected-mean range regularity supplied locally. This combines `History.measurable_pairHistoryRewardProjection` with the bounded-regularity wrapper, so callers only provide reward-history context/state measurability plus the concrete next-pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-755896764edd","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3518,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11486"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected_of_context_state_measurable_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_projected_of_context_state_measurable_rawrangemeasurablemeanrangebounded projected trace-pair route with projection measurability and raw reward/selected-mean range regularity supplied locally. this combines `history.measurable_pairhistoryrewardprojection` with the bounded-regularity wrapper, so callers only provide reward-history context/state measurability plus the concrete next-pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Consume a full finite-pair-trace `partialTraj` law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the finite-pair `partialTraj` consumer: callers no longer need to pass centered-reward integrability separately when the usual bounded reward/mean evidence is available.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-81a5dc7deff1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3519,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11606"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Actio…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded consume a full finite-pair-trace `partialtraj` law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the finite-pair `partialtraj` consumer: callers no longer need to pass centered-reward integrability separately when the usual bounded reward/mean evidence is available. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Consume a frozen-prefix extension-map `partialTraj` law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the narrower extension-map `partialTraj` consumer. It keeps the trajectory-law gap at the extension-map surface while removing the separate centered-reward integrability argument.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2704a8f8feab","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3520,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11744"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded consume a frozen-prefix extension-map `partialtraj` law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the narrower extension-map `partialtraj` consumer. it keeps the trajectory-law gap at the extension-map surface while removing the separate centered-reward integrability argument. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Consume a direct canonical history-step next-pair law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the finite-pair next-pair consumer: callers can provide the `RewardKernel.actionRewardHistoryStepKernelFamily` pair law directly, without separately proving centered-reward integrability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4644805bc71a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3521,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:11885"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context ->…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded consume a direct canonical history-step next-pair law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the finite-pair next-pair consumer: callers can provide the `rewardkernel.actionrewardhistorystepkernelfamily` pair law directly, without separately proving centered-reward integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Generated-action surface for the canonical history-step next-pair raw-range consumer. The underlying direct pair-law consumer already has enough information to prove conditional mean-zero; this wrapper records the common generated-action calling convention used by the adjacent reward-map and `partialTraj` routes.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7761b4972e02","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3522,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12019"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (stat…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded generated-action surface for the canonical history-step next-pair raw-range consumer. the underlying direct pair-law consumer already has enough information to prove conditional mean-zero; this wrapper records the common generated-action calling convention used by the adjacent reward-map and `partialtraj` routes. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Consume the split next-pair law assumptions plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the mean-zero surface for the split-law builder: conditional a.e. action equality and the reward-coordinate selected-measure law first build the canonical history-step next-pair law, then the finite-pair raw-range consumer supplies integrability and conditional…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d426507cdee8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3523,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12138"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat)…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded consume the split next-pair law assumptions plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the mean-zero surface for the split-law builder: conditional a.e. action equality and the reward-coordinate selected-measure law first build the canonical history-step next-pair law, then the finite-pair raw-range consumer supplies integrability and conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_selected_policy_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_selected_policy_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_selected_policy_rawRangeMeasurableMeanRangeBounded","description":"Consume a generated-action trace plus a policy-selected reward-coordinate law and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. Compared with the actual-action reward-law wrapper, this surface matches the policy-selected reward side consumed by the split-law builder directly; the generated trace supplies the conditional action a.e. equality.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b9fa705d709f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3524,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12334"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_selected_policy_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varia…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_selected_policy_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_selected_policy_rawrangemeasurablemeanrangebounded consume a generated-action trace plus a policy-selected reward-coordinate law and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. compared with the actual-action reward-law wrapper, this surface matches the policy-selected reward side consumed by the split-law builder directly; the generated trace supplies the conditional action a.e. equality. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawMeanBoundedSource_of_rawBoundMeanBoundedSource","label":"generatedActionRandomPairRawMeanBoundedSource_of_rawBoundMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawMeanBoundedSource_of_rawBoundMeanBoundedSource","description":"Turn raw-reward bounds plus selected-mean measurability/bounds into the raw/mean bounded source by deriving raw reward a.e. measurability from `hreward`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4a5daee468e0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3525,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12491"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairRawMeanBoundedSource_of_rawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega…","missing":[],"search":"generatedactionrandompairrawmeanboundedsource_of_rawboundmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawmeanboundedsource_of_rawboundmeanboundedsource turn raw-reward bounds plus selected-mean measurability/bounds into the raw/mean bounded source by deriving raw reward a.e. measurability from `hreward`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeanBoundedSource_of_rawBoundMeasurableMeanBoundedSource","label":"generatedActionRandomPairRawBoundMeanBoundedSource_of_rawBoundMeasurableMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeanBoundedSource_of_rawBoundMeasurableMeanBoundedSource","description":"Turn raw-reward bounds plus a measurable selected-mean surface into the raw-bound/mean-bounded source by deriving selected-mean a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-76fdd9e76679","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3526,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12534"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairRawBoundMeanBoundedSource_of_rawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"generatedactionrandompairrawboundmeanboundedsource_of_rawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawboundmeanboundedsource_of_rawboundmeasurablemeanboundedsource turn raw-reward bounds plus a measurable selected-mean surface into the raw-bound/mean-bounded source by deriving selected-mean a.e. measurability. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeasurableMeanBoundedSource_of_rawBoundMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairRawBoundMeasurableMeanBoundedSource_of_rawBoundMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeasurableMeanBoundedSource_of_rawBoundMeasurableMeanRangeBoundedSource","description":"Turn raw-reward bounds plus a measurable mean surface and deterministic mean range bounds into the measurable-mean bounded source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-08bbdf5f1590","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3527,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12578"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairRawBoundMeasurableMeanBoundedSource_of_rawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat,…","missing":[],"search":"generatedactionrandompairrawboundmeasurablemeanboundedsource_of_rawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawboundmeasurablemeanboundedsource_of_rawboundmeasurablemeanrangeboundedsource turn raw-reward bounds plus a measurable mean surface and deterministic mean range bounds into the measurable-mean bounded source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource_of_rawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource_of_rawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource_of_rawRangeMeasurableMeanRangeBoundedSource","description":"Turn pointwise raw-reward range bounds plus a measurable mean surface and deterministic mean range bounds into the raw-bound/measurable-mean-range source.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8ddad968435e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3528,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12623"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource_of_rawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t :…","missing":[],"search":"generatedactionrandompairrawboundmeasurablemeanrangeboundedsource_of_rawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawboundmeasurablemeanrangeboundedsource_of_rawrangemeasurablemeanrangeboundedsource turn pointwise raw-reward range bounds plus a measurable mean surface and deterministic mean range bounds into the raw-bound/measurable-mean-range source. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Turn a definitional generated-action raw-range source into the existing explicit-action raw-range source by deriving the action trace and its measurability from the reward-history state.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e9fc4bcf413b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3529,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12667"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega…","missing":[],"search":"generatedactionrandompairrawrangemeasurablemeanrangeboundedsource_of_definitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairrawrangemeasurablemeanrangeboundedsource_of_definitionalrawrangemeasurablemeanrangeboundedsource turn a definitional generated-action raw-range source into the existing explicit-action raw-range source by deriving the action trace and its measurability from the reward-history state. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairMapSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Project a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source directly to the explicit generated random-pair map source whose action trace is `generatedActionFromRewardHistory`. This is the named map-source projection sitting below the stronger raw-range and actual-reward-map source conversions.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4f5f2637302f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3530,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12722"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (re…","missing":[],"search":"generatedactionrandompairmapsource_of_definitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_definitionalrawrangemeasurablemeanrangeboundedsource project a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source directly to the explicit generated random-pair map source whose action trace is `generatedactionfromrewardhistory`. this is the named map-source projection sitting below the stronger raw-range and actual-reward-map source conversions. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawMeanBoundedSource","label":"centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawMeanBoundedSource","description":"Raw reward and selected-mean a.e. measurability imply centered generated reward a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ac6ed4599f37","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3531,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12766"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun ome…","missing":[],"search":"centeredreward_succ_aemeasurable_of_generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_aemeasurable_of_generatedactionrandompairrawmeanboundedsource raw reward and selected-mean a.e. measurability imply centered generated reward a.e. measurability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawMeanBoundedSource","label":"centeredReward_succ_bound_of_generatedActionRandomPairRawMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawMeanBoundedSource","description":"Raw reward bounds and selected-mean bounds imply a centered generated reward bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-587689116961","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3532,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12810"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_bound_of_generatedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Om…","missing":[],"search":"centeredreward_succ_bound_of_generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_bound_of_generatedactionrandompairrawmeanboundedsource raw reward bounds and selected-mean bounds imply a centered generated reward bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_rawMeanBoundedSource","label":"generatedActionRandomPairBoundedCenteredSource_of_rawMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_rawMeanBoundedSource","description":"Turn raw reward and selected-mean bounds into the bounded centered source consumed by the generated-policy conditional reward-law route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8e47698db61d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3533,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12885"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairBoundedCenteredSource_of_rawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => a…","missing":[],"search":"generatedactionrandompairboundedcenteredsource_of_rawmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairboundedcenteredsource_of_rawmeanboundedsource turn raw reward and selected-mean bounds into the bounded centered source consumed by the generated-policy conditional reward-law route. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairBoundedCenteredSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairBoundedCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairBoundedCenteredSource","description":"Bounded generated centered rewards are ambient-integrable.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d680ad029d6f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3534,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:12961"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omeg…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairboundedcenteredsource bounded generated centered rewards are ambient-integrable. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.integrable_exp_mul_of_mem_Icc","label":"integrable_exp_mul_of_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.integrable_exp_mul_of_mem_Icc","description":"A real random variable with an a.e. interval bound has every exponential tilt integrable on a finite measure space.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-054778a22c6b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3535,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13006"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_mul_of_mem_Icc {Omega : Type u} [MeasurableSpace Omega] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (X : Omega -> Real) (lo hi : Real) (hX : AEMeasurable X mu) (hbound : Filter.Eventually (fun omega : Omega => Set.Icc lo hi (X omega)) (ae mu)) (t : Real) : MeasureTheory.Integrable (fun omega : Omega => Real.exp (t * X omega)) mu","missing":[],"search":"integrable_exp_mul_of_mem_icc banditrlproof.conditionalexpectationreward.integrable_exp_mul_of_mem_icc a real random variable with an a.e. interval bound has every exponential tilt integrable on a finite measure space. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_exp_of_generatedActionRandomPairBoundedCenteredSource","label":"centeredReward_succ_integrable_exp_of_generatedActionRandomPairBoundedCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_exp_of_generatedActionRandomPairBoundedCenteredSource","description":"Bounded generated centered rewards have integrable exponential tilts.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9f425b3205f7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3536,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13043"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_exp_of_generatedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"centeredreward_succ_integrable_exp_of_generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_exp_of_generatedactionrandompairboundedcenteredsource bounded generated centered rewards have integrable exponential tilts. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_boundedCenteredSource","label":"generatedActionRandomPairCenteredSource_of_boundedCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_boundedCenteredSource","description":"Turn a bounded centered generated source into the integrability-based centered source consumed by the existing conditional mean-zero route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-abd539ca48a5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3537,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13104"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairCenteredSource_of_boundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action…","missing":[],"search":"generatedactionrandompaircenteredsource_of_boundedcenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompaircenteredsource_of_boundedcenteredsource turn a bounded centered generated source into the integrability-based centered source consumed by the existing conditional mean-zero route. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairBoundedCenteredSource","label":"generatedActionRandomPairMapSource_of_randomPairBoundedCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairBoundedCenteredSource","description":"Project a bounded centered generated random next-pair source directly to its packaged random-pair map source. The bounded source keeps a.e. measurability and interval-bound evidence for integrability consumers. This wrapper records the weaker map-source interface under a stable name for downstream law consumers that do not need those bounds.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-328f86bfb99d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3538,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13162"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => ac…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairboundedcenteredsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairboundedcenteredsource project a bounded centered generated random next-pair source directly to its packaged random-pair map source. the bounded source keeps a.e. measurability and interval-bound evidence for integrability consumers. this wrapper records the weaker map-source interface under a stable name for downstream law consumers that do not need those bounds. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairBoundedCenteredSource","label":"generatedActionActualRewardMapSource_of_randomPairBoundedCenteredSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairBoundedCenteredSource","description":"Convert a bounded centered generated random next-pair source into the weaker generated actual-action reward-coordinate source. The bounded source already contains the random-pair source and state measurability required by the reward-map conversion; the a.e. bound evidence is kept for consumers that need integrability, but is not needed by this weaker interface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bbc4a57575ec","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3539,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13200"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairboundedcenteredsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairboundedcenteredsource convert a bounded centered generated random next-pair source into the weaker generated actual-action reward-coordinate source. the bounded source already contains the random-pair source and state measurability required by the reward-map conversion; the a.e. bound evidence is kept for consumers that need integrability, but is not needed by this weaker interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawMeanBoundedSource","label":"generatedActionRandomPairMapSource_of_randomPairRawMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawMeanBoundedSource","description":"Project a raw-reward/selected-mean bounded generated random next-pair source directly to its packaged random-pair map source. This source already contains the map source used by its history-step and actual-reward projections. The wrapper records that weaker interface under a stable name for downstream law consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2df5933cd76c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3540,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13249"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => act…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairrawmeanboundedsource project a raw-reward/selected-mean bounded generated random next-pair source directly to its packaged random-pair map source. this source already contains the map source used by its history-step and actual-reward projections. the wrapper records that weaker interface under a stable name for downstream law consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawMeanBoundedSource","label":"generatedActionActualRewardMapSource_of_randomPairRawMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawMeanBoundedSource","description":"Convert a raw-reward/selected-mean bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. The raw/mean bounded source already packages the random-pair source and state measurability required by the reward-map conversion. Its raw and selected-mean regularity fields are used by integrability and centered-bound consumers, but not by this weaker source interface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-38283b73d852","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3541,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13287"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => a…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairrawmeanboundedsource convert a raw-reward/selected-mean bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. the raw/mean bounded source already packages the random-pair source and state measurability required by the reward-map conversion. its raw and selected-mean regularity fields are used by integrability and centered-bound consumers, but not by this weaker source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeanBoundedSource","label":"generatedActionRandomPairMapSource_of_randomPairRawBoundMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeanBoundedSource","description":"Project a raw-reward-bound/selected-mean bounded generated random next-pair source directly to its packaged random-pair map source. This source already contains the map source used by its history-step and actual-reward projections. The wrapper records that weaker interface under a stable name for downstream law consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-14021cb65a2b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3542,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13336"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega =…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairrawboundmeanboundedsource project a raw-reward-bound/selected-mean bounded generated random next-pair source directly to its packaged random-pair map source. this source already contains the map source used by its history-step and actual-reward projections. the wrapper records that weaker interface under a stable name for downstream law consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeanBoundedSource","label":"generatedActionActualRewardMapSource_of_randomPairRawBoundMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeanBoundedSource","description":"Convert a raw-reward-bound/selected-mean bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. The raw-bound source already packages the random-pair source and state measurability required by the reward-map conversion. Its raw reward bound and selected-mean regularity fields remain for centered-bound and integrability consumers, but are not needed by this weaker…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e9231563ded1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3543,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13374"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairrawboundmeanboundedsource convert a raw-reward-bound/selected-mean bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. the raw-bound source already packages the random-pair source and state measurability required by the reward-map conversion. its raw reward bound and selected-mean regularity fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","label":"generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","description":"Project a raw-reward-bound/measurable-selected-mean generated random next-pair source directly to its packaged random-pair map source. This source already contains the map source used by its history-step and actual-reward projections. The wrapper records that weaker interface under a stable name for downstream law consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4e6db050ba85","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3544,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13423"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairrawboundmeasurablemeanboundedsource project a raw-reward-bound/measurable-selected-mean generated random next-pair source directly to its packaged random-pair map source. this source already contains the map source used by its history-step and actual-reward projections. the wrapper records that weaker interface under a stable name for downstream law consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","label":"generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","description":"Convert a raw-reward-bound/measurable-selected-mean generated random next-pair source into the weaker generated actual-action reward-coordinate source. This source already packages the random-pair source and state measurability required by the reward-map conversion. The measurable-mean and bound fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-12813159f8eb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3545,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13461"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun ome…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairrawboundmeasurablemeanboundedsource convert a raw-reward-bound/measurable-selected-mean generated random next-pair source into the weaker generated actual-action reward-coordinate source. this source already packages the random-pair source and state measurability required by the reward-map conversion. the measurable-mean and bound fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Project a raw-reward-bound/measurable-mean-range bounded generated random next-pair source directly to its packaged random-pair map source. This is the source-interface counterpart of the weaker actual-reward projection below: the range-bounded source already contains the map source, and this wrapper gives downstream law consumers a stable named entry point without unpacking the structure fields.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ed809421ba1c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3546,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13511"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairrawboundmeasurablemeanrangeboundedsource project a raw-reward-bound/measurable-mean-range bounded generated random next-pair source directly to its packaged random-pair map source. this is the source-interface counterpart of the weaker actual-reward projection below: the range-bounded source already contains the map source, and this wrapper gives downstream law consumers a stable named entry point without unpacking the structure fields. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","label":"generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Convert a raw-reward-bound/measurable-mean-range bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. This range-bounded source already packages the random-pair source and state measurability required by the reward-map conversion. The measurable-mean and range-bound fields remain for centered-bound and integrability consumers, but are not needed by this weaker s…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3b6ebbefca84","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3547,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13551"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fu…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairrawboundmeasurablemeanrangeboundedsource convert a raw-reward-bound/measurable-mean-range bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. this range-bounded source already packages the random-pair source and state measurability required by the reward-map conversion. the measurable-mean and range-bound fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Project a raw-reward-range/measurable-mean-range bounded generated random next-pair source directly to its packaged random-pair map source. This is the explicit-action counterpart of `generatedActionRandomPairMapSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource`: the top practical regularity source already contains the map source, and this wrapper gives downstream law consumers a stable named entry poin…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9e5a0c55c9e5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3548,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13602"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"generatedactionrandompairmapsource_of_randompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_randompairrawrangemeasurablemeanrangeboundedsource project a raw-reward-range/measurable-mean-range bounded generated random next-pair source directly to its packaged random-pair map source. this is the explicit-action counterpart of `generatedactionrandompairmapsource_of_definitionalrawrangemeasurablemeanrangeboundedsource`: the top practical regularity source already contains the map source, and this wrapper gives downstream law consumers a stable named entry point. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionActualRewardMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Convert a raw-reward-range/measurable-mean-range bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. This top explicit-action source already packages the random-pair source and state measurability required by the reward-map conversion. Its deterministic raw-reward and mean range fields remain for centered-bound and integrability consumers, but are not needed by…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b013fdcbb160","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3549,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13642"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fu…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairrawrangemeasurablemeanrangeboundedsource convert a raw-reward-range/measurable-mean-range bounded generated random next-pair source into the weaker generated actual-action reward-coordinate source. this top explicit-action source already packages the random-pair source and state measurability required by the reward-map conversion. its deterministic raw-reward and mean range fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Convert a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source into the weaker definitional generated actual-action reward-coordinate source. This top definitional source already packages the definitional random-pair source. Its deterministic raw-reward and mean range fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interf…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-fc601f04a3ac","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3550,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13694"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource convert a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source into the weaker definitional generated actual-action reward-coordinate source. this top definitional source already packages the definitional random-pair source. its deterministic raw-reward and mean range fields remain for centered-bound and integrability consumers, but are not needed by this weaker source interface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Convert a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source into the explicit generated actual-action reward-coordinate source whose action trace is `generatedActionFromRewardHistory`. This is a convenience projection for consumers that use `GeneratedActionActualRewardMapSource` rather than the definitional source surface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9bab5f0bdeda","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3551,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13739"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward o…","missing":[],"search":"generatedactionactualrewardmapsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource convert a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source into the explicit generated actual-action reward-coordinate source whose action trace is `generatedactionfromrewardhistory`. this is a convenience projection for consumers that use `generatedactionactualrewardmapsource` rather than the definitional source surface. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Convert a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source into the generated finite-pair `partialTraj` source. The practical raw-range source already contains the definitional random-pair map source and context measurability. This wrapper exposes the weaker `GeneratedActionPartialTrajectoryPairLawSource` surface directly. It is only a source conversion: the packaged rand…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e5bf88708866","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3552,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13805"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource convert a definitional raw-reward-range/measurable-mean-range bounded generated random next-pair source into the generated finite-pair `partialtraj` source. the practical raw-range source already contains the definitional random-pair map source and context measurability. this wrapper exposes the weaker `generatedactionpartialtrajectorypairlawsource` surface directly. it is only a source conversion: the packaged random next-pair law is still assumed. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Convert a practical definitional raw-range/measurable-mean-range bounded generated random next-pair source into the generated selected-reward finite-pair-history source. This is a source-conversion theorem route for downstream selected-reward consumers: the practical source already exposes the full finite-pair `partialTraj` law, and the selected finite-pair-history source is obtained by projecting that law to the po…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7fe38dca11d6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3553,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13853"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_randompairdefinitionalrawrangemeasurablemeanrangeboundedsource convert a practical definitional raw-range/measurable-mean-range bounded generated random next-pair source into the generated selected-reward finite-pair-history source. this is a source-conversion theorem route for downstream selected-reward consumers: the practical source already exposes the full finite-pair `partialtraj` law, and the selected finite-pair-history source is obtained by projecting that law to the policy-selected reward coordinate. it still consumes the packaged random next-pair law rather than proving the ambient trajectory identification. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_uniformVarianceBoundedSource","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_uniformVarianceBoundedSource","description":"Project a practical uniform-variance source to the generated selected-reward finite-pair-history law source. The variance ceiling is retained by the original source and can be supplied to conditional-MGF consumers after this selected-law projection.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b0fef9360c95","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3554,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13911"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewar…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_uniformvarianceboundedsource project a practical uniform-variance source to the generated selected-reward finite-pair-history law source. the variance ceiling is retained by the original source and can be supplied to conditional-mgf consumers after this selected-law projection. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_historyVarianceBoundedSource","label":"generatedActionSelectedRewardFinitePairHistoryLawSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_historyVarianceBoundedSource","description":"Project a practical selected-history variance source to the generated selected-reward finite-pair-history law source. The time-indexed variance ceilings remain source-side regularity that can be fed to the selected finite-pair-history conditional-MGF consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-fd9c5b5effa7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3555,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:13960"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionSelectedRewardFinitePairHistoryLawSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewar…","missing":[],"search":"generatedactionselectedrewardfinitepairhistorylawsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionselectedrewardfinitepairhistorylawsource_of_historyvarianceboundedsource project a practical selected-history variance source to the generated selected-reward finite-pair-history law source. the time-indexed variance ceilings remain source-side regularity that can be fed to the selected finite-pair-history conditional-mgf consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","description":"Consume a generated-policy actual reward-coordinate law source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-687d2120f988","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3556,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14006"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward :…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionactualrewardmapsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionactualrewardmapsource consume a generated-policy actual reward-coordinate law source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","description":"Consume a generated-policy actual reward-coordinate law source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8ce7280bc143","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3557,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14100"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omeg…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionactualrewardmapsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionactualrewardmapsource consume a generated-policy actual reward-coordinate law source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource","description":"Consume a generated-policy actual reward-coordinate law source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-135e06306bda","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3558,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14175"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : Reward…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionactualrewardmapsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionactualrewardmapsource consume a generated-policy actual reward-coordinate law source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","description":"Consume an explicit generated-action equality and actual-action reward-coordinate law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-free raw-range wrapper for the narrow reward-coordinate route: callers can provide the generated-action identity and the one-step reward map law directly, without first packaging them as `GeneratedActionActua…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d5f8a6f1f45f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3559,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14260"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianc…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_actual_action_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_reward_map_eq_actual_action_rawrangemeasurablemeanrangebounded consume an explicit generated-action equality and actual-action reward-coordinate law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-free raw-range wrapper for the narrow reward-coordinate route: callers can provide the generated-action identity and the one-step reward map law directly, without first packaging them as `generatedactionactualrewardmapsource`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Consume a full finite-pair `partialTraj` law plus generated-action equality and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the raw-range consumer for the reward-map adapter leaf: the full finite-pair trace law first projects to the actual-action reward-coordinate law, then the existing raw-range reward-coordinate route supplies integrability and condition…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-cb5ec87fcbeb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3560,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14393"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Me…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded consume a full finite-pair `partialtraj` law plus generated-action equality and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the raw-range consumer for the reward-map adapter leaf: the full finite-pair trace law first projects to the actual-action reward-coordinate law, then the existing raw-range reward-coordinate route supplies integrability and conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","description":"Consume a frozen-prefix extension-map `partialTraj` law plus generated-action equality and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the raw-range consumer for the extension-map reward-map adapter leaf: the extension-map law first projects to the actual-action reward-coordinate law, then the existing raw-range reward-coordinate route supplies integrabili…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e4e62f447343","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3561,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14551"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n :…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_actionrewardpartialtrajectorykernel_extend_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_rawrangemeasurablemeanrangebounded consume a frozen-prefix extension-map `partialtraj` law plus generated-action equality and raw reward/selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the raw-range consumer for the extension-map reward-map adapter leaf: the extension-map law first projects to the actual-action reward-coordinate law, then the existing raw-range reward-coordinate route supplies integrability and conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","description":"Consume an explicit generated-action equality and actual-action pair-product law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is one law shape upstream from the reward-coordinate wrapper: the pair-product law is marginalized through `Prod.snd` by the existing generated action route, while the raw/mean range contract supplies integrability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-15859c4c925e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3562,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14710"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceP…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_pair_map_eq_actual_action_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_pair_map_eq_actual_action_rawrangemeasurablemeanrangebounded consume an explicit generated-action equality and actual-action pair-product law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is one law shape upstream from the reward-coordinate wrapper: the pair-product law is marginalized through `prod.snd` by the existing generated action route, while the raw/mean range contract supplies integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","description":"Consume an explicit generated-action equality and fully random next-pair law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the most trajectory-facing direct map-law consumer in this file: the law may keep both successor coordinates random under `condExpKernel`; the existing generated-action route freezes the action coordinate, while the raw/mean ran…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0dd64fa18b1d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3563,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14846"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (va…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_random_pair_map_eq_actual_action_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiontracesucc_random_pair_map_eq_actual_action_rawrangemeasurablemeanrangebounded consume an explicit generated-action equality and fully random next-pair law plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the most trajectory-facing direct map-law consumer in this file: the law may keep both successor coordinates random under `condexpkernel`; the existing generated-action route freezes the action coordinate, while the raw/mean range contract supplies integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource_rawRangeMeasurableMeanRangeBounded","description":"Consume a generated-policy actual reward-coordinate law source plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the explicit-action counterpart of the definitional raw-range wrapper: the source supplies the reward-coordinate conditional law and generated-action identity, while the raw/mean range contract supplies integrability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b2801459eb8e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3564,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:14981"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionactualrewardmapsource_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionactualrewardmapsource_rawrangemeasurablemeanrangebounded consume a generated-policy actual reward-coordinate law source plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the explicit-action counterpart of the definitional raw-range wrapper: the source supplies the reward-coordinate conditional law and generated-action identity, while the raw/mean range contract supplies integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","description":"Consume a definitional generated-action actual reward-coordinate law source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7b3df985b251","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3565,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15088"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiondefinitionalactualrewardmapsource consume a definitional generated-action actual reward-coordinate law source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","description":"Consume a definitional generated-action actual reward-coordinate law source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b374d709d9a7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3566,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15190"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Om…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactiondefinitionalactualrewardmapsource consume a definitional generated-action actual reward-coordinate law source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_generatedActionDefinitionalActualRewardMapSource","label":"generatedActionRandomPairDefinitionalMapSource_of_generatedActionDefinitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_generatedActionDefinitionalActualRewardMapSource","description":"Upgrade a definitional generated-action actual reward-coordinate law source to the stronger definitional generated random next-pair source. The proof factors through the existing full finite-pair `partialTraj` consumer: the actual reward-coordinate source supplies the full trace law, and the definitional random-pair source constructor projects that law back to the next action/reward pair.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-df4b1e105c07","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3567,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15299"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalMapSource_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (source : GeneratedActi…","missing":[],"search":"generatedactionrandompairdefinitionalmapsource_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalmapsource_of_generatedactiondefinitionalactualrewardmapsource upgrade a definitional generated-action actual reward-coordinate law source to the stronger definitional generated random next-pair source. the proof factors through the existing full finite-pair `partialtraj` consumer: the actual reward-coordinate source supplies the full trace law, and the definitional random-pair source constructor projects that law back to the next action/reward pair. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource","description":"Consume a definitional generated-action actual reward-coordinate law source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ba5b26baab46","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3568,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15349"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (defaultAct…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiondefinitionalactualrewardmapsource consume a definitional generated-action actual reward-coordinate law source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","description":"Consume a generated-policy random next-pair law source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9224014e99bf","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3569,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15445"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : O…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairmapsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairmapsource consume a generated-policy random next-pair law source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","description":"Consume a generated-policy random next-pair law source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6f44894311a8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3570,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15539"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairmapsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairmapsource consume a generated-policy random next-pair law source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource","description":"Consume a generated-policy random next-pair law source to obtain ordinary succ-indexed conditional mean-zero for the centered reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5009224e950b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3571,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15614"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKe…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairmapsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairmapsource consume a generated-policy random next-pair law source to obtain ordinary succ-indexed conditional mean-zero for the centered reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource_rawRangeMeasurableMeanRangeBounded","description":"Consume a generated-policy random next-pair law source plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This is the source-packaged companion to the explicit random-pair map-law range wrapper: the source supplies generated-action equality and the random next-pair law, while the raw/mean range contract supplies integrability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7a0baef1e5ac","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3572,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15698"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairmapsource_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairmapsource_rawrangemeasurablemeanrangebounded consume a generated-policy random next-pair law source plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this is the source-packaged companion to the explicit random-pair map-law range wrapper: the source supplies generated-action equality and the random next-pair law, while the raw/mean range contract supplies integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","description":"Consume the definitional generated-action random next-pair source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-730cc6a3e5e8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3573,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15783"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalmapsource consume the definitional generated-action random next-pair source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","description":"Consume the definitional generated-action random next-pair source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-74394203149b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3574,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15885"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omeg…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalmapsource consume the definitional generated-action random next-pair source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource","description":"Consume the definitional generated-action random next-pair source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6ba4d690131c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3575,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:15990"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean varianceProxy) (defaultActio…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalmapsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalmapsource consume the definitional generated-action random next-pair source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource_rawRangeMeasurableMeanRangeBounded","description":"Consume the definitional generated-action random next-pair source plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. This removes the explicit generated action trace and timewise action measurability inputs from the range-based random-pair source consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c2b81142c3e3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3576,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16089"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hmean : Measurable (fun pair : Prod Context Action =>…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalmapsource_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalmapsource_rawrangemeasurablemeanrangebounded consume the definitional generated-action random next-pair source plus raw reward and selected-mean range regularity to obtain ordinary succ-indexed conditional mean-zero. this removes the explicit generated action trace and timewise action measurability inputs from the range-based random-pair source consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeBounded","description":"Consume a policy-selected reward-coordinate law plus raw reward and selected-mean range regularity through the bare definitional random-pair map source. This is the source-route companion to `centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded`: the reward law first builds `GeneratedActionRandomPairDefinitionalMapSource`, then the existing source-level…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-79d6fdb042f0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3577,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16197"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (h…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangebounded consume a policy-selected reward-coordinate law plus raw reward and selected-mean range regularity through the bare definitional random-pair map source. this is the source-route companion to `centeredreward_succ_condexp_eq_zero_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded`: the reward law first builds `generatedactionrandompairdefinitionalmapsource`, then the existing source-level raw/mean range consumer supplies ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","description":"Consume a centered generated-policy random next-pair source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-12bfa6d1a84f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3578,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16328"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardT…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompaircenteredsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompaircenteredsource consume a centered generated-policy random next-pair source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","description":"Consume a centered generated-policy random next-pair source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c6f379a0cc5f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3579,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16400"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompaircenteredsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompaircenteredsource consume a centered generated-policy random next-pair source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairCenteredSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairCenteredSource","description":"Consume a centered generated-policy random next-pair source to obtain ordinary succ-indexed conditional mean-zero for the centered reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5b46478f8c5b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3580,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16474"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompaircenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompaircenteredsource consume a centered generated-policy random next-pair source to obtain ordinary succ-indexed conditional mean-zero for the centered reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairCenteredSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairCenteredSource","description":"Consume a centered generated-policy random next-pair source to obtain the succ-indexed conditional sub-Gaussian MGF witness for the centered reward. The source supplies the generated-action law, the canonical next-pair map law, the centered reward-kernel law, and context/state measurability. The analytic regularity contracts that remain explicit are ambient centered-reward measurability and the deterministic varianc…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5eb0bbab73dd","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3581,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16543"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompaircenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompaircenteredsource consume a centered generated-policy random next-pair source to obtain the succ-indexed conditional sub-gaussian mgf witness for the centered reward. the source supplies the generated-action law, the canonical next-pair map law, the centered reward-kernel law, and context/state measurability. the analytic regularity contracts that remain explicit are ambient centered-reward measurability and the deterministic variance-proxy upper bound; exponential integrability is derived by the integrated target-law transfer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairBoundedCenteredSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairBoundedCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairBoundedCenteredSource","description":"Consume a bounded centered generated-policy random next-pair source to obtain the succ-indexed conditional sub-Gaussian MGF witness for the centered reward. This is the bounded-source wrapper around the centered-source MGF consumer. It uses the bounded source only to lower into `GeneratedActionRandomPairCenteredSource`; ambient centered-reward measurability and variance-proxy domination remain explicit, while expone…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c2a8ee60fc14","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3582,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16771"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurabl…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairboundedcenteredsource consume a bounded centered generated-policy random next-pair source to obtain the succ-indexed conditional sub-gaussian mgf witness for the centered reward. this is the bounded-source wrapper around the centered-source mgf consumer. it uses the bounded source only to lower into `generatedactionrandompaircenteredsource`; ambient centered-reward measurability and variance-proxy domination remain explicit, while exponential integrability is derived. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","description":"Consume a definitional centered generated-policy random next-pair source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6e4f5b9088d2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3583,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16876"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalcenteredsource consume a definitional centered generated-policy random next-pair source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","description":"Consume a definitional centered generated-policy random next-pair source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-562339518499","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3584,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:16981"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t :…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalcenteredsource consume a definitional centered generated-policy random next-pair source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalCenteredSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalCenteredSource","description":"Consume a definitional centered generated-policy random next-pair source to obtain ordinary succ-indexed conditional mean-zero for the centered reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e76fe84d3111","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3585,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17089"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t))…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalcenteredsource consume a definitional centered generated-policy random next-pair source to obtain ordinary succ-indexed conditional mean-zero for the centered reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalCenteredSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalCenteredSource","description":"Consume a definitional centered generated-policy random next-pair source to obtain the succ-indexed conditional sub-Gaussian MGF witness for the centered reward. This is the definitional-action wrapper around the explicit centered-source MGF consumer: the action trace is fixed to `generatedActionFromRewardHistory`, and the definitional source is lowered to the explicit centered source. The centered-reward measurabil…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-23b3a42c376e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3586,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17175"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward ome…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalcenteredsource consume a definitional centered generated-policy random next-pair source to obtain the succ-indexed conditional sub-gaussian mgf witness for the centered reward. this is the definitional-action wrapper around the explicit centered-source mgf consumer: the action trace is fixed to `generatedactionfromrewardhistory`, and the definitional source is lowered to the explicit centered source. the centered-reward measurability and variance ceiling remain explicit; exponential integrability is derived. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","description":"Consume a bounded centered generated-policy random next-pair source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-25ec10ec05d2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3587,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17295"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega ->…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairboundedcenteredsource consume a bounded centered generated-policy random next-pair source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","description":"Consume a bounded centered generated-policy random next-pair source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5fb6694fe39b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3588,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17385"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> Rewar…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairboundedcenteredsource consume a bounded centered generated-policy random next-pair source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairBoundedCenteredSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairBoundedCenteredSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairBoundedCenteredSource","description":"Consume a bounded centered generated-policy random next-pair source to obtain ordinary succ-indexed conditional mean-zero for the centered reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b6cce38a2087","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3589,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17477"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairBoundedCenteredSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairboundedcenteredsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairboundedcenteredsource consume a bounded centered generated-policy random next-pair source to obtain ordinary succ-indexed conditional mean-zero for the centered reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawMeanBoundedSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairRawMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawMeanBoundedSource","description":"Raw reward and selected-mean bounded sources give ambient integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0632128723bf","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3590,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17554"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairrawmeanboundedsource raw reward and selected-mean bounded sources give ambient integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","description":"Consume a raw reward and selected-mean bounded generated-policy source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9e47107c036b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3591,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17631"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> R…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawmeanboundedsource consume a raw reward and selected-mean bounded generated-policy source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","description":"Consume a raw reward and selected-mean bounded generated-policy source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d98988233b27","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3592,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17725"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> Reward…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawmeanboundedsource consume a raw reward and selected-mean bounded generated-policy source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawMeanBoundedSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawMeanBoundedSource","description":"Consume a raw reward and selected-mean bounded generated-policy source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2261c19cbb34","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3593,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17821"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawmeanboundedsource consume a raw reward and selected-mean bounded generated-policy source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeanBoundedSource","label":"centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeanBoundedSource","description":"Raw-reward bounds plus selected-mean measurability/bounds give centered successor reward a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c2f980049741","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3594,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17902"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fu…","missing":[],"search":"centeredreward_succ_aemeasurable_of_generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_aemeasurable_of_generatedactionrandompairrawboundmeanboundedsource raw-reward bounds plus selected-mean measurability/bounds give centered successor reward a.e. measurability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeanBoundedSource","label":"centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeanBoundedSource","description":"Raw-reward bounds plus selected-mean bounds give a centered successor reward interval bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-04679b8b5bf5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3595,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:17981"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega…","missing":[],"search":"centeredreward_succ_bound_of_generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_bound_of_generatedactionrandompairrawboundmeanboundedsource raw-reward bounds plus selected-mean bounds give a centered successor reward interval bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeanBoundedSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeanBoundedSource","description":"Raw-reward bounds plus selected-mean measurability/bounds give ambient integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-90b2255d45b8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3596,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18062"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairrawboundmeanboundedsource raw-reward bounds plus selected-mean measurability/bounds give ambient integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","description":"Consume a raw-reward-bound and selected-mean-bounded generated-policy source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-457a82b5b10c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3597,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18141"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeanboundedsource consume a raw-reward-bound and selected-mean-bounded generated-policy source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","description":"Consume a raw-reward-bound and selected-mean-bounded generated-policy source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0fe091882979","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3598,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18237"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> R…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeanboundedsource consume a raw-reward-bound and selected-mean-bounded generated-policy source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeanBoundedSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeanBoundedSource","description":"Consume a raw-reward-bound and selected-mean-bounded generated-policy source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5aeab6027924","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3599,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18335"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawboundmeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawboundmeanboundedsource consume a raw-reward-bound and selected-mean-bounded generated-policy source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Measurable selected-mean surface plus raw-reward/mean bounds give centered successor reward a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a5772203cc17","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3600,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18418"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Meas…","missing":[],"search":"centeredreward_succ_aemeasurable_of_generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_aemeasurable_of_generatedactionrandompairrawboundmeasurablemeanboundedsource measurable selected-mean surface plus raw-reward/mean bounds give centered successor reward a.e. measurability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Measurable selected-mean surface plus raw-reward/mean bounds give a centered successor reward interval bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d809394f4916","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3601,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18497"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable…","missing":[],"search":"centeredreward_succ_bound_of_generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_bound_of_generatedactionrandompairrawboundmeasurablemeanboundedsource measurable selected-mean surface plus raw-reward/mean bounds give a centered successor reward interval bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Measurable selected-mean surface plus raw-reward/mean bounds give ambient integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6283607490ec","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3602,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18578"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measur…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairrawboundmeasurablemeanboundedsource measurable selected-mean surface plus raw-reward/mean bounds give ambient integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Consume a raw-reward-bound and measurable selected-mean generated-policy source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b758311dd6eb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3603,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18657"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (rewa…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanboundedsource consume a raw-reward-bound and measurable selected-mean generated-policy source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Consume a raw-reward-bound and measurable selected-mean generated-policy source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-899b2307d3de","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3604,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18753"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward :…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanboundedsource consume a raw-reward-bound and measurable selected-mean generated-policy source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","description":"Consume a raw-reward-bound and measurable selected-mean generated-policy source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c6b5df110401","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3605,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18851"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, M…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawboundmeasurablemeanboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawboundmeasurablemeanboundedsource consume a raw-reward-bound and measurable selected-mean generated-policy source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Measurable mean surface plus deterministic mean range bounds give centered successor reward a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0192f8c1d55d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3606,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:18934"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat,…","missing":[],"search":"centeredreward_succ_aemeasurable_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_aemeasurable_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource measurable mean surface plus deterministic mean range bounds give centered successor reward a.e. measurability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Measurable mean surface plus deterministic mean range bounds give a centered successor reward interval bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d0f0520c5b22","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3607,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19013"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measur…","missing":[],"search":"centeredreward_succ_bound_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_bound_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource measurable mean surface plus deterministic mean range bounds give a centered successor reward interval bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Measurable mean surface plus deterministic mean range bounds give ambient integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5bfda9442b10","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3608,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19094"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, M…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource measurable mean surface plus deterministic mean range bounds give ambient integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Consume a raw-reward-bound and deterministic mean-range generated-policy source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-52edf4d2b405","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3609,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19173"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action)…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource consume a raw-reward-bound and deterministic mean-range generated-policy source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Consume a raw-reward-bound and deterministic mean-range generated-policy source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-20bae8933810","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3610,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19269"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (rewa…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource consume a raw-reward-bound and deterministic mean-range generated-policy source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","description":"Consume a raw-reward-bound and deterministic mean-range generated-policy source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4b00e500e884","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3611,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19367"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : N…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawboundmeasurablemeanrangeboundedsource consume a raw-reward-bound and deterministic mean-range generated-policy source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Pointwise raw-reward and mean range bounds give centered successor reward a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bcd60fff2bfe","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3612,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19450"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat,…","missing":[],"search":"centeredreward_succ_aemeasurable_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_aemeasurable_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource pointwise raw-reward and mean range bounds give centered successor reward a.e. measurability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_bound_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Pointwise raw-reward and mean range bounds give a centered successor reward interval bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-71cbd1c2cd87","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3613,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19529"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_bound_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measur…","missing":[],"search":"centeredreward_succ_bound_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_bound_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource pointwise raw-reward and mean range bounds give a centered successor reward interval bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Pointwise raw-reward and mean range bounds give ambient integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5cd1ce560bf6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3614,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19610"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, M…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource pointwise raw-reward and mean range bounds give ambient integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a pointwise raw-reward and mean-range generated-policy source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9705019fb736","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3615,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19689"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action)…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource consume a pointwise raw-reward and mean-range generated-policy source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a pointwise raw-reward and mean-range generated-policy source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c019387e8ecb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3616,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19785"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (rewa…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource consume a pointwise raw-reward and mean-range generated-policy source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a pointwise raw-reward and mean-range generated-policy source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4e0a0823a7ac","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3617,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19883"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (haction : forall t : N…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairrawrangemeasurablemeanrangeboundedsource consume a pointwise raw-reward and mean-range generated-policy source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_aemeasurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain centered successor reward a.e. measurability.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-39068e656d5a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3618,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:19966"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_aemeasurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Om…","missing":[],"search":"centeredreward_succ_aemeasurable_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_aemeasurable_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain centered successor reward a.e. measurability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_measurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_measurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_measurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain full measurability for the generated centered successor reward. This strengthens the a.e.-measurable regularity surface when the source carries timewise reward measurability, context/state measurability, and a measurable mean surface.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b31efffd1b1c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3619,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20049"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_measurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omeg…","missing":[],"search":"centeredreward_succ_measurable_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_measurable_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain full measurability for the generated centered successor reward. this strengthens the a.e.-measurable regularity surface when the source carries timewise reward measurability, context/state measurability, and a measurable mean surface. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_bound_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain centered successor reward interval bounds.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8ef07cb5e890","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3620,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20137"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_bound_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"centeredreward_succ_bound_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_bound_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain centered successor reward interval bounds. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairBoundedCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Turn a definitional generated-action raw-range source into the bounded centered source consumed by the generated-policy conditional reward-law route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8572fb2a8b55","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3621,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20218"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairBoundedCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward o…","missing":[],"search":"generatedactionrandompairboundedcenteredsource_of_definitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairboundedcenteredsource_of_definitionalrawrangemeasurablemeanrangeboundedsource turn a definitional generated-action raw-range source into the bounded centered source consumed by the generated-policy conditional reward-law route. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Turn a definitional generated-action raw-range source into the integrability-based centered source consumed by the generated-policy conditional reward-law route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3d4c5efc4a8c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3622,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20309"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)…","missing":[],"search":"generatedactionrandompaircenteredsource_of_definitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompaircenteredsource_of_definitionalrawrangemeasurablemeanrangeboundedsource turn a definitional generated-action raw-range source into the integrability-based centered source consumed by the generated-policy conditional reward-law route. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain ambient integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d1d9d591170a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3623,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20381"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omeg…","missing":[],"search":"centeredreward_succ_integrable_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain ambient integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_exp_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_integrable_exp_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_exp_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain exponential integrability for the generated centered successor reward.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0ba8cf82ae5a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3624,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20460"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_integrable_exp_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"centeredreward_succ_integrable_exp_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_integrable_exp_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain exponential integrability for the generated centered successor reward. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","label":"generatedActionRandomPairDefinitionalCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Turn a definitional generated-action raw-range source into the definitional integrability-based centered source. This keeps the generated action trace implicit through `generatedActionFromRewardHistory`, instead of first lowering to the explicit centered source over that trace.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5ef35022442e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3625,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20544"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => rew…","missing":[],"search":"generatedactionrandompairdefinitionalcenteredsource_of_definitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalcenteredsource_of_definitionalrawrangemeasurablemeanrangeboundedsource turn a definitional generated-action raw-range source into the definitional integrability-based centered source. this keeps the generated action trace implicit through `generatedactionfromrewardhistory`, instead of first lowering to the explicit centered source over that trace. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_historyVarianceBoundedSource","label":"generatedActionRandomPairDefinitionalCenteredSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the definitional centered-source interface. The history-variance wrapper carries the practical raw/mean range regularity and selected-history variance ceilings for MGF consumers. Consumers that only need the centered law and bounded-derived integrability can use this projection without unpacking the history-variance source manually.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3cc043d4df76","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3626,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20598"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalCenteredSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo…","missing":[],"search":"generatedactionrandompairdefinitionalcenteredsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalcenteredsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the definitional centered-source interface. the history-variance wrapper carries the practical raw/mean range regularity and selected-history variance ceilings for mgf consumers. consumers that only need the centered law and bounded-derived integrability can use this projection without unpacking the history-variance source manually. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_historyVarianceBoundedSource","label":"generatedActionRandomPairBoundedCenteredSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the bounded centered-source interface. This keeps the deterministic centered reward bounds available for downstream integrability and tail consumers while hiding the selected-history variance wrapper when those consumers only need the bounded-centered contract.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-da1747aefbe0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3627,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20664"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairBoundedCenteredSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewar…","missing":[],"search":"generatedactionrandompairboundedcenteredsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairboundedcenteredsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the bounded centered-source interface. this keeps the deterministic centered reward bounds available for downstream integrability and tail consumers while hiding the selected-history variance wrapper when those consumers only need the bounded-centered contract. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_historyVarianceBoundedSource","label":"generatedActionRandomPairCenteredSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the integrability-based centered-source interface. This is the direct projection for consumers that need the centered-source contract rather than the stronger bounded-centered or definitional interfaces.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-31d4ebc410c8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3628,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20737"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairCenteredSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi mea…","missing":[],"search":"generatedactionrandompaircenteredsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompaircenteredsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the integrability-based centered-source interface. this is the direct projection for consumers that need the centered-source contract rather than the stronger bounded-centered or definitional interfaces. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_historyVarianceBoundedSource","label":"generatedActionDefinitionalActualRewardMapSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the weaker definitional actual-action reward-map source. This keeps consumers on the definitional `generatedActionFromRewardHistory` surface while hiding the selected-history variance wrapper and the stronger random-pair law package.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c926686403c3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3629,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20813"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rew…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the weaker definitional actual-action reward-map source. this keeps consumers on the definitional `generatedactionfromrewardhistory` surface while hiding the selected-history variance wrapper and the stronger random-pair law package. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_historyVarianceBoundedSource","label":"generatedActionActualRewardMapSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the explicit generated actual-action reward-map source. This is the direct projection for consumers that only need the selected reward coordinate law over `generatedActionFromRewardHistory`, not the full random next-pair or centered-source interfaces.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-12a40518d2f3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3630,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20879"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo…","missing":[],"search":"generatedactionactualrewardmapsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the explicit generated actual-action reward-map source. this is the direct projection for consumers that only need the selected reward coordinate law over `generatedactionfromrewardhistory`, not the full random next-pair or centered-source interfaces. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_historyVarianceBoundedSource","label":"generatedActionRandomPairMapSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the explicit generated random-pair map source. This exposes the full random next-pair law over `generatedActionFromRewardHistory` for downstream history-step and `partialTraj` consumers while hiding the selected-history variance wrapper.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c308551a1327","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3631,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:20951"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo m…","missing":[],"search":"generatedactionrandompairmapsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the explicit generated random-pair map source. this exposes the full random next-pair law over `generatedactionfromrewardhistory` for downstream history-step and `partialtraj` consumers while hiding the selected-history variance wrapper. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_historyVarianceBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_historyVarianceBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_historyVarianceBoundedSource","description":"Consume a definitional generated-action raw-range/history-variance source to obtain the canonical history-step pair law. This first exposes the packaged generated random-pair map source, then reuses the generic source-level history-step consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-60f688c09afb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3632,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21022"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (f…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_historyvarianceboundedsource consume a definitional generated-action raw-range/history-variance source to obtain the canonical history-step pair law. this first exposes the packaged generated random-pair map source, then reuses the generic source-level history-step consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_uniformVarianceBoundedSource","label":"generatedActionRandomPairMapSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the explicit generated random-pair map source. This exposes the full random next-pair law over `generatedActionFromRewardHistory` for downstream history-step and `partialTraj` consumers while hiding the uniform variance wrapper.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-56e3d00d7675","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3633,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21137"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairMapSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo m…","missing":[],"search":"generatedactionrandompairmapsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairmapsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the explicit generated random-pair map source. this exposes the full random next-pair law over `generatedactionfromrewardhistory` for downstream history-step and `partialtraj` consumers while hiding the uniform variance wrapper. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_uniformVarianceBoundedSource","label":"generatedActionPartialTrajectoryPairLawSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the generated finite-pair `partialTraj` source. The uniform-variance wrapper only adds a global variance ceiling. Consumers that need the weaker full finite-pair source can first project to the packaged raw-range source and then reuse the raw-range-to-`partialTraj` source conversion.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bf62e2f562ba","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3634,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21210"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo reward…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the generated finite-pair `partialtraj` source. the uniform-variance wrapper only adds a global variance ceiling. consumers that need the weaker full finite-pair source can first project to the packaged raw-range source and then reuse the raw-range-to-`partialtraj` source conversion. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_historyVarianceBoundedSource","label":"generatedActionPartialTrajectoryPairLawSource_of_historyVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_historyVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/history-variance source into the generated finite-pair `partialTraj` source. The history-variance wrapper only adds time-indexed selected-history variance ceilings. Consumers that need the weaker full finite-pair source can first project to the packaged raw-range source and then reuse the raw-range-to- `partialTraj` source conversion.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2b3fded556d1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3635,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21277"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionPartialTrajectoryPairLawSource_of_historyVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo reward…","missing":[],"search":"generatedactionpartialtrajectorypairlawsource_of_historyvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionpartialtrajectorypairlawsource_of_historyvarianceboundedsource lower a definitional generated-action raw-range/history-variance source into the generated finite-pair `partialtraj` source. the history-variance wrapper only adds time-indexed selected-history variance ceilings. consumers that need the weaker full finite-pair source can first project to the packaged raw-range source and then reuse the raw-range-to- `partialtraj` source conversion. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_uniformVarianceBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_uniformVarianceBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_uniformVarianceBoundedSource","description":"Consume a definitional generated-action raw-range/uniform-variance source to obtain the canonical history-step pair law. This first exposes the packaged generated random-pair map source, then reuses the generic source-level history-step consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f7e724ad6e60","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3636,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21342"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (f…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_uniformvarianceboundedsource consume a definitional generated-action raw-range/uniform-variance source to obtain the canonical history-step pair law. this first exposes the packaged generated random-pair map source, then reuses the generic source-level history-step consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_uniformVarianceBoundedSource","label":"generatedActionDefinitionalActualRewardMapSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the weaker definitional actual-action reward-map source. This keeps consumers on the definitional `generatedActionFromRewardHistory` surface while hiding the uniform variance wrapper and the stronger random-pair law package.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f585ed1660a8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3637,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21457"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionDefinitionalActualRewardMapSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rew…","missing":[],"search":"generatedactiondefinitionalactualrewardmapsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactiondefinitionalactualrewardmapsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the weaker definitional actual-action reward-map source. this keeps consumers on the definitional `generatedactionfromrewardhistory` surface while hiding the uniform variance wrapper and the stronger random-pair law package. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_uniformVarianceBoundedSource","label":"generatedActionActualRewardMapSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the explicit generated actual-action reward-map source. This is the direct projection for consumers that only need the selected reward coordinate law over `generatedActionFromRewardHistory`, not the full random next-pair or centered-source interfaces.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-fad178e91ce1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3638,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21523"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionActualRewardMapSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi meanLo…","missing":[],"search":"generatedactionactualrewardmapsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionactualrewardmapsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the explicit generated actual-action reward-map source. this is the direct projection for consumers that only need the selected reward coordinate law over `generatedactionfromrewardhistory`, not the full random next-pair or centered-source interfaces. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_uniformVarianceBoundedSource","label":"generatedActionRandomPairBoundedCenteredSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the bounded centered-source interface. This keeps deterministic centered reward bounds available for downstream integrability and tail consumers while hiding the uniform variance wrapper when those consumers only need the bounded-centered contract.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2d92d514b0f1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3639,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21595"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairBoundedCenteredSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewar…","missing":[],"search":"generatedactionrandompairboundedcenteredsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairboundedcenteredsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the bounded centered-source interface. this keeps deterministic centered reward bounds available for downstream integrability and tail consumers while hiding the uniform variance wrapper when those consumers only need the bounded-centered contract. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_uniformVarianceBoundedSource","label":"generatedActionRandomPairCenteredSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the integrability-based centered-source interface. This is the direct projection for consumers that need the centered-source contract rather than the stronger bounded-centered or definitional interfaces.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9f6fcd336ebf","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3640,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21668"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairCenteredSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo rewardHi mea…","missing":[],"search":"generatedactionrandompaircenteredsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompaircenteredsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the integrability-based centered-source interface. this is the direct projection for consumers that need the centered-source contract rather than the stronger bounded-centered or definitional interfaces. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_uniformVarianceBoundedSource","label":"generatedActionRandomPairDefinitionalCenteredSource_of_uniformVarianceBoundedSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_uniformVarianceBoundedSource","description":"Lower a definitional generated-action raw-range/uniform-variance source into the definitional centered-source interface. The uniform-variance wrapper carries the practical raw/mean range regularity and a global variance ceiling for MGF consumers. Consumers that only need the centered law and bounded-derived integrability can use this projection without unpacking the uniform-variance source manually.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-adcf5829f94a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3641,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21745"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalCenteredSource_of_uniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (rewardLo…","missing":[],"search":"generatedactionrandompairdefinitionalcenteredsource_of_uniformvarianceboundedsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalcenteredsource_of_uniformvarianceboundedsource lower a definitional generated-action raw-range/uniform-variance source into the definitional centered-source interface. the uniform-variance wrapper carries the practical raw/mean range regularity and a global variance ceiling for mgf consumers. consumers that only need the centered law and bounded-derived integrability can use this projection without unpacking the uniform-variance source manually. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain the canonical history-step pair law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9ef4beb7ce74","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3642,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21807"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTr…","missing":[],"search":"actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_pair_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain the canonical history-step pair law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain the full finite-pair-trace `partialTraj` law.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b40941b8aac3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3643,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:21922"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace R…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain the full finite-pair-trace `partialtraj` law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a definitional generated-action raw-range source to obtain ordinary succ-indexed conditional mean-zero.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0e757974701a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3644,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22039"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a definitional generated-action raw-range source to obtain ordinary succ-indexed conditional mean-zero. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","description":"Consume a practical definitional raw-range/measurable-mean-range generated random-pair source to obtain the succ-indexed conditional sub-Gaussian MGF witness for the centered reward. This exposes the newest definitional centered-source MGF consumer at the top-level bounded practical source surface. The range evidence is used through the existing conversion into `GeneratedActionRandomPairDefinitionalCenteredSource`;…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-11ac36389f13","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3645,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22136"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun o…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource consume a practical definitional raw-range/measurable-mean-range generated random-pair source to obtain the succ-indexed conditional sub-gaussian mgf witness for the centered reward. this exposes the newest definitional centered-source mgf consumer at the top-level bounded practical source surface. the range evidence is used through the existing conversion into `generatedactionrandompairdefinitionalcenteredsource`; centered-reward measurability and the variance ceiling remain explicit, while the converted selected laws derive exponential integrability. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_centered_meas","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_centered_meas","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_centered_meas","description":"Consume the practical definitional raw-range source for the conditional MGF route from a supplied centered-reward measurability witness. The selected-law MGF transfer now derives exponential integrability directly.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c9d9e5f52711","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3646,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22256"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_centered_meas {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat,…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_centered_meas banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_centered_meas consume the practical definitional raw-range source for the conditional mgf route from a supplied centered-reward measurability witness. the selected-law mgf transfer now derives exponential integrability directly. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_variance_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_variance_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_variance_le","description":"Consume the practical definitional raw-range source for the conditional MGF route while deriving centered-reward measurability from the source regularity fields; exponential integrability follows from the selected laws.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d06ecabb0e5f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3647,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22364"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_variance_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_variance_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_variance_le consume the practical definitional raw-range source for the conditional mgf route while deriving centered-reward measurability from the source regularity fields; exponential integrability follows from the selected laws. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniform_variance_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniform_variance_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniform_variance_le","description":"Consume the practical definitional raw-range source for the conditional MGF route under a deterministic variance-proxy ceiling. This replaces the trimmed-a.e. selected-history variance domination hypothesis with a pointwise kernel-level ceiling on `varianceProxy`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-00d8e919559a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3648,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22479"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniform_variance_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_uniform_variance_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_uniform_variance_le consume the practical definitional raw-range source for the conditional mgf route under a deterministic variance-proxy ceiling. this replaces the trimmed-a.e. selected-history variance domination hypothesis with a pointwise kernel-level ceiling on `varianceproxy`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_history_variance_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_history_variance_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_history_variance_le","description":"Consume the practical definitional raw-range source for the conditional MGF route under a deterministic variance-proxy ceiling on the selected finite reward histories at this time. This is weaker than a global context/action ceiling and still removes the trimmed-a.e. selected-history variance domination side condition.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c67769c12492","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3649,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22567"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_history_variance_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_history_variance_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_history_variance_le consume the practical definitional raw-range source for the conditional mgf route under a deterministic variance-proxy ceiling on the selected finite reward histories at this time. this is weaker than a global context/action ceiling and still removes the trimmed-a.e. selected-history variance domination side condition. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","description":"Consume a practical definitional raw-range source with a packaged deterministic variance-proxy ceiling for the conditional MGF route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9628276bbd6e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3650,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22648"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource consume a practical definitional raw-range source with a packaged deterministic variance-proxy ceiling for the conditional mgf route. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","description":"Consume a practical definitional raw-range source with packaged time-indexed selected-history variance-proxy ceilings for the conditional MGF route.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f97d7b153ffc","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3651,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22724"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource consume a practical definitional raw-range source with packaged time-indexed selected-history variance-proxy ceilings for the conditional mgf route. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_varianceCeiling_le","description":"Consume a packaged selected-history variance source with any deterministic proxy that dominates the selected ceiling at the requested time. This is useful when downstream tail APIs use a coarser shared proxy than the model-side time-indexed variance schedule.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5a6fc102bfa8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3652,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22803"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hrewar…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_varianceceiling_le consume a packaged selected-history variance source with any deterministic proxy that dominates the selected ceiling at the requested time. this is useful when downstream tail apis use a coarser shared proxy than the model-side time-indexed variance schedule. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_historyVarianceSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_historyVarianceSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_historyVarianceSource","description":"Consume a uniform-variance source through the weaker selected-history variance source interface. This is a convenience wrapper for downstream callers that standardize on the history-variance source API while their model supplies a global variance ceiling.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-464dc4c4afed","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3653,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22886"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_historyVarianceSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hr…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_via_historyvariancesource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_via_historyvariancesource consume a uniform-variance source through the weaker selected-history variance source interface. this is a convenience wrapper for downstream callers that standardize on the history-variance source api while their model supplies a global variance ceiling. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_varianceCeiling_le","description":"Consume a packaged uniform-variance source with any deterministic proxy that dominates the global variance ceiling. This is the uniform-source companion to the selected-history larger-proxy consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-aedd00709f81","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3654,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:22986"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hrewar…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_varianceceiling_le consume a packaged uniform-variance source with any deterministic proxy that dominates the global variance ceiling. this is the uniform-source companion to the selected-history larger-proxy consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume the canonical history-step next-pair law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the history-step law-surface companion to the packaged uniform-variance source consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8fdf600b987f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3655,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23091"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : f…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume the canonical history-step next-pair law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the history-step law-surface companion to the packaged uniform-variance source consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume the canonical history-step next-pair law plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-96c9c32e69cf","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3656,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23255"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardT…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume the canonical history-step next-pair law plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume the full finite-pair-trace `partialTraj` law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the full-trace law-surface companion to the packaged uniform-variance source consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-49d85ac989ee","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3657,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23427"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume the full finite-pair-trace `partialtraj` law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the full-trace law-surface companion to the packaged uniform-variance source consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Consume a generated-history `partialTraj` pair-law source plus raw/mean range regularity and a global variance ceiling directly into the succ-indexed conditional MGF witness. This is the source-contract version of `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`: the packaged source supplies the context/state measu…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5ec2201d73de","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3658,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23595"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded consume a generated-history `partialtraj` pair-law source plus raw/mean range regularity and a global variance ceiling directly into the succ-indexed conditional mgf witness. this is the source-contract version of `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`: the packaged source supplies the context/state measurability and full finite-pair partial-trajectory law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume the full finite-pair-trace `partialTraj` law plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-00ed87398937","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3659,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23693"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume the full finite-pair-trace `partialtraj` law plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Consume a generated-history `partialTraj` pair-law source plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. This is the source-contract version of `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le`: the packaged source supplies the context/…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-54921f071ca1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3660,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23869"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hrewar…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le consume a generated-history `partialtraj` pair-law source plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. this is the source-contract version of `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le`: the packaged source supplies the context/state measurability and full finite-pair partial-trajectory law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume the full finite-pair-trace `partialTraj` law plus the practical raw/mean range regularity package and a selected-history variance ceiling to obtain the succ-indexed conditional MGF witness. This is the full-trace law-surface companion to the packaged history-variance source consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f3746ce3c3f2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3661,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:23971"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume the full finite-pair-trace `partialtraj` law plus the practical raw/mean range regularity package and a selected-history variance ceiling to obtain the succ-indexed conditional mgf witness. this is the full-trace law-surface companion to the packaged history-variance source consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Consume a generated-history `partialTraj` pair-law source plus raw/mean range regularity and selected-history variance ceilings directly into the succ-indexed conditional MGF witness. This is the source-contract version of `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`: the packaged source supplies the context/st…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-f4c4d776b1cd","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3662,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:24140"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded consume a generated-history `partialtraj` pair-law source plus raw/mean range regularity and selected-history variance ceilings directly into the succ-indexed conditional mgf witness. this is the source-contract version of `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`: the packaged source supplies the context/state measurability and full finite-pair partial-trajectory law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume the full finite-pair-trace `partialTraj` law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a90299ef1e24","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3663,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:24239"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume the full finite-pair-trace `partialtraj` law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Consume a generated-history `partialTraj` pair-law source plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. This is the source-contract version of `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le`: the pack…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a20cc6840aa5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3664,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:24416"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hrewar…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le consume a generated-history `partialtraj` pair-law source plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. this is the source-contract version of `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le`: the packaged source supplies the context/state measurability and full finite-pair partial-trajectory law. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume the frozen-prefix extension-map `partialTraj` law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the extension-map law-surface companion to the packaged uniform-variance source consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-efd99cac830f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3665,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:24519"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume the frozen-prefix extension-map `partialtraj` law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the extension-map law-surface companion to the packaged uniform-variance source consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume the frozen-prefix extension-map `partialTraj` law plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5fc8e45a192b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3666,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:24689"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> Rewar…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume the frozen-prefix extension-map `partialtraj` law plus the practical uniform variance package at any deterministic proxy that dominates the global ceiling. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume the frozen-prefix extension-map `partialTraj` law plus the practical raw/mean range regularity package and a selected-history variance ceiling to obtain the succ-indexed conditional MGF witness. This is the extension-map law-surface companion to the packaged history-variance source consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bb6221f608ee","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3667,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:24867"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume the frozen-prefix extension-map `partialtraj` law plus the practical raw/mean range regularity package and a selected-history variance ceiling to obtain the succ-indexed conditional mgf witness. this is the extension-map law-surface companion to the packaged history-variance source consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume the frozen-prefix extension-map `partialTraj` law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a38f7e4ff438","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3668,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25038"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> Rewar…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume the frozen-prefix extension-map `partialtraj` law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume the canonical history-step next-pair law plus the practical raw/mean range regularity package and a selected-history variance ceiling to obtain the succ-indexed conditional MGF witness. This is the history-step law-surface companion to the packaged history-variance source consumer above.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ecd31ca277a2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3669,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25217"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : f…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume the canonical history-step next-pair law plus the practical raw/mean range regularity package and a selected-history variance ceiling to obtain the succ-indexed conditional mgf witness. this is the history-step law-surface companion to the packaged history-variance source consumer above. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume the canonical history-step next-pair law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-36d5d781c803","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3670,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25382"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardT…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume the canonical history-step next-pair law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume the full finite-pair-trace `partialTraj` law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This is the full-trace companion to the frozen-prefix extension-map wrapper below.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9ddc9e64dec6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3671,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25555"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_map_eq_definitionalrawrangemeasurablemeanrangebounded directly consume the full finite-pair-trace `partialtraj` law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this is the full-trace companion to the frozen-prefix extension-map wrapper below. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeBounded","description":"Consume a packaged full finite-pair-trace `partialTraj` law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This is the source-level companion to the full-trace law-surface wrapper above: the packaged source supplies context/state measurability and the `partialTraj`/`condExpKernel` law field.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-42e27d732760","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3672,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25711"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_partialtrajectorypairlawsource_definitionalrawrangemeasurablemeanrangebounded consume a packaged full finite-pair-trace `partialtraj` law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this is the source-level companion to the full-trace law-surface wrapper above: the packaged source supplies context/state measurability and the `partialtraj`/`condexpkernel` law field. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeBounded","description":"Consume the generated selected-reward finite-pair-history source plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This is the direct mean-zero consumer for `GeneratedActionSelectedRewardFinitePairHistoryLawSource`: the source is first converted to `GeneratedActionPartialTrajectoryPairLawSource`, then the existing full finite-pair-trace consumer applies. The…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-86cc914b70c5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3673,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25802"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (f…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangebounded consume the generated selected-reward finite-pair-history source plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this is the direct mean-zero consumer for `generatedactionselectedrewardfinitepairhistorylawsource`: the source is first converted to `generatedactionpartialtrajectorypairlawsource`, then the existing full finite-pair-trace consumer applies. the selected-reward law is still a source field, not proved here. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Consume the generated selected-reward finite-pair-history source plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the selected-reward source-level companion to the packaged `partialTraj` source consumer. The source is lowered through `GeneratedActionPartialTrajectoryPairLawSource`; it still consumes the selected-reward law…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7c2fbaad6a0c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3674,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:25905"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded consume the generated selected-reward finite-pair-history source plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the selected-reward source-level companion to the packaged `partialtraj` source consumer. the source is lowered through `generatedactionpartialtrajectorypairlawsource`; it still consumes the selected-reward law field and does not prove ambient trajectory transport. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Consume the generated selected-reward finite-pair-history source plus the practical uniform variance package at any deterministic proxy dominating the global ceiling. This is the coarser-proxy selected-source wrapper for the existing packaged `partialTraj` source MGF consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b81528511593","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3675,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26016"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Ra…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le consume the generated selected-reward finite-pair-history source plus the practical uniform variance package at any deterministic proxy dominating the global ceiling. this is the coarser-proxy selected-source wrapper for the existing packaged `partialtraj` source mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Consume the generated selected-reward finite-pair-history source plus the practical raw/mean range regularity package and selected-history variance ceilings to obtain the succ-indexed conditional MGF witness. This is the selected-source wrapper for the existing packaged `partialTraj` source history-variance consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-dbd15e510974","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3676,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26131"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded consume the generated selected-reward finite-pair-history source plus the practical raw/mean range regularity package and selected-history variance ceilings to obtain the succ-indexed conditional mgf witness. this is the selected-source wrapper for the existing packaged `partialtraj` source history-variance consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Consume the generated selected-reward finite-pair-history source plus the practical selected-history variance package at any deterministic proxy dominating the requested time's ceiling. This is the coarser-proxy selected-source wrapper for the existing packaged `partialTraj` source history-variance MGF consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-22b66a2df21e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3677,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26243"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Ra…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_selectedrewardfinitepairhistorylawsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le consume the generated selected-reward finite-pair-history source plus the practical selected-history variance package at any deterministic proxy dominating the requested time's ceiling. this is the coarser-proxy selected-source wrapper for the existing packaged `partialtraj` source history-variance mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_via_selectedRewardFinitePairHistoryLawSource","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_via_selectedRewardFinitePairHistoryLawSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_via_selectedRewardFinitePairHistoryLawSource","description":"Consume the practical definitional raw-range source through the generated selected-reward finite-pair-history source route to obtain ordinary succ-indexed conditional mean-zero. This records the end-to-end composition used by selected-reward theorem-card routes: the practical package is projected to `GeneratedActionSelectedRewardFinitePairHistoryLawSource`, then the selected source mean-zero consumer applies. The pa…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9c29465c4e1c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3678,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26362"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_via_selectedRewardFinitePairHistoryLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hrew…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_via_selectedrewardfinitepairhistorylawsource banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_via_selectedrewardfinitepairhistorylawsource consume the practical definitional raw-range source through the generated selected-reward finite-pair-history source route to obtain ordinary succ-indexed conditional mean-zero. this records the end-to-end composition used by selected-reward theorem-card routes: the practical package is projected to `generatedactionselectedrewardfinitepairhistorylawsource`, then the selected source mean-zero consumer applies. the packaged random next-pair law is still a source field, not proved here. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","description":"Consume the practical uniform-variance source through the generated selected-reward finite-pair-history source route to obtain the conditional MGF witness at the packaged global variance ceiling.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-31b0c2f5032d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3679,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26457"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> R…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource consume the practical uniform-variance source through the generated selected-reward finite-pair-history source route to obtain the conditional mgf witness at the packaged global variance ceiling. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","description":"Consume the practical uniform-variance source through the selected-source route at any deterministic proxy dominating the packaged global ceiling.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0c04e9c762b4","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3680,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26562"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource_of_varianceceiling_le consume the practical uniform-variance source through the selected-source route at any deterministic proxy dominating the packaged global ceiling. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","description":"Consume the practical selected-history variance source through the generated selected-reward finite-pair-history source route to obtain the conditional MGF witness at the requested time-indexed ceiling.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-cf70161a6397","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3681,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26672"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> R…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource consume the practical selected-history variance source through the generated selected-reward finite-pair-history source route to obtain the conditional mgf witness at the requested time-indexed ceiling. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","description":"Consume the practical selected-history variance source through the selected-source route at any deterministic proxy dominating the requested time-indexed ceiling.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d8709b98938d","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3682,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26778"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_via_selectedrewardfinitepairhistorylawsource_of_varianceceiling_le consume the practical selected-history variance source through the selected-source route at any deterministic proxy dominating the requested time-indexed ceiling. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume a selected-reward law stated at the finite pair-prefix comap conditioning surface plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This composes the comap-law source constructor with the selected-reward source mean-zero consumer. It still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disi…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-80766f5e37fe","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3683,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:26892"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defau…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded directly consume a selected-reward law stated at the finite pair-prefix comap conditioning surface plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this composes the comap-law source constructor with the selected-reward source mean-zero consumer. it still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This is the comap-trim companion to the generated-history-trim mean-zero wrapper. It only changes the input law surface; the proof still routes through `GeneratedAction…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-37c9092b7098","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3684,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:27034"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this is the comap-trim companion to the generated-history-trim mean-zero wrapper. it only changes the input law surface; the proof still routes through `generatedactionselectedrewardfinitepairhistorylawsource` and does not prove the selected-reward law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume a selected-reward law stated at the finite pair-prefix comap conditioning surface plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the uniform-variance MGF companion to the comap mean-zero wrapper: it constructs the full generated finite-pair `partialTraj` source from the comap selected-reward law, then inv…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-d9cb2243ad6c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3685,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:27230"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measura…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume a selected-reward law stated at the finite pair-prefix comap conditioning surface plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the uniform-variance mgf companion to the comap mean-zero wrapper: it constructs the full generated finite-pair `partialtraj` source from the comap selected-reward law, then invokes the existing source-level conditional mgf consumer. it still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the comap-trim companion to the uniform-variance MGF wrapper. It constructs the full generated finite-pair `partialTraj` source from…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0d9d4815c9bc","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3686,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:27384"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the comap-trim companion to the uniform-variance mgf wrapper. it constructs the full generated finite-pair `partialtraj` source from the comap-trim selected-reward law, then invokes the existing source-level conditional mgf consumer. it still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus a global variance ceiling at any deterministic coarser proxy. This is the comap-trim companion to the coarser-proxy uniform-variance MGF wrapper. It constructs the full generated finite-pair `partialTraj` source from the comap-trim selected-reward law, then invokes the e…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-69f98a8f2a4a","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3687,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:27591"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstat…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus a global variance ceiling at any deterministic coarser proxy. this is the comap-trim companion to the coarser-proxy uniform-variance mgf wrapper. it constructs the full generated finite-pair `partialtraj` source from the comap-trim selected-reward law, then invokes the existing source-level larger-proxy conditional mgf consumer. it still consumes the selected-reward conditional law and the proxy-domination proof; it does not prove either one from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume a selected-reward law stated at the finite pair-prefix comap plus a global variance ceiling at any deterministic coarser proxy. This is the coarser-proxy companion to the comap uniform-variance wrapper. It constructs the full generated finite-pair `partialTraj` source from the comap selected-reward law, then invokes the existing source-level larger-proxy conditional MGF consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6118c6a18910","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3688,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:27799"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : f…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume a selected-reward law stated at the finite pair-prefix comap plus a global variance ceiling at any deterministic coarser proxy. this is the coarser-proxy companion to the comap uniform-variance wrapper. it constructs the full generated finite-pair `partialtraj` source from the comap selected-reward law, then invokes the existing source-level larger-proxy conditional mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume a selected-reward law stated at the finite pair-prefix comap, raw/mean range regularity, and selected-history variance ceilings to obtain the succ-indexed conditional MGF witness. This is the history-variance companion to the comap uniform-variance wrapper: it constructs the full generated finite-pair `partialTraj` source from the comap selected-reward law, then invokes the existing source-level con…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-216a49884c0b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3689,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:27955"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measura…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume a selected-reward law stated at the finite pair-prefix comap, raw/mean range regularity, and selected-history variance ceilings to obtain the succ-indexed conditional mgf witness. this is the history-variance companion to the comap uniform-variance wrapper: it constructs the full generated finite-pair `partialtraj` source from the comap selected-reward law, then invokes the existing source-level conditional mgf consumer. it still consumes the selected-reward conditional law; it does not prove that law from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, raw/mean range regularity, and selected-history variance ceilings to obtain the succ-indexed conditional MGF witness. This is the comap-trim companion to the history-variance MGF wrapper. It constructs the full generated finite-pair `partialTraj` source from the comap-trim se…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b67108bda718","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3690,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:28110"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Me…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, raw/mean range regularity, and selected-history variance ceilings to obtain the succ-indexed conditional mgf witness. this is the comap-trim companion to the history-variance mgf wrapper. it constructs the full generated finite-pair `partialtraj` source from the comap-trim selected-reward law, then invokes the existing source-level conditional mgf consumer. it still consumes the selected-reward conditional law and selected-history ceilings; it does not prove either from an ambient trajectory/disintegration construction. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume a selected-reward law stated at the finite pair-prefix comap plus selected-history variance ceilings at any deterministic coarser proxy. This is the coarser-proxy companion to the comap history-variance wrapper. It constructs the full generated finite-pair `partialTraj` source from the comap selected-reward law, then invokes the existing source-level larger-proxy conditional MGF consumer.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-71d0e5da36a2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3691,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:28315"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstate : f…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume a selected-reward law stated at the finite pair-prefix comap plus selected-history variance ceilings at any deterministic coarser proxy. this is the coarser-proxy companion to the comap history-variance wrapper. it constructs the full generated finite-pair `partialtraj` source from the comap selected-reward law, then invokes the existing source-level larger-proxy conditional mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus selected-history variance ceilings at any deterministic coarser proxy. This is the comap-trim companion to the coarser-proxy history-variance MGF wrapper. It constructs the full generated finite-pair `partialTraj` source from the comap-trim selected-reward law, then invo…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-7fb24bf78588","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3692,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:28471"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (hcontext : forall n : Nat, Measurable (context n)) (hstat…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_comap_trim_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume a selected-reward law stated entirely at the finite pair-prefix comap conditioning surface, including the `trim` a.e. filter, plus selected-history variance ceilings at any deterministic coarser proxy. this is the comap-trim companion to the coarser-proxy history-variance mgf wrapper. it constructs the full generated finite-pair `partialtraj` source from the comap-trim selected-reward law, then invokes the existing source-level larger-proxy conditional mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume the frozen-prefix extension-map `partialTraj` law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This avoids requiring callers to first build `GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource` by hand.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3f1f26c4a2aa","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3693,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:28680"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Meas…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardpartialtrajectorykernel_extend_map_eq_definitionalrawrangemeasurablemeanrangebounded directly consume the frozen-prefix extension-map `partialtraj` law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this avoids requiring callers to first build `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource` by hand. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume the canonical history-step next-pair law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This moves the practical generated-action interface one step before the extension-map `partialTraj` surface: callers can provide the next-pair `RewardKernel.actionRewardHistoryStepKernelFamily` law directly.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-30f93673fa8c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3694,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:28840"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measur…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_actionrewardhistorystepkernelfamily_pair_map_eq_definitionalrawrangemeasurablemeanrangebounded directly consume the canonical history-step next-pair law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this moves the practical generated-action interface one step before the extension-map `partialtraj` surface: callers can provide the next-pair `rewardkernel.actionrewardhistorystepkernelfamily` law directly. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume the actual-action reward-coordinate selected-measure law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. The generated action trace supplies the action side of the next-pair split, so the remaining law input is only the reward-coordinate conditional kernel map to `RewardKernel.selectedMeasure` at the generated next action.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-28eef15b8442","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3695,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29050"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Om…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangebounded directly consume the actual-action reward-coordinate selected-measure law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. the generated action trace supplies the action side of the next-pair split, so the remaining law input is only the reward-coordinate conditional kernel map to `rewardkernel.selectedmeasure` at the generated next action. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","description":"Directly consume the policy-selected reward-coordinate law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This is the definitional generated-action counterpart of the explicit policy-selected reward-coordinate wrapper: callers can state the reward law at `(policy i).action (state i history)` without first rewriting it to the generated successor action.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-903c61f711c4","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3696,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29252"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega :…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangebounded directly consume the policy-selected reward-coordinate law plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this is the definitional generated-action counterpart of the explicit policy-selected reward-coordinate wrapper: callers can state the reward law at `(policy i).action (state i history)` without first rewriting it to the generated successor action. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeBounded","label":"centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeBounded","description":"Consume a definitional actual-action reward-coordinate source plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. This is the source-level wrapper for the direct reward-coordinate selected measure law consumer: callers provide `GeneratedActionDefinitionalActualRewardMapSource` instead of separately threading its state measurability and reward-map field.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-425d31dd3549","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3697,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29382"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measu…","missing":[],"search":"centeredreward_succ_condexp_eq_zero_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_condexp_eq_zero_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangebounded consume a definitional actual-action reward-coordinate source plus the practical raw/mean range regularity package to obtain ordinary succ-indexed conditional mean-zero. this is the source-level wrapper for the direct reward-coordinate selected measure law consumer: callers provide `generatedactiondefinitionalactualrewardmapsource` instead of separately threading its state measurability and reward-map field. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","description":"Build the practical definitional raw-range/measurable-mean-range source from a packaged definitional actual-action reward-coordinate law source. This is a source-level constructor: the actual reward-coordinate source already packages the conditional reward law and state measurability, while the caller adds only the context, centered-kernel, mean-measurability, and deterministic raw/mean range regularity needed by th…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b8abac9d211e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3698,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29472"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fu…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_generatedactiondefinitionalactualrewardmapsource build the practical definitional raw-range/measurable-mean-range source from a packaged definitional actual-action reward-coordinate law source. this is a source-level constructor: the actual reward-coordinate source already packages the conditional reward law and state measurability, while the caller adds only the context, centered-kernel, mean-measurability, and deterministic raw/mean range regularity needed by the top raw-range layer. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","description":"Build the packaged uniform-variance source from a packaged definitional actual-action reward-coordinate source. This is the uniform-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource`: the reward-coordinate source and raw/mean range regularity build the base source, while `hvariance` supplies the global deterministi…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4e05fd1fedd8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3699,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29537"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat,…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_generatedactiondefinitionalactualrewardmapsource build the packaged uniform-variance source from a packaged definitional actual-action reward-coordinate source. this is the uniform-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_generatedactiondefinitionalactualrewardmapsource`: the reward-coordinate source and raw/mean range regularity build the base source, while `hvariance` supplies the global deterministic variance proxy ceiling used by downstream conditional mgf consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Consume a packaged definitional actual-action reward-coordinate source plus uniform variance regularity through the practical uniform-variance conditional MGF route. The packaged actual source supplies the reward-coordinate law and state measurability. The caller adds raw/mean range regularity and the global variance ceiling; the proof builds the packaged uniform-variance source and reuses its source-level MGF consu…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-49d54d413416","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3700,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29612"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded consume a packaged definitional actual-action reward-coordinate source plus uniform variance regularity through the practical uniform-variance conditional mgf route. the packaged actual source supplies the reward-coordinate law and state measurability. the caller adds raw/mean range regularity and the global variance ceiling; the proof builds the packaged uniform-variance source and reuses its source-level mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Consume a packaged definitional actual-action reward-coordinate source plus uniform variance regularity at a coarser deterministic proxy. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`: the proof builds the same packaged uniform-variance source and then reuses the sour…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-82c3734404ad","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3701,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29732"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> Reward…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le consume a packaged definitional actual-action reward-coordinate source plus uniform variance regularity at a coarser deterministic proxy. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`: the proof builds the same packaged uniform-variance source and then reuses the source-level larger-proxy consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","description":"Build the packaged selected-history variance source from a packaged definitional actual-action reward-coordinate source. This is the history-variance companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource`: the reward-coordinate source and raw/mean range regularity build the base source, while `hvariance` supplies the time-index…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-277d3ef2b63c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3702,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29857"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat,…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_generatedactiondefinitionalactualrewardmapsource banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_generatedactiondefinitionalactualrewardmapsource build the packaged selected-history variance source from a packaged definitional actual-action reward-coordinate source. this is the history-variance companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_generatedactiondefinitionalactualrewardmapsource`: the reward-coordinate source and raw/mean range regularity build the base source, while `hvariance` supplies the time-indexed selected-history variance ceiling used by downstream conditional mgf consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Consume a packaged definitional actual-action reward-coordinate source plus selected-history variance regularity through the practical history-variance conditional MGF route. The packaged actual source supplies the reward-coordinate law and state measurability. The caller adds raw/mean range regularity and the time-indexed selected-history variance ceilings; the proof builds the packaged history-variance source and…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-2f5f59af0fe3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3703,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:29933"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded consume a packaged definitional actual-action reward-coordinate source plus selected-history variance regularity through the practical history-variance conditional mgf route. the packaged actual source supplies the reward-coordinate law and state measurability. the caller adds raw/mean range regularity and the time-indexed selected-history variance ceilings; the proof builds the packaged history-variance source and reuses its source-level mgf consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Consume a packaged definitional actual-action reward-coordinate source plus selected-history variance regularity at a coarser deterministic proxy. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`: the proof builds the same packaged history-variance source and then reuses…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-0aadd1d2425f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3704,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:30054"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> Reward…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le consume a packaged definitional actual-action reward-coordinate source plus selected-history variance regularity at a coarser deterministic proxy. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_generatedactiondefinitionalactualrewardmapsource_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`: the proof builds the same packaged history-variance source and then reuses the source-level larger-proxy consumer. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_actual_action","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_actual_action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_actual_action","description":"Build the practical definitional generated-policy raw-range source from the actual-action reward-coordinate selected-measure law. This is the base source-level companion to the uniform/history variance source constructors: the reward-coordinate law is first lifted to the frozen-prefix extension-map `partialTraj` law, then packaged into `GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a41ddd3d7559","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3705,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:30179"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega => re…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_reward_map_eq_actual_action build the practical definitional generated-policy raw-range source from the actual-action reward-coordinate selected-measure law. this is the base source-level companion to the uniform/history variance source constructors: the reward-coordinate law is first lifted to the frozen-prefix extension-map `partialtraj` law, then packaged into `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_selected_policy","description":"Build the practical definitional generated-policy raw-range source from the policy-selected reward-coordinate selected-measure law. For `generatedActionFromRewardHistory`, the successor coordinate is definitionally the policy-selected action. This wrapper rewrites the policy-facing reward law into the actual successor-action law and reuses the base actual-action source constructor.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c09a5b865f92","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3706,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:30361"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omega : Omega =>…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeboundedsource_of_reward_map_eq_selected_policy build the practical definitional generated-policy raw-range source from the policy-selected reward-coordinate selected-measure law. for `generatedactionfromrewardhistory`, the successor coordinate is definitionally the policy-selected action. this wrapper rewrites the policy-facing reward law into the actual successor-action law and reuses the base actual-action source constructor. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume the actual-action reward-coordinate selected-measure law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the reward-coordinate law-surface companion to the frozen-prefix extension-map uniform-variance conditional MGF wrapper.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-63aa2b072663","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3707,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:30498"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measu…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume the actual-action reward-coordinate selected-measure law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the reward-coordinate law-surface companion to the frozen-prefix extension-map uniform-variance conditional mgf wrapper. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_actual_action","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_actual_action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_actual_action","description":"Build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the actual-action reward-coordinate selected-measure law. This source-level wrapper preserves the reward-map law surface while exposing the reusable packaged uniform-variance source used by downstream conditional MGF consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ecc41ac0cbdb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3708,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:30709"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omeg…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_reward_map_eq_actual_action build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the actual-action reward-coordinate selected-measure law. this source-level wrapper preserves the reward-map law surface while exposing the reusable packaged uniform-variance source used by downstream conditional mgf consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume the actual-action reward-coordinate selected-measure law plus the practical uniform-variance regularity package at any deterministic proxy that dominates the global variance ceiling. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b33dfa2401b3","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3709,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:30896"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume the actual-action reward-coordinate selected-measure law plus the practical uniform-variance regularity package at any deterministic proxy that dominates the global variance ceiling. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Directly consume the policy-selected reward-coordinate law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional MGF witness. This is the policy-facing counterpart of the actual-action reward-map wrapper: callers can state the reward law at `(policy i).action (state i history)` without first rewriting it to the generated successor action.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-467d78900c14","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3710,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:31052"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Mea…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded directly consume the policy-selected reward-coordinate law plus the practical raw/mean range regularity package and a global variance ceiling to obtain the succ-indexed conditional mgf witness. this is the policy-facing counterpart of the actual-action reward-map wrapper: callers can state the reward law at `(policy i).action (state i history)` without first rewriting it to the generated successor action. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded","description":"Consume the policy-selected reward-coordinate law plus the practical uniform-variance regularity package through the bare definitional random-pair map source. The reward law first builds `GeneratedActionRandomPairDefinitionalMapSource`; that bare source is then wrapped with raw/mean range regularity and the global variance ceiling before the source-level conditional MGF consumer is applied.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a7f7ce868a9c","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3711,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:31219"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangeuniformvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangeuniformvariancebounded consume the policy-selected reward-coordinate law plus the practical uniform-variance regularity package through the bare definitional random-pair map source. the reward law first builds `generatedactionrandompairdefinitionalmapsource`; that bare source is then wrapped with raw/mean range regularity and the global variance ceiling before the source-level conditional mgf consumer is applied. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Consume the policy-selected reward-coordinate law plus uniform-variance regularity through the bare definitional random-pair map source at a coarser deterministic proxy. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-81b893f04afb","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3712,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:31376"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le consume the policy-selected reward-coordinate law plus uniform-variance regularity through the bare definitional random-pair map source at a coarser deterministic proxy. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangeuniformvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Consume the policy-selected reward-coordinate law plus the practical selected-history variance regularity package through the bare definitional random-pair map source. The reward law first builds `GeneratedActionRandomPairDefinitionalMapSource`; that bare source is then wrapped with raw/mean range regularity and the time-indexed selected-history variance ceiling before the source-level conditional MGF consumer is ap…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-1d72b5f02549","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3713,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:31539"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangehistoryvariancebounded consume the policy-selected reward-coordinate law plus the practical selected-history variance regularity package through the bare definitional random-pair map source. the reward law first builds `generatedactionrandompairdefinitionalmapsource`; that bare source is then wrapped with raw/mean range regularity and the time-indexed selected-history variance ceiling before the source-level conditional mgf consumer is applied. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Consume the policy-selected reward-coordinate law plus selected-history variance regularity through the bare definitional random-pair map source at a coarser deterministic proxy. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c766b0b21044","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3714,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:31697"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le consume the policy-selected reward-coordinate law plus selected-history variance regularity through the bare definitional random-pair map source at a coarser deterministic proxy. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalmapsource_rawrangemeasurablemeanrangehistoryvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_selected_policy","description":"Build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the policy-selected reward-coordinate selected-measure law. This is the policy-facing source-level companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_actual_action`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-4217ac61fd13","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3715,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:31859"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun om…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_reward_map_eq_selected_policy build the practical definitional generated-policy raw-range source with a packaged uniform variance ceiling from the policy-selected reward-coordinate selected-measure law. this is the policy-facing source-level companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangeuniformvarianceboundedsource_of_reward_map_eq_actual_action`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","description":"Directly consume the policy-selected reward-coordinate selected-measure law plus the practical uniform-variance regularity package at any deterministic proxy that dominates the global variance ceiling. This is the policy-facing coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-b8434565e084","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3716,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:32002"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded_of_varianceceiling_le directly consume the policy-selected reward-coordinate selected-measure law plus the practical uniform-variance regularity package at any deterministic proxy that dominates the global variance ceiling. this is the policy-facing coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangeuniformvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume the actual-action reward-coordinate selected-measure law plus the practical raw/mean range regularity package and selected-history variance ceilings to obtain the succ-indexed conditional MGF witness. This is the reward-coordinate law-surface companion to the frozen-prefix extension-map history-variance conditional MGF wrapper.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e7a02a7f02e0","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3717,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:32158"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measu…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume the actual-action reward-coordinate selected-measure law plus the practical raw/mean range regularity package and selected-history variance ceilings to obtain the succ-indexed conditional mgf witness. this is the reward-coordinate law-surface companion to the frozen-prefix extension-map history-variance conditional mgf wrapper. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_actual_action","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_actual_action","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_actual_action","description":"Build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the actual-action reward-coordinate selected-measure law. This source-level wrapper preserves the reward-map law surface while exposing the reusable packaged history-variance source used by downstream conditional MGF consumers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-cd74390a8ad6","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3718,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:32370"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_actual_action {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun omeg…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_reward_map_eq_actual_action banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_reward_map_eq_actual_action build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the actual-action reward-coordinate selected-measure law. this source-level wrapper preserves the reward-map law surface while exposing the reusable packaged history-variance source used by downstream conditional mgf consumers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume the actual-action reward-coordinate selected-measure law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. This is the coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-59d8cc1958c7","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3719,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:32558"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume the actual-action reward-coordinate selected-measure law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. this is the coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_actual_action_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Directly consume the policy-selected reward-coordinate law plus the practical raw/mean range regularity package and selected-history variance ceilings to obtain the succ-indexed conditional MGF witness. This is the policy-facing history-variance counterpart of the actual-action reward-map wrapper: callers can state the reward law at `(policy i).action (state i history)` without first rewriting it to the generated su…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-54da225d1068","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3720,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:32716"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Mea…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded directly consume the policy-selected reward-coordinate law plus the practical raw/mean range regularity package and selected-history variance ceilings to obtain the succ-indexed conditional mgf witness. this is the policy-facing history-variance counterpart of the actual-action reward-map wrapper: callers can state the reward law at `(policy i).action (state i history)` without first rewriting it to the generated successor action. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_selected_policy","label":"generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_selected_policy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_selected_policy","description":"Build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the policy-selected reward-coordinate selected-measure law. This is the policy-facing source-level companion to `generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_actual_action`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-cce54fbf050f","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3721,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:32883"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_selected_policy {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Nat, Measurable (fun om…","missing":[],"search":"generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_reward_map_eq_selected_policy banditrlproof.conditionalexpectationreward.generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_reward_map_eq_selected_policy build the practical definitional generated-policy raw-range source with a packaged selected-history variance ceiling from the policy-selected reward-coordinate selected-measure law. this is the policy-facing source-level companion to `generatedactionrandompairdefinitionalrawrangemeasurablemeanrangehistoryvarianceboundedsource_of_reward_map_eq_actual_action`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","label":"centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","description":"Directly consume the policy-selected reward-coordinate selected-measure law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. This is the policy-facing coarser-proxy companion to `centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-a578758fd770","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3722,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33027"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded_of_varianceceiling_le directly consume the policy-selected reward-coordinate selected-measure law plus the practical selected-history variance package at any deterministic proxy that dominates the selected ceiling at the requested time. this is the policy-facing coarser-proxy companion to `centeredreward_succ_hascondsubgaussianmgf_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredRewardSuccProcess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Fixed-horizon Azuma-Hoeffding tail for an ambient generated reward process whose successor reward-coordinate conditional laws are the policy-selected kernel laws. The process is zero at index zero and contains the centered rewards at indices `1, ..., n - 1`. Raw reward and selected-mean range contracts supply the regularity needed by the practical one-step conditional-MGF producer, while `varianceCeiling i` bounds i…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-25b691e832fd","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3723,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33188"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredRewardSuccProcess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : forall t : Na…","missing":[],"search":"centeredrewardsuccprocess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredrewardsuccprocess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded fixed-horizon azuma-hoeffding tail for an ambient generated reward process whose successor reward-coordinate conditional laws are the policy-selected kernel laws. the process is zero at index zero and contains the centered rewards at indices `1, ..., n - 1`. raw reward and selected-mean range contracts supply the regularity needed by the practical one-step conditional-mgf producer, while `varianceceiling i` bounds its selected history-dependent variance proxy. this is a fixed-horizon aggregate tail, not an arm-wise empirical-mean, anytime, or regret theorem. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Delta-calibrated two-sided confidence bound for the ambient practical selected-policy reward-law route. The zero-initialized finite sum contains centered successor rewards at indices `1, ..., n - 1`. The all-time selected reward-coordinate conditional laws and practical raw/mean/history-variance contracts construct every conditional MGF witness; the generic conditional sub-Gaussian confidence theorem then uses the r…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bf3173794d06","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3724,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33372"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward : for…","missing":[],"search":"centeredrewardsuccprocess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredrewardsuccprocess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded delta-calibrated two-sided confidence bound for the ambient practical selected-policy reward-law route. the zero-initialized finite sum contains centered successor rewards at indices `1, ..., n - 1`. the all-time selected reward-coordinate conditional laws and practical raw/mean/history-variance contracts construct every conditional mgf witness; the generic conditional sub-gaussian confidence theorem then uses the radius `sqrt (2 * totalvariance * log (2 / delta))`. positive total proxy variance is explicit because the bad event uses non-strict comparison. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmPullCount","label":"successorArmPullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmPullCount","description":"Number of selections of `arm` among successor coordinates `1, ..., n-1`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e2e33fc16541","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3725,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33553"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def successorArmPullCount {Action : Type x} [DecidableEq Action] (action : ActionTrace Action) (arm : Action) (n : Nat) : Nat","missing":[],"search":"successorarmpullcount banditrlproof.conditionalexpectationreward.successorarmpullcount number of selections of `arm` among successor coordinates `1, ..., n-1`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmPullCount_le_horizon","label":"successorArmPullCount_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmPullCount_le_horizon","description":"The successor pull count is bounded by the enclosing process horizon.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-5be06f03f673","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3726,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33559"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem successorArmPullCount_le_horizon {Action : Type x} [DecidableEq Action] (action : ActionTrace Action) (arm : Action) (n : Nat) : successorArmPullCount action arm n <= n","missing":[],"search":"successorarmpullcount_le_horizon banditrlproof.conditionalexpectationreward.successorarmpullcount_le_horizon the successor pull count is bounded by the enclosing process horizon. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmRewardSum","label":"successorArmRewardSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmRewardSum","description":"Real reward sum from `arm` among successor coordinates `1, ..., n-1`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-10f034e12846","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3727,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33567"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def successorArmRewardSum {Action : Type x} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (arm : Action) (n : Nat) : Real","missing":[],"search":"successorarmrewardsum banditrlproof.conditionalexpectationreward.successorarmrewardsum real reward sum from `arm` among successor coordinates `1, ..., n-1`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean","label":"successorArmEmpiricalMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean","description":"Empirical mean of `arm` over the positive successor pull count.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-e5d7610c714e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3728,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33575"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMean {Action : Type x} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (arm : Action) (n : Nat) : Real","missing":[],"search":"successorarmempiricalmean banditrlproof.conditionalexpectationreward.successorarmempiricalmean empirical mean of `arm` over the positive successor pull count. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_eq_successorArmRewardSum_sub_pullCount_mul","label":"armMaskedCenteredRewardSuccProcess_sum_eq_successorArmRewardSum_sub_pullCount_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_eq_successorArmRewardSum_sub_pullCount_mul","description":"The zero-initialized fixed-arm masked centered sum is the successor selected reward sum minus successor pull count times a stationary arm mean. Only selected coordinates need identify their centering value with `armMean`; the centering surface away from the arm is erased by the mask.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-9a8be8c2a9a1","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3729,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33589"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem armMaskedCenteredRewardSuccProcess_sum_eq_successorArmRewardSum_sub_pullCount_mul {Action : Type x} [DecidableEq Action] (action : ActionTrace Action) (reward : RewardTrace Rat) (center : Nat -> Rat) (arm : Action) (armMean : Rat) (hcenter : forall i : Nat, action (i + 1) = arm -> center i = armMean) (n : Nat) : (Finset.range n).sum (fun t => match t with | 0 => 0 | i + 1 => if action (i + 1) = arm then (((reward (i + 1) - center i : Rat) : Real)) else 0) = successorArmRewardSum action reward arm n - (successorArmPullCount action arm n : Real) * (armMean : Real)","missing":[],"search":"armmaskedcenteredrewardsuccprocess_sum_eq_successorarmrewardsum_sub_pullcount_mul banditrlproof.conditionalexpectationreward.armmaskedcenteredrewardsuccprocess_sum_eq_successorarmrewardsum_sub_pullcount_mul the zero-initialized fixed-arm masked centered sum is the successor selected reward sum minus successor pull count times a stationary arm mean. only selected coordinates need identify their centering value with `armmean`; the centering surface away from the arm is erased by the mask. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedVarianceSuccProcess_sum_eq_mul_successorArmPullCount","label":"armMaskedVarianceSuccProcess_sum_eq_mul_successorArmPullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.armMaskedVarianceSuccProcess_sum_eq_mul_successorArmPullCount","description":"The masked constant successor proxy sums to the proxy times pull count.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-c25552931976","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3730,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33675"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem armMaskedVarianceSuccProcess_sum_eq_mul_successorArmPullCount {Action : Type x} [DecidableEq Action] (action : ActionTrace Action) (arm : Action) (sigma2 : NNReal) (n : Nat) : (Finset.range n).sum (fun t => match t with | 0 => 0 | i + 1 => if action (i + 1) = arm then (((sigma2 : NNReal) : Real)) else 0) = (((sigma2 : NNReal) : Real)) * (successorArmPullCount action arm n : Real)","missing":[],"search":"armmaskedvariancesuccprocess_sum_eq_mul_successorarmpullcount banditrlproof.conditionalexpectationreward.armmaskedvariancesuccprocess_sum_eq_mul_successorarmpullcount the masked constant successor proxy sums to the proxy times pull count. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"armMaskedCenteredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Fixed-arm, fixed-horizon two-sided confidence bound for the practical selected-policy reward-law route. Each successor centered reward is masked by the predictable event that the generated policy selected `arm`. The deterministic proxy remains the full `varianceCeiling i`, so this theorem controls an arm-masked finite sum but does not yet normalize by the random pull count or provide an anytime bound.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-bb4143113047","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3731,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33717"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem armMaskedCenteredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (reward : Omega -> RewardTrace Rat) (…","missing":[],"search":"armmaskedcenteredrewardsuccprocess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.armmaskedcenteredrewardsuccprocess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded fixed-arm, fixed-horizon two-sided confidence bound for the practical selected-policy reward-law route. each successor centered reward is masked by the predictable event that the generated policy selected `arm`. the deterministic proxy remains the full `varianceceiling i`, so this theorem controls an arm-masked finite sum but does not yet normalize by the random pull count or provide an anytime bound. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Practical fixed-arm two-sided tail retaining the random cumulative masked sub-Gaussian proxy. A constant selected-history ceiling `sigma2` is charged only at successor times when the generated policy selects `arm`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-044809dee24e","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3732,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:33930"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (reward : Omega -…","missing":[],"search":"armmaskedcenteredrewardsuccprocess_sum_abs_tail_predictablevariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.armmaskedcenteredrewardsuccprocess_sum_abs_tail_predictablevariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded practical fixed-arm two-sided tail retaining the random cumulative masked sub-gaussian proxy. a constant selected-history ceiling `sigma2` is charged only at successor times when the generated policy selects `arm`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanExactCountRadius","label":"successorArmEmpiricalMeanExactCountRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanExactCountRadius","description":"Count-adaptive radius on the fiber where the successor pull count is `k`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-3a5d94e533d2","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3733,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34134"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMeanExactCountRadius (sigma2 : NNReal) (k : Nat) (delta : Real) : Real","missing":[],"search":"successorarmempiricalmeanexactcountradius banditrlproof.conditionalexpectationreward.successorarmempiricalmeanexactcountradius count-adaptive radius on the fiber where the successor pull count is `k`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Fixed-arm empirical-mean confidence on one exact positive successor-pull-count fiber. Unlike the full-horizon proxy theorem, this endpoint charges exactly `k * sigma2` on the fiber `successorArmPullCount = k`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-527dfd38a2c5","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3734,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34145"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (armMean : Ra…","missing":[],"search":"successorarmempiricalmean_abs_tail_exact_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.successorarmempiricalmean_abs_tail_exact_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded fixed-arm empirical-mean confidence on one exact positive successor-pull-count fiber. unlike the full-horizon proxy theorem, this endpoint charges exactly `k * sigma2` on the fiber `successorarmpullcount = k`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanPeelingRadius","label":"successorArmEmpiricalMeanPeelingRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanPeelingRadius","description":"Equal-share radius after peeling over the at most `n` positive successor pull-count fibers.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-38c3b11ed58b","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3735,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34369"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMeanPeelingRadius (sigma2 : NNReal) (k n : Nat) (delta : Real) : Real","missing":[],"search":"successorarmempiricalmeanpeelingradius banditrlproof.conditionalexpectationreward.successorarmempiricalmeanpeelingradius equal-share radius after peeling over the at most `n` positive successor pull-count fibers. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Count-adaptive fixed-arm empirical-mean confidence at a positive random successor pull count. The exact-count theorem is applied at confidence `delta / n` on every positive fiber and the finite outer-measure union bound combines the fibers. The radius therefore keeps the realized count-adaptive proxy `count * sigma2`; this is a fixed-horizon peeling theorem, not an anytime confidence sequence.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-8a46e585b951","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3736,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34383"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (armMean : R…","missing":[],"search":"successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded count-adaptive fixed-arm empirical-mean confidence at a positive random successor pull count. the exact-count theorem is applied at confidence `delta / n` on every positive fiber and the finite outer-measure union bound combines the fibers. the radius therefore keeps the realized count-adaptive proxy `count * sigma2`; this is a fixed-horizon peeling theorem, not an anytime confidence sequence. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeConfidenceShare","label":"successorArmEmpiricalMeanFiniteArmTimeConfidenceShare","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeConfidenceShare","description":"Equal confidence share for one member of a finite arm/time family.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-57fe17bb09d8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3737,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34524"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMeanFiniteArmTimeConfidenceShare {Action : Type x} [DecidableEq Action] (arms : Finset Action) (T : Nat) (delta : Real) : Real","missing":[],"search":"successorarmempiricalmeanfinitearmtimeconfidenceshare banditrlproof.conditionalexpectationreward.successorarmempiricalmeanfinitearmtimeconfidenceshare equal confidence share for one member of a finite arm/time family. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius","label":"successorArmEmpiricalMeanFiniteArmTimePeelingRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius","description":"Realized-count radius after equal sharing over a finite arm/time family.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-282ba0200bdc","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3738,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34530"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMeanFiniteArmTimePeelingRadius {Action : Type x} [DecidableEq Action] (sigma2 : NNReal) (k n : Nat) (arms : Finset Action) (T : Nat) (delta : Real) : Real","missing":[],"search":"successorarmempiricalmeanfinitearmtimepeelingradius banditrlproof.conditionalexpectationreward.successorarmempiricalmeanfinitearmtimepeelingradius realized-count radius after equal sharing over a finite arm/time family. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeBadEvent","label":"successorArmEmpiricalMeanFiniteArmTimeBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeBadEvent","description":"The union of positive-count empirical-mean failures for every explicit arm and every successor horizon `i + 1`, `i < T`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-6a16231fdb71","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3739,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34541"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def successorArmEmpiricalMeanFiniteArmTimeBadEvent {Omega : Type u} {Action : Type x} [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (arms : Finset Action) (armMean : Action -> Rat) (sigma2 : NNReal) (T : Nat) (delta : Real) : Set Omega","missing":[],"search":"successorarmempiricalmeanfinitearmtimebadevent banditrlproof.conditionalexpectationreward.successorarmempiricalmeanfinitearmtimebadevent the union of positive-count empirical-mean failures for every explicit arm and every successor horizon `i + 1`, `i < t`. definition compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Simultaneous finite-arm, finite-time empirical-mean confidence for the generated selected-policy process. The total confidence budget is first shared uniformly over `arms × Finset.range T`. Each member then invokes the random-pull-count theorem, which internally peels its positive realized-count fibers. This is a fixed finite-horizon union theorem, not an anytime confidence sequence.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-12b21828dfca","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3740,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34564"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (arms…","missing":[],"search":"successorarmempiricalmean_simultaneous_finitearmtime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.successorarmempiricalmean_simultaneous_finitearmtime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded simultaneous finite-arm, finite-time empirical-mean confidence for the generated selected-policy process. the total confidence budget is first shared uniformly over `arms × finset.range t`. each member then invokes the random-pull-count theorem, which internally peels its positive realized-count fibers. this is a fixed finite-horizon union theorem, not an anytime confidence sequence. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"successorArmEmpiricalMean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Fixed-arm empirical-mean confidence at a positive random successor pull count. The arm mean is assumed stationary across the model contexts visited by the fixed arm. The denominator counts exactly successor selections `1, ..., n-1`, matching the zero-initialized masked centered sum. The confidence proxy remains the full deterministic horizon proxy; a count-adaptive proxy requires a different predictable-variance con…","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-96462d03f2b8","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3741,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34725"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem successorArmEmpiricalMean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [DecidableEq Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction arm : Action) (armMean : Rat) (reward : Ome…","missing":[],"search":"successorarmempiricalmean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.successorarmempiricalmean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded fixed-arm empirical-mean confidence at a positive random successor pull count. the arm mean is assumed stationary across the model contexts visited by the fixed arm. the denominator counts exactly successor selections `1, ..., n-1`, matching the zero-initialized masked centered sum. the confidence proxy remains the full deterministic horizon proxy; a count-adaptive proxy requires a different predictable-variance concentration interface. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","label":"centeredRewardSuccProcess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","description":"Delta-calibrated two-sided confidence bound for the empirical average of exactly `m` ambient practical selected-policy centered rewards. The zero-initialized prefix is `Finset.range (m + 1)`: slot zero contributes zero and slots `1, ..., m` contribute the `m` centered successor rewards. The proof reuses the practical sum-confidence theorem and divides its event by the positive deterministic sample count `m`.","url":"../modules/banditrlproof-conditionalrewardlawsource/index.html#decl-ff0b10720015","parent":"module:BanditRLProof.ConditionalRewardLawSource","order":3742,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardLawSource"],["Source","BanditRLProof/ConditionalRewardLawSource.lean:34929"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredRewardSuccProcess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (mu : MeasureTheory.Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (defaultAction : Action) (reward : Omega -> RewardTrace Rat) (hreward :…","missing":[],"search":"centeredrewardsuccprocess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded banditrlproof.conditionalexpectationreward.centeredrewardsuccprocess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalrawrangemeasurablemeanrangehistoryvariancebounded delta-calibrated two-sided confidence bound for the empirical average of exactly `m` ambient practical selected-policy centered rewards. the zero-initialized prefix is `finset.range (m + 1)`: slot zero contributes zero and slots `1, ..., m` contribute the `m` centered successor rewards. the proof reuses the practical sum-confidence theorem and divides its event by the positive deterministic sample count `m`. theorem compiled","shard":"modules/3bcc27636785d725.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeGeometricAllTimeBadEvent","label":"successorArmEmpiricalMeanFintypeGeometricAllTimeBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeGeometricAllTimeBadEvent","description":"The union, over every positive successor horizon and every finite arm, of the canonical positive-random-count empirical-mean failures at equal per-arm geometric confidence shares.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorygeometricalltime/index.html#decl-e888594ecbdd","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","order":3743,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryGeometricAllTime.lean:29"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMeanFintypeGeometricAllTimeBadEvent {Omega : Type u} {Action : Type x} [Fintype Action] [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (armMean : Action -> Rat) (sigma2 : NNReal) (delta : Real) : Set Omega","missing":[],"search":"successorarmempiricalmeanfintypegeometricalltimebadevent banditrlproof.conditionalexpectationreward.successorarmempiricalmeanfintypegeometricalltimebadevent the union, over every positive successor horizon and every finite arm, of the canonical positive-random-count empirical-mean failures at equal per-arm geometric confidence shares. definition compiled","shard":"modules/d795266350ff4336.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","label":"actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","description":"On one canonical generated action/reward trajectory, all positive successor horizons and all arms in a nonempty finite action type satisfy the random-count empirical-mean confidence radius outside one event of outer measure at most `ENNReal.ofReal delta`.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorygeometricalltime/index.html#decl-ffd763bcc092","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","order":3744,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryGeometricAllTime.lean:49"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Fintype Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardK…","missing":[],"search":"actionrewardhistorystepkernelfamily_successorarmempiricalmean_simultaneous_fintype_geometricalltime_abs_tail_ennreal_delta_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_successorarmempiricalmean_simultaneous_fintype_geometricalltime_abs_tail_ennreal_delta_trajmeasure on one canonical generated action/reward trajectory, all positive successor horizons and all arms in a nonempty finite action type satisfy the random-count empirical-mean confidence radius outside one event of outer measure at most `ennreal.ofreal delta`. theorem compiled","shard":"modules/d795266350ff4336.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_condDistrib","label":"historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_condDistrib","description":"An ambient reward process satisfying the configured initial law and successor conditional-distribution recursion has the generated finite-pair `partialTraj` law on the generated history filtration. Unlike the unrestricted theorem card, the action trace here is the policy action generated from the reward history, and the model-side trajectory law is supplied by `hzero` and `hcond` rather than assumed through the conc…","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-dd2ebd3672aa","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3745,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:17"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Rat) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (defaultAction : Action) (reward : Omega ->…","missing":[],"search":"historystepkernelfamily_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_conddistrib banditrlproof.conditionalexpectationreward.historystepkernelfamily_actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_conddistrib an ambient reward process satisfying the configured initial law and successor conditional-distribution recursion has the generated finite-pair `partialtraj` law on the generated history filtration. unlike the unrestricted theorem card, the action trace here is the policy action generated from the reward history, and the model-side trajectory law is supplied by `hzero` and `hcond` rather than assumed through the conclusion. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_pair_condDistrib","label":"actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_pair_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_pair_condDistrib","description":"An arbitrary measurable action/reward process has the full finite-pair `partialTraj` law at time `i` when its successor pair regular conditional distribution given the observed finite pair prefix is the configured history-step action/reward kernel. This is the unrestricted-action theorem-card route under a genuine upstream pair `condDistrib` law. It needs neither a generated-action equality nor a complete trajectory…","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-e14b631eeb20","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3746,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:112"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_pair_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (reward : Omega -> Rewar…","missing":[],"search":"actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_pair_conddistrib banditrlproof.conditionalexpectationreward.actionrewardpartialtrajectorykernel_map_eq_historyfiltrationsucc_finitepairhistoryoftrace_of_pair_conddistrib an arbitrary measurable action/reward process has the full finite-pair `partialtraj` law at time `i` when its successor pair regular conditional distribution given the observed finite pair prefix is the configured history-step action/reward kernel. this is the unrestricted-action theorem-card route under a genuine upstream pair `conddistrib` law. it needs neither a generated-action equality nor a complete trajectory-law assumption. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_stronglyAdapted_historyFiltrationSucc","label":"centeredRewardSuccProcess_stronglyAdapted_historyFiltrationSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_stronglyAdapted_historyFiltrationSucc","description":"The zero-initialized successor centered-reward process is strongly adapted to the full action/reward history filtration for any measurable action trace.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-87963c5d8606","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3747,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:322"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredRewardSuccProcess_stronglyAdapted_historyFiltrationSucc {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSingletonClass Action] (action : Omega -> ActionTrace Action) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (hmean : Measurable (fun pair : Prod Context Action => mean pair.1 pair.2)) (reward : Omega -> RewardTrace Rat) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega…","missing":[],"search":"centeredrewardsuccprocess_stronglyadapted_historyfiltrationsucc banditrlproof.conditionalexpectationreward.centeredrewardsuccprocess_stronglyadapted_historyfiltrationsucc the zero-initialized successor centered-reward process is strongly adapted to the full action/reward history filtration for any measurable action trace. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib_of_ae_variance","label":"centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib_of_ae_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib_of_ae_variance","description":"An explicit successor-pair conditional distribution and the native trimmed-a.e. selected-variance bound yield the centered-reward conditional MGF witness for an arbitrary measurable action trace.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-d8446355b04b","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3748,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:456"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib_of_ae_variance {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Contex…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_pair_conddistrib_of_ae_variance banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_pair_conddistrib_of_ae_variance an explicit successor-pair conditional distribution and the native trimmed-a.e. selected-variance bound yield the centered-reward conditional mgf witness for an arbitrary measurable action trace. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib","label":"centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib","description":"Compatibility wrapper for a pointwise selected-history variance ceiling. The core pair-conditional-law transfer only needs the corresponding trimmed-a.e. variance event.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-e304541d713c","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3749,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:657"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action ->…","missing":[],"search":"centeredreward_succ_hascondsubgaussianmgf_of_pair_conddistrib banditrlproof.conditionalexpectationreward.centeredreward_succ_hascondsubgaussianmgf_of_pair_conddistrib compatibility wrapper for a pointwise selected-history variance ceiling. the core pair-conditional-law transfer only needs the corresponding trimmed-a.e. variance event. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib","label":"actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib","description":"Azuma-Hoeffding upper tail for an arbitrary measurable action/reward process whose every successor-pair conditional law is the configured history-step kernel.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-c55a98c8e6c3","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3750,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:742"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] [Nonempty Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> R…","missing":[],"search":"actionrewardhistorystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_of_pair_conddistrib banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_of_pair_conddistrib azuma-hoeffding upper tail for an arbitrary measurable action/reward process whose every successor-pair conditional law is the configured history-step kernel. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib_on_horizon","label":"actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib_on_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib_on_horizon","description":"Finite-horizon Azuma-Hoeffding upper tail for an arbitrary measurable action/reward process. Only successor-pair laws and trimmed-a.e. selected-variance bounds at indices `i < n - 1` are required. This is the native contract consumed by the finite-sum conditional sub-Gaussian assembler.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-e50334a0089d","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3751,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:868"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib_on_horizon {Omega : Type u} {Context : Type v} {State : Type w} {Action : Type x} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Omega] [Nonempty Action] (mu : Measure Omega) [IsProbabilityMeasure mu] (action : Omega -> ActionTrace Action) (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context ->…","missing":[],"search":"actionrewardhistorystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_of_pair_conddistrib_on_horizon banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_of_pair_conddistrib_on_horizon finite-horizon azuma-hoeffding upper tail for an arbitrary measurable action/reward process. only successor-pair laws and trimmed-a.e. selected-variance bounds at indices `i < n - 1` are required. this is the native contract consumed by the finite-sum conditional sub-gaussian assembler. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure_on_horizon","label":"actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure_on_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure_on_horizon","description":"Canonical action/reward `trajMeasure` Azuma-Hoeffding upper tail. The Mathlib Ionescu--Tulcea trajectory law supplies every successor-pair conditional distribution. The caller only provides the centered reward kernel law and horizon-local selected-history variance ceilings.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorylaw/index.html#decl-b58c2b33fa65","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","order":3752,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryLaw.lean:1003"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure_on_horizon {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.CenteredRewardKernelLaw rewardKernel mean vari…","missing":[],"search":"actionrewardhistorystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_trajmeasure_on_horizon banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_centeredrewardsuccprocess_sum_tail_ennreal_trajmeasure_on_horizon canonical action/reward `trajmeasure` azuma-hoeffding upper tail. the mathlib ionescu--tulcea trajectory law supplies every successor-pair conditional distribution. the caller only provides the centered reward kernel law and horizon-local selected-history variance ceilings. theorem compiled","shard":"modules/60d0c5c2fec7b756.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_action_succ_ae_eq_policy_trajMeasure","label":"actionRewardHistoryStepKernelFamily_action_succ_ae_eq_policy_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_action_succ_ae_eq_policy_trajMeasure","description":"On the canonical action/reward trajectory, every sampled successor action is almost surely the action selected by the policy from the frozen pair prefix. This is an ambient trajectory statement, not merely an a.e. statement inside the regular conditional kernel.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html#decl-c0f1ac7b2dde","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","order":3753,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean:30"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_action_succ_ae_eq_policy_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} {Reward : Type*} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [MeasurableSingletonClass Action] [Countable Action] (mu0 : Measure (Prod Action Reward)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : let stepKernel := RewardKernel.actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate let mu := Prob…","missing":[],"search":"actionrewardhistorystepkernelfamily_action_succ_ae_eq_policy_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_action_succ_ae_eq_policy_trajmeasure on the canonical action/reward trajectory, every sampled successor action is almost surely the action selected by the policy from the frozen pair prefix. this is an ambient trajectory statement, not merely an a.e. statement inside the regular conditional kernel. theorem compiled","shard":"modules/1d074e4c5894e81f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_policyArmMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_trajMeasure_on_horizon","label":"actionRewardHistoryStepKernelFamily_policyArmMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_trajMeasure_on_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_policyArmMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_trajMeasure_on_horizon","description":"Canonical action/reward `trajMeasure` two-sided tail for one policy-selected arm under a random cumulative predictable-variance budget. The mask uses the action selected from the observed reward history, so it is measurable at filtration level `i`. Identifying this mask with the sampled next-action coordinate is intentionally left to a separate transport theorem.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html#decl-bffc43498857","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","order":3754,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean:112"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_policyArmMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_trajMeasure_on_horizon {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.C…","missing":[],"search":"actionrewardhistorystepkernelfamily_policyarmmaskedcenteredrewardsuccprocess_sum_abs_tail_predictablevariance_ennreal_delta_trajmeasure_on_horizon banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_policyarmmaskedcenteredrewardsuccprocess_sum_abs_tail_predictablevariance_ennreal_delta_trajmeasure_on_horizon canonical action/reward `trajmeasure` two-sided tail for one policy-selected arm under a random cumulative predictable-variance budget. the mask uses the action selected from the observed reward history, so it is measurable at filtration level `i`. identifying this mask with the sampled next-action coordinate is intentionally left to a separate transport theorem. theorem compiled","shard":"modules/1d074e4c5894e81f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_sampledArmMaskedCenteredRewardSuccProcess_sum_abs_tail_successorPullCount_ennreal_delta_trajMeasure_on_horizon","label":"actionRewardHistoryStepKernelFamily_sampledArmMaskedCenteredRewardSuccProcess_sum_abs_tail_successorPullCount_ennreal_delta_trajMeasure_on_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_sampledArmMaskedCenteredRewardSuccProcess_sum_abs_tail_successorPullCount_ennreal_delta_trajMeasure_on_horizon","description":"Canonical sampled-arm masked centered-reward tail with the random predictable proxy written as `sigma2` times the actual successor pull count. The proof transports the history-policy mask through the canonical ambient a.e. successor-action law. It remains a fixed-horizon joint event, before exact-count peeling or empirical-mean normalization.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html#decl-3e1db3f00399","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","order":3755,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean:341"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_sampledArmMaskedCenteredRewardSuccProcess_sum_abs_tail_successorPullCount_ennreal_delta_trajMeasure_on_horizon {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal)…","missing":[],"search":"actionrewardhistorystepkernelfamily_sampledarmmaskedcenteredrewardsuccprocess_sum_abs_tail_successorpullcount_ennreal_delta_trajmeasure_on_horizon banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_sampledarmmaskedcenteredrewardsuccprocess_sum_abs_tail_successorpullcount_ennreal_delta_trajmeasure_on_horizon canonical sampled-arm masked centered-reward tail with the random predictable proxy written as `sigma2` times the actual successor pull count. the proof transports the history-policy mask through the canonical ambient a.e. successor-action law. it remains a fixed-horizon joint event, before exact-count peeling or empirical-mean normalization. theorem compiled","shard":"modules/1d074e4c5894e81f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_trajMeasure_on_horizon","label":"actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_trajMeasure_on_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_trajMeasure_on_horizon","description":"Canonical fixed-arm empirical-mean confidence on one exact positive successor pull-count fiber. The canonical sampled-arm tail is instantiated with budget `sigma2 * k`, and the centered sum is divided by the positive count `k`.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html#decl-153314f56f62","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","order":3756,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean:582"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_trajMeasure_on_horizon {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.Cen…","missing":[],"search":"actionrewardhistorystepkernelfamily_successorarmempiricalmean_abs_tail_exact_pullcount_ennreal_delta_trajmeasure_on_horizon banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_successorarmempiricalmean_abs_tail_exact_pullcount_ennreal_delta_trajmeasure_on_horizon canonical fixed-arm empirical-mean confidence on one exact positive successor pull-count fiber. the canonical sampled-arm tail is instantiated with budget `sigma2 * k`, and the centered sum is divided by the positive count `k`. theorem compiled","shard":"modules/1d074e4c5894e81f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon","label":"actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon","description":"Canonical positive random-pull-count empirical-mean confidence obtained by peeling the exact-count theorem over the at most `n` successor count fibers. This is fixed-horizon confidence, not an anytime confidence sequence.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html#decl-400a7af101fa","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","order":3757,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean:817"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.Ce…","missing":[],"search":"actionrewardhistorystepkernelfamily_successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_trajmeasure_on_horizon banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_trajmeasure_on_horizon canonical positive random-pull-count empirical-mean confidence obtained by peeling the exact-count theorem over the at most `n` successor count fibers. this is fixed-horizon confidence, not an anytime confidence sequence. theorem compiled","shard":"modules/1d074e4c5894e81f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_trajMeasure","label":"actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_trajMeasure","description":"Canonical positive random-pull-count empirical-mean confidence obtained by peeling the exact-count theorem over the at most `n` successor count fibers. This is fixed-horizon confidence, not an anytime confidence sequence. -/ theorem actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon {Context : Type v} {State : Type w} {Action : Type x} [Measur…","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorymaskedlaw/index.html#decl-4c2a6c6e92a3","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","order":3758,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryMaskedLaw.lean:956"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Countable Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : RewardKernel.Cen…","missing":[],"search":"actionrewardhistorystepkernelfamily_successorarmempiricalmean_simultaneous_finitearmtime_abs_tail_ennreal_delta_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_successorarmempiricalmean_simultaneous_finitearmtime_abs_tail_ennreal_delta_trajmeasure canonical positive random-pull-count empirical-mean confidence obtained by peeling the exact-count theorem over the at most `n` successor count fibers. this is fixed-horizon confidence, not an anytime confidence sequence. -/ theorem actionrewardhistorystepkernelfamily_successorarmempiricalmean_abs_tail_random_pullcount_ennreal_delta_trajmeasure_on_horizon {context : type v} {state : type w} {action : type x} [measurablespace context] [measurablespace state] [measurablespace action] [standardborelspace action] [measurablesingletonclass action] [countable action] [nonempty action] [decidableeq action] (mu0 : measure (prod action rat)) [isprobabilitymeasure mu0] (rewardkernel : rewardkernel.markovrewardkernel (prod context action) rat) (policy : nat -> policy.measurablepolicy state action) (context : (n : nat) -> ((j : finset.iic n) -> rat) -> context) (state : (n : nat) -> ((j : finset.iic n) -> rat) -> state) (hcontext : forall n : nat, measurable (context n)) (hstate : forall n : nat, measurable (state n)) (mean : context -> action -> rat) (varianceproxy : context -> action -> nnreal) (law :…","shard":"modules/1d074e4c5894e81f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent","label":"successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent","description":"The union, over every positive successor horizon and finite arm, of the canonical positive-random-count empirical-mean failures at equal per-arm telescoping confidence shares.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorytelescopingalltime/index.html#decl-119c2a232994","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","order":3759,"meta":[["Kind","definition"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryTelescopingAllTime.lean:29"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent {Omega : Type u} {Action : Type x} [Fintype Action] [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Rat) (armMean : Action -> Rat) (sigma2 : NNReal) (delta : Real) : Set Omega","missing":[],"search":"successorarmempiricalmeanfintypetelescopingalltimebadevent banditrlproof.conditionalexpectationreward.successorarmempiricalmeanfintypetelescopingalltimebadevent the union, over every positive successor horizon and finite arm, of the canonical positive-random-count empirical-mean failures at equal per-arm telescoping confidence shares. definition compiled","shard":"modules/3386baaf829e7267.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","label":"actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","description":"On one canonical generated action/reward trajectory, every positive successor horizon and every arm satisfies the random-count empirical-mean radius outside one telescoping-schedule event of outer measure at most `ENNReal.ofReal delta`.","url":"../modules/banditrlproof-conditionalrewardpartialtrajectorytelescopingalltime/index.html#decl-3e08f74cd5ef","parent":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","order":3760,"meta":[["Kind","theorem"],["Module","BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime"],["Source","BanditRLProof/ConditionalRewardPartialTrajectoryTelescopingAllTime.lean:49"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure {Context : Type v} {State : Type w} {Action : Type x} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Fintype Action] [Nonempty Action] [DecidableEq Action] (mu0 : Measure (Prod Action Rat)) [IsProbabilityMeasure mu0] (rewardKernel : RewardKernel.MarkovRewardKernel (Prod Context Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((j : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : Rewar…","missing":[],"search":"actionrewardhistorystepkernelfamily_successorarmempiricalmean_simultaneous_fintype_telescopingalltime_abs_tail_ennreal_delta_trajmeasure banditrlproof.conditionalexpectationreward.actionrewardhistorystepkernelfamily_successorarmempiricalmean_simultaneous_fintype_telescopingalltime_abs_tail_ennreal_delta_trajmeasure on one canonical generated action/reward trajectory, every positive successor horizon and every arm satisfies the random-count empirical-mean radius outside one telescoping-schedule event of outer measure at most `ennreal.ofreal delta`. theorem compiled","shard":"modules/3386baaf829e7267.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ActionTrace","label":"ActionTrace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.ActionTrace","description":"A sequence of actions chosen by a bandit algorithm.","url":"../modules/banditrlproof-core/index.html#decl-7a003cb611ce","parent":"module:BanditRLProof.Core","order":3761,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev ActionTrace (Action : Type u)","missing":[],"search":"actiontrace banditrlproof.actiontrace a sequence of actions chosen by a bandit algorithm. abbreviation compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.RewardTrace","label":"RewardTrace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.RewardTrace","description":"A sequence of observed rewards.","url":"../modules/banditrlproof-core/index.html#decl-dc57df983a26","parent":"module:BanditRLProof.Core","order":3762,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev RewardTrace (Reward : Type v)","missing":[],"search":"rewardtrace banditrlproof.rewardtrace a sequence of observed rewards. abbreviation compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount","label":"pullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.pullCount","description":"Number of pulls of action `a` before time `t`.","url":"../modules/banditrlproof-core/index.html#decl-2fa48f75a8f1","parent":"module:BanditRLProof.Core","order":3763,"meta":[["Kind","definition"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def pullCount [DecidableEq Action] (action : ActionTrace Action) (a : Action) : Nat → Nat | 0 => 0 | t + 1 => pullCount action a t + if action t = a then 1 else 0 @[simp] theorem pullCount_zero [DecidableEq Action] (action : ActionTrace Action) (a : Action) : pullCount action a 0 = 0","missing":[],"search":"pullcount banditrlproof.pullcount number of pulls of action `a` before time `t`. definition compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_zero","label":"pullCount_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_zero","description":"@[simp] theorem pullCount_zero [DecidableEq Action] (action : ActionTrace Action) (a : Action) : pullCount action a 0 = 0","url":"../modules/banditrlproof-core/index.html#decl-f508b637a367","parent":"module:BanditRLProof.Core","order":3764,"meta":[["Kind","theorem"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pullCount_zero [DecidableEq Action] (action : ActionTrace Action) (a : Action) : pullCount action a 0 = 0","missing":[],"search":"pullcount_zero banditrlproof.pullcount_zero @[simp] theorem pullcount_zero [decidableeq action] (action : actiontrace action) (a : action) : pullcount action a 0 = 0 theorem compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_succ","label":"pullCount_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_succ","description":"@[simp] theorem pullCount_succ [DecidableEq Action] (action : ActionTrace Action) (a : Action) (t : Nat) : pullCount action a (t + 1) = pullCount action a t + if action t = a then 1 else 0","url":"../modules/banditrlproof-core/index.html#decl-21160c59a557","parent":"module:BanditRLProof.Core","order":3765,"meta":[["Kind","theorem"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pullCount_succ [DecidableEq Action] (action : ActionTrace Action) (a : Action) (t : Nat) : pullCount action a (t + 1) = pullCount action a t + if action t = a then 1 else 0","missing":[],"search":"pullcount_succ banditrlproof.pullcount_succ @[simp] theorem pullcount_succ [decidableeq action] (action : actiontrace action) (a : action) (t : nat) : pullcount action a (t + 1) = pullcount action a t + if action t = a then 1 else 0 theorem compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards","label":"sumRewards","kind":"definition","status":"compiled","subtitle":"BanditRLProof.sumRewards","description":"Sum of rewards obtained from action `a` before time `t`.","url":"../modules/banditrlproof-core/index.html#decl-124205994101","parent":"module:BanditRLProof.Core","order":3766,"meta":[["Kind","definition"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def sumRewards [DecidableEq Action] [OfNat Reward 0] [HAdd Reward Reward Reward] (action : ActionTrace Action) (reward : RewardTrace Reward) (a : Action) : Nat → Reward | 0 => 0 | t + 1 => sumRewards action reward a t + if action t = a then reward t else 0 @[simp] theorem sumRewards_zero [DecidableEq Action] [OfNat Reward 0] [HAdd Reward Reward Reward] (action : ActionTrace Action) (reward : RewardTrace Reward) (a : Action) : sumRewards action reward a 0 = 0","missing":[],"search":"sumrewards banditrlproof.sumrewards sum of rewards obtained from action `a` before time `t`. definition compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_zero","label":"sumRewards_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_zero","description":"@[simp] theorem sumRewards_zero [DecidableEq Action] [OfNat Reward 0] [HAdd Reward Reward Reward] (action : ActionTrace Action) (reward : RewardTrace Reward) (a : Action) : sumRewards action reward a 0 = 0","url":"../modules/banditrlproof-core/index.html#decl-5b0ddbde3ab5","parent":"module:BanditRLProof.Core","order":3767,"meta":[["Kind","theorem"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem sumRewards_zero [DecidableEq Action] [OfNat Reward 0] [HAdd Reward Reward Reward] (action : ActionTrace Action) (reward : RewardTrace Reward) (a : Action) : sumRewards action reward a 0 = 0","missing":[],"search":"sumrewards_zero banditrlproof.sumrewards_zero @[simp] theorem sumrewards_zero [decidableeq action] [ofnat reward 0] [hadd reward reward reward] (action : actiontrace action) (reward : rewardtrace reward) (a : action) : sumrewards action reward a 0 = 0 theorem compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_succ","label":"sumRewards_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_succ","description":"@[simp] theorem sumRewards_succ [DecidableEq Action] [OfNat Reward 0] [HAdd Reward Reward Reward] (action : ActionTrace Action) (reward : RewardTrace Reward) (a : Action) (t : Nat) : sumRewards action reward a (t + 1) = sumRewards action reward a t + if action t = a then reward t else 0","url":"../modules/banditrlproof-core/index.html#decl-ea872371d2c5","parent":"module:BanditRLProof.Core","order":3768,"meta":[["Kind","theorem"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem sumRewards_succ [DecidableEq Action] [OfNat Reward 0] [HAdd Reward Reward Reward] (action : ActionTrace Action) (reward : RewardTrace Reward) (a : Action) (t : Nat) : sumRewards action reward a (t + 1) = sumRewards action reward a t + if action t = a then reward t else 0","missing":[],"search":"sumrewards_succ banditrlproof.sumrewards_succ @[simp] theorem sumrewards_succ [decidableeq action] [ofnat reward 0] [hadd reward reward reward] (action : actiontrace action) (reward : rewardtrace reward) (a : action) (t : nat) : sumrewards action reward a (t + 1) = sumrewards action reward a t + if action t = a then reward t else 0 theorem compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel","label":"FiniteBanditModel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel","description":"A finite stochastic bandit model represented by the mean reward of each arm.","url":"../modules/banditrlproof-core/index.html#decl-b3327136d3fc","parent":"module:BanditRLProof.Core","order":3769,"meta":[["Kind","structure"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure FiniteBanditModel (K : Nat) where","missing":[],"search":"finitebanditmodel banditrlproof.finitebanditmodel a finite stochastic bandit model represented by the mean reward of each arm. structure compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.bestArm","label":"bestArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.bestArm","description":"A computable argmax-style selector for the best arm.","url":"../modules/banditrlproof-core/index.html#decl-5ccc7a23c686","parent":"module:BanditRLProof.Core","order":3770,"meta":[["Kind","definition"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bestArm (model : FiniteBanditModel K) : Fin K","missing":[],"search":"bestarm banditrlproof.finitebanditmodel.bestarm a computable argmax-style selector for the best arm. definition compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.bestMean","label":"bestMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.bestMean","description":"The mean reward of `bestArm`.","url":"../modules/banditrlproof-core/index.html#decl-cadfbf421df7","parent":"module:BanditRLProof.Core","order":3771,"meta":[["Kind","definition"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bestMean (model : FiniteBanditModel K) : Rat","missing":[],"search":"bestmean banditrlproof.finitebanditmodel.bestmean the mean reward of `bestarm`. definition compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.gap","label":"gap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.gap","description":"The gap of an arm relative to the selected best arm.","url":"../modules/banditrlproof-core/index.html#decl-b3eaeba47176","parent":"module:BanditRLProof.Core","order":3772,"meta":[["Kind","definition"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gap (model : FiniteBanditModel K) (arm : Fin K) : Rat","missing":[],"search":"gap banditrlproof.finitebanditmodel.gap the gap of an arm relative to the selected best arm. definition compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.gap_bestArm","label":"gap_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.gap_bestArm","description":"@[simp] theorem gap_bestArm (model : FiniteBanditModel K) : model.gap model.bestArm = 0","url":"../modules/banditrlproof-core/index.html#decl-6916ddd174fb","parent":"module:BanditRLProof.Core","order":3773,"meta":[["Kind","theorem"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem gap_bestArm (model : FiniteBanditModel K) : model.gap model.bestArm = 0","missing":[],"search":"gap_bestarm banditrlproof.finitebanditmodel.gap_bestarm @[simp] theorem gap_bestarm (model : finitebanditmodel k) : model.gap model.bestarm = 0 theorem compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.PolicySketch","label":"PolicySketch","kind":"structure","status":"compiled","subtitle":"BanditRLProof.PolicySketch","description":"A named policy surface that agents can map to a concrete Lean definition.","url":"../modules/banditrlproof-core/index.html#decl-b42faa605c20","parent":"module:BanditRLProof.Core","order":3774,"meta":[["Kind","structure"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure PolicySketch (K : Nat) where","missing":[],"search":"policysketch banditrlproof.policysketch a named policy surface that agents can map to a concrete lean definition. structure compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CertificateStatus","label":"CertificateStatus","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.CertificateStatus","description":"Status of a theorem or candidate in the harness memory.","url":"../modules/banditrlproof-core/index.html#decl-bacf50ec5424","parent":"module:BanditRLProof.Core","order":3775,"meta":[["Kind","inductive type"],["Module","BanditRLProof.Core"],["Source","BanditRLProof/Core.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"inductive CertificateStatus where","missing":[],"search":"certificatestatus banditrlproof.certificatestatus status of a theorem or candidate in the harness memory. inductive type compiled","shard":"modules/384912856a4c897c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.IsSimplexTangent","label":"IsSimplexTangent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.IsSimplexTangent","description":"A finite direction tangent to the affine simplex hyperplane.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-fbb87090e2f9","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3776,"meta":[["Kind","definition"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def IsSimplexTangent {Index : Type*} [DecidableEq Index] (indices : Finset Index) (direction : Index -> Real) : Prop","missing":[],"search":"issimplextangent banditrlproof.curvaturenoisegap.issimplextangent a finite direction tangent to the affine simplex hyperplane. definition compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.tangentPairing","label":"tangentPairing","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.tangentPairing","description":"The finite pairing used to test a signal against a tangent direction.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-6965ed252ebb","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3777,"meta":[["Kind","definition"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def tangentPairing {Index : Type*} [DecidableEq Index] (indices : Finset Index) (signal direction : Index -> Real) : Real","missing":[],"search":"tangentpairing banditrlproof.curvaturenoisegap.tangentpairing the finite pairing used to test a signal against a tangent direction. definition compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.tangentPairing_add_const_of_isSimplexTangent","label":"tangentPairing_add_const_of_isSimplexTangent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.tangentPairing_add_const_of_isSimplexTangent","description":"Adding a constant to a signal is invisible on the simplex tangent space.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-3e13cc71a596","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3778,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem tangentPairing_add_const_of_isSimplexTangent {Index : Type*} [DecidableEq Index] (indices : Finset Index) (signal direction : Index -> Real) (shift : Real) (hdirection : IsSimplexTangent indices direction) : tangentPairing indices (fun i => signal i + shift) direction = tangentPairing indices signal direction","missing":[],"search":"tangentpairing_add_const_of_issimplextangent banditrlproof.curvaturenoisegap.tangentpairing_add_const_of_issimplextangent adding a constant to a signal is invisible on the simplex tangent space. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy","label":"weightedShiftEnergy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy","description":"Weighted squared energy after removing one common scalar shift.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-b0ae712f8ae9","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3779,"meta":[["Kind","definition"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def weightedShiftEnergy {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) (shift : Real) : Real","missing":[],"search":"weightedshiftenergy banditrlproof.curvaturenoisegap.weightedshiftenergy weighted squared energy after removing one common scalar shift. definition compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedCenter","label":"weightedCenter","kind":"definition","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedCenter","description":"The scalar shift selected by a nondegenerate weighted quadratic energy.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-5b05db076378","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3780,"meta":[["Kind","definition"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def weightedCenter {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) : Real","missing":[],"search":"weightedcenter banditrlproof.curvaturenoisegap.weightedcenter the scalar shift selected by a nondegenerate weighted quadratic energy. definition compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.sum_weight_mul_sub_weightedCenter_eq_zero","label":"sum_weight_mul_sub_weightedCenter_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.sum_weight_mul_sub_weightedCenter_eq_zero","description":"The weighted residual about `weightedCenter` has zero weighted sum.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-9c24ca49db51","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3781,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_weight_mul_sub_weightedCenter_eq_zero {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) (hweightSum : indices.sum weight ≠ 0) : indices.sum (fun i => weight i * (signal i - weightedCenter indices weight signal)) = 0","missing":[],"search":"sum_weight_mul_sub_weightedcenter_eq_zero banditrlproof.curvaturenoisegap.sum_weight_mul_sub_weightedcenter_eq_zero the weighted residual about `weightedcenter` has zero weighted sum. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition_of_centered","label":"weightedShiftEnergy_decomposition_of_centered","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition_of_centered","description":"Completing the square around any weighted-centered scalar.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-00f34f32a6aa","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3782,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedShiftEnergy_decomposition_of_centered {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) (center shift : Real) (hcenter : indices.sum (fun i => weight i * (signal i - center)) = 0) : weightedShiftEnergy indices weight signal shift = weightedShiftEnergy indices weight signal center + indices.sum weight * (shift - center) ^ 2","missing":[],"search":"weightedshiftenergy_decomposition_of_centered banditrlproof.curvaturenoisegap.weightedshiftenergy_decomposition_of_centered completing the square around any weighted-centered scalar. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition","label":"weightedShiftEnergy_decomposition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition","description":"Exact min-shift decomposition at the weighted center.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-204229299496","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3783,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedShiftEnergy_decomposition {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) (shift : Real) (hweightSum : indices.sum weight ≠ 0) : weightedShiftEnergy indices weight signal shift = weightedShiftEnergy indices weight signal (weightedCenter indices weight signal) + indices.sum weight * (shift - weightedCenter indices weight signal) ^ 2","missing":[],"search":"weightedshiftenergy_decomposition banditrlproof.curvaturenoisegap.weightedshiftenergy_decomposition exact min-shift decomposition at the weighted center. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_center_le","label":"weightedShiftEnergy_center_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_center_le","description":"The weighted center minimizes the quadratic shift energy.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-4f1a52eb8580","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3784,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedShiftEnergy_center_le {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) (shift : Real) (hweight : forall i, i ∈ indices -> 0 <= weight i) (hweightSum : 0 < indices.sum weight) : weightedShiftEnergy indices weight signal (weightedCenter indices weight signal) <= weightedShiftEnergy indices weight signal shift","missing":[],"search":"weightedshiftenergy_center_le banditrlproof.curvaturenoisegap.weightedshiftenergy_center_le the weighted center minimizes the quadratic shift energy. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_eq_center_iff","label":"weightedShiftEnergy_eq_center_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_eq_center_iff","description":"Strict positivity of the total weight makes the minimizing shift unique.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-aff241116d3a","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3785,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:129"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedShiftEnergy_eq_center_iff {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal : Index -> Real) (shift : Real) (hweightSum : 0 < indices.sum weight) : weightedShiftEnergy indices weight signal shift = weightedShiftEnergy indices weight signal (weightedCenter indices weight signal) <-> shift = weightedCenter indices weight signal","missing":[],"search":"weightedshiftenergy_eq_center_iff banditrlproof.curvaturenoisegap.weightedshiftenergy_eq_center_iff strict positivity of the total weight makes the minimizing shift unique. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_decomposition","label":"weightedShiftEnergy_add_decomposition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_decomposition","description":"Exact signal--noise decomposition, including the interaction term.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-3530332e4676","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3786,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:153"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedShiftEnergy_add_decomposition {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal noise : Index -> Real) (signalShift noiseShift : Real) : weightedShiftEnergy indices weight (fun i => signal i + noise i) (signalShift + noiseShift) = weightedShiftEnergy indices weight signal signalShift + weightedShiftEnergy indices weight noise noiseShift + 2 * indices.sum (fun i => weight i * (signal i - signalShift) * (noise i - noiseShift))","missing":[],"search":"weightedshiftenergy_add_decomposition banditrlproof.curvaturenoisegap.weightedshiftenergy_add_decomposition exact signal--noise decomposition, including the interaction term. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_le_two","label":"weightedShiftEnergy_add_le_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_le_two","description":"A conservative two-term signal--noise energy bound.","url":"../modules/banditrlproof-curvaturenoisegapgeometry/index.html#decl-de1665837868","parent":"module:BanditRLProof.CurvatureNoiseGapGeometry","order":3787,"meta":[["Kind","theorem"],["Module","BanditRLProof.CurvatureNoiseGapGeometry"],["Source","BanditRLProof/CurvatureNoiseGapGeometry.lean:182"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weightedShiftEnergy_add_le_two {Index : Type*} [DecidableEq Index] (indices : Finset Index) (weight signal noise : Index -> Real) (signalShift noiseShift : Real) (hweight : forall i, i ∈ indices -> 0 <= weight i) : weightedShiftEnergy indices weight (fun i => signal i + noise i) (signalShift + noiseShift) <= 2 * weightedShiftEnergy indices weight signal signalShift + 2 * weightedShiftEnergy indices weight noise noiseShift","missing":[],"search":"weightedshiftenergy_add_le_two banditrlproof.curvaturenoisegap.weightedshiftenergy_add_le_two a conservative two-term signal--noise energy bound. theorem compiled","shard":"modules/1762b537fe4635c9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.observedBefore","label":"observedBefore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.observedBefore","description":"Source rounds whose delayed feedback is available before the action at round `t`. The strict inequality matches the NeurIPS 2025 delayed-SAPO source: feedback generated at `s` arrives at the end of `s + delay s`, hence it can be used for the next action only when `s + delay s < t`.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-68c24970f6cf","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3788,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:12"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def observedBefore (delay : Nat → Nat) (t : Nat) : Finset Nat","missing":[],"search":"observedbefore banditrlproof.delayedfeedback.observedbefore source rounds whose delayed feedback is available before the action at round `t`. the strict inequality matches the neurips 2025 delayed-sapo source: feedback generated at `s` arrives at the end of `s + delay s`, hence it can be used for the next action only when `s + delay s < t`. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.outstandingAt","label":"outstandingAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.outstandingAt","description":"Source rounds before `t` whose feedback is not yet available when the action at `t` is chosen. Future source rounds are excluded by `Finset.range t`.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-88221740fd84","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3789,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:18"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def outstandingAt (delay : Nat → Nat) (t : Nat) : Finset Nat","missing":[],"search":"outstandingat banditrlproof.delayedfeedback.outstandingat source rounds before `t` whose feedback is not yet available when the action at `t` is chosen. future source rounds are excluded by `finset.range t`. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.observedBefore_disjoint_outstandingAt","label":"observedBefore_disjoint_outstandingAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.observedBefore_disjoint_outstandingAt","description":"Available and outstanding feedback form disjoint parts of the past.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-68cd74058f4d","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3790,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:22"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem observedBefore_disjoint_outstandingAt (delay : Nat → Nat) (t : Nat) : Disjoint (observedBefore delay t) (outstandingAt delay t)","missing":[],"search":"observedbefore_disjoint_outstandingat banditrlproof.delayedfeedback.observedbefore_disjoint_outstandingat available and outstanding feedback form disjoint parts of the past. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.observedBefore_union_outstandingAt","label":"observedBefore_union_outstandingAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.observedBefore_union_outstandingAt","description":"Every source round before `t` is either available or outstanding.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-f4a2613cace5","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3791,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:30"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem observedBefore_union_outstandingAt (delay : Nat → Nat) (t : Nat) : observedBefore delay t ∪ outstandingAt delay t = Finset.range t","missing":[],"search":"observedbefore_union_outstandingat banditrlproof.delayedfeedback.observedbefore_union_outstandingat every source round before `t` is either available or outstanding. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.card_observedBefore_add_card_outstandingAt","label":"card_observedBefore_add_card_outstandingAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.card_observedBefore_add_card_outstandingAt","description":"The available and outstanding counts add up to the number of past source rounds.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-b73229dcb4a5","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3792,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:46"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","delayed-nonstationary"]],"statement":"theorem card_observedBefore_add_card_outstandingAt (delay : Nat → Nat) (t : Nat) : (observedBefore delay t).card + (outstandingAt delay t).card = t","missing":[],"search":"card_observedbefore_add_card_outstandingat banditrlproof.delayedfeedback.card_observedbefore_add_card_outstandingat the available and outstanding counts add up to the number of past source rounds. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.DelayedFeedback.outstandingCount","label":"outstandingCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.outstandingCount","description":"Number of source rounds before `t` whose feedback cannot yet be used by the action at `t`. This action-time surface is kept separate from the paper's end-of-round `sigma(t)` until their one-based/zero-based index bridge is proved.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-bcedc40b759b","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3793,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:58"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def outstandingCount (delay : Nat → Nat) (t : Nat) : Nat","missing":[],"search":"outstandingcount banditrlproof.delayedfeedback.outstandingcount number of source rounds before `t` whose feedback cannot yet be used by the action at `t`. this action-time surface is kept separate from the paper's end-of-round `sigma(t)` until their one-based/zero-based index bridge is proved. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.maxOutstandingBeforeThrough","label":"maxOutstandingBeforeThrough","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.maxOutstandingBeforeThrough","description":"Largest action-time outstanding count through the inclusive horizon.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-d13f76f0c072","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3794,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:62"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def maxOutstandingBeforeThrough (delay : Nat → Nat) (horizon : Nat) : Nat","missing":[],"search":"maxoutstandingbeforethrough banditrlproof.delayedfeedback.maxoutstandingbeforethrough largest action-time outstanding count through the inclusive horizon. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.outstandingCount_le_round","label":"outstandingCount_le_round","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.outstandingCount_le_round","description":"An action-time outstanding count never exceeds the number of past source rounds.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-fb74dedf0c92","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3795,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:67"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem outstandingCount_le_round (delay : Nat → Nat) (t : Nat) : outstandingCount delay t ≤ t","missing":[],"search":"outstandingcount_le_round banditrlproof.delayedfeedback.outstandingcount_le_round an action-time outstanding count never exceeds the number of past source rounds. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.outstandingCount_le_maxOutstandingBeforeThrough","label":"outstandingCount_le_maxOutstandingBeforeThrough","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.outstandingCount_le_maxOutstandingBeforeThrough","description":"Every action-time outstanding count inside the horizon is bounded by the finite maximum surface.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-d1c53a121d6f","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3796,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:75"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem outstandingCount_le_maxOutstandingBeforeThrough (delay : Nat → Nat) {t horizon : Nat} (ht : t ≤ horizon) : outstandingCount delay t ≤ maxOutstandingBeforeThrough delay horizon","missing":[],"search":"outstandingcount_le_maxoutstandingbeforethrough banditrlproof.delayedfeedback.outstandingcount_le_maxoutstandingbeforethrough every action-time outstanding count inside the horizon is bounded by the finite maximum surface. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.oneBasedDelayShift","label":"oneBasedDelayShift","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.oneBasedDelayShift","description":"Convert a one-based paper delay sequence into the zero-based source carrier used by `outstandingAt`: zero-based source `s` represents paper source round `s + 1`.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-453d84b6201e","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3797,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:86"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def oneBasedDelayShift (delay : Nat → Nat) (s : Nat) : Nat","missing":[],"search":"onebaseddelayshift banditrlproof.delayedfeedback.onebaseddelayshift convert a one-based paper delay sequence into the zero-based source carrier used by `outstandingat`: zero-based source `s` represents paper source round `s + 1`. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperMissingAtEnd","label":"paperMissingAtEnd","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperMissingAtEnd","description":"The paper's end-of-round missing-feedback set, represented on a zero-based finite carrier. An element `s` denotes paper round `s + 1`, and the predicate is exactly `(s + 1) + d_(s+1) > t` for source rounds at most `t`.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-d473ba895fbb","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3798,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:92"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def paperMissingAtEnd (delay : Nat → Nat) (t : Nat) : Finset Nat","missing":[],"search":"papermissingatend banditrlproof.delayedfeedback.papermissingatend the paper's end-of-round missing-feedback set, represented on a zero-based finite carrier. an element `s` denotes paper round `s + 1`, and the predicate is exactly `(s + 1) + d_(s+1) > t` for source rounds at most `t`. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperMissingAtEnd_eq_outstandingAt_oneBasedDelayShift","label":"paperMissingAtEnd_eq_outstandingAt_oneBasedDelayShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperMissingAtEnd_eq_outstandingAt_oneBasedDelayShift","description":"The source paper's one-based end-of-round missing set is exactly the action-time outstanding set after reindexing source rounds and delays. This lemma is the explicit off-by-one bridge; the two surfaces are not identified by notation alone.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-c687b9ec4d42","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3799,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:99"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem paperMissingAtEnd_eq_outstandingAt_oneBasedDelayShift (delay : Nat → Nat) (t : Nat) : paperMissingAtEnd delay t = outstandingAt (oneBasedDelayShift delay) t","missing":[],"search":"papermissingatend_eq_outstandingat_onebaseddelayshift banditrlproof.delayedfeedback.papermissingatend_eq_outstandingat_onebaseddelayshift the source paper's one-based end-of-round missing set is exactly the action-time outstanding set after reindexing source rounds and delays. this lemma is the explicit off-by-one bridge; the two surfaces are not identified by notation alone. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount","label":"paperMissingCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperMissingCount","description":"Paper-facing count corresponding to `sigma(t)`, with paper rounds `1, ..., t` represented by zero-based source indices.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-9eacfe13003e","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3800,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:113"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def paperMissingCount (delay : Nat → Nat) (t : Nat) : Nat","missing":[],"search":"papermissingcount banditrlproof.delayedfeedback.papermissingcount paper-facing count corresponding to `sigma(t)`, with paper rounds `1, ..., t` represented by zero-based source indices. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_eq_outstandingCount_oneBasedDelayShift","label":"paperMissingCount_eq_outstandingCount_oneBasedDelayShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperMissingCount_eq_outstandingCount_oneBasedDelayShift","description":"Cardinal form of the one-based/end-of-round indexing bridge.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-b4d80ee61ba9","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3801,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:117"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem paperMissingCount_eq_outstandingCount_oneBasedDelayShift (delay : Nat → Nat) (t : Nat) : paperMissingCount delay t = outstandingCount (oneBasedDelayShift delay) t","missing":[],"search":"papermissingcount_eq_outstandingcount_onebaseddelayshift banditrlproof.delayedfeedback.papermissingcount_eq_outstandingcount_onebaseddelayshift cardinal form of the one-based/end-of-round indexing bridge. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_le_round","label":"paperMissingCount_le_round","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperMissingCount_le_round","description":"The paper-facing missing count at round `t` is at most `t`.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-a9d66b0cfc39","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3802,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:124"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem paperMissingCount_le_round (delay : Nat → Nat) (t : Nat) : paperMissingCount delay t ≤ t","missing":[],"search":"papermissingcount_le_round banditrlproof.delayedfeedback.papermissingcount_le_round the paper-facing missing count at round `t` is at most `t`. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperSigmaMaxThrough","label":"paperSigmaMaxThrough","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperSigmaMaxThrough","description":"Finite maximum of the paper-facing missing-count surface through an inclusive horizon.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-d8b4e69194be","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3803,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:131"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def paperSigmaMaxThrough (delay : Nat → Nat) (horizon : Nat) : Nat","missing":[],"search":"papersigmamaxthrough banditrlproof.delayedfeedback.papersigmamaxthrough finite maximum of the paper-facing missing-count surface through an inclusive horizon. definition compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_le_paperSigmaMaxThrough","label":"paperMissingCount_le_paperSigmaMaxThrough","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.paperMissingCount_le_paperSigmaMaxThrough","description":"Each paper-facing missing count is bounded by its finite maximum through the declared horizon.","url":"../modules/banditrlproof-delayedfeedback-accounting/index.html#decl-4800f6fb7ea8","parent":"module:BanditRLProof.DelayedFeedback.Accounting","order":3804,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Accounting"],["Source","BanditRLProof/DelayedFeedback/Accounting.lean:136"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem paperMissingCount_le_paperSigmaMaxThrough (delay : Nat → Nat) {t horizon : Nat} (ht : t ≤ horizon) : paperMissingCount delay t ≤ paperSigmaMaxThrough delay horizon","missing":[],"search":"papermissingcount_le_papersigmamaxthrough banditrlproof.delayedfeedback.papermissingcount_le_papersigmamaxthrough each paper-facing missing count is bounded by its finite maximum through the declared horizon. theorem compiled","shard":"modules/8ce75c939e50ce73.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation","label":"DelayedSAPOAllocation","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation","description":"A line-15 allocation together with the exact hypotheses needed to make its coordinates a probability distribution. EAP must eventually construct these fields; this structure does not assume that obligation away.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-5156a640cd85","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3805,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:13"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOAllocation (K : Nat) where","missing":[],"search":"delayedsapoallocation banditrlproof.delayedfeedback.delayedsapoallocation a line-15 allocation together with the exact hypotheses needed to make its coordinates a probability distribution. eap must eventually construct these fields; this structure does not assume that obligation away. structure compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability","label":"probability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability","description":"The Algorithm 5 line-15 probability coordinate associated with a certified allocation.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-00f8bbdd0f7f","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3806,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:26"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def probability {K : Nat} (allocation : DelayedSAPOAllocation K) (i : Fin K) : ℝ","missing":[],"search":"probability banditrlproof.delayedfeedback.delayedsapoallocation.probability the algorithm 5 line-15 probability coordinate associated with a certified allocation. definition compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability_nonnegative","label":"probability_nonnegative","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability_nonnegative","description":"theorem probability_nonnegative {K : Nat} (allocation : DelayedSAPOAllocation K) (i : Fin K) : 0 ≤ allocation.probability i","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-c1a18fb449d7","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3807,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:30"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem probability_nonnegative {K : Nat} (allocation : DelayedSAPOAllocation K) (i : Fin K) : 0 ≤ allocation.probability i","missing":[],"search":"probability_nonnegative banditrlproof.delayedfeedback.delayedsapoallocation.probability_nonnegative theorem probability_nonnegative {k : nat} (allocation : delayedsapoallocation k) (i : fin k) : 0 ≤ allocation.probability i theorem compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.sum_probability_eq_one","label":"sum_probability_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.sum_probability_eq_one","description":"theorem sum_probability_eq_one {K : Nat} (allocation : DelayedSAPOAllocation K) : ∑ i, allocation.probability i = 1","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-4a7d14ed99ad","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3808,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:37"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sum_probability_eq_one {K : Nat} (allocation : DelayedSAPOAllocation K) : ∑ i, allocation.probability i = 1","missing":[],"search":"sum_probability_eq_one banditrlproof.delayedfeedback.delayedsapoallocation.sum_probability_eq_one theorem sum_probability_eq_one {k : nat} (allocation : delayedsapoallocation k) : ∑ i, allocation.probability i = 1 theorem compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.finiteActionDistribution","label":"finiteActionDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.finiteActionDistribution","description":"Reuse the library's finite-action probability-vector interface instead of introducing a second ad hoc sampling semantics for Delayed SAPO.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-f0b6ff989571","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3809,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:45"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem finiteActionDistribution {K : Nat} (allocation : DelayedSAPOAllocation K) : Exp3.FiniteActionDistribution (Finset.univ : Finset (Fin K)) allocation.probability where","missing":[],"search":"finiteactiondistribution banditrlproof.delayedfeedback.delayedsapoallocation.finiteactiondistribution reuse the library's finite-action probability-vector interface instead of introducing a second ad hoc sampling semantics for delayed sapo. theorem compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure","label":"actionMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure","description":"One-round randomized action law induced by the certified line-15 allocation. A measurable history kernel and recursive generated trajectory remain separate obligations.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-1c4ab16e05e2","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3810,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:57"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def actionMeasure {K : Nat} (allocation : DelayedSAPOAllocation K) : Measure (Fin K)","missing":[],"search":"actionmeasure banditrlproof.delayedfeedback.delayedsapoallocation.actionmeasure one-round randomized action law induced by the certified line-15 allocation. a measurable history kernel and recursive generated trajectory remain separate obligations. definition compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure_isProbabilityMeasure","label":"actionMeasure_isProbabilityMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure_isProbabilityMeasure","description":"theorem actionMeasure_isProbabilityMeasure {K : Nat} (allocation : DelayedSAPOAllocation K) : IsProbabilityMeasure allocation.actionMeasure","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-f1b2f04e4dde","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3811,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:61"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionMeasure_isProbabilityMeasure {K : Nat} (allocation : DelayedSAPOAllocation K) : IsProbabilityMeasure allocation.actionMeasure","missing":[],"search":"actionmeasure_isprobabilitymeasure banditrlproof.delayedfeedback.delayedsapoallocation.actionmeasure_isprobabilitymeasure theorem actionmeasure_isprobabilitymeasure {k : nat} (allocation : delayedsapoallocation k) : isprobabilitymeasure allocation.actionmeasure theorem compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule","label":"causalDelayedSAPOActionMeasureRule","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule","description":"Convert a causal allocation rule into a causal, measure-valued action rule. The only input remains `ActionTimeView`; hidden delays and unobserved losses are not added to the algorithm interface.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-e8665e536404","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3812,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:72"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def causalDelayedSAPOActionMeasureRule {K : Nat} {Loss : Type*} (rule : CausalDecisionRule (Fin K) Loss (DelayedSAPOAllocation K)) : CausalDecisionRule (Fin K) Loss (Measure (Fin K))","missing":[],"search":"causaldelayedsapoactionmeasurerule banditrlproof.delayedfeedback.causaldelayedsapoactionmeasurerule convert a causal allocation rule into a causal, measure-valued action rule. the only input remains `actiontimeview`; hidden delays and unobserved losses are not added to the algorithm interface. definition compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_isProbabilityMeasure","label":"causalDelayedSAPOActionMeasureRule_isProbabilityMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_isProbabilityMeasure","description":"Every output of the causal measure-valued rule is a probability measure.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-d1169a8a12eb","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3813,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:79"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem causalDelayedSAPOActionMeasureRule_isProbabilityMeasure {K : Nat} {Loss : Type*} (rule : CausalDecisionRule (Fin K) Loss (DelayedSAPOAllocation K)) (t : Nat) (view : ActionTimeView (Fin K) Loss) : IsProbabilityMeasure (causalDelayedSAPOActionMeasureRule rule t view)","missing":[],"search":"causaldelayedsapoactionmeasurerule_isprobabilitymeasure banditrlproof.delayedfeedback.causaldelayedsapoactionmeasurerule_isprobabilitymeasure every output of the causal measure-valued rule is a probability measure. theorem compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_eq_of_observation_equivalent","label":"causalDelayedSAPOActionMeasureRule_eq_of_observation_equivalent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_eq_of_observation_equivalent","description":"Observation-equivalent hidden worlds induce exactly the same one-round Delayed SAPO action law for every allocation rule typed on the causal view.","url":"../modules/banditrlproof-delayedfeedback-actionlaw/index.html#decl-85b2308b3ae4","parent":"module:BanditRLProof.DelayedFeedback.ActionLaw","order":3814,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActionLaw"],["Source","BanditRLProof/DelayedFeedback/ActionLaw.lean:89"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem causalDelayedSAPOActionMeasureRule_eq_of_observation_equivalent {K : Nat} {Loss : Type*} (rule : CausalDecisionRule (Fin K) Loss (DelayedSAPOAllocation K)) (delay₁ delay₂ : Nat → Nat) (action₁ action₂ : Nat → Fin K) (loss₁ loss₂ : Nat → Loss) (t : Nat) (hvisible : observedBefore delay₁ t = observedBefore delay₂ t) (haction : ∀ s, s < t → action₁ s = action₂ s) (hloss : ∀ s, s ∈ observedBefore delay₁ t → loss₁ s = loss₂ s) : causalDelayedSAPOActionMeasureRule rule t (actionTimeViewAt delay₁ action₁ loss₁ t) = causalDelayedSAPOActionMeasureRule rule t (actionTimeViewAt delay₂ action₂ loss₂ t)","missing":[],"search":"causaldelayedsapoactionmeasurerule_eq_of_observation_equivalent banditrlproof.delayedfeedback.causaldelayedsapoactionmeasurerule_eq_of_observation_equivalent observation-equivalent hidden worlds induce exactly the same one-round delayed sapo action law for every allocation rule typed on the causal view. theorem compiled","shard":"modules/96574241bd97d88f.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.inactiveArms","label":"inactiveArms","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.inactiveArms","description":"Arms outside the current active set in Algorithm 5.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-82a4fbc680bb","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3815,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:11"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def inactiveArms {K : Nat} (active : Finset (Fin K)) : Finset (Fin K)","missing":[],"search":"inactivearms banditrlproof.delayedfeedback.inactivearms arms outside the current active set in algorithm 5. definition compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.activeEqualShare","label":"activeEqualShare","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.activeEqualShare","description":"Algorithm 5 line 15 assigns the residual mass equally to active arms.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-201e9fae84ea","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3816,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:15"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def activeEqualShare {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) : ℝ","missing":[],"search":"activeequalshare banditrlproof.delayedfeedback.activeequalshare algorithm 5 line 15 assigns the residual mass equally to active arms. definition compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability","label":"delayedSAPOProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.delayedSAPOProbability","description":"Full probability vector obtained from externally maintained probabilities on eliminated arms and equal residual mass on active arms.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-6f47f79ef279","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3817,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:22"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","delayed-nonstationary"]],"statement":"noncomputable def delayedSAPOProbability {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) (i : Fin K) : ℝ","missing":[],"search":"delayedsapoprobability banditrlproof.delayedfeedback.delayedsapoprobability full probability vector obtained from externally maintained probabilities on eliminated arms and equal residual mass on active arms. definition compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_active","label":"delayedSAPOProbability_of_active","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_active","description":"theorem delayedSAPOProbability_of_active {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) {i : Fin K} (hi : i ∈ active) : delayedSAPOProbability active inactiveProbability i = activeEqualShare active inactiveProbability","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-5ff2e6c1aeb9","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3818,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:29"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem delayedSAPOProbability_of_active {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) {i : Fin K} (hi : i ∈ active) : delayedSAPOProbability active inactiveProbability i = activeEqualShare active inactiveProbability","missing":[],"search":"delayedsapoprobability_of_active banditrlproof.delayedfeedback.delayedsapoprobability_of_active theorem delayedsapoprobability_of_active {k : nat} (active : finset (fin k)) (inactiveprobability : fin k → ℝ) {i : fin k} (hi : i ∈ active) : delayedsapoprobability active inactiveprobability i = activeequalshare active inactiveprobability theorem compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_inactive","label":"delayedSAPOProbability_of_inactive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_inactive","description":"theorem delayedSAPOProbability_of_inactive {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) {i : Fin K} (hi : i ∉ active) : delayedSAPOProbability active inactiveProbability i = inactiveProbability i","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-c463363d0b6a","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3819,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:37"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem delayedSAPOProbability_of_inactive {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) {i : Fin K} (hi : i ∉ active) : delayedSAPOProbability active inactiveProbability i = inactiveProbability i","missing":[],"search":"delayedsapoprobability_of_inactive banditrlproof.delayedfeedback.delayedsapoprobability_of_inactive theorem delayedsapoprobability_of_inactive {k : nat} (active : finset (fin k)) (inactiveprobability : fin k → ℝ) {i : fin k} (hi : i ∉ active) : delayedsapoprobability active inactiveprobability i = inactiveprobability i theorem compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.activeEqualShare_nonneg","label":"activeEqualShare_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.activeEqualShare_nonneg","description":"The residual active-arm share is nonnegative when eliminated-arm mass does not exceed one.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-c8c7e9fb3f46","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3820,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:46"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem activeEqualShare_nonneg {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) (hmass : (∑ i ∈ inactiveArms active, inactiveProbability i) ≤ 1) : 0 ≤ activeEqualShare active inactiveProbability","missing":[],"search":"activeequalshare_nonneg banditrlproof.delayedfeedback.activeequalshare_nonneg the residual active-arm share is nonnegative when eliminated-arm mass does not exceed one. theorem compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_nonneg","label":"delayedSAPOProbability_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.delayedSAPOProbability_nonneg","description":"Every coordinate of the line-15 allocation is nonnegative under the source-side residual-mass and inactive-coordinate hypotheses.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-72095bcff657","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3821,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:54"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem delayedSAPOProbability_nonneg {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) (hinactive : ∀ i ∈ inactiveArms active, 0 ≤ inactiveProbability i) (hmass : (∑ i ∈ inactiveArms active, inactiveProbability i) ≤ 1) (i : Fin K) : 0 ≤ delayedSAPOProbability active inactiveProbability i","missing":[],"search":"delayedsapoprobability_nonneg banditrlproof.delayedfeedback.delayedsapoprobability_nonneg every coordinate of the line-15 allocation is nonnegative under the source-side residual-mass and inactive-coordinate hypotheses. theorem compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sum_delayedSAPOProbability_eq_one","label":"sum_delayedSAPOProbability_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sum_delayedSAPOProbability_eq_one","description":"With at least one active arm, the equal-residual allocation has total mass exactly one.","url":"../modules/banditrlproof-delayedfeedback-activeallocation/index.html#decl-36eb003e08d9","parent":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","order":3822,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ActiveAllocation"],["Source","BanditRLProof/DelayedFeedback/ActiveAllocation.lean:68"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sum_delayedSAPOProbability_eq_one {K : Nat} (active : Finset (Fin K)) (inactiveProbability : Fin K → ℝ) (hactive : active.Nonempty) : ∑ i, delayedSAPOProbability active inactiveProbability i = 1","missing":[],"search":"sum_delayedsapoprobability_eq_one banditrlproof.delayedfeedback.sum_delayedsapoprobability_eq_one with at least one active arm, the equal-residual allocation has total mass exactly one. theorem compiled","shard":"modules/f0f22e9a6d0b345e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.ActionTimeView","label":"ActionTimeView","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.ActionTimeView","description":"The information exposed to an action rule immediately before a delayed bandit action. Past actions are visible, while a loss is exposed only when its source round belongs to the source-faithful strict-availability set. The view deliberately has no delay field and no total loss trace field. A future Delayed SAPO implementation must consume this view (or a proved equivalent), rather than the environment's hidden delay…","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-fb3db054613e","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3823,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:16"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure ActionTimeView (Action : Type uAction) (Loss : Type uLoss) where","missing":[],"search":"actiontimeview banditrlproof.delayedfeedback.actiontimeview the information exposed to an action rule immediately before a delayed bandit action. past actions are visible, while a loss is exposed only when its source round belongs to the source-faithful strict-availability set. the view deliberately has no delay field and no total loss trace field. a future delayed sapo implementation must consume this view (or a proved equivalent), rather than the environment's hidden delay/loss functions. structure compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.ActionTimeView.ext","label":"ext","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.ActionTimeView.ext","description":"theorem ActionTimeView.ext {Action : Type uAction} {Loss : Type uLoss} {left right : ActionTimeView Action Loss} (hpast : left.pastAction = right.pastAction) (hloss : left.observedLoss = right.observedLoss) : left = right","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-204a97806945","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3824,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:21"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ActionTimeView.ext {Action : Type uAction} {Loss : Type uLoss} {left right : ActionTimeView Action Loss} (hpast : left.pastAction = right.pastAction) (hloss : left.observedLoss = right.observedLoss) : left = right","missing":[],"search":"ext banditrlproof.delayedfeedback.actiontimeview.ext theorem actiontimeview.ext {action : type uaction} {loss : type uloss} {left right : actiontimeview action loss} (hpast : left.pastaction = right.pastaction) (hloss : left.observedloss = right.observedloss) : left = right theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt","label":"actionTimeViewAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt","description":"Construct the pre-action view at round `t` from an environment trace. Actions before `t` are known to the learner. A source loss is known exactly when `s + delay s < t`; future and outstanding losses return `none`.","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-55cbb54dc125","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3825,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:36"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","delayed-nonstationary"]],"statement":"def actionTimeViewAt {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) (t : Nat) : ActionTimeView Action Loss where","missing":[],"search":"actiontimeviewat banditrlproof.delayedfeedback.actiontimeviewat construct the pre-action view at round `t` from an environment trace. actions before `t` are known to the learner. a source loss is known exactly when `s + delay s < t`; future and outstanding losses return `none`. definition compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.DelayedFeedback.CausalDecisionRule","label":"CausalDecisionRule","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.CausalDecisionRule","description":"A causal decision rule receives only the round number and its action-time view. `Decision` can later be instantiated by an action distribution, a kernel, or a deterministic action.","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-02e8756b639f","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3826,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:47"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"abbrev CausalDecisionRule (Action : Type uAction) (Loss : Type uLoss) (Decision : Type uDecision)","missing":[],"search":"causaldecisionrule banditrlproof.delayedfeedback.causaldecisionrule a causal decision rule receives only the round number and its action-time view. `decision` can later be instantiated by an action distribution, a kernel, or a deterministic action. abbreviation compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_lt","label":"actionTimeViewAt_pastAction_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_lt","description":"theorem actionTimeViewAt_pastAction_of_lt {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s < t) : (actionTimeViewAt delay action loss t).pastAction s = some (action s)","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-479a6bab941b","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3827,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:52"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionTimeViewAt_pastAction_of_lt {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s < t) : (actionTimeViewAt delay action loss t).pastAction s = some (action s)","missing":[],"search":"actiontimeviewat_pastaction_of_lt banditrlproof.delayedfeedback.actiontimeviewat_pastaction_of_lt theorem actiontimeviewat_pastaction_of_lt {action : type uaction} {loss : type uloss} (delay : nat → nat) (action : nat → action) (loss : nat → loss) {s t : nat} (hs : s < t) : (actiontimeviewat delay action loss t).pastaction s = some (action s) theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_not_lt","label":"actionTimeViewAt_pastAction_of_not_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_not_lt","description":"theorem actionTimeViewAt_pastAction_of_not_lt {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : ¬ s < t) : (actionTimeViewAt delay action loss t).pastAction s = none","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-3fd92d1a45e4","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3828,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:60"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionTimeViewAt_pastAction_of_not_lt {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : ¬ s < t) : (actionTimeViewAt delay action loss t).pastAction s = none","missing":[],"search":"actiontimeviewat_pastaction_of_not_lt banditrlproof.delayedfeedback.actiontimeviewat_pastaction_of_not_lt theorem actiontimeviewat_pastaction_of_not_lt {action : type uaction} {loss : type uloss} (delay : nat → nat) (action : nat → action) (loss : nat → loss) {s t : nat} (hs : ¬ s < t) : (actiontimeviewat delay action loss t).pastaction s = none theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_mem","label":"actionTimeViewAt_observedLoss_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_mem","description":"theorem actionTimeViewAt_observedLoss_of_mem {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s ∈ observedBefore delay t) : (actionTimeViewAt delay action loss t).observedLoss s = some (loss s)","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-f77b976f0cba","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3829,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:68"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionTimeViewAt_observedLoss_of_mem {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s ∈ observedBefore delay t) : (actionTimeViewAt delay action loss t).observedLoss s = some (loss s)","missing":[],"search":"actiontimeviewat_observedloss_of_mem banditrlproof.delayedfeedback.actiontimeviewat_observedloss_of_mem theorem actiontimeviewat_observedloss_of_mem {action : type uaction} {loss : type uloss} (delay : nat → nat) (action : nat → action) (loss : nat → loss) {s t : nat} (hs : s ∈ observedbefore delay t) : (actiontimeviewat delay action loss t).observedloss s = some (loss s) theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_not_mem","label":"actionTimeViewAt_observedLoss_of_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_not_mem","description":"theorem actionTimeViewAt_observedLoss_of_not_mem {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s ∉ observedBefore delay t) : (actionTimeViewAt delay action loss t).observedLoss s = none","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-97518353070a","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3830,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:76"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionTimeViewAt_observedLoss_of_not_mem {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s ∉ observedBefore delay t) : (actionTimeViewAt delay action loss t).observedLoss s = none","missing":[],"search":"actiontimeviewat_observedloss_of_not_mem banditrlproof.delayedfeedback.actiontimeviewat_observedloss_of_not_mem theorem actiontimeviewat_observedloss_of_not_mem {action : type uaction} {loss : type uloss} (delay : nat → nat) (action : nat → action) (loss : nat → loss) {s t : nat} (hs : s ∉ observedbefore delay t) : (actiontimeviewat delay action loss t).observedloss s = none theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_outstanding_loss_hidden","label":"actionTimeViewAt_outstanding_loss_hidden","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt_outstanding_loss_hidden","description":"Outstanding feedback is absent from the action-time view.","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-e156acd99278","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3831,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:84"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionTimeViewAt_outstanding_loss_hidden {Action : Type uAction} {Loss : Type uLoss} (delay : Nat → Nat) (action : Nat → Action) (loss : Nat → Loss) {s t : Nat} (hs : s ∈ outstandingAt delay t) : (actionTimeViewAt delay action loss t).observedLoss s = none","missing":[],"search":"actiontimeviewat_outstanding_loss_hidden banditrlproof.delayedfeedback.actiontimeviewat_outstanding_loss_hidden outstanding feedback is absent from the action-time view. theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_eq_of_observation_equivalent","label":"actionTimeViewAt_eq_of_observation_equivalent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.actionTimeViewAt_eq_of_observation_equivalent","description":"Two environment traces yield exactly the same pre-action view whenever their visible source sets, past actions, and revealed losses agree. Hidden delays and unobserved losses may differ.","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-6c138a4819c8","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3832,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:96"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem actionTimeViewAt_eq_of_observation_equivalent {Action : Type uAction} {Loss : Type uLoss} (delay₁ delay₂ : Nat → Nat) (action₁ action₂ : Nat → Action) (loss₁ loss₂ : Nat → Loss) (t : Nat) (hvisible : observedBefore delay₁ t = observedBefore delay₂ t) (haction : ∀ s, s < t → action₁ s = action₂ s) (hloss : ∀ s, s ∈ observedBefore delay₁ t → loss₁ s = loss₂ s) : actionTimeViewAt delay₁ action₁ loss₁ t = actionTimeViewAt delay₂ action₂ loss₂ t","missing":[],"search":"actiontimeviewat_eq_of_observation_equivalent banditrlproof.delayedfeedback.actiontimeviewat_eq_of_observation_equivalent two environment traces yield exactly the same pre-action view whenever their visible source sets, past actions, and revealed losses agree. hidden delays and unobserved losses may differ. theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.causalDecision_eq_of_observation_equivalent","label":"causalDecision_eq_of_observation_equivalent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.causalDecision_eq_of_observation_equivalent","description":"A decision rule consuming only `ActionTimeView` cannot distinguish two worlds that agree on all information visible before the action.","url":"../modules/banditrlproof-delayedfeedback-causalview/index.html#decl-59a9887cb4e2","parent":"module:BanditRLProof.DelayedFeedback.CausalView","order":3833,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.CausalView"],["Source","BanditRLProof/DelayedFeedback/CausalView.lean:122"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem causalDecision_eq_of_observation_equivalent {Action : Type uAction} {Loss : Type uLoss} {Decision : Type uDecision} (rule : CausalDecisionRule Action Loss Decision) (delay₁ delay₂ : Nat → Nat) (action₁ action₂ : Nat → Action) (loss₁ loss₂ : Nat → Loss) (t : Nat) (hvisible : observedBefore delay₁ t = observedBefore delay₂ t) (haction : ∀ s, s < t → action₁ s = action₂ s) (hloss : ∀ s, s ∈ observedBefore delay₁ t → loss₁ s = loss₂ s) : rule t (actionTimeViewAt delay₁ action₁ loss₁ t) = rule t (actionTimeViewAt delay₂ action₂ loss₂ t)","missing":[],"search":"causaldecision_eq_of_observation_equivalent banditrlproof.delayedfeedback.causaldecision_eq_of_observation_equivalent a decision rule consuming only `actiontimeview` cannot distinguish two worlds that agree on all information visible before the action. theorem compiled","shard":"modules/e941e09c0a7ab1e9.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOInitialEliminatedProbability","label":"delayedSAPOInitialEliminatedProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.delayedSAPOInitialEliminatedProbability","description":"Algorithm 5 line 10's initial sampling probability `p_i^1 = 1/(2K) + n_i(S)/(2T)`.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-ac0ed65bef79","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3834,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:27"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def delayedSAPOInitialEliminatedProbability (armCount horizon pullCount : Nat) : Real","missing":[],"search":"delayedsapoinitialeliminatedprobability banditrlproof.delayedfeedback.delayedsapoinitialeliminatedprobability algorithm 5 line 10's initial sampling probability `p_i^1 = 1/(2k) + n_i(s)/(2t)`. definition compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOInitialPhaseTarget","label":"delayedSAPOInitialPhaseTarget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.delayedSAPOInitialPhaseTarget","description":"Algorithm 5 line 10's first EAP phase target `N_i^1 = 1280/(p_i^1 * (Delta-tilde_i)^2)`.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-cca63e7bd8c3","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3835,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:34"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def delayedSAPOInitialPhaseTarget (probability surrogateGap : Real) : Real","missing":[],"search":"delayedsapoinitialphasetarget banditrlproof.delayedfeedback.delayedsapoinitialphasetarget algorithm 5 line 10's first eap phase target `n_i^1 = 1280/(p_i^1 * (delta-tilde_i)^2)`. definition compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_pos","label":"sourceEmpiricalWidthScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_pos","description":"A positive source scale makes the capped inverse-square-root width strictly positive, including the zero-count branch where the width is one.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-4acbb8615a2d","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3836,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:40"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_pos (scale count : Real) (hscale : 0 < scale) : 0 < sourceEmpiricalWidthScale scale count","missing":[],"search":"sourceempiricalwidthscale_pos banditrlproof.delayedfeedback.sourceempiricalwidthscale_pos a positive source scale makes the capped inverse-square-root width strictly positive, including the zero-count branch where the width is one. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization","label":"DelayedSAPOEliminatedArmInitialization","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization","description":"Per-arm data created by Algorithm 5 line 10. `processedAtProbabilityLevel j` represents the source sets `C_i^(p_i^1 * 2^(-j))`; line 10 initializes every represented level to the empty set. Later EAP calls and their phase transitions remain a separate state-machine leaf.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-2e6f04468698","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3837,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:55"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOEliminatedArmInitialization (K : Nat) where","missing":[],"search":"delayedsapoeliminatedarminitialization banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization per-arm data created by algorithm 5 line 10. `processedatprobabilitylevel j` represents the source sets `c_i^(p_i^1 * 2^(-j))`; line 10 initializes every represented level to the empty set. later eap calls and their phase transitions remain a separate state-machine leaf. structure compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmBank","label":"DelayedSAPOEliminatedArmBank","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmBank","description":"Registry of EAP data. Active arms have not yet entered EAP and are therefore represented by `none`.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-620073522fe6","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3838,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:70"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"abbrev DelayedSAPOEliminatedArmBank (K : Nat)","missing":[],"search":"delayedsapoeliminatedarmbank banditrlproof.delayedfeedback.delayedsapoeliminatedarmbank registry of eap data. active arms have not yet entered eap and are therefore represented by `none`. abbreviation compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.ActiveArmsUninitialized","label":"ActiveArmsUninitialized","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.ActiveArmsUninitialized","description":"Round-start invariant needed to justify that line 10 initializes rather than resets every arm in the newly eliminated set.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-58cf8bec3ec0","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3839,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:75"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def ActiveArmsUninitialized {K : Nat} (state : DelayedSAPOStructuralRoundState K) (bank : DelayedSAPOEliminatedArmBank K) : Prop","missing":[],"search":"activearmsuninitialized banditrlproof.delayedfeedback.activearmsuninitialized round-start invariant needed to justify that line 10 initializes rather than resets every arm in the newly eliminated set. definition compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne","label":"ofProcessOne","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne","description":"Construct exactly the line-10 data for one candidate arm from the post-line-4 numerical snapshot. Membership in the line-7 elimination set is enforced by `initializeIfEliminated` below.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-d7ec9b9f25ae","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3840,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:85"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def ofProcessOne {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : DelayedSAPOEliminatedArmInitialization K","missing":[],"search":"ofprocessone banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone construct exactly the line-10 data for one candidate arm from the post-line-4 numerical snapshot. membership in the line-7 elimination set is enforced by `initializeifeliminated` below. definition compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated","label":"initializeIfEliminated","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated","description":"Algorithm 5 lines 9--10 initialize precisely the arms selected by the line-7 elimination set. Returning `Option` keeps that domain visible instead of silently producing EAP state for an arm that remains active.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-2cb7ae0e9173","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3841,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:112"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def initializeIfEliminated {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : Option (DelayedSAPOEliminatedArmInitialization K)","missing":[],"search":"initializeifeliminated banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializeifeliminated algorithm 5 lines 9--10 initialize precisely the arms selected by the line-7 elimination set. returning `option` keeps that domain visible instead of silently producing eap state for an arm that remains active. definition compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated","label":"initializeNewlyEliminated","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated","description":"Pointwise line-10 update of an existing EAP registry. Arms in the new line-7 elimination set receive the literal source initializer; every other arm keeps its prior state.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-03259dcf0b87","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3842,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:124"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def initializeNewlyEliminated {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : DelayedSAPOEliminatedArmBank K) : DelayedSAPOEliminatedArmBank K","missing":[],"search":"initializenewlyeliminated banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializenewlyeliminated pointwise line-10 update of an existing eap registry. arms in the new line-7 elimination set receive the literal source initializer; every other arm keeps its prior state. definition compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_some_iff","label":"initializeIfEliminated_eq_some_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_some_iff","description":"theorem initializeIfEliminated_eq_some_iff {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : initializeIfEliminated step horizon i = some (ofProcessOne step horizon i) <-> i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-fbd6f0aab817","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3843,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:136"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initializeIfEliminated_eq_some_iff {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : initializeIfEliminated step horizon i = some (ofProcessOne step horizon i) <-> i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated","missing":[],"search":"initializeifeliminated_eq_some_iff banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializeifeliminated_eq_some_iff theorem initializeifeliminated_eq_some_iff {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : initializeifeliminated step horizon i = some (ofprocessone step horizon i) <-> i ∈ (step.topreeliminationsummary.toconfidencesnapshot horizon).eliminated theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_none_iff","label":"initializeIfEliminated_eq_none_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_none_iff","description":"theorem initializeIfEliminated_eq_none_iff {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : initializeIfEliminated step horizon i = none <-> i ∉ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-7184bd119485","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3844,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:159"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initializeIfEliminated_eq_none_iff {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : initializeIfEliminated step horizon i = none <-> i ∉ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated","missing":[],"search":"initializeifeliminated_eq_none_iff banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializeifeliminated_eq_none_iff theorem initializeifeliminated_eq_none_iff {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : initializeifeliminated step horizon i = none <-> i ∉ (step.topreeliminationsummary.toconfidencesnapshot horizon).eliminated theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_mem","label":"initializeNewlyEliminated_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_mem","description":"theorem initializeNewlyEliminated_of_mem {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hi : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : initializeNewlyEliminated step horizon prior i = some (ofProcessOne step horizon i)","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-3e214586b15a","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3845,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:181"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initializeNewlyEliminated_of_mem {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hi : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : initializeNewlyEliminated step horizon prior i = some (ofProcessOne step horizon i)","missing":[],"search":"initializenewlyeliminated_of_mem banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializenewlyeliminated_of_mem theorem initializenewlyeliminated_of_mem {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (prior : fin k -> option (delayedsapoeliminatedarminitialization k)) (i : fin k) (hi : i ∈ (step.topreeliminationsummary.toconfidencesnapshot horizon).eliminated) : initializenewlyeliminated step horizon prior i = some (ofprocessone step horizon i) theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_not_mem","label":"initializeNewlyEliminated_of_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_not_mem","description":"theorem initializeNewlyEliminated_of_not_mem {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hi : i ∉ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : initializeNewlyEliminated step horizon prior i = prior i","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-0e5b7eb0bc57","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3846,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:194"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initializeNewlyEliminated_of_not_mem {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hi : i ∉ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : initializeNewlyEliminated step horizon prior i = prior i","missing":[],"search":"initializenewlyeliminated_of_not_mem banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializenewlyeliminated_of_not_mem theorem initializenewlyeliminated_of_not_mem {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (prior : fin k -> option (delayedsapoeliminatedarminitialization k)) (i : fin k) (hi : i ∉ (step.topreeliminationsummary.toconfidencesnapshot horizon).eliminated) : initializenewlyeliminated step horizon prior i = prior i theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.mem_eliminated_of_initializeNewlyEliminated_ne_prior","label":"mem_eliminated_of_initializeNewlyEliminated_ne_prior","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.mem_eliminated_of_initializeNewlyEliminated_ne_prior","description":"Line 10 can change an arm's EAP-bank entry only when line 7 selected that arm for elimination in the same snapshot.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-3d6fa0fb47d5","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3847,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:207"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_eliminated_of_initializeNewlyEliminated_ne_prior {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hchanged : initializeNewlyEliminated step horizon prior i ≠ prior i) : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated","missing":[],"search":"mem_eliminated_of_initializenewlyeliminated_ne_prior banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.mem_eliminated_of_initializenewlyeliminated_ne_prior line 10 can change an arm's eap-bank entry only when line 7 selected that arm for elimination in the same snapshot. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_eq_prior_of_mem_remainingActive","label":"initializeNewlyEliminated_eq_prior_of_mem_remainingActive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_eq_prior_of_mem_remainingActive","description":"An arm that remains active after line 8 is not reset by line 10.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-57eb80c3be45","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3848,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:220"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initializeNewlyEliminated_eq_prior_of_mem_remainingActive {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hi : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).remainingActive) : initializeNewlyEliminated step horizon prior i = prior i","missing":[],"search":"initializenewlyeliminated_eq_prior_of_mem_remainingactive banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializenewlyeliminated_eq_prior_of_mem_remainingactive an arm that remains active after line 8 is not reset by line 10. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.prior_eq_none_of_mem_eliminated","label":"prior_eq_none_of_mem_eliminated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.prior_eq_none_of_mem_eliminated","description":"Every arm selected by line 7 was active immediately before line 8, so a valid prior bank has no EAP state for it yet.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-c2077011b3c3","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3849,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:236"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem prior_eq_none_of_mem_eliminated {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : DelayedSAPOEliminatedArmBank K) (hprior : ActiveArmsUninitialized state prior) (i : Fin K) (hi : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : prior i = none","missing":[],"search":"prior_eq_none_of_mem_eliminated banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.prior_eq_none_of_mem_eliminated every arm selected by line 7 was active immediately before line 8, so a valid prior bank has no eap state for it yet. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.remainingActive_uninitialized_after_initialize","label":"remainingActive_uninitialized_after_initialize","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.remainingActive_uninitialized_after_initialize","description":"After line 10, every arm that survives line 8 is still uninitialized in the EAP bank.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-c7d435b016ec","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3850,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:254"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem remainingActive_uninitialized_after_initialize {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (prior : DelayedSAPOEliminatedArmBank K) (hprior : ActiveArmsUninitialized state prior) (i : Fin K) (hi : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).remainingActive) : initializeNewlyEliminated step horizon prior i = none","missing":[],"search":"remainingactive_uninitialized_after_initialize banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.remainingactive_uninitialized_after_initialize after line 10, every arm that survives line 8 is still uninitialized in the eap bank. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_arm","label":"ofProcessOne_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_arm","description":"theorem ofProcessOne_arm {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).arm = i","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-d8134f9540de","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3851,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:273"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_arm {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).arm = i","missing":[],"search":"ofprocessone_arm banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_arm theorem ofprocessone_arm {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : (ofprocessone step horizon i).arm = i theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationRound","label":"ofProcessOne_eliminationRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationRound","description":"theorem ofProcessOne_eliminationRound {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).eliminationRound = state.currentActionRound","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-66e5eef919a8","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3852,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:280"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_eliminationRound {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).eliminationRound = state.currentActionRound","missing":[],"search":"ofprocessone_eliminationround banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_eliminationround theorem ofprocessone_eliminationround {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : (ofprocessone step horizon i).eliminationround = state.currentactionround theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationProcessedOrder","label":"ofProcessOne_eliminationProcessedOrder","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationProcessedOrder","description":"theorem ofProcessOne_eliminationProcessedOrder {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).eliminationProcessedOrder = step.extendedOrder","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-815fee939577","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3853,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:288"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_eliminationProcessedOrder {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).eliminationProcessedOrder = step.extendedOrder","missing":[],"search":"ofprocessone_eliminationprocessedorder banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_eliminationprocessedorder theorem ofprocessone_eliminationprocessedorder {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : (ofprocessone step horizon i).eliminationprocessedorder = step.extendedorder theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_errorCount","label":"ofProcessOne_errorCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_errorCount","description":"theorem ofProcessOne_errorCount {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).errorCount = 0","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-1d8d7153aec1","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3854,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:296"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_errorCount {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).errorCount = 0","missing":[],"search":"ofprocessone_errorcount banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_errorcount theorem ofprocessone_errorcount {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : (ofprocessone step horizon i).errorcount = 0 theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseIndex","label":"ofProcessOne_phaseIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseIndex","description":"theorem ofProcessOne_phaseIndex {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).phaseIndex = 1","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-7c65f757ebb7","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3855,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:303"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_phaseIndex {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).phaseIndex = 1","missing":[],"search":"ofprocessone_phaseindex banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_phaseindex theorem ofprocessone_phaseindex {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : (ofprocessone step horizon i).phaseindex = 1 theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseSamples","label":"ofProcessOne_phaseSamples","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseSamples","description":"theorem ofProcessOne_phaseSamples {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).phaseSamples = []","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-3a8263a1f4f3","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3856,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:310"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_phaseSamples {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : (ofProcessOne step horizon i).phaseSamples = []","missing":[],"search":"ofprocessone_phasesamples banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_phasesamples theorem ofprocessone_phasesamples {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) : (ofprocessone step horizon i).phasesamples = [] theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_processedAtProbabilityLevel","label":"ofProcessOne_processedAtProbabilityLevel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_processedAtProbabilityLevel","description":"theorem ofProcessOne_processedAtProbabilityLevel {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) (j : Nat) : (ofProcessOne step horizon i).processedAtProbabilityLevel j = {}","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-86bfab77ae7c","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3857,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:317"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ofProcessOne_processedAtProbabilityLevel {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) (j : Nat) : (ofProcessOne step horizon i).processedAtProbabilityLevel j = {}","missing":[],"search":"ofprocessone_processedatprobabilitylevel banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.ofprocessone_processedatprobabilitylevel theorem ofprocessone_processedatprobabilitylevel {k : nat} {state : delayedsapostructuralroundstate k} (step : delayedsaponoswitchprocessone state) (horizon : nat) (i : fin k) (j : nat) : (ofprocessone step horizon i).processedatprobabilitylevel j = {} theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_pos","label":"initialProbability_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_pos","description":"The source initial eliminated-arm probability is strictly positive when there is an arm and the horizon is positive. No count upper bound is needed for this lower endpoint.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-3b3f677a88e1","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3858,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:326"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initialProbability_pos {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : 0 < (ofProcessOne step horizon i).initialProbability","missing":[],"search":"initialprobability_pos banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initialprobability_pos the source initial eliminated-arm probability is strictly positive when there is an arm and the horizon is positive. no count upper bound is needed for this lower endpoint. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_le_one","label":"initialProbability_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_le_one","description":"With the source-side count bound `n_i(S) <= T`, the initial probability is at most one. The count bound will be generated by the full Algorithm-5 trajectory; it is not manufactured by this numerical initializer.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-f68fc90cee29","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3859,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:347"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initialProbability_le_one {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) (hhorizon : 0 < horizon) (hcount : step.toPreEliminationSummary.toProcessedPrefix.processedPullCount i <= horizon) : (ofProcessOne step horizon i).initialProbability <= 1","missing":[],"search":"initialprobability_le_one banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initialprobability_le_one with the source-side count bound `n_i(s) <= t`, the initial probability is at most one. the count bound will be generated by the full algorithm-5 trajectory; it is not manufactured by this numerical initializer. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_nonneg","label":"surrogateGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_nonneg","description":"The line-10 surrogate gap is nonnegative because the exact source width is capped below by zero.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-7c74b1065972","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3860,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:378"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem surrogateGap_nonneg {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : 0 <= (ofProcessOne step horizon i).surrogateGap","missing":[],"search":"surrogategap_nonneg banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.surrogategap_nonneg the line-10 surrogate gap is nonnegative because the exact source width is capped below by zero. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_pos","label":"surrogateGap_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_pos","description":"For the nontrivial source regime `1 < T`, the frozen line-10 surrogate gap is strictly positive for every processed count.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-308e4d169236","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3861,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:390"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem surrogateGap_pos {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) (hhorizon : 1 < horizon) : 0 < (ofProcessOne step horizon i).surrogateGap","missing":[],"search":"surrogategap_pos banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.surrogategap_pos for the nontrivial source regime `1 < t`, the frozen line-10 surrogate gap is strictly positive for every processed count. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_nonneg","label":"initialPhaseTarget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_nonneg","description":"The first EAP phase target is nonnegative without hiding Lean's totalized-division boundary.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-b5dcbf2459b1","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3862,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:408"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initialPhaseTarget_nonneg {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) : 0 <= (ofProcessOne step horizon i).initialPhaseTarget","missing":[],"search":"initialphasetarget_nonneg banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initialphasetarget_nonneg the first eap phase target is nonnegative without hiding lean's totalized-division boundary. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_pos","label":"initialPhaseTarget_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_pos","description":"Under the explicit positive-width condition used by the source analysis, the first EAP phase target has a genuinely positive denominator and value.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-62ba8c1d5893","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3863,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:421"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem initialPhaseTarget_pos {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (i : Fin K) (hhorizon : 1 < horizon) : 0 < (ofProcessOne step horizon i).initialPhaseTarget","missing":[],"search":"initialphasetarget_pos banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initialphasetarget_pos under the explicit positive-width condition used by the source analysis, the first eap phase target has a genuinely positive denominator and value. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_spec_of_mem","label":"initializeNewlyEliminated_spec_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_spec_of_mem","description":"Bundled deterministic producer for the next EAP leaf: a newly eliminated arm receives a source-exact state with positive sampling probability, surrogate gap, and first phase target in the nontrivial horizon regime.","url":"../modules/banditrlproof-delayedfeedback-eliminatedarminitialization/index.html#decl-07add3be2dc8","parent":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","order":3864,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.EliminatedArmInitialization"],["Source","BanditRLProof/DelayedFeedback/EliminatedArmInitialization.lean:436"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","delayed-nonstationary"]],"statement":"theorem initializeNewlyEliminated_spec_of_mem {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) (hhorizon : 1 < horizon) (prior : Fin K -> Option (DelayedSAPOEliminatedArmInitialization K)) (i : Fin K) (hi : i ∈ (step.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : initializeNewlyEliminated step horizon prior i = some (ofProcessOne step horizon i) ∧ 0 < (ofProcessOne step horizon i).initialProbability ∧ 0 < (ofProcessOne step horizon i).surrogateGap ∧ 0 < (ofProcessOne step horizon i).initialPhaseTarget","missing":[],"search":"initializenewlyeliminated_spec_of_mem banditrlproof.delayedfeedback.delayedsapoeliminatedarminitialization.initializenewlyeliminated_spec_of_mem bundled deterministic producer for the next eap leaf: a newly eliminated arm receives a source-exact state with positive sampling probability, surrogate gap, and first phase target in the nontrivial horizon regime. theorem compiled","shard":"modules/93a392d1ccd0bdca.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot","label":"DelayedSAPOEliminationSnapshot","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot","description":"Inputs read by Algorithm 5 line 7 when it processes one newly available feedback item. This is deliberately an elimination snapshot rather than a claim that the full Delayed SAPO state machine has already been implemented.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-e52d922b0020","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3865,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:12"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOEliminationSnapshot (K : Nat) where","missing":[],"search":"delayedsapoeliminationsnapshot banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot inputs read by algorithm 5 line 7 when it processes one newly available feedback item. this is deliberately an elimination snapshot rather than a claim that the full delayed sapo state machine has already been implemented. structure compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.eliminated","label":"eliminated","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.eliminated","description":"Arms selected by the strict elimination test in Algorithm 5 line 7.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-52cce7356cb9","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3866,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:21"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def eliminated {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) : Finset (Fin K)","missing":[],"search":"eliminated banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.eliminated arms selected by the strict elimination test in algorithm 5 line 7. definition compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive","label":"remainingActive","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive","description":"Active set after Algorithm 5 line 8 removes every arm selected by line 7.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-1e1324ec2ffa","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3867,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:29"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def remainingActive {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) : Finset (Fin K)","missing":[],"search":"remainingactive banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.remainingactive active set after algorithm 5 line 8 removes every arm selected by line 7. definition compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_eliminated_iff","label":"mem_eliminated_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_eliminated_iff","description":"theorem mem_eliminated_iff {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (i : Fin K) : i ∈ snapshot.eliminated ↔ i ∈ snapshot.active ∧ snapshot.ucbStar < snapshot.empiricalMean i - (9 : ℝ) * snapshot.empiricalWidth i","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-0fb46344d866","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3868,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:35"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_eliminated_iff {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (i : Fin K) : i ∈ snapshot.eliminated ↔ i ∈ snapshot.active ∧ snapshot.ucbStar < snapshot.empiricalMean i - (9 : ℝ) * snapshot.empiricalWidth i","missing":[],"search":"mem_eliminated_iff banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.mem_eliminated_iff theorem mem_eliminated_iff {k : nat} (snapshot : delayedsapoeliminationsnapshot k) (i : fin k) : i ∈ snapshot.eliminated ↔ i ∈ snapshot.active ∧ snapshot.ucbstar < snapshot.empiricalmean i - (9 : ℝ) * snapshot.empiricalwidth i theorem compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_remainingActive_iff","label":"mem_remainingActive_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_remainingActive_iff","description":"theorem mem_remainingActive_iff {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (i : Fin K) : i ∈ snapshot.remainingActive ↔ i ∈ snapshot.active ∧ snapshot.empiricalMean i - (9 : ℝ) * snapshot.empiricalWidth i ≤ snapshot.ucbStar","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-f3f299ca68b6","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3869,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:45"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_remainingActive_iff {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (i : Fin K) : i ∈ snapshot.remainingActive ↔ i ∈ snapshot.active ∧ snapshot.empiricalMean i - (9 : ℝ) * snapshot.empiricalWidth i ≤ snapshot.ucbStar","missing":[],"search":"mem_remainingactive_iff banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.mem_remainingactive_iff theorem mem_remainingactive_iff {k : nat} (snapshot : delayedsapoeliminationsnapshot k) (i : fin k) : i ∈ snapshot.remainingactive ↔ i ∈ snapshot.active ∧ snapshot.empiricalmean i - (9 : ℝ) * snapshot.empiricalwidth i ≤ snapshot.ucbstar theorem compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.OptimalArmSurvivalCertificate","label":"OptimalArmSurvivalCertificate","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.OptimalArmSurvivalCertificate","description":"The two source conditions used in the deterministic core of Lemma D.9: the empirical mean of the optimal arm lies in its good-event confidence interval, and the source's `ucbStar` remains an upper certificate for the optimal mean. The event probability establishing these fields is a separate concentration obligation.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-ad6b2dbf9bc4","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3870,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:61"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure OptimalArmSurvivalCertificate {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) : Prop where","missing":[],"search":"optimalarmsurvivalcertificate banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.optimalarmsurvivalcertificate the two source conditions used in the deterministic core of lemma d.9: the empirical mean of the optimal arm lies in its good-event confidence interval, and the source's `ucbstar` remains an upper certificate for the optimal mean. the event probability establishing these fields is a separate concentration obligation. structure compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.optimal_mem_remainingActive_of_certificate","label":"optimal_mem_remainingActive_of_certificate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.optimal_mem_remainingActive_of_certificate","description":"Deterministic core of source Lemma D.9: on the relevant stochastic good-event projections, Algorithm 5's strict line-7 test cannot eliminate the certified optimal arm.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-83947802ca0a","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3871,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:74"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem optimal_mem_remainingActive_of_certificate {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (certificate : OptimalArmSurvivalCertificate snapshot mean optimal) : optimal ∈ snapshot.remainingActive","missing":[],"search":"optimal_mem_remainingactive_of_certificate banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.optimal_mem_remainingactive_of_certificate deterministic core of source lemma d.9: on the relevant stochastic good-event projections, algorithm 5's strict line-7 test cannot eliminate the certified optimal arm. theorem compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive_nonempty_of_certificate","label":"remainingActive_nonempty_of_certificate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive_nonempty_of_certificate","description":"Consequently the post-elimination active set is nonempty. This closes the exact nonemptiness premise required by Algorithm 5 line 15's residual allocation.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-d9078494a923","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3872,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:92"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem remainingActive_nonempty_of_certificate {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (certificate : OptimalArmSurvivalCertificate snapshot mean optimal) : snapshot.remainingActive.Nonempty","missing":[],"search":"remainingactive_nonempty_of_certificate banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.remainingactive_nonempty_of_certificate consequently the post-elimination active set is nonempty. this closes the exact nonemptiness premise required by algorithm 5 line 15's residual allocation. theorem compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.sum_delayedSAPOProbability_after_elimination_eq_one","label":"sum_delayedSAPOProbability_after_elimination_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.sum_delayedSAPOProbability_after_elimination_eq_one","description":"Under the same optimal-arm survival certificate, the probability vector computed after line 8 and line 15 has total mass exactly one. EAP still has to establish coordinate nonnegativity and the inactive-mass upper bound.","url":"../modules/banditrlproof-delayedfeedback-elimination/index.html#decl-2e019f70ad68","parent":"module:BanditRLProof.DelayedFeedback.Elimination","order":3873,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Elimination"],["Source","BanditRLProof/DelayedFeedback/Elimination.lean:103"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sum_delayedSAPOProbability_after_elimination_eq_one {K : Nat} (snapshot : DelayedSAPOEliminationSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (certificate : OptimalArmSurvivalCertificate snapshot mean optimal) (inactiveProbability : Fin K → ℝ) : ∑ i, delayedSAPOProbability snapshot.remainingActive inactiveProbability i = 1","missing":[],"search":"sum_delayedsapoprobability_after_elimination_eq_one banditrlproof.delayedfeedback.delayedsapoeliminationsnapshot.sum_delayedsapoprobability_after_elimination_eq_one under the same optimal-arm survival certificate, the probability vector computed after line 8 and line 15 has total mass exactly one. eap still has to establish coordinate nonnegativity and the inactive-mass upper bound. theorem compiled","shard":"modules/ee44238470f31349.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract","label":"SameAlgorithmMultiRegimeContract","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract","description":"A source-facing interface for a theorem that evaluates one algorithmic object in two environment regimes. Both endpoint predicates receive the same algorithm, initialization, tuning, information structure, and comparator. This structure is only a target contract. Constructing its data does not prove either endpoint, identify Delayed SAPO, or establish best-of-both-worlds regret.","url":"../modules/banditrlproof-delayedfeedback-multiregimecontract/index.html#decl-624e226ccef3","parent":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","order":3874,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.MultiRegimeContract"],["Source","BanditRLProof/DelayedFeedback/MultiRegimeContract.lean:15"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure SameAlgorithmMultiRegimeContract (Algorithm : Type uAlgorithm) (Initialization : Type uInitialization) (Tuning : Type uTuning) (Information : Type uInformation) (Comparator : Type uComparator) (StochasticEnvironment : Type uStochasticEnvironment) (AdversarialEnvironment : Type uAdversarialEnvironment) where","missing":[],"search":"samealgorithmmultiregimecontract banditrlproof.delayedfeedback.samealgorithmmultiregimecontract a source-facing interface for a theorem that evaluates one algorithmic object in two environment regimes. both endpoint predicates receive the same algorithm, initialization, tuning, information structure, and comparator. this structure is only a target contract. constructing its data does not prove either endpoint, identify delayed sapo, or establish best-of-both-worlds regret. structure compiled","shard":"modules/a0c7dad176accccb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim","label":"stochasticClaim","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim","description":"The stochastic endpoint instantiated with the contract's shared identity fields.","url":"../modules/banditrlproof-delayedfeedback-multiregimecontract/index.html#decl-0b7138065095","parent":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","order":3875,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.MultiRegimeContract"],["Source","BanditRLProof/DelayedFeedback/MultiRegimeContract.lean:39"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def stochasticClaim {Algorithm : Type uAlgorithm} {Initialization : Type uInitialization} {Tuning : Type uTuning} {Information : Type uInformation} {Comparator : Type uComparator} {StochasticEnvironment : Type uStochasticEnvironment} {AdversarialEnvironment : Type uAdversarialEnvironment} (contract : SameAlgorithmMultiRegimeContract Algorithm Initialization Tuning Information Comparator StochasticEnvironment AdversarialEnvironment) (environment : StochasticEnvironment) : Prop","missing":[],"search":"stochasticclaim banditrlproof.delayedfeedback.samealgorithmmultiregimecontract.stochasticclaim the stochastic endpoint instantiated with the contract's shared identity fields. definition compiled","shard":"modules/a0c7dad176accccb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim","label":"adversarialClaim","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim","description":"The adversarial endpoint instantiated with exactly the same shared algorithm, initialization, tuning, information, and comparator fields.","url":"../modules/banditrlproof-delayedfeedback-multiregimecontract/index.html#decl-c471e7046705","parent":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","order":3876,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.MultiRegimeContract"],["Source","BanditRLProof/DelayedFeedback/MultiRegimeContract.lean:56"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def adversarialClaim {Algorithm : Type uAlgorithm} {Initialization : Type uInitialization} {Tuning : Type uTuning} {Information : Type uInformation} {Comparator : Type uComparator} {StochasticEnvironment : Type uStochasticEnvironment} {AdversarialEnvironment : Type uAdversarialEnvironment} (contract : SameAlgorithmMultiRegimeContract Algorithm Initialization Tuning Information Comparator StochasticEnvironment AdversarialEnvironment) (environment : AdversarialEnvironment) : Prop","missing":[],"search":"adversarialclaim banditrlproof.delayedfeedback.samealgorithmmultiregimecontract.adversarialclaim the adversarial endpoint instantiated with exactly the same shared algorithm, initialization, tuning, information, and comparator fields. definition compiled","shard":"modules/a0c7dad176accccb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim_iff_shared_fields","label":"stochasticClaim_iff_shared_fields","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim_iff_shared_fields","description":"theorem stochasticClaim_iff_shared_fields {Algorithm : Type uAlgorithm} {Initialization : Type uInitialization} {Tuning : Type uTuning} {Information : Type uInformation} {Comparator : Type uComparator} {StochasticEnvironment : Type uStochasticEnvironment} {AdversarialEnvironment : Type uAdversarialEnvironment} (contract : SameAlgorithmMultiRegimeContract Algorithm Initialization Tuning Information Comparator Stochas…","url":"../modules/banditrlproof-delayedfeedback-multiregimecontract/index.html#decl-96f14b7d9235","parent":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","order":3877,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.MultiRegimeContract"],["Source","BanditRLProof/DelayedFeedback/MultiRegimeContract.lean:71"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem stochasticClaim_iff_shared_fields {Algorithm : Type uAlgorithm} {Initialization : Type uInitialization} {Tuning : Type uTuning} {Information : Type uInformation} {Comparator : Type uComparator} {StochasticEnvironment : Type uStochasticEnvironment} {AdversarialEnvironment : Type uAdversarialEnvironment} (contract : SameAlgorithmMultiRegimeContract Algorithm Initialization Tuning Information Comparator StochasticEnvironment AdversarialEnvironment) (environment : StochasticEnvironment) : stochasticClaim contract environment ↔ contract.stochasticEndpoint contract.algorithm contract.initialization contract.tuning contract.information contract.comparator environment","missing":[],"search":"stochasticclaim_iff_shared_fields banditrlproof.delayedfeedback.samealgorithmmultiregimecontract.stochasticclaim_iff_shared_fields theorem stochasticclaim_iff_shared_fields {algorithm : type ualgorithm} {initialization : type uinitialization} {tuning : type utuning} {information : type uinformation} {comparator : type ucomparator} {stochasticenvironment : type ustochasticenvironment} {adversarialenvironment : type uadversarialenvironment} (contract : samealgorithmmultiregimecontract algorithm initialization tuning information comparator stochasticenvironment adversarialenvironment) (environment : stochasticenvironment) : stochasticclaim contract environment ↔ contract.stochasticendpoint contract.algorithm contract.initialization contract.tuning contract.information contract.comparator environment theorem compiled","shard":"modules/a0c7dad176accccb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim_iff_shared_fields","label":"adversarialClaim_iff_shared_fields","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim_iff_shared_fields","description":"theorem adversarialClaim_iff_shared_fields {Algorithm : Type uAlgorithm} {Initialization : Type uInitialization} {Tuning : Type uTuning} {Information : Type uInformation} {Comparator : Type uComparator} {StochasticEnvironment : Type uStochasticEnvironment} {AdversarialEnvironment : Type uAdversarialEnvironment} (contract : SameAlgorithmMultiRegimeContract Algorithm Initialization Tuning Information Comparator Stocha…","url":"../modules/banditrlproof-delayedfeedback-multiregimecontract/index.html#decl-60d1b0d45d52","parent":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","order":3878,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.MultiRegimeContract"],["Source","BanditRLProof/DelayedFeedback/MultiRegimeContract.lean:88"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem adversarialClaim_iff_shared_fields {Algorithm : Type uAlgorithm} {Initialization : Type uInitialization} {Tuning : Type uTuning} {Information : Type uInformation} {Comparator : Type uComparator} {StochasticEnvironment : Type uStochasticEnvironment} {AdversarialEnvironment : Type uAdversarialEnvironment} (contract : SameAlgorithmMultiRegimeContract Algorithm Initialization Tuning Information Comparator StochasticEnvironment AdversarialEnvironment) (environment : AdversarialEnvironment) : adversarialClaim contract environment ↔ contract.adversarialEndpoint contract.algorithm contract.initialization contract.tuning contract.information contract.comparator environment","missing":[],"search":"adversarialclaim_iff_shared_fields banditrlproof.delayedfeedback.samealgorithmmultiregimecontract.adversarialclaim_iff_shared_fields theorem adversarialclaim_iff_shared_fields {algorithm : type ualgorithm} {initialization : type uinitialization} {tuning : type utuning} {information : type uinformation} {comparator : type ucomparator} {stochasticenvironment : type ustochasticenvironment} {adversarialenvironment : type uadversarialenvironment} (contract : samealgorithmmultiregimecontract algorithm initialization tuning information comparator stochasticenvironment adversarialenvironment) (environment : adversarialenvironment) : adversarialclaim contract environment ↔ contract.adversarialendpoint contract.algorithm contract.initialization contract.tuning contract.information contract.comparator environment theorem compiled","shard":"modules/a0c7dad176accccb.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose","label":"DelayedSAPONoSwitchRoundClose","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose","description":"Certificate that the no-switch inner loop for one action round is exhausted and that its final intra-round active set is the line-15 active set recorded by the source-round trace. The second field is a structural consistency contract. It is not a claim that EAP has already constructed a valid probability vector.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-ad4fbff547fe","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3879,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:32"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPONoSwitchRoundClose {K : Nat} (state : DelayedSAPOStructuralRoundState K) : Prop where","missing":[],"search":"delayedsaponoswitchroundclose banditrlproof.delayedfeedback.delayedsaponoswitchroundclose certificate that the no-switch inner loop for one action round is exhausted and that its final intra-round active set is the line-15 active set recorded by the source-round trace. the second field is a structural consistency contract. it is not a claim that eap has already constructed a valid probability vector. structure compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.processedOrder_toFinset_eq_observedBefore","label":"processedOrder_toFinset_eq_observedBefore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.processedOrder_toFinset_eq_observedBefore","description":"Exhausting `B(t) \\ S` means that the ordered ledger contains exactly the feedback available before action `t`. The equality is set-level only and does not impose a chronological order on simultaneous arrivals.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-2150d9404dc3","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3880,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:46"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem processedOrder_toFinset_eq_observedBefore {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : state.processedOrder.toFinset = observedBefore state.delayAt state.currentActionRound","missing":[],"search":"processedorder_tofinset_eq_observedbefore banditrlproof.delayedfeedback.delayedsaponoswitchroundclose.processedorder_tofinset_eq_observedbefore exhausting `b(t) \\ s` means that the ordered ledger contains exactly the feedback available before action `t`. the equality is set-level only and does not impose a chronological order on simultaneous arrivals. theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState","label":"nextRoundState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState","description":"Structural state at the start of the next action round. The processed ledger and active set are unchanged; the round-start invariant follows from the close certificate's identification with the just-finished source-round active set.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-e4b670d62f96","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3881,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:71"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def nextRoundState {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : DelayedSAPOStructuralRoundState K where","missing":[],"search":"nextroundstate banditrlproof.delayedfeedback.delayedsaponoswitchroundclose.nextroundstate structural state at the start of the next action round. the processed ledger and active set are unchanged; the round-start invariant follows from the close certificate's identification with the just-finished source-round active set. definition compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActionRound","label":"nextRoundState_currentActionRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActionRound","description":"theorem nextRoundState_currentActionRound {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : closed.nextRoundState.currentActionRound = state.currentActionRound + 1","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-dedd96499538","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3882,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:97"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem nextRoundState_currentActionRound {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : closed.nextRoundState.currentActionRound = state.currentActionRound + 1","missing":[],"search":"nextroundstate_currentactionround banditrlproof.delayedfeedback.delayedsaponoswitchroundclose.nextroundstate_currentactionround theorem nextroundstate_currentactionround {k : nat} {state : delayedsapostructuralroundstate k} (closed : delayedsaponoswitchroundclose state) : closed.nextroundstate.currentactionround = state.currentactionround + 1 theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_processedOrder","label":"nextRoundState_processedOrder","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_processedOrder","description":"theorem nextRoundState_processedOrder {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : closed.nextRoundState.processedOrder = state.processedOrder","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-ea9559079a4b","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3883,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:104"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem nextRoundState_processedOrder {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : closed.nextRoundState.processedOrder = state.processedOrder","missing":[],"search":"nextroundstate_processedorder banditrlproof.delayedfeedback.delayedsaponoswitchroundclose.nextroundstate_processedorder theorem nextroundstate_processedorder {k : nat} {state : delayedsapostructuralroundstate k} (closed : delayedsaponoswitchroundclose state) : closed.nextroundstate.processedorder = state.processedorder theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActive","label":"nextRoundState_currentActive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActive","description":"theorem nextRoundState_currentActive {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : closed.nextRoundState.currentActive = state.currentActive","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-3555278bb680","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3884,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:110"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem nextRoundState_currentActive {K : Nat} {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : closed.nextRoundState.currentActive = state.currentActive","missing":[],"search":"nextroundstate_currentactive banditrlproof.delayedfeedback.delayedsaponoswitchroundclose.nextroundstate_currentactive theorem nextroundstate_currentactive {k : nat} {state : delayedsapostructuralroundstate k} (closed : delayedsaponoswitchroundclose state) : closed.nextroundstate.currentactive = state.currentactive theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralStep","label":"DelayedSAPONoSwitchStructuralStep","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralStep","description":"One edge of the deterministic no-switch structural trace. Processing uses the exact line-8 successor; advancing rounds requires an exhausted-loop certificate.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-61304bd70bc0","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3885,"meta":[["Kind","inductive type"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:120"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive DelayedSAPONoSwitchStructuralStep {K : Nat} (horizon : Nat) : DelayedSAPOStructuralRoundState K → DelayedSAPOStructuralRoundState K → Prop | process {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) : DelayedSAPONoSwitchStructuralStep horizon state (step.afterLine8 horizon) | nextRound {state : DelayedSAPOStructuralRoundState K} (closed : DelayedSAPONoSwitchRoundClose state) : DelayedSAPONoSwitchStructuralStep horizon state closed.nextRoundState /-- Finite reflexive-transitive no-switch reachability. It is a structural relation, not a generated stochastic trajectory. -/ abbrev DelayedSAPONoSwitchStructuralReachable {K : Nat} (horizon : Nat)","missing":[],"search":"delayedsaponoswitchstructuralstep banditrlproof.delayedfeedback.delayedsaponoswitchstructuralstep one edge of the deterministic no-switch structural trace. processing uses the exact line-8 successor; advancing rounds requires an exhausted-loop certificate. inductive type compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralReachable","label":"DelayedSAPONoSwitchStructuralReachable","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralReachable","description":"Finite reflexive-transitive no-switch reachability. It is a structural relation, not a generated stochastic trajectory.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-eb611d6772fb","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3886,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:135"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"abbrev DelayedSAPONoSwitchStructuralReachable {K : Nat} (horizon : Nat)","missing":[],"search":"delayedsaponoswitchstructuralreachable banditrlproof.delayedfeedback.delayedsaponoswitchstructuralreachable finite reflexive-transitive no-switch reachability. it is a structural relation, not a generated stochastic trajectory. abbreviation compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralStep","label":"currentActive_subset_of_structuralStep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralStep","description":"Every primitive no-switch structural edge can only remove arms.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-8d6d320c1387","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3887,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:141"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem currentActive_subset_of_structuralStep {K horizon : Nat} {initial final : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchStructuralStep horizon initial final) : final.currentActive <= initial.currentActive","missing":[],"search":"currentactive_subset_of_structuralstep banditrlproof.delayedfeedback.currentactive_subset_of_structuralstep every primitive no-switch structural edge can only remove arms. theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralReachable","label":"currentActive_subset_of_structuralReachable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralReachable","description":"Active sets are antitone along every finite no-switch structural trace.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-b7c6577709f7","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3888,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:153"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem currentActive_subset_of_structuralReachable {K horizon : Nat} {initial final : DelayedSAPOStructuralRoundState K} (run : DelayedSAPONoSwitchStructuralReachable horizon initial final) : final.currentActive <= initial.currentActive","missing":[],"search":"currentactive_subset_of_structuralreachable banditrlproof.delayedfeedback.currentactive_subset_of_structuralreachable active sets are antitone along every finite no-switch structural trace. theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.mem_earlierRemainingActive_of_laterEliminated","label":"mem_earlierRemainingActive_of_laterEliminated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.mem_earlierRemainingActive_of_laterEliminated","description":"An arm eliminated by a later processing step was still present after an earlier line-8 removal whenever the two steps are connected by a no-switch structural trace. This is the temporal premise previously left to callers.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-e0bdcce54222","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3889,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:169"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem mem_earlierRemainingActive_of_laterEliminated {K horizon : Nat} {initial laterState : DelayedSAPOStructuralRoundState K} (earlierStep : DelayedSAPONoSwitchProcessOne initial) (between : DelayedSAPONoSwitchStructuralReachable horizon (earlierStep.afterLine8 horizon) laterState) (laterStep : DelayedSAPONoSwitchProcessOne laterState) (iLater : Fin K) (hLaterEliminated : iLater ∈ (laterStep.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) : iLater ∈ (earlierStep.toPreEliminationSummary.toConfidenceSnapshot horizon).remainingActive","missing":[],"search":"mem_earlierremainingactive_of_latereliminated banditrlproof.delayedfeedback.mem_earlierremainingactive_of_latereliminated an arm eliminated by a later processing step was still present after an earlier line-8 removal whenever the two steps are connected by a no-switch structural trace. this is the temporal premise previously left to callers. theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations","label":"gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations","description":"Ordered two-elimination factor-20 consumer. Unlike the one-snapshot consumer, this theorem derives later-arm survival at the earlier snapshot from the exact structural trace. It remains conditional on the earlier D.4 count clause and elimination-good projection.","url":"../modules/banditrlproof-delayedfeedback-orderednoswitchtrace/index.html#decl-0500e0abbc9b","parent":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","order":3890,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace"],["Source","BanditRLProof/DelayedFeedback/OrderedNoSwitchTrace.lean:196"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations {K : Nat} [Nonempty (Fin K)] {initial laterState : DelayedSAPOStructuralRoundState K} (horizon : Nat) (hhorizon : 1 < horizon) (earlierStep : DelayedSAPONoSwitchProcessOne initial) (between : DelayedSAPONoSwitchStructuralReachable horizon (earlierStep.afterLine8 horizon) laterState) (laterStep : DelayedSAPONoSwitchProcessOne laterState) (mean : Fin K → Real) (optimal iEarlier iLater : Fin K) (hoptimal : ∀ i, mean optimal <= mean i) (hmeanBounds : ∀ i, mean i ∈ Set.Icc (0 : Real) 1) (hD4 : earlierStep.toPreEliminationSummary.D4CountClause horizon) (hgood : (earlierStep.toPreEliminationSummary.toConfidenceSnapshot horizon).EliminationGoodEvent mean) (hoptimalActive : optimal ∈ initial.currentActive) (hEarlierEliminated : iEarlier ∈ (earlierStep.toPreEliminationSummary.toConfidenceSnapshot horizon).eliminated) (hLaterEliminate…","missing":[],"search":"gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations banditrlproof.delayedfeedback.gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations ordered two-elimination factor-20 consumer. unlike the one-snapshot consumer, this theorem derives later-arm survival at the earlier snapshot from the exact structural trace. it remains conditional on the earlier d.4 count clause and elimination-good projection. theorem compiled","shard":"modules/e439675c7b7680f1.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState","label":"DelayedSAPOStructuralRoundState","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState","description":"Structural state while Algorithm 5 is processing newly observed feedback before action `currentActionRound`. `processedOrder` is the paper sequence `S`, not a sorted set of source rounds. The current intra-round active set is contained in the previous action round's line-15 active set. Together with the antitone source-round trace, this is the primitive invariant from which a line-7 trace summary obtains current-to-…","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-27f114e40b1a","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3891,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:33"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOStructuralRoundState (K : Nat) where","missing":[],"search":"delayedsapostructuralroundstate banditrlproof.delayedfeedback.delayedsapostructuralroundstate structural state while algorithm 5 is processing newly observed feedback before action `currentactionround`. `processedorder` is the paper sequence `s`, not a sorted set of source rounds. the current intra-round active set is contained in the previous action round's line-15 active set. together with the antitone source-round trace, this is the primitive invariant from which a line-7 trace summary obtains current-to-source active persistence. structure compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.source_le_roundStart","label":"source_le_roundStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.source_le_roundStart","description":"Every source already in the processed order lies no later than the action round immediately preceding the current one.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-927b95916661","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3892,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:53"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem source_le_roundStart {K : Nat} (state : DelayedSAPOStructuralRoundState K) {s : Nat} (hs : s ∈ state.processedOrder) : s <= state.currentActionRound - 1","missing":[],"search":"source_le_roundstart banditrlproof.delayedfeedback.delayedsapostructuralroundstate.source_le_roundstart every source already in the processed order lies no later than the action round immediately preceding the current one. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.currentActive_subset_activeAtSourceRound","label":"currentActive_subset_activeAtSourceRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.currentActive_subset_activeAtSourceRound","description":"The current intra-round active set is contained in the source-round active set of every previously processed item.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-365113f68682","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3893,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:62"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem currentActive_subset_activeAtSourceRound {K : Nat} (state : DelayedSAPOStructuralRoundState K) {s : Nat} (hs : s ∈ state.processedOrder) : state.currentActive <= state.activeAtSourceRound s","missing":[],"search":"currentactive_subset_activeatsourceround banditrlproof.delayedfeedback.delayedsapostructuralroundstate.currentactive_subset_activeatsourceround the current intra-round active set is contained in the source-round active set of every previously processed item. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne","label":"DelayedSAPONoSwitchProcessOne","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne","description":"One source-faithful no-switch iteration of Algorithm 5's inner processing loop. `source_new` is exactly line 3's membership in `B(t) \\ S`. The numerical fields are the values read by line 7 after line 4 has appended `sourceRound`; constructing them from observed losses is a separate recursive-state leaf.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-0fa1b6f2e66b","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3894,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:75"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPONoSwitchProcessOne {K : Nat} (state : DelayedSAPOStructuralRoundState K) where","missing":[],"search":"delayedsaponoswitchprocessone banditrlproof.delayedfeedback.delayedsaponoswitchprocessone one source-faithful no-switch iteration of algorithm 5's inner processing loop. `source_new` is exactly line 3's membership in `b(t) \\ s`. the numerical fields are the values read by line 7 after line 4 has appended `sourceround`; constructing them from observed losses is a separate recursive-state leaf. structure compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder","label":"extendedOrder","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder","description":"Algorithm 5 line 4: append the selected source to the end of the processing sequence.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-fb52042f811a","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3895,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:89"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def extendedOrder {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) : List Nat","missing":[],"search":"extendedorder banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.extendedorder algorithm 5 line 4: append the selected source to the end of the processing sequence. definition compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.sourceRound_not_mem","label":"sourceRound_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.sourceRound_not_mem","description":"The line-3 source was not already present in the processed sequence.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-5b4c82be166d","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3896,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:94"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceRound_not_mem {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) : step.sourceRound ∉ state.processedOrder","missing":[],"search":"sourceround_not_mem banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.sourceround_not_mem the line-3 source was not already present in the processed sequence. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_nodup","label":"extendedOrder_nodup","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_nodup","description":"Appending one genuinely new source preserves duplicate freedom.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-ce0bbd0f73df","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3897,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:102"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem extendedOrder_nodup {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) : step.extendedOrder.Nodup","missing":[],"search":"extendedorder_nodup banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.extendedorder_nodup appending one genuinely new source preserves duplicate freedom. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_available","label":"extendedOrder_available","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_available","description":"Every source in the extended sequence satisfies the paper's exact strict availability condition at the current action round.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-5dc8aeb7ed6b","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3898,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:111"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem extendedOrder_available {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) {s : Nat} (hs : s ∈ step.extendedOrder) : s + state.delayAt s < state.currentActionRound","missing":[],"search":"extendedorder_available banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.extendedorder_available every source in the extended sequence satisfies the paper's exact strict availability condition at the current action round. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedSource_le_roundStart","label":"extendedSource_le_roundStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedSource_le_roundStart","description":"Every source in the extended sequence is at most the previous action round.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-4b891f3614e5","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3899,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:124"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem extendedSource_le_roundStart {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) {s : Nat} (hs : s ∈ step.extendedOrder) : s <= state.currentActionRound - 1","missing":[],"search":"extendedsource_le_roundstart banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.extendedsource_le_roundstart every source in the extended sequence is at most the previous action round. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.currentActive_subset_extendedSourceActive","label":"currentActive_subset_extendedSourceActive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.currentActive_subset_extendedSourceActive","description":"The line-7 active set is contained in every source-time line-15 active set represented by the extended processing sequence.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-b2ac56a7521c","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3900,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:134"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem currentActive_subset_extendedSourceActive {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) {s : Nat} (hs : s ∈ step.extendedOrder) : state.currentActive <= state.activeAtSourceRound s","missing":[],"search":"currentactive_subset_extendedsourceactive banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.currentactive_subset_extendedsourceactive the line-7 active set is contained in every source-time line-15 active set represented by the extended processing sequence. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.toPreEliminationSummary","label":"toPreEliminationSummary","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.toPreEliminationSummary","description":"Line-7 trace summary after, not before, Algorithm 5 line 4 appends the newly observed source. Its source-index injectivity, strict availability, and current-to-source containment are derived from the ordered transition state.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-5c69c0cd3581","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3901,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:145"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def toPreEliminationSummary {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) : DelayedSAPOProcessedTraceSummary K where","missing":[],"search":"topreeliminationsummary banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.topreeliminationsummary line-7 trace summary after, not before, algorithm 5 line 4 appends the newly observed source. its source-index injectivity, strict availability, and current-to-source containment are derived from the ordered transition state. definition compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.line8RemainingActive_subset_currentActive","label":"line8RemainingActive_subset_currentActive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.line8RemainingActive_subset_currentActive","description":"The exact Algorithm-5 line-8 removal is contained in the active set read by the line-7 snapshot.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-5a966ad5d95a","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3902,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:173"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem line8RemainingActive_subset_currentActive {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) : (step.toPreEliminationSummary.toConfidenceSnapshot horizon).remainingActive <= state.currentActive","missing":[],"search":"line8remainingactive_subset_currentactive banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.line8remainingactive_subset_currentactive the exact algorithm-5 line-8 removal is contained in the active set read by the line-7 snapshot. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8","label":"afterLine8","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8","description":"Algorithm 5 line 8 successor for the same action round. It preserves the ordered source sequence and updates only the intra-round active set to the exact line-7 complement.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-b423d03f6469","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3903,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:189"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def afterLine8 {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) : DelayedSAPOStructuralRoundState K where","missing":[],"search":"afterline8 banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.afterline8 algorithm 5 line 8 successor for the same action round. it preserves the ordered source sequence and updates only the intra-round active set to the exact line-7 complement. definition compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_currentActive_subset_before","label":"afterLine8_currentActive_subset_before","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_currentActive_subset_before","description":"Line 8 can only remove arms from the current intra-round active set.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-3d1a0019952e","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3904,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:212"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem afterLine8_currentActive_subset_before {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) : (step.afterLine8 horizon).currentActive <= state.currentActive","missing":[],"search":"afterline8_currentactive_subset_before banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.afterline8_currentactive_subset_before line 8 can only remove arms from the current intra-round active set. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_preserves_roundStart","label":"afterLine8_preserves_roundStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_preserves_roundStart","description":"The line-8 successor remains below the active set at the start of the current action round, so another arbitrary new arrival can be processed.","url":"../modules/banditrlproof-delayedfeedback-orderedprocessingtransition/index.html#decl-49ca2ad65536","parent":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","order":3905,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.OrderedProcessingTransition"],["Source","BanditRLProof/DelayedFeedback/OrderedProcessingTransition.lean:220"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem afterLine8_preserves_roundStart {K : Nat} {state : DelayedSAPOStructuralRoundState K} (step : DelayedSAPONoSwitchProcessOne state) (horizon : Nat) : (step.afterLine8 horizon).currentActive <= state.activeAtSourceRound (state.currentActionRound - 1)","missing":[],"search":"afterline8_preserves_roundstart banditrlproof.delayedfeedback.delayedsaponoswitchprocessone.afterline8_preserves_roundstart the line-8 successor remains below the active set at the start of the current action round, so another arbitrary new arrival can be processed. theorem compiled","shard":"modules/df0e9f28799a2ec2.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix","label":"DelayedSAPOProcessedPrefix","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix","description":"A ledger for one ordered ledger of source rounds whose feedback has been processed by delayed SAPO. Each entry stores the arm chosen at that source round and the Algorithm-5 allocation that was in force when the action was sampled. In particular, `inactiveProbabilityAtSource` is source-time data; it must not be reconstructed from the later processing-time state.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-f695882d30da","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3906,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:14"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOProcessedPrefix (K : Nat) where","missing":[],"search":"delayedsapoprocessedprefix banditrlproof.delayedfeedback.delayedsapoprocessedprefix a ledger for one ordered ledger of source rounds whose feedback has been processed by delayed sapo. each entry stores the arm chosen at that source round and the algorithm-5 allocation that was in force when the action was sampled. in particular, `inactiveprobabilityatsource` is source-time data; it must not be reconstructed from the later processing-time state. structure compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.processedPullCount","label":"processedPullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.processedPullCount","description":"The source count `n_i(S)` on the processed ledger.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-e1539204f65b","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3907,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:24"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def processedPullCount {K : Nat} (ledger : DelayedSAPOProcessedPrefix K) (i : Fin K) : Nat","missing":[],"search":"processedpullcount banditrlproof.delayedfeedback.delayedsapoprocessedprefix.processedpullcount the source count `n_i(s)` on the processed ledger. definition compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass","label":"expectedPullMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass","description":"The source conditional mass `sum_{s in S} p_i(s)` on the same ledger.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-aa0f41b63aec","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3908,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:29"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedPullMass {K : Nat} (ledger : DelayedSAPOProcessedPrefix K) (i : Fin K) : Real","missing":[],"search":"expectedpullmass banditrlproof.delayedfeedback.delayedsapoprocessedprefix.expectedpullmass the source conditional mass `sum_{s in s} p_i(s)` on the same ledger. definition compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass_eq_of_active_throughout","label":"expectedPullMass_eq_of_active_throughout","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass_eq_of_active_throughout","description":"Two arms that were active at every source round represented by the processed ledger receive the same cumulative Algorithm-5 probability mass. This is the deterministic line-15 producer used by the D.1 count clause.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-cb1ddef0bdbb","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3909,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:37"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expectedPullMass_eq_of_active_throughout {K : Nat} (ledger : DelayedSAPOProcessedPrefix K) (i j : Fin K) (hi : forall s, i ∈ ledger.activeAtSource s) (hj : forall s, j ∈ ledger.activeAtSource s) : ledger.expectedPullMass i = ledger.expectedPullMass j","missing":[],"search":"expectedpullmass_eq_of_active_throughout banditrlproof.delayedfeedback.delayedsapoprocessedprefix.expectedpullmass_eq_of_active_throughout two arms that were active at every source round represented by the processed ledger receive the same cumulative algorithm-5 probability mass. this is the deterministic line-15 producer used by the d.1 count clause. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate","label":"DelayedSAPOProcessedPrefixCountCertificate","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate","description":"The exact deterministic projection needed from Algorithm 5 and the pull-count clause of source Definition D.1 at one processed ledger. The certificate keeps the source-time allocation ledger explicit. It assumes the D.1 count inequalities and the source definitions of the width and the recursive empirical UCB, but it does not assume either of the width-comparison conclusions that it is designed to prove. Constructin…","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-ba2c1f2287e3","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3910,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:61"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOProcessedPrefixCountCertificate {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (ledger : DelayedSAPOProcessedPrefix K) (horizon : Nat) : Prop where","missing":[],"search":"delayedsapoprocessedprefixcountcertificate banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate the exact deterministic projection needed from algorithm 5 and the pull-count clause of source definition d.1 at one processed ledger. the certificate keeps the source-time allocation ledger explicit. it assumes the d.1 count inequalities and the source definitions of the width and the recursive empirical ucb, but it does not assume either of the width-comparison conclusions that it is designed to prove. constructing this projection from the full recursive delayed-sapo state and proving its probability are separate obligations. structure compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_nonneg","label":"sourceEmpiricalWidthScale_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_nonneg","description":"The capped source empirical width is always nonnegative.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-9d793dd85c97","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3911,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:86"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_nonneg (scale count : Real) : 0 <= sourceEmpiricalWidthScale scale count","missing":[],"search":"sourceempiricalwidthscale_nonneg banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.sourceempiricalwidthscale_nonneg the capped source empirical width is always nonnegative. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_one","label":"sourceEmpiricalWidthScale_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_one","description":"The capped source empirical width is at most one.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-cb80eb5d0d3f","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3912,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:94"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_le_one (scale count : Real) : sourceEmpiricalWidthScale scale count <= 1","missing":[],"search":"sourceempiricalwidthscale_le_one banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.sourceempiricalwidthscale_le_one the capped source empirical width is at most one. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_three_of_count_le_eight_mul","label":"sourceEmpiricalWidthScale_le_three_of_count_le_eight_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_three_of_count_le_eight_mul","description":"If one positive count is at most eight times another, then the other arm's inverse-square-root width is at most three times the reference width. The factor three is the integer relaxation of `sqrt 8` used in source D.10.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-847bf86c288a","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3913,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:104"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_le_three_of_count_le_eight_mul (scale countReference countOther : Real) (hscale : 0 <= scale) (hreference : 0 < countReference) (hcount : countReference <= 8 * countOther) : sourceEmpiricalWidthScale scale countOther <= 3 * sourceEmpiricalWidthScale scale countReference","missing":[],"search":"sourceempiricalwidthscale_le_three_of_count_le_eight_mul banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.sourceempiricalwidthscale_le_three_of_count_le_eight_mul if one positive count is at most eight times another, then the other arm's inverse-square-root width is at most three times the reference width. the factor three is the integer relaxation of `sqrt 8` used in source d.10. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.expectedPullMass_eq_of_mem_active","label":"expectedPullMass_eq_of_mem_active","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.expectedPullMass_eq_of_mem_active","description":"Active arms have equal source-time expected pull mass on the ledger.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-5bd1cc101564","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3914,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:148"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem expectedPullMass_eq_of_mem_active {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) {i j : Fin K} (hi : i ∈ snapshot.active) (hj : j ∈ snapshot.active) : ledger.expectedPullMass i = ledger.expectedPullMass j","missing":[],"search":"expectedpullmass_eq_of_mem_active banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.expectedpullmass_eq_of_mem_active active arms have equal source-time expected pull mass on the ledger. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.quarter_count_sub_six_log_le_count_of_mem_active","label":"quarter_count_sub_six_log_le_count_of_mem_active","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.quarter_count_sub_six_log_le_count_of_mem_active","description":"Combining the two exact D.1 count inequalities with equal active-arm probability mass gives the stronger `n_j >= n_i/4 - 6 log T` comparison.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-5f4ca5995fbc","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3915,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:164"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem quarter_count_sub_six_log_le_count_of_mem_active {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) {i j : Fin K} (hi : i ∈ snapshot.active) (hj : j ∈ snapshot.active) : (ledger.processedPullCount i : Real) / 4 - 6 * Real.log (horizon : Real) <= (ledger.processedPullCount j : Real)","missing":[],"search":"quarter_count_sub_six_log_le_count_of_mem_active banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.quarter_count_sub_six_log_le_count_of_mem_active combining the two exact d.1 count inequalities with equal active-arm probability mass gives the stronger `n_j >= n_i/4 - 6 log t` comparison. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.eighth_count_le_count_of_large_count","label":"eighth_count_le_count_of_large_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.eighth_count_le_count_of_large_count","description":"In the source large-count branch, any other active arm has at least one eighth of the reference arm's processed count.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-f15ff0227712","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3916,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:181"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem eighth_count_le_count_of_large_count {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) (hhorizon : 1 < horizon) {i j : Fin K} (hi : i ∈ snapshot.active) (hj : j ∈ snapshot.active) (hlarge : 192 * Real.log (horizon : Real) < (ledger.processedPullCount i : Real)) : (ledger.processedPullCount i : Real) / 8 <= (ledger.processedPullCount j : Real)","missing":[],"search":"eighth_count_le_count_of_large_count banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.eighth_count_le_count_of_large_count in the source large-count branch, any other active arm has at least one eighth of the reference arm's processed count. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_three_of_large_count","label":"empiricalWidth_le_three_of_large_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_three_of_large_count","description":"The source D.10 factor-three width producer in the large-count branch.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-398081476740","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3917,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:200"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem empiricalWidth_le_three_of_large_count {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) (hhorizon : 1 < horizon) {iReference iOther : Fin K} (hreferenceActive : iReference ∈ snapshot.active) (hotherActive : iOther ∈ snapshot.active) (hlarge : 192 * Real.log (horizon : Real) < (ledger.processedPullCount iReference : Real)) : snapshot.empiricalWidth iOther <= 3 * snapshot.empiricalWidth iReference","missing":[],"search":"empiricalwidth_le_three_of_large_count banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.empiricalwidth_le_three_of_large_count the source d.10 factor-three width producer in the large-count branch. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_ten_of_mem_active","label":"empiricalWidth_le_ten_of_mem_active","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_ten_of_mem_active","description":"The unconditional same-ledger D.10 width comparison. It is derived from the source-time allocation ledger, the D.1 count event, and the exact source width formula; no factor-ten comparison is assumed as a contract field.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-0cf5687e4860","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3918,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:230"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem empiricalWidth_le_ten_of_mem_active {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) (hhorizon : 1 < horizon) {iReference iOther : Fin K} (hreferenceActive : iReference ∈ snapshot.active) (hotherActive : iOther ∈ snapshot.active) : snapshot.empiricalWidth iOther <= 10 * snapshot.empiricalWidth iReference","missing":[],"search":"empiricalwidth_le_ten_of_mem_active banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.empiricalwidth_le_ten_of_mem_active the unconditional same-ledger d.10 width comparison. it is derived from the source-time allocation ledger, the d.1 count event, and the exact source width formula; no factor-ten comparison is assumed as a contract field. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.ucbStar_le_empiricalMean_add_width","label":"ucbStar_le_empiricalMean_add_width","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.ucbStar_le_empiricalMean_add_width","description":"The recursive empirical UCB is no larger than the current empirical mean-plus-width surface, so the source minimum `ucbStar` has the current-UCB upper edge needed by the active-arm gap proof.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-ba260142a2ea","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3919,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:266"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem ucbStar_le_empiricalMean_add_width {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) (mean : Fin K -> Real) (hgood : snapshot.EliminationGoodEvent mean) (i : Fin K) : snapshot.ucbStar <= snapshot.empiricalMean i + snapshot.empiricalWidth i","missing":[],"search":"ucbstar_le_empiricalmean_add_width banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.ucbstar_le_empiricalmean_add_width the recursive empirical ucb is no larger than the current empirical mean-plus-width surface, so the source minimum `ucbstar` has the current-ucb upper edge needed by the active-arm gap proof. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.activeArmGapBranch","label":"activeArmGapBranch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.activeArmGapBranch","description":"The exact large/small branch expected by the existing active-arm gap consumer is now produced from the processed-ledger certificate.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-ef463754ecd4","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3920,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:289"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem activeArmGapBranch {K : Nat} [Nonempty (Fin K)] {snapshot : DelayedSAPOSourceConfidenceSnapshot K} {ledger : DelayedSAPOProcessedPrefix K} {horizon : Nat} (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) (hhorizon : 1 < horizon) (mean : Fin K -> Real) (hgood : snapshot.EliminationGoodEvent mean) {optimal i : Fin K} (hoptimalActive : optimal ∈ snapshot.active) (hiActive : i ∈ snapshot.active) : (snapshot.ucbStar <= snapshot.empiricalMean optimal + snapshot.empiricalWidth optimal /\\ snapshot.empiricalWidth optimal <= 3 * snapshot.empiricalWidth i) \\/ (exists scale count : Real, 0 < scale /\\ count <= 96 * scale /\\ snapshot.empiricalWidth i = sourceEmpiricalWidthScale scale count)","missing":[],"search":"activearmgapbranch banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.activearmgapbranch the exact large/small branch expected by the existing active-arm gap consumer is now produced from the processed-ledger certificate. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countCertificate","label":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countCertificate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countCertificate","description":"Repaired same-snapshot D.12 / main-text Lemma-4.2 deterministic slice. The theorem removes the manually supplied branch and pair-width premises from the earlier consumer: both are generated by the actual Algorithm-5 source-time allocation ledger and the D.1 count clause. The full recursive state projection and probability of the source good event remain open.","url":"../modules/banditrlproof-delayedfeedback-processedprefixcounts/index.html#decl-36413eec5395","parent":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","order":3921,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.ProcessedPrefixCounts"],["Source","BanditRLProof/DelayedFeedback/ProcessedPrefixCounts.lean:329"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countCertificate {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (ledger : DelayedSAPOProcessedPrefix K) (horizon : Nat) (certificate : DelayedSAPOProcessedPrefixCountCertificate snapshot ledger horizon) (hhorizon : 1 < horizon) (mean : Fin K -> Real) (optimal iEarlier iLater : Fin K) (hoptimal : forall j, mean optimal <= mean j) (hmeanBounds : forall j, mean j ∈ Set.Icc (0 : Real) 1) (hgood : snapshot.EliminationGoodEvent mean) (hoptimalActive : optimal ∈ snapshot.active) (hEarlierEliminated : iEarlier ∈ snapshot.eliminated) (hLaterRemaining : iLater ∈ snapshot.remainingActive) : mean iLater - mean optimal <= 20 * (mean iEarlier - mean optimal)","missing":[],"search":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countcertificate banditrlproof.delayedfeedback.delayedsapoprocessedprefixcountcertificate.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countcertificate repaired same-snapshot d.12 / main-text lemma-4.2 deterministic slice. the theorem removes the manually supplied branch and pair-width premises from the earlier consumer: both are generated by the actual algorithm-5 source-time allocation ledger and the d.1 count clause. the full recursive state projection and probability of the source good event remain open. theorem compiled","shard":"modules/02feff4ea12028f0.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.newlyObservedBefore","label":"newlyObservedBefore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.newlyObservedBefore","description":"Feedback source rounds that are available before action `t` but have not yet been processed. This is the set-level content of Algorithm 5's `B(t) \\ S`; the source sequence order remains a later algorithm-state choice.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-cd8eb50d9686","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3922,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:10"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def newlyObservedBefore (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) : Finset Nat","missing":[],"search":"newlyobservedbefore banditrlproof.delayedfeedback.newlyobservedbefore feedback source rounds that are available before action `t` but have not yet been processed. this is the set-level content of algorithm 5's `b(t) \\ s`; the source sequence order remains a later algorithm-state choice. definition compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.observedBefore_mono","label":"observedBefore_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.observedBefore_mono","description":"Strictly available feedback remains available at every later action round.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-efd1aa1a3018","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3923,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:16"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem observedBefore_mono (delay : Nat → Nat) {t u : Nat} (htu : t ≤ u) : observedBefore delay t ⊆ observedBefore delay u","missing":[],"search":"observedbefore_mono banditrlproof.delayedfeedback.observedbefore_mono strictly available feedback remains available at every later action round. theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.processed_disjoint_newlyObservedBefore","label":"processed_disjoint_newlyObservedBefore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.processed_disjoint_newlyObservedBefore","description":"A source round is never both already processed and newly observed.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-f15f14abe3e5","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3924,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:25"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem processed_disjoint_newlyObservedBefore (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) : Disjoint processed (newlyObservedBefore delay processed t)","missing":[],"search":"processed_disjoint_newlyobservedbefore banditrlproof.delayedfeedback.processed_disjoint_newlyobservedbefore a source round is never both already processed and newly observed. theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.processed_union_newlyObservedBefore","label":"processed_union_newlyObservedBefore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.processed_union_newlyObservedBefore","description":"If all processed rounds were legitimately available, adjoining every new arrival yields exactly the current available set.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-e5be5861e576","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3925,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:34"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem processed_union_newlyObservedBefore (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) (hprocessed : processed ⊆ observedBefore delay t) : processed ∪ newlyObservedBefore delay processed t = observedBefore delay t","missing":[],"search":"processed_union_newlyobservedbefore banditrlproof.delayedfeedback.processed_union_newlyobservedbefore if all processed rounds were legitimately available, adjoining every new arrival yields exactly the current available set. theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.processAllNew","label":"processAllNew","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.processAllNew","description":"The update obtained by processing every currently new arrival is the current available set.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-1f078775a64c","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3926,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:52"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def processAllNew (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) : Finset Nat","missing":[],"search":"processallnew banditrlproof.delayedfeedback.processallnew the update obtained by processing every currently new arrival is the current available set. definition compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.processAllNew_eq_observedBefore","label":"processAllNew_eq_observedBefore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.processAllNew_eq_observedBefore","description":"theorem processAllNew_eq_observedBefore (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) (hprocessed : processed ⊆ observedBefore delay t) : processAllNew delay processed t = observedBefore delay t","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-5dca1337fc6d","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3927,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:56"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem processAllNew_eq_observedBefore (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) (hprocessed : processed ⊆ observedBefore delay t) : processAllNew delay processed t = observedBefore delay t","missing":[],"search":"processallnew_eq_observedbefore banditrlproof.delayedfeedback.processallnew_eq_observedbefore theorem processallnew_eq_observedbefore (delay : nat → nat) (processed : finset nat) (t : nat) (hprocessed : processed ⊆ observedbefore delay t) : processallnew delay processed t = observedbefore delay t theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.previousObservedBefore_subset_current","label":"previousObservedBefore_subset_current","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.previousObservedBefore_subset_current","description":"A completed earlier available set is a valid processed prefix later.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-adc7c4d7e942","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3928,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:63"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem previousObservedBefore_subset_current (delay : Nat → Nat) {t u : Nat} (htu : t ≤ u) : observedBefore delay t ⊆ observedBefore delay u","missing":[],"search":"previousobservedbefore_subset_current banditrlproof.delayedfeedback.previousobservedbefore_subset_current a completed earlier available set is a valid processed prefix later. theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.processAllNew_from_previous_eq_current","label":"processAllNew_from_previous_eq_current","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.processAllNew_from_previous_eq_current","description":"Starting from the fully processed set at an earlier round and processing all arrivals through a later round yields exactly the later available set.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-a2d6c5bbfbe2","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3929,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:70"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem processAllNew_from_previous_eq_current (delay : Nat → Nat) {t u : Nat} (htu : t ≤ u) : processAllNew delay (observedBefore delay t) u = observedBefore delay u","missing":[],"search":"processallnew_from_previous_eq_current banditrlproof.delayedfeedback.processallnew_from_previous_eq_current starting from the fully processed set at an earlier round and processing all arrivals through a later round yields exactly the later available set. theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.outstandingAt_disjoint_newlyObservedBefore","label":"outstandingAt_disjoint_newlyObservedBefore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.outstandingAt_disjoint_newlyObservedBefore","description":"Outstanding feedback cannot appear in the newly observed batch at the same action time.","url":"../modules/banditrlproof-delayedfeedback-processing/index.html#decl-3879d92ee192","parent":"module:BanditRLProof.DelayedFeedback.Processing","order":3930,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.Processing"],["Source","BanditRLProof/DelayedFeedback/Processing.lean:79"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem outstandingAt_disjoint_newlyObservedBefore (delay : Nat → Nat) (processed : Finset Nat) (t : Nat) : Disjoint (outstandingAt delay t) (newlyObservedBefore delay processed t)","missing":[],"search":"outstandingat_disjoint_newlyobservedbefore banditrlproof.delayedfeedback.outstandingat_disjoint_newlyobservedbefore outstanding feedback cannot appear in the newly observed batch at the same action time. theorem compiled","shard":"modules/be65cd8268927f5e.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary","label":"DelayedSAPOProcessedTraceSummary","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary","description":"Source-shaped deterministic summary of a Delayed-SAPO processed trace. `sourceIndex` is an ordered processing ledger, not a chronological source-round prefix: feedback from a later source round may be processed before feedback from an earlier source round. Its entries are distinct and satisfy the exact strict availability test at `currentActionRound`. `activeAtSourceRound` records the line-15 sampling set at past so…","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-bd3f46728108","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3931,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:42"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOProcessedTraceSummary (K : Nat) where","missing":[],"search":"delayedsapoprocessedtracesummary banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary source-shaped deterministic summary of a delayed-sapo processed trace. `sourceindex` is an ordered processing ledger, not a chronological source-round prefix: feedback from a later source round may be processed before feedback from an earlier source round. its entries are distinct and satisfy the exact strict availability test at `currentactionround`. `activeatsourceround` records the line-15 sampling set at past source rounds and is antitone because algorithm 5 only removes arms. the possibly intra-round `currentactive` set is stored separately. its containment in every ledger source set is an explicit trace-summary invariant that a future algorithm-5 producer must prove. the current empirical surfaces are likewise separate from the source-time allocations used by the ledger. this record is an interface summary, not a proof that algorithm 5 generates the supplied fields. structure compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefix","label":"toProcessedPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefix","description":"Read every processed entry at its recorded source round.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-1b7d5be0bc86","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3932,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:65"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def toProcessedPrefix {K : Nat} (state : DelayedSAPOProcessedTraceSummary K) : DelayedSAPOProcessedPrefix K where","missing":[],"search":"toprocessedprefix banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.toprocessedprefix read every processed entry at its recorded source round. definition compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalWidthAt","label":"empiricalWidthAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalWidthAt","description":"The exact source empirical width at the current processed state.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-876a8fec370b","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3933,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:76"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalWidthAt {K : Nat} (state : DelayedSAPOProcessedTraceSummary K) (horizon : Nat) (i : Fin K) : Real","missing":[],"search":"empiricalwidthat banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.empiricalwidthat the exact source empirical width at the current processed state. definition compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalUpperAt","label":"empiricalUpperAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalUpperAt","description":"The recursive empirical UCB printed by the source, evaluated from the current empirical mean, the current processed count, and the preceding UCB.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-f6b6ca2e95ca","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3934,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:84"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalUpperAt {K : Nat} (state : DelayedSAPOProcessedTraceSummary K) (horizon : Nat) (i : Fin K) : Real","missing":[],"search":"empiricalupperat banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.empiricalupperat the recursive empirical ucb printed by the source, evaluated from the current empirical mean, the current processed count, and the preceding ucb. definition compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toConfidenceSnapshot","label":"toConfidenceSnapshot","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toConfidenceSnapshot","description":"Confidence snapshot projected from the supplied trace summary. The width and recursive empirical UCB are definitions here, rather than certificate premises.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-dff2a3d92b87","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3935,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:93"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def toConfidenceSnapshot {K : Nat} (state : DelayedSAPOProcessedTraceSummary K) (horizon : Nat) : DelayedSAPOSourceConfidenceSnapshot K where","missing":[],"search":"toconfidencesnapshot banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.toconfidencesnapshot confidence snapshot projected from the supplied trace summary. the width and recursive empirical ucb are definitions here, rather than certificate premises. definition compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.D4CountClause","label":"D4CountClause","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.D4CountClause","description":"The two armwise count inequalities printed in source Definition D.1 and used by source Lemma D.4. This is the stochastic boundary of the present module: proving that it holds simultaneously with probability at least `1 - 2 / T` on the generated trajectory remains open.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-855338998120","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3936,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:107"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure D4CountClause {K : Nat} (state : DelayedSAPOProcessedTraceSummary K) (horizon : Nat) : Prop where","missing":[],"search":"d4countclause banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.d4countclause the two armwise count inequalities printed in source definition d.1 and used by source lemma d.4. this is the stochastic boundary of the present module: proving that it holds simultaneously with probability at least `1 - 2 / t` on the generated trajectory remains open. structure compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.currentActive_subset_activeAt_sourceIndex","label":"currentActive_subset_activeAt_sourceIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.currentActive_subset_activeAt_sourceIndex","description":"Expose the trace-summary invariant that every arm in the intra-round current set was active at every source round in the processed ledger. Producing this invariant from Algorithm 5 is deliberately outside this adapter.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-2ebf1f55b874","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3937,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:122"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem currentActive_subset_activeAt_sourceIndex {K : Nat} (state : DelayedSAPOProcessedTraceSummary K) (q : Fin state.length) : state.currentActive <= state.activeAtSourceRound (state.sourceIndex q)","missing":[],"search":"currentactive_subset_activeat_sourceindex banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.currentactive_subset_activeat_sourceindex expose the trace-summary invariant that every arm in the intra-round current set was active at every source round in the processed ledger. producing this invariant from algorithm 5 is deliberately outside this adapter. theorem compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefixCountCertificate","label":"toProcessedPrefixCountCertificate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefixCountCertificate","description":"Deterministic trace-summary adapter for the processed-prefix count certificate. Source-time allocation data come from `toProcessedPrefix`, active persistence is an explicit trace-summary invariant, and the two count bounds are the explicit D.4 boundary. No width comparison or gap conclusion is assumed.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-0f65b002c905","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3938,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:133"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem toProcessedPrefixCountCertificate {K : Nat} [Nonempty (Fin K)] (state : DelayedSAPOProcessedTraceSummary K) (horizon : Nat) (hD4 : state.D4CountClause horizon) : DelayedSAPOProcessedPrefixCountCertificate (state.toConfidenceSnapshot horizon) state.toProcessedPrefix horizon where","missing":[],"search":"toprocessedprefixcountcertificate banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.toprocessedprefixcountcertificate deterministic trace-summary adapter for the processed-prefix count certificate. source-time allocation data come from `toprocessedprefix`, active persistence is an explicit trace-summary invariant, and the two count bounds are the explicit d.4 boundary. no width comparison or gap conclusion is assumed. theorem compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_traceSummary","label":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_traceSummary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_traceSummary","description":"Downstream same-snapshot factor-twenty consumer reached from a processed trace summary plus the explicit D.4 count clause. This is still conditional on the elimination projection of the source good event; neither its probability nor an ordered multi-snapshot elimination theorem is claimed.","url":"../modules/banditrlproof-delayedfeedback-recursiveprocessedstate/index.html#decl-85142a0b4e27","parent":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","order":3939,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.RecursiveProcessedState"],["Source","BanditRLProof/DelayedFeedback/RecursiveProcessedState.lean:157"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_traceSummary {K : Nat} [Nonempty (Fin K)] (state : DelayedSAPOProcessedTraceSummary K) (horizon : Nat) (hD4 : state.D4CountClause horizon) (hhorizon : 1 < horizon) (mean : Fin K -> Real) (optimal iEarlier iLater : Fin K) (hoptimal : forall j, mean optimal <= mean j) (hmeanBounds : forall j, mean j ∈ Set.Icc (0 : Real) 1) (hgood : (state.toConfidenceSnapshot horizon).EliminationGoodEvent mean) (hoptimalActive : optimal ∈ (state.toConfidenceSnapshot horizon).active) (hEarlierEliminated : iEarlier ∈ (state.toConfidenceSnapshot horizon).eliminated) (hLaterRemaining : iLater ∈ (state.toConfidenceSnapshot horizon).remainingActive) : mean iLater - mean optimal <= 20 * (mean iEarlier - mean optimal)","missing":[],"search":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_tracesummary banditrlproof.delayedfeedback.delayedsapoprocessedtracesummary.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_tracesummary downstream same-snapshot factor-twenty consumer reached from a processed trace summary plus the explicit d.4 count clause. this is still conditional on the elimination projection of the source good event; neither its probability nor an ordered multi-snapshot elimination theorem is claimed. theorem compiled","shard":"modules/03ff1d65e572fce8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.finiteAverageGap","label":"finiteAverageGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.finiteAverageGap","description":"Arithmetic mean of a finite family of real gaps. Lean's total division makes this definition equal to zero when `K = 0`; the counting theorem below handles that empty case before using the denominator.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html#decl-5b492dfd4f7c","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","order":3940,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGapHalfSet"],["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean:26"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteAverageGap {K : Nat} (gap : Fin K -> Real) : Real","missing":[],"search":"finiteaveragegap banditrlproof.delayedfeedback.finiteaveragegap arithmetic mean of a finite family of real gaps. lean's total division makes this definition equal to zero when `k = 0`; the counting theorem below handles that empty case before using the denominator. definition compiled","shard":"modules/fc10501d3a23a119.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.aboveTwiceAverageGap","label":"aboveTwiceAverageGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.aboveTwiceAverageGap","description":"Indices whose gaps are strictly greater than twice the finite average, matching the strict word \"greater\" in the source statement of Lemma D.11.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html#decl-f94e0ba63eec","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","order":3941,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGapHalfSet"],["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean:31"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def aboveTwiceAverageGap {K : Nat} (gap : Fin K -> Real) : Finset (Fin K)","missing":[],"search":"abovetwiceaveragegap banditrlproof.delayedfeedback.abovetwiceaveragegap indices whose gaps are strictly greater than twice the finite average, matching the strict word \"greater\" in the source statement of lemma d.11. definition compiled","shard":"modules/fc10501d3a23a119.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.two_mul_card_aboveTwiceAverageGap_le","label":"two_mul_card_aboveTwiceAverageGap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.two_mul_card_aboveTwiceAverageGap_le","description":"Nonnegative-domain deterministic content used by the source's Lemma D.11. Nonnegativity is explicit rather than extending the promoted contract to an arbitrary signed family. The theorem includes `K = 0`; for positive `K`, its proof separately closes the zero-average branch before cancelling the positive average in the usual counting/Markov argument.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html#decl-431f01f66ed0","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","order":3942,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapHalfSet"],["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean:41"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem two_mul_card_aboveTwiceAverageGap_le {K : Nat} (gap : Fin K -> Real) (hgap : forall i, 0 <= gap i) : 2 * (aboveTwiceAverageGap gap).card <= K","missing":[],"search":"two_mul_card_abovetwiceaveragegap_le banditrlproof.delayedfeedback.two_mul_card_abovetwiceaveragegap_le nonnegative-domain deterministic content used by the source's lemma d.11. nonnegativity is explicit rather than extending the promoted contract to an arbitrary signed family. the theorem includes `k = 0`; for positive `k`, its proof separately closes the zero-average branch before cancelling the positive average in the usual counting/markov argument. theorem compiled","shard":"modules/fc10501d3a23a119.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceStochasticLossGap","label":"sourceStochasticLossGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceStochasticLossGap","description":"The stochastic loss gap convention used by the delayed-SAPO source: smaller mean loss is better, so the gap of arm `i` is `mean i - mean optimal`.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html#decl-79376c3f7c7c","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","order":3943,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGapHalfSet"],["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean:110"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def sourceStochasticLossGap {K : Nat} (mean : Fin K -> Real) (optimal i : Fin K) : Real","missing":[],"search":"sourcestochasticlossgap banditrlproof.delayedfeedback.sourcestochasticlossgap the stochastic loss gap convention used by the delayed-sapo source: smaller mean loss is better, so the gap of arm `i` is `mean i - mean optimal`. definition compiled","shard":"modules/fc10501d3a23a119.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceStochasticLossGap_nonneg","label":"sourceStochasticLossGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceStochasticLossGap_nonneg","description":"An optimal arm makes every source stochastic loss gap nonnegative.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html#decl-f27be598ee0c","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","order":3944,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapHalfSet"],["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean:115"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceStochasticLossGap_nonneg {K : Nat} (mean : Fin K -> Real) (optimal : Fin K) (hoptimal : forall i, mean optimal <= mean i) (i : Fin K) : 0 <= sourceStochasticLossGap mean optimal i","missing":[],"search":"sourcestochasticlossgap_nonneg banditrlproof.delayedfeedback.sourcestochasticlossgap_nonneg an optimal arm makes every source stochastic loss gap nonnegative. theorem compiled","shard":"modules/fc10501d3a23a119.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.two_mul_card_sourceStochasticLossGap_aboveTwiceAverage_le","label":"two_mul_card_sourceStochasticLossGap_aboveTwiceAverage_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.two_mul_card_sourceStochasticLossGap_aboveTwiceAverage_le","description":"Bandit specialization of the nonnegative-domain D.11 counting statement. Among the `K` stochastic loss gaps from an optimal arm, strictly fewer than half can exceed twice their average (and hence their cardinality is at most half). This is a deterministic producer; it does not assume or claim any generated delayed-feedback probability law.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaphalfset/index.html#decl-9e61f4924159","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","order":3945,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapHalfSet"],["Source","BanditRLProof/DelayedFeedback/StochasticGapHalfSet.lean:128"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","delayed-nonstationary"]],"statement":"theorem two_mul_card_sourceStochasticLossGap_aboveTwiceAverage_le {K : Nat} (mean : Fin K -> Real) (optimal : Fin K) (hoptimal : forall i, mean optimal <= mean i) : 2 * (aboveTwiceAverageGap (sourceStochasticLossGap mean optimal)).card <= K","missing":[],"search":"two_mul_card_sourcestochasticlossgap_abovetwiceaverage_le banditrlproof.delayedfeedback.two_mul_card_sourcestochasticlossgap_abovetwiceaverage_le bandit specialization of the nonnegative-domain d.11 counting statement. among the `k` stochastic loss gaps from an optimal arm, strictly fewer than half can exceed twice their average (and hence their cardinality is at most half). this is a deterministic producer; it does not assume or claim any generated delayed-feedback probability law. theorem compiled","shard":"modules/fc10501d3a23a119.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale","label":"sourceEmpiricalWidthScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale","description":"The scalar form of the empirical width printed in the delayed-SAPO source, with `scale = 2 * log T` and `count = n_i(S)`. A nonpositive count uses the capped width `1`; this prevents Lean's totalized real division at zero from manufacturing a zero-width observation. Keeping the scale explicit isolates the positive-count order issue from logarithmic side conditions.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-9514fe1c7d3b","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3946,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:12"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceEmpiricalWidthScale (scale count : Real) : Real","missing":[],"search":"sourceempiricalwidthscale banditrlproof.delayedfeedback.sourceempiricalwidthscale the scalar form of the empirical width printed in the delayed-sapo source, with `scale = 2 * log t` and `count = n_i(s)`. a nonpositive count uses the capped width `1`; this prevents lean's totalized real division at zero from manufacturing a zero-width observation. keeping the scale explicit isolates the positive-count order issue from logarithmic side conditions. definition compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_antitone","label":"sourceEmpiricalWidthScale_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_antitone","description":"For nonnegative scale and positive counts, the source empirical width is antitone in the count. Thus a later state with at least as many pulls has no larger width than an earlier prefix.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-f542d5195f47","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3947,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:18"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_antitone (scale countEarlier countLater : Real) (hscale : 0 <= scale) (hcountEarlier : 0 < countEarlier) (hcount : countEarlier <= countLater) : sourceEmpiricalWidthScale scale countLater <= sourceEmpiricalWidthScale scale countEarlier","missing":[],"search":"sourceempiricalwidthscale_antitone banditrlproof.delayedfeedback.sourceempiricalwidthscale_antitone for nonnegative scale and positive counts, the source empirical width is antitone in the count. thus a later state with at least as many pulls has no larger width than an earlier prefix. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_of_count_le_96_mul_scale","label":"one_le_ten_mul_sourceEmpiricalWidthScale_of_count_le_96_mul_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_of_count_le_96_mul_scale","description":"The small-count branch used in the source proof of Lemma D.10. If the count is at most `96 * scale`, then the capped inverse-square-root width is at least one tenth. For the printed choice `scale = 2 * log T`, this is exactly the implication from `count <= 192 * log T` to `1 <= 10 * width`.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-3b8a78697092","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3948,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:35"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem one_le_ten_mul_sourceEmpiricalWidthScale_of_count_le_96_mul_scale (scale count : Real) (hscale : 0 < scale) (hcount : count <= 96 * scale) : 1 <= 10 * sourceEmpiricalWidthScale scale count","missing":[],"search":"one_le_ten_mul_sourceempiricalwidthscale_of_count_le_96_mul_scale banditrlproof.delayedfeedback.one_le_ten_mul_sourceempiricalwidthscale_of_count_le_96_mul_scale the small-count branch used in the source proof of lemma d.10. if the count is at most `96 * scale`, then the capped inverse-square-root width is at least one tenth. for the printed choice `scale = 2 * log t`, this is exactly the implication from `count <= 192 * log t` to `1 <= 10 * width`. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_two_log_of_small_count","label":"one_le_ten_mul_sourceEmpiricalWidthScale_two_log_of_small_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_two_log_of_small_count","description":"Source-parameter specialization of the preceding small-count lemma.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-20d0715b4716","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3949,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:58"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem one_le_ten_mul_sourceEmpiricalWidthScale_two_log_of_small_count (horizon count : Real) (hhorizon : 1 < horizon) (hcount : count <= 192 * Real.log horizon) : 1 <= 10 * sourceEmpiricalWidthScale (2 * Real.log horizon) count","missing":[],"search":"one_le_ten_mul_sourceempiricalwidthscale_two_log_of_small_count banditrlproof.delayedfeedback.one_le_ten_mul_sourceempiricalwidthscale_two_log_of_small_count source-parameter specialization of the preceding small-count lemma. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_one","label":"sourceEmpiricalWidthScale_one_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_one","description":"Exact small instance used to audit the direction of the displayed D.10 prefix-to-elimination inequality.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-396d21fff494","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3950,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:70"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_one_one : sourceEmpiricalWidthScale 1 1 = 1","missing":[],"search":"sourceempiricalwidthscale_one_one banditrlproof.delayedfeedback.sourceempiricalwidthscale_one_one exact small instance used to audit the direction of the displayed d.10 prefix-to-elimination inequality. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_four","label":"sourceEmpiricalWidthScale_one_four","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_four","description":"Four times the count gives half the uncapped width in the same exact instance.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-16f78dc38c24","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3951,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:77"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceEmpiricalWidthScale_one_four : sourceEmpiricalWidthScale 1 4 = (1 / 2 : Real)","missing":[],"search":"sourceempiricalwidthscale_one_four banditrlproof.delayedfeedback.sourceempiricalwidthscale_one_four four times the count gives half the uncapped width in the same exact instance. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_one_le_four","label":"not_sourceEmpiricalWidthScale_one_le_four","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_one_le_four","description":"The reverse inequality used in the displayed D.10 proof is not a generic consequence of prefix count growth: it already fails at scale one between counts one and four. This diagnoses an edge of the frozen proof, not a counterexample to every possible repair of Lemma D.10.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-deb477c66b53","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3952,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:85"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem not_sourceEmpiricalWidthScale_one_le_four : not (sourceEmpiricalWidthScale 1 1 <= sourceEmpiricalWidthScale 1 4)","missing":[],"search":"not_sourceempiricalwidthscale_one_le_four banditrlproof.delayedfeedback.not_sourceempiricalwidthscale_one_le_four the reverse inequality used in the displayed d.10 proof is not a generic consequence of prefix count growth: it already fails at scale one between counts one and four. this diagnoses an edge of the frozen proof, not a counterexample to every possible repair of lemma d.10. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_horizon_four_one_le_four","label":"not_sourceEmpiricalWidthScale_horizon_four_one_le_four","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_horizon_four_one_le_four","description":"A literal source-width instance at integer horizon `T = 4`. Moving from one to four processed pulls strictly decreases the printed capped radius, so the reverse transport fails inside the paper's own parameter domain rather than only for the normalized scale-one diagnostic above.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-10a2c77670fb","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3953,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:94"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem not_sourceEmpiricalWidthScale_horizon_four_one_le_four : ¬ (sourceEmpiricalWidthScale (2 * Real.log 4) 1 <= sourceEmpiricalWidthScale (2 * Real.log 4) 4)","missing":[],"search":"not_sourceempiricalwidthscale_horizon_four_one_le_four banditrlproof.delayedfeedback.not_sourceempiricalwidthscale_horizon_four_one_le_four a literal source-width instance at integer horizon `t = 4`. moving from one to four processed pulls strictly decreases the printed capped radius, so the reverse transport fails inside the paper's own parameter domain rather than only for the normalized scale-one diagnostic above. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.eight_mul_empiricalWidth_lt_gap_of_mem_eliminated","label":"eight_mul_empiricalWidth_lt_gap_of_mem_eliminated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.eight_mul_empiricalWidth_lt_gap_of_mem_eliminated","description":"At one source elimination snapshot, the line-7 strict test and the elimination projection of the stochastic good event put the eliminated arm's true gap strictly above eight empirical widths. This is the lower endpoint used by the repaired D.12 route below; it is derived from the actual source snapshot rather than supplied as a scalar contract field.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-94edae056fd8","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3954,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:133"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem eight_mul_empiricalWidth_lt_gap_of_mem_eliminated {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (optimal i : Fin K) (hoptimal : forall j, mean optimal <= mean j) (hgood : snapshot.EliminationGoodEvent mean) (hi : i ∈ snapshot.eliminated) : 8 * snapshot.empiricalWidth i < mean i - mean optimal","missing":[],"search":"eight_mul_empiricalwidth_lt_gap_of_mem_eliminated banditrlproof.delayedfeedback.eight_mul_empiricalwidth_lt_gap_of_mem_eliminated at one source elimination snapshot, the line-7 strict test and the elimination projection of the stochastic good event put the eliminated arm's true gap strictly above eight empirical widths. this is the lower endpoint used by the repaired d.12 route below; it is derived from the actual source snapshot rather than supplied as a scalar contract field. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive","label":"gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive","description":"A still-active arm's gap is at most sixteen of its widths at the same processed prefix, once the source upper-confidence surface is bounded by the current optimal-arm empirical radius and the optimal width is within factor three. These are exactly the algebraic facts used inside the displayed D.10 proof; no transport to a later elimination prefix occurs here.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-8888f98f9fae","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3955,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:158"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (optimal i : Fin K) (hgood : snapshot.EliminationGoodEvent mean) (hi : i ∈ snapshot.remainingActive) (hucbCurrent : snapshot.ucbStar <= snapshot.empiricalMean optimal + snapshot.empiricalWidth optimal) (hoptimalWidth : snapshot.empiricalWidth optimal <= 3 * snapshot.empiricalWidth i) : mean i - mean optimal <= 16 * snapshot.empiricalWidth i","missing":[],"search":"gap_le_sixteen_mul_empiricalwidth_of_mem_remainingactive banditrlproof.delayedfeedback.gap_le_sixteen_mul_empiricalwidth_of_mem_remainingactive a still-active arm's gap is at most sixteen of its widths at the same processed prefix, once the source upper-confidence surface is bounded by the current optimal-arm empirical radius and the optimal width is within factor three. these are exactly the algebraic facts used inside the displayed d.10 proof; no transport to a later elimination prefix occurs here. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive_of_large_or_small_count","label":"gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive_of_large_or_small_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive_of_large_or_small_count","description":"Source-faithful case split behind the D.10 active-arm upper endpoint. The large-count branch supplies the current-UCB and factor-three width edges. The small-count branch supplies the printed width formula together with `count <= 96 * scale`; bounded losses then give `gap <= 1 <= 10 * width`. This removes the unconditional factor-three assumption from the combined consumer while leaving the recursive count and width…","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-cb8ebacd8789","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3956,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:191"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive_of_large_or_small_count {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (optimal i : Fin K) (hmeanBounds : forall j, mean j ∈ Set.Icc (0 : Real) 1) (hgood : snapshot.EliminationGoodEvent mean) (hi : i ∈ snapshot.remainingActive) (hbranch : (snapshot.ucbStar <= snapshot.empiricalMean optimal + snapshot.empiricalWidth optimal /\\ snapshot.empiricalWidth optimal <= 3 * snapshot.empiricalWidth i) \\/ (exists scale count : Real, 0 < scale /\\ count <= 96 * scale /\\ snapshot.empiricalWidth i = sourceEmpiricalWidthScale scale count)) : mean i - mean optimal <= 16 * snapshot.empiricalWidth i","missing":[],"search":"gap_le_sixteen_mul_empiricalwidth_of_mem_remainingactive_of_large_or_small_count banditrlproof.delayedfeedback.gap_le_sixteen_mul_empiricalwidth_of_mem_remainingactive_of_large_or_small_count source-faithful case split behind the d.10 active-arm upper endpoint. the large-count branch supplies the current-ucb and factor-three width edges. the small-count branch supplies the printed width formula together with `count <= 96 * scale`; bounded losses then give `gap <= 1 <= 10 * width`. this removes the unconditional factor-three assumption from the combined consumer while leaving the recursive count and width-shape producers explicit. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot","label":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot","description":"Same-snapshot repair of the displayed D.12 / main-text Lemma 4.2 chain. When `iEarlier` is eliminated and `iLater` remains active in that very update, the active-prefix D.10 gap upper bound can be consumed before any later elimination snapshot is mentioned. The proof therefore avoids the reversed prefix-to-elimination width inequality diagnosed above. The two width comparison premises still have to be produced by th…","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-f7417e6b51c4","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3957,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:229"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_twenty_mul_gap_at_earlier_elimination_snapshot {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (optimal iEarlier iLater : Fin K) (hoptimal : forall j, mean optimal <= mean j) (hgood : snapshot.EliminationGoodEvent mean) (hEarlierEliminated : iEarlier ∈ snapshot.eliminated) (hLaterRemaining : iLater ∈ snapshot.remainingActive) (hucbCurrent : snapshot.ucbStar <= snapshot.empiricalMean optimal + snapshot.empiricalWidth optimal) (hoptimalWidth : snapshot.empiricalWidth optimal <= 3 * snapshot.empiricalWidth iLater) (hpairWidth : snapshot.empiricalWidth iLater <= 10 * snapshot.empiricalWidth iEarlier) : mean iLater - mean optimal <= 20 * (mean iEarlier - mean optimal)","missing":[],"search":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot banditrlproof.delayedfeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot same-snapshot repair of the displayed d.12 / main-text lemma 4.2 chain. when `iearlier` is eliminated and `ilater` remains active in that very update, the active-prefix d.10 gap upper bound can be consumed before any later elimination snapshot is mentioned. the proof therefore avoids the reversed prefix-to-elimination width inequality diagnosed above. the two width comparison premises still have to be produced by the source count event on a recursive delayed sapo trajectory. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count","label":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count","description":"The same-snapshot factor-twenty consumer with the source's large/small- count split exposed. In the small-count branch the factor-three premise is replaced by bounded means and the exact source-width lower bound. The theorem still requires a recursive producer for the selected branch and the same- prefix factor-ten comparison; it is not an unconditional port of D.10/D.12.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-c2295de5f85e","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3958,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:260"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (optimal iEarlier iLater : Fin K) (hoptimal : forall j, mean optimal <= mean j) (hmeanBounds : forall j, mean j ∈ Set.Icc (0 : Real) 1) (hgood : snapshot.EliminationGoodEvent mean) (hEarlierEliminated : iEarlier ∈ snapshot.eliminated) (hLaterRemaining : iLater ∈ snapshot.remainingActive) (hbranch : (snapshot.ucbStar <= snapshot.empiricalMean optimal + snapshot.empiricalWidth optimal /\\ snapshot.empiricalWidth optimal <= 3 * snapshot.empiricalWidth iLater) \\/ (exists scale count : Real, 0 < scale /\\ count <= 96 * scale /\\ snapshot.empiricalWidth iLater = sourceEmpiricalWidthScale scale count)) (hpairWidth : snapshot.empiricalWidth iLater <= 10 * snapshot.empiricalWidth iEarlier) : mean iLater - mean op…","missing":[],"search":"gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count banditrlproof.delayedfeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count the same-snapshot factor-twenty consumer with the source's large/small- count split exposed. in the small-count branch the factor-three premise is replaced by bounded means and the exact source-width lower bound. the theorem still requires a recursive producer for the selected branch and the same- prefix factor-ten comparison; it is not an unconditional port of d.10/d.12. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract","label":"DelayedSAPOD10D12GapOrderingContract","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract","description":"Exact deterministic interface needed by the displayed proof of source Lemma D.12 (main-text Lemma 4.2). Its index is a shared processed-sequence prefix length, not wall-clock action time. The fields deliberately name the four edges consumed by D.12: D.10 supplies the gap endpoints and cross-arm comparison, while the width definition supplies the time transport.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-a97f7132b8d1","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3959,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:295"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOD10D12GapOrderingContract (K : Nat) where","missing":[],"search":"delayedsapod10d12gaporderingcontract banditrlproof.delayedfeedback.delayedsapod10d12gaporderingcontract exact deterministic interface needed by the displayed proof of source lemma d.12 (main-text lemma 4.2). its index is a shared processed-sequence prefix length, not wall-clock action time. the fields deliberately name the four edges consumed by d.12: d.10 supplies the gap endpoints and cross-arm comparison, while the width definition supplies the time transport. structure compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap","label":"surrogateGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap","description":"The source surrogate gap `Delta-tilde_i = 8 width_i(S-tilde_i)`.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-e69c14dce811","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3960,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:312"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def surrogateGap {K : Nat} (contract : DelayedSAPOD10D12GapOrderingContract K) (i : Fin K) : Real","missing":[],"search":"surrogategap banditrlproof.delayedfeedback.delayedsapod10d12gaporderingcontract.surrogategap the source surrogate gap `delta-tilde_i = 8 width_i(s-tilde_i)`. definition compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap_le_gap","label":"surrogateGap_le_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap_le_gap","description":"Lower half of the displayed D.10 two-sided surrogate-gap comparison.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-72c08e59043b","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3961,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:318"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem surrogateGap_le_gap {K : Nat} (contract : DelayedSAPOD10D12GapOrderingContract K) (i : Fin K) : contract.surrogateGap i <= contract.gap i","missing":[],"search":"surrogategap_le_gap banditrlproof.delayedfeedback.delayedsapod10d12gaporderingcontract.surrogategap_le_gap lower half of the displayed d.10 two-sided surrogate-gap comparison. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_two_mul_surrogateGap","label":"gap_le_two_mul_surrogateGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_two_mul_surrogateGap","description":"Upper half of the displayed D.10 two-sided comparison, exposed from the factor-16 endpoint rather than assumed in factor-two form.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-d9bdb5344427","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3962,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:326"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_two_mul_surrogateGap {K : Nat} (contract : DelayedSAPOD10D12GapOrderingContract K) (i : Fin K) : contract.gap i <= 2 * contract.surrogateGap i","missing":[],"search":"gap_le_two_mul_surrogategap banditrlproof.delayedfeedback.delayedsapod10d12gaporderingcontract.gap_le_two_mul_surrogategap upper half of the displayed d.10 two-sided comparison, exposed from the factor-16 endpoint rather than assumed in factor-two form. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.d12_gap_ordering_chain","label":"d12_gap_ordering_chain","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.d12_gap_ordering_chain","description":"The four inequalities in the source's displayed D.12 chain, kept separate so an audit can identify which edge is missing from a recursive Delayed SAPO implementation.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-cd61e0f1307b","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3963,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:337"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem d12_gap_ordering_chain {K : Nat} (contract : DelayedSAPOD10D12GapOrderingContract K) (iEarlier iLater : Fin K) (horder : contract.eliminationPrefixIndex iEarlier <= contract.eliminationPrefixIndex iLater) : contract.gap iLater <= 16 * contract.widthAt iLater (contract.eliminationPrefixIndex iLater) /\\ 16 * contract.widthAt iLater (contract.eliminationPrefixIndex iLater) <= 16 * contract.widthAt iLater (contract.eliminationPrefixIndex iEarlier) /\\ 16 * contract.widthAt iLater (contract.eliminationPrefixIndex iEarlier) <= 160 * contract.widthAt iEarlier (contract.eliminationPrefixIndex iEarlier) /\\ 160 * contract.widthAt iEarlier (contract.eliminationPrefixIndex iEarlier) <= 20 * contract.gap iEarlier","missing":[],"search":"d12_gap_ordering_chain banditrlproof.delayedfeedback.delayedsapod10d12gaporderingcontract.d12_gap_ordering_chain the four inequalities in the source's displayed d.12 chain, kept separate so an audit can identify which edge is missing from a recursive delayed sapo implementation. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_twenty_mul_gap_of_eliminationPrefixIndex_le","label":"gap_le_twenty_mul_gap_of_eliminationPrefixIndex_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_twenty_mul_gap_of_eliminationPrefixIndex_le","description":"Conditional source Lemma D.12 / main-text Lemma 4.2 consumer. It proves the factor-20 gap ordering once the exact D.10 endpoints and the correctly oriented width transport are supplied; it does not claim those disputed inputs follow from the current one-snapshot library.","url":"../modules/banditrlproof-delayedfeedback-stochasticgaporderingaudit/index.html#decl-092f3676d5d5","parent":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","order":3964,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit"],["Source","BanditRLProof/DelayedFeedback/StochasticGapOrderingAudit.lean:367"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem gap_le_twenty_mul_gap_of_eliminationPrefixIndex_le {K : Nat} (contract : DelayedSAPOD10D12GapOrderingContract K) (iEarlier iLater : Fin K) (horder : contract.eliminationPrefixIndex iEarlier <= contract.eliminationPrefixIndex iLater) : contract.gap iLater <= 20 * contract.gap iEarlier","missing":[],"search":"gap_le_twenty_mul_gap_of_eliminationprefixindex_le banditrlproof.delayedfeedback.delayedsapod10d12gaporderingcontract.gap_le_twenty_mul_gap_of_eliminationprefixindex_le conditional source lemma d.12 / main-text lemma 4.2 consumer. it proves the factor-20 gap ordering once the exact d.10 endpoints and the correctly oriented width transport are supplied; it does not claim those disputed inputs follow from the current one-snapshot library. theorem compiled","shard":"modules/a204c8cacf49e908.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot","label":"DelayedSAPOSourceConfidenceSnapshot","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot","description":"The two armwise upper-confidence surfaces entering the source definition of `ucbStar`. The recursive construction of those surfaces from a processed history remains outside this snapshot.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-76471a7cfe13","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3965,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:13"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOSourceConfidenceSnapshot (K : Nat) extends DelayedSAPOEliminationSnapshot K where","missing":[],"search":"delayedsaposourceconfidencesnapshot banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot the two armwise upper-confidence surfaces entering the source definition of `ucbstar`. the recursive construction of those surfaces from a processed history remains outside this snapshot. structure compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.sourceUcbStar","label":"sourceUcbStar","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.sourceUcbStar","description":"Source-shaped projection of `ucbStar(S) = min_i {ucb_i(S), overline-ucb_i(S)}`.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-02056dda6d88","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3966,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:22"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceUcbStar {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) : ℝ","missing":[],"search":"sourceucbstar banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.sourceucbstar source-shaped projection of `ucbstar(s) = min_i {ucb_i(s), overline-ucb_i(s)}`. definition compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.EliminationGoodEvent","label":"EliminationGoodEvent","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.EliminationGoodEvent","description":"The elimination-relevant projection of the stochastic good event in Definition D.1. It records both empirical-mean confidence and the two upper confidence surfaces used by `ucbStar`. Count/phase/error/delay clauses from the full source event are deliberately not included here.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-b69b5dadb171","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3967,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:31"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure EliminationGoodEvent {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) : Prop where","missing":[],"search":"eliminationgoodevent banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.eliminationgoodevent the elimination-relevant projection of the stochastic good event in definition d.1. it records both empirical-mean confidence and the two upper confidence surfaces used by `ucbstar`. count/phase/error/delay clauses from the full source event are deliberately not included here. structure compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalMean_le_ucbStar_of_eliminationGoodEvent","label":"optimalMean_le_ucbStar_of_eliminationGoodEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalMean_le_ucbStar_of_eliminationGoodEvent","description":"On the source-shaped elimination good event, the best loss mean is below the exact minimum of both armwise upper-confidence surfaces.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-27ea0d0eb69d","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3968,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:43"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem optimalMean_le_ucbStar_of_eliminationGoodEvent {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (hoptimal : ∀ i, mean optimal ≤ mean i) (hgood : EliminationGoodEvent snapshot mean) : mean optimal ≤ snapshot.ucbStar","missing":[],"search":"optimalmean_le_ucbstar_of_eliminationgoodevent banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.optimalmean_le_ucbstar_of_eliminationgoodevent on the source-shaped elimination good event, the best loss mean is below the exact minimum of both armwise upper-confidence surfaces. theorem compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalArmSurvivalCertificate_of_eliminationGoodEvent","label":"optimalArmSurvivalCertificate_of_eliminationGoodEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalArmSurvivalCertificate_of_eliminationGoodEvent","description":"The source-shaped elimination projection constructs the exact certificate consumed by the deterministic core of Lemma D.9; no independent `mean optimal <= ucbStar` premise remains.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-9819b35aa8f5","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3969,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:60"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem optimalArmSurvivalCertificate_of_eliminationGoodEvent {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (hoptimal : ∀ i, mean optimal ≤ mean i) (hactive : optimal ∈ snapshot.active) (hgood : EliminationGoodEvent snapshot mean) : DelayedSAPOEliminationSnapshot.OptimalArmSurvivalCertificate snapshot.toDelayedSAPOEliminationSnapshot mean optimal where","missing":[],"search":"optimalarmsurvivalcertificate_of_eliminationgoodevent banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.optimalarmsurvivalcertificate_of_eliminationgoodevent the source-shaped elimination projection constructs the exact certificate consumed by the deterministic core of lemma d.9; no independent `mean optimal <= ucbstar` premise remains. theorem compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimal_mem_remainingActive_of_eliminationGoodEvent","label":"optimal_mem_remainingActive_of_eliminationGoodEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimal_mem_remainingActive_of_eliminationGoodEvent","description":"Source Lemma D.9 at one elimination snapshot, conditional only on the elimination projection of Definition D.1 and current activity of the optimal arm. Establishing the full good-event probability and persistence over the recursive state machine remain separate obligations.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-56a3cd78ffc8","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3970,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:80"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem optimal_mem_remainingActive_of_eliminationGoodEvent {K : Nat} [Nonempty (Fin K)] (snapshot : DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (hoptimal : ∀ i, mean optimal ≤ mean i) (hactive : optimal ∈ snapshot.active) (hgood : EliminationGoodEvent snapshot mean) : optimal ∈ snapshot.remainingActive","missing":[],"search":"optimal_mem_remainingactive_of_eliminationgoodevent banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.optimal_mem_remainingactive_of_eliminationgoodevent source lemma d.9 at one elimination snapshot, conditional only on the elimination projection of definition d.1 and current activity of the optimal arm. establishing the full good-event probability and persistence over the recursive state machine remain separate obligations. theorem compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet","label":"eliminationGoodEventSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet","description":"Random-state event on which the elimination projection of Definition D.1 holds.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-daa13ea1cea9","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3971,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:95"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def eliminationGoodEventSet {Ω : Type*} {K : Nat} [Nonempty (Fin K)] (snapshot : Ω → DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) : Set Ω","missing":[],"search":"eliminationgoodeventset banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.eliminationgoodeventset random-state event on which the elimination projection of definition d.1 holds. definition compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalSurvivalEventSet","label":"optimalSurvivalEventSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalSurvivalEventSet","description":"Event that the optimal arm survives the current elimination update.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-bcc25788803d","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3972,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:101"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def optimalSurvivalEventSet {Ω : Type*} {K : Nat} (snapshot : Ω → DelayedSAPOSourceConfidenceSnapshot K) (optimal : Fin K) : Set Ω","missing":[],"search":"optimalsurvivaleventset banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.optimalsurvivaleventset event that the optimal arm survives the current elimination update. definition compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet_subset_optimalSurvivalEventSet","label":"eliminationGoodEventSet_subset_optimalSurvivalEventSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet_subset_optimalSurvivalEventSet","description":"The elimination good event is contained in the optimal-arm survival event when the optimal arm is active in every input snapshot.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-bb05b7b4facd","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3973,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:108"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem eliminationGoodEventSet_subset_optimalSurvivalEventSet {Ω : Type*} {K : Nat} [Nonempty (Fin K)] (snapshot : Ω → DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (hoptimal : ∀ i, mean optimal ≤ mean i) (hactive : ∀ ω, optimal ∈ (snapshot ω).active) : eliminationGoodEventSet snapshot mean ⊆ optimalSurvivalEventSet snapshot optimal","missing":[],"search":"eliminationgoodeventset_subset_optimalsurvivaleventset banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.eliminationgoodeventset_subset_optimalsurvivaleventset the elimination good event is contained in the optimal-arm survival event when the optimal arm is active in every input snapshot. theorem compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le","label":"measure_optimalSurvivalEventSet_compl_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le","description":"Any tail bound for the complement of the source-shaped good event immediately controls the probability that the optimal arm is eliminated.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-73d222e205a6","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3974,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:122"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measure_optimalSurvivalEventSet_compl_le {Ω : Type*} [MeasurableSpace Ω] {K : Nat} [Nonempty (Fin K)] (mu : Measure Ω) (snapshot : Ω → DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (hoptimal : ∀ i, mean optimal ≤ mean i) (hactive : ∀ ω, optimal ∈ (snapshot ω).active) : mu (optimalSurvivalEventSet snapshot optimal)ᶜ ≤ mu (eliminationGoodEventSet snapshot mean)ᶜ","missing":[],"search":"measure_optimalsurvivaleventset_compl_le banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.measure_optimalsurvivaleventset_compl_le any tail bound for the complement of the source-shaped good event immediately controls the probability that the optimal arm is eliminated. theorem compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le_of_goodEvent","label":"measure_optimalSurvivalEventSet_compl_le_of_goodEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le_of_goodEvent","description":"Failure-budget consumer for the D.9 projection. The hypothesis must be discharged by the D.2--D.7 concentration/counting development and the Corollary-D.8 union assembly; this theorem does not manufacture that probability estimate.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodevent/index.html#decl-54f8107a4890","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","order":3975,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEvent"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEvent.lean:141"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measure_optimalSurvivalEventSet_compl_le_of_goodEvent {Ω : Type*} [MeasurableSpace Ω] {K : Nat} [Nonempty (Fin K)] (mu : Measure Ω) (snapshot : Ω → DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K → ℝ) (optimal : Fin K) (delta : ℝ) (hoptimal : ∀ i, mean optimal ≤ mean i) (hactive : ∀ ω, optimal ∈ (snapshot ω).active) (hgoodProbability : mu (eliminationGoodEventSet snapshot mean)ᶜ ≤ ENNReal.ofReal delta) : mu (optimalSurvivalEventSet snapshot optimal)ᶜ ≤ ENNReal.ofReal delta","missing":[],"search":"measure_optimalsurvivaleventset_compl_le_of_goodevent banditrlproof.delayedfeedback.delayedsaposourceconfidencesnapshot.measure_optimalsurvivaleventset_compl_le_of_goodevent failure-budget consumer for the d.9 projection. the hypothesis must be discharged by the d.2--d.7 concentration/counting development and the corollary-d.8 union assembly; this theorem does not manufacture that probability estimate. theorem compiled","shard":"modules/4a04fac997a67e65.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventComponent","label":"DelayedSAPOGoodEventComponent","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventComponent","description":"The six source components combined by Corollary D.8. These constructors name the failure events proved separately in Lemmas D.2--D.7; they do not assert those concentration results.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-2a91dfafa227","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3976,"meta":[["Kind","inductive type"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:13"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive DelayedSAPOGoodEventComponent where","missing":[],"search":"delayedsapogoodeventcomponent banditrlproof.delayedfeedback.delayedsapogoodeventcomponent the six source components combined by corollary d.8. these constructors name the failure events proved separately in lemmas d.2--d.7; they do not assert those concentration results. inductive type compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily","label":"DelayedSAPOGoodEventFailureFamily","kind":"structure","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily","description":"Failure-event family corresponding to the six clauses combined in source Corollary D.8. The stochastic-delay component may be set to the empty event when only the oblivious-delay version is used.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-49680efab6e1","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3977,"meta":[["Kind","structure"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:25"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure DelayedSAPOGoodEventFailureFamily (Omega : Type*) where","missing":[],"search":"delayedsapogoodeventfailurefamily banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily failure-event family corresponding to the six clauses combined in source corollary d.8. the stochastic-delay component may be set to the empty event when only the oblivious-delay version is used. structure compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.componentFailure","label":"componentFailure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.componentFailure","description":"Select the failure event named by a source good-event component.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-a319f12e8485","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3978,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:36"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def componentFailure {Omega : Type*} (family : DelayedSAPOGoodEventFailureFamily Omega) : DelayedSAPOGoodEventComponent -> Set Omega | .bscConfidence => family.bscConfidence | .eapConfidence => family.eapConfidence | .pullCount => family.pullCount | .eliminatedDelay => family.eliminatedDelay | .lossDifference => family.lossDifference | .stochasticDelay => family.stochasticDelay /-- The bad event appearing in the union-bound proof of Corollary D.8. -/ def failureSet {Omega : Type*} (family : DelayedSAPOGoodEventFailureFamily Omega) : Set Omega","missing":[],"search":"componentfailure banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.componentfailure select the failure event named by a source good-event component. definition compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.failureSet","label":"failureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.failureSet","description":"The bad event appearing in the union-bound proof of Corollary D.8.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-14edfb7f5662","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3979,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:47"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def failureSet {Omega : Type*} (family : DelayedSAPOGoodEventFailureFamily Omega) : Set Omega","missing":[],"search":"failureset banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.failureset the bad event appearing in the union-bound proof of corollary d.8. definition compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet","label":"sourceGoodEventSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet","description":"Source good event assembled from the complements of the six D.2--D.7 failure events.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-1928b1b63f10","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3980,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:53"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def sourceGoodEventSet {Omega : Type*} (family : DelayedSAPOGoodEventFailureFamily Omega) : Set Omega","missing":[],"search":"sourcegoodeventset banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.sourcegoodeventset source good event assembled from the complements of the six d.2--d.7 failure events. definition compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet_compl","label":"sourceGoodEventSet_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet_compl","description":"The complement of the assembled source good event is exactly the union of the six named failure events.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-15283b27e1f5","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3981,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:59"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceGoodEventSet_compl {Omega : Type*} (family : DelayedSAPOGoodEventFailureFamily Omega) : family.sourceGoodEventSetᶜ = family.failureSet","missing":[],"search":"sourcegoodeventset_compl banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.sourcegoodeventset_compl the complement of the assembled source good event is exactly the union of the six named failure events. theorem compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_sum","label":"measure_sourceGoodEventSet_compl_le_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_sum","description":"Finite outer-measure union bound for the six D.2--D.7 components.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-4b070f6bfa95","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3982,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:65"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measure_sourceGoodEventSet_compl_le_sum {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) (family : DelayedSAPOGoodEventFailureFamily Omega) : mu family.sourceGoodEventSetᶜ <= (Finset.univ : Finset DelayedSAPOGoodEventComponent).sum (fun component => mu (family.componentFailure component))","missing":[],"search":"measure_sourcegoodeventset_compl_le_sum banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.measure_sourcegoodeventset_compl_le_sum finite outer-measure union bound for the six d.2--d.7 components. theorem compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.linearFailureBudget","label":"linearFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.linearFailureBudget","description":"The source `1 / T` budget used for Lemmas D.5--D.7.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-7b62509a9a3d","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3983,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:77"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def linearFailureBudget (horizon : Nat) : Real","missing":[],"search":"linearfailurebudget banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.linearfailurebudget the source `1 / t` budget used for lemmas d.5--d.7. definition compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.doubleLinearFailureBudget","label":"doubleLinearFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.doubleLinearFailureBudget","description":"The source `2 / T` budget used for Lemmas D.2--D.4.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-f3bac5fd0cfb","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3984,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:81"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def doubleLinearFailureBudget (horizon : Nat) : Real","missing":[],"search":"doublelinearfailurebudget banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.doublelinearfailurebudget the source `2 / t` budget used for lemmas d.2--d.4. definition compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceComponentFailureBudget","label":"sourceComponentFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceComponentFailureBudget","description":"The source-exact failure share assigned to each clause in Corollary D.8: D.2--D.4 contribute `2 / T` each and D.5--D.7 contribute `1 / T` each.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-4346780c0130","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3985,"meta":[["Kind","definition"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:86"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceComponentFailureBudget (horizon : Nat) : DelayedSAPOGoodEventComponent -> Real | .bscConfidence => doubleLinearFailureBudget horizon | .eapConfidence => doubleLinearFailureBudget horizon | .pullCount => doubleLinearFailureBudget horizon | .eliminatedDelay => linearFailureBudget horizon | .lossDifference => linearFailureBudget horizon | .stochasticDelay => linearFailureBudget horizon /-- The six source-exact shares in Corollary D.8 sum to `9 / T`. -/ theorem sum_sourceComponentFailureBudget_eq_nine_div (horizon : Nat) : (Finset.univ : Finset DelayedSAPOGoodEventComponent).sum (sourceComponentFailureBudget horizon) = 9 / (horizon : Real)","missing":[],"search":"sourcecomponentfailurebudget banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.sourcecomponentfailurebudget the source-exact failure share assigned to each clause in corollary d.8: d.2--d.4 contribute `2 / t` each and d.5--d.7 contribute `1 / t` each. definition compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sum_sourceComponentFailureBudget_eq_nine_div","label":"sum_sourceComponentFailureBudget_eq_nine_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sum_sourceComponentFailureBudget_eq_nine_div","description":"The six source-exact shares in Corollary D.8 sum to `9 / T`.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-0f0ad107ccc2","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3986,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:96"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sum_sourceComponentFailureBudget_eq_nine_div (horizon : Nat) : (Finset.univ : Finset DelayedSAPOGoodEventComponent).sum (sourceComponentFailureBudget horizon) = 9 / (horizon : Real)","missing":[],"search":"sum_sourcecomponentfailurebudget_eq_nine_div banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.sum_sourcecomponentfailurebudget_eq_nine_div the six source-exact shares in corollary d.8 sum to `9 / t`. theorem compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_nine_div","label":"measure_sourceGoodEventSet_compl_le_nine_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_nine_div","description":"Corollary-D.8 union assembly. The six hypotheses are precisely the probability-producing obligations of Lemmas D.2--D.7. This theorem combines them and proves the paper's deliberately loose `9 / T` failure budget; it does not prove the six component concentration lemmas.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-96b360716658","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3987,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:115"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measure_sourceGoodEventSet_compl_le_nine_div {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) (family : DelayedSAPOGoodEventFailureFamily Omega) (horizon : Nat) (hhorizon : 0 < horizon) (hbsc : mu family.bscConfidence <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (heap : mu family.eapConfidence <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (hpull : mu family.pullCount <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (heliminated : mu family.eliminatedDelay <= ENNReal.ofReal (linearFailureBudget horizon)) (hloss : mu family.lossDifference <= ENNReal.ofReal (linearFailureBudget horizon)) (hdelay : mu family.stochasticDelay <= ENNReal.ofReal (linearFailureBudget horizon)) : mu family.sourceGoodEventSetᶜ <= ENNReal.ofReal (9 / (horizon : Real))","missing":[],"search":"measure_sourcegoodeventset_compl_le_nine_div banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.measure_sourcegoodeventset_compl_le_nine_div corollary-d.8 union assembly. the six hypotheses are precisely the probability-producing obligations of lemmas d.2--d.7. this theorem combines them and proves the paper's deliberately loose `9 / t` failure budget; it does not prove the six component concentration lemmas. theorem compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_eliminationGoodEventSet_compl_le_nine_div","label":"measure_eliminationGoodEventSet_compl_le_nine_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_eliminationGoodEventSet_compl_le_nine_div","description":"The full source good event implies the already compiled elimination slice when the projection relation is recorded explicitly.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-45a6c3cb9d05","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3988,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:169"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measure_eliminationGoodEventSet_compl_le_nine_div {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [Nonempty (Fin K)] (mu : Measure Omega) (family : DelayedSAPOGoodEventFailureFamily Omega) (snapshot : Omega -> DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (horizon : Nat) (hhorizon : 0 < horizon) (hprojection : family.sourceGoodEventSet ⊆ DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet snapshot mean) (hbsc : mu family.bscConfidence <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (heap : mu family.eapConfidence <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (hpull : mu family.pullCount <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (heliminated : mu family.eliminatedDelay <= ENNReal.ofReal (linearFailureBudget horizon)) (hloss : mu family.lossDifference <= ENNReal.ofReal (linearFailureBudget horizon)) (hdelay : mu family.stochas…","missing":[],"search":"measure_eliminationgoodeventset_compl_le_nine_div banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.measure_eliminationgoodeventset_compl_le_nine_div the full source good event implies the already compiled elimination slice when the projection relation is recorded explicitly. theorem compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_optimalSurvivalEventSet_compl_le_nine_div","label":"measure_optimalSurvivalEventSet_compl_le_nine_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_optimalSurvivalEventSet_compl_le_nine_div","description":"Corollary D.8 composed with the compiled one-snapshot D.9 projection: component failure budgets control optimal-arm elimination. Recursive persistence and both regret endpoints remain separate obligations.","url":"../modules/banditrlproof-delayedfeedback-stochasticgoodeventassembly/index.html#decl-ec6aa351c990","parent":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","order":3989,"meta":[["Kind","theorem"],["Module","BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly"],["Source","BanditRLProof/DelayedFeedback/StochasticGoodEventAssembly.lean:203"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem measure_optimalSurvivalEventSet_compl_le_nine_div {Omega : Type*} [MeasurableSpace Omega] {K : Nat} [Nonempty (Fin K)] (mu : Measure Omega) (family : DelayedSAPOGoodEventFailureFamily Omega) (snapshot : Omega -> DelayedSAPOSourceConfidenceSnapshot K) (mean : Fin K -> Real) (optimal : Fin K) (horizon : Nat) (hhorizon : 0 < horizon) (hoptimal : forall i, mean optimal <= mean i) (hactive : forall omega, optimal ∈ (snapshot omega).active) (hprojection : family.sourceGoodEventSet ⊆ DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet snapshot mean) (hbsc : mu family.bscConfidence <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (heap : mu family.eapConfidence <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (hpull : mu family.pullCount <= ENNReal.ofReal (doubleLinearFailureBudget horizon)) (heliminated : mu family.eliminatedDelay <= ENNReal.ofReal (linearFailureBud…","missing":[],"search":"measure_optimalsurvivaleventset_compl_le_nine_div banditrlproof.delayedfeedback.delayedsapogoodeventfailurefamily.measure_optimalsurvivaleventset_compl_le_nine_div corollary d.8 composed with the compiled one-snapshot d.9 projection: component failure budgets control optimal-arm elimination. recursive persistence and both regret endpoints remain separate obligations. theorem compiled","shard":"modules/9ae235d6048c94ed.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.MeasurableFiniteActionDistribution","label":"MeasurableFiniteActionDistribution","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Exp3.MeasurableFiniteActionDistribution","description":"Pointwise probability-vector and coordinate-measurability contracts.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-1096e7a5aa9f","parent":"module:BanditRLProof.Exp3ActionProcess","order":3990,"meta":[["Kind","structure"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"structure MeasurableFiniteActionDistribution {History : Type u} {Action : Type v} [MeasurableSpace History] (arms : Finset Action) (prob : History -> Action -> Real) : Prop where","missing":[],"search":"measurablefiniteactiondistribution banditrlproof.exp3.measurablefiniteactiondistribution pointwise probability-vector and coordinate-measurability contracts. structure compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionKernel","label":"finiteActionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionKernel","description":"The history-adaptive finite action kernel generated by `prob`.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-6a80285fcc5e","parent":"module:BanditRLProof.Exp3ActionProcess","order":3991,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteActionKernel {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) : Kernel History Action where","missing":[],"search":"finiteactionkernel banditrlproof.exp3.finiteactionkernel the history-adaptive finite action kernel generated by `prob`. definition compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionKernel_apply","label":"finiteActionKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionKernel_apply","description":"theorem finiteActionKernel_apply {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (history : History) : finiteActionKernel arms prob source history = finiteActionMeasure arms (prob history)","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-7fb76b2fbb20","parent":"module:BanditRLProof.Exp3ActionProcess","order":3992,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:48"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionKernel_apply {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (history : History) : finiteActionKernel arms prob source history = finiteActionMeasure arms (prob history)","missing":[],"search":"finiteactionkernel_apply banditrlproof.exp3.finiteactionkernel_apply theorem finiteactionkernel_apply {history : type u} {action : type v} [measurablespace history] [measurablespace action] [measurablesingletonclass action] (arms : finset action) (prob : history -> action -> real) (source : measurablefiniteactiondistribution arms prob) (history : history) : finiteactionkernel arms prob source history = finiteactionmeasure arms (prob history) theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcessMeasure","label":"actionProcessMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcessMeasure","description":"Canonical joint history/action law generated by the adaptive policy.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-b68651a60b11","parent":"module:BanditRLProof.Exp3ActionProcess","order":3993,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:72"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def actionProcessMeasure {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) : Measure (History × Action)","missing":[],"search":"actionprocessmeasure banditrlproof.exp3.actionprocessmeasure canonical joint history/action law generated by the adaptive policy. definition compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcessHistory","label":"actionProcessHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcessHistory","description":"History coordinate of the generated one-round process.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-c4ec3cf4ba37","parent":"module:BanditRLProof.Exp3ActionProcess","order":3994,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:105"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def actionProcessHistory {History Action : Type*} : History × Action -> History","missing":[],"search":"actionprocesshistory banditrlproof.exp3.actionprocesshistory history coordinate of the generated one-round process. definition compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcessAction","label":"actionProcessAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcessAction","description":"Sampled action coordinate of the generated one-round process.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-955ceece2b1b","parent":"module:BanditRLProof.Exp3ActionProcess","order":3995,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:109"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def actionProcessAction {History Action : Type*} : History × Action -> Action","missing":[],"search":"actionprocessaction banditrlproof.exp3.actionprocessaction sampled action coordinate of the generated one-round process. definition compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcessHistory_measurable","label":"actionProcessHistory_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcessHistory_measurable","description":"theorem actionProcessHistory_measurable {History Action : Type*} [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@actionProcessHistory History Action)","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-d163ae58e6d7","parent":"module:BanditRLProof.Exp3ActionProcess","order":3996,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:112"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcessHistory_measurable {History Action : Type*} [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@actionProcessHistory History Action)","missing":[],"search":"actionprocesshistory_measurable banditrlproof.exp3.actionprocesshistory_measurable theorem actionprocesshistory_measurable {history action : type*} [measurablespace history] [measurablespace action] : measurable (@actionprocesshistory history action) theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcessAction_measurable","label":"actionProcessAction_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcessAction_measurable","description":"theorem actionProcessAction_measurable {History Action : Type*} [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@actionProcessAction History Action)","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-7859870cf305","parent":"module:BanditRLProof.Exp3ActionProcess","order":3997,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:118"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcessAction_measurable {History Action : Type*} [MeasurableSpace History] [MeasurableSpace Action] : Measurable (@actionProcessAction History Action)","missing":[],"search":"actionprocessaction_measurable banditrlproof.exp3.actionprocessaction_measurable theorem actionprocessaction_measurable {history action : type*} [measurablespace history] [measurablespace action] : measurable (@actionprocessaction history action) theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_history_map_eq","label":"actionProcess_history_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_history_map_eq","description":"The generated process preserves the supplied history marginal.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-af0dafcf8005","parent":"module:BanditRLProof.Exp3ActionProcess","order":3998,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:125"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_history_map_eq {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) : (actionProcessMeasure historyMu arms prob source).map actionProcessHistory = historyMu","missing":[],"search":"actionprocess_history_map_eq banditrlproof.exp3.actionprocess_history_map_eq the generated process preserves the supplied history marginal. theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_condDistrib_action_ae_eq_finiteActionKernel","label":"actionProcess_condDistrib_action_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_condDistrib_action_ae_eq_finiteActionKernel","description":"The sampled action has the generated policy as its a.e. conditional law.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-1d094b862324","parent":"module:BanditRLProof.Exp3ActionProcess","order":3999,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:138"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_condDistrib_action_ae_eq_finiteActionKernel {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) : condDistrib actionProcessAction actionProcessHistory (actionProcessMeasure historyMu arms prob source) =ᵐ[ (actionProcessMeasure historyMu arms prob source).map actionProcessHistory] finiteActionKernel arms prob source","missing":[],"search":"actionprocess_conddistrib_action_ae_eq_finiteactionkernel banditrlproof.exp3.actionprocess_conddistrib_action_ae_eq_finiteactionkernel the sampled action has the generated policy as its a.e. conditional law. theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionKernel_ae_eq_finiteActionMeasure","label":"finiteActionKernel_ae_eq_finiteActionMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionKernel_ae_eq_finiteActionMeasure","description":"The generated policy is pointwise the explicit finite Dirac action law.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-23809a0d4a31","parent":"module:BanditRLProof.Exp3ActionProcess","order":4000,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:160"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionKernel_ae_eq_finiteActionMeasure {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] (historyMu : Measure History) (arms : Finset Action) (prob : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) : finiteActionKernel arms prob source =ᵐ[historyMu] fun history => finiteActionMeasure arms (prob history)","missing":[],"search":"finiteactionkernel_ae_eq_finiteactionmeasure banditrlproof.exp3.finiteactionkernel_ae_eq_finiteactionmeasure the generated policy is pointwise the explicit finite dirac action law. theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_integral_importanceWeightedLoss_eq_integral_loss","label":"actionProcess_integral_importanceWeightedLoss_eq_integral_loss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_integral_importanceWeightedLoss_eq_integral_loss","description":"Canonical generated-process armwise importance-weighted identity.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-8a4e3bbbd1bb","parent":"module:BanditRLProof.Exp3ActionProcess","order":4001,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:173"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_importanceWeightedLoss_eq_integral_loss {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (hprob : forall history action, action ∈ arms -> 0 < prob history action) (comparator : Action) (hcomparator : comparator ∈ arms) (hscore : Measurable (fun z : History × Action => importanceWeightedLoss (prob z.1) (loss z.1) z.2 comparator)) (hIntegrable : Integrable (fun z : History × Action => importanceWeightedLoss (prob z.1) (loss z.1) z.2 comparator) (historyMu ⊗ₘ finiteActionKernel arms prob source)) : integral (actionProcessMeasure historyMu arms prob…","missing":[],"search":"actionprocess_integral_importanceweightedloss_eq_integral_loss banditrlproof.exp3.actionprocess_integral_importanceweightedloss_eq_integral_loss canonical generated-process armwise importance-weighted identity. theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss","label":"actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss","description":"Canonical generated-process mixed first-moment identity.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-69b5a18bfbc5","parent":"module:BanditRLProof.Exp3ActionProcess","order":4002,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:208"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (hprob : forall history action, action ∈ arms -> 0 < prob history action) (hscore : Measurable (fun z : History × Action => mixedImportanceWeightedLoss arms (prob z.1) (loss z.1) z.2)) (hIntegrable : Integrable (fun z : History × Action => mixedImportanceWeightedLoss arms (prob z.1) (loss z.1) z.2) (historyMu ⊗ₘ finiteActionKernel arms prob source)) : integral (actionProcessMeasure historyMu arms prob source) (fun sample => mixedImportanceWeightedL…","missing":[],"search":"actionprocess_integral_mixedimportanceweightedloss_eq_integral_mixedloss banditrlproof.exp3.actionprocess_integral_mixedimportanceweightedloss_eq_integral_mixedloss canonical generated-process mixed first-moment identity. theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq","label":"actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq","description":"Canonical generated-process mixed second-moment identity.","url":"../modules/banditrlproof-exp3actionprocess/index.html#decl-766997e7d125","parent":"module:BanditRLProof.Exp3ActionProcess","order":4003,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ActionProcess"],["Source","BanditRLProof/Exp3ActionProcess.lean:242"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (hprob : forall history action, action ∈ arms -> 0 < prob history action) (hscore : Measurable (fun z : History × Action => mixedSquaredImportanceWeightedLoss arms (prob z.1) (loss z.1) z.2)) (hIntegrable : Integrable (fun z : History × Action => mixedSquaredImportanceWeightedLoss arms (prob z.1) (loss z.1) z.2) (historyMu ⊗ₘ finiteActionKernel arms prob source)) : integral (actionProcessMeasure historyMu arms prob source) (fun sample => m…","missing":[],"search":"actionprocess_integral_mixedsquaredimportanceweightedloss_eq_integral_sum_loss_sq banditrlproof.exp3.actionprocess_integral_mixedsquaredimportanceweightedloss_eq_integral_sum_loss_sq canonical generated-process mixed second-moment identity. theorem compiled","shard":"modules/cf33b341e5d9121d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.predictableLossAt_nonneg","label":"predictableLossAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.predictableLossAt_nonneg","description":"Every predictable comparator coordinate is nonnegative.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-081a31927705","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4004,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (action : Action) : 0 <= predictableLossAt loss t sample action","missing":[],"search":"predictablelossat_nonneg banditrlproof.exp3.predictablelossat_nonneg every predictable comparator coordinate is nonnegative. theorem compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_realizedLoss_le_one_ae","label":"sampledPredictableTrajectoryMeasure_realizedLoss_le_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_realizedLoss_le_one_ae","description":"Each generated realized scalar loss is at most one almost surely.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-2027dd6b0016","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4005,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_realizedLoss_le_one_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, sampledTrajectoryRealizedLossAt t sample <= 1","missing":[],"search":"sampledpredictabletrajectorymeasure_realizedloss_le_one_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_realizedloss_le_one_ae each generated realized scalar loss is at most one almost surely. theorem compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_finiteHorizon_realizedLoss_le_one_ae","label":"sampledPredictableTrajectoryMeasure_finiteHorizon_realizedLoss_le_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_finiteHorizon_realizedLoss_le_one_ae","description":"One common almost-sure event bounds every realized scalar loss strictly before a finite horizon.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-7d341015fdf0","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4006,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:59"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_finiteHorizon_realizedLoss_le_one_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, ∀ t, t < horizon -> sampledTrajectoryRealizedLossAt t sample <= 1","missing":[],"search":"sampledpredictabletrajectorymeasure_finitehorizon_realizedloss_le_one_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_finitehorizon_realizedloss_le_one_ae one common almost-sure event bounds every realized scalar loss strictly before a finite horizon. theorem compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedRegret_le_horizon_ae","label":"sampledPredictable_realizedRegret_le_horizon_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_realizedRegret_le_horizon_ae","description":"Generated selected-loss regret against any comparator is at most the horizon almost surely.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-11c41a2d361f","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4007,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:85"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_realizedRegret_le_horizon_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator) <= (horizon : Real)","missing":[],"search":"sampledpredictable_realizedregret_le_horizon_ae banditrlproof.exp3.sampledpredictable_realizedregret_le_horizon_ae generated selected-loss regret against any comparator is at most the horizon almost surely. theorem compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_trivialRealizedRegret_tail","label":"sampledPredictable_trivialRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_trivialRealizedRegret_tail","description":"The strict `T + 1` threshold has zero failure probability under the generated trajectory law.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-163c09428f44","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4008,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:130"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_trivialRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (horizon : Nat) (delta : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment mu {sample | (horizon : Real) + 1 <= (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)} <= ENNReal.ofReal delta","missing":[],"search":"sampledpredictable_trivialrealizedregret_tail banditrlproof.exp3.sampledpredictable_trivialrealizedregret_tail the strict `t + 1` threshold has zero failure probability under the generated trajectory law. theorem compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinLargeHorizonCondition","label":"bernsteinLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinLargeHorizonCondition","description":"The regime in which the clipped explicit schedule satisfies the Bernstein dominance contracts without activating its clip.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-f08ff33bfdd7","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4009,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:172"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def bernsteinLargeHorizonCondition (K T delta : Real) : Prop","missing":[],"search":"bernsteinlargehorizoncondition banditrlproof.exp3.bernsteinlargehorizoncondition the regime in which the clipped explicit schedule satisfies the bernstein dominance contracts without activating its clip. definition compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinAllHorizonRegretThreshold","label":"bernsteinAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinAllHorizonRegretThreshold","description":"All-horizon threshold: use the explicit Bernstein rate in its valid regime and the strict pathwise horizon fallback otherwise.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-d654a72131bb","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4010,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:180"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinAllHorizonRegretThreshold (K T delta : Real) : Real","missing":[],"search":"bernsteinallhorizonregretthreshold banditrlproof.exp3.bernsteinallhorizonregretthreshold all-horizon threshold: use the explicit bernstein rate in its valid regime and the strict pathwise horizon fallback otherwise. definition compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinRealizedRegret_tail","label":"sampledPredictable_allHorizonBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinRealizedRegret_tail","description":"Generated realized-regret tail for every positive horizon. In the explicit large-horizon regime the threshold is `11 * gamma * T`; otherwise the theorem uses the genuine almost-sure `T` regret bound and threshold `T + 1`.","url":"../modules/banditrlproof-exp3bernsteinallhorizon/index.html#decl-d8f6c2867a2c","parent":"module:BanditRLProof.Exp3BernsteinAllHorizon","order":4011,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinAllHorizon"],["Source","BanditRLProof/Exp3BernsteinAllHorizon.lean:191"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := bernsteinClippedExplorationRate (arms.card : Real) (horizon : Real) delta let eta := bernsteinHighProbabilityLearningRate (arms.card : Real) (horizon : Real) gamma let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma (bernsteinClippedExp…","missing":[],"search":"sampledpredictable_allhorizonbernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonbernsteinrealizedregret_tail generated realized-regret tail for every positive horizon. in the explicit large-horizon regime the threshold is `11 * gamma * t`; otherwise the theorem uses the genuine almost-sure `t` regret bound and threshold `t + 1`. theorem compiled","shard":"modules/5b5bbb53d9e8c6d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinArmEntropyExplorationScale","label":"bernsteinArmEntropyExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinArmEntropyExplorationScale","description":"Cube-root scale required by the arm-entropy term.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-e9ece48915e2","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4012,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinArmEntropyExplorationScale (K T : Real) : Real","missing":[],"search":"bernsteinarmentropyexplorationscale banditrlproof.exp3.bernsteinarmentropyexplorationscale cube-root scale required by the arm-entropy term. definition compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinConfidenceExplorationScale","label":"bernsteinConfidenceExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinConfidenceExplorationScale","description":"Cube-root scale required by both importance-weighted Bernstein radii.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-f93972a73ac0","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4013,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinConfidenceExplorationScale (K T delta : Real) : Real","missing":[],"search":"bernsteinconfidenceexplorationscale banditrlproof.exp3.bernsteinconfidenceexplorationscale cube-root scale required by both importance-weighted bernstein radii. definition compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinRealizedExplorationScale","label":"bernsteinRealizedExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinRealizedExplorationScale","description":"Square-root scale required by the bounded realized-deviation radius.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-7192d5e1207f","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4014,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:29"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinRealizedExplorationScale (T delta : Real) : Real","missing":[],"search":"bernsteinrealizedexplorationscale banditrlproof.exp3.bernsteinrealizedexplorationscale square-root scale required by the bounded realized-deviation radius. definition compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinRawExplorationRate","label":"bernsteinRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinRawExplorationRate","description":"Unclipped exploration scale simultaneously covering arm entropy, the two importance-weighted confidence radii, and the realized-deviation radius.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-9ceefbc75b15","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4015,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:37"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinRawExplorationRate (K T delta : Real) : Real","missing":[],"search":"bernsteinrawexplorationrate banditrlproof.exp3.bernsteinrawexplorationrate unclipped exploration scale simultaneously covering arm entropy, the two importance-weighted confidence radii, and the realized-deviation radius. definition compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate","label":"bernsteinClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinClippedExplorationRate","description":"Explicit exploration schedule, clipped into the stability regime `gamma <= 1 / 2`.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-2624a8532d10","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4016,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:45"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinClippedExplorationRate (K T delta : Real) : Real","missing":[],"search":"bernsteinclippedexplorationrate banditrlproof.exp3.bernsteinclippedexplorationrate explicit exploration schedule, clipped into the stability regime `gamma <= 1 / 2`. definition compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.rpow_inv_three_le_half_of_eight_mul_le","label":"rpow_inv_three_le_half_of_eight_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.rpow_inv_three_le_half_of_eight_mul_le","description":"A cube-root scale is at most one half when its numerator is at most one eighth of the positive horizon.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-dbe58445ee99","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4017,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:51"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem rpow_inv_three_le_half_of_eight_mul_le (numerator T : Real) (hnumerator : 0 <= numerator) (hT : 0 < T) (hlarge : 8 * numerator <= T) : (numerator / T) ^ (3 : Real)⁻¹ <= 1 / 2","missing":[],"search":"rpow_inv_three_le_half_of_eight_mul_le banditrlproof.exp3.rpow_inv_three_le_half_of_eight_mul_le a cube-root scale is at most one half when its numerator is at most one eighth of the positive horizon. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sqrt_div_le_half_of_four_mul_le","label":"sqrt_div_le_half_of_four_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sqrt_div_le_half_of_four_mul_le","description":"A square-root scale is at most one half when its numerator is at most one quarter of the positive horizon.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-afcf71ac77f1","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4018,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:67"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sqrt_div_le_half_of_four_mul_le (numerator T : Real) (hT : 0 < T) (hlarge : 4 * numerator <= T) : Real.sqrt (numerator / T) <= 1 / 2","missing":[],"search":"sqrt_div_le_half_of_four_mul_le banditrlproof.exp3.sqrt_div_le_half_of_four_mul_le a square-root scale is at most one half when its numerator is at most one quarter of the positive horizon. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.numerator_le_cube_mul_of_rpow_inv_three_le","label":"numerator_le_cube_mul_of_rpow_inv_three_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.numerator_le_cube_mul_of_rpow_inv_three_le","description":"If the cube-root scale is below `gamma`, then the corresponding numerator obeys the cubic dominance contract.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-2a530ec06c82","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4019,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:80"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem numerator_le_cube_mul_of_rpow_inv_three_le (numerator T gamma : Real) (hnumerator : 0 <= numerator) (hT : 0 < T) (hroot : (numerator / T) ^ (3 : Real)⁻¹ <= gamma) : numerator <= gamma ^ 3 * T","missing":[],"search":"numerator_le_cube_mul_of_rpow_inv_three_le banditrlproof.exp3.numerator_le_cube_mul_of_rpow_inv_three_le if the cube-root scale is below `gamma`, then the corresponding numerator obeys the cubic dominance contract. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.numerator_le_sq_mul_of_sqrt_div_le","label":"numerator_le_sq_mul_of_sqrt_div_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.numerator_le_sq_mul_of_sqrt_div_le","description":"If the square-root scale is below `gamma`, then the corresponding numerator obeys the quadratic dominance contract.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-6008f8202b3a","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4020,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:99"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem numerator_le_sq_mul_of_sqrt_div_le (numerator T gamma : Real) (hnumerator : 0 <= numerator) (hT : 0 < T) (hroot : Real.sqrt (numerator / T) <= gamma) : numerator <= gamma ^ 2 * T","missing":[],"search":"numerator_le_sq_mul_of_sqrt_div_le banditrlproof.exp3.numerator_le_sq_mul_of_sqrt_div_le if the square-root scale is below `gamma`, then the corresponding numerator obeys the quadratic dominance contract. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_le_half","label":"bernsteinClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinClippedExplorationRate_le_half","description":"The clipped schedule never exceeds the stability threshold.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-e4857d3aa29b","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4021,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:111"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinClippedExplorationRate_le_half (K T delta : Real) : bernsteinClippedExplorationRate K T delta <= 1 / 2","missing":[],"search":"bernsteinclippedexplorationrate_le_half banditrlproof.exp3.bernsteinclippedexplorationrate_le_half the clipped schedule never exceeds the stability threshold. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_eq_raw","label":"bernsteinClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinClippedExplorationRate_eq_raw","description":"If the raw schedule is already stable, clipping leaves it unchanged.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-44ca2edd4752","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4022,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:116"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinClippedExplorationRate_eq_raw (K T delta : Real) (hraw : bernsteinRawExplorationRate K T delta <= 1 / 2) : bernsteinClippedExplorationRate K T delta = bernsteinRawExplorationRate K T delta","missing":[],"search":"bernsteinclippedexplorationrate_eq_raw banditrlproof.exp3.bernsteinclippedexplorationrate_eq_raw if the raw schedule is already stable, clipping leaves it unchanged. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinRawExplorationRate_pos","label":"bernsteinRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinRawExplorationRate_pos","description":"The raw schedule is positive as soon as there are at least two arms and the horizon is positive.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-7c0cbaaf8bd7","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4023,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:125"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinRawExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < bernsteinRawExplorationRate K T delta","missing":[],"search":"bernsteinrawexplorationrate_pos banditrlproof.exp3.bernsteinrawexplorationrate_pos the raw schedule is positive as soon as there are at least two arms and the horizon is positive. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_pos","label":"bernsteinClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinClippedExplorationRate_pos","description":"The clipped schedule remains positive in the nondegenerate finite-arm, positive-horizon regime.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-32956a7110c3","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4024,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:137"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinClippedExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < bernsteinClippedExplorationRate K T delta","missing":[],"search":"bernsteinclippedexplorationrate_pos banditrlproof.exp3.bernsteinclippedexplorationrate_pos the clipped schedule remains positive in the nondegenerate finite-arm, positive-horizon regime. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinRawExplorationRate_le_half_of_horizon_contracts","label":"bernsteinRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinRawExplorationRate_le_half_of_horizon_contracts","description":"Three transparent horizon inequalities ensure that every raw component is at most one half, so the clipping branch is inactive.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-a1006005b75b","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4025,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:146"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinRawExplorationRate_le_half_of_horizon_contracts (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 8 * (K * Real.log K) <= T) (hlarge_confidence : 8 * (K * Real.log (3 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (3 / delta) <= T) : bernsteinRawExplorationRate K T delta <= 1 / 2","missing":[],"search":"bernsteinrawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.bernsteinrawexplorationrate_le_half_of_horizon_contracts three transparent horizon inequalities ensure that every raw component is at most one half, so the clipping branch is inactive. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_contracts","label":"bernsteinClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinClippedExplorationRate_contracts","description":"Under the large-horizon regime, the clipped schedule satisfies positivity, stability, both cubic dominance contracts, and the realized quadratic dominance contract required by the tuned tail theorem.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-da0fc26c91e5","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4026,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:181"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinClippedExplorationRate_contracts (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 8 * (K * Real.log K) <= T) (hlarge_confidence : 8 * (K * Real.log (3 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (3 / delta) <= T) : let gamma := bernsteinClippedExplorationRate K T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ K * Real.log K <= gamma ^ 3 * T ∧ K * Real.log (3 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (3 / delta) <= gamma ^ 2 * T","missing":[],"search":"bernsteinclippedexplorationrate_contracts banditrlproof.exp3.bernsteinclippedexplorationrate_contracts under the large-horizon regime, the clipped schedule satisfies positivity, stability, both cubic dominance contracts, and the realized quadratic dominance contract required by the tuned tail theorem. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitBernsteinRealizedHighProbabilityRegret_tail","label":"sampledPredictable_explicitBernsteinRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitBernsteinRealizedHighProbabilityRegret_tail","description":"Generated realized-regret tail for the explicit clipped maximum of the two cube-root Bernstein scales and the realized square-root scale. The three large-horizon premises are sufficient conditions ensuring that clipping is inactive and all contracts of the characterized `11 * gamma * T` theorem hold.","url":"../modules/banditrlproof-exp3bernsteinexplicittuning/index.html#decl-57949c47292e","parent":"module:BanditRLProof.Exp3BernsteinExplicitTuning","order":4027,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinExplicitTuning"],["Source","BanditRLProof/Exp3BernsteinExplicitTuning.lean:250"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitBernsteinRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 8 * ((arms.card : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_confidence : 8 * ((arms.card : Real) * Real.log (3 / delta)) <= (horizon : Real)) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Rea…","missing":[],"search":"sampledpredictable_explicitbernsteinrealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_explicitbernsteinrealizedhighprobabilityregret_tail generated realized-regret tail for the explicit clipped maximum of the two cube-root bernstein scales and the realized square-root scale. the three large-horizon premises are sufficient conditions ensuring that clipping is inactive and all contracts of the characterized `11 * gamma * t` theorem hold. theorem compiled","shard":"modules/254fb3c4a0a0b351.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinHighProbabilityRegretBudget","label":"sampledPredictableBernsteinHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBernsteinHighProbabilityRegretBudget","description":"Predictable generated-regret budget with variance-sensitive confidence radii and the existing pathwise estimator-square contribution.","url":"../modules/banditrlproof-exp3bernsteinhighprobabilityregret/index.html#decl-c869c42b05e0","parent":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","order":4028,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinHighProbabilityRegret"],["Source","BanditRLProof/Exp3BernsteinHighProbabilityRegret.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableBernsteinHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablebernsteinhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablebernsteinhighprobabilityregretbudget predictable generated-regret budget with variance-sensitive confidence radii and the existing pathwise estimator-square contribution. definition compiled","shard":"modules/c7ceb5537c8e4d02.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinHighProbabilityRegret_tail_delta","label":"sampledPredictable_bernsteinHighProbabilityRegret_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinHighProbabilityRegret_tail_delta","description":"Generated predictable EXP3 regret with both importance-weighted confidence events using their variance-sensitive fixed-tilt routes.","url":"../modules/banditrlproof-exp3bernsteinhighprobabilityregret/index.html#decl-00f3675a4e21","parent":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","order":4029,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinHighProbabilityRegret"],["Source","BanditRLProof/Exp3BernsteinHighProbabilityRegret.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinHighProbabilityRegret_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinHighProbabilityRegretBudget arms eta gamma horizon delta <= (Finset.range horizon).sum (fun t => sam…","missing":[],"search":"sampledpredictable_bernsteinhighprobabilityregret_tail_delta banditrlproof.exp3.sampledpredictable_bernsteinhighprobabilityregret_tail_delta generated predictable exp3 regret with both importance-weighted confidence events using their variance-sensitive fixed-tilt routes. theorem compiled","shard":"modules/c7ceb5537c8e4d02.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_bernsteinHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinHighProbabilityRegret_tail_total_delta","description":"Total-failure form: each variance-sensitive confidence event receives `delta / 2`. Unlike the range-Hoeffding predecessor, no positive-horizon premise is needed.","url":"../modules/banditrlproof-exp3bernsteinhighprobabilityregret/index.html#decl-fe40ff1b0b5a","parent":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","order":4030,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinHighProbabilityRegret"],["Source","BanditRLProof/Exp3BernsteinHighProbabilityRegret.lean:183"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinHighProbabilityRegretBudget arms eta gamma horizon (delta / 2) <= (Finset.range horizon).sum (…","missing":[],"search":"sampledpredictable_bernsteinhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_bernsteinhighprobabilityregret_tail_total_delta total-failure form: each variance-sensitive confidence event receives `delta / 2`. unlike the range-hoeffding predecessor, no positive-horizon premise is needed. theorem compiled","shard":"modules/c7ceb5537c8e4d02.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinRealizedHighProbabilityRegretBudget","label":"sampledPredictableBernsteinRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBernsteinRealizedHighProbabilityRegretBudget","description":"Realized selected-loss regret budget whose predictable component uses the two variance-sensitive Bernstein confidence radii.","url":"../modules/banditrlproof-exp3bernsteinrealizedhighprobabilityregret/index.html#decl-9cc6a8b36389","parent":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","order":4031,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3BernsteinRealizedHighProbabilityRegret.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableBernsteinRealizedHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablebernsteinrealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablebernsteinrealizedhighprobabilityregretbudget realized selected-loss regret budget whose predictable component uses the two variance-sensitive bernstein confidence radii. definition compiled","shard":"modules/51275df4f4321a21.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_delta","label":"sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_delta","description":"Raw three-event form. The predictable component contributes the pure-cross and fixed-comparator Bernstein events, while the third event is the bounded realized-minus-predictable deviation.","url":"../modules/banditrlproof-exp3bernsteinrealizedhighprobabilityregret/index.html#decl-81766b981b35","parent":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","order":4032,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3BernsteinRealizedHighProbabilityRegret.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinRealizedHighProbabilityRegretBudget arms eta gamma horizon delta <=…","missing":[],"search":"sampledpredictable_bernsteinrealizedhighprobabilityregret_tail_delta banditrlproof.exp3.sampledpredictable_bernsteinrealizedhighprobabilityregret_tail_delta raw three-event form. the predictable component contributes the pure-cross and fixed-comparator bernstein events, while the third event is the bounded realized-minus-predictable deviation. theorem compiled","shard":"modules/51275df4f4321a21.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_total_delta","description":"Total-failure form: the pure-cross Bernstein, fixed-comparator Bernstein, and realized-deviation events each receive `delta / 3`.","url":"../modules/banditrlproof-exp3bernsteinrealizedhighprobabilityregret/index.html#decl-3822d00e99a4","parent":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","order":4033,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3BernsteinRealizedHighProbabilityRegret.lean:148"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinRealizedHighProbabilityRegretBudget arms eta gamma horizon (d…","missing":[],"search":"sampledpredictable_bernsteinrealizedhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_bernsteinrealizedhighprobabilityregret_tail_total_delta total-failure form: the pure-cross bernstein, fixed-comparator bernstein, and realized-deviation events each receive `delta / 3`. theorem compiled","shard":"modules/51275df4f4321a21.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate","label":"bernsteinHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate","description":"Learning rate balancing the entropy and pathwise estimator-square terms for the current high-probability budget.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-0df7090c013c","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4034,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinHighProbabilityLearningRate (K T gamma : Real) : Real","missing":[],"search":"bernsteinhighprobabilitylearningrate banditrlproof.exp3.bernsteinhighprobabilitylearningrate learning rate balancing the entropy and pathwise estimator-square terms for the current high-probability budget. definition compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinConfidenceRadius_le_three_mul_gamma_mul_horizon","label":"bernsteinConfidenceRadius_le_three_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinConfidenceRadius_le_three_mul_gamma_mul_horizon","description":"A cubic exploration budget makes one current Bernstein confidence radius at most `3 * gamma * T`.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-ab89d81f51d2","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4035,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinConfidenceRadius_le_three_mul_gamma_mul_horizon (K T budget gamma : Real) (hK : 0 < K) (hT : 0 < T) (_hbudget : 0 <= budget) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (hcubic : K * budget <= gamma ^ 3 * T) : 2 * Real.sqrt (T * budget / (gamma / K)) + budget / (gamma / K) <= 3 * gamma * T","missing":[],"search":"bernsteinconfidenceradius_le_three_mul_gamma_mul_horizon banditrlproof.exp3.bernsteinconfidenceradius_le_three_mul_gamma_mul_horizon a cubic exploration budget makes one current bernstein confidence radius at most `3 * gamma * t`. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.realizedDeviationRadius_le_mul_gamma_mul_horizon","label":"realizedDeviationRadius_le_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.realizedDeviationRadius_le_mul_gamma_mul_horizon","description":"The bounded realized-loss deviation radius is at most `gamma * T` under its matching quadratic exploration budget.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-a8c6519d3e4e","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4036,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:69"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem realizedDeviationRadius_le_mul_gamma_mul_horizon (T budget variance gamma : Real) (hT : 0 < T) (_hbudget : 0 <= budget) (_hvariance : 0 <= variance) (hgamma_pos : 0 < gamma) (hquadratic : 2 * variance * budget <= gamma ^ 2 * T) : Real.sqrt (2 * (T * variance) * budget) <= gamma * T","missing":[],"search":"realizeddeviationradius_le_mul_gamma_mul_horizon banditrlproof.exp3.realizeddeviationradius_le_mul_gamma_mul_horizon the bounded realized-loss deviation radius is at most `gamma * t` under its matching quadratic exploration budget. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_pos","label":"bernsteinHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_pos","description":"The tuned learning rate is positive in the nondegenerate finite-arm, positive-horizon regime.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-f3956460b141","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4037,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:83"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinHighProbabilityLearningRate_pos (K T gamma : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) : 0 < bernsteinHighProbabilityLearningRate K T gamma","missing":[],"search":"bernsteinhighprobabilitylearningrate_pos banditrlproof.exp3.bernsteinhighprobabilitylearningrate_pos the tuned learning rate is positive in the nondegenerate finite-arm, positive-horizon regime. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_sq_mul","label":"bernsteinHighProbabilityLearningRate_sq_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_sq_mul","description":"Squared balance identity for the tuned learning rate.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-7788a4bb6e7d","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4038,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:92"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinHighProbabilityLearningRate_sq_mul (K T gamma : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) : bernsteinHighProbabilityLearningRate K T gamma ^ 2 * (T * K) = Real.log K * gamma","missing":[],"search":"bernsteinhighprobabilitylearningrate_sq_mul banditrlproof.exp3.bernsteinhighprobabilitylearningrate_sq_mul squared balance identity for the tuned learning rate. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_le_sq_div","label":"bernsteinHighProbabilityLearningRate_le_sq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_le_sq_div","description":"The cubic exploration contract places the tuned learning rate below the scale `gamma ^ 2 / K`.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-baafbedbf009","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4039,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:105"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinHighProbabilityLearningRate_le_sq_div (K T gamma : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) (hcubic_log : K * Real.log K <= gamma ^ 3 * T) : bernsteinHighProbabilityLearningRate K T gamma <= gamma ^ 2 / K","missing":[],"search":"bernsteinhighprobabilitylearningrate_le_sq_div banditrlproof.exp3.bernsteinhighprobabilitylearningrate_le_sq_div the cubic exploration contract places the tuned learning rate below the scale `gamma ^ 2 / k`. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinEntropyBudget_le_mul_gamma_mul_horizon","label":"bernsteinEntropyBudget_le_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinEntropyBudget_le_mul_gamma_mul_horizon","description":"The entropy contribution is at most `gamma * T` under the cubic arm-log budget.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-80a6ca4f4d7e","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4040,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:126"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinEntropyBudget_le_mul_gamma_mul_horizon (K T gamma : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) (hcubic_log : K * Real.log K <= gamma ^ 3 * T) : Real.log K / bernsteinHighProbabilityLearningRate K T gamma <= gamma * T","missing":[],"search":"bernsteinentropybudget_le_mul_gamma_mul_horizon banditrlproof.exp3.bernsteinentropybudget_le_mul_gamma_mul_horizon the entropy contribution is at most `gamma * t` under the cubic arm-log budget. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinUnscaledSquareBudget_le_mul_gamma_mul_horizon","label":"bernsteinUnscaledSquareBudget_le_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinUnscaledSquareBudget_le_mul_gamma_mul_horizon","description":"Before the stability factor `1 / (1 - gamma)`, the pathwise square term has the same tuned scale as the entropy term.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-7393ac273e2f","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4041,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:157"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinUnscaledSquareBudget_le_mul_gamma_mul_horizon (K T gamma : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) (hcubic_log : K * Real.log K <= gamma ^ 3 * T) : bernsteinHighProbabilityLearningRate K T gamma * (T * (1 / (gamma / K))) <= gamma * T","missing":[],"search":"bernsteinunscaledsquarebudget_le_mul_gamma_mul_horizon banditrlproof.exp3.bernsteinunscaledsquarebudget_le_mul_gamma_mul_horizon before the stability factor `1 / (1 - gamma)`, the pathwise square term has the same tuned scale as the entropy term. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinHedgeBudget_le_three_mul_gamma_mul_horizon","label":"bernsteinHedgeBudget_le_three_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinHedgeBudget_le_three_mul_gamma_mul_horizon","description":"The entropy and pathwise square terms together cost at most `3 * gamma * T` when `gamma <= 1 / 2`.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-4be32865ac2f","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4042,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:182"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinHedgeBudget_le_three_mul_gamma_mul_horizon (K T gamma : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hcubic_log : K * Real.log K <= gamma ^ 3 * T) : Real.log K / bernsteinHighProbabilityLearningRate K T gamma + (bernsteinHighProbabilityLearningRate K T gamma * (1 / (1 - gamma))) * (T * (1 / (gamma / K))) <= 3 * gamma * T","missing":[],"search":"bernsteinhedgebudget_le_three_mul_gamma_mul_horizon banditrlproof.exp3.bernsteinhedgebudget_le_three_mul_gamma_mul_horizon the entropy and pathwise square terms together cost at most `3 * gamma * t` when `gamma <= 1 / 2`. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_one_div_third_eq_log_three_div","label":"log_one_div_third_eq_log_three_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_one_div_third_eq_log_three_div","description":"Dividing a total failure probability by three changes the logarithmic budget to `log (3 / delta)`.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-db39f1a62e91","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4043,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:218"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_one_div_third_eq_log_three_div (delta : Real) (hdelta : 0 < delta) : Real.log (1 / (delta / 3)) = Real.log (3 / delta)","missing":[],"search":"log_one_div_third_eq_log_three_div banditrlproof.exp3.log_one_div_third_eq_log_three_div dividing a total failure probability by three changes the logarithmic budget to `log (3 / delta)`. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinRealizedHighProbabilityRegretBudget_le_eleven_mul","label":"sampledPredictableBernsteinRealizedHighProbabilityRegretBudget_le_eleven_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBernsteinRealizedHighProbabilityRegretBudget_le_eleven_mul","description":"Under the explicit cubic and quadratic dominance contracts, the complete three-event realized Bernstein budget is at most `11 * gamma * T`.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-4f68cc1157a4","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4044,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:225"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableBernsteinRealizedHighProbabilityRegretBudget_le_eleven_mul {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon : Nat) (hhorizon : 0 < horizon) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcubic_log : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 3 * (horizon : Real)) (hcubic_confidence : (arms.card : Real) * Real.log (3 / delta) <= gamma ^ 3 * (horizon : Real)) (hquadratic_realized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (3 / delta) <= gamma ^ 2 * (horizon : Real)) : sampledPredictableBernsteinRealizedHighProbabilityRegretBudget arms (bernsteinHighProbabilityLearningRate (arms.card : Real) (horizon : Real) gamma) gamma horizon (delta / 3) <= 11 * gamma * (horizon : Real)","missing":[],"search":"sampledpredictablebernsteinrealizedhighprobabilityregretbudget_le_eleven_mul banditrlproof.exp3.sampledpredictablebernsteinrealizedhighprobabilityregretbudget_le_eleven_mul under the explicit cubic and quadratic dominance contracts, the complete three-event realized bernstein budget is at most `11 * gamma * t`. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedBernsteinRealizedHighProbabilityRegret_tail","label":"sampledPredictable_tunedBernsteinRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedBernsteinRealizedHighProbabilityRegret_tail","description":"Tuned generated realized-regret tail with an explicit `11 * gamma * T` threshold. The current confidence route yields a `T^(2/3)`-type contract: the arm entropy and both importance-weighted confidence budgets must be dominated by `gamma ^ 3 * T`, while the bounded realized deviation uses the displayed quadratic contract.","url":"../modules/banditrlproof-exp3bernsteintuning/index.html#decl-830e002ad16b","parent":"module:BanditRLProof.Exp3BernsteinTuning","order":4045,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BernsteinTuning"],["Source","BanditRLProof/Exp3BernsteinTuning.lean:316"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedBernsteinRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcubic_log : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 3 * (horizon : Real)) (hcubic_confidence : (arms.card : Real) * Real.log (3 / delta) <= gamma ^ 3 * (horizon : Real)) (hqua…","missing":[],"search":"sampledpredictable_tunedbernsteinrealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_tunedbernsteinrealizedhighprobabilityregret_tail tuned generated realized-regret tail with an explicit `11 * gamma * t` threshold. the current confidence route yields a `t^(2/3)`-type contract: the arm entropy and both importance-weighted confidence budgets must be dominated by `gamma ^ 3 * t`, while the bounded realized deviation uses the displayed quadratic contract. theorem compiled","shard":"modules/232ae82a12ecc628.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBestArmCumulativeLoss","label":"sampledPredictableBestArmCumulativeLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBestArmCumulativeLoss","description":"Cumulative predictable loss of the best supported arm in hindsight.","url":"../modules/banditrlproof-exp3bestarm/index.html#decl-e5718bd0fdab","parent":"module:BanditRLProof.Exp3BestArm","order":4046,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3BestArm"],["Source","BanditRLProof/Exp3BestArm.lean:17"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableBestArmCumulativeLoss {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (harms : arms.Nonempty) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : Real","missing":[],"search":"sampledpredictablebestarmcumulativeloss banditrlproof.exp3.sampledpredictablebestarmcumulativeloss cumulative predictable loss of the best supported arm in hindsight. definition compiled","shard":"modules/0d63c9d1d6b80df1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.threshold_le_sampledPredictableRealizedLoss_sub_bestArmCumulativeLoss_iff","label":"threshold_le_sampledPredictableRealizedLoss_sub_bestArmCumulativeLoss_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.threshold_le_sampledPredictableRealizedLoss_sub_bestArmCumulativeLoss_iff","description":"The best-arm regret event is exactly the finite existential union of the fixed-comparator regret events.","url":"../modules/banditrlproof-exp3bestarm/index.html#decl-2151bd69dc8a","parent":"module:BanditRLProof.Exp3BestArm","order":4047,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3BestArm"],["Source","BanditRLProof/Exp3BestArm.lean:29"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem threshold_le_sampledPredictableRealizedLoss_sub_bestArmCumulativeLoss_iff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (harms : arms.Nonempty) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) (threshold : Real) : threshold <= (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - sampledPredictableBestArmCumulativeLoss arms harms loss horizon sample ↔ ∃ comparator ∈ arms, threshold <= (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)","missing":[],"search":"threshold_le_sampledpredictablerealizedloss_sub_bestarmcumulativeloss_iff banditrlproof.exp3.threshold_le_sampledpredictablerealizedloss_sub_bestarmcumulativeloss_iff the best-arm regret event is exactly the finite existential union of the fixed-comparator regret events. theorem compiled","shard":"modules/0d63c9d1d6b80df1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.exp_le_one_add_self_add_sq_of_abs_le_one","label":"exp_le_one_add_self_add_sq_of_abs_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.exp_le_one_add_self_add_sq_of_abs_le_one","description":"On `[-1, 1]`, the exponential remainder is bounded by the square.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-a869a6027216","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4048,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exp_le_one_add_self_add_sq_of_abs_le_one {x : Real} (hx : |x| <= 1) : Real.exp x <= 1 + x + x ^ 2","missing":[],"search":"exp_le_one_add_self_add_sq_of_abs_le_one banditrlproof.concentration.exp_le_one_add_self_add_sq_of_abs_le_one on `[-1, 1]`, the exponential remainder is bounded by the square. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.exists_tilt_fixedMGF_exponent_le_neg","label":"exists_tilt_fixedMGF_exponent_le_neg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.exists_tilt_fixedMGF_exponent_le_neg","description":"Optimize a quadratic fixed-tilt MGF budget under the hard constraint `tilt <= epsilon`.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-20e738992703","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4049,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exists_tilt_fixedMGF_exponent_le_neg (horizon epsilon budget : Real) (hhorizon : 0 <= horizon) (hepsilon : 0 < epsilon) (hbudget : 0 <= budget) : exists tilt : Real, 0 <= tilt ∧ tilt <= epsilon ∧ -tilt * (2 * Real.sqrt (horizon * budget / epsilon) + budget / epsilon) + horizon * (tilt ^ 2 / epsilon) <= -budget","missing":[],"search":"exists_tilt_fixedmgf_exponent_le_neg banditrlproof.concentration.exists_tilt_fixedmgf_exponent_le_neg optimize a quadratic fixed-tilt mgf budget under the hard constraint `tilt <= epsilon`. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_comparatorEstimatorDeviation_eq","label":"sum_prob_mul_sq_comparatorEstimatorDeviation_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_sq_comparatorEstimatorDeviation_eq","description":"Exact centered second moment of one fixed-arm importance-weighted estimator.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-ff37c2c809d2","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4050,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:115"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_sq_comparatorEstimatorDeviation_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hdist : FiniteActionDistribution arms prob) (comparator : Action) (hcomparator : comparator ∈ arms) (hprob : prob comparator ≠ 0) : arms.sum (fun chosen => prob chosen * (importanceWeightedLoss prob loss chosen comparator - loss comparator) ^ 2) = (loss comparator) ^ 2 / prob comparator - (loss comparator) ^ 2","missing":[],"search":"sum_prob_mul_sq_comparatorestimatordeviation_eq banditrlproof.exp3.sum_prob_mul_sq_comparatorestimatordeviation_eq exact centered second moment of one fixed-arm importance-weighted estimator. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_comparatorEstimatorDeviation_le_inv_floor","label":"sum_prob_mul_sq_comparatorEstimatorDeviation_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_sq_comparatorEstimatorDeviation_le_inv_floor","description":"The centered comparator estimator has second moment at most the reciprocal probability floor.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-f783c74953b9","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4051,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:151"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_sq_comparatorEstimatorDeviation_le_inv_floor {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hdist : FiniteActionDistribution arms prob) (epsilon : Real) (hepsilon : 0 < epsilon) (comparator : Action) (hcomparator : comparator ∈ arms) (hfloor : epsilon <= prob comparator) (hloss : loss comparator ∈ Set.Icc (0 : Real) 1) : arms.sum (fun chosen => prob chosen * (importanceWeightedLoss prob loss chosen comparator - loss comparator) ^ 2) <= 1 / epsilon","missing":[],"search":"sum_prob_mul_sq_comparatorestimatordeviation_le_inv_floor banditrlproof.exp3.sum_prob_mul_sq_comparatorestimatordeviation_le_inv_floor the centered comparator estimator has second moment at most the reciprocal probability floor. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionComparatorEstimator_hasMGFUpperBoundAt","label":"finiteActionComparatorEstimator_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionComparatorEstimator_hasMGFUpperBoundAt","description":"Fixed-tilt MGF budget for a finite-law fixed-comparator estimator.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-88a0b9237b26","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4052,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:176"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionComparatorEstimator_hasMGFUpperBoundAt {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hdist : FiniteActionDistribution arms prob) (epsilon : Real) (hepsilon : 0 < epsilon) (comparator : Action) (hcomparator : comparator ∈ arms) (hfloor : epsilon <= prob comparator) (hloss : loss comparator ∈ Set.Icc (0 : Real) 1) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) : Concentration.HasMGFUpperBoundAt (fun chosen => importanceWeightedLoss prob loss chosen comparator - loss comparator) tilt (tilt ^ 2 / epsilon) (finiteActionMeasure arms prob)","missing":[],"search":"finiteactioncomparatorestimator_hasmgfupperboundat banditrlproof.exp3.finiteactioncomparatorestimator_hasmgfupperboundat fixed-tilt mgf budget for a finite-law fixed-comparator estimator. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.comparatorEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","label":"comparatorEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.comparatorEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","description":"A finite conditional action law supplies the variance-sensitive fixed-tilt MGF budget for one fixed comparator estimator.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-1a24f69bcf07","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4053,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:283"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem comparatorEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (comparator : Action) (hcomparator : comparator ∈ arms) (tilt : Real) (htilt_nonneg : 0 <= tilt) (ht…","missing":[],"search":"comparatorestimator_hascondmgfupperboundat_of_conddistrib_ae_eq_finiteactionkernel banditrlproof.exp3.comparatorestimator_hascondmgfupperboundat_of_conddistrib_ae_eq_finiteactionkernel a finite conditional action law supplies the variance-sensitive fixed-tilt mgf budget for one fixed comparator estimator. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","label":"sampledPredictableComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","description":"theorem sampledPredictableComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hg…","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-e006952e0eb6","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4054,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:421"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat)…","missing":[],"search":"sampledpredictablecomparatorestimatordeviation_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablecomparatorestimatordeviation_zero_hascondmgfupperboundat theorem sampledpredictablecomparatorestimatordeviation_zero_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment concentration.hascondmgfupperboundat ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypredictablecomparatorestimatordeviationat arms eta gamma loss comparator 0) tilt (tilt ^ 2 / (gamma / (arms.card : real))) mu theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","label":"sampledPredictableComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","description":"theorem sampledPredictableComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hg…","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-4ebad783fc37","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4055,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:478"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (n : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sam…","missing":[],"search":"sampledpredictablecomparatorestimatordeviation_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablecomparatorestimatordeviation_succ_hascondmgfupperboundat theorem sampledpredictablecomparatorestimatordeviation_succ_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (n : nat) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) concentration.hascondmgfupperboundat ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypredictablecomparatorestimatordeviationat arms eta gamma loss comparator (n + 1)…","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","label":"sampledObservedComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","description":"theorem sampledObservedComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamm…","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-f9780deb1594","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4056,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:535"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) ->…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_zero_hascondmgfupperboundat banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_zero_hascondmgfupperboundat theorem sampledobservedcomparatorestimatordeviation_zero_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment concentration.hascondmgfupperboundat ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator 0) tilt (tilt ^ 2 / (gamma / (arms.card : real))) mu theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","label":"sampledObservedComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","description":"theorem sampledObservedComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamm…","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-bc1f561c2058","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4057,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:600"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (n : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_succ_hascondmgfupperboundat banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_succ_hascondmgfupperboundat theorem sampledobservedcomparatorestimatordeviation_succ_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (n : nat) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) concentration.hascondmgfupperboundat ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator (n + 1)) tilt (tilt…","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_fixedTilt","label":"sampledObservedComparatorEstimatorDeviation_sum_tail_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_fixedTilt","description":"Variance-sensitive finite-horizon Chernoff bound for one fixed comparator on the generated EXP3 trajectory. The per-round budget is linear, rather than quadratic, in the reciprocal exploration floor.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-b87c3fc127f8","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4058,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:674"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_sum_tail_fixedTilt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) (threshold : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu.real {sample | threshold <= (Finset.range horizon).sum (fun i => sampledTrajectory…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_sum_tail_fixedtilt banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_sum_tail_fixedtilt variance-sensitive finite-horizon chernoff bound for one fixed comparator on the generated exp3 trajectory. the per-round budget is linear, rather than quadratic, in the reciprocal exploration floor. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorBernsteinConfidenceRadius","label":"sampledComparatorEstimatorBernsteinConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledComparatorEstimatorBernsteinConfidenceRadius","description":"Variance-sensitive confidence radius obtained by optimizing the fixed-tilt comparator tail.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-2cece2bcc685","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4059,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:752"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledComparatorEstimatorBernsteinConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledcomparatorestimatorbernsteinconfidenceradius banditrlproof.exp3.sampledcomparatorestimatorbernsteinconfidenceradius variance-sensitive confidence radius obtained by optimizing the fixed-tilt comparator tail. definition compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_bernstein_delta","label":"sampledObservedComparatorEstimatorDeviation_sum_tail_bernstein_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_bernstein_delta","description":"Delta-shaped variance-sensitive confidence bound for one fixed comparator's observed importance-weighted estimator on the generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3comparatorbernstein/index.html#decl-f0a8b8161dd9","parent":"module:BanditRLProof.Exp3ComparatorBernstein","order":4060,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorBernstein"],["Source","BanditRLProof/Exp3ComparatorBernstein.lean:762"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_sum_tail_bernstein_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledComparatorEstimatorBernsteinConfidenceRadius arms gamma horizon delta <= (Finset.range horizon).sum (fun i => sampledTrajectoryObse…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_sum_tail_bernstein_delta banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_sum_tail_bernstein_delta delta-shaped variance-sensitive confidence bound for one fixed comparator's observed importance-weighted estimator on the generated exp3 trajectory. theorem compiled","shard":"modules/96ce3115ea0c36d6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.comparatorEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","label":"comparatorEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.comparatorEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","description":"theorem comparatorEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Acti…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-f4c7fd71e01e","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4061,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:9"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem comparatorEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (comparator : Action) (hcomparator : comparator ∈ arms) (hcond : condDistrib action history mu =ᵐ[mu…","missing":[],"search":"comparatorestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel banditrlproof.exp3.comparatorestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel theorem comparatorestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel {omega : type u} {history : type v} {action : type w} [momega : measurablespace omega] [standardborelspace omega] [nonempty omega] [mhistory : measurablespace history] [standardborelspace history] [maction : measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (mu : measure omega) [isfinitemeasure mu] (history : omega -> history) (hhistory : measurable history) (action : omega -> action) (haction : measurable action) (arms : finset action) (prob loss : history -> action -> real) (source : measurablefiniteactiondistribution arms prob) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) (comparator : action) (hcomparator : comparator ∈ arms) (hcond : conddistrib action history mu =ᵐ[mu.map history] finiteactionkernel arms prob source) : probabilitytheory.hascondsubgaussianmgf (mhistory.comap history) hhistory.comap_le (fun omega => importanceweightedloss (prob (history omega)) (loss (history omega)) (action omega) comparator - loss (history omega) comparator) (concentration.intervalvarianceproxy…","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableComparatorEstimatorDeviationAt","label":"sampledTrajectoryPredictableComparatorEstimatorDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableComparatorEstimatorDeviationAt","description":"noncomputable def sampledTrajectoryPredictableComparatorEstimatorDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-16c8433472ee","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4062,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:159"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPredictableComparatorEstimatorDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypredictablecomparatorestimatordeviationat banditrlproof.exp3.sampledtrajectorypredictablecomparatorestimatordeviationat noncomputable def sampledtrajectorypredictablecomparatorestimatordeviationat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (comparator : action) (t : nat) (sample : env × ((k : nat) -> action × real)) : real definition compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryObservedComparatorEstimatorDeviationAt","label":"sampledTrajectoryObservedComparatorEstimatorDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryObservedComparatorEstimatorDeviationAt","description":"noncomputable def sampledTrajectoryObservedComparatorEstimatorDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-30502d80eaaf","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4063,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:170"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryObservedComparatorEstimatorDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectoryobservedcomparatorestimatordeviationat banditrlproof.exp3.sampledtrajectoryobservedcomparatorestimatordeviationat noncomputable def sampledtrajectoryobservedcomparatorestimatordeviationat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (comparator : action) (t : nat) (sample : env × ((k : nat) -> action × real)) : real definition compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","label":"sampledPredictableComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledPredictableComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hga…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-562dee49fed1","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4064,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:179"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryPredictableCo…","missing":[],"search":"sampledpredictablecomparatorestimatordeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictablecomparatorestimatordeviation_zero_hascondsubgaussianmgf theorem sampledpredictablecomparatorestimatordeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypredictablecomparatorestimatordeviationat arms eta gamma loss comparator 0) (concentration.intervalvarianceproxy 0 (1 / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","label":"sampledPredictableComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledPredictableComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hga…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-5991f75c8998","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4065,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:234"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × Hi…","missing":[],"search":"sampledpredictablecomparatorestimatordeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictablecomparatorestimatordeviation_succ_hascondsubgaussianmgf theorem sampledpredictablecomparatorestimatordeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypredictablecomparatorestimatordeviationat arms eta gamma loss comparator (n + 1)) (concentration.intervalvarianceproxy 0 (1 / (gamma / (arms.card : real)))) mu theorem c…","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","label":"sampledObservedComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledObservedComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-b74a6d1ab2dc","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4066,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:289"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryObservedComparat…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_zero_hascondsubgaussianmgf theorem sampledobservedcomparatorestimatordeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator 0) (concentration.intervalvarianceproxy 0 (1 / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","label":"sampledObservedComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledObservedComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-79bcadecd323","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4067,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:354"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × Histo…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_succ_hascondsubgaussianmgf theorem sampledobservedcomparatorestimatordeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator (n + 1)) (concentration.intervalvarianceproxy 0 (1 / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess","label":"sampledObservedComparatorEstimatorDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess","description":"noncomputable def sampledObservedComparatorEstimatorDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryObservedComparatorEstimatorDeviationA…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-220cd64d60fa","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4068,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:426"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledObservedComparatorEstimatorDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryObservedComparatorEstimatorDeviationAt arms eta gamma loss comparator i sample noncomputable def sampledComparatorEstimatorVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","missing":[],"search":"sampledobservedcomparatorestimatordeviationprocess banditrlproof.exp3.sampledobservedcomparatorestimatordeviationprocess noncomputable def sampledobservedcomparatorestimatordeviationprocess {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (comparator : action) : nat -> env × ((k : nat) -> action × real) -> real | 0, _sample => 0 | i + 1, sample => sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator i sample noncomputable def sampledcomparatorestimatorvarianceproxy {action : type v} (arms : finset action) (gamma : real) : nnreal definition compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorVarianceProxy","label":"sampledComparatorEstimatorVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledComparatorEstimatorVarianceProxy","description":"noncomputable def sampledComparatorEstimatorVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-049f3a1ea232","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4069,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:437"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledComparatorEstimatorVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","missing":[],"search":"sampledcomparatorestimatorvarianceproxy banditrlproof.exp3.sampledcomparatorestimatorvarianceproxy noncomputable def sampledcomparatorestimatorvarianceproxy {action : type v} (arms : finset action) (gamma : real) : nnreal definition compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorDeviationProxy","label":"sampledComparatorEstimatorDeviationProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledComparatorEstimatorDeviationProxy","description":"noncomputable def sampledComparatorEstimatorDeviationProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : Nat -> NNReal | 0 => 0 | _i + 1 => sampledComparatorEstimatorVarianceProxy arms gamma theorem sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-f80a78415eb3","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4070,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:442"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledComparatorEstimatorDeviationProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : Nat -> NNReal | 0 => 0 | _i + 1 => sampledComparatorEstimatorVarianceProxy arms gamma theorem sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledObservedComparatorEstimatorDeviationProcess arms eta gamma loss comparator)","missing":[],"search":"sampledcomparatorestimatordeviationproxy banditrlproof.exp3.sampledcomparatorestimatordeviationproxy noncomputable def sampledcomparatorestimatordeviationproxy {action : type v} (arms : finset action) (gamma : real) : nat -> nnreal | 0 => 0 | _i + 1 => sampledcomparatorestimatorvarianceproxy arms gamma theorem sampledobservedcomparatorestimatordeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledobservedcomparatorestimatordeviationprocess arms eta gamma loss comparator) definition compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted","label":"sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted","description":"theorem sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : compar…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-36aacb0faee6","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4071,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:447"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledObservedComparatorEstimatorDeviationProcess arms eta gamma loss comparator)","missing":[],"search":"sampledobservedcomparatorestimatordeviationprocess_stronglyadapted banditrlproof.exp3.sampledobservedcomparatorestimatordeviationprocess_stronglyadapted theorem sampledobservedcomparatorestimatordeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledobservedcomparatorestimatordeviationprocess arms eta gamma loss comparator) theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess_sum_range_succ","label":"sampledObservedComparatorEstimatorDeviationProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess_sum_range_succ","description":"theorem sampledObservedComparatorEstimatorDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledObservedComparatorEstima…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-99ff7a029939","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4072,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:600"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledObservedComparatorEstimatorDeviationProcess arms eta gamma loss comparator i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryObservedComparatorEstimatorDeviationAt arms eta gamma loss comparator i sample)","missing":[],"search":"sampledobservedcomparatorestimatordeviationprocess_sum_range_succ banditrlproof.exp3.sampledobservedcomparatorestimatordeviationprocess_sum_range_succ theorem sampledobservedcomparatorestimatordeviationprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (comparator : action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledobservedcomparatorestimatordeviationprocess arms eta gamma loss comparator i sample) = (finset.range horizon).sum (fun i => sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator i sample) theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorDeviationProxy_sum_range_succ","label":"sampledComparatorEstimatorDeviationProxy_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledComparatorEstimatorDeviationProxy_sum_range_succ","description":"theorem sampledComparatorEstimatorDeviationProxy_sum_range_succ {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) : (Finset.range (horizon + 1)).sum (sampledComparatorEstimatorDeviationProxy arms gamma) = (horizon : NNReal) * sampledComparatorEstimatorVarianceProxy arms gamma","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-dfbb414dc0f5","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4073,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:638"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledComparatorEstimatorDeviationProxy_sum_range_succ {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) : (Finset.range (horizon + 1)).sum (sampledComparatorEstimatorDeviationProxy arms gamma) = (horizon : NNReal) * sampledComparatorEstimatorVarianceProxy arms gamma","missing":[],"search":"sampledcomparatorestimatordeviationproxy_sum_range_succ banditrlproof.exp3.sampledcomparatorestimatordeviationproxy_sum_range_succ theorem sampledcomparatorestimatordeviationproxy_sum_range_succ {action : type v} (arms : finset action) (gamma : real) (horizon : nat) : (finset.range (horizon + 1)).sum (sampledcomparatorestimatordeviationproxy arms gamma) = (horizon : nnreal) * sampledcomparatorestimatorvarianceproxy arms gamma theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_ennreal","label":"sampledObservedComparatorEstimatorDeviation_sum_tail_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_ennreal","description":"Fixed-comparator concentration for the observed importance-weighted EXP3 estimator minus its true predictable comparator loss.","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-79b71deab19a","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4074,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:655"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_sum_tail_ennreal {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | eps <= (Finset.range horizon).sum (fun i => sampledTrajectoryObservedComparatorEstimatorDeviationAt arms eta gamma loss comparator i sample)} <= ENNRea…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_sum_tail_ennreal banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_sum_tail_ennreal fixed-comparator concentration for the observed importance-weighted exp3 estimator minus its true predictable comparator loss. theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorConfidenceRadius","label":"sampledComparatorEstimatorConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledComparatorEstimatorConfidenceRadius","description":"noncomputable def sampledComparatorEstimatorConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-aadecee17d20","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4075,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:725"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledComparatorEstimatorConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledcomparatorestimatorconfidenceradius banditrlproof.exp3.sampledcomparatorestimatorconfidenceradius noncomputable def sampledcomparatorestimatorconfidenceradius {action : type v} (arms : finset action) (gamma : real) (horizon : nat) (delta : real) : real definition compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorVarianceProxy_pos","label":"sampledComparatorEstimatorVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledComparatorEstimatorVarianceProxy_pos","description":"theorem sampledComparatorEstimatorVarianceProxy_pos {Action : Type v} (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) : 0 < ((sampledComparatorEstimatorVarianceProxy arms gamma : NNReal) : Real)","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-cf9c89fc9867","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4076,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:733"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledComparatorEstimatorVarianceProxy_pos {Action : Type v} (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) : 0 < ((sampledComparatorEstimatorVarianceProxy arms gamma : NNReal) : Real)","missing":[],"search":"sampledcomparatorestimatorvarianceproxy_pos banditrlproof.exp3.sampledcomparatorestimatorvarianceproxy_pos theorem sampledcomparatorestimatorvarianceproxy_pos {action : type v} (arms : finset action) (harms : arms.nonempty) (gamma : real) (hgamma_pos : 0 < gamma) : 0 < ((sampledcomparatorestimatorvarianceproxy arms gamma : nnreal) : real) theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_exp_neg_budget","label":"sampledObservedComparatorEstimatorDeviation_sum_tail_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_exp_neg_budget","description":"theorem sampledObservedComparatorEstimatorDeviation_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgam…","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-7ec68c234482","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4077,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:750"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (budget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | Real.sqrt (2 * ((((horizon : NNReal) * sampledComparatorEstimatorVarianceProxy arms gamma : NNReal)) : Real) * budget) <= (Finset.rang…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_sum_tail_exp_neg_budget banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_sum_tail_exp_neg_budget theorem sampledobservedcomparatorestimatordeviation_sum_tail_exp_neg_budget {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (horizon : nat) (hhorizon : 0 < horizon) (budget : real) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | real.sqrt (2 * ((((horizon : nnreal) * sampledcomparatorestimatorvarianceproxy arms gamma : nnreal)) : real) * budget) <= (finset.range horizon).sum (fun i => sampledtrajectoryobservedcomparatorestimatordeviationat arms eta gamma loss comparator i sample)} <= ennreal.ofreal (real.exp (-budget)) theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_delta","label":"sampledObservedComparatorEstimatorDeviation_sum_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_delta","description":"Delta-shaped one-sided confidence bound for a fixed comparator's observed importance-weighted estimator against its true predictable loss.","url":"../modules/banditrlproof-exp3comparatorconfidence/index.html#decl-4d9e6790811a","parent":"module:BanditRLProof.Exp3ComparatorConfidence","order":4078,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ComparatorConfidence"],["Source","BanditRLProof/Exp3ComparatorConfidence.lean:800"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedComparatorEstimatorDeviation_sum_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledComparatorEstimatorConfidenceRadius arms gamma horizon delta <= (Finset.range horizon).sum (fun i => observedImporta…","missing":[],"search":"sampledobservedcomparatorestimatordeviation_sum_tail_delta banditrlproof.exp3.sampledobservedcomparatorestimatordeviation_sum_tail_delta delta-shaped one-sided confidence bound for a fixed comparator's observed importance-weighted estimator against its true predictable loss. theorem compiled","shard":"modules/1cfd6644c7b64780.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.FiniteActionDistribution","label":"FiniteActionDistribution","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Exp3.FiniteActionDistribution","description":"Probability-vector contracts on an explicit finite action support.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-fdfb9052e832","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4079,"meta":[["Kind","structure"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"structure FiniteActionDistribution {Action : Type u} (arms : Finset Action) (prob : Action -> Real) : Prop where","missing":[],"search":"finiteactiondistribution banditrlproof.exp3.finiteactiondistribution probability-vector contracts on an explicit finite action support. structure compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionMeasure","label":"finiteActionMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionMeasure","description":"The finite action law generated by a probability vector on `arms`.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-5316884508ff","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4080,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:28"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteActionMeasure {Action : Type u} [MeasurableSpace Action] (arms : Finset Action) (prob : Action -> Real) : Measure Action","missing":[],"search":"finiteactionmeasure banditrlproof.exp3.finiteactionmeasure the finite action law generated by a probability vector on `arms`. definition compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionMeasure_isProbabilityMeasure","label":"finiteActionMeasure_isProbabilityMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionMeasure_isProbabilityMeasure","description":"theorem finiteActionMeasure_isProbabilityMeasure {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : Action -> Real) (hdist : FiniteActionDistribution arms prob) : IsProbabilityMeasure (finiteActionMeasure arms prob)","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-ca2a53191d64","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4081,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionMeasure_isProbabilityMeasure {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : Action -> Real) (hdist : FiniteActionDistribution arms prob) : IsProbabilityMeasure (finiteActionMeasure arms prob)","missing":[],"search":"finiteactionmeasure_isprobabilitymeasure banditrlproof.exp3.finiteactionmeasure_isprobabilitymeasure theorem finiteactionmeasure_isprobabilitymeasure {action : type u} [measurablespace action] [measurablesingletonclass action] (arms : finset action) (prob : action -> real) (hdist : finiteactiondistribution arms prob) : isprobabilitymeasure (finiteactionmeasure arms prob) theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_finiteActionMeasure_eq_sum","label":"integral_finiteActionMeasure_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_finiteActionMeasure_eq_sum","description":"theorem integral_finiteActionMeasure_eq_sum {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : Action -> Real) (hdist : FiniteActionDistribution arms prob) (f : Action -> Real) : integral (finiteActionMeasure arms prob) f = arms.sum (fun action => prob action * f action)","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-f6bfb05fa056","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4082,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:42"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_finiteActionMeasure_eq_sum {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : Action -> Real) (hdist : FiniteActionDistribution arms prob) (f : Action -> Real) : integral (finiteActionMeasure arms prob) f = arms.sum (fun action => prob action * f action)","missing":[],"search":"integral_finiteactionmeasure_eq_sum banditrlproof.exp3.integral_finiteactionmeasure_eq_sum theorem integral_finiteactionmeasure_eq_sum {action : type u} [measurablespace action] [measurablesingletonclass action] (arms : finset action) (prob : action -> real) (hdist : finiteactiondistribution arms prob) (f : action -> real) : integral (finiteactionmeasure arms prob) f = arms.sum (fun action => prob action * f action) theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_historyAction_eq_integral_sum_of_condDistrib_ae_eq_finiteActionMeasure","label":"integral_historyAction_eq_integral_sum_of_condDistrib_ae_eq_finiteActionMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_historyAction_eq_integral_sum_of_condDistrib_ae_eq_finiteActionMeasure","description":"An identified finite conditional action law converts any integrable history/action score into the corresponding probability-weighted finite sum.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-a243992bf803","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4083,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:60"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_historyAction_eq_integral_sum_of_condDistrib_ae_eq_finiteActionMeasure {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob : History -> Action -> Real) (hdist : forall h, FiniteActionDistribution arms (prob h)) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (score : History × Action -> Real) (hscore : Measurable score) (hIntegrable : Integrable score (mu.map history ⊗ₘ policy)…","missing":[],"search":"integral_historyaction_eq_integral_sum_of_conddistrib_ae_eq_finiteactionmeasure banditrlproof.exp3.integral_historyaction_eq_integral_sum_of_conddistrib_ae_eq_finiteactionmeasure an identified finite conditional action law converts any integrable history/action score into the corresponding probability-weighted finite sum. theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_importanceWeightedLoss_eq_integral_loss_of_condDistrib","label":"integral_importanceWeightedLoss_eq_integral_loss_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_importanceWeightedLoss_eq_integral_loss_of_condDistrib","description":"One-round armwise unbiasedness under the identified conditional action law.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-afd3eb8a11db","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4084,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:100"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_importanceWeightedLoss_eq_integral_loss_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (hdist : forall h, FiniteActionDistribution arms (prob h)) (hprob : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (comparator : Action) (hcomparator : com…","missing":[],"search":"integral_importanceweightedloss_eq_integral_loss_of_conddistrib banditrlproof.exp3.integral_importanceweightedloss_eq_integral_loss_of_conddistrib one-round armwise unbiasedness under the identified conditional action law. theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_condDistrib","label":"integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_condDistrib","description":"The mixed estimated-loss integral equals the true mixed-loss integral.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-068e992f1c15","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4085,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:138"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (hdist : forall h, FiniteActionDistribution arms (prob h)) (hprob : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (hscore : Measurable (fun z :…","missing":[],"search":"integral_mixedimportanceweightedloss_eq_integral_mixedloss_of_conddistrib banditrlproof.exp3.integral_mixedimportanceweightedloss_eq_integral_mixedloss_of_conddistrib the mixed estimated-loss integral equals the true mixed-loss integral. theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_weightedImportanceWeightedLoss_eq_integral_weightedLoss_of_condDistrib","label":"integral_weightedImportanceWeightedLoss_eq_integral_weightedLoss_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_weightedImportanceWeightedLoss_eq_integral_weightedLoss_of_condDistrib","description":"A second predictable finite distribution may weight the estimator without changing the sampling law used by the conditional transport.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-c47dc4f5f82f","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4086,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:177"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_weightedImportanceWeightedLoss_eq_integral_weightedLoss_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob weight loss : History -> Action -> Real) (hdist : forall h, FiniteActionDistribution arms (prob h)) (hprob : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (hscore : Measura…","missing":[],"search":"integral_weightedimportanceweightedloss_eq_integral_weightedloss_of_conddistrib banditrlproof.exp3.integral_weightedimportanceweightedloss_eq_integral_weightedloss_of_conddistrib a second predictable finite distribution may weight the estimator without changing the sampling law used by the conditional transport. theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_condDistrib","label":"integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_condDistrib","description":"The mixed estimator-square integral equals the integral of summed loss squares.","url":"../modules/banditrlproof-exp3conditionalmoments/index.html#decl-c7e7e9744ff7","parent":"module:BanditRLProof.Exp3ConditionalMoments","order":4087,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ConditionalMoments"],["Source","BanditRLProof/Exp3ConditionalMoments.lean:216"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (hdist : forall h, FiniteActionDistribution arms (prob h)) (hprob : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (hscore : Measurable…","missing":[],"search":"integral_mixedsquaredimportanceweightedloss_eq_integral_sum_loss_sq_of_conddistrib banditrlproof.exp3.integral_mixedsquaredimportanceweightedloss_eq_integral_sum_loss_sq_of_conddistrib the mixed estimator-square integral equals the integral of summed loss squares. theorem compiled","shard":"modules/b9c971db79a12e84.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.eventually_const_le_natCast","label":"eventually_const_le_natCast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.eventually_const_le_natCast","description":"private theorem eventually_const_le_natCast (c : Real) : ∀ᶠ n : Nat in atTop, c <= (n : Real)","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-b665fa6cfe38","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4088,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"private theorem eventually_const_le_natCast (c : Real) : ∀ᶠ n : Nat in atTop, c <= (n : Real)","missing":[],"search":"eventually_const_le_natcast banditrlproof.exp3.eventually_const_le_natcast private theorem eventually_const_le_natcast (c : real) : ∀ᶠ n : nat in attop, c <= (n : real) theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.eventually_const_le_natCast_pow_three","label":"eventually_const_le_natCast_pow_three","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.eventually_const_le_natCast_pow_three","description":"private theorem eventually_const_le_natCast_pow_three (c : Real) : ∀ᶠ n : Nat in atTop, c <= (n : Real) ^ 3","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-3f550216384f","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4089,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"private theorem eventually_const_le_natCast_pow_three (c : Real) : ∀ᶠ n : Nat in atTop, c <= (n : Real) ^ 3","missing":[],"search":"eventually_const_le_natcast_pow_three banditrlproof.exp3.eventually_const_le_natcast_pow_three private theorem eventually_const_le_natcast_pow_three (c : real) : ∀ᶠ n : nat in attop, c <= (n : real) ^ 3 theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.eventually_doubleVarianceProbabilisticSparseLossLargeHorizonCondition","label":"eventually_doubleVarianceProbabilisticSparseLossLargeHorizonCondition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.eventually_doubleVarianceProbabilisticSparseLossLargeHorizonCondition","description":"Every fixed choice of the four real-valued scale parameters eventually satisfies the deterministic large-horizon inequalities.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-75d3436db964","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4090,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:45"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem eventually_doubleVarianceProbabilisticSparseLossLargeHorizonCondition (K S delta : Real) : ∀ᶠ horizon : Nat in atTop, doubleVarianceProbabilisticSparseLossLargeHorizonCondition K S (horizon : Real) delta","missing":[],"search":"eventually_doublevarianceprobabilisticsparselosslargehorizoncondition banditrlproof.exp3.eventually_doublevarianceprobabilisticsparselosslargehorizoncondition every fixed choice of the four real-valued scale parameters eventually satisfies the deterministic large-horizon inequalities. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold_eq_explicit_of_largeHorizon","label":"doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold_eq_explicit_of_largeHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold_eq_explicit_of_largeHorizon","description":"On the large-horizon branch, the best-arm all-horizon threshold is exactly the existing explicit double-variance threshold `16 * gamma * horizon`.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-f7bb83330915","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4091,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:61"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold_eq_explicit_of_largeHorizon {Action : Type*} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) (hlarge : doubleVarianceProbabilisticSparseLossLargeHorizonCondition (arms.card : Real) (sparsity : Real) (horizon : Real) (delta / (arms.card : Real))) : doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold arms horizon sparsity delta = pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedExplicitThreshold arms (doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) (delta / (arms.card : Real))) horizon sparsity (delta / (arms.card : Real))","missing":[],"search":"doublevarianceprobabilisticsparselossbestarmallhorizonregretthreshold_eq_explicit_of_largehorizon banditrlproof.exp3.doublevarianceprobabilisticsparselossbestarmallhorizonregretthreshold_eq_explicit_of_largehorizon on the large-horizon branch, the best-arm all-horizon threshold is exactly the existing explicit double-variance threshold `16 * gamma * horizon`. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure_of_largeHorizon","label":"sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure_of_largeHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure_of_largeHorizon","description":"Pointwise off-sparsity tail with the explicit threshold, obtained without changing the generated trajectory measure selected by the parent theorem.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-e95e68ea312b","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4092,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:84"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure_of_largeHorizon {Env Action : Type*} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : MeasureTheory.Measure Env) [MeasureTheory.IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge : doubleVarianceProbabilisticSparseLossLargeHorizonCondition (arms.card : Real) (sparsity : Real) (horizon : Real) (delta / (arms.card : Real))) : let deltaArm := delta / (arms.card : Real) let gamma := doubl…","missing":[],"search":"sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_off_sparsityfailure_of_largehorizon banditrlproof.exp3.sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_off_sparsityfailure_of_largehorizon pointwise off-sparsity tail with the explicit threshold, obtained without changing the generated trajectory measure selected by the parent theorem. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_largeHorizon","label":"sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_largeHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_largeHorizon","description":"Pointwise residual tail with the explicit threshold. The sparsity-failure term is retained exactly as in the all-horizon parent theorem.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-9272a4795c4c","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4093,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:143"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_largeHorizon {Env Action : Type*} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge : doubleVarianceProbabilisticSparseLossLargeHorizonCondition (arms.card : Real) (sparsity : Real) (horizon : Real) (delta / (arms.card : Real))) : let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorati…","missing":[],"search":"sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_largehorizon banditrlproof.exp3.sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_largehorizon pointwise residual tail with the explicit threshold. the sparsity-failure term is retained exactly as in the all-horizon parent theorem. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le_of_largeHorizon","label":"sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le_of_largeHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le_of_largeHorizon","description":"Pointwise practical tail after supplying an outer-measure bound for the sparsity-failure event.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-0e508d69daf2","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4094,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:203"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le_of_largeHorizon {Env Action : Type*} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge : doubleVarianceProbabilisticSparseLossLargeHorizonCondition (arms.card : Real) (sparsity : Real) (horizon : Real) (delta / (arms.card : Real))) : let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabili…","missing":[],"search":"sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_sparsityfailure_le_of_largehorizon banditrlproof.exp3.sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_sparsityfailure_le_of_largehorizon pointwise practical tail after supplying an outer-measure bound for the sparsity-failure event. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","label":"eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","description":"Eventually, every positive horizon has the explicit off-sparsity tail. Each horizon retains its own internally selected rates and trajectory measure.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-c3e210af7cbc","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4095,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:266"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure {Env Action : Type*} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (sparsity : Nat) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : ∀ᶠ horizon : Nat in atTop, 0 < horizon ∧ ∀ hhorizon : 0 < horizon, let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabi…","missing":[],"search":"eventually_sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.eventually_sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_off_sparsityfailure eventually, every positive horizon has the explicit off-sparsity tail. each horizon retains its own internally selected rates and trajectory measure. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","label":"eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","description":"Eventually, every positive horizon has the explicit residual tail. This is an at-top statement about horizon-indexed measures, not an anytime event.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-b29f3cc8fb23","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4096,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:326"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail {Env Action : Type*} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (sparsity : Nat) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : ∀ᶠ horizon : Nat in atTop, 0 < horizon ∧ ∀ hhorizon : 0 < horizon, let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHigh…","missing":[],"search":"eventually_sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail banditrlproof.exp3.eventually_sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail eventually, every positive horizon has the explicit residual tail. this is an at-top statement about horizon-indexed measures, not an anytime event. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","label":"eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","description":"Eventually, an external sparsity-failure bound yields the explicit practical tail under the same horizon-indexed generated trajectory measures.","url":"../modules/banditrlproof-exp3doublevariancesparsebestarmeventualrefinedregret/index.html#decl-d9041e55230b","parent":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","order":4097,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret"],["Source","BanditRLProof/Exp3DoubleVarianceSparseBestArmEventualRefinedRegret.lean:387"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le {Env Action : Type*} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (sparsity : Nat) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : ∀ᶠ horizon : Nat in atTop, 0 < horizon ∧ ∀ hhorizon : 0 < horizon, let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVaria…","missing":[],"search":"eventually_sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.eventually_sampledpredictable_explicitdoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_sparsityfailure_le eventually, an external sparsity-failure bound yields the explicit practical tail under the same horizon-indexed generated trajectory measures. theorem compiled","shard":"modules/4d795fab0daae3b8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.expectedRegretBudget_le_four_mul_gamma_mul_horizon","label":"expectedRegretBudget_le_four_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.expectedRegretBudget_le_four_mul_gamma_mul_horizon","description":"Deterministic EXP3 parameter algebra. If `eta = gamma / K`, exploration is at most one half, and `gamma^2 T` covers `K log K`, then the unoptimized budget is at most `4 gamma T`.","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-cf07d6ef251b","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4098,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem expectedRegretBudget_le_four_mul_gamma_mul_horizon (K T logK eta gamma : Real) (hK : 0 < K) (hT : 0 <= T) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (heta : eta = gamma / K) (hlog_budget : K * logK <= gamma ^ 2 * T) : logK / eta + (eta * (1 / (1 - gamma))) * (K * T) + gamma * T <= 4 * gamma * T","missing":[],"search":"expectedregretbudget_le_four_mul_gamma_mul_horizon banditrlproof.exp3.expectedregretbudget_le_four_mul_gamma_mul_horizon deterministic exp3 parameter algebra. if `eta = gamma / k`, exploration is at most one half, and `gamma^2 t` covers `k log k`, then the unoptimized budget is at most `4 gamma t`. theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedExplorationRate","label":"tunedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedExplorationRate","description":"Square-root exploration scale used by the tuned corollary.","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-202d7fa95ad4","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4099,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:54"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def tunedExplorationRate (K T : Real) : Real","missing":[],"search":"tunedexplorationrate banditrlproof.exp3.tunedexplorationrate square-root exploration scale used by the tuned corollary. definition compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedLearningRate","label":"tunedLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedLearningRate","description":"Learning rate paired with `tunedExplorationRate`.","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-ff1ad557bb73","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4100,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:58"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def tunedLearningRate (K T : Real) : Real","missing":[],"search":"tunedlearningrate banditrlproof.exp3.tunedlearningrate learning rate paired with `tunedexplorationrate`. definition compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedExplorationRate_pos","label":"tunedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedExplorationRate_pos","description":"theorem tunedExplorationRate_pos (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < tunedExplorationRate K T","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-7e8aeb206984","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4101,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:61"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem tunedExplorationRate_pos (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < tunedExplorationRate K T","missing":[],"search":"tunedexplorationrate_pos banditrlproof.exp3.tunedexplorationrate_pos theorem tunedexplorationrate_pos (k t : real) (hk_one : 1 < k) (ht : 0 < t) : 0 < tunedexplorationrate k t theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedExplorationRate_le_half","label":"tunedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedExplorationRate_le_half","description":"theorem tunedExplorationRate_le_half (K T : Real) (hT : 0 < T) (hscale : 4 * K * Real.log K <= T) : tunedExplorationRate K T <= 1 / 2","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-c4204728e568","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4102,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:67"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem tunedExplorationRate_le_half (K T : Real) (hT : 0 < T) (hscale : 4 * K * Real.log K <= T) : tunedExplorationRate K T <= 1 / 2","missing":[],"search":"tunedexplorationrate_le_half banditrlproof.exp3.tunedexplorationrate_le_half theorem tunedexplorationrate_le_half (k t : real) (ht : 0 < t) (hscale : 4 * k * real.log k <= t) : tunedexplorationrate k t <= 1 / 2 theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedLearningRate_pos","label":"tunedLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedLearningRate_pos","description":"theorem tunedLearningRate_pos (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < tunedLearningRate K T","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-abdf0eed5f8b","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4103,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:77"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem tunedLearningRate_pos (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < tunedLearningRate K T","missing":[],"search":"tunedlearningrate_pos banditrlproof.exp3.tunedlearningrate_pos theorem tunedlearningrate_pos (k t : real) (hk_one : 1 < k) (ht : 0 < t) : 0 < tunedlearningrate k t theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedExplorationRate_sq_mul_eq","label":"tunedExplorationRate_sq_mul_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedExplorationRate_sq_mul_eq","description":"theorem tunedExplorationRate_sq_mul_eq (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : tunedExplorationRate K T ^ 2 * T = K * Real.log K","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-25916b3be89c","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4104,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:82"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem tunedExplorationRate_sq_mul_eq (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : tunedExplorationRate K T ^ 2 * T = K * Real.log K","missing":[],"search":"tunedexplorationrate_sq_mul_eq banditrlproof.exp3.tunedexplorationrate_sq_mul_eq theorem tunedexplorationrate_sq_mul_eq (k t : real) (hk_one : 1 < k) (ht : 0 < t) : tunedexplorationrate k t ^ 2 * t = k * real.log k theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedExplorationRate_mul_eq_sqrt_mul","label":"tunedExplorationRate_mul_eq_sqrt_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedExplorationRate_mul_eq_sqrt_mul","description":"theorem tunedExplorationRate_mul_eq_sqrt_mul (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : tunedExplorationRate K T * T = Real.sqrt (K * T * Real.log K)","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-49c54e144271","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4105,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:90"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem tunedExplorationRate_mul_eq_sqrt_mul (K T : Real) (hK_one : 1 < K) (hT : 0 < T) : tunedExplorationRate K T * T = Real.sqrt (K * T * Real.log K)","missing":[],"search":"tunedexplorationrate_mul_eq_sqrt_mul banditrlproof.exp3.tunedexplorationrate_mul_eq_sqrt_mul theorem tunedexplorationrate_mul_eq_sqrt_mul (k t : real) (hk_one : 1 < k) (ht : 0 < t) : tunedexplorationrate k t * t = real.sqrt (k * t * real.log k) theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_gamma_mul_horizon","label":"sampledPredictable_expectedRegret_le_four_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_gamma_mul_horizon","description":"The generated predictable EXP3 trajectory has regret at most `4 gamma horizon` when `eta = gamma / |arms|` and the exploration budget dominates `|arms| log |arms|`.","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-1a7ec4028370","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4106,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:110"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_expectedRegret_le_four_mul_gamma_mul_horizon {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (horizon : Nat) (hlog_budget : (arms.card : Real) * Real.log arms.card <= gamma ^ 2 * (horizon : Real)) (comparator : Action) (hcomparator : comparator ∈ arms) : let eta := gamma / (arms.card : Real) let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le (by linarith) loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajector…","missing":[],"search":"sampledpredictable_expectedregret_le_four_mul_gamma_mul_horizon banditrlproof.exp3.sampledpredictable_expectedregret_le_four_mul_gamma_mul_horizon the generated predictable exp3 trajectory has regret at most `4 gamma horizon` when `eta = gamma / |arms|` and the exploration budget dominates `|arms| log |arms|`. theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.tunedPredictableTrajectoryKernel","label":"tunedPredictableTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.tunedPredictableTrajectoryKernel","description":"Generated predictable EXP3 kernel at the square-root exploration and learning rates. The large-horizon hypotheses discharge the kernel's `0 <= gamma <= 1` contract internally.","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-1bc066d57ca9","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4107,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:153"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def tunedPredictableTrajectoryKernel {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon_pos : 0 < horizon) (hscale : 4 * (arms.card : Real) * Real.log arms.card <= (horizon : Real)) : Kernel Env (Nat -> Action × Real)","missing":[],"search":"tunedpredictabletrajectorykernel banditrlproof.exp3.tunedpredictabletrajectorykernel generated predictable exp3 kernel at the square-root exploration and learning rates. the large-horizon hypotheses discharge the kernel's `0 <= gamma <= 1` contract internally. definition compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_sqrt","label":"sampledPredictable_expectedRegret_le_four_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_sqrt","description":"Tuned generated-trajectory expected predictable-regret bound. Under the large-horizon regime `4 |A| log |A| <= T`, the square-root exploration and learning rates give the classical `sqrt(|A| T log |A|)` scale.","url":"../modules/banditrlproof-exp3expectedregret/index.html#decl-143a57ec9c49","parent":"module:BanditRLProof.Exp3ExpectedRegret","order":4108,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExpectedRegret"],["Source","BanditRLProof/Exp3ExpectedRegret.lean:181"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","adversarial-bobw"]],"statement":"theorem sampledPredictable_expectedRegret_le_four_mul_sqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon_pos : 0 < horizon) (hscale : 4 * (arms.card : Real) * Real.log arms.card <= (horizon : Real)) (comparator : Action) (hcomparator : comparator ∈ arms) : let K := (arms.card : Real) let T := (horizon : Real) let gamma := tunedExplorationRate K T let eta := tunedLearningRate K T let mu := prior ⊗ₘ tunedPredictableTrajectoryKernel arms harms hcard_two loss horizon hhorizon_pos hscale integral mu (fun sample => (Finset.range horizon).sum (f…","missing":[],"search":"sampledpredictable_expectedregret_le_four_mul_sqrt banditrlproof.exp3.sampledpredictable_expectedregret_le_four_mul_sqrt tuned generated-trajectory expected predictable-regret bound. under the large-horizon regime `4 |a| log |a| <= t`, the square-root exploration and learning rates give the classical `sqrt(|a| t log |a|)` scale. theorem compiled","shard":"modules/5875be88110a5986.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":["adversarial-bobw"]},{"id":"declaration:BanditRLProof.Exp3.distribution_le_sampledTrajectoryProbabilityAt_div_one_sub_gamma","label":"distribution_le_sampledTrajectoryProbabilityAt_div_one_sub_gamma","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.distribution_le_sampledTrajectoryProbabilityAt_div_one_sub_gamma","description":"A pure Hedge coordinate is at most the corresponding explored probability divided by `1 - gamma`.","url":"../modules/banditrlproof-exp3explorationbias/index.html#decl-31da692af62a","parent":"module:BanditRLProof.Exp3ExplorationBias","order":4109,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExplorationBias"],["Source","BanditRLProof/Exp3ExplorationBias.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem distribution_le_sampledTrajectoryProbabilityAt_div_one_sub_gamma {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_lt_one : gamma < 1) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (action : Action) : distribution arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t action <= sampledTrajectoryProbabilityAt arms eta gamma t sample action / (1 - gamma)","missing":[],"search":"distribution_le_sampledtrajectoryprobabilityat_div_one_sub_gamma banditrlproof.exp3.distribution_le_sampledtrajectoryprobabilityat_div_one_sub_gamma a pure hedge coordinate is at most the corresponding explored probability divided by `1 - gamma`. theorem compiled","shard":"modules/45e8ae535c89af6c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredLoss_sampledTrajectoryObservedLoss_le_inv_one_sub_gamma","label":"mixedSquaredLoss_sampledTrajectoryObservedLoss_le_inv_one_sub_gamma","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredLoss_sampledTrajectoryObservedLoss_le_inv_one_sub_gamma","description":"The pure-Hedge estimator square is bounded by the explored-probability mixed square with the standard `1 / (1 - gamma)` factor.","url":"../modules/banditrlproof-exp3explorationbias/index.html#decl-f5e8b06c5be7","parent":"module:BanditRLProof.Exp3ExplorationBias","order":4110,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExplorationBias"],["Source","BanditRLProof/Exp3ExplorationBias.lean:50"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredLoss_sampledTrajectoryObservedLoss_le_inv_one_sub_gamma {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_lt_one : gamma < 1) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : mixedSquaredLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t <= (1 / (1 - gamma)) * observedMixedSquaredImportanceWeightedLossAt arms eta gamma t sample","missing":[],"search":"mixedsquaredloss_sampledtrajectoryobservedloss_le_inv_one_sub_gamma banditrlproof.exp3.mixedsquaredloss_sampledtrajectoryobservedloss_le_inv_one_sub_gamma the pure-hedge estimator square is bounded by the explored-probability mixed square with the standard `1 / (1 - gamma)` factor. theorem compiled","shard":"modules/45e8ae535c89af6c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedLoss_le_pure_add_gamma","label":"sampledTrajectoryPredictableMixedLoss_le_pure_add_gamma","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableMixedLoss_le_pure_add_gamma","description":"Exploration changes one predictable `[0,1]` mixed loss by at most `gamma` relative to the pure Hedge distribution.","url":"../modules/banditrlproof-exp3explorationbias/index.html#decl-5e69a223e23c","parent":"module:BanditRLProof.Exp3ExplorationBias","order":4111,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExplorationBias"],["Source","BanditRLProof/Exp3ExplorationBias.lean:90"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableMixedLoss_le_pure_add_gamma {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : arms.sum (fun action => sampledTrajectoryProbabilityAt arms eta gamma t sample action * predictableLossAt loss t sample action) <= arms.sum (fun action => distribution arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t action * predictableLossAt loss t sample action) + gamma","missing":[],"search":"sampledtrajectorypredictablemixedloss_le_pure_add_gamma banditrlproof.exp3.sampledtrajectorypredictablemixedloss_le_pure_add_gamma exploration changes one predictable `[0,1]` mixed loss by at most `gamma` relative to the pure hedge distribution. theorem compiled","shard":"modules/45e8ae535c89af6c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectory_finiteHorizon_explorationBias_secondMoment","label":"sampledTrajectory_finiteHorizon_explorationBias_secondMoment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectory_finiteHorizon_explorationBias_secondMoment","description":"Finite-horizon exploration bias and second-moment comparison on one concrete sampled trajectory.","url":"../modules/banditrlproof-exp3explorationbias/index.html#decl-d9dc29f55452","parent":"module:BanditRLProof.Exp3ExplorationBias","order":4112,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ExplorationBias"],["Source","BanditRLProof/Exp3ExplorationBias.lean:180"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectory_finiteHorizon_explorationBias_secondMoment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : ((Finset.range horizon).sum (fun t => arms.sum (fun action => sampledTrajectoryProbabilityAt arms eta gamma t sample action * predictableLossAt loss t sample action)) <= (Finset.range horizon).sum (fun t => arms.sum (fun action => distribution arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t action * predictableLossAt loss t sample action)) + gamma * (horizon : Real)) ∧ ((Finset.range horizon).sum (fun t => mixedSquaredLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma s…","missing":[],"search":"sampledtrajectory_finitehorizon_explorationbias_secondmoment banditrlproof.exp3.sampledtrajectory_finitehorizon_explorationbias_secondmoment finite-horizon exploration bias and second-moment comparison on one concrete sampled trajectory. theorem compiled","shard":"modules/45e8ae535c89af6c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.cumulativeLoss","label":"cumulativeLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.cumulativeLoss","description":"Cumulative loss of one action before round `t`.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-8ef24797a264","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4113,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeLoss {Action : Type u} (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : Real","missing":[],"search":"cumulativeloss banditrlproof.exp3.cumulativeloss cumulative loss of one action before round `t`. definition compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weight","label":"weight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.weight","description":"Exponential weight generated by the cumulative loss before round `t`.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-7688b144da49","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4114,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def weight {Action : Type u} (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : Real","missing":[],"search":"weight banditrlproof.exp3.weight exponential weight generated by the cumulative loss before round `t`. definition compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.totalWeight","label":"totalWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.totalWeight","description":"Total exponential weight on an explicit finite action set.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-7381d6b3d7bd","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4115,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:40"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def totalWeight {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : Real","missing":[],"search":"totalweight banditrlproof.exp3.totalweight total exponential weight on an explicit finite action set. definition compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.distribution","label":"distribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.distribution","description":"Normalized exponential weight at round `t`.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-1c765b8402d1","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4116,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:46"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def distribution {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : Real","missing":[],"search":"distribution banditrlproof.exp3.distribution normalized exponential weight at round `t`. definition compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedLoss","label":"mixedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedLoss","description":"Mixed loss incurred by the normalized exponential weights at round `t`.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-01a54d13a0d3","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4117,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:52"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def mixedLoss {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : Real","missing":[],"search":"mixedloss banditrlproof.exp3.mixedloss mixed loss incurred by the normalized exponential weights at round `t`. definition compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredLoss","label":"mixedSquaredLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredLoss","description":"Mixed second moment of the losses at round `t`.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-99709d95cca8","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4118,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:58"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def mixedSquaredLoss {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : Real","missing":[],"search":"mixedsquaredloss banditrlproof.exp3.mixedsquaredloss mixed second moment of the losses at round `t`. definition compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.cumulativeLoss_succ","label":"cumulativeLoss_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.cumulativeLoss_succ","description":"theorem cumulativeLoss_succ {Action : Type u} (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : cumulativeLoss loss (t + 1) a = cumulativeLoss loss t a + loss t a","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-6a695d25ba2e","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4119,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:63"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLoss_succ {Action : Type u} (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : cumulativeLoss loss (t + 1) a = cumulativeLoss loss t a + loss t a","missing":[],"search":"cumulativeloss_succ banditrlproof.exp3.cumulativeloss_succ theorem cumulativeloss_succ {action : type u} (loss : nat -> action -> real) (t : nat) (a : action) : cumulativeloss loss (t + 1) a = cumulativeloss loss t a + loss t a theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weight_pos","label":"weight_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.weight_pos","description":"theorem weight_pos {Action : Type u} (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : 0 < weight eta loss t a","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-69c62930cc62","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4120,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:68"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem weight_pos {Action : Type u} (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : 0 < weight eta loss t a","missing":[],"search":"weight_pos banditrlproof.exp3.weight_pos theorem weight_pos {action : type u} (eta : real) (loss : nat -> action -> real) (t : nat) (a : action) : 0 < weight eta loss t a theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weight_succ","label":"weight_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.weight_succ","description":"theorem weight_succ {Action : Type u} (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : weight eta loss (t + 1) a = weight eta loss t a * Real.exp (-eta * loss t a)","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-9721b2b292f9","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4121,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:73"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem weight_succ {Action : Type u} (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : weight eta loss (t + 1) a = weight eta loss t a * Real.exp (-eta * loss t a)","missing":[],"search":"weight_succ banditrlproof.exp3.weight_succ theorem weight_succ {action : type u} (eta : real) (loss : nat -> action -> real) (t : nat) (a : action) : weight eta loss (t + 1) a = weight eta loss t a * real.exp (-eta * loss t a) theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.totalWeight_pos","label":"totalWeight_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.totalWeight_pos","description":"theorem totalWeight_pos {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : 0 < totalWeight arms eta loss t","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-e325529c662d","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4122,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:83"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem totalWeight_pos {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : 0 < totalWeight arms eta loss t","missing":[],"search":"totalweight_pos banditrlproof.exp3.totalweight_pos theorem totalweight_pos {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : nat -> action -> real) (t : nat) : 0 < totalweight arms eta loss t theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.distribution_nonneg","label":"distribution_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.distribution_nonneg","description":"theorem distribution_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : 0 <= distribution arms eta loss t a","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-632a431925b3","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4123,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:92"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem distribution_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : 0 <= distribution arms eta loss t a","missing":[],"search":"distribution_nonneg banditrlproof.exp3.distribution_nonneg theorem distribution_nonneg {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : nat -> action -> real) (t : nat) (a : action) : 0 <= distribution arms eta loss t a theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.distribution_pos","label":"distribution_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.distribution_pos","description":"theorem distribution_pos {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : 0 < distribution arms eta loss t a","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-a4678bf9b3f3","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4124,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:99"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem distribution_pos {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (a : Action) : 0 < distribution arms eta loss t a","missing":[],"search":"distribution_pos banditrlproof.exp3.distribution_pos theorem distribution_pos {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : nat -> action -> real) (t : nat) (a : action) : 0 < distribution arms eta loss t a theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_distribution","label":"sum_distribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_distribution","description":"theorem sum_distribution {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : arms.sum (distribution arms eta loss t) = 1","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-ec399e2d0217","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4125,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:106"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_distribution {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : arms.sum (distribution arms eta loss t) = 1","missing":[],"search":"sum_distribution banditrlproof.exp3.sum_distribution theorem sum_distribution {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : nat -> action -> real) (t : nat) : arms.sum (distribution arms eta loss t) = 1 theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.totalWeight_zero","label":"totalWeight_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.totalWeight_zero","description":"theorem totalWeight_zero {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) : totalWeight arms eta loss 0 = arms.card","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-b6fb74151c2a","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4126,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:117"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem totalWeight_zero {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) : totalWeight arms eta loss 0 = arms.card","missing":[],"search":"totalweight_zero banditrlproof.exp3.totalweight_zero theorem totalweight_zero {action : type u} (arms : finset action) (eta : real) (loss : nat -> action -> real) : totalweight arms eta loss 0 = arms.card theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exp_neg_le_one_sub_add_sq","label":"exp_neg_le_one_sub_add_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exp_neg_le_one_sub_add_sq","description":"Global quadratic upper bound for a nonnegative exponential argument.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-39829dd3792f","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4127,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:123"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_le_one_sub_add_sq {x : Real} (hx : 0 <= x) : Real.exp (-x) <= 1 - x + x ^ 2","missing":[],"search":"exp_neg_le_one_sub_add_sq banditrlproof.exp3.exp_neg_le_one_sub_add_sq global quadratic upper bound for a nonnegative exponential argument. theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exp_neg_mul_le_one_sub_add_sq_of_nonneg","label":"exp_neg_mul_le_one_sub_add_sq_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exp_neg_mul_le_one_sub_add_sq_of_nonneg","description":"Quadratic exponential bound for arbitrary nonnegative learning rates and losses.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-6ac98e9c6c3b","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4128,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:140"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_mul_le_one_sub_add_sq_of_nonneg {eta ell : Real} (heta : 0 <= eta) (hell : 0 <= ell) : Real.exp (-eta * ell) <= 1 - eta * ell + eta ^ 2 * ell ^ 2","missing":[],"search":"exp_neg_mul_le_one_sub_add_sq_of_nonneg banditrlproof.exp3.exp_neg_mul_le_one_sub_add_sq_of_nonneg quadratic exponential bound for arbitrary nonnegative learning rates and losses. theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exp_neg_mul_le_one_sub_add_sq","label":"exp_neg_mul_le_one_sub_add_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exp_neg_mul_le_one_sub_add_sq","description":"Backwards-compatible bounded-loss specialization of the global bound.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-5e613bd7f4eb","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4129,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:147"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_mul_le_one_sub_add_sq {eta ell : Real} (heta : 0 <= eta) (_heta_le : eta <= 1) (hell : 0 <= ell) (_hell_le : ell <= 1) : Real.exp (-eta * ell) <= 1 - eta * ell + eta ^ 2 * ell ^ 2","missing":[],"search":"exp_neg_mul_le_one_sub_add_sq banditrlproof.exp3.exp_neg_mul_le_one_sub_add_sq backwards-compatible bounded-loss specialization of the global bound. theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.totalWeight_succ_div_eq_sum_distribution_mul_exp","label":"totalWeight_succ_div_eq_sum_distribution_mul_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.totalWeight_succ_div_eq_sum_distribution_mul_exp","description":"theorem totalWeight_succ_div_eq_sum_distribution_mul_exp {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : totalWeight arms eta loss (t + 1) / totalWeight arms eta loss t = arms.sum (fun a => distribution arms eta loss t a * Real.exp (-eta * loss t a))","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-dd09daa5e508","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4130,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:153"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem totalWeight_succ_div_eq_sum_distribution_mul_exp {Action : Type u} (arms : Finset Action) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : totalWeight arms eta loss (t + 1) / totalWeight arms eta loss t = arms.sum (fun a => distribution arms eta loss t a * Real.exp (-eta * loss t a))","missing":[],"search":"totalweight_succ_div_eq_sum_distribution_mul_exp banditrlproof.exp3.totalweight_succ_div_eq_sum_distribution_mul_exp theorem totalweight_succ_div_eq_sum_distribution_mul_exp {action : type u} (arms : finset action) (eta : real) (loss : nat -> action -> real) (t : nat) : totalweight arms eta loss (t + 1) / totalweight arms eta loss t = arms.sum (fun a => distribution arms eta loss t a * real.exp (-eta * loss t a)) theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.totalWeight_succ_div_le_of_nonneg","label":"totalWeight_succ_div_le_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.totalWeight_succ_div_le_of_nonneg","description":"theorem totalWeight_succ_div_le_of_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a) : totalWeight arms eta loss (t + 1) / totalWeight arms eta loss t <= 1 - eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-8646314044fb","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4131,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:169"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem totalWeight_succ_div_le_of_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a) : totalWeight arms eta loss (t + 1) / totalWeight arms eta loss t <= 1 - eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","missing":[],"search":"totalweight_succ_div_le_of_nonneg banditrlproof.exp3.totalweight_succ_div_le_of_nonneg theorem totalweight_succ_div_le_of_nonneg {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (heta : 0 <= eta) (loss : nat -> action -> real) (t : nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a) : totalweight arms eta loss (t + 1) / totalweight arms eta loss t <= 1 - eta * mixedloss arms eta loss t + eta ^ 2 * mixedsquaredloss arms eta loss t theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.totalWeight_succ_div_le","label":"totalWeight_succ_div_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.totalWeight_succ_div_le","description":"theorem totalWeight_succ_div_le {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (_heta_le : eta <= 1) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : totalWeight arms eta loss (t + 1) / totalWeight arms eta loss t <= 1 - eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-3773196ebcd1","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4132,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:218"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem totalWeight_succ_div_le {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (_heta_le : eta <= 1) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : totalWeight arms eta loss (t + 1) / totalWeight arms eta loss t <= 1 - eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","missing":[],"search":"totalweight_succ_div_le banditrlproof.exp3.totalweight_succ_div_le theorem totalweight_succ_div_le {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (heta : 0 <= eta) (_heta_le : eta <= 1) (loss : nat -> action -> real) (t : nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : totalweight arms eta loss (t + 1) / totalweight arms eta loss t <= 1 - eta * mixedloss arms eta loss t + eta ^ 2 * mixedsquaredloss arms eta loss t theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_totalWeight_succ_sub_le_of_nonneg","label":"log_totalWeight_succ_sub_le_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_totalWeight_succ_sub_le_of_nonneg","description":"theorem log_totalWeight_succ_sub_le_of_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a) : Real.log (totalWeight arms eta loss (t + 1)) - Real.log (totalWeight arms eta loss t) <= -eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-b7120b22d7e0","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4133,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:231"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_totalWeight_succ_sub_le_of_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a) : Real.log (totalWeight arms eta loss (t + 1)) - Real.log (totalWeight arms eta loss t) <= -eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","missing":[],"search":"log_totalweight_succ_sub_le_of_nonneg banditrlproof.exp3.log_totalweight_succ_sub_le_of_nonneg theorem log_totalweight_succ_sub_le_of_nonneg {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (heta : 0 <= eta) (loss : nat -> action -> real) (t : nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a) : real.log (totalweight arms eta loss (t + 1)) - real.log (totalweight arms eta loss t) <= -eta * mixedloss arms eta loss t + eta ^ 2 * mixedsquaredloss arms eta loss t theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_totalWeight_succ_sub_le","label":"log_totalWeight_succ_sub_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_totalWeight_succ_sub_le","description":"theorem log_totalWeight_succ_sub_le {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (_heta_le : eta <= 1) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : Real.log (totalWeight arms eta loss (t + 1)) - Real.log (totalWeight arms eta loss t) <= -eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta…","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-4102743bf0fc","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4134,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:249"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_totalWeight_succ_sub_le {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 <= eta) (_heta_le : eta <= 1) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : Real.log (totalWeight arms eta loss (t + 1)) - Real.log (totalWeight arms eta loss t) <= -eta * mixedLoss arms eta loss t + eta ^ 2 * mixedSquaredLoss arms eta loss t","missing":[],"search":"log_totalweight_succ_sub_le banditrlproof.exp3.log_totalweight_succ_sub_le theorem log_totalweight_succ_sub_le {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (heta : 0 <= eta) (_heta_le : eta <= 1) (loss : nat -> action -> real) (t : nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : real.log (totalweight arms eta loss (t + 1)) - real.log (totalweight arms eta loss t) <= -eta * mixedloss arms eta loss t + eta ^ 2 * mixedsquaredloss arms eta loss t theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredLoss_le_one","label":"mixedSquaredLoss_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredLoss_le_one","description":"theorem mixedSquaredLoss_le_one {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : mixedSquaredLoss arms eta loss t <= 1","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-e9c74b0d2952","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4135,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:262"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredLoss_le_one {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : mixedSquaredLoss arms eta loss t <= 1","missing":[],"search":"mixedsquaredloss_le_one banditrlproof.exp3.mixedsquaredloss_le_one theorem mixedsquaredloss_le_one {action : type u} (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : nat -> action -> real) (t : nat) (hloss : forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) : mixedsquaredloss arms eta loss t <= 1 theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss_of_nonneg","label":"hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss_of_nonneg","description":"Second-order finite-horizon Hedge regret bound against one comparator action.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-1c1b7b7f2961","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4136,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:283"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss_of_nonneg {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (loss : Nat -> Action -> Real) (T : Nat) (hloss : forall t, t < T -> forall a, a ∈ arms -> 0 <= loss t a) (comparator : Action) (hcomparator : comparator ∈ arms) : (Finset.range T).sum (fun t => mixedLoss arms eta loss t) - cumulativeLoss loss T comparator <= Real.log arms.card / eta + eta * (Finset.range T).sum (fun t => mixedSquaredLoss arms eta loss t)","missing":[],"search":"hedge_regret_le_log_card_div_add_eta_mul_mixedsquaredloss_of_nonneg banditrlproof.exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedsquaredloss_of_nonneg second-order finite-horizon hedge regret bound against one comparator action. theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss","label":"hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss","description":"Bounded-loss compatibility wrapper for the generalized second-order theorem.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-36052c1873b7","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4137,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:359"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (_heta_le : eta <= 1) (loss : Nat -> Action -> Real) (T : Nat) (hloss : forall t, t < T -> forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) (comparator : Action) (hcomparator : comparator ∈ arms) : (Finset.range T).sum (fun t => mixedLoss arms eta loss t) - cumulativeLoss loss T comparator <= Real.log arms.card / eta + eta * (Finset.range T).sum (fun t => mixedSquaredLoss arms eta loss t)","missing":[],"search":"hedge_regret_le_log_card_div_add_eta_mul_mixedsquaredloss banditrlproof.exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedsquaredloss bounded-loss compatibility wrapper for the generalized second-order theorem. theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_horizon","label":"hedge_regret_le_log_card_div_add_eta_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_horizon","description":"Finite-horizon Hedge regret for losses in `[0,1]`. This is the deterministic full-information theorem consumed by the future importance-weighted EXP3 route.","url":"../modules/banditrlproof-exp3hedgeregret/index.html#decl-80ae8edd2c79","parent":"module:BanditRLProof.Exp3HedgeRegret","order":4138,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HedgeRegret"],["Source","BanditRLProof/Exp3HedgeRegret.lean:382"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem hedge_regret_le_log_card_div_add_eta_mul_horizon {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (heta_le : eta <= 1) (loss : Nat -> Action -> Real) (T : Nat) (hloss : forall t, t < T -> forall a, a ∈ arms -> 0 <= loss t a ∧ loss t a <= 1) (comparator : Action) (hcomparator : comparator ∈ arms) : (Finset.range T).sum (fun t => mixedLoss arms eta loss t) - cumulativeLoss loss T comparator <= Real.log arms.card / eta + eta * T","missing":[],"search":"hedge_regret_le_log_card_div_add_eta_mul_horizon banditrlproof.exp3.hedge_regret_le_log_card_div_add_eta_mul_horizon finite-horizon hedge regret for losses in `[0,1]`. this is the deterministic full-information theorem consumed by the future importance-weighted exp3 route. theorem compiled","shard":"modules/ac9368b3cb1556bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt_le_inv_explorationFloor","label":"observedMixedSquaredImportanceWeightedLossAt_le_inv_explorationFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt_le_inv_explorationFloor","description":"theorem observedMixedSquaredImportanceWeightedLossAt_le_inv_explorationFloor {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Actio…","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html#decl-4b531d14d464","parent":"module:BanditRLProof.Exp3HighProbabilityRegret","order":4139,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HighProbabilityRegret"],["Source","BanditRLProof/Exp3HighProbabilityRegret.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem observedMixedSquaredImportanceWeightedLossAt_le_inv_explorationFloor {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (hreward : (sample.2 t).2 ∈ Set.Icc (0 : Real) 1) : observedMixedSquaredImportanceWeightedLossAt arms eta gamma t sample <= 1 / (gamma / (arms.card : Real))","missing":[],"search":"observedmixedsquaredimportanceweightedlossat_le_inv_explorationfloor banditrlproof.exp3.observedmixedsquaredimportanceweightedlossat_le_inv_explorationfloor theorem observedmixedsquaredimportanceweightedlossat_le_inv_explorationfloor {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) (hreward : (sample.2 t).2 ∈ set.icc (0 : real) 1) : observedmixedsquaredimportanceweightedlossat arms eta gamma t sample <= 1 / (gamma / (arms.card : real)) theorem compiled","shard":"modules/0823a558d097504b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_mem_unitInterval_ae","label":"sampledPredictableTrajectoryMeasure_reward_mem_unitInterval_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_mem_unitInterval_ae","description":"theorem sampledPredictableTrajectoryMeasure_reward_mem_unitInterval_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (…","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html#decl-6545144e0379","parent":"module:BanditRLProof.Exp3HighProbabilityRegret","order":4140,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HighProbabilityRegret"],["Source","BanditRLProof/Exp3HighProbabilityRegret.lean:70"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_reward_mem_unitInterval_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (sample.2 t).2 ∈ Set.Icc (0 : Real) 1","missing":[],"search":"sampledpredictabletrajectorymeasure_reward_mem_unitinterval_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_reward_mem_unitinterval_ae theorem sampledpredictabletrajectorymeasure_reward_mem_unitinterval_ae {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (sample.2 t).2 ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/0823a558d097504b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_observedMixedSquared_sum_le_ae","label":"sampledPredictableTrajectoryMeasure_observedMixedSquared_sum_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_observedMixedSquared_sum_le_ae","description":"theorem sampledPredictableTrajectoryMeasure_observedMixedSquared_sum_le_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (…","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html#decl-a465392c61d2","parent":"module:BanditRLProof.Exp3HighProbabilityRegret","order":4141,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HighProbabilityRegret"],["Source","BanditRLProof/Exp3HighProbabilityRegret.lean:113"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_observedMixedSquared_sum_le_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (Finset.range horizon).sum (fun t => observedMixedSquaredImportanceWeightedLossAt arms eta gamma t sample) <= (horizon : Real) * (1 / (gamma / (arms.card : Real)))","missing":[],"search":"sampledpredictabletrajectorymeasure_observedmixedsquared_sum_le_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_observedmixedsquared_sum_le_ae theorem sampledpredictabletrajectorymeasure_observedmixedsquared_sum_le_ae {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (finset.range horizon).sum (fun t => observedmixedsquaredimportanceweightedlossat arms eta gamma t sample) <= (horizon : real) * (1 / (gamma / (arms.card : real))) theorem compiled","shard":"modules/0823a558d097504b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableHighProbabilityRegretBudget","label":"sampledPredictableHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableHighProbabilityRegretBudget","description":"noncomputable def sampledPredictableHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html#decl-edbf47317b75","parent":"module:BanditRLProof.Exp3HighProbabilityRegret","order":4142,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3HighProbabilityRegret"],["Source","BanditRLProof/Exp3HighProbabilityRegret.lean:158"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablehighprobabilityregretbudget banditrlproof.exp3.sampledpredictablehighprobabilityregretbudget noncomputable def sampledpredictablehighprobabilityregretbudget {action : type v} (arms : finset action) (eta gamma : real) (horizon : nat) (delta : real) : real definition compiled","shard":"modules/0823a558d097504b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_delta","label":"sampledPredictable_highProbabilityRegret_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_delta","description":"Generated predictable EXP3 pseudo-regret exceeds the explicit Hedge, exploration, estimator-square, and two confidence-radius budget with probability at most the sum of the pure-q and comparator failure probabilities.","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html#decl-43df2c8d58ee","parent":"module:BanditRLProof.Exp3HighProbabilityRegret","order":4143,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HighProbabilityRegret"],["Source","BanditRLProof/Exp3HighProbabilityRegret.lean:171"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_highProbabilityRegret_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableHighProbabilityRegretBudget arms eta gamma horizon delta <= (Finset.range horizon).sum (fun t…","missing":[],"search":"sampledpredictable_highprobabilityregret_tail_delta banditrlproof.exp3.sampledpredictable_highprobabilityregret_tail_delta generated predictable exp3 pseudo-regret exceeds the explicit hedge, exploration, estimator-square, and two confidence-radius budget with probability at most the sum of the pure-q and comparator failure probabilities. theorem compiled","shard":"modules/0823a558d097504b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_total_delta","label":"sampledPredictable_highProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_total_delta","description":"Standard total-failure-probability form of the generated predictable EXP3 pseudo-regret bound. Each of the two confidence events receives `delta / 2`, so their union has probability at most `delta`.","url":"../modules/banditrlproof-exp3highprobabilityregret/index.html#decl-99f8fec2f7a3","parent":"module:BanditRLProof.Exp3HighProbabilityRegret","order":4144,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3HighProbabilityRegret"],["Source","BanditRLProof/Exp3HighProbabilityRegret.lean:319"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_highProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableHighProbabilityRegretBudget arms eta gamma horizon (delta / 2) <= (Finset.range horizon…","missing":[],"search":"sampledpredictable_highprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_highprobabilityregret_tail_total_delta standard total-failure-probability form of the generated predictable exp3 pseudo-regret bound. each of the two confidence events receives `delta / 2`, so their union has probability at most `delta`. theorem compiled","shard":"modules/0823a558d097504b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.importanceWeightedLoss","label":"importanceWeightedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.importanceWeightedLoss","description":"Loss estimate that reveals only the loss of the sampled action.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-81a40157f5e0","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4145,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def importanceWeightedLoss {Action : Type u} (prob loss : Action -> Real) (chosen action : Action) : Real","missing":[],"search":"importanceweightedloss banditrlproof.exp3.importanceweightedloss loss estimate that reveals only the loss of the sampled action. definition compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedImportanceWeightedLoss","label":"mixedImportanceWeightedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedImportanceWeightedLoss","description":"The loss mixed by `prob` after replacing losses by one sampled estimate.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-30ed388e0860","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4146,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def mixedImportanceWeightedLoss {Action : Type u} (arms : Finset Action) (prob loss : Action -> Real) (chosen : Action) : Real","missing":[],"search":"mixedimportanceweightedloss banditrlproof.exp3.mixedimportanceweightedloss the loss mixed by `prob` after replacing losses by one sampled estimate. definition compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weightedImportanceWeightedLoss","label":"weightedImportanceWeightedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.weightedImportanceWeightedLoss","description":"An importance-weighted estimate mixed by weights that may differ from the sampling probabilities used in the estimator denominator.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-bba12250b4f7","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4147,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:38"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def weightedImportanceWeightedLoss {Action : Type u} (arms : Finset Action) (prob weight loss : Action -> Real) (chosen : Action) : Real","missing":[],"search":"weightedimportanceweightedloss banditrlproof.exp3.weightedimportanceweightedloss an importance-weighted estimate mixed by weights that may differ from the sampling probabilities used in the estimator denominator. definition compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss","label":"mixedSquaredImportanceWeightedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss","description":"The mixed square of one sampled importance-weighted loss vector.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-3b18dbe7d93b","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4148,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:45"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def mixedSquaredImportanceWeightedLoss {Action : Type u} (arms : Finset Action) (prob loss : Action -> Real) (chosen : Action) : Real","missing":[],"search":"mixedsquaredimportanceweightedloss banditrlproof.exp3.mixedsquaredimportanceweightedloss the mixed square of one sampled importance-weighted loss vector. definition compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.importanceWeightedLoss_nonneg","label":"importanceWeightedLoss_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.importanceWeightedLoss_nonneg","description":"theorem importanceWeightedLoss_nonneg {Action : Type u} [DecidableEq Action] {prob loss : Action -> Real} {chosen action : Action} (hprob : 0 <= prob action) (hloss : 0 <= loss action) : 0 <= importanceWeightedLoss prob loss chosen action","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-aac4bcac2b4f","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4149,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:50"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem importanceWeightedLoss_nonneg {Action : Type u} [DecidableEq Action] {prob loss : Action -> Real} {chosen action : Action} (hprob : 0 <= prob action) (hloss : 0 <= loss action) : 0 <= importanceWeightedLoss prob loss chosen action","missing":[],"search":"importanceweightedloss_nonneg banditrlproof.exp3.importanceweightedloss_nonneg theorem importanceweightedloss_nonneg {action : type u} [decidableeq action] {prob loss : action -> real} {chosen action : action} (hprob : 0 <= prob action) (hloss : 0 <= loss action) : 0 <= importanceweightedloss prob loss chosen action theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_importanceWeightedLoss_eq_loss","label":"sum_prob_mul_importanceWeightedLoss_eq_loss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_importanceWeightedLoss_eq_loss","description":"The probability-weighted finite sum of one coordinate recovers its loss.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-1804073231d4","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4150,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:60"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_importanceWeightedLoss_eq_loss {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (action : Action) (haction : action ∈ arms) (hprob : prob action ≠ 0) : arms.sum (fun chosen => prob chosen * importanceWeightedLoss prob loss chosen action) = loss action","missing":[],"search":"sum_prob_mul_importanceweightedloss_eq_loss banditrlproof.exp3.sum_prob_mul_importanceweightedloss_eq_loss the probability-weighted finite sum of one coordinate recovers its loss. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedImportanceWeightedLoss_eq_selectedLoss","label":"mixedImportanceWeightedLoss_eq_selectedLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedImportanceWeightedLoss_eq_selectedLoss","description":"Pathwise cancellation: the mixed estimate is the sampled loss.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-9d33480ee9cc","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4151,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:77"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedImportanceWeightedLoss_eq_selectedLoss {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (chosen : Action) (hchosen : chosen ∈ arms) (hprob : prob chosen ≠ 0) : mixedImportanceWeightedLoss arms prob loss chosen = loss chosen","missing":[],"search":"mixedimportanceweightedloss_eq_selectedloss banditrlproof.exp3.mixedimportanceweightedloss_eq_selectedloss pathwise cancellation: the mixed estimate is the sampled loss. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_mixedImportanceWeightedLoss_eq_mixedLoss","label":"sum_prob_mul_mixedImportanceWeightedLoss_eq_mixedLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_mixedImportanceWeightedLoss_eq_mixedLoss","description":"Averaging the pathwise mixed estimate recovers the true mixed loss.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-215bdc5844b5","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4152,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:93"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_mixedImportanceWeightedLoss_eq_mixedLoss {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hprob : forall action, action ∈ arms -> prob action ≠ 0) : arms.sum (fun chosen => prob chosen * mixedImportanceWeightedLoss arms prob loss chosen) = arms.sum (fun action => prob action * loss action)","missing":[],"search":"sum_prob_mul_mixedimportanceweightedloss_eq_mixedloss banditrlproof.exp3.sum_prob_mul_mixedimportanceweightedloss_eq_mixedloss averaging the pathwise mixed estimate recovers the true mixed loss. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_weightedImportanceWeightedLoss_eq_weightedLoss","label":"sum_prob_mul_weightedImportanceWeightedLoss_eq_weightedLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_weightedImportanceWeightedLoss_eq_weightedLoss","description":"Averaging an estimator mixed by arbitrary predictable weights recovers the same weighted true loss.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-6fb73b0e4048","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4153,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:107"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_weightedImportanceWeightedLoss_eq_weightedLoss {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob weight loss : Action -> Real) (hprob : forall action, action ∈ arms -> prob action ≠ 0) : arms.sum (fun chosen => prob chosen * weightedImportanceWeightedLoss arms prob weight loss chosen) = arms.sum (fun action => weight action * loss action)","missing":[],"search":"sum_prob_mul_weightedimportanceweightedloss_eq_weightedloss banditrlproof.exp3.sum_prob_mul_weightedimportanceweightedloss_eq_weightedloss averaging an estimator mixed by arbitrary predictable weights recovers the same weighted true loss. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss_eq_selectedLoss_sq_div","label":"mixedSquaredImportanceWeightedLoss_eq_selectedLoss_sq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss_eq_selectedLoss_sq_div","description":"Exact pathwise mixed-square formula for a sampled action.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-96ea42c57335","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4154,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:135"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredImportanceWeightedLoss_eq_selectedLoss_sq_div {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (chosen : Action) (hchosen : chosen ∈ arms) (hprob : prob chosen ≠ 0) : mixedSquaredImportanceWeightedLoss arms prob loss chosen = (loss chosen) ^ 2 / prob chosen","missing":[],"search":"mixedsquaredimportanceweightedloss_eq_selectedloss_sq_div banditrlproof.exp3.mixedsquaredimportanceweightedloss_eq_selectedloss_sq_div exact pathwise mixed-square formula for a sampled action. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_mixedSquaredImportanceWeightedLoss_eq_sum_loss_sq","label":"sum_prob_mul_mixedSquaredImportanceWeightedLoss_eq_sum_loss_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_mixedSquaredImportanceWeightedLoss_eq_sum_loss_sq","description":"Exact probability-weighted finite sum of the mixed estimator's square.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-ade174aaf6c2","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4155,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:151"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_mixedSquaredImportanceWeightedLoss_eq_sum_loss_sq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hprob : forall action, action ∈ arms -> prob action ≠ 0) : arms.sum (fun chosen => prob chosen * mixedSquaredImportanceWeightedLoss arms prob loss chosen) = arms.sum (fun action => (loss action) ^ 2)","missing":[],"search":"sum_prob_mul_mixedsquaredimportanceweightedloss_eq_sum_loss_sq banditrlproof.exp3.sum_prob_mul_mixedsquaredimportanceweightedloss_eq_sum_loss_sq exact probability-weighted finite sum of the mixed estimator's square. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_mixedSquaredImportanceWeightedLoss_le_card","label":"sum_prob_mul_mixedSquaredImportanceWeightedLoss_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_mixedSquaredImportanceWeightedLoss_le_card","description":"For losses in `[0,1]`, the probability-weighted mixed square is at most the arm count.","url":"../modules/banditrlproof-exp3importanceweighted/index.html#decl-b5a99ad46a97","parent":"module:BanditRLProof.Exp3ImportanceWeighted","order":4156,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ImportanceWeighted"],["Source","BanditRLProof/Exp3ImportanceWeighted.lean:166"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_mixedSquaredImportanceWeightedLoss_le_card {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hprob : forall action, action ∈ arms -> prob action ≠ 0) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => prob chosen * mixedSquaredImportanceWeightedLoss arms prob loss chosen) <= arms.card","missing":[],"search":"sum_prob_mul_mixedsquaredimportanceweightedloss_le_card banditrlproof.exp3.sum_prob_mul_mixedsquaredimportanceweightedloss_le_card for losses in `[0,1]`, the probability-weighted mixed square is at most the arm count. theorem compiled","shard":"modules/436ffdc4a38efa23.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredImportanceWeightedLoss_eq","label":"sum_prob_mul_sq_mixedSquaredImportanceWeightedLoss_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredImportanceWeightedLoss_eq","description":"Exact uncentered second moment of one mixed importance-weighted square.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-42043ef2427f","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4157,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_sq_mixedSquaredImportanceWeightedLoss_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hprob : forall action, action ∈ arms -> prob action ≠ 0) : arms.sum (fun chosen => prob chosen * (mixedSquaredImportanceWeightedLoss arms prob loss chosen) ^ 2) = arms.sum (fun action => (loss action) ^ 4 / prob action)","missing":[],"search":"sum_prob_mul_sq_mixedsquaredimportanceweightedloss_eq banditrlproof.exp3.sum_prob_mul_sq_mixedsquaredimportanceweightedloss_eq exact uncentered second moment of one mixed importance-weighted square. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_card_div_floor","label":"sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_card_div_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_card_div_floor","description":"The centered mixed-square score has second moment at most `K / epsilon` under a uniform probability floor.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-e3f12b46274e","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4158,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:41"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_card_div_floor {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hdist : FiniteActionDistribution arms prob) (epsilon : Real) (hepsilon : 0 < epsilon) (hfloor : forall action, action ∈ arms -> epsilon <= prob action) (hloss : forall action, action ∈ arms -> loss action ∈ Set.Icc (0 : Real) 1) : arms.sum (fun chosen => prob chosen * (mixedSquaredImportanceWeightedLoss arms prob loss chosen - arms.sum (fun action => (loss action) ^ 2)) ^ 2) <= (arms.card : Real) / epsilon","missing":[],"search":"sum_prob_mul_sq_mixedsquaredestimatordeviation_le_card_div_floor banditrlproof.exp3.sum_prob_mul_sq_mixedsquaredestimatordeviation_le_card_div_floor the centered mixed-square score has second moment at most `k / epsilon` under a uniform probability floor. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt","label":"finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt","description":"Fixed-tilt MGF budget for a centered mixed estimator square under a finite sampling law. The quadratic coefficient is `K / epsilon`, while the admissible tilt is controlled by the sharper range cap `epsilon`.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-733c482b33fb","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4159,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:111"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hdist : FiniteActionDistribution arms prob) (epsilon : Real) (hepsilon : 0 < epsilon) (hfloor : forall action, action ∈ arms -> epsilon <= prob action) (hloss : forall action, action ∈ arms -> loss action ∈ Set.Icc (0 : Real) 1) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) : Concentration.HasMGFUpperBoundAt (fun chosen => mixedSquaredImportanceWeightedLoss arms prob loss chosen - arms.sum (fun action => (loss action) ^ 2)) tilt (tilt ^ 2 * ((arms.card : Real) / epsilon)) (finiteActionMeasure arms prob)","missing":[],"search":"finiteactionmixedsquaredestimator_hasmgfupperboundat banditrlproof.exp3.finiteactionmixedsquaredestimator_hasmgfupperboundat fixed-tilt mgf budget for a centered mixed estimator square under a finite sampling law. the quadratic coefficient is `k / epsilon`, while the admissible tilt is controlled by the sharper range cap `epsilon`. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","label":"mixedSquaredEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","description":"A finite conditional action law supplies the fixed-tilt mixed-square MGF budget with second-moment coefficient `K / epsilon`.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-499be6151c89","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4160,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:232"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) (hcond : condDistrib action…","missing":[],"search":"mixedsquaredestimator_hascondmgfupperboundat_of_conddistrib_ae_eq_finiteactionkernel banditrlproof.exp3.mixedsquaredestimator_hascondmgfupperboundat_of_conddistrib_ae_eq_finiteactionkernel a finite conditional action law supplies the fixed-tilt mixed-square mgf budget with second-moment coefficient `k / epsilon`. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_hasCondMGFUpperBoundAt","label":"sampledPredictableMixedSquaredDeviation_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_hasCondMGFUpperBoundAt","description":"theorem sampledPredictableMixedSquaredDeviation_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_po…","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-300e7628e070","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4161,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:372"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sample…","missing":[],"search":"sampledpredictablemixedsquareddeviation_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablemixedsquareddeviation_zero_hascondmgfupperboundat theorem sampledpredictablemixedsquareddeviation_zero_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment concentration.hascondmgfupperboundat ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss 0) tilt (tilt ^ 2 * ((arms.card : real) / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_hasCondMGFUpperBoundAt","label":"sampledPredictableMixedSquaredDeviation_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_hasCondMGFUpperBoundAt","description":"theorem sampledPredictableMixedSquaredDeviation_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_po…","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-d58cea5bcda0","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4162,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:429"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Concentration.HasCondMGFUpperBoundAt ((inferInstance : M…","missing":[],"search":"sampledpredictablemixedsquareddeviation_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablemixedsquareddeviation_succ_hascondmgfupperboundat theorem sampledpredictablemixedsquareddeviation_succ_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) concentration.hascondmgfupperboundat ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss (n + 1)) tilt (tilt ^ 2 * ((arms.card : real) / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_fixedTilt","label":"sampledPredictableMixedSquaredDeviation_sum_tail_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_fixedTilt","description":"Fixed-tilt Bernstein tail for the centered predictable mixed-square process on the generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-e4a214521645","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4163,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:488"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_sum_tail_fixedTilt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) (threshold : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | threshold <= (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss i sample)}…","missing":[],"search":"sampledpredictablemixedsquareddeviation_sum_tail_fixedtilt banditrlproof.exp3.sampledpredictablemixedsquareddeviation_sum_tail_fixedtilt fixed-tilt bernstein tail for the centered predictable mixed-square process on the generated exp3 trajectory. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_bernstein_fixedTilt","label":"sampledPredictableObservedMixedSquared_sum_tail_bernstein_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_bernstein_fixedTilt","description":"Fixed-tilt Bernstein tail for the observed, uncentered mixed estimator-square sum on the generated trajectory.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-5c08d64c85ce","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4164,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:582"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_bernstein_fixedTilt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) (threshold : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + threshold <= sampledObservedMixedSquaredSum arms eta gamma horizon sample} <= ENNRe…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_bernstein_fixedtilt banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_bernstein_fixedtilt fixed-tilt bernstein tail for the observed, uncentered mixed estimator-square sum on the generated trajectory. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredBernsteinVarianceCoefficient","label":"sampledMixedSquaredBernsteinVarianceCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredBernsteinVarianceCoefficient","description":"Variance coefficient in the mixed-square fixed-tilt MGF budget.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-916ff2d4ff57","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4165,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:628"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledMixedSquaredBernsteinVarianceCoefficient {Action : Type v} (arms : Finset Action) (gamma : Real) : Real","missing":[],"search":"sampledmixedsquaredbernsteinvariancecoefficient banditrlproof.exp3.sampledmixedsquaredbernsteinvariancecoefficient variance coefficient in the mixed-square fixed-tilt mgf budget. definition compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredBernsteinConfidenceRadius","label":"sampledMixedSquaredBernsteinConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredBernsteinConfidenceRadius","description":"Optimized Bernstein radius for the observed mixed estimator-square sum.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-787ef9892204","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4166,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:633"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledMixedSquaredBernsteinConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledmixedsquaredbernsteinconfidenceradius banditrlproof.exp3.sampledmixedsquaredbernsteinconfidenceradius optimized bernstein radius for the observed mixed estimator-square sum. definition compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_bernstein_delta","label":"sampledPredictableObservedMixedSquared_sum_tail_bernstein_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_bernstein_delta","description":"Delta-shaped variance-sensitive confidence bound for the generated observed mixed estimator-square sum. The square-root term uses the second-moment coefficient `K / epsilon`, while the linear correction uses the reciprocal range cap `1 / epsilon`.","url":"../modules/banditrlproof-exp3mixedsquarebernstein/index.html#decl-8eea48538266","parent":"module:BanditRLProof.Exp3MixedSquareBernstein","order":4167,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernstein"],["Source","BanditRLProof/Exp3MixedSquareBernstein.lean:645"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_bernstein_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + sampledMixedSquaredBernsteinConfidenceRadius arms gamma horizon delta <= sampledObservedMixedSquaredSum arms eta gamma horizon sample} <= ENNReal.ofReal delta","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_bernstein_delta banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_bernstein_delta delta-shaped variance-sensitive confidence bound for the generated observed mixed estimator-square sum. the square-root term uses the second-moment coefficient `k / epsilon`, while the linear correction uses the reciprocal range cap `1 / epsilon`. theorem compiled","shard":"modules/4f2dbc86b81509b4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinSquareHighProbabilityRegretBudget","label":"sampledPredictableBernsteinSquareHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBernsteinSquareHighProbabilityRegretBudget","description":"Predictable regret budget using the variance-sensitive mixed-square Bernstein radius.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinhighprobabilityregret/index.html#decl-571d1ad55137","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","order":4168,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareBernsteinHighProbabilityRegret.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableBernsteinSquareHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (deltaSquare deltaConfidence : Real) : Real","missing":[],"search":"sampledpredictablebernsteinsquarehighprobabilityregretbudget banditrlproof.exp3.sampledpredictablebernsteinsquarehighprobabilityregretbudget predictable regret budget using the variance-sensitive mixed-square bernstein radius. definition compiled","shard":"modules/e567f4449dfa2df4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareHighProbabilityRegret_tail","label":"sampledPredictable_bernsteinSquareHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinSquareHighProbabilityRegret_tail","description":"Generated predictable EXP3 regret with the mixed-square, pure-cross, and fixed-comparator Bernstein events.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinhighprobabilityregret/index.html#decl-3d412168b74b","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","order":4169,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareBernsteinHighProbabilityRegret.lean:37"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinSquareHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (deltaSquare deltaConfidence : Real) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinSquareHighProbabilityRegr…","missing":[],"search":"sampledpredictable_bernsteinsquarehighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_bernsteinsquarehighprobabilityregret_tail generated predictable exp3 regret with the mixed-square, pure-cross, and fixed-comparator bernstein events. theorem compiled","shard":"modules/e567f4449dfa2df4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_bernsteinSquareHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinSquareHighProbabilityRegret_tail_total_delta","description":"Total-failure form with all three Bernstein events allocated `delta / 3`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinhighprobabilityregret/index.html#decl-d330d52efd24","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","order":4170,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareBernsteinHighProbabilityRegret.lean:199"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinSquareHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinSquareHighProbabilityRegretBudget arms eta gamma horizon (delta / 3) (delta / 3) <= (Fin…","missing":[],"search":"sampledpredictable_bernsteinsquarehighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_bernsteinsquarehighprobabilityregret_tail_total_delta total-failure form with all three bernstein events allocated `delta / 3`. theorem compiled","shard":"modules/e567f4449dfa2df4.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareLargeHorizonCondition","label":"bernsteinSquareLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareLargeHorizonCondition","description":"The regime in which all four components of the reused variance-sensitive exploration schedule are at most one half.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedallhorizon/index.html#decl-6ed45d4bec34","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","order":4171,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedAllHorizon.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def bernsteinSquareLargeHorizonCondition (K T delta : Real) : Prop","missing":[],"search":"bernsteinsquarelargehorizoncondition banditrlproof.exp3.bernsteinsquarelargehorizoncondition the regime in which all four components of the reused variance-sensitive exploration schedule are at most one half. definition compiled","shard":"modules/c9017662babd9653.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareAllHorizonRegretThreshold","label":"bernsteinSquareAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareAllHorizonRegretThreshold","description":"All-horizon threshold for the variance-sensitive mixed-square route: use the explicit large-horizon rate in its valid regime and `T + 1` otherwise.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedallhorizon/index.html#decl-172761ca4388","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","order":4172,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedAllHorizon.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"bernsteinsquareallhorizonregretthreshold banditrlproof.exp3.bernsteinsquareallhorizonregretthreshold all-horizon threshold for the variance-sensitive mixed-square route: use the explicit large-horizon rate in its valid regime and `t + 1` otherwise. definition compiled","shard":"modules/c9017662babd9653.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareRealizedRegret_tail","label":"sampledPredictable_allHorizonBernsteinSquareRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareRealizedRegret_tail","description":"Generated realized-regret tail for every positive horizon under the exact variance-sensitive learning rate and reused clipped exploration schedule. The refined threshold is used precisely in the four-contract regime; the complementary branch is the genuine zero-probability `T + 1` fallback.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedallhorizon/index.html#decl-8da7c90c15cf","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","order":4173,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedAllHorizon.lean:49"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonBernsteinSquareRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := bernsteinSquareClippedExplorationRate (arms.card : Real) (horizon : Real) delta let eta := bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma (bernsteinSquareCli…","missing":[],"search":"sampledpredictable_allhorizonbernsteinsquarerealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonbernsteinsquarerealizedregret_tail generated realized-regret tail for every positive horizon under the exact variance-sensitive learning rate and reused clipped exploration schedule. the refined threshold is used precisely in the four-contract regime; the complementary branch is the genuine zero-probability `t + 1` fallback. theorem compiled","shard":"modules/c9017662babd9653.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareBestArmAllHorizonRegretThreshold","label":"bernsteinSquareBestArmAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareBestArmAllHorizonRegretThreshold","description":"Best-arm Bernstein mixed-square all-horizon threshold. The underlying fixed-comparator schedule receives the confidence share `delta / K`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedbestarmallhorizon/index.html#decl-78c9ef0ed9cc","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","order":4174,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedBestArmAllHorizon.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareBestArmAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"bernsteinsquarebestarmallhorizonregretthreshold banditrlproof.exp3.bernsteinsquarebestarmallhorizonregretthreshold best-arm bernstein mixed-square all-horizon threshold. the underlying fixed-comparator schedule receives the confidence share `delta / k`. definition compiled","shard":"modules/b6cad411093bcbaa.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail","label":"sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail","description":"Generated realized-regret tail against the best supported arm in hindsight for every positive horizon. The finite comparator union spends `delta / K` on each arm and therefore has total failure probability at most `delta`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedbestarmallhorizon/index.html#decl-d4cd9cffe66a","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","order":4175,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedBestArmAllHorizon.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := bernsteinSquareClippedExplorationRate (arms.card : Real) (horizon : Real) deltaArm let eta := bernsteinSquareHighProbabilityLearningRate arms gamma horizon deltaArm let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma (bernsteinSquareCli…","missing":[],"search":"sampledpredictable_allhorizonbernsteinsquarebestarmrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonbernsteinsquarebestarmrealizedregret_tail generated realized-regret tail against the best supported arm in hindsight for every positive horizon. the finite comparator union spends `delta / k` on each arm and therefore has total failure probability at most `delta`. theorem compiled","shard":"modules/b6cad411093bcbaa.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredBernsteinVarianceCoefficient_eq_card_sq_div_gamma","label":"sampledMixedSquaredBernsteinVarianceCoefficient_eq_card_sq_div_gamma","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredBernsteinVarianceCoefficient_eq_card_sq_div_gamma","description":"The deterministic mixed-square Bernstein variance coefficient is exactly `K^2 / gamma`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-63bdeebb112e","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4176,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledMixedSquaredBernsteinVarianceCoefficient_eq_card_sq_div_gamma {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) : sampledMixedSquaredBernsteinVarianceCoefficient arms gamma = (arms.card : Real) ^ 2 / gamma","missing":[],"search":"sampledmixedsquaredbernsteinvariancecoefficient_eq_card_sq_div_gamma banditrlproof.exp3.sampledmixedsquaredbernsteinvariancecoefficient_eq_card_sq_div_gamma the deterministic mixed-square bernstein variance coefficient is exactly `k^2 / gamma`. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_mul_sampledMixedSquaredBernsteinConfidenceRadius_le_three_mul_sq_mul_horizon_sq","label":"log_mul_sampledMixedSquaredBernsteinConfidenceRadius_le_three_mul_sq_mul_horizon_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_mul_sampledMixedSquaredBernsteinConfidenceRadius_le_three_mul_sq_mul_horizon_sq","description":"Under the existing arm, sixth-power mixed, and confidence contracts, the log-weighted Bernstein mixed-square radius is at most `3 * gamma^2 * T^2`. The first two copies control the square-root term; the third controls the linear `log_+ / epsilon` correction.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-20d34582aa10","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4177,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:38"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_mul_sampledMixedSquaredBernsteinConfidenceRadius_le_three_mul_sq_mul_horizon_sq {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.log (arms.card : Real) * sampledMixedSquaredBernsteinConfidenceRadius arms gamma horizon (delta / 4) <= 3 * gamma ^ 2 * (horizon : Real) ^ 2","missing":[],"search":"log_mul_sampledmixedsquaredbernsteinconfidenceradius_le_three_mul_sq_mul_horizon_sq banditrlproof.exp3.log_mul_sampledmixedsquaredbernsteinconfidenceradius_le_three_mul_sq_mul_horizon_sq under the existing arm, sixth-power mixed, and confidence contracts, the log-weighted bernstein mixed-square radius is at most `3 * gamma^2 * t^2`. the first two copies control the square-root term; the third controls the linear `log_+ / epsilon` correction. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","label":"bernsteinSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","description":"The armwise base term and the variance-sensitive mixed-square radius make the learning-rate-balanced square root at most `2 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-a1b71a54dec1","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4178,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:162"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareBalancedSqrt_le_two_mul_gamma_mul_horizon {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.sqrt (Real.log (arms.card : Real) * bernsteinSquareHighProbabilityScale arms gamma horizon delta) <= 2 * gamma * (horizon : Real)","missing":[],"search":"bernsteinsquarebalancedsqrt_le_two_mul_gamma_mul_horizon banditrlproof.exp3.bernsteinsquarebalancedsqrt_le_two_mul_gamma_mul_horizon the armwise base term and the variance-sensitive mixed-square radius make the learning-rate-balanced square root at most `2 * gamma * t`. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareRealizedExplicitThreshold","label":"bernsteinSquareRealizedExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareRealizedExplicitThreshold","description":"Explicit threshold after controlling the balanced square root and all three remaining confidence contributions by the exploration scale.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-8df4eb08a6c8","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4179,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:220"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareRealizedExplicitThreshold {Action : Type v} (_arms : Finset Action) (gamma : Real) (horizon : Nat) (_delta : Real) : Real","missing":[],"search":"bernsteinsquarerealizedexplicitthreshold banditrlproof.exp3.bernsteinsquarerealizedexplicitthreshold explicit threshold after controlling the balanced square root and all three remaining confidence contributions by the exploration scale. definition compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareRealizedTunedThreshold_le_explicitThreshold","label":"bernsteinSquareRealizedTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareRealizedTunedThreshold_le_explicitThreshold","description":"The existing quadratic, sixth-power, cubic, and realized quadratic contracts reduce the variance-sensitive tuned threshold to `14 * gamma * T`. The sixth-power contract is conservative for the new square-root term, while the arm and confidence contracts jointly control the linear correction.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-65774346f074","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4180,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:229"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareRealizedTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon : Nat) (hhorizon : 0 < horizon) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) (hrealized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * (horizon : Real)) : bernsteinSquareRealizedTunedThreshold arms gamma horizon delta <= bernsteinSquareRealizedExplicitThreshold arms gamma horiz…","missing":[],"search":"bernsteinsquarerealizedtunedthreshold_le_explicitthreshold banditrlproof.exp3.bernsteinsquarerealizedtunedthreshold_le_explicitthreshold the existing quadratic, sixth-power, cubic, and realized quadratic contracts reduce the variance-sensitive tuned threshold to `14 * gamma * t`. the sixth-power contract is conservative for the new square-root term, while the arm and confidence contracts jointly control the linear correction. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedBernsteinSquareRealizedRegret_tail","label":"sampledPredictable_gammaCharacterizedBernsteinSquareRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedBernsteinSquareRealizedRegret_tail","description":"Generated realized-regret tail under the four algebraic exploration contracts.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-0c3d65bee0b0","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4181,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:318"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedBernsteinSquareRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma…","missing":[],"search":"sampledpredictable_gammacharacterizedbernsteinsquarerealizedregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizedbernsteinsquarerealizedregret_tail generated realized-regret tail under the four algebraic exploration contracts. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate","label":"bernsteinSquareClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate","description":"The variance-sensitive route reuses the already compiled four-scale clipped schedule.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-0a33eaf3f78a","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4182,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:393"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareClippedExplorationRate (K T delta : Real) : Real","missing":[],"search":"bernsteinsquareclippedexplorationrate banditrlproof.exp3.bernsteinsquareclippedexplorationrate the variance-sensitive route reuses the already compiled four-scale clipped schedule. definition compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_pos","label":"bernsteinSquareClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_pos","description":"theorem bernsteinSquareClippedExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < bernsteinSquareClippedExplorationRate K T delta","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-10c388fc49e4","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4183,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:397"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareClippedExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < bernsteinSquareClippedExplorationRate K T delta","missing":[],"search":"bernsteinsquareclippedexplorationrate_pos banditrlproof.exp3.bernsteinsquareclippedexplorationrate_pos theorem bernsteinsquareclippedexplorationrate_pos (k t delta : real) (hk_one : 1 < k) (ht : 0 < t) : 0 < bernsteinsquareclippedexplorationrate k t delta theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_le_half","label":"bernsteinSquareClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_le_half","description":"theorem bernsteinSquareClippedExplorationRate_le_half (K T delta : Real) : bernsteinSquareClippedExplorationRate K T delta <= 1 / 2","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-713b762a2316","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4184,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:404"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareClippedExplorationRate_le_half (K T delta : Real) : bernsteinSquareClippedExplorationRate K T delta <= 1 / 2","missing":[],"search":"bernsteinsquareclippedexplorationrate_le_half banditrlproof.exp3.bernsteinsquareclippedexplorationrate_le_half theorem bernsteinsquareclippedexplorationrate_le_half (k t delta : real) : bernsteinsquareclippedexplorationrate k t delta <= 1 / 2 theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_contracts","label":"bernsteinSquareClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_contracts","description":"The reused clipped schedule satisfies all four contracts needed by the variance-sensitive gamma-characterized theorem.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-a6ffc3f4f4d0","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4185,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:412"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareClippedExplorationRate_contracts (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (K * Real.log K) <= T) (hlarge_mixed : 64 * (K ^ 2 * Real.log K ^ 2 * Real.log (4 / delta) / 2) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : let gamma := bernsteinSquareClippedExplorationRate K T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ K * Real.log K <= gamma ^ 2 * T ∧ K ^ 2 * Real.log K ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * T ^ 3 ∧ K * Real.log (4 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * T","missing":[],"search":"bernsteinsquareclippedexplorationrate_contracts banditrlproof.exp3.bernsteinsquareclippedexplorationrate_contracts the reused clipped schedule satisfies all four contracts needed by the variance-sensitive gamma-characterized theorem. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitBernsteinSquareRealizedRegret_tail","label":"sampledPredictable_explicitBernsteinSquareRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitBernsteinSquareRealizedRegret_tail","description":"Fully explicit generated realized-regret tail for the variance-sensitive route, using the reused clipped maximum of four exploration scales.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedexplicittuning/index.html#decl-395ab7c684dc","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","order":4186,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedExplicitTuning.lean:438"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitBernsteinSquareRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((arms.card : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 64 * ((arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2) <= (horizon : Real) ^ 3) (hlarge_confidence : 8 * ((arms.card : Real) * Real.log…","missing":[],"search":"sampledpredictable_explicitbernsteinsquarerealizedregret_tail banditrlproof.exp3.sampledpredictable_explicitbernsteinsquarerealizedregret_tail fully explicit generated realized-regret tail for the variance-sensitive route, using the reused clipped maximum of four exploration scales. theorem compiled","shard":"modules/a4e6b128fc438c5d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget","label":"sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget","description":"Realized selected-loss regret budget whose predictable component uses the variance-sensitive mixed estimator-square event and two Bernstein confidence radii.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedhighprobabilityregret/index.html#decl-f587c2ef0cae","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","order":4187,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedHighProbabilityRegret.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictablebernsteinsquarerealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablebernsteinsquarerealizedhighprobabilityregretbudget realized selected-loss regret budget whose predictable component uses the variance-sensitive mixed estimator-square event and two bernstein confidence radii. definition compiled","shard":"modules/6024f47d9033dc0e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail","label":"sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail","description":"Raw four-event form. The predictable component contributes the variance-sensitive mixed-square event and two Bernstein confidence events; the fourth event is the bounded realized-minus-predictable deviation.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedhighprobabilityregret/index.html#decl-cc5f94210a68","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","order":4188,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedHighProbabilityRegret.lean:36"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (deltaSquare deltaConfidence deltaRealized : Real) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.l…","missing":[],"search":"sampledpredictable_bernsteinsquarerealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_bernsteinsquarerealizedhighprobabilityregret_tail raw four-event form. the predictable component contributes the variance-sensitive mixed-square event and two bernstein confidence events; the fourth event is the bounded realized-minus-predictable deviation. theorem compiled","shard":"modules/6024f47d9033dc0e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail_total_delta","description":"Total-failure form: the mixed-square, pure-cross Bernstein, fixed-comparator Bernstein, and realized-deviation events each receive `delta / 4`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedhighprobabilityregret/index.html#decl-b07b4e5e1696","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","order":4189,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedHighProbabilityRegret.lean:157"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget arms eta gamm…","missing":[],"search":"sampledpredictable_bernsteinsquarerealizedhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_bernsteinsquarerealizedhighprobabilityregret_tail_total_delta total-failure form: the mixed-square, pure-cross bernstein, fixed-comparator bernstein, and realized-deviation events each receive `delta / 4`. theorem compiled","shard":"modules/6024f47d9033dc0e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityScale","label":"bernsteinSquareHighProbabilityScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareHighProbabilityScale","description":"Positive scale appearing in the Bernstein-square Hedge term when the square event receives `delta / 4`.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-214ae76b37f1","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4190,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareHighProbabilityScale {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"bernsteinsquarehighprobabilityscale banditrlproof.exp3.bernsteinsquarehighprobabilityscale positive scale appearing in the bernstein-square hedge term when the square event receives `delta / 4`. definition compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityScale_pos","label":"bernsteinSquareHighProbabilityScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareHighProbabilityScale_pos","description":"theorem bernsteinSquareHighProbabilityScale_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < bernsteinSquareHighProbabilityScale arms gamma horizon delta","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-6e54efab1d27","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4191,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareHighProbabilityScale_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < bernsteinSquareHighProbabilityScale arms gamma horizon delta","missing":[],"search":"bernsteinsquarehighprobabilityscale_pos banditrlproof.exp3.bernsteinsquarehighprobabilityscale_pos theorem bernsteinsquarehighprobabilityscale_pos {action : type v} (arms : finset action) (hcard_two : 2 <= arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) : 0 < bernsteinsquarehighprobabilityscale arms gamma horizon delta theorem compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate","label":"bernsteinSquareHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate","description":"Learning rate balancing entropy against the complete Bernstein-square stability scale at the public `delta / 4` square allocation.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-616d0c068642","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4192,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:61"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareHighProbabilityLearningRate {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"bernsteinsquarehighprobabilitylearningrate banditrlproof.exp3.bernsteinsquarehighprobabilitylearningrate learning rate balancing entropy against the complete bernstein-square stability scale at the public `delta / 4` square allocation. definition compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate_pos","label":"bernsteinSquareHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate_pos","description":"theorem bernsteinSquareHighProbabilityLearningRate_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-db9c9a9d641c","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4193,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:68"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareHighProbabilityLearningRate_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta","missing":[],"search":"bernsteinsquarehighprobabilitylearningrate_pos banditrlproof.exp3.bernsteinsquarehighprobabilitylearningrate_pos theorem bernsteinsquarehighprobabilitylearningrate_pos {action : type v} (arms : finset action) (hcard_two : 2 <= arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) : 0 < bernsteinsquarehighprobabilitylearningrate arms gamma horizon delta theorem compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate_sq_mul_scale","label":"bernsteinSquareHighProbabilityLearningRate_sq_mul_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate_sq_mul_scale","description":"theorem bernsteinSquareHighProbabilityLearningRate_sq_mul_scale {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta ^ 2 * bernsteinSquareHighProbabilityScale arms gamma horizon delta = Real.log (arms.card : Real)","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-7ef3443d34ce","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4194,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:81"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareHighProbabilityLearningRate_sq_mul_scale {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta ^ 2 * bernsteinSquareHighProbabilityScale arms gamma horizon delta = Real.log (arms.card : Real)","missing":[],"search":"bernsteinsquarehighprobabilitylearningrate_sq_mul_scale banditrlproof.exp3.bernsteinsquarehighprobabilitylearningrate_sq_mul_scale theorem bernsteinsquarehighprobabilitylearningrate_sq_mul_scale {action : type v} (arms : finset action) (hcard_two : 2 <= arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) : bernsteinsquarehighprobabilitylearningrate arms gamma horizon delta ^ 2 * bernsteinsquarehighprobabilityscale arms gamma horizon delta = real.log (arms.card : real) theorem compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","label":"bernsteinSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","description":"With `gamma <= 1/2`, entropy and the stability-amplified Bernstein-square scale cost at most three copies of their balanced square-root scale.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-75ca167f44b5","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4195,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:99"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem bernsteinSquareHighProbabilityHedgeBudget_le_three_mul_sqrt {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : Real.log (arms.card : Real) / bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta + (bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta * (1 / (1 - gamma))) * bernsteinSquareHighProbabilityScale arms gamma horizon delta <= 3 * Real.sqrt (Real.log (arms.card : Real) * bernsteinSquareHighProbabilityScale arms gamma horizon delta)","missing":[],"search":"bernsteinsquarehighprobabilityhedgebudget_le_three_mul_sqrt banditrlproof.exp3.bernsteinsquarehighprobabilityhedgebudget_le_three_mul_sqrt with `gamma <= 1/2`, entropy and the stability-amplified bernstein-square scale cost at most three copies of their balanced square-root scale. theorem compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.bernsteinSquareRealizedTunedThreshold","label":"bernsteinSquareRealizedTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.bernsteinSquareRealizedTunedThreshold","description":"Explicit threshold after tuning all learning-rate-dependent terms. The exploration and three confidence contributions remain visible.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-0186959948cf","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4196,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:170"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinSquareRealizedTunedThreshold {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"bernsteinsquarerealizedtunedthreshold banditrlproof.exp3.bernsteinsquarerealizedtunedthreshold explicit threshold after tuning all learning-rate-dependent terms. the exploration and three confidence contributions remain visible. definition compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget_le_tunedThreshold","description":"The complete Bernstein-square four-event realized budget is bounded by the learning-rate-tuned threshold.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-2bc9a1f73edf","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4197,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:185"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget arms (bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta) gamma horizon (delta / 4) (delta / 4) (delta / 4) <= bernsteinSquareRealizedTunedThreshold arms gamma horizon delta","missing":[],"search":"sampledpredictablebernsteinsquarerealizedhighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictablebernsteinsquarerealizedhighprobabilityregretbudget_le_tunedthreshold the complete bernstein-square four-event realized budget is bounded by the learning-rate-tuned threshold. theorem compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedBernsteinSquareRealizedRegret_tail","label":"sampledPredictable_tunedBernsteinSquareRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedBernsteinSquareRealizedRegret_tail","description":"Generated realized-regret tail with the Bernstein-square-balanced learning rate. Gamma scheduling and confidence-radius simplification remain for downstream consumers.","url":"../modules/banditrlproof-exp3mixedsquarebernsteinrealizedtuning/index.html#decl-bbe5708c773b","parent":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","order":4198,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareBernsteinRealizedTuning.lean:211"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedBernsteinSquareRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let eta := bernsteinSquareHighProbabilityLearningRate arms gamma horizon delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le (by linarith : gamma <= 1) loss.environment mu {sample | b…","missing":[],"search":"sampledpredictable_tunedbernsteinsquarerealizedregret_tail banditrlproof.exp3.sampledpredictable_tunedbernsteinsquarerealizedregret_tail generated realized-regret tail with the bernstein-square-balanced learning rate. gamma scheduling and confidence-radius simplification remain for downstream consumers. theorem compiled","shard":"modules/f519189cf91b7ea6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss_le_inv_floor","label":"mixedSquaredImportanceWeightedLoss_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss_le_inv_floor","description":"Under a positive probability floor and unit losses, the mixed estimator square has the sharper reciprocal-floor bound, rather than the generic square of that reciprocal.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-472366d14ac2","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4199,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredImportanceWeightedLoss_le_inv_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (chosen : Action) : mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) chosen <= 1 / epsilon","missing":[],"search":"mixedsquaredimportanceweightedloss_le_inv_floor banditrlproof.exp3.mixedsquaredimportanceweightedloss_le_inv_floor under a positive probability floor and unit losses, the mixed estimator square has the sharper reciprocal-floor bound, rather than the generic square of that reciprocal. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","label":"mixedSquaredEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","description":"A finite-action mixed estimator square, centered by its exact conditional mean, is conditionally sub-Gaussian whenever the selected action has the stated finite-action conditional distribution.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-420b8d27b65b","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4200,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:64"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (hcond : condDistrib action history mu =ᵐ[mu.map history] finiteActionKernel arms prob source) : P…","missing":[],"search":"mixedsquaredestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel banditrlproof.exp3.mixedsquaredestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel a finite-action mixed estimator square, centered by its exact conditional mean, is conditionally sub-gaussian whenever the selected action has the stated finite-action conditional distribution. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredDeviationAt","label":"sampledTrajectoryPredictableMixedSquaredDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredDeviationAt","description":"The predictable mixed-square deviation at an actual generated time.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-3396d73d84fc","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4201,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:218"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPredictableMixedSquaredDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypredictablemixedsquareddeviationat banditrlproof.exp3.sampledtrajectorypredictablemixedsquareddeviationat the predictable mixed-square deviation at an actual generated time. definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_hasCondSubgaussianMGF","label":"sampledPredictableMixedSquaredDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledPredictableMixedSquaredDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-81e8cc1403d3","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4202,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:229"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss 0) (Concentration.int…","missing":[],"search":"sampledpredictablemixedsquareddeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictablemixedsquareddeviation_zero_hascondsubgaussianmgf theorem sampledpredictablemixedsquareddeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss 0) (concentration.intervalvarianceproxy 0 (1 / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_hasCondSubgaussianMGF","label":"sampledPredictableMixedSquaredDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledPredictableMixedSquaredDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-199a8051f0d1","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4203,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:283"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real n)).comap history) (measura…","missing":[],"search":"sampledpredictablemixedsquareddeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictablemixedsquareddeviation_succ_hascondsubgaussianmgf theorem sampledpredictablemixedsquareddeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss (n + 1)) (concentration.intervalvarianceproxy 0 (1 / (gamma / (arms.card : real)))) mu theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess","label":"sampledPredictableMixedSquaredDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess","description":"Shift the actual-time predictable mixed-square deviations by one so that the process starts with the deterministic zero required by the finite-sum conditional concentration API.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-b79566c9260a","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4204,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:340"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableMixedSquaredDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss i sample noncomputable def sampledMixedSquaredVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","missing":[],"search":"sampledpredictablemixedsquareddeviationprocess banditrlproof.exp3.sampledpredictablemixedsquareddeviationprocess shift the actual-time predictable mixed-square deviations by one so that the process starts with the deterministic zero required by the finite-sum conditional concentration api. definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy","label":"sampledMixedSquaredVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy","description":"noncomputable def sampledMixedSquaredVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-032fbcff1e96","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4205,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:351"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledMixedSquaredVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","missing":[],"search":"sampledmixedsquaredvarianceproxy banditrlproof.exp3.sampledmixedsquaredvarianceproxy noncomputable def sampledmixedsquaredvarianceproxy {action : type v} (arms : finset action) (gamma : real) : nnreal definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredDeviationProxy","label":"sampledMixedSquaredDeviationProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredDeviationProxy","description":"noncomputable def sampledMixedSquaredDeviationProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : Nat -> NNReal | 0 => 0 | _i + 1 => sampledMixedSquaredVarianceProxy arms gamma set_option maxHeartbeats 800000 in theorem sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Decidable…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-c1147a93854b","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4206,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:356"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledMixedSquaredDeviationProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : Nat -> NNReal | 0 => 0 | _i + 1 => sampledMixedSquaredVarianceProxy arms gamma set_option maxHeartbeats 800000 in theorem sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPredictableMixedSquaredDeviationProcess arms eta gamma loss)","missing":[],"search":"sampledmixedsquareddeviationproxy banditrlproof.exp3.sampledmixedsquareddeviationproxy noncomputable def sampledmixedsquareddeviationproxy {action : type v} (arms : finset action) (gamma : real) : nat -> nnreal | 0 => 0 | _i + 1 => sampledmixedsquaredvarianceproxy arms gamma set_option maxheartbeats 800000 in theorem sampledpredictablemixedsquareddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablemixedsquareddeviationprocess arms eta gamma loss) definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted","label":"sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted","description":"theorem sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltr…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-45ad546dd691","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4207,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:362"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPredictableMixedSquaredDeviationProcess arms eta gamma loss)","missing":[],"search":"sampledpredictablemixedsquareddeviationprocess_stronglyadapted banditrlproof.exp3.sampledpredictablemixedsquareddeviationprocess_stronglyadapted theorem sampledpredictablemixedsquareddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablemixedsquareddeviationprocess arms eta gamma loss) theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_sum_range_succ","label":"sampledPredictableMixedSquaredDeviationProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_sum_range_succ","description":"theorem sampledPredictableMixedSquaredDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableMixedSquaredDeviationProcess arms eta g…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-1e4fb5458550","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4208,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:521"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableMixedSquaredDeviationProcess arms eta gamma loss i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss i sample)","missing":[],"search":"sampledpredictablemixedsquareddeviationprocess_sum_range_succ banditrlproof.exp3.sampledpredictablemixedsquareddeviationprocess_sum_range_succ theorem sampledpredictablemixedsquareddeviationprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpredictablemixedsquareddeviationprocess arms eta gamma loss i sample) = (finset.range horizon).sum (fun i => sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss i sample) theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredDeviationProxy_sum_range_succ","label":"sampledMixedSquaredDeviationProxy_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredDeviationProxy_sum_range_succ","description":"theorem sampledMixedSquaredDeviationProxy_sum_range_succ {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) : (Finset.range (horizon + 1)).sum (sampledMixedSquaredDeviationProxy arms gamma) = (horizon : NNReal) * sampledMixedSquaredVarianceProxy arms gamma","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-ac7de73e7cc6","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4209,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:541"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledMixedSquaredDeviationProxy_sum_range_succ {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) : (Finset.range (horizon + 1)).sum (sampledMixedSquaredDeviationProxy arms gamma) = (horizon : NNReal) * sampledMixedSquaredVarianceProxy arms gamma","missing":[],"search":"sampledmixedsquareddeviationproxy_sum_range_succ banditrlproof.exp3.sampledmixedsquareddeviationproxy_sum_range_succ theorem sampledmixedsquareddeviationproxy_sum_range_succ {action : type v} (arms : finset action) (gamma : real) (horizon : nat) : (finset.range (horizon + 1)).sum (sampledmixedsquareddeviationproxy arms gamma) = (horizon : nnreal) * sampledmixedsquaredvarianceproxy arms gamma theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_ennreal","label":"sampledPredictableMixedSquaredDeviation_sum_tail_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_ennreal","description":"Exponential tail for the centered predictable mixed estimator-square sum.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-c009a6c0a79a","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4210,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:556"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_sum_tail_ennreal {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | eps <= (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss i sample)} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((((horizon : NNReal) * sampledMixedSqu…","missing":[],"search":"sampledpredictablemixedsquareddeviation_sum_tail_ennreal banditrlproof.exp3.sampledpredictablemixedsquareddeviation_sum_tail_ennreal exponential tail for the centered predictable mixed estimator-square sum. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredSum","label":"sampledPredictableMixedSquaredSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredSum","description":"The latent predictable mixed estimator-square sum.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-e1cdeed69131","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4211,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:623"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableMixedSquaredSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledpredictablemixedsquaredsum banditrlproof.exp3.sampledpredictablemixedsquaredsum the latent predictable mixed estimator-square sum. definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredSum","label":"sampledPredictableLossSquaredSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossSquaredSum","description":"The sum of exact conditional means of the latent mixed-square scores.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-9dcef53b6da0","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4212,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:634"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableLossSquaredSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledpredictablelosssquaredsum banditrlproof.exp3.sampledpredictablelosssquaredsum the sum of exact conditional means of the latent mixed-square scores. definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredAt_le_card","label":"sampledPredictableLossSquaredAt_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossSquaredAt_le_card","description":"theorem sampledPredictableLossSquaredAt_le_card {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : arms.sum (fun candidate => (predictableLossAt loss t sample candidate) ^ 2) <= (arms.card : Real)","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-842525f43680","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4213,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:642"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossSquaredAt_le_card {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : arms.sum (fun candidate => (predictableLossAt loss t sample candidate) ^ 2) <= (arms.card : Real)","missing":[],"search":"sampledpredictablelosssquaredat_le_card banditrlproof.exp3.sampledpredictablelosssquaredat_le_card theorem sampledpredictablelosssquaredat_le_card {env : type u} {action : type v} [measurablespace env] [measurablespace action] (arms : finset action) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : arms.sum (fun candidate => (predictablelossat loss t sample candidate) ^ 2) <= (arms.card : real) theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredSum_le_card_mul","label":"sampledPredictableLossSquaredSum_le_card_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossSquaredSum_le_card_mul","description":"theorem sampledPredictableLossSquaredSum_le_card_mul {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledPredictableLossSquaredSum arms loss horizon sample <= (arms.card : Real) * (horizon : Real)","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-9c4d533444ea","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4214,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:666"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossSquaredSum_le_card_mul {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledPredictableLossSquaredSum arms loss horizon sample <= (arms.card : Real) * (horizon : Real)","missing":[],"search":"sampledpredictablelosssquaredsum_le_card_mul banditrlproof.exp3.sampledpredictablelosssquaredsum_le_card_mul theorem sampledpredictablelosssquaredsum_le_card_mul {env : type u} {action : type v} [measurablespace env] [measurablespace action] (arms : finset action) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : sampledpredictablelosssquaredsum arms loss horizon sample <= (arms.card : real) * (horizon : real) theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_eq","label":"sampledPredictableMixedSquaredDeviation_sum_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_eq","description":"theorem sampledPredictableMixedSquaredDeviation_sum_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range horizon).sum (fun t => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss t samp…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-5de34d7ed91b","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4215,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:683"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_sum_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range horizon).sum (fun t => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss t sample) = sampledPredictableMixedSquaredSum arms eta gamma loss horizon sample - sampledPredictableLossSquaredSum arms loss horizon sample","missing":[],"search":"sampledpredictablemixedsquareddeviation_sum_eq banditrlproof.exp3.sampledpredictablemixedsquareddeviation_sum_eq theorem sampledpredictablemixedsquareddeviation_sum_eq {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range horizon).sum (fun t => sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss t sample) = sampledpredictablemixedsquaredsum arms eta gamma loss horizon sample - sampledpredictablelosssquaredsum arms loss horizon sample theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquared_sum_tail_ennreal","label":"sampledPredictableMixedSquared_sum_tail_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquared_sum_tail_ennreal","description":"Exponential tail for the latent, uncentered mixed estimator-square sum.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-280ff83a717a","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4216,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:698"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquared_sum_tail_ennreal {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + eps <= sampledPredictableMixedSquaredSum arms eta gamma loss horizon sample} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((((horizon : NNReal) * sampledMixedSquaredVarianceProxy a…","missing":[],"search":"sampledpredictablemixedsquared_sum_tail_ennreal banditrlproof.exp3.sampledpredictablemixedsquared_sum_tail_ennreal exponential tail for the latent, uncentered mixed estimator-square sum. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedMixedSquaredSum_eq_predictable_ae","label":"sampledObservedMixedSquaredSum_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedMixedSquaredSum_eq_predictable_ae","description":"The observed scalar-feedback mixed-square sum agrees almost everywhere with the latent predictable mixed-square sum.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-3dd9043ad621","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4217,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:744"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedMixedSquaredSum_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment sampledObservedMixedSquaredSum arms eta gamma horizon =ᵐ[mu] sampledPredictableMixedSquaredSum arms eta gamma loss horizon","missing":[],"search":"sampledobservedmixedsquaredsum_eq_predictable_ae banditrlproof.exp3.sampledobservedmixedsquaredsum_eq_predictable_ae the observed scalar-feedback mixed-square sum agrees almost everywhere with the latent predictable mixed-square sum. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_exponential","label":"sampledPredictableObservedMixedSquared_sum_tail_exponential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_exponential","description":"Exponential confidence for the observed mixed estimator-square sum.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-f79fc4cd36f9","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4218,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:775"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_exponential {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + eps <= sampledObservedMixedSquaredSum arms eta gamma horizon sample} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((((horizon : NNReal) * sampledMixedSquaredVariancePro…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_exponential banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_exponential exponential confidence for the observed mixed estimator-square sum. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredConfidenceRadius","label":"sampledMixedSquaredConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredConfidenceRadius","description":"noncomputable def sampledMixedSquaredConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-4a41ff8826db","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4219,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:816"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledMixedSquaredConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledmixedsquaredconfidenceradius banditrlproof.exp3.sampledmixedsquaredconfidenceradius noncomputable def sampledmixedsquaredconfidenceradius {action : type v} (arms : finset action) (gamma : real) (horizon : nat) (delta : real) : real definition compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy_pos","label":"sampledMixedSquaredVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy_pos","description":"theorem sampledMixedSquaredVarianceProxy_pos {Action : Type v} (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) : 0 < ((sampledMixedSquaredVarianceProxy arms gamma : NNReal) : Real)","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-ba2564faa75d","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4220,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:824"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledMixedSquaredVarianceProxy_pos {Action : Type v} (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) : 0 < ((sampledMixedSquaredVarianceProxy arms gamma : NNReal) : Real)","missing":[],"search":"sampledmixedsquaredvarianceproxy_pos banditrlproof.exp3.sampledmixedsquaredvarianceproxy_pos theorem sampledmixedsquaredvarianceproxy_pos {action : type v} (arms : finset action) (harms : arms.nonempty) (gamma : real) (hgamma_pos : 0 < gamma) : 0 < ((sampledmixedsquaredvarianceproxy arms gamma : nnreal) : real) theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_exp_neg_budget","label":"sampledPredictableObservedMixedSquared_sum_tail_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_exp_neg_budget","description":"theorem sampledPredictableObservedMixedSquared_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_po…","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-80726d7e69c0","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4221,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:841"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (budget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + Real.sqrt (2 * ((((horizon : NNReal) * sampledMixedSquaredVarianceProxy arms gamma : NNReal)) : Real) * budget) <= sampledObservedMixedSquaredSum arms eta…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_exp_neg_budget banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_exp_neg_budget theorem sampledpredictableobservedmixedsquared_sum_tail_exp_neg_budget {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) (hhorizon : 0 < horizon) (budget : real) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : real) * (horizon : real) + real.sqrt (2 * ((((horizon : nnreal) * sampledmixedsquaredvarianceproxy arms gamma : nnreal)) : real) * budget) <= sampledobservedmixedsquaredsum arms eta gamma horizon sample} <= ennreal.ofreal (real.exp (-budget)) theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_delta","label":"sampledPredictableObservedMixedSquared_sum_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_delta","description":"Delta-shaped exponential confidence bound for the observed finite-horizon mixed estimator-square sum.","url":"../modules/banditrlproof-exp3mixedsquareconfidence/index.html#decl-190d256e33d5","parent":"module:BanditRLProof.Exp3MixedSquareConfidence","order":4222,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareConfidence"],["Source","BanditRLProof/Exp3MixedSquareConfidence.lean:889"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + sampledMixedSquaredConfidenceRadius arms gamma horizon delta <= sampledObservedMixedSquaredSum arms eta gamma horizon sample} <= ENNReal.ofReal…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_delta banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_delta delta-shaped exponential confidence bound for the observed finite-horizon mixed estimator-square sum. theorem compiled","shard":"modules/d70f7e49468531f1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinHighProbabilityRegretBudget","label":"sampledPredictableExponentialSquareBernsteinHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinHighProbabilityRegretBudget","description":"Predictable regret budget using the exponential mixed-square threshold instead of the Markov `|arms| * T / deltaSquare` threshold.","url":"../modules/banditrlproof-exp3mixedsquareexponentialhighprobabilityregret/index.html#decl-1cced0789f88","parent":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","order":4223,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareExponentialHighProbabilityRegret.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableExponentialSquareBernsteinHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (deltaSquare deltaConfidence : Real) : Real","missing":[],"search":"sampledpredictableexponentialsquarebernsteinhighprobabilityregretbudget banditrlproof.exp3.sampledpredictableexponentialsquarebernsteinhighprobabilityregretbudget predictable regret budget using the exponential mixed-square threshold instead of the markov `|arms| * t / deltasquare` threshold. definition compiled","shard":"modules/8e28484e2139b0f5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail","label":"sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail","description":"Generated predictable EXP3 regret with an exponential estimator-square event and the two compiled Bernstein confidence events.","url":"../modules/banditrlproof-exp3mixedsquareexponentialhighprobabilityregret/index.html#decl-05d405e228e5","parent":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","order":4224,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareExponentialHighProbabilityRegret.lean:39"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (deltaSquare deltaConfidence : Real) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictab…","missing":[],"search":"sampledpredictable_exponentialsquarebernsteinhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_exponentialsquarebernsteinhighprobabilityregret_tail generated predictable exp3 regret with an exponential estimator-square event and the two compiled bernstein confidence events. theorem compiled","shard":"modules/8e28484e2139b0f5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail_total_delta","description":"Total-failure form: the exponential square event and both Bernstein confidence events each receive `delta / 3`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialhighprobabilityregret/index.html#decl-393cdf0eaa56","parent":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","order":4225,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareExponentialHighProbabilityRegret.lean:201"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableExponentialSquareBernsteinHighProbabilityRegretBudget arms et…","missing":[],"search":"sampledpredictable_exponentialsquarebernsteinhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_exponentialsquarebernsteinhighprobabilityregret_tail_total_delta total-failure form: the exponential square event and both bernstein confidence events each receive `delta / 3`. theorem compiled","shard":"modules/8e28484e2139b0f5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinLargeHorizonCondition","label":"exponentialSquareBernsteinLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinLargeHorizonCondition","description":"The regime in which all four components of the exponential-square exploration schedule are at most one half.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedallhorizon/index.html#decl-96b9da68e582","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","order":4226,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedAllHorizon.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def exponentialSquareBernsteinLargeHorizonCondition (K T delta : Real) : Prop","missing":[],"search":"exponentialsquarebernsteinlargehorizoncondition banditrlproof.exp3.exponentialsquarebernsteinlargehorizoncondition the regime in which all four components of the exponential-square exploration schedule are at most one half. definition compiled","shard":"modules/5bb3146d79843e5b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinAllHorizonRegretThreshold","label":"exponentialSquareBernsteinAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinAllHorizonRegretThreshold","description":"All-horizon threshold for the exponential-square route: use the explicit large-horizon rate in its valid regime and `T + 1` otherwise.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedallhorizon/index.html#decl-92e034497bfc","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","order":4227,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedAllHorizon.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareBernsteinAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"exponentialsquarebernsteinallhorizonregretthreshold banditrlproof.exp3.exponentialsquarebernsteinallhorizonregretthreshold all-horizon threshold for the exponential-square route: use the explicit large-horizon rate in its valid regime and `t + 1` otherwise. definition compiled","shard":"modules/5bb3146d79843e5b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonExponentialSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_allHorizonExponentialSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonExponentialSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail for every positive horizon under the exact exponential-square learning rate and clipped exploration schedule. The refined threshold is used precisely in the four-contract regime; the complementary branch is the genuine zero-probability `T + 1` fallback.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedallhorizon/index.html#decl-b4dad78a7b64","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","order":4228,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedAllHorizon.lean:50"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonExponentialSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := exponentialSquareBernsteinClippedExplorationRate (arms.card : Real) (horizon : Real) delta let eta := exponentialSquareHighProbabilityLearningRate arms gamma horizon delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta g…","missing":[],"search":"sampledpredictable_allhorizonexponentialsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonexponentialsquarebernsteinrealizedregret_tail generated realized-regret tail for every positive horizon under the exact exponential-square learning rate and clipped exploration schedule. the refined threshold is used precisely in the four-contract regime; the complementary branch is the genuine zero-probability `t + 1` fallback. theorem compiled","shard":"modules/5bb3146d79843e5b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy_coe_eq_card_div_two_gamma_sq","label":"sampledMixedSquaredVarianceProxy_coe_eq_card_div_two_gamma_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy_coe_eq_card_div_two_gamma_sq","description":"The mixed-square interval proxy is exactly `(K / (2 * gamma))^2`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-4fba5b758af3","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4229,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledMixedSquaredVarianceProxy_coe_eq_card_div_two_gamma_sq {Action : Type v} (arms : Finset Action) (gamma : Real) (hgamma_pos : 0 < gamma) : ((sampledMixedSquaredVarianceProxy arms gamma : NNReal) : Real) = ((arms.card : Real) / (2 * gamma)) ^ 2","missing":[],"search":"sampledmixedsquaredvarianceproxy_coe_eq_card_div_two_gamma_sq banditrlproof.exp3.sampledmixedsquaredvarianceproxy_coe_eq_card_div_two_gamma_sq the mixed-square interval proxy is exactly `(k / (2 * gamma))^2`. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_mul_sampledMixedSquaredConfidenceRadius_le_sq_mul_horizon_sq","label":"log_mul_sampledMixedSquaredConfidenceRadius_le_sq_mul_horizon_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_mul_sampledMixedSquaredConfidenceRadius_le_sq_mul_horizon_sq","description":"The logarithmically weighted mixed-square confidence radius is controlled by `gamma^2 T^2` under the sixth-power dominance contract.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-5811664db08d","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4230,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:39"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_mul_sampledMixedSquaredConfidenceRadius_le_sq_mul_horizon_sq {Action : Type v} (arms : Finset Action) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * (horizon : Real) ^ 3) : Real.log (arms.card : Real) * sampledMixedSquaredConfidenceRadius arms gamma horizon (delta / 4) <= gamma ^ 2 * (horizon : Real) ^ 2","missing":[],"search":"log_mul_sampledmixedsquaredconfidenceradius_le_sq_mul_horizon_sq banditrlproof.exp3.log_mul_sampledmixedsquaredconfidenceradius_le_sq_mul_horizon_sq the logarithmically weighted mixed-square confidence radius is controlled by `gamma^2 t^2` under the sixth-power dominance contract. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","label":"exponentialSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","description":"The armwise base term and the mixed-square confidence radius together make the learning-rate-balanced square root at most `2 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-78d564b0ff58","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4231,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:105"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBalancedSqrt_le_two_mul_gamma_mul_horizon {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * (horizon : Real) ^ 3) : Real.sqrt (Real.log (arms.card : Real) * exponentialSquareHighProbabilityScale arms gamma horizon delta) <= 2 * gamma * (horizon : Real)","missing":[],"search":"exponentialsquarebalancedsqrt_le_two_mul_gamma_mul_horizon banditrlproof.exp3.exponentialsquarebalancedsqrt_le_two_mul_gamma_mul_horizon the armwise base term and the mixed-square confidence radius together make the learning-rate-balanced square root at most `2 * gamma * t`. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRealizedExplicitThreshold","label":"exponentialSquareBernsteinRealizedExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinRealizedExplicitThreshold","description":"Explicit threshold after controlling the balanced square root and all three confidence contributions by the exploration scale.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-f59c6a7099ce","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4232,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:151"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareBernsteinRealizedExplicitThreshold {Action : Type v} (_arms : Finset Action) (gamma : Real) (horizon : Nat) (_delta : Real) : Real","missing":[],"search":"exponentialsquarebernsteinrealizedexplicitthreshold banditrlproof.exp3.exponentialsquarebernsteinrealizedexplicitthreshold explicit threshold after controlling the balanced square root and all three confidence contributions by the exploration scale. definition compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","label":"exponentialSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","description":"Quadratic, sixth-power, cubic, and realized quadratic contracts reduce the learning-rate-tuned threshold to `14 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-9b801fb1d4fc","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4233,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:158"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinRealizedTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon : Nat) (hhorizon : 0 < horizon) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) (hrealized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * (horizon : Real)) : exponentialSquareBernsteinRealizedTunedThreshold arms gamma horizon delta <= exponentialSquareBernsteinRealizedE…","missing":[],"search":"exponentialsquarebernsteinrealizedtunedthreshold_le_explicitthreshold banditrlproof.exp3.exponentialsquarebernsteinrealizedtunedthreshold_le_explicitthreshold quadratic, sixth-power, cubic, and realized quadratic contracts reduce the learning-rate-tuned threshold to `14 * gamma * t`. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedExponentialSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_gammaCharacterizedExponentialSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedExponentialSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail under the four algebraic exploration contracts consumed by the explicit schedule below.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-18aad0a3505d","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4234,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:249"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedExponentialSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (arms.card : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) /…","missing":[],"search":"sampledpredictable_gammacharacterizedexponentialsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizedexponentialsquarebernsteinrealizedregret_tail generated realized-regret tail under the four algebraic exploration contracts consumed by the explicit schedule below. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareArmExplorationScale","label":"exponentialSquareArmExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareArmExplorationScale","description":"Square-root component required by the armwise `K * T` part of the exponential-square scale.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-a9adc8043004","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4235,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:329"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareArmExplorationScale (K T : Real) : Real","missing":[],"search":"exponentialsquarearmexplorationscale banditrlproof.exp3.exponentialsquarearmexplorationscale square-root component required by the armwise `k * t` part of the exponential-square scale. definition compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareMixedExplorationScale","label":"exponentialSquareMixedExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareMixedExplorationScale","description":"Sixth-root component forced by the current interval variance proxy `(K / (2 * gamma))^2`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-119b6bb44d36","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4236,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:335"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareMixedExplorationScale (K T delta : Real) : Real","missing":[],"search":"exponentialsquaremixedexplorationscale banditrlproof.exp3.exponentialsquaremixedexplorationscale sixth-root component forced by the current interval variance proxy `(k / (2 * gamma))^2`. definition compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate","label":"exponentialSquareBernsteinRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate","description":"Unclipped maximum of the arm, mixed-square, Bernstein-confidence, and realized-deviation exploration scales.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-cd49ccbd8698","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4237,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:342"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareBernsteinRawExplorationRate (K T delta : Real) : Real","missing":[],"search":"exponentialsquarebernsteinrawexplorationrate banditrlproof.exp3.exponentialsquarebernsteinrawexplorationrate unclipped maximum of the arm, mixed-square, bernstein-confidence, and realized-deviation exploration scales. definition compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate","label":"exponentialSquareBernsteinClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate","description":"Explicit exploration schedule clipped into the Hedge stability regime.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-72d04ebbc994","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4238,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:350"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareBernsteinClippedExplorationRate (K T delta : Real) : Real","missing":[],"search":"exponentialsquarebernsteinclippedexplorationrate banditrlproof.exp3.exponentialsquarebernsteinclippedexplorationrate explicit exploration schedule clipped into the hedge stability regime. definition compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.rpow_inv_six_le_half_of_sixtyfour_mul_le","label":"rpow_inv_six_le_half_of_sixtyfour_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.rpow_inv_six_le_half_of_sixtyfour_mul_le","description":"A nonnegative sixth-root scale is at most one half when its numerator is at most one sixty-fourth of its positive denominator.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-bcfde006bace","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4239,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:356"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem rpow_inv_six_le_half_of_sixtyfour_mul_le (numerator denominator : Real) (hnumerator : 0 <= numerator) (hdenominator : 0 < denominator) (hlarge : 64 * numerator <= denominator) : (numerator / denominator) ^ (6 : Real)⁻¹ <= 1 / 2","missing":[],"search":"rpow_inv_six_le_half_of_sixtyfour_mul_le banditrlproof.exp3.rpow_inv_six_le_half_of_sixtyfour_mul_le a nonnegative sixth-root scale is at most one half when its numerator is at most one sixty-fourth of its positive denominator. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.numerator_le_pow_six_mul_of_rpow_inv_six_le","label":"numerator_le_pow_six_mul_of_rpow_inv_six_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.numerator_le_pow_six_mul_of_rpow_inv_six_le","description":"If a sixth-root scale is below `gamma`, its numerator satisfies the corresponding sixth-power dominance contract.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-f6aa439ffa52","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4240,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:374"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem numerator_le_pow_six_mul_of_rpow_inv_six_le (numerator denominator gamma : Real) (hnumerator : 0 <= numerator) (hdenominator : 0 < denominator) (hroot : (numerator / denominator) ^ (6 : Real)⁻¹ <= gamma) : numerator <= gamma ^ 6 * denominator","missing":[],"search":"numerator_le_pow_six_mul_of_rpow_inv_six_le banditrlproof.exp3.numerator_le_pow_six_mul_of_rpow_inv_six_le if a sixth-root scale is below `gamma`, its numerator satisfies the corresponding sixth-power dominance contract. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_le_half","label":"exponentialSquareBernsteinClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_le_half","description":"theorem exponentialSquareBernsteinClippedExplorationRate_le_half (K T delta : Real) : exponentialSquareBernsteinClippedExplorationRate K T delta <= 1 / 2","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-3c06f837a7b3","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4241,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:395"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinClippedExplorationRate_le_half (K T delta : Real) : exponentialSquareBernsteinClippedExplorationRate K T delta <= 1 / 2","missing":[],"search":"exponentialsquarebernsteinclippedexplorationrate_le_half banditrlproof.exp3.exponentialsquarebernsteinclippedexplorationrate_le_half theorem exponentialsquarebernsteinclippedexplorationrate_le_half (k t delta : real) : exponentialsquarebernsteinclippedexplorationrate k t delta <= 1 / 2 theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_eq_raw","label":"exponentialSquareBernsteinClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_eq_raw","description":"theorem exponentialSquareBernsteinClippedExplorationRate_eq_raw (K T delta : Real) (hraw : exponentialSquareBernsteinRawExplorationRate K T delta <= 1 / 2) : exponentialSquareBernsteinClippedExplorationRate K T delta = exponentialSquareBernsteinRawExplorationRate K T delta","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-ae889d7301f4","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4242,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:400"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinClippedExplorationRate_eq_raw (K T delta : Real) (hraw : exponentialSquareBernsteinRawExplorationRate K T delta <= 1 / 2) : exponentialSquareBernsteinClippedExplorationRate K T delta = exponentialSquareBernsteinRawExplorationRate K T delta","missing":[],"search":"exponentialsquarebernsteinclippedexplorationrate_eq_raw banditrlproof.exp3.exponentialsquarebernsteinclippedexplorationrate_eq_raw theorem exponentialsquarebernsteinclippedexplorationrate_eq_raw (k t delta : real) (hraw : exponentialsquarebernsteinrawexplorationrate k t delta <= 1 / 2) : exponentialsquarebernsteinclippedexplorationrate k t delta = exponentialsquarebernsteinrawexplorationrate k t delta theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate_pos","label":"exponentialSquareBernsteinRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate_pos","description":"theorem exponentialSquareBernsteinRawExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < exponentialSquareBernsteinRawExplorationRate K T delta","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-f685808764ea","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4243,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:407"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinRawExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < exponentialSquareBernsteinRawExplorationRate K T delta","missing":[],"search":"exponentialsquarebernsteinrawexplorationrate_pos banditrlproof.exp3.exponentialsquarebernsteinrawexplorationrate_pos theorem exponentialsquarebernsteinrawexplorationrate_pos (k t delta : real) (hk_one : 1 < k) (ht : 0 < t) : 0 < exponentialsquarebernsteinrawexplorationrate k t delta theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_pos","label":"exponentialSquareBernsteinClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_pos","description":"theorem exponentialSquareBernsteinClippedExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < exponentialSquareBernsteinClippedExplorationRate K T delta","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-5a012a7f9018","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4244,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:416"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinClippedExplorationRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) : 0 < exponentialSquareBernsteinClippedExplorationRate K T delta","missing":[],"search":"exponentialsquarebernsteinclippedexplorationrate_pos banditrlproof.exp3.exponentialsquarebernsteinclippedexplorationrate_pos theorem exponentialsquarebernsteinclippedexplorationrate_pos (k t delta : real) (hk_one : 1 < k) (ht : 0 < t) : 0 < exponentialsquarebernsteinclippedexplorationrate k t delta theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","label":"exponentialSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","description":"Four transparent horizon contracts ensure every raw schedule component is at most one half, hence clipping is inactive.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-3cd88f65dca3","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4245,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:426"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (K * Real.log K) <= T) (hlarge_mixed : 64 * (K ^ 2 * Real.log K ^ 2 * Real.log (4 / delta) / 2) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : exponentialSquareBernsteinRawExplorationRate K T delta <= 1 / 2","missing":[],"search":"exponentialsquarebernsteinrawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.exponentialsquarebernsteinrawexplorationrate_le_half_of_horizon_contracts four transparent horizon contracts ensure every raw schedule component is at most one half, hence clipping is inactive. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_contracts","label":"exponentialSquareBernsteinClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_contracts","description":"The clipped maximum satisfies exactly the four contracts consumed by the gamma-characterized tail theorem.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-f46639cbe3e6","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4246,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:473"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareBernsteinClippedExplorationRate_contracts (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (K * Real.log K) <= T) (hlarge_mixed : 64 * (K ^ 2 * Real.log K ^ 2 * Real.log (4 / delta) / 2) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : let gamma := exponentialSquareBernsteinClippedExplorationRate K T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ K * Real.log K <= gamma ^ 2 * T ∧ K ^ 2 * Real.log K ^ 2 * Real.log (4 / delta) / 2 <= gamma ^ 6 * T ^ 3 ∧ K * Real.log (4 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * T","missing":[],"search":"exponentialsquarebernsteinclippedexplorationrate_contracts banditrlproof.exp3.exponentialsquarebernsteinclippedexplorationrate_contracts the clipped maximum satisfies exactly the four contracts consumed by the gamma-characterized tail theorem. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitExponentialSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_explicitExponentialSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitExponentialSquareBernsteinRealizedRegret_tail","description":"Fully explicit generated realized-regret tail for the clipped maximum of the four exploration scales.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedexplicittuning/index.html#decl-ed67290c7488","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","order":4247,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedExplicitTuning.lean:571"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitExponentialSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((arms.card : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 64 * ((arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) / 2) <= (horizon : Real) ^ 3) (hlarge_confidence : 8 * ((arms.card : Real)…","missing":[],"search":"sampledpredictable_explicitexponentialsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_explicitexponentialsquarebernsteinrealizedregret_tail fully explicit generated realized-regret tail for the clipped maximum of the four exploration scales. theorem compiled","shard":"modules/6d24f3b068b5394d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget","label":"sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget","description":"Realized selected-loss regret budget whose predictable component uses the exponential mixed estimator-square event and two Bernstein confidence radii.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedhighprobabilityregret/index.html#decl-1569ce222177","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","order":4248,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedHighProbabilityRegret.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictableexponentialsquarebernsteinrealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictableexponentialsquarebernsteinrealizedhighprobabilityregretbudget realized selected-loss regret budget whose predictable component uses the exponential mixed estimator-square event and two bernstein confidence radii. definition compiled","shard":"modules/6bf101da91579426.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail","label":"sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail","description":"Raw four-event form. The predictable component contributes the exponential estimator-square event and two Bernstein confidence events; the fourth event is the bounded realized-minus-predictable deviation.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedhighprobabilityregret/index.html#decl-006364fa940e","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","order":4249,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedHighProbabilityRegret.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (deltaSquare deltaConfidence deltaRealized : Real) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgam…","missing":[],"search":"sampledpredictable_exponentialsquarebernsteinrealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_exponentialsquarebernsteinrealizedhighprobabilityregret_tail raw four-event form. the predictable component contributes the exponential estimator-square event and two bernstein confidence events; the fourth event is the bounded realized-minus-predictable deviation. theorem compiled","shard":"modules/6bf101da91579426.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","description":"Total-failure form: the exponential square, pure-cross Bernstein, fixed-comparator Bernstein, and realized-deviation events each receive `delta / 4`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedhighprobabilityregret/index.html#decl-488af8f22351","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","order":4250,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedHighProbabilityRegret.lean:157"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegr…","missing":[],"search":"sampledpredictable_exponentialsquarebernsteinrealizedhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_exponentialsquarebernsteinrealizedhighprobabilityregret_tail_total_delta total-failure form: the exponential square, pure-cross bernstein, fixed-comparator bernstein, and realized-deviation events each receive `delta / 4`. theorem compiled","shard":"modules/6bf101da91579426.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityScale","label":"exponentialSquareHighProbabilityScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareHighProbabilityScale","description":"Positive scale appearing in the exponential-square Hedge term when the square event receives `delta / 4`.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-8d05dc4c7e99","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4251,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareHighProbabilityScale {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"exponentialsquarehighprobabilityscale banditrlproof.exp3.exponentialsquarehighprobabilityscale positive scale appearing in the exponential-square hedge term when the square event receives `delta / 4`. definition compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityScale_pos","label":"exponentialSquareHighProbabilityScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareHighProbabilityScale_pos","description":"theorem exponentialSquareHighProbabilityScale_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < exponentialSquareHighProbabilityScale arms gamma horizon delta","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-bd67a8ac34d8","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4252,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareHighProbabilityScale_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < exponentialSquareHighProbabilityScale arms gamma horizon delta","missing":[],"search":"exponentialsquarehighprobabilityscale_pos banditrlproof.exp3.exponentialsquarehighprobabilityscale_pos theorem exponentialsquarehighprobabilityscale_pos {action : type v} (arms : finset action) (hcard_two : 2 <= arms.card) (gamma : real) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) : 0 < exponentialsquarehighprobabilityscale arms gamma horizon delta theorem compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate","label":"exponentialSquareHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate","description":"Learning rate balancing entropy against the complete exponential-square stability scale at the public `delta / 4` square allocation.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-f03eb531f51f","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4253,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:49"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareHighProbabilityLearningRate {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"exponentialsquarehighprobabilitylearningrate banditrlproof.exp3.exponentialsquarehighprobabilitylearningrate learning rate balancing entropy against the complete exponential-square stability scale at the public `delta / 4` square allocation. definition compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate_pos","label":"exponentialSquareHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate_pos","description":"theorem exponentialSquareHighProbabilityLearningRate_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < exponentialSquareHighProbabilityLearningRate arms gamma horizon delta","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-0dd3fd92b908","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4254,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:56"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareHighProbabilityLearningRate_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : 0 < exponentialSquareHighProbabilityLearningRate arms gamma horizon delta","missing":[],"search":"exponentialsquarehighprobabilitylearningrate_pos banditrlproof.exp3.exponentialsquarehighprobabilitylearningrate_pos theorem exponentialsquarehighprobabilitylearningrate_pos {action : type v} (arms : finset action) (hcard_two : 2 <= arms.card) (gamma : real) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) : 0 < exponentialsquarehighprobabilitylearningrate arms gamma horizon delta theorem compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate_sq_mul_scale","label":"exponentialSquareHighProbabilityLearningRate_sq_mul_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate_sq_mul_scale","description":"theorem exponentialSquareHighProbabilityLearningRate_sq_mul_scale {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : exponentialSquareHighProbabilityLearningRate arms gamma horizon delta ^ 2 * exponentialSquareHighProbabilityScale arms gamma horizon delta = Real.log (arms.card : Real)","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-e03e315c7003","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4255,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:68"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareHighProbabilityLearningRate_sq_mul_scale {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : exponentialSquareHighProbabilityLearningRate arms gamma horizon delta ^ 2 * exponentialSquareHighProbabilityScale arms gamma horizon delta = Real.log (arms.card : Real)","missing":[],"search":"exponentialsquarehighprobabilitylearningrate_sq_mul_scale banditrlproof.exp3.exponentialsquarehighprobabilitylearningrate_sq_mul_scale theorem exponentialsquarehighprobabilitylearningrate_sq_mul_scale {action : type v} (arms : finset action) (hcard_two : 2 <= arms.card) (gamma : real) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) : exponentialsquarehighprobabilitylearningrate arms gamma horizon delta ^ 2 * exponentialsquarehighprobabilityscale arms gamma horizon delta = real.log (arms.card : real) theorem compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","label":"exponentialSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","description":"With `gamma <= 1/2`, entropy and the stability-amplified exponential-square scale cost at most three copies of their balanced square-root scale.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-22f765406d8c","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4256,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:85"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exponentialSquareHighProbabilityHedgeBudget_le_three_mul_sqrt {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : Real.log (arms.card : Real) / exponentialSquareHighProbabilityLearningRate arms gamma horizon delta + (exponentialSquareHighProbabilityLearningRate arms gamma horizon delta * (1 / (1 - gamma))) * exponentialSquareHighProbabilityScale arms gamma horizon delta <= 3 * Real.sqrt (Real.log (arms.card : Real) * exponentialSquareHighProbabilityScale arms gamma horizon delta)","missing":[],"search":"exponentialsquarehighprobabilityhedgebudget_le_three_mul_sqrt banditrlproof.exp3.exponentialsquarehighprobabilityhedgebudget_le_three_mul_sqrt with `gamma <= 1/2`, entropy and the stability-amplified exponential-square scale cost at most three copies of their balanced square-root scale. theorem compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRealizedTunedThreshold","label":"exponentialSquareBernsteinRealizedTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exponentialSquareBernsteinRealizedTunedThreshold","description":"Explicit threshold after tuning all learning-rate-dependent terms. The exploration and three confidence contributions remain visible.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-ee72251e5d9a","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4257,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:156"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exponentialSquareBernsteinRealizedTunedThreshold {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"exponentialsquarebernsteinrealizedtunedthreshold banditrlproof.exp3.exponentialsquarebernsteinrealizedtunedthreshold explicit threshold after tuning all learning-rate-dependent terms. the exploration and three confidence contributions remain visible. definition compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","description":"The complete exponential-square four-event realized budget is bounded by the learning-rate-tuned threshold.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-f2e0a8ecfeca","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4258,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:171"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) : sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget arms (exponentialSquareHighProbabilityLearningRate arms gamma horizon delta) gamma horizon (delta / 4) (delta / 4) (delta / 4) <= exponentialSquareBernsteinRealizedTunedThreshold arms gamma horizon delta","missing":[],"search":"sampledpredictableexponentialsquarebernsteinrealizedhighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictableexponentialsquarebernsteinrealizedhighprobabilityregretbudget_le_tunedthreshold the complete exponential-square four-event realized budget is bounded by the learning-rate-tuned threshold. theorem compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedExponentialSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_tunedExponentialSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedExponentialSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail with the exponential-square-balanced learning rate. This removes the old Markov tuning and leaves only the gamma schedule and confidence-radius simplification for downstream consumers.","url":"../modules/banditrlproof-exp3mixedsquareexponentialrealizedtuning/index.html#decl-40af2e31c005","parent":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","order":4259,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquareExponentialRealizedTuning"],["Source","BanditRLProof/Exp3MixedSquareExponentialRealizedTuning.lean:198"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedExponentialSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let eta := exponentialSquareHighProbabilityLearningRate arms gamma horizon delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le (by linarith : gamma <= 1) loss.environment m…","missing":[],"search":"sampledpredictable_tunedexponentialsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_tunedexponentialsquarebernsteinrealizedregret_tail generated realized-regret tail with the exponential-square-balanced learning rate. this removes the old markov tuning and leaves only the gamma schedule and confidence-radius simplification for downstream consumers. theorem compiled","shard":"modules/23c5a73a7619a967.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment","label":"mixedSquaredEstimatorCenteredSecondMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment","description":"Exact centered second moment of the mixed importance-weighted square under a finite sampling distribution.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-5844873e123f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4260,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def mixedSquaredEstimatorCenteredSecondMoment {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) : Real","missing":[],"search":"mixedsquaredestimatorcenteredsecondmoment banditrlproof.exp3.mixedsquaredestimatorcenteredsecondmoment exact centered second moment of the mixed importance-weighted square under a finite sampling distribution. definition compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_mixedSquaredEstimatorCenteredSecondMoment","label":"measurable_mixedSquaredEstimatorCenteredSecondMoment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_mixedSquaredEstimatorCenteredSecondMoment","description":"theorem measurable_mixedSquaredEstimatorCenteredSecondMoment {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon)…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-9248c2f2086b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4261,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_mixedSquaredEstimatorCenteredSecondMoment {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Measurable (mixedSquaredEstimatorCenteredSecondMoment arms prob loss)","missing":[],"search":"measurable_mixedsquaredestimatorcenteredsecondmoment banditrlproof.exp3.measurable_mixedsquaredestimatorcenteredsecondmoment theorem measurable_mixedsquaredestimatorcenteredsecondmoment {history : type u} {action : type v} [measurablespace history] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (prob loss : history -> action -> real) (source : measurablefiniteactiondistribution arms prob) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) : measurable (mixedsquaredestimatorcenteredsecondmoment arms prob loss) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_nonneg","label":"mixedSquaredEstimatorCenteredSecondMoment_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_nonneg","description":"theorem mixedSquaredEstimatorCenteredSecondMoment_nonneg {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) : 0 <= mixedSquaredEstimatorCenteredSecondMoment arms prob loss history","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-efa2744ec529","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4262,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:58"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimatorCenteredSecondMoment_nonneg {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) : 0 <= mixedSquaredEstimatorCenteredSecondMoment arms prob loss history","missing":[],"search":"mixedsquaredestimatorcenteredsecondmoment_nonneg banditrlproof.exp3.mixedsquaredestimatorcenteredsecondmoment_nonneg theorem mixedsquaredestimatorcenteredsecondmoment_nonneg {history : type u} {action : type v} [decidableeq action] (arms : finset action) (prob loss : history -> action -> real) (history : history) (hdist : finiteactiondistribution arms (prob history)) : 0 <= mixedsquaredestimatorcenteredsecondmoment arms prob loss history theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_le_card_div_floor","label":"mixedSquaredEstimatorCenteredSecondMoment_le_card_div_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_le_card_div_floor","description":"theorem mixedSquaredEstimatorCenteredSecondMoment_le_card_div_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : mixedSquaredEstimatorCenteredS…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-4b785917c6a8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4263,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:68"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimatorCenteredSecondMoment_le_card_div_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : mixedSquaredEstimatorCenteredSecondMoment arms prob loss history <= (arms.card : Real) / epsilon","missing":[],"search":"mixedsquaredestimatorcenteredsecondmoment_le_card_div_floor banditrlproof.exp3.mixedsquaredestimatorcenteredsecondmoment_le_card_div_floor theorem mixedsquaredestimatorcenteredsecondmoment_le_card_div_floor {history : type u} {action : type v} [measurablespace history] [decidableeq action] (arms : finset action) (prob loss : history -> action -> real) (history : history) (hdist : finiteactiondistribution arms (prob history)) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) : mixedsquaredestimatorcenteredsecondmoment arms prob loss history <= (arms.card : real) / epsilon theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_sq_mixedSquaredEstimatorDeviation_finiteActionMeasure_eq","label":"integral_sq_mixedSquaredEstimatorDeviation_finiteActionMeasure_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_sq_mixedSquaredEstimatorDeviation_finiteActionMeasure_eq","description":"theorem integral_sq_mixedSquaredEstimatorDeviation_finiteActionMeasure_eq {History : Type u} {Action : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) : integral (finiteActionMeasure arms (prob history)) (fun chosen => (mixedSquaredImportanc…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-642eced9bfd0","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4264,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:84"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_mixedSquaredEstimatorDeviation_finiteActionMeasure_eq {History : Type u} {Action : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) : integral (finiteActionMeasure arms (prob history)) (fun chosen => (mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) chosen - arms.sum (fun action => (loss history action) ^ 2)) ^ 2) = mixedSquaredEstimatorCenteredSecondMoment arms prob loss history","missing":[],"search":"integral_sq_mixedsquaredestimatordeviation_finiteactionmeasure_eq banditrlproof.exp3.integral_sq_mixedsquaredestimatordeviation_finiteactionmeasure_eq theorem integral_sq_mixedsquaredestimatordeviation_finiteactionmeasure_eq {history : type u} {action : type v} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (prob loss : history -> action -> real) (history : history) (hdist : finiteactiondistribution arms (prob history)) : integral (finiteactionmeasure arms (prob history)) (fun chosen => (mixedsquaredimportanceweightedloss arms (prob history) (loss history) chosen - arms.sum (fun action => (loss history action) ^ 2)) ^ 2) = mixedsquaredestimatorcenteredsecondmoment arms prob loss history theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorDeviation_condExpKernel_map_eq_finiteActionMeasure_of_condDistrib","label":"mixedSquaredEstimatorDeviation_condExpKernel_map_eq_finiteActionMeasure_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimatorDeviation_condExpKernel_map_eq_finiteActionMeasure_of_condDistrib","description":"Transport the centered mixed-square score law from an identified finite conditional action distribution into the ambient conditional-expectation kernel.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-5496cd57efab","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4265,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:105"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimatorDeviation_condExpKernel_map_eq_finiteActionMeasure_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (hcond : condDistrib action history mu =ᵐ[mu.map history] finiteActionKernel arms prob source) :…","missing":[],"search":"mixedsquaredestimatordeviation_condexpkernel_map_eq_finiteactionmeasure_of_conddistrib banditrlproof.exp3.mixedsquaredestimatordeviation_condexpkernel_map_eq_finiteactionmeasure_of_conddistrib transport the centered mixed-square score law from an identified finite conditional action distribution into the ambient conditional-expectation kernel. theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integral_sq_mixedSquaredEstimatorDeviation_condExpKernel_eq_of_condDistrib","label":"integral_sq_mixedSquaredEstimatorDeviation_condExpKernel_eq_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integral_sq_mixedSquaredEstimatorDeviation_condExpKernel_eq_of_condDistrib","description":"The ambient conditional-expectation kernel integrates the squared centered mixed-square increment to the explicit finite-law centered second moment.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-0ea6bff58a6b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4266,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:219"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_mixedSquaredEstimatorDeviation_condExpKernel_eq_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (hcond : condDistrib action history mu =ᵐ[mu.map history] finiteActionKernel arms prob source) : Filter.Even…","missing":[],"search":"integral_sq_mixedsquaredestimatordeviation_condexpkernel_eq_of_conddistrib banditrlproof.exp3.integral_sq_mixedsquaredestimatordeviation_condexpkernel_eq_of_conddistrib the ambient conditional-expectation kernel integrates the squared centered mixed-square increment to the explicit finite-law centered second moment. theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt","label":"sampledTrajectoryPredictableMixedSquaredVarianceAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt","description":"Exact finite-law predictable variance at an actual generated time.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-35af561ff29b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4267,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:300"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPredictableMixedSquaredVarianceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) : Env × ((k : Nat) -> Action × Real) -> Real","missing":[],"search":"sampledtrajectorypredictablemixedsquaredvarianceat banditrlproof.exp3.sampledtrajectorypredictablemixedsquaredvarianceat exact finite-law predictable variance at an actual generated time. definition compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt","label":"measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt","description":"theorem measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryPredictable…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-0aa5ec49a8a4","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4268,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:310"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectorypredictablemixedsquaredvarianceat banditrlproof.exp3.measurable_sampledtrajectorypredictablemixedsquaredvarianceat theorem measurable_sampledtrajectorypredictablemixedsquaredvarianceat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : measurable (sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss t) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_nonneg","label":"sampledTrajectoryPredictableMixedSquaredVarianceAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_nonneg","description":"theorem sampledTrajectoryPredictableMixedSquaredVarianceAt_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Rea…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-98d3bc61420e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4269,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:330"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableMixedSquaredVarianceAt_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : 0 <= sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss t sample","missing":[],"search":"sampledtrajectorypredictablemixedsquaredvarianceat_nonneg banditrlproof.exp3.sampledtrajectorypredictablemixedsquaredvarianceat_nonneg theorem sampledtrajectorypredictablemixedsquaredvarianceat_nonneg {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : 0 <= sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss t sample theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_le","label":"sampledTrajectoryPredictableMixedSquaredVarianceAt_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_le","description":"theorem sampledTrajectoryPredictableMixedSquaredVarianceAt_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sa…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-a03a2f74eb75","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4270,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:348"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableMixedSquaredVarianceAt_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss t sample <= (arms.card : Real) / (gamma / (arms.card : Real))","missing":[],"search":"sampledtrajectorypredictablemixedsquaredvarianceat_le banditrlproof.exp3.sampledtrajectorypredictablemixedsquaredvarianceat_le theorem sampledtrajectorypredictablemixedsquaredvarianceat_le {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss t sample <= (arms.card : real) / (gamma / (arms.card : real)) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_condExpKernel_integral_sq_eq_variance","label":"sampledPredictableMixedSquaredDeviation_zero_condExpKernel_integral_sq_eq_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_condExpKernel_integral_sq_eq_variance","description":"At generated time zero, the ambient conditional-expectation kernel given the environment integrates the squared centered mixed-square increment to the explicit predictable variance.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-3288b777d4d8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4271,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:373"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_zero_condExpKernel_integral_sq_eq_variance {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Filter.Eventually (fun omega => integral (@condExpKernel (Env × ((k : Nat) -> Action × Real)) inferInstance _ mu _ ((inferInstance : MeasurableSpace Env).comap (fun sample => sample.1)) omega) (fun y => (sampledTrajectoryPredictableMixedSquaredDeviat…","missing":[],"search":"sampledpredictablemixedsquareddeviation_zero_condexpkernel_integral_sq_eq_variance banditrlproof.exp3.sampledpredictablemixedsquareddeviation_zero_condexpkernel_integral_sq_eq_variance at generated time zero, the ambient conditional-expectation kernel given the environment integrates the squared centered mixed-square increment to the explicit predictable variance. theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_condExpKernel_integral_sq_eq_variance","label":"sampledPredictableMixedSquaredDeviation_succ_condExpKernel_integral_sq_eq_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_condExpKernel_integral_sq_eq_variance","description":"At a generated successor time, the ambient conditional-expectation kernel given the environment and preceding finite prefix integrates the squared centered mixed-square increment to the explicit predictable variance.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-287e14d97616","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4272,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:436"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_succ_condExpKernel_integral_sq_eq_variance {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Filter.Eventually (fun omega => integral (@condExpKernel (Env × ((k : Nat) -> Action × Real)) inferInstance _ mu _ ((inferInstance…","missing":[],"search":"sampledpredictablemixedsquareddeviation_succ_condexpkernel_integral_sq_eq_variance banditrlproof.exp3.sampledpredictablemixedsquareddeviation_succ_condexpkernel_integral_sq_eq_variance at a generated successor time, the ambient conditional-expectation kernel given the environment and preceding finite prefix integrates the squared centered mixed-square increment to the explicit predictable variance. theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess","label":"sampledPredictableMixedSquaredVarianceProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess","description":"Shift the actual-time conditional variances by one. The resulting process is predictable for the generated deviation filtration.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-45eec30ba07a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4273,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:500"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableMixedSquaredVarianceProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real","missing":[],"search":"sampledpredictablemixedsquaredvarianceprocess banditrlproof.exp3.sampledpredictablemixedsquaredvarianceprocess shift the actual-time conditional variances by one. the resulting process is predictable for the generated deviation filtration. definition compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_condExpKernel_integral_sq_eq_varianceProcess","label":"sampledPredictableMixedSquaredDeviationProcess_condExpKernel_integral_sq_eq_varianceProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_condExpKernel_integral_sq_eq_varianceProcess","description":"Every shifted mixed-square increment has conditional square integral equal to the matching shifted predictable variance under the existing generated filtration.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-1083583c9eba","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4274,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:515"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviationProcess_condExpKernel_integral_sq_eq_varianceProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Filter.Eventually (fun omega => integral (@condExpKernel (Env × ((k : Nat) -> Action × Real)) inferInstance _ mu _ (sampledPredictableDeviationFiltration Env Action n) omega) (fun y => (sampledPredictableMixedSquaredDeviationProces…","missing":[],"search":"sampledpredictablemixedsquareddeviationprocess_condexpkernel_integral_sq_eq_varianceprocess banditrlproof.exp3.sampledpredictablemixedsquareddeviationprocess_condexpkernel_integral_sq_eq_varianceprocess every shifted mixed-square increment has conditional square integral equal to the matching shifted predictable variance under the existing generated filtration. theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt_filtration","label":"measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt_filtration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt_filtration","description":"theorem measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt_filtration {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable[sampledPredictable…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-2e9894fb1f69","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4275,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:559"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt_filtration {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable[sampledPredictableDeviationFiltration Env Action t] (sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectorypredictablemixedsquaredvarianceat_filtration banditrlproof.exp3.measurable_sampledtrajectorypredictablemixedsquaredvarianceat_filtration theorem measurable_sampledtrajectorypredictablemixedsquaredvarianceat_filtration {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : measurable[sampledpredictabledeviationfiltration env action t] (sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss t) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess_isPredictable","label":"sampledPredictableMixedSquaredVarianceProcess_isPredictable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess_isPredictable","description":"theorem sampledPredictableMixedSquaredVarianceProcess_isPredictable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : IsPredictable (sampledPredictableDeviationFiltration…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-492c038f4745","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4276,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:620"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceProcess_isPredictable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : IsPredictable (sampledPredictableDeviationFiltration Env Action) (sampledPredictableMixedSquaredVarianceProcess arms eta gamma loss)","missing":[],"search":"sampledpredictablemixedsquaredvarianceprocess_ispredictable banditrlproof.exp3.sampledpredictablemixedsquaredvarianceprocess_ispredictable theorem sampledpredictablemixedsquaredvarianceprocess_ispredictable {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : ispredictable (sampledpredictabledeviationfiltration env action) (sampledpredictablemixedsquaredvarianceprocess arms eta gamma loss) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess_sum_range_succ","label":"sampledPredictableMixedSquaredVarianceProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess_sum_range_succ","description":"theorem sampledPredictableMixedSquaredVarianceProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableMixedSquaredVarianceProcess arms eta gam…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-c64f64088f4c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4277,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:640"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableMixedSquaredVarianceProcess arms eta gamma loss i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss i sample)","missing":[],"search":"sampledpredictablemixedsquaredvarianceprocess_sum_range_succ banditrlproof.exp3.sampledpredictablemixedsquaredvarianceprocess_sum_range_succ theorem sampledpredictablemixedsquaredvarianceprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpredictablemixedsquaredvarianceprocess arms eta gamma loss i sample) = (finset.range horizon).sum (fun i => sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss i sample) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVariance_sum_le","label":"sampledPredictableMixedSquaredVariance_sum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVariance_sum_le","description":"theorem sampledPredictableMixedSquaredVariance_sum_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Fin…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariance/index.html#decl-e8df89345c42","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","order":4278,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVariance"],["Source","BanditRLProof/Exp3MixedSquarePredictableVariance.lean:660"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVariance_sum_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss i sample) <= (horizon : Real) * ((arms.card : Real) / (gamma / (arms.card : Real)))","missing":[],"search":"sampledpredictablemixedsquaredvariance_sum_le banditrlproof.exp3.sampledpredictablemixedsquaredvariance_sum_le theorem sampledpredictablemixedsquaredvariance_sum_le {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range horizon).sum (fun i => sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss i sample) <= (horizon : real) * ((arms.card : real) / (gamma / (arms.card : real))) theorem compiled","shard":"modules/cd8ef79537333a57.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_delta","label":"sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_delta","description":"The observed mixed estimator-square sum has a random predictable-variance tail on the event that the cumulative variance is at most `varianceBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html#decl-f0c24ca471f7","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","order":4279,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) + sampledMixedSquaredPredictableVarianceRadius arms gamma varianceBudget delta <= sampledObserved…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_delta banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_delta the observed mixed estimator-square sum has a random predictable-variance tail on the event that the cumulative variance is at most `variancebudget`. theorem compiled","shard":"modules/8179b50f04b867f0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareHighProbabilityRegretBudget","description":"Predictable regret budget with a caller-supplied cumulative predictable variance budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html#decl-b2155b11795d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","order":4280,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean:80"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (varianceBudget deltaSquare deltaConfidence : Real) : Real","missing":[],"search":"sampledpredictablevariancesquarehighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquarehighprobabilityregretbudget predictable regret budget with a caller-supplied cumulative predictable variance budget. definition compiled","shard":"modules/8179b50f04b867f0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint","label":"sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint","description":"Generated predictable EXP3 regret on the event that the cumulative predictable mixed-square variance stays below `varianceBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html#decl-899e2f1aa634","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","order":4281,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean:97"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (varianceBudget deltaSquare deltaConfidence : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environ…","missing":[],"search":"sampledpredictable_predictablevariancesquarehighprobabilityregret_tail_joint banditrlproof.exp3.sampledpredictable_predictablevariancesquarehighprobabilityregret_tail_joint generated predictable exp3 regret on the event that the cumulative predictable mixed-square variance stays below `variancebudget`. theorem compiled","shard":"modules/8179b50f04b867f0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail","label":"sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail","description":"Unconditional predictable-regret bound with the cumulative predictable variance overflow probability left explicit.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html#decl-7b9f0147bf17","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","order":4282,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean:281"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (varianceBudget deltaSquare deltaConfidence : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment m…","missing":[],"search":"sampledpredictable_predictablevariancesquarehighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_predictablevariancesquarehighprobabilityregret_tail unconditional predictable-regret bound with the cumulative predictable variance overflow probability left explicit. theorem compiled","shard":"modules/8179b50f04b867f0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint_total_delta","label":"sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint_total_delta","description":"Total-failure joint-event form with the square, pure-cross, and comparator events allocated `delta / 3`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html#decl-5cd6999c05ba","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","order":4283,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean:365"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableVarianceSquareHighProbabili…","missing":[],"search":"sampledpredictable_predictablevariancesquarehighprobabilityregret_tail_joint_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquarehighprobabilityregret_tail_joint_total_delta total-failure joint-event form with the square, pure-cross, and comparator events allocated `delta / 3`. theorem compiled","shard":"modules/8179b50f04b867f0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_total_delta","description":"Primary residual-variance form: total confidence failure is `delta`, and the only remaining term is the probability that cumulative predictable variance exceeds `varianceBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancehighprobabilityregret/index.html#decl-c86095b5a853","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","order":4284,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceHighProbabilityRegret.lean:413"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableVarianceSquareHighProbabilityRegr…","missing":[],"search":"sampledpredictable_predictablevariancesquarehighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquarehighprobabilityregret_tail_total_delta primary residual-variance form: total confidence failure is `delta`, and the only remaining term is the probability that cumulative predictable variance exceeds `variancebudget`. theorem compiled","shard":"modules/8179b50f04b867f0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_inv_floor_mul_sum_loss_sq","label":"sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_inv_floor_mul_sum_loss_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_inv_floor_mul_sum_loss_sq","description":"A finite-action centered mixed-square estimator has variance at most the inverse probability floor times the armwise loss-square energy.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-b200d6d71b6a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4285,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_inv_floor_mul_sum_loss_sq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action → Real) (hdist : FiniteActionDistribution arms prob) (epsilon : Real) (hepsilon : 0 < epsilon) (hfloor : ∀ action, action ∈ arms → epsilon ≤ prob action) (hloss : ∀ action, action ∈ arms → loss action ∈ Set.Icc (0 : Real) 1) : arms.sum (fun chosen => prob chosen * (mixedSquaredImportanceWeightedLoss arms prob loss chosen - arms.sum (fun action => (loss action) ^ 2)) ^ 2) ≤ (1 / epsilon) * arms.sum (fun action => (loss action) ^ 2)","missing":[],"search":"sum_prob_mul_sq_mixedsquaredestimatordeviation_le_inv_floor_mul_sum_loss_sq banditrlproof.exp3.sum_prob_mul_sq_mixedsquaredestimatordeviation_le_inv_floor_mul_sum_loss_sq a finite-action centered mixed-square estimator has variance at most the inverse probability floor times the armwise loss-square energy. theorem compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_le_inv_floor_mul_sum_loss_sq","label":"mixedSquaredEstimatorCenteredSecondMoment_le_inv_floor_mul_sum_loss_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_le_inv_floor_mul_sum_loss_sq","description":"theorem mixedSquaredEstimatorCenteredSecondMoment_le_inv_floor_mul_sum_loss_sq {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : mixedSquaredEstimator…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-cc638b1a822d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4286,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:92"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimatorCenteredSecondMoment_le_inv_floor_mul_sum_loss_sq {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : mixedSquaredEstimatorCenteredSecondMoment arms prob loss history ≤ (1 / epsilon) * arms.sum (fun action => (loss history action) ^ 2)","missing":[],"search":"mixedsquaredestimatorcenteredsecondmoment_le_inv_floor_mul_sum_loss_sq banditrlproof.exp3.mixedsquaredestimatorcenteredsecondmoment_le_inv_floor_mul_sum_loss_sq theorem mixedsquaredestimatorcenteredsecondmoment_le_inv_floor_mul_sum_loss_sq {history : type u} {action : type v} [measurablespace history] [decidableeq action] (arms : finset action) (prob loss : history → action → real) (history : history) (hdist : finiteactiondistribution arms (prob history)) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) : mixedsquaredestimatorcenteredsecondmoment arms prob loss history ≤ (1 / epsilon) * arms.sum (fun action => (loss history action) ^ 2) theorem compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_le_inv_floor_mul_lossSquaredAt","label":"sampledTrajectoryPredictableMixedSquaredVarianceAt_le_inv_floor_mul_lossSquaredAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_le_inv_floor_mul_lossSquaredAt","description":"Pointwise generated-time specialization of the finite loss-energy bound.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-b28451c1282c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4287,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:108"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableMixedSquaredVarianceAt_le_inv_floor_mul_lossSquaredAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss t sample ≤ (1 / (gamma / (arms.card : Real))) * arms.sum (fun action => (predictableLossAt loss t sample action) ^ 2)","missing":[],"search":"sampledtrajectorypredictablemixedsquaredvarianceat_le_inv_floor_mul_losssquaredat banditrlproof.exp3.sampledtrajectorypredictablemixedsquaredvarianceat_le_inv_floor_mul_losssquaredat pointwise generated-time specialization of the finite loss-energy bound. theorem compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossSquaredSum","label":"sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossSquaredSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossSquaredSum","description":"Cumulative predictable variance is controlled by cumulative armwise predictable loss-square energy.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-c79810eb0e3a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4288,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:134"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossSquaredSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableMixedSquaredVarianceSum arms eta gamma loss horizon sample ≤ (1 / (gamma / (arms.card : Real))) * sampledPredictableLossSquaredSum arms loss horizon sample","missing":[],"search":"sampledpredictablemixedsquaredvariancesum_le_inv_floor_mul_losssquaredsum banditrlproof.exp3.sampledpredictablemixedsquaredvariancesum_le_inv_floor_mul_losssquaredsum cumulative predictable variance is controlled by cumulative armwise predictable loss-square energy. theorem compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossSquaredSum_le","label":"sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossSquaredSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossSquaredSum_le","description":"A pathwise predictable loss-square budget yields the generated cumulative-variance `lintegral` budget required by the Markov route.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-e46248f605bf","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4289,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:169"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossSquaredSum_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (lossSquaredBudget : Real) (henergy : ∀ sample, sampledPredictableLossSquaredSum arms loss horizon sample ≤ lossSquaredBudget) : sampledPredictableMixedSquaredVarianceLIntegral mu arms eta gamma loss horizon ≤ ENNReal.ofReal ((1 / (gamma / (arms.card : Real))) * lossSquaredBudget)","missing":[],"search":"sampledpredictablemixedsquaredvariancelintegral_le_of_losssquaredsum_le banditrlproof.exp3.sampledpredictablemixedsquaredvariancelintegral_le_of_losssquaredsum_le a pathwise predictable loss-square budget yields the generated cumulative-variance `lintegral` budget required by the markov route. theorem compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegretBudget","description":"Realized Markov budget specialized to a cumulative predictable loss-square energy budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-691fd61f352f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4290,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:215"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (lossSquaredBudget delta : Real) : Real","missing":[],"search":"sampledpredictablevariancesquarelossenergyrealizedmarkovhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquarelossenergyrealizedmarkovhighprobabilityregretbudget realized markov budget specialized to a cumulative predictable loss-square energy budget. definition compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_predictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegret_tail_total_delta","description":"Primary small-loss-energy specialization of the Markov-closed realized EXP3 theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancelossenergyrealizedmarkovhighprobabilityregret/index.html#decl-7483b7d707b3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","order":4291,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret.lean:225"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (lossSquaredBudget delta : Real) (hlossSquaredBudget : 0 < lossSquaredBudget) (hdelta : 0 < delta) (henergy : ∀ sample, sampledPredictableLossSquaredSum arms loss horizon sample ≤ lossSquaredBudget) : let mu := prior ⊗ₘ sampledImportance…","missing":[],"search":"sampledpredictable_predictablevariancesquarelossenergyrealizedmarkovhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquarelossenergyrealizedmarkovhighprobabilityregret_tail_total_delta primary small-loss-energy specialization of the markov-closed realized exp3 theorem. theorem compiled","shard":"modules/92b6fc3e407e4601.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableDoubleVarianceRealizedHighProbabilityRegretBudget","label":"sampledPredictableDoubleVarianceRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableDoubleVarianceRealizedHighProbabilityRegretBudget","description":"noncomputable def sampledPredictableDoubleVarianceRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (mixedVarianceBudget realizedVarianceBudget deltaSquare deltaConfidence deltaRealized : Real) : Real","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizeddoublepredictablevariancehi-73ffccd3bfbc/index.html#decl-d20704eb6914","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","order":4292,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableDoubleVarianceRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (mixedVarianceBudget realizedVarianceBudget deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictabledoublevariancerealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictabledoublevariancerealizedhighprobabilityregretbudget noncomputable def sampledpredictabledoublevariancerealizedhighprobabilityregretbudget {action : type v} [decidableeq action] (arms : finset action) (eta gamma : real) (horizon : nat) (mixedvariancebudget realizedvariancebudget deltasquare deltaconfidence deltarealized : real) : real definition compiled","shard":"modules/6275faacb1d275a2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_doublePredictableVarianceRealizedHighProbabilityRegret_tail_joint","label":"sampledPredictable_doublePredictableVarianceRealizedHighProbabilityRegret_tail_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_doublePredictableVarianceRealizedHighProbabilityRegret_tail_joint","description":"Joint realized-regret tail on simultaneous pathwise budgets for the mixed-square and selected-loss predictable variances.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizeddoublepredictablevariancehi-73ffccd3bfbc/index.html#decl-659bea338627","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","order":4293,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_doublePredictableVarianceRealizedHighProbabilityRegret_tail_joint {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (mixedVarianceBudget realizedVarianceBudget deltaSquare deltaConfidence deltaRealized : Real) (hmixedVarianceBudget : 0 < mixedVarianceBudget) (hrealizedVarianceBudget : 0 < realizedVarianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaReali…","missing":[],"search":"sampledpredictable_doublepredictablevariancerealizedhighprobabilityregret_tail_joint banditrlproof.exp3.sampledpredictable_doublepredictablevariancerealizedhighprobabilityregret_tail_joint joint realized-regret tail on simultaneous pathwise budgets for the mixed-square and selected-loss predictable variances. theorem compiled","shard":"modules/6275faacb1d275a2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareRealizedHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareRealizedHighProbabilityRegretBudget","description":"Realized selected-loss regret budget with a caller-supplied cumulative predictable mixed-square variance budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedhighprobabilityregret/index.html#decl-aeec76b47e12","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","order":4294,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (varianceBudget deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictablevariancesquarerealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquarerealizedhighprobabilityregretbudget realized selected-loss regret budget with a caller-supplied cumulative predictable mixed-square variance budget. definition compiled","shard":"modules/022f2888bf1bc5ee.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint","label":"sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint","description":"Realized selected-loss regret on the event that cumulative predictable mixed-square variance stays below `varianceBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedhighprobabilityregret/index.html#decl-990e0d28e148","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","order":4295,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (varianceBudget deltaSquare deltaConfidence deltaRealized : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) : let mu := prior ⊗ₘ sampledImportanceWeigh…","missing":[],"search":"sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_joint banditrlproof.exp3.sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_joint realized selected-loss regret on the event that cumulative predictable mixed-square variance stays below `variancebudget`. theorem compiled","shard":"modules/022f2888bf1bc5ee.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail","label":"sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail","description":"Unconditional realized selected-loss regret with the cumulative predictable-variance overflow probability left explicit.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedhighprobabilityregret/index.html#decl-7ffc90f823d1","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","order":4296,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret.lean:169"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (varianceBudget deltaSquare deltaConfidence deltaRealized : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) : let mu := prior ⊗ₘ sampledImportanceWeightedTra…","missing":[],"search":"sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail unconditional realized selected-loss regret with the cumulative predictable-variance overflow probability left explicit. theorem compiled","shard":"modules/022f2888bf1bc5ee.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint_total_delta","label":"sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint_total_delta","description":"Total-failure joint-event form with all four confidence events allocated `delta / 4`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedhighprobabilityregret/index.html#decl-d8d44162dd8c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","order":4297,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret.lean:255"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredi…","missing":[],"search":"sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_joint_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_joint_total_delta total-failure joint-event form with all four confidence events allocated `delta / 4`. theorem compiled","shard":"modules/022f2888bf1bc5ee.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_total_delta","description":"Primary residual-variance realized-regret theorem. The four explicit confidence failures total `delta`; only predictable-variance overflow remains.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedhighprobabilityregret/index.html#decl-bf70595c2612","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","order":4298,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret.lean:307"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictable…","missing":[],"search":"sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_total_delta primary residual-variance realized-regret theorem. the four explicit confidence failures total `delta`; only predictable-variance overflow remains. theorem compiled","shard":"modules/022f2888bf1bc5ee.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum","label":"sampledPredictableMixedSquaredVarianceSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum","description":"Cumulative predictable mixed-square variance on a generated trajectory.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-53e5169f3ad8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4299,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableMixedSquaredVarianceSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) : Env × ((k : Nat) → Action × Real) → Real","missing":[],"search":"sampledpredictablemixedsquaredvariancesum banditrlproof.exp3.sampledpredictablemixedsquaredvariancesum cumulative predictable mixed-square variance on a generated trajectory. definition compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledPredictableMixedSquaredVarianceSum","label":"measurable_sampledPredictableMixedSquaredVarianceSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledPredictableMixedSquaredVarianceSum","description":"theorem measurable_sampledPredictableMixedSquaredVarianceSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : Measurable (sampledPredictableMixedSquaredVa…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-7f6534ab3ce2","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4300,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledPredictableMixedSquaredVarianceSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : Measurable (sampledPredictableMixedSquaredVarianceSum arms eta gamma loss horizon)","missing":[],"search":"measurable_sampledpredictablemixedsquaredvariancesum banditrlproof.exp3.measurable_sampledpredictablemixedsquaredvariancesum theorem measurable_sampledpredictablemixedsquaredvariancesum {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (horizon : nat) : measurable (sampledpredictablemixedsquaredvariancesum arms eta gamma loss horizon) theorem compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_nonneg","label":"sampledPredictableMixedSquaredVarianceSum_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_nonneg","description":"theorem sampledPredictableMixedSquaredVarianceSum_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) :…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-8a03f3e7d255","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4301,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:45"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceSum_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : 0 ≤ sampledPredictableMixedSquaredVarianceSum arms eta gamma loss horizon sample","missing":[],"search":"sampledpredictablemixedsquaredvariancesum_nonneg banditrlproof.exp3.sampledpredictablemixedsquaredvariancesum_nonneg theorem sampledpredictablemixedsquaredvariancesum_nonneg {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) → action × real)) : 0 ≤ sampledpredictablemixedsquaredvariancesum arms eta gamma loss horizon sample theorem compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral","label":"sampledPredictableMixedSquaredVarianceLIntegral","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral","description":"`lintegral` form of the cumulative predictable mixed-square variance.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-00f92359d927","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4302,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:62"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableMixedSquaredVarianceLIntegral {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) : ENNReal","missing":[],"search":"sampledpredictablemixedsquaredvariancelintegral banditrlproof.exp3.sampledpredictablemixedsquaredvariancelintegral `lintegral` form of the cumulative predictable mixed-square variance. definition compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measure_sampledPredictableMixedSquaredVarianceSum_gt_le_lintegral_div","label":"measure_sampledPredictableMixedSquaredVarianceSum_gt_le_lintegral_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measure_sampledPredictableMixedSquaredVarianceSum_gt_le_lintegral_div","description":"Mathlib-backed Markov tail for cumulative predictable mixed-square variance. The measure need not be finite.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-a111f743c8cb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4303,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:74"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measure_sampledPredictableMixedSquaredVarianceSum_gt_le_lintegral_div {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (varianceBudget : Real) (hvarianceBudget : 0 < varianceBudget) : mu {sample | varianceBudget < sampledPredictableMixedSquaredVarianceSum arms eta gamma loss horizon sample} ≤ sampledPredictableMixedSquaredVarianceLIntegral mu arms eta gamma loss horizon / ENNReal.ofReal varianceBudget","missing":[],"search":"measure_sampledpredictablemixedsquaredvariancesum_gt_le_lintegral_div banditrlproof.exp3.measure_sampledpredictablemixedsquaredvariancesum_gt_le_lintegral_div mathlib-backed markov tail for cumulative predictable mixed-square variance. the measure need not be finite. theorem compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_of_lintegral_variance_le","label":"sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_of_lintegral_variance_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_of_lintegral_variance_le","description":"Consume a cumulative predictable-variance `lintegral` budget in the realized-regret residual theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-699f5fc4694d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4304,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:118"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_of_lintegral_variance_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (varianceBudget deltaSquare deltaConfidence deltaRealized : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) (varianceMeanBudget : Re…","missing":[],"search":"sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_of_lintegral_variance_le banditrlproof.exp3.sampledpredictable_predictablevariancesquarerealizedhighprobabilityregret_tail_of_lintegral_variance_le consume a cumulative predictable-variance `lintegral` budget in the realized-regret residual theorem. theorem compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareRealizedMarkovHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareRealizedMarkovHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareRealizedMarkovHighProbabilityRegretBudget","description":"Five-event realized-regret budget: four confidence failures and one Markov predictable-variance overflow failure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-a0681368ba44","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4305,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:191"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareRealizedMarkovHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (varianceMeanBudget delta : Real) : Real","missing":[],"search":"sampledpredictablevariancesquarerealizedmarkovhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquarerealizedmarkovhighprobabilityregretbudget five-event realized-regret budget: four confidence failures and one markov predictable-variance overflow failure. definition compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedMarkovHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_predictableVarianceSquareRealizedMarkovHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedMarkovHighProbabilityRegret_tail_total_delta","description":"Primary Markov-closed realized-regret theorem. A cumulative predictable variance `lintegral` bound is allocated the fifth failure probability.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancerealizedmarkovhighprobabilityregret/index.html#decl-ab2f5f6e8265","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","order":4306,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret.lean:201"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareRealizedMarkovHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (varianceMeanBudget delta : Real) (hvarianceMeanBudget : 0 < varianceMeanBudget) (hdelta : 0 < delta) (hvarianceLIntegral : sampledPredictableMixedSquaredVarianceLIntegral (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hg…","missing":[],"search":"sampledpredictable_predictablevariancesquarerealizedmarkovhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquarerealizedmarkovhighprobabilityregret_tail_total_delta primary markov-closed realized-regret theorem. a cumulative predictable variance `lintegral` bound is allocated the fifth failure probability. theorem compiled","shard":"modules/176571e6347c5e89.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossDoubleVarianceRealizedHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareSmallLossDoubleVarianceRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossDoubleVarianceRealizedHighProbabilityRegretBudget","description":"noncomputable def sampledPredictableVarianceSquareSmallLossDoubleVarianceRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (lossMassBudget mixedVarianceBudget realizedVarianceBudget deltaSquare deltaConfidence deltaRealized : Real) : Real","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizeddoublepredictablev-6b1be70bfcd6/index.html#decl-11dedc9f81b5","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","order":4307,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareSmallLossDoubleVarianceRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (lossMassBudget mixedVarianceBudget realizedVarianceBudget deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictablevariancesquaresmalllossdoublevariancerealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquaresmalllossdoublevariancerealizedhighprobabilityregretbudget noncomputable def sampledpredictablevariancesquaresmalllossdoublevariancerealizedhighprobabilityregretbudget {action : type v} [decidableeq action] (arms : finset action) (eta gamma : real) (horizon : nat) (lossmassbudget mixedvariancebudget realizedvariancebudget deltasquare deltaconfidence deltarealized : real) : real definition compiled","shard":"modules/e1c5c3939b6c0e15.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_smallLossDoublePredictableVarianceRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","label":"sampledPredictable_smallLossDoublePredictableVarianceRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_smallLossDoublePredictableVarianceRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","description":"Joint realized-regret tail away from an explicit loss-mass bad set, with simultaneous mixed-square and selected-loss predictable-variance budgets.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizeddoublepredictablev-6b1be70bfcd6/index.html#decl-0cb38d76c706","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","order":4308,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_smallLossDoublePredictableVarianceRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (lossMassBudget mixedVarianceBudget realizedVarianceBudget deltaSquare deltaConfidence deltaRealized : Real) (hmixedVarianceBudget : 0 < mixedVarianceBudget) (hrealizedVarianceBudget : 0 < realizedVarianceBudget) (hdeltaSquare : 0 < deltaSqua…","missing":[],"search":"sampledpredictable_smalllossdoublepredictablevariancerealizedhighprobabilityregret_tail_joint_off_bad_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictable_smalllossdoublepredictablevariancerealizedhighprobabilityregret_tail_joint_off_bad_of_lossmasssum_le_or_mem joint realized-regret tail away from an explicit loss-mass bad set, with simultaneous mixed-square and selected-loss predictable-variance budgets. theorem compiled","shard":"modules/e1c5c3939b6c0e15.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.predictableLossAt_mem_unitInterval","label":"predictableLossAt_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.predictableLossAt_mem_unitInterval","description":"Every generated predictable loss coordinate remains in the unit interval.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-083d054f1280","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4309,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) (action : Action) : predictableLossAt loss t sample action ∈ Set.Icc (0 : Real) 1","missing":[],"search":"predictablelossat_mem_unitinterval banditrlproof.exp3.predictablelossat_mem_unitinterval every generated predictable loss coordinate remains in the unit interval. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum","label":"sampledPredictableLossMassSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossMassSum","description":"Cumulative armwise predictable loss mass along a generated trajectory.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-36f8a2329c6f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4310,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:36"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableLossMassSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : Real","missing":[],"search":"sampledpredictablelossmasssum banditrlproof.exp3.sampledpredictablelossmasssum cumulative armwise predictable loss mass along a generated trajectory. definition compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredAt_le_lossMassAt","label":"sampledPredictableLossSquaredAt_le_lossMassAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossSquaredAt_le_lossMassAt","description":"Unit-interval losses have armwise square mass at most armwise loss mass.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-3b28cad0f016","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4311,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:45"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossSquaredAt_le_lossMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : arms.sum (fun action => (predictableLossAt loss t sample action) ^ 2) ≤ arms.sum (fun action => predictableLossAt loss t sample action)","missing":[],"search":"sampledpredictablelosssquaredat_le_lossmassat banditrlproof.exp3.sampledpredictablelosssquaredat_le_lossmassat unit-interval losses have armwise square mass at most armwise loss mass. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredSum_le_lossMassSum","label":"sampledPredictableLossSquaredSum_le_lossMassSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossSquaredSum_le_lossMassSum","description":"Cumulative predictable loss-square energy is at most cumulative armwise predictable loss mass.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-2a6193af8c28","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4312,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:63"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossSquaredSum_le_lossMassSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableLossSquaredSum arms loss horizon sample ≤ sampledPredictableLossMassSum arms loss horizon sample","missing":[],"search":"sampledpredictablelosssquaredsum_le_lossmasssum banditrlproof.exp3.sampledpredictablelosssquaredsum_le_lossmasssum cumulative predictable loss-square energy is at most cumulative armwise predictable loss mass. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossMassSum","label":"sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossMassSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossMassSum","description":"Cumulative predictable mixed-square variance is controlled by armwise predictable loss mass.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-ee795034a77f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4313,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:76"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossMassSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableMixedSquaredVarianceSum arms eta gamma loss horizon sample ≤ (1 / (gamma / (arms.card : Real))) * sampledPredictableLossMassSum arms loss horizon sample","missing":[],"search":"sampledpredictablemixedsquaredvariancesum_le_inv_floor_mul_lossmasssum banditrlproof.exp3.sampledpredictablemixedsquaredvariancesum_le_inv_floor_mul_lossmasssum cumulative predictable mixed-square variance is controlled by armwise predictable loss mass. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossMassSum_le","label":"sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossMassSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossMassSum_le","description":"An almost-everywhere armwise loss-mass budget supplies the cumulative-variance `lintegral` contract required by the Markov route.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-6f1e3b2986f6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4314,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:102"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossMassSum_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (lossMassBudget : Real) (hmass : ∀ᵐ sample ∂mu, sampledPredictableLossMassSum arms loss horizon sample ≤ lossMassBudget) : sampledPredictableMixedSquaredVarianceLIntegral mu arms eta gamma loss horizon ≤ ENNReal.ofReal ((1 / (gamma / (arms.card : Real))) * lossMassBudget)","missing":[],"search":"sampledpredictablemixedsquaredvariancelintegral_le_of_lossmasssum_le banditrlproof.exp3.sampledpredictablemixedsquaredvariancelintegral_le_of_lossmasssum_le an almost-everywhere armwise loss-mass budget supplies the cumulative-variance `lintegral` contract required by the markov route. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_off_bad_of_lossMassSum_le_or_mem","label":"sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_off_bad_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_off_bad_of_lossMassSum_le_or_mem","description":"Off-bad observed-square tail when the armwise loss-mass budget may fail on an explicit bad set. No measurability assumption on that set is needed.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-e706f9511bc8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4315,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:150"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_off_bad_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (lossMassBudget varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) (bad : Set (Env × ((k : Nat) → Action × Real))) (hmass : ∀ᵐ sample ∂(prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment), sampledPredictableLossMassSum arms loss horizon…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_off_bad_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_off_bad_of_lossmasssum_le_or_mem off-bad observed-square tail when the armwise loss-mass budget may fail on an explicit bad set. no measurability assumption on that set is needed. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le_or_mem","label":"sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le_or_mem","description":"Residual observed-square tail when the armwise loss-mass budget may fail on an explicit bad set.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-4d14cb6eb09c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4316,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:233"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (lossMassBudget varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) (bad : Set (Env × ((k : Nat) → Action × Real))) (hmass : ∀ᵐ sample ∂(prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment), sampledPredictableLossMassSum arms loss horizon sample ≤…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_of_lossmasssum_le_or_mem residual observed-square tail when the armwise loss-mass budget may fail on an explicit bad set. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le","label":"sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le","description":"The observed mixed estimator-square sum has its predictable mean bounded by the supplied almost-everywhere armwise loss-mass budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-20f61e629ced","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4317,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:316"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (lossMassBudget varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) (hmass : ∀ᵐ sample ∂(prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment), sampledPredictableLossMassSum arms loss horizon sample ≤ lossMassBudget) : let mu := prior ⊗ₘ sampledImportance…","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_of_lossmasssum_le banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_predictablevariance_of_lossmasssum_le the observed mixed estimator-square sum has its predictable mean bounded by the supplied almost-everywhere armwise loss-mass budget. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareSmallLossHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossHighProbabilityRegretBudget","description":"Predictable regret budget whose mixed-square mean upper bound is the supplied armwise loss-mass budget rather than `K * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-c419f293f911","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4318,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:353"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareSmallLossHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (lossMassBudget varianceBudget deltaSquare deltaConfidence : Real) : Real","missing":[],"search":"sampledpredictablevariancesquaresmalllosshighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquaresmalllosshighprobabilityregretbudget predictable regret budget whose mixed-square mean upper bound is the supplied armwise loss-mass budget rather than `k * t`. definition compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","label":"sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","description":"Off-bad predictable small-loss regret on the variance-good event. The explicit bad set is removed from the source event, so it does not enter the confidence allocation.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-e09a756207f9","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4319,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:371"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (lossMassBudget varianceBudget : Real) (deltaSquare deltaConfidence : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (bad : Set (Env × ((k : Nat) → Action × Real))) (hmass : ∀ᵐ s…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllosshighprobabilityregret_tail_joint_off_bad_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllosshighprobabilityregret_tail_joint_off_bad_of_lossmasssum_le_or_mem off-bad predictable small-loss regret on the variance-good event. the explicit bad set is removed from the source event, so it does not enter the confidence allocation. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","label":"sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","description":"Residual small-loss predictable regret on the variance-good event when the loss-mass budget may fail on an explicit bad set.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-65a88dfe3897","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4320,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:559"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (lossMassBudget varianceBudget : Real) (deltaSquare deltaConfidence : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (bad : Set (Env × ((k : Nat) → Action × Real))) (hmass : ∀ᵐ sample ∂(…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllosshighprobabilityregret_tail_joint_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllosshighprobabilityregret_tail_joint_of_lossmasssum_le_or_mem residual small-loss predictable regret on the variance-good event when the loss-mass budget may fail on an explicit bad set. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint","label":"sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint","description":"Small-loss predictable regret on the event that cumulative predictable mixed-square variance stays below `varianceBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-abe701704f09","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4321,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:763"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (lossMassBudget varianceBudget : Real) (deltaSquare deltaConfidence : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hmass : ∀ᵐ sample ∂(prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma h…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllosshighprobabilityregret_tail_joint banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllosshighprobabilityregret_tail_joint small-loss predictable regret on the event that cumulative predictable mixed-square variance stays below `variancebudget`. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossRealizedHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareSmallLossRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossRealizedHighProbabilityRegretBudget","description":"Realized small-loss regret budget with a caller-supplied predictable variance budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-94e0d797f78c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4322,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:810"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareSmallLossRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (lossMassBudget varianceBudget deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictablevariancesquaresmalllossrealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquaresmalllossrealizedhighprobabilityregretbudget realized small-loss regret budget with a caller-supplied predictable variance budget. definition compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","label":"sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","description":"Off-bad realized small-loss regret on the predictable-variance-good event. The four confidence events are charged here; the explicit bad set can be charged once by a downstream pathwise decomposition.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-06f861419730","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4323,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:822"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (lossMassBudget varianceBudget : Real) (deltaSquare deltaConfidence deltaRealized : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealize…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllossrealizedhighprobabilityregret_tail_joint_off_bad_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllossrealizedhighprobabilityregret_tail_joint_off_bad_of_lossmasssum_le_or_mem off-bad realized small-loss regret on the predictable-variance-good event. the four confidence events are charged here; the explicit bad set can be charged once by a downstream pathwise decomposition. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","label":"sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","description":"Residual realized small-loss regret on the predictable-variance-good event when the loss-mass budget may fail on an explicit bad set.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-a1e8dd59815f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4324,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:960"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (lossMassBudget varianceBudget : Real) (deltaSquare deltaConfidence deltaRealized : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 <…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllossrealizedhighprobabilityregret_tail_joint_of_lossmasssum_le_or_mem banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllossrealizedhighprobabilityregret_tail_joint_of_lossmasssum_le_or_mem residual realized small-loss regret on the predictable-variance-good event when the loss-mass budget may fail on an explicit bad set. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint","label":"sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint","description":"Realized small-loss regret on the predictable-variance-good event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-0465a6602094","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4325,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:1113"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (lossMassBudget varianceBudget : Real) (deltaSquare deltaConfidence deltaRealized : Real) (hvarianceBudget : 0 < varianceBudget) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) (hmass : ∀…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllossrealizedhighprobabilityregret_tail_joint banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllossrealizedhighprobabilityregret_tail_joint realized small-loss regret on the predictable-variance-good event. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegretBudget","description":"Five-event small-loss realized budget: the mixed-square predictable mean is `lossMassBudget`, while the Markov variance mean is `(1 / (gamma / K)) * lossMassBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-b3232d2cdae4","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4326,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:1163"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (lossMassBudget delta : Real) : Real","missing":[],"search":"sampledpredictablevariancesquaresmalllossrealizedmarkovhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquaresmalllossrealizedmarkovhighprobabilityregretbudget five-event small-loss realized budget: the mixed-square predictable mean is `lossmassbudget`, while the markov variance mean is `(1 / (gamma / k)) * lossmassbudget`. definition compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_predictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegret_tail_total_delta","description":"Primary armwise small-loss specialization of the Markov-closed realized EXP3 theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesmalllossrealizedmarkovhighprobabilityregret/index.html#decl-fce74a7e329e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","order":4327,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret.lean:1174"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (lossMassBudget delta : Real) (hlossMassBudget : 0 < lossMassBudget) (hdelta : 0 < delta) (hmass : ∀ᵐ sample ∂(prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment), sampl…","missing":[],"search":"sampledpredictable_predictablevariancesquaresmalllossrealizedmarkovhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquaresmalllossrealizedmarkovhighprobabilityregret_tail_total_delta primary armwise small-loss specialization of the markov-closed realized exp3 theorem. theorem compiled","shard":"modules/fcce993f0a19b8c3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSparseRealizedVarianceBudget","label":"sampledPredictableSparseRealizedVarianceBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSparseRealizedVarianceBudget","description":"noncomputable def sampledPredictableSparseRealizedVarianceBudget (horizon sparsity : Nat) : Real","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html#decl-d1fe63eba742","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","order":4328,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableSparseRealizedVarianceBudget (horizon sparsity : Nat) : Real","missing":[],"search":"sampledpredictablesparserealizedvariancebudget banditrlproof.exp3.sampledpredictablesparserealizedvariancebudget noncomputable def sampledpredictablesparserealizedvariancebudget (horizon sparsity : nat) : real definition compiled","shard":"modules/659a04d8d9ce0435.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceSum_le_sparseRealizedVarianceBudget_or_mem_sparsityFailure","label":"sampledPredictableRealizedVarianceSum_le_sparseRealizedVarianceBudget_or_mem_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedVarianceSum_le_sparseRealizedVarianceBudget_or_mem_sparsityFailure","description":"theorem sampledPredictableRealizedVarianceSum_le_sparseRealizedVarianceBudget_or_mem_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon sparsity :…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html#decl-ad14fbf318dc","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","order":4329,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedVarianceSum_le_sparseRealizedVarianceBudget_or_mem_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample) ≤ sampledPredictableSparseRealizedVarianceBudget horizon sparsity ∨ sample ∈ sampledPredictableSparsityFailure arms loss horizon sparsity","missing":[],"search":"sampledpredictablerealizedvariancesum_le_sparserealizedvariancebudget_or_mem_sparsityfailure banditrlproof.exp3.sampledpredictablerealizedvariancesum_le_sparserealizedvariancebudget_or_mem_sparsityfailure theorem sampledpredictablerealizedvariancesum_le_sparserealizedvariancebudget_or_mem_sparsityfailure {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (horizon sparsity : nat) (sample : env × ((k : nat) → action × real)) : (finset.range horizon).sum (fun i => sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss i sample) ≤ sampledpredictablesparserealizedvariancebudget horizon sparsity ∨ sample ∈ sampledpredictablesparsityfailure arms loss horizon sparsity theorem compiled","shard":"modules/659a04d8d9ce0435.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget","label":"sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget","description":"noncomputable def sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html#decl-b44acfe8863c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","order":4330,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean:52"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget noncomputable def sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget {action : type v} [decidableeq action] (arms : finset action) (eta gamma : real) (horizon sparsity : nat) (delta : real) : real definition compiled","shard":"modules/659a04d8d9ce0435.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_off_sparsityFailure","label":"sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_off_sparsityFailure","description":"theorem sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harm…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html#decl-9bda879dd662","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","order":4331,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean:64"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu ({sample | sampledPredicta…","missing":[],"search":"sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail_off_sparsityfailure theorem sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail_off_sparsityfailure {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu ({sample | sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget arms eta gamma horizon sparsity delta ≤ (finset.range horizon).sum (fun t => sampledtrajectoryrealizedlossat t sample) - (finset.range horizon).sum (fun t => predictablelossat loss t…","shard":"modules/659a04d8d9ce0435.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail","label":"sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail","description":"theorem sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html#decl-2148ea64d3ad","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","order":4332,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean:214"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableDoubleVarianceProb…","missing":[],"search":"sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail theorem sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget arms eta gamma horizon sparsity delta ≤ (finset.range horizon).sum (fun t => sampledtrajectoryrealizedlossat t sample) - (finset.range horizon).sum (fun t => predictablelossat loss t sample comparator)} ≤ ennreal.ofreal delta + mu (sampledpredi…","shard":"modules/659a04d8d9ce0435.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_of_sparsityFailure_le","description":"theorem sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (ha…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-6cdcca1a3d1b/index.html#decl-743b3277f52e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","order":4333,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity.lean:271"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hfailure : (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment) (sampledPredictab…","missing":[],"search":"sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail_of_sparsityfailure_le theorem sampledpredictable_doublevarianceprobabilisticsparselossrealizedhighprobabilityregret_tail_of_sparsityfailure_le {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : predictablelossvector env action) (comparator : action) (hcomparator : comparator ∈ arms) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : real) (hdelta : 0 < delta) (hfailure : (prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment) (sampledpredictablesparsityfailure arms loss horizon sparsity) ≤ ennreal.ofreal epsilon) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledpredictabledoublevarianceprobab…","shard":"modules/659a04d8d9ce0435.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossLargeHorizonCondition","label":"doubleVarianceProbabilisticSparseLossLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossLargeHorizonCondition","description":"Regime in which every component of the exact double-variance exploration schedule is at most one half.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f7cd8884d456/index.html#decl-aa16d47b765f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","order":4334,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def doubleVarianceProbabilisticSparseLossLargeHorizonCondition (K S T delta : Real) : Prop","missing":[],"search":"doublevarianceprobabilisticsparselosslargehorizoncondition banditrlproof.exp3.doublevarianceprobabilisticsparselosslargehorizoncondition regime in which every component of the exact double-variance exploration schedule is at most one half. definition compiled","shard":"modules/877b96a01f3788a5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossAllHorizonRegretThreshold","label":"doubleVarianceProbabilisticSparseLossAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossAllHorizonRegretThreshold","description":"All-horizon threshold for the exact double-variance sparse route: use the refined threshold in its valid regime and strict `T + 1` otherwise.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f7cd8884d456/index.html#decl-7d57ce01158d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","order":4335,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon.lean:34"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def doubleVarianceProbabilisticSparseLossAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"doublevarianceprobabilisticsparselossallhorizonregretthreshold banditrlproof.exp3.doublevarianceprobabilisticsparselossallhorizonregretthreshold all-horizon threshold for the exact double-variance sparse route: use the refined threshold in its valid regime and strict `t + 1` otherwise. definition compiled","shard":"modules/877b96a01f3788a5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","description":"Generated all-horizon exact double-variance regret away from the common sparsity-failure event. Both branches use the same eta, gamma, and measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f7cd8884d456/index.html#decl-cb0fdb2892bd","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","order":4336,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon.lean:50"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := pathwiseVarianceProbabilisticSparseLossHighProbabil…","missing":[],"search":"sampledpredictable_allhorizondoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_allhorizondoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure generated all-horizon exact double-variance regret away from the common sparsity-failure event. both branches use the same eta, gamma, and measure. theorem compiled","shard":"modules/877b96a01f3788a5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","label":"sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","description":"Generated all-horizon exact double-variance regret with the common sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f7cd8884d456/index.html#decl-bceebd62c54b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","order":4337,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon.lean:133"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms…","missing":[],"search":"sampledpredictable_allhorizondoublevarianceprobabilisticsparselossrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizondoublevarianceprobabilisticsparselossrealizedregret_tail generated all-horizon exact double-variance regret with the common sparsity-failure residual. theorem compiled","shard":"modules/877b96a01f3788a5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","description":"Practical all-horizon `delta + epsilon` theorem under the exact same-measure sparsity-failure bound.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f7cd8884d456/index.html#decl-9874ec4930de","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","order":4338,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon.lean:215"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := pathwiseVarianceProbabilisticSparseLossHi…","missing":[],"search":"sampledpredictable_allhorizondoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_allhorizondoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le practical all-horizon `delta + epsilon` theorem under the exact same-measure sparsity-failure bound. theorem compiled","shard":"modules/877b96a01f3788a5.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","label":"doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","description":"Best-arm exact double-variance threshold. The fixed-comparator schedule receives the armwise confidence share `delta / K`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f70e76a4b1d2/index.html#decl-ca9b468123c6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4339,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"doublevarianceprobabilisticsparselossbestarmallhorizonregretthreshold banditrlproof.exp3.doublevarianceprobabilisticsparselossbestarmallhorizonregretthreshold best-arm exact double-variance threshold. the fixed-comparator schedule receives the armwise confidence share `delta / k`. definition compiled","shard":"modules/529aad5bcb19c46d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","description":"Best-arm all-horizon tail away from the common support-sparsity failure event. The comparator union spends only the armwise confidence shares.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f70e76a4b1d2/index.html#decl-9170364091cd","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4340,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:29"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHighProbability…","missing":[],"search":"sampledpredictable_allhorizondoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_allhorizondoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_off_sparsityfailure best-arm all-horizon tail away from the common support-sparsity failure event. the comparator union spends only the armwise confidence shares. theorem compiled","shard":"modules/529aad5bcb19c46d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","label":"sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","description":"Best-arm all-horizon residual theorem. The common support-sparsity failure event is charged exactly once after the comparator union.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f70e76a4b1d2/index.html#decl-18ee8ce0cdc5","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4341,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:178"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms ga…","missing":[],"search":"sampledpredictable_allhorizondoublevarianceprobabilisticsparselossbestarmrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizondoublevarianceprobabilisticsparselossbestarmrealizedregret_tail best-arm all-horizon residual theorem. the common support-sparsity failure event is charged exactly once after the comparator union. theorem compiled","shard":"modules/529aad5bcb19c46d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","description":"Practical exact double-variance best-arm theorem with a single charge for the common support-sparsity failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-f70e76a4b1d2/index.html#decl-7eaa23bf9549","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4342,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:275"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHighP…","missing":[],"search":"sampledpredictable_allhorizondoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_allhorizondoublevarianceprobabilisticsparselossbestarmrealizedregret_tail_of_sparsityfailure_le practical exact double-variance best-arm theorem with a single charge for the common support-sparsity failure event. theorem compiled","shard":"modules/529aad5bcb19c46d.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRealizedExplorationScale","label":"doubleVarianceProbabilisticSparseLossRealizedExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRealizedExplorationScale","description":"Exploration scale forced by the exact selected-loss predictable-variance radius with budget `S * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-cab299f664d3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4343,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def doubleVarianceProbabilisticSparseLossRealizedExplorationScale (S T delta : Real) : Real","missing":[],"search":"doublevarianceprobabilisticsparselossrealizedexplorationscale banditrlproof.exp3.doublevarianceprobabilisticsparselossrealizedexplorationscale exploration scale forced by the exact selected-loss predictable-variance radius with budget `s * t`. definition compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate","label":"doubleVarianceProbabilisticSparseLossRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate","description":"The previous pathwise raw schedule augmented by the exact selected-loss predictable-variance scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-3b80d9a7b2d9","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4344,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def doubleVarianceProbabilisticSparseLossRawExplorationRate (K S T delta : Real) : Real","missing":[],"search":"doublevarianceprobabilisticsparselossrawexplorationrate banditrlproof.exp3.doublevarianceprobabilisticsparselossrawexplorationrate the previous pathwise raw schedule augmented by the exact selected-loss predictable-variance scale. definition compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate","label":"doubleVarianceProbabilisticSparseLossClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate","description":"Explicit double-variance exploration schedule clipped into the Hedge stability regime.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-d94d40a95362","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4345,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:41"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def doubleVarianceProbabilisticSparseLossClippedExplorationRate (K S T delta : Real) : Real","missing":[],"search":"doublevarianceprobabilisticsparselossclippedexplorationrate banditrlproof.exp3.doublevarianceprobabilisticsparselossclippedexplorationrate explicit double-variance exploration schedule clipped into the hedge stability regime. definition compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_le_half","label":"doubleVarianceProbabilisticSparseLossClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_le_half","description":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_le_half (K S T delta : Real) : doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta <= 1 / 2","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-baf318170986","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4346,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:46"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_le_half (K S T delta : Real) : doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta <= 1 / 2","missing":[],"search":"doublevarianceprobabilisticsparselossclippedexplorationrate_le_half banditrlproof.exp3.doublevarianceprobabilisticsparselossclippedexplorationrate_le_half theorem doublevarianceprobabilisticsparselossclippedexplorationrate_le_half (k s t delta : real) : doublevarianceprobabilisticsparselossclippedexplorationrate k s t delta <= 1 / 2 theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","label":"doubleVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","description":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta <= 1 / 2) : doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta = doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-3baaadc573b8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4347,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:53"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta <= 1 / 2) : doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta = doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta","missing":[],"search":"doublevarianceprobabilisticsparselossclippedexplorationrate_eq_raw banditrlproof.exp3.doublevarianceprobabilisticsparselossclippedexplorationrate_eq_raw theorem doublevarianceprobabilisticsparselossclippedexplorationrate_eq_raw (k s t delta : real) (hraw : doublevarianceprobabilisticsparselossrawexplorationrate k s t delta <= 1 / 2) : doublevarianceprobabilisticsparselossclippedexplorationrate k s t delta = doublevarianceprobabilisticsparselossrawexplorationrate k s t delta theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate_pos","label":"doubleVarianceProbabilisticSparseLossRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate_pos","description":"theorem doubleVarianceProbabilisticSparseLossRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-bfc111ea21c9","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4348,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:65"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta","missing":[],"search":"doublevarianceprobabilisticsparselossrawexplorationrate_pos banditrlproof.exp3.doublevarianceprobabilisticsparselossrawexplorationrate_pos theorem doublevarianceprobabilisticsparselossrawexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < doublevarianceprobabilisticsparselossrawexplorationrate k s t delta theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_pos","label":"doubleVarianceProbabilisticSparseLossClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_pos","description":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-ca7744af6168","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4349,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:73"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta","missing":[],"search":"doublevarianceprobabilisticsparselossclippedexplorationrate_pos banditrlproof.exp3.doublevarianceprobabilisticsparselossclippedexplorationrate_pos theorem doublevarianceprobabilisticsparselossclippedexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < doublevarianceprobabilisticsparselossclippedexplorationrate k s t delta theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.boundedRealizedLargeHorizon_of_doubleVarianceLargeHorizon","label":"boundedRealizedLargeHorizon_of_doubleVarianceLargeHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.boundedRealizedLargeHorizon_of_doubleVarianceLargeHorizon","description":"The selected-loss horizon contract also controls the older bounded realized-deviation component retained by the shared raw schedule.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-c5ba9ccc74c6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4350,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:86"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem boundedRealizedLargeHorizon_of_doubleVarianceLargeHorizon (S T delta : Real) (hS_one : 1 <= S) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_realized : 4 * (S * Real.log (4 / delta)) <= T) : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T","missing":[],"search":"boundedrealizedlargehorizon_of_doublevariancelargehorizon banditrlproof.exp3.boundedrealizedlargehorizon_of_doublevariancelargehorizon the selected-loss horizon contract also controls the older bounded realized-deviation component retained by the shared raw schedule. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","label":"doubleVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","description":"Four horizon contracts place every component of the double-variance raw schedule below one half.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-71a2c61f7b35","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4351,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:106"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts (K S T delta : Real) (hK_one : 1 < K) (hS_one : 1 <= S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (K * S * Real.log K ^ 2 * Real.log (4 / delta)) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 4 * (S * Real.log (4 / delta)) <= T) : doubleVarianceProbabilisticSparseLossRawExplorationRate K S T delta <= 1 / 2","missing":[],"search":"doublevarianceprobabilisticsparselossrawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.doublevarianceprobabilisticsparselossrawexplorationrate_le_half_of_horizon_contracts four horizon contracts place every component of the double-variance raw schedule below one half. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_contracts","label":"doubleVarianceProbabilisticSparseLossClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_contracts","description":"The clipped schedule supplies the base, mixed-square, Bernstein, and selected-loss predictable-variance contracts.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-1d2b0c6d0356","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4352,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:144"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem doubleVarianceProbabilisticSparseLossClippedExplorationRate_contracts (K S T delta : Real) (hK_one : 1 < K) (hS_one : 1 <= S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (K * S * Real.log K ^ 2 * Real.log (4 / delta)) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 4 * (S * Real.log (4 / delta)) <= T) : let gamma := doubleVarianceProbabilisticSparseLossClippedExplorationRate K S T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ S * Real.log K <= gamma ^ 2 * T ∧ K * S * Real.log K ^ 2 * Real.log (4 / delta) <= gamma ^ 5 * T ^ 3 ∧ K * Real.log (4 / delta) <= gamma ^ 3 * T ∧ S * Real.log (4 / delta) <= gamma ^ 2 * T","missing":[],"search":"doublevarianceprobabilisticsparselossclippedexplorationrate_contracts banditrlproof.exp3.doublevarianceprobabilisticsparselossclippedexplorationrate_contracts the clipped schedule supplies the base, mixed-square, bernstein, and selected-loss predictable-variance contracts. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedExplicitThreshold","label":"pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedExplicitThreshold","description":"Explicit double-variance threshold after tuning eta and gamma.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-f79c806a06f4","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4353,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:267"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedExplicitThreshold {Action : Type v} (_arms : Finset Action) (gamma : Real) (horizon _sparsity : Nat) (_delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossdoublevariancerealizedexplicitthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossdoublevariancerealizedexplicitthreshold explicit double-variance threshold after tuning eta and gamma. definition compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseRealizedPredictableVarianceRadius_le_three_mul_gamma_mul_horizon","label":"sparseRealizedPredictableVarianceRadius_le_three_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseRealizedPredictableVarianceRadius_le_three_mul_gamma_mul_horizon","description":"The exact selected-loss predictable-variance radius is at most `3 * gamma * T` under its sparse variance contract.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-5d4805a6db9c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4354,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:274"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseRealizedPredictableVarianceRadius_le_three_mul_gamma_mul_horizon (S T gamma delta : Real) (hS_one : 1 <= S) (hT : 0 < T) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrealized : S * Real.log (4 / delta) <= gamma ^ 2 * T) : sampledRealizedPredictableVarianceRadius (S * T) (delta / 4) <= 3 * gamma * T","missing":[],"search":"sparserealizedpredictablevarianceradius_le_three_mul_gamma_mul_horizon banditrlproof.exp3.sparserealizedpredictablevarianceradius_le_three_mul_gamma_mul_horizon the exact selected-loss predictable-variance radius is at most `3 * gamma * t` under its sparse variance contract. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold_le_explicitThreshold","label":"pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold_le_explicitThreshold","description":"The four exploration contracts reduce the eta-tuned double-variance threshold to `16 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-237b56778da0","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4355,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:321"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) <= gamma ^ 5 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) (hrealized : (sparsity : Real) * Real.log (4 / delta) <= gamma ^ 2 * (horizon : Real)) : pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold arms gam…","missing":[],"search":"pathwisevarianceprobabilisticsparselossdoublevariancerealizedtunedthreshold_le_explicitthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossdoublevariancerealizedtunedthreshold_le_explicitthreshold the four exploration contracts reduce the eta-tuned double-variance threshold to `16 * gamma * t`. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","description":"Gamma-characterized double-variance regret tail away from the exact sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-8dbfa84914d9","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4356,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:419"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Re…","missing":[],"search":"sampledpredictable_gammacharacterizeddoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_gammacharacterizeddoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure gamma-characterized double-variance regret tail away from the exact sparsity-failure event. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","label":"sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","description":"Gamma-characterized double-variance regret with the exact sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-cda16f5e277d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4357,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:505"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) * (sparsity : Re…","missing":[],"search":"sampledpredictable_gammacharacterizeddoublevarianceprobabilisticsparselossrealizedregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizeddoublevarianceprobabilisticsparselossrealizedregret_tail gamma-characterized double-variance regret with the exact sparsity-failure residual. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","description":"Practical gamma-characterized endpoint under the exact generated-measure bound on the sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-c8db007083a0","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4358,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:564"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms…","missing":[],"search":"sampledpredictable_gammacharacterizeddoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_gammacharacterizeddoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le practical gamma-characterized endpoint under the exact generated-measure bound on the sparsity-failure event. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","description":"Fully explicit double-variance regret tail away from the exact sparsity-failure event for the clipped schedule.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-357b55da9ba6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4359,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:620"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * ((arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 /…","missing":[],"search":"sampledpredictable_explicitdoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_explicitdoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure fully explicit double-variance regret tail away from the exact sparsity-failure event for the clipped schedule. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","label":"sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","description":"Fully explicit double-variance regret with the exact sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-54f354b05413","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4360,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:697"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * ((arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta)) <= (horizon…","missing":[],"search":"sampledpredictable_explicitdoublevarianceprobabilisticsparselossrealizedregret_tail banditrlproof.exp3.sampledpredictable_explicitdoublevarianceprobabilisticsparselossrealizedregret_tail fully explicit double-variance regret with the exact sparsity-failure residual. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","description":"Fully explicit practical `delta + epsilon` theorem under the exact sparsity-failure bound for the internally eta/gamma-tuned measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-12b618e35279/index.html#decl-0e92eb96573e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","order":4361,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning.lean:774"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * ((arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Rea…","missing":[],"search":"sampledpredictable_explicitdoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_explicitdoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le fully explicit practical `delta + epsilon` theorem under the exact sparsity-failure bound for the internally eta/gamma-tuned measure. theorem compiled","shard":"modules/2410943acc92ee4e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold","label":"pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold","description":"Eta-tuned threshold with both sparse pathwise predictable variances.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-86f08ff5ad1a/index.html#decl-232ae2f612a3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","order":4362,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossdoublevariancerealizedtunedthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossdoublevariancerealizedtunedthreshold eta-tuned threshold with both sparse pathwise predictable variances. definition compiled","shard":"modules/507ceac1b8ba411e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget_le_tunedThreshold","description":"The sparse small-loss double-variance budget is bounded by the eta-tuned threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-86f08ff5ad1a/index.html#decl-53f92fbbdf30","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","order":4363,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning.lean:39"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget arms (pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta) gamma horizon sparsity delta ≤ pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold arms gamma horizon sparsity delta","missing":[],"search":"sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictabledoublevarianceprobabilisticsparselossrealizedhighprobabilityregretbudget_le_tunedthreshold the sparse small-loss double-variance budget is bounded by the eta-tuned threshold. theorem compiled","shard":"modules/507ceac1b8ba411e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","description":"Eta-tuned generated regret away from the common sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-86f08ff5ad1a/index.html#decl-ad26a6068077","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","order":4364,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning.lean:67"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImportanceWeightedT…","missing":[],"search":"sampledpredictable_tuneddoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_tuneddoublevarianceprobabilisticsparselossrealizedregret_tail_off_sparsityfailure eta-tuned generated regret away from the common sparsity-failure event. theorem compiled","shard":"modules/507ceac1b8ba411e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","label":"sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","description":"Eta-tuned generated regret with the exact sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-86f08ff5ad1a/index.html#decl-0806ae2c3fde","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","order":4365,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning.lean:150"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms…","missing":[],"search":"sampledpredictable_tuneddoublevarianceprobabilisticsparselossrealizedregret_tail banditrlproof.exp3.sampledpredictable_tuneddoublevarianceprobabilisticsparselossrealizedregret_tail eta-tuned generated regret with the exact sparsity-failure residual. theorem compiled","shard":"modules/507ceac1b8ba411e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","description":"Practical eta-tuned `delta + epsilon` theorem under the internally tuned generated measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizeddoublepathwisevar-86f08ff5ad1a/index.html#decl-1f3df69d5d4e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","order":4366,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning.lean:204"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) : let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImportanc…","missing":[],"search":"sampledpredictable_tuneddoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_tuneddoublevarianceprobabilisticsparselossrealizedregret_tail_of_sparsityfailure_le practical eta-tuned `delta + epsilon` theorem under the internally tuned generated measure. theorem compiled","shard":"modules/507ceac1b8ba411e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_ae_sparsity","label":"sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_ae_sparsity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_ae_sparsity","description":"Generated sparse-loss realized-regret tail for every positive horizon when support sparsity holds almost everywhere under the exact internally tuned trajectory measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovaesparsityallhorizon/index.html#decl-e5550b0530ff","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","order":4367,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_ae_sparsity {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := sparseLossPredictableVarianceClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := sparseLossPredictableVarianceHighProbabilityLearningRate arms gamm…","missing":[],"search":"sampledpredictable_allhorizonsparselosspredictablevariancerealizedmarkovregret_tail_of_ae_sparsity banditrlproof.exp3.sampledpredictable_allhorizonsparselosspredictablevariancerealizedmarkovregret_tail_of_ae_sparsity generated sparse-loss realized-regret tail for every positive horizon when support sparsity holds almost everywhere under the exact internally tuned trajectory measure. theorem compiled","shard":"modules/fae895a5c7b5f65c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceLargeHorizonCondition","label":"sparseLossPredictableVarianceLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceLargeHorizonCondition","description":"The regime in which all four components of the sparse-loss predictable-variance exploration schedule are at most one half.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovallhorizon/index.html#decl-30f1fdd344ee","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","order":4368,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def sparseLossPredictableVarianceLargeHorizonCondition (K S T delta : Real) : Prop","missing":[],"search":"sparselosspredictablevariancelargehorizoncondition banditrlproof.exp3.sparselosspredictablevariancelargehorizoncondition the regime in which all four components of the sparse-loss predictable-variance exploration schedule are at most one half. definition compiled","shard":"modules/95cf2e92e47865e3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceAllHorizonRegretThreshold","label":"sparseLossPredictableVarianceAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceAllHorizonRegretThreshold","description":"All-horizon threshold for the sparse-loss predictable-variance route: use the explicit large-horizon rate in its valid regime and `T + 1` otherwise.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovallhorizon/index.html#decl-7431d1e853a8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","order":4369,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon.lean:37"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sparselosspredictablevarianceallhorizonregretthreshold banditrlproof.exp3.sparselosspredictablevarianceallhorizonregretthreshold all-horizon threshold for the sparse-loss predictable-variance route: use the explicit large-horizon rate in its valid regime and `t + 1` otherwise. definition compiled","shard":"modules/95cf2e92e47865e3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Generated sparse-loss realized-regret tail for every positive horizon under the exact eta and clipped gamma schedules. The refined threshold is used precisely in the four-contract regime; the complementary branch is the genuine zero-probability `T + 1` fallback.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovallhorizon/index.html#decl-33f84382a3c7","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","order":4370,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon.lean:54"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hsparse : ∀ sample t, t < horizon → (sampledPredictableLossSupport arms loss t sample).card <= sparsity) : let gamma := sparseLossPredictableVarianceClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon…","missing":[],"search":"sampledpredictable_allhorizonsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_allhorizonsparselosspredictablevariancerealizedmarkovregret_tail generated sparse-loss realized-regret tail for every positive horizon under the exact eta and clipped gamma schedules. the refined threshold is used precisely in the four-contract regime; the complementary branch is the genuine zero-probability `t + 1` fallback. theorem compiled","shard":"modules/95cf2e92e47865e3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_one_div_fifth_eq_log_five_div","label":"log_one_div_fifth_eq_log_five_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_one_div_fifth_eq_log_five_div","description":"The five-way confidence allocation has logarithmic budget `log (5 / delta)`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-c0cc27cc5192","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4371,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_one_div_fifth_eq_log_five_div (delta : Real) (hdelta : 0 < delta) : Real.log (1 / (delta / 5)) = Real.log (5 / delta)","missing":[],"search":"log_one_div_fifth_eq_log_five_div banditrlproof.exp3.log_one_div_fifth_eq_log_five_div the five-way confidence allocation has logarithmic budget `log (5 / delta)`. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBudget_eq","label":"sparseLossPredictableVarianceBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceBudget_eq","description":"Closed form of the sparse-loss Markov predictable-variance budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-208f54685c41","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4372,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:38"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceBudget_eq {Action : Type v} (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (delta : Real) (hdelta : 0 < delta) : sparseLossPredictableVarianceBudget arms gamma horizon sparsity delta = 5 * (arms.card : Real) * (sparsity : Real) * (horizon : Real) / (gamma * delta)","missing":[],"search":"sparselosspredictablevariancebudget_eq banditrlproof.exp3.sparselosspredictablevariancebudget_eq closed form of the sparse-loss markov predictable-variance budget. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_mul_sparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","label":"log_mul_sparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_mul_sparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","description":"Under the sparse base, fifth-power Markov, and cubic confidence contracts, the log-weighted mixed-square radius is at most `3 * gamma^2 * T^2`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-f987fca5022b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4373,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:52"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_mul_sparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (5 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.log (arms.card : Real) * sampledMixedSquaredPredictableVarianceRadius arms gamma (sparseLossPredictableVarianceBudget arms gamma horizon sparsity delta) (delta / 5) <=…","missing":[],"search":"log_mul_sparselosspredictablevarianceradius_le_three_mul_sq_mul_horizon_sq banditrlproof.exp3.log_mul_sparselosspredictablevarianceradius_le_three_mul_sq_mul_horizon_sq under the sparse base, fifth-power markov, and cubic confidence contracts, the log-weighted mixed-square radius is at most `3 * gamma^2 * t^2`. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","label":"sparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","description":"The sparse base term and Markov predictable-variance radius make the learning-rate-balanced square root at most `2 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-1243920b6dc1","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4374,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:175"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (5 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.sqrt (Real.log (arms.card : Real) * sparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta) <= 2 * gamma * (horizon : Real)","missing":[],"search":"sparselosspredictablevariancebalancedsqrt_le_two_mul_gamma_mul_horizon banditrlproof.exp3.sparselosspredictablevariancebalancedsqrt_le_two_mul_gamma_mul_horizon the sparse base term and markov predictable-variance radius make the learning-rate-balanced square root at most `2 * gamma * t`. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovExplicitThreshold","label":"sparseLossPredictableVarianceRealizedMarkovExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovExplicitThreshold","description":"Explicit sparse-loss realized-regret threshold after exploration tuning.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-dc61f6507bd2","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4375,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:242"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceRealizedMarkovExplicitThreshold {Action : Type v} (_arms : Finset Action) (gamma : Real) (horizon _sparsity : Nat) (_delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancerealizedmarkovexplicitthreshold banditrlproof.exp3.sparselosspredictablevariancerealizedmarkovexplicitthreshold explicit sparse-loss realized-regret threshold after exploration tuning. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","label":"sparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","description":"The four algebraic exploration contracts reduce the eta-tuned threshold to `14 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-225c072a48c4","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4376,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:249"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (5 / delta) <= gamma ^ 3 * (horizon : Real)) (hrealized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= gamma ^ 2 * (horizon : Real)) : sparseLossPredictableVarianceRealizedMarkovT…","missing":[],"search":"sparselosspredictablevariancerealizedmarkovtunedthreshold_le_explicitthreshold banditrlproof.exp3.sparselosspredictablevariancerealizedmarkovtunedthreshold_le_explicitthreshold the four algebraic exploration contracts reduce the eta-tuned threshold to `14 * gamma * t`. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_gammaCharacterizedSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Generated sparse-loss realized-regret tail under four algebraic exploration contracts.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-3ac86b5a7dcb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4377,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:345"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) * (sparsity :…","missing":[],"search":"sampledpredictable_gammacharacterizedsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizedsparselosspredictablevariancerealizedmarkovregret_tail generated sparse-loss realized-regret tail under four algebraic exploration contracts. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceArmExplorationScale","label":"sparseLossPredictableVarianceArmExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceArmExplorationScale","description":"Square-root component required by the pathwise sparse-loss base term.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-b8d13e646b67","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4378,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:429"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceArmExplorationScale (K S T : Real) : Real","missing":[],"search":"sparselosspredictablevariancearmexplorationscale banditrlproof.exp3.sparselosspredictablevariancearmexplorationscale square-root component required by the pathwise sparse-loss base term. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceMarkovExplorationScale","label":"sparseLossPredictableVarianceMarkovExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceMarkovExplorationScale","description":"Fifth-root component forced by the sparse-loss Markov variance threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-b380179968cd","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4379,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:434"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceMarkovExplorationScale (K S T delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancemarkovexplorationscale banditrlproof.exp3.sparselosspredictablevariancemarkovexplorationscale fifth-root component forced by the sparse-loss markov variance threshold. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceConfidenceExplorationScale","label":"sparseLossPredictableVarianceConfidenceExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceConfidenceExplorationScale","description":"Cube-root component required by both Bernstein confidence radii.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-886857a93975","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4380,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:440"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceConfidenceExplorationScale (K T delta : Real) : Real","missing":[],"search":"sparselosspredictablevarianceconfidenceexplorationscale banditrlproof.exp3.sparselosspredictablevarianceconfidenceexplorationscale cube-root component required by both bernstein confidence radii. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedExplorationScale","label":"sparseLossPredictableVarianceRealizedExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedExplorationScale","description":"Square-root component required by the bounded realized-deviation radius.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-aad60e66df72","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4381,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:445"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceRealizedExplorationScale (T delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancerealizedexplorationscale banditrlproof.exp3.sparselosspredictablevariancerealizedexplorationscale square-root component required by the bounded realized-deviation radius. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate","label":"sparseLossPredictableVarianceRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate","description":"Unclipped maximum of the four sparse-loss exploration scales.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-ef757e0bd771","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4382,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:452"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceRawExplorationRate (K S T delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancerawexplorationrate banditrlproof.exp3.sparselosspredictablevariancerawexplorationrate unclipped maximum of the four sparse-loss exploration scales. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate","label":"sparseLossPredictableVarianceClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate","description":"Explicit exploration schedule clipped into the Hedge stability regime.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-92d7cf92985b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4383,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:460"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceClippedExplorationRate (K S T delta : Real) : Real","missing":[],"search":"sparselosspredictablevarianceclippedexplorationrate banditrlproof.exp3.sparselosspredictablevarianceclippedexplorationrate explicit exploration schedule clipped into the hedge stability regime. definition compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.rpow_inv_five_le_half_of_thirtytwo_mul_le","label":"rpow_inv_five_le_half_of_thirtytwo_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.rpow_inv_five_le_half_of_thirtytwo_mul_le","description":"A nonnegative fifth-root scale is at most one half when its numerator is at most one thirty-second of its positive denominator.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-dc9ec2767ce3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4384,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:467"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem rpow_inv_five_le_half_of_thirtytwo_mul_le (numerator denominator : Real) (hnumerator : 0 <= numerator) (hdenominator : 0 < denominator) (hlarge : 32 * numerator <= denominator) : (numerator / denominator) ^ (5 : Real)⁻¹ <= 1 / 2","missing":[],"search":"rpow_inv_five_le_half_of_thirtytwo_mul_le banditrlproof.exp3.rpow_inv_five_le_half_of_thirtytwo_mul_le a nonnegative fifth-root scale is at most one half when its numerator is at most one thirty-second of its positive denominator. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.numerator_le_pow_five_mul_of_rpow_inv_five_le","label":"numerator_le_pow_five_mul_of_rpow_inv_five_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.numerator_le_pow_five_mul_of_rpow_inv_five_le","description":"If a fifth-root scale is below `gamma`, its numerator satisfies the corresponding fifth-power dominance contract.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-944f67fa5fed","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4385,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:485"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem numerator_le_pow_five_mul_of_rpow_inv_five_le (numerator denominator gamma : Real) (hnumerator : 0 <= numerator) (hdenominator : 0 < denominator) (hroot : (numerator / denominator) ^ (5 : Real)⁻¹ <= gamma) : numerator <= gamma ^ 5 * denominator","missing":[],"search":"numerator_le_pow_five_mul_of_rpow_inv_five_le banditrlproof.exp3.numerator_le_pow_five_mul_of_rpow_inv_five_le if a fifth-root scale is below `gamma`, its numerator satisfies the corresponding fifth-power dominance contract. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_le_half","label":"sparseLossPredictableVarianceClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_le_half","description":"theorem sparseLossPredictableVarianceClippedExplorationRate_le_half (K S T delta : Real) : sparseLossPredictableVarianceClippedExplorationRate K S T delta <= 1 / 2","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-ade75c5f07c1","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4386,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:506"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceClippedExplorationRate_le_half (K S T delta : Real) : sparseLossPredictableVarianceClippedExplorationRate K S T delta <= 1 / 2","missing":[],"search":"sparselosspredictablevarianceclippedexplorationrate_le_half banditrlproof.exp3.sparselosspredictablevarianceclippedexplorationrate_le_half theorem sparselosspredictablevarianceclippedexplorationrate_le_half (k s t delta : real) : sparselosspredictablevarianceclippedexplorationrate k s t delta <= 1 / 2 theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_eq_raw","label":"sparseLossPredictableVarianceClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_eq_raw","description":"theorem sparseLossPredictableVarianceClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : sparseLossPredictableVarianceRawExplorationRate K S T delta <= 1 / 2) : sparseLossPredictableVarianceClippedExplorationRate K S T delta = sparseLossPredictableVarianceRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-4072c1f71a43","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4387,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:512"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : sparseLossPredictableVarianceRawExplorationRate K S T delta <= 1 / 2) : sparseLossPredictableVarianceClippedExplorationRate K S T delta = sparseLossPredictableVarianceRawExplorationRate K S T delta","missing":[],"search":"sparselosspredictablevarianceclippedexplorationrate_eq_raw banditrlproof.exp3.sparselosspredictablevarianceclippedexplorationrate_eq_raw theorem sparselosspredictablevarianceclippedexplorationrate_eq_raw (k s t delta : real) (hraw : sparselosspredictablevariancerawexplorationrate k s t delta <= 1 / 2) : sparselosspredictablevarianceclippedexplorationrate k s t delta = sparselosspredictablevariancerawexplorationrate k s t delta theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate_pos","label":"sparseLossPredictableVarianceRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate_pos","description":"theorem sparseLossPredictableVarianceRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < sparseLossPredictableVarianceRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-d0a7706e765d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4388,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:520"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < sparseLossPredictableVarianceRawExplorationRate K S T delta","missing":[],"search":"sparselosspredictablevariancerawexplorationrate_pos banditrlproof.exp3.sparselosspredictablevariancerawexplorationrate_pos theorem sparselosspredictablevariancerawexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < sparselosspredictablevariancerawexplorationrate k s t delta theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_pos","label":"sparseLossPredictableVarianceClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_pos","description":"theorem sparseLossPredictableVarianceClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < sparseLossPredictableVarianceClippedExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-ee7b97f25a98","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4389,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:530"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < sparseLossPredictableVarianceClippedExplorationRate K S T delta","missing":[],"search":"sparselosspredictablevarianceclippedexplorationrate_pos banditrlproof.exp3.sparselosspredictablevarianceclippedexplorationrate_pos theorem sparselosspredictablevarianceclippedexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < sparselosspredictablevarianceclippedexplorationrate k s t delta theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","label":"sparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","description":"Four transparent horizon contracts ensure every raw schedule component is at most one half, so clipping is inactive.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-b66217db9b1d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4390,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:540"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (5 * K * S * Real.log K ^ 2 * Real.log (5 / delta)) <= delta * T ^ 3) (hlarge_confidence : 8 * (K * Real.log (5 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= T) : sparseLossPredictableVarianceRawExplorationRate K S T delta <= 1 / 2","missing":[],"search":"sparselosspredictablevariancerawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.sparselosspredictablevariancerawexplorationrate_le_half_of_horizon_contracts four transparent horizon contracts ensure every raw schedule component is at most one half, so clipping is inactive. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_contracts","label":"sparseLossPredictableVarianceClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_contracts","description":"The clipped maximum satisfies exactly the four contracts consumed by the gamma-characterized sparse-loss theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-77a172a8259d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4391,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:582"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceClippedExplorationRate_contracts (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (5 * K * S * Real.log K ^ 2 * Real.log (5 / delta)) <= delta * T ^ 3) (hlarge_confidence : 8 * (K * Real.log (5 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= T) : let gamma := sparseLossPredictableVarianceClippedExplorationRate K S T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ S * Real.log K <= gamma ^ 2 * T ∧ 5 * K * S * Real.log K ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * T ^ 3 ∧ K * Real.log (5 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= gamma ^ 2 * T","missing":[],"search":"sparselosspredictablevarianceclippedexplorationrate_contracts banditrlproof.exp3.sparselosspredictablevarianceclippedexplorationrate_contracts the clipped maximum satisfies exactly the four contracts consumed by the gamma-characterized sparse-loss theorem. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_explicitSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Fully explicit generated sparse-loss realized-regret tail for the clipped maximum of the four exploration scales.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovexplicittuning/index.html#decl-f546d895d095","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","order":4392,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning.lean:688"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * (5 * (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta)) <= delta…","missing":[],"search":"sampledpredictable_explicitsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_explicitsparselosspredictablevariancerealizedmarkovregret_tail fully explicit generated sparse-loss realized-regret tail for the clipped maximum of the four exploration scales. theorem compiled","shard":"modules/a2135a1b34dce681.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossSupport","label":"sampledPredictableLossSupport","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossSupport","description":"Nonzero predictable-loss coordinates among the active arms at one time.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-0a61a0dfe20b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4393,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:29"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableLossSupport {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : Finset Action","missing":[],"search":"sampledpredictablelosssupport banditrlproof.exp3.sampledpredictablelosssupport nonzero predictable-loss coordinates among the active arms at one time. definition compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassAt_le_supportCard","label":"sampledPredictableLossMassAt_le_supportCard","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossMassAt_le_supportCard","description":"The predictable loss mass at one time is at most the number of active nonzero loss coordinates.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-24a036d4e7a0","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4394,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:39"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossMassAt_le_supportCard {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : arms.sum (fun action => predictableLossAt loss t sample action) ≤ (sampledPredictableLossSupport arms loss t sample).card","missing":[],"search":"sampledpredictablelossmassat_le_supportcard banditrlproof.exp3.sampledpredictablelossmassat_le_supportcard the predictable loss mass at one time is at most the number of active nonzero loss coordinates. theorem compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_sparsity_mul_horizon_of_sample","label":"sampledPredictableLossMassSum_le_sparsity_mul_horizon_of_sample","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossMassSum_le_sparsity_mul_horizon_of_sample","description":"A per-round support-cardinality bound for one trajectory supplies its armwise loss-mass budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-36bddf95ce38","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4395,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:69"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossMassSum_le_sparsity_mul_horizon_of_sample {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (sample : Env × ((k : Nat) → Action × Real)) (hsparse : ∀ t, t < horizon → (sampledPredictableLossSupport arms loss t sample).card ≤ sparsity) : sampledPredictableLossMassSum arms loss horizon sample ≤ (sparsity : Real) * (horizon : Real)","missing":[],"search":"sampledpredictablelossmasssum_le_sparsity_mul_horizon_of_sample banditrlproof.exp3.sampledpredictablelossmasssum_le_sparsity_mul_horizon_of_sample a per-round support-cardinality bound for one trajectory supplies its armwise loss-mass budget. theorem compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_sparsity_mul_horizon","label":"sampledPredictableLossMassSum_le_sparsity_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossMassSum_le_sparsity_mul_horizon","description":"A uniform per-round support-cardinality bound supplies the pathwise armwise loss-mass budget used by the small-loss theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-63bdc0828fe3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4396,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:94"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossMassSum_le_sparsity_mul_horizon {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hsparse : ∀ sample t, t < horizon → (sampledPredictableLossSupport arms loss t sample).card ≤ sparsity) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableLossMassSum arms loss horizon sample ≤ (sparsity : Real) * (horizon : Real)","missing":[],"search":"sampledpredictablelossmasssum_le_sparsity_mul_horizon banditrlproof.exp3.sampledpredictablelossmasssum_le_sparsity_mul_horizon a uniform per-round support-cardinality bound supplies the pathwise armwise loss-mass budget used by the small-loss theorem. theorem compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget","description":"Sparse-loss specialization of the five-event realized regret budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-1571a946cdb0","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4397,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:108"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablevariancesquaresparselossrealizedmarkovhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquaresparselossrealizedmarkovhighprobabilityregretbudget sparse-loss specialization of the five-event realized regret budget. definition compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta_of_ae_sparsity","label":"sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta_of_ae_sparsity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta_of_ae_sparsity","description":"Realized predictable-variance EXP3 regret when the per-round nonzero-loss support-cardinality bound holds almost everywhere under the exact generated trajectory measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-4082d1ce31e7","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4398,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:118"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta_of_ae_sparsity {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hsparse : ∀ᵐ sample ∂(prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment), ∀ t,…","missing":[],"search":"sampledpredictable_predictablevariancesquaresparselossrealizedmarkovhighprobabilityregret_tail_total_delta_of_ae_sparsity banditrlproof.exp3.sampledpredictable_predictablevariancesquaresparselossrealizedmarkovhighprobabilityregret_tail_total_delta_of_ae_sparsity realized predictable-variance exp3 regret when the per-round nonzero-loss support-cardinality bound holds almost everywhere under the exact generated trajectory measure. theorem compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta","description":"Backward-compatible pathwise sparse-loss wrapper. The stronger universal contract is converted to the generated-measure almost-everywhere contract used by the primary theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovhighprobabilityregret/index.html#decl-5f0c41d2f02e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","order":4399,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret.lean:169"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hsparse : ∀ sample t, t < horizon → (sampledPredictableLossSupport arms loss t sample).card ≤ sparsity) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKern…","missing":[],"search":"sampledpredictable_predictablevariancesquaresparselossrealizedmarkovhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_predictablevariancesquaresparselossrealizedmarkovhighprobabilityregret_tail_total_delta backward-compatible pathwise sparse-loss wrapper. the stronger universal contract is converted to the generated-measure almost-everywhere contract used by the primary theorem. theorem compiled","shard":"modules/9f2f32b15d579ccb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSparsityFailure","label":"sampledPredictableSparsityFailure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSparsityFailure","description":"Generated trajectories on which the requested per-round support cap fails at some time before the horizon.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-27ddece43c65","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4400,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableSparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) : Set (Env × ((k : Nat) → Action × Real))","missing":[],"search":"sampledpredictablesparsityfailure banditrlproof.exp3.sampledpredictablesparsityfailure generated trajectories on which the requested per-round support cap fails at some time before the horizon. definition compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_or_mem_sparsityFailure","label":"sampledPredictableLossMassSum_le_or_mem_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossMassSum_le_or_mem_sparsityFailure","description":"Every trajectory either obeys the sparse armwise loss-mass budget or lies in the explicit sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-9466d3739793","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4401,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:37"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossMassSum_le_or_mem_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableLossMassSum arms loss horizon sample ≤ (sparsity : Real) * (horizon : Real) ∨ sample ∈ sampledPredictableSparsityFailure arms loss horizon sparsity","missing":[],"search":"sampledpredictablelossmasssum_le_or_mem_sparsityfailure banditrlproof.exp3.sampledpredictablelossmasssum_le_or_mem_sparsityfailure every trajectory either obeys the sparse armwise loss-mass budget or lies in the explicit sparsity-failure event. theorem compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_card_mul_horizon","label":"sampledPredictableLossMassSum_le_card_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableLossMassSum_le_card_mul_horizon","description":"The number of nonzero active coordinates is always at most the number of active arms, giving a global pathwise loss-mass envelope.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-0672b2e4cf1c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4402,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:59"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableLossMassSum_le_card_mul_horizon {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableLossMassSum arms loss horizon sample ≤ (arms.card : Real) * (horizon : Real)","missing":[],"search":"sampledpredictablelossmasssum_le_card_mul_horizon banditrlproof.exp3.sampledpredictablelossmasssum_le_card_mul_horizon the number of nonzero active coordinates is always at most the number of active arms, giving a global pathwise loss-mass envelope. theorem compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableGlobalVarianceMeanBudget","label":"sampledPredictableGlobalVarianceMeanBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableGlobalVarianceMeanBudget","description":"Global predictable-variance mean envelope used when sparsity may fail with positive probability.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-19cbb4141ba0","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4403,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:73"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableGlobalVarianceMeanBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon : Nat) : Real","missing":[],"search":"sampledpredictableglobalvariancemeanbudget banditrlproof.exp3.sampledpredictableglobalvariancemeanbudget global predictable-variance mean envelope used when sparsity may fail with positive probability. definition compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_globalLossMass","label":"sampledPredictableMixedSquaredVarianceLIntegral_le_globalLossMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_globalLossMass","description":"The global `arms.card * horizon` loss-mass envelope closes the cumulative predictable-variance `lintegral` without any sparsity assumption.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-3b1fc4351186","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4404,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:81"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceLIntegral_le_globalLossMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : sampledPredictableMixedSquaredVarianceLIntegral mu arms eta gamma loss horizon ≤ ENNReal.ofReal (sampledPredictableGlobalVarianceMeanBudget arms gamma horizon)","missing":[],"search":"sampledpredictablemixedsquaredvariancelintegral_le_globallossmass banditrlproof.exp3.sampledpredictablemixedsquaredvariancelintegral_le_globallossmass the global `arms.card * horizon` loss-mass envelope closes the cumulative predictable-variance `lintegral` without any sparsity assumption. theorem compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget","description":"Realized sparse-loss budget with a global Markov variance envelope. The observed-square mean remains `sparsity * horizon`, while the variance overflow threshold uses `arms.card * horizon`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-9def9076ad98","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4405,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:107"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregretbudget realized sparse-loss budget with a global markov variance envelope. the observed-square mean remains `sparsity * horizon`, while the variance overflow threshold uses `arms.card * horizon`. definition compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail","label":"sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail","description":"Generated realized regret with an explicit positive-probability sparsity failure event. The total tail is the ordinary five-event budget plus the exact generated measure of that event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-c2bdaf154cc4","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4406,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:119"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableVarianceSquareProbabilisticS…","missing":[],"search":"sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregret_tail generated realized regret with an explicit positive-probability sparsity failure event. the total tail is the ordinary five-event budget plus the exact generated measure of that event. theorem compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail_of_sparsityFailure_le","description":"Practical `delta + epsilon` consumer of the explicit sparsity-failure residual theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilisticsparsity/index.html#decl-27fd05d65d9c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","order":4407,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity.lean:272"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (delta epsilon : Real) (hdelta : 0 < delta) (hfailure : (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment) (sampledPredictableSparsity…","missing":[],"search":"sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregret_tail_of_sparsityfailure_le practical `delta + epsilon` consumer of the explicit sparsity-failure residual theorem. theorem compiled","shard":"modules/f101f1a70bd37105.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceLargeHorizonCondition","label":"probabilisticSparseLossPredictableVarianceLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceLargeHorizonCondition","description":"Regime in which all four components of the probabilistic-sparsity predictable-variance exploration schedule are at most one half.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-306e3ec30e1f/index.html#decl-124d7b40351f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","order":4408,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def probabilisticSparseLossPredictableVarianceLargeHorizonCondition (K S T delta : Real) : Prop","missing":[],"search":"probabilisticsparselosspredictablevariancelargehorizoncondition banditrlproof.exp3.probabilisticsparselosspredictablevariancelargehorizoncondition regime in which all four components of the probabilistic-sparsity predictable-variance exploration schedule are at most one half. definition compiled","shard":"modules/62f8e0413cc173dd.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceAllHorizonRegretThreshold","label":"probabilisticSparseLossPredictableVarianceAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceAllHorizonRegretThreshold","description":"All-horizon threshold for the probabilistic-sparsity route: use the explicit large-horizon threshold in its valid regime and `T + 1` otherwise.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-306e3ec30e1f/index.html#decl-3e0f56607d5b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","order":4409,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon.lean:42"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevarianceallhorizonregretthreshold banditrlproof.exp3.probabilisticsparselosspredictablevarianceallhorizonregretthreshold all-horizon threshold for the probabilistic-sparsity route: use the explicit large-horizon threshold in its valid regime and `t + 1` otherwise. definition compiled","shard":"modules/62f8e0413cc173dd.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Generated all-horizon realized-regret theorem with the exact support-sparsity failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-306e3ec30e1f/index.html#decl-6333dc1f3b9f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","order":4410,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon.lean:58"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := probabilisticSparseLossPredictableVarianceClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := probabilisticSparseLossPredictableVarianceHighProbabili…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail generated all-horizon realized-regret theorem with the exact support-sparsity failure residual. theorem compiled","shard":"modules/62f8e0413cc173dd.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","description":"Practical all-horizon `delta + epsilon` theorem under an exact bound on the support-sparsity failure event for the same internally tuned measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-306e3ec30e1f/index.html#decl-2140dd14d376","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","order":4411,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon.lean:140"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := probabilisticSparseLossPredictableVarianceClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := probabilisticSparseLossPr…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le practical all-horizon `delta + epsilon` theorem under an exact bound on the support-sparsity failure event for the same internally tuned measure. theorem compiled","shard":"modules/62f8e0413cc173dd.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget_eq","label":"probabilisticSparseLossPredictableVarianceBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget_eq","description":"Closed form of the global-envelope Markov variance threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-c9d3beccc688","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4412,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceBudget_eq {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : probabilisticSparseLossPredictableVarianceBudget arms gamma horizon delta = 5 * (arms.card : Real) ^ 2 * (horizon : Real) / (gamma * delta)","missing":[],"search":"probabilisticsparselosspredictablevariancebudget_eq banditrlproof.exp3.probabilisticsparselosspredictablevariancebudget_eq closed form of the global-envelope markov variance threshold. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_mul_probabilisticSparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","label":"log_mul_probabilisticSparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_mul_probabilisticSparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","description":"Under the sparse-base, global-envelope fifth-power, and cubic confidence contracts, the log-weighted mixed-square radius is at most `3 * gamma^2 * T^2`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-d750cc92107b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4413,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:48"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_mul_probabilisticSparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (5 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.log (arms.card : Real) * sampledMixedSquaredPredictableVarianceRadius arms gamma (probabilisticSparseLossPredictableVarianceBudget arms gamma horizon delta) (delta / 5) <…","missing":[],"search":"log_mul_probabilisticsparselosspredictablevarianceradius_le_three_mul_sq_mul_horizon_sq banditrlproof.exp3.log_mul_probabilisticsparselosspredictablevarianceradius_le_three_mul_sq_mul_horizon_sq under the sparse-base, global-envelope fifth-power, and cubic confidence contracts, the log-weighted mixed-square radius is at most `3 * gamma^2 * t^2`. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","label":"probabilisticSparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","description":"The sparse base and global-envelope predictable-variance radius make the learning-rate-balanced square root at most `2 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-e543dea2e41f","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4414,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:172"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (5 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.sqrt (Real.log (arms.card : Real) * probabilisticSparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta) <= 2 * gamma * (horizon : Real)","missing":[],"search":"probabilisticsparselosspredictablevariancebalancedsqrt_le_two_mul_gamma_mul_horizon banditrlproof.exp3.probabilisticsparselosspredictablevariancebalancedsqrt_le_two_mul_gamma_mul_horizon the sparse base and global-envelope predictable-variance radius make the learning-rate-balanced square root at most `2 * gamma * t`. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovExplicitThreshold","label":"probabilisticSparseLossPredictableVarianceRealizedMarkovExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovExplicitThreshold","description":"Explicit probabilistic-sparsity realized-regret threshold after tuning both eta and gamma.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-081ec6d40308","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4415,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:240"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceRealizedMarkovExplicitThreshold {Action : Type v} (_arms : Finset Action) (gamma : Real) (horizon _sparsity : Nat) (_delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancerealizedmarkovexplicitthreshold banditrlproof.exp3.probabilisticsparselosspredictablevariancerealizedmarkovexplicitthreshold explicit probabilistic-sparsity realized-regret threshold after tuning both eta and gamma. definition compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","label":"probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","description":"The four algebraic exploration contracts reduce the probabilistic- sparsity eta-tuned threshold to `14 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-675bec8a1e4e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4416,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:247"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (5 / delta) <= gamma ^ 3 * (horizon : Real)) (hrealized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= gamma ^ 2 * (horizon : Real)) : probabilisticSparseLossPredictableVarianceReali…","missing":[],"search":"probabilisticsparselosspredictablevariancerealizedmarkovtunedthreshold_le_explicitthreshold banditrlproof.exp3.probabilisticsparselosspredictablevariancerealizedmarkovtunedthreshold_le_explicitthreshold the four algebraic exploration contracts reduce the probabilistic- sparsity eta-tuned threshold to `14 * gamma * t`. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Generated probabilistic-sparsity regret under four algebraic exploration contracts, retaining the exact sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-dbf2afd2f156","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4417,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:344"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : 5 * (arms.card : Real) ^…","missing":[],"search":"sampledpredictable_gammacharacterizedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail generated probabilistic-sparsity regret under four algebraic exploration contracts, retaining the exact sparsity-failure residual. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","description":"Practical gamma-characterized theorem under an exact generated-measure bound on the sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-b671bc199dcf","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4418,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:434"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmi…","missing":[],"search":"sampledpredictable_gammacharacterizedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_gammacharacterizedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le practical gamma-characterized theorem under an exact generated-measure bound on the sparsity-failure event. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceMarkovExplorationScale","label":"probabilisticSparseLossPredictableVarianceMarkovExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceMarkovExplorationScale","description":"Fifth-root component forced by the global `K * T` Markov envelope.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-b7535d292c03","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4419,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:490"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceMarkovExplorationScale (K T delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancemarkovexplorationscale banditrlproof.exp3.probabilisticsparselosspredictablevariancemarkovexplorationscale fifth-root component forced by the global `k * t` markov envelope. definition compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate","label":"probabilisticSparseLossPredictableVarianceRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate","description":"Unclipped maximum of the sparse base, global Markov, Bernstein, and realized-deviation exploration scales.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-1ea46b24099d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4420,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:497"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceRawExplorationRate (K S T delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancerawexplorationrate banditrlproof.exp3.probabilisticsparselosspredictablevariancerawexplorationrate unclipped maximum of the sparse base, global markov, bernstein, and realized-deviation exploration scales. definition compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate","label":"probabilisticSparseLossPredictableVarianceClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate","description":"Explicit probabilistic-sparsity exploration schedule clipped into the Hedge stability regime.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-e6b7ac78956a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4421,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:509"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceClippedExplorationRate (K S T delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevarianceclippedexplorationrate banditrlproof.exp3.probabilisticsparselosspredictablevarianceclippedexplorationrate explicit probabilistic-sparsity exploration schedule clipped into the hedge stability regime. definition compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_le_half","label":"probabilisticSparseLossPredictableVarianceClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_le_half","description":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_le_half (K S T delta : Real) : probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta <= 1 / 2","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-fa62299bf61c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4422,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:515"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_le_half (K S T delta : Real) : probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta <= 1 / 2","missing":[],"search":"probabilisticsparselosspredictablevarianceclippedexplorationrate_le_half banditrlproof.exp3.probabilisticsparselosspredictablevarianceclippedexplorationrate_le_half theorem probabilisticsparselosspredictablevarianceclippedexplorationrate_le_half (k s t delta : real) : probabilisticsparselosspredictablevarianceclippedexplorationrate k s t delta <= 1 / 2 theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_eq_raw","label":"probabilisticSparseLossPredictableVarianceClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_eq_raw","description":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta <= 1 / 2) : probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta = probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-69db913cbc1a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4423,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:522"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta <= 1 / 2) : probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta = probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta","missing":[],"search":"probabilisticsparselosspredictablevarianceclippedexplorationrate_eq_raw banditrlproof.exp3.probabilisticsparselosspredictablevarianceclippedexplorationrate_eq_raw theorem probabilisticsparselosspredictablevarianceclippedexplorationrate_eq_raw (k s t delta : real) (hraw : probabilisticsparselosspredictablevariancerawexplorationrate k s t delta <= 1 / 2) : probabilisticsparselosspredictablevarianceclippedexplorationrate k s t delta = probabilisticsparselosspredictablevariancerawexplorationrate k s t delta theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate_pos","label":"probabilisticSparseLossPredictableVarianceRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate_pos","description":"theorem probabilisticSparseLossPredictableVarianceRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-34e0b8b68e57","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4424,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:534"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta","missing":[],"search":"probabilisticsparselosspredictablevariancerawexplorationrate_pos banditrlproof.exp3.probabilisticsparselosspredictablevariancerawexplorationrate_pos theorem probabilisticsparselosspredictablevariancerawexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < probabilisticsparselosspredictablevariancerawexplorationrate k s t delta theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_pos","label":"probabilisticSparseLossPredictableVarianceClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_pos","description":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-8ac103df4e55","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4425,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:546"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta","missing":[],"search":"probabilisticsparselosspredictablevarianceclippedexplorationrate_pos banditrlproof.exp3.probabilisticsparselosspredictablevarianceclippedexplorationrate_pos theorem probabilisticsparselosspredictablevarianceclippedexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < probabilisticsparselosspredictablevarianceclippedexplorationrate k s t delta theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","label":"probabilisticSparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","description":"Transparent horizon contracts ensure every raw schedule component is at most one half, so clipping is inactive.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-aa2ceae98709","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4426,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:559"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts (K S T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (5 * K ^ 2 * Real.log K ^ 2 * Real.log (5 / delta)) <= delta * T ^ 3) (hlarge_confidence : 8 * (K * Real.log (5 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= T) : probabilisticSparseLossPredictableVarianceRawExplorationRate K S T delta <= 1 / 2","missing":[],"search":"probabilisticsparselosspredictablevariancerawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.probabilisticsparselosspredictablevariancerawexplorationrate_le_half_of_horizon_contracts transparent horizon contracts ensure every raw schedule component is at most one half, so clipping is inactive. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_contracts","label":"probabilisticSparseLossPredictableVarianceClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_contracts","description":"The clipped maximum supplies exactly the four contracts consumed by the gamma-characterized probabilistic-sparsity theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-e99bc0c28ceb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4427,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:602"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceClippedExplorationRate_contracts (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (5 * K ^ 2 * Real.log K ^ 2 * Real.log (5 / delta)) <= delta * T ^ 3) (hlarge_confidence : 8 * (K * Real.log (5 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= T) : let gamma := probabilisticSparseLossPredictableVarianceClippedExplorationRate K S T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ S * Real.log K <= gamma ^ 2 * T ∧ 5 * K ^ 2 * Real.log K ^ 2 * Real.log (5 / delta) <= gamma ^ 5 * delta * T ^ 3 ∧ K * Real.log (5 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (5 / delta) <= gamma ^ 2 * T","missing":[],"search":"probabilisticsparselosspredictablevarianceclippedexplorationrate_contracts banditrlproof.exp3.probabilisticsparselosspredictablevarianceclippedexplorationrate_contracts the clipped maximum supplies exactly the four contracts consumed by the gamma-characterized probabilistic-sparsity theorem. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Fully explicit generated probabilistic-sparsity regret tail for the clipped maximum schedule, retaining the exact failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-01df475fd6a6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4428,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:728"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * (5 * (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real.log (5 / delta)) <= delta * (…","missing":[],"search":"sampledpredictable_explicitprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_explicitprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail fully explicit generated probabilistic-sparsity regret tail for the clipped maximum schedule, retaining the exact failure residual. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","description":"Fully explicit practical `delta + epsilon` theorem under an exact bound on the sparsity-failure event for the internally eta/gamma-tuned measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-683bfad2c843/index.html#decl-d90f34767cd7","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","order":4429,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning.lean:805"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * (5 * (arms.card : Real) ^ 2 * Real.log (arms.card : Real) ^ 2 * Real…","missing":[],"search":"sampledpredictable_explicitprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_explicitprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le fully explicit practical `delta + epsilon` theorem under an exact bound on the sparsity-failure event for the internally eta/gamma-tuned measure. theorem compiled","shard":"modules/71993992e89b3475.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget","label":"probabilisticSparseLossPredictableVarianceBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget","description":"Markov threshold for cumulative predictable mixed-square variance when positive-probability sparsity failures require the global `K * T` envelope.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-8c0fb77c9620","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4430,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:29"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancebudget banditrlproof.exp3.probabilisticsparselosspredictablevariancebudget markov threshold for cumulative predictable mixed-square variance when positive-probability sparsity failures require the global `k * t` envelope. definition compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget_pos","label":"probabilisticSparseLossPredictableVarianceBudget_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget_pos","description":"theorem probabilisticSparseLossPredictableVarianceBudget_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : 0 < probabilisticSparseLossPredictableVarianceBudget arms gamma horizon delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-a2020509542b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4431,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceBudget_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : 0 < probabilisticSparseLossPredictableVarianceBudget arms gamma horizon delta","missing":[],"search":"probabilisticsparselosspredictablevariancebudget_pos banditrlproof.exp3.probabilisticsparselosspredictablevariancebudget_pos theorem probabilisticsparselosspredictablevariancebudget_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) (hdelta : 0 < delta) : 0 < probabilisticsparselosspredictablevariancebudget arms gamma horizon delta theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityScale","label":"probabilisticSparseLossPredictableVarianceHighProbabilityScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityScale","description":"Complete Hedge scale with sparse observed-square mean and global Markov variance control.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-3c9f87b9ac85","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4432,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:57"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceHighProbabilityScale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancehighprobabilityscale banditrlproof.exp3.probabilisticsparselosspredictablevariancehighprobabilityscale complete hedge scale with sparse observed-square mean and global markov variance control. definition compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityScale_pos","label":"probabilisticSparseLossPredictableVarianceHighProbabilityScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityScale_pos","description":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityScale_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < probabilisticSparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-ddf74e4623b3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4433,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:67"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityScale_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < probabilisticSparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta","missing":[],"search":"probabilisticsparselosspredictablevariancehighprobabilityscale_pos banditrlproof.exp3.probabilisticsparselosspredictablevariancehighprobabilityscale_pos theorem probabilisticsparselosspredictablevariancehighprobabilityscale_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) : 0 < probabilisticsparselosspredictablevariancehighprobabilityscale arms gamma horizon sparsity delta theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate","label":"probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate","description":"Learning rate balancing entropy against the probabilistic-sparsity scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-75e0b01e1891","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4434,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:96"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancehighprobabilitylearningrate banditrlproof.exp3.probabilisticsparselosspredictablevariancehighprobabilitylearningrate learning rate balancing entropy against the probabilistic-sparsity scale. definition compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_pos","label":"probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_pos","description":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-c822c3d151e2","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4435,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:105"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta","missing":[],"search":"probabilisticsparselosspredictablevariancehighprobabilitylearningrate_pos banditrlproof.exp3.probabilisticsparselosspredictablevariancehighprobabilitylearningrate_pos theorem probabilisticsparselosspredictablevariancehighprobabilitylearningrate_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) : 0 < probabilisticsparselosspredictablevariancehighprobabilitylearningrate arms gamma horizon sparsity delta theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","label":"probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","description":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-9859c0869cb1","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4436,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:121"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta ^ 2 * probabilisticSparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta = Real.log (arms.card : Real)","missing":[],"search":"probabilisticsparselosspredictablevariancehighprobabilitylearningrate_sq_mul_scale banditrlproof.exp3.probabilisticsparselosspredictablevariancehighprobabilitylearningrate_sq_mul_scale theorem probabilisticsparselosspredictablevariancehighprobabilitylearningrate_sq_mul_scale {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) : probabilisticsparselosspredictablevariancehighprobabilitylearningrate arms gamma horizon sparsity delta ^ 2 * probabilisticsparselosspredictablevariancehighprobabilityscale arms gamma horizon sparsity delta = real.log (arms.card : real) theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","label":"probabilisticSparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","description":"Under `gamma ≤ 1/2`, entropy and stability cost at most three copies of the balanced probabilistic-sparsity scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-cfaf24632e04","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4437,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:144"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem probabilisticSparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : Real.log (arms.card : Real) / probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta + (probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta * (1 / (1 - gamma))) * probabilisticSparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta ≤ 3 * Real.sqrt (Real.log (arms.card : Real) * probabilisticSparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta)","missing":[],"search":"probabilisticsparselosspredictablevariancehighprobabilityhedgebudget_le_three_mul_sqrt banditrlproof.exp3.probabilisticsparselosspredictablevariancehighprobabilityhedgebudget_le_three_mul_sqrt under `gamma ≤ 1/2`, entropy and stability cost at most three copies of the balanced probabilistic-sparsity scale. theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold","label":"probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold","description":"Tuned regret threshold retaining the global Markov variance radius and the three non-Hedge confidence terms.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-35d0d1714481","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4438,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:224"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"probabilisticsparselosspredictablevariancerealizedmarkovtunedthreshold banditrlproof.exp3.probabilisticsparselosspredictablevariancerealizedmarkovtunedthreshold tuned regret threshold retaining the global markov variance radius and the three non-hedge confidence terms. definition compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","description":"The raw probabilistic-sparsity budget is bounded by the eta-tuned threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-5468bb7932f9","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4439,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:241"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget arms (probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta) gamma horizon sparsity delta ≤ probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold arms gamma horizon sparsity delta","missing":[],"search":"sampledpredictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictablevariancesquareprobabilisticsparselossrealizedmarkovhighprobabilityregretbudget_le_tunedthreshold the raw probabilistic-sparsity budget is bounded by the eta-tuned threshold. theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Eta-tuned generated regret with the exact sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-525958583429","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4440,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:270"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let eta := probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImportanceWeightedTraject…","missing":[],"search":"sampledpredictable_tunedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_tunedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail eta-tuned generated regret with the exact sparsity-failure residual. theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","description":"Practical eta-tuned `delta + epsilon` theorem under an exact generated- measure bound on the sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovprobabilist-7403bec8a860/index.html#decl-2d9b58f71baa","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","order":4441,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning.lean:357"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) : let eta := probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sa…","missing":[],"search":"sampledpredictable_tunedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_tunedprobabilisticsparselosspredictablevariancerealizedmarkovregret_tail_of_sparsityfailure_le practical eta-tuned `delta + epsilon` theorem under an exact generated- measure bound on the sparsity-failure event. theorem compiled","shard":"modules/f5373e7fc179aad6.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBudget","label":"sparseLossPredictableVarianceBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceBudget","description":"Markov threshold for cumulative predictable mixed-square variance under the sparse armwise loss-mass budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-04ec5a61328a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4442,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:32"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceBudget {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancebudget banditrlproof.exp3.sparselosspredictablevariancebudget markov threshold for cumulative predictable mixed-square variance under the sparse armwise loss-mass budget. definition compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBudget_pos","label":"sparseLossPredictableVarianceBudget_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceBudget_pos","description":"theorem sparseLossPredictableVarianceBudget_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : 0 < sparseLossPredictableVarianceBudget arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-9c15b821c176","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4443,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:38"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceBudget_pos {Action : Type v} (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : 0 < sparseLossPredictableVarianceBudget arms gamma horizon sparsity delta","missing":[],"search":"sparselosspredictablevariancebudget_pos banditrlproof.exp3.sparselosspredictablevariancebudget_pos theorem sparselosspredictablevariancebudget_pos {action : type v} (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) (hdelta : 0 < delta) : 0 < sparselosspredictablevariancebudget arms gamma horizon sparsity delta theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityScale","label":"sparseLossPredictableVarianceHighProbabilityScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityScale","description":"Complete sparse-loss Hedge scale at the public five-event allocation.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-adfcc09a1945","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4444,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:56"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceHighProbabilityScale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancehighprobabilityscale banditrlproof.exp3.sparselosspredictablevariancehighprobabilityscale complete sparse-loss hedge scale at the public five-event allocation. definition compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityScale_pos","label":"sparseLossPredictableVarianceHighProbabilityScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityScale_pos","description":"theorem sparseLossPredictableVarianceHighProbabilityScale_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : 0 < sparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-89f655bb993d","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4445,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:66"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceHighProbabilityScale_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : 0 < sparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta","missing":[],"search":"sparselosspredictablevariancehighprobabilityscale_pos banditrlproof.exp3.sparselosspredictablevariancehighprobabilityscale_pos theorem sparselosspredictablevariancehighprobabilityscale_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) (hdelta : 0 < delta) : 0 < sparselosspredictablevariancehighprobabilityscale arms gamma horizon sparsity delta theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate","label":"sparseLossPredictableVarianceHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate","description":"Learning rate balancing entropy against the complete sparse-loss predictable-variance scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-c3bf715742a4","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4446,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:100"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceHighProbabilityLearningRate {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancehighprobabilitylearningrate banditrlproof.exp3.sparselosspredictablevariancehighprobabilitylearningrate learning rate balancing entropy against the complete sparse-loss predictable-variance scale. definition compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate_pos","label":"sparseLossPredictableVarianceHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate_pos","description":"theorem sparseLossPredictableVarianceHighProbabilityLearningRate_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : 0 < sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-57c5c41754a6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4447,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:109"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceHighProbabilityLearningRate_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : 0 < sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta","missing":[],"search":"sparselosspredictablevariancehighprobabilitylearningrate_pos banditrlproof.exp3.sparselosspredictablevariancehighprobabilitylearningrate_pos theorem sparselosspredictablevariancehighprobabilitylearningrate_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) (hdelta : 0 < delta) : 0 < sparselosspredictablevariancehighprobabilitylearningrate arms gamma horizon sparsity delta theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","label":"sparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","description":"theorem sparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta ^ 2 *…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-9b9d6571b5c8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4448,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:126"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta ^ 2 * sparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta = Real.log (arms.card : Real)","missing":[],"search":"sparselosspredictablevariancehighprobabilitylearningrate_sq_mul_scale banditrlproof.exp3.sparselosspredictablevariancehighprobabilitylearningrate_sq_mul_scale theorem sparselosspredictablevariancehighprobabilitylearningrate_sq_mul_scale {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) (hdelta : 0 < delta) : sparselosspredictablevariancehighprobabilitylearningrate arms gamma horizon sparsity delta ^ 2 * sparselosspredictablevariancehighprobabilityscale arms gamma horizon sparsity delta = real.log (arms.card : real) theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","label":"sparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","description":"With `gamma <= 1/2`, entropy and sparse-loss stability cost at most three copies of their balanced square-root scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-1e5b229ea8a7","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4449,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:150"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : Real.log (arms.card : Real) / sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta + (sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta * (1 / (1 - gamma))) * sparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta ≤ 3 * Real.sqrt (Real.log (arms.card : Real) * sparseLossPredictableVarianceHighProbabilityScale arms gamma horizon sparsity delta)","missing":[],"search":"sparselosspredictablevariancehighprobabilityhedgebudget_le_three_mul_sqrt banditrlproof.exp3.sparselosspredictablevariancehighprobabilityhedgebudget_le_three_mul_sqrt with `gamma <= 1/2`, entropy and sparse-loss stability cost at most three copies of their balanced square-root scale. theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovTunedThreshold","label":"sparseLossPredictableVarianceRealizedMarkovTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovTunedThreshold","description":"Explicit sparse-loss threshold after tuning every learning-rate-dependent term. Gamma and the three non-square confidence contributions remain visible.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-422fb74d462c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4450,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:229"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sparseLossPredictableVarianceRealizedMarkovTunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sparselosspredictablevariancerealizedmarkovtunedthreshold banditrlproof.exp3.sparselosspredictablevariancerealizedmarkovtunedthreshold explicit sparse-loss threshold after tuning every learning-rate-dependent term. gamma and the three non-square confidence contributions remain visible. definition compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","description":"The complete sparse-loss Markov budget is bounded by the eta-tuned threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-78bbfbfa7a8a","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4451,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:246"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget arms (sparseLossPredictableVarianceHighProbabilityLearningRate arms gamma horizon sparsity delta) gamma horizon sparsity delta ≤ sparseLossPredictableVarianceRealizedMarkovTunedThreshold arms gamma horizon sparsity delta","missing":[],"search":"sampledpredictablevariancesquaresparselossrealizedmarkovhighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictablevariancesquaresparselossrealizedmarkovhighprobabilityregretbudget_le_tunedthreshold the complete sparse-loss markov budget is bounded by the eta-tuned threshold. theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedSparseLossPredictableVarianceRealizedMarkovRegret_tail","label":"sampledPredictable_tunedSparseLossPredictableVarianceRealizedMarkovRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedSparseLossPredictableVarianceRealizedMarkovRegret_tail","description":"Generated sparse-loss realized-regret tail with the exact predictable-variance-balanced learning rate.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedmarkovtuning/index.html#decl-14b46db8a887","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","order":4452,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning.lean:276"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedSparseLossPredictableVarianceRealizedMarkovRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hsparse : ∀ sample t, t < horizon → (sampledPredictableLossSupport arms loss t sample).card ≤ sparsity) : let eta := sparseLossPredictableVarianceHighProbabilityLearningRate arms g…","missing":[],"search":"sampledpredictable_tunedsparselosspredictablevariancerealizedmarkovregret_tail banditrlproof.exp3.sampledpredictable_tunedsparselosspredictablevariancerealizedmarkovregret_tail generated sparse-loss realized-regret tail with the exact predictable-variance-balanced learning rate. theorem compiled","shard":"modules/24c4d202a35479be.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSparsePathwiseVarianceBudget","label":"sampledPredictableSparsePathwiseVarianceBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSparsePathwiseVarianceBudget","description":"Deterministic cumulative predictable-variance budget on trajectories that obey the requested support-cardinality cap.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html#decl-2b73325e7624","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","order":4453,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableSparsePathwiseVarianceBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) : Real","missing":[],"search":"sampledpredictablesparsepathwisevariancebudget banditrlproof.exp3.sampledpredictablesparsepathwisevariancebudget deterministic cumulative predictable-variance budget on trajectories that obey the requested support-cardinality cap. definition compiled","shard":"modules/626461cb8ae51bf1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_sparsePathwiseVarianceBudget_or_mem_sparsityFailure","label":"sampledPredictableMixedSquaredVarianceSum_le_sparsePathwiseVarianceBudget_or_mem_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_sparsePathwiseVarianceBudget_or_mem_sparsityFailure","description":"Every trajectory either obeys the sparse pathwise variance budget or lies in the explicit sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html#decl-eb2fac96b0f1","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","order":4454,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredVarianceSum_le_sparsePathwiseVarianceBudget_or_mem_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledPredictableMixedSquaredVarianceSum arms eta gamma loss horizon sample ≤ sampledPredictableSparsePathwiseVarianceBudget arms gamma horizon sparsity ∨ sample ∈ sampledPredictableSparsityFailure arms loss horizon sparsity","missing":[],"search":"sampledpredictablemixedsquaredvariancesum_le_sparsepathwisevariancebudget_or_mem_sparsityfailure banditrlproof.exp3.sampledpredictablemixedsquaredvariancesum_le_sparsepathwisevariancebudget_or_mem_sparsityfailure every trajectory either obeys the sparse pathwise variance budget or lies in the explicit sparsity-failure event. theorem compiled","shard":"modules/626461cb8ae51bf1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget","label":"sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget","description":"Four-event realized sparse-loss budget with deterministic pathwise predictable variance on the sparsity-good event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html#decl-88e61374ae6e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","order":4455,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean:65"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregretbudget banditrlproof.exp3.sampledpredictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregretbudget four-event realized sparse-loss budget with deterministic pathwise predictable variance on the sparsity-good event. definition compiled","shard":"modules/626461cb8ae51bf1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_off_sparsityFailure","label":"sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_off_sparsityFailure","description":"Generated realized regret away from the exact support-sparsity failure event. Only the four confidence events remain after removing the common sparsity-failure set.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html#decl-7f1633e4d729","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","order":4456,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean:78"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu…","missing":[],"search":"sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregret_tail_off_sparsityfailure generated realized regret away from the exact support-sparsity failure event. only the four confidence events remain after removing the common sparsity-failure set. theorem compiled","shard":"modules/626461cb8ae51bf1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail","label":"sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail","description":"Generated realized regret under probabilistic sparsity and deterministic pathwise variance on the good event. The common support-sparsity failure set is added exactly once to the off-bad confidence tail.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html#decl-4464fdc7645e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","order":4457,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean:199"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPr…","missing":[],"search":"sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregret_tail generated realized regret under probabilistic sparsity and deterministic pathwise variance on the good event. the common support-sparsity failure set is added exactly once to the off-bad confidence tail. theorem compiled","shard":"modules/626461cb8ae51bf1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_of_sparsityFailure_le","description":"Practical `delta + epsilon` consumer of the pathwise-variance residual theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-a99cbedfeaa5/index.html#decl-0392f204eb21","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","order":4458,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity.lean:258"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hfailure : (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.env…","missing":[],"search":"sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_predictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregret_tail_of_sparsityfailure_le practical `delta + epsilon` consumer of the pathwise-variance residual theorem. theorem compiled","shard":"modules/626461cb8ae51bf1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossLargeHorizonCondition","label":"pathwiseVarianceProbabilisticSparseLossLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossLargeHorizonCondition","description":"Regime in which all four components of the pathwise probabilistic-sparsity exploration schedule are at most one half.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-40d17aa0b72d/index.html#decl-ca4e9115f8b2","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","order":4459,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def pathwiseVarianceProbabilisticSparseLossLargeHorizonCondition (K S T delta : Real) : Prop","missing":[],"search":"pathwisevarianceprobabilisticsparselosslargehorizoncondition banditrlproof.exp3.pathwisevarianceprobabilisticsparselosslargehorizoncondition regime in which all four components of the pathwise probabilistic-sparsity exploration schedule are at most one half. definition compiled","shard":"modules/0aa582151645075c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossAllHorizonRegretThreshold","label":"pathwiseVarianceProbabilisticSparseLossAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossAllHorizonRegretThreshold","description":"All-horizon threshold for the pathwise probabilistic-sparsity route: use the explicit large-horizon threshold in its valid regime and `T + 1` otherwise.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-40d17aa0b72d/index.html#decl-11cab659c384","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","order":4460,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon.lean:41"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossallhorizonregretthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossallhorizonregretthreshold all-horizon threshold for the pathwise probabilistic-sparsity route: use the explicit large-horizon threshold in its valid regime and `t + 1` otherwise. definition compiled","shard":"modules/0aa582151645075c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","description":"Generated all-horizon realized-regret theorem away from the exact support-sparsity failure event. Both branches use the same internally selected eta, gamma, and generated trajectory measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-40d17aa0b72d/index.html#decl-71bd37bdcfed","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","order":4461,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon.lean:58"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := pathwiseVarianceProbabilisticSparseLossHighProb…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure generated all-horizon realized-regret theorem away from the exact support-sparsity failure event. both branches use the same internally selected eta, gamma, and generated trajectory measure. theorem compiled","shard":"modules/0aa582151645075c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","description":"Generated all-horizon realized-regret theorem with the exact support-sparsity failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-40d17aa0b72d/index.html#decl-86f73dc08083","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","order":4462,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon.lean:141"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancerealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancerealizedregret_tail generated all-horizon realized-regret theorem with the exact support-sparsity failure residual. theorem compiled","shard":"modules/0aa582151645075c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","description":"Practical all-horizon `delta + epsilon` theorem under an exact bound on the support-sparsity failure event for the same internally tuned measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-40d17aa0b72d/index.html#decl-b88970a79a1c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","order":4463,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon.lean:223"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) delta let eta := pathwiseVarianceProbabilisticSparseLo…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le practical all-horizon `delta + epsilon` theorem under an exact bound on the support-sparsity failure event for the same internally tuned measure. theorem compiled","shard":"modules/0aa582151645075c.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","label":"pathwiseVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","description":"Best-arm all-horizon threshold. The fixed-comparator schedule receives the armwise confidence share `delta / K`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html#decl-42208dff1fdb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4464,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossbestarmallhorizonregretthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossbestarmallhorizonregretthreshold best-arm all-horizon threshold. the fixed-comparator schedule receives the armwise confidence share `delta / k`. definition compiled","shard":"modules/ecf5383114beb843.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_off_sparsityFailure","description":"Best-arm all-horizon tail away from the common support-sparsity failure event. The comparator union spends only the armwise confidence shares.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html#decl-e3eaddb49146","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4465,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:34"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHighProbabi…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_off_sparsityfailure best-arm all-horizon tail away from the common support-sparsity failure event. the comparator union spends only the armwise confidence shares. theorem compiled","shard":"modules/ecf5383114beb843.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail","description":"All-horizon best-arm residual theorem. The common sparsity-failure event is charged once for every arm because this wrapper consumes only the compiled fixed-comparator residual surface.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html#decl-037d0f621b8e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4466,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:184"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arm…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail all-horizon best-arm residual theorem. the common sparsity-failure event is charged once for every arm because this wrapper consumes only the compiled fixed-comparator residual surface. theorem compiled","shard":"modules/ecf5383114beb843.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le","description":"Practical best-arm all-horizon theorem. Per-arm calibration of both confidence and sparsity-failure budgets yields total failure `delta + epsilon`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html#decl-ad800aa5efb6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4467,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:334"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossH…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_of_sparsityfailure_le practical best-arm all-horizon theorem. per-arm calibration of both confidence and sparsity-failure budgets yields total failure `delta + epsilon`. theorem compiled","shard":"modules/ecf5383114beb843.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_single_sparsityFailure","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_single_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_single_sparsityFailure","description":"All-horizon best-arm residual theorem that charges the common support-sparsity failure event exactly once.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html#decl-5517a4715f5c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4468,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:428"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_single_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilisticSparseLossHighProb…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_single_sparsityfailure banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_single_sparsityfailure all-horizon best-arm residual theorem that charges the common support-sparsity failure event exactly once. theorem compiled","shard":"modules/ecf5383114beb843.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le_single_charge","label":"sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le_single_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le_single_charge","description":"Practical best-arm all-horizon theorem with a single charge for the common support-sparsity failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-3921c950650b/index.html#decl-ca64e010391e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","order":4469,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon.lean:525"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le_single_charge {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let deltaArm := delta / (arms.card : Real) let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (arms.card : Real) (sparsity : Real) (horizon : Real) deltaArm let eta := pathwiseVarianceProbabilis…","missing":[],"search":"sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_of_sparsityfailure_le_single_charge banditrlproof.exp3.sampledpredictable_allhorizonprobabilisticsparselosspathwisevariancebestarmrealizedregret_tail_of_sparsityfailure_le_single_charge practical best-arm all-horizon theorem with a single charge for the common support-sparsity failure event. theorem compiled","shard":"modules/ecf5383114beb843.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSparsePathwiseVarianceBudget_eq","label":"sampledPredictableSparsePathwiseVarianceBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSparsePathwiseVarianceBudget_eq","description":"Closed form of the deterministic sparse pathwise variance budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-e5ffcef7c475","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4470,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:34"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSparsePathwiseVarianceBudget_eq {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) : sampledPredictableSparsePathwiseVarianceBudget arms gamma horizon sparsity = (arms.card : Real) * (sparsity : Real) * (horizon : Real) / gamma","missing":[],"search":"sampledpredictablesparsepathwisevariancebudget_eq banditrlproof.exp3.sampledpredictablesparsepathwisevariancebudget_eq closed form of the deterministic sparse pathwise variance budget. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_mul_pathwiseVarianceProbabilisticSparseLossRadius_le_three_mul_sq_mul_horizon_sq","label":"log_mul_pathwiseVarianceProbabilisticSparseLossRadius_le_three_mul_sq_mul_horizon_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_mul_pathwiseVarianceProbabilisticSparseLossRadius_le_three_mul_sq_mul_horizon_sq","description":"The sparse base, pathwise fifth-power, and cubic confidence contracts bound the log-weighted predictable-variance radius by `3 * gamma^2 * T^2`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-37b8e6ffae84","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4471,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:49"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_mul_pathwiseVarianceProbabilisticSparseLossRadius_le_three_mul_sq_mul_horizon_sq {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) <= gamma ^ 5 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.log (arms.card : Real) * sampledMixedSquaredPredictableVarianceRadius arms gamma (sampledPredictableSparsePathwiseVarianceBudget arms gamma horizon sparsity) (delta / 4)…","missing":[],"search":"log_mul_pathwisevarianceprobabilisticsparselossradius_le_three_mul_sq_mul_horizon_sq banditrlproof.exp3.log_mul_pathwisevarianceprobabilisticsparselossradius_le_three_mul_sq_mul_horizon_sq the sparse base, pathwise fifth-power, and cubic confidence contracts bound the log-weighted predictable-variance radius by `3 * gamma^2 * t^2`. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossBalancedSqrt_le_two_mul_gamma_mul_horizon","label":"pathwiseVarianceProbabilisticSparseLossBalancedSqrt_le_two_mul_gamma_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossBalancedSqrt_le_two_mul_gamma_mul_horizon","description":"The sparse base and pathwise predictable-variance radius make the learning-rate-balanced square root at most `2 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-bf0e84953717","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4472,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:172"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossBalancedSqrt_le_two_mul_gamma_mul_horizon {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) <= gamma ^ 5 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) : Real.sqrt (Real.log (arms.card : Real) * pathwiseVarianceProbabilisticSparseLossHighProbabilityScale arms gamma horizon sparsity delta) <= 2 * gamma * (horizon : Real)","missing":[],"search":"pathwisevarianceprobabilisticsparselossbalancedsqrt_le_two_mul_gamma_mul_horizon banditrlproof.exp3.pathwisevarianceprobabilisticsparselossbalancedsqrt_le_two_mul_gamma_mul_horizon the sparse base and pathwise predictable-variance radius make the learning-rate-balanced square root at most `2 * gamma * t`. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedExplicitThreshold","label":"pathwiseVarianceProbabilisticSparseLossRealizedExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedExplicitThreshold","description":"Explicit realized-regret threshold after tuning eta and gamma.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-b557c1e2612e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4473,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:239"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossRealizedExplicitThreshold {Action : Type v} (_arms : Finset Action) (gamma : Real) (horizon _sparsity : Nat) (_delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossrealizedexplicitthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossrealizedexplicitthreshold explicit realized-regret threshold after tuning eta and gamma. definition compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold_le_explicitThreshold","label":"pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold_le_explicitThreshold","description":"Four algebraic exploration contracts reduce the eta-tuned pathwise threshold to `14 * gamma * T`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-5d1bf41112a6","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4474,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:246"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta) <= gamma ^ 5 * (horizon : Real) ^ 3) (hconfidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) (hrealized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * (horizon : Real)) : pathwiseVarianceProbabilisticSparseLossRealizedTuned…","missing":[],"search":"pathwisevarianceprobabilisticsparselossrealizedtunedthreshold_le_explicitthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossrealizedtunedthreshold_le_explicitthreshold four algebraic exploration contracts reduce the eta-tuned pathwise threshold to `14 * gamma * t`. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","description":"Gamma-characterized generated regret away from the exact sparsity-failure event and without a global-envelope Markov term.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-98a08bb92e95","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4475,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:343"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card :…","missing":[],"search":"sampledpredictable_gammacharacterizedprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_gammacharacterizedprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure gamma-characterized generated regret away from the exact sparsity-failure event and without a global-envelope markov term. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","label":"sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","description":"Gamma-characterized generated regret with the exact sparsity-failure residual and no global-envelope Markov term.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-73212d310e86","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4476,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:430"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (arms.card : Real) * (sparsity :…","missing":[],"search":"sampledpredictable_gammacharacterizedprobabilisticsparselosspathwisevariancerealizedregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizedprobabilisticsparselosspathwisevariancerealizedregret_tail gamma-characterized generated regret with the exact sparsity-failure residual and no global-envelope markov term. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","description":"Practical gamma-characterized endpoint under the exact generated-measure bound on the sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-89edf7eaecea","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4477,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:510"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hbase : (sparsity : Real) * Real.log (arms.card : Real) <= gamma ^ 2 * (horizon : Real)) (hmixed : (ar…","missing":[],"search":"sampledpredictable_gammacharacterizedprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_gammacharacterizedprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le practical gamma-characterized endpoint under the exact generated-measure bound on the sparsity-failure event. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossMixedExplorationScale","label":"pathwiseVarianceProbabilisticSparseLossMixedExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossMixedExplorationScale","description":"Fifth-root scale forced by the sparse pathwise predictable-variance radius. Unlike the Markov scale, it has no polynomial `1 / delta` factor.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-1ff99120d59c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4478,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:567"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossMixedExplorationScale (K S T delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossmixedexplorationscale banditrlproof.exp3.pathwisevarianceprobabilisticsparselossmixedexplorationscale fifth-root scale forced by the sparse pathwise predictable-variance radius. unlike the markov scale, it has no polynomial `1 / delta` factor. definition compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate","label":"pathwiseVarianceProbabilisticSparseLossRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate","description":"Unclipped maximum of the sparse arm, pathwise mixed-square, Bernstein, and realized-deviation exploration scales.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-05f1365cb8cf","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4479,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:574"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossRawExplorationRate (K S T delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossrawexplorationrate banditrlproof.exp3.pathwisevarianceprobabilisticsparselossrawexplorationrate unclipped maximum of the sparse arm, pathwise mixed-square, bernstein, and realized-deviation exploration scales. definition compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate","label":"pathwiseVarianceProbabilisticSparseLossClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate","description":"Explicit pathwise-variance exploration schedule clipped into the Hedge stability regime.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-ff4d4b215b88","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4480,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:586"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossClippedExplorationRate (K S T delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossclippedexplorationrate banditrlproof.exp3.pathwisevarianceprobabilisticsparselossclippedexplorationrate explicit pathwise-variance exploration schedule clipped into the hedge stability regime. definition compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_le_half","label":"pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_le_half","description":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_le_half (K S T delta : Real) : pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta <= 1 / 2","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-24f57ada5cc8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4481,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:591"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_le_half (K S T delta : Real) : pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta <= 1 / 2","missing":[],"search":"pathwisevarianceprobabilisticsparselossclippedexplorationrate_le_half banditrlproof.exp3.pathwisevarianceprobabilisticsparselossclippedexplorationrate_le_half theorem pathwisevarianceprobabilisticsparselossclippedexplorationrate_le_half (k s t delta : real) : pathwisevarianceprobabilisticsparselossclippedexplorationrate k s t delta <= 1 / 2 theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","label":"pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","description":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta <= 1 / 2) : pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta = pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-288e628caa18","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4482,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:598"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw (K S T delta : Real) (hraw : pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta <= 1 / 2) : pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta = pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta","missing":[],"search":"pathwisevarianceprobabilisticsparselossclippedexplorationrate_eq_raw banditrlproof.exp3.pathwisevarianceprobabilisticsparselossclippedexplorationrate_eq_raw theorem pathwisevarianceprobabilisticsparselossclippedexplorationrate_eq_raw (k s t delta : real) (hraw : pathwisevarianceprobabilisticsparselossrawexplorationrate k s t delta <= 1 / 2) : pathwisevarianceprobabilisticsparselossclippedexplorationrate k s t delta = pathwisevarianceprobabilisticsparselossrawexplorationrate k s t delta theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate_pos","label":"pathwiseVarianceProbabilisticSparseLossRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate_pos","description":"theorem pathwiseVarianceProbabilisticSparseLossRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-5819f21490f3","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4483,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:610"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossRawExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta","missing":[],"search":"pathwisevarianceprobabilisticsparselossrawexplorationrate_pos banditrlproof.exp3.pathwisevarianceprobabilisticsparselossrawexplorationrate_pos theorem pathwisevarianceprobabilisticsparselossrawexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < pathwisevarianceprobabilisticsparselossrawexplorationrate k s t delta theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_pos","label":"pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_pos","description":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-1d5a2df340bd","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4484,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:622"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_pos (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) : 0 < pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta","missing":[],"search":"pathwisevarianceprobabilisticsparselossclippedexplorationrate_pos banditrlproof.exp3.pathwisevarianceprobabilisticsparselossclippedexplorationrate_pos theorem pathwisevarianceprobabilisticsparselossclippedexplorationrate_pos (k s t delta : real) (hk_one : 1 < k) (hs : 0 < s) (ht : 0 < t) : 0 < pathwisevarianceprobabilisticsparselossclippedexplorationrate k s t delta theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","label":"pathwiseVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","description":"Transparent horizon contracts ensure every raw schedule component is at most one half, so clipping is inactive.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-f9a6f537150e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4485,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:635"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (K * S * Real.log K ^ 2 * Real.log (4 / delta)) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : pathwiseVarianceProbabilisticSparseLossRawExplorationRate K S T delta <= 1 / 2","missing":[],"search":"pathwisevarianceprobabilisticsparselossrawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.pathwisevarianceprobabilisticsparselossrawexplorationrate_le_half_of_horizon_contracts transparent horizon contracts ensure every raw schedule component is at most one half, so clipping is inactive. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_contracts","label":"pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_contracts","description":"The clipped maximum supplies exactly the four contracts consumed by the gamma-characterized pathwise-variance theorem.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-b40feae4b883","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4486,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:677"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_contracts (K S T delta : Real) (hK_one : 1 < K) (hS : 0 < S) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * (S * Real.log K) <= T) (hlarge_mixed : 32 * (K * S * Real.log K ^ 2 * Real.log (4 / delta)) <= T ^ 3) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : let gamma := pathwiseVarianceProbabilisticSparseLossClippedExplorationRate K S T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ S * Real.log K <= gamma ^ 2 * T ∧ K * S * Real.log K ^ 2 * Real.log (4 / delta) <= gamma ^ 5 * T ^ 3 ∧ K * Real.log (4 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * T","missing":[],"search":"pathwisevarianceprobabilisticsparselossclippedexplorationrate_contracts banditrlproof.exp3.pathwisevarianceprobabilisticsparselossclippedexplorationrate_contracts the clipped maximum supplies exactly the four contracts consumed by the gamma-characterized pathwise-variance theorem. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","description":"Fully explicit generated pathwise-variance regret tail away from the exact sparsity-failure event for the clipped maximum schedule.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-85d6148a947c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4487,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:798"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * ((arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4…","missing":[],"search":"sampledpredictable_explicitprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_explicitprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure fully explicit generated pathwise-variance regret tail away from the exact sparsity-failure event for the clipped maximum schedule. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","label":"sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","description":"Fully explicit generated pathwise-variance regret tail for the clipped maximum schedule, retaining the exact sparsity-failure residual.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-59ac81a0b864","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4488,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:875"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * ((arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * Real.log (4 / delta)) <= (horiz…","missing":[],"search":"sampledpredictable_explicitprobabilisticsparselosspathwisevariancerealizedregret_tail banditrlproof.exp3.sampledpredictable_explicitprobabilisticsparselosspathwisevariancerealizedregret_tail fully explicit generated pathwise-variance regret tail for the clipped maximum schedule, retaining the exact sparsity-failure residual. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","description":"Fully explicit practical `delta + epsilon` theorem under the exact sparsity-failure bound for the internally eta/gamma-tuned measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-4203b800b493/index.html#decl-1a62041672af","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","order":4489,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning.lean:952"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_arm : 4 * ((sparsity : Real) * Real.log (arms.card : Real)) <= (horizon : Real)) (hlarge_mixed : 32 * ((arms.card : Real) * (sparsity : Real) * Real.log (arms.card : Real) ^ 2 * R…","missing":[],"search":"sampledpredictable_explicitprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_explicitprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le fully explicit practical `delta + epsilon` theorem under the exact sparsity-failure bound for the internally eta/gamma-tuned measure. theorem compiled","shard":"modules/1a76e426f30074c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityScale","label":"pathwiseVarianceProbabilisticSparseLossHighProbabilityScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityScale","description":"Complete Hedge scale using sparse loss mass and sparse pathwise variance.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-dd47c21f7ef8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4490,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossHighProbabilityScale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselosshighprobabilityscale banditrlproof.exp3.pathwisevarianceprobabilisticsparselosshighprobabilityscale complete hedge scale using sparse loss mass and sparse pathwise variance. definition compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityScale_pos","label":"pathwiseVarianceProbabilisticSparseLossHighProbabilityScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityScale_pos","description":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityScale_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < pathwiseVarianceProbabilisticSparseLossHighProbabilityScale arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-13d77a8ed6eb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4491,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityScale_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < pathwiseVarianceProbabilisticSparseLossHighProbabilityScale arms gamma horizon sparsity delta","missing":[],"search":"pathwisevarianceprobabilisticsparselosshighprobabilityscale_pos banditrlproof.exp3.pathwisevarianceprobabilisticsparselosshighprobabilityscale_pos theorem pathwisevarianceprobabilisticsparselosshighprobabilityscale_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) : 0 < pathwisevarianceprobabilisticsparselosshighprobabilityscale arms gamma horizon sparsity delta theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate","label":"pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate","description":"Learning rate balancing entropy against the four-event sparse scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-05fe0455b33e","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4492,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:73"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate banditrlproof.exp3.pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate learning rate balancing entropy against the four-event sparse scale. definition compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_pos","label":"pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_pos","description":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-9b41613e3ddb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4493,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:82"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_pos {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : 0 < pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta","missing":[],"search":"pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate_pos banditrlproof.exp3.pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate_pos theorem pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate_pos {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) : 0 < pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate arms gamma horizon sparsity delta theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_sq_mul_scale","label":"pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_sq_mul_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_sq_mul_scale","description":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_sq_mul_scale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta ^ 2 *…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-1a01a15b199c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4494,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:98"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_sq_mul_scale {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta ^ 2 * pathwiseVarianceProbabilisticSparseLossHighProbabilityScale arms gamma horizon sparsity delta = Real.log (arms.card : Real)","missing":[],"search":"pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate_sq_mul_scale banditrlproof.exp3.pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate_sq_mul_scale theorem pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate_sq_mul_scale {action : type v} [decidableeq action] (arms : finset action) (hcard_two : 2 ≤ arms.card) (gamma : real) (hgamma_pos : 0 < gamma) (horizon sparsity : nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : real) : pathwisevarianceprobabilisticsparselosshighprobabilitylearningrate arms gamma horizon sparsity delta ^ 2 * pathwisevarianceprobabilisticsparselosshighprobabilityscale arms gamma horizon sparsity delta = real.log (arms.card : real) theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityHedgeBudget_le_three_mul_sqrt","label":"pathwiseVarianceProbabilisticSparseLossHighProbabilityHedgeBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityHedgeBudget_le_three_mul_sqrt","description":"Under `gamma ≤ 1/2`, entropy and stability cost at most three balanced copies of the four-event sparse scale.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-3cedefb498bb","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4495,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:121"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem pathwiseVarianceProbabilisticSparseLossHighProbabilityHedgeBudget_le_three_mul_sqrt {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : Real.log (arms.card : Real) / pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta + (pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta * (1 / (1 - gamma))) * pathwiseVarianceProbabilisticSparseLossHighProbabilityScale arms gamma horizon sparsity delta ≤ 3 * Real.sqrt (Real.log (arms.card : Real) * pathwiseVarianceProbabilisticSparseLossHighProbabilityScale arms gamma horizon sparsity delta)","missing":[],"search":"pathwisevarianceprobabilisticsparselosshighprobabilityhedgebudget_le_three_mul_sqrt banditrlproof.exp3.pathwisevarianceprobabilisticsparselosshighprobabilityhedgebudget_le_three_mul_sqrt under `gamma ≤ 1/2`, entropy and stability cost at most three balanced copies of the four-event sparse scale. theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold","label":"pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold","description":"Tuned regret threshold for the four-event pathwise-variance route.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-d631bc5250e7","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4496,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:200"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma : Real) (horizon sparsity : Nat) (delta : Real) : Real","missing":[],"search":"pathwisevarianceprobabilisticsparselossrealizedtunedthreshold banditrlproof.exp3.pathwisevarianceprobabilisticsparselossrealizedtunedthreshold tuned regret threshold for the four-event pathwise-variance route. definition compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget_le_tunedThreshold","description":"The raw four-event budget is bounded by the eta-tuned threshold.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-aa583e25ba6b","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4497,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:216"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) : sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget arms (pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta) gamma horizon sparsity delta ≤ pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold arms gamma horizon sparsity delta","missing":[],"search":"sampledpredictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictablevariancesquareprobabilisticsparselossrealizedpathwisevariancehighprobabilityregretbudget_le_tunedthreshold the raw four-event budget is bounded by the eta-tuned threshold. theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","label":"sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","description":"Eta-tuned generated regret away from the exact sparsity-failure event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-b8267c8b5f98","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4498,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:244"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImportanceWeighte…","missing":[],"search":"sampledpredictable_tunedprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure banditrlproof.exp3.sampledpredictable_tunedprobabilisticsparselosspathwisevariancerealizedregret_tail_off_sparsityfailure eta-tuned generated regret away from the exact sparsity-failure event. theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","label":"sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","description":"Eta-tuned generated regret with the exact sparsity-failure residual and the sparse pathwise variance budget.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-8126d9718fdd","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4499,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:328"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta : Real) (hdelta : 0 < delta) : let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel ar…","missing":[],"search":"sampledpredictable_tunedprobabilisticsparselosspathwisevariancerealizedregret_tail banditrlproof.exp3.sampledpredictable_tunedprobabilisticsparselosspathwisevariancerealizedregret_tail eta-tuned generated regret with the exact sparsity-failure residual and the sparse pathwise variance budget. theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","label":"sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","description":"Practical eta-tuned `delta + epsilon` theorem under the exact internally tuned generated measure.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancesparselossrealizedpathwisevariancep-95f6c41637e8/index.html#decl-d1d2920e5f24","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","order":4500,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning.lean:392"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 ≤ arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma ≤ 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon sparsity : Nat) (hhorizon : 0 < horizon) (hsparsity : 0 < sparsity) (delta epsilon : Real) (hdelta : 0 < delta) : let eta := pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate arms gamma horizon sparsity delta let mu := prior ⊗ₘ sampledImporta…","missing":[],"search":"sampledpredictable_tunedprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le banditrlproof.exp3.sampledpredictable_tunedprobabilisticsparselosspathwisevariancerealizedregret_tail_of_sparsityfailure_le practical eta-tuned `delta + epsilon` theorem under the exact internally tuned generated measure. theorem compiled","shard":"modules/26a04dea238c920a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.abs_mixedSquaredEstimatorDeviation_le_inv_floor","label":"abs_mixedSquaredEstimatorDeviation_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.abs_mixedSquaredEstimatorDeviation_le_inv_floor","description":"A centered mixed-square score is bounded by the reciprocal probability floor on the support of the finite sampling distribution.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-158b9cb77731","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4501,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem abs_mixedSquaredEstimatorDeviation_le_inv_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (chosen : Action) (hchosen : chosen ∈ arms) : |mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) chosen - arms.sum (fun action => (loss history action) ^ 2)| <= 1 / epsilon","missing":[],"search":"abs_mixedsquaredestimatordeviation_le_inv_floor banditrlproof.exp3.abs_mixedsquaredestimatordeviation_le_inv_floor a centered mixed-square score is bounded by the reciprocal probability floor on the support of the finite sampling distribution. theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt_variance","label":"finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt_variance","description":"Exact finite-law fixed-tilt MGF bound. Unlike the deterministic Bernstein wrapper, its exponent retains the actual centered second moment.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-09e32ef2bd0c","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4502,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:87"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt_variance {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) : Concentration.HasMGFUpperBoundAt (fun chosen => mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) chosen - arms.sum (fun action => (loss history action) ^ 2)) tilt (tilt ^ 2 * mixedSquaredEstimatorCenteredSecondMoment arms prob loss history) (finiteActionMeasure arms (prob history))","missing":[],"search":"finiteactionmixedsquaredestimator_hasmgfupperboundat_variance banditrlproof.exp3.finiteactionmixedsquaredestimator_hasmgfupperboundat_variance exact finite-law fixed-tilt mgf bound. unlike the deterministic bernstein wrapper, its exponent retains the actual centered second moment. theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_compensated_hasMGFUpperBoundAt","label":"finiteActionMixedSquaredEstimator_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_compensated_hasMGFUpperBoundAt","description":"Finite-law exponential-supermartingale increment obtained by subtracting the exact variance budget from the centered mixed-square score.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-bd045da8f2f8","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4503,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:179"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionMixedSquaredEstimator_compensated_hasMGFUpperBoundAt {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) : Concentration.HasMGFUpperBoundAt (fun chosen => tilt * (mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) chosen - arms.sum (fun action => (loss history action) ^ 2)) - tilt ^ 2 * mixedSquaredEstimatorCenteredSecondMoment arms prob loss history) 1 0 (finiteActionMeasure arms (prob history))","missing":[],"search":"finiteactionmixedsquaredestimator_compensated_hasmgfupperboundat banditrlproof.exp3.finiteactionmixedsquaredestimator_compensated_hasmgfupperboundat finite-law exponential-supermartingale increment obtained by subtracting the exact variance budget from the centered mixed-square score. theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mixedSquaredEstimator_compensated_hasCondMGFUpperBoundAt_of_condDistrib","label":"mixedSquaredEstimator_compensated_hasCondMGFUpperBoundAt_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mixedSquaredEstimator_compensated_hasCondMGFUpperBoundAt_of_condDistrib","description":"An identified finite conditional action law yields a zero-budget conditional MGF for the exact variance-compensated mixed-square increment.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-c69beccd1290","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4504,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:206"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mixedSquaredEstimator_compensated_hasCondMGFUpperBoundAt_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) (hcond : condDistrib action history mu =ᵐ…","missing":[],"search":"mixedsquaredestimator_compensated_hascondmgfupperboundat_of_conddistrib banditrlproof.exp3.mixedsquaredestimator_compensated_hascondmgfupperboundat_of_conddistrib an identified finite conditional action law yields a zero-budget conditional mgf for the exact variance-compensated mixed-square increment. theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensated_zero_hasCondMGFUpperBoundAt","label":"sampledPredictableMixedSquaredCompensated_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensated_zero_hasCondMGFUpperBoundAt","description":"theorem sampledPredictableMixedSquaredCompensated_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-22ab28ae0871","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4505,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:405"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredCompensated_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (fun…","missing":[],"search":"sampledpredictablemixedsquaredcompensated_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablemixedsquaredcompensated_zero_hascondmgfupperboundat theorem sampledpredictablemixedsquaredcompensated_zero_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment concentration.hascondmgfupperboundat ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (fun sample => tilt * sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss 0 sample - tilt ^ 2 * sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss 0 sample) 1 0 mu theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensated_succ_hasCondMGFUpperBoundAt","label":"sampledPredictableMixedSquaredCompensated_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensated_succ_hasCondMGFUpperBoundAt","description":"theorem sampledPredictableMixedSquaredCompensated_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-55a6691a81ff","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4506,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:465"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredCompensated_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Concentration.HasCondMGFUpperBoundAt ((inferInstance :…","missing":[],"search":"sampledpredictablemixedsquaredcompensated_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablemixedsquaredcompensated_succ_hascondmgfupperboundat theorem sampledpredictablemixedsquaredcompensated_succ_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) concentration.hascondmgfupperboundat ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (fun sample => tilt * sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss (n + 1) sample - tilt ^ 2 * sampledtrajectorypredictablemixedsquaredvarianc…","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess","label":"sampledPredictableMixedSquaredCompensatedProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess","description":"Shifted variance-compensated increment process. Index `i + 1` contains the actual-time `i` centered increment and its predictable variance.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-3a9e3eadd349","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4507,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:527"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableMixedSquaredCompensatedProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma tilt : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real","missing":[],"search":"sampledpredictablemixedsquaredcompensatedprocess banditrlproof.exp3.sampledpredictablemixedsquaredcompensatedprocess shifted variance-compensated increment process. index `i + 1` contains the actual-time `i` centered increment and its predictable variance. definition compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess_stronglyAdapted","label":"sampledPredictableMixedSquaredCompensatedProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess_stronglyAdapted","description":"theorem sampledPredictableMixedSquaredCompensatedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma tilt : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviati…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-4c33340ff491","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4508,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:539"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredCompensatedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma tilt : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPredictableMixedSquaredCompensatedProcess arms eta gamma tilt loss)","missing":[],"search":"sampledpredictablemixedsquaredcompensatedprocess_stronglyadapted banditrlproof.exp3.sampledpredictablemixedsquaredcompensatedprocess_stronglyadapted theorem sampledpredictablemixedsquaredcompensatedprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma tilt : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablemixedsquaredcompensatedprocess arms eta gamma tilt loss) theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess_sum_range_succ","label":"sampledPredictableMixedSquaredCompensatedProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess_sum_range_succ","description":"theorem sampledPredictableMixedSquaredCompensatedProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma tilt : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableMixedSquaredCompensatedProcess a…","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-fd6c0b06cba2","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4509,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:561"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredCompensatedProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma tilt : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableMixedSquaredCompensatedProcess arms eta gamma tilt loss i sample) = tilt * (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredDeviationAt arms eta gamma loss i sample) - tilt ^ 2 * (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredVarianceAt arms eta gamma loss i sample)","missing":[],"search":"sampledpredictablemixedsquaredcompensatedprocess_sum_range_succ banditrlproof.exp3.sampledpredictablemixedsquaredcompensatedprocess_sum_range_succ theorem sampledpredictablemixedsquaredcompensatedprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma tilt : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpredictablemixedsquaredcompensatedprocess arms eta gamma tilt loss i sample) = tilt * (finset.range horizon).sum (fun i => sampledtrajectorypredictablemixedsquareddeviationat arms eta gamma loss i sample) - tilt ^ 2 * (finset.range horizon).sum (fun i => sampledtrajectorypredictablemixedsquaredvarianceat arms eta gamma loss i sample) theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_fixedTilt","label":"sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_fixedTilt","description":"Fixed-tilt predictable-variance tail for the generated centered mixed-square process. The variance is kept random and enters through the event `sum V <= varianceBudget`; it is not replaced by the deterministic `horizon * K / epsilon` envelope.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-52a52d7b3028","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4510,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:585"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_fixedTilt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) (threshold varianceBudget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | threshold <= (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableMixedSquaredDeviati…","missing":[],"search":"sampledpredictablemixedsquareddeviation_sum_tail_predictablevariance_fixedtilt banditrlproof.exp3.sampledpredictablemixedsquareddeviation_sum_tail_predictablevariance_fixedtilt fixed-tilt predictable-variance tail for the generated centered mixed-square process. the variance is kept random and enters through the event `sum v <= variancebudget`; it is not replaced by the deterministic `horizon * k / epsilon` envelope. theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledMixedSquaredPredictableVarianceRadius","label":"sampledMixedSquaredPredictableVarianceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledMixedSquaredPredictableVarianceRadius","description":"Optimized radius for the mixed-square predictable-variance event.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-270d9d263857","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4511,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:714"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledMixedSquaredPredictableVarianceRadius {Action : Type v} [DecidableEq Action] (arms : Finset Action) (gamma varianceBudget delta : Real) : Real","missing":[],"search":"sampledmixedsquaredpredictablevarianceradius banditrlproof.exp3.sampledmixedsquaredpredictablevarianceradius optimized radius for the mixed-square predictable-variance event. definition compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_delta","label":"sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_delta","description":"Delta-shaped predictable-variance Bernstein/Freedman bound for the generated centered mixed-square process. This controls the deviation jointly with the event that its cumulative predictable variance is at most `varianceBudget`.","url":"../modules/banditrlproof-exp3mixedsquarepredictablevariancetail/index.html#decl-6a5e36f395c9","parent":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","order":4512,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3MixedSquarePredictableVarianceTail"],["Source","BanditRLProof/Exp3MixedSquarePredictableVarianceTail.lean:725"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledMixedSquaredPredictableVarianceRadius arms gamma varianceBudget delta <= (Finset.range horizon).sum (fun i => sampledTrajectory…","missing":[],"search":"sampledpredictablemixedsquareddeviation_sum_tail_predictablevariance_delta banditrlproof.exp3.sampledpredictablemixedsquareddeviation_sum_tail_predictablevariance_delta delta-shaped predictable-variance bernstein/freedman bound for the generated centered mixed-square process. this controls the deviation jointly with the event that its cumulative predictable variance is at most `variancebudget`. theorem compiled","shard":"modules/0e68ce7abd9ca78a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.potential","label":"potential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.potential","description":"Finite-action exponential-weights potential.","url":"../modules/banditrlproof-exp3potential/index.html#decl-f9536eba07ee","parent":"module:BanditRLProof.Exp3Potential","order":4513,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def potential {Action : Type u} (arms : Finset Action) (w : Action -> Real) : Real","missing":[],"search":"potential banditrlproof.exp3potential.potential finite-action exponential-weights potential. definition compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.updatedWeight","label":"updatedWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.updatedWeight","description":"Multiplicative exponential-weights update for one action.","url":"../modules/banditrlproof-exp3potential/index.html#decl-3200270c9742","parent":"module:BanditRLProof.Exp3Potential","order":4514,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def updatedWeight {Action : Type u} (eta : Real) (w loss : Action -> Real) (a : Action) : Real","missing":[],"search":"updatedweight banditrlproof.exp3potential.updatedweight multiplicative exponential-weights update for one action. definition compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.updatedPotential","label":"updatedPotential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.updatedPotential","description":"Potential after applying the exponential-weights update on each action.","url":"../modules/banditrlproof-exp3potential/index.html#decl-8b4ed8510666","parent":"module:BanditRLProof.Exp3Potential","order":4515,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def updatedPotential {Action : Type u} (arms : Finset Action) (eta : Real) (w loss : Action -> Real) : Real","missing":[],"search":"updatedpotential banditrlproof.exp3potential.updatedpotential potential after applying the exponential-weights update on each action. definition compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.updatedPotential_eq_sum","label":"updatedPotential_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.updatedPotential_eq_sum","description":"Unfold the updated potential as an explicit finite sum.","url":"../modules/banditrlproof-exp3potential/index.html#decl-f6fe1d86db92","parent":"module:BanditRLProof.Exp3Potential","order":4516,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:36"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem updatedPotential_eq_sum {Action : Type u} (arms : Finset Action) (eta : Real) (w loss : Action -> Real) : updatedPotential arms eta w loss = arms.sum (fun a => w a * Real.exp (-eta * loss a))","missing":[],"search":"updatedpotential_eq_sum banditrlproof.exp3potential.updatedpotential_eq_sum unfold the updated potential as an explicit finite sum. theorem compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.updatedWeight_nonneg_of_nonneg","label":"updatedWeight_nonneg_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.updatedWeight_nonneg_of_nonneg","description":"Exponential updating preserves nonnegative weights.","url":"../modules/banditrlproof-exp3potential/index.html#decl-294ef3a6124e","parent":"module:BanditRLProof.Exp3Potential","order":4517,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:44"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem updatedWeight_nonneg_of_nonneg {Action : Type u} (eta : Real) (w loss : Action -> Real) (a : Action) (hw : 0 <= w a) : 0 <= updatedWeight eta w loss a","missing":[],"search":"updatedweight_nonneg_of_nonneg banditrlproof.exp3potential.updatedweight_nonneg_of_nonneg exponential updating preserves nonnegative weights. theorem compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.updatedPotential_nonneg_of_nonneg","label":"updatedPotential_nonneg_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.updatedPotential_nonneg_of_nonneg","description":"The updated potential is nonnegative when all current finite weights are.","url":"../modules/banditrlproof-exp3potential/index.html#decl-062bc48b3185","parent":"module:BanditRLProof.Exp3Potential","order":4518,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:51"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem updatedPotential_nonneg_of_nonneg {Action : Type u} (arms : Finset Action) (eta : Real) (w loss : Action -> Real) (hw : forall a, a ∈ arms -> 0 <= w a) : 0 <= updatedPotential arms eta w loss","missing":[],"search":"updatedpotential_nonneg_of_nonneg banditrlproof.exp3potential.updatedpotential_nonneg_of_nonneg the updated potential is nonnegative when all current finite weights are. theorem compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.updatedPotential_sub_potential_eq_sum_weight_mul_exp_sub_one","label":"updatedPotential_sub_potential_eq_sum_weight_mul_exp_sub_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.updatedPotential_sub_potential_eq_sum_weight_mul_exp_sub_one","description":"One-step potential increment identity. This is the algebraic finite-sum surface used before applying any EXP/log inequality such as `exp x <= 1 + x + x^2`.","url":"../modules/banditrlproof-exp3potential/index.html#decl-ca5622006e10","parent":"module:BanditRLProof.Exp3Potential","order":4519,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:65"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem updatedPotential_sub_potential_eq_sum_weight_mul_exp_sub_one {Action : Type u} (arms : Finset Action) (eta : Real) (w loss : Action -> Real) : updatedPotential arms eta w loss - potential arms w = arms.sum (fun a => w a * (Real.exp (-eta * loss a) - 1))","missing":[],"search":"updatedpotential_sub_potential_eq_sum_weight_mul_exp_sub_one banditrlproof.exp3potential.updatedpotential_sub_potential_eq_sum_weight_mul_exp_sub_one one-step potential increment identity. this is the algebraic finite-sum surface used before applying any exp/log inequality such as `exp x <= 1 + x + x^2`. theorem compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.sum_range_forward_difference","label":"sum_range_forward_difference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.sum_range_forward_difference","description":"Finite-horizon telescoping for a real-valued potential process.","url":"../modules/banditrlproof-exp3potential/index.html#decl-9a9402c872e5","parent":"module:BanditRLProof.Exp3Potential","order":4520,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:78"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_range_forward_difference (Phi : Nat -> Real) (T : Nat) : (Finset.range T).sum (fun t => Phi (t + 1) - Phi t) = Phi T - Phi 0","missing":[],"search":"sum_range_forward_difference banditrlproof.exp3potential.sum_range_forward_difference finite-horizon telescoping for a real-valued potential process. theorem compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.potentialProcess","label":"potentialProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.potentialProcess","description":"Potential process induced by a time-indexed finite-action weight family.","url":"../modules/banditrlproof-exp3potential/index.html#decl-63d422a540e3","parent":"module:BanditRLProof.Exp3Potential","order":4521,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:90"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def potentialProcess {Action : Type u} (arms : Finset Action) (w : Nat -> Action -> Real) (t : Nat) : Real","missing":[],"search":"potentialprocess banditrlproof.exp3potential.potentialprocess potential process induced by a time-indexed finite-action weight family. definition compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3Potential.potentialProcess_telescope_sum_range","label":"potentialProcess_telescope_sum_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3Potential.potentialProcess_telescope_sum_range","description":"Finite-horizon telescope specialized to exponential-weights potentials. This is the compiled local replacement for treating \"the potential telescopes\" as only a proof weapon.","url":"../modules/banditrlproof-exp3potential/index.html#decl-66e2694dc5ac","parent":"module:BanditRLProof.Exp3Potential","order":4522,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3Potential"],["Source","BanditRLProof/Exp3Potential.lean:100"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem potentialProcess_telescope_sum_range {Action : Type u} (arms : Finset Action) (w : Nat -> Action -> Real) (T : Nat) : (Finset.range T).sum (fun t => potentialProcess arms w (t + 1) - potentialProcess arms w t) = potentialProcess arms w T - potentialProcess arms w 0","missing":[],"search":"potentialprocess_telescope_sum_range banditrlproof.exp3potential.potentialprocess_telescope_sum_range finite-horizon telescope specialized to exponential-weights potentials. this is the compiled local replacement for treating \"the potential telescopes\" as only a proof weapon. theorem compiled","shard":"modules/a90d8336912dfc27.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.PredictableLossVector","label":"PredictableLossVector","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Exp3.PredictableLossVector","description":"A measurable adversarial loss-vector family selected before the current action. The successor loss vector may depend on the environment and the preceding pair history, but not on the action sampled at that successor step.","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-9745f58bbd35","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4523,"meta":[["Kind","structure"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"structure PredictableLossVector (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] where","missing":[],"search":"predictablelossvector banditrlproof.exp3.predictablelossvector a measurable adversarial loss-vector family selected before the current action. the successor loss vector may depend on the environment and the preceding pair history, but not on the action sampled at that successor step. structure compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.PredictableLossVector.environment","label":"environment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.PredictableLossVector.environment","description":"Deterministic chosen-coordinate feedback generated by a predictable loss family.","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-ad3aaa505b2e","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4524,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:43"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def PredictableLossVector.environment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) : Thompson.MeasurableHistoryEnvironment Env Action Real where","missing":[],"search":"environment banditrlproof.exp3.predictablelossvector.environment deterministic chosen-coordinate feedback generated by a predictable loss family. definition compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.PredictableLossVector.environment_initialFeedback_apply","label":"environment_initialFeedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.PredictableLossVector.environment_initialFeedback_apply","description":"theorem PredictableLossVector.environment_initialFeedback_apply {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (env : Env) (action : Action) : loss.environment.initialFeedback (env, action) = Measure.dirac (loss.initial env action)","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-f4436ce67b99","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4525,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:57"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem PredictableLossVector.environment_initialFeedback_apply {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (env : Env) (action : Action) : loss.environment.initialFeedback (env, action) = Measure.dirac (loss.initial env action)","missing":[],"search":"environment_initialfeedback_apply banditrlproof.exp3.predictablelossvector.environment_initialfeedback_apply theorem predictablelossvector.environment_initialfeedback_apply {env : type u} {action : type v} [measurablespace env] [measurablespace action] (loss : predictablelossvector env action) (env : env) (action : action) : loss.environment.initialfeedback (env, action) = measure.dirac (loss.initial env action) theorem compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.PredictableLossVector.environment_feedback_apply","label":"environment_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.PredictableLossVector.environment_feedback_apply","description":"theorem PredictableLossVector.environment_feedback_apply {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (action : Action) : loss.environment.feedback n (env, (history, action)) = Measure.dirac (loss.successor n env history action)","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-221807470193","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4526,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:66"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem PredictableLossVector.environment_feedback_apply {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (action : Action) : loss.environment.feedback n (env, (history, action)) = Measure.dirac (loss.successor n env history action)","missing":[],"search":"environment_feedback_apply banditrlproof.exp3.predictablelossvector.environment_feedback_apply theorem predictablelossvector.environment_feedback_apply {env : type u} {action : type v} [measurablespace env] [measurablespace action] (loss : predictablelossvector env action) (n : nat) (env : env) (history : history.finitepairhistory action real n) (action : action) : loss.environment.feedback n (env, (history, action)) = measure.dirac (loss.successor n env history action) theorem compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.PredictableLossVector.initial_mem_unitInterval","label":"initial_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.PredictableLossVector.initial_mem_unitInterval","description":"theorem PredictableLossVector.initial_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (env : Env) (action : Action) : loss.initial env action ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-0f8d1e347dad","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4527,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:76"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem PredictableLossVector.initial_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (env : Env) (action : Action) : loss.initial env action ∈ Set.Icc (0 : Real) 1","missing":[],"search":"initial_mem_unitinterval banditrlproof.exp3.predictablelossvector.initial_mem_unitinterval theorem predictablelossvector.initial_mem_unitinterval {env : type u} {action : type v} [measurablespace env] [measurablespace action] (loss : predictablelossvector env action) (env : env) (action : action) : loss.initial env action ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.PredictableLossVector.successor_mem_unitInterval","label":"successor_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.PredictableLossVector.successor_mem_unitInterval","description":"theorem PredictableLossVector.successor_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (action : Action) : loss.successor n env history action ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-985a9f014981","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4528,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:83"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem PredictableLossVector.successor_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (n : Nat) (env : Env) (history : History.FinitePairHistory Action Real n) (action : Action) : loss.successor n env history action ∈ Set.Icc (0 : Real) 1","missing":[],"search":"successor_mem_unitinterval banditrlproof.exp3.predictablelossvector.successor_mem_unitinterval theorem predictablelossvector.successor_mem_unitinterval {env : type u} {action : type v} [measurablespace env] [measurablespace action] (loss : predictablelossvector env action) (n : nat) (env : env) (history : history.finitepairhistory action real n) (action : action) : loss.successor n env history action ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.trajectoryMixture_condDistrib_action_given_environment_history","label":"trajectoryMixture_condDistrib_action_given_environment_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.trajectoryMixture_condDistrib_action_given_environment_history","description":"Mixing fixed-environment trajectory laws preserves a common action kernel when the conditioning variable retains the environment coordinate.","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-5d8e2788e8eb","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4529,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:96"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMixture_condDistrib_action_given_environment_history {Env : Type u} {Omega : Type v} {History : Type w} {Action : Type x} [MeasurableSpace Env] [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (prior : Measure Env) [IsFiniteMeasure prior] (trajectory : Kernel Env Omega) [IsMarkovKernel trajectory] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (policy : Kernel History Action) [IsMarkovKernel policy] (hlaw : forall env, (trajectory env).map (fun omega => (history omega, action omega)) = (trajectory env).map history ⊗ₘ policy) : condDistrib (action ∘ Prod.snd) (fun sample : Env × Omega => (sample.1, history sample.2)) (prior ⊗ₘ trajectory) =ᵐ[ (prior ⊗ₘ trajectory).map (fun sample : Env × Omega => (sample.1, history sample.2))] po…","missing":[],"search":"trajectorymixture_conddistrib_action_given_environment_history banditrlproof.exp3.trajectorymixture_conddistrib_action_given_environment_history mixing fixed-environment trajectory laws preserves a common action kernel when the conditioning variable retains the environment coordinate. theorem compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryMeasure_condDistrib_action_given_environment","label":"sampledImportanceWeightedTrajectoryMeasure_condDistrib_action_given_environment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryMeasure_condDistrib_action_given_environment","description":"The concrete sampled-score EXP3 action remains independent of the current predictable loss vector after conditioning on the environment and prefix.","url":"../modules/banditrlproof-exp3predictableadversary/index.html#decl-ea020233cccd","parent":"module:BanditRLProof.Exp3PredictableAdversary","order":4530,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableAdversary"],["Source","BanditRLProof/Exp3PredictableAdversary.lean:219"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledImportanceWeightedTrajectoryMeasure_condDistrib_action_given_environment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2)) (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment) =ᵐ[ (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hg…","missing":[],"search":"sampledimportanceweightedtrajectorymeasure_conddistrib_action_given_environment banditrlproof.exp3.sampledimportanceweightedtrajectorymeasure_conddistrib_action_given_environment the concrete sampled-score exp3 action remains independent of the current predictable loss vector after conditioning on the environment and prefix. theorem compiled","shard":"modules/71f59f5ffeeaf39a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_nonneg_ae","label":"sampledPredictableTrajectoryMeasure_reward_nonneg_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_nonneg_ae","description":"Every observed scalar reward is nonnegative almost surely under the generated predictable sampled-EXP3 trajectory law.","url":"../modules/banditrlproof-exp3predictablehedge/index.html#decl-e19593bb5a15","parent":"module:BanditRLProof.Exp3PredictableHedge","order":4531,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableHedge"],["Source","BanditRLProof/Exp3PredictableHedge.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_reward_nonneg_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, 0 <= (sample.2 t).2","missing":[],"search":"sampledpredictabletrajectorymeasure_reward_nonneg_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_reward_nonneg_ae every observed scalar reward is nonnegative almost surely under the generated predictable sampled-exp3 trajectory law. theorem compiled","shard":"modules/5f867901da88b643.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_finiteHorizon_reward_nonneg_ae","label":"sampledPredictableTrajectoryMeasure_finiteHorizon_reward_nonneg_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_finiteHorizon_reward_nonneg_ae","description":"One common almost-sure event supplies reward nonnegativity at every time strictly before a finite horizon.","url":"../modules/banditrlproof-exp3predictablehedge/index.html#decl-27597fe74218","parent":"module:BanditRLProof.Exp3PredictableHedge","order":4532,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableHedge"],["Source","BanditRLProof/Exp3PredictableHedge.lean:71"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_finiteHorizon_reward_nonneg_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, ∀ t, t < horizon -> 0 <= (sample.2 t).2","missing":[],"search":"sampledpredictabletrajectorymeasure_finitehorizon_reward_nonneg_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_finitehorizon_reward_nonneg_ae one common almost-sure event supplies reward nonnegativity at every time strictly before a finite horizon. theorem compiled","shard":"modules/5f867901da88b643.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_hedge_regret_le_ae","label":"sampledPredictableTrajectoryMeasure_hedge_regret_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_hedge_regret_le_ae","description":"The concrete finite-horizon Hedge inequality holds almost surely on the generated predictable sampled-EXP3 trajectory.","url":"../modules/banditrlproof-exp3predictablehedge/index.html#decl-a57cfe62d5f7","parent":"module:BanditRLProof.Exp3PredictableHedge","order":4533,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableHedge"],["Source","BanditRLProof/Exp3PredictableHedge.lean:96"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_hedge_regret_le_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (Finset.range horizon).sum (fun t => mixedLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t) - cumulativeLoss (sampledTrajectoryObservedLoss arms eta gamma sample) h…","missing":[],"search":"sampledpredictabletrajectorymeasure_hedge_regret_le_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_hedge_regret_le_ae the concrete finite-horizon hedge inequality holds almost surely on the generated predictable sampled-exp3 trajectory. theorem compiled","shard":"modules/5f867901da88b643.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableScoreHedge_ae","label":"sampledPredictableScoreHedge_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableScoreHedge_ae","description":"The same almost-sure inequality with the comparator cumulative estimator exposed as the inclusive concrete sampled history score.","url":"../modules/banditrlproof-exp3predictablehedge/index.html#decl-2faeb49d4a91","parent":"module:BanditRLProof.Exp3PredictableHedge","order":4534,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableHedge"],["Source","BanditRLProof/Exp3PredictableHedge.lean:130"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableScoreHedge_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ∀ᵐ sample ∂mu, (Finset.range (n + 1)).sum (fun t => mixedLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t) - sampledHistoryScore arms eta gamma n (Preorder.frestrictLe n sample.2) comparator <= Real.log arms.…","missing":[],"search":"sampledpredictablescorehedge_ae banditrlproof.exp3.sampledpredictablescorehedge_ae the same almost-sure inequality with the comparator cumulative estimator exposed as the inclusive concrete sampled history score. theorem compiled","shard":"modules/5f867901da88b643.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureProbabilityAt","label":"sampledTrajectoryPureProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPureProbabilityAt","description":"The pure exponential-weights probability used by Hedge at an actual sampled-trajectory time.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-8837e8535d1b","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4535,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPureProbabilityAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledtrajectorypureprobabilityat banditrlproof.exp3.sampledtrajectorypureprobabilityat the pure exponential-weights probability used by hedge at an actual sampled-trajectory time. definition compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureProbabilitySourceAt","label":"sampledTrajectoryPureProbabilitySourceAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPureProbabilitySourceAt","description":"Measurable finite-distribution source for the pure Hedge probabilities at every actual trajectory time.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-1d0c513b26d4","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4536,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:34"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPureProbabilitySourceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (t : Nat) : MeasurableFiniteActionDistribution arms (sampledTrajectoryPureProbabilityAt (Env := Env) arms eta gamma t)","missing":[],"search":"sampledtrajectorypureprobabilitysourceat banditrlproof.exp3.sampledtrajectorypureprobabilitysourceat measurable finite-distribution source for the pure hedge probabilities at every actual trajectory time. definition compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableLossAt","label":"sampledTrajectoryPurePredictableLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPurePredictableLossAt","description":"Pure-Hedge predictable loss at one actual trajectory time.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-436955c7c0b2","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4537,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:69"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPurePredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypurepredictablelossat banditrlproof.exp3.sampledtrajectorypurepredictablelossat pure-hedge predictable loss at one actual trajectory time. definition compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryExploredPredictableLossAt","label":"sampledTrajectoryExploredPredictableLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryExploredPredictableLossAt","description":"Exploration-mixed predictable loss at one actual trajectory time.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-4fb4029ad05d","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4538,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:80"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryExploredPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectoryexploredpredictablelossat banditrlproof.exp3.sampledtrajectoryexploredpredictablelossat exploration-mixed predictable loss at one actual trajectory time. definition compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureObservedLossAt","label":"sampledTrajectoryPureObservedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPureObservedLossAt","description":"The pure-Hedge mixed observed estimator at one actual trajectory time.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-86cdf0732eb7","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4539,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:91"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPureObservedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypureobservedlossat banditrlproof.exp3.sampledtrajectorypureobservedlossat the pure-hedge mixed observed estimator at one actual trajectory time. definition compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPurePredictableLossAt","label":"measurable_sampledTrajectoryPurePredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectoryPurePredictableLossAt","description":"theorem measurable_sampledTrajectoryPurePredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryPurePredictableLossAt arms eta gamma loss t)","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-8ad70626c0ca","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4540,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:98"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectoryPurePredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryPurePredictableLossAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectorypurepredictablelossat banditrlproof.exp3.measurable_sampledtrajectorypurepredictablelossat theorem measurable_sampledtrajectorypurepredictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (loss : predictablelossvector env action) (t : nat) : measurable (sampledtrajectorypurepredictablelossat arms eta gamma loss t) theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableLossAt_mem_unitInterval","label":"sampledTrajectoryPurePredictableLossAt_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPurePredictableLossAt_mem_unitInterval","description":"theorem sampledTrajectoryPurePredictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledTrajectoryPurePredictableLossAt arms eta gamma…","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-fb5c5a3a30f5","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4541,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:114"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPurePredictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledTrajectoryPurePredictableLossAt arms eta gamma loss t sample ∈ Set.Icc (0 : Real) 1","missing":[],"search":"sampledtrajectorypurepredictablelossat_mem_unitinterval banditrlproof.exp3.sampledtrajectorypurepredictablelossat_mem_unitinterval theorem sampledtrajectorypurepredictablelossat_mem_unitinterval {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : sampledtrajectorypurepredictablelossat arms eta gamma loss t sample ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryPurePredictableLossAt","label":"integrable_sampledTrajectoryPurePredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_sampledTrajectoryPurePredictableLossAt","description":"theorem integrable_sampledTrajectoryPurePredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (sampledTrajectoryPure…","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-f81471b21d0e","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4542,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:149"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledTrajectoryPurePredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (sampledTrajectoryPurePredictableLossAt arms eta gamma loss t) mu","missing":[],"search":"integrable_sampledtrajectorypurepredictablelossat banditrlproof.exp3.integrable_sampledtrajectorypurepredictablelossat theorem integrable_sampledtrajectorypurepredictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) -> action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (loss : predictablelossvector env action) (t : nat) : integrable (sampledtrajectorypurepredictablelossat arms eta gamma loss t) mu theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryExploredPredictableLossAt","label":"measurable_sampledTrajectoryExploredPredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectoryExploredPredictableLossAt","description":"theorem measurable_sampledTrajectoryExploredPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryExploredPredict…","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-c5ae451900b0","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4543,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:167"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectoryExploredPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectoryexploredpredictablelossat banditrlproof.exp3.measurable_sampledtrajectoryexploredpredictablelossat theorem measurable_sampledtrajectoryexploredpredictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : measurable (sampledtrajectoryexploredpredictablelossat arms eta gamma loss t) theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryExploredPredictableLossAt_mem_unitInterval","label":"sampledTrajectoryExploredPredictableLossAt_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryExploredPredictableLossAt_mem_unitInterval","description":"theorem sampledTrajectoryExploredPredictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × R…","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-c8d6902fb806","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4544,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:185"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryExploredPredictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample ∈ Set.Icc (0 : Real) 1","missing":[],"search":"sampledtrajectoryexploredpredictablelossat_mem_unitinterval banditrlproof.exp3.sampledtrajectoryexploredpredictablelossat_mem_unitinterval theorem sampledtrajectoryexploredpredictablelossat_mem_unitinterval {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : sampledtrajectoryexploredpredictablelossat arms eta gamma loss t sample ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryExploredPredictableLossAt","label":"integrable_sampledTrajectoryExploredPredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_sampledTrajectoryExploredPredictableLossAt","description":"theorem integrable_sampledTrajectoryExploredPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVe…","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-6d88c5c417c4","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4545,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:222"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledTrajectoryExploredPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t) mu","missing":[],"search":"integrable_sampledtrajectoryexploredpredictablelossat banditrlproof.exp3.integrable_sampledtrajectoryexploredpredictablelossat theorem integrable_sampledtrajectoryexploredpredictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) -> action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : integrable (sampledtrajectoryexploredpredictablelossat arms eta gamma loss t) mu theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureObservedLossAt_ae_eq_weightedPredictable","label":"sampledTrajectoryPureObservedLossAt_ae_eq_weightedPredictable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPureObservedLossAt_ae_eq_weightedPredictable","description":"The observed scalar reward can be replaced almost surely by its selected predictable coordinate inside the pure-Hedge mixed estimator.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-c3d5f1e36613","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4546,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:244"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPureObservedLossAt_ae_eq_weightedPredictable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment (fun sample => sampledTrajectoryPureObservedLossAt arms eta gamma t sample) =ᵐ[mu] (fun sample => weightedImportanceWeightedLoss arms (sampledTrajectoryProbabilityAt arms eta gamma t sample) (sampledTrajectoryPureProbabilityAt arms eta gamma t sample) (predictableLossAt l…","missing":[],"search":"sampledtrajectorypureobservedlossat_ae_eq_weightedpredictable banditrlproof.exp3.sampledtrajectorypureobservedlossat_ae_eq_weightedpredictable the observed scalar reward can be replaced almost surely by its selected predictable coordinate inside the pure-hedge mixed estimator. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryPureObservedLossAt","label":"integrable_sampledTrajectoryPureObservedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_sampledTrajectoryPureObservedLossAt","description":"The pure-Hedge observed mixed estimator is integrable on the generated trajectory law.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-28eab881198b","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4547,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:303"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledTrajectoryPureObservedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Integrable (sampledTrajectoryPureObservedLossAt arms eta gamma t) mu","missing":[],"search":"integrable_sampledtrajectorypureobservedlossat banditrlproof.exp3.integrable_sampledtrajectorypureobservedlossat the pure-hedge observed mixed estimator is integrable on the generated trajectory law. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictablePureObservedInitial_integral_eq","label":"sampledPredictablePureObservedInitial_integral_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictablePureObservedInitial_integral_eq","description":"At time zero, the pure-Hedge mixed observed estimator has the same integral as the pure-Hedge predictable loss.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-0ae6ee9f7007","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4548,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:346"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictablePureObservedInitial_integral_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment integral mu (sampledTrajectoryPureObservedLossAt arms eta gamma 0) = integral mu (sampledTrajectoryPurePredictableLossAt arms eta gamma loss 0)","missing":[],"search":"sampledpredictablepureobservedinitial_integral_eq banditrlproof.exp3.sampledpredictablepureobservedinitial_integral_eq at time zero, the pure-hedge mixed observed estimator has the same integral as the pure-hedge predictable loss. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictablePureObservedSuccessor_integral_eq","label":"sampledPredictablePureObservedSuccessor_integral_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictablePureObservedSuccessor_integral_eq","description":"At a successor time, the pure-Hedge mixed observed estimator has the same integral as the pure-Hedge predictable loss.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-b440de0e17a5","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4549,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:475"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictablePureObservedSuccessor_integral_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment integral mu (sampledTrajectoryPureObservedLossAt arms eta gamma (n + 1)) = integral mu (sampledTrajectoryPurePredictableLossAt arms eta gamma loss (n + 1))","missing":[],"search":"sampledpredictablepureobservedsuccessor_integral_eq banditrlproof.exp3.sampledpredictablepureobservedsuccessor_integral_eq at a successor time, the pure-hedge mixed observed estimator has the same integral as the pure-hedge predictable loss. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictablePureObservedAt_integral_eq","label":"sampledPredictablePureObservedAt_integral_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictablePureObservedAt_integral_eq","description":"Every actual time satisfies the adaptive pure-Hedge first-moment identity.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-e061e7f5faa4","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4550,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:627"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictablePureObservedAt_integral_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment integral mu (sampledTrajectoryPureObservedLossAt arms eta gamma t) = integral mu (sampledTrajectoryPurePredictableLossAt arms eta gamma loss t)","missing":[],"search":"sampledpredictablepureobservedat_integral_eq banditrlproof.exp3.sampledpredictablepureobservedat_integral_eq every actual time satisfies the adaptive pure-hedge first-moment identity. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictablePureObserved_finiteHorizon_integral_eq","label":"sampledPredictablePureObserved_finiteHorizon_integral_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictablePureObserved_finiteHorizon_integral_eq","description":"The adaptive pure-Hedge first-moment identity summed over a finite horizon.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-ed6360a5c851","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4551,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:652"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictablePureObserved_finiteHorizon_integral_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryPureObservedLossAt arms eta gamma t sample)) = integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryPurePredictableLossAt arms eta gamma loss t sample))","missing":[],"search":"sampledpredictablepureobserved_finitehorizon_integral_eq banditrlproof.exp3.sampledpredictablepureobserved_finitehorizon_integral_eq the adaptive pure-hedge first-moment identity summed over a finite horizon. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_hedge_exploredSecondMoment_le_ae","label":"sampledPredictableTrajectoryMeasure_hedge_exploredSecondMoment_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_hedge_exploredSecondMoment_le_ae","description":"The a.e. sampled-Hedge inequality with its pure estimator-square term replaced by the exploration-mixed square.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-936ce0e94c74","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4552,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:703"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_hedge_exploredSecondMoment_le_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment ∀ᵐ sample ∂mu, (Finset.range horizon).sum (fun t => sampledTrajectoryPureObservedLossAt arms eta gamma t sample) - (Finset.range horizon).sum (fun t => observedImportanceWeightedLossAt arm…","missing":[],"search":"sampledpredictabletrajectorymeasure_hedge_exploredsecondmoment_le_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_hedge_exploredsecondmoment_le_ae the a.e. sampled-hedge inequality with its pure estimator-square term replaced by the exploration-mixed square. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_integral_pureHedge_le_exploredSecondMoment","label":"sampledPredictable_integral_pureHedge_le_exploredSecondMoment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_integral_pureHedge_le_exploredSecondMoment","description":"Integrated sampled-Hedge control with the exploration-mixed second moment on the generated predictable trajectory law.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-064a57575c27","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4553,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:763"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_integral_pureHedge_le_exploredSecondMoment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryPureObservedLossAt arms eta gamma t sample)) - integral mu (fun sample => (Finset.range horizon).sum (fun t =…","missing":[],"search":"sampledpredictable_integral_purehedge_le_exploredsecondmoment banditrlproof.exp3.sampledpredictable_integral_purehedge_le_exploredsecondmoment integrated sampled-hedge control with the exploration-mixed second moment on the generated predictable trajectory law. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_integral_exploredLoss_le_pure_add_gamma","label":"sampledPredictable_integral_exploredLoss_le_pure_add_gamma","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_integral_exploredLoss_le_pure_add_gamma","description":"Expected exploration-mixed predictable loss is at most expected pure predictable loss plus `gamma` per round.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-6cf36b2cd97d","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4554,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:838"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_integral_exploredLoss_le_pure_add_gamma {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_lt_one.le loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample)) <= integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryPurePredictableLossAt arms eta gamma los…","missing":[],"search":"sampledpredictable_integral_exploredloss_le_pure_add_gamma banditrlproof.exp3.sampledpredictable_integral_exploredloss_le_pure_add_gamma expected exploration-mixed predictable loss is at most expected pure predictable loss plus `gamma` per round. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObserved_finiteHorizon_secondMoment_integral_le_card_mul","label":"sampledPredictableObserved_finiteHorizon_secondMoment_integral_le_card_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObserved_finiteHorizon_secondMoment_integral_le_card_mul","description":"The explored probability-mixed estimator square has expectation at most `|arms| * horizon` under predictable `[0,1]` losses.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-a11820f24b5a","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4555,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:907"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObserved_finiteHorizon_secondMoment_integral_le_card_mul {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => observedMixedSquaredImportanceWeightedLossAt arms eta gamma t sample)) <= (arms.card : Real) * (horizon : Real)","missing":[],"search":"sampledpredictableobserved_finitehorizon_secondmoment_integral_le_card_mul banditrlproof.exp3.sampledpredictableobserved_finitehorizon_secondmoment_integral_le_card_mul the explored probability-mixed estimator square has expectation at most `|arms| * horizon` under predictable `[0,1]` losses. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le","label":"sampledPredictable_expectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_expectedRegret_le","description":"Unoptimized expected predictable EXP3 regret bound. This is the first complete generated-trajectory theorem on the route; eta/gamma optimization is kept as a separate deterministic parameter leaf.","url":"../modules/banditrlproof-exp3predictableintegration/index.html#decl-3ce1bf2fe44e","parent":"module:BanditRLProof.Exp3PredictableIntegration","order":4556,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableIntegration"],["Source","BanditRLProof/Exp3PredictableIntegration.lean:978"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_expectedRegret_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample compara…","missing":[],"search":"sampledpredictable_expectedregret_le banditrlproof.exp3.sampledpredictable_expectedregret_le unoptimized expected predictable exp3 regret bound. this is the first complete generated-trajectory theorem on the route; eta/gamma optimization is kept as a separate deterministic parameter leaf. theorem compiled","shard":"modules/da27cf4f49cd00c0.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.trajectoryMixture_map_environment_history_output_eq_compProd","label":"trajectoryMixture_map_environment_history_output_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.trajectoryMixture_map_environment_history_output_eq_compProd","description":"Mixing fixed-environment trajectory laws preserves a history-dependent output kernel when the conditioning variable retains the environment coordinate.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-572ee45a8f9d","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4557,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMixture_map_environment_history_output_eq_compProd {Env : Type u} {Omega : Type v} {History : Type w} {Output : Type x} [MeasurableSpace Env] [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Output] (prior : Measure Env) [IsFiniteMeasure prior] (trajectory : Kernel Env Omega) [IsMarkovKernel trajectory] (history : Omega -> History) (hhistory : Measurable history) (output : Omega -> Output) (houtput : Measurable output) (outputKernel : Kernel (Env × History) Output) [IsMarkovKernel outputKernel] (hlaw : forall env, (trajectory env).map (fun omega => (history omega, output omega)) = (trajectory env).map history ⊗ₘ outputKernel.comap (fun h => (env, h)) (measurable_const.prodMk measurable_id)) : (prior ⊗ₘ trajectory).map (fun sample : Env × Omega => ((sample.1, history sample.2), output sample.2)) = (prior ⊗ₘ trajectory).map (fun sample : Env × Omega =>…","missing":[],"search":"trajectorymixture_map_environment_history_output_eq_compprod banditrlproof.exp3.trajectorymixture_map_environment_history_output_eq_compprod mixing fixed-environment trajectory laws preserves a history-dependent output kernel when the conditioning variable retains the environment coordinate. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_prefix_next_eq_compProd","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_prefix_next_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_prefix_next_eq_compProd","description":"The canonical measurable trajectory, mixed over an environment prior, has the joint law obtained by adjoining one global measurable history-step kernel to the retained environment/prefix history.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-a1954361fb83","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4558,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:150"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_prefix_next_eq_compProd {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Reward) (environment : Thompson.MeasurableHistoryEnvironment Env Action Reward) (n : Nat) : (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample : Env × ((k : Nat) -> Action × Reward) => ((sample.1, Preorder.frestrictLe n sample.2), sample.2 (n + 1))) = (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample : Env × ((k : Nat) -> Action × Reward) => (sample.1, Preorde…","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_map_environment_prefix_next_eq_compprod banditrlproof.exp3.canonicalmeasurableenvironmenttrajectorymeasure_map_environment_prefix_next_eq_compprod the canonical measurable trajectory, mixed over an environment prior, has the joint law obtained by adjoining one global measurable history-step kernel to the retained environment/prefix history. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_nextPair_given_environment_prefix","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_nextPair_given_environment_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_nextPair_given_environment_prefix","description":"Conditional on the latent environment and the preceding finite pair history, the next canonical trajectory pair follows the global measurable history-step kernel.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-224cd32b4976","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4559,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:190"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_nextPair_given_environment_prefix {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Reward) (environment : Thompson.MeasurableHistoryEnvironment Env Action Reward) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Reward) => sample.2 (n + 1)) (fun sample : Env × ((k : Nat) -> Action × Reward) => (sample.1, Preorder.frestrictLe n sample.2)) (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment) =ᵐ[ (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun…","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_conddistrib_nextpair_given_environment_prefix banditrlproof.exp3.canonicalmeasurableenvironmenttrajectorymeasure_conddistrib_nextpair_given_environment_prefix conditional on the latent environment and the preceding finite pair history, the next canonical trajectory pair follows the global measurable history-step kernel. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledEnvironmentHistoryDistributionSource","label":"sampledEnvironmentHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledEnvironmentHistoryDistributionSource","description":"The sampled EXP3 distribution viewed on a retained environment/prefix history.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-5ae91d71b1c9","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4560,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:223"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledEnvironmentHistoryDistributionSource {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : Nat) : MeasurableFiniteActionDistribution arms (fun input : Env × History.FinitePairHistory Action Real n => sampledHistoryDistribution arms eta gamma n input.2) where","missing":[],"search":"sampledenvironmenthistorydistributionsource banditrlproof.exp3.sampledenvironmenthistorydistributionsource the sampled exp3 distribution viewed on a retained environment/prefix history. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSuccessorLossRegularity","label":"sampledPredictableSuccessorLossRegularity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSuccessorLossRegularity","description":"Predictable successor losses satisfy the regularity contract required by the EXP3 conditional first- and second-moment transport layer.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-b26ae2c63639","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4561,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:251"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSuccessorLossRegularity {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : BoundedMeasurableLossWithProbabilityFloor arms (fun input : Env × History.FinitePairHistory Action Real n => sampledHistoryDistribution arms eta gamma n input.2) (fun input : Env × History.FinitePairHistory Action Real n => loss.successor n input.1 input.2) (gamma / (arms.card : Real)) where","missing":[],"search":"sampledpredictablesuccessorlossregularity banditrlproof.exp3.sampledpredictablesuccessorlossregularity predictable successor losses satisfy the regularity contract required by the exp3 conditional first- and second-moment transport layer. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledInitialEnvironmentDistributionSource","label":"sampledInitialEnvironmentDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledInitialEnvironmentDistributionSource","description":"The time-zero sampled EXP3 distribution viewed as a constant kernel on environments.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-9a948fe5f1bc","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4562,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:278"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledInitialEnvironmentDistributionSource {Env : Type u} {Action : Type v} [MeasurableSpace Env] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) : MeasurableFiniteActionDistribution arms (fun _env : Env => initialExploredDistribution arms eta gamma) where","missing":[],"search":"sampledinitialenvironmentdistributionsource banditrlproof.exp3.sampledinitialenvironmentdistributionsource the time-zero sampled exp3 distribution viewed as a constant kernel on environments. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableInitialLossRegularity","label":"sampledPredictableInitialLossRegularity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableInitialLossRegularity","description":"Initial predictable losses satisfy the sampled EXP3 regularity contract.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-24333c616667","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4563,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:292"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableInitialLossRegularity {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : BoundedMeasurableLossWithProbabilityFloor arms (fun _env : Env => initialExploredDistribution arms eta gamma) loss.initial (gamma / (arms.card : Real)) where","missing":[],"search":"sampledpredictableinitiallossregularity banditrlproof.exp3.sampledpredictableinitiallossregularity initial predictable losses satisfy the sampled exp3 regularity contract. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_eval_zero","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_eval_zero","description":"Mixing the canonical trajectory through a prior preserves its global initial-pair law.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-133460a1855b","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4564,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:315"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_eval_zero {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [Nonempty Action] [MeasurableSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Reward) (environment : Thompson.MeasurableHistoryEnvironment Env Action Reward) : (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample : Env × ((k : Nat) -> Action × Reward) => (sample.1, sample.2 0)) = prior ⊗ₘ Thompson.measurableEnvironmentInitialPairKernel algorithm environment","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_map_environment_eval_zero banditrlproof.exp3.canonicalmeasurableenvironmenttrajectorymeasure_map_environment_eval_zero mixing the canonical trajectory through a prior preserves its global initial-pair law. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_action_zero","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_action_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_action_zero","description":"The prior-mixed canonical trajectory retains the environment beside its initial action law.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-71354a44bab7","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4565,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:347"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_action_zero {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [Nonempty Action] [MeasurableSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Reward) (environment : Thompson.MeasurableHistoryEnvironment Env Action Reward) : (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample : Env × ((k : Nat) -> Action × Reward) => (sample.1, (sample.2 0).1)) = prior ⊗ₘ Kernel.const Env algorithm.initialAction","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_map_environment_action_zero banditrlproof.exp3.canonicalmeasurableenvironmenttrajectorymeasure_map_environment_action_zero the prior-mixed canonical trajectory retains the environment beside its initial action law. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action_zero_given_environment","label":"canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action_zero_given_environment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action_zero_given_environment","description":"Conditional on the retained environment, the canonical initial action follows `initialAction`.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-3d69c10570ce","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4566,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:389"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action_zero_given_environment {Env : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Reward] [Nonempty Reward] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Reward) (environment : Thompson.MeasurableHistoryEnvironment Env Action Reward) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Reward) => (sample.2 0).1) (fun sample : Env × ((k : Nat) -> Action × Reward) => sample.1) (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment) =ᵐ[ (prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm environment).map (fun sample : Env × ((k : Nat) -> Action × Reward) => sample.1)] Kernel.const Env algorithm.initialAction","missing":[],"search":"canonicalmeasurableenvironmenttrajectorymeasure_conddistrib_action_zero_given_environment banditrlproof.exp3.canonicalmeasurableenvironmenttrajectorymeasure_conddistrib_action_zero_given_environment conditional on the retained environment, the canonical initial action follows `initialaction`. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalPredictableTrajectoryMeasure_reward_zero_eq_initialLoss_ae","label":"canonicalPredictableTrajectoryMeasure_reward_zero_eq_initialLoss_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalPredictableTrajectoryMeasure_reward_zero_eq_initialLoss_ae","description":"Under predictable deterministic feedback, the observed initial reward is the initial loss-vector coordinate selected by the initial action.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-0e5092cf61ba","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4567,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:446"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalPredictableTrajectoryMeasure_reward_zero_eq_initialLoss_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [Nonempty Action] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Real) (loss : PredictableLossVector Env Action) : (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 0).2) =ᵐ[ prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm loss.environment] (fun sample : Env × ((k : Nat) -> Action × Real) => loss.initial sample.1 (sample.2 0).1)","missing":[],"search":"canonicalpredictabletrajectorymeasure_reward_zero_eq_initialloss_ae banditrlproof.exp3.canonicalpredictabletrajectorymeasure_reward_zero_eq_initialloss_ae under predictable deterministic feedback, the observed initial reward is the initial loss-vector coordinate selected by the initial action. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedInitial_first_second_moment","label":"sampledPredictableObservedInitial_first_second_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedInitial_first_second_moment","description":"The time-zero sampled EXP3 estimator has the observed-scalar armwise first moment and exact probability-mixed estimator-square moment.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-87ea522b6dfc","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4568,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:513"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedInitial_first_second_moment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let prob := initialExploredDistribution arms eta gamma (integral mu (fun sample => importanceWeightedLoss prob (fun _ => (sample.2 0).2) (sample.2 0).1 comparator) = integral prior (fun env => loss.initial env comparator)) ∧ (integral mu…","missing":[],"search":"sampledpredictableobservedinitial_first_second_moment banditrlproof.exp3.sampledpredictableobservedinitial_first_second_moment the time-zero sampled exp3 estimator has the observed-scalar armwise first moment and exact probability-mixed estimator-square moment. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.canonicalPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","label":"canonicalPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.canonicalPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","description":"Under predictable deterministic feedback, the observed successor reward is the loss-vector coordinate selected by the action in the same successor pair.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-8a5776da9b75","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4569,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:675"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem canonicalPredictableTrajectoryMeasure_reward_eq_successorLoss_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (prior : Measure Env) [IsFiniteMeasure prior] (algorithm : Thompson.HistoryAlgorithm Action Real) (loss : PredictableLossVector Env Action) (n : Nat) : (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).2) =ᵐ[ prior ⊗ₘ Thompson.canonicalMeasurableEnvironmentTrajectoryKernel algorithm loss.environment] (fun sample : Env × ((k : Nat) -> Action × Real) => loss.successor n sample.1 (Preorder.frestrictLe n sample.2) (sample.2 (n + 1)).1)","missing":[],"search":"canonicalpredictabletrajectorymeasure_reward_eq_successorloss_ae banditrlproof.exp3.canonicalpredictabletrajectorymeasure_reward_eq_successorloss_ae under predictable deterministic feedback, the observed successor reward is the loss-vector coordinate selected by the action in the same successor pair. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","label":"sampledPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","description":"Concrete sampled-loss EXP3 observes exactly the selected predictable successor loss in every round, almost surely under the environment/trajectory mixture.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-123b15de5709","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4570,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:758"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryMeasure_reward_eq_successorLoss_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).2) =ᵐ[ prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment] (fun sample : Env × ((k : Nat) -> Action × Real) => loss.successor n sample.1 (Preorder.frestrictLe n sample.2) (sample.2 (n + 1)).1)","missing":[],"search":"sampledpredictabletrajectorymeasure_reward_eq_successorloss_ae banditrlproof.exp3.sampledpredictabletrajectorymeasure_reward_eq_successorloss_ae concrete sampled-loss exp3 observes exactly the selected predictable successor loss in every round, almost surely under the environment/trajectory mixture. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSuccessorLoss_first_second_moment","label":"sampledPredictableSuccessorLoss_first_second_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSuccessorLoss_first_second_moment","description":"The concrete sampled EXP3 successor round has the armwise unbiased first moment and the exact probability-mixed estimator-square moment for every predictable loss vector with positive exploration.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-3ff871358dc3","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4571,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:787"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSuccessorLoss_first_second_moment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) let prob := fun input : Env × History.FinitePairHistory Action Real n => sampledHistoryDistribution arms eta gamma n inp…","missing":[],"search":"sampledpredictablesuccessorloss_first_second_moment banditrlproof.exp3.sampledpredictablesuccessorloss_first_second_moment the concrete sampled exp3 successor round has the armwise unbiased first moment and the exact probability-mixed estimator-square moment for every predictable loss vector with positive exploration. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedSuccessor_first_second_moment","label":"sampledPredictableObservedSuccessor_first_second_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedSuccessor_first_second_moment","description":"Observed-scalar form of the sampled EXP3 roundwise moment theorem. The score uses only the reward coordinate stored in the generated trajectory; the right sides expose the full predictable loss vector required by regret analysis.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-a7a60aa54938","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4572,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:914"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedSuccessor_first_second_moment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) let prob := fun input : Env × History.FinitePairHistory Action Real n => sampledHistoryDistribution arms eta gamma n…","missing":[],"search":"sampledpredictableobservedsuccessor_first_second_moment banditrlproof.exp3.sampledpredictableobservedsuccessor_first_second_moment observed-scalar form of the sampled exp3 roundwise moment theorem. the score uses only the reward coordinate stored in the generated trajectory; the right sides expose the full predictable loss vector required by regret analysis. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilityAt","label":"sampledTrajectoryProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryProbabilityAt","description":"Sampling probabilities used by the concrete trajectory at every actual time index.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-4317ff7ffaa8","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4573,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1008"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryProbabilityAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledtrajectoryprobabilityat banditrlproof.exp3.sampledtrajectoryprobabilityat sampling probabilities used by the concrete trajectory at every actual time index. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.predictableLossAt","label":"predictableLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.predictableLossAt","description":"Predictable loss vector selected before the action at every actual time index.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-b328bb459b3a","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4574,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1018"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def predictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"predictablelossat banditrlproof.exp3.predictablelossat predictable loss vector selected before the action at every actual time index. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.observedImportanceWeightedLossAt","label":"observedImportanceWeightedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.observedImportanceWeightedLossAt","description":"The scalar-feedback importance-weighted coordinate used at an actual time.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-97f872266c7d","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4575,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1029"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def observedImportanceWeightedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (comparator : Action) : Real","missing":[],"search":"observedimportanceweightedlossat banditrlproof.exp3.observedimportanceweightedlossat the scalar-feedback importance-weighted coordinate used at an actual time. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt","label":"observedMixedSquaredImportanceWeightedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt","description":"The scalar-feedback probability-mixed estimator square used at an actual time.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-35648c9e385d","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4576,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1038"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def observedMixedSquaredImportanceWeightedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"observedmixedsquaredimportanceweightedlossat banditrlproof.exp3.observedmixedsquaredimportanceweightedlossat the scalar-feedback probability-mixed estimator square used at an actual time. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilitySourceAt","label":"sampledTrajectoryProbabilitySourceAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryProbabilitySourceAt","description":"Measurable finite-action source for the sampled probability vector at any time.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-eec01506b696","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4577,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1047"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryProbabilitySourceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) : MeasurableFiniteActionDistribution arms (sampledTrajectoryProbabilityAt (Env := Env) arms eta gamma t)","missing":[],"search":"sampledtrajectoryprobabilitysourceat banditrlproof.exp3.sampledtrajectoryprobabilitysourceat measurable finite-action source for the sampled probability vector at any time. definition compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryLossRegularityAt","label":"sampledPredictableTrajectoryLossRegularityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableTrajectoryLossRegularityAt","description":"Predictable trajectory losses satisfy one uniform regularity interface at every time.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-be244618598d","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4578,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1078"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableTrajectoryLossRegularityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : BoundedMeasurableLossWithProbabilityFloor arms (sampledTrajectoryProbabilityAt (Env := Env) arms eta gamma t) (predictableLossAt loss t) (gamma / (arms.card : Real))","missing":[],"search":"sampledpredictabletrajectorylossregularityat banditrlproof.exp3.sampledpredictabletrajectorylossregularityat predictable trajectory losses satisfy one uniform regularity interface at every time. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_predictableLossAt","label":"measurable_predictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_predictableLossAt","description":"theorem measurable_predictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (action : Action) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => predictableLossAt loss t sample action)","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-0b420e6dabb6","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4579,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1128"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_predictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (action : Action) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => predictableLossAt loss t sample action)","missing":[],"search":"measurable_predictablelossat banditrlproof.exp3.measurable_predictablelossat theorem measurable_predictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] (loss : predictablelossvector env action) (t : nat) (action : action) : measurable (fun sample : env × ((k : nat) -> action × real) => predictablelossat loss t sample action) theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_observedImportanceWeightedLossAt","label":"measurable_observedImportanceWeightedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_observedImportanceWeightedLossAt","description":"theorem measurable_observedImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : Measurable (observedImportanceWeightedLo…","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-7fc5fc1eeef9","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4580,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1146"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_observedImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : Measurable (observedImportanceWeightedLossAt (Env := Env) arms eta gamma t · comparator)","missing":[],"search":"measurable_observedimportanceweightedlossat banditrlproof.exp3.measurable_observedimportanceweightedlossat theorem measurable_observedimportanceweightedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : nat) (comparator : action) (hcomparator : comparator ∈ arms) : measurable (observedimportanceweightedlossat (env := env) arms eta gamma t · comparator) theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_observedMixedSquaredImportanceWeightedLossAt","label":"measurable_observedMixedSquaredImportanceWeightedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_observedMixedSquaredImportanceWeightedLossAt","description":"theorem measurable_observedMixedSquaredImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) : Measurable (observedMixedSquaredImportanceWeightedLossAt (Env := Env) arms eta gamma…","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-acea796c77c4","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4581,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1164"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_observedMixedSquaredImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) : Measurable (observedMixedSquaredImportanceWeightedLossAt (Env := Env) arms eta gamma t)","missing":[],"search":"measurable_observedmixedsquaredimportanceweightedlossat banditrlproof.exp3.measurable_observedmixedsquaredimportanceweightedlossat theorem measurable_observedmixedsquaredimportanceweightedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : nat) : measurable (observedmixedsquaredimportanceweightedlossat (env := env) arms eta gamma t) theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_predictableImportanceWeightedLossAt","label":"integrable_predictableImportanceWeightedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_predictableImportanceWeightedLossAt","description":"theorem integrable_predictableImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Act…","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-1d193fcc9708","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4582,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1184"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_predictableImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let prob := sampledTrajectoryProbabilityAt (Env := Env) arms eta gamma t let roundLoss := predictableLossAt loss t Integrable (fun sample => importanceWeightedLoss (prob sample) (roundLoss sample) (sample.2 t).1 comparator) mu","missing":[],"search":"integrable_predictableimportanceweightedlossat banditrlproof.exp3.integrable_predictableimportanceweightedlossat theorem integrable_predictableimportanceweightedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) → action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) (comparator : action) (hcomparator : comparator ∈ arms) : let prob := sampledtrajectoryprobabilityat (env := env) arms eta gamma t let roundloss := predictablelossat loss t integrable (fun sample => importanceweightedloss (prob sample) (roundloss sample) (sample.2 t).1 comparator) mu theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_predictableMixedSquaredImportanceWeightedLossAt","label":"integrable_predictableMixedSquaredImportanceWeightedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_predictableMixedSquaredImportanceWeightedLossAt","description":"theorem integrable_predictableMixedSquaredImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVe…","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-5d67abfe927e","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4583,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1217"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_predictableMixedSquaredImportanceWeightedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let prob := sampledTrajectoryProbabilityAt (Env := Env) arms eta gamma t let roundLoss := predictableLossAt loss t Integrable (fun sample => mixedSquaredImportanceWeightedLoss arms (prob sample) (roundLoss sample) (sample.2 t).1) mu","missing":[],"search":"integrable_predictablemixedsquaredimportanceweightedlossat banditrlproof.exp3.integrable_predictablemixedsquaredimportanceweightedlossat theorem integrable_predictablemixedsquaredimportanceweightedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) → action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : let prob := sampledtrajectoryprobabilityat (env := env) arms eta gamma t let roundloss := predictablelossat loss t integrable (fun sample => mixedsquaredimportanceweightedloss arms (prob sample) (roundloss sample) (sample.2 t).1) mu theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_predictableLossAt","label":"integrable_predictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_predictableLossAt","description":"theorem integrable_predictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (loss : PredictableLossVector Env Action) (t : Nat) (action : Action) : Integrable (fun sample => predictableLossAt loss t sample action) mu","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-4d71bf7770c1","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4584,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1249"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_predictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (loss : PredictableLossVector Env Action) (t : Nat) (action : Action) : Integrable (fun sample => predictableLossAt loss t sample action) mu","missing":[],"search":"integrable_predictablelossat banditrlproof.exp3.integrable_predictablelossat theorem integrable_predictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] (mu : measure (env × ((k : nat) → action × real))) [isfinitemeasure mu] (loss : predictablelossvector env action) (t : nat) (action : action) : integrable (fun sample => predictablelossat loss t sample action) mu theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_predictableLossSqSumAt","label":"integrable_predictableLossSqSumAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_predictableLossSqSumAt","description":"theorem integrable_predictableLossSqSumAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (fun sample => arms.sum (fun action => (predictableLossAt loss t sample action) ^ 2)) mu","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-ba9b55596b1d","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4585,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1269"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_predictableLossSqSumAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (fun sample => arms.sum (fun action => (predictableLossAt loss t sample action) ^ 2)) mu","missing":[],"search":"integrable_predictablelosssqsumat banditrlproof.exp3.integrable_predictablelosssqsumat theorem integrable_predictablelosssqsumat {env : type u} {action : type v} [measurablespace env] [measurablespace action] (mu : measure (env × ((k : nat) → action × real))) [isfinitemeasure mu] (arms : finset action) (loss : predictablelossvector env action) (t : nat) : integrable (fun sample => arms.sum (fun action => (predictablelossat loss t sample action) ^ 2)) mu theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.observedAt_eq_predictableAt_ae","label":"observedAt_eq_predictableAt_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.observedAt_eq_predictableAt_ae","description":"On the generated predictable trajectory, observed scalar scores agree almost everywhere with their latent predictable-loss counterparts.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-f82f04355d5a","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4586,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1297"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem observedAt_eq_predictableAt_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (comparator : Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ((fun sample => observedImportanceWeightedLossAt arms eta gamma t sample comparator) =ᵐ[mu] (fun sample => importanceWeightedLoss (sampledTrajectoryProbabilityAt arms eta gamma t sample) (predictableLossAt loss t sample) (sample.2 t).1 comparator)) ∧ ((fun sample => observedMixedS…","missing":[],"search":"observedat_eq_predictableat_ae banditrlproof.exp3.observedat_eq_predictableat_ae on the generated predictable trajectory, observed scalar scores agree almost everywhere with their latent predictable-loss counterparts. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_observedAt","label":"integrable_observedAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_observedAt","description":"The two observed score families are integrable under the generated predictable trajectory law at every actual time index.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-cafdae7832d2","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4587,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1375"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_observedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Integrable (fun sample => observedImportanceWeightedLossAt arms eta gamma t sample comparator) mu ∧ Integrable (fun sample => observedMixedSquaredImportanceWeightedLossAt arms eta gamma t sample) mu","missing":[],"search":"integrable_observedat banditrlproof.exp3.integrable_observedat the two observed score families are integrable under the generated predictable trajectory law at every actual time index. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedAt_first_second_moment","label":"sampledPredictableObservedAt_first_second_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedAt_first_second_moment","description":"At every actual time, including time zero, the observed armwise first moment and probability-mixed estimator-square moment equal the corresponding predictable loss-vector moments on the common full trajectory law.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-0ebdf8a94385","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4588,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1408"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedAt_first_second_moment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment (integral mu (fun sample => observedImportanceWeightedLossAt arms eta gamma t sample comparator) = integral mu (fun sample => predictableLossAt loss t sample comparator)) ∧ (integral mu (fun sample => observedMixedSquaredImportanceWe…","missing":[],"search":"sampledpredictableobservedat_first_second_moment banditrlproof.exp3.sampledpredictableobservedat_first_second_moment at every actual time, including time zero, the observed armwise first moment and probability-mixed estimator-square moment equal the corresponding predictable loss-vector moments on the common full trajectory law. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObserved_finiteHorizon_first_second_moment","label":"sampledPredictableObserved_finiteHorizon_first_second_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObserved_finiteHorizon_first_second_moment","description":"Finite-horizon observed EXP3 first and mixed-second moments over `t < horizon`; when the horizon is positive this range includes time zero.","url":"../modules/banditrlproof-exp3predictablemoments/index.html#decl-6fab7b0624da","parent":"module:BanditRLProof.Exp3PredictableMoments","order":4589,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableMoments"],["Source","BanditRLProof/Exp3PredictableMoments.lean:1553"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObserved_finiteHorizon_first_second_moment {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment (integral mu (fun sample => (Finset.range horizon).sum (fun t => observedImportanceWeightedLossAt arms eta gamma t sample comparator)) = integral mu (fun sample => (Finset.range horizon).sum (fun t => predictableLos…","missing":[],"search":"sampledpredictableobserved_finitehorizon_first_second_moment banditrlproof.exp3.sampledpredictableobserved_finitehorizon_first_second_moment finite-horizon observed exp3 first and mixed-second moments over `t < horizon`; when the horizon is positive this range includes time zero. theorem compiled","shard":"modules/43b280a37cc70ee8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRegretGeometricAllTimeBudget","label":"sampledPredictableRegretGeometricAllTimeBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRegretGeometricAllTimeBudget","description":"The fixed-horizon predictable-regret budget at prefix `n+1`. The outer geometric share is divided by two because the parent total-delta theorem allocates equal shares to its pure-cross and comparator-estimator events.","url":"../modules/banditrlproof-exp3predictableregretalltime/index.html#decl-0b4310ee45a4","parent":"module:BanditRLProof.Exp3PredictableRegretAllTime","order":4590,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableRegretAllTime"],["Source","BanditRLProof/Exp3PredictableRegretAllTime.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRegretGeometricAllTimeBudget {Action : Type v} (arms : Finset Action) (eta gamma delta : Real) (n : Nat) : Real","missing":[],"search":"sampledpredictableregretgeometricalltimebudget banditrlproof.exp3.sampledpredictableregretgeometricalltimebudget the fixed-horizon predictable-regret budget at prefix `n+1`. the outer geometric share is divided by two because the parent total-delta theorem allocates equal shares to its pure-cross and comparator-estimator events. definition compiled","shard":"modules/092d1623420c79bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRegretGeometricAllTimeFailureSet","label":"sampledPredictableRegretGeometricAllTimeFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRegretGeometricAllTimeFailureSet","description":"Countable predictable-regret failure event over every positive prefix of one generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3predictableregretalltime/index.html#decl-d2e1863f50bd","parent":"module:BanditRLProof.Exp3PredictableRegretAllTime","order":4591,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PredictableRegretAllTime"],["Source","BanditRLProof/Exp3PredictableRegretAllTime.lean:32"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRegretGeometricAllTimeFailureSet {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (delta : Real) : Set (Env × ((k : Nat) -> Action × Real))","missing":[],"search":"sampledpredictableregretgeometricalltimefailureset banditrlproof.exp3.sampledpredictableregretgeometricalltimefailureset countable predictable-regret failure event over every positive prefix of one generated exp3 trajectory. definition compiled","shard":"modules/092d1623420c79bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mem_sampledPredictableRegretGeometricAllTimeFailureSet_iff","label":"mem_sampledPredictableRegretGeometricAllTimeFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mem_sampledPredictableRegretGeometricAllTimeFailureSet_iff","description":"Membership is predictable-regret failure at at least one positive prefix.","url":"../modules/banditrlproof-exp3predictableregretalltime/index.html#decl-150f3aee8953","parent":"module:BanditRLProof.Exp3PredictableRegretAllTime","order":4592,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableRegretAllTime"],["Source","BanditRLProof/Exp3PredictableRegretAllTime.lean:49"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mem_sampledPredictableRegretGeometricAllTimeFailureSet_iff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (delta : Real) (sample : Env × ((k : Nat) -> Action × Real)) : sample ∈ sampledPredictableRegretGeometricAllTimeFailureSet arms eta gamma loss comparator delta ↔ ∃ n, sampledPredictableRegretGeometricAllTimeBudget arms eta gamma delta n <= (Finset.range (n + 1)).sum (fun t => sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample) - (Finset.range (n + 1)).sum (fun t => predictableLossAt loss t sample comparator)","missing":[],"search":"mem_sampledpredictableregretgeometricalltimefailureset_iff banditrlproof.exp3.mem_sampledpredictableregretgeometricalltimefailureset_iff membership is predictable-regret failure at at least one positive prefix. theorem compiled","shard":"modules/092d1623420c79bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","label":"measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","description":"On one fixed generated EXP3 process and against one fixed supported comparator, predictable-regret failures over all positive prefixes have outer measure at most the geometric confidence budget.","url":"../modules/banditrlproof-exp3predictableregretalltime/index.html#decl-5fe4c02a5855","parent":"module:BanditRLProof.Exp3PredictableRegretAllTime","order":4593,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PredictableRegretAllTime"],["Source","BanditRLProof/Exp3PredictableRegretAllTime.lean:70"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measure_sampledPredictableRegretGeometricAllTimeFailureSet_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu (sampledPredictableRegretGeometricAllTimeFailureSet arms eta gamma loss comparator delta) <= ENNReal.ofReal delta","missing":[],"search":"measure_sampledpredictableregretgeometricalltimefailureset_le banditrlproof.exp3.measure_sampledpredictableregretgeometricalltimefailureset_le on one fixed generated exp3 process and against one fixed supported comparator, predictable-regret failures over all positive prefixes have outer measure at most the geometric confidence budget. theorem compiled","shard":"modules/092d1623420c79bb.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weightedImportanceWeightedLoss_eq_selected","label":"weightedImportanceWeightedLoss_eq_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.weightedImportanceWeightedLoss_eq_selected","description":"On the finite support, the cross-weighted estimator has only the sampled coordinate as a nonzero summand.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-b679739b87a3","parent":"module:BanditRLProof.Exp3PureBernstein","order":4594,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem weightedImportanceWeightedLoss_eq_selected {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob weight loss : Action -> Real) (chosen : Action) (hchosen : chosen ∈ arms) : weightedImportanceWeightedLoss arms prob weight loss chosen = weight chosen * loss chosen / prob chosen","missing":[],"search":"weightedimportanceweightedloss_eq_selected banditrlproof.exp3.weightedimportanceweightedloss_eq_selected on the finite support, the cross-weighted estimator has only the sampled coordinate as a nonzero summand. theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_weightedEstimatorMeanMinusRaw_le_inv_floor","label":"sum_prob_mul_sq_weightedEstimatorMeanMinusRaw_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_prob_mul_sq_weightedEstimatorMeanMinusRaw_le_inv_floor","description":"The centered pure-Hedge cross-weighted estimator has second moment at most the reciprocal exploration floor.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-d8c09ef13f1f","parent":"module:BanditRLProof.Exp3PureBernstein","order":4595,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:38"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_sq_weightedEstimatorMeanMinusRaw_le_inv_floor {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob weight loss : Action -> Real) (hprob : FiniteActionDistribution arms prob) (hweight : FiniteActionDistribution arms weight) (epsilon : Real) (hepsilon : 0 < epsilon) (hfloor : forall action, action ∈ arms -> epsilon <= prob action) (hloss : forall action, action ∈ arms -> loss action ∈ Set.Icc (0 : Real) 1) : let mean := arms.sum (fun action => weight action * loss action) arms.sum (fun chosen => prob chosen * (mean - weightedImportanceWeightedLoss arms prob weight loss chosen) ^ 2) <= 1 / epsilon","missing":[],"search":"sum_prob_mul_sq_weightedestimatormeanminusraw_le_inv_floor banditrlproof.exp3.sum_prob_mul_sq_weightedestimatormeanminusraw_le_inv_floor the centered pure-hedge cross-weighted estimator has second moment at most the reciprocal exploration floor. theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionWeightedEstimatorMeanMinusRaw_hasMGFUpperBoundAt","label":"finiteActionWeightedEstimatorMeanMinusRaw_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionWeightedEstimatorMeanMinusRaw_hasMGFUpperBoundAt","description":"Fixed-tilt MGF budget for the sign used by the pure-Hedge regret route.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-cb4ee4e6f569","parent":"module:BanditRLProof.Exp3PureBernstein","order":4596,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:118"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionWeightedEstimatorMeanMinusRaw_hasMGFUpperBoundAt {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (prob weight loss : Action -> Real) (hprob : FiniteActionDistribution arms prob) (hweight : FiniteActionDistribution arms weight) (epsilon : Real) (hepsilon : 0 < epsilon) (hfloor : forall action, action ∈ arms -> epsilon <= prob action) (hloss : forall action, action ∈ arms -> loss action ∈ Set.Icc (0 : Real) 1) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= epsilon) : let mean := arms.sum (fun action => weight action * loss action) Concentration.HasMGFUpperBoundAt (fun chosen => mean - weightedImportanceWeightedLoss arms prob weight loss chosen) tilt (tilt ^ 2 / epsilon) (finiteActionMeasure arms prob)","missing":[],"search":"finiteactionweightedestimatormeanminusraw_hasmgfupperboundat banditrlproof.exp3.finiteactionweightedestimatormeanminusraw_hasmgfupperboundat fixed-tilt mgf budget for the sign used by the pure-hedge regret route. theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weightedEstimatorMeanMinusRaw_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","label":"weightedEstimatorMeanMinusRaw_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.weightedEstimatorMeanMinusRaw_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","description":"An identified finite conditional action law transports the variance-sensitive fixed-tilt budget for `mean - cross-weighted estimator`.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-278356e663e6","parent":"module:BanditRLProof.Exp3PureBernstein","order":4597,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:260"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem weightedEstimatorMeanMinusRaw_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (harms : arms.Nonempty) (prob weight loss : History -> Action -> Real) (probSource : MeasurableFiniteActionDistribution arms prob) (weightSource : MeasurableFiniteActionDistribution arms weight) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss…","missing":[],"search":"weightedestimatormeanminusraw_hascondmgfupperboundat_of_conddistrib_ae_eq_finiteactionkernel banditrlproof.exp3.weightedestimatormeanminusraw_hascondmgfupperboundat_of_conddistrib_ae_eq_finiteactionkernel an identified finite conditional action law transports the variance-sensitive fixed-tilt budget for `mean - cross-weighted estimator`. theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableMinusWeightedAt","label":"sampledTrajectoryPurePredictableMinusWeightedAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPurePredictableMinusWeightedAt","description":"Latent predictable form of the sign-correct pure-Hedge deviation.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-b99efe92707d","parent":"module:BanditRLProof.Exp3PureBernstein","order":4598,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:412"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPurePredictableMinusWeightedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypurepredictableminusweightedat banditrlproof.exp3.sampledtrajectorypurepredictableminusweightedat latent predictable form of the sign-correct pure-hedge deviation. definition compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusWeighted_zero_hasCondMGFUpperBoundAt","label":"sampledPurePredictableMinusWeighted_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusWeighted_zero_hasCondMGFUpperBoundAt","description":"theorem sampledPurePredictableMinusWeighted_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos :…","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-402918d45818","parent":"module:BanditRLProof.Exp3PureBernstein","order":4599,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:424"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusWeighted_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTra…","missing":[],"search":"sampledpurepredictableminusweighted_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpurepredictableminusweighted_zero_hascondmgfupperboundat theorem sampledpurepredictableminusweighted_zero_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment concentration.hascondmgfupperboundat ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypurepredictableminusweightedat arms eta gamma loss 0) tilt (tilt ^ 2 / (gamma / (arms.card : real))) mu theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusWeighted_succ_hasCondMGFUpperBoundAt","label":"sampledPurePredictableMinusWeighted_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusWeighted_succ_hasCondMGFUpperBoundAt","description":"theorem sampledPurePredictableMinusWeighted_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos :…","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-0d20cab74347","parent":"module:BanditRLProof.Exp3PureBernstein","order":4600,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:488"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusWeighted_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Concentration.HasCondMGFUpperBoundAt ((inferInstance : Measu…","missing":[],"search":"sampledpurepredictableminusweighted_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpurepredictableminusweighted_succ_hascondmgfupperboundat theorem sampledpurepredictableminusweighted_succ_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) concentration.hascondmgfupperboundat ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypurepredictableminusweightedat arms eta gamma loss (n + 1)) tilt (tilt ^ 2 / (gamma / (arms.card : real))) mu theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_zero_hasCondMGFUpperBoundAt","label":"sampledPurePredictableMinusObserved_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_zero_hasCondMGFUpperBoundAt","description":"theorem sampledPurePredictableMinusObserved_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos :…","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-e6bc79192b97","parent":"module:BanditRLProof.Exp3PureBernstein","order":4601,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:565"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTra…","missing":[],"search":"sampledpurepredictableminusobserved_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpurepredictableminusobserved_zero_hascondmgfupperboundat theorem sampledpurepredictableminusobserved_zero_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment concentration.hascondmgfupperboundat ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypurepredictableminusobservedat arms eta gamma loss 0) tilt (tilt ^ 2 / (gamma / (arms.card : real))) mu theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_succ_hasCondMGFUpperBoundAt","label":"sampledPurePredictableMinusObserved_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_succ_hasCondMGFUpperBoundAt","description":"theorem sampledPurePredictableMinusObserved_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos :…","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-d437b746103c","parent":"module:BanditRLProof.Exp3PureBernstein","order":4602,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:625"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Concentration.HasCondMGFUpperBoundAt ((inferInstance : Measu…","missing":[],"search":"sampledpurepredictableminusobserved_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpurepredictableminusobserved_succ_hascondmgfupperboundat theorem sampledpurepredictableminusobserved_succ_hascondmgfupperboundat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) (tilt : real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : real)) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) concentration.hascondmgfupperboundat ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypurepredictableminusobservedat arms eta gamma loss (n + 1)) tilt (tilt ^ 2 / (gamma / (arms.card : real))) mu theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_fixedTilt","label":"sampledPurePredictableMinusObserved_sum_tail_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_fixedTilt","description":"Variance-sensitive fixed-tilt tail for the pure-Hedge predictable loss minus its observed cross-weighted estimator.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-00de6aff7419","parent":"module:BanditRLProof.Exp3PureBernstein","order":4603,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:696"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_sum_tail_fixedTilt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (tilt : Real) (htilt_nonneg : 0 <= tilt) (htilt_le : tilt <= gamma / (arms.card : Real)) (threshold : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu.real {sample | threshold <= (Finset.range horizon).sum (fun i => sampledTrajectoryPurePredictableMinusObservedAt arms eta gamma loss i sample)} <=…","missing":[],"search":"sampledpurepredictableminusobserved_sum_tail_fixedtilt banditrlproof.exp3.sampledpurepredictableminusobserved_sum_tail_fixedtilt variance-sensitive fixed-tilt tail for the pure-hedge predictable loss minus its observed cross-weighted estimator. theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedBernsteinConfidenceRadius","label":"sampledPurePredictableMinusObservedBernsteinConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObservedBernsteinConfidenceRadius","description":"Variance-sensitive confidence radius for the pure-Hedge cross-weighted deviation in the sign consumed by the regret decomposition.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-6706af481f1a","parent":"module:BanditRLProof.Exp3PureBernstein","order":4604,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:773"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPurePredictableMinusObservedBernsteinConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpurepredictableminusobservedbernsteinconfidenceradius banditrlproof.exp3.sampledpurepredictableminusobservedbernsteinconfidenceradius variance-sensitive confidence radius for the pure-hedge cross-weighted deviation in the sign consumed by the regret decomposition. definition compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_bernstein_delta","label":"sampledPurePredictableMinusObserved_sum_tail_bernstein_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_bernstein_delta","description":"Delta-shaped variance-sensitive confidence bound for the pure-Hedge predictable loss minus its observed cross-weighted estimator.","url":"../modules/banditrlproof-exp3purebernstein/index.html#decl-9b2e7f56dad8","parent":"module:BanditRLProof.Exp3PureBernstein","order":4605,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureBernstein"],["Source","BanditRLProof/Exp3PureBernstein.lean:783"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_sum_tail_bernstein_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledPurePredictableMinusObservedBernsteinConfidenceRadius arms gamma horizon delta <= (Finset.range horizon).sum (fun i => sampledTrajectoryPurePredictableLossAt arms eta gamma loss i sample - sample…","missing":[],"search":"sampledpurepredictableminusobserved_sum_tail_bernstein_delta banditrlproof.exp3.sampledpurepredictableminusobserved_sum_tail_bernstein_delta delta-shaped variance-sensitive confidence bound for the pure-hedge predictable loss minus its observed cross-weighted estimator. theorem compiled","shard":"modules/10b14b64d48b6652.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.weightedEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","label":"weightedEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.weightedEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","description":"theorem weightedEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-50b2e01b53d4","parent":"module:BanditRLProof.Exp3PureConfidence","order":4606,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem weightedEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel {Omega : Type u} {History : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob weight loss : History -> Action -> Real) (probSource : MeasurableFiniteActionDistribution arms prob) (weightSource : MeasurableFiniteActionDistribution arms weight) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (hcond : condDistrib action…","missing":[],"search":"weightedestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel banditrlproof.exp3.weightedestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel theorem weightedestimator_hascondsubgaussianmgf_of_conddistrib_ae_eq_finiteactionkernel {omega : type u} {history : type v} {action : type w} [momega : measurablespace omega] [standardborelspace omega] [nonempty omega] [mhistory : measurablespace history] [standardborelspace history] [maction : measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (mu : measure omega) [isfinitemeasure mu] (history : omega -> history) (hhistory : measurable history) (action : omega -> action) (haction : measurable action) (arms : finset action) (prob weight loss : history -> action -> real) (probsource : measurablefiniteactiondistribution arms prob) (weightsource : measurablefiniteactiondistribution arms weight) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) (hcond : conddistrib action history mu =ᵐ[mu.map history] finiteactionkernel arms prob probsource) : probabilitytheory.hascondsubgaussianmgf (mhistory.comap history) hhistory.comap_le (fun omega => weightedimportanceweightedloss arms (prob (history omega)) (weight (history omega)) (loss (history omega)) (action omega) - arms.sum (fun candidate =>…","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryWeightedPurePredictableDeviationAt","label":"sampledTrajectoryWeightedPurePredictableDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryWeightedPurePredictableDeviationAt","description":"noncomputable def sampledTrajectoryWeightedPurePredictableDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-21621e6c465a","parent":"module:BanditRLProof.Exp3PureConfidence","order":4607,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:193"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryWeightedPurePredictableDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectoryweightedpurepredictabledeviationat banditrlproof.exp3.sampledtrajectoryweightedpurepredictabledeviationat noncomputable def sampledtrajectoryweightedpurepredictabledeviationat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : real definition compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureObservedDeviationAt","label":"sampledTrajectoryPureObservedDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPureObservedDeviationAt","description":"noncomputable def sampledTrajectoryPureObservedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-9923cdec4399","parent":"module:BanditRLProof.Exp3PureConfidence","order":4608,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:205"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPureObservedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypureobserveddeviationat banditrlproof.exp3.sampledtrajectorypureobserveddeviationat noncomputable def sampledtrajectorypureobserveddeviationat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : real definition compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledWeightedPurePredictableDeviation_zero_hasCondSubgaussianMGF","label":"sampledWeightedPurePredictableDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledWeightedPurePredictableDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledWeightedPurePredictableDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-e8783367e7c4","parent":"module:BanditRLProof.Exp3PureConfidence","order":4609,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:214"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledWeightedPurePredictableDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryWeightedPurePredictableDeviationAt arms eta gamma loss 0) (sampledComparator…","missing":[],"search":"sampledweightedpurepredictabledeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledweightedpurepredictabledeviation_zero_hascondsubgaussianmgf theorem sampledweightedpurepredictabledeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectoryweightedpurepredictabledeviationat arms eta gamma loss 0) (sampledcomparatorestimatorvarianceproxy arms gamma) mu theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledWeightedPurePredictableDeviation_succ_hasCondSubgaussianMGF","label":"sampledWeightedPurePredictableDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledWeightedPurePredictableDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledWeightedPurePredictableDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-fab92310e7c1","parent":"module:BanditRLProof.Exp3PureConfidence","order":4610,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:275"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledWeightedPurePredictableDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real n)).comap history) (measura…","missing":[],"search":"sampledweightedpurepredictabledeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledweightedpurepredictabledeviation_succ_hascondsubgaussianmgf theorem sampledweightedpurepredictabledeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectoryweightedpurepredictabledeviationat arms eta gamma loss (n + 1)) (sampledcomparatorestimatorvarianceproxy arms gamma) mu theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_zero_hasCondSubgaussianMGF","label":"sampledPureObservedDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledPureObservedDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamm…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-95392e270ba4","parent":"module:BanditRLProof.Exp3PureConfidence","order":4611,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:349"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryPureObservedDeviationAt arms eta gamma loss 0) (sampledComparatorEstimatorVarianceProxy…","missing":[],"search":"sampledpureobserveddeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledpureobserveddeviation_zero_hascondsubgaussianmgf theorem sampledpureobserveddeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypureobserveddeviationat arms eta gamma loss 0) (sampledcomparatorestimatorvarianceproxy arms gamma) mu theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_succ_hasCondSubgaussianMGF","label":"sampledPureObservedDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledPureObservedDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamm…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-35ec85bbc58d","parent":"module:BanditRLProof.Exp3PureConfidence","order":4612,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:404"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real n)).comap history) (measurable_fst.pro…","missing":[],"search":"sampledpureobserveddeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledpureobserveddeviation_succ_hascondsubgaussianmgf theorem sampledpureobserveddeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypureobserveddeviationat arms eta gamma loss (n + 1)) (sampledcomparatorestimatorvarianceproxy arms gamma) mu theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProcess","label":"sampledPureObservedDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviationProcess","description":"noncomputable def sampledPureObservedDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryPureObservedDeviationAt arms eta gamma loss i sample theorem sampledPureOb…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-110f711fc515","parent":"module:BanditRLProof.Exp3PureConfidence","order":4613,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:467"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPureObservedDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryPureObservedDeviationAt arms eta gamma loss i sample theorem sampledPureObservedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPureObservedDeviationProcess arms eta gamma loss)","missing":[],"search":"sampledpureobserveddeviationprocess banditrlproof.exp3.sampledpureobserveddeviationprocess noncomputable def sampledpureobserveddeviationprocess {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) : nat -> env × ((k : nat) -> action × real) -> real | 0, _sample => 0 | i + 1, sample => sampledtrajectorypureobserveddeviationat arms eta gamma loss i sample theorem sampledpureobserveddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpureobserveddeviationprocess arms eta gamma loss) definition compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProcess_stronglyAdapted","label":"sampledPureObservedDeviationProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviationProcess_stronglyAdapted","description":"theorem sampledPureObservedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration E…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-79a5ceb3e0c3","parent":"module:BanditRLProof.Exp3PureConfidence","order":4614,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:477"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPureObservedDeviationProcess arms eta gamma loss)","missing":[],"search":"sampledpureobserveddeviationprocess_stronglyadapted banditrlproof.exp3.sampledpureobserveddeviationprocess_stronglyadapted theorem sampledpureobserveddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpureobserveddeviationprocess arms eta gamma loss) theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationVarianceProxy","label":"sampledPureObservedDeviationVarianceProxy","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviationVarianceProxy","description":"noncomputable abbrev sampledPureObservedDeviationVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-be9653c1cdb3","parent":"module:BanditRLProof.Exp3PureConfidence","order":4615,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:668"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable abbrev sampledPureObservedDeviationVarianceProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : NNReal","missing":[],"search":"sampledpureobserveddeviationvarianceproxy banditrlproof.exp3.sampledpureobserveddeviationvarianceproxy noncomputable abbrev sampledpureobserveddeviationvarianceproxy {action : type v} (arms : finset action) (gamma : real) : nnreal abbreviation compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProxy","label":"sampledPureObservedDeviationProxy","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviationProxy","description":"noncomputable abbrev sampledPureObservedDeviationProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : Nat -> NNReal","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-0cde78bbf20e","parent":"module:BanditRLProof.Exp3PureConfidence","order":4616,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:672"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable abbrev sampledPureObservedDeviationProxy {Action : Type v} (arms : Finset Action) (gamma : Real) : Nat -> NNReal","missing":[],"search":"sampledpureobserveddeviationproxy banditrlproof.exp3.sampledpureobserveddeviationproxy noncomputable abbrev sampledpureobserveddeviationproxy {action : type v} (arms : finset action) (gamma : real) : nat -> nnreal abbreviation compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProcess_sum_range_succ","label":"sampledPureObservedDeviationProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviationProcess_sum_range_succ","description":"theorem sampledPureObservedDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPureObservedDeviationProcess arms eta gamma loss i sample) =…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-1fca2cf0595d","parent":"module:BanditRLProof.Exp3PureConfidence","order":4617,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:676"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPureObservedDeviationProcess arms eta gamma loss i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryPureObservedDeviationAt arms eta gamma loss i sample)","missing":[],"search":"sampledpureobserveddeviationprocess_sum_range_succ banditrlproof.exp3.sampledpureobserveddeviationprocess_sum_range_succ theorem sampledpureobserveddeviationprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpureobserveddeviationprocess arms eta gamma loss i sample) = (finset.range horizon).sum (fun i => sampledtrajectorypureobserveddeviationat arms eta gamma loss i sample) theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_ennreal","label":"sampledPureObservedDeviation_sum_tail_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_ennreal","description":"One-sided concentration for the pure-Hedge cross-weighted observed EXP3 estimator minus its true pure-Hedge predictable loss.","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-3c34f249edd4","parent":"module:BanditRLProof.Exp3PureConfidence","order":4618,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:713"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviation_sum_tail_ennreal {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | eps <= (Finset.range horizon).sum (fun i => sampledTrajectoryPureObservedDeviationAt arms eta gamma loss i sample)} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((((horizon : NNReal) * sampledPureObservedDeviationVarianceP…","missing":[],"search":"sampledpureobserveddeviation_sum_tail_ennreal banditrlproof.exp3.sampledpureobserveddeviation_sum_tail_ennreal one-sided concentration for the pure-hedge cross-weighted observed exp3 estimator minus its true pure-hedge predictable loss. theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationConfidenceRadius","label":"sampledPureObservedDeviationConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviationConfidenceRadius","description":"noncomputable def sampledPureObservedDeviationConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-38d7f2a19d42","parent":"module:BanditRLProof.Exp3PureConfidence","order":4619,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:784"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPureObservedDeviationConfidenceRadius {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpureobserveddeviationconfidenceradius banditrlproof.exp3.sampledpureobserveddeviationconfidenceradius noncomputable def sampledpureobserveddeviationconfidenceradius {action : type v} (arms : finset action) (gamma : real) (horizon : nat) (delta : real) : real definition compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_exp_neg_budget","label":"sampledPureObservedDeviation_sum_tail_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_exp_neg_budget","description":"theorem sampledPureObservedDeviation_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < ga…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-d2e901d6b0c0","parent":"module:BanditRLProof.Exp3PureConfidence","order":4620,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:792"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviation_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (budget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | Real.sqrt (2 * ((((horizon : NNReal) * sampledPureObservedDeviationVarianceProxy arms gamma : NNReal)) : Real) * budget) <= (Finset.range horizon).sum (fun i => sampledTrajectoryPureObservedDeviationAt arm…","missing":[],"search":"sampledpureobserveddeviation_sum_tail_exp_neg_budget banditrlproof.exp3.sampledpureobserveddeviation_sum_tail_exp_neg_budget theorem sampledpureobserveddeviation_sum_tail_exp_neg_budget {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) (hhorizon : 0 < horizon) (budget : real) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | real.sqrt (2 * ((((horizon : nnreal) * sampledpureobserveddeviationvarianceproxy arms gamma : nnreal)) : real) * budget) <= (finset.range horizon).sum (fun i => sampledtrajectorypureobserveddeviationat arms eta gamma loss i sample)} <= ennreal.ofreal (real.exp (-budget)) theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_delta","label":"sampledPureObservedDeviation_sum_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_delta","description":"Delta-shaped one-sided confidence bound for the pure-Hedge cross-weighted observed EXP3 estimator against its true pure-Hedge predictable loss.","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-9a82bc8325d0","parent":"module:BanditRLProof.Exp3PureConfidence","order":4621,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:841"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPureObservedDeviation_sum_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledPureObservedDeviationConfidenceRadius arms gamma horizon delta <= (Finset.range horizon).sum (fun i => sampledTrajectoryPureObservedLossAt arms eta gamma i sample - sampledTrajectoryPureP…","missing":[],"search":"sampledpureobserveddeviation_sum_tail_delta banditrlproof.exp3.sampledpureobserveddeviation_sum_tail_delta delta-shaped one-sided confidence bound for the pure-hedge cross-weighted observed exp3 estimator against its true pure-hedge predictable loss. theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableMinusObservedAt","label":"sampledTrajectoryPurePredictableMinusObservedAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPurePredictableMinusObservedAt","description":"The pure-Hedge predictable loss minus its observed cross-weighted estimator. This is the sign needed when the sampled Hedge inequality is converted to true predictable regret.","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-68fa39c36e2a","parent":"module:BanditRLProof.Exp3PureConfidence","order":4622,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:879"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPurePredictableMinusObservedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectorypurepredictableminusobservedat banditrlproof.exp3.sampledtrajectorypurepredictableminusobservedat the pure-hedge predictable loss minus its observed cross-weighted estimator. this is the sign needed when the sampled hedge inequality is converted to true predictable regret. definition compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_zero_hasCondSubgaussianMGF","label":"sampledPurePredictableMinusObserved_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_zero_hasCondSubgaussianMGF","description":"theorem sampledPurePredictableMinusObserved_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-3d83d2b592d2","parent":"module:BanditRLProof.Exp3PureConfidence","order":4623,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:888"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryPurePredictableMinusObservedAt arms eta gamma loss 0) (sampledPureObservedDeviat…","missing":[],"search":"sampledpurepredictableminusobserved_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledpurepredictableminusobserved_zero_hascondsubgaussianmgf theorem sampledpurepredictableminusobserved_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectorypurepredictableminusobservedat arms eta gamma loss 0) (sampledpureobserveddeviationvarianceproxy arms gamma) mu theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_succ_hasCondSubgaussianMGF","label":"sampledPurePredictableMinusObserved_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_succ_hasCondSubgaussianMGF","description":"theorem sampledPurePredictableMinusObserved_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-a89584924d24","parent":"module:BanditRLProof.Exp3PureConfidence","order":4624,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:917"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real n)).comap history) (measurable_…","missing":[],"search":"sampledpurepredictableminusobserved_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledpurepredictableminusobserved_succ_hascondsubgaussianmgf theorem sampledpurepredictableminusobserved_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectorypurepredictableminusobservedat arms eta gamma loss (n + 1)) (sampledpureobserveddeviationvarianceproxy arms gamma) mu theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess","label":"sampledPurePredictableMinusObservedProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess","description":"noncomputable def sampledPurePredictableMinusObservedProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryPurePredictableMinusObservedAt arms eta gamma loss i sample theorem…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-daa211282630","parent":"module:BanditRLProof.Exp3PureConfidence","order":4625,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:949"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPurePredictableMinusObservedProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryPurePredictableMinusObservedAt arms eta gamma loss i sample theorem sampledPurePredictableMinusObservedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPurePredictableMinusObservedProcess arms e…","missing":[],"search":"sampledpurepredictableminusobservedprocess banditrlproof.exp3.sampledpurepredictableminusobservedprocess noncomputable def sampledpurepredictableminusobservedprocess {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) : nat -> env × ((k : nat) -> action × real) -> real | 0, _sample => 0 | i + 1, sample => sampledtrajectorypurepredictableminusobservedat arms eta gamma loss i sample theorem sampledpurepredictableminusobservedprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpurepredictableminusobservedprocess arms eta gamma loss) definition compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess_stronglyAdapted","label":"sampledPurePredictableMinusObservedProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess_stronglyAdapted","description":"theorem sampledPurePredictableMinusObservedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltr…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-4c5dcd4104b5","parent":"module:BanditRLProof.Exp3PureConfidence","order":4626,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:960"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObservedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPurePredictableMinusObservedProcess arms eta gamma loss)","missing":[],"search":"sampledpurepredictableminusobservedprocess_stronglyadapted banditrlproof.exp3.sampledpurepredictableminusobservedprocess_stronglyadapted theorem sampledpurepredictableminusobservedprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpurepredictableminusobservedprocess arms eta gamma loss) theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess_sum_range_succ","label":"sampledPurePredictableMinusObservedProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess_sum_range_succ","description":"theorem sampledPurePredictableMinusObservedProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPurePredictableMinusObservedProcess arms eta gamma los…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-22d7d7829b63","parent":"module:BanditRLProof.Exp3PureConfidence","order":4627,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:987"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObservedProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPurePredictableMinusObservedProcess arms eta gamma loss i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryPurePredictableMinusObservedAt arms eta gamma loss i sample)","missing":[],"search":"sampledpurepredictableminusobservedprocess_sum_range_succ banditrlproof.exp3.sampledpurepredictableminusobservedprocess_sum_range_succ theorem sampledpurepredictableminusobservedprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpurepredictableminusobservedprocess arms eta gamma loss i sample) = (finset.range horizon).sum (fun i => sampledtrajectorypurepredictableminusobservedat arms eta gamma loss i sample) theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_ennreal","label":"sampledPurePredictableMinusObserved_sum_tail_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_ennreal","description":"One-sided concentration in the sign required by the predictable-regret decomposition: true pure-Hedge predictable loss minus its observed estimator.","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-2c566afe157b","parent":"module:BanditRLProof.Exp3PureConfidence","order":4628,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:1027"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_sum_tail_ennreal {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | eps <= (Finset.range horizon).sum (fun i => sampledTrajectoryPurePredictableMinusObservedAt arms eta gamma loss i sample)} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((((horizon : NNReal) * sampledPureObservedDevi…","missing":[],"search":"sampledpurepredictableminusobserved_sum_tail_ennreal banditrlproof.exp3.sampledpurepredictableminusobserved_sum_tail_ennreal one-sided concentration in the sign required by the predictable-regret decomposition: true pure-hedge predictable loss minus its observed estimator. theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_exp_neg_budget","label":"sampledPurePredictableMinusObserved_sum_tail_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_exp_neg_budget","description":"theorem sampledPurePredictableMinusObserved_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos :…","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-c92390b1ca31","parent":"module:BanditRLProof.Exp3PureConfidence","order":4629,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:1098"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (budget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | Real.sqrt (2 * ((((horizon : NNReal) * sampledPureObservedDeviationVarianceProxy arms gamma : NNReal)) : Real) * budget) <= (Finset.range horizon).sum (fun i => sampledTrajectoryPurePredictableMinus…","missing":[],"search":"sampledpurepredictableminusobserved_sum_tail_exp_neg_budget banditrlproof.exp3.sampledpurepredictableminusobserved_sum_tail_exp_neg_budget theorem sampledpurepredictableminusobserved_sum_tail_exp_neg_budget {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) (hhorizon : 0 < horizon) (budget : real) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | real.sqrt (2 * ((((horizon : nnreal) * sampledpureobserveddeviationvarianceproxy arms gamma : nnreal)) : real) * budget) <= (finset.range horizon).sum (fun i => sampledtrajectorypurepredictableminusobservedat arms eta gamma loss i sample)} <= ennreal.ofreal (real.exp (-budget)) theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_delta","label":"sampledPurePredictableMinusObserved_sum_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_delta","description":"Delta-shaped lower-tail counterpart in the sign needed by the generated predictable EXP3 regret decomposition.","url":"../modules/banditrlproof-exp3pureconfidence/index.html#decl-1ab22cfef15a","parent":"module:BanditRLProof.Exp3PureConfidence","order":4630,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3PureConfidence"],["Source","BanditRLProof/Exp3PureConfidence.lean:1147"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPurePredictableMinusObserved_sum_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledPureObservedDeviationConfidenceRadius arms gamma horizon delta <= (Finset.range horizon).sum (fun i => sampledTrajectoryPurePredictableLossAt arms eta gamma loss i sample - sampled…","missing":[],"search":"sampledpurepredictableminusobserved_sum_tail_delta banditrlproof.exp3.sampledpurepredictableminusobserved_sum_tail_delta delta-shaped lower-tail counterpart in the sign needed by the generated predictable exp3 regret decomposition. theorem compiled","shard":"modules/895c11c7bda6720a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinLargeHorizonCondition","label":"randomSquareBernsteinLargeHorizonCondition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinLargeHorizonCondition","description":"The regime in which the explicit random-square exploration schedule satisfies both confidence dominance contracts without activating its clip.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedallhorizon/index.html#decl-3b79b5576b0a","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","order":4631,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedAllHorizon.lean:22"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def randomSquareBernsteinLargeHorizonCondition (K T delta : Real) : Prop","missing":[],"search":"randomsquarebernsteinlargehorizoncondition banditrlproof.exp3.randomsquarebernsteinlargehorizoncondition the regime in which the explicit random-square exploration schedule satisfies both confidence dominance contracts without activating its clip. definition compiled","shard":"modules/52284718dbde5c44.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinAllHorizonRegretThreshold","label":"randomSquareBernsteinAllHorizonRegretThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinAllHorizonRegretThreshold","description":"All-horizon threshold for the random-square route: use the explicit large-horizon rate in its valid regime and `T + 1` otherwise.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedallhorizon/index.html#decl-3aa5f0c9796b","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","order":4632,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedAllHorizon.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinAllHorizonRegretThreshold {Action : Type v} (arms : Finset Action) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"randomsquarebernsteinallhorizonregretthreshold banditrlproof.exp3.randomsquarebernsteinallhorizonregretthreshold all-horizon threshold for the random-square route: use the explicit large-horizon rate in its valid regime and `t + 1` otherwise. definition compiled","shard":"modules/52284718dbde5c44.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonRandomSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_allHorizonRandomSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_allHorizonRandomSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail for every positive horizon under the random-square learning rate and clipped exploration schedule. The refined threshold is used exactly in the two-contract large-horizon regime; the other branch is the genuine zero-probability `T + 1` fallback.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedallhorizon/index.html#decl-eae221b77a6e","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","order":4633,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedAllHorizon.lean:46"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_allHorizonRandomSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let gamma := randomSquareBernsteinClippedExplorationRate (arms.card : Real) (horizon : Real) delta let eta := randomSquareHighProbabilityLearningRate (arms.card : Real) (horizon : Real) delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta…","missing":[],"search":"sampledpredictable_allhorizonrandomsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_allhorizonrandomsquarebernsteinrealizedregret_tail generated realized-regret tail for every positive horizon under the random-square learning rate and clipped exploration schedule. the refined threshold is used exactly in the two-contract large-horizon regime; the other branch is the genuine zero-probability `t + 1` fallback. theorem compiled","shard":"modules/52284718dbde5c44.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.log_one_div_fourth_eq_log_four_div","label":"log_one_div_fourth_eq_log_four_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.log_one_div_fourth_eq_log_four_div","description":"Dividing a total failure probability by four changes the logarithmic budget to `log (4 / delta)`.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-f72c39877968","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4634,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:21"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem log_one_div_fourth_eq_log_four_div (delta : Real) (hdelta : 0 < delta) : Real.log (1 / (delta / 4)) = Real.log (4 / delta)","missing":[],"search":"log_one_div_fourth_eq_log_four_div banditrlproof.exp3.log_one_div_fourth_eq_log_four_div dividing a total failure probability by four changes the logarithmic budget to `log (4 / delta)`. theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedExplicitThreshold","label":"randomSquareBernsteinRealizedExplicitThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRealizedExplicitThreshold","description":"The threshold after controlling both exploration-floor Bernstein radii and the realized-deviation radius by the exploration scale.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-b6e936fdd011","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4635,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:29"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinRealizedExplicitThreshold {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"randomsquarebernsteinrealizedexplicitthreshold banditrlproof.exp3.randomsquarebernsteinrealizedexplicitthreshold the threshold after controlling both exploration-floor bernstein radii and the realized-deviation radius by the exploration scale. definition compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","label":"randomSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","description":"Cubic and quadratic exploration contracts turn the three remaining confidence radii in the learning-rate-tuned threshold into `7 * gamma * T`; together with exploration bias this contributes `8 * gamma * T`.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-b115174d7286","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4636,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:40"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinRealizedTunedThreshold_le_explicitThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon : Nat) (hhorizon : 0 < horizon) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcubic_confidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) (hquadratic_realized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * (horizon : Real)) : randomSquareBernsteinRealizedTunedThreshold arms gamma horizon delta <= randomSquareBernsteinRealizedExplicitThreshold arms gamma horizon delta","missing":[],"search":"randomsquarebernsteinrealizedtunedthreshold_le_explicitthreshold banditrlproof.exp3.randomsquarebernsteinrealizedtunedthreshold_le_explicitthreshold cubic and quadratic exploration contracts turn the three remaining confidence radii in the learning-rate-tuned threshold into `7 * gamma * t`; together with exploration bias this contributes `8 * gamma * t`. theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedRandomSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_gammaCharacterizedRandomSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedRandomSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail after characterizing the remaining exploration parameter by one cubic and one quadratic dominance contract.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-ea1d2178d2d9","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4637,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:114"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_gammaCharacterizedRandomSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcubic_confidence : (arms.card : Real) * Real.log (4 / delta) <= gamma ^ 3 * (horizon : Real)) (hquadratic_realized : 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Re…","missing":[],"search":"sampledpredictable_gammacharacterizedrandomsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_gammacharacterizedrandomsquarebernsteinrealizedregret_tail generated realized-regret tail after characterizing the remaining exploration parameter by one cubic and one quadratic dominance contract. theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinConfidenceExplorationScale","label":"randomSquareBernsteinConfidenceExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinConfidenceExplorationScale","description":"Cube-root scale required by both `delta / 4` importance-weighted Bernstein radii.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-d60b3ee4dab0","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4638,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:187"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinConfidenceExplorationScale (K T delta : Real) : Real","missing":[],"search":"randomsquarebernsteinconfidenceexplorationscale banditrlproof.exp3.randomsquarebernsteinconfidenceexplorationscale cube-root scale required by both `delta / 4` importance-weighted bernstein radii. definition compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedExplorationScale","label":"randomSquareBernsteinRealizedExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRealizedExplorationScale","description":"Square-root scale required by the `delta / 4` realized-deviation radius.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-8becda6fbd7a","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4639,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:192"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinRealizedExplorationScale (T delta : Real) : Real","missing":[],"search":"randomsquarebernsteinrealizedexplorationscale banditrlproof.exp3.randomsquarebernsteinrealizedexplorationscale square-root scale required by the `delta / 4` realized-deviation radius. definition compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate","label":"randomSquareBernsteinRawExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate","description":"Unclipped exploration scale covering both remaining confidence contracts.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-0fbe59d9bb88","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4640,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:199"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinRawExplorationRate (K T delta : Real) : Real","missing":[],"search":"randomsquarebernsteinrawexplorationrate banditrlproof.exp3.randomsquarebernsteinrawexplorationrate unclipped exploration scale covering both remaining confidence contracts. definition compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate","label":"randomSquareBernsteinClippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate","description":"Explicit exploration schedule clipped into the Hedge stability regime.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-5552d2ec3b1f","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4641,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:205"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinClippedExplorationRate (K T delta : Real) : Real","missing":[],"search":"randomsquarebernsteinclippedexplorationrate banditrlproof.exp3.randomsquarebernsteinclippedexplorationrate explicit exploration schedule clipped into the hedge stability regime. definition compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_le_half","label":"randomSquareBernsteinClippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_le_half","description":"theorem randomSquareBernsteinClippedExplorationRate_le_half (K T delta : Real) : randomSquareBernsteinClippedExplorationRate K T delta <= 1 / 2","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-cd8d6f4bcd60","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4642,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:209"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinClippedExplorationRate_le_half (K T delta : Real) : randomSquareBernsteinClippedExplorationRate K T delta <= 1 / 2","missing":[],"search":"randomsquarebernsteinclippedexplorationrate_le_half banditrlproof.exp3.randomsquarebernsteinclippedexplorationrate_le_half theorem randomsquarebernsteinclippedexplorationrate_le_half (k t delta : real) : randomsquarebernsteinclippedexplorationrate k t delta <= 1 / 2 theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_eq_raw","label":"randomSquareBernsteinClippedExplorationRate_eq_raw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_eq_raw","description":"theorem randomSquareBernsteinClippedExplorationRate_eq_raw (K T delta : Real) (hraw : randomSquareBernsteinRawExplorationRate K T delta <= 1 / 2) : randomSquareBernsteinClippedExplorationRate K T delta = randomSquareBernsteinRawExplorationRate K T delta","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-1083bdb74bb6","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4643,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:214"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinClippedExplorationRate_eq_raw (K T delta : Real) (hraw : randomSquareBernsteinRawExplorationRate K T delta <= 1 / 2) : randomSquareBernsteinClippedExplorationRate K T delta = randomSquareBernsteinRawExplorationRate K T delta","missing":[],"search":"randomsquarebernsteinclippedexplorationrate_eq_raw banditrlproof.exp3.randomsquarebernsteinclippedexplorationrate_eq_raw theorem randomsquarebernsteinclippedexplorationrate_eq_raw (k t delta : real) (hraw : randomsquarebernsteinrawexplorationrate k t delta <= 1 / 2) : randomsquarebernsteinclippedexplorationrate k t delta = randomsquarebernsteinrawexplorationrate k t delta theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate_pos","label":"randomSquareBernsteinRawExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate_pos","description":"theorem randomSquareBernsteinRawExplorationRate_pos (K T delta : Real) (hK : 0 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < randomSquareBernsteinRawExplorationRate K T delta","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-a1751e2ac567","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4644,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:221"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinRawExplorationRate_pos (K T delta : Real) (hK : 0 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < randomSquareBernsteinRawExplorationRate K T delta","missing":[],"search":"randomsquarebernsteinrawexplorationrate_pos banditrlproof.exp3.randomsquarebernsteinrawexplorationrate_pos theorem randomsquarebernsteinrawexplorationrate_pos (k t delta : real) (hk : 0 < k) (ht : 0 < t) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < randomsquarebernsteinrawexplorationrate k t delta theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_pos","label":"randomSquareBernsteinClippedExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_pos","description":"theorem randomSquareBernsteinClippedExplorationRate_pos (K T delta : Real) (hK : 0 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < randomSquareBernsteinClippedExplorationRate K T delta","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-53817f627f87","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4645,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:235"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinClippedExplorationRate_pos (K T delta : Real) (hK : 0 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < randomSquareBernsteinClippedExplorationRate K T delta","missing":[],"search":"randomsquarebernsteinclippedexplorationrate_pos banditrlproof.exp3.randomsquarebernsteinclippedexplorationrate_pos theorem randomsquarebernsteinclippedexplorationrate_pos (k t delta : real) (hk : 0 < k) (ht : 0 < t) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < randomsquarebernsteinclippedexplorationrate k t delta theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","label":"randomSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","description":"Transparent large-horizon conditions ensure clipping is inactive.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-76e3fb8986af","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4646,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:245"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts (K T delta : Real) (hK : 0 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : randomSquareBernsteinRawExplorationRate K T delta <= 1 / 2","missing":[],"search":"randomsquarebernsteinrawexplorationrate_le_half_of_horizon_contracts banditrlproof.exp3.randomsquarebernsteinrawexplorationrate_le_half_of_horizon_contracts transparent large-horizon conditions ensure clipping is inactive. theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_contracts","label":"randomSquareBernsteinClippedExplorationRate_contracts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_contracts","description":"The clipped schedule satisfies the exact cubic and quadratic contracts consumed by the characterized random-square theorem.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-14b1f3c141fd","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4647,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:273"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareBernsteinClippedExplorationRate_contracts (K T delta : Real) (hK : 0 < K) (hT : 0 < T) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_confidence : 8 * (K * Real.log (4 / delta)) <= T) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= T) : let gamma := randomSquareBernsteinClippedExplorationRate K T delta 0 < gamma ∧ gamma <= 1 / 2 ∧ K * Real.log (4 / delta) <= gamma ^ 3 * T ∧ 2 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= gamma ^ 2 * T","missing":[],"search":"randomsquarebernsteinclippedexplorationrate_contracts banditrlproof.exp3.randomsquarebernsteinclippedexplorationrate_contracts the clipped schedule satisfies the exact cubic and quadratic contracts consumed by the characterized random-square theorem. theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitRandomSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_explicitRandomSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_explicitRandomSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail for the explicit clipped maximum of the confidence cube-root and realized square-root scales.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedexplicittuning/index.html#decl-a0a00ae86d3b","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","order":4648,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedExplicitTuning.lean:328"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_explicitRandomSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hlarge_confidence : 8 * ((arms.card : Real) * Real.log (4 / delta)) <= (horizon : Real)) (hlarge_realized : 8 * ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real) * Real.log (4 / delta) <= (horizon : Real)) : let gamma := randomSquareBernsteinClippedExploration…","missing":[],"search":"sampledpredictable_explicitrandomsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_explicitrandomsquarebernsteinrealizedregret_tail generated realized-regret tail for the explicit clipped maximum of the confidence cube-root and realized square-root scales. theorem compiled","shard":"modules/263eb29d89a3e112.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget","label":"sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget","description":"Realized selected-loss regret budget whose predictable component uses the random estimator-square event and two variance-sensitive Bernstein radii.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedhighprobabilityregret/index.html#decl-9f82d0766b02","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","order":4649,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedHighProbabilityRegret.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (deltaSquare deltaConfidence deltaRealized : Real) : Real","missing":[],"search":"sampledpredictablerandomsquarebernsteinrealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablerandomsquarebernsteinrealizedhighprobabilityregretbudget realized selected-loss regret budget whose predictable component uses the random estimator-square event and two variance-sensitive bernstein radii. definition compiled","shard":"modules/4004bf5482c4c440.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail","label":"sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail","description":"Raw four-event form. The predictable component contributes the random estimator-square event and the two Bernstein confidence events; the fourth event is the bounded realized-minus-predictable deviation.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedhighprobabilityregret/index.html#decl-a8dce480a0b1","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","order":4650,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedHighProbabilityRegret.lean:36"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (deltaSquare deltaConfidence deltaRealized : Real) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) (hdeltaRealized : 0 < deltaRealized) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt…","missing":[],"search":"sampledpredictable_randomsquarebernsteinrealizedhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_randomsquarebernsteinrealizedhighprobabilityregret_tail raw four-event form. the predictable component contributes the random estimator-square event and the two bernstein confidence events; the fourth event is the bounded realized-minus-predictable deviation. theorem compiled","shard":"modules/4004bf5482c4c440.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","description":"Total-failure form: the estimator-square, pure-cross Bernstein, fixed-comparator Bernstein, and realized-deviation events each receive `delta / 4`.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedhighprobabilityregret/index.html#decl-95caacb18b1d","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","order":4651,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedHighProbabilityRegret.lean:158"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget a…","missing":[],"search":"sampledpredictable_randomsquarebernsteinrealizedhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_randomsquarebernsteinrealizedhighprobabilityregret_tail_total_delta total-failure form: the estimator-square, pure-cross bernstein, fixed-comparator bernstein, and realized-deviation events each receive `delta / 4`. theorem compiled","shard":"modules/4004bf5482c4c440.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate","label":"randomSquareHighProbabilityLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate","description":"Learning rate balancing entropy against the Markov estimator-square term when the square event receives `delta / 4`.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-8fc5b8696118","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4652,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareHighProbabilityLearningRate (K T delta : Real) : Real","missing":[],"search":"randomsquarehighprobabilitylearningrate banditrlproof.exp3.randomsquarehighprobabilitylearningrate learning rate balancing entropy against the markov estimator-square term when the square event receives `delta / 4`. definition compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate_pos","label":"randomSquareHighProbabilityLearningRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate_pos","description":"theorem randomSquareHighProbabilityLearningRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) : 0 < randomSquareHighProbabilityLearningRate K T delta","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-60b571ca7c5e","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4653,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareHighProbabilityLearningRate_pos (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) : 0 < randomSquareHighProbabilityLearningRate K T delta","missing":[],"search":"randomsquarehighprobabilitylearningrate_pos banditrlproof.exp3.randomsquarehighprobabilitylearningrate_pos theorem randomsquarehighprobabilitylearningrate_pos (k t delta : real) (hk_one : 1 < k) (ht : 0 < t) (hdelta : 0 < delta) : 0 < randomsquarehighprobabilitylearningrate k t delta theorem compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate_sq_mul","label":"randomSquareHighProbabilityLearningRate_sq_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate_sq_mul","description":"theorem randomSquareHighProbabilityLearningRate_sq_mul (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) : randomSquareHighProbabilityLearningRate K T delta ^ 2 * (T * K) = Real.log K * (delta / 4)","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-6a5ffbe334f5","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4654,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:37"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareHighProbabilityLearningRate_sq_mul (K T delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hdelta : 0 < delta) : randomSquareHighProbabilityLearningRate K T delta ^ 2 * (T * K) = Real.log K * (delta / 4)","missing":[],"search":"randomsquarehighprobabilitylearningrate_sq_mul banditrlproof.exp3.randomsquarehighprobabilitylearningrate_sq_mul theorem randomsquarehighprobabilitylearningrate_sq_mul (k t delta : real) (hk_one : 1 < k) (ht : 0 < t) (hdelta : 0 < delta) : randomsquarehighprobabilitylearningrate k t delta ^ 2 * (t * k) = real.log k * (delta / 4) theorem compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","label":"randomSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","description":"With `gamma <= 1/2`, the entropy and stability-amplified random-square terms cost at most three copies of their balanced square-root scale.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-79e719c7cad8","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4655,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:47"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem randomSquareHighProbabilityHedgeBudget_le_three_mul_sqrt (K T gamma delta : Real) (hK_one : 1 < K) (hT : 0 < T) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) : Real.log K / randomSquareHighProbabilityLearningRate K T delta + (randomSquareHighProbabilityLearningRate K T delta * (1 / (1 - gamma))) * (K * T / (delta / 4)) <= 3 * Real.sqrt (4 * K * T * Real.log K / delta)","missing":[],"search":"randomsquarehighprobabilityhedgebudget_le_three_mul_sqrt banditrlproof.exp3.randomsquarehighprobabilityhedgebudget_le_three_mul_sqrt with `gamma <= 1/2`, the entropy and stability-amplified random-square terms cost at most three copies of their balanced square-root scale. theorem compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedTunedThreshold","label":"randomSquareBernsteinRealizedTunedThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.randomSquareBernsteinRealizedTunedThreshold","description":"Explicit threshold after tuning only the learning-rate-dependent terms. The exploration and three confidence contributions remain visible.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-bd62734c31b3","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4656,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:113"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def randomSquareBernsteinRealizedTunedThreshold {Action : Type v} (arms : Finset Action) (gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"randomsquarebernsteinrealizedtunedthreshold banditrlproof.exp3.randomsquarebernsteinrealizedtunedthreshold explicit threshold after tuning only the learning-rate-dependent terms. the exploration and three confidence contributions remain visible. definition compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","label":"sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","description":"The complete four-event realized budget is bounded by the explicit learning-rate-tuned threshold.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-8f963df66ed8","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4657,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:128"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold {Action : Type v} [DecidableEq Action] (arms : Finset Action) (hcard_two : 2 <= arms.card) (horizon : Nat) (hhorizon : 0 < horizon) (gamma delta : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (hdelta : 0 < delta) : sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget arms (randomSquareHighProbabilityLearningRate (arms.card : Real) (horizon : Real) delta) gamma horizon (delta / 4) (delta / 4) (delta / 4) <= randomSquareBernsteinRealizedTunedThreshold arms gamma horizon delta","missing":[],"search":"sampledpredictablerandomsquarebernsteinrealizedhighprobabilityregretbudget_le_tunedthreshold banditrlproof.exp3.sampledpredictablerandomsquarebernsteinrealizedhighprobabilityregretbudget_le_tunedthreshold the complete four-event realized budget is bounded by the explicit learning-rate-tuned threshold. theorem compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedRandomSquareBernsteinRealizedRegret_tail","label":"sampledPredictable_tunedRandomSquareBernsteinRealizedRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_tunedRandomSquareBernsteinRealizedRegret_tail","description":"Generated realized-regret tail with the random-square learning rate. This optimizes the entropy/Markov-square pair while leaving the exploration-floor and realized-deviation confidence terms explicit.","url":"../modules/banditrlproof-exp3randomsquarebernsteinrealizedtuning/index.html#decl-12405a37a6f2","parent":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","order":4658,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning"],["Source","BanditRLProof/Exp3RandomSquareBernsteinRealizedTuning.lean:158"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_tunedRandomSquareBernsteinRealizedRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_half : gamma <= 1 / 2) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let eta := randomSquareHighProbabilityLearningRate (arms.card : Real) (horizon : Real) delta let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le (by linarith : gamma <= 1) loss.enviro…","missing":[],"search":"sampledpredictable_tunedrandomsquarebernsteinrealizedregret_tail banditrlproof.exp3.sampledpredictable_tunedrandomsquarebernsteinrealizedregret_tail generated realized-regret tail with the random-square learning rate. this optimizes the entropy/markov-square pair while leaving the exploration-floor and realized-deviation confidence terms explicit. theorem compiled","shard":"modules/84268fd8bd9855c7.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedMixedSquaredSum","label":"sampledObservedMixedSquaredSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedMixedSquaredSum","description":"Finite-horizon sum of the probability-mixed squared importance-weighted loss estimates observed on a sampled trajectory.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-e3399cd0fc9c","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4659,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:23"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledObservedMixedSquaredSum {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledobservedmixedsquaredsum banditrlproof.exp3.sampledobservedmixedsquaredsum finite-horizon sum of the probability-mixed squared importance-weighted loss estimates observed on a sampled trajectory. definition compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt_nonneg","label":"observedMixedSquaredImportanceWeightedLossAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt_nonneg","description":"Each observed mixed estimator square is pointwise nonnegative.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-73338501701c","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4660,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem observedMixedSquaredImportanceWeightedLossAt_nonneg {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : 0 <= observedMixedSquaredImportanceWeightedLossAt arms eta gamma t sample","missing":[],"search":"observedmixedsquaredimportanceweightedlossat_nonneg banditrlproof.exp3.observedmixedsquaredimportanceweightedlossat_nonneg each observed mixed estimator square is pointwise nonnegative. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledObservedMixedSquaredSum_nonneg","label":"sampledObservedMixedSquaredSum_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledObservedMixedSquaredSum_nonneg","description":"The finite-horizon observed mixed estimator-square sum is pointwise nonnegative.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-bb8e422c27c2","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4661,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:49"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledObservedMixedSquaredSum_nonneg {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : 0 <= sampledObservedMixedSquaredSum arms eta gamma horizon sample","missing":[],"search":"sampledobservedmixedsquaredsum_nonneg banditrlproof.exp3.sampledobservedmixedsquaredsum_nonneg the finite-horizon observed mixed estimator-square sum is pointwise nonnegative. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledObservedMixedSquaredSum","label":"measurable_sampledObservedMixedSquaredSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledObservedMixedSquaredSum","description":"The finite-horizon observed mixed estimator-square sum is measurable.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-28ea6b97e378","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4662,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:62"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledObservedMixedSquaredSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (horizon : Nat) : Measurable (sampledObservedMixedSquaredSum arms eta gamma horizon : (Env × ((k : Nat) -> Action × Real)) -> Real)","missing":[],"search":"measurable_sampledobservedmixedsquaredsum banditrlproof.exp3.measurable_sampledobservedmixedsquaredsum the finite-horizon observed mixed estimator-square sum is measurable. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_sampledPredictableObservedMixedSquaredSum","label":"integrable_sampledPredictableObservedMixedSquaredSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_sampledPredictableObservedMixedSquaredSum","description":"The generated finite-horizon observed mixed estimator-square sum is integrable.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-592f617ada92","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4663,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:79"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledPredictableObservedMixedSquaredSum {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Integrable (sampledObservedMixedSquaredSum arms eta gamma horizon) mu","missing":[],"search":"integrable_sampledpredictableobservedmixedsquaredsum banditrlproof.exp3.integrable_sampledpredictableobservedmixedsquaredsum the generated finite-horizon observed mixed estimator-square sum is integrable. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_markov","label":"sampledPredictableObservedMixedSquared_sum_tail_markov","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_markov","description":"Markov tail for the random finite-horizon estimator-square sum. Its threshold is `|arms| * T / deltaSquare`, with no reciprocal exploration-rate factor.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-73e1118835f5","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4664,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:108"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableObservedMixedSquared_sum_tail_markov {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (deltaSquare : Real) (hdeltaSquare : 0 < deltaSquare) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | (arms.card : Real) * (horizon : Real) / deltaSquare <= sampledObservedMixedSquaredSum arms eta gamma horizon sample} <= ENNReal.ofReal deltaSquare","missing":[],"search":"sampledpredictableobservedmixedsquared_sum_tail_markov banditrlproof.exp3.sampledpredictableobservedmixedsquared_sum_tail_markov markov tail for the random finite-horizon estimator-square sum. its threshold is `|arms| * t / deltasquare`, with no reciprocal exploration-rate factor. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinHighProbabilityRegretBudget","label":"sampledPredictableRandomSquareBernsteinHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinHighProbabilityRegretBudget","description":"Predictable regret budget with a caller-visible random-square failure allocation and the two variance-sensitive confidence radii.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-5038544a8b67","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4665,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:184"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRandomSquareBernsteinHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (deltaSquare deltaConfidence : Real) : Real","missing":[],"search":"sampledpredictablerandomsquarebernsteinhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablerandomsquarebernsteinhighprobabilityregretbudget predictable regret budget with a caller-visible random-square failure allocation and the two variance-sensitive confidence radii. definition compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail","label":"sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail","description":"Generated predictable EXP3 regret with a Markov estimator-square event and the two compiled Bernstein confidence events.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-895edef0811e","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4666,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:198"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (deltaSquare deltaConfidence : Real) (hdeltaSquare : 0 < deltaSquare) (hdeltaConfidence : 0 < deltaConfidence) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableRan…","missing":[],"search":"sampledpredictable_randomsquarebernsteinhighprobabilityregret_tail banditrlproof.exp3.sampledpredictable_randomsquarebernsteinhighprobabilityregret_tail generated predictable exp3 regret with a markov estimator-square event and the two compiled bernstein confidence events. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail_total_delta","description":"Total-failure form: the estimator-square event and both Bernstein confidence events each receive `delta / 3`.","url":"../modules/banditrlproof-exp3randomsquarehighprobabilityregret/index.html#decl-0755fe8df14b","parent":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","order":4667,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RandomSquareHighProbabilityRegret"],["Source","BanditRLProof/Exp3RandomSquareHighProbabilityRegret.lean:359"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableRandomSquareBernsteinHighProbabilityRegretBudget arms eta gamma ho…","missing":[],"search":"sampledpredictable_randomsquarebernsteinhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_randomsquarebernsteinhighprobabilityregret_tail_total_delta total-failure form: the estimator-square event and both bernstein confidence events each receive `delta / 3`. theorem compiled","shard":"modules/d605b90fd453a37a.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.condExpKernel_map_eq_finiteActionMeasure_of_condDistrib_ae_eq","label":"condExpKernel_map_eq_finiteActionMeasure_of_condDistrib_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.condExpKernel_map_eq_finiteActionMeasure_of_condDistrib_ae_eq","description":"theorem condExpKernel_map_eq_finiteActionMeasure_of_condDistrib_ae_eq {Omega : Type u} {Condition : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mCondition : MeasurableSpace Condition] [mAction : MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (X : Omega -> Acti…","url":"../modules/banditrlproof-exp3realizedconcentration/index.html#decl-c7f27f2f6fd6","parent":"module:BanditRLProof.Exp3RealizedConcentration","order":4668,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConcentration"],["Source","BanditRLProof/Exp3RealizedConcentration.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_eq_finiteActionMeasure_of_condDistrib_ae_eq {Omega : Type u} {Condition : Type v} {Action : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mCondition : MeasurableSpace Condition] [mAction : MeasurableSpace Action] [StandardBorelSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (X : Omega -> Action) (Y : Omega -> Condition) (hX : @Measurable Omega Action mOmega mAction X) (hY : @Measurable Omega Condition mOmega mCondition Y) (arms : Finset Action) (prob : Condition -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (hcond : Filter.EventuallyEq (ae (mu.map Y)) (condDistrib X Y mu) (finiteActionKernel arms prob source)) : Filter.Eventually (fun omega => @Measure.map Omega Action mOmega mAction X (@condExpKernel Omega mOmega _ mu _ (mCondition.c…","missing":[],"search":"condexpkernel_map_eq_finiteactionmeasure_of_conddistrib_ae_eq banditrlproof.exp3.condexpkernel_map_eq_finiteactionmeasure_of_conddistrib_ae_eq theorem condexpkernel_map_eq_finiteactionmeasure_of_conddistrib_ae_eq {omega : type u} {condition : type v} {action : type w} [momega : measurablespace omega] [standardborelspace omega] [nonempty omega] [mcondition : measurablespace condition] [maction : measurablespace action] [standardborelspace action] [measurablesingletonclass action] [nonempty action] (mu : measure omega) [isfinitemeasure mu] (x : omega -> action) (y : omega -> condition) (hx : @measurable omega action momega maction x) (hy : @measurable omega condition momega mcondition y) (arms : finset action) (prob : condition -> action -> real) (source : measurablefiniteactiondistribution arms prob) (hcond : filter.eventuallyeq (ae (mu.map y)) (conddistrib x y mu) (finiteactionkernel arms prob source)) : filter.eventually (fun omega => @measure.map omega action momega maction x (@condexpkernel omega momega _ mu _ (mcondition.comap y) omega) = finiteactionmeasure arms (prob (y omega))) (ae (mu.trim hy.comap_le)) theorem compiled","shard":"modules/de2dee354111cae8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectorySelectedDeviationAt","label":"sampledTrajectorySelectedDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectorySelectedDeviationAt","description":"Selected predictable loss minus its exploration-mixed conditional mean.","url":"../modules/banditrlproof-exp3realizedconcentration/index.html#decl-3efc7cc0effb","parent":"module:BanditRLProof.Exp3RealizedConcentration","order":4669,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedConcentration"],["Source","BanditRLProof/Exp3RealizedConcentration.lean:180"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectorySelectedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectoryselecteddeviationat banditrlproof.exp3.sampledtrajectoryselecteddeviationat selected predictable loss minus its exploration-mixed conditional mean. definition compiled","shard":"modules/de2dee354111cae8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectorySelectedDeviationAt","label":"measurable_sampledTrajectorySelectedDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectorySelectedDeviationAt","description":"theorem measurable_sampledTrajectorySelectedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectorySelectedDeviationAt a…","url":"../modules/banditrlproof-exp3realizedconcentration/index.html#decl-612359118b6f","parent":"module:BanditRLProof.Exp3RealizedConcentration","order":4670,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConcentration"],["Source","BanditRLProof/Exp3RealizedConcentration.lean:189"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectorySelectedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectorySelectedDeviationAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectoryselecteddeviationat banditrlproof.exp3.measurable_sampledtrajectoryselecteddeviationat theorem measurable_sampledtrajectoryselecteddeviationat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : measurable (sampledtrajectoryselecteddeviationat arms eta gamma loss t) theorem compiled","shard":"modules/de2dee354111cae8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedDeviation_succ_hasCondSubgaussianMGF","label":"sampledPredictableSelectedDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSelectedDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledPredictableSelectedDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg…","url":"../modules/banditrlproof-exp3realizedconcentration/index.html#decl-cb5e81e5ae88","parent":"module:BanditRLProof.Exp3RealizedConcentration","order":4671,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConcentration"],["Source","BanditRLProof/Exp3RealizedConcentration.lean:204"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSelectedDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real n)).comap history) (measura…","missing":[],"search":"sampledpredictableselecteddeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictableselecteddeviation_succ_hascondsubgaussianmgf theorem sampledpredictableselecteddeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectoryselecteddeviationat arms eta gamma loss (n + 1)) (concentration.intervalvarianceproxy 0 1) mu theorem compiled","shard":"modules/de2dee354111cae8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedDeviationAt","label":"sampledTrajectoryRealizedDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryRealizedDeviationAt","description":"noncomputable def sampledTrajectoryRealizedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","url":"../modules/banditrlproof-exp3realizedconcentration/index.html#decl-286716242224","parent":"module:BanditRLProof.Exp3RealizedConcentration","order":4672,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedConcentration"],["Source","BanditRLProof/Exp3RealizedConcentration.lean:367"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryRealizedDeviationAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectoryrealizeddeviationat banditrlproof.exp3.sampledtrajectoryrealizeddeviationat noncomputable def sampledtrajectoryrealizeddeviationat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : real definition compiled","shard":"modules/de2dee354111cae8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_succ_hasCondSubgaussianMGF","label":"sampledPredictableRealizedDeviation_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_succ_hasCondSubgaussianMGF","description":"theorem sampledPredictableRealizedDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg…","url":"../modules/banditrlproof-exp3realizedconcentration/index.html#decl-30892947d569","parent":"module:BanditRLProof.Exp3RealizedConcentration","order":4673,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConcentration"],["Source","BanditRLProof/Exp3RealizedConcentration.lean:376"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_succ_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real n)).comap history) (measura…","missing":[],"search":"sampledpredictablerealizeddeviation_succ_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictablerealizeddeviation_succ_hascondsubgaussianmgf theorem sampledpredictablerealizeddeviation_succ_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (n : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment let history := fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle n sample.2) probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace (env × history.finitepairhistory action real n)).comap history) (measurable_fst.prodmk ((preorder.measurable_frestrictle n).comp measurable_snd)).comap_le (sampledtrajectoryrealizeddeviationat arms eta gamma loss (n + 1)) (concentration.intervalvarianceproxy 0 1) mu theorem compiled","shard":"modules/de2dee354111cae8.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius","label":"sampledPredictableRealizedDeviationConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius","description":"noncomputable def sampledPredictableRealizedDeviationConfidenceRadius (horizon : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3realizedconfidence/index.html#decl-88e69c544a36","parent":"module:BanditRLProof.Exp3RealizedConfidence","order":4674,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedConfidence"],["Source","BanditRLProof/Exp3RealizedConfidence.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedDeviationConfidenceRadius (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablerealizeddeviationconfidenceradius banditrlproof.exp3.sampledpredictablerealizeddeviationconfidenceradius noncomputable def sampledpredictablerealizeddeviationconfidenceradius (horizon : nat) (delta : real) : real definition compiled","shard":"modules/52f01b4c40547c85.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.intervalVarianceProxy_zero_one_pos","label":"intervalVarianceProxy_zero_one_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.intervalVarianceProxy_zero_one_pos","description":"theorem intervalVarianceProxy_zero_one_pos : 0 < ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real)","url":"../modules/banditrlproof-exp3realizedconfidence/index.html#decl-a0029bf8bcf1","parent":"module:BanditRLProof.Exp3RealizedConfidence","order":4675,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConfidence"],["Source","BanditRLProof/Exp3RealizedConfidence.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem intervalVarianceProxy_zero_one_pos : 0 < ((Concentration.intervalVarianceProxy 0 1 : NNReal) : Real)","missing":[],"search":"intervalvarianceproxy_zero_one_pos banditrlproof.exp3.intervalvarianceproxy_zero_one_pos theorem intervalvarianceproxy_zero_one_pos : 0 < ((concentration.intervalvarianceproxy 0 1 : nnreal) : real) theorem compiled","shard":"modules/52f01b4c40547c85.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius_nonneg","label":"sampledPredictableRealizedDeviationConfidenceRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius_nonneg","description":"theorem sampledPredictableRealizedDeviationConfidenceRadius_nonneg (horizon : Nat) (delta : Real) : 0 <= sampledPredictableRealizedDeviationConfidenceRadius horizon delta","url":"../modules/banditrlproof-exp3realizedconfidence/index.html#decl-acaa151524a1","parent":"module:BanditRLProof.Exp3RealizedConfidence","order":4676,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConfidence"],["Source","BanditRLProof/Exp3RealizedConfidence.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviationConfidenceRadius_nonneg (horizon : Nat) (delta : Real) : 0 <= sampledPredictableRealizedDeviationConfidenceRadius horizon delta","missing":[],"search":"sampledpredictablerealizeddeviationconfidenceradius_nonneg banditrlproof.exp3.sampledpredictablerealizeddeviationconfidenceradius_nonneg theorem sampledpredictablerealizeddeviationconfidenceradius_nonneg (horizon : nat) (delta : real) : 0 <= sampledpredictablerealizeddeviationconfidenceradius horizon delta theorem compiled","shard":"modules/52f01b4c40547c85.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius_sq_domination","label":"sampledPredictableRealizedDeviationConfidenceRadius_sq_domination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius_sq_domination","description":"theorem sampledPredictableRealizedDeviationConfidenceRadius_sq_domination (horizon : Nat) (budget : Real) : 2 * ((((horizon : NNReal) * Concentration.intervalVarianceProxy 0 1 : NNReal)) : Real) * budget <= (Real.sqrt (2 * ((((horizon : NNReal) * Concentration.intervalVarianceProxy 0 1 : NNReal)) : Real) * budget)) ^ 2","url":"../modules/banditrlproof-exp3realizedconfidence/index.html#decl-b005dab35768","parent":"module:BanditRLProof.Exp3RealizedConfidence","order":4677,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConfidence"],["Source","BanditRLProof/Exp3RealizedConfidence.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviationConfidenceRadius_sq_domination (horizon : Nat) (budget : Real) : 2 * ((((horizon : NNReal) * Concentration.intervalVarianceProxy 0 1 : NNReal)) : Real) * budget <= (Real.sqrt (2 * ((((horizon : NNReal) * Concentration.intervalVarianceProxy 0 1 : NNReal)) : Real) * budget)) ^ 2","missing":[],"search":"sampledpredictablerealizeddeviationconfidenceradius_sq_domination banditrlproof.exp3.sampledpredictablerealizeddeviationconfidenceradius_sq_domination theorem sampledpredictablerealizeddeviationconfidenceradius_sq_domination (horizon : nat) (budget : real) : 2 * ((((horizon : nnreal) * concentration.intervalvarianceproxy 0 1 : nnreal)) : real) * budget <= (real.sqrt (2 * ((((horizon : nnreal) * concentration.intervalvarianceproxy 0 1 : nnreal)) : real) * budget)) ^ 2 theorem compiled","shard":"modules/52f01b4c40547c85.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_exp_neg_budget","label":"sampledPredictableRealizedDeviation_sum_tail_exp_neg_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_exp_neg_budget","description":"theorem sampledPredictableRealizedDeviation_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonne…","url":"../modules/banditrlproof-exp3realizedconfidence/index.html#decl-1555bbcf0419","parent":"module:BanditRLProof.Exp3RealizedConfidence","order":4678,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConfidence"],["Source","BanditRLProof/Exp3RealizedConfidence.lean:45"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_sum_tail_exp_neg_budget {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (budget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment mu {sample | Real.sqrt (2 * ((((horizon : NNReal) * Concentration.intervalVarianceProxy 0 1 : NNReal)) : Real) * budget) <= (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta…","missing":[],"search":"sampledpredictablerealizeddeviation_sum_tail_exp_neg_budget banditrlproof.exp3.sampledpredictablerealizeddeviation_sum_tail_exp_neg_budget theorem sampledpredictablerealizeddeviation_sum_tail_exp_neg_budget {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) (hhorizon : 0 < horizon) (budget : real) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment mu {sample | real.sqrt (2 * ((((horizon : nnreal) * concentration.intervalvarianceproxy 0 1 : nnreal)) : real) * budget) <= (finset.range horizon).sum (fun i => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample)} <= ennreal.ofreal (real.exp (-budget)) theorem compiled","shard":"modules/52f01b4c40547c85.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_delta","label":"sampledPredictableRealizedDeviation_sum_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_delta","description":"theorem sampledPredictableRealizedDeviation_sum_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <=…","url":"../modules/banditrlproof-exp3realizedconfidence/index.html#decl-41518bca0bc4","parent":"module:BanditRLProof.Exp3RealizedConfidence","order":4679,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedConfidence"],["Source","BanditRLProof/Exp3RealizedConfidence.lean:90"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_sum_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment mu {sample | sampledPredictableRealizedDeviationConfidenceRadius horizon delta <= (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample)} <= ENNRea…","missing":[],"search":"sampledpredictablerealizeddeviation_sum_tail_delta banditrlproof.exp3.sampledpredictablerealizeddeviation_sum_tail_delta theorem sampledpredictablerealizeddeviation_sum_tail_delta {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (horizon : nat) (hhorizon : 0 < horizon) (delta : real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment mu {sample | sampledpredictablerealizeddeviationconfidenceradius horizon delta <= (finset.range horizon).sum (fun i => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample)} <= ennreal.ofreal delta theorem compiled","shard":"modules/52f01b4c40547c85.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceLinearBudget","label":"sampledRealizedPredictableVarianceLinearBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedPredictableVarianceLinearBudget","description":"Deterministic selected-loss predictable-variance budget at prefix `n+1`.","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-aa62687058af","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4680,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedPredictableVarianceLinearBudget (n : Nat) : Real","missing":[],"search":"sampledrealizedpredictablevariancelinearbudget banditrlproof.exp3.sampledrealizedpredictablevariancelinearbudget deterministic selected-loss predictable-variance budget at prefix `n+1`. definition compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceLinearBudget_pos","label":"sampledRealizedPredictableVarianceLinearBudget_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedPredictableVarianceLinearBudget_pos","description":"theorem sampledRealizedPredictableVarianceLinearBudget_pos (n : Nat) : 0 < sampledRealizedPredictableVarianceLinearBudget n","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-a5154a80fc80","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4681,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:23"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledRealizedPredictableVarianceLinearBudget_pos (n : Nat) : 0 < sampledRealizedPredictableVarianceLinearBudget n","missing":[],"search":"sampledrealizedpredictablevariancelinearbudget_pos banditrlproof.exp3.sampledrealizedpredictablevariancelinearbudget_pos theorem sampledrealizedpredictablevariancelinearbudget_pos (n : nat) : 0 < sampledrealizedpredictablevariancelinearbudget n theorem compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedDeviationGeometricAllTimeRadius","label":"sampledRealizedDeviationGeometricAllTimeRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedDeviationGeometricAllTimeRadius","description":"Geometric-share deviation radius with deterministic variance budget `n+1`.","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-742ad15930eb","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4682,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:30"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedDeviationGeometricAllTimeRadius (delta : Real) (n : Nat) : Real","missing":[],"search":"sampledrealizeddeviationgeometricalltimeradius banditrlproof.exp3.sampledrealizeddeviationgeometricalltimeradius geometric-share deviation radius with deterministic variance budget `n+1`. definition compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedDeviationGeometricAllTimeFailureSet","label":"sampledRealizedDeviationGeometricAllTimeFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedDeviationGeometricAllTimeFailureSet","description":"Pure realized-deviation failure event over every positive prefix of one generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-e6096cb6ba4f","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4683,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:37"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedDeviationGeometricAllTimeFailureSet {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (delta : Real) : Set (Env × ((k : Nat) → Action × Real))","missing":[],"search":"sampledrealizeddeviationgeometricalltimefailureset banditrlproof.exp3.sampledrealizeddeviationgeometricalltimefailureset pure realized-deviation failure event over every positive prefix of one generated exp3 trajectory. definition compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mem_sampledRealizedDeviationGeometricAllTimeFailureSet_iff","label":"mem_sampledRealizedDeviationGeometricAllTimeFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mem_sampledRealizedDeviationGeometricAllTimeFailureSet_iff","description":"theorem mem_sampledRealizedDeviationGeometricAllTimeFailureSet_iff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (delta : Real) (sample : Env × ((k : Nat) → Action × Real)) : sample ∈ sampledRealizedDeviationGeometricAllTimeFailureSet arms eta gamma loss delta ↔ ∃ n, sampledReali…","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-a272e2a7f012","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4684,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:48"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mem_sampledRealizedDeviationGeometricAllTimeFailureSet_iff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (delta : Real) (sample : Env × ((k : Nat) → Action × Real)) : sample ∈ sampledRealizedDeviationGeometricAllTimeFailureSet arms eta gamma loss delta ↔ ∃ n, sampledRealizedDeviationGeometricAllTimeRadius delta n ≤ (Finset.range (n + 1)).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample)","missing":[],"search":"mem_sampledrealizeddeviationgeometricalltimefailureset_iff banditrlproof.exp3.mem_sampledrealizeddeviationgeometricalltimefailureset_iff theorem mem_sampledrealizeddeviationgeometricalltimefailureset_iff {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (delta : real) (sample : env × ((k : nat) → action × real)) : sample ∈ sampledrealizeddeviationgeometricalltimefailureset arms eta gamma loss delta ↔ ∃ n, sampledrealizeddeviationgeometricalltimeradius delta n ≤ (finset.range (n + 1)).sum (fun i => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample) theorem compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationAllTimeFailureSet_linearBudget_eq","label":"sampledPredictableRealizedDeviationAllTimeFailureSet_linearBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationAllTimeFailureSet_linearBudget_eq","description":"With the deterministic budget `n+1`, the prior joint failure event is exactly the pure realized-deviation crossing event.","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-ed865699bcf2","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4685,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:65"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviationAllTimeFailureSet_linearBudget_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (delta : Real) : sampledPredictableRealizedDeviationAllTimeFailureSet arms eta gamma loss sampledRealizedPredictableVarianceLinearBudget delta = sampledRealizedDeviationGeometricAllTimeFailureSet arms eta gamma loss delta","missing":[],"search":"sampledpredictablerealizeddeviationalltimefailureset_linearbudget_eq banditrlproof.exp3.sampledpredictablerealizeddeviationalltimefailureset_linearbudget_eq with the deterministic budget `n+1`, the prior joint failure event is exactly the pure realized-deviation crossing event. theorem compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","label":"measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","description":"On one fixed generated EXP3 trajectory law, the realized selected-loss deviation stays below its geometric-share quadratic radius at every positive prefix outside a set of mass at most the outer confidence budget.","url":"../modules/banditrlproof-exp3realizeddeviationalltime/index.html#decl-01d8fa6ab9ae","parent":"module:BanditRLProof.Exp3RealizedDeviationAllTime","order":4686,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationAllTime"],["Source","BanditRLProof/Exp3RealizedDeviationAllTime.lean:97"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu (sampledRealizedDeviationGeometricAllTimeFailureSet arms eta gamma loss delta) ≤ ENNReal.ofReal delta","missing":[],"search":"measure_sampledrealizeddeviationgeometricalltimefailureset_le banditrlproof.exp3.measure_sampledrealizeddeviationgeometricalltimefailureset_le on one fixed generated exp3 trajectory law, the realized selected-loss deviation stays below its geometric-share quadratic radius at every positive prefix outside a set of mass at most the outer confidence budget. theorem compiled","shard":"modules/274f260682a6b672.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedDeviation_zero_hasCondSubgaussianMGF","label":"sampledPredictableSelectedDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSelectedDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledPredictableSelectedDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg…","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-46a3610e90a1","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4687,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSelectedDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectorySelectedDeviationAt arms eta gamma loss 0) (Concentration.intervalVariancePr…","missing":[],"search":"sampledpredictableselecteddeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictableselecteddeviation_zero_hascondsubgaussianmgf theorem sampledpredictableselecteddeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectoryselecteddeviationat arms eta gamma loss 0) (concentration.intervalvarianceproxy 0 1) mu theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_zero_hasCondSubgaussianMGF","label":"sampledPredictableRealizedDeviation_zero_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_zero_hasCondSubgaussianMGF","description":"theorem sampledPredictableRealizedDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg…","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-c7157477c3b9","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4688,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:178"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_zero_hasCondSubgaussianMGF {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment ProbabilityTheory.HasCondSubgaussianMGF ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)) measurable_fst.comap_le (sampledTrajectoryRealizedDeviationAt arms eta gamma loss 0) (Concentration.intervalVariancePr…","missing":[],"search":"sampledpredictablerealizeddeviation_zero_hascondsubgaussianmgf banditrlproof.exp3.sampledpredictablerealizeddeviation_zero_hascondsubgaussianmgf theorem sampledpredictablerealizeddeviation_zero_hascondsubgaussianmgf {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment probabilitytheory.hascondsubgaussianmgf ((inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1)) measurable_fst.comap_le (sampledtrajectoryrealizeddeviationat arms eta gamma loss 0) (concentration.intervalvarianceproxy 0 1) mu theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableDeviationFiltration","label":"sampledPredictableDeviationFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableDeviationFiltration","description":"def sampledPredictableDeviationFiltration (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] : Filtration Nat (inferInstance : MeasurableSpace (Env × ((k : Nat) -> Action × Real))) where","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-02abc9027d38","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4689,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:235"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def sampledPredictableDeviationFiltration (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] : Filtration Nat (inferInstance : MeasurableSpace (Env × ((k : Nat) -> Action × Real))) where","missing":[],"search":"sampledpredictabledeviationfiltration banditrlproof.exp3.sampledpredictabledeviationfiltration def sampledpredictabledeviationfiltration (env : type u) (action : type v) [measurablespace env] [measurablespace action] : filtration nat (inferinstance : measurablespace (env × ((k : nat) -> action × real))) where definition compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableDeviationFiltration_zero","label":"sampledPredictableDeviationFiltration_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableDeviationFiltration_zero","description":"theorem sampledPredictableDeviationFiltration_zero (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] : sampledPredictableDeviationFiltration Env Action 0 = (inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-3105bb7c6399","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4690,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:312"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableDeviationFiltration_zero (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] : sampledPredictableDeviationFiltration Env Action 0 = (inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) -> Action × Real) => sample.1)","missing":[],"search":"sampledpredictabledeviationfiltration_zero banditrlproof.exp3.sampledpredictabledeviationfiltration_zero theorem sampledpredictabledeviationfiltration_zero (env : type u) (action : type v) [measurablespace env] [measurablespace action] : sampledpredictabledeviationfiltration env action 0 = (inferinstance : measurablespace env).comap (fun sample : env × ((k : nat) -> action × real) => sample.1) theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableDeviationFiltration_succ","label":"sampledPredictableDeviationFiltration_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableDeviationFiltration_succ","description":"theorem sampledPredictableDeviationFiltration_succ (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] (i : Nat) : sampledPredictableDeviationFiltration Env Action (i + 1) = (inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real i)).comap (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe i sample.2))","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-819d5335832e","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4691,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:320"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableDeviationFiltration_succ (Env : Type u) (Action : Type v) [MeasurableSpace Env] [MeasurableSpace Action] (i : Nat) : sampledPredictableDeviationFiltration Env Action (i + 1) = (inferInstance : MeasurableSpace (Env × History.FinitePairHistory Action Real i)).comap (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe i sample.2))","missing":[],"search":"sampledpredictabledeviationfiltration_succ banditrlproof.exp3.sampledpredictabledeviationfiltration_succ theorem sampledpredictabledeviationfiltration_succ (env : type u) (action : type v) [measurablespace env] [measurablespace action] (i : nat) : sampledpredictabledeviationfiltration env action (i + 1) = (inferinstance : measurablespace (env × history.finitepairhistory action real i)).comap (fun sample : env × ((k : nat) -> action × real) => (sample.1, preorder.frestrictle i sample.2)) theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess","label":"sampledPredictableRealizedDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess","description":"noncomputable def sampledPredictableRealizedDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample noncomputable def…","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-8df9e108c827","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4692,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:329"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedDeviationProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat -> Env × ((k : Nat) -> Action × Real) -> Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample noncomputable def sampledPredictableRealizedDeviationProxy : Nat -> NNReal | 0 => 0 | _i + 1 => Concentration.intervalVarianceProxy 0 1 theorem sampledPredictableRealizedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Ac…","missing":[],"search":"sampledpredictablerealizeddeviationprocess banditrlproof.exp3.sampledpredictablerealizeddeviationprocess noncomputable def sampledpredictablerealizeddeviationprocess {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) : nat -> env × ((k : nat) -> action × real) -> real | 0, _sample => 0 | i + 1, sample => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample noncomputable def sampledpredictablerealizeddeviationproxy : nat -> nnreal | 0 => 0 | _i + 1 => concentration.intervalvarianceproxy 0 1 theorem sampledpredictablerealizeddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablerealizeddeviationprocess arms eta gamma loss) definition compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProxy","label":"sampledPredictableRealizedDeviationProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationProxy","description":"noncomputable def sampledPredictableRealizedDeviationProxy : Nat -> NNReal | 0 => 0 | _i + 1 => Concentration.intervalVarianceProxy 0 1 theorem sampledPredictableRealizedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg…","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-23e4a81e4162","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4693,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:339"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedDeviationProxy : Nat -> NNReal | 0 => 0 | _i + 1 => Concentration.intervalVarianceProxy 0 1 theorem sampledPredictableRealizedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPredictableRealizedDeviationProcess arms eta gamma loss)","missing":[],"search":"sampledpredictablerealizeddeviationproxy banditrlproof.exp3.sampledpredictablerealizeddeviationproxy noncomputable def sampledpredictablerealizeddeviationproxy : nat -> nnreal | 0 => 0 | _i + 1 => concentration.intervalvarianceproxy 0 1 theorem sampledpredictablerealizeddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablerealizeddeviationprocess arms eta gamma loss) definition compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess_stronglyAdapted","label":"sampledPredictableRealizedDeviationProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess_stronglyAdapted","description":"theorem sampledPredictableRealizedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltr…","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-dcd4b2a29d57","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4694,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:343"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviationProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPredictableRealizedDeviationProcess arms eta gamma loss)","missing":[],"search":"sampledpredictablerealizeddeviationprocess_stronglyadapted banditrlproof.exp3.sampledpredictablerealizeddeviationprocess_stronglyadapted theorem sampledpredictablerealizeddeviationprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablerealizeddeviationprocess arms eta gamma loss) theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess_sum_range_succ","label":"sampledPredictableRealizedDeviationProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess_sum_range_succ","description":"theorem sampledPredictableRealizedDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableRealizedDeviationProcess arms eta gamma los…","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-2afcd43575ba","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4695,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:476"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviationProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableRealizedDeviationProcess arms eta gamma loss i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample)","missing":[],"search":"sampledpredictablerealizeddeviationprocess_sum_range_succ banditrlproof.exp3.sampledpredictablerealizeddeviationprocess_sum_range_succ theorem sampledpredictablerealizeddeviationprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpredictablerealizeddeviationprocess arms eta gamma loss i sample) = (finset.range horizon).sum (fun i => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample) theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProxy_sum_range_succ","label":"sampledPredictableRealizedDeviationProxy_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationProxy_sum_range_succ","description":"theorem sampledPredictableRealizedDeviationProxy_sum_range_succ (horizon : Nat) : (Finset.range (horizon + 1)).sum sampledPredictableRealizedDeviationProxy = (horizon : NNReal) * Concentration.intervalVarianceProxy 0 1","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-f1cd4b7d617a","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4696,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:513"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviationProxy_sum_range_succ (horizon : Nat) : (Finset.range (horizon + 1)).sum sampledPredictableRealizedDeviationProxy = (horizon : NNReal) * Concentration.intervalVarianceProxy 0 1","missing":[],"search":"sampledpredictablerealizeddeviationproxy_sum_range_succ banditrlproof.exp3.sampledpredictablerealizeddeviationproxy_sum_range_succ theorem sampledpredictablerealizeddeviationproxy_sum_range_succ (horizon : nat) : (finset.range (horizon + 1)).sum sampledpredictablerealizeddeviationproxy = (horizon : nnreal) * concentration.intervalvarianceproxy 0 1 theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_ennreal","label":"sampledPredictableRealizedDeviation_sum_tail_ennreal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_ennreal","description":"Finite-horizon one-sided concentration for the realized predictable EXP3 deviation from its exploration-mixed conditional mean.","url":"../modules/banditrlproof-exp3realizeddeviationtail/index.html#decl-f0061caf9dc4","parent":"module:BanditRLProof.Exp3RealizedDeviationTail","order":4697,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedDeviationTail"],["Source","BanditRLProof/Exp3RealizedDeviationTail.lean:530"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_sum_tail_ennreal {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) {eps : Real} (heps : 0 <= eps) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment mu {sample | eps <= (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample)} <= ENNReal.ofReal (Real.exp (-eps ^ 2 / (2 * ((((horizon : NNReal) * Concentration.intervalVariance…","missing":[],"search":"sampledpredictablerealizeddeviation_sum_tail_ennreal banditrlproof.exp3.sampledpredictablerealizeddeviation_sum_tail_ennreal finite-horizon one-sided concentration for the realized predictable exp3 deviation from its exploration-mixed conditional mean. theorem compiled","shard":"modules/ab3091012b05494b.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedHighProbabilityRegretBudget","label":"sampledPredictableRealizedHighProbabilityRegretBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedHighProbabilityRegretBudget","description":"noncomputable def sampledPredictableRealizedHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (delta : Real) : Real","url":"../modules/banditrlproof-exp3realizedhighprobabilityregret/index.html#decl-85c1ee1f25db","parent":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","order":4698,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3RealizedHighProbabilityRegret.lean:24"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedHighProbabilityRegretBudget {Action : Type v} (arms : Finset Action) (eta gamma : Real) (horizon : Nat) (delta : Real) : Real","missing":[],"search":"sampledpredictablerealizedhighprobabilityregretbudget banditrlproof.exp3.sampledpredictablerealizedhighprobabilityregretbudget noncomputable def sampledpredictablerealizedhighprobabilityregretbudget {action : type v} (arms : finset action) (eta gamma : real) (horizon : nat) (delta : real) : real definition compiled","shard":"modules/469a5d2c5a4fca15.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedHighProbabilityRegret_tail_delta","label":"sampledPredictable_realizedHighProbabilityRegret_tail_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_realizedHighProbabilityRegret_tail_delta","description":"Raw three-event form of the generated realized EXP3 regret theorem. The same `delta` is used for the pure-q, comparator-estimator, and realized deviation tails, so the displayed failure probability is their three-term sum.","url":"../modules/banditrlproof-exp3realizedhighprobabilityregret/index.html#decl-87fb97397afb","parent":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","order":4699,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3RealizedHighProbabilityRegret.lean:35"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_realizedHighProbabilityRegret_tail_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableRealizedHighProbabilityRegretBudget arms eta gamma horizon delta <= (Finset.range hor…","missing":[],"search":"sampledpredictable_realizedhighprobabilityregret_tail_delta banditrlproof.exp3.sampledpredictable_realizedhighprobabilityregret_tail_delta raw three-event form of the generated realized exp3 regret theorem. the same `delta` is used for the pure-q, comparator-estimator, and realized deviation tails, so the displayed failure probability is their three-term sum. theorem compiled","shard":"modules/469a5d2c5a4fca15.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedHighProbabilityRegret_tail_total_delta","label":"sampledPredictable_realizedHighProbabilityRegret_tail_total_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_realizedHighProbabilityRegret_tail_total_delta","description":"Standard total-failure-probability form of generated realized EXP3 regret. Each of the three underlying confidence events receives `delta / 3`.","url":"../modules/banditrlproof-exp3realizedhighprobabilityregret/index.html#decl-5384d2dfca91","parent":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","order":4700,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedHighProbabilityRegret"],["Source","BanditRLProof/Exp3RealizedHighProbabilityRegret.lean:147"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_realizedHighProbabilityRegret_tail_total_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (horizon : Nat) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu {sample | sampledPredictableRealizedHighProbabilityRegretBudget arms eta gamma horizon (delta / 3) <= (Fins…","missing":[],"search":"sampledpredictable_realizedhighprobabilityregret_tail_total_delta banditrlproof.exp3.sampledpredictable_realizedhighprobabilityregret_tail_total_delta standard total-failure-probability form of generated realized exp3 regret. each of the three underlying confidence events receives `delta / 3`. theorem compiled","shard":"modules/469a5d2c5a4fca15.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment","label":"selectedLossCenteredSecondMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.selectedLossCenteredSecondMoment","description":"Exact centered second moment of a bounded loss under a finite action law.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-5acb34830474","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4701,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedLossCenteredSecondMoment {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) : Real","missing":[],"search":"selectedlosscenteredsecondmoment banditrlproof.exp3.selectedlosscenteredsecondmoment exact centered second moment of a bounded loss under a finite action law. definition compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_selectedLossCenteredSecondMoment","label":"measurable_selectedLossCenteredSecondMoment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_selectedLossCenteredSecondMoment","description":"theorem measurable_selectedLossCenteredSecondMoment {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Measurab…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-db28c6b79baa","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4702,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedLossCenteredSecondMoment {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Measurable (selectedLossCenteredSecondMoment arms prob loss)","missing":[],"search":"measurable_selectedlosscenteredsecondmoment banditrlproof.exp3.measurable_selectedlosscenteredsecondmoment theorem measurable_selectedlosscenteredsecondmoment {history : type u} {action : type v} [measurablespace history] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (prob loss : history → action → real) (source : measurablefiniteactiondistribution arms prob) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) : measurable (selectedlosscenteredsecondmoment arms prob loss) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_nonneg","label":"selectedLossCenteredSecondMoment_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.selectedLossCenteredSecondMoment_nonneg","description":"theorem selectedLossCenteredSecondMoment_nonneg {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) : 0 ≤ selectedLossCenteredSecondMoment arms prob loss history","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-4260e2d5942a","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4703,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:49"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem selectedLossCenteredSecondMoment_nonneg {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) : 0 ≤ selectedLossCenteredSecondMoment arms prob loss history","missing":[],"search":"selectedlosscenteredsecondmoment_nonneg banditrlproof.exp3.selectedlosscenteredsecondmoment_nonneg theorem selectedlosscenteredsecondmoment_nonneg {history : type u} {action : type v} [decidableeq action] (arms : finset action) (prob loss : history → action → real) (history : history) (hdist : finiteactiondistribution arms (prob history)) : 0 ≤ selectedlosscenteredsecondmoment arms prob loss history theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_one","label":"selectedLossCenteredSecondMoment_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_one","description":"A probability-weighted centered second moment of `[0,1]` losses is at most one.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-df3220339784","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4704,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:60"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem selectedLossCenteredSecondMoment_le_one {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (hloss : ∀ action ∈ arms, loss history action ∈ Set.Icc (0 : Real) 1) : selectedLossCenteredSecondMoment arms prob loss history ≤ 1","missing":[],"search":"selectedlosscenteredsecondmoment_le_one banditrlproof.exp3.selectedlosscenteredsecondmoment_le_one a probability-weighted centered second moment of `[0,1]` losses is at most one. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_lossMass","label":"selectedLossCenteredSecondMoment_le_lossMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_lossMass","description":"For `[0,1]` losses, the exact selected-loss variance is at most the unweighted armwise loss mass.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-28211ca2a42c","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4705,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:106"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem selectedLossCenteredSecondMoment_le_lossMass {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (hloss : ∀ action ∈ arms, loss history action ∈ Set.Icc (0 : Real) 1) : selectedLossCenteredSecondMoment arms prob loss history ≤ arms.sum fun action => loss history action","missing":[],"search":"selectedlosscenteredsecondmoment_le_lossmass banditrlproof.exp3.selectedlosscenteredsecondmoment_le_lossmass for `[0,1]` losses, the exact selected-loss variance is at most the unweighted armwise loss mass. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionSelectedLossDeviation_hasMGFUpperBoundAt_variance","label":"finiteActionSelectedLossDeviation_hasMGFUpperBoundAt_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionSelectedLossDeviation_hasMGFUpperBoundAt_variance","description":"Fixed-tilt MGF budget retaining the exact selected-loss variance.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-3ace2ce65af4","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4706,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:174"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionSelectedLossDeviation_hasMGFUpperBoundAt_variance {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mean := arms.sum fun action => prob history action * loss history action Concentration.HasMGFUpperBoundAt (fun selected => loss history selected - mean) tilt (tilt ^ 2 * selectedLossCenteredSecondMoment arms prob loss history) (finiteActionMeasure arms (prob history))","missing":[],"search":"finiteactionselectedlossdeviation_hasmgfupperboundat_variance banditrlproof.exp3.finiteactionselectedlossdeviation_hasmgfupperboundat_variance fixed-tilt mgf budget retaining the exact selected-loss variance. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionSelectedLossDeviation_compensated_hasMGFUpperBoundAt","label":"finiteActionSelectedLossDeviation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionSelectedLossDeviation_compensated_hasMGFUpperBoundAt","description":"theorem finiteActionSelectedLossDeviation_compensated_hasMGFUpperBoundAt {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbability…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-02d9c94a37ca","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4707,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:273"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionSelectedLossDeviation_compensated_hasMGFUpperBoundAt {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History → Action → Real) (history : History) (hdist : FiniteActionDistribution arms (prob history)) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mean := arms.sum fun action => prob history action * loss history action Concentration.HasMGFUpperBoundAt (fun selected => tilt * (loss history selected - mean) - tilt ^ 2 * selectedLossCenteredSecondMoment arms prob loss history) 1 0 (finiteActionMeasure arms (prob history))","missing":[],"search":"finiteactionselectedlossdeviation_compensated_hasmgfupperboundat banditrlproof.exp3.finiteactionselectedlossdeviation_compensated_hasmgfupperboundat theorem finiteactionselectedlossdeviation_compensated_hasmgfupperboundat {history : type u} {action : type v} [measurablespace history] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (prob loss : history → action → real) (history : history) (hdist : finiteactiondistribution arms (prob history)) (epsilon : real) (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) (tilt : real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mean := arms.sum fun action => prob history action * loss history action concentration.hasmgfupperboundat (fun selected => tilt * (loss history selected - mean) - tilt ^ 2 * selectedlosscenteredsecondmoment arms prob loss history) 1 0 (finiteactionmeasure arms (prob history)) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt","label":"sampledTrajectoryPredictableRealizedVarianceAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt","description":"Exact selected-loss predictable variance at an actual generated time.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-d77dc28f089c","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4708,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:298"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryPredictableRealizedVarianceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (t : Nat) : Env × ((k : Nat) → Action × Real) → Real","missing":[],"search":"sampledtrajectorypredictablerealizedvarianceat banditrlproof.exp3.sampledtrajectorypredictablerealizedvarianceat exact selected-loss predictable variance at an actual generated time. definition compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableRealizedVarianceAt","label":"measurable_sampledTrajectoryPredictableRealizedVarianceAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableRealizedVarianceAt","description":"theorem measurable_sampledTrajectoryPredictableRealizedVarianceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryPredictableReali…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-78fd72ce69ad","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4709,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:308"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectoryPredictableRealizedVarianceAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectorypredictablerealizedvarianceat banditrlproof.exp3.measurable_sampledtrajectorypredictablerealizedvarianceat theorem measurable_sampledtrajectorypredictablerealizedvarianceat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (t : nat) : measurable (sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss t) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_nonneg","label":"sampledTrajectoryPredictableRealizedVarianceAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_nonneg","description":"theorem sampledTrajectoryPredictableRealizedVarianceAt_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : 0…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-22cf468fd488","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4710,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:328"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableRealizedVarianceAt_nonneg {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : 0 ≤ sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss t sample","missing":[],"search":"sampledtrajectorypredictablerealizedvarianceat_nonneg banditrlproof.exp3.sampledtrajectorypredictablerealizedvarianceat_nonneg theorem sampledtrajectorypredictablerealizedvarianceat_nonneg {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) → action × real)) : 0 ≤ sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss t sample theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_le_lossMassAt","label":"sampledTrajectoryPredictableRealizedVarianceAt_le_lossMassAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_le_lossMassAt","description":"theorem sampledTrajectoryPredictableRealizedVarianceAt_le_lossMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Rea…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-6685df0c4c99","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4711,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:346"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableRealizedVarianceAt_le_lossMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss t sample ≤ arms.sum fun action => predictableLossAt loss t sample action","missing":[],"search":"sampledtrajectorypredictablerealizedvarianceat_le_lossmassat banditrlproof.exp3.sampledtrajectorypredictablerealizedvarianceat_le_lossmassat theorem sampledtrajectorypredictablerealizedvarianceat_le_lossmassat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) → action × real)) : sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss t sample ≤ arms.sum fun action => predictablelossat loss t sample action theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_le_one","label":"sampledTrajectoryPredictableRealizedVarianceAt_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_le_one","description":"Every generated selected-loss predictable variance is at most one.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-687f3fdb8391","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4712,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:375"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryPredictableRealizedVarianceAt_le_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) → Action × Real)) : sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss t sample ≤ 1","missing":[],"search":"sampledtrajectorypredictablerealizedvarianceat_le_one banditrlproof.exp3.sampledtrajectorypredictablerealizedvarianceat_le_one every generated selected-loss predictable variance is at most one. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.selectedLossDeviation_compensated_hasCondMGFUpperBoundAt_of_condDistrib","label":"selectedLossDeviation_compensated_hasCondMGFUpperBoundAt_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.selectedLossDeviation_compensated_hasCondMGFUpperBoundAt_of_condDistrib","description":"An identified finite conditional action law supplies a zero-budget MGF for the exact selected-loss variance-compensated increment.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-74ab5abce2d1","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4713,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:404"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem selectedLossDeviation_compensated_hasCondMGFUpperBoundAt_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Nonempty Omega] [mHistory : MeasurableSpace History] [StandardBorelSpace History] [mAction : MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega → History) (hhistory : Measurable history) (action : Omega → Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History → Action → Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (hloss : Measurable (fun input : History × Action => loss input.1 input.2)) (hloss_mem : ∀ history action, loss his…","missing":[],"search":"selectedlossdeviation_compensated_hascondmgfupperboundat_of_conddistrib banditrlproof.exp3.selectedlossdeviation_compensated_hascondmgfupperboundat_of_conddistrib an identified finite conditional action law supplies a zero-budget mgf for the exact selected-loss variance-compensated increment. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedLossCompensated_zero_hasCondMGFUpperBoundAt","label":"sampledPredictableSelectedLossCompensated_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSelectedLossCompensated_zero_hasCondMGFUpperBoundAt","description":"Initial generated selected-loss increment with its exact predictable variance compensation.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-6ced29a7c59e","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4714,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:691"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSelectedLossCompensated_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) → Action × Real) => sample.1)) measurable_fst.comap_le (fun sample => tilt * sampledT…","missing":[],"search":"sampledpredictableselectedlosscompensated_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpredictableselectedlosscompensated_zero_hascondmgfupperboundat initial generated selected-loss increment with its exact predictable variance compensation. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedLossCompensated_succ_hasCondMGFUpperBoundAt","label":"sampledPredictableSelectedLossCompensated_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableSelectedLossCompensated_succ_hasCondMGFUpperBoundAt","description":"Successor generated selected-loss increment with its exact predictable variance compensation.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-9b716efdafc6","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4715,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:753"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableSelectedLossCompensated_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (n : Nat) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) → Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace (Env × H…","missing":[],"search":"sampledpredictableselectedlosscompensated_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpredictableselectedlosscompensated_succ_hascondmgfupperboundat successor generated selected-loss increment with its exact predictable variance compensation. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensated_zero_hasCondMGFUpperBoundAt","label":"sampledPredictableRealizedCompensated_zero_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedCompensated_zero_hasCondMGFUpperBoundAt","description":"The deterministic-feedback realization transports the initial selected loss compensation to the observed realized-loss increment.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-e317c3b072dd","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4716,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:822"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedCompensated_zero_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace Env).comap (fun sample : Env × ((k : Nat) → Action × Real) => sample.1)) measurable_fst.comap_le (fun sample => tilt * sampledTraje…","missing":[],"search":"sampledpredictablerealizedcompensated_zero_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablerealizedcompensated_zero_hascondmgfupperboundat the deterministic-feedback realization transports the initial selected loss compensation to the observed realized-loss increment. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensated_succ_hasCondMGFUpperBoundAt","label":"sampledPredictableRealizedCompensated_succ_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedCompensated_succ_hasCondMGFUpperBoundAt","description":"The deterministic-feedback realization transports each successor selected loss compensation to the observed realized-loss increment.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-8309d41a4c3f","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4717,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:903"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedCompensated_succ_hasCondMGFUpperBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (n : Nat) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment let history := fun sample : Env × ((k : Nat) → Action × Real) => (sample.1, Preorder.frestrictLe n sample.2) Concentration.HasCondMGFUpperBoundAt ((inferInstance : MeasurableSpace (Env × Histo…","missing":[],"search":"sampledpredictablerealizedcompensated_succ_hascondmgfupperboundat banditrlproof.exp3.sampledpredictablerealizedcompensated_succ_hascondmgfupperboundat the deterministic-feedback realization transports each successor selected loss compensation to the observed realized-loss increment. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess","label":"sampledPredictableRealizedVarianceProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess","description":"Shift actual-time selected-loss variances by one so that the process is predictable for the generated deviation filtration.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-f55818acd811","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4718,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:990"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedVarianceProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) : Nat → Env × ((k : Nat) → Action × Real) → Real | 0, _sample => 0 | i + 1, sample => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample theorem measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable[sampledPredictableDeviationFiltration Env Action t] (sampledTrajectoryPredictableRealizedVarianc…","missing":[],"search":"sampledpredictablerealizedvarianceprocess banditrlproof.exp3.sampledpredictablerealizedvarianceprocess shift actual-time selected-loss variances by one so that the process is predictable for the generated deviation filtration. definition compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration","label":"measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration","description":"theorem measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable[sampledPredictableDevia…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-2840b335ea13","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4719,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:1001"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (t : Nat) : Measurable[sampledPredictableDeviationFiltration Env Action t] (sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss t)","missing":[],"search":"measurable_sampledtrajectorypredictablerealizedvarianceat_filtration banditrlproof.exp3.measurable_sampledtrajectorypredictablerealizedvarianceat_filtration theorem measurable_sampledtrajectorypredictablerealizedvarianceat_filtration {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (t : nat) : measurable[sampledpredictabledeviationfiltration env action t] (sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss t) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess_isPredictable","label":"sampledPredictableRealizedVarianceProcess_isPredictable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess_isPredictable","description":"theorem sampledPredictableRealizedVarianceProcess_isPredictable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) : IsPredictable (sampledPredictableDeviationFiltration Env…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-77db4b676e5c","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4720,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:1062"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedVarianceProcess_isPredictable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) : IsPredictable (sampledPredictableDeviationFiltration Env Action) (sampledPredictableRealizedVarianceProcess arms eta gamma loss)","missing":[],"search":"sampledpredictablerealizedvarianceprocess_ispredictable banditrlproof.exp3.sampledpredictablerealizedvarianceprocess_ispredictable theorem sampledpredictablerealizedvarianceprocess_ispredictable {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) : ispredictable (sampledpredictabledeviationfiltration env action) (sampledpredictablerealizedvarianceprocess arms eta gamma loss) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess_sum_range_succ","label":"sampledPredictableRealizedVarianceProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess_sum_range_succ","description":"theorem sampledPredictableRealizedVarianceProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableRealizedVarianceProcess arms eta gamma loss i…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-36e4701d762b","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4721,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:1082"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedVarianceProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableRealizedVarianceProcess arms eta gamma loss i sample) = (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample)","missing":[],"search":"sampledpredictablerealizedvarianceprocess_sum_range_succ banditrlproof.exp3.sampledpredictablerealizedvarianceprocess_sum_range_succ theorem sampledpredictablerealizedvarianceprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) → action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpredictablerealizedvarianceprocess arms eta gamma loss i sample) = (finset.range horizon).sum (fun i => sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss i sample) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_lossMass","label":"sampledPredictableRealizedVariance_sum_le_lossMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_lossMass","description":"theorem sampledPredictableRealizedVariance_sum_le_lossMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real))…","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-34c87828e5c6","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4722,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:1102"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedVariance_sum_le_lossMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample) ≤ (Finset.range horizon).sum (fun i => arms.sum fun action => predictableLossAt loss i sample action)","missing":[],"search":"sampledpredictablerealizedvariance_sum_le_lossmass banditrlproof.exp3.sampledpredictablerealizedvariance_sum_le_lossmass theorem sampledpredictablerealizedvariance_sum_le_lossmass {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) → action × real)) : (finset.range horizon).sum (fun i => sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss i sample) ≤ (finset.range horizon).sum (fun i => arms.sum fun action => predictablelossat loss i sample action) theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_horizon","label":"sampledPredictableRealizedVariance_sum_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_horizon","description":"The first `horizon` generated selected-loss predictable variances have deterministic total budget `horizon`.","url":"../modules/banditrlproof-exp3realizedpredictablevariance/index.html#decl-a36b8a77ce40","parent":"module:BanditRLProof.Exp3RealizedPredictableVariance","order":4723,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVariance"],["Source","BanditRLProof/Exp3RealizedPredictableVariance.lean:1123"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedVariance_sum_le_horizon {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 ≤ gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample) ≤ (horizon : Real)","missing":[],"search":"sampledpredictablerealizedvariance_sum_le_horizon banditrlproof.exp3.sampledpredictablerealizedvariance_sum_le_horizon the first `horizon` generated selected-loss predictable variances have deterministic total budget `horizon`. theorem compiled","shard":"modules/6d2efe210f9cede2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceGeometricRadius","label":"sampledRealizedPredictableVarianceGeometricRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedPredictableVarianceGeometricRadius","description":"The selected-loss predictable-variance radius at prefix `n+1` under the geometric confidence schedule.","url":"../modules/banditrlproof-exp3realizedpredictablevariancealltime/index.html#decl-ab5e9e14affa","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","order":4724,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceAllTime"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceAllTime.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedPredictableVarianceGeometricRadius (varianceBudget : Nat -> Real) (delta : Real) (n : Nat) : Real","missing":[],"search":"sampledrealizedpredictablevariancegeometricradius banditrlproof.exp3.sampledrealizedpredictablevariancegeometricradius the selected-loss predictable-variance radius at prefix `n+1` under the geometric confidence schedule. definition compiled","shard":"modules/3dec3ce7e621a217.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationAllTimeFailureSet","label":"sampledPredictableRealizedDeviationAllTimeFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviationAllTimeFailureSet","description":"Countable failure event over every positive prefix of one generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3realizedpredictablevariancealltime/index.html#decl-c1a2c8b900af","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","order":4725,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceAllTime"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceAllTime.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedDeviationAllTimeFailureSet {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (varianceBudget : Nat -> Real) (delta : Real) : Set (Env × ((k : Nat) -> Action × Real))","missing":[],"search":"sampledpredictablerealizeddeviationalltimefailureset banditrlproof.exp3.sampledpredictablerealizeddeviationalltimefailureset countable failure event over every positive prefix of one generated exp3 trajectory. definition compiled","shard":"modules/3dec3ce7e621a217.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mem_sampledPredictableRealizedDeviationAllTimeFailureSet_iff","label":"mem_sampledPredictableRealizedDeviationAllTimeFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mem_sampledPredictableRealizedDeviationAllTimeFailureSet_iff","description":"Membership is failure at at least one positive prefix.","url":"../modules/banditrlproof-exp3realizedpredictablevariancealltime/index.html#decl-4d48560789ad","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","order":4726,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceAllTime"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceAllTime.lean:51"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mem_sampledPredictableRealizedDeviationAllTimeFailureSet_iff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (varianceBudget : Nat -> Real) (delta : Real) (sample : Env × ((k : Nat) -> Action × Real)) : sample ∈ sampledPredictableRealizedDeviationAllTimeFailureSet arms eta gamma loss varianceBudget delta ↔ ∃ n, sampledRealizedPredictableVarianceGeometricRadius varianceBudget delta n <= (Finset.range (n + 1)).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample) ∧ (Finset.range (n + 1)).sum (fun i => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample) <= varianceBudget n","missing":[],"search":"mem_sampledpredictablerealizeddeviationalltimefailureset_iff banditrlproof.exp3.mem_sampledpredictablerealizeddeviationalltimefailureset_iff membership is failure at at least one positive prefix. theorem compiled","shard":"modules/3dec3ce7e621a217.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","label":"measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","description":"On one generated EXP3 trajectory law, the joint deviation/predictable-variance failures over all positive prefixes have total mass at most the outer confidence budget.","url":"../modules/banditrlproof-exp3realizedpredictablevariancealltime/index.html#decl-3b915a242637","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","order":4727,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceAllTime"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceAllTime.lean:74"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (varianceBudget : Nat -> Real) (hvarianceBudget : forall n, 0 < varianceBudget n) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu (sampledPredictableRealizedDeviationAllTimeFailureSet arms eta gamma loss varianceBudget delta) <= ENNReal.ofReal delta","missing":[],"search":"measure_sampledpredictablerealizeddeviationalltimefailureset_le banditrlproof.exp3.measure_sampledpredictablerealizeddeviationalltimefailureset_le on one generated exp3 trajectory law, the joint deviation/predictable-variance failures over all positive prefixes have total mass at most the outer confidence budget. theorem compiled","shard":"modules/3dec3ce7e621a217.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceMaximalRadius","label":"sampledRealizedPredictableVarianceMaximalRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedPredictableVarianceMaximalRadius","description":"Equal-share radius for indices `t < horizon`, hence prefix lengths `1` through `horizon` inclusive.","url":"../modules/banditrlproof-exp3realizedpredictablevariancemaximal/index.html#decl-21f4244e5220","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","order":4728,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceMaximal"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceMaximal.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedPredictableVarianceMaximalRadius (horizon : Nat) (varianceBudget delta : Real) : Real","missing":[],"search":"sampledrealizedpredictablevariancemaximalradius banditrlproof.exp3.sampledrealizedpredictablevariancemaximalradius equal-share radius for indices `t < horizon`, hence prefix lengths `1` through `horizon` inclusive. definition compiled","shard":"modules/f84a454b7f2d26d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_prefix_max_tail_predictableVariance_delta","label":"sampledPredictableRealizedDeviation_prefix_max_tail_predictableVariance_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_prefix_max_tail_predictableVariance_delta","description":"Finite maximal predictable-variance tail for the realized selected-loss deviation over every positive prefix of the generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3realizedpredictablevariancemaximal/index.html#decl-d35944598dbb","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","order":4729,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceMaximal"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceMaximal.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_prefix_max_tail_predictableVariance_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon : 0 < horizon) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu (⋃ t ∈ Finset.range horizon, {sample | sampledRealizedPredictableVarianceMaximalRadius horizon varianceBudget delta…","missing":[],"search":"sampledpredictablerealizeddeviation_prefix_max_tail_predictablevariance_delta banditrlproof.exp3.sampledpredictablerealizeddeviation_prefix_max_tail_predictablevariance_delta finite maximal predictable-variance tail for the realized selected-loss deviation over every positive prefix of the generated exp3 trajectory. theorem compiled","shard":"modules/f84a454b7f2d26d3.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess","label":"sampledPredictableRealizedCompensatedProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess","description":"Shifted exact-variance compensated realized-loss process.","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html#decl-12a5a3e6c782","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","order":4730,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceTail"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean:19"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledPredictableRealizedCompensatedProcess {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma tilt : Real) (loss : PredictableLossVector Env Action) : Nat → Env × ((k : Nat) → Action × Real) → Real","missing":[],"search":"sampledpredictablerealizedcompensatedprocess banditrlproof.exp3.sampledpredictablerealizedcompensatedprocess shifted exact-variance compensated realized-loss process. definition compiled","shard":"modules/5a031350159eb85e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess_stronglyAdapted","label":"sampledPredictableRealizedCompensatedProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess_stronglyAdapted","description":"theorem sampledPredictableRealizedCompensatedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma tilt : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFil…","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html#decl-f7e873803cc7","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","order":4731,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceTail"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean:31"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedCompensatedProcess_stronglyAdapted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma tilt : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) : StronglyAdapted (sampledPredictableDeviationFiltration Env Action) (sampledPredictableRealizedCompensatedProcess arms eta gamma tilt loss)","missing":[],"search":"sampledpredictablerealizedcompensatedprocess_stronglyadapted banditrlproof.exp3.sampledpredictablerealizedcompensatedprocess_stronglyadapted theorem sampledpredictablerealizedcompensatedprocess_stronglyadapted {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma tilt : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) : stronglyadapted (sampledpredictabledeviationfiltration env action) (sampledpredictablerealizedcompensatedprocess arms eta gamma tilt loss) theorem compiled","shard":"modules/5a031350159eb85e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess_sum_range_succ","label":"sampledPredictableRealizedCompensatedProcess_sum_range_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess_sum_range_succ","description":"theorem sampledPredictableRealizedCompensatedProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma tilt : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableRealizedCompensatedProcess arms eta g…","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html#decl-21fe53ffa5a4","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","order":4732,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceTail"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean:53"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedCompensatedProcess_sum_range_succ {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma tilt : Real) (loss : PredictableLossVector Env Action) (horizon : Nat) (sample : Env × ((k : Nat) → Action × Real)) : (Finset.range (horizon + 1)).sum (fun i => sampledPredictableRealizedCompensatedProcess arms eta gamma tilt loss i sample) = tilt * (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample) - tilt ^ 2 * (Finset.range horizon).sum (fun i => sampledTrajectoryPredictableRealizedVarianceAt arms eta gamma loss i sample)","missing":[],"search":"sampledpredictablerealizedcompensatedprocess_sum_range_succ banditrlproof.exp3.sampledpredictablerealizedcompensatedprocess_sum_range_succ theorem sampledpredictablerealizedcompensatedprocess_sum_range_succ {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (eta gamma tilt : real) (loss : predictablelossvector env action) (horizon : nat) (sample : env × ((k : nat) → action × real)) : (finset.range (horizon + 1)).sum (fun i => sampledpredictablerealizedcompensatedprocess arms eta gamma tilt loss i sample) = tilt * (finset.range horizon).sum (fun i => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample) - tilt ^ 2 * (finset.range horizon).sum (fun i => sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss i sample) theorem compiled","shard":"modules/5a031350159eb85e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_fixedTilt","label":"sampledPredictableRealizedDeviation_sum_tail_predictableVariance_fixedTilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_fixedTilt","description":"Fixed-tilt upper tail retaining the random cumulative selected-loss predictable variance.","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html#decl-ab792c19df4f","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","order":4733,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceTail"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean:75"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_sum_tail_predictableVariance_fixedTilt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (tilt : Real) (htilt_nonneg : 0 ≤ tilt) (htilt_le_one : tilt ≤ 1) (threshold varianceBudget : Real) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | threshold ≤ (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt arms eta gamma loss i sample) ∧ (Finset…","missing":[],"search":"sampledpredictablerealizeddeviation_sum_tail_predictablevariance_fixedtilt banditrlproof.exp3.sampledpredictablerealizeddeviation_sum_tail_predictablevariance_fixedtilt fixed-tilt upper tail retaining the random cumulative selected-loss predictable variance. theorem compiled","shard":"modules/5a031350159eb85e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceRadius","label":"sampledRealizedPredictableVarianceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedPredictableVarianceRadius","description":"Bernstein radius for realized selected-loss deviation under a pathwise predictable-variance budget.","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html#decl-108765b84ed3","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","order":4734,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceTail"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean:168"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedPredictableVarianceRadius (varianceBudget delta : Real) : Real","missing":[],"search":"sampledrealizedpredictablevarianceradius banditrlproof.exp3.sampledrealizedpredictablevarianceradius bernstein radius for realized selected-loss deviation under a pathwise predictable-variance budget. definition compiled","shard":"modules/5a031350159eb85e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_delta","label":"sampledPredictableRealizedDeviation_sum_tail_predictableVariance_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_delta","description":"theorem sampledPredictableRealizedDeviation_sum_tail_predictableVariance_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (h…","url":"../modules/banditrlproof-exp3realizedpredictablevariancetail/index.html#decl-6d56e0887b65","parent":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","order":4735,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedPredictableVarianceTail"],["Source","BanditRLProof/Exp3RealizedPredictableVarianceTail.lean:173"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedDeviation_sum_tail_predictableVariance_delta {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (varianceBudget delta : Real) (hvarianceBudget : 0 < varianceBudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledRealizedPredictableVarianceRadius varianceBudget delta ≤ (Finset.range horizon).sum (fun i => sampledTrajectoryRealizedDeviationAt a…","missing":[],"search":"sampledpredictablerealizeddeviation_sum_tail_predictablevariance_delta banditrlproof.exp3.sampledpredictablerealizeddeviation_sum_tail_predictablevariance_delta theorem sampledpredictablerealizeddeviation_sum_tail_predictablevariance_delta {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [nonempty env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isprobabilitymeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_pos : 0 < gamma) (hgamma_le_one : gamma ≤ 1) (loss : predictablelossvector env action) (horizon : nat) (variancebudget delta : real) (hvariancebudget : 0 < variancebudget) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_pos.le hgamma_le_one loss.environment mu {sample | sampledrealizedpredictablevarianceradius variancebudget delta ≤ (finset.range horizon).sum (fun i => sampledtrajectoryrealizeddeviationat arms eta gamma loss i sample) ∧ (finset.range horizon).sum (fun i => sampledtrajectorypredictablerealizedvarianceat arms eta gamma loss i sample) ≤ variancebudget} ≤ ennreal.ofreal delta theorem compiled","shard":"modules/5a031350159eb85e.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedLossAt","label":"sampledTrajectoryRealizedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryRealizedLossAt","description":"The scalar loss realized at an actual generated-trajectory time.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-91e49809a0e3","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4736,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def sampledTrajectoryRealizedLossAt {Env : Type u} {Action : Type v} (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledtrajectoryrealizedlossat banditrlproof.exp3.sampledtrajectoryrealizedlossat the scalar loss realized at an actual generated-trajectory time. definition compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedLossAt_ae_eq_selectedPredictable","label":"sampledTrajectoryRealizedLossAt_ae_eq_selectedPredictable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryRealizedLossAt_ae_eq_selectedPredictable","description":"Generated predictable feedback identifies the realized scalar loss with the predictable coordinate selected by the sampled action.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-14229fe35dbd","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4737,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:27"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryRealizedLossAt_ae_eq_selectedPredictable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment sampledTrajectoryRealizedLossAt t =ᵐ[mu] (fun sample => predictableLossAt loss t sample (sample.2 t).1)","missing":[],"search":"sampledtrajectoryrealizedlossat_ae_eq_selectedpredictable banditrlproof.exp3.sampledtrajectoryrealizedlossat_ae_eq_selectedpredictable generated predictable feedback identifies the realized scalar loss with the predictable coordinate selected by the sampled action. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectorySelectedPredictableLossAt","label":"measurable_sampledTrajectorySelectedPredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledTrajectorySelectedPredictableLossAt","description":"The selected predictable coordinate is measurable on the full trajectory space, even though the coordinate itself varies with the sampled action.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-038ab30ee41d","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4738,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:58"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledTrajectorySelectedPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => predictableLossAt loss t sample (sample.2 t).1)","missing":[],"search":"measurable_sampledtrajectoryselectedpredictablelossat banditrlproof.exp3.measurable_sampledtrajectoryselectedpredictablelossat the selected predictable coordinate is measurable on the full trajectory space, even though the coordinate itself varies with the sampled action. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectorySelectedPredictableLossAt_mem_unitInterval","label":"sampledTrajectorySelectedPredictableLossAt_mem_unitInterval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectorySelectedPredictableLossAt_mem_unitInterval","description":"theorem sampledTrajectorySelectedPredictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : predictableLossAt loss t sample (sample.2 t).1 ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-b49f1b77efb7","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4739,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:80"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectorySelectedPredictableLossAt_mem_unitInterval {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : predictableLossAt loss t sample (sample.2 t).1 ∈ Set.Icc (0 : Real) 1","missing":[],"search":"sampledtrajectoryselectedpredictablelossat_mem_unitinterval banditrlproof.exp3.sampledtrajectoryselectedpredictablelossat_mem_unitinterval theorem sampledtrajectoryselectedpredictablelossat_mem_unitinterval {env : type u} {action : type v} [measurablespace env] [measurablespace action] (loss : predictablelossvector env action) (t : nat) (sample : env × ((k : nat) -> action × real)) : predictablelossat loss t sample (sample.2 t).1 ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectorySelectedPredictableLossAt","label":"integrable_sampledTrajectorySelectedPredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_sampledTrajectorySelectedPredictableLossAt","description":"theorem integrable_sampledTrajectorySelectedPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (fun sample => predictableLossAt loss t sample (sample.2 t).1) mu","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-51c302c29f56","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4740,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:93"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledTrajectorySelectedPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (loss : PredictableLossVector Env Action) (t : Nat) : Integrable (fun sample => predictableLossAt loss t sample (sample.2 t).1) mu","missing":[],"search":"integrable_sampledtrajectoryselectedpredictablelossat banditrlproof.exp3.integrable_sampledtrajectoryselectedpredictablelossat theorem integrable_sampledtrajectoryselectedpredictablelossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] (mu : measure (env × ((k : nat) -> action × real))) [isfinitemeasure mu] (loss : predictablelossvector env action) (t : nat) : integrable (fun sample => predictablelossat loss t sample (sample.2 t).1) mu theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryRealizedLossAt","label":"integrable_sampledTrajectoryRealizedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_sampledTrajectoryRealizedLossAt","description":"theorem integrable_sampledTrajectoryRealizedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamm…","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-1b17ccca99c6","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4741,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:108"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledTrajectoryRealizedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment Integrable (sampledTrajectoryRealizedLossAt t) mu","missing":[],"search":"integrable_sampledtrajectoryrealizedlossat banditrlproof.exp3.integrable_sampledtrajectoryrealizedlossat theorem integrable_sampledtrajectoryrealizedlossat {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : predictablelossvector env action) (t : nat) : let mu := prior ⊗ₘ sampledimportanceweightedtrajectorykernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integrable (sampledtrajectoryrealizedlossat t) mu theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedInitial_integral_eq_explored","label":"sampledPredictableRealizedInitial_integral_eq_explored","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedInitial_integral_eq_explored","description":"At time zero, the expected realized loss is the exploration-distribution mixed predictable loss.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-f2b380d39627","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4742,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:130"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedInitial_integral_eq_explored {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integral mu (sampledTrajectoryRealizedLossAt 0) = integral mu (sampledTrajectoryExploredPredictableLossAt arms eta gamma loss 0)","missing":[],"search":"sampledpredictablerealizedinitial_integral_eq_explored banditrlproof.exp3.sampledpredictablerealizedinitial_integral_eq_explored at time zero, the expected realized loss is the exploration-distribution mixed predictable loss. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedSuccessor_integral_eq_explored","label":"sampledPredictableRealizedSuccessor_integral_eq_explored","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedSuccessor_integral_eq_explored","description":"At a successor time, conditioning on the retained environment/history prefix converts the expected selected predictable coordinate to the `p_t` finite sum.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-292530eff83e","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4743,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:217"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedSuccessor_integral_eq_explored {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integral mu (sampledTrajectoryRealizedLossAt (n + 1)) = integral mu (sampledTrajectoryExploredPredictableLossAt arms eta gamma loss (n + 1))","missing":[],"search":"sampledpredictablerealizedsuccessor_integral_eq_explored banditrlproof.exp3.sampledpredictablerealizedsuccessor_integral_eq_explored at a successor time, conditioning on the retained environment/history prefix converts the expected selected predictable coordinate to the `p_t` finite sum. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedAt_integral_eq_explored","label":"sampledPredictableRealizedAt_integral_eq_explored","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealizedAt_integral_eq_explored","description":"Every actual time has the realized-to-`p_t`-mixed first-moment identity.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-e84489e23ddd","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4744,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:320"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealizedAt_integral_eq_explored {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integral mu (sampledTrajectoryRealizedLossAt t) = integral mu (sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t)","missing":[],"search":"sampledpredictablerealizedat_integral_eq_explored banditrlproof.exp3.sampledpredictablerealizedat_integral_eq_explored every actual time has the realized-to-`p_t`-mixed first-moment identity. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictableRealized_finiteHorizon_integral_eq_explored","label":"sampledPredictableRealized_finiteHorizon_integral_eq_explored","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictableRealized_finiteHorizon_integral_eq_explored","description":"The realized and exploration-mixed predictable cumulative losses have the same expectation over every finite horizon.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-f21f4416deec","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4745,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:346"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictableRealized_finiteHorizon_integral_eq_explored {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample)) = integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample))","missing":[],"search":"sampledpredictablerealized_finitehorizon_integral_eq_explored banditrlproof.exp3.sampledpredictablerealized_finitehorizon_integral_eq_explored the realized and exploration-mixed predictable cumulative losses have the same expectation over every finite horizon. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le","label":"sampledPredictable_realizedExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le","description":"Unoptimized expected regret for the scalar losses actually realized by the sampled predictable EXP3 trajectory.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-f77869bc8eb5","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4746,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:396"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_realizedExpectedRegret_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)) <= Real.log arms.…","missing":[],"search":"sampledpredictable_realizedexpectedregret_le banditrlproof.exp3.sampledpredictable_realizedexpectedregret_le unoptimized expected regret for the scalar losses actually realized by the sampled predictable exp3 trajectory. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le_four_mul_sqrt","label":"sampledPredictable_realizedExpectedRegret_le_four_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le_four_mul_sqrt","description":"Tuned large-horizon expected regret for the scalar loss actually realized by the sampled predictable EXP3 trajectory.","url":"../modules/banditrlproof-exp3realizedregret/index.html#decl-db75ad9e83ed","parent":"module:BanditRLProof.Exp3RealizedRegret","order":4747,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegret"],["Source","BanditRLProof/Exp3RealizedRegret.lean:457"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_realizedExpectedRegret_le_four_mul_sqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon : Nat) (hhorizon_pos : 0 < horizon) (hscale : 4 * (arms.card : Real) * Real.log arms.card <= (horizon : Real)) (comparator : Action) (hcomparator : comparator ∈ arms) : let K := (arms.card : Real) let T := (horizon : Real) let mu := prior ⊗ₘ tunedPredictableTrajectoryKernel arms harms hcard_two loss horizon hhorizon_pos hscale integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.ran…","missing":[],"search":"sampledpredictable_realizedexpectedregret_le_four_mul_sqrt banditrlproof.exp3.sampledpredictable_realizedexpectedregret_le_four_mul_sqrt tuned large-horizon expected regret for the scalar loss actually realized by the sampled predictable exp3 trajectory. theorem compiled","shard":"modules/33851dc85a03b8c2.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation","label":"sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation","description":"Exact finite-prefix decomposition of realized selected-loss regret into exploration-mixed predictable regret and realized deviation.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html#decl-6f038e36fd29","parent":"module:BanditRLProof.Exp3RealizedRegretAllTime","order":4748,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegretAllTime"],["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean:23"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator) = ((Finset.range horizon).sum (fun t => sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)) + (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedDeviationAt arms eta gamma loss t sample)","missing":[],"search":"sampledtrajectoryrealizedregret_eq_predictableregret_add_realizeddeviation banditrlproof.exp3.sampledtrajectoryrealizedregret_eq_predictableregret_add_realizeddeviation exact finite-prefix decomposition of realized selected-loss regret into exploration-mixed predictable regret and realized deviation. theorem compiled","shard":"modules/b2fdeaca9a90732f.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeBudget","label":"sampledRealizedRegretGeometricAllTimeBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeBudget","description":"At prefix `n+1`, add the predictable-regret and pure realized-deviation schedules after assigning half of the total confidence budget to each event family.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html#decl-c9c001880892","parent":"module:BanditRLProof.Exp3RealizedRegretAllTime","order":4749,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedRegretAllTime"],["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean:47"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedRegretGeometricAllTimeBudget {Action : Type v} (arms : Finset Action) (eta gamma delta : Real) (n : Nat) : Real","missing":[],"search":"sampledrealizedregretgeometricalltimebudget banditrlproof.exp3.sampledrealizedregretgeometricalltimebudget at prefix `n+1`, add the predictable-regret and pure realized-deviation schedules after assigning half of the total confidence budget to each event family. definition compiled","shard":"modules/b2fdeaca9a90732f.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeFailureSet","label":"sampledRealizedRegretGeometricAllTimeFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeFailureSet","description":"Countable realized selected-loss regret failure event over every positive prefix of one generated EXP3 trajectory.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html#decl-f3c0ae4de8ff","parent":"module:BanditRLProof.Exp3RealizedRegretAllTime","order":4750,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RealizedRegretAllTime"],["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean:56"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledRealizedRegretGeometricAllTimeFailureSet {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (delta : Real) : Set (Env × ((k : Nat) -> Action × Real))","missing":[],"search":"sampledrealizedregretgeometricalltimefailureset banditrlproof.exp3.sampledrealizedregretgeometricalltimefailureset countable realized selected-loss regret failure event over every positive prefix of one generated exp3 trajectory. definition compiled","shard":"modules/b2fdeaca9a90732f.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.mem_sampledRealizedRegretGeometricAllTimeFailureSet_iff","label":"mem_sampledRealizedRegretGeometricAllTimeFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.mem_sampledRealizedRegretGeometricAllTimeFailureSet_iff","description":"Membership is realized selected-loss regret failure at at least one positive prefix.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html#decl-ab345cf9e2f7","parent":"module:BanditRLProof.Exp3RealizedRegretAllTime","order":4751,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegretAllTime"],["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean:72"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem mem_sampledRealizedRegretGeometricAllTimeFailureSet_iff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (delta : Real) (sample : Env × ((k : Nat) -> Action × Real)) : sample ∈ sampledRealizedRegretGeometricAllTimeFailureSet arms eta gamma loss comparator delta ↔ ∃ n, sampledRealizedRegretGeometricAllTimeBudget arms eta gamma delta n <= (Finset.range (n + 1)).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range (n + 1)).sum (fun t => predictableLossAt loss t sample comparator)","missing":[],"search":"mem_sampledrealizedregretgeometricalltimefailureset_iff banditrlproof.exp3.mem_sampledrealizedregretgeometricalltimefailureset_iff membership is realized selected-loss regret failure at at least one positive prefix. theorem compiled","shard":"modules/b2fdeaca9a90732f.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeFailureSet_subset","label":"sampledRealizedRegretGeometricAllTimeFailureSet_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeFailureSet_subset","description":"A combined realized-regret crossing forces either a predictable-regret crossing or a pure realized-deviation crossing at the same prefix.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html#decl-3ed2e0810496","parent":"module:BanditRLProof.Exp3RealizedRegretAllTime","order":4752,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegretAllTime"],["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean:91"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledRealizedRegretGeometricAllTimeFailureSet_subset {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (loss : PredictableLossVector Env Action) (comparator : Action) (delta : Real) : sampledRealizedRegretGeometricAllTimeFailureSet arms eta gamma loss comparator delta ⊆ sampledPredictableRegretGeometricAllTimeFailureSet arms eta gamma loss comparator (delta / 2) ∪ sampledRealizedDeviationGeometricAllTimeFailureSet arms eta gamma loss (delta / 2)","missing":[],"search":"sampledrealizedregretgeometricalltimefailureset_subset banditrlproof.exp3.sampledrealizedregretgeometricalltimefailureset_subset a combined realized-regret crossing forces either a predictable-regret crossing or a pure realized-deviation crossing at the same prefix. theorem compiled","shard":"modules/b2fdeaca9a90732f.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","label":"measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","description":"On one fixed generated EXP3 process and against one fixed supported comparator, realized selected-loss regret stays below the sum of the scheduled predictable-regret and realized-deviation budgets at every positive prefix, outside a set of outer measure at most the total confidence budget.","url":"../modules/banditrlproof-exp3realizedregretalltime/index.html#decl-173071bb2290","parent":"module:BanditRLProof.Exp3RealizedRegretAllTime","order":4753,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RealizedRegretAllTime"],["Source","BanditRLProof/Exp3RealizedRegretAllTime.lean:145"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measure_sampledRealizedRegretGeometricAllTimeFailureSet_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [Nonempty Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_pos : 0 < gamma) (hgamma_lt_one : gamma < 1) (loss : PredictableLossVector Env Action) (comparator : Action) (hcomparator : comparator ∈ arms) (delta : Real) (hdelta : 0 < delta) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_pos.le hgamma_lt_one.le loss.environment mu (sampledRealizedRegretGeometricAllTimeFailureSet arms eta gamma loss comparator delta) <= ENNReal.ofReal delta","missing":[],"search":"measure_sampledrealizedregretgeometricalltimefailureset_le banditrlproof.exp3.measure_sampledrealizedregretgeometricalltimefailureset_le on one fixed generated exp3 process and against one fixed supported comparator, realized selected-loss regret stays below the sum of the scheduled predictable-regret and realized-deviation budgets at every positive prefix, outside a set of outer measure at most the total confidence budget. theorem compiled","shard":"modules/b2fdeaca9a90732f.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.historyWeight","label":"historyWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.historyWeight","description":"Exponential weight generated by a finite-history cumulative score.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-f5aeecfc6ed2","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4754,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:26"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def historyWeight {Action History : Type*} (eta : Real) (score : History -> Action -> Real) (history : History) (action : Action) : Real","missing":[],"search":"historyweight banditrlproof.exp3.historyweight exponential weight generated by a finite-history cumulative score. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.historyTotalWeight","label":"historyTotalWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.historyTotalWeight","description":"Total exponential weight on the explicit finite support.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-8fdc0d5f20dc","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4755,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:32"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def historyTotalWeight {Action History : Type*} (arms : Finset Action) (eta : Real) (score : History -> Action -> Real) (history : History) : Real","missing":[],"search":"historytotalweight banditrlproof.exp3.historytotalweight total exponential weight on the explicit finite support. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.normalizedHistoryDistribution","label":"normalizedHistoryDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.normalizedHistoryDistribution","description":"Normalized exponential history policy before explicit exploration mixing.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-236d9a86f7e8","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4756,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:38"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedHistoryDistribution {Action History : Type*} (arms : Finset Action) (eta : Real) (score : History -> Action -> Real) (history : History) (action : Action) : Real","missing":[],"search":"normalizedhistorydistribution banditrlproof.exp3.normalizedhistorydistribution normalized exponential history policy before explicit exploration mixing. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredHistoryDistribution","label":"exploredHistoryDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredHistoryDistribution","description":"EXP3 policy: normalized exponential weights mixed with uniform exploration.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-d56daf391939","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4757,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:46"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exploredHistoryDistribution {Action History : Type*} (arms : Finset Action) (eta gamma : Real) (score : History -> Action -> Real) (history : History) (action : Action) : Real","missing":[],"search":"exploredhistorydistribution banditrlproof.exp3.exploredhistorydistribution exp3 policy: normalized exponential weights mixed with uniform exploration. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.historyTotalWeight_pos","label":"historyTotalWeight_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.historyTotalWeight_pos","description":"theorem historyTotalWeight_pos {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (history : History) : 0 < historyTotalWeight arms eta score history","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-27a7aedf8139","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4758,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:53"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem historyTotalWeight_pos {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (history : History) : 0 < historyTotalWeight arms eta score history","missing":[],"search":"historytotalweight_pos banditrlproof.exp3.historytotalweight_pos theorem historytotalweight_pos {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta : real) (score : history -> action -> real) (history : history) : 0 < historytotalweight arms eta score history theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.normalizedHistoryDistribution_nonneg","label":"normalizedHistoryDistribution_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.normalizedHistoryDistribution_nonneg","description":"theorem normalizedHistoryDistribution_nonneg {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (history : History) (action : Action) : 0 <= normalizedHistoryDistribution arms eta score history action","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-6cf85693daea","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4759,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:63"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem normalizedHistoryDistribution_nonneg {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (history : History) (action : Action) : 0 <= normalizedHistoryDistribution arms eta score history action","missing":[],"search":"normalizedhistorydistribution_nonneg banditrlproof.exp3.normalizedhistorydistribution_nonneg theorem normalizedhistorydistribution_nonneg {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta : real) (score : history -> action -> real) (history : history) (action : action) : 0 <= normalizedhistorydistribution arms eta score history action theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_normalizedHistoryDistribution","label":"sum_normalizedHistoryDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_normalizedHistoryDistribution","description":"theorem sum_normalizedHistoryDistribution {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (history : History) : arms.sum (normalizedHistoryDistribution arms eta score history) = 1","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-6e73e6dfae65","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4760,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:72"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_normalizedHistoryDistribution {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (history : History) : arms.sum (normalizedHistoryDistribution arms eta score history) = 1","missing":[],"search":"sum_normalizedhistorydistribution banditrlproof.exp3.sum_normalizedhistorydistribution theorem sum_normalizedhistorydistribution {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta : real) (score : history -> action -> real) (history : history) : arms.sum (normalizedhistorydistribution arms eta score history) = 1 theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredHistoryDistribution_nonneg","label":"exploredHistoryDistribution_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredHistoryDistribution_nonneg","description":"theorem exploredHistoryDistribution_nonneg {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (history : History) (action : Action) : 0 <= exploredHistoryDistribution arms eta gamma score history action","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-981da004d600","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4761,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:87"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exploredHistoryDistribution_nonneg {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (history : History) (action : Action) : 0 <= exploredHistoryDistribution arms eta gamma score history action","missing":[],"search":"exploredhistorydistribution_nonneg banditrlproof.exp3.exploredhistorydistribution_nonneg theorem exploredhistorydistribution_nonneg {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (score : history -> action -> real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (history : history) (action : action) : 0 <= exploredhistorydistribution arms eta gamma score history action theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sum_exploredHistoryDistribution","label":"sum_exploredHistoryDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sum_exploredHistoryDistribution","description":"theorem sum_exploredHistoryDistribution {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (history : History) : arms.sum (exploredHistoryDistribution arms eta gamma score history) = 1","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-f4b6c9bf3cc1","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4762,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:100"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sum_exploredHistoryDistribution {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (history : History) : arms.sum (exploredHistoryDistribution arms eta gamma score history) = 1","missing":[],"search":"sum_exploredhistorydistribution banditrlproof.exp3.sum_exploredhistorydistribution theorem sum_exploredhistorydistribution {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (score : history -> action -> real) (history : history) : arms.sum (exploredhistorydistribution arms eta gamma score history) = 1 theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredHistoryDistribution_floor","label":"exploredHistoryDistribution_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredHistoryDistribution_floor","description":"theorem exploredHistoryDistribution_floor {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (hgamma_le_one : gamma <= 1) (history : History) (action : Action) : gamma / (arms.card : Real) <= exploredHistoryDistribution arms eta gamma score history action","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-93745aac5d08","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4763,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:115"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exploredHistoryDistribution_floor {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (hgamma_le_one : gamma <= 1) (history : History) (action : Action) : gamma / (arms.card : Real) <= exploredHistoryDistribution arms eta gamma score history action","missing":[],"search":"exploredhistorydistribution_floor banditrlproof.exp3.exploredhistorydistribution_floor theorem exploredhistorydistribution_floor {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (score : history -> action -> real) (hgamma_le_one : gamma <= 1) (history : history) (action : action) : gamma / (arms.card : real) <= exploredhistorydistribution arms eta gamma score history action theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.explorationFloor_pos","label":"explorationFloor_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.explorationFloor_pos","description":"theorem explorationFloor_pos {Action : Type*} (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) : 0 < gamma / (arms.card : Real)","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-fb826314cc92","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4764,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:128"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem explorationFloor_pos {Action : Type*} (arms : Finset Action) (harms : arms.Nonempty) (gamma : Real) (hgamma_pos : 0 < gamma) : 0 < gamma / (arms.card : Real)","missing":[],"search":"explorationfloor_pos banditrlproof.exp3.explorationfloor_pos theorem explorationfloor_pos {action : type*} (arms : finset action) (harms : arms.nonempty) (gamma : real) (hgamma_pos : 0 < gamma) : 0 < gamma / (arms.card : real) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionDistribution_exploredHistoryDistribution","label":"finiteActionDistribution_exploredHistoryDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionDistribution_exploredHistoryDistribution","description":"theorem finiteActionDistribution_exploredHistoryDistribution {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (history : History) : FiniteActionDistribution arms (exploredHistoryDistribution arms eta gamma score history) where","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-9669912a76dd","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4765,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:135"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionDistribution_exploredHistoryDistribution {Action History : Type*} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : History -> Action -> Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (history : History) : FiniteActionDistribution arms (exploredHistoryDistribution arms eta gamma score history) where","missing":[],"search":"finiteactiondistribution_exploredhistorydistribution banditrlproof.exp3.finiteactiondistribution_exploredhistorydistribution theorem finiteactiondistribution_exploredhistorydistribution {action history : type*} (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (score : history -> action -> real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (history : history) : finiteactiondistribution arms (exploredhistorydistribution arms eta gamma score history) where theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_historyWeight","label":"measurable_historyWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_historyWeight","description":"theorem measurable_historyWeight {Action History : Type*} [MeasurableSpace History] (eta : Real) (score : History -> Action -> Real) (action : Action) (hscore : Measurable (fun history => score history action)) : Measurable (fun history => historyWeight eta score history action)","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-9cd0b51bae70","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4766,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:148"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyWeight {Action History : Type*} [MeasurableSpace History] (eta : Real) (score : History -> Action -> Real) (action : Action) (hscore : Measurable (fun history => score history action)) : Measurable (fun history => historyWeight eta score history action)","missing":[],"search":"measurable_historyweight banditrlproof.exp3.measurable_historyweight theorem measurable_historyweight {action history : type*} [measurablespace history] (eta : real) (score : history -> action -> real) (action : action) (hscore : measurable (fun history => score history action)) : measurable (fun history => historyweight eta score history action) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_historyTotalWeight","label":"measurable_historyTotalWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_historyTotalWeight","description":"theorem measurable_historyTotalWeight {Action History : Type*} [MeasurableSpace History] (arms : Finset Action) (eta : Real) (score : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) : Measurable (historyTotalWeight arms eta score)","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-1d8f355cb6a7","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4767,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:155"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyTotalWeight {Action History : Type*} [MeasurableSpace History] (arms : Finset Action) (eta : Real) (score : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) : Measurable (historyTotalWeight arms eta score)","missing":[],"search":"measurable_historytotalweight banditrlproof.exp3.measurable_historytotalweight theorem measurable_historytotalweight {action history : type*} [measurablespace history] (arms : finset action) (eta : real) (score : history -> action -> real) (hscore : forall action, action ∈ arms -> measurable (fun history => score history action)) : measurable (historytotalweight arms eta score) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_normalizedHistoryDistribution","label":"measurable_normalizedHistoryDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_normalizedHistoryDistribution","description":"theorem measurable_normalizedHistoryDistribution {Action History : Type*} [MeasurableSpace History] (arms : Finset Action) (eta : Real) (score : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) (action : Action) (haction : action ∈ arms) : Measurable (fun history => normalizedHistoryDistribution arms eta score history action)","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-76e28b686c76","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4768,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:166"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_normalizedHistoryDistribution {Action History : Type*} [MeasurableSpace History] (arms : Finset Action) (eta : Real) (score : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) (action : Action) (haction : action ∈ arms) : Measurable (fun history => normalizedHistoryDistribution arms eta score history action)","missing":[],"search":"measurable_normalizedhistorydistribution banditrlproof.exp3.measurable_normalizedhistorydistribution theorem measurable_normalizedhistorydistribution {action history : type*} [measurablespace history] (arms : finset action) (eta : real) (score : history -> action -> real) (hscore : forall action, action ∈ arms -> measurable (fun history => score history action)) (action : action) (haction : action ∈ arms) : measurable (fun history => normalizedhistorydistribution arms eta score history action) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_exploredHistoryDistribution","label":"measurable_exploredHistoryDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_exploredHistoryDistribution","description":"theorem measurable_exploredHistoryDistribution {Action History : Type*} [MeasurableSpace History] (arms : Finset Action) (eta gamma : Real) (score : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) (action : Action) (haction : action ∈ arms) : Measurable (fun history => exploredHistoryDistribution arms eta gamma score history action)","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-2a96c18e5404","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4769,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:178"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_exploredHistoryDistribution {Action History : Type*} [MeasurableSpace History] (arms : Finset Action) (eta gamma : Real) (score : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) (action : Action) (haction : action ∈ arms) : Measurable (fun history => exploredHistoryDistribution arms eta gamma score history action)","missing":[],"search":"measurable_exploredhistorydistribution banditrlproof.exp3.measurable_exploredhistorydistribution theorem measurable_exploredhistorydistribution {action history : type*} [measurablespace history] (arms : finset action) (eta gamma : real) (score : history -> action -> real) (hscore : forall action, action ∈ arms -> measurable (fun history => score history action)) (action : action) (haction : action ∈ arms) : measurable (fun history => exploredhistorydistribution arms eta gamma score history action) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.MeasurableFiniteHistoryScore","label":"MeasurableFiniteHistoryScore","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Exp3.MeasurableFiniteHistoryScore","description":"Coordinate measurability for a score indexed by inclusive pair histories.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-34fa28c70a11","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4770,"meta":[["Kind","structure"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:192"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"structure MeasurableFiniteHistoryScore {Action : Type u} {Loss : Type v} [MeasurableSpace Action] [MeasurableSpace Loss] (arms : Finset Action) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) : Prop where","missing":[],"search":"measurablefinitehistoryscore banditrlproof.exp3.measurablefinitehistoryscore coordinate measurability for a score indexed by inclusive pair histories. structure compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.normalizedHistoryDistributionSource","label":"normalizedHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.normalizedHistoryDistributionSource","description":"The un-explored normalized exponential-weights policy supplies measurable probability vectors on every finite-history level.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-f8840a3d1555","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4771,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:203"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def normalizedHistoryDistributionSource {Action : Type u} {Loss : Type v} [MeasurableSpace Action] [MeasurableSpace Loss] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (n : Nat) : MeasurableFiniteActionDistribution arms (normalizedHistoryDistribution arms eta (score n)) where","missing":[],"search":"normalizedhistorydistributionsource banditrlproof.exp3.normalizedhistorydistributionsource the un-explored normalized exponential-weights policy supplies measurable probability vectors on every finite-history level. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredHistoryDistributionSource","label":"exploredHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredHistoryDistributionSource","description":"The exploration-mixed history policy supplies measurable probability vectors.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-792bcdf11929","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4772,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:226"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def exploredHistoryDistributionSource {Action : Type u} {Loss : Type v} [MeasurableSpace Action] [MeasurableSpace Loss] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : Nat) : MeasurableFiniteActionDistribution arms (exploredHistoryDistribution arms eta gamma (score n)) where","missing":[],"search":"exploredhistorydistributionsource banditrlproof.exp3.exploredhistorydistributionsource the exploration-mixed history policy supplies measurable probability vectors. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.initialExploredDistribution","label":"initialExploredDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.initialExploredDistribution","description":"Initial EXP3 probabilities, obtained from zero cumulative scores.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-4a38e6698864","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4773,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:247"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def initialExploredDistribution {Action : Type u} (arms : Finset Action) (eta gamma : Real) : Action -> Real","missing":[],"search":"initialexploreddistribution banditrlproof.exp3.initialexploreddistribution initial exp3 probabilities, obtained from zero cumulative scores. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteActionDistribution_initialExploredDistribution","label":"finiteActionDistribution_initialExploredDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteActionDistribution_initialExploredDistribution","description":"theorem finiteActionDistribution_initialExploredDistribution {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) : FiniteActionDistribution arms (initialExploredDistribution arms eta gamma)","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-5670ee76b7f5","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4774,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:252"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem finiteActionDistribution_initialExploredDistribution {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) : FiniteActionDistribution arms (initialExploredDistribution arms eta gamma)","missing":[],"search":"finiteactiondistribution_initialexploreddistribution banditrlproof.exp3.finiteactiondistribution_initialexploreddistribution theorem finiteactiondistribution_initialexploreddistribution {action : type u} (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) : finiteactiondistribution arms (initialexploreddistribution arms eta gamma) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredHistoryAlgorithm","label":"exploredHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredHistoryAlgorithm","description":"Stochastic finite-history algorithm generated by the EXP3 score policy.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-f2c5af60262e","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4775,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:263"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exploredHistoryAlgorithm {Action : Type u} {Loss : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Loss] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) : Thompson.HistoryAlgorithm Action Loss where","missing":[],"search":"exploredhistoryalgorithm banditrlproof.exp3.exploredhistoryalgorithm stochastic finite-history algorithm generated by the exp3 score policy. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredHistoryAlgorithm_policy","label":"exploredHistoryAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredHistoryAlgorithm_policy","description":"theorem exploredHistoryAlgorithm_policy {Action : Type u} {Loss : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Loss] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : Na…","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-2f15b688e154","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4776,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:290"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exploredHistoryAlgorithm_policy {Action : Type u} {Loss : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Loss] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : Nat) : (exploredHistoryAlgorithm arms harms eta gamma score hscore hgamma_nonneg hgamma_le_one).policy n = finiteActionKernel arms (exploredHistoryDistribution arms eta gamma (score n)) (exploredHistoryDistributionSource arms harms eta gamma score hscore hgamma_nonneg hgamma_le_one n)","missing":[],"search":"exploredhistoryalgorithm_policy banditrlproof.exp3.exploredhistoryalgorithm_policy theorem exploredhistoryalgorithm_policy {action : type u} {loss : type v} [measurablespace action] [measurablesingletonclass action] [measurablespace loss] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (score : (n : nat) -> history.finitepairhistory action loss n -> action -> real) (hscore : measurablefinitehistoryscore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : nat) : (exploredhistoryalgorithm arms harms eta gamma score hscore hgamma_nonneg hgamma_le_one).policy n = finiteactionkernel arms (exploredhistorydistribution arms eta gamma (score n)) (exploredhistorydistributionsource arms harms eta gamma score hscore hgamma_nonneg hgamma_le_one n) theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredTrajectoryKernel","label":"exploredTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredTrajectoryKernel","description":"Complete environment-indexed recursive EXP3 action/loss trajectory kernel.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-75173d5e45a8","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4777,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:310"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def exploredTrajectoryKernel {Env : Type w} {Action : Type u} {Loss : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Loss] [Nonempty Action] [Nonempty Loss] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (environment : Thompson.MeasurableHistoryEnvironment Env Action Loss) : Kernel Env ((n : Nat) -> Action × Loss)","missing":[],"search":"exploredtrajectorykernel banditrlproof.exp3.exploredtrajectorykernel complete environment-indexed recursive exp3 action/loss trajectory kernel. definition compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.exploredTrajectoryMeasure_condDistrib_action","label":"exploredTrajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.exploredTrajectoryMeasure_condDistrib_action","description":"Every successor action of the recursive trajectory has the explicit exploration-mixed EXP3 policy as its conditional law given the finite history.","url":"../modules/banditrlproof-exp3recursivetrajectory/index.html#decl-2afc8ae67b9a","parent":"module:BanditRLProof.Exp3RecursiveTrajectory","order":4778,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3RecursiveTrajectory"],["Source","BanditRLProof/Exp3RecursiveTrajectory.lean:349"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem exploredTrajectoryMeasure_condDistrib_action {Env : Type w} {Action : Type u} {Loss : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableSpace Loss] [StandardBorelSpace Loss] [Nonempty Loss] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (score : (n : Nat) -> History.FinitePairHistory Action Loss n -> Action -> Real) (hscore : MeasurableFiniteHistoryScore arms score) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (environment : Thompson.MeasurableHistoryEnvironment Env Action Loss) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Loss) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ exploredTrajectoryKernel arms harms eta gamma score…","missing":[],"search":"exploredtrajectorymeasure_conddistrib_action banditrlproof.exp3.exploredtrajectorymeasure_conddistrib_action every successor action of the recursive trajectory has the explicit exploration-mixed exp3 policy as its conditional law given the finite history. theorem compiled","shard":"modules/7b2b37874141e198.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryObservedLoss","label":"sampledTrajectoryObservedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryObservedLoss","description":"The complete importance-weighted loss vector observed at an actual time.","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-e8fa00fceb05","parent":"module:BanditRLProof.Exp3SampledHedge","order":4779,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:23"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledTrajectoryObservedLoss {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (sample : Env × ((k : Nat) -> Action × Real)) : Nat -> Action -> Real","missing":[],"search":"sampledtrajectoryobservedloss banditrlproof.exp3.sampledtrajectoryobservedloss the complete importance-weighted loss vector observed at an actual time. definition compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.previousPairHistory_frestrictLe","label":"previousPairHistory_frestrictLe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.previousPairHistory_frestrictLe","description":"theorem previousPairHistory_frestrictLe {Action : Type v} (n : Nat) (trajectory : (k : Nat) -> Action × Real) : previousPairHistory (Preorder.frestrictLe (n + 1) trajectory) = Preorder.frestrictLe n trajectory","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-40ad1622700b","parent":"module:BanditRLProof.Exp3SampledHedge","order":4780,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:32"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem previousPairHistory_frestrictLe {Action : Type v} (n : Nat) (trajectory : (k : Nat) -> Action × Real) : previousPairHistory (Preorder.frestrictLe (n + 1) trajectory) = Preorder.frestrictLe n trajectory","missing":[],"search":"previouspairhistory_frestrictle banditrlproof.exp3.previouspairhistory_frestrictle theorem previouspairhistory_frestrictle {action : type v} (n : nat) (trajectory : (k : nat) -> action × real) : previouspairhistory (preorder.frestrictle (n + 1) trajectory) = preorder.frestrictle n trajectory theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryScore_frestrictLe_eq_cumulativeLoss","label":"sampledHistoryScore_frestrictLe_eq_cumulativeLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryScore_frestrictLe_eq_cumulativeLoss","description":"The inclusive sampled score through `n` is Hedge cumulative loss at `n+1`.","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-a852f3b81845","parent":"module:BanditRLProof.Exp3SampledHedge","order":4781,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:41"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledHistoryScore_frestrictLe_eq_cumulativeLoss {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (sample : Env × ((k : Nat) -> Action × Real)) (n : Nat) (action : Action) : sampledHistoryScore arms eta gamma n (Preorder.frestrictLe n sample.2) action = cumulativeLoss (sampledTrajectoryObservedLoss arms eta gamma sample) (n + 1) action","missing":[],"search":"sampledhistoryscore_frestrictle_eq_cumulativeloss banditrlproof.exp3.sampledhistoryscore_frestrictle_eq_cumulativeloss the inclusive sampled score through `n` is hedge cumulative loss at `n+1`. theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.distribution_sampledTrajectoryObservedLoss_succ","label":"distribution_sampledTrajectoryObservedLoss_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.distribution_sampledTrajectoryObservedLoss_succ","description":"At a successor time, the Hedge distribution is the normalized sampled score.","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-3fda22d4960a","parent":"module:BanditRLProof.Exp3SampledHedge","order":4782,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:66"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem distribution_sampledTrajectoryObservedLoss_succ {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (sample : Env × ((k : Nat) -> Action × Real)) (n : Nat) (action : Action) : distribution arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) (n + 1) action = normalizedHistoryDistribution arms eta (sampledHistoryScore arms eta gamma n) (Preorder.frestrictLe n sample.2) action","missing":[],"search":"distribution_sampledtrajectoryobservedloss_succ banditrlproof.exp3.distribution_sampledtrajectoryobservedloss_succ at a successor time, the hedge distribution is the normalized sampled score. theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilityAt_eq_mix_distribution","label":"sampledTrajectoryProbabilityAt_eq_mix_distribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryProbabilityAt_eq_mix_distribution","description":"The concrete sampling law is uniform exploration mixed with the Hedge law.","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-2e053c03e175","parent":"module:BanditRLProof.Exp3SampledHedge","order":4783,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:86"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryProbabilityAt_eq_mix_distribution {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (action : Action) : sampledTrajectoryProbabilityAt arms eta gamma t sample action = (1 - gamma) * distribution arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t action + gamma / (arms.card : Real)","missing":[],"search":"sampledtrajectoryprobabilityat_eq_mix_distribution banditrlproof.exp3.sampledtrajectoryprobabilityat_eq_mix_distribution the concrete sampling law is uniform exploration mixed with the hedge law. theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilityAt_nonneg","label":"sampledTrajectoryProbabilityAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryProbabilityAt_nonneg","description":"theorem sampledTrajectoryProbabilityAt_nonneg {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (action : Action) : 0 <= sampledTrajectoryProbabilityAt arms eta gamma t sample action","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-d4e9144fe86f","parent":"module:BanditRLProof.Exp3SampledHedge","order":4784,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:103"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryProbabilityAt_nonneg {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (action : Action) : 0 <= sampledTrajectoryProbabilityAt arms eta gamma t sample action","missing":[],"search":"sampledtrajectoryprobabilityat_nonneg banditrlproof.exp3.sampledtrajectoryprobabilityat_nonneg theorem sampledtrajectoryprobabilityat_nonneg {env : type u} {action : type v} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (t : nat) (sample : env × ((k : nat) -> action × real)) (action : action) : 0 <= sampledtrajectoryprobabilityat arms eta gamma t sample action theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectoryObservedLoss_nonneg","label":"sampledTrajectoryObservedLoss_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectoryObservedLoss_nonneg","description":"theorem sampledTrajectoryObservedLoss_nonneg {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) (action : Action) (hreward : 0 <= (sample.2 t).2) : 0 <= sampledTrajectoryObservedLoss arms eta gamma sample t action","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-af189153956e","parent":"module:BanditRLProof.Exp3SampledHedge","order":4785,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:121"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectoryObservedLoss_nonneg {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) (action : Action) (hreward : 0 <= (sample.2 t).2) : 0 <= sampledTrajectoryObservedLoss arms eta gamma sample t action","missing":[],"search":"sampledtrajectoryobservedloss_nonneg banditrlproof.exp3.sampledtrajectoryobservedloss_nonneg theorem sampledtrajectoryobservedloss_nonneg {env : type u} {action : type v} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (sample : env × ((k : nat) -> action × real)) (t : nat) (action : action) (hreward : 0 <= (sample.2 t).2) : 0 <= sampledtrajectoryobservedloss arms eta gamma sample t action theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledTrajectory_hedge_regret_le","label":"sampledTrajectory_hedge_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledTrajectory_hedge_regret_le","description":"Concrete finite-horizon sampled-trajectory specialization of Hedge.","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-5a6dc5372101","parent":"module:BanditRLProof.Exp3SampledHedge","order":4786,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:135"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledTrajectory_hedge_regret_le {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (sample : Env × ((k : Nat) -> Action × Real)) (horizon : Nat) (hreward_nonneg : forall t, t < horizon -> 0 <= (sample.2 t).2) (comparator : Action) (hcomparator : comparator ∈ arms) : (Finset.range horizon).sum (fun t => mixedLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t) - cumulativeLoss (sampledTrajectoryObservedLoss arms eta gamma sample) horizon comparator <= Real.log arms.card / eta + eta * (Finset.range horizon).sum (fun t => mixedSquaredLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t)","missing":[],"search":"sampledtrajectory_hedge_regret_le banditrlproof.exp3.sampledtrajectory_hedge_regret_le concrete finite-horizon sampled-trajectory specialization of hedge. theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryScore_hedge_regret_le","label":"sampledHistoryScore_hedge_regret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryScore_hedge_regret_le","description":"The same pathwise Hedge bound with its comparator term exposed as the concrete inclusive `sampledHistoryScore`.","url":"../modules/banditrlproof-exp3sampledhedge/index.html#decl-a3c66f43fdae","parent":"module:BanditRLProof.Exp3SampledHedge","order":4787,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHedge"],["Source","BanditRLProof/Exp3SampledHedge.lean:162"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledHistoryScore_hedge_regret_le {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (heta : 0 < eta) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (sample : Env × ((k : Nat) -> Action × Real)) (n : Nat) (hreward_nonneg : forall t, t < n + 1 -> 0 <= (sample.2 t).2) (comparator : Action) (hcomparator : comparator ∈ arms) : (Finset.range (n + 1)).sum (fun t => mixedLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t) - sampledHistoryScore arms eta gamma n (Preorder.frestrictLe n sample.2) comparator <= Real.log arms.card / eta + eta * (Finset.range (n + 1)).sum (fun t => mixedSquaredLoss arms eta (sampledTrajectoryObservedLoss arms eta gamma sample) t)","missing":[],"search":"sampledhistoryscore_hedge_regret_le banditrlproof.exp3.sampledhistoryscore_hedge_regret_le the same pathwise hedge bound with its comparator term exposed as the concrete inclusive `sampledhistoryscore`. theorem compiled","shard":"modules/3a1d59db36248686.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.previousPairHistory","label":"previousPairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.previousPairHistory","description":"Remove the newest coordinate from an inclusive successor pair history.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-a0b6d211e9d6","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4788,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:25"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"def previousPairHistory {Action : Type u} {n : Nat} (history : History.FinitePairHistory Action Real (n + 1)) : History.FinitePairHistory Action Real n","missing":[],"search":"previouspairhistory banditrlproof.exp3.previouspairhistory remove the newest coordinate from an inclusive successor pair history. definition compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_previousPairHistory","label":"measurable_previousPairHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_previousPairHistory","description":"theorem measurable_previousPairHistory {Action : Type u} [MeasurableSpace Action] {n : Nat} : Measurable (previousPairHistory (Action := Action) (n := n))","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-891ab6edb1e2","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4789,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:32"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_previousPairHistory {Action : Type u} [MeasurableSpace Action] {n : Nat} : Measurable (previousPairHistory (Action := Action) (n := n))","missing":[],"search":"measurable_previouspairhistory banditrlproof.exp3.measurable_previouspairhistory theorem measurable_previouspairhistory {action : type u} [measurablespace action] {n : nat} : measurable (previouspairhistory (action := action) (n := n)) theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_observedImportanceWeightedLoss","label":"measurable_observedImportanceWeightedLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_observedImportanceWeightedLoss","description":"Measurability of one importance-weighted coordinate when only the sampled action and its scalar observed loss are available.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-59cd597e2765","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4790,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:47"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_observedImportanceWeightedLoss {History : Type*} {Action : Type u} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (prob : History -> Action -> Real) (chosen : History -> Action) (observedLoss : History -> Real) (action : Action) (hprob : Measurable (fun history => prob history action)) (hchosen : Measurable chosen) (hloss : Measurable observedLoss) : Measurable (fun history => importanceWeightedLoss (prob history) (fun _ => observedLoss history) (chosen history) action)","missing":[],"search":"measurable_observedimportanceweightedloss banditrlproof.exp3.measurable_observedimportanceweightedloss measurability of one importance-weighted coordinate when only the sampled action and its scalar observed loss are available. theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryScore","label":"sampledHistoryScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryScore","description":"The cumulative sampled importance-weighted loss through an inclusive history. At time zero the estimator uses the initial action law. At time `n + 1`, it uses the exploration-mixed law generated from the score on the prefix through `n`, exactly matching `HistoryAlgorithm.policy n`.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-9e6c66f41987","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4791,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:71"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHistoryScore {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) : (n : Nat) -> History.FinitePairHistory Action Real n -> Action -> Real | 0, history, action => importanceWeightedLoss (initialExploredDistribution arms eta gamma) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action | n + 1, history, action => let previous := previousPairHistory history sampledHistoryScore arms eta gamma n previous action + importanceWeightedLoss (exploredHistoryDistribution arms eta gamma (sampledHistoryScore arms eta gamma n) previous) (fun _ => (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 action @[simp] theorem sampledHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (history : History.FinitePairHisto…","missing":[],"search":"sampledhistoryscore banditrlproof.exp3.sampledhistoryscore the cumulative sampled importance-weighted loss through an inclusive history. at time zero the estimator uses the initial action law. at time `n + 1`, it uses the exploration-mixed law generated from the score on the prefix through `n`, exactly matching `historyalgorithm.policy n`. definition compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryScore_zero","label":"sampledHistoryScore_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryScore_zero","description":"theorem sampledHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (history : History.FinitePairHistory Action Real 0) (action : Action) : sampledHistoryScore arms eta gamma 0 history action = importanceWeightedLoss (initialExploredDistribution arms eta gamma) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-d04e12aa077d","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4792,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:91"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (history : History.FinitePairHistory Action Real 0) (action : Action) : sampledHistoryScore arms eta gamma 0 history action = importanceWeightedLoss (initialExploredDistribution arms eta gamma) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action","missing":[],"search":"sampledhistoryscore_zero banditrlproof.exp3.sampledhistoryscore_zero theorem sampledhistoryscore_zero {action : type u} [decidableeq action] (arms : finset action) (eta gamma : real) (history : history.finitepairhistory action real 0) (action : action) : sampledhistoryscore arms eta gamma 0 history action = importanceweightedloss (initialexploreddistribution arms eta gamma) (fun _ => (history ⟨0, finset.mem_iic.mpr le_rfl⟩).2) (history ⟨0, finset.mem_iic.mpr le_rfl⟩).1 action theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryScore_succ","label":"sampledHistoryScore_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryScore_succ","description":"theorem sampledHistoryScore_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (action : Action) : sampledHistoryScore arms eta gamma (n + 1) history action = sampledHistoryScore arms eta gamma n (previousPairHistory history) action + importanceWeightedLoss (exploredHistoryDistribution arms eta gamma (sampledHistor…","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-f66e3501bc48","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4793,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:104"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledHistoryScore_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (action : Action) : sampledHistoryScore arms eta gamma (n + 1) history action = sampledHistoryScore arms eta gamma n (previousPairHistory history) action + importanceWeightedLoss (exploredHistoryDistribution arms eta gamma (sampledHistoryScore arms eta gamma n) (previousPairHistory history)) (fun _ => (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 action","missing":[],"search":"sampledhistoryscore_succ banditrlproof.exp3.sampledhistoryscore_succ theorem sampledhistoryscore_succ {action : type u} [decidableeq action] (arms : finset action) (eta gamma : real) (n : nat) (history : history.finitepairhistory action real (n + 1)) (action : action) : sampledhistoryscore arms eta gamma (n + 1) history action = sampledhistoryscore arms eta gamma n (previouspairhistory history) action + importanceweightedloss (exploredhistorydistribution arms eta gamma (sampledhistoryscore arms eta gamma n) (previouspairhistory history)) (fun _ => (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).2) (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).1 action theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_sampledHistoryScore","label":"measurable_sampledHistoryScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_sampledHistoryScore","description":"theorem measurable_sampledHistoryScore {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) : forall n action, action ∈ arms -> Measurable (fun history : History.FinitePairHistory Action Real n => sampledHistoryScore arms eta gamma n history action)","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-256a1d66c40e","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4794,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:121"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHistoryScore {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) : forall n action, action ∈ arms -> Measurable (fun history : History.FinitePairHistory Action Real n => sampledHistoryScore arms eta gamma n history action)","missing":[],"search":"measurable_sampledhistoryscore banditrlproof.exp3.measurable_sampledhistoryscore theorem measurable_sampledhistoryscore {action : type u} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (eta gamma : real) : forall n action, action ∈ arms -> measurable (fun history : history.finitepairhistory action real n => sampledhistoryscore arms eta gamma n history action) theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurableFiniteHistoryScore_sampledHistoryScore","label":"measurableFiniteHistoryScore_sampledHistoryScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurableFiniteHistoryScore_sampledHistoryScore","description":"The concrete sampled score satisfies the generic measurable-score API.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-464294dcb5b8","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4795,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:199"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurableFiniteHistoryScore_sampledHistoryScore {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) : MeasurableFiniteHistoryScore arms (sampledHistoryScore arms eta gamma) where","missing":[],"search":"measurablefinitehistoryscore_sampledhistoryscore banditrlproof.exp3.measurablefinitehistoryscore_sampledhistoryscore the concrete sampled score satisfies the generic measurable-score api. theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryDistribution","label":"sampledHistoryDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryDistribution","description":"Concrete exploration-mixed probabilities generated by sampled losses.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-3bb8c49f8dc0","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4796,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:210"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHistoryDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta gamma : Real) (n : Nat) : History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledhistorydistribution banditrlproof.exp3.sampledhistorydistribution concrete exploration-mixed probabilities generated by sampled losses. definition compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledHistoryDistribution_floor","label":"sampledHistoryDistribution_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledHistoryDistribution_floor","description":"theorem sampledHistoryDistribution_floor {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_le_one : gamma <= 1) (n : Nat) (history : History.FinitePairHistory Action Real n) (action : Action) : gamma / (arms.card : Real) <= sampledHistoryDistribution arms eta gamma n history action","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-0cc78deecf5c","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4797,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:217"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledHistoryDistribution_floor {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_le_one : gamma <= 1) (n : Nat) (history : History.FinitePairHistory Action Real n) (action : Action) : gamma / (arms.card : Real) <= sampledHistoryDistribution arms eta gamma n history action","missing":[],"search":"sampledhistorydistribution_floor banditrlproof.exp3.sampledhistorydistribution_floor theorem sampledhistorydistribution_floor {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_le_one : gamma <= 1) (n : nat) (history : history.finitepairhistory action real n) (action : action) : gamma / (arms.card : real) <= sampledhistorydistribution arms eta gamma n history action theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedHistoryAlgorithm","label":"sampledImportanceWeightedHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledImportanceWeightedHistoryAlgorithm","description":"The concrete stochastic history algorithm for sampled-loss EXP3.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-06f89737be42","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4798,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:229"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledImportanceWeightedHistoryAlgorithm {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) : Thompson.HistoryAlgorithm Action Real","missing":[],"search":"sampledimportanceweightedhistoryalgorithm banditrlproof.exp3.sampledimportanceweightedhistoryalgorithm the concrete stochastic history algorithm for sampled-loss exp3. definition compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedHistoryAlgorithm_policy","label":"sampledImportanceWeightedHistoryAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledImportanceWeightedHistoryAlgorithm_policy","description":"theorem sampledImportanceWeightedHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : Nat) : (sampledImportanceWeightedHistoryAlgorithm arms harms eta gamma hgamma_nonneg hgamma_le_one).policy n = finiteActionKernel arms…","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-72908a8d01bf","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4799,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:243"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledImportanceWeightedHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : Nat) : (sampledImportanceWeightedHistoryAlgorithm arms harms eta gamma hgamma_nonneg hgamma_le_one).policy n = finiteActionKernel arms (sampledHistoryDistribution arms eta gamma n) (exploredHistoryDistributionSource arms harms eta gamma (sampledHistoryScore arms eta gamma) (measurableFiniteHistoryScore_sampledHistoryScore arms eta gamma) hgamma_nonneg hgamma_le_one n)","missing":[],"search":"sampledimportanceweightedhistoryalgorithm_policy banditrlproof.exp3.sampledimportanceweightedhistoryalgorithm_policy theorem sampledimportanceweightedhistoryalgorithm_policy {action : type u} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta gamma : real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (n : nat) : (sampledimportanceweightedhistoryalgorithm arms harms eta gamma hgamma_nonneg hgamma_le_one).policy n = finiteactionkernel arms (sampledhistorydistribution arms eta gamma n) (exploredhistorydistributionsource arms harms eta gamma (sampledhistoryscore arms eta gamma) (measurablefinitehistoryscore_sampledhistoryscore arms eta gamma) hgamma_nonneg hgamma_le_one n) theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryKernel","label":"sampledImportanceWeightedTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryKernel","description":"Complete recursive sampled-loss EXP3 trajectory kernel.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-579997277126","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4800,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:262"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledImportanceWeightedTrajectoryKernel {Env : Type w} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] [Nonempty Action] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) : Kernel Env ((n : Nat) -> Action × Real)","missing":[],"search":"sampledimportanceweightedtrajectorykernel banditrlproof.exp3.sampledimportanceweightedtrajectorykernel complete recursive sampled-loss exp3 trajectory kernel. definition compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryMeasure_condDistrib_action","label":"sampledImportanceWeightedTrajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryMeasure_condDistrib_action","description":"Every successor action of the concrete sampled-loss EXP3 trajectory has the exploration-mixed law generated from its recursively accumulated importance-weighted score.","url":"../modules/banditrlproof-exp3sampledhistoryscore/index.html#decl-292f333c9113","parent":"module:BanditRLProof.Exp3SampledHistoryScore","order":4801,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3SampledHistoryScore"],["Source","BanditRLProof/Exp3SampledHistoryScore.lean:297"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledImportanceWeightedTrajectoryMeasure_condDistrib_action {Env : Type w} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one environment) =ᵐ[ (prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one environment).map (…","missing":[],"search":"sampledimportanceweightedtrajectorymeasure_conddistrib_action banditrlproof.exp3.sampledimportanceweightedtrajectorymeasure_conddistrib_action every successor action of the concrete sampled-loss exp3 trajectory has the exploration-mixed law generated from its recursively accumulated importance-weighted score. theorem compiled","shard":"modules/5c770646523f2639.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.BoundedMeasurableLossWithProbabilityFloor","label":"BoundedMeasurableLossWithProbabilityFloor","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Exp3.BoundedMeasurableLossWithProbabilityFloor","description":"Measurable bounded losses and a uniform exploration floor on the finite support.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-1fbfd0341eff","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4802,"meta":[["Kind","structure"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"structure BoundedMeasurableLossWithProbabilityFloor {History : Type u} {Action : Type v} [MeasurableSpace History] (arms : Finset Action) (prob loss : History -> Action -> Real) (epsilon : Real) : Prop where","missing":[],"search":"boundedmeasurablelosswithprobabilityfloor banditrlproof.exp3.boundedmeasurablelosswithprobabilityfloor measurable bounded losses and a uniform exploration floor on the finite support. structure compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.BoundedMeasurableLossWithProbabilityFloor.prob_pos","label":"prob_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.BoundedMeasurableLossWithProbabilityFloor.prob_pos","description":"theorem BoundedMeasurableLossWithProbabilityFloor.prob_pos {History : Type u} {Action : Type v} [MeasurableSpace History] {arms : Finset Action} {prob loss : History -> Action -> Real} {epsilon : Real} (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (action : Action) (haction : action ∈ arms) : 0 < prob history action","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-3f5949202d3e","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4803,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:33"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem BoundedMeasurableLossWithProbabilityFloor.prob_pos {History : Type u} {Action : Type v} [MeasurableSpace History] {arms : Finset Action} {prob loss : History -> Action -> Real} {epsilon : Real} (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (action : Action) (haction : action ∈ arms) : 0 < prob history action","missing":[],"search":"prob_pos banditrlproof.exp3.boundedmeasurablelosswithprobabilityfloor.prob_pos theorem boundedmeasurablelosswithprobabilityfloor.prob_pos {history : type u} {action : type v} [measurablespace history] {arms : finset action} {prob loss : history -> action -> real} {epsilon : real} (regularity : boundedmeasurablelosswithprobabilityfloor arms prob loss epsilon) (history : history) (action : action) (haction : action ∈ arms) : 0 < prob history action theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_importanceWeightedLoss_score","label":"measurable_importanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_importanceWeightedLoss_score","description":"A fixed-arm importance-weighted score is measurable on history/action pairs.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-3b99ad9a0cd6","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4804,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:46"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_importanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (action : Action) (haction : action ∈ arms) : Measurable (fun sample : History × Action => importanceWeightedLoss (prob sample.1) (loss sample.1) sample.2 action)","missing":[],"search":"measurable_importanceweightedloss_score banditrlproof.exp3.measurable_importanceweightedloss_score a fixed-arm importance-weighted score is measurable on history/action pairs. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_mixedImportanceWeightedLoss_score","label":"measurable_mixedImportanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_mixedImportanceWeightedLoss_score","description":"The probability-mixed first-moment score is measurable.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-1a3506f4cbb9","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4805,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:68"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_mixedImportanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Measurable (fun sample : History × Action => mixedImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2)","missing":[],"search":"measurable_mixedimportanceweightedloss_score banditrlproof.exp3.measurable_mixedimportanceweightedloss_score the probability-mixed first-moment score is measurable. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_weightedImportanceWeightedLoss_score","label":"measurable_weightedImportanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_weightedImportanceWeightedLoss_score","description":"A score mixed by a second measurable finite distribution is measurable.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-dc61790d1273","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4806,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:87"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_weightedImportanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob weight loss : History -> Action -> Real) (probSource : MeasurableFiniteActionDistribution arms prob) (weightSource : MeasurableFiniteActionDistribution arms weight) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Measurable (fun sample : History × Action => weightedImportanceWeightedLoss arms (prob sample.1) (weight sample.1) (loss sample.1) sample.2)","missing":[],"search":"measurable_weightedimportanceweightedloss_score banditrlproof.exp3.measurable_weightedimportanceweightedloss_score a score mixed by a second measurable finite distribution is measurable. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.measurable_mixedSquaredImportanceWeightedLoss_score","label":"measurable_mixedSquaredImportanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.measurable_mixedSquaredImportanceWeightedLoss_score","description":"The probability-mixed second-moment score is measurable.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-2a7bdf3718fb","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4807,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:107"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem measurable_mixedSquaredImportanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Measurable (fun sample : History × Action => mixedSquaredImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2)","missing":[],"search":"measurable_mixedsquaredimportanceweightedloss_score banditrlproof.exp3.measurable_mixedsquaredimportanceweightedloss_score the probability-mixed second-moment score is measurable. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.norm_importanceWeightedLoss_score_le_inv_floor","label":"norm_importanceWeightedLoss_score_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.norm_importanceWeightedLoss_score_le_inv_floor","description":"A fixed-arm importance-weighted score is bounded by the reciprocal floor.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-dc662a3bd0c3","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4808,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:126"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem norm_importanceWeightedLoss_score_le_inv_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (chosen action : Action) (haction : action ∈ arms) : ‖importanceWeightedLoss (prob history) (loss history) chosen action‖ <= 1 / epsilon","missing":[],"search":"norm_importanceweightedloss_score_le_inv_floor banditrlproof.exp3.norm_importanceweightedloss_score_le_inv_floor a fixed-arm importance-weighted score is bounded by the reciprocal floor. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.norm_mixedImportanceWeightedLoss_score_le_inv_floor","label":"norm_mixedImportanceWeightedLoss_score_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.norm_mixedImportanceWeightedLoss_score_le_inv_floor","description":"The mixed first-moment score is bounded by the reciprocal floor.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-4d3e0c9a4de9","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4809,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:148"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem norm_mixedImportanceWeightedLoss_score_le_inv_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (chosen : Action) : ‖mixedImportanceWeightedLoss arms (prob history) (loss history) chosen‖ <= 1 / epsilon","missing":[],"search":"norm_mixedimportanceweightedloss_score_le_inv_floor banditrlproof.exp3.norm_mixedimportanceweightedloss_score_le_inv_floor the mixed first-moment score is bounded by the reciprocal floor. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.norm_weightedImportanceWeightedLoss_score_le_inv_floor","label":"norm_weightedImportanceWeightedLoss_score_le_inv_floor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.norm_weightedImportanceWeightedLoss_score_le_inv_floor","description":"Mixing by any finite probability vector preserves the reciprocal-floor bound for the importance-weighted score.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-8068626acd84","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4810,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:186"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem norm_weightedImportanceWeightedLoss_score_le_inv_floor {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob weight loss : History -> Action -> Real) (weightSource : MeasurableFiniteActionDistribution arms weight) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (chosen : Action) : ‖weightedImportanceWeightedLoss arms (prob history) (weight history) (loss history) chosen‖ <= 1 / epsilon","missing":[],"search":"norm_weightedimportanceweightedloss_score_le_inv_floor banditrlproof.exp3.norm_weightedimportanceweightedloss_score_le_inv_floor mixing by any finite probability vector preserves the reciprocal-floor bound for the importance-weighted score. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.norm_mixedSquaredImportanceWeightedLoss_score_le_inv_floor_sq","label":"norm_mixedSquaredImportanceWeightedLoss_score_le_inv_floor_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.norm_mixedSquaredImportanceWeightedLoss_score_le_inv_floor_sq","description":"The mixed second-moment score is bounded by the square reciprocal floor.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-b00bb1594e0b","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4811,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:224"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem norm_mixedSquaredImportanceWeightedLoss_score_le_inv_floor_sq {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (history : History) (chosen : Action) : ‖mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) chosen‖ <= (1 / epsilon) ^ 2","missing":[],"search":"norm_mixedsquaredimportanceweightedloss_score_le_inv_floor_sq banditrlproof.exp3.norm_mixedsquaredimportanceweightedloss_score_le_inv_floor_sq the mixed second-moment score is bounded by the square reciprocal floor. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_importanceWeightedLoss_score","label":"integrable_importanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_importanceWeightedLoss_score","description":"The armwise score is integrable under the generated history/action law.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-63b9f40cc0ea","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4812,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:268"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_importanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (action : Action) (haction : action ∈ arms) : Integrable (fun sample : History × Action => importanceWeightedLoss (prob sample.1) (loss sample.1) sample.2 action) (Measure.compProd historyMu (finiteActionKernel arms prob source))","missing":[],"search":"integrable_importanceweightedloss_score banditrlproof.exp3.integrable_importanceweightedloss_score the armwise score is integrable under the generated history/action law. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_mixedImportanceWeightedLoss_score","label":"integrable_mixedImportanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_mixedImportanceWeightedLoss_score","description":"The mixed first-moment score is integrable under the generated law.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-75765a97e353","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4813,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:292"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_mixedImportanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Integrable (fun sample : History × Action => mixedImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2) (Measure.compProd historyMu (finiteActionKernel arms prob source))","missing":[],"search":"integrable_mixedimportanceweightedloss_score banditrlproof.exp3.integrable_mixedimportanceweightedloss_score the mixed first-moment score is integrable under the generated law. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_weightedImportanceWeightedLoss_score","label":"integrable_weightedImportanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_weightedImportanceWeightedLoss_score","description":"A predictably weighted importance-weighted score is integrable under the sampling law when both finite distributions are measurable.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-4c613dd4c188","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4814,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:316"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_weightedImportanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob weight loss : History -> Action -> Real) (probSource : MeasurableFiniteActionDistribution arms prob) (weightSource : MeasurableFiniteActionDistribution arms weight) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Integrable (fun sample : History × Action => weightedImportanceWeightedLoss arms (prob sample.1) (weight sample.1) (loss sample.1) sample.2) (Measure.compProd historyMu (finiteActionKernel arms prob probSource))","missing":[],"search":"integrable_weightedimportanceweightedloss_score banditrlproof.exp3.integrable_weightedimportanceweightedloss_score a predictably weighted importance-weighted score is integrable under the sampling law when both finite distributions are measurable. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_mixedSquaredImportanceWeightedLoss_score","label":"integrable_mixedSquaredImportanceWeightedLoss_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_mixedSquaredImportanceWeightedLoss_score","description":"The mixed second-moment score is integrable under the generated law.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-26251520be58","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4815,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:340"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_mixedSquaredImportanceWeightedLoss_score {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : Integrable (fun sample : History × Action => mixedSquaredImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2) (Measure.compProd historyMu (finiteActionKernel arms prob source))","missing":[],"search":"integrable_mixedsquaredimportanceweightedloss_score banditrlproof.exp3.integrable_mixedsquaredimportanceweightedloss_score the mixed second-moment score is integrable under the generated law. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_importanceWeightedLoss_selected_of_isFiniteMeasure","label":"integrable_importanceWeightedLoss_selected_of_isFiniteMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_importanceWeightedLoss_selected_of_isFiniteMeasure","description":"A bounded armwise score remains integrable when the sampled action is any measurable function on a finite history measure.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-4f2affcd0b1a","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4816,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:364"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_importanceWeightedLoss_selected_of_isFiniteMeasure {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure History) [IsFiniteMeasure mu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (chosen : History -> Action) (hchosen : Measurable chosen) (comparator : Action) (hcomparator : comparator ∈ arms) : Integrable (fun history => importanceWeightedLoss (prob history) (loss history) (chosen history) comparator) mu","missing":[],"search":"integrable_importanceweightedloss_selected_of_isfinitemeasure banditrlproof.exp3.integrable_importanceweightedloss_selected_of_isfinitemeasure a bounded armwise score remains integrable when the sampled action is any measurable function on a finite history measure. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_weightedImportanceWeightedLoss_selected_of_isFiniteMeasure","label":"integrable_weightedImportanceWeightedLoss_selected_of_isFiniteMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_weightedImportanceWeightedLoss_selected_of_isFiniteMeasure","description":"A predictably weighted score remains integrable when the sampled action is an arbitrary measurable function on a finite measure.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-5b429d3966c6","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4817,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:391"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_weightedImportanceWeightedLoss_selected_of_isFiniteMeasure {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure History) [IsFiniteMeasure mu] (arms : Finset Action) (prob weight loss : History -> Action -> Real) (probSource : MeasurableFiniteActionDistribution arms prob) (weightSource : MeasurableFiniteActionDistribution arms weight) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (chosen : History -> Action) (hchosen : Measurable chosen) : Integrable (fun history => weightedImportanceWeightedLoss arms (prob history) (weight history) (loss history) (chosen history)) mu","missing":[],"search":"integrable_weightedimportanceweightedloss_selected_of_isfinitemeasure banditrlproof.exp3.integrable_weightedimportanceweightedloss_selected_of_isfinitemeasure a predictably weighted score remains integrable when the sampled action is an arbitrary measurable function on a finite measure. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.integrable_mixedSquaredImportanceWeightedLoss_selected_of_isFiniteMeasure","label":"integrable_mixedSquaredImportanceWeightedLoss_selected_of_isFiniteMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.integrable_mixedSquaredImportanceWeightedLoss_selected_of_isFiniteMeasure","description":"A bounded mixed second-moment score remains integrable when the sampled action is any measurable function on a finite history measure.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-51232aae749d","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4818,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:418"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem integrable_mixedSquaredImportanceWeightedLoss_selected_of_isFiniteMeasure {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure History) [IsFiniteMeasure mu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (chosen : History -> Action) (hchosen : Measurable chosen) : Integrable (fun history => mixedSquaredImportanceWeightedLoss arms (prob history) (loss history) (chosen history)) mu","missing":[],"search":"integrable_mixedsquaredimportanceweightedloss_selected_of_isfinitemeasure banditrlproof.exp3.integrable_mixedsquaredimportanceweightedloss_selected_of_isfinitemeasure a bounded mixed second-moment score remains integrable when the sampled action is any measurable function on a finite history measure. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_integral_importanceWeightedLoss_eq_integral_loss_of_regularity","label":"actionProcess_integral_importanceWeightedLoss_eq_integral_loss_of_regularity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_integral_importanceWeightedLoss_eq_integral_loss_of_regularity","description":"Canonical armwise identity with score regularity inferred from bounded losses.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-fc8ab7845baf","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4819,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:443"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_importanceWeightedLoss_eq_integral_loss_of_regularity {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) (comparator : Action) (hcomparator : comparator ∈ arms) : integral (actionProcessMeasure historyMu arms prob source) (fun sample => importanceWeightedLoss (prob (actionProcessHistory sample)) (loss (actionProcessHistory sample)) (actionProcessAction sample) comparator) = integral historyMu (fun history => loss history comparator)","missing":[],"search":"actionprocess_integral_importanceweightedloss_eq_integral_loss_of_regularity banditrlproof.exp3.actionprocess_integral_importanceweightedloss_eq_integral_loss_of_regularity canonical armwise identity with score regularity inferred from bounded losses. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_regularity","label":"actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_regularity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_regularity","description":"Canonical mixed first-moment identity with score regularity inferred.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-f7f46be82aad","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4820,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:470"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_regularity {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : integral (actionProcessMeasure historyMu arms prob source) (fun sample => mixedImportanceWeightedLoss arms (prob (actionProcessHistory sample)) (loss (actionProcessHistory sample)) (actionProcessAction sample)) = integral historyMu (fun history => arms.sum (fun action => prob history action * loss history action))","missing":[],"search":"actionprocess_integral_mixedimportanceweightedloss_eq_integral_mixedloss_of_regularity banditrlproof.exp3.actionprocess_integral_mixedimportanceweightedloss_eq_integral_mixedloss_of_regularity canonical mixed first-moment identity with score regularity inferred. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_regularity","label":"actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_regularity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_regularity","description":"Canonical mixed second-moment identity with score regularity inferred.","url":"../modules/banditrlproof-exp3scoreregularity/index.html#decl-8c186670d17b","parent":"module:BanditRLProof.Exp3ScoreRegularity","order":4821,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3ScoreRegularity"],["Source","BanditRLProof/Exp3ScoreRegularity.lean:496"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_regularity {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : MeasurableFiniteActionDistribution arms prob) (epsilon : Real) (regularity : BoundedMeasurableLossWithProbabilityFloor arms prob loss epsilon) : integral (actionProcessMeasure historyMu arms prob source) (fun sample => mixedSquaredImportanceWeightedLoss arms (prob (actionProcessHistory sample)) (loss (actionProcessHistory sample)) (actionProcessAction sample)) = integral historyMu (fun history => arms.sum (fun action => (loss history action) ^ 2))","missing":[],"search":"actionprocess_integral_mixedsquaredimportanceweightedloss_eq_integral_sum_loss_sq_of_regularity banditrlproof.exp3.actionprocess_integral_mixedsquaredimportanceweightedloss_eq_integral_sum_loss_sq_of_regularity canonical mixed second-moment identity with score regularity inferred. theorem compiled","shard":"modules/48999943f50a0b47.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_exploredExpectedRegret_le_horizon","label":"sampledPredictable_exploredExpectedRegret_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_exploredExpectedRegret_le_horizon","description":"Any generated predictable EXP3 process has expected exploration-mixed regret at most the horizon, independently of the learning rate.","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-001cddac7520","parent":"module:BanditRLProof.Exp3UniformRegret","order":4822,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:20"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_exploredExpectedRegret_le_horizon {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryExploredPredictableLossAt arms eta gamma loss t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)) <= (horizon : Real)","missing":[],"search":"sampledpredictable_exploredexpectedregret_le_horizon banditrlproof.exp3.sampledpredictable_exploredexpectedregret_le_horizon any generated predictable exp3 process has expected exploration-mixed regret at most the horizon, independently of the learning rate. theorem compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le_horizon","label":"sampledPredictable_realizedExpectedRegret_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le_horizon","description":"The same horizon bound for the scalar losses actually generated by the predictable environment.","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-e6b327ea5a3a","parent":"module:BanditRLProof.Exp3UniformRegret","order":4823,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:93"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_realizedExpectedRegret_le_horizon {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta gamma : Real) (hgamma_nonneg : 0 <= gamma) (hgamma_le_one : gamma <= 1) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) : let mu := prior ⊗ₘ sampledImportanceWeightedTrajectoryKernel arms harms eta gamma hgamma_nonneg hgamma_le_one loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)) <= (horizon : Real)","missing":[],"search":"sampledpredictable_realizedexpectedregret_le_horizon banditrlproof.exp3.sampledpredictable_realizedexpectedregret_le_horizon the same horizon bound for the scalar losses actually generated by the predictable environment. theorem compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.clippedExplorationRate","label":"clippedExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.clippedExplorationRate","description":"Exploration rate clipped to the range supported uniformly by the compiled expected-regret route.","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-8388577db443","parent":"module:BanditRLProof.Exp3UniformRegret","order":4824,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:150"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedExplorationRate (K T : Real) : Real","missing":[],"search":"clippedexplorationrate banditrlproof.exp3.clippedexplorationrate exploration rate clipped to the range supported uniformly by the compiled expected-regret route. definition compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.clippedLearningRate","label":"clippedLearningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.clippedLearningRate","description":"Learning rate paired with `clippedExplorationRate`.","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-5bacce24749c","parent":"module:BanditRLProof.Exp3UniformRegret","order":4825,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:154"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedLearningRate (K T : Real) : Real","missing":[],"search":"clippedlearningrate banditrlproof.exp3.clippedlearningrate learning rate paired with `clippedexplorationrate`. definition compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.clippedExplorationRate_nonneg","label":"clippedExplorationRate_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.clippedExplorationRate_nonneg","description":"theorem clippedExplorationRate_nonneg (K T : Real) : 0 <= clippedExplorationRate K T","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-fbe5d22c2360","parent":"module:BanditRLProof.Exp3UniformRegret","order":4826,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:157"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem clippedExplorationRate_nonneg (K T : Real) : 0 <= clippedExplorationRate K T","missing":[],"search":"clippedexplorationrate_nonneg banditrlproof.exp3.clippedexplorationrate_nonneg theorem clippedexplorationrate_nonneg (k t : real) : 0 <= clippedexplorationrate k t theorem compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.clippedExplorationRate_le_half","label":"clippedExplorationRate_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.clippedExplorationRate_le_half","description":"theorem clippedExplorationRate_le_half (K T : Real) : clippedExplorationRate K T <= 1 / 2","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-b58c4aced7dc","parent":"module:BanditRLProof.Exp3UniformRegret","order":4827,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:162"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem clippedExplorationRate_le_half (K T : Real) : clippedExplorationRate K T <= 1 / 2","missing":[],"search":"clippedexplorationrate_le_half banditrlproof.exp3.clippedexplorationrate_le_half theorem clippedexplorationrate_le_half (k t : real) : clippedexplorationrate k t <= 1 / 2 theorem compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.clippedExplorationRate_eq_tuned","label":"clippedExplorationRate_eq_tuned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.clippedExplorationRate_eq_tuned","description":"theorem clippedExplorationRate_eq_tuned (K T : Real) (h : tunedExplorationRate K T <= 1 / 2) : clippedExplorationRate K T = tunedExplorationRate K T","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-3ff61f99d5b4","parent":"module:BanditRLProof.Exp3UniformRegret","order":4828,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:166"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem clippedExplorationRate_eq_tuned (K T : Real) (h : tunedExplorationRate K T <= 1 / 2) : clippedExplorationRate K T = tunedExplorationRate K T","missing":[],"search":"clippedexplorationrate_eq_tuned banditrlproof.exp3.clippedexplorationrate_eq_tuned theorem clippedexplorationrate_eq_tuned (k t : real) (h : tunedexplorationrate k t <= 1 / 2) : clippedexplorationrate k t = tunedexplorationrate k t theorem compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.clippedPredictableTrajectoryKernel","label":"clippedPredictableTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.clippedPredictableTrajectoryKernel","description":"Generated predictable EXP3 kernel using the clipped all-horizon rates.","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-c8bc17c9b8ba","parent":"module:BanditRLProof.Exp3UniformRegret","order":4829,"meta":[["Kind","definition"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:172"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedPredictableTrajectoryKernel {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (loss : PredictableLossVector Env Action) (horizon : Nat) : Kernel Env (Nat -> Action × Real)","missing":[],"search":"clippedpredictabletrajectorykernel banditrlproof.exp3.clippedpredictabletrajectorykernel generated predictable exp3 kernel using the clipped all-horizon rates. definition compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.sampledPredictable_clippedRealizedExpectedRegret_le_min","label":"sampledPredictable_clippedRealizedExpectedRegret_le_min","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.sampledPredictable_clippedRealizedExpectedRegret_le_min","description":"Uniform-horizon expected regret for the scalar loss actually generated by predictable EXP3. The minimum records the trivial short-horizon budget and the large-horizon square-root theorem without imposing a regime assumption.","url":"../modules/banditrlproof-exp3uniformregret/index.html#decl-8c53c84b8c6c","parent":"module:BanditRLProof.Exp3UniformRegret","order":4830,"meta":[["Kind","theorem"],["Module","BanditRLProof.Exp3UniformRegret"],["Source","BanditRLProof/Exp3UniformRegret.lean:190"],["Chapter","EXP3"],["Used in books","bandit, online-learning"],["Reading references","teaching:exp3"],["Indexed settings","None registered"]],"statement":"theorem sampledPredictable_clippedRealizedExpectedRegret_le_min {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (hcard_two : 2 <= arms.card) (loss : PredictableLossVector Env Action) (horizon : Nat) (comparator : Action) (hcomparator : comparator ∈ arms) : let K := (arms.card : Real) let T := (horizon : Real) let mu := prior ⊗ₘ clippedPredictableTrajectoryKernel arms harms loss horizon integral mu (fun sample => (Finset.range horizon).sum (fun t => sampledTrajectoryRealizedLossAt t sample) - (Finset.range horizon).sum (fun t => predictableLossAt loss t sample comparator)) <= min T (4 * Real.sqrt (K * T * Real.log K))","missing":[],"search":"sampledpredictable_clippedrealizedexpectedregret_le_min banditrlproof.exp3.sampledpredictable_clippedrealizedexpectedregret_le_min uniform-horizon expected regret for the scalar loss actually generated by predictable exp3. the minimum records the trivial short-horizon budget and the large-horizon square-root theorem without imposing a regime assumption. theorem compiled","shard":"modules/12e447357fda0ba1.json","books":["bandit","online-learning"],"chapters":["teaching:exp3"],"settings":[]},{"id":"declaration:BanditRLProof.ExpectationBochnerSums.integral_finset_sum","label":"integral_finset_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ExpectationBochnerSums.integral_finset_sum","description":"The Bochner integral distributes over a finite sum of integrable terms. This is the `EXP-FINITE-SUM` import wrapper. It is polymorphic in the Bochner codomain; bandit expected-regret applications typically instantiate `E := Real`.","url":"../modules/banditrlproof-expectationbochnersums/index.html#decl-98c91b8d1eac","parent":"module:BanditRLProof.ExpectationBochnerSums","order":4831,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationBochnerSums"],["Source","BanditRLProof/ExpectationBochnerSums.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_finset_sum {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} {E : Type w} [NormedAddCommGroup E] [NormedSpace Real E] (mu : Measure Omega) (s : Finset Idx) (f : Idx -> Omega -> E) (hf : forall i, i ∈ s -> Integrable (f i) mu) : MeasureTheory.integral mu (fun omega : Omega => s.sum (fun i => f i omega)) = s.sum (fun i => MeasureTheory.integral mu (f i))","missing":[],"search":"integral_finset_sum banditrlproof.expectationbochnersums.integral_finset_sum the bochner integral distributes over a finite sum of integrable terms. this is the `exp-finite-sum` import wrapper. it is polymorphic in the bochner codomain; bandit expected-regret applications typically instantiate `e := real`. theorem compiled","shard":"modules/a7c5ae2e45e6eff4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ExpectationBochnerSums.integral_univ_sum","label":"integral_univ_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ExpectationBochnerSums.integral_univ_sum","description":"Finite-type specialization of `integral_finset_sum`. The statement exposes the common finite-arm `(Finset.univ : Finset Idx)` shape used by regret decompositions.","url":"../modules/banditrlproof-expectationbochnersums/index.html#decl-a567e47a2adc","parent":"module:BanditRLProof.ExpectationBochnerSums","order":4832,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationBochnerSums"],["Source","BanditRLProof/ExpectationBochnerSums.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_univ_sum {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} [Fintype Idx] {E : Type w} [NormedAddCommGroup E] [NormedSpace Real E] (mu : Measure Omega) (f : Idx -> Omega -> E) (hf : forall i : Idx, Integrable (f i) mu) : MeasureTheory.integral mu (fun omega : Omega => (Finset.univ : Finset Idx).sum (fun i => f i omega)) = (Finset.univ : Finset Idx).sum (fun i => MeasureTheory.integral mu (f i))","missing":[],"search":"integral_univ_sum banditrlproof.expectationbochnersums.integral_univ_sum finite-type specialization of `integral_finset_sum`. the statement exposes the common finite-arm `(finset.univ : finset idx)` shape used by regret decompositions. theorem compiled","shard":"modules/a7c5ae2e45e6eff4.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_univ_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","label":"lintegral_univ_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_univ_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","description":"The `Fin K`/`Finset.univ` specialization of the `ENNReal` weighted pull-count budget bound under a probability measure. This is the `EXP-WEIGHTED-PULLCOUNT-LE-TIME-FIN` bridge. It is still only an `ENNReal` finite-action probability-count bound; it does not introduce `FiniteBanditModel`, `Rat`/`Real`, or Bochner expectation.","url":"../modules/banditrlproof-expectationfinitebanditbounds/index.html#decl-7334862e3b11","parent":"module:BanditRLProof.ExpectationFiniteBanditBounds","order":4833,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationFiniteBanditBounds"],["Source","BanditRLProof/ExpectationFiniteBanditBounds.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_univ_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (gap : Fin K -> ENNReal) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => gap a * ((pullCount (action omega) a n : Nat) : ENNReal))) <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => gap a * (n : ENNReal))","missing":[],"search":"lintegral_univ_sum_gap_mul_natcast_pullcount_le_sum_gap_mul_time banditrlproof.lintegral_univ_sum_gap_mul_natcast_pullcount_le_sum_gap_mul_time the `fin k`/`finset.univ` specialization of the `ennreal` weighted pull-count budget bound under a probability measure. this is the `exp-weighted-pullcount-le-time-fin` bridge. it is still only an `ennreal` finite-action probability-count bound; it does not introduce `finitebanditmodel`, `rat`/`real`, or bochner expectation. theorem compiled","shard":"modules/69a6cc9ead26a426.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_univ_sum_model_gap_ofReal_mul_natCast_pullCount_le_sum_model_gap_ofReal_mul_time","label":"lintegral_univ_sum_model_gap_ofReal_mul_natCast_pullCount_le_sum_model_gap_ofReal_mul_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_univ_sum_model_gap_ofReal_mul_natCast_pullCount_le_sum_model_gap_ofReal_mul_time","description":"The finite-arm weighted pull-count budget bound instantiated with the `ENNReal.ofReal` image of the local model gap. This is the `EXP-MODEL-GAP-OFREAL-BOUND` bridge. It is an executable Rat-to-`ENNReal` model-gap wrapper, not a Bochner expected-regret theorem and not a proof that Rat-valued pseudo-regret equals this lower integral.","url":"../modules/banditrlproof-expectationfinitebanditmodelbounds/index.html#decl-394c25bcfed8","parent":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","order":4834,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationFiniteBanditModelBounds"],["Source","BanditRLProof/ExpectationFiniteBanditModelBounds.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_univ_sum_model_gap_ofReal_mul_natCast_pullCount_le_sum_model_gap_ofReal_mul_time {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ENNReal.ofReal (((model.gap a : Rat) : Real)) * ((pullCount (action omega) a n : Nat) : ENNReal))) <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ENNReal.ofReal (((model.gap a : Rat) : Real)) * (n : ENNReal))","missing":[],"search":"lintegral_univ_sum_model_gap_ofreal_mul_natcast_pullcount_le_sum_model_gap_ofreal_mul_time banditrlproof.lintegral_univ_sum_model_gap_ofreal_mul_natcast_pullcount_le_sum_model_gap_ofreal_mul_time the finite-arm weighted pull-count budget bound instantiated with the `ennreal.ofreal` image of the local model gap. this is the `exp-model-gap-ofreal-bound` bridge. it is an executable rat-to-`ennreal` model-gap wrapper, not a bochner expected-regret theorem and not a proof that rat-valued pseudo-regret equals this lower integral. theorem compiled","shard":"modules/2070dac73dddce67.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_actionTrace_eval_eq_indicator_one","label":"lintegral_actionTrace_eval_eq_indicator_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_actionTrace_eval_eq_indicator_one","description":"The lower integral of the indicator of a measurable pull event is the measure of that event. This is the `EXP-INDICATOR-PULL` canary. It uses an arbitrary measure and an `ENNReal` indicator, so it does not choose a Bochner expectation or probability measure interface yet.","url":"../modules/banditrlproof-expectationfoundation/index.html#decl-536f57bb6d23","parent":"module:BanditRLProof.ExpectationFoundation","order":4835,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationFoundation"],["Source","BanditRLProof/ExpectationFoundation.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_actionTrace_eval_eq_indicator_one {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (t : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => (({omega' : Omega | action omega' t = a} : Set Omega).indicator (1 : Omega -> ENNReal)) omega) = mu {omega : Omega | action omega t = a}","missing":[],"search":"lintegral_actiontrace_eval_eq_indicator_one banditrlproof.lintegral_actiontrace_eval_eq_indicator_one the lower integral of the indicator of a measurable pull event is the measure of that event. this is the `exp-indicator-pull` canary. it uses an arbitrary measure and an `ennreal` indicator, so it does not choose a bochner expectation or probability measure interface yet. theorem compiled","shard":"modules/6fbaa91da864ae05.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_nonneg","label":"lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_nonneg","description":"Under explicit nonnegativity of model gaps, the lower integral of `ENNReal.ofReal` pseudo-regret is bounded by the finite-arm model-gap horizon budget. This is the `EXP-OFREAL-PSEUDOREGRET-BOUND` leaf. It consumes the pointwise scalar/model pseudo-regret bridge and the existing finite-bandit model-gap lower-integral bound; it does not prove gap nonnegativity, Bochner expectation, filtration, or concentration results.","url":"../modules/banditrlproof-expectationpseudoregretofrealbounds/index.html#decl-e0f47b77e765","parent":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","order":4836,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationPseudoRegretOfRealBounds"],["Source","BanditRLProof/ExpectationPseudoRegretOfRealBounds.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_nonneg {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hgap : forall a : Fin K, 0 <= (((model.gap a : Rat) : Real))) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => ENNReal.ofReal (((pseudoRegret model (action omega) n : Rat) : Real))) <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ENNReal.ofReal (((model.gap a : Rat) : Real)) * (n : ENNReal))","missing":[],"search":"lintegral_ofreal_pseudoregret_le_sum_model_gap_ofreal_mul_time_of_nonneg banditrlproof.lintegral_ofreal_pseudoregret_le_sum_model_gap_ofreal_mul_time_of_nonneg under explicit nonnegativity of model gaps, the lower integral of `ennreal.ofreal` pseudo-regret is bounded by the finite-arm model-gap horizon budget. this is the `exp-ofreal-pseudoregret-bound` leaf. it consumes the pointwise scalar/model pseudo-regret bridge and the existing finite-bandit model-gap lower-integral bound; it does not prove gap nonnegativity, bochner expectation, filtration, or concentration results. theorem compiled","shard":"modules/4aea156c48d03fe0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_rat_gap_nonneg","label":"lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_rat_gap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_rat_gap_nonneg","description":"The `ENNReal.ofReal` lower-integral pseudo-regret bound under a Rat-level nonnegativity contract for model gaps. This is the `EXP-OFREAL-PSEUDOREGRET-BOUND-OF-RAT-GAP-NONNEG` adapter. It does not prove model-derived gap nonnegativity and is not a Bochner expected regret theorem.","url":"../modules/banditrlproof-expectationpseudoregretratbounds/index.html#decl-30e3840496f2","parent":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","order":4837,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationPseudoRegretRatBounds"],["Source","BanditRLProof/ExpectationPseudoRegretRatBounds.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_rat_gap_nonneg {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hgap : forall a : Fin K, (0 : Rat) <= model.gap a) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => ENNReal.ofReal (((pseudoRegret model (action omega) n : Rat) : Real))) <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ENNReal.ofReal (((model.gap a : Rat) : Real)) * (n : ENNReal))","missing":[],"search":"lintegral_ofreal_pseudoregret_le_sum_model_gap_ofreal_mul_time_of_rat_gap_nonneg banditrlproof.lintegral_ofreal_pseudoregret_le_sum_model_gap_ofreal_mul_time_of_rat_gap_nonneg the `ennreal.ofreal` lower-integral pseudo-regret bound under a rat-level nonnegativity contract for model gaps. this is the `exp-ofreal-pseudoregret-bound-of-rat-gap-nonneg` adapter. it does not prove model-derived gap nonnegativity and is not a bochner expected regret theorem. theorem compiled","shard":"modules/0f71a05fd49771da.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time","label":"lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time","description":"The `ENNReal.ofReal` lower-integral pseudo-regret bound with model-derived gap nonnegativity. This is the `EXP-OFREAL-PSEUDOREGRET-BOUND-MODEL-GAP` adapter. It consumes `FiniteBanditModel.gap_nonneg`; it is still an `ENNReal.ofReal` lower-integral surrogate, not a Rat-valued or Bochner expected-regret theorem.","url":"../modules/banditrlproof-expectationpseudoregretratbounds/index.html#decl-c7c37ababe0d","parent":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","order":4838,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationPseudoRegretRatBounds"],["Source","BanditRLProof/ExpectationPseudoRegretRatBounds.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time {Omega : Type u} {K : Nat} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => ENNReal.ofReal (((pseudoRegret model (action omega) n : Rat) : Real))) <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ENNReal.ofReal (((model.gap a : Rat) : Real)) * (n : ENNReal))","missing":[],"search":"lintegral_ofreal_pseudoregret_le_sum_model_gap_ofreal_mul_time banditrlproof.lintegral_ofreal_pseudoregret_le_sum_model_gap_ofreal_mul_time the `ennreal.ofreal` lower-integral pseudo-regret bound with model-derived gap nonnegativity. this is the `exp-ofreal-pseudoregret-bound-model-gap` adapter. it consumes `finitebanditmodel.gap_nonneg`; it is still an `ennreal.ofreal` lower-integral surrogate, not a rat-valued or bochner expected-regret theorem. theorem compiled","shard":"modules/0f71a05fd49771da.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ennreal_natCast_pullCount_eq_finset_range_indicator_one","label":"ennreal_natCast_pullCount_eq_finset_range_indicator_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ennreal_natCast_pullCount_eq_finset_range_indicator_one","description":"private theorem ennreal_natCast_pullCount_eq_finset_range_indicator_one {Omega : Type u} {Action : Type v} [DecidableEq Action] (action : Omega -> ActionTrace Action) (a : Action) (n : Nat) (omega : Omega) : ((pullCount (action omega) a n : Nat) : ENNReal) = (Finset.range n).sum (fun t : Nat => (({omega' : Omega | action omega' t = a} : Set Omega).indicator (1 : Omega -> ENNReal)) omega)","url":"../modules/banditrlproof-expectationpullcount/index.html#decl-dd71e2109056","parent":"module:BanditRLProof.ExpectationPullCount","order":4839,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationPullCount"],["Source","BanditRLProof/ExpectationPullCount.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem ennreal_natCast_pullCount_eq_finset_range_indicator_one {Omega : Type u} {Action : Type v} [DecidableEq Action] (action : Omega -> ActionTrace Action) (a : Action) (n : Nat) (omega : Omega) : ((pullCount (action omega) a n : Nat) : ENNReal) = (Finset.range n).sum (fun t : Nat => (({omega' : Omega | action omega' t = a} : Set Omega).indicator (1 : Omega -> ENNReal)) omega)","missing":[],"search":"ennreal_natcast_pullcount_eq_finset_range_indicator_one banditrlproof.ennreal_natcast_pullcount_eq_finset_range_indicator_one private theorem ennreal_natcast_pullcount_eq_finset_range_indicator_one {omega : type u} {action : type v} [decidableeq action] (action : omega -> actiontrace action) (a : action) (n : nat) (omega : omega) : ((pullcount (action omega) a n : nat) : ennreal) = (finset.range n).sum (fun t : nat => (({omega' : omega | action omega' t = a} : set omega).indicator (1 : omega -> ennreal)) omega) theorem compiled","shard":"modules/766c259f05a14bce.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_natCast_pullCount_eq_sum_measure_actionTrace_eval_eq","label":"lintegral_natCast_pullCount_eq_sum_measure_actionTrace_eval_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_natCast_pullCount_eq_sum_measure_actionTrace_eval_eq","description":"The lower integral of the scalar-casted local pull count is the finite sum of the corresponding action-event measures. This is the `EXP-PULLCOUNT-LINTEGRAL` bridge. It connects the compiled finite-sum lower-integral identity to the recursive `pullCount` surface without choosing a Bochner expectation or probability-measure interface.","url":"../modules/banditrlproof-expectationpullcount/index.html#decl-08015ec50017","parent":"module:BanditRLProof.ExpectationPullCount","order":4840,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationPullCount"],["Source","BanditRLProof/ExpectationPullCount.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_natCast_pullCount_eq_sum_measure_actionTrace_eval_eq {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure Omega) (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => ((pullCount (action omega) a n : Nat) : ENNReal)) = (Finset.range n).sum (fun t : Nat => mu {omega : Omega | action omega t = a})","missing":[],"search":"lintegral_natcast_pullcount_eq_sum_measure_actiontrace_eval_eq banditrlproof.lintegral_natcast_pullcount_eq_sum_measure_actiontrace_eval_eq the lower integral of the scalar-casted local pull count is the finite sum of the corresponding action-event measures. this is the `exp-pullcount-lintegral` bridge. it connects the compiled finite-sum lower-integral identity to the recursive `pullcount` surface without choosing a bochner expectation or probability-measure interface. theorem compiled","shard":"modules/766c259f05a14bce.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_natCast_pullCount_le_time","label":"lintegral_natCast_pullCount_le_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_natCast_pullCount_le_time","description":"The lower integral of a scalar-casted pull count is bounded by the horizon under a probability measure. This is the `EXP-PULLCOUNT-LE-TIME` bridge. It is a probability-facing budget bound for expected pull counts, not a Bochner expectation or expected-regret theorem.","url":"../modules/banditrlproof-expectationpullcountbounds/index.html#decl-8d57e2559a3b","parent":"module:BanditRLProof.ExpectationPullCountBounds","order":4841,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationPullCountBounds"],["Source","BanditRLProof/ExpectationPullCountBounds.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_natCast_pullCount_le_time {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => ((pullCount (action omega) a n : Nat) : ENNReal)) <= (n : ENNReal)","missing":[],"search":"lintegral_natcast_pullcount_le_time banditrlproof.lintegral_natcast_pullcount_le_time the lower integral of a scalar-casted pull count is bounded by the horizon under a probability measure. this is the `exp-pullcount-le-time` bridge. it is a probability-facing budget bound for expected pull counts, not a bochner expectation or expected-regret theorem. theorem compiled","shard":"modules/610130ea3000fa86.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.real_pseudoRegret_eq_univ_sum_gap_mul_natCast_pullCount","label":"real_pseudoRegret_eq_univ_sum_gap_mul_natCast_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.real_pseudoRegret_eq_univ_sum_gap_mul_natCast_pullCount","description":"private theorem real_pseudoRegret_eq_univ_sum_gap_mul_natCast_pullCount {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n : Nat) : ((pseudoRegret model action n : Rat) : Real) = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ((model.gap a : Rat) : Real) * ((pullCount action a n : Nat) : Real))","url":"../modules/banditrlproof-expectationregretpullcount/index.html#decl-0c59a09119fe","parent":"module:BanditRLProof.ExpectationRegretPullCount","order":4842,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationRegretPullCount"],["Source","BanditRLProof/ExpectationRegretPullCount.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem real_pseudoRegret_eq_univ_sum_gap_mul_natCast_pullCount {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n : Nat) : ((pseudoRegret model action n : Rat) : Real) = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ((model.gap a : Rat) : Real) * ((pullCount action a n : Nat) : Real))","missing":[],"search":"real_pseudoregret_eq_univ_sum_gap_mul_natcast_pullcount banditrlproof.real_pseudoregret_eq_univ_sum_gap_mul_natcast_pullcount private theorem real_pseudoregret_eq_univ_sum_gap_mul_natcast_pullcount {k : nat} (model : finitebanditmodel k) (action : actiontrace (fin k)) (n : nat) : ((pseudoregret model action n : rat) : real) = (finset.univ : finset (fin k)).sum (fun a : fin k => ((model.gap a : rat) : real) * ((pullcount action a n : nat) : real)) theorem compiled","shard":"modules/97cd0850e41034cf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integrable_real_pullCount_of_measurable_action","label":"integrable_real_pullCount_of_measurable_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integrable_real_pullCount_of_measurable_action","description":"The Real cast of a finite-horizon pull count is integrable under any finite measure when the action trace is measurable one time coordinate at a time. This is the generic regularity adapter behind Real expected-regret wrappers: measurability comes from `measurable_natCast_pullCount`, while `pullCount_le_time` supplies the deterministic integrable bound.","url":"../modules/banditrlproof-expectationregretpullcount/index.html#decl-b9fd49698e4e","parent":"module:BanditRLProof.ExpectationRegretPullCount","order":4843,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationRegretPullCount"],["Source","BanditRLProof/ExpectationRegretPullCount.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_real_pullCount_of_measurable_action {Omega : Type u} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure Omega) [MeasureTheory.IsFiniteMeasure mu] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (arm : Action) (n : Nat) : Integrable (fun omega : Omega => ((pullCount (action omega) arm n : Nat) : Real)) mu","missing":[],"search":"integrable_real_pullcount_of_measurable_action banditrlproof.integrable_real_pullcount_of_measurable_action the real cast of a finite-horizon pull count is integrable under any finite measure when the action trace is measurable one time coordinate at a time. this is the generic regularity adapter behind real expected-regret wrappers: measurability comes from `measurable_natcast_pullcount`, while `pullcount_le_time` supplies the deterministic integrable bound. theorem compiled","shard":"modules/97cd0850e41034cf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integrable_real_pseudoRegret_of_integrable_pullCount","label":"integrable_real_pseudoRegret_of_integrable_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integrable_real_pseudoRegret_of_integrable_pullCount","description":"If every finite-horizon pull count is integrable after casting to `Real`, then the corresponding Real-valued pseudo-regret is integrable. This is the regularity adapter used by the Bochner expected-regret decomposition. It does not prove measurability or integrability from a policy model; callers provide the pull-count integrability witnesses.","url":"../modules/banditrlproof-expectationregretpullcount/index.html#decl-d4be02f452d6","parent":"module:BanditRLProof.ExpectationRegretPullCount","order":4844,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationRegretPullCount"],["Source","BanditRLProof/ExpectationRegretPullCount.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_real_pseudoRegret_of_integrable_pullCount {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (n : Nat) (hcount : forall a : Fin K, Integrable (fun omega : Omega => ((pullCount (action omega) a n : Nat) : Real)) mu) : Integrable (fun omega : Omega => ((pseudoRegret model (action omega) n : Rat) : Real)) mu","missing":[],"search":"integrable_real_pseudoregret_of_integrable_pullcount banditrlproof.integrable_real_pseudoregret_of_integrable_pullcount if every finite-horizon pull count is integrable after casting to `real`, then the corresponding real-valued pseudo-regret is integrable. this is the regularity adapter used by the bochner expected-regret decomposition. it does not prove measurability or integrability from a policy model; callers provide the pull-count integrability witnesses. theorem compiled","shard":"modules/97cd0850e41034cf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integral_real_pseudoRegret_eq_sum_gap_mul_integral_pullCount","label":"integral_real_pseudoRegret_eq_sum_gap_mul_integral_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_real_pseudoRegret_eq_sum_gap_mul_integral_pullCount","description":"The Real-valued Bochner expectation of pseudo-regret is the finite sum of each arm gap multiplied by the Bochner expectation of that arm's pull count. This is the local `EXP-REGRET-PULLCOUNT` leaf. It consumes the deterministic `REGRET-PULLCOUNT` bridge and the Mathlib-backed `EXP-FINITE-SUM` wrapper.","url":"../modules/banditrlproof-expectationregretpullcount/index.html#decl-0abd1f3ae32c","parent":"module:BanditRLProof.ExpectationRegretPullCount","order":4845,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationRegretPullCount"],["Source","BanditRLProof/ExpectationRegretPullCount.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_real_pseudoRegret_eq_sum_gap_mul_integral_pullCount {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (n : Nat) (hcount : forall a : Fin K, Integrable (fun omega : Omega => ((pullCount (action omega) a n : Nat) : Real)) mu) : MeasureTheory.integral mu (fun omega : Omega => ((pseudoRegret model (action omega) n : Rat) : Real)) = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ((model.gap a : Rat) : Real) * MeasureTheory.integral mu (fun omega : Omega => ((pullCount (action omega) a n : Nat) : Real)))","missing":[],"search":"integral_real_pseudoregret_eq_sum_gap_mul_integral_pullcount banditrlproof.integral_real_pseudoregret_eq_sum_gap_mul_integral_pullcount the real-valued bochner expectation of pseudo-regret is the finite sum of each arm gap multiplied by the bochner expectation of that arm's pull count. this is the local `exp-regret-pullcount` leaf. it consumes the deterministic `regret-pullcount` bridge and the mathlib-backed `exp-finite-sum` wrapper. theorem compiled","shard":"modules/97cd0850e41034cf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_finset_sum_actionTrace_eval_eq_indicator_one","label":"lintegral_finset_sum_actionTrace_eval_eq_indicator_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_finset_sum_actionTrace_eval_eq_indicator_one","description":"The lower integral of a finite sum of pull-event indicators is the finite sum of the corresponding event measures. This is the `EXP-FINSET-INDICATOR-PULL` bridge. It isolates lower-integral finite-additivity before connecting the finite sum to `pullCount`.","url":"../modules/banditrlproof-expectationsums/index.html#decl-873224c0cbf3","parent":"module:BanditRLProof.ExpectationSums","order":4846,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationSums"],["Source","BanditRLProof/ExpectationSums.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_finset_sum_actionTrace_eval_eq_indicator_one {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] (mu : Measure Omega) (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (s : Finset Nat) : MeasureTheory.lintegral mu (fun omega : Omega => s.sum (fun t : Nat => (({omega' : Omega | action omega' t = a} : Set Omega).indicator (1 : Omega -> ENNReal)) omega)) = s.sum (fun t : Nat => mu {omega : Omega | action omega t = a})","missing":[],"search":"lintegral_finset_sum_actiontrace_eval_eq_indicator_one banditrlproof.lintegral_finset_sum_actiontrace_eval_eq_indicator_one the lower integral of a finite sum of pull-event indicators is the finite sum of the corresponding event measures. this is the `exp-finset-indicator-pull` bridge. it isolates lower-integral finite-additivity before connecting the finite sum to `pullcount`. theorem compiled","shard":"modules/9048db3abdecc673.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_finset_sum_gap_mul_natCast_pullCount_eq","label":"lintegral_finset_sum_gap_mul_natCast_pullCount_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_finset_sum_gap_mul_natCast_pullCount_eq","description":"The lower integral of a finite weighted sum of scalar-casted pull counts equals the same finite weighted sum of the corresponding action-event measures. This is the `EXP-WEIGHTED-PULLCOUNT-LINTEGRAL` bridge. It is shaped like a nonnegative expected-regret identity, but it deliberately avoids `FiniteBanditModel`, `Rat`, `Real`, Bochner expectation, filtrations, kernels, and concentration assumptions.","url":"../modules/banditrlproof-expectationweightedpullcount/index.html#decl-123dc1f11dac","parent":"module:BanditRLProof.ExpectationWeightedPullCount","order":4847,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationWeightedPullCount"],["Source","BanditRLProof/ExpectationWeightedPullCount.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_finset_sum_gap_mul_natCast_pullCount_eq {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure Omega) (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (gap : Action -> ENNReal) (arms : Finset Action) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => arms.sum (fun a : Action => gap a * ((pullCount (action omega) a n : Nat) : ENNReal))) = arms.sum (fun a : Action => gap a * (Finset.range n).sum (fun t : Nat => mu {omega : Omega | action omega t = a}))","missing":[],"search":"lintegral_finset_sum_gap_mul_natcast_pullcount_eq banditrlproof.lintegral_finset_sum_gap_mul_natcast_pullcount_eq the lower integral of a finite weighted sum of scalar-casted pull counts equals the same finite weighted sum of the corresponding action-event measures. this is the `exp-weighted-pullcount-lintegral` bridge. it is shaped like a nonnegative expected-regret identity, but it deliberately avoids `finitebanditmodel`, `rat`, `real`, bochner expectation, filtrations, kernels, and concentration assumptions. theorem compiled","shard":"modules/2b7426220aad76b5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.lintegral_finset_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","label":"lintegral_finset_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.lintegral_finset_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","description":"The lower integral of a finite weighted sum of scalar-casted pull counts is bounded by the corresponding weighted horizon budget under a probability measure. This is the `EXP-WEIGHTED-PULLCOUNT-LE-TIME` bridge. It is an `ENNReal` probability-count budget bound, not a Bochner expected-regret theorem.","url":"../modules/banditrlproof-expectationweightedpullcountbounds/index.html#decl-c7fb76b430de","parent":"module:BanditRLProof.ExpectationWeightedPullCountBounds","order":4848,"meta":[["Kind","theorem"],["Module","BanditRLProof.ExpectationWeightedPullCountBounds"],["Source","BanditRLProof/ExpectationWeightedPullCountBounds.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_finset_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure Omega) [MeasureTheory.IsProbabilityMeasure mu] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (gap : Action -> ENNReal) (arms : Finset Action) (n : Nat) : MeasureTheory.lintegral mu (fun omega : Omega => arms.sum (fun a : Action => gap a * ((pullCount (action omega) a n : Nat) : ENNReal))) <= arms.sum (fun a : Action => gap a * (n : ENNReal))","missing":[],"search":"lintegral_finset_sum_gap_mul_natcast_pullcount_le_sum_gap_mul_time banditrlproof.lintegral_finset_sum_gap_mul_natcast_pullcount_le_sum_gap_mul_time the lower integral of a finite weighted sum of scalar-casted pull counts is bounded by the corresponding weighted horizon budget under a probability measure. this is the `exp-weighted-pullcount-le-time` bridge. it is an `ennreal` probability-count budget bound, not a bochner expected-regret theorem. theorem compiled","shard":"modules/9c04fa97f43b1e61.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.linearLoss","label":"linearLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FTRL.linearLoss","description":"Finite-action linear loss of a weight vector against a loss vector.","url":"../modules/banditrlproof-ftrlonestep/index.html#decl-8db2b3a9446f","parent":"module:BanditRLProof.FTRLOneStep","order":4849,"meta":[["Kind","definition"],["Module","BanditRLProof.FTRLOneStep"],["Source","BanditRLProof/FTRLOneStep.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def linearLoss {Action : Type u} (arms : Finset Action) (p loss : Action -> Real) : Real","missing":[],"search":"linearloss banditrlproof.ftrl.linearloss finite-action linear loss of a weight vector against a loss vector. definition compiled","shard":"modules/6ac6b62dc49dfa98.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.finiteSimplex","label":"finiteSimplex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FTRL.finiteSimplex","description":"The finite probability-simplex predicate over an explicit action set.","url":"../modules/banditrlproof-ftrlonestep/index.html#decl-c13b3a4ca87a","parent":"module:BanditRLProof.FTRLOneStep","order":4850,"meta":[["Kind","definition"],["Module","BanditRLProof.FTRLOneStep"],["Source","BanditRLProof/FTRLOneStep.lean:30"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def finiteSimplex {Action : Type u} (arms : Finset Action) (p : Action -> Real) : Prop","missing":[],"search":"finitesimplex banditrlproof.ftrl.finitesimplex the finite probability-simplex predicate over an explicit action set. definition compiled","shard":"modules/6ac6b62dc49dfa98.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.regularizedObjective","label":"regularizedObjective","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FTRL.regularizedObjective","description":"Regularized one-round FTRL objective `eta * <p, loss> + R p`. The learning-rate positivity contract is kept on the theorem, not the definition, so future leaves can reuse the objective algebraically.","url":"../modules/banditrlproof-ftrlonestep/index.html#decl-699c480246ef","parent":"module:BanditRLProof.FTRLOneStep","order":4851,"meta":[["Kind","definition"],["Module","BanditRLProof.FTRLOneStep"],["Source","BanditRLProof/FTRLOneStep.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def regularizedObjective {Action : Type u} (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Action -> Real) (p : Action -> Real) : Real","missing":[],"search":"regularizedobjective banditrlproof.ftrl.regularizedobjective regularized one-round ftrl objective `eta * <p, loss> + r p`. the learning-rate positivity contract is kept on the theorem, not the definition, so future leaves can reuse the objective algebraically. definition compiled","shard":"modules/6ac6b62dc49dfa98.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.IsRegularizedMinimizer","label":"IsRegularizedMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FTRL.IsRegularizedMinimizer","description":"A point `p` minimizes the regularized objective over a feasible predicate.","url":"../modules/banditrlproof-ftrlonestep/index.html#decl-6a1735bd44e4","parent":"module:BanditRLProof.FTRLOneStep","order":4852,"meta":[["Kind","definition"],["Module","BanditRLProof.FTRLOneStep"],["Source","BanditRLProof/FTRLOneStep.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def IsRegularizedMinimizer {Action : Type u} (feasible : (Action -> Real) -> Prop) (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Action -> Real) (p : Action -> Real) : Prop","missing":[],"search":"isregularizedminimizer banditrlproof.ftrl.isregularizedminimizer a point `p` minimizes the regularized objective over a feasible predicate. definition compiled","shard":"modules/6ac6b62dc49dfa98.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.linearLoss_sub_le_regularizer_sub_div_of_isRegularizedMinimizer","label":"linearLoss_sub_le_regularizer_sub_div_of_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.linearLoss_sub_le_regularizer_sub_div_of_isRegularizedMinimizer","description":"FTRL one-step inequality from an explicit regularized-objective minimizer. The conclusion is the deterministic algebraic form used before later leaves choose a concrete regularizer or prove a stability/penalty sum.","url":"../modules/banditrlproof-ftrlonestep/index.html#decl-e6cd659f4b73","parent":"module:BanditRLProof.FTRLOneStep","order":4853,"meta":[["Kind","theorem"],["Module","BanditRLProof.FTRLOneStep"],["Source","BanditRLProof/FTRLOneStep.lean:63"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_sub_le_regularizer_sub_div_of_isRegularizedMinimizer {Action : Type u} (feasible : (Action -> Real) -> Prop) (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Action -> Real) (p q : Action -> Real) (heta : 0 < eta) (hp : IsRegularizedMinimizer feasible arms eta regularizer loss p) (hq : feasible q) : linearLoss arms p loss - linearLoss arms q loss <= (regularizer q - regularizer p) / eta","missing":[],"search":"linearloss_sub_le_regularizer_sub_div_of_isregularizedminimizer banditrlproof.ftrl.linearloss_sub_le_regularizer_sub_div_of_isregularizedminimizer ftrl one-step inequality from an explicit regularized-objective minimizer. the conclusion is the deterministic algebraic form used before later leaves choose a concrete regularizer or prove a stability/penalty sum. theorem compiled","shard":"modules/6ac6b62dc49dfa98.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.linearLoss_sub_le_regularizer_sub_div_of_simplex_minimizer","label":"linearLoss_sub_le_regularizer_sub_div_of_simplex_minimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.linearLoss_sub_le_regularizer_sub_div_of_simplex_minimizer","description":"FTRL one-step inequality specialized to the finite simplex.","url":"../modules/banditrlproof-ftrlonestep/index.html#decl-2a5502439d73","parent":"module:BanditRLProof.FTRLOneStep","order":4854,"meta":[["Kind","theorem"],["Module","BanditRLProof.FTRLOneStep"],["Source","BanditRLProof/FTRLOneStep.lean:94"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_sub_le_regularizer_sub_div_of_simplex_minimizer {Action : Type u} (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Action -> Real) (p q : Action -> Real) (heta : 0 < eta) (hp : IsRegularizedMinimizer (finiteSimplex arms) arms eta regularizer loss p) (hq : finiteSimplex arms q) : linearLoss arms p loss - linearLoss arms q loss <= (regularizer q - regularizer p) / eta","missing":[],"search":"linearloss_sub_le_regularizer_sub_div_of_simplex_minimizer banditrlproof.ftrl.linearloss_sub_le_regularizer_sub_div_of_simplex_minimizer ftrl one-step inequality specialized to the finite simplex. theorem compiled","shard":"modules/6ac6b62dc49dfa98.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteArmIntervalVarianceProxy","label":"finiteArmIntervalVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteArmIntervalVarianceProxy","description":"The largest Hoeffding proxy among a finite family of armwise intervals.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-79dacaf78d3e","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4855,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:18"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIntervalVarianceProxy {K : Nat} (lo hi : Fin K -> Real) : NNReal","missing":[],"search":"finitearmintervalvarianceproxy banditrlproof.concentration.finitearmintervalvarianceproxy the largest hoeffding proxy among a finite family of armwise intervals. definition compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.intervalVarianceProxy_le_finiteArmIntervalVarianceProxy","label":"intervalVarianceProxy_le_finiteArmIntervalVarianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.intervalVarianceProxy_le_finiteArmIntervalVarianceProxy","description":"Every armwise interval proxy is bounded by the finite-arm maximum proxy.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-52ca75c80514","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4856,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:23"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem intervalVarianceProxy_le_finiteArmIntervalVarianceProxy {K : Nat} (lo hi : Fin K -> Real) (arm : Fin K) : intervalVarianceProxy (lo arm) (hi arm) <= finiteArmIntervalVarianceProxy lo hi","missing":[],"search":"intervalvarianceproxy_le_finitearmintervalvarianceproxy banditrlproof.concentration.intervalvarianceproxy_le_finitearmintervalvarianceproxy every armwise interval proxy is bounded by the finite-arm maximum proxy. theorem compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteArmIntervalVarianceProxy_pos","label":"finiteArmIntervalVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteArmIntervalVarianceProxy_pos","description":"Nondegenerate armwise intervals give a positive finite-arm maximum proxy.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-9bbd530ec47d","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4857,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:36"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIntervalVarianceProxy_pos {K : Nat} (hK : 0 < K) (lo hi : Fin K -> Real) (hlohi : forall arm, lo arm < hi arm) : 0 < ((finiteArmIntervalVarianceProxy lo hi : NNReal) : Real)","missing":[],"search":"finitearmintervalvarianceproxy_pos banditrlproof.concentration.finitearmintervalvarianceproxy_pos nondegenerate armwise intervals give a positive finite-arm maximum proxy. theorem compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteArmVarianceProxy","label":"finiteArmVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteArmVarianceProxy","description":"The largest variance proxy in a finite family of arms.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-75090b2c7179","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4858,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:55"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmVarianceProxy {K : Nat} (varianceProxy : Fin K -> NNReal) : NNReal","missing":[],"search":"finitearmvarianceproxy banditrlproof.concentration.finitearmvarianceproxy the largest variance proxy in a finite family of arms. definition compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteArmVarianceProxy","label":"varianceProxy_le_finiteArmVarianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.varianceProxy_le_finiteArmVarianceProxy","description":"Every arm proxy is bounded by the finite-arm maximum proxy.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-644f9289a550","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4859,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:60"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem varianceProxy_le_finiteArmVarianceProxy {K : Nat} (varianceProxy : Fin K -> NNReal) (arm : Fin K) : varianceProxy arm <= finiteArmVarianceProxy varianceProxy","missing":[],"search":"varianceproxy_le_finitearmvarianceproxy banditrlproof.concentration.varianceproxy_le_finitearmvarianceproxy every arm proxy is bounded by the finite-arm maximum proxy. theorem compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteArmVarianceProxy_pos_of_exists","label":"finiteArmVarianceProxy_pos_of_exists","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteArmVarianceProxy_pos_of_exists","description":"A positive member makes the finite-arm maximum proxy positive.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-5a2c22722924","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4860,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:71"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteArmVarianceProxy_pos_of_exists {K : Nat} (varianceProxy : Fin K -> NNReal) (hpos : exists arm, 0 < ((varianceProxy arm : NNReal) : Real)) : 0 < ((finiteArmVarianceProxy varianceProxy : NNReal) : Real)","missing":[],"search":"finitearmvarianceproxy_pos_of_exists banditrlproof.concentration.finitearmvarianceproxy_pos_of_exists a positive member makes the finite-arm maximum proxy positive. theorem compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteArmPositiveVarianceProxy","label":"finiteArmPositiveVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteArmPositiveVarianceProxy","description":"The finite-arm maximum padded by one, providing a strictly positive tuning proxy even when all genuine armwise proxies are zero.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-239f7b7fba82","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4861,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:84"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmPositiveVarianceProxy {K : Nat} (varianceProxy : Fin K -> NNReal) : NNReal","missing":[],"search":"finitearmpositivevarianceproxy banditrlproof.concentration.finitearmpositivevarianceproxy the finite-arm maximum padded by one, providing a strictly positive tuning proxy even when all genuine armwise proxies are zero. definition compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteArmPositiveVarianceProxy","label":"varianceProxy_le_finiteArmPositiveVarianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.varianceProxy_le_finiteArmPositiveVarianceProxy","description":"Every armwise proxy is bounded by the positive padded finite-arm proxy.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-b8894dea7f00","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4862,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:89"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem varianceProxy_le_finiteArmPositiveVarianceProxy {K : Nat} (varianceProxy : Fin K -> NNReal) (arm : Fin K) : varianceProxy arm <= finiteArmPositiveVarianceProxy varianceProxy","missing":[],"search":"varianceproxy_le_finitearmpositivevarianceproxy banditrlproof.concentration.varianceproxy_le_finitearmpositivevarianceproxy every armwise proxy is bounded by the positive padded finite-arm proxy. theorem compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteArmPositiveVarianceProxy_pos","label":"finiteArmPositiveVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteArmPositiveVarianceProxy_pos","description":"The padded finite-arm proxy is always strictly positive.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-4e62bcf972ef","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4863,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:97"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteArmPositiveVarianceProxy_pos {K : Nat} (varianceProxy : Fin K -> NNReal) : 0 < ((finiteArmPositiveVarianceProxy varianceProxy : NNReal) : Real)","missing":[],"search":"finitearmpositivevarianceproxy_pos banditrlproof.concentration.finitearmpositivevarianceproxy_pos the padded finite-arm proxy is always strictly positive. theorem compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.contextIndependentCenteredRewardKernelLaw_of_hasSubgaussianMGF","label":"contextIndependentCenteredRewardKernelLaw_of_hasSubgaussianMGF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.contextIndependentCenteredRewardKernelLaw_of_hasSubgaussianMGF","description":"Action-indexed probability laws with exact means and direct centered sub-Gaussian witnesses form a context-independent centered reward-kernel law.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-e14066e62043","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4864,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:112"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def contextIndependentCenteredRewardKernelLaw_of_hasSubgaussianMGF {Context Action : Type} [MeasurableSpace Context] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (armLaw : Action -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (armMean : Action -> Rat) (varianceProxy : Action -> NNReal) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((armMean arm : Rat) : Real)) (hsubG : forall arm, HasSubgaussianMGF (fun reward : Rat => (((reward - armMean arm : Rat) : Real))) (varianceProxy arm) (armLaw arm)) : CenteredRewardKernelLaw (contextIndependentOfActionLaws (Context := Context) armLaw hprob) (fun _ arm => armMean arm) (fun _ arm => varianceProxy arm) where","missing":[],"search":"contextindependentcenteredrewardkernellaw_of_hassubgaussianmgf banditrlproof.rewardkernel.contextindependentcenteredrewardkernellaw_of_hassubgaussianmgf action-indexed probability laws with exact means and direct centered sub-gaussian witnesses form a context-independent centered reward-kernel law. definition compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.contextIndependentBoundedCenteredRewardKernelLaw","label":"contextIndependentBoundedCenteredRewardKernelLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.contextIndependentBoundedCenteredRewardKernelLaw","description":"Common almost-sure interval bounds and exact means form a context-independent centered reward-kernel law with the Hoeffding proxy.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-3cf2abad1d0f","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4865,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:162"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def contextIndependentBoundedCenteredRewardKernelLaw {Context Action : Type} [MeasurableSpace Context] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (armLaw : Action -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (armMean : Action -> Rat) (lo hi : Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc lo hi ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((armMean arm : Rat) : Real)) : CenteredRewardKernelLaw (contextIndependentOfActionLaws (Context := Context) armLaw hprob) (fun _ arm => armMean arm) (fun _ _ => Concentration.intervalVarianceProxy lo hi)","missing":[],"search":"contextindependentboundedcenteredrewardkernellaw banditrlproof.rewardkernel.contextindependentboundedcenteredrewardkernellaw common almost-sure interval bounds and exact means form a context-independent centered reward-kernel law with the hoeffding proxy. definition compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.contextIndependentArmwiseBoundedCenteredRewardKernelLaw","label":"contextIndependentArmwiseBoundedCenteredRewardKernelLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.contextIndependentArmwiseBoundedCenteredRewardKernelLaw","description":"Arm-dependent almost-sure interval bounds and exact means form a context-independent centered reward-kernel law with armwise Hoeffding proxies.","url":"../modules/banditrlproof-finitearmrewardkernellaw/index.html#decl-93a4bf019b57","parent":"module:BanditRLProof.FiniteArmRewardKernelLaw","order":4866,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteArmRewardKernelLaw"],["Source","BanditRLProof/FiniteArmRewardKernelLaw.lean:203"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def contextIndependentArmwiseBoundedCenteredRewardKernelLaw {Context Action : Type} [MeasurableSpace Context] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] (armLaw : Action -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (armMean : Action -> Rat) (lo hi : Action -> Real) (hmeas : forall arm, AEMeasurable (fun reward : Rat => ((reward : Rat) : Real)) (armLaw arm)) (hbound : forall arm, Filter.Eventually (fun reward : Rat => Set.Icc (lo arm) (hi arm) ((reward : Rat) : Real)) (ae (armLaw arm))) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((armMean arm : Rat) : Real)) : CenteredRewardKernelLaw (contextIndependentOfActionLaws (Context := Context) armLaw hprob) (fun _ arm => armMean arm) (fun _ arm => Concentration.intervalVarianceProxy (lo arm) (hi arm))","missing":[],"search":"contextindependentarmwiseboundedcenteredrewardkernellaw banditrlproof.rewardkernel.contextindependentarmwiseboundedcenteredrewardkernellaw arm-dependent almost-sure interval bounds and exact means form a context-independent centered reward-kernel law with armwise hoeffding proxies. definition compiled","shard":"modules/a4afe9b44c4fbe74.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.mean_le_foldl_select","label":"mean_le_foldl_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.mean_le_foldl_select","description":"private theorem mean_le_foldl_select {K : Nat} (mean : Fin K -> Rat) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, a ∈ l -> mean a <= mean (l.foldl (fun best arm : Fin K => if mean best < mean arm then arm else best) init)) /\\ mean init <= mean (l.foldl (fun best arm : Fin K => if mean best < mean arm then arm else best) init) | [] => by simp | arm :: rest => by let select := fun best arm : Fin K => i…","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html#decl-e8ad6493425d","parent":"module:BanditRLProof.FiniteBanditModelInvariants","order":4867,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteBanditModelInvariants"],["Source","BanditRLProof/FiniteBanditModelInvariants.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem mean_le_foldl_select {K : Nat} (mean : Fin K -> Rat) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, a ∈ l -> mean a <= mean (l.foldl (fun best arm : Fin K => if mean best < mean arm then arm else best) init)) /\\ mean init <= mean (l.foldl (fun best arm : Fin K => if mean best < mean arm then arm else best) init) | [] => by simp | arm :: rest => by let select := fun best arm : Fin K => if mean best < mean arm then arm else best let next := select init arm have ih","missing":[],"search":"mean_le_foldl_select banditrlproof.finitebanditmodel.mean_le_foldl_select private theorem mean_le_foldl_select {k : nat} (mean : fin k -> rat) (init : fin k) : forall l : list (fin k), (forall a : fin k, a ∈ l -> mean a <= mean (l.foldl (fun best arm : fin k => if mean best < mean arm then arm else best) init)) /\\ mean init <= mean (l.foldl (fun best arm : fin k => if mean best < mean arm then arm else best) init) | [] => by simp | arm :: rest => by let select := fun best arm : fin k => if mean best < mean arm then arm else best let next := select init arm have ih theorem compiled","shard":"modules/f277ccb6e7cb9000.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.mean_le_bestArm_mean","label":"mean_le_bestArm_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.mean_le_bestArm_mean","description":"The mean of every arm is at most the mean of the model's selected best arm. This is the `FINITE-BANDIT-BESTARM-DOMINATES` model-invariant leaf. It is a semantic fact about the local finite-arm selector only; it does not prove gap nonnegativity or any expectation/concentration statement.","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html#decl-78ad630be0f0","parent":"module:BanditRLProof.FiniteBanditModelInvariants","order":4868,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteBanditModelInvariants"],["Source","BanditRLProof/FiniteBanditModelInvariants.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_le_bestArm_mean {K : Nat} (model : FiniteBanditModel K) (a : Fin K) : model.mean a <= model.mean model.bestArm","missing":[],"search":"mean_le_bestarm_mean banditrlproof.finitebanditmodel.mean_le_bestarm_mean the mean of every arm is at most the mean of the model's selected best arm. this is the `finite-bandit-bestarm-dominates` model-invariant leaf. it is a semantic fact about the local finite-arm selector only; it does not prove gap nonnegativity or any expectation/concentration statement. theorem compiled","shard":"modules/f277ccb6e7cb9000.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.gap_nonneg","label":"gap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.gap_nonneg","description":"Every local model gap is nonnegative. This is the `FINITE-BANDIT-GAP-NONNEG` model-invariant leaf. It consumes only `FiniteBanditModel.mean_le_bestArm_mean` and the local `gap` definition; it does not prove or use any expectation, filtration, or concentration statement.","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html#decl-dd65770230f9","parent":"module:BanditRLProof.FiniteBanditModelInvariants","order":4869,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteBanditModelInvariants"],["Source","BanditRLProof/FiniteBanditModelInvariants.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_nonneg {K : Nat} (model : FiniteBanditModel K) (a : Fin K) : (0 : Rat) <= model.gap a","missing":[],"search":"gap_nonneg banditrlproof.finitebanditmodel.gap_nonneg every local model gap is nonnegative. this is the `finite-bandit-gap-nonneg` model-invariant leaf. it consumes only `finitebanditmodel.mean_le_bestarm_mean` and the local `gap` definition; it does not prove or use any expectation, filtration, or concentration statement. theorem compiled","shard":"modules/f277ccb6e7cb9000.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.maxGap","label":"maxGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.maxGap","description":"The maximum local arm gap over the finite arm set. This is a deterministic finite-model constant only. It does not introduce probability, expectation, concentration, or algorithmic behavior.","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html#decl-d76d585d9c30","parent":"module:BanditRLProof.FiniteBanditModelInvariants","order":4870,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteBanditModelInvariants"],["Source","BanditRLProof/FiniteBanditModelInvariants.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def maxGap {K : Nat} (model : FiniteBanditModel K) : Rat","missing":[],"search":"maxgap banditrlproof.finitebanditmodel.maxgap the maximum local arm gap over the finite arm set. this is a deterministic finite-model constant only. it does not introduce probability, expectation, concentration, or algorithmic behavior. definition compiled","shard":"modules/f277ccb6e7cb9000.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.gap_le_maxGap","label":"gap_le_maxGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.gap_le_maxGap","description":"Every local model gap is bounded by `FiniteBanditModel.maxGap`. This is the finite max-gap adapter used by sharper ETC suffix bounds.","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html#decl-e123d23a55fe","parent":"module:BanditRLProof.FiniteBanditModelInvariants","order":4871,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteBanditModelInvariants"],["Source","BanditRLProof/FiniteBanditModelInvariants.lean:114"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_le_maxGap {K : Nat} (model : FiniteBanditModel K) (a : Fin K) : model.gap a <= model.maxGap","missing":[],"search":"gap_le_maxgap banditrlproof.finitebanditmodel.gap_le_maxgap every local model gap is bounded by `finitebanditmodel.maxgap`. this is the finite max-gap adapter used by sharper etc suffix bounds. theorem compiled","shard":"modules/f277ccb6e7cb9000.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.maxGap_nonneg","label":"maxGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.maxGap_nonneg","description":"The finite maximum gap is nonnegative.","url":"../modules/banditrlproof-finitebanditmodelinvariants/index.html#decl-da7a5825bb1c","parent":"module:BanditRLProof.FiniteBanditModelInvariants","order":4872,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteBanditModelInvariants"],["Source","BanditRLProof/FiniteBanditModelInvariants.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem maxGap_nonneg {K : Nat} (model : FiniteBanditModel K) : (0 : Rat) <= model.maxGap","missing":[],"search":"maxgap_nonneg banditrlproof.finitebanditmodel.maxgap_nonneg the finite maximum gap is nonnegative. theorem compiled","shard":"modules/f277ccb6e7cb9000.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteContextArmVarianceProxy","label":"finiteContextArmVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteContextArmVarianceProxy","description":"The largest variance proxy over a finite context-action family.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html#decl-c357d32eaf91","parent":"module:BanditRLProof.FiniteContextVarianceProxy","order":4873,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteContextVarianceProxy"],["Source","BanditRLProof/FiniteContextVarianceProxy.lean:14"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteContextArmVarianceProxy {Context : Type} [Fintype Context] {K : Nat} (varianceProxy : Context -> Fin K -> NNReal) : NNReal","missing":[],"search":"finitecontextarmvarianceproxy banditrlproof.concentration.finitecontextarmvarianceproxy the largest variance proxy over a finite context-action family. definition compiled","shard":"modules/6fb7932810277bae.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteContextArmVarianceProxy","label":"varianceProxy_le_finiteContextArmVarianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.varianceProxy_le_finiteContextArmVarianceProxy","description":"Every context-action proxy is bounded by the finite-family maximum.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html#decl-7c2d66b62e02","parent":"module:BanditRLProof.FiniteContextVarianceProxy","order":4874,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteContextVarianceProxy"],["Source","BanditRLProof/FiniteContextVarianceProxy.lean:20"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem varianceProxy_le_finiteContextArmVarianceProxy {Context : Type} [Fintype Context] {K : Nat} (varianceProxy : Context -> Fin K -> NNReal) (context : Context) (arm : Fin K) : varianceProxy context arm <= finiteContextArmVarianceProxy varianceProxy","missing":[],"search":"varianceproxy_le_finitecontextarmvarianceproxy banditrlproof.concentration.varianceproxy_le_finitecontextarmvarianceproxy every context-action proxy is bounded by the finite-family maximum. theorem compiled","shard":"modules/6fb7932810277bae.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteContextArmVarianceProxy_pos_of_exists","label":"finiteContextArmVarianceProxy_pos_of_exists","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteContextArmVarianceProxy_pos_of_exists","description":"A positive member makes the finite context-action maximum positive.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html#decl-2589df4db83a","parent":"module:BanditRLProof.FiniteContextVarianceProxy","order":4875,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteContextVarianceProxy"],["Source","BanditRLProof/FiniteContextVarianceProxy.lean:37"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteContextArmVarianceProxy_pos_of_exists {Context : Type} [Fintype Context] {K : Nat} (varianceProxy : Context -> Fin K -> NNReal) (hpos : exists context arm, 0 < ((varianceProxy context arm : NNReal) : Real)) : 0 < ((finiteContextArmVarianceProxy varianceProxy : NNReal) : Real)","missing":[],"search":"finitecontextarmvarianceproxy_pos_of_exists banditrlproof.concentration.finitecontextarmvarianceproxy_pos_of_exists a positive member makes the finite context-action maximum positive. theorem compiled","shard":"modules/6fb7932810277bae.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteContextArmPositiveVarianceProxy","label":"finiteContextArmPositiveVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteContextArmPositiveVarianceProxy","description":"The finite context-action maximum padded by one. This gives algorithms that require a strictly positive tuning parameter a uniform proxy even when every genuine variance proxy is zero.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html#decl-fc73b5d9bb47","parent":"module:BanditRLProof.FiniteContextVarianceProxy","order":4876,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteContextVarianceProxy"],["Source","BanditRLProof/FiniteContextVarianceProxy.lean:53"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteContextArmPositiveVarianceProxy {Context : Type} [Fintype Context] {K : Nat} (varianceProxy : Context -> Fin K -> NNReal) : NNReal","missing":[],"search":"finitecontextarmpositivevarianceproxy banditrlproof.concentration.finitecontextarmpositivevarianceproxy the finite context-action maximum padded by one. this gives algorithms that require a strictly positive tuning parameter a uniform proxy even when every genuine variance proxy is zero. definition compiled","shard":"modules/6fb7932810277bae.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteContextArmPositiveVarianceProxy","label":"varianceProxy_le_finiteContextArmPositiveVarianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.varianceProxy_le_finiteContextArmPositiveVarianceProxy","description":"Every genuine proxy is bounded by the positive padded proxy.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html#decl-9341e76ece89","parent":"module:BanditRLProof.FiniteContextVarianceProxy","order":4877,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteContextVarianceProxy"],["Source","BanditRLProof/FiniteContextVarianceProxy.lean:59"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem varianceProxy_le_finiteContextArmPositiveVarianceProxy {Context : Type} [Fintype Context] {K : Nat} (varianceProxy : Context -> Fin K -> NNReal) (context : Context) (arm : Fin K) : varianceProxy context arm <= finiteContextArmPositiveVarianceProxy varianceProxy","missing":[],"search":"varianceproxy_le_finitecontextarmpositivevarianceproxy banditrlproof.concentration.varianceproxy_le_finitecontextarmpositivevarianceproxy every genuine proxy is bounded by the positive padded proxy. theorem compiled","shard":"modules/6fb7932810277bae.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.finiteContextArmPositiveVarianceProxy_pos","label":"finiteContextArmPositiveVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.finiteContextArmPositiveVarianceProxy_pos","description":"The padded finite context-action proxy is always strictly positive.","url":"../modules/banditrlproof-finitecontextvarianceproxy/index.html#decl-b6636311beed","parent":"module:BanditRLProof.FiniteContextVarianceProxy","order":4878,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteContextVarianceProxy"],["Source","BanditRLProof/FiniteContextVarianceProxy.lean:70"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteContextArmPositiveVarianceProxy_pos {Context : Type} [Fintype Context] {K : Nat} (varianceProxy : Context -> Fin K -> NNReal) : 0 < ((finiteContextArmPositiveVarianceProxy varianceProxy : NNReal) : Real)","missing":[],"search":"finitecontextarmpositivevarianceproxy_pos banditrlproof.concentration.finitecontextarmpositivevarianceproxy_pos the padded finite context-action proxy is always strictly positive. theorem compiled","shard":"modules/6fb7932810277bae.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteGapLayerCake.sum_le_cutoff_integral","label":"sum_le_cutoff_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake.sum_le_cutoff_integral","description":"theorem sum_le_cutoff_integral {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) (hab : a≤b) (hd : ∀t∈s, d t≤b) (ell : ℝ → ℝ) (hi : IntervalIntegrable ell volume a b) (hc : ∀x∈Set.Icc a b, ((s.filter (fun t => x≤d t)).card:ℝ)≤ell x+1) : ∑t∈s, d t ≤ a*(s.card:ℝ)+(∫x in a..b, ell x)+(b-a)","url":"../modules/banditrlproof-finitegapcutoff/index.html#decl-e03d8a6157ce","parent":"module:BanditRLProof.FiniteGapCutoff","order":4879,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteGapCutoff"],["Source","BanditRLProof/FiniteGapCutoff.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_le_cutoff_integral {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) (hab : a≤b) (hd : ∀t∈s, d t≤b) (ell : ℝ → ℝ) (hi : IntervalIntegrable ell volume a b) (hc : ∀x∈Set.Icc a b, ((s.filter (fun t => x≤d t)).card:ℝ)≤ell x+1) : ∑t∈s, d t ≤ a*(s.card:ℝ)+(∫x in a..b, ell x)+(b-a)","missing":[],"search":"sum_le_cutoff_integral banditrlproof.finitegaplayercake.sum_le_cutoff_integral theorem sum_le_cutoff_integral {ι : type*} (s : finset ι) (d : ι → ℝ) (a b : ℝ) (hab : a≤b) (hd : ∀t∈s, d t≤b) (ell : ℝ → ℝ) (hi : intervalintegrable ell volume a b) (hc : ∀x∈set.icc a b, ((s.filter (fun t => x≤d t)).card:ℝ)≤ell x+1) : ∑t∈s, d t ≤ a*(s.card:ℝ)+(∫x in a..b, ell x)+(b-a) theorem compiled","shard":"modules/d903dd5bd78971a1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteGapLayerCake.intervalIntegrable_step","label":"intervalIntegrable_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake.intervalIntegrable_step","description":"theorem intervalIntegrable_step (d a b : ℝ) : IntervalIntegrable (fun x : ℝ => if x≤d then (1:ℝ) else 0) volume a b","url":"../modules/banditrlproof-finitegaplayercake/index.html#decl-a82431ce550b","parent":"module:BanditRLProof.FiniteGapLayerCake","order":4880,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteGapLayerCake"],["Source","BanditRLProof/FiniteGapLayerCake.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem intervalIntegrable_step (d a b : ℝ) : IntervalIntegrable (fun x : ℝ => if x≤d then (1:ℝ) else 0) volume a b","missing":[],"search":"intervalintegrable_step banditrlproof.finitegaplayercake.intervalintegrable_step theorem intervalintegrable_step (d a b : ℝ) : intervalintegrable (fun x : ℝ => if x≤d then (1:ℝ) else 0) volume a b theorem compiled","shard":"modules/147f959a8571008f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteGapLayerCake.integral_step","label":"integral_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake.integral_step","description":"theorem integral_step {a b d : ℝ} (hd : d∈Set.Icc a b) : (∫x in a..b, if x≤d then (1:ℝ) else 0)=d-a","url":"../modules/banditrlproof-finitegaplayercake/index.html#decl-27bedd573897","parent":"module:BanditRLProof.FiniteGapLayerCake","order":4881,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteGapLayerCake"],["Source","BanditRLProof/FiniteGapLayerCake.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_step {a b d : ℝ} (hd : d∈Set.Icc a b) : (∫x in a..b, if x≤d then (1:ℝ) else 0)=d-a","missing":[],"search":"integral_step banditrlproof.finitegaplayercake.integral_step theorem integral_step {a b d : ℝ} (hd : d∈set.icc a b) : (∫x in a..b, if x≤d then (1:ℝ) else 0)=d-a theorem compiled","shard":"modules/147f959a8571008f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteGapLayerCake.intervalIntegrable_card","label":"intervalIntegrable_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake.intervalIntegrable_card","description":"theorem intervalIntegrable_card {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) : IntervalIntegrable (fun x => ((s.filter (fun t => x≤d t)).card:ℝ)) volume a b","url":"../modules/banditrlproof-finitegaplayercake/index.html#decl-931ee960766d","parent":"module:BanditRLProof.FiniteGapLayerCake","order":4882,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteGapLayerCake"],["Source","BanditRLProof/FiniteGapLayerCake.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem intervalIntegrable_card {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) : IntervalIntegrable (fun x => ((s.filter (fun t => x≤d t)).card:ℝ)) volume a b","missing":[],"search":"intervalintegrable_card banditrlproof.finitegaplayercake.intervalintegrable_card theorem intervalintegrable_card {ι : type*} (s : finset ι) (d : ι → ℝ) (a b : ℝ) : intervalintegrable (fun x => ((s.filter (fun t => x≤d t)).card:ℝ)) volume a b theorem compiled","shard":"modules/147f959a8571008f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteGapLayerCake.sum_eq_layerCake","label":"sum_eq_layerCake","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake.sum_eq_layerCake","description":"theorem sum_eq_layerCake {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) (hd : ∀t∈s, d t∈Set.Icc a b) : ∑t∈s, d t = a*(s.card:ℝ)+(∫x in a..b, ((s.filter (fun t => x≤d t)).card:ℝ))","url":"../modules/banditrlproof-finitegaplayercake/index.html#decl-5983f95d007a","parent":"module:BanditRLProof.FiniteGapLayerCake","order":4883,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteGapLayerCake"],["Source","BanditRLProof/FiniteGapLayerCake.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_eq_layerCake {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) (hd : ∀t∈s, d t∈Set.Icc a b) : ∑t∈s, d t = a*(s.card:ℝ)+(∫x in a..b, ((s.filter (fun t => x≤d t)).card:ℝ))","missing":[],"search":"sum_eq_layercake banditrlproof.finitegaplayercake.sum_eq_layercake theorem sum_eq_layercake {ι : type*} (s : finset ι) (d : ι → ℝ) (a b : ℝ) (hd : ∀t∈s, d t∈set.icc a b) : ∑t∈s, d t = a*(s.card:ℝ)+(∫x in a..b, ((s.filter (fun t => x≤d t)).card:ℝ)) theorem compiled","shard":"modules/147f959a8571008f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteGapLayerCake.sum_le_refined_integral","label":"sum_le_refined_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteGapLayerCake.sum_le_refined_integral","description":"theorem sum_le_refined_integral {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) (ha : 0≤a) (hab : a≤b) (hd : ∀t∈s, d t∈Set.Icc a b) (ell : ℝ → ℝ) (hi : IntervalIntegrable ell volume a b) (hc : ∀x∈Set.Icc a b, ((s.filter (fun t => x≤d t)).card:ℝ)≤ell x+1) : ∑t∈s, d t ≤ a*ell a+(∫x in a..b, ell x)+b","url":"../modules/banditrlproof-finitegaplayercake/index.html#decl-914139c6e57e","parent":"module:BanditRLProof.FiniteGapLayerCake","order":4884,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteGapLayerCake"],["Source","BanditRLProof/FiniteGapLayerCake.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_le_refined_integral {ι : Type*} (s : Finset ι) (d : ι → ℝ) (a b : ℝ) (ha : 0≤a) (hab : a≤b) (hd : ∀t∈s, d t∈Set.Icc a b) (ell : ℝ → ℝ) (hi : IntervalIntegrable ell volume a b) (hc : ∀x∈Set.Icc a b, ((s.filter (fun t => x≤d t)).card:ℝ)≤ell x+1) : ∑t∈s, d t ≤ a*ell a+(∫x in a..b, ell x)+b","missing":[],"search":"sum_le_refined_integral banditrlproof.finitegaplayercake.sum_le_refined_integral theorem sum_le_refined_integral {ι : type*} (s : finset ι) (d : ι → ℝ) (a b : ℝ) (ha : 0≤a) (hab : a≤b) (hd : ∀t∈s, d t∈set.icc a b) (ell : ℝ → ℝ) (hi : intervalintegrable ell volume a b) (hc : ∀x∈set.icc a b, ((s.filter (fun t => x≤d t)).card:ℝ)≤ell x+1) : ∑t∈s, d t ≤ a*ell a+(∫x in a..b, ell x)+b theorem compiled","shard":"modules/147f959a8571008f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.score_le_foldl_select","label":"score_le_foldl_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.score_le_foldl_select","description":"private theorem score_le_foldl_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha; cases…","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-1712b493f409","parent":"module:BanditRLProof.FiniteRealArgmax","order":4885,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem score_le_foldl_select {K : Nat} (scores : Fin K -> Real) (init : Fin K) : forall l : List (Fin K), (forall a : Fin K, List.Mem a l -> scores a <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : Fin K => if scores best < scores arm then arm else best) init) | [] => by exact And.intro (by intro _ ha; cases ha) (by simp) | arm :: rest => by let select := fun best arm : Fin K => if scores best < scores arm then arm else best let next := select init arm have ih","missing":[],"search":"score_le_foldl_select banditrlproof.finiterealargmax.score_le_foldl_select private theorem score_le_foldl_select {k : nat} (scores : fin k -> real) (init : fin k) : forall l : list (fin k), (forall a : fin k, list.mem a l -> scores a <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init)) /\\ scores init <= scores (l.foldl (fun best arm : fin k => if scores best < scores arm then arm else best) init) | [] => by exact and.intro (by intro _ ha; cases ha) (by simp) | arm :: rest => by let select := fun best arm : fin k => if scores best < scores arm then arm else best let next := select init arm have ih theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.chooseFin","label":"chooseFin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.chooseFin","description":"private noncomputable def chooseFin {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Fin K","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-99013c2e0435","parent":"module:BanditRLProof.FiniteRealArgmax","order":4886,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private noncomputable def chooseFin {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) : Fin K","missing":[],"search":"choosefin banditrlproof.finiterealargmax.choosefin private noncomputable def choosefin {k : nat} (hk : 0 < k) (scores : fin k -> real) : fin k definition compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.score_le_chooseFin","label":"score_le_chooseFin","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.score_le_chooseFin","description":"private theorem score_le_chooseFin {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) : scores a <= scores (chooseFin hK scores)","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-55aef7af4a88","parent":"module:BanditRLProof.FiniteRealArgmax","order":4887,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem score_le_chooseFin {K : Nat} (hK : 0 < K) (scores : Fin K -> Real) (a : Fin K) : scores a <= scores (chooseFin hK scores)","missing":[],"search":"score_le_choosefin banditrlproof.finiterealargmax.score_le_choosefin private theorem score_le_choosefin {k : nat} (hk : 0 < k) (scores : fin k -> real) (a : fin k) : scores a <= scores (choosefin hk scores) theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.measurable_selected_score","label":"measurable_selected_score","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.measurable_selected_score","description":"private theorem measurable_selected_score {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : Measurable (fun omega : Omega => scores omega (best omega))","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-4664171b0d93","parent":"module:BanditRLProof.FiniteRealArgmax","order":4888,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem measurable_selected_score {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : Measurable (fun omega : Omega => scores omega (best omega))","missing":[],"search":"measurable_selected_score banditrlproof.finiterealargmax.measurable_selected_score private theorem measurable_selected_score {omega : type u} {k : nat} [measurablespace omega] (scores : omega -> fin k -> real) (hscores : forall a : fin k, measurable (fun omega : omega => scores omega a)) (best : omega -> fin k) (hbest : measurable best) : measurable (fun omega : omega => scores omega (best omega)) theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.measurable_foldl_select","label":"measurable_foldl_select","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.measurable_foldl_select","description":"private theorem measurable_foldl_select {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : forall l : List (Fin K), Measurable (fun omega : Omega => l.foldl (fun best arm : Fin K => if scores omega best < scores omega arm then arm else best) (best omega)…","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-a278eaccc4e6","parent":"module:BanditRLProof.FiniteRealArgmax","order":4889,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem measurable_foldl_select {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) (best : Omega -> Fin K) (hbest : Measurable best) : forall l : List (Fin K), Measurable (fun omega : Omega => l.foldl (fun best arm : Fin K => if scores omega best < scores omega arm then arm else best) (best omega)) | [] => by simpa using hbest | arm :: rest => by have hselected : Measurable (fun omega : Omega => scores omega (best omega))","missing":[],"search":"measurable_foldl_select banditrlproof.finiterealargmax.measurable_foldl_select private theorem measurable_foldl_select {omega : type u} {k : nat} [measurablespace omega] (scores : omega -> fin k -> real) (hscores : forall a : fin k, measurable (fun omega : omega => scores omega a)) (best : omega -> fin k) (hbest : measurable best) : forall l : list (fin k), measurable (fun omega : omega => l.foldl (fun best arm : fin k => if scores omega best < scores omega arm then arm else best) (best omega)) | [] => by simpa using hbest | arm :: rest => by have hselected : measurable (fun omega : omega => scores omega (best omega)) theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.measurable_chooseFin_of_forall_measurable","label":"measurable_chooseFin_of_forall_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.measurable_chooseFin_of_forall_measurable","description":"private theorem measurable_chooseFin_of_forall_measurable {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) : Measurable (fun omega : Omega => chooseFin hK (scores omega))","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-0655d5c2fcd0","parent":"module:BanditRLProof.FiniteRealArgmax","order":4890,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem measurable_chooseFin_of_forall_measurable {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (hK : 0 < K) (scores : Omega -> Fin K -> Real) (hscores : forall a : Fin K, Measurable (fun omega : Omega => scores omega a)) : Measurable (fun omega : Omega => chooseFin hK (scores omega))","missing":[],"search":"measurable_choosefin_of_forall_measurable banditrlproof.finiterealargmax.measurable_choosefin_of_forall_measurable private theorem measurable_choosefin_of_forall_measurable {omega : type u} {k : nat} [measurablespace omega] (hk : 0 < k) (scores : omega -> fin k -> real) (hscores : forall a : fin k, measurable (fun omega : omega => scores omega a)) : measurable (fun omega : omega => choosefin hk (scores omega)) theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.choose","label":"choose","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.choose","description":"A fixed-enumeration maximizer on a nonempty finite type.","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-c7d94c389795","parent":"module:BanditRLProof.FiniteRealArgmax","order":4891,"meta":[["Kind","definition"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def choose {alpha : Type v} [Fintype alpha] [Nonempty alpha] (scores : alpha -> Real) : alpha","missing":[],"search":"choose banditrlproof.finiterealargmax.choose a fixed-enumeration maximizer on a nonempty finite type. definition compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.score_le_choose","label":"score_le_choose","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.score_le_choose","description":"Every score is bounded by the score selected by `choose`.","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-aa848020b018","parent":"module:BanditRLProof.FiniteRealArgmax","order":4892,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem score_le_choose {alpha : Type v} [Fintype alpha] [Nonempty alpha] (scores : alpha -> Real) (a : alpha) : scores a <= scores (choose scores)","missing":[],"search":"score_le_choose banditrlproof.finiterealargmax.score_le_choose every score is bounded by the score selected by `choose`. theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.measurable_choose_of_forall_measurable","label":"measurable_choose_of_forall_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.measurable_choose_of_forall_measurable","description":"The fixed-enumeration maximizer is measurable in external parameters.","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-7fc37f73110f","parent":"module:BanditRLProof.FiniteRealArgmax","order":4893,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:161"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_choose_of_forall_measurable {Omega : Type u} {alpha : Type v} [MeasurableSpace Omega] [MeasurableSpace alpha] [Fintype alpha] [Nonempty alpha] (scores : Omega -> alpha -> Real) (hscores : forall a : alpha, Measurable (fun omega : Omega => scores omega a)) : Measurable (fun omega : Omega => choose (scores omega))","missing":[],"search":"measurable_choose_of_forall_measurable banditrlproof.finiterealargmax.measurable_choose_of_forall_measurable the fixed-enumeration maximizer is measurable in external parameters. theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteRealArgmax.measurable_selected_score_of_forall_measurable","label":"measurable_selected_score_of_forall_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteRealArgmax.measurable_selected_score_of_forall_measurable","description":"Evaluation at a measurable finite-valued selector preserves measurability.","url":"../modules/banditrlproof-finiterealargmax/index.html#decl-9b312a30bc0d","parent":"module:BanditRLProof.FiniteRealArgmax","order":4894,"meta":[["Kind","theorem"],["Module","BanditRLProof.FiniteRealArgmax"],["Source","BanditRLProof/FiniteRealArgmax.lean:180"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_selected_score_of_forall_measurable {Omega : Type u} {alpha : Type v} [MeasurableSpace Omega] [MeasurableSpace alpha] [Fintype alpha] [MeasurableSingletonClass alpha] (scores : Omega -> alpha -> Real) (hscores : forall a : alpha, Measurable (fun omega : Omega => scores omega a)) (selected : Omega -> alpha) (hselected : Measurable selected) : Measurable (fun omega : Omega => scores omega (selected omega))","missing":[],"search":"measurable_selected_score_of_forall_measurable banditrlproof.finiterealargmax.measurable_selected_score_of_forall_measurable evaluation at a measurable finite-valued selector preserves measurability. theorem compiled","shard":"modules/aa256c808ed3b441.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.Arm","label":"Arm","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.Arm","description":"abbrev Arm","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-22e490a2abf2","parent":"module:BanditRLProof.HOOCantorModel","order":4895,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev Arm","missing":[],"search":"arm banditrlproof.hoo.cantormodel.arm abbrev arm abbreviation compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.center","label":"center","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.center","description":"def center (v : Node) : Arm","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-3faae78c4807","parent":"module:BanditRLProof.HOOCantorModel","order":4896,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def center (v : Node) : Arm","missing":[],"search":"center banditrlproof.hoo.cantormodel.center def center (v : node) : arm definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.region","label":"region","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.region","description":"def region (v : Node) : Set Arm","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-dfefbfea0c26","parent":"module:BanditRLProof.HOOCantorModel","order":4897,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def region (v : Node) : Set Arm","missing":[],"search":"region banditrlproof.hoo.cantormodel.region def region (v : node) : set arm definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.ell","label":"ell","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.ell","description":"noncomputable def ell (x y : Arm) : ℝ","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-00cf9ed62842","parent":"module:BanditRLProof.HOOCantorModel","order":4898,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def ell (x y : Arm) : ℝ","missing":[],"search":"ell banditrlproof.hoo.cantormodel.ell noncomputable def ell (x y : arm) : ℝ definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.center_child_of_lt","label":"center_child_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.center_child_of_lt","description":"theorem center_child_of_lt (v : Node) (b : Bool) (i : ℕ) (hi : i < v.length) : center (child v b) i = center v i","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-4cdc65742d1d","parent":"module:BanditRLProof.HOOCantorModel","order":4899,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem center_child_of_lt (v : Node) (b : Bool) (i : ℕ) (hi : i < v.length) : center (child v b) i = center v i","missing":[],"search":"center_child_of_lt banditrlproof.hoo.cantormodel.center_child_of_lt theorem center_child_of_lt (v : node) (b : bool) (i : ℕ) (hi : i < v.length) : center (child v b) i = center v i theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.center_child_last","label":"center_child_last","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.center_child_last","description":"@[simp] theorem center_child_last (v : Node) (b : Bool) : center (child v b) v.length = b","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-d2f40fb1be60","parent":"module:BanditRLProof.HOOCantorModel","order":4900,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem center_child_last (v : Node) (b : Bool) : center (child v b) v.length = b","missing":[],"search":"center_child_last banditrlproof.hoo.cantormodel.center_child_last @[simp] theorem center_child_last (v : node) (b : bool) : center (child v b) v.length = b theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.mem_child_iff","label":"mem_child_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.mem_child_iff","description":"theorem mem_child_iff (v : Node) (b : Bool) (x : Arm) : x ∈ region (child v b) ↔ x ∈ region v ∧ x v.length = b","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-ddd18f0e0ec7","parent":"module:BanditRLProof.HOOCantorModel","order":4901,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_child_iff (v : Node) (b : Bool) (x : Arm) : x ∈ region (child v b) ↔ x ∈ region v ∧ x v.length = b","missing":[],"search":"mem_child_iff banditrlproof.hoo.cantormodel.mem_child_iff theorem mem_child_iff (v : node) (b : bool) (x : arm) : x ∈ region (child v b) ↔ x ∈ region v ∧ x v.length = b theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.region_children","label":"region_children","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.region_children","description":"theorem region_children (v : Node) : region v = region (child v false) ∪ region (child v true)","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-e2b2befe614a","parent":"module:BanditRLProof.HOOCantorModel","order":4902,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_children (v : Node) : region v = region (child v false) ∪ region (child v true)","missing":[],"search":"region_children banditrlproof.hoo.cantormodel.region_children theorem region_children (v : node) : region v = region (child v false) ∪ region (child v true) theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.region_measurable","label":"region_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.region_measurable","description":"theorem region_measurable (v : Node) : MeasurableSet (region v)","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-25640aa8031e","parent":"module:BanditRLProof.HOOCantorModel","order":4903,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_measurable (v : Node) : MeasurableSet (region v)","missing":[],"search":"region_measurable banditrlproof.hoo.cantormodel.region_measurable theorem region_measurable (v : node) : measurableset (region v) theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.region_diameter","label":"region_diameter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.region_diameter","description":"theorem region_diameter (v : Node) (x : Arm) (hx : x ∈ region v) (y : Arm) (hy : y ∈ region v) : ell x y ≤ (1/2:ℝ)^v.length","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-292b9c8c98ae","parent":"module:BanditRLProof.HOOCantorModel","order":4904,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_diameter (v : Node) (x : Arm) (hx : x ∈ region v) (y : Arm) (hy : y ∈ region v) : ell x y ≤ (1/2:ℝ)^v.length","missing":[],"search":"region_diameter banditrlproof.hoo.cantormodel.region_diameter theorem region_diameter (v : node) (x : arm) (hx : x ∈ region v) (y : arm) (hy : y ∈ region v) : ell x y ≤ (1/2:ℝ)^v.length theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.ball_subset_region","label":"ball_subset_region","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.ball_subset_region","description":"theorem ball_subset_region (v : Node) : {y : Arm | ell (center v) y < (1/2:ℝ)^v.length} ⊆ region v","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-bfee2a297958","parent":"module:BanditRLProof.HOOCantorModel","order":4905,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ball_subset_region (v : Node) : {y : Arm | ell (center v) y < (1/2:ℝ)^v.length} ⊆ region v","missing":[],"search":"ball_subset_region banditrlproof.hoo.cantormodel.ball_subset_region theorem ball_subset_region (v : node) : {y : arm | ell (center v) y < (1/2:ℝ)^v.length} ⊆ region v theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.region_disjoint","label":"region_disjoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.region_disjoint","description":"theorem region_disjoint {v w : Node} (hl : v.length=w.length) (hne : v≠w) : Disjoint (region v) (region w)","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-994e170c404a","parent":"module:BanditRLProof.HOOCantorModel","order":4906,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_disjoint {v w : Node} (hl : v.length=w.length) (hne : v≠w) : Disjoint (region v) (region w)","missing":[],"search":"region_disjoint banditrlproof.hoo.cantormodel.region_disjoint theorem region_disjoint {v w : node} (hl : v.length=w.length) (hne : v≠w) : disjoint (region v) (region w) theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.covering","label":"covering","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.covering","description":"noncomputable def covering : RegularCovering Arm where","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-f996c2e773ec","parent":"module:BanditRLProof.HOOCantorModel","order":4907,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def covering : RegularCovering Arm where","missing":[],"search":"covering banditrlproof.hoo.cantormodel.covering noncomputable def covering : regularcovering arm where definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.mean","label":"mean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.mean","description":"noncomputable def mean (x : Arm) : ℝ","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-074d63215a58","parent":"module:BanditRLProof.HOOCantorModel","order":4908,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def mean (x : Arm) : ℝ","missing":[],"search":"mean banditrlproof.hoo.cantormodel.mean noncomputable def mean (x : arm) : ℝ definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.mean_le_best","label":"mean_le_best","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.mean_le_best","description":"theorem mean_le_best (x : Arm) : mean x ≤ 1/2","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-13157151f1c3","parent":"module:BanditRLProof.HOOCantorModel","order":4909,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mean_le_best (x : Arm) : mean x ≤ 1/2","missing":[],"search":"mean_le_best banditrlproof.hoo.cantormodel.mean_le_best theorem mean_le_best (x : arm) : mean x ≤ 1/2 theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.first_coordinate_distance","label":"first_coordinate_distance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.first_coordinate_distance","description":"theorem first_coordinate_distance (x y : Arm) (h : x 0 ≠ y 0) : ell x y = 1","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-d1d9e2eca922","parent":"module:BanditRLProof.HOOCantorModel","order":4910,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem first_coordinate_distance (x y : Arm) (h : x 0 ≠ y 0) : ell x y = 1","missing":[],"search":"first_coordinate_distance banditrlproof.hoo.cantormodel.first_coordinate_distance theorem first_coordinate_distance (x y : arm) (h : x 0 ≠ y 0) : ell x y = 1 theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.weaklyLipschitz","label":"weaklyLipschitz","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.weaklyLipschitz","description":"theorem weaklyLipschitz : WeaklyLipschitz mean ell (1/2)","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-878dbae698cf","parent":"module:BanditRLProof.HOOCantorModel","order":4911,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weaklyLipschitz : WeaklyLipschitz mean ell (1/2)","missing":[],"search":"weaklylipschitz banditrlproof.hoo.cantormodel.weaklylipschitz theorem weaklylipschitz : weaklylipschitz mean ell (1/2) theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.bitReward","label":"bitReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.bitReward","description":"noncomputable def bitReward (b : Bool) : Measure ℝ","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-e27259cb3392","parent":"module:BanditRLProof.HOOCantorModel","order":4912,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bitReward (b : Bool) : Measure ℝ","missing":[],"search":"bitreward banditrlproof.hoo.cantormodel.bitreward noncomputable def bitreward (b : bool) : measure ℝ definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.bitKernel","label":"bitKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.bitKernel","description":"noncomputable def bitKernel : Kernel Bool ℝ","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-1a51379d183f","parent":"module:BanditRLProof.HOOCantorModel","order":4913,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def bitKernel : Kernel Bool ℝ","missing":[],"search":"bitkernel banditrlproof.hoo.cantormodel.bitkernel noncomputable def bitkernel : kernel bool ℝ definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.law","label":"law","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.law","description":"noncomputable def law : Kernel Arm ℝ","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-1ced11d05cd9","parent":"module:BanditRLProof.HOOCantorModel","order":4914,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:134"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def law : Kernel Arm ℝ","missing":[],"search":"law banditrlproof.hoo.cantormodel.law noncomputable def law : kernel arm ℝ definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.law_bounded","label":"law_bounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.law_bounded","description":"theorem law_bounded (x : Arm) : ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-be8d89878feb","parent":"module:BanditRLProof.HOOCantorModel","order":4915,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:137"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem law_bounded (x : Arm) : ∀ᵐ y ∂law x, y ∈ Set.Icc (0 : ℝ) 1","missing":[],"search":"law_bounded banditrlproof.hoo.cantormodel.law_bounded theorem law_bounded (x : arm) : ∀ᵐ y ∂law x, y ∈ set.icc (0 : ℝ) 1 theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.law_zero_mass","label":"law_zero_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.law_zero_mass","description":"theorem law_zero_mass (x : Arm) : law x {0} = (1/2 : ENNReal)","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-72f7b8181f25","parent":"module:BanditRLProof.HOOCantorModel","order":4916,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:144"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem law_zero_mass (x : Arm) : law x {0} = (1/2 : ENNReal)","missing":[],"search":"law_zero_mass banditrlproof.hoo.cantormodel.law_zero_mass theorem law_zero_mass (x : arm) : law x {0} = (1/2 : ennreal) theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.law_not_dirac","label":"law_not_dirac","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.law_not_dirac","description":"theorem law_not_dirac (x : Arm) (r : ℝ) : law x ≠ Measure.dirac r","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-92b391f82c7e","parent":"module:BanditRLProof.HOOCantorModel","order":4917,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem law_not_dirac (x : Arm) (r : ℝ) : law x ≠ Measure.dirac r","missing":[],"search":"law_not_dirac banditrlproof.hoo.cantormodel.law_not_dirac theorem law_not_dirac (x : arm) (r : ℝ) : law x ≠ measure.dirac r theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.law_mean","label":"law_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.law_mean","description":"theorem law_mean (x : Arm) : (∫ y, y ∂law x) = mean x","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-2c56f139673c","parent":"module:BanditRLProof.HOOCantorModel","order":4918,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem law_mean (x : Arm) : (∫ y, y ∂law x) = mean x","missing":[],"search":"law_mean banditrlproof.hoo.cantormodel.law_mean theorem law_mean (x : arm) : (∫ y, y ∂law x) = mean x theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.global_sup","label":"global_sup","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.global_sup","description":"theorem global_sup : regionSup mean Set.univ = 1/2","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-7c6e8e0345cf","parent":"module:BanditRLProof.HOOCantorModel","order":4919,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem global_sup : regionSup mean Set.univ = 1/2","missing":[],"search":"global_sup banditrlproof.hoo.cantormodel.global_sup theorem global_sup : regionsup mean set.univ = 1/2 theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.poorNode","label":"poorNode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.poorNode","description":"def poorNode : Node","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-4f7ed1b89c07","parent":"module:BanditRLProof.HOOCantorModel","order":4920,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def poorNode : Node","missing":[],"search":"poornode banditrlproof.hoo.cantormodel.poornode def poornode : node definition compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.poor_region_mean","label":"poor_region_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.poor_region_mean","description":"theorem poor_region_mean (x : Arm) (hx : x ∈ region poorNode) : mean x = 1/4","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-0da16ccd6665","parent":"module:BanditRLProof.HOOCantorModel","order":4921,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:176"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem poor_region_mean (x : Arm) (hx : x ∈ region poorNode) : mean x = 1/4","missing":[],"search":"poor_region_mean banditrlproof.hoo.cantormodel.poor_region_mean theorem poor_region_mean (x : arm) (hx : x ∈ region poornode) : mean x = 1/4 theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.poor_sup","label":"poor_sup","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.poor_sup","description":"theorem poor_sup : regionSup mean (region poorNode) = 1/4","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-78839f11dea5","parent":"module:BanditRLProof.HOOCantorModel","order":4922,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:181"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem poor_sup : regionSup mean (region poorNode) = 1/4","missing":[],"search":"poor_sup banditrlproof.hoo.cantormodel.poor_sup theorem poor_sup : regionsup mean (region poornode) = 1/4 theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.expected_poor_visits","label":"expected_poor_visits","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.expected_poor_visits","description":"A concrete nontrivial poor region in the infinite-arm model consumes the complete algorithm-to-expected-visits chain.","url":"../modules/banditrlproof-hoocantormodel/index.html#decl-5233c2d4352b","parent":"module:BanditRLProof.HOOCantorModel","order":4923,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorModel"],["Source","BanditRLProof/HOOCantorModel.lean:195"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expected_poor_visits (N : ℕ) : (∫ Y, (visits (history 1 (1/2) Y N) poorNode : ℝ) ∂trajectory 1 (1/2) (covering.toCovering.nodeLaw law)) ≤ 512 * Real.log (max (N:ℝ) 2) + 4","missing":[],"search":"expected_poor_visits banditrlproof.hoo.cantormodel.expected_poor_visits a concrete nontrivial poor region in the infinite-arm model consumes the complete algorithm-to-expected-visits chain. theorem compiled","shard":"modules/5ed75fc881ffdeb7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.packing_le_two_div","label":"packing_le_two_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.packing_le_two_div","description":"theorem packing_le_two_div (A : Set Arm) {ε : ℝ} (hε : 0<ε) (hε1 : ε≤1) : (covering.packingNumber A ε : ℝ) ≤ 2/ε","url":"../modules/banditrlproof-hoocantorrate/index.html#decl-d28845a98ac9","parent":"module:BanditRLProof.HOOCantorRate","order":4924,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorRate"],["Source","BanditRLProof/HOOCantorRate.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem packing_le_two_div (A : Set Arm) {ε : ℝ} (hε : 0<ε) (hε1 : ε≤1) : (covering.packingNumber A ε : ℝ) ≤ 2/ε","missing":[],"search":"packing_le_two_div banditrlproof.hoo.cantormodel.packing_le_two_div theorem packing_le_two_div (a : set arm) {ε : ℝ} (hε : 0<ε) (hε1 : ε≤1) : (covering.packingnumber a ε : ℝ) ≤ 2/ε theorem compiled","shard":"modules/74a1f773f2dc1a6f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.dimension_le_two","label":"dimension_le_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.dimension_le_two","description":"Conservative upper certificate, sufficient for a nonvacuous full-rate canary. It follows from ambient packing, not a postulated dimension value.","url":"../modules/banditrlproof-hoocantorrate/index.html#decl-eadb525fb57d","parent":"module:BanditRLProof.HOOCantorRate","order":4925,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorRate"],["Source","BanditRLProof/HOOCantorRate.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dimension_le_two (c : ℝ) : covering.nearOptimalityDimension mean (1/2) c ≤ (2:EReal)","missing":[],"search":"dimension_le_two banditrlproof.hoo.cantormodel.dimension_le_two conservative upper certificate, sufficient for a nonvacuous full-rate canary. it follows from ambient packing, not a postulated dimension value. theorem compiled","shard":"modules/74a1f773f2dc1a6f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.CantorModel.expected_actual_rate","label":"expected_actual_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.CantorModel.expected_actual_rate","description":"A concrete full-rate consequence with d'=3>dimension. The general theorem retains every d' strictly above the actual dimension; this model certificate is deliberately conservative and does not claim the sharp model exponent.","url":"../modules/banditrlproof-hoocantorrate/index.html#decl-ae185a03b39b","parent":"module:BanditRLProof.HOOCantorRate","order":4926,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOCantorRate"],["Source","BanditRLProof/HOOCantorRate.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"theorem expected_actual_rate : ∃ γ : ℝ, 0<γ ∧ ∀ N : ℕ, 1≤N → (∫ Y, (∑ n ∈ Finset.range N, ((1/2:ℝ)-Y n)) ∂trajectory 1 (1/2) (covering.toCovering.nodeLaw law)) ≤ γ*(N:ℝ)^(4/5:ℝ)*(Real.log (max (N:ℝ) 2))^(1/5:ℝ)","missing":[],"search":"expected_actual_rate banditrlproof.hoo.cantormodel.expected_actual_rate a concrete full-rate consequence with d'=3>dimension. the general theorem retains every d' strictly above the actual dimension; this model certificate is deliberately conservative and does not claim the sharp model exponent. theorem compiled","shard":"modules/74a1f773f2dc1a6f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalPacking","label":"nearOptimalPacking","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimalPacking","description":"noncomputable def RegularCovering.nearOptimalPacking {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c ε : ℝ) : ℕ","url":"../modules/banditrlproof-hoodimension/index.html#decl-cdfd4be64514","parent":"module:BanditRLProof.HOODimension","order":4927,"meta":[["Kind","definition"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def RegularCovering.nearOptimalPacking {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c ε : ℝ) : ℕ","missing":[],"search":"nearoptimalpacking banditrlproof.hoo.regularcovering.nearoptimalpacking noncomputable def regularcovering.nearoptimalpacking {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best c ε : ℝ) : ℕ definition compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.packingExponent","label":"packingExponent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.packingExponent","description":"The normalized extended logarithm. At zero packing the value is minus infinity; this does not invoke the totalized real logarithm at zero.","url":"../modules/banditrlproof-hoodimension/index.html#decl-b77d0ea5f50d","parent":"module:BanditRLProof.HOODimension","order":4928,"meta":[["Kind","definition"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def packingExponent (N : ℕ) (ε : ℝ) : EReal","missing":[],"search":"packingexponent banditrlproof.hoo.packingexponent the normalized extended logarithm. at zero packing the value is minus infinity; this does not invoke the totalized real logarithm at zero. definition compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.packingExponent_zero","label":"packingExponent_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.packingExponent_zero","description":"@[simp] theorem packingExponent_zero (ε : ℝ) : packingExponent 0 ε = ⊥","url":"../modules/banditrlproof-hoodimension/index.html#decl-24fafd621ef6","parent":"module:BanditRLProof.HOODimension","order":4929,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem packingExponent_zero (ε : ℝ) : packingExponent 0 ε = ⊥","missing":[],"search":"packingexponent_zero banditrlproof.hoo.packingexponent_zero @[simp] theorem packingexponent_zero (ε : ℝ) : packingexponent 0 ε = ⊥ theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.packingExponent_positive","label":"packingExponent_positive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.packingExponent_positive","description":"theorem packingExponent_positive {N : ℕ} (hN : 0<N) (ε : ℝ) : packingExponent N ε = (Real.log (N:ℝ) / Real.log (1/ε) : ℝ)","url":"../modules/banditrlproof-hoodimension/index.html#decl-0256b4e5d151","parent":"module:BanditRLProof.HOODimension","order":4930,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem packingExponent_positive {N : ℕ} (hN : 0<N) (ε : ℝ) : packingExponent N ε = (Real.log (N:ℝ) / Real.log (1/ε) : ℝ)","missing":[],"search":"packingexponent_positive banditrlproof.hoo.packingexponent_positive theorem packingexponent_positive {n : ℕ} (hn : 0<n) (ε : ℝ) : packingexponent n ε = (real.log (n:ℝ) / real.log (1/ε) : ℝ) theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalityDimension","label":"nearOptimalityDimension","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimalityDimension","description":"noncomputable def RegularCovering.nearOptimalityDimension {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c : ℝ) : EReal","url":"../modules/banditrlproof-hoodimension/index.html#decl-98b62ebc33a6","parent":"module:BanditRLProof.HOODimension","order":4931,"meta":[["Kind","definition"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def RegularCovering.nearOptimalityDimension {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c : ℝ) : EReal","missing":[],"search":"nearoptimalitydimension banditrlproof.hoo.regularcovering.nearoptimalitydimension noncomputable def regularcovering.nearoptimalitydimension {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best c : ℝ) : ereal definition compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalityDimension_nonneg","label":"nearOptimalityDimension_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimalityDimension_nonneg","description":"theorem RegularCovering.nearOptimalityDimension_nonneg {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c : ℝ) : 0 ≤ C.nearOptimalityDimension f best c","url":"../modules/banditrlproof-hoodimension/index.html#decl-2d2245103093","parent":"module:BanditRLProof.HOODimension","order":4932,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.nearOptimalityDimension_nonneg {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c : ℝ) : 0 ≤ C.nearOptimalityDimension f best c","missing":[],"search":"nearoptimalitydimension_nonneg banditrlproof.hoo.regularcovering.nearoptimalitydimension_nonneg theorem regularcovering.nearoptimalitydimension_nonneg {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best c : ℝ) : 0 ≤ c.nearoptimalitydimension f best c theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.packing_le_rpow_of_exponent_lt","label":"packing_le_rpow_of_exponent_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.packing_le_rpow_of_exponent_lt","description":"theorem packing_le_rpow_of_exponent_lt {N : ℕ} {ε d : ℝ} (hε : 0<ε) (hε1 : ε<1) (h : packingExponent N ε < (d:EReal)) : (N:ℝ) ≤ ε^(-d)","url":"../modules/banditrlproof-hoodimension/index.html#decl-e405025921d2","parent":"module:BanditRLProof.HOODimension","order":4933,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem packing_le_rpow_of_exponent_lt {N : ℕ} {ε d : ℝ} (hε : 0<ε) (hε1 : ε<1) (h : packingExponent N ε < (d:EReal)) : (N:ℝ) ≤ ε^(-d)","missing":[],"search":"packing_le_rpow_of_exponent_lt banditrlproof.hoo.packing_le_rpow_of_exponent_lt theorem packing_le_rpow_of_exponent_lt {n : ℕ} {ε d : ℝ} (hε : 0<ε) (hε1 : ε<1) (h : packingexponent n ε < (d:ereal)) : (n:ℝ) ≤ ε^(-d) theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.eventually_nearOptimalPacking_le","label":"eventually_nearOptimalPacking_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.eventually_nearOptimalPacking_le","description":"Strictly exceeding the actual limsup dimension produces a fine-scale packing bound; no power-law packing premise is assumed.","url":"../modules/banditrlproof-hoodimension/index.html#decl-880a6744189a","parent":"module:BanditRLProof.HOODimension","order":4934,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.eventually_nearOptimalPacking_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c d : ℝ) (hd : C.nearOptimalityDimension f best c < (d:EReal)) : ∀ᶠ ε in 𝓝[>] (0:ℝ), (C.nearOptimalPacking f best c ε : ℝ) ≤ ε^(-d)","missing":[],"search":"eventually_nearoptimalpacking_le banditrlproof.hoo.regularcovering.eventually_nearoptimalpacking_le strictly exceeding the actual limsup dimension produces a fine-scale packing bound; no power-law packing premise is assumed. theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.uniform_nearOptimalPacking_le","label":"uniform_nearOptimalPacking_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.uniform_nearOptimalPacking_le","description":"Source Theorem 6's uniform constant, including all coarse scales up to R. The fine-scale bound comes from Definition 5, the coarse bound from A1.","url":"../modules/banditrlproof-hoodimension/index.html#decl-0a5d6d4eb0f5","parent":"module:BanditRLProof.HOODimension","order":4935,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.uniform_nearOptimalPacking_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c d R : ℝ) (hR : 0<R) (hd : C.nearOptimalityDimension f best c < (d:EReal)) : ∃ K : ℝ, 0<K ∧ ∀ ε : ℝ, 0<ε → ε≤R → (C.nearOptimalPacking f best c ε : ℝ) ≤ K * ε^(-d)","missing":[],"search":"uniform_nearoptimalpacking_le banditrlproof.hoo.regularcovering.uniform_nearoptimalpacking_le source theorem 6's uniform constant, including all coarse scales up to r. the fine-scale bound comes from definition 5, the coarse bound from a1. theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalNodes_power_bound","label":"nearOptimalNodes_power_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimalNodes_power_bound","description":"Actual near-optimal tree levels inherit the power bound at the exact source parameter c=4*nu1/nu2. The constant is independent of depth.","url":"../modules/banditrlproof-hoodimension/index.html#decl-7786292722e0","parent":"module:BanditRLProof.HOODimension","order":4936,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOODimension"],["Source","BanditRLProof/HOODimension.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"theorem RegularCovering.nearOptimalNodes_power_bound {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best d : ℝ) (hw : WeaklyLipschitz f C.ell best) (hd : C.nearOptimalityDimension f best (4*C.nu1/C.nu2) < (d:EReal)) : ∃ K : ℝ, 0<K ∧ ∀ h : ℕ, ((C.nearOptimalNodes f best h).card : ℝ) ≤ K * (C.nu2*C.rho^h)^(-d)","missing":[],"search":"nearoptimalnodes_power_bound banditrlproof.hoo.regularcovering.nearoptimalnodes_power_bound actual near-optimal tree levels inherit the power bound at the exact source parameter c=4*nu1/nu2. the constant is independent of depth. theorem compiled","shard":"modules/ce32616e42a5f606.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.WeaklyLipschitz","label":"WeaklyLipschitz","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.WeaklyLipschitz","description":"def WeaklyLipschitz {X : Type*} (f : X → ℝ) (ell : X → X → ℝ) (best : ℝ) : Prop","url":"../modules/banditrlproof-hoogeometry/index.html#decl-c095db548e18","parent":"module:BanditRLProof.HOOGeometry","order":4937,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOGeometry"],["Source","BanditRLProof/HOOGeometry.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def WeaklyLipschitz {X : Type*} (f : X → ℝ) (ell : X → X → ℝ) (best : ℝ) : Prop","missing":[],"search":"weaklylipschitz banditrlproof.hoo.weaklylipschitz def weaklylipschitz {x : type*} (f : x → ℝ) (ell : x → x → ℝ) (best : ℝ) : prop definition compiled","shard":"modules/6b3741ed62ab0f12.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.regionSup","label":"regionSup","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.regionSup","description":"noncomputable def regionSup {X : Type*} (f : X → ℝ) (A : Set X) : ℝ","url":"../modules/banditrlproof-hoogeometry/index.html#decl-54f201871988","parent":"module:BanditRLProof.HOOGeometry","order":4938,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOGeometry"],["Source","BanditRLProof/HOOGeometry.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def regionSup {X : Type*} (f : X → ℝ) (A : Set X) : ℝ","missing":[],"search":"regionsup banditrlproof.hoo.regionsup noncomputable def regionsup {x : type*} (f : x → ℝ) (a : set x) : ℝ definition compiled","shard":"modules/6b3741ed62ab0f12.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.region_gap_le","label":"region_gap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.region_gap_le","description":"First part of source Lemma 3, retaining the region's exact suboptimality.","url":"../modules/banditrlproof-hoogeometry/index.html#decl-98e11f256070","parent":"module:BanditRLProof.HOOGeometry","order":4939,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOGeometry"],["Source","BanditRLProof/HOOGeometry.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem region_gap_le {X : Type*} (f : X → ℝ) (ell : X → X → ℝ) (best D : ℝ) (A : Set X) (hA : A.Nonempty) (hw : WeaklyLipschitz f ell best) (hdiam : ∀ x ∈ A, ∀ y ∈ A, ell x y ≤ D) (y : X) (hy : y ∈ A) : best - f y ≤ (best - regionSup f A) + max (best - regionSup f A) D","missing":[],"search":"region_gap_le banditrlproof.hoo.region_gap_le first part of source lemma 3, retaining the region's exact suboptimality. theorem compiled","shard":"modules/6b3741ed62ab0f12.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.near_optimal_region","label":"near_optimal_region","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.near_optimal_region","description":"Source Lemma 3, including c=0 and the max(2c,c+1) constant.","url":"../modules/banditrlproof-hoogeometry/index.html#decl-df483642e6d0","parent":"module:BanditRLProof.HOOGeometry","order":4940,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOGeometry"],["Source","BanditRLProof/HOOGeometry.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem near_optimal_region {X : Type*} (f : X → ℝ) (ell : X → X → ℝ) (best D c : ℝ) (A : Set X) (hA : A.Nonempty) (hD : 0 ≤ D) (hw : WeaklyLipschitz f ell best) (hdiam : ∀ x ∈ A, ∀ y ∈ A, ell x y ≤ D) (hgap : best - regionSup f A ≤ c*D) (y : X) (hy : y ∈ A) : best - f y ≤ max (2*c) (c+1)*D","missing":[],"search":"near_optimal_region banditrlproof.hoo.near_optimal_region source lemma 3, including c=0 and the max(2c,c+1) constant. theorem compiled","shard":"modules/6b3741ed62ab0f12.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.nodesAtDepth","label":"nodesAtDepth","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.nodesAtDepth","description":"def nodesAtDepth (n : ℕ) : Finset Node","url":"../modules/banditrlproof-hoolevels/index.html#decl-d7a62b5f997c","parent":"module:BanditRLProof.HOOLevels","order":4941,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOLevels"],["Source","BanditRLProof/HOOLevels.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def nodesAtDepth (n : ℕ) : Finset Node","missing":[],"search":"nodesatdepth banditrlproof.hoo.nodesatdepth def nodesatdepth (n : ℕ) : finset node definition compiled","shard":"modules/794e45f5a86bd3d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.mem_nodesAtDepth","label":"mem_nodesAtDepth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.mem_nodesAtDepth","description":"@[simp] theorem mem_nodesAtDepth (v : Node) (n : ℕ) : v ∈ nodesAtDepth n ↔ v.length=n","url":"../modules/banditrlproof-hoolevels/index.html#decl-6722a7c0bd6d","parent":"module:BanditRLProof.HOOLevels","order":4942,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOLevels"],["Source","BanditRLProof/HOOLevels.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem mem_nodesAtDepth (v : Node) (n : ℕ) : v ∈ nodesAtDepth n ↔ v.length=n","missing":[],"search":"mem_nodesatdepth banditrlproof.hoo.mem_nodesatdepth @[simp] theorem mem_nodesatdepth (v : node) (n : ℕ) : v ∈ nodesatdepth n ↔ v.length=n theorem compiled","shard":"modules/794e45f5a86bd3d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.card_nodesAtDepth","label":"card_nodesAtDepth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.card_nodesAtDepth","description":"@[simp] theorem card_nodesAtDepth (n : ℕ) : (nodesAtDepth n).card = 2^n","url":"../modules/banditrlproof-hoolevels/index.html#decl-5699da905278","parent":"module:BanditRLProof.HOOLevels","order":4943,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOLevels"],["Source","BanditRLProof/HOOLevels.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem card_nodesAtDepth (n : ℕ) : (nodesAtDepth n).card = 2^n","missing":[],"search":"card_nodesatdepth banditrlproof.hoo.card_nodesatdepth @[simp] theorem card_nodesatdepth (n : ℕ) : (nodesatdepth n).card = 2^n theorem compiled","shard":"modules/794e45f5a86bd3d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.exists_region_at_depth","label":"exists_region_at_depth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.exists_region_at_depth","description":"theorem Covering.exists_region_at_depth {X : Type*} (C : Covering X) (n : ℕ) (x : X) : ∃ v ∈ nodesAtDepth n, x ∈ C.region v","url":"../modules/banditrlproof-hoolevels/index.html#decl-598616627481","parent":"module:BanditRLProof.HOOLevels","order":4944,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOLevels"],["Source","BanditRLProof/HOOLevels.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.exists_region_at_depth {X : Type*} (C : Covering X) (n : ℕ) (x : X) : ∃ v ∈ nodesAtDepth n, x ∈ C.region v","missing":[],"search":"exists_region_at_depth banditrlproof.hoo.covering.exists_region_at_depth theorem covering.exists_region_at_depth {x : type*} (c : covering x) (n : ℕ) (x : x) : ∃ v ∈ nodesatdepth n, x ∈ c.region v theorem compiled","shard":"modules/794e45f5a86bd3d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.disjoint_ball_family_card_le","label":"disjoint_ball_family_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.disjoint_ball_family_card_le","description":"Coarse-scale finite packing bound from A1 alone, valid for asymmetric dissimilarities without a triangle inequality.","url":"../modules/banditrlproof-hoolevels/index.html#decl-ecfb20b2b345","parent":"module:BanditRLProof.HOOLevels","order":4945,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOLevels"],["Source","BanditRLProof/HOOLevels.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.disjoint_ball_family_card_le {X I : Type*} [MeasurableSpace X] [Fintype I] (C : RegularCovering X) (centers : I → X) (ε : ℝ) (h : ℕ) (hε : C.nu1*C.rho^h < ε) (hd : Pairwise (fun i j => Disjoint {y | C.ell (centers i) y < ε} {y | C.ell (centers j) y < ε})) : Fintype.card I ≤ 2^h","missing":[],"search":"disjoint_ball_family_card_le banditrlproof.hoo.regularcovering.disjoint_ball_family_card_le coarse-scale finite packing bound from a1 alone, valid for asymmetric dissimilarities without a triangle inequality. theorem compiled","shard":"modules/794e45f5a86bd3d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.exists_finite_packing_bound","label":"exists_finite_packing_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.exists_finite_packing_bound","description":"theorem RegularCovering.exists_finite_packing_bound {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (ε : ℝ) (hε : 0 < ε) : ∃ M : ℕ, ∀ (I : Type) [Fintype I] (centers : I → X), Pairwise (fun i j => Disjoint {y | C.ell (centers i) y < ε} {y | C.ell (centers j) y < ε}) → Fintype.card I ≤ M","url":"../modules/banditrlproof-hoolevels/index.html#decl-faf598d964ac","parent":"module:BanditRLProof.HOOLevels","order":4946,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOLevels"],["Source","BanditRLProof/HOOLevels.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.exists_finite_packing_bound {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (ε : ℝ) (hε : 0 < ε) : ∃ M : ℕ, ∀ (I : Type) [Fintype I] (centers : I → X), Pairwise (fun i j => Disjoint {y | C.ell (centers i) y < ε} {y | C.ell (centers j) y < ε}) → Fintype.card I ≤ M","missing":[],"search":"exists_finite_packing_bound banditrlproof.hoo.regularcovering.exists_finite_packing_bound theorem regularcovering.exists_finite_packing_bound {x : type*} [measurablespace x] (c : regularcovering x) (ε : ℝ) (hε : 0 < ε) : ∃ m : ℕ, ∀ (i : type) [fintype i] (centers : i → x), pairwise (fun i j => disjoint {y | c.ell (centers i) y < ε} {y | c.ell (centers j) y < ε}) → fintype.card i ≤ m theorem compiled","shard":"modules/794e45f5a86bd3d2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering","label":"Covering","kind":"structure","status":"compiled","subtitle":"BanditRLProof.HOO.Covering","description":"structure Covering (X : Type*) where","url":"../modules/banditrlproof-hoomodel/index.html#decl-937470b1df80","parent":"module:BanditRLProof.HOOModel","order":4947,"meta":[["Kind","structure"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure Covering (X : Type*) where","missing":[],"search":"covering banditrlproof.hoo.covering structure covering (x : type*) where structure compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.child_subset","label":"child_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.child_subset","description":"theorem Covering.child_subset {X : Type*} (C : Covering X) (v : Node) (b : Bool) : C.region (child v b) ⊆ C.region v","url":"../modules/banditrlproof-hoomodel/index.html#decl-bc3a852f0d2c","parent":"module:BanditRLProof.HOOModel","order":4948,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.child_subset {X : Type*} (C : Covering X) (v : Node) (b : Bool) : C.region (child v b) ⊆ C.region v","missing":[],"search":"child_subset banditrlproof.hoo.covering.child_subset theorem covering.child_subset {x : type*} (c : covering x) (v : node) (b : bool) : c.region (child v b) ⊆ c.region v theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.append_subset","label":"append_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.append_subset","description":"theorem Covering.append_subset {X : Type*} (C : Covering X) (v u : Node) : C.region (v ++ u) ⊆ C.region v","url":"../modules/banditrlproof-hoomodel/index.html#decl-847e8b740094","parent":"module:BanditRLProof.HOOModel","order":4949,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.append_subset {X : Type*} (C : Covering X) (v u : Node) : C.region (v ++ u) ⊆ C.region v","missing":[],"search":"append_subset banditrlproof.hoo.covering.append_subset theorem covering.append_subset {x : type*} (c : covering x) (v u : node) : c.region (v ++ u) ⊆ c.region v theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.descendant_subset","label":"descendant_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.descendant_subset","description":"theorem Covering.descendant_subset {X : Type*} (C : Covering X) {v w : Node} (h : v <+: w) : C.region w ⊆ C.region v","url":"../modules/banditrlproof-hoomodel/index.html#decl-03f96c9eb8f2","parent":"module:BanditRLProof.HOOModel","order":4950,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.descendant_subset {X : Type*} (C : Covering X) {v w : Node} (h : v <+: w) : C.region w ⊆ C.region v","missing":[],"search":"descendant_subset banditrlproof.hoo.covering.descendant_subset theorem covering.descendant_subset {x : type*} (c : covering x) {v w : node} (h : v <+: w) : c.region w ⊆ c.region v theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.representative","label":"representative","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.representative","description":"noncomputable def Covering.representative {X : Type*} (C : Covering X) (v : Node) : X","url":"../modules/banditrlproof-hoomodel/index.html#decl-6805fdb4d6eb","parent":"module:BanditRLProof.HOOModel","order":4951,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def Covering.representative {X : Type*} (C : Covering X) (v : Node) : X","missing":[],"search":"representative banditrlproof.hoo.covering.representative noncomputable def covering.representative {x : type*} (c : covering x) (v : node) : x definition compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.representative_mem","label":"representative_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.representative_mem","description":"theorem Covering.representative_mem {X : Type*} (C : Covering X) (v : Node) : C.representative v ∈ C.region v","url":"../modules/banditrlproof-hoomodel/index.html#decl-71e7ab42380c","parent":"module:BanditRLProof.HOOModel","order":4952,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.representative_mem {X : Type*} (C : Covering X) (v : Node) : C.representative v ∈ C.region v","missing":[],"search":"representative_mem banditrlproof.hoo.covering.representative_mem theorem covering.representative_mem {x : type*} (c : covering x) (v : node) : c.representative v ∈ c.region v theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering","label":"RegularCovering","kind":"structure","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering","description":"A1, stated with pointwise diameter bounds equivalent to the source bound. Inner balls need not be metric balls: asymmetry and lack of triangle inequality are retained.","url":"../modules/banditrlproof-hoomodel/index.html#decl-76ea58441407","parent":"module:BanditRLProof.HOOModel","order":4953,"meta":[["Kind","structure"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","lipschitz"]],"statement":"structure RegularCovering (X : Type*) [MeasurableSpace X] extends Covering X where","missing":[],"search":"regularcovering banditrlproof.hoo.regularcovering a1, stated with pointwise diameter bounds equivalent to the source bound. inner balls need not be metric balls: asymmetry and lack of triangle inequality are retained. structure compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["lipschitz"]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.region_near_optimal","label":"region_near_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.region_near_optimal","description":"theorem RegularCovering.region_near_optimal {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c : ℝ) (v : Node) (hw : WeaklyLipschitz f C.ell best) (hgap : best - regionSup f (C.region v) ≤ c*(C.nu1*C.rho^v.length)) (y : X) (hy : y ∈ C.region v) : best - f y ≤ max (2*c) (c+1)*(C.nu1*C.rho^v.length)","url":"../modules/banditrlproof-hoomodel/index.html#decl-29865291bf00","parent":"module:BanditRLProof.HOOModel","order":4954,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.region_near_optimal {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best c : ℝ) (v : Node) (hw : WeaklyLipschitz f C.ell best) (hgap : best - regionSup f (C.region v) ≤ c*(C.nu1*C.rho^v.length)) (y : X) (hy : y ∈ C.region v) : best - f y ≤ max (2*c) (c+1)*(C.nu1*C.rho^v.length)","missing":[],"search":"region_near_optimal banditrlproof.hoo.regularcovering.region_near_optimal theorem regularcovering.region_near_optimal {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best c : ℝ) (v : node) (hw : weaklylipschitz f c.ell best) (hgap : best - regionsup f (c.region v) ≤ c*(c.nu1*c.rho^v.length)) (y : x) (hy : y ∈ c.region v) : best - f y ≤ max (2*c) (c+1)*(c.nu1*c.rho^v.length) theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.nodeLaw","label":"nodeLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.nodeLaw","description":"Fixed representatives are allowed by source Algorithm 1. Their countable domain makes the arm-to-law map measurable without requiring a continuous selector on the entire arm space.","url":"../modules/banditrlproof-hoomodel/index.html#decl-1798b495a4be","parent":"module:BanditRLProof.HOOModel","order":4955,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def Covering.nodeLaw {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) : Kernel Node ℝ","missing":[],"search":"nodelaw banditrlproof.hoo.covering.nodelaw fixed representatives are allowed by source algorithm 1. their countable domain makes the arm-to-law map measurable without requiring a continuous selector on the entire arm space. definition compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.arm","label":"arm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.arm","description":"noncomputable def Covering.arm {X : Type*} (C : Covering X) (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : X","url":"../modules/banditrlproof-hoomodel/index.html#decl-3ccc0fc03b7a","parent":"module:BanditRLProof.HOOModel","order":4956,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def Covering.arm {X : Type*} (C : Covering X) (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : X","missing":[],"search":"arm banditrlproof.hoo.covering.arm noncomputable def covering.arm {x : type*} (c : covering x) (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : x definition compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.arm_mem","label":"arm_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.arm_mem","description":"theorem Covering.arm_mem {X : Type*} (C : Covering X) (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : C.arm ν ρ Y n ∈ C.region (action ν ρ Y n)","url":"../modules/banditrlproof-hoomodel/index.html#decl-20b9a6ea28d8","parent":"module:BanditRLProof.HOOModel","order":4957,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.arm_mem {X : Type*} (C : Covering X) (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : C.arm ν ρ Y n ∈ C.region (action ν ρ Y n)","missing":[],"search":"arm_mem banditrlproof.hoo.covering.arm_mem theorem covering.arm_mem {x : type*} (c : covering x) (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : c.arm ν ρ y n ∈ c.region (action ν ρ y n) theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.measurable_arm","label":"measurable_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.measurable_arm","description":"theorem Covering.measurable_arm {X : Type*} [MeasurableSpace X] (C : Covering X) (ν ρ : ℝ) (n : ℕ) : Measurable (fun Y => C.arm ν ρ Y n)","url":"../modules/banditrlproof-hoomodel/index.html#decl-fd10abacdb01","parent":"module:BanditRLProof.HOOModel","order":4958,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:92"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.measurable_arm {X : Type*} [MeasurableSpace X] (C : Covering X) (ν ρ : ℝ) (n : ℕ) : Measurable (fun Y => C.arm ν ρ Y n)","missing":[],"search":"measurable_arm banditrlproof.hoo.covering.measurable_arm theorem covering.measurable_arm {x : type*} [measurablespace x] (c : covering x) (ν ρ : ℝ) (n : ℕ) : measurable (fun y => c.arm ν ρ y n) theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.stepKernel_actual_arm","label":"stepKernel_actual_arm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.stepKernel_actual_arm","description":"theorem Covering.stepKernel_actual_arm {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : stepKernel ν ρ (C.nodeLaw law) n (Preorder.frestrictLe n Y) = law (C.arm ν ρ Y (n+1))","url":"../modules/banditrlproof-hoomodel/index.html#decl-f5fde623385d","parent":"module:BanditRLProof.HOOModel","order":4959,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOModel"],["Source","BanditRLProof/HOOModel.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.stepKernel_actual_arm {X : Type*} [MeasurableSpace X] (C : Covering X) (law : Kernel X ℝ) (ν ρ : ℝ) (Y : ℕ → ℝ) (n : ℕ) : stepKernel ν ρ (C.nodeLaw law) n (Preorder.frestrictLe n Y) = law (C.arm ν ρ Y (n+1))","missing":[],"search":"stepkernel_actual_arm banditrlproof.hoo.covering.stepkernel_actual_arm theorem covering.stepkernel_actual_arm {x : type*} [measurablespace x] (c : covering x) (law : kernel x ℝ) (ν ρ : ℝ) (y : ℕ → ℝ) (n : ℕ) : stepkernel ν ρ (c.nodelaw law) n (preorder.frestrictle n y) = law (c.arm ν ρ y (n+1)) theorem compiled","shard":"modules/1509ac4b769586f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.regionSup_children","label":"regionSup_children","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.regionSup_children","description":"theorem Covering.regionSup_children {X : Type*} (C : Covering X) (f : X → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (v : Node) : regionSup f (C.region v) = max (regionSup f (C.region (child v false))) (regionSup f (C.region (child v true)))","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-19d84873f8f9","parent":"module:BanditRLProof.HOOOptimalBranch","order":4960,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.regionSup_children {X : Type*} (C : Covering X) (f : X → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (v : Node) : regionSup f (C.region v) = max (regionSup f (C.region (child v false))) (regionSup f (C.region (child v true)))","missing":[],"search":"regionsup_children banditrlproof.hoo.covering.regionsup_children theorem covering.regionsup_children {x : type*} (c : covering x) (f : x → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (v : node) : regionsup f (c.region v) = max (regionsup f (c.region (child v false))) (regionsup f (c.region (child v true))) theorem compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.optimalChild","label":"optimalChild","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.optimalChild","description":"noncomputable def Covering.optimalChild {X : Type*} (C : Covering X) (f : X → ℝ) (v : Node) : Bool","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-bf5f02c16f2a","parent":"module:BanditRLProof.HOOOptimalBranch","order":4961,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def Covering.optimalChild {X : Type*} (C : Covering X) (f : X → ℝ) (v : Node) : Bool","missing":[],"search":"optimalchild banditrlproof.hoo.covering.optimalchild noncomputable def covering.optimalchild {x : type*} (c : covering x) (f : x → ℝ) (v : node) : bool definition compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.optimalChild_sup","label":"optimalChild_sup","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.optimalChild_sup","description":"theorem Covering.optimalChild_sup {X : Type*} (C : Covering X) (f : X → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (v : Node) : regionSup f (C.region (child v (C.optimalChild f v))) = regionSup f (C.region v)","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-1494cbbdf51f","parent":"module:BanditRLProof.HOOOptimalBranch","order":4962,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.optimalChild_sup {X : Type*} (C : Covering X) (f : X → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (v : Node) : regionSup f (C.region (child v (C.optimalChild f v))) = regionSup f (C.region v)","missing":[],"search":"optimalchild_sup banditrlproof.hoo.covering.optimalchild_sup theorem covering.optimalchild_sup {x : type*} (c : covering x) (f : x → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (v : node) : regionsup f (c.region (child v (c.optimalchild f v))) = regionsup f (c.region v) theorem compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.optimalPath","label":"optimalPath","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.optimalPath","description":"noncomputable def Covering.optimalPath {X : Type*} (C : Covering X) (f : X → ℝ) : ℕ → Node | 0 => [] | n+1 => child (C.optimalPath f n) (C.optimalChild f (C.optimalPath f n)) @[simp] theorem Covering.optimalPath_length {X : Type*} (C : Covering X) (f : X → ℝ) (n : ℕ) : (C.optimalPath f n).length = n","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-76758071229e","parent":"module:BanditRLProof.HOOOptimalBranch","order":4963,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def Covering.optimalPath {X : Type*} (C : Covering X) (f : X → ℝ) : ℕ → Node | 0 => [] | n+1 => child (C.optimalPath f n) (C.optimalChild f (C.optimalPath f n)) @[simp] theorem Covering.optimalPath_length {X : Type*} (C : Covering X) (f : X → ℝ) (n : ℕ) : (C.optimalPath f n).length = n","missing":[],"search":"optimalpath banditrlproof.hoo.covering.optimalpath noncomputable def covering.optimalpath {x : type*} (c : covering x) (f : x → ℝ) : ℕ → node | 0 => [] | n+1 => child (c.optimalpath f n) (c.optimalchild f (c.optimalpath f n)) @[simp] theorem covering.optimalpath_length {x : type*} (c : covering x) (f : x → ℝ) (n : ℕ) : (c.optimalpath f n).length = n definition compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.optimalPath_length","label":"optimalPath_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.optimalPath_length","description":"@[simp] theorem Covering.optimalPath_length {X : Type*} (C : Covering X) (f : X → ℝ) (n : ℕ) : (C.optimalPath f n).length = n","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-fadfc62ac006","parent":"module:BanditRLProof.HOOOptimalBranch","order":4964,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem Covering.optimalPath_length {X : Type*} (C : Covering X) (f : X → ℝ) (n : ℕ) : (C.optimalPath f n).length = n","missing":[],"search":"optimalpath_length banditrlproof.hoo.covering.optimalpath_length @[simp] theorem covering.optimalpath_length {x : type*} (c : covering x) (f : x → ℝ) (n : ℕ) : (c.optimalpath f n).length = n theorem compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.optimalPath_sup","label":"optimalPath_sup","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.optimalPath_sup","description":"theorem Covering.optimalPath_sup {X : Type*} (C : Covering X) (f : X → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (n : ℕ) : regionSup f (C.region (C.optimalPath f n)) = regionSup f Set.univ","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-1898d15b513b","parent":"module:BanditRLProof.HOOOptimalBranch","order":4965,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.optimalPath_sup {X : Type*} (C : Covering X) (f : X → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (n : ℕ) : regionSup f (C.region (C.optimalPath f n)) = regionSup f Set.univ","missing":[],"search":"optimalpath_sup banditrlproof.hoo.covering.optimalpath_sup theorem covering.optimalpath_sup {x : type*} (c : covering x) (f : x → ℝ) (best : ℝ) (hf : ∀ x, f x ≤ best) (n : ℕ) : regionsup f (c.region (c.optimalpath f n)) = regionsup f set.univ theorem compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.optimalPath_prefix","label":"optimalPath_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.optimalPath_prefix","description":"theorem Covering.optimalPath_prefix {X : Type*} (C : Covering X) (f : X → ℝ) {i j : ℕ} (hij : i ≤ j) : C.optimalPath f i <+: C.optimalPath f j","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-9632f591a30e","parent":"module:BanditRLProof.HOOOptimalBranch","order":4966,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.optimalPath_prefix {X : Type*} (C : Covering X) (f : X → ℝ) {i j : ℕ} (hij : i ≤ j) : C.optimalPath f i <+: C.optimalPath f j","missing":[],"search":"optimalpath_prefix banditrlproof.hoo.covering.optimalpath_prefix theorem covering.optimalpath_prefix {x : type*} (c : covering x) (f : x → ℝ) {i j : ℕ} (hij : i ≤ j) : c.optimalpath f i <+: c.optimalpath f j theorem compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.Covering.exists_first_unexpanded","label":"exists_first_unexpanded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.Covering.exists_first_unexpanded","description":"Every finite expanded tree has a first missing node on this infinite branch, bounded by its maximum depth plus one.","url":"../modules/banditrlproof-hoooptimalbranch/index.html#decl-6dededa79803","parent":"module:BanditRLProof.HOOOptimalBranch","order":4967,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOOptimalBranch"],["Source","BanditRLProof/HOOOptimalBranch.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem Covering.exists_first_unexpanded {X : Type*} (C : Covering X) (f : X → ℝ) (S : Finset Node) : ∃ k ≤ depthBound S + 1, C.optimalPath f k ∉ S ∧ ∀ j < k, C.optimalPath f j ∈ S","missing":[],"search":"exists_first_unexpanded banditrlproof.hoo.covering.exists_first_unexpanded every finite expanded tree has a first missing node on this infinite branch, bounded by its maximum depth plus one. theorem compiled","shard":"modules/18b477f5f279c715.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.ContainedBallPacking","label":"ContainedBallPacking","kind":"structure","status":"compiled","subtitle":"BanditRLProof.HOO.ContainedBallPacking","description":"structure ContainedBallPacking {X I : Type*} (ell : X → X → ℝ) (A : Set X) (ε : ℝ) (centers : I → X) : Prop where","url":"../modules/banditrlproof-hoopacking/index.html#decl-6c60ed5b6f46","parent":"module:BanditRLProof.HOOPacking","order":4968,"meta":[["Kind","structure"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure ContainedBallPacking {X I : Type*} (ell : X → X → ℝ) (A : Set X) (ε : ℝ) (centers : I → X) : Prop where","missing":[],"search":"containedballpacking banditrlproof.hoo.containedballpacking structure containedballpacking {x i : type*} (ell : x → x → ℝ) (a : set x) (ε : ℝ) (centers : i → x) : prop where structure compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.packingSizes","label":"packingSizes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.packingSizes","description":"def packingSizes {X : Type*} (ell : X → X → ℝ) (A : Set X) (ε : ℝ) : Set ℕ","url":"../modules/banditrlproof-hoopacking/index.html#decl-44f993e46c2b","parent":"module:BanditRLProof.HOOPacking","order":4969,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def packingSizes {X : Type*} (ell : X → X → ℝ) (A : Set X) (ε : ℝ) : Set ℕ","missing":[],"search":"packingsizes banditrlproof.hoo.packingsizes def packingsizes {x : type*} (ell : x → x → ℝ) (a : set x) (ε : ℝ) : set ℕ definition compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.zero_mem_packingSizes","label":"zero_mem_packingSizes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.zero_mem_packingSizes","description":"theorem zero_mem_packingSizes {X : Type*} (ell : X → X → ℝ) (A : Set X) (ε : ℝ) : 0 ∈ packingSizes ell A ε","url":"../modules/banditrlproof-hoopacking/index.html#decl-ed6f1415bed5","parent":"module:BanditRLProof.HOOPacking","order":4970,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem zero_mem_packingSizes {X : Type*} (ell : X → X → ℝ) (A : Set X) (ε : ℝ) : 0 ∈ packingSizes ell A ε","missing":[],"search":"zero_mem_packingsizes banditrlproof.hoo.zero_mem_packingsizes theorem zero_mem_packingsizes {x : type*} (ell : x → x → ℝ) (a : set x) (ε : ℝ) : 0 ∈ packingsizes ell a ε theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.packingSizes_bddAbove","label":"packingSizes_bddAbove","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.packingSizes_bddAbove","description":"theorem RegularCovering.packingSizes_bddAbove {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (A : Set X) (ε : ℝ) (hε : 0 < ε) : BddAbove (packingSizes C.ell A ε)","url":"../modules/banditrlproof-hoopacking/index.html#decl-ee73b5e5e2ac","parent":"module:BanditRLProof.HOOPacking","order":4971,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.packingSizes_bddAbove {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (A : Set X) (ε : ℝ) (hε : 0 < ε) : BddAbove (packingSizes C.ell A ε)","missing":[],"search":"packingsizes_bddabove banditrlproof.hoo.regularcovering.packingsizes_bddabove theorem regularcovering.packingsizes_bddabove {x : type*} [measurablespace x] (c : regularcovering x) (a : set x) (ε : ℝ) (hε : 0 < ε) : bddabove (packingsizes c.ell a ε) theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber","label":"packingNumber","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.packingNumber","description":"Natural-valued source packing number on an A1 space. All semantic theorems below require positive radius, where finiteness is proved.","url":"../modules/banditrlproof-hoopacking/index.html#decl-2d0a15246547","parent":"module:BanditRLProof.HOOPacking","order":4972,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def RegularCovering.packingNumber {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (A : Set X) (ε : ℝ) : ℕ","missing":[],"search":"packingnumber banditrlproof.hoo.regularcovering.packingnumber natural-valued source packing number on an a1 space. all semantic theorems below require positive radius, where finiteness is proved. definition compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber_attained","label":"packingNumber_attained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.packingNumber_attained","description":"theorem RegularCovering.packingNumber_attained {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (A : Set X) (ε : ℝ) (hε : 0 < ε) : ∃ centers : Fin (C.packingNumber A ε) → X, ContainedBallPacking C.ell A ε centers","url":"../modules/banditrlproof-hoopacking/index.html#decl-3941dae515d9","parent":"module:BanditRLProof.HOOPacking","order":4973,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.packingNumber_attained {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (A : Set X) (ε : ℝ) (hε : 0 < ε) : ∃ centers : Fin (C.packingNumber A ε) → X, ContainedBallPacking C.ell A ε centers","missing":[],"search":"packingnumber_attained banditrlproof.hoo.regularcovering.packingnumber_attained theorem regularcovering.packingnumber_attained {x : type*} [measurablespace x] (c : regularcovering x) (a : set x) (ε : ℝ) (hε : 0 < ε) : ∃ centers : fin (c.packingnumber a ε) → x, containedballpacking c.ell a ε centers theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber_mono","label":"packingNumber_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.packingNumber_mono","description":"theorem RegularCovering.packingNumber_mono {X : Type*} [MeasurableSpace X] (C : RegularCovering X) {A B : Set X} (hAB : A ⊆ B) (ε : ℝ) (hε : 0 < ε) : C.packingNumber A ε ≤ C.packingNumber B ε","url":"../modules/banditrlproof-hoopacking/index.html#decl-01b77ffed646","parent":"module:BanditRLProof.HOOPacking","order":4974,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.packingNumber_mono {X : Type*} [MeasurableSpace X] (C : RegularCovering X) {A B : Set X} (hAB : A ⊆ B) (ε : ℝ) (hε : 0 < ε) : C.packingNumber A ε ≤ C.packingNumber B ε","missing":[],"search":"packingnumber_mono banditrlproof.hoo.regularcovering.packingnumber_mono theorem regularcovering.packingnumber_mono {x : type*} [measurablespace x] (c : regularcovering x) {a b : set x} (hab : a ⊆ b) (ε : ℝ) (hε : 0 < ε) : c.packingnumber a ε ≤ c.packingnumber b ε theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.containedPacking_card_le","label":"containedPacking_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.containedPacking_card_le","description":"theorem RegularCovering.containedPacking_card_le {X I : Type*} [MeasurableSpace X] [Fintype I] (C : RegularCovering X) (A : Set X) (ε : ℝ) (hε : 0 < ε) (centers : I → X) (hc : ContainedBallPacking C.ell A ε centers) : Fintype.card I ≤ C.packingNumber A ε","url":"../modules/banditrlproof-hoopacking/index.html#decl-f88a6a6010c9","parent":"module:BanditRLProof.HOOPacking","order":4975,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.containedPacking_card_le {X I : Type*} [MeasurableSpace X] [Fintype I] (C : RegularCovering X) (A : Set X) (ε : ℝ) (hε : 0 < ε) (centers : I → X) (hc : ContainedBallPacking C.ell A ε centers) : Fintype.card I ≤ C.packingNumber A ε","missing":[],"search":"containedpacking_card_le banditrlproof.hoo.regularcovering.containedpacking_card_le theorem regularcovering.containedpacking_card_le {x i : type*} [measurablespace x] [fintype i] (c : regularcovering x) (a : set x) (ε : ℝ) (hε : 0 < ε) (centers : i → x) (hc : containedballpacking c.ell a ε centers) : fintype.card i ≤ c.packingnumber a ε theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalNodes","label":"nearOptimalNodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimalNodes","description":"noncomputable def RegularCovering.nearOptimalNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) : Finset Node","url":"../modules/banditrlproof-hoopacking/index.html#decl-d33ac21a8e43","parent":"module:BanditRLProof.HOOPacking","order":4976,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def RegularCovering.nearOptimalNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) : Finset Node","missing":[],"search":"nearoptimalnodes banditrlproof.hoo.regularcovering.nearoptimalnodes noncomputable def regularcovering.nearoptimalnodes {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) : finset node definition compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalNodes_card_le_packing","label":"nearOptimalNodes_card_le_packing","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimalNodes_card_le_packing","description":"Source Theorem 6, second-step packing producer for the actual near-optimal nodes. It constructs contained balls via A1 and Lemma 3 with c=2.","url":"../modules/banditrlproof-hoopacking/index.html#decl-f4f72ce2da27","parent":"module:BanditRLProof.HOOPacking","order":4977,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.nearOptimalNodes_card_le_packing {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) (h : ℕ) : (C.nearOptimalNodes f best h).card ≤ C.packingNumber {y | best-f y ≤ 4*(C.nu1*C.rho^h)} (C.nu2*C.rho^h)","missing":[],"search":"nearoptimalnodes_card_le_packing banditrlproof.hoo.regularcovering.nearoptimalnodes_card_le_packing source theorem 6, second-step packing producer for the actual near-optimal nodes. it constructs contained balls via a1 and lemma 3 with c=2. theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber_le_ambient","label":"packingNumber_le_ambient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.packingNumber_le_ambient","description":"Larger-radius contained packings inject into an ambient smaller-radius packing by shrinking every ball, without symmetry or a triangle inequality.","url":"../modules/banditrlproof-hoopacking/index.html#decl-75d81631fb27","parent":"module:BanditRLProof.HOOPacking","order":4978,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPacking"],["Source","BanditRLProof/HOOPacking.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.packingNumber_le_ambient {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (A : Set X) {δ ε : ℝ} (hδ : 0<δ) (hδε : δ≤ε) : C.packingNumber A ε ≤ C.packingNumber Set.univ δ","missing":[],"search":"packingnumber_le_ambient banditrlproof.hoo.regularcovering.packingnumber_le_ambient larger-radius contained packings inject into an ambient smaller-radius packing by shrinking every ball, without symmetry or a triangle inequality. theorem compiled","shard":"modules/e619d0a70af0d280.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.mem_nearOptimalNodes","label":"mem_nearOptimalNodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.mem_nearOptimalNodes","description":"@[simp] theorem RegularCovering.mem_nearOptimalNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (v : Node) (h : ℕ) : v ∈ C.nearOptimalNodes f best h ↔ v.length=h ∧ best-regionSup f (C.region v) ≤ 2*(C.nu1*C.rho^h)","url":"../modules/banditrlproof-hoopartition/index.html#decl-93a8ee99c2cf","parent":"module:BanditRLProof.HOOPartition","order":4979,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem RegularCovering.mem_nearOptimalNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (v : Node) (h : ℕ) : v ∈ C.nearOptimalNodes f best h ↔ v.length=h ∧ best-regionSup f (C.region v) ≤ 2*(C.nu1*C.rho^h)","missing":[],"search":"mem_nearoptimalnodes banditrlproof.hoo.regularcovering.mem_nearoptimalnodes @[simp] theorem regularcovering.mem_nearoptimalnodes {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (v : node) (h : ℕ) : v ∈ c.nearoptimalnodes f best h ↔ v.length=h ∧ best-regionsup f (c.region v) ≤ 2*(c.nu1*c.rho^h) theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.root_nearOptimal","label":"root_nearOptimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.root_nearOptimal","description":"theorem RegularCovering.root_nearOptimal {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hbest : regionSup f Set.univ = best) : [] ∈ C.nearOptimalNodes f best 0","url":"../modules/banditrlproof-hoopartition/index.html#decl-27acf428f093","parent":"module:BanditRLProof.HOOPartition","order":4980,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.root_nearOptimal {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hbest : regionSup f Set.univ = best) : [] ∈ C.nearOptimalNodes f best 0","missing":[],"search":"root_nearoptimal banditrlproof.hoo.regularcovering.root_nearoptimal theorem regularcovering.root_nearoptimal {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hbest : regionsup f set.univ = best) : [] ∈ c.nearoptimalnodes f best 0 theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimal_prefix","label":"nearOptimal_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.nearOptimal_prefix","description":"theorem RegularCovering.nearOptimal_prefix {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x ≤ best) {v w : Node} (hp : v <+: w) (hw : w ∈ C.nearOptimalNodes f best w.length) : v ∈ C.nearOptimalNodes f best v.length","url":"../modules/banditrlproof-hoopartition/index.html#decl-08b1cfdf7ea2","parent":"module:BanditRLProof.HOOPartition","order":4981,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.nearOptimal_prefix {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x ≤ best) {v w : Node} (hp : v <+: w) (hw : w ∈ C.nearOptimalNodes f best w.length) : v ∈ C.nearOptimalNodes f best v.length","missing":[],"search":"nearoptimal_prefix banditrlproof.hoo.regularcovering.nearoptimal_prefix theorem regularcovering.nearoptimal_prefix {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hf : ∀x, f x ≤ best) {v w : node} (hp : v <+: w) (hw : w ∈ c.nearoptimalnodes f best w.length) : v ∈ c.nearoptimalnodes f best v.length theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.boundaryNodes","label":"boundaryNodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.boundaryNodes","description":"Indexed by the parent's depth: this is source J_(h+1).","url":"../modules/banditrlproof-hoopartition/index.html#decl-c1b48878602f","parent":"module:BanditRLProof.HOOPartition","order":4982,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def RegularCovering.boundaryNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) : Finset Node","missing":[],"search":"boundarynodes banditrlproof.hoo.regularcovering.boundarynodes indexed by the parent's depth: this is source j_(h+1). definition compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.mem_boundaryNodes","label":"mem_boundaryNodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.mem_boundaryNodes","description":"theorem RegularCovering.mem_boundaryNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) (v : Node) : v ∈ C.boundaryNodes f best h ↔ (∃ p ∈ C.nearOptimalNodes f best h, ∃ b, child p b=v) ∧ v ∉ C.nearOptimalNodes f best (h+1)","url":"../modules/banditrlproof-hoopartition/index.html#decl-8b33c738c37b","parent":"module:BanditRLProof.HOOPartition","order":4983,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.mem_boundaryNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) (v : Node) : v ∈ C.boundaryNodes f best h ↔ (∃ p ∈ C.nearOptimalNodes f best h, ∃ b, child p b=v) ∧ v ∉ C.nearOptimalNodes f best (h+1)","missing":[],"search":"mem_boundarynodes banditrlproof.hoo.regularcovering.mem_boundarynodes theorem regularcovering.mem_boundarynodes {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) (v : node) : v ∈ c.boundarynodes f best h ↔ (∃ p ∈ c.nearoptimalnodes f best h, ∃ b, child p b=v) ∧ v ∉ c.nearoptimalnodes f best (h+1) theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.boundaryNodes_card_le","label":"boundaryNodes_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.boundaryNodes_card_le","description":"theorem RegularCovering.boundaryNodes_card_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) : (C.boundaryNodes f best h).card ≤ 2*(C.nearOptimalNodes f best h).card","url":"../modules/banditrlproof-hoopartition/index.html#decl-c93a74a1b371","parent":"module:BanditRLProof.HOOPartition","order":4984,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.boundaryNodes_card_le {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (h : ℕ) : (C.boundaryNodes f best h).card ≤ 2*(C.nearOptimalNodes f best h).card","missing":[],"search":"boundarynodes_card_le banditrlproof.hoo.regularcovering.boundarynodes_card_le theorem regularcovering.boundarynodes_card_le {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) : (c.boundarynodes f best h).card ≤ 2*(c.nearoptimalnodes f best h).card theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.boundaryNodes_poor","label":"boundaryNodes_poor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.boundaryNodes_poor","description":"theorem RegularCovering.boundaryNodes_poor {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) {h : ℕ} {v : Node} (hv : v ∈ C.boundaryNodes f best h) : v.length=h+1 ∧ 2*(C.nu1*C.rho^(h+1)) < best-regionSup f (C.region v)","url":"../modules/banditrlproof-hoopartition/index.html#decl-0b08ca6e6f7a","parent":"module:BanditRLProof.HOOPartition","order":4985,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:60"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.boundaryNodes_poor {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) {h : ℕ} {v : Node} (hv : v ∈ C.boundaryNodes f best h) : v.length=h+1 ∧ 2*(C.nu1*C.rho^(h+1)) < best-regionSup f (C.region v)","missing":[],"search":"boundarynodes_poor banditrlproof.hoo.regularcovering.boundarynodes_poor theorem regularcovering.boundarynodes_poor {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) {h : ℕ} {v : node} (hv : v ∈ c.boundarynodes f best h) : v.length=h+1 ∧ 2*(c.nu1*c.rho^(h+1)) < best-regionsup f (c.region v) theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.node_partition_cover","label":"node_partition_cover","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.node_partition_cover","description":"Any word either lies below a good depth-H prefix, is itself a shallow near-optimal word, or lies below a first bad child of a near-optimal parent.","url":"../modules/banditrlproof-hoopartition/index.html#decl-e072745440f9","parent":"module:BanditRLProof.HOOPartition","order":4986,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.node_partition_cover {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hbest : regionSup f Set.univ=best) (H : ℕ) (v : Node) : (∃ p ∈ C.nearOptimalNodes f best H, p <+: v) ∨ (v.length<H ∧ v ∈ C.nearOptimalNodes f best v.length) ∨ (∃ h < H, ∃ p ∈ C.boundaryNodes f best h, p <+: v)","missing":[],"search":"node_partition_cover banditrlproof.hoo.regularcovering.node_partition_cover any word either lies below a good depth-h prefix, is itself a shallow near-optimal word, or lies below a first bad child of a near-optimal parent. theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.descendant_nearOptimal_gap","label":"descendant_nearOptimal_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.descendant_nearOptimal_gap","description":"The deep-good and shallow-good regret bounds use Lemma 3 on the actual representative in a descendant region.","url":"../modules/banditrlproof-hoopartition/index.html#decl-a7f0ba1c29da","parent":"module:BanditRLProof.HOOPartition","order":4987,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.descendant_nearOptimal_gap {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) {p v : Node} {h : ℕ} (hp : p ∈ C.nearOptimalNodes f best h) (hv : p <+: v) : best-f (C.toCovering.representative v) ≤ 4*(C.nu1*C.rho^h)","missing":[],"search":"descendant_nearoptimal_gap banditrlproof.hoo.regularcovering.descendant_nearoptimal_gap the deep-good and shallow-good regret bounds use lemma 3 on the actual representative in a descendant region. theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.descendant_boundary_gap","label":"descendant_boundary_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.descendant_boundary_gap","description":"Poor subtrees inherit their parent's gap bound, not the poor child's gap.","url":"../modules/banditrlproof-hoopartition/index.html#decl-ef348cdf83d9","parent":"module:BanditRLProof.HOOPartition","order":4988,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.descendant_boundary_gap {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hw : WeaklyLipschitz f C.ell best) {p v : Node} {h : ℕ} (hp : p ∈ C.boundaryNodes f best h) (hv : p <+: v) : best-f (C.toCovering.representative v) ≤ 4*(C.nu1*C.rho^h)","missing":[],"search":"descendant_boundary_gap banditrlproof.hoo.regularcovering.descendant_boundary_gap poor subtrees inherit their parent's gap bound, not the poor child's gap. theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.prefix_of_prefixes_length","label":"prefix_of_prefixes_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.prefix_of_prefixes_length","description":"private theorem prefix_of_prefixes_length {p q v : Node} (hp : p <+: v) (hq : q <+: v) (hl : p.length≤q.length) : p <+: q","url":"../modules/banditrlproof-hoopartition/index.html#decl-78316dbb09ea","parent":"module:BanditRLProof.HOOPartition","order":4989,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem prefix_of_prefixes_length {p q v : Node} (hp : p <+: v) (hq : q <+: v) (hl : p.length≤q.length) : p <+: q","missing":[],"search":"prefix_of_prefixes_length banditrlproof.hoo.prefix_of_prefixes_length private theorem prefix_of_prefixes_length {p q v : node} (hp : p <+: v) (hq : q <+: v) (hl : p.length≤q.length) : p <+: q theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.deepGoodNodes","label":"deepGoodNodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.deepGoodNodes","description":"def RegularCovering.deepGoodNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Set Node","url":"../modules/banditrlproof-hoopartition/index.html#decl-0bc6b39905d2","parent":"module:BanditRLProof.HOOPartition","order":4990,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def RegularCovering.deepGoodNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Set Node","missing":[],"search":"deepgoodnodes banditrlproof.hoo.regularcovering.deepgoodnodes def regularcovering.deepgoodnodes {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) : set node definition compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.shallowGoodNodes","label":"shallowGoodNodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.shallowGoodNodes","description":"def RegularCovering.shallowGoodNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Set Node","url":"../modules/banditrlproof-hoopartition/index.html#decl-5c6d95be5909","parent":"module:BanditRLProof.HOOPartition","order":4991,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def RegularCovering.shallowGoodNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Set Node","missing":[],"search":"shallowgoodnodes banditrlproof.hoo.regularcovering.shallowgoodnodes def regularcovering.shallowgoodnodes {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) : set node definition compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.badSubtreeNodes","label":"badSubtreeNodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.badSubtreeNodes","description":"def RegularCovering.badSubtreeNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Set Node","url":"../modules/banditrlproof-hoopartition/index.html#decl-4e4f97062543","parent":"module:BanditRLProof.HOOPartition","order":4992,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:147"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def RegularCovering.badSubtreeNodes {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Set Node","missing":[],"search":"badsubtreenodes banditrlproof.hoo.regularcovering.badsubtreenodes def regularcovering.badsubtreenodes {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) : set node definition compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.deep_shallow_disjoint","label":"deep_shallow_disjoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.deep_shallow_disjoint","description":"theorem RegularCovering.deep_shallow_disjoint {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Disjoint (C.deepGoodNodes f best H) (C.shallowGoodNodes f best H)","url":"../modules/banditrlproof-hoopartition/index.html#decl-6c93f8cb23fb","parent":"module:BanditRLProof.HOOPartition","order":4993,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.deep_shallow_disjoint {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (H : ℕ) : Disjoint (C.deepGoodNodes f best H) (C.shallowGoodNodes f best H)","missing":[],"search":"deep_shallow_disjoint banditrlproof.hoo.regularcovering.deep_shallow_disjoint theorem regularcovering.deep_shallow_disjoint {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (h : ℕ) : disjoint (c.deepgoodnodes f best h) (c.shallowgoodnodes f best h) theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.deep_bad_disjoint","label":"deep_bad_disjoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.deep_bad_disjoint","description":"theorem RegularCovering.deep_bad_disjoint {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (H : ℕ) : Disjoint (C.deepGoodNodes f best H) (C.badSubtreeNodes f best H)","url":"../modules/banditrlproof-hoopartition/index.html#decl-ac033f3a430c","parent":"module:BanditRLProof.HOOPartition","order":4994,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.deep_bad_disjoint {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (H : ℕ) : Disjoint (C.deepGoodNodes f best H) (C.badSubtreeNodes f best H)","missing":[],"search":"deep_bad_disjoint banditrlproof.hoo.regularcovering.deep_bad_disjoint theorem regularcovering.deep_bad_disjoint {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (h : ℕ) : disjoint (c.deepgoodnodes f best h) (c.badsubtreenodes f best h) theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.shallow_bad_disjoint","label":"shallow_bad_disjoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.shallow_bad_disjoint","description":"theorem RegularCovering.shallow_bad_disjoint {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (H : ℕ) : Disjoint (C.shallowGoodNodes f best H) (C.badSubtreeNodes f best H)","url":"../modules/banditrlproof-hoopartition/index.html#decl-0c66dfb1f30f","parent":"module:BanditRLProof.HOOPartition","order":4995,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:171"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.shallow_bad_disjoint {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (H : ℕ) : Disjoint (C.shallowGoodNodes f best H) (C.badSubtreeNodes f best H)","missing":[],"search":"shallow_bad_disjoint banditrlproof.hoo.regularcovering.shallow_bad_disjoint theorem regularcovering.shallow_bad_disjoint {x : type*} [measurablespace x] (c : regularcovering x) (f : x → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (h : ℕ) : disjoint (c.shallowgoodnodes f best h) (c.badsubtreenodes f best h) theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.RegularCovering.actual_regret_partition","label":"actual_regret_partition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.RegularCovering.actual_regret_partition","description":"Exact three-part identity for the actual HOO action trace, before any inequality or expectation. Cover and disjointness are derived above.","url":"../modules/banditrlproof-hoopartition/index.html#decl-c4a917a7a5c2","parent":"module:BanditRLProof.HOOPartition","order":4996,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOPartition"],["Source","BanditRLProof/HOOPartition.lean:182"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem RegularCovering.actual_regret_partition {X : Type*} [MeasurableSpace X] (C : RegularCovering X) (f : X → ℝ) (best : ℝ) (hf : ∀x, f x≤best) (hbest : regionSup f Set.univ=best) (H N : ℕ) (Y : ℕ → ℝ) : (∑ n ∈ Finset.range N, (best-f (C.toCovering.arm C.nu1 C.rho Y n))) = (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.deepGoodNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) + (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.shallowGoodNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0) + (∑ n ∈ Finset.range N, if action C.nu1 C.rho Y n ∈ C.badSubtreeNodes f best H then best-f (C.toCovering.arm C.nu1 C.rho Y n) else 0)","missing":[],"search":"actual_regret_partition banditrlproof.hoo.regularcovering.actual_regret_partition exact three-part identity for the actual hoo action trace, before any inequality or expectation. cover and disjointness are derived above. theorem compiled","shard":"modules/29186340eeacbfed.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.selectionFailureBudget","label":"selectionFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HOO.selectionFailureBudget","description":"noncomputable def selectionFailureBudget (n : ℕ) : ℝ","url":"../modules/banditrlproof-hootailsum/index.html#decl-9f3dd8890ad8","parent":"module:BanditRLProof.HOOTailSum","order":4997,"meta":[["Kind","definition"],["Module","BanditRLProof.HOOTailSum"],["Source","BanditRLProof/HOOTailSum.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def selectionFailureBudget (n : ℕ) : ℝ","missing":[],"search":"selectionfailurebudget banditrlproof.hoo.selectionfailurebudget noncomputable def selectionfailurebudget (n : ℕ) : ℝ definition compiled","shard":"modules/122daf1ff2aaebcb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.selection_failure_exp_eq","label":"selection_failure_exp_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.selection_failure_exp_eq","description":"theorem selection_failure_exp_eq (n : ℕ) : Real.exp (-4*Real.log (max (n:ℝ) 2)) = 1/(max (n:ℝ) 2)^4","url":"../modules/banditrlproof-hootailsum/index.html#decl-84e2fd48e777","parent":"module:BanditRLProof.HOOTailSum","order":4998,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOTailSum"],["Source","BanditRLProof/HOOTailSum.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem selection_failure_exp_eq (n : ℕ) : Real.exp (-4*Real.log (max (n:ℝ) 2)) = 1/(max (n:ℝ) 2)^4","missing":[],"search":"selection_failure_exp_eq banditrlproof.hoo.selection_failure_exp_eq theorem selection_failure_exp_eq (n : ℕ) : real.exp (-4*real.log (max (n:ℝ) 2)) = 1/(max (n:ℝ) 2)^4 theorem compiled","shard":"modules/122daf1ff2aaebcb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.selection_failure_le_telescope","label":"selection_failure_le_telescope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.selection_failure_le_telescope","description":"theorem selection_failure_le_telescope (n : ℕ) (hn : 2 ≤ n) : selectionFailureBudget n ≤ 2 * (1/((n:ℝ)-1) - 1/n)","url":"../modules/banditrlproof-hootailsum/index.html#decl-3478b1164428","parent":"module:BanditRLProof.HOOTailSum","order":4999,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOTailSum"],["Source","BanditRLProof/HOOTailSum.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem selection_failure_le_telescope (n : ℕ) (hn : 2 ≤ n) : selectionFailureBudget n ≤ 2 * (1/((n:ℝ)-1) - 1/n)","missing":[],"search":"selection_failure_le_telescope banditrlproof.hoo.selection_failure_le_telescope theorem selection_failure_le_telescope (n : ℕ) (hn : 2 ≤ n) : selectionfailurebudget n ≤ 2 * (1/((n:ℝ)-1) - 1/n) theorem compiled","shard":"modules/122daf1ff2aaebcb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HOO.selection_failure_sum_le_three","label":"selection_failure_sum_le_three","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HOO.selection_failure_sum_le_three","description":"theorem selection_failure_sum_le_three (N : ℕ) : (∑ n ∈ Finset.range N, selectionFailureBudget n) ≤ 3","url":"../modules/banditrlproof-hootailsum/index.html#decl-69099f14ef44","parent":"module:BanditRLProof.HOOTailSum","order":5000,"meta":[["Kind","theorem"],["Module","BanditRLProof.HOOTailSum"],["Source","BanditRLProof/HOOTailSum.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem selection_failure_sum_le_three (N : ℕ) : (∑ n ∈ Finset.range N, selectionFailureBudget n) ≤ 3","missing":[],"search":"selection_failure_sum_le_three banditrlproof.hoo.selection_failure_sum_le_three theorem selection_failure_sum_le_three (n : ℕ) : (∑ n ∈ finset.range n, selectionfailurebudget n) ≤ 3 theorem compiled","shard":"modules/122daf1ff2aaebcb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.arm_coordinate_integral","label":"arm_coordinate_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.arm_coordinate_integral","description":"theorem arm_coordinate_integral (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (i : ℕ) (g : ℝ → ℝ) (hg : Measurable g) : (∫ stream, g (stream i arm) ∂UCB.armStreamMeasure ν) = ∫ x, g x ∂ν arm","url":"../modules/banditrlproof-heavytailarmlaw/index.html#decl-8a298aec1095","parent":"module:BanditRLProof.HeavyTailArmLaw","order":5001,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailArmLaw"],["Source","BanditRLProof/HeavyTailArmLaw.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arm_coordinate_integral (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (i : ℕ) (g : ℝ → ℝ) (hg : Measurable g) : (∫ stream, g (stream i arm) ∂UCB.armStreamMeasure ν) = ∫ x, g x ∂ν arm","missing":[],"search":"arm_coordinate_integral banditrlproof.heavytail.arm_coordinate_integral theorem arm_coordinate_integral (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (i : ℕ) (g : ℝ → ℝ) (hg : measurable g) : (∫ stream, g (stream i arm) ∂ucb.armstreammeasure ν) = ∫ x, g x ∂ν arm theorem compiled","shard":"modules/905d14b915b1c1b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.arm_coordinate_integrable","label":"arm_coordinate_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.arm_coordinate_integrable","description":"theorem arm_coordinate_integrable (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (i : ℕ) (g : ℝ → ℝ) (hg : Integrable g (ν arm)) : Integrable (fun stream => g (stream i arm)) (UCB.armStreamMeasure ν)","url":"../modules/banditrlproof-heavytailarmlaw/index.html#decl-47f7bd71ba87","parent":"module:BanditRLProof.HeavyTailArmLaw","order":5002,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailArmLaw"],["Source","BanditRLProof/HeavyTailArmLaw.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arm_coordinate_integrable (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (i : ℕ) (g : ℝ → ℝ) (hg : Integrable g (ν arm)) : Integrable (fun stream => g (stream i arm)) (UCB.armStreamMeasure ν)","missing":[],"search":"arm_coordinate_integrable banditrlproof.heavytail.arm_coordinate_integrable theorem arm_coordinate_integrable (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (i : ℕ) (g : ℝ → ℝ) (hg : integrable g (ν arm)) : integrable (fun stream => g (stream i arm)) (ucb.armstreammeasure ν) theorem compiled","shard":"modules/905d14b915b1c1b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.arm_adaptive_mean_tail","label":"arm_adaptive_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.arm_adaptive_mean_tail","description":"theorem arm_adaptive_mean_tail (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ confidenceRadius ε u t (coun…","url":"../modules/banditrlproof-heavytailarmlaw/index.html#decl-72884078c203","parent":"module:BanditRLProof.HeavyTailArmLaw","order":5003,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailArmLaw"],["Source","BanditRLProof/HeavyTailArmLaw.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arm_adaptive_mean_tail (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : Integrable (fun x : ℝ => x) (ν arm)) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ confidenceRadius ε u t (count stream) ≤ |(∑ s ∈ Finset.range (count stream), truncate (sampleThreshold ε u t s) (stream s arm)) / count stream - ∫ x, x ∂ν arm|} ≤ t * (2 * Real.exp (-confidenceLog t))","missing":[],"search":"arm_adaptive_mean_tail banditrlproof.heavytail.arm_adaptive_mean_tail theorem arm_adaptive_mean_tail (ν : kernel (fin k) ℝ) [ismarkovkernel ν] (arm : fin k) (count : ucb.armrewardstream k → ℕ) (ε u : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : integrable (fun x : ℝ => x) (ν arm)) (hm : integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) : (ucb.armstreammeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ confidenceradius ε u t (count stream) ≤ |(∑ s ∈ finset.range (count stream), truncate (samplethreshold ε u t s) (stream s arm)) / count stream - ∫ x, x ∂ν arm|} ≤ t * (2 * real.exp (-confidencelog t)) theorem compiled","shard":"modules/905d14b915b1c1b2.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clipped_centered_mgf","label":"clipped_centered_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clipped_centered_mgf","description":"theorem clipped_centered_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (B ε u tilt : ℝ) (hXm : Measurable X) (hB : 0 < B) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) (hu : (∫ ω, |X ω|^(1+ε) ∂μ) ≤ u) (hsmall : |tilt| * (2*B) ≤ 1) : Concentration.HasMGFUpperBoundAt (fun ω => clip B (X ω) - ∫ ω, clip B (X ω) ∂μ) tilt (tilt^2*(u*B^(1-ε))) μ","url":"../modules/banditrlproof-heavytailclippedconfidence/index.html#decl-823553b1df23","parent":"module:BanditRLProof.HeavyTailClippedConfidence","order":5004,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedConfidence"],["Source","BanditRLProof/HeavyTailClippedConfidence.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipped_centered_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (B ε u tilt : ℝ) (hXm : Measurable X) (hB : 0 < B) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) (hu : (∫ ω, |X ω|^(1+ε) ∂μ) ≤ u) (hsmall : |tilt| * (2*B) ≤ 1) : Concentration.HasMGFUpperBoundAt (fun ω => clip B (X ω) - ∫ ω, clip B (X ω) ∂μ) tilt (tilt^2*(u*B^(1-ε))) μ","missing":[],"search":"clipped_centered_mgf banditrlproof.heavytail.clipped_centered_mgf theorem clipped_centered_mgf {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ω → ℝ) (b ε u tilt : ℝ) (hxm : measurable x) (hb : 0 < b) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hm : integrable (fun ω => |x ω|^(1+ε)) μ) (hu : (∫ ω, |x ω|^(1+ε) ∂μ) ≤ u) (hsmall : |tilt| * (2*b) ≤ 1) : concentration.hasmgfupperboundat (fun ω => clip b (x ω) - ∫ ω, clip b (x ω) ∂μ) tilt (tilt^2*(u*b^(1-ε))) μ theorem compiled","shard":"modules/51d6d2fcfa2e68d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clipped_sum_abs_tail","label":"clipped_sum_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clipped_sum_abs_tail","description":"theorem clipped_sum_abs_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L : ℝ) (n : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2*B i ≤ b) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤…","url":"../modules/banditrlproof-heavytailclippedconfidence/index.html#decl-1a4971dbe997","parent":"module:BanditRLProof.HeavyTailClippedConfidence","order":5005,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedConfidence"],["Source","BanditRLProof/HeavyTailClippedConfidence.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipped_sum_abs_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L : ℝ) (n : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2*B i ≤ b) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 2*Real.sqrt ((∑ i ∈ Finset.range n, u*(B i)^(1-ε))*L)+b*L ≤ |∑ i ∈ Finset.range n, (clip (B i) (X i ω) - ∫ ω, clip (B i) (X i ω) ∂μ)|} ≤ 2*Real.exp (-L)","missing":[],"search":"clipped_sum_abs_tail banditrlproof.heavytail.clipped_sum_abs_tail theorem clipped_sum_abs_tail {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (b : ℕ → ℝ) (ε u b l : ℝ) (n : ℕ) (hxm : ∀ i, measurable (x i)) (hi : iindepfun x μ) (hb : ∀ i, 0 < b i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hl : 0 ≤ l) (hbound : ∀ i ∈ finset.range n, 2*b i ≤ b) (hm : ∀ i, integrable (fun ω => |x i ω|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |x i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 2*real.sqrt ((∑ i ∈ finset.range n, u*(b i)^(1-ε))*l)+b*l ≤ |∑ i ∈ finset.range n, (clip (b i) (x i ω) - ∫ ω, clip (b i) (x i ω) ∂μ)|} ≤ 2*real.exp (-l) theorem compiled","shard":"modules/51d6d2fcfa2e68d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clipped_mean_tail","label":"clipped_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clipped_mean_tail","description":"theorem clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2*B i ≤ b) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (h…","url":"../modules/banditrlproof-heavytailclippedconfidence/index.html#decl-bfe0661640bf","parent":"module:BanditRLProof.HeavyTailClippedConfidence","order":5006,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedConfidence"],["Source","BanditRLProof/HeavyTailClippedConfidence.lean:50"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2*B i ≤ b) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | ((∑ i ∈ Finset.range n, u/(B i)^ε) + (2*Real.sqrt ((∑ i ∈ Finset.range n, u*(B i)^(1-ε))*L)+b*L))/n ≤ |prefixMean (fun i => clip (B i) (X i ω)) n - mean|} ≤ 2*Real.exp (-L)","missing":[],"search":"clipped_mean_tail banditrlproof.heavytail.clipped_mean_tail theorem clipped_mean_tail {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (b : ℕ → ℝ) (ε u b l mean : ℝ) (n : ℕ) (hn : 0 < n) (hxm : ∀ i, measurable (x i)) (hi : iindepfun x μ) (hb : ∀ i, 0 < b i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hl : 0 ≤ l) (hbound : ∀ i ∈ finset.range n, 2*b i ≤ b) (hx : ∀ i, integrable (x i) μ) (hmean : ∀ i, (∫ ω, x i ω ∂μ) = mean) (hm : ∀ i, integrable (fun ω => |x i ω|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |x i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | ((∑ i ∈ finset.range n, u/(b i)^ε) + (2*real.sqrt ((∑ i ∈ finset.range n, u*(b i)^(1-ε))*l)+b*l))/n ≤ |prefixmean (fun i => clip (b i) (x i ω)) n - mean|} ≤ 2*real.exp (-l) theorem compiled","shard":"modules/51d6d2fcfa2e68d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clip_eq_self","label":"clip_eq_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clip_eq_self","description":"theorem clip_eq_self (B x : ℝ) (hx : |x| ≤ B) : clip B x = x","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-d69ac72d658f","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5007,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clip_eq_self (B x : ℝ) (hx : |x| ≤ B) : clip B x = x","missing":[],"search":"clip_eq_self banditrlproof.heavytail.clip_eq_self theorem clip_eq_self (b x : ℝ) (hx : |x| ≤ b) : clip b x = x theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_clip_le","label":"abs_clip_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_clip_le","description":"theorem abs_clip_le (B x : ℝ) (hB : 0 ≤ B) : |clip B x| ≤ B","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-40f87f2453ce","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5008,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_clip_le (B x : ℝ) (hB : 0 ≤ B) : |clip B x| ≤ B","missing":[],"search":"abs_clip_le banditrlproof.heavytail.abs_clip_le theorem abs_clip_le (b x : ℝ) (hb : 0 ≤ b) : |clip b x| ≤ b theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_clip_le_abs","label":"abs_clip_le_abs","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_clip_le_abs","description":"theorem abs_clip_le_abs (B x : ℝ) (hB : 0 ≤ B) : |clip B x| ≤ |x|","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-a32b5e995a24","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5009,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_clip_le_abs (B x : ℝ) (hB : 0 ≤ B) : |clip B x| ≤ |x|","missing":[],"search":"abs_clip_le_abs banditrlproof.heavytail.abs_clip_le_abs theorem abs_clip_le_abs (b x : ℝ) (hb : 0 ≤ b) : |clip b x| ≤ |x| theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_sub_clip_le_truncate","label":"abs_sub_clip_le_truncate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_sub_clip_le_truncate","description":"theorem abs_sub_clip_le_truncate (B x : ℝ) (hB : 0 ≤ B) : |x - clip B x| ≤ |x - truncate B x|","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-227a4bb53312","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5010,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_sub_clip_le_truncate (B x : ℝ) (hB : 0 ≤ B) : |x - clip B x| ≤ |x - truncate B x|","missing":[],"search":"abs_sub_clip_le_truncate banditrlproof.heavytail.abs_sub_clip_le_truncate theorem abs_sub_clip_le_truncate (b x : ℝ) (hb : 0 ≤ b) : |x - clip b x| ≤ |x - truncate b x| theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_sub_clip_moment_le","label":"abs_sub_clip_moment_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_sub_clip_moment_le","description":"theorem abs_sub_clip_moment_le (B x ε : ℝ) (hB : 0 < B) (hε : 0 ≤ ε) : |x - clip B x| ≤ |x|^(1+ε) / B^ε","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-1c288ccbe340","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5011,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_sub_clip_moment_le (B x ε : ℝ) (hB : 0 < B) (hε : 0 ≤ ε) : |x - clip B x| ≤ |x|^(1+ε) / B^ε","missing":[],"search":"abs_sub_clip_moment_le banditrlproof.heavytail.abs_sub_clip_moment_le theorem abs_sub_clip_moment_le (b x ε : ℝ) (hb : 0 < b) (hε : 0 ≤ ε) : |x - clip b x| ≤ |x|^(1+ε) / b^ε theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sq_clip_moment_le","label":"sq_clip_moment_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sq_clip_moment_le","description":"theorem sq_clip_moment_le (B x ε : ℝ) (hB : 0 < B) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) : (clip B x)^2 ≤ |x|^(1+ε) * B^(1-ε)","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-279db942854a","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5012,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sq_clip_moment_le (B x ε : ℝ) (hB : 0 < B) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) : (clip B x)^2 ≤ |x|^(1+ε) * B^(1-ε)","missing":[],"search":"sq_clip_moment_le banditrlproof.heavytail.sq_clip_moment_le theorem sq_clip_moment_le (b x ε : ℝ) (hb : 0 < b) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) : (clip b x)^2 ≤ |x|^(1+ε) * b^(1-ε) theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_clip","label":"measurable_clip","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_clip","description":"theorem measurable_clip (B : ℝ) : Measurable (clip B)","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-9cc5b984aecd","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5013,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_clip (B : ℝ) : Measurable (clip B)","missing":[],"search":"measurable_clip banditrlproof.heavytail.measurable_clip theorem measurable_clip (b : ℝ) : measurable (clip b) theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.integral_clip_bias_le","label":"integral_clip_bias_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.integral_clip_bias_le","description":"theorem integral_clip_bias_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (B ε u : ℝ) (hB : 0 < B) (hε : 0 ≤ ε) (hXm : Measurable X) (hX : Integrable X μ) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) (hu : (∫ ω, |X ω|^(1+ε) ∂μ) ≤ u) : |(∫ ω, X ω ∂μ) - ∫ ω, clip B (X ω) ∂μ| ≤ u/B^ε","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-f0ba560ed411","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5014,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_clip_bias_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (B ε u : ℝ) (hB : 0 < B) (hε : 0 ≤ ε) (hXm : Measurable X) (hX : Integrable X μ) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) (hu : (∫ ω, |X ω|^(1+ε) ∂μ) ≤ u) : |(∫ ω, X ω ∂μ) - ∫ ω, clip B (X ω) ∂μ| ≤ u/B^ε","missing":[],"search":"integral_clip_bias_le banditrlproof.heavytail.integral_clip_bias_le theorem integral_clip_bias_le {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ω → ℝ) (b ε u : ℝ) (hb : 0 < b) (hε : 0 ≤ ε) (hxm : measurable x) (hx : integrable x μ) (hm : integrable (fun ω => |x ω|^(1+ε)) μ) (hu : (∫ ω, |x ω|^(1+ε) ∂μ) ≤ u) : |(∫ ω, x ω ∂μ) - ∫ ω, clip b (x ω) ∂μ| ≤ u/b^ε theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.integral_sq_clip_le","label":"integral_sq_clip_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.integral_sq_clip_le","description":"theorem integral_sq_clip_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) (X : Ω → ℝ) (B ε u : ℝ) (hB : 0 < B) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hXm : Measurable X) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) (hu : (∫ ω, |X ω|^(1+ε) ∂μ) ≤ u) : (∫ ω, (clip B (X ω))^2 ∂μ) ≤ u*B^(1-ε)","url":"../modules/banditrlproof-heavytailclippedmoments/index.html#decl-961ea1d2a7ed","parent":"module:BanditRLProof.HeavyTailClippedMoments","order":5015,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedMoments"],["Source","BanditRLProof/HeavyTailClippedMoments.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_clip_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) (X : Ω → ℝ) (B ε u : ℝ) (hB : 0 < B) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hXm : Measurable X) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) (hu : (∫ ω, |X ω|^(1+ε) ∂μ) ≤ u) : (∫ ω, (clip B (X ω))^2 ∂μ) ≤ u*B^(1-ε)","missing":[],"search":"integral_sq_clip_le banditrlproof.heavytail.integral_sq_clip_le theorem integral_sq_clip_le {ω : type*} [measurablespace ω] (μ : measure ω) (x : ω → ℝ) (b ε u : ℝ) (hb : 0 < b) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hxm : measurable x) (hm : integrable (fun ω => |x ω|^(1+ε)) μ) (hu : (∫ ω, |x ω|^(1+ε) ∂μ) ≤ u) : (∫ ω, (clip b (x ω))^2 ∂μ) ≤ u*b^(1-ε) theorem compiled","shard":"modules/e297979e39992423.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.scheduled_clipped_mean_tail","label":"scheduled_clipped_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.scheduled_clipped_mean_tail","description":"theorem scheduled_clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u mean : ℝ) (t n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤…","url":"../modules/banditrlproof-heavytailclippedscheduled/index.html#decl-6bd973705ef2","parent":"module:BanditRLProof.HeavyTailClippedScheduled","order":5016,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedScheduled"],["Source","BanditRLProof/HeavyTailClippedScheduled.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem scheduled_clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u mean : ℝ) (t n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | confidenceRadius ε u t n ≤ |(∑ s ∈ Finset.range n, clip (sampleThreshold ε u t s) (X s ω)) / n - mean|} ≤ 2 * Real.exp (-confidenceLog t)","missing":[],"search":"scheduled_clipped_mean_tail banditrlproof.heavytail.scheduled_clipped_mean_tail theorem scheduled_clipped_mean_tail {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (ε u mean : ℝ) (t n : ℕ) (hn : 0 < n) (hxm : ∀ i, measurable (x i)) (hi : iindepfun x μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : ∀ i, integrable (x i) μ) (hmean : ∀ i, (∫ ω, x i ω ∂μ) = mean) (hm : ∀ i, integrable (fun ω => |x i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |x i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | confidenceradius ε u t n ≤ |(∑ s ∈ finset.range n, clip (samplethreshold ε u t s) (x s ω)) / n - mean|} ≤ 2 * real.exp (-confidencelog t) theorem compiled","shard":"modules/36c1a5551903455f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.scheduled_adaptive_clipped_mean_tail","label":"scheduled_adaptive_clipped_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.scheduled_adaptive_clipped_mean_tail","description":"Union over deterministic prefix sizes. The adaptive count is never asserted IID.","url":"../modules/banditrlproof-heavytailclippedscheduled/index.html#decl-74d67d3a60b4","parent":"module:BanditRLProof.HeavyTailClippedScheduled","order":5017,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedScheduled"],["Source","BanditRLProof/HeavyTailClippedScheduled.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem scheduled_adaptive_clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (count : Ω → ℕ) (ε u mean : ℝ) (t : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | 0 < count ω ∧ count ω ≤ t ∧ confidenceRadius ε u t (count ω) ≤ |(∑ s ∈ Finset.range (count ω), clip (sampleThreshold ε u t s) (X s ω)) / count ω - mean|} ≤ t * (2 * Real.exp (-confidenceLog t))","missing":[],"search":"scheduled_adaptive_clipped_mean_tail banditrlproof.heavytail.scheduled_adaptive_clipped_mean_tail union over deterministic prefix sizes. the adaptive count is never asserted iid. theorem compiled","shard":"modules/36c1a5551903455f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.integrable_of_raw_moment","label":"integrable_of_raw_moment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.integrable_of_raw_moment","description":"theorem integrable_of_raw_moment {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (ε : ℝ) (hXm : Measurable X) (hε : 0 ≤ ε) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) : Integrable X μ","url":"../modules/banditrlproof-heavytailclippedtransfer/index.html#decl-d898a061c933","parent":"module:BanditRLProof.HeavyTailClippedTransfer","order":5018,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedTransfer"],["Source","BanditRLProof/HeavyTailClippedTransfer.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_of_raw_moment {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (ε : ℝ) (hXm : Measurable X) (hε : 0 ≤ ε) (hm : Integrable (fun ω => |X ω|^(1+ε)) μ) : Integrable X μ","missing":[],"search":"integrable_of_raw_moment banditrlproof.heavytail.integrable_of_raw_moment theorem integrable_of_raw_moment {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ω → ℝ) (ε : ℝ) (hxm : measurable x) (hε : 0 ≤ ε) (hm : integrable (fun ω => |x ω|^(1+ε)) μ) : integrable x μ theorem compiled","shard":"modules/e380497bdb656e0d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.adaptive_corrupted_clipped_mean_tail","label":"adaptive_corrupted_clipped_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.adaptive_corrupted_clipped_mean_tail","description":"No clean-confidence or fluctuation premise: these are produced from the independent raw-moment stream. Only the actually consumed prefix is budgeted.","url":"../modules/banditrlproof-heavytailclippedtransfer/index.html#decl-cbe3cb6c9e5a","parent":"module:BanditRLProof.HeavyTailClippedTransfer","order":5019,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedTransfer"],["Source","BanditRLProof/HeavyTailClippedTransfer.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem adaptive_corrupted_clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X c : ℕ → Ω → ℝ) (count : Ω → ℕ) (ε u mean C : ℝ) (t : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) (hC : ∀ ω, (∑ s ∈ Finset.range (count ω), |c s ω|) ≤ C) : μ.real {ω | 0 < count ω ∧ count ω ≤ t ∧ confidenceRadius ε u t (count ω) + C / count ω ≤ |prefixMean (fun s => clip (sampleThreshold ε u t s) (X s ω + c s ω)) (count ω) - mean|} ≤ t * (2 * Real.exp (-confidenceLog t))","missing":[],"search":"adaptive_corrupted_clipped_mean_tail banditrlproof.heavytail.adaptive_corrupted_clipped_mean_tail no clean-confidence or fluctuation premise: these are produced from the independent raw-moment stream. only the actually consumed prefix is budgeted. theorem compiled","shard":"modules/e380497bdb656e0d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.observed_corrupted_clipped_mean_tail","label":"observed_corrupted_clipped_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.observed_corrupted_clipped_mean_tail","description":"Confidence for the actual clipped observations along an arbitrary action trace, including a policy driven by corrupted observations. This compares clean and corrupted rewards along that same trace, not two different policies.","url":"../modules/banditrlproof-heavytailclippedtransfer/index.html#decl-c4f4c692631a","parent":"module:BanditRLProof.HeavyTailClippedTransfer","order":5020,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedTransfer"],["Source","BanditRLProof/HeavyTailClippedTransfer.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem observed_corrupted_clipped_mean_tail {Ω : Type*} [MeasurableSpace Ω] {K : ℕ} (μ : Measure Ω) [IsProbabilityMeasure μ] (action : Ω → ActionTrace (Fin K)) (stream corruption : Ω → UCB.ArmRewardStream K) (arm : Fin K) (ε u mean C : ℝ) (t : ℕ) (hXm : ∀ i, Measurable (fun ω => stream ω i arm)) (hi : iIndepFun (fun i ω => stream ω i arm) μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hmean : ∀ i, (∫ ω, stream ω i arm ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |stream ω i arm|^(1+ε)) μ) (hu : ∀ i, (∫ ω, |stream ω i arm|^(1+ε) ∂μ) ≤ u) (hC : ∀ ω, (∑ s ∈ Finset.range (pullCount (action ω) arm t), |corruption ω s arm|) ≤ C) : μ.real {ω | 0 < pullCount (action ω) arm t ∧ pullCount (action ω) arm t ≤ t ∧ confidenceRadius ε u t (pullCount (action ω) arm t) + C / pullCount (action ω) arm t ≤ |sumRewards (action ω) (fun s => clip (sampleThreshold ε u t (pullCount (action ω) (action ω s) s)) (UC…","missing":[],"search":"observed_corrupted_clipped_mean_tail banditrlproof.heavytail.observed_corrupted_clipped_mean_tail confidence for the actual clipped observations along an arbitrary action trace, including a policy driven by corrupted observations. this compares clean and corrupted rewards along that same trace, not two different policies. theorem compiled","shard":"modules/e380497bdb656e0d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.arm_corrupted_clipped_mean_tail","label":"arm_corrupted_clipped_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.arm_corrupted_clipped_mean_tail","description":"Stationary arm laws supply the coordinate independence and moments for the reserved transfer endpoint. No confidence bound is supplied by the caller.","url":"../modules/banditrlproof-heavytailclippedtransfer/index.html#decl-7201a31fc91a","parent":"module:BanditRLProof.HeavyTailClippedTransfer","order":5021,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClippedTransfer"],["Source","BanditRLProof/HeavyTailClippedTransfer.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arm_corrupted_clipped_mean_tail {K : ℕ} (ν : Kernel (Fin K) ℝ) [IsMarkovKernel ν] (arm : Fin K) (count : UCB.ArmRewardStream K → ℕ) (c : ℕ → UCB.ArmRewardStream K → ℝ) (ε u C : ℝ) (t : ℕ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hm : Integrable (fun x : ℝ => |x|^(1+ε)) (ν arm)) (hu : (∫ x, |x|^(1+ε) ∂ν arm) ≤ u) (hC : ∀ stream, (∑ s ∈ Finset.range (count stream), |c s stream|) ≤ C) : (UCB.armStreamMeasure ν).real {stream | 0 < count stream ∧ count stream ≤ t ∧ confidenceRadius ε u t (count stream) + C / count stream ≤ |prefixMean (fun s => clip (sampleThreshold ε u t s) (stream s arm + c s stream)) (count stream) - ∫ x, x ∂ν arm|} ≤ t * (2 * Real.exp (-confidenceLog t))","missing":[],"search":"arm_corrupted_clipped_mean_tail banditrlproof.heavytail.arm_corrupted_clipped_mean_tail stationary arm laws supply the coordinate independence and moments for the reserved transfer endpoint. no confidence bound is supplied by the caller. theorem compiled","shard":"modules/e380497bdb656e0d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clip","label":"clip","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clip","description":"noncomputable def clip (B x : ℝ) : ℝ","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-478cf2399a8b","parent":"module:BanditRLProof.HeavyTailClipping","order":5022,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def clip (B x : ℝ) : ℝ","missing":[],"search":"clip banditrlproof.heavytail.clip noncomputable def clip (b x : ℝ) : ℝ definition compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_clip_sub_clip_le","label":"abs_clip_sub_clip_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_clip_sub_clip_le","description":"theorem abs_clip_sub_clip_le (B x y : ℝ) : |clip B x - clip B y| ≤ |x - y|","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-29a7f1507c06","parent":"module:BanditRLProof.HeavyTailClipping","order":5023,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_clip_sub_clip_le (B x y : ℝ) : |clip B x - clip B y| ≤ |x - y|","missing":[],"search":"abs_clip_sub_clip_le banditrlproof.heavytail.abs_clip_sub_clip_le theorem abs_clip_sub_clip_le (b x y : ℝ) : |clip b x - clip b y| ≤ |x - y| theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clip_corruption_le","label":"clip_corruption_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clip_corruption_le","description":"Unlike hard truncation, clipping is stable even when corruption crosses B.","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-735fc71df8b1","parent":"module:BanditRLProof.HeavyTailClipping","order":5024,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clip_corruption_le (B x c : ℝ) : |clip B (x + c) - clip B x| ≤ |c|","missing":[],"search":"clip_corruption_le banditrlproof.heavytail.clip_corruption_le unlike hard truncation, clipping is stable even when corruption crosses b. theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.prefixMean","label":"prefixMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.prefixMean","description":"noncomputable def prefixMean (X : ℕ → ℝ) (n : ℕ) : ℝ","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-2b486e4f5ac7","parent":"module:BanditRLProof.HeavyTailClipping","order":5025,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def prefixMean (X : ℕ → ℝ) (n : ℕ) : ℝ","missing":[],"search":"prefixmean banditrlproof.heavytail.prefixmean noncomputable def prefixmean (x : ℕ → ℝ) (n : ℕ) : ℝ definition compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clipped_prefix_corruption_le","label":"clipped_prefix_corruption_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clipped_prefix_corruption_le","description":"The finite-prefix corruption producer allows index-dependent clipping.","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-dbca7e467bb9","parent":"module:BanditRLProof.HeavyTailClipping","order":5026,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipped_prefix_corruption_le (X c B : ℕ → ℝ) (n : ℕ) (C : ℝ) (hC : (∑ s ∈ Finset.range n, |c s|) ≤ C) : |prefixMean (fun s => clip (B s) (X s + c s)) n - prefixMean (fun s => clip (B s) (X s)) n| ≤ C / n","missing":[],"search":"clipped_prefix_corruption_le banditrlproof.heavytail.clipped_prefix_corruption_le the finite-prefix corruption producer allows index-dependent clipping. theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.clipped_observed_prefix","label":"clipped_observed_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.clipped_observed_prefix","description":"Actual adaptively consumed clipped observations retain the same prefix. The corruption stream can be arbitrary; this is a pathwise statement.","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-5b511752e376","parent":"module:BanditRLProof.HeavyTailClipping","order":5027,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipped_observed_prefix {Ω : Type*} {K : ℕ} (action : Ω → ActionTrace (Fin K)) (stream corruption : Ω → UCB.ArmRewardStream K) (B : ℕ → Fin K → ℝ) (ω : Ω) (arm : Fin K) (t : ℕ) : sumRewards (action ω) (fun s => clip (B (pullCount (action ω) (action ω s) s) (action ω s)) (UCB.rewardFromArmStream action (fun ω j a => stream ω j a + corruption ω j a) ω s)) arm t = ∑ j ∈ Finset.range (pullCount (action ω) arm t), clip (B j arm) (stream ω j arm + corruption ω j arm)","missing":[],"search":"clipped_observed_prefix banditrlproof.heavytail.clipped_observed_prefix actual adaptively consumed clipped observations retain the same prefix. the corruption stream can be arbitrary; this is a pathwise statement. theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.corrupted_clipped_estimator_error_le","label":"corrupted_clipped_estimator_error_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.corrupted_clipped_estimator_error_le","description":"Combine the new corruption producer with the common error assembly. Bias and fluctuation are still supplied; no concentration is asserted here.","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-286a76e910dc","parent":"module:BanditRLProof.HeavyTailClipping","order":5028,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem corrupted_clipped_estimator_error_le (X c B : ℕ → ℝ) (n : ℕ) (C center mean bias fluctuation : ℝ) (hC : (∑ s ∈ Finset.range n, |c s|) ≤ C) (hb : |center - mean| ≤ bias) (hf : |prefixMean (fun s => clip (B s) (X s)) n - center| ≤ fluctuation) : |prefixMean (fun s => clip (B s) (X s + c s)) n - mean| ≤ bias + fluctuation + C / n","missing":[],"search":"corrupted_clipped_estimator_error_le banditrlproof.heavytail.corrupted_clipped_estimator_error_le combine the new corruption producer with the common error assembly. bias and fluctuation are still supplied; no concentration is asserted here. theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.actual_clipped_corruption_le","label":"actual_clipped_corruption_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.actual_clipped_corruption_le","description":"Full pathwise transfer to the actually observed estimator. Both reward streams are compared along the SAME action trace; this does not compare two policies whose actions changed in response to the corruption.","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-20db797bfcad","parent":"module:BanditRLProof.HeavyTailClipping","order":5029,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem actual_clipped_corruption_le {Ω : Type*} {K : ℕ} (action : Ω → ActionTrace (Fin K)) (stream corruption : Ω → UCB.ArmRewardStream K) (B : ℕ → Fin K → ℝ) (ω : Ω) (arm : Fin K) (t : ℕ) (C : ℝ) (hC : (∑ j ∈ Finset.range (pullCount (action ω) arm t), |corruption ω j arm|) ≤ C) : |sumRewards (action ω) (fun s => clip (B (pullCount (action ω) (action ω s) s) (action ω s)) (UCB.rewardFromArmStream action (fun ω j a => stream ω j a + corruption ω j a) ω s)) arm t / pullCount (action ω) arm t - sumRewards (action ω) (fun s => clip (B (pullCount (action ω) (action ω s) s) (action ω s)) (UCB.rewardFromArmStream action stream ω s)) arm t / pullCount (action ω) arm t| ≤ C / pullCount (action ω) arm t","missing":[],"search":"actual_clipped_corruption_le banditrlproof.heavytail.actual_clipped_corruption_le full pathwise transfer to the actually observed estimator. both reward streams are compared along the same action trace; this does not compare two policies whose actions changed in response to the corruption. theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncate_not_unit_lipschitz","label":"truncate_not_unit_lipschitz","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncate_not_unit_lipschitz","description":"A concrete discontinuity diagnostic: hard truncation cannot use the same unit-Lipschitz corruption producer.","url":"../modules/banditrlproof-heavytailclipping/index.html#decl-fc8cdfdd941d","parent":"module:BanditRLProof.HeavyTailClipping","order":5030,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailClipping"],["Source","BanditRLProof/HeavyTailClipping.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncate_not_unit_lipschitz : ¬ (∀ x y : ℝ, |truncate 1 x - truncate 1 y| ≤ |x - y|)","missing":[],"search":"truncate_not_unit_lipschitz banditrlproof.heavytail.truncate_not_unit_lipschitz a concrete discontinuity diagnostic: hard truncation cannot use the same unit-lipschitz corruption producer. theorem compiled","shard":"modules/403e67ada2edd0be.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.exists_variance_tilt","label":"exists_variance_tilt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.exists_variance_tilt","description":"theorem exists_variance_tilt (V b L : ℝ) (hV : 0 ≤ V) (hb : 0 < b) (hL : 0 ≤ L) : ∃ t : ℝ, 0 ≤ t ∧ t * b ≤ 1 ∧ -t * (2 * Real.sqrt (V * L) + b * L) + t^2 * V ≤ -L","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-29aa039b9fbc","parent":"module:BanditRLProof.HeavyTailConfidence","order":5031,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_variance_tilt (V b L : ℝ) (hV : 0 ≤ V) (hb : 0 < b) (hL : 0 ≤ L) : ∃ t : ℝ, 0 ≤ t ∧ t * b ≤ 1 ∧ -t * (2 * Real.sqrt (V * L) + b * L) + t^2 * V ≤ -L","missing":[],"search":"exists_variance_tilt banditrlproof.heavytail.exists_variance_tilt theorem exists_variance_tilt (v b l : ℝ) (hv : 0 ≤ v) (hb : 0 < b) (hl : 0 ≤ l) : ∃ t : ℝ, 0 ≤ t ∧ t * b ≤ 1 ∧ -t * (2 * real.sqrt (v * l) + b * l) + t^2 * v ≤ -l theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.neg_mgf","label":"neg_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.neg_mgf","description":"theorem neg_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) (X : Ω → ℝ) (t v : ℝ) (h : Concentration.HasMGFUpperBoundAt X (-t) v μ) : Concentration.HasMGFUpperBoundAt (fun ω => -X ω) t v μ","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-ec8a7f3ef2d1","parent":"module:BanditRLProof.HeavyTailConfidence","order":5032,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem neg_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) (X : Ω → ℝ) (t v : ℝ) (h : Concentration.HasMGFUpperBoundAt X (-t) v μ) : Concentration.HasMGFUpperBoundAt (fun ω => -X ω) t v μ","missing":[],"search":"neg_mgf banditrlproof.heavytail.neg_mgf theorem neg_mgf {ω : type*} [measurablespace ω] (μ : measure ω) (x : ω → ℝ) (t v : ℝ) (h : concentration.hasmgfupperboundat x (-t) v μ) : concentration.hasmgfupperboundat (fun ω => -x ω) t v μ theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.fixed_mgf_abs_tail","label":"fixed_mgf_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.fixed_mgf_abs_tail","description":"theorem fixed_mgf_abs_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (V b L : ℝ) (hV : 0 ≤ V) (hb : 0 < b) (hL : 0 ≤ L) (h : ∀ t : ℝ, |t| * b ≤ 1 → Concentration.HasMGFUpperBoundAt X t (t^2 * V) μ) : μ.real {ω | 2 * Real.sqrt (V * L) + b * L ≤ |X ω|} ≤ 2 * Real.exp (-L)","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-deaeec0d4536","parent":"module:BanditRLProof.HeavyTailConfidence","order":5033,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem fixed_mgf_abs_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (V b L : ℝ) (hV : 0 ≤ V) (hb : 0 < b) (hL : 0 ≤ L) (h : ∀ t : ℝ, |t| * b ≤ 1 → Concentration.HasMGFUpperBoundAt X t (t^2 * V) μ) : μ.real {ω | 2 * Real.sqrt (V * L) + b * L ≤ |X ω|} ≤ 2 * Real.exp (-L)","missing":[],"search":"fixed_mgf_abs_tail banditrlproof.heavytail.fixed_mgf_abs_tail theorem fixed_mgf_abs_tail {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ω → ℝ) (v b l : ℝ) (hv : 0 ≤ v) (hb : 0 < b) (hl : 0 ≤ l) (h : ∀ t : ℝ, |t| * b ≤ 1 → concentration.hasmgfupperboundat x t (t^2 * v) μ) : μ.real {ω | 2 * real.sqrt (v * l) + b * l ≤ |x ω|} ≤ 2 * real.exp (-l) theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncated_sum_abs_tail","label":"truncated_sum_abs_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncated_sum_abs_tail","description":"Two-sided centered sum, with variance budget and maximum increment bound derived from the raw moments and deterministic truncation thresholds.","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-429722cdee5f","parent":"module:BanditRLProof.HeavyTailConfidence","order":5034,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncated_sum_abs_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L : ℝ) (n : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2 * B i ≤ b) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | 2 * Real.sqrt ((∑ i ∈ Finset.range n, u * (B i)^(1-ε)) * L) + b * L ≤ |∑ i ∈ Finset.range n, (truncate (B i) (X i ω) - ∫ ω, truncate (B i) (X i ω) ∂μ)|} ≤ 2 * Real.exp (-L)","missing":[],"search":"truncated_sum_abs_tail banditrlproof.heavytail.truncated_sum_abs_tail two-sided centered sum, with variance budget and maximum increment bound derived from the raw moments and deterministic truncation thresholds. theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sum_mean_tail_of_centered","label":"sum_mean_tail_of_centered","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sum_mean_tail_of_centered","description":"Shared bias-plus-fluctuation assembly for independent transformed estimators. The truncation and clipping producers discharge both premises separately.","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-6b77972c47ff","parent":"module:BanditRLProof.HeavyTailConfidence","order":5035,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_mean_tail_of_centered {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (Y : ℕ → Ω → ℝ) (mean bias fluctuation δ : ℝ) (n : ℕ) (hbias : |∑ i ∈ Finset.range n, ((∫ ω, Y i ω ∂μ) - mean)| ≤ bias) (htail : μ.real {ω | fluctuation ≤ |∑ i ∈ Finset.range n, (Y i ω - ∫ ω, Y i ω ∂μ)|} ≤ δ) : μ.real {ω | bias + fluctuation ≤ |(∑ i ∈ Finset.range n, Y i ω) - n*mean|} ≤ δ","missing":[],"search":"sum_mean_tail_of_centered banditrlproof.heavytail.sum_mean_tail_of_centered shared bias-plus-fluctuation assembly for independent transformed estimators. the truncation and clipping producers discharge both premises separately. theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncated_sum_mean_tail","label":"truncated_sum_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncated_sum_mean_tail","description":"The bias is produced from the same raw moment hypotheses as the fluctuation. The common mean is a distributional assumption, not a confidence assumption.","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-ca7e16d905a3","parent":"module:BanditRLProof.HeavyTailConfidence","order":5036,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncated_sum_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L mean : ℝ) (n : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2 * B i ≤ b) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | (∑ i ∈ Finset.range n, u / (B i)^ε) + (2 * Real.sqrt ((∑ i ∈ Finset.range n, u * (B i)^(1-ε)) * L) + b * L) ≤ |(∑ i ∈ Finset.range n, truncate (B i) (X i ω)) - n * mean|} ≤ 2 * Real.exp (-L)","missing":[],"search":"truncated_sum_mean_tail banditrlproof.heavytail.truncated_sum_mean_tail the bias is produced from the same raw moment hypotheses as the fluctuation. the common mean is a distributional assumption, not a confidence assumption. theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncated_mean_tail","label":"truncated_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncated_mean_tail","description":"Fixed positive sample-size confidence for the actual truncated empirical mean.","url":"../modules/banditrlproof-heavytailconfidence/index.html#decl-de32a96e15e8","parent":"module:BanditRLProof.HeavyTailConfidence","order":5037,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailConfidence"],["Source","BanditRLProof/HeavyTailConfidence.lean:137"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncated_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u b L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 ≤ u) (hb : 0 < b) (hL : 0 ≤ L) (hbound : ∀ i ∈ Finset.range n, 2 * B i ≤ b) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | ((∑ i ∈ Finset.range n, u / (B i)^ε) + (2 * Real.sqrt ((∑ i ∈ Finset.range n, u * (B i)^(1-ε)) * L) + b * L)) / n ≤ |(∑ i ∈ Finset.range n, truncate (B i) (X i ω)) / n - mean|} ≤ 2 * Real.exp (-L)","missing":[],"search":"truncated_mean_tail banditrlproof.heavytail.truncated_mean_tail fixed positive sample-size confidence for the actual truncated empirical mean. theorem compiled","shard":"modules/0f956effc6634cf3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.bounded_centered_mgf","label":"bounded_centered_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.bounded_centered_mgf","description":"theorem bounded_centered_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (b v tilt : ℝ) (hXm : Measurable X) (hb : ∀ ω, |X ω| ≤ b) (hmean : (∫ ω, X ω ∂μ) = 0) (hv : (∫ ω, (X ω)^2 ∂μ) ≤ v) (hsmall : |tilt| * b ≤ 1) : Concentration.HasMGFUpperBoundAt X tilt (tilt^2 * v) μ","url":"../modules/banditrlproof-heavytailfixedtilt/index.html#decl-f378cdfda04f","parent":"module:BanditRLProof.HeavyTailFixedTilt","order":5038,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailFixedTilt"],["Source","BanditRLProof/HeavyTailFixedTilt.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bounded_centered_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (b v tilt : ℝ) (hXm : Measurable X) (hb : ∀ ω, |X ω| ≤ b) (hmean : (∫ ω, X ω ∂μ) = 0) (hv : (∫ ω, (X ω)^2 ∂μ) ≤ v) (hsmall : |tilt| * b ≤ 1) : Concentration.HasMGFUpperBoundAt X tilt (tilt^2 * v) μ","missing":[],"search":"bounded_centered_mgf banditrlproof.heavytail.bounded_centered_mgf theorem bounded_centered_mgf {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ω → ℝ) (b v tilt : ℝ) (hxm : measurable x) (hb : ∀ ω, |x ω| ≤ b) (hmean : (∫ ω, x ω ∂μ) = 0) (hv : (∫ ω, (x ω)^2 ∂μ) ≤ v) (hsmall : |tilt| * b ≤ 1) : concentration.hasmgfupperboundat x tilt (tilt^2 * v) μ theorem compiled","shard":"modules/d2e8745d536fa2e6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.bounded_centering_mgf","label":"bounded_centering_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.bounded_centering_mgf","description":"Shared centering producer for bounded transformed rewards. Both hard truncation and clipping supply their own second-moment proofs to this interface.","url":"../modules/banditrlproof-heavytailfixedtilt/index.html#decl-472dcbe76dd1","parent":"module:BanditRLProof.HeavyTailFixedTilt","order":5039,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailFixedTilt"],["Source","BanditRLProof/HeavyTailFixedTilt.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bounded_centering_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (Y : Ω → ℝ) (B v tilt : ℝ) (hYm : Measurable Y) (hbound : ∀ ω, |Y ω| ≤ B) (hv : (∫ ω, (Y ω)^2 ∂μ) ≤ v) (hsmall : |tilt| * (2 * B) ≤ 1) : Concentration.HasMGFUpperBoundAt (fun ω => Y ω - ∫ ω, Y ω ∂μ) tilt (tilt^2 * v) μ","missing":[],"search":"bounded_centering_mgf banditrlproof.heavytail.bounded_centering_mgf shared centering producer for bounded transformed rewards. both hard truncation and clipping supply their own second-moment proofs to this interface. theorem compiled","shard":"modules/d2e8745d536fa2e6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncated_centered_mgf","label":"truncated_centered_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncated_centered_mgf","description":"Moment assumptions produce a centered truncated MGF at every admissible tilt. Signed rewards require the factor 2 in the centering range.","url":"../modules/banditrlproof-heavytailfixedtilt/index.html#decl-e92e15b23fea","parent":"module:BanditRLProof.HeavyTailFixedTilt","order":5040,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailFixedTilt"],["Source","BanditRLProof/HeavyTailFixedTilt.lean:87"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncated_centered_mgf {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : Ω → ℝ) (B ε u tilt : ℝ) (hXm : Measurable X) (hB : 0 < B) (hε : ε ≤ 1) (hm : Integrable (fun ω => |X ω| ^ (1 + ε)) μ) (hu : (∫ ω, |X ω| ^ (1 + ε) ∂μ) ≤ u) (hsmall : |tilt| * (2 * B) ≤ 1) : Concentration.HasMGFUpperBoundAt (fun ω => truncate B (X ω) - ∫ ω, truncate B (X ω) ∂μ) tilt (tilt^2 * (u * B^(1-ε))) μ","missing":[],"search":"truncated_centered_mgf banditrlproof.heavytail.truncated_centered_mgf moment assumptions produce a centered truncated mgf at every admissible tilt. signed rewards require the factor 2 in the centering range. theorem compiled","shard":"modules/d2e8745d536fa2e6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.independent_sum_mgf","label":"independent_sum_mgf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.independent_sum_mgf","description":"Independent composition preserves the individual admissible tilt, instead of imposing a range bound on the whole sum.","url":"../modules/banditrlproof-heavytailfixedtilt/index.html#decl-ff200292133a","parent":"module:BanditRLProof.HeavyTailFixedTilt","order":5041,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailFixedTilt"],["Source","BanditRLProof/HeavyTailFixedTilt.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem independent_sum_mgf {Ω ι : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ι → Ω → ℝ) (s : Finset ι) (tilt : ℝ) (budget : ι → ℝ) (hi : iIndepFun X μ) (hm : ∀ i, Measurable (X i)) (h : ∀ i ∈ s, Concentration.HasMGFUpperBoundAt (X i) tilt (budget i) μ) : Concentration.HasMGFUpperBoundAt (∑ i ∈ s, X i) tilt (∑ i ∈ s, budget i) μ","missing":[],"search":"independent_sum_mgf banditrlproof.heavytail.independent_sum_mgf independent composition preserves the individual admissible tilt, instead of imposing a range bound on the whole sum. theorem compiled","shard":"modules/d2e8745d536fa2e6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncated_sum_tail","label":"truncated_sum_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncated_sum_tail","description":"Fixed-prefix one-sided concentration from independent raw-moment data. This is an actual tail producer, before bias assembly and adaptive-count peeling.","url":"../modules/banditrlproof-heavytailfixedtilt/index.html#decl-b9319a403d3d","parent":"module:BanditRLProof.HeavyTailFixedTilt","order":5042,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailFixedTilt"],["Source","BanditRLProof/HeavyTailFixedTilt.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncated_sum_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (B : ℕ → ℝ) (ε u tilt r : ℝ) (n : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hB : ∀ i, 0 < B i) (hε : ε ≤ 1) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) (ht : 0 ≤ tilt) (hsmall : ∀ i ∈ Finset.range n, |tilt| * (2 * B i) ≤ 1) : μ.real {ω | r ≤ ∑ i ∈ Finset.range n, (truncate (B i) (X i ω) - ∫ ω, truncate (B i) (X i ω) ∂μ)} ≤ Real.exp (-tilt * r + ∑ i ∈ Finset.range n, tilt^2 * (u * (B i)^(1-ε)))","missing":[],"search":"truncated_sum_tail banditrlproof.heavytail.truncated_sum_tail fixed-prefix one-sided concentration from independent raw-moment data. this is an actual tail producer, before bias assembly and adaptive-count peeling. theorem compiled","shard":"modules/d2e8745d536fa2e6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.gapThreshold","label":"gapThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.gapThreshold","description":"noncomputable def gapThreshold (ε u gap : ℝ) (T : ℕ) : ℕ","url":"../modules/banditrlproof-heavytailgapthreshold/index.html#decl-0993164d3e8e","parent":"module:BanditRLProof.HeavyTailGapThreshold","order":5043,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailGapThreshold"],["Source","BanditRLProof/HeavyTailGapThreshold.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapThreshold (ε u gap : ℝ) (T : ℕ) : ℕ","missing":[],"search":"gapthreshold banditrlproof.heavytail.gapthreshold noncomputable def gapthreshold (ε u gap : ℝ) (t : ℕ) : ℕ definition compiled","shard":"modules/5cdaeedbc5ec6757.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.gapThreshold_pos","label":"gapThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.gapThreshold_pos","description":"theorem gapThreshold_pos (ε u gap : ℝ) (T : ℕ) : 0 < gapThreshold ε u gap T","url":"../modules/banditrlproof-heavytailgapthreshold/index.html#decl-33d382201847","parent":"module:BanditRLProof.HeavyTailGapThreshold","order":5044,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailGapThreshold"],["Source","BanditRLProof/HeavyTailGapThreshold.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_pos (ε u gap : ℝ) (T : ℕ) : 0 < gapThreshold ε u gap T","missing":[],"search":"gapthreshold_pos banditrlproof.heavytail.gapthreshold_pos theorem gapthreshold_pos (ε u gap : ℝ) (t : ℕ) : 0 < gapthreshold ε u gap t theorem compiled","shard":"modules/5cdaeedbc5ec6757.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.confidenceLog_mono","label":"confidenceLog_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.confidenceLog_mono","description":"theorem confidenceLog_mono {t T : ℕ} (ht : t ≤ T) : confidenceLog t ≤ confidenceLog T","url":"../modules/banditrlproof-heavytailgapthreshold/index.html#decl-146c3aff22b9","parent":"module:BanditRLProof.HeavyTailGapThreshold","order":5045,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailGapThreshold"],["Source","BanditRLProof/HeavyTailGapThreshold.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem confidenceLog_mono {t T : ℕ} (ht : t ≤ T) : confidenceLog t ≤ confidenceLog T","missing":[],"search":"confidencelog_mono banditrlproof.heavytail.confidencelog_mono theorem confidencelog_mono {t t : ℕ} (ht : t ≤ t) : confidencelog t ≤ confidencelog t theorem compiled","shard":"modules/5cdaeedbc5ec6757.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.twice_radius_lt_gap","label":"twice_radius_lt_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.twice_radius_lt_gap","description":"theorem twice_radius_lt_gap (ε u gap : ℝ) (hε : 0 < ε) (hu : 0 < u) (hg : 0 < gap) (t T n : ℕ) (ht : t ≤ T) (hn : gapThreshold ε u gap T ≤ n) : 2 * confidenceRadius ε u t n < gap","url":"../modules/banditrlproof-heavytailgapthreshold/index.html#decl-c413b3046295","parent":"module:BanditRLProof.HeavyTailGapThreshold","order":5046,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailGapThreshold"],["Source","BanditRLProof/HeavyTailGapThreshold.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem twice_radius_lt_gap (ε u gap : ℝ) (hε : 0 < ε) (hu : 0 < u) (hg : 0 < gap) (t T n : ℕ) (ht : t ≤ T) (hn : gapThreshold ε u gap T ≤ n) : 2 * confidenceRadius ε u t n < gap","missing":[],"search":"twice_radius_lt_gap banditrlproof.heavytail.twice_radius_lt_gap theorem twice_radius_lt_gap (ε u gap : ℝ) (hε : 0 < ε) (hu : 0 < u) (hg : 0 < gap) (t t n : ℕ) (ht : t ≤ t) (hn : gapthreshold ε u gap t ≤ n) : 2 * confidenceradius ε u t n < gap theorem compiled","shard":"modules/5cdaeedbc5ec6757.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.rpow_increment_lower","label":"rpow_increment_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.rpow_increment_lower","description":"theorem rpow_increment_lower (x a : ℝ) (hx : 0 ≤ x) (ha : 0 ≤ a) (ha1 : a ≤ 1) : a * (x+1)^(a-1) ≤ (x+1)^a - x^a","url":"../modules/banditrlproof-heavytailpowersum/index.html#decl-cd4ab3abdc6e","parent":"module:BanditRLProof.HeavyTailPowerSum","order":5047,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailPowerSum"],["Source","BanditRLProof/HeavyTailPowerSum.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rpow_increment_lower (x a : ℝ) (hx : 0 ≤ x) (ha : 0 ≤ a) (ha1 : a ≤ 1) : a * (x+1)^(a-1) ≤ (x+1)^a - x^a","missing":[],"search":"rpow_increment_lower banditrlproof.heavytail.rpow_increment_lower theorem rpow_increment_lower (x a : ℝ) (hx : 0 ≤ x) (ha : 0 ≤ a) (ha1 : a ≤ 1) : a * (x+1)^(a-1) ≤ (x+1)^a - x^a theorem compiled","shard":"modules/f7af9e80a1514c97.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sum_shifted_rpow_le","label":"sum_shifted_rpow_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sum_shifted_rpow_le","description":"theorem sum_shifted_rpow_le (a : ℝ) (ha : 0 < a) (ha1 : a ≤ 1) (n : ℕ) : (∑ s ∈ Finset.range n, ((s : ℝ)+1)^(a-1)) ≤ (n : ℝ)^a / a","url":"../modules/banditrlproof-heavytailpowersum/index.html#decl-f130e201e385","parent":"module:BanditRLProof.HeavyTailPowerSum","order":5048,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailPowerSum"],["Source","BanditRLProof/HeavyTailPowerSum.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_shifted_rpow_le (a : ℝ) (ha : 0 < a) (ha1 : a ≤ 1) (n : ℕ) : (∑ s ∈ Finset.range n, ((s : ℝ)+1)^(a-1)) ≤ (n : ℝ)^a / a","missing":[],"search":"sum_shifted_rpow_le banditrlproof.heavytail.sum_shifted_rpow_le theorem sum_shifted_rpow_le (a : ℝ) (ha : 0 < a) (ha1 : a ≤ 1) (n : ℕ) : (∑ s ∈ finset.range n, ((s : ℝ)+1)^(a-1)) ≤ (n : ℝ)^a / a theorem compiled","shard":"modules/f7af9e80a1514c97.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.scheduled_mean_tail","label":"scheduled_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.scheduled_mean_tail","description":"theorem scheduled_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u mean : ℝ) (t n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.r…","url":"../modules/banditrlproof-heavytailscheduledconfidence/index.html#decl-a6d7e9700cc9","parent":"module:BanditRLProof.HeavyTailScheduledConfidence","order":5049,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailScheduledConfidence"],["Source","BanditRLProof/HeavyTailScheduledConfidence.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem scheduled_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u mean : ℝ) (t n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | confidenceRadius ε u t n ≤ |(∑ s ∈ Finset.range n, truncate (sampleThreshold ε u t s) (X s ω)) / n - mean|} ≤ 2 * Real.exp (-confidenceLog t)","missing":[],"search":"scheduled_mean_tail banditrlproof.heavytail.scheduled_mean_tail theorem scheduled_mean_tail {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (ε u mean : ℝ) (t n : ℕ) (hn : 0 < n) (hxm : ∀ i, measurable (x i)) (hi : iindepfun x μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hx : ∀ i, integrable (x i) μ) (hmean : ∀ i, (∫ ω, x i ω ∂μ) = mean) (hm : ∀ i, integrable (fun ω => |x i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |x i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | confidenceradius ε u t n ≤ |(∑ s ∈ finset.range n, truncate (samplethreshold ε u t s) (x s ω)) / n - mean|} ≤ 2 * real.exp (-confidencelog t) theorem compiled","shard":"modules/7e10c23dd66ba6aa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.scheduled_adaptive_mean_tail","label":"scheduled_adaptive_mean_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.scheduled_adaptive_mean_tail","description":"Union over deterministic prefix sizes. The adaptive count is never asserted IID.","url":"../modules/banditrlproof-heavytailscheduledconfidence/index.html#decl-27a1bec7a873","parent":"module:BanditRLProof.HeavyTailScheduledConfidence","order":5050,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailScheduledConfidence"],["Source","BanditRLProof/HeavyTailScheduledConfidence.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem scheduled_adaptive_mean_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (count : Ω → ℕ) (ε u mean : ℝ) (t : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu0 : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω| ^ (1 + ε)) μ) (hu : ∀ i, (∫ ω, |X i ω| ^ (1 + ε) ∂μ) ≤ u) : μ.real {ω | 0 < count ω ∧ count ω ≤ t ∧ confidenceRadius ε u t (count ω) ≤ |(∑ s ∈ Finset.range (count ω), truncate (sampleThreshold ε u t s) (X s ω)) / count ω - mean|} ≤ t * (2 * Real.exp (-confidenceLog t))","missing":[],"search":"scheduled_adaptive_mean_tail banditrlproof.heavytail.scheduled_adaptive_mean_tail union over deterministic prefix sizes. the adaptive count is never asserted iid. theorem compiled","shard":"modules/7e10c23dd66ba6aa.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceTruncationThreshold","label":"sourceTruncationThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceTruncationThreshold","description":"noncomputable def sourceTruncationThreshold (ε u L : ℝ) (s : ℕ) : ℝ","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-49a3fa5f1a22","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5051,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceTruncationThreshold (ε u L : ℝ) (s : ℕ) : ℝ","missing":[],"search":"sourcetruncationthreshold banditrlproof.heavytail.sourcetruncationthreshold noncomputable def sourcetruncationthreshold (ε u l : ℝ) (s : ℕ) : ℝ definition compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceThreshold_pos","label":"sourceThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceThreshold_pos","description":"theorem sourceThreshold_pos (ε u L : ℝ) (hu : 0 < u) (hL : 0 < L) (s : ℕ) : 0 < sourceTruncationThreshold ε u L s","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-193c689be21f","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5052,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_pos (ε u L : ℝ) (hu : 0 < u) (hL : 0 < L) (s : ℕ) : 0 < sourceTruncationThreshold ε u L s","missing":[],"search":"sourcethreshold_pos banditrlproof.heavytail.sourcethreshold_pos theorem sourcethreshold_pos (ε u l : ℝ) (hu : 0 < u) (hl : 0 < l) (s : ℕ) : 0 < sourcetruncationthreshold ε u l s theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceThreshold_le_terminal","label":"sourceThreshold_le_terminal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceThreshold_le_terminal","description":"theorem sourceThreshold_le_terminal (ε u L : ℝ) (hε : 0 ≤ ε) (hu : 0 ≤ u) (hL : 0 < L) (n s : ℕ) (hs : s < n) : sourceTruncationThreshold ε u L s ≤ (u*n/L)^(1/(1+ε))","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-db867b197f06","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5053,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_le_terminal (ε u L : ℝ) (hε : 0 ≤ ε) (hu : 0 ≤ u) (hL : 0 < L) (n s : ℕ) (hs : s < n) : sourceTruncationThreshold ε u L s ≤ (u*n/L)^(1/(1+ε))","missing":[],"search":"sourcethreshold_le_terminal banditrlproof.heavytail.sourcethreshold_le_terminal theorem sourcethreshold_le_terminal (ε u l : ℝ) (hε : 0 ≤ ε) (hu : 0 ≤ u) (hl : 0 < l) (n s : ℕ) (hs : s < n) : sourcetruncationthreshold ε u l s ≤ (u*n/l)^(1/(1+ε)) theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceThreshold_bias_average","label":"sourceThreshold_bias_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceThreshold_bias_average","description":"theorem sourceThreshold_bias_average (ε u L : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hL : 0 < L) (n : ℕ) (hn : 0 < n) : (∑ s ∈ Finset.range n, u / (sourceTruncationThreshold ε u L s)^ε) / n ≤ (1+ε)*u^(1/(1+ε))*(L/n)^(ε/(1+ε))","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-c4208c538ad5","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5054,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_bias_average (ε u L : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hL : 0 < L) (n : ℕ) (hn : 0 < n) : (∑ s ∈ Finset.range n, u / (sourceTruncationThreshold ε u L s)^ε) / n ≤ (1+ε)*u^(1/(1+ε))*(L/n)^(ε/(1+ε))","missing":[],"search":"sourcethreshold_bias_average banditrlproof.heavytail.sourcethreshold_bias_average theorem sourcethreshold_bias_average (ε u l : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (hl : 0 < l) (n : ℕ) (hn : 0 < n) : (∑ s ∈ finset.range n, u / (sourcetruncationthreshold ε u l s)^ε) / n ≤ (1+ε)*u^(1/(1+ε))*(l/n)^(ε/(1+ε)) theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceThreshold_variance_sum","label":"sourceThreshold_variance_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceThreshold_variance_sum","description":"theorem sourceThreshold_variance_sum (ε u L : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (n : ℕ) : (∑ s ∈ Finset.range n, u*(sourceTruncationThreshold ε u L s)^(1-ε)) ≤ n*u*((u*n/L)^(1/(1+ε)))^(1-ε)","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-427d127bc8f5","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5055,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceThreshold_variance_sum (ε u L : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (n : ℕ) : (∑ s ∈ Finset.range n, u*(sourceTruncationThreshold ε u L s)^(1-ε)) ≤ n*u*((u*n/L)^(1/(1+ε)))^(1-ε)","missing":[],"search":"sourcethreshold_variance_sum banditrlproof.heavytail.sourcethreshold_variance_sum theorem sourcethreshold_variance_sum (ε u l : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hl : 0 < l) (n : ℕ) : (∑ s ∈ finset.range n, u*(sourcetruncationthreshold ε u l s)^(1-ε)) ≤ n*u*((u*n/l)^(1/(1+ε)))^(1-ε) theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_centered_sum_upper_tail_sharp","label":"source_centered_sum_upper_tail_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_centered_sum_upper_tail_sharp","description":"One-sided centered-sum bound at the full raw-variable tilt 1/B.","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-f9f7e98194d9","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5056,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_centered_sum_upper_tail_sharp {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 2*(u*n/L)^(1/(1+ε))*L ≤ ∑ i ∈ Finset.range n, (truncate (sourceTruncationThreshold ε u L i) (X i ω) - ∫ ω, truncate (sourceTruncationThreshold ε u L i) (X i ω) ∂μ)} ≤ Real.exp (-(5/4 : ℝ)*L)","missing":[],"search":"source_centered_sum_upper_tail_sharp banditrlproof.heavytail.source_centered_sum_upper_tail_sharp one-sided centered-sum bound at the full raw-variable tilt 1/b. theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_centered_sum_upper_tail","label":"source_centered_sum_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_centered_sum_upper_tail","description":"theorem source_centered_sum_upper_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 2*(u*n/L)^(1/(1+ε))*L ≤ ∑ i ∈ Finset.range n, (t…","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-4afbe6d5588d","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5057,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_centered_sum_upper_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 2*(u*n/L)^(1/(1+ε))*L ≤ ∑ i ∈ Finset.range n, (truncate (sourceTruncationThreshold ε u L i) (X i ω) - ∫ ω, truncate (sourceTruncationThreshold ε u L i) (X i ω) ∂μ)} ≤ Real.exp (-L)","missing":[],"search":"source_centered_sum_upper_tail banditrlproof.heavytail.source_centered_sum_upper_tail theorem source_centered_sum_upper_tail {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (ε u l : ℝ) (n : ℕ) (hn : 0 < n) (hxm : ∀ i, measurable (x i)) (hi : iindepfun x μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hl : 0 < l) (hm : ∀ i, integrable (fun ω => |x i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |x i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 2*(u*n/l)^(1/(1+ε))*l ≤ ∑ i ∈ finset.range n, (truncate (sourcetruncationthreshold ε u l i) (x i ω) - ∫ ω, truncate (sourcetruncationthreshold ε u l i) (x i ω) ∂μ)} ≤ real.exp (-l) theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log_sharp","label":"source_truncated_mean_upper_tail_log_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log_sharp","description":"Constant-four upper deviation for arbitrary positive log confidence.","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-08cfd66f7384","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5058,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:137"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_truncated_mean_upper_tail_log_sharp {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 4*u^(1/(1+ε))*(L/n)^(ε/(1+ε)) ≤ (∑ i ∈ Finset.range n, truncate (sourceTruncationThreshold ε u L i) (X i ω))/n - mean} ≤ Real.exp (-(5/4 : ℝ)*L)","missing":[],"search":"source_truncated_mean_upper_tail_log_sharp banditrlproof.heavytail.source_truncated_mean_upper_tail_log_sharp constant-four upper deviation for arbitrary positive log confidence. theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log","label":"source_truncated_mean_upper_tail_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log","description":"theorem source_truncated_mean_upper_tail_log {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i…","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-e20652759c38","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5059,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:193"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_truncated_mean_upper_tail_log {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 4*u^(1/(1+ε))*(L/n)^(ε/(1+ε)) ≤ (∑ i ∈ Finset.range n, truncate (sourceTruncationThreshold ε u L i) (X i ω))/n - mean} ≤ Real.exp (-L)","missing":[],"search":"source_truncated_mean_upper_tail_log banditrlproof.heavytail.source_truncated_mean_upper_tail_log theorem source_truncated_mean_upper_tail_log {ω : type*} [measurablespace ω] (μ : measure ω) [isprobabilitymeasure μ] (x : ℕ → ω → ℝ) (ε u l mean : ℝ) (n : ℕ) (hn : 0 < n) (hxm : ∀ i, measurable (x i)) (hi : iindepfun x μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hl : 0 < l) (hx : ∀ i, integrable (x i) μ) (hmean : ∀ i, (∫ ω, x i ω ∂μ) = mean) (hm : ∀ i, integrable (fun ω => |x i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |x i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 4*u^(1/(1+ε))*(l/n)^(ε/(1+ε)) ≤ (∑ i ∈ finset.range n, truncate (sourcetruncationthreshold ε u l i) (x i ω))/n - mean} ≤ real.exp (-l) theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail","label":"source_truncated_mean_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_truncated_mean_upper_tail","description":"BCL 2013 Lemma 1 upper deviation, retaining its radius constant four. The non-strict bad event proved here is stronger than a strict upper-tail event.","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-6d11ee534ce9","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5060,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:210"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem source_truncated_mean_upper_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u δ mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu : 0 < u) (hδ : 0 < δ) (hδ1 : δ < 1) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 4*u^(1/(1+ε))*(Real.log (1/δ)/n)^(ε/(1+ε)) ≤ (∑ i ∈ Finset.range n, truncate (sourceTruncationThreshold ε u (Real.log (1/δ)) i) (X i ω))/n - mean} ≤ δ","missing":[],"search":"source_truncated_mean_upper_tail banditrlproof.heavytail.source_truncated_mean_upper_tail bcl 2013 lemma 1 upper deviation, retaining its radius constant four. the non-strict bad event proved here is stronger than a strict upper-tail event. theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.truncate_neg","label":"truncate_neg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncate_neg","description":"theorem truncate_neg (B x : ℝ) : truncate B (-x) = -truncate B x","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-2a21d58b6e01","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5061,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:232"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem truncate_neg (B x : ℝ) : truncate B (-x) = -truncate B x","missing":[],"search":"truncate_neg banditrlproof.heavytail.truncate_neg theorem truncate_neg (b x : ℝ) : truncate b (-x) = -truncate b x theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_lower_tail","label":"source_truncated_mean_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_truncated_mean_lower_tail","description":"Reflection supplies the other one-sided source confidence statement.","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-8df3b9ea69f7","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5062,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:237"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","heavy-tailed"]],"statement":"theorem source_truncated_mean_lower_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u δ mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 < ε) (hε : ε ≤ 1) (hu : 0 < u) (hδ : 0 < δ) (hδ1 : δ < 1) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 4*u^(1/(1+ε))*(Real.log (1/δ)/n)^(ε/(1+ε)) ≤ mean - (∑ i ∈ Finset.range n, truncate (sourceTruncationThreshold ε u (Real.log (1/δ)) i) (X i ω))/n} ≤ δ","missing":[],"search":"source_truncated_mean_lower_tail banditrlproof.heavytail.source_truncated_mean_lower_tail reflection supplies the other one-sided source confidence statement. theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":["heavy-tailed"]},{"id":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_lower_tail_log_sharp","label":"source_truncated_mean_lower_tail_log_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_truncated_mean_lower_tail_log_sharp","description":"Sharper lower log-confidence tail, obtained by reflection.","url":"../modules/banditrlproof-heavytailsourceconfidence/index.html#decl-ca23750bd3cb","parent":"module:BanditRLProof.HeavyTailSourceConfidence","order":5063,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceConfidence"],["Source","BanditRLProof/HeavyTailSourceConfidence.lean:260"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_truncated_mean_lower_tail_log_sharp {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (ε u L mean : ℝ) (n : ℕ) (hn : 0 < n) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hL : 0 < L) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 4*u^(1/(1+ε))*(L/n)^(ε/(1+ε)) ≤ mean - (∑ i ∈ Finset.range n, truncate (sourceTruncationThreshold ε u (L) i) (X i ω))/n} ≤ Real.exp (-(5/4 : ℝ)*L)","missing":[],"search":"source_truncated_mean_lower_tail_log_sharp banditrlproof.heavytail.source_truncated_mean_lower_tail_log_sharp sharper lower log-confidence tail, obtained by reflection. theorem compiled","shard":"modules/2ec6f853c774d792.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.horizonLog","label":"horizonLog","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.horizonLog","description":"noncomputable def horizonLog (T : ℕ) : ℝ","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-5213c9fbc581","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5064,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def horizonLog (T : ℕ) : ℝ","missing":[],"search":"horizonlog banditrlproof.heavytail.sourcepolicy.horizonlog noncomputable def horizonlog (t : ℕ) : ℝ definition compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapBudget","label":"gapBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.gapBudget","description":"noncomputable def gapBudget (ε u gap : ℝ) (T : ℕ) : ℝ","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-d331cd7be87f","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5065,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapBudget (ε u gap : ℝ) (T : ℕ) : ℝ","missing":[],"search":"gapbudget banditrlproof.heavytail.sourcepolicy.gapbudget noncomputable def gapbudget (ε u gap : ℝ) (t : ℕ) : ℝ definition compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapThreshold","label":"gapThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.gapThreshold","description":"noncomputable def gapThreshold (ε u gap : ℝ) (T : ℕ) : ℕ","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-f34527944243","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5066,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapThreshold (ε u gap : ℝ) (T : ℕ) : ℕ","missing":[],"search":"gapthreshold banditrlproof.heavytail.sourcepolicy.gapthreshold noncomputable def gapthreshold (ε u gap : ℝ) (t : ℕ) : ℕ definition compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.horizonLog_nonneg","label":"horizonLog_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.horizonLog_nonneg","description":"theorem horizonLog_nonneg (T : ℕ) : 0 ≤ horizonLog T","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-5d04af0a4294","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5067,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem horizonLog_nonneg (T : ℕ) : 0 ≤ horizonLog T","missing":[],"search":"horizonlog_nonneg banditrlproof.heavytail.sourcepolicy.horizonlog_nonneg theorem horizonlog_nonneg (t : ℕ) : 0 ≤ horizonlog t theorem compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.horizonLog_pos","label":"horizonLog_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.horizonLog_pos","description":"theorem horizonLog_pos (T : ℕ) (hT : 2 ≤ T) : 0 < horizonLog T","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-554113be3399","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5068,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem horizonLog_pos (T : ℕ) (hT : 2 ≤ T) : 0 < horizonLog T","missing":[],"search":"horizonlog_pos banditrlproof.heavytail.sourcepolicy.horizonlog_pos theorem horizonlog_pos (t : ℕ) (ht : 2 ≤ t) : 0 < horizonlog t theorem compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapBudget_nonneg","label":"gapBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.gapBudget_nonneg","description":"theorem gapBudget_nonneg (ε u gap : ℝ) (T : ℕ) (hu : 0 < u) (hg : 0 < gap) : 0 ≤ gapBudget ε u gap T","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-e4d05a85c644","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5069,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapBudget_nonneg (ε u gap : ℝ) (T : ℕ) (hu : 0 < u) (hg : 0 < gap) : 0 ≤ gapBudget ε u gap T","missing":[],"search":"gapbudget_nonneg banditrlproof.heavytail.sourcepolicy.gapbudget_nonneg theorem gapbudget_nonneg (ε u gap : ℝ) (t : ℕ) (hu : 0 < u) (hg : 0 < gap) : 0 ≤ gapbudget ε u gap t theorem compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapThreshold_pos","label":"gapThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.gapThreshold_pos","description":"theorem gapThreshold_pos (ε u gap : ℝ) (T : ℕ) (hT : 2 ≤ T) (hu : 0 < u) (hg : 0 < gap) : 0 < gapThreshold ε u gap T","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-4072ebebb94e","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5070,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapThreshold_pos (ε u gap : ℝ) (T : ℕ) (hT : 2 ≤ T) (hu : 0 < u) (hg : 0 < gap) : 0 < gapThreshold ε u gap T","missing":[],"search":"gapthreshold_pos banditrlproof.heavytail.sourcepolicy.gapthreshold_pos theorem gapthreshold_pos (ε u gap : ℝ) (t : ℕ) (ht : 2 ≤ t) (hu : 0 < u) (hg : 0 < gap) : 0 < gapthreshold ε u gap t theorem compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.sourceLog_le_horizon","label":"sourceLog_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.sourceLog_le_horizon","description":"theorem sourceLog_le_horizon {t T : ℕ} (ht : t < T) : sourceConfidenceLog t ≤ horizonLog T","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-2acb91174e47","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5071,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceLog_le_horizon {t T : ℕ} (ht : t < T) : sourceConfidenceLog t ≤ horizonLog T","missing":[],"search":"sourcelog_le_horizon banditrlproof.heavytail.sourcepolicy.sourcelog_le_horizon theorem sourcelog_le_horizon {t t : ℕ} (ht : t < t) : sourceconfidencelog t ≤ horizonlog t theorem compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.SourcePolicy.twice_radius_le_gap","label":"twice_radius_le_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.SourcePolicy.twice_radius_le_gap","description":"theorem twice_radius_le_gap (ε u gap : ℝ) (hε : 0 < ε) (hu : 0 < u) (hg : 0 < gap) (t T n : ℕ) (hT : 2 ≤ T) (ht : t < T) (hn : gapThreshold ε u gap T ≤ n) : 2*sourceConfidenceRadius ε u t n ≤ gap","url":"../modules/banditrlproof-heavytailsourcegap/index.html#decl-d3fdc0e6ed4e","parent":"module:BanditRLProof.HeavyTailSourceGap","order":5072,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceGap"],["Source","BanditRLProof/HeavyTailSourceGap.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem twice_radius_le_gap (ε u gap : ℝ) (hε : 0 < ε) (hu : 0 < u) (hg : 0 < gap) (t T n : ℕ) (hT : 2 ≤ T) (ht : t < T) (hn : gapThreshold ε u gap T ≤ n) : 2*sourceConfidenceRadius ε u t n ≤ gap","missing":[],"search":"twice_radius_le_gap banditrlproof.heavytail.sourcepolicy.twice_radius_le_gap theorem twice_radius_le_gap (ε u gap : ℝ) (hε : 0 < ε) (hu : 0 < u) (hg : 0 < gap) (t t n : ℕ) (ht : 2 ≤ t) (ht : t < t) (hn : gapthreshold ε u gap t ≤ n) : 2*sourceconfidenceradius ε u t n ≤ gap theorem compiled","shard":"modules/0b05d49bb6e92bb8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceConfidenceLog","label":"sourceConfidenceLog","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceConfidenceLog","description":"noncomputable def sourceConfidenceLog (t : ℕ) : ℝ","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-5d783599dfd0","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5073,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceConfidenceLog (t : ℕ) : ℝ","missing":[],"search":"sourceconfidencelog banditrlproof.heavytail.sourceconfidencelog noncomputable def sourceconfidencelog (t : ℕ) : ℝ definition compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceConfidenceRadius","label":"sourceConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceConfidenceRadius","description":"noncomputable def sourceConfidenceRadius (ε u : ℝ) (t n : ℕ) : ℝ","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-2764283b5291","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5074,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def sourceConfidenceRadius (ε u : ℝ) (t n : ℕ) : ℝ","missing":[],"search":"sourceconfidenceradius banditrlproof.heavytail.sourceconfidenceradius noncomputable def sourceconfidenceradius (ε u : ℝ) (t n : ℕ) : ℝ definition compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sourceConfidenceLog_pos","label":"sourceConfidenceLog_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sourceConfidenceLog_pos","description":"theorem sourceConfidenceLog_pos (t : ℕ) (ht : 0 < t) : 0 < sourceConfidenceLog t","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-92cbaae1671a","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5075,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceConfidenceLog_pos (t : ℕ) (ht : 0 < t) : 0 < sourceConfidenceLog t","missing":[],"search":"sourceconfidencelog_pos banditrlproof.heavytail.sourceconfidencelog_pos theorem sourceconfidencelog_pos (t : ℕ) (ht : 0 < t) : 0 < sourceconfidencelog t theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_adaptive_mean_upper_tail","label":"source_adaptive_mean_upper_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_adaptive_mean_upper_tail","description":"The count is arbitrary; the set explicitly restricts it to the available prefixes.","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-c69705018aa8","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5076,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_adaptive_mean_upper_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (count : Ω → ℕ) (ε u mean : ℝ) (t : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 0 < count ω ∧ count ω ≤ t ∧ sourceConfidenceRadius ε u t (count ω) ≤ (∑ s ∈ Finset.range (count ω), truncate (sourceTruncationThreshold ε u (sourceConfidenceLog t) s) (X s ω)) / count ω - mean} ≤ t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"source_adaptive_mean_upper_tail banditrlproof.heavytail.source_adaptive_mean_upper_tail the count is arbitrary; the set explicitly restricts it to the available prefixes. theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_adaptive_mean_lower_tail","label":"source_adaptive_mean_lower_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_adaptive_mean_lower_tail","description":"Reflection retains the identical count, schedule and raw moment hypotheses.","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-7af0a4ecf74a","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5077,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_adaptive_mean_lower_tail {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (X : ℕ → Ω → ℝ) (count : Ω → ℕ) (ε u mean : ℝ) (t : ℕ) (hXm : ∀ i, Measurable (X i)) (hi : iIndepFun X μ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (hX : ∀ i, Integrable (X i) μ) (hmean : ∀ i, (∫ ω, X i ω ∂μ) = mean) (hm : ∀ i, Integrable (fun ω => |X i ω|^(1+ε)) μ) (hraw : ∀ i, (∫ ω, |X i ω|^(1+ε) ∂μ) ≤ u) : μ.real {ω | 0 < count ω ∧ count ω ≤ t ∧ sourceConfidenceRadius ε u t (count ω) ≤ mean - (∑ s ∈ Finset.range (count ω), truncate (sourceTruncationThreshold ε u (sourceConfidenceLog t) s) (X s ω)) / count ω} ≤ t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)","missing":[],"search":"source_adaptive_mean_lower_tail banditrlproof.heavytail.source_adaptive_mean_lower_tail reflection retains the identical count, schedule and raw moment hypotheses. theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.inverse_sqrt_step","label":"inverse_sqrt_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.inverse_sqrt_step","description":"theorem inverse_sqrt_step (x : ℝ) (hx : 0 < x) : 1 / (Real.sqrt (x+1))^3 ≤ 2*(1/Real.sqrt x - 1/Real.sqrt (x+1))","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-387827a79a65","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5078,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inverse_sqrt_step (x : ℝ) (hx : 0 < x) : 1 / (Real.sqrt (x+1))^3 ≤ 2*(1/Real.sqrt x - 1/Real.sqrt (x+1))","missing":[],"search":"inverse_sqrt_step banditrlproof.heavytail.inverse_sqrt_step theorem inverse_sqrt_step (x : ℝ) (hx : 0 < x) : 1 / (real.sqrt (x+1))^3 ≤ 2*(1/real.sqrt x - 1/real.sqrt (x+1)) theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_schedule_exp_eq","label":"source_schedule_exp_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_schedule_exp_eq","description":"theorem source_schedule_exp_eq (t : ℕ) : Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t) = 1/(Real.sqrt ((t : ℝ)+1))^5","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-7a6f548b76fd","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5079,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_schedule_exp_eq (t : ℕ) : Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t) = 1/(Real.sqrt ((t : ℝ)+1))^5","missing":[],"search":"source_schedule_exp_eq banditrlproof.heavytail.source_schedule_exp_eq theorem source_schedule_exp_eq (t : ℕ) : real.exp (-(5/4 : ℝ)*sourceconfidencelog t) = 1/(real.sqrt ((t : ℝ)+1))^5 theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_schedule_tail_le_telescope","label":"source_schedule_tail_le_telescope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_schedule_tail_le_telescope","description":"theorem source_schedule_tail_le_telescope (t : ℕ) (ht : 0 < t) : t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t) ≤ 2*(1/Real.sqrt (t : ℝ)-1/Real.sqrt ((t : ℝ)+1))","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-57d8f19246e8","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5080,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_schedule_tail_le_telescope (t : ℕ) (ht : 0 < t) : t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t) ≤ 2*(1/Real.sqrt (t : ℝ)-1/Real.sqrt ((t : ℝ)+1))","missing":[],"search":"source_schedule_tail_le_telescope banditrlproof.heavytail.source_schedule_tail_le_telescope theorem source_schedule_tail_le_telescope (t : ℕ) (ht : 0 < t) : t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t) ≤ 2*(1/real.sqrt (t : ℝ)-1/real.sqrt ((t : ℝ)+1)) theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.source_schedule_tail_sum_le_two","label":"source_schedule_tail_sum_le_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.source_schedule_tail_sum_le_two","description":"theorem source_schedule_tail_sum_le_two (T : ℕ) : (∑ t ∈ Finset.range T, t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)) ≤ 2","url":"../modules/banditrlproof-heavytailsourceschedule/index.html#decl-15afb8b4514c","parent":"module:BanditRLProof.HeavyTailSourceSchedule","order":5081,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailSourceSchedule"],["Source","BanditRLProof/HeavyTailSourceSchedule.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_schedule_tail_sum_le_two (T : ℕ) : (∑ t ∈ Finset.range T, t*Real.exp (-(5/4 : ℝ)*sourceConfidenceLog t)) ≤ 2","missing":[],"search":"source_schedule_tail_sum_le_two banditrlproof.heavytail.source_schedule_tail_sum_le_two theorem source_schedule_tail_sum_le_two (t : ℕ) : (∑ t ∈ finset.range t, t*real.exp (-(5/4 : ℝ)*sourceconfidencelog t)) ≤ 2 theorem compiled","shard":"modules/548e0673851b3aee.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.scheduled_exp_eq","label":"scheduled_exp_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.scheduled_exp_eq","description":"theorem scheduled_exp_eq (t : ℕ) : Real.exp (-confidenceLog t) = 1 / (max (t : ℝ) 2)^4","url":"../modules/banditrlproof-heavytailtailsum/index.html#decl-0bddb9899f9a","parent":"module:BanditRLProof.HeavyTailTailSum","order":5082,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTailSum"],["Source","BanditRLProof/HeavyTailTailSum.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem scheduled_exp_eq (t : ℕ) : Real.exp (-confidenceLog t) = 1 / (max (t : ℝ) 2)^4","missing":[],"search":"scheduled_exp_eq banditrlproof.heavytail.scheduled_exp_eq theorem scheduled_exp_eq (t : ℕ) : real.exp (-confidencelog t) = 1 / (max (t : ℝ) 2)^4 theorem compiled","shard":"modules/3f46588cba47b9f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.cubic_tail_le_telescope","label":"cubic_tail_le_telescope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.cubic_tail_le_telescope","description":"theorem cubic_tail_le_telescope (t : ℕ) (ht : 2 ≤ t) : 4*t*Real.exp (-confidenceLog t) ≤ 1 / ((t : ℝ)-1) - 1/t","url":"../modules/banditrlproof-heavytailtailsum/index.html#decl-bfecd6d30ca2","parent":"module:BanditRLProof.HeavyTailTailSum","order":5083,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTailSum"],["Source","BanditRLProof/HeavyTailTailSum.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cubic_tail_le_telescope (t : ℕ) (ht : 2 ≤ t) : 4*t*Real.exp (-confidenceLog t) ≤ 1 / ((t : ℝ)-1) - 1/t","missing":[],"search":"cubic_tail_le_telescope banditrlproof.heavytail.cubic_tail_le_telescope theorem cubic_tail_le_telescope (t : ℕ) (ht : 2 ≤ t) : 4*t*real.exp (-confidencelog t) ≤ 1 / ((t : ℝ)-1) - 1/t theorem compiled","shard":"modules/3f46588cba47b9f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.reciprocal_telescope","label":"reciprocal_telescope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.reciprocal_telescope","description":"theorem reciprocal_telescope (n : ℕ) : (∑ s ∈ Finset.range n, (1 / ((s : ℝ)+1) - 1/((s : ℝ)+2))) = 1 - 1/((n : ℝ)+1)","url":"../modules/banditrlproof-heavytailtailsum/index.html#decl-ae8f60b4e4e9","parent":"module:BanditRLProof.HeavyTailTailSum","order":5084,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTailSum"],["Source","BanditRLProof/HeavyTailTailSum.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem reciprocal_telescope (n : ℕ) : (∑ s ∈ Finset.range n, (1 / ((s : ℝ)+1) - 1/((s : ℝ)+2))) = 1 - 1/((n : ℝ)+1)","missing":[],"search":"reciprocal_telescope banditrlproof.heavytail.reciprocal_telescope theorem reciprocal_telescope (n : ℕ) : (∑ s ∈ finset.range n, (1 / ((s : ℝ)+1) - 1/((s : ℝ)+2))) = 1 - 1/((n : ℝ)+1) theorem compiled","shard":"modules/3f46588cba47b9f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.scheduled_tail_sum_le_two","label":"scheduled_tail_sum_le_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.scheduled_tail_sum_le_two","description":"theorem scheduled_tail_sum_le_two (T : ℕ) : (∑ t ∈ Finset.range T, 4*t*Real.exp (-confidenceLog t)) ≤ 2","url":"../modules/banditrlproof-heavytailtailsum/index.html#decl-47cf9e2c5ee6","parent":"module:BanditRLProof.HeavyTailTailSum","order":5085,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTailSum"],["Source","BanditRLProof/HeavyTailTailSum.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem scheduled_tail_sum_le_two (T : ℕ) : (∑ t ∈ Finset.range T, 4*t*Real.exp (-confidenceLog t)) ≤ 2","missing":[],"search":"scheduled_tail_sum_le_two banditrlproof.heavytail.scheduled_tail_sum_le_two theorem scheduled_tail_sum_le_two (t : ℕ) : (∑ t ∈ finset.range t, 4*t*real.exp (-confidencelog t)) ≤ 2 theorem compiled","shard":"modules/3f46588cba47b9f6.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.truncate","label":"truncate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.HeavyTail.truncate","description":"noncomputable def truncate (B x : ℝ) : ℝ","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-cfd9593ce788","parent":"module:BanditRLProof.HeavyTailTruncation","order":5086,"meta":[["Kind","definition"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def truncate (B x : ℝ) : ℝ","missing":[],"search":"truncate banditrlproof.heavytail.truncate noncomputable def truncate (b x : ℝ) : ℝ definition compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_truncate_le","label":"abs_truncate_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_truncate_le","description":"theorem abs_truncate_le (B x : ℝ) (hB : 0 ≤ B) : |truncate B x| ≤ B","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-223430beba4f","parent":"module:BanditRLProof.HeavyTailTruncation","order":5087,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_truncate_le (B x : ℝ) (hB : 0 ≤ B) : |truncate B x| ≤ B","missing":[],"search":"abs_truncate_le banditrlproof.heavytail.abs_truncate_le theorem abs_truncate_le (b x : ℝ) (hb : 0 ≤ b) : |truncate b x| ≤ b theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.measurable_truncate","label":"measurable_truncate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.measurable_truncate","description":"theorem measurable_truncate (B : ℝ) : Measurable (truncate B)","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-d9e73e81ea68","parent":"module:BanditRLProof.HeavyTailTruncation","order":5088,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_truncate (B : ℝ) : Measurable (truncate B)","missing":[],"search":"measurable_truncate banditrlproof.heavytail.measurable_truncate theorem measurable_truncate (b : ℝ) : measurable (truncate b) theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.abs_sub_truncate_le","label":"abs_sub_truncate_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.abs_sub_truncate_le","description":"The discarded tail is controlled by a raw (1+epsilon)-moment.","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-1c6f3f3cfe7b","parent":"module:BanditRLProof.HeavyTailTruncation","order":5089,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem abs_sub_truncate_le (B x ε : ℝ) (hB : 0 < B) (hε : 0 ≤ ε) : |x - truncate B x| ≤ |x| ^ (1 + ε) / B ^ ε","missing":[],"search":"abs_sub_truncate_le banditrlproof.heavytail.abs_sub_truncate_le the discarded tail is controlled by a raw (1+epsilon)-moment. theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sq_truncate_le","label":"sq_truncate_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sq_truncate_le","description":"Truncation supplies the variance-scale envelope needed by Bernstein.","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-3c541d0a6630","parent":"module:BanditRLProof.HeavyTailTruncation","order":5090,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sq_truncate_le (B x ε : ℝ) (hB : 0 < B) (hε : ε ≤ 1) : (truncate B x) ^ 2 ≤ |x| ^ (1 + ε) * B ^ (1 - ε)","missing":[],"search":"sq_truncate_le banditrlproof.heavytail.sq_truncate_le truncation supplies the variance-scale envelope needed by bernstein. theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.integral_truncate_bias_le","label":"integral_truncate_bias_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.integral_truncate_bias_le","description":"Integrating the pointwise tail inequality produces an actual bias bound.","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-679592c17f4e","parent":"module:BanditRLProof.HeavyTailTruncation","order":5091,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_truncate_bias_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) (X : Ω → ℝ) (B ε u : ℝ) (hB : 0 < B) (hε : 0 ≤ ε) (hXm : Measurable X) (hX : Integrable X μ) (hm : Integrable (fun ω => |X ω| ^ (1 + ε)) μ) (hu : (∫ ω, |X ω| ^ (1 + ε) ∂μ) ≤ u) : |(∫ ω, X ω ∂μ) - ∫ ω, truncate B (X ω) ∂μ| ≤ u / B ^ ε","missing":[],"search":"integral_truncate_bias_le banditrlproof.heavytail.integral_truncate_bias_le integrating the pointwise tail inequality produces an actual bias bound. theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.integral_sq_truncate_le","label":"integral_sq_truncate_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.integral_sq_truncate_le","description":"The second moment envelope is integrable and bounded by the raw moment.","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-0b9a4be3acfe","parent":"module:BanditRLProof.HeavyTailTruncation","order":5092,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:87"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_truncate_le {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) (X : Ω → ℝ) (B ε u : ℝ) (hB : 0 < B) (hε : ε ≤ 1) (hXm : Measurable X) (hm : Integrable (fun ω => |X ω| ^ (1 + ε)) μ) (hu : (∫ ω, |X ω| ^ (1 + ε) ∂μ) ≤ u) : (∫ ω, (truncate B (X ω)) ^ 2 ∂μ) ≤ u * B ^ (1 - ε)","missing":[],"search":"integral_sq_truncate_le banditrlproof.heavytail.integral_sq_truncate_le the second moment envelope is integrable and bounded by the raw moment. theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.transformed_observed_prefix","label":"transformed_observed_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.transformed_observed_prefix","description":"A sample-index transform of actual observations is the same transform of the consumed latent prefix. No IID claim is made about adaptively selected data.","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-308c0aeeb61d","parent":"module:BanditRLProof.HeavyTailTruncation","order":5093,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem transformed_observed_prefix {Ω : Type*} {K : ℕ} (action : Ω → ActionTrace (Fin K)) (stream : Ω → UCB.ArmRewardStream K) (transform : ℕ → Fin K → ℝ → ℝ) (ω : Ω) (arm : Fin K) (n : ℕ) : sumRewards (action ω) (fun t => transform (pullCount (action ω) (action ω t) t) (action ω t) (UCB.rewardFromArmStream action stream ω t)) arm n = (Finset.range (pullCount (action ω) arm n)).sum (fun s => transform s arm (stream ω s arm))","missing":[],"search":"transformed_observed_prefix banditrlproof.heavytail.transformed_observed_prefix a sample-index transform of actual observations is the same transform of the consumed latent prefix. no iid claim is made about adaptively selected data. theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.estimator_error_le","label":"estimator_error_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.estimator_error_le","description":"Common estimator assembly: deterministic bias and stochastic fluctuation remain separate obligations, rather than assuming the desired confidence event.","url":"../modules/banditrlproof-heavytailtruncation/index.html#decl-82cc04e6d20d","parent":"module:BanditRLProof.HeavyTailTruncation","order":5094,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTruncation"],["Source","BanditRLProof/HeavyTailTruncation.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem estimator_error_le (estimate center mean bias fluctuation : ℝ) (hb : |center - mean| ≤ bias) (hf : |estimate - center| ≤ fluctuation) : |estimate - mean| ≤ bias + fluctuation","missing":[],"search":"estimator_error_le banditrlproof.heavytail.estimator_error_le common estimator assembly: deterministic bias and stochastic fluctuation remain separate obligations, rather than assuming the desired confidence event. theorem compiled","shard":"modules/90be129d6eb38816.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.power_threshold_bias_term","label":"power_threshold_bias_term","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.power_threshold_bias_term","description":"theorem power_threshold_bias_term (u c a ε x : ℝ) (hc : 0 < c) (hx : 0 < x) (haε : a * ε = 1-a) : u / (c * x^a)^ε = (u / c^ε) * x^(a-1)","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-cb1b9c0fea1f","parent":"module:BanditRLProof.HeavyTailTuning","order":5095,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem power_threshold_bias_term (u c a ε x : ℝ) (hc : 0 < c) (hx : 0 < x) (haε : a * ε = 1-a) : u / (c * x^a)^ε = (u / c^ε) * x^(a-1)","missing":[],"search":"power_threshold_bias_term banditrlproof.heavytail.power_threshold_bias_term theorem power_threshold_bias_term (u c a ε x : ℝ) (hc : 0 < c) (hx : 0 < x) (haε : a * ε = 1-a) : u / (c * x^a)^ε = (u / c^ε) * x^(a-1) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.power_threshold_bias_sum","label":"power_threshold_bias_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.power_threshold_bias_sum","description":"theorem power_threshold_bias_sum (u c a ε : ℝ) (hu : 0 ≤ u) (hc : 0 < c) (ha : 0 < a) (ha1 : a ≤ 1) (haε : a * ε = 1-a) (n : ℕ) : (∑ s ∈ Finset.range n, u / (c * ((s : ℝ)+1)^a)^ε) ≤ (u / c^ε) * ((n : ℝ)^a / a)","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-1ada51b0b616","parent":"module:BanditRLProof.HeavyTailTuning","order":5096,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem power_threshold_bias_sum (u c a ε : ℝ) (hu : 0 ≤ u) (hc : 0 < c) (ha : 0 < a) (ha1 : a ≤ 1) (haε : a * ε = 1-a) (n : ℕ) : (∑ s ∈ Finset.range n, u / (c * ((s : ℝ)+1)^a)^ε) ≤ (u / c^ε) * ((n : ℝ)^a / a)","missing":[],"search":"power_threshold_bias_sum banditrlproof.heavytail.power_threshold_bias_sum theorem power_threshold_bias_sum (u c a ε : ℝ) (hu : 0 ≤ u) (hc : 0 < c) (ha : 0 < a) (ha1 : a ≤ 1) (haε : a * ε = 1-a) (n : ℕ) : (∑ s ∈ finset.range n, u / (c * ((s : ℝ)+1)^a)^ε) ≤ (u / c^ε) * ((n : ℝ)^a / a) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold_factor","label":"sampleThreshold_factor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold_factor","description":"theorem sampleThreshold_factor (ε u : ℝ) (hu : 0 ≤ u) (t s : ℕ) : sampleThreshold ε u t s = (u / confidenceLog t)^(1/(1+ε)) * ((s : ℝ)+1)^(1/(1+ε))","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-0b89b3548b25","parent":"module:BanditRLProof.HeavyTailTuning","order":5097,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleThreshold_factor (ε u : ℝ) (hu : 0 ≤ u) (t s : ℕ) : sampleThreshold ε u t s = (u / confidenceLog t)^(1/(1+ε)) * ((s : ℝ)+1)^(1/(1+ε))","missing":[],"search":"samplethreshold_factor banditrlproof.heavytail.samplethreshold_factor theorem samplethreshold_factor (ε u : ℝ) (hu : 0 ≤ u) (t s : ℕ) : samplethreshold ε u t s = (u / confidencelog t)^(1/(1+ε)) * ((s : ℝ)+1)^(1/(1+ε)) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold_bias_sum","label":"sampleThreshold_bias_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold_bias_sum","description":"theorem sampleThreshold_bias_sum (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (t n : ℕ) : (∑ s ∈ Finset.range n, u / (sampleThreshold ε u t s)^ε) ≤ (u / ((u / confidenceLog t)^(1/(1+ε)))^ε) * ((n : ℝ)^(1/(1+ε)) / (1/(1+ε)))","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-53f44ab9d805","parent":"module:BanditRLProof.HeavyTailTuning","order":5098,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleThreshold_bias_sum (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (t n : ℕ) : (∑ s ∈ Finset.range n, u / (sampleThreshold ε u t s)^ε) ≤ (u / ((u / confidenceLog t)^(1/(1+ε)))^ε) * ((n : ℝ)^(1/(1+ε)) / (1/(1+ε)))","missing":[],"search":"samplethreshold_bias_sum banditrlproof.heavytail.samplethreshold_bias_sum theorem samplethreshold_bias_sum (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (t n : ℕ) : (∑ s ∈ finset.range n, u / (samplethreshold ε u t s)^ε) ≤ (u / ((u / confidencelog t)^(1/(1+ε)))^ε) * ((n : ℝ)^(1/(1+ε)) / (1/(1+ε))) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.power_scale_bias","label":"power_scale_bias","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.power_scale_bias","description":"theorem power_scale_bias (u L a ε : ℝ) (hu : 0 < u) (hL : 0 < L) (haε : a * ε = 1-a) : u / ((u/L)^a)^ε = u^a * L^(1-a)","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-79be30d67971","parent":"module:BanditRLProof.HeavyTailTuning","order":5099,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem power_scale_bias (u L a ε : ℝ) (hu : 0 < u) (hL : 0 < L) (haε : a * ε = 1-a) : u / ((u/L)^a)^ε = u^a * L^(1-a)","missing":[],"search":"power_scale_bias banditrlproof.heavytail.power_scale_bias theorem power_scale_bias (u l a ε : ℝ) (hu : 0 < u) (hl : 0 < l) (haε : a * ε = 1-a) : u / ((u/l)^a)^ε = u^a * l^(1-a) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.power_bias_normalization","label":"power_bias_normalization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.power_bias_normalization","description":"theorem power_bias_normalization (u L a N : ℝ) (hL : 0 < L) (hN : 0 < N) : (u^a * L^(1-a)) * (N^a / a) / N = (1/a) * u^a * (L/N)^(1-a)","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-0dcfc4f0cf97","parent":"module:BanditRLProof.HeavyTailTuning","order":5100,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem power_bias_normalization (u L a N : ℝ) (hL : 0 < L) (hN : 0 < N) : (u^a * L^(1-a)) * (N^a / a) / N = (1/a) * u^a * (L/N)^(1-a)","missing":[],"search":"power_bias_normalization banditrlproof.heavytail.power_bias_normalization theorem power_bias_normalization (u l a n : ℝ) (hl : 0 < l) (hn : 0 < n) : (u^a * l^(1-a)) * (n^a / a) / n = (1/a) * u^a * (l/n)^(1-a) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold_bias_average","label":"sampleThreshold_bias_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold_bias_average","description":"theorem sampleThreshold_bias_average (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (t n : ℕ) (hn : 0 < n) : (∑ s ∈ Finset.range n, u / (sampleThreshold ε u t s)^ε) / n ≤ (1+ε) * u^(1/(1+ε)) * (confidenceLog t / n)^(ε/(1+ε))","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-f577e34dc3e0","parent":"module:BanditRLProof.HeavyTailTuning","order":5101,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleThreshold_bias_average (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (t n : ℕ) (hn : 0 < n) : (∑ s ∈ Finset.range n, u / (sampleThreshold ε u t s)^ε) / n ≤ (1+ε) * u^(1/(1+ε)) * (confidenceLog t / n)^(ε/(1+ε))","missing":[],"search":"samplethreshold_bias_average banditrlproof.heavytail.samplethreshold_bias_average theorem samplethreshold_bias_average (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 < u) (t n : ℕ) (hn : 0 < n) : (∑ s ∈ finset.range n, u / (samplethreshold ε u t s)^ε) / n ≤ (1+ε) * u^(1/(1+ε)) * (confidencelog t / n)^(ε/(1+ε)) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.threshold_scale_identity","label":"threshold_scale_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.threshold_scale_identity","description":"theorem threshold_scale_identity (u L N a : ℝ) (hu : 0 ≤ u) (hL : 0 < L) (hN : 0 < N) : (u*N/L)^a * L = N * (u^a * (L/N)^(1-a))","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-fa5b7215a7a5","parent":"module:BanditRLProof.HeavyTailTuning","order":5102,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem threshold_scale_identity (u L N a : ℝ) (hu : 0 ≤ u) (hL : 0 < L) (hN : 0 < N) : (u*N/L)^a * L = N * (u^a * (L/N)^(1-a))","missing":[],"search":"threshold_scale_identity banditrlproof.heavytail.threshold_scale_identity theorem threshold_scale_identity (u l n a : ℝ) (hu : 0 ≤ u) (hl : 0 < l) (hn : 0 < n) : (u*n/l)^a * l = n * (u^a * (l/n)^(1-a)) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.threshold_variance_identity","label":"threshold_variance_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.threshold_variance_identity","description":"theorem threshold_variance_identity (u L N ε : ℝ) (hu : 0 < u) (hL : 0 < L) (hN : 0 < N) (hp : 0 < 1+ε) : N*u*((u*N/L)^(1/(1+ε)))^(1-ε)*L = ((u*N/L)^(1/(1+ε))*L)^2","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-b767a45d5670","parent":"module:BanditRLProof.HeavyTailTuning","order":5103,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem threshold_variance_identity (u L N ε : ℝ) (hu : 0 < u) (hL : 0 < L) (hN : 0 < N) (hp : 0 < 1+ε) : N*u*((u*N/L)^(1/(1+ε)))^(1-ε)*L = ((u*N/L)^(1/(1+ε))*L)^2","missing":[],"search":"threshold_variance_identity banditrlproof.heavytail.threshold_variance_identity theorem threshold_variance_identity (u l n ε : ℝ) (hu : 0 < u) (hl : 0 < l) (hn : 0 < n) (hp : 0 < 1+ε) : n*u*((u*n/l)^(1/(1+ε)))^(1-ε)*l = ((u*n/l)^(1/(1+ε))*l)^2 theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.confidenceLog_pos","label":"confidenceLog_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.confidenceLog_pos","description":"theorem confidenceLog_pos (t : ℕ) : 0 < confidenceLog t","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-4559d7c4e2f8","parent":"module:BanditRLProof.HeavyTailTuning","order":5104,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem confidenceLog_pos (t : ℕ) : 0 < confidenceLog t","missing":[],"search":"confidencelog_pos banditrlproof.heavytail.confidencelog_pos theorem confidencelog_pos (t : ℕ) : 0 < confidencelog t theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold_pos","label":"sampleThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold_pos","description":"theorem sampleThreshold_pos (ε u : ℝ) (hu : 0 < u) (t s : ℕ) : 0 < sampleThreshold ε u t s","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-0d0ff535f3b8","parent":"module:BanditRLProof.HeavyTailTuning","order":5105,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleThreshold_pos (ε u : ℝ) (hu : 0 < u) (t s : ℕ) : 0 < sampleThreshold ε u t s","missing":[],"search":"samplethreshold_pos banditrlproof.heavytail.samplethreshold_pos theorem samplethreshold_pos (ε u : ℝ) (hu : 0 < u) (t s : ℕ) : 0 < samplethreshold ε u t s theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold_le_terminal","label":"sampleThreshold_le_terminal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold_le_terminal","description":"theorem sampleThreshold_le_terminal (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 ≤ u) (t n s : ℕ) (hs : s < n) : sampleThreshold ε u t s ≤ (u*n/confidenceLog t)^(1/(1+ε))","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-3624b0585fe4","parent":"module:BanditRLProof.HeavyTailTuning","order":5106,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:134"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleThreshold_le_terminal (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 ≤ u) (t n s : ℕ) (hs : s < n) : sampleThreshold ε u t s ≤ (u*n/confidenceLog t)^(1/(1+ε))","missing":[],"search":"samplethreshold_le_terminal banditrlproof.heavytail.samplethreshold_le_terminal theorem samplethreshold_le_terminal (ε u : ℝ) (hε : 0 ≤ ε) (hu : 0 ≤ u) (t n s : ℕ) (hs : s < n) : samplethreshold ε u t s ≤ (u*n/confidencelog t)^(1/(1+ε)) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.sampleThreshold_variance_sum","label":"sampleThreshold_variance_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.sampleThreshold_variance_sum","description":"theorem sampleThreshold_variance_sum (ε u : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (t n : ℕ) : (∑ s ∈ Finset.range n, u * (sampleThreshold ε u t s)^(1-ε)) ≤ n*u*((u*n/confidenceLog t)^(1/(1+ε)))^(1-ε)","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-640a760a5613","parent":"module:BanditRLProof.HeavyTailTuning","order":5107,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sampleThreshold_variance_sum (ε u : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (t n : ℕ) : (∑ s ∈ Finset.range n, u * (sampleThreshold ε u t s)^(1-ε)) ≤ n*u*((u*n/confidenceLog t)^(1/(1+ε)))^(1-ε)","missing":[],"search":"samplethreshold_variance_sum banditrlproof.heavytail.samplethreshold_variance_sum theorem samplethreshold_variance_sum (ε u : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (t n : ℕ) : (∑ s ∈ finset.range n, u * (samplethreshold ε u t s)^(1-ε)) ≤ n*u*((u*n/confidencelog t)^(1/(1+ε)))^(1-ε) theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.tuned_radius_le","label":"tuned_radius_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.tuned_radius_le","description":"theorem tuned_radius_le (ε u : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (t n : ℕ) (hn : 0 < n) : ((∑ s ∈ Finset.range n, u / (sampleThreshold ε u t s)^ε) + (2 * Real.sqrt ((∑ s ∈ Finset.range n, u * (sampleThreshold ε u t s)^(1-ε)) * confidenceLog t) + 2 * (u*n/confidenceLog t)^(1/(1+ε)) * confidenceLog t)) / n ≤ confidenceRadius ε u t n","url":"../modules/banditrlproof-heavytailtuning/index.html#decl-0ff917129693","parent":"module:BanditRLProof.HeavyTailTuning","order":5108,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailTuning"],["Source","BanditRLProof/HeavyTailTuning.lean:157"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem tuned_radius_le (ε u : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (t n : ℕ) (hn : 0 < n) : ((∑ s ∈ Finset.range n, u / (sampleThreshold ε u t s)^ε) + (2 * Real.sqrt ((∑ s ∈ Finset.range n, u * (sampleThreshold ε u t s)^(1-ε)) * confidenceLog t) + 2 * (u*n/confidenceLog t)^(1/(1+ε)) * confidenceLog t)) / n ≤ confidenceRadius ε u t n","missing":[],"search":"tuned_radius_le banditrlproof.heavytail.tuned_radius_le theorem tuned_radius_le (ε u : ℝ) (hε0 : 0 ≤ ε) (hε : ε ≤ 1) (hu : 0 < u) (t n : ℕ) (hn : 0 < n) : ((∑ s ∈ finset.range n, u / (samplethreshold ε u t s)^ε) + (2 * real.sqrt ((∑ s ∈ finset.range n, u * (samplethreshold ε u t s)^(1-ε)) * confidencelog t) + 2 * (u*n/confidencelog t)^(1/(1+ε)) * confidencelog t)) / n ≤ confidenceradius ε u t n theorem compiled","shard":"modules/52d27245a640134c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.exp_le_one_add_self_add_three_quarters_sq","label":"exp_le_one_add_self_add_three_quarters_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.exp_le_one_add_self_add_three_quarters_sq","description":"theorem exp_le_one_add_self_add_three_quarters_sq {x : ℝ} (hx : |x| ≤ 1) : Real.exp x ≤ 1 + x + (3/4 : ℝ)*x^2","url":"../modules/banditrlproof-heavytailunshiftedmgf/index.html#decl-1c0a3fe9349c","parent":"module:BanditRLProof.HeavyTailUnshiftedMGF","order":5109,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailUnshiftedMGF"],["Source","BanditRLProof/HeavyTailUnshiftedMGF.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_le_one_add_self_add_three_quarters_sq {x : ℝ} (hx : |x| ≤ 1) : Real.exp x ≤ 1 + x + (3/4 : ℝ)*x^2","missing":[],"search":"exp_le_one_add_self_add_three_quarters_sq banditrlproof.heavytail.exp_le_one_add_self_add_three_quarters_sq theorem exp_le_one_add_self_add_three_quarters_sq {x : ℝ} (hx : |x| ≤ 1) : real.exp x ≤ 1 + x + (3/4 : ℝ)*x^2 theorem compiled","shard":"modules/3d2a5ac4eb840c7a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.bounded_centering_mgf_unshifted_sharp","label":"bounded_centering_mgf_unshifted_sharp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.bounded_centering_mgf_unshifted_sharp","description":"Sharper raw-second-moment MGF at the unchanged full raw-variable tilt.","url":"../modules/banditrlproof-heavytailunshiftedmgf/index.html#decl-b0ad1680f127","parent":"module:BanditRLProof.HeavyTailUnshiftedMGF","order":5110,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailUnshiftedMGF"],["Source","BanditRLProof/HeavyTailUnshiftedMGF.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bounded_centering_mgf_unshifted_sharp {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (Y : Ω → ℝ) (B v tilt : ℝ) (hYm : Measurable Y) (hbound : ∀ ω, |Y ω| ≤ B) (hv : (∫ ω, (Y ω)^2 ∂μ) ≤ v) (hsmall : |tilt| * B ≤ 1) : Concentration.HasMGFUpperBoundAt (fun ω => Y ω - ∫ ω, Y ω ∂μ) tilt ((3/4 : ℝ)*tilt^2 * v) μ","missing":[],"search":"bounded_centering_mgf_unshifted_sharp banditrlproof.heavytail.bounded_centering_mgf_unshifted_sharp sharper raw-second-moment mgf at the unchanged full raw-variable tilt. theorem compiled","shard":"modules/3d2a5ac4eb840c7a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.HeavyTail.bounded_centering_mgf_unshifted","label":"bounded_centering_mgf_unshifted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.HeavyTail.bounded_centering_mgf_unshifted","description":"Compatibility weakening of the sharper raw-second-moment producer.","url":"../modules/banditrlproof-heavytailunshiftedmgf/index.html#decl-c3bf79a8f8a2","parent":"module:BanditRLProof.HeavyTailUnshiftedMGF","order":5111,"meta":[["Kind","theorem"],["Module","BanditRLProof.HeavyTailUnshiftedMGF"],["Source","BanditRLProof/HeavyTailUnshiftedMGF.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bounded_centering_mgf_unshifted {Ω : Type*} [MeasurableSpace Ω] (μ : Measure Ω) [IsProbabilityMeasure μ] (Y : Ω → ℝ) (B v tilt : ℝ) (hYm : Measurable Y) (hbound : ∀ ω, |Y ω| ≤ B) (hv : (∫ ω, (Y ω)^2 ∂μ) ≤ v) (hsmall : |tilt| * B ≤ 1) : Concentration.HasMGFUpperBoundAt (fun ω => Y ω - ∫ ω, Y ω ∂μ) tilt (tilt^2 * v) μ","missing":[],"search":"bounded_centering_mgf_unshifted banditrlproof.heavytail.bounded_centering_mgf_unshifted compatibility weakening of the sharper raw-second-moment producer. theorem compiled","shard":"modules/3d2a5ac4eb840c7a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.History.FiniteActionHistory","label":"FiniteActionHistory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.History.FiniteActionHistory","description":"Finite action histories through index `t`, represented on Mathlib's `Finset.Iic t` finite-prefix index type.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-770020465cc0","parent":"module:BanditRLProof.HistoryFiltration","order":5112,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:22"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"abbrev FiniteActionHistory (Action : Type v) (t : Nat)","missing":[],"search":"finiteactionhistory banditrlproof.history.finiteactionhistory finite action histories through index `t`, represented on mathlib's `finset.iic t` finite-prefix index type. abbreviation compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.FiniteRewardHistory","label":"FiniteRewardHistory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.History.FiniteRewardHistory","description":"Finite reward histories through index `t`, represented on Mathlib's `Finset.Iic t` finite-prefix index type.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-9087f6b857f9","parent":"module:BanditRLProof.HistoryFiltration","order":5113,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:27"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"abbrev FiniteRewardHistory (Reward : Type w) (t : Nat)","missing":[],"search":"finiterewardhistory banditrlproof.history.finiterewardhistory finite reward histories through index `t`, represented on mathlib's `finset.iic t` finite-prefix index type. abbreviation compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.FinitePairHistory","label":"FinitePairHistory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.History.FinitePairHistory","description":"Pair-coordinate finite action/reward histories through index `t`.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-56f314647881","parent":"module:BanditRLProof.HistoryFiltration","order":5114,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:31"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"abbrev FinitePairHistory (Action : Type v) (Reward : Type w) (t : Nat)","missing":[],"search":"finitepairhistory banditrlproof.history.finitepairhistory pair-coordinate finite action/reward histories through index `t`. abbreviation compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.FiniteHistory","label":"FiniteHistory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.History.FiniteHistory","description":"The paired finite action/reward history object at a finite prefix.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-c0905b6162f7","parent":"module:BanditRLProof.HistoryFiltration","order":5115,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:35"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"abbrev FiniteHistory (Action : Type v) (Reward : Type w) (t : Nat)","missing":[],"search":"finitehistory banditrlproof.history.finitehistory the paired finite action/reward history object at a finite prefix. abbreviation compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteActionHistoryOfTrace","label":"finiteActionHistoryOfTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.finiteActionHistoryOfTrace","description":"Restrict an infinite action trace to the finite prefix indexed by `Finset.Iic t`.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-dab5b2794990","parent":"module:BanditRLProof.HistoryFiltration","order":5116,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:40"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def finiteActionHistoryOfTrace {Action : Type v} (action : ActionTrace Action) (t : Nat) : FiniteActionHistory Action t","missing":[],"search":"finiteactionhistoryoftrace banditrlproof.history.finiteactionhistoryoftrace restrict an infinite action trace to the finite prefix indexed by `finset.iic t`. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteRewardHistoryOfTrace","label":"finiteRewardHistoryOfTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.finiteRewardHistoryOfTrace","description":"Restrict an infinite reward trace to the finite prefix indexed by `Finset.Iic t`.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-6c3280deda6a","parent":"module:BanditRLProof.HistoryFiltration","order":5117,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:48"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def finiteRewardHistoryOfTrace {Reward : Type w} (reward : RewardTrace Reward) (t : Nat) : FiniteRewardHistory Reward t","missing":[],"search":"finiterewardhistoryoftrace banditrlproof.history.finiterewardhistoryoftrace restrict an infinite reward trace to the finite prefix indexed by `finset.iic t`. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.completeRewardTrace","label":"completeRewardTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.completeRewardTrace","description":"Complete a finite reward history to an infinite trace, using `default` outside the observed prefix. The completion agrees with `history` at every coordinate through `t`. It is a deterministic history-reconstruction helper; it does not assert that the defaulted future coordinates agree with an ambient reward process.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-b831db96b0ca","parent":"module:BanditRLProof.HistoryFiltration","order":5118,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:62"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def completeRewardTrace {Reward : Type w} (t : Nat) (history : FiniteRewardHistory Reward t) (default : Reward) : RewardTrace Reward","missing":[],"search":"completerewardtrace banditrlproof.history.completerewardtrace complete a finite reward history to an infinite trace, using `default` outside the observed prefix. the completion agrees with `history` at every coordinate through `t`. it is a deterministic history-reconstruction helper; it does not assert that the defaulted future coordinates agree with an ambient reward process. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteHistoryOfTrace","label":"finiteHistoryOfTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.finiteHistoryOfTrace","description":"Restrict infinite action and reward traces to a paired finite history.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-985d22ba91cc","parent":"module:BanditRLProof.HistoryFiltration","order":5119,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:69"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def finiteHistoryOfTrace {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : FiniteHistory Action Reward t","missing":[],"search":"finitehistoryoftrace banditrlproof.history.finitehistoryoftrace restrict infinite action and reward traces to a paired finite history. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finitePairHistoryOfTrace","label":"finitePairHistoryOfTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.finitePairHistoryOfTrace","description":"Restrict infinite action and reward traces to pair-coordinate history.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-2bebbc2b01ce","parent":"module:BanditRLProof.HistoryFiltration","order":5120,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:78"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def finitePairHistoryOfTrace {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : FinitePairHistory Action Reward t","missing":[],"search":"finitepairhistoryoftrace banditrlproof.history.finitepairhistoryoftrace restrict infinite action and reward traces to pair-coordinate history. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.extendPairHistorySucc","label":"extendPairHistorySucc","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.extendPairHistorySucc","description":"Extend a finite pair history from `t` to `t + 1` by appending the next pair.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-c8d1619d6b2d","parent":"module:BanditRLProof.HistoryFiltration","order":5121,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:87"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def extendPairHistorySucc {Action : Type v} {Reward : Type w} {t : Nat} (history : FinitePairHistory Action Reward t) (next : Prod Action Reward) : FinitePairHistory Action Reward (t + 1)","missing":[],"search":"extendpairhistorysucc banditrlproof.history.extendpairhistorysucc extend a finite pair history from `t` to `t + 1` by appending the next pair. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteActionHistoryOfTrace_apply","label":"finiteActionHistoryOfTrace_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.finiteActionHistoryOfTrace_apply","description":"theorem finiteActionHistoryOfTrace_apply {Action : Type v} (action : ActionTrace Action) (t : Nat) (i : Finset.Iic t) : finiteActionHistoryOfTrace action t i = action i.1","url":"../modules/banditrlproof-historyfiltration/index.html#decl-d3a34e56a2f9","parent":"module:BanditRLProof.HistoryFiltration","order":5122,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:99"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteActionHistoryOfTrace_apply {Action : Type v} (action : ActionTrace Action) (t : Nat) (i : Finset.Iic t) : finiteActionHistoryOfTrace action t i = action i.1","missing":[],"search":"finiteactionhistoryoftrace_apply banditrlproof.history.finiteactionhistoryoftrace_apply theorem finiteactionhistoryoftrace_apply {action : type v} (action : actiontrace action) (t : nat) (i : finset.iic t) : finiteactionhistoryoftrace action t i = action i.1 theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteRewardHistoryOfTrace_apply","label":"finiteRewardHistoryOfTrace_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.finiteRewardHistoryOfTrace_apply","description":"theorem finiteRewardHistoryOfTrace_apply {Reward : Type w} (reward : RewardTrace Reward) (t : Nat) (i : Finset.Iic t) : finiteRewardHistoryOfTrace reward t i = reward i.1","url":"../modules/banditrlproof-historyfiltration/index.html#decl-c60cfe2bc20f","parent":"module:BanditRLProof.HistoryFiltration","order":5123,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:105"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteRewardHistoryOfTrace_apply {Reward : Type w} (reward : RewardTrace Reward) (t : Nat) (i : Finset.Iic t) : finiteRewardHistoryOfTrace reward t i = reward i.1","missing":[],"search":"finiterewardhistoryoftrace_apply banditrlproof.history.finiterewardhistoryoftrace_apply theorem finiterewardhistoryoftrace_apply {reward : type w} (reward : rewardtrace reward) (t : nat) (i : finset.iic t) : finiterewardhistoryoftrace reward t i = reward i.1 theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.completeRewardTrace_finiteRewardHistoryOfTrace_apply_of_le","label":"completeRewardTrace_finiteRewardHistoryOfTrace_apply_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.completeRewardTrace_finiteRewardHistoryOfTrace_apply_of_le","description":"Completing the actual finite reward prefix recovers every original coordinate through that prefix.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-d7b50c609fcc","parent":"module:BanditRLProof.HistoryFiltration","order":5124,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:115"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem completeRewardTrace_finiteRewardHistoryOfTrace_apply_of_le {Reward : Type w} (reward : RewardTrace Reward) (t s : Nat) (default : Reward) (hs : s <= t) : completeRewardTrace t (finiteRewardHistoryOfTrace reward t) default s = reward s","missing":[],"search":"completerewardtrace_finiterewardhistoryoftrace_apply_of_le banditrlproof.history.completerewardtrace_finiterewardhistoryoftrace_apply_of_le completing the actual finite reward prefix recovers every original coordinate through that prefix. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteHistoryOfTrace_fst","label":"finiteHistoryOfTrace_fst","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.finiteHistoryOfTrace_fst","description":"theorem finiteHistoryOfTrace_fst {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : (finiteHistoryOfTrace action reward t).1 = finiteActionHistoryOfTrace action t","url":"../modules/banditrlproof-historyfiltration/index.html#decl-d5f1f6b871d9","parent":"module:BanditRLProof.HistoryFiltration","order":5125,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:124"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryOfTrace_fst {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : (finiteHistoryOfTrace action reward t).1 = finiteActionHistoryOfTrace action t","missing":[],"search":"finitehistoryoftrace_fst banditrlproof.history.finitehistoryoftrace_fst theorem finitehistoryoftrace_fst {action : type v} {reward : type w} (action : actiontrace action) (reward : rewardtrace reward) (t : nat) : (finitehistoryoftrace action reward t).1 = finiteactionhistoryoftrace action t theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finiteHistoryOfTrace_snd","label":"finiteHistoryOfTrace_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.finiteHistoryOfTrace_snd","description":"theorem finiteHistoryOfTrace_snd {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : (finiteHistoryOfTrace action reward t).2 = finiteRewardHistoryOfTrace reward t","url":"../modules/banditrlproof-historyfiltration/index.html#decl-a14316f8dbfc","parent":"module:BanditRLProof.HistoryFiltration","order":5126,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:133"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryOfTrace_snd {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : (finiteHistoryOfTrace action reward t).2 = finiteRewardHistoryOfTrace reward t","missing":[],"search":"finitehistoryoftrace_snd banditrlproof.history.finitehistoryoftrace_snd theorem finitehistoryoftrace_snd {action : type v} {reward : type w} (action : actiontrace action) (reward : rewardtrace reward) (t : nat) : (finitehistoryoftrace action reward t).2 = finiterewardhistoryoftrace reward t theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finitePairHistoryOfTrace_apply","label":"finitePairHistoryOfTrace_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.finitePairHistoryOfTrace_apply","description":"theorem finitePairHistoryOfTrace_apply {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) (i : Finset.Iic t) : finitePairHistoryOfTrace action reward t i = (action i.1, reward i.1)","url":"../modules/banditrlproof-historyfiltration/index.html#decl-5515a8fde601","parent":"module:BanditRLProof.HistoryFiltration","order":5127,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:142"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistoryOfTrace_apply {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) (i : Finset.Iic t) : finitePairHistoryOfTrace action reward t i = (action i.1, reward i.1)","missing":[],"search":"finitepairhistoryoftrace_apply banditrlproof.history.finitepairhistoryoftrace_apply theorem finitepairhistoryoftrace_apply {action : type v} {reward : type w} (action : actiontrace action) (reward : rewardtrace reward) (t : nat) (i : finset.iic t) : finitepairhistoryoftrace action reward t i = (action i.1, reward i.1) theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.extendPairHistorySucc_apply_of_le","label":"extendPairHistorySucc_apply_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.extendPairHistorySucc_apply_of_le","description":"theorem extendPairHistorySucc_apply_of_le {Action : Type v} {Reward : Type w} {t : Nat} (history : FinitePairHistory Action Reward t) (next : Prod Action Reward) (i : Finset.Iic (t + 1)) (hi : i.1 <= t) : extendPairHistorySucc history next i = history ⟨i.1, Finset.mem_Iic.mpr hi⟩","url":"../modules/banditrlproof-historyfiltration/index.html#decl-f75c5b79d131","parent":"module:BanditRLProof.HistoryFiltration","order":5128,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:151"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem extendPairHistorySucc_apply_of_le {Action : Type v} {Reward : Type w} {t : Nat} (history : FinitePairHistory Action Reward t) (next : Prod Action Reward) (i : Finset.Iic (t + 1)) (hi : i.1 <= t) : extendPairHistorySucc history next i = history ⟨i.1, Finset.mem_Iic.mpr hi⟩","missing":[],"search":"extendpairhistorysucc_apply_of_le banditrlproof.history.extendpairhistorysucc_apply_of_le theorem extendpairhistorysucc_apply_of_le {action : type v} {reward : type w} {t : nat} (history : finitepairhistory action reward t) (next : prod action reward) (i : finset.iic (t + 1)) (hi : i.1 <= t) : extendpairhistorysucc history next i = history ⟨i.1, finset.mem_iic.mpr hi⟩ theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.extendPairHistorySucc_apply_succ","label":"extendPairHistorySucc_apply_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.extendPairHistorySucc_apply_succ","description":"theorem extendPairHistorySucc_apply_succ {Action : Type v} {Reward : Type w} {t : Nat} (history : FinitePairHistory Action Reward t) (next : Prod Action Reward) : extendPairHistorySucc history next ⟨t + 1, Finset.mem_Iic.mpr le_rfl⟩ = next","url":"../modules/banditrlproof-historyfiltration/index.html#decl-b55347a20669","parent":"module:BanditRLProof.HistoryFiltration","order":5129,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:161"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem extendPairHistorySucc_apply_succ {Action : Type v} {Reward : Type w} {t : Nat} (history : FinitePairHistory Action Reward t) (next : Prod Action Reward) : extendPairHistorySucc history next ⟨t + 1, Finset.mem_Iic.mpr le_rfl⟩ = next","missing":[],"search":"extendpairhistorysucc_apply_succ banditrlproof.history.extendpairhistorysucc_apply_succ theorem extendpairhistorysucc_apply_succ {action : type v} {reward : type w} {t : nat} (history : finitepairhistory action reward t) (next : prod action reward) : extendpairhistorysucc history next ⟨t + 1, finset.mem_iic.mpr le_rfl⟩ = next theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.finitePairHistoryOfTrace_succ","label":"finitePairHistoryOfTrace_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.finitePairHistoryOfTrace_succ","description":"The trace prefix at `t + 1` is the old pair prefix extended by the next action/reward pair.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-b84aa29a15a8","parent":"module:BanditRLProof.HistoryFiltration","order":5130,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:175"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem finitePairHistoryOfTrace_succ {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : finitePairHistoryOfTrace action reward (t + 1) = extendPairHistorySucc (finitePairHistoryOfTrace action reward t) (action (t + 1), reward (t + 1))","missing":[],"search":"finitepairhistoryoftrace_succ banditrlproof.history.finitepairhistoryoftrace_succ the trace prefix at `t + 1` is the old pair prefix extended by the next action/reward pair. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.pairHistoryRewardProjection","label":"pairHistoryRewardProjection","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.pairHistoryRewardProjection","description":"Project the reward coordinates from a prefix of `(Action, Reward)` pairs.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-bbb0cc7ab24d","parent":"module:BanditRLProof.HistoryFiltration","order":5131,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:193"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def pairHistoryRewardProjection {Action : Type v} {Reward : Type w} {t : Nat} (history : (i : Finset.Iic t) -> Prod Action Reward) : FiniteRewardHistory Reward t","missing":[],"search":"pairhistoryrewardprojection banditrlproof.history.pairhistoryrewardprojection project the reward coordinates from a prefix of `(action, reward)` pairs. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.pairHistoryRewardProjection_apply","label":"pairHistoryRewardProjection_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.pairHistoryRewardProjection_apply","description":"theorem pairHistoryRewardProjection_apply {Action : Type v} {Reward : Type w} {t : Nat} (history : (i : Finset.Iic t) -> Prod Action Reward) (i : Finset.Iic t) : pairHistoryRewardProjection history i = (history i).2","url":"../modules/banditrlproof-historyfiltration/index.html#decl-2ee7694a7f89","parent":"module:BanditRLProof.HistoryFiltration","order":5132,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:200"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem pairHistoryRewardProjection_apply {Action : Type v} {Reward : Type w} {t : Nat} (history : (i : Finset.Iic t) -> Prod Action Reward) (i : Finset.Iic t) : pairHistoryRewardProjection history i = (history i).2","missing":[],"search":"pairhistoryrewardprojection_apply banditrlproof.history.pairhistoryrewardprojection_apply theorem pairhistoryrewardprojection_apply {action : type v} {reward : type w} {t : nat} (history : (i : finset.iic t) -> prod action reward) (i : finset.iic t) : pairhistoryrewardprojection history i = (history i).2 theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.pairHistoryRewardProjection_finitePairHistoryOfTrace","label":"pairHistoryRewardProjection_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.pairHistoryRewardProjection_finitePairHistoryOfTrace","description":"theorem pairHistoryRewardProjection_finitePairHistoryOfTrace {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : pairHistoryRewardProjection (finitePairHistoryOfTrace action reward t) = finiteRewardHistoryOfTrace reward t","url":"../modules/banditrlproof-historyfiltration/index.html#decl-afcb80355ce4","parent":"module:BanditRLProof.HistoryFiltration","order":5133,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:207"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem pairHistoryRewardProjection_finitePairHistoryOfTrace {Action : Type v} {Reward : Type w} (action : ActionTrace Action) (reward : RewardTrace Reward) (t : Nat) : pairHistoryRewardProjection (finitePairHistoryOfTrace action reward t) = finiteRewardHistoryOfTrace reward t","missing":[],"search":"pairhistoryrewardprojection_finitepairhistoryoftrace banditrlproof.history.pairhistoryrewardprojection_finitepairhistoryoftrace theorem pairhistoryrewardprojection_finitepairhistoryoftrace {action : type v} {reward : type w} (action : actiontrace action) (reward : rewardtrace reward) (t : nat) : pairhistoryrewardprojection (finitepairhistoryoftrace action reward t) = finiterewardhistoryoftrace reward t theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteActionHistory_eval","label":"measurable_finiteActionHistory_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteActionHistory_eval","description":"Coordinate evaluation on finite action histories is measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-a1e3e64357cc","parent":"module:BanditRLProof.HistoryFiltration","order":5134,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:217"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteActionHistory_eval {Action : Type v} [MeasurableSpace Action] (t : Nat) (i : Finset.Iic t) : Measurable (fun history : FiniteActionHistory Action t => history i)","missing":[],"search":"measurable_finiteactionhistory_eval banditrlproof.history.measurable_finiteactionhistory_eval coordinate evaluation on finite action histories is measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteRewardHistory_eval","label":"measurable_finiteRewardHistory_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteRewardHistory_eval","description":"Coordinate evaluation on finite reward histories is measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-32e60a2df988","parent":"module:BanditRLProof.HistoryFiltration","order":5135,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:224"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteRewardHistory_eval {Reward : Type w} [MeasurableSpace Reward] (t : Nat) (i : Finset.Iic t) : Measurable (fun history : FiniteRewardHistory Reward t => history i)","missing":[],"search":"measurable_finiterewardhistory_eval banditrlproof.history.measurable_finiterewardhistory_eval coordinate evaluation on finite reward histories is measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteHistory_action_eval","label":"measurable_finiteHistory_action_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteHistory_action_eval","description":"Action-coordinate evaluation on paired finite histories is measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-820e3f5f9383","parent":"module:BanditRLProof.HistoryFiltration","order":5136,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:231"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistory_action_eval {Action : Type v} {Reward : Type w} [MeasurableSpace Action] [MeasurableSpace Reward] (t : Nat) (i : Finset.Iic t) : Measurable (fun history : FiniteHistory Action Reward t => history.1 i)","missing":[],"search":"measurable_finitehistory_action_eval banditrlproof.history.measurable_finitehistory_action_eval action-coordinate evaluation on paired finite histories is measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteHistory_reward_eval","label":"measurable_finiteHistory_reward_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteHistory_reward_eval","description":"Reward-coordinate evaluation on paired finite histories is measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-137d5e280c61","parent":"module:BanditRLProof.HistoryFiltration","order":5137,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:241"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistory_reward_eval {Action : Type v} {Reward : Type w} [MeasurableSpace Action] [MeasurableSpace Reward] (t : Nat) (i : Finset.Iic t) : Measurable (fun history : FiniteHistory Action Reward t => history.2 i)","missing":[],"search":"measurable_finitehistory_reward_eval banditrlproof.history.measurable_finitehistory_reward_eval reward-coordinate evaluation on paired finite histories is measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_pairHistoryRewardProjection","label":"measurable_pairHistoryRewardProjection","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_pairHistoryRewardProjection","description":"The reward projection from pair-coordinate histories is measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-523b1adb198d","parent":"module:BanditRLProof.HistoryFiltration","order":5138,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:251"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_pairHistoryRewardProjection {Action : Type v} {Reward : Type w} [MeasurableSpace Action] [MeasurableSpace Reward] (t : Nat) : Measurable (fun history : (i : Finset.Iic t) -> Prod Action Reward => pairHistoryRewardProjection history)","missing":[],"search":"measurable_pairhistoryrewardprojection banditrlproof.history.measurable_pairhistoryrewardprojection the reward projection from pair-coordinate histories is measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finitePairHistoryOfTrace","label":"measurable_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finitePairHistoryOfTrace","description":"Timewise measurable action and reward traces restrict to a measurable pair-coordinate finite history object.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-f64c32c2719e","parent":"module:BanditRLProof.HistoryFiltration","order":5139,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:266"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finitePairHistoryOfTrace {Omega : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega : Omega => finitePairHistoryOfTrace (action omega) (reward omega) t)","missing":[],"search":"measurable_finitepairhistoryoftrace banditrlproof.history.measurable_finitepairhistoryoftrace timewise measurable action and reward traces restrict to a measurable pair-coordinate finite history object. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_extendPairHistorySucc","label":"measurable_extendPairHistorySucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_extendPairHistorySucc","description":"The successor-extension map for pair histories is measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-f1a019ed3ba2","parent":"module:BanditRLProof.HistoryFiltration","order":5140,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:285"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_extendPairHistorySucc {Action : Type v} {Reward : Type w} {t : Nat} [MeasurableSpace Action] [MeasurableSpace Reward] : Measurable (fun input : Prod (FinitePairHistory Action Reward t) (Prod Action Reward) => extendPairHistorySucc input.1 input.2)","missing":[],"search":"measurable_extendpairhistorysucc banditrlproof.history.measurable_extendpairhistorysucc the successor-extension map for pair histories is measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteActionHistoryOfTrace","label":"measurable_finiteActionHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteActionHistoryOfTrace","description":"Timewise measurable action traces restrict to measurable finite action-history objects.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-5ac8f4023c2c","parent":"module:BanditRLProof.HistoryFiltration","order":5141,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:310"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteActionHistoryOfTrace {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (t : Nat) : Measurable (fun omega : Omega => finiteActionHistoryOfTrace (action omega) t)","missing":[],"search":"measurable_finiteactionhistoryoftrace banditrlproof.history.measurable_finiteactionhistoryoftrace timewise measurable action traces restrict to measurable finite action-history objects. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteRewardHistoryOfTrace","label":"measurable_finiteRewardHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteRewardHistoryOfTrace","description":"Timewise measurable reward traces restrict to measurable finite reward-history objects.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-5e38a74c9a53","parent":"module:BanditRLProof.HistoryFiltration","order":5142,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:326"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteRewardHistoryOfTrace {Omega : Type u} {Reward : Type w} [MeasurableSpace Omega] [MeasurableSpace Reward] (reward : Omega -> RewardTrace Reward) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega : Omega => finiteRewardHistoryOfTrace (reward omega) t)","missing":[],"search":"measurable_finiterewardhistoryoftrace banditrlproof.history.measurable_finiterewardhistoryoftrace timewise measurable reward traces restrict to measurable finite reward-history objects. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finiteHistoryOfTrace","label":"measurable_finiteHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finiteHistoryOfTrace","description":"Timewise measurable action and reward traces restrict to a measurable paired finite history object.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-ee20f4ff1b5d","parent":"module:BanditRLProof.HistoryFiltration","order":5143,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:342"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryOfTrace {Omega : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSpace Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : Measurable (fun omega : Omega => finiteHistoryOfTrace (action omega) (reward omega) t)","missing":[],"search":"measurable_finitehistoryoftrace banditrlproof.history.measurable_finitehistoryoftrace timewise measurable action and reward traces restrict to a measurable paired finite history object. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyGenerators","label":"historyGenerators","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.historyGenerators","description":"The singleton-valued action/reward events visible before time `t`. This is deliberately discrete: it is shaped for finite arms and discrete local reward traces such as `Rat`, where singleton events are the current compiled measurability surface.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-6c55665b94d8","parent":"module:BanditRLProof.HistoryFiltration","order":5144,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:367"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyGenerators {Omega : Type u} {Action : Type v} {Reward : Type w} (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (t : Nat) : Set (Set Omega)","missing":[],"search":"historygenerators banditrlproof.history.historygenerators the singleton-valued action/reward events visible before time `t`. this is deliberately discrete: it is shaped for finite arms and discrete local reward traces such as `rat`, where singleton events are the current compiled measurability surface. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyGenerators_mono","label":"historyGenerators_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyGenerators_mono","description":"Past history generators are monotone in the horizon.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-7579ebb695cf","parent":"module:BanditRLProof.HistoryFiltration","order":5145,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:381"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyGenerators_mono {Omega : Type u} {Action : Type v} {Reward : Type w} (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) {t u : Nat} (htu : t <= u) : Set.Subset (historyGenerators action reward t) (historyGenerators action reward u)","missing":[],"search":"historygenerators_mono banditrlproof.history.historygenerators_mono past history generators are monotone in the horizon. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyMeasurableSpace","label":"historyMeasurableSpace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.historyMeasurableSpace","description":"The sigma-algebra generated by past action/reward singleton events.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-036b114e5a51","parent":"module:BanditRLProof.HistoryFiltration","order":5146,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:413"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyMeasurableSpace {Omega : Type u} {Action : Type v} {Reward : Type w} (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (t : Nat) : MeasurableSpace Omega","missing":[],"search":"historymeasurablespace banditrlproof.history.historymeasurablespace the sigma-algebra generated by past action/reward singleton events. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyMeasurableSpace_mono","label":"historyMeasurableSpace_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyMeasurableSpace_mono","description":"The history sigma-algebras are monotone.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-043125922767","parent":"module:BanditRLProof.HistoryFiltration","order":5147,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:421"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyMeasurableSpace_mono {Omega : Type u} {Action : Type v} {Reward : Type w} (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) {t u : Nat} (htu : t <= u) : historyMeasurableSpace action reward t <= historyMeasurableSpace action reward u","missing":[],"search":"historymeasurablespace_mono banditrlproof.history.historymeasurablespace_mono the history sigma-algebras are monotone. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyMeasurableSpace_le","label":"historyMeasurableSpace_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyMeasurableSpace_le","description":"The generated history sigma-algebra is a sub-sigma-algebra of the ambient measurable space when action and reward coordinates are timewise measurable.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-5c58bb30804c","parent":"module:BanditRLProof.HistoryFiltration","order":5148,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:435"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyMeasurableSpace_le {Omega : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : historyMeasurableSpace action reward t <= (inferInstance : MeasurableSpace Omega)","missing":[],"search":"historymeasurablespace_le banditrlproof.history.historymeasurablespace_le the generated history sigma-algebra is a sub-sigma-algebra of the ambient measurable space when action and reward coordinates are timewise measurable. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltration","label":"historyFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.historyFiltration","description":"The local history filtration generated by past action/reward singleton events. This is the compiled `FILTRATION-HISTORY` canary. It is intentionally only a filtration construction, not a policy/predictability or conditional-expectation theorem.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-6203f87e8593","parent":"module:BanditRLProof.HistoryFiltration","order":5149,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:476"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyFiltration {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Filtration Nat mOmega where","missing":[],"search":"historyfiltration banditrlproof.history.historyfiltration the local history filtration generated by past action/reward singleton events. this is the compiled `filtration-history` canary. it is intentionally only a filtration construction, not a policy/predictability or conditional-expectation theorem. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltration_apply","label":"historyFiltration_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyFiltration_apply","description":"theorem historyFiltration_apply {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, M…","url":"../modules/banditrlproof-historyfiltration/index.html#decl-5b6586100d13","parent":"module:BanditRLProof.HistoryFiltration","order":5150,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:493"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyFiltration_apply {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : (historyFiltration action reward haction hreward t : MeasurableSpace Omega) = historyMeasurableSpace action reward t","missing":[],"search":"historyfiltration_apply banditrlproof.history.historyfiltration_apply theorem historyfiltration_apply {omega : type u} {action : type v} {reward : type w} [momega : measurablespace omega] [measurablespace action] [measurablesingletonclass action] [measurablespace reward] [measurablesingletonclass reward] (action : omega -> actiontrace action) (reward : omega -> rewardtrace reward) (haction : forall t : nat, measurable (fun omega : omega => action omega t)) (hreward : forall t : nat, measurable (fun omega : omega => reward omega t)) (t : nat) : (historyfiltration action reward haction hreward t : measurablespace omega) = historymeasurablespace action reward t theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltrationSucc","label":"historyFiltrationSucc","kind":"definition","status":"compiled","subtitle":"BanditRLProof.History.historyFiltrationSucc","description":"The one-step shifted history filtration. `historyFiltration` at index `t` contains observations with index `< t`. For adapted reward increments it is often more convenient to index the filtration by observations available after time `t`, i.e. `< t + 1`.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-e98c0a78ec2f","parent":"module:BanditRLProof.HistoryFiltration","order":5151,"meta":[["Kind","definition"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:516"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyFiltrationSucc {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) : Filtration Nat mOmega where","missing":[],"search":"historyfiltrationsucc banditrlproof.history.historyfiltrationsucc the one-step shifted history filtration. `historyfiltration` at index `t` contains observations with index `< t`. for adapted reward increments it is often more convenient to index the filtration by observations available after time `t`, i.e. `< t + 1`. definition compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltrationSucc_apply","label":"historyFiltrationSucc_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyFiltrationSucc_apply","description":"theorem historyFiltrationSucc_apply {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Na…","url":"../modules/banditrlproof-historyfiltration/index.html#decl-b92cb8621bed","parent":"module:BanditRLProof.HistoryFiltration","order":5152,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:536"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyFiltrationSucc_apply {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (t : Nat) : (historyFiltrationSucc action reward haction hreward t : MeasurableSpace Omega) = historyFiltration action reward haction hreward (t + 1)","missing":[],"search":"historyfiltrationsucc_apply banditrlproof.history.historyfiltrationsucc_apply theorem historyfiltrationsucc_apply {omega : type u} {action : type v} {reward : type w} [momega : measurablespace omega] [measurablespace action] [measurablesingletonclass action] [measurablespace reward] [measurablesingletonclass reward] (action : omega -> actiontrace action) (reward : omega -> rewardtrace reward) (haction : forall t : nat, measurable (fun omega : omega => action omega t)) (hreward : forall t : nat, measurable (fun omega : omega => reward omega t)) (t : nat) : (historyfiltrationsucc action reward haction hreward t : measurablespace omega) = historyfiltration action reward haction hreward (t + 1) theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurableSet_action_mem_historyFiltration","label":"measurableSet_action_mem_historyFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurableSet_action_mem_historyFiltration","description":"Past action singleton events are measurable in the generated history.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-a5bec4162ae0","parent":"module:BanditRLProof.HistoryFiltration","order":5153,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:553"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_action_mem_historyFiltration {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) {i t : Nat} (hit : i < t) (a : Action) : MeasurableSet[historyFiltration action reward haction hreward t] (Set.preimage (fun omega => action omega i) (Set.singleton a))","missing":[],"search":"measurableset_action_mem_historyfiltration banditrlproof.history.measurableset_action_mem_historyfiltration past action singleton events are measurable in the generated history. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_action_mem_historyFiltration_of_lt","label":"measurable_action_mem_historyFiltration_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_action_mem_historyFiltration_of_lt","description":"Past action coordinates are measurable with respect to the generated history. This is a deliberately discrete `ADAPTED-ACTION` canary: it uses countability and singleton-event measurability, not a full policy-predictability contract.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-b8ddd6fbaf61","parent":"module:BanditRLProof.HistoryFiltration","order":5154,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:579"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_action_mem_historyFiltration_of_lt {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) {i t : Nat} (hit : i < t) : @Measurable Omega Action (historyFiltration action reward haction hreward t) inferInstance (fun omega => action omega i)","missing":[],"search":"measurable_action_mem_historyfiltration_of_lt banditrlproof.history.measurable_action_mem_historyfiltration_of_lt past action coordinates are measurable with respect to the generated history. this is a deliberately discrete `adapted-action` canary: it uses countability and singleton-event measurability, not a full policy-predictability contract. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurableSet_reward_mem_historyFiltration","label":"measurableSet_reward_mem_historyFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurableSet_reward_mem_historyFiltration","description":"Past reward singleton events are measurable in the generated history.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-6b1306c80b99","parent":"module:BanditRLProof.HistoryFiltration","order":5155,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:603"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_reward_mem_historyFiltration {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) {i t : Nat} (hit : i < t) (r : Reward) : MeasurableSet[historyFiltration action reward haction hreward t] (Set.preimage (fun omega => reward omega i) (Set.singleton r))","missing":[],"search":"measurableset_reward_mem_historyfiltration banditrlproof.history.measurableset_reward_mem_historyfiltration past reward singleton events are measurable in the generated history. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_reward_mem_historyFiltration_of_lt","label":"measurable_reward_mem_historyFiltration_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_reward_mem_historyFiltration_of_lt","description":"Past reward coordinates are measurable with respect to the generated history. This is the reward-side companion canary for the local history filtration; it still does not instantiate conditional reward laws or martingale differences.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-0e120c0e772d","parent":"module:BanditRLProof.HistoryFiltration","order":5156,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:629"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_reward_mem_historyFiltration_of_lt {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [Countable Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) {i t : Nat} (hit : i < t) : @Measurable Omega Reward (historyFiltration action reward haction hreward t) inferInstance (fun omega => reward omega i)","missing":[],"search":"measurable_reward_mem_historyfiltration_of_lt banditrlproof.history.measurable_reward_mem_historyfiltration_of_lt past reward coordinates are measurable with respect to the generated history. this is the reward-side companion canary for the local history filtration; it still does not instantiate conditional reward laws or martingale differences. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.measurable_finitePairHistoryOfTrace_mem_historyFiltration_of_lt","label":"measurable_finitePairHistoryOfTrace_mem_historyFiltration_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.measurable_finitePairHistoryOfTrace_mem_historyFiltration_of_lt","description":"Finite pair histories up to `n` are measurable with respect to any generated history filtration level strictly after `n`. This is the product-valued counterpart of the coordinate measurability canaries above. It is intentionally countable/discrete, matching the singleton-event definition of `historyMeasurableSpace`.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-7529531f5dab","parent":"module:BanditRLProof.HistoryFiltration","order":5157,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:661"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finitePairHistoryOfTrace_mem_historyFiltration_of_lt {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [Countable Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) {n t : Nat} (hnt : n < t) : @Measurable Omega ((i : Finset.Iic n) -> Prod Action Reward) (historyFiltration (mOmega := mOmega) action reward haction hreward t) inferInstance (fun omega : Omega => finitePairHistoryOfTrace (action omega) (reward omega) n)","missing":[],"search":"measurable_finitepairhistoryoftrace_mem_historyfiltration_of_lt banditrlproof.history.measurable_finitepairhistoryoftrace_mem_historyfiltration_of_lt finite pair histories up to `n` are measurable with respect to any generated history filtration level strictly after `n`. this is the product-valued counterpart of the coordinate measurability canaries above. it is intentionally countable/discrete, matching the singleton-event definition of `historymeasurablespace`. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltration_succ_eq_comap_finitePairHistoryOfTrace","label":"historyFiltration_succ_eq_comap_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyFiltration_succ_eq_comap_finitePairHistoryOfTrace","description":"The generated history filtration after observing indices `<= n` is exactly the comap of the finite pair-history restriction map. The forward inclusion follows from the singleton generators. The reverse inclusion follows from the previous product-valued measurability wrapper. This equality is a local bridge between the hand-rolled history filtration and Mathlib conditional-distribution statements conditioned on finit…","url":"../modules/banditrlproof-historyfiltration/index.html#decl-09ba9536feec","parent":"module:BanditRLProof.HistoryFiltration","order":5158,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:716"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyFiltration_succ_eq_comap_finitePairHistoryOfTrace {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [Countable Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (n : Nat) : (historyFiltration action reward haction hreward (n + 1) : MeasurableSpace Omega) = (inferInstance : MeasurableSpace ((i : Finset.Iic n) -> Prod Action Reward)).comap (fun omega : Omega => finitePairHistoryOfTrace (action omega) (reward omega) n)","missing":[],"search":"historyfiltration_succ_eq_comap_finitepairhistoryoftrace banditrlproof.history.historyfiltration_succ_eq_comap_finitepairhistoryoftrace the generated history filtration after observing indices `<= n` is exactly the comap of the finite pair-history restriction map. the forward inclusion follows from the singleton generators. the reverse inclusion follows from the previous product-valued measurability wrapper. this equality is a local bridge between the hand-rolled history filtration and mathlib conditional-distribution statements conditioned on finite prefixes. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltrationSucc_eq_comap_finitePairHistoryOfTrace","label":"historyFiltrationSucc_eq_comap_finitePairHistoryOfTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyFiltrationSucc_eq_comap_finitePairHistoryOfTrace","description":"The shifted generated-history filtration at time `n` is exactly the comap of the finite pair-history restriction through index `n`. This is the `historyFiltrationSucc`-indexed form used by the conditional-kernel source contracts.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-7da091a6a8cf","parent":"module:BanditRLProof.HistoryFiltration","order":5159,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:819"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyFiltrationSucc_eq_comap_finitePairHistoryOfTrace {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [Countable Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (n : Nat) : (historyFiltrationSucc action reward haction hreward n : MeasurableSpace Omega) = (inferInstance : MeasurableSpace ((i : Finset.Iic n) -> Prod Action Reward)).comap (fun omega : Omega => finitePairHistoryOfTrace (action omega) (reward omega) n)","missing":[],"search":"historyfiltrationsucc_eq_comap_finitepairhistoryoftrace banditrlproof.history.historyfiltrationsucc_eq_comap_finitepairhistoryoftrace the shifted generated-history filtration at time `n` is exactly the comap of the finite pair-history restriction through index `n`. this is the `historyfiltrationsucc`-indexed form used by the conditional-kernel source contracts. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.History.historyFiltrationSucc_eq_of_action_eq_on_prefix","label":"historyFiltrationSucc_eq_of_action_eq_on_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.History.historyFiltrationSucc_eq_of_action_eq_on_prefix","description":"Two shifted history filtrations agree at time `n` when their action traces agree pointwise through `n` and they use the same reward trace. The proof passes through the finite-pair-history comap characterization. It is useful when an adaptive policy has a deterministic exploration prefix.","url":"../modules/banditrlproof-historyfiltration/index.html#decl-8ba37d4521b9","parent":"module:BanditRLProof.HistoryFiltration","order":5160,"meta":[["Kind","theorem"],["Module","BanditRLProof.HistoryFiltration"],["Source","BanditRLProof/HistoryFiltration.lean:849"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyFiltrationSucc_eq_of_action_eq_on_prefix {Omega : Type u} {Action : Type v} {Reward : Type w} [mOmega : MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Countable Action] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [Countable Reward] (action0 action1 : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction0 : forall t : Nat, Measurable (fun omega : Omega => action0 omega t)) (haction1 : forall t : Nat, Measurable (fun omega : Omega => action1 omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (n : Nat) (haction_eq : forall omega i, i <= n -> action0 omega i = action1 omega i) : (historyFiltrationSucc action0 reward haction0 hreward n : MeasurableSpace Omega) = (historyFiltrationSucc action1 reward haction1 hreward n : MeasurableSpace Omega)","missing":[],"search":"historyfiltrationsucc_eq_of_action_eq_on_prefix banditrlproof.history.historyfiltrationsucc_eq_of_action_eq_on_prefix two shifted history filtrations agree at time `n` when their action traces agree pointwise through `n` and they use the same reward trace. the proof passes through the finite-pair-history comap characterization. it is useful when an adaptive policy has a deterministic exploration prefix. theorem compiled","shard":"modules/024bfdf00d4ef14a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.IndependenceFoundation.iIndepFun_infinitePi_coord","label":"iIndepFun_infinitePi_coord","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.IndependenceFoundation.iIndepFun_infinitePi_coord","description":"Coordinate transforms under an infinite product measure form an independent family. This is the generic `IID-REWARD-FAMILY` import wrapper. It is a project-local surface over Mathlib's `ProbabilityTheory.iIndepFun_infinitePi`.","url":"../modules/banditrlproof-independencefoundation/index.html#decl-bcde1d9d1249","parent":"module:BanditRLProof.IndependenceFoundation","order":5161,"meta":[["Kind","theorem"],["Module","BanditRLProof.IndependenceFoundation"],["Source","BanditRLProof/IndependenceFoundation.lean:22"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_infinitePi_coord {Idx : Type u} {Omega : Idx -> Type v} {Target : Idx -> Type w} [mOmega : forall i, MeasurableSpace (Omega i)] [mTarget : forall i, MeasurableSpace (Target i)] (coordLaw : forall i, MeasureTheory.Measure (Omega i)) [forall i, MeasureTheory.IsProbabilityMeasure (coordLaw i)] (X : forall i, Omega i -> Target i) (hX : forall i, Measurable (X i)) : ProbabilityTheory.iIndepFun (fun i omega => X i (omega i)) (MeasureTheory.Measure.infinitePi coordLaw)","missing":[],"search":"iindepfun_infinitepi_coord banditrlproof.independencefoundation.iindepfun_infinitepi_coord coordinate transforms under an infinite product measure form an independent family. this is the generic `iid-reward-family` import wrapper. it is a project-local surface over mathlib's `probabilitytheory.iindepfun_infinitepi`. theorem compiled","shard":"modules/4fba937e92935ca6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.IndependenceFoundation.iIndepFun_rewardTrace_infinitePi","label":"iIndepFun_rewardTrace_infinitePi","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.IndependenceFoundation.iIndepFun_rewardTrace_infinitePi","description":"The coordinate projections of an infinite product reward trace are independent. This is the reward-trace specialization of `iIndepFun_infinitePi_coord`.","url":"../modules/banditrlproof-independencefoundation/index.html#decl-fc8b2e0ce12b","parent":"module:BanditRLProof.IndependenceFoundation","order":5162,"meta":[["Kind","theorem"],["Module","BanditRLProof.IndependenceFoundation"],["Source","BanditRLProof/IndependenceFoundation.lean:44"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_rewardTrace_infinitePi {Reward : Type u} [MeasurableSpace Reward] (coordLaw : Nat -> MeasureTheory.Measure Reward) [forall t : Nat, MeasureTheory.IsProbabilityMeasure (coordLaw t)] : ProbabilityTheory.iIndepFun (fun t (omega : RewardTrace Reward) => omega t) (MeasureTheory.Measure.infinitePi coordLaw)","missing":[],"search":"iindepfun_rewardtrace_infinitepi banditrlproof.independencefoundation.iindepfun_rewardtrace_infinitepi the coordinate projections of an infinite product reward trace are independent. this is the reward-trace specialization of `iindepfun_infinitepi_coord`. theorem compiled","shard":"modules/4fba937e92935ca6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.IntegrabilitySums.integrable_finset_sum","label":"integrable_finset_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.IntegrabilitySums.integrable_finset_sum","description":"Finite sums of integrable terms are integrable. This is the `INT-FINITE-SUM` import wrapper. It is polymorphic in the codomain, following Mathlib's `integrable_finset_sum'`; in bandit applications the codomain is usually `Real`.","url":"../modules/banditrlproof-integrabilitysums/index.html#decl-0f6733a3e68f","parent":"module:BanditRLProof.IntegrabilitySums","order":5163,"meta":[["Kind","theorem"],["Module","BanditRLProof.IntegrabilitySums"],["Source","BanditRLProof/IntegrabilitySums.lean:26"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_finset_sum {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} {E : Type w} [TopologicalSpace E] [ESeminormedAddCommMonoid E] [ContinuousAdd E] (mu : Measure Omega) (s : Finset Idx) (f : Idx -> Omega -> E) (hf : forall i, i ∈ s -> Integrable (f i) mu) : Integrable (fun omega : Omega => s.sum (fun i => f i omega)) mu","missing":[],"search":"integrable_finset_sum banditrlproof.integrabilitysums.integrable_finset_sum finite sums of integrable terms are integrable. this is the `int-finite-sum` import wrapper. it is polymorphic in the codomain, following mathlib's `integrable_finset_sum'`; in bandit applications the codomain is usually `real`. theorem compiled","shard":"modules/5ad915b30d58d24a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.IntegrabilitySums.integrable_univ_sum","label":"integrable_univ_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.IntegrabilitySums.integrable_univ_sum","description":"Finite-type specialization of `integrable_finset_sum`. This version exposes the common finite-arm shape with `(Finset.univ : Finset Idx)`.","url":"../modules/banditrlproof-integrabilitysums/index.html#decl-0b87783c8540","parent":"module:BanditRLProof.IntegrabilitySums","order":5164,"meta":[["Kind","theorem"],["Module","BanditRLProof.IntegrabilitySums"],["Source","BanditRLProof/IntegrabilitySums.lean:49"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_univ_sum {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} [Fintype Idx] {E : Type w} [TopologicalSpace E] [ESeminormedAddCommMonoid E] [ContinuousAdd E] (mu : Measure Omega) (f : Idx -> Omega -> E) (hf : forall i : Idx, Integrable (f i) mu) : Integrable (fun omega : Omega => (Finset.univ : Finset Idx).sum (fun i => f i omega)) mu","missing":[],"search":"integrable_univ_sum banditrlproof.integrabilitysums.integrable_univ_sum finite-type specialization of `integrable_finset_sum`. this version exposes the common finite-arm shape with `(finset.univ : finset idx)`. theorem compiled","shard":"modules/5ad915b30d58d24a.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Measure.compProd_restrict_prod","label":"compProd_restrict_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Measure.compProd_restrict_prod","description":"Restricting both coordinates of a semidirect product is the semidirect product of the restricted base measure and restricted fiber kernel.","url":"../modules/banditrlproof-kernelindependentextension/index.html#decl-d0c941245d10","parent":"module:BanditRLProof.KernelIndependentExtension","order":5165,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelIndependentExtension"],["Source","BanditRLProof/KernelIndependentExtension.lean:21"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem compProd_restrict_prod {A B : Type*} [MeasurableSpace A] [MeasurableSpace B] (mu : Measure A) [SFinite mu] (kernel : Kernel A B) [IsSFiniteKernel kernel] {s : Set A} {t : Set B} (hs : MeasurableSet s) (ht : MeasurableSet t) : (mu ⊗ₘ kernel).restrict (s ×ˢ t) = mu.restrict s ⊗ₘ kernel.restrict ht","missing":[],"search":"compprod_restrict_prod banditrlproof.measure.compprod_restrict_prod restricting both coordinates of a semidirect product is the semidirect product of the restricted base measure and restricted fiber kernel. theorem compiled","shard":"modules/cf893df6918904a0.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Measure.compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","label":"compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Measure.compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","description":"A semidirect-product law restricted to a measurable safe set depends only on the base law on a measurable base safe set and on each kernel's restriction to the corresponding safe fiber.","url":"../modules/banditrlproof-kernelindependentextension/index.html#decl-cfef0d5ecf71","parent":"module:BanditRLProof.KernelIndependentExtension","order":5166,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelIndependentExtension"],["Source","BanditRLProof/KernelIndependentExtension.lean:44"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq {A B : Type*} [MeasurableSpace A] [MeasurableSpace B] {mu nu : Measure A} [SFinite mu] [SFinite nu] {kernel eta : Kernel A B} [IsSFiniteKernel kernel] [IsSFiniteKernel eta] {baseSafe : Set A} {safe : Set (A × B)} (hbaseSafe : MeasurableSet baseSafe) (hsafe : MeasurableSet safe) (hsafe_base : safe ⊆ baseSafe ×ˢ Set.univ) (hbase : mu.restrict baseSafe = nu.restrict baseSafe) (hfiber : ∀ a ∈ baseSafe, (kernel a).restrict (Prod.mk a ⁻¹' safe) = (eta a).restrict (Prod.mk a ⁻¹' safe)) : (mu ⊗ₘ kernel).restrict safe = (nu ⊗ₘ eta).restrict safe","missing":[],"search":"compprod_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq banditrlproof.measure.compprod_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq a semidirect-product law restricted to a measurable safe set depends only on the base law on a measurable base safe set and on each kernel's restriction to the corresponding safe fiber. theorem compiled","shard":"modules/cf893df6918904a0.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Measure.map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","label":"map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Measure.map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","description":"Mapped form of `compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq`: a measurable successor map preserves the safe-set equality.","url":"../modules/banditrlproof-kernelindependentextension/index.html#decl-443d5d20b6a7","parent":"module:BanditRLProof.KernelIndependentExtension","order":5167,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelIndependentExtension"],["Source","BanditRLProof/KernelIndependentExtension.lean:116"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq {A B C : Type*} [MeasurableSpace A] [MeasurableSpace B] [MeasurableSpace C] {mu nu : Measure A} [SFinite mu] [SFinite nu] {kernel eta : Kernel A B} [IsSFiniteKernel kernel] [IsSFiniteKernel eta] (successor : A × B → C) (hsuccessor : Measurable successor) {baseSafe : Set A} {successorSafe : Set C} (hbaseSafe : MeasurableSet baseSafe) (hsuccessorSafe : MeasurableSet successorSafe) (hpreimage_base : successor ⁻¹' successorSafe ⊆ baseSafe ×ˢ Set.univ) (hbase : mu.restrict baseSafe = nu.restrict baseSafe) (hfiber : ∀ a ∈ baseSafe, (kernel a).restrict (Prod.mk a ⁻¹' (successor ⁻¹' successorSafe)) = (eta a).restrict (Prod.mk a ⁻¹' (successor ⁻¹' successorSafe))) : ((mu ⊗ₘ kernel).map successor).restrict successorSafe = ((nu ⊗ₘ eta).map successor).restrict successorSafe","missing":[],"search":"map_compprod_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq banditrlproof.measure.map_compprod_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq mapped form of `compprod_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq`: a measurable successor map preserves the safe-set equality. theorem compiled","shard":"modules/cf893df6918904a0.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.IndepFun.comp_of_map","label":"comp_of_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.IndepFun.comp_of_map","description":"Pull independence on a pushforward measure back along the measurable map that produced that pushforward.","url":"../modules/banditrlproof-kernelindependentextension/index.html#decl-6c6c730b17c6","parent":"module:BanditRLProof.KernelIndependentExtension","order":5168,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelIndependentExtension"],["Source","BanditRLProof/KernelIndependentExtension.lean:145"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem IndepFun.comp_of_map {Omega Sample X Y : Type*} [MeasurableSpace Omega] [MeasurableSpace Sample] [MeasurableSpace X] [MeasurableSpace Y] {mu : Measure Omega} {z : Omega -> Sample} {x : Sample -> X} {y : Sample -> Y} (hz : Measurable z) (hx : Measurable x) (hy : Measurable y) (hindep : IndepFun x y (mu.map z)) : IndepFun (x ∘ z) (y ∘ z) mu","missing":[],"search":"comp_of_map banditrlproof.indepfun.comp_of_map pull independence on a pushforward measure back along the measurable map that produced that pushforward. theorem compiled","shard":"modules/cf893df6918904a0.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.indepFun_fst_snd_compProd_comap_of_indepFun","label":"indepFun_fst_snd_compProd_comap_of_indepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.indepFun_fst_snd_compProd_comap_of_indepFun","description":"If `x` is independent of `past`, adjoining a Markov-kernel output whose law only depends on `past` leaves `x` independent of that output.","url":"../modules/banditrlproof-kernelindependentextension/index.html#decl-4b21c0ca5a1c","parent":"module:BanditRLProof.KernelIndependentExtension","order":5169,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelIndependentExtension"],["Source","BanditRLProof/KernelIndependentExtension.lean:165"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem indepFun_fst_snd_compProd_comap_of_indepFun {Omega X Past Output : Type*} [MeasurableSpace Omega] [MeasurableSpace X] [MeasurableSpace Past] [MeasurableSpace Output] (mu : Measure Omega) [IsProbabilityMeasure mu] (x : Omega -> X) (hx : Measurable x) (past : Omega -> Past) (hpast : Measurable past) (kernel : Kernel Past Output) [IsMarkovKernel kernel] (hindep : IndepFun x past mu) : IndepFun (x ∘ Prod.fst) Prod.snd (mu ⊗ₘ kernel.comap past hpast)","missing":[],"search":"indepfun_fst_snd_compprod_comap_of_indepfun banditrlproof.indepfun_fst_snd_compprod_comap_of_indepfun if `x` is independent of `past`, adjoining a markov-kernel output whose law only depends on `past` leaves `x` independent of that output. theorem compiled","shard":"modules/cf893df6918904a0.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.map_snd_x_compProd_comap_eq_prod_map_of_indepFun","label":"map_snd_x_compProd_comap_eq_prod_map_of_indepFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.map_snd_x_compProd_comap_eq_prod_map_of_indepFun","description":"If `x` is independent of `past`, adjoining an output through a finite kernel that only sees `past` gives the joint output/`x` law as the output marginal times the original `x` marginal. Unlike the preceding `IndepFun` wrapper, this statement remains valid for subprobability kernels such as a branch-restricted Markov kernel.","url":"../modules/banditrlproof-kernelindependentextension/index.html#decl-a91a32de1587","parent":"module:BanditRLProof.KernelIndependentExtension","order":5170,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelIndependentExtension"],["Source","BanditRLProof/KernelIndependentExtension.lean:215"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem map_snd_x_compProd_comap_eq_prod_map_of_indepFun {Omega X Past Output : Type*} [MeasurableSpace Omega] [MeasurableSpace X] [MeasurableSpace Past] [MeasurableSpace Output] (mu : Measure Omega) [IsProbabilityMeasure mu] (x : Omega -> X) (hx : Measurable x) (past : Omega -> Past) (hpast : Measurable past) (kernel : Kernel Past Output) [IsFiniteKernel kernel] (hindep : IndepFun x past mu) : Measure.map (fun sample : Omega × Output => (sample.2, x sample.1)) (mu ⊗ₘ kernel.comap past hpast) = (Measure.map Prod.snd (mu ⊗ₘ kernel.comap past hpast)).prod (Measure.map x mu)","missing":[],"search":"map_snd_x_compprod_comap_eq_prod_map_of_indepfun banditrlproof.map_snd_x_compprod_comap_eq_prod_map_of_indepfun if `x` is independent of `past`, adjoining an output through a finite kernel that only sees `past` gives the joint output/`x` law as the output marginal times the original `x` marginal. unlike the preceding `indepfun` wrapper, this statement remains valid for subprobability kernels such as a branch-restricted markov kernel. theorem compiled","shard":"modules/cf893df6918904a0.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.KernelTrajectoryPrefix.partialTraj_zero_congr","label":"partialTraj_zero_congr","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KernelTrajectoryPrefix.partialTraj_zero_congr","description":"Two partial trajectories from time zero agree through `n` when their step kernels agree strictly before `n`.","url":"../modules/banditrlproof-kerneltrajectoryprefix/index.html#decl-5cf5329e8f42","parent":"module:BanditRLProof.KernelTrajectoryPrefix","order":5171,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelTrajectoryPrefix"],["Source","BanditRLProof/KernelTrajectoryPrefix.lean:22"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem partialTraj_zero_congr {X : Nat -> Type u} [forall n, MeasurableSpace (X n)] (kappa eta : (n : Nat) -> Kernel ((i : Finset.Iic n) -> X i) (X (n + 1))) [forall n, IsMarkovKernel (kappa n)] [forall n, IsMarkovKernel (eta n)] (n : Nat) (hstep : forall k, k < n -> kappa k = eta k) : Kernel.partialTraj kappa 0 n = Kernel.partialTraj eta 0 n","missing":[],"search":"partialtraj_zero_congr banditrlproof.kerneltrajectoryprefix.partialtraj_zero_congr two partial trajectories from time zero agree through `n` when their step kernels agree strictly before `n`. theorem compiled","shard":"modules/3a808c9e36eff4fd.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.KernelTrajectoryPrefix.trajMeasure_map_frestrictLe_congr","label":"trajMeasure_map_frestrictLe_congr","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.KernelTrajectoryPrefix.trajMeasure_map_frestrictLe_congr","description":"The finite marginal of an Ionescu-Tulcea trajectory depends only on its initial measure and the step kernels strictly before the endpoint.","url":"../modules/banditrlproof-kerneltrajectoryprefix/index.html#decl-5e0402ba09b9","parent":"module:BanditRLProof.KernelTrajectoryPrefix","order":5172,"meta":[["Kind","theorem"],["Module","BanditRLProof.KernelTrajectoryPrefix"],["Source","BanditRLProof/KernelTrajectoryPrefix.lean:41"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem trajMeasure_map_frestrictLe_congr {X : Nat -> Type u} [forall n, MeasurableSpace (X n)] (mu0 nu0 : Measure (X 0)) [IsProbabilityMeasure mu0] [IsProbabilityMeasure nu0] (kappa eta : (n : Nat) -> Kernel ((i : Finset.Iic n) -> X i) (X (n + 1))) [forall n, IsMarkovKernel (kappa n)] [forall n, IsMarkovKernel (eta n)] (n : Nat) (hinitial : mu0 = nu0) (hstep : forall k, k < n -> kappa k = eta k) : (Kernel.trajMeasure mu0 kappa).map (Preorder.frestrictLe n) = (Kernel.trajMeasure nu0 eta).map (Preorder.frestrictLe n)","missing":[],"search":"trajmeasure_map_frestrictle_congr banditrlproof.kerneltrajectoryprefix.trajmeasure_map_frestrictle_congr the finite marginal of an ionescu-tulcea trajectory depends only on its initial measure and the step kernels strictly before the endpoint. theorem compiled","shard":"modules/3a808c9e36eff4fd.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_one","label":"pullCount_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_one","description":"@[simp] theorem pullCount_one : pullCount action a 1 = if action 0 = a then 1 else 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-f22cfc1186c8","parent":"module:BanditRLProof.LeafLemmas","order":5173,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pullCount_one : pullCount action a 1 = if action 0 = a then 1 else 0","missing":[],"search":"pullcount_one banditrlproof.pullcount_one @[simp] theorem pullcount_one : pullcount action a 1 = if action 0 = a then 1 else 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_succ_of_eq","label":"pullCount_succ_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_succ_of_eq","description":"theorem pullCount_succ_of_eq (h : action t = a) : pullCount action a (t + 1) = pullCount action a t + 1","url":"../modules/banditrlproof-leaflemmas/index.html#decl-a983db10bccf","parent":"module:BanditRLProof.LeafLemmas","order":5174,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_succ_of_eq (h : action t = a) : pullCount action a (t + 1) = pullCount action a t + 1","missing":[],"search":"pullcount_succ_of_eq banditrlproof.pullcount_succ_of_eq theorem pullcount_succ_of_eq (h : action t = a) : pullcount action a (t + 1) = pullcount action a t + 1 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_succ_of_ne","label":"pullCount_succ_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_succ_of_ne","description":"theorem pullCount_succ_of_ne (h : action t ≠ a) : pullCount action a (t + 1) = pullCount action a t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-9e1cb61221fb","parent":"module:BanditRLProof.LeafLemmas","order":5175,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_succ_of_ne (h : action t ≠ a) : pullCount action a (t + 1) = pullCount action a t","missing":[],"search":"pullcount_succ_of_ne banditrlproof.pullcount_succ_of_ne theorem pullcount_succ_of_ne (h : action t ≠ a) : pullcount action a (t + 1) = pullcount action a t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_le_succ","label":"pullCount_le_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_le_succ","description":"theorem pullCount_le_succ : pullCount action a t ≤ pullCount action a (t + 1)","url":"../modules/banditrlproof-leaflemmas/index.html#decl-29f3cfecba80","parent":"module:BanditRLProof.LeafLemmas","order":5176,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_succ : pullCount action a t ≤ pullCount action a (t + 1)","missing":[],"search":"pullcount_le_succ banditrlproof.pullcount_le_succ theorem pullcount_le_succ : pullcount action a t ≤ pullcount action a (t + 1) theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_succ_le_succ","label":"pullCount_succ_le_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_succ_le_succ","description":"theorem pullCount_succ_le_succ : pullCount action a (t + 1) ≤ pullCount action a t + 1","url":"../modules/banditrlproof-leaflemmas/index.html#decl-242e82d199b9","parent":"module:BanditRLProof.LeafLemmas","order":5177,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_succ_le_succ : pullCount action a (t + 1) ≤ pullCount action a t + 1","missing":[],"search":"pullcount_succ_le_succ banditrlproof.pullcount_succ_le_succ theorem pullcount_succ_le_succ : pullcount action a (t + 1) ≤ pullcount action a t + 1 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_lt_of_forall_succ_ne","label":"pullCount_lt_of_forall_succ_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_lt_of_forall_succ_ne","description":"If a positive count level is never hit at a successor time, every finite pull count stays strictly below that level. The proof uses only that `pullCount` starts at zero and grows by at most one per round.","url":"../modules/banditrlproof-leaflemmas/index.html#decl-f9d0282ec437","parent":"module:BanditRLProof.LeafLemmas","order":5178,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_lt_of_forall_succ_ne (target : Nat) (htarget : 0 < target) (hnever : ∀ chron, pullCount action a (chron + 1) ≠ target) : pullCount action a t < target","missing":[],"search":"pullcount_lt_of_forall_succ_ne banditrlproof.pullcount_lt_of_forall_succ_ne if a positive count level is never hit at a successor time, every finite pull count stays strictly below that level. the proof uses only that `pullcount` starts at zero and grows by at most one per round. theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_mono","label":"pullCount_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_mono","description":"theorem pullCount_mono {s t : Nat} (h : s ≤ t) : pullCount action a s ≤ pullCount action a t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-d1458009bbe5","parent":"module:BanditRLProof.LeafLemmas","order":5179,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_mono {s t : Nat} (h : s ≤ t) : pullCount action a s ≤ pullCount action a t","missing":[],"search":"pullcount_mono banditrlproof.pullcount_mono theorem pullcount_mono {s t : nat} (h : s ≤ t) : pullcount action a s ≤ pullcount action a t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_le_time","label":"pullCount_le_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_le_time","description":"theorem pullCount_le_time : pullCount action a t ≤ t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-4fb1a2bee033","parent":"module:BanditRLProof.LeafLemmas","order":5180,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_time : pullCount action a t ≤ t","missing":[],"search":"pullcount_le_time banditrlproof.pullcount_le_time theorem pullcount_le_time : pullcount action a t ≤ t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_add_le","label":"pullCount_add_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_add_le","description":"theorem pullCount_add_le (n : Nat) : pullCount action a (t + n) ≤ pullCount action a t + n","url":"../modules/banditrlproof-leaflemmas/index.html#decl-a0c08e2c7170","parent":"module:BanditRLProof.LeafLemmas","order":5181,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:73"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_add_le (n : Nat) : pullCount action a (t + n) ≤ pullCount action a t + n","missing":[],"search":"pullcount_add_le banditrlproof.pullcount_add_le theorem pullcount_add_le (n : nat) : pullcount action a (t + n) ≤ pullcount action a t + n theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_le_add","label":"pullCount_le_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_le_add","description":"theorem pullCount_le_add : pullCount action a t ≤ pullCount action a (t + n)","url":"../modules/banditrlproof-leaflemmas/index.html#decl-9495e71cf3f4","parent":"module:BanditRLProof.LeafLemmas","order":5182,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_add : pullCount action a t ≤ pullCount action a (t + n)","missing":[],"search":"pullcount_le_add banditrlproof.pullcount_le_add theorem pullcount_le_add : pullcount action a t ≤ pullcount action a (t + n) theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_eq_zero_of_forall_ne","label":"pullCount_eq_zero_of_forall_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_eq_zero_of_forall_ne","description":"theorem pullCount_eq_zero_of_forall_ne (h : ∀ s, s < t → action s ≠ a) : pullCount action a t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-ed50736c90bd","parent":"module:BanditRLProof.LeafLemmas","order":5183,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:86"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_eq_zero_of_forall_ne (h : ∀ s, s < t → action s ≠ a) : pullCount action a t = 0","missing":[],"search":"pullcount_eq_zero_of_forall_ne banditrlproof.pullcount_eq_zero_of_forall_ne theorem pullcount_eq_zero_of_forall_ne (h : ∀ s, s < t → action s ≠ a) : pullcount action a t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_eq_time_of_forall_eq","label":"pullCount_eq_time_of_forall_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_eq_time_of_forall_eq","description":"theorem pullCount_eq_time_of_forall_eq (h : ∀ s, s < t → action s = a) : pullCount action a t = t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-dfc0c3ae8eb7","parent":"module:BanditRLProof.LeafLemmas","order":5184,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_eq_time_of_forall_eq (h : ∀ s, s < t → action s = a) : pullCount action a t = t","missing":[],"search":"pullcount_eq_time_of_forall_eq banditrlproof.pullcount_eq_time_of_forall_eq theorem pullcount_eq_time_of_forall_eq (h : ∀ s, s < t → action s = a) : pullcount action a t = t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_pos_of_eq_before","label":"pullCount_pos_of_eq_before","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_pos_of_eq_before","description":"theorem pullCount_pos_of_eq_before {s t : Nat} (hst : s < t) (h : action s = a) : 0 < pullCount action a t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-3077511280e5","parent":"module:BanditRLProof.LeafLemmas","order":5185,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_pos_of_eq_before {s t : Nat} (hst : s < t) (h : action s = a) : 0 < pullCount action a t","missing":[],"search":"pullcount_pos_of_eq_before banditrlproof.pullcount_pos_of_eq_before theorem pullcount_pos_of_eq_before {s t : nat} (hst : s < t) (h : action s = a) : 0 < pullcount action a t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_eq_of_forall_lt","label":"pullCount_eq_of_forall_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_eq_of_forall_lt","description":"Pull counts depend only on the half-open action prefix `0, ..., t - 1`. This keeps later adaptive-policy wrappers from reproving the same induction when a history-generated trace is known to agree pointwise with an index-policy trace up to a finite horizon.","url":"../modules/banditrlproof-leaflemmas/index.html#decl-d79ef2e13388","parent":"module:BanditRLProof.LeafLemmas","order":5186,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_eq_of_forall_lt (action action' : ActionTrace Action) (a : Action) : forall t : Nat, (forall s : Nat, s < t -> action s = action' s) -> pullCount action a t = pullCount action' a t","missing":[],"search":"pullcount_eq_of_forall_lt banditrlproof.pullcount_eq_of_forall_lt pull counts depend only on the half-open action prefix `0, ..., t - 1`. this keeps later adaptive-policy wrappers from reproving the same induction when a history-generated trace is known to agree pointwise with an index-policy trace up to a finite horizon. theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_const_self","label":"pullCount_const_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_const_self","description":"@[simp] theorem pullCount_const_self (a : Action) (t : Nat) : pullCount (fun _ => a) a t = t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-8925fcd4f9a7","parent":"module:BanditRLProof.LeafLemmas","order":5187,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pullCount_const_self (a : Action) (t : Nat) : pullCount (fun _ => a) a t = t","missing":[],"search":"pullcount_const_self banditrlproof.pullcount_const_self @[simp] theorem pullcount_const_self (a : action) (t : nat) : pullcount (fun _ => a) a t = t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_const_of_ne","label":"pullCount_const_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_const_of_ne","description":"theorem pullCount_const_of_ne (b : Action) (h : b ≠ a) (t : Nat) : pullCount (fun _ => b) a t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-460aa31c1bcd","parent":"module:BanditRLProof.LeafLemmas","order":5188,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:145"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_const_of_ne (b : Action) (h : b ≠ a) (t : Nat) : pullCount (fun _ => b) a t = 0","missing":[],"search":"pullcount_const_of_ne banditrlproof.pullcount_const_of_ne theorem pullcount_const_of_ne (b : action) (h : b ≠ a) (t : nat) : pullcount (fun _ => b) a t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_add_eq_of_forall_ne_between","label":"pullCount_add_eq_of_forall_ne_between","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_add_eq_of_forall_ne_between","description":"theorem pullCount_add_eq_of_forall_ne_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s ≠ a) : pullCount action a (t + n) = pullCount action a t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-5b93e91e22f1","parent":"module:BanditRLProof.LeafLemmas","order":5189,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_add_eq_of_forall_ne_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s ≠ a) : pullCount action a (t + n) = pullCount action a t","missing":[],"search":"pullcount_add_eq_of_forall_ne_between banditrlproof.pullcount_add_eq_of_forall_ne_between theorem pullcount_add_eq_of_forall_ne_between (n : nat) (h : ∀ s, t ≤ s → s < t + n → action s ≠ a) : pullcount action a (t + n) = pullcount action a t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_add_eq_add_of_forall_eq_between","label":"pullCount_add_eq_add_of_forall_eq_between","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_add_eq_add_of_forall_eq_between","description":"theorem pullCount_add_eq_add_of_forall_eq_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s = a) : pullCount action a (t + n) = pullCount action a t + n","url":"../modules/banditrlproof-leaflemmas/index.html#decl-b99401859987","parent":"module:BanditRLProof.LeafLemmas","order":5190,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:163"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_add_eq_add_of_forall_eq_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s = a) : pullCount action a (t + n) = pullCount action a t + n","missing":[],"search":"pullcount_add_eq_add_of_forall_eq_between banditrlproof.pullcount_add_eq_add_of_forall_eq_between theorem pullcount_add_eq_add_of_forall_eq_between (n : nat) (h : ∀ s, t ≤ s → s < t + n → action s = a) : pullcount action a (t + n) = pullcount action a t + n theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_eq_list_filter_length","label":"pullCount_eq_list_filter_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_eq_list_filter_length","description":"The recursive pull count equals the number of matching actions in the half-open time prefix `0, ..., t - 1`. This is intentionally a dependency-light `List.range` bridge. The Mathlib `Finset.range` cardinality wrapper is a separate downstream leaf.","url":"../modules/banditrlproof-leaflemmas/index.html#decl-7895c29e9d49","parent":"module:BanditRLProof.LeafLemmas","order":5191,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:183"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_eq_list_filter_length : pullCount action a t = ((List.range t).filter (fun s : Nat => decide (action s = a))).length","missing":[],"search":"pullcount_eq_list_filter_length banditrlproof.pullcount_eq_list_filter_length the recursive pull count equals the number of matching actions in the half-open time prefix `0, ..., t - 1`. this is intentionally a dependency-light `list.range` bridge. the mathlib `finset.range` cardinality wrapper is a separate downstream leaf. theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_succ_of_eq","label":"sumRewards_succ_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_succ_of_eq","description":"theorem sumRewards_succ_of_eq (h : action t = a) : sumRewards action reward a (t + 1) = sumRewards action reward a t + reward t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-687148a61da3","parent":"module:BanditRLProof.LeafLemmas","order":5192,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:204"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_succ_of_eq (h : action t = a) : sumRewards action reward a (t + 1) = sumRewards action reward a t + reward t","missing":[],"search":"sumrewards_succ_of_eq banditrlproof.sumrewards_succ_of_eq theorem sumrewards_succ_of_eq (h : action t = a) : sumrewards action reward a (t + 1) = sumrewards action reward a t + reward t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_succ_of_ne","label":"sumRewards_succ_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_succ_of_ne","description":"theorem sumRewards_succ_of_ne (hzero : ∀ x : Reward, x + 0 = x) (h : action t ≠ a) : sumRewards action reward a (t + 1) = sumRewards action reward a t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-4a061607bfbe","parent":"module:BanditRLProof.LeafLemmas","order":5193,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:209"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_succ_of_ne (hzero : ∀ x : Reward, x + 0 = x) (h : action t ≠ a) : sumRewards action reward a (t + 1) = sumRewards action reward a t","missing":[],"search":"sumrewards_succ_of_ne banditrlproof.sumrewards_succ_of_ne theorem sumrewards_succ_of_ne (hzero : ∀ x : reward, x + 0 = x) (h : action t ≠ a) : sumrewards action reward a (t + 1) = sumrewards action reward a t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_eq_zero_of_forall_ne","label":"sumRewards_eq_zero_of_forall_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_eq_zero_of_forall_ne","description":"theorem sumRewards_eq_zero_of_forall_ne (hzero : ∀ x : Reward, x + 0 = x) (h : ∀ s, s < t → action s ≠ a) : sumRewards action reward a t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-3ab51b50e9de","parent":"module:BanditRLProof.LeafLemmas","order":5194,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:215"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_eq_zero_of_forall_ne (hzero : ∀ x : Reward, x + 0 = x) (h : ∀ s, s < t → action s ≠ a) : sumRewards action reward a t = 0","missing":[],"search":"sumrewards_eq_zero_of_forall_ne banditrlproof.sumrewards_eq_zero_of_forall_ne theorem sumrewards_eq_zero_of_forall_ne (hzero : ∀ x : reward, x + 0 = x) (h : ∀ s, s < t → action s ≠ a) : sumrewards action reward a t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_const_of_ne","label":"sumRewards_const_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_const_of_ne","description":"theorem sumRewards_const_of_ne (hzero : ∀ x : Reward, x + 0 = x) (b : Action) (h : b ≠ a) (t : Nat) : sumRewards (fun _ => b) reward a t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-1f73079aab66","parent":"module:BanditRLProof.LeafLemmas","order":5195,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:224"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_const_of_ne (hzero : ∀ x : Reward, x + 0 = x) (b : Action) (h : b ≠ a) (t : Nat) : sumRewards (fun _ => b) reward a t = 0","missing":[],"search":"sumrewards_const_of_ne banditrlproof.sumrewards_const_of_ne theorem sumrewards_const_of_ne (hzero : ∀ x : reward, x + 0 = x) (b : action) (h : b ≠ a) (t : nat) : sumrewards (fun _ => b) reward a t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_add_eq_of_forall_ne_between","label":"sumRewards_add_eq_of_forall_ne_between","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_add_eq_of_forall_ne_between","description":"theorem sumRewards_add_eq_of_forall_ne_between (hzero : ∀ x : Reward, x + 0 = x) (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s ≠ a) : sumRewards action reward a (t + n) = sumRewards action reward a t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-2a57a446e0a8","parent":"module:BanditRLProof.LeafLemmas","order":5196,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:232"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_add_eq_of_forall_ne_between (hzero : ∀ x : Reward, x + 0 = x) (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s ≠ a) : sumRewards action reward a (t + n) = sumRewards action reward a t","missing":[],"search":"sumrewards_add_eq_of_forall_ne_between banditrlproof.sumrewards_add_eq_of_forall_ne_between theorem sumrewards_add_eq_of_forall_ne_between (hzero : ∀ x : reward, x + 0 = x) (n : nat) (h : ∀ s, t ≤ s → s < t + n → action s ≠ a) : sumrewards action reward a (t + n) = sumrewards action reward a t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_eq_list_range_foldl","label":"sumRewards_eq_list_range_foldl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_eq_list_range_foldl","description":"The recursive reward sum equals a left fold over the half-open time prefix. This bridge deliberately keeps the false-branch `+ 0` steps in the fold, so it does not require additive laws beyond the weak assumptions used by `sumRewards`.","url":"../modules/banditrlproof-leaflemmas/index.html#decl-a6252cce7376","parent":"module:BanditRLProof.LeafLemmas","order":5197,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:250"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_eq_list_range_foldl : sumRewards action reward a t = (List.range t).foldl (fun acc s => acc + if action s = a then reward s else 0) 0","missing":[],"search":"sumrewards_eq_list_range_foldl banditrlproof.sumrewards_eq_list_range_foldl the recursive reward sum equals a left fold over the half-open time prefix. this bridge deliberately keeps the false-branch `+ 0` steps in the fold, so it does not require additive laws beyond the weak assumptions used by `sumrewards`. theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_eq_list_range_filter_foldl","label":"sumRewards_eq_list_range_filter_foldl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_eq_list_range_filter_foldl","description":"The reward sum can also drop nonmatching time steps from the list fold when the accumulator has a right-zero law. This is still dependency-light: it uses `List.range` and `List.filter`, not a Mathlib `Finset` sum.","url":"../modules/banditrlproof-leaflemmas/index.html#decl-05c07fbd2547","parent":"module:BanditRLProof.LeafLemmas","order":5198,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:268"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_eq_list_range_filter_foldl (hzero : ∀ x : Reward, x + 0 = x) : sumRewards action reward a t = ((List.range t).filter (fun s : Nat => decide (action s = a))).foldl (fun acc s => acc + reward s) 0","missing":[],"search":"sumrewards_eq_list_range_filter_foldl banditrlproof.sumrewards_eq_list_range_filter_foldl the reward sum can also drop nonmatching time steps from the list fold when the accumulator has a right-zero law. this is still dependency-light: it uses `list.range` and `list.filter`, not a mathlib `finset` sum. theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.bestMean_eq_mean_bestArm","label":"bestMean_eq_mean_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.bestMean_eq_mean_bestArm","description":"@[simp] theorem bestMean_eq_mean_bestArm (model : FiniteBanditModel K) : model.bestMean = model.mean model.bestArm","url":"../modules/banditrlproof-leaflemmas/index.html#decl-013a419b6b91","parent":"module:BanditRLProof.LeafLemmas","order":5199,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:287"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem bestMean_eq_mean_bestArm (model : FiniteBanditModel K) : model.bestMean = model.mean model.bestArm","missing":[],"search":"bestmean_eq_mean_bestarm banditrlproof.finitebanditmodel.bestmean_eq_mean_bestarm @[simp] theorem bestmean_eq_mean_bestarm (model : finitebanditmodel k) : model.bestmean = model.mean model.bestarm theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteBanditModel.gap_of_ne_bestArm","label":"gap_of_ne_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteBanditModel.gap_of_ne_bestArm","description":"theorem gap_of_ne_bestArm (model : FiniteBanditModel K) (arm : Fin K) (h : arm ≠ model.bestArm) : model.gap arm = model.bestMean - model.mean arm","url":"../modules/banditrlproof-leaflemmas/index.html#decl-725a420fd59f","parent":"module:BanditRLProof.LeafLemmas","order":5200,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:290"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gap_of_ne_bestArm (model : FiniteBanditModel K) (arm : Fin K) (h : arm ≠ model.bestArm) : model.gap arm = model.bestMean - model.mean arm","missing":[],"search":"gap_of_ne_bestarm banditrlproof.finitebanditmodel.gap_of_ne_bestarm theorem gap_of_ne_bestarm (model : finitebanditmodel k) (arm : fin k) (h : arm ≠ model.bestarm) : model.gap arm = model.bestmean - model.mean arm theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_one","label":"pseudoRegret_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_one","description":"@[simp] theorem pseudoRegret_one : pseudoRegret model action 1 = model.gap (action 0)","url":"../modules/banditrlproof-leaflemmas/index.html#decl-c406067214e2","parent":"module:BanditRLProof.LeafLemmas","order":5201,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:301"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pseudoRegret_one : pseudoRegret model action 1 = model.gap (action 0)","missing":[],"search":"pseudoregret_one banditrlproof.pseudoregret_one @[simp] theorem pseudoregret_one : pseudoregret model action 1 = model.gap (action 0) theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_succ_of_bestArm","label":"pseudoRegret_succ_of_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_succ_of_bestArm","description":"theorem pseudoRegret_succ_of_bestArm (h : action t = model.bestArm) : pseudoRegret model action (t + 1) = pseudoRegret model action t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-e8fe5277f15a","parent":"module:BanditRLProof.LeafLemmas","order":5202,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:306"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_succ_of_bestArm (h : action t = model.bestArm) : pseudoRegret model action (t + 1) = pseudoRegret model action t","missing":[],"search":"pseudoregret_succ_of_bestarm banditrlproof.pseudoregret_succ_of_bestarm theorem pseudoregret_succ_of_bestarm (h : action t = model.bestarm) : pseudoregret model action (t + 1) = pseudoregret model action t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_succ_of_gap_zero","label":"pseudoRegret_succ_of_gap_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_succ_of_gap_zero","description":"theorem pseudoRegret_succ_of_gap_zero (h : model.gap (action t) = 0) : pseudoRegret model action (t + 1) = pseudoRegret model action t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-11d4f5f5177d","parent":"module:BanditRLProof.LeafLemmas","order":5203,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:311"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_succ_of_gap_zero (h : model.gap (action t) = 0) : pseudoRegret model action (t + 1) = pseudoRegret model action t","missing":[],"search":"pseudoregret_succ_of_gap_zero banditrlproof.pseudoregret_succ_of_gap_zero theorem pseudoregret_succ_of_gap_zero (h : model.gap (action t) = 0) : pseudoregret model action (t + 1) = pseudoregret model action t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_eq_zero_of_forall_bestArm","label":"pseudoRegret_eq_zero_of_forall_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_eq_zero_of_forall_bestArm","description":"theorem pseudoRegret_eq_zero_of_forall_bestArm (h : ∀ s, s < t → action s = model.bestArm) : pseudoRegret model action t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-0758075753de","parent":"module:BanditRLProof.LeafLemmas","order":5204,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:316"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_eq_zero_of_forall_bestArm (h : ∀ s, s < t → action s = model.bestArm) : pseudoRegret model action t = 0","missing":[],"search":"pseudoregret_eq_zero_of_forall_bestarm banditrlproof.pseudoregret_eq_zero_of_forall_bestarm theorem pseudoregret_eq_zero_of_forall_bestarm (h : ∀ s, s < t → action s = model.bestarm) : pseudoregret model action t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_eq_zero_of_forall_gap_zero","label":"pseudoRegret_eq_zero_of_forall_gap_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_eq_zero_of_forall_gap_zero","description":"theorem pseudoRegret_eq_zero_of_forall_gap_zero (h : ∀ s, s < t → model.gap (action s) = 0) : pseudoRegret model action t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-ef8fd252c5fd","parent":"module:BanditRLProof.LeafLemmas","order":5205,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:325"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_eq_zero_of_forall_gap_zero (h : ∀ s, s < t → model.gap (action s) = 0) : pseudoRegret model action t = 0","missing":[],"search":"pseudoregret_eq_zero_of_forall_gap_zero banditrlproof.pseudoregret_eq_zero_of_forall_gap_zero theorem pseudoregret_eq_zero_of_forall_gap_zero (h : ∀ s, s < t → model.gap (action s) = 0) : pseudoregret model action t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_const_bestArm","label":"pseudoRegret_const_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_const_bestArm","description":"@[simp] theorem pseudoRegret_const_bestArm : pseudoRegret model (fun _ => model.bestArm) t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-dc5ccd7709a0","parent":"module:BanditRLProof.LeafLemmas","order":5206,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:334"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pseudoRegret_const_bestArm : pseudoRegret model (fun _ => model.bestArm) t = 0","missing":[],"search":"pseudoregret_const_bestarm banditrlproof.pseudoregret_const_bestarm @[simp] theorem pseudoregret_const_bestarm : pseudoregret model (fun _ => model.bestarm) t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_const_of_gap_zero","label":"pseudoRegret_const_of_gap_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_const_of_gap_zero","description":"theorem pseudoRegret_const_of_gap_zero (arm : Fin K) (h : model.gap arm = 0) : pseudoRegret model (fun _ => arm) t = 0","url":"../modules/banditrlproof-leaflemmas/index.html#decl-333612519dd0","parent":"module:BanditRLProof.LeafLemmas","order":5207,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:340"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_const_of_gap_zero (arm : Fin K) (h : model.gap arm = 0) : pseudoRegret model (fun _ => arm) t = 0","missing":[],"search":"pseudoregret_const_of_gap_zero banditrlproof.pseudoregret_const_of_gap_zero theorem pseudoregret_const_of_gap_zero (arm : fin k) (h : model.gap arm = 0) : pseudoregret model (fun _ => arm) t = 0 theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_add_eq_of_forall_bestArm_between","label":"pseudoRegret_add_eq_of_forall_bestArm_between","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_add_eq_of_forall_bestArm_between","description":"theorem pseudoRegret_add_eq_of_forall_bestArm_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s = model.bestArm) : pseudoRegret model action (t + n) = pseudoRegret model action t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-0b4d7a856986","parent":"module:BanditRLProof.LeafLemmas","order":5208,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:346"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_add_eq_of_forall_bestArm_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → action s = model.bestArm) : pseudoRegret model action (t + n) = pseudoRegret model action t","missing":[],"search":"pseudoregret_add_eq_of_forall_bestarm_between banditrlproof.pseudoregret_add_eq_of_forall_bestarm_between theorem pseudoregret_add_eq_of_forall_bestarm_between (n : nat) (h : ∀ s, t ≤ s → s < t + n → action s = model.bestarm) : pseudoregret model action (t + n) = pseudoregret model action t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_add_eq_of_forall_gap_zero_between","label":"pseudoRegret_add_eq_of_forall_gap_zero_between","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_add_eq_of_forall_gap_zero_between","description":"theorem pseudoRegret_add_eq_of_forall_gap_zero_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → model.gap (action s) = 0) : pseudoRegret model action (t + n) = pseudoRegret model action t","url":"../modules/banditrlproof-leaflemmas/index.html#decl-f1ae889461ed","parent":"module:BanditRLProof.LeafLemmas","order":5209,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:358"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_add_eq_of_forall_gap_zero_between (n : Nat) (h : ∀ s, t ≤ s → s < t + n → model.gap (action s) = 0) : pseudoRegret model action (t + n) = pseudoRegret model action t","missing":[],"search":"pseudoregret_add_eq_of_forall_gap_zero_between banditrlproof.pseudoregret_add_eq_of_forall_gap_zero_between theorem pseudoregret_add_eq_of_forall_gap_zero_between (n : nat) (h : ∀ s, t ≤ s → s < t + n → model.gap (action s) = 0) : pseudoregret model action (t + n) = pseudoregret model action t theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_eq_list_range_foldl","label":"pseudoRegret_eq_list_range_foldl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_eq_list_range_foldl","description":"The recursive pseudo-regret equals a left fold over the half-open time prefix. This is the dependency-light `List.range` bridge for the Rat-valued regret accumulator. It matches the recursive bracketing directly.","url":"../modules/banditrlproof-leaflemmas/index.html#decl-59956f26db5b","parent":"module:BanditRLProof.LeafLemmas","order":5210,"meta":[["Kind","theorem"],["Module","BanditRLProof.LeafLemmas"],["Source","BanditRLProof/LeafLemmas.lean:376"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_eq_list_range_foldl : pseudoRegret model action t = (List.range t).foldl (fun acc s => acc + model.gap (action s)) 0","missing":[],"search":"pseudoregret_eq_list_range_foldl banditrlproof.pseudoregret_eq_list_range_foldl the recursive pseudo-regret equals a left fold over the half-open time prefix. this is the dependency-light `list.range` bridge for the rat-valued regret accumulator. it matches the recursive bracketing directly. theorem compiled","shard":"modules/5ecd1e8ad608dabc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.UpstreamRef","label":"UpstreamRef","kind":"structure","status":"compiled","subtitle":"BanditRLProof.UpstreamRef","description":"Public upstream source used by the memory layer.","url":"../modules/banditrlproof-literature/index.html#decl-d05455d3cf1b","parent":"module:BanditRLProof.Literature","order":5211,"meta":[["Kind","structure"],["Module","BanditRLProof.Literature"],["Source","BanditRLProof/Literature.lean:12"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure UpstreamRef where","missing":[],"search":"upstreamref banditrlproof.upstreamref public upstream source used by the memory layer. structure compiled","shard":"modules/1da684ee192cb448.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.lmlRef","label":"lmlRef","kind":"definition","status":"compiled","subtitle":"BanditRLProof.lmlRef","description":"Main upstream Lean library for bandit formalization.","url":"../modules/banditrlproof-literature/index.html#decl-bab031418e36","parent":"module:BanditRLProof.Literature","order":5212,"meta":[["Kind","definition"],["Module","BanditRLProof.Literature"],["Source","BanditRLProof/Literature.lean:22"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def lmlRef : UpstreamRef where","missing":[],"search":"lmlref banditrlproof.lmlref main upstream lean library for bandit formalization. definition compiled","shard":"modules/1da684ee192cb448.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.lmlBanditDeclarationCards","label":"lmlBanditDeclarationCards","kind":"definition","status":"compiled","subtitle":"BanditRLProof.lmlBanditDeclarationCards","description":"Selected LML declarations that seed the retrieval memory.","url":"../modules/banditrlproof-literature/index.html#decl-c3c0ac180e7c","parent":"module:BanditRLProof.Literature","order":5213,"meta":[["Kind","definition"],["Module","BanditRLProof.Literature"],["Source","BanditRLProof/Literature.lean:31"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def lmlBanditDeclarationCards : List RegretBoundCard","missing":[],"search":"lmlbanditdeclarationcards banditrlproof.lmlbanditdeclarationcards selected lml declarations that seed the retrieval memory. definition compiled","shard":"modules/1da684ee192cb448.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.mul_exp_neg_half_log_eq_sqrt","label":"mul_exp_neg_half_log_eq_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.mul_exp_neg_half_log_eq_sqrt","description":"theorem mul_exp_neg_half_log_eq_sqrt {r : ℝ} (hr : 0 ≤ r) : r * Real.exp (-Real.log r / 2) = Real.sqrt r","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-57eefe94b80f","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5214,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mul_exp_neg_half_log_eq_sqrt {r : ℝ} (hr : 0 ≤ r) : r * Real.exp (-Real.log r / 2) = Real.sqrt r","missing":[],"search":"mul_exp_neg_half_log_eq_sqrt banditrlproof.lowerbounds.mul_exp_neg_half_log_eq_sqrt theorem mul_exp_neg_half_log_eq_sqrt {r : ℝ} (hr : 0 ≤ r) : r * real.exp (-real.log r / 2) = real.sqrt r theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_sqrt_rnDeriv","label":"integrable_sqrt_rnDeriv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_sqrt_rnDeriv","description":"theorem integrable_sqrt_rnDeriv {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : Integrable (fun x => Real.sqrt (P.rnDeriv Q x).toReal) Q","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-236c952609d2","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5215,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_sqrt_rnDeriv {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : Integrable (fun x => Real.sqrt (P.rnDeriv Q x).toReal) Q","missing":[],"search":"integrable_sqrt_rnderiv banditrlproof.lowerbounds.integrable_sqrt_rnderiv theorem integrable_sqrt_rnderiv {α : type*} [measurablespace α] (p q : measure α) [isfinitemeasure p] [isfinitemeasure q] : integrable (fun x => real.sqrt (p.rnderiv q x).toreal) q theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_exp_neg_half_llr","label":"integrable_exp_neg_half_llr","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_exp_neg_half_llr","description":"theorem integrable_exp_neg_half_llr {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (hPQ : P ≪ Q) : Integrable (fun x => Real.exp (-llr P Q x / 2)) P","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-52d8a7a02dae","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5216,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_neg_half_llr {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (hPQ : P ≪ Q) : Integrable (fun x => Real.exp (-llr P Q x / 2)) P","missing":[],"search":"integrable_exp_neg_half_llr banditrlproof.lowerbounds.integrable_exp_neg_half_llr theorem integrable_exp_neg_half_llr {α : type*} [measurablespace α] (p q : measure α) [isfinitemeasure p] [isfinitemeasure q] (hpq : p ≪ q) : integrable (fun x => real.exp (-llr p q x / 2)) p theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_exp_neg_half_llr_eq","label":"integral_exp_neg_half_llr_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_exp_neg_half_llr_eq","description":"theorem integral_exp_neg_half_llr_eq {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (hPQ : P ≪ Q) : (∫ x, Real.exp (-llr P Q x / 2) ∂P) = ∫ x, Real.sqrt (P.rnDeriv Q x).toReal ∂Q","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-92889e29dd1c","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5217,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:40"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_neg_half_llr_eq {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (hPQ : P ≪ Q) : (∫ x, Real.exp (-llr P Q x / 2) ∂P) = ∫ x, Real.sqrt (P.rnDeriv Q x).toReal ∂Q","missing":[],"search":"integral_exp_neg_half_llr_eq banditrlproof.lowerbounds.integral_exp_neg_half_llr_eq theorem integral_exp_neg_half_llr_eq {α : type*} [measurablespace α] (p q : measure α) [isfinitemeasure p] [isfinitemeasure q] (hpq : p ≪ q) : (∫ x, real.exp (-llr p q x / 2) ∂p) = ∫ x, real.sqrt (p.rnderiv q x).toreal ∂q theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exp_neg_half_integral_llr_le_rnAffinity","label":"exp_neg_half_integral_llr_le_rnAffinity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exp_neg_half_integral_llr_le_rnAffinity","description":"Jensen's source proof step, in the RN representation and finite-KL branch.","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-f47e895acf47","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5218,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_half_integral_llr_le_rnAffinity {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] (hPQ : P ≪ Q) (hi : Integrable (llr P Q) P) : Real.exp (-(∫ x, llr P Q x ∂P) / 2) ≤ ∫ x, Real.sqrt (P.rnDeriv Q x).toReal ∂Q","missing":[],"search":"exp_neg_half_integral_llr_le_rnaffinity banditrlproof.lowerbounds.exp_neg_half_integral_llr_le_rnaffinity jensen's source proof step, in the rn representation and finite-kl branch. theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.rnAffinity_eq_commonDensityAffinity","label":"rnAffinity_eq_commonDensityAffinity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.rnAffinity_eq_commonDensityAffinity","description":"Change the RN affinity to any common sigma-finite dominating measure.","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-f645fc567b62","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5219,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rnAffinity_eq_commonDensityAffinity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] [SigmaFinite μ] (hPQ : P ≪ Q) (hQ : Q ≪ μ) : (∫ x, Real.sqrt (P.rnDeriv Q x).toReal ∂Q) = commonDensityAffinity P Q μ","missing":[],"search":"rnaffinity_eq_commondensityaffinity banditrlproof.lowerbounds.rnaffinity_eq_commondensityaffinity change the rn affinity to any common sigma-finite dominating measure. theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exp_neg_half_integral_llr_le_commonDensityAffinity","label":"exp_neg_half_integral_llr_le_commonDensityAffinity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exp_neg_half_integral_llr_le_commonDensityAffinity","description":"The source's measure-level Jensen step with arbitrary common domination.","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-d15f9c310ec3","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5220,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_half_integral_llr_le_commonDensityAffinity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] [SigmaFinite μ] (hPQ : P ≪ Q) (hQ : Q ≪ μ) (hi : Integrable (llr P Q) P) : Real.exp (-(∫ x, llr P Q x ∂P) / 2) ≤ commonDensityAffinity P Q μ","missing":[],"search":"exp_neg_half_integral_llr_le_commondensityaffinity banditrlproof.lowerbounds.exp_neg_half_integral_llr_le_commondensityaffinity the source's measure-level jensen step with arbitrary common domination. theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_le_half_commonDensityAffinity_sq","label":"bretagnolleHuberScale_le_half_commonDensityAffinity_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale_le_half_commonDensityAffinity_sq","description":"The squared-affinity/KL bound, with the infinite-KL branch explicit.","url":"../modules/banditrlproof-lowerbounds-affinitykl/index.html#decl-c98aaffbd8de","parent":"module:BanditRLProof.LowerBounds.AffinityKL","order":5221,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.AffinityKL"],["Source","BanditRLProof/LowerBounds/AffinityKL.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuberScale_le_half_commonDensityAffinity_sq {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] [SigmaFinite μ] (hQ : Q ≪ μ) : bretagnolleHuberScale (relativeEntropy P Q) ≤ (1 / 2 : ℝ) * commonDensityAffinity P Q μ ^ 2","missing":[],"search":"bretagnollehuberscale_le_half_commondensityaffinity_sq banditrlproof.lowerbounds.bretagnollehuberscale_le_half_commondensityaffinity_sq the squared-affinity/kl bound, with the infinite-kl branch explicit. theorem compiled","shard":"modules/00f13da6645c930b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlockList","label":"sourceBlockList","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlockList","description":"def sourceBlockList {α : Type*} : (n : ℕ) → SourceBlock α n → List α | 0, _ => [] | n + 1, x => x.1 :: sourceBlockList n x.2 theorem sourceBlockList_length {α : Type*} (n : ℕ) (x : SourceBlock α n) : (sourceBlockList n x).length = n","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-d4d6dad7e778","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5222,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def sourceBlockList {α : Type*} : (n : ℕ) → SourceBlock α n → List α | 0, _ => [] | n + 1, x => x.1 :: sourceBlockList n x.2 theorem sourceBlockList_length {α : Type*} (n : ℕ) (x : SourceBlock α n) : (sourceBlockList n x).length = n","missing":[],"search":"sourceblocklist banditrlproof.lowerbounds.sourceblocklist def sourceblocklist {α : type*} : (n : ℕ) → sourceblock α n → list α | 0, _ => [] | n + 1, x => x.1 :: sourceblocklist n x.2 theorem sourceblocklist_length {α : type*} (n : ℕ) (x : sourceblock α n) : (sourceblocklist n x).length = n definition compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlockList_length","label":"sourceBlockList_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlockList_length","description":"theorem sourceBlockList_length {α : Type*} (n : ℕ) (x : SourceBlock α n) : (sourceBlockList n x).length = n","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-e66569973947","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5223,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceBlockList_length {α : Type*} (n : ℕ) (x : SourceBlock α n) : (sourceBlockList n x).length = n","missing":[],"search":"sourceblocklist_length banditrlproof.lowerbounds.sourceblocklist_length theorem sourceblocklist_length {α : type*} (n : ℕ) (x : sourceblock α n) : (sourceblocklist n x).length = n theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlockList_injective","label":"sourceBlockList_injective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlockList_injective","description":"theorem sourceBlockList_injective {α : Type*} (n : ℕ) : Function.Injective (sourceBlockList (α := α) n)","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-1865dd67e138","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5224,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceBlockList_injective {α : Type*} (n : ℕ) : Function.Injective (sourceBlockList (α := α) n)","missing":[],"search":"sourceblocklist_injective banditrlproof.lowerbounds.sourceblocklist_injective theorem sourceblocklist_injective {α : type*} (n : ℕ) : function.injective (sourceblocklist (α := α) n) theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlockList_mass","label":"sourceBlockList_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlockList_mass","description":"theorem sourceBlockList_mass {α : Type*} (p : α → ℝ) (n : ℕ) (x : SourceBlock α n) : ((sourceBlockList n x).map p).prod = sourceBlockMass p n x","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-89d4dd7ae32c","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5225,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceBlockList_mass {α : Type*} (p : α → ℝ) (n : ℕ) (x : SourceBlock α n) : ((sourceBlockList n x).map p).prod = sourceBlockMass p n x","missing":[],"search":"sourceblocklist_mass banditrlproof.lowerbounds.sourceblocklist_mass theorem sourceblocklist_mass {α : type*} (p : α → ℝ) (n : ℕ) (x : sourceblock α n) : ((sourceblocklist n x).map p).prod = sourceblockmass p n x theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_arithmeticBlockSupport","label":"exists_arithmeticBlockSupport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_arithmeticBlockSupport","description":"theorem exists_arithmeticBlockSupport {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : ∃ positive : BinaryPrefixCode {x : SourceBlock.{0,0} (Fin k) n // 0 < sourceBlockMass p n x}, (∀ x, (positive.encode x).length = arithmeticLength (sourceBlockMass p n x.val)) ∧ (∀ x, (arithmeticInterval p (sourceBlockList n x.val)).1 ≤ dyadicAddressLower (positive.encode x) ∧ dyadicAddressUpper (positive.e…","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-1e04f6157f0a","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5226,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_arithmeticBlockSupport {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : ∃ positive : BinaryPrefixCode {x : SourceBlock.{0,0} (Fin k) n // 0 < sourceBlockMass p n x}, (∀ x, (positive.encode x).length = arithmeticLength (sourceBlockMass p n x.val)) ∧ (∀ x, (arithmeticInterval p (sourceBlockList n x.val)).1 ≤ dyadicAddressLower (positive.encode x) ∧ dyadicAddressUpper (positive.encode x) < (arithmeticInterval p (sourceBlockList n x.val)).2) ∧ expectedCodeLength (sourceBlockMass p n) (positive.extendZeroMass (sourceBlockMass p n) (huffmanCode (sourceBlockMass p n) (sourceBlockMass_nonneg p hp n))) ≤ n * discreteEntropyBaseTwo Finset.univ p + 3","missing":[],"search":"exists_arithmeticblocksupport banditrlproof.lowerbounds.exists_arithmeticblocksupport theorem exists_arithmeticblocksupport {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : ∃ positive : binaryprefixcode {x : sourceblock.{0,0} (fin k) n // 0 < sourceblockmass p n x}, (∀ x, (positive.encode x).length = arithmeticlength (sourceblockmass p n x.val)) ∧ (∀ x, (arithmeticinterval p (sourceblocklist n x.val)).1 ≤ dyadicaddresslower (positive.encode x) ∧ dyadicaddressupper (positive.encode x) < (arithmeticinterval p (sourceblocklist n x.val)).2) ∧ expectedcodelength (sourceblockmass p n) (positive.extendzeromass (sourceblockmass p n) (huffmancode (sourceblockmass p n) (sourceblockmass_nonneg p hp n))) ≤ n * discreteentropybasetwo finset.univ p + 3 theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode","label":"arithmeticBlockCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticBlockCode","description":"The named arithmetic block code: interval addresses on positive blocks, with the one-bit tagged fallback on zero-mass blocks.","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-f523f2074cc3","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5227,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def arithmeticBlockCode {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : BinaryPrefixCode (SourceBlock.{0,0} (Fin k) n)","missing":[],"search":"arithmeticblockcode banditrlproof.lowerbounds.arithmeticblockcode the named arithmetic block code: interval addresses on positive blocks, with the one-bit tagged fallback on zero-mass blocks. definition compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_expected_length_le","label":"arithmeticBlockCode_expected_length_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticBlockCode_expected_length_le","description":"theorem arithmeticBlockCode_expected_length_le {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : expectedCodeLength (sourceBlockMass p n) (arithmeticBlockCode p hp hs n) ≤ n * discreteEntropyBaseTwo Finset.univ p + 3","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-ae0256130f53","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5228,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticBlockCode_expected_length_le {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : expectedCodeLength (sourceBlockMass p n) (arithmeticBlockCode p hp hs n) ≤ n * discreteEntropyBaseTwo Finset.univ p + 3","missing":[],"search":"arithmeticblockcode_expected_length_le banditrlproof.lowerbounds.arithmeticblockcode_expected_length_le theorem arithmeticblockcode_expected_length_le {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) : expectedcodelength (sourceblockmass p n) (arithmeticblockcode p hp hs n) ≤ n * discreteentropybasetwo finset.univ p + 3 theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_payload_interval","label":"arithmeticBlockCode_payload_interval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticBlockCode_payload_interval","description":"Removing the support tag from a positive block yields an address inside its actual arithmetic interval. This property belongs to the named code, not merely to an unrelated witness with the same length.","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-27c63fe10217","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5229,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem arithmeticBlockCode_payload_interval {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) (x : SourceBlock.{0,0} (Fin k) n) (hx : 0 < sourceBlockMass p n x) : (arithmeticInterval p (sourceBlockList n x)).1 ≤ dyadicAddressLower ((arithmeticBlockCode p hp hs n).encode x).tail ∧ dyadicAddressUpper ((arithmeticBlockCode p hp hs n).encode x).tail < (arithmeticInterval p (sourceBlockList n x)).2","missing":[],"search":"arithmeticblockcode_payload_interval banditrlproof.lowerbounds.arithmeticblockcode_payload_interval removing the support tag from a positive block yields an address inside its actual arithmetic interval. this property belongs to the named code, not merely to an unrelated witness with the same length. theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_rate_sandwich","label":"arithmeticBlockCode_rate_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticBlockCode_rate_sandwich","description":"theorem arithmeticBlockCode_rate_sandwich {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) (hn : 0 < n) : discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength (sourceBlockMass p n) (arithmeticBlockCode p hp hs n) / n ∧ expectedCodeLength (sourceBlockMass p n) (arithmeticBlockCode p hp hs n) / n ≤ discreteEntropyBaseTwo Finset.univ p + 3 / n","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-110fc670f1eb","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5230,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticBlockCode_rate_sandwich {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) (hn : 0 < n) : discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength (sourceBlockMass p n) (arithmeticBlockCode p hp hs n) / n ∧ expectedCodeLength (sourceBlockMass p n) (arithmeticBlockCode p hp hs n) / n ≤ discreteEntropyBaseTwo Finset.univ p + 3 / n","missing":[],"search":"arithmeticblockcode_rate_sandwich banditrlproof.lowerbounds.arithmeticblockcode_rate_sandwich theorem arithmeticblockcode_rate_sandwich {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) (hn : 0 < n) : discreteentropybasetwo finset.univ p ≤ expectedcodelength (sourceblockmass p n) (arithmeticblockcode p hp hs n) / n ∧ expectedcodelength (sourceblockmass p n) (arithmeticblockcode p hp hs n) / n ≤ discreteentropybasetwo finset.univ p + 3 / n theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_rate_tendsto_entropy","label":"arithmeticBlockCode_rate_tendsto_entropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticBlockCode_rate_tendsto_entropy","description":"The constructed arithmetic code's expected bits per symbol tend to entropy, including sources with zero-probability symbols.","url":"../modules/banditrlproof-lowerbounds-arithmeticblockcoding/index.html#decl-9fdfa11ea805","parent":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","order":5231,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticBlockCoding"],["Source","BanditRLProof/LowerBounds/ArithmeticBlockCoding.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem arithmeticBlockCode_rate_tendsto_entropy {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : Filter.Tendsto (fun n : ℕ => expectedCodeLength (sourceBlockMass p (n + 1)) (arithmeticBlockCode p hp hs (n + 1)) / (n + 1)) Filter.atTop (nhds (discreteEntropyBaseTwo Finset.univ p))","missing":[],"search":"arithmeticblockcode_rate_tendsto_entropy banditrlproof.lowerbounds.arithmeticblockcode_rate_tendsto_entropy the constructed arithmetic code's expected bits per symbol tend to entropy, including sources with zero-probability symbols. theorem compiled","shard":"modules/c863a1e7129c5a32.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticOffset","label":"arithmeticOffset","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticOffset","description":"Left endpoint of a symbol's arithmetic-coding partition cell.","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-fc3e0aa4a46a","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5232,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def arithmeticOffset {k : ℕ} (p : Fin k → ℝ) (a : Fin k) : ℝ","missing":[],"search":"arithmeticoffset banditrlproof.lowerbounds.arithmeticoffset left endpoint of a symbol's arithmetic-coding partition cell. definition compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticOffset_nonneg","label":"arithmeticOffset_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticOffset_nonneg","description":"theorem arithmeticOffset_nonneg {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (a : Fin k) : 0 ≤ arithmeticOffset p a","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-ffe8aac09d8d","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5233,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticOffset_nonneg {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (a : Fin k) : 0 ≤ arithmeticOffset p a","missing":[],"search":"arithmeticoffset_nonneg banditrlproof.lowerbounds.arithmeticoffset_nonneg theorem arithmeticoffset_nonneg {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (a : fin k) : 0 ≤ arithmeticoffset p a theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticOffset_add_le_one","label":"arithmeticOffset_add_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticOffset_add_le_one","description":"theorem arithmeticOffset_add_le_one {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (a : Fin k) : arithmeticOffset p a + p a ≤ 1","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-2b05a4e3b0c7","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5234,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticOffset_add_le_one {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (a : Fin k) : arithmeticOffset p a + p a ≤ 1","missing":[],"search":"arithmeticoffset_add_le_one banditrlproof.lowerbounds.arithmeticoffset_add_le_one theorem arithmeticoffset_add_le_one {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (a : fin k) : arithmeticoffset p a + p a ≤ 1 theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticInterval","label":"arithmeticInterval","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticInterval","description":"Successive affine subdivisions for an arithmetic-coded word.","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-256218f3edfa","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5235,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def arithmeticInterval {k : ℕ} (p : Fin k → ℝ) : List (Fin k) → ℝ × ℝ | [] => (0, 1) | a :: w => (arithmeticOffset p a + p a * (arithmeticInterval p w).1, arithmeticOffset p a + p a * (arithmeticInterval p w).2) theorem arithmeticInterval_width {k : ℕ} (p : Fin k → ℝ) (w : List (Fin k)) : (arithmeticInterval p w).2 - (arithmeticInterval p w).1 = (w.map p).prod","missing":[],"search":"arithmeticinterval banditrlproof.lowerbounds.arithmeticinterval successive affine subdivisions for an arithmetic-coded word. definition compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_width","label":"arithmeticInterval_width","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticInterval_width","description":"theorem arithmeticInterval_width {k : ℕ} (p : Fin k → ℝ) (w : List (Fin k)) : (arithmeticInterval p w).2 - (arithmeticInterval p w).1 = (w.map p).prod","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-eb276ae66440","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5236,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticInterval_width {k : ℕ} (p : Fin k → ℝ) (w : List (Fin k)) : (arithmeticInterval p w).2 - (arithmeticInterval p w).1 = (w.map p).prod","missing":[],"search":"arithmeticinterval_width banditrlproof.lowerbounds.arithmeticinterval_width theorem arithmeticinterval_width {k : ℕ} (p : fin k → ℝ) (w : list (fin k)) : (arithmeticinterval p w).2 - (arithmeticinterval p w).1 = (w.map p).prod theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_bounds","label":"arithmeticInterval_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticInterval_bounds","description":"theorem arithmeticInterval_bounds {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (w : List (Fin k)) : 0 ≤ (arithmeticInterval p w).1 ∧ (arithmeticInterval p w).1 ≤ (arithmeticInterval p w).2 ∧ (arithmeticInterval p w).2 ≤ 1","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-d052b92dc29b","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5237,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticInterval_bounds {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (w : List (Fin k)) : 0 ≤ (arithmeticInterval p w).1 ∧ (arithmeticInterval p w).1 ≤ (arithmeticInterval p w).2 ∧ (arithmeticInterval p w).2 ≤ 1","missing":[],"search":"arithmeticinterval_bounds banditrlproof.lowerbounds.arithmeticinterval_bounds theorem arithmeticinterval_bounds {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (w : list (fin k)) : 0 ≤ (arithmeticinterval p w).1 ∧ (arithmeticinterval p w).1 ≤ (arithmeticinterval p w).2 ∧ (arithmeticinterval p w).2 ≤ 1 theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticOffset_separated","label":"arithmeticOffset_separated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticOffset_separated","description":"theorem arithmeticOffset_separated {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (a b : Fin k) (hab : a < b) : arithmeticOffset p a + p a ≤ arithmeticOffset p b","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-48538899fac0","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5238,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticOffset_separated {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (a b : Fin k) (hab : a < b) : arithmeticOffset p a + p a ≤ arithmeticOffset p b","missing":[],"search":"arithmeticoffset_separated banditrlproof.lowerbounds.arithmeticoffset_separated theorem arithmeticoffset_separated {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (a b : fin k) (hab : a < b) : arithmeticoffset p a + p a ≤ arithmeticoffset p b theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_head_separated","label":"arithmeticInterval_head_separated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticInterval_head_separated","description":"theorem arithmeticInterval_head_separated {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (a b : Fin k) (u v : List (Fin k)) (hab : a < b) : (arithmeticInterval p (a :: u)).2 ≤ (arithmeticInterval p (b :: v)).1","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-9ad5d094c2cf","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5239,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticInterval_head_separated {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (a b : Fin k) (u v : List (Fin k)) (hab : a < b) : (arithmeticInterval p (a :: u)).2 ≤ (arithmeticInterval p (b :: v)).1","missing":[],"search":"arithmeticinterval_head_separated banditrlproof.lowerbounds.arithmeticinterval_head_separated theorem arithmeticinterval_head_separated {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (a b : fin k) (u v : list (fin k)) (hab : a < b) : (arithmeticinterval p (a :: u)).2 ≤ (arithmeticinterval p (b :: v)).1 theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_separated","label":"arithmeticInterval_separated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticInterval_separated","description":"Different equal-length messages occupy cells with disjoint interiors.","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-b16818c82086","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5240,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticInterval_separated {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (u v : List (Fin k)) (hlen : u.length = v.length) (hne : u ≠ v) : (arithmeticInterval p u).2 ≤ (arithmeticInterval p v).1 ∨ (arithmeticInterval p v).2 ≤ (arithmeticInterval p u).1","missing":[],"search":"arithmeticinterval_separated banditrlproof.lowerbounds.arithmeticinterval_separated different equal-length messages occupy cells with disjoint interiors. theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_interior_unique","label":"arithmeticInterval_interior_unique","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticInterval_interior_unique","description":"Any point strictly inside a message cell identifies that message uniquely among messages of the same length.","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-3418930d249e","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5241,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:114"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticInterval_interior_unique {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (u v : List (Fin k)) (hlen : u.length = v.length) (x : ℝ) (hu : (arithmeticInterval p u).1 < x ∧ x < (arithmeticInterval p u).2) (hv : (arithmeticInterval p v).1 < x ∧ x < (arithmeticInterval p v).2) : u = v","missing":[],"search":"arithmeticinterval_interior_unique banditrlproof.lowerbounds.arithmeticinterval_interior_unique any point strictly inside a message cell identifies that message uniquely among messages of the same length. theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_grid_cell_inside","label":"exists_grid_cell_inside","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_grid_cell_inside","description":"A positive grid cell fits in any nonnegative interval at least twice as wide. Mathlib-candidate scalar rounding leaf for arithmetic-code dyadic selection.","url":"../modules/banditrlproof-lowerbounds-arithmeticintervals/index.html#decl-072260cfd61d","parent":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","order":5242,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticIntervals"],["Source","BanditRLProof/LowerBounds/ArithmeticIntervals.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_grid_cell_inside (L U δ : ℝ) (hL : 0 ≤ L) (hδ : 0 < δ) (hwidth : 2 * δ ≤ U - L) : ∃ m : ℕ, L ≤ (m : ℝ) * δ ∧ ((m : ℝ) + 1) * δ < U","missing":[],"search":"exists_grid_cell_inside banditrlproof.lowerbounds.exists_grid_cell_inside a positive grid cell fits in any nonnegative interval at least twice as wide. mathlib-candidate scalar rounding leaf for arithmetic-code dyadic selection. theorem compiled","shard":"modules/baa92f181aa335e8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticAddress_prefix_forces_eq","label":"arithmeticAddress_prefix_forces_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticAddress_prefix_forces_eq","description":"theorem arithmeticAddress_prefix_forces_eq {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (u v : List (Fin k)) (hlen : u.length = v.length) (cu cv : List Bool) (huL : (arithmeticInterval p u).1 ≤ dyadicAddressLower cu) (huU : dyadicAddressUpper cu < (arithmeticInterval p u).2) (hvL : (arithmeticInterval p v).1 ≤ dyadicAddressLower cv) (hvU : dyadicAddressUpper cv < (arithmeticInterval p v).2) (hpref…","url":"../modules/banditrlproof-lowerbounds-arithmeticprefixcode/index.html#decl-c15dcf21a823","parent":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","order":5243,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticPrefixCode"],["Source","BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticAddress_prefix_forces_eq {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (u v : List (Fin k)) (hlen : u.length = v.length) (cu cv : List Bool) (huL : (arithmeticInterval p u).1 ≤ dyadicAddressLower cu) (huU : dyadicAddressUpper cu < (arithmeticInterval p u).2) (hvL : (arithmeticInterval p v).1 ≤ dyadicAddressLower cv) (hvU : dyadicAddressUpper cv < (arithmeticInterval p v).2) (hprefix : cu <+: cv) : u = v","missing":[],"search":"arithmeticaddress_prefix_forces_eq banditrlproof.lowerbounds.arithmeticaddress_prefix_forces_eq theorem arithmeticaddress_prefix_forces_eq {k : ℕ} (p : fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (u v : list (fin k)) (hlen : u.length = v.length) (cu cv : list bool) (hul : (arithmeticinterval p u).1 ≤ dyadicaddresslower cu) (huu : dyadicaddressupper cu < (arithmeticinterval p u).2) (hvl : (arithmeticinterval p v).1 ≤ dyadicaddresslower cv) (hvu : dyadicaddressupper cv < (arithmeticinterval p v).2) (hprefix : cu <+: cv) : u = v theorem compiled","shard":"modules/529bb9fb74a6dab5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_arithmeticPrefixCode","label":"exists_arithmeticPrefixCode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_arithmeticPrefixCode","description":"Assemble arithmetic-cell addresses into a genuine prefix code, preserving the supplied bit lengths. The width budgets are discharged by later allocation.","url":"../modules/banditrlproof-lowerbounds-arithmeticprefixcode/index.html#decl-ae53816b75dd","parent":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","order":5244,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticPrefixCode"],["Source","BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_arithmeticPrefixCode {α : Type*} {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (message : α → List (Fin k)) (hinj : Function.Injective message) (hlen : ∀ a b, (message a).length = (message b).length) (bits : α → ℕ) (hbudget : ∀ a, 2 * (1 / (2 : ℝ) ^ bits a) ≤ (arithmeticInterval p (message a)).2 - (arithmeticInterval p (message a)).1) : ∃ code : BinaryPrefixCode α, (∀ a, (code.encode a).length = bits a) ∧ ∀ a, (arithmeticInterval p (message a)).1 ≤ dyadicAddressLower (code.encode a) ∧ dyadicAddressUpper (code.encode a) < (arithmeticInterval p (message a)).2","missing":[],"search":"exists_arithmeticprefixcode banditrlproof.lowerbounds.exists_arithmeticprefixcode assemble arithmetic-cell addresses into a genuine prefix code, preserving the supplied bit lengths. the width budgets are discharged by later allocation. theorem compiled","shard":"modules/529bb9fb74a6dab5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticLength","label":"arithmeticLength","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticLength","description":"One extra bit beyond strict Shannon length fits a dyadic cell inside the arithmetic interval, rather than merely meeting a Kraft budget.","url":"../modules/banditrlproof-lowerbounds-arithmeticprefixcode/index.html#decl-4b24d66b372d","parent":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","order":5245,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticPrefixCode"],["Source","BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def arithmeticLength (mass : ℝ) : ℕ","missing":[],"search":"arithmeticlength banditrlproof.lowerbounds.arithmeticlength one extra bit beyond strict shannon length fits a dyadic cell inside the arithmetic interval, rather than merely meeting a kraft budget. definition compiled","shard":"modules/529bb9fb74a6dab5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticLength_width_budget","label":"arithmeticLength_width_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticLength_width_budget","description":"theorem arithmeticLength_width_budget {mass : ℝ} (hm : 0 < mass) : 2 * (1 / (2 : ℝ) ^ arithmeticLength mass) < mass","url":"../modules/banditrlproof-lowerbounds-arithmeticprefixcode/index.html#decl-96ee40ffe022","parent":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","order":5246,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticPrefixCode"],["Source","BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticLength_width_budget {mass : ℝ} (hm : 0 < mass) : 2 * (1 / (2 : ℝ) ^ arithmeticLength mass) < mass","missing":[],"search":"arithmeticlength_width_budget banditrlproof.lowerbounds.arithmeticlength_width_budget theorem arithmeticlength_width_budget {mass : ℝ} (hm : 0 < mass) : 2 * (1 / (2 : ℝ) ^ arithmeticlength mass) < mass theorem compiled","shard":"modules/529bb9fb74a6dab5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.arithmeticLength_le_information_add_two","label":"arithmeticLength_le_information_add_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.arithmeticLength_le_information_add_two","description":"theorem arithmeticLength_le_information_add_two {mass : ℝ} (hm : 0 < mass) (hm1 : mass ≤ 1) : (arithmeticLength mass : ℝ) ≤ Real.log mass⁻¹ / Real.log 2 + 2","url":"../modules/banditrlproof-lowerbounds-arithmeticprefixcode/index.html#decl-cf3e5ec156fd","parent":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","order":5247,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticPrefixCode"],["Source","BanditRLProof/LowerBounds/ArithmeticPrefixCode.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem arithmeticLength_le_information_add_two {mass : ℝ} (hm : 0 < mass) (hm1 : mass ≤ 1) : (arithmeticLength mass : ℝ) ≤ Real.log mass⁻¹ / Real.log 2 + 2","missing":[],"search":"arithmeticlength_le_information_add_two banditrlproof.lowerbounds.arithmeticlength_le_information_add_two theorem arithmeticlength_le_information_add_two {mass : ℝ} (hm : 0 < mass) (hm1 : mass ≤ 1) : (arithmeticlength mass : ℝ) ≤ real.log mass⁻¹ / real.log 2 + 2 theorem compiled","shard":"modules/529bb9fb74a6dab5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.supportTaggedWord","label":"supportTaggedWord","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.supportTaggedWord","description":"noncomputable def supportTaggedWord {α : Type*} (p : α → ℝ) (positive : BinaryPrefixCode {a // 0 < p a}) (fallback : BinaryPrefixCode α) (a : α) : List Bool","url":"../modules/banditrlproof-lowerbounds-arithmeticzeroextension/index.html#decl-220ca08a2f43","parent":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","order":5248,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticZeroExtension"],["Source","BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def supportTaggedWord {α : Type*} (p : α → ℝ) (positive : BinaryPrefixCode {a // 0 < p a}) (fallback : BinaryPrefixCode α) (a : α) : List Bool","missing":[],"search":"supporttaggedword banditrlproof.lowerbounds.supporttaggedword noncomputable def supporttaggedword {α : type*} (p : α → ℝ) (positive : binaryprefixcode {a // 0 < p a}) (fallback : binaryprefixcode α) (a : α) : list bool definition compiled","shard":"modules/36b23875b3139ebf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.supportTaggedWord_prefixFree","label":"supportTaggedWord_prefixFree","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.supportTaggedWord_prefixFree","description":"theorem supportTaggedWord_prefixFree {α : Type*} (p : α → ℝ) (positive : BinaryPrefixCode {a // 0 < p a}) (fallback : BinaryPrefixCode α) {a b : α} (h : supportTaggedWord p positive fallback a <+: supportTaggedWord p positive fallback b) : a = b","url":"../modules/banditrlproof-lowerbounds-arithmeticzeroextension/index.html#decl-ec7956337cd1","parent":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","order":5249,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticZeroExtension"],["Source","BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem supportTaggedWord_prefixFree {α : Type*} (p : α → ℝ) (positive : BinaryPrefixCode {a // 0 < p a}) (fallback : BinaryPrefixCode α) {a b : α} (h : supportTaggedWord p positive fallback a <+: supportTaggedWord p positive fallback b) : a = b","missing":[],"search":"supporttaggedword_prefixfree banditrlproof.lowerbounds.supporttaggedword_prefixfree theorem supporttaggedword_prefixfree {α : type*} (p : α → ℝ) (positive : binaryprefixcode {a // 0 < p a}) (fallback : binaryprefixcode α) {a b : α} (h : supporttaggedword p positive fallback a <+: supporttaggedword p positive fallback b) : a = b theorem compiled","shard":"modules/36b23875b3139ebf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.extendZeroMass","label":"extendZeroMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.extendZeroMass","description":"Extend an arithmetic support code to all symbols, with a one-bit escape tag.","url":"../modules/banditrlproof-lowerbounds-arithmeticzeroextension/index.html#decl-c331b27a59ff","parent":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","order":5250,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ArithmeticZeroExtension"],["Source","BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def BinaryPrefixCode.extendZeroMass {α : Type*} (p : α → ℝ) (positive : BinaryPrefixCode {a // 0 < p a}) (fallback : BinaryPrefixCode α) : BinaryPrefixCode α where","missing":[],"search":"extendzeromass banditrlproof.lowerbounds.binaryprefixcode.extendzeromass extend an arithmetic support code to all symbols, with a one-bit escape tag. definition compiled","shard":"modules/36b23875b3139ebf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_extendZeroMass_le","label":"expectedCodeLength_extendZeroMass_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_extendZeroMass_le","description":"Zero-mass fallback words cost nothing in expectation. The extra support tag adds only one bit to a support code with information-plus-two length bound.","url":"../modules/banditrlproof-lowerbounds-arithmeticzeroextension/index.html#decl-7872be25632c","parent":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","order":5251,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticZeroExtension"],["Source","BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_extendZeroMass_le {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ a, 0 ≤ p a) (hs : ∑ a, p a = 1) (positive : BinaryPrefixCode {a // 0 < p a}) (fallback : BinaryPrefixCode α) (hpos : ∀ a (ha : 0 < p a), ((positive.encode ⟨a, ha⟩).length : ℝ) ≤ Real.log (p a)⁻¹ / Real.log 2 + 2) : expectedCodeLength p (positive.extendZeroMass p fallback) ≤ discreteEntropyBaseTwo Finset.univ p + 3","missing":[],"search":"expectedcodelength_extendzeromass_le banditrlproof.lowerbounds.expectedcodelength_extendzeromass_le zero-mass fallback words cost nothing in expectation. the extra support tag adds only one bit to a support code with information-plus-two length bound. theorem compiled","shard":"modules/36b23875b3139ebf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_zeroSafe_arithmeticCode","label":"exists_zeroSafe_arithmeticCode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_zeroSafe_arithmeticCode","description":"Arithmetic coding on the positive support, extended to every message. Only the zero-mass escape branch uses the supplied total Huffman fallback.","url":"../modules/banditrlproof-lowerbounds-arithmeticzeroextension/index.html#decl-369d064cdcdb","parent":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","order":5252,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ArithmeticZeroExtension"],["Source","BanditRLProof/LowerBounds/ArithmeticZeroExtension.lean:61"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_zeroSafe_arithmeticCode {α : Type*} [Fintype α] {k : ℕ} (p : Fin k → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (message : α → List (Fin k)) (hinj : Function.Injective message) (hlen : ∀ a b, (message a).length = (message b).length) (q : α → ℝ) (hq : ∀ a, 0 ≤ q a) (hqs : ∑ a, q a = 1) (hmass : ∀ a, ((message a).map p).prod = q a) : ∃ positive : BinaryPrefixCode {a // 0 < q a}, (∀ a, (positive.encode a).length = arithmeticLength (q a.val)) ∧ (∀ a, (arithmeticInterval p (message a.val)).1 ≤ dyadicAddressLower (positive.encode a) ∧ dyadicAddressUpper (positive.encode a) < (arithmeticInterval p (message a.val)).2) ∧ expectedCodeLength q (positive.extendZeroMass q (huffmanCode q hq)) ≤ discreteEntropyBaseTwo Finset.univ q + 3","missing":[],"search":"exists_zerosafe_arithmeticcode banditrlproof.lowerbounds.exists_zerosafe_arithmeticcode arithmetic coding on the positive support, extended to every message. only the zero-mass escape branch uses the supplied total huffman fallback. theorem compiled","shard":"modules/36b23875b3139ebf.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_map_le","label":"klDiv_map_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_map_le","description":"Kullback--Leibler divergence cannot increase under a measurable observation. The proof reuses Chapter 14's map/trim identity and sub-sigma-algebra KL contraction. The infinite-KL branch is explicit, so absolute continuity is derived only in the finite branch and is not a caller assumption.","url":"../modules/banditrlproof-lowerbounds-bandithistorydataprocessing/index.html#decl-0a7e108e6f47","parent":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","order":5253,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryDataProcessing"],["Source","BanditRLProof/LowerBounds/BanditHistoryDataProcessing.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem klDiv_map_le {Source Target : Type*} [MeasurableSpace Source] [MeasurableSpace Target] (mu nu : Measure Source) [IsFiniteMeasure mu] [IsFiniteMeasure nu] (observe : Source -> Target) (hobserve : Measurable observe) : InformationTheory.klDiv (mu.map observe) (nu.map observe) <= InformationTheory.klDiv mu nu","missing":[],"search":"kldiv_map_le banditrlproof.lowerbounds.kldiv_map_le kullback--leibler divergence cannot increase under a measurable observation. the proof reuses chapter 14's map/trim identity and sub-sigma-algebra kl contraction. the infinite-kl branch is explicit, so absolute continuity is derived only in the finite branch and is not a caller assumption. theorem compiled","shard":"modules/281da1b93079ff83.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_observedBanditHistory_le_expectedPulls_sum","label":"klDiv_observedBanditHistory_le_expectedPulls_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_observedBanditHistory_le_expectedPulls_sum","description":"Any measurable statistic of a deterministic finite bandit history has KL at most the first-law expected pull-count-weighted arm information. This is the data-processing half of Exercise 15.7 at a fixed horizon. It is not the stopping-time theorem because the right-hand side still counts all pulls through `lastRound`.","url":"../modules/banditrlproof-lowerbounds-bandithistorydataprocessing/index.html#decl-6dd7c19bc8ba","parent":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","order":5254,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryDataProcessing"],["Source","BanditRLProof/LowerBounds/BanditHistoryDataProcessing.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem klDiv_observedBanditHistory_le_expectedPulls_sum {K : Nat} {Reward : Type v} {Observation : Type w} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] [MeasurableSpace Observation] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (lastRound : Nat) (observe : History.FinitePairHistory (Fin K) Reward lastRound -> Observation) (hobserve : Measurable observe) : InformationTheory.klDiv ((canonicalBanditHistoryMeasure algorithm armLaw lastRound).map observe) ((canonicalBanditHistoryMeasure algorithm referenceArmLaw lastRound).map observe) <= ∑ arm : Fin K, canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm * InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm)","missing":[],"search":"kldiv_observedbandithistory_le_expectedpulls_sum banditrlproof.lowerbounds.kldiv_observedbandithistory_le_expectedpulls_sum any measurable statistic of a deterministic finite bandit history has kl at most the first-law expected pull-count-weighted arm information. this is the data-processing half of exercise 15.7 at a fixed horizon. it is not the stopping-time theorem because the right-hand side still counts all pulls through `lastround`. theorem compiled","shard":"modules/281da1b93079ff83.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.pairHistoryZeroMeasurableEquiv","label":"pairHistoryZeroMeasurableEquiv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.pairHistoryZeroMeasurableEquiv","description":"The time-zero action/reward pair is measurably equivalent to its singleton history.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-1e14131db0b2","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5255,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def pairHistoryZeroMeasurableEquiv (Action : Type u) (Reward : Type v) [MeasurableSpace Action] [MeasurableSpace Reward] : (Action × Reward) ≃ᵐ History.FinitePairHistory Action Reward 0 where","missing":[],"search":"pairhistoryzeromeasurableequiv banditrlproof.lowerbounds.pairhistoryzeromeasurableequiv the time-zero action/reward pair is measurably equivalent to its singleton history. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.pairHistorySuccMeasurableEquiv","label":"pairHistorySuccMeasurableEquiv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.pairHistorySuccMeasurableEquiv","description":"A finite history and one next pair are measurably equivalent to the successor history.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-348323349f46","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5256,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def pairHistorySuccMeasurableEquiv (Action : Type u) (Reward : Type v) [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) : (History.FinitePairHistory Action Reward n × (Action × Reward)) ≃ᵐ History.FinitePairHistory Action Reward (n + 1)","missing":[],"search":"pairhistorysuccmeasurableequiv banditrlproof.lowerbounds.pairhistorysuccmeasurableequiv a finite history and one next pair are measurably equivalent to the successor history. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.pairHistoryZeroMeasurableEquiv_apply","label":"pairHistoryZeroMeasurableEquiv_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.pairHistoryZeroMeasurableEquiv_apply","description":"theorem pairHistoryZeroMeasurableEquiv_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (pair : Action × Reward) : pairHistoryZeroMeasurableEquiv Action Reward pair = Thompson.singletonPairHistory pair","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-8999ffda2d22","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5257,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pairHistoryZeroMeasurableEquiv_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (pair : Action × Reward) : pairHistoryZeroMeasurableEquiv Action Reward pair = Thompson.singletonPairHistory pair","missing":[],"search":"pairhistoryzeromeasurableequiv_apply banditrlproof.lowerbounds.pairhistoryzeromeasurableequiv_apply theorem pairhistoryzeromeasurableequiv_apply {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] (pair : action × reward) : pairhistoryzeromeasurableequiv action reward pair = thompson.singletonpairhistory pair theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.pairHistorySuccMeasurableEquiv_apply","label":"pairHistorySuccMeasurableEquiv_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.pairHistorySuccMeasurableEquiv_apply","description":"theorem pairHistorySuccMeasurableEquiv_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) (input : History.FinitePairHistory Action Reward n × (Action × Reward)) : pairHistorySuccMeasurableEquiv Action Reward n input = History.extendPairHistorySucc input.1 input.2","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-a7c2d0a88173","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5258,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pairHistorySuccMeasurableEquiv_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (n : Nat) (input : History.FinitePairHistory Action Reward n × (Action × Reward)) : pairHistorySuccMeasurableEquiv Action Reward n input = History.extendPairHistorySucc input.1 input.2","missing":[],"search":"pairhistorysuccmeasurableequiv_apply banditrlproof.lowerbounds.pairhistorysuccmeasurableequiv_apply theorem pairhistorysuccmeasurableequiv_apply {action : type u} {reward : type v} [measurablespace action] [measurablespace reward] (n : nat) (input : history.finitepairhistory action reward n × (action × reward)) : pairhistorysuccmeasurableequiv action reward n input = history.extendpairhistorysucc input.1 input.2 theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment","label":"stationaryBanditHistoryEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment","description":"A stationary arm-indexed reward kernel viewed as a history environment.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-0e4431a95134","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5259,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def stationaryBanditHistoryEnvironment {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] : Thompson.HistoryEnvironment (Fin K) Reward where","missing":[],"search":"stationarybandithistoryenvironment banditrlproof.lowerbounds.stationarybandithistoryenvironment a stationary arm-indexed reward kernel viewed as a history environment. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment_initialFeedback","label":"stationaryBanditHistoryEnvironment_initialFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment_initialFeedback","description":"theorem stationaryBanditHistoryEnvironment_initialFeedback {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] : (stationaryBanditHistoryEnvironment armLaw).initialFeedback = armLaw","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-ad16fb13c7bb","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5260,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stationaryBanditHistoryEnvironment_initialFeedback {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] : (stationaryBanditHistoryEnvironment armLaw).initialFeedback = armLaw","missing":[],"search":"stationarybandithistoryenvironment_initialfeedback banditrlproof.lowerbounds.stationarybandithistoryenvironment_initialfeedback theorem stationarybandithistoryenvironment_initialfeedback {k : nat} {reward : type v} [measurablespace reward] (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] : (stationarybandithistoryenvironment armlaw).initialfeedback = armlaw theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment_feedback_apply","label":"stationaryBanditHistoryEnvironment_feedback_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment_feedback_apply","description":"theorem stationaryBanditHistoryEnvironment_feedback_apply {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : (stationaryBanditHistoryEnvironment armLaw).feedback n (history, arm) = armLaw arm","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-2dffc10867f5","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5261,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stationaryBanditHistoryEnvironment_feedback_apply {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : (stationaryBanditHistoryEnvironment armLaw).feedback n (history, arm) = armLaw arm","missing":[],"search":"stationarybandithistoryenvironment_feedback_apply banditrlproof.lowerbounds.stationarybandithistoryenvironment_feedback_apply theorem stationarybandithistoryenvironment_feedback_apply {k : nat} {reward : type v} [measurablespace reward] (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (n : nat) (history : history.finitepairhistory (fin k) reward n) (arm : fin k) : (stationarybandithistoryenvironment armlaw).feedback n (history, arm) = armlaw arm theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure","label":"canonicalBanditHistoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure","description":"Law of the observable action/reward history through the inclusive round `n`.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-a6407cdc4388","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5262,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:116"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBanditHistoryMeasure {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) : Measure (History.FinitePairHistory (Fin K) Reward n)","missing":[],"search":"canonicalbandithistorymeasure banditrlproof.lowerbounds.canonicalbandithistorymeasure law of the observable action/reward history through the inclusive round `n`. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_zero","label":"canonicalBanditHistoryMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_zero","description":"The inclusive time-zero history is the initial action/reward law in singleton form.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-3cb9c40963b6","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5263,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalBanditHistoryMeasure_zero {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] : canonicalBanditHistoryMeasure algorithm armLaw 0 = (algorithm.initialAction ⊗ₘ armLaw).map (pairHistoryZeroMeasurableEquiv (Fin K) Reward)","missing":[],"search":"canonicalbandithistorymeasure_zero banditrlproof.lowerbounds.canonicalbandithistorymeasure_zero the inclusive time-zero history is the initial action/reward law in singleton form. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_succ","label":"canonicalBanditHistoryMeasure_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_succ","description":"The inclusive successor history is obtained from the prefix history and the canonical same-policy action/reward step, then re-encoded by a measurable equivalence.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-5c53b9cec65a","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5264,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:181"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalBanditHistoryMeasure_succ {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) : canonicalBanditHistoryMeasure algorithm armLaw (n + 1) = (canonicalBanditHistoryMeasure algorithm armLaw n ⊗ₘ Thompson.historyStepKernel algorithm (stationaryBanditHistoryEnvironment armLaw) n).map (pairHistorySuccMeasurableEquiv (Fin K) Reward n)","missing":[],"search":"canonicalbandithistorymeasure_succ banditrlproof.lowerbounds.canonicalbandithistorymeasure_succ the inclusive successor history is obtained from the prefix history and the canonical same-policy action/reward step, then re-encoded by a measurable equivalence. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_zero","label":"klDiv_canonicalBanditHistoryMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_zero","description":"At the first pull, same-policy history KL is the initial-action average of the arm-law KL divergence. The expectation is under the first environment.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-8f1b0f2611e1","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5265,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:240"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_zero {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (h_ac : ∀ arm, armLaw arm ≪ referenceArmLaw arm) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw 0) (canonicalBanditHistoryMeasure algorithm referenceArmLaw 0) = ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂algorithm.initialAction","missing":[],"search":"kldiv_canonicalbandithistorymeasure_zero banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_zero at the first pull, same-policy history kl is the initial-action average of the arm-law kl divergence. the expectation is under the first environment. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_succ","label":"klDiv_canonicalBanditHistoryMeasure_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_succ","description":"One successor round adds the first-environment policy average of the selected arm KL divergence. This is the exact adaptive-history chain-rule step; the possibly randomized policy is shared and contributes no additional KL term.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-7cc6816db062","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5266,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:263"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_succ {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (h_ac : ∀ arm, armLaw arm ≪ referenceArmLaw arm) (n : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw (n + 1)) (canonicalBanditHistoryMeasure algorithm referenceArmLaw (n + 1)) = InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw n) (canonicalBanditHistoryMeasure algorithm referenceArmLaw n) + ∫⁻ history, ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂algorithm.policy n history ∂canonicalBanditHistoryMeasure algorithm armLaw n","missing":[],"search":"kldiv_canonicalbandithistorymeasure_succ banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_succ one successor round adds the first-environment policy average of the selected arm kl divergence. this is the exact adaptive-history chain-rule step; the possibly randomized policy is shared and contributes no additional kl term. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalPolicyArmMass","label":"canonicalPolicyArmMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalPolicyArmMass","description":"First-environment probability mass assigned to one arm at successor round `n + 1`.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-2a3e485ecaeb","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5267,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:294"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalPolicyArmMass {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : ENNReal","missing":[],"search":"canonicalpolicyarmmass banditrlproof.lowerbounds.canonicalpolicyarmmass first-environment probability mass assigned to one arm at successor round `n + 1`. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough","label":"canonicalExpectedPullCountThrough","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough","description":"Expected number of selections of `arm` through inclusive round `n`: the initial-action mass plus the successor policy masses for rounds `1, ..., n`. The equality with the lower integral of the realized finite-history pull count is proved below; this definition exposes the recurrence used by the KL proof.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-f5817c56b136","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5268,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:310"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalExpectedPullCountThrough {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : ENNReal","missing":[],"search":"canonicalexpectedpullcountthrough banditrlproof.lowerbounds.canonicalexpectedpullcountthrough expected number of selections of `arm` through inclusive round `n`: the initial-action mass plus the successor policy masses for rounds `1, ..., n`. the equality with the lower integral of the realized finite-history pull count is proved below; this definition exposes the recurrence used by the kl proof. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough_zero","label":"canonicalExpectedPullCountThrough_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough_zero","description":"theorem canonicalExpectedPullCountThrough_zero {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (arm : Fin K) : canonicalExpectedPullCountThrough algorithm armLaw 0 arm = algorithm.initialAction ({arm} : Set (Fin K))","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-903c9b5d0810","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5269,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:321"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalExpectedPullCountThrough_zero {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (arm : Fin K) : canonicalExpectedPullCountThrough algorithm armLaw 0 arm = algorithm.initialAction ({arm} : Set (Fin K))","missing":[],"search":"canonicalexpectedpullcountthrough_zero banditrlproof.lowerbounds.canonicalexpectedpullcountthrough_zero theorem canonicalexpectedpullcountthrough_zero {k : nat} {reward : type v} [measurablespace reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (arm : fin k) : canonicalexpectedpullcountthrough algorithm armlaw 0 arm = algorithm.initialaction ({arm} : set (fin k)) theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough_succ","label":"canonicalExpectedPullCountThrough_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough_succ","description":"theorem canonicalExpectedPullCountThrough_succ {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : canonicalExpectedPullCountThrough algorithm armLaw (n + 1) arm = canonicalExpectedPullCountThrough algorithm armLaw n arm + canonicalPolicyArmMass algorithm armLaw n arm","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-a6711c8ff41d","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5270,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:331"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalExpectedPullCountThrough_succ {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : canonicalExpectedPullCountThrough algorithm armLaw (n + 1) arm = canonicalExpectedPullCountThrough algorithm armLaw n arm + canonicalPolicyArmMass algorithm armLaw n arm","missing":[],"search":"canonicalexpectedpullcountthrough_succ banditrlproof.lowerbounds.canonicalexpectedpullcountthrough_succ theorem canonicalexpectedpullcountthrough_succ {k : nat} {reward : type v} [measurablespace reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (n : nat) (arm : fin k) : canonicalexpectedpullcountthrough algorithm armlaw (n + 1) arm = canonicalexpectedpullcountthrough algorithm armlaw n arm + canonicalpolicyarmmass algorithm armlaw n arm theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_lintegral_fin_eq_sum_armMass_mul","label":"lintegral_lintegral_fin_eq_sum_armMass_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_lintegral_fin_eq_sum_armMass_mul","description":"A finite-action conditional cost integral regroups by action mass.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-fdc4bc967820","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5271,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:344"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_lintegral_fin_eq_sum_armMass_mul {K : Nat} {History : Type*} [MeasurableSpace History] (historyLaw : Measure History) (policy : Kernel History (Fin K)) [IsMarkovKernel policy] (cost : Fin K → ENNReal) : (∫⁻ history, ∫⁻ arm, cost arm ∂policy history ∂historyLaw) = ∑ arm : Fin K, (∫⁻ history, policy history ({arm} : Set (Fin K)) ∂historyLaw) * cost arm","missing":[],"search":"lintegral_lintegral_fin_eq_sum_armmass_mul banditrlproof.lowerbounds.lintegral_lintegral_fin_eq_sum_armmass_mul a finite-action conditional cost integral regroups by action mass. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL","label":"klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL","description":"Policy-mass recurrence form of the finite-history KL decomposition.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-be0daf571eb2","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5272,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:367"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (h_ac : ∀ arm, armLaw arm ≪ referenceArmLaw arm) (n : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw n) (canonicalBanditHistoryMeasure algorithm referenceArmLaw n) = ∑ arm : Fin K, canonicalExpectedPullCountThrough algorithm armLaw n arm * InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm)","missing":[],"search":"kldiv_canonicalbandithistorymeasure_eq_sum_expectedpullcount_mul_armkl banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_eq_sum_expectedpullcount_mul_armkl policy-mass recurrence form of the finite-history kl decomposition. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal","label":"finiteHistoryPullCountENNReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal","description":"Realized pull count encoded directly on an inclusive finite history.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-294e793e8702","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5273,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:401"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryPullCountENNReal {K : Nat} {Reward : Type v} : (n : Nat) → History.FinitePairHistory (Fin K) Reward n → Fin K → ENNReal | 0, history, arm => if (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 = arm then 1 else 0 | n + 1, history, arm => finiteHistoryPullCountENNReal n (Thompson.pairHistoryPrefix history) arm + if (Thompson.pairHistoryLast history).1 = arm then 1 else 0 theorem measurable_finiteHistoryPullCountENNReal {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryPullCountENNReal n history arm)","missing":[],"search":"finitehistorypullcountennreal banditrlproof.lowerbounds.finitehistorypullcountennreal realized pull count encoded directly on an inclusive finite history. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryPullCountENNReal","label":"measurable_finiteHistoryPullCountENNReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_finiteHistoryPullCountENNReal","description":"theorem measurable_finiteHistoryPullCountENNReal {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryPullCountENNReal n history arm)","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-7818e163614c","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5274,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:410"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryPullCountENNReal {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryPullCountENNReal n history arm)","missing":[],"search":"measurable_finitehistorypullcountennreal banditrlproof.lowerbounds.measurable_finitehistorypullcountennreal theorem measurable_finitehistorypullcountennreal {k : nat} {reward : type v} [measurablespace reward] (n : nat) (arm : fin k) : measurable (fun history : history.finitepairhistory (fin k) reward n => finitehistorypullcountennreal n history arm) theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_pairHistoryZeroMeasurableEquiv","label":"finiteHistoryPullCountENNReal_pairHistoryZeroMeasurableEquiv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_pairHistoryZeroMeasurableEquiv","description":"theorem finiteHistoryPullCountENNReal_pairHistoryZeroMeasurableEquiv {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (pair : Fin K × Reward) (arm : Fin K) : finiteHistoryPullCountENNReal 0 (pairHistoryZeroMeasurableEquiv (Fin K) Reward pair) arm = if pair.1 = arm then 1 else 0","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-e4e59297b7f2","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5275,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:431"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPullCountENNReal_pairHistoryZeroMeasurableEquiv {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (pair : Fin K × Reward) (arm : Fin K) : finiteHistoryPullCountENNReal 0 (pairHistoryZeroMeasurableEquiv (Fin K) Reward pair) arm = if pair.1 = arm then 1 else 0","missing":[],"search":"finitehistorypullcountennreal_pairhistoryzeromeasurableequiv banditrlproof.lowerbounds.finitehistorypullcountennreal_pairhistoryzeromeasurableequiv theorem finitehistorypullcountennreal_pairhistoryzeromeasurableequiv {k : nat} {reward : type v} [measurablespace reward] (pair : fin k × reward) (arm : fin k) : finitehistorypullcountennreal 0 (pairhistoryzeromeasurableequiv (fin k) reward pair) arm = if pair.1 = arm then 1 else 0 theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_pairHistorySuccMeasurableEquiv","label":"finiteHistoryPullCountENNReal_pairHistorySuccMeasurableEquiv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_pairHistorySuccMeasurableEquiv","description":"theorem finiteHistoryPullCountENNReal_pairHistorySuccMeasurableEquiv {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (n : Nat) (input : History.FinitePairHistory (Fin K) Reward n × (Fin K × Reward)) (arm : Fin K) : finiteHistoryPullCountENNReal (n + 1) (pairHistorySuccMeasurableEquiv (Fin K) Reward n input) arm = finiteHistoryPullCountENNReal n input.1 arm + if input.2.1 = arm then 1 else 0","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-a167a453febf","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5276,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:443"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPullCountENNReal_pairHistorySuccMeasurableEquiv {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (n : Nat) (input : History.FinitePairHistory (Fin K) Reward n × (Fin K × Reward)) (arm : Fin K) : finiteHistoryPullCountENNReal (n + 1) (pairHistorySuccMeasurableEquiv (Fin K) Reward n input) arm = finiteHistoryPullCountENNReal n input.1 arm + if input.2.1 = arm then 1 else 0","missing":[],"search":"finitehistorypullcountennreal_pairhistorysuccmeasurableequiv banditrlproof.lowerbounds.finitehistorypullcountennreal_pairhistorysuccmeasurableequiv theorem finitehistorypullcountennreal_pairhistorysuccmeasurableequiv {k : nat} {reward : type v} [measurablespace reward] (n : nat) (input : history.finitepairhistory (fin k) reward n × (fin k × reward)) (arm : fin k) : finitehistorypullcountennreal (n + 1) (pairhistorysuccmeasurableequiv (fin k) reward n input) arm = finitehistorypullcountennreal n input.1 arm + if input.2.1 = arm then 1 else 0 theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough","label":"canonicalRealizedExpectedPullCountThrough","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough","description":"Lower integral of the realized pull count on the generated finite history.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-3ee0d50bfdc2","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5277,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:457"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalRealizedExpectedPullCountThrough {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : ENNReal","missing":[],"search":"canonicalrealizedexpectedpullcountthrough banditrlproof.lowerbounds.canonicalrealizedexpectedpullcountthrough lower integral of the realized pull count on the generated finite history. definition compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_historyStepKernel_armIndicator_eq_policy_mass","label":"lintegral_historyStepKernel_armIndicator_eq_policy_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_historyStepKernel_armIndicator_eq_policy_mass","description":"The action indicator under a history-step kernel integrates to policy mass.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-10d7cdfe1418","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5278,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:468"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_historyStepKernel_armIndicator_eq_policy_mass {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (environment : Thompson.HistoryEnvironment (Fin K) Reward) (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : (∫⁻ pair, ({arm} : Set (Fin K)).indicator (fun _ => (1 : ENNReal)) pair.1 ∂Thompson.historyStepKernel algorithm environment n history) = algorithm.policy n history ({arm} : Set (Fin K))","missing":[],"search":"lintegral_historystepkernel_armindicator_eq_policy_mass banditrlproof.lowerbounds.lintegral_historystepkernel_armindicator_eq_policy_mass the action indicator under a history-step kernel integrates to policy mass. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_zero","label":"canonicalRealizedExpectedPullCountThrough_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_zero","description":"theorem canonicalRealizedExpectedPullCountThrough_zero {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw 0 arm = algorithm.initialAction ({arm} : Set (Fin K))","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-dc69d74d9135","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5279,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:505"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalRealizedExpectedPullCountThrough_zero {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw 0 arm = algorithm.initialAction ({arm} : Set (Fin K))","missing":[],"search":"canonicalrealizedexpectedpullcountthrough_zero banditrlproof.lowerbounds.canonicalrealizedexpectedpullcountthrough_zero theorem canonicalrealizedexpectedpullcountthrough_zero {k : nat} {reward : type v} [measurablespace reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (arm : fin k) : canonicalrealizedexpectedpullcountthrough algorithm armlaw 0 arm = algorithm.initialaction ({arm} : set (fin k)) theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_succ","label":"canonicalRealizedExpectedPullCountThrough_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_succ","description":"theorem canonicalRealizedExpectedPullCountThrough_succ {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw (n + 1) arm = canonicalRealizedExpectedPullCountThrough algorithm armLaw n arm + canonicalPolicyArmMass algorithm…","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-e6413e858b14","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5280,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:538"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalRealizedExpectedPullCountThrough_succ {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw (n + 1) arm = canonicalRealizedExpectedPullCountThrough algorithm armLaw n arm + canonicalPolicyArmMass algorithm armLaw n arm","missing":[],"search":"canonicalrealizedexpectedpullcountthrough_succ banditrlproof.lowerbounds.canonicalrealizedexpectedpullcountthrough_succ theorem canonicalrealizedexpectedpullcountthrough_succ {k : nat} {reward : type v} [measurablespace reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (n : nat) (arm : fin k) : canonicalrealizedexpectedpullcountthrough algorithm armlaw (n + 1) arm = canonicalrealizedexpectedpullcountthrough algorithm armlaw n arm + canonicalpolicyarmmass algorithm armlaw n arm theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough","label":"canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough","description":"The policy-mass recurrence is exactly the lower integral of the realized pull count on the first-environment finite history.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-d596c5e2ab3c","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5281,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:597"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough {K : Nat} {Reward : Type v} [MeasurableSpace Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (n : Nat) (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw n arm = canonicalExpectedPullCountThrough algorithm armLaw n arm","missing":[],"search":"canonicalrealizedexpectedpullcountthrough_eq_expectedpullcountthrough banditrlproof.lowerbounds.canonicalrealizedexpectedpullcountthrough_eq_expectedpullcountthrough the policy-mass recurrence is exactly the lower integral of the realized pull count on the first-environment finite history. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_zero_general","label":"klDiv_canonicalBanditHistoryMeasure_zero_general","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_zero_general","description":"The first-pull KL identity, including singular and infinite-divergence arms.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-03c5b932910d","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5282,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:614"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_zero_general {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw 0) (canonicalBanditHistoryMeasure algorithm referenceArmLaw 0) = ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂algorithm.initialAction","missing":[],"search":"kldiv_canonicalbandithistorymeasure_zero_general banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_zero_general the first-pull kl identity, including singular and infinite-divergence arms. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_succ_general","label":"klDiv_canonicalBanditHistoryMeasure_succ_general","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_succ_general","description":"General successor KL recursion for a common randomized history policy.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-99b85de3885f","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5283,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:632"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_succ_general {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (n : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw (n + 1)) (canonicalBanditHistoryMeasure algorithm referenceArmLaw (n + 1)) = InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw n) (canonicalBanditHistoryMeasure algorithm referenceArmLaw n) + ∫⁻ history, ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂algorithm.policy n history ∂canonicalBanditHistoryMeasure algorithm armLaw n","missing":[],"search":"kldiv_canonicalbandithistorymeasure_succ_general banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_succ_general general successor kl recursion for a common randomized history policy. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL_general","label":"klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL_general","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL_general","description":"Policy-mass form of the unrestricted finite-history decomposition.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-85b987806216","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5284,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:662"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL_general {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (n : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw n) (canonicalBanditHistoryMeasure algorithm referenceArmLaw n) = ∑ arm : Fin K, canonicalExpectedPullCountThrough algorithm armLaw n arm * InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm)","missing":[],"search":"kldiv_canonicalbandithistorymeasure_eq_sum_expectedpullcount_mul_armkl_general banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_eq_sum_expectedpullcount_mul_armkl_general policy-mass form of the unrestricted finite-history decomposition. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_realizedExpectedPullCount_mul_armKL","label":"klDiv_canonicalBanditHistoryMeasure_eq_sum_realizedExpectedPullCount_mul_armKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_realizedExpectedPullCount_mul_armKL","description":"Lattimore--Szepesvari Lemma 15.1 in the repository's inclusive-round convention. `canonicalBanditHistoryMeasure ... n` contains exactly `n + 1` action/reward pairs. Its directed KL divergence is the finite-arm sum of the first-environment lower integrals of the realized pull counts through round `n`, multiplied by the arm-law KL divergences. The algorithm is one common, possibly randomized, nonanticipating history p…","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-cea746c89ce4","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5285,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:702"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_canonicalBanditHistoryMeasure_eq_sum_realizedExpectedPullCount_mul_armKL {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (n : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw n) (canonicalBanditHistoryMeasure algorithm referenceArmLaw n) = ∑ arm : Fin K, canonicalRealizedExpectedPullCountThrough algorithm armLaw n arm * InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm)","missing":[],"search":"kldiv_canonicalbandithistorymeasure_eq_sum_realizedexpectedpullcount_mul_armkl banditrlproof.lowerbounds.kldiv_canonicalbandithistorymeasure_eq_sum_realizedexpectedpullcount_mul_armkl lattimore--szepesvari lemma 15.1 in the repository's inclusive-round convention. `canonicalbandithistorymeasure ... n` contains exactly `n + 1` action/reward pairs. its directed kl divergence is the finite-arm sum of the first-environment lower integrals of the realized pull counts through round `n`, multiplied by the arm-law kl divergences. the algorithm is one common, possibly randomized, nonanticipating history policy in both environments. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","label":"banditHistoryRelativeEntropy_eq_expectedPulls_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","description":"Source-facing name for Lattimore--Szepesvari Lemma 15.1 / Eq. (15.1). The local history index `lastRound` represents exactly `lastRound + 1` pulls.","url":"../modules/banditrlproof-lowerbounds-bandithistorykl/index.html#decl-09d5b6143563","parent":"module:BanditRLProof.LowerBounds.BanditHistoryKL","order":5286,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BanditHistoryKL"],["Source","BanditRLProof/LowerBounds/BanditHistoryKL.lean:726"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem banditHistoryRelativeEntropy_eq_expectedPulls_sum {K : Nat} {Reward : Type v} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (lastRound : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw lastRound) (canonicalBanditHistoryMeasure algorithm referenceArmLaw lastRound) = ∑ arm : Fin K, canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm * InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm)","missing":[],"search":"bandithistoryrelativeentropy_eq_expectedpulls_sum banditrlproof.lowerbounds.bandithistoryrelativeentropy_eq_expectedpulls_sum source-facing name for lattimore--szepesvari lemma 15.1 / eq. (15.1). the local history index `lastround` represents exactly `lastround + 1` pulls. theorem compiled","shard":"modules/4185177a88ce70a9.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.worstCaseExpectedRegret","label":"worstCaseExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.worstCaseExpectedRegret","description":"Worst-case expected regret of one policy over an explicit environment class. The `ENNReal` codomain supplies the source-level supremum without a hidden boundedness hypothesis. If `environmentClass` is empty, this retains the standard complete-lattice value `bot`; meaningful bandit consumers should prove their intended class is nonempty.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-dcd474eaafe4","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5287,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"noncomputable def worstCaseExpectedRegret {Policy : Type u} {Environment : Type v} (regret : Policy -> Environment -> ENNReal) (environmentClass : Set Environment) (policy : Policy) : ENNReal","missing":[],"search":"worstcaseexpectedregret banditrlproof.lowerbounds.worstcaseexpectedregret worst-case expected regret of one policy over an explicit environment class. the `ennreal` codomain supplies the source-level supremum without a hidden boundedness hypothesis. if `environmentclass` is empty, this retains the standard complete-lattice value `bot`; meaningful bandit consumers should prove their intended class is nonempty. definition compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret","label":"minimaxExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.minimaxExpectedRegret","description":"Minimax expected regret over explicit policy and environment classes. The definition mirrors `inf_pi sup_nu R_n(pi,nu)`. Nonemptiness of the classes is intentionally a consumer-side semantic contract rather than a hidden assumption of the definition.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-512d89b0f3f1","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5288,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"noncomputable def minimaxExpectedRegret {Policy : Type u} {Environment : Type v} (regret : Policy -> Environment -> ENNReal) (policyClass : Set Policy) (environmentClass : Set Environment) : ENNReal","missing":[],"search":"minimaxexpectedregret banditrlproof.lowerbounds.minimaxexpectedregret minimax expected regret over explicit policy and environment classes. the definition mirrors `inf_pi sup_nu r_n(pi,nu)`. nonemptiness of the classes is intentionally a consumer-side semantic contract rather than a hidden assumption of the definition. definition compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal","label":"IsMinimaxOptimal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsMinimaxOptimal","description":"A policy is minimax optimal for an explicit policy class, environment class, and fixed-horizon regret functional when it is admissible and its worst-case regret attains the minimax value. The horizon is carried by `regret`; keeping the two classes explicit records the source warning that minimax optimality is not a property of a policy alone.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-250ad508cc6e","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5289,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"def IsMinimaxOptimal {Policy : Type u} {Environment : Type v} (regret : Policy -> Environment -> ENNReal) (policyClass : Set Policy) (environmentClass : Set Environment) (policy : Policy) : Prop","missing":[],"search":"isminimaxoptimal banditrlproof.lowerbounds.isminimaxoptimal a policy is minimax optimal for an explicit policy class, environment class, and fixed-horizon regret functional when it is admissible and its worst-case regret attains the minimax value. the horizon is carried by `regret`; keeping the two classes explicit records the source warning that minimax optimality is not a property of a policy alone. definition compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal.mem_policyClass","label":"mem_policyClass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsMinimaxOptimal.mem_policyClass","description":"A minimax-optimal policy belongs to the policy class over which the infimum is taken.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-8c1eaa33d51a","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5290,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsMinimaxOptimal.mem_policyClass {Policy : Type u} {Environment : Type v} {regret : Policy -> Environment -> ENNReal} {policyClass : Set Policy} {environmentClass : Set Environment} {policy : Policy} (hpolicy : IsMinimaxOptimal regret policyClass environmentClass policy) : policy ∈ policyClass","missing":[],"search":"mem_policyclass banditrlproof.lowerbounds.isminimaxoptimal.mem_policyclass a minimax-optimal policy belongs to the policy class over which the infimum is taken. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal.eq_minimaxExpectedRegret","label":"eq_minimaxExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsMinimaxOptimal.eq_minimaxExpectedRegret","description":"A minimax-optimal policy attains the fixed-class minimax value.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-61667f78f62f","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5291,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsMinimaxOptimal.eq_minimaxExpectedRegret {Policy : Type u} {Environment : Type v} {regret : Policy -> Environment -> ENNReal} {policyClass : Set Policy} {environmentClass : Set Environment} {policy : Policy} (hpolicy : IsMinimaxOptimal regret policyClass environmentClass policy) : worstCaseExpectedRegret regret environmentClass policy = minimaxExpectedRegret regret policyClass environmentClass","missing":[],"search":"eq_minimaxexpectedregret banditrlproof.lowerbounds.isminimaxoptimal.eq_minimaxexpectedregret a minimax-optimal policy attains the fixed-class minimax value. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedRegret_le_worstCaseExpectedRegret","label":"expectedRegret_le_worstCaseExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedRegret_le_worstCaseExpectedRegret","description":"One environment's regret is below the worst case over any class containing it.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-562534467a54","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5292,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_le_worstCaseExpectedRegret {Policy : Type u} {Environment : Type v} (regret : Policy -> Environment -> ENNReal) (environmentClass : Set Environment) (policy : Policy) (environment : Environment) (henvironment : environment ∈ environmentClass) : regret policy environment ≤ worstCaseExpectedRegret regret environmentClass policy","missing":[],"search":"expectedregret_le_worstcaseexpectedregret banditrlproof.lowerbounds.expectedregret_le_worstcaseexpectedregret one environment's regret is below the worst case over any class containing it. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret_le_worstCaseExpectedRegret","label":"minimaxExpectedRegret_le_worstCaseExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.minimaxExpectedRegret_le_worstCaseExpectedRegret","description":"The minimax value is below the worst-case value of each admissible policy.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-e826611df866","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5293,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem minimaxExpectedRegret_le_worstCaseExpectedRegret {Policy : Type u} {Environment : Type v} (regret : Policy -> Environment -> ENNReal) (policyClass : Set Policy) (environmentClass : Set Environment) (policy : Policy) (hpolicy : policy ∈ policyClass) : minimaxExpectedRegret regret policyClass environmentClass ≤ worstCaseExpectedRegret regret environmentClass policy","missing":[],"search":"minimaxexpectedregret_le_worstcaseexpectedregret banditrlproof.lowerbounds.minimaxexpectedregret_le_worstcaseexpectedregret the minimax value is below the worst-case value of each admissible policy. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.le_minimaxExpectedRegret","label":"le_minimaxExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.le_minimaxExpectedRegret","description":"A uniform lower bound on every admissible policy is a minimax lower bound.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-023a08905f78","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5294,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:120"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem le_minimaxExpectedRegret {Policy : Type u} {Environment : Type v} (regret : Policy -> Environment -> ENNReal) (policyClass : Set Policy) (environmentClass : Set Environment) (lower : ENNReal) (hlower : ∀ policy : policyClass, lower ≤ worstCaseExpectedRegret regret environmentClass policy.1) : lower ≤ minimaxExpectedRegret regret policyClass environmentClass","missing":[],"search":"le_minimaxexpectedregret banditrlproof.lowerbounds.le_minimaxexpectedregret a uniform lower bound on every admissible policy is a minimax lower bound. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_alternative_le_average","label":"exists_alternative_le_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_alternative_le_average","description":"Finite averaging: among a nonempty family whose sum is at most `budget`, one coordinate is at most `budget / m`. This is a Mathlib-composed deterministic leaf. It contains no bandit law, expectation, measurability, or concentration assumption.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-3433e3e67817","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5295,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem exists_alternative_le_average {m : Nat} (hm : 0 < m) (alternativeExpectedPulls : Fin m -> Real) (budget : Real) (hbudget : ∑ i : Fin m, alternativeExpectedPulls i ≤ budget) : ∃ i : Fin m, alternativeExpectedPulls i ≤ budget / (m : Real)","missing":[],"search":"exists_alternative_le_average banditrlproof.lowerbounds.exists_alternative_le_average finite averaging: among a nonempty family whose sum is at most `budget`, one coordinate is at most `budget / m`. this is a mathlib-composed deterministic leaf. it contains no bandit law, expectation, measurability, or concentration assumption. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.alternativeExpectedPullBudget_le","label":"alternativeExpectedPullBudget_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.alternativeExpectedPullBudget_le","description":"Removing the distinguished arm zero from an exact nonnegative pull budget leaves at most the full budget on the `Fin.succ` alternative arms.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-c161c41634a8","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5296,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:164"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem alternativeExpectedPullBudget_le {m : Nat} (expectedPulls : Fin (m + 1) -> Real) (budget : Real) (hnonneg : ∀ arm, 0 ≤ expectedPulls arm) (htotal : ∑ arm : Fin (m + 1), expectedPulls arm = budget) : (∑ i : Fin m, expectedPulls i.succ) ≤ budget","missing":[],"search":"alternativeexpectedpullbudget_le banditrlproof.lowerbounds.alternativeexpectedpullbudget_le removing the distinguished arm zero from an exact nonnegative pull budget leaves at most the full budget on the `fin.succ` alternative arms. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_leastExploredAlternative","label":"exists_leastExploredAlternative","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_leastExploredAlternative","description":"Chapter 13's least-explored alternative arm. There are `m + 1` arms: arm zero is the base arm and `i.succ`, for `i : Fin m`, are the alternatives. The hypotheses are precisely the expected pull-count nonnegativity and total-budget identity needed by the averaging argument.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-a194d361149a","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5297,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:181"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem exists_leastExploredAlternative {m : Nat} (hm : 0 < m) (expectedPulls : Fin (m + 1) -> Real) (horizon : Nat) (hnonneg : ∀ arm, 0 ≤ expectedPulls arm) (htotal : ∑ arm : Fin (m + 1), expectedPulls arm = (horizon : Real)) : ∃ i : Fin m, expectedPulls i.succ ≤ (horizon : Real) / (m : Real)","missing":[],"search":"exists_leastexploredalternative banditrlproof.lowerbounds.exists_leastexploredalternative chapter 13's least-explored alternative arm. there are `m + 1` arms: arm zero is the base arm and `i.succ`, for `i : fin m`, are the alternatives. the hypotheses are precisely the expected pull-count nonnegativity and total-budget identity needed by the averaging argument. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.baseEnvironmentRegret","label":"baseEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.baseEnvironmentRegret","description":"The exact deterministic expression in the base-environment identity (13.2).","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-9a6cc4a86c94","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5298,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:194"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"def baseEnvironmentRegret (horizon : Nat) (gap baseFirstExpectedPulls : Real) : Real","missing":[],"search":"baseenvironmentregret banditrlproof.lowerbounds.baseenvironmentregret the exact deterministic expression in the base-environment identity (13.2). definition compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.changedEnvironmentRegretLowerBound","label":"changedEnvironmentRegretLowerBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.changedEnvironmentRegretLowerBound","description":"The changed-environment regret lower expression on the right of (13.3).","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-5bcaeef5449a","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5299,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:199"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"def changedEnvironmentRegretLowerBound (gap changedFirstExpectedPulls : Real) : Real","missing":[],"search":"changedenvironmentregretlowerbound banditrlproof.lowerbounds.changedenvironmentregretlowerbound the changed-environment regret lower expression on the right of (13.3). definition compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","label":"max_base_changed_regretLowerBound_ge_half_sub_error","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","description":"Quantitative algebraic core of the two-environment heuristic. The premise bounds the cross-environment discrepancy in the expected number of base-arm pulls. Chapter 13 writes these expectations as approximately equal; later information-theoretic chapters must supply an actual value of `error`. This theorem neither derives that premise nor proves Theorem 13.1.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-e094b7afa46b","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5300,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:211"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem max_base_changed_regretLowerBound_ge_half_sub_error (horizon : Nat) (gap baseFirstExpectedPulls changedFirstExpectedPulls error : Real) (hgap : 0 ≤ gap) (hpullDifference : baseFirstExpectedPulls - changedFirstExpectedPulls ≤ error) : gap * ((horizon : Real) - error) / 2 ≤ max (baseEnvironmentRegret horizon gap baseFirstExpectedPulls) (changedEnvironmentRegretLowerBound gap changedFirstExpectedPulls)","missing":[],"search":"max_base_changed_regretlowerbound_ge_half_sub_error banditrlproof.lowerbounds.max_base_changed_regretlowerbound_ge_half_sub_error quantitative algebraic core of the two-environment heuristic. the premise bounds the cross-environment discrepancy in the expected number of base-arm pulls. chapter 13 writes these expectations as approximately equal; later information-theoretic chapters must supply an actual value of `error`. this theorem neither derives that premise nor proves theorem 13.1. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half","label":"max_base_changed_regretLowerBound_ge_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half","description":"Zero-error corollary of the quantitative two-environment algebra. The premise `baseFirstExpectedPulls ≤ changedFirstExpectedPulls` specializes the pull discrepancy to at most zero. The quantitative predecessor is the intended interface for later information theorems; neither declaration is the Gaussian minimax lower bound.","url":"../modules/banditrlproof-lowerbounds-basicideas/index.html#decl-5573ca9f2d05","parent":"module:BanditRLProof.LowerBounds.BasicIdeas","order":5301,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BasicIdeas"],["Source","BanditRLProof/LowerBounds/BasicIdeas.lean:243"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem max_base_changed_regretLowerBound_ge_half (horizon : Nat) (gap baseFirstExpectedPulls changedFirstExpectedPulls : Real) (hgap : 0 ≤ gap) (htransport : baseFirstExpectedPulls ≤ changedFirstExpectedPulls) : gap * (horizon : Real) / 2 ≤ max (baseEnvironmentRegret horizon gap baseFirstExpectedPulls) (changedEnvironmentRegretLowerBound gap changedFirstExpectedPulls)","missing":[],"search":"max_base_changed_regretlowerbound_ge_half banditrlproof.lowerbounds.max_base_changed_regretlowerbound_ge_half zero-error corollary of the quantitative two-environment algebra. the premise `basefirstexpectedpulls ≤ changedfirstexpectedpulls` specializes the pull discrepancy to at most zero. the quantitative predecessor is the intended interface for later information theorems; neither declaration is the gaussian minimax lower bound. theorem compiled","shard":"modules/7efce33e480c5631.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.entropy_product_term","label":"entropy_product_term","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.entropy_product_term","description":"theorem entropy_product_term (p q : ℝ) : (p * q) * Real.log (p * q)⁻¹ = q * (p * Real.log p⁻¹) + p * (q * Real.log q⁻¹)","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-26b5b2c8a076","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5302,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem entropy_product_term (p q : ℝ) : (p * q) * Real.log (p * q)⁻¹ = q * (p * Real.log p⁻¹) + p * (q * Real.log q⁻¹)","missing":[],"search":"entropy_product_term banditrlproof.lowerbounds.entropy_product_term theorem entropy_product_term (p q : ℝ) : (p * q) * real.log (p * q)⁻¹ = q * (p * real.log p⁻¹) + p * (q * real.log q⁻¹) theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropy_prod","label":"discreteEntropy_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropy_prod","description":"Entropy of a product mass function before normalization.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-8b5d9112b787","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5303,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropy_prod {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (q : β → ℝ) : discreteEntropy Finset.univ (fun x : α × β => p x.1 * q x.2) = (∑ j, q j) * discreteEntropy Finset.univ p + (∑ i, p i) * discreteEntropy Finset.univ q","missing":[],"search":"discreteentropy_prod banditrlproof.lowerbounds.discreteentropy_prod entropy of a product mass function before normalization. theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropy_prod_probability","label":"discreteEntropy_prod_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropy_prod_probability","description":"theorem discreteEntropy_prod_probability {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (q : β → ℝ) (hp : ∑ i, p i = 1) (hq : ∑ j, q j = 1) : discreteEntropy Finset.univ (fun x : α × β => p x.1 * q x.2) = discreteEntropy Finset.univ p + discreteEntropy Finset.univ q","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-3638aaac89f9","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5304,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropy_prod_probability {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (q : β → ℝ) (hp : ∑ i, p i = 1) (hq : ∑ j, q j = 1) : discreteEntropy Finset.univ (fun x : α × β => p x.1 * q x.2) = discreteEntropy Finset.univ p + discreteEntropy Finset.univ q","missing":[],"search":"discreteentropy_prod_probability banditrlproof.lowerbounds.discreteentropy_prod_probability theorem discreteentropy_prod_probability {α β : type*} [fintype α] [fintype β] (p : α → ℝ) (q : β → ℝ) (hp : ∑ i, p i = 1) (hq : ∑ j, q j = 1) : discreteentropy finset.univ (fun x : α × β => p x.1 * q x.2) = discreteentropy finset.univ p + discreteentropy finset.univ q theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_prod_probability","label":"discreteEntropyBaseTwo_prod_probability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropyBaseTwo_prod_probability","description":"theorem discreteEntropyBaseTwo_prod_probability {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (q : β → ℝ) (hp : ∑ i, p i = 1) (hq : ∑ j, q j = 1) : discreteEntropyBaseTwo Finset.univ (fun x : α × β => p x.1 * q x.2) = discreteEntropyBaseTwo Finset.univ p + discreteEntropyBaseTwo Finset.univ q","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-5a77fa1ba61e","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5305,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropyBaseTwo_prod_probability {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (q : β → ℝ) (hp : ∑ i, p i = 1) (hq : ∑ j, q j = 1) : discreteEntropyBaseTwo Finset.univ (fun x : α × β => p x.1 * q x.2) = discreteEntropyBaseTwo Finset.univ p + discreteEntropyBaseTwo Finset.univ q","missing":[],"search":"discreteentropybasetwo_prod_probability banditrlproof.lowerbounds.discreteentropybasetwo_prod_probability theorem discreteentropybasetwo_prod_probability {α β : type*} [fintype α] [fintype β] (p : α → ℝ) (q : β → ℝ) (hp : ∑ i, p i = 1) (hq : ∑ j, q j = 1) : discreteentropybasetwo finset.univ (fun x : α × β => p x.1 * q x.2) = discreteentropybasetwo finset.univ p + discreteentropybasetwo finset.univ q theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.SourceBlock","label":"SourceBlock","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.SourceBlock","description":"An n-symbol source block, represented by a nested product.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-abf9184b2451","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5306,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def SourceBlock (α : Type*) : ℕ → Type _ | 0 => PUnit | n + 1 => α × SourceBlock α n instance sourceBlockFintype {α : Type*} [Fintype α] (n : ℕ) : Fintype (SourceBlock α n)","missing":[],"search":"sourceblock banditrlproof.lowerbounds.sourceblock an n-symbol source block, represented by a nested product. definition compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlockMass","label":"sourceBlockMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlockMass","description":"IID product mass, including the empty block of mass one.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-5aee27ce68ea","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5307,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def sourceBlockMass {α : Type*} (p : α → ℝ) : (n : ℕ) → SourceBlock α n → ℝ | 0, _ => 1 | n + 1, x => p x.1 * sourceBlockMass p n x.2 theorem sourceBlockMass_nonneg {α : Type*} (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (n : ℕ) (x : SourceBlock α n) : 0 ≤ sourceBlockMass p n x","missing":[],"search":"sourceblockmass banditrlproof.lowerbounds.sourceblockmass iid product mass, including the empty block of mass one. definition compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlockMass_nonneg","label":"sourceBlockMass_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlockMass_nonneg","description":"theorem sourceBlockMass_nonneg {α : Type*} (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (n : ℕ) (x : SourceBlock α n) : 0 ≤ sourceBlockMass p n x","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-0e051035a8fb","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5308,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceBlockMass_nonneg {α : Type*} (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (n : ℕ) (x : SourceBlock α n) : 0 ≤ sourceBlockMass p n x","missing":[],"search":"sourceblockmass_nonneg banditrlproof.lowerbounds.sourceblockmass_nonneg theorem sourceblockmass_nonneg {α : type*} (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (n : ℕ) (x : sourceblock α n) : 0 ≤ sourceblockmass p n x theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_sourceBlockMass","label":"sum_sourceBlockMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_sourceBlockMass","description":"theorem sum_sourceBlockMass {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∑ i, p i = 1) (n : ℕ) : ∑ x, sourceBlockMass p n x = 1","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-2c530cefbf02","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5309,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_sourceBlockMass {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∑ i, p i = 1) (n : ℕ) : ∑ x, sourceBlockMass p n x = 1","missing":[],"search":"sum_sourceblockmass banditrlproof.lowerbounds.sum_sourceblockmass theorem sum_sourceblockmass {α : type*} [fintype α] (p : α → ℝ) (hp : ∑ i, p i = 1) (n : ℕ) : ∑ x, sourceblockmass p n x = 1 theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_sourceBlockMass","label":"discreteEntropyBaseTwo_sourceBlockMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropyBaseTwo_sourceBlockMass","description":"theorem discreteEntropyBaseTwo_sourceBlockMass {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∑ i, p i = 1) (n : ℕ) : discreteEntropyBaseTwo Finset.univ (sourceBlockMass p n) = n * discreteEntropyBaseTwo Finset.univ p","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-5266ea90812f","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5310,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropyBaseTwo_sourceBlockMass {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∑ i, p i = 1) (n : ℕ) : discreteEntropyBaseTwo Finset.univ (sourceBlockMass p n) = n * discreteEntropyBaseTwo Finset.univ p","missing":[],"search":"discreteentropybasetwo_sourceblockmass banditrlproof.lowerbounds.discreteentropybasetwo_sourceblockmass theorem discreteentropybasetwo_sourceblockmass {α : type*} [fintype α] (p : α → ℝ) (hp : ∑ i, p i = 1) (n : ℕ) : discreteentropybasetwo finset.univ (sourceblockmass p n) = n * discreteentropybasetwo finset.univ p theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_sourceBlock_code_rate_sandwich","label":"exists_sourceBlock_code_rate_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_sourceBlock_code_rate_sandwich","description":"Finite-block source coding with an explicit one-bit total overhead.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-cc54acee69fa","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5311,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_sourceBlock_code_rate_sandwich {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) (hn : 0 < n) : ∃ code : BinaryPrefixCode (SourceBlock α n), discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength (sourceBlockMass p n) code / n ∧ expectedCodeLength (sourceBlockMass p n) code / n ≤ discreteEntropyBaseTwo Finset.univ p + 1 / n","missing":[],"search":"exists_sourceblock_code_rate_sandwich banditrlproof.lowerbounds.exists_sourceblock_code_rate_sandwich finite-block source coding with an explicit one-bit total overhead. theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlock_code_rate_lower_bound","label":"sourceBlock_code_rate_lower_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlock_code_rate_lower_bound","description":"Every finite block prefix code has rate at least the source entropy.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-632778be7da6","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5312,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sourceBlock_code_rate_lower_bound {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (n : ℕ) (hn : 0 < n) (code : BinaryPrefixCode (SourceBlock α n)) : discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength (sourceBlockMass p n) code / n","missing":[],"search":"sourceblock_code_rate_lower_bound banditrlproof.lowerbounds.sourceblock_code_rate_lower_bound every finite block prefix code has rate at least the source entropy. theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_sourceBlock_code_family_tendsto_entropy","label":"exists_sourceBlock_code_family_tendsto_entropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_sourceBlock_code_family_tendsto_entropy","description":"An actual family of block prefix codes has rate tending to entropy.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-4f79ababdc13","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5313,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:122"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_sourceBlock_code_family_tendsto_entropy {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : ∃ code : (n : ℕ) → BinaryPrefixCode (SourceBlock α (n + 1)), Filter.Tendsto (fun n => expectedCodeLength (sourceBlockMass p (n + 1)) (code n) / (n + 1)) Filter.atTop (nhds (discreteEntropyBaseTwo Finset.univ p))","missing":[],"search":"exists_sourceblock_code_family_tendsto_entropy banditrlproof.lowerbounds.exists_sourceblock_code_family_tendsto_entropy an actual family of block prefix codes has rate tending to entropy. theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sourceBlock_code_family_limit_ge_entropy","label":"sourceBlock_code_family_limit_ge_entropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sourceBlock_code_family_limit_ge_entropy","description":"No convergent family of block prefix codes has limiting rate below entropy.","url":"../modules/banditrlproof-lowerbounds-blockentropy/index.html#decl-a56e8ffa2fec","parent":"module:BanditRLProof.LowerBounds.BlockEntropy","order":5314,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.BlockEntropy"],["Source","BanditRLProof/LowerBounds/BlockEntropy.lean:142"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem sourceBlock_code_family_limit_ge_entropy {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (code : (n : ℕ) → BinaryPrefixCode (SourceBlock α (n + 1))) (r : ℝ) (hr : Filter.Tendsto (fun n => expectedCodeLength (sourceBlockMass p (n + 1)) (code n) / (n + 1)) Filter.atTop (nhds r)) : discreteEntropyBaseTwo Finset.univ p ≤ r","missing":[],"search":"sourceblock_code_family_limit_ge_entropy banditrlproof.lowerbounds.sourceblock_code_family_limit_ge_entropy no convergent family of block prefix codes has limiting rate below entropy. theorem compiled","shard":"modules/766f791ee94508ad.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.entropy_term_le_codeLength_remainder","label":"entropy_term_le_codeLength_remainder","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.entropy_term_le_codeLength_remainder","description":"theorem entropy_term_le_codeLength_remainder {p : ℝ} (hp : 0 ≤ p) (l : ℕ) : p * Real.log p⁻¹ ≤ p * l * Real.log 2 + (1 / 2 : ℝ) ^ l - p","url":"../modules/banditrlproof-lowerbounds-codingentropybound/index.html#decl-d0f1536ee35f","parent":"module:BanditRLProof.LowerBounds.CodingEntropyBound","order":5315,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CodingEntropyBound"],["Source","BanditRLProof/LowerBounds/CodingEntropyBound.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem entropy_term_le_codeLength_remainder {p : ℝ} (hp : 0 ≤ p) (l : ℕ) : p * Real.log p⁻¹ ≤ p * l * Real.log 2 + (1 / 2 : ℝ) ^ l - p","missing":[],"search":"entropy_term_le_codelength_remainder banditrlproof.lowerbounds.entropy_term_le_codelength_remainder theorem entropy_term_le_codelength_remainder {p : ℝ} (hp : 0 ≤ p) (l : ℕ) : p * real.log p⁻¹ ≤ p * l * real.log 2 + (1 / 2 : ℝ) ^ l - p theorem compiled","shard":"modules/70bffeb024e0eca8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_le_expectedCodeLength","label":"discreteEntropyBaseTwo_le_expectedCodeLength","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropyBaseTwo_le_expectedCodeLength","description":"The lower half of Eq. (14.2), for every finite binary prefix code.","url":"../modules/banditrlproof-lowerbounds-codingentropybound/index.html#decl-bd02aaeb3283","parent":"module:BanditRLProof.LowerBounds.CodingEntropyBound","order":5316,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CodingEntropyBound"],["Source","BanditRLProof/LowerBounds/CodingEntropyBound.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropyBaseTwo_le_expectedCodeLength {Symbol : Type*} [Fintype Symbol] [DecidableEq Symbol] (p : Symbol → ℝ) (hp : ∀ i, 0 ≤ p i) (hsum : ∑ i, p i = 1) (code : BinaryPrefixCode Symbol) : discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength p code","missing":[],"search":"discreteentropybasetwo_le_expectedcodelength banditrlproof.lowerbounds.discreteentropybasetwo_le_expectedcodelength the lower half of eq. (14.2), for every finite binary prefix code. theorem compiled","shard":"modules/70bffeb024e0eca8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.llr_ae_eq_log_commonDensity","label":"llr_ae_eq_log_commonDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.llr_ae_eq_log_commonDensity","description":"The log likelihood ratio is the log ratio of two common-measure densities.","url":"../modules/banditrlproof-lowerbounds-commondensitykl/index.html#decl-6c3075c1d5dc","parent":"module:BanditRLProof.LowerBounds.CommonDensityKL","order":5317,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityKL"],["Source","BanditRLProof/LowerBounds/CommonDensityKL.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem llr_ae_eq_log_commonDensity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) (hPQ : P ≪ Q) : llr P Q =ᵐ[P] fun x => Real.log ((P.rnDeriv μ x).toReal / (Q.rnDeriv μ x).toReal)","missing":[],"search":"llr_ae_eq_log_commondensity banditrlproof.lowerbounds.llr_ae_eq_log_commondensity the log likelihood ratio is the log ratio of two common-measure densities. theorem compiled","shard":"modules/63200a032e0cec96.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_commonDensity_iff","label":"integrable_commonDensity_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_commonDensity_iff","description":"Integrability of the common-density logarithmic integrand is exactly that of LLR.","url":"../modules/banditrlproof-lowerbounds-commondensitykl/index.html#decl-1bd9015556ee","parent":"module:BanditRLProof.LowerBounds.CommonDensityKL","order":5318,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityKL"],["Source","BanditRLProof/LowerBounds/CommonDensityKL.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_commonDensity_iff {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) (hPQ : P ≪ Q) : Integrable (fun x => (P.rnDeriv μ x).toReal * Real.log ((P.rnDeriv μ x).toReal / (Q.rnDeriv μ x).toReal)) μ ↔ Integrable (llr P Q) P","missing":[],"search":"integrable_commondensity_iff banditrlproof.lowerbounds.integrable_commondensity_iff integrability of the common-density logarithmic integrand is exactly that of llr. theorem compiled","shard":"modules/63200a032e0cec96.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_of_integrable","label":"relativeEntropy_commonDensity_of_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_commonDensity_of_integrable","description":"Eq. (14.6) in its supported, integrable probability-measure branch.","url":"../modules/banditrlproof-lowerbounds-commondensitykl/index.html#decl-1ffb807b9085","parent":"module:BanditRLProof.LowerBounds.CommonDensityKL","order":5319,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityKL"],["Source","BanditRLProof/LowerBounds/CommonDensityKL.lean:33"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_commonDensity_of_integrable {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) (hPQ : P ≪ Q) (hi : Integrable (fun x => (P.rnDeriv μ x).toReal * Real.log ((P.rnDeriv μ x).toReal / (Q.rnDeriv μ x).toReal)) μ) : relativeEntropy P Q = ENNReal.ofReal (∫ x, (P.rnDeriv μ x).toReal * Real.log ((P.rnDeriv μ x).toReal / (Q.rnDeriv μ x).toReal) ∂μ)","missing":[],"search":"relativeentropy_commondensity_of_integrable banditrlproof.lowerbounds.relativeentropy_commondensity_of_integrable eq. (14.6) in its supported, integrable probability-measure branch. theorem compiled","shard":"modules/63200a032e0cec96.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_eq_if","label":"relativeEntropy_commonDensity_eq_if","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_commonDensity_eq_if","description":"Common-density formula with singular and nonintegrable branches kept infinite.","url":"../modules/banditrlproof-lowerbounds-commondensitykl/index.html#decl-1250d3170511","parent":"module:BanditRLProof.LowerBounds.CommonDensityKL","order":5320,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityKL"],["Source","BanditRLProof/LowerBounds/CommonDensityKL.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_commonDensity_eq_if {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) : relativeEntropy P Q = if P ≪ Q ∧ Integrable (fun x => (P.rnDeriv μ x).toReal * Real.log ((P.rnDeriv μ x).toReal / (Q.rnDeriv μ x).toReal)) μ then ENNReal.ofReal (∫ x, (P.rnDeriv μ x).toReal * Real.log ((P.rnDeriv μ x).toReal / (Q.rnDeriv μ x).toReal) ∂μ) else (⊤ : ENNReal)","missing":[],"search":"relativeentropy_commondensity_eq_if banditrlproof.lowerbounds.relativeentropy_commondensity_eq_if common-density formula with singular and nonintegrable branches kept infinite. theorem compiled","shard":"modules/63200a032e0cec96.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_klFun","label":"relativeEntropy_commonDensity_klFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_commonDensity_klFun","description":"Nonnegative common-density KL integral, valid even at infinite KL under AC.","url":"../modules/banditrlproof-lowerbounds-commondensitykl/index.html#decl-7dbd18e2e941","parent":"module:BanditRLProof.LowerBounds.CommonDensityKL","order":5321,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityKL"],["Source","BanditRLProof/LowerBounds/CommonDensityKL.lean:68"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_commonDensity_klFun {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) (hPQ : P ≪ Q) : relativeEntropy P Q = ∫⁻ x, Q.rnDeriv μ x * ENNReal.ofReal (InformationTheory.klFun ((P.rnDeriv μ x / Q.rnDeriv μ x).toReal)) ∂μ","missing":[],"search":"relativeentropy_commondensity_klfun banditrlproof.lowerbounds.relativeentropy_commondensity_klfun nonnegative common-density kl integral, valid even at infinite kl under ac. theorem compiled","shard":"modules/63200a032e0cec96.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.memLp_sqrt_of_integrable_nonneg","label":"memLp_sqrt_of_integrable_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.memLp_sqrt_of_integrable_nonneg","description":"Square roots of nonnegative integrable functions belong to L2.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-7770d5e9e371","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5322,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem memLp_sqrt_of_integrable_nonneg {α : Type*} [MeasurableSpace α] {μ : Measure α} {f : α → ℝ} (hf : Integrable f μ) (hpos : ∀ x, 0 ≤ f x) : MemLp (fun x => Real.sqrt (f x)) 2 μ","missing":[],"search":"memlp_sqrt_of_integrable_nonneg banditrlproof.lowerbounds.memlp_sqrt_of_integrable_nonneg square roots of nonnegative integrable functions belong to l2. theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_sqrt_mul_sq_le","label":"integral_sqrt_mul_sq_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_sqrt_mul_sq_le","description":"Cauchy--Schwarz for the square-root affinity integrand.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-04f996d1a3f7","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5323,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_sqrt_mul_sq_le {α : Type*} [MeasurableSpace α] {μ : Measure α} {p q : α → ℝ} (hp : Integrable p μ) (hq : Integrable q μ) (hp0 : ∀ x, 0 ≤ p x) (hq0 : ∀ x, 0 ≤ q x) : (∫ x, Real.sqrt (p x * q x) ∂μ) ^ 2 ≤ (∫ x, p x ∂μ) * ∫ x, q x ∂μ","missing":[],"search":"integral_sqrt_mul_sq_le banditrlproof.lowerbounds.integral_sqrt_mul_sq_le cauchy--schwarz for the square-root affinity integrand. theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap","label":"commonDensityOverlap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.commonDensityOverlap","description":"Integral of the pointwise minimum of common RN densities.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-0f604586953a","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5324,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def commonDensityOverlap {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : ℝ","missing":[],"search":"commondensityoverlap banditrlproof.lowerbounds.commondensityoverlap integral of the pointwise minimum of common rn densities. definition compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.commonDensityAffinity","label":"commonDensityAffinity","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.commonDensityAffinity","description":"Hellinger affinity of two densities relative to a common measure.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-6f60c504c338","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5325,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:55"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def commonDensityAffinity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : ℝ","missing":[],"search":"commondensityaffinity banditrlproof.lowerbounds.commondensityaffinity hellinger affinity of two densities relative to a common measure. definition compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_commonDensityAffinity","label":"integrable_commonDensityAffinity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_commonDensityAffinity","description":"theorem integrable_commonDensityAffinity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : Integrable (fun x => Real.sqrt ((P.rnDeriv μ x).toReal * (Q.rnDeriv μ x).toReal)) μ","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-c4d7c928df07","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5326,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_commonDensityAffinity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : Integrable (fun x => Real.sqrt ((P.rnDeriv μ x).toReal * (Q.rnDeriv μ x).toReal)) μ","missing":[],"search":"integrable_commondensityaffinity banditrlproof.lowerbounds.integrable_commondensityaffinity theorem integrable_commondensityaffinity {α : type*} [measurablespace α] (p q μ : measure α) [isfinitemeasure p] [isfinitemeasure q] : integrable (fun x => real.sqrt ((p.rnderiv μ x).toreal * (q.rnderiv μ x).toreal)) μ theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.half_commonDensityAffinity_sq_le_overlap","label":"half_commonDensityAffinity_sq_le_overlap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.half_commonDensityAffinity_sq_le_overlap","description":"Eq. (14.9), including integrable densities with zeros.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-3a69a4a14dd9","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5327,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem half_commonDensityAffinity_sq_le_overlap {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) : (1 / 2 : ℝ) * commonDensityAffinity P Q μ ^ 2 ≤ commonDensityOverlap P Q μ","missing":[],"search":"half_commondensityaffinity_sq_le_overlap banditrlproof.lowerbounds.half_commondensityaffinity_sq_le_overlap eq. (14.9), including integrable densities with zeros. theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.commonDensityComparisonEvent","label":"commonDensityComparisonEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.commonDensityComparisonEvent","description":"The event selecting the smaller source density.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-94942bcafecd","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5328,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def commonDensityComparisonEvent {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : Set α","missing":[],"search":"commondensitycomparisonevent banditrlproof.lowerbounds.commondensitycomparisonevent the event selecting the smaller source density. definition compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurableSet_commonDensityComparisonEvent","label":"measurableSet_commonDensityComparisonEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurableSet_commonDensityComparisonEvent","description":"theorem measurableSet_commonDensityComparisonEvent {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : MeasurableSet (commonDensityComparisonEvent P Q μ)","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-e5494d240250","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5329,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_commonDensityComparisonEvent {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : MeasurableSet (commonDensityComparisonEvent P Q μ)","missing":[],"search":"measurableset_commondensitycomparisonevent banditrlproof.lowerbounds.measurableset_commondensitycomparisonevent theorem measurableset_commondensitycomparisonevent {α : type*} [measurablespace α] (p q μ : measure α) : measurableset (commondensitycomparisonevent p q μ) theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_min_commonDensity","label":"integrable_min_commonDensity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_min_commonDensity","description":"theorem integrable_min_commonDensity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : Integrable (fun x => min (P.rnDeriv μ x).toReal (Q.rnDeriv μ x).toReal) μ","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-9027872a7f2e","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5330,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_min_commonDensity {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : Integrable (fun x => min (P.rnDeriv μ x).toReal (Q.rnDeriv μ x).toReal) μ","missing":[],"search":"integrable_min_commondensity banditrlproof.lowerbounds.integrable_min_commondensity theorem integrable_min_commondensity {α : type*} [measurablespace α] (p q μ : measure α) [isfinitemeasure p] [isfinitemeasure q] : integrable (fun x => min (p.rnderiv μ x).toreal (q.rnderiv μ x).toreal) μ theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_nonneg","label":"commonDensityOverlap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.commonDensityOverlap_nonneg","description":"theorem commonDensityOverlap_nonneg {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : 0 ≤ commonDensityOverlap P Q μ","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-9a205f89b4cb","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5331,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:124"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem commonDensityOverlap_nonneg {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) : 0 ≤ commonDensityOverlap P Q μ","missing":[],"search":"commondensityoverlap_nonneg banditrlproof.lowerbounds.commondensityoverlap_nonneg theorem commondensityoverlap_nonneg {α : type*} [measurablespace α] (p q μ : measure α) : 0 ≤ commondensityoverlap p q μ theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_eq_testingError","label":"commonDensityOverlap_eq_testingError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.commonDensityOverlap_eq_testingError","description":"The likelihood comparison event attains the density-overlap testing error.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-2842bfe051a8","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5332,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem commonDensityOverlap_eq_testingError {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) : commonDensityOverlap P Q μ = P.real (commonDensityComparisonEvent P Q μ) + Q.real (commonDensityComparisonEvent P Q μ)ᶜ","missing":[],"search":"commondensityoverlap_eq_testingerror banditrlproof.lowerbounds.commondensityoverlap_eq_testingerror the likelihood comparison event attains the density-overlap testing error. theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_le_testingError","label":"commonDensityOverlap_le_testingError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.commonDensityOverlap_le_testingError","description":"The overlap is no larger than the testing error of any measurable event.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-8f6252dd49c1","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5333,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:155"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem commonDensityOverlap_le_testingError {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) {A : Set α} (hA : MeasurableSet A) : commonDensityOverlap P Q μ ≤ P.real A + Q.real Aᶜ","missing":[],"search":"commondensityoverlap_le_testingerror banditrlproof.lowerbounds.commondensityoverlap_le_testingerror the overlap is no larger than the testing error of any measurable event. theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_le_commonDensityOverlap","label":"bretagnolleHuberScale_le_commonDensityOverlap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale_le_commonDensityOverlap","description":"Eq. (14.8): the measure overlap is bounded below by the BH exponential scale.","url":"../modules/banditrlproof-lowerbounds-commondensityoverlap/index.html#decl-31eb011a8273","parent":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","order":5334,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDensityOverlap"],["Source","BanditRLProof/LowerBounds/CommonDensityOverlap.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuberScale_le_commonDensityOverlap {α : Type*} [MeasurableSpace α] (P Q μ : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] [SigmaFinite μ] (hP : P ≪ μ) (hQ : Q ≪ μ) : bretagnolleHuberScale (relativeEntropy P Q) ≤ commonDensityOverlap P Q μ","missing":[],"search":"bretagnollehuberscale_le_commondensityoverlap banditrlproof.lowerbounds.bretagnollehuberscale_le_commondensityoverlap eq. (14.8): the measure overlap is bounded below by the bh exponential scale. theorem compiled","shard":"modules/a9cb2127ecf6e8de.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_commonFiniteDominatingMeasure","label":"exists_commonFiniteDominatingMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_commonFiniteDominatingMeasure","description":"The sum of two finite laws supplies the common dominating measure in the source.","url":"../modules/banditrlproof-lowerbounds-commondomination/index.html#decl-d991db70947e","parent":"module:BanditRLProof.LowerBounds.CommonDomination","order":5335,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDomination"],["Source","BanditRLProof/LowerBounds/CommonDomination.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_commonFiniteDominatingMeasure {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : ∃ μ : Measure α, IsFiniteMeasure μ ∧ P ≪ μ ∧ Q ≪ μ","missing":[],"search":"exists_commonfinitedominatingmeasure banditrlproof.lowerbounds.exists_commonfinitedominatingmeasure the sum of two finite laws supplies the common dominating measure in the source. theorem compiled","shard":"modules/fd48243cca1b4852.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_commonSigmaFiniteDominatingMeasure","label":"exists_commonSigmaFiniteDominatingMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_commonSigmaFiniteDominatingMeasure","description":"theorem exists_commonSigmaFiniteDominatingMeasure {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : ∃ μ : Measure α, SigmaFinite μ ∧ P ≪ μ ∧ Q ≪ μ","url":"../modules/banditrlproof-lowerbounds-commondomination/index.html#decl-2fdf1a0452f4","parent":"module:BanditRLProof.LowerBounds.CommonDomination","order":5336,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDomination"],["Source","BanditRLProof/LowerBounds/CommonDomination.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem exists_commonSigmaFiniteDominatingMeasure {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : ∃ μ : Measure α, SigmaFinite μ ∧ P ≪ μ ∧ Q ≪ μ","missing":[],"search":"exists_commonsigmafinitedominatingmeasure banditrlproof.lowerbounds.exists_commonsigmafinitedominatingmeasure theorem exists_commonsigmafinitedominatingmeasure {α : type*} [measurablespace α] (p q : measure α) [isfinitemeasure p] [isfinitemeasure q] : ∃ μ : measure α, sigmafinite μ ∧ p ≪ μ ∧ q ≪ μ theorem compiled","shard":"modules/fd48243cca1b4852.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_lt_top_iff_ac","label":"relativeEntropy_finite_lt_top_iff_ac","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_lt_top_iff_ac","description":"Absolute continuity suffices for finite KL on a finite alphabet, not in general.","url":"../modules/banditrlproof-lowerbounds-commondomination/index.html#decl-222194b5978a","parent":"module:BanditRLProof.LowerBounds.CommonDomination","order":5337,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CommonDomination"],["Source","BanditRLProof/LowerBounds/CommonDomination.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_lt_top_iff_ac {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] : relativeEntropy P Q < ⊤ ↔ P ≪ Q","missing":[],"search":"relativeentropy_finite_lt_top_iff_ac banditrlproof.lowerbounds.relativeentropy_finite_lt_top_iff_ac absolute continuity suffices for finite kl on a finite alphabet, not in general. theorem compiled","shard":"modules/fd48243cca1b4852.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_kernelRN_ae","label":"klDiv_compProd_same_left_eq_lintegral_kernelRN_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_kernelRN_ae","description":"The conditional KL term for two composition products with one common left measure is the iterated integral of `klFun` applied to the measurable kernel Radon--Nikodym derivative. This is the measurable formulation of the conditional-KL integral needed by the Chapter 15 history chain rule. Pointwise absolute continuity is explicit; without it, individual conditional KL terms may be infinite because of support mismatch.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-561c87f0592f","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5338,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_compProd_same_left_eq_lintegral_kernelRN_ae {Base Target : Type*} [MeasurableSpace Base] [MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Base) [IsFiniteMeasure mu] (kappa eta : Kernel Base Target) [IsMarkovKernel kappa] [IsMarkovKernel eta] (h_ac : ∀ᵐ base ∂mu, kappa base ≪ eta base) : InformationTheory.klDiv (mu ⊗ₘ kappa) (mu ⊗ₘ eta) = ∫⁻ x, ∫⁻ y, ENNReal.ofReal (InformationTheory.klFun ((kappa.rnDeriv eta x y).toReal)) ∂eta x ∂mu","missing":[],"search":"kldiv_compprod_same_left_eq_lintegral_kernelrn_ae banditrlproof.lowerbounds.kldiv_compprod_same_left_eq_lintegral_kernelrn_ae the conditional kl term for two composition products with one common left measure is the iterated integral of `klfun` applied to the measurable kernel radon--nikodym derivative. this is the measurable formulation of the conditional-kl integral needed by the chapter 15 history chain rule. pointwise absolute continuity is explicit; without it, individual conditional kl terms may be infinite because of support mismatch. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_kernelRN","label":"klDiv_compProd_same_left_eq_lintegral_kernelRN","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_kernelRN","description":"Pointwise absolute continuity is a convenient sufficient specialization.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-3ab81498e944","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5339,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_compProd_same_left_eq_lintegral_kernelRN {Base Target : Type*} [MeasurableSpace Base] [MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Base) [IsFiniteMeasure mu] (kappa eta : Kernel Base Target) [IsMarkovKernel kappa] [IsMarkovKernel eta] (h_ac : forall base, kappa base ≪ eta base) : InformationTheory.klDiv (mu ⊗ₘ kappa) (mu ⊗ₘ eta) = ∫⁻ x, ∫⁻ y, ENNReal.ofReal (InformationTheory.klFun ((kappa.rnDeriv eta x y).toReal)) ∂eta x ∂mu","missing":[],"search":"kldiv_compprod_same_left_eq_lintegral_kernelrn banditrlproof.lowerbounds.kldiv_compprod_same_left_eq_lintegral_kernelrn pointwise absolute continuity is a convenient sufficient specialization. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv","label":"klDiv_compProd_same_left_eq_lintegral_klDiv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv","description":"Under pointwise absolute continuity, the measurable kernel-RN formulation is equal to the familiar conditional-KL integral. This closes Mathlib's stated `klDiv_compProd_eq_add` TODO for the countably generated target-space branch.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-2c43ba877ffb","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5340,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:120"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_compProd_same_left_eq_lintegral_klDiv {Base Target : Type*} [MeasurableSpace Base] [MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Base) [IsFiniteMeasure mu] (kappa eta : Kernel Base Target) [IsMarkovKernel kappa] [IsMarkovKernel eta] (h_ac : forall base, kappa base ≪ eta base) : InformationTheory.klDiv (mu ⊗ₘ kappa) (mu ⊗ₘ eta) = ∫⁻ base, InformationTheory.klDiv (kappa base) (eta base) ∂mu","missing":[],"search":"kldiv_compprod_same_left_eq_lintegral_kldiv banditrlproof.lowerbounds.kldiv_compprod_same_left_eq_lintegral_kldiv under pointwise absolute continuity, the measurable kernel-rn formulation is equal to the familiar conditional-kl integral. this closes mathlib's stated `kldiv_compprod_eq_add` todo for the countably generated target-space branch. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_ae","label":"klDiv_compProd_same_left_eq_lintegral_klDiv_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_ae","description":"Conditional-KL integral under almost-everywhere conditional absolute continuity.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-7c987eb5392e","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5341,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:155"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_compProd_same_left_eq_lintegral_klDiv_ae {Base Target : Type*} [MeasurableSpace Base] [MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Base) [IsFiniteMeasure mu] (kappa eta : Kernel Base Target) [IsMarkovKernel kappa] [IsMarkovKernel eta] (h_ac : ∀ᵐ base ∂mu, kappa base ≪ eta base) : InformationTheory.klDiv (mu ⊗ₘ kappa) (mu ⊗ₘ eta) = ∫⁻ base, InformationTheory.klDiv (kappa base) (eta base) ∂mu","missing":[],"search":"kldiv_compprod_same_left_eq_lintegral_kldiv_ae banditrlproof.lowerbounds.kldiv_compprod_same_left_eq_lintegral_kldiv_ae conditional-kl integral under almost-everywhere conditional absolute continuity. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","label":"klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","description":"If the conditional KL function is measurable, the same-left composition-product identity holds without an absolute-continuity hypothesis. Singular conditional fibres correctly contribute `∞` on sets of positive base measure.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-a55eb6dae7d3","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5342,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:195"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable {Base Target : Type*} [MeasurableSpace Base] [MeasurableSpace Target] [MeasurableSpace.CountablyGenerated Target] (mu : Measure Base) [IsFiniteMeasure mu] (kappa eta : Kernel Base Target) [IsMarkovKernel kappa] [IsMarkovKernel eta] (hmeasurable : Measurable (fun base => InformationTheory.klDiv (kappa base) (eta base))) : InformationTheory.klDiv (mu ⊗ₘ kappa) (mu ⊗ₘ eta) = ∫⁻ base, InformationTheory.klDiv (kappa base) (eta base) ∂mu","missing":[],"search":"kldiv_compprod_same_left_eq_lintegral_kldiv_of_measurable banditrlproof.lowerbounds.kldiv_compprod_same_left_eq_lintegral_kldiv_of_measurable if the conditional kl function is measurable, the same-left composition-product identity holds without an absolute-continuity hypothesis. singular conditional fibres correctly contribute `∞` on sets of positive base measure. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_map_measurableEquiv","label":"klDiv_map_measurableEquiv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_map_measurableEquiv","description":"Relative entropy is invariant under a measurable equivalence.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-f461297677ca","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5343,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:227"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_map_measurableEquiv {Source Target : Type*} [MeasurableSpace Source] [MeasurableSpace Target] (mu nu : Measure Source) [IsFiniteMeasure mu] [IsFiniteMeasure nu] (equiv : Source ≃ᵐ Target) : InformationTheory.klDiv (mu.map equiv) (nu.map equiv) = InformationTheory.klDiv mu nu","missing":[],"search":"kldiv_map_measurableequiv banditrlproof.lowerbounds.kldiv_map_measurableequiv relative entropy is invariant under a measurable equivalence. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL","label":"klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL","description":"For one adaptive bandit round, a common randomized policy contributes no KL cost of its own. The conditional history-extension cost is the first-law policy average of the selected arm reward-law KL divergences. The policy is an arbitrary Markov kernel from the complete visible history to the action space; in particular, this statement is not restricted to a deterministic action rule.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-2c18d439f847","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5344,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:290"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL {History Action Reward : Type*} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSpace.CountablyGenerated Action] [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (historyLaw : Measure History) [IsFiniteMeasure historyLaw] (policy : Kernel History Action) [IsMarkovKernel policy] (armLaw referenceArmLaw : Kernel Action Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (h_ac : forall arm, armLaw arm ≪ referenceArmLaw arm) : InformationTheory.klDiv (historyLaw ⊗ₘ (policy ⊗ₖ armLaw.comap Prod.snd measurable_snd)) (historyLaw ⊗ₘ (policy ⊗ₖ referenceArmLaw.comap Prod.snd measurable_snd)) = ∫⁻ history, ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂policy history ∂historyLaw","missing":[],"search":"kldiv_historystep_samepolicy_eq_iterated_lintegral_armkl banditrlproof.lowerbounds.kldiv_historystep_samepolicy_eq_iterated_lintegral_armkl for one adaptive bandit round, a common randomized policy contributes no kl cost of its own. the conditional history-extension cost is the first-law policy average of the selected arm reward-law kl divergences. the policy is an arbitrary markov kernel from the complete visible history to the action space; in particular, this statement is not restricted to a deterministic action rule. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_finiteAction_compProd_eq_lintegral_armKL","label":"klDiv_finiteAction_compProd_eq_lintegral_armKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_finiteAction_compProd_eq_lintegral_armKL","description":"A finite action law admits the conditional-KL identity without an AC hypothesis.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-cd8653a66838","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5345,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:341"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_finiteAction_compProd_eq_lintegral_armKL {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (actionLaw : Measure (Fin K)) [IsFiniteMeasure actionLaw] (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] : InformationTheory.klDiv (actionLaw ⊗ₘ armLaw) (actionLaw ⊗ₘ referenceArmLaw) = ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂actionLaw","missing":[],"search":"kldiv_finiteaction_compprod_eq_lintegral_armkl banditrlproof.lowerbounds.kldiv_finiteaction_compprod_eq_lintegral_armkl a finite action law admits the conditional-kl identity without an ac hypothesis. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_finiteAction_compProd_eq_sum_mass_mul_armKL","label":"klDiv_finiteAction_compProd_eq_sum_mass_mul_armKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_finiteAction_compProd_eq_sum_mass_mul_armKL","description":"Finite-action conditional KL written as a weighted finite sum.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-ba01fe2a29c2","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5346,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:356"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_finiteAction_compProd_eq_sum_mass_mul_armKL {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (actionLaw : Measure (Fin K)) [IsFiniteMeasure actionLaw] (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] : InformationTheory.klDiv (actionLaw ⊗ₘ armLaw) (actionLaw ⊗ₘ referenceArmLaw) = ∑ arm : Fin K, actionLaw ({arm} : Set (Fin K)) * InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm)","missing":[],"search":"kldiv_finiteaction_compprod_eq_sum_mass_mul_armkl banditrlproof.lowerbounds.kldiv_finiteaction_compprod_eq_sum_mass_mul_armkl finite-action conditional kl written as a weighted finite sum. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","label":"klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","description":"General finite-action same-policy one-round KL identity. Unlike the AC specialization above, this theorem also covers infinite arm divergences and zero-probability singular arms, with the standard `ENNReal` convention `0 * ∞ = 0`.","url":"../modules/banditrlproof-lowerbounds-conditionalkernelkl/index.html#decl-270b7863ad77","parent":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","order":5347,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ConditionalKernelKL"],["Source","BanditRLProof/LowerBounds/ConditionalKernelKL.lean:379"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general {History Reward : Type*} {K : Nat} [MeasurableSpace History] [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (historyLaw : Measure History) [IsFiniteMeasure historyLaw] (policy : Kernel History (Fin K)) [IsMarkovKernel policy] (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] : InformationTheory.klDiv (historyLaw ⊗ₘ (policy ⊗ₖ armLaw.comap Prod.snd measurable_snd)) (historyLaw ⊗ₘ (policy ⊗ₖ referenceArmLaw.comap Prod.snd measurable_snd)) = ∫⁻ history, ∫⁻ arm, InformationTheory.klDiv (armLaw arm) (referenceArmLaw arm) ∂policy history ∂historyLaw","missing":[],"search":"kldiv_historystep_samepolicy_eq_iterated_lintegral_armkl_general banditrlproof.lowerbounds.kldiv_historystep_samepolicy_eq_iterated_lintegral_armkl_general general finite-action same-policy one-round kl identity. unlike the ac specialization above, this theorem also covers infinite arm divergences and zero-probability singular arms, with the standard `ennreal` convention `0 * ∞ = 0`. theorem compiled","shard":"modules/d9b7b3d2d0ef672c.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteCrossEntropy","label":"discreteCrossEntropy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteCrossEntropy","description":"noncomputable def discreteCrossEntropy {α : Type*} [Fintype α] (p q : α → ℝ) : ℝ","url":"../modules/banditrlproof-lowerbounds-crossentropy/index.html#decl-9c06f0376912","parent":"module:BanditRLProof.LowerBounds.CrossEntropy","order":5348,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.CrossEntropy"],["Source","BanditRLProof/LowerBounds/CrossEntropy.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def discreteCrossEntropy {α : Type*} [Fintype α] (p q : α → ℝ) : ℝ","missing":[],"search":"discretecrossentropy banditrlproof.lowerbounds.discretecrossentropy noncomputable def discretecrossentropy {α : type*} [fintype α] (p q : α → ℝ) : ℝ definition compiled","shard":"modules/0bfae14292396ff9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteCrossEntropy_sub_entropy","label":"discreteCrossEntropy_sub_entropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteCrossEntropy_sub_entropy","description":"The unrounded information-cost interpretation of Eq. (14.4).","url":"../modules/banditrlproof-lowerbounds-crossentropy/index.html#decl-03417d561c85","parent":"module:BanditRLProof.LowerBounds.CrossEntropy","order":5349,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CrossEntropy"],["Source","BanditRLProof/LowerBounds/CrossEntropy.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteCrossEntropy_sub_entropy {α : Type*} [Fintype α] (p q : α → ℝ) (hsupport : ∀ i, p i ≠ 0 → q i ≠ 0) : discreteCrossEntropy p q - discreteEntropy Finset.univ p = ∑ i, p i * Real.log (p i / q i)","missing":[],"search":"discretecrossentropy_sub_entropy banditrlproof.lowerbounds.discretecrossentropy_sub_entropy the unrounded information-cost interpretation of eq. (14.4). theorem compiled","shard":"modules/0bfae14292396ff9.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_crossEntropy","label":"relativeEntropy_finite_crossEntropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_crossEntropy","description":"theorem relativeEntropy_finite_crossEntropy {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] (h : P ≪ Q) : relativeEntropy P Q = ENNReal.ofReal (discreteCrossEntropy (fun i => (P {i}).toReal) (fun i => (Q {i}).toReal) - discreteEntropy Finset.univ (fun i => (P {i}).toReal))","url":"../modules/banditrlproof-lowerbounds-crossentropy/index.html#decl-d7bfdc48bfdf","parent":"module:BanditRLProof.LowerBounds.CrossEntropy","order":5350,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CrossEntropy"],["Source","BanditRLProof/LowerBounds/CrossEntropy.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_crossEntropy {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] (h : P ≪ Q) : relativeEntropy P Q = ENNReal.ofReal (discreteCrossEntropy (fun i => (P {i}).toReal) (fun i => (Q {i}).toReal) - discreteEntropy Finset.univ (fun i => (P {i}).toReal))","missing":[],"search":"relativeentropy_finite_crossentropy banditrlproof.lowerbounds.relativeentropy_finite_crossentropy theorem relativeentropy_finite_crossentropy {α : type*} [fintype α] [measurablespace α] [measurablesingletonclass α] (p q : measure α) [isprobabilitymeasure p] [isprobabilitymeasure q] (h : p ≪ q) : relativeentropy p q = ennreal.ofreal (discretecrossentropy (fun i => (p {i}).toreal) (fun i => (q {i}).toreal) - discreteentropy finset.univ (fun i => (p {i}).toreal)) theorem compiled","shard":"modules/0bfae14292396ff9.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.entropyTerm_tendsto_zero_right","label":"entropyTerm_tendsto_zero_right","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.entropyTerm_tendsto_zero_right","description":"The source's zero-mass entropy convention agrees with the right-hand limit.","url":"../modules/banditrlproof-lowerbounds-crossentropy/index.html#decl-78478b47b5be","parent":"module:BanditRLProof.LowerBounds.CrossEntropy","order":5351,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.CrossEntropy"],["Source","BanditRLProof/LowerBounds/CrossEntropy.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem entropyTerm_tendsto_zero_right : Filter.Tendsto (fun x : ℝ => x * Real.log x⁻¹) (nhdsWithin 0 (Set.Ioi 0)) (nhds 0)","missing":[],"search":"entropyterm_tendsto_zero_right banditrlproof.lowerbounds.entropyterm_tendsto_zero_right the source's zero-mass entropy convention agrees with the right-hand limit. theorem compiled","shard":"modules/0bfae14292396ff9.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryAddressValue","label":"binaryAddressValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryAddressValue","description":"Big-endian binary address, with leading zeroes retained by the word length.","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-a1006dabe5d7","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5352,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def binaryAddressValue : List Bool → ℕ | [] => 0 | b :: w => (if b then 2 ^ w.length else 0) + binaryAddressValue w theorem binaryAddressValue_lt (w : List Bool) : binaryAddressValue w < 2 ^ w.length","missing":[],"search":"binaryaddressvalue banditrlproof.lowerbounds.binaryaddressvalue big-endian binary address, with leading zeroes retained by the word length. definition compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryAddressValue_lt","label":"binaryAddressValue_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryAddressValue_lt","description":"theorem binaryAddressValue_lt (w : List Bool) : binaryAddressValue w < 2 ^ w.length","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-7a9e68dcae11","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5353,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem binaryAddressValue_lt (w : List Bool) : binaryAddressValue w < 2 ^ w.length","missing":[],"search":"binaryaddressvalue_lt banditrlproof.lowerbounds.binaryaddressvalue_lt theorem binaryaddressvalue_lt (w : list bool) : binaryaddressvalue w < 2 ^ w.length theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binaryAddress","label":"exists_binaryAddress","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binaryAddress","description":"Every dyadic cell index has a binary address of the specified length.","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-e4db77087476","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5354,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binaryAddress (n m : ℕ) (hm : m < 2 ^ n) : ∃ w : List Bool, w.length = n ∧ binaryAddressValue w = m","missing":[],"search":"exists_binaryaddress banditrlproof.lowerbounds.exists_binaryaddress every dyadic cell index has a binary address of the specified length. theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryAddressValue_append","label":"binaryAddressValue_append","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryAddressValue_append","description":"theorem binaryAddressValue_append (u v : List Bool) : binaryAddressValue (u ++ v) = binaryAddressValue u * 2 ^ v.length + binaryAddressValue v","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-a4965a0d647d","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5355,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem binaryAddressValue_append (u v : List Bool) : binaryAddressValue (u ++ v) = binaryAddressValue u * 2 ^ v.length + binaryAddressValue v","missing":[],"search":"binaryaddressvalue_append banditrlproof.lowerbounds.binaryaddressvalue_append theorem binaryaddressvalue_append (u v : list bool) : binaryaddressvalue (u ++ v) = binaryaddressvalue u * 2 ^ v.length + binaryaddressvalue v theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.dyadicAddressLower","label":"dyadicAddressLower","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.dyadicAddressLower","description":"noncomputable def dyadicAddressLower (w : List Bool) : ℝ","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-4a3f38f01670","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5356,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def dyadicAddressLower (w : List Bool) : ℝ","missing":[],"search":"dyadicaddresslower banditrlproof.lowerbounds.dyadicaddresslower noncomputable def dyadicaddresslower (w : list bool) : ℝ definition compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.dyadicAddressUpper","label":"dyadicAddressUpper","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.dyadicAddressUpper","description":"noncomputable def dyadicAddressUpper (w : List Bool) : ℝ","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-b9c1c4156eb3","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5357,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def dyadicAddressUpper (w : List Bool) : ℝ","missing":[],"search":"dyadicaddressupper banditrlproof.lowerbounds.dyadicaddressupper noncomputable def dyadicaddressupper (w : list bool) : ℝ definition compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.dyadicAddress_width","label":"dyadicAddress_width","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.dyadicAddress_width","description":"theorem dyadicAddress_width (w : List Bool) : dyadicAddressUpper w - dyadicAddressLower w = 1 / (2 : ℝ) ^ w.length","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-9ac24e7d4099","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5358,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dyadicAddress_width (w : List Bool) : dyadicAddressUpper w - dyadicAddressLower w = 1 / (2 : ℝ) ^ w.length","missing":[],"search":"dyadicaddress_width banditrlproof.lowerbounds.dyadicaddress_width theorem dyadicaddress_width (w : list bool) : dyadicaddressupper w - dyadicaddresslower w = 1 / (2 : ℝ) ^ w.length theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.dyadicAddress_nonempty","label":"dyadicAddress_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.dyadicAddress_nonempty","description":"theorem dyadicAddress_nonempty (w : List Bool) : dyadicAddressLower w < dyadicAddressUpper w","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-021eeea51ae8","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5359,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dyadicAddress_nonempty (w : List Bool) : dyadicAddressLower w < dyadicAddressUpper w","missing":[],"search":"dyadicaddress_nonempty banditrlproof.lowerbounds.dyadicaddress_nonempty theorem dyadicaddress_nonempty (w : list bool) : dyadicaddresslower w < dyadicaddressupper w theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.dyadicAddress_append_contained","label":"dyadicAddress_append_contained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.dyadicAddress_append_contained","description":"theorem dyadicAddress_append_contained (u v : List Bool) : dyadicAddressLower u ≤ dyadicAddressLower (u ++ v) ∧ dyadicAddressUpper (u ++ v) ≤ dyadicAddressUpper u","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-5bda8f8f271c","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5360,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dyadicAddress_append_contained (u v : List Bool) : dyadicAddressLower u ≤ dyadicAddressLower (u ++ v) ∧ dyadicAddressUpper (u ++ v) ≤ dyadicAddressUpper u","missing":[],"search":"dyadicaddress_append_contained banditrlproof.lowerbounds.dyadicaddress_append_contained theorem dyadicaddress_append_contained (u v : list bool) : dyadicaddresslower u ≤ dyadicaddresslower (u ++ v) ∧ dyadicaddressupper (u ++ v) ≤ dyadicaddressupper u theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.dyadicAddress_prefix_contained","label":"dyadicAddress_prefix_contained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.dyadicAddress_prefix_contained","description":"theorem dyadicAddress_prefix_contained (u v : List Bool) (h : u <+: v) : dyadicAddressLower u ≤ dyadicAddressLower v ∧ dyadicAddressUpper v ≤ dyadicAddressUpper u","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-26063eaffbff","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5361,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem dyadicAddress_prefix_contained (u v : List Bool) (h : u <+: v) : dyadicAddressLower u ≤ dyadicAddressLower v ∧ dyadicAddressUpper v ≤ dyadicAddressUpper u","missing":[],"search":"dyadicaddress_prefix_contained banditrlproof.lowerbounds.dyadicaddress_prefix_contained theorem dyadicaddress_prefix_contained (u v : list bool) (h : u <+: v) : dyadicaddresslower u ≤ dyadicaddresslower v ∧ dyadicaddressupper v ≤ dyadicaddressupper u theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_dyadicAddress_inside","label":"exists_dyadicAddress_inside","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_dyadicAddress_inside","description":"Select an actual binary word whose dyadic cell fits inside the given interval.","url":"../modules/banditrlproof-lowerbounds-dyadicaddresses/index.html#decl-0c8c26820e8c","parent":"module:BanditRLProof.LowerBounds.DyadicAddresses","order":5362,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.DyadicAddresses"],["Source","BanditRLProof/LowerBounds/DyadicAddresses.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_dyadicAddress_inside (L U : ℝ) (n : ℕ) (hL : 0 ≤ L) (hU : U ≤ 1) (hwidth : 2 * (1 / (2 : ℝ) ^ n) ≤ U - L) : ∃ w : List Bool, w.length = n ∧ L ≤ dyadicAddressLower w ∧ dyadicAddressUpper w < U","missing":[],"search":"exists_dyadicaddress_inside banditrlproof.lowerbounds.exists_dyadicaddress_inside select an actual binary word whose dyadic cell fits inside the given interval. theorem compiled","shard":"modules/8f71f7f78637e4d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.absolutelyContinuous_iff_atom_support","label":"absolutelyContinuous_iff_atom_support","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.absolutelyContinuous_iff_atom_support","description":"On a finite alphabet, absolute continuity is exactly atomwise support inclusion.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-a29df3fff947","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5363,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem absolutelyContinuous_iff_atom_support {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) : P ≪ Q ↔ ∀ x, Q {x} = 0 → P {x} = 0","missing":[],"search":"absolutelycontinuous_iff_atom_support banditrlproof.lowerbounds.absolutelycontinuous_iff_atom_support on a finite alphabet, absolute continuity is exactly atomwise support inclusion. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.rnDeriv_mul_atom","label":"rnDeriv_mul_atom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.rnDeriv_mul_atom","description":"Atomwise density identity, retaining zero-mass atoms.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-2a3c6ac0f5fb","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5364,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rnDeriv_mul_atom {α : Type*} [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) (x : α) : P.rnDeriv Q x * Q {x} = P {x}","missing":[],"search":"rnderiv_mul_atom banditrlproof.lowerbounds.rnderiv_mul_atom atomwise density identity, retaining zero-mass atoms. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.rnDeriv_atom_eq_div","label":"rnDeriv_atom_eq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.rnDeriv_atom_eq_div","description":"On a positive reference atom, the RN density is the atom-mass ratio.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-81ae0cc481ae","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5365,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem rnDeriv_atom_eq_div {α : Type*} [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) (x : α) (hq : Q {x} ≠ 0) : P.rnDeriv Q x = P {x} / Q {x}","missing":[],"search":"rnderiv_atom_eq_div banditrlproof.lowerbounds.rnderiv_atom_eq_div on a positive reference atom, the rn density is the atom-mass ratio. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_atom_support_mismatch","label":"relativeEntropy_eq_top_of_atom_support_mismatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_atom_support_mismatch","description":"A positive source atom absent from the reference law forces infinite KL.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-88df91341eea","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5366,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_eq_top_of_atom_support_mismatch {α : Type*} [MeasurableSpace α] (P Q : Measure α) (x : α) (hp : P {x} ≠ 0) (hq : Q {x} = 0) : relativeEntropy P Q = ∞","missing":[],"search":"relativeentropy_eq_top_of_atom_support_mismatch banditrlproof.lowerbounds.relativeentropy_eq_top_of_atom_support_mismatch a positive source atom absent from the reference law forces infinite kl. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_klFun","label":"relativeEntropy_finite_klFun","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_klFun","description":"Finite-alphabet KL in its nonnegative convex-integrand form.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-96da648a0248","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5367,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_klFun {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) : relativeEntropy P Q = ∑ x, ENNReal.ofReal (InformationTheory.klFun ((P {x} / Q {x}).toReal)) * Q {x}","missing":[],"search":"relativeentropy_finite_klfun banditrlproof.lowerbounds.relativeentropy_finite_klfun finite-alphabet kl in its nonnegative convex-integrand form. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_sum_log","label":"relativeEntropy_finite_sum_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_sum_log","description":"Textbook Eq. (14.4) on any finite alphabet in the supported branch.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-78d5bc4d7825","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5368,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_sum_log {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] (h : P ≪ Q) : relativeEntropy P Q = ENNReal.ofReal (∑ x, (P {x}).toReal * Real.log ((P {x}).toReal / (Q {x}).toReal))","missing":[],"search":"relativeentropy_finite_sum_log banditrlproof.lowerbounds.relativeentropy_finite_sum_log textbook eq. (14.4) on any finite alphabet in the supported branch. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_if","label":"relativeEntropy_finite_eq_if","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_eq_if","description":"Eq. (14.4), including the infinite branch when atomwise support fails.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-5ed14a9f76ad","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5369,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_eq_if {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] : relativeEntropy P Q = if ∀ x, Q {x} = 0 → P {x} = 0 then ENNReal.ofReal (∑ x, (P {x}).toReal * Real.log ((P {x}).toReal / (Q {x}).toReal)) else ∞","missing":[],"search":"relativeentropy_finite_eq_if banditrlproof.lowerbounds.relativeentropy_finite_eq_if eq. (14.4), including the infinite branch when atomwise support fails. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_top_iff","label":"relativeEntropy_finite_eq_top_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_eq_top_iff","description":"Finite alphabets have infinite KL exactly at a support mismatch.","url":"../modules/banditrlproof-lowerbounds-finitediscretekl/index.html#decl-7f4f5a42d8c7","parent":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","order":5370,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FiniteDiscreteKL"],["Source","BanditRLProof/LowerBounds/FiniteDiscreteKL.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_eq_top_iff {α : Type*} [Fintype α] [MeasurableSpace α] [MeasurableSingletonClass α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] : relativeEntropy P Q = ∞ ↔ ∃ x, P {x} ≠ 0 ∧ Q {x} = 0","missing":[],"search":"relativeentropy_finite_eq_top_iff banditrlproof.lowerbounds.relativeentropy_finite_eq_top_iff finite alphabets have infinite kl exactly at a support mismatch. theorem compiled","shard":"modules/724de932ea7680fc.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.totalMass_klFun_le_relativeEntropy","label":"totalMass_klFun_le_relativeEntropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.totalMass_klFun_le_relativeEntropy","description":"Convexity bounds the divergence of total masses by finite-measure KL.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-a7d4e89e6bfc","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5371,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem totalMass_klFun_le_relativeEntropy {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : ENNReal.ofReal (InformationTheory.klFun ((P univ / Q univ).toReal)) * Q univ ≤ relativeEntropy P Q","missing":[],"search":"totalmass_klfun_le_relativeentropy banditrlproof.lowerbounds.totalmass_klfun_le_relativeentropy convexity bounds the divergence of total masses by finite-measure kl. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_relativeEntropy_restrict_fibers","label":"sum_relativeEntropy_restrict_fibers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_relativeEntropy_restrict_fibers","description":"KL splits into the restrictions to all cells of a finite observation.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-e0aaf56f7606","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5372,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_relativeEntropy_restrict_fibers {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) {n : ℕ} (f : α → Fin n) (hf : Measurable f) : (∑ i, relativeEntropy (P.restrict (f ⁻¹' {i})) (Q.restrict (f ⁻¹' {i}))) = relativeEntropy P Q","missing":[],"search":"sum_relativeentropy_restrict_fibers banditrlproof.lowerbounds.sum_relativeentropy_restrict_fibers kl splits into the restrictions to all cells of a finite observation. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_map_le","label":"relativeEntropy_finite_map_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_map_le","description":"Finite-valued measurable observations cannot increase KL.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-f811da38a65c","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5373,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_map_le {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] {n : ℕ} (f : α → Fin n) (hf : Measurable f) : relativeEntropy (P.map f) (Q.map f) ≤ relativeEntropy P Q","missing":[],"search":"relativeentropy_finite_map_le banditrlproof.lowerbounds.relativeentropy_finite_map_le finite-valued measurable observations cannot increase kl. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy","label":"finitePartitionRelativeEntropy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finitePartitionRelativeEntropy","description":"Relative entropy defined by the supremum over finite measurable observations.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-3554f8bc7192","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5374,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:84"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"def finitePartitionRelativeEntropy {α : Type*} [MeasurableSpace α] (P Q : Measure α) : ENNReal","missing":[],"search":"finitepartitionrelativeentropy banditrlproof.lowerbounds.finitepartitionrelativeentropy relative entropy defined by the supremum over finite measurable observations. definition compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_le_relativeEntropy","label":"finitePartitionRelativeEntropy_le_relativeEntropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_le_relativeEntropy","description":"The finite-discretisation supremum never exceeds RN relative entropy.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-29608e7e2430","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5375,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:90"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finitePartitionRelativeEntropy_le_relativeEntropy {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : finitePartitionRelativeEntropy P Q ≤ relativeEntropy P Q","missing":[],"search":"finitepartitionrelativeentropy_le_relativeentropy banditrlproof.lowerbounds.finitepartitionrelativeentropy_le_relativeentropy the finite-discretisation supremum never exceeds rn relative entropy. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_map_le_finitePartitionRelativeEntropy","label":"relativeEntropy_map_le_finitePartitionRelativeEntropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_map_le_finitePartitionRelativeEntropy","description":"Every finite measurable observation is included in the defining supremum.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-f9b98f5f530f","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5376,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_map_le_finitePartitionRelativeEntropy {α : Type*} [MeasurableSpace α] (P Q : Measure α) {n : ℕ} (f : α → Fin n) (hf : Measurable f) : relativeEntropy (P.map f) (Q.map f) ≤ finitePartitionRelativeEntropy P Q","missing":[],"search":"relativeentropy_map_le_finitepartitionrelativeentropy banditrlproof.lowerbounds.relativeentropy_map_le_finitepartitionrelativeentropy every finite measurable observation is included in the defining supremum. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_fin_eq","label":"finitePartitionRelativeEntropy_fin_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_fin_eq","description":"On an already finite observation space, the identity partition loses nothing.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-406061568682","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5377,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finitePartitionRelativeEntropy_fin_eq {n : ℕ} (P Q : Measure (Fin n)) [IsFiniteMeasure P] [IsFiniteMeasure Q] : finitePartitionRelativeEntropy P Q = relativeEntropy P Q","missing":[],"search":"finitepartitionrelativeentropy_fin_eq banditrlproof.lowerbounds.finitepartitionrelativeentropy_fin_eq on an already finite observation space, the identity partition loses nothing. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_map_eq_if","label":"relativeEntropy_finite_map_eq_if","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_finite_map_eq_if","description":"The discrete KL of an observation is the exact cell-mass formula.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-cfdb2d1fa594","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5378,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:112"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_finite_map_eq_if {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] {n : ℕ} (f : α → Fin n) (hf : Measurable f) : relativeEntropy (P.map f) (Q.map f) = if ∀ i, Q (f ⁻¹' {i}) = 0 → P (f ⁻¹' {i}) = 0 then ENNReal.ofReal (∑ i, (P (f ⁻¹' {i})).toReal * Real.log ((P (f ⁻¹' {i})).toReal / (Q (f ⁻¹' {i})).toReal)) else (⊤ : ENNReal)","missing":[],"search":"relativeentropy_finite_map_eq_if banditrlproof.lowerbounds.relativeentropy_finite_map_eq_if the discrete kl of an observation is the exact cell-mass formula. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binary_map_relativeEntropy_eq_top_of_event","label":"exists_binary_map_relativeEntropy_eq_top_of_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binary_map_relativeEntropy_eq_top_of_event","description":"A measurable support mismatch is detected by a two-cell observation.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-b3de281cdfda","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5379,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:128"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binary_map_relativeEntropy_eq_top_of_event {α : Type*} [MeasurableSpace α] (P Q : Measure α) {A : Set α} (hA : MeasurableSet A) (hp : P A ≠ 0) (hq : Q A = 0) : ∃ f : α → Fin 2, Measurable f ∧ relativeEntropy (P.map f) (Q.map f) = (⊤ : ENNReal)","missing":[],"search":"exists_binary_map_relativeentropy_eq_top_of_event banditrlproof.lowerbounds.exists_binary_map_relativeentropy_eq_top_of_event a measurable support mismatch is detected by a two-cell observation. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_top_of_not_absolutelyContinuous","label":"finitePartitionRelativeEntropy_eq_top_of_not_absolutelyContinuous","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_top_of_not_absolutelyContinuous","description":"Non-absolute-continuity forces infinite finite-partition relative entropy.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-bebb973fdfde","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5380,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:144"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finitePartitionRelativeEntropy_eq_top_of_not_absolutelyContinuous {α : Type*} [MeasurableSpace α] (P Q : Measure α) (h : ¬ P ≪ Q) : finitePartitionRelativeEntropy P Q = (⊤ : ENNReal)","missing":[],"search":"finitepartitionrelativeentropy_eq_top_of_not_absolutelycontinuous banditrlproof.lowerbounds.finitepartitionrelativeentropy_eq_top_of_not_absolutelycontinuous non-absolute-continuity forces infinite finite-partition relative entropy. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy_of_not_absolutelyContinuous","label":"finitePartitionRelativeEntropy_eq_relativeEntropy_of_not_absolutelyContinuous","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy_of_not_absolutelyContinuous","description":"The finite-partition and RN definitions agree in the singular branch.","url":"../modules/banditrlproof-lowerbounds-finitepartitionkl/index.html#decl-dbd604f36f0c","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKL","order":5381,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKL"],["Source","BanditRLProof/LowerBounds/FinitePartitionKL.lean:162"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finitePartitionRelativeEntropy_eq_relativeEntropy_of_not_absolutelyContinuous {α : Type*} [MeasurableSpace α] (P Q : Measure α) (h : ¬ P ≪ Q) : finitePartitionRelativeEntropy P Q = relativeEntropy P Q","missing":[],"search":"finitepartitionrelativeentropy_eq_relativeentropy_of_not_absolutelycontinuous banditrlproof.lowerbounds.finitepartitionrelativeentropy_eq_relativeentropy_of_not_absolutelycontinuous the finite-partition and rn definitions agree in the singular branch. theorem compiled","shard":"modules/6dea14d3460ba942.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_fin_encoding_of_finite_range","label":"exists_fin_encoding_of_finite_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_fin_encoding_of_finite_range","description":"A measurable finite-range map admits a finite code and an exact decoder.","url":"../modules/banditrlproof-lowerbounds-finitepartitionklrecovery/index.html#decl-4d7ff17f8d91","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","order":5382,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKLRecovery"],["Source","BanditRLProof/LowerBounds/FinitePartitionKLRecovery.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_fin_encoding_of_finite_range {α β : Type*} [MeasurableSpace α] [MeasurableSpace β] [MeasurableSingletonClass β] (f : α → β) (hf : Measurable f) (hfin : (Set.range f).Finite) : ∃ (n : ℕ) (g : α → Fin n), Measurable g ∧ ∃ d : Fin n → β, f = d ∘ g","missing":[],"search":"exists_fin_encoding_of_finite_range banditrlproof.lowerbounds.exists_fin_encoding_of_finite_range a measurable finite-range map admits a finite code and an exact decoder. theorem compiled","shard":"modules/096ca9a547e4399c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_fin_observation_densityApproximation","label":"exists_fin_observation_densityApproximation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_fin_observation_densityApproximation","description":"Each density-approximation layer is contained in a finite observation sigma-algebra.","url":"../modules/banditrlproof-lowerbounds-finitepartitionklrecovery/index.html#decl-7aee1c198c97","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","order":5383,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKLRecovery"],["Source","BanditRLProof/LowerBounds/FinitePartitionKLRecovery.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_fin_observation_densityApproximation {α : Type*} [m : MeasurableSpace α] (r : α → ENNReal) (n : ℕ) : ∃ (k : ℕ) (g : α → Fin k) (_hg : Measurable g), densityApproximationFiltration r n ≤ (inferInstance : MeasurableSpace (Fin k)).comap g","missing":[],"search":"exists_fin_observation_densityapproximation banditrlproof.lowerbounds.exists_fin_observation_densityapproximation each density-approximation layer is contained in a finite observation sigma-algebra. theorem compiled","shard":"modules/096ca9a547e4399c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_mono","label":"relativeEntropy_trim_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_trim_mono","description":"Refining a sub-sigma-algebra can only increase its retained KL.","url":"../modules/banditrlproof-lowerbounds-finitepartitionklrecovery/index.html#decl-0559a15c84de","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","order":5384,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKLRecovery"],["Source","BanditRLProof/LowerBounds/FinitePartitionKLRecovery.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_trim_mono {α : Type*} {m₁ m₂ m₀ : MeasurableSpace α} (P Q : @Measure α m₀) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h₁₂ : m₁ ≤ m₂) (h₂ : m₂ ≤ m₀) : @relativeEntropy α m₁ (P.trim (h₁₂.trans h₂)) (Q.trim (h₁₂.trans h₂)) ≤ @relativeEntropy α m₂ (P.trim h₂) (Q.trim h₂)","missing":[],"search":"relativeentropy_trim_mono banditrlproof.lowerbounds.relativeentropy_trim_mono refining a sub-sigma-algebra can only increase its retained kl. theorem compiled","shard":"modules/096ca9a547e4399c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy","label":"finitePartitionRelativeEntropy_eq_relativeEntropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy","description":"Textbook Eq. (14.5) and Theorem 14.1: finite discretisations recover RN KL.","url":"../modules/banditrlproof-lowerbounds-finitepartitionklrecovery/index.html#decl-52bbd3d96100","parent":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","order":5385,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FinitePartitionKLRecovery"],["Source","BanditRLProof/LowerBounds/FinitePartitionKLRecovery.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem finitePartitionRelativeEntropy_eq_relativeEntropy {α : Type*} [m : MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] : finitePartitionRelativeEntropy P Q = relativeEntropy P Q","missing":[],"search":"finitepartitionrelativeentropy_eq_relativeentropy banditrlproof.lowerbounds.finitepartitionrelativeentropy_eq_relativeentropy textbook eq. (14.5) and theorem 14.1: finite discretisations recover rn kl. theorem compiled","shard":"modules/096ca9a547e4399c.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_fixedLengthPrefixCode","label":"exists_fixedLengthPrefixCode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_fixedLengthPrefixCode","description":"A finite alphabet fitting in n bits has an actual fixed-length prefix code.","url":"../modules/banditrlproof-lowerbounds-fixedlengthcoding/index.html#decl-e26f2f17b5f4","parent":"module:BanditRLProof.LowerBounds.FixedLengthCoding","order":5386,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FixedLengthCoding"],["Source","BanditRLProof/LowerBounds/FixedLengthCoding.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_fixedLengthPrefixCode {α : Type*} [Fintype α] (n : ℕ) (hn : 0 < n) (hcapacity : Fintype.card α ≤ 2 ^ n) : ∃ code : BinaryPrefixCode α, ∀ a, (code.encode a).length = n","missing":[],"search":"exists_fixedlengthprefixcode banditrlproof.lowerbounds.exists_fixedlengthprefixcode a finite alphabet fitting in n bits has an actual fixed-length prefix code. theorem compiled","shard":"modules/83fdd270b00e8a08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_ceilingLogPrefixCode","label":"exists_ceilingLogPrefixCode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_ceilingLogPrefixCode","description":"The source's ceiling-log binary code, for at least two symbols.","url":"../modules/banditrlproof-lowerbounds-fixedlengthcoding/index.html#decl-dac95bf981ed","parent":"module:BanditRLProof.LowerBounds.FixedLengthCoding","order":5387,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FixedLengthCoding"],["Source","BanditRLProof/LowerBounds/FixedLengthCoding.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem exists_ceilingLogPrefixCode {α : Type*} [Fintype α] (hcard : 1 < Fintype.card α) : ∃ code : BinaryPrefixCode α, ∀ a, (code.encode a).length = Nat.clog 2 (Fintype.card α)","missing":[],"search":"exists_ceilinglogprefixcode banditrlproof.lowerbounds.exists_ceilinglogprefixcode the source's ceiling-log binary code, for at least two symbols. theorem compiled","shard":"modules/83fdd270b00e8a08.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_fixedLength","label":"expectedCodeLength_fixedLength","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_fixedLength","description":"theorem expectedCodeLength_fixedLength {α : Type*} [Fintype α] (p : α → ℝ) (hs : ∑ a, p a = 1) (code : BinaryPrefixCode α) (n : ℕ) (hlen : ∀ a, (code.encode a).length = n) : expectedCodeLength p code = n","url":"../modules/banditrlproof-lowerbounds-fixedlengthcoding/index.html#decl-db5b47d2dace","parent":"module:BanditRLProof.LowerBounds.FixedLengthCoding","order":5388,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.FixedLengthCoding"],["Source","BanditRLProof/LowerBounds/FixedLengthCoding.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_fixedLength {α : Type*} [Fintype α] (p : α → ℝ) (hs : ∑ a, p a = 1) (code : BinaryPrefixCode α) (n : ℕ) (hlen : ∀ a, (code.encode a).length = n) : expectedCodeLength p code = n","missing":[],"search":"expectedcodelength_fixedlength banditrlproof.lowerbounds.expectedcodelength_fixedlength theorem expectedcodelength_fixedlength {α : type*} [fintype α] (p : α → ℝ) (hs : ∑ a, p a = 1) (code : binaryprefixcode α) (n : ℕ) (hlen : ∀ a, (code.encode a).length = n) : expectedcodelength p code = n theorem compiled","shard":"modules/83fdd270b00e8a08.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance","label":"gaussianSampleMeanVariance","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanVariance","description":"Variance `1 / n` of the mean of `n` unit-variance observations.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-2eee5d51467a","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5389,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def gaussianSampleMeanVariance (sampleSize : Nat) : NNReal","missing":[],"search":"gaussiansamplemeanvariance banditrlproof.lowerbounds.gaussiansamplemeanvariance variance `1 / n` of the mean of `n` unit-variance observations. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance_pos","label":"gaussianSampleMeanVariance_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanVariance_pos","description":"The source's sample-mean variance is nondegenerate for a positive sample size.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-4b5a7fb7e8fa","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5390,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianSampleMeanVariance_pos (sampleSize : Nat) (hsampleSize : 0 < sampleSize) : 0 < gaussianSampleMeanVariance sampleSize","missing":[],"search":"gaussiansamplemeanvariance_pos banditrlproof.lowerbounds.gaussiansamplemeanvariance_pos the source's sample-mean variance is nondegenerate for a positive sample size. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanLaw","label":"gaussianSampleMeanLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanLaw","description":"The Gaussian law stated in Chapter 13.1 for the sample mean observation.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-66beb36e772e","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5391,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianSampleMeanLaw (sampleSize : Nat) (mean : Real) : Measure Real","missing":[],"search":"gaussiansamplemeanlaw banditrlproof.lowerbounds.gaussiansamplemeanlaw the gaussian law stated in chapter 13.1 for the sample mean observation. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianIIDObservationLaw","label":"gaussianIIDObservationLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianIIDObservationLaw","description":"Canonical joint law of `n` independent `N(mean, 1)` observations.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-9dda41b43de5","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5392,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianIIDObservationLaw (sampleSize : Nat) (mean : Real) : Measure (Fin sampleSize → Real)","missing":[],"search":"gaussianiidobservationlaw banditrlproof.lowerbounds.gaussianiidobservationlaw canonical joint law of `n` independent `n(mean, 1)` observations. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianCoordinateAverage","label":"gaussianCoordinateAverage","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianCoordinateAverage","description":"Arithmetic mean of a finite coordinate family.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-019be016f529","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5393,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def gaussianCoordinateAverage (sampleSize : Nat) (observations : Fin sampleSize → Real) : Real","missing":[],"search":"gaussiancoordinateaverage banditrlproof.lowerbounds.gaussiancoordinateaverage arithmetic mean of a finite coordinate family. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianIIDSumLaw","label":"gaussianIIDSumLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianIIDSumLaw","description":"The sum of the canonical iid observations has mean `n * mean` and variance `n`.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-a2cf7a31072c","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5394,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianIIDSumLaw (sampleSize : Nat) (mean : Real) : (gaussianIIDObservationLaw sampleSize mean).map (fun observations => ∑ i, observations i) = gaussianReal ((sampleSize : Real) * mean) (sampleSize : NNReal)","missing":[],"search":"gaussianiidsumlaw banditrlproof.lowerbounds.gaussianiidsumlaw the sum of the canonical iid observations has mean `n * mean` and variance `n`. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianIIDSampleMeanLaw","label":"gaussianIIDSampleMeanLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianIIDSampleMeanLaw","description":"The arithmetic mean under the canonical product law of `n > 0` independent `N(mean, 1)` observations has exactly the source law `N(mean, 1 / n)`.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-0a06ebce5491","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5395,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:77"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianIIDSampleMeanLaw (sampleSize : Nat) (mean : Real) (hsampleSize : 0 < sampleSize) : (gaussianIIDObservationLaw sampleSize mean).map (gaussianCoordinateAverage sampleSize) = gaussianSampleMeanLaw sampleSize mean","missing":[],"search":"gaussianiidsamplemeanlaw banditrlproof.lowerbounds.gaussianiidsamplemeanlaw the arithmetic mean under the canonical product law of `n > 0` independent `n(mean, 1)` observations has exactly the source law `n(mean, 1 / n)`. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision","label":"twoPointGaussianThresholdDecision","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision","description":"Threshold rule for the two hypotheses `mean = 0` and `mean = gap`. Ties are assigned to `gap`, matching the source's zero-mean error event `sampleMean >= gap / 2`; for a nondegenerate Gaussian law, the tie has mass zero.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-f4e075dc5168","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5396,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:101"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def twoPointGaussianThresholdDecision (gap observation : Real) : Real","missing":[],"search":"twopointgaussianthresholddecision banditrlproof.lowerbounds.twopointgaussianthresholddecision threshold rule for the two hypotheses `mean = 0` and `mean = gap`. ties are assigned to `gap`, matching the source's zero-mean error event `samplemean >= gap / 2`; for a nondegenerate gaussian law, the tie has mass zero. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_zero_error_event","label":"twoPointGaussianThresholdDecision_zero_error_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_zero_error_event","description":"Under the zero-mean hypothesis, the threshold rule errs exactly above the midpoint.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-e102d69ebe6c","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5397,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:105"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem twoPointGaussianThresholdDecision_zero_error_event {gap : Real} (hgap : 0 < gap) : {observation | twoPointGaussianThresholdDecision gap observation ≠ 0} = Set.Ici (gap / 2)","missing":[],"search":"twopointgaussianthresholddecision_zero_error_event banditrlproof.lowerbounds.twopointgaussianthresholddecision_zero_error_event under the zero-mean hypothesis, the threshold rule errs exactly above the midpoint. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_gap_error_event","label":"twoPointGaussianThresholdDecision_gap_error_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_gap_error_event","description":"Under the positive-mean hypothesis, errors occur exactly below the midpoint.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-b92947c2d8e0","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5398,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem twoPointGaussianThresholdDecision_gap_error_event {gap : Real} (hgap : 0 < gap) : {observation | twoPointGaussianThresholdDecision gap observation ≠ gap} = Set.Iio (gap / 2)","missing":[],"search":"twopointgaussianthresholddecision_gap_error_event banditrlproof.lowerbounds.twopointgaussianthresholddecision_gap_error_event under the positive-mean hypothesis, errors occur exactly below the midpoint. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability","label":"gaussianSampleMeanZeroErrorProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability","description":"Error probability of the threshold decision under the zero-mean sample-mean law.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-aaebaabcf9d6","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5399,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianSampleMeanZeroErrorProbability (sampleSize : Nat) (gap : Real) : Real","missing":[],"search":"gaussiansamplemeanzeroerrorprobability banditrlproof.lowerbounds.gaussiansamplemeanzeroerrorprobability error probability of the threshold decision under the zero-mean sample-mean law. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability","label":"gaussianSampleMeanGapErrorProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability","description":"Error probability of the threshold decision under the positive-mean sample-mean law.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-7d675e966d2f","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5400,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:137"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianSampleMeanGapErrorProbability (sampleSize : Nat) (gap : Real) : Real","missing":[],"search":"gaussiansamplemeangaperrorprobability banditrlproof.lowerbounds.gaussiansamplemeangaperrorprobability error probability of the threshold decision under the positive-mean sample-mean law. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_id_gaussianReal_zero","label":"hasSubgaussianMGF_id_gaussianReal_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.hasSubgaussianMGF_id_gaussianReal_zero","description":"A centered real Gaussian has its variance as a sub-Gaussian proxy.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-292fbf282bde","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5401,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_id_gaussianReal_zero (variance : NNReal) : HasSubgaussianMGF id variance (gaussianReal 0 variance)","missing":[],"search":"hassubgaussianmgf_id_gaussianreal_zero banditrlproof.lowerbounds.hassubgaussianmgf_id_gaussianreal_zero a centered real gaussian has its variance as a sub-gaussian proxy. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_gap_sub_id_gaussianReal","label":"hasSubgaussianMGF_gap_sub_id_gaussianReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.hasSubgaussianMGF_gap_sub_id_gaussianReal","description":"Reflection around the mean turns `N(gap, variance)` into a centered sub-Gaussian.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-6c45bf8ef251","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5402,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:153"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_gap_sub_id_gaussianReal (gap : Real) (variance : NNReal) : HasSubgaussianMGF (fun observation => gap - observation) variance (gaussianReal gap variance)","missing":[],"search":"hassubgaussianmgf_gap_sub_id_gaussianreal banditrlproof.lowerbounds.hassubgaussianmgf_gap_sub_id_gaussianreal reflection around the mean turns `n(gap, variance)` into a centered sub-gaussian. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_Ici_le_exp_neg_sq_div_two_variance","label":"gaussianReal_zero_Ici_le_exp_neg_sq_div_two_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_zero_Ici_le_exp_neg_sq_div_two_variance","description":"Chernoff upper bound for a centered Gaussian right tail.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-143ec84c6968","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5403,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:162"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_zero_Ici_le_exp_neg_sq_div_two_variance (variance : NNReal) (threshold : Real) (hthreshold : 0 ≤ threshold) : (gaussianReal 0 variance).real (Set.Ici threshold) ≤ Real.exp (-threshold ^ 2 / (2 * (variance : Real)))","missing":[],"search":"gaussianreal_zero_ici_le_exp_neg_sq_div_two_variance banditrlproof.lowerbounds.gaussianreal_zero_ici_le_exp_neg_sq_div_two_variance chernoff upper bound for a centered gaussian right tail. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_gap_Iio_half_le_exp_neg_sq_div_two_variance","label":"gaussianReal_gap_Iio_half_le_exp_neg_sq_div_two_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_gap_Iio_half_le_exp_neg_sq_div_two_variance","description":"The positive-mean midpoint error has the same Chernoff exponent by reflection.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-93ebe9c8b95e","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5404,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:170"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_gap_Iio_half_le_exp_neg_sq_div_two_variance (gap : Real) (variance : NNReal) (hgap : 0 < gap) : (gaussianReal gap variance).real (Set.Iio (gap / 2)) ≤ Real.exp (-(gap / 2) ^ 2 / (2 * (variance : Real)))","missing":[],"search":"gaussianreal_gap_iio_half_le_exp_neg_sq_div_two_variance banditrlproof.lowerbounds.gaussianreal_gap_iio_half_le_exp_neg_sq_div_two_variance the positive-mean midpoint error has the same chernoff exponent by reflection. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_le_exp","label":"gaussianSampleMeanZeroErrorProbability_le_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_le_exp","description":"For the Chapter 13.1 zero-mean branch, the midpoint threshold has error at most `exp (-n * gap^2 / 8)` under the stated `N(0, 1/n)` sample-mean law.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-4e2df497737b","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5405,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:194"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianSampleMeanZeroErrorProbability_le_exp (sampleSize : Nat) (gap : Real) (hgap : 0 < gap) : gaussianSampleMeanZeroErrorProbability sampleSize gap ≤ Real.exp (-(sampleSize : Real) * gap ^ 2 / 8)","missing":[],"search":"gaussiansamplemeanzeroerrorprobability_le_exp banditrlproof.lowerbounds.gaussiansamplemeanzeroerrorprobability_le_exp for the chapter 13.1 zero-mean branch, the midpoint threshold has error at most `exp (-n * gap^2 / 8)` under the stated `n(0, 1/n)` sample-mean law. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability_le_exp","label":"gaussianSampleMeanGapErrorProbability_le_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability_le_exp","description":"For the positive-mean branch, the same midpoint rule has the identical `exp (-n * gap^2 / 8)` Chernoff upper bound.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-c530543cf19a","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5406,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:219"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianSampleMeanGapErrorProbability_le_exp (sampleSize : Nat) (gap : Real) (hgap : 0 < gap) : gaussianSampleMeanGapErrorProbability sampleSize gap ≤ Real.exp (-(sampleSize : Real) * gap ^ 2 / 8)","missing":[],"search":"gaussiansamplemeangaperrorprobability_le_exp banditrlproof.lowerbounds.gaussiansamplemeangaperrorprobability_le_exp for the positive-mean branch, the same midpoint rule has the identical `exp (-n * gap^2 / 8)` chernoff upper bound. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk","label":"gaussianSampleMeanThresholdRisk","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk","description":"Worst of the two midpoint-decision error probabilities.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-f64b6d98adc2","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5407,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:240"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianSampleMeanThresholdRisk (sampleSize : Nat) (gap : Real) : Real","missing":[],"search":"gaussiansamplemeanthresholdrisk banditrlproof.lowerbounds.gaussiansamplemeanthresholdrisk worst of the two midpoint-decision error probabilities. definition compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","label":"gaussianSampleMeanThresholdRisk_le_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","description":"Both Gaussian hypotheses obey the same source-shaped Chernoff exponent.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-9aa2e47c8d05","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5408,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:246"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianSampleMeanThresholdRisk_le_exp (sampleSize : Nat) (gap : Real) (hgap : 0 < gap) : gaussianSampleMeanThresholdRisk sampleSize gap ≤ Real.exp (-(sampleSize : Real) * gap ^ 2 / 8)","missing":[],"search":"gaussiansamplemeanthresholdrisk_le_exp banditrlproof.lowerbounds.gaussiansamplemeanthresholdrisk_le_exp both gaussian hypotheses obey the same source-shaped chernoff exponent. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_mills_bounds","label":"gaussianSampleMeanZeroErrorProbability_mills_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_mills_bounds","description":"Exact Mills bounds for the source error probability, in standardized coordinates.","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-6f18d79c2b06","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5409,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:255"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianSampleMeanZeroErrorProbability_mills_bounds (sampleSize : Nat) (hsampleSize : 0 < sampleSize) (gap : Real) (hgap : 0 < gap) : let z := (gap / 2) / Real.sqrt (2 * (gaussianSampleMeanVariance sampleSize : Real)) Real.exp (-z ^ 2) / (z + Real.sqrt (z ^ 2 + 2)) / Real.sqrt Real.pi ≤ gaussianSampleMeanZeroErrorProbability sampleSize gap ∧ gaussianSampleMeanZeroErrorProbability sampleSize gap ≤ Real.exp (-z ^ 2) / (z + Real.sqrt (z ^ 2 + 4 / Real.pi)) / Real.sqrt Real.pi","missing":[],"search":"gaussiansamplemeanzeroerrorprobability_mills_bounds banditrlproof.lowerbounds.gaussiansamplemeanzeroerrorprobability_mills_bounds exact mills bounds for the source error probability, in standardized coordinates. theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_source_bounds","label":"gaussianSampleMeanZeroErrorProbability_source_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_source_bounds","description":"The exact printed two-sided Gaussian testing inequality, Eq. (13.1).","url":"../modules/banditrlproof-lowerbounds-gaussianhypothesistesting/index.html#decl-e48c13abd773","parent":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","order":5410,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianHypothesisTesting"],["Source","BanditRLProof/LowerBounds/GaussianHypothesisTesting.lean:270"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianSampleMeanZeroErrorProbability_source_bounds (sampleSize : Nat) (hsampleSize : 0 < sampleSize) (gap : Real) (hgap : 0 < gap) : let q := (sampleSize : Real) * gap ^ 2 Real.sqrt (8 / Real.pi) * Real.exp (-q / 8) / (Real.sqrt q + Real.sqrt (q + 16)) ≤ gaussianSampleMeanZeroErrorProbability sampleSize gap ∧ gaussianSampleMeanZeroErrorProbability sampleSize gap ≤ Real.sqrt (8 / Real.pi) * Real.exp (-q / 8) / (Real.sqrt q + Real.sqrt (q + 32 / Real.pi))","missing":[],"search":"gaussiansamplemeanzeroerrorprobability_source_bounds banditrlproof.lowerbounds.gaussiansamplemeanzeroerrorprobability_source_bounds the exact printed two-sided gaussian testing inequality, eq. (13.1). theorem compiled","shard":"modules/6d573e942ee2905b.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison","label":"gaussianMillsComparison","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsComparison","description":"noncomputable def gaussianMillsComparison (c x : ℝ) : ℝ","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-118e8b6ba707","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5411,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianMillsComparison (c x : ℝ) : ℝ","missing":[],"search":"gaussianmillscomparison banditrlproof.lowerbounds.gaussianmillscomparison noncomputable def gaussianmillscomparison (c x : ℝ) : ℝ definition compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison_denominator_pos","label":"gaussianMillsComparison_denominator_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsComparison_denominator_pos","description":"theorem gaussianMillsComparison_denominator_pos {c x : ℝ} (hc : 0 < c) : 0 < x + Real.sqrt (x ^ 2 + c)","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-01f9e53a2d82","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5412,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsComparison_denominator_pos {c x : ℝ} (hc : 0 < c) : 0 < x + Real.sqrt (x ^ 2 + c)","missing":[],"search":"gaussianmillscomparison_denominator_pos banditrlproof.lowerbounds.gaussianmillscomparison_denominator_pos theorem gaussianmillscomparison_denominator_pos {c x : ℝ} (hc : 0 < c) : 0 < x + real.sqrt (x ^ 2 + c) theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.hasDerivAt_gaussianMillsComparison","label":"hasDerivAt_gaussianMillsComparison","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.hasDerivAt_gaussianMillsComparison","description":"theorem hasDerivAt_gaussianMillsComparison {c x : ℝ} (hc : 0 < c) : HasDerivAt (gaussianMillsComparison c) (-gaussianMillsComparison c x * (2 * x + 1 / Real.sqrt (x ^ 2 + c))) x","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-5d3b73b7af1f","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5413,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_gaussianMillsComparison {c x : ℝ} (hc : 0 < c) : HasDerivAt (gaussianMillsComparison c) (-gaussianMillsComparison c x * (2 * x + 1 / Real.sqrt (x ^ 2 + c))) x","missing":[],"search":"hasderivat_gaussianmillscomparison banditrlproof.lowerbounds.hasderivat_gaussianmillscomparison theorem hasderivat_gaussianmillscomparison {c x : ℝ} (hc : 0 < c) : hasderivat (gaussianmillscomparison c) (-gaussianmillscomparison c x * (2 * x + 1 / real.sqrt (x ^ 2 + c))) x theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison_lower_derivative_bound","label":"gaussianMillsComparison_lower_derivative_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsComparison_lower_derivative_bound","description":"theorem gaussianMillsComparison_lower_derivative_bound (x : ℝ) : gaussianMillsComparison 2 x * (2 * x + 1 / Real.sqrt (x ^ 2 + 2)) ≤ Real.exp (-x ^ 2)","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-2429f8e43009","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5414,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsComparison_lower_derivative_bound (x : ℝ) : gaussianMillsComparison 2 x * (2 * x + 1 / Real.sqrt (x ^ 2 + 2)) ≤ Real.exp (-x ^ 2)","missing":[],"search":"gaussianmillscomparison_lower_derivative_bound banditrlproof.lowerbounds.gaussianmillscomparison_lower_derivative_bound theorem gaussianmillscomparison_lower_derivative_bound (x : ℝ) : gaussianmillscomparison 2 x * (2 * x + 1 / real.sqrt (x ^ 2 + 2)) ≤ real.exp (-x ^ 2) theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison_pos","label":"gaussianMillsComparison_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsComparison_pos","description":"theorem gaussianMillsComparison_pos {c : ℝ} (hc : 0 < c) (x : ℝ) : 0 < gaussianMillsComparison c x","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-40b95070a0fc","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5415,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:59"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsComparison_pos {c : ℝ} (hc : 0 < c) (x : ℝ) : 0 < gaussianMillsComparison c x","missing":[],"search":"gaussianmillscomparison_pos banditrlproof.lowerbounds.gaussianmillscomparison_pos theorem gaussianmillscomparison_pos {c : ℝ} (hc : 0 < c) (x : ℝ) : 0 < gaussianmillscomparison c x theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.tendsto_gaussianMillsComparison","label":"tendsto_gaussianMillsComparison","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.tendsto_gaussianMillsComparison","description":"theorem tendsto_gaussianMillsComparison {c : ℝ} (hc : 0 < c) : Tendsto (gaussianMillsComparison c) atTop (𝓝 0)","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-8b53f844e97d","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5416,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem tendsto_gaussianMillsComparison {c : ℝ} (hc : 0 < c) : Tendsto (gaussianMillsComparison c) atTop (𝓝 0)","missing":[],"search":"tendsto_gaussianmillscomparison banditrlproof.lowerbounds.tendsto_gaussianmillscomparison theorem tendsto_gaussianmillscomparison {c : ℝ} (hc : 0 < c) : tendsto (gaussianmillscomparison c) attop (𝓝 0) theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMills_lower_integral","label":"gaussianMills_lower_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMills_lower_integral","description":"The exact lower half of Lattimore--Szepesvari Eq. (13.4).","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-e75b8b8803f2","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5417,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:79"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianMills_lower_integral {x : ℝ} (hx : 0 ≤ x) : Real.exp (-x ^ 2) / (x + Real.sqrt (x ^ 2 + 2)) ≤ ∫ t in Ioi x, Real.exp (-t ^ 2)","missing":[],"search":"gaussianmills_lower_integral banditrlproof.lowerbounds.gaussianmills_lower_integral the exact lower half of lattimore--szepesvari eq. (13.4). theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMills_sign_iff","label":"gaussianMills_sign_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMills_sign_iff","description":"Algebraic sign test for the derivative of the upper comparison error.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-bdd0650d0105","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5418,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:107"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMills_sign_iff {c x : ℝ} (hc : 1 < c) (hx : 0 ≤ x) : 0 ≤ x ^ 2 + c - 1 - x * Real.sqrt (x ^ 2 + c) ↔ x ^ 2 * (2 - c) ≤ (c - 1) ^ 2","missing":[],"search":"gaussianmills_sign_iff banditrlproof.lowerbounds.gaussianmills_sign_iff algebraic sign test for the derivative of the upper comparison error. theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMills_sign_threshold","label":"gaussianMills_sign_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMills_sign_threshold","description":"The sign changes at exactly one nonnegative threshold when `1<c<2`.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-cbf9a71928a1","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5419,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:126"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMills_sign_threshold {c x : ℝ} (hc : 1 < c) (hc2 : c < 2) (hx : 0 ≤ x) : 0 ≤ x ^ 2 + c - 1 - x * Real.sqrt (x ^ 2 + c) ↔ x ≤ (c - 1) / Real.sqrt (2 - c)","missing":[],"search":"gaussianmills_sign_threshold banditrlproof.lowerbounds.gaussianmills_sign_threshold the sign changes at exactly one nonnegative threshold when `1<c<2`. theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative","label":"gaussianMillsErrorDerivative","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsErrorDerivative","description":"The derivative of comparison minus Gaussian tail has this explicit value.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-a5080a1b5e01","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5420,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:144"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianMillsErrorDerivative (c x : ℝ) : ℝ","missing":[],"search":"gaussianmillserrorderivative banditrlproof.lowerbounds.gaussianmillserrorderivative the derivative of comparison minus gaussian tail has this explicit value. definition compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_factor","label":"gaussianMillsErrorDerivative_factor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_factor","description":"theorem gaussianMillsErrorDerivative_factor {c x : ℝ} (hc : 0 < c) : gaussianMillsErrorDerivative c x = Real.exp (-x ^ 2) * (x ^ 2 + c - 1 - x * Real.sqrt (x ^ 2 + c)) / ((x + Real.sqrt (x ^ 2 + c)) * Real.sqrt (x ^ 2 + c))","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-553a22117a16","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5421,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsErrorDerivative_factor {c x : ℝ} (hc : 0 < c) : gaussianMillsErrorDerivative c x = Real.exp (-x ^ 2) * (x ^ 2 + c - 1 - x * Real.sqrt (x ^ 2 + c)) / ((x + Real.sqrt (x ^ 2 + c)) * Real.sqrt (x ^ 2 + c))","missing":[],"search":"gaussianmillserrorderivative_factor banditrlproof.lowerbounds.gaussianmillserrorderivative_factor theorem gaussianmillserrorderivative_factor {c x : ℝ} (hc : 0 < c) : gaussianmillserrorderivative c x = real.exp (-x ^ 2) * (x ^ 2 + c - 1 - x * real.sqrt (x ^ 2 + c)) / ((x + real.sqrt (x ^ 2 + c)) * real.sqrt (x ^ 2 + c)) theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_nonneg_iff","label":"gaussianMillsErrorDerivative_nonneg_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_nonneg_iff","description":"theorem gaussianMillsErrorDerivative_nonneg_iff {c x : ℝ} (hc : 1 < c) (hc2 : c < 2) (hx : 0 ≤ x) : 0 ≤ gaussianMillsErrorDerivative c x ↔ x ≤ (c - 1) / Real.sqrt (2 - c)","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-8fcbfe9f438c","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5422,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsErrorDerivative_nonneg_iff {c x : ℝ} (hc : 1 < c) (hc2 : c < 2) (hx : 0 ≤ x) : 0 ≤ gaussianMillsErrorDerivative c x ↔ x ≤ (c - 1) / Real.sqrt (2 - c)","missing":[],"search":"gaussianmillserrorderivative_nonneg_iff banditrlproof.lowerbounds.gaussianmillserrorderivative_nonneg_iff theorem gaussianmillserrorderivative_nonneg_iff {c x : ℝ} (hc : 1 < c) (hc2 : c < 2) (hx : 0 ≤ x) : 0 ≤ gaussianmillserrorderivative c x ↔ x ≤ (c - 1) / real.sqrt (2 - c) theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_source_nonneg_iff","label":"gaussianMillsErrorDerivative_source_nonneg_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_source_nonneg_iff","description":"Specialization to the exact upper-bound constant in source Eq. (13.4).","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-3810a7c191b3","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5423,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsErrorDerivative_source_nonneg_iff {x : ℝ} (hx : 0 ≤ x) : 0 ≤ gaussianMillsErrorDerivative (4 / Real.pi) x ↔ x ≤ (4 / Real.pi - 1) / Real.sqrt (2 - 4 / Real.pi)","missing":[],"search":"gaussianmillserrorderivative_source_nonneg_iff banditrlproof.lowerbounds.gaussianmillserrorderivative_source_nonneg_iff specialization to the exact upper-bound constant in source eq. (13.4). theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsError","label":"gaussianMillsError","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsError","description":"Comparison error expressed using a finite-interval Gaussian integral.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-58731d2aaf51","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5424,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:184"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianMillsError (c x : ℝ) : ℝ","missing":[],"search":"gaussianmillserror banditrlproof.lowerbounds.gaussianmillserror comparison error expressed using a finite-interval gaussian integral. definition compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.hasDerivAt_gaussianMillsError","label":"hasDerivAt_gaussianMillsError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.hasDerivAt_gaussianMillsError","description":"theorem hasDerivAt_gaussianMillsError {c x : ℝ} (hc : 0 < c) : HasDerivAt (gaussianMillsError c) (gaussianMillsErrorDerivative c x) x","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-8f3dd0baaecd","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5425,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:188"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_gaussianMillsError {c x : ℝ} (hc : 0 < c) : HasDerivAt (gaussianMillsError c) (gaussianMillsErrorDerivative c x) x","missing":[],"search":"hasderivat_gaussianmillserror banditrlproof.lowerbounds.hasderivat_gaussianmillserror theorem hasderivat_gaussianmillserror {c x : ℝ} (hc : 0 < c) : hasderivat (gaussianmillserror c) (gaussianmillserrorderivative c x) x theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsError_source_zero","label":"gaussianMillsError_source_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsError_source_zero","description":"theorem gaussianMillsError_source_zero : gaussianMillsError (4 / Real.pi) 0 = 0","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-bf12b840b85b","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5426,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:198"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsError_source_zero : gaussianMillsError (4 / Real.pi) 0 = 0","missing":[],"search":"gaussianmillserror_source_zero banditrlproof.lowerbounds.gaussianmillserror_source_zero theorem gaussianmillserror_source_zero : gaussianmillserror (4 / real.pi) 0 = 0 theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.tendsto_gaussianMillsError","label":"tendsto_gaussianMillsError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.tendsto_gaussianMillsError","description":"theorem tendsto_gaussianMillsError {c : ℝ} (hc : 0 < c) : Tendsto (gaussianMillsError c) atTop (𝓝 0)","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-ba11a1ef9c8c","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5427,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:205"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem tendsto_gaussianMillsError {c : ℝ} (hc : 0 < c) : Tendsto (gaussianMillsError c) atTop (𝓝 0)","missing":[],"search":"tendsto_gaussianmillserror banditrlproof.lowerbounds.tendsto_gaussianmillserror theorem tendsto_gaussianmillserror {c : ℝ} (hc : 0 < c) : tendsto (gaussianmillserror c) attop (𝓝 0) theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMillsError_source_nonneg","label":"gaussianMillsError_source_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMillsError_source_nonneg","description":"theorem gaussianMillsError_source_nonneg {x : ℝ} (hx : 0 ≤ x) : 0 ≤ gaussianMillsError (4 / Real.pi) x","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-3b3de4f947f5","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5428,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:216"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMillsError_source_nonneg {x : ℝ} (hx : 0 ≤ x) : 0 ≤ gaussianMillsError (4 / Real.pi) x","missing":[],"search":"gaussianmillserror_source_nonneg banditrlproof.lowerbounds.gaussianmillserror_source_nonneg theorem gaussianmillserror_source_nonneg {x : ℝ} (hx : 0 ≤ x) : 0 ≤ gaussianmillserror (4 / real.pi) x theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussian_integral_split","label":"gaussian_integral_split","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussian_integral_split","description":"theorem gaussian_integral_split (x : ℝ) : (∫ t in (0 : ℝ)..x, Real.exp (-t ^ 2)) + (∫ t in Ioi x, Real.exp (-t ^ 2)) = Real.sqrt Real.pi / 2","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-44d32ea82578","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5429,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:253"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussian_integral_split (x : ℝ) : (∫ t in (0 : ℝ)..x, Real.exp (-t ^ 2)) + (∫ t in Ioi x, Real.exp (-t ^ 2)) = Real.sqrt Real.pi / 2","missing":[],"search":"gaussian_integral_split banditrlproof.lowerbounds.gaussian_integral_split theorem gaussian_integral_split (x : ℝ) : (∫ t in (0 : ℝ)..x, real.exp (-t ^ 2)) + (∫ t in ioi x, real.exp (-t ^ 2)) = real.sqrt real.pi / 2 theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMills_upper_integral","label":"gaussianMills_upper_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMills_upper_integral","description":"The exact upper half of Lattimore--Szepesvari Eq. (13.4).","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-4ad0cb9da188","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5430,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:276"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem gaussianMills_upper_integral {x : ℝ} (hx : 0 ≤ x) : (∫ t in Ioi x, Real.exp (-t ^ 2)) ≤ Real.exp (-x ^ 2) / (x + Real.sqrt (x ^ 2 + 4 / Real.pi))","missing":[],"search":"gaussianmills_upper_integral banditrlproof.lowerbounds.gaussianmills_upper_integral the exact upper half of lattimore--szepesvari eq. (13.4). theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_tail_integral","label":"gaussianReal_zero_tail_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_zero_tail_integral","description":"Exact density-integral representation for a centered Gaussian tail.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-3c64ecf6f9cf","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5431,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:288"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_zero_tail_integral (v : ℝ≥0) (hv : 0 < v) (a : ℝ) : (gaussianReal 0 v).real (Ici a) = (Real.sqrt (2 * Real.pi * (v : ℝ)))⁻¹ * ∫ t in Ioi a, Real.exp (-t ^ 2 / (2 * (v : ℝ)))","missing":[],"search":"gaussianreal_zero_tail_integral banditrlproof.lowerbounds.gaussianreal_zero_tail_integral exact density-integral representation for a centered gaussian tail. theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_half_tail_integral","label":"gaussianReal_half_tail_integral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_half_tail_integral","description":"The variance-one-half normal law is exactly the normalized source integral.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-9f1e017f2992","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5432,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:299"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_half_tail_integral (a : ℝ) : (gaussianReal 0 (1 / 2 : ℝ≥0)).real (Ici a) = (∫ t in Ioi a, Real.exp (-t ^ 2)) / Real.sqrt Real.pi","missing":[],"search":"gaussianreal_half_tail_integral banditrlproof.lowerbounds.gaussianreal_half_tail_integral the variance-one-half normal law is exactly the normalized source integral. theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_half_mills_bounds","label":"gaussianReal_half_mills_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_half_mills_bounds","description":"Both exact Mills bounds for the normalized variance-one-half Gaussian.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-927bdc7f4f02","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5433,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:307"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_half_mills_bounds {a : ℝ} (ha : 0 ≤ a) : Real.exp (-a ^ 2) / (a + Real.sqrt (a ^ 2 + 2)) / Real.sqrt Real.pi ≤ (gaussianReal 0 (1 / 2 : ℝ≥0)).real (Ici a) ∧ (gaussianReal 0 (1 / 2 : ℝ≥0)).real (Ici a) ≤ Real.exp (-a ^ 2) / (a + Real.sqrt (a ^ 2 + 4 / Real.pi)) / Real.sqrt Real.pi","missing":[],"search":"gaussianreal_half_mills_bounds banditrlproof.lowerbounds.gaussianreal_half_mills_bounds both exact mills bounds for the normalized variance-one-half gaussian. theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_standardized_tail","label":"gaussianReal_zero_standardized_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_zero_standardized_tail","description":"theorem gaussianReal_zero_standardized_tail (v : ℝ≥0) (hv : 0 < v) (a : ℝ) : (gaussianReal 0 v).real (Ici a) = (∫ t in Ioi (a / Real.sqrt (2 * (v : ℝ))), Real.exp (-t ^ 2)) / Real.sqrt Real.pi","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-88ef0efaadab","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5434,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:316"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_zero_standardized_tail (v : ℝ≥0) (hv : 0 < v) (a : ℝ) : (gaussianReal 0 v).real (Ici a) = (∫ t in Ioi (a / Real.sqrt (2 * (v : ℝ))), Real.exp (-t ^ 2)) / Real.sqrt Real.pi","missing":[],"search":"gaussianreal_zero_standardized_tail banditrlproof.lowerbounds.gaussianreal_zero_standardized_tail theorem gaussianreal_zero_standardized_tail (v : ℝ≥0) (hv : 0 < v) (a : ℝ) : (gaussianreal 0 v).real (ici a) = (∫ t in ioi (a / real.sqrt (2 * (v : ℝ))), real.exp (-t ^ 2)) / real.sqrt real.pi theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_mills_bounds","label":"gaussianReal_zero_mills_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_zero_mills_bounds","description":"Exact standardized Mills bounds for every positive Gaussian variance.","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-0bf738440e23","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5435,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:342"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_zero_mills_bounds (v : ℝ≥0) (hv : 0 < v) (a : ℝ) (ha : 0 ≤ a) : let z := a / Real.sqrt (2 * (v : ℝ)) Real.exp (-z ^ 2) / (z + Real.sqrt (z ^ 2 + 2)) / Real.sqrt Real.pi ≤ (gaussianReal 0 v).real (Ici a) ∧ (gaussianReal 0 v).real (Ici a) ≤ Real.exp (-z ^ 2) / (z + Real.sqrt (z ^ 2 + 4 / Real.pi)) / Real.sqrt Real.pi","missing":[],"search":"gaussianreal_zero_mills_bounds banditrlproof.lowerbounds.gaussianreal_zero_mills_bounds exact standardized mills bounds for every positive gaussian variance. theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMills_expression_rescale","label":"gaussianMills_expression_rescale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMills_expression_rescale","description":"Rescale a Mills expression to the square-root denominators in Eq. (13.1).","url":"../modules/banditrlproof-lowerbounds-gaussianmillsratio/index.html#decl-ba29faf54646","parent":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","order":5436,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMillsRatio"],["Source","BanditRLProof/LowerBounds/GaussianMillsRatio.lean:356"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMills_expression_rescale {z c q : ℝ} (hz : 0 ≤ z) (hc : 0 < c) (hq : q = 8 * z ^ 2) : Real.exp (-z ^ 2) / (z + Real.sqrt (z ^ 2 + c)) / Real.sqrt Real.pi = Real.sqrt (8 / Real.pi) * Real.exp (-q / 8) / (Real.sqrt q + Real.sqrt (q + 8 * c))","missing":[],"search":"gaussianmills_expression_rescale banditrlproof.lowerbounds.gaussianmills_expression_rescale rescale a mills expression to the square-root denominators in eq. (13.1). theorem compiled","shard":"modules/578fb1a5c76d9667.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountReal","label":"finiteHistoryPullCountReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryPullCountReal","description":"A realized pull count, converted to `Real` for the source's regret algebra.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-55791a777347","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5437,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryPullCountReal {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : Real","missing":[],"search":"finitehistorypullcountreal banditrlproof.lowerbounds.finitehistorypullcountreal a realized pull count, converted to `real` for the source's regret algebra. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_ne_top","label":"finiteHistoryPullCountENNReal_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_ne_top","description":"theorem finiteHistoryPullCountENNReal_ne_top {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : finiteHistoryPullCountENNReal n history arm ≠ ∞","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-f770a1c708b3","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5438,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPullCountENNReal_ne_top {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : finiteHistoryPullCountENNReal n history arm ≠ ∞","missing":[],"search":"finitehistorypullcountennreal_ne_top banditrlproof.lowerbounds.finitehistorypullcountennreal_ne_top theorem finitehistorypullcountennreal_ne_top {k : nat} {reward : type*} (n : nat) (history : history.finitepairhistory (fin k) reward n) (arm : fin k) : finitehistorypullcountennreal n history arm ≠ ∞ theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_finiteHistoryPullCountENNReal","label":"sum_finiteHistoryPullCountENNReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_finiteHistoryPullCountENNReal","description":"theorem sum_finiteHistoryPullCountENNReal {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : ∑ arm : Fin K, finiteHistoryPullCountENNReal n history arm = n + 1","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-ec3fa5f7c605","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5439,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_finiteHistoryPullCountENNReal {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : ∑ arm : Fin K, finiteHistoryPullCountENNReal n history arm = n + 1","missing":[],"search":"sum_finitehistorypullcountennreal banditrlproof.lowerbounds.sum_finitehistorypullcountennreal theorem sum_finitehistorypullcountennreal {k : nat} {reward : type*} (n : nat) (history : history.finitepairhistory (fin k) reward n) : ∑ arm : fin k, finitehistorypullcountennreal n history arm = n + 1 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_finiteHistoryPullCountReal","label":"sum_finiteHistoryPullCountReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_finiteHistoryPullCountReal","description":"theorem sum_finiteHistoryPullCountReal {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : ∑ arm : Fin K, finiteHistoryPullCountReal n history arm = n + 1","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-ed940b43ae5a","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5440,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_finiteHistoryPullCountReal {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : ∑ arm : Fin K, finiteHistoryPullCountReal n history arm = n + 1","missing":[],"search":"sum_finitehistorypullcountreal banditrlproof.lowerbounds.sum_finitehistorypullcountreal theorem sum_finitehistorypullcountreal {k : nat} {reward : type*} (n : nat) (history : history.finitepairhistory (fin k) reward n) : ∑ arm : fin k, finitehistorypullcountreal n history arm = n + 1 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountReal_nonneg","label":"finiteHistoryPullCountReal_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryPullCountReal_nonneg","description":"theorem finiteHistoryPullCountReal_nonneg {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : 0 ≤ finiteHistoryPullCountReal n history arm","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-966212b000bf","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5441,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:78"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPullCountReal_nonneg {K : Nat} {Reward : Type*} (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (arm : Fin K) : 0 ≤ finiteHistoryPullCountReal n history arm","missing":[],"search":"finitehistorypullcountreal_nonneg banditrlproof.lowerbounds.finitehistorypullcountreal_nonneg theorem finitehistorypullcountreal_nonneg {k : nat} {reward : type*} (n : nat) (history : history.finitepairhistory (fin k) reward n) (arm : fin k) : 0 ≤ finitehistorypullcountreal n history arm theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryPullCountReal","label":"measurable_finiteHistoryPullCountReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_finiteHistoryPullCountReal","description":"theorem measurable_finiteHistoryPullCountReal {K : Nat} {Reward : Type*} [MeasurableSpace Reward] (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryPullCountReal n history arm)","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-e9e6cd89015b","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5442,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryPullCountReal {K : Nat} {Reward : Type*} [MeasurableSpace Reward] (n : Nat) (arm : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryPullCountReal n history arm)","missing":[],"search":"measurable_finitehistorypullcountreal banditrlproof.lowerbounds.measurable_finitehistorypullcountreal theorem measurable_finitehistorypullcountreal {k : nat} {reward : type*} [measurablespace reward] (n : nat) (arm : fin k) : measurable (fun history : history.finitepairhistory (fin k) reward n => finitehistorypullcountreal n history arm) theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.ofReal_mul_probReal_le_lintegral_of_event","label":"ofReal_mul_probReal_le_lintegral_of_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ofReal_mul_probReal_le_lintegral_of_event","description":"Event integration lower bound in the exact real-probability convention used by the Bretagnolle--Huber theorem.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-2cf9e60aa70e","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5443,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:94"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ofReal_mul_probReal_le_lintegral_of_event {alpha : Type*} [MeasurableSpace alpha] {mu : Measure alpha} [IsFiniteMeasure mu] {A : Set alpha} (hA : MeasurableSet A) {c : Real} (hc : 0 ≤ c) {f : alpha -> ENNReal} (hf : forall x, x ∈ A -> ENNReal.ofReal c ≤ f x) : ENNReal.ofReal (c * mu.real A) ≤ ∫⁻ x, f x ∂mu","missing":[],"search":"ofreal_mul_probreal_le_lintegral_of_event banditrlproof.lowerbounds.ofreal_mul_probreal_le_lintegral_of_event event integration lower bound in the exact real-probability convention used by the bretagnolle--huber theorem. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sixteen_div_twentySeven_le_exp_neg_half","label":"sixteen_div_twentySeven_le_exp_neg_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sixteen_div_twentySeven_le_exp_neg_half","description":"A rigorous rational lower bound for the testing constant in the source proof.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-fbe60463f85e","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5444,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:116"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem sixteen_div_twentySeven_le_exp_neg_half : (16 / 27 : Real) ≤ Real.exp (-(1 / 2 : Real))","missing":[],"search":"sixteen_div_twentyseven_le_exp_neg_half banditrlproof.lowerbounds.sixteen_div_twentyseven_le_exp_neg_half a rigorous rational lower bound for the testing constant in the source proof. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment","label":"UnitGaussianBanditEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment","description":"A unit-cube Gaussian environment together with a certified optimal arm. The optimal-arm field is proof data; the induced reward kernel depends only on `mean`.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-54030d2f1783","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5445,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:135"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"structure UnitGaussianBanditEnvironment (K : Nat) where","missing":[],"search":"unitgaussianbanditenvironment banditrlproof.lowerbounds.unitgaussianbanditenvironment a unit-cube gaussian environment together with a certified optimal arm. the optimal-arm field is proof data; the induced reward kernel depends only on `mean`. structure compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianKernel","label":"unitGaussianKernel","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianKernel","description":"The stationary reward kernel induced by a finite vector of Gaussian means.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-a39ab5076890","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5446,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable abbrev unitGaussianKernel {K : Nat} (mean : Fin K -> Real) : Kernel (Fin K) Real","missing":[],"search":"unitgaussiankernel banditrlproof.lowerbounds.unitgaussiankernel the stationary reward kernel induced by a finite vector of gaussian means. abbreviation compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianKernel_apply","label":"unitGaussianKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianKernel_apply","description":"theorem unitGaussianKernel_apply {K : Nat} (mean : Fin K -> Real) (arm : Fin K) : unitGaussianKernel mean arm = unitGaussianArm (mean arm)","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-c8a3038d7299","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5447,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianKernel_apply {K : Nat} (mean : Fin K -> Real) (arm : Fin K) : unitGaussianKernel mean arm = unitGaussianArm (mean arm)","missing":[],"search":"unitgaussiankernel_apply banditrlproof.lowerbounds.unitgaussiankernel_apply theorem unitgaussiankernel_apply {k : nat} (mean : fin k -> real) (arm : fin k) : unitgaussiankernel mean arm = unitgaussianarm (mean arm) theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret","label":"finiteHistoryGaussianPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret","description":"Realized pseudo-regret on an inclusive finite history.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-5c35e9dc18ca","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5448,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryGaussianPseudoRegret {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : ENNReal","missing":[],"search":"finitehistorygaussianpseudoregret banditrlproof.lowerbounds.finitehistorygaussianpseudoregret realized pseudo-regret on an inclusive finite history. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret","label":"gaussianExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret","description":"Expected pseudo-regret under the canonical history law. This is the Lemma 4.5-equivalent gap-times-pull-count form of the source quantity `R_n`, represented in `ENNReal` to align with Mathlib's measure-KL API. The separate reward-sum-regret equality is not asserted by this definition.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-ac1c058dd060","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5449,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:173"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianExpectedPseudoRegret {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : ENNReal","missing":[],"search":"gaussianexpectedpseudoregret banditrlproof.lowerbounds.gaussianexpectedpseudoregret expected pseudo-regret under the canonical history law. this is the lemma 4.5-equivalent gap-times-pull-count form of the source quantity `r_n`, represented in `ennreal` to align with mathlib's measure-kl api. the separate reward-sum-regret equality is not asserted by this definition. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryGaussianPseudoRegret","label":"measurable_finiteHistoryGaussianPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_finiteHistoryGaussianPseudoRegret","description":"theorem measurable_finiteHistoryGaussianPseudoRegret {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Measurable (finiteHistoryGaussianPseudoRegret environment lastRound)","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-e2000cab341c","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5450,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:182"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryGaussianPseudoRegret {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Measurable (finiteHistoryGaussianPseudoRegret environment lastRound)","missing":[],"search":"measurable_finitehistorygaussianpseudoregret banditrlproof.lowerbounds.measurable_finitehistorygaussianpseudoregret theorem measurable_finitehistorygaussianpseudoregret {k : nat} (environment : unitgaussianbanditenvironment k) (lastround : nat) : measurable (finitehistorygaussianpseudoregret environment lastround) theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret_eq_sum_expectedPulls","label":"gaussianExpectedPseudoRegret_eq_sum_expectedPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret_eq_sum_expectedPulls","description":"Regrouping expected Gaussian pseudo-regret by arm pulls.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-f1a6c27e944e","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5451,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:193"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianExpectedPseudoRegret_eq_sum_expectedPulls {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : gaussianExpectedPseudoRegret algorithm environment lastRound = ∑ arm : Fin K, ENNReal.ofReal (environment.mean environment.bestArm - environment.mean arm) * canonicalRealizedExpectedPullCountThrough algorithm (unitGaussianKernel environment.mean) lastRound arm","missing":[],"search":"gaussianexpectedpseudoregret_eq_sum_expectedpulls banditrlproof.lowerbounds.gaussianexpectedpseudoregret_eq_sum_expectedpulls regrouping expected gaussian pseudo-regret by arm pulls. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough","label":"sum_canonicalRealizedExpectedPullCountThrough","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough","description":"Every inclusive history contains exactly `lastRound + 1` pulls, hence so do the expected realized counts.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-ff4f2dc79f27","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5452,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:217"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_canonicalRealizedExpectedPullCountThrough {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (mean : Fin K -> Real) (lastRound : Nat) : ∑ arm : Fin K, canonicalRealizedExpectedPullCountThrough algorithm (unitGaussianKernel mean) lastRound arm = lastRound + 1","missing":[],"search":"sum_canonicalrealizedexpectedpullcountthrough banditrlproof.lowerbounds.sum_canonicalrealizedexpectedpullcountthrough every inclusive history contains exactly `lastround + 1` pulls, hence so do the expected realized counts. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPullCountReal","label":"gaussianExpectedPullCountReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPullCountReal","description":"Real-valued first-environment expected pull count.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-5e0ad062b462","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5453,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:232"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianExpectedPullCountReal {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (mean : Fin K -> Real) (lastRound : Nat) (arm : Fin K) : Real","missing":[],"search":"gaussianexpectedpullcountreal banditrlproof.lowerbounds.gaussianexpectedpullcountreal real-valued first-environment expected pull count. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPullCountReal_nonneg","label":"gaussianExpectedPullCountReal_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPullCountReal_nonneg","description":"theorem gaussianExpectedPullCountReal_nonneg {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (mean : Fin K -> Real) (lastRound : Nat) (arm : Fin K) : 0 ≤ gaussianExpectedPullCountReal algorithm mean lastRound arm","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-927e4a42e753","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5454,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:238"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianExpectedPullCountReal_nonneg {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (mean : Fin K -> Real) (lastRound : Nat) (arm : Fin K) : 0 ≤ gaussianExpectedPullCountReal algorithm mean lastRound arm","missing":[],"search":"gaussianexpectedpullcountreal_nonneg banditrlproof.lowerbounds.gaussianexpectedpullcountreal_nonneg theorem gaussianexpectedpullcountreal_nonneg {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (mean : fin k -> real) (lastround : nat) (arm : fin k) : 0 ≤ gaussianexpectedpullcountreal algorithm mean lastround arm theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_gaussianExpectedPullCountReal","label":"sum_gaussianExpectedPullCountReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_gaussianExpectedPullCountReal","description":"theorem sum_gaussianExpectedPullCountReal {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (mean : Fin K -> Real) (lastRound : Nat) : ∑ arm : Fin K, gaussianExpectedPullCountReal algorithm mean lastRound arm = lastRound + 1","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-e4113d1fb417","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5455,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:245"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_gaussianExpectedPullCountReal {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (mean : Fin K -> Real) (lastRound : Nat) : ∑ arm : Fin K, gaussianExpectedPullCountReal algorithm mean lastRound arm = lastRound + 1","missing":[],"search":"sum_gaussianexpectedpullcountreal banditrlproof.lowerbounds.sum_gaussianexpectedpullcountreal theorem sum_gaussianexpectedpullcountreal {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (mean : fin k -> real) (lastround : nat) : ∑ arm : fin k, gaussianexpectedpullcountreal algorithm mean lastround arm = lastround + 1 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseMean","label":"gaussianMinimaxBaseMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxBaseMean","description":"Base mean vector in the proof of Theorem 15.2: arm zero has mean `gap`, and every alternative has mean zero.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-920667482ed7","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5456,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:271"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def gaussianMinimaxBaseMean {m : Nat} (gap : Real) : Fin (m + 1) -> Real","missing":[],"search":"gaussianminimaxbasemean banditrlproof.lowerbounds.gaussianminimaxbasemean base mean vector in the proof of theorem 15.2: arm zero has mean `gap`, and every alternative has mean zero. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseMean_zero","label":"gaussianMinimaxBaseMean_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxBaseMean_zero","description":"theorem gaussianMinimaxBaseMean_zero {m : Nat} (gap : Real) : gaussianMinimaxBaseMean (m := m) gap 0 = gap","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-e88104ee3ac6","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5457,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:275"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxBaseMean_zero {m : Nat} (gap : Real) : gaussianMinimaxBaseMean (m := m) gap 0 = gap","missing":[],"search":"gaussianminimaxbasemean_zero banditrlproof.lowerbounds.gaussianminimaxbasemean_zero theorem gaussianminimaxbasemean_zero {m : nat} (gap : real) : gaussianminimaxbasemean (m := m) gap 0 = gap theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseMean_succ","label":"gaussianMinimaxBaseMean_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxBaseMean_succ","description":"theorem gaussianMinimaxBaseMean_succ {m : Nat} (gap : Real) (i : Fin m) : gaussianMinimaxBaseMean (m := m) gap i.succ = 0","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-b7bfe41fc6cb","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5458,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:280"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxBaseMean_succ {m : Nat} (gap : Real) (i : Fin m) : gaussianMinimaxBaseMean (m := m) gap i.succ = 0","missing":[],"search":"gaussianminimaxbasemean_succ banditrlproof.lowerbounds.gaussianminimaxbasemean_succ theorem gaussianminimaxbasemean_succ {m : nat} (gap : real) (i : fin m) : gaussianminimaxbasemean (m := m) gap i.succ = 0 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean","label":"gaussianMinimaxChangedMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxChangedMean","description":"Changed mean vector: the selected alternative `i.succ` is raised from zero to `2*gap`, while arm zero remains at `gap`.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-51ac1b830a6d","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5459,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:286"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def gaussianMinimaxChangedMean {m : Nat} (gap : Real) (i : Fin m) : Fin (m + 1) -> Real","missing":[],"search":"gaussianminimaxchangedmean banditrlproof.lowerbounds.gaussianminimaxchangedmean changed mean vector: the selected alternative `i.succ` is raised from zero to `2*gap`, while arm zero remains at `gap`. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_selected","label":"gaussianMinimaxChangedMean_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_selected","description":"theorem gaussianMinimaxChangedMean_selected {m : Nat} (gap : Real) (i : Fin m) : gaussianMinimaxChangedMean gap i i.succ = 2 * gap","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-66c1ab60c98c","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5460,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:292"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxChangedMean_selected {m : Nat} (gap : Real) (i : Fin m) : gaussianMinimaxChangedMean gap i i.succ = 2 * gap","missing":[],"search":"gaussianminimaxchangedmean_selected banditrlproof.lowerbounds.gaussianminimaxchangedmean_selected theorem gaussianminimaxchangedmean_selected {m : nat} (gap : real) (i : fin m) : gaussianminimaxchangedmean gap i i.succ = 2 * gap theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_zero","label":"gaussianMinimaxChangedMean_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_zero","description":"theorem gaussianMinimaxChangedMean_zero {m : Nat} (gap : Real) (i : Fin m) : gaussianMinimaxChangedMean gap i 0 = gap","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-36d260a5860e","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5461,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:298"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxChangedMean_zero {m : Nat} (gap : Real) (i : Fin m) : gaussianMinimaxChangedMean gap i 0 = gap","missing":[],"search":"gaussianminimaxchangedmean_zero banditrlproof.lowerbounds.gaussianminimaxchangedmean_zero theorem gaussianminimaxchangedmean_zero {m : nat} (gap : real) (i : fin m) : gaussianminimaxchangedmean gap i 0 = gap theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_other","label":"gaussianMinimaxChangedMean_other","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_other","description":"theorem gaussianMinimaxChangedMean_other {m : Nat} (gap : Real) {i j : Fin m} (hji : j ≠ i) : gaussianMinimaxChangedMean gap i j.succ = 0","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-6346dcebb4fb","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5462,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:305"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxChangedMean_other {m : Nat} (gap : Real) {i j : Fin m} (hji : j ≠ i) : gaussianMinimaxChangedMean gap i j.succ = 0","missing":[],"search":"gaussianminimaxchangedmean_other banditrlproof.lowerbounds.gaussianminimaxchangedmean_other theorem gaussianminimaxchangedmean_other {m : nat} (gap : real) {i j : fin m} (hji : j ≠ i) : gaussianminimaxchangedmean gap i j.succ = 0 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseEnvironment","label":"gaussianMinimaxBaseEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxBaseEnvironment","description":"Certified base environment.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-92eb850735e7","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5463,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:311"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianMinimaxBaseEnvironment {m : Nat} (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) : UnitGaussianBanditEnvironment (m + 1) where","missing":[],"search":"gaussianminimaxbaseenvironment banditrlproof.lowerbounds.gaussianminimaxbaseenvironment certified base environment. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedEnvironment","label":"gaussianMinimaxChangedEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxChangedEnvironment","description":"Certified changed environment with `i.succ` optimal.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-6fc7251c7c05","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5464,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:329"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianMinimaxChangedEnvironment {m : Nat} (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) : UnitGaussianBanditEnvironment (m + 1) where","missing":[],"search":"gaussianminimaxchangedenvironment banditrlproof.lowerbounds.gaussianminimaxchangedenvironment certified changed environment with `i.succ` optimal. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_ne_top","label":"finiteHistoryGaussianPseudoRegret_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_ne_top","description":"theorem finiteHistoryGaussianPseudoRegret_ne_top {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : finiteHistoryGaussianPseudoRegret environment lastRound history ≠ ∞","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-e35375aef845","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5465,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:358"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryGaussianPseudoRegret_ne_top {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : finiteHistoryGaussianPseudoRegret environment lastRound history ≠ ∞","missing":[],"search":"finitehistorygaussianpseudoregret_ne_top banditrlproof.lowerbounds.finitehistorygaussianpseudoregret_ne_top theorem finitehistorygaussianpseudoregret_ne_top {k : nat} (environment : unitgaussianbanditenvironment k) (lastround : nat) (history : history.finitepairhistory (fin k) real lastround) : finitehistorygaussianpseudoregret environment lastround history ≠ ∞ theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_toReal","label":"finiteHistoryGaussianPseudoRegret_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_toReal","description":"theorem finiteHistoryGaussianPseudoRegret_toReal {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : (finiteHistoryGaussianPseudoRegret environment lastRound history).toReal = ∑ arm : Fin K, (environment.mean environment.bestArm - environment.mean arm) * finiteHistoryPullCountReal lastRound history arm","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-b567fb7a9aff","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5466,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:369"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryGaussianPseudoRegret_toReal {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : (finiteHistoryGaussianPseudoRegret environment lastRound history).toReal = ∑ arm : Fin K, (environment.mean environment.bestArm - environment.mean arm) * finiteHistoryPullCountReal lastRound history arm","missing":[],"search":"finitehistorygaussianpseudoregret_toreal banditrlproof.lowerbounds.finitehistorygaussianpseudoregret_toreal theorem finitehistorygaussianpseudoregret_toreal {k : nat} (environment : unitgaussianbanditenvironment k) (lastround : nat) (history : history.finitepairhistory (fin k) real lastround) : (finitehistorygaussianpseudoregret environment lastround history).toreal = ∑ arm : fin k, (environment.mean environment.bestarm - environment.mean arm) * finitehistorypullcountreal lastround history arm theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseSmallPullEvent","label":"gaussianMinimaxBaseSmallPullEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxBaseSmallPullEvent","description":"The source event `A={T_0(n) <= n/2}`, in the repository's inclusive history convention.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-a9909262bf7d","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5467,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:391"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def gaussianMinimaxBaseSmallPullEvent {m : Nat} (lastRound : Nat) : Set (History.FinitePairHistory (Fin (m + 1)) Real lastRound)","missing":[],"search":"gaussianminimaxbasesmallpullevent banditrlproof.lowerbounds.gaussianminimaxbasesmallpullevent the source event `a={t_0(n) <= n/2}`, in the repository's inclusive history convention. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurableSet_gaussianMinimaxBaseSmallPullEvent","label":"measurableSet_gaussianMinimaxBaseSmallPullEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurableSet_gaussianMinimaxBaseSmallPullEvent","description":"theorem measurableSet_gaussianMinimaxBaseSmallPullEvent {m : Nat} (lastRound : Nat) : MeasurableSet (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-3dde021c8b75","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5468,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:397"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_gaussianMinimaxBaseSmallPullEvent {m : Nat} (lastRound : Nat) : MeasurableSet (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)","missing":[],"search":"measurableset_gaussianminimaxbasesmallpullevent banditrlproof.lowerbounds.measurableset_gaussianminimaxbasesmallpullevent theorem measurableset_gaussianminimaxbasesmallpullevent {m : nat} (lastround : nat) : measurableset (gaussianminimaxbasesmallpullevent (m := m) lastround) theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_base_toReal","label":"finiteHistoryGaussianPseudoRegret_base_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_base_toReal","description":"theorem finiteHistoryGaussianPseudoRegret_base_toReal {m : Nat} (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) : (finiteHistoryGaussianPseudoRegret (gaussianMinimaxBaseEnvironment gap hgap hgap_le) lastRound history).toReal = gap * ((lastRound + 1 : Nat) - finiteHistoryPullCountReal lastRound history 0)","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-edc4dd50001d","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5469,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:405"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryGaussianPseudoRegret_base_toReal {m : Nat} (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) : (finiteHistoryGaussianPseudoRegret (gaussianMinimaxBaseEnvironment gap hgap hgap_le) lastRound history).toReal = gap * ((lastRound + 1 : Nat) - finiteHistoryPullCountReal lastRound history 0)","missing":[],"search":"finitehistorygaussianpseudoregret_base_toreal banditrlproof.lowerbounds.finitehistorygaussianpseudoregret_base_toreal theorem finitehistorygaussianpseudoregret_base_toreal {m : nat} (gap : real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastround : nat) (history : history.finitepairhistory (fin (m + 1)) real lastround) : (finitehistorygaussianpseudoregret (gaussianminimaxbaseenvironment gap hgap hgap_le) lastround history).toreal = gap * ((lastround + 1 : nat) - finitehistorypullcountreal lastround history 0) theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_changed_toReal_lower","label":"finiteHistoryGaussianPseudoRegret_changed_toReal_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_changed_toReal_lower","description":"theorem finiteHistoryGaussianPseudoRegret_changed_toReal_lower {m : Nat} (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) : gap * finiteHistoryPullCountReal lastRound history 0 ≤ (finiteHistoryGaussianPseudoRegret (gaussianMinimaxChangedEnvironment gap i hgap hgap_le) lastRound history).toReal","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-9001434fbd84","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5470,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:431"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryGaussianPseudoRegret_changed_toReal_lower {m : Nat} (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) : gap * finiteHistoryPullCountReal lastRound history 0 ≤ (finiteHistoryGaussianPseudoRegret (gaussianMinimaxChangedEnvironment gap i hgap hgap_le) lastRound history).toReal","missing":[],"search":"finitehistorygaussianpseudoregret_changed_toreal_lower banditrlproof.lowerbounds.finitehistorygaussianpseudoregret_changed_toreal_lower theorem finitehistorygaussianpseudoregret_changed_toreal_lower {m : nat} (gap : real) (i : fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastround : nat) (history : history.finitepairhistory (fin (m + 1)) real lastround) : gap * finitehistorypullcountreal lastround history 0 ≤ (finitehistorygaussianpseudoregret (gaussianminimaxchangedenvironment gap i hgap hgap_le) lastround history).toreal theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.base_event_forces_gaussianPseudoRegret","label":"base_event_forces_gaussianPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.base_event_forces_gaussianPseudoRegret","description":"theorem base_event_forces_gaussianPseudoRegret {m : Nat} (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) (hA : history ∈ gaussianMinimaxBaseSmallPullEvent (m := m) lastRound) : ENNReal.ofReal (((lastRound + 1 : Nat) : Real) * gap / 2) ≤ finiteHistoryGaussianPseudoRegret (gaussianMinimaxBaseEnvironment gap hgap hgap_le) lastRou…","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-0e5642d188ae","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5471,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:457"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem base_event_forces_gaussianPseudoRegret {m : Nat} (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) (hA : history ∈ gaussianMinimaxBaseSmallPullEvent (m := m) lastRound) : ENNReal.ofReal (((lastRound + 1 : Nat) : Real) * gap / 2) ≤ finiteHistoryGaussianPseudoRegret (gaussianMinimaxBaseEnvironment gap hgap hgap_le) lastRound history","missing":[],"search":"base_event_forces_gaussianpseudoregret banditrlproof.lowerbounds.base_event_forces_gaussianpseudoregret theorem base_event_forces_gaussianpseudoregret {m : nat} (gap : real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastround : nat) (history : history.finitepairhistory (fin (m + 1)) real lastround) (ha : history ∈ gaussianminimaxbasesmallpullevent (m := m) lastround) : ennreal.ofreal (((lastround + 1 : nat) : real) * gap / 2) ≤ finitehistorygaussianpseudoregret (gaussianminimaxbaseenvironment gap hgap hgap_le) lastround history theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.changed_complement_forces_gaussianPseudoRegret","label":"changed_complement_forces_gaussianPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.changed_complement_forces_gaussianPseudoRegret","description":"theorem changed_complement_forces_gaussianPseudoRegret {m : Nat} (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) (hAc : history ∈ (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)ᶜ) : ENNReal.ofReal (((lastRound + 1 : Nat) : Real) * gap / 2) ≤ finiteHistoryGaussianPseudoRegret (gaussianMinimaxChangedEnvironmen…","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-75ad3ec402f9","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5472,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:473"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem changed_complement_forces_gaussianPseudoRegret {m : Nat} (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) (history : History.FinitePairHistory (Fin (m + 1)) Real lastRound) (hAc : history ∈ (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)ᶜ) : ENNReal.ofReal (((lastRound + 1 : Nat) : Real) * gap / 2) ≤ finiteHistoryGaussianPseudoRegret (gaussianMinimaxChangedEnvironment gap i hgap hgap_le) lastRound history","missing":[],"search":"changed_complement_forces_gaussianpseudoregret banditrlproof.lowerbounds.changed_complement_forces_gaussianpseudoregret theorem changed_complement_forces_gaussianpseudoregret {m : nat} (gap : real) (i : fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastround : nat) (history : history.finitepairhistory (fin (m + 1)) real lastround) (hac : history ∈ (gaussianminimaxbasesmallpullevent (m := m) lastround)ᶜ) : ennreal.ofreal (((lastround + 1 : nat) : real) * gap / 2) ≤ finitehistorygaussianpseudoregret (gaussianminimaxchangedenvironment gap i hgap hgap_le) lastround history theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.base_event_probability_lower_bound","label":"base_event_probability_lower_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.base_event_probability_lower_bound","description":"theorem base_event_probability_lower_bound {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) : ENNReal.ofReal ((((lastRound + 1 : Nat) : Real) * gap / 2) * (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxBaseMean gap)) lastRound).real (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)) ≤ gaussi…","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-3dbb126d3ff7","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5473,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:493"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem base_event_probability_lower_bound {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) : ENNReal.ofReal ((((lastRound + 1 : Nat) : Real) * gap / 2) * (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxBaseMean gap)) lastRound).real (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)) ≤ gaussianExpectedPseudoRegret algorithm (gaussianMinimaxBaseEnvironment gap hgap hgap_le) lastRound","missing":[],"search":"base_event_probability_lower_bound banditrlproof.lowerbounds.base_event_probability_lower_bound theorem base_event_probability_lower_bound {m : nat} (algorithm : thompson.historyalgorithm (fin (m + 1)) real) (gap : real) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastround : nat) : ennreal.ofreal ((((lastround + 1 : nat) : real) * gap / 2) * (canonicalbandithistorymeasure algorithm (unitgaussiankernel (gaussianminimaxbasemean gap)) lastround).real (gaussianminimaxbasesmallpullevent (m := m) lastround)) ≤ gaussianexpectedpseudoregret algorithm (gaussianminimaxbaseenvironment gap hgap hgap_le) lastround theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.changed_complement_probability_lower_bound","label":"changed_complement_probability_lower_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.changed_complement_probability_lower_bound","description":"theorem changed_complement_probability_lower_bound {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) : ENNReal.ofReal ((((lastRound + 1 : Nat) : Real) * gap / 2) * (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxChangedMean gap i)) lastRound).real (gaussianMinimaxBaseSmallPullEvent (m :…","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-c169def54aa3","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5474,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:510"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem changed_complement_probability_lower_bound {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (i : Fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastRound : Nat) : ENNReal.ofReal ((((lastRound + 1 : Nat) : Real) * gap / 2) * (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxChangedMean gap i)) lastRound).real (gaussianMinimaxBaseSmallPullEvent (m := m) lastRound)ᶜ) ≤ gaussianExpectedPseudoRegret algorithm (gaussianMinimaxChangedEnvironment gap i hgap hgap_le) lastRound","missing":[],"search":"changed_complement_probability_lower_bound banditrlproof.lowerbounds.changed_complement_probability_lower_bound theorem changed_complement_probability_lower_bound {m : nat} (algorithm : thompson.historyalgorithm (fin (m + 1)) real) (gap : real) (i : fin m) (hgap : 0 ≤ gap) (hgap_le : gap ≤ 1 / 2) (lastround : nat) : ennreal.ofreal ((((lastround + 1 : nat) : real) * gap / 2) * (canonicalbandithistorymeasure algorithm (unitgaussiankernel (gaussianminimaxchangedmean gap i)) lastround).real (gaussianminimaxbasesmallpullevent (m := m) lastround)ᶜ) ≤ gaussianexpectedpseudoregret algorithm (gaussianminimaxchangedenvironment gap i hgap hgap_le) lastround theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianKernel_base_changed","label":"klDiv_unitGaussianKernel_base_changed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_unitGaussianKernel_base_changed","description":"theorem klDiv_unitGaussianKernel_base_changed {m : Nat} (gap : Real) (i j : Fin m) : InformationTheory.klDiv (unitGaussianKernel (gaussianMinimaxBaseMean gap) j.succ) (unitGaussianKernel (gaussianMinimaxChangedMean gap i) j.succ) = if j = i then ENNReal.ofReal (2 * gap ^ 2) else 0","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-a7e01a3a44d0","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5475,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:529"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_unitGaussianKernel_base_changed {m : Nat} (gap : Real) (i j : Fin m) : InformationTheory.klDiv (unitGaussianKernel (gaussianMinimaxBaseMean gap) j.succ) (unitGaussianKernel (gaussianMinimaxChangedMean gap i) j.succ) = if j = i then ENNReal.ofReal (2 * gap ^ 2) else 0","missing":[],"search":"kldiv_unitgaussiankernel_base_changed banditrlproof.lowerbounds.kldiv_unitgaussiankernel_base_changed theorem kldiv_unitgaussiankernel_base_changed {m : nat} (gap : real) (i j : fin m) : informationtheory.kldiv (unitgaussiankernel (gaussianminimaxbasemean gap) j.succ) (unitgaussiankernel (gaussianminimaxchangedmean gap i) j.succ) = if j = i then ennreal.ofreal (2 * gap ^ 2) else 0 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianMinimax_base_changed_history","label":"klDiv_gaussianMinimax_base_changed_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_gaussianMinimax_base_changed_history","description":"Lemma 15.1 specialized to the source's base/changed Gaussian pair: only the selected alternative contributes to the directed history KL.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-d2d51d0ea6ce","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5476,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:547"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_gaussianMinimax_base_changed_history {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (i : Fin m) (lastRound : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxBaseMean gap)) lastRound) (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxChangedMean gap i)) lastRound) = canonicalRealizedExpectedPullCountThrough algorithm (unitGaussianKernel (gaussianMinimaxBaseMean gap)) lastRound i.succ * ENNReal.ofReal (2 * gap ^ 2)","missing":[],"search":"kldiv_gaussianminimax_base_changed_history banditrlproof.lowerbounds.kldiv_gaussianminimax_base_changed_history lemma 15.1 specialized to the source's base/changed gaussian pair: only the selected alternative contributes to the directed history kl. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_leastExploredAlternative","label":"exists_gaussianMinimax_leastExploredAlternative","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_gaussianMinimax_leastExploredAlternative","description":"Least-explored alternative under the actual base Gaussian history law.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-6a30f314726b","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5477,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:572"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_gaussianMinimax_leastExploredAlternative {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (lastRound : Nat) : ∃ i : Fin m, gaussianExpectedPullCountReal algorithm (gaussianMinimaxBaseMean gap) lastRound i.succ ≤ ((lastRound + 1 : Nat) : Real) / (m : Real)","missing":[],"search":"exists_gaussianminimax_leastexploredalternative banditrlproof.lowerbounds.exists_gaussianminimax_leastexploredalternative least-explored alternative under the actual base gaussian history law. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_historyKL_le_half","label":"exists_gaussianMinimax_historyKL_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_gaussianMinimax_historyKL_le_half","description":"The source gap choice makes the selected base-to-changed history KL at most `1/2`.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-8fdede289dcd","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5478,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:592"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem exists_gaussianMinimax_historyKL_le_half {m horizon : Nat} (hm : 0 < m) (hmhorizon : m ≤ horizon) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) : let gap := gaussianMinimaxGap (m : Real) (horizon : Real) ∃ i : Fin m, InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxBaseMean gap)) (horizon - 1)) (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel (gaussianMinimaxChangedMean gap i)) (horizon - 1)) ≤ ENNReal.ofReal (1 / 2 : Real)","missing":[],"search":"exists_gaussianminimax_historykl_le_half banditrlproof.lowerbounds.exists_gaussianminimax_historykl_le_half the source gap choice makes the selected base-to-changed history kl at most `1/2`. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.horizon_mul_gaussianMinimaxGap_eq_half_sqrt","label":"horizon_mul_gaussianMinimaxGap_eq_half_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.horizon_mul_gaussianMinimaxGap_eq_half_sqrt","description":"theorem horizon_mul_gaussianMinimaxGap_eq_half_sqrt {alternativeCount horizon : Real} (halternatives : 0 ≤ alternativeCount) (hhorizon : 0 < horizon) : horizon * gaussianMinimaxGap alternativeCount horizon = Real.sqrt (alternativeCount * horizon) / 2","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-61d315c614f2","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5479,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:675"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem horizon_mul_gaussianMinimaxGap_eq_half_sqrt {alternativeCount horizon : Real} (halternatives : 0 ≤ alternativeCount) (hhorizon : 0 < horizon) : horizon * gaussianMinimaxGap alternativeCount horizon = Real.sqrt (alternativeCount * horizon) / 2","missing":[],"search":"horizon_mul_gaussianminimaxgap_eq_half_sqrt banditrlproof.lowerbounds.horizon_mul_gaussianminimaxgap_eq_half_sqrt theorem horizon_mul_gaussianminimaxgap_eq_half_sqrt {alternativecount horizon : real} (halternatives : 0 ≤ alternativecount) (hhorizon : 0 < horizon) : horizon * gaussianminimaxgap alternativecount horizon = real.sqrt (alternativecount * horizon) / 2 theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_unitGaussianBandit_expectedPseudoRegret_ge_sqrt_div_twentySeven","label":"exists_unitGaussianBandit_expectedPseudoRegret_ge_sqrt_div_twentySeven","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_unitGaussianBandit_expectedPseudoRegret_ge_sqrt_div_twentySeven","description":"Quantitative two-environment conclusion in the exact source constant. For every policy, one of the source's base/changed unit-Gaussian bandits has expected pseudo-regret at least `sqrt(m*n)/27`.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-a185972c8c52","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5480,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:710"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_unitGaussianBandit_expectedPseudoRegret_ge_sqrt_div_twentySeven {m horizon : Nat} (hm : 0 < m) (hmhorizon : m ≤ horizon) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) : let gap := gaussianMinimaxGap (m : Real) (horizon : Real) ∃ i : Fin m, ENNReal.ofReal ((1 / 27 : Real) * Real.sqrt ((m : Real) * (horizon : Real))) ≤ max (gaussianExpectedPseudoRegret algorithm (gaussianMinimaxBaseEnvironment gap (show 0 ≤ gaussianMinimaxGap (m : Real) (horizon : Real) from Real.sqrt_nonneg _) (gaussianMinimaxGap_le_half (show (0 : Real) < (horizon : Real) by exact_mod_cast lt_of_lt_of_le hm hmhorizon) (by exact_mod_cast hmhorizon))) (horizon - 1)) (gaussianExpectedPseudoRegret algorithm (gaussianMinimaxChangedEnvironment gap i (show 0 ≤ gaussianMinimaxGap (m : Real) (horizon : Real) from Real.sqrt_nonneg _) (gaussianMinimaxGap_le_half (show (0 : Real) < (horizon : Real) by ex…","missing":[],"search":"exists_unitgaussianbandit_expectedpseudoregret_ge_sqrt_div_twentyseven banditrlproof.lowerbounds.exists_unitgaussianbandit_expectedpseudoregret_ge_sqrt_div_twentyseven quantitative two-environment conclusion in the exact source constant. for every policy, one of the source's base/changed unit-gaussian bandits has expected pseudo-regret at least `sqrt(m*n)/27`. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_unitGaussianBanditEnvironment_expectedPseudoRegret_ge","label":"exists_unitGaussianBanditEnvironment_expectedPseudoRegret_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_unitGaussianBanditEnvironment_expectedPseudoRegret_ge","description":"Source-facing existence form of Theorem 15.2 with `m=k-1`. The returned environment records both the unit-cube mean vector and a genuinely optimal arm; its reward law is the corresponding unit-variance Gaussian bandit.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-f8f45bb20911","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5481,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:834"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_unitGaussianBanditEnvironment_expectedPseudoRegret_ge {m horizon : Nat} (hm : 0 < m) (hmhorizon : m ≤ horizon) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) : ∃ environment : UnitGaussianBanditEnvironment (m + 1), ENNReal.ofReal ((1 / 27 : Real) * Real.sqrt ((m : Real) * (horizon : Real))) ≤ gaussianExpectedPseudoRegret algorithm environment (horizon - 1)","missing":[],"search":"exists_unitgaussianbanditenvironment_expectedpseudoregret_ge banditrlproof.lowerbounds.exists_unitgaussianbanditenvironment_expectedpseudoregret_ge source-facing existence form of theorem 15.2 with `m=k-1`. the returned environment records both the unit-cube mean vector and a genuinely optimal arm; its reward law is the corresponding unit-variance gaussian bandit. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","label":"finiteArmedGaussianMinimaxLowerBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","description":"*Lattimore--Szepesvari, Theorem 15.2.** For `k>1`, `n≥k-1`, and every possibly randomized nonanticipating policy, there exists a mean vector in `[0,1]^k` for which the unit-variance Gaussian bandit has expected pseudo-regret at least `sqrt((k-1)n)/27`. The local history parameter is `n-1`, hence contains exactly `n` observations.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-4c43658e3020","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5482,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:864"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem finiteArmedGaussianMinimaxLowerBound {k horizon : Nat} (hk : 1 < k) (hkhorizon : k - 1 ≤ horizon) (algorithm : Thompson.HistoryAlgorithm (Fin k) Real) : ∃ environment : UnitGaussianBanditEnvironment k, ENNReal.ofReal ((1 / 27 : Real) * Real.sqrt (((k - 1 : Nat) : Real) * (horizon : Real))) ≤ gaussianExpectedPseudoRegret algorithm environment (horizon - 1)","missing":[],"search":"finitearmedgaussianminimaxlowerbound banditrlproof.lowerbounds.finitearmedgaussianminimaxlowerbound *lattimore--szepesvari, theorem 15.2.** for `k>1`, `n≥k-1`, and every possibly randomized nonanticipating policy, there exists a mean vector in `[0,1]^k` for which the unit-variance gaussian bandit has expected pseudo-regret at least `sqrt((k-1)n)/27`. the local history parameter is `n-1`, hence contains exactly `n` observations. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret","label":"unitGaussianWorstCaseExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret","description":"Worst-case expected pseudo-regret over all unit-cube, unit-variance Gaussian environments.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-9ee95b0c556f","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5483,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:882"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def unitGaussianWorstCaseExpectedPseudoRegret (K : Nat) (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (lastRound : Nat) : ENNReal","missing":[],"search":"unitgaussianworstcaseexpectedpseudoregret banditrlproof.lowerbounds.unitgaussianworstcaseexpectedpseudoregret worst-case expected pseudo-regret over all unit-cube, unit-variance gaussian environments. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret","label":"unitGaussianMinimaxExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret","description":"Minimax expected pseudo-regret over stochastic finite-history policies.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-1f514abf4eac","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5484,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:889"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def unitGaussianMinimaxExpectedPseudoRegret (K lastRound : Nat) : ENNReal","missing":[],"search":"unitgaussianminimaxexpectedpseudoregret banditrlproof.lowerbounds.unitgaussianminimaxexpectedpseudoregret minimax expected pseudo-regret over stochastic finite-history policies. definition compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret_ge","label":"unitGaussianWorstCaseExpectedPseudoRegret_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret_ge","description":"theorem unitGaussianWorstCaseExpectedPseudoRegret_ge {k horizon : Nat} (hk : 1 < k) (hkhorizon : k - 1 ≤ horizon) (algorithm : Thompson.HistoryAlgorithm (Fin k) Real) : ENNReal.ofReal ((1 / 27 : Real) * Real.sqrt (((k - 1 : Nat) : Real) * (horizon : Real))) ≤ unitGaussianWorstCaseExpectedPseudoRegret k algorithm (horizon - 1)","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-600fce63f71a","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5485,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:894"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianWorstCaseExpectedPseudoRegret_ge {k horizon : Nat} (hk : 1 < k) (hkhorizon : k - 1 ≤ horizon) (algorithm : Thompson.HistoryAlgorithm (Fin k) Real) : ENNReal.ofReal ((1 / 27 : Real) * Real.sqrt (((k - 1 : Nat) : Real) * (horizon : Real))) ≤ unitGaussianWorstCaseExpectedPseudoRegret k algorithm (horizon - 1)","missing":[],"search":"unitgaussianworstcaseexpectedpseudoregret_ge banditrlproof.lowerbounds.unitgaussianworstcaseexpectedpseudoregret_ge theorem unitgaussianworstcaseexpectedpseudoregret_ge {k horizon : nat} (hk : 1 < k) (hkhorizon : k - 1 ≤ horizon) (algorithm : thompson.historyalgorithm (fin k) real) : ennreal.ofreal ((1 / 27 : real) * real.sqrt (((k - 1 : nat) : real) * (horizon : real))) ≤ unitgaussianworstcaseexpectedpseudoregret k algorithm (horizon - 1) theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge","label":"unitGaussianMinimaxExpectedPseudoRegret_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge","description":"Minimax form of Theorem 15.2.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-7d888e72081c","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5486,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:909"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianMinimaxExpectedPseudoRegret_ge {k horizon : Nat} (hk : 1 < k) (hkhorizon : k - 1 ≤ horizon) : ENNReal.ofReal ((1 / 27 : Real) * Real.sqrt (((k - 1 : Nat) : Real) * (horizon : Real))) ≤ unitGaussianMinimaxExpectedPseudoRegret k (horizon - 1)","missing":[],"search":"unitgaussianminimaxexpectedpseudoregret_ge banditrlproof.lowerbounds.unitgaussianminimaxexpectedpseudoregret_ge minimax form of theorem 15.2. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt","label":"unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt","description":"Chapter 13's coarser `c*sqrt(k*n)` statement, now discharged from the exact Chapter 15 constant. The explicit universal choice here is `c=1/54`; the source only asks for existence of a positive universal constant.","url":"../modules/banditrlproof-lowerbounds-gaussianminimax/index.html#decl-ae72137aa924","parent":"module:BanditRLProof.LowerBounds.GaussianMinimax","order":5487,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianMinimax"],["Source","BanditRLProof/LowerBounds/GaussianMinimax.lean:922"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","stochastic-finite"]],"statement":"theorem unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt {k horizon : Nat} (hk : 1 < k) (hkhorizon : k ≤ horizon) : ENNReal.ofReal ((1 / 54 : Real) * Real.sqrt ((k : Real) * (horizon : Real))) ≤ unitGaussianMinimaxExpectedPseudoRegret k (horizon - 1)","missing":[],"search":"unitgaussianminimaxexpectedpseudoregret_ge_one_div_fiftyfour_sqrt banditrlproof.lowerbounds.unitgaussianminimaxexpectedpseudoregret_ge_one_div_fiftyfour_sqrt chapter 13's coarser `c*sqrt(k*n)` statement, now discharged from the exact chapter 15 constant. the explicit universal choice here is `c=1/54`; the source only asks for existence of a positive universal constant. theorem compiled","shard":"modules/585f4c0b4877ae79.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":["stochastic-finite"]},{"id":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_same_variance","label":"log_gaussianPDFReal_div_same_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.log_gaussianPDFReal_div_same_variance","description":"theorem log_gaussianPDFReal_div_same_variance (m n x : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : Real.log (gaussianPDFReal m v x / gaussianPDFReal n v x) = ((m - n) * x + (n ^ 2 - m ^ 2) / 2) / v","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-e0f73eb6900c","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5488,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem log_gaussianPDFReal_div_same_variance (m n x : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : Real.log (gaussianPDFReal m v x / gaussianPDFReal n v x) = ((m - n) * x + (n ^ 2 - m ^ 2) / 2) / v","missing":[],"search":"log_gaussianpdfreal_div_same_variance banditrlproof.lowerbounds.log_gaussianpdfreal_div_same_variance theorem log_gaussianpdfreal_div_same_variance (m n x : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : real.log (gaussianpdfreal m v x / gaussianpdfreal n v x) = ((m - n) * x + (n ^ 2 - m ^ 2) / 2) / v theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_same_variance_ae","label":"llr_gaussianReal_same_variance_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.llr_gaussianReal_same_variance_ae","description":"theorem llr_gaussianReal_same_variance_ae (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : llr (gaussianReal m v) (gaussianReal n v) =ᵐ[gaussianReal m v] fun x => ((m - n) * x + (n ^ 2 - m ^ 2) / 2) / v","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-354019074e59","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5489,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem llr_gaussianReal_same_variance_ae (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : llr (gaussianReal m v) (gaussianReal n v) =ᵐ[gaussianReal m v] fun x => ((m - n) * x + (n ^ 2 - m ^ 2) / 2) / v","missing":[],"search":"llr_gaussianreal_same_variance_ae banditrlproof.lowerbounds.llr_gaussianreal_same_variance_ae theorem llr_gaussianreal_same_variance_ae (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : llr (gaussianreal m v) (gaussianreal n v) =ᵐ[gaussianreal m v] fun x => ((m - n) * x + (n ^ 2 - m ^ 2) / 2) / v theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_same_variance","label":"integrable_llr_gaussianReal_same_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_llr_gaussianReal_same_variance","description":"theorem integrable_llr_gaussianReal_same_variance (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : Integrable (llr (gaussianReal m v) (gaussianReal n v)) (gaussianReal m v)","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-f5a10c60214f","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5490,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:35"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_llr_gaussianReal_same_variance (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : Integrable (llr (gaussianReal m v) (gaussianReal n v)) (gaussianReal m v)","missing":[],"search":"integrable_llr_gaussianreal_same_variance banditrlproof.lowerbounds.integrable_llr_gaussianreal_same_variance theorem integrable_llr_gaussianreal_same_variance (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : integrable (llr (gaussianreal m v) (gaussianreal n v)) (gaussianreal m v) theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_same_variance","label":"klDiv_gaussianReal_same_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_gaussianReal_same_variance","description":"The common positive variance Gaussian KL formula in Chapter 14.","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-d54e1c463605","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5491,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem klDiv_gaussianReal_same_variance (m n : ℝ) (v : ℝ≥0) (hv : v ≠ 0) : InformationTheory.klDiv (gaussianReal m v) (gaussianReal n v) = ENNReal.ofReal ((m - n) ^ 2 / (2 * v))","missing":[],"search":"kldiv_gaussianreal_same_variance banditrlproof.lowerbounds.kldiv_gaussianreal_same_variance the common positive variance gaussian kl formula in chapter 14. theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.three_fifths_le_exp_neg_half","label":"three_fifths_le_exp_neg_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.three_fifths_le_exp_neg_half","description":"Exact rational certification of the displayed Gaussian testing constant.","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-5ca39360ad6a","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5492,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:63"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem three_fifths_le_exp_neg_half : (3 / 5 : ℝ) ≤ Real.exp (-(1 / 2 : ℝ))","missing":[],"search":"three_fifths_le_exp_neg_half banditrlproof.lowerbounds.three_fifths_le_exp_neg_half exact rational certification of the displayed gaussian testing constant. theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussian_testing_error_lower_bound","label":"gaussian_testing_error_lower_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussian_testing_error_lower_bound","description":"Testing two Gaussian means from one observation, with arbitrary positive variance.","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-d058c39e97e3","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5493,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:74"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussian_testing_error_lower_bound (Δ : ℝ) (v : ℝ≥0) (hv : v ≠ 0) {A : Set ℝ} (hA : MeasurableSet A) : (1 / 2 : ℝ) * Real.exp (-(Δ ^ 2 / (2 * v))) ≤ (gaussianReal 0 v).real A + (gaussianReal Δ v).real Aᶜ","missing":[],"search":"gaussian_testing_error_lower_bound banditrlproof.lowerbounds.gaussian_testing_error_lower_bound testing two gaussian means from one observation, with arbitrary positive variance. theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussian_testing_error_three_tenths","label":"gaussian_testing_error_three_tenths","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussian_testing_error_three_tenths","description":"Under signal-to-noise ratio at most one, the sum of errors is at least 3/10.","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-b14511b212f8","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5494,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussian_testing_error_three_tenths (Δ : ℝ) (v : ℝ≥0) (hv : v ≠ 0) (hsnr : Δ ^ 2 / (v : ℝ) ≤ 1) {A : Set ℝ} (hA : MeasurableSet A) : (3 / 10 : ℝ) ≤ (gaussianReal 0 v).real A + (gaussianReal Δ v).real Aᶜ","missing":[],"search":"gaussian_testing_error_three_tenths banditrlproof.lowerbounds.gaussian_testing_error_three_tenths under signal-to-noise ratio at most one, the sum of errors is at least 3/10. theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussian_testing_max_error_three_twentieths","label":"gaussian_testing_max_error_three_twentieths","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussian_testing_max_error_three_twentieths","description":"No measurable decision rule has both errors below 3/20 in the low-SNR regime.","url":"../modules/banditrlproof-lowerbounds-gaussiantesting/index.html#decl-287ff64336ad","parent":"module:BanditRLProof.LowerBounds.GaussianTesting","order":5495,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.GaussianTesting"],["Source","BanditRLProof/LowerBounds/GaussianTesting.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem gaussian_testing_max_error_three_twentieths (Δ : ℝ) (v : ℝ≥0) (hv : v ≠ 0) (hsnr : Δ ^ 2 / (v : ℝ) ≤ 1) {A : Set ℝ} (hA : MeasurableSet A) : (3 / 20 : ℝ) ≤ max ((gaussianReal 0 v).real A) ((gaussianReal Δ v).real Aᶜ)","missing":[],"search":"gaussian_testing_max_error_three_twentieths banditrlproof.lowerbounds.gaussian_testing_max_error_three_twentieths no measurable decision rule has both errors below 3/20 in the low-snr regime. theorem compiled","shard":"modules/b880d6ac0ebab727.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.tailAtLeast","label":"tailAtLeast","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.tailAtLeast","description":"The tail event used throughout Chapter 17: the realized quantity is at least the displayed lower-bound threshold.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f4b3d13b0c12","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5496,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:39"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"def tailAtLeast {Omega : Type*} (quantity : Omega -> Real) (threshold : Real) : Set Omega","missing":[],"search":"tailatleast banditrlproof.lowerbounds.tailatleast the tail event used throughout chapter 17: the realized quantity is at least the displayed lower-bound threshold. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold","label":"stochasticHighProbabilityThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold","description":"The exact threshold inside Theorem 17.1, with `alternativeArms = k - 1`. The theorem's factor `1/4` multiplies the whole minimum.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-4bb7c14eda2f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5497,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"def stochasticHighProbabilityThreshold (horizon alternativeArms : Nat) (B delta : Real) : Real","missing":[],"search":"stochastichighprobabilitythreshold banditrlproof.lowerbounds.stochastichighprobabilitythreshold the exact threshold inside theorem 17.1, with `alternativearms = k - 1`. the theorem's factor `1/4` multiplies the whole minimum. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold","label":"stochasticMinimaxHighProbabilityThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold","description":"The exact threshold in Corollary 17.2, again with `alternativeArms = k - 1`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-d4aa01caa584","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5498,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"def stochasticMinimaxHighProbabilityThreshold (horizon alternativeArms : Nat) (delta : Real) : Real","missing":[],"search":"stochasticminimaxhighprobabilitythreshold banditrlproof.lowerbounds.stochasticminimaxhighprobabilitythreshold the exact threshold in corollary 17.2, again with `alternativearms = k - 1`. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHighProbabilityThreshold","label":"adversarialHighProbabilityThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHighProbabilityThreshold","description":"The threshold shape in Theorem 17.4. The universal constant `c` remains an explicit argument, and the logarithm is `log (1 / (2 * delta))`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-abad2730b98f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5499,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:64"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"def adversarialHighProbabilityThreshold (horizon arms : Nat) (c delta : Real) : Real","missing":[],"search":"adversarialhighprobabilitythreshold banditrlproof.lowerbounds.adversarialhighprobabilitythreshold the threshold shape in theorem 17.4. the universal constant `c` remains an explicit argument, and the logarithm is `log (1 / (2 * delta))`. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret","label":"gaussianRandomPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret","description":"The stochastic random pseudo-regret from Section 17.1, evaluated on one realized canonical finite history. It is deliberately separate from the deterministic expected pseudo-regret below.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-012409374629","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5500,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:72"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianRandomPseudoRegret {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : Real","missing":[],"search":"gaussianrandompseudoregret banditrlproof.lowerbounds.gaussianrandompseudoregret the stochastic random pseudo-regret from section 17.1, evaluated on one realized canonical finite history. it is deliberately separate from the deterministic expected pseudo-regret below. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegretReal","label":"gaussianExpectedPseudoRegretReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPseudoRegretReal","description":"The deterministic expected pseudo-regret presentation used in Eq. (17.4), written as expected pull counts times gaps (the Lemma 4.5 identity). This is not the random pseudo-regret and is not adversarial random regret.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-57e8233fb96f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5501,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianExpectedPseudoRegretReal {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Real","missing":[],"search":"gaussianexpectedpseudoregretreal banditrlproof.lowerbounds.gaussianexpectedpseudoregretreal the deterministic expected pseudo-regret presentation used in eq. (17.4), written as expected pull counts times gaps (the lemma 4.5 identity). this is not the random pseudo-regret and is not adversarial random regret. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.GapOneGaussianBanditEnvironment","label":"GapOneGaussianBanditEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.GapOneGaussianBanditEnvironment","description":"The source class `E^k`: unit-variance Gaussian arms with gaps bounded by one, without an artificial absolute restriction on the means.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-13c4a0ea8a67","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5502,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:91"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure GapOneGaussianBanditEnvironment (K : Nat) where","missing":[],"search":"gaponegaussianbanditenvironment banditrlproof.lowerbounds.gaponegaussianbanditenvironment the source class `e^k`: unit-variance gaussian arms with gaps bounded by one, without an artificial absolute restriction on the means. structure compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gapOneGaussianExpectedPseudoRegretReal","label":"gapOneGaussianExpectedPseudoRegretReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gapOneGaussianExpectedPseudoRegretReal","description":"Deterministic expected pseudo-regret for the full source class `E^k`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ca2430d0386d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5503,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:98"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapOneGaussianExpectedPseudoRegretReal {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : GapOneGaussianBanditEnvironment K) (lastRound : Nat) : Real","missing":[],"search":"gaponegaussianexpectedpseudoregretreal banditrlproof.lowerbounds.gaponegaussianexpectedpseudoregretreal deterministic expected pseudo-regret for the full source class `e^k`. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gapOneGaussianRandomPseudoRegret","label":"gapOneGaussianRandomPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gapOneGaussianRandomPseudoRegret","description":"The stochastic random pseudo-regret on the full source class `E^k`. Unlike `gapOneGaussianExpectedPseudoRegretReal`, this is a random variable on the realized finite history.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3ab8d4efd866","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5504,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def gapOneGaussianRandomPseudoRegret {K : Nat} (environment : GapOneGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : Real","missing":[],"search":"gaponegaussianrandompseudoregret banditrlproof.lowerbounds.gaponegaussianrandompseudoregret the stochastic random pseudo-regret on the full source class `e^k`. unlike `gaponegaussianexpectedpseudoregretreal`, this is a random variable on the realized finite history. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment.toGapOne","label":"toGapOne","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment.toGapOne","description":"def UnitGaussianBanditEnvironment.toGapOne {K : Nat} (environment : UnitGaussianBanditEnvironment K) : GapOneGaussianBanditEnvironment K where","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-72ff0152254f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5505,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def UnitGaussianBanditEnvironment.toGapOne {K : Nat} (environment : UnitGaussianBanditEnvironment K) : GapOneGaussianBanditEnvironment K where","missing":[],"search":"togapone banditrlproof.lowerbounds.unitgaussianbanditenvironment.togapone def unitgaussianbanditenvironment.togapone {k : nat} (environment : unitgaussianbanditenvironment k) : gaponegaussianbanditenvironment k where definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gapOneGaussianExpectedPseudoRegretReal_toGapOne","label":"gapOneGaussianExpectedPseudoRegretReal_toGapOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gapOneGaussianExpectedPseudoRegretReal_toGapOne","description":"theorem gapOneGaussianExpectedPseudoRegretReal_toGapOne {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : gapOneGaussianExpectedPseudoRegretReal algorithm environment.toGapOne lastRound = gaussianExpectedPseudoRegretReal algorithm environment lastRound","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-edc0c147cdc3","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5506,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:130"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapOneGaussianExpectedPseudoRegretReal_toGapOne {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : gapOneGaussianExpectedPseudoRegretReal algorithm environment.toGapOne lastRound = gaussianExpectedPseudoRegretReal algorithm environment lastRound","missing":[],"search":"gaponegaussianexpectedpseudoregretreal_togapone banditrlproof.lowerbounds.gaponegaussianexpectedpseudoregretreal_togapone theorem gaponegaussianexpectedpseudoregretreal_togapone {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (environment : unitgaussianbanditenvironment k) (lastround : nat) : gaponegaussianexpectedpseudoregretreal algorithm environment.togapone lastround = gaussianexpectedpseudoregretreal algorithm environment lastround theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gapOneGaussianRandomPseudoRegret_toGapOne","label":"gapOneGaussianRandomPseudoRegret_toGapOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gapOneGaussianRandomPseudoRegret_toGapOne","description":"theorem gapOneGaussianRandomPseudoRegret_toGapOne {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : gapOneGaussianRandomPseudoRegret environment.toGapOne lastRound history = gaussianRandomPseudoRegret environment lastRound history","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e54f5f4a6a82","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5507,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:137"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapOneGaussianRandomPseudoRegret_toGapOne {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : gapOneGaussianRandomPseudoRegret environment.toGapOne lastRound history = gaussianRandomPseudoRegret environment lastRound history","missing":[],"search":"gaponegaussianrandompseudoregret_togapone banditrlproof.lowerbounds.gaponegaussianrandompseudoregret_togapone theorem gaponegaussianrandompseudoregret_togapone {k : nat} (environment : unitgaussianbanditenvironment k) (lastround : nat) (history : history.finitepairhistory (fin k) real lastround) : gaponegaussianrandompseudoregret environment.togapone lastround history = gaussianrandompseudoregret environment lastround history theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityGap","label":"stochasticHighProbabilityGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticHighProbabilityGap","description":"The exact gap selected in the proof of Theorem 17.1.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1be4b5c2221e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5508,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:148"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticHighProbabilityGap (horizon alternativeArms : Nat) (B delta : Real) : Real","missing":[],"search":"stochastichighprobabilitygap banditrlproof.lowerbounds.stochastichighprobabilitygap the exact gap selected in the proof of theorem 17.1. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_nonneg","label":"gaussianRandomPseudoRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_nonneg","description":"theorem gaussianRandomPseudoRegret_nonneg {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : 0 <= gaussianRandomPseudoRegret environment lastRound history","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5e5f3d1aa365","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5509,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:155"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianRandomPseudoRegret_nonneg {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : 0 <= gaussianRandomPseudoRegret environment lastRound history","missing":[],"search":"gaussianrandompseudoregret_nonneg banditrlproof.lowerbounds.gaussianrandompseudoregret_nonneg theorem gaussianrandompseudoregret_nonneg {k : nat} (environment : unitgaussianbanditenvironment k) (lastround : nat) (history : history.finitepairhistory (fin k) real lastround) : 0 <= gaussianrandompseudoregret environment lastround history theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_gaussianRandomPseudoRegret","label":"measurable_gaussianRandomPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_gaussianRandomPseudoRegret","description":"theorem measurable_gaussianRandomPseudoRegret {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Measurable (gaussianRandomPseudoRegret environment lastRound)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-0f74790561ba","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5510,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:162"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_gaussianRandomPseudoRegret {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Measurable (gaussianRandomPseudoRegret environment lastRound)","missing":[],"search":"measurable_gaussianrandompseudoregret banditrlproof.lowerbounds.measurable_gaussianrandompseudoregret theorem measurable_gaussianrandompseudoregret {k : nat} (environment : unitgaussianbanditenvironment k) (lastround : nat) : measurable (gaussianrandompseudoregret environment lastround) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_le_horizon","label":"gaussianRandomPseudoRegret_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_le_horizon","description":"Unit-cube gaps and the exact pull-count sum bound stochastic random pseudo-regret by the number of observed rounds.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1cc804d0e96d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5511,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:171"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianRandomPseudoRegret_le_horizon {K : Nat} (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Real lastRound) : gaussianRandomPseudoRegret environment lastRound history <= (lastRound + 1 : Nat)","missing":[],"search":"gaussianrandompseudoregret_le_horizon banditrlproof.lowerbounds.gaussianrandompseudoregret_le_horizon unit-cube gaps and the exact pull-count sum bound stochastic random pseudo-regret by the number of observed rounds. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_gaussianRandomPseudoRegret","label":"integrable_gaussianRandomPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_gaussianRandomPseudoRegret","description":"theorem integrable_gaussianRandomPseudoRegret {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Integrable (gaussianRandomPseudoRegret environment lastRound) (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) lastRound)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-000d237a77e5","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5512,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:197"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_gaussianRandomPseudoRegret {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : Integrable (gaussianRandomPseudoRegret environment lastRound) (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) lastRound)","missing":[],"search":"integrable_gaussianrandompseudoregret banditrlproof.lowerbounds.integrable_gaussianrandompseudoregret theorem integrable_gaussianrandompseudoregret {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (environment : unitgaussianbanditenvironment k) (lastround : nat) : integrable (gaussianrandompseudoregret environment lastround) (canonicalbandithistorymeasure algorithm (unitgaussiankernel environment.mean) lastround) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret_toReal_eq","label":"gaussianExpectedPseudoRegret_toReal_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret_toReal_eq","description":"theorem gaussianExpectedPseudoRegret_toReal_eq {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : (gaussianExpectedPseudoRegret algorithm environment lastRound).toReal = gaussianExpectedPseudoRegretReal algorithm environment lastRound","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-d368e8771930","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5513,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:213"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianExpectedPseudoRegret_toReal_eq {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : (gaussianExpectedPseudoRegret algorithm environment lastRound).toReal = gaussianExpectedPseudoRegretReal algorithm environment lastRound","missing":[],"search":"gaussianexpectedpseudoregret_toreal_eq banditrlproof.lowerbounds.gaussianexpectedpseudoregret_toreal_eq theorem gaussianexpectedpseudoregret_toreal_eq {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (environment : unitgaussianbanditenvironment k) (lastround : nat) : (gaussianexpectedpseudoregret algorithm environment lastround).toreal = gaussianexpectedpseudoregretreal algorithm environment lastround theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_gaussianRandomPseudoRegret_eq_expected","label":"integral_gaussianRandomPseudoRegret_eq_expected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_gaussianRandomPseudoRegret_eq_expected","description":"The deterministic expected pseudo-regret surface is the integral of the random pseudo-regret under the same policy/environment history law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a16171acf085","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5514,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:235"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_gaussianRandomPseudoRegret_eq_expected {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitGaussianBanditEnvironment K) (lastRound : Nat) : ∫ history, gaussianRandomPseudoRegret environment lastRound history ∂canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) lastRound = gaussianExpectedPseudoRegretReal algorithm environment lastRound","missing":[],"search":"integral_gaussianrandompseudoregret_eq_expected banditrlproof.lowerbounds.integral_gaussianrandompseudoregret_eq_expected the deterministic expected pseudo-regret surface is the integral of the random pseudo-regret under the same policy/environment history law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_le_threshold_add_bound_mul_tailMass","label":"integral_le_threshold_add_bound_mul_tailMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_le_threshold_add_bound_mul_tailMass","description":"Integrating a nonnegative bounded random variable after splitting at one measurable upper-tail event.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-b3eb9442608a","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5515,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:273"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_le_threshold_add_bound_mul_tailMass {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (quantity : Omega -> Real) (threshold bound : Real) (hmeas : Measurable quantity) (hintegrable : Integrable quantity mu) (hthreshold : 0 <= threshold) (hbound : forall omega, quantity omega <= bound) : ∫ omega, quantity omega ∂mu <= threshold + bound * mu.real (tailAtLeast quantity threshold)","missing":[],"search":"integral_le_threshold_add_bound_mul_tailmass banditrlproof.lowerbounds.integral_le_threshold_add_bound_mul_tailmass integrating a nonnegative bounded random variable after splitting at one measurable upper-tail event. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegretReal_base_eq","label":"gaussianExpectedPseudoRegretReal_base_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPseudoRegretReal_base_eq","description":"In the base environment, expected pseudo-regret is exactly the common alternative gap times the sum of the alternative expected pull counts.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e22bb1a3053a","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5516,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:307"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianExpectedPseudoRegretReal_base_eq {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (hgap : 0 <= gap) (hgap_le : gap <= 1 / 2) (lastRound : Nat) : gaussianExpectedPseudoRegretReal algorithm (gaussianMinimaxBaseEnvironment gap hgap hgap_le) lastRound = gap * ∑ i : Fin m, gaussianExpectedPullCountReal algorithm (gaussianMinimaxBaseMean gap) lastRound i.succ","missing":[],"search":"gaussianexpectedpseudoregretreal_base_eq banditrlproof.lowerbounds.gaussianexpectedpseudoregretreal_base_eq in the base environment, expected pseudo-regret is exactly the common alternative gap times the sum of the alternative expected pull counts. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.horizon_mul_sqrt_div_eq_sqrt_mul","label":"horizon_mul_sqrt_div_eq_sqrt_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.horizon_mul_sqrt_div_eq_sqrt_mul","description":"Positive-horizon square-root normalization used by the exact Chapter 17 gap calibration.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-04cd179549ff","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5517,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:325"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem horizon_mul_sqrt_div_eq_sqrt_mul {horizon alternatives : Real} (hhorizon : 0 < horizon) (halternatives : 0 <= alternatives) : horizon * Real.sqrt (alternatives / horizon) = Real.sqrt (horizon * alternatives)","missing":[],"search":"horizon_mul_sqrt_div_eq_sqrt_mul banditrlproof.lowerbounds.horizon_mul_sqrt_div_eq_sqrt_mul positive-horizon square-root normalization used by the exact chapter 17 gap calibration. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sqrt_mul_mul_sqrt_div_eq_alternatives","label":"sqrt_mul_mul_sqrt_div_eq_alternatives","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sqrt_mul_mul_sqrt_div_eq_alternatives","description":"The two square-root factors in the information exponent cancel exactly.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-fe66299a6a60","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5518,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:346"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sqrt_mul_mul_sqrt_div_eq_alternatives {horizon alternatives : Real} (hhorizon : 0 < horizon) (halternatives : 0 <= alternatives) : Real.sqrt (alternatives * horizon) * Real.sqrt (alternatives / horizon) = alternatives","missing":[],"search":"sqrt_mul_mul_sqrt_div_eq_alternatives banditrlproof.lowerbounds.sqrt_mul_mul_sqrt_div_eq_alternatives the two square-root factors in the information exponent cancel exactly. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.horizon_mul_stochasticHighProbabilityGap_div_two","label":"horizon_mul_stochasticHighProbabilityGap_div_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.horizon_mul_stochasticHighProbabilityGap_div_two","description":"The chosen source gap times half the horizon is exactly the threshold displayed in Theorem 17.1, including the outer factor `1/4`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f2d01a9128b1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5519,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:370"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem horizon_mul_stochasticHighProbabilityGap_div_two (horizon alternatives : Nat) (B delta : Real) (hhorizon : 0 < horizon) (hB : 0 < B) : (horizon : Real) * stochasticHighProbabilityGap horizon alternatives B delta / 2 = stochasticHighProbabilityThreshold horizon alternatives B delta","missing":[],"search":"horizon_mul_stochastichighprobabilitygap_div_two banditrlproof.lowerbounds.horizon_mul_stochastichighprobabilitygap_div_two the chosen source gap times half the horizon is exactly the threshold displayed in theorem 17.1, including the outer factor `1/4`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticHighProbability_informationExponent_le_log","label":"stochasticHighProbability_informationExponent_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticHighProbability_informationExponent_le_log","description":"Exact scalar information calibration in the positive-logarithm branch of Theorem 17.1.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-c090534c34bf","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5520,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:413"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stochasticHighProbability_informationExponent_le_log (horizon alternatives : Nat) (B delta : Real) (hhorizon : 0 < horizon) (halternatives : 0 < alternatives) (hB : 0 < B) (hlog : 0 < Real.log (1 / (4 * delta))) : let gap := stochasticHighProbabilityGap horizon alternatives B delta (B * Real.sqrt ((alternatives : Real) * (horizon : Real)) / (gap * (alternatives : Real))) * (2 * gap ^ 2) <= Real.log (1 / (4 * delta))","missing":[],"search":"stochastichighprobability_informationexponent_le_log banditrlproof.lowerbounds.stochastichighprobability_informationexponent_le_log exact scalar information calibration in the positive-logarithm branch of theorem 17.1. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1_of_four_mul_delta_lt_one","label":"gaussianRandomPseudoRegret_ge_theorem17_1_of_four_mul_delta_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1_of_four_mul_delta_lt_one","description":"The positive-logarithm branch of Lattimore--Szepesvari Theorem 17.1. The algorithm is one common randomized nonanticipating history policy in the base and changed environments. The expected-regret premise is deterministic; the conclusion is a tail probability for random pseudo-regret.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-9f09cfce88f8","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5521,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:466"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianRandomPseudoRegret_ge_theorem17_1_of_four_mul_delta_lt_one {alternatives horizon : Nat} (halternatives : 0 < alternatives) (hhorizon : 0 < horizon) (B delta : Real) (hB : 0 < B) (hdelta : 0 < delta) (hfourDelta : 4 * delta < 1) (algorithm : Thompson.HistoryAlgorithm (Fin (alternatives + 1)) Real) (hExpected : forall environment : UnitGaussianBanditEnvironment (alternatives + 1), gaussianExpectedPseudoRegretReal algorithm environment (horizon - 1) <= B * Real.sqrt ((alternatives : Real) * (horizon : Real))) : exists environment : UnitGaussianBanditEnvironment (alternatives + 1), delta <= (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) (horizon - 1)).real (tailAtLeast (gaussianRandomPseudoRegret environment (horizon - 1)) (stochasticHighProbabilityThreshold horizon alternatives B delta))","missing":[],"search":"gaussianrandompseudoregret_ge_theorem17_1_of_four_mul_delta_lt_one banditrlproof.lowerbounds.gaussianrandompseudoregret_ge_theorem17_1_of_four_mul_delta_lt_one the positive-logarithm branch of lattimore--szepesvari theorem 17.1. the algorithm is one common randomized nonanticipating history policy in the base and changed environments. the expected-regret premise is deterministic; the conclusion is a tail probability for random pseudo-regret. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1_unitCube","label":"gaussianRandomPseudoRegret_ge_theorem17_1_unitCube","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1_unitCube","description":"*Lattimore--Szepesvari, Theorem 17.1.** A single policy whose deterministic expected pseudo-regret is uniformly at most `B * sqrt ((k-1) * n)` on the unit-cube subfamily has a unit-Gaussian environment whose random pseudo-regret exceeds the exact source threshold with probability at least `delta`. This internal theorem is stronger than the printed result because it assumes the uniform bound only on the unit-cube Gau…","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1ba15bcae6d4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5522,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:683"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianRandomPseudoRegret_ge_theorem17_1_unitCube {alternatives horizon : Nat} (halternatives : 0 < alternatives) (hhorizon : 0 < horizon) (B delta : Real) (hB : 0 < B) (hdelta : 0 < delta) (hdelta_one : delta < 1) (algorithm : Thompson.HistoryAlgorithm (Fin (alternatives + 1)) Real) (hExpected : forall environment : UnitGaussianBanditEnvironment (alternatives + 1), gaussianExpectedPseudoRegretReal algorithm environment (horizon - 1) <= B * Real.sqrt ((alternatives : Real) * (horizon : Real))) : exists environment : UnitGaussianBanditEnvironment (alternatives + 1), delta <= (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) (horizon - 1)).real (tailAtLeast (gaussianRandomPseudoRegret environment (horizon - 1)) (stochasticHighProbabilityThreshold horizon alternatives B delta))","missing":[],"search":"gaussianrandompseudoregret_ge_theorem17_1_unitcube banditrlproof.lowerbounds.gaussianrandompseudoregret_ge_theorem17_1_unitcube *lattimore--szepesvari, theorem 17.1.** a single policy whose deterministic expected pseudo-regret is uniformly at most `b * sqrt ((k-1) * n)` on the unit-cube subfamily has a unit-gaussian environment whose random pseudo-regret exceeds the exact source threshold with probability at least `delta`. this internal theorem is stronger than the printed result because it assumes the uniform bound only on the unit-cube gaussian subfamily used by the proof. the public source-class wrapper below restores the printed `e^k` premise. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1","label":"gaussianRandomPseudoRegret_ge_theorem17_1","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1","description":"Theorem 17.1 with the premise quantified over the full source class `E^k` of unit-variance Gaussian bandits whose gaps are at most one. The hard witness lies in the unit-cube subfamily, which is embedded into `E^k`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ef466200fc8d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5523,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:743"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"theorem gaussianRandomPseudoRegret_ge_theorem17_1 {alternatives horizon : Nat} (halternatives : 0 < alternatives) (hhorizon : 0 < horizon) (B delta : Real) (hB : 0 < B) (hdelta : 0 < delta) (hdelta_one : delta < 1) (algorithm : Thompson.HistoryAlgorithm (Fin (alternatives + 1)) Real) (hExpected : forall environment : GapOneGaussianBanditEnvironment (alternatives + 1), gapOneGaussianExpectedPseudoRegretReal algorithm environment (horizon - 1) <= B * Real.sqrt ((alternatives : Real) * (horizon : Real))) : exists environment : UnitGaussianBanditEnvironment (alternatives + 1), delta <= (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) (horizon - 1)).real (tailAtLeast (gaussianRandomPseudoRegret environment (horizon - 1)) (stochasticHighProbabilityThreshold horizon alternatives B delta))","missing":[],"search":"gaussianrandompseudoregret_ge_theorem17_1 banditrlproof.lowerbounds.gaussianrandompseudoregret_ge_theorem17_1 theorem 17.1 with the premise quantified over the full source class `e^k` of unit-variance gaussian bandits whose gaps are at most one. the hard witness lies in the unit-cube subfamily, which is embedded into `e^k`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticMinimax_sourceTerm_eq","label":"stochasticMinimax_sourceTerm_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticMinimax_sourceTerm_eq","description":"The square-root identity used to specialize Theorem 17.1 in Corollary 17.2. Keeping it separate makes the source constant `B = sqrt (2 log (1/(4δ)))` mechanically visible.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-78cc56396100","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5524,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:770"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stochasticMinimax_sourceTerm_eq {horizon alternatives : Nat} (delta : Real) (hlog : 0 < Real.log (1 / (4 * delta))) : (1 / Real.sqrt (2 * Real.log (1 / (4 * delta)))) * Real.sqrt ((alternatives : Real) * (horizon : Real)) * Real.log (1 / (4 * delta)) = Real.sqrt (((horizon : Real) * (alternatives : Real) / 2) * Real.log (1 / (4 * delta)))","missing":[],"search":"stochasticminimax_sourceterm_eq banditrlproof.lowerbounds.stochasticminimax_sourceterm_eq the square-root identity used to specialize theorem 17.1 in corollary 17.2. keeping it separate makes the source constant `b = sqrt (2 log (1/(4δ)))` mechanically visible. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold_at_minimax_scale","label":"stochasticHighProbabilityThreshold_at_minimax_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold_at_minimax_scale","description":"The Theorem 17.1 threshold at the source choice of `B` is exactly the Corollary 17.2 minimax threshold.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-9a79b3513a96","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5525,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:796"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stochasticHighProbabilityThreshold_at_minimax_scale {horizon alternatives : Nat} (delta : Real) (hlog : 0 < Real.log (1 / (4 * delta))) : stochasticHighProbabilityThreshold horizon alternatives (Real.sqrt (2 * Real.log (1 / (4 * delta)))) delta = stochasticMinimaxHighProbabilityThreshold horizon alternatives delta","missing":[],"search":"stochastichighprobabilitythreshold_at_minimax_scale banditrlproof.lowerbounds.stochastichighprobabilitythreshold_at_minimax_scale the theorem 17.1 threshold at the source choice of `b` is exactly the corollary 17.2 minimax threshold. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold_le_quarter_root","label":"stochasticMinimaxHighProbabilityThreshold_le_quarter_root","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold_le_quarter_root","description":"theorem stochasticMinimaxHighProbabilityThreshold_le_quarter_root {horizon alternatives : Nat} (delta : Real) (hlog : 0 <= Real.log (1 / (4 * delta))) : stochasticMinimaxHighProbabilityThreshold horizon alternatives delta <= (1 / 4 : Real) * Real.sqrt ((horizon : Real) * (alternatives : Real) * Real.log (1 / (4 * delta)))","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a3dda0d3df60","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5526,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:806"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem stochasticMinimaxHighProbabilityThreshold_le_quarter_root {horizon alternatives : Nat} (delta : Real) (hlog : 0 <= Real.log (1 / (4 * delta))) : stochasticMinimaxHighProbabilityThreshold horizon alternatives delta <= (1 / 4 : Real) * Real.sqrt ((horizon : Real) * (alternatives : Real) * Real.log (1 / (4 * delta)))","missing":[],"search":"stochasticminimaxhighprobabilitythreshold_le_quarter_root banditrlproof.lowerbounds.stochasticminimaxhighprobabilitythreshold_le_quarter_root theorem stochasticminimaxhighprobabilitythreshold_le_quarter_root {horizon alternatives : nat} (delta : real) (hlog : 0 <= real.log (1 / (4 * delta))) : stochasticminimaxhighprobabilitythreshold horizon alternatives delta <= (1 / 4 : real) * real.sqrt ((horizon : real) * (alternatives : real) * real.log (1 / (4 * delta))) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.minimax_expected_scale_identity","label":"minimax_expected_scale_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.minimax_expected_scale_identity","description":"theorem minimax_expected_scale_identity {horizon alternatives : Nat} (delta : Real) (hlog : 0 <= Real.log (1 / (4 * delta))) : Real.sqrt (2 * Real.log (1 / (4 * delta))) * Real.sqrt ((alternatives : Real) * (horizon : Real)) = Real.sqrt 2 * Real.sqrt ((horizon : Real) * (alternatives : Real) * Real.log (1 / (4 * delta)))","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-86abc024f642","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5527,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:824"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem minimax_expected_scale_identity {horizon alternatives : Nat} (delta : Real) (hlog : 0 <= Real.log (1 / (4 * delta))) : Real.sqrt (2 * Real.log (1 / (4 * delta))) * Real.sqrt ((alternatives : Real) * (horizon : Real)) = Real.sqrt 2 * Real.sqrt ((horizon : Real) * (alternatives : Real) * Real.log (1 / (4 * delta)))","missing":[],"search":"minimax_expected_scale_identity banditrlproof.lowerbounds.minimax_expected_scale_identity theorem minimax_expected_scale_identity {horizon alternatives : nat} (delta : real) (hlog : 0 <= real.log (1 / (4 * delta))) : real.sqrt (2 * real.log (1 / (4 * delta))) * real.sqrt ((alternatives : real) * (horizon : real)) = real.sqrt 2 * real.sqrt ((horizon : real) * (alternatives : real) * real.log (1 / (4 * delta))) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_exp_neg_rpow_inv_le_one","label":"integral_exp_neg_rpow_inv_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_exp_neg_rpow_inv_le_one","description":"The analytic inequality quoted in Corollary 17.3: `integral_0^infinity exp (-x^(1/p)) dx <= 1` for `0<p<1`. The change of variables identifies the integral with `Gamma (p+1)`; convexity of Gamma between one and two gives the bound.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-74b3b1545f10","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5528,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:842"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_neg_rpow_inv_le_one {p : Real} (hp : 0 < p) (hp_one : p < 1) : (∫ x : Real in Set.Ioi 0, Real.exp (-(x ^ (1 / p)))) <= 1","missing":[],"search":"integral_exp_neg_rpow_inv_le_one banditrlproof.lowerbounds.integral_exp_neg_rpow_inv_le_one the analytic inequality quoted in corollary 17.3: `integral_0^infinity exp (-x^(1/p)) dx <= 1` for `0<p<1`. the change of variables identifies the integral with `gamma (p+1)`; convexity of gamma between one and two gives the bound. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_le_scale_of_all_rpow_log_tail","label":"integral_le_scale_of_all_rpow_log_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_le_scale_of_all_rpow_log_tail","description":"Integrating an all-confidence stretched-exponential tail gives the first-moment bound used by Corollary 17.3. The strict source tail is retained at the caller; only its weak consequence is needed under the integral.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1eceb4a23164","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5529,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:873"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_le_scale_of_all_rpow_log_tail {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (quantity : Omega -> Real) (scale p : Real) (hscale : 0 < scale) (hp : 0 < p) (hp_one : p < 1) (hintegrable : Integrable quantity mu) (hnonneg : forall omega, 0 <= quantity omega) (htail : forall delta : Real, 0 < delta -> delta < 1 -> mu.real (tailAtLeast quantity (scale * (Real.log (1 / delta)) ^ p)) < delta) : (∫ omega, quantity omega ∂mu) <= scale","missing":[],"search":"integral_le_scale_of_all_rpow_log_tail banditrlproof.lowerbounds.integral_le_scale_of_all_rpow_log_tail integrating an all-confidence stretched-exponential tail gives the first-moment bound used by corollary 17.3. the strict source tail is retained at the caller; only its weak consequence is needed under the integral. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_corollary17_2","label":"gaussianRandomPseudoRegret_ge_corollary17_2","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_corollary17_2","description":"*Lattimore--Szepesvari, Corollary 17.2.** Under the exact side condition in Eq. (17.6), every policy has a unit-Gaussian instance whose random pseudo-regret reaches the minimax high-probability threshold with probability at least `delta`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-c093fd82ca2f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5530,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:953"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"theorem gaussianRandomPseudoRegret_ge_corollary17_2 {alternatives horizon : Nat} (halternatives : 0 < alternatives) (hhorizon : 0 < horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta < 1) (hside : (horizon : Real) * delta <= Real.sqrt ((horizon : Real) * (alternatives : Real) * Real.log (1 / (4 * delta)))) (algorithm : Thompson.HistoryAlgorithm (Fin (alternatives + 1)) Real) : exists environment : UnitGaussianBanditEnvironment (alternatives + 1), delta <= (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) (horizon - 1)).real (tailAtLeast (gaussianRandomPseudoRegret environment (horizon - 1)) (stochasticMinimaxHighProbabilityThreshold horizon alternatives delta))","missing":[],"search":"gaussianrandompseudoregret_ge_corollary17_2 banditrlproof.lowerbounds.gaussianrandompseudoregret_ge_corollary17_2 *lattimore--szepesvari, corollary 17.2.** under the exact side condition in eq. (17.6), every policy has a unit-gaussian instance whose random pseudo-regret reaches the minimax high-probability threshold with probability at least `delta`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.noUniformGaussianRandomPseudoRegretTail_corollary17_3","label":"noUniformGaussianRandomPseudoRegretTail_corollary17_3","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.noUniformGaussianRandomPseudoRegretTail_corollary17_3","description":"*Lattimore--Szepesvari, Corollary 17.3.** No single policy has a strictly smaller than `delta` random-pseudo-regret tail at every horizon, confidence level, and environment in the full gap-at-most-one Gaussian class when the logarithmic exponent lies in `(0,1)`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-7f3bbcfa6913","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5531,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1072"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"theorem noUniformGaussianRandomPseudoRegretTail_corollary17_3 {alternatives : Nat} (halternatives : 0 < alternatives) (p B : Real) (hp : 0 < p) (hp_one : p < 1) (hB : 0 < B) : ¬ exists algorithm : Thompson.HistoryAlgorithm (Fin (alternatives + 1)) Real, forall horizon : Nat, 0 < horizon -> forall delta : Real, 0 < delta -> delta < 1 -> forall environment : GapOneGaussianBanditEnvironment (alternatives + 1), (canonicalBanditHistoryMeasure algorithm (unitGaussianKernel environment.mean) (horizon - 1)).real (tailAtLeast (gapOneGaussianRandomPseudoRegret environment (horizon - 1)) (B * Real.sqrt ((alternatives : Real) * (horizon : Real)) * (Real.log (1 / delta)) ^ p)) < delta","missing":[],"search":"nouniformgaussianrandompseudoregrettail_corollary17_3 banditrlproof.lowerbounds.nouniformgaussianrandompseudoregrettail_corollary17_3 *lattimore--szepesvari, corollary 17.3.** no single policy has a strictly smaller than `delta` random-pseudo-regret tail at every horizon, confidence level, and environment in the full gap-at-most-one gaussian class when the logarithmic exponent lies in `(0,1)`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.exists_tailMass_ge_of_integral_ge","label":"exists_tailMass_ge_of_integral_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_tailMass_ge_of_integral_ge","description":"Claim 17.5 in its abstract first-moment form. If the average tail mass is at least `delta`, some deterministic instance has tail mass at least `delta`. The textbook suppresses the regularity needed to write the expectation. Lean makes it explicit as `Integrable tailMass Q`; `Q` is explicitly a probability measure.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3083ff541e28","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5532,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1247"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_tailMass_ge_of_integral_ge {Instance : Type*} [MeasurableSpace Instance] (Q : Measure Instance) [IsProbabilityMeasure Q] (tailMass : Instance -> Real) (delta : Real) (hIntegrable : Integrable tailMass Q) (hAverage : delta <= ∫ x, tailMass x ∂Q) : exists x, delta <= tailMass x","missing":[],"search":"exists_tailmass_ge_of_integral_ge banditrlproof.lowerbounds.exists_tailmass_ge_of_integral_ge claim 17.5 in its abstract first-moment form. if the average tail mass is at least `delta`, some deterministic instance has tail mass at least `delta`. the textbook suppresses the regularity needed to write the expectation. lean makes it explicit as `integrable tailmass q`; `q` is explicitly a probability measure. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_cdfTail_ge_of_integral_ge","label":"exists_cdfTail_ge_of_integral_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_cdfTail_ge_of_integral_ge","description":"Claim 17.5 specialized to the source notation `1 - F_x(u)`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5c2f374b84c4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5533,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1258"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"theorem exists_cdfTail_ge_of_integral_ge {Instance : Type*} [MeasurableSpace Instance] (Q : Measure Instance) [IsProbabilityMeasure Q] (cdf : Instance -> Real -> Real) (threshold delta : Real) (hIntegrable : Integrable (fun x => 1 - cdf x threshold) Q) (hAverage : delta <= ∫ x, 1 - cdf x threshold ∂Q) : exists x, delta <= 1 - cdf x threshold","missing":[],"search":"exists_cdftail_ge_of_integral_ge banditrlproof.lowerbounds.exists_cdftail_ge_of_integral_ge claim 17.5 specialized to the source notation `1 - f_x(u)`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.measureReal_diff_ge_delta","label":"measureReal_diff_ge_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measureReal_diff_ge_delta","description":"Probability subtraction used after Claims 17.6 and 17.7. If the pull-count event has probability at least `2 * delta` and the clipping event has probability at most `delta`, their good difference has probability at least `delta`. No measurability hypothesis is hidden: `Measure.real` is defined for all sets, and `le_measureReal_diff` is an outer-measure inequality.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-8701bd2a5e4a","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5534,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1276"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"theorem measureReal_diff_ge_delta {Omega : Type*} [MeasurableSpace Omega] (P : Measure Omega) [IsFiniteMeasure P] (pullSmall clippingBad : Set Omega) (delta : Real) (hPullSmall : 2 * delta <= P.real pullSmall) (hClippingBad : P.real clippingBad <= delta) : delta <= P.real (pullSmall \\ clippingBad)","missing":[],"search":"measurereal_diff_ge_delta banditrlproof.lowerbounds.measurereal_diff_ge_delta probability subtraction used after claims 17.6 and 17.7. if the pull-count event has probability at least `2 * delta` and the clipping event has probability at most `delta`, their good difference has probability at least `delta`. no measurability hypothesis is hidden: `measure.real` is defined for all sets, and `le_measurereal_diff` is an outer-measure inequality. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression","label":"adversarialRegretLowerExpression","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialRegretLowerExpression","description":"The deterministic lower expression on the right-hand side of Eq. (17.8), after writing the pull and clipping counts as real numbers.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f26682e6ecab","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5535,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1291"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialRegretLowerExpression (horizon pullCount clippingCount : Nat) (gap : Real) : Real","missing":[],"search":"adversarialregretlowerexpression banditrlproof.lowerbounds.adversarialregretlowerexpression the deterministic lower expression on the right-hand side of eq. (17.8), after writing the pull and clipping counts as real numbers. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression_ge_quarter","label":"adversarialRegretLowerExpression_ge_quarter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialRegretLowerExpression_ge_quarter","description":"If fewer than half of the rounds pull the distinguished arm and at most a quarter are clipped, the Eq. (17.8) lower expression is at least one quarter of `gap * horizon`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-fe0048f14f14","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5536,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1299"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"theorem adversarialRegretLowerExpression_ge_quarter (horizon pullCount clippingCount : Nat) (gap : Real) (hGap : 0 <= gap) (hPull : (pullCount : Real) <= (horizon : Real) / 2) (hClipping : (clippingCount : Real) <= (horizon : Real) / 4) : gap * ((horizon : Real) / 4) <= adversarialRegretLowerExpression horizon pullCount clippingCount gap","missing":[],"search":"adversarialregretlowerexpression_ge_quarter banditrlproof.lowerbounds.adversarialregretlowerexpression_ge_quarter if fewer than half of the rounds pull the distinguished arm and at most a quarter are clipped, the eq. (17.8) lower expression is at least one quarter of `gap * horizon`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.randomRegret_ge_quarter_of_clippingDecomposition","label":"randomRegret_ge_quarter_of_clippingDecomposition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.randomRegret_ge_quarter_of_clippingDecomposition","description":"The explicit transfer from Eq. (17.8) to the quarter-horizon regret threshold. The premise `hSource` is exactly the construction-specific part that Chapter 17 must still supply.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-2ddbbaa6b8b9","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5537,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1313"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"theorem randomRegret_ge_quarter_of_clippingDecomposition (horizon pullCount clippingCount : Nat) (gap randomRegret : Real) (hGap : 0 <= gap) (hPull : (pullCount : Real) <= (horizon : Real) / 2) (hClipping : (clippingCount : Real) <= (horizon : Real) / 4) (hSource : adversarialRegretLowerExpression horizon pullCount clippingCount gap <= randomRegret) : gap * ((horizon : Real) / 4) <= randomRegret","missing":[],"search":"randomregret_ge_quarter_of_clippingdecomposition banditrlproof.lowerbounds.randomregret_ge_quarter_of_clippingdecomposition the explicit transfer from eq. (17.8) to the quarter-horizon regret threshold. the premise `hsource` is exactly the construction-specific part that chapter 17 must still supply. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.clipUnitReward","label":"clipUnitReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.clipUnitReward","description":"Clipping to the reward interval `[0,1]`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e6385d317954","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5538,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1328"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def clipUnitReward (x : Real) : Real","missing":[],"search":"clipunitreward banditrlproof.lowerbounds.clipunitreward clipping to the reward interval `[0,1]`. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.clipUnitReward_mono","label":"clipUnitReward_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.clipUnitReward_mono","description":"theorem clipUnitReward_mono {x y : Real} (hxy : x <= y) : clipUnitReward x <= clipUnitReward y","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-4c99eb02f2fc","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5539,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1330"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipUnitReward_mono {x y : Real} (hxy : x <= y) : clipUnitReward x <= clipUnitReward y","missing":[],"search":"clipunitreward_mono banditrlproof.lowerbounds.clipunitreward_mono theorem clipunitreward_mono {x y : real} (hxy : x <= y) : clipunitreward x <= clipunitreward y theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.clipUnitReward_eq_self","label":"clipUnitReward_eq_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.clipUnitReward_eq_self","description":"theorem clipUnitReward_eq_self {x : Real} (hx0 : 0 <= x) (hx1 : x <= 1) : clipUnitReward x = x","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-00e35f2dac10","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5540,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1335"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem clipUnitReward_eq_self {x : Real} (hx0 : 0 <= x) (hx1 : x <= 1) : clipUnitReward x = x","missing":[],"search":"clipunitreward_eq_self banditrlproof.lowerbounds.clipunitreward_eq_self theorem clipunitreward_eq_self {x : real} (hx0 : 0 <= x) (hx1 : x <= 1) : clipunitreward x = x theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHardShift","label":"adversarialHardShift","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHardShift","description":"The source hard-family shift: arm zero receives `gap`, the distinguished nonzero arm receives `2*gap`, and every other arm receives zero.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f39850e0b401","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5541,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1341"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialHardShift {alternatives : Nat} (gap : Real) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) : Real","missing":[],"search":"adversarialhardshift banditrlproof.lowerbounds.adversarialhardshift the source hard-family shift: arm zero receives `gap`, the distinguished nonzero arm receives `2*gap`, and every other arm receives zero. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward","label":"adversarialClippedGaussianReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedGaussianReward","description":"The exact pre-sampled construction underlying Theorem 17.4. One scalar `eta t` is shared by every arm at round `t`; hence the arm rewards are correlated within a round. Independence and standard-Gaussian assumptions belong to the law of the path `eta`, not to this pathwise definition.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3b845e2f3f68","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5542,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1350"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialClippedGaussianReward {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (distinguished : Fin alternatives) (t : Fin horizon) (arm : Fin (alternatives + 1)) : Real","missing":[],"search":"adversarialclippedgaussianreward banditrlproof.lowerbounds.adversarialclippedgaussianreward the exact pre-sampled construction underlying theorem 17.4. one scalar `eta t` is shared by every arm at round `t`; hence the arm rewards are correlated within a round. independence and standard-gaussian assumptions belong to the law of the path `eta`, not to this pathwise definition. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw","label":"adversarialCenteredNoiseLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw","description":"The IID centered-Gaussian path law used by the equivalent centered form of the source construction. Adding `1/2` in `adversarialClippedGaussianReward` makes `1/2 + eta t` have the source law `N(1/2, sigma^2)`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-2402a4c5afa4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5543,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1361"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialCenteredNoiseLaw (horizon : Nat) (sigma : Real) : Measure (Fin horizon -> Real)","missing":[],"search":"adversarialcenterednoiselaw banditrlproof.lowerbounds.adversarialcenterednoiselaw the iid centered-gaussian path law used by the equivalent centered form of the source construction. adding `1/2` in `adversarialclippedgaussianreward` makes `1/2 + eta t` have the source law `n(1/2, sigma^2)`. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedArmLaw","label":"adversarialClippedArmLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedArmLaw","description":"One observed arm's marginal in the clipped Gaussian construction. The joint reward matrix still uses a shared noise coordinate for all arms.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-c24f2f8f20fb","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5544,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1368"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialClippedArmLaw (sigma shift : Real) : Measure Real","missing":[],"search":"adversarialclippedarmlaw banditrlproof.lowerbounds.adversarialclippedarmlaw one observed arm's marginal in the clipped gaussian construction. the joint reward matrix still uses a shared noise coordinate for all arms. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClippedArmMap","label":"measurable_adversarialClippedArmMap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialClippedArmMap","description":"theorem measurable_adversarialClippedArmMap (shift : Real) : Measurable (fun x : Real => clipUnitReward (1 / 2 + x + shift))","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-162735b6a104","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5545,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1372"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialClippedArmMap (shift : Real) : Measurable (fun x : Real => clipUnitReward (1 / 2 + x + shift))","missing":[],"search":"measurable_adversarialclippedarmmap banditrlproof.lowerbounds.measurable_adversarialclippedarmmap theorem measurable_adversarialclippedarmmap (shift : real) : measurable (fun x : real => clipunitreward (1 / 2 + x + shift)) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedKernel","label":"adversarialClippedKernel","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedKernel","description":"Finite-arm observation kernel, parameterized also for the base family used in Claim 17.6.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-75d356229e73","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5546,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1385"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable abbrev adversarialClippedKernel {K : Nat} (sigma : Real) (shift : Fin K -> Real) : Kernel (Fin K) Real","missing":[],"search":"adversarialclippedkernel banditrlproof.lowerbounds.adversarialclippedkernel finite-arm observation kernel, parameterized also for the base family used in claim 17.6. abbreviation compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw","label":"adversarialClippedHistoryLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedHistoryLaw","description":"The observable history law under a fixed randomized policy and the clipped Gaussian feedback kernel. This is not a law on full reward matrices.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-8b3380a0da30","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5547,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1396"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialClippedHistoryLaw {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) (lastRound : Nat) : Measure (History.FinitePairHistory (Fin K) Real lastRound)","missing":[],"search":"adversarialclippedhistorylaw banditrlproof.lowerbounds.adversarialclippedhistorylaw the observable history law under a fixed randomized policy and the clipped gaussian feedback kernel. this is not a law on full reward matrices. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_reward_marginal","label":"adversarialCenteredNoiseLaw_reward_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_reward_marginal","description":"A coordinate of the shared-noise matrix has exactly the arm marginal used by the observation kernel. This retains the original product noise law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-fcc56c9aa7a4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5548,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1411"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialCenteredNoiseLaw_reward_marginal {horizon alternatives : Nat} (sigma gap : Real) (distinguished : Fin alternatives) (t : Fin horizon) (arm : Fin (alternatives + 1)) : (adversarialCenteredNoiseLaw horizon sigma).map (fun eta => adversarialClippedGaussianReward eta gap distinguished t arm) = adversarialClippedArmLaw sigma (adversarialHardShift gap distinguished arm)","missing":[],"search":"adversarialcenterednoiselaw_reward_marginal banditrlproof.lowerbounds.adversarialcenterednoiselaw_reward_marginal a coordinate of the shared-noise matrix has exactly the arm marginal used by the observation kernel. this retains the original product noise law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipHistory","label":"adversarialClipHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipHistory","description":"Clip the observed reward coordinates while retaining every action.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1c5d6c747f16","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5549,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1428"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialClipHistory {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : History.FinitePairHistory (Fin K) Real n","missing":[],"search":"adversarialcliphistory banditrlproof.lowerbounds.adversarialcliphistory clip the observed reward coordinates while retaining every action. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClipHistory","label":"measurable_adversarialClipHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialClipHistory","description":"theorem measurable_adversarialClipHistory {K : Nat} (n : Nat) : Measurable (adversarialClipHistory (K := K) n)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-235ae44cb0b1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5550,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1433"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialClipHistory {K : Nat} (n : Nat) : Measurable (adversarialClipHistory (K := K) n)","missing":[],"search":"measurable_adversarialcliphistory banditrlproof.lowerbounds.measurable_adversarialcliphistory theorem measurable_adversarialcliphistory {k : nat} (n : nat) : measurable (adversarialcliphistory (k := k) n) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipHistoryAlgorithm","label":"adversarialClipHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipHistoryAlgorithm","description":"Lift the original policy to unbounded observations by feeding it only the clipped history. This construction is independent of the hard instance.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a92c435cebdb","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5551,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1440"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialClipHistoryAlgorithm {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) : Thompson.HistoryAlgorithm (Fin K) Real where","missing":[],"search":"adversarialcliphistoryalgorithm banditrlproof.lowerbounds.adversarialcliphistoryalgorithm lift the original policy to unbounded observations by feeding it only the clipped history. this construction is independent of the hard instance. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipHistoryAlgorithm_policy_apply","label":"adversarialClipHistoryAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipHistoryAlgorithm_policy_apply","description":"theorem adversarialClipHistoryAlgorithm_policy_apply {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (adversarialClipHistoryAlgorithm algorithm).policy n history = algorithm.policy n (adversarialClipHistory n history)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-14c82ea9c43c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5552,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1447"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipHistoryAlgorithm_policy_apply {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (adversarialClipHistoryAlgorithm algorithm).policy n history = algorithm.policy n (adversarialClipHistory n history)","missing":[],"search":"adversarialcliphistoryalgorithm_policy_apply banditrlproof.lowerbounds.adversarialcliphistoryalgorithm_policy_apply theorem adversarialcliphistoryalgorithm_policy_apply {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (n : nat) (history : history.finitepairhistory (fin k) real n) : (adversarialcliphistoryalgorithm algorithm).policy n history = algorithm.policy n (adversarialcliphistory n history) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipHistory_pullCount","label":"adversarialClipHistory_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipHistory_pullCount","description":"theorem adversarialClipHistory_pullCount {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountENNReal n (adversarialClipHistory n history) arm = finiteHistoryPullCountENNReal n history arm","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-dab38b39e72f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5553,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1454"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipHistory_pullCount {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountENNReal n (adversarialClipHistory n history) arm = finiteHistoryPullCountENNReal n history arm","missing":[],"search":"adversarialcliphistory_pullcount banditrlproof.lowerbounds.adversarialcliphistory_pullcount theorem adversarialcliphistory_pullcount {k : nat} (n : nat) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : finitehistorypullcountennreal n (adversarialcliphistory n history) arm = finitehistorypullcountennreal n history arm theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipHistory_pullCountReal","label":"adversarialClipHistory_pullCountReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipHistory_pullCountReal","description":"theorem adversarialClipHistory_pullCountReal {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountReal n (adversarialClipHistory n history) arm = finiteHistoryPullCountReal n history arm","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-d345cf5b7bc9","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5554,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1466"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipHistory_pullCountReal {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountReal n (adversarialClipHistory n history) arm = finiteHistoryPullCountReal n history arm","missing":[],"search":"adversarialcliphistory_pullcountreal banditrlproof.lowerbounds.adversarialcliphistory_pullcountreal theorem adversarialcliphistory_pullcountreal {k : nat} (n : nat) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : finitehistorypullcountreal n (adversarialcliphistory n history) arm = finitehistorypullcountreal n history arm theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialUnclippedKernel","label":"adversarialUnclippedKernel","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialUnclippedKernel","description":"Unclipped feedback law used with the lifted policy.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e02bff8309e6","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5555,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1474"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable abbrev adversarialUnclippedKernel {K : Nat} (sigma : Real) (shift : Fin K -> Real) : Kernel (Fin K) Real","missing":[],"search":"adversarialunclippedkernel banditrlproof.lowerbounds.adversarialunclippedkernel unclipped feedback law used with the lifted policy. abbreviation compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedKernel_eq_map","label":"adversarialClippedKernel_eq_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedKernel_eq_map","description":"theorem adversarialClippedKernel_eq_map {K : Nat} (sigma : Real) (shift : Fin K -> Real) : adversarialClippedKernel sigma shift = (adversarialUnclippedKernel sigma shift).map clipUnitReward","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-af1d89d19422","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5556,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1485"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedKernel_eq_map {K : Nat} (sigma : Real) (shift : Fin K -> Real) : adversarialClippedKernel sigma shift = (adversarialUnclippedKernel sigma shift).map clipUnitReward","missing":[],"search":"adversarialclippedkernel_eq_map banditrlproof.lowerbounds.adversarialclippedkernel_eq_map theorem adversarialclippedkernel_eq_map {k : nat} (sigma : real) (shift : fin k -> real) : adversarialclippedkernel sigma shift = (adversarialunclippedkernel sigma shift).map clipunitreward theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipped_initialPairLaw","label":"adversarialClipped_initialPairLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipped_initialPairLaw","description":"Initial action/reward law transport for the original and lifted policy.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5d351aa7c4cf","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5557,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1500"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipped_initialPairLaw {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) : algorithm.initialAction ⊗ₘ adversarialClippedKernel sigma shift = ((adversarialClipHistoryAlgorithm algorithm).initialAction ⊗ₘ adversarialUnclippedKernel sigma shift).map (Prod.map id clipUnitReward)","missing":[],"search":"adversarialclipped_initialpairlaw banditrlproof.lowerbounds.adversarialclipped_initialpairlaw initial action/reward law transport for the original and lifted policy. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_zero","label":"adversarialClippedHistoryLaw_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_zero","description":"The exact observable history transport at the first observation.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-d8e58a042b14","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5558,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1510"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedHistoryLaw_zero {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) : adversarialClippedHistoryLaw algorithm sigma shift 0 = (canonicalBanditHistoryMeasure (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma shift) 0).map (adversarialClipHistory 0)","missing":[],"search":"adversarialclippedhistorylaw_zero banditrlproof.lowerbounds.adversarialclippedhistorylaw_zero the exact observable history transport at the first observation. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipped_historyStepLaw","label":"adversarialClipped_historyStepLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipped_historyStepLaw","description":"Pointwise transport of the next action/reward pair after a history.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f8984d2166cf","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5559,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1528"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipped_historyStepLaw {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Thompson.historyStepKernel algorithm (stationaryBanditHistoryEnvironment (adversarialClippedKernel sigma shift)) n (adversarialClipHistory n history) = (Thompson.historyStepKernel (adversarialClipHistoryAlgorithm algorithm) (stationaryBanditHistoryEnvironment (adversarialUnclippedKernel sigma shift)) n history).map (Prod.map id clipUnitReward)","missing":[],"search":"adversarialclipped_historysteplaw banditrlproof.lowerbounds.adversarialclipped_historysteplaw pointwise transport of the next action/reward pair after a history. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipped_prefixStepLaw","label":"adversarialClipped_prefixStepLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipped_prefixStepLaw","description":"Integrate the pointwise next-pair transport over any prefix law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-affc2dc2b9e1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5560,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1550"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipped_prefixStepLaw {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) (n : Nat) (P : Measure (History.FinitePairHistory (Fin K) Real n)) [IsProbabilityMeasure P] : P.map (adversarialClipHistory n) ⊗ₘ Thompson.historyStepKernel algorithm (stationaryBanditHistoryEnvironment (adversarialClippedKernel sigma shift)) n = (P ⊗ₘ Thompson.historyStepKernel (adversarialClipHistoryAlgorithm algorithm) (stationaryBanditHistoryEnvironment (adversarialUnclippedKernel sigma shift)) n).map (Prod.map (adversarialClipHistory n) (Prod.map id clipUnitReward))","missing":[],"search":"adversarialclipped_prefixsteplaw banditrlproof.lowerbounds.adversarialclipped_prefixsteplaw integrate the pointwise next-pair transport over any prefix law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_eq_map","label":"adversarialClippedHistoryLaw_eq_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_eq_map","description":"Exact transport of the entire finite observed history under the same original policy and its clipped-history lift.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ce502eb70815","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5561,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1583"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedHistoryLaw_eq_map {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) (n : Nat) : adversarialClippedHistoryLaw algorithm sigma shift n = (canonicalBanditHistoryMeasure (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma shift) n).map (adversarialClipHistory n)","missing":[],"search":"adversarialclippedhistorylaw_eq_map banditrlproof.lowerbounds.adversarialclippedhistorylaw_eq_map exact transport of the entire finite observed history under the same original policy and its clipped-history lift. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_pullSmall","label":"adversarialClippedHistoryLaw_pullSmall","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_pullSmall","description":"Pull-count events are preserved by the full history transport.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a3a9c339c6a8","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5562,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1614"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedHistoryLaw_pullSmall {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) (n : Nat) (arm : Fin K) (threshold : Real) : (adversarialClippedHistoryLaw algorithm sigma shift n) {h | finiteHistoryPullCountReal n h arm < threshold} = (canonicalBanditHistoryMeasure (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma shift) n) {h | finiteHistoryPullCountReal n h arm < threshold}","missing":[],"search":"adversarialclippedhistorylaw_pullsmall banditrlproof.lowerbounds.adversarialclippedhistorylaw_pullsmall pull-count events are preserved by the full history transport. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialUnclippedKernel_apply","label":"adversarialUnclippedKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialUnclippedKernel_apply","description":"Identify the unbounded observation marginal as the exact shifted Gaussian.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a3cdc3c7e5da","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5563,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1631"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialUnclippedKernel_apply {K : Nat} (sigma : Real) (shift : Fin K -> Real) (arm : Fin K) : adversarialUnclippedKernel sigma shift arm = gaussianReal (1 / 2 + shift arm) ⟨sigma ^ 2, sq_nonneg sigma⟩","missing":[],"search":"adversarialunclippedkernel_apply banditrlproof.lowerbounds.adversarialunclippedkernel_apply identify the unbounded observation marginal as the exact shifted gaussian. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_common_scale","label":"klDiv_gaussianReal_common_scale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_gaussianReal_common_scale","description":"Equal nonzero variance Gaussian KL, obtained by scaling the unit variance theorem through a measurable equivalence. Mathlib candidate.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-7825434292a5","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5564,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1644"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_gaussianReal_common_scale (sigma mu nu : Real) (hs : sigma ≠ 0) : InformationTheory.klDiv (gaussianReal mu ⟨sigma ^ 2, sq_nonneg sigma⟩) (gaussianReal nu ⟨sigma ^ 2, sq_nonneg sigma⟩) = ENNReal.ofReal ((mu - nu) ^ 2 / (2 * sigma ^ 2))","missing":[],"search":"kldiv_gaussianreal_common_scale banditrlproof.lowerbounds.kldiv_gaussianreal_common_scale equal nonzero variance gaussian kl, obtained by scaling the unit variance theorem through a measurable equivalence. mathlib candidate. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_adversarialUnclippedKernel","label":"klDiv_adversarialUnclippedKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_adversarialUnclippedKernel","description":"Directed per-arm information for the unbounded hard-family observations.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ea3eeaefac2f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5565,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1661"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_adversarialUnclippedKernel {K : Nat} (sigma : Real) (hs : sigma ≠ 0) (shift referenceShift : Fin K -> Real) (arm : Fin K) : InformationTheory.klDiv (adversarialUnclippedKernel sigma shift arm) (adversarialUnclippedKernel sigma referenceShift arm) = ENNReal.ofReal ((shift arm - referenceShift arm) ^ 2 / (2 * sigma ^ 2))","missing":[],"search":"kldiv_adversarialunclippedkernel banditrlproof.lowerbounds.kldiv_adversarialunclippedkernel directed per-arm information for the unbounded hard-family observations. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_adversarialUnclipped_base_changed_history","label":"klDiv_adversarialUnclipped_base_changed_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_adversarialUnclipped_base_changed_history","description":"Exact first-law pull-count information identity for Claim 17.6's unclipped hard family. The same lifted policy is used on both sides.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-bd9089a4b538","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5566,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1674"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem klDiv_adversarialUnclipped_base_changed_history {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (hs : sigma ≠ 0) (i : Fin m) (n : Nat) : InformationTheory.klDiv (canonicalBanditHistoryMeasure (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma (gaussianMinimaxBaseMean gap)) n) (canonicalBanditHistoryMeasure (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma (adversarialHardShift gap i)) n) = canonicalRealizedExpectedPullCountThrough (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma (gaussianMinimaxBaseMean gap)) n i.succ * ENNReal.ofReal (2 * gap ^ 2 / sigma ^ 2)","missing":[],"search":"kldiv_adversarialunclipped_base_changed_history banditrlproof.lowerbounds.kldiv_adversarialunclipped_base_changed_history exact first-law pull-count information identity for claim 17.6's unclipped hard family. the same lifted policy is used on both sides. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClaim17_6Gap","label":"adversarialClaim17_6Gap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClaim17_6Gap","description":"The exact source tuning from Claim 17.6, with `alternatives = k - 1`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-198a66b8b6ac","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5567,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1707"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialClaim17_6Gap (horizon alternatives : Nat) (sigma delta : Real) : Real","missing":[],"search":"adversarialclaim17_6gap banditrlproof.lowerbounds.adversarialclaim17_6gap the exact source tuning from claim 17.6, with `alternatives = k - 1`. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift","label":"adversarialFullHardShift","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullHardShift","description":"Include the base instance as arm zero, as required by Claim 17.6.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-15dbdc829c4b","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5568,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1714"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialFullHardShift {m : Nat} (gap : Real) (distinguished arm : Fin (m + 1)) : Real","missing":[],"search":"adversarialfullhardshift banditrlproof.lowerbounds.adversarialfullhardshift include the base instance as arm zero, as required by claim 17.6. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_zero","label":"adversarialFullHardShift_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullHardShift_zero","description":"theorem adversarialFullHardShift_zero {m : Nat} (gap : Real) : adversarialFullHardShift (m := m) gap 0 = gaussianMinimaxBaseMean gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ad499679cfe5","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5569,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1718"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullHardShift_zero {m : Nat} (gap : Real) : adversarialFullHardShift (m := m) gap 0 = gaussianMinimaxBaseMean gap","missing":[],"search":"adversarialfullhardshift_zero banditrlproof.lowerbounds.adversarialfullhardshift_zero theorem adversarialfullhardshift_zero {m : nat} (gap : real) : adversarialfullhardshift (m := m) gap 0 = gaussianminimaxbasemean gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_succ","label":"adversarialFullHardShift_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullHardShift_succ","description":"theorem adversarialFullHardShift_succ {m : Nat} (gap : Real) (i : Fin m) : adversarialFullHardShift gap i.succ = adversarialHardShift gap i","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-cf984004dedb","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5570,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1723"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullHardShift_succ {m : Nat} (gap : Real) (i : Fin m) : adversarialFullHardShift gap i.succ = adversarialHardShift gap i","missing":[],"search":"adversarialfullhardshift_succ banditrlproof.lowerbounds.adversarialfullhardshift_succ theorem adversarialfullhardshift_succ {m : nat} (gap : real) (i : fin m) : adversarialfullhardshift gap i.succ = adversarialhardshift gap i theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClaim17_6Gap_information_calibration","label":"adversarialClaim17_6Gap_information_calibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClaim17_6Gap_information_calibration","description":"Source gap tuning cancels the least-arm information bound exactly.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-598a10cc50e5","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5571,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1727"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClaim17_6Gap_information_calibration {horizon m : Nat} (hn : 0 < horizon) (hm : 0 < m) (sigma delta : Real) (hs : sigma ≠ 0) (hd : 0 < delta) (hd8 : delta < 1 / 8) : ((horizon : Real) / m) * (2 * (adversarialClaim17_6Gap horizon m sigma delta) ^ 2 / sigma ^ 2) = Real.log (1 / (8 * delta))","missing":[],"search":"adversarialclaim17_6gap_information_calibration banditrlproof.lowerbounds.adversarialclaim17_6gap_information_calibration source gap tuning cancels the least-arm information bound exactly. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_adversarialUnclipped_expectedPulls","label":"sum_adversarialUnclipped_expectedPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_adversarialUnclipped_expectedPulls","description":"Conservation of expected pulls for the lifted policy's unbounded law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5e9f93fb1392","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5572,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1744"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_adversarialUnclipped_expectedPulls {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (sigma : Real) (shift : Fin K -> Real) (n : Nat) : (∑ arm : Fin K, canonicalRealizedExpectedPullCountThrough (adversarialClipHistoryAlgorithm algorithm) (adversarialUnclippedKernel sigma shift) n arm) = n + 1","missing":[],"search":"sum_adversarialunclipped_expectedpulls banditrlproof.lowerbounds.sum_adversarialunclipped_expectedpulls conservation of expected pulls for the lifted policy's unbounded law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistory_pull_le_half_claim17_6","label":"adversarialClippedHistory_pull_le_half_claim17_6","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedHistory_pull_le_half_claim17_6","description":"Corrected Claim 17.6: the source strict inequality must be non-strict. The witness ranges over the base instance as well as all changed instances.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-70fb79a495aa","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5573,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1760"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedHistory_pull_le_half_claim17_6 {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (sigma delta : Real) (hs : sigma ≠ 0) (hd : 0 < delta) (hd8 : delta < 1 / 8) : ∃ arm : Fin (m + 1), 2 * delta <= (adversarialClippedHistoryLaw algorithm sigma (adversarialFullHardShift (adversarialClaim17_6Gap (n + 1) m sigma delta) arm) n).real {h | finiteHistoryPullCountReal n h arm <= ((n + 1 : Nat) : Real) / 2}","missing":[],"search":"adversarialclippedhistory_pull_le_half_claim17_6 banditrlproof.lowerbounds.adversarialclippedhistory_pull_le_half_claim17_6 corrected claim 17.6: the source strict inequality must be non-strict. the witness ranges over the base instance as well as all changed instances. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_distinguished","label":"adversarialHardShift_distinguished","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHardShift_distinguished","description":"theorem adversarialHardShift_distinguished {alternatives : Nat} (gap : Real) (distinguished : Fin alternatives) : adversarialHardShift gap distinguished distinguished.succ = 2 * gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-0280ecea9169","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5574,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1854"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHardShift_distinguished {alternatives : Nat} (gap : Real) (distinguished : Fin alternatives) : adversarialHardShift gap distinguished distinguished.succ = 2 * gap","missing":[],"search":"adversarialhardshift_distinguished banditrlproof.lowerbounds.adversarialhardshift_distinguished theorem adversarialhardshift_distinguished {alternatives : nat} (gap : real) (distinguished : fin alternatives) : adversarialhardshift gap distinguished distinguished.succ = 2 * gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_nonneg","label":"adversarialHardShift_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHardShift_nonneg","description":"theorem adversarialHardShift_nonneg {alternatives : Nat} (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) : 0 <= adversarialHardShift gap distinguished arm","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-2f2268d8a66e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5575,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1859"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHardShift_nonneg {alternatives : Nat} (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) : 0 <= adversarialHardShift gap distinguished arm","missing":[],"search":"adversarialhardshift_nonneg banditrlproof.lowerbounds.adversarialhardshift_nonneg theorem adversarialhardshift_nonneg {alternatives : nat} (gap : real) (hgap : 0 <= gap) (distinguished : fin alternatives) (arm : fin (alternatives + 1)) : 0 <= adversarialhardshift gap distinguished arm theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_le_two_mul","label":"adversarialHardShift_le_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHardShift_le_two_mul","description":"theorem adversarialHardShift_le_two_mul {alternatives : Nat} (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) : adversarialHardShift gap distinguished arm <= 2 * gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3a4620fae2a7","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5576,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1869"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHardShift_le_two_mul {alternatives : Nat} (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) : adversarialHardShift gap distinguished arm <= 2 * gap","missing":[],"search":"adversarialhardshift_le_two_mul banditrlproof.lowerbounds.adversarialhardshift_le_two_mul theorem adversarialhardshift_le_two_mul {alternatives : nat} (gap : real) (hgap : 0 <= gap) (distinguished : fin alternatives) (arm : fin (alternatives + 1)) : adversarialhardshift gap distinguished arm <= 2 * gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_le_gap_of_ne","label":"adversarialHardShift_le_gap_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHardShift_le_gap_of_ne","description":"theorem adversarialHardShift_le_gap_of_ne {alternatives : Nat} (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) (hne : arm ≠ distinguished.succ) : adversarialHardShift gap distinguished arm <= gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-32b3642178c9","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5577,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1881"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHardShift_le_gap_of_ne {alternatives : Nat} (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (arm : Fin (alternatives + 1)) (hne : arm ≠ distinguished.succ) : adversarialHardShift gap distinguished arm <= gap","missing":[],"search":"adversarialhardshift_le_gap_of_ne banditrlproof.lowerbounds.adversarialhardshift_le_gap_of_ne theorem adversarialhardshift_le_gap_of_ne {alternatives : nat} (gap : real) (hgap : 0 <= gap) (distinguished : fin alternatives) (arm : fin (alternatives + 1)) (hne : arm ≠ distinguished.succ) : adversarialhardshift gap distinguished arm <= gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward_distinguished_mono","label":"adversarialClippedGaussianReward_distinguished_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedGaussianReward_distinguished_mono","description":"theorem adversarialClippedGaussianReward_distinguished_mono {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (t : Fin horizon) (arm : Fin (alternatives + 1)) : adversarialClippedGaussianReward eta gap distinguished t arm <= adversarialClippedGaussianReward eta gap distinguished t distinguished.succ","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-10ad45517924","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5578,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1889"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedGaussianReward_distinguished_mono {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (t : Fin horizon) (arm : Fin (alternatives + 1)) : adversarialClippedGaussianReward eta gap distinguished t arm <= adversarialClippedGaussianReward eta gap distinguished t distinguished.succ","missing":[],"search":"adversarialclippedgaussianreward_distinguished_mono banditrlproof.lowerbounds.adversarialclippedgaussianreward_distinguished_mono theorem adversarialclippedgaussianreward_distinguished_mono {horizon alternatives : nat} (eta : fin horizon -> real) (gap : real) (hgap : 0 <= gap) (distinguished : fin alternatives) (t : fin horizon) (arm : fin (alternatives + 1)) : adversarialclippedgaussianreward eta gap distinguished t arm <= adversarialclippedgaussianreward eta gap distinguished t distinguished.succ theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward_gap_of_not_clipped","label":"adversarialClippedGaussianReward_gap_of_not_clipped","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippedGaussianReward_gap_of_not_clipped","description":"Away from clipping, the distinguished arm beats every other arm by at least `gap`. This is the pointwise engine of Eq. (17.8).","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-852e83d01a20","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5579,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1904"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippedGaussianReward_gap_of_not_clipped {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (t : Fin horizon) (arm : Fin (alternatives + 1)) (hne : arm ≠ distinguished.succ) (hgood : |eta t| < 1 / 2 - 2 * gap) : gap <= adversarialClippedGaussianReward eta gap distinguished t distinguished.succ - adversarialClippedGaussianReward eta gap distinguished t arm","missing":[],"search":"adversarialclippedgaussianreward_gap_of_not_clipped banditrlproof.lowerbounds.adversarialclippedgaussianreward_gap_of_not_clipped away from clipping, the distinguished arm beats every other arm by at least `gap`. this is the pointwise engine of eq. (17.8). theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialPullCountReal","label":"adversarialPullCountReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialPullCountReal","description":"def adversarialPullCountReal {horizon alternatives : Nat} (actions : Fin horizon -> Fin (alternatives + 1)) (distinguished : Fin alternatives) : Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-007d0ad46a84","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5580,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1938"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialPullCountReal {horizon alternatives : Nat} (actions : Fin horizon -> Fin (alternatives + 1)) (distinguished : Fin alternatives) : Real","missing":[],"search":"adversarialpullcountreal banditrlproof.lowerbounds.adversarialpullcountreal def adversarialpullcountreal {horizon alternatives : nat} (actions : fin horizon -> fin (alternatives + 1)) (distinguished : fin alternatives) : real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippingCountReal","label":"adversarialClippingCountReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippingCountReal","description":"def adversarialClippingCountReal {horizon : Nat} (eta : Fin horizon -> Real) (gap : Real) : Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-7f7278e23f6d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5581,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1944"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialClippingCountReal {horizon : Nat} (eta : Fin horizon -> Real) (gap : Real) : Real","missing":[],"search":"adversarialclippingcountreal banditrlproof.lowerbounds.adversarialclippingcountreal def adversarialclippingcountreal {horizon : nat} (eta : fin horizon -> real) (gap : real) : real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipIndicator","label":"adversarialClipIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipIndicator","description":"The Bernoulli indicator of a clipped round in the construction for Theorem 17.4.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-8108056c1ace","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5582,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1950"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialClipIndicator (gap x : Real) : Real","missing":[],"search":"adversarialclipindicator banditrlproof.lowerbounds.adversarialclipindicator the bernoulli indicator of a clipped round in the construction for theorem 17.4. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClipIndicator","label":"measurable_adversarialClipIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialClipIndicator","description":"theorem measurable_adversarialClipIndicator (gap : Real) : Measurable (adversarialClipIndicator gap)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1d6c81ef4f6d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5583,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1953"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialClipIndicator (gap : Real) : Measurable (adversarialClipIndicator gap)","missing":[],"search":"measurable_adversarialclipindicator banditrlproof.lowerbounds.measurable_adversarialclipindicator theorem measurable_adversarialclipindicator (gap : real) : measurable (adversarialclipindicator gap) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClipIndicator_mem_Icc","label":"adversarialClipIndicator_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClipIndicator_mem_Icc","description":"theorem adversarialClipIndicator_mem_Icc (gap x : Real) : adversarialClipIndicator gap x ∈ Set.Icc 0 1","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-9ada07ad3a83","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5584,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1959"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClipIndicator_mem_Icc (gap x : Real) : adversarialClipIndicator gap x ∈ Set.Icc 0 1","missing":[],"search":"adversarialclipindicator_mem_icc banditrlproof.lowerbounds.adversarialclipindicator_mem_icc theorem adversarialclipindicator_mem_icc (gap x : real) : adversarialclipindicator gap x ∈ set.icc 0 1 theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianReal_tenth_abs_quarter_le_eighth","label":"gaussianReal_tenth_abs_quarter_le_eighth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianReal_tenth_abs_quarter_le_eighth","description":"The one-dimensional Gaussian tail estimate used in Claim 17.7. It is proved from Mathlib's sub-Gaussian Chernoff bound; the final numerical step is the degree-five lower Taylor bound for `exp (25/8)`.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-71587aa9d86c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5585,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:1967"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianReal_tenth_abs_quarter_le_eighth : (gaussianReal 0 ⟨(1 / 10 : Real) ^ 2, sq_nonneg _⟩).real {x : Real | (1 / 4 : Real) <= |x|} <= 1 / 8","missing":[],"search":"gaussianreal_tenth_abs_quarter_le_eighth banditrlproof.lowerbounds.gaussianreal_tenth_abs_quarter_le_eighth the one-dimensional gaussian tail estimate used in claim 17.7. it is proved from mathlib's sub-gaussian chernoff bound; the final numerical step is the degree-five lower taylor bound for `exp (25/8)`. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integral_adversarialClipIndicator_eq","label":"integral_adversarialClipIndicator_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integral_adversarialClipIndicator_eq","description":"Each coordinate under the IID product law has the stated Gaussian marginal, specialized to the clipping indicator needed for Claim 17.7.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-cffa1a2a9388","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5586,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2018"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_adversarialClipIndicator_eq {horizon : Nat} (gap : Real) (t : Fin horizon) : ∫ eta, adversarialClipIndicator gap (eta t) ∂(adversarialCenteredNoiseLaw horizon (1 / 10)) = (gaussianReal 0 ⟨(1 / 10 : Real) ^ 2, sq_nonneg _⟩).real {x : Real | 1 / 2 - 2 * gap <= |x|}","missing":[],"search":"integral_adversarialclipindicator_eq banditrlproof.lowerbounds.integral_adversarialclipindicator_eq each coordinate under the iid product law has the stated gaussian marginal, specialized to the clipping indicator needed for claim 17.7. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClippingCount_tail_claim17_7","label":"adversarialClippingCount_tail_claim17_7","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClippingCount_tail_claim17_7","description":"*Claim 17.7.** Under the source variance `sigma = 1/10`, if `gap < 1/8` and `n >= 32 log(1/delta)`, at most a `delta` fraction of noise paths have at least `n/4` clipped rounds.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3064f73bc4d1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5587,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2050"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClippingCount_tail_claim17_7 {horizon : Nat} (hhorizon : 0 < horizon) (delta gap : Real) (hdelta : 0 < delta) (hdelta1 : delta < 1) (hgap_lt : gap < 1 / 8) (horizon_condition : 32 * Real.log (1 / delta) <= horizon) : (adversarialCenteredNoiseLaw horizon (1 / 10)).real {eta | (horizon : Real) / 4 <= adversarialClippingCountReal eta gap} <= delta","missing":[],"search":"adversarialclippingcount_tail_claim17_7 banditrlproof.lowerbounds.adversarialclippingcount_tail_claim17_7 *claim 17.7.** under the source variance `sigma = 1/10`, if `gap < 1/8` and `n >= 32 log(1/delta)`, at most a `delta` fraction of noise paths have at least `n/4` clipped rounds. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialBoundaryClippingCountReal","label":"adversarialBoundaryClippingCountReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialBoundaryClippingCountReal","description":"The textbook clipping count: a round is counted when at least one arm has reward at a boundary of the unit interval.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-bf362a4e4a5f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5588,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2160"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialBoundaryClippingCountReal {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (distinguished : Fin alternatives) : Real","missing":[],"search":"adversarialboundaryclippingcountreal banditrlproof.lowerbounds.adversarialboundaryclippingcountreal the textbook clipping count: a round is counted when at least one arm has reward at a boundary of the unit interval. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialBoundaryClippingCountReal_le","label":"adversarialBoundaryClippingCountReal_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialBoundaryClippingCountReal_le","description":"theorem adversarialBoundaryClippingCountReal_le {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) : adversarialBoundaryClippingCountReal eta gap distinguished <= adversarialClippingCountReal eta gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-57c677301623","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5589,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2168"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialBoundaryClippingCountReal_le {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) : adversarialBoundaryClippingCountReal eta gap distinguished <= adversarialClippingCountReal eta gap","missing":[],"search":"adversarialboundaryclippingcountreal_le banditrlproof.lowerbounds.adversarialboundaryclippingcountreal_le theorem adversarialboundaryclippingcountreal_le {horizon alternatives : nat} (eta : fin horizon -> real) (gap : real) (hgap : 0 <= gap) (distinguished : fin alternatives) : adversarialboundaryclippingcountreal eta gap distinguished <= adversarialclippingcountreal eta gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialBoundaryClippingCount_tail_claim17_7","label":"adversarialBoundaryClippingCount_tail_claim17_7","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialBoundaryClippingCount_tail_claim17_7","description":"Claim 17.7 for the literal boundary event in the textbook.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-077761c3b379","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5590,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2198"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialBoundaryClippingCount_tail_claim17_7 {horizon alternatives : Nat} (hhorizon : 0 < horizon) (delta gap : Real) (hdelta : 0 < delta) (hdelta1 : delta < 1) (hgap : 0 <= gap) (hgap_lt : gap < 1 / 8) (distinguished : Fin alternatives) (horizon_condition : 32 * Real.log (1 / delta) <= horizon) : (adversarialCenteredNoiseLaw horizon (1 / 10)).real {eta | (horizon : Real) / 4 <= adversarialBoundaryClippingCountReal eta gap distinguished} <= delta","missing":[],"search":"adversarialboundaryclippingcount_tail_claim17_7 banditrlproof.lowerbounds.adversarialboundaryclippingcount_tail_claim17_7 claim 17.7 for the literal boundary event in the textbook. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialComparatorRegret","label":"adversarialComparatorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialComparatorRegret","description":"def adversarialComparatorRegret {horizon alternatives : Nat} (reward : Fin horizon -> Fin (alternatives + 1) -> Real) (actions : Fin horizon -> Fin (alternatives + 1)) (comparator : Fin (alternatives + 1)) : Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-172771c40e0c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5591,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2217"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialComparatorRegret {horizon alternatives : Nat} (reward : Fin horizon -> Fin (alternatives + 1) -> Real) (actions : Fin horizon -> Fin (alternatives + 1)) (comparator : Fin (alternatives + 1)) : Real","missing":[],"search":"adversarialcomparatorregret banditrlproof.lowerbounds.adversarialcomparatorregret def adversarialcomparatorregret {horizon alternatives : nat} (reward : fin horizon -> fin (alternatives + 1) -> real) (actions : fin horizon -> fin (alternatives + 1)) (comparator : fin (alternatives + 1)) : real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret","label":"adversarialRandomRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialRandomRegret","description":"Adversarial random regret: the best fixed arm in hindsight minus the reward collected along the realized action path.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-368e2041a275","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5592,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2226"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialRandomRegret {horizon alternatives : Nat} (reward : Fin horizon -> Fin (alternatives + 1) -> Real) (actions : Fin horizon -> Fin (alternatives + 1)) : Real","missing":[],"search":"adversarialrandomregret banditrlproof.lowerbounds.adversarialrandomregret adversarial random regret: the best fixed arm in hindsight minus the reward collected along the realized action path. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialComparatorRegret_le_randomRegret","label":"adversarialComparatorRegret_le_randomRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialComparatorRegret_le_randomRegret","description":"theorem adversarialComparatorRegret_le_randomRegret {horizon alternatives : Nat} (reward : Fin horizon -> Fin (alternatives + 1) -> Real) (actions : Fin horizon -> Fin (alternatives + 1)) (comparator : Fin (alternatives + 1)) : adversarialComparatorRegret reward actions comparator <= adversarialRandomRegret reward actions","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-d111e526b466","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5593,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2234"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialComparatorRegret_le_randomRegret {horizon alternatives : Nat} (reward : Fin horizon -> Fin (alternatives + 1) -> Real) (actions : Fin horizon -> Fin (alternatives + 1)) (comparator : Fin (alternatives + 1)) : adversarialComparatorRegret reward actions comparator <= adversarialRandomRegret reward actions","missing":[],"search":"adversarialcomparatorregret_le_randomregret banditrlproof.lowerbounds.adversarialcomparatorregret_le_randomregret theorem adversarialcomparatorregret_le_randomregret {horizon alternatives : nat} (reward : fin horizon -> fin (alternatives + 1) -> real) (actions : fin horizon -> fin (alternatives + 1)) (comparator : fin (alternatives + 1)) : adversarialcomparatorregret reward actions comparator <= adversarialrandomregret reward actions theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialComparatorRegret_ge_eq17_8","label":"adversarialComparatorRegret_ge_eq17_8","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialComparatorRegret_ge_eq17_8","description":"*Equation (17.8), construction level.** For every realized shared-noise path and every action path, regret against the distinguished arm is bounded below by `gap * (n - T_i(n) - C)`, where `C` counts clipping rounds.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a702cb85834c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5594,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2247"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialComparatorRegret_ge_eq17_8 {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (actions : Fin horizon -> Fin (alternatives + 1)) : gap * ((horizon : Real) - adversarialPullCountReal actions distinguished - adversarialClippingCountReal eta gap) <= adversarialComparatorRegret (adversarialClippedGaussianReward eta gap distinguished) actions distinguished.succ","missing":[],"search":"adversarialcomparatorregret_ge_eq17_8 banditrlproof.lowerbounds.adversarialcomparatorregret_ge_eq17_8 *equation (17.8), construction level.** for every realized shared-noise path and every action path, regret against the distinguished arm is bounded below by `gap * (n - t_i(n) - c)`, where `c` counts clipping rounds. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_eq17_8","label":"adversarialRandomRegret_ge_eq17_8","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialRandomRegret_ge_eq17_8","description":"Eq. (17.8) in the textbook's actual random-regret form.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-2baa1151442d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5595,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2296"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","distributional-lower-bounds"]],"statement":"theorem adversarialRandomRegret_ge_eq17_8 {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (actions : Fin horizon -> Fin (alternatives + 1)) : gap * ((horizon : Real) - adversarialPullCountReal actions distinguished - adversarialClippingCountReal eta gap) <= adversarialRandomRegret (adversarialClippedGaussianReward eta gap distinguished) actions","missing":[],"search":"adversarialrandomregret_ge_eq17_8 banditrlproof.lowerbounds.adversarialrandomregret_ge_eq17_8 eq. (17.8) in the textbook's actual random-regret form. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":["distributional-lower-bounds"]},{"id":"declaration:BanditRLProof.LowerBounds.clipUnitReward_eq_self_of_ne_endpoints","label":"clipUnitReward_eq_self_of_ne_endpoints","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.clipUnitReward_eq_self_of_ne_endpoints","description":"private theorem clipUnitReward_eq_self_of_ne_endpoints (x : Real) (h0 : clipUnitReward x ≠ 0) (h1 : clipUnitReward x ≠ 1) : clipUnitReward x = x","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-63966550de76","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5596,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2312"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem clipUnitReward_eq_self_of_ne_endpoints (x : Real) (h0 : clipUnitReward x ≠ 0) (h1 : clipUnitReward x ≠ 1) : clipUnitReward x = x","missing":[],"search":"clipunitreward_eq_self_of_ne_endpoints banditrlproof.lowerbounds.clipunitreward_eq_self_of_ne_endpoints private theorem clipunitreward_eq_self_of_ne_endpoints (x : real) (h0 : clipunitreward x ≠ 0) (h1 : clipunitreward x ≠ 1) : clipunitreward x = x theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_boundary_eq17_8","label":"adversarialRandomRegret_ge_boundary_eq17_8","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialRandomRegret_ge_boundary_eq17_8","description":"Exact Eq. (17.8), with the textbook's actual boundary count.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-06d532a51146","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5597,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2330"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialRandomRegret_ge_boundary_eq17_8 {horizon alternatives : Nat} (eta : Fin horizon -> Real) (gap : Real) (hgap : 0 <= gap) (distinguished : Fin alternatives) (actions : Fin horizon -> Fin (alternatives + 1)) : gap * ((horizon : Real) - adversarialPullCountReal actions distinguished - adversarialBoundaryClippingCountReal eta gap distinguished) <= adversarialRandomRegret (adversarialClippedGaussianReward eta gap distinguished) actions","missing":[],"search":"adversarialrandomregret_ge_boundary_eq17_8 banditrlproof.lowerbounds.adversarialrandomregret_ge_boundary_eq17_8 exact eq. (17.8), with the textbook's actual boundary count. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullClippedReward","label":"adversarialFullClippedReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullClippedReward","description":"The shared-noise matrix for every witness, including the base arm.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-2e6f631725f4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5598,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2379"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialFullClippedReward {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (i : Fin (m + 1)) (t : Fin horizon) (arm : Fin (m + 1)) : Real","missing":[],"search":"adversarialfullclippedreward banditrlproof.lowerbounds.adversarialfullclippedreward the shared-noise matrix for every witness, including the base arm. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_bounds","label":"adversarialFullHardShift_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullHardShift_bounds","description":"theorem adversarialFullHardShift_bounds {m : Nat} (gap : Real) (hg : 0 <= gap) (i arm : Fin (m + 1)) : 0 <= adversarialFullHardShift gap i arm ∧ adversarialFullHardShift gap i arm <= 2 * gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-c000ae83c662","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5599,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2384"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullHardShift_bounds {m : Nat} (gap : Real) (hg : 0 <= gap) (i arm : Fin (m + 1)) : 0 <= adversarialFullHardShift gap i arm ∧ adversarialFullHardShift gap i arm <= 2 * gap","missing":[],"search":"adversarialfullhardshift_bounds banditrlproof.lowerbounds.adversarialfullhardshift_bounds theorem adversarialfullhardshift_bounds {m : nat} (gap : real) (hg : 0 <= gap) (i arm : fin (m + 1)) : 0 <= adversarialfullhardshift gap i arm ∧ adversarialfullhardshift gap i arm <= 2 * gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_separation","label":"adversarialFullHardShift_separation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullHardShift_separation","description":"theorem adversarialFullHardShift_separation {m : Nat} (gap : Real) (hg : 0 <= gap) (i arm : Fin (m + 1)) (hi : arm ≠ i) : gap <= adversarialFullHardShift gap i i - adversarialFullHardShift gap i arm","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5829c51f2bf3","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5600,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2391"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullHardShift_separation {m : Nat} (gap : Real) (hg : 0 <= gap) (i arm : Fin (m + 1)) (hi : arm ≠ i) : gap <= adversarialFullHardShift gap i i - adversarialFullHardShift gap i arm","missing":[],"search":"adversarialfullhardshift_separation banditrlproof.lowerbounds.adversarialfullhardshift_separation theorem adversarialfullhardshift_separation {m : nat} (gap : real) (hg : 0 <= gap) (i arm : fin (m + 1)) (hi : arm ≠ i) : gap <= adversarialfullhardshift gap i i - adversarialfullhardshift gap i arm theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullClippedReward_best","label":"adversarialFullClippedReward_best","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullClippedReward_best","description":"theorem adversarialFullClippedReward_best {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (hg : 0 <= gap) (i : Fin (m + 1)) (t : Fin horizon) (arm : Fin (m + 1)) : adversarialFullClippedReward eta gap i t arm <= adversarialFullClippedReward eta gap i t i","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-555aaee3cb5f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5601,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2403"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullClippedReward_best {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (hg : 0 <= gap) (i : Fin (m + 1)) (t : Fin horizon) (arm : Fin (m + 1)) : adversarialFullClippedReward eta gap i t arm <= adversarialFullClippedReward eta gap i t i","missing":[],"search":"adversarialfullclippedreward_best banditrlproof.lowerbounds.adversarialfullclippedreward_best theorem adversarialfullclippedreward_best {horizon m : nat} (eta : fin horizon -> real) (gap : real) (hg : 0 <= gap) (i : fin (m + 1)) (t : fin horizon) (arm : fin (m + 1)) : adversarialfullclippedreward eta gap i t arm <= adversarialfullclippedreward eta gap i t i theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount","label":"adversarialFullBoundaryCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullBoundaryCount","description":"noncomputable def adversarialFullBoundaryCount {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (i : Fin (m + 1)) : Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f51b768c9be1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5602,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2414"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialFullBoundaryCount {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (i : Fin (m + 1)) : Real","missing":[],"search":"adversarialfullboundarycount banditrlproof.lowerbounds.adversarialfullboundarycount noncomputable def adversarialfullboundarycount {horizon m : nat} (eta : fin horizon -> real) (gap : real) (i : fin (m + 1)) : real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullRandomRegret_ge_boundary_eq17_8","label":"adversarialFullRandomRegret_ge_boundary_eq17_8","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullRandomRegret_ge_boundary_eq17_8","description":"Literal boundary-count Eq. (17.8), now also valid for the base witness.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e362e36e34c4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5603,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2420"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullRandomRegret_ge_boundary_eq17_8 {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (hg : 0 <= gap) (i : Fin (m + 1)) (actions : Fin horizon -> Fin (m + 1)) : gap * ((horizon : Real) - (∑ t, if actions t = i then 1 else 0) - adversarialFullBoundaryCount eta gap i) <= adversarialRandomRegret (adversarialFullClippedReward eta gap i) actions","missing":[],"search":"adversarialfullrandomregret_ge_boundary_eq17_8 banditrlproof.lowerbounds.adversarialfullrandomregret_ge_boundary_eq17_8 literal boundary-count eq. (17.8), now also valid for the base witness. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount_le","label":"adversarialFullBoundaryCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullBoundaryCount_le","description":"theorem adversarialFullBoundaryCount_le {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (hg : 0 <= gap) (i : Fin (m + 1)) : adversarialFullBoundaryCount eta gap i <= adversarialClippingCountReal eta gap","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-a72243040087","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5604,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2459"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullBoundaryCount_le {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (hg : 0 <= gap) (i : Fin (m + 1)) : adversarialFullBoundaryCount eta gap i <= adversarialClippingCountReal eta gap","missing":[],"search":"adversarialfullboundarycount_le banditrlproof.lowerbounds.adversarialfullboundarycount_le theorem adversarialfullboundarycount_le {horizon m : nat} (eta : fin horizon -> real) (gap : real) (hg : 0 <= gap) (i : fin (m + 1)) : adversarialfullboundarycount eta gap i <= adversarialclippingcountreal eta gap theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount_tail_claim17_7","label":"adversarialFullBoundaryCount_tail_claim17_7","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullBoundaryCount_tail_claim17_7","description":"Claim 17.7 for every member of the corrected Claim 17.6 witness family.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-14e3cc084435","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5605,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2483"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullBoundaryCount_tail_claim17_7 {horizon m : Nat} (hn : 0 < horizon) (delta gap : Real) (hd : 0 < delta) (hd1 : delta < 1) (hg : 0 <= gap) (hg8 : gap < 1 / 8) (i : Fin (m + 1)) (horizon_condition : 32 * Real.log (1 / delta) <= horizon) : (adversarialCenteredNoiseLaw horizon (1 / 10)).real {eta | (horizon : Real) / 4 <= adversarialFullBoundaryCount eta gap i} <= delta","missing":[],"search":"adversarialfullboundarycount_tail_claim17_7 banditrlproof.lowerbounds.adversarialfullboundarycount_tail_claim17_7 claim 17.7 for every member of the corrected claim 17.6 witness family. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.AdversarialRewardTable","label":"AdversarialRewardTable","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.AdversarialRewardTable","description":"A deterministic oblivious reward table. Only its finite prefix is used.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f1ce119a9b6c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5606,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2499"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"abbrev AdversarialRewardTable (K : Nat)","missing":[],"search":"adversarialrewardtable banditrlproof.lowerbounds.adversarialrewardtable a deterministic oblivious reward table. only its finite prefix is used. abbreviation compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableInitialFeedback","label":"adversarialTableInitialFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableInitialFeedback","description":"Feedback from a fixed table, with the table retained as a kernel parameter.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-4fda9ff20b19","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5607,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2502"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialTableInitialFeedback {K : Nat} : Kernel (AdversarialRewardTable K × Fin K) Real","missing":[],"search":"adversarialtableinitialfeedback banditrlproof.lowerbounds.adversarialtableinitialfeedback feedback from a fixed table, with the table retained as a kernel parameter. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableNextFeedback","label":"adversarialTableNextFeedback","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableNextFeedback","description":"noncomputable def adversarialTableNextFeedback {K : Nat} (n : Nat) : Kernel ((AdversarialRewardTable K × History.FinitePairHistory (Fin K) Real n) × Fin K) Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-bbc7b6d2619c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5608,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2513"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialTableNextFeedback {K : Nat} (n : Nat) : Kernel ((AdversarialRewardTable K × History.FinitePairHistory (Fin K) Real n) × Fin K) Real","missing":[],"search":"adversarialtablenextfeedback banditrlproof.lowerbounds.adversarialtablenextfeedback noncomputable def adversarialtablenextfeedback {k : nat} (n : nat) : kernel ((adversarialrewardtable k × history.finitepairhistory (fin k) real n) × fin k) real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableStepKernel","label":"adversarialTableStepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableStepKernel","description":"The original policy observes exactly its action/reward prefix.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5eabe3f8ba94","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5609,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2525"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialTableStepKernel {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) : Kernel (AdversarialRewardTable K × History.FinitePairHistory (Fin K) Real n) (Fin K × Real)","missing":[],"search":"adversarialtablestepkernel banditrlproof.lowerbounds.adversarialtablestepkernel the original policy observes exactly its action/reward prefix. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel","label":"adversarialTableHistoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableHistoryKernel","description":"Conditional history distribution under each fixed oblivious table. This is a measurable kernel, so averaging over random tables is well-defined.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-fa09aa6887c4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5610,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2538"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialTableHistoryKernel {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) : (n : Nat) -> Kernel (AdversarialRewardTable K) (History.FinitePairHistory (Fin K) Real n) | 0 => ((Kernel.const (AdversarialRewardTable K) algorithm.initialAction) ⊗ₖ adversarialTableInitialFeedback).map (pairHistoryZeroMeasurableEquiv (Fin K) Real) | n + 1 => ((adversarialTableHistoryKernel algorithm n) ⊗ₖ adversarialTableStepKernel algorithm n).map (pairHistorySuccMeasurableEquiv (Fin K) Real n) instance adversarialTableHistoryKernel_isMarkov {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) : IsMarkovKernel (adversarialTableHistoryKernel algorithm n)","missing":[],"search":"adversarialtablehistorykernel banditrlproof.lowerbounds.adversarialtablehistorykernel conditional history distribution under each fixed oblivious table. this is a measurable kernel, so averaging over random tables is well-defined. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableStepKernel_apply","label":"adversarialTableStepKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableStepKernel_apply","description":"Fixing the table gives exactly the original policy and deterministic feedback.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-90da877f3af5","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5611,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2559"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialTableStepKernel_apply {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (table : AdversarialRewardTable K) (history : History.FinitePairHistory (Fin K) Real n) : adversarialTableStepKernel algorithm n (table, history) = algorithm.policy n history ⊗ₘ Kernel.deterministic (table (n + 1)) (measurable_of_countable _)","missing":[],"search":"adversarialtablestepkernel_apply banditrlproof.lowerbounds.adversarialtablestepkernel_apply fixing the table gives exactly the original policy and deterministic feedback. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel_zero","label":"adversarialTableHistoryKernel_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableHistoryKernel_zero","description":"theorem adversarialTableHistoryKernel_zero {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (table : AdversarialRewardTable K) : adversarialTableHistoryKernel algorithm 0 table = (algorithm.initialAction ⊗ₘ Kernel.deterministic (table 0) (measurable_of_countable _)).map (pairHistoryZeroMeasurableEquiv (Fin K) Real)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-f2510d69bfff","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5612,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2569"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialTableHistoryKernel_zero {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (table : AdversarialRewardTable K) : adversarialTableHistoryKernel algorithm 0 table = (algorithm.initialAction ⊗ₘ Kernel.deterministic (table 0) (measurable_of_countable _)).map (pairHistoryZeroMeasurableEquiv (Fin K) Real)","missing":[],"search":"adversarialtablehistorykernel_zero banditrlproof.lowerbounds.adversarialtablehistorykernel_zero theorem adversarialtablehistorykernel_zero {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (table : adversarialrewardtable k) : adversarialtablehistorykernel algorithm 0 table = (algorithm.initialaction ⊗ₘ kernel.deterministic (table 0) (measurable_of_countable _)).map (pairhistoryzeromeasurableequiv (fin k) real) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel_succ","label":"adversarialTableHistoryKernel_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableHistoryKernel_succ","description":"theorem adversarialTableHistoryKernel_succ {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (table : AdversarialRewardTable K) : adversarialTableHistoryKernel algorithm (n + 1) table = ((adversarialTableHistoryKernel algorithm n table) ⊗ₘ Kernel.sectR (adversarialTableStepKernel algorithm n) table).map (pairHistorySuccMeasurableEquiv (Fin K) Real n)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-31b59e6b10fe","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5613,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2580"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialTableHistoryKernel_succ {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (table : AdversarialRewardTable K) : adversarialTableHistoryKernel algorithm (n + 1) table = ((adversarialTableHistoryKernel algorithm n table) ⊗ₘ Kernel.sectR (adversarialTableStepKernel algorithm n) table).map (pairHistorySuccMeasurableEquiv (Fin K) Real n)","missing":[],"search":"adversarialtablehistorykernel_succ banditrlproof.lowerbounds.adversarialtablehistorykernel_succ theorem adversarialtablehistorykernel_succ {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (n : nat) (table : adversarialrewardtable k) : adversarialtablehistorykernel algorithm (n + 1) table = ((adversarialtablehistorykernel algorithm n table) ⊗ₘ kernel.sectr (adversarialtablestepkernel algorithm n) table).map (pairhistorysuccmeasurableequiv (fin k) real n) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel_prefix_congr","label":"adversarialTableHistoryKernel_prefix_congr","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableHistoryKernel_prefix_congr","description":"Future reward rows cannot affect an already observed history.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-77d2d6d38a92","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5614,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2592"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialTableHistoryKernel_prefix_congr {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (table other : AdversarialRewardTable K) (heq : ∀ t, t <= n -> table t = other t) : adversarialTableHistoryKernel algorithm n table = adversarialTableHistoryKernel algorithm n other","missing":[],"search":"adversarialtablehistorykernel_prefix_congr banditrlproof.lowerbounds.adversarialtablehistorykernel_prefix_congr future reward rows cannot affect an already observed history. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullRewardTable","label":"adversarialFullRewardTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullRewardTable","description":"Extend the finite shared-noise matrix by zero rows after its horizon.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-47901a1ffc1e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5615,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2613"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialFullRewardTable {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (i : Fin (m + 1)) : AdversarialRewardTable (m + 1)","missing":[],"search":"adversarialfullrewardtable banditrlproof.lowerbounds.adversarialfullrewardtable extend the finite shared-noise matrix by zero rows after its horizon. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialFullRewardTable","label":"measurable_adversarialFullRewardTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialFullRewardTable","description":"theorem measurable_adversarialFullRewardTable {horizon m : Nat} (gap : Real) (i : Fin (m + 1)) : Measurable (fun eta : Fin horizon -> Real => adversarialFullRewardTable eta gap i)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-548e5fd941a1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5616,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2618"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialFullRewardTable {horizon m : Nat} (gap : Real) (i : Fin (m + 1)) : Measurable (fun eta : Fin horizon -> Real => adversarialFullRewardTable eta gap i)","missing":[],"search":"measurable_adversarialfullrewardtable banditrlproof.lowerbounds.measurable_adversarialfullrewardtable theorem measurable_adversarialfullrewardtable {horizon m : nat} (gap : real) (i : fin (m + 1)) : measurable (fun eta : fin horizon -> real => adversarialfullrewardtable eta gap i) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialFullRewardTable_at","label":"adversarialFullRewardTable_at","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialFullRewardTable_at","description":"theorem adversarialFullRewardTable_at {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (i : Fin (m + 1)) (t : Fin horizon) : adversarialFullRewardTable eta gap i t = adversarialFullClippedReward eta gap i t","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-b9419bbabaf9","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5617,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2631"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialFullRewardTable_at {horizon m : Nat} (eta : Fin horizon -> Real) (gap : Real) (i : Fin (m + 1)) (t : Fin horizon) : adversarialFullRewardTable eta gap i t = adversarialFullClippedReward eta gap i t","missing":[],"search":"adversarialfullrewardtable_at banditrlproof.lowerbounds.adversarialfullrewardtable_at theorem adversarialfullrewardtable_at {horizon m : nat} (eta : fin horizon -> real) (gap : real) (i : fin (m + 1)) (t : fin horizon) : adversarialfullrewardtable eta gap i t = adversarialfullclippedreward eta gap i t theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel","label":"adversarialNoiseHistoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel","description":"Conditional policy history kernel parameterized by the finite noise path.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3d11627c30a0","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5618,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2638"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialNoiseHistoryKernel {horizon m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (i : Fin (m + 1)) (n : Nat) : Kernel (Fin horizon -> Real) (History.FinitePairHistory (Fin (m + 1)) Real n)","missing":[],"search":"adversarialnoisehistorykernel banditrlproof.lowerbounds.adversarialnoisehistorykernel conditional policy history kernel parameterized by the finite noise path. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint","label":"adversarialNoiseHistoryJoint","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint","description":"Shared-noise matrix and the original randomized policy on one joint space.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-b4f778e5de4d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5619,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2653"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialNoiseHistoryJoint {horizon m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) : Measure ((Fin horizon -> Real) × History.FinitePairHistory (Fin (m + 1)) Real n)","missing":[],"search":"adversarialnoisehistoryjoint banditrlproof.lowerbounds.adversarialnoisehistoryjoint shared-noise matrix and the original randomized policy on one joint space. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_noise_marginal","label":"adversarialNoiseHistoryJoint_noise_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_noise_marginal","description":"theorem adversarialNoiseHistoryJoint_noise_marginal {horizon m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) : (adversarialNoiseHistoryJoint (horizon := horizon) algorithm sigma gap i n).fst = adversarialCenteredNoiseLaw horizon sigma","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-175f127b9243","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5620,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2666"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_noise_marginal {horizon m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) : (adversarialNoiseHistoryJoint (horizon := horizon) algorithm sigma gap i n).fst = adversarialCenteredNoiseLaw horizon sigma","missing":[],"search":"adversarialnoisehistoryjoint_noise_marginal banditrlproof.lowerbounds.adversarialnoisehistoryjoint_noise_marginal theorem adversarialnoisehistoryjoint_noise_marginal {horizon m : nat} (algorithm : thompson.historyalgorithm (fin (m + 1)) real) (sigma gap : real) (i : fin (m + 1)) (n : nat) : (adversarialnoisehistoryjoint (horizon := horizon) algorithm sigma gap i n).fst = adversarialcenterednoiselaw horizon sigma theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel_update_future","label":"adversarialNoiseHistoryKernel_update_future","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel_update_future","description":"Resampling an unobserved noise coordinate leaves the prefix law unchanged.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-69b8f8e5bf40","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5621,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2678"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryKernel_update_future {horizon m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (i : Fin (m + 1)) (n : Nat) (eta : Fin horizon -> Real) (j : Fin horizon) (hj : n < j.val) (x : Real) : adversarialNoiseHistoryKernel algorithm gap i n (Function.update eta j x) = adversarialNoiseHistoryKernel algorithm gap i n eta","missing":[],"search":"adversarialnoisehistorykernel_update_future banditrlproof.lowerbounds.adversarialnoisehistorykernel_update_future resampling an unobserved noise coordinate leaves the prefix law unchanged. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_split","label":"adversarialCenteredNoiseLaw_split","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_split","description":"Isolate any coordinate of the shared Gaussian noise as a product factor.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-24779499231f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5622,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2701"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialCenteredNoiseLaw_split {N : Nat} (sigma : Real) (j : Fin (N + 1)) : MeasurePreserving (MeasurableEquiv.piFinSuccAbove (fun _ : Fin (N + 1) => Real) j) (adversarialCenteredNoiseLaw (N + 1) sigma) ((gaussianReal 0 ⟨sigma ^ 2, sq_nonneg sigma⟩).prod (adversarialCenteredNoiseLaw N sigma))","missing":[],"search":"adversarialcenterednoiselaw_split banditrlproof.lowerbounds.adversarialcenterednoiselaw_split isolate any coordinate of the shared gaussian noise as a product factor. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel_split_future","label":"adversarialNoiseHistoryKernel_split_future","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel_split_future","description":"In split coordinates, the prefix history law is independent of the future factor.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-6a3230244ca3","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5623,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2710"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryKernel_split_future {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (gap : Real) (i : Fin (m + 1)) (n : Nat) (j : Fin (N + 1)) (hj : n < j.val) (rest : Fin N -> Real) (x y : Real) : adversarialNoiseHistoryKernel algorithm gap i n (j.insertNth x rest) = adversarialNoiseHistoryKernel algorithm gap i n (j.insertNth y rest)","missing":[],"search":"adversarialnoisehistorykernel_split_future banditrlproof.lowerbounds.adversarialnoisehistorykernel_split_future in split coordinates, the prefix history law is independent of the future factor. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialCenteredNoiseLaw_split","label":"lintegral_adversarialCenteredNoiseLaw_split","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialCenteredNoiseLaw_split","description":"Tonelli in the independent-coordinate representation of the hard noise law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ed6b6f1973c3","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5624,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2731"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialCenteredNoiseLaw_split {N : Nat} (sigma : Real) (j : Fin (N + 1)) (f : (Fin (N + 1) -> Real) -> ENNReal) (hf : Measurable f) : (∫⁻ eta, f eta ∂adversarialCenteredNoiseLaw (N + 1) sigma) = ∫⁻ x, ∫⁻ rest, f (j.insertNth x rest) ∂adversarialCenteredNoiseLaw N sigma ∂gaussianReal 0 ⟨sigma ^ 2, sq_nonneg sigma⟩","missing":[],"search":"lintegral_adversarialcenterednoiselaw_split banditrlproof.lowerbounds.lintegral_adversarialcenterednoiselaw_split tonelli in the independent-coordinate representation of the hard noise law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_full_reward_marginal","label":"adversarialCenteredNoiseLaw_full_reward_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_full_reward_marginal","description":"theorem adversarialCenteredNoiseLaw_full_reward_marginal {horizon m : Nat} (sigma gap : Real) (i : Fin (m + 1)) (t : Fin horizon) (arm : Fin (m + 1)) : (adversarialCenteredNoiseLaw horizon sigma).map (fun eta => adversarialFullClippedReward eta gap i t arm) = adversarialClippedArmLaw sigma (adversarialFullHardShift gap i arm)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-3c9be5cfe864","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5625,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2745"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialCenteredNoiseLaw_full_reward_marginal {horizon m : Nat} (sigma gap : Real) (i : Fin (m + 1)) (t : Fin horizon) (arm : Fin (m + 1)) : (adversarialCenteredNoiseLaw horizon sigma).map (fun eta => adversarialFullClippedReward eta gap i t arm) = adversarialClippedArmLaw sigma (adversarialFullHardShift gap i arm)","missing":[],"search":"adversarialcenterednoiselaw_full_reward_marginal banditrlproof.lowerbounds.adversarialcenterednoiselaw_full_reward_marginal theorem adversarialcenterednoiselaw_full_reward_marginal {horizon m : nat} (sigma gap : real) (i : fin (m + 1)) (t : fin horizon) (arm : fin (m + 1)) : (adversarialcenterednoiselaw horizon sigma).map (fun eta => adversarialfullclippedreward eta gap i t arm) = adversarialclippedarmlaw sigma (adversarialfullhardshift gap i arm) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialTableHistoryKernel_zero","label":"lintegral_adversarialTableHistoryKernel_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialTableHistoryKernel_zero","description":"theorem lintegral_adversarialTableHistoryKernel_zero {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (table : AdversarialRewardTable K) (f : History.FinitePairHistory (Fin K) Real 0 -> ENNReal) (hf : Measurable f) : (∫⁻ h, f h ∂adversarialTableHistoryKernel algorithm 0 table) = ∫⁻ arm, f (pairHistoryZeroMeasurableEquiv (Fin K) Real (arm, table 0 arm)) ∂algorithm.initialAction","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-05b68ab82b1f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5626,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2758"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialTableHistoryKernel_zero {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (table : AdversarialRewardTable K) (f : History.FinitePairHistory (Fin K) Real 0 -> ENNReal) (hf : Measurable f) : (∫⁻ h, f h ∂adversarialTableHistoryKernel algorithm 0 table) = ∫⁻ arm, f (pairHistoryZeroMeasurableEquiv (Fin K) Real (arm, table 0 arm)) ∂algorithm.initialAction","missing":[],"search":"lintegral_adversarialtablehistorykernel_zero banditrlproof.lowerbounds.lintegral_adversarialtablehistorykernel_zero theorem lintegral_adversarialtablehistorykernel_zero {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (table : adversarialrewardtable k) (f : history.finitepairhistory (fin k) real 0 -> ennreal) (hf : measurable f) : (∫⁻ h, f h ∂adversarialtablehistorykernel algorithm 0 table) = ∫⁻ arm, f (pairhistoryzeromeasurableequiv (fin k) real (arm, table 0 arm)) ∂algorithm.initialaction theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_zero","label":"lintegral_adversarialNoiseHistoryKernel_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_zero","description":"The initial matrix mixture has exactly the canonical clipped observation law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-cdde4efa80d9","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5627,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2772"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialNoiseHistoryKernel_zero {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (f : History.FinitePairHistory (Fin (m + 1)) Real 0 -> ENNReal) (hf : Measurable f) : (∫⁻ eta, ∫⁻ h, f h ∂adversarialNoiseHistoryKernel (horizon := N + 1) algorithm gap i 0 eta ∂adversarialCenteredNoiseLaw (N + 1) sigma) = ∫⁻ h, f h ∂adversarialClippedHistoryLaw algorithm sigma (adversarialFullHardShift gap i) 0","missing":[],"search":"lintegral_adversarialnoisehistorykernel_zero banditrlproof.lowerbounds.lintegral_adversarialnoisehistorykernel_zero the initial matrix mixture has exactly the canonical clipped observation law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal_zero","label":"adversarialNoiseHistoryJoint_history_marginal_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal_zero","description":"Initial case of the joint-space history marginal identification.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-6c79581b1be4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5628,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2816"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_history_marginal_zero {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) : (adversarialNoiseHistoryJoint (horizon := N + 1) algorithm sigma gap i 0).snd = adversarialClippedHistoryLaw algorithm sigma (adversarialFullHardShift gap i) 0","missing":[],"search":"adversarialnoisehistoryjoint_history_marginal_zero banditrlproof.lowerbounds.adversarialnoisehistoryjoint_history_marginal_zero initial case of the joint-space history marginal identification. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialTableHistoryKernel_succ","label":"lintegral_adversarialTableHistoryKernel_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialTableHistoryKernel_succ","description":"Conditional table-history recursion, with deterministic feedback integrated out.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-54473d8e453f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5629,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2834"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialTableHistoryKernel_succ {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (table : AdversarialRewardTable K) (n : Nat) (f : History.FinitePairHistory (Fin K) Real (n + 1) -> ENNReal) (hf : Measurable f) : (∫⁻ h, f h ∂adversarialTableHistoryKernel algorithm (n + 1) table) = ∫⁻ h, ∫⁻ arm, f (pairHistorySuccMeasurableEquiv (Fin K) Real n (h, (arm, table (n + 1) arm))) ∂algorithm.policy n h ∂adversarialTableHistoryKernel algorithm n table","missing":[],"search":"lintegral_adversarialtablehistorykernel_succ banditrlproof.lowerbounds.lintegral_adversarialtablehistorykernel_succ conditional table-history recursion, with deterministic feedback integrated out. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialFreshNoise_step","label":"lintegral_adversarialFreshNoise_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialFreshNoise_step","description":"Integrate fresh shared noise against a fixed prefix law and the same policy.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e8f1c2896d8c","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5630,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2854"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialFreshNoise_step {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (n : Nat) (P : Measure (History.FinitePairHistory (Fin K) Real n)) [IsProbabilityMeasure P] (sigma : Real) (shift : Fin K -> Real) (f : History.FinitePairHistory (Fin K) Real n × (Fin K × Real) -> ENNReal) (hf : Measurable f) : (∫⁻ x, ∫⁻ h, ∫⁻ arm, f (h, (arm, clipUnitReward (1 / 2 + x + shift arm))) ∂algorithm.policy n h ∂P ∂gaussianReal 0 ⟨sigma ^ 2, sq_nonneg sigma⟩) = ∫⁻ h, ∫⁻ pair, f (h, pair) ∂Thompson.historyStepKernel algorithm (stationaryBanditHistoryEnvironment (adversarialClippedKernel sigma shift)) n h ∂P","missing":[],"search":"lintegral_adversarialfreshnoise_step banditrlproof.lowerbounds.lintegral_adversarialfreshnoise_step integrate fresh shared noise against a fixed prefix law and the same policy. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_succ_slice","label":"lintegral_adversarialNoiseHistoryKernel_succ_slice","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_succ_slice","description":"Successor history integration on each fixed remaining-noise slice.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-8b96f415b417","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5631,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2901"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialNoiseHistoryKernel_succ_slice {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) (hn : n + 1 < N + 1) (rest : Fin N -> Real) (f : History.FinitePairHistory (Fin (m + 1)) Real (n + 1) -> ENNReal) (hf : Measurable f) : (∫⁻ x, ∫⁻ h, f h ∂adversarialNoiseHistoryKernel algorithm gap i (n + 1) ((⟨n + 1, hn⟩ : Fin (N + 1)).insertNth x rest) ∂gaussianReal 0 ⟨sigma ^ 2, sq_nonneg sigma⟩) = ∫⁻ h, ∫⁻ pair, f (pairHistorySuccMeasurableEquiv (Fin (m + 1)) Real n (h, pair)) ∂Thompson.historyStepKernel algorithm (stationaryBanditHistoryEnvironment (adversarialClippedKernel sigma (adversarialFullHardShift gap i))) n h ∂adversarialNoiseHistoryKernel algorithm gap i n ((⟨n + 1, hn⟩ : Fin (N + 1)).insertNth 0 rest)","missing":[],"search":"lintegral_adversarialnoisehistorykernel_succ_slice banditrlproof.lowerbounds.lintegral_adversarialnoisehistorykernel_succ_slice successor history integration on each fixed remaining-noise slice. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_succ","label":"lintegral_adversarialNoiseHistoryKernel_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_succ","description":"Full-noise successor recursion after integrating the independent next coordinate.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-6d320a06f7e8","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5632,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2945"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialNoiseHistoryKernel_succ {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) (hn : n + 1 < N + 1) (f : History.FinitePairHistory (Fin (m + 1)) Real (n + 1) -> ENNReal) (hf : Measurable f) : (∫⁻ eta, ∫⁻ h, f h ∂adversarialNoiseHistoryKernel (horizon := N + 1) algorithm gap i (n + 1) eta ∂adversarialCenteredNoiseLaw (N + 1) sigma) = ∫⁻ eta, ∫⁻ h, ∫⁻ pair, f (pairHistorySuccMeasurableEquiv (Fin (m + 1)) Real n (h, pair)) ∂Thompson.historyStepKernel algorithm (stationaryBanditHistoryEnvironment (adversarialClippedKernel sigma (adversarialFullHardShift gap i))) n h ∂adversarialNoiseHistoryKernel algorithm gap i n eta ∂adversarialCenteredNoiseLaw (N + 1) sigma","missing":[],"search":"lintegral_adversarialnoisehistorykernel_succ banditrlproof.lowerbounds.lintegral_adversarialnoisehistorykernel_succ full-noise successor recursion after integrating the independent next coordinate. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_eq_clipped","label":"lintegral_adversarialNoiseHistoryKernel_eq_clipped","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_eq_clipped","description":"All observed finite prefixes have the canonical clipped history law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-0ca618f8327e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5633,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:2994"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem lintegral_adversarialNoiseHistoryKernel_eq_clipped {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) (hn : n < N + 1) (f : History.FinitePairHistory (Fin (m + 1)) Real n -> ENNReal) (hf : Measurable f) : (∫⁻ eta, ∫⁻ h, f h ∂adversarialNoiseHistoryKernel (horizon := N + 1) algorithm gap i n eta ∂adversarialCenteredNoiseLaw (N + 1) sigma) = ∫⁻ h, f h ∂adversarialClippedHistoryLaw algorithm sigma (adversarialFullHardShift gap i) n","missing":[],"search":"lintegral_adversarialnoisehistorykernel_eq_clipped banditrlproof.lowerbounds.lintegral_adversarialnoisehistorykernel_eq_clipped all observed finite prefixes have the canonical clipped history law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal","label":"adversarialNoiseHistoryJoint_history_marginal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal","description":"Full matrix-policy coupling: its history marginal is exactly Claim 17.6's law.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-495751989bff","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5634,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3020"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_history_marginal {N m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (sigma gap : Real) (i : Fin (m + 1)) (n : Nat) (hn : n < N + 1) : (adversarialNoiseHistoryJoint (horizon := N + 1) algorithm sigma gap i n).snd = adversarialClippedHistoryLaw algorithm sigma (adversarialFullHardShift gap i) n","missing":[],"search":"adversarialnoisehistoryjoint_history_marginal banditrlproof.lowerbounds.adversarialnoisehistoryjoint_history_marginal full matrix-policy coupling: its history marginal is exactly claim 17.6's law. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_pull_le_half_claim17_6","label":"adversarialNoiseHistoryJoint_pull_le_half_claim17_6","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_pull_le_half_claim17_6","description":"Corrected Claim 17.6 on the shared-noise matrix and policy joint space.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-18fcf088163f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5635,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3038"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_pull_le_half_claim17_6 {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (sigma delta : Real) (hs : sigma ≠ 0) (hd : 0 < delta) (hd8 : delta < 1 / 8) : ∃ i : Fin (m + 1), 2 * delta <= (adversarialNoiseHistoryJoint (horizon := n + 1) algorithm sigma (adversarialClaim17_6Gap (n + 1) m sigma delta) i n).real {p | finiteHistoryPullCountReal n p.2 i <= ((n + 1 : Nat) : Real) / 2}","missing":[],"search":"adversarialnoisehistoryjoint_pull_le_half_claim17_6 banditrlproof.lowerbounds.adversarialnoisehistoryjoint_pull_le_half_claim17_6 corrected claim 17.6 on the shared-noise matrix and policy joint space. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClippingCountReal","label":"measurable_adversarialClippingCountReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialClippingCountReal","description":"theorem measurable_adversarialClippingCountReal {horizon : Nat} (gap : Real) : Measurable (fun eta : Fin horizon -> Real => adversarialClippingCountReal eta gap)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-38a658f23bcf","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5636,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3062"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialClippingCountReal {horizon : Nat} (gap : Real) : Measurable (fun eta : Fin horizon -> Real => adversarialClippingCountReal eta gap)","missing":[],"search":"measurable_adversarialclippingcountreal banditrlproof.lowerbounds.measurable_adversarialclippingcountreal theorem measurable_adversarialclippingcountreal {horizon : nat} (gap : real) : measurable (fun eta : fin horizon -> real => adversarialclippingcountreal eta gap) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_clipping_tail","label":"adversarialNoiseHistoryJoint_clipping_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_clipping_tail","description":"theorem adversarialNoiseHistoryJoint_clipping_tail {horizon m : Nat} (hn : 0 < horizon) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (delta gap : Real) (hd : 0 < delta) (hd1 : delta < 1) (hg8 : gap < 1 / 8) (i : Fin (m + 1)) (n : Nat) (horizon_condition : 32 * Real.log (1 / delta) <= horizon) : (adversarialNoiseHistoryJoint (horizon := horizon) algorithm (1 / 10) gap i n).real {p | (horizon : Real) / 4…","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-ca0c2c388727","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5637,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3068"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_clipping_tail {horizon m : Nat} (hn : 0 < horizon) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (delta gap : Real) (hd : 0 < delta) (hd1 : delta < 1) (hg8 : gap < 1 / 8) (i : Fin (m + 1)) (n : Nat) (horizon_condition : 32 * Real.log (1 / delta) <= horizon) : (adversarialNoiseHistoryJoint (horizon := horizon) algorithm (1 / 10) gap i n).real {p | (horizon : Real) / 4 <= adversarialClippingCountReal p.1 gap} <= delta","missing":[],"search":"adversarialnoisehistoryjoint_clipping_tail banditrlproof.lowerbounds.adversarialnoisehistoryjoint_clipping_tail theorem adversarialnoisehistoryjoint_clipping_tail {horizon m : nat} (hn : 0 < horizon) (algorithm : thompson.historyalgorithm (fin (m + 1)) real) (delta gap : real) (hd : 0 < delta) (hd1 : delta < 1) (hg8 : gap < 1 / 8) (i : fin (m + 1)) (n : nat) (horizon_condition : 32 * real.log (1 / delta) <= horizon) : (adversarialnoisehistoryjoint (horizon := horizon) algorithm (1 / 10) gap i n).real {p | (horizon : real) / 4 <= adversarialclippingcountreal p.1 gap} <= delta theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_good_event","label":"adversarialNoiseHistoryJoint_good_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_good_event","description":"Pull-small and literally few clipped rounds hold jointly with probability at least delta.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1030c2360e3d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5638,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3087"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_good_event {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (delta : Real) (hd : 0 < delta) (hd8 : delta < 1 / 8) (hg8 : adversarialClaim17_6Gap (n + 1) m (1 / 10) delta < 1 / 8) (horizon_condition : 32 * Real.log (1 / delta) <= ((n + 1 : Nat) : Real)) : ∃ i : Fin (m + 1), delta <= (adversarialNoiseHistoryJoint (horizon := n + 1) algorithm (1 / 10) (adversarialClaim17_6Gap (n + 1) m (1 / 10) delta) i n).real {p | finiteHistoryPullCountReal n p.2 i <= ((n + 1 : Nat) : Real) / 2 ∧ adversarialFullBoundaryCount p.1 (adversarialClaim17_6Gap (n + 1) m (1 / 10) delta) i < ((n + 1 : Nat) : Real) / 4}","missing":[],"search":"adversarialnoisehistoryjoint_good_event banditrlproof.lowerbounds.adversarialnoisehistoryjoint_good_event pull-small and literally few clipped rounds hold jointly with probability at least delta. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHistoryActions","label":"adversarialHistoryActions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHistoryActions","description":"The realized action path, indexed by the actual number of observations.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-7f164efac27f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5639,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3118"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialHistoryActions {K : Nat} (n : Nat) (h : History.FinitePairHistory (Fin K) Real n) : Fin (n + 1) -> Fin K","missing":[],"search":"adversarialhistoryactions banditrlproof.lowerbounds.adversarialhistoryactions the realized action path, indexed by the actual number of observations. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHistoryActions_pullCountENNReal","label":"adversarialHistoryActions_pullCountENNReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHistoryActions_pullCountENNReal","description":"theorem adversarialHistoryActions_pullCountENNReal {K : Nat} (n : Nat) (h : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountENNReal n h arm = ∑ t, if adversarialHistoryActions n h t = arm then (1 : ENNReal) else 0","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-5afc549c717f","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5640,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3122"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHistoryActions_pullCountENNReal {K : Nat} (n : Nat) (h : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountENNReal n h arm = ∑ t, if adversarialHistoryActions n h t = arm then (1 : ENNReal) else 0","missing":[],"search":"adversarialhistoryactions_pullcountennreal banditrlproof.lowerbounds.adversarialhistoryactions_pullcountennreal theorem adversarialhistoryactions_pullcountennreal {k : nat} (n : nat) (h : history.finitepairhistory (fin k) real n) (arm : fin k) : finitehistorypullcountennreal n h arm = ∑ t, if adversarialhistoryactions n h t = arm then (1 : ennreal) else 0 theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHistoryActions_pullCountReal","label":"adversarialHistoryActions_pullCountReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHistoryActions_pullCountReal","description":"theorem adversarialHistoryActions_pullCountReal {K : Nat} (n : Nat) (h : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountReal n h arm = ∑ t, if adversarialHistoryActions n h t = arm then (1 : Real) else 0","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-fc6761e1f84e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5641,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHistoryActions_pullCountReal {K : Nat} (n : Nat) (h : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteHistoryPullCountReal n h arm = ∑ t, if adversarialHistoryActions n h t = arm then (1 : Real) else 0","missing":[],"search":"adversarialhistoryactions_pullcountreal banditrlproof.lowerbounds.adversarialhistoryactions_pullcountreal theorem adversarialhistoryactions_pullcountreal {k : nat} (n : nat) (h : history.finitepairhistory (fin k) real n) (arm : fin k) : finitehistorypullcountreal n h arm = ∑ t, if adversarialhistoryactions n h t = arm then (1 : real) else 0 theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHistory_randomRegret_ge_quarter","label":"adversarialHistory_randomRegret_ge_quarter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHistory_randomRegret_ge_quarter","description":"Eq. (17.8) on the same history coordinates as the joint good event.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-b9149831980d","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5642,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3151"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHistory_randomRegret_ge_quarter {m : Nat} (n : Nat) (eta : Fin (n + 1) -> Real) (gap : Real) (hg : 0 <= gap) (i : Fin (m + 1)) (h : History.FinitePairHistory (Fin (m + 1)) Real n) (hp : finiteHistoryPullCountReal n h i <= ((n + 1 : Nat) : Real) / 2) (hc : adversarialFullBoundaryCount eta gap i <= ((n + 1 : Nat) : Real) / 4) : gap * (((n + 1 : Nat) : Real) / 4) <= adversarialRandomRegret (adversarialFullClippedReward eta gap i) (adversarialHistoryActions n h)","missing":[],"search":"adversarialhistory_randomregret_ge_quarter banditrlproof.lowerbounds.adversarialhistory_randomregret_ge_quarter eq. (17.8) on the same history coordinates as the joint good event. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_randomRegret_tail","label":"adversarialNoiseHistoryJoint_randomRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_randomRegret_tail","description":"Random-regret tail on the coupled hard matrix, before final constant calibration.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-166761f8378e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5643,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3166"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialNoiseHistoryJoint_randomRegret_tail {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (delta : Real) (hd : 0 < delta) (hd8 : delta < 1 / 8) (hg8 : adversarialClaim17_6Gap (n + 1) m (1 / 10) delta < 1 / 8) (horizon_condition : 32 * Real.log (1 / delta) <= ((n + 1 : Nat) : Real)) : ∃ i : Fin (m + 1), delta <= (adversarialNoiseHistoryJoint (horizon := n + 1) algorithm (1 / 10) (adversarialClaim17_6Gap (n + 1) m (1 / 10) delta) i n).real {p | adversarialClaim17_6Gap (n + 1) m (1 / 10) delta * (((n + 1 : Nat) : Real) / 4) <= adversarialRandomRegret (adversarialFullClippedReward p.1 (adversarialClaim17_6Gap (n + 1) m (1 / 10) delta) i) (adversarialHistoryActions n p.2)}","missing":[],"search":"adversarialnoisehistoryjoint_randomregret_tail banditrlproof.lowerbounds.adversarialnoisehistoryjoint_randomregret_tail random-regret tail on the coupled hard matrix, before final constant calibration. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_kernel_section_mass_ge","label":"exists_kernel_section_mass_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_kernel_section_mass_ge","description":"First-moment extraction for a measurable kernel event. Mathlib candidate.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-8526a3a9e99a","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5644,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3186"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_kernel_section_mass_ge {X Y : Type*} [MeasurableSpace X] [MeasurableSpace Y] (μ : Measure X) [IsProbabilityMeasure μ] (κ : Kernel X Y) [IsMarkovKernel κ] (E : Set (X × Y)) (hE : MeasurableSet E) (delta : Real) (hd : delta <= (μ ⊗ₘ κ).real E) : ∃ x, delta <= (κ x).real (Prod.mk x ⁻¹' E)","missing":[],"search":"exists_kernel_section_mass_ge banditrlproof.lowerbounds.exists_kernel_section_mass_ge first-moment extraction for a measurable kernel event. mathlib candidate. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialRandomRegret","label":"measurable_adversarialRandomRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialRandomRegret","description":"theorem measurable_adversarialRandomRegret {horizon m : Nat} (actions : Fin horizon -> Fin (m + 1)) : Measurable (fun reward : Fin horizon -> Fin (m + 1) -> Real => adversarialRandomRegret reward actions)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-76bc57bda9b6","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5645,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3205"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialRandomRegret {horizon m : Nat} (actions : Fin horizon -> Fin (m + 1)) : Measurable (fun reward : Fin horizon -> Fin (m + 1) -> Real => adversarialRandomRegret reward actions)","missing":[],"search":"measurable_adversarialrandomregret banditrlproof.lowerbounds.measurable_adversarialrandomregret theorem measurable_adversarialrandomregret {horizon m : nat} (actions : fin horizon -> fin (m + 1)) : measurable (fun reward : fin horizon -> fin (m + 1) -> real => adversarialrandomregret reward actions) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialJointRandomRegret","label":"measurable_adversarialJointRandomRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialJointRandomRegret","description":"theorem measurable_adversarialJointRandomRegret {m : Nat} (n : Nat) (gap : Real) (i : Fin (m + 1)) : Measurable (fun p : (Fin (n + 1) -> Real) × History.FinitePairHistory (Fin (m + 1)) Real n => adversarialRandomRegret (adversarialFullClippedReward p.1 gap i) (adversarialHistoryActions n p.2))","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-783f307caa0b","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5646,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3223"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialJointRandomRegret {m : Nat} (n : Nat) (gap : Real) (i : Fin (m + 1)) : Measurable (fun p : (Fin (n + 1) -> Real) × History.FinitePairHistory (Fin (m + 1)) Real n => adversarialRandomRegret (adversarialFullClippedReward p.1 gap i) (adversarialHistoryActions n p.2))","missing":[],"search":"measurable_adversarialjointrandomregret banditrlproof.lowerbounds.measurable_adversarialjointrandomregret theorem measurable_adversarialjointrandomregret {m : nat} (n : nat) (gap : real) (i : fin (m + 1)) : measurable (fun p : (fin (n + 1) -> real) × history.finitepairhistory (fin (m + 1)) real n => adversarialrandomregret (adversarialfullclippedreward p.1 gap i) (adversarialhistoryactions n p.2)) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_adversarialTable_randomRegret_tail","label":"exists_adversarialTable_randomRegret_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_adversarialTable_randomRegret_tail","description":"A deterministic bounded reward table realizes the uncalibrated random-regret tail.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-1e85f9966728","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5647,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3254"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_adversarialTable_randomRegret_tail {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (delta : Real) (hd : 0 < delta) (hd8 : delta < 1 / 8) (hg8 : adversarialClaim17_6Gap (n + 1) m (1 / 10) delta < 1 / 8) (horizon_condition : 32 * Real.log (1 / delta) <= ((n + 1 : Nat) : Real)) : ∃ table : AdversarialRewardTable (m + 1), (∀ t arm, table t arm ∈ Set.Icc (0 : Real) 1) ∧ delta <= (adversarialTableHistoryKernel algorithm n table).real {h | adversarialClaim17_6Gap (n + 1) m (1 / 10) delta * (((n + 1 : Nat) : Real) / 4) <= adversarialRandomRegret (fun t : Fin (n + 1) => table t.val) (adversarialHistoryActions n h)}","missing":[],"search":"exists_adversarialtable_randomregret_tail banditrlproof.lowerbounds.exists_adversarialtable_randomregret_tail a deterministic bounded reward table realizes the uncalibrated random-regret tail. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialConfidence_log_calibration","label":"adversarialConfidence_log_calibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialConfidence_log_calibration","description":"Explicit logarithmic comparison on the corrected confidence domain.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-27928e5956c4","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5648,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3290"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialConfidence_log_calibration (delta : Real) (hd : 0 < delta) (hd32 : delta <= 1 / 32) : 0 < Real.log (1 / (2 * delta)) ∧ Real.log (1 / (2 * delta)) / 2 <= Real.log (1 / (8 * delta)) ∧ Real.log (1 / delta) <= 2 * Real.log (1 / (2 * delta)) ∧ Real.log (1 / (8 * delta)) <= Real.log (1 / (2 * delta))","missing":[],"search":"adversarialconfidence_log_calibration banditrlproof.lowerbounds.adversarialconfidence_log_calibration explicit logarithmic comparison on the corrected confidence domain. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialClaim17_6Gap_tenth_sq","label":"adversarialClaim17_6Gap_tenth_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialClaim17_6Gap_tenth_sq","description":"theorem adversarialClaim17_6Gap_tenth_sq {N m : Nat} (hN : 0 < N) (hm : 0 < m) (delta : Real) (hd : 0 < delta) (hd8 : delta < 1 / 8) : (adversarialClaim17_6Gap N m (1 / 10) delta) ^ 2 = (m : Real) * Real.log (1 / (8 * delta)) / (200 * N)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-92100a5cf327","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5649,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3318"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialClaim17_6Gap_tenth_sq {N m : Nat} (hN : 0 < N) (hm : 0 < m) (delta : Real) (hd : 0 < delta) (hd8 : delta < 1 / 8) : (adversarialClaim17_6Gap N m (1 / 10) delta) ^ 2 = (m : Real) * Real.log (1 / (8 * delta)) / (200 * N)","missing":[],"search":"adversarialclaim17_6gap_tenth_sq banditrlproof.lowerbounds.adversarialclaim17_6gap_tenth_sq theorem adversarialclaim17_6gap_tenth_sq {n m : nat} (hn : 0 < n) (hm : 0 < m) (delta : real) (hd : 0 < delta) (hd8 : delta < 1 / 8) : (adversarialclaim17_6gap n m (1 / 10) delta) ^ 2 = (m : real) * real.log (1 / (8 * delta)) / (200 * n) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialHorizon_calibration","label":"adversarialHorizon_calibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialHorizon_calibration","description":"The source clipping and gap conditions follow from a source-form horizon bound.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-01482c9bbb34","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5650,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3332"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialHorizon_calibration {N m : Nat} (hN : 0 < N) (hm : 0 < m) (delta : Real) (hd : 0 < delta) (hd32 : delta <= 1 / 32) (horizon : 64 * ((m + 1 : Nat) : Real) * Real.log (1 / (2 * delta)) <= N) : adversarialClaim17_6Gap N m (1 / 10) delta < 1 / 8 ∧ 32 * Real.log (1 / delta) <= N","missing":[],"search":"adversarialhorizon_calibration banditrlproof.lowerbounds.adversarialhorizon_calibration the source clipping and gap conditions follow from a source-form horizon bound. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialThreshold_calibration","label":"adversarialThreshold_calibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialThreshold_calibration","description":"Strict slack converts the construction's non-strict event to a CDF-complement event.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-b956edd28977","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5651,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3356"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialThreshold_calibration {N m : Nat} (hN : 0 < N) (hm : 0 < m) (delta : Real) (hd : 0 < delta) (hd32 : delta <= 1 / 32) : adversarialHighProbabilityThreshold N (m + 1) (1 / 160) delta < adversarialClaim17_6Gap N m (1 / 10) delta * ((N : Real) / 4)","missing":[],"search":"adversarialthreshold_calibration banditrlproof.lowerbounds.adversarialthreshold_calibration strict slack converts the construction's non-strict event to a cdf-complement event. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_adversarialTable_randomRegret_gt_theorem17_4","label":"exists_adversarialTable_randomRegret_gt_theorem17_4","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_adversarialTable_randomRegret_gt_theorem17_4","description":"*Corrected Theorem 17.4.** Explicit constants and confidence domain; the event is strict random regret, for a deterministic bounded reward table.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-187117f58ccf","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5652,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3394"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_adversarialTable_randomRegret_gt_theorem17_4 {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (delta : Real) (hd : 0 < delta) (hd32 : delta <= 1 / 32) (horizon : 64 * ((m + 1 : Nat) : Real) * Real.log (1 / (2 * delta)) <= ((n + 1 : Nat) : Real)) : ∃ table : AdversarialRewardTable (m + 1), (∀ t arm, table t arm ∈ Set.Icc (0 : Real) 1) ∧ delta <= (adversarialTableHistoryKernel algorithm n table).real {h | adversarialHighProbabilityThreshold (n + 1) (m + 1) (1 / 160) delta < adversarialRandomRegret (fun t : Fin (n + 1) => table t.val) (adversarialHistoryActions n h)}","missing":[],"search":"exists_adversarialtable_randomregret_gt_theorem17_4 banditrlproof.lowerbounds.exists_adversarialtable_randomregret_gt_theorem17_4 *corrected theorem 17.4.** explicit constants and confidence domain; the event is strict random regret, for a deterministic bounded reward table. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableRandomRegret","label":"adversarialTableRandomRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableRandomRegret","description":"def adversarialTableRandomRegret {m : Nat} (table : AdversarialRewardTable (m + 1)) (n : Nat) (h : History.FinitePairHistory (Fin (m + 1)) Real n) : Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-55090aa0842e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5653,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3410"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def adversarialTableRandomRegret {m : Nat} (table : AdversarialRewardTable (m + 1)) (n : Nat) (h : History.FinitePairHistory (Fin (m + 1)) Real n) : Real","missing":[],"search":"adversarialtablerandomregret banditrlproof.lowerbounds.adversarialtablerandomregret def adversarialtablerandomregret {m : nat} (table : adversarialrewardtable (m + 1)) (n : nat) (h : history.finitepairhistory (fin (m + 1)) real n) : real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableExpectedRegret","label":"adversarialTableExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableExpectedRegret","description":"Deterministic expectation, separate from the pathwise random-regret variable.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-e56ced69789e","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5654,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3415"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialTableExpectedRegret {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (table : AdversarialRewardTable (m + 1)) (n : Nat) : Real","missing":[],"search":"adversarialtableexpectedregret banditrlproof.lowerbounds.adversarialtableexpectedregret deterministic expectation, separate from the pathwise random-regret variable. definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTableCDF","label":"adversarialTableCDF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTableCDF","description":"noncomputable def adversarialTableCDF {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (table : AdversarialRewardTable (m + 1)) (n : Nat) (u : Real) : Real","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-698d219227f1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5655,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3420"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def adversarialTableCDF {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (table : AdversarialRewardTable (m + 1)) (n : Nat) (u : Real) : Real","missing":[],"search":"adversarialtablecdf banditrlproof.lowerbounds.adversarialtablecdf noncomputable def adversarialtablecdf {m : nat} (algorithm : thompson.historyalgorithm (fin (m + 1)) real) (table : adversarialrewardtable (m + 1)) (n : nat) (u : real) : real definition compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_adversarialTableRandomRegret","label":"measurable_adversarialTableRandomRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_adversarialTableRandomRegret","description":"theorem measurable_adversarialTableRandomRegret {m : Nat} (table : AdversarialRewardTable (m + 1)) (n : Nat) : Measurable (adversarialTableRandomRegret table n)","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-aa854dc61724","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5656,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3425"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_adversarialTableRandomRegret {m : Nat} (table : AdversarialRewardTable (m + 1)) (n : Nat) : Measurable (adversarialTableRandomRegret table n)","missing":[],"search":"measurable_adversarialtablerandomregret banditrlproof.lowerbounds.measurable_adversarialtablerandomregret theorem measurable_adversarialtablerandomregret {m : nat} (table : adversarialrewardtable (m + 1)) (n : nat) : measurable (adversarialtablerandomregret table n) theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_adversarialTableRandomRegret","label":"integrable_adversarialTableRandomRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_adversarialTableRandomRegret","description":"A fixed finite reward table has only finitely many possible action-path regrets, so its expectation is a genuine integrable random variable.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-63c77f9fd832","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5657,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3449"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_adversarialTableRandomRegret {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (table : AdversarialRewardTable (m + 1)) (n : Nat) : Integrable (adversarialTableRandomRegret table n) (adversarialTableHistoryKernel algorithm n table)","missing":[],"search":"integrable_adversarialtablerandomregret banditrlproof.lowerbounds.integrable_adversarialtablerandomregret a fixed finite reward table has only finitely many possible action-path regrets, so its expectation is a genuine integrable random variable. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialTable_strictTail_eq_one_sub_CDF","label":"adversarialTable_strictTail_eq_one_sub_CDF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialTable_strictTail_eq_one_sub_CDF","description":"theorem adversarialTable_strictTail_eq_one_sub_CDF {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (table : AdversarialRewardTable (m + 1)) (n : Nat) (u : Real) : (adversarialTableHistoryKernel algorithm n table).real {h | u < adversarialTableRandomRegret table n h} = 1 - adversarialTableCDF algorithm table n u","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-35d898832212","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5658,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3462"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem adversarialTable_strictTail_eq_one_sub_CDF {m : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (table : AdversarialRewardTable (m + 1)) (n : Nat) (u : Real) : (adversarialTableHistoryKernel algorithm n table).real {h | u < adversarialTableRandomRegret table n h} = 1 - adversarialTableCDF algorithm table n u","missing":[],"search":"adversarialtable_stricttail_eq_one_sub_cdf banditrlproof.lowerbounds.adversarialtable_stricttail_eq_one_sub_cdf theorem adversarialtable_stricttail_eq_one_sub_cdf {m : nat} (algorithm : thompson.historyalgorithm (fin (m + 1)) real) (table : adversarialrewardtable (m + 1)) (n : nat) (u : real) : (adversarialtablehistorykernel algorithm n table).real {h | u < adversarialtablerandomregret table n h} = 1 - adversarialtablecdf algorithm table n u theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_theorem17_4","label":"adversarialRandomRegret_ge_theorem17_4","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.adversarialRandomRegret_ge_theorem17_4","description":"Corrected Theorem 17.4 in the source's CDF-complement notation.","url":"../modules/banditrlproof-lowerbounds-highprobability/index.html#decl-01eaad33d4e1","parent":"module:BanditRLProof.LowerBounds.HighProbability","order":5659,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HighProbability"],["Source","BanditRLProof/LowerBounds/HighProbability.lean:3473"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-17-high-probability"],["Indexed settings","None registered"]],"statement":"theorem adversarialRandomRegret_ge_theorem17_4 {m : Nat} (hm : 0 < m) (algorithm : Thompson.HistoryAlgorithm (Fin (m + 1)) Real) (n : Nat) (delta : Real) (hd : 0 < delta) (hd32 : delta <= 1 / 32) (horizon : 64 * ((m + 1 : Nat) : Real) * Real.log (1 / (2 * delta)) <= ((n + 1 : Nat) : Real)) : ∃ table : AdversarialRewardTable (m + 1), (∀ t arm, table t arm ∈ Set.Icc (0 : Real) 1) ∧ delta <= 1 - adversarialTableCDF algorithm table n (adversarialHighProbabilityThreshold (n + 1) (m + 1) (1 / 160) delta)","missing":[],"search":"adversarialrandomregret_ge_theorem17_4 banditrlproof.lowerbounds.adversarialrandomregret_ge_theorem17_4 corrected theorem 17.4 in the source's cdf-complement notation. theorem compiled","shard":"modules/c48a84f74677737d.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-17-high-probability"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_relabel","label":"expectedCodeLength_relabel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_relabel","description":"theorem expectedCodeLength_relabel {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (code : BinaryPrefixCode α) (e : β ≃ α) : expectedCodeLength (p ∘ e) (code.relabel e) = expectedCodeLength p code","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-478dd8aff25d","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5660,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_relabel {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (code : BinaryPrefixCode α) (e : β ≃ α) : expectedCodeLength (p ∘ e) (code.relabel e) = expectedCodeLength p code","missing":[],"search":"expectedcodelength_relabel banditrlproof.lowerbounds.expectedcodelength_relabel theorem expectedcodelength_relabel {α β : type*} [fintype α] [fintype β] (p : α → ℝ) (code : binaryprefixcode α) (e : β ≃ α) : expectedcodelength (p ∘ e) (code.relabel e) = expectedcodelength p code theorem compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.relabel","label":"relabel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsOptimalPrefixCode.relabel","description":"Global optimality is independent of the names of the alphabet symbols.","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-bc8dd165f649","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5661,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsOptimalPrefixCode.relabel {α β : Type*} [Fintype α] [Fintype β] (p : α → ℝ) (code : BinaryPrefixCode α) (e : β ≃ α) (hopt : IsOptimalPrefixCode p code) : IsOptimalPrefixCode (p ∘ e) (code.relabel e)","missing":[],"search":"relabel banditrlproof.lowerbounds.isoptimalprefixcode.relabel global optimality is independent of the names of the alphabet symbols. theorem compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.HuffmanRemainder","label":"HuffmanRemainder","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.HuffmanRemainder","description":"The unmerged symbols, retaining their original labels.","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-7b9477765b11","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5662,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def HuffmanRemainder {α : Type*} (a b : α)","missing":[],"search":"huffmanremainder banditrlproof.lowerbounds.huffmanremainder the unmerged symbols, retaining their original labels. definition compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanSplitEquiv","label":"huffmanSplitEquiv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanSplitEquiv","description":"Separate two selected symbols as the false and true leaves.","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-ebcd539e1ee8","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5663,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def huffmanSplitEquiv {α : Type*} [DecidableEq α] (a b : α) (hab : a ≠ b) : HuffmanRemainder a b ⊕ Bool ≃ α where","missing":[],"search":"huffmansplitequiv banditrlproof.lowerbounds.huffmansplitequiv separate two selected symbols as the false and true leaves. definition compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanSplitEquiv_false","label":"huffmanSplitEquiv_false","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanSplitEquiv_false","description":"@[simp] theorem huffmanSplitEquiv_false {α : Type*} [DecidableEq α] (a b : α) (hab : a ≠ b) : huffmanSplitEquiv a b hab (.inr false) = a","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-39e72fdfeecb","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5664,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem huffmanSplitEquiv_false {α : Type*} [DecidableEq α] (a b : α) (hab : a ≠ b) : huffmanSplitEquiv a b hab (.inr false) = a","missing":[],"search":"huffmansplitequiv_false banditrlproof.lowerbounds.huffmansplitequiv_false @[simp] theorem huffmansplitequiv_false {α : type*} [decidableeq α] (a b : α) (hab : a ≠ b) : huffmansplitequiv a b hab (.inr false) = a theorem compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanSplitEquiv_true","label":"huffmanSplitEquiv_true","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanSplitEquiv_true","description":"@[simp] theorem huffmanSplitEquiv_true {α : Type*} [DecidableEq α] (a b : α) (hab : a ≠ b) : huffmanSplitEquiv a b hab (.inr true) = b","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-24f6d977ac88","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5665,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem huffmanSplitEquiv_true {α : Type*} [DecidableEq α] (a b : α) (hab : a ≠ b) : huffmanSplitEquiv a b hab (.inr true) = b","missing":[],"search":"huffmansplitequiv_true banditrlproof.lowerbounds.huffmansplitequiv_true @[simp] theorem huffmansplitequiv_true {α : type*} [decidableeq α] (a b : α) (hab : a ≠ b) : huffmansplitequiv a b hab (.inr true) = b theorem compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffman_merged_card_lt","label":"huffman_merged_card_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffman_merged_card_lt","description":"Merging two distinct symbols reduces the recursive alphabet size by one.","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-c2b08b7b6b34","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5666,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem huffman_merged_card_lt {α : Type*} [Fintype α] [DecidableEq α] (a b : α) (hab : a ≠ b) : Fintype.card (Option (HuffmanRemainder a b)) < Fintype.card α","missing":[],"search":"huffman_merged_card_lt banditrlproof.lowerbounds.huffman_merged_card_lt merging two distinct symbols reduces the recursive alphabet size by one. theorem compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_two_least_weights","label":"exists_two_least_weights","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_two_least_weights","description":"A finite nontrivial alphabet has two least weights, including ties.","url":"../modules/banditrlproof-lowerbounds-huffmanalphabet/index.html#decl-62fdf8d814f2","parent":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","order":5667,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanAlphabet"],["Source","BanditRLProof/LowerBounds/HuffmanAlphabet.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_two_least_weights {α : Type*} [Fintype α] [DecidableEq α] [Nontrivial α] (p : α → ℝ) : ∃ a b, a ≠ b ∧ (∀ i, p a ≤ p i) ∧ (∀ i, i ≠ a → p b ≤ p i)","missing":[],"search":"exists_two_least_weights banditrlproof.lowerbounds.exists_two_least_weights a finite nontrivial alphabet has two least weights, including ties. theorem compiled","shard":"modules/b32960fded537975.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneBitCode_optimal","label":"oneBitCode_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneBitCode_optimal","description":"theorem oneBitCode_optimal {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (code : BinaryPrefixCode α) (hlen : ∀ i, (code.encode i).length = 1) : IsOptimalPrefixCode p code","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html#decl-5374f102b474","parent":"module:BanditRLProof.LowerBounds.HuffmanConstruction","order":5668,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanConstruction"],["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oneBitCode_optimal {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (code : BinaryPrefixCode α) (hlen : ∀ i, (code.encode i).length = 1) : IsOptimalPrefixCode p code","missing":[],"search":"onebitcode_optimal banditrlproof.lowerbounds.onebitcode_optimal theorem onebitcode_optimal {α : type*} [fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (code : binaryprefixcode α) (hlen : ∀ i, (code.encode i).length = 1) : isoptimalprefixcode p code theorem compiled","shard":"modules/6f010ee120079082.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.emptyRemainderRoot","label":"emptyRemainderRoot","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.emptyRemainderRoot","description":"def emptyRemainderRoot {α : Type*} [IsEmpty α] : BinaryPrefixCode (α ⊕ Bool) where","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html#decl-d3fa57425419","parent":"module:BanditRLProof.LowerBounds.HuffmanConstruction","order":5669,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HuffmanConstruction"],["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def emptyRemainderRoot {α : Type*} [IsEmpty α] : BinaryPrefixCode (α ⊕ Bool) where","missing":[],"search":"emptyremainderroot banditrlproof.lowerbounds.emptyremainderroot def emptyremainderroot {α : type*} [isempty α] : binaryprefixcode (α ⊕ bool) where definition compiled","shard":"modules/6f010ee120079082.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanOptimalCode","label":"huffmanOptimalCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanOptimalCode","description":"Huffman's recursive merge-two-least construction, with its correctness proof. Real-weight choices are classical; the code itself is assembled recursively.","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html#decl-0ba1fc5f6500","parent":"module:BanditRLProof.LowerBounds.HuffmanConstruction","order":5670,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HuffmanConstruction"],["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def huffmanOptimalCode {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) : {code : BinaryPrefixCode α // IsOptimalPrefixCode p code}","missing":[],"search":"huffmanoptimalcode banditrlproof.lowerbounds.huffmanoptimalcode huffman's recursive merge-two-least construction, with its correctness proof. real-weight choices are classical; the code itself is assembled recursively. definition compiled","shard":"modules/6f010ee120079082.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanCode","label":"huffmanCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanCode","description":"The prefix code produced by recursive Huffman merging.","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html#decl-ddef97de2fd6","parent":"module:BanditRLProof.LowerBounds.HuffmanConstruction","order":5671,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.HuffmanConstruction"],["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def huffmanCode {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) : BinaryPrefixCode α","missing":[],"search":"huffmancode banditrlproof.lowerbounds.huffmancode the prefix code produced by recursive huffman merging. definition compiled","shard":"modules/6f010ee120079082.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanCode_optimal","label":"huffmanCode_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanCode_optimal","description":"theorem huffmanCode_optimal {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) : IsOptimalPrefixCode p (huffmanCode p hp)","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html#decl-fbcab29817b6","parent":"module:BanditRLProof.LowerBounds.HuffmanConstruction","order":5672,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanConstruction"],["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem huffmanCode_optimal {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) : IsOptimalPrefixCode p (huffmanCode p hp)","missing":[],"search":"huffmancode_optimal banditrlproof.lowerbounds.huffmancode_optimal theorem huffmancode_optimal {α : type*} [fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) : isoptimalprefixcode p (huffmancode p hp) theorem compiled","shard":"modules/6f010ee120079082.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.huffmanCode_entropy_sandwich","label":"huffmanCode_entropy_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.huffmanCode_entropy_sandwich","description":"Chapter 14, Eq. (14.2), for the recursively constructed Huffman code.","url":"../modules/banditrlproof-lowerbounds-huffmanconstruction/index.html#decl-c55c19f6aa65","parent":"module:BanditRLProof.LowerBounds.HuffmanConstruction","order":5673,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanConstruction"],["Source","BanditRLProof/LowerBounds/HuffmanConstruction.lean:108"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem huffmanCode_entropy_sandwich {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength p (huffmanCode p hp) ∧ expectedCodeLength p (huffmanCode p hp) ≤ discreteEntropyBaseTwo Finset.univ p + 1","missing":[],"search":"huffmancode_entropy_sandwich banditrlproof.lowerbounds.huffmancode_entropy_sandwich chapter 14, eq. (14.2), for the recursively constructed huffman code. theorem compiled","shard":"modules/6f010ee120079082.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_oriented_sibling_code","label":"exists_oriented_sibling_code","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_oriented_sibling_code","description":"Orient an actual pair of sibling leaves without changing its cost.","url":"../modules/banditrlproof-lowerbounds-huffmanstep/index.html#decl-86b0193cc23b","parent":"module:BanditRLProof.LowerBounds.HuffmanStep","order":5674,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanStep"],["Source","BanditRLProof/LowerBounds/HuffmanStep.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_oriented_sibling_code {α : Type*} [Fintype α] [DecidableEq α] (p : α ⊕ Bool → ℝ) (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (bit : Bool) (hf : code.encode (.inr false) = w ++ [bit]) (ht : code.encode (.inr true) = w ++ [!bit]) : ∃ other : BinaryPrefixCode (α ⊕ Bool), expectedCodeLength p other = expectedCodeLength p code ∧ ∀ b, other.encode (.inr b) = w ++ [b]","missing":[],"search":"exists_oriented_sibling_code banditrlproof.lowerbounds.exists_oriented_sibling_code orient an actual pair of sibling leaves without changing its cost. theorem compiled","shard":"modules/02427867ffc46e50.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.expand_least_weights","label":"expand_least_weights","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsOptimalPrefixCode.expand_least_weights","description":"The Huffman induction step: expanding an optimal merged code is globally optimal when the split symbols are the two least weights.","url":"../modules/banditrlproof-lowerbounds-huffmanstep/index.html#decl-efd9a2512758","parent":"module:BanditRLProof.LowerBounds.HuffmanStep","order":5675,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.HuffmanStep"],["Source","BanditRLProof/LowerBounds/HuffmanStep.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsOptimalPrefixCode.expand_least_weights {α : Type*} [Fintype α] [DecidableEq α] [Nonempty α] (p : α → ℝ) (q r : ℝ) (hp : ∀ i, 0 ≤ p i) (hq : 0 ≤ q) (hqr : q ≤ r) (hr : ∀ i, r ≤ p i) (code : BinaryPrefixCode (Option α)) (hopt : IsOptimalPrefixCode (fun a => a.elim (q + r) p) code) : IsOptimalPrefixCode (Sum.elim p (fun b => if b then r else q)) code.expandSibling","missing":[],"search":"expand_least_weights banditrlproof.lowerbounds.isoptimalprefixcode.expand_least_weights the huffman induction step: expanding an optimal merged code is globally optimal when the split symbols are the two least weights. theorem compiled","shard":"modules/02427867ffc46e50.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode","label":"BinaryPrefixCode","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode","description":"A finite-alphabet binary prefix code. Excluding the empty codeword is the regularity condition needed for concatenations of repeated messages to be uniquely decodable.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-0dedbeba2d1c","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5676,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:34"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"structure BinaryPrefixCode (Symbol : Type*) where","missing":[],"search":"binaryprefixcode banditrlproof.lowerbounds.binaryprefixcode a finite-alphabet binary prefix code. excluding the empty codeword is the regularity condition needed for concatenations of repeated messages to be uniquely decodable. structure compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.uniquelyDecodable_range","label":"uniquelyDecodable_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.uniquelyDecodable_range","description":"A prefix-free codebook with no empty codeword is uniquely decodable.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-eaf078025321","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5677,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniquelyDecodable_range (code : BinaryPrefixCode Symbol) : InformationTheory.UniquelyDecodable (Set.range code.encode)","missing":[],"search":"uniquelydecodable_range banditrlproof.lowerbounds.binaryprefixcode.uniquelydecodable_range a prefix-free codebook with no empty codeword is uniquely decodable. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.codebook","label":"codebook","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.codebook","description":"The finite set of codewords induced by a finite source alphabet.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-72c7add430b6","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5678,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def codebook [Fintype Symbol] [DecidableEq Symbol] (code : BinaryPrefixCode Symbol) : Finset (List Bool)","missing":[],"search":"codebook banditrlproof.lowerbounds.binaryprefixcode.codebook the finite set of codewords induced by a finite source alphabet. definition compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.coe_codebook","label":"coe_codebook","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.coe_codebook","description":"theorem coe_codebook [Fintype Symbol] [DecidableEq Symbol] (code : BinaryPrefixCode Symbol) : (code.codebook : Set (List Bool)) = Set.range code.encode","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-987fba051eec","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5679,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:103"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem coe_codebook [Fintype Symbol] [DecidableEq Symbol] (code : BinaryPrefixCode Symbol) : (code.codebook : Set (List Bool)) = Set.range code.encode","missing":[],"search":"coe_codebook banditrlproof.lowerbounds.binaryprefixcode.coe_codebook theorem coe_codebook [fintype symbol] [decidableeq symbol] (code : binaryprefixcode symbol) : (code.codebook : set (list bool)) = set.range code.encode theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.kraft_inequality","label":"kraft_inequality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.kraft_inequality","description":"Kraft--McMillan for a finite binary prefix code, obtained by adapting the codebook to Mathlib's uniquely-decodable-code theorem.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-ccdf46d24fa0","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5680,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:111"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem kraft_inequality [Fintype Symbol] [DecidableEq Symbol] (code : BinaryPrefixCode Symbol) : ∑ word ∈ code.codebook, (1 / 2 : Real) ^ word.length ≤ 1","missing":[],"search":"kraft_inequality banditrlproof.lowerbounds.binaryprefixcode.kraft_inequality kraft--mcmillan for a finite binary prefix code, obtained by adapting the codebook to mathlib's uniquely-decodable-code theorem. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropy","label":"discreteEntropy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropy","description":"Natural-log entropy (nats) of a finite supported mass function, Eq. (14.3).","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-847aa1096514","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5681,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:121"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def discreteEntropy (support : Finset Symbol) (probability : Symbol → Real) : Real","missing":[],"search":"discreteentropy banditrlproof.lowerbounds.discreteentropy natural-log entropy (nats) of a finite supported mass function, eq. (14.3). definition compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo","label":"discreteEntropyBaseTwo","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropyBaseTwo","description":"Base-two entropy (bits) of a finite supported mass function.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-17eb3ee1926a","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5682,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:126"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def discreteEntropyBaseTwo (support : Finset Symbol) (probability : Symbol → Real) : Real","missing":[],"search":"discreteentropybasetwo banditrlproof.lowerbounds.discreteentropybasetwo base-two entropy (bits) of a finite supported mass function. definition compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","label":"discreteEntropyBaseTwo_eq_div_log_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","description":"theorem discreteEntropyBaseTwo_eq_div_log_two (support : Finset Symbol) (probability : Symbol → Real) : discreteEntropyBaseTwo support probability = discreteEntropy support probability / Real.log 2","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-20c3bf9feeb6","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5683,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:131"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropyBaseTwo_eq_div_log_two (support : Finset Symbol) (probability : Symbol → Real) : discreteEntropyBaseTwo support probability = discreteEntropy support probability / Real.log 2","missing":[],"search":"discreteentropybasetwo_eq_div_log_two banditrlproof.lowerbounds.discreteentropybasetwo_eq_div_log_two theorem discreteentropybasetwo_eq_div_log_two (support : finset symbol) (probability : symbol → real) : discreteentropybasetwo support probability = discreteentropy support probability / real.log 2 theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.discreteEntropy_nonneg","label":"discreteEntropy_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.discreteEntropy_nonneg","description":"theorem discreteEntropy_nonneg (support : Finset Symbol) (probability : Symbol → Real) (hprobability : ∀ symbol ∈ support, 0 ≤ probability symbol ∧ probability symbol ≤ 1) : 0 ≤ discreteEntropy support probability","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-1a617df72f09","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5684,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:138"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem discreteEntropy_nonneg (support : Finset Symbol) (probability : Symbol → Real) (hprobability : ∀ symbol ∈ support, 0 ≤ probability symbol ∧ probability symbol ≤ 1) : 0 ≤ discreteEntropy support probability","missing":[],"search":"discreteentropy_nonneg banditrlproof.lowerbounds.discreteentropy_nonneg theorem discreteentropy_nonneg (support : finset symbol) (probability : symbol → real) (hprobability : ∀ symbol ∈ support, 0 ≤ probability symbol ∧ probability symbol ≤ 1) : 0 ≤ discreteentropy support probability theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength","label":"expectedCodeLength","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength","description":"Expected binary codeword length, the objective in Eq. (14.1).","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-51734c977ce9","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5685,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:152"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedCodeLength [Fintype Symbol] (probability : Symbol → Real) (code : BinaryPrefixCode Symbol) : Real","missing":[],"search":"expectedcodelength banditrlproof.lowerbounds.expectedcodelength expected binary codeword length, the objective in eq. (14.1). definition compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_nonneg","label":"expectedCodeLength_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_nonneg","description":"theorem expectedCodeLength_nonneg [Fintype Symbol] (probability : Symbol → Real) (code : BinaryPrefixCode Symbol) (hprobability : ∀ symbol, 0 ≤ probability symbol) : 0 ≤ expectedCodeLength probability code","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-5673fdc29dcc","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5686,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:156"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_nonneg [Fintype Symbol] (probability : Symbol → Real) (code : BinaryPrefixCode Symbol) (hprobability : ∀ symbol, 0 ≤ probability symbol) : 0 ≤ expectedCodeLength probability code","missing":[],"search":"expectedcodelength_nonneg banditrlproof.lowerbounds.expectedcodelength_nonneg theorem expectedcodelength_nonneg [fintype symbol] (probability : symbol → real) (code : binaryprefixcode symbol) (hprobability : ∀ symbol, 0 ≤ probability symbol) : 0 ≤ expectedcodelength probability code theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy","label":"relativeEntropy","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy","description":"Chapter 14 relative entropy, with value `∞` on support mismatch or a non-integrable log-likelihood ratio.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-e58d2272a11a","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5687,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:168"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"abbrev relativeEntropy {α : Type*} [MeasurableSpace α] (P Q : Measure α) : ENNReal","missing":[],"search":"relativeentropy banditrlproof.lowerbounds.relativeentropy chapter 14 relative entropy, with value `∞` on support mismatch or a non-integrable log-likelihood ratio. abbreviation compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_absolutelyContinuous_of_integrable","label":"relativeEntropy_of_absolutelyContinuous_of_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_of_absolutelyContinuous_of_integrable","description":"The finite regular branch of the Radon--Nikodym representation. The mass correction vanishes when `P` and `Q` are probability measures.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-927db8ea2ef8","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5688,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:174"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_of_absolutelyContinuous_of_integrable {α : Type*} [MeasurableSpace α] (P Q : Measure α) (hPQ : P ≪ Q) (hInt : Integrable (llr P Q) P) : relativeEntropy P Q = ENNReal.ofReal (∫ x, llr P Q x ∂P + Q.real univ - P.real univ)","missing":[],"search":"relativeentropy_of_absolutelycontinuous_of_integrable banditrlproof.lowerbounds.relativeentropy_of_absolutelycontinuous_of_integrable the finite regular branch of the radon--nikodym representation. the mass correction vanishes when `p` and `q` are probability measures. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_probability_absolutelyContinuous_of_integrable","label":"relativeEntropy_of_probability_absolutelyContinuous_of_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_of_probability_absolutelyContinuous_of_integrable","description":"Probability-measure specialization of Theorem 14.1: the finite relative entropy is exactly the expected log likelihood ratio.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-33bebf79a4d8","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5689,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:184"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_of_probability_absolutelyContinuous_of_integrable {α : Type*} [MeasurableSpace α] (P Q : Measure α) [IsProbabilityMeasure P] [IsProbabilityMeasure Q] (hPQ : P ≪ Q) (hInt : Integrable (llr P Q) P) : relativeEntropy P Q = ENNReal.ofReal (∫ x, llr P Q x ∂P)","missing":[],"search":"relativeentropy_of_probability_absolutelycontinuous_of_integrable banditrlproof.lowerbounds.relativeentropy_of_probability_absolutelycontinuous_of_integrable probability-measure specialization of theorem 14.1: the finite relative entropy is exactly the expected log likelihood ratio. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_not_absolutelyContinuous","label":"relativeEntropy_eq_top_of_not_absolutelyContinuous","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_not_absolutelyContinuous","description":"The singular branch of Theorem 14.1.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-af80a4c0d103","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5690,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:194"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_eq_top_of_not_absolutelyContinuous {α : Type*} [MeasurableSpace α] {P Q : Measure α} (hPQ : ¬ P ≪ Q) : relativeEntropy P Q = ∞","missing":[],"search":"relativeentropy_eq_top_of_not_absolutelycontinuous banditrlproof.lowerbounds.relativeentropy_eq_top_of_not_absolutelycontinuous the singular branch of theorem 14.1. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","label":"relativeEntropy_ne_top_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","description":"Exact finiteness contract for the Mathlib representation of Chapter 14 relative entropy.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-2fe4612a71ee","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5691,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:202"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_ne_top_iff {α : Type*} [MeasurableSpace α] {P Q : Measure α} : relativeEntropy P Q ≠ ∞ ↔ P ≪ Q ∧ Integrable (llr P Q) P","missing":[],"search":"relativeentropy_ne_top_iff banditrlproof.lowerbounds.relativeentropy_ne_top_iff exact finiteness contract for the mathlib representation of chapter 14 relative entropy. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_zero_iff","label":"relativeEntropy_eq_zero_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_eq_zero_iff","description":"Relative entropy vanishes exactly when the finite measures agree.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-abc675b64adb","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5692,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:208"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_eq_zero_iff {α : Type*} [MeasurableSpace α] {P Q : Measure α} [IsFiniteMeasure P] [IsFiniteMeasure Q] : relativeEntropy P Q = 0 ↔ P = Q","missing":[],"search":"relativeentropy_eq_zero_iff banditrlproof.lowerbounds.relativeentropy_eq_zero_iff relative entropy vanishes exactly when the finite measures agree. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_le","label":"relativeEntropy_trim_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_trim_le","description":"Exercise 14.10 in its full sub-sigma-algebra form: forgetting measurable sets cannot increase relative entropy. The proof uses the Radon--Nikodym conditional-expectation identity and conditional Jensen for `klFun`.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-705c5b623c15","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5693,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:217"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_trim_le {α : Type*} {m m₀ : MeasurableSpace α} {P Q : @Measure α m₀} [IsFiniteMeasure P] [IsFiniteMeasure Q] (hm : m ≤ m₀) : @relativeEntropy α m (P.trim hm) (Q.trim hm) ≤ @relativeEntropy α m₀ P Q","missing":[],"search":"relativeentropy_trim_le banditrlproof.lowerbounds.relativeentropy_trim_le exercise 14.10 in its full sub-sigma-algebra form: forgetting measurable sets cannot increase relative entropy. the proof uses the radon--nikodym conditional-expectation identity and conditional jensen for `klfun`. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy","label":"bernoulliRelativeEntropy","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bernoulliRelativeEntropy","description":"The Bernoulli relative entropy from Eq. (14.4), reusing the project's exact support and endpoint convention.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-06fde70784d0","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5694,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:293"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"abbrev bernoulliRelativeEntropy (p q : Real) : ENNReal","missing":[],"search":"bernoullirelativeentropy banditrlproof.lowerbounds.bernoullirelativeentropy the bernoulli relative entropy from eq. (14.4), reusing the project's exact support and endpoint convention. abbreviation compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.rnDeriv_restrict_restrict","label":"rnDeriv_restrict_restrict","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.rnDeriv_restrict_restrict","description":"Restricting both laws to a measurable event preserves the original Radon--Nikodym derivative almost everywhere on that event.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-a24c9968dfd1","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5695,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:298"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem rnDeriv_restrict_restrict {α : Type*} [MeasurableSpace α] {P Q : Measure α} [SigmaFinite P] [SigmaFinite Q] (hPQ : P ≪ Q) {A : Set α} (hA : MeasurableSet A) : (P.restrict A).rnDeriv (Q.restrict A) =ᵐ[Q.restrict A] P.rnDeriv Q","missing":[],"search":"rnderiv_restrict_restrict banditrlproof.lowerbounds.rnderiv_restrict_restrict restricting both laws to a measurable event preserves the original radon--nikodym derivative almost everywhere on that event. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_restrict_add_compl","label":"relativeEntropy_restrict_add_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_restrict_add_compl","description":"Relative entropy splits exactly across an event and its complement. This is the two-cell partition identity used by the event data-processing proof.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-62ce216779f3","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5696,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:314"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_restrict_add_compl {α : Type*} [MeasurableSpace α] {P Q : Measure α} [IsFiniteMeasure P] [IsFiniteMeasure Q] (hPQ : P ≪ Q) {A : Set α} (hA : MeasurableSet A) : relativeEntropy P Q = relativeEntropy (P.restrict A) (Q.restrict A) + relativeEntropy (P.restrict Aᶜ) (Q.restrict Aᶜ)","missing":[],"search":"relativeentropy_restrict_add_compl banditrlproof.lowerbounds.relativeentropy_restrict_add_compl relative entropy splits exactly across an event and its complement. this is the two-cell partition identity used by the event data-processing proof. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bernoulliKLCore_event_le","label":"bernoulliKLCore_event_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bernoulliKLCore_event_le","description":"Event data processing in the finite, non-singular Bernoulli branch. This is the quantitative core of Exercise 14.10 specialized to the sigma-algebra generated by one event.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-52c575587539","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5697,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:357"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bernoulliKLCore_event_le {α : Type*} [MeasurableSpace α] {P Q : Measure α} [IsProbabilityMeasure P] [IsProbabilityMeasure Q] {A : Set α} (hA : MeasurableSet A) (hKL : relativeEntropy P Q ≠ ∞) (hQ0 : 0 < Q.real A) (hQ1 : Q.real A < 1) : KLUCB.bernoulliKLCore (P.real A) (Q.real A) ≤ (relativeEntropy P Q).toReal","missing":[],"search":"bernoulliklcore_event_le banditrlproof.lowerbounds.bernoulliklcore_event_le event data processing in the finite, non-singular bernoulli branch. this is the quantitative core of exercise 14.10 specialized to the sigma-algebra generated by one event. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.mul_sqrt_div_eq_sqrt_mul","label":"mul_sqrt_div_eq_sqrt_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.mul_sqrt_div_eq_sqrt_mul","description":"private theorem mul_sqrt_div_eq_sqrt_mul {a b : Real} (ha : 0 < a) (hb : 0 ≤ b) : a * Real.sqrt (b / a) = Real.sqrt (a * b)","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-461464768003","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5698,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:410"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem mul_sqrt_div_eq_sqrt_mul {a b : Real} (ha : 0 < a) (hb : 0 ≤ b) : a * Real.sqrt (b / a) = Real.sqrt (a * b)","missing":[],"search":"mul_sqrt_div_eq_sqrt_mul banditrlproof.lowerbounds.mul_sqrt_div_eq_sqrt_mul private theorem mul_sqrt_div_eq_sqrt_mul {a b : real} (ha : 0 < a) (hb : 0 ≤ b) : a * real.sqrt (b / a) = real.sqrt (a * b) theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exp_neg_half_bernoulliKLCore_le_affinity","label":"exp_neg_half_bernoulliKLCore_le_affinity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exp_neg_half_bernoulliKLCore_le_affinity","description":"The binary likelihood affinity dominates `exp(-d/2)`. This is the two-atom Jensen step in the source proof of Theorem 14.2.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-67a47e9f1b61","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5699,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:428"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_half_bernoulliKLCore_le_affinity {p q : Real} (hp0 : 0 < p) (hp1 : p < 1) (hq0 : 0 < q) (hq1 : q < 1) : Real.exp (-(KLUCB.bernoulliKLCore p q) / 2) ≤ Real.sqrt (p * q) + Real.sqrt ((1 - p) * (1 - q))","missing":[],"search":"exp_neg_half_bernoulliklcore_le_affinity banditrlproof.lowerbounds.exp_neg_half_bernoulliklcore_le_affinity the binary likelihood affinity dominates `exp(-d/2)`. this is the two-atom jensen step in the source proof of theorem 14.2. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.half_binaryAffinity_sq_le_eventError","label":"half_binaryAffinity_sq_le_eventError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.half_binaryAffinity_sq_le_eventError","description":"The two-atom Le Cam overlap inequality in the orientation needed for an event `A`: the affinity squared, divided by two, is bounded by `p + (1 - q)`.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-6d81c05b8be1","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5700,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:496"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem half_binaryAffinity_sq_le_eventError {p q : Real} (hp : KLUCB.IsBernoulliParameter p) (hq : KLUCB.IsBernoulliParameter q) : (1 / 2 : Real) * (Real.sqrt (p * q) + Real.sqrt ((1 - p) * (1 - q))) ^ 2 ≤ p + (1 - q)","missing":[],"search":"half_binaryaffinity_sq_le_eventerror banditrlproof.lowerbounds.half_binaryaffinity_sq_le_eventerror the two-atom le cam overlap inequality in the orientation needed for an event `a`: the affinity squared, divided by two, is bounded by `p + (1 - q)`. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuberCore","label":"binaryBretagnolleHuberCore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryBretagnolleHuberCore","description":"Bretagnolle--Huber for the finite analytic Bernoulli KL expression.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-68a478d6b831","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5701,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:531"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem binaryBretagnolleHuberCore {p q : Real} (hp : KLUCB.IsBernoulliParameter p) (hq0 : 0 < q) (hq1 : q < 1) : (1 / 2 : Real) * Real.exp (-KLUCB.bernoulliKLCore p q) ≤ p + (1 - q)","missing":[],"search":"binarybretagnollehubercore banditrlproof.lowerbounds.binarybretagnollehubercore bretagnolle--huber for the finite analytic bernoulli kl expression. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale","label":"bretagnolleHuberScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale","description":"The source convention `exp(-∞)=0`, exposed as a real-valued testing scale so Theorem 14.2 remains unconditional.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-605c0b611253","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5702,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:576"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"noncomputable def bretagnolleHuberScale (d : ENNReal) : Real","missing":[],"search":"bretagnollehuberscale banditrlproof.lowerbounds.bretagnollehuberscale the source convention `exp(-∞)=0`, exposed as a real-valued testing scale so theorem 14.2 remains unconditional. definition compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_nonneg","label":"bretagnolleHuberScale_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale_nonneg","description":"theorem bretagnolleHuberScale_nonneg (d : ENNReal) : 0 ≤ bretagnolleHuberScale d","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-9c783ef9e84c","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5703,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:579"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuberScale_nonneg (d : ENNReal) : 0 ≤ bretagnolleHuberScale d","missing":[],"search":"bretagnollehuberscale_nonneg banditrlproof.lowerbounds.bretagnollehuberscale_nonneg theorem bretagnollehuberscale_nonneg (d : ennreal) : 0 ≤ bretagnollehuberscale d theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","label":"binaryBretagnolleHuber","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryBretagnolleHuber","description":"Exact two-atom Bretagnolle--Huber inequality, including singular Bernoulli endpoints through the extended-real testing scale.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-0001f3c4602b","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5704,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:588"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem binaryBretagnolleHuber {p q : Real} (hp : KLUCB.IsBernoulliParameter p) (hq : KLUCB.IsBernoulliParameter q) : bretagnolleHuberScale (bernoulliRelativeEntropy p q) ≤ p + (1 - q)","missing":[],"search":"binarybretagnollehuber banditrlproof.lowerbounds.binarybretagnollehuber exact two-atom bretagnolle--huber inequality, including singular bernoulli endpoints through the extended-real testing scale. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","label":"bernoulliRelativeEntropy_event_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","description":"Event-level binary data processing: observing only membership in `A` cannot increase the relative entropy.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-56fb0cd3cb0f","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5705,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:621"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem bernoulliRelativeEntropy_event_le {α : Type*} [MeasurableSpace α] {P Q : Measure α} [IsProbabilityMeasure P] [IsProbabilityMeasure Q] {A : Set α} (hA : MeasurableSet A) : bernoulliRelativeEntropy (P.real A) (Q.real A) ≤ relativeEntropy P Q","missing":[],"search":"bernoullirelativeentropy_event_le banditrlproof.lowerbounds.bernoullirelativeentropy_event_le event-level binary data processing: observing only membership in `a` cannot increase the relative entropy. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","label":"bretagnolleHuberScale_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","description":"The Bretagnolle--Huber testing scale is antitone in its information argument.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-c3963a54a3ae","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5706,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:669"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuberScale_antitone {d D : ENNReal} (h : d ≤ D) : bretagnolleHuberScale D ≤ bretagnolleHuberScale d","missing":[],"search":"bretagnollehuberscale_antitone banditrlproof.lowerbounds.bretagnollehuberscale_antitone the bretagnolle--huber testing scale is antitone in its information argument. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","label":"bretagnolleHuber","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuber","description":"*Bretagnolle--Huber inequality** (Lattimore--Szepesvári, Theorem 14.2). For any measurable event, the two testing errors are bounded below in the source KL direction `D(P,Q)`. The infinite-divergence case is included by `bretagnolleHuberScale`.","url":"../modules/banditrlproof-lowerbounds-informationtheory/index.html#decl-8e5b7ced3e82","parent":"module:BanditRLProof.LowerBounds.InformationTheory","order":5707,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InformationTheory"],["Source","BanditRLProof/LowerBounds/InformationTheory.lean:688"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuber {α : Type*} [MeasurableSpace α] {P Q : Measure α} [IsProbabilityMeasure P] [IsProbabilityMeasure Q] {A : Set α} (hA : MeasurableSet A) : bretagnolleHuberScale (relativeEntropy P Q) ≤ P.real A + Q.real Aᶜ","missing":[],"search":"bretagnollehuber banditrlproof.lowerbounds.bretagnollehuber *bretagnolle--huber inequality** (lattimore--szepesvári, theorem 14.2). for any measurable event, the two testing errors are bounded below in the source kl direction `d(p,q)`. the infinite-divergence case is included by `bretagnollehuberscale`. theorem compiled","shard":"modules/71eb4f20c1620f5f.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret","label":"IsConsistentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentRegret","description":"The exact scalar asymptotic quantifier in Definition 16.1: a nonnegative regret sequence is smaller than every positive polynomial order. Regret nonnegativity is kept outside this analytic predicate so callers must expose the model-specific fact explicitly.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-135b07c6a59d","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5708,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"def IsConsistentRegret (regret : Nat -> Real) : Prop","missing":[],"search":"isconsistentregret banditrlproof.lowerbounds.isconsistentregret the exact scalar asymptotic quantifier in definition 16.1: a nonnegative regret sequence is smaller than every positive polynomial order. regret nonnegativity is kept outside this analytic predicate so callers must expose the model-specific fact explicitly. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentPolicyOver","label":"IsConsistentPolicyOver","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentPolicyOver","description":"Definition 16.1 over an abstract policy/environment regret interface. This preserves the source quantifier order: one policy, every environment in the class, and every positive real exponent.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ee26e984d278","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5709,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"def IsConsistentPolicyOver {Policy Environment : Type*} (environmentClass : Set Environment) (regret : Policy -> Environment -> Nat -> Real) (policy : Policy) : Prop","missing":[],"search":"isconsistentpolicyover banditrlproof.lowerbounds.isconsistentpolicyover definition 16.1 over an abstract policy/environment regret interface. this preserves the source quantifier order: one policy, every environment in the class, and every positive real exponent. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment","label":"FiniteMeanBanditEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment","description":"A finite-armed product environment whose arm laws have certified finite means. The `mean` field is intentionally tied to the Bochner integral rather than treated as an unrelated parameter: this is the source-level bridge used when an unchanged arm law must imply an unchanged arm mean.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-3eff2544aeaa","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5710,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:56"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"structure FiniteMeanBanditEnvironment (K : Nat) where","missing":[],"search":"finitemeanbanditenvironment banditrlproof.lowerbounds.finitemeanbanditenvironment a finite-armed product environment whose arm laws have certified finite means. the `mean` field is intentionally tied to the bochner integral rather than treated as an unrelated parameter: this is the source-level bridge used when an unchanged arm law must imply an unchanged arm mean. structure compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap","label":"gap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap","description":"The source gap `Delta_i(nu) = muStar(nu) - mu_i(nu)`.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-77ab9457e03d","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5711,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:71"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def FiniteMeanBanditEnvironment.gap {K : Nat} (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) : Real","missing":[],"search":"gap banditrlproof.lowerbounds.finitemeanbanditenvironment.gap the source gap `delta_i(nu) = mustar(nu) - mu_i(nu)`. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap_nonneg","label":"gap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap_nonneg","description":"theorem FiniteMeanBanditEnvironment.gap_nonneg {K : Nat} (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) : 0 <= environment.gap arm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-8e03e20fa09a","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5712,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.gap_nonneg {K : Nat} (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) : 0 <= environment.gap arm","missing":[],"search":"gap_nonneg banditrlproof.lowerbounds.finitemeanbanditenvironment.gap_nonneg theorem finitemeanbanditenvironment.gap_nonneg {k : nat} (environment : finitemeanbanditenvironment k) (arm : fin k) : 0 <= environment.gap arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap_bestArm","label":"gap_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap_bestArm","description":"theorem FiniteMeanBanditEnvironment.gap_bestArm {K : Nat} (environment : FiniteMeanBanditEnvironment K) : environment.gap environment.bestArm = 0","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-14b50b8db723","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5713,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:83"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.gap_bestArm {K : Nat} (environment : FiniteMeanBanditEnvironment K) : environment.gap environment.bestArm = 0","missing":[],"search":"gap_bestarm banditrlproof.lowerbounds.finitemeanbanditenvironment.gap_bestarm theorem finitemeanbanditenvironment.gap_bestarm {k : nat} (environment : finitemeanbanditenvironment k) : environment.gap environment.bestarm = 0 theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMeanIncrease","label":"oneArmMeanIncrease","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMeanIncrease","description":"The mean increase `lambda` of the single changed arm in Lemma 16.3.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ff71a201ce52","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5714,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def oneArmMeanIncrease {K : Nat} (original reference : FiniteMeanBanditEnvironment K) (changedArm : Fin K) : Real","missing":[],"search":"onearmmeanincrease banditrlproof.lowerbounds.onearmmeanincrease the mean increase `lambda` of the single changed arm in lemma 16.3. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmChangedMargin","label":"oneArmChangedMargin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmChangedMargin","description":"The changed environment's advantage over the original optimal mean, `lambda - Delta_i(nu)`, in the exact second branch of Lemma 16.3's minimum.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-036d8c0e9983","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5715,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:96"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def oneArmChangedMargin {K : Nat} (original reference : FiniteMeanBanditEnvironment K) (changedArm : Fin K) : Real","missing":[],"search":"onearmchangedmargin banditrlproof.lowerbounds.onearmchangedmargin the changed environment's advantage over the original optimal mean, `lambda - delta_i(nu)`, in the exact second branch of lemma 16.3's minimum. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.mean_eq_of_armLaw_eq","label":"mean_eq_of_armLaw_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.mean_eq_of_armLaw_eq","description":"Equal arm laws have equal certified finite means.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-a64a13f78485","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5716,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:102"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.mean_eq_of_armLaw_eq {K : Nat} (first second : FiniteMeanBanditEnvironment K) (arm : Fin K) (hlaw : first.armLaw arm = second.armLaw arm) : first.mean arm = second.mean arm","missing":[],"search":"mean_eq_of_armlaw_eq banditrlproof.lowerbounds.finitemeanbanditenvironment.mean_eq_of_armlaw_eq equal arm laws have equal certified finite means. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.InUnstructuredClass","label":"InUnstructuredClass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.InUnstructuredClass","description":"The unstructured product class in Theorem 16.2: each arm law belongs to its specified component class. Finite means are certified by the environment.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-8fec358357f4","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5717,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:110"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def FiniteMeanBanditEnvironment.InUnstructuredClass {K : Nat} (environment : FiniteMeanBanditEnvironment K) (componentClass : Fin K → Set (Measure Real)) : Prop","missing":[],"search":"inunstructuredclass banditrlproof.lowerbounds.finitemeanbanditenvironment.inunstructuredclass the unstructured product class in theorem 16.2: each arm law belongs to its specified component class. finite means are certified by the environment. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm","label":"withImprovedArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm","description":"Replace one arm by a finite-mean probability law whose mean exceeds the old optimum. This constructs the source alternative environment explicitly.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-fc7ff1af52e6","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5718,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def FiniteMeanBanditEnvironment.withImprovedArm {K : Nat} (environment : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (alternative : Measure Real) [IsProbabilityMeasure alternative] (hintegrable : Integrable id alternative) (hbetter : environment.mean environment.bestArm < ∫ x, x ∂alternative) : FiniteMeanBanditEnvironment K where","missing":[],"search":"withimprovedarm banditrlproof.lowerbounds.finitemeanbanditenvironment.withimprovedarm replace one arm by a finite-mean probability law whose mean exceeds the old optimum. this constructs the source alternative environment explicitly. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_law","label":"withImprovedArm_law","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_law","description":"theorem FiniteMeanBanditEnvironment.withImprovedArm_law {K : Nat} (environment : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (alternative : Measure Real) [IsProbabilityMeasure alternative] (hi : Integrable id alternative) (hb : environment.mean environment.bestArm < ∫ x, x ∂alternative) (arm : Fin K) : (environment.withImprovedArm changedArm alternative hi hb).armLaw arm = if arm = changedArm then alternativ…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-de8dcc198bfd","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5719,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:154"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.withImprovedArm_law {K : Nat} (environment : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (alternative : Measure Real) [IsProbabilityMeasure alternative] (hi : Integrable id alternative) (hb : environment.mean environment.bestArm < ∫ x, x ∂alternative) (arm : Fin K) : (environment.withImprovedArm changedArm alternative hi hb).armLaw arm = if arm = changedArm then alternative else environment.armLaw arm","missing":[],"search":"withimprovedarm_law banditrlproof.lowerbounds.finitemeanbanditenvironment.withimprovedarm_law theorem finitemeanbanditenvironment.withimprovedarm_law {k : nat} (environment : finitemeanbanditenvironment k) (changedarm : fin k) (alternative : measure real) [isprobabilitymeasure alternative] (hi : integrable id alternative) (hb : environment.mean environment.bestarm < ∫ x, x ∂alternative) (arm : fin k) : (environment.withimprovedarm changedarm alternative hi hb).armlaw arm = if arm = changedarm then alternative else environment.armlaw arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_unique","label":"withImprovedArm_unique","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_unique","description":"The replacement arm is uniquely optimal, as required by Lemma 16.3.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-f97173fbf283","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5720,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:164"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.withImprovedArm_unique {K : Nat} (environment : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (alternative : Measure Real) [IsProbabilityMeasure alternative] (hi : Integrable id alternative) (hb : environment.mean environment.bestArm < ∫ x, x ∂alternative) (arm : Fin K) (hne : arm ≠ changedArm) : (environment.withImprovedArm changedArm alternative hi hb).mean arm < (environment.withImprovedArm changedArm alternative hi hb).mean changedArm","missing":[],"search":"withimprovedarm_unique banditrlproof.lowerbounds.finitemeanbanditenvironment.withimprovedarm_unique the replacement arm is uniquely optimal, as required by lemma 16.3. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_mem","label":"withImprovedArm_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_mem","description":"Unstructured classes permit exactly the single-component replacement used by the source change-of-measure argument.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-afe8102bb518","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5721,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:177"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.withImprovedArm_mem {K : Nat} (environment : FiniteMeanBanditEnvironment K) (componentClass : Fin K → Set (Measure Real)) (hclass : environment.InUnstructuredClass componentClass) (changedArm : Fin K) (alternative : Measure Real) [IsProbabilityMeasure alternative] (hi : Integrable id alternative) (hb : environment.mean environment.bestArm < ∫ x, x ∂alternative) (halt : alternative ∈ componentClass changedArm) : (environment.withImprovedArm changedArm alternative hi hb).InUnstructuredClass componentClass","missing":[],"search":"withimprovedarm_mem banditrlproof.lowerbounds.finitemeanbanditenvironment.withimprovedarm_mem unstructured classes permit exactly the single-component replacement used by the source change-of-measure argument. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMeanIncrease_sub_gap_eq_changedMargin","label":"oneArmMeanIncrease_sub_gap_eq_changedMargin","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMeanIncrease_sub_gap_eq_changedMargin","description":"The source identity behind Lemma 16.3: `lambda - Delta_i(nu) = mu_i(nu') - muStar(nu)`.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-4c402f65496c","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5722,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:195"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oneArmMeanIncrease_sub_gap_eq_changedMargin {K : Nat} (original reference : FiniteMeanBanditEnvironment K) (changedArm : Fin K) : oneArmMeanIncrease original reference changedArm - original.gap changedArm = oneArmChangedMargin original reference changedArm","missing":[],"search":"onearmmeanincrease_sub_gap_eq_changedmargin banditrlproof.lowerbounds.onearmmeanincrease_sub_gap_eq_changedmargin the source identity behind lemma 16.3: `lambda - delta_i(nu) = mu_i(nu') - mustar(nu)`. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMeanChange_produces_gap_contract","label":"oneArmMeanChange_produces_gap_contract","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMeanChange_produces_gap_contract","description":"Source mean-to-gap producer for a one-arm change. If the changed arm is suboptimal originally, uniquely optimal after the change, and every other arm law is unchanged, then it produces all sign and comparison obligations needed by the exact majority-event consumer for Lemma 16.3.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-3cb21c9f9058","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5723,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:209"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem oneArmMeanChange_produces_gap_contract {K : Nat} (original reference : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (hsuboptimal : original.mean changedArm < original.mean original.bestArm) (hunique : forall arm, arm ≠ changedArm -> reference.mean arm < reference.mean changedArm) (hsame : forall arm, arm ≠ changedArm -> original.armLaw arm = reference.armLaw arm) : 0 < original.gap changedArm /\\ 0 < oneArmChangedMargin original reference changedArm /\\ (forall arm, arm ≠ changedArm -> oneArmChangedMargin original reference changedArm <= reference.gap arm)","missing":[],"search":"onearmmeanchange_produces_gap_contract banditrlproof.lowerbounds.onearmmeanchange_produces_gap_contract source mean-to-gap producer for a one-arm change. if the changed arm is suboptimal originally, uniquely optimal after the change, and every other arm law is unchanged, then it produces all sign and comparison obligations needed by the exact majority-event consumer for lemma 16.3. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment","label":"UnitVarianceGaussianBanditEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment","description":"An unrestricted finite vector of unit-variance Gaussian means, together with a certified optimal arm. Unlike the Chapter 15 minimax cube, Chapter 16 uses all mean vectors in `Real^k`.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-f9f5e26ada50","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5724,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:253"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"structure UnitVarianceGaussianBanditEnvironment (K : Nat) where","missing":[],"search":"unitvariancegaussianbanditenvironment banditrlproof.lowerbounds.unitvariancegaussianbanditenvironment an unrestricted finite vector of unit-variance gaussian means, together with a certified optimal arm. unlike the chapter 15 minimax cube, chapter 16 uses all mean vectors in `real^k`. structure compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.gap","label":"gap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.gap","description":"def UnitVarianceGaussianBanditEnvironment.gap {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (arm : Fin K) : Real","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-342045be5ead","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5725,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:258"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def UnitVarianceGaussianBanditEnvironment.gap {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (arm : Fin K) : Real","missing":[],"search":"gap banditrlproof.lowerbounds.unitvariancegaussianbanditenvironment.gap def unitvariancegaussianbanditenvironment.gap {k : nat} (environment : unitvariancegaussianbanditenvironment k) (arm : fin k) : real definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.gap_nonneg","label":"gap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.gap_nonneg","description":"theorem UnitVarianceGaussianBanditEnvironment.gap_nonneg {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (arm : Fin K) : 0 <= environment.gap arm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-6f9b830dc76e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5726,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:263"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem UnitVarianceGaussianBanditEnvironment.gap_nonneg {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (arm : Fin K) : 0 <= environment.gap arm","missing":[],"search":"gap_nonneg banditrlproof.lowerbounds.unitvariancegaussianbanditenvironment.gap_nonneg theorem unitvariancegaussianbanditenvironment.gap_nonneg {k : nat} (environment : unitvariancegaussianbanditenvironment k) (arm : fin k) : 0 <= environment.gap arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean","label":"toFiniteMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean","description":"The finite-mean product environment induced by arbitrary unit-variance Gaussian arms.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-8c52a084296a","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5727,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:271"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def UnitVarianceGaussianBanditEnvironment.toFiniteMean {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : FiniteMeanBanditEnvironment K where","missing":[],"search":"tofinitemean banditrlproof.lowerbounds.unitvariancegaussianbanditenvironment.tofinitemean the finite-mean product environment induced by arbitrary unit-variance gaussian arms. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean_mean","label":"toFiniteMean_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean_mean","description":"theorem UnitVarianceGaussianBanditEnvironment.toFiniteMean_mean {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : environment.toFiniteMean.mean = environment.mean","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ad90064c2eb8","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5728,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:289"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem UnitVarianceGaussianBanditEnvironment.toFiniteMean_mean {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : environment.toFiniteMean.mean = environment.mean","missing":[],"search":"tofinitemean_mean banditrlproof.lowerbounds.unitvariancegaussianbanditenvironment.tofinitemean_mean theorem unitvariancegaussianbanditenvironment.tofinitemean_mean {k : nat} (environment : unitvariancegaussianbanditenvironment k) : environment.tofinitemean.mean = environment.mean theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean_gap","label":"toFiniteMean_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean_gap","description":"theorem UnitVarianceGaussianBanditEnvironment.toFiniteMean_gap {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : environment.toFiniteMean.gap = environment.gap","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-56afb20d8209","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5729,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:294"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem UnitVarianceGaussianBanditEnvironment.toFiniteMean_gap {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : environment.toFiniteMean.gap = environment.gap","missing":[],"search":"tofinitemean_gap banditrlproof.lowerbounds.unitvariancegaussianbanditenvironment.tofinitemean_gap theorem unitvariancegaussianbanditenvironment.tofinitemean_gap {k : nat} (environment : unitvariancegaussianbanditenvironment k) : environment.tofinitemean.gap = environment.gap theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean","label":"chapter16GaussianChangedMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedMean","description":"Mean vector obtained by increasing exactly one Gaussian arm by `(1 + epsilon) Delta_i`, as in the proof of Theorem 16.4.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ac4e9cbaeb9e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5730,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:300"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def chapter16GaussianChangedMean {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (arm : Fin K) : Real","missing":[],"search":"chapter16gaussianchangedmean banditrlproof.lowerbounds.chapter16gaussianchangedmean mean vector obtained by increasing exactly one gaussian arm by `(1 + epsilon) delta_i`, as in the proof of theorem 16.4. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_changed","label":"chapter16GaussianChangedMean_changed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedMean_changed","description":"theorem chapter16GaussianChangedMean_changed {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) : chapter16GaussianChangedMean environment changedArm epsilon changedArm = environment.mean changedArm + (1 + epsilon) * environment.gap changedArm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-61b6c275b819","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5731,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:309"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedMean_changed {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) : chapter16GaussianChangedMean environment changedArm epsilon changedArm = environment.mean changedArm + (1 + epsilon) * environment.gap changedArm","missing":[],"search":"chapter16gaussianchangedmean_changed banditrlproof.lowerbounds.chapter16gaussianchangedmean_changed theorem chapter16gaussianchangedmean_changed {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) : chapter16gaussianchangedmean environment changedarm epsilon changedarm = environment.mean changedarm + (1 + epsilon) * environment.gap changedarm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_other","label":"chapter16GaussianChangedMean_other","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedMean_other","description":"theorem chapter16GaussianChangedMean_other {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (arm : Fin K) (harm : arm ≠ changedArm) : chapter16GaussianChangedMean environment changedArm epsilon arm = environment.mean arm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-63993b4cc10f","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5732,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:317"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedMean_other {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (arm : Fin K) (harm : arm ≠ changedArm) : chapter16GaussianChangedMean environment changedArm epsilon arm = environment.mean arm","missing":[],"search":"chapter16gaussianchangedmean_other banditrlproof.lowerbounds.chapter16gaussianchangedmean_other theorem chapter16gaussianchangedmean_other {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (arm : fin k) (harm : arm ≠ changedarm) : chapter16gaussianchangedmean environment changedarm epsilon arm = environment.mean arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_uniqueBest","label":"chapter16GaussianChangedMean_uniqueBest","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedMean_uniqueBest","description":"The shifted Gaussian arm is uniquely optimal whenever its original gap and `epsilon` are positive.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-0c0db9b16998","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5733,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:327"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedMean_uniqueBest {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedArm -> chapter16GaussianChangedMean environment changedArm epsilon arm < chapter16GaussianChangedMean environment changedArm epsilon changedArm","missing":[],"search":"chapter16gaussianchangedmean_uniquebest banditrlproof.lowerbounds.chapter16gaussianchangedmean_uniquebest the shifted gaussian arm is uniquely optimal whenever its original gap and `epsilon` are positive. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment","label":"chapter16GaussianChangedEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment","description":"The shifted mean vector, certified with the changed arm as an optimum.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-a6a7571071a2","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5734,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:343"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def chapter16GaussianChangedEnvironment {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : UnitVarianceGaussianBanditEnvironment K where","missing":[],"search":"chapter16gaussianchangedenvironment banditrlproof.lowerbounds.chapter16gaussianchangedenvironment the shifted mean vector, certified with the changed arm as an optimum. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_mem_localBox","label":"chapter16GaussianChangedMean_mem_localBox","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedMean_mem_localBox","description":"The Gaussian shift stays in the source box `[mu_j, mu_j + 2 Delta_j]` when `epsilon <= 1`.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-6d97fbf3578e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5735,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:360"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedMean_mem_localBox {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) (hepsilon_one : epsilon <= 1) : forall arm, chapter16GaussianChangedMean environment changedArm epsilon arm ∈ Set.Icc (environment.mean arm) (environment.mean arm + 2 * environment.gap arm)","missing":[],"search":"chapter16gaussianchangedmean_mem_localbox banditrlproof.lowerbounds.chapter16gaussianchangedmean_mem_localbox the gaussian shift stays in the source box `[mu_j, mu_j + 2 delta_j]` when `epsilon <= 1`. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_uniqueBest","label":"chapter16GaussianChangedEnvironment_uniqueBest","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_uniqueBest","description":"theorem chapter16GaussianChangedEnvironment_uniqueBest {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedArm -> (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).mean arm < (chapter16GaussianChangedEnvironment environment changedArm epsilon…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-3e8433502f62","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5736,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:379"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedEnvironment_uniqueBest {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedArm -> (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).mean arm < (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).mean changedArm","missing":[],"search":"chapter16gaussianchangedenvironment_uniquebest banditrlproof.lowerbounds.chapter16gaussianchangedenvironment_uniquebest theorem chapter16gaussianchangedenvironment_uniquebest {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedarm -> (chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).mean arm < (chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).mean changedarm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_sameArmLaw","label":"chapter16GaussianChangedEnvironment_sameArmLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_sameArmLaw","description":"theorem chapter16GaussianChangedEnvironment_sameArmLaw {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedArm -> environment.toFiniteMean.armLaw arm = (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean.armLaw arm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-454e8e5fb447","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5737,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:392"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedEnvironment_sameArmLaw {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedArm -> environment.toFiniteMean.armLaw arm = (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean.armLaw arm","missing":[],"search":"chapter16gaussianchangedenvironment_samearmlaw banditrlproof.lowerbounds.chapter16gaussianchangedenvironment_samearmlaw theorem chapter16gaussianchangedenvironment_samearmlaw {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) : forall arm, arm ≠ changedarm -> environment.tofinitemean.armlaw arm = (chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).tofinitemean.armlaw arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_meanIncrease","label":"chapter16GaussianChangedEnvironment_meanIncrease","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_meanIncrease","description":"theorem chapter16GaussianChangedEnvironment_meanIncrease {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : oneArmMeanIncrease environment.toFiniteMean (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean changedArm = (1 + epsilon) * environment.gap change…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-6a0f75d0c438","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5738,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:409"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedEnvironment_meanIncrease {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : oneArmMeanIncrease environment.toFiniteMean (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean changedArm = (1 + epsilon) * environment.gap changedArm","missing":[],"search":"chapter16gaussianchangedenvironment_meanincrease banditrlproof.lowerbounds.chapter16gaussianchangedenvironment_meanincrease theorem chapter16gaussianchangedenvironment_meanincrease {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) : onearmmeanincrease environment.tofinitemean (chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).tofinitemean changedarm = (1 + epsilon) * environment.gap changedarm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_changedMargin","label":"chapter16GaussianChangedEnvironment_changedMargin","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_changedMargin","description":"theorem chapter16GaussianChangedEnvironment_changedMargin {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : oneArmChangedMargin environment.toFiniteMean (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean changedArm = epsilon * environment.gap changedArm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-da084a9c3f55","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5739,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:421"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedEnvironment_changedMargin {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : oneArmChangedMargin environment.toFiniteMean (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean changedArm = epsilon * environment.gap changedArm","missing":[],"search":"chapter16gaussianchangedenvironment_changedmargin banditrlproof.lowerbounds.chapter16gaussianchangedenvironment_changedmargin theorem chapter16gaussianchangedenvironment_changedmargin {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) : onearmchangedmargin environment.tofinitemean (chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).tofinitemean changedarm = epsilon * environment.gap changedarm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL","label":"chapter16GaussianChangedEnvironment_armKL","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL","description":"theorem chapter16GaussianChangedEnvironment_armKL {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : InformationTheory.klDiv (environment.toFiniteMean.armLaw changedArm) ((chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean.armLaw changedArm) = ENNReal.ofR…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-a92958d91eed","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5740,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:436"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedEnvironment_armKL {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : InformationTheory.klDiv (environment.toFiniteMean.armLaw changedArm) ((chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean.armLaw changedArm) = ENNReal.ofReal (((1 + epsilon) * environment.gap changedArm) ^ 2 / 2)","missing":[],"search":"chapter16gaussianchangedenvironment_armkl banditrlproof.lowerbounds.chapter16gaussianchangedenvironment_armkl theorem chapter16gaussianchangedenvironment_armkl {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) : informationtheory.kldiv (environment.tofinitemean.armlaw changedarm) ((chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).tofinitemean.armlaw changedarm) = ennreal.ofreal (((1 + epsilon) * environment.gap changedarm) ^ 2 / 2) theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL_toReal","label":"chapter16GaussianChangedEnvironment_armKL_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL_toReal","description":"theorem chapter16GaussianChangedEnvironment_armKL_toReal {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : (InformationTheory.klDiv (environment.toFiniteMean.armLaw changedArm) ((chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean.armLaw changedArm)).toRe…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-62d079debd73","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5741,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:453"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16GaussianChangedEnvironment_armKL_toReal {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) : (InformationTheory.klDiv (environment.toFiniteMean.armLaw changedArm) ((chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon).toFiniteMean.armLaw changedArm)).toReal = ((1 + epsilon) * environment.gap changedArm) ^ 2 / 2","missing":[],"search":"chapter16gaussianchangedenvironment_armkl_toreal banditrlproof.lowerbounds.chapter16gaussianchangedenvironment_armkl_toreal theorem chapter16gaussianchangedenvironment_armkl_toreal {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) : (informationtheory.kldiv (environment.tofinitemean.armlaw changedarm) ((chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon).tofinitemean.armlaw changedarm)).toreal = ((1 + epsilon) * environment.gap changedarm) ^ 2 / 2 theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.InChapter16GaussianLocalClass","label":"InChapter16GaussianLocalClass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.InChapter16GaussianLocalClass","description":"Membership in the local Gaussian class `E(nu)` of Theorem 16.4.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-753482915e6b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5742,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:468"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def InChapter16GaussianLocalClass {K : Nat} (base candidate : UnitVarianceGaussianBanditEnvironment K) : Prop","missing":[],"search":"inchapter16gaussianlocalclass banditrlproof.lowerbounds.inchapter16gaussianlocalclass membership in the local gaussian class `e(nu)` of theorem 16.4. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.inChapter16GaussianLocalClass_self","label":"inChapter16GaussianLocalClass_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.inChapter16GaussianLocalClass_self","description":"theorem inChapter16GaussianLocalClass_self {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : InChapter16GaussianLocalClass environment environment","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-cf85e61320d5","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5743,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:473"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inChapter16GaussianLocalClass_self {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) : InChapter16GaussianLocalClass environment environment","missing":[],"search":"inchapter16gaussianlocalclass_self banditrlproof.lowerbounds.inchapter16gaussianlocalclass_self theorem inchapter16gaussianlocalclass_self {k : nat} (environment : unitvariancegaussianbanditenvironment k) : inchapter16gaussianlocalclass environment environment theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.inChapter16GaussianLocalClass_changed","label":"inChapter16GaussianLocalClass_changed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.inChapter16GaussianLocalClass_changed","description":"theorem inChapter16GaussianLocalClass_changed {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) (hepsilon_one : epsilon <= 1) : InChapter16GaussianLocalClass environment (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon)","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-132e0658185e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5744,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:481"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem inChapter16GaussianLocalClass_changed {K : Nat} (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) (hepsilon_one : epsilon <= 1) : InChapter16GaussianLocalClass environment (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon)","missing":[],"search":"inchapter16gaussianlocalclass_changed banditrlproof.lowerbounds.inchapter16gaussianlocalclass_changed theorem inchapter16gaussianlocalclass_changed {k : nat} (environment : unitvariancegaussianbanditenvironment k) (changedarm : fin k) (epsilon : real) (hgap : 0 < environment.gap changedarm) (hepsilon : 0 < epsilon) (hepsilon_one : epsilon <= 1) : inchapter16gaussianlocalclass environment (chapter16gaussianchangedenvironment environment changedarm epsilon hgap hepsilon) theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.add","label":"add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentRegret.add","description":"The sum of two source-consistent regret sequences is still consistent. This is the closure step used for `R_n(nu) + R_n(nu')` in the proof of Theorem 16.2.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-87b6bd67bb73","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5745,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:495"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem IsConsistentRegret.add {first second : Nat -> Real} (hfirst : IsConsistentRegret first) (hsecond : IsConsistentRegret second) : IsConsistentRegret (fun n => first n + second n)","missing":[],"search":"add banditrlproof.lowerbounds.isconsistentregret.add the sum of two source-consistent regret sequences is still consistent. this is the closure step used for `r_n(nu) + r_n(nu')` in the proof of theorem 16.2. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_add_le_rpow","label":"eventually_add_le_rpow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentRegret.eventually_add_le_rpow","description":"Two consistent regret sequences are eventually at most `n^p` for every positive exponent `p`. This is a stronger eventual form of the source's auxiliary `C_p n^p` bound and avoids introducing an opaque constant.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-49ab3ef5d04e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5746,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:506"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem IsConsistentRegret.eventually_add_le_rpow {first second : Nat -> Real} (hfirst : IsConsistentRegret first) (hsecond : IsConsistentRegret second) {p : Real} (hp : 0 < p) : ∀ᶠ n : Nat in atTop, first n + second n <= (n : Real) ^ p","missing":[],"search":"eventually_add_le_rpow banditrlproof.lowerbounds.isconsistentregret.eventually_add_le_rpow two consistent regret sequences are eventually at most `n^p` for every positive exponent `p`. this is a stronger eventual form of the source's auxiliary `c_p n^p` bound and avoids introducing an opaque constant. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","label":"eventually_log_add_div_log_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","description":"Eventually, the logarithmic growth ratio of a positive sum of two consistent regrets is at most every positive exponent. This is the direction-correct analytic leaf used before the source takes a limsup.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-cbc3039e957d","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5747,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:526"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem IsConsistentRegret.eventually_log_add_div_log_le {first second : Nat -> Real} (hfirst : IsConsistentRegret first) (hsecond : IsConsistentRegret second) (hpositive : ∀ᶠ n : Nat in atTop, 0 < first n + second n) {p : Real} (hp : 0 < p) : ∀ᶠ n : Nat in atTop, Real.log (first n + second n) / Real.log n <= p","missing":[],"search":"eventually_log_add_div_log_le banditrlproof.lowerbounds.isconsistentregret.eventually_log_add_div_log_le eventually, the logarithmic growth ratio of a positive sum of two consistent regrets is at most every positive exponent. this is the direction-correct analytic leaf used before the source takes a limsup. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_pull_div_log_ge","label":"eventually_pull_div_log_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentRegret.eventually_pull_div_log_ge","description":"Consistency extracts every strict reciprocal-information lower bound from the source logarithmic inequality. This eventual form does not assume that the normalized pull counts converge.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-b39dea25f193","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5748,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:547"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsConsistentRegret.eventually_pull_div_log_ge {first second pulls : Nat → Real} {c d r : Real} (hfirst : IsConsistentRegret first) (hsecond : IsConsistentRegret second) (hpositive : ∀ᶠ n in atTop, 0 < first n + second n) (hd : 0 < d) (hr : r < 1 / d) (hsource : ∀ᶠ n : Nat in atTop, (c + Real.log n - Real.log (first n + second n)) / d ≤ pulls n) : ∀ᶠ n in atTop, r ≤ pulls n / Real.log n","missing":[],"search":"eventually_pull_div_log_ge banditrlproof.lowerbounds.isconsistentregret.eventually_pull_div_log_ge consistency extracts every strict reciprocal-information lower bound from the source logarithmic inequality. this eventual form does not assume that the normalized pull counts converge. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.liminf_pull_div_log_ge","label":"liminf_pull_div_log_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsConsistentRegret.liminf_pull_div_log_ge","description":"The finite-positive information branch in extended-real liminf form. The conclusion allows infinite normalized pull growth.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-f32e88f44643","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5749,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:586"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsConsistentRegret.liminf_pull_div_log_ge {first second pulls : Nat → Real} {c d : Real} (hfirst : IsConsistentRegret first) (hsecond : IsConsistentRegret second) (hpositive : ∀ᶠ n in atTop, 0 < first n + second n) (hd : 0 < d) (hsource : ∀ᶠ n : Nat in atTop, (c + Real.log n - Real.log (first n + second n)) / d ≤ pulls n) : ENNReal.ofReal (1 / d) ≤ liminf (fun n : Nat => ENNReal.ofReal (pulls n / Real.log n)) atTop","missing":[],"search":"liminf_pull_div_log_ge banditrlproof.lowerbounds.isconsistentregret.liminf_pull_div_log_ge the finite-positive information branch in extended-real liminf form. the conclusion allows infinite normalized pull growth. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.divergenceInfimum","label":"divergenceInfimum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.divergenceInfimum","description":"The source quantity `d_inf(P, muStar, M) = inf {D(P,P') : P' in M, mean(P') > muStar}`. The value is extended-real, so an empty alternative set has infimum `∞` and support failures remain visible.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-82fc8658ce21","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5750,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:608"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"def divergenceInfimum {Reward : Type*} [MeasurableSpace Reward] (P : Measure Reward) (muStar : Real) (distributionClass : Set (Measure Reward)) (mean : Measure Reward -> Real) : ENNReal","missing":[],"search":"divergenceinfimum banditrlproof.lowerbounds.divergenceinfimum the source quantity `d_inf(p, mustar, m) = inf {d(p,p') : p' in m, mean(p') > mustar}`. the value is extended-real, so an empty alternative set has infimum `∞` and support failures remain visible. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_le","label":"divergenceInfimum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.divergenceInfimum_le","description":"Any admissible confusing alternative upper-bounds `d_inf`, with KL in the source direction from the original law to the alternative law.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-76ffee0d67d1","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5751,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:620"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem divergenceInfimum_le {Reward : Type*} [MeasurableSpace Reward] {P P' : Measure Reward} {muStar : Real} {distributionClass : Set (Measure Reward)} {mean : Measure Reward -> Real} (hclass : P' ∈ distributionClass) (hbetter : muStar < mean P') : divergenceInfimum P muStar distributionClass mean <= relativeEntropy P P'","missing":[],"search":"divergenceinfimum_le banditrlproof.lowerbounds.divergenceinfimum_le any admissible confusing alternative upper-bounds `d_inf`, with kl in the source direction from the original law to the alternative law. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_exists_alternative_lt","label":"divergenceInfimum_exists_alternative_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.divergenceInfimum_exists_alternative_lt","description":"Every strict upper bound on `d_inf` admits a confusing alternative below that bound. No minimizer or finite positive infimum is assumed.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-2d6e8df5adf1","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5752,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:634"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem divergenceInfimum_exists_alternative_lt {Reward : Type*} [MeasurableSpace Reward] {P : Measure Reward} {muStar : Real} {distributionClass : Set (Measure Reward)} {mean : Measure Reward -> Real} {bound : ENNReal} (hbound : divergenceInfimum P muStar distributionClass mean < bound) : ∃ P', P' ∈ distributionClass ∧ muStar < mean P' ∧ relativeEntropy P P' < bound","missing":[],"search":"divergenceinfimum_exists_alternative_lt banditrlproof.lowerbounds.divergenceinfimum_exists_alternative_lt every strict upper bound on `d_inf` admits a confusing alternative below that bound. no minimizer or finite positive infimum is assumed. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.exists_confusingEnvironment_lt","label":"exists_confusingEnvironment_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.exists_confusingEnvironment_lt","description":"Lift a near-infimum arm law to an admissible product environment. The component classes consist of finite-mean probability laws, exactly as in Theorem 16.2; the information cost stays in `ENNReal`.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-90c45ccbd250","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5753,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:653"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem FiniteMeanBanditEnvironment.exists_confusingEnvironment_lt {K : Nat} (environment : FiniteMeanBanditEnvironment K) (componentClass : Fin K → Set (Measure Real)) (hclass : environment.InUnstructuredClass componentClass) (hfinite : ∀ arm P, P ∈ componentClass arm → IsProbabilityMeasure P ∧ Integrable id P) (changedArm : Fin K) {bound : ENNReal} (hbound : divergenceInfimum (environment.armLaw changedArm) (environment.mean environment.bestArm) (componentClass changedArm) (fun P => ∫ x, x ∂P) < bound) : ∃ reference : FiniteMeanBanditEnvironment K, reference.InUnstructuredClass componentClass ∧ (∀ arm, arm ≠ changedArm → environment.armLaw arm = reference.armLaw arm) ∧ (∀ arm, arm ≠ changedArm → reference.mean arm < reference.mean changedArm) ∧ relativeEntropy (environment.armLaw changedArm) (reference.armLaw changedArm) < bound","missing":[],"search":"exists_confusingenvironment_lt banditrlproof.lowerbounds.finitemeanbanditenvironment.exists_confusingenvironment_lt lift a near-infimum arm law to an admissible product environment. the component classes consist of finite-mean probability laws, exactly as in theorem 16.2; the information cost stays in `ennreal`. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_eq_top_iff","label":"divergenceInfimum_eq_top_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.divergenceInfimum_eq_top_iff","description":"The infinite branch includes both an empty alternative class and classes whose every confusing alternative has infinite directed KL.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-c0d5daabe553","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5754,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:683"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem divergenceInfimum_eq_top_iff {Reward : Type*} [MeasurableSpace Reward] {P : Measure Reward} {muStar : Real} {distributionClass : Set (Measure Reward)} {mean : Measure Reward -> Real} : divergenceInfimum P muStar distributionClass mean = ⊤ ↔ ∀ P', P' ∈ distributionClass → muStar < mean P' → relativeEntropy P P' = ⊤","missing":[],"search":"divergenceinfimum_eq_top_iff banditrlproof.lowerbounds.divergenceinfimum_eq_top_iff the infinite branch includes both an empty alternative class and classes whose every confusing alternative has infinite directed kl. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum","label":"parametricDivergenceInfimum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.parametricDivergenceInfimum","description":"Parameterized form of `d_inf`, useful when a distribution class is presented by a family of laws rather than an injective set-level mean map.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-c3220e9f5fad","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5755,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:704"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"def parametricDivergenceInfimum {Reward Parameter : Type*} [MeasurableSpace Reward] (law : Parameter -> Measure Reward) (mean : Parameter -> Real) (parameter : Parameter) (muStar : Real) : ENNReal","missing":[],"search":"parametricdivergenceinfimum banditrlproof.lowerbounds.parametricdivergenceinfimum parameterized form of `d_inf`, useful when a distribution class is presented by a family of laws rather than an injective set-level mean map. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum_le","label":"parametricDivergenceInfimum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.parametricDivergenceInfimum_le","description":"Candidate inequality for the parameterized `d_inf` surface.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-992480c18d2e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5756,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:714"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem parametricDivergenceInfimum_le {Reward Parameter : Type*} [MeasurableSpace Reward] {law : Parameter -> Measure Reward} {mean : Parameter -> Real} {parameter alternative : Parameter} {muStar : Real} (hbetter : muStar < mean alternative) : parametricDivergenceInfimum law mean parameter muStar <= relativeEntropy (law parameter) (law alternative)","missing":[],"search":"parametricdivergenceinfimum_le banditrlproof.lowerbounds.parametricdivergenceinfimum_le candidate inequality for the parameterized `d_inf` surface. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum","label":"unitGaussianDivergenceInfimum","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum","description":"Unit-variance Gaussian specialization of the source `d_inf`.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-221b64b14e33","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5757,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:725"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"abbrev unitGaussianDivergenceInfimum (mu muStar : Real) : ENNReal","missing":[],"search":"unitgaussiandivergenceinfimum banditrlproof.lowerbounds.unitgaussiandivergenceinfimum unit-variance gaussian specialization of the source `d_inf`. abbreviation compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_le_perturbed","label":"unitGaussianDivergenceInfimum_le_perturbed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_le_perturbed","description":"A Gaussian alternative with mean `muStar + epsilon` is admissible and has the exact arm-level information cost shown here. Taking `epsilon -> 0` is a separate infimum/limit leaf and is not hidden in this theorem.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ce7925b1b6e9","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5758,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:731"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianDivergenceInfimum_le_perturbed (mu muStar epsilon : Real) (hepsilon : 0 < epsilon) : unitGaussianDivergenceInfimum mu muStar <= ENNReal.ofReal (((muStar - mu) + epsilon) ^ 2 / 2)","missing":[],"search":"unitgaussiandivergenceinfimum_le_perturbed banditrlproof.lowerbounds.unitgaussiandivergenceinfimum_le_perturbed a gaussian alternative with mean `mustar + epsilon` is admissible and has the exact arm-level information cost shown here. taking `epsilon -> 0` is a separate infimum/limit leaf and is not hidden in this theorem. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_ge","label":"unitGaussianDivergenceInfimum_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_ge","description":"Every strictly better unit-Gaussian alternative costs at least the boundary value from Table 16.1.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-956e65922747","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5759,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:749"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianDivergenceInfimum_ge (mu muStar : Real) (hmu : mu < muStar) : ENNReal.ofReal ((muStar - mu) ^ 2 / 2) <= unitGaussianDivergenceInfimum mu muStar","missing":[],"search":"unitgaussiandivergenceinfimum_ge banditrlproof.lowerbounds.unitgaussiandivergenceinfimum_ge every strictly better unit-gaussian alternative costs at least the boundary value from table 16.1. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","label":"unitGaussianDivergenceInfimum_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","description":"Exact unit-variance Gaussian row of Table 16.1. The strict alternative mean means the boundary law is not itself admissible; the reverse inequality is obtained from positive perturbations tending to zero.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-18acde29d1db","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5760,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:773"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianDivergenceInfimum_eq (mu muStar : Real) (hmu : mu < muStar) : unitGaussianDivergenceInfimum mu muStar = ENNReal.ofReal ((muStar - mu) ^ 2 / 2)","missing":[],"search":"unitgaussiandivergenceinfimum_eq banditrlproof.lowerbounds.unitgaussiandivergenceinfimum_eq exact unit-variance gaussian row of table 16.1. the strict alternative mean means the boundary law is not itself admissible; the reverse inequality is obtained from positive perturbations tending to zero. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","label":"banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","description":"Chapter 16's one-arm change-of-measure specialization of Lemma 15.1. When two stationary bandit environments differ only at `changedArm`, the directed relative entropy of their finite histories is exactly the first-environment expected number of pulls of that arm times its arm-law relative entropy. The algorithm is one common, possibly randomized, nonanticipating history policy in both environments.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-b2459945a469","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5761,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:804"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (changedArm : Fin K) (lastRound : Nat) (hsame : forall arm, arm ≠ changedArm -> armLaw arm = referenceArmLaw arm) : InformationTheory.klDiv (canonicalBanditHistoryMeasure algorithm armLaw lastRound) (canonicalBanditHistoryMeasure algorithm referenceArmLaw lastRound) = canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound changedArm * InformationTheory.klDiv (armLaw changedArm) (referenceArmLaw changedArm)","missing":[],"search":"bandithistoryrelativeentropy_eq_expectedpulls_mul_of_only_arm_changed banditrlproof.lowerbounds.bandithistoryrelativeentropy_eq_expectedpulls_mul_of_only_arm_changed chapter 16's one-arm change-of-measure specialization of lemma 15.1. when two stationary bandit environments differ only at `changedarm`, the directed relative entropy of their finite histories is exactly the first-environment expected number of pulls of that arm times its arm-law relative entropy. the algorithm is one common, possibly randomized, nonanticipating history policy in both environments. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMajorityPullEvent","label":"oneArmMajorityPullEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMajorityPullEvent","description":"The source majority event `A = {T_i(n) > n/2}` in the repository's inclusive convention, where `lastRound` contains `lastRound + 1` pulls.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-d3d74200e186","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5762,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:829"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"def oneArmMajorityPullEvent {K : Nat} {Reward : Type*} (changedArm : Fin K) (lastRound : Nat) : Set (History.FinitePairHistory (Fin K) Reward lastRound)","missing":[],"search":"onearmmajoritypullevent banditrlproof.lowerbounds.onearmmajoritypullevent the source majority event `a = {t_i(n) > n/2}` in the repository's inclusive convention, where `lastround` contains `lastround + 1` pulls. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurableSet_oneArmMajorityPullEvent","label":"measurableSet_oneArmMajorityPullEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurableSet_oneArmMajorityPullEvent","description":"theorem measurableSet_oneArmMajorityPullEvent {K : Nat} {Reward : Type*} [MeasurableSpace Reward] (changedArm : Fin K) (lastRound : Nat) : MeasurableSet (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound)","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-e0724693944a","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5763,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:837"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_oneArmMajorityPullEvent {K : Nat} {Reward : Type*} [MeasurableSpace Reward] (changedArm : Fin K) (lastRound : Nat) : MeasurableSet (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound)","missing":[],"search":"measurableset_onearmmajoritypullevent banditrlproof.lowerbounds.measurableset_onearmmajoritypullevent theorem measurableset_onearmmajoritypullevent {k : nat} {reward : type*} [measurablespace reward] (changedarm : fin k) (lastround : nat) : measurableset (onearmmajoritypullevent (reward := reward) changedarm lastround) theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret","label":"finiteHistoryGapPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret","description":"Gap-times-pull-count pseudo-regret on one realized finite history. The gap vector is kept explicit so the later Chapter 16 environment layer must identify it with `muStar - mu_i`; no scalar regret hypothesis is hidden here.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-35031cdafb04","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5764,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:849"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryGapPseudoRegret {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) : ENNReal","missing":[],"search":"finitehistorygappseudoregret banditrlproof.lowerbounds.finitehistorygappseudoregret gap-times-pull-count pseudo-regret on one realized finite history. the gap vector is kept explicit so the later chapter 16 environment layer must identify it with `mustar - mu_i`; no scalar regret hypothesis is hidden here. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret","label":"canonicalGapExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret","description":"Expected gap pseudo-regret under the canonical law generated by one possibly randomized history policy and one stationary arm kernel.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-6ceee186a579","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5765,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:859"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalGapExpectedPseudoRegret {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (lastRound : Nat) : ENNReal","missing":[],"search":"canonicalgapexpectedpseudoregret banditrlproof.lowerbounds.canonicalgapexpectedpseudoregret expected gap pseudo-regret under the canonical law generated by one possibly randomized history policy and one stationary arm kernel. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryGapPseudoRegret","label":"measurable_finiteHistoryGapPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_finiteHistoryGapPseudoRegret","description":"theorem measurable_finiteHistoryGapPseudoRegret {K : Nat} {Reward : Type*} [MeasurableSpace Reward] (gap : Fin K -> Real) (lastRound : Nat) : Measurable (finiteHistoryGapPseudoRegret (Reward := Reward) gap lastRound)","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-5d76d137ea14","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5766,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:869"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryGapPseudoRegret {K : Nat} {Reward : Type*} [MeasurableSpace Reward] (gap : Fin K -> Real) (lastRound : Nat) : Measurable (finiteHistoryGapPseudoRegret (Reward := Reward) gap lastRound)","missing":[],"search":"measurable_finitehistorygappseudoregret banditrlproof.lowerbounds.measurable_finitehistorygappseudoregret theorem measurable_finitehistorygappseudoregret {k : nat} {reward : type*} [measurablespace reward] (gap : fin k -> real) (lastround : nat) : measurable (finitehistorygappseudoregret (reward := reward) gap lastround) theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_ne_top","label":"finiteHistoryGapPseudoRegret_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_ne_top","description":"theorem finiteHistoryGapPseudoRegret_ne_top {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) : finiteHistoryGapPseudoRegret gap lastRound history ≠ ∞","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-03820caebb69","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5767,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:880"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryGapPseudoRegret_ne_top {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) : finiteHistoryGapPseudoRegret gap lastRound history ≠ ∞","missing":[],"search":"finitehistorygappseudoregret_ne_top banditrlproof.lowerbounds.finitehistorygappseudoregret_ne_top theorem finitehistorygappseudoregret_ne_top {k : nat} {reward : type*} (gap : fin k -> real) (lastround : nat) (history : history.finitepairhistory (fin k) reward lastround) : finitehistorygappseudoregret gap lastround history ≠ ∞ theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_toReal","label":"finiteHistoryGapPseudoRegret_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_toReal","description":"theorem finiteHistoryGapPseudoRegret_toReal {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) (hgap : forall arm, 0 <= gap arm) : (finiteHistoryGapPseudoRegret gap lastRound history).toReal = ∑ arm : Fin K, gap arm * finiteHistoryPullCountReal lastRound history arm","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-67d98b763d01","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5768,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:891"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryGapPseudoRegret_toReal {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) (hgap : forall arm, 0 <= gap arm) : (finiteHistoryGapPseudoRegret gap lastRound history).toReal = ∑ arm : Fin K, gap arm * finiteHistoryPullCountReal lastRound history arm","missing":[],"search":"finitehistorygappseudoregret_toreal banditrlproof.lowerbounds.finitehistorygappseudoregret_toreal theorem finitehistorygappseudoregret_toreal {k : nat} {reward : type*} (gap : fin k -> real) (lastround : nat) (history : history.finitepairhistory (fin k) reward lastround) (hgap : forall arm, 0 <= gap arm) : (finitehistorygappseudoregret gap lastround history).toreal = ∑ arm : fin k, gap arm * finitehistorypullcountreal lastround history arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough_general","label":"sum_canonicalRealizedExpectedPullCountThrough_general","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough_general","description":"The expected realized pull counts of all arms sum to the inclusive horizon for every Markov arm kernel, not only for the Gaussian Chapter 15 instance.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-3ed2b203b210","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5769,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:912"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_canonicalRealizedExpectedPullCountThrough_general {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (lastRound : Nat) : ∑ arm : Fin K, canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm = lastRound + 1","missing":[],"search":"sum_canonicalrealizedexpectedpullcountthrough_general banditrlproof.lowerbounds.sum_canonicalrealizedexpectedpullcountthrough_general the expected realized pull counts of all arms sum to the inclusive horizon for every markov arm kernel, not only for the gaussian chapter 15 instance. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_ne_top","label":"canonicalRealizedExpectedPullCountThrough_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_ne_top","description":"theorem canonicalRealizedExpectedPullCountThrough_ne_top {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (lastRound : Nat) (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm ≠ ∞","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-c9cbe54edf7e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5770,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:929"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalRealizedExpectedPullCountThrough_ne_top {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (lastRound : Nat) (arm : Fin K) : canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm ≠ ∞","missing":[],"search":"canonicalrealizedexpectedpullcountthrough_ne_top banditrlproof.lowerbounds.canonicalrealizedexpectedpullcountthrough_ne_top theorem canonicalrealizedexpectedpullcountthrough_ne_top {k : nat} {reward : type*} [measurablespace reward] [measurablespace.countablygenerated reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (lastround : nat) (arm : fin k) : canonicalrealizedexpectedpullcountthrough algorithm armlaw lastround arm ≠ ∞ theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls","label":"canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls","description":"theorem canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (lastRound : Nat) : canonicalGapExpectedPseudoRegret algorithm armLaw gap lastRound = ∑ arm : Fin K, ENNReal.ofReal (gap arm) *…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-7693bcaa101b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5771,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:946"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (lastRound : Nat) : canonicalGapExpectedPseudoRegret algorithm armLaw gap lastRound = ∑ arm : Fin K, ENNReal.ofReal (gap arm) * canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm","missing":[],"search":"canonicalgapexpectedpseudoregret_eq_sum_expectedpulls banditrlproof.lowerbounds.canonicalgapexpectedpseudoregret_eq_sum_expectedpulls theorem canonicalgapexpectedpseudoregret_eq_sum_expectedpulls {k : nat} {reward : type*} [measurablespace reward] [measurablespace.countablygenerated reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (gap : fin k -> real) (lastround : nat) : canonicalgapexpectedpseudoregret algorithm armlaw gap lastround = ∑ arm : fin k, ennreal.ofreal (gap arm) * canonicalrealizedexpectedpullcountthrough algorithm armlaw lastround arm theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_ne_top","label":"canonicalGapExpectedPseudoRegret_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_ne_top","description":"theorem canonicalGapExpectedPseudoRegret_ne_top {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (lastRound : Nat) : canonicalGapExpectedPseudoRegret algorithm armLaw gap lastRound ≠ ∞","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-55939c800084","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5772,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:969"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalGapExpectedPseudoRegret_ne_top {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (lastRound : Nat) : canonicalGapExpectedPseudoRegret algorithm armLaw gap lastRound ≠ ∞","missing":[],"search":"canonicalgapexpectedpseudoregret_ne_top banditrlproof.lowerbounds.canonicalgapexpectedpseudoregret_ne_top theorem canonicalgapexpectedpseudoregret_ne_top {k : nat} {reward : type*} [measurablespace reward] [measurablespace.countablygenerated reward] (algorithm : thompson.historyalgorithm (fin k) reward) (armlaw : kernel (fin k) reward) [ismarkovkernel armlaw] (gap : fin k -> real) (lastround : nat) : canonicalgapexpectedpseudoregret algorithm armlaw gap lastround ≠ ∞ theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal","label":"canonicalGapExpectedPseudoRegretReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal","description":"Real-valued presentation of the finite expected pseudo-regret. Finiteness is proved above rather than assumed by the Chapter 16 consumer.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ae60d736f100","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5773,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:984"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalGapExpectedPseudoRegretReal {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (lastRound : Nat) : Real","missing":[],"search":"canonicalgapexpectedpseudoregretreal banditrlproof.lowerbounds.canonicalgapexpectedpseudoregretreal real-valued presentation of the finite expected pseudo-regret. finiteness is proved above rather than assumed by the chapter 16 consumer. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal_eq_sum_expectedPulls","label":"canonicalGapExpectedPseudoRegretReal_eq_sum_expectedPulls","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal_eq_sum_expectedPulls","description":"Real-valued gap-times-expected-pulls decomposition. This is the precise finite-horizon form used when Theorem 16.4 sums its per-arm lower bounds.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-520ea307f312","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5774,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:994"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem canonicalGapExpectedPseudoRegretReal_eq_sum_expectedPulls {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (hgap : forall arm, 0 <= gap arm) (lastRound : Nat) : canonicalGapExpectedPseudoRegretReal algorithm armLaw gap lastRound = ∑ arm : Fin K, gap arm * (canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound arm).toReal","missing":[],"search":"canonicalgapexpectedpseudoregretreal_eq_sum_expectedpulls banditrlproof.lowerbounds.canonicalgapexpectedpseudoregretreal_eq_sum_expectedpulls real-valued gap-times-expected-pulls decomposition. this is the precise finite-horizon form used when theorem 16.4 sums its per-arm lower bounds. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitVarianceGaussianExpectedPseudoRegret","label":"unitVarianceGaussianExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitVarianceGaussianExpectedPseudoRegret","description":"Expected pseudo-regret of an unrestricted unit-variance Gaussian environment in the Chapter 16 gap-times-pulls convention.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-66b31b21949e","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5775,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1019"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def unitVarianceGaussianExpectedPseudoRegret {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitVarianceGaussianBanditEnvironment K) (lastRound : Nat) : Real","missing":[],"search":"unitvariancegaussianexpectedpseudoregret banditrlproof.lowerbounds.unitvariancegaussianexpectedpseudoregret expected pseudo-regret of an unrestricted unit-variance gaussian environment in the chapter 16 gap-times-pulls convention. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMajority_forces_gapPseudoRegret","label":"oneArmMajority_forces_gapPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMajority_forces_gapPseudoRegret","description":"On the source majority event, the original environment pays at least half the horizon times the changed arm's positive gap.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-008525bcb85b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5776,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1028"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oneArmMajority_forces_gapPseudoRegret {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (hgap : forall arm, 0 <= gap arm) (changedArm : Fin K) (hchanged : 0 < gap changedArm) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) (hA : history ∈ oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound) : ENNReal.ofReal (((lastRound + 1 : Nat) : Real) * gap changedArm / 2) <= finiteHistoryGapPseudoRegret gap lastRound history","missing":[],"search":"onearmmajority_forces_gappseudoregret banditrlproof.lowerbounds.onearmmajority_forces_gappseudoregret on the source majority event, the original environment pays at least half the horizon times the changed arm's positive gap. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_forces_gapPseudoRegret","label":"oneArmMajority_compl_forces_gapPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMajority_compl_forces_gapPseudoRegret","description":"On the complement of the majority event, every non-changed arm charged by at least `changedMargin` forces half-horizon pseudo-regret.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-80c6cacb7fef","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5777,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1061"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem oneArmMajority_compl_forces_gapPseudoRegret {K : Nat} {Reward : Type*} (gap : Fin K -> Real) (hgap : forall arm, 0 <= gap arm) (changedArm : Fin K) (changedMargin : Real) (hmargin : 0 < changedMargin) (hother : forall arm, arm ≠ changedArm -> changedMargin <= gap arm) (lastRound : Nat) (history : History.FinitePairHistory (Fin K) Reward lastRound) (hAc : history ∈ (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound)ᶜ) : ENNReal.ofReal (((lastRound + 1 : Nat) : Real) * changedMargin / 2) <= finiteHistoryGapPseudoRegret gap lastRound history","missing":[],"search":"onearmmajority_compl_forces_gappseudoregret banditrlproof.lowerbounds.onearmmajority_compl_forces_gappseudoregret on the complement of the majority event, every non-changed arm charged by at least `changedmargin` forces half-horizon pseudo-regret. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMajority_probability_charge_le_expectedPseudoRegret","label":"oneArmMajority_probability_charge_le_expectedPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMajority_probability_charge_le_expectedPseudoRegret","description":"Original-environment event probability charged to the actual canonical gap pseudo-regret.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-8248bb64380d","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5778,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem oneArmMajority_probability_charge_le_expectedPseudoRegret {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (hgap : forall arm, 0 <= gap arm) (changedArm : Fin K) (hchanged : 0 < gap changedArm) (lastRound : Nat) : ((lastRound + 1 : Nat) : Real) * gap changedArm / 2 * (canonicalBanditHistoryMeasure algorithm armLaw lastRound).real (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound) <= canonicalGapExpectedPseudoRegretReal algorithm armLaw gap lastRound","missing":[],"search":"onearmmajority_probability_charge_le_expectedpseudoregret banditrlproof.lowerbounds.onearmmajority_probability_charge_le_expectedpseudoregret original-environment event probability charged to the actual canonical gap pseudo-regret. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","label":"oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","description":"Changed-environment complement probability charged to the actual canonical gap pseudo-regret.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-c37dc2f3ea20","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5779,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1162"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem oneArmMajority_compl_probability_charge_le_expectedPseudoRegret {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] (gap : Fin K -> Real) (hgap : forall arm, 0 <= gap arm) (changedArm : Fin K) (changedMargin : Real) (hmargin : 0 < changedMargin) (hother : forall arm, arm ≠ changedArm -> changedMargin <= gap arm) (lastRound : Nat) : ((lastRound + 1 : Nat) : Real) * changedMargin / 2 * (canonicalBanditHistoryMeasure algorithm armLaw lastRound).real (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound)ᶜ <= canonicalGapExpectedPseudoRegretReal algorithm armLaw gap lastRound","missing":[],"search":"onearmmajority_compl_probability_charge_le_expectedpseudoregret banditrlproof.lowerbounds.onearmmajority_compl_probability_charge_le_expectedpseudoregret changed-environment complement probability charged to the actual canonical gap pseudo-regret. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","label":"bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","description":"Chapter 16's measurable-event information constraint before regret calibration. It instantiates Bretagnolle--Huber on the exact majority event and rewrites the full history KL using the compiled one-arm specialization of Lemma 15.1.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-793485755554","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5780,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1201"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (changedArm : Fin K) (lastRound : Nat) (hsame : forall arm, arm ≠ changedArm -> armLaw arm = referenceArmLaw arm) : bretagnolleHuberScale (canonicalRealizedExpectedPullCountThrough algorithm armLaw lastRound changedArm * InformationTheory.klDiv (armLaw changedArm) (referenceArmLaw changedArm)) <= (canonicalBanditHistoryMeasure algorithm armLaw lastRound).real (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound) + (canonicalBanditHistoryMeasure algorithm referenceArmLaw lastRound).real (oneArmMajorityPullEvent (Reward := Reward) changedArm lastRound)ᶜ","missing":[],"search":"bretagnollehuberscale_expectedpulls_mul_armkl_le_majorityerrors banditrlproof.lowerbounds.bretagnollehuberscale_expectedpulls_mul_armkl_le_majorityerrors chapter 16's measurable-event information constraint before regret calibration. it instantiates bretagnolle--huber on the exact majority event and rewrites the full history kl using the compiled one-arm specialization of lemma 15.1. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_mul_eq_exp","label":"bretagnolleHuberScale_mul_eq_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bretagnolleHuberScale_mul_eq_exp","description":"Finite-information evaluation of the testing scale used when Chapter 16 passes from extended-real KL to the real logarithmic inequality.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-d2dad5210b99","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5781,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1237"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem bretagnolleHuberScale_mul_eq_exp {expectedPull armInformation : ENNReal} (hpull : expectedPull ≠ ∞) (hinformation : armInformation ≠ ∞) : bretagnolleHuberScale (expectedPull * armInformation) = (1 / 2 : Real) * Real.exp (-(expectedPull.toReal * armInformation.toReal))","missing":[],"search":"bretagnollehuberscale_mul_eq_exp banditrlproof.lowerbounds.bretagnollehuberscale_mul_eq_exp finite-information evaluation of the testing scale used when chapter 16 passes from extended-real kl to the real logarithmic inequality. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exp_testing_bound_of_majority_regret_bounds","label":"exp_testing_bound_of_majority_regret_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exp_testing_bound_of_majority_regret_bounds","description":"Deterministic assembly of the two majority-event regret charges with the Bretagnolle--Huber testing error.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-a3d9c9f12fc8","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5782,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1249"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exp_testing_bound_of_majority_regret_bounds (expectedPull information gap changedMargin horizon originalError changedError originalRegret changedRegret : Real) (hgap : 0 < gap) (hmargin : 0 < changedMargin) (hhorizon : 0 < horizon) (horiginalError : 0 <= originalError) (hchangedError : 0 <= changedError) (htesting : (1 / 2 : Real) * Real.exp (-(expectedPull * information)) <= originalError + changedError) (horiginalRegret : horizon * gap / 2 * originalError <= originalRegret) (hchangedRegret : horizon * changedMargin / 2 * changedError <= changedRegret) : horizon * min gap changedMargin / 4 * Real.exp (-(expectedPull * information)) <= originalRegret + changedRegret","missing":[],"search":"exp_testing_bound_of_majority_regret_bounds banditrlproof.lowerbounds.exp_testing_bound_of_majority_regret_bounds deterministic assembly of the two majority-event regret charges with the bretagnolle--huber testing error. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_of_exp_testing_bound","label":"expectedPullCount_ge_log_regret_of_exp_testing_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_of_exp_testing_bound","description":"The exact logarithmic rearrangement used in Lemma 16.3. This theorem is only the scalar consumer of the testing inequality; the bandit theorem must still produce `htesting` from the common-policy history law, the majority event, and the two expected pseudo-regrets.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-dd1f673b9bd7","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5783,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1287"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem expectedPullCount_ge_log_regret_of_exp_testing_bound (expectedPull information gap changedMargin horizon regretSum : Real) (hinformation : 0 < information) (hgap : 0 < gap) (hmargin : 0 < changedMargin) (hhorizon : 0 < horizon) (htesting : horizon * min gap changedMargin / 4 * Real.exp (-(expectedPull * information)) <= regretSum) : (Real.log (min gap changedMargin / 4) + Real.log horizon - Real.log regretSum) / information <= expectedPull","missing":[],"search":"expectedpullcount_ge_log_regret_of_exp_testing_bound banditrlproof.lowerbounds.expectedpullcount_ge_log_regret_of_exp_testing_bound the exact logarithmic rearrangement used in lemma 16.3. this theorem is only the scalar consumer of the testing inequality; the bandit theorem must still produce `htesting` from the common-policy history law, the majority event, and the two expected pseudo-regrets. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","label":"expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","description":"Chapter 16's finite-time change-of-measure calculation with the actual canonical gap pseudo-regrets produced above. This closes the common-policy, one-arm-KL, majority-event, exact `1/4`, and logarithmic assembly route for explicit nonnegative gap vectors. It is intentionally not named as source Lemma 16.3: a later environment layer must still prove that these vectors are the mean gaps of finite-mean arm laws and di…","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-2e1714674420","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5784,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1324"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","stochastic-finite"]],"statement":"theorem expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (originalGap referenceGap : Fin K -> Real) (horiginalGap : forall arm, 0 <= originalGap arm) (hreferenceGap : forall arm, 0 <= referenceGap arm) (changedArm : Fin K) (changedMargin : Real) (hchangedGap : 0 < originalGap changedArm) (hmargin : 0 < changedMargin) (hother : forall arm, arm ≠ changedArm -> changedMargin <= referenceGap arm) (lastRound : Nat) (hsame : forall arm, arm ≠ changedArm -> armLaw arm = referenceArmLaw arm) (hinformation_ne_top : InformationTheory.klDiv (armLaw changedArm) (referenceArmLaw changedArm) ≠ ∞) (hinformation_pos : 0 < (InformationTheo…","missing":[],"search":"expectedpullcount_ge_log_gappseudoregret_of_only_arm_changed banditrlproof.lowerbounds.expectedpullcount_ge_log_gappseudoregret_of_only_arm_changed chapter 16's finite-time change-of-measure calculation with the actual canonical gap pseudo-regrets produced above. this closes the common-policy, one-arm-kl, majority-event, exact `1/4`, and logarithmic assembly route for explicit nonnegative gap vectors. it is intentionally not named as source lemma 16.3: a later environment layer must still prove that these vectors are the mean gaps of finite-mean arm laws and discharge the finite-positive-kl branch from that source contract. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":["stochastic-finite"]},{"id":"declaration:BanditRLProof.LowerBounds.gapPseudoRegret_add_pos_of_only_arm_changed","label":"gapPseudoRegret_add_pos_of_only_arm_changed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gapPseudoRegret_add_pos_of_only_arm_changed","description":"The two regrets in the one-arm testing construction cannot both vanish when arm KL is finite and both source gaps are positive.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-a2a9b110757b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5785,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1413"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gapPseudoRegret_add_pos_of_only_arm_changed {K : Nat} {Reward : Type*} [MeasurableSpace Reward] [MeasurableSpace.CountablyGenerated Reward] (algorithm : Thompson.HistoryAlgorithm (Fin K) Reward) (armLaw referenceArmLaw : Kernel (Fin K) Reward) [IsMarkovKernel armLaw] [IsMarkovKernel referenceArmLaw] (originalGap referenceGap : Fin K -> Real) (horiginalGap : forall arm, 0 <= originalGap arm) (hreferenceGap : forall arm, 0 <= referenceGap arm) (changedArm : Fin K) (changedMargin : Real) (hchangedGap : 0 < originalGap changedArm) (hmargin : 0 < changedMargin) (hother : forall arm, arm ≠ changedArm -> changedMargin <= referenceGap arm) (lastRound : Nat) (hsame : forall arm, arm ≠ changedArm -> armLaw arm = referenceArmLaw arm) (hinformation_ne_top : InformationTheory.klDiv (armLaw changedArm) (referenceArmLaw changedArm) ≠ ∞) : 0 < canonicalGapExpectedPseudoRegretReal algorithm armL…","missing":[],"search":"gappseudoregret_add_pos_of_only_arm_changed banditrlproof.lowerbounds.gappseudoregret_add_pos_of_only_arm_changed the two regrets in the one-arm testing construction cannot both vanish when arm kl is finite and both source gaps are positive. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","label":"expectedPullCount_ge_log_regret_changeOfMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","description":"*Lattimore--Szepesvári, Lemma 16.3.** This is the exact source one-arm change-of-measure inequality for finite-mean product environments. The denominator is presented through `ENNReal.toReal`; when arm KL is infinite this evaluates the source convention `x / ∞ = 0`. The zero-KL case is impossible here because the changed arm has different certified finite means.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-2330742f98a2","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5786,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1497"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem expectedPullCount_ge_log_regret_changeOfMeasure {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (original reference : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (hsuboptimal : original.mean changedArm < original.mean original.bestArm) (hunique : forall arm, arm ≠ changedArm -> reference.mean arm < reference.mean changedArm) (hsame : forall arm, arm ≠ changedArm -> original.armLaw arm = reference.armLaw arm) (lastRound : Nat) : (Real.log (min (oneArmMeanIncrease original reference changedArm - original.gap changedArm) (original.gap changedArm) / 4) + Real.log ((lastRound + 1 : Nat) : Real) - Real.log (canonicalGapExpectedPseudoRegretReal algorithm original.armLaw original.gap lastRound + canonicalGapExpectedPseudoRegretReal algorithm reference.armLaw reference.gap lastRound)) / (InformationTheory.klDiv (original.armLaw changedArm) (reference.armLaw changed…","missing":[],"search":"expectedpullcount_ge_log_regret_changeofmeasure banditrlproof.lowerbounds.expectedpullcount_ge_log_regret_changeofmeasure *lattimore--szepesvári, lemma 16.3.** this is the exact source one-arm change-of-measure inequality for finite-mean product environments. the denominator is presented through `ennreal.toreal`; when arm kl is infinite this evaluates the source convention `x / ∞ = 0`. the zero-kl case is impossible here because the changed arm has different certified finite means. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedRegret","label":"finiteMeanExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteMeanExpectedRegret","description":"Expected regret after exactly `n` pulls, including the empty horizon.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-47c8e0c8eedd","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5787,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1554"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def finiteMeanExpectedRegret {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) : Nat → Real | 0 => 0 | n + 1 => canonicalGapExpectedPseudoRegretReal algorithm environment.armLaw environment.gap n /-- Expected pull count after exactly `n` pulls, including the empty horizon. -/ def finiteMeanExpectedPullCount {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) : Nat → Real | 0 => 0 | n + 1 => (canonicalRealizedExpectedPullCountThrough algorithm environment.armLaw n arm).toReal theorem finiteMeanExpectedPullCount_nonneg {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) (n : Nat) : 0 ≤ finiteMeanExpectedPullCount algorithm environment arm n","missing":[],"search":"finitemeanexpectedregret banditrlproof.lowerbounds.finitemeanexpectedregret expected regret after exactly `n` pulls, including the empty horizon. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedPullCount","label":"finiteMeanExpectedPullCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteMeanExpectedPullCount","description":"Expected pull count after exactly `n` pulls, including the empty horizon.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-9dfda554b2b4","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5788,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1562"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def finiteMeanExpectedPullCount {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) : Nat → Real | 0 => 0 | n + 1 => (canonicalRealizedExpectedPullCountThrough algorithm environment.armLaw n arm).toReal theorem finiteMeanExpectedPullCount_nonneg {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) (n : Nat) : 0 ≤ finiteMeanExpectedPullCount algorithm environment arm n","missing":[],"search":"finitemeanexpectedpullcount banditrlproof.lowerbounds.finitemeanexpectedpullcount expected pull count after exactly `n` pulls, including the empty horizon. definition compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedPullCount_nonneg","label":"finiteMeanExpectedPullCount_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteMeanExpectedPullCount_nonneg","description":"theorem finiteMeanExpectedPullCount_nonneg {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) (n : Nat) : 0 ≤ finiteMeanExpectedPullCount algorithm environment arm n","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-f7b82317293a","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5789,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1569"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteMeanExpectedPullCount_nonneg {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (arm : Fin K) (n : Nat) : 0 ≤ finiteMeanExpectedPullCount algorithm environment arm n","missing":[],"search":"finitemeanexpectedpullcount_nonneg banditrlproof.lowerbounds.finitemeanexpectedpullcount_nonneg theorem finitemeanexpectedpullcount_nonneg {k : nat} (algorithm : thompson.historyalgorithm (fin k) real) (environment : finitemeanbanditenvironment k) (arm : fin k) (n : nat) : 0 ≤ finitemeanexpectedpullcount algorithm environment arm n theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedRegret_eq_sum","label":"finiteMeanExpectedRegret_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteMeanExpectedRegret_eq_sum","description":"Source regret decomposition with the exact number-of-pulls horizon.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-d43abedd30bf","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5790,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1576"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteMeanExpectedRegret_eq_sum {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (n : Nat) : finiteMeanExpectedRegret algorithm environment n = ∑ arm, environment.gap arm * finiteMeanExpectedPullCount algorithm environment arm n","missing":[],"search":"finitemeanexpectedregret_eq_sum banditrlproof.lowerbounds.finitemeanexpectedregret_eq_sum source regret decomposition with the exact number-of-pulls horizon. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.finiteMeanNormalizedRegret_eq_sum","label":"finiteMeanNormalizedRegret_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.finiteMeanNormalizedRegret_eq_sum","description":"Extended-real normalization of the source regret decomposition.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-790e1242e007","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5791,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1588"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finiteMeanNormalizedRegret_eq_sum {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : FiniteMeanBanditEnvironment K) (n : Nat) (hn : 1 < n) : ENNReal.ofReal (finiteMeanExpectedRegret algorithm environment n / Real.log n) = ∑ arm, ENNReal.ofReal (environment.gap arm) * ENNReal.ofReal (finiteMeanExpectedPullCount algorithm environment arm n / Real.log n)","missing":[],"search":"finitemeannormalizedregret_eq_sum banditrlproof.lowerbounds.finitemeannormalizedregret_eq_sum extended-real normalization of the source regret decomposition. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.consistentRegret_liminf_expectedPull_div_log_ge_of_alternative","label":"consistentRegret_liminf_expectedPull_div_log_ge_of_alternative","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.consistentRegret_liminf_expectedPull_div_log_ge_of_alternative","description":"Theorem 16.2's per-alternative information constraint for finite KL. Consistency is imposed on the actual regret sequences of the two environments, and the conclusion uses the original-law pull count.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-ffe9cfe9c23b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5792,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1607"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem consistentRegret_liminf_expectedPull_div_log_ge_of_alternative {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (original reference : FiniteMeanBanditEnvironment K) (changedArm : Fin K) (hsuboptimal : original.mean changedArm < original.mean original.bestArm) (hunique : ∀ arm, arm ≠ changedArm → reference.mean arm < reference.mean changedArm) (hsame : ∀ arm, arm ≠ changedArm → original.armLaw arm = reference.armLaw arm) (hfirst : IsConsistentRegret (finiteMeanExpectedRegret algorithm original)) (hsecond : IsConsistentRegret (finiteMeanExpectedRegret algorithm reference)) (hfinite : InformationTheory.klDiv (original.armLaw changedArm) (reference.armLaw changedArm) ≠ ∞) : ENNReal.ofReal (1 / (InformationTheory.klDiv (original.armLaw changedArm) (reference.armLaw changedArm)).toReal) ≤ liminf (fun n : Nat => ENNReal.ofReal (finiteMeanExpectedPullCount algorithm origin…","missing":[],"search":"consistentregret_liminf_expectedpull_div_log_ge_of_alternative banditrlproof.lowerbounds.consistentregret_liminf_expectedpull_div_log_ge_of_alternative theorem 16.2's per-alternative information constraint for finite kl. consistency is imposed on the actual regret sequences of the two environments, and the conclusion uses the original-law pull count. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedPull_div_log_ge_inv_dInf","label":"consistentPolicy_liminf_expectedPull_div_log_ge_inv_dInf","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedPull_div_log_ge_inv_dInf","description":"The exact per-arm information constraint supporting Theorem 16.2. The inverse infimum remains extended-real: zero infimum forces infinite liminf, while an empty or all-infinite alternative class contributes zero.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-9dba7d263a6a","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5793,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1653"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem consistentPolicy_liminf_expectedPull_div_log_ge_inv_dInf {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (componentClass : Fin K → Set (Measure Real)) (hfinite : ∀ arm P, P ∈ componentClass arm → IsProbabilityMeasure P ∧ Integrable id P) (hconsistent : IsConsistentPolicyOver {environment : FiniteMeanBanditEnvironment K | environment.InUnstructuredClass componentClass} finiteMeanExpectedRegret algorithm) (original : FiniteMeanBanditEnvironment K) (hclass : original.InUnstructuredClass componentClass) (changedArm : Fin K) (hsuboptimal : original.mean changedArm < original.mean original.bestArm) : (divergenceInfimum (original.armLaw changedArm) (original.mean original.bestArm) (componentClass changedArm) (fun P => ∫ x, x ∂P))⁻¹ ≤ liminf (fun n : Nat => ENNReal.ofReal (finiteMeanExpectedPullCount algorithm original changedArm n / Real.log n)) atTop","missing":[],"search":"consistentpolicy_liminf_expectedpull_div_log_ge_inv_dinf banditrlproof.lowerbounds.consistentpolicy_liminf_expectedpull_div_log_ge_inv_dinf the exact per-arm information constraint supporting theorem 16.2. the inverse infimum remains extended-real: zero infimum forces infinite liminf, while an empty or all-infinite alternative class contributes zero. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedRegret_div_log_ge","label":"consistentPolicy_liminf_expectedRegret_div_log_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedRegret_div_log_ge","description":"*Lattimore--Szepesvári, Theorem 16.2.** The unstructured finite-mean product-class regret lower bound, including zero and infinite information costs. All quotients and the liminf are extended-real.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-2d8cc2fd21b7","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5794,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1705"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem consistentPolicy_liminf_expectedRegret_div_log_ge {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (componentClass : Fin K → Set (Measure Real)) (hfinite : ∀ arm P, P ∈ componentClass arm → IsProbabilityMeasure P ∧ Integrable id P) (hconsistent : IsConsistentPolicyOver {environment : FiniteMeanBanditEnvironment K | environment.InUnstructuredClass componentClass} finiteMeanExpectedRegret algorithm) (original : FiniteMeanBanditEnvironment K) (hclass : original.InUnstructuredClass componentClass) : (∑ arm : Fin K with 0 < original.gap arm, ENNReal.ofReal (original.gap arm) / divergenceInfimum (original.armLaw arm) (original.mean original.bestArm) (componentClass arm) (fun P => ∫ x, x ∂P)) ≤ liminf (fun n : Nat => ENNReal.ofReal (finiteMeanExpectedRegret algorithm original n / Real.log n)) atTop","missing":[],"search":"consistentpolicy_liminf_expectedregret_div_log_ge banditrlproof.lowerbounds.consistentpolicy_liminf_expectedregret_div_log_ge *lattimore--szepesvári, theorem 16.2.** the unstructured finite-mean product-class regret lower bound, including zero and infinite information costs. all quotients and the liminf are extended-real. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPullCount_ge_finiteTimeInstanceDependent","label":"gaussianExpectedPullCount_ge_finiteTimeInstanceDependent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedPullCount_ge_finiteTimeInstanceDependent","description":"Per-arm finite-time Gaussian consequence used in Theorem 16.4, before summing and taking positive parts. Both regret bounds are retained explicitly so the published `2 C n^p` denominator is visible.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-e68b4dc8f03b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5795,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1767"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem gaussianExpectedPullCount_ge_finiteTimeInstanceDependent {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitVarianceGaussianBanditEnvironment K) (changedArm : Fin K) (epsilon C p : Real) (hgap : 0 < environment.gap changedArm) (hepsilon : 0 < epsilon) (hepsilon_one : epsilon <= 1) (_hC : 0 < C) (lastRound : Nat) (hbase : unitVarianceGaussianExpectedPseudoRegret algorithm environment lastRound <= C * (((lastRound + 1 : Nat) : Real) ^ p)) (hchanged : unitVarianceGaussianExpectedPseudoRegret algorithm (chapter16GaussianChangedEnvironment environment changedArm epsilon hgap hepsilon) lastRound <= C * (((lastRound + 1 : Nat) : Real) ^ p)) : (Real.log (epsilon * environment.gap changedArm / 4) + Real.log ((lastRound + 1 : Nat) : Real) - Real.log (2 * C * (((lastRound + 1 : Nat) : Real) ^ p))) / (((1 + epsilon) * environment.gap changedArm) ^ 2 / 2) <= (c…","missing":[],"search":"gaussianexpectedpullcount_ge_finitetimeinstancedependent banditrlproof.lowerbounds.gaussianexpectedpullcount_ge_finitetimeinstancedependent per-arm finite-time gaussian consequence used in theorem 16.4, before summing and taking positive parts. both regret bounds are retained explicitly so the published `2 c n^p` denominator is visible. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.chapter16Gaussian_finiteTime_log_identity","label":"chapter16Gaussian_finiteTime_log_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.chapter16Gaussian_finiteTime_log_identity","description":"Exact logarithmic normalization in the displayed bound (16.5).","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-3c6342a48925","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5796,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1879"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem chapter16Gaussian_finiteTime_log_identity (epsilon gap C p horizon : Real) (hepsilon : 0 < epsilon) (hgap : 0 < gap) (hC : 0 < C) (hhorizon : 0 < horizon) : Real.log (epsilon * gap / 4) + Real.log horizon - Real.log (2 * C * horizon ^ p) = (1 - p) * Real.log horizon + Real.log (epsilon * gap / (8 * C))","missing":[],"search":"chapter16gaussian_finitetime_log_identity banditrlproof.lowerbounds.chapter16gaussian_finitetime_log_identity exact logarithmic normalization in the displayed bound (16.5). theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","label":"gaussianExpectedRegret_ge_finiteTimeInstanceDependent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","description":"*Lattimore--Szepesvári, Theorem 16.4.** Finite-time instance-dependent lower bound for arbitrary unit-variance Gaussian means. `lastRound + 1` is the source horizon `n`; the supplied set `N` is therefore stated on positive horizon lengths.","url":"../modules/banditrlproof-lowerbounds-instancedependent/index.html#decl-31841dba3b9b","parent":"module:BanditRLProof.LowerBounds.InstanceDependent","order":5797,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.InstanceDependent"],["Source","BanditRLProof/LowerBounds/InstanceDependent.lean:1909"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-16-instance-dependent"],["Indexed settings","None registered"]],"statement":"theorem gaussianExpectedRegret_ge_finiteTimeInstanceDependent {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : UnitVarianceGaussianBanditEnvironment K) (horizons : Set Nat) (hhorizons : horizons.Nonempty) (C p : Real) (hC : 0 < C) (_hp : p ∈ Set.Ioo (0 : Real) 1) (hregret : forall (lastRound : Nat), lastRound + 1 ∈ horizons -> forall candidate : UnitVarianceGaussianBanditEnvironment K, InChapter16GaussianLocalClass environment candidate -> unitVarianceGaussianExpectedPseudoRegret algorithm candidate lastRound <= C * (((lastRound + 1 : Nat) : Real) ^ p)) (epsilon : Real) (hepsilon : epsilon ∈ Set.Ioc (0 : Real) 1) (lastRound : Nat) (hhorizon : lastRound + 1 ∈ horizons) : unitVarianceGaussianExpectedPseudoRegret algorithm environment lastRound >= 2 / (1 + epsilon) ^ 2 * ∑ arm ∈ Finset.univ.filter (fun arm : Fin K => 0 < environment.gap arm), max (((1 - p) * Re…","missing":[],"search":"gaussianexpectedregret_ge_finitetimeinstancedependent banditrlproof.lowerbounds.gaussianexpectedregret_ge_finitetimeinstancedependent *lattimore--szepesvári, theorem 16.4.** finite-time instance-dependent lower bound for arbitrary unit-variance gaussian means. `lastround + 1` is the source horizon `n`; the supplied set `n` is therefore stated on positive horizon lengths. theorem compiled","shard":"modules/361365b9df4da239.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-16-instance-dependent"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianArm","label":"unitGaussianArm","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianArm","description":"A unit-variance Gaussian arm with mean `mu`.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-7612ccae3d47","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5798,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:23"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"abbrev unitGaussianArm (mu : Real) : Measure Real","missing":[],"search":"unitgaussianarm banditrlproof.lowerbounds.unitgaussianarm a unit-variance gaussian arm with mean `mu`. abbreviation compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianBandit","label":"unitGaussianBandit","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianBandit","description":"The finite family of unit-variance Gaussian arms indexed by a mean vector.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-416248b74706","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5799,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"abbrev unitGaussianBandit {k : Nat} (mean : Fin k -> Real) : Fin k -> Measure Real","missing":[],"search":"unitgaussianbandit banditrlproof.lowerbounds.unitgaussianbandit the finite family of unit-variance gaussian arms indexed by a mean vector. abbreviation compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_gaussianPDFReal_one","label":"log_gaussianPDFReal_div_gaussianPDFReal_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.log_gaussianPDFReal_div_gaussianPDFReal_one","description":"Pointwise log-density ratio for two unit-variance Gaussian laws.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-b35b711abcf8","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5800,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem log_gaussianPDFReal_div_gaussianPDFReal_one (mu nu x : Real) : Real.log (gaussianPDFReal mu (1 : NNReal) x / gaussianPDFReal nu (1 : NNReal) x) = (mu - nu) * x + (nu ^ 2 - mu ^ 2) / 2","missing":[],"search":"log_gaussianpdfreal_div_gaussianpdfreal_one banditrlproof.lowerbounds.log_gaussianpdfreal_div_gaussianpdfreal_one pointwise log-density ratio for two unit-variance gaussian laws. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","label":"llr_gaussianReal_one_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","description":"The log Radon--Nikodym derivative has the expected affine form under the first Gaussian law. The direction is `N(mu,1)` to `N(nu,1)`.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-b8c0eca2b278","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5801,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:49"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem llr_gaussianReal_one_ae (mu nu : Real) : llr (unitGaussianArm mu) (unitGaussianArm nu) =ᵐ[unitGaussianArm mu] fun x => (mu - nu) * x + (nu ^ 2 - mu ^ 2) / 2","missing":[],"search":"llr_gaussianreal_one_ae banditrlproof.lowerbounds.llr_gaussianreal_one_ae the log radon--nikodym derivative has the expected affine form under the first gaussian law. the direction is `n(mu,1)` to `n(nu,1)`. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_one","label":"integrable_llr_gaussianReal_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.integrable_llr_gaussianReal_one","description":"The Gaussian log likelihood ratio is integrable under its first law.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-ed027cfd8488","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5802,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:69"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem integrable_llr_gaussianReal_one (mu nu : Real) : Integrable (llr (unitGaussianArm mu) (unitGaussianArm nu)) (unitGaussianArm mu)","missing":[],"search":"integrable_llr_gaussianreal_one banditrlproof.lowerbounds.integrable_llr_gaussianreal_one the gaussian log likelihood ratio is integrable under its first law. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","label":"klDiv_gaussianReal_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_gaussianReal_one","description":"Exact KL divergence between unit-variance Gaussian laws: `D(N(mu,1),N(nu,1))=(mu-nu)^2/2`. The result is stated in `ENNReal`, matching Mathlib's measure-KL API and retaining the source direction even though this equal-variance value happens to be symmetric in the two means.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-c9c669679982","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5803,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:85"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem klDiv_gaussianReal_one (mu nu : Real) : InformationTheory.klDiv (unitGaussianArm mu) (unitGaussianArm nu) = ENNReal.ofReal ((mu - nu) ^ 2 / 2)","missing":[],"search":"kldiv_gaussianreal_one banditrlproof.lowerbounds.kldiv_gaussianreal_one exact kl divergence between unit-variance gaussian laws: `d(n(mu,1),n(nu,1))=(mu-nu)^2/2`. the result is stated in `ennreal`, matching mathlib's measure-kl api and retaining the source direction even though this equal-variance value happens to be symmetric in the two means. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianArm_zero_two_mul","label":"klDiv_unitGaussianArm_zero_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.klDiv_unitGaussianArm_zero_two_mul","description":"Source-specialized Gaussian KL value for the changed arm in the proof of Theorem 15.2.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-b1a94cf47ba2","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5804,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:117"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem klDiv_unitGaussianArm_zero_two_mul (gap : Real) : InformationTheory.klDiv (unitGaussianArm 0) (unitGaussianArm (2 * gap)) = ENNReal.ofReal (2 * gap ^ 2)","missing":[],"search":"kldiv_unitgaussianarm_zero_two_mul banditrlproof.lowerbounds.kldiv_unitgaussianarm_zero_two_mul source-specialized gaussian kl value for the changed arm in the proof of theorem 15.2. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap","label":"gaussianMinimaxGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxGap","description":"The source tuning `Delta=sqrt(m/(4n))`, written for positive real-valued alternative count `m` and horizon `n`. Natural-count consumers must discharge their cast and positivity obligations explicitly.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-f8a30ac23a7d","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5805,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:127"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianMinimaxGap (alternativeCount horizon : Real) : Real","missing":[],"search":"gaussianminimaxgap banditrlproof.lowerbounds.gaussianminimaxgap the source tuning `delta=sqrt(m/(4n))`, written for positive real-valued alternative count `m` and horizon `n`. natural-count consumers must discharge their cast and positivity obligations explicitly. definition compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_sq","label":"gaussianMinimaxGap_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxGap_sq","description":"Squared form of the Chapter 15 minimax gap choice.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-e3c060e284df","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5806,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:132"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxGap_sq {alternativeCount horizon : Real} (halternatives : 0 ≤ alternativeCount) (hhorizon : 0 ≤ horizon) : gaussianMinimaxGap alternativeCount horizon ^ 2 = alternativeCount / (4 * horizon)","missing":[],"search":"gaussianminimaxgap_sq banditrlproof.lowerbounds.gaussianminimaxgap_sq squared form of the chapter 15 minimax gap choice. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","label":"gaussianMinimaxGap_informationExponent_eq_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","description":"With the source tuning, the history-KL upper exponent `2*n*Delta^2/m` is exactly `1/2`. This is a numeric dependency only: it does not supply the history-KL upper bound itself.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-98c8a32e8df5","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5807,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:143"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxGap_informationExponent_eq_half {alternativeCount horizon : Real} (halternatives : 0 < alternativeCount) (hhorizon : 0 < horizon) : 2 * horizon * gaussianMinimaxGap alternativeCount horizon ^ 2 / alternativeCount = 1 / 2","missing":[],"search":"gaussianminimaxgap_informationexponent_eq_half banditrlproof.lowerbounds.gaussianminimaxgap_informationexponent_eq_half with the source tuning, the history-kl upper exponent `2*n*delta^2/m` is exactly `1/2`. this is a numeric dependency only: it does not supply the history-kl upper bound itself. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_le_half","label":"gaussianMinimaxGap_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.gaussianMinimaxGap_le_half","description":"If the number of alternative arms is at most the horizon, the tuned gap is at most one half, so the source means `Delta` and `2*Delta` lie in the unit interval after nonnegativity is combined with this leaf.","url":"../modules/banditrlproof-lowerbounds-minimax/index.html#decl-bcf6bc862a01","parent":"module:BanditRLProof.LowerBounds.Minimax","order":5808,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.Minimax"],["Source","BanditRLProof/LowerBounds/Minimax.lean:155"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-15-minimax-lower-bounds"],["Indexed settings","None registered"]],"statement":"theorem gaussianMinimaxGap_le_half {alternativeCount horizon : Real} (hhorizon : 0 < horizon) (hcount_le : alternativeCount ≤ horizon) : gaussianMinimaxGap alternativeCount horizon ≤ 1 / 2","missing":[],"search":"gaussianminimaxgap_le_half banditrlproof.lowerbounds.gaussianminimaxgap_le_half if the number of alternative arms is at most the horizon, the tuned gap is at most one half, so the source means `delta` and `2*delta` lie in the unit interval after nonnegativity is combined with this leaf. theorem compiled","shard":"modules/6f019e70c12d35c3.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-15-minimax-lower-bounds"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryWords","label":"binaryWords","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryWords","description":"All words at depth n in the full binary tree.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-b70e489afe85","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5809,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def binaryWords (n : ℕ) : Finset (List Bool)","missing":[],"search":"binarywords banditrlproof.lowerbounds.binarywords all words at depth n in the full binary tree. definition compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.card_binaryWords","label":"card_binaryWords","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.card_binaryWords","description":"theorem card_binaryWords (n : ℕ) : (binaryWords n).card = 2 ^ n","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-1ff8550e3c04","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5810,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem card_binaryWords (n : ℕ) : (binaryWords n).card = 2 ^ n","missing":[],"search":"card_binarywords banditrlproof.lowerbounds.card_binarywords theorem card_binarywords (n : ℕ) : (binarywords n).card = 2 ^ n theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.mem_binaryWords_iff","label":"mem_binaryWords_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.mem_binaryWords_iff","description":"theorem mem_binaryWords_iff (w : List Bool) (n : ℕ) : w ∈ binaryWords n ↔ w.length = n","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-0ceacd55e470","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5811,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_binaryWords_iff (w : List Bool) (n : ℕ) : w ∈ binaryWords n ↔ w.length = n","missing":[],"search":"mem_binarywords_iff banditrlproof.lowerbounds.mem_binarywords_iff theorem mem_binarywords_iff (w : list bool) (n : ℕ) : w ∈ binarywords n ↔ w.length = n theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryExtensions","label":"binaryExtensions","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryExtensions","description":"The descendants of a prefix after n additional bits.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-63f65f4da027","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5812,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def binaryExtensions (w : List Bool) (n : ℕ) : Finset (List Bool)","missing":[],"search":"binaryextensions banditrlproof.lowerbounds.binaryextensions the descendants of a prefix after n additional bits. definition compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.card_binaryExtensions","label":"card_binaryExtensions","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.card_binaryExtensions","description":"theorem card_binaryExtensions (w : List Bool) (n : ℕ) : (binaryExtensions w n).card = 2 ^ n","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-6c822f07a58b","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5813,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:29"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem card_binaryExtensions (w : List Bool) (n : ℕ) : (binaryExtensions w n).card = 2 ^ n","missing":[],"search":"card_binaryextensions banditrlproof.lowerbounds.card_binaryextensions theorem card_binaryextensions (w : list bool) (n : ℕ) : (binaryextensions w n).card = 2 ^ n theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.mem_binaryExtensions_iff","label":"mem_binaryExtensions_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.mem_binaryExtensions_iff","description":"theorem mem_binaryExtensions_iff (w v : List Bool) (n : ℕ) : v ∈ binaryExtensions w n ↔ w <+: v ∧ v.length = w.length + n","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-7807d1b55063","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5814,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem mem_binaryExtensions_iff (w v : List Bool) (n : ℕ) : v ∈ binaryExtensions w n ↔ w <+: v ∧ v.length = w.length + n","missing":[],"search":"mem_binaryextensions_iff banditrlproof.lowerbounds.mem_binaryextensions_iff theorem mem_binaryextensions_iff (w v : list bool) (n : ℕ) : v ∈ binaryextensions w n ↔ w <+: v ∧ v.length = w.length + n theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryExtensions_disjoint_of_incomparable","label":"binaryExtensions_disjoint_of_incomparable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryExtensions_disjoint_of_incomparable","description":"Cylinders from incomparable prefixes are disjoint.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-8e33b19da815","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5815,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem binaryExtensions_disjoint_of_incomparable (u v : List Bool) (m n : ℕ) (huv : ¬ u <+: v) (hvu : ¬ v <+: u) : Disjoint (binaryExtensions u m) (binaryExtensions v n)","missing":[],"search":"binaryextensions_disjoint_of_incomparable banditrlproof.lowerbounds.binaryextensions_disjoint_of_incomparable cylinders from incomparable prefixes are disjoint. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binaryWord_avoiding_prefixes","label":"exists_binaryWord_avoiding_prefixes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binaryWord_avoiding_prefixes","description":"A free word exists whenever previous prefixes occupy fewer than all level nodes.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-362009dd3006","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5816,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binaryWord_avoiding_prefixes (S : Finset (List Bool)) (n : ℕ) (hlen : ∀ w ∈ S, w.length ≤ n) (hbudget : (∑ w ∈ S, 2 ^ (n - w.length)) < 2 ^ n) : ∃ v : List Bool, v.length = n ∧ ∀ w ∈ S, ¬ w <+: v","missing":[],"search":"exists_binaryword_avoiding_prefixes banditrlproof.lowerbounds.exists_binaryword_avoiding_prefixes a free word exists whenever previous prefixes occupy fewer than all level nodes. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binary_level_mul_kraft_weight","label":"binary_level_mul_kraft_weight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binary_level_mul_kraft_weight","description":"theorem binary_level_mul_kraft_weight {k n : ℕ} (h : k ≤ n) : (2 : ℝ) ^ n * (1 / 2 : ℝ) ^ k = 2 ^ (n - k)","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-0ae5c1ff20fd","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5817,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:81"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem binary_level_mul_kraft_weight {k n : ℕ} (h : k ≤ n) : (2 : ℝ) ^ n * (1 / 2 : ℝ) ^ k = 2 ^ (n - k)","missing":[],"search":"binary_level_mul_kraft_weight banditrlproof.lowerbounds.binary_level_mul_kraft_weight theorem binary_level_mul_kraft_weight {k n : ℕ} (h : k ≤ n) : (2 : ℝ) ^ n * (1 / 2 : ℝ) ^ k = 2 ^ (n - k) theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binaryWord_of_kraft_lt_one","label":"exists_binaryWord_of_kraft_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binaryWord_of_kraft_lt_one","description":"Real Kraft slack supplies the integer capacity needed to insert a word.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-3b58c662ef20","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5818,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:89"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binaryWord_of_kraft_lt_one (S : Finset (List Bool)) (n : ℕ) (hlen : ∀ w ∈ S, w.length ≤ n) (hk : (∑ w ∈ S, (1 / 2 : ℝ) ^ w.length) < 1) : ∃ v : List Bool, v.length = n ∧ ∀ w ∈ S, ¬ w <+: v","missing":[],"search":"exists_binaryword_of_kraft_lt_one banditrlproof.lowerbounds.exists_binaryword_of_kraft_lt_one real kraft slack supplies the integer capacity needed to insert a word. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_prefixFree_insert_of_kraft_lt_one","label":"exists_prefixFree_insert_of_kraft_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_prefixFree_insert_of_kraft_lt_one","description":"The greedy insertion preserves prefix freedom in both directions.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-5ba43cf9ebf7","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5819,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:104"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_prefixFree_insert_of_kraft_lt_one (S : Finset (List Bool)) (n : ℕ) (hfree : ∀ a ∈ S, ∀ b ∈ S, a <+: b → a = b) (hlen : ∀ w ∈ S, w.length ≤ n) (hk : (∑ w ∈ S, (1 / 2 : ℝ) ^ w.length) < 1) : ∃ v : List Bool, v.length = n ∧ v ∉ S ∧ ∀ a ∈ insert v S, ∀ b ∈ insert v S, a <+: b → a = b","missing":[],"search":"exists_prefixfree_insert_of_kraft_lt_one banditrlproof.lowerbounds.exists_prefixfree_insert_of_kraft_lt_one the greedy insertion preserves prefix freedom in both directions. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_prefix_encoding_of_kraft_le_one","label":"exists_prefix_encoding_of_kraft_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_prefix_encoding_of_kraft_le_one","description":"Finite Kraft converse, retaining each prescribed length, including equality.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-d37ba87472d6","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5820,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:139"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_prefix_encoding_of_kraft_le_one {α : Type*} [DecidableEq α] (s : Finset α) (l : α → ℕ) (hk : (∑ i ∈ s, (1 / 2 : ℝ) ^ l i) ≤ 1) : ∃ c : α → List Bool, (∀ i ∈ s, (c i).length = l i) ∧ (∀ i ∈ s, ∀ j ∈ s, c i <+: c j → i = j)","missing":[],"search":"exists_prefix_encoding_of_kraft_le_one banditrlproof.lowerbounds.exists_prefix_encoding_of_kraft_le_one finite kraft converse, retaining each prescribed length, including equality. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_prefix_encoding_of_kraft_lt_one","label":"exists_prefix_encoding_of_kraft_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_prefix_encoding_of_kraft_lt_one","description":"Backwards-compatible strict version of the finite Kraft converse.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-b04d6f278535","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5821,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:217"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_prefix_encoding_of_kraft_lt_one {α : Type*} [DecidableEq α] (s : Finset α) (l : α → ℕ) (hk : (∑ i ∈ s, (1 / 2 : ℝ) ^ l i) < 1) : ∃ c : α → List Bool, (∀ i ∈ s, (c i).length = l i) ∧ (∀ i ∈ s, ∀ j ∈ s, c i <+: c j → i = j)","missing":[],"search":"exists_prefix_encoding_of_kraft_lt_one banditrlproof.lowerbounds.exists_prefix_encoding_of_kraft_lt_one backwards-compatible strict version of the finite kraft converse. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binaryPrefixCode_of_kraft_le_one","label":"exists_binaryPrefixCode_of_kraft_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binaryPrefixCode_of_kraft_le_one","description":"Non-strict Kraft converse packaged as an actual binary prefix code.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-cdf4be736364","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5822,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:225"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binaryPrefixCode_of_kraft_le_one {α : Type*} [Fintype α] [DecidableEq α] (l : α → ℕ) (hl : ∀ i, 0 < l i) (hk : (∑ i, (1 / 2 : ℝ) ^ l i) ≤ 1) : ∃ code : BinaryPrefixCode α, ∀ i, (code.encode i).length = l i","missing":[],"search":"exists_binaryprefixcode_of_kraft_le_one banditrlproof.lowerbounds.exists_binaryprefixcode_of_kraft_le_one non-strict kraft converse packaged as an actual binary prefix code. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binaryPrefixCode_of_kraft_lt_one","label":"exists_binaryPrefixCode_of_kraft_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binaryPrefixCode_of_kraft_lt_one","description":"Backwards-compatible strict Kraft packaging.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-ba12bec415d8","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5823,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:247"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binaryPrefixCode_of_kraft_lt_one {α : Type*} [Fintype α] [DecidableEq α] (l : α → ℕ) (hl : ∀ i, 0 < l i) (hk : (∑ i, (1 / 2 : ℝ) ^ l i) < 1) : ∃ code : BinaryPrefixCode α, ∀ i, (code.encode i).length = l i","missing":[],"search":"exists_binaryprefixcode_of_kraft_lt_one banditrlproof.lowerbounds.exists_binaryprefixcode_of_kraft_lt_one backwards-compatible strict kraft packaging. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_prefixCode_of_uniquelyDecodable","label":"exists_prefixCode_of_uniquelyDecodable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_prefixCode_of_uniquelyDecodable","description":"Every finite uniquely decodable encoder has a prefix code with exactly the same symbol lengths (the boxed assertion in Chapter 14, Section 14.1).","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-8de6b7c62c08","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5824,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:255"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem exists_prefixCode_of_uniquelyDecodable {α : Type*} [Fintype α] (c : α → List Bool) (hinj : Function.Injective c) (hud : InformationTheory.UniquelyDecodable (Set.range c)) : ∃ code : BinaryPrefixCode α, ∀ i, (code.encode i).length = (c i).length","missing":[],"search":"exists_prefixcode_of_uniquelydecodable banditrlproof.lowerbounds.exists_prefixcode_of_uniquelydecodable every finite uniquely decodable encoder has a prefix code with exactly the same symbol lengths (the boxed assertion in chapter 14, section 14.1). theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_binaryPrefixCode_entropy_sandwich","label":"exists_binaryPrefixCode_entropy_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_binaryPrefixCode_entropy_sandwich","description":"A realizable finite prefix code attains the one-bit entropy sandwich. This proves existence, not Huffman's algorithm or optimality.","url":"../modules/banditrlproof-lowerbounds-prefixcodeconstruction/index.html#decl-6c9bcc0b914c","parent":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","order":5825,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeConstruction"],["Source","BanditRLProof/LowerBounds/PrefixCodeConstruction.lean:274"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_binaryPrefixCode_entropy_sandwich {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : ∃ code : BinaryPrefixCode α, discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength p code ∧ expectedCodeLength p code ≤ discreteEntropyBaseTwo Finset.univ p + 1","missing":[],"search":"exists_binaryprefixcode_entropy_sandwich banditrlproof.lowerbounds.exists_binaryprefixcode_entropy_sandwich a realizable finite prefix code attains the one-bit entropy sandwich. this proves existence, not huffman's algorithm or optimality. theorem compiled","shard":"modules/468516b3f46bed90.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.relabel","label":"relabel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.relabel","description":"def BinaryPrefixCode.relabel {α β : Type*} (code : BinaryPrefixCode α) (e : β ≃ α) : BinaryPrefixCode β where","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-5fb4f9fd634e","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5826,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def BinaryPrefixCode.relabel {α β : Type*} (code : BinaryPrefixCode α) (e : β ≃ α) : BinaryPrefixCode β where","missing":[],"search":"relabel banditrlproof.lowerbounds.binaryprefixcode.relabel def binaryprefixcode.relabel {α β : type*} (code : binaryprefixcode α) (e : β ≃ α) : binaryprefixcode β where definition compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode","label":"IsOptimalPrefixCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsOptimalPrefixCode","description":"Full optimality against every code, not merely against a chosen candidate family.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-9b41a69f3a7c","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5827,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:13"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def IsOptimalPrefixCode {α : Type*} [Fintype α] (p : α → ℝ) (code : BinaryPrefixCode α) : Prop","missing":[],"search":"isoptimalprefixcode banditrlproof.lowerbounds.isoptimalprefixcode full optimality against every code, not merely against a chosen candidate family. definition compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_swap","label":"expectedCodeLength_swap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_swap","description":"theorem expectedCodeLength_swap {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a b : α) (hab : a ≠ b) : expectedCodeLength p (code.relabel (Equiv.swap a b)) = expectedCodeLength p code + (p a - p b) * ((code.encode b).length - (code.encode a).length : ℝ)","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-fb00eb92a757","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5828,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_swap {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a b : α) (hab : a ≠ b) : expectedCodeLength p (code.relabel (Equiv.swap a b)) = expectedCodeLength p code + (p a - p b) * ((code.encode b).length - (code.encode a).length : ℝ)","missing":[],"search":"expectedcodelength_swap banditrlproof.lowerbounds.expectedcodelength_swap theorem expectedcodelength_swap {α : type*} [fintype α] [decidableeq α] (p : α → ℝ) (code : binaryprefixcode α) (a b : α) (hab : a ≠ b) : expectedcodelength p (code.relabel (equiv.swap a b)) = expectedcodelength p code + (p a - p b) * ((code.encode b).length - (code.encode a).length : ℝ) theorem compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_swap_le","label":"expectedCodeLength_swap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_swap_le","description":"Assigning a shorter word to a higher-probability symbol cannot increase cost.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-b45d4720c0bb","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5829,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_swap_le {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a b : α) (hab : a ≠ b) (hp : p a ≤ p b) (hl : (code.encode a).length ≤ (code.encode b).length) : expectedCodeLength p (code.relabel (Equiv.swap a b)) ≤ expectedCodeLength p code","missing":[],"search":"expectedcodelength_swap_le banditrlproof.lowerbounds.expectedcodelength_swap_le assigning a shorter word to a higher-probability symbol cannot increase cost. theorem compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.length_antitone","label":"length_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsOptimalPrefixCode.length_antitone","description":"An optimal code orders lengths opposite to strictly ordered probabilities.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-efd0ffd841ea","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5830,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem IsOptimalPrefixCode.length_antitone {α : Type*} [Fintype α] [DecidableEq α] {p : α → ℝ} {code : BinaryPrefixCode α} (hopt : IsOptimalPrefixCode p code) (a b : α) (hp : p a < p b) : (code.encode b).length ≤ (code.encode a).length","missing":[],"search":"length_antitone banditrlproof.lowerbounds.isoptimalprefixcode.length_antitone an optimal code orders lengths opposite to strictly ordered probabilities. theorem compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.entropy_sandwich","label":"entropy_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.IsOptimalPrefixCode.entropy_sandwich","description":"Any global minimizer inherits the entropy sandwich; this does not assert existence.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-aa8c1d7e078e","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5831,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:66"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem IsOptimalPrefixCode.entropy_sandwich {α : Type*} [Fintype α] [DecidableEq α] {p : α → ℝ} {code : BinaryPrefixCode α} (hopt : IsOptimalPrefixCode p code) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : discreteEntropyBaseTwo Finset.univ p ≤ expectedCodeLength p code ∧ expectedCodeLength p code ≤ discreteEntropyBaseTwo Finset.univ p + 1","missing":[],"search":"entropy_sandwich banditrlproof.lowerbounds.isoptimalprefixcode.entropy_sandwich any global minimizer inherits the entropy sandwich; this does not assert existence. theorem compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.one_le_expectedCodeLength","label":"one_le_expectedCodeLength","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.one_le_expectedCodeLength","description":"The local nonempty-codeword convention forces at least one expected bit.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-0843289eeb67","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5832,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:75"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem one_le_expectedCodeLength {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) (code : BinaryPrefixCode α) : 1 ≤ expectedCodeLength p code","missing":[],"search":"one_le_expectedcodelength banditrlproof.lowerbounds.one_le_expectedcodelength the local nonempty-codeword convention forces at least one expected bit. theorem compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.singletonPrefixCode","label":"singletonPrefixCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.singletonPrefixCode","description":"The singleton alphabet's nonempty one-bit code.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-a3c0e1101d5a","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5833,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:90"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def singletonPrefixCode (α : Type*) [Subsingleton α] : BinaryPrefixCode α where","missing":[],"search":"singletonprefixcode banditrlproof.lowerbounds.singletonprefixcode the singleton alphabet's nonempty one-bit code. definition compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.singletonPrefixCode_optimal","label":"singletonPrefixCode_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.singletonPrefixCode_optimal","description":"The singleton base case has a genuine global optimum under the local convention.","url":"../modules/banditrlproof-lowerbounds-prefixcodeexchange/index.html#decl-05f3f22f4a59","parent":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","order":5834,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeExchange"],["Source","BanditRLProof/LowerBounds/PrefixCodeExchange.lean:97"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem singletonPrefixCode_optimal {α : Type*} [Fintype α] [Subsingleton α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : IsOptimalPrefixCode p (singletonPrefixCode α)","missing":[],"search":"singletonprefixcode_optimal banditrlproof.lowerbounds.singletonprefixcode_optimal the singleton base case has a genuine global optimum under the local convention. theorem compiled","shard":"modules/b426b006e5df32f0.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_swap_le_allow_eq","label":"expectedCodeLength_swap_le_allow_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_swap_le_allow_eq","description":"The exchange bound also permits an already correctly placed label.","url":"../modules/banditrlproof-lowerbounds-prefixcodegreedy/index.html#decl-42aa59c905c6","parent":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","order":5835,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeGreedy"],["Source","BanditRLProof/LowerBounds/PrefixCodeGreedy.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_swap_le_allow_eq {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a b : α) (hp : p a ≤ p b) (hl : (code.encode a).length ≤ (code.encode b).length) : expectedCodeLength p (code.relabel (Equiv.swap a b)) ≤ expectedCodeLength p code","missing":[],"search":"expectedcodelength_swap_le_allow_eq banditrlproof.lowerbounds.expectedcodelength_swap_le_allow_eq the exchange bound also permits an already correctly placed label. theorem compiled","shard":"modules/5d36df5c68c900cb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_no_worse_least_weight_siblings","label":"exists_no_worse_least_weight_siblings","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_no_worse_least_weight_siblings","description":"Huffman's greedy choice: any competitor can place two specified least weights at deepest sibling leaves without increasing expected length. Ties and zero weights are permitted, and no existence of an optimal code is assumed.","url":"../modules/banditrlproof-lowerbounds-prefixcodegreedy/index.html#decl-666a538f3372","parent":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","order":5836,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeGreedy"],["Source","BanditRLProof/LowerBounds/PrefixCodeGreedy.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_no_worse_least_weight_siblings {α : Type*} [Fintype α] [DecidableEq α] [Nontrivial α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (original : BinaryPrefixCode α) (a b : α) (hab : a ≠ b) (ha : ∀ i, p a ≤ p i) (hb : ∀ i, i ≠ a → p b ≤ p i) : ∃ code : BinaryPrefixCode α, expectedCodeLength p code ≤ expectedCodeLength p original ∧ ∃ w bit, code.encode a = w ++ [bit] ∧ code.encode b = w ++ [!bit] ∧ ∀ i, (code.encode i).length ≤ (code.encode a).length","missing":[],"search":"exists_no_worse_least_weight_siblings banditrlproof.lowerbounds.exists_no_worse_least_weight_siblings huffman's greedy choice: any competitor can place two specified least weights at deepest sibling leaves without increasing expected length. ties and zero weights are permitted, and no existence of an optimal code is assumed. theorem compiled","shard":"modules/5d36df5c68c900cb.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.deepest_parent_incomparable","label":"deepest_parent_incomparable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.deepest_parent_incomparable","description":"A deepest leaf with no sibling has a parent incomparable with every other word.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-0e28fcf9ce0a","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5837,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem deepest_parent_incomparable {α : Type*} (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : ∀ i, i ≠ a → (¬ code.encode i <+: w) ∧ (¬ w <+: code.encode i)","missing":[],"search":"deepest_parent_incomparable banditrlproof.lowerbounds.deepest_parent_incomparable a deepest leaf with no sibling has a parent incomparable with every other word. theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.replaceWord","label":"replaceWord","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.replaceWord","description":"Replace one word by an incomparable nonempty word, preserving code validity.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-a32a1dad4e50","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5838,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def BinaryPrefixCode.replaceWord {α : Type*} [DecidableEq α] (code : BinaryPrefixCode α) (a : α) (w : List Bool) (hw : w ≠ []) (hsep : ∀ i, i ≠ a → (¬ code.encode i <+: w) ∧ (¬ w <+: code.encode i)) : BinaryPrefixCode α","missing":[],"search":"replaceword banditrlproof.lowerbounds.binaryprefixcode.replaceword replace one word by an incomparable nonempty word, preserving code validity. definition compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.pruneDeepest","label":"pruneDeepest","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.pruneDeepest","description":"Legal deletion of the last bit of a deepest sibling-free leaf.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-303a9e7e3226","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5839,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def BinaryPrefixCode.pruneDeepest {α : Type*} [DecidableEq α] (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : BinaryPrefixCode α","missing":[],"search":"prunedeepest banditrlproof.lowerbounds.binaryprefixcode.prunedeepest legal deletion of the last bit of a deepest sibling-free leaf. definition compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_replaceWord","label":"expectedCodeLength_replaceWord","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_replaceWord","description":"theorem expectedCodeLength_replaceWord {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a : α) (w : List Bool) (hw : w ≠ []) (hsep : ∀ i, i ≠ a → (¬ code.encode i <+: w) ∧ (¬ w <+: code.encode i)) : expectedCodeLength p (code.replaceWord a w hw hsep) = expectedCodeLength p code + p a * (w.length - (code.encode a).length : ℝ)","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-d0b6e59bc095","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5840,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_replaceWord {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a : α) (w : List Bool) (hw : w ≠ []) (hsep : ∀ i, i ≠ a → (¬ code.encode i <+: w) ∧ (¬ w <+: code.encode i)) : expectedCodeLength p (code.replaceWord a w hw hsep) = expectedCodeLength p code + p a * (w.length - (code.encode a).length : ℝ)","missing":[],"search":"expectedcodelength_replaceword banditrlproof.lowerbounds.expectedcodelength_replaceword theorem expectedcodelength_replaceword {α : type*} [fintype α] [decidableeq α] (p : α → ℝ) (code : binaryprefixcode α) (a : α) (w : list bool) (hw : w ≠ []) (hsep : ∀ i, i ≠ a → (¬ code.encode i <+: w) ∧ (¬ w <+: code.encode i)) : expectedcodelength p (code.replaceword a w hw hsep) = expectedcodelength p code + p a * (w.length - (code.encode a).length : ℝ) theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_pruneDeepest","label":"expectedCodeLength_pruneDeepest","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_pruneDeepest","description":"theorem expectedCodeLength_pruneDeepest {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : expectedCodeLength p (code.pruneDeepest a w b hw ha hmax hmissing) = expectedCodeLength p code - p a","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-580425da5789","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5841,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:82"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_pruneDeepest {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : expectedCodeLength p (code.pruneDeepest a w b hw ha hmax hmissing) = expectedCodeLength p code - p a","missing":[],"search":"expectedcodelength_prunedeepest banditrlproof.lowerbounds.expectedcodelength_prunedeepest theorem expectedcodelength_prunedeepest {α : type*} [fintype α] [decidableeq α] (p : α → ℝ) (code : binaryprefixcode α) (a : α) (w : list bool) (b : bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : expectedcodelength p (code.prunedeepest a w b hw ha hmax hmissing) = expectedcodelength p code - p a theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_pruneDeepest_le","label":"expectedCodeLength_pruneDeepest_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_pruneDeepest_le","description":"Pruning remains cost-nonincreasing when the removed symbol has zero mass.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-7b05e4142840","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5842,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:95"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_pruneDeepest_le {α : Type*} [Fintype α] [DecidableEq α] (p : α → ℝ) (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (hp : 0 ≤ p a) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : expectedCodeLength p (code.pruneDeepest a w b hw ha hmax hmissing) ≤ expectedCodeLength p code","missing":[],"search":"expectedcodelength_prunedeepest_le banditrlproof.lowerbounds.expectedcodelength_prunedeepest_le pruning remains cost-nonincreasing when the removed symbol has zero mass. theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.totalCodeLength","label":"totalCodeLength","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.totalCodeLength","description":"Structural termination measure, independent of source probabilities.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-b3e3c30e248c","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5843,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:106"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def totalCodeLength {α : Type*} [Fintype α] (code : BinaryPrefixCode α) : ℕ","missing":[],"search":"totalcodelength banditrlproof.lowerbounds.totalcodelength structural termination measure, independent of source probabilities. definition compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.totalCodeLength_pruneDeepest","label":"totalCodeLength_pruneDeepest","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.totalCodeLength_pruneDeepest","description":"theorem totalCodeLength_pruneDeepest {α : Type*} [Fintype α] [DecidableEq α] (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : totalCodeLength (code.pruneDeepest a w b hw ha hmax hmissing) + 1 = totalCodeLength code","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-fed80886c2b3","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5844,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:109"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem totalCodeLength_pruneDeepest {α : Type*} [Fintype α] [DecidableEq α] (code : BinaryPrefixCode α) (a : α) (w : List Bool) (b : Bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : totalCodeLength (code.pruneDeepest a w b hw ha hmax hmissing) + 1 = totalCodeLength code","missing":[],"search":"totalcodelength_prunedeepest banditrlproof.lowerbounds.totalcodelength_prunedeepest theorem totalcodelength_prunedeepest {α : type*} [fintype α] [decidableeq α] (code : binaryprefixcode α) (a : α) (w : list bool) (b : bool) (hw : w ≠ []) (ha : code.encode a = w ++ [b]) (hmax : ∀ i, (code.encode i).length ≤ (code.encode a).length) (hmissing : ∀ i, code.encode i ≠ w ++ [!b]) : totalcodelength (code.prunedeepest a w b hw ha hmax hmissing) + 1 = totalcodelength code theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_minimal_totalCodeLength_competitor","label":"exists_minimal_totalCodeLength_competitor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_minimal_totalCodeLength_competitor","description":"Choose a structurally minimal no-worse competitor without assuming cost-minimizer existence.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-5a2e3a8bdcfb","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5845,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:123"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_minimal_totalCodeLength_competitor {α : Type*} [Fintype α] (p : α → ℝ) (original : BinaryPrefixCode α) : ∃ code : BinaryPrefixCode α, expectedCodeLength p code ≤ expectedCodeLength p original ∧ ∀ other : BinaryPrefixCode α, expectedCodeLength p other ≤ expectedCodeLength p original → totalCodeLength code ≤ totalCodeLength other","missing":[],"search":"exists_minimal_totalcodelength_competitor banditrlproof.lowerbounds.exists_minimal_totalcodelength_competitor choose a structurally minimal no-worse competitor without assuming cost-minimizer existence. theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_competitor_with_deepest_siblings","label":"exists_competitor_with_deepest_siblings","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_competitor_with_deepest_siblings","description":"Normalize any competitor so that each deepest leaf has its sibling present.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-e2a1cbbed3a9","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5846,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:141"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_competitor_with_deepest_siblings {α : Type*} [Fintype α] [DecidableEq α] [Nontrivial α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (original : BinaryPrefixCode α) : ∃ code : BinaryPrefixCode α, expectedCodeLength p code ≤ expectedCodeLength p original ∧ ∀ a w b, code.encode a = w ++ [b] → (∀ i, (code.encode i).length ≤ (code.encode a).length) → ∃ j, code.encode j = w ++ [!b]","missing":[],"search":"exists_competitor_with_deepest_siblings banditrlproof.lowerbounds.exists_competitor_with_deepest_siblings normalize any competitor so that each deepest leaf has its sibling present. theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_no_worse_deepest_sibling_pair","label":"exists_no_worse_deepest_sibling_pair","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_no_worse_deepest_sibling_pair","description":"Every competitor has a no-worse code containing a deepest sibling pair.","url":"../modules/banditrlproof-lowerbounds-prefixcodepruning/index.html#decl-02ad158c5bb8","parent":"module:BanditRLProof.LowerBounds.PrefixCodePruning","order":5847,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodePruning"],["Source","BanditRLProof/LowerBounds/PrefixCodePruning.lean:167"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_no_worse_deepest_sibling_pair {α : Type*} [Fintype α] [DecidableEq α] [Nontrivial α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (original : BinaryPrefixCode α) : ∃ code : BinaryPrefixCode α, expectedCodeLength p code ≤ expectedCodeLength p original ∧ ∃ a j w b, a ≠ j ∧ code.encode a = w ++ [b] ∧ code.encode j = w ++ [!b] ∧ ∀ i, (code.encode i).length ≤ (code.encode a).length","missing":[],"search":"exists_no_worse_deepest_sibling_pair banditrlproof.lowerbounds.exists_no_worse_deepest_sibling_pair every competitor has a no-worse code containing a deepest sibling pair. theorem compiled","shard":"modules/0ebf053c16ab922f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.extended_prefix_parent_eq","label":"extended_prefix_parent_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.extended_prefix_parent_eq","description":"theorem BinaryPrefixCode.extended_prefix_parent_eq {α : Type*} (code : BinaryPrefixCode α) (a b : α) (u v : List Bool) (h : code.encode a ++ u <+: code.encode b ++ v) : a = b","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-7fcb3b646faf","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5848,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:5"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem BinaryPrefixCode.extended_prefix_parent_eq {α : Type*} (code : BinaryPrefixCode α) (a b : α) (u v : List Bool) (h : code.encode a ++ u <+: code.encode b ++ v) : a = b","missing":[],"search":"extended_prefix_parent_eq banditrlproof.lowerbounds.binaryprefixcode.extended_prefix_parent_eq theorem binaryprefixcode.extended_prefix_parent_eq {α : type*} (code : binaryprefixcode α) (a b : α) (u v : list bool) (h : code.encode a ++ u <+: code.encode b ++ v) : a = b theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.siblingExpandedWord","label":"siblingExpandedWord","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.siblingExpandedWord","description":"def siblingExpandedWord {α : Type*} (code : BinaryPrefixCode (Option α)) : α ⊕ Bool → List Bool | .inl a => code.encode (some a) | .inr b => code.encode none ++ [b] theorem siblingExpandedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (Option α)) {a b : α ⊕ Bool} (h : siblingExpandedWord code a <+: siblingExpandedWord code b) : a = b","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-059553eaadf9","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5849,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def siblingExpandedWord {α : Type*} (code : BinaryPrefixCode (Option α)) : α ⊕ Bool → List Bool | .inl a => code.encode (some a) | .inr b => code.encode none ++ [b] theorem siblingExpandedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (Option α)) {a b : α ⊕ Bool} (h : siblingExpandedWord code a <+: siblingExpandedWord code b) : a = b","missing":[],"search":"siblingexpandedword banditrlproof.lowerbounds.siblingexpandedword def siblingexpandedword {α : type*} (code : binaryprefixcode (option α)) : α ⊕ bool → list bool | .inl a => code.encode (some a) | .inr b => code.encode none ++ [b] theorem siblingexpandedword_prefixfree {α : type*} (code : binaryprefixcode (option α)) {a b : α ⊕ bool} (h : siblingexpandedword code a <+: siblingexpandedword code b) : a = b definition compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.siblingExpandedWord_prefixFree","label":"siblingExpandedWord_prefixFree","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.siblingExpandedWord_prefixFree","description":"theorem siblingExpandedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (Option α)) {a b : α ⊕ Bool} (h : siblingExpandedWord code a <+: siblingExpandedWord code b) : a = b","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-4b3f93d46640","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5850,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:18"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem siblingExpandedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (Option α)) {a b : α ⊕ Bool} (h : siblingExpandedWord code a <+: siblingExpandedWord code b) : a = b","missing":[],"search":"siblingexpandedword_prefixfree banditrlproof.lowerbounds.siblingexpandedword_prefixfree theorem siblingexpandedword_prefixfree {α : type*} (code : binaryprefixcode (option α)) {a b : α ⊕ bool} (h : siblingexpandedword code a <+: siblingexpandedword code b) : a = b theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.expandSibling","label":"expandSibling","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.expandSibling","description":"Split a designated merged leaf into two siblings.","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-5b60ef19ca98","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5851,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:46"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def BinaryPrefixCode.expandSibling {α : Type*} (code : BinaryPrefixCode (Option α)) : BinaryPrefixCode (α ⊕ Bool) where","missing":[],"search":"expandsibling banditrlproof.lowerbounds.binaryprefixcode.expandsibling split a designated merged leaf into two siblings. definition compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_expandSibling","label":"expectedCodeLength_expandSibling","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_expandSibling","description":"Exact cost recurrence of splitting a merged symbol.","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-57cc7a9f16cd","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5852,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_expandSibling {α : Type*} [Fintype α] (code : BinaryPrefixCode (Option α)) (p : α → ℝ) (q r : ℝ) : expectedCodeLength (Sum.elim p (fun b => if b then r else q)) code.expandSibling = expectedCodeLength (fun a => a.elim (q + r) p) code + q + r","missing":[],"search":"expectedcodelength_expandsibling banditrlproof.lowerbounds.expectedcodelength_expandsibling exact cost recurrence of splitting a merged symbol. theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sibling_parent_not_prefix_other","label":"sibling_parent_not_prefix_other","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sibling_parent_not_prefix_other","description":"theorem sibling_parent_not_prefix_other {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) (a : α) : ¬ w <+: code.encode (.inl a)","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-c4f236a5f432","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5853,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sibling_parent_not_prefix_other {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) (a : α) : ¬ w <+: code.encode (.inl a)","missing":[],"search":"sibling_parent_not_prefix_other banditrlproof.lowerbounds.sibling_parent_not_prefix_other theorem sibling_parent_not_prefix_other {α : type*} (code : binaryprefixcode (α ⊕ bool)) (w : list bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) (a : α) : ¬ w <+: code.encode (.inl a) theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.siblingContractedWord","label":"siblingContractedWord","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.siblingContractedWord","description":"def siblingContractedWord {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) : Option α → List Bool | none => w | some a => code.encode (.inl a) theorem siblingContractedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) {a b : Option α} (h : siblingContractedWord code w a <+: siblingContractedWord code w b) : a = b","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-141970602341","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5854,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:87"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def siblingContractedWord {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) : Option α → List Bool | none => w | some a => code.encode (.inl a) theorem siblingContractedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) {a b : Option α} (h : siblingContractedWord code w a <+: siblingContractedWord code w b) : a = b","missing":[],"search":"siblingcontractedword banditrlproof.lowerbounds.siblingcontractedword def siblingcontractedword {α : type*} (code : binaryprefixcode (α ⊕ bool)) (w : list bool) : option α → list bool | none => w | some a => code.encode (.inl a) theorem siblingcontractedword_prefixfree {α : type*} (code : binaryprefixcode (α ⊕ bool)) (w : list bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) {a b : option α} (h : siblingcontractedword code w a <+: siblingcontractedword code w b) : a = b definition compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.siblingContractedWord_prefixFree","label":"siblingContractedWord_prefixFree","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.siblingContractedWord_prefixFree","description":"theorem siblingContractedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) {a b : Option α} (h : siblingContractedWord code w a <+: siblingContractedWord code w b) : a = b","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-9ef015904545","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5855,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:92"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem siblingContractedWord_prefixFree {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) {a b : Option α} (h : siblingContractedWord code w a <+: siblingContractedWord code w b) : a = b","missing":[],"search":"siblingcontractedword_prefixfree banditrlproof.lowerbounds.siblingcontractedword_prefixfree theorem siblingcontractedword_prefixfree {α : type*} (code : binaryprefixcode (α ⊕ bool)) (w : list bool) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) {a b : option α} (h : siblingcontractedword code w a <+: siblingcontractedword code w b) : a = b theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.contractSibling","label":"contractSibling","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.BinaryPrefixCode.contractSibling","description":"Merge two actual sibling leaves whose parent is nonempty.","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-029061fa13ab","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5856,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:113"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def BinaryPrefixCode.contractSibling {α : Type*} (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hw : w ≠ []) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) : BinaryPrefixCode (Option α) where","missing":[],"search":"contractsibling banditrlproof.lowerbounds.binaryprefixcode.contractsibling merge two actual sibling leaves whose parent is nonempty. definition compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_contractSibling","label":"expectedCodeLength_contractSibling","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.expectedCodeLength_contractSibling","description":"theorem expectedCodeLength_contractSibling {α : Type*} [Fintype α] (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hw : w ≠ []) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) (p : α → ℝ) (q r : ℝ) : expectedCodeLength (fun a => a.elim (q + r) p) (code.contractSibling w hw hs) + q + r = expectedCodeLength (Sum.elim p (fun b => if b then r else q)) code","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-a3f5b0860941","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5857,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:125"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem expectedCodeLength_contractSibling {α : Type*} [Fintype α] (code : BinaryPrefixCode (α ⊕ Bool)) (w : List Bool) (hw : w ≠ []) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) (p : α → ℝ) (q r : ℝ) : expectedCodeLength (fun a => a.elim (q + r) p) (code.contractSibling w hw hs) + q + r = expectedCodeLength (Sum.elim p (fun b => if b then r else q)) code","missing":[],"search":"expectedcodelength_contractsibling banditrlproof.lowerbounds.expectedcodelength_contractsibling theorem expectedcodelength_contractsibling {α : type*} [fintype α] (code : binaryprefixcode (α ⊕ bool)) (w : list bool) (hw : w ≠ []) (hs : ∀ b, code.encode (.inr b) = w ++ [b]) (p : α → ℝ) (q r : ℝ) : expectedcodelength (fun a => a.elim (q + r) p) (code.contractsibling w hw hs) + q + r = expectedcodelength (sum.elim p (fun b => if b then r else q)) code theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryRootPrefixCode","label":"binaryRootPrefixCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryRootPrefixCode","description":"The two-symbol root code, avoiding an invalid empty parent codeword.","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-ef73ee8e487b","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5858,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:136"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def binaryRootPrefixCode : BinaryPrefixCode Bool where","missing":[],"search":"binaryrootprefixcode banditrlproof.lowerbounds.binaryrootprefixcode the two-symbol root code, avoiding an invalid empty parent codeword. definition compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.binaryRootPrefixCode_optimal","label":"binaryRootPrefixCode_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.binaryRootPrefixCode_optimal","description":"theorem binaryRootPrefixCode_optimal (p : Bool → ℝ) (hp : ∀ b, 0 ≤ p b) (hs : ∑ b, p b = 1) : IsOptimalPrefixCode p binaryRootPrefixCode","url":"../modules/banditrlproof-lowerbounds-prefixcodesiblings/index.html#decl-9740e74e4df8","parent":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","order":5859,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.PrefixCodeSiblings"],["Source","BanditRLProof/LowerBounds/PrefixCodeSiblings.lean:144"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem binaryRootPrefixCode_optimal (p : Bool → ℝ) (hp : ∀ b, 0 ≤ p b) (hs : ∑ b, p b = 1) : IsOptimalPrefixCode p binaryRootPrefixCode","missing":[],"search":"binaryrootprefixcode_optimal banditrlproof.lowerbounds.binaryrootprefixcode_optimal theorem binaryrootprefixcode_optimal (p : bool → ℝ) (hp : ∀ b, 0 ≤ p b) (hs : ∑ b, p b = 1) : isoptimalprefixcode p binaryrootprefixcode theorem compiled","shard":"modules/e00e04f22275603a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_eq_lintegral_condExp","label":"relativeEntropy_trim_eq_lintegral_condExp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_trim_eq_lintegral_condExp","description":"Trimmed KL as the convex lower integral of the conditional RN density.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html#decl-c39c97be3f72","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","order":5860,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyFiltration"],["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean:14"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_trim_eq_lintegral_condExp {α : Type*} {m m₀ : MeasurableSpace α} (P Q : @Measure α m₀) [IsFiniteMeasure P] [IsFiniteMeasure Q] (hm : m ≤ m₀) (h : P ≪ Q) : @relativeEntropy α m (P.trim hm) (Q.trim hm) = ∫⁻ x, ENNReal.ofReal (InformationTheory.klFun (Q[fun y => (P.rnDeriv Q y).toReal | m] x)) ∂Q","missing":[],"search":"relativeentropy_trim_eq_lintegral_condexp banditrlproof.lowerbounds.relativeentropy_trim_eq_lintegral_condexp trimmed kl as the convex lower integral of the conditional rn density. theorem compiled","shard":"modules/a00cc609302d8339.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_map_eq_trim_of_absolutelyContinuous","label":"relativeEntropy_map_eq_trim_of_absolutelyContinuous","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_map_eq_trim_of_absolutelyContinuous","description":"A measurable observation has the same KL as restriction to its generated sigma-algebra.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html#decl-e614f9741c2e","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","order":5861,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyFiltration"],["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_map_eq_trim_of_absolutelyContinuous {α β : Type*} [mα : MeasurableSpace α] [mβ : MeasurableSpace β] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) (f : α → β) (hf : Measurable f) : relativeEntropy (P.map f) (Q.map f) = @relativeEntropy α (mβ.comap f) (P.trim hf.comap_le) (Q.trim hf.comap_le)","missing":[],"search":"relativeentropy_map_eq_trim_of_absolutelycontinuous banditrlproof.lowerbounds.relativeentropy_map_eq_trim_of_absolutelycontinuous a measurable observation has the same kl as restriction to its generated sigma-algebra. theorem compiled","shard":"modules/a00cc609302d8339.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_trim_of_density_measurable","label":"relativeEntropy_eq_iSup_trim_of_density_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_trim_of_density_measurable","description":"Resolving the RN density along a filtration recovers KL, including infinite KL.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html#decl-f1a57aa7516b","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","order":5862,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyFiltration"],["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean:53"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_eq_iSup_trim_of_density_measurable {α : Type*} [m₀ : MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) (F : Filtration ℕ m₀) (hDensity : StronglyMeasurable[⨆ n, F n] (fun x => (P.rnDeriv Q x).toReal)) : relativeEntropy P Q = ⨆ n, @relativeEntropy α (F n) (P.trim (F.le n)) (Q.trim (F.le n))","missing":[],"search":"relativeentropy_eq_isup_trim_of_density_measurable banditrlproof.lowerbounds.relativeentropy_eq_isup_trim_of_density_measurable resolving the rn density along a filtration recovers kl, including infinite kl. theorem compiled","shard":"modules/a00cc609302d8339.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.densityApproximationFiltration","label":"densityApproximationFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.densityApproximationFiltration","description":"Natural filtration of the finite-valued lower approximations to a density.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html#decl-b739f7e50a6c","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","order":5863,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.RelativeEntropyFiltration"],["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def densityApproximationFiltration {α : Type*} [m : MeasurableSpace α] (r : α → ENNReal) : Filtration ℕ m","missing":[],"search":"densityapproximationfiltration banditrlproof.lowerbounds.densityapproximationfiltration natural filtration of the finite-valued lower approximations to a density. definition compiled","shard":"modules/a00cc609302d8339.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.measurable_density_iSup_approximationFiltration","label":"measurable_density_iSup_approximationFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.measurable_density_iSup_approximationFiltration","description":"The limiting sigma-algebra of the approximation filtration resolves the density.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html#decl-0ec2e1095e84","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","order":5864,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyFiltration"],["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean:99"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_density_iSup_approximationFiltration {α : Type*} [m : MeasurableSpace α] (r : α → ENNReal) (hr : Measurable r) : Measurable[⨆ n, densityApproximationFiltration r n] r","missing":[],"search":"measurable_density_isup_approximationfiltration banditrlproof.lowerbounds.measurable_density_isup_approximationfiltration the limiting sigma-algebra of the approximation filtration resolves the density. theorem compiled","shard":"modules/a00cc609302d8339.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_densityApproximation_trim","label":"relativeEntropy_eq_iSup_densityApproximation_trim","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_densityApproximation_trim","description":"RN KL is recovered along a concretely constructed simple-approximation filtration.","url":"../modules/banditrlproof-lowerbounds-relativeentropyfiltration/index.html#decl-7d8ce68289c4","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","order":5865,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyFiltration"],["Source","BanditRLProof/LowerBounds/RelativeEntropyFiltration.lean:115"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_eq_iSup_densityApproximation_trim {α : Type*} [m : MeasurableSpace α] (P Q : Measure α) [IsFiniteMeasure P] [IsFiniteMeasure Q] (h : P ≪ Q) : relativeEntropy P Q = ⨆ n, @relativeEntropy α (densityApproximationFiltration (P.rnDeriv Q) n) (P.trim ((densityApproximationFiltration (P.rnDeriv Q)).le n)) (Q.trim ((densityApproximationFiltration (P.rnDeriv Q)).le n))","missing":[],"search":"relativeentropy_eq_isup_densityapproximation_trim banditrlproof.lowerbounds.relativeentropy_eq_isup_densityapproximation_trim rn kl is recovered along a concretely constructed simple-approximation filtration. theorem compiled","shard":"modules/a00cc609302d8339.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.relativeEntropy_triangle_counterexample","label":"relativeEntropy_triangle_counterexample","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.relativeEntropy_triangle_counterexample","description":"A finite-valued Gaussian counterexample to the triangle inequality.","url":"../modules/banditrlproof-lowerbounds-relativeentropynonmetric/index.html#decl-95f06a2c6905","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","order":5866,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyNonMetric"],["Source","BanditRLProof/LowerBounds/RelativeEntropyNonMetric.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem relativeEntropy_triangle_counterexample : relativeEntropy (gaussianReal 0 1) (gaussianReal 1 1) + relativeEntropy (gaussianReal 1 1) (gaussianReal 2 1) < relativeEntropy (gaussianReal 0 1) (gaussianReal 2 1)","missing":[],"search":"relativeentropy_triangle_counterexample banditrlproof.lowerbounds.relativeentropy_triangle_counterexample a finite-valued gaussian counterexample to the triangle inequality. theorem compiled","shard":"modules/d2261a7f5cc626f8.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_asymmetry","label":"bernoulliRelativeEntropy_asymmetry","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.bernoulliRelativeEntropy_asymmetry","description":"Reversing the Bernoulli comparison can change finite KL to infinite KL.","url":"../modules/banditrlproof-lowerbounds-relativeentropynonmetric/index.html#decl-cfd08db16876","parent":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","order":5867,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.RelativeEntropyNonMetric"],["Source","BanditRLProof/LowerBounds/RelativeEntropyNonMetric.lean:19"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem bernoulliRelativeEntropy_asymmetry : bernoulliRelativeEntropy 0 (1 / 2) ≠ bernoulliRelativeEntropy (1 / 2) 0","missing":[],"search":"bernoullirelativeentropy_asymmetry banditrlproof.lowerbounds.bernoullirelativeentropy_asymmetry reversing the bernoulli comparison can change finite kl to infinite kl. theorem compiled","shard":"modules/d2261a7f5cc626f8.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.shannonLength","label":"shannonLength","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.shannonLength","description":"Strict Shannon length, leaving Kraft slack on every positive-mass symbol. This definition is not a prefix-code constructor.","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-9f158e174f3d","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5868,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:9"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def shannonLength (p : ℝ) : ℕ","missing":[],"search":"shannonlength banditrlproof.lowerbounds.shannonlength strict shannon length, leaving kraft slack on every positive-mass symbol. this definition is not a prefix-code constructor. definition compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.shannonLength_pos","label":"shannonLength_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.shannonLength_pos","description":"theorem shannonLength_pos (p : ℝ) : 0 < shannonLength p","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-3148179ce943","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5869,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem shannonLength_pos (p : ℝ) : 0 < shannonLength p","missing":[],"search":"shannonlength_pos banditrlproof.lowerbounds.shannonlength_pos theorem shannonlength_pos (p : ℝ) : 0 < shannonlength p theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.shannonLength_kraft_weight_lt","label":"shannonLength_kraft_weight_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.shannonLength_kraft_weight_lt","description":"Each positive-mass symbol leaves strict Kraft slack.","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-9f96ffb72d75","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5870,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:15"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem shannonLength_kraft_weight_lt {p : ℝ} (hp : 0 < p) : (1 / 2 : ℝ) ^ shannonLength p < p","missing":[],"search":"shannonlength_kraft_weight_lt banditrlproof.lowerbounds.shannonlength_kraft_weight_lt each positive-mass symbol leaves strict kraft slack. theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.shannonLength_le_information_add_one","label":"shannonLength_le_information_add_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.shannonLength_le_information_add_one","description":"theorem shannonLength_le_information_add_one {p : ℝ} (hp : 0 < p) (hp1 : p ≤ 1) : (shannonLength p : ℝ) ≤ Real.log p⁻¹ / Real.log 2 + 1","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-bf1f996b4d13","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5871,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem shannonLength_le_information_add_one {p : ℝ} (hp : 0 < p) (hp1 : p ≤ 1) : (shannonLength p : ℝ) ≤ Real.log p⁻¹ / Real.log 2 + 1","missing":[],"search":"shannonlength_le_information_add_one banditrlproof.lowerbounds.shannonlength_le_information_add_one theorem shannonlength_le_information_add_one {p : ℝ} (hp : 0 < p) (hp1 : p ≤ 1) : (shannonlength p : ℝ) ≤ real.log p⁻¹ / real.log 2 + 1 theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.weighted_shannonLength_le","label":"weighted_shannonLength_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.weighted_shannonLength_le","description":"theorem weighted_shannonLength_le {p : ℝ} (hp : 0 ≤ p) (hp1 : p ≤ 1) : p * shannonLength p ≤ p * (Real.log p⁻¹ / Real.log 2) + p","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-cfcf0efab248","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5872,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem weighted_shannonLength_le {p : ℝ} (hp : 0 ≤ p) (hp1 : p ≤ 1) : p * shannonLength p ≤ p * (Real.log p⁻¹ / Real.log 2) + p","missing":[],"search":"weighted_shannonlength_le banditrlproof.lowerbounds.weighted_shannonlength_le theorem weighted_shannonlength_le {p : ℝ} (hp : 0 ≤ p) (hp1 : p ≤ 1) : p * shannonlength p ≤ p * (real.log p⁻¹ / real.log 2) + p theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_weighted_shannonLength_le_entropy_add_one","label":"sum_weighted_shannonLength_le_entropy_add_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_weighted_shannonLength_le_entropy_add_one","description":"A numerical expected-length bound; realizability by a prefix code is separate.","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-98ec9b328ac7","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5873,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:45"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_weighted_shannonLength_le_entropy_add_one {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : (∑ i, p i * shannonLength (p i)) ≤ discreteEntropyBaseTwo Finset.univ p + 1","missing":[],"search":"sum_weighted_shannonlength_le_entropy_add_one banditrlproof.lowerbounds.sum_weighted_shannonlength_le_entropy_add_one a numerical expected-length bound; realizability by a prefix code is separate. theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.sum_positive_shannon_weights_lt_one","label":"sum_positive_shannon_weights_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.sum_positive_shannon_weights_lt_one","description":"The positive support occupies strictly less than the available Kraft mass.","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-3562dffabe9a","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5874,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:57"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_positive_shannon_weights_lt_one {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : (∑ i, if 0 < p i then (1 / 2 : ℝ) ^ shannonLength (p i) else 0) < 1","missing":[],"search":"sum_positive_shannon_weights_lt_one banditrlproof.lowerbounds.sum_positive_shannon_weights_lt_one the positive support occupies strictly less than the available kraft mass. theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.exists_lengths_kraft_lt_one_entropy_bound","label":"exists_lengths_kraft_lt_one_entropy_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.exists_lengths_kraft_lt_one_entropy_bound","description":"Complete length assignment, including zero-mass symbols; codewords are not yet constructed.","url":"../modules/banditrlproof-lowerbounds-shannonlengths/index.html#decl-f5ec971b44ec","parent":"module:BanditRLProof.LowerBounds.ShannonLengths","order":5875,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.ShannonLengths"],["Source","BanditRLProof/LowerBounds/ShannonLengths.lean:76"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem exists_lengths_kraft_lt_one_entropy_bound {α : Type*} [Fintype α] (p : α → ℝ) (hp : ∀ i, 0 ≤ p i) (hs : ∑ i, p i = 1) : ∃ l : α → ℕ, (∀ i, 0 < l i) ∧ (∑ i, (1 / 2 : ℝ) ^ l i) < 1 ∧ (∑ i, p i * l i) ≤ discreteEntropyBaseTwo Finset.univ p + 1","missing":[],"search":"exists_lengths_kraft_lt_one_entropy_bound banditrlproof.lowerbounds.exists_lengths_kraft_lt_one_entropy_bound complete length assignment, including zero-mass symbols; codewords are not yet constructed. theorem compiled","shard":"modules/6911224bad3322b1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitSubgaussianBanditEnvironment","label":"UnitSubgaussianBanditEnvironment","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitSubgaussianBanditEnvironment","description":"The class in Chapter 13's main prose: gaps, not means, lie in [0,1].","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-a5530e89edc7","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5876,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:10"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure UnitSubgaussianBanditEnvironment (k : ℕ) where","missing":[],"search":"unitsubgaussianbanditenvironment banditrlproof.lowerbounds.unitsubgaussianbanditenvironment the class in chapter 13's main prose: gaps, not means, lie in [0,1]. structure compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment.toSubgaussian","label":"toSubgaussian","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment.toSubgaussian","description":"def UnitGaussianBanditEnvironment.toSubgaussian {k : ℕ} (e : UnitGaussianBanditEnvironment k) : UnitSubgaussianBanditEnvironment k where","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-4de02965ad85","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5877,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def UnitGaussianBanditEnvironment.toSubgaussian {k : ℕ} (e : UnitGaussianBanditEnvironment k) : UnitSubgaussianBanditEnvironment k where","missing":[],"search":"tosubgaussian banditrlproof.lowerbounds.unitgaussianbanditenvironment.tosubgaussian def unitgaussianbanditenvironment.tosubgaussian {k : ℕ} (e : unitgaussianbanditenvironment k) : unitsubgaussianbanditenvironment k where definition compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.subgaussianExpectedPseudoRegret","label":"subgaussianExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.subgaussianExpectedPseudoRegret","description":"def subgaussianExpectedPseudoRegret {k : ℕ} (algorithm : Thompson.HistoryAlgorithm (Fin k) ℝ) (e : UnitSubgaussianBanditEnvironment k) (t : ℕ) : ℝ≥0∞","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-8ec40c69faf3","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5878,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:38"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def subgaussianExpectedPseudoRegret {k : ℕ} (algorithm : Thompson.HistoryAlgorithm (Fin k) ℝ) (e : UnitSubgaussianBanditEnvironment k) (t : ℕ) : ℝ≥0∞","missing":[],"search":"subgaussianexpectedpseudoregret banditrlproof.lowerbounds.subgaussianexpectedpseudoregret def subgaussianexpectedpseudoregret {k : ℕ} (algorithm : thompson.historyalgorithm (fin k) ℝ) (e : unitsubgaussianbanditenvironment k) (t : ℕ) : ℝ≥0∞ definition compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.subgaussianWorstCaseExpectedPseudoRegret","label":"subgaussianWorstCaseExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.subgaussianWorstCaseExpectedPseudoRegret","description":"def subgaussianWorstCaseExpectedPseudoRegret (k : ℕ) (algorithm : Thompson.HistoryAlgorithm (Fin k) ℝ) (t : ℕ) : ℝ≥0∞","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-48e3ee9b0ab1","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5879,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:43"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def subgaussianWorstCaseExpectedPseudoRegret (k : ℕ) (algorithm : Thompson.HistoryAlgorithm (Fin k) ℝ) (t : ℕ) : ℝ≥0∞","missing":[],"search":"subgaussianworstcaseexpectedpseudoregret banditrlproof.lowerbounds.subgaussianworstcaseexpectedpseudoregret def subgaussianworstcaseexpectedpseudoregret (k : ℕ) (algorithm : thompson.historyalgorithm (fin k) ℝ) (t : ℕ) : ℝ≥0∞ definition compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.subgaussianMinimaxExpectedPseudoRegret","label":"subgaussianMinimaxExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.subgaussianMinimaxExpectedPseudoRegret","description":"def subgaussianMinimaxExpectedPseudoRegret (k t : ℕ) : ℝ≥0∞","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-1331f26d7318","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5880,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:47"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def subgaussianMinimaxExpectedPseudoRegret (k t : ℕ) : ℝ≥0∞","missing":[],"search":"subgaussianminimaxexpectedpseudoregret banditrlproof.lowerbounds.subgaussianminimaxexpectedpseudoregret def subgaussianminimaxexpectedpseudoregret (k t : ℕ) : ℝ≥0∞ definition compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.subgaussianExpectedPseudoRegret_gaussian","label":"subgaussianExpectedPseudoRegret_gaussian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.subgaussianExpectedPseudoRegret_gaussian","description":"theorem subgaussianExpectedPseudoRegret_gaussian {k : ℕ} (algorithm : Thompson.HistoryAlgorithm (Fin k) ℝ) (e : UnitGaussianBanditEnvironment k) (t : ℕ) : subgaussianExpectedPseudoRegret algorithm e.toSubgaussian t = gaussianExpectedPseudoRegret algorithm e t","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-f018dd284238","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5881,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:51"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem subgaussianExpectedPseudoRegret_gaussian {k : ℕ} (algorithm : Thompson.HistoryAlgorithm (Fin k) ℝ) (e : UnitGaussianBanditEnvironment k) (t : ℕ) : subgaussianExpectedPseudoRegret algorithm e.toSubgaussian t = gaussianExpectedPseudoRegret algorithm e t","missing":[],"search":"subgaussianexpectedpseudoregret_gaussian banditrlproof.lowerbounds.subgaussianexpectedpseudoregret_gaussian theorem subgaussianexpectedpseudoregret_gaussian {k : ℕ} (algorithm : thompson.historyalgorithm (fin k) ℝ) (e : unitgaussianbanditenvironment k) (t : ℕ) : subgaussianexpectedpseudoregret algorithm e.tosubgaussian t = gaussianexpectedpseudoregret algorithm e t theorem compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimax_le_subgaussianMinimax","label":"unitGaussianMinimax_le_subgaussianMinimax","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.unitGaussianMinimax_le_subgaussianMinimax","description":"Gaussian subclass inclusion is on the identical policy and regret functional.","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-7de340da9ff4","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5882,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:58"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem unitGaussianMinimax_le_subgaussianMinimax (k t : ℕ) : unitGaussianMinimaxExpectedPseudoRegret k t ≤ subgaussianMinimaxExpectedPseudoRegret k t","missing":[],"search":"unitgaussianminimax_le_subgaussianminimax banditrlproof.lowerbounds.unitgaussianminimax_le_subgaussianminimax gaussian subclass inclusion is on the identical policy and regret functional. theorem compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.moss_subgaussianExpectedPseudoRegret_le","label":"moss_subgaussianExpectedPseudoRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.moss_subgaussianExpectedPseudoRegret_le","description":"theorem moss_subgaussianExpectedPseudoRegret_le {k : ℕ} [NeZero k] (hk : 0 < k) (t : ℕ) (hkt : k ≤ t+1) (e : UnitSubgaussianBanditEnvironment k) : subgaussianExpectedPseudoRegret (MOSS.historyAlgorithm hk (t+1)) e t ≤ ENNReal.ofReal (40*Real.sqrt ((k : ℝ)*(t+1)))","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-f1a93d3ac318","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5883,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:67"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem moss_subgaussianExpectedPseudoRegret_le {k : ℕ} [NeZero k] (hk : 0 < k) (t : ℕ) (hkt : k ≤ t+1) (e : UnitSubgaussianBanditEnvironment k) : subgaussianExpectedPseudoRegret (MOSS.historyAlgorithm hk (t+1)) e t ≤ ENNReal.ofReal (40*Real.sqrt ((k : ℝ)*(t+1)))","missing":[],"search":"moss_subgaussianexpectedpseudoregret_le banditrlproof.lowerbounds.moss_subgaussianexpectedpseudoregret_le theorem moss_subgaussianexpectedpseudoregret_le {k : ℕ} [nezero k] (hk : 0 < k) (t : ℕ) (hkt : k ≤ t+1) (e : unitsubgaussianbanditenvironment k) : subgaussianexpectedpseudoregret (moss.historyalgorithm hk (t+1)) e t ≤ ennreal.ofreal (40*real.sqrt ((k : ℝ)*(t+1))) theorem compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.subgaussianMinimax_sandwich","label":"subgaussianMinimax_sandwich","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.subgaussianMinimax_sandwich","description":"Chapter 13's broader-class minimax sandwich with universal constants.","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-4cfeeef1fdc0","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5884,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:92"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem subgaussianMinimax_sandwich {k : ℕ} [NeZero k] (hk : 1 < k) (t : ℕ) (hkt : k ≤ t+1) : ENNReal.ofReal ((1/54 : ℝ)*Real.sqrt ((k : ℝ)*(t+1))) ≤ subgaussianMinimaxExpectedPseudoRegret k t ∧ subgaussianMinimaxExpectedPseudoRegret k t ≤ subgaussianWorstCaseExpectedPseudoRegret k (MOSS.historyAlgorithm (by omega) (t+1)) t ∧ subgaussianWorstCaseExpectedPseudoRegret k (MOSS.historyAlgorithm (by omega) (t+1)) t ≤ ENNReal.ofReal (40*Real.sqrt ((k : ℝ)*(t+1)))","missing":[],"search":"subgaussianminimax_sandwich banditrlproof.lowerbounds.subgaussianminimax_sandwich chapter 13's broader-class minimax sandwich with universal constants. theorem compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.moss_nearMinimax","label":"moss_nearMinimax","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.moss_nearMinimax","description":"Algorithm 7 is minimax optimal up to one universal multiplicative factor.","url":"../modules/banditrlproof-lowerbounds-subgaussianminimax/index.html#decl-b657879f704f","parent":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","order":5885,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SubgaussianMinimax"],["Source","BanditRLProof/LowerBounds/SubgaussianMinimax.lean:107"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-13-basic-ideas"],["Indexed settings","None registered"]],"statement":"theorem moss_nearMinimax {k : ℕ} [NeZero k] (hk : 1 < k) (t : ℕ) (hkt : k ≤ t+1) : subgaussianWorstCaseExpectedPseudoRegret k (MOSS.historyAlgorithm (by omega) (t+1)) t ≤ 2160 * subgaussianMinimaxExpectedPseudoRegret k t","missing":[],"search":"moss_nearminimax banditrlproof.lowerbounds.moss_nearminimax algorithm 7 is minimax optimal up to one universal multiplicative factor. theorem compiled","shard":"modules/a86ee01fe165bae1.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-13-basic-ideas"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem","label":"SuccinctUnitSystem","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem","description":"Source Axiom 3.1: a nonempty collection of unit atoms closed under negation. No spanning assumption is added.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-bb712a25313a","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5886,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:33"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure SuccinctUnitSystem (V : Type*) [NormedAddCommGroup V] [InnerProductSpace ℝ V] where","missing":[],"search":"succinctunitsystem banditrlproof.lowerbounds.succinct.succinctunitsystem source axiom 3.1: a nonempty collection of unit atoms closed under negation. no spanning assumption is added. structure compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ","label":"sourceQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ","description":"The source quantity `Q(X)=sup_{E in U} <X,E>`.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-38e8c8ea9d99","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5887,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:45"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def sourceQ (system : SuccinctUnitSystem V) (x : V) : ℝ","missing":[],"search":"sourceq banditrlproof.lowerbounds.succinct.succinctunitsystem.sourceq the source quantity `q(x)=sup_{e in u} <x,e>`. definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQSet_bddAbove","label":"sourceQSet_bddAbove","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQSet_bddAbove","description":"theorem sourceQSet_bddAbove (system : SuccinctUnitSystem V) (x : V) : BddAbove ((fun e : V => ⟪x, e⟫_ℝ) '' system.atoms)","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-44d7cde95d8c","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5888,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:48"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQSet_bddAbove (system : SuccinctUnitSystem V) (x : V) : BddAbove ((fun e : V => ⟪x, e⟫_ℝ) '' system.atoms)","missing":[],"search":"sourceqset_bddabove banditrlproof.lowerbounds.succinct.succinctunitsystem.sourceqset_bddabove theorem sourceqset_bddabove (system : succinctunitsystem v) (x : v) : bddabove ((fun e : v => ⟪x, e⟫_ℝ) '' system.atoms) theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.le_sourceQ_of_mem","label":"le_sourceQ_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.le_sourceQ_of_mem","description":"theorem le_sourceQ_of_mem (system : SuccinctUnitSystem V) {x e : V} (he : e ∈ system.atoms) : ⟪x, e⟫_ℝ ≤ system.sourceQ x","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-032cfb5b01b0","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5889,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:57"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem le_sourceQ_of_mem (system : SuccinctUnitSystem V) {x e : V} (he : e ∈ system.atoms) : ⟪x, e⟫_ℝ ≤ system.sourceQ x","missing":[],"search":"le_sourceq_of_mem banditrlproof.lowerbounds.succinct.succinctunitsystem.le_sourceq_of_mem theorem le_sourceq_of_mem (system : succinctunitsystem v) {x e : v} (he : e ∈ system.atoms) : ⟪x, e⟫_ℝ ≤ system.sourceq x theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_le_norm","label":"sourceQ_le_norm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_le_norm","description":"theorem sourceQ_le_norm (system : SuccinctUnitSystem V) (x : V) : system.sourceQ x ≤ ‖x‖","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-8b730499e417","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5890,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:62"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_le_norm (system : SuccinctUnitSystem V) (x : V) : system.sourceQ x ≤ ‖x‖","missing":[],"search":"sourceq_le_norm banditrlproof.lowerbounds.succinct.succinctunitsystem.sourceq_le_norm theorem sourceq_le_norm (system : succinctunitsystem v) (x : v) : system.sourceq x ≤ ‖x‖ theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_nonneg","label":"sourceQ_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_nonneg","description":"theorem sourceQ_nonneg (system : SuccinctUnitSystem V) (x : V) : 0 ≤ system.sourceQ x","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-3a55acceac77","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5891,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:72"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_nonneg (system : SuccinctUnitSystem V) (x : V) : 0 ≤ system.sourceQ x","missing":[],"search":"sourceq_nonneg banditrlproof.lowerbounds.succinct.succinctunitsystem.sourceq_nonneg theorem sourceq_nonneg (system : succinctunitsystem v) (x : v) : 0 ≤ system.sourceq x theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_zero","label":"sourceQ_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_zero","description":"theorem sourceQ_zero (system : SuccinctUnitSystem V) : system.sourceQ 0 = 0","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-53b9725770cc","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5892,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:83"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_zero (system : SuccinctUnitSystem V) : system.sourceQ 0 = 0","missing":[],"search":"sourceq_zero banditrlproof.lowerbounds.succinct.succinctunitsystem.sourceq_zero theorem sourceq_zero (system : succinctunitsystem v) : system.sourceq 0 = 0 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.abs_inner_le_sourceQ_of_mem","label":"abs_inner_le_sourceQ_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.abs_inner_le_sourceQ_of_mem","description":"theorem abs_inner_le_sourceQ_of_mem (system : SuccinctUnitSystem V) {x e : V} (he : e ∈ system.atoms) : |⟪x, e⟫_ℝ| ≤ system.sourceQ x","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-526741d76fd4","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5893,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:88"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_inner_le_sourceQ_of_mem (system : SuccinctUnitSystem V) {x e : V} (he : e ∈ system.atoms) : |⟪x, e⟫_ℝ| ≤ system.sourceQ x","missing":[],"search":"abs_inner_le_sourceq_of_mem banditrlproof.lowerbounds.succinct.succinctunitsystem.abs_inner_le_sourceq_of_mem theorem abs_inner_le_sourceq_of_mem (system : succinctunitsystem v) {x e : v} (he : e ∈ system.atoms) : |⟪x, e⟫_ℝ| ≤ system.sourceq x theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_eq_zero_of_atom_orthogonal","label":"sourceQ_eq_zero_of_atom_orthogonal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_eq_zero_of_atom_orthogonal","description":"theorem sourceQ_eq_zero_of_atom_orthogonal (system : SuccinctUnitSystem V) {x : V} (horthogonal : ∀ e ∈ system.atoms, ⟪x, e⟫_ℝ = 0) : system.sourceQ x = 0","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-18bb4a41ad8b","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5894,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:98"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_eq_zero_of_atom_orthogonal (system : SuccinctUnitSystem V) {x : V} (horthogonal : ∀ e ∈ system.atoms, ⟪x, e⟫_ℝ = 0) : system.sourceQ x = 0","missing":[],"search":"sourceq_eq_zero_of_atom_orthogonal banditrlproof.lowerbounds.succinct.succinctunitsystem.sourceq_eq_zero_of_atom_orthogonal theorem sourceq_eq_zero_of_atom_orthogonal (system : succinctunitsystem v) {x : v} (horthogonal : ∀ e ∈ system.atoms, ⟪x, e⟫_ℝ = 0) : system.sourceq x = 0 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceR","label":"sourceR","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceR","description":"The source quantity `R(X)=sup_{Q(Y)<=1} <X,Y>`, retained as a real-valued `sSup`. Consumers must separately prove the defining set is bounded.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-38136d758ac1","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5895,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:110"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def sourceR (system : SuccinctUnitSystem V) (x : V) : ℝ","missing":[],"search":"sourcer banditrlproof.lowerbounds.succinct.succinctunitsystem.sourcer the source quantity `r(x)=sup_{q(y)<=1} <x,y>`, retained as a real-valued `ssup`. consumers must separately prove the defining set is bounded. definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceRSet_not_bddAbove_of_nonzero_atom_orthogonal","label":"sourceRSet_not_bddAbove_of_nonzero_atom_orthogonal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceRSet_not_bddAbove_of_nonzero_atom_orthogonal","description":"If a nonzero ambient direction is orthogonal to every atom, the set used to define its source `R` is unbounded. This exposes the hidden global regularity/codomain obligation in Definition 3.2.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-c19b11b37671","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5896,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:116"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceRSet_not_bddAbove_of_nonzero_atom_orthogonal (system : SuccinctUnitSystem V) {x : V} (hx : x ≠ 0) (horthogonal : ∀ e ∈ system.atoms, ⟪x, e⟫_ℝ = 0) : ¬ BddAbove ((fun y : V => ⟪x, y⟫_ℝ) '' {y | system.sourceQ y ≤ 1})","missing":[],"search":"sourcerset_not_bddabove_of_nonzero_atom_orthogonal banditrlproof.lowerbounds.succinct.succinctunitsystem.sourcerset_not_bddabove_of_nonzero_atom_orthogonal if a nonzero ambient direction is orthogonal to every atom, the set used to define its source `r` is unbounded. this exposes the hidden global regularity/codomain obligation in definition 3.2. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport","label":"IsSuccinctSupport","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport","description":"Source Definition 3.1. The explicit boundedness field is the ordinary mathematical meaning of the displayed finite real supremum, not a replacement for the source equality.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-640f4434b380","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5897,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:142"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure IsSuccinctSupport (system : SuccinctUnitSystem V) {s : Nat} (basis : Fin s → V) : Prop where","missing":[],"search":"issuccinctsupport banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport source definition 3.1. the explicit boundedness field is the ordinary mathematical meaning of the displayed finite real supremum, not a replacement for the source equality. structure compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.correlationSum_le_one","label":"correlationSum_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.correlationSum_le_one","description":"theorem correlationSum_le_one (support : IsSuccinctSupport system basis) {e : V} (he : e ∈ system.atoms) : (∑ i, |⟪e, basis i⟫_ℝ|) ≤ 1","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-8a9c23280ae2","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5898,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:154"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem correlationSum_le_one (support : IsSuccinctSupport system basis) {e : V} (he : e ∈ system.atoms) : (∑ i, |⟪e, basis i⟫_ℝ|) ≤ 1","missing":[],"search":"correlationsum_le_one banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.correlationsum_le_one theorem correlationsum_le_one (support : issuccinctsupport system basis) {e : v} (he : e ∈ system.atoms) : (∑ i, |⟪e, basis i⟫_ℝ|) ≤ 1 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_basis_basis","label":"inner_basis_basis","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_basis_basis","description":"The mutual-orthogonality consequence stated after source Definition 3.1.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-ea48fb9d518d","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5899,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:164"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_basis_basis (support : IsSuccinctSupport system basis) (i j : Fin s) : ⟪basis i, basis j⟫_ℝ = if i = j then 1 else 0","missing":[],"search":"inner_basis_basis banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.inner_basis_basis the mutual-orthogonality consequence stated after source definition 3.1. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.orthonormal","label":"orthonormal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.orthonormal","description":"A source succinct support is an orthonormal family. This is the Mathlib interface consumed by the finite Bessel step in source Lemma 3.3.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-afc626dac822","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5900,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:202"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem orthonormal (support : IsSuccinctSupport system basis) : Orthonormal ℝ basis","missing":[],"search":"orthonormal banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.orthonormal a source succinct support is an orthonormal family. this is the mathlib interface consumed by the finite bessel step in source lemma 3.3. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient","label":"maxAbsCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient","description":"The finite maximum appearing in source Lemma 3.1.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-81d6d99d485c","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5901,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:208"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def maxAbsCoefficient [Nonempty (Fin s)] (a : Fin s → ℝ) : ℝ","missing":[],"search":"maxabscoefficient banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.maxabscoefficient the finite maximum appearing in source lemma 3.1. definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_le_maxAbsCoefficient","label":"abs_le_maxAbsCoefficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_le_maxAbsCoefficient","description":"theorem abs_le_maxAbsCoefficient [Nonempty (Fin s)] (a : Fin s → ℝ) (i : Fin s) : |a i| ≤ maxAbsCoefficient a","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-eedfb19f377e","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5902,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:211"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_le_maxAbsCoefficient [Nonempty (Fin s)] (a : Fin s → ℝ) (i : Fin s) : |a i| ≤ maxAbsCoefficient a","missing":[],"search":"abs_le_maxabscoefficient banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.abs_le_maxabscoefficient theorem abs_le_maxabscoefficient [nonempty (fin s)] (a : fin s → ℝ) (i : fin s) : |a i| ≤ maxabscoefficient a theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_nonneg","label":"maxAbsCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_nonneg","description":"theorem maxAbsCoefficient_nonneg [Nonempty (Fin s)] (a : Fin s → ℝ) : 0 ≤ maxAbsCoefficient a","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-80b984a466d8","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5903,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:215"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem maxAbsCoefficient_nonneg [Nonempty (Fin s)] (a : Fin s → ℝ) : 0 ≤ maxAbsCoefficient a","missing":[],"search":"maxabscoefficient_nonneg banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.maxabscoefficient_nonneg theorem maxabscoefficient_nonneg [nonempty (fin s)] (a : fin s → ℝ) : 0 ≤ maxabscoefficient a theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.exists_abs_eq_maxAbsCoefficient","label":"exists_abs_eq_maxAbsCoefficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.exists_abs_eq_maxAbsCoefficient","description":"theorem exists_abs_eq_maxAbsCoefficient [Nonempty (Fin s)] (a : Fin s → ℝ) : ∃ i : Fin s, |a i| = maxAbsCoefficient a","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-6675070460bc","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5904,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:220"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem exists_abs_eq_maxAbsCoefficient [Nonempty (Fin s)] (a : Fin s → ℝ) : ∃ i : Fin s, |a i| = maxAbsCoefficient a","missing":[],"search":"exists_abs_eq_maxabscoefficient banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.exists_abs_eq_maxabscoefficient theorem exists_abs_eq_maxabscoefficient [nonempty (fin s)] (a : fin s → ℝ) : ∃ i : fin s, |a i| = maxabscoefficient a theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportCombination","label":"supportCombination","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportCombination","description":"def supportCombination (basis : Fin s → V) (a : Fin s → ℝ) : V","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-e34530d7d7db","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5905,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:227"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def supportCombination (basis : Fin s → V) (a : Fin s → ℝ) : V","missing":[],"search":"supportcombination banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.supportcombination def supportcombination (basis : fin s → v) (a : fin s → ℝ) : v definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom","label":"signedSupportAtom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom","description":"def signedSupportAtom (coefficient : ℝ) (atom : V) : V","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-ce9c3bcae8ca","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5906,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:230"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def signedSupportAtom (coefficient : ℝ) (atom : V) : V","missing":[],"search":"signedsupportatom banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.signedsupportatom def signedsupportatom (coefficient : ℝ) (atom : v) : v definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom_mem","label":"signedSupportAtom_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom_mem","description":"theorem signedSupportAtom_mem (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (i : Fin s) : signedSupportAtom (a i) (basis i) ∈ system.atoms","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-ad118e08f268","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5907,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:233"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem signedSupportAtom_mem (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (i : Fin s) : signedSupportAtom (a i) (basis i) ∈ system.atoms","missing":[],"search":"signedsupportatom_mem banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.signedsupportatom_mem theorem signedsupportatom_mem (support : issuccinctsupport system basis) (a : fin s → ℝ) (i : fin s) : signedsupportatom (a i) (basis i) ∈ system.atoms theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_le","label":"sourceQ_supportCombination_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_le","description":"theorem sourceQ_supportCombination_le [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : system.sourceQ (supportCombination basis a) ≤ maxAbsCoefficient a","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-daeed4f6ee9c","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5908,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:240"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_supportCombination_le [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : system.sourceQ (supportCombination basis a) ≤ maxAbsCoefficient a","missing":[],"search":"sourceq_supportcombination_le banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.sourceq_supportcombination_le theorem sourceq_supportcombination_le [nonempty (fin s)] (support : issuccinctsupport system basis) (a : fin s → ℝ) : system.sourceq (supportcombination basis a) ≤ maxabscoefficient a theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_basis","label":"inner_supportCombination_basis","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_basis","description":"theorem inner_supportCombination_basis (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (m : Fin s) : ⟪supportCombination basis a, basis m⟫_ℝ = a m","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-dcf65ecd8d29","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5909,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:268"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_supportCombination_basis (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (m : Fin s) : ⟪supportCombination basis a, basis m⟫_ℝ = a m","missing":[],"search":"inner_supportcombination_basis banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.inner_supportcombination_basis theorem inner_supportcombination_basis (support : issuccinctsupport system basis) (a : fin s → ℝ) (m : fin s) : ⟪supportcombination basis a, basis m⟫_ℝ = a m theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_signedSupportAtom","label":"inner_supportCombination_signedSupportAtom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_signedSupportAtom","description":"theorem inner_supportCombination_signedSupportAtom (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (m : Fin s) : ⟪supportCombination basis a, signedSupportAtom (a m) (basis m)⟫_ℝ = |a m|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-680f30a2bd52","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5910,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:281"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_supportCombination_signedSupportAtom (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (m : Fin s) : ⟪supportCombination basis a, signedSupportAtom (a m) (basis m)⟫_ℝ = |a m|","missing":[],"search":"inner_supportcombination_signedsupportatom banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.inner_supportcombination_signedsupportatom theorem inner_supportcombination_signedsupportatom (support : issuccinctsupport system basis) (a : fin s → ℝ) (m : fin s) : ⟪supportcombination basis a, signedsupportatom (a m) (basis m)⟫_ℝ = |a m| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","label":"sourceQ_supportCombination_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","description":"Source Lemma 3.1: `Q` is the coefficient maximum on a succinct support.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-5a766a24832a","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5911,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:292"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_supportCombination_eq [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : system.sourceQ (supportCombination basis a) = maxAbsCoefficient a","missing":[],"search":"sourceq_supportcombination_eq banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.sourceq_supportcombination_eq source lemma 3.1: `q` is the coefficient maximum on a succinct support. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign","label":"coefficientSign","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign","description":"def coefficientSign (coefficient : ℝ) : ℝ","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-e7928bcd16e7","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5912,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:305"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def coefficientSign (coefficient : ℝ) : ℝ","missing":[],"search":"coefficientsign banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.coefficientsign def coefficientsign (coefficient : ℝ) : ℝ definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_coefficientSign","label":"abs_coefficientSign","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_coefficientSign","description":"theorem abs_coefficientSign (coefficient : ℝ) : |coefficientSign coefficient| = 1","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-350e84eeda96","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5913,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:309"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_coefficientSign (coefficient : ℝ) : |coefficientSign coefficient| = 1","missing":[],"search":"abs_coefficientsign banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.abs_coefficientsign theorem abs_coefficientsign (coefficient : ℝ) : |coefficientsign coefficient| = 1 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul","label":"coefficientSign_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul","description":"theorem coefficientSign_mul (coefficient : ℝ) : coefficientSign coefficient * coefficient = |coefficient|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-ceb735ecc604","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5914,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:312"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem coefficientSign_mul (coefficient : ℝ) : coefficientSign coefficient * coefficient = |coefficient|","missing":[],"search":"coefficientsign_mul banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.coefficientsign_mul theorem coefficientsign_mul (coefficient : ℝ) : coefficientsign coefficient * coefficient = |coefficient| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_coefficientSign","label":"coefficientSign_coefficientSign","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_coefficientSign","description":"theorem coefficientSign_coefficientSign (coefficient : ℝ) : coefficientSign (coefficientSign coefficient) = coefficientSign coefficient","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-df2587052778","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5915,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:320"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem coefficientSign_coefficientSign (coefficient : ℝ) : coefficientSign (coefficientSign coefficient) = coefficientSign coefficient","missing":[],"search":"coefficientsign_coefficientsign banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.coefficientsign_coefficientsign theorem coefficientsign_coefficientsign (coefficient : ℝ) : coefficientsign (coefficientsign coefficient) = coefficientsign coefficient theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul_self","label":"coefficientSign_mul_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul_self","description":"theorem coefficientSign_mul_self (coefficient : ℝ) : coefficientSign coefficient * coefficientSign coefficient = 1","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-f58d852f536f","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5916,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:327"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem coefficientSign_mul_self (coefficient : ℝ) : coefficientSign coefficient * coefficientSign coefficient = 1","missing":[],"search":"coefficientsign_mul_self banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.coefficientsign_mul_self theorem coefficientsign_mul_self (coefficient : ℝ) : coefficientsign coefficient * coefficientsign coefficient = 1 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportSignCombination","label":"supportSignCombination","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportSignCombination","description":"def supportSignCombination (basis : Fin s → V) (a : Fin s → ℝ) : V","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-8d28a052a66b","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5917,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:331"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def supportSignCombination (basis : Fin s → V) (a : Fin s → ℝ) : V","missing":[],"search":"supportsigncombination banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.supportsigncombination def supportsigncombination (basis : fin s → v) (a : fin s → ℝ) : v definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_coefficientSign","label":"maxAbsCoefficient_coefficientSign","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_coefficientSign","description":"theorem maxAbsCoefficient_coefficientSign [Nonempty (Fin s)] (a : Fin s → ℝ) : maxAbsCoefficient (fun i => coefficientSign (a i)) = 1","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-b085461df573","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5918,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:334"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem maxAbsCoefficient_coefficientSign [Nonempty (Fin s)] (a : Fin s → ℝ) : maxAbsCoefficient (fun i => coefficientSign (a i)) = 1","missing":[],"search":"maxabscoefficient_coefficientsign banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.maxabscoefficient_coefficientsign theorem maxabscoefficient_coefficientsign [nonempty (fin s)] (a : fin s → ℝ) : maxabscoefficient (fun i => coefficientsign (a i)) = 1 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportSignCombination","label":"sourceQ_supportSignCombination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportSignCombination","description":"theorem sourceQ_supportSignCombination [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : system.sourceQ (supportSignCombination basis a) = 1","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-4a2152967941","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5919,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:338"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_supportSignCombination [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : system.sourceQ (supportSignCombination basis a) = 1","missing":[],"search":"sourceq_supportsigncombination banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.sourceq_supportsigncombination theorem sourceq_supportsigncombination [nonempty (fin s)] (support : issuccinctsupport system basis) (a : fin s → ℝ) : system.sourceq (supportsigncombination basis a) = 1 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.norm_sq_supportSignCombination","label":"norm_sq_supportSignCombination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.norm_sq_supportSignCombination","description":"The sign sum used in Appendix A.3 has squared norm equal to the support size.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-2cf345cd61dc","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5920,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:346"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem norm_sq_supportSignCombination [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : ‖supportSignCombination basis a‖ ^ 2 = (s : ℝ)","missing":[],"search":"norm_sq_supportsigncombination banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.norm_sq_supportsigncombination the sign sum used in appendix a.3 has squared norm equal to the support size. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_supportSignCombination","label":"inner_supportCombination_supportSignCombination","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_supportSignCombination","description":"theorem inner_supportCombination_supportSignCombination (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : ⟪supportCombination basis a, supportSignCombination basis a⟫_ℝ = ∑ i, |a i|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-40152037d2bb","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5921,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:359"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_supportCombination_supportSignCombination (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : ⟪supportCombination basis a, supportSignCombination basis a⟫_ℝ = ∑ i, |a i|","missing":[],"search":"inner_supportcombination_supportsigncombination banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.inner_supportcombination_supportsigncombination theorem inner_supportcombination_supportsigncombination (support : issuccinctsupport system basis) (a : fin s → ℝ) : ⟪supportcombination basis a, supportsigncombination basis a⟫_ℝ = ∑ i, |a i| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_mul_sourceQ","label":"inner_supportCombination_le_sumAbs_mul_sourceQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_mul_sourceQ","description":"theorem inner_supportCombination_le_sumAbs_mul_sourceQ (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (y : V) : |⟪supportCombination basis a, y⟫_ℝ| ≤ (∑ i, |a i|) * system.sourceQ y","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-5b21881b6f7b","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5922,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:378"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_supportCombination_le_sumAbs_mul_sourceQ (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) (y : V) : |⟪supportCombination basis a, y⟫_ℝ| ≤ (∑ i, |a i|) * system.sourceQ y","missing":[],"search":"inner_supportcombination_le_sumabs_mul_sourceq banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.inner_supportcombination_le_sumabs_mul_sourceq theorem inner_supportcombination_le_sumabs_mul_sourceq (support : issuccinctsupport system basis) (a : fin s → ℝ) (y : v) : |⟪supportcombination basis a, y⟫_ℝ| ≤ (∑ i, |a i|) * system.sourceq y theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_of_sourceQ_le_one","label":"inner_supportCombination_le_sumAbs_of_sourceQ_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_of_sourceQ_le_one","description":"theorem inner_supportCombination_le_sumAbs_of_sourceQ_le_one (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) {y : V} (hy : system.sourceQ y ≤ 1) : ⟪supportCombination basis a, y⟫_ℝ ≤ ∑ i, |a i|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-e7e38fcf731c","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5923,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:400"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_supportCombination_le_sumAbs_of_sourceQ_le_one (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) {y : V} (hy : system.sourceQ y ≤ 1) : ⟪supportCombination basis a, y⟫_ℝ ≤ ∑ i, |a i|","missing":[],"search":"inner_supportcombination_le_sumabs_of_sourceq_le_one banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.inner_supportcombination_le_sumabs_of_sourceq_le_one theorem inner_supportcombination_le_sumabs_of_sourceq_le_one (support : issuccinctsupport system basis) (a : fin s → ℝ) {y : v} (hy : system.sourceq y ≤ 1) : ⟪supportcombination basis a, y⟫_ℝ ≤ ∑ i, |a i| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceRSet_bddAbove","label":"sourceRSet_bddAbove","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceRSet_bddAbove","description":"theorem sourceRSet_bddAbove (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : BddAbove ((fun y : V => ⟪supportCombination basis a, y⟫_ℝ) '' {y | system.sourceQ y ≤ 1})","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-c738beee7d0b","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5924,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:413"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceRSet_bddAbove (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : BddAbove ((fun y : V => ⟪supportCombination basis a, y⟫_ℝ) '' {y | system.sourceQ y ≤ 1})","missing":[],"search":"sourcerset_bddabove banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.sourcerset_bddabove theorem sourcerset_bddabove (support : issuccinctsupport system basis) (a : fin s → ℝ) : bddabove ((fun y : v => ⟪supportcombination basis a, y⟫_ℝ) '' {y | system.sourceq y ≤ 1}) theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceR_supportCombination_eq","label":"sourceR_supportCombination_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceR_supportCombination_eq","description":"Source Lemma 3.2 for a succinct vector. The proof also supplies the boundedness evidence missing from a bare use of real `sSup`.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-b16fd92d44b6","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5925,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:423"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceR_supportCombination_eq [Nonempty (Fin s)] (support : IsSuccinctSupport system basis) (a : Fin s → ℝ) : system.sourceR (supportCombination basis a) = ∑ i, |a i|","missing":[],"search":"sourcer_supportcombination_eq banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctsupport.sourcer_supportcombination_eq source lemma 3.2 for a succinct vector. the proof also supplies the boundedness evidence missing from a bare use of real `ssup`. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation","label":"SuccinctRepresentation","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation","description":"Source Definition 3.3, represented with its witness data exposed. The positive support-size premise is recorded explicitly.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-c5a3300ab528","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5926,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:447"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure SuccinctRepresentation (system : SuccinctUnitSystem V) (x : V) (s : Nat) where","missing":[],"search":"succinctrepresentation banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation source definition 3.3, represented with its witness data exposed. the positive support-size premise is recorded explicitly. structure compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation","label":"StrictSuccinctRepresentation","kind":"structure","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation","description":"The strict clause in source Definition 3.3: every coefficient in the chosen succinct representation is nonzero.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-2e7f68a8c31f","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5927,"meta":[["Kind","structure"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:457"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure StrictSuccinctRepresentation (system : SuccinctUnitSystem V) (x : V) (s : Nat) extends SuccinctRepresentation system x s where","missing":[],"search":"strictsuccinctrepresentation banditrlproof.lowerbounds.succinct.succinctunitsystem.strictsuccinctrepresentation the strict clause in source definition 3.3: every coefficient in the chosen succinct representation is nonzero. structure compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctAt","label":"IsSuccinctAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctAt","description":"Proposition-level source wording: `x` admits an `s`-succinct representation.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-b0052c09bd2b","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5928,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:463"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def IsSuccinctAt (system : SuccinctUnitSystem V) (x : V) (s : Nat) : Prop","missing":[],"search":"issuccinctat banditrlproof.lowerbounds.succinct.succinctunitsystem.issuccinctat proposition-level source wording: `x` admits an `s`-succinct representation. definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsStrictlySuccinctAt","label":"IsStrictlySuccinctAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsStrictlySuccinctAt","description":"Proposition-level source wording: `x` admits a strictly `s`-succinct representation.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-1db24e5934cf","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5929,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:468"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def IsStrictlySuccinctAt (system : SuccinctUnitSystem V) (x : V) (s : Nat) : Prop","missing":[],"search":"isstrictlysuccinctat banditrlproof.lowerbounds.succinct.succinctunitsystem.isstrictlysuccinctat proposition-level source wording: `x` admits a strictly `s`-succinct representation. definition compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceR_eq_sumAbs","label":"sourceR_eq_sumAbs","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceR_eq_sumAbs","description":"theorem sourceR_eq_sumAbs (representation : SuccinctRepresentation system x s) : system.sourceR x = ∑ i, |representation.coefficients i|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-741b96628266","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5930,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:475"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceR_eq_sumAbs (representation : SuccinctRepresentation system x s) : system.sourceR x = ∑ i, |representation.coefficients i|","missing":[],"search":"sourcer_eq_sumabs banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.sourcer_eq_sumabs theorem sourcer_eq_sumabs (representation : succinctrepresentation system x s) : system.sourcer x = ∑ i, |representation.coefficients i| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.inner_supportSignCombination_eq_sumAbs","label":"inner_supportSignCombination_eq_sumAbs","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.inner_supportSignCombination_eq_sumAbs","description":"theorem inner_supportSignCombination_eq_sumAbs (representation : SuccinctRepresentation system x s) : ⟪x, IsSuccinctSupport.supportSignCombination representation.basis representation.coefficients⟫_ℝ = ∑ i, |representation.coefficients i|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-62ab0c3a28fa","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5931,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:486"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem inner_supportSignCombination_eq_sumAbs (representation : SuccinctRepresentation system x s) : ⟪x, IsSuccinctSupport.supportSignCombination representation.basis representation.coefficients⟫_ℝ = ∑ i, |representation.coefficients i|","missing":[],"search":"inner_supportsigncombination_eq_sumabs banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.inner_supportsigncombination_eq_sumabs theorem inner_supportsigncombination_eq_sumabs (representation : succinctrepresentation system x s) : ⟪x, issuccinctsupport.supportsigncombination representation.basis representation.coefficients⟫_ℝ = ∑ i, |representation.coefficients i| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceQ_supportSignCombination_eq_one","label":"sourceQ_supportSignCombination_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceQ_supportSignCombination_eq_one","description":"theorem sourceQ_supportSignCombination_eq_one (representation : SuccinctRepresentation system x s) : system.sourceQ (IsSuccinctSupport.supportSignCombination representation.basis representation.coefficients) = 1","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-8e27761c92a7","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5932,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:505"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sourceQ_supportSignCombination_eq_one (representation : SuccinctRepresentation system x s) : system.sourceQ (IsSuccinctSupport.supportSignCombination representation.basis representation.coefficients) = 1","missing":[],"search":"sourceq_supportsigncombination_eq_one banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.sourceq_supportsigncombination_eq_one theorem sourceq_supportsigncombination_eq_one (representation : succinctrepresentation system x s) : system.sourceq (issuccinctsupport.supportsigncombination representation.basis representation.coefficients) = 1 theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.norm_sq_supportSignCombination_eq_size","label":"norm_sq_supportSignCombination_eq_size","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.norm_sq_supportSignCombination_eq_size","description":"theorem norm_sq_supportSignCombination_eq_size (representation : SuccinctRepresentation system x s) : ‖IsSuccinctSupport.supportSignCombination representation.basis representation.coefficients‖ ^ 2 = (s : ℝ)","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-67f86d36d56c","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5933,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:512"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem norm_sq_supportSignCombination_eq_size (representation : SuccinctRepresentation system x s) : ‖IsSuccinctSupport.supportSignCombination representation.basis representation.coefficients‖ ^ 2 = (s : ℝ)","missing":[],"search":"norm_sq_supportsigncombination_eq_size banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.norm_sq_supportsigncombination_eq_size theorem norm_sq_supportsigncombination_eq_size (representation : succinctrepresentation system x s) : ‖issuccinctsupport.supportsigncombination representation.basis representation.coefficients‖ ^ 2 = (s : ℝ) theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sumAbs_eq","label":"sumAbs_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sumAbs_eq","description":"The two local Lemma-3.2 identities give equality of coefficient `l1` sums for any two succinct representations of the same vector.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-fe55ff0308b3","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5934,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:521"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem sumAbs_eq (first : SuccinctRepresentation system x s) (second : SuccinctRepresentation system x z) : (∑ i, |first.coefficients i|) = ∑ j, |second.coefficients j|","missing":[],"search":"sumabs_eq banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.sumabs_eq the two local lemma-3.2 identities give equality of coefficient `l1` sums for any two succinct representations of the same vector. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.abs_inner_strictBasis_supportSignCombination_eq_one","label":"abs_inner_strictBasis_supportSignCombination_eq_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.abs_inner_strictBasis_supportSignCombination_eq_one","description":"Appendix A.3 equality case: if the second representation is strict, every second-support atom has unit absolute correlation with the sign sum of the first support.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-b5243e1d5757","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5935,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:532"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_inner_strictBasis_supportSignCombination_eq_one (first : SuccinctRepresentation system x s) (second : StrictSuccinctRepresentation system x z) (j : Fin z) : |⟪second.toSuccinctRepresentation.basis j, IsSuccinctSupport.supportSignCombination first.basis first.coefficients⟫_ℝ| = 1","missing":[],"search":"abs_inner_strictbasis_supportsigncombination_eq_one banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.abs_inner_strictbasis_supportsigncombination_eq_one appendix a.3 equality case: if the second representation is strict, every second-support atom has unit absolute correlation with the sign sum of the first support. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.strictSize_le","label":"strictSize_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.strictSize_le","description":"Source Lemma 3.3 on explicit witnesses: if the same vector has an `s`-succinct representation and a strictly `z`-succinct representation, then `z <= s`.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-1e568861c361","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5936,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:598"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem strictSize_le (first : SuccinctRepresentation system x s) (second : StrictSuccinctRepresentation system x z) : z ≤ s","missing":[],"search":"strictsize_le banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctrepresentation.strictsize_le source lemma 3.3 on explicit witnesses: if the same vector has an `s`-succinct representation and a strictly `z`-succinct representation, then `z <= s`. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation.abs_coefficient_pos","label":"abs_coefficient_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation.abs_coefficient_pos","description":"theorem abs_coefficient_pos (representation : StrictSuccinctRepresentation system x s) (i : Fin s) : 0 < |representation.toSuccinctRepresentation.coefficients i|","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-8d781aa2815f","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5937,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:632"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem abs_coefficient_pos (representation : StrictSuccinctRepresentation system x s) (i : Fin s) : 0 < |representation.toSuccinctRepresentation.coefficients i|","missing":[],"search":"abs_coefficient_pos banditrlproof.lowerbounds.succinct.succinctunitsystem.strictsuccinctrepresentation.abs_coefficient_pos theorem abs_coefficient_pos (representation : strictsuccinctrepresentation system x s) (i : fin s) : 0 < |representation.tosuccinctrepresentation.coefficients i| theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","label":"succinctSize_ge_strictSize","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","description":"Source Lemma 3.3 in proposition-level Definition-3.3 wording.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-46a3880e28c0","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5938,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:640"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem succinctSize_ge_strictSize {system : SuccinctUnitSystem V} {x : V} {s z : Nat} (hs : IsSuccinctAt system x s) (hz : IsStrictlySuccinctAt system x z) : z ≤ s","missing":[],"search":"succinctsize_ge_strictsize banditrlproof.lowerbounds.succinct.succinctunitsystem.succinctsize_ge_strictsize source lemma 3.3 in proposition-level definition-3.3 wording. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.strictlySuccinctSize_unique","label":"strictlySuccinctSize_unique","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.strictlySuccinctSize_unique","description":"Source Lemma 3.4: two strict succinct representations of the same vector have the same size.","url":"../modules/banditrlproof-lowerbounds-succinctgeometryaudit/index.html#decl-26966174c834","parent":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","order":5939,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.SuccinctGeometryAudit"],["Source","BanditRLProof/LowerBounds/SuccinctGeometryAudit.lean:649"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"theorem strictlySuccinctSize_unique {system : SuccinctUnitSystem V} {x : V} {s z : Nat} (hs : IsStrictlySuccinctAt system x s) (hz : IsStrictlySuccinctAt system x z) : s = z","missing":[],"search":"strictlysuccinctsize_unique banditrlproof.lowerbounds.succinct.succinctunitsystem.strictlysuccinctsize_unique source lemma 3.4: two strict succinct representations of the same vector have the same size. theorem compiled","shard":"modules/c766d95e3c82faab.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.uniformPowerTwo_entropy","label":"uniformPowerTwo_entropy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.uniformPowerTwo_entropy","description":"theorem uniformPowerTwo_entropy {α : Type*} [Fintype α] (n : ℕ) (hcard : Fintype.card α = 2 ^ n) : discreteEntropyBaseTwo Finset.univ (fun _ : α => (1 / (2 : ℝ) ^ n)) = n","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html#decl-ac32994a9911","parent":"module:BanditRLProof.LowerBounds.UniformCoding","order":5940,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.UniformCoding"],["Source","BanditRLProof/LowerBounds/UniformCoding.lean:6"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem uniformPowerTwo_entropy {α : Type*} [Fintype α] (n : ℕ) (hcard : Fintype.card α = 2 ^ n) : discreteEntropyBaseTwo Finset.univ (fun _ : α => (1 / (2 : ℝ) ^ n)) = n","missing":[],"search":"uniformpowertwo_entropy banditrlproof.lowerbounds.uniformpowertwo_entropy theorem uniformpowertwo_entropy {α : type*} [fintype α] (n : ℕ) (hcard : fintype.card α = 2 ^ n) : discreteentropybasetwo finset.univ (fun _ : α => (1 / (2 : ℝ) ^ n)) = n theorem compiled","shard":"modules/6e5ad5e5e5d2a9a5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.fixedLength_uniformPowerTwo_optimal","label":"fixedLength_uniformPowerTwo_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.fixedLength_uniformPowerTwo_optimal","description":"theorem fixedLength_uniformPowerTwo_optimal {α : Type*} [Fintype α] (n : ℕ) (hcard : Fintype.card α = 2 ^ n) (code : BinaryPrefixCode α) (hlen : ∀ a, (code.encode a).length = n) : IsOptimalPrefixCode (fun _ : α => 1 / (2 : ℝ) ^ n) code","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html#decl-4d37734e1e15","parent":"module:BanditRLProof.LowerBounds.UniformCoding","order":5941,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.UniformCoding"],["Source","BanditRLProof/LowerBounds/UniformCoding.lean:12"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem fixedLength_uniformPowerTwo_optimal {α : Type*} [Fintype α] (n : ℕ) (hcard : Fintype.card α = 2 ^ n) (code : BinaryPrefixCode α) (hlen : ∀ a, (code.encode a).length = n) : IsOptimalPrefixCode (fun _ : α => 1 / (2 : ℝ) ^ n) code","missing":[],"search":"fixedlength_uniformpowertwo_optimal banditrlproof.lowerbounds.fixedlength_uniformpowertwo_optimal theorem fixedlength_uniformpowertwo_optimal {α : type*} [fintype α] (n : ℕ) (hcard : fintype.card α = 2 ^ n) (code : binaryprefixcode α) (hlen : ∀ a, (code.encode a).length = n) : isoptimalprefixcode (fun _ : α => 1 / (2 : ℝ) ^ n) code theorem compiled","shard":"modules/6e5ad5e5e5d2a9a5.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.ternaryPrefixWord","label":"ternaryPrefixWord","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ternaryPrefixWord","description":"def ternaryPrefixWord (a : Fin 3) : List Bool","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html#decl-041a09eeafed","parent":"module:BanditRLProof.LowerBounds.UniformCoding","order":5942,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.UniformCoding"],["Source","BanditRLProof/LowerBounds/UniformCoding.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def ternaryPrefixWord (a : Fin 3) : List Bool","missing":[],"search":"ternaryprefixword banditrlproof.lowerbounds.ternaryprefixword def ternaryprefixword (a : fin 3) : list bool definition compiled","shard":"modules/6e5ad5e5e5d2a9a5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.ternaryPrefixCode","label":"ternaryPrefixCode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ternaryPrefixCode","description":"def ternaryPrefixCode : BinaryPrefixCode (Fin 3) where","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html#decl-1e345eed03f6","parent":"module:BanditRLProof.LowerBounds.UniformCoding","order":5943,"meta":[["Kind","definition"],["Module","BanditRLProof.LowerBounds.UniformCoding"],["Source","BanditRLProof/LowerBounds/UniformCoding.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"def ternaryPrefixCode : BinaryPrefixCode (Fin 3) where","missing":[],"search":"ternaryprefixcode banditrlproof.lowerbounds.ternaryprefixcode def ternaryprefixcode : binaryprefixcode (fin 3) where definition compiled","shard":"modules/6e5ad5e5e5d2a9a5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.ternaryPrefixCode_uniform_length","label":"ternaryPrefixCode_uniform_length","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.ternaryPrefixCode_uniform_length","description":"theorem ternaryPrefixCode_uniform_length : expectedCodeLength (fun _ : Fin 3 => (1 / 3 : ℝ)) ternaryPrefixCode = 5 / 3","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html#decl-e6c94f8b6e40","parent":"module:BanditRLProof.LowerBounds.UniformCoding","order":5944,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.UniformCoding"],["Source","BanditRLProof/LowerBounds/UniformCoding.lean:37"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ternaryPrefixCode_uniform_length : expectedCodeLength (fun _ : Fin 3 => (1 / 3 : ℝ)) ternaryPrefixCode = 5 / 3","missing":[],"search":"ternaryprefixcode_uniform_length banditrlproof.lowerbounds.ternaryprefixcode_uniform_length theorem ternaryprefixcode_uniform_length : expectedcodelength (fun _ : fin 3 => (1 / 3 : ℝ)) ternaryprefixcode = 5 / 3 theorem compiled","shard":"modules/6e5ad5e5e5d2a9a5.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.LowerBounds.uniform_three_fixedLength_not_optimal","label":"uniform_three_fixedLength_not_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.LowerBounds.uniform_three_fixedLength_not_optimal","description":"Uniform masses do not make a constant-length code optimal for every alphabet size.","url":"../modules/banditrlproof-lowerbounds-uniformcoding/index.html#decl-93b69ea440f6","parent":"module:BanditRLProof.LowerBounds.UniformCoding","order":5945,"meta":[["Kind","theorem"],["Module","BanditRLProof.LowerBounds.UniformCoding"],["Source","BanditRLProof/LowerBounds/UniformCoding.lean:42"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations, spine:chapter-14-information-theory"],["Indexed settings","None registered"]],"statement":"theorem uniform_three_fixedLength_not_optimal (code : BinaryPrefixCode (Fin 3)) (hlen : ∀ a, (code.encode a).length = 2) : ¬ IsOptimalPrefixCode (fun _ : Fin 3 => (1 / 3 : ℝ)) code","missing":[],"search":"uniform_three_fixedlength_not_optimal banditrlproof.lowerbounds.uniform_three_fixedlength_not_optimal uniform masses do not make a constant-length code optimal for every alphabet size. theorem compiled","shard":"modules/6e5ad5e5e5d2a9a5.json","books":["bandit"],"chapters":["teaching:foundations","spine:chapter-14-information-theory"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference","label":"SuccMartingaleDifference","kind":"structure","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifference","description":"Global succ-indexed martingale-difference witness. This is the all-time version of `SuccMartingaleDifferencePrefix`. It records the local hypotheses needed to build the Mathlib martingale of partial sums.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-0a4f4ccbdcfe","parent":"module:BanditRLProof.MartingaleDifference","order":5946,"meta":[["Kind","structure"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:24"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure SuccMartingaleDifference {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) (F : Filtration Nat mOmega) (Y : Nat -> Omega -> Real) : Prop where","missing":[],"search":"succmartingaledifference banditrlproof.martingalediff.succmartingaledifference global succ-indexed martingale-difference witness. this is the all-time version of `succmartingaledifferenceprefix`. it records the local hypotheses needed to build the mathlib martingale of partial sums. structure compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix","label":"SuccMartingaleDifferencePrefix","kind":"structure","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix","description":"Finite-prefix succ-indexed martingale-difference witness. For a process `Y`, this records that the prefix up to `n` is integrable and adapted, and that each later increment `Y (i + 1)` has zero conditional expectation against filtration level `i`. This is the local bridge between the compiled conditional-mean-zero leaves and later martingale or concentration consumers.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-e3dcfc46a6c1","parent":"module:BanditRLProof.MartingaleDifference","order":5947,"meta":[["Kind","structure"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:46"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure SuccMartingaleDifferencePrefix {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) (F : Filtration Nat mOmega) (Y : Nat -> Omega -> Real) (n : Nat) : Prop where","missing":[],"search":"succmartingaledifferenceprefix banditrlproof.martingalediff.succmartingaledifferenceprefix finite-prefix succ-indexed martingale-difference witness. for a process `y`, this records that the prefix up to `n` is integrable and adapted, and that each later increment `y (i + 1)` has zero conditional expectation against filtration level `i`. this is the local bridge between the compiled conditional-mean-zero leaves and later martingale or concentration consumers. structure compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.toPrefix","label":"toPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifference.toPrefix","description":"theorem toPrefix {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) (n : Nat) : SuccMartingaleDifferencePrefix mu F Y n where","url":"../modules/banditrlproof-martingaledifference/index.html#decl-ba2bcf606aa7","parent":"module:BanditRLProof.MartingaleDifference","order":5948,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:62"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem toPrefix {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) (n : Nat) : SuccMartingaleDifferencePrefix mu F Y n where","missing":[],"search":"toprefix banditrlproof.martingalediff.succmartingaledifference.toprefix theorem toprefix {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} (h : succmartingaledifference mu f y) (n : nat) : succmartingaledifferenceprefix mu f y n where theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.stronglyAdapted'","label":"stronglyAdapted'","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifference.stronglyAdapted'","description":"theorem stronglyAdapted' {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) : StronglyAdapted F Y","url":"../modules/banditrlproof-martingaledifference/index.html#decl-f4cb25d2a6b9","parent":"module:BanditRLProof.MartingaleDifference","order":5949,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:74"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem stronglyAdapted' {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) : StronglyAdapted F Y","missing":[],"search":"stronglyadapted' banditrlproof.martingalediff.succmartingaledifference.stronglyadapted' theorem stronglyadapted' {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} (h : succmartingaledifference mu f y) : stronglyadapted f y theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.integrable'","label":"integrable'","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifference.integrable'","description":"theorem integrable' {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) (t : Nat) : Integrable (Y t) mu","url":"../modules/banditrlproof-martingaledifference/index.html#decl-c389b769cf07","parent":"module:BanditRLProof.MartingaleDifference","order":5950,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:83"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable' {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) (t : Nat) : Integrable (Y t) mu","missing":[],"search":"integrable' banditrlproof.martingalediff.succmartingaledifference.integrable' theorem integrable' {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} (h : succmartingaledifference mu f y) (t : nat) : integrable (y t) mu theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.condExp_succ_ae_eq_zero","label":"condExp_succ_ae_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifference.condExp_succ_ae_eq_zero","description":"theorem condExp_succ_ae_eq_zero {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) (i : Nat) : Filter.EventuallyEq (ae mu) (condExp (F i) mu (Y (i + 1))) (fun _omega => (0 : Real))","url":"../modules/banditrlproof-martingaledifference/index.html#decl-f4977976709b","parent":"module:BanditRLProof.MartingaleDifference","order":5951,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:93"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExp_succ_ae_eq_zero {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) (i : Nat) : Filter.EventuallyEq (ae mu) (condExp (F i) mu (Y (i + 1))) (fun _omega => (0 : Real))","missing":[],"search":"condexp_succ_ae_eq_zero banditrlproof.martingalediff.succmartingaledifference.condexp_succ_ae_eq_zero theorem condexp_succ_ae_eq_zero {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} (h : succmartingaledifference mu f y) (i : nat) : filter.eventuallyeq (ae mu) (condexp (f i) mu (y (i + 1))) (fun _omega => (0 : real)) theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.stronglyAdapted'","label":"stronglyAdapted'","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.stronglyAdapted'","description":"theorem stronglyAdapted' {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} {n : Nat} (h : SuccMartingaleDifferencePrefix mu F Y n) : StronglyAdapted F Y","url":"../modules/banditrlproof-martingaledifference/index.html#decl-b35144190e5a","parent":"module:BanditRLProof.MartingaleDifference","order":5952,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:109"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem stronglyAdapted' {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} {n : Nat} (h : SuccMartingaleDifferencePrefix mu F Y n) : StronglyAdapted F Y","missing":[],"search":"stronglyadapted' banditrlproof.martingalediff.succmartingaledifferenceprefix.stronglyadapted' theorem stronglyadapted' {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} {n : nat} (h : succmartingaledifferenceprefix mu f y n) : stronglyadapted f y theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.integrable_of_lt","label":"integrable_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.integrable_of_lt","description":"theorem integrable_of_lt {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} {n : Nat} (h : SuccMartingaleDifferencePrefix mu F Y n) {t : Nat} (ht : t < n) : Integrable (Y t) mu","url":"../modules/banditrlproof-martingaledifference/index.html#decl-ec5ad344e259","parent":"module:BanditRLProof.MartingaleDifference","order":5953,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:119"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_of_lt {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} {n : Nat} (h : SuccMartingaleDifferencePrefix mu F Y n) {t : Nat} (ht : t < n) : Integrable (Y t) mu","missing":[],"search":"integrable_of_lt banditrlproof.martingalediff.succmartingaledifferenceprefix.integrable_of_lt theorem integrable_of_lt {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} {n : nat} (h : succmartingaledifferenceprefix mu f y n) {t : nat} (ht : t < n) : integrable (y t) mu theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.condExp_succ_ae_eq_zero","label":"condExp_succ_ae_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.condExp_succ_ae_eq_zero","description":"theorem condExp_succ_ae_eq_zero {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} {n : Nat} (h : SuccMartingaleDifferencePrefix mu F Y n) {i : Nat} (hi : i + 1 < n) : Filter.EventuallyEq (ae mu) (condExp (F i) mu (Y (i + 1))) (fun _omega => (0 : Real))","url":"../modules/banditrlproof-martingaledifference/index.html#decl-0819d4625c27","parent":"module:BanditRLProof.MartingaleDifference","order":5954,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:130"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem condExp_succ_ae_eq_zero {Omega : Type u} [mOmega : MeasurableSpace Omega] {mu : Measure Omega} {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} {n : Nat} (h : SuccMartingaleDifferencePrefix mu F Y n) {i : Nat} (hi : i + 1 < n) : Filter.EventuallyEq (ae mu) (condExp (F i) mu (Y (i + 1))) (fun _omega => (0 : Real))","missing":[],"search":"condexp_succ_ae_eq_zero banditrlproof.martingalediff.succmartingaledifferenceprefix.condexp_succ_ae_eq_zero theorem condexp_succ_ae_eq_zero {omega : type u} [momega : measurablespace omega] {mu : measure omega} {f : filtration nat momega} {y : nat -> omega -> real} {n : nat} (h : succmartingaledifferenceprefix mu f y n) {i : nat} (hi : i + 1 < n) : filter.eventuallyeq (ae mu) (condexp (f i) mu (y (i + 1))) (fun _omega => (0 : real)) theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.centeredRewardProcess","label":"centeredRewardProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.centeredRewardProcess","description":"Centered reward process obtained by subtracting a predictable or otherwise chosen baseline from a raw reward process. The baseline is allowed to be random; the martingale-difference builder below therefore asks directly for adaptedness, integrability, and conditional mean-zero of this centered process.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-8f4a905f6c98","parent":"module:BanditRLProof.MartingaleDifference","order":5955,"meta":[["Kind","definition"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:153"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def centeredRewardProcess {Omega : Type u} (reward baseline : Nat -> Omega -> Real) : Nat -> Omega -> Real","missing":[],"search":"centeredrewardprocess banditrlproof.martingalediff.centeredrewardprocess centered reward process obtained by subtracting a predictable or otherwise chosen baseline from a raw reward process. the baseline is allowed to be random; the martingale-difference builder below therefore asks directly for adaptedness, integrability, and conditional mean-zero of this centered process. definition compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.succMartingaleDifference_centeredRewardProcess_of_condExp","label":"succMartingaleDifference_centeredRewardProcess_of_condExp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.succMartingaleDifference_centeredRewardProcess_of_condExp","description":"Build the global succ-indexed martingale-difference witness for a centered reward process from the exact three local contracts used by Mathlib: adaptedness, integrability, and succ-indexed conditional mean zero.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-cff22aded0c6","parent":"module:BanditRLProof.MartingaleDifference","order":5956,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:164"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem succMartingaleDifference_centeredRewardProcess_of_condExp {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) {F : Filtration Nat mOmega} (reward baseline : Nat -> Omega -> Real) (hadapted : StronglyAdapted F (centeredRewardProcess reward baseline)) (hintegrable : forall t, Integrable (centeredRewardProcess reward baseline t) mu) (hcond : forall i, Filter.EventuallyEq (ae mu) (condExp (F i) mu (centeredRewardProcess reward baseline (i + 1))) (fun _omega => (0 : Real))) : SuccMartingaleDifference mu F (centeredRewardProcess reward baseline) where","missing":[],"search":"succmartingaledifference_centeredrewardprocess_of_condexp banditrlproof.martingalediff.succmartingaledifference_centeredrewardprocess_of_condexp build the global succ-indexed martingale-difference witness for a centered reward process from the exact three local contracts used by mathlib: adaptedness, integrability, and succ-indexed conditional mean zero. theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.succMartingaleDifferencePrefix_centeredRewardProcess_of_condExp","label":"succMartingaleDifferencePrefix_centeredRewardProcess_of_condExp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.succMartingaleDifferencePrefix_centeredRewardProcess_of_condExp","description":"Finite-prefix version of `succMartingaleDifference_centeredRewardProcess_of_condExp`.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-d2108f847651","parent":"module:BanditRLProof.MartingaleDifference","order":5957,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:189"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem succMartingaleDifferencePrefix_centeredRewardProcess_of_condExp {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) {F : Filtration Nat mOmega} (reward baseline : Nat -> Omega -> Real) (n : Nat) (hadapted : StronglyAdapted F (centeredRewardProcess reward baseline)) (hintegrable : forall t, t < n -> Integrable (centeredRewardProcess reward baseline t) mu) (hcond : forall i, i + 1 < n -> Filter.EventuallyEq (ae mu) (condExp (F i) mu (centeredRewardProcess reward baseline (i + 1))) (fun _omega => (0 : Real))) : SuccMartingaleDifferencePrefix mu F (centeredRewardProcess reward baseline) n where","missing":[],"search":"succmartingaledifferenceprefix_centeredrewardprocess_of_condexp banditrlproof.martingalediff.succmartingaledifferenceprefix_centeredrewardprocess_of_condexp finite-prefix version of `succmartingaledifference_centeredrewardprocess_of_condexp`. theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.partialSumsSucc","label":"partialSumsSucc","kind":"definition","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.partialSumsSucc","description":"Partial sums of a succ-indexed martingale-difference process. The sum starts at `Y 1`, so the increment from time `i` to `i + 1` is definitionally `Y (i + 1)`, matching Mathlib's martingale theorem `martingale_of_condExp_sub_eq_zero_nat`.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-d9c245a0071a","parent":"module:BanditRLProof.MartingaleDifference","order":5958,"meta":[["Kind","definition"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:219"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def partialSumsSucc {Omega : Type u} (Y : Nat -> Omega -> Real) : Nat -> Omega -> Real","missing":[],"search":"partialsumssucc banditrlproof.martingalediff.partialsumssucc partial sums of a succ-indexed martingale-difference process. the sum starts at `y 1`, so the increment from time `i` to `i + 1` is definitionally `y (i + 1)`, matching mathlib's martingale theorem `martingale_of_condexp_sub_eq_zero_nat`. definition compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.martingale_partialSumsSucc_of_succMartingaleDifference","label":"martingale_partialSumsSucc_of_succMartingaleDifference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.martingale_partialSumsSucc_of_succMartingaleDifference","description":"The partial sums of a global succ-indexed martingale-difference process form a Mathlib martingale. This is a thin wrapper around `MeasureTheory.martingale_of_condExp_sub_eq_zero_nat`. It does not assert optional stopping, concentration, or final regret; it only upgrades the local conditional-mean-zero increment contract to Mathlib's martingale API.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-17188820c890","parent":"module:BanditRLProof.MartingaleDifference","order":5959,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:234"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem martingale_partialSumsSucc_of_succMartingaleDifference {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} {Y : Nat -> Omega -> Real} (h : SuccMartingaleDifference mu F Y) : Martingale (partialSumsSucc Y) F mu","missing":[],"search":"martingale_partialsumssucc_of_succmartingaledifference banditrlproof.martingalediff.martingale_partialsumssucc_of_succmartingaledifference the partial sums of a global succ-indexed martingale-difference process form a mathlib martingale. this is a thin wrapper around `measuretheory.martingale_of_condexp_sub_eq_zero_nat`. it does not assert optional stopping, concentration, or final regret; it only upgrades the local conditional-mean-zero increment contract to mathlib's martingale api. theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.MartingaleDiff.martingale_partialSumsSucc_centeredRewardProcess_of_condExp","label":"martingale_partialSumsSucc_centeredRewardProcess_of_condExp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.MartingaleDiff.martingale_partialSumsSucc_centeredRewardProcess_of_condExp","description":"The partial sums of a global centered reward process form a Mathlib martingale once the centered process satisfies the martingale-difference contracts.","url":"../modules/banditrlproof-martingaledifference/index.html#decl-0a990dc3485f","parent":"module:BanditRLProof.MartingaleDifference","order":5960,"meta":[["Kind","theorem"],["Module","BanditRLProof.MartingaleDifference"],["Source","BanditRLProof/MartingaleDifference.lean:270"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem martingale_partialSumsSucc_centeredRewardProcess_of_condExp {Omega : Type u} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (reward baseline : Nat -> Omega -> Real) (hadapted : StronglyAdapted F (centeredRewardProcess reward baseline)) (hintegrable : forall t, Integrable (centeredRewardProcess reward baseline t) mu) (hcond : forall i, Filter.EventuallyEq (ae mu) (condExp (F i) mu (centeredRewardProcess reward baseline (i + 1))) (fun _omega => (0 : Real))) : Martingale (partialSumsSucc (centeredRewardProcess reward baseline)) F mu","missing":[],"search":"martingale_partialsumssucc_centeredrewardprocess_of_condexp banditrlproof.martingalediff.martingale_partialsumssucc_centeredrewardprocess_of_condexp the partial sums of a global centered reward process form a mathlib martingale once the centered process satisfies the martingale-difference contracts. theorem compiled","shard":"modules/a0be58a134567394.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_eq_finset_filter_card","label":"pullCount_eq_finset_filter_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_eq_finset_filter_card","description":"The recursive pull count equals the cardinality of the matching arm times in `Finset.range t`. This is the first Mathlib-backed wrapper leaf. It is intentionally kept separate from the dependency-light `List.range` bridge.","url":"../modules/banditrlproof-mathlibwrappers/index.html#decl-80536ec1c719","parent":"module:BanditRLProof.MathlibWrappers","order":5961,"meta":[["Kind","theorem"],["Module","BanditRLProof.MathlibWrappers"],["Source","BanditRLProof/MathlibWrappers.lean:28"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_eq_finset_filter_card : pullCount action a t = ((Finset.range t).filter (fun s : Nat => action s = a)).card","missing":[],"search":"pullcount_eq_finset_filter_card banditrlproof.pullcount_eq_finset_filter_card the recursive pull count equals the cardinality of the matching arm times in `finset.range t`. this is the first mathlib-backed wrapper leaf. it is intentionally kept separate from the dependency-light `list.range` bridge. theorem compiled","shard":"modules/bf55ef0c3ad0f92f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_eq_finset_filter_sum","label":"sumRewards_eq_finset_filter_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_eq_finset_filter_sum","description":"The recursive selected reward sum equals the Mathlib finite sum over selected time points in `Finset.range t`. The local recursive definition only needs weak `0` and `+` operations. This Mathlib-facing wrapper strengthens the algebra contract to `AddCommMonoid` because `Finset.sum` is commutative and the insertion proof swaps summand order.","url":"../modules/banditrlproof-mathlibwrappers/index.html#decl-39f5b12da36d","parent":"module:BanditRLProof.MathlibWrappers","order":5962,"meta":[["Kind","theorem"],["Module","BanditRLProof.MathlibWrappers"],["Source","BanditRLProof/MathlibWrappers.lean:62"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sumRewards_eq_finset_filter_sum : sumRewards action reward a t = ((Finset.range t).filter (fun s : Nat => action s = a)).sum (fun s : Nat => reward s)","missing":[],"search":"sumrewards_eq_finset_filter_sum banditrlproof.sumrewards_eq_finset_filter_sum the recursive selected reward sum equals the mathlib finite sum over selected time points in `finset.range t`. the local recursive definition only needs weak `0` and `+` operations. this mathlib-facing wrapper strengthens the algebra contract to `addcommmonoid` because `finset.sum` is commutative and the insertion proof swaps summand order. theorem compiled","shard":"modules/bf55ef0c3ad0f92f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum","label":"pseudoRegret_eq_finset_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_eq_finset_sum","description":"The recursive pseudo-regret equals the Mathlib finite sum of selected gaps over `Finset.range t`. This is the Rat-valued finite-sum wrapper. It deliberately comes before the generic reward-sum wrapper because `Finset.sum` only needs the Mathlib `AddCommMonoid Rat` instance supplied by the Rat algebra imports.","url":"../modules/banditrlproof-mathlibwrappers/index.html#decl-795abbb71163","parent":"module:BanditRLProof.MathlibWrappers","order":5963,"meta":[["Kind","theorem"],["Module","BanditRLProof.MathlibWrappers"],["Source","BanditRLProof/MathlibWrappers.lean:93"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_eq_finset_sum : pseudoRegret model action t = (Finset.range t).sum (fun s : Nat => model.gap (action s))","missing":[],"search":"pseudoregret_eq_finset_sum banditrlproof.pseudoregret_eq_finset_sum the recursive pseudo-regret equals the mathlib finite sum of selected gaps over `finset.range t`. this is the rat-valued finite-sum wrapper. it deliberately comes before the generic reward-sum wrapper because `finset.sum` only needs the mathlib `addcommmonoid rat` instance supplied by the rat algebra imports. theorem compiled","shard":"modules/bf55ef0c3ad0f92f.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sumRewards_eq_finset_range_indicator_reward","label":"sumRewards_eq_finset_range_indicator_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sumRewards_eq_finset_range_indicator_reward","description":"private theorem sumRewards_eq_finset_range_indicator_reward {Omega : Type u} {Action : Type v} {Reward : Type v} [DecidableEq Action] [AddCommMonoid Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (a : Action) (n : Nat) (omega : Omega) : sumRewards (action omega) (reward omega) a n = (Finset.range n).sum (fun t : Nat => (({omega' : Omega | action omega' t = a} : Set Omega).indic…","url":"../modules/banditrlproof-measurablelocalquantities/index.html#decl-92efccfcb58a","parent":"module:BanditRLProof.MeasurableLocalQuantities","order":5964,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasurableLocalQuantities"],["Source","BanditRLProof/MeasurableLocalQuantities.lean:16"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"private theorem sumRewards_eq_finset_range_indicator_reward {Omega : Type u} {Action : Type v} {Reward : Type v} [DecidableEq Action] [AddCommMonoid Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (a : Action) (n : Nat) (omega : Omega) : sumRewards (action omega) (reward omega) a n = (Finset.range n).sum (fun t : Nat => (({omega' : Omega | action omega' t = a} : Set Omega).indicator (fun omega' : Omega => reward omega' t)) omega)","missing":[],"search":"sumrewards_eq_finset_range_indicator_reward banditrlproof.sumrewards_eq_finset_range_indicator_reward private theorem sumrewards_eq_finset_range_indicator_reward {omega : type u} {action : type v} {reward : type v} [decidableeq action] [addcommmonoid reward] (action : omega -> actiontrace action) (reward : omega -> rewardtrace reward) (a : action) (n : nat) (omega : omega) : sumrewards (action omega) (reward omega) a n = (finset.range n).sum (fun t : nat => (({omega' : omega | action omega' t = a} : set omega).indicator (fun omega' : omega => reward omega' t)) omega) theorem compiled","shard":"modules/5382c55ae34cc823.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_sumRewards","label":"measurable_sumRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_sumRewards","description":"The local recursive selected-reward accumulator is measurable when action and reward traces are timewise measurable. This is the `MEAS-SUMREWARDS` bridge. It is still only a measurability result: no expectation, probability measure, filtration, or concentration structure is introduced here.","url":"../modules/banditrlproof-measurablelocalquantities/index.html#decl-bfffa74b93f4","parent":"module:BanditRLProof.MeasurableLocalQuantities","order":5965,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasurableLocalQuantities"],["Source","BanditRLProof/MeasurableLocalQuantities.lean:60"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_sumRewards {Omega : Type u} {Action : Type v} {Reward : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [AddCommMonoid Reward] [MeasurableAdd₂ Reward] [DecidableEq Action] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Action) (n : Nat) : Measurable (fun omega : Omega => sumRewards (action omega) (reward omega) a n)","missing":[],"search":"measurable_sumrewards banditrlproof.measurable_sumrewards the local recursive selected-reward accumulator is measurable when action and reward traces are timewise measurable. this is the `meas-sumrewards` bridge. it is still only a measurability result: no expectation, probability measure, filtration, or concentration structure is introduced here. theorem compiled","shard":"modules/5382c55ae34cc823.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_pullCount","label":"measurable_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_pullCount","description":"The local recursive pull count is measurable at each finite horizon. This is the `MEAS-PULLCOUNT` bridge. It prepares expected pull-count leaves without introducing measures, integration, filtration, or concentration.","url":"../modules/banditrlproof-measurablepullcount/index.html#decl-4a4ccbff1833","parent":"module:BanditRLProof.MeasurablePullCount","order":5966,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasurablePullCount"],["Source","BanditRLProof/MeasurablePullCount.lean:23"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_pullCount {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Nat] [MeasurableAdd₂ Nat] [DecidableEq Action] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (n : Nat) : Measurable (fun omega : Omega => pullCount (action omega) a n)","missing":[],"search":"measurable_pullcount banditrlproof.measurable_pullcount the local recursive pull count is measurable at each finite horizon. this is the `meas-pullcount` bridge. it prepares expected pull-count leaves without introducing measures, integration, filtration, or concentration. theorem compiled","shard":"modules/ebd93cbb17611ff1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_natCast_pullCount","label":"measurable_natCast_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_natCast_pullCount","description":"The local recursive pull count, cast into an additive scalar, is measurable at each finite horizon. This is the `MEAS-PULLCOUNT-CAST` bridge. Later expectation leaves can instantiate `Beta := Rat` without changing the proof.","url":"../modules/banditrlproof-measurablepullcountcast/index.html#decl-e9e04c38e68a","parent":"module:BanditRLProof.MeasurablePullCountCast","order":5967,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasurablePullCountCast"],["Source","BanditRLProof/MeasurablePullCountCast.lean:23"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_natCast_pullCount {Omega : Type u} {Action : Type v} {Beta : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Beta] [AddCommMonoidWithOne Beta] [MeasurableAdd₂ Beta] [DecidableEq Action] (action : Omega -> ActionTrace Action) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (n : Nat) : Measurable (fun omega : Omega => ((pullCount (action omega) a n : Nat) : Beta))","missing":[],"search":"measurable_natcast_pullcount banditrlproof.measurable_natcast_pullcount the local recursive pull count, cast into an additive scalar, is measurable at each finite horizon. this is the `meas-pullcount-cast` bridge. later expectation leaves can instantiate `beta := rat` without changing the proof. theorem compiled","shard":"modules/67083779814e664f.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_pseudoRegret","label":"measurable_pseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_pseudoRegret","description":"The local pseudo-regret process is measurable as a random variable at each finite horizon. This is the narrow `MEAS-REGRET` bridge. It does not introduce probability measures, expectations, filtrations, or concentration assumptions.","url":"../modules/banditrlproof-measurableregret/index.html#decl-d6e1ae4205c5","parent":"module:BanditRLProof.MeasurableRegret","order":5968,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasurableRegret"],["Source","BanditRLProof/MeasurableRegret.lean:24"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_pseudoRegret {Omega : Type u} [MeasurableSpace Omega] [MeasurableSpace (Fin K)] [MeasurableSingletonClass (Fin K)] [MeasurableSpace Rat] [MeasurableAdd₂ Rat] (model : FiniteBanditModel K) (action : Omega -> ActionTrace (Fin K)) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (n : Nat) : Measurable (fun omega : Omega => pseudoRegret model (action omega) n)","missing":[],"search":"measurable_pseudoregret banditrlproof.measurable_pseudoregret the local pseudo-regret process is measurable as a random variable at each finite horizon. this is the narrow `meas-regret` bridge. it does not introduce probability measures, expectations, filtrations, or concentration assumptions. theorem compiled","shard":"modules/5385f9f8f4af17ff.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_finset_sum_indicator_reward","label":"measurable_finset_sum_indicator_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_finset_sum_indicator_reward","description":"Finite sums of selected-reward indicator contributions are measurable. This is the `MEAS-SELECTED-REWARD-FINITE-SUM` bridge. The statement is over an arbitrary finite set of times so later range, window, or stopped-prefix corollaries can instantiate the same local API.","url":"../modules/banditrlproof-measurablesums/index.html#decl-1d9c3d0ddf5a","parent":"module:BanditRLProof.MeasurableSums","order":5969,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasurableSums"],["Source","BanditRLProof/MeasurableSums.lean:24"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_finset_sum_indicator_reward {Omega : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [AddCommMonoid Reward] [MeasurableAdd₂ Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Action) (s : Finset Nat) : Measurable (fun omega : Omega => s.sum (fun t : Nat => (({omega' : Omega | action omega' t = a} : Set Omega).indicator (fun omega' : Omega => reward omega' t)) omega))","missing":[],"search":"measurable_finset_sum_indicator_reward banditrlproof.measurable_finset_sum_indicator_reward finite sums of selected-reward indicator contributions are measurable. this is the `meas-selected-reward-finite-sum` bridge. the statement is over an arbitrary finite set of times so later range, window, or stopped-prefix corollaries can instantiate the same local api. theorem compiled","shard":"modules/b5c7f53727a4fe1c.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurableSet_actionTrace_eval_eq","label":"measurableSet_actionTrace_eval_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurableSet_actionTrace_eval_eq","description":"If every time-indexed action random variable is measurable, then the event that the action at a fixed time equals a fixed arm is measurable. This is the `MEAS-FIN-ACTION` canary. The statement is more general than finite actions: it only needs singleton measurability of the action space.","url":"../modules/banditrlproof-measurefoundation/index.html#decl-ca3f57d58252","parent":"module:BanditRLProof.MeasureFoundation","order":5970,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasureFoundation"],["Source","BanditRLProof/MeasureFoundation.lean:25"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_actionTrace_eval_eq {Omega : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] (action : Omega -> ActionTrace Action) (hmeas : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (t : Nat) : MeasurableSet {omega : Omega | action omega t = a}","missing":[],"search":"measurableset_actiontrace_eval_eq banditrlproof.measurableset_actiontrace_eval_eq if every time-indexed action random variable is measurable, then the event that the action at a fixed time equals a fixed arm is measurable. this is the `meas-fin-action` canary. the statement is more general than finite actions: it only needs singleton measurability of the action space. theorem compiled","shard":"modules/d5dcabb5695737f5.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_actionTrace_eval_eq_indicator_const","label":"measurable_actionTrace_eval_eq_indicator_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_actionTrace_eval_eq_indicator_const","description":"The indicator of a measurable action-equality event with a constant value is measurable. This is the `MEAS-PULL-INDICATOR` bridge. It remains scalar-agnostic so later expectation work can choose the codomain deliberately.","url":"../modules/banditrlproof-measurefoundation/index.html#decl-10a170cf497a","parent":"module:BanditRLProof.MeasureFoundation","order":5971,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasureFoundation"],["Source","BanditRLProof/MeasureFoundation.lean:43"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_actionTrace_eval_eq_indicator_const {Omega : Type u} {Action : Type v} {Beta : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Beta] [Zero Beta] (action : Omega -> ActionTrace Action) (hmeas : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (a : Action) (t : Nat) (c : Beta) : Measurable (({omega : Omega | action omega t = a} : Set Omega).indicator (fun _ : Omega => c))","missing":[],"search":"measurable_actiontrace_eval_eq_indicator_const banditrlproof.measurable_actiontrace_eval_eq_indicator_const the indicator of a measurable action-equality event with a constant value is measurable. this is the `meas-pull-indicator` bridge. it remains scalar-agnostic so later expectation work can choose the codomain deliberately. theorem compiled","shard":"modules/d5dcabb5695737f5.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_actionTrace_eval_eq_indicator_reward","label":"measurable_actionTrace_eval_eq_indicator_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_actionTrace_eval_eq_indicator_reward","description":"The selected-reward contribution for a fixed action event is measurable when the action and reward traces are timewise measurable. This is the `MEAS-REWARD` bridge. It deliberately stays at the measurability layer and does not choose an expectation or scalar algebra route.","url":"../modules/banditrlproof-measurefoundation/index.html#decl-7753d56ee225","parent":"module:BanditRLProof.MeasureFoundation","order":5972,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasureFoundation"],["Source","BanditRLProof/MeasureFoundation.lean:64"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_actionTrace_eval_eq_indicator_reward {Omega : Type u} {Action : Type v} {Reward : Type w} [MeasurableSpace Omega] [MeasurableSpace Action] [MeasurableSingletonClass Action] [MeasurableSpace Reward] [Zero Reward] (action : Omega -> ActionTrace Action) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => action omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (a : Action) (t : Nat) : Measurable (({omega : Omega | action omega t = a} : Set Omega).indicator (fun omega : Omega => reward omega t))","missing":[],"search":"measurable_actiontrace_eval_eq_indicator_reward banditrlproof.measurable_actiontrace_eval_eq_indicator_reward the selected-reward contribution for a fixed action event is measurable when the action and reward traces are timewise measurable. this is the `meas-reward` bridge. it deliberately stays at the measurability layer and does not choose an expectation or scalar algebra route. theorem compiled","shard":"modules/d5dcabb5695737f5.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","label":"integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","description":"A nonnegative `L2` function restricted to a measurable event is controlled by its exact second moment and the square root of the event's real mass.","url":"../modules/banditrlproof-measurel2indicator/index.html#decl-21bf83855e73","parent":"module:BanditRLProof.MeasureL2Indicator","order":5973,"meta":[["Kind","theorem"],["Module","BanditRLProof.MeasureL2Indicator"],["Source","BanditRLProof/MeasureL2Indicator.lean:21"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (f : Omega -> Real) (hf_nonneg : forall omega, 0 <= f omega) (hf : MemLp f 2 mu) (event : Set Omega) (hevent : MeasurableSet event) : integral mu (event.indicator f) <= Real.sqrt (integral mu (fun omega => f omega ^ 2)) * Real.sqrt (mu.real event)","missing":[],"search":"integral_indicator_le_sqrt_secondmoment_mul_sqrt_real_measure banditrlproof.integral_indicator_le_sqrt_secondmoment_mul_sqrt_real_measure a nonnegative `l2` function restricted to a measurable event is controlled by its exact second moment and the square root of the event's real mass. theorem compiled","shard":"modules/735ac88bb6fc3ec6.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeScheduledScalarRidgeConfidenceFailureSet","label":"allTimeScheduledScalarRidgeConfidenceFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeScheduledScalarRidgeConfidenceFailureSet","description":"Countable union of scalar-ridge confidence failures over every horizon.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-8bd1b32e2529","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5974,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:28"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def allTimeScheduledScalarRidgeConfidenceFailureSet {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R : Real) (deltaAt : Nat -> Real) : Set Omega","missing":[],"search":"alltimescheduledscalarridgeconfidencefailureset banditrlproof.oful.alltimescheduledscalarridgeconfidencefailureset countable union of scalar-ridge confidence failures over every horizon. definition compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","label":"mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","description":"Membership means failure at at least one deterministic horizon.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-dc2c88162b10","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5975,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:42"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R : Real) (deltaAt : Nat -> Real) (omega : Omega) : omega ∈ allTimeScheduledScalarRidgeConfidenceFailureSet lambda thetaStar S feature response R deltaAt ↔ ∃ n, omega ∈ scalarRidgeConfidenceFailureAt lambda thetaStar S feature response R (deltaAt n) n","missing":[],"search":"mem_alltimescheduledscalarridgeconfidencefailureset_iff banditrlproof.oful.mem_alltimescheduledscalarridgeconfidencefailureset_iff membership means failure at at least one deterministic horizon. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.not_mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","label":"not_mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.not_mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","description":"Outside the countable failure union, every deterministic-horizon confidence ellipsoid holds simultaneously.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-f255d051edf0","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5976,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:63"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem not_mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R : Real) (deltaAt : Nat -> Real) (omega : Omega) : omega ∉ allTimeScheduledScalarRidgeConfidenceFailureSet lambda thetaStar S feature response R deltaAt ↔ ∀ n, matrixNorm (Matrix.scalar Feature lambda + finiteHorizonFeatureGram feature n omega) (finiteHorizonRidgeEstimate (Matrix.scalar Feature lambda) feature response n omega - thetaStar) <= finiteHorizonScalarConfidenceRadius feature R (deltaAt n) lambda S n omega","missing":[],"search":"not_mem_alltimescheduledscalarridgeconfidencefailureset_iff banditrlproof.oful.not_mem_alltimescheduledscalarridgeconfidencefailureset_iff outside the countable failure union, every deterministic-horizon confidence ellipsoid holds simultaneously. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_allTimeScheduledScalarRidgeConfidenceFailureSet_le_tsum","label":"measure_allTimeScheduledScalarRidgeConfidenceFailureSet_le_tsum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_allTimeScheduledScalarRidgeConfidenceFailureSet_le_tsum","description":"All-time scheduled confidence: countable subadditivity bounds the failure probability by the `ENNReal` sum of the fixed-time budgets.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-6ae5949f34bc","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5977,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:92"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_allTimeScheduledScalarRidgeConfidenceFailureSet_le_tsum {Omega : Type u} {Feature : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (W…","missing":[],"search":"measure_alltimescheduledscalarridgeconfidencefailureset_le_tsum banditrlproof.oful.measure_alltimescheduledscalarridgeconfidencefailureset_le_tsum all-time scheduled confidence: countable subadditivity bounds the failure probability by the `ennreal` sum of the fixed-time budgets. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight","label":"allTimeTelescopingWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingWeight","description":"Positive telescoping weight `1 / ((n+1)(n+2))`.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-a93d2b30d923","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5978,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:149"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def allTimeTelescopingWeight (n : Nat) : Real","missing":[],"search":"alltimetelescopingweight banditrlproof.oful.alltimetelescopingweight positive telescoping weight `1 / ((n+1)(n+2))`. definition compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight_eq_sub","label":"allTimeTelescopingWeight_eq_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingWeight_eq_sub","description":"The telescoping weight is a difference of consecutive reciprocals.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-0b8f9792faeb","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5979,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:153"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingWeight_eq_sub (n : Nat) : allTimeTelescopingWeight n = 1 / (((n + 1 : Nat) : Real)) - 1 / (((n + 2 : Nat) : Real))","missing":[],"search":"alltimetelescopingweight_eq_sub banditrlproof.oful.alltimetelescopingweight_eq_sub the telescoping weight is a difference of consecutive reciprocals. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_allTimeTelescopingWeight","label":"sum_range_allTimeTelescopingWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_allTimeTelescopingWeight","description":"Exact finite partial sum of the telescoping weights.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-182a45ec7fe5","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5980,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:162"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_allTimeTelescopingWeight (n : Nat) : (Finset.range n).sum allTimeTelescopingWeight = 1 - 1 / (((n + 1 : Nat) : Real))","missing":[],"search":"sum_range_alltimetelescopingweight banditrlproof.oful.sum_range_alltimetelescopingweight exact finite partial sum of the telescoping weights. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight_nonneg","label":"allTimeTelescopingWeight_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingWeight_nonneg","description":"Every telescoping confidence weight is nonnegative.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-5b3621ac9018","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5981,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:171"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingWeight_nonneg (n : Nat) : 0 <= allTimeTelescopingWeight n","missing":[],"search":"alltimetelescopingweight_nonneg banditrlproof.oful.alltimetelescopingweight_nonneg every telescoping confidence weight is nonnegative. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight_le_one","label":"allTimeTelescopingWeight_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingWeight_le_one","description":"Every telescoping confidence weight is at most one.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-d64b290c590e","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5982,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:177"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingWeight_le_one (n : Nat) : allTimeTelescopingWeight n <= 1","missing":[],"search":"alltimetelescopingweight_le_one banditrlproof.oful.alltimetelescopingweight_le_one every telescoping confidence weight is at most one. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.hasSum_allTimeTelescopingWeight","label":"hasSum_allTimeTelescopingWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.hasSum_allTimeTelescopingWeight","description":"The telescoping weights sum exactly to one.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-abf7a5f369e9","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5983,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:197"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem hasSum_allTimeTelescopingWeight : HasSum allTimeTelescopingWeight 1","missing":[],"search":"hassum_alltimetelescopingweight banditrlproof.oful.hassum_alltimetelescopingweight the telescoping weights sum exactly to one. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta","label":"allTimeTelescopingDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingDelta","description":"Time-`n` share of an all-time failure budget.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-6cd3bde16549","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5984,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:215"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def allTimeTelescopingDelta (delta : Real) (n : Nat) : Real","missing":[],"search":"alltimetelescopingdelta banditrlproof.oful.alltimetelescopingdelta time-`n` share of an all-time failure budget. definition compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_eq_div","label":"allTimeTelescopingDelta_eq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingDelta_eq_div","description":"Display the scheduled budget as `delta / ((n+1)(n+2))`.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-56d53464bb1e","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5985,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:220"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingDelta_eq_div (delta : Real) (n : Nat) : allTimeTelescopingDelta delta n = delta / (((n + 1 : Nat) : Real) * ((n + 2 : Nat) : Real))","missing":[],"search":"alltimetelescopingdelta_eq_div banditrlproof.oful.alltimetelescopingdelta_eq_div display the scheduled budget as `delta / ((n+1)(n+2))`. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_pos","label":"allTimeTelescopingDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingDelta_pos","description":"Positive outer budget gives positive timewise budgets.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-b3f712c1ffa7","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5986,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:227"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingDelta_pos {delta : Real} (hdelta : 0 < delta) (n : Nat) : 0 < allTimeTelescopingDelta delta n","missing":[],"search":"alltimetelescopingdelta_pos banditrlproof.oful.alltimetelescopingdelta_pos positive outer budget gives positive timewise budgets. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_le_one","label":"allTimeTelescopingDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingDelta_le_one","description":"If `delta <= 1`, every timewise budget is also at most one.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-eaf4b6e42ae9","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5987,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:235"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingDelta_le_one {delta : Real} (hdelta : 0 < delta) (hdelta_one : delta <= 1) (n : Nat) : allTimeTelescopingDelta delta n <= 1","missing":[],"search":"alltimetelescopingdelta_le_one banditrlproof.oful.alltimetelescopingdelta_le_one if `delta <= 1`, every timewise budget is also at most one. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.tsum_ofReal_allTimeTelescopingDelta","label":"tsum_ofReal_allTimeTelescopingDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.tsum_ofReal_allTimeTelescopingDelta","description":"The `ENNReal` scheduled failure budgets sum exactly to `ofReal delta`.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-86cbd756151e","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5988,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:248"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem tsum_ofReal_allTimeTelescopingDelta {delta : Real} (hdelta : 0 <= delta) : ∑' n, ENNReal.ofReal (allTimeTelescopingDelta delta n) = ENNReal.ofReal delta","missing":[],"search":"tsum_ofreal_alltimetelescopingdelta banditrlproof.oful.tsum_ofreal_alltimetelescopingdelta the `ennreal` scheduled failure budgets sum exactly to `ofreal delta`. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingScalarRidgeConfidenceFailureSet","label":"allTimeTelescopingScalarRidgeConfidenceFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingScalarRidgeConfidenceFailureSet","description":"Failure set for the exact telescoping all-time confidence schedule.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-b08e0b33746e","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5989,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:278"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def allTimeTelescopingScalarRidgeConfidenceFailureSet {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R delta : Real) : Set Omega","missing":[],"search":"alltimetelescopingscalarridgeconfidencefailureset banditrlproof.oful.alltimetelescopingscalarridgeconfidencefailureset failure set for the exact telescoping all-time confidence schedule. definition compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","label":"measure_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","description":"All-time scalar-ridge confidence under the exact telescoping failure schedule. The result is one countable event on one process, not a family of horizon-dependent generated algorithms.","url":"../modules/banditrlproof-ofulalltimeconfidence/index.html#decl-bbfd5c5a7827","parent":"module:BanditRLProof.OFULAllTimeConfidence","order":5990,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULAllTimeConfidence"],["Source","BanditRLProof/OFULAllTimeConfidence.lean:296"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_allTimeTelescopingScalarRidgeConfidenceFailureSet_le {Omega : Type u} {Feature : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (With…","missing":[],"search":"measure_alltimetelescopingscalarridgeconfidencefailureset_le banditrlproof.oful.measure_alltimetelescopingscalarridgeconfidencefailureset_le all-time scalar-ridge confidence under the exact telescoping failure schedule. the result is one countable event on one process, not a family of horizon-dependent generated algorithms. theorem compiled","shard":"modules/acf5d661f53afac0.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_matrix_det_of_apply","label":"measurable_matrix_det_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_matrix_det_of_apply","description":"A determinant is measurable when every random matrix entry is measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-5c84f240948c","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5991,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_matrix_det_of_apply {Omega Index : Type*} [MeasurableSpace Omega] [Fintype Index] [DecidableEq Index] (A : Omega -> Matrix Index Index Real) (hA : forall i j, Measurable (fun omega => A omega i j)) : Measurable (fun omega => Matrix.det (A omega))","missing":[],"search":"measurable_matrix_det_of_apply banditrlproof.oful.measurable_matrix_det_of_apply a determinant is measurable when every random matrix entry is measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_matrix_adjugate_apply_of_apply","label":"measurable_matrix_adjugate_apply_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_matrix_adjugate_apply_of_apply","description":"Every adjugate entry is measurable under coordinatewise matrix measurability.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-76f08458a800","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5992,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:35"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_matrix_adjugate_apply_of_apply {Omega Index : Type*} [MeasurableSpace Omega] [Fintype Index] [DecidableEq Index] (A : Omega -> Matrix Index Index Real) (hA : forall i j, Measurable (fun omega => A omega i j)) (i j : Index) : Measurable (fun omega => Matrix.adjugate (A omega) i j)","missing":[],"search":"measurable_matrix_adjugate_apply_of_apply banditrlproof.oful.measurable_matrix_adjugate_apply_of_apply every adjugate entry is measurable under coordinatewise matrix measurability. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_matrix_nonsingInv_apply_of_apply","label":"measurable_matrix_nonsingInv_apply_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_matrix_nonsingInv_apply_of_apply","description":"Every entry of Mathlib's nonsingular inverse is measurable. The proof uses the determinant/adjugate formula, including its zero-at-singular convention.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-4a965e3690e5","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5993,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:57"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_matrix_nonsingInv_apply_of_apply {Omega Index : Type*} [MeasurableSpace Omega] [Fintype Index] [DecidableEq Index] (A : Omega -> Matrix Index Index Real) (hA : forall i j, Measurable (fun omega => A omega i j)) (i j : Index) : Measurable (fun omega => (A omega)⁻¹ i j)","missing":[],"search":"measurable_matrix_nonsinginv_apply_of_apply banditrlproof.oful.measurable_matrix_nonsinginv_apply_of_apply every entry of mathlib's nonsingular inverse is measurable. the proof uses the determinant/adjugate formula, including its zero-at-singular convention. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_matrix_mulVec_apply_of_apply","label":"measurable_matrix_mulVec_apply_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_matrix_mulVec_apply_of_apply","description":"A random matrix-vector product is coordinatewise measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-364e65025831","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5994,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:73"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_matrix_mulVec_apply_of_apply {Omega Index : Type*} [MeasurableSpace Omega] [Fintype Index] (A : Omega -> Matrix Index Index Real) (x : Omega -> Index -> Real) (hA : forall i j, Measurable (fun omega => A omega i j)) (hx : forall i, Measurable (fun omega => x omega i)) (i : Index) : Measurable (fun omega => (A omega).mulVec (x omega) i)","missing":[],"search":"measurable_matrix_mulvec_apply_of_apply banditrlproof.oful.measurable_matrix_mulvec_apply_of_apply a random matrix-vector product is coordinatewise measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_dotProduct_of_apply","label":"measurable_dotProduct_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_dotProduct_of_apply","description":"A random finite dot product is measurable coordinatewise.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-c54c21a8bfac","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5995,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:85"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_dotProduct_of_apply {Omega Index : Type*} [MeasurableSpace Omega] [Fintype Index] (x y : Omega -> Index -> Real) (hx : forall i, Measurable (fun omega => x omega i)) (hy : forall i, Measurable (fun omega => y omega i)) : Measurable (fun omega => dotProduct (x omega) (y omega))","missing":[],"search":"measurable_dotproduct_of_apply banditrlproof.oful.measurable_dotproduct_of_apply a random finite dot product is measurable coordinatewise. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_confidenceWidth_of_apply","label":"measurable_confidenceWidth_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_confidenceWidth_of_apply","description":"The OFUL confidence width is measurable from matrix and feature coordinates.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-56d1025ae4d8","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5996,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:95"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_confidenceWidth_of_apply {Omega Feature : Type*} [MeasurableSpace Omega] [Fintype Feature] [DecidableEq Feature] (V : Omega -> Matrix Feature Feature Real) (x : Omega -> Feature -> Real) (hV : forall i j, Measurable (fun omega => V omega i j)) (hx : forall i, Measurable (fun omega => x omega i)) : Measurable (fun omega => confidenceWidth (V omega) (x omega))","missing":[],"search":"measurable_confidencewidth_of_apply banditrlproof.oful.measurable_confidencewidth_of_apply the oful confidence width is measurable from matrix and feature coordinates. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_optimisticScore_of_apply","label":"measurable_optimisticScore_of_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_optimisticScore_of_apply","description":"The OFUL optimistic score is measurable from its scalar coordinates.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-674b3b40f71f","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5997,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:114"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimisticScore_of_apply {Omega Feature : Type*} [MeasurableSpace Omega] [Fintype Feature] [DecidableEq Feature] (thetaHat : Omega -> Feature -> Real) (V : Omega -> Matrix Feature Feature Real) (beta : Omega -> Real) (x : Omega -> Feature -> Real) (hthetaHat : forall i, Measurable (fun omega => thetaHat omega i)) (hV : forall i j, Measurable (fun omega => V omega i j)) (hbeta : Measurable beta) (hx : forall i, Measurable (fun omega => x omega i)) : Measurable (fun omega => optimisticScore (thetaHat omega) (V omega) (beta omega) (x omega))","missing":[],"search":"measurable_optimisticscore_of_apply banditrlproof.oful.measurable_optimisticscore_of_apply the oful optimistic score is measurable from its scalar coordinates. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonFeatureGram_apply","label":"measurable_finiteHorizonFeatureGram_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonFeatureGram_apply","description":"Entries of a finite-horizon feature Gram are measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-c51a68827848","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5998,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:134"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonFeatureGram_apply {Omega Feature : Type*} [MeasurableSpace Omega] (feature : Nat -> Omega -> Feature -> Real) (hfeature : forall t i, Measurable (fun omega => feature t omega i)) (n : Nat) (i j : Feature) : Measurable (fun omega => finiteHorizonFeatureGram feature n omega i j)","missing":[],"search":"measurable_finitehorizonfeaturegram_apply banditrlproof.oful.measurable_finitehorizonfeaturegram_apply entries of a finite-horizon feature gram are measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonResponseVector_apply","label":"measurable_finiteHorizonResponseVector_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonResponseVector_apply","description":"Coordinates of the finite-horizon response vector are measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-0be529ebc760","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":5999,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:147"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonResponseVector_apply {Omega Feature : Type*} [MeasurableSpace Omega] [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (hfeature : forall t i, Measurable (fun omega => feature t omega i)) (hresponse : forall t, Measurable (response t)) (n : Nat) (i : Feature) : Measurable (fun omega => finiteHorizonResponseVector feature response n omega i)","missing":[],"search":"measurable_finitehorizonresponsevector_apply banditrlproof.oful.measurable_finitehorizonresponsevector_apply coordinates of the finite-horizon response vector are measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonRidgeEstimate_apply","label":"measurable_finiteHorizonRidgeEstimate_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonRidgeEstimate_apply","description":"Every coordinate of the finite-horizon ridge estimate is measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-fbc4e76c3aba","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6000,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:163"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonRidgeEstimate_apply {Omega Feature : Type*} [MeasurableSpace Omega] [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (hfeature : forall t i, Measurable (fun omega => feature t omega i)) (hresponse : forall t, Measurable (response t)) (n : Nat) (i : Feature) : Measurable (fun omega => finiteHorizonRidgeEstimate V0 feature response n omega i)","missing":[],"search":"measurable_finitehorizonridgeestimate_apply banditrlproof.oful.measurable_finitehorizonridgeestimate_apply every coordinate of the finite-horizon ridge estimate is measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonScalarConfidenceRadius","label":"measurable_finiteHorizonScalarConfidenceRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonScalarConfidenceRadius","description":"The scalar-ridge finite-horizon confidence radius is measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-31838aa9f593","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6001,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:188"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonScalarConfidenceRadius {Omega Feature : Type*} [MeasurableSpace Omega] [Fintype Feature] [DecidableEq Feature] (feature : Nat -> Omega -> Feature -> Real) (hfeature : forall t i, Measurable (fun omega => feature t omega i)) (R delta lambda S : Real) (n : Nat) : Measurable (finiteHorizonScalarConfidenceRadius feature R delta lambda S n)","missing":[],"search":"measurable_finitehorizonscalarconfidenceradius banditrlproof.oful.measurable_finitehorizonscalarconfidenceradius the scalar-ridge finite-horizon confidence radius is measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryObservedFeature","label":"finiteHistoryObservedFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryObservedFeature","description":"Feature observed at a history coordinate, with zero outside the prefix.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-9cfa4bec280f","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6002,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:217"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryObservedFeature {K : Nat} {Feature : Type u} (actionFeature : Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (t : Nat) : Feature -> Real","missing":[],"search":"finitehistoryobservedfeature banditrlproof.oful.finitehistoryobservedfeature feature observed at a history coordinate, with zero outside the prefix. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryObservedResponse","label":"finiteHistoryObservedResponse","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryObservedResponse","description":"Reward observed at a history coordinate, with zero outside the prefix.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-1854624d3f86","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6003,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:228"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryObservedResponse {K : Nat} (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (t : Nat) : Real","missing":[],"search":"finitehistoryobservedresponse banditrlproof.oful.finitehistoryobservedresponse reward observed at a history coordinate, with zero outside the prefix. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryObservedFeature_apply","label":"measurable_finiteHistoryObservedFeature_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryObservedFeature_apply","description":"Every coordinate of the history-observed feature process is measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-b8cdd910b87b","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6004,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:238"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryObservedFeature_apply {K : Nat} {Feature : Type u} (actionFeature : Fin K -> Feature -> Real) (n t : Nat) (i : Feature) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => finiteHistoryObservedFeature actionFeature n history t i)","missing":[],"search":"measurable_finitehistoryobservedfeature_apply banditrlproof.oful.measurable_finitehistoryobservedfeature_apply every coordinate of the history-observed feature process is measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryObservedResponse","label":"measurable_finiteHistoryObservedResponse","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryObservedResponse","description":"The history-observed response process is measurable.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-07fc37823ba8","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6005,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:255"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryObservedResponse {K : Nat} (n t : Nat) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => finiteHistoryObservedResponse n history t)","missing":[],"search":"measurable_finitehistoryobservedresponse banditrlproof.oful.measurable_finitehistoryobservedresponse the history-observed response process is measurable. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeEstimate","label":"finiteHistoryScalarRidgeEstimate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeEstimate","description":"Scalar-ridge estimate reconstructed from the inclusive finite history.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-f9a961cd5b04","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6006,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:268"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeEstimate {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Feature -> Real","missing":[],"search":"finitehistoryscalarridgeestimate banditrlproof.oful.finitehistoryscalarridgeestimate scalar-ridge estimate reconstructed from the inclusive finite history. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeDesign","label":"finiteHistoryScalarRidgeDesign","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeDesign","description":"Scalar-ridge design matrix reconstructed from the inclusive finite history.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-d72d8de91c55","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6007,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:284"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeDesign {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Matrix Feature Feature Real","missing":[],"search":"finitehistoryscalarridgedesign banditrlproof.oful.finitehistoryscalarridgedesign scalar-ridge design matrix reconstructed from the inclusive finite history. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeRadius","label":"finiteHistoryScalarRidgeRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeRadius","description":"Scalar confidence radius reconstructed from the inclusive finite history.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-c49f0a46c076","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6008,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:298"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeRadius {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (actionFeature : Fin K -> Feature -> Real) (R delta lambda S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Real","missing":[],"search":"finitehistoryscalarridgeradius banditrlproof.oful.finitehistoryscalarridgeradius scalar confidence radius reconstructed from the inclusive finite history. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryFixedActionFeature","label":"finiteHistoryFixedActionFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryFixedActionFeature","description":"The fixed feature of an arm, exposed on the finite-history component API.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-644a96b5bfdf","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6009,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:311"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def finiteHistoryFixedActionFeature {K : Nat} {Feature : Type u} (actionFeature : Fin K -> Feature -> Real) (n : Nat) (_history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : Feature -> Real","missing":[],"search":"finitehistoryfixedactionfeature banditrlproof.oful.finitehistoryfixedactionfeature the fixed feature of an arm, exposed on the finite-history component api. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticScore","label":"finiteHistoryScalarRidgeOptimisticScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticScore","description":"Concrete scalar-ridge OFUL score computed from an inclusive finite history.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-87840a0e8ef9","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6010,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:319"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeOptimisticScore {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : Real","missing":[],"search":"finitehistoryscalarridgeoptimisticscore banditrlproof.oful.finitehistoryscalarridgeoptimisticscore concrete scalar-ridge oful score computed from an inclusive finite history. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScalarRidgeOptimisticScore","label":"measurable_finiteHistoryScalarRidgeOptimisticScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryScalarRidgeOptimisticScore","description":"Every fixed-arm concrete scalar-ridge score is measurable. In particular, callers do not need to supply a score-measurability premise.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-649b1e7e5103","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6011,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:338"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryScalarRidgeOptimisticScore {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (action : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => finiteHistoryScalarRidgeOptimisticScore lambda actionFeature R delta S n history action)","missing":[],"search":"measurable_finitehistoryscalarridgeoptimisticscore banditrlproof.oful.measurable_finitehistoryscalarridgeoptimisticscore every fixed-arm concrete scalar-ridge score is measurable. in particular, callers do not need to supply a score-measurability premise. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction","label":"finiteHistoryScalarRidgeOptimisticAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction","description":"Deterministic score-maximizing arm for the concrete scalar-ridge state.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-30a7468fcb8d","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6012,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:386"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"finitehistoryscalarridgeoptimisticaction banditrlproof.oful.finitehistoryscalarridgeoptimisticaction deterministic score-maximizing arm for the concrete scalar-ridge state. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction_score_max","label":"finiteHistoryScalarRidgeOptimisticAction_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction_score_max","description":"The concrete finite-history selector maximizes its scalar-ridge score.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-8fb9713e9dfc","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6013,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:404"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScalarRidgeOptimisticAction_score_max {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : finiteHistoryScalarRidgeOptimisticScore lambda actionFeature R delta S n history action <= finiteHistoryScalarRidgeOptimisticScore lambda actionFeature R delta S n history (finiteHistoryScalarRidgeOptimisticAction hK lambda actionFeature R delta S n history)","missing":[],"search":"finitehistoryscalarridgeoptimisticaction_score_max banditrlproof.oful.finitehistoryscalarridgeoptimisticaction_score_max the concrete finite-history selector maximizes its scalar-ridge score. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAlgorithm","label":"finiteHistoryScalarRidgeOptimisticAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAlgorithm","description":"Concrete measurable history algorithm obtained from the reconstructed scalar-ridge state, with score measurability discharged internally.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-32eb58d852cd","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6014,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:431"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeOptimisticAlgorithm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) : Thompson.HistoryAlgorithm (Fin K) Real","missing":[],"search":"finitehistoryscalarridgeoptimisticalgorithm banditrlproof.oful.finitehistoryscalarridgeoptimisticalgorithm concrete measurable history algorithm obtained from the reconstructed scalar-ridge state, with score measurability discharged internally. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeSelectedFeature","label":"finiteHistoryScalarRidgeSelectedFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeSelectedFeature","description":"Feature selected by the concrete scalar-ridge history selector.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-94560445cf21","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6015,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:449"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScalarRidgeSelectedFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Feature -> Real","missing":[],"search":"finitehistoryscalarridgeselectedfeature banditrlproof.oful.finitehistoryscalarridgeselectedfeature feature selected by the concrete scalar-ridge history selector. definition compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_observedFeature_succ_ae_eq_finiteHistoryScalarRidgeSelectedFeature","label":"canonicalHistoryTrajectory_observedFeature_succ_ae_eq_finiteHistoryScalarRidgeSelectedFeature","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_observedFeature_succ_ae_eq_finiteHistoryScalarRidgeSelectedFeature","description":"On the canonical recursive trajectory, the feature of the actual successor arm is almost surely the feature selected from the realized scalar-ridge state.","url":"../modules/banditrlproof-ofulconcretehistoryridgeselection/index.html#decl-c8b3b148ba40","parent":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","order":6016,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConcreteHistoryRidgeSelection"],["Source","BanditRLProof/OFULConcreteHistoryRidgeSelection.lean:466"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_observedFeature_succ_ae_eq_finiteHistoryScalarRidgeSelectedFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory (n + 1)) = finiteHistoryScalarRidgeSelectedFeature hK lambda actionFeature R delta S n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n)","missing":[],"search":"canonicalhistorytrajectory_observedfeature_succ_ae_eq_finitehistoryscalarridgeselectedfeature banditrlproof.oful.canonicalhistorytrajectory_observedfeature_succ_ae_eq_finitehistoryscalarridgeselectedfeature on the canonical recursive trajectory, the feature of the actual successor arm is almost surely the feature selected from the realized scalar-ridge state. theorem compiled","shard":"modules/9372d8d027df328b.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.matrixNorm","label":"matrixNorm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.matrixNorm","description":"Quadratic-form square root; it is a norm when the matrix is positive definite.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-51b80aea2543","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6017,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def matrixNorm [Fintype Feature] (A : Matrix Feature Feature Real) (x : Feature -> Real) : Real","missing":[],"search":"matrixnorm banditrlproof.oful.matrixnorm quadratic-form square root; it is a norm when the matrix is positive definite. definition compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.matrixNorm_add_le","label":"matrixNorm_add_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.matrixNorm_add_le","description":"Triangle inequality for the norm induced by a positive-definite matrix.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-1344005ccfde","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6018,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:26"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem matrixNorm_add_le [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : A.PosDef) (x y : Feature -> Real) : matrixNorm A (x + y) <= matrixNorm A x + matrixNorm A y","missing":[],"search":"matrixnorm_add_le banditrlproof.oful.matrixnorm_add_le triangle inequality for the norm induced by a positive-definite matrix. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.matrixNorm_sub_le","label":"matrixNorm_sub_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.matrixNorm_sub_le","description":"Triangle inequality for subtraction in a positive-definite matrix norm.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-3e6c8514d7e4","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6019,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:43"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem matrixNorm_sub_le [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : A.PosDef) (x y : Feature -> Real) : matrixNorm A (x - y) <= matrixNorm A x + matrixNorm A y","missing":[],"search":"matrixnorm_sub_le banditrlproof.oful.matrixnorm_sub_le triangle inequality for subtraction in a positive-definite matrix norm. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.matrixNorm_sq","label":"matrixNorm_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.matrixNorm_sq","description":"Squaring the positive-definite matrix norm recovers its quadratic form.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-25521b7aa9de","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6020,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:60"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem matrixNorm_sq [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : A.PosDef) (x : Feature -> Real) : matrixNorm A x ^ 2 = x ⬝ᵥ A.mulVec x","missing":[],"search":"matrixnorm_sq banditrlproof.oful.matrixnorm_sq squaring the positive-definite matrix norm recovers its quadratic form. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonResponseVector","label":"finiteHorizonResponseVector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonResponseVector","description":"The finite-horizon sufficient statistic `sum_{i<n} x_i y_i`.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-37b6f3069bd5","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6021,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:68"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonResponseVector [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (n : Nat) (omega : Omega) : Feature -> Real","missing":[],"search":"finitehorizonresponsevector banditrlproof.oful.finitehorizonresponsevector the finite-horizon sufficient statistic `sum_{i<n} x_i y_i`. definition compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonRidgeEstimate","label":"finiteHorizonRidgeEstimate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonRidgeEstimate","description":"Ridge least-squares estimate with deterministic positive-definite base.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-1bda6d80155d","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6022,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:77"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonRidgeEstimate [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (n : Nat) (omega : Omega) : Feature -> Real","missing":[],"search":"finitehorizonridgeestimate banditrlproof.oful.finitehorizonridgeestimate ridge least-squares estimate with deterministic positive-definite base. definition compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonConfidenceThreshold","label":"finiteHorizonConfidenceThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonConfidenceThreshold","description":"The squared self-normalized radius from the common-`R` Markov tail.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-0ac7df49290c","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6023,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:87"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonConfidenceThreshold [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (feature : Nat -> Omega -> Feature -> Real) (R delta : Real) (n : Nat) (omega : Omega) : Real","missing":[],"search":"finitehorizonconfidencethreshold banditrlproof.oful.finitehorizonconfidencethreshold the squared self-normalized radius from the common-`r` markov tail. definition compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonConfidenceRadius","label":"finiteHorizonConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonConfidenceRadius","description":"Confidence radius with an explicit deterministic regularization-bias cap.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-6bf514e2a2d5","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6024,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:100"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonConfidenceRadius [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (feature : Nat -> Omega -> Feature -> Real) (R delta biasRadius : Real) (n : Nat) (omega : Omega) : Real","missing":[],"search":"finitehorizonconfidenceradius banditrlproof.oful.finitehorizonconfidenceradius confidence radius with an explicit deterministic regularization-bias cap. definition compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.one_le_det_add_posSemidef_div_det","label":"one_le_det_add_posSemidef_div_det","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.one_le_det_add_posSemidef_div_det","description":"Adding a positive-semidefinite Gram to a positive-definite base cannot reduce the determinant ratio below one.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-d357585d008c","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6025,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:113"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem one_le_det_add_posSemidef_div_det [Fintype Feature] [DecidableEq Feature] (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) : 1 <= Matrix.det (V0 + G) / Matrix.det V0","missing":[],"search":"one_le_det_add_possemidef_div_det banditrlproof.oful.one_le_det_add_possemidef_div_det adding a positive-semidefinite gram to a positive-definite base cannot reduce the determinant ratio below one. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonConfidenceThreshold_nonneg","label":"finiteHorizonConfidenceThreshold_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonConfidenceThreshold_nonneg","description":"The squared confidence threshold is nonnegative for `0 < delta <= 1`.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-900ce8351a6e","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6026,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:130"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonConfidenceThreshold_nonneg [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (feature : Nat -> Omega -> Feature -> Real) (R delta : Real) (n : Nat) (omega : Omega) (hdelta : 0 < delta) (hdelta_one : delta <= 1) : 0 <= finiteHorizonConfidenceThreshold V0 feature R delta n omega","missing":[],"search":"finitehorizonconfidencethreshold_nonneg banditrlproof.oful.finitehorizonconfidencethreshold_nonneg the squared confidence threshold is nonnegative for `0 < delta <= 1`. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonResponseVector_eq_featureGram_mulVec_add_noiseScore","label":"finiteHorizonResponseVector_eq_featureGram_mulVec_add_noiseScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonResponseVector_eq_featureGram_mulVec_add_noiseScore","description":"Under the linear observation model, the response sufficient statistic is the feature Gram applied to the true parameter plus the martingale-noise score.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-a83a25970cf4","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6027,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:162"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonResponseVector_eq_featureGram_mulVec_add_noiseScore [Fintype Feature] (thetaStar : Feature -> Real) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (n : Nat) (omega : Omega) (hresponse : forall i, i < n -> response i omega = dotProduct thetaStar (feature i omega) + noise i omega) : finiteHorizonResponseVector feature response n omega = (finiteHorizonFeatureGram feature n omega).mulVec thetaStar + WithLp.ofLp (finiteHorizonNoiseScore feature noise n omega)","missing":[],"search":"finitehorizonresponsevector_eq_featuregram_mulvec_add_noisescore banditrlproof.oful.finitehorizonresponsevector_eq_featuregram_mulvec_add_noisescore under the linear observation model, the response sufficient statistic is the feature gram applied to the true parameter plus the martingale-noise score. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.posDef_nonsingInv_mulVec_mulVec","label":"posDef_nonsingInv_mulVec_mulVec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.posDef_nonsingInv_mulVec_mulVec","description":"Applying the nonsingular inverse of a positive-definite matrix cancels it.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-6bc61966bbf7","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6028,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:197"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem posDef_nonsingInv_mulVec_mulVec [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : A.PosDef) (x : Feature -> Real) : A⁻¹.mulVec (A.mulVec x) = x","missing":[],"search":"posdef_nonsinginv_mulvec_mulvec banditrlproof.oful.posdef_nonsinginv_mulvec_mulvec applying the nonsingular inverse of a positive-definite matrix cancels it. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.posDef_mulVec_nonsingInv_mulVec","label":"posDef_mulVec_nonsingInv_mulVec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.posDef_mulVec_nonsingInv_mulVec","description":"A positive-definite matrix also cancels its inverse on the left.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-5abe1be584dd","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6029,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:207"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem posDef_mulVec_nonsingInv_mulVec [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : A.PosDef) (x : Feature -> Real) : A.mulVec (A⁻¹.mulVec x) = x","missing":[],"search":"posdef_mulvec_nonsinginv_mulvec banditrlproof.oful.posdef_mulvec_nonsinginv_mulvec a positive-definite matrix also cancels its inverse on the left. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.matrixNorm_nonsingInv_mulVec_sq","label":"matrixNorm_nonsingInv_mulVec_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.matrixNorm_nonsingInv_mulVec_sq","description":"The squared `A`-norm of `A⁻¹x` is the inverse quadratic form of `x`.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-1c646b41c356","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6030,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:219"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem matrixNorm_nonsingInv_mulVec_sq [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : A.PosDef) (x : Feature -> Real) : matrixNorm A (A⁻¹.mulVec x) ^ 2 = x ⬝ᵥ A⁻¹.mulVec x","missing":[],"search":"matrixnorm_nonsinginv_mulvec_sq banditrlproof.oful.matrixnorm_nonsinginv_mulvec_sq the squared `a`-norm of `a⁻¹x` is the inverse quadratic form of `x`. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonRidgeEstimate_sub_eq_inverseScore_sub_inverseBias","label":"finiteHorizonRidgeEstimate_sub_eq_inverseScore_sub_inverseBias","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonRidgeEstimate_sub_eq_inverseScore_sub_inverseBias","description":"Exact ridge-estimation error decomposition into inverse-Gram noise score and inverse-Gram regularization bias.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-b0805b425f7d","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6031,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:232"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonRidgeEstimate_sub_eq_inverseScore_sub_inverseBias [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (thetaStar : Feature -> Real) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (n : Nat) (omega : Omega) (hresponse : forall i, i < n -> response i omega = dotProduct thetaStar (feature i omega) + noise i omega) : finiteHorizonRidgeEstimate V0 feature response n omega - thetaStar = (V0 + finiteHorizonFeatureGram feature n omega)⁻¹.mulVec (WithLp.ofLp (finiteHorizonNoiseScore feature noise n omega)) - (V0 + finiteHorizonFeatureGram feature n omega)⁻¹.mulVec (V0.mulVec thetaStar)","missing":[],"search":"finitehorizonridgeestimate_sub_eq_inversescore_sub_inversebias banditrlproof.oful.finitehorizonridgeestimate_sub_eq_inversescore_sub_inversebias exact ridge-estimation error decomposition into inverse-gram noise score and inverse-gram regularization bias. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonRidgeEstimate_error_matrixNorm_le_score_add_bias","label":"finiteHorizonRidgeEstimate_error_matrixNorm_le_score_add_bias","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonRidgeEstimate_error_matrixNorm_le_score_add_bias","description":"Deterministic confidence decomposition: estimator error is bounded by the self-normalized score plus the regularization-bias norm.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-e80266e8b099","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6032,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:294"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonRidgeEstimate_error_matrixNorm_le_score_add_bias [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (thetaStar : Feature -> Real) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (n : Nat) (omega : Omega) (hresponse : forall i, i < n -> response i omega = dotProduct thetaStar (feature i omega) + noise i omega) : matrixNorm (V0 + finiteHorizonFeatureGram feature n omega) (finiteHorizonRidgeEstimate V0 feature response n omega - thetaStar) <= matrixNorm (V0 + finiteHorizonFeatureGram feature n omega) ((V0 + finiteHorizonFeatureGram feature n omega)⁻¹.mulVec (WithLp.ofLp (finiteHorizonNoiseScore feature noise n omega))) + matrixNorm (V0 + finiteHorizonFeatureGram feature n omega) ((V0 + finiteHorizonFeatureGram feature n omega)⁻¹.mulVec (V0.mulVec thetaStar))","missing":[],"search":"finitehorizonridgeestimate_error_matrixnorm_le_score_add_bias banditrlproof.oful.finitehorizonridgeestimate_error_matrixnorm_le_score_add_bias deterministic confidence decomposition: estimator error is bounded by the self-normalized score plus the regularization-bias norm. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.selfNormalizedQuadratic_gt_of_ridgeEstimate_error_matrixNorm_gt_confidenceRadius","label":"selfNormalizedQuadratic_gt_of_ridgeEstimate_error_matrixNorm_gt_confidenceRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.selfNormalizedQuadratic_gt_of_ridgeEstimate_error_matrixNorm_gt_confidenceRadius","description":"Pointwise transport from confidence-ellipsoid failure to the compiled self-normalized score bad event.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-899d7e78810a","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6033,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:325"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem selfNormalizedQuadratic_gt_of_ridgeEstimate_error_matrixNorm_gt_confidenceRadius [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (thetaStar : Feature -> Real) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R delta biasRadius : Real) (n : Nat) (omega : Omega) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (hresponse : forall i, i < n -> response i omega = dotProduct thetaStar (feature i omega) + noise i omega) (hbias : matrixNorm (V0 + finiteHorizonFeatureGram feature n omega) ((V0 + finiteHorizonFeatureGram feature n omega)⁻¹.mulVec (V0.mulVec thetaStar)) <= biasRadius) (hbad : matrixNorm (V0 + finiteHorizonFeatureGram feature n omega) (finiteHorizonRidgeEstimate V0 feature response n omega - thetaStar) > finiteHorizonConfidenceRadius V0 feature R delta biasRadius n omega) : (finiteHorizonNoiseSco…","missing":[],"search":"selfnormalizedquadratic_gt_of_ridgeestimate_error_matrixnorm_gt_confidenceradius banditrlproof.oful.selfnormalizedquadratic_gt_of_ridgeestimate_error_matrixnorm_gt_confidenceradius pointwise transport from confidence-ellipsoid failure to the compiled self-normalized score bad event. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","label":"measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","description":"Finite-horizon OFUL least-squares confidence ellipsoid. With probability at least `1 - delta`, the ridge-estimation error lies in the current Gram norm inside the self-normalized radius plus the supplied regularization-bias cap.","url":"../modules/banditrlproof-ofulconfidenceellipsoid/index.html#decl-4e17fa3cd9a3","parent":"module:BanditRLProof.OFULConfidenceEllipsoid","order":6034,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULConfidenceEllipsoid"],["Source","BanditRLProof/OFULConfidenceEllipsoid.lean:401"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (thetaStar : Feature -> Real) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i)…","missing":[],"search":"measure_finitehorizonridgeestimate_error_matrixnorm_gt_confidenceradius_le banditrlproof.oful.measure_finitehorizonridgeestimate_error_matrixnorm_gt_confidenceradius_le finite-horizon oful least-squares confidence ellipsoid. with probability at least `1 - delta`, the ridge-estimation error lies in the current gram norm inside the self-normalized radius plus the supplied regularization-bias cap. theorem compiled","shard":"modules/2e6fae67537e0551.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.rankOneGram","label":"rankOneGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.rankOneGram","description":"Rank-one Gram matrix generated by one feature vector.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-70dafc2e2fa6","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6035,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:29"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def rankOneGram {Feature : Type u} (x : Feature -> Real) : Matrix Feature Feature Real","missing":[],"search":"rankonegram banditrlproof.oful.rankonegram rank-one gram matrix generated by one feature vector. definition compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.rankOneGram_eq_replicateCol_mul_replicateRow","label":"rankOneGram_eq_replicateCol_mul_replicateRow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.rankOneGram_eq_replicateCol_mul_replicateRow","description":"The local rank-one Gram matrix is Mathlib's column-row product shape.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-ce11fa536691","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6036,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:34"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem rankOneGram_eq_replicateCol_mul_replicateRow {Feature : Type u} [Fintype Feature] (x : Feature -> Real) : rankOneGram x = Matrix.replicateCol Unit x * Matrix.replicateRow Unit x","missing":[],"search":"rankonegram_eq_replicatecol_mul_replicaterow banditrlproof.oful.rankonegram_eq_replicatecol_mul_replicaterow the local rank-one gram matrix is mathlib's column-row product shape. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.rankOneGram_isHermitian","label":"rankOneGram_isHermitian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.rankOneGram_isHermitian","description":"Rank-one Gram matrices are Hermitian.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-b6fc357cbec4","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6037,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:43"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem rankOneGram_isHermitian {Feature : Type u} [Fintype Feature] (x : Feature -> Real) : (rankOneGram x).IsHermitian","missing":[],"search":"rankonegram_ishermitian banditrlproof.oful.rankonegram_ishermitian rank-one gram matrices are hermitian. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_one_add_rankOneGram","label":"det_one_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_one_add_rankOneGram","description":"Rank-one determinant lemma at the identity, specialized to feature Gram updates.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-fa81e804ef74","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6038,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:55"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_one_add_rankOneGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (x : Feature -> Real) : (1 + rankOneGram x).det = 1 + dotProduct x x","missing":[],"search":"det_one_add_rankonegram banditrlproof.oful.det_one_add_rankonegram rank-one determinant lemma at the identity, specialized to feature gram updates. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_rankOne_update_factor_eq_one_add_dotProduct_inv_mulVec","label":"det_rankOne_update_factor_eq_one_add_dotProduct_inv_mulVec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_rankOne_update_factor_eq_one_add_dotProduct_inv_mulVec","description":"The one-dimensional determinant factor in the rank-one matrix determinant lemma is the scalar `1 + x^T A^{-1} x`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-6aa4e82bb2b5","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6039,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:68"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_rankOne_update_factor_eq_one_add_dotProduct_inv_mulVec {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (x : Feature -> Real) : (1 + Matrix.replicateRow Unit x * A⁻¹ * Matrix.replicateCol Unit x).det = 1 + dotProduct x (Matrix.mulVec (A⁻¹) x)","missing":[],"search":"det_rankone_update_factor_eq_one_add_dotproduct_inv_mulvec banditrlproof.oful.det_rankone_update_factor_eq_one_add_dotproduct_inv_mulvec the one-dimensional determinant factor in the rank-one matrix determinant lemma is the scalar `1 + x^t a^{-1} x`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_add_rankOneGram","label":"det_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_add_rankOneGram","description":"Rank-one matrix determinant lemma in OFUL feature notation. The invertibility side condition is kept as Mathlib's `IsUnit A.det`, matching the upstream Schur-complement API.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-50bd5504254d","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6040,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:90"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_add_rankOneGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (A : Matrix Feature Feature Real) (hA : IsUnit A.det) (x : Feature -> Real) : (A + rankOneGram x).det = A.det * (1 + dotProduct x (Matrix.mulVec (A⁻¹) x))","missing":[],"search":"det_add_rankonegram banditrlproof.oful.det_add_rankonegram rank-one matrix determinant lemma in oful feature notation. the invertibility side condition is kept as mathlib's `isunit a.det`, matching the upstream schur-complement api. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_scalar_identity","label":"det_scalar_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_scalar_identity","description":"Determinant of the scalar regularization matrix `lambda I`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-3ec861a33f39","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6041,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:102"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_scalar_identity {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) : (Matrix.scalar Feature lambda : Matrix Feature Feature Real).det = lambda ^ Fintype.card Feature","missing":[],"search":"det_scalar_identity banditrlproof.oful.det_scalar_identity determinant of the scalar regularization matrix `lambda i`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_scalar_identity_ne_zero","label":"det_scalar_identity_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_scalar_identity_ne_zero","description":"Nonzero scalar regularization has nonzero determinant.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-3526a0c47cb4","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6042,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:112"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_scalar_identity_ne_zero {Feature : Type u} [Fintype Feature] [DecidableEq Feature] {lambda : Real} (hlambda : lambda ≠ 0) : (Matrix.scalar Feature lambda : Matrix Feature Feature Real).det ≠ 0","missing":[],"search":"det_scalar_identity_ne_zero banditrlproof.oful.det_scalar_identity_ne_zero nonzero scalar regularization has nonzero determinant. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.isUnit_det_scalar_identity","label":"isUnit_det_scalar_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.isUnit_det_scalar_identity","description":"Nonzero scalar regularization satisfies Mathlib's determinant-unit side condition for the rank-one determinant update lemma.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-d22703c848ad","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6043,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:123"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem isUnit_det_scalar_identity {Feature : Type u} [Fintype Feature] [DecidableEq Feature] {lambda : Real} (hlambda : lambda ≠ 0) : IsUnit (Matrix.scalar Feature lambda : Matrix Feature Feature Real).det","missing":[],"search":"isunit_det_scalar_identity banditrlproof.oful.isunit_det_scalar_identity nonzero scalar regularization satisfies mathlib's determinant-unit side condition for the rank-one determinant update lemma. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_scalar_add_rankOneGram","label":"det_scalar_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_scalar_add_rankOneGram","description":"First determinant update from the regularized scalar base `lambda I`. This is the one-step OFUL/LinUCB determinant recursion specialized to the empty-history base matrix. General log-det telescoping and determinant-growth bounds remain separate leaves.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-afd1e399816b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6044,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:136"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_scalar_add_rankOneGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] {lambda : Real} (hlambda : lambda ≠ 0) (x : Feature -> Real) : ((Matrix.scalar Feature lambda : Matrix Feature Feature Real) + rankOneGram x).det = (Matrix.scalar Feature lambda : Matrix Feature Feature Real).det * (1 + dotProduct x (Matrix.mulVec ((Matrix.scalar Feature lambda : Matrix Feature Feature Real)⁻¹) x))","missing":[],"search":"det_scalar_add_rankonegram banditrlproof.oful.det_scalar_add_rankonegram first determinant update from the regularized scalar base `lambda i`. this is the one-step oful/linucb determinant recursion specialized to the empty-history base matrix. general log-det telescoping and determinant-growth bounds remain separate leaves. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.quadraticForm","label":"quadraticForm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.quadraticForm","description":"Quadratic form associated with a finite real matrix.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-3eb8490d212e","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6045,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:151"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def quadraticForm {Feature : Type u} [Fintype Feature] (A : Matrix Feature Feature Real) (y : Feature -> Real) : Real","missing":[],"search":"quadraticform banditrlproof.oful.quadraticform quadratic form associated with a finite real matrix. definition compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.quadraticForm_add","label":"quadraticForm_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.quadraticForm_add","description":"Quadratic forms distribute over matrix addition.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-42991a94191e","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6046,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:158"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem quadraticForm_add {Feature : Type u} [Fintype Feature] (A B : Matrix Feature Feature Real) (y : Feature -> Real) : quadraticForm (A + B) y = quadraticForm A y + quadraticForm B y","missing":[],"search":"quadraticform_add banditrlproof.oful.quadraticform_add quadratic forms distribute over matrix addition. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.dotProduct_mulVec_eq_quadraticForm","label":"dotProduct_mulVec_eq_quadraticForm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.dotProduct_mulVec_eq_quadraticForm","description":"The Mathlib dot-product/mulVec form agrees with the local quadratic form.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-2f22897e722b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6047,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:195"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem dotProduct_mulVec_eq_quadraticForm {Feature : Type u} [Fintype Feature] (A : Matrix Feature Feature Real) (y : Feature -> Real) : dotProduct y (Matrix.mulVec A y) = quadraticForm A y","missing":[],"search":"dotproduct_mulvec_eq_quadraticform banditrlproof.oful.dotproduct_mulvec_eq_quadraticform the mathlib dot-product/mulvec form agrees with the local quadratic form. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.quadraticForm_scalar_identity","label":"quadraticForm_scalar_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.quadraticForm_scalar_identity","description":"Quadratic form of the scalar regularization matrix `lambda I`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-26c6fbf23736","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6048,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:202"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem quadraticForm_scalar_identity {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (y : Feature -> Real) : quadraticForm (Matrix.scalar Feature lambda : Matrix Feature Feature Real) y = lambda * (Finset.univ : Finset Feature).sum (fun i => y i ^ 2)","missing":[],"search":"quadraticform_scalar_identity banditrlproof.oful.quadraticform_scalar_identity quadratic form of the scalar regularization matrix `lambda i`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_sq_pos_of_exists_ne_zero","label":"sum_sq_pos_of_exists_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_sq_pos_of_exists_ne_zero","description":"A finite sum of coordinate squares is positive for a nonzero vector.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-7e91ce0061ef","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6049,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:224"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_sq_pos_of_exists_ne_zero {Feature : Type u} [Fintype Feature] (y : Feature -> Real) (hy : ∃ i : Feature, y i ≠ 0) : 0 < (Finset.univ : Finset Feature).sum (fun i => y i ^ 2)","missing":[],"search":"sum_sq_pos_of_exists_ne_zero banditrlproof.oful.sum_sq_pos_of_exists_ne_zero a finite sum of coordinate squares is positive for a nonzero vector. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.exists_coord_ne_zero_of_ne_zero","label":"exists_coord_ne_zero_of_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.exists_coord_ne_zero_of_ne_zero","description":"A nonzero vector has a nonzero coordinate.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-6f98f9920973","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6050,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:236"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem exists_coord_ne_zero_of_ne_zero {Feature : Type u} (y : Feature -> Real) (hy : y ≠ 0) : ∃ i : Feature, y i ≠ 0","missing":[],"search":"exists_coord_ne_zero_of_ne_zero banditrlproof.oful.exists_coord_ne_zero_of_ne_zero a nonzero vector has a nonzero coordinate. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.featureGram","label":"featureGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.featureGram","description":"Finite-history feature Gram matrix `sum_t x_t x_t^T`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-7de48710333c","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6051,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:246"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def featureGram {Time : Type v} {Feature : Type u} [Fintype Time] (x : Time -> Feature -> Real) : Matrix Feature Feature Real","missing":[],"search":"featuregram banditrlproof.oful.featuregram finite-history feature gram matrix `sum_t x_t x_t^t`. definition compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.featureGram_isHermitian","label":"featureGram_isHermitian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.featureGram_isHermitian","description":"Finite-history feature Gram matrices are Hermitian.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-cbf50cc5a255","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6052,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:253"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem featureGram_isHermitian {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] (x : Time -> Feature -> Real) : (featureGram x).IsHermitian","missing":[],"search":"featuregram_ishermitian banditrlproof.oful.featuregram_ishermitian finite-history feature gram matrices are hermitian. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram","label":"regularizedFeatureGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram","description":"Regularized finite-history Gram matrix `lambda I + sum_t x_t x_t^T`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c21c3735f658","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6053,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:263"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def regularizedFeatureGram {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Time -> Feature -> Real) : Matrix Feature Feature Real","missing":[],"search":"regularizedfeaturegram banditrlproof.oful.regularizedfeaturegram regularized finite-history gram matrix `lambda i + sum_t x_t x_t^t`. definition compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_eq_scalar_add_featureGram","label":"regularizedFeatureGram_eq_scalar_add_featureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_eq_scalar_add_featureGram","description":"The named regularized Gram unfolds to scalar regularization plus Gram.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c47141313e0d","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6054,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:271"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_eq_scalar_add_featureGram {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Time -> Feature -> Real) : regularizedFeatureGram lambda x = Matrix.scalar Feature lambda + featureGram x","missing":[],"search":"regularizedfeaturegram_eq_scalar_add_featuregram banditrlproof.oful.regularizedfeaturegram_eq_scalar_add_featuregram the named regularized gram unfolds to scalar regularization plus gram. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_isHermitian","label":"regularizedFeatureGram_isHermitian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_isHermitian","description":"Regularized finite-history feature Gram matrices are Hermitian.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1715dad6cdb3","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6055,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:280"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_isHermitian {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Time -> Feature -> Real) : (regularizedFeatureGram lambda x).IsHermitian","missing":[],"search":"regularizedfeaturegram_ishermitian banditrlproof.oful.regularizedfeaturegram_ishermitian regularized finite-history feature gram matrices are hermitian. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prefixFeatureGram","label":"prefixFeatureGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.prefixFeatureGram","description":"Prefix Gram matrix `sum_{t < T} x_t x_t^T` for a Nat-indexed history.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-e6d586432e22","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6056,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:291"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def prefixFeatureGram {Feature : Type u} (x : Nat -> Feature -> Real) (T : Nat) : Matrix Feature Feature Real","missing":[],"search":"prefixfeaturegram banditrlproof.oful.prefixfeaturegram prefix gram matrix `sum_{t < t} x_t x_t^t` for a nat-indexed history. definition compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram","label":"regularizedPrefixFeatureGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram","description":"Regularized prefix Gram matrix `lambda I + sum_{t < T} x_t x_t^T`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-930f0b48aa67","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6057,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:297"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def regularizedPrefixFeatureGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) : Matrix Feature Feature Real","missing":[],"search":"regularizedprefixfeaturegram banditrlproof.oful.regularizedprefixfeaturegram regularized prefix gram matrix `lambda i + sum_{t < t} x_t x_t^t`. definition compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prefixFeatureGram_succ","label":"prefixFeatureGram_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.prefixFeatureGram_succ","description":"Prefix Grams grow by one rank-one feature update.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-bf5529f5a224","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6058,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:304"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem prefixFeatureGram_succ {Feature : Type u} (x : Nat -> Feature -> Real) (T : Nat) : prefixFeatureGram x (T + 1) = prefixFeatureGram x T + rankOneGram (x T)","missing":[],"search":"prefixfeaturegram_succ banditrlproof.oful.prefixfeaturegram_succ prefix grams grow by one rank-one feature update. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_succ","label":"regularizedPrefixFeatureGram_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_succ","description":"Regularized prefix Grams grow by one rank-one feature update.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-5d3251c3030c","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6059,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:313"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_succ {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) : regularizedPrefixFeatureGram lambda x (T + 1) = regularizedPrefixFeatureGram lambda x T + rankOneGram (x T)","missing":[],"search":"regularizedprefixfeaturegram_succ banditrlproof.oful.regularizedprefixfeaturegram_succ regularized prefix grams grow by one rank-one feature update. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prefixFeatureGram_zero","label":"prefixFeatureGram_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.prefixFeatureGram_zero","description":"The zero-horizon prefix Gram is the zero matrix.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-a56a92452a54","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6060,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:324"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem prefixFeatureGram_zero {Feature : Type u} (x : Nat -> Feature -> Real) : prefixFeatureGram x 0 = 0","missing":[],"search":"prefixfeaturegram_zero banditrlproof.oful.prefixfeaturegram_zero the zero-horizon prefix gram is the zero matrix. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_zero","label":"regularizedPrefixFeatureGram_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_zero","description":"The zero-horizon regularized prefix Gram is the scalar base `lambda I`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c1266e67fde9","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6061,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:331"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_zero {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) : regularizedPrefixFeatureGram lambda x 0 = (Matrix.scalar Feature lambda : Matrix Feature Feature Real)","missing":[],"search":"regularizedprefixfeaturegram_zero banditrlproof.oful.regularizedprefixfeaturegram_zero the zero-horizon regularized prefix gram is the scalar base `lambda i`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_zero","label":"det_regularizedPrefixFeatureGram_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_zero","description":"The zero-horizon regularized prefix Gram determinant is `lambda^d`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c56e82730cdb","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6062,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:340"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_zero {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) : (regularizedPrefixFeatureGram lambda x 0).det = lambda ^ Fintype.card Feature","missing":[],"search":"det_regularizedprefixfeaturegram_zero banditrlproof.oful.det_regularizedprefixfeaturegram_zero the zero-horizon regularized prefix gram determinant is `lambda^d`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prefixFeatureGram_isHermitian","label":"prefixFeatureGram_isHermitian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.prefixFeatureGram_isHermitian","description":"Prefix Gram matrices are Hermitian.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c12a4826503f","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6063,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:349"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem prefixFeatureGram_isHermitian {Feature : Type u} [Fintype Feature] (x : Nat -> Feature -> Real) (T : Nat) : (prefixFeatureGram x T).IsHermitian","missing":[],"search":"prefixfeaturegram_ishermitian banditrlproof.oful.prefixfeaturegram_ishermitian prefix gram matrices are hermitian. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_isHermitian","label":"regularizedPrefixFeatureGram_isHermitian","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_isHermitian","description":"Regularized prefix Gram matrices are Hermitian.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c4a7cac5efe4","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6064,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:358"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_isHermitian {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) : (regularizedPrefixFeatureGram lambda x T).IsHermitian","missing":[],"search":"regularizedprefixfeaturegram_ishermitian banditrlproof.oful.regularizedprefixfeaturegram_ishermitian regularized prefix gram matrices are hermitian. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_rankOneGram","label":"trace_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_rankOneGram","description":"The trace of a rank-one Gram is the squared feature norm.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-3a7164bc8a6f","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6065,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:368"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_rankOneGram {Feature : Type u} [Fintype Feature] (x : Feature -> Real) : (rankOneGram x).trace = dotProduct x x","missing":[],"search":"trace_rankonegram banditrlproof.oful.trace_rankonegram the trace of a rank-one gram is the squared feature norm. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_scalar_identity","label":"trace_scalar_identity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_scalar_identity","description":"The trace of the scalar regularization matrix is `d * lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-80b2aad69299","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6066,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:375"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_scalar_identity {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) : (Matrix.scalar Feature lambda : Matrix Feature Feature Real).trace = (Fintype.card Feature : Real) * lambda","missing":[],"search":"trace_scalar_identity banditrlproof.oful.trace_scalar_identity the trace of the scalar regularization matrix is `d * lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_prefixFeatureGram","label":"trace_prefixFeatureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_prefixFeatureGram","description":"Prefix Gram trace as the finite sum of squared feature norms.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-8db0c86cfe29","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6067,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:384"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_prefixFeatureGram {Feature : Type u} [Fintype Feature] (x : Nat -> Feature -> Real) (T : Nat) : (prefixFeatureGram x T).trace = (Finset.range T).sum (fun t => dotProduct (x t) (x t))","missing":[],"search":"trace_prefixfeaturegram banditrlproof.oful.trace_prefixfeaturegram prefix gram trace as the finite sum of squared feature norms. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_regularizedPrefixFeatureGram","label":"trace_regularizedPrefixFeatureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_regularizedPrefixFeatureGram","description":"Regularized prefix Gram trace as scalar base plus squared feature norms.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-6d99dd2d6cfb","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6068,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:394"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_regularizedPrefixFeatureGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) : (regularizedPrefixFeatureGram lambda x T).trace = (Fintype.card Feature : Real) * lambda + (Finset.range T).sum (fun t => dotProduct (x t) (x t))","missing":[],"search":"trace_regularizedprefixfeaturegram banditrlproof.oful.trace_regularizedprefixfeaturegram regularized prefix gram trace as scalar base plus squared feature norms. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_regularizedPrefixFeatureGram_le","label":"trace_regularizedPrefixFeatureGram_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_regularizedPrefixFeatureGram_le","description":"If each feature vector has squared norm at most `L2`, the regularized prefix Gram trace is bounded by `d * lambda + T * L2`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-af4ea9974e1e","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6069,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:408"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_regularizedPrefixFeatureGram_le {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hbound : forall t : Nat, t < T -> dotProduct (x t) (x t) <= L2) : (regularizedPrefixFeatureGram lambda x T).trace <= (Fintype.card Feature : Real) * lambda + T * L2","missing":[],"search":"trace_regularizedprefixfeaturegram_le banditrlproof.oful.trace_regularizedprefixfeaturegram_le if each feature vector has squared norm at most `l2`, the regularized prefix gram trace is bounded by `d * lambda + t * l2`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finset_prod_le_pow_sum_div_card_of_nonneg","label":"finset_prod_le_pow_sum_div_card_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finset_prod_le_pow_sum_div_card_of_nonneg","description":"Finite nonnegative product is bounded by the corresponding arithmetic mean power.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-2f775536cf51","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6070,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:424"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finset_prod_le_pow_sum_div_card_of_nonneg {ι : Type v} (s : Finset ι) (hs : s.Nonempty) (z : ι -> Real) (hz : forall i : ι, i ∈ s -> 0 <= z i) : (s.prod z) <= ((s.sum z) / (s.card : Real)) ^ s.card","missing":[],"search":"finset_prod_le_pow_sum_div_card_of_nonneg banditrlproof.oful.finset_prod_le_pow_sum_div_card_of_nonneg finite nonnegative product is bounded by the corresponding arithmetic mean power. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prod_univ_le_pow_sum_div_card_of_nonneg","label":"prod_univ_le_pow_sum_div_card_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.prod_univ_le_pow_sum_div_card_of_nonneg","description":"Fintype specialization of the finite AM-GM product bound.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-faf68b13153f","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6071,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:448"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem prod_univ_le_pow_sum_div_card_of_nonneg {ι : Type v} [Fintype ι] [Nonempty ι] (z : ι -> Real) (hz : forall i : ι, 0 <= z i) : (Finset.univ.prod z) <= ((Finset.univ.sum z) / (Fintype.card ι : Real)) ^ Fintype.card ι","missing":[],"search":"prod_univ_le_pow_sum_div_card_of_nonneg banditrlproof.oful.prod_univ_le_pow_sum_div_card_of_nonneg fintype specialization of the finite am-gm product bound. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_posSemidef_le_pow_trace_div_card","label":"det_posSemidef_le_pow_trace_div_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_posSemidef_le_pow_trace_div_card","description":"Positive-semidefinite determinant upper bound from AM-GM on eigenvalues.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-a8e1972f1291","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6072,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:458"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_posSemidef_le_pow_trace_div_card {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (A : Matrix Feature Feature Real) (hA : A.PosSemidef) : A.det <= (A.trace / (Fintype.card Feature : Real)) ^ Fintype.card Feature","missing":[],"search":"det_possemidef_le_pow_trace_div_card banditrlproof.oful.det_possemidef_le_pow_trace_div_card positive-semidefinite determinant upper bound from am-gm on eigenvalues. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_bound_average","label":"det_regularizedPrefixFeatureGram_le_pow_trace_bound_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_bound_average","description":"If a separate AM-GM/eigenvalue route supplies `det(V_T) <= (trace(V_T) / d)^d`, the prefix trace bound converts it into a dimension/radius determinant upper bound.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c58e2c211773","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6073,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:475"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_pow_trace_bound_average {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hdet_by_trace : (regularizedPrefixFeatureGram lambda x T).det <= ((regularizedPrefixFeatureGram lambda x T).trace / (Fintype.card Feature : Real)) ^ Fintype.card Feature) (havg_nonneg : 0 <= (regularizedPrefixFeatureGram lambda x T).trace / (Fintype.card Feature : Real)) (hbound : forall t : Nat, t < T -> dotProduct (x t) (x t) <= L2) : (regularizedPrefixFeatureGram lambda x T).det <= (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature","missing":[],"search":"det_regularizedprefixfeaturegram_le_pow_trace_bound_average banditrlproof.oful.det_regularizedprefixfeaturegram_le_pow_trace_bound_average if a separate am-gm/eigenvalue route supplies `det(v_t) <= (trace(v_t) / d)^d`, the prefix trace bound converts it into a dimension/radius determinant upper bound. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.rankOneGram_quadraticForm_eq_sq","label":"rankOneGram_quadraticForm_eq_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.rankOneGram_quadraticForm_eq_sq","description":"The quadratic form of a rank-one Gram matrix is the square of the projection of the query vector onto the feature vector.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-464e59094f8e","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6074,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:506"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem rankOneGram_quadraticForm_eq_sq {Feature : Type u} [Fintype Feature] (x y : Feature -> Real) : quadraticForm (rankOneGram x) y = ((Finset.univ : Finset Feature).sum (fun i => x i * y i)) ^ 2","missing":[],"search":"rankonegram_quadraticform_eq_sq banditrlproof.oful.rankonegram_quadraticform_eq_sq the quadratic form of a rank-one gram matrix is the square of the projection of the query vector onto the feature vector. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.rankOneGram_posSemidef","label":"rankOneGram_posSemidef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.rankOneGram_posSemidef","description":"Rank-one Gram matrices are Mathlib-positive semidefinite.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c59428da2299","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6075,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:550"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem rankOneGram_posSemidef {Feature : Type u} [Fintype Feature] (x : Feature -> Real) : (rankOneGram x).PosSemidef","missing":[],"search":"rankonegram_possemidef banditrlproof.oful.rankonegram_possemidef rank-one gram matrices are mathlib-positive semidefinite. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.rankOneGram_quadraticForm_nonneg","label":"rankOneGram_quadraticForm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.rankOneGram_quadraticForm_nonneg","description":"Rank-one Gram matrices have nonnegative quadratic forms.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-23891049da36","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6076,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:566"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem rankOneGram_quadraticForm_nonneg {Feature : Type u} [Fintype Feature] (x y : Feature -> Real) : 0 <= quadraticForm (rankOneGram x) y","missing":[],"search":"rankonegram_quadraticform_nonneg banditrlproof.oful.rankonegram_quadraticform_nonneg rank-one gram matrices have nonnegative quadratic forms. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.featureGram_quadraticForm_eq_sum_sq","label":"featureGram_quadraticForm_eq_sum_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.featureGram_quadraticForm_eq_sum_sq","description":"The quadratic form of a finite-history Gram matrix is a finite sum of squared feature projections.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-6c4392d35807","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6077,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:577"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem featureGram_quadraticForm_eq_sum_sq {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] (x : Time -> Feature -> Real) (y : Feature -> Real) : quadraticForm (featureGram x) y = (Finset.univ : Finset Time).sum (fun t => ((Finset.univ : Finset Feature).sum (fun i => x t i * y i)) ^ 2)","missing":[],"search":"featuregram_quadraticform_eq_sum_sq banditrlproof.oful.featuregram_quadraticform_eq_sum_sq the quadratic form of a finite-history gram matrix is a finite sum of squared feature projections. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prefixFeatureGram_quadraticForm_eq_sum_sq","label":"prefixFeatureGram_quadraticForm_eq_sum_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.prefixFeatureGram_quadraticForm_eq_sum_sq","description":"The quadratic form of a Nat-prefix Gram matrix is a finite range sum of squared feature projections.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-98fbfd8800b2","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6078,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:632"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem prefixFeatureGram_quadraticForm_eq_sum_sq {Feature : Type u} [Fintype Feature] (x : Nat -> Feature -> Real) (T : Nat) (y : Feature -> Real) : quadraticForm (prefixFeatureGram x T) y = (Finset.range T).sum (fun t => ((Finset.univ : Finset Feature).sum (fun i => x t i * y i)) ^ 2)","missing":[],"search":"prefixfeaturegram_quadraticform_eq_sum_sq banditrlproof.oful.prefixfeaturegram_quadraticform_eq_sum_sq the quadratic form of a nat-prefix gram matrix is a finite range sum of squared feature projections. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_eq_sum_sq","label":"regularizedFeatureGram_quadraticForm_eq_sum_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_eq_sum_sq","description":"The regularized Gram quadratic form is the scalar regularization term plus the finite-history Gram sum of squared feature projections.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-8741ea0d319b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6079,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:687"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_quadraticForm_eq_sum_sq {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Time -> Feature -> Real) (y : Feature -> Real) : quadraticForm (regularizedFeatureGram lambda x) y = lambda * (Finset.univ : Finset Feature).sum (fun i => y i ^ 2) + (Finset.univ : Finset Time).sum (fun t => ((Finset.univ : Finset Feature).sum (fun i => x t i * y i)) ^ 2)","missing":[],"search":"regularizedfeaturegram_quadraticform_eq_sum_sq banditrlproof.oful.regularizedfeaturegram_quadraticform_eq_sum_sq the regularized gram quadratic form is the scalar regularization term plus the finite-history gram sum of squared feature projections. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_eq_sum_sq","label":"regularizedPrefixFeatureGram_quadraticForm_eq_sum_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_eq_sum_sq","description":"The regularized prefix Gram quadratic form is the scalar regularization term plus the prefix sum of squared feature projections.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-7aae94ec85a0","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6080,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:706"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_quadraticForm_eq_sum_sq {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (x : Nat -> Feature -> Real) (T : Nat) (y : Feature -> Real) : quadraticForm (regularizedPrefixFeatureGram lambda x T) y = lambda * (Finset.univ : Finset Feature).sum (fun i => y i ^ 2) + (Finset.range T).sum (fun t => ((Finset.univ : Finset Feature).sum (fun i => x t i * y i)) ^ 2)","missing":[],"search":"regularizedprefixfeaturegram_quadraticform_eq_sum_sq banditrlproof.oful.regularizedprefixfeaturegram_quadraticform_eq_sum_sq the regularized prefix gram quadratic form is the scalar regularization term plus the prefix sum of squared feature projections. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.featureGram_quadraticForm_nonneg","label":"featureGram_quadraticForm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.featureGram_quadraticForm_nonneg","description":"Finite-history Gram matrices have nonnegative quadratic forms.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-f094624caef6","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6081,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:722"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem featureGram_quadraticForm_nonneg {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] (x : Time -> Feature -> Real) (y : Feature -> Real) : 0 <= quadraticForm (featureGram x) y","missing":[],"search":"featuregram_quadraticform_nonneg banditrlproof.oful.featuregram_quadraticform_nonneg finite-history gram matrices have nonnegative quadratic forms. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.prefixFeatureGram_quadraticForm_nonneg","label":"prefixFeatureGram_quadraticForm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.prefixFeatureGram_quadraticForm_nonneg","description":"Nat-prefix Gram matrices have nonnegative quadratic forms.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-f74dd9897d94","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6082,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:730"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem prefixFeatureGram_quadraticForm_nonneg {Feature : Type u} [Fintype Feature] (x : Nat -> Feature -> Real) (T : Nat) (y : Feature -> Real) : 0 <= quadraticForm (prefixFeatureGram x T) y","missing":[],"search":"prefixfeaturegram_quadraticform_nonneg banditrlproof.oful.prefixfeaturegram_quadraticform_nonneg nat-prefix gram matrices have nonnegative quadratic forms. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_nonneg","label":"regularizedFeatureGram_quadraticForm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_nonneg","description":"Regularized finite-history Gram matrices are PSD when `0 <= lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-5ec1c957a1ef","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6083,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:738"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_quadraticForm_nonneg {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 <= lambda) (x : Time -> Feature -> Real) (y : Feature -> Real) : 0 <= quadraticForm (regularizedFeatureGram lambda x) y","missing":[],"search":"regularizedfeaturegram_quadraticform_nonneg banditrlproof.oful.regularizedfeaturegram_quadraticform_nonneg regularized finite-history gram matrices are psd when `0 <= lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_nonneg","label":"regularizedPrefixFeatureGram_quadraticForm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_nonneg","description":"Regularized Nat-prefix Gram matrices are PSD when `0 <= lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-363bc861ebfc","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6084,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:752"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_quadraticForm_nonneg {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 <= lambda) (x : Nat -> Feature -> Real) (T : Nat) (y : Feature -> Real) : 0 <= quadraticForm (regularizedPrefixFeatureGram lambda x T) y","missing":[],"search":"regularizedprefixfeaturegram_quadraticform_nonneg banditrlproof.oful.regularizedprefixfeaturegram_quadraticform_nonneg regularized nat-prefix gram matrices are psd when `0 <= lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_pos_of_pos_lambda","label":"regularizedFeatureGram_quadraticForm_pos_of_pos_lambda","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_pos_of_pos_lambda","description":"Regularized finite-history Gram matrices have strictly positive quadratic forms on nonzero vectors when `0 < lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-e62d45169221","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6085,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:768"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_quadraticForm_pos_of_pos_lambda {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Time -> Feature -> Real) (y : Feature -> Real) (hy : ∃ i : Feature, y i ≠ 0) : 0 < quadraticForm (regularizedFeatureGram lambda x) y","missing":[],"search":"regularizedfeaturegram_quadraticform_pos_of_pos_lambda banditrlproof.oful.regularizedfeaturegram_quadraticform_pos_of_pos_lambda regularized finite-history gram matrices have strictly positive quadratic forms on nonzero vectors when `0 < lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_pos_of_pos_lambda","label":"regularizedPrefixFeatureGram_quadraticForm_pos_of_pos_lambda","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_pos_of_pos_lambda","description":"Regularized Nat-prefix Gram matrices have strictly positive quadratic forms on nonzero vectors when `0 < lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-bd10dc131fbc","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6086,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:784"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_quadraticForm_pos_of_pos_lambda {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) (y : Feature -> Real) (hy : ∃ i : Feature, y i ≠ 0) : 0 < quadraticForm (regularizedPrefixFeatureGram lambda x T) y","missing":[],"search":"regularizedprefixfeaturegram_quadraticform_pos_of_pos_lambda banditrlproof.oful.regularizedprefixfeaturegram_quadraticform_pos_of_pos_lambda regularized nat-prefix gram matrices have strictly positive quadratic forms on nonzero vectors when `0 < lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_posDef","label":"regularizedFeatureGram_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_posDef","description":"Regularized finite-history Gram matrices are Mathlib-positive definite.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-4714282905fb","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6087,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:796"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_posDef {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Time -> Feature -> Real) : (regularizedFeatureGram lambda x).PosDef","missing":[],"search":"regularizedfeaturegram_posdef banditrlproof.oful.regularizedfeaturegram_posdef regularized finite-history gram matrices are mathlib-positive definite. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_posDef","label":"regularizedPrefixFeatureGram_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_posDef","description":"Regularized Nat-prefix Gram matrices are Mathlib-positive definite.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-3052e54e3c7b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6088,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:814"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_posDef {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) : (regularizedPrefixFeatureGram lambda x T).PosDef","missing":[],"search":"regularizedprefixfeaturegram_posdef banditrlproof.oful.regularizedprefixfeaturegram_posdef regularized nat-prefix gram matrices are mathlib-positive definite. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_div_card","label":"det_regularizedPrefixFeatureGram_le_pow_trace_div_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_div_card","description":"Regularized Nat-prefix Gram determinant is bounded by trace average power.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-e7e6fa4b4e8e","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6089,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:831"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_pow_trace_div_card {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) : (regularizedPrefixFeatureGram lambda x T).det <= ((regularizedPrefixFeatureGram lambda x T).trace / (Fintype.card Feature : Real)) ^ Fintype.card Feature","missing":[],"search":"det_regularizedprefixfeaturegram_le_pow_trace_div_card banditrlproof.oful.det_regularizedprefixfeaturegram_le_pow_trace_div_card regularized nat-prefix gram determinant is bounded by trace average power. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_bound_average_of_pos_lambda","label":"det_regularizedPrefixFeatureGram_le_pow_trace_bound_average_of_pos_lambda","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_bound_average_of_pos_lambda","description":"Concrete trace/radius determinant upper bound for regularized Nat-prefix Grams.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-d9f995666d23","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6090,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:847"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_pow_trace_bound_average_of_pos_lambda {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hbound : forall t : Nat, t < T -> dotProduct (x t) (x t) <= L2) : (regularizedPrefixFeatureGram lambda x T).det <= (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature","missing":[],"search":"det_regularizedprefixfeaturegram_le_pow_trace_bound_average_of_pos_lambda banditrlproof.oful.det_regularizedprefixfeaturegram_le_pow_trace_bound_average_of_pos_lambda concrete trace/radius determinant upper bound for regularized nat-prefix grams. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_posDef","label":"regularizedPrefixFeatureGram_inv_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_posDef","description":"Inverses of regularized Nat-prefix Gram matrices are positive definite.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c6e5dcc3df1b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6091,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:870"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_inv_posDef {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) : ((regularizedPrefixFeatureGram lambda x T)⁻¹).PosDef","missing":[],"search":"regularizedprefixfeaturegram_inv_posdef banditrlproof.oful.regularizedprefixfeaturegram_inv_posdef inverses of regularized nat-prefix gram matrices are positive definite. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_quadratic_nonneg","label":"regularizedPrefixFeatureGram_inv_quadratic_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_quadratic_nonneg","description":"Inverse-quadratic scalars of regularized Nat-prefix Gram matrices are nonnegative.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1aba8fd7c123","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6092,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:882"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_inv_quadratic_nonneg {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) (y : Feature -> Real) : 0 <= dotProduct y (Matrix.mulVec ((regularizedPrefixFeatureGram lambda x T)⁻¹) y)","missing":[],"search":"regularizedprefixfeaturegram_inv_quadratic_nonneg banditrlproof.oful.regularizedprefixfeaturegram_inv_quadratic_nonneg inverse-quadratic scalars of regularized nat-prefix gram matrices are nonnegative. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_det_pos","label":"regularizedFeatureGram_det_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_det_pos","description":"Regularized finite-history Gram matrices have positive determinant.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-89630a3d804c","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6093,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:899"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_det_pos {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Time -> Feature -> Real) : 0 < (regularizedFeatureGram lambda x).det","missing":[],"search":"regularizedfeaturegram_det_pos banditrlproof.oful.regularizedfeaturegram_det_pos regularized finite-history gram matrices have positive determinant. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_det_pos","label":"regularizedPrefixFeatureGram_det_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_det_pos","description":"Regularized Nat-prefix Gram matrices have positive determinant.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-847a026e03f7","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6094,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:908"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_det_pos {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) : 0 < (regularizedPrefixFeatureGram lambda x T).det","missing":[],"search":"regularizedprefixfeaturegram_det_pos banditrlproof.oful.regularizedprefixfeaturegram_det_pos regularized nat-prefix gram matrices have positive determinant. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_det_ne_zero","label":"regularizedFeatureGram_det_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_det_ne_zero","description":"Regularized finite-history Gram matrices have nonzero determinant.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-71b890436bff","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6095,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:917"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_det_ne_zero {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Time -> Feature -> Real) : (regularizedFeatureGram lambda x).det ≠ 0","missing":[],"search":"regularizedfeaturegram_det_ne_zero banditrlproof.oful.regularizedfeaturegram_det_ne_zero regularized finite-history gram matrices have nonzero determinant. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_det_ne_zero","label":"regularizedPrefixFeatureGram_det_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_det_ne_zero","description":"Regularized Nat-prefix Gram matrices have nonzero determinant.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-f0e1603dec4b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6096,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:926"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_det_ne_zero {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) : (regularizedPrefixFeatureGram lambda x T).det ≠ 0","missing":[],"search":"regularizedprefixfeaturegram_det_ne_zero banditrlproof.oful.regularizedprefixfeaturegram_det_ne_zero regularized nat-prefix gram matrices have nonzero determinant. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.isUnit_det_regularizedFeatureGram","label":"isUnit_det_regularizedFeatureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.isUnit_det_regularizedFeatureGram","description":"Positive regularization supplies Mathlib's determinant-unit side condition for arbitrary finite-history regularized Gram matrices.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-b2661112f579","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6097,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:937"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem isUnit_det_regularizedFeatureGram {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Time -> Feature -> Real) : IsUnit (regularizedFeatureGram lambda x).det","missing":[],"search":"isunit_det_regularizedfeaturegram banditrlproof.oful.isunit_det_regularizedfeaturegram positive regularization supplies mathlib's determinant-unit side condition for arbitrary finite-history regularized gram matrices. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.isUnit_det_regularizedPrefixFeatureGram","label":"isUnit_det_regularizedPrefixFeatureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.isUnit_det_regularizedPrefixFeatureGram","description":"Positive regularization supplies Mathlib's determinant-unit side condition for Nat-prefix regularized Gram matrices.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-dfe130ea7115","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6098,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:949"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem isUnit_det_regularizedPrefixFeatureGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (x : Nat -> Feature -> Real) (T : Nat) : IsUnit (regularizedPrefixFeatureGram lambda x T).det","missing":[],"search":"isunit_det_regularizedprefixfeaturegram banditrlproof.oful.isunit_det_regularizedprefixfeaturegram positive regularization supplies mathlib's determinant-unit side condition for nat-prefix regularized gram matrices. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedFeatureGram_add_rankOneGram","label":"det_regularizedFeatureGram_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedFeatureGram_add_rankOneGram","description":"One-step determinant recursion for an arbitrary finite-history regularized Gram matrix. This is the determinant identity consumed by OFUL/LinUCB log-det telescoping: the `IsUnit det` side condition is discharged from positive regularization.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-8c0223506912","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6099,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:964"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedFeatureGram_add_rankOneGram {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Time -> Feature -> Real) (x : Feature -> Real) : (regularizedFeatureGram lambda history + rankOneGram x).det = (regularizedFeatureGram lambda history).det * (1 + dotProduct x (Matrix.mulVec ((regularizedFeatureGram lambda history)⁻¹) x))","missing":[],"search":"det_regularizedfeaturegram_add_rankonegram banditrlproof.oful.det_regularizedfeaturegram_add_rankonegram one-step determinant recursion for an arbitrary finite-history regularized gram matrix. this is the determinant identity consumed by oful/linucb log-det telescoping: the `isunit det` side condition is discharged from positive regularization. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_add_rankOneGram","label":"det_regularizedPrefixFeatureGram_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_add_rankOneGram","description":"One-step determinant recursion for Nat-prefix regularized Gram matrices. This keeps the growing-history surface in a single Nat-indexed feature stream.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-6c88092822b5","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6100,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:982"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_add_rankOneGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) : (regularizedPrefixFeatureGram lambda history T + rankOneGram x).det = (regularizedPrefixFeatureGram lambda history T).det * (1 + dotProduct x (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history T)⁻¹) x))","missing":[],"search":"det_regularizedprefixfeaturegram_add_rankonegram banditrlproof.oful.det_regularizedprefixfeaturegram_add_rankonegram one-step determinant recursion for nat-prefix regularized gram matrices. this keeps the growing-history surface in a single nat-indexed feature stream. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_add_rankOneGram_posDef","label":"regularizedFeatureGram_add_rankOneGram_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_add_rankOneGram_posDef","description":"Positive regularized Grams stay positive definite after a rank-one update.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-ac8e6a656bb1","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6101,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:995"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_add_rankOneGram_posDef {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Time -> Feature -> Real) (x : Feature -> Real) : (regularizedFeatureGram lambda history + rankOneGram x).PosDef","missing":[],"search":"regularizedfeaturegram_add_rankonegram_posdef banditrlproof.oful.regularizedfeaturegram_add_rankonegram_posdef positive regularized grams stay positive definite after a rank-one update. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_add_rankOneGram_posDef","label":"regularizedPrefixFeatureGram_add_rankOneGram_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_add_rankOneGram_posDef","description":"Positive regularized prefix Grams stay positive definite after a rank-one update.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-d5775a69e241","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6102,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1006"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_add_rankOneGram_posDef {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) : (regularizedPrefixFeatureGram lambda history T + rankOneGram x).PosDef","missing":[],"search":"regularizedprefixfeaturegram_add_rankonegram_posdef banditrlproof.oful.regularizedprefixfeaturegram_add_rankonegram_posdef positive regularized prefix grams stay positive definite after a rank-one update. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedFeatureGram_add_rankOneGram_pos","label":"det_regularizedFeatureGram_add_rankOneGram_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedFeatureGram_add_rankOneGram_pos","description":"The determinant after a positive regularized rank-one update is positive.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1bfa552d0dce","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6103,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1016"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedFeatureGram_add_rankOneGram_pos {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Time -> Feature -> Real) (x : Feature -> Real) : 0 < (regularizedFeatureGram lambda history + rankOneGram x).det","missing":[],"search":"det_regularizedfeaturegram_add_rankonegram_pos banditrlproof.oful.det_regularizedfeaturegram_add_rankonegram_pos the determinant after a positive regularized rank-one update is positive. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_add_rankOneGram_pos","label":"det_regularizedPrefixFeatureGram_add_rankOneGram_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_add_rankOneGram_pos","description":"The determinant after a positive regularized prefix rank-one update is positive.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-cfa75c9b8136","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6104,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1027"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_add_rankOneGram_pos {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) : 0 < (regularizedPrefixFeatureGram lambda history T + rankOneGram x).det","missing":[],"search":"det_regularizedprefixfeaturegram_add_rankonegram_pos banditrlproof.oful.det_regularizedprefixfeaturegram_add_rankonegram_pos the determinant after a positive regularized prefix rank-one update is positive. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_rankOne_update_factor_pos","label":"regularizedFeatureGram_rankOne_update_factor_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedFeatureGram_rankOne_update_factor_pos","description":"The scalar rank-one determinant-update factor is positive.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-5c252906a5a0","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6105,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1037"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedFeatureGram_rankOne_update_factor_pos {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Time -> Feature -> Real) (x : Feature -> Real) : 0 < 1 + dotProduct x (Matrix.mulVec ((regularizedFeatureGram lambda history)⁻¹) x)","missing":[],"search":"regularizedfeaturegram_rankone_update_factor_pos banditrlproof.oful.regularizedfeaturegram_rankone_update_factor_pos the scalar rank-one determinant-update factor is positive. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_rankOne_update_factor_pos","label":"regularizedPrefixFeatureGram_rankOne_update_factor_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_rankOne_update_factor_pos","description":"The scalar Nat-prefix rank-one determinant-update factor is positive.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-811572493218","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6106,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1063"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_rankOne_update_factor_pos {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) : 0 < 1 + dotProduct x (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history T)⁻¹) x)","missing":[],"search":"regularizedprefixfeaturegram_rankone_update_factor_pos banditrlproof.oful.regularizedprefixfeaturegram_rankone_update_factor_pos the scalar nat-prefix rank-one determinant-update factor is positive. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.log_det_regularizedFeatureGram_add_rankOneGram","label":"log_det_regularizedFeatureGram_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.log_det_regularizedFeatureGram_add_rankOneGram","description":"Logarithmic one-step determinant recursion for regularized Grams.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-abdeed086b90","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6107,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1088"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem log_det_regularizedFeatureGram_add_rankOneGram {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Time -> Feature -> Real) (x : Feature -> Real) : Real.log (regularizedFeatureGram lambda history + rankOneGram x).det = Real.log (regularizedFeatureGram lambda history).det + Real.log (1 + dotProduct x (Matrix.mulVec ((regularizedFeatureGram lambda history)⁻¹) x))","missing":[],"search":"log_det_regularizedfeaturegram_add_rankonegram banditrlproof.oful.log_det_regularizedfeaturegram_add_rankonegram logarithmic one-step determinant recursion for regularized grams. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_add_rankOneGram","label":"log_det_regularizedPrefixFeatureGram_add_rankOneGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_add_rankOneGram","description":"Logarithmic one-step determinant recursion for regularized prefix Grams.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-c8867bb299b5","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6108,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1105"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem log_det_regularizedPrefixFeatureGram_add_rankOneGram {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) : Real.log (regularizedPrefixFeatureGram lambda history T + rankOneGram x).det = Real.log (regularizedPrefixFeatureGram lambda history T).det + Real.log (1 + dotProduct x (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history T)⁻¹) x))","missing":[],"search":"log_det_regularizedprefixfeaturegram_add_rankonegram banditrlproof.oful.log_det_regularizedprefixfeaturegram_add_rankonegram logarithmic one-step determinant recursion for regularized prefix grams. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.log_det_regularizedFeatureGram_add_rankOneGram_sub","label":"log_det_regularizedFeatureGram_add_rankOneGram_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.log_det_regularizedFeatureGram_add_rankOneGram_sub","description":"Increment form of the logarithmic determinant recursion.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-a221f6b9e93c","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6109,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1122"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem log_det_regularizedFeatureGram_add_rankOneGram_sub {Time : Type v} {Feature : Type u} [Fintype Time] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Time -> Feature -> Real) (x : Feature -> Real) : Real.log (regularizedFeatureGram lambda history + rankOneGram x).det - Real.log (regularizedFeatureGram lambda history).det = Real.log (1 + dotProduct x (Matrix.mulVec ((regularizedFeatureGram lambda history)⁻¹) x))","missing":[],"search":"log_det_regularizedfeaturegram_add_rankonegram_sub banditrlproof.oful.log_det_regularizedfeaturegram_add_rankonegram_sub increment form of the logarithmic determinant recursion. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_add_rankOneGram_sub","label":"log_det_regularizedPrefixFeatureGram_add_rankOneGram_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_add_rankOneGram_sub","description":"Increment form of the logarithmic determinant recursion for prefix Grams.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-a63739f136af","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6110,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1136"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem log_det_regularizedPrefixFeatureGram_add_rankOneGram_sub {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) : Real.log (regularizedPrefixFeatureGram lambda history T + rankOneGram x).det - Real.log (regularizedPrefixFeatureGram lambda history T).det = Real.log (1 + dotProduct x (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history T)⁻¹) x))","missing":[],"search":"log_det_regularizedprefixfeaturegram_add_rankonegram_sub banditrlproof.oful.log_det_regularizedprefixfeaturegram_add_rankonegram_sub increment form of the logarithmic determinant recursion for prefix grams. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_succ","label":"det_regularizedPrefixFeatureGram_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_succ","description":"Determinant recursion for the concrete prefix update `T -> T + 1`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-f58fe0721431","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6111,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1150"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_succ {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) : (regularizedPrefixFeatureGram lambda history (T + 1)).det = (regularizedPrefixFeatureGram lambda history T).det * (1 + dotProduct (history T) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history T)⁻¹) (history T)))","missing":[],"search":"det_regularizedprefixfeaturegram_succ banditrlproof.oful.det_regularizedprefixfeaturegram_succ determinant recursion for the concrete prefix update `t -> t + 1`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_succ_sub","label":"log_det_regularizedPrefixFeatureGram_succ_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_succ_sub","description":"Log-det increment recursion for the concrete prefix update `T -> T + 1`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-724e6cca7ce8","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6112,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1165"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem log_det_regularizedPrefixFeatureGram_succ_sub {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) : Real.log (regularizedPrefixFeatureGram lambda history (T + 1)).det - Real.log (regularizedPrefixFeatureGram lambda history T).det = Real.log (1 + dotProduct (history T) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history T)⁻¹) (history T)))","missing":[],"search":"log_det_regularizedprefixfeaturegram_succ_sub banditrlproof.oful.log_det_regularizedprefixfeaturegram_succ_sub log-det increment recursion for the concrete prefix update `t -> t + 1`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_forward_difference","label":"sum_range_forward_difference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_forward_difference","description":"Finite forward-difference telescope used by log-det recursions.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-3ee8874019b3","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6113,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1180"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_forward_difference (Phi : Nat -> Real) (T : Nat) : (Finset.range T).sum (fun t => Phi (t + 1) - Phi t) = Phi T - Phi 0","missing":[],"search":"sum_range_forward_difference banditrlproof.oful.sum_range_forward_difference finite forward-difference telescope used by log-det recursions. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_log_update_factor_eq_log_det_ratio","label":"sum_range_log_update_factor_eq_log_det_ratio","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_log_update_factor_eq_log_det_ratio","description":"Abstract finite-horizon log-det telescope from one-step log-update factors. This packages the shape needed after instantiating `detSeq` with an OFUL regularized Gram determinant process and `factor` with the corresponding rank-one update factor.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-904970f61da3","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6114,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1197"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_log_update_factor_eq_log_det_ratio (detSeq factor : Nat -> Real) (T : Nat) (hstep : forall t : Nat, t < T -> Real.log (detSeq (t + 1)) - Real.log (detSeq t) = Real.log (factor t)) : (Finset.range T).sum (fun t => Real.log (factor t)) = Real.log (detSeq T) - Real.log (detSeq 0)","missing":[],"search":"sum_range_log_update_factor_eq_log_det_ratio banditrlproof.oful.sum_range_log_update_factor_eq_log_det_ratio abstract finite-horizon log-det telescope from one-step log-update factors. this packages the shape needed after instantiating `detseq` with an oful regularized gram determinant process and `factor` with the corresponding rank-one update factor. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.div_one_add_self_le_log_one_add","label":"div_one_add_self_le_log_one_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.div_one_add_self_le_log_one_add","description":"Lower bound for `log (1 + z)` used in elliptical-potential estimates.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1e93aa301235","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6115,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1216"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem div_one_add_self_le_log_one_add {z : Real} (hz : 0 <= z) : z / (1 + z) <= Real.log (1 + z)","missing":[],"search":"div_one_add_self_le_log_one_add banditrlproof.oful.div_one_add_self_le_log_one_add lower bound for `log (1 + z)` used in elliptical-potential estimates. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.self_le_two_log_one_add_of_le_one","label":"self_le_two_log_one_add_of_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.self_le_two_log_one_add_of_le_one","description":"For `0 <= z <= 1`, `z` is bounded by twice the log update.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-df2eb5acc2c6","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6116,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1234"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem self_le_two_log_one_add_of_le_one {z : Real} (hz0 : 0 <= z) (hz1 : z <= 1) : z <= 2 * Real.log (1 + z)","missing":[],"search":"self_le_two_log_one_add_of_le_one banditrlproof.oful.self_le_two_log_one_add_of_le_one for `0 <= z <= 1`, `z` is bounded by twice the log update. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.one_le_two_log_two","label":"one_le_two_log_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.one_le_two_log_two","description":"Numeric endpoint: `1 <= 2 * log 2`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1000fcaa117a","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6117,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1247"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem one_le_two_log_two : (1 : Real) <= 2 * Real.log 2","missing":[],"search":"one_le_two_log_two banditrlproof.oful.one_le_two_log_two numeric endpoint: `1 <= 2 * log 2`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.min_one_le_two_log_one_add","label":"min_one_le_two_log_one_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.min_one_le_two_log_one_add","description":"Elliptical-potential scalar inequality: `min 1 z <= 2 log (1+z)`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-ee82aaf1d27d","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6118,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1254"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem min_one_le_two_log_one_add {z : Real} (hz : 0 <= z) : min 1 z <= 2 * Real.log (1 + z)","missing":[],"search":"min_one_le_two_log_one_add banditrlproof.oful.min_one_le_two_log_one_add elliptical-potential scalar inequality: `min 1 z <= 2 log (1+z)`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_one_le_two_sum_log_one_add","label":"sum_range_min_one_le_two_sum_log_one_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_one_le_two_sum_log_one_add","description":"Finite-sum scalar elliptical-potential wrapper. This consumes nonnegativity of each update scalar and leaves the log telescope or determinant-ratio identity as a separate input.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-a94c3f398fdb","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6119,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1272"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_one_le_two_sum_log_one_add (u : Nat -> Real) (T : Nat) (hu : forall t : Nat, t < T -> 0 <= u t) : (Finset.range T).sum (fun t => min 1 (u t)) <= 2 * (Finset.range T).sum (fun t => Real.log (1 + u t))","missing":[],"search":"sum_range_min_one_le_two_sum_log_one_add banditrlproof.oful.sum_range_min_one_le_two_sum_log_one_add finite-sum scalar elliptical-potential wrapper. this consumes nonnegativity of each update scalar and leaves the log telescope or determinant-ratio identity as a separate input. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_one_le_two_of_sum_log_one_add_le","label":"sum_range_min_one_le_two_of_sum_log_one_add_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_one_le_two_of_sum_log_one_add_le","description":"Clipped finite-sum upper-bound handoff under an explicit log-sum certificate. This is the scalar clipped counterpart of the later small-update raw-sum handoff and does not mention any determinant or matrix route.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-0173e9d71fde","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6120,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1292"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_one_le_two_of_sum_log_one_add_le (u : Nat -> Real) (T : Nat) (B : Real) (hu : forall t : Nat, t < T -> 0 <= u t) (hlog_sum_le : (Finset.range T).sum (fun t => Real.log (1 + u t)) <= B) : (Finset.range T).sum (fun t => min 1 (u t)) <= 2 * B","missing":[],"search":"sum_range_min_one_le_two_of_sum_log_one_add_le banditrlproof.oful.sum_range_min_one_le_two_of_sum_log_one_add_le clipped finite-sum upper-bound handoff under an explicit log-sum certificate. this is the scalar clipped counterpart of the later small-update raw-sum handoff and does not mention any determinant or matrix route. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_le_two_sum_log_one_add_of_le_one","label":"sum_range_le_two_sum_log_one_add_of_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_le_two_sum_log_one_add_of_le_one","description":"Raw finite-sum log upper bound under an explicit small-update contract. For nonnegative update scalars bounded by one, the unclipped sum agrees with the clipped sum consumed by the determinant-growth route.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-29f0e9e22681","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6121,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1307"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_le_two_sum_log_one_add_of_le_one (u : Nat -> Real) (T : Nat) (hu_nonneg : forall t : Nat, t < T -> 0 <= u t) (hu_le_one : forall t : Nat, t < T -> u t <= 1) : (Finset.range T).sum (fun t => u t) <= 2 * (Finset.range T).sum (fun t => Real.log (1 + u t))","missing":[],"search":"sum_range_le_two_sum_log_one_add_of_le_one banditrlproof.oful.sum_range_le_two_sum_log_one_add_of_le_one raw finite-sum log upper bound under an explicit small-update contract. for nonnegative update scalars bounded by one, the unclipped sum agrees with the clipped sum consumed by the determinant-growth route. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_le_two_of_sum_log_one_add_le_of_le_one","label":"sum_range_le_two_of_sum_log_one_add_le_of_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_le_two_of_sum_log_one_add_le_of_le_one","description":"Raw finite-sum upper-bound handoff under an explicit log-sum certificate. This keeps the scalar small-update proof independent from any determinant or matrix route that might later prove the log-sum upper bound.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-09ed31ed7d59","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6122,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1330"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_le_two_of_sum_log_one_add_le_of_le_one (u : Nat -> Real) (T : Nat) (B : Real) (hu_nonneg : forall t : Nat, t < T -> 0 <= u t) (hu_le_one : forall t : Nat, t < T -> u t <= 1) (hlog_sum_le : (Finset.range T).sum (fun t => Real.log (1 + u t)) <= B) : (Finset.range T).sum (fun t => u t) <= 2 * B","missing":[],"search":"sum_range_le_two_of_sum_log_one_add_le_of_le_one banditrlproof.oful.sum_range_le_two_of_sum_log_one_add_le_of_le_one raw finite-sum upper-bound handoff under an explicit log-sum certificate. this keeps the scalar small-update proof independent from any determinant or matrix route that might later prove the log-sum upper bound. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_ratio","label":"sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_ratio","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_ratio","description":"Concrete finite-horizon log-det telescope for Nat-prefix regularized Grams. This is the growing-history instantiation of the abstract telescope. It does not prove any determinant-growth upper bound for the resulting ratio.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-d95bdc53dabf","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6123,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1347"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_ratio {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) : (Finset.range T).sum (fun t => Real.log (1 + dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) = Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (regularizedPrefixFeatureGram lambda history 0).det","missing":[],"search":"sum_range_log_regularizedprefixfeaturegram_update_factor_eq_log_det_ratio banditrlproof.oful.sum_range_log_regularizedprefixfeaturegram_update_factor_eq_log_det_ratio concrete finite-horizon log-det telescope for nat-prefix regularized grams. this is the growing-history instantiation of the abstract telescope. it does not prove any determinant-growth upper bound for the resulting ratio. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_sub_base","label":"sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_sub_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_sub_base","description":"Concrete prefix log-det telescope with the scalar-base determinant expanded as `lambda ^ d`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-b129d502d1c4","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6124,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1373"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_sub_base {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) : (Finset.range T).sum (fun t => Real.log (1 + dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) = Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature)","missing":[],"search":"sum_range_log_regularizedprefixfeaturegram_update_factor_eq_log_det_sub_base banditrlproof.oful.sum_range_log_regularizedprefixfeaturegram_update_factor_eq_log_det_sub_base concrete prefix log-det telescope with the scalar-base determinant expanded as `lambda ^ d`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_sub_base","label":"sum_range_min_prefix_update_le_two_log_det_sub_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_sub_base","description":"Concrete prefix determinant-growth consumer with the inverse-quadratic nonnegativity contract left explicit. This is the finite-sum elliptical-potential inequality once each update scalar `x_t^T V_t^{-1} x_t` is known nonnegative.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-2e7c21c7ef95","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6125,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1395"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_log_det_sub_base {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (hquad_nonneg : forall t : Nat, t < T -> 0 <= dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * (Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature))","missing":[],"search":"sum_range_min_prefix_update_le_two_log_det_sub_base banditrlproof.oful.sum_range_min_prefix_update_le_two_log_det_sub_base concrete prefix determinant-growth consumer with the inverse-quadratic nonnegativity contract left explicit. this is the finite-sum elliptical-potential inequality once each update scalar `x_t^t v_t^{-1} x_t` is known nonnegative. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_sub_base","label":"sum_range_prefix_update_le_two_log_det_sub_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_sub_base","description":"Raw prefix inverse-quadratic sum bound from the terminal log-determinant ratio, with inverse-quadratic nonnegativity and small-update contracts left explicit.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-0c5e3025d185","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6126,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1425"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_log_det_sub_base {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (hquad_nonneg : forall t : Nat, t < T -> 0 <= dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * (Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature))","missing":[],"search":"sum_range_prefix_update_le_two_log_det_sub_base banditrlproof.oful.sum_range_prefix_update_le_two_log_det_sub_base raw prefix inverse-quadratic sum bound from the terminal log-determinant ratio, with inverse-quadratic nonnegativity and small-update contracts left explicit. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda","label":"sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda","description":"Concrete prefix determinant-growth consumer with inverse-quadratic nonnegativity discharged from the positive-definite inverse of the regularized Gram matrix.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-0f60fdd9b327","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6127,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1462"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * (Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature))","missing":[],"search":"sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda banditrlproof.oful.sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda concrete prefix determinant-growth consumer with inverse-quadratic nonnegativity discharged from the positive-definite inverse of the regularized gram matrix. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one","label":"sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one","description":"Raw prefix inverse-quadratic sum bound from the terminal log-determinant ratio, under a small-update contract. When every update scalar is at most one, the raw update sum agrees with the clipped sum controlled by the log-det telescope endpoint.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-20aab9b101ab","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6128,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1486"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * (Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature))","missing":[],"search":"sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one raw prefix inverse-quadratic sum bound from the terminal log-determinant ratio, under a small-update contract. when every update scalar is at most one, the raw update sum agrees with the clipped sum controlled by the log-det telescope endpoint. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_upper","label":"sum_range_min_prefix_update_le_two_log_det_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_upper","description":"If a separate route supplies a terminal log-determinant upper bound, the clipped prefix inverse-quadratic sum inherits the corresponding bound.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-0b4c9c82a42b","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6129,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1513"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_log_det_upper {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (B : Real) (hlog_upper : Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature) <= B) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * B","missing":[],"search":"sum_range_min_prefix_update_le_two_log_det_upper banditrlproof.oful.sum_range_min_prefix_update_le_two_log_det_upper if a separate route supplies a terminal log-determinant upper bound, the clipped prefix inverse-quadratic sum inherits the corresponding bound. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_upper_of_update_le_one","label":"sum_range_prefix_update_le_two_log_det_upper_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_upper_of_update_le_one","description":"Raw prefix inverse-quadratic sum bound from a terminal log-determinant upper bound, under a small-update contract. When every update scalar is at most one, the raw update sum agrees with the clipped sum consumed by the determinant-growth route.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-fc8ef6346965","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6130,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1537"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_log_det_upper_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (B : Real) (hlog_upper : Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature) <= B) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * B","missing":[],"search":"sum_range_prefix_update_le_two_log_det_upper_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_log_det_upper_of_update_le_one raw prefix inverse-quadratic sum bound from a terminal log-determinant upper bound, under a small-update contract. when every update scalar is at most one, the raw update sum agrees with the clipped sum consumed by the determinant-growth route. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_sub_base_le_of_det_le_mul_exp","label":"log_det_regularizedPrefixFeatureGram_sub_base_le_of_det_le_mul_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_sub_base_le_of_det_le_mul_exp","description":"A multiplicative determinant upper bound of the form `det(V_T) <= lambda^d * exp(B)` supplies the terminal log-determinant upper bound used by the clipped elliptical-potential consumer.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-2f76c24d5e66","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6131,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1564"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem log_det_regularizedPrefixFeatureGram_sub_base_le_of_det_le_mul_exp {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (B : Real) (hdet_upper : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp B) : Real.log (regularizedPrefixFeatureGram lambda history T).det - Real.log (lambda ^ Fintype.card Feature) <= B","missing":[],"search":"log_det_regularizedprefixfeaturegram_sub_base_le_of_det_le_mul_exp banditrlproof.oful.log_det_regularizedprefixfeaturegram_sub_base_le_of_det_le_mul_exp a multiplicative determinant upper bound of the form `det(v_t) <= lambda^d * exp(b)` supplies the terminal log-determinant upper bound used by the clipped elliptical-potential consumer. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_det_mul_exp_upper","label":"sum_range_min_prefix_update_le_two_det_mul_exp_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_det_mul_exp_upper","description":"Clipped prefix inverse-quadratic sum bound from a multiplicative terminal determinant upper bound.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-ad8e77dcae9c","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6132,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1591"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_det_mul_exp_upper {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (B : Real) (hdet_upper : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp B) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * B","missing":[],"search":"sum_range_min_prefix_update_le_two_det_mul_exp_upper banditrlproof.oful.sum_range_min_prefix_update_le_two_det_mul_exp_upper clipped prefix inverse-quadratic sum bound from a multiplicative terminal determinant upper bound. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one","label":"sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one","description":"Raw prefix inverse-quadratic sum bound from a multiplicative determinant upper bound, under a small-update contract.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-d9dd52d9e60a","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6133,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1613"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (B : Real) (hdet_upper : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp B) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * B","missing":[],"search":"sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one raw prefix inverse-quadratic sum bound from a multiplicative determinant upper bound, under a small-update contract. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp_dim_scaled","label":"trace_average_pow_le_lambda_pow_mul_exp_dim_scaled","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp_dim_scaled","description":"Scalar simplification for the AM-GM trace/radius determinant upper bound. With `d = Fintype.card Feature` and squared-radius bound `L2 >= 0`, the trace-average expression is bounded by the multiplicative exponential form with exponent `d * (T * L2 / (d * lambda))`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1758c5b5bd8f","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6134,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1644"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_average_pow_le_lambda_pow_mul_exp_dim_scaled {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) : (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature <= lambda ^ Fintype.card Feature * Real.exp ((Fintype.card Feature : Real) * ((T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"trace_average_pow_le_lambda_pow_mul_exp_dim_scaled banditrlproof.oful.trace_average_pow_le_lambda_pow_mul_exp_dim_scaled scalar simplification for the am-gm trace/radius determinant upper bound. with `d = fintype.card feature` and squared-radius bound `l2 >= 0`, the trace-average expression is bounded by the multiplicative exponential form with exponent `d * (t * l2 / (d * lambda))`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp_log","label":"trace_average_pow_le_lambda_pow_mul_exp_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp_log","description":"Scalar AM-GM trace/radius simplification in the standard logarithmic determinant-growth form. This keeps the textbook `d * log (1 + T L2 / (d lambda))` exponent instead of linearizing it by `log (1 + x) <= x`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-ad3527048f93","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6135,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1705"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_average_pow_le_lambda_pow_mul_exp_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) : (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature <= lambda ^ Fintype.card Feature * Real.exp ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"trace_average_pow_le_lambda_pow_mul_exp_log banditrlproof.oful.trace_average_pow_le_lambda_pow_mul_exp_log scalar am-gm trace/radius simplification in the standard logarithmic determinant-growth form. this keeps the textbook `d * log (1 + t l2 / (d lambda))` exponent instead of linearizing it by `log (1 + x) <= x`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_average_exp_exponent_dim_cancel","label":"trace_average_exp_exponent_dim_cancel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_average_exp_exponent_dim_cancel","description":"Dimension cancellation in the scalar OFUL trace-average exponent.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-d2b7d5969802","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6136,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1755"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_average_exp_exponent_dim_cancel {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (T : Nat) (L2 : Real) : (Fintype.card Feature : Real) * ((T * L2) / ((Fintype.card Feature : Real) * lambda)) = (T * L2) / lambda","missing":[],"search":"trace_average_exp_exponent_dim_cancel banditrlproof.oful.trace_average_exp_exponent_dim_cancel dimension cancellation in the scalar oful trace-average exponent. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp","label":"trace_average_pow_le_lambda_pow_mul_exp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp","description":"Scalar AM-GM trace/radius simplification with the dimension-cancelled exponent `T * L2 / lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-386bca480acc","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6137,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1770"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem trace_average_pow_le_lambda_pow_mul_exp {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) : (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature <= lambda ^ Fintype.card Feature * Real.exp ((T * L2) / lambda)","missing":[],"search":"trace_average_pow_le_lambda_pow_mul_exp banditrlproof.oful.trace_average_pow_le_lambda_pow_mul_exp scalar am-gm trace/radius simplification with the dimension-cancelled exponent `t * l2 / lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_of_trace_average_bound","label":"det_regularizedPrefixFeatureGram_le_mul_exp_of_trace_average_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_of_trace_average_bound","description":"If the trace-average determinant upper bound is simplified to the standard multiplicative `lambda^d * exp(B)` form, regularized Nat-prefix Grams inherit that multiplicative determinant upper bound.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-0d8e0a28c105","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6138,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1787"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_mul_exp_of_trace_average_bound {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 B : Real) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) (haverage_to_exp : (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature <= lambda ^ Fintype.card Feature * Real.exp B) : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp B","missing":[],"search":"det_regularizedprefixfeaturegram_le_mul_exp_of_trace_average_bound banditrlproof.oful.det_regularizedprefixfeaturegram_le_mul_exp_of_trace_average_bound if the trace-average determinant upper bound is simplified to the standard multiplicative `lambda^d * exp(b)` form, regularized nat-prefix grams inherit that multiplicative determinant upper bound. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_of_trace_average_bound","label":"sum_range_min_prefix_update_le_two_of_trace_average_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_of_trace_average_bound","description":"Clipped prefix inverse-quadratic sum bound from the AM-GM determinant trace route plus a separate scalar simplification to `lambda^d * exp(B)`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-7333cacc3885","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6139,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1806"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_of_trace_average_bound {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 B : Real) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) (haverage_to_exp : (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature <= lambda ^ Fintype.card Feature * Real.exp B) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * B","missing":[],"search":"sum_range_min_prefix_update_le_two_of_trace_average_bound banditrlproof.oful.sum_range_min_prefix_update_le_two_of_trace_average_bound clipped prefix inverse-quadratic sum bound from the am-gm determinant trace route plus a separate scalar simplification to `lambda^d * exp(b)`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one","label":"sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one","description":"Unclipped prefix inverse-quadratic sum bound from a scalar trace-average certificate, under a small-update contract. When each update scalar is at most one, the raw update sum agrees with the clipped sum used by the standard elliptical-potential route.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-76b632651bfd","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6140,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1834"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 B : Real) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) (haverage_to_exp : (((Fintype.card Feature : Real) * lambda + T * L2) / (Fintype.card Feature : Real)) ^ Fintype.card Feature <= lambda ^ Fintype.card Feature * Real.exp B) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * B","missing":[],"search":"sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one unclipped prefix inverse-quadratic sum bound from a scalar trace-average certificate, under a small-update contract. when each update scalar is at most one, the raw update sum agrees with the clipped sum used by the standard elliptical-potential route. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_dim_scaled","label":"det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_dim_scaled","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_dim_scaled","description":"Concrete determinant upper bound obtained by combining the AM-GM trace/radius route with the scalar exponential simplification.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-2dd6c335185e","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6141,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1877"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_dim_scaled {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp ((Fintype.card Feature : Real) * ((T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"det_regularizedprefixfeaturegram_le_mul_exp_trace_average_dim_scaled banditrlproof.oful.det_regularizedprefixfeaturegram_le_mul_exp_trace_average_dim_scaled concrete determinant upper bound obtained by combining the am-gm trace/radius route with the scalar exponential simplification. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_dim_scaled","label":"sum_range_min_prefix_update_le_two_trace_average_dim_scaled","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_dim_scaled","description":"Concrete clipped elliptical-potential sum bound obtained from the trace/radius route and scalar exponential simplification.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-ec2720f44807","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6142,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1900"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_trace_average_dim_scaled {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * ((Fintype.card Feature : Real) * ((T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"sum_range_min_prefix_update_le_two_trace_average_dim_scaled banditrlproof.oful.sum_range_min_prefix_update_le_two_trace_average_dim_scaled concrete clipped elliptical-potential sum bound obtained from the trace/radius route and scalar exponential simplification. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one","label":"sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one","description":"Concrete raw small-update elliptical-potential sum bound with the dimension-scaled exponent `d * (T * L2 / (d * lambda))`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-41b13c08a468","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6143,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1926"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * ((Fintype.card Feature : Real) * ((T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one concrete raw small-update elliptical-potential sum bound with the dimension-scaled exponent `d * (t * l2 / (d * lambda))`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average","label":"det_regularizedPrefixFeatureGram_le_mul_exp_trace_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average","description":"Concrete determinant upper bound from trace/radius with the dimension-cancelled exponent `T * L2 / lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-7124df153bd5","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6144,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1958"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_mul_exp_trace_average {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp ((T * L2) / lambda)","missing":[],"search":"det_regularizedprefixfeaturegram_le_mul_exp_trace_average banditrlproof.oful.det_regularizedprefixfeaturegram_le_mul_exp_trace_average concrete determinant upper bound from trace/radius with the dimension-cancelled exponent `t * l2 / lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_log","label":"det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_log","description":"Concrete determinant upper bound from trace/radius with the standard logarithmic exponent `d * log (1 + T L2 / (d lambda))`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-b11f827ab8a8","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6145,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:1976"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_log {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"det_regularizedprefixfeaturegram_le_mul_exp_trace_average_log banditrlproof.oful.det_regularizedprefixfeaturegram_le_mul_exp_trace_average_log concrete determinant upper bound from trace/radius with the standard logarithmic exponent `d * log (1 + t l2 / (d lambda))`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average","label":"sum_range_min_prefix_update_le_two_trace_average","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average","description":"Concrete clipped elliptical-potential sum bound with the dimension-cancelled exponent `T * L2 / lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-f4dee826ad4f","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6146,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:2001"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_trace_average {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * ((T * L2) / lambda)","missing":[],"search":"sum_range_min_prefix_update_le_two_trace_average banditrlproof.oful.sum_range_min_prefix_update_le_two_trace_average concrete clipped elliptical-potential sum bound with the dimension-cancelled exponent `t * l2 / lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_of_update_le_one","label":"sum_range_prefix_update_le_two_trace_average_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_of_update_le_one","description":"Concrete raw small-update elliptical-potential sum bound with the dimension-cancelled exponent `T * L2 / lambda`.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-1ea0880d6eed","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6147,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:2023"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_trace_average_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * ((T * L2) / lambda)","missing":[],"search":"sum_range_prefix_update_le_two_trace_average_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_trace_average_of_update_le_one concrete raw small-update elliptical-potential sum bound with the dimension-cancelled exponent `t * l2 / lambda`. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_log","label":"sum_range_min_prefix_update_le_two_trace_average_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_log","description":"Concrete clipped elliptical-potential sum bound with the standard logarithmic trace/radius endpoint. This is the deterministic OFUL/LinUCB textbook shape before self-normalized concentration and confidence-ellipsoid arguments are introduced.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-494ba5b93d87","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6148,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:2054"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_prefix_update_le_two_trace_average_log {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"sum_range_min_prefix_update_le_two_trace_average_log banditrlproof.oful.sum_range_min_prefix_update_le_two_trace_average_log concrete clipped elliptical-potential sum bound with the standard logarithmic trace/radius endpoint. this is the deterministic oful/linucb textbook shape before self-normalized concentration and confidence-ellipsoid arguments are introduced. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_log_of_update_le_one","label":"sum_range_prefix_update_le_two_trace_average_log_of_update_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_log_of_update_le_one","description":"Unclipped logarithmic elliptical-potential bound for small update scalars. If every inverse-quadratic update scalar is already at most one, the clipped sum bound applies to the raw update sum.","url":"../modules/banditrlproof-ofulellipticalpotential/index.html#decl-9fcf007e77cb","parent":"module:BanditRLProof.OFULEllipticalPotential","order":6149,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotential"],["Source","BanditRLProof/OFULEllipticalPotential.lean:2084"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_prefix_update_le_two_trace_average_log_of_update_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) (hupdate_le_one : forall t : Nat, t < T -> dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)) <= 1) : (Finset.range T).sum (fun t => dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t))) <= 2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"sum_range_prefix_update_le_two_trace_average_log_of_update_le_one banditrlproof.oful.sum_range_prefix_update_le_two_trace_average_log_of_update_le_one unclipped logarithmic elliptical-potential bound for small update scalars. if every inverse-quadratic update scalar is already at most one, the clipped sum bound applies to the raw update sum. theorem compiled","shard":"modules/889e011214c8f29f.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardLogDeterminantAndEllipticalPotential","label":"standardLogDeterminantAndEllipticalPotential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardLogDeterminantAndEllipticalPotential","description":"Standard logarithmic determinant and elliptical-potential bounds. For a positive scalar regularization and uniformly bounded squared feature norms, the terminal regularized Gram determinant and the cumulative clipped inverse-quadratic updates are controlled by `d * log (1 + T * L2 / (d * lambda))`.","url":"../modules/banditrlproof-ofulellipticalpotentialfoundation/index.html#decl-8f1e9d0db503","parent":"module:BanditRLProof.OFULEllipticalPotentialFoundation","order":6150,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULEllipticalPotentialFoundation"],["Source","BanditRLProof/OFULEllipticalPotentialFoundation.lean:25"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardLogDeterminantAndEllipticalPotential {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t : Nat, t < T -> dotProduct (history t) (history t) <= L2) : (regularizedPrefixFeatureGram lambda history T).det <= lambda ^ Fintype.card Feature * Real.exp ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda))) /\\ (Finset.range T).sum (fun t => min 1 (dotProduct (history t) (Matrix.mulVec ((regularizedPrefixFeatureGram lambda history t)⁻¹) (history t)))) <= 2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda)))","missing":[],"search":"standardlogdeterminantandellipticalpotential banditrlproof.oful.standardlogdeterminantandellipticalpotential standard logarithmic determinant and elliptical-potential bounds. for a positive scalar regularization and uniformly bounded squared feature norms, the terminal regularized gram determinant and the cumulative clipped inverse-quadratic updates are controlled by `d * log (1 + t * l2 / (d * lambda))`. theorem compiled","shard":"modules/28ca18f59b716403.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGap","label":"canonicalHistoryTrajectorySumRangeAllGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGap","description":"The complete finite-window linear gap along one canonical trajectory.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-df30f9c54447","parent":"module:BanditRLProof.OFULExpectedRegret","order":6151,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:18"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryTrajectorySumRangeAllGap {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (horizon : Nat) (comparator : Nat -> Fin K) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"canonicalhistorytrajectorysumrangeallgap banditrlproof.oful.canonicalhistorytrajectorysumrangeallgap the complete finite-window linear gap along one canonical trajectory. definition compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.IsOptimalLinearArm","label":"IsOptimalLinearArm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.IsOptimalLinearArm","description":"A fixed arm maximizes the true linear value over the finite action set.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-1b12df2edb3a","parent":"module:BanditRLProof.OFULExpectedRegret","order":6152,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:33"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def IsOptimalLinearArm {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) : Prop","missing":[],"search":"isoptimallineararm banditrlproof.oful.isoptimallineararm a fixed arm maximizes the true linear value over the finite action set. definition compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_nonneg","label":"canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_nonneg","description":"Gap to an optimal fixed arm is pointwise nonnegative.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-c59551f7485a","parent":"module:BanditRLProof.OFULExpectedRegret","order":6153,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:44"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_nonneg {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (horizon : Nat) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (trajectory : Nat -> Fin K × Real) : 0 <= canonicalHistoryTrajectorySumRangeAllGap thetaStar actionFeature horizon (fun _t => best) trajectory","missing":[],"search":"canonicalhistorytrajectorysumrangeallfixedcomparatorgap_nonneg banditrlproof.oful.canonicalhistorytrajectorysumrangeallfixedcomparatorgap_nonneg gap to an optimal fixed arm is pointwise nonnegative. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_canonicalHistoryTrajectorySumRangeAllGap","label":"measurable_canonicalHistoryTrajectorySumRangeAllGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_canonicalHistoryTrajectorySumRangeAllGap","description":"The all-round cumulative linear gap is measurable on trajectory space.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-5ed2534c13ac","parent":"module:BanditRLProof.OFULExpectedRegret","order":6154,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:61"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalHistoryTrajectorySumRangeAllGap {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (horizon : Nat) (comparator : Nat -> Fin K) : Measurable (canonicalHistoryTrajectorySumRangeAllGap thetaStar actionFeature horizon comparator)","missing":[],"search":"measurable_canonicalhistorytrajectorysumrangeallgap banditrlproof.oful.measurable_canonicalhistorytrajectorysumrangeallgap the all-round cumulative linear gap is measurable on trajectory space. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","label":"abs_linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","description":"The absolute linear-value difference between any two bounded arm features is at most `2*S*sqrt L2`.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-0a4a2eed09dd","parent":"module:BanditRLProof.OFULExpectedRegret","order":6155,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:82"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_linearValue_sub_linearValue_le_two_mul_parameterFeatureBound {Feature : Type u} [Fintype Feature] (theta x y : Feature -> Real) (S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength theta <= S) (hx : dotProduct x x <= L2) (hy : dotProduct y y <= L2) : |linearValue theta x - linearValue theta y| <= 2 * S * Real.sqrt L2","missing":[],"search":"abs_linearvalue_sub_linearvalue_le_two_mul_parameterfeaturebound banditrlproof.oful.abs_linearvalue_sub_linearvalue_le_two_mul_parameterfeaturebound the absolute linear-value difference between any two bounded arm features is at most `2*s*sqrt l2`. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope","label":"standardScalarAllRoundGapEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope","description":"Uniform absolute envelope for the complete finite-window cumulative gap.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-38522056938d","parent":"module:BanditRLProof.OFULExpectedRegret","order":6156,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:105"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardScalarAllRoundGapEnvelope (S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"standardscalarallroundgapenvelope banditrlproof.oful.standardscalarallroundgapenvelope uniform absolute envelope for the complete finite-window cumulative gap. definition compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_canonicalHistoryTrajectorySumRangeAllGap_le_envelope","label":"abs_canonicalHistoryTrajectorySumRangeAllGap_le_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_canonicalHistoryTrajectorySumRangeAllGap_le_envelope","description":"Every trajectory satisfies the deterministic all-round gap envelope.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-6609c717115a","parent":"module:BanditRLProof.OFULExpectedRegret","order":6157,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:110"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_canonicalHistoryTrajectorySumRangeAllGap_le_envelope {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (htheta : euclideanLength thetaStar <= S) (trajectory : Nat -> Fin K × Real) : |canonicalHistoryTrajectorySumRangeAllGap thetaStar actionFeature horizon comparator trajectory| <= standardScalarAllRoundGapEnvelope S horizon L2","missing":[],"search":"abs_canonicalhistorytrajectorysumrangeallgap_le_envelope banditrlproof.oful.abs_canonicalhistorytrajectorysumrangeallgap_le_envelope every trajectory satisfies the deterministic all-round gap envelope. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_nonneg","label":"standardScalarAllRoundGapEnvelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_nonneg","description":"The deterministic all-round gap envelope is nonnegative.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-7d95dd089457","parent":"module:BanditRLProof.OFULExpectedRegret","order":6158,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:157"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarAllRoundGapEnvelope_nonneg (S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) : 0 <= standardScalarAllRoundGapEnvelope S horizon L2","missing":[],"search":"standardscalarallroundgapenvelope_nonneg banditrlproof.oful.standardscalarallroundgapenvelope_nonneg the deterministic all-round gap envelope is nonnegative. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound_nonneg","label":"standardScalarAllRoundGapBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapBound_nonneg","description":"The standard all-round high-probability budget is nonnegative.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-fe8626358bab","parent":"module:BanditRLProof.OFULExpectedRegret","order":6159,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:165"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarAllRoundGapBound_nonneg {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) : 0 <= standardScalarAllRoundGapBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"standardscalarallroundgapbound_nonneg banditrlproof.oful.standardscalarallroundgapbound_nonneg the standard all-round high-probability budget is nonnegative. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integrable_canonicalHistoryTrajectorySumRangeAllGap","label":"integrable_canonicalHistoryTrajectorySumRangeAllGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integrable_canonicalHistoryTrajectorySumRangeAllGap","description":"The complete finite-window cumulative gap is integrable under a finite measure.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-307dc00c9900","parent":"module:BanditRLProof.OFULExpectedRegret","order":6160,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:185"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integrable_canonicalHistoryTrajectorySumRangeAllGap {K : Nat} {Feature : Type u} [Fintype Feature] (mu : Measure (Nat -> Fin K × Real)) [IsFiniteMeasure mu] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (htheta : euclideanLength thetaStar <= S) : Integrable (canonicalHistoryTrajectorySumRangeAllGap thetaStar actionFeature horizon comparator) mu","missing":[],"search":"integrable_canonicalhistorytrajectorysumrangeallgap banditrlproof.oful.integrable_canonicalhistorytrajectorysumrangeallgap the complete finite-window cumulative gap is integrable under a finite measure. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurableSet_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","label":"measurableSet_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurableSet_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","description":"The named violation set is measurable.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-a07ede5cf69c","parent":"module:BanditRLProof.OFULExpectedRegret","order":6161,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:211"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (horizon : Nat) (L2 : Real) (comparator : Nat -> Fin K) : MeasurableSet (canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet lambda thetaStar actionFeature R delta S horizon L2 comparator)","missing":[],"search":"measurableset_canonicalhistorytrajectorysumrangeallgapstandardviolationset banditrlproof.oful.measurableset_canonicalhistorytrajectorysumrangeallgapstandardviolationset the named violation set is measurable. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_real_measure","label":"integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_real_measure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_real_measure","description":"Generic expectation assembly: the all-round gap is charged by the standard budget off the violation set and by the deterministic envelope on it.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-4289c2f639be","parent":"module:BanditRLProof.OFULExpectedRegret","order":6162,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:231"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_real_measure {K : Nat} {Feature : Type u} [Fintype Feature] (mu : Measure (Nat -> Fin K × Real)) [IsProbabilityMeasure mu] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (htheta : euclideanLength thetaStar <= S) : integral mu (canonicalHistoryTrajectorySumRangeAllGap thetaStar actionFeature horizon comparator) <= standardScalarAllRoundGapBound (Feature := Feature) R delta lambda S horizon L2 + standardScalarAllRoundGapEnvelope S horizon L2 * mu.real (canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet lambda thetaStar actionFeature R delta S horizon L2 comparator)","missing":[],"search":"integral_canonicalhistorytrajectorysumrangeallgap_le_standard_add_envelope_mul_real_measure banditrlproof.oful.integral_canonicalhistorytrajectorysumrangeallgap_le_standard_add_envelope_mul_real_measure generic expectation assembly: the all-round gap is charged by the standard budget off the violation set and by the deterministic envelope on it. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Finite-window expected cumulative linear-gap theorem for the canonical OFUL trajectory. The additive bad-event contribution is the deterministic all-round envelope times `delta`.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-6f4acc856756","parent":"module:BanditRLProof.OFULExpectedRegret","order":6163,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:322"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : integral (Thompson.canonicalHistoryTrajectoryMeasure (finiteH…","missing":[],"search":"integral_canonicalhistorytrajectorysumrangeallgap_le_standard_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_canonicalhistorytrajectorysumrangeallgap_le_standard_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization finite-window expected cumulative linear-gap theorem for the canonical oful trajectory. the additive bad-event contribution is the deterministic all-round envelope times `delta`. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Fixed-comparator form of the expected cumulative linear-gap theorem, matching the usual stochastic linear-bandit pseudo-regret surface.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-bff4478b9118","parent":"module:BanditRLProof.OFULExpectedRegret","order":6164,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:393"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : integral (Thompson.canonicalHistoryTrajectoryMeasure (finit…","missing":[],"search":"integral_canonicalhistorytrajectorysumrangeallfixedcomparatorgap_le_standard_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_canonicalhistorytrajectorysumrangeallfixedcomparatorgap_le_standard_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization fixed-comparator form of the expected cumulative linear-gap theorem, matching the usual stochastic linear-bandit pseudo-regret surface. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Expected finite-window OFUL pseudo-regret theorem for a certified optimal fixed arm. The conjunction records both nonnegativity and the explicit upper bound.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-5d0209e8bc7a","parent":"module:BanditRLProof.OFULExpectedRegret","order":6165,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:434"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : let exp…","missing":[],"search":"integral_canonicalhistorytrajectorypseudoregret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_canonicalhistorytrajectorypseudoregret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization expected finite-window oful pseudo-regret theorem for a certified optimal fixed arm. the conjunction records both nonnegativity and the explicit upper bound. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedRegretDelta","label":"standardExpectedRegretDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedRegretDelta","description":"Canonical horizon-tuned outer failure budget for expected regret.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-1c7a2c3b302f","parent":"module:BanditRLProof.OFULExpectedRegret","order":6166,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:481"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardExpectedRegretDelta (horizon : Nat) : Real","missing":[],"search":"standardexpectedregretdelta banditrlproof.oful.standardexpectedregretdelta canonical horizon-tuned outer failure budget for expected regret. definition compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedRegretDelta_pos","label":"standardExpectedRegretDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedRegretDelta_pos","description":"theorem standardExpectedRegretDelta_pos (horizon : Nat) : 0 < standardExpectedRegretDelta horizon","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-8b463043ffa1","parent":"module:BanditRLProof.OFULExpectedRegret","order":6167,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:484"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardExpectedRegretDelta_pos (horizon : Nat) : 0 < standardExpectedRegretDelta horizon","missing":[],"search":"standardexpectedregretdelta_pos banditrlproof.oful.standardexpectedregretdelta_pos theorem standardexpectedregretdelta_pos (horizon : nat) : 0 < standardexpectedregretdelta horizon theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedRegretDelta_le_one","label":"standardExpectedRegretDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedRegretDelta_le_one","description":"theorem standardExpectedRegretDelta_le_one (horizon : Nat) : standardExpectedRegretDelta horizon <= 1","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-97ff03de6f8d","parent":"module:BanditRLProof.OFULExpectedRegret","order":6168,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:489"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardExpectedRegretDelta_le_one (horizon : Nat) : standardExpectedRegretDelta horizon <= 1","missing":[],"search":"standardexpectedregretdelta_le_one banditrlproof.oful.standardexpectedregretdelta_le_one theorem standardexpectedregretdelta_le_one (horizon : nat) : standardexpectedregretdelta horizon <= 1 theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_mul_standardExpectedRegretDelta","label":"standardScalarAllRoundGapEnvelope_mul_standardExpectedRegretDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_mul_standardExpectedRegretDelta","description":"At `delta_T=1/(T+1)`, the bad-event envelope charge is one arm-gap envelope.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-e871beadf648","parent":"module:BanditRLProof.OFULExpectedRegret","order":6169,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:498"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarAllRoundGapEnvelope_mul_standardExpectedRegretDelta (S : Real) (horizon : Nat) (L2 : Real) : standardScalarAllRoundGapEnvelope S horizon L2 * standardExpectedRegretDelta horizon = standardScalarInitialGapBound S L2","missing":[],"search":"standardscalarallroundgapenvelope_mul_standardexpectedregretdelta banditrlproof.oful.standardscalarallroundgapenvelope_mul_standardexpectedregretdelta at `delta_t=1/(t+1)`, the bad-event envelope charge is one arm-gap envelope. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Horizon-tuned finite-window expected OFUL pseudo-regret corollary obtained by choosing the outer failure budget `delta_T=1/(T+1)`. The algorithm receives the local parameter `delta_T/(T+1)=1/(T+1)^2`, and the bad-event expectation contributes exactly one additional `2*S*sqrt L2` charge.","url":"../modules/banditrlproof-ofulexpectedregret/index.html#decl-8240a5714224","parent":"module:BanditRLProof.OFULExpectedRegret","order":6170,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegret"],["Source","BanditRLProof/OFULExpectedRegret.lean:515"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : let delta := standardExpectedRegretDelta horizon let expectedPseudoRegret := in…","missing":[],"search":"integral_canonicalhistorytrajectorypseudoregret_nonneg_and_le_standardexpectedbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_canonicalhistorytrajectorypseudoregret_nonneg_and_le_standardexpectedbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization horizon-tuned finite-window expected oful pseudo-regret corollary obtained by choosing the outer failure budget `delta_t=1/(t+1)`. the algorithm receives the local parameter `delta_t/(t+1)=1/(t+1)^2`, and the bad-event expectation contributes exactly one additional `2*s*sqrt l2` charge. theorem compiled","shard":"modules/b138104eaf9568de.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_isBigO_log","label":"standardScalarLogDetBudget_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarLogDetBudget_isBigO_log","description":"theorem standardScalarLogDetBudget_isBigO_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardScalarLogDetBudget (Feature := Feature) lambda (horizon + 1) L2) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-23df20aa43ae","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6171,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarLogDetBudget_isBigO_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardScalarLogDetBudget (Feature := Feature) lambda (horizon + 1) L2) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"standardscalarlogdetbudget_isbigo_log banditrlproof.oful.standardscalarlogdetbudget_isbigo_log theorem standardscalarlogdetbudget_isbigo_log {feature : type u} [fintype feature] [nonempty feature] (lambda : real) (hlambda : 0 < lambda) (l2 : real) (hl2 : 0 <= l2) : (fun horizon : nat => standardscalarlogdetbudget (feature := feature) lambda (horizon + 1) l2) =o[attop] (fun horizon : nat => real.log (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedRegretLogBudget_isBigO_log","label":"standardExpectedRegretLogBudget_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedRegretLogBudget_isBigO_log","description":"theorem standardExpectedRegretLogBudget_isBigO_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardExpectedRegretLogBudget (Feature := Feature) lambda horizon L2) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-103d85b401c1","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6172,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:80"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardExpectedRegretLogBudget_isBigO_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardExpectedRegretLogBudget (Feature := Feature) lambda horizon L2) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"standardexpectedregretlogbudget_isbigo_log banditrlproof.oful.standardexpectedregretlogbudget_isbigo_log theorem standardexpectedregretlogbudget_isbigo_log {feature : type u} [fintype feature] [nonempty feature] (lambda : real) (hlambda : 0 < lambda) (l2 : real) (hl2 : 0 <= l2) : (fun horizon : nat => standardexpectedregretlogbudget (feature := feature) lambda horizon l2) =o[attop] (fun horizon : nat => real.log (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sqrt_log_succ_isBigO_log_succ","label":"sqrt_log_succ_isBigO_log_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sqrt_log_succ_isBigO_log_succ","description":"theorem sqrt_log_succ_isBigO_log_succ : (fun horizon : Nat => Real.sqrt (Real.log (((horizon + 1 : Nat) : Real)))) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-bee80f9962a0","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6173,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:101"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sqrt_log_succ_isBigO_log_succ : (fun horizon : Nat => Real.sqrt (Real.log (((horizon + 1 : Nat) : Real)))) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"sqrt_log_succ_isbigo_log_succ banditrlproof.oful.sqrt_log_succ_isbigo_log_succ theorem sqrt_log_succ_isbigo_log_succ : (fun horizon : nat => real.sqrt (real.log (((horizon + 1 : nat) : real)))) =o[attop] (fun horizon : nat => real.log (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","label":"standardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","description":"theorem standardExpectedPseudoRegretBound_isBigO_sqrt_mul_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2) =O[atTop] (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Re…","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-491c7a61c912","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6174,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:126"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardExpectedPseudoRegretBound_isBigO_sqrt_mul_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2) =O[atTop] (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"standardexpectedpseudoregretbound_isbigo_sqrt_mul_log banditrlproof.oful.standardexpectedpseudoregretbound_isbigo_sqrt_mul_log theorem standardexpectedpseudoregretbound_isbigo_sqrt_mul_log {feature : type u} [fintype feature] [nonempty feature] (r : real) (lambda : real) (hlambda : 0 < lambda) (s : real) (l2 : real) (hl2 : 0 <= l2) : (fun horizon : nat => standardexpectedpseudoregretbound (feature := feature) r lambda s horizon l2) =o[attop] (fun horizon : nat => real.sqrt (((horizon + 1 : nat) : real)) * real.log (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret","label":"canonicalStandardExpectedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret","description":"noncomputable def canonicalStandardExpectedPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (best : Fin K) (horizon : Nat) : Real","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-e64597cfb362","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6175,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:247"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalStandardExpectedPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (best : Fin K) (horizon : Nat) : Real","missing":[],"search":"canonicalstandardexpectedpseudoregret banditrlproof.oful.canonicalstandardexpectedpseudoregret noncomputable def canonicalstandardexpectedpseudoregret {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r s : real) (environment : thompson.historyenvironment (fin k) real) (best : fin k) (horizon : nat) : real definition compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_nonneg_and_le","label":"canonicalStandardExpectedPseudoRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_nonneg_and_le","description":"theorem canonicalStandardExpectedPseudoRegret_nonneg_and_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <=…","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-5062cfbfe992","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6176,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:267"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardExpectedPseudoRegret_nonneg_and_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : 0 <= canonicalStandardExpectedPseudoRegret hK lambda thetaStar actionFeature R S environment best horizon ∧ canonicalStandardExpectedPseudoRegret hK lambda thetaStar actionFeatu…","missing":[],"search":"canonicalstandardexpectedpseudoregret_nonneg_and_le banditrlproof.oful.canonicalstandardexpectedpseudoregret_nonneg_and_le theorem canonicalstandardexpectedpseudoregret_nonneg_and_le {k : nat} {feature : type u} [fintype feature] [decidableeq feature] [nonempty feature] (hk : 0 < k) (lambda : real) (hlambda : 0 < lambda) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r : real) (hr : 0 < r) (s : real) (hs : 0 <= s) (environment : thompson.historyenvironment (fin k) real) (horizon : nat) (l2 : real) (hl2 : 0 <= l2) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (best : fin k) (hbest : isoptimallineararm thetastar actionfeature best) (source : canonicallinearsubgaussianenvironmentlaw hk thetastar actionfeature r s environment) : 0 <= canonicalstandardexpectedpseudoregret hk lambda thetastar actionfeature r s environment best horizon ∧ canonicalstandardexpectedpseudoregret hk lambda thetastar actionfeature r s environment best horizon <= standardexpectedpseudoregretbound (feature := feature) r lambda s horizon l2 theorem compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_isBigO_sqrt_mul_log","label":"canonicalStandardExpectedPseudoRegret_isBigO_sqrt_mul_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_isBigO_sqrt_mul_log","description":"theorem canonicalStandardExpectedPseudoRegret_isBigO_sqrt_mul_log {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hacti…","url":"../modules/banditrlproof-ofulexpectedregretasymptotics/index.html#decl-de35c992201f","parent":"module:BanditRLProof.OFULExpectedRegretAsymptotics","order":6177,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULExpectedRegretAsymptotics.lean:296"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardExpectedPseudoRegret_isBigO_sqrt_mul_log {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : (canonicalStandardExpectedPseudoRegret hK lambda thetaStar actionFeature R S environment best) =O[atTop] (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horiz…","missing":[],"search":"canonicalstandardexpectedpseudoregret_isbigo_sqrt_mul_log banditrlproof.oful.canonicalstandardexpectedpseudoregret_isbigo_sqrt_mul_log theorem canonicalstandardexpectedpseudoregret_isbigo_sqrt_mul_log {k : nat} {feature : type u} [fintype feature] [decidableeq feature] [nonempty feature] (hk : 0 < k) (lambda : real) (hlambda : 0 < lambda) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r : real) (hr : 0 < r) (s : real) (hs : 0 <= s) (environment : thompson.historyenvironment (fin k) real) (l2 : real) (hl2 : 0 <= l2) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (best : fin k) (hbest : isoptimallineararm thetastar actionfeature best) (source : canonicallinearsubgaussianenvironmentlaw hk thetastar actionfeature r s environment) : (canonicalstandardexpectedpseudoregret hk lambda thetastar actionfeature r s environment best) =o[attop] (fun horizon : nat => real.sqrt (((horizon + 1 : nat) : real)) * real.log (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/2291c64b9b778162.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sqrt_mul_log_succ_isLittleO_natCast_succ","label":"sqrt_mul_log_succ_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sqrt_mul_log_succ_isLittleO_natCast_succ","description":"theorem sqrt_mul_log_succ_isLittleO_natCast_succ : (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Real))) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","url":"../modules/banditrlproof-ofulexpectedregretconsistency/index.html#decl-e9ca74d65a41","parent":"module:BanditRLProof.OFULExpectedRegretConsistency","order":6178,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretConsistency"],["Source","BanditRLProof/OFULExpectedRegretConsistency.lean:18"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sqrt_mul_log_succ_isLittleO_natCast_succ : (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Real))) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","missing":[],"search":"sqrt_mul_log_succ_islittleo_natcast_succ banditrlproof.oful.sqrt_mul_log_succ_islittleo_natcast_succ theorem sqrt_mul_log_succ_islittleo_natcast_succ : (fun horizon : nat => real.sqrt (((horizon + 1 : nat) : real)) * real.log (((horizon + 1 : nat) : real))) =o[attop] (fun horizon : nat => (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/c08bbca52c3ee0df.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedPseudoRegretBound_isLittleO_natCast_succ","label":"standardExpectedPseudoRegretBound_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedPseudoRegretBound_isLittleO_natCast_succ","description":"theorem standardExpectedPseudoRegretBound_isLittleO_natCast_succ {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","url":"../modules/banditrlproof-ofulexpectedregretconsistency/index.html#decl-07a88b0e82f6","parent":"module:BanditRLProof.OFULExpectedRegretConsistency","order":6179,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretConsistency"],["Source","BanditRLProof/OFULExpectedRegretConsistency.lean:40"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardExpectedPseudoRegretBound_isLittleO_natCast_succ {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => standardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","missing":[],"search":"standardexpectedpseudoregretbound_islittleo_natcast_succ banditrlproof.oful.standardexpectedpseudoregretbound_islittleo_natcast_succ theorem standardexpectedpseudoregretbound_islittleo_natcast_succ {feature : type u} [fintype feature] [nonempty feature] (r : real) (lambda : real) (hlambda : 0 < lambda) (s : real) (l2 : real) (hl2 : 0 <= l2) : (fun horizon : nat => standardexpectedpseudoregretbound (feature := feature) r lambda s horizon l2) =o[attop] (fun horizon : nat => (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/c08bbca52c3ee0df.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_isLittleO_natCast_succ","label":"canonicalStandardExpectedPseudoRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_isLittleO_natCast_succ","description":"theorem canonicalStandardExpectedPseudoRegret_isLittleO_natCast_succ {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (ha…","url":"../modules/banditrlproof-ofulexpectedregretconsistency/index.html#decl-beb6dbbaa3e6","parent":"module:BanditRLProof.OFULExpectedRegretConsistency","order":6180,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretConsistency"],["Source","BanditRLProof/OFULExpectedRegretConsistency.lean:53"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardExpectedPseudoRegret_isLittleO_natCast_succ {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : (canonicalStandardExpectedPseudoRegret hK lambda thetaStar actionFeature R S environment best) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","missing":[],"search":"canonicalstandardexpectedpseudoregret_islittleo_natcast_succ banditrlproof.oful.canonicalstandardexpectedpseudoregret_islittleo_natcast_succ theorem canonicalstandardexpectedpseudoregret_islittleo_natcast_succ {k : nat} {feature : type u} [fintype feature] [decidableeq feature] [nonempty feature] (hk : 0 < k) (lambda : real) (hlambda : 0 < lambda) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r : real) (hr : 0 < r) (s : real) (hs : 0 <= s) (environment : thompson.historyenvironment (fin k) real) (l2 : real) (hl2 : 0 <= l2) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (best : fin k) (hbest : isoptimallineararm thetastar actionfeature best) (source : canonicallinearsubgaussianenvironmentlaw hk thetastar actionfeature r s environment) : (canonicalstandardexpectedpseudoregret hk lambda thetastar actionfeature r s environment best) =o[attop] (fun horizon : nat => (((horizon + 1 : nat) : real))) theorem compiled","shard":"modules/c08bbca52c3ee0df.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret","label":"canonicalStandardExpectedAveragePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret","description":"noncomputable def canonicalStandardExpectedAveragePseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (best : Fin K) (horizon : Nat) : Real","url":"../modules/banditrlproof-ofulexpectedregretconsistency/index.html#decl-6107bcaba3d6","parent":"module:BanditRLProof.OFULExpectedRegretConsistency","order":6181,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegretConsistency"],["Source","BanditRLProof/OFULExpectedRegretConsistency.lean:79"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalStandardExpectedAveragePseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (best : Fin K) (horizon : Nat) : Real","missing":[],"search":"canonicalstandardexpectedaveragepseudoregret banditrlproof.oful.canonicalstandardexpectedaveragepseudoregret noncomputable def canonicalstandardexpectedaveragepseudoregret {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r s : real) (environment : thompson.historyenvironment (fin k) real) (best : fin k) (horizon : nat) : real definition compiled","shard":"modules/c08bbca52c3ee0df.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret_tendsto_zero","label":"canonicalStandardExpectedAveragePseudoRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret_tendsto_zero","description":"theorem canonicalStandardExpectedAveragePseudoRegret_tendsto_zero {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hacti…","url":"../modules/banditrlproof-ofulexpectedregretconsistency/index.html#decl-63deb95a019a","parent":"module:BanditRLProof.OFULExpectedRegretConsistency","order":6182,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretConsistency"],["Source","BanditRLProof/OFULExpectedRegretConsistency.lean:94"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardExpectedAveragePseudoRegret_tendsto_zero {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Tendsto (canonicalStandardExpectedAveragePseudoRegret hK lambda thetaStar actionFeature R S environment best) atTop (nhds 0)","missing":[],"search":"canonicalstandardexpectedaveragepseudoregret_tendsto_zero banditrlproof.oful.canonicalstandardexpectedaveragepseudoregret_tendsto_zero theorem canonicalstandardexpectedaveragepseudoregret_tendsto_zero {k : nat} {feature : type u} [fintype feature] [decidableeq feature] [nonempty feature] (hk : 0 < k) (lambda : real) (hlambda : 0 < lambda) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r : real) (hr : 0 < r) (s : real) (hs : 0 <= s) (environment : thompson.historyenvironment (fin k) real) (l2 : real) (hl2 : 0 <= l2) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (best : fin k) (hbest : isoptimallineararm thetastar actionfeature best) (source : canonicallinearsubgaussianenvironmentlaw hk thetastar actionfeature r s environment) : tendsto (canonicalstandardexpectedaveragepseudoregret hk lambda thetastar actionfeature r s environment best) attop (nhds 0) theorem compiled","shard":"modules/c08bbca52c3ee0df.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedRegretLogBudget","label":"standardExpectedRegretLogBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedRegretLogBudget","description":"The log budget in the horizon-tuned expected-regret radius: the standard log-determinant term plus the `4 * log (T+1)` confidence contribution.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html#decl-1889d170a416","parent":"module:BanditRLProof.OFULExpectedRegretRate","order":6183,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegretRate"],["Source","BanditRLProof/OFULExpectedRegretRate.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardExpectedRegretLogBudget {Feature : Type u} [Fintype Feature] (lambda : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"standardexpectedregretlogbudget banditrlproof.oful.standardexpectedregretlogbudget the log budget in the horizon-tuned expected-regret radius: the standard log-determinant term plus the `4 * log (t+1)` confidence contribution. definition compiled","shard":"modules/2dbf1f19a2a258b4.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedRegretAlgorithmDelta_eq_inv_sq","label":"standardExpectedRegretAlgorithmDelta_eq_inv_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedRegretAlgorithmDelta_eq_inv_sq","description":"The outer failure budget `1/(T+1)`, divided once more by the uniform-time schedule size, is exactly the algorithm parameter `1/(T+1)^2`.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html#decl-044096782b31","parent":"module:BanditRLProof.OFULExpectedRegretRate","order":6184,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretRate"],["Source","BanditRLProof/OFULExpectedRegretRate.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardExpectedRegretAlgorithmDelta_eq_inv_sq (horizon : Nat) : standardExpectedRegretDelta horizon / (((horizon + 1 : Nat) : Real)) = 1 / (((horizon + 1 : Nat) : Real) ^ 2)","missing":[],"search":"standardexpectedregretalgorithmdelta_eq_inv_sq banditrlproof.oful.standardexpectedregretalgorithmdelta_eq_inv_sq the outer failure budget `1/(t+1)`, divided once more by the uniform-time schedule size, is exactly the algorithm parameter `1/(t+1)^2`. theorem compiled","shard":"modules/2dbf1f19a2a258b4.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_standardExpectedRegret","label":"standardScalarConfidenceRadiusUpper_standardExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_standardExpectedRegret","description":"At the canonical horizon-tuned algorithm parameter, the standard confidence radius has the explicit form `R * sqrt (B_T + 4 * log (T+1)) + sqrt lambda * S`.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html#decl-1bdd2574bad1","parent":"module:BanditRLProof.OFULExpectedRegretRate","order":6185,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretRate"],["Source","BanditRLProof/OFULExpectedRegretRate.lean:45"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarConfidenceRadiusUpper_standardExpectedRegret {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 < R) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : standardScalarConfidenceRadiusUpper (Feature := Feature) R (standardExpectedRegretDelta horizon / (((horizon + 1 : Nat) : Real))) lambda S (horizon + 1) L2 = R * Real.sqrt (standardExpectedRegretLogBudget (Feature := Feature) lambda horizon L2) + Real.sqrt lambda * S","missing":[],"search":"standardscalarconfidenceradiusupper_standardexpectedregret banditrlproof.oful.standardscalarconfidenceradiusupper_standardexpectedregret at the canonical horizon-tuned algorithm parameter, the standard confidence radius has the explicit form `r * sqrt (b_t + 4 * log (t+1)) + sqrt lambda * s`. theorem compiled","shard":"modules/2dbf1f19a2a258b4.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardExpectedPseudoRegretBound","label":"standardExpectedPseudoRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardExpectedPseudoRegretBound","description":"Explicit logarithmic square-root bound for expected OFUL pseudo-regret.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html#decl-302679a03318","parent":"module:BanditRLProof.OFULExpectedRegretRate","order":6186,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULExpectedRegretRate"],["Source","BanditRLProof/OFULExpectedRegretRate.lean:110"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardExpectedPseudoRegretBound {Feature : Type u} [Fintype Feature] (R lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"standardexpectedpseudoregretbound banditrlproof.oful.standardexpectedpseudoregretbound explicit logarithmic square-root bound for expected oful pseudo-regret. definition compiled","shard":"modules/2dbf1f19a2a258b4.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound_add_initial_eq_standardExpectedPseudoRegretBound","label":"standardScalarAllRoundGapBound_add_initial_eq_standardExpectedPseudoRegretBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapBound_add_initial_eq_standardExpectedPseudoRegretBound","description":"The prior named expected bound is exactly the explicit rate expression.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html#decl-bfc37c93fa29","parent":"module:BanditRLProof.OFULExpectedRegretRate","order":6187,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretRate"],["Source","BanditRLProof/OFULExpectedRegretRate.lean:126"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarAllRoundGapBound_add_initial_eq_standardExpectedPseudoRegretBound {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 < R) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : standardScalarAllRoundGapBound (Feature := Feature) R (standardExpectedRegretDelta horizon) lambda S horizon L2 + standardScalarInitialGapBound S L2 = standardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2","missing":[],"search":"standardscalarallroundgapbound_add_initial_eq_standardexpectedpseudoregretbound banditrlproof.oful.standardscalarallroundgapbound_add_initial_eq_standardexpectedpseudoregretbound the prior named expected bound is exactly the explicit rate expression. theorem compiled","shard":"modules/2dbf1f19a2a258b4.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_explicitStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_explicitStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_explicitStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Explicit finite-window expected pseudo-regret theorem for the canonical scalar-ridge OFUL trajectory. The algorithm parameter is displayed directly as `1/(T+1)^2`, and the upper bound is the logarithmic square-root expression `standardExpectedPseudoRegretBound`.","url":"../modules/banditrlproof-ofulexpectedregretrate/index.html#decl-910de2501119","parent":"module:BanditRLProof.OFULExpectedRegretRate","order":6188,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULExpectedRegretRate"],["Source","BanditRLProof/OFULExpectedRegretRate.lean:152"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_explicitStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : let algorithmDelta := 1 / (((horizon + 1 : Nat) : Real) ^ 2) let expect…","missing":[],"search":"integral_canonicalhistorytrajectorypseudoregret_nonneg_and_le_explicitstandardexpectedbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_canonicalhistorytrajectorypseudoregret_nonneg_and_le_explicitstandardexpectedbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization explicit finite-window expected pseudo-regret theorem for the canonical scalar-ridge oful trajectory. the algorithm parameter is displayed directly as `1/(t+1)^2`, and the upper bound is the logarithmic square-root expression `standardexpectedpseudoregretbound`. theorem compiled","shard":"modules/2dbf1f19a2a258b4.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.confidenceWidth","label":"confidenceWidth","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.confidenceWidth","description":"The inverse-Gram confidence width of one feature vector.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-ec205485e89d","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6189,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:18"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def confidenceWidth [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (x : Feature -> Real) : Real","missing":[],"search":"confidencewidth banditrlproof.oful.confidencewidth the inverse-gram confidence width of one feature vector. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.linearValue","label":"linearValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.linearValue","description":"Linear value of a feature vector under a parameter.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-497bb97424ad","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6190,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def linearValue [Fintype Feature] (theta x : Feature -> Real) : Real","missing":[],"search":"linearvalue banditrlproof.oful.linearvalue linear value of a feature vector under a parameter. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.optimisticScore","label":"optimisticScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.optimisticScore","description":"OFUL upper-confidence score `thetaHat dot x + beta * ||x||_(V⁻¹)`.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-698975e16aec","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6191,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:30"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticScore [Fintype Feature] [DecidableEq Feature] (thetaHat : Feature -> Real) (V : Matrix Feature Feature Real) (beta : Real) (x : Feature -> Real) : Real","missing":[],"search":"optimisticscore banditrlproof.oful.optimisticscore oful upper-confidence score `thetahat dot x + beta * ||x||_(v⁻¹)`. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteActionArgmax","label":"finiteActionArgmax","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteActionArgmax","description":"A finite nonempty type admits a score-maximizing action.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-cf16f4570ae2","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6192,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:38"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteActionArgmax [Finite Action] [Nonempty Action] (score : Action -> Real) : Action","missing":[],"search":"finiteactionargmax banditrlproof.oful.finiteactionargmax a finite nonempty type admits a score-maximizing action. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteActionArgmax_spec","label":"finiteActionArgmax_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteActionArgmax_spec","description":"The finite-action argmax dominates every action score.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-50c87f7eacd5","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6193,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:44"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteActionArgmax_spec [Finite Action] [Nonempty Action] (score : Action -> Real) (action : Action) : score action <= score (finiteActionArgmax score)","missing":[],"search":"finiteactionargmax_spec banditrlproof.oful.finiteactionargmax_spec the finite-action argmax dominates every action score. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_dotProduct_le_matrixNorm_mul_confidenceWidth","label":"abs_dotProduct_le_matrixNorm_mul_confidenceWidth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_dotProduct_le_matrixNorm_mul_confidenceWidth","description":"Weighted Cauchy-Schwarz: the ordinary dot product is controlled by the `V`-norm and inverse-`V` confidence width.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-91758f6599a8","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6194,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:54"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_dotProduct_le_matrixNorm_mul_confidenceWidth [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (hV : V.PosDef) (error x : Feature -> Real) : |dotProduct error x| <= matrixNorm V error * confidenceWidth V x","missing":[],"search":"abs_dotproduct_le_matrixnorm_mul_confidencewidth banditrlproof.oful.abs_dotproduct_le_matrixnorm_mul_confidencewidth weighted cauchy-schwarz: the ordinary dot product is controlled by the `v`-norm and inverse-`v` confidence width. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.linearValue_le_optimisticScore_of_matrixNorm_sub_le","label":"linearValue_le_optimisticScore_of_matrixNorm_sub_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.linearValue_le_optimisticScore_of_matrixNorm_sub_le","description":"On a confidence ellipsoid, every true linear action value lies below its upper-confidence score.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-35775843daff","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6195,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:93"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem linearValue_le_optimisticScore_of_matrixNorm_sub_le [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (hV : V.PosDef) (thetaHat thetaStar : Feature -> Real) (beta : Real) (x : Feature -> Real) (hconfidence : matrixNorm V (thetaHat - thetaStar) <= beta) : linearValue thetaStar x <= optimisticScore thetaHat V beta x","missing":[],"search":"linearvalue_le_optimisticscore_of_matrixnorm_sub_le banditrlproof.oful.linearvalue_le_optimisticscore_of_matrixnorm_sub_le on a confidence ellipsoid, every true linear action value lies below its upper-confidence score. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.optimisticScore_le_linearValue_add_two_mul_bonus_of_matrixNorm_sub_le","label":"optimisticScore_le_linearValue_add_two_mul_bonus_of_matrixNorm_sub_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.optimisticScore_le_linearValue_add_two_mul_bonus_of_matrixNorm_sub_le","description":"On the same confidence ellipsoid, every upper-confidence score is at most the true value plus twice its confidence bonus.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-c144611b4cbc","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6196,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:120"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem optimisticScore_le_linearValue_add_two_mul_bonus_of_matrixNorm_sub_le [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (hV : V.PosDef) (thetaHat thetaStar : Feature -> Real) (beta : Real) (x : Feature -> Real) (hconfidence : matrixNorm V (thetaHat - thetaStar) <= beta) : optimisticScore thetaHat V beta x <= linearValue thetaStar x + 2 * beta * confidenceWidth V x","missing":[],"search":"optimisticscore_le_linearvalue_add_two_mul_bonus_of_matrixnorm_sub_le banditrlproof.oful.optimisticscore_le_linearvalue_add_two_mul_bonus_of_matrixnorm_sub_le on the same confidence ellipsoid, every upper-confidence score is at most the true value plus twice its confidence bonus. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteActionOptimisticChoice","label":"finiteActionOptimisticChoice","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteActionOptimisticChoice","description":"The score-maximizing finite action for the OFUL upper-confidence score.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-f0fc0d397896","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6197,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:145"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteActionOptimisticChoice [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] (thetaHat : Feature -> Real) (V : Matrix Feature Feature Real) (beta : Real) (actionFeature : Action -> Feature -> Real) : Action","missing":[],"search":"finiteactionoptimisticchoice banditrlproof.oful.finiteactionoptimisticchoice the score-maximizing finite action for the oful upper-confidence score. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteActionOptimisticChoice_score_max","label":"finiteActionOptimisticChoice_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteActionOptimisticChoice_score_max","description":"The finite OFUL choice maximizes the optimistic score.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-314d27d0442b","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6198,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:155"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteActionOptimisticChoice_score_max [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] (thetaHat : Feature -> Real) (V : Matrix Feature Feature Real) (beta : Real) (actionFeature : Action -> Feature -> Real) (action : Action) : optimisticScore thetaHat V beta (actionFeature action) <= optimisticScore thetaHat V beta (actionFeature (finiteActionOptimisticChoice thetaHat V beta actionFeature))","missing":[],"search":"finiteactionoptimisticchoice_score_max banditrlproof.oful.finiteactionoptimisticchoice_score_max the finite oful choice maximizes the optimistic score. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.linearValue_sub_finiteActionOptimisticChoice_le_two_mul_bonus","label":"linearValue_sub_finiteActionOptimisticChoice_le_two_mul_bonus","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.linearValue_sub_finiteActionOptimisticChoice_le_two_mul_bonus","description":"On the confidence ellipsoid, the true value gap between any comparator and the finite optimistic choice is at most twice the chosen action's bonus.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-aba71e241e15","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6199,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:174"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem linearValue_sub_finiteActionOptimisticChoice_le_two_mul_bonus [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (hV : V.PosDef) (thetaHat thetaStar : Feature -> Real) (beta : Real) (actionFeature : Action -> Feature -> Real) (hconfidence : matrixNorm V (thetaHat - thetaStar) <= beta) (action : Action) : linearValue thetaStar (actionFeature action) - linearValue thetaStar (actionFeature (finiteActionOptimisticChoice thetaHat V beta actionFeature)) <= 2 * beta * confidenceWidth V (actionFeature (finiteActionOptimisticChoice thetaHat V beta actionFeature))","missing":[],"search":"linearvalue_sub_finiteactionoptimisticchoice_le_two_mul_bonus banditrlproof.oful.linearvalue_sub_finiteactionoptimisticchoice_le_two_mul_bonus on the confidence ellipsoid, the true value gap between any comparator and the finite optimistic choice is at most twice the chosen action's bonus. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarGram","label":"finiteHorizonScalarGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarGram","description":"Scalar-regularized finite-horizon Gram matrix used by OFUL selection.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-4a4db96105b0","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6200,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:212"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonScalarGram [Fintype Feature] [DecidableEq Feature] (lambda : Real) (feature : Nat -> Omega -> Feature -> Real) (n : Nat) (omega : Omega) : Matrix Feature Feature Real","missing":[],"search":"finitehorizonscalargram banditrlproof.oful.finitehorizonscalargram scalar-regularized finite-horizon gram matrix used by oful selection. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarOptimisticAction","label":"finiteHorizonScalarOptimisticAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarOptimisticAction","description":"Finite-action OFUL choice using the scalar-ridge estimate and the compiled scalar confidence radius.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-89bc35098b29","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6201,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:223"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonScalarOptimisticAction [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (candidateFeature : Omega -> Action -> Feature -> Real) (R delta S : Real) (n : Nat) (omega : Omega) : Action","missing":[],"search":"finitehorizonscalaroptimisticaction banditrlproof.oful.finitehorizonscalaroptimisticaction finite-action oful choice using the scalar-ridge estimate and the compiled scalar confidence radius. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarOptimisticAction_gap_le","label":"finiteHorizonScalarOptimisticAction_gap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarOptimisticAction_gap_le","description":"Pointwise finite-horizon optimism: on the scalar confidence ellipsoid, every comparator's true linear value exceeds the selected value by at most twice the selected confidence bonus.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-ee6e3b6e82e7","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6202,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:244"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonScalarOptimisticAction_gap_le [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (candidateFeature : Omega -> Action -> Feature -> Real) (R delta : Real) (n : Nat) (omega : Omega) (hconfidence : matrixNorm (finiteHorizonScalarGram lambda feature n omega) (finiteHorizonRidgeEstimate (Matrix.scalar Feature lambda) feature response n omega - thetaStar) <= finiteHorizonScalarConfidenceRadius feature R delta lambda S n omega) (action : Action) : linearValue thetaStar (candidateFeature omega action) - linearValue thetaStar (candidateFeature omega (finiteHorizonScalarOptimisticAction lambda feature response candidateFeature R delta S n omega)) <= 2 * finiteHorizonScalarConfidenceRadius feature R…","missing":[],"search":"finitehorizonscalaroptimisticaction_gap_le banditrlproof.oful.finitehorizonscalaroptimisticaction_gap_le pointwise finite-horizon optimism: on the scalar confidence ellipsoid, every comparator's true linear value exceeds the selected value by at most twice the selected confidence bonus. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarOptimismViolationSet","label":"finiteHorizonScalarOptimismViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarOptimismViolationSet","description":"The event that some finite candidate action violates the one-step OFUL gap certificate.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-423f965f2241","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6203,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:295"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonScalarOptimismViolationSet [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (candidateFeature : Omega -> Action -> Feature -> Real) (R delta : Real) (n : Nat) : Set Omega","missing":[],"search":"finitehorizonscalaroptimismviolationset banditrlproof.oful.finitehorizonscalaroptimismviolationset the event that some finite candidate action violates the one-step oful gap certificate. definition compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizonScalarOptimismViolationSet_le","label":"measure_finiteHorizonScalarOptimismViolationSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizonScalarOptimismViolationSet_le","description":"Finite-action OFUL optimism violation has probability at most `delta`. No measurability of the candidate features or selected action is required: the violation event is included pointwise in the already controlled scalar confidence-ellipsoid bad event.","url":"../modules/banditrlproof-ofulfiniteactionoptimism/index.html#decl-b86f39b0e131","parent":"module:BanditRLProof.OFULFiniteActionOptimism","order":6204,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteActionOptimism"],["Source","BanditRLProof/OFULFiniteActionOptimism.lean:325"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonScalarOptimismViolationSet_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Finite Action] [Nonempty Action] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (candidateFeature : Omega -> Action -> Feature -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBou…","missing":[],"search":"measure_finitehorizonscalaroptimismviolationset_le banditrlproof.oful.measure_finitehorizonscalaroptimismviolationset_le finite-action oful optimism violation has probability at most `delta`. no measurability of the candidate features or selected action is required: the violation event is included pointwise in the already controlled scalar confidence-ellipsoid bad event. theorem compiled","shard":"modules/1e1baacfe4aed984.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonNoiseScore","label":"finiteHorizonNoiseScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonNoiseScore","description":"The finite-horizon martingale score `sum_{i<n} x_i eta_i`.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-972cbf28dcdb","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6205,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonNoiseScore [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (n : Nat) (omega : Omega) : EuclideanSpace Real Feature","missing":[],"search":"finitehorizonnoisescore banditrlproof.oful.finitehorizonnoisescore the finite-horizon martingale score `sum_{i<n} x_i eta_i`. definition compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.varianceWeightedFeature","label":"varianceWeightedFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.varianceWeightedFeature","description":"Feature rescaling whose ordinary Gram is weighted by the variance proxy.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-714069ad1f6c","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6206,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:29"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def varianceWeightedFeature (feature : Nat -> Omega -> Feature -> Real) (varianceProxy : Nat -> NNReal) (i : Nat) (omega : Omega) (j : Feature) : Real","missing":[],"search":"varianceweightedfeature banditrlproof.oful.varianceweightedfeature feature rescaling whose ordinary gram is weighted by the variance proxy. definition compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram","label":"finiteHorizonVarianceGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonVarianceGram","description":"The finite-horizon variance-weighted Gram `sum_{i<n} c_i x_i x_i^T`.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-0d2020f04d1d","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6207,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:36"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonVarianceGram [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (omega : Omega) : Matrix Feature Feature Real","missing":[],"search":"finitehorizonvariancegram banditrlproof.oful.finitehorizonvariancegram the finite-horizon variance-weighted gram `sum_{i<n} c_i x_i x_i^t`. definition compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.inner_finiteHorizonNoiseScore_eq_sum_projection_mul_noise","label":"inner_finiteHorizonNoiseScore_eq_sum_projection_mul_noise","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.inner_finiteHorizonNoiseScore_eq_sum_projection_mul_noise","description":"The score inner product is the finite sum of projected noise increments.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-78dffdde3be3","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6208,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:45"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem inner_finiteHorizonNoiseScore_eq_sum_projection_mul_noise [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (n : Nat) (omega : Omega) (theta : EuclideanSpace Real Feature) : ⟪finiteHorizonNoiseScore feature noise n omega, theta⟫_Real = (Finset.range n).sum (fun i => dotProduct (WithLp.ofLp theta) (feature i omega) * noise i omega)","missing":[],"search":"inner_finitehorizonnoisescore_eq_sum_projection_mul_noise banditrlproof.oful.inner_finitehorizonnoisescore_eq_sum_projection_mul_noise the score inner product is the finite sum of projected noise increments. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram_quadraticForm_eq_sum","label":"finiteHorizonVarianceGram_quadraticForm_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonVarianceGram_quadraticForm_eq_sum","description":"The weighted Gram quadratic form is the variance-weighted projection sum.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-d441497d66e8","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6209,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:71"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonVarianceGram_quadraticForm_eq_sum [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (omega : Omega) (theta : Feature -> Real) : quadraticForm (finiteHorizonVarianceGram feature varianceProxy n omega) theta = (Finset.range n).sum (fun i => (((varianceProxy i : NNReal) : Real)) * (dotProduct theta (feature i omega)) ^ 2)","missing":[],"search":"finitehorizonvariancegram_quadraticform_eq_sum banditrlproof.oful.finitehorizonvariancegram_quadraticform_eq_sum the weighted gram quadratic form is the variance-weighted projection sum. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram_posSemidef","label":"finiteHorizonVarianceGram_posSemidef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonVarianceGram_posSemidef","description":"The finite-horizon variance Gram is positive semidefinite.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-ec211d8330aa","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6210,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:100"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonVarianceGram_posSemidef [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (omega : Omega) : (finiteHorizonVarianceGram feature varianceProxy n omega).PosSemidef","missing":[],"search":"finitehorizonvariancegram_possemidef banditrlproof.oful.finitehorizonvariancegram_possemidef the finite-horizon variance gram is positive semidefinite. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonVarianceGram_apply","label":"measurable_finiteHorizonVarianceGram_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonVarianceGram_apply","description":"Coordinatewise measurability of the finite-horizon variance Gram.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-cc9192409110","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6211,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:119"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonVarianceGram_apply [MeasurableSpace Omega] [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (hfeature : forall i j, Measurable (fun omega => feature i omega j)) : forall j k, Measurable (fun omega => finiteHorizonVarianceGram feature varianceProxy n omega j k)","missing":[],"search":"measurable_finitehorizonvariancegram_apply banditrlproof.oful.measurable_finitehorizonvariancegram_apply coordinatewise measurability of the finite-horizon variance gram. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonNoiseScore","label":"measurable_finiteHorizonNoiseScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonNoiseScore","description":"Measurability of the finite-horizon score from coordinatewise inputs.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-a6aa85e709a3","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6212,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:135"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonNoiseScore [MeasurableSpace Omega] [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (n : Nat) (hfeature : forall i j, Measurable (fun omega => feature i omega j)) (hnoise : forall i, Measurable (noise i)) : Measurable (finiteHorizonNoiseScore feature noise n)","missing":[],"search":"measurable_finitehorizonnoisescore banditrlproof.oful.measurable_finitehorizonnoisescore measurability of the finite-horizon score from coordinatewise inputs. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.compensatedScore_eq_inner_sub_varianceGram","label":"compensatedScore_eq_inner_sub_varianceGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.compensatedScore_eq_inner_sub_varianceGram","description":"Pointwise identification of the fixed-direction compensated finite sum with the random score/random-Gram quadratic exponent.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-963529a0445e","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6213,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:150"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem compensatedScore_eq_inner_sub_varianceGram [Fintype Feature] [DecidableEq Feature] (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (omega : Omega) (theta : EuclideanSpace Real Feature) : (Finset.range (n + 1)).sum (fun t => match t with | 0 => 0 | i + 1 => dotProduct (WithLp.ofLp theta) (feature i omega) * noise i omega - (((varianceProxy i : NNReal) : Real) * (dotProduct (WithLp.ofLp theta) (feature i omega)) ^ 2 / 2)) = ⟪finiteHorizonNoiseScore feature noise n omega, theta⟫_Real - ⟪theta, Matrix.toEuclideanCLM (𝕜 := Real) (finiteHorizonVarianceGram feature varianceProxy n omega) theta⟫_Real / 2","missing":[],"search":"compensatedscore_eq_inner_sub_variancegram banditrlproof.oful.compensatedscore_eq_inner_sub_variancegram pointwise identification of the fixed-direction compensated finite sum with the random score/random-gram quadratic exponent. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScoreVarianceGram_hasMGFUpperBoundAt","label":"finiteHorizonScoreVarianceGram_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScoreVarianceGram_hasMGFUpperBoundAt","description":"The fixed-direction conditional-MGF endpoint transported to the random finite-horizon score and variance Gram.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-eeb591957b04","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6214,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:180"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonScoreVarianceGram_hasMGFUpperBoundAt [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (theta : EuclideanSpace Real Feature) (projectionBound : Nat -> Real) (hprojection : forall i, StronglyMeasurable[F i] (fun omega => dotProduct (WithLp.ofLp theta) (feature i omega))) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall i, 0 <= projectionBound i) (hprojectionBound : forall i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound i) (n : Nat) (hsubGaussian : forall i, i < n -> HasCondSubgaussianMGF (F i) (F.le i) (noise i) (…","missing":[],"search":"finitehorizonscorevariancegram_hasmgfupperboundat banditrlproof.oful.finitehorizonscorevariancegram_hasmgfupperboundat the fixed-direction conditional-mgf endpoint transported to the random finite-horizon score and variance gram. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_inner_finiteHorizonScore_sub_varianceGram_le_one","label":"integral_exp_inner_finiteHorizonScore_sub_varianceGram_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_inner_finiteHorizonScore_sub_varianceGram_le_one","description":"Explicit unit expectation bound for the score/Gram quadratic exponential.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-d1d89df53495","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6215,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:257"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_inner_finiteHorizonScore_sub_varianceGram_le_one [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (theta : EuclideanSpace Real Feature) (projectionBound : Nat -> Real) (hprojection : forall i, StronglyMeasurable[F i] (fun omega => dotProduct (WithLp.ofLp theta) (feature i omega))) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall i, 0 <= projectionBound i) (hprojectionBound : forall i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound i) (n : Nat) (hsubGaussian : forall i, i < n -> HasCondSubgaussianMGF (F i) (F.le i)…","missing":[],"search":"integral_exp_inner_finitehorizonscore_sub_variancegram_le_one banditrlproof.oful.integral_exp_inner_finitehorizonscore_sub_variancegram_le_one explicit unit expectation bound for the score/gram quadratic exponential. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_gaussianQuadraticExponential_finiteHorizon_le_one","label":"integral_gaussianQuadraticExponential_finiteHorizon_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_gaussianQuadraticExponential_finiteHorizon_le_one","description":"The fixed-direction expectation bound on the exact Gaussian-mixture consumer surface.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-a9251cb85b2d","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6216,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:299"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_gaussianQuadraticExponential_finiteHorizon_le_one [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (theta : EuclideanSpace Real Feature) (projectionBound : Nat -> Real) (hprojection : forall i, StronglyMeasurable[F i] (fun omega => dotProduct (WithLp.ofLp theta) (feature i omega))) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall i, 0 <= projectionBound i) (hprojectionBound : forall i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound i) (n : Nat) (hsubGaussian : forall i, i < n -> HasCondSubgaussianMGF (F i) (F.le i) (n…","missing":[],"search":"integral_gaussianquadraticexponential_finitehorizon_le_one banditrlproof.oful.integral_gaussianquadraticexponential_finitehorizon_le_one the fixed-direction expectation bound on the exact gaussian-mixture consumer surface. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_le_one","label":"lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_le_one","description":"Tonelli transport of all fixed-direction bounds through an arbitrary probability law on Gaussian directions.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-a7b336f66e98","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6217,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:338"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_le_one [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (nu : Measure (EuclideanSpace Real Feature)) [IsProbabilityMeasure nu] (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsub…","missing":[],"search":"lintegral_gaussianquadraticexponentialennreal_finitehorizon_prod_le_one banditrlproof.oful.lintegral_gaussianquadraticexponentialennreal_finitehorizon_prod_le_one tonelli transport of all fixed-direction bounds through an arbitrary probability law on gaussian directions. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_multivariateGaussian_zero_inv_le_one","label":"lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_multivariateGaussian_zero_inv_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_multivariateGaussian_zero_inv_le_one","description":"Gaussian-direction specialization of the finite-horizon Tonelli bound.","url":"../modules/banditrlproof-ofulfinitehorizonscoregram/index.html#decl-321b7f46f3ed","parent":"module:BanditRLProof.OFULFiniteHorizonScoreGram","order":6218,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULFiniteHorizonScoreGram"],["Source","BanditRLProof/OFULFiniteHorizonScoreGram.lean:456"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_multivariateGaussian_zero_inv_le_one [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsubGaussi…","missing":[],"search":"lintegral_gaussianquadraticexponentialennreal_finitehorizon_prod_multivariategaussian_zero_inv_le_one banditrlproof.oful.lintegral_gaussianquadraticexponentialennreal_finitehorizon_prod_multivariategaussian_zero_inv_le_one gaussian-direction specialization of the finite-horizon tonelli bound. theorem compiled","shard":"modules/9111c1c9811fcb37.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.inner_toEuclideanCLM_selfAdjoint","label":"inner_toEuclideanCLM_selfAdjoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.inner_toEuclideanCLM_selfAdjoint","description":"private theorem inner_toEuclideanCLM_selfAdjoint (C : Matrix Feature Feature Real) (hC : C.IsHermitian) (x y : EuclideanSpace Real Feature) : ⟪x, Matrix.toEuclideanCLM (𝕜 := Real) C y⟫_ℝ = ⟪Matrix.toEuclideanCLM (𝕜 := Real) C x, y⟫_ℝ","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-16f9baa9c3e9","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6219,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem inner_toEuclideanCLM_selfAdjoint (C : Matrix Feature Feature Real) (hC : C.IsHermitian) (x y : EuclideanSpace Real Feature) : ⟪x, Matrix.toEuclideanCLM (𝕜 := Real) C y⟫_ℝ = ⟪Matrix.toEuclideanCLM (𝕜 := Real) C x, y⟫_ℝ","missing":[],"search":"inner_toeuclideanclm_selfadjoint banditrlproof.oful.inner_toeuclideanclm_selfadjoint private theorem inner_toeuclideanclm_selfadjoint (c : matrix feature feature real) (hc : c.ishermitian) (x y : euclideanspace real feature) : ⟪x, matrix.toeuclideanclm (𝕜 := real) c y⟫_ℝ = ⟪matrix.toeuclideanclm (𝕜 := real) c x, y⟫_ℝ theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.inner_toEuclideanCLM_congruence","label":"inner_toEuclideanCLM_congruence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.inner_toEuclideanCLM_congruence","description":"private theorem inner_toEuclideanCLM_congruence (C G : Matrix Feature Feature Real) (hC : C.IsHermitian) (z : EuclideanSpace Real Feature) : ⟪Matrix.toEuclideanCLM (𝕜 := Real) C z, Matrix.toEuclideanCLM (𝕜 := Real) G (Matrix.toEuclideanCLM (𝕜 := Real) C z)⟫_ℝ = ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) (C * G * C) z⟫_ℝ","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-e1bb75232ec8","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6220,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:35"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem inner_toEuclideanCLM_congruence (C G : Matrix Feature Feature Real) (hC : C.IsHermitian) (z : EuclideanSpace Real Feature) : ⟪Matrix.toEuclideanCLM (𝕜 := Real) C z, Matrix.toEuclideanCLM (𝕜 := Real) G (Matrix.toEuclideanCLM (𝕜 := Real) C z)⟫_ℝ = ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) (C * G * C) z⟫_ℝ","missing":[],"search":"inner_toeuclideanclm_congruence banditrlproof.oful.inner_toeuclideanclm_congruence private theorem inner_toeuclideanclm_congruence (c g : matrix feature feature real) (hc : c.ishermitian) (z : euclideanspace real feature) : ⟪matrix.toeuclideanclm (𝕜 := real) c z, matrix.toeuclideanclm (𝕜 := real) g (matrix.toeuclideanclm (𝕜 := real) c z)⟫_ℝ = ⟪z, matrix.toeuclideanclm (𝕜 := real) (c * g * c) z⟫_ℝ theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_transformed","label":"integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_transformed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_transformed","description":"private theorem integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_transformed (V0 G : Matrix Feature Feature Real) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : let C := CFC.sqrt V0⁻¹ integral (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) = (Real.sqrt (Matrix.det (1 + C * G…","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-73436c508b77","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6221,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:57"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_transformed (V0 G : Matrix Feature Feature Real) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : let C := CFC.sqrt V0⁻¹ integral (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) = (Real.sqrt (Matrix.det (1 + C * G * C)))⁻¹ * Real.exp ((Matrix.toEuclideanCLM (𝕜 := Real) C score) ⬝ᵥ (1 + C * G * C)⁻¹.mulVec (Matrix.toEuclideanCLM (𝕜 := Real) C score) / 2)","missing":[],"search":"integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv_transformed banditrlproof.oful.integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv_transformed private theorem integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv_transformed (v0 g : matrix feature feature real) (hg : g.possemidef) (score : euclideanspace real feature) : let c := cfc.sqrt v0⁻¹ integral (probabilitytheory.multivariategaussian 0 v0⁻¹) (fun z : euclideanspace real feature => real.exp (⟪score, z⟫_ℝ - ⟪z, matrix.toeuclideanclm (𝕜 := real) g z⟫_ℝ / 2)) = (real.sqrt (matrix.det (1 + c * g * c)))⁻¹ * real.exp ((matrix.toeuclideanclm (𝕜 := real) c score) ⬝ᵥ (1 + c * g * c)⁻¹.mulvec (matrix.toeuclideanclm (𝕜 := real) c score) / 2) theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sqrt_inv_congruence_factorization","label":"sqrt_inv_congruence_factorization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sqrt_inv_congruence_factorization","description":"private theorem sqrt_inv_congruence_factorization (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) : let D := CFC.sqrt V0 let C := CFC.sqrt V0⁻¹ D * (1 + C * G * C) * D = V0 + G","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-18b05bc96533","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6222,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:103"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_inv_congruence_factorization (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) : let D := CFC.sqrt V0 let C := CFC.sqrt V0⁻¹ D * (1 + C * G * C) * D = V0 + G","missing":[],"search":"sqrt_inv_congruence_factorization banditrlproof.oful.sqrt_inv_congruence_factorization private theorem sqrt_inv_congruence_factorization (v0 g : matrix feature feature real) (hv0 : v0.posdef) : let d := cfc.sqrt v0 let c := cfc.sqrt v0⁻¹ d * (1 + c * g * c) * d = v0 + g theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_one_add_sqrt_inv_congruence_eq_ratio","label":"det_one_add_sqrt_inv_congruence_eq_ratio","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_one_add_sqrt_inv_congruence_eq_ratio","description":"theorem det_one_add_sqrt_inv_congruence_eq_ratio (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) : let C := CFC.sqrt V0⁻¹ Matrix.det (1 + C * G * C) = Matrix.det (V0 + G) / Matrix.det V0","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-feafcc13fbcb","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6223,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:130"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_one_add_sqrt_inv_congruence_eq_ratio (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) : let C := CFC.sqrt V0⁻¹ Matrix.det (1 + C * G * C) = Matrix.det (V0 + G) / Matrix.det V0","missing":[],"search":"det_one_add_sqrt_inv_congruence_eq_ratio banditrlproof.oful.det_one_add_sqrt_inv_congruence_eq_ratio theorem det_one_add_sqrt_inv_congruence_eq_ratio (v0 g : matrix feature feature real) (hv0 : v0.posdef) : let c := cfc.sqrt v0⁻¹ matrix.det (1 + c * g * c) = matrix.det (v0 + g) / matrix.det v0 theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sqrt_inv_congruence_inverse","label":"sqrt_inv_congruence_inverse","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sqrt_inv_congruence_inverse","description":"theorem sqrt_inv_congruence_inverse (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) : let C := CFC.sqrt V0⁻¹ C * (1 + C * G * C)⁻¹ * C = (V0 + G)⁻¹","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-f0a5e42b522e","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6224,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:162"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sqrt_inv_congruence_inverse (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) : let C := CFC.sqrt V0⁻¹ C * (1 + C * G * C)⁻¹ * C = (V0 + G)⁻¹","missing":[],"search":"sqrt_inv_congruence_inverse banditrlproof.oful.sqrt_inv_congruence_inverse theorem sqrt_inv_congruence_inverse (v0 g : matrix feature feature real) (hv0 : v0.posdef) : let c := cfc.sqrt v0⁻¹ c * (1 + c * g * c)⁻¹ * c = (v0 + g)⁻¹ theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.dotProduct_sqrt_inv_congruence_inverse","label":"dotProduct_sqrt_inv_congruence_inverse","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.dotProduct_sqrt_inv_congruence_inverse","description":"theorem dotProduct_sqrt_inv_congruence_inverse (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (score : EuclideanSpace Real Feature) : let C := CFC.sqrt V0⁻¹ (Matrix.toEuclideanCLM (𝕜 := Real) C score) ⬝ᵥ (1 + C * G * C)⁻¹.mulVec (Matrix.toEuclideanCLM (𝕜 := Real) C score) = score ⬝ᵥ (V0 + G)⁻¹.mulVec score","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-eea29733dd02","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6225,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:187"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem dotProduct_sqrt_inv_congruence_inverse (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (score : EuclideanSpace Real Feature) : let C := CFC.sqrt V0⁻¹ (Matrix.toEuclideanCLM (𝕜 := Real) C score) ⬝ᵥ (1 + C * G * C)⁻¹.mulVec (Matrix.toEuclideanCLM (𝕜 := Real) C score) = score ⬝ᵥ (V0 + G)⁻¹.mulVec score","missing":[],"search":"dotproduct_sqrt_inv_congruence_inverse banditrlproof.oful.dotproduct_sqrt_inv_congruence_inverse theorem dotproduct_sqrt_inv_congruence_inverse (v0 g : matrix feature feature real) (hv0 : v0.posdef) (score : euclideanspace real feature) : let c := cfc.sqrt v0⁻¹ (matrix.toeuclideanclm (𝕜 := real) c score) ⬝ᵥ (1 + c * g * c)⁻¹.mulvec (matrix.toeuclideanclm (𝕜 := real) c score) = score ⬝ᵥ (v0 + g)⁻¹.mulvec score theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_detRatio","label":"integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_detRatio","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_detRatio","description":"theorem integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_detRatio (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : integral (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) = (Real.sqrt (Matrix.det (V0 + G) / Matrix.det V0))…","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-230e44a946eb","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6226,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:222"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_detRatio (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : integral (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) = (Real.sqrt (Matrix.det (V0 + G) / Matrix.det V0))⁻¹ * Real.exp (score ⬝ᵥ (V0 + G)⁻¹.mulVec score / 2)","missing":[],"search":"integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv_detratio banditrlproof.oful.integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv_detratio theorem integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv_detratio (v0 g : matrix feature feature real) (hv0 : v0.posdef) (hg : g.possemidef) (score : euclideanspace real feature) : integral (probabilitytheory.multivariategaussian 0 v0⁻¹) (fun z : euclideanspace real feature => real.exp (⟪score, z⟫_ℝ - ⟪z, matrix.toeuclideanclm (𝕜 := real) g z⟫_ℝ / 2)) = (real.sqrt (matrix.det (v0 + g) / matrix.det v0))⁻¹ * real.exp (score ⬝ᵥ (v0 + g)⁻¹.mulvec score / 2) theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","label":"integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","description":"theorem integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : integral (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) = Real.sqrt (Matrix.det V0 / Matrix.det (V0 + G)) * Real.exp…","url":"../modules/banditrlproof-ofulgaussiancovariancemixture/index.html#decl-523d9ec812c9","parent":"module:BanditRLProof.OFULGaussianCovarianceMixture","order":6227,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianCovarianceMixture"],["Source","BanditRLProof/OFULGaussianCovarianceMixture.lean:240"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : integral (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) = Real.sqrt (Matrix.det V0 / Matrix.det (V0 + G)) * Real.exp (score ⬝ᵥ (V0 + G)⁻¹.mulVec score / 2)","missing":[],"search":"integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv banditrlproof.oful.integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv theorem integral_exp_inner_sub_quadratic_multivariategaussian_zero_inv (v0 g : matrix feature feature real) (hv0 : v0.posdef) (hg : g.possemidef) (score : euclideanspace real feature) : integral (probabilitytheory.multivariategaussian 0 v0⁻¹) (fun z : euclideanspace real feature => real.exp (⟪score, z⟫_ℝ - ⟪z, matrix.toeuclideanclm (𝕜 := real) g z⟫_ℝ / 2)) = real.sqrt (matrix.det v0 / matrix.det (v0 + g)) * real.exp (score ⬝ᵥ (v0 + g)⁻¹.mulvec score / 2) theorem compiled","shard":"modules/183124b07255e581.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integrable_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","label":"integrable_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integrable_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","description":"The positive-semidefinite Gaussian quadratic exponential is integrable. Fernique supplies a square-exponential envelope for the linear term.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html#decl-f642cebbf466","parent":"module:BanditRLProof.OFULGaussianEvaluatedMixture","order":6228,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianEvaluatedMixture"],["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_inner_sub_quadratic_multivariateGaussian_zero_inv [Fintype Feature] [DecidableEq Feature] (V0 G : Matrix Feature Feature Real) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : Integrable (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) (ProbabilityTheory.multivariateGaussian 0 V0⁻¹)","missing":[],"search":"integrable_exp_inner_sub_quadratic_multivariategaussian_zero_inv banditrlproof.oful.integrable_exp_inner_sub_quadratic_multivariategaussian_zero_inv the positive-semidefinite gaussian quadratic exponential is integrable. fernique supplies a square-exponential envelope for the linear term. theorem compiled","shard":"modules/4479e55dbdf79eaf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_multivariateGaussian_zero_inv","label":"lintegral_gaussianQuadraticExponentialENNReal_multivariateGaussian_zero_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_multivariateGaussian_zero_inv","description":"The `ENNReal` Gaussian direction integral equals the completed-square determinant-ratio expression.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html#decl-ba2578c20b12","parent":"module:BanditRLProof.OFULGaussianEvaluatedMixture","order":6229,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianEvaluatedMixture"],["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean:103"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_gaussianQuadraticExponentialENNReal_multivariateGaussian_zero_inv [Fintype Feature] [DecidableEq Feature] (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) : ∫⁻ z : EuclideanSpace Real Feature, ENNReal.ofReal (Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) G z⟫_ℝ / 2)) ∂ProbabilityTheory.multivariateGaussian 0 V0⁻¹ = ENNReal.ofReal (Real.sqrt (Matrix.det V0 / Matrix.det (V0 + G)) * Real.exp (score ⬝ᵥ (V0 + G)⁻¹.mulVec score / 2))","missing":[],"search":"lintegral_gaussianquadraticexponentialennreal_multivariategaussian_zero_inv banditrlproof.oful.lintegral_gaussianquadraticexponentialennreal_multivariategaussian_zero_inv the `ennreal` gaussian direction integral equals the completed-square determinant-ratio expression. theorem compiled","shard":"modules/4479e55dbdf79eaf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonDetRatioInvGramExponential","label":"finiteHorizonDetRatioInvGramExponential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonDetRatioInvGramExponential","description":"The evaluated determinant-ratio inverse-Gram exponential for a finite-horizon score and variance Gram.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html#decl-b9f861de5aee","parent":"module:BanditRLProof.OFULGaussianEvaluatedMixture","order":6230,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGaussianEvaluatedMixture"],["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean:128"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonDetRatioInvGramExponential [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (omega : Omega) : ENNReal","missing":[],"search":"finitehorizondetratioinvgramexponential banditrlproof.oful.finitehorizondetratioinvgramexponential the evaluated determinant-ratio inverse-gram exponential for a finite-horizon score and variance gram. definition compiled","shard":"modules/4479e55dbdf79eaf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_multivariateGaussian_zero_inv","label":"lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_multivariateGaussian_zero_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_multivariateGaussian_zero_inv","description":"Samplewise evaluation of the finite-horizon Gaussian direction integral.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html#decl-d9ddbb7c5e5d","parent":"module:BanditRLProof.OFULGaussianEvaluatedMixture","order":6231,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianEvaluatedMixture"],["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean:151"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_multivariateGaussian_zero_inv [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (omega : Omega) : ∫⁻ theta : EuclideanSpace Real Feature, gaussianQuadraticExponentialENNReal (finiteHorizonNoiseScore feature noise n) (finiteHorizonVarianceGram feature varianceProxy n) (omega, theta) ∂ProbabilityTheory.multivariateGaussian 0 V0⁻¹ = finiteHorizonDetRatioInvGramExponential V0 feature noise varianceProxy n omega","missing":[],"search":"lintegral_gaussianquadraticexponentialennreal_finitehorizon_multivariategaussian_zero_inv banditrlproof.oful.lintegral_gaussianquadraticexponentialennreal_finitehorizon_multivariategaussian_zero_inv samplewise evaluation of the finite-horizon gaussian direction integral. theorem compiled","shard":"modules/4479e55dbdf79eaf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonDetRatioInvGramExponential","label":"measurable_finiteHorizonDetRatioInvGramExponential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHorizonDetRatioInvGramExponential","description":"The evaluated finite-horizon mixture is measurable in the sample.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html#decl-7139eed0903f","parent":"module:BanditRLProof.OFULGaussianEvaluatedMixture","order":6232,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianEvaluatedMixture"],["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean:180"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHorizonDetRatioInvGramExponential [MeasurableSpace Omega] [Fintype Feature] [DecidableEq Feature] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (n : Nat) (hfeature : forall i j, Measurable (fun omega => feature i omega j)) (hnoise : forall i, Measurable (noise i)) : Measurable (finiteHorizonDetRatioInvGramExponential V0 feature noise varianceProxy n)","missing":[],"search":"measurable_finitehorizondetratioinvgramexponential banditrlproof.oful.measurable_finitehorizondetratioinvgramexponential the evaluated finite-horizon mixture is measurable in the sample. theorem compiled","shard":"modules/4479e55dbdf79eaf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_finiteHorizon_detRatio_invGramExponential_le_one","label":"lintegral_finiteHorizon_detRatio_invGramExponential_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_finiteHorizon_detRatio_invGramExponential_le_one","description":"The evaluated finite-horizon Gaussian mixture has `ENNReal` expectation at most one.","url":"../modules/banditrlproof-ofulgaussianevaluatedmixture/index.html#decl-6a182495852d","parent":"module:BanditRLProof.OFULGaussianEvaluatedMixture","order":6233,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianEvaluatedMixture"],["Source","BanditRLProof/OFULGaussianEvaluatedMixture.lean:230"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_finiteHorizon_detRatio_invGramExponential_le_one [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsubGaussian : forall i, i < n -> H…","missing":[],"search":"lintegral_finitehorizon_detratio_invgramexponential_le_one banditrlproof.oful.lintegral_finitehorizon_detratio_invgramexponential_le_one the evaluated finite-horizon gaussian mixture has `ennreal` expectation at most one. theorem compiled","shard":"modules/4479e55dbdf79eaf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_linear_sub_quadratic_gaussianReal_zero_one","label":"integral_exp_linear_sub_quadratic_gaussianReal_zero_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_linear_sub_quadratic_gaussianReal_zero_one","description":"The exact scalar quadratic-exponential integral under a standard Gaussian. This is the one-coordinate completed-square identity used by the Gaussian method of mixtures. The nonnegative quadratic coefficient is the scalar eigenvalue contract that will arise from a positive-semidefinite Gram matrix.","url":"../modules/banditrlproof-ofulgaussianmixture/index.html#decl-b75c724a450e","parent":"module:BanditRLProof.OFULGaussianMixture","order":6234,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixture"],["Source","BanditRLProof/OFULGaussianMixture.lean:27"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_linear_sub_quadratic_gaussianReal_zero_one (s q : Real) (hq : 0 <= q) : integral (ProbabilityTheory.gaussianReal 0 1) (fun x : Real => Real.exp (s * x - q * x ^ 2 / 2)) = (Real.sqrt (1 + q))⁻¹ * Real.exp (s ^ 2 / (2 * (1 + q)))","missing":[],"search":"integral_exp_linear_sub_quadratic_gaussianreal_zero_one banditrlproof.oful.integral_exp_linear_sub_quadratic_gaussianreal_zero_one the exact scalar quadratic-exponential integral under a standard gaussian. this is the one-coordinate completed-square identity used by the gaussian method of mixtures. the nonnegative quadratic coefficient is the scalar eigenvalue contract that will arise from a positive-semidefinite gram matrix. theorem compiled","shard":"modules/d5becfe86da00ed6.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal","label":"integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal","description":"Finite independent-coordinate version of the scalar completed-square identity. The right-hand side is deliberately left as a finite product. The following matrix-facing lemmas identify that product with the square-root determinant factor and the exponential term with an inverse-diagonal quadratic form.","url":"../modules/banditrlproof-ofulgaussianmixture/index.html#decl-036a09491f90","parent":"module:BanditRLProof.OFULGaussianMixture","order":6235,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixture"],["Source","BanditRLProof/OFULGaussianMixture.lean:120"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal {Feature : Type*} [Fintype Feature] (score quadratic : Feature -> Real) (hquadratic : forall i, 0 <= quadratic i) : integral (Measure.pi (fun _ : Feature => ProbabilityTheory.gaussianReal 0 1)) (fun z : Feature -> Real => Real.exp (Finset.univ.sum (fun i => score i * z i - quadratic i * (z i) ^ 2 / 2))) = Finset.univ.prod (fun i => (Real.sqrt (1 + quadratic i))⁻¹ * Real.exp (score i ^ 2 / (2 * (1 + quadratic i))))","missing":[],"search":"integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianreal banditrlproof.oful.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianreal finite independent-coordinate version of the scalar completed-square identity. the right-hand side is deliberately left as a finite product. the following matrix-facing lemmas identify that product with the square-root determinant factor and the exponential term with an inverse-diagonal quadratic form. theorem compiled","shard":"modules/d5becfe86da00ed6.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_eq","label":"integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_eq","description":"Finite product identity with the normalization and exponential factors collected separately.","url":"../modules/banditrlproof-ofulgaussianmixture/index.html#decl-60cea2a11e84","parent":"module:BanditRLProof.OFULGaussianMixture","order":6236,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixture"],["Source","BanditRLProof/OFULGaussianMixture.lean:161"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_eq {Feature : Type*} [Fintype Feature] (score quadratic : Feature -> Real) (hquadratic : forall i, 0 <= quadratic i) : integral (Measure.pi (fun _ : Feature => ProbabilityTheory.gaussianReal 0 1)) (fun z : Feature -> Real => Real.exp (Finset.univ.sum (fun i => score i * z i - quadratic i * (z i) ^ 2 / 2))) = (Finset.univ.prod (fun i => Real.sqrt (1 + quadratic i)))⁻¹ * Real.exp (Finset.univ.sum (fun i => score i ^ 2 / (2 * (1 + quadratic i))))","missing":[],"search":"integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianreal_eq banditrlproof.oful.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianreal_eq finite product identity with the normalization and exponential factors collected separately. theorem compiled","shard":"modules/d5becfe86da00ed6.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_det","label":"integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_det","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_det","description":"Diagonal matrix form of the finite Gaussian quadratic-exponential identity. This is the diagonal-coordinate determinant-ratio and inverse-quadratic expression intended for later transport through a PSD matrix eigenbasis.","url":"../modules/banditrlproof-ofulgaussianmixture/index.html#decl-18d038f51fca","parent":"module:BanditRLProof.OFULGaussianMixture","order":6237,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixture"],["Source","BanditRLProof/OFULGaussianMixture.lean:187"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_det {Feature : Type*} [Fintype Feature] [DecidableEq Feature] (score quadratic : Feature -> Real) (hquadratic : forall i, 0 <= quadratic i) : integral (Measure.pi (fun _ : Feature => ProbabilityTheory.gaussianReal 0 1)) (fun z : Feature -> Real => Real.exp (Finset.univ.sum (fun i => score i * z i - quadratic i * (z i) ^ 2 / 2))) = (Real.sqrt (Matrix.det (Matrix.diagonal (fun i : Feature => 1 + quadratic i))))⁻¹ * Real.exp (score ⬝ᵥ ((Matrix.diagonal (fun i : Feature => 1 + quadratic i))⁻¹).mulVec score / 2)","missing":[],"search":"integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianreal_det banditrlproof.oful.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianreal_det diagonal matrix form of the finite gaussian quadratic-exponential identity. this is the diagonal-coordinate determinant-ratio and inverse-quadratic expression intended for later transport through a psd matrix eigenbasis. theorem compiled","shard":"modules/d5becfe86da00ed6.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.gaussianQuadraticExponential","label":"gaussianQuadraticExponential","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.gaussianQuadraticExponential","description":"The quadratic exponential mixed over the Gaussian direction in OFUL.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-6adb4f0e384b","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6238,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianQuadraticExponential (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (p : Omega × EuclideanSpace Real Feature) : Real","missing":[],"search":"gaussianquadraticexponential banditrlproof.oful.gaussianquadraticexponential the quadratic exponential mixed over the gaussian direction in oful. definition compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_gaussianQuadraticExponential_dot","label":"measurable_gaussianQuadraticExponential_dot","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_gaussianQuadraticExponential_dot","description":"private theorem measurable_gaussianQuadraticExponential_dot (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (hscore : Measurable score) (hgram : forall i j, Measurable (fun omega => gram omega i j)) : Measurable (fun p : Omega × EuclideanSpace Real Feature => Real.exp ((WithLp.ofLp (score p.1)) ⬝ᵥ (WithLp.ofLp p.2) - (WithLp.ofLp p.2) ⬝ᵥ (gram p.1).mulVec (WithLp.ofLp p.2…","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-93928c61294d","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6239,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:33"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem measurable_gaussianQuadraticExponential_dot (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (hscore : Measurable score) (hgram : forall i j, Measurable (fun omega => gram omega i j)) : Measurable (fun p : Omega × EuclideanSpace Real Feature => Real.exp ((WithLp.ofLp (score p.1)) ⬝ᵥ (WithLp.ofLp p.2) - (WithLp.ofLp p.2) ⬝ᵥ (gram p.1).mulVec (WithLp.ofLp p.2) / 2))","missing":[],"search":"measurable_gaussianquadraticexponential_dot banditrlproof.oful.measurable_gaussianquadraticexponential_dot private theorem measurable_gaussianquadraticexponential_dot (score : omega -> euclideanspace real feature) (gram : omega -> matrix feature feature real) (hscore : measurable score) (hgram : forall i j, measurable (fun omega => gram omega i j)) : measurable (fun p : omega × euclideanspace real feature => real.exp ((withlp.oflp (score p.1)) ⬝ᵥ (withlp.oflp p.2) - (withlp.oflp p.2) ⬝ᵥ (gram p.1).mulvec (withlp.oflp p.2) / 2)) theorem compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_gaussianQuadraticExponential","label":"measurable_gaussianQuadraticExponential","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_gaussianQuadraticExponential","description":"Joint measurability of the quadratic exponential from measurable random scores and coordinatewise measurable random Gram matrices.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-40865d5a168f","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6240,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:51"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_gaussianQuadraticExponential (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (hscore : Measurable score) (hgram : forall i j, Measurable (fun omega => gram omega i j)) : Measurable (gaussianQuadraticExponential score gram)","missing":[],"search":"measurable_gaussianquadraticexponential banditrlproof.oful.measurable_gaussianquadraticexponential joint measurability of the quadratic exponential from measurable random scores and coordinatewise measurable random gram matrices. theorem compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.gaussianQuadraticExponentialENNReal","label":"gaussianQuadraticExponentialENNReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.gaussianQuadraticExponentialENNReal","description":"The nonnegative extended-real surface used by Tonelli.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-c4f32cafd87d","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6241,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:68"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def gaussianQuadraticExponentialENNReal (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (p : Omega × EuclideanSpace Real Feature) : ENNReal","missing":[],"search":"gaussianquadraticexponentialennreal banditrlproof.oful.gaussianquadraticexponentialennreal the nonnegative extended-real surface used by tonelli. definition compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_gaussianQuadraticExponentialENNReal","label":"measurable_gaussianQuadraticExponentialENNReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_gaussianQuadraticExponentialENNReal","description":"Joint measurability of the `ENNReal` Tonelli surface.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-0f3a3e1501c7","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6242,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:75"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_gaussianQuadraticExponentialENNReal (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (hscore : Measurable score) (hgram : forall i j, Measurable (fun omega => gram omega i j)) : Measurable (gaussianQuadraticExponentialENNReal score gram)","missing":[],"search":"measurable_gaussianquadraticexponentialennreal banditrlproof.oful.measurable_gaussianquadraticexponentialennreal joint measurability of the `ennreal` tonelli surface. theorem compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_prod","label":"lintegral_gaussianQuadraticExponentialENNReal_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_prod","description":"Tonelli for the quadratic exponential under an arbitrary `SFinite` parameter law.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-4985607ddc37","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6243,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:84"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_gaussianQuadraticExponentialENNReal_prod (mu : Measure Omega) (nu : Measure (EuclideanSpace Real Feature)) [SFinite nu] (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (hscore : Measurable score) (hgram : forall i j, Measurable (fun omega => gram omega i j)) : ∫⁻ p, gaussianQuadraticExponentialENNReal score gram p ∂mu.prod nu = ∫⁻ omega, ∫⁻ theta, gaussianQuadraticExponentialENNReal score gram (omega, theta) ∂nu ∂mu","missing":[],"search":"lintegral_gaussianquadraticexponentialennreal_prod banditrlproof.oful.lintegral_gaussianquadraticexponentialennreal_prod tonelli for the quadratic exponential under an arbitrary `sfinite` parameter law. theorem compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_prod_multivariateGaussian_zero_inv","label":"lintegral_gaussianQuadraticExponentialENNReal_prod_multivariateGaussian_zero_inv","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_prod_multivariateGaussian_zero_inv","description":"Tonelli specialized to the `N(0, V0⁻¹)` parameter law used by OFUL.","url":"../modules/banditrlproof-ofulgaussianmixturemeasurability/index.html#decl-612994a83f22","parent":"module:BanditRLProof.OFULGaussianMixtureMeasurability","order":6244,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianMixtureMeasurability"],["Source","BanditRLProof/OFULGaussianMixtureMeasurability.lean:99"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem lintegral_gaussianQuadraticExponentialENNReal_prod_multivariateGaussian_zero_inv (mu : Measure Omega) (V0 : Matrix Feature Feature Real) (score : Omega -> EuclideanSpace Real Feature) (gram : Omega -> Matrix Feature Feature Real) (hscore : Measurable score) (hgram : forall i j, Measurable (fun omega => gram omega i j)) : ∫⁻ p, gaussianQuadraticExponentialENNReal score gram p ∂mu.prod (ProbabilityTheory.multivariateGaussian 0 V0⁻¹) = ∫⁻ omega, ∫⁻ theta, gaussianQuadraticExponentialENNReal score gram (omega, theta) ∂ProbabilityTheory.multivariateGaussian 0 V0⁻¹ ∂mu","missing":[],"search":"lintegral_gaussianquadraticexponentialennreal_prod_multivariategaussian_zero_inv banditrlproof.oful.lintegral_gaussianquadraticexponentialennreal_prod_multivariategaussian_zero_inv tonelli specialized to the `n(0, v0⁻¹)` parameter law used by oful. theorem compiled","shard":"modules/4559641663483d22.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.inner_sum_smul_orthonormalBasis","label":"inner_sum_smul_orthonormalBasis","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.inner_sum_smul_orthonormalBasis","description":"private theorem inner_sum_smul_orthonormalBasis (b : OrthonormalBasis Feature Real (EuclideanSpace Real Feature)) (x : Feature -> Real) (i : Feature) : ⟪b i, ∑ j, x j • b j⟫_ℝ = x i","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-f5ac66f6401f","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6245,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem inner_sum_smul_orthonormalBasis (b : OrthonormalBasis Feature Real (EuclideanSpace Real Feature)) (x : Feature -> Real) (i : Feature) : ⟪b i, ∑ j, x j • b j⟫_ℝ = x i","missing":[],"search":"inner_sum_smul_orthonormalbasis banditrlproof.oful.inner_sum_smul_orthonormalbasis private theorem inner_sum_smul_orthonormalbasis (b : orthonormalbasis feature real (euclideanspace real feature)) (x : feature -> real) (i : feature) : ⟪b i, ∑ j, x j • b j⟫_ℝ = x i theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.inner_sum_smul_orthonormalBasis_left","label":"inner_sum_smul_orthonormalBasis_left","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.inner_sum_smul_orthonormalBasis_left","description":"private theorem inner_sum_smul_orthonormalBasis_left (b : OrthonormalBasis Feature Real (EuclideanSpace Real Feature)) (score : EuclideanSpace Real Feature) (x : Feature -> Real) : ⟪score, ∑ i, x i • b i⟫_ℝ = ∑ i, ⟪score, b i⟫_ℝ * x i","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-b6ae794884af","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6246,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:28"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem inner_sum_smul_orthonormalBasis_left (b : OrthonormalBasis Feature Real (EuclideanSpace Real Feature)) (score : EuclideanSpace Real Feature) (x : Feature -> Real) : ⟪score, ∑ i, x i • b i⟫_ℝ = ∑ i, ⟪score, b i⟫_ℝ * x i","missing":[],"search":"inner_sum_smul_orthonormalbasis_left banditrlproof.oful.inner_sum_smul_orthonormalbasis_left private theorem inner_sum_smul_orthonormalbasis_left (b : orthonormalbasis feature real (euclideanspace real feature)) (score : euclideanspace real feature) (x : feature -> real) : ⟪score, ∑ i, x i • b i⟫_ℝ = ∑ i, ⟪score, b i⟫_ℝ * x i theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.inner_toEuclideanCLM_sum_eigenvectorBasis","label":"inner_toEuclideanCLM_sum_eigenvectorBasis","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.inner_toEuclideanCLM_sum_eigenvectorBasis","description":"private theorem inner_toEuclideanCLM_sum_eigenvectorBasis (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (x : Feature -> Real) : let b := hA.isHermitian.eigenvectorBasis let v : EuclideanSpace Real Feature := ∑ i, x i • b i ⟪v, Matrix.toEuclideanCLM (𝕜 := Real) A v⟫_ℝ = ∑ i, hA.isHermitian.eigenvalues i * x i ^ 2","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-3be3c78c959c","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6247,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:39"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem inner_toEuclideanCLM_sum_eigenvectorBasis (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (x : Feature -> Real) : let b := hA.isHermitian.eigenvectorBasis let v : EuclideanSpace Real Feature := ∑ i, x i • b i ⟪v, Matrix.toEuclideanCLM (𝕜 := Real) A v⟫_ℝ = ∑ i, hA.isHermitian.eigenvalues i * x i ^ 2","missing":[],"search":"inner_toeuclideanclm_sum_eigenvectorbasis banditrlproof.oful.inner_toeuclideanclm_sum_eigenvectorbasis private theorem inner_toeuclideanclm_sum_eigenvectorbasis (a : matrix feature feature real) (ha : a.possemidef) (x : feature -> real) : let b := ha.ishermitian.eigenvectorbasis let v : euclideanspace real feature := ∑ i, x i • b i ⟪v, matrix.toeuclideanclm (𝕜 := real) a v⟫_ℝ = ∑ i, ha.ishermitian.eigenvalues i * x i ^ 2 theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_stdGaussian_eigenvalues","label":"integral_exp_inner_sub_quadratic_stdGaussian_eigenvalues","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_stdGaussian_eigenvalues","description":"The standard-Gaussian quadratic-exponential integral in the orthonormal eigenbasis of a positive-semidefinite matrix. This is the explicit spectral-coordinate transport bridge from the diagonal product identity. The following lemmas collect its right-hand side into matrix determinant and inverse-quadratic notation.","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-34a06e089ffa","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6248,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:108"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_inner_sub_quadratic_stdGaussian_eigenvalues (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (score : EuclideanSpace Real Feature) : integral (ProbabilityTheory.stdGaussian (EuclideanSpace Real Feature)) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) A z⟫_ℝ / 2)) = (Finset.univ.prod (fun i => Real.sqrt (1 + hA.isHermitian.eigenvalues i)))⁻¹ * Real.exp (Finset.univ.sum (fun i => ⟪score, hA.isHermitian.eigenvectorBasis i⟫_ℝ ^ 2 / (2 * (1 + hA.isHermitian.eigenvalues i))))","missing":[],"search":"integral_exp_inner_sub_quadratic_stdgaussian_eigenvalues banditrlproof.oful.integral_exp_inner_sub_quadratic_stdgaussian_eigenvalues the standard-gaussian quadratic-exponential integral in the orthonormal eigenbasis of a positive-semidefinite matrix. this is the explicit spectral-coordinate transport bridge from the diagonal product identity. the following lemmas collect its right-hand side into matrix determinant and inverse-quadratic notation. theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_one_add_posSemidef_eq_prod_eigenvalues","label":"det_one_add_posSemidef_eq_prod_eigenvalues","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_one_add_posSemidef_eq_prod_eigenvalues","description":"The determinant of `1 + A` is the product of one plus the eigenvalues of a real positive-semidefinite matrix.","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-10392f97796d","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6249,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:175"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_one_add_posSemidef_eq_prod_eigenvalues (A : Matrix Feature Feature Real) (hA : A.PosSemidef) : Matrix.det (1 + A) = Finset.univ.prod (fun i => 1 + hA.isHermitian.eigenvalues i)","missing":[],"search":"det_one_add_possemidef_eq_prod_eigenvalues banditrlproof.oful.det_one_add_possemidef_eq_prod_eigenvalues the determinant of `1 + a` is the product of one plus the eigenvalues of a real positive-semidefinite matrix. theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.toEuclideanCLM_one_add_posSemidef_inv_eigenvectorBasis","label":"toEuclideanCLM_one_add_posSemidef_inv_eigenvectorBasis","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.toEuclideanCLM_one_add_posSemidef_inv_eigenvectorBasis","description":"private theorem toEuclideanCLM_one_add_posSemidef_inv_eigenvectorBasis (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (i : Feature) : Matrix.toEuclideanCLM (𝕜 := Real) (1 + A)⁻¹ (hA.isHermitian.eigenvectorBasis i) = (1 + hA.isHermitian.eigenvalues i)⁻¹ • hA.isHermitian.eigenvectorBasis i","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-42a09aa319da","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6250,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:233"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"private theorem toEuclideanCLM_one_add_posSemidef_inv_eigenvectorBasis (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (i : Feature) : Matrix.toEuclideanCLM (𝕜 := Real) (1 + A)⁻¹ (hA.isHermitian.eigenvectorBasis i) = (1 + hA.isHermitian.eigenvalues i)⁻¹ • hA.isHermitian.eigenvectorBasis i","missing":[],"search":"toeuclideanclm_one_add_possemidef_inv_eigenvectorbasis banditrlproof.oful.toeuclideanclm_one_add_possemidef_inv_eigenvectorbasis private theorem toeuclideanclm_one_add_possemidef_inv_eigenvectorbasis (a : matrix feature feature real) (ha : a.possemidef) (i : feature) : matrix.toeuclideanclm (𝕜 := real) (1 + a)⁻¹ (ha.ishermitian.eigenvectorbasis i) = (1 + ha.ishermitian.eigenvalues i)⁻¹ • ha.ishermitian.eigenvectorbasis i theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.dotProduct_one_add_posSemidef_inv_mulVec_eq_sum_eigenvalues","label":"dotProduct_one_add_posSemidef_inv_mulVec_eq_sum_eigenvalues","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.dotProduct_one_add_posSemidef_inv_mulVec_eq_sum_eigenvalues","description":"The inverse quadratic form of `1 + A` equals its spectral-coordinate sum.","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-d2bae4c4d9a6","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6251,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:285"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem dotProduct_one_add_posSemidef_inv_mulVec_eq_sum_eigenvalues (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (score : EuclideanSpace Real Feature) : score ⬝ᵥ (1 + A)⁻¹.mulVec score = Finset.univ.sum (fun i => ⟪score, hA.isHermitian.eigenvectorBasis i⟫_ℝ ^ 2 / (1 + hA.isHermitian.eigenvalues i))","missing":[],"search":"dotproduct_one_add_possemidef_inv_mulvec_eq_sum_eigenvalues banditrlproof.oful.dotproduct_one_add_possemidef_inv_mulvec_eq_sum_eigenvalues the inverse quadratic form of `1 + a` equals its spectral-coordinate sum. theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_stdGaussian_det","label":"integral_exp_inner_sub_quadratic_stdGaussian_det","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_stdGaussian_det","description":"Exact standard-Gaussian quadratic-exponential identity for an arbitrary real positive-semidefinite matrix. This closes the orthonormal spectral transport from the diagonal-coordinate Gaussian mixture leaf. Transport from a nonstandard initial covariance and the stochastic Tonelli/Markov assembly remain separate obligations.","url":"../modules/banditrlproof-ofulgaussianspectralmixture/index.html#decl-dc755411fb75","parent":"module:BanditRLProof.OFULGaussianSpectralMixture","order":6252,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGaussianSpectralMixture"],["Source","BanditRLProof/OFULGaussianSpectralMixture.lean:351"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_exp_inner_sub_quadratic_stdGaussian_det (A : Matrix Feature Feature Real) (hA : A.PosSemidef) (score : EuclideanSpace Real Feature) : integral (ProbabilityTheory.stdGaussian (EuclideanSpace Real Feature)) (fun z : EuclideanSpace Real Feature => Real.exp (⟪score, z⟫_ℝ - ⟪z, Matrix.toEuclideanCLM (𝕜 := Real) A z⟫_ℝ / 2)) = (Real.sqrt (Matrix.det (1 + A)))⁻¹ * Real.exp (score ⬝ᵥ (1 + A)⁻¹.mulVec score / 2)","missing":[],"search":"integral_exp_inner_sub_quadratic_stdgaussian_det banditrlproof.oful.integral_exp_inner_sub_quadratic_stdgaussian_det exact standard-gaussian quadratic-exponential identity for an arbitrary real positive-semidefinite matrix. this closes the orthonormal spectral transport from the diagonal-coordinate gaussian mixture leaf. transport from a nonstandard initial covariance and the stochastic tonelli/markov assembly remain separate obligations. theorem compiled","shard":"modules/5d04c0c7389d56b5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.linearValue_sub_selected_le_two_mul_bonus_of_score_max","label":"linearValue_sub_selected_le_two_mul_bonus_of_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.linearValue_sub_selected_le_two_mul_bonus_of_score_max","description":"Any action whose optimistic score dominates a comparator satisfies the usual two-bonus OFUL gap bound on the confidence ellipsoid.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-ae3b12ab0a88","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6253,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:25"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem linearValue_sub_selected_le_two_mul_bonus_of_score_max {Feature Action : Type*} [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (hV : V.PosDef) (thetaHat thetaStar : Feature -> Real) (beta : Real) (actionFeature : Action -> Feature -> Real) (selected comparator : Action) (hconfidence : matrixNorm V (thetaHat - thetaStar) <= beta) (hscoreMax : optimisticScore thetaHat V beta (actionFeature comparator) <= optimisticScore thetaHat V beta (actionFeature selected)) : linearValue thetaStar (actionFeature comparator) - linearValue thetaStar (actionFeature selected) <= 2 * beta * confidenceWidth V (actionFeature selected)","missing":[],"search":"linearvalue_sub_selected_le_two_mul_bonus_of_score_max banditrlproof.oful.linearvalue_sub_selected_le_two_mul_bonus_of_score_max any action whose optimistic score dominates a comparator satisfies the usual two-bonus oful gap bound on the confidence ellipsoid. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction_gap_le","label":"finiteHistoryScalarRidgeOptimisticAction_gap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction_gap_le","description":"The concrete strict-fold history selector has the OFUL one-step gap certificate. No identification with the nonconstructive finite argmax is used.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-62d1d06aaaab","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6254,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:56"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScalarRidgeOptimisticAction_gap_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hconfidence : matrixNorm (finiteHistoryScalarRidgeDesign lambda actionFeature n history) (finiteHistoryScalarRidgeEstimate lambda actionFeature n history - thetaStar) <= finiteHistoryScalarRidgeRadius actionFeature R delta lambda S n history) (comparator : Fin K) : linearValue thetaStar (actionFeature comparator) - linearValue thetaStar (actionFeature (finiteHistoryScalarRidgeOptimisticAction hK lambda actionFeature R delta S n history)) <= 2 * finiteHistoryScalarRidgeRadius actionFeature R delta lambda S n history * confidenceWidth (finiteHistoryScalarRidgeDe…","missing":[],"search":"finitehistoryscalarridgeoptimisticaction_gap_le banditrlproof.oful.finitehistoryscalarridgeoptimisticaction_gap_le the concrete strict-fold history selector has the oful one-step gap certificate. no identification with the nonconstructive finite argmax is used. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature","label":"canonicalHistoryTrajectoryFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature","description":"Feature process obtained by projecting actions from a canonical trajectory.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-6d809d283e85","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6255,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:113"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectoryFeature {K : Nat} {Feature : Type u} (actionFeature : Fin K -> Feature -> Real) (t : Nat) (trajectory : (k : Nat) -> Fin K × Real) : Feature -> Real","missing":[],"search":"canonicalhistorytrajectoryfeature banditrlproof.oful.canonicalhistorytrajectoryfeature feature process obtained by projecting actions from a canonical trajectory. definition compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryResponse","label":"canonicalHistoryTrajectoryResponse","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryResponse","description":"Response process obtained by projecting rewards from a canonical trajectory.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-94fe1a9846b2","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6256,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:121"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectoryResponse {K : Nat} (t : Nat) (trajectory : (k : Nat) -> Fin K × Real) : Real","missing":[],"search":"canonicalhistorytrajectoryresponse banditrlproof.oful.canonicalhistorytrajectoryresponse response process obtained by projecting rewards from a canonical trajectory. definition compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryObservedFeature_finitePairHistoryOfTrace_of_le","label":"finiteHistoryObservedFeature_finitePairHistoryOfTrace_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryObservedFeature_finitePairHistoryOfTrace_of_le","description":"A trace prefix exposes the original action feature at every stored index.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-6bf7e331e617","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6257,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:127"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"@[simp] theorem finiteHistoryObservedFeature_finitePairHistoryOfTrace_of_le {K : Nat} {Feature : Type u} (actionFeature : Fin K -> Feature -> Real) (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n t : Nat) (ht : t <= n) : finiteHistoryObservedFeature actionFeature n (History.finitePairHistoryOfTrace action reward n) t = actionFeature (action t)","missing":[],"search":"finitehistoryobservedfeature_finitepairhistoryoftrace_of_le banditrlproof.oful.finitehistoryobservedfeature_finitepairhistoryoftrace_of_le a trace prefix exposes the original action feature at every stored index. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryObservedResponse_finitePairHistoryOfTrace_of_le","label":"finiteHistoryObservedResponse_finitePairHistoryOfTrace_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryObservedResponse_finitePairHistoryOfTrace_of_le","description":"A trace prefix exposes the original response at every stored index.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-2c96baa3f65f","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6258,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:138"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"@[simp] theorem finiteHistoryObservedResponse_finitePairHistoryOfTrace_of_le {K : Nat} (action : ActionTrace (Fin K)) (reward : RewardTrace Real) (n t : Nat) (ht : t <= n) : finiteHistoryObservedResponse n (History.finitePairHistoryOfTrace action reward n) t = reward t","missing":[],"search":"finitehistoryobservedresponse_finitepairhistoryoftrace_of_le banditrlproof.oful.finitehistoryobservedresponse_finitepairhistoryoftrace_of_le a trace prefix exposes the original response at every stored index. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonFeatureGram_finitePairHistoryOfTrace_eq","label":"finiteHorizonFeatureGram_finitePairHistoryOfTrace_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonFeatureGram_finitePairHistoryOfTrace_eq","description":"The feature Gram reconstructed from inclusive history `n` is exactly the trajectory feature Gram at horizon `n + 1`.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-3f0e4897f6d5","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6259,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:151"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonFeatureGram_finitePairHistoryOfTrace_eq {K : Nat} {Feature : Type u} [Fintype Feature] (actionFeature : Fin K -> Feature -> Real) (trajectory : (k : Nat) -> Fin K × Real) (n : Nat) : finiteHorizonFeatureGram (fun t historyValue => finiteHistoryObservedFeature actionFeature n historyValue t) (n + 1) (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n) = finiteHorizonFeatureGram (canonicalHistoryTrajectoryFeature actionFeature) (n + 1) trajectory","missing":[],"search":"finitehorizonfeaturegram_finitepairhistoryoftrace_eq banditrlproof.oful.finitehorizonfeaturegram_finitepairhistoryoftrace_eq the feature gram reconstructed from inclusive history `n` is exactly the trajectory feature gram at horizon `n + 1`. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonResponseVector_finitePairHistoryOfTrace_eq","label":"finiteHorizonResponseVector_finitePairHistoryOfTrace_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonResponseVector_finitePairHistoryOfTrace_eq","description":"The response vector reconstructed from inclusive history `n` is exactly the trajectory response vector at horizon `n + 1`.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-f094b4e6b390","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6260,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:178"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonResponseVector_finitePairHistoryOfTrace_eq {K : Nat} {Feature : Type u} [Fintype Feature] (actionFeature : Fin K -> Feature -> Real) (trajectory : (k : Nat) -> Fin K × Real) (n : Nat) : finiteHorizonResponseVector (fun t historyValue => finiteHistoryObservedFeature actionFeature n historyValue t) (fun t historyValue => finiteHistoryObservedResponse n historyValue t) (n + 1) (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n) = finiteHorizonResponseVector (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse (n + 1) trajectory","missing":[],"search":"finitehorizonresponsevector_finitepairhistoryoftrace_eq banditrlproof.oful.finitehorizonresponsevector_finitepairhistoryoftrace_eq the response vector reconstructed from inclusive history `n` is exactly the trajectory response vector at horizon `n + 1`. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeDesign_finitePairHistoryOfTrace_eq","label":"finiteHistoryScalarRidgeDesign_finitePairHistoryOfTrace_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeDesign_finitePairHistoryOfTrace_eq","description":"Exact design-matrix alignment between finite history and trajectory state.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-1b9e69e4c853","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6261,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:206"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScalarRidgeDesign_finitePairHistoryOfTrace_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (trajectory : (k : Nat) -> Fin K × Real) (n : Nat) : finiteHistoryScalarRidgeDesign lambda actionFeature n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n) = finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) (n + 1) trajectory","missing":[],"search":"finitehistoryscalarridgedesign_finitepairhistoryoftrace_eq banditrlproof.oful.finitehistoryscalarridgedesign_finitepairhistoryoftrace_eq exact design-matrix alignment between finite history and trajectory state. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeEstimate_finitePairHistoryOfTrace_eq","label":"finiteHistoryScalarRidgeEstimate_finitePairHistoryOfTrace_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeEstimate_finitePairHistoryOfTrace_eq","description":"Exact ridge-estimate alignment between finite history and trajectory state.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-c4f1633a4b11","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6262,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:223"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScalarRidgeEstimate_finitePairHistoryOfTrace_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (trajectory : (k : Nat) -> Fin K × Real) (n : Nat) : finiteHistoryScalarRidgeEstimate lambda actionFeature n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n) = finiteHorizonRidgeEstimate (Matrix.scalar Feature lambda) (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse (n + 1) trajectory","missing":[],"search":"finitehistoryscalarridgeestimate_finitepairhistoryoftrace_eq banditrlproof.oful.finitehistoryscalarridgeestimate_finitepairhistoryoftrace_eq exact ridge-estimate alignment between finite history and trajectory state. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeRadius_finitePairHistoryOfTrace_eq","label":"finiteHistoryScalarRidgeRadius_finitePairHistoryOfTrace_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidgeRadius_finitePairHistoryOfTrace_eq","description":"Exact confidence-radius alignment between finite history and trajectory.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-5b0ee7447bb0","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6263,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:243"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScalarRidgeRadius_finitePairHistoryOfTrace_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (trajectory : (k : Nat) -> Fin K × Real) (n : Nat) : finiteHistoryScalarRidgeRadius actionFeature R delta lambda S n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n) = finiteHorizonScalarConfidenceRadius (canonicalHistoryTrajectoryFeature actionFeature) R delta lambda S (n + 1) trajectory","missing":[],"search":"finitehistoryscalarridgeradius_finitepairhistoryoftrace_eq banditrlproof.oful.finitehistoryscalarridgeradius_finitepairhistoryoftrace_eq exact confidence-radius alignment between finite history and trajectory. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","label":"canonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","description":"Outside the generic fixed-time confidence failure set, the actual canonical successor action satisfies the concrete OFUL one-step gap certificate.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-5936a0853606","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6264,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:268"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (comparator : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∉ scalarRidgeConfidenceFailureAt lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta (n + 1) -> linearValue thetaStar (actionFeature comparator) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory (n + 1))) <= 2 * finiteHorizonSca…","missing":[],"search":"canonicalhistorytrajectory_action_succ_gap_le_of_not_mem_confidencefailure banditrlproof.oful.canonicalhistorytrajectory_action_succ_gap_le_of_not_mem_confidencefailure outside the generic fixed-time confidence failure set, the actual canonical successor action satisfies the concrete oful one-step gap certificate. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_gap_le_on_uniformConfidence","label":"canonicalHistoryTrajectory_sum_range_succ_gap_le_on_uniformConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_gap_le_on_uniformConfidence","description":"On the equal-share uniform confidence event, all canonical successor rounds `1, ..., horizon` satisfy the cumulative OFUL bonus bound. The fixed initial round `0` is not part of this sum.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryconfidencegap/index.html#decl-91a8b2494cd8","parent":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","order":6265,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryConfidenceGap"],["Source","BanditRLProof/OFULGeneratedTrajectoryConfidenceGap.lean:371"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_sum_range_succ_gap_le_on_uniformConfidence {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (comparator : Nat -> Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment, trajectory ∉ finiteHorizonUniformScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta horizon -> (Finset.range horizon).sum (fun n => linearValue thetaStar (actionFeature (comparator (n + 1))) - linearValue the…","missing":[],"search":"canonicalhistorytrajectory_sum_range_succ_gap_le_on_uniformconfidence banditrlproof.oful.canonicalhistorytrajectory_sum_range_succ_gap_le_on_uniformconfidence on the equal-share uniform confidence event, all canonical successor rounds `1, ..., horizon` satisfy the cumulative oful bonus bound. the fixed initial round `0` is not part of this sum. theorem compiled","shard":"modules/69bb82127102175c.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration","label":"canonicalHistoryTrajectoryBeforeFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration","description":"The canonical trajectory filtration strictly before the current coordinate. Level zero is trivial, while level `n + 1` contains exactly coordinates `0, ..., n`.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-d4129c6ceb45","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6266,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:28"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectoryBeforeFiltration {K : Nat} : Filtration Nat (inferInstance : MeasurableSpace (Nat -> Fin K × Real)) where","missing":[],"search":"canonicalhistorytrajectorybeforefiltration banditrlproof.oful.canonicalhistorytrajectorybeforefiltration the canonical trajectory filtration strictly before the current coordinate. level zero is trivial, while level `n + 1` contains exactly coordinates `0, ..., n`. definition compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration_zero","label":"canonicalHistoryTrajectoryBeforeFiltration_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration_zero","description":"theorem canonicalHistoryTrajectoryBeforeFiltration_zero {K : Nat} : (canonicalHistoryTrajectoryBeforeFiltration (K := K) 0 : MeasurableSpace (Nat -> Fin K × Real)) = ⊥","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-b5eb660ff313","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6267,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:56"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryBeforeFiltration_zero {K : Nat} : (canonicalHistoryTrajectoryBeforeFiltration (K := K) 0 : MeasurableSpace (Nat -> Fin K × Real)) = ⊥","missing":[],"search":"canonicalhistorytrajectorybeforefiltration_zero banditrlproof.oful.canonicalhistorytrajectorybeforefiltration_zero theorem canonicalhistorytrajectorybeforefiltration_zero {k : nat} : (canonicalhistorytrajectorybeforefiltration (k := k) 0 : measurablespace (nat -> fin k × real)) = ⊥ theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration_succ","label":"canonicalHistoryTrajectoryBeforeFiltration_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration_succ","description":"theorem canonicalHistoryTrajectoryBeforeFiltration_succ {K : Nat} (n : Nat) : (canonicalHistoryTrajectoryBeforeFiltration (K := K) (n + 1) : MeasurableSpace (Nat -> Fin K × Real)) = Filtration.piLE (X := fun _ : Nat => Fin K × Real) n","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-cb46b40cdcfa","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6268,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:62"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryBeforeFiltration_succ {K : Nat} (n : Nat) : (canonicalHistoryTrajectoryBeforeFiltration (K := K) (n + 1) : MeasurableSpace (Nat -> Fin K × Real)) = Filtration.piLE (X := fun _ : Nat => Fin K × Real) n","missing":[],"search":"canonicalhistorytrajectorybeforefiltration_succ banditrlproof.oful.canonicalhistorytrajectorybeforefiltration_succ theorem canonicalhistorytrajectorybeforefiltration_succ {k : nat} (n : nat) : (canonicalhistorytrajectorybeforefiltration (k := k) (n + 1) : measurablespace (nat -> fin k × real)) = filtration.pile (x := fun _ : nat => fin k × real) n theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScalarRidgeOptimisticAction","label":"measurable_finiteHistoryScalarRidgeOptimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryScalarRidgeOptimisticAction","description":"The concrete scalar-ridge strict-fold selector is measurable in its history.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-5f79d3f4170a","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6269,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:69"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) : Measurable (finiteHistoryScalarRidgeOptimisticAction hK lambda actionFeature R delta S n)","missing":[],"search":"measurable_finitehistoryscalarridgeoptimisticaction banditrlproof.oful.measurable_finitehistoryscalarridgeoptimisticaction the concrete scalar-ridge strict-fold selector is measurable in its history. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableFeature","label":"canonicalHistoryTrajectoryPredictableFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableFeature","description":"Pointwise predictable feature used by scalar-ridge confidence. At time zero it uses the deterministic initial arm. At time `n + 1` it uses the strict-fold selector evaluated only on coordinates through `n`.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-71133e1e9f09","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6270,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:98"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryTrajectoryPredictableFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) : Nat -> (Nat -> Fin K × Real) -> Feature -> Real | 0, _trajectory => actionFeature ⟨0, hK⟩ | n + 1, trajectory => finiteHistoryScalarRidgeSelectedFeature hK lambda actionFeature R delta S n (Preorder.frestrictLe n trajectory) /-- Every coordinate of the predictable feature is strict-past measurable. -/ theorem canonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (t : Nat) (j : Feature) : StronglyMeasurable[ canonicalHistoryTrajectoryBeforeFiltration (K := K) t] (fun trajectory => canonicalHistoryTra…","missing":[],"search":"canonicalhistorytrajectorypredictablefeature banditrlproof.oful.canonicalhistorytrajectorypredictablefeature pointwise predictable feature used by scalar-ridge confidence. at time zero it uses the deterministic initial arm. at time `n + 1` it uses the strict-fold selector evaluated only on coordinates through `n`. definition compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","label":"canonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","description":"Every coordinate of the predictable feature is strict-past measurable.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-2c21120eee3b","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6271,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:114"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (t : Nat) (j : Feature) : StronglyMeasurable[ canonicalHistoryTrajectoryBeforeFiltration (K := K) t] (fun trajectory => canonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R delta S t trajectory j)","missing":[],"search":"canonicalhistorytrajectorypredictablefeature_stronglymeasurable banditrlproof.oful.canonicalhistorytrajectorypredictablefeature_stronglymeasurable every coordinate of the predictable feature is strict-past measurable. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_zero_ae_eq_initialArm","label":"canonicalHistoryTrajectory_action_zero_ae_eq_initialArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_zero_ae_eq_initialArm","description":"The initial canonical action of the concrete OFUL algorithm is its Dirac arm a.e.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-8c97a34f2c7f","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6272,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:159"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_zero_ae_eq_initialArm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, Thompson.canonicalHistoryTrajectoryAction trajectory 0 = ⟨0, hK⟩","missing":[],"search":"canonicalhistorytrajectory_action_zero_ae_eq_initialarm banditrlproof.oful.canonicalhistorytrajectory_action_zero_ae_eq_initialarm the initial canonical action of the concrete oful algorithm is its dirac arm a.e. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature","label":"canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature","description":"Actual and pointwise predictable canonical features agree a.e. at each time.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-39df76fc9311","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6273,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:213"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (t : Nat) : canonicalHistoryTrajectoryFeature actionFeature t =ᵐ[ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment] canonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R delta S t","missing":[],"search":"canonicalhistorytrajectoryfeature_ae_eq_predictablefeature banditrlproof.oful.canonicalhistorytrajectoryfeature_ae_eq_predictablefeature actual and pointwise predictable canonical features agree a.e. at each time. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature_all","label":"canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature_all","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature_all","description":"Actual and predictable canonical features agree simultaneously at all times.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-b6b16a735275","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6274,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:249"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature_all {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, ∀ t, canonicalHistoryTrajectoryFeature actionFeature t trajectory = canonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R delta S t trajectory","missing":[],"search":"canonicalhistorytrajectoryfeature_ae_eq_predictablefeature_all banditrlproof.oful.canonicalhistorytrajectoryfeature_ae_eq_predictablefeature_all actual and predictable canonical features agree simultaneously at all times. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableResidual","label":"canonicalHistoryTrajectoryPredictableResidual","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableResidual","description":"Canonical reward residual around the pointwise predictable linear response.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-0ec8d61ec73d","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6275,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:273"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryTrajectoryPredictableResidual {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (i : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"canonicalhistorytrajectorypredictableresidual banditrlproof.oful.canonicalhistorytrajectorypredictableresidual canonical reward residual around the pointwise predictable linear response. definition compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_canonicalHistoryTrajectoryResponse_before_succ","label":"measurable_canonicalHistoryTrajectoryResponse_before_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_canonicalHistoryTrajectoryResponse_before_succ","description":"The canonical reward coordinate at time `i` is measurable at level `i + 1`.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-80c790a39c26","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6276,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:288"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalHistoryTrajectoryResponse_before_succ {K : Nat} (i : Nat) : @Measurable (Nat -> Fin K × Real) Real (canonicalHistoryTrajectoryBeforeFiltration (K := K) (i + 1)) inferInstance (canonicalHistoryTrajectoryResponse i)","missing":[],"search":"measurable_canonicalhistorytrajectoryresponse_before_succ banditrlproof.oful.measurable_canonicalhistorytrajectoryresponse_before_succ the canonical reward coordinate at time `i` is measurable at level `i + 1`. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","label":"canonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","description":"The zero-initialized predictable residual process is strongly adapted.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-b2450313f244","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6277,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:313"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryPredictableResidual_stronglyAdapted {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) : StronglyAdapted (canonicalHistoryTrajectoryBeforeFiltration (K := K)) (fun t trajectory => match t with | 0 => 0 | i + 1 => canonicalHistoryTrajectoryPredictableResidual hK lambda thetaStar actionFeature R delta S i trajectory)","missing":[],"search":"canonicalhistorytrajectorypredictableresidual_stronglyadapted banditrlproof.oful.canonicalhistorytrajectorypredictableresidual_stronglyadapted the zero-initialized predictable residual process is strongly adapted. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteActionProjectionBound","label":"finiteActionProjectionBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteActionProjectionBound","description":"Maximum absolute projection over the finite arm set.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-78b161e6d232","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6278,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:371"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteActionProjectionBound {K : Nat} {Feature : Type u} [Fintype Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (theta : EuclideanSpace Real Feature) : Real","missing":[],"search":"finiteactionprojectionbound banditrlproof.oful.finiteactionprojectionbound maximum absolute projection over the finite arm set. definition compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteActionProjectionBound_nonneg","label":"finiteActionProjectionBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteActionProjectionBound_nonneg","description":"theorem finiteActionProjectionBound_nonneg {K : Nat} {Feature : Type u} [Fintype Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (theta : EuclideanSpace Real Feature) : 0 <= finiteActionProjectionBound hK actionFeature theta","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-dfaa76e196cf","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6279,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:383"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteActionProjectionBound_nonneg {K : Nat} {Feature : Type u} [Fintype Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (theta : EuclideanSpace Real Feature) : 0 <= finiteActionProjectionBound hK actionFeature theta","missing":[],"search":"finiteactionprojectionbound_nonneg banditrlproof.oful.finiteactionprojectionbound_nonneg theorem finiteactionprojectionbound_nonneg {k : nat} {feature : type u} [fintype feature] (hk : 0 < k) (actionfeature : fin k -> feature -> real) (theta : euclideanspace real feature) : 0 <= finiteactionprojectionbound hk actionfeature theta theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.predictableFeature_projection_le_finiteActionProjectionBound","label":"predictableFeature_projection_le_finiteActionProjectionBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.predictableFeature_projection_le_finiteActionProjectionBound","description":"theorem predictableFeature_projection_le_finiteActionProjectionBound {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (theta : EuclideanSpace Real Feature) (i : Nat) (trajectory : Nat -> Fin K × Real) : |dotProduct (WithLp.ofLp theta) (canonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R d…","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-7c49f7b6cae2","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6280,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:403"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem predictableFeature_projection_le_finiteActionProjectionBound {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (theta : EuclideanSpace Real Feature) (i : Nat) (trajectory : Nat -> Fin K × Real) : |dotProduct (WithLp.ofLp theta) (canonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R delta S i trajectory)| <= finiteActionProjectionBound hK actionFeature theta","missing":[],"search":"predictablefeature_projection_le_finiteactionprojectionbound banditrlproof.oful.predictablefeature_projection_le_finiteactionprojectionbound theorem predictablefeature_projection_le_finiteactionprojectionbound {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (actionfeature : fin k -> feature -> real) (r delta s : real) (theta : euclideanspace real feature) (i : nat) (trajectory : nat -> fin k × real) : |dotproduct (withlp.oflp theta) (canonicalhistorytrajectorypredictablefeature hk lambda actionfeature r delta s i trajectory)| <= finiteactionprojectionbound hk actionfeature theta theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.CanonicalPredictableScalarRidgeResidualLaw","label":"CanonicalPredictableScalarRidgeResidualLaw","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.CanonicalPredictableScalarRidgeResidualLaw","description":"The remaining stochastic law for predictable canonical confidence. Measurability, adaptedness, deterministic finite-arm projection bounds, and the pointwise response identity are derived by this module. A concrete environment producer therefore only needs the parameter norm and the strict-past conditional MGF of the canonical predictable residual.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-53ce721f3674","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6281,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:447"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure CanonicalPredictableScalarRidgeResidualLaw {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) : Prop where","missing":[],"search":"canonicalpredictablescalarridgeresiduallaw banditrlproof.oful.canonicalpredictablescalarridgeresiduallaw the remaining stochastic law for predictable canonical confidence. measurability, adaptedness, deterministic finite-arm projection bounds, and the pointwise response identity are derived by this module. a concrete environment producer therefore only needs the parameter norm and the strict-past conditional mgf of the canonical predictable residual. structure compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.mem_scalarRidgeConfidenceFailureAt_iff_of_feature_eq","label":"mem_scalarRidgeConfidenceFailureAt_iff_of_feature_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.mem_scalarRidgeConfidenceFailureAt_iff_of_feature_eq","description":"Pointwise feature equality preserves every fixed-time scalar-ridge event.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-c9973fe58805","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6282,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:471"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem mem_scalarRidgeConfidenceFailureAt_iff_of_feature_eq {Omega : Type*} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature feature' : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R delta : Real) (n : Nat) (omega : Omega) (hfeature : forall i, feature i omega = feature' i omega) : omega ∈ scalarRidgeConfidenceFailureAt lambda thetaStar S feature response R delta n ↔ omega ∈ scalarRidgeConfidenceFailureAt lambda thetaStar S feature' response R delta n","missing":[],"search":"mem_scalarridgeconfidencefailureat_iff_of_feature_eq banditrlproof.oful.mem_scalarridgeconfidencefailureat_iff_of_feature_eq pointwise feature equality preserves every fixed-time scalar-ridge event. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.mem_finiteHorizonUniformScalarRidgeConfidenceFailureSet_iff_of_feature_eq","label":"mem_finiteHorizonUniformScalarRidgeConfidenceFailureSet_iff_of_feature_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.mem_finiteHorizonUniformScalarRidgeConfidenceFailureSet_iff_of_feature_eq","description":"Pointwise feature equality preserves the finite-window uniform failure event.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-1a7202f17a20","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6283,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:510"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem mem_finiteHorizonUniformScalarRidgeConfidenceFailureSet_iff_of_feature_eq {Omega : Type*} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature feature' : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R delta : Real) (horizon : Nat) (omega : Omega) (hfeature : forall i, feature i omega = feature' i omega) : omega ∈ finiteHorizonUniformScalarRidgeConfidenceFailureSet lambda thetaStar S feature response R delta horizon ↔ omega ∈ finiteHorizonUniformScalarRidgeConfidenceFailureSet lambda thetaStar S feature' response R delta horizon","missing":[],"search":"mem_finitehorizonuniformscalarridgeconfidencefailureset_iff_of_feature_eq banditrlproof.oful.mem_finitehorizonuniformscalarridgeconfidencefailureset_iff_of_feature_eq pointwise feature equality preserves the finite-window uniform failure event. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le_of_predictableResidualLaw","label":"measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le_of_predictableResidualLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le_of_predictableResidualLaw","description":"Uniform scalar-ridge confidence on the actual canonical selected features, derived from only the predictable residual conditional law.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-bbf97b4c0849","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6284,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:549"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le_of_predictableResidualLaw {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (source : CanonicalPredictableScalarRidgeResidualLaw hK lambda thetaStar actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S environment horizon) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment (finiteHorizonUniformScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHi…","missing":[],"search":"measure_canonicalhistorytrajectory_uniformscalarridgeconfidencefailureset_le_of_predictableresiduallaw banditrlproof.oful.measure_canonicalhistorytrajectory_uniformscalarridgeconfidencefailureset_le_of_predictableresiduallaw uniform scalar-ridge confidence on the actual canonical selected features, derived from only the predictable residual conditional law. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_predictableResidualLaw","label":"measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_predictableResidualLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_predictableResidualLaw","description":"Canonical successor-gap tail derived from the strict-past predictable residual law. Unlike the earlier source theorem, no pointwise measurability of the actual selected feature is assumed.","url":"../modules/banditrlproof-ofulgeneratedtrajectorypredictableconfidence/index.html#decl-1001526943ca","parent":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","order":6285,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryPredictableConfidence.lean:661"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_predictableResidualLaw {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (comparator : Nat -> Fin K) (source : CanonicalPredictableScalarRidgeResidualLaw hK lambda thetaStar actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S environment horizon) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment (canonicalHistoryTrajectorySumRangeSuccGapViolationSet lambda thetaS…","missing":[],"search":"measure_canonicalhistorytrajectorysumrangesuccgapviolationset_le_of_predictableresiduallaw banditrlproof.oful.measure_canonicalhistorytrajectorysumrangesuccgapviolationset_le_of_predictableresiduallaw canonical successor-gap tail derived from the strict-past predictable residual law. unlike the earlier source theorem, no pointwise measurability of the actual selected feature is assumed. theorem compiled","shard":"modules/8f536a33f6c96662.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget","label":"standardScalarLogDetBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarLogDetBudget","description":"The standard log-determinant budget for a bounded scalar-ridge prefix.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-352859af05e7","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6286,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardScalarLogDetBudget {Feature : Type u} [Fintype Feature] (lambda : Real) (T : Nat) (L2 : Real) : Real","missing":[],"search":"standardscalarlogdetbudget banditrlproof.oful.standardscalarlogdetbudget the standard log-determinant budget for a bounded scalar-ridge prefix. definition compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper","label":"standardScalarConfidenceRadiusUpper","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper","description":"A deterministic upper radius obtained by replacing the random determinant ratio with `exp (standardScalarLogDetBudget ...)`.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-e80b5dbb2428","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6287,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:34"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardScalarConfidenceRadiusUpper {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (T : Nat) (L2 : Real) : Real","missing":[],"search":"standardscalarconfidenceradiusupper banditrlproof.oful.standardscalarconfidenceradiusupper a deterministic upper radius obtained by replacing the random determinant ratio with `exp (standardscalarlogdetbudget ...)`. definition compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardSelectedWidthBudget","label":"standardSelectedWidthBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardSelectedWidthBudget","description":"The selected-width budget paired with the standard log-determinant term.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-e82df5ad7f9d","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6288,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:48"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardSelectedWidthBudget {Feature : Type u} [Fintype Feature] (lambda : Real) (T : Nat) (L2 : Real) : Real","missing":[],"search":"standardselectedwidthbudget banditrlproof.oful.standardselectedwidthbudget the selected-width budget paired with the standard log-determinant term. definition compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarRadiusWidthBound","label":"standardScalarRadiusWidthBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarRadiusWidthBound","description":"Deterministic radius-times-width budget used by the canonical gap theorem.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-6da74a7193ba","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6289,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:58"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardScalarRadiusWidthBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (T : Nat) (L2 : Real) : Real","missing":[],"search":"standardscalarradiuswidthbound banditrlproof.oful.standardscalarradiuswidthbound deterministic radius-times-width budget used by the canonical gap theorem. definition compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarGram_eq_regularizedPrefixFeatureGram","label":"finiteHorizonScalarGram_eq_regularizedPrefixFeatureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarGram_eq_regularizedPrefixFeatureGram","description":"The process scalar Gram is definitionally the regularized feature prefix.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-000597fb5258","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6290,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:68"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonScalarGram_eq_regularizedPrefixFeatureGram {Omega Feature : Type*} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (feature : Nat -> Omega -> Feature -> Real) (n : Nat) (omega : Omega) : finiteHorizonScalarGram lambda feature n omega = regularizedPrefixFeatureGram lambda (fun t => feature t omega) n","missing":[],"search":"finitehorizonscalargram_eq_regularizedprefixfeaturegram banditrlproof.oful.finitehorizonscalargram_eq_regularizedprefixfeaturegram the process scalar gram is definitionally the regularized feature prefix. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_mono","label":"standardScalarLogDetBudget_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarLogDetBudget_mono","description":"The standard log-determinant budget is monotone in the horizon.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-914b2a06cbe0","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6291,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:79"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarLogDetBudget_mono {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (L2 : Real) (hL2 : 0 <= L2) {n T : Nat} (hnT : n <= T) : standardScalarLogDetBudget (Feature := Feature) lambda n L2 <= standardScalarLogDetBudget (Feature := Feature) lambda T L2","missing":[],"search":"standardscalarlogdetbudget_mono banditrlproof.oful.standardscalarlogdetbudget_mono the standard log-determinant budget is monotone in the horizon. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_nonneg","label":"standardScalarLogDetBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarLogDetBudget_nonneg","description":"The standard log-determinant budget is nonnegative.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-464757f00985","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6292,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:118"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarLogDetBudget_nonneg {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) : 0 <= standardScalarLogDetBudget (Feature := Feature) lambda T L2","missing":[],"search":"standardscalarlogdetbudget_nonneg banditrlproof.oful.standardscalarlogdetbudget_nonneg the standard log-determinant budget is nonnegative. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_le_standardUpper","label":"finiteHorizonScalarConfidenceRadius_le_standardUpper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_le_standardUpper","description":"Every prefix scalar confidence radius is bounded by the standard deterministic radius at a later horizon.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-c54d588fe0fa","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6293,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:142"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonScalarConfidenceRadius_le_standardUpper {Omega Feature : Type*} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Omega -> Feature -> Real) (R delta S : Real) (omega : Omega) (n T : Nat) (hnT : n <= T) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (feature t omega) (feature t omega) <= L2) (hdelta : 0 < delta) : finiteHorizonScalarConfidenceRadius feature R delta lambda S n omega <= standardScalarConfidenceRadiusUpper (Feature := Feature) R delta lambda S T L2","missing":[],"search":"finitehorizonscalarconfidenceradius_le_standardupper banditrlproof.oful.finitehorizonscalarconfidenceradius_le_standardupper every prefix scalar confidence radius is bounded by the standard deterministic radius at a later horizon. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_nonneg","label":"standardScalarConfidenceRadiusUpper_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_nonneg","description":"The standard deterministic confidence-radius upper bound is nonnegative.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-0f4f9112d02c","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6294,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:251"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarConfidenceRadiusUpper_nonneg {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (T : Nat) (L2 : Real) (hS : 0 <= S) : 0 <= standardScalarConfidenceRadiusUpper (Feature := Feature) R delta lambda S T L2","missing":[],"search":"standardscalarconfidenceradiusupper_nonneg banditrlproof.oful.standardscalarconfidenceradiusupper_nonneg the standard deterministic confidence-radius upper bound is nonnegative. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardSelectedWidthBudget_nonneg","label":"standardSelectedWidthBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardSelectedWidthBudget_nonneg","description":"The standard selected-width budget is nonnegative.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-474ec1210867","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6295,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:261"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardSelectedWidthBudget_nonneg {Feature : Type u} [Fintype Feature] (lambda : Real) (T : Nat) (L2 : Real) : 0 <= standardSelectedWidthBudget (Feature := Feature) lambda T L2","missing":[],"search":"standardselectedwidthbudget_nonneg banditrlproof.oful.standardselectedwidthbudget_nonneg the standard selected-width budget is nonnegative. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard","label":"canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard","description":"Pointwise canonical successor bonus bound. The full action prefix through `horizon` is charged to the selected-width theorem, while the gap sum only uses successor rounds `1, ..., horizon`.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-0070bfe257da","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6296,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:273"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (trajectory : Nat -> Fin K × Real) (hwidth : forall t, t < horizon + 1 -> confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) t trajectory) (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory t)) <= 1) : (Finset.range horizon).sum (fun n => 2 * finiteHorizonScalarConfidenceRadius (canonicalHistoryTrajectoryFeature actionFeature) R (delta / ((horizon + 1 : Nat) : Real)) lambda S (n + 1…","missing":[],"search":"canonicalhistorytrajectory_sum_range_succ_radius_mul_width_le_standard banditrlproof.oful.canonicalhistorytrajectory_sum_range_succ_radius_mul_width_le_standard pointwise canonical successor bonus bound. the full action prefix through `horizon` is charged to the selected-width theorem, while the gap sum only uses successor rounds `1, ..., horizon`. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet","label":"canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet","description":"Violation of the deterministic standard radius-times-width successor-gap budget. The fixed initial gap is still excluded.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-49873aa8ee8f","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6297,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:432"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (horizon : Nat) (L2 : Real) (comparator : Nat -> Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"canonicalhistorytrajectorysumrangesuccgapstandardviolationset banditrlproof.oful.canonicalhistorytrajectorysumrangesuccgapstandardviolationset violation of the deterministic standard radius-times-width successor-gap budget. the fixed initial gap is still excluded. definition compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_subset","label":"canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_subset","description":"The deterministic-budget violation event is contained in the compiled random radius-times-width violation event.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-6c14bc77878c","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6298,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:458"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_subset {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (hwidth : forall trajectory t, t < horizon + 1 -> confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) t trajectory) (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory t)) <= 1) : canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet lambda thetaStar actionFeature R delta S horizon L2 comparator <= canonicalHistoryTraject…","missing":[],"search":"canonicalhistorytrajectorysumrangesuccgapstandardviolationset_subset banditrlproof.oful.canonicalhistorytrajectorysumrangesuccgapstandardviolationset_subset the deterministic-budget violation event is contained in the compiled random radius-times-width violation event. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment","label":"measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment","description":"Concrete high-probability successor-gap theorem with a deterministic standard radius-times-width budget under the linear-sub-Gaussian environment law.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryradiuswidth/index.html#decl-d68fca064734","parent":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","order":6299,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryRadiusWidth"],["Source","BanditRLProof/OFULGeneratedTrajectoryRadiusWidth.lean:524"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (hwidth : forall trajectory t, t < horizon + 1 -> confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) t trajectory) (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory…","missing":[],"search":"measure_canonicalhistorytrajectorysumrangesuccgapstandardviolationset_le_of_linearsubgaussianenvironment banditrlproof.oful.measure_canonicalhistorytrajectorysumrangesuccgapstandardviolationset_le_of_linearsubgaussianenvironment concrete high-probability successor-gap theorem with a deterministic standard radius-times-width budget under the linear-sub-gaussian environment law. theorem compiled","shard":"modules/e3fcc58bd11cf3fd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.CanonicalScalarRidgeConfidenceSource","label":"CanonicalScalarRidgeConfidenceSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.CanonicalScalarRidgeConfidenceSource","description":"Regularity source needed to apply scalar-ridge uniform confidence to one canonical history-algorithm trajectory. The selected feature is the actual canonical action feature. Constructing this source from a concrete reward environment therefore remains a genuine predictability and conditional-law obligation.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryuniformconfidence/index.html#decl-d3f43ebf2923","parent":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","order":6300,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULGeneratedTrajectoryUniformConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryUniformConfidence.lean:28"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure CanonicalScalarRidgeConfidenceSource {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (horizon : Nat) where","missing":[],"search":"canonicalscalarridgeconfidencesource banditrlproof.oful.canonicalscalarridgeconfidencesource regularity source needed to apply scalar-ridge uniform confidence to one canonical history-algorithm trajectory. the selected feature is the actual canonical action feature. constructing this source from a concrete reward environment therefore remains a genuine predictability and conditional-law obligation. structure compiled","shard":"modules/79c711e35b34d2bf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le","label":"measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le","description":"The generic equal-share scalar-ridge confidence theorem specialized to one canonical history-algorithm trajectory.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryuniformconfidence/index.html#decl-009709e43cca","parent":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","order":6301,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryUniformConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryUniformConfidence.lean:72"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (horizon : Nat) (source : CanonicalScalarRidgeConfidenceSource algorithm environment thetaStar actionFeature R S horizon) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) : Thompson.canonicalHistoryTrajectoryMeasure algorithm environment (finiteHorizonUniformScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta horizon) <= ENNReal.ofReal delta","missing":[],"search":"measure_canonicalhistorytrajectory_uniformscalarridgeconfidencefailureset_le banditrlproof.oful.measure_canonicalhistorytrajectory_uniformscalarridgeconfidencefailureset_le the generic equal-share scalar-ridge confidence theorem specialized to one canonical history-algorithm trajectory. theorem compiled","shard":"modules/79c711e35b34d2bf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapViolationSet","label":"canonicalHistoryTrajectorySumRangeSuccGapViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapViolationSet","description":"Strict violation of the compiled successor-window OFUL good-event gap bound. The range index `n` is charged to the actual action at time `n + 1`, so time zero is not part of this event.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryuniformconfidence/index.html#decl-36aa09392900","parent":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","order":6302,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULGeneratedTrajectoryUniformConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryUniformConfidence.lean:116"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectorySumRangeSuccGapViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (horizon : Nat) (comparator : Nat -> Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"canonicalhistorytrajectorysumrangesuccgapviolationset banditrlproof.oful.canonicalhistorytrajectorysumrangesuccgapviolationset strict violation of the compiled successor-window oful good-event gap bound. the range index `n` is charged to the actual action at time `n + 1`, so time zero is not part of this event. definition compiled","shard":"modules/79c711e35b34d2bf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le","label":"measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le","description":"Canonical high-probability successor-gap theorem. Under the exact generated feature/noise/response regularity source, the probability that the cumulative true linear gap over rounds `1, ..., horizon` exceeds the compiled radius-times-width sum is at most the total confidence budget `delta`.","url":"../modules/banditrlproof-ofulgeneratedtrajectoryuniformconfidence/index.html#decl-811cc599cdbc","parent":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","order":6303,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULGeneratedTrajectoryUniformConfidence"],["Source","BanditRLProof/OFULGeneratedTrajectoryUniformConfidence.lean:153"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (comparator : Nat -> Fin K) (source : CanonicalScalarRidgeConfidenceSource (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment thetaStar actionFeature R S horizon) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment (canonicalHistoryTrajectorySumRangeS…","missing":[],"search":"measure_canonicalhistorytrajectorysumrangesuccgapviolationset_le banditrlproof.oful.measure_canonicalhistorytrajectorysumrangesuccgapviolationset_le canonical high-probability successor-gap theorem. under the exact generated feature/noise/response regularity source, the probability that the cumulative true linear gap over rounds `1, ..., horizon` exceeds the compiled radius-times-width sum is at most the total confidence budget `delta`. theorem compiled","shard":"modules/79c711e35b34d2bf.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardHighProbabilityRegretLogBudget","label":"standardHighProbabilityRegretLogBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardHighProbabilityRegretLogBudget","description":"The confidence logarithm for outer failure probability `delta` over the complete finite window `0, ..., horizon`.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html#decl-c4226eb54ff0","parent":"module:BanditRLProof.OFULHighProbabilityRegretRate","order":6304,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULHighProbabilityRegretRate"],["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean:22"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardHighProbabilityRegretLogBudget {Feature : Type u} [Fintype Feature] (lambda delta : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"standardhighprobabilityregretlogbudget banditrlproof.oful.standardhighprobabilityregretlogbudget the confidence logarithm for outer failure probability `delta` over the complete finite window `0, ..., horizon`. definition compiled","shard":"modules/f8cc3317a82480ed.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_highProbabilityRegret","label":"standardScalarConfidenceRadiusUpper_highProbabilityRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_highProbabilityRegret","description":"At algorithm parameter `delta / (T+1)`, the standard confidence radius has the explicit form `R * sqrt (B_T + 2 * log ((T+1)/delta)) + sqrt lambda * S`.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html#decl-c5cbea7389c5","parent":"module:BanditRLProof.OFULHighProbabilityRegretRate","order":6305,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHighProbabilityRegretRate"],["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean:34"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarConfidenceRadiusUpper_highProbabilityRegret {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : standardScalarConfidenceRadiusUpper (Feature := Feature) R (delta / (((horizon + 1 : Nat) : Real))) lambda S (horizon + 1) L2 = R * Real.sqrt (standardHighProbabilityRegretLogBudget (Feature := Feature) lambda delta horizon L2) + Real.sqrt lambda * S","missing":[],"search":"standardscalarconfidenceradiusupper_highprobabilityregret banditrlproof.oful.standardscalarconfidenceradiusupper_highprobabilityregret at algorithm parameter `delta / (t+1)`, the standard confidence radius has the explicit form `r * sqrt (b_t + 2 * log ((t+1)/delta)) + sqrt lambda * s`. theorem compiled","shard":"modules/f8cc3317a82480ed.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardHighProbabilityPseudoRegretBound","label":"standardHighProbabilityPseudoRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardHighProbabilityPseudoRegretBound","description":"Explicit complete finite-window high-probability pseudo-regret budget.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html#decl-f252e692ceb2","parent":"module:BanditRLProof.OFULHighProbabilityRegretRate","order":6306,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULHighProbabilityRegretRate"],["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean:98"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardHighProbabilityPseudoRegretBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"standardhighprobabilitypseudoregretbound banditrlproof.oful.standardhighprobabilitypseudoregretbound explicit complete finite-window high-probability pseudo-regret budget. definition compiled","shard":"modules/f8cc3317a82480ed.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound_eq_standardHighProbabilityPseudoRegretBound","label":"standardScalarAllRoundGapBound_eq_standardHighProbabilityPseudoRegretBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapBound_eq_standardHighProbabilityPseudoRegretBound","description":"The named all-round gap budget is exactly the explicit rate expression.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html#decl-64bef486f704","parent":"module:BanditRLProof.OFULHighProbabilityRegretRate","order":6307,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHighProbabilityRegretRate"],["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean:114"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarAllRoundGapBound_eq_standardHighProbabilityPseudoRegretBound {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : standardScalarAllRoundGapBound (Feature := Feature) R delta lambda S horizon L2 = standardHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"standardscalarallroundgapbound_eq_standardhighprobabilitypseudoregretbound banditrlproof.oful.standardscalarallroundgapbound_eq_standardhighprobabilitypseudoregretbound the named all-round gap budget is exactly the explicit rate expression. theorem compiled","shard":"modules/f8cc3317a82480ed.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret","label":"canonicalStandardHighProbabilityPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret","description":"Complete fixed-optimal-arm pseudo-regret along one canonical trajectory.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html#decl-1f3c02f3ec52","parent":"module:BanditRLProof.OFULHighProbabilityRegretRate","order":6308,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULHighProbabilityRegretRate"],["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean:134"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalStandardHighProbabilityPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret complete fixed-optimal-arm pseudo-regret along one canonical trajectory. definition compiled","shard":"modules/f8cc3317a82480ed.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"canonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Explicit finite-window high-probability pseudo-regret theorem for the canonical scalar-ridge OFUL trajectory. The horizon-`T` algorithm is run at `delta / (T+1)`. The result controls the complete rounds `0, ..., T`; it is not an anytime or all-horizon statement.","url":"../modules/banditrlproof-ofulhighprobabilityregretrate/index.html#decl-9234c1a7569f","parent":"module:BanditRLProof.OFULHighProbabilityRegretRate","order":6309,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHighProbabilityRegretRate"],["Source","BanditRLProof/OFULHighProbabilityRegretRate.lean:150"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : let algorithmDelta := d…","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_nonneg_and_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_nonneg_and_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization explicit finite-window high-probability pseudo-regret theorem for the canonical scalar-ridge oful trajectory. the horizon-`t` algorithm is run at `delta / (t+1)`. the result controls the complete rounds `0, ..., t`; it is not an anytime or all-horizon statement. theorem compiled","shard":"modules/f8cc3317a82480ed.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_map_eq_historyAlgorithmInitialFeedback","label":"canonicalHistoryTrajectory_initialReward_map_eq_historyAlgorithmInitialFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_map_eq_historyAlgorithmInitialFeedback","description":"For any canonical history algorithm, the time-zero reward marginal is the environment feedback law at that algorithm's initial action.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-b0e44a99c891","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6310,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_initialReward_map_eq_historyAlgorithmInitialFeedback {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (initialAction : Fin K) (hinitial : algorithm.initialAction = Measure.dirac initialAction) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Measure.map (fun trajectory : Nat -> Fin K × Real => trajectory 0 |>.2) (Thompson.canonicalHistoryTrajectoryMeasure algorithm environment) = environment.initialFeedback initialAction","missing":[],"search":"canonicalhistorytrajectory_initialreward_map_eq_historyalgorithminitialfeedback banditrlproof.oful.canonicalhistorytrajectory_initialreward_map_eq_historyalgorithminitialfeedback for any canonical history algorithm, the time-zero reward marginal is the environment feedback law at that algorithm's initial action. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condDistrib_eq_historyStepReward","label":"canonicalHistoryTrajectory_reward_succ_condDistrib_eq_historyStepReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condDistrib_eq_historyStepReward","description":"For any canonical history algorithm, the successor reward conditioned on its finite pair prefix has the reward marginal of the history-step kernel.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-0f58f0cbe89e","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6311,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:61"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_reward_succ_condDistrib_eq_historyStepReward {K : Nat} (hK : 0 < K) (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : Nat -> Fin K × Real => (trajectory (n + 1)).2) (Preorder.frestrictLe n) (Thompson.canonicalHistoryTrajectoryMeasure algorithm environment) =ᵐ[ (Thompson.canonicalHistoryTrajectoryMeasure algorithm environment).map (Preorder.frestrictLe n)] (Thompson.historyStepKernel algorithm environment n).map Prod.snd","missing":[],"search":"canonicalhistorytrajectory_reward_succ_conddistrib_eq_historystepreward banditrlproof.oful.canonicalhistorytrajectory_reward_succ_conddistrib_eq_historystepreward for any canonical history algorithm, the successor reward conditioned on its finite pair prefix has the reward marginal of the history-step kernel. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_historyStepReward_comap","label":"canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_historyStepReward_comap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_historyStepReward_comap","description":"Algorithm-parametric trimmed conditional-expectation reward law at successor time for a canonical history trajectory.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-119bc8e1dca6","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6312,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:117"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_historyStepReward_comap {K : Nat} (hK : 0 < K) (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : Filter.Eventually (fun trajectory : Nat -> Fin K × Real => Measure.map (fun y : Nat -> Fin K × Real => (y (n + 1)).2) (ProbabilityTheory.condExpKernel (Thompson.canonicalHistoryTrajectoryMeasure algorithm environment) ((inferInstance : MeasurableSpace (History.FinitePairHistory (Fin K) Real n)).comap (Preorder.frestrictLe n)) trajectory) = (Thompson.historyStepKernel algorithm environment n).map Prod.snd (Preorder.frestrictLe n trajectory)) (ae ((Thompson.canonicalHistoryTrajectoryMeasure algorithm environment).trim ((History.measurable_finitePairHistoryOfTrace Thompson.canonicalHistoryTrajectoryAction Thompson.canonicalHistoryTrajectoryReward Thompson.m…","missing":[],"search":"canonicalhistorytrajectory_reward_succ_condexpkernel_map_eq_historystepreward_comap banditrlproof.oful.canonicalhistorytrajectory_reward_succ_condexpkernel_map_eq_historystepreward_comap algorithm-parametric trimmed conditional-expectation reward law at successor time for a canonical history trajectory. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_historyAlgorithmInitialFeedback_unitComap","label":"canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_historyAlgorithmInitialFeedback_unitComap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_historyAlgorithmInitialFeedback_unitComap","description":"Algorithm-parametric time-zero conditional reward law on the trivial sigma-algebra.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-7c1cc6d29f05","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6313,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:170"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_historyAlgorithmInitialFeedback_unitComap {K : Nat} (algorithm : Thompson.HistoryAlgorithm (Fin K) Real) (initialAction : Fin K) (hinitial : algorithm.initialAction = Measure.dirac initialAction) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Filter.Eventually (fun trajectory : Nat -> Fin K × Real => Measure.map (fun y : Nat -> Fin K × Real => (y 0).2) (ProbabilityTheory.condExpKernel (Thompson.canonicalHistoryTrajectoryMeasure algorithm environment) ((inferInstance : MeasurableSpace Unit).comap (fun _trajectory : Nat -> Fin K × Real => ())) trajectory) = environment.initialFeedback initialAction) (ae ((Thompson.canonicalHistoryTrajectoryMeasure algorithm environment).trim ((show Measurable (fun _trajectory : Nat -> Fin K × Real => ()) from measurable_const).comap_le)))","missing":[],"search":"canonicalhistorytrajectory_initialreward_condexpkernel_map_eq_historyalgorithminitialfeedback_unitcomap banditrlproof.oful.canonicalhistorytrajectory_initialreward_condexpkernel_map_eq_historyalgorithminitialfeedback_unitcomap algorithm-parametric time-zero conditional reward law on the trivial sigma-algebra. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidge_historyStepKernel_map_snd","label":"finiteHistoryScalarRidge_historyStepKernel_map_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScalarRidge_historyStepKernel_map_snd","description":"For the deterministic finite-history OFUL policy, the reward marginal of the next pair kernel is exactly the environment feedback kernel selected by the current history.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-53137783c16c","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6314,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:244"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScalarRidge_historyStepKernel_map_snd {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (Thompson.historyStepKernel (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment n).map Prod.snd history = environment.feedback n (history, finiteHistoryScalarRidgeOptimisticAction hK lambda actionFeature R algorithmDelta S n history)","missing":[],"search":"finitehistoryscalarridge_historystepkernel_map_snd banditrlproof.oful.finitehistoryscalarridge_historystepkernel_map_snd for the deterministic finite-history oful policy, the reward marginal of the next pair kernel is exactly the environment feedback kernel selected by the current history. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_map_eq_scalarRidgeInitialFeedback","label":"canonicalHistoryTrajectory_initialReward_map_eq_scalarRidgeInitialFeedback","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_map_eq_scalarRidgeInitialFeedback","description":"The initial canonical reward marginal is the initial-arm feedback law.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-53ceabdfc896","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6315,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:288"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_initialReward_map_eq_scalarRidgeInitialFeedback {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Measure.map (fun trajectory : Nat -> Fin K × Real => trajectory 0 |>.2) (Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment) = environment.initialFeedback ⟨0, hK⟩","missing":[],"search":"canonicalhistorytrajectory_initialreward_map_eq_scalarridgeinitialfeedback banditrlproof.oful.canonicalhistorytrajectory_initialreward_map_eq_scalarridgeinitialfeedback the initial canonical reward marginal is the initial-arm feedback law. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condDistrib_eq_scalarRidgeStepReward","label":"canonicalHistoryTrajectory_reward_succ_condDistrib_eq_scalarRidgeStepReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condDistrib_eq_scalarRidgeStepReward","description":"The successor canonical reward conditioned on the finite pair prefix has the reward marginal of the concrete history step kernel.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-2547ecff27f0","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6316,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:333"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_reward_succ_condDistrib_eq_scalarRidgeStepReward {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : Nat -> Fin K × Real => (trajectory (n + 1)).2) (Preorder.frestrictLe n) (Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment) =ᵐ[ (Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment).map (Preorder.frestrictLe n)] (Thompson.historyStepKernel (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment n)…","missing":[],"search":"canonicalhistorytrajectory_reward_succ_conddistrib_eq_scalarridgestepreward banditrlproof.oful.canonicalhistorytrajectory_reward_succ_conddistrib_eq_scalarridgestepreward the successor canonical reward conditioned on the finite pair prefix has the reward marginal of the concrete history step kernel. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_scalarRidgeStepReward_comap","label":"canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_scalarRidgeStepReward_comap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_scalarRidgeStepReward_comap","description":"Trimmed conditional-expectation reward law at successor time for the concrete canonical OFUL trajectory.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-927bd838f45d","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6317,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:403"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_scalarRidgeStepReward_comap {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : Filter.Eventually (fun trajectory : Nat -> Fin K × Real => Measure.map (fun y : Nat -> Fin K × Real => (y (n + 1)).2) (ProbabilityTheory.condExpKernel (Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment) ((inferInstance : MeasurableSpace (History.FinitePairHistory (Fin K) Real n)).comap (Preorder.frestrictLe n)) trajectory) = (Thompson.historyStepKernel (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment n).map Prod.snd…","missing":[],"search":"canonicalhistorytrajectory_reward_succ_condexpkernel_map_eq_scalarridgestepreward_comap banditrlproof.oful.canonicalhistorytrajectory_reward_succ_condexpkernel_map_eq_scalarridgestepreward_comap trimmed conditional-expectation reward law at successor time for the concrete canonical oful trajectory. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_scalarRidgeInitialFeedback_unitComap","label":"canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_scalarRidgeInitialFeedback_unitComap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_scalarRidgeInitialFeedback_unitComap","description":"At time zero, conditioning on the trivial sigma-algebra leaves the canonical initial reward law equal to the initial-arm feedback measure.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-df91a7cf409d","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6318,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:469"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_scalarRidgeInitialFeedback_unitComap {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Filter.Eventually (fun trajectory : Nat -> Fin K × Real => Measure.map (fun y : Nat -> Fin K × Real => (y 0).2) (ProbabilityTheory.condExpKernel (Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment) ((inferInstance : MeasurableSpace Unit).comap (fun _trajectory : Nat -> Fin K × Real => ())) trajectory) = environment.initialFeedback ⟨0, hK⟩) (ae ((Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithm…","missing":[],"search":"canonicalhistorytrajectory_initialreward_condexpkernel_map_eq_scalarridgeinitialfeedback_unitcomap banditrlproof.oful.canonicalhistorytrajectory_initialreward_condexpkernel_map_eq_scalarridgeinitialfeedback_unitcomap at time zero, conditioning on the trivial sigma-algebra leaves the canonical initial reward law equal to the initial-arm feedback measure. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.CanonicalLinearSubgaussianEnvironmentLaw","label":"CanonicalLinearSubgaussianEnvironmentLaw","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.CanonicalLinearSubgaussianEnvironmentLaw","description":"A concrete linear reward contract on the kernels of a history environment. The initial and successor fields are unconditional sub-Gaussian laws of each kernel section around the linear response of the supplied action.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-e3d6ace524ac","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6319,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:554"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure CanonicalLinearSubgaussianEnvironmentLaw {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Prop where","missing":[],"search":"canonicallinearsubgaussianenvironmentlaw banditrlproof.oful.canonicallinearsubgaussianenvironmentlaw a concrete linear reward contract on the kernels of a history environment. the initial and successor fields are unconditional sub-gaussian laws of each kernel section around the linear response of the supplied action. structure compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","label":"canonicalPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","description":"The kernel-level linear sub-Gaussian environment contract constructs the strict-past predictable residual law required by the OFUL confidence route.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-287041bbafb0","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6320,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:582"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : CanonicalPredictableScalarRidgeResidualLaw hK lambda thetaStar actionFeature R algorithmDelta S environment horizon where","missing":[],"search":"canonicalpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment banditrlproof.oful.canonicalpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment the kernel-level linear sub-gaussian environment contract constructs the strict-past predictable residual law required by the oful confidence route. definition compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","label":"measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","description":"Concrete successor-gap tail obtained directly from a linear sub-Gaussian `HistoryEnvironment` contract.","url":"../modules/banditrlproof-ofulhistoryenvironmentrewardlaw/index.html#decl-7e152ee63dcf","parent":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","order":6321,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULHistoryEnvironmentRewardLaw"],["Source","BanditRLProof/OFULHistoryEnvironmentRewardLaw.lean:787"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment (canonicalHistoryTrajectorySumRangeSuccGapViolationSet lambda thetaStar actionFeature R delta S horizon comparator) <=…","missing":[],"search":"measure_canonicalhistorytrajectorysumrangesuccgapviolationset_le_of_linearsubgaussianenvironment banditrlproof.oful.measure_canonicalhistorytrajectorysumrangesuccgapviolationset_le_of_linearsubgaussianenvironment concrete successor-gap tail obtained directly from a linear sub-gaussian `historyenvironment` contract. theorem compiled","shard":"modules/be65fb476a795048.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_linearValue_le_euclideanLength_mul_euclideanLength","label":"abs_linearValue_le_euclideanLength_mul_euclideanLength","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_linearValue_le_euclideanLength_mul_euclideanLength","description":"Ordinary finite-dimensional Cauchy--Schwarz on the local linear-value surface.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-52a0da5f04dd","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6322,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:19"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_linearValue_le_euclideanLength_mul_euclideanLength {Feature : Type u} [Fintype Feature] (theta x : Feature -> Real) : |linearValue theta x| <= euclideanLength theta * euclideanLength x","missing":[],"search":"abs_linearvalue_le_euclideanlength_mul_euclideanlength banditrlproof.oful.abs_linearvalue_le_euclideanlength_mul_euclideanlength ordinary finite-dimensional cauchy--schwarz on the local linear-value surface. theorem compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_linearValue_le_parameterFeatureBound","label":"abs_linearValue_le_parameterFeatureBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_linearValue_le_parameterFeatureBound","description":"A common parameter/feature envelope bounds the absolute linear value.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-ea91a9b29f48","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6323,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:40"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_linearValue_le_parameterFeatureBound {Feature : Type u} [Fintype Feature] (theta x : Feature -> Real) (S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength theta <= S) (hx : dotProduct x x <= L2) : |linearValue theta x| <= S * Real.sqrt L2","missing":[],"search":"abs_linearvalue_le_parameterfeaturebound banditrlproof.oful.abs_linearvalue_le_parameterfeaturebound a common parameter/feature envelope bounds the absolute linear value. theorem compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","label":"linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","description":"Two arms sharing the same feature envelope differ by at most `2*S*sqrt L2`.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-afbebcf3527e","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6324,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:54"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem linearValue_sub_linearValue_le_two_mul_parameterFeatureBound {Feature : Type u} [Fintype Feature] (theta x y : Feature -> Real) (S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength theta <= S) (hx : dotProduct x x <= L2) (hy : dotProduct y y <= L2) : linearValue theta x - linearValue theta y <= 2 * S * Real.sqrt L2","missing":[],"search":"linearvalue_sub_linearvalue_le_two_mul_parameterfeaturebound banditrlproof.oful.linearvalue_sub_linearvalue_le_two_mul_parameterfeaturebound two arms sharing the same feature envelope differ by at most `2*s*sqrt l2`. theorem compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarInitialGapBound","label":"standardScalarInitialGapBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarInitialGapBound","description":"Deterministic charge used for the canonical time-zero arm.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-0844d611ff67","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6325,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:72"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardScalarInitialGapBound (S L2 : Real) : Real","missing":[],"search":"standardscalarinitialgapbound banditrlproof.oful.standardscalarinitialgapbound deterministic charge used for the canonical time-zero arm. definition compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialGap_le_ae","label":"canonicalHistoryTrajectory_initialGap_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_initialGap_le_ae","description":"The canonical OFUL initial action is the fixed arm `0` almost surely, so its linear gap to any comparator is bounded by the common parameter/feature envelope.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-0c416450477c","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6326,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:80"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_initialGap_le_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R algorithmDelta S : Real) (hS : 0 <= S) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (environment : Thompson.HistoryEnvironment (Fin K) Real) (comparator : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R algorithmDelta S) environment, linearValue thetaStar (actionFeature comparator) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0)) <= standardScalarInitialGapBound S L2","missing":[],"search":"canonicalhistorytrajectory_initialgap_le_ae banditrlproof.oful.canonicalhistorytrajectory_initialgap_le_ae the canonical oful initial action is the fixed arm `0` almost surely, so its linear gap to any comparator is bounded by the common parameter/feature envelope. theorem compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound","label":"standardScalarAllRoundGapBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapBound","description":"Standard finite-window cumulative-gap budget including time zero.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-a6861b74b869","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6327,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:117"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def standardScalarAllRoundGapBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"standardscalarallroundgapbound banditrlproof.oful.standardscalarallroundgapbound standard finite-window cumulative-gap budget including time zero. definition compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","label":"canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","description":"Violation of the standard OFUL cumulative-gap budget over rounds `0, ..., horizon`.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-6666236fe390","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6328,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:127"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (horizon : Nat) (L2 : Real) (comparator : Nat -> Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"canonicalhistorytrajectorysumrangeallgapstandardviolationset banditrlproof.oful.canonicalhistorytrajectorysumrangeallgapstandardviolationset violation of the standard oful cumulative-gap budget over rounds `0, ..., horizon`. definition compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_ae_le_succ","label":"canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_ae_le_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_ae_le_succ","description":"Almost surely, an all-round violation implies the previously compiled successor-only standard-budget violation.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-3953403c2f31","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6329,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:150"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_ae_le_succ {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimisticAlgorithm hK lambda actionFeature R (delta / ((horizon + 1 : Nat) : Real)) S) environment, trajectory ∈ canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet lambda thetaStar actionFeature R delta S horizon L2 comparator -> trajectory ∈ canonicalHistoryTr…","missing":[],"search":"canonicalhistorytrajectorysumrangeallgapstandardviolationset_ae_le_succ banditrlproof.oful.canonicalhistorytrajectorysumrangeallgapstandardviolationset_ae_le_succ almost surely, an all-round violation implies the previously compiled successor-only standard-budget violation. theorem compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"measure_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Concrete high-probability standard cumulative-gap theorem over all rounds `0, ..., horizon`, with automatic normalized width discharge.","url":"../modules/banditrlproof-ofulinitialroundgap/index.html#decl-b417765c0adf","parent":"module:BanditRLProof.OFULInitialRoundGap","order":6330,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULInitialRoundGap"],["Source","BanditRLProof/OFULInitialRoundGap.lean:211"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptimi…","missing":[],"search":"measure_canonicalhistorytrajectorysumrangeallgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.measure_canonicalhistorytrajectorysumrangeallgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization concrete high-probability standard cumulative-gap theorem over all rounds `0, ..., horizon`, with automatic normalized width discharge. theorem compiled","shard":"modules/ad567d08ec5c40f2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticScore","label":"finiteHistoryOptimisticScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryOptimisticScore","description":"The OFUL score of one action at one inclusive finite pair history.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-bcc6ae240455","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6331,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryOptimisticScore {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (action : Fin K) : Real","missing":[],"search":"finitehistoryoptimisticscore banditrlproof.oful.finitehistoryoptimisticscore the oful score of one action at one inclusive finite pair history. definition compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAction","label":"finiteHistoryOptimisticAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryOptimisticAction","description":"Deterministic history OFUL selector with the strict-update finite fold. The strict comparison gives a fixed deterministic tie behavior.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-a5b251761567","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6332,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:46"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryOptimisticAction {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : Fin K","missing":[],"search":"finitehistoryoptimisticaction banditrlproof.oful.finitehistoryoptimisticaction deterministic history oful selector with the strict-update finite fold. the strict comparison gives a fixed deterministic tie behavior. definition compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAction_score_max","label":"finiteHistoryOptimisticAction_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryOptimisticAction_score_max","description":"The history selector maximizes the current OFUL score.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-aa9bd43c5b26","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6333,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:67"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryOptimisticAction_score_max {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) (action : Fin K) : finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history action <= finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history (finiteHistoryOptimisticAction hK thetaHat V beta candidateFeature n history)","missing":[],"search":"finitehistoryoptimisticaction_score_max banditrlproof.oful.finitehistoryoptimisticaction_score_max the history selector maximizes the current oful score. theorem compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryOptimisticAction","label":"measurable_finiteHistoryOptimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryOptimisticAction","description":"The strict-fold history selector is measurable when each fixed-action score coordinate is measurable.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-4481b0385a5f","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6334,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:97"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryOptimisticAction {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (n : Nat) (hscores : forall action : Fin K, Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history action)) : Measurable (finiteHistoryOptimisticAction hK thetaHat V beta candidateFeature n)","missing":[],"search":"measurable_finitehistoryoptimisticaction banditrlproof.oful.measurable_finitehistoryoptimisticaction the strict-fold history selector is measurable when each fixed-action score coordinate is measurable. theorem compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAlgorithm","label":"finiteHistoryOptimisticAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryOptimisticAlgorithm","description":"The measurable history selector packaged as a deterministic algorithm.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-f6d5f709ffd3","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6335,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:126"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryOptimisticAlgorithm {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (hscores : forall (n : Nat) (action : Fin K), Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history action)) : Thompson.HistoryAlgorithm (Fin K) Reward where","missing":[],"search":"finitehistoryoptimisticalgorithm banditrlproof.oful.finitehistoryoptimisticalgorithm the measurable history selector packaged as a deterministic algorithm. definition compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAlgorithm_policy_apply","label":"finiteHistoryOptimisticAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryOptimisticAlgorithm_policy_apply","description":"Every policy section of the history OFUL algorithm is the selector Dirac law.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-b2c93541dd34","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6336,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:155"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"@[simp] theorem finiteHistoryOptimisticAlgorithm_policy_apply {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (hscores : forall (n : Nat) (action : Fin K), Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history action)) (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : (finiteHistoryOptimisticAlgorithm hK thetaHat V beta candidateFeature hscores).policy n…","missing":[],"search":"finitehistoryoptimisticalgorithm_policy_apply banditrlproof.oful.finitehistoryoptimisticalgorithm_policy_apply every policy section of the history oful algorithm is the selector dirac law. theorem compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryOptimisticAction","label":"canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryOptimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryOptimisticAction","description":"Along the canonical recursive trajectory, the successor action is almost surely the measurable OFUL selector evaluated on the realized finite history.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-eb4d5f31810f","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6337,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:187"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryOptimisticAction {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (hscores : forall (n : Nat) (action : Fin K), Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history action)) (environment : Thompson.HistoryEnvironment (Fin K) Reward) (n : Nat) : ∀ᵐ trajectory ∂…","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_finitehistoryoptimisticaction banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_finitehistoryoptimisticaction along the canonical recursive trajectory, the successor action is almost surely the measurable oful selector evaluated on the realized finite history. theorem compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticSelectedFeature","label":"finiteHistoryOptimisticSelectedFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryOptimisticSelectedFeature","description":"Candidate feature selected by the measurable OFUL history action.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-eb832f0dc879","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6338,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:276"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryOptimisticSelectedFeature {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Reward n) : Feature -> Real","missing":[],"search":"finitehistoryoptimisticselectedfeature banditrlproof.oful.finitehistoryoptimisticselectedfeature candidate feature selected by the measurable oful history action. definition compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_candidateFeature_succ_ae_eq_selectedFeature","label":"canonicalHistoryTrajectory_candidateFeature_succ_ae_eq_selectedFeature","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_candidateFeature_succ_ae_eq_selectedFeature","description":"The feature indexed by the actual canonical successor action agrees almost surely with the feature selected from the realized history.","url":"../modules/banditrlproof-ofulmeasurablerecursiveselection/index.html#decl-9d59726b2c99","parent":"module:BanditRLProof.OFULMeasurableRecursiveSelection","order":6339,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULMeasurableRecursiveSelection"],["Source","BanditRLProof/OFULMeasurableRecursiveSelection.lean:301"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_candidateFeature_succ_ae_eq_selectedFeature {K : Nat} {Reward : Type u} {Feature : Type v} [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaHat : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Feature -> Real) (V : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Matrix Feature Feature Real) (beta : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Real) (candidateFeature : (n : Nat) -> History.FinitePairHistory (Fin K) Reward n -> Fin K -> Feature -> Real) (hscores : forall (n : Nat) (action : Fin K), Measurable (fun history : History.FinitePairHistory (Fin K) Reward n => finiteHistoryOptimisticScore thetaHat V beta candidateFeature n history action)) (environment : Thompson.HistoryEnvironment (Fin K) Reward) (n : Nat) : ∀ᵐ trajectory ∂ Thom…","missing":[],"search":"canonicalhistorytrajectory_candidatefeature_succ_ae_eq_selectedfeature banditrlproof.oful.canonicalhistorytrajectory_candidatefeature_succ_ae_eq_selectedfeature the feature indexed by the actual canonical successor action agrees almost surely with the feature selected from the realized history. theorem compiled","shard":"modules/f8e33688f91d2e87.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_quadratic_le_one","label":"regularizedPrefixFeatureGram_inv_quadratic_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_quadratic_le_one","description":"theorem regularizedPrefixFeatureGram_inv_quadratic_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) (hx : dotProduct x x <= lambda) : dotProduct x ((regularizedPrefixFeatureGram lambda feature T)⁻¹.mulVec x) <= 1","url":"../modules/banditrlproof-ofulnormalizedradiuswidth/index.html#decl-94fc43096dab","parent":"module:BanditRLProof.OFULNormalizedRadiusWidth","order":6340,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULNormalizedRadiusWidth"],["Source","BanditRLProof/OFULNormalizedRadiusWidth.lean:18"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem regularizedPrefixFeatureGram_inv_quadratic_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) (hx : dotProduct x x <= lambda) : dotProduct x ((regularizedPrefixFeatureGram lambda feature T)⁻¹.mulVec x) <= 1","missing":[],"search":"regularizedprefixfeaturegram_inv_quadratic_le_one banditrlproof.oful.regularizedprefixfeaturegram_inv_quadratic_le_one theorem regularizedprefixfeaturegram_inv_quadratic_le_one {feature : type u} [fintype feature] [decidableeq feature] (lambda : real) (hlambda : 0 < lambda) (feature : nat -> feature -> real) (t : nat) (x : feature -> real) (hx : dotproduct x x <= lambda) : dotproduct x ((regularizedprefixfeaturegram lambda feature t)⁻¹.mulvec x) <= 1 theorem compiled","shard":"modules/a72da139f18910bb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.confidenceWidth_regularizedPrefixFeatureGram_le_one","label":"confidenceWidth_regularizedPrefixFeatureGram_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.confidenceWidth_regularizedPrefixFeatureGram_le_one","description":"theorem confidenceWidth_regularizedPrefixFeatureGram_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) (hx : dotProduct x x <= lambda) : confidenceWidth (regularizedPrefixFeatureGram lambda feature T) x <= 1","url":"../modules/banditrlproof-ofulnormalizedradiuswidth/index.html#decl-30ff4f5691b6","parent":"module:BanditRLProof.OFULNormalizedRadiusWidth","order":6341,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULNormalizedRadiusWidth"],["Source","BanditRLProof/OFULNormalizedRadiusWidth.lean:80"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem confidenceWidth_regularizedPrefixFeatureGram_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Feature -> Real) (T : Nat) (x : Feature -> Real) (hx : dotProduct x x <= lambda) : confidenceWidth (regularizedPrefixFeatureGram lambda feature T) x <= 1","missing":[],"search":"confidencewidth_regularizedprefixfeaturegram_le_one banditrlproof.oful.confidencewidth_regularizedprefixfeaturegram_le_one theorem confidencewidth_regularizedprefixfeaturegram_le_one {feature : type u} [fintype feature] [decidableeq feature] (lambda : real) (hlambda : 0 < lambda) (feature : nat -> feature -> real) (t : nat) (x : feature -> real) (hx : dotproduct x x <= lambda) : confidencewidth (regularizedprefixfeaturegram lambda feature t) x <= 1 theorem compiled","shard":"modules/a72da139f18910bb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_confidenceWidth_le_one","label":"canonicalHistoryTrajectory_confidenceWidth_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_confidenceWidth_le_one","description":"theorem canonicalHistoryTrajectory_confidenceWidth_le_one {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (trajectory : Nat -> Fin K × Real) (t : Nat) : confidenceWidth (finit…","url":"../modules/banditrlproof-ofulnormalizedradiuswidth/index.html#decl-0eb010c02d7e","parent":"module:BanditRLProof.OFULNormalizedRadiusWidth","order":6342,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULNormalizedRadiusWidth"],["Source","BanditRLProof/OFULNormalizedRadiusWidth.lean:92"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_confidenceWidth_le_one {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (trajectory : Nat -> Fin K × Real) (t : Nat) : confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) t trajectory) (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory t)) <= 1","missing":[],"search":"canonicalhistorytrajectory_confidencewidth_le_one banditrlproof.oful.canonicalhistorytrajectory_confidencewidth_le_one theorem canonicalhistorytrajectory_confidencewidth_le_one {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (lambda : real) (hlambda : 0 < lambda) (actionfeature : fin k -> feature -> real) (l2 : real) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (trajectory : nat -> fin k × real) (t : nat) : confidencewidth (finitehorizonscalargram lambda (canonicalhistorytrajectoryfeature actionfeature) t trajectory) (actionfeature (thompson.canonicalhistorytrajectoryaction trajectory t)) <= 1 theorem compiled","shard":"modules/a72da139f18910bb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard_of_featureBound_le_regularization","label":"canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard_of_featureBound_le_regularization","description":"theorem canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action,…","url":"../modules/banditrlproof-ofulnormalizedradiuswidth/index.html#decl-f5e35caab564","parent":"module:BanditRLProof.OFULNormalizedRadiusWidth","order":6343,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULNormalizedRadiusWidth"],["Source","BanditRLProof/OFULNormalizedRadiusWidth.lean:120"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (trajectory : Nat -> Fin K × Real) : (Finset.range horizon).sum (fun n => 2 * finiteHorizonScalarConfidenceRadius (canonicalHistoryTrajectoryFeature actionFeature) R (delta / ((horizon + 1 : Nat) : Real)) lambda S (n + 1) trajectory * confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) (n + 1) trajectory) (actionFeature (Thompson.canonicalHi…","missing":[],"search":"canonicalhistorytrajectory_sum_range_succ_radius_mul_width_le_standard_of_featurebound_le_regularization banditrlproof.oful.canonicalhistorytrajectory_sum_range_succ_radius_mul_width_le_standard_of_featurebound_le_regularization theorem canonicalhistorytrajectory_sum_range_succ_radius_mul_width_le_standard_of_featurebound_le_regularization {k : nat} {feature : type u} [fintype feature] [decidableeq feature] [nonempty feature] (lambda : real) (hlambda : 0 < lambda) (actionfeature : fin k -> feature -> real) (r delta s : real) (hdelta : 0 < delta) (hs : 0 <= s) (horizon : nat) (l2 : real) (hl2 : 0 <= l2) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (trajectory : nat -> fin k × real) : (finset.range horizon).sum (fun n => 2 * finitehorizonscalarconfidenceradius (canonicalhistorytrajectoryfeature actionfeature) r (delta / ((horizon + 1 : nat) : real)) lambda s (n + 1) trajectory * confidencewidth (finitehorizonscalargram lambda (canonicalhistorytrajectoryfeature actionfeature) (n + 1) trajectory) (actionfeature (thompson.canonicalhistorytrajectoryaction trajectory (n + 1)))) <= standardscalarradiuswidthbound (feature := feature) r (delta / ((horizon + 1 : nat) : real)) lambda s (horizon + 1) l2 theorem compiled","shard":"modules/a72da139f18910bb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"theorem measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta)…","url":"../modules/banditrlproof-ofulnormalizedradiuswidth/index.html#decl-7c589c0a15d3","parent":"module:BanditRLProof.OFULNormalizedRadiusWidth","order":6344,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULNormalizedRadiusWidth"],["Source","BanditRLProof/OFULNormalizedRadiusWidth.lean:158"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScalarRidgeOptim…","missing":[],"search":"measure_canonicalhistorytrajectorysumrangesuccgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.measure_canonicalhistorytrajectorysumrangesuccgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization theorem measure_canonicalhistorytrajectorysumrangesuccgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization {k : nat} {feature : type u} [fintype feature] [decidableeq feature] [nonempty feature] (hk : 0 < k) (lambda : real) (hlambda : 0 < lambda) (thetastar : feature -> real) (actionfeature : fin k -> feature -> real) (r : real) (hr : 0 < r) (delta : real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (s : real) (hs : 0 <= s) (environment : thompson.historyenvironment (fin k) real) (horizon : nat) (l2 : real) (hl2 : 0 <= l2) (hactionfeaturebound : forall action, dotproduct (actionfeature action) (actionfeature action) <= l2) (hl2lambda : l2 <= lambda) (comparator : nat -> fin k) (source : canonicallinearsubgaussianenvironmentlaw hk thetastar actionfeature r s environment) : thompson.canonicalhistorytrajectorymeasure (finitehistoryscalarridgeoptimisticalgorithm hk lambda actionfeature r (delta / ((horizon + 1 : nat) : real)) s) environment (canonicalhistorytrajectorysumrangesuccgapstandardviolationset lambda thetastar actionfeature r delta s hori…","shard":"modules/a72da139f18910bb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.euclideanLength","label":"euclideanLength","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.euclideanLength","description":"The Euclidean length written on the same finite-coordinate surface as `dotProduct`.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html#decl-7cafb67ae18b","parent":"module:BanditRLProof.OFULScalarRegularizationBias","order":6345,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScalarRegularizationBias"],["Source","BanditRLProof/OFULScalarRegularizationBias.lean:18"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def euclideanLength [Fintype Feature] (theta : Feature -> Real) : Real","missing":[],"search":"euclideanlength banditrlproof.oful.euclideanlength the euclidean length written on the same finite-coordinate surface as `dotproduct`. definition compiled","shard":"modules/b7ab38cd1a862f90.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scalarIdentity_posDef","label":"scalarIdentity_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.scalarIdentity_posDef","description":"A positive scalar multiple of the identity is positive definite.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html#decl-f425dbd66d4a","parent":"module:BanditRLProof.OFULScalarRegularizationBias","order":6346,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScalarRegularizationBias"],["Source","BanditRLProof/OFULScalarRegularizationBias.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem scalarIdentity_posDef [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) : (Matrix.scalar Feature lambda : Matrix Feature Feature Real).PosDef","missing":[],"search":"scalaridentity_posdef banditrlproof.oful.scalaridentity_posdef a positive scalar multiple of the identity is positive definite. theorem compiled","shard":"modules/b7ab38cd1a862f90.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.matrixNorm_nonsingInv_scalar_mulVec_le_sqrt_mul_euclideanLength","label":"matrixNorm_nonsingInv_scalar_mulVec_le_sqrt_mul_euclideanLength","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.matrixNorm_nonsingInv_scalar_mulVec_le_sqrt_mul_euclideanLength","description":"For `V = lambda I + G` with `G` positive semidefinite, the `V`-norm of the ridge bias `V⁻¹ (lambda theta)` is at most `sqrt lambda * ‖theta‖₂`.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html#decl-44e655d24d9e","parent":"module:BanditRLProof.OFULScalarRegularizationBias","order":6347,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScalarRegularizationBias"],["Source","BanditRLProof/OFULScalarRegularizationBias.lean:34"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem matrixNorm_nonsingInv_scalar_mulVec_le_sqrt_mul_euclideanLength [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (G : Matrix Feature Feature Real) (hG : G.PosSemidef) (theta : Feature -> Real) : matrixNorm (Matrix.scalar Feature lambda + G) ((Matrix.scalar Feature lambda + G)⁻¹.mulVec ((Matrix.scalar Feature lambda).mulVec theta)) <= Real.sqrt lambda * euclideanLength theta","missing":[],"search":"matrixnorm_nonsinginv_scalar_mulvec_le_sqrt_mul_euclideanlength banditrlproof.oful.matrixnorm_nonsinginv_scalar_mulvec_le_sqrt_mul_euclideanlength for `v = lambda i + g` with `g` positive semidefinite, the `v`-norm of the ridge bias `v⁻¹ (lambda theta)` is at most `sqrt lambda * ‖theta‖₂`. theorem compiled","shard":"modules/b7ab38cd1a862f90.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius","label":"finiteHorizonScalarConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius","description":"The standard scalar-regularized OFUL radius `noise radius + sqrt lambda * S`.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html#decl-f68b0beaa208","parent":"module:BanditRLProof.OFULScalarRegularizationBias","order":6348,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScalarRegularizationBias"],["Source","BanditRLProof/OFULScalarRegularizationBias.lean:161"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonScalarConfidenceRadius [Fintype Feature] [DecidableEq Feature] (feature : Nat -> Omega -> Feature -> Real) (R delta lambda S : Real) (n : Nat) (omega : Omega) : Real","missing":[],"search":"finitehorizonscalarconfidenceradius banditrlproof.oful.finitehorizonscalarconfidenceradius the standard scalar-regularized oful radius `noise radius + sqrt lambda * s`. definition compiled","shard":"modules/b7ab38cd1a862f90.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizon_scalarRegularizationBias_le","label":"finiteHorizon_scalarRegularizationBias_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizon_scalarRegularizationBias_le","description":"The scalar regularization bias is uniformly bounded by `sqrt lambda * S` whenever the true parameter has Euclidean length at most `S`.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html#decl-260773e245d9","parent":"module:BanditRLProof.OFULScalarRegularizationBias","order":6349,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScalarRegularizationBias"],["Source","BanditRLProof/OFULScalarRegularizationBias.lean:173"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizon_scalarRegularizationBias_le [Fintype Feature] [DecidableEq Feature] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (feature : Nat -> Omega -> Feature -> Real) (n : Nat) (omega : Omega) : matrixNorm (Matrix.scalar Feature lambda + finiteHorizonFeatureGram feature n omega) ((Matrix.scalar Feature lambda + finiteHorizonFeatureGram feature n omega)⁻¹.mulVec ((Matrix.scalar Feature lambda).mulVec thetaStar)) <= Real.sqrt lambda * S","missing":[],"search":"finitehorizon_scalarregularizationbias_le banditrlproof.oful.finitehorizon_scalarregularizationbias_le the scalar regularization bias is uniformly bounded by `sqrt lambda * s` whenever the true parameter has euclidean length at most `s`. theorem compiled","shard":"modules/b7ab38cd1a862f90.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizonScalarRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","label":"measure_finiteHorizonScalarRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizonScalarRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","description":"Finite-horizon OFUL confidence ellipsoid with scalar ridge regularization. The explicit bias premise of the general theorem is discharged from `euclideanLength thetaStar <= S`.","url":"../modules/banditrlproof-ofulscalarregularizationbias/index.html#decl-731eea74655e","parent":"module:BanditRLProof.OFULScalarRegularizationBias","order":6350,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScalarRegularizationBias"],["Source","BanditRLProof/OFULScalarRegularizationBias.lean:199"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonScalarRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (…","missing":[],"search":"measure_finitehorizonscalarridgeestimate_error_matrixnorm_gt_confidenceradius_le banditrlproof.oful.measure_finitehorizonscalarridgeestimate_error_matrixnorm_gt_confidenceradius_le finite-horizon oful confidence ellipsoid with scalar ridge regularization. the explicit bias premise of the general theorem is discharged from `euclideanlength thetastar <= s`. theorem compiled","shard":"modules/b7ab38cd1a862f90.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.squareIntegrableFiniteStoppingTime_of_bounded_ae","label":"squareIntegrableFiniteStoppingTime_of_bounded_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.squareIntegrableFiniteStoppingTime_of_bounded_ae","description":"An almost-everywhere deterministic bound supplies the local square-integrable finite-stopping contract. Values outside the support may remain unbounded.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-2ee10187f42d","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6351,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem squareIntegrableFiniteStoppingTime_of_bounded_ae {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (tau : Omega -> WithTop Nat) (htau : IsStoppingTime F tau) (bound : Nat) (htau_le : ∀ᵐ omega ∂mu, tau omega <= (bound : WithTop Nat)) : SquareIntegrableFiniteStoppingTime mu tau","missing":[],"search":"squareintegrablefinitestoppingtime_of_bounded_ae banditrlproof.oful.squareintegrablefinitestoppingtime_of_bounded_ae an almost-everywhere deterministic bound supplies the local square-integrable finite-stopping contract. values outside the support may remain unbounded. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_le_sq_of_bounded_ae","label":"stoppingTimeRoundSecondMoment_le_sq_of_bounded_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_le_sq_of_bounded_ae","description":"The exact round-count second moment obeys the same numerical square bound when the stopping-time bound is only almost everywhere.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-2d2c512f05b1","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6352,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:63"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_le_sq_of_bounded_ae {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (tau : Omega -> WithTop Nat) (hstop : SquareIntegrableFiniteStoppingTime mu tau) (bound : Nat) (htau_le : ∀ᵐ omega ∂mu, tau omega <= (bound : WithTop Nat)) : stoppingTimeRoundSecondMoment mu tau hstop <= (((bound + 1 : Nat) : Real)) ^ 2","missing":[],"search":"stoppingtimeroundsecondmoment_le_sq_of_bounded_ae banditrlproof.oful.stoppingtimeroundsecondmoment_le_sq_of_bounded_ae the exact round-count second moment obeys the same numerical square bound when the stopping-time bound is only almost everywhere. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_reachHorizon_ae_of_spent_reach","label":"budgetExhaustionTime_le_reachHorizon_ae_of_spent_reach","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.budgetExhaustionTime_le_reachHorizon_ae_of_spent_reach","description":"Almost-everywhere resource reach gives an almost-everywhere deterministic bound on the first budget-exhaustion time.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-d3c82f7a2169","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6353,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:110"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem budgetExhaustionTime_le_reachHorizon_ae_of_spent_reach {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) (spent : Nat -> Omega -> Nat) (budget reachHorizon : Nat) (hreach : ∀ᵐ omega ∂mu, budget <= spent reachHorizon omega) : ∀ᵐ omega ∂mu, budgetExhaustionTime spent budget omega <= (reachHorizon : WithTop Nat)","missing":[],"search":"budgetexhaustiontime_le_reachhorizon_ae_of_spent_reach banditrlproof.budget.budgetexhaustiontime_le_reachhorizon_ae_of_spent_reach almost-everywhere resource reach gives an almost-everywhere deterministic bound on the first budget-exhaustion time. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon_ae","label":"squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon_ae","description":"An adapted budget-exhaustion time reached almost surely by a deterministic horizon is square-integrable under every finite measure.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-977b5d6b079c","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6354,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:129"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon_ae {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget reachHorizon : Nat) (hspent : Adapted F spent) (hreach : ∀ᵐ omega ∂mu, budget <= spent reachHorizon omega) : OFUL.SquareIntegrableFiniteStoppingTime mu (budgetExhaustionTime spent budget)","missing":[],"search":"squareintegrablefinitestoppingtime_budgetexhaustiontime_of_reachhorizon_ae banditrlproof.budget.squareintegrablefinitestoppingtime_budgetexhaustiontime_of_reachhorizon_ae an adapted budget-exhaustion time reached almost surely by a deterministic horizon is square-integrable under every finite measure. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon_ae","label":"stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon_ae","description":"The round-count second moment of budget exhaustion reached almost surely by `reachHorizon` is at most `(reachHorizon + 1)^2`.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-1dcc066e27c1","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6355,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:152"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon_ae {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget reachHorizon : Nat) (hspent : Adapted F spent) (hreach : ∀ᵐ omega ∂mu, budget <= spent reachHorizon omega) : let tau := budgetExhaustionTime spent budget let hstop := squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon_ae mu spent budget reachHorizon hspent hreach OFUL.stoppingTimeRoundSecondMoment mu tau hstop <= (((reachHorizon + 1 : Nat) : Real)) ^ 2","missing":[],"search":"stoppingtimeroundsecondmoment_budgetexhaustiontime_le_of_reachhorizon_ae banditrlproof.budget.stoppingtimeroundsecondmoment_budgetexhaustiontime_le_of_reachhorizon_ae the round-count second moment of budget exhaustion reached almost surely by `reachhorizon` is at most `(reachhorizon + 1)^2`. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy_ae","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy_ae","description":"Canonical scheduled OFUL expected pseudo-regret when budget exhaustion occurs by a deterministic reach horizon almost surely under the canonical measure.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-c00e7531bf61","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6356,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:181"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (source : OFUL.CanonicalLi…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_reachhorizonroundssq_add_initialgap_mul_reachhorizonrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime_reachedby_ae banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_reachhorizonroundssq_add_initialgap_mul_reachhorizonrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime_reachedby_ae canonical scheduled oful expected pseudo-regret when budget exhaustion occurs by a deterministic reach horizon almost surely under the canonical measure. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.AlignedWindowPositiveActionCostAE","label":"AlignedWindowPositiveActionCostAE","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Budget.AlignedWindowPositiveActionCostAE","description":"Almost every trajectory selects an action with positive deterministic cost at least once in every aligned half-open block.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-562afa0cb9ec","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6357,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:280"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def AlignedWindowPositiveActionCostAE {K : Nat} (mu : Measure (Nat -> Fin K × Real)) (actionCost : Fin K -> Nat) (window : Nat) : Prop","missing":[],"search":"alignedwindowpositiveactioncostae banditrlproof.budget.alignedwindowpositiveactioncostae almost every trajectory selects an action with positive deterministic cost at least once in every aligned half-open block. definition compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_iff_forall_ae","label":"alignedWindowPositiveActionCostAE_iff_forall_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.alignedWindowPositiveActionCostAE_iff_forall_ae","description":"The all-block a.e. contract is equivalent to proving the selected positive-cost witness almost surely for each natural block separately.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-c5e2a69f1e38","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6358,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:294"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem alignedWindowPositiveActionCostAE_iff_forall_ae {K : Nat} (mu : Measure (Nat -> Fin K × Real)) (actionCost : Fin K -> Nat) (window : Nat) : AlignedWindowPositiveActionCostAE mu actionCost window ↔ forall block, ∀ᵐ trajectory ∂mu, exists s, s ∈ Finset.Ico (block * window) ((block + 1) * window) /\\ 1 <= actionCost (trajectory s).1","missing":[],"search":"alignedwindowpositiveactioncostae_iff_forall_ae banditrlproof.budget.alignedwindowpositiveactioncostae_iff_forall_ae the all-block a.e. contract is equivalent to proving the selected positive-cost witness almost surely for each natural block separately. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.canonicalActionCostProcess_alignedWindowPositive_ae","label":"canonicalActionCostProcess_alignedWindowPositive_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.canonicalActionCostProcess_alignedWindowPositive_ae","description":"One positive selected-action cost in each aligned block makes the whole block sum positive almost surely.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-0814ce120b08","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6359,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:311"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalActionCostProcess_alignedWindowPositive_ae {K : Nat} (mu : Measure (Nat -> Fin K × Real)) (actionCost : Fin K -> Nat) (window : Nat) (hwindow : AlignedWindowPositiveActionCostAE mu actionCost window) : ∀ᵐ trajectory ∂mu, forall block, 1 <= (Finset.Ico (block * window) ((block + 1) * window)).sum (fun s => canonicalActionCostProcess actionCost s trajectory)","missing":[],"search":"canonicalactioncostprocess_alignedwindowpositive_ae banditrlproof.budget.canonicalactioncostprocess_alignedwindowpositive_ae one positive selected-action cost in each aligned block makes the whole block sum positive almost surely. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budget_le_cumulativeActionCost_budget_mul_ae","label":"budget_le_cumulativeActionCost_budget_mul_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.budget_le_cumulativeActionCost_budget_mul_ae","description":"The a.e. aligned selected-action contract reaches cumulative action-cost budget by completed round `budget * window`.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-64bbd0f2687c","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6360,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:342"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem budget_le_cumulativeActionCost_budget_mul_ae {K : Nat} (mu : Measure (Nat -> Fin K × Real)) (actionCost : Fin K -> Nat) (window budget : Nat) (hwindow : AlignedWindowPositiveActionCostAE mu actionCost window) : ∀ᵐ trajectory ∂mu, budget <= cumulativeActionCost actionCost (budget * window) trajectory","missing":[],"search":"budget_le_cumulativeactioncost_budget_mul_ae banditrlproof.budget.budget_le_cumulativeactioncost_budget_mul_ae the a.e. aligned selected-action contract reaches cumulative action-cost budget by completed round `budget * window`. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_aeAlignedWindowPositiveActionCostBudgetExhaustionTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_aeAlignedWindowPositiveActionCostBudgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_aeAlignedWindowPositiveActionCostBudgetExhaustionTime","description":"Canonical scheduled OFUL expected pseudo-regret stopped at deterministic selected-action cost budget exhaustion under an a.e. aligned-window positive cost contract.","url":"../modules/banditrlproof-ofulscheduledaealignedwindowpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-a1cddca4cb4f","parent":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","order":6361,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret.lean:379"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_aeAlignedWindowPositiveActionCostBudgetExhaustionTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (sourc…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetwindowroundssq_add_initialgap_mul_budgetwindowrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_aealignedwindowpositiveactioncostbudgetexhaustiontime banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetwindowroundssq_add_initialgap_mul_budgetwindowrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_aealignedwindowpositiveactioncostbudgetexhaustiontime canonical scheduled oful expected pseudo-regret stopped at deterministic selected-action cost budget exhaustion under an a.e. aligned-window positive cost contract. theorem compiled","shard":"modules/3afa54f7197ac783.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardScalarAllRoundGapBound","label":"telescopingStandardScalarAllRoundGapBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardScalarAllRoundGapBound","description":"Scheduled deterministic cumulative-gap budget including time zero.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-56cb24acdfe3","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6362,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingStandardScalarAllRoundGapBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"telescopingstandardscalarallroundgapbound banditrlproof.oful.telescopingstandardscalarallroundgapbound scheduled deterministic cumulative-gap budget including time zero. definition compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_initialGap_le_ae","label":"telescopingCanonicalHistoryTrajectory_initialGap_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_initialGap_le_ae","description":"The time-zero gap envelope holds on the trajectory generated by the single telescoping-schedule policy.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-7fc1c11a9dd0","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6363,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectory_initialGap_le_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hS : 0 <= S) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (environment : Thompson.HistoryEnvironment (Fin K) Real) (comparator : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, linearValue thetaStar (actionFeature comparator) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0)) <= standardScalarInitialGapBound S L2","missing":[],"search":"telescopingcanonicalhistorytrajectory_initialgap_le_ae banditrlproof.oful.telescopingcanonicalhistorytrajectory_initialgap_le_ae the time-zero gap envelope holds on the trajectory generated by the single telescoping-schedule policy. theorem compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet","label":"telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet","description":"There exists a finite horizon whose complete gap over rounds `0, ..., horizon` exceeds the corresponding scheduled initial-plus-radius-width budget.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-ae421ada92dd","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6364,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:72"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (comparator : Nat -> Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"telescopingcanonicalhistorytrajectoryallhorizonallroundgapstandardviolationset banditrlproof.oful.telescopingcanonicalhistorytrajectoryallhorizonallroundgapstandardviolationset there exists a finite horizon whose complete gap over rounds `0, ..., horizon` exceeds the corresponding scheduled initial-plus-radius-width budget. definition compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_subset_succ_ae","label":"telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_subset_succ_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_subset_succ_ae","description":"Almost surely, every all-round all-horizon violation gives a successor-only all-horizon violation at the same witness horizon.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-3a627d4f7a06","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6365,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:95"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_subset_succ_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (comparator : Nat -> Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∈ telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet lambda thetaStar actionFeature R delta S L2 comparator -> trajectory ∈ telescopingCanonicalHist…","missing":[],"search":"telescopingcanonicalhistorytrajectoryallhorizonallroundgapstandardviolationset_subset_succ_ae banditrlproof.oful.telescopingcanonicalhistorytrajectoryallhorizonallroundgapstandardviolationset_subset_succ_ae almost surely, every all-round all-horizon violation gives a successor-only all-horizon violation at the same witness horizon. theorem compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"measure_telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"The probability of any finite-horizon complete cumulative-gap violation for the one telescoping-schedule policy is at most `delta`.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-154451968e72","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6366,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:136"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScal…","missing":[],"search":"measure_telescopingcanonicalhistorytrajectoryallhorizonallroundgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.measure_telescopingcanonicalhistorytrajectoryallhorizonallroundgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization the probability of any finite-horizon complete cumulative-gap violation for the one telescoping-schedule policy is at most `delta`. theorem compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","label":"telescopingCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","description":"All-horizon violation event for the named fixed-best pseudo-regret.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-c384ba9acc61","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6367,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:187"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"telescopingcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset banditrlproof.oful.telescopingcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset all-horizon violation event for the named fixed-best pseudo-regret. definition compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete one-policy all-horizon high-probability pseudo-regret theorem. The fixed-best pseudo-regret is nonnegative at every finite horizon, and the probability that any horizon exceeds its scheduled standard budget is at most `delta`.","url":"../modules/banditrlproof-ofulscheduledallhorizonallroundgap/index.html#decl-a24463ffc52e","parent":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","order":6368,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonAllRoundGap"],["Source","BanditRLProof/OFULScheduledAllHorizonAllRoundGap.lean:209"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : (forall horizon trajectory, 0 <…","missing":[],"search":"telescopingcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.telescopingcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete one-policy all-horizon high-probability pseudo-regret theorem. the fixed-best pseudo-regret is nonnegative at every finite horizon, and the probability that any horizon exceeds its scheduled standard budget is at most `delta`. theorem compiled","shard":"modules/8fb09f760fbe9d13.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_antitone","label":"allTimeTelescopingDelta_antitone","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingDelta_antitone","description":"The telescoping confidence schedule is antitone in its time index.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-47a8cf4e2754","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6369,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:22"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingDelta_antitone {delta : Real} (hdelta : 0 <= delta) {n T : Nat} (hnT : n <= T) : allTimeTelescopingDelta delta T <= allTimeTelescopingDelta delta n","missing":[],"search":"alltimetelescopingdelta_antitone banditrlproof.oful.alltimetelescopingdelta_antitone the telescoping confidence schedule is antitone in its time index. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_antitone_delta","label":"standardScalarConfidenceRadiusUpper_antitone_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_antitone_delta","description":"For fixed determinant budget, the standard confidence-radius upper bound is antitone in the positive confidence level.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-5901e3069fa7","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6370,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:42"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarConfidenceRadiusUpper_antitone_delta {Feature : Type u} [Fintype Feature] (R lambda S : Real) (T : Nat) (L2 : Real) {deltaSmall deltaLarge : Real} (hdeltaSmall : 0 < deltaSmall) (hdelta : deltaSmall <= deltaLarge) : standardScalarConfidenceRadiusUpper (Feature := Feature) R deltaLarge lambda S T L2 <= standardScalarConfidenceRadiusUpper (Feature := Feature) R deltaSmall lambda S T L2","missing":[],"search":"standardscalarconfidenceradiusupper_antitone_delta banditrlproof.oful.standardscalarconfidenceradiusupper_antitone_delta for fixed determinant budget, the standard confidence-radius upper bound is antitone in the positive confidence level. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper_of_indices","label":"finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper_of_indices","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper_of_indices","description":"Every prefix radius using the varying telescoping schedule is bounded by one standard radius using separate terminal schedule and Gram-matrix budgets.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-0a36361ffce7","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6371,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:119"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper_of_indices {Omega Feature : Type*} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Omega -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (omega : Omega) (n scheduleT gramT : Nat) (hnSchedule : n <= scheduleT) (hnGram : n <= gramT) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < gramT -> dotProduct (feature t omega) (feature t omega) <= L2) : finiteHorizonScalarConfidenceRadius feature R (allTimeTelescopingDelta delta n) lambda S n omega <= standardScalarConfidenceRadiusUpper (Feature := Feature) R (allTimeTelescopingDelta delta scheduleT) lambda S gramT L2","missing":[],"search":"finitehorizonscalarconfidenceradius_telescoping_le_standardupper_of_indices banditrlproof.oful.finitehorizonscalarconfidenceradius_telescoping_le_standardupper_of_indices every prefix radius using the varying telescoping schedule is bounded by one standard radius using separate terminal schedule and gram-matrix budgets. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper","label":"finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper","description":"Every prefix radius using the varying telescoping schedule is bounded by one standard radius using the same terminal index for both budgets.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-801ce468c906","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6372,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:165"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper {Omega Feature : Type*} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (feature : Nat -> Omega -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (omega : Omega) (n T : Nat) (hnT : n <= T) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (feature t omega) (feature t omega) <= L2) : finiteHorizonScalarConfidenceRadius feature R (allTimeTelescopingDelta delta n) lambda S n omega <= standardScalarConfidenceRadiusUpper (Feature := Feature) R (allTimeTelescopingDelta delta T) lambda S T L2","missing":[],"search":"finitehorizonscalarconfidenceradius_telescoping_le_standardupper banditrlproof.oful.finitehorizonscalarconfidenceradius_telescoping_le_standardupper every prefix radius using the varying telescoping schedule is bounded by one standard radius using the same terminal index for both budgets. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardScalarRadiusWidthBound","label":"telescopingStandardScalarRadiusWidthBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardScalarRadiusWidthBound","description":"Standard varying-budget radius-width envelope at a finite horizon.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-ff753aa26369","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6373,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:189"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingStandardScalarRadiusWidthBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"telescopingstandardscalarradiuswidthbound banditrlproof.oful.telescopingstandardscalarradiuswidthbound standard varying-budget radius-width envelope at a finite horizon. definition compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard","label":"canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard","description":"The scheduled successor bonuses up to a fixed horizon obey one deterministic terminal radius-times-width budget.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-6e355afe84eb","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6374,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:201"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (trajectory : Nat -> Fin K × Real) (hwidth : forall t, t < horizon + 1 -> confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) t trajectory) (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory t)) <= 1) : (Finset.range horizon).sum (fun n => 2 * finiteHorizonScalarConfidenceRadius (canonicalHistoryTrajectoryFeature actionFeature) R (allTimeTelescopingDelta delta (n + 1)) la…","missing":[],"search":"canonicalhistorytrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard banditrlproof.oful.canonicalhistorytrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard the scheduled successor bonuses up to a fixed horizon obey one deterministic terminal radius-times-width budget. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featureBound_le_regularization","label":"canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featureBound_le_regularization","description":"The explicit feature normalization `L2 <= lambda` discharges every scheduled selected-width premise.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-a3ffabff832d","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6375,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:355"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (trajectory : Nat -> Fin K × Real) : (Finset.range horizon).sum (fun n => 2 * finiteHorizonScalarConfidenceRadius (canonicalHistoryTrajectoryFeature actionFeature) R (allTimeTelescopingDelta delta (n + 1)) lambda S (n + 1) trajectory * confidenceWidth (finiteHorizonScalarGram lambda (canonicalHistoryTrajectoryFeature actionFeature) (n + 1) trajectory) (actionFeature (Thompso…","missing":[],"search":"canonicalhistorytrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featurebound_le_regularization banditrlproof.oful.canonicalhistorytrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featurebound_le_regularization the explicit feature normalization `l2 <= lambda` discharges every scheduled selected-width premise. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_radius_mul_width_ae","label":"telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_radius_mul_width_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_radius_mul_width_ae","description":"Outside the single all-time confidence failure event, every fixed-horizon successor-gap sum is bounded by its scheduled radius-times-width sum.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-35d141caaf58","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6376,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:395"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_radius_mul_width_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (comparator : Nat -> Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> (Finset.range horizon).sum (fun n => linearValue thetaStar (actionFeature (comparator (n + 1))) - linearValue thetaStar (actionFeature…","missing":[],"search":"telescopingcanonicalhistorytrajectory_sum_range_succ_gap_le_radius_mul_width_ae banditrlproof.oful.telescopingcanonicalhistorytrajectory_sum_range_succ_gap_le_radius_mul_width_ae outside the single all-time confidence failure event, every fixed-horizon successor-gap sum is bounded by its scheduled radius-times-width sum. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_standard_ae_of_featureBound_le_regularization","label":"telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_standard_ae_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_standard_ae_of_featureBound_le_regularization","description":"The normalized feature contract turns the scheduled fixed-horizon gap sum into the deterministic terminal standard budget on the same all-time good event.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-24bec5d2e39b","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6377,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:493"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_standard_ae_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHis…","missing":[],"search":"telescopingcanonicalhistorytrajectory_sum_range_succ_gap_le_standard_ae_of_featurebound_le_regularization banditrlproof.oful.telescopingcanonicalhistorytrajectory_sum_range_succ_gap_le_standard_ae_of_featurebound_le_regularization the normalized feature contract turns the scheduled fixed-horizon gap sum into the deterministic terminal standard budget on the same all-time good event. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet","label":"telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet","description":"There exists a finite horizon whose cumulative successor gap exceeds the corresponding scheduled deterministic terminal budget.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-aaa75f9666c2","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6378,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:541"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (comparator : Nat -> Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"telescopingcanonicalhistorytrajectoryallhorizonsuccgapstandardviolationset banditrlproof.oful.telescopingcanonicalhistorytrajectoryallhorizonsuccgapstandardviolationset there exists a finite horizon whose cumulative successor gap exceeds the corresponding scheduled deterministic terminal budget. definition compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_subset_confidenceFailure_ae","label":"telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_subset_confidenceFailure_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_subset_confidenceFailure_ae","description":"Any all-horizon cumulative successor-gap violation forces the one-policy all-time confidence failure event almost surely.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-ea3efe6e6a24","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6379,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:565"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_subset_confidenceFailure_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∈ telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet lambda thetaStar actionFea…","missing":[],"search":"telescopingcanonicalhistorytrajectoryallhorizonsuccgapstandardviolationset_subset_confidencefailure_ae banditrlproof.oful.telescopingcanonicalhistorytrajectoryallhorizonsuccgapstandardviolationset_subset_confidencefailure_ae any all-horizon cumulative successor-gap violation forces the one-policy all-time confidence failure event almost surely. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"measure_telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"For the one telescoping-schedule generated policy, the probability that any finite-horizon cumulative successor gap exceeds its scheduled deterministic standard budget is at most `delta`.","url":"../modules/banditrlproof-ofulscheduledallhorizoncumulativegap/index.html#decl-66a3840d82cc","parent":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","order":6380,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonCumulativeGap"],["Source","BanditRLProof/OFULScheduledAllHorizonCumulativeGap.lean:630"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRi…","missing":[],"search":"measure_telescopingcanonicalhistorytrajectoryallhorizonsuccgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.measure_telescopingcanonicalhistorytrajectoryallhorizonsuccgapstandardviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization for the one telescoping-schedule generated policy, the probability that any finite-horizon cumulative successor gap exceeds its scheduled deterministic standard budget is at most `delta`. theorem compiled","shard":"modules/46980bd208ebf135.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget","label":"telescopingHighProbabilityRegretLogBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget","description":"The explicit confidence logarithm induced at horizon `T` by the telescoping failure share `delta / ((T+1)(T+2))`.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-ca3548291ecb","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6381,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:22"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingHighProbabilityRegretLogBudget {Feature : Type u} [Fintype Feature] (lambda delta : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"telescopinghighprobabilityregretlogbudget banditrlproof.oful.telescopinghighprobabilityregretlogbudget the explicit confidence logarithm induced at horizon `t` by the telescoping failure share `delta / ((t+1)(t+2))`. definition compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound","label":"telescopingHighProbabilityPseudoRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound","description":"Explicit complete all-round pseudo-regret budget for the telescoping-schedule policy at a finite horizon.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-ca0096228057","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6382,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:35"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingHighProbabilityPseudoRegretBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"telescopinghighprobabilitypseudoregretbound banditrlproof.oful.telescopinghighprobabilitypseudoregretbound explicit complete all-round pseudo-regret budget for the telescoping-schedule policy at a finite horizon. definition compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_eq_outerBudget_div_succ","label":"allTimeTelescopingDelta_eq_outerBudget_div_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.allTimeTelescopingDelta_eq_outerBudget_div_succ","description":"The time-`T` telescoping share is a finite-window confidence parameter with outer budget `delta / (T+2)`.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-6c362421e385","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6383,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:54"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem allTimeTelescopingDelta_eq_outerBudget_div_succ (delta : Real) (horizon : Nat) : allTimeTelescopingDelta delta horizon = (delta / (((horizon + 2 : Nat) : Real))) / (((horizon + 1 : Nat) : Real))","missing":[],"search":"alltimetelescopingdelta_eq_outerbudget_div_succ banditrlproof.oful.alltimetelescopingdelta_eq_outerbudget_div_succ the time-`t` telescoping share is a finite-window confidence parameter with outer budget `delta / (t+2)`. theorem compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardHighProbabilityRegretLogBudget_outerBudget_eq_telescoping","label":"standardHighProbabilityRegretLogBudget_outerBudget_eq_telescoping","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardHighProbabilityRegretLogBudget_outerBudget_eq_telescoping","description":"The finite-window confidence logarithm at outer budget `delta / (T+2)` is the explicit telescoping confidence logarithm.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-303dcb6b83d8","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6384,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:66"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardHighProbabilityRegretLogBudget_outerBudget_eq_telescoping {Feature : Type u} [Fintype Feature] (lambda delta : Real) (hdelta : 0 < delta) (horizon : Nat) (L2 : Real) : standardHighProbabilityRegretLogBudget (Feature := Feature) lambda (delta / (((horizon + 2 : Nat) : Real))) horizon L2 = telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda delta horizon L2","missing":[],"search":"standardhighprobabilityregretlogbudget_outerbudget_eq_telescoping banditrlproof.oful.standardhighprobabilityregretlogbudget_outerbudget_eq_telescoping the finite-window confidence logarithm at outer budget `delta / (t+2)` is the explicit telescoping confidence logarithm. theorem compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardHighProbabilityPseudoRegretBound_outerBudget_eq_telescoping","label":"standardHighProbabilityPseudoRegretBound_outerBudget_eq_telescoping","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardHighProbabilityPseudoRegretBound_outerBudget_eq_telescoping","description":"The finite-window explicit pseudo-regret budget at outer confidence `delta / (T+2)` is exactly the explicit telescoping budget.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-7f243be134cf","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6385,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:90"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardHighProbabilityPseudoRegretBound_outerBudget_eq_telescoping {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (hdelta : 0 < delta) (horizon : Nat) (L2 : Real) : standardHighProbabilityPseudoRegretBound (Feature := Feature) R (delta / (((horizon + 2 : Nat) : Real))) lambda S horizon L2 = telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"standardhighprobabilitypseudoregretbound_outerbudget_eq_telescoping banditrlproof.oful.standardhighprobabilitypseudoregretbound_outerbudget_eq_telescoping the finite-window explicit pseudo-regret budget at outer confidence `delta / (t+2)` is exactly the explicit telescoping budget. theorem compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardScalarAllRoundGapBound_eq_telescopingHighProbabilityPseudoRegretBound","label":"telescopingStandardScalarAllRoundGapBound_eq_telescopingHighProbabilityPseudoRegretBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardScalarAllRoundGapBound_eq_telescopingHighProbabilityPseudoRegretBound","description":"The named scheduled all-round budget is exactly the explicit telescoping pseudo-regret rate at every finite horizon.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-adc6460e9bd1","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6386,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:111"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingStandardScalarAllRoundGapBound_eq_telescopingHighProbabilityPseudoRegretBound {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : telescopingStandardScalarAllRoundGapBound (Feature := Feature) R delta lambda S horizon L2 = telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"telescopingstandardscalarallroundgapbound_eq_telescopinghighprobabilitypseudoregretbound banditrlproof.oful.telescopingstandardscalarallroundgapbound_eq_telescopinghighprobabilitypseudoregretbound the named scheduled all-round budget is exactly the explicit telescoping pseudo-regret rate at every finite horizon. theorem compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet","label":"telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet","description":"Explicit all-horizon violation event for the named fixed-best pseudo-regret under the one telescoping-schedule policy.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-20a11b23dae7","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6387,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:169"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) : Set (Nat -> Prod (Fin K) Real)","missing":[],"search":"telescopingcanonicalexplicithighprobabilitypseudoregretallhorizonviolationset banditrlproof.oful.telescopingcanonicalexplicithighprobabilitypseudoregretallhorizonviolationset explicit all-horizon violation event for the named fixed-best pseudo-regret under the one telescoping-schedule policy. definition compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet_eq_standard","label":"telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet_eq_standard","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet_eq_standard","description":"The explicit violation event is the compiled abstract scheduled violation event.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-9245f17cff20","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6388,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:189"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet_eq_standard {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S L2 : Real) (hL2 : 0 <= L2) (best : Fin K) : telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best = telescopingCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best","missing":[],"search":"telescopingcanonicalexplicithighprobabilitypseudoregretallhorizonviolationset_eq_standard banditrlproof.oful.telescopingcanonicalexplicithighprobabilitypseudoregretallhorizonviolationset_eq_standard the explicit violation event is the compiled abstract scheduled violation event. theorem compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete explicit one-policy all-horizon high-probability pseudo-regret theorem for telescoping-schedule OFUL.","url":"../modules/banditrlproof-ofulscheduledallhorizonhighprobabilityregretrate/index.html#decl-b9038b9a0219","parent":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","order":6389,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledAllHorizonHighProbabilityRegretRate.lean:225"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","linear-contextual"]],"statement":"theorem telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : (forall horizon t…","missing":[],"search":"telescopingcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.telescopingcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete explicit one-policy all-horizon high-probability pseudo-regret theorem for telescoping-schedule oful. theorem compiled","shard":"modules/7a9a85d762279023.json","books":["bandit"],"chapters":["teaching:oful"],"settings":["linear-contextual"]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeRadius","label":"finiteHistoryScheduledScalarRidgeRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeRadius","description":"History-indexed radius using the budget assigned to successor time `n+1`.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-81efdeb932b3","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6390,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:27"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScheduledScalarRidgeRadius {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (lambda S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Real","missing":[],"search":"finitehistoryscheduledscalarridgeradius banditrlproof.oful.finitehistoryscheduledscalarridgeradius history-indexed radius using the budget assigned to successor time `n+1`. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticScore","label":"finiteHistoryScheduledScalarRidgeOptimisticScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticScore","description":"Scalar-ridge score with the successor-time confidence schedule.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-71526332110a","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6391,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:37"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScheduledScalarRidgeOptimisticScore {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : Real","missing":[],"search":"finitehistoryscheduledscalarridgeoptimisticscore banditrlproof.oful.finitehistoryscheduledscalarridgeoptimisticscore scalar-ridge score with the successor-time confidence schedule. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticScore_eq","label":"finiteHistoryScheduledScalarRidgeOptimisticScore_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticScore_eq","description":"theorem finiteHistoryScheduledScalarRidgeOptimisticScore_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : finiteHistoryScheduledScalarRidgeOptimisticScore lambda actionFeature R deltaAt S n history action = fi…","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-da14b4d40c5c","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6392,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:53"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScheduledScalarRidgeOptimisticScore_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : finiteHistoryScheduledScalarRidgeOptimisticScore lambda actionFeature R deltaAt S n history action = finiteHistoryScalarRidgeOptimisticScore lambda actionFeature R (deltaAt (n + 1)) S n history action","missing":[],"search":"finitehistoryscheduledscalarridgeoptimisticscore_eq banditrlproof.oful.finitehistoryscheduledscalarridgeoptimisticscore_eq theorem finitehistoryscheduledscalarridgeoptimisticscore_eq {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (lambda : real) (actionfeature : fin k -> feature -> real) (r : real) (deltaat : nat -> real) (s : real) (n : nat) (history : history.finitepairhistory (fin k) real n) (action : fin k) : finitehistoryscheduledscalarridgeoptimisticscore lambda actionfeature r deltaat s n history action = finitehistoryscalarridgeoptimisticscore lambda actionfeature r (deltaat (n + 1)) s n history action theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScheduledScalarRidgeOptimisticScore","label":"measurable_finiteHistoryScheduledScalarRidgeOptimisticScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryScheduledScalarRidgeOptimisticScore","description":"Every fixed-action scheduled score is measurable in finite history.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-88c5233a1acc","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6393,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:67"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryScheduledScalarRidgeOptimisticScore {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (action : Fin K) : Measurable (fun history : History.FinitePairHistory (Fin K) Real n => finiteHistoryScheduledScalarRidgeOptimisticScore lambda actionFeature R deltaAt S n history action)","missing":[],"search":"measurable_finitehistoryscheduledscalarridgeoptimisticscore banditrlproof.oful.measurable_finitehistoryscheduledscalarridgeoptimisticscore every fixed-action scheduled score is measurable in finite history. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction","label":"finiteHistoryScheduledScalarRidgeOptimisticAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction","description":"Deterministic strict-fold selector for the scheduled scalar-ridge score.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-bd41fae8b467","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6394,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:83"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScheduledScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"finitehistoryscheduledscalarridgeoptimisticaction banditrlproof.oful.finitehistoryscheduledscalarridgeoptimisticaction deterministic strict-fold selector for the scheduled scalar-ridge score. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction_eq","label":"finiteHistoryScheduledScalarRidgeOptimisticAction_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction_eq","description":"theorem finiteHistoryScheduledScalarRidgeOptimisticAction_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : finiteHistoryScheduledScalarRidgeOptimisticAction hK lambda actionFeature R deltaAt S n history = finiteHi…","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-4a6a071fa85e","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6395,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:101"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScheduledScalarRidgeOptimisticAction_eq {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : finiteHistoryScheduledScalarRidgeOptimisticAction hK lambda actionFeature R deltaAt S n history = finiteHistoryScalarRidgeOptimisticAction hK lambda actionFeature R (deltaAt (n + 1)) S n history","missing":[],"search":"finitehistoryscheduledscalarridgeoptimisticaction_eq banditrlproof.oful.finitehistoryscheduledscalarridgeoptimisticaction_eq theorem finitehistoryscheduledscalarridgeoptimisticaction_eq {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (actionfeature : fin k -> feature -> real) (r : real) (deltaat : nat -> real) (s : real) (n : nat) (history : history.finitepairhistory (fin k) real n) : finitehistoryscheduledscalarridgeoptimisticaction hk lambda actionfeature r deltaat s n history = finitehistoryscalarridgeoptimisticaction hk lambda actionfeature r (deltaat (n + 1)) s n history theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction_score_max","label":"finiteHistoryScheduledScalarRidgeOptimisticAction_score_max","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction_score_max","description":"The scheduled selector maximizes its time-indexed score.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-e1c59911f44e","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6396,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:115"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScheduledScalarRidgeOptimisticAction_score_max {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (action : Fin K) : finiteHistoryScheduledScalarRidgeOptimisticScore lambda actionFeature R deltaAt S n history action <= finiteHistoryScheduledScalarRidgeOptimisticScore lambda actionFeature R deltaAt S n history (finiteHistoryScheduledScalarRidgeOptimisticAction hK lambda actionFeature R deltaAt S n history)","missing":[],"search":"finitehistoryscheduledscalarridgeoptimisticaction_score_max banditrlproof.oful.finitehistoryscheduledscalarridgeoptimisticaction_score_max the scheduled selector maximizes its time-indexed score. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAlgorithm","label":"finiteHistoryScheduledScalarRidgeOptimisticAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAlgorithm","description":"One measurable history algorithm using the supplied all-time schedule.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-853d70458b66","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6397,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:140"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScheduledScalarRidgeOptimisticAlgorithm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) : Thompson.HistoryAlgorithm (Fin K) Real","missing":[],"search":"finitehistoryscheduledscalarridgeoptimisticalgorithm banditrlproof.oful.finitehistoryscheduledscalarridgeoptimisticalgorithm one measurable history algorithm using the supplied all-time schedule. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","label":"finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","description":"Scheduled algorithm specialized to the exact telescoping confidence budget.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-56dcc6963ebf","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6398,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:159"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) : Thompson.HistoryAlgorithm (Fin K) Real","missing":[],"search":"finitehistorytelescopingscalarridgeoptimisticalgorithm banditrlproof.oful.finitehistorytelescopingscalarridgeoptimisticalgorithm scheduled algorithm specialized to the exact telescoping confidence budget. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_initialAction","label":"finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_initialAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_initialAction","description":"The telescoping scheduled policy starts from the fixed canonical arm.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-322c684d0d94","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6399,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:171"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_initialAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) : (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S).initialAction = Measure.dirac ⟨0, hK⟩","missing":[],"search":"finitehistorytelescopingscalarridgeoptimisticalgorithm_initialaction banditrlproof.oful.finitehistorytelescopingscalarridgeoptimisticalgorithm_initialaction the telescoping scheduled policy starts from the fixed canonical arm. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidge_historyStepKernel_map_snd","label":"finiteHistoryScheduledScalarRidge_historyStepKernel_map_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidge_historyStepKernel_map_snd","description":"The scheduled history-step reward marginal is the environment feedback law at the scheduled selected action.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-9c90f0969825","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6400,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:186"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryScheduledScalarRidge_historyStepKernel_map_snd {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (Thompson.historyStepKernel (finiteHistoryScheduledScalarRidgeOptimisticAlgorithm hK lambda actionFeature R deltaAt S) environment n).map Prod.snd history = environment.feedback n (history, finiteHistoryScheduledScalarRidgeOptimisticAction hK lambda actionFeature R deltaAt S n history)","missing":[],"search":"finitehistoryscheduledscalarridge_historystepkernel_map_snd banditrlproof.oful.finitehistoryscheduledscalarridge_historystepkernel_map_snd the scheduled history-step reward marginal is the environment feedback law at the scheduled selected action. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeSelectedFeature","label":"finiteHistoryScheduledScalarRidgeSelectedFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeSelectedFeature","description":"Feature selected from the scheduled finite-history ridge state.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-17bbba17922f","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6401,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:229"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryScheduledScalarRidgeSelectedFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Feature -> Real","missing":[],"search":"finitehistoryscheduledscalarridgeselectedfeature banditrlproof.oful.finitehistoryscheduledscalarridgeselectedfeature feature selected from the scheduled finite-history ridge state. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_observedFeature_succ_ae_eq_scheduledSelectedFeature","label":"canonicalHistoryTrajectory_observedFeature_succ_ae_eq_scheduledSelectedFeature","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_observedFeature_succ_ae_eq_scheduledSelectedFeature","description":"The actual canonical successor feature agrees with the scheduled selector.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-833c35504e40","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6402,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:243"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_observedFeature_succ_ae_eq_scheduledSelectedFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScheduledScalarRidgeOptimisticAlgorithm hK lambda actionFeature R deltaAt S) environment, actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory (n + 1)) = finiteHistoryScheduledScalarRidgeSelectedFeature hK lambda actionFeature R deltaAt S n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n)","missing":[],"search":"canonicalhistorytrajectory_observedfeature_succ_ae_eq_scheduledselectedfeature banditrlproof.oful.canonicalhistorytrajectory_observedfeature_succ_ae_eq_scheduledselectedfeature the actual canonical successor feature agrees with the scheduled selector. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScheduledScalarRidgeOptimisticAction","label":"measurable_finiteHistoryScheduledScalarRidgeOptimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryScheduledScalarRidgeOptimisticAction","description":"The scheduled selector is measurable in its finite history.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-e997bacb3517","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6403,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:282"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryScheduledScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (n : Nat) : Measurable (finiteHistoryScheduledScalarRidgeOptimisticAction hK lambda actionFeature R deltaAt S n)","missing":[],"search":"measurable_finitehistoryscheduledscalarridgeoptimisticaction banditrlproof.oful.measurable_finitehistoryscheduledscalarridgeoptimisticaction the scheduled selector is measurable in its finite history. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableFeature","label":"scheduledCanonicalHistoryTrajectoryPredictableFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableFeature","description":"Strict-past scheduled feature process on the canonical trajectory space.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-0829faefa178","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6404,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:307"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def scheduledCanonicalHistoryTrajectoryPredictableFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) : Nat -> (Nat -> Fin K × Real) -> Feature -> Real | 0, _trajectory => actionFeature ⟨0, hK⟩ | n + 1, trajectory => finiteHistoryScheduledScalarRidgeSelectedFeature hK lambda actionFeature R deltaAt S n (Preorder.frestrictLe n trajectory) /-- Every scheduled predictable-feature coordinate is strict-past measurable. -/ theorem scheduledCanonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (t : Nat) (j : Feature) : StronglyMeasurable[ canonic…","missing":[],"search":"scheduledcanonicalhistorytrajectorypredictablefeature banditrlproof.oful.scheduledcanonicalhistorytrajectorypredictablefeature strict-past scheduled feature process on the canonical trajectory space. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","label":"scheduledCanonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","description":"Every scheduled predictable-feature coordinate is strict-past measurable.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-ea7d1576b375","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6405,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:323"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem scheduledCanonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (t : Nat) (j : Feature) : StronglyMeasurable[ canonicalHistoryTrajectoryBeforeFiltration (K := K) t] (fun trajectory => scheduledCanonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R deltaAt S t trajectory j)","missing":[],"search":"scheduledcanonicalhistorytrajectorypredictablefeature_stronglymeasurable banditrlproof.oful.scheduledcanonicalhistorytrajectorypredictablefeature_stronglymeasurable every scheduled predictable-feature coordinate is strict-past measurable. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectory_action_zero_ae_eq_initialArm","label":"scheduledCanonicalHistoryTrajectory_action_zero_ae_eq_initialArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectory_action_zero_ae_eq_initialArm","description":"The scheduled algorithm starts from the same deterministic arm.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-5981bcde4ad8","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6406,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:367"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem scheduledCanonicalHistoryTrajectory_action_zero_ae_eq_initialArm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScheduledScalarRidgeOptimisticAlgorithm hK lambda actionFeature R deltaAt S) environment, Thompson.canonicalHistoryTrajectoryAction trajectory 0 = ⟨0, hK⟩","missing":[],"search":"scheduledcanonicalhistorytrajectory_action_zero_ae_eq_initialarm banditrlproof.oful.scheduledcanonicalhistorytrajectory_action_zero_ae_eq_initialarm the scheduled algorithm starts from the same deterministic arm. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature","label":"canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature","description":"Actual and scheduled predictable features agree at each time.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-868cd1cddfac","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6407,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:419"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (t : Nat) : canonicalHistoryTrajectoryFeature actionFeature t =ᵐ[ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScheduledScalarRidgeOptimisticAlgorithm hK lambda actionFeature R deltaAt S) environment] scheduledCanonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R deltaAt S t","missing":[],"search":"canonicalhistorytrajectoryfeature_ae_eq_scheduledpredictablefeature banditrlproof.oful.canonicalhistorytrajectoryfeature_ae_eq_scheduledpredictablefeature actual and scheduled predictable features agree at each time. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature_all","label":"canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature_all","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature_all","description":"Actual and scheduled predictable features agree simultaneously.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-19c52aed058b","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6408,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:455"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature_all {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryScheduledScalarRidgeOptimisticAlgorithm hK lambda actionFeature R deltaAt S) environment, ∀ t, canonicalHistoryTrajectoryFeature actionFeature t trajectory = scheduledCanonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R deltaAt S t trajectory","missing":[],"search":"canonicalhistorytrajectoryfeature_ae_eq_scheduledpredictablefeature_all banditrlproof.oful.canonicalhistorytrajectoryfeature_ae_eq_scheduledpredictablefeature_all actual and scheduled predictable features agree simultaneously. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableResidual","label":"scheduledCanonicalHistoryTrajectoryPredictableResidual","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableResidual","description":"Reward residual around the scheduled strict-past selected feature.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-3cf95330e949","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6409,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:479"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def scheduledCanonicalHistoryTrajectoryPredictableResidual {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (i : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"scheduledcanonicalhistorytrajectorypredictableresidual banditrlproof.oful.scheduledcanonicalhistorytrajectorypredictableresidual reward residual around the scheduled strict-past selected feature. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","label":"scheduledCanonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","description":"The zero-initialized scheduled residual process is strongly adapted.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-4423a04724b2","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6410,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:494"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem scheduledCanonicalHistoryTrajectoryPredictableResidual_stronglyAdapted {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) : StronglyAdapted (canonicalHistoryTrajectoryBeforeFiltration (K := K)) (fun t trajectory => match t with | 0 => 0 | i + 1 => scheduledCanonicalHistoryTrajectoryPredictableResidual hK lambda thetaStar actionFeature R deltaAt S i trajectory)","missing":[],"search":"scheduledcanonicalhistorytrajectorypredictableresidual_stronglyadapted banditrlproof.oful.scheduledcanonicalhistorytrajectorypredictableresidual_stronglyadapted the zero-initialized scheduled residual process is strongly adapted. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scheduledPredictableFeature_projection_le_finiteActionProjectionBound","label":"scheduledPredictableFeature_projection_le_finiteActionProjectionBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.scheduledPredictableFeature_projection_le_finiteActionProjectionBound","description":"Scheduled selected features obey the same finite-action projection cap.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-1db21cda464f","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6411,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:554"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem scheduledPredictableFeature_projection_le_finiteActionProjectionBound {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (deltaAt : Nat -> Real) (S : Real) (theta : EuclideanSpace Real Feature) (i : Nat) (trajectory : Nat -> Fin K × Real) : |dotProduct (WithLp.ofLp theta) (scheduledCanonicalHistoryTrajectoryPredictableFeature hK lambda actionFeature R deltaAt S i trajectory)| <= finiteActionProjectionBound hK actionFeature theta","missing":[],"search":"scheduledpredictablefeature_projection_le_finiteactionprojectionbound banditrlproof.oful.scheduledpredictablefeature_projection_le_finiteactionprojectionbound scheduled selected features obey the same finite-action projection cap. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.CanonicalScheduledPredictableScalarRidgeResidualLaw","label":"CanonicalScheduledPredictableScalarRidgeResidualLaw","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.CanonicalScheduledPredictableScalarRidgeResidualLaw","description":"All stochastic input needed by scheduled canonical all-time confidence. The field is intentionally all-time and is stated on the one trajectory measure generated by the telescoping-schedule algorithm.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-1794731f552e","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6412,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:597"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure CanonicalScheduledPredictableScalarRidgeResidualLaw {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Prop where","missing":[],"search":"canonicalscheduledpredictablescalarridgeresiduallaw banditrlproof.oful.canonicalscheduledpredictablescalarridgeresiduallaw all stochastic input needed by scheduled canonical all-time confidence. the field is intentionally all-time and is stated on the one trajectory measure generated by the telescoping-schedule algorithm. structure compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalScheduledPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","label":"canonicalScheduledPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalScheduledPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","description":"The kernel-level linear sub-Gaussian environment contract constructs the all-time predictable residual law for the single telescoping scheduled policy.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-4e8bc74009b5","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6413,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:624"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalScheduledPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : CanonicalScheduledPredictableScalarRidgeResidualLaw hK lambda thetaStar actionFeature R delta S environment where","missing":[],"search":"canonicalscheduledpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment banditrlproof.oful.canonicalscheduledpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment the kernel-level linear sub-gaussian environment contract constructs the all-time predictable residual law for the single telescoping scheduled policy. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.mem_allTimeTelescopingScalarRidgeConfidenceFailureSet_iff_of_feature_eq","label":"mem_allTimeTelescopingScalarRidgeConfidenceFailureSet_iff_of_feature_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.mem_allTimeTelescopingScalarRidgeConfidenceFailureSet_iff_of_feature_eq","description":"Pointwise feature equality preserves the countable telescoping event.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-58dcebd37dfc","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6414,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:832"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem mem_allTimeTelescopingScalarRidgeConfidenceFailureSet_iff_of_feature_eq {Omega : Type*} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature feature' : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R delta : Real) (omega : Omega) (hfeature : forall i, feature i omega = feature' i omega) : omega ∈ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S feature response R delta ↔ omega ∈ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S feature' response R delta","missing":[],"search":"mem_alltimetelescopingscalarridgeconfidencefailureset_iff_of_feature_eq banditrlproof.oful.mem_alltimetelescopingscalarridgeconfidencefailureset_iff_of_feature_eq pointwise feature equality preserves the countable telescoping event. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","label":"measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","description":"The scheduled canonical trajectory satisfies all deterministic-horizon scalar-ridge confidence ellipsoids outside one event of probability `delta`.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-8cb95ebeb8ce","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6415,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:868"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalScheduledPredictableScalarRidgeResidualLaw hK lambda thetaStar actionFeature R delta S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment (allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta) <= ENNReal.ofRea…","missing":[],"search":"measure_telescopingcanonicalhistorytrajectory_alltimeconfidencefailureset_le banditrlproof.oful.measure_telescopingcanonicalhistorytrajectory_alltimeconfidencefailureset_le the scheduled canonical trajectory satisfies all deterministic-horizon scalar-ridge confidence ellipsoids outside one event of probability `delta`. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_action_succ_ae_eq_scheduledAction","label":"telescopingCanonicalHistoryTrajectory_action_succ_ae_eq_scheduledAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_action_succ_ae_eq_scheduledAction","description":"The actual successor action agrees with the scheduled strict-fold action.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-8c428cc4ed3c","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6416,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:980"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectory_action_succ_ae_eq_scheduledAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, Thompson.canonicalHistoryTrajectoryAction trajectory (n + 1) = finiteHistoryScheduledScalarRidgeOptimisticAction hK lambda actionFeature R (allTimeTelescopingDelta delta) S n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n)","missing":[],"search":"telescopingcanonicalhistorytrajectory_action_succ_ae_eq_scheduledaction banditrlproof.oful.telescopingcanonicalhistorytrajectory_action_succ_ae_eq_scheduledaction the actual successor action agrees with the scheduled strict-fold action. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","label":"telescopingCanonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","description":"Outside the scheduled fixed-time confidence failure, the actual successor action satisfies the matching time-indexed optimism-gap certificate.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-dc4090ba207b","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6417,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:1020"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (comparator : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∉ scalarRidgeConfidenceFailureAt lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R (allTimeTelescopingDelta delta (n + 1)) (n + 1) -> linearValue thetaStar (actionFeature comparator) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTraje…","missing":[],"search":"telescopingcanonicalhistorytrajectory_action_succ_gap_le_of_not_mem_confidencefailure banditrlproof.oful.telescopingcanonicalhistorytrajectory_action_succ_gap_le_of_not_mem_confidencefailure outside the scheduled fixed-time confidence failure, the actual successor action satisfies the matching time-indexed optimism-gap certificate. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet","label":"telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet","description":"Event that some successor action violates its scheduled optimism bound.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-8825ef570898","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6418,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:1115"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (comparator : Nat -> Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset banditrlproof.oful.telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset event that some successor action violates its scheduled optimism bound. definition compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_subset_confidenceFailure_ae","label":"telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_subset_confidenceFailure_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_subset_confidenceFailure_ae","description":"Any scheduled successor-gap violation forces a confidence failure a.e.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-83784efcef70","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6419,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:1144"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_subset_confidenceFailure_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (comparator : Nat -> Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, trajectory ∈ telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet lambda thetaStar actionFeature R delta S comparator -> trajectory ∈ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta","missing":[],"search":"telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset_subset_confidencefailure_ae banditrlproof.oful.telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset_subset_confidencefailure_ae any scheduled successor-gap violation forces a confidence failure a.e. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le","label":"measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le","description":"Generated one-policy all-time successor-gap tail under the explicit scheduled predictable-residual law.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-24c857f7ba6e","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6420,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:1229"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (comparator : Nat -> Fin K) (source : CanonicalScheduledPredictableScalarRidgeResidualLaw hK lambda thetaStar actionFeature R delta S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment (telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet lambda thetaStar actionFeature R delta S comparator) <= ENNReal.ofReal delta","missing":[],"search":"measure_telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset_le banditrlproof.oful.measure_telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset_le generated one-policy all-time successor-gap tail under the explicit scheduled predictable-residual law. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","label":"measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","description":"Generated one-policy all-time confidence tail obtained directly from the linear sub-Gaussian history-environment contract.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-ec9575f57e21","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6421,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:1279"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment (allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta) <= ENNRe…","missing":[],"search":"measure_telescopingcanonicalhistorytrajectory_alltimeconfidencefailureset_le_of_linearsubgaussianenvironment banditrlproof.oful.measure_telescopingcanonicalhistorytrajectory_alltimeconfidencefailureset_le_of_linearsubgaussianenvironment generated one-policy all-time confidence tail obtained directly from the linear sub-gaussian history-environment contract. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","label":"measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","description":"Generated one-policy all-time successor-gap tail obtained directly from the linear sub-Gaussian history-environment contract.","url":"../modules/banditrlproof-ofulscheduledalltimeconfidence/index.html#decl-2c08f488d87f","parent":"module:BanditRLProof.OFULScheduledAllTimeConfidence","order":6422,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledAllTimeConfidence.lean:1313"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (comparator : Nat -> Fin K) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment (telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet lambda thetaStar actionFeature R delta S comparator) <= ENNReal.ofReal delta","missing":[],"search":"measure_telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset_le_of_linearsubgaussianenvironment banditrlproof.oful.measure_telescopingcanonicalhistorytrajectoryalltimesuccgapviolationset_le_of_linearsubgaussianenvironment generated one-policy all-time successor-gap tail obtained directly from the linear sub-gaussian history-environment contract. theorem compiled","shard":"modules/1d83be5691f4b15d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedIndexSet","label":"blockStartForcedIndexSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedIndexSet","description":"History indices whose successor action is prescribed by the block-start schedule.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-925460fcebae","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6423,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:11"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def blockStartForcedIndexSet (window horizon : Nat) : Finset Nat","missing":[],"search":"blockstartforcedindexset banditrlproof.oful.blockstartforcedindexset history indices whose successor action is prescribed by the block-start schedule. definition compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.card_blockStartForcedIndexSet_le_div_add_one","label":"card_blockStartForcedIndexSet_le_div_add_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.card_blockStartForcedIndexSet_le_div_add_one","description":"Every forced index below `horizon` belongs to the image of the quotient-block range under `block ↦ block * window`, so its cardinality is at most the number of quotient blocks. With `window = 0`, the modulo condition reduces to `n = 0`, and the same conservative bound remains valid.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-90bdd38984e6","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6424,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem card_blockStartForcedIndexSet_le_div_add_one (window horizon : Nat) : (blockStartForcedIndexSet window horizon).card <= horizon / window + 1","missing":[],"search":"card_blockstartforcedindexset_le_div_add_one banditrlproof.oful.card_blockstartforcedindexset_le_div_add_one every forced index below `horizon` belongs to the image of the quotient-block range under `block ↦ block * window`, so its cardinality is at most the number of quotient blocks. with `window = 0`, the modulo condition reduces to `n = 0`, and the same conservative bound remains valid. theorem compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_card_mul","label":"blockStartForcedActionSuccessorPseudoRegret_le_card_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_card_mul","description":"A pointwise forced-arm gap ceiling bounds the forced charge by its cardinality.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-9193c330f2c7","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6425,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:45"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedActionSuccessorPseudoRegret_le_card_mul {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) (window horizon : Nat) (forcedGapBound : Real) (hforcedGap : forall block, linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (forcedAction block)) <= forcedGapBound) : blockStartForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction window horizon <= ((blockStartForcedIndexSet window horizon).card : Real) * forcedGapBound","missing":[],"search":"blockstartforcedactionsuccessorpseudoregret_le_card_mul banditrlproof.oful.blockstartforcedactionsuccessorpseudoregret_le_card_mul a pointwise forced-arm gap ceiling bounds the forced charge by its cardinality. theorem compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul","label":"blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul","description":"The forced charge is at most the number of quotient blocks times a gap ceiling.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-d9191fc45b52","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6426,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:77"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) (window horizon : Nat) (forcedGapBound : Real) (hforcedGapBound : 0 <= forcedGapBound) (hforcedGap : forall block, linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (forcedAction block)) <= forcedGapBound) : blockStartForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction window horizon <= ((horizon / window + 1 : Nat) : Real) * forcedGapBound","missing":[],"search":"blockstartforcedactionsuccessorpseudoregret_le_div_add_one_mul banditrlproof.oful.blockstartforcedactionsuccessorpseudoregret_le_div_add_one_mul the forced charge is at most the number of quotient blocks times a gap ceiling. theorem compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul_two_mul_parameterFeatureBound","label":"blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul_two_mul_parameterFeatureBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul_two_mul_parameterFeatureBound","description":"Linear parameter and arm envelopes instantiate the generic forced-gap ceiling. No separate `0 <= L2` hypothesis is needed here: `Real.sqrt L2` is nonnegative, and the arm bound at `best` already rules out a negative feasible `L2`.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-9dd63c783f71","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6427,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:111"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul_two_mul_parameterFeatureBound {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength thetaStar <= S) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (forcedAction : Nat -> Fin K) (window horizon : Nat) : blockStartForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction window horizon <= ((horizon / window + 1 : Nat) : Real) * (2 * S * Real.sqrt L2)","missing":[],"search":"blockstartforcedactionsuccessorpseudoregret_le_div_add_one_mul_two_mul_parameterfeaturebound banditrlproof.oful.blockstartforcedactionsuccessorpseudoregret_le_div_add_one_mul_two_mul_parameterfeaturebound linear parameter and arm envelopes instantiate the generic forced-gap ceiling. no separate `0 <= l2` hypothesis is needed here: `real.sqrt l2` is nonnegative, and the arm bound at `best` already rules out a negative feasible `l2`. theorem compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","description":"Fully scalar all-horizon violation event for the block-start forced policy.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-085da941c08c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6428,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:142"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (window : Nat) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset fully scalar all-horizon violation event for the block-start forced policy. definition compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","description":"The scalar-budget violation event is contained in the forced-charge event.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-d11272eb01ca","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6429,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:162"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength thetaStar <= S) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (forcedAction : Nat -> Fin K) (window : Nat) (best : Fin K) : blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 window best ⊆ blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 forcedAction window best","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset_subset banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset_subset the scalar-budget violation event is contained in the forced-charge event. theorem compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete all-horizon theorem with the forced-action charge replaced by the scalar quotient-block count and the common linear arm-gap envelope.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedactionchargebound/index.html#decl-c061ee90ab83","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","order":6430,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound"],["Source","BanditRLProof/OFULScheduledBlockStartForcedActionChargeBound.lean:198"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thet…","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_scalarallhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_scalarallhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete all-horizon theorem with the forced-action charge replaced by the scalar quotient-block count and the common linear arm-gap envelope. theorem compiled","shard":"modules/2d0b08cb2da05bd2.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableFeature","label":"deterministicHistoryCanonicalPredictableFeature","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableFeature","description":"Strict-past feature process generated by a measurable deterministic selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-930e87c8d962","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6431,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def deterministicHistoryCanonicalPredictableFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) : Nat -> (Nat -> Fin K × Real) -> Feature -> Real | 0, _trajectory => actionFeature ⟨0, hK⟩ | n + 1, trajectory => actionFeature (selector n (Preorder.frestrictLe n trajectory)) /-- Every coordinate of the deterministic-selector feature is strict-past measurable. -/ theorem deterministicHistoryCanonicalPredictableFeature_stronglyMeasurable {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (t : Nat) (j : Feature) : StronglyMeasura…","missing":[],"search":"deterministichistorycanonicalpredictablefeature banditrlproof.oful.deterministichistorycanonicalpredictablefeature strict-past feature process generated by a measurable deterministic selector. definition compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableFeature_stronglyMeasurable","label":"deterministicHistoryCanonicalPredictableFeature_stronglyMeasurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableFeature_stronglyMeasurable","description":"Every coordinate of the deterministic-selector feature is strict-past measurable.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-897aa901ed83","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6432,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:37"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem deterministicHistoryCanonicalPredictableFeature_stronglyMeasurable {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (t : Nat) (j : Feature) : StronglyMeasurable[ canonicalHistoryTrajectoryBeforeFiltration (K := K) t] (fun trajectory => deterministicHistoryCanonicalPredictableFeature hK actionFeature selector t trajectory j)","missing":[],"search":"deterministichistorycanonicalpredictablefeature_stronglymeasurable banditrlproof.oful.deterministichistorycanonicalpredictablefeature_stronglymeasurable every coordinate of the deterministic-selector feature is strict-past measurable. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonical_action_zero_ae_eq_initialArm","label":"deterministicHistoryCanonical_action_zero_ae_eq_initialArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistoryCanonical_action_zero_ae_eq_initialArm","description":"A deterministic history algorithm starts from its prescribed fixed arm.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-83eb2fb258dc","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6433,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:73"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem deterministicHistoryCanonical_action_zero_ae_eq_initialArm {K : Nat} (hK : 0 < K) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (Thompson.deterministicHistoryAlgorithm (Fin.mk 0 hK) selector hselector) environment, Thompson.canonicalHistoryTrajectoryAction trajectory 0 = Fin.mk 0 hK","missing":[],"search":"deterministichistorycanonical_action_zero_ae_eq_initialarm banditrlproof.oful.deterministichistorycanonical_action_zero_ae_eq_initialarm a deterministic history algorithm starts from its prescribed fixed arm. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature","label":"canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature","description":"Actual and predictable features agree at every fixed time.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-f12751f2ff38","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6434,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:121"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) (t : Nat) : canonicalHistoryTrajectoryFeature actionFeature t =ᵐ[ Thompson.canonicalHistoryTrajectoryMeasure (Thompson.deterministicHistoryAlgorithm (Fin.mk 0 hK) selector hselector) environment] deterministicHistoryCanonicalPredictableFeature hK actionFeature selector t","missing":[],"search":"canonicalhistorytrajectoryfeature_ae_eq_deterministichistorypredictablefeature banditrlproof.oful.canonicalhistorytrajectoryfeature_ae_eq_deterministichistorypredictablefeature actual and predictable features agree at every fixed time. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature_all","label":"canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature_all","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature_all","description":"Actual and predictable features agree simultaneously at all times.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-6aabf1fc65e3","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6435,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:159"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature_all {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (Thompson.deterministicHistoryAlgorithm (Fin.mk 0 hK) selector hselector) environment, forall t, canonicalHistoryTrajectoryFeature actionFeature t trajectory = deterministicHistoryCanonicalPredictableFeature hK actionFeature selector t trajectory","missing":[],"search":"canonicalhistorytrajectoryfeature_ae_eq_deterministichistorypredictablefeature_all banditrlproof.oful.canonicalhistorytrajectoryfeature_ae_eq_deterministichistorypredictablefeature_all actual and predictable features agree simultaneously at all times. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableResidual","label":"deterministicHistoryCanonicalPredictableResidual","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableResidual","description":"Reward residual centered at the deterministic selector's strict-past feature.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-bb266c8b0cfd","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6436,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:184"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def deterministicHistoryCanonicalPredictableResidual {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (i : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"deterministichistorycanonicalpredictableresidual banditrlproof.oful.deterministichistorycanonicalpredictableresidual reward residual centered at the deterministic selector's strict-past feature. definition compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableResidual_stronglyAdapted","label":"deterministicHistoryCanonicalPredictableResidual_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableResidual_stronglyAdapted","description":"The zero-initialized deterministic-selector residual is strongly adapted.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-0e124a664e15","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6437,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:199"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem deterministicHistoryCanonicalPredictableResidual_stronglyAdapted {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) : StronglyAdapted (canonicalHistoryTrajectoryBeforeFiltration (K := K)) (fun t trajectory => match t with | 0 => 0 | i + 1 => deterministicHistoryCanonicalPredictableResidual hK thetaStar actionFeature selector i trajectory)","missing":[],"search":"deterministichistorycanonicalpredictableresidual_stronglyadapted banditrlproof.oful.deterministichistorycanonicalpredictableresidual_stronglyadapted the zero-initialized deterministic-selector residual is strongly adapted. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistoryPredictableFeature_projection_le_finiteActionProjectionBound","label":"deterministicHistoryPredictableFeature_projection_le_finiteActionProjectionBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistoryPredictableFeature_projection_le_finiteActionProjectionBound","description":"Deterministic-selector features obey the finite-action projection cap.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-8855738a1024","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6438,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:258"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem deterministicHistoryPredictableFeature_projection_le_finiteActionProjectionBound {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (actionFeature : Fin K -> Feature -> Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (theta : EuclideanSpace Real Feature) (i : Nat) (trajectory : Nat -> Fin K × Real) : |dotProduct (WithLp.ofLp theta) (deterministicHistoryCanonicalPredictableFeature hK actionFeature selector i trajectory)| <= finiteActionProjectionBound hK actionFeature theta","missing":[],"search":"deterministichistorypredictablefeature_projection_le_finiteactionprojectionbound banditrlproof.oful.deterministichistorypredictablefeature_projection_le_finiteactionprojectionbound deterministic-selector features obey the finite-action projection cap. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.deterministicHistory_historyStepKernel_map_snd","label":"deterministicHistory_historyStepKernel_map_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.deterministicHistory_historyStepKernel_map_snd","description":"The reward marginal of a deterministic history step is its selected feedback law.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-644e3e8f6521","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6439,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:293"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem deterministicHistory_historyStepKernel_map_snd {K : Nat} (initialAction : Fin K) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (Thompson.historyStepKernel (Thompson.deterministicHistoryAlgorithm initialAction selector hselector) environment n).map Prod.snd history = environment.feedback n (history, selector n history)","missing":[],"search":"deterministichistory_historystepkernel_map_snd banditrlproof.oful.deterministichistory_historystepkernel_map_snd the reward marginal of a deterministic history step is its selected feedback law. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw","label":"CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw","description":"Stochastic law required by the deterministic-selector all-time confidence layer.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-0397ef2edd9e","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6440,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:320"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) : Prop where","missing":[],"search":"canonicaldeterministichistorypredictablescalarridgeresiduallaw banditrlproof.oful.canonicaldeterministichistorypredictablescalarridgeresiduallaw stochastic law required by the deterministic-selector all-time confidence layer. structure compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalDeterministicHistoryPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","label":"canonicalDeterministicHistoryPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalDeterministicHistoryPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","description":"The kernel-level linear sub-Gaussian environment law constructs the predictable residual law for any measurable deterministic history selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-840618843a99","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6441,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:348"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalDeterministicHistoryPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw hK thetaStar actionFeature R S selector hselector environment where","missing":[],"search":"canonicaldeterministichistorypredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment banditrlproof.oful.canonicaldeterministichistorypredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment the kernel-level linear sub-gaussian environment law constructs the predictable residual law for any measurable deterministic history selector. definition compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_deterministicHistoryCanonical_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","label":"measure_deterministicHistoryCanonical_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_deterministicHistoryCanonical_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","description":"All-time scalar-ridge confidence for any measurable deterministic history selector, under that selector's own canonical trajectory measure.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-5043c42e1916","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6442,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:545"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_deterministicHistoryCanonical_allTimeTelescopingScalarRidgeConfidenceFailureSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (selector : (n : Nat) -> History.FinitePairHistory (Fin K) Real n -> Fin K) (hselector : forall n, Measurable (selector n)) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw hK thetaStar actionFeature R S selector hselector environment) : Thompson.canonicalHistoryTrajectoryMeasure (Thompson.deterministicHistoryAlgorithm (Fin.mk 0 hK) selector hselector) environment (allTimeTelescopingScalarRidgeConfidenceFa…","missing":[],"search":"measure_deterministichistorycanonical_alltimetelescopingscalarridgeconfidencefailureset_le banditrlproof.oful.measure_deterministichistorycanonical_alltimetelescopingscalarridgeconfidencefailureset_le all-time scalar-ridge confidence for any measurable deterministic history selector, under that selector's own canonical trajectory measure. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalBlockStartForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","label":"canonicalBlockStartForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalBlockStartForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","description":"The linear environment law supplies the residual law for the forced selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-0e385466867c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6443,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:658"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalBlockStartForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw hK thetaStar actionFeature R S (finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window) (measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window) environment","missing":[],"search":"canonicalblockstartforcedpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment banditrlproof.oful.canonicalblockstartforcedpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment the linear environment law supplies the residual law for the forced selector. definition compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","label":"measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","description":"Source-level all-time confidence tail for the block-start forced policy.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-ab7ad6200463","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6444,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:688"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw hK thetaStar actionFeature R S (finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window) (measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window) environment) : Thompson.canonicalHistoryTrajectoryMeas…","missing":[],"search":"measure_blockstartforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le banditrlproof.oful.measure_blockstartforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le source-level all-time confidence tail for the block-start forced policy. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","label":"measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","description":"The concrete block-start forced policy has the same all-time confidence tail under its own canonical measure as any other measurable deterministic selector driven by the same linear sub-Gaussian environment.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-951c2b3fde24","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6445,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:733"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment (allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajector…","missing":[],"search":"measure_blockstartforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le_of_linearsubgaussianenvironment banditrlproof.oful.measure_blockstartforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le_of_linearsubgaussianenvironment the concrete block-start forced policy has the same all-time confidence tail under its own canonical measure as any other measurable deterministic selector driven by the same linear sub-gaussian environment. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","label":"blockStartOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","description":"The nonforced scheduled charge is bounded by the full pathwise width budget.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-c57f1780da0d","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6446,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:767"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (window horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (trajectory : Nat -> Fin K × Real) : blockStartOptimisticSuccessorRadiusWidthCharge lambda actionFeature R delta S window horizon trajectory <= telescopingStandardScalarRadiusWidthBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"blockstartoptimisticsuccessorradiuswidthcharge_le_telescopingstandardscalarradiuswidthbound banditrlproof.oful.blockstartoptimisticsuccessorradiuswidthcharge_le_telescopingstandardscalarradiuswidthbound the nonforced scheduled charge is bounded by the full pathwise width budget. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalHistoryTrajectory_initialGap_le_ae","label":"blockStartForcedCanonicalHistoryTrajectory_initialGap_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalHistoryTrajectory_initialGap_le_ae","description":"The fixed initial arm has the standard deterministic linear-gap envelope.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-791cd5c62b8f","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6447,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:844"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalHistoryTrajectory_initialGap_le_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0)) <= standardScalarInitia…","missing":[],"search":"blockstartforcedcanonicalhistorytrajectory_initialgap_le_ae banditrlproof.oful.blockstartforcedcanonicalhistorytrajectory_initialgap_le_ae the fixed initial arm has the standard deterministic linear-gap envelope. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","description":"On the one all-time confidence event, every finite-horizon pseudo-regret is bounded by the deterministic forced-arm charge plus the explicit telescoping OFUL rate.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-399c5d58892c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6448,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:891"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature…","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregret_le_forcedactioncharge_add_explicitbound_ae banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregret_le_forcedactioncharge_add_explicitbound_ae on the one all-time confidence event, every finite-horizon pseudo-regret is bounded by the deterministic forced-arm charge plus the explicit telescoping oful rate. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","description":"All-horizon violation event for the block-start forced pseudo-regret rate.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-af3173a9db81","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6449,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:952"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (forcedAction : Nat -> Fin K) (window : Nat) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset all-horizon violation event for the block-start forced pseudo-regret rate. definition compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","description":"Every all-horizon forced-policy regret violation is a confidence failure a.e.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-a517c9836f64","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6450,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:973"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionF…","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset_subset_confidencefailure_ae banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset_subset_confidencefailure_ae every all-horizon forced-policy regret violation is a confidence failure a.e. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete all-horizon high-probability pseudo-regret theorem for the block-start forced policy. The ordinary OFUL rate is augmented only by the deterministic forced-arm gap charge.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedalltimeconfidence/index.html#decl-d94003237453","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","order":6451,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledBlockStartForcedAllTimeConfidence.lean:1022"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar…","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete all-horizon high-probability pseudo-regret theorem for the block-start forced policy. the ordinary oful rate is augmented only by the deterministic forced-arm gap charge. theorem compiled","shard":"modules/2a3e8e8910a8b322.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.deterministicHistoryAlgorithm","label":"deterministicHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Thompson.deterministicHistoryAlgorithm","description":"Package a measurable finite-history selector as a deterministic algorithm.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-74e8b799d1ea","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6452,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:28"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def deterministicHistoryAlgorithm {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (initialAction : Action) (selector : (n : Nat) -> History.FinitePairHistory Action Reward n -> Action) (hselector : forall n, Measurable (selector n)) : HistoryAlgorithm Action Reward where","missing":[],"search":"deterministichistoryalgorithm banditrlproof.thompson.deterministichistoryalgorithm package a measurable finite-history selector as a deterministic algorithm. definition compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.deterministicHistoryAlgorithm_policy_apply","label":"deterministicHistoryAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.deterministicHistoryAlgorithm_policy_apply","description":"Every policy section of a deterministic history algorithm is a Dirac law.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-1f4256e6f8e0","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6453,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:41"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem deterministicHistoryAlgorithm_policy_apply {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [MeasurableSpace Reward] (initialAction : Action) (selector : (n : Nat) -> History.FinitePairHistory Action Reward n -> Action) (hselector : forall n, Measurable (selector n)) (n : Nat) (history : History.FinitePairHistory Action Reward n) : (deterministicHistoryAlgorithm initialAction selector hselector).policy n history = Measure.dirac (selector n history)","missing":[],"search":"deterministichistoryalgorithm_policy_apply banditrlproof.thompson.deterministichistoryalgorithm_policy_apply every policy section of a deterministic history algorithm is a dirac law. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectory_action_succ_ae_eq_deterministicHistorySelector","label":"canonicalHistoryTrajectory_action_succ_ae_eq_deterministicHistorySelector","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Thompson.canonicalHistoryTrajectory_action_succ_ae_eq_deterministicHistorySelector","description":"The canonical successor action of a deterministic history algorithm lies on the graph of its measurable selector almost surely.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-400e34a38abd","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6454,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:59"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_deterministicHistorySelector {Action : Type u} {Reward : Type v} [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] [MeasurableEq Action] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (initialAction : Action) (selector : (n : Nat) -> History.FinitePairHistory Action Reward n -> Action) (hselector : forall n, Measurable (selector n)) (environment : HistoryEnvironment Action Reward) (n : Nat) : ∀ᵐ trajectory ∂ canonicalHistoryTrajectoryMeasure (deterministicHistoryAlgorithm initialAction selector hselector) environment, canonicalHistoryTrajectoryAction trajectory (n + 1) = selector n (History.finitePairHistoryOfTrace (canonicalHistoryTrajectoryAction trajectory) (canonicalHistoryTrajectoryReward trajectory) n)","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_deterministichistoryselector banditrlproof.thompson.canonicalhistorytrajectory_action_succ_ae_eq_deterministichistoryselector the canonical successor action of a deterministic history algorithm lies on the graph of its measurable selector almost surely. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","label":"finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","description":"Use the forced block action at divisible history indices and the ordinary telescoping OFUL selector at all other indices.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-4c1e48c9a415","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6455,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:140"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryBlockStartForcedTelescopingScalarRidgeAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"finitehistoryblockstartforcedtelescopingscalarridgeaction banditrlproof.oful.finitehistoryblockstartforcedtelescopingscalarridgeaction use the forced block action at divisible history indices and the ordinary telescoping oful selector at all other indices. definition compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","label":"measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","description":"The block-start forced selector is measurable in its finite history.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-6b3d62f267ce","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6456,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:158"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window n : Nat) : Measurable (finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window n)","missing":[],"search":"measurable_finitehistoryblockstartforcedtelescopingscalarridgeaction banditrlproof.oful.measurable_finitehistoryblockstartforcedtelescopingscalarridgeaction the block-start forced selector is measurable in its finite history. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","label":"finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","description":"The block-start forced selector packaged as a fully specified history algorithm.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-2fc62a811163","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6457,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:181"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) : Thompson.HistoryAlgorithm (Fin K) Real","missing":[],"search":"finitehistoryblockstartforcedtelescopingscalarridgealgorithm banditrlproof.oful.finitehistoryblockstartforcedtelescopingscalarridgealgorithm the block-start forced selector packaged as a fully specified history algorithm. definition compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm_policy_apply","label":"finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm_policy_apply","description":"Every policy section is the Dirac law at the modified selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-8ef15834431f","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6458,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:200"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm_policy_apply {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window).policy n history = Measure.dirac (finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window n history)","missing":[],"search":"finitehistoryblockstartforcedtelescopingscalarridgealgorithm_policy_apply banditrlproof.oful.finitehistoryblockstartforcedtelescopingscalarridgealgorithm_policy_apply every policy section is the dirac law at the modified selector. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_mul_window","label":"finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_mul_window","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_mul_window","description":"At a block-multiple history index, the modified selector uses that block's forced arm.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-f6916c5b189c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6459,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:226"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_mul_window {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window block : Nat) (hwindow : 0 < window) (history : History.FinitePairHistory (Fin K) Real (block * window)) : finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window (block * window) history = forcedAction block","missing":[],"search":"finitehistoryblockstartforcedtelescopingscalarridgeaction_mul_window banditrlproof.oful.finitehistoryblockstartforcedtelescopingscalarridgeaction_mul_window at a block-multiple history index, the modified selector uses that block's forced arm. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","label":"canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","description":"Generated successor actions follow the modified selector almost surely.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-3ab14527b81e","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6460,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:248"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, (trajectory (n + 1)).1 = finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window n (Preorder.frestrictLe n trajectory)","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_finitehistoryblockstartforcedtelescopingscalarridgeaction banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_finitehistoryblockstartforcedtelescopingscalarridgeaction generated successor actions follow the modified selector almost surely. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_blockStart_succ_ae_eq_forcedAction","label":"canonicalHistoryTrajectory_action_blockStart_succ_ae_eq_forcedAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_blockStart_succ_ae_eq_forcedAction","description":"The generated action after a block-start history is the prescribed forced arm.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-b319a15d0fad","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6461,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:283"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_blockStart_succ_ae_eq_forcedAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (hwindow : 0 < window) (environment : Thompson.HistoryEnvironment (Fin K) Real) (block : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, (trajectory (block * window + 1)).1 = forcedAction block","missing":[],"search":"canonicalhistorytrajectory_action_blockstart_succ_ae_eq_forcedaction banditrlproof.oful.canonicalhistorytrajectory_action_blockstart_succ_ae_eq_forcedaction the generated action after a block-start history is the prescribed forced arm. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","label":"alignedWindowPositiveActionCostAE_finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.alignedWindowPositiveActionCostAE_finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","description":"Positive forced-arm costs produce one positive selected action in every aligned block under the modified algorithm's canonical law.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhistoryalgorithm/index.html#decl-b3f4de336bc7","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","order":6462,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHistoryAlgorithm.lean:320"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem alignedWindowPositiveActionCostAE_finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (hwindow : 2 <= window) (environment : Thompson.HistoryEnvironment (Fin K) Real) (actionCost : Fin K -> Nat) (hpositive : forall block, 1 <= actionCost (forcedAction block)) : AlignedWindowPositiveActionCostAE (Thompson.canonicalHistoryTrajectoryMeasure (OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment) actionCost window","missing":[],"search":"alignedwindowpositiveactioncostae_finitehistoryblockstartforcedtelescopingscalarridgealgorithm banditrlproof.budget.alignedwindowpositiveactioncostae_finitehistoryblockstartforcedtelescopingscalarridgealgorithm positive forced-arm costs produce one positive selected action in every aligned block under the modified algorithm's canonical law. theorem compiled","shard":"modules/d489a2ed3f54e800.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure","label":"blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure","description":"Canonical trajectory measure of the member indexed by `horizon` in the horizon-window forced-policy family.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate/index.html#decl-9253dca4258d","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","order":6463,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean:14"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedActionAtHorizon : Nat -> Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) : Measure (Nat -> Fin K × Real)","missing":[],"search":"blockstartforcedhorizonindexedcanonicaltrajectorymeasure banditrlproof.oful.blockstartforcedhorizonindexedcanonicaltrajectorymeasure canonical trajectory measure of the member indexed by `horizon` in the horizon-window forced-policy family. definition compiled","shard":"modules/49e027c27d113924.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure_apply","label":"blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure_apply","description":"Unfold the policy and window selected by the outer horizon index.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate/index.html#decl-e81bb1e4a32c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","order":6464,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure_apply {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedActionAtHorizon : Nat -> Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) : blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure hK lambda actionFeature R delta S forcedActionAtHorizon environment horizon = Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S (forcedActionAtHorizon horizon) horizon) environment","missing":[],"search":"blockstartforcedhorizonindexedcanonicaltrajectorymeasure_apply banditrlproof.oful.blockstartforcedhorizonindexedcanonicaltrajectorymeasure_apply unfold the policy and window selected by the outer horizon index. theorem compiled","shard":"modules/49e027c27d113924.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet","label":"blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet","description":"One-envelope fixed-horizon bad-event family.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate/index.html#decl-c848c697982e","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","order":6465,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean:53"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) : Nat -> Set (Nat -> Fin K × Real)","missing":[],"search":"blockstartforcedhorizonindexedcanonicalstandardhighprobabilitypseudoregretviolationset banditrlproof.oful.blockstartforcedhorizonindexedcanonicalstandardhighprobabilitypseudoregretviolationset one-envelope fixed-horizon bad-event family. definition compiled","shard":"modules/49e027c27d113924.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet_apply","label":"blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet_apply","description":"Unfold the one-envelope bad event selected by the outer horizon index.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate/index.html#decl-0500aa250f26","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","order":6466,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean:68"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet_apply {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) (horizon : Nat) : blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet lambda thetaStar actionFeature R delta S L2 best horizon = blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet lambda thetaStar actionFeature R delta S L2 horizon best","missing":[],"search":"blockstartforcedhorizonindexedcanonicalstandardhighprobabilitypseudoregretviolationset_apply banditrlproof.oful.blockstartforcedhorizonindexedcanonicalstandardhighprobabilitypseudoregretviolationset_apply unfold the one-envelope bad event selected by the outer horizon index. theorem compiled","shard":"modules/49e027c27d113924.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Every positive member of the horizon-indexed forced-policy family has the one-envelope fixed-horizon pseudo-regret tail. Each member has its own measure.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonindexedhighprobabilityregretrate/index.html#decl-c810ec6ca16b","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","order":6467,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate.lean:88"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedActionAtHorizon : Nat -> Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaS…","missing":[],"search":"blockstartforcedhorizonindexedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.blockstartforcedhorizonindexedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization every positive member of the horizon-indexed forced-policy family has the one-envelope fixed-horizon pseudo-regret tail. each member has its own measure. theorem compiled","shard":"modules/49e027c27d113924.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedIndexSet_horizon_self_eq_singleton","label":"blockStartForcedIndexSet_horizon_self_eq_singleton","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedIndexSet_horizon_self_eq_singleton","description":"A positive horizon used as its own block window has one forced successor index.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html#decl-146c6cbc4826","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","order":6468,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean:11"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedIndexSet_horizon_self_eq_singleton {horizon : Nat} (hhorizon : 0 < horizon) : blockStartForcedIndexSet horizon horizon = {0}","missing":[],"search":"blockstartforcedindexset_horizon_self_eq_singleton banditrlproof.oful.blockstartforcedindexset_horizon_self_eq_singleton a positive horizon used as its own block window has one forced successor index. theorem compiled","shard":"modules/fde36a91ea377d79.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_horizon_self","label":"blockStartForcedActionSuccessorPseudoRegret_horizon_self","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_horizon_self","description":"With `window = horizon > 0`, the forced successor charge is its first gap.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html#decl-0c5834762da2","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","order":6469,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean:29"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedActionSuccessorPseudoRegret_horizon_self {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) {horizon : Nat} (hhorizon : 0 < horizon) : blockStartForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction horizon horizon = linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (forcedAction 0))","missing":[],"search":"blockstartforcedactionsuccessorpseudoregret_horizon_self banditrlproof.oful.blockstartforcedactionsuccessorpseudoregret_horizon_self with `window = horizon > 0`, the forced successor charge is its first gap. theorem compiled","shard":"modules/fde36a91ea377d79.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_horizon_self_le_two_mul_parameterFeatureBound","label":"blockStartForcedActionSuccessorPseudoRegret_horizon_self_le_two_mul_parameterFeatureBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_horizon_self_le_two_mul_parameterFeatureBound","description":"The exact horizon-window forced charge is bounded by one linear arm-gap envelope.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html#decl-db2cbc294ea8","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","order":6470,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean:53"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedActionSuccessorPseudoRegret_horizon_self_le_two_mul_parameterFeatureBound {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength thetaStar <= S) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (forcedAction : Nat -> Fin K) {horizon : Nat} (hhorizon : 0 < horizon) : blockStartForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction horizon horizon <= 2 * S * Real.sqrt L2","missing":[],"search":"blockstartforcedactionsuccessorpseudoregret_horizon_self_le_two_mul_parameterfeaturebound banditrlproof.oful.blockstartforcedactionsuccessorpseudoregret_horizon_self_le_two_mul_parameterfeaturebound the exact horizon-window forced charge is bounded by one linear arm-gap envelope. theorem compiled","shard":"modules/fde36a91ea377d79.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet","description":"Fixed-horizon scalar violation event for the policy whose block window is that horizon.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html#decl-6f38f59c9bcd","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","order":6471,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean:80"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (horizon : Nat) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregrethorizonwindowviolationset banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregrethorizonwindowviolationset fixed-horizon scalar violation event for the policy whose block window is that horizon. definition compiled","shard":"modules/fde36a91ea377d79.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet_subset","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet_subset","description":"The one-charge fixed-horizon violation event is contained in the exact forced-charge all-horizon event for the same horizon-window policy.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html#decl-c9a8ce058c6d","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","order":6472,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean:102"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet_subset {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength thetaStar <= S) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (forcedAction : Nat -> Fin K) {horizon : Nat} (hhorizon : 0 < horizon) (best : Fin K) : blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet lambda thetaStar actionFeature R delta S L2 horizon best ⊆ blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 forcedAction horizon best","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregrethorizonwindowviolationset_subset banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregrethorizonwindowviolationset_subset the one-charge fixed-horizon violation event is contained in the exact forced-charge all-horizon event for the same horizon-window policy. theorem compiled","shard":"modules/fde36a91ea377d79.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_horizonWindow_nonneg_and_finiteHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_horizonWindow_nonneg_and_finiteHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_horizonWindow_nonneg_and_finiteHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Fixed-horizon high-probability pseudo-regret theorem for the horizon-indexed block-start policy family. The forced charge is one linear arm-gap envelope.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedhorizonwindowfinitehorizontail/index.html#decl-46a1fcad3d56","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","order":6473,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail"],["Source","BanditRLProof/OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail.lean:137"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_horizonWindow_nonneg_and_finiteHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (horizon : Nat) (hhorizon : 0 < horizon) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLi…","missing":[],"search":"blockstartforcedcanonicalstandardhighprobabilitypseudoregret_horizonwindow_nonneg_and_finitehorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.blockstartforcedcanonicalstandardhighprobabilitypseudoregret_horizonwindow_nonneg_and_finitehorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization fixed-horizon high-probability pseudo-regret theorem for the horizon-indexed block-start policy family. the forced charge is one linear arm-gap envelope. theorem compiled","shard":"modules/fde36a91ea377d79.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAction","label":"finiteHistoryTelescopingScalarRidgeOptimisticAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAction","description":"The action selected by the telescoping-confidence OFUL policy.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-24e2908adcfd","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","order":6474,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryTelescopingScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"finitehistorytelescopingscalarridgeoptimisticaction banditrlproof.oful.finitehistorytelescopingscalarridgeoptimisticaction the action selected by the telescoping-confidence oful policy. definition compiled","shard":"modules/984cbb88235ee835.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_policy_apply","label":"finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_policy_apply","description":"theorem finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_policy_apply {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S).policy n history = Measure.…","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-7ee9f1768023","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","order":6475,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean:46"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_policy_apply {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S).policy n history = Measure.dirac (finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n history)","missing":[],"search":"finitehistorytelescopingscalarridgeoptimisticalgorithm_policy_apply banditrlproof.oful.finitehistorytelescopingscalarridgeoptimisticalgorithm_policy_apply theorem finitehistorytelescopingscalarridgeoptimisticalgorithm_policy_apply {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (actionfeature : fin k -> feature -> real) (r delta s : real) (n : nat) (history : history.finitepairhistory (fin k) real n) : (finitehistorytelescopingscalarridgeoptimisticalgorithm hk lambda actionfeature r delta s).policy n history = measure.dirac (finitehistorytelescopingscalarridgeoptimisticaction hk lambda actionfeature r delta s n history) theorem compiled","shard":"modules/984cbb88235ee835.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction","label":"canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction","description":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingS…","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-b6cf3ccb1116","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","order":6476,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean:71"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment, (trajectory (n + 1)).1 = finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n (Preorder.frestrictLe n trajectory)","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction theorem canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (actionfeature : fin k -> feature -> real) (r delta s : real) (environment : thompson.historyenvironment (fin k) real) (n : nat) : ∀ᵐ trajectory ∂ thompson.canonicalhistorytrajectorymeasure (finitehistorytelescopingscalarridgeoptimisticalgorithm hk lambda actionfeature r delta s) environment, (trajectory (n + 1)).1 = finitehistorytelescopingscalarridgeoptimisticaction hk lambda actionfeature r delta s n (preorder.frestrictle n trajectory) theorem compiled","shard":"modules/984cbb88235ee835.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_of_blockStartTelescopingActionCostPositive","label":"alignedWindowPositiveActionCostAE_of_blockStartTelescopingActionCostPositive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.alignedWindowPositiveActionCostAE_of_blockStartTelescopingActionCostPositive","description":"theorem alignedWindowPositiveActionCostAE_of_blockStartTelescopingActionCostPositive {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (actionCost : Fin K -> Nat) (window : Nat) (hwindow : 2 <= window) (hpositive : forall block (history : History.Finit…","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-003f990b7ff7","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","order":6477,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean:111"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem alignedWindowPositiveActionCostAE_of_blockStartTelescopingActionCostPositive {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (actionCost : Fin K -> Nat) (window : Nat) (hwindow : 2 <= window) (hpositive : forall block (history : History.FinitePairHistory (Fin K) Real (block * window)), 1 <= actionCost (OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S (block * window) history)) : AlignedWindowPositiveActionCostAE (Thompson.canonicalHistoryTrajectoryMeasure (OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment) actionCost window","missing":[],"search":"alignedwindowpositiveactioncostae_of_blockstarttelescopingactioncostpositive banditrlproof.budget.alignedwindowpositiveactioncostae_of_blockstarttelescopingactioncostpositive theorem alignedwindowpositiveactioncostae_of_blockstarttelescopingactioncostpositive {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (actionfeature : fin k -> feature -> real) (r delta s : real) (environment : thompson.historyenvironment (fin k) real) (actioncost : fin k -> nat) (window : nat) (hwindow : 2 <= window) (hpositive : forall block (history : history.finitepairhistory (fin k) real (block * window)), 1 <= actioncost (oful.finitehistorytelescopingscalarridgeoptimisticaction hk lambda actionfeature r delta s (block * window) history)) : alignedwindowpositiveactioncostae (thompson.canonicalhistorytrajectorymeasure (oful.finitehistorytelescopingscalarridgeoptimisticalgorithm hk lambda actionfeature r delta s) environment) actioncost window theorem compiled","shard":"modules/984cbb88235ee835.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_of_blockStartForcedTelescopingAction","label":"alignedWindowPositiveActionCostAE_of_blockStartForcedTelescopingAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.alignedWindowPositiveActionCostAE_of_blockStartForcedTelescopingAction","description":"theorem alignedWindowPositiveActionCostAE_of_blockStartForcedTelescopingAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (actionCost : Fin K -> Nat) (forcedAction : Nat -> Fin K) (window : Nat) (hwindow : 2 <= window) (hforced : forall block (h…","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-73d4a727930f","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","order":6478,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean:153"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem alignedWindowPositiveActionCostAE_of_blockStartForcedTelescopingAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (actionCost : Fin K -> Nat) (forcedAction : Nat -> Fin K) (window : Nat) (hwindow : 2 <= window) (hforced : forall block (history : History.FinitePairHistory (Fin K) Real (block * window)), OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S (block * window) history = forcedAction block) (hpositive : forall block, 1 <= actionCost (forcedAction block)) : AlignedWindowPositiveActionCostAE (Thompson.canonicalHistoryTrajectoryMeasure (OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm hK lambda actionFeature R delta S) environment) actionCost window","missing":[],"search":"alignedwindowpositiveactioncostae_of_blockstartforcedtelescopingaction banditrlproof.budget.alignedwindowpositiveactioncostae_of_blockstartforcedtelescopingaction theorem alignedwindowpositiveactioncostae_of_blockstartforcedtelescopingaction {k : nat} {feature : type u} [fintype feature] [decidableeq feature] (hk : 0 < k) (lambda : real) (actionfeature : fin k -> feature -> real) (r delta s : real) (environment : thompson.historyenvironment (fin k) real) (actioncost : fin k -> nat) (forcedaction : nat -> fin k) (window : nat) (hwindow : 2 <= window) (hforced : forall block (history : history.finitepairhistory (fin k) real (block * window)), oful.finitehistorytelescopingscalarridgeoptimisticaction hk lambda actionfeature r delta s (block * window) history = forcedaction block) (hpositive : forall block, 1 <= actioncost (forcedaction block)) : alignedwindowpositiveactioncostae (thompson.canonicalhistorytrajectorymeasure (oful.finitehistorytelescopingscalarridgeoptimisticalgorithm hk lambda actionfeature r delta s) environment) actioncost window theorem compiled","shard":"modules/984cbb88235ee835.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_blockStartForcedPositiveActionCostBudgetExhaustionTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_blockStartForcedPositiveActionCostBudgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_blockStartForcedPositiveActionCostBudgetExhaustionTime","description":"Canonical scheduled OFUL expected pseudo-regret stopped at action-cost budget exhaustion when every aligned block contains a history-independent forced positive-cost action immediately after its block start.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-888940ecdd91","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","order":6479,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret.lean:194"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_blockStartForcedPositiveActionCostBudgetExhaustionTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (sour…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetwindowroundssq_add_initialgap_mul_budgetwindowrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_blockstartforcedpositiveactioncostbudgetexhaustiontime banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetwindowroundssq_add_initialgap_mul_budgetwindowrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_blockstartforcedpositiveactioncostbudgetexhaustiontime canonical scheduled oful expected pseudo-regret stopped at action-cost budget exhaustion when every aligned block contains a history-independent forced positive-cost action immediately after its block start. theorem compiled","shard":"modules/984cbb88235ee835.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_of_mod_ne_zero","label":"finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_of_mod_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_of_mod_ne_zero","description":"Away from block starts, the modified selector is the telescoping OFUL selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-ba9002021bec","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6480,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:22"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_of_mod_ne_zero {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hnonforced : n % window ≠ 0) : finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window n history = finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n history","missing":[],"search":"finitehistoryblockstartforcedtelescopingscalarridgeaction_eq_of_mod_ne_zero banditrlproof.oful.finitehistoryblockstartforcedtelescopingscalarridgeaction_eq_of_mod_ne_zero away from block starts, the modified selector is the telescoping oful selector. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_forcedAction_of_mod_eq_zero","label":"finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_forcedAction_of_mod_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_forcedAction_of_mod_eq_zero","description":"At a forced index, the modified selector is the prescribed block action.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-71e6967af2d1","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6481,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:41"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_forcedAction_of_mod_eq_zero {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hforced : n % window = 0) : finiteHistoryBlockStartForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction window n history = forcedAction (n / window)","missing":[],"search":"finitehistoryblockstartforcedtelescopingscalarridgeaction_eq_forcedaction_of_mod_eq_zero banditrlproof.oful.finitehistoryblockstartforcedtelescopingscalarridgeaction_eq_forcedaction_of_mod_eq_zero at a forced index, the modified selector is the prescribed block action. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_mod_ne_zero","label":"canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_mod_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_mod_ne_zero","description":"Under the modified policy's own trajectory law, every nonforced successor action agrees almost surely with the ordinary telescoping OFUL selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-f1815e0cde3c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6482,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:62"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_mod_ne_zero {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (hnonforced : n % window ≠ 0) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, (trajectory (n + 1)).1 = finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n (Preorder.frestrictLe n trajectory)","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction_of_mod_ne_zero banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction_of_mod_ne_zero under the modified policy's own trajectory law, every nonforced successor action agrees almost surely with the ordinary telescoping oful selector. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedSuccessorPseudoRegret","label":"blockStartForcedSuccessorPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedSuccessorPseudoRegret","description":"Successor pseudo-regret charged to block-start forced actions.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-202b4a364cb6","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6483,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:95"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedSuccessorPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (window horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"blockstartforcedsuccessorpseudoregret banditrlproof.oful.blockstartforcedsuccessorpseudoregret successor pseudo-regret charged to block-start forced actions. definition compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret","label":"blockStartForcedActionSuccessorPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret","description":"Deterministic successor charge of the prescribed forced actions.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-1c615e74001d","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6484,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:110"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartForcedActionSuccessorPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) (window horizon : Nat) : Real","missing":[],"search":"blockstartforcedactionsuccessorpseudoregret banditrlproof.oful.blockstartforcedactionsuccessorpseudoregret deterministic successor charge of the prescribed forced actions. definition compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorPseudoRegret","label":"blockStartOptimisticSuccessorPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartOptimisticSuccessorPseudoRegret","description":"Successor pseudo-regret charged to nonforced optimistic actions.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-4cec59c1bb79","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6485,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:123"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartOptimisticSuccessorPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (window horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"blockstartoptimisticsuccessorpseudoregret banditrlproof.oful.blockstartoptimisticsuccessorpseudoregret successor pseudo-regret charged to nonforced optimistic actions. definition compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorRadiusWidthCharge","label":"blockStartOptimisticSuccessorRadiusWidthCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartOptimisticSuccessorRadiusWidthCharge","description":"Scheduled radius-width charge over the nonforced successor actions.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-fa2943f92de7","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6486,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:138"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def blockStartOptimisticSuccessorRadiusWidthCharge {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (window horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"blockstartoptimisticsuccessorradiuswidthcharge banditrlproof.oful.blockstartoptimisticsuccessorradiuswidthcharge scheduled radius-width charge over the nonforced successor actions. definition compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_gap_le_of_eq_telescopingAction_of_not_mem_confidenceFailure","label":"canonicalHistoryTrajectory_action_succ_gap_le_of_eq_telescopingAction_of_not_mem_confidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_gap_le_of_eq_telescopingAction_of_not_mem_confidenceFailure","description":"Pathwise one-step optimism transport. The policy-specific input is only the equality between the observed successor action and the telescoping selector.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-96ac0c379f6b","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6487,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:163"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_gap_le_of_eq_telescopingAction_of_not_mem_confidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (trajectory : Nat -> Fin K × Real) (n : Nat) (comparator : Fin K) (haction : Thompson.canonicalHistoryTrajectoryAction trajectory (n + 1) = finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n (History.finitePairHistoryOfTrace (Thompson.canonicalHistoryTrajectoryAction trajectory) (Thompson.canonicalHistoryTrajectoryReward trajectory) n)) (hgood : trajectory ∉ scalarRidgeConfidenceFailureAt lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R (allTimeTelescopingDelta delta (n + 1)) (n +…","missing":[],"search":"canonicalhistorytrajectory_action_succ_gap_le_of_eq_telescopingaction_of_not_mem_confidencefailure banditrlproof.oful.canonicalhistorytrajectory_action_succ_gap_le_of_eq_telescopingaction_of_not_mem_confidencefailure pathwise one-step optimism transport. the policy-specific input is only the equality between the observed successor action and the telescoping selector. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_blockStartForced_add_optimistic","label":"canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_blockStartForced_add_optimistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_blockStartForced_add_optimistic","description":"The complete pseudo-regret through action `horizon` is the initial gap plus the forced and nonforced successor charges. This identity is pathwise and does not require a probability or confidence assumption.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-7e2fa93037ca","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6488,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:259"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_blockStartForced_add_optimistic {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (window horizon : Nat) (trajectory : Nat -> Fin K × Real) : canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory = (linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0))) + blockStartForcedSuccessorPseudoRegret thetaStar actionFeature best window horizon trajectory + blockStartOptimisticSuccessorPseudoRegret thetaStar actionFeature best window horizon trajectory","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_eq_initial_add_blockstartforced_add_optimistic banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_eq_initial_add_blockstartforced_add_optimistic the complete pseudo-regret through action `horizon` is the initial gap plus the forced and nonforced successor charges. this identity is pathwise and does not require a probability or confidence assumption. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","label":"blockStartForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","description":"Under the modified policy, the trajectory-valued forced successor charge is almost surely the deterministic charge of `forcedAction (n / window)`.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-bad93f49f28c","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6489,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:290"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : blockStartForcedSuccessorPseudoRegret thetaStar actionFeature best window horizon =ᵐ[ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment] fun _trajectory => blockStartForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction window horizon","missing":[],"search":"blockstartforcedsuccessorpseudoregret_ae_eq_forcedactioncharge banditrlproof.oful.blockstartforcedsuccessorpseudoregret_ae_eq_forcedactioncharge under the modified policy, the trajectory-valued forced successor charge is almost surely the deterministic charge of `forcedaction (n / window)`. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","label":"blockStartOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.blockStartOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","description":"On the all-time confidence event, the modified policy's nonforced successor pseudo-regret is bounded by its matching scheduled radius-width charge.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-1a4be5a4b9ff","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6490,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:358"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem blockStartOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> blockStartOptimisticSuccessorPseudoRegret thetaStar act…","missing":[],"search":"blockstartoptimisticsuccessorpseudoregret_le_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure banditrlproof.oful.blockstartoptimisticsuccessorpseudoregret_le_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure on the all-time confidence event, the modified policy's nonforced successor pseudo-regret is bounded by its matching scheduled radius-width charge. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_blockStartForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","label":"canonicalStandardHighProbabilityPseudoRegret_le_initial_add_blockStartForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_blockStartForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","description":"Complete finite-horizon pseudo-regret bound for the modified policy on the all-time confidence event. Forced-round regret remains an explicit charge.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-0e1a1aea6cd7","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6491,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:440"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_le_initial_add_blockStartForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> canonicalStandardHi…","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_le_initial_add_blockstartforced_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_le_initial_add_blockstartforced_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure complete finite-horizon pseudo-regret bound for the modified policy on the all-time confidence event. forced-round regret remains an explicit charge. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_forcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","label":"canonicalStandardHighProbabilityPseudoRegret_le_initial_add_forcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_forcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","description":"The modified-policy confidence-event bound with the forced successor charge written directly in terms of the prescribed forced actions.","url":"../modules/banditrlproof-ofulscheduledblockstartforcedpseudoregretdecomposition/index.html#decl-d0ec5cf594f9","parent":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","order":6492,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledBlockStartForcedPseudoRegretDecomposition.lean:488"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_le_initial_add_forcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (window : Nat) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction window) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> canonicalStandard…","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_le_initial_add_forcedactioncharge_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_le_initial_add_forcedactioncharge_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure the modified-policy confidence-event bound with the forced successor charge written directly in terms of the prescribed forced actions. theorem compiled","shard":"modules/1355bf4b8c8863f8.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_nonneg","label":"telescopingHighProbabilityRegretLogBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_nonneg","description":"The explicit telescoping confidence logarithm is nonnegative.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-5ead56fbb1d1","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6493,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityRegretLogBudget_nonneg {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : 0 <= telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda delta horizon L2","missing":[],"search":"telescopinghighprobabilityregretlogbudget_nonneg banditrlproof.oful.telescopinghighprobabilityregretlogbudget_nonneg the explicit telescoping confidence logarithm is nonnegative. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_mono","label":"telescopingHighProbabilityRegretLogBudget_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_mono","description":"The explicit telescoping confidence logarithm is monotone in the horizon.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-741c194a1799","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6494,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:59"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityRegretLogBudget_mono {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (delta : Real) (hdelta : 0 < delta) (L2 : Real) (hL2 : 0 <= L2) {n horizon : Nat} (hn : n <= horizon) : telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda delta n L2 <= telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda delta horizon L2","missing":[],"search":"telescopinghighprobabilityregretlogbudget_mono banditrlproof.oful.telescopinghighprobabilityregretlogbudget_mono the explicit telescoping confidence logarithm is monotone in the horizon. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_nonneg","label":"telescopingHighProbabilityPseudoRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_nonneg","description":"The explicit telescoping pseudo-regret budget is nonnegative.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-ad6d10cb5d6b","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6495,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:110"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretBound_nonneg {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : 0 <= telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"telescopinghighprobabilitypseudoregretbound_nonneg banditrlproof.oful.telescopinghighprobabilitypseudoregretbound_nonneg the explicit telescoping pseudo-regret budget is nonnegative. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_mono","label":"telescopingHighProbabilityPseudoRegretBound_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_mono","description":"The explicit telescoping pseudo-regret budget is monotone in the horizon.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-da4a371a5db8","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6496,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:127"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretBound_mono {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (hS : 0 <= S) (L2 : Real) (hL2 : 0 <= L2) {n horizon : Nat} (hn : n <= horizon) : telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S n L2 <= telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"telescopinghighprobabilitypseudoregretbound_mono banditrlproof.oful.telescopinghighprobabilitypseudoregretbound_mono the explicit telescoping pseudo-regret budget is monotone in the horizon. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_mono","label":"standardScalarAllRoundGapEnvelope_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_mono","description":"The deterministic all-round gap envelope is monotone in the horizon.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-c714056aaa29","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6497,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:206"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarAllRoundGapEnvelope_mono (S : Real) (hS : 0 <= S) (L2 : Real) {n horizon : Nat} (hn : n <= horizon) : standardScalarAllRoundGapEnvelope S n L2 <= standardScalarAllRoundGapEnvelope S horizon L2","missing":[],"search":"standardscalarallroundgapenvelope_mono banditrlproof.oful.standardscalarallroundgapenvelope_mono the deterministic all-round gap envelope is monotone in the horizon. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_envelope","label":"abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_envelope","description":"A bounded stopped pseudo-regret obeys the endpoint gap envelope.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-656142f14f8b","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6498,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:222"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_envelope {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S : Real) (hS : 0 <= S) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (maxHorizon : Nat) (htau_le : forall trajectory, tau trajectory <= (maxHorizon : WithTop Nat)) (trajectory : Nat -> Fin K × Real) : |stoppedValue (fun horizon trajectory => canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory) tau trajectory| <= standardScalarAllRoundGapEnvelope S maxHorizon L2","missing":[],"search":"abs_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_envelope banditrlproof.oful.abs_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_envelope a bounded stopped pseudo-regret obeys the endpoint gap envelope. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_boundedStoppingTime","label":"integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_boundedStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_boundedStoppingTime","description":"A bounded stopped pseudo-regret is integrable under a finite measure.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-3a199f282413","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6499,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:268"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_boundedStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] (mu : Measure (Nat -> Fin K × Real)) [IsFiniteMeasure mu] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S : Real) (hS : 0 <= S) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (maxHorizon : Nat) (htau_le : forall trajectory, tau trajectory <= (maxHorizon : WithTop Nat)) : Integrable (stoppedValue (fun horizon trajectory => canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory) tau) mu","missing":[],"search":"integrable_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_of_boundedstoppingtime banditrlproof.oful.integrable_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_of_boundedstoppingtime a bounded stopped pseudo-regret is integrable under a finite measure. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_endpoint_add_envelope_mul_real_measure","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_endpoint_add_envelope_mul_real_measure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_endpoint_add_envelope_mul_real_measure","description":"Generic stopped expectation assembly: off the stopped violation event, the stopped regret is below the endpoint explicit budget; on the event it is charged by the endpoint deterministic envelope.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-b8f231665195","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6500,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:325"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_endpoint_add_envelope_mul_real_measure {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (mu : Measure (Nat -> Fin K × Real)) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (maxHorizon : Nat) (htau_le : forall trajectory, tau trajectory <= (maxHorizon : WithTop Nat)) : integral mu (sto…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_endpoint_add_envelope_mul_real_measure banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_endpoint_add_envelope_mul_real_measure generic stopped expectation assembly: off the stopped violation event, the stopped regret is below the endpoint explicit budget; on the event it is charged by the endpoint deterministic envelope. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete bounded-stopping-time expected pseudo-regret theorem for the single telescoping-schedule OFUL policy.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregret/index.html#decl-ad6ff9793e6c","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","order":6501,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegret.lean:491"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau :…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete bounded-stopping-time expected pseudo-regret theorem for the single telescoping-schedule oful policy. theorem compiled","shard":"modules/e464efabbb613273.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedConfidenceLog_isBigO_log_succ","label":"telescopingStandardExpectedConfidenceLog_isBigO_log_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardExpectedConfidenceLog_isBigO_log_succ","description":"The additional scheduled confidence logarithm at outer budget `1 / (T + 1)` is `O(log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretasymptotics/index.html#decl-364778fd1ea2","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","order":6502,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingStandardExpectedConfidenceLog_isBigO_log_succ : (fun horizon : Nat => 2 * Real.log ((((horizon + 1 : Nat) : Real) ^ 2) * ((horizon + 2 : Nat) : Real))) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"telescopingstandardexpectedconfidencelog_isbigo_log_succ banditrlproof.oful.telescopingstandardexpectedconfidencelog_isbigo_log_succ the additional scheduled confidence logarithm at outer budget `1 / (t + 1)` is `o(log (t + 1))`. theorem compiled","shard":"modules/2e58e522b3f30601.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedRegretLogBudget_isBigO_log","label":"telescopingStandardExpectedRegretLogBudget_isBigO_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardExpectedRegretLogBudget_isBigO_log","description":"The complete scheduled confidence log budget at outer budget `1 / (T + 1)` is `O(log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretasymptotics/index.html#decl-188116d02560","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","order":6503,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics.lean:81"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingStandardExpectedRegretLogBudget_isBigO_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => telescopingStandardExpectedRegretLogBudget (Feature := Feature) lambda horizon L2) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"telescopingstandardexpectedregretlogbudget_isbigo_log banditrlproof.oful.telescopingstandardexpectedregretlogbudget_isbigo_log the complete scheduled confidence log budget at outer budget `1 / (t + 1)` is `o(log (t + 1))`. theorem compiled","shard":"modules/2e58e522b3f30601.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","label":"telescopingStandardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","description":"The explicit tuned bounded-stopping-time expected pseudo-regret budget is `O(sqrt (T + 1) * log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretasymptotics/index.html#decl-1763028967a6","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","order":6504,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics.lean:101"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingStandardExpectedPseudoRegretBound_isBigO_sqrt_mul_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => telescopingStandardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2) =O[atTop] (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"telescopingstandardexpectedpseudoregretbound_isbigo_sqrt_mul_log banditrlproof.oful.telescopingstandardexpectedpseudoregretbound_isbigo_sqrt_mul_log the explicit tuned bounded-stopping-time expected pseudo-regret budget is `o(sqrt (t + 1) * log (t + 1))`. theorem compiled","shard":"modules/2e58e522b3f30601.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_isBigO_sqrt_mul_log","label":"canonicalTelescopingStandardExpectedStoppedPseudoRegret_isBigO_sqrt_mul_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_isBigO_sqrt_mul_log","description":"For fixed model parameters and any horizon-indexed canonical stopping-time family bounded pointwise by its horizon, the named expected stopped pseudo-regret is `O(sqrt (T + 1) * log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretasymptotics/index.html#decl-513940c93053","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","order":6505,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics.lean:229"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalTelescopingStandardExpectedStoppedPseudoRegret_isBigO_sqrt_mul_log {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : Nat -> (Nat -> Fin K × Real) -> WithTop Nat) (htau : forall maxHorizon, IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (tau maxHorizon)) (ht…","missing":[],"search":"canonicaltelescopingstandardexpectedstoppedpseudoregret_isbigo_sqrt_mul_log banditrlproof.oful.canonicaltelescopingstandardexpectedstoppedpseudoregret_isbigo_sqrt_mul_log for fixed model parameters and any horizon-indexed canonical stopping-time family bounded pointwise by its horizon, the named expected stopped pseudo-regret is `o(sqrt (t + 1) * log (t + 1))`. theorem compiled","shard":"modules/2e58e522b3f30601.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound_isLittleO_natCast_succ","label":"telescopingStandardExpectedPseudoRegretBound_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound_isLittleO_natCast_succ","description":"The explicit tuned bounded-stopping-time expected pseudo-regret budget is `o(T + 1)`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretconsistency/index.html#decl-5a3200c24613","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","order":6506,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretConsistency.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingStandardExpectedPseudoRegretBound_isLittleO_natCast_succ {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => telescopingStandardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","missing":[],"search":"telescopingstandardexpectedpseudoregretbound_islittleo_natcast_succ banditrlproof.oful.telescopingstandardexpectedpseudoregretbound_islittleo_natcast_succ the explicit tuned bounded-stopping-time expected pseudo-regret budget is `o(t + 1)`. theorem compiled","shard":"modules/40002839ba7154da.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_isLittleO_natCast_succ","label":"canonicalTelescopingStandardExpectedStoppedPseudoRegret_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_isLittleO_natCast_succ","description":"For fixed model parameters and a horizon-indexed stopping-time family bounded by its horizon, the named expected stopped pseudo-regret is `o(T + 1)`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretconsistency/index.html#decl-aeffc3df0bba","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","order":6507,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretConsistency.lean:41"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalTelescopingStandardExpectedStoppedPseudoRegret_isLittleO_natCast_succ {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : Nat -> (Nat -> Fin K × Real) -> WithTop Nat) (htau : forall maxHorizon, IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (tau maxHorizon))…","missing":[],"search":"canonicaltelescopingstandardexpectedstoppedpseudoregret_islittleo_natcast_succ banditrlproof.oful.canonicaltelescopingstandardexpectedstoppedpseudoregret_islittleo_natcast_succ for fixed model parameters and a horizon-indexed stopping-time family bounded by its horizon, the named expected stopped pseudo-regret is `o(t + 1)`. theorem compiled","shard":"modules/40002839ba7154da.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret","label":"canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret","description":"The expected stopped pseudo-regret per available round for the horizon-indexed telescoping-schedule policy and stopping-time family.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretconsistency/index.html#decl-2c34a7c87b3b","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","order":6508,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretConsistency.lean:80"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (best : Fin K) (tau : Nat -> (Nat -> Fin K × Real) -> WithTop Nat) (horizon : Nat) : Real","missing":[],"search":"canonicaltelescopingstandardexpectedaveragestoppedpseudoregret banditrlproof.oful.canonicaltelescopingstandardexpectedaveragestoppedpseudoregret the expected stopped pseudo-regret per available round for the horizon-indexed telescoping-schedule policy and stopping-time family. definition compiled","shard":"modules/40002839ba7154da.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret_tendsto_zero","label":"canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret_tendsto_zero","description":"For fixed model parameters and any horizon-indexed canonical stopping-time family bounded pointwise by its horizon, expected stopped pseudo-regret per available round converges to zero.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretconsistency/index.html#decl-d9b3f3416fd9","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","order":6509,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretConsistency.lean:101"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret_tendsto_zero {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : Nat -> (Nat -> Fin K × Real) -> WithTop Nat) (htau : forall maxHorizon, IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (tau maxHorizon)) (ht…","missing":[],"search":"canonicaltelescopingstandardexpectedaveragestoppedpseudoregret_tendsto_zero banditrlproof.oful.canonicaltelescopingstandardexpectedaveragestoppedpseudoregret_tendsto_zero for fixed model parameters and any horizon-indexed canonical stopping-time family bounded pointwise by its horizon, expected stopped pseudo-regret per available round converges to zero. theorem compiled","shard":"modules/40002839ba7154da.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedRegretLogBudget","label":"telescopingStandardExpectedRegretLogBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardExpectedRegretLogBudget","description":"The explicit scheduled confidence logarithm after choosing the outer expected-regret budget `delta_T = 1 / (T + 1)`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-ff49157d5334","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6510,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingStandardExpectedRegretLogBudget {Feature : Type u} [Fintype Feature] (lambda : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"telescopingstandardexpectedregretlogbudget banditrlproof.oful.telescopingstandardexpectedregretlogbudget the explicit scheduled confidence logarithm after choosing the outer expected-regret budget `delta_t = 1 / (t + 1)`. definition compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_standardExpectedRegretDelta","label":"telescopingHighProbabilityRegretLogBudget_standardExpectedRegretDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_standardExpectedRegretDelta","description":"Substituting `delta_T = 1 / (T + 1)` into the scheduled confidence logarithm gives the explicit cubic-in-horizon logarithmic scale.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-9f11ac25795e","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6511,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:37"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityRegretLogBudget_standardExpectedRegretDelta {Feature : Type u} [Fintype Feature] (lambda : Real) (horizon : Nat) (L2 : Real) : telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda (standardExpectedRegretDelta horizon) horizon L2 = telescopingStandardExpectedRegretLogBudget (Feature := Feature) lambda horizon L2","missing":[],"search":"telescopinghighprobabilityregretlogbudget_standardexpectedregretdelta banditrlproof.oful.telescopinghighprobabilityregretlogbudget_standardexpectedregretdelta substituting `delta_t = 1 / (t + 1)` into the scheduled confidence logarithm gives the explicit cubic-in-horizon logarithmic scale. theorem compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound","label":"telescopingStandardExpectedPseudoRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound","description":"Explicit expected pseudo-regret budget for a stopping time bounded by `T` under the single scheduled policy tuned with outer budget `1 / (T + 1)`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-9f500c7423a5","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6512,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:62"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingStandardExpectedPseudoRegretBound {Feature : Type u} [Fintype Feature] (R lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"telescopingstandardexpectedpseudoregretbound banditrlproof.oful.telescopingstandardexpectedpseudoregretbound explicit expected pseudo-regret budget for a stopping time bounded by `t` under the single scheduled policy tuned with outer budget `1 / (t + 1)`. definition compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_add_initial_standardExpectedRegretDelta_eq","label":"telescopingHighProbabilityPseudoRegretBound_add_initial_standardExpectedRegretDelta_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_add_initial_standardExpectedRegretDelta_eq","description":"The endpoint high-probability budget plus the tuned bad-event envelope charge is exactly the explicit scheduled expected pseudo-regret rate.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-aae763903cd6","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6513,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:81"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretBound_add_initial_standardExpectedRegretDelta_eq {Feature : Type u} [Fintype Feature] (R lambda S : Real) (horizon : Nat) (L2 : Real) : telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R (standardExpectedRegretDelta horizon) lambda S horizon L2 + standardScalarInitialGapBound S L2 = telescopingStandardExpectedPseudoRegretBound (Feature := Feature) R lambda S horizon L2","missing":[],"search":"telescopinghighprobabilitypseudoregretbound_add_initial_standardexpectedregretdelta_eq banditrlproof.oful.telescopinghighprobabilitypseudoregretbound_add_initial_standardexpectedregretdelta_eq the endpoint high-probability budget plus the tuned bad-event envelope charge is exactly the explicit scheduled expected pseudo-regret rate. theorem compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_telescopingStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_telescopingStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_telescopingStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete explicit expected pseudo-regret theorem for one bounded stopping time under the horizon-tuned telescoping-schedule OFUL policy.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-bd0bcebbcbff","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6514,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:104"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_telescopingStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_telescopingstandardexpectedbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_telescopingstandardexpectedbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete explicit expected pseudo-regret theorem for one bounded stopping time under the horizon-tuned telescoping-schedule oful policy. theorem compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret","label":"canonicalTelescopingStandardExpectedStoppedPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret","description":"Expected stopped pseudo-regret for the horizon-indexed family of scheduled policies tuned at `delta_T = 1 / (T + 1)`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-b028db5f010f","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6515,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:164"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalTelescopingStandardExpectedStoppedPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R S : Real) (environment : Thompson.HistoryEnvironment (Fin K) Real) (best : Fin K) (tau : Nat -> (Nat -> Fin K × Real) -> WithTop Nat) (maxHorizon : Nat) : Real","missing":[],"search":"canonicaltelescopingstandardexpectedstoppedpseudoregret banditrlproof.oful.canonicaltelescopingstandardexpectedstoppedpseudoregret expected stopped pseudo-regret for the horizon-indexed family of scheduled policies tuned at `delta_t = 1 / (t + 1)`. definition compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_nonneg_and_le","label":"canonicalTelescopingStandardExpectedStoppedPseudoRegret_nonneg_and_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_nonneg_and_le","description":"The named horizon-indexed expected stopped pseudo-regret is nonnegative and bounded pointwise by the explicit tuned scheduled rate.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimeexpectedregretrate/index.html#decl-d96a2a18a466","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","order":6516,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeExpectedRegretRate.lean:192"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalTelescopingStandardExpectedStoppedPseudoRegret_nonneg_and_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : Nat -> (Nat -> Fin K × Real) -> WithTop Nat) (htau : forall maxHorizon, IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (tau maxHorizon)) (htau_le…","missing":[],"search":"canonicaltelescopingstandardexpectedstoppedpseudoregret_nonneg_and_le banditrlproof.oful.canonicaltelescopingstandardexpectedstoppedpseudoregret_nonneg_and_le the named horizon-indexed expected stopped pseudo-regret is nonnegative and bounded pointwise by the explicit tuned scheduled rate. theorem compiled","shard":"modules/3886869feb38791d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryAllRoundFiltration","label":"canonicalHistoryTrajectoryAllRoundFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryAllRoundFiltration","description":"The canonical trajectory filtration through the current round. At level `T` it contains exactly the coordinates `0, ..., T`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-20a1450ff7c1","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6517,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalHistoryTrajectoryAllRoundFiltration {K : Nat} : Filtration Nat (inferInstance : MeasurableSpace (Nat -> Fin K × Real))","missing":[],"search":"canonicalhistorytrajectoryallroundfiltration banditrlproof.oful.canonicalhistorytrajectoryallroundfiltration the canonical trajectory filtration through the current round. at level `t` it contains exactly the coordinates `0, ..., t`. definition compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryAllRoundFiltration_apply","label":"canonicalHistoryTrajectoryAllRoundFiltration_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectoryAllRoundFiltration_apply","description":"theorem canonicalHistoryTrajectoryAllRoundFiltration_apply {K : Nat} (horizon : Nat) : (canonicalHistoryTrajectoryAllRoundFiltration (K := K) horizon : MeasurableSpace (Nat -> Fin K × Real)) = Filtration.piLE (X := fun _ : Nat => Fin K × Real) horizon","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-b023afe5c157","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6518,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:31"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectoryAllRoundFiltration_apply {K : Nat} (horizon : Nat) : (canonicalHistoryTrajectoryAllRoundFiltration (K := K) horizon : MeasurableSpace (Nat -> Fin K × Real)) = Filtration.piLE (X := fun _ : Nat => Fin K × Real) horizon","missing":[],"search":"canonicalhistorytrajectoryallroundfiltration_apply banditrlproof.oful.canonicalhistorytrajectoryallroundfiltration_apply theorem canonicalhistorytrajectoryallroundfiltration_apply {k : nat} (horizon : nat) : (canonicalhistorytrajectoryallroundfiltration (k := k) horizon : measurablespace (nat -> fin k × real)) = filtration.pile (x := fun _ : nat => fin k × real) horizon theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_canonicalHistoryTrajectory_coordinate_allRound","label":"measurable_canonicalHistoryTrajectory_coordinate_allRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_canonicalHistoryTrajectory_coordinate_allRound","description":"Every coordinate `t <= T` is measurable at all-round level `T`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-84c7d2f5ce6e","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6519,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:39"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalHistoryTrajectory_coordinate_allRound {K : Nat} {t horizon : Nat} (ht : t <= horizon) : @Measurable (Nat -> Fin K × Real) (Fin K × Real) (canonicalHistoryTrajectoryAllRoundFiltration (K := K) horizon) inferInstance (fun trajectory => trajectory t)","missing":[],"search":"measurable_canonicalhistorytrajectory_coordinate_allround banditrlproof.oful.measurable_canonicalhistorytrajectory_coordinate_allround every coordinate `t <= t` is measurable at all-round level `t`. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.boundedTrajectoryTime_ne_top","label":"boundedTrajectoryTime_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.boundedTrajectoryTime_ne_top","description":"A trajectory time below a finite deterministic horizon cannot be `top`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-802985d88807","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6520,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:63"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem boundedTrajectoryTime_ne_top {K : Nat} (tau : (Nat -> Fin K × Real) -> WithTop Nat) (maxHorizon : Nat) (htau_le : forall trajectory, tau trajectory <= (maxHorizon : WithTop Nat)) : forall trajectory, tau trajectory ≠ (⊤ : WithTop Nat)","missing":[],"search":"boundedtrajectorytime_ne_top banditrlproof.oful.boundedtrajectorytime_ne_top a trajectory time below a finite deterministic horizon cannot be `top`. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.coe_untopA_boundedTrajectoryTime","label":"coe_untopA_boundedTrajectoryTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.coe_untopA_boundedTrajectoryTime","description":"Under a finite bound, `untopA` recovers the actual stopping-time value.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-cbb14a5c81eb","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6521,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:75"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem coe_untopA_boundedTrajectoryTime {K : Nat} (tau : (Nat -> Fin K × Real) -> WithTop Nat) (maxHorizon : Nat) (htau_le : forall trajectory, tau trajectory <= (maxHorizon : WithTop Nat)) (trajectory : Nat -> Fin K × Real) : ((tau trajectory).untopA : WithTop Nat) = tau trajectory","missing":[],"search":"coe_untopa_boundedtrajectorytime banditrlproof.oful.coe_untopa_boundedtrajectorytime under a finite bound, `untopa` recovers the actual stopping-time value. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_canonicalStandardHighProbabilityPseudoRegret_allRound","label":"measurable_canonicalStandardHighProbabilityPseudoRegret_allRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_canonicalStandardHighProbabilityPseudoRegret_allRound","description":"Complete fixed-best pseudo-regret through horizon `T` is measurable using only trajectory coordinates `0, ..., T`.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-8a1d0238536d","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6522,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:92"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_canonicalStandardHighProbabilityPseudoRegret_allRound {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (horizon : Nat) : @Measurable (Nat -> Fin K × Real) Real (canonicalHistoryTrajectoryAllRoundFiltration (K := K) horizon) inferInstance (canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon)","missing":[],"search":"measurable_canonicalstandardhighprobabilitypseudoregret_allround banditrlproof.oful.measurable_canonicalstandardhighprobabilitypseudoregret_allround complete fixed-best pseudo-regret through horizon `t` is measurable using only trajectory coordinates `0, ..., t`. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_stronglyAdapted_allRound","label":"canonicalStandardHighProbabilityPseudoRegret_stronglyAdapted_allRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_stronglyAdapted_allRound","description":"The complete fixed-best pseudo-regret process is strongly adapted to the canonical all-round trajectory filtration.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-88e60546ef6d","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6523,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:150"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_stronglyAdapted_allRound {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) : StronglyAdapted (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (fun horizon trajectory => canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory)","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_stronglyadapted_allround banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_stronglyadapted_allround the complete fixed-best pseudo-regret process is strongly adapted to the canonical all-round trajectory filtration. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_stronglyAdapted_allRound","label":"telescopingHighProbabilityPseudoRegretBound_stronglyAdapted_allRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_stronglyAdapted_allRound","description":"The deterministic explicit rate process is strongly adapted.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-de8722cb9262","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6524,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:168"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretBound_stronglyAdapted_allRound {K : Nat} {Feature : Type u} [Fintype Feature] (R delta lambda S L2 : Real) : StronglyAdapted (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (fun horizon (_trajectory : Nat -> Fin K × Real) => telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2)","missing":[],"search":"telescopinghighprobabilitypseudoregretbound_stronglyadapted_allround banditrlproof.oful.telescopinghighprobabilitypseudoregretbound_stronglyadapted_allround the deterministic explicit rate process is strongly adapted. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet","label":"telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet","description":"Violation of the explicit scheduled pseudo-regret rate after evaluating both the deterministic budget and pseudo-regret process at a Mathlib stopped value.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-8c69b4b7046d","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6525,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:185"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) (tau : (Nat -> Fin K × Real) -> WithTop Nat) : Set (Nat -> Fin K × Real)","missing":[],"search":"telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset banditrlproof.oful.telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset violation of the explicit scheduled pseudo-regret rate after evaluating both the deterministic budget and pseudo-regret process at a mathlib stopped value. definition compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_subset_allHorizon","label":"telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_subset_allHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_subset_allHorizon","description":"Every stopped-value violation is already an all-horizon violation, with the same trajectory and witness horizon.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-13c9807cbc23","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6526,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:212"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_subset_allHorizon {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) (tau : (Nat -> Fin K × Real) -> WithTop Nat) : telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet lambda thetaStar actionFeature R delta S L2 best tau ⊆ telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best","missing":[],"search":"telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_subset_allhorizon banditrlproof.oful.telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_subset_allhorizon every stopped-value violation is already an all-horizon violation, with the same trajectory and witness horizon. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_boundedStoppingTime","label":"measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_boundedStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_boundedStoppingTime","description":"For a bounded stopping time, the stopped explicit violation event is measurable at the deterministic bound.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-dee51f451849","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6527,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:236"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_boundedStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (maxHorizon : Nat) (htau_le : forall trajectory, tau trajectory <= (maxHorizon : WithTop Nat)) : MeasurableSet[ canonicalHistoryTrajectoryAllRoundFiltration (K := K) maxHorizon] (telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet lambda thetaStar actionFeature R delta S L2 best tau)","missing":[],"search":"measurableset_telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_of_boundedstoppingtime banditrlproof.oful.measurableset_telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_of_boundedstoppingtime for a bounded stopping time, the stopped explicit violation event is measurable at the deterministic bound. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"measure_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"The stopped-value violation event inherits the all-horizon `delta` tail. This pathwise transport does not require the index to be a stopping time.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-2262045bcff2","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6528,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:297"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : (Nat -> Fin K × Real) -> W…","missing":[],"search":"measure_telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.measure_telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization the stopped-value violation event inherits the all-horizon `delta` tail. this pathwise transport does not require the index to be a stopping time. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_boundedStoppingTime_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_boundedStoppingTime_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_boundedStoppingTime_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete bounded-stopping-time high-probability pseudo-regret theorem for the single telescoping-schedule OFUL policy.","url":"../modules/banditrlproof-ofulscheduledboundedstoppingtimehighprobabilityregretrate/index.html#decl-b8347ddf4024","parent":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","order":6529,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate.lean:351"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_boundedStoppingTime_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) (tau : (Na…","missing":[],"search":"telescopingcanonicalstandardhighprobabilitypseudoregret_nonneg_and_boundedstoppingtime_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.telescopingcanonicalstandardhighprobabilitypseudoregret_nonneg_and_boundedstoppingtime_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete bounded-stopping-time high-probability pseudo-regret theorem for the single telescoping-schedule oful policy. theorem compiled","shard":"modules/db54f0cbc6436ea5.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.squareIntegrableFiniteStoppingTime_of_bounded","label":"squareIntegrableFiniteStoppingTime_of_bounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.squareIntegrableFiniteStoppingTime_of_bounded","description":"A stopping time bounded by a deterministic natural horizon has the local square-integrable finite-stopping contract under any finite measure.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html#decl-68483d0c7248","parent":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","order":6530,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean:26"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem squareIntegrableFiniteStoppingTime_of_bounded {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (tau : Omega -> WithTop Nat) (htau : IsStoppingTime F tau) (bound : Nat) (htau_le : forall omega, tau omega <= (bound : WithTop Nat)) : SquareIntegrableFiniteStoppingTime mu tau","missing":[],"search":"squareintegrablefinitestoppingtime_of_bounded banditrlproof.oful.squareintegrablefinitestoppingtime_of_bounded a stopping time bounded by a deterministic natural horizon has the local square-integrable finite-stopping contract under any finite measure. theorem compiled","shard":"modules/be274911e9fa198a.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_le_sq_of_bounded","label":"stoppingTimeRoundSecondMoment_le_sq_of_bounded","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_le_sq_of_bounded","description":"The exact second moment of a deterministically bounded square-integrable stopping time is at most the square of the corresponding round-count bound.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html#decl-aa2b65dccc16","parent":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","order":6531,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean:57"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_le_sq_of_bounded {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (tau : Omega -> WithTop Nat) (hstop : SquareIntegrableFiniteStoppingTime mu tau) (bound : Nat) (htau_le : forall omega, tau omega <= (bound : WithTop Nat)) : stoppingTimeRoundSecondMoment mu tau hstop <= (((bound + 1 : Nat) : Real)) ^ 2","missing":[],"search":"stoppingtimeroundsecondmoment_le_sq_of_bounded banditrlproof.oful.stoppingtimeroundsecondmoment_le_sq_of_bounded the exact second moment of a deterministically bounded square-integrable stopping time is at most the square of the corresponding round-count bound. theorem compiled","shard":"modules/be274911e9fa198a.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_budget_of_spent_budget","label":"budgetExhaustionTime_le_budget_of_spent_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.budgetExhaustionTime_le_budget_of_spent_budget","description":"If the accumulated resource has reached `budget` by index `budget`, its first budget-exhaustion time is pointwise at most `budget`.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html#decl-130bd47c13bb","parent":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","order":6532,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean:104"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem budgetExhaustionTime_le_budget_of_spent_budget {Omega : Type v} (spent : Nat -> Omega -> Nat) (budget : Nat) (hreach : forall omega, budget <= spent budget omega) : forall omega, budgetExhaustionTime spent budget omega <= (budget : WithTop Nat)","missing":[],"search":"budgetexhaustiontime_le_budget_of_spent_budget banditrlproof.budget.budgetexhaustiontime_le_budget_of_spent_budget if the accumulated resource has reached `budget` by index `budget`, its first budget-exhaustion time is pointwise at most `budget`. theorem compiled","shard":"modules/be274911e9fa198a.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime","label":"squareIntegrableFiniteStoppingTime_budgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime","description":"An adapted budget-exhaustion time reached by index `budget` is a square-integrable finite stopping time under every finite measure.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html#decl-637f1f643777","parent":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","order":6533,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean:121"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem squareIntegrableFiniteStoppingTime_budgetExhaustionTime {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget : Nat) (hspent : Adapted F spent) (hreach : forall omega, budget <= spent budget omega) : OFUL.SquareIntegrableFiniteStoppingTime mu (budgetExhaustionTime spent budget)","missing":[],"search":"squareintegrablefinitestoppingtime_budgetexhaustiontime banditrlproof.budget.squareintegrablefinitestoppingtime_budgetexhaustiontime an adapted budget-exhaustion time reached by index `budget` is a square-integrable finite stopping time under every finite measure. theorem compiled","shard":"modules/be274911e9fa198a.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le","label":"stoppingTimeRoundSecondMoment_budgetExhaustionTime_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le","description":"The exact round-count second moment of the reached-by-budget exhaustion time is at most `(budget + 1)^2`.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html#decl-b5a43a04cdd7","parent":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","order":6534,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean:143"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_budgetExhaustionTime_le {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget : Nat) (hspent : Adapted F spent) (hreach : forall omega, budget <= spent budget omega) : let tau := budgetExhaustionTime spent budget let hstop := squareIntegrableFiniteStoppingTime_budgetExhaustionTime mu spent budget hspent hreach OFUL.stoppingTimeRoundSecondMoment mu tau hstop <= (((budget + 1 : Nat) : Real)) ^ 2","missing":[],"search":"stoppingtimeroundsecondmoment_budgetexhaustiontime_le banditrlproof.budget.stoppingtimeroundsecondmoment_budgetexhaustiontime_le the exact round-count second moment of the reached-by-budget exhaustion time is at most `(budget + 1)^2`. theorem compiled","shard":"modules/be274911e9fa198a.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime","description":"Canonical expected pseudo-regret bound for the single telescoping-schedule OFUL policy stopped at a reached-by-budget resource exhaustion time.","url":"../modules/banditrlproof-ofulscheduledbudgetexhaustionexpectedregret/index.html#decl-a0d56cd26bdd","parent":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","order":6535,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledBudgetExhaustionExpectedRegret.lean:171"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (source : OFUL.CanonicalLinearSubgaussianEnvironmen…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime canonical expected pseudo-regret bound for the single telescoping-schedule oful policy stopped at a reached-by-budget resource exhaustion time. theorem compiled","shard":"modules/be274911e9fa198a.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_reachHorizon_of_spent_reach","label":"budgetExhaustionTime_le_reachHorizon_of_spent_reach","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.budgetExhaustionTime_le_reachHorizon_of_spent_reach","description":"If accumulated resource reaches `budget` by `reachHorizon`, its first budget-exhaustion time is pointwise at most `reachHorizon`.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-0b4c64db10a4","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6536,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem budgetExhaustionTime_le_reachHorizon_of_spent_reach {Omega : Type v} (spent : Nat -> Omega -> Nat) (budget reachHorizon : Nat) (hreach : forall omega, budget <= spent reachHorizon omega) : forall omega, budgetExhaustionTime spent budget omega <= (reachHorizon : WithTop Nat)","missing":[],"search":"budgetexhaustiontime_le_reachhorizon_of_spent_reach banditrlproof.budget.budgetexhaustiontime_le_reachhorizon_of_spent_reach if accumulated resource reaches `budget` by `reachhorizon`, its first budget-exhaustion time is pointwise at most `reachhorizon`. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon","label":"squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon","description":"An adapted budget-exhaustion time reached by an arbitrary deterministic horizon is square-integrable under every finite measure.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-8738d62b26c9","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6537,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:42"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget reachHorizon : Nat) (hspent : Adapted F spent) (hreach : forall omega, budget <= spent reachHorizon omega) : OFUL.SquareIntegrableFiniteStoppingTime mu (budgetExhaustionTime spent budget)","missing":[],"search":"squareintegrablefinitestoppingtime_budgetexhaustiontime_of_reachhorizon banditrlproof.budget.squareintegrablefinitestoppingtime_budgetexhaustiontime_of_reachhorizon an adapted budget-exhaustion time reached by an arbitrary deterministic horizon is square-integrable under every finite measure. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon","label":"stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon","description":"The round-count second moment of a budget-exhaustion time reached by `reachHorizon` is at most `(reachHorizon + 1)^2`.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-eb2e09fa81c1","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6538,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:64"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget reachHorizon : Nat) (hspent : Adapted F spent) (hreach : forall omega, budget <= spent reachHorizon omega) : let tau := budgetExhaustionTime spent budget let hstop := squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon mu spent budget reachHorizon hspent hreach OFUL.stoppingTimeRoundSecondMoment mu tau hstop <= (((reachHorizon + 1 : Nat) : Real)) ^ 2","missing":[],"search":"stoppingtimeroundsecondmoment_budgetexhaustiontime_le_of_reachhorizon banditrlproof.budget.stoppingtimeroundsecondmoment_budgetexhaustiontime_le_of_reachhorizon the round-count second moment of a budget-exhaustion time reached by `reachhorizon` is at most `(reachhorizon + 1)^2`. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy","description":"Canonical scheduled OFUL expected pseudo-regret when budget exhaustion is known to occur by a separately supplied deterministic reach horizon.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-a8589de9a763","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6539,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:93"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (source : OFUL.CanonicalLinea…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_reachhorizonroundssq_add_initialgap_mul_reachhorizonrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime_reachedby banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_reachhorizonroundssq_add_initialgap_mul_reachhorizonrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime_reachedby canonical scheduled oful expected pseudo-regret when budget exhaustion is known to occur by a separately supplied deterministic reach horizon. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.block_le_cumulativeSpent_mul_of_alignedWindowPositive","label":"block_le_cumulativeSpent_mul_of_alignedWindowPositive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.block_le_cumulativeSpent_mul_of_alignedWindowPositive","description":"If every aligned block of length `window` costs at least one, cumulative spend after `block * window` completed rounds is at least `block`.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-b8fd53d8bee8","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6540,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:187"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem block_le_cumulativeSpent_mul_of_alignedWindowPositive {Omega : Type v} (cost : Nat -> Omega -> Nat) (window : Nat) (haligned : forall block omega, 1 <= (Finset.Ico (block * window) ((block + 1) * window)).sum (fun s => cost s omega)) : forall block omega, block <= cumulativeSpent cost (block * window) omega","missing":[],"search":"block_le_cumulativespent_mul_of_alignedwindowpositive banditrlproof.budget.block_le_cumulativespent_mul_of_alignedwindowpositive if every aligned block of length `window` costs at least one, cumulative spend after `block * window` completed rounds is at least `block`. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budget_le_cumulativeSpent_budget_mul_of_alignedWindowPositive","label":"budget_le_cumulativeSpent_budget_mul_of_alignedWindowPositive","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.budget_le_cumulativeSpent_budget_mul_of_alignedWindowPositive","description":"Aligned-window positivity reaches resource threshold `budget` by completed round `budget * window`.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-069ac6456816","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6541,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:235"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem budget_le_cumulativeSpent_budget_mul_of_alignedWindowPositive {Omega : Type v} (cost : Nat -> Omega -> Nat) (window budget : Nat) (haligned : forall block omega, 1 <= (Finset.Ico (block * window) ((block + 1) * window)).sum (fun s => cost s omega)) : forall omega, budget <= cumulativeSpent cost (budget * window) omega","missing":[],"search":"budget_le_cumulativespent_budget_mul_of_alignedwindowpositive banditrlproof.budget.budget_le_cumulativespent_budget_mul_of_alignedwindowpositive aligned-window positivity reaches resource threshold `budget` by completed round `budget * window`. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativeAlignedWindowPositiveCostBudgetExhaustionTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativeAlignedWindowPositiveCostBudgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativeAlignedWindowPositiveCostBudgetExhaustionTime","description":"Canonical scheduled OFUL expected pseudo-regret under aligned-window positive costs. Individual rounds may have zero cost; the deterministic reach horizon is `budget * window`.","url":"../modules/banditrlproof-ofulscheduledcumulativealignedwindowpositivecostbudgetexhaustionexpectedregret/index.html#decl-89146dcc0d56","parent":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","order":6542,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret.lean:253"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativeAlignedWindowPositiveCostBudgetExhaustionTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (sou…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetwindowroundssq_add_initialgap_mul_budgetwindowrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_cumulativealignedwindowpositivecostbudgetexhaustiontime banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetwindowroundssq_add_initialgap_mul_budgetwindowrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_cumulativealignedwindowpositivecostbudgetexhaustiontime canonical scheduled oful expected pseudo-regret under aligned-window positive costs. individual rounds may have zero cost; the deterministic reach horizon is `budget * window`. theorem compiled","shard":"modules/813bbf66cd6c8851.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeSpent","label":"cumulativeSpent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeSpent","description":"Resource spent after `t` completed rounds, using the half-open index set `{0, ..., t - 1}`.","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html#decl-f15df1bc6f50","parent":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","order":6543,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def cumulativeSpent {Omega : Type v} (cost : Nat -> Omega -> Nat) (t : Nat) (omega : Omega) : Nat","missing":[],"search":"cumulativespent banditrlproof.budget.cumulativespent resource spent after `t` completed rounds, using the half-open index set `{0, ..., t - 1}`. definition compiled","shard":"modules/4ac1df2117868a14.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeSpent_zero","label":"cumulativeSpent_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeSpent_zero","description":"theorem cumulativeSpent_zero {Omega : Type v} (cost : Nat -> Omega -> Nat) (omega : Omega) : cumulativeSpent cost 0 omega = 0","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html#decl-8e1cde0260b4","parent":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","order":6544,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSpent_zero {Omega : Type v} (cost : Nat -> Omega -> Nat) (omega : Omega) : cumulativeSpent cost 0 omega = 0","missing":[],"search":"cumulativespent_zero banditrlproof.budget.cumulativespent_zero theorem cumulativespent_zero {omega : type v} (cost : nat -> omega -> nat) (omega : omega) : cumulativespent cost 0 omega = 0 theorem compiled","shard":"modules/4ac1df2117868a14.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeSpent_succ","label":"cumulativeSpent_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeSpent_succ","description":"theorem cumulativeSpent_succ {Omega : Type v} (cost : Nat -> Omega -> Nat) (t : Nat) (omega : Omega) : cumulativeSpent cost (t + 1) omega = cumulativeSpent cost t omega + cost t omega","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html#decl-0f83dedd5420","parent":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","order":6545,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean:40"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSpent_succ {Omega : Type v} (cost : Nat -> Omega -> Nat) (t : Nat) (omega : Omega) : cumulativeSpent cost (t + 1) omega = cumulativeSpent cost t omega + cost t omega","missing":[],"search":"cumulativespent_succ banditrlproof.budget.cumulativespent_succ theorem cumulativespent_succ {omega : type v} (cost : nat -> omega -> nat) (t : nat) (omega : omega) : cumulativespent cost (t + 1) omega = cumulativespent cost t omega + cost t omega theorem compiled","shard":"modules/4ac1df2117868a14.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.adapted_cumulativeSpent","label":"adapted_cumulativeSpent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.adapted_cumulativeSpent","description":"Finite prefix sums of an adapted per-round Nat-valued cost process remain adapted. A cost observed at `s < t` is promoted from `F s` to `F t`.","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html#decl-af6b36c5b2d0","parent":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","order":6546,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean:53"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem adapted_cumulativeSpent {Omega : Type v} [mOmega : MeasurableSpace Omega] {F : Filtration Nat mOmega} {cost : Nat -> Omega -> Nat} (hcost : Adapted F cost) : Adapted F (cumulativeSpent cost)","missing":[],"search":"adapted_cumulativespent banditrlproof.budget.adapted_cumulativespent finite prefix sums of an adapted per-round nat-valued cost process remain adapted. a cost observed at `s < t` is promoted from `f s` to `f t`. theorem compiled","shard":"modules/4ac1df2117868a14.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeSpent_unitGrowth_of_one_le","label":"cumulativeSpent_unitGrowth_of_one_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeSpent_unitGrowth_of_one_le","description":"If every completed round costs at least one, cumulative spend grows by at least one at each step.","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html#decl-cdc625d1d797","parent":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","order":6547,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean:69"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSpent_unitGrowth_of_one_le {Omega : Type v} (cost : Nat -> Omega -> Nat) (hcost : forall t omega, 1 <= cost t omega) : forall t omega, cumulativeSpent cost t omega + 1 <= cumulativeSpent cost (t + 1) omega","missing":[],"search":"cumulativespent_unitgrowth_of_one_le banditrlproof.budget.cumulativespent_unitgrowth_of_one_le if every completed round costs at least one, cumulative spend grows by at least one at each step. theorem compiled","shard":"modules/4ac1df2117868a14.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativePositiveCostBudgetExhaustionTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativePositiveCostBudgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativePositiveCostBudgetExhaustionTime","description":"Canonical expected pseudo-regret bound for the single telescoping-schedule OFUL policy stopped when the cumulative positive per-round cost reaches the budget.","url":"../modules/banditrlproof-ofulscheduledcumulativepositivecostbudgetexhaustionexpectedregret/index.html#decl-45d3ced6ab4a","parent":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","order":6548,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret.lean:85"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativePositiveCostBudgetExhaustionTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (source : OFUL.CanonicalLinea…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_cumulativepositivecostbudgetexhaustiontime banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_cumulativepositivecostbudgetexhaustiontime canonical expected pseudo-regret bound for the single telescoping-schedule oful policy stopped when the cumulative positive per-round cost reaches the budget. theorem compiled","shard":"modules/4ac1df2117868a14.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.canonicalActionCostProcess","label":"canonicalActionCostProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Budget.canonicalActionCostProcess","description":"The deterministic cost of the current canonical trajectory action.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-c3813ebece33","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6549,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def canonicalActionCostProcess {K : Nat} (actionCost : Fin K -> Nat) (t : Nat) (trajectory : Nat -> Fin K × Real) : Nat","missing":[],"search":"canonicalactioncostprocess banditrlproof.budget.canonicalactioncostprocess the deterministic cost of the current canonical trajectory action. definition compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.adapted_canonicalActionCostProcess","label":"adapted_canonicalActionCostProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.adapted_canonicalActionCostProcess","description":"The current-action cost is adapted to the canonical all-round filtration: the current trajectory coordinate is measurable, as are all maps between countable measurable spaces.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-09401be0b117","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6550,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:33"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem adapted_canonicalActionCostProcess {K : Nat} (actionCost : Fin K -> Nat) : Adapted (OFUL.canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (canonicalActionCostProcess actionCost)","missing":[],"search":"adapted_canonicalactioncostprocess banditrlproof.budget.adapted_canonicalactioncostprocess the current-action cost is adapted to the canonical all-round filtration: the current trajectory coordinate is measurable, as are all maps between countable measurable spaces. theorem compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.canonicalActionCostProcess_one_le","label":"canonicalActionCostProcess_one_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.canonicalActionCostProcess_one_le","description":"Positive arm costs give a pointwise positive trajectory cost process.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-3e609de054d3","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6551,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:47"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalActionCostProcess_one_le {K : Nat} (actionCost : Fin K -> Nat) (hpositive : forall action, 1 <= actionCost action) : forall t trajectory, 1 <= canonicalActionCostProcess actionCost t trajectory","missing":[],"search":"canonicalactioncostprocess_one_le banditrlproof.budget.canonicalactioncostprocess_one_le positive arm costs give a pointwise positive trajectory cost process. theorem compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeActionCost","label":"cumulativeActionCost","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeActionCost","description":"Action cost spent after `t` completed rounds, using the half-open index set `{0, ..., t - 1}`.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-1a872e823842","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6552,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:59"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def cumulativeActionCost {K : Nat} (actionCost : Fin K -> Nat) (t : Nat) (trajectory : Nat -> Fin K × Real) : Nat","missing":[],"search":"cumulativeactioncost banditrlproof.budget.cumulativeactioncost action cost spent after `t` completed rounds, using the half-open index set `{0, ..., t - 1}`. definition compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeActionCost_apply","label":"cumulativeActionCost_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeActionCost_apply","description":"theorem cumulativeActionCost_apply {K : Nat} (actionCost : Fin K -> Nat) (t : Nat) (trajectory : Nat -> Fin K × Real) : cumulativeActionCost actionCost t trajectory = (Finset.range t).sum (fun s => actionCost (trajectory s).1)","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-d2bf266fd401","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6553,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:67"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem cumulativeActionCost_apply {K : Nat} (actionCost : Fin K -> Nat) (t : Nat) (trajectory : Nat -> Fin K × Real) : cumulativeActionCost actionCost t trajectory = (Finset.range t).sum (fun s => actionCost (trajectory s).1)","missing":[],"search":"cumulativeactioncost_apply banditrlproof.budget.cumulativeactioncost_apply theorem cumulativeactioncost_apply {k : nat} (actioncost : fin k -> nat) (t : nat) (trajectory : nat -> fin k × real) : cumulativeactioncost actioncost t trajectory = (finset.range t).sum (fun s => actioncost (trajectory s).1) theorem compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.adapted_cumulativeActionCost","label":"adapted_cumulativeActionCost","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.adapted_cumulativeActionCost","description":"The half-open cumulative action-cost process remains adapted.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-ac50a1a0435e","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6554,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:77"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem adapted_cumulativeActionCost {K : Nat} (actionCost : Fin K -> Nat) : Adapted (OFUL.canonicalHistoryTrajectoryAllRoundFiltration (K := K)) (cumulativeActionCost actionCost)","missing":[],"search":"adapted_cumulativeactioncost banditrlproof.budget.adapted_cumulativeactioncost the half-open cumulative action-cost process remains adapted. theorem compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.cumulativeActionCost_unitGrowth_of_one_le","label":"cumulativeActionCost_unitGrowth_of_one_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.cumulativeActionCost_unitGrowth_of_one_le","description":"Positive arm costs make cumulative action cost grow by at least one.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-4e1b0e1b325e","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6555,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:88"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem cumulativeActionCost_unitGrowth_of_one_le {K : Nat} (actionCost : Fin K -> Nat) (hpositive : forall action, 1 <= actionCost action) : forall t trajectory, cumulativeActionCost actionCost t trajectory + 1 <= cumulativeActionCost actionCost (t + 1) trajectory","missing":[],"search":"cumulativeactioncost_unitgrowth_of_one_le banditrlproof.budget.cumulativeactioncost_unitgrowth_of_one_le positive arm costs make cumulative action cost grow by at least one. theorem compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_positiveActionCostBudgetExhaustionTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_positiveActionCostBudgetExhaustionTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_positiveActionCostBudgetExhaustionTime","description":"Canonical scheduled OFUL expected pseudo-regret stopped when the half-open cumulative deterministic action cost reaches the budget.","url":"../modules/banditrlproof-ofulscheduledpositiveactioncostbudgetexhaustionexpectedregret/index.html#decl-58d6b863d606","parent":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","order":6556,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret.lean:104"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_positiveActionCostBudgetExhaustionTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (source : OFUL.CanonicalLinearSub…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_positiveactioncostbudgetexhaustiontime banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_positiveactioncostbudgetexhaustiontime canonical scheduled oful expected pseudo-regret stopped when the half-open cumulative deterministic action cost reaches the budget. theorem compiled","shard":"modules/4fb104d9366fa1a9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorRadiusWidthCharge","label":"powerOfTwoOptimisticSuccessorRadiusWidthCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorRadiusWidthCharge","description":"Scheduled radius-width charge over nonforced successor actions.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-2648f5feb6f6","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6557,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoOptimisticSuccessorRadiusWidthCharge {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"poweroftwooptimisticsuccessorradiuswidthcharge banditrlproof.oful.poweroftwooptimisticsuccessorradiuswidthcharge scheduled radius-width charge over nonforced successor actions. definition compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","label":"powerOfTwoOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","description":"On the all-time confidence event, the nonforced successor pseudo-regret is bounded by its matching scheduled radius-width charge.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-25f2dc9363b7","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6558,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:50"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> powerOfTwoOptimisticSuccessorPseudoRegret thetaStar actionFeature best horizo…","missing":[],"search":"poweroftwooptimisticsuccessorpseudoregret_le_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure banditrlproof.oful.poweroftwooptimisticsuccessorpseudoregret_le_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure on the all-time confidence event, the nonforced successor pseudo-regret is bounded by its matching scheduled radius-width charge. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","label":"canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","description":"On the confidence event, complete pseudo-regret is bounded by the initial gap, generated forced charge, and nonforced radius-width charge.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-e3a17def01c9","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6559,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:132"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> canonicalStandardHighProbabilityPseudoReg…","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_le_initial_add_poweroftwoforced_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_le_initial_add_poweroftwoforced_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure on the confidence event, complete pseudo-regret is bounded by the initial gap, generated forced charge, and nonforced radius-width charge. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","label":"canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","description":"The confidence-event finite-horizon bound with the forced charge written directly in terms of the prescribed actions.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-988be8d6c47e","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6560,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:180"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, trajectory ∉ allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature) canonicalHistoryTrajectoryResponse R delta -> canonicalStandardHighProbabil…","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_le_initial_add_poweroftwoforcedactioncharge_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_le_initial_add_poweroftwoforcedactioncharge_add_radiuswidthcharge_ae_of_not_mem_alltimeconfidencefailure the confidence-event finite-horizon bound with the forced charge written directly in terms of the prescribed actions. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalPowerOfTwoForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","label":"canonicalPowerOfTwoForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalPowerOfTwoForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","description":"The linear environment law supplies the residual law for this selector.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-d768cacd86c7","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6561,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:224"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalPowerOfTwoForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw hK thetaStar actionFeature R S (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction) (measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction) environment","missing":[],"search":"canonicalpoweroftwoforcedpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment banditrlproof.oful.canonicalpoweroftwoforcedpredictablescalarridgeresiduallaw_of_linearsubgaussianenvironment the linear environment law supplies the residual law for this selector. definition compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","label":"measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","description":"Source-level all-time confidence tail for the power-of-two policy.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-57308c0885e7","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6562,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:253"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw hK thetaStar actionFeature R S (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction) (measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction) environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoF…","missing":[],"search":"measure_poweroftwoforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le banditrlproof.oful.measure_poweroftwoforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le source-level all-time confidence tail for the power-of-two policy. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","label":"measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","description":"Environment-backed all-time confidence tail for the power-of-two policy.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-f24f5dd552c4","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6563,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:293"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R S environment) : Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment (allTimeTelescopingScalarRidgeConfidenceFailureSet lambda thetaStar S (canonicalHistoryTrajectoryFeature actionFeature…","missing":[],"search":"measure_poweroftwoforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le_of_linearsubgaussianenvironment banditrlproof.oful.measure_poweroftwoforcedcanonicalhistorytrajectory_alltimeconfidencefailureset_le_of_linearsubgaussianenvironment environment-backed all-time confidence tail for the power-of-two policy. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","label":"powerOfTwoOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","description":"The nonforced charge is bounded by the full pathwise width budget.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-a5da9901b766","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6564,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:326"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hdelta : 0 < delta) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (trajectory : Nat -> Fin K × Real) : powerOfTwoOptimisticSuccessorRadiusWidthCharge lambda actionFeature R delta S horizon trajectory <= telescopingStandardScalarRadiusWidthBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"poweroftwooptimisticsuccessorradiuswidthcharge_le_telescopingstandardscalarradiuswidthbound banditrlproof.oful.poweroftwooptimisticsuccessorradiuswidthcharge_le_telescopingstandardscalarradiuswidthbound the nonforced charge is bounded by the full pathwise width budget. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalHistoryTrajectory_initialGap_le_ae","label":"powerOfTwoForcedCanonicalHistoryTrajectory_initialGap_le_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalHistoryTrajectory_initialGap_le_ae","description":"The fixed initial arm has the standard deterministic gap envelope.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-5eface199e8c","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6565,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:404"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalHistoryTrajectory_initialGap_le_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0)) <= standardScalarInitialGapBound S L2","missing":[],"search":"poweroftwoforcedcanonicalhistorytrajectory_initialgap_le_ae banditrlproof.oful.poweroftwoforcedcanonicalhistorytrajectory_initialgap_le_ae the fixed initial arm has the standard deterministic gap envelope. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","description":"On one all-time confidence event, every horizon is bounded by the deterministic forced charge plus the explicit telescoping OFUL rate.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-92f8e3a3b7c5","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6566,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:449"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S force…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_le_forcedactioncharge_add_explicitbound_ae banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_le_forcedactioncharge_add_explicitbound_ae on one all-time confidence event, every horizon is bounded by the deterministic forced charge plus the explicit telescoping oful rate. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","description":"All-horizon violation event for the power-of-two forced policy.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-bc9ca1beee1a","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6567,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:509"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (forcedAction : Nat -> Fin K) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset all-horizon violation event for the power-of-two forced policy. definition compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","description":"Every all-horizon violation is a confidence failure almost surely.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-3897cf35037b","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6568,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:529"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (htheta : euclideanLength thetaStar <= S) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset_subset_confidencefailure_ae banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretallhorizonviolationset_subset_confidencefailure_ae every all-horizon violation is a confidence failure almost surely. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete all-horizon high-probability pseudo-regret theorem for the fixed power-of-two forced policy, with deterministic forced charge explicit.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedalltimeconfidence/index.html#decl-cfba313671c5","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","order":6569,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedAllTimeConfidence.lean:575"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature R…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_allhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete all-horizon high-probability pseudo-regret theorem for the fixed power-of-two forced policy, with deterministic forced charge explicit. theorem compiled","shard":"modules/4f4deace0566166d.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_of_not_mem_allHorizonViolationSet","label":"powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_of_not_mem_allHorizonViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_of_not_mem_allHorizonViolationSet","description":"Complete fixed-best average pseudo-regret tends to zero on every trajectory outside the fixed-model power-of-two forced all-horizon violation event.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency/index.html#decl-f36df487ff08","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","order":6570,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_of_not_mem_allHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (delta : Real) (hdelta : 0 < delta) (S L2 : Real) (hL2 : 0 <= L2) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (trajectory : Nat -> Fin K × Real) (hnot : trajectory ∉ powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best) : Tendsto (fun horizon => powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret thetaStar actionFeature best horizon trajectory) atTop (nhds 0)","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret_tendsto_zero_of_not_mem_allhorizonviolationset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret_tendsto_zero_of_not_mem_allhorizonviolationset complete fixed-best average pseudo-regret tends to zero on every trajectory outside the fixed-model power-of-two forced all-horizon violation event. theorem compiled","shard":"modules/5040ce7e3935c6d7.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet","label":"powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet","description":"The trajectories where complete fixed-best average pseudo-regret does not converge to zero.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency/index.html#decl-f18633b01e0a","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","order":6571,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency.lean:66"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregretconsistencyfailureset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregretconsistencyfailureset the trajectories where complete fixed-best average pseudo-regret does not converge to zero. definition compiled","shard":"modules/5040ce7e3935c6d7.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet_subset_allHorizonViolationSet","label":"powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet_subset_allHorizonViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet_subset_allHorizonViolationSet","description":"Failure of trajectory-level average consistency can only occur inside the existing all-horizon violation event.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency/index.html#decl-4ee37d2e8ecd","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","order":6572,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency.lean:84"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet_subset_allHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (delta : Real) (hdelta : 0 < delta) (S L2 : Real) (hL2 : 0 <= L2) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) : powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet thetaStar actionFeature best ⊆ powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregretconsistencyfailureset_subset_allhorizonviolationset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregretconsistencyfailureset_subset_allhorizonviolationset failure of trajectory-level average consistency can only occur inside the existing all-horizon violation event. theorem compiled","shard":"modules/5040ce7e3935c6d7.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_off_violation_and_consistencyFailure_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_off_violation_and_consistencyFailure_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_off_violation_and_consistencyFailure_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete fixed-confidence trajectory-level average-consistency theorem for one power-of-two forced policy. Outside the unchanged all-horizon event average pseudo-regret tends to zero, and the set of trajectories where this convergence fails has outer measure at most `delta` under the same canonical measure.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageconsistency/index.html#decl-c648a2dd34fa","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","order":6573,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency.lean:119"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_off_violation_and_consistencyFailure_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thet…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret_tendsto_zero_off_violation_and_consistencyfailure_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret_tendsto_zero_off_violation_and_consistencyfailure_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete fixed-confidence trajectory-level average-consistency theorem for one power-of-two forced policy. outside the unchanged all-horizon event average pseudo-regret tends to zero, and the set of trajectories where this convergence fails has outer measure at most `delta` under the same canonical measure. theorem compiled","shard":"modules/5040ce7e3935c6d7.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isLittleO_natCast_succ","label":"powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isLittleO_natCast_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isLittleO_natCast_succ","description":"The exact scalar power-of-two forced high-probability budget is `o(T + 1)` for fixed model parameters and fixed positive outer confidence budget.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html#decl-8edeab6ba6d6","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","order":6574,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isLittleO_natCast_succ {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (delta : Real) (hdelta : 0 < delta) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => powerOfTwoForcedScalarHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2) =o[atTop] (fun horizon : Nat => (((horizon + 1 : Nat) : Real)))","missing":[],"search":"poweroftwoforcedscalarhighprobabilitypseudoregretbound_islittleo_natcast_succ banditrlproof.oful.poweroftwoforcedscalarhighprobabilitypseudoregretbound_islittleo_natcast_succ the exact scalar power-of-two forced high-probability budget is `o(t + 1)` for fixed model parameters and fixed positive outer confidence budget. theorem compiled","shard":"modules/78456b0d1df0a2bd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound","label":"powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound","description":"The exact scalar high-probability budget per available round.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html#decl-44824d36fc13","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","order":6575,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean:40"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"poweroftwoforcedscalarhighprobabilityaveragepseudoregretbound banditrlproof.oful.poweroftwoforcedscalarhighprobabilityaveragepseudoregretbound the exact scalar high-probability budget per available round. definition compiled","shard":"modules/78456b0d1df0a2bd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound_tendsto_zero","label":"powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound_tendsto_zero","description":"The exact scalar high-probability budget per round converges to zero.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html#decl-dc26ca51d1e6","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","order":6576,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean:48"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound_tendsto_zero {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (delta : Real) (hdelta : 0 < delta) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : Tendsto (powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound (Feature := Feature) R delta lambda S · L2) atTop (nhds 0)","missing":[],"search":"poweroftwoforcedscalarhighprobabilityaveragepseudoregretbound_tendsto_zero banditrlproof.oful.poweroftwoforcedscalarhighprobabilityaveragepseudoregretbound_tendsto_zero the exact scalar high-probability budget per round converges to zero. theorem compiled","shard":"modules/78456b0d1df0a2bd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret","label":"powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret","description":"Complete fixed-best pseudo-regret per available round.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html#decl-cbea2a81f382","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","order":6577,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean:66"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret complete fixed-best pseudo-regret per available round. definition compiled","shard":"modules/78456b0d1df0a2bd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_le_averageBound_of_not_mem_allHorizonViolationSet","label":"powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_le_averageBound_of_not_mem_allHorizonViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_le_averageBound_of_not_mem_allHorizonViolationSet","description":"Outside the named all-horizon violation event, complete pseudo-regret per round is bounded by the exact scalar average budget at every horizon.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html#decl-871953ad23fa","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","order":6578,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean:83"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_le_averageBound_of_not_mem_allHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) (trajectory : Nat -> Fin K × Real) (hnot : trajectory ∉ powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best) (horizon : Nat) : powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret thetaStar actionFeature best horizon trajectory <= powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound (Feature := Feature) R delta lambda S horizon L2","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret_le_averagebound_of_not_mem_allhorizonviolationset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilityaveragepseudoregret_le_averagebound_of_not_mem_allhorizonviolationset outside the named all-horizon violation event, complete pseudo-regret per round is bounded by the exact scalar average budget at every horizon. theorem compiled","shard":"modules/78456b0d1df0a2bd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_averageBudget_tendsto_zero_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_averageBudget_tendsto_zero_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_averageBudget_tendsto_zero_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete fixed-model average-envelope theorem for one power-of-two forced policy: the exact scalar average budget tends to zero, it bounds complete pseudo-regret per round outside the unchanged all-horizon event, pseudo-regret is nonnegative, and the event has probability at most `delta`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityaverageregret/index.html#decl-7ef3c0aec71e","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","order":6579,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret.lean:122"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_averageBudget_tendsto_zero_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar ac…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_averagebudget_tendsto_zero_nonneg_and_allhorizon_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_averagebudget_tendsto_zero_nonneg_and_allhorizon_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete fixed-model average-envelope theorem for one power-of-two forced policy: the exact scalar average budget tends to zero, it bounds complete pseudo-regret per round outside the unchanged all-horizon event, pseudo-regret is nonnegative, and the event has probability at most `delta`. theorem compiled","shard":"modules/78456b0d1df0a2bd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound","label":"powerOfTwoForcedScalarHighProbabilityPseudoRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound","description":"The complete scalar budget displayed by the power-of-two forced tail.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-b18b339ebff2","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6580,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:22"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedScalarHighProbabilityPseudoRegretBound {Feature : Type u} [Fintype Feature] (R delta lambda S : Real) (horizon : Nat) (L2 : Real) : Real","missing":[],"search":"poweroftwoforcedscalarhighprobabilitypseudoregretbound banditrlproof.oful.poweroftwoforcedscalarhighprobabilitypseudoregretbound the complete scalar budget displayed by the power-of-two forced tail. definition compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_anti_delta","label":"telescopingHighProbabilityRegretLogBudget_anti_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_anti_delta","description":"Increasing the outer confidence budget decreases the explicit telescoping confidence logarithm.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-181449518aa5","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6581,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:34"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityRegretLogBudget_anti_delta {Feature : Type u} [Fintype Feature] (lambda L2 : Real) {deltaSmall deltaLarge : Real} (hdeltaSmall : 0 < deltaSmall) (hdelta : deltaSmall <= deltaLarge) (horizon : Nat) : telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda deltaLarge horizon L2 <= telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda deltaSmall horizon L2","missing":[],"search":"telescopinghighprobabilityregretlogbudget_anti_delta banditrlproof.oful.telescopinghighprobabilityregretlogbudget_anti_delta increasing the outer confidence budget decreases the explicit telescoping confidence logarithm. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_isBigO_log_succ","label":"telescopingHighProbabilityRegretLogBudget_isBigO_log_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_isBigO_log_succ","description":"For fixed positive `delta`, the explicit telescoping confidence logarithm is `O(log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-62b96d6d326d","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6582,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:78"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityRegretLogBudget_isBigO_log_succ {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (delta : Real) (hdelta : 0 < delta) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda delta horizon L2) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"telescopinghighprobabilityregretlogbudget_isbigo_log_succ banditrlproof.oful.telescopinghighprobabilityregretlogbudget_isbigo_log_succ for fixed positive `delta`, the explicit telescoping confidence logarithm is `o(log (t + 1))`. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","label":"telescopingHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","description":"For fixed model parameters and confidence level, the explicit telescoping high-probability budget is `O(sqrt (T + 1) * log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-745387e72335","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6583,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:194"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (delta : Real) (hdelta : 0 < delta) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2) =O[atTop] (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"telescopinghighprobabilitypseudoregretbound_isbigo_sqrt_mul_log banditrlproof.oful.telescopinghighprobabilitypseudoregretbound_isbigo_sqrt_mul_log for fixed model parameters and confidence level, the explicit telescoping high-probability budget is `o(sqrt (t + 1) * log (t + 1))`. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.natCast_log2_add_one_isBigO_log_succ","label":"natCast_log2_add_one_isBigO_log_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.natCast_log2_add_one_isBigO_log_succ","description":"The cast of `Nat.log2 T + 1` is `O(log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-507c89991815","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6584,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:321"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem natCast_log2_add_one_isBigO_log_succ : (fun horizon : Nat => ((Nat.log2 horizon + 1 : Nat) : Real)) =O[atTop] (fun horizon : Nat => Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"natcast_log2_add_one_isbigo_log_succ banditrlproof.oful.natcast_log2_add_one_isbigo_log_succ the cast of `nat.log2 t + 1` is `o(log (t + 1))`. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","label":"powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","description":"The exact scalar power-of-two forced budget is `O(sqrt (T + 1) * log (T + 1))`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-b48f6468aa17","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6585,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:393"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (delta : Real) (hdelta : 0 < delta) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (L2 : Real) (hL2 : 0 <= L2) : (fun horizon : Nat => powerOfTwoForcedScalarHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2) =O[atTop] (fun horizon : Nat => Real.sqrt (((horizon + 1 : Nat) : Real)) * Real.log (((horizon + 1 : Nat) : Real)))","missing":[],"search":"poweroftwoforcedscalarhighprobabilitypseudoregretbound_isbigo_sqrt_mul_log banditrlproof.oful.poweroftwoforcedscalarhighprobabilitypseudoregretbound_isbigo_sqrt_mul_log the exact scalar power-of-two forced budget is `o(sqrt (t + 1) * log (t + 1))`. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet","label":"powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet","description":"Named all-horizon violation event for the exact scalar asymptotic budget.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-0a254f896940","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6586,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:455"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"poweroftwoforcedcanonicalasymptotichighprobabilitypseudoregretallhorizonviolationset banditrlproof.oful.poweroftwoforcedcanonicalasymptotichighprobabilitypseudoregretallhorizonviolationset named all-horizon violation event for the exact scalar asymptotic budget. definition compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet_eq_scalar","label":"powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet_eq_scalar","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet_eq_scalar","description":"The asymptotic wrapper uses exactly the compiled scalar violation event.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-3ae20ad529a0","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6587,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:472"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet_eq_scalar {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) : powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best = powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best","missing":[],"search":"poweroftwoforcedcanonicalasymptotichighprobabilitypseudoregretallhorizonviolationset_eq_scalar banditrlproof.oful.poweroftwoforcedcanonicalasymptotichighprobabilitypseudoregretallhorizonviolationset_eq_scalar the asymptotic wrapper uses exactly the compiled scalar violation event. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_asymptoticRate_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_asymptoticRate_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_asymptoticRate_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete one-policy asymptotic high-probability theorem: the exact scalar budget has the displayed fixed-model Big-O rate, complete pseudo-regret is nonnegative, and the unchanged all-horizon violation event has probability at most `delta`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhighprobabilityregretrate/index.html#decl-dfa2cd049ae2","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","order":6588,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate.lean:493"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_asymptoticRate_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFeature…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_asymptoticrate_nonneg_and_allhorizon_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_asymptoticrate_nonneg_and_allhorizon_tail_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete one-policy asymptotic high-probability theorem: the exact scalar budget has the displayed fixed-model big-o rate, complete pseudo-regret is nonnegative, and the unchanged all-horizon violation event has probability at most `delta`. theorem compiled","shard":"modules/d1e2a899a845ad30.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","label":"finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","description":"Use the exponent-indexed forced action at power-of-two successor indices and the ordinary telescoping OFUL selector at all other history indices.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-16c0bd0c0fa3","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6589,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:26"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : Fin K","missing":[],"search":"finitehistorypoweroftwoforcedtelescopingscalarridgeaction banditrlproof.oful.finitehistorypoweroftwoforcedtelescopingscalarridgeaction use the exponent-indexed forced action at power-of-two successor indices and the ordinary telescoping oful selector at all other history indices. definition compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","label":"measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","description":"The power-of-two forced selector is measurable in its finite history.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-7b23f6a8c79b","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6590,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:44"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (n : Nat) : Measurable (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction n)","missing":[],"search":"measurable_finitehistorypoweroftwoforcedtelescopingscalarridgeaction banditrlproof.oful.measurable_finitehistorypoweroftwoforcedtelescopingscalarridgeaction the power-of-two forced selector is measurable in its finite history. theorem compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm","label":"finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm","description":"The power-of-two forced selector as one deterministic history algorithm.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-85ee25272d68","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6591,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:67"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) : Thompson.HistoryAlgorithm (Fin K) Real","missing":[],"search":"finitehistorypoweroftwoforcedtelescopingscalarridgealgorithm banditrlproof.oful.finitehistorypoweroftwoforcedtelescopingscalarridgealgorithm the power-of-two forced selector as one deterministic history algorithm. definition compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm_policy_apply","label":"finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm_policy_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm_policy_apply","description":"Every policy section is the Dirac law at the power-of-two selector.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-f8a01678433b","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6592,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:85"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm_policy_apply {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) : (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction).policy n history = Measure.dirac (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction n history)","missing":[],"search":"finitehistorypoweroftwoforcedtelescopingscalarridgealgorithm_policy_apply banditrlproof.oful.finitehistorypoweroftwoforcedtelescopingscalarridgealgorithm_policy_apply every policy section is the dirac law at the power-of-two selector. theorem compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_pow_sub_one","label":"finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_pow_sub_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_pow_sub_one","description":"At history index `2 ^ exponent - 1`, the selector uses the action prescribed for that exponent.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-3a66b2d744c1","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6593,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:114"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_pow_sub_one {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (exponent : Nat) (history : History.FinitePairHistory (Fin K) Real (2 ^ exponent - 1)) : finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction (2 ^ exponent - 1) history = forcedAction exponent","missing":[],"search":"finitehistorypoweroftwoforcedtelescopingscalarridgeaction_pow_sub_one banditrlproof.oful.finitehistorypoweroftwoforcedtelescopingscalarridgeaction_pow_sub_one at history index `2 ^ exponent - 1`, the selector uses the action prescribed for that exponent. theorem compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","label":"canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","description":"Generated successor actions follow the power-of-two selector almost surely.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-8bffd7d7866c","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6594,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:138"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, (trajectory (n + 1)).1 = finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction n (Preorder.frestrictLe n trajectory)","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_finitehistorypoweroftwoforcedtelescopingscalarridgeaction banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_finitehistorypoweroftwoforcedtelescopingscalarridgeaction generated successor actions follow the power-of-two selector almost surely. theorem compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_pow_ae_eq_forcedAction","label":"canonicalHistoryTrajectory_action_pow_ae_eq_forcedAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_pow_ae_eq_forcedAction","description":"At successor round `2 ^ exponent`, the generated action is prescribed.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedhistoryalgorithm/index.html#decl-bfeee8fd10d0","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","order":6595,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedHistoryAlgorithm.lean:172"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_pow_ae_eq_forcedAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (exponent : Nat) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, (trajectory (2 ^ exponent)).1 = forcedAction exponent","missing":[],"search":"canonicalhistorytrajectory_action_pow_ae_eq_forcedaction banditrlproof.oful.canonicalhistorytrajectory_action_pow_ae_eq_forcedaction at successor round `2 ^ exponent`, the generated action is prescribed. theorem compiled","shard":"modules/b055bad5ea7bf136.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.isPowerOfTwoForcedIndex","label":"isPowerOfTwoForcedIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.isPowerOfTwoForcedIndex","description":"A horizon-independent forcing predicate for successor indices one below powers of two. The `Nat.log2` equality makes the predicate decidable without a classical search over exponents.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-0ae425becdca","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6596,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:12"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def isPowerOfTwoForcedIndex (n : Nat) : Prop","missing":[],"search":"ispoweroftwoforcedindex banditrlproof.oful.ispoweroftwoforcedindex a horizon-independent forcing predicate for successor indices one below powers of two. the `nat.log2` equality makes the predicate decidable without a classical search over exponents. definition compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.isPowerOfTwoForcedIndex_iff","label":"isPowerOfTwoForcedIndex_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.isPowerOfTwoForcedIndex_iff","description":"The computable predicate has the intended existential power-of-two semantics.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-543a089acf2c","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6597,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem isPowerOfTwoForcedIndex_iff {n : Nat} : isPowerOfTwoForcedIndex n <-> exists k, n + 1 = 2 ^ k","missing":[],"search":"ispoweroftwoforcedindex_iff banditrlproof.oful.ispoweroftwoforcedindex_iff the computable predicate has the intended existential power-of-two semantics. theorem compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedIndexSet","label":"powerOfTwoForcedIndexSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedIndexSet","description":"Power-of-two forced successor indices strictly below `horizon`.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-219680dc692a","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6598,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:31"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def powerOfTwoForcedIndexSet (horizon : Nat) : Finset Nat","missing":[],"search":"poweroftwoforcedindexset banditrlproof.oful.poweroftwoforcedindexset power-of-two forced successor indices strictly below `horizon`. definition compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.mem_powerOfTwoForcedIndexSet_iff","label":"mem_powerOfTwoForcedIndexSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.mem_powerOfTwoForcedIndexSet_iff","description":"Membership combines the prefix bound with the intended power-of-two equation.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-956acf69ff47","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6599,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:35"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem mem_powerOfTwoForcedIndexSet_iff {horizon n : Nat} : n ∈ powerOfTwoForcedIndexSet horizon <-> n < horizon ∧ exists k, n + 1 = 2 ^ k","missing":[],"search":"mem_poweroftwoforcedindexset_iff banditrlproof.oful.mem_poweroftwoforcedindexset_iff membership combines the prefix bound with the intended power-of-two equation. theorem compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedIndexSet_zero","label":"powerOfTwoForcedIndexSet_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedIndexSet_zero","description":"theorem powerOfTwoForcedIndexSet_zero : powerOfTwoForcedIndexSet 0 = ∅","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-c6e04572496b","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6600,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:45"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedIndexSet_zero : powerOfTwoForcedIndexSet 0 = ∅","missing":[],"search":"poweroftwoforcedindexset_zero banditrlproof.oful.poweroftwoforcedindexset_zero theorem poweroftwoforcedindexset_zero : poweroftwoforcedindexset 0 = ∅ theorem compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.zero_mem_powerOfTwoForcedIndexSet_iff","label":"zero_mem_powerOfTwoForcedIndexSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.zero_mem_powerOfTwoForcedIndexSet_iff","description":"Index zero is forced exactly in nonempty horizon prefixes.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-e61fbba9554a","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6601,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:51"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem zero_mem_powerOfTwoForcedIndexSet_iff {horizon : Nat} : 0 ∈ powerOfTwoForcedIndexSet horizon <-> 0 < horizon","missing":[],"search":"zero_mem_poweroftwoforcedindexset_iff banditrlproof.oful.zero_mem_poweroftwoforcedindexset_iff index zero is forced exactly in nonempty horizon prefixes. theorem compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.card_powerOfTwoForcedIndexSet_le_log2_add_one","label":"card_powerOfTwoForcedIndexSet_le_log2_add_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.card_powerOfTwoForcedIndexSet_le_log2_add_one","description":"There are at most `Nat.log2 horizon + 1` power-of-two forced indices below a horizon. Each member embeds into the image of the admissible exponent range.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedindexcount/index.html#decl-560a7766a3e0","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","order":6602,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedIndexCount.lean:63"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem card_powerOfTwoForcedIndexSet_le_log2_add_one (horizon : Nat) : (powerOfTwoForcedIndexSet horizon).card <= Nat.log2 horizon + 1","missing":[],"search":"card_poweroftwoforcedindexset_le_log2_add_one banditrlproof.oful.card_poweroftwoforcedindexset_le_log2_add_one there are at most `nat.log2 horizon + 1` power-of-two forced indices below a horizon. each member embeds into the image of the admissible exponent range. theorem compiled","shard":"modules/0c6215825c60c5b3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_of_not_forced","label":"finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_of_not_forced","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_of_not_forced","description":"Away from forced indices, the modified selector is the telescoping selector.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-ccfa76607cb8","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6603,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_of_not_forced {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hnonforced : ¬ isPowerOfTwoForcedIndex n) : finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction n history = finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n history","missing":[],"search":"finitehistorypoweroftwoforcedtelescopingscalarridgeaction_eq_of_not_forced banditrlproof.oful.finitehistorypoweroftwoforcedtelescopingscalarridgeaction_eq_of_not_forced away from forced indices, the modified selector is the telescoping selector. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_forcedAction","label":"finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_forcedAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_forcedAction","description":"At a forced index, the selector uses the action indexed by its exponent.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-0b6f8e1171ef","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6604,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:44"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_forcedAction {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (hforced : isPowerOfTwoForcedIndex n) : finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction hK lambda actionFeature R delta S forcedAction n history = forcedAction (Nat.log2 (n + 1))","missing":[],"search":"finitehistorypoweroftwoforcedtelescopingscalarridgeaction_eq_forcedaction banditrlproof.oful.finitehistorypoweroftwoforcedtelescopingscalarridgeaction_eq_forcedaction at a forced index, the selector uses the action indexed by its exponent. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_not_powerOfTwoForced","label":"canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_not_powerOfTwoForced","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_not_powerOfTwoForced","description":"Under the modified policy's own trajectory law, every nonforced successor action agrees almost surely with the ordinary telescoping OFUL selector.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-d10d827e20a2","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6605,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:67"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_not_powerOfTwoForced {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (hnonforced : ¬ isPowerOfTwoForcedIndex n) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, (trajectory (n + 1)).1 = finiteHistoryTelescopingScalarRidgeOptimisticAction hK lambda actionFeature R delta S n (Preorder.frestrictLe n trajectory)","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction_of_not_poweroftwoforced banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_finitehistorytelescopingscalarridgeoptimisticaction_of_not_poweroftwoforced under the modified policy's own trajectory law, every nonforced successor action agrees almost surely with the ordinary telescoping oful selector. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_forcedAction_of_powerOfTwoForced","label":"canonicalHistoryTrajectory_action_succ_ae_eq_forcedAction_of_powerOfTwoForced","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_forcedAction_of_powerOfTwoForced","description":"At any forced successor index, the generated action is the indexed arm.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-0b4cdbb50dca","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6606,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:99"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalHistoryTrajectory_action_succ_ae_eq_forcedAction_of_powerOfTwoForced {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (n : Nat) (hforced : isPowerOfTwoForcedIndex n) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, (trajectory (n + 1)).1 = forcedAction (Nat.log2 (n + 1))","missing":[],"search":"canonicalhistorytrajectory_action_succ_ae_eq_forcedaction_of_poweroftwoforced banditrlproof.oful.canonicalhistorytrajectory_action_succ_ae_eq_forcedaction_of_poweroftwoforced at any forced successor index, the generated action is the indexed arm. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedSuccessorPseudoRegret","label":"powerOfTwoForcedSuccessorPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedSuccessorPseudoRegret","description":"Successor pseudo-regret charged at power-of-two forced indices.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-44bb6668bbf7","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6607,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:129"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedSuccessorPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"poweroftwoforcedsuccessorpseudoregret banditrlproof.oful.poweroftwoforcedsuccessorpseudoregret successor pseudo-regret charged at power-of-two forced indices. definition compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret","label":"powerOfTwoForcedActionSuccessorPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret","description":"Deterministic successor charge of the prescribed forced actions.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-81292ddbe7e3","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6608,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:144"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedActionSuccessorPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) (horizon : Nat) : Real","missing":[],"search":"poweroftwoforcedactionsuccessorpseudoregret banditrlproof.oful.poweroftwoforcedactionsuccessorpseudoregret deterministic successor charge of the prescribed forced actions. definition compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorPseudoRegret","label":"powerOfTwoOptimisticSuccessorPseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorPseudoRegret","description":"Successor pseudo-regret charged to nonforced optimistic actions.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-be27313d1d4e","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6609,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:158"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoOptimisticSuccessorPseudoRegret {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (horizon : Nat) (trajectory : Nat -> Fin K × Real) : Real","missing":[],"search":"poweroftwooptimisticsuccessorpseudoregret banditrlproof.oful.poweroftwooptimisticsuccessorpseudoregret successor pseudo-regret charged to nonforced optimistic actions. definition compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_powerOfTwoForced_add_optimistic","label":"canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_powerOfTwoForced_add_optimistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_powerOfTwoForced_add_optimistic","description":"Complete pseudo-regret through action `horizon` is the initial gap plus the power-of-two forced and nonforced optimistic successor charges.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-38a51498dae8","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6610,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:177"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_powerOfTwoForced_add_optimistic {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (horizon : Nat) (trajectory : Nat -> Fin K × Real) : canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory = (linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0))) + powerOfTwoForcedSuccessorPseudoRegret thetaStar actionFeature best horizon trajectory + powerOfTwoOptimisticSuccessorPseudoRegret thetaStar actionFeature best horizon trajectory","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_eq_initial_add_poweroftwoforced_add_optimistic banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_eq_initial_add_poweroftwoforced_add_optimistic complete pseudo-regret through action `horizon` is the initial gap plus the power-of-two forced and nonforced optimistic successor charges. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","label":"powerOfTwoForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","description":"Under the power-of-two forced policy, its trajectory-valued forced successor charge is almost surely the deterministic prescribed-action charge.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-57210783fe1d","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6611,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:210"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : powerOfTwoForcedSuccessorPseudoRegret thetaStar actionFeature best horizon =ᵐ[ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment] fun _trajectory => powerOfTwoForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction horizon","missing":[],"search":"poweroftwoforcedsuccessorpseudoregret_ae_eq_forcedactioncharge banditrlproof.oful.poweroftwoforcedsuccessorpseudoregret_ae_eq_forcedactioncharge under the power-of-two forced policy, its trajectory-valued forced successor charge is almost surely the deterministic prescribed-action charge. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_ae_eq_initial_add_powerOfTwoForcedAction_add_optimistic","label":"canonicalStandardHighProbabilityPseudoRegret_ae_eq_initial_add_powerOfTwoForcedAction_add_optimistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_ae_eq_initial_add_powerOfTwoForcedAction_add_optimistic","description":"The complete pathwise decomposition with its forced term already replaced by the deterministic prescribed-action charge almost surely.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedpseudoregretdecomposition/index.html#decl-3305cfa70a3d","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","order":6612,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition.lean:272"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem canonicalStandardHighProbabilityPseudoRegret_ae_eq_initial_add_powerOfTwoForcedAction_add_optimistic {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (hK : 0 < K) (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S : Real) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (horizon : Nat) (best : Fin K) : ∀ᵐ trajectory ∂ Thompson.canonicalHistoryTrajectoryMeasure (finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm hK lambda actionFeature R delta S forcedAction) environment, canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory = (linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (Thompson.canonicalHistoryTrajectoryAction trajectory 0))) + powerOfTwoForcedActionSuccessorPseudoRegret thetaStar actio…","missing":[],"search":"canonicalstandardhighprobabilitypseudoregret_ae_eq_initial_add_poweroftwoforcedaction_add_optimistic banditrlproof.oful.canonicalstandardhighprobabilitypseudoregret_ae_eq_initial_add_poweroftwoforcedaction_add_optimistic the complete pathwise decomposition with its forced term already replaced by the deterministic prescribed-action charge almost surely. theorem compiled","shard":"modules/138f6a849a8fc2c9.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_card_mul","label":"powerOfTwoForcedActionSuccessorPseudoRegret_le_card_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_card_mul","description":"A pointwise prescribed-arm gap ceiling bounds the forced charge by its cardinality.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html#decl-b7843b6f585d","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","order":6613,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean:22"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedActionSuccessorPseudoRegret_le_card_mul {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) (horizon : Nat) (forcedGapBound : Real) (hforcedGap : forall exponent, linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (forcedAction exponent)) <= forcedGapBound) : powerOfTwoForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction horizon <= ((powerOfTwoForcedIndexSet horizon).card : Real) * forcedGapBound","missing":[],"search":"poweroftwoforcedactionsuccessorpseudoregret_le_card_mul banditrlproof.oful.poweroftwoforcedactionsuccessorpseudoregret_le_card_mul a pointwise prescribed-arm gap ceiling bounds the forced charge by its cardinality. theorem compiled","shard":"modules/af164904f3022bdb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul","label":"powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul","description":"The forced charge is at most `Nat.log2 horizon + 1` times a gap ceiling.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html#decl-aa4447ed7e6e","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","order":6614,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean:54"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (forcedAction : Nat -> Fin K) (horizon : Nat) (forcedGapBound : Real) (hforcedGapBound : 0 <= forcedGapBound) (hforcedGap : forall exponent, linearValue thetaStar (actionFeature best) - linearValue thetaStar (actionFeature (forcedAction exponent)) <= forcedGapBound) : powerOfTwoForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction horizon <= ((Nat.log2 horizon + 1 : Nat) : Real) * forcedGapBound","missing":[],"search":"poweroftwoforcedactionsuccessorpseudoregret_le_log2_add_one_mul banditrlproof.oful.poweroftwoforcedactionsuccessorpseudoregret_le_log2_add_one_mul the forced charge is at most `nat.log2 horizon + 1` times a gap ceiling. theorem compiled","shard":"modules/af164904f3022bdb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul_two_mul_parameterFeatureBound","label":"powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul_two_mul_parameterFeatureBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul_two_mul_parameterFeatureBound","description":"The standard parameter and arm envelopes instantiate the generic gap ceiling for every prescribed arm.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html#decl-7f085c470a80","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","order":6615,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean:86"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul_two_mul_parameterFeatureBound {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength thetaStar <= S) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (forcedAction : Nat -> Fin K) (horizon : Nat) : powerOfTwoForcedActionSuccessorPseudoRegret thetaStar actionFeature best forcedAction horizon <= ((Nat.log2 horizon + 1 : Nat) : Real) * (2 * S * Real.sqrt L2)","missing":[],"search":"poweroftwoforcedactionsuccessorpseudoregret_le_log2_add_one_mul_two_mul_parameterfeaturebound banditrlproof.oful.poweroftwoforcedactionsuccessorpseudoregret_le_log2_add_one_mul_two_mul_parameterfeaturebound the standard parameter and arm envelopes instantiate the generic gap ceiling for every prescribed arm. theorem compiled","shard":"modules/af164904f3022bdb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","description":"Fully scalar all-horizon violation event for the power-of-two forced policy. At horizon zero the forced set is empty while `Nat.log2 0 + 1 = 1`, so this uses a harmless conservative envelope rather than an exact zero charge.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html#decl-4545e892be6e","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","order":6616,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean:121"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) : Set (Nat -> Fin K × Real)","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset fully scalar all-horizon violation event for the power-of-two forced policy. at horizon zero the forced set is empty while `nat.log2 0 + 1 = 1`, so this uses a harmless conservative envelope rather than an exact zero charge. definition compiled","shard":"modules/af164904f3022bdb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","description":"The logarithmic scalar-budget violation event is contained in the explicit-charge event.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html#decl-8420d66d06b4","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","order":6617,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean:140"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (hS : 0 <= S) (htheta : euclideanLength thetaStar <= S) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (forcedAction : Nat -> Fin K) (best : Fin K) : powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 best ⊆ powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet lambda thetaStar actionFeature R delta S L2 forcedAction best","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset_subset banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregretscalarallhorizonviolationset_subset the logarithmic scalar-budget violation event is contained in the explicit-charge event. theorem compiled","shard":"modules/af164904f3022bdb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Complete all-horizon theorem with the prescribed-action charge replaced by the logarithmic power-of-two count and the common linear arm-gap envelope.","url":"../modules/banditrlproof-ofulscheduledpoweroftwoforcedscalarchargebound/index.html#decl-c57cce18dfb8","parent":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","order":6618,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound"],["Source","BanditRLProof/OFULScheduledPowerOfTwoForcedScalarChargeBound.lean:173"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (forcedAction : Nat -> Fin K) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubgaussianEnvironmentLaw hK thetaStar actionFea…","missing":[],"search":"poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_scalarallhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.poweroftwoforcedcanonicalstandardhighprobabilitypseudoregret_nonneg_and_scalarallhorizon_tail_le_explicitbound_of_linearsubgaussianenvironment_of_featurebound_le_regularization complete all-horizon theorem with the prescribed-action charge replaced by the logarithmic power-of-two count and the common linear arm-gap envelope. theorem compiled","shard":"modules/af164904f3022bdb.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.IntegrableFiniteStoppingTime","label":"IntegrableFiniteStoppingTime","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.IntegrableFiniteStoppingTime","description":"Semantic finiteness and first-moment regularity for a `WithTop Nat` stopping time. The a.e. finiteness field rules out interpreting `untopA` at `top`; the integrability field controls the random all-round gap envelope.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-1eeb3acd5755","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6619,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:26"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure IntegrableFiniteStoppingTime {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) (tau : Omega -> WithTop Nat) : Prop where","missing":[],"search":"integrablefinitestoppingtime banditrlproof.oful.integrablefinitestoppingtime semantic finiteness and first-moment regularity for a `withtop nat` stopping time. the a.e. finiteness field rules out interpreting `untopa` at `top`; the integrability field controls the random all-round gap envelope. structure compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_randomHorizonEnvelope","label":"abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_randomHorizonEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_randomHorizonEnvelope","description":"A stopped pseudo-regret obeys the gap envelope at its random horizon.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-996750151739","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6620,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:34"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_randomHorizonEnvelope {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S : Real) (hS : 0 <= S) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (trajectory : Nat -> Fin K × Real) : |stoppedValue (fun horizon trajectory => canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory) tau trajectory| <= standardScalarAllRoundGapEnvelope S (tau trajectory).untopA L2","missing":[],"search":"abs_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_randomhorizonenvelope banditrlproof.oful.abs_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_randomhorizonenvelope a stopped pseudo-regret obeys the gap envelope at its random horizon. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_stoppingTime","label":"measurable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_stoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_stoppingTime","description":"A progressively measurable pseudo-regret process remains measurable at an arbitrary stopping time.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-95c10968c17c","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6621,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:61"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_stoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (best : Fin K) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) : Measurable (stoppedValue (fun horizon trajectory => canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory) tau)","missing":[],"search":"measurable_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_of_stoppingtime banditrlproof.oful.measurable_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_of_stoppingtime a progressively measurable pseudo-regret process remains measurable at an arbitrary stopping time. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_stoppingTime","label":"measurable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_stoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_stoppingTime","description":"The explicit deterministic budget process remains measurable at an arbitrary stopping time.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-4f9a982a91c7","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6622,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:84"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_stoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] (R delta lambda S L2 : Real) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) : Measurable (stoppedValue (fun horizon (_trajectory : Nat -> Fin K × Real) => telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2) tau)","missing":[],"search":"measurable_stoppedvalue_telescopinghighprobabilitypseudoregretbound_of_stoppingtime banditrlproof.oful.measurable_stoppedvalue_telescopinghighprobabilitypseudoregretbound_of_stoppingtime the explicit deterministic budget process remains measurable at an arbitrary stopping time. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integrable_standardScalarAllRoundGapEnvelope_at_stoppingTime","label":"integrable_standardScalarAllRoundGapEnvelope_at_stoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integrable_standardScalarAllRoundGapEnvelope_at_stoppingTime","description":"First-moment stopping-time regularity makes the random gap envelope integrable.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-0f42727622eb","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6623,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:105"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integrable_standardScalarAllRoundGapEnvelope_at_stoppingTime {K : Nat} (mu : Measure (Nat -> Fin K × Real)) (S L2 : Real) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (hstop : IntegrableFiniteStoppingTime mu tau) : Integrable (fun trajectory => standardScalarAllRoundGapEnvelope S (tau trajectory).untopA L2) mu","missing":[],"search":"integrable_standardscalarallroundgapenvelope_at_stoppingtime banditrlproof.oful.integrable_standardscalarallroundgapenvelope_at_stoppingtime first-moment stopping-time regularity makes the random gap envelope integrable. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_integrableFiniteStoppingTime","label":"integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_integrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_integrableFiniteStoppingTime","description":"The stopped pseudo-regret is integrable under first-moment stopping-time regularity. No deterministic stopping-horizon bound is used.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-f9bac085a197","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6624,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:122"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_integrableFiniteStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] (mu : Measure (Nat -> Fin K × Real)) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (S : Real) (hS : 0 <= S) (L2 : Real) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (hstop : IntegrableFiniteStoppingTime mu tau) : Integrable (stoppedValue (fun horizon trajectory => canonicalStandardHighProbabilityPseudoRegret thetaStar actionFeature best horizon trajectory) tau) mu","missing":[],"search":"integrable_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_of_integrablefinitestoppingtime banditrlproof.oful.integrable_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_of_integrablefinitestoppingtime the stopped pseudo-regret is integrable under first-moment stopping-time regularity. no deterministic stopping-horizon bound is used. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_stoppingTime","label":"measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_stoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_stoppingTime","description":"The stopped explicit violation event is measurable without a deterministic stopping bound.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-cc4c35f8306a","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6625,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:157"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_stoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] (lambda : Real) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R delta S L2 : Real) (best : Fin K) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) : MeasurableSet (telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet lambda thetaStar actionFeature R delta S L2 best tau)","missing":[],"search":"measurableset_telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_of_stoppingtime banditrlproof.oful.measurableset_telescopingcanonicalexplicithighprobabilitypseudoregretstoppedviolationset_of_stoppingtime the stopped explicit violation event is measurable without a deterministic stopping bound. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope","description":"Exact unbounded-stopping expectation decomposition. The bad-event term remains an integral of the random horizon envelope.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-c312cf38faae","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6626,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:182"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (mu : Measure (Nat -> Fin K × Real)) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (best : Fin K) (htheta : euclideanLength thetaStar <= S) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (hstop : IntegrableFiniteStoppingTime mu tau) (hbudgetIntegrable : Integrable (stoppedValue (fun horizon (_…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_integral_stoppedbudget_add_integral_badindicator_randomhorizonenvelope banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_le_integral_stoppedbudget_add_integral_badindicator_randomhorizonenvelope exact unbounded-stopping expectation decomposition. the bad-event term remains an integral of the random horizon envelope. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Canonical generated-trajectory unbounded-stopping theorem. The same stopped violation event has the compiled `delta` tail, while its random-envelope overflow remains explicit in the expectation bound.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregret/index.html#decl-9b3bb004913e","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","order":6627,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegret.lean:303"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubga…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_integral_stoppedbudget_add_integral_badindicator_randomhorizonenvelope_and_stoppedviolation_measure_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_integral_stoppedbudget_add_integral_badindicator_randomhorizonenvelope_and_stoppedviolation_measure_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization canonical generated-trajectory unbounded-stopping theorem. the same stopped violation event has the compiled `delta` tail, while its random-envelope overflow remains explicit in the expectation bound. theorem compiled","shard":"modules/0174dd6d0f187983.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_le_rounds_mul_div","label":"standardScalarLogDetBudget_le_rounds_mul_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.standardScalarLogDetBudget_le_rounds_mul_div","description":"The scalar log-determinant budget is at most linear in the round count.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-21cc4d89202e","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6628,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem standardScalarLogDetBudget_le_rounds_mul_div {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (rounds : Nat) (L2 : Real) (hL2 : 0 <= L2) : standardScalarLogDetBudget (Feature := Feature) lambda rounds L2 <= (rounds : Real) * (L2 / lambda)","missing":[],"search":"standardscalarlogdetbudget_le_rounds_mul_div banditrlproof.oful.standardscalarlogdetbudget_le_rounds_mul_div the scalar log-determinant budget is at most linear in the round count. theorem compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretQuadraticCoefficient","label":"telescopingHighProbabilityPseudoRegretQuadraticCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretQuadraticCoefficient","description":"Parameter-only coefficient in the quadratic envelope for the explicit telescoping pseudo-regret budget.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-72cd516dfa09","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6629,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:57"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def telescopingHighProbabilityPseudoRegretQuadraticCoefficient {Feature : Type u} [Fintype Feature] (R delta lambda S L2 : Real) : Real","missing":[],"search":"telescopinghighprobabilitypseudoregretquadraticcoefficient banditrlproof.oful.telescopinghighprobabilitypseudoregretquadraticcoefficient parameter-only coefficient in the quadratic envelope for the explicit telescoping pseudo-regret budget. definition compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretQuadraticCoefficient_nonneg","label":"telescopingHighProbabilityPseudoRegretQuadraticCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretQuadraticCoefficient_nonneg","description":"The quadratic-envelope coefficient is nonnegative.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-789f25b9e181","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6630,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:67"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretQuadraticCoefficient_nonneg {Feature : Type u} [Fintype Feature] (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (hS : 0 <= S) (L2 : Real) (hL2 : 0 <= L2) : 0 <= telescopingHighProbabilityPseudoRegretQuadraticCoefficient (Feature := Feature) R delta lambda S L2","missing":[],"search":"telescopinghighprobabilitypseudoregretquadraticcoefficient_nonneg banditrlproof.oful.telescopinghighprobabilitypseudoregretquadraticcoefficient_nonneg the quadratic-envelope coefficient is nonnegative. theorem compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_le_rounds_sq_mul","label":"telescopingHighProbabilityRegretLogBudget_le_rounds_sq_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_le_rounds_sq_mul","description":"The telescoping confidence logarithm is bounded by a parameter-only coefficient times the square of the round count.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-5343f6cd1be1","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6631,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:86"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityRegretLogBudget_le_rounds_sq_mul {Feature : Type u} [Fintype Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (delta : Real) (hdelta : 0 < delta) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : telescopingHighProbabilityRegretLogBudget (Feature := Feature) lambda delta horizon L2 <= (((horizon + 1 : Nat) : Real) ^ 2) * (L2 / lambda + 4 / delta)","missing":[],"search":"telescopinghighprobabilityregretlogbudget_le_rounds_sq_mul banditrlproof.oful.telescopinghighprobabilityregretlogbudget_le_rounds_sq_mul the telescoping confidence logarithm is bounded by a parameter-only coefficient times the square of the round count. theorem compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_le_rounds_sq_mul_coefficient","label":"telescopingHighProbabilityPseudoRegretBound_le_rounds_sq_mul_coefficient","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_le_rounds_sq_mul_coefficient","description":"The explicit telescoping pseudo-regret budget grows at most quadratically in the round count.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-c3545dedc0fd","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6632,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:163"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem telescopingHighProbabilityPseudoRegretBound_le_rounds_sq_mul_coefficient {Feature : Type u} [Fintype Feature] [Nonempty Feature] (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (hS : 0 <= S) (horizon : Nat) (L2 : Real) (hL2 : 0 <= L2) : telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2 <= (((horizon + 1 : Nat) : Real) ^ 2) * telescopingHighProbabilityPseudoRegretQuadraticCoefficient (Feature := Feature) R delta lambda S L2","missing":[],"search":"telescopinghighprobabilitypseudoregretbound_le_rounds_sq_mul_coefficient banditrlproof.oful.telescopinghighprobabilitypseudoregretbound_le_rounds_sq_mul_coefficient the explicit telescoping pseudo-regret budget grows at most quadratically in the round count. theorem compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integrable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_squareIntegrableFiniteStoppingTime","label":"integrable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_squareIntegrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integrable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_squareIntegrableFiniteStoppingTime","description":"Square-integrability of the stopping-time round count automatically makes the stopped explicit telescoping pseudo-regret budget integrable.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-40a3f2e9af5a","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6633,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:320"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integrable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_squareIntegrableFiniteStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (mu : Measure (Nat -> Fin K × Real)) (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (hS : 0 <= S) (L2 : Real) (hL2 : 0 <= L2) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (hstop : SquareIntegrableFiniteStoppingTime mu tau) : Integrable (stoppedValue (fun horizon (_trajectory : Nat -> Fin K × Real) => telescopingHighProbabilityPseudoRegretBound (Feature := Feature) R delta lambda S horizon L2) tau) mu","missing":[],"search":"integrable_stoppedvalue_telescopinghighprobabilitypseudoregretbound_of_squareintegrablefinitestoppingtime banditrlproof.oful.integrable_stoppedvalue_telescopinghighprobabilitypseudoregretbound_of_squareintegrablefinitestoppingtime square-integrability of the stopping-time round count automatically makes the stopped explicit telescoping pseudo-regret budget integrable. theorem compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime_automaticBudgetIntegrability","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime_automaticBudgetIntegrability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime_automaticBudgetIntegrability","description":"Canonical square-integrable unbounded-stopping expected pseudo-regret rate with stopped-budget integrability discharged automatically.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretclosed/index.html#decl-fb50094c4af9","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","order":6634,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretClosed.lean:371"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime_automaticBudgetIntegrability {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalL…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_integral_stoppedbudget_add_initialgap_mul_sqrt_roundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_squareintegrablefinitestoppingtime_automaticbudgetintegrability banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_integral_stoppedbudget_add_initialgap_mul_sqrt_roundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_squareintegrablefinitestoppingtime_automaticbudgetintegrability canonical square-integrable unbounded-stopping expected pseudo-regret rate with stopped-budget integrability discharged automatically. theorem compiled","shard":"modules/f9942c9d2c8fd217.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment","label":"stoppingTimeRoundSecondMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.stoppingTimeRoundSecondMoment","description":"The actual second moment of the real round count `tau.untopA + 1` under the square-integrable finite-stopping contract.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretexactmoment/index.html#decl-336d02683b54","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","order":6635,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def stoppingTimeRoundSecondMoment {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) (tau : Omega -> WithTop Nat) (_hstop : SquareIntegrableFiniteStoppingTime mu tau) : Real","missing":[],"search":"stoppingtimeroundsecondmoment banditrlproof.oful.stoppingtimeroundsecondmoment the actual second moment of the real round count `tau.untopa + 1` under the square-integrable finite-stopping contract. definition compiled","shard":"modules/1175000d7c838a28.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_nonneg","label":"stoppingTimeRoundSecondMoment_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_nonneg","description":"The exact stopping-time round-count second moment is nonnegative.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretexactmoment/index.html#decl-0bb0ea030d84","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","order":6636,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment.lean:32"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_nonneg {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) (tau : Omega -> WithTop Nat) (hstop : SquareIntegrableFiniteStoppingTime mu tau) : 0 <= stoppingTimeRoundSecondMoment mu tau hstop","missing":[],"search":"stoppingtimeroundsecondmoment_nonneg banditrlproof.oful.stoppingtimeroundsecondmoment_nonneg the exact stopping-time round-count second moment is nonnegative. theorem compiled","shard":"modules/1175000d7c838a28.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","description":"Canonical generated-trajectory unbounded-stopping expected pseudo-regret bound stated directly with the actual round-count second moment.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretexactmoment/index.html#decl-2464b5657159","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","order":6637,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment.lean:44"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (sour…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_stoppingtimeroundsecondmoment_add_initialgap_mul_sqrt_stoppingtimeroundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_squareintegrablefinitestoppingtime banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_stoppingtimeroundsecondmoment_add_initialgap_mul_sqrt_stoppingtimeroundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_squareintegrablefinitestoppingtime canonical generated-trajectory unbounded-stopping expected pseudo-regret bound stated directly with the actual round-count second moment. theorem compiled","shard":"modules/1175000d7c838a28.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.SquareIntegrableFiniteStoppingTime","label":"SquareIntegrableFiniteStoppingTime","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OFUL.SquareIntegrableFiniteStoppingTime","description":"Semantic finiteness and second-moment regularity for a `WithTop Nat` stopping time. Under a finite measure this contract implies `IntegrableFiniteStoppingTime`.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretrate/index.html#decl-e7fb2080405a","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","order":6638,"meta":[["Kind","structure"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretRate.lean:25"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"structure SquareIntegrableFiniteStoppingTime {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) (tau : Omega -> WithTop Nat) : Prop where","missing":[],"search":"squareintegrablefinitestoppingtime banditrlproof.oful.squareintegrablefinitestoppingtime semantic finiteness and second-moment regularity for a `withtop nat` stopping time. under a finite measure this contract implies `integrablefinitestoppingtime`. structure compiled","shard":"modules/59628a6743778641.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.SquareIntegrableFiniteStoppingTime.toIntegrableFiniteStoppingTime","label":"toIntegrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.SquareIntegrableFiniteStoppingTime.toIntegrableFiniteStoppingTime","description":"An `L2` finite stopping-time contract supplies the earlier `L1` contract.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretrate/index.html#decl-1e8aed45c450","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","order":6639,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretRate.lean:33"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem SquareIntegrableFiniteStoppingTime.toIntegrableFiniteStoppingTime {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (hstop : SquareIntegrableFiniteStoppingTime mu tau) : IntegrableFiniteStoppingTime mu tau","missing":[],"search":"tointegrablefinitestoppingtime banditrlproof.oful.squareintegrablefinitestoppingtime.tointegrablefinitestoppingtime an `l2` finite stopping-time contract supplies the earlier `l1` contract. theorem compiled","shard":"modules/59628a6743778641.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","label":"integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","description":"Project-local `L2` indicator bound. This is the nonnegative `2,2` Holder specialization needed by the random-horizon overflow consumer.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretrate/index.html#decl-95216a42d2a6","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","order":6640,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretRate.lean:46"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure {Omega : Type v} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (f : Omega -> Real) (hf_nonneg : forall omega, 0 <= f omega) (hf : MemLp f 2 mu) (bad : Set Omega) (hbad : MeasurableSet bad) : integral mu (bad.indicator f) <= Real.sqrt (integral mu (fun omega => f omega ^ 2)) * Real.sqrt (mu.real bad)","missing":[],"search":"integral_indicator_le_sqrt_secondmoment_mul_sqrt_real_measure banditrlproof.oful.integral_indicator_le_sqrt_secondmoment_mul_sqrt_real_measure project-local `l2` indicator bound. this is the nonnegative `2,2` holder specialization needed by the random-horizon overflow consumer. theorem compiled","shard":"modules/59628a6743778641.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_badIndicator_standardScalarAllRoundGapEnvelope_at_stoppingTime_le","label":"integral_badIndicator_standardScalarAllRoundGapEnvelope_at_stoppingTime_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_badIndicator_standardScalarAllRoundGapEnvelope_at_stoppingTime_le","description":"The bad-event random-horizon gap envelope is controlled by the stopping-time second moment and the square root of the event probability.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretrate/index.html#decl-80baa8507c32","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","order":6641,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretRate.lean:98"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_badIndicator_standardScalarAllRoundGapEnvelope_at_stoppingTime_le {K : Nat} (mu : Measure (Nat -> Fin K × Real)) [IsFiniteMeasure mu] (S : Real) (hS : 0 <= S) (L2 : Real) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (hstop : SquareIntegrableFiniteStoppingTime mu tau) (roundSecondMoment : Real) (hroundSecondMoment : integral mu (fun trajectory => ((((tau trajectory).untopA + 1 : Nat) : Real)) ^ 2) <= roundSecondMoment) (bad : Set (Nat -> Fin K × Real)) (hbad : MeasurableSet bad) : integral mu (bad.indicator (fun trajectory => standardScalarAllRoundGapEnvelope S (tau trajectory).untopA L2)) <= standardScalarInitialGapBound S L2 * Real.sqrt roundSecondMoment * Real.sqrt (mu.real bad)","missing":[],"search":"integral_badindicator_standardscalarallroundgapenvelope_at_stoppingtime_le banditrlproof.oful.integral_badindicator_standardscalarallroundgapenvelope_at_stoppingtime_le the bad-event random-horizon gap envelope is controlled by the stopping-time second moment and the square root of the event probability. theorem compiled","shard":"modules/59628a6743778641.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","description":"Canonical generated-trajectory unbounded-stopping expected pseudo-regret rate under a round-count second-moment bound. The `sqrt delta` term comes from the compiled stopped-event tail and the `L2` indicator inequality above.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretrate/index.html#decl-a9bbaa2977ee","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","order":6642,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretRate.lean:174"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLi…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_integral_stoppedbudget_add_initialgap_mul_sqrt_roundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_integral_stoppedbudget_add_initialgap_mul_sqrt_roundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_linearsubgaussianenvironment_of_featurebound_le_regularization canonical generated-trajectory unbounded-stopping expected pseudo-regret rate under a round-count second-moment bound. the `sqrt delta` term comes from the compiled stopped-event tail and the `l2` indicator inequality above. theorem compiled","shard":"modules/59628a6743778641.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_telescopingHighProbabilityPseudoRegretBound_le_quadraticCoefficient_mul_roundSecondMoment_of_squareIntegrableFiniteStoppingTime","label":"integral_stoppedValue_telescopingHighProbabilityPseudoRegretBound_le_quadraticCoefficient_mul_roundSecondMoment_of_squareIntegrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_telescopingHighProbabilityPseudoRegretBound_le_quadraticCoefficient_mul_roundSecondMoment_of_squareIntegrableFiniteStoppingTime","description":"The expected stopped explicit telescoping budget is controlled by the same round-count second moment used by the bad-event overflow bound.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretsecondmoment/index.html#decl-36c32b6632a7","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","order":6643,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_telescopingHighProbabilityPseudoRegretBound_le_quadraticCoefficient_mul_roundSecondMoment_of_squareIntegrableFiniteStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] [Nonempty Feature] (mu : Measure (Nat -> Fin K × Real)) (R : Real) (hR : 0 <= R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (lambda : Real) (hlambda : 0 < lambda) (S : Real) (hS : 0 <= S) (L2 : Real) (hL2 : 0 <= L2) (tau : (Nat -> Fin K × Real) -> WithTop Nat) (htau : IsStoppingTime (canonicalHistoryTrajectoryAllRoundFiltration (K := K)) tau) (hstop : SquareIntegrableFiniteStoppingTime mu tau) (roundSecondMoment : Real) (hroundSecondMoment : integral mu (fun trajectory => ((((tau trajectory).untopA + 1 : Nat) : Real)) ^ 2) <= roundSecondMoment) : integral mu (stoppedValue (fun horizon (_trajectory : Nat -> Fin K × Real) => telescopingHighProbabilityPseudoRegretBound…","missing":[],"search":"integral_stoppedvalue_telescopinghighprobabilitypseudoregretbound_le_quadraticcoefficient_mul_roundsecondmoment_of_squareintegrablefinitestoppingtime banditrlproof.oful.integral_stoppedvalue_telescopinghighprobabilitypseudoregretbound_le_quadraticcoefficient_mul_roundsecondmoment_of_squareintegrablefinitestoppingtime the expected stopped explicit telescoping budget is controlled by the same round-count second moment used by the bad-event overflow bound. theorem compiled","shard":"modules/d44da43c79de5426.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_roundSecondMoment_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_roundSecondMoment_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_roundSecondMoment_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","description":"Canonical generated-trajectory unbounded-stopping expected pseudo-regret bound with every stopped-budget term replaced by an explicit second-moment charge.","url":"../modules/banditrlproof-ofulscheduledunboundedstoppingtimeexpectedregretsecondmoment/index.html#decl-55946261e16a","parent":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","order":6644,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment"],["Source","BanditRLProof/OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment.lean:114"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_roundSecondMoment_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : IsOptimalLinearArm thetaStar actionFeature best) (source : CanonicalLinearSubg…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_roundsecondmoment_add_initialgap_mul_sqrt_roundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_squareintegrablefinitestoppingtime banditrlproof.oful.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_roundsecondmoment_add_initialgap_mul_sqrt_roundsecondmoment_mul_sqrt_delta_and_stoppedviolation_measure_le_of_squareintegrablefinitestoppingtime canonical generated-trajectory unbounded-stopping expected pseudo-regret bound with every stopped-budget term replaced by an explicit second-moment charge. theorem compiled","shard":"modules/d44da43c79de5426.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.index_le_spent_of_unitGrowth","label":"index_le_spent_of_unitGrowth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.index_le_spent_of_unitGrowth","description":"If a Nat-valued resource process grows by at least one at every step, then its value at every index dominates that index.","url":"../modules/banditrlproof-ofulscheduledunitgrowthbudgetexhaustionexpectedregret/index.html#decl-2ffabe2956cd","parent":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","order":6645,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem index_le_spent_of_unitGrowth {Omega : Type v} (spent : Nat -> Omega -> Nat) (hunit : forall t omega, spent t omega + 1 <= spent (t + 1) omega) : forall t omega, t <= spent t omega","missing":[],"search":"index_le_spent_of_unitgrowth banditrlproof.budget.index_le_spent_of_unitgrowth if a nat-valued resource process grows by at least one at every step, then its value at every index dominates that index. theorem compiled","shard":"modules/0283d78b22bb68d3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_budget_of_unitGrowth","label":"budgetExhaustionTime_le_budget_of_unitGrowth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.budgetExhaustionTime_le_budget_of_unitGrowth","description":"Unit pathwise resource growth makes the budget-exhaustion time at most the budget index.","url":"../modules/banditrlproof-ofulscheduledunitgrowthbudgetexhaustionexpectedregret/index.html#decl-7bdbc2a4d275","parent":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","order":6646,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret.lean:43"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem budgetExhaustionTime_le_budget_of_unitGrowth {Omega : Type v} (spent : Nat -> Omega -> Nat) (budget : Nat) (hunit : forall t omega, spent t omega + 1 <= spent (t + 1) omega) : forall omega, budgetExhaustionTime spent budget omega <= (budget : WithTop Nat)","missing":[],"search":"budgetexhaustiontime_le_budget_of_unitgrowth banditrlproof.budget.budgetexhaustiontime_le_budget_of_unitgrowth unit pathwise resource growth makes the budget-exhaustion time at most the budget index. theorem compiled","shard":"modules/0283d78b22bb68d3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_unitGrowth","label":"squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_unitGrowth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_unitGrowth","description":"An adapted unit-growth resource process has a square-integrable finite budget-exhaustion time under every finite measure.","url":"../modules/banditrlproof-ofulscheduledunitgrowthbudgetexhaustionexpectedregret/index.html#decl-92c0b7b6e073","parent":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","order":6647,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret.lean:59"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_unitGrowth {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget : Nat) (hspent : Adapted F spent) (hunit : forall t omega, spent t omega + 1 <= spent (t + 1) omega) : OFUL.SquareIntegrableFiniteStoppingTime mu (budgetExhaustionTime spent budget)","missing":[],"search":"squareintegrablefinitestoppingtime_budgetexhaustiontime_of_unitgrowth banditrlproof.budget.squareintegrablefinitestoppingtime_budgetexhaustiontime_of_unitgrowth an adapted unit-growth resource process has a square-integrable finite budget-exhaustion time under every finite measure. theorem compiled","shard":"modules/0283d78b22bb68d3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_unitGrowth","label":"stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_unitGrowth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_unitGrowth","description":"The exact round-count second moment of a unit-growth budget-exhaustion time is at most `(budget + 1)^2`.","url":"../modules/banditrlproof-ofulscheduledunitgrowthbudgetexhaustionexpectedregret/index.html#decl-7953b98e7224","parent":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","order":6648,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret.lean:79"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_unitGrowth {Omega : Type v} [mOmega : MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] {F : Filtration Nat mOmega} (spent : Nat -> Omega -> Nat) (budget : Nat) (hspent : Adapted F spent) (hunit : forall t omega, spent t omega + 1 <= spent (t + 1) omega) : let tau := budgetExhaustionTime spent budget let hstop := squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_unitGrowth mu spent budget hspent hunit OFUL.stoppingTimeRoundSecondMoment mu tau hstop <= (((budget + 1 : Nat) : Real)) ^ 2","missing":[],"search":"stoppingtimeroundsecondmoment_budgetexhaustiontime_le_of_unitgrowth banditrlproof.budget.stoppingtimeroundsecondmoment_budgetexhaustiontime_le_of_unitgrowth the exact round-count second moment of a unit-growth budget-exhaustion time is at most `(budget + 1)^2`. theorem compiled","shard":"modules/0283d78b22bb68d3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_of_unitGrowth","label":"integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_of_unitGrowth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_of_unitGrowth","description":"Canonical expected pseudo-regret bound for the single telescoping-schedule OFUL policy stopped at the budget-exhaustion time of an adapted unit-growth resource process.","url":"../modules/banditrlproof-ofulscheduledunitgrowthbudgetexhaustionexpectedregret/index.html#decl-bf50b20f44f9","parent":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","order":6649,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret"],["Source","BanditRLProof/OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret.lean:105"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_of_unitGrowth {K : Nat} {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (hK : 0 < K) (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (actionFeature : Fin K -> Feature -> Real) (R : Real) (hR : 0 < R) (delta : Real) (hdelta : 0 < delta) (hdelta_one : delta <= 1) (S : Real) (hS : 0 <= S) (environment : Thompson.HistoryEnvironment (Fin K) Real) (L2 : Real) (hL2 : 0 <= L2) (hactionFeatureBound : forall action, dotProduct (actionFeature action) (actionFeature action) <= L2) (hL2lambda : L2 <= lambda) (best : Fin K) (hbest : OFUL.IsOptimalLinearArm thetaStar actionFeature best) (source : OFUL.CanonicalLinearSubgaus…","missing":[],"search":"integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime_of_unitgrowth banditrlproof.budget.integral_stoppedvalue_canonicalstandardhighprobabilitypseudoregret_nonneg_and_le_quadraticcoefficient_mul_budgetroundssq_add_initialgap_mul_budgetrounds_mul_sqrt_delta_and_stoppedviolation_measure_le_of_budgetexhaustiontime_of_unitgrowth canonical expected pseudo-regret bound for the single telescoping-schedule oful policy stopped at the budget-exhaustion time of an adapted unit-growth resource process. theorem compiled","shard":"modules/0283d78b22bb68d3.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.clippedConfidenceWidth","label":"clippedConfidenceWidth","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.clippedConfidenceWidth","description":"The confidence width clipped at the inverse-quadratic level. This equals `min 1 (confidenceWidth V x)` when `V` is positive definite.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-d0e8198629cf","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6650,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:21"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedConfidenceWidth {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (x : Feature -> Real) : Real","missing":[],"search":"clippedconfidencewidth banditrlproof.oful.clippedconfidencewidth the confidence width clipped at the inverse-quadratic level. this equals `min 1 (confidencewidth v x)` when `v` is positive definite. definition compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.clippedConfidenceWidth_eq_min_one_confidenceWidth","label":"clippedConfidenceWidth_eq_min_one_confidenceWidth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.clippedConfidenceWidth_eq_min_one_confidenceWidth","description":"Clipping before the square root agrees with clipping the confidence width.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-6c70efdcc220","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6651,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:27"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem clippedConfidenceWidth_eq_min_one_confidenceWidth {Feature : Type u} [Fintype Feature] [DecidableEq Feature] (V : Matrix Feature Feature Real) (hV : V.PosDef) (x : Feature -> Real) : clippedConfidenceWidth V x = min 1 (confidenceWidth V x)","missing":[],"search":"clippedconfidencewidth_eq_min_one_confidencewidth banditrlproof.oful.clippedconfidencewidth_eq_min_one_confidencewidth clipping before the square root agrees with clipping the confidence width. theorem compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_clippedConfidenceWidth_le_sqrt_mul_sqrt_log","label":"sum_range_clippedConfidenceWidth_le_sqrt_mul_sqrt_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_clippedConfidenceWidth_le_sqrt_mul_sqrt_log","description":"The cumulative clipped widths of a bounded feature sequence are controlled by the square root of the standard logarithmic elliptical-potential budget.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-d3ecfd50d840","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6652,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:46"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_clippedConfidenceWidth_le_sqrt_mul_sqrt_log {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (history t) (history t) <= L2) : (Finset.range T).sum (fun t => clippedConfidenceWidth (regularizedPrefixFeatureGram lambda history t) (history t)) <= Real.sqrt (T : Real) * Real.sqrt (2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda))))","missing":[],"search":"sum_range_clippedconfidencewidth_le_sqrt_mul_sqrt_log banditrlproof.oful.sum_range_clippedconfidencewidth_le_sqrt_mul_sqrt_log the cumulative clipped widths of a bounded feature sequence are controlled by the square root of the standard logarithmic elliptical-potential budget. theorem compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","label":"sum_range_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","description":"Public confidence-width form of the clipped selected-feature sum bound.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-0d8f61627f54","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6653,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:112"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_min_one_confidenceWidth_le_sqrt_mul_sqrt_log {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (history t) (history t) <= L2) : (Finset.range T).sum (fun t => min 1 (confidenceWidth (regularizedPrefixFeatureGram lambda history t) (history t))) <= Real.sqrt (T : Real) * Real.sqrt (2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda))))","missing":[],"search":"sum_range_min_one_confidencewidth_le_sqrt_mul_sqrt_log banditrlproof.oful.sum_range_min_one_confidencewidth_le_sqrt_mul_sqrt_log public confidence-width form of the clipped selected-feature sum bound. theorem compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","label":"sum_range_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","description":"Raw selected-feature widths satisfy the same bound when every charged width is at most one.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-f3a6cc4f980d","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6654,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:155"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one {Feature : Type u} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (history : Nat -> Feature -> Real) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (history t) (history t) <= L2) (hwidth : forall t, t < T -> confidenceWidth (regularizedPrefixFeatureGram lambda history t) (history t) <= 1) : (Finset.range T).sum (fun t => confidenceWidth (regularizedPrefixFeatureGram lambda history t) (history t)) <= Real.sqrt (T : Real) * Real.sqrt (2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda))))","missing":[],"search":"sum_range_confidencewidth_le_sqrt_mul_sqrt_log_of_width_le_one banditrlproof.oful.sum_range_confidencewidth_le_sqrt_mul_sqrt_log_of_width_le_one raw selected-feature widths satisfy the same bound when every charged width is at most one. theorem compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_selectedAction_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","label":"sum_range_selectedAction_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_selectedAction_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","description":"Selected-action specialization of the cumulative clipped confidence-width bound.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-3fc24ce50ade","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6655,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:194"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_selectedAction_min_one_confidenceWidth_le_sqrt_mul_sqrt_log {Feature : Type u} {Action : Type v} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Action -> Feature -> Real) (selectedAction : Nat -> Action) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (actionFeature (selectedAction t)) (actionFeature (selectedAction t)) <= L2) : (Finset.range T).sum (fun t => min 1 (confidenceWidth (regularizedPrefixFeatureGram lambda (fun s => actionFeature (selectedAction s)) t) (actionFeature (selectedAction t)))) <= Real.sqrt (T : Real) * Real.sqrt (2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintype.card Feature : Real) * lambda))))","missing":[],"search":"sum_range_selectedaction_min_one_confidencewidth_le_sqrt_mul_sqrt_log banditrlproof.oful.sum_range_selectedaction_min_one_confidencewidth_le_sqrt_mul_sqrt_log selected-action specialization of the cumulative clipped confidence-width bound. theorem compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sum_range_selectedAction_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","label":"sum_range_selectedAction_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sum_range_selectedAction_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","description":"Raw selected-action width sum under an explicit small-width contract.","url":"../modules/banditrlproof-ofulselectedwidthsummation/index.html#decl-b3cf52010a46","parent":"module:BanditRLProof.OFULSelectedWidthSummation","order":6656,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelectedWidthSummation"],["Source","BanditRLProof/OFULSelectedWidthSummation.lean:222"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sum_range_selectedAction_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one {Feature : Type u} {Action : Type v} [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (lambda : Real) (hlambda : 0 < lambda) (actionFeature : Action -> Feature -> Real) (selectedAction : Nat -> Action) (T : Nat) (L2 : Real) (hL2 : 0 <= L2) (hbound : forall t, t < T -> dotProduct (actionFeature (selectedAction t)) (actionFeature (selectedAction t)) <= L2) (hwidth : forall t, t < T -> confidenceWidth (regularizedPrefixFeatureGram lambda (fun s => actionFeature (selectedAction s)) t) (actionFeature (selectedAction t)) <= 1) : (Finset.range T).sum (fun t => confidenceWidth (regularizedPrefixFeatureGram lambda (fun s => actionFeature (selectedAction s)) t) (actionFeature (selectedAction t))) <= Real.sqrt (T : Real) * Real.sqrt (2 * ((Fintype.card Feature : Real) * Real.log (1 + (T * L2) / ((Fintyp…","missing":[],"search":"sum_range_selectedaction_confidencewidth_le_sqrt_mul_sqrt_log_of_width_le_one banditrlproof.oful.sum_range_selectedaction_confidencewidth_le_sqrt_mul_sqrt_log_of_width_le_one raw selected-action width sum under an explicit small-width contract. theorem compiled","shard":"modules/d258cb96e21daf12.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt","label":"predictable_mul_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt","description":"Freeze a conditioning-measurable multiplier inside the conditional law and compensate its conditionally sub-Gaussian MGF. The explicit exponential-integrability premise is the regularity required by the local fixed-tilt composition API. A bounded-predictable-multiplier wrapper will discharge it for the OFUL feature process.","url":"../modules/banditrlproof-ofulselfnormalizedconfidence/index.html#decl-7d459c58b233","parent":"module:BanditRLProof.OFULSelfNormalizedConfidence","order":6657,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedConfidence"],["Source","BanditRLProof/OFULSelfNormalizedConfidence.lean:31"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt {Omega : Type u} {m mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : Measure Omega} [IsProbabilityMeasure mu] {X A : Omega -> Real} {c : NNReal} (hm : m <= mOmega) (hX : HasCondSubgaussianMGF m hm X c mu) (hA : @Measurable Omega Real m inferInstance A) (hintegrable : forall s : Real, Integrable (fun omega => Real.exp (s * (A omega * X omega - (((c : NNReal) : Real) * A omega ^ 2 / 2)))) mu) : BanditRLProof.Concentration.HasCondMGFUpperBoundAt m hm (fun omega => A omega * X omega - (((c : NNReal) : Real) * A omega ^ 2 / 2)) 1 0 mu","missing":[],"search":"predictable_mul_compensated_hascondmgfupperboundat probabilitytheory.hascondsubgaussianmgf.predictable_mul_compensated_hascondmgfupperboundat freeze a conditioning-measurable multiplier inside the conditional law and compensate its conditionally sub-gaussian mgf. the explicit exponential-integrability premise is the regularity required by the local fixed-tilt composition api. a bounded-predictable-multiplier wrapper will discharge it for the oful feature process. theorem compiled","shard":"modules/ec572ee7cb9ffe63.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.integrable_exp_mul_predictable_mul_compensated_of_abs_le","label":"integrable_exp_mul_predictable_mul_compensated_of_abs_le","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.integrable_exp_mul_predictable_mul_compensated_of_abs_le","description":"Uniform boundedness of a predictable multiplier discharges the exponential integrability contract of the compensated increment.","url":"../modules/banditrlproof-ofulselfnormalizedconfidence/index.html#decl-2e79335045ea","parent":"module:BanditRLProof.OFULSelfNormalizedConfidence","order":6658,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedConfidence"],["Source","BanditRLProof/OFULSelfNormalizedConfidence.lean:134"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.integrable_exp_mul_predictable_mul_compensated_of_abs_le {Omega : Type u} {m mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : Measure Omega} [IsProbabilityMeasure mu] {X A : Omega -> Real} {c : NNReal} (hm : m <= mOmega) (hX : HasCondSubgaussianMGF m hm X c mu) (hA : @Measurable Omega Real m inferInstance A) (B : Real) (hB : 0 <= B) (hAbound : forall omega, |A omega| <= B) : forall s : Real, Integrable (fun omega => Real.exp (s * (A omega * X omega - (((c : NNReal) : Real) * A omega ^ 2 / 2)))) mu","missing":[],"search":"integrable_exp_mul_predictable_mul_compensated_of_abs_le probabilitytheory.hascondsubgaussianmgf.integrable_exp_mul_predictable_mul_compensated_of_abs_le uniform boundedness of a predictable multiplier discharges the exponential integrability contract of the compensated increment. theorem compiled","shard":"modules/ec572ee7cb9ffe63.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt_of_abs_le","label":"predictable_mul_compensated_hasCondMGFUpperBoundAt_of_abs_le","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt_of_abs_le","description":"Bounded predictable multipliers satisfy the compensated conditional MGF contract without a caller-supplied exponential-integrability proof.","url":"../modules/banditrlproof-ofulselfnormalizedconfidence/index.html#decl-4012c9957de4","parent":"module:BanditRLProof.OFULSelfNormalizedConfidence","order":6659,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedConfidence"],["Source","BanditRLProof/OFULSelfNormalizedConfidence.lean:221"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt_of_abs_le {Omega : Type u} {m mOmega : MeasurableSpace Omega} [StandardBorelSpace Omega] {mu : Measure Omega} [IsProbabilityMeasure mu] {X A : Omega -> Real} {c : NNReal} (hm : m <= mOmega) (hX : HasCondSubgaussianMGF m hm X c mu) (hA : @Measurable Omega Real m inferInstance A) (B : Real) (hB : 0 <= B) (hAbound : forall omega, |A omega| <= B) : BanditRLProof.Concentration.HasCondMGFUpperBoundAt m hm (fun omega => A omega * X omega - (((c : NNReal) : Real) * A omega ^ 2 / 2)) 1 0 mu","missing":[],"search":"predictable_mul_compensated_hascondmgfupperboundat_of_abs_le probabilitytheory.hascondsubgaussianmgf.predictable_mul_compensated_hascondmgfupperboundat_of_abs_le bounded predictable multipliers satisfy the compensated conditional mgf contract without a caller-supplied exponential-integrability proof. theorem compiled","shard":"modules/ec572ee7cb9ffe63.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.fixedDirectionCompensatedScore_hasMGFUpperBoundAt","label":"fixedDirectionCompensatedScore_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.fixedDirectionCompensatedScore_hasMGFUpperBoundAt","description":"Finite-horizon fixed-direction exponential-supermartingale endpoint for predictable vector features and conditionally sub-Gaussian scalar noise. This is the deterministic-horizon local form of Lemma 1 in Abbasi-Yadkori, Pal, and Szepesvari (2011). It is the input to the Gaussian mixture step, not yet the vector self-normalized determinant-ratio theorem.","url":"../modules/banditrlproof-ofulselfnormalizedconfidence/index.html#decl-dc27489e9957","parent":"module:BanditRLProof.OFULSelfNormalizedConfidence","order":6660,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedConfidence"],["Source","BanditRLProof/OFULSelfNormalizedConfidence.lean:259"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem fixedDirectionCompensatedScore_hasMGFUpperBoundAt {Omega : Type v} {Feature : Type w} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (theta : Feature -> Real) (projectionBound : Nat -> Real) (hprojection : forall i, StronglyMeasurable[F i] (fun omega => dotProduct theta (feature i omega))) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall i, 0 <= projectionBound i) (hprojectionBound : forall i omega, |dotProduct theta (feature i omega)| <= projectionBound i) (n : Nat) (hsubGaussian : forall i, i < n -> HasCondSubgaussianMGF (F i) (F.le i) (noise i) (vari…","missing":[],"search":"fixeddirectioncompensatedscore_hasmgfupperboundat banditrlproof.oful.fixeddirectioncompensatedscore_hasmgfupperboundat finite-horizon fixed-direction exponential-supermartingale endpoint for predictable vector features and conditionally sub-gaussian scalar noise. this is the deterministic-horizon local form of lemma 1 in abbasi-yadkori, pal, and szepesvari (2011). it is the input to the gaussian mixture step, not yet the vector self-normalized determinant-ratio theorem. theorem compiled","shard":"modules/ec572ee7cb9ffe63.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.constantSquaredVarianceProxy","label":"constantSquaredVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.constantSquaredVarianceProxy","description":"The common conditional sub-Gaussian variance proxy `R²`.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-ba3db7c82f00","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6661,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:20"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def constantSquaredVarianceProxy (R : Real) : Nat -> NNReal","missing":[],"search":"constantsquaredvarianceproxy banditrlproof.oful.constantsquaredvarianceproxy the common conditional sub-gaussian variance proxy `r²`. definition compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonFeatureGram","label":"finiteHorizonFeatureGram","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonFeatureGram","description":"The unweighted finite-horizon feature Gram `sum_{i<n} x_i x_i^T`.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-aa3477770858","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6662,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:24"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteHorizonFeatureGram (feature : Nat -> Omega -> Feature -> Real) (n : Nat) (omega : Omega) : Matrix Feature Feature Real","missing":[],"search":"finitehorizonfeaturegram banditrlproof.oful.finitehorizonfeaturegram the unweighted finite-horizon feature gram `sum_{i<n} x_i x_i^t`. definition compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram_constantSquared_eq_smul_featureGram","label":"finiteHorizonVarianceGram_constantSquared_eq_smul_featureGram","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonVarianceGram_constantSquared_eq_smul_featureGram","description":"At common proxy `R²`, the variance Gram is `R²` times the unweighted feature Gram.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-10d0c39a5abf","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6663,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:33"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonVarianceGram_constantSquared_eq_smul_featureGram [Fintype Feature] (R : Real) (hR : 0 < R) (feature : Nat -> Omega -> Feature -> Real) (n : Nat) (omega : Omega) : finiteHorizonVarianceGram feature (constantSquaredVarianceProxy R) n omega = R ^ 2 • finiteHorizonFeatureGram feature n omega","missing":[],"search":"finitehorizonvariancegram_constantsquared_eq_smul_featuregram banditrlproof.oful.finitehorizonvariancegram_constantsquared_eq_smul_featuregram at common proxy `r²`, the variance gram is `r²` times the unweighted feature gram. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonFeatureGram_posSemidef","label":"finiteHorizonFeatureGram_posSemidef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonFeatureGram_posSemidef","description":"The unweighted finite-horizon feature Gram is positive semidefinite.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-171d06aeebd4","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6664,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:51"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonFeatureGram_posSemidef [Fintype Feature] (feature : Nat -> Omega -> Feature -> Real) (n : Nat) (omega : Omega) : (finiteHorizonFeatureGram feature n omega).PosSemidef","missing":[],"search":"finitehorizonfeaturegram_possemidef banditrlproof.oful.finitehorizonfeaturegram_possemidef the unweighted finite-horizon feature gram is positive semidefinite. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.sq_smul_posDef","label":"sq_smul_posDef","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.sq_smul_posDef","description":"Positive scalar-square rescaling preserves positive definiteness.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-75496f73f831","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6665,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:64"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem sq_smul_posDef [Fintype Feature] [DecidableEq Feature] (R : Real) (hR : 0 < R) (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) : (R ^ 2 • V0).PosDef","missing":[],"search":"sq_smul_posdef banditrlproof.oful.sq_smul_posdef positive scalar-square rescaling preserves positive definiteness. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.det_sq_smul_div_det_sq_smul","label":"det_sq_smul_div_det_sq_smul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.det_sq_smul_div_det_sq_smul","description":"A common nonzero scalar factor cancels from a determinant ratio.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-3f6693a07b40","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6666,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:77"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem det_sq_smul_div_det_sq_smul [Fintype Feature] [DecidableEq Feature] (R : Real) (hR : 0 < R) (A B : Matrix Feature Feature Real) : Matrix.det (R ^ 2 • A) / Matrix.det (R ^ 2 • B) = Matrix.det A / Matrix.det B","missing":[],"search":"det_sq_smul_div_det_sq_smul banditrlproof.oful.det_sq_smul_div_det_sq_smul a common nonzero scalar factor cancels from a determinant ratio. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.dotProduct_sq_smul_inv_mulVec","label":"dotProduct_sq_smul_inv_mulVec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.dotProduct_sq_smul_inv_mulVec","description":"The inverse quadratic form of a matrix scaled by `R²` is the original quadratic form divided by `R²`.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-4d2294c942ae","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6667,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:91"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem dotProduct_sq_smul_inv_mulVec [Fintype Feature] [DecidableEq Feature] (R : Real) (hR : 0 < R) (A : Matrix Feature Feature Real) (hA : A.PosDef) (score : EuclideanSpace Real Feature) : score ⬝ᵥ (R ^ 2 • A)⁻¹.mulVec score = (score ⬝ᵥ A⁻¹.mulVec score) / R ^ 2","missing":[],"search":"dotproduct_sq_smul_inv_mulvec banditrlproof.oful.dotproduct_sq_smul_inv_mulvec the inverse quadratic form of a matrix scaled by `r²` is the original quadratic form divided by `r²`. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_inv_le_finiteHorizonDetRatioInvGramExponential_le","label":"measure_inv_le_finiteHorizonDetRatioInvGramExponential_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_inv_le_finiteHorizonDetRatioInvGramExponential_le","description":"Markov's inequality for the evaluated finite-horizon Gaussian-mixture surface, with an arbitrary nonzero finite `ENNReal` confidence level.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-e821b5287dc6","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6668,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:109"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_inv_le_finiteHorizonDetRatioInvGramExponential_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsubGaussian : forall i, i < n -> Ha…","missing":[],"search":"measure_inv_le_finitehorizondetratioinvgramexponential_le banditrlproof.oful.measure_inv_le_finitehorizondetratioinvgramexponential_le markov's inequality for the evaluated finite-horizon gaussian-mixture surface, with an arbitrary nonzero finite `ennreal` confidence level. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_invOfReal_le_finiteHorizonDetRatioInvGramExponential_le","label":"measure_invOfReal_le_finiteHorizonDetRatioInvGramExponential_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_invOfReal_le_finiteHorizonDetRatioInvGramExponential_le","description":"Real-confidence specialization of the evaluated-mixture Markov bound.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-be7f04977380","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6669,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:182"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_invOfReal_le_finiteHorizonDetRatioInvGramExponential_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsubGaussian : forall i, i < n…","missing":[],"search":"measure_invofreal_le_finitehorizondetratioinvgramexponential_le banditrlproof.oful.measure_invofreal_le_finitehorizondetratioinvgramexponential_le real-confidence specialization of the evaluated-mixture markov bound. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.invOfReal_le_gaussianDetRatioExponential_of_invGramQuadratic_gt","label":"invOfReal_le_gaussianDetRatioExponential_of_invGramQuadratic_gt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.invOfReal_le_gaussianDetRatioExponential_of_invGramQuadratic_gt","description":"The inverse-Gram/log-determinant bad-event inequality implies the evaluated Gaussian-mixture Markov threshold, pointwise.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-7b03d0ed7631","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6670,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:225"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem invOfReal_le_gaussianDetRatioExponential_of_invGramQuadratic_gt [Fintype Feature] [DecidableEq Feature] (V0 G : Matrix Feature Feature Real) (hV0 : V0.PosDef) (hG : G.PosSemidef) (score : EuclideanSpace Real Feature) (delta : Real) (hdelta : 0 < delta) (hbad : score ⬝ᵥ (V0 + G)⁻¹.mulVec score > 2 * Real.log (Real.sqrt (Matrix.det (V0 + G) / Matrix.det V0) / delta)) : (ENNReal.ofReal delta)⁻¹ <= ENNReal.ofReal (Real.sqrt (Matrix.det V0 / Matrix.det (V0 + G)) * Real.exp (score ⬝ᵥ (V0 + G)⁻¹.mulVec score / 2))","missing":[],"search":"invofreal_le_gaussiandetratioexponential_of_invgramquadratic_gt banditrlproof.oful.invofreal_le_gaussiandetratioexponential_of_invgramquadratic_gt the inverse-gram/log-determinant bad-event inequality implies the evaluated gaussian-mixture markov threshold, pointwise. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizon_invGramQuadratic_gt_two_log_detRatio_div_le","label":"measure_finiteHorizon_invGramQuadratic_gt_two_log_detRatio_div_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizon_invGramQuadratic_gt_two_log_detRatio_div_le","description":"Finite-horizon inverse-Gram/log-determinant bad-event probability bound for the variance-weighted Gram produced by the conditional-MGF route.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-a072d866ef29","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6671,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:296"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizon_invGramQuadratic_gt_two_log_detRatio_div_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (varianceProxy : Nat -> NNReal) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsubGaussian : forall i, i <…","missing":[],"search":"measure_finitehorizon_invgramquadratic_gt_two_log_detratio_div_le banditrlproof.oful.measure_finitehorizon_invgramquadratic_gt_two_log_detratio_div_le finite-horizon inverse-gram/log-determinant bad-event probability bound for the variance-weighted gram produced by the conditional-mgf route. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizon_selfNormalizedQuadratic_gt_two_mul_sq_mul_log_detRatio_div_le","label":"measure_finiteHorizon_selfNormalizedQuadratic_gt_two_mul_sq_mul_log_detRatio_div_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizon_selfNormalizedQuadratic_gt_two_mul_sq_mul_log_detRatio_div_le","description":"Finite-horizon self-normalized tail bound with the conventional common conditional sub-Gaussian scale `R`. The matrix in the reported norm is the unweighted feature Gram, while the confidence radius carries the factor `R²`.","url":"../modules/banditrlproof-ofulselfnormalizedmarkov/index.html#decl-f0dc224ebe02","parent":"module:BanditRLProof.OFULSelfNormalizedMarkov","order":6672,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULSelfNormalizedMarkov"],["Source","BanditRLProof/OFULSelfNormalizedMarkov.lean:380"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizon_selfNormalizedQuadratic_gt_two_mul_sq_mul_log_detRatio_div_le [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (V0 : Matrix Feature Feature Real) (hV0 : V0.PosDef) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (WithLp.ofLp theta) (feature i omega)| <= projectionBound theta i) (n : Nat) (hsubGaussian : for…","missing":[],"search":"measure_finitehorizon_selfnormalizedquadratic_gt_two_mul_sq_mul_log_detratio_div_le banditrlproof.oful.measure_finitehorizon_selfnormalizedquadratic_gt_two_mul_sq_mul_log_detratio_div_le finite-horizon self-normalized tail bound with the conventional common conditional sub-gaussian scale `r`. the matrix in the reported norm is the unweighted feature gram, while the confidence radius carries the factor `r²`. theorem compiled","shard":"modules/61b98a7b71d70c75.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.scalarRidgeConfidenceFailureAt","label":"scalarRidgeConfidenceFailureAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.scalarRidgeConfidenceFailureAt","description":"Failure of the scalar-ridge confidence ellipsoid at one horizon.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-f6886da0b7aa","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6673,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:23"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def scalarRidgeConfidenceFailureAt {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R delta : Real) (n : Nat) : Set Omega","missing":[],"search":"scalarridgeconfidencefailureat banditrlproof.oful.scalarridgeconfidencefailureat failure of the scalar-ridge confidence ellipsoid at one horizon. definition compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonScheduledScalarRidgeConfidenceFailureSet","label":"finiteHorizonScheduledScalarRidgeConfidenceFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonScheduledScalarRidgeConfidenceFailureSet","description":"Union of scalar-ridge confidence failures over all `n <= horizon`, with the confidence level selected by `deltaAt n`.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-e96341a4f65e","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6674,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:47"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def finiteHorizonScheduledScalarRidgeConfidenceFailureSet {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R : Real) (deltaAt : Nat -> Real) (horizon : Nat) : Set Omega","missing":[],"search":"finitehorizonscheduledscalarridgeconfidencefailureset banditrlproof.oful.finitehorizonscheduledscalarridgeconfidencefailureset union of scalar-ridge confidence failures over all `n <= horizon`, with the confidence level selected by `deltaat n`. definition compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","label":"mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","description":"Membership in the scheduled failure set is failure at some `n <= horizon`.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-3278f359db5e","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6675,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:63"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R : Real) (deltaAt : Nat -> Real) (horizon : Nat) (omega : Omega) : omega ∈ finiteHorizonScheduledScalarRidgeConfidenceFailureSet lambda thetaStar S feature response R deltaAt horizon ↔ ∃ n, n <= horizon ∧ omega ∈ scalarRidgeConfidenceFailureAt lambda thetaStar S feature response R (deltaAt n) n","missing":[],"search":"mem_finitehorizonscheduledscalarridgeconfidencefailureset_iff banditrlproof.oful.mem_finitehorizonscheduledscalarridgeconfidencefailureset_iff membership in the scheduled failure set is failure at some `n <= horizon`. theorem compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.not_mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","label":"not_mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.not_mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","description":"Outside the scheduled failure union, every scalar-ridge confidence ellipsoid in the inclusive finite window holds simultaneously.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-3aa31771004a","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6676,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:86"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem not_mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R : Real) (deltaAt : Nat -> Real) (horizon : Nat) (omega : Omega) : omega ∉ finiteHorizonScheduledScalarRidgeConfidenceFailureSet lambda thetaStar S feature response R deltaAt horizon ↔ ∀ n, n <= horizon -> matrixNorm (Matrix.scalar Feature lambda + finiteHorizonFeatureGram feature n omega) (finiteHorizonRidgeEstimate (Matrix.scalar Feature lambda) feature response n omega - thetaStar) <= finiteHorizonScalarConfidenceRadius feature R (deltaAt n) lambda S n omega","missing":[],"search":"not_mem_finitehorizonscheduledscalarridgeconfidencefailureset_iff banditrlproof.oful.not_mem_finitehorizonscheduledscalarridgeconfidencefailureset_iff outside the scheduled failure union, every scalar-ridge confidence ellipsoid in the inclusive finite window holds simultaneously. theorem compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_le_sum","label":"measure_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_le_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_le_sum","description":"Scheduled finite-window confidence: the failure-union probability is bounded by the sum of the fixed-time confidence budgets.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-a439e5c18819","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6677,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:116"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_le_sum {Omega : Type u} {Feature : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProdu…","missing":[],"search":"measure_finitehorizonscheduledscalarridgeconfidencefailureset_le_sum banditrlproof.oful.measure_finitehorizonscheduledscalarridgeconfidencefailureset_le_sum scheduled finite-window confidence: the failure-union probability is bounded by the sum of the fixed-time confidence budgets. theorem compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.finiteHorizonUniformScalarRidgeConfidenceFailureSet","label":"finiteHorizonUniformScalarRidgeConfidenceFailureSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.OFUL.finiteHorizonUniformScalarRidgeConfidenceFailureSet","description":"Equal-share finite-window failure union. The inclusive window has `horizon + 1` horizons, so each fixed-time theorem receives `delta / (horizon + 1)`.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-1a7f8e0502bb","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6678,"meta":[["Kind","definition"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:188"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"def finiteHorizonUniformScalarRidgeConfidenceFailureSet {Omega : Type u} {Feature : Type v} [Fintype Feature] [DecidableEq Feature] (lambda : Real) (thetaStar : Feature -> Real) (S : Real) (feature : Nat -> Omega -> Feature -> Real) (response : Nat -> Omega -> Real) (R delta : Real) (horizon : Nat) : Set Omega","missing":[],"search":"finitehorizonuniformscalarridgeconfidencefailureset banditrlproof.oful.finitehorizonuniformscalarridgeconfidencefailureset equal-share finite-window failure union. the inclusive window has `horizon + 1` horizons, so each fixed-time theorem receives `delta / (horizon + 1)`. definition compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.OFUL.measure_finiteHorizonUniformScalarRidgeConfidenceFailureSet_le","label":"measure_finiteHorizonUniformScalarRidgeConfidenceFailureSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.OFUL.measure_finiteHorizonUniformScalarRidgeConfidenceFailureSet_le","description":"Uniform scalar-ridge confidence over every deterministic horizon `n <= horizon`, obtained by equal allocation of the total failure budget.","url":"../modules/banditrlproof-ofuluniformtimeconfidence/index.html#decl-ded226b4ffc8","parent":"module:BanditRLProof.OFULUniformTimeConfidence","order":6679,"meta":[["Kind","theorem"],["Module","BanditRLProof.OFULUniformTimeConfidence"],["Source","BanditRLProof/OFULUniformTimeConfidence.lean:206"],["Chapter","OFUL"],["Used in books","bandit"],["Reading references","teaching:oful"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonUniformScalarRidgeConfidenceFailureSet_le {Omega : Type u} {Feature : Type v} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] [Fintype Feature] [DecidableEq Feature] [Nonempty Feature] (mu : Measure Omega) [IsProbabilityMeasure mu] (lambda : Real) (hlambda : 0 < lambda) (thetaStar : Feature -> Real) (S : Real) (htheta : euclideanLength thetaStar <= S) (F : Filtration Nat mOmega) (feature : Nat -> Omega -> Feature -> Real) (response noise : Nat -> Omega -> Real) (R : Real) (hR : 0 < R) (projectionBound : EuclideanSpace Real Feature -> Nat -> Real) (hfeature : forall i j, StronglyMeasurable[F i] (fun omega => feature i omega j)) (hnoise : StronglyAdapted F (fun t omega => match t with | 0 => 0 | i + 1 => noise i omega)) (hprojectionBound_nonneg : forall theta i, 0 <= projectionBound theta i) (hprojectionBound : forall theta i omega, |dotProduct (Wi…","missing":[],"search":"measure_finitehorizonuniformscalarridgeconfidencefailureset_le banditrlproof.oful.measure_finitehorizonuniformscalarridgeconfidencefailureset_le uniform scalar-ridge confidence over every deterministic horizon `n <= horizon`, obtained by equal allocation of the total failure budget. theorem compiled","shard":"modules/50b065f3471e8cdd.json","books":["bandit"],"chapters":["teaching:oful"],"settings":[]},{"id":"declaration:BanditRLProof.ProblemArea","label":"ProblemArea","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.ProblemArea","description":"inductive ProblemArea where","url":"../modules/banditrlproof-openproblems/index.html#decl-cb39115fb721","parent":"module:BanditRLProof.OpenProblems","order":6680,"meta":[["Kind","inductive type"],["Module","BanditRLProof.OpenProblems"],["Source","BanditRLProof/OpenProblems.lean:9"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"inductive ProblemArea where","missing":[],"search":"problemarea banditrlproof.problemarea inductive problemarea where inductive type compiled","shard":"modules/173ca7392389ffa8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.OpenProblem","label":"OpenProblem","kind":"structure","status":"compiled","subtitle":"BanditRLProof.OpenProblem","description":"structure OpenProblem where","url":"../modules/banditrlproof-openproblems/index.html#decl-3296114c4ab4","parent":"module:BanditRLProof.OpenProblems","order":6681,"meta":[["Kind","structure"],["Module","BanditRLProof.OpenProblems"],["Source","BanditRLProof/OpenProblems.lean:19"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"structure OpenProblem where","missing":[],"search":"openproblem banditrlproof.openproblem structure openproblem where structure compiled","shard":"modules/173ca7392389ffa8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.seedOpenProblems","label":"seedOpenProblems","kind":"definition","status":"compiled","subtitle":"BanditRLProof.seedOpenProblems","description":"def seedOpenProblems : List OpenProblem","url":"../modules/banditrlproof-openproblems/index.html#decl-64ff803056b1","parent":"module:BanditRLProof.OpenProblems","order":6682,"meta":[["Kind","definition"],["Module","BanditRLProof.OpenProblems"],["Source","BanditRLProof/OpenProblems.lean:28"],["Chapter","Frontier"],["Used in books","bandit"],["Reading references","teaching:frontier"],["Indexed settings","None registered"]],"statement":"def seedOpenProblems : List OpenProblem","missing":[],"search":"seedopenproblems banditrlproof.seedopenproblems def seedopenproblems : list openproblem definition compiled","shard":"modules/173ca7392389ffa8.json","books":["bandit"],"chapters":["teaching:frontier"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.MeasurablePolicy","label":"MeasurablePolicy","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Policy.MeasurablePolicy","description":"A policy whose action map is measurable from a history/context state into the action space.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-42431bda83c8","parent":"module:BanditRLProof.PolicyMeasurability","order":6683,"meta":[["Kind","structure"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:21"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure MeasurablePolicy (State : Type u) (Action : Type v) [MeasurableSpace State] [MeasurableSpace Action] where","missing":[],"search":"measurablepolicy banditrlproof.policy.measurablepolicy a policy whose action map is measurable from a history/context state into the action space. structure compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_action_of_measurable_state","label":"measurable_action_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_action_of_measurable_state","description":"A measurable history/context state composed with a measurable policy is a measurable action random variable.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-0d8031bc04e6","parent":"module:BanditRLProof.PolicyMeasurability","order":6684,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:29"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_action_of_measurable_state {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (policy : MeasurablePolicy State Action) (state : Omega -> State) (hstate : Measurable state) : Measurable (fun omega : Omega => policy.action (state omega))","missing":[],"search":"measurable_action_of_measurable_state banditrlproof.policy.measurable_action_of_measurable_state a measurable history/context state composed with a measurable policy is a measurable action random variable. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_action_mem_filtration_of_measurable_state","label":"measurable_action_mem_filtration_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_action_mem_filtration_of_measurable_state","description":"If a history/context state is measurable with respect to a filtration at time `t`, then the action selected by a measurable policy from that state is also measurable with respect to the same filtration.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-0a9d67e2d89c","parent":"module:BanditRLProof.PolicyMeasurability","order":6685,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:43"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_action_mem_filtration_of_measurable_state {Omega : Type w} {State : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (F : MeasureTheory.Filtration Nat mOmega) (policy : MeasurablePolicy State Action) (state : Omega -> State) (t : Nat) (hstate : @Measurable Omega State (F t) inferInstance state) : @Measurable Omega Action (F t) inferInstance (fun omega : Omega => policy.action (state omega))","missing":[],"search":"measurable_action_mem_filtration_of_measurable_state banditrlproof.policy.measurable_action_mem_filtration_of_measurable_state if a history/context state is measurable with respect to a filtration at time `t`, then the action selected by a measurable policy from that state is also measurable with respect to the same filtration. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_action_mem_historyFiltration_of_measurable_state","label":"measurable_action_mem_historyFiltration_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_action_mem_historyFiltration_of_measurable_state","description":"Specialization of policy measurability to the generated history filtration: once a policy state is measurable from the past action/reward history at time `t`, the policy-selected action is measurable from that same history.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-caecb09971d7","parent":"module:BanditRLProof.PolicyMeasurability","order":6686,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:64"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_action_mem_historyFiltration_of_measurable_state {Omega : Type w} {TraceAction : Type v} {Reward : Type x} {State : Type u} {PolicyAction : Type v} [mOmega : MeasurableSpace Omega] [MeasurableSpace TraceAction] [MeasurableSingletonClass TraceAction] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [MeasurableSpace State] [MeasurableSpace PolicyAction] (traceAction : Omega -> ActionTrace TraceAction) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => traceAction omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (policy : MeasurablePolicy State PolicyAction) (state : Omega -> State) (t : Nat) (hstate : @Measurable Omega State (History.historyFiltration traceAction reward haction hreward t) inferInstance state) : @Measurable Omega PolicyAction (History.historyFiltration traceAc…","missing":[],"search":"measurable_action_mem_historyfiltration_of_measurable_state banditrlproof.policy.measurable_action_mem_historyfiltration_of_measurable_state specialization of policy measurability to the generated history filtration: once a policy state is measurable from the past action/reward history at time `t`, the policy-selected action is measurable from that same history. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.generatedActionTrace","label":"generatedActionTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Policy.generatedActionTrace","description":"The action trace generated by applying a fixed measurable policy to a time-indexed history/context state process.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-3dd11bfaeaff","parent":"module:BanditRLProof.PolicyMeasurability","order":6687,"meta":[["Kind","definition"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:98"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionTrace {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace State] [MeasurableSpace Action] (policy : MeasurablePolicy State Action) (state : Nat -> Omega -> State) : Omega -> ActionTrace Action","missing":[],"search":"generatedactiontrace banditrlproof.policy.generatedactiontrace the action trace generated by applying a fixed measurable policy to a time-indexed history/context state process. definition compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_generatedActionTrace_eval_of_measurable_state","label":"measurable_generatedActionTrace_eval_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_generatedActionTrace_eval_of_measurable_state","description":"If every history/context state coordinate is ambient-measurable, every coordinate of the generated policy action trace is ambient-measurable.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-c45ac8abdf64","parent":"module:BanditRLProof.PolicyMeasurability","order":6688,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:110"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedActionTrace_eval_of_measurable_state {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (policy : MeasurablePolicy State Action) (state : Nat -> Omega -> State) (hstate : forall t : Nat, Measurable (state t)) (t : Nat) : Measurable (fun omega : Omega => (generatedActionTrace policy state omega) t)","missing":[],"search":"measurable_generatedactiontrace_eval_of_measurable_state banditrlproof.policy.measurable_generatedactiontrace_eval_of_measurable_state if every history/context state coordinate is ambient-measurable, every coordinate of the generated policy action trace is ambient-measurable. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.generatedActionTraceSucc","label":"generatedActionTraceSucc","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Policy.generatedActionTraceSucc","description":"The shifted action trace generated by a time-indexed policy. The value at time `t + 1` is selected by `policy t` from the state available at time `t`. This matches the one-step kernel convention used later by `RewardKernel.actionRewardHistoryStepKernelFamily`.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-67faa7b78994","parent":"module:BanditRLProof.PolicyMeasurability","order":6689,"meta":[["Kind","definition"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:130"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def generatedActionTraceSucc {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> MeasurablePolicy State Action) (state : Nat -> Omega -> State) (defaultAction : Action) : Omega -> ActionTrace Action","missing":[],"search":"generatedactiontracesucc banditrlproof.policy.generatedactiontracesucc the shifted action trace generated by a time-indexed policy. the value at time `t + 1` is selected by `policy t` from the state available at time `t`. this matches the one-step kernel convention used later by `rewardkernel.actionrewardhistorystepkernelfamily`. definition compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.generatedActionTraceSucc_zero","label":"generatedActionTraceSucc_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.generatedActionTraceSucc_zero","description":"The initial value of the shifted generated trace is the chosen default.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-bd842ed9ba82","parent":"module:BanditRLProof.PolicyMeasurability","order":6690,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:144"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionTraceSucc_zero {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> MeasurablePolicy State Action) (state : Nat -> Omega -> State) (defaultAction : Action) (omega : Omega) : (generatedActionTraceSucc policy state defaultAction omega) 0 = defaultAction","missing":[],"search":"generatedactiontracesucc_zero banditrlproof.policy.generatedactiontracesucc_zero the initial value of the shifted generated trace is the chosen default. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.generatedActionTraceSucc_succ","label":"generatedActionTraceSucc_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.generatedActionTraceSucc_succ","description":"At time `t + 1`, the shifted generated trace is exactly the action selected by `policy t`.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-71bfad075503","parent":"module:BanditRLProof.PolicyMeasurability","order":6691,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:159"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionTraceSucc_succ {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> MeasurablePolicy State Action) (state : Nat -> Omega -> State) (defaultAction : Action) (omega : Omega) (t : Nat) : (generatedActionTraceSucc policy state defaultAction omega) (t + 1) = (policy t).action (state t omega)","missing":[],"search":"generatedactiontracesucc_succ banditrlproof.policy.generatedactiontracesucc_succ at time `t + 1`, the shifted generated trace is exactly the action selected by `policy t`. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.generatedActionTraceSucc_succ_eq","label":"generatedActionTraceSucc_succ_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.generatedActionTraceSucc_succ_eq","description":"Function-level equality for the predictable `t + 1` coordinate.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-9f14243ba6b0","parent":"module:BanditRLProof.PolicyMeasurability","order":6692,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:171"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem generatedActionTraceSucc_succ_eq {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> MeasurablePolicy State Action) (state : Nat -> Omega -> State) (defaultAction : Action) (t : Nat) : (fun omega : Omega => (generatedActionTraceSucc policy state defaultAction omega) (t + 1)) = (fun omega : Omega => (policy t).action (state t omega))","missing":[],"search":"generatedactiontracesucc_succ_eq banditrlproof.policy.generatedactiontracesucc_succ_eq function-level equality for the predictable `t + 1` coordinate. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_generatedActionTraceSucc_eval_of_measurable_state","label":"measurable_generatedActionTraceSucc_eval_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_generatedActionTraceSucc_eval_of_measurable_state","description":"If every state coordinate is ambient-measurable, every coordinate of the shifted generated policy action trace is ambient-measurable.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-06df906b52eb","parent":"module:BanditRLProof.PolicyMeasurability","order":6693,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:188"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedActionTraceSucc_eval_of_measurable_state {Omega : Type w} {State : Type u} {Action : Type v} [MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (policy : Nat -> MeasurablePolicy State Action) (state : Nat -> Omega -> State) (defaultAction : Action) (hstate : forall t : Nat, Measurable (state t)) (t : Nat) : Measurable (fun omega : Omega => (generatedActionTraceSucc policy state defaultAction omega) t)","missing":[],"search":"measurable_generatedactiontracesucc_eval_of_measurable_state banditrlproof.policy.measurable_generatedactiontracesucc_eval_of_measurable_state if every state coordinate is ambient-measurable, every coordinate of the shifted generated policy action trace is ambient-measurable. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_generatedActionTraceSucc_succ_mem_filtration_of_measurable_state","label":"measurable_generatedActionTraceSucc_succ_mem_filtration_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_generatedActionTraceSucc_succ_mem_filtration_of_measurable_state","description":"Predictable-coordinate measurability for the shifted generated action trace: if the state at time `t` is measurable in `F t`, then action coordinate `t + 1` is also measurable in `F t`.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-bdd1be6c3699","parent":"module:BanditRLProof.PolicyMeasurability","order":6694,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:211"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedActionTraceSucc_succ_mem_filtration_of_measurable_state {Omega : Type w} {State : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (F : MeasureTheory.Filtration Nat mOmega) (policy : Nat -> MeasurablePolicy State Action) (state : Nat -> Omega -> State) (defaultAction : Action) (hstate : forall t : Nat, @Measurable Omega State (F t) inferInstance (state t)) (t : Nat) : @Measurable Omega Action (F t) inferInstance (fun omega : Omega => (generatedActionTraceSucc policy state defaultAction omega) (t + 1))","missing":[],"search":"measurable_generatedactiontracesucc_succ_mem_filtration_of_measurable_state banditrlproof.policy.measurable_generatedactiontracesucc_succ_mem_filtration_of_measurable_state predictable-coordinate measurability for the shifted generated action trace: if the state at time `t` is measurable in `f t`, then action coordinate `t + 1` is also measurable in `f t`. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_generatedActionTrace_eval_mem_filtration_of_measurable_state","label":"measurable_generatedActionTrace_eval_mem_filtration_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_generatedActionTrace_eval_mem_filtration_of_measurable_state","description":"If every history/context state coordinate is measurable with respect to the filtration at the same time, every coordinate of the generated policy action trace is measurable with respect to that filtration.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-48b61d706105","parent":"module:BanditRLProof.PolicyMeasurability","order":6695,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:237"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedActionTrace_eval_mem_filtration_of_measurable_state {Omega : Type w} {State : Type u} {Action : Type v} [mOmega : MeasurableSpace Omega] [MeasurableSpace State] [MeasurableSpace Action] (F : MeasureTheory.Filtration Nat mOmega) (policy : MeasurablePolicy State Action) (state : Nat -> Omega -> State) (hstate : forall t : Nat, @Measurable Omega State (F t) inferInstance (state t)) (t : Nat) : @Measurable Omega Action (F t) inferInstance (fun omega : Omega => (generatedActionTrace policy state omega) t)","missing":[],"search":"measurable_generatedactiontrace_eval_mem_filtration_of_measurable_state banditrlproof.policy.measurable_generatedactiontrace_eval_mem_filtration_of_measurable_state if every history/context state coordinate is measurable with respect to the filtration at the same time, every coordinate of the generated policy action trace is measurable with respect to that filtration. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.Policy.measurable_generatedActionTrace_eval_mem_historyFiltration_of_measurable_state","label":"measurable_generatedActionTrace_eval_mem_historyFiltration_of_measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Policy.measurable_generatedActionTrace_eval_mem_historyFiltration_of_measurable_state","description":"Specialization of generated policy trace measurability to the generated history filtration.","url":"../modules/banditrlproof-policymeasurability/index.html#decl-a1225d8453dd","parent":"module:BanditRLProof.PolicyMeasurability","order":6696,"meta":[["Kind","theorem"],["Module","BanditRLProof.PolicyMeasurability"],["Source","BanditRLProof/PolicyMeasurability.lean:261"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedActionTrace_eval_mem_historyFiltration_of_measurable_state {Omega : Type w} {TraceAction : Type v} {Reward : Type x} {State : Type u} {PolicyAction : Type v} [mOmega : MeasurableSpace Omega] [MeasurableSpace TraceAction] [MeasurableSingletonClass TraceAction] [MeasurableSpace Reward] [MeasurableSingletonClass Reward] [MeasurableSpace State] [MeasurableSpace PolicyAction] (traceAction : Omega -> ActionTrace TraceAction) (reward : Omega -> RewardTrace Reward) (haction : forall t : Nat, Measurable (fun omega : Omega => traceAction omega t)) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (policy : MeasurablePolicy State PolicyAction) (state : Nat -> Omega -> State) (hstate : forall t : Nat, @Measurable Omega State (History.historyFiltration traceAction reward haction hreward t) inferInstance (state t)) (t : Nat) : @Measurable Omega P…","missing":[],"search":"measurable_generatedactiontrace_eval_mem_historyfiltration_of_measurable_state banditrlproof.policy.measurable_generatedactiontrace_eval_mem_historyfiltration_of_measurable_state specialization of generated policy trace measurability to the generated history filtration. theorem compiled","shard":"modules/bfb5aa112ca61654.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.MarkovPosteriorKernel","label":"MarkovPosteriorKernel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.MarkovPosteriorKernel","description":"A posterior distribution over environments indexed by observed histories. The underlying object is Mathlib's `ProbabilityTheory.Kernel`; the local wrapper gives Thompson-sampling and Bayesian-regret leaves a stable project name for the regularity contract.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-c1752f2bdf4b","parent":"module:BanditRLProof.PosteriorKernel","order":6697,"meta":[["Kind","structure"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:29"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure MarkovPosteriorKernel (History : Type u) (Env : Type v) [MeasurableSpace History] [MeasurableSpace Env] where","missing":[],"search":"markovposteriorkernel banditrlproof.posteriorkernel.markovposteriorkernel a posterior distribution over environments indexed by observed histories. the underlying object is mathlib's `probabilitytheory.kernel`; the local wrapper gives thompson-sampling and bayesian-regret leaves a stable project name for the regularity contract. structure compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.ofKernel","label":"ofKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.ofKernel","description":"Build the local posterior-kernel contract from an existing Mathlib kernel.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-cff773c8f06d","parent":"module:BanditRLProof.PosteriorKernel","order":6698,"meta":[["Kind","definition"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:44"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def ofKernel (kernel : ProbabilityTheory.Kernel History Env) (hkernel : ProbabilityTheory.IsMarkovKernel kernel) : MarkovPosteriorKernel History Env where","missing":[],"search":"ofkernel banditrlproof.posteriorkernel.ofkernel build the local posterior-kernel contract from an existing mathlib kernel. definition compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.ofMeasureSelector","label":"ofMeasureSelector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.ofMeasureSelector","description":"Build a posterior kernel from a measurable posterior-measure selector.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-ca60b7009484","parent":"module:BanditRLProof.PosteriorKernel","order":6699,"meta":[["Kind","definition"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:52"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def ofMeasureSelector (posterior : History -> Measure Env) (hposterior : Measurable posterior) (hprob : forall history, IsProbabilityMeasure (posterior history)) : MarkovPosteriorKernel History Env where","missing":[],"search":"ofmeasureselector banditrlproof.posteriorkernel.ofmeasureselector build a posterior kernel from a measurable posterior-measure selector. definition compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.ofMeasureSelector_apply","label":"ofMeasureSelector_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.ofMeasureSelector_apply","description":"theorem ofMeasureSelector_apply (posterior : History -> Measure Env) (hposterior : Measurable posterior) (hprob : forall history, IsProbabilityMeasure (posterior history)) (history : History) : (ofMeasureSelector posterior hposterior hprob).kernel history = posterior history","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-04035144034d","parent":"module:BanditRLProof.PosteriorKernel","order":6700,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:61"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem ofMeasureSelector_apply (posterior : History -> Measure Env) (hposterior : Measurable posterior) (hprob : forall history, IsProbabilityMeasure (posterior history)) (history : History) : (ofMeasureSelector posterior hposterior hprob).kernel history = posterior history","missing":[],"search":"ofmeasureselector_apply banditrlproof.posteriorkernel.ofmeasureselector_apply theorem ofmeasureselector_apply (posterior : history -> measure env) (hposterior : measurable posterior) (hprob : forall history, isprobabilitymeasure (posterior history)) (history : history) : (ofmeasureselector posterior hposterior hprob).kernel history = posterior history theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.ofCountableHistorySelector","label":"ofCountableHistorySelector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.ofCountableHistorySelector","description":"Build a posterior kernel on a countable/discrete history space. Finite histories in the bandit development are typically countable/discrete, so Mathlib's `Kernel.ofFunOfCountable` can turn any probability-valued posterior selector into a Markov kernel without a separate measurability proof.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-f66a4f22c32b","parent":"module:BanditRLProof.PosteriorKernel","order":6701,"meta":[["Kind","definition"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:76"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def ofCountableHistorySelector [Countable History] [MeasurableSingletonClass History] (posterior : History -> Measure Env) (hprob : forall history, IsProbabilityMeasure (posterior history)) : MarkovPosteriorKernel History Env where","missing":[],"search":"ofcountablehistoryselector banditrlproof.posteriorkernel.ofcountablehistoryselector build a posterior kernel on a countable/discrete history space. finite histories in the bandit development are typically countable/discrete, so mathlib's `kernel.offunofcountable` can turn any probability-valued posterior selector into a markov kernel without a separate measurability proof. definition compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.ofCountableHistorySelector_apply","label":"ofCountableHistorySelector_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.ofCountableHistorySelector_apply","description":"theorem ofCountableHistorySelector_apply [Countable History] [MeasurableSingletonClass History] (posterior : History -> Measure Env) (hprob : forall history, IsProbabilityMeasure (posterior history)) (history : History) : (ofCountableHistorySelector posterior hprob).kernel history = posterior history","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-ee086a6e0105","parent":"module:BanditRLProof.PosteriorKernel","order":6702,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:85"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem ofCountableHistorySelector_apply [Countable History] [MeasurableSingletonClass History] (posterior : History -> Measure Env) (hprob : forall history, IsProbabilityMeasure (posterior history)) (history : History) : (ofCountableHistorySelector posterior hprob).kernel history = posterior history","missing":[],"search":"ofcountablehistoryselector_apply banditrlproof.posteriorkernel.ofcountablehistoryselector_apply theorem ofcountablehistoryselector_apply [countable history] [measurablesingletonclass history] (posterior : history -> measure env) (hprob : forall history, isprobabilitymeasure (posterior history)) (history : history) : (ofcountablehistoryselector posterior hprob).kernel history = posterior history theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.measurable_kernel","label":"measurable_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.measurable_kernel","description":"The posterior kernel is measurable as a map from histories to measures.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-836d63c930d0","parent":"module:BanditRLProof.PosteriorKernel","order":6703,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:94"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_kernel (posterior : MarkovPosteriorKernel History Env) : Measurable posterior.kernel","missing":[],"search":"measurable_kernel banditrlproof.posteriorkernel.measurable_kernel the posterior kernel is measurable as a map from histories to measures. theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.measurable_apply_of_measurable_history","label":"measurable_apply_of_measurable_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.measurable_apply_of_measurable_history","description":"A measurable random history selects a measurable random posterior measure.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-50560ada8fd2","parent":"module:BanditRLProof.PosteriorKernel","order":6704,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:100"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_apply_of_measurable_history {Omega : Type u} [MeasurableSpace Omega] (posterior : MarkovPosteriorKernel History Env) (history : Omega -> History) (hhistory : Measurable history) : Measurable (fun omega : Omega => posterior.kernel (history omega))","missing":[],"search":"measurable_apply_of_measurable_history banditrlproof.posteriorkernel.measurable_apply_of_measurable_history a measurable random history selects a measurable random posterior measure. theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.measurable_eventProbability_of_measurable_history","label":"measurable_eventProbability_of_measurable_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.measurable_eventProbability_of_measurable_history","description":"For every measurable environment event, the posterior event probability is a measurable scalar function of a measurable random history.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-b25e8550af2e","parent":"module:BanditRLProof.PosteriorKernel","order":6705,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:112"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_eventProbability_of_measurable_history {Omega : Type u} [MeasurableSpace Omega] (posterior : MarkovPosteriorKernel History Env) (history : Omega -> History) (hhistory : Measurable history) {event : Set Env} (hevent : MeasurableSet event) : Measurable (fun omega : Omega => posterior.kernel (history omega) event)","missing":[],"search":"measurable_eventprobability_of_measurable_history banditrlproof.posteriorkernel.measurable_eventprobability_of_measurable_history for every measurable environment event, the posterior event probability is a measurable scalar function of a measurable random history. theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.isProbabilityMeasure_apply","label":"isProbabilityMeasure_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.isProbabilityMeasure_apply","description":"Every measure selected by a posterior kernel is a probability measure.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-a090c3afc6fd","parent":"module:BanditRLProof.PosteriorKernel","order":6706,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:125"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isProbabilityMeasure_apply (posterior : MarkovPosteriorKernel History Env) (history : History) : IsProbabilityMeasure (posterior.kernel history)","missing":[],"search":"isprobabilitymeasure_apply banditrlproof.posteriorkernel.isprobabilitymeasure_apply every measure selected by a posterior kernel is a probability measure. theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.apply_univ","label":"apply_univ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.apply_univ","description":"theorem apply_univ (posterior : MarkovPosteriorKernel History Env) (history : History) : posterior.kernel history Set.univ = 1","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-75e66db0880b","parent":"module:BanditRLProof.PosteriorKernel","order":6707,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:134"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem apply_univ (posterior : MarkovPosteriorKernel History Env) (history : History) : posterior.kernel history Set.univ = 1","missing":[],"search":"apply_univ banditrlproof.posteriorkernel.apply_univ theorem apply_univ (posterior : markovposteriorkernel history env) (history : history) : posterior.kernel history set.univ = 1 theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior","label":"canonicalPosterior","kind":"definition","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.canonicalPosterior","description":"The Mathlib posterior of a likelihood kernel under a prior measure.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-7a992909c401","parent":"module:BanditRLProof.PosteriorKernel","order":6708,"meta":[["Kind","definition"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:153"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalPosterior [StandardBorelSpace Env] [Nonempty Env] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] : MarkovPosteriorKernel History Env","missing":[],"search":"canonicalposterior banditrlproof.posteriorkernel.canonicalposterior the mathlib posterior of a likelihood kernel under a prior measure. definition compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel","label":"canonicalPosterior_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.canonicalPosterior_kernel","description":"theorem canonicalPosterior_kernel [StandardBorelSpace Env] [Nonempty Env] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] : (canonicalPosterior prior likelihood).kernel = ProbabilityTheory.posterior likelihood prior","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-46c205938bd5","parent":"module:BanditRLProof.PosteriorKernel","order":6709,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:162"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem canonicalPosterior_kernel [StandardBorelSpace Env] [Nonempty Env] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] : (canonicalPosterior prior likelihood).kernel = ProbabilityTheory.posterior likelihood prior","missing":[],"search":"canonicalposterior_kernel banditrlproof.posteriorkernel.canonicalposterior_kernel theorem canonicalposterior_kernel [standardborelspace env] [nonempty env] (prior : measure env) [isprobabilitymeasure prior] (likelihood : probabilitytheory.kernel env history) [probabilitytheory.ismarkovkernel likelihood] : (canonicalposterior prior likelihood).kernel = probabilitytheory.posterior likelihood prior theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.canonicalJointMeasure","label":"canonicalJointMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.canonicalJointMeasure","description":"The canonical joint law generated by a prior followed by a likelihood.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-335be9e2672c","parent":"module:BanditRLProof.PosteriorKernel","order":6710,"meta":[["Kind","definition"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:171"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalJointMeasure (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] : Measure (Env × History)","missing":[],"search":"canonicaljointmeasure banditrlproof.posteriorkernel.canonicaljointmeasure the canonical joint law generated by a prior followed by a likelihood. definition compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq","label":"canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq","description":"Identify the canonical posterior on any source with the prescribed Bayesian environment/history pair law. The proof uses the defining composition-product identity for Mathlib's `posterior`. Thus the posterior conditional-law equality is produced from a joint-law transport, rather than assumed as a separate Bayes-law field.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-104ff246858b","parent":"module:BanditRLProof.PosteriorKernel","order":6711,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:194"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq {Omega : Type*} [MeasurableSpace Omega] [StandardBorelSpace Env] [Nonempty Env] (mu : Measure Omega) [IsFiniteMeasure mu] (env : Omega -> Env) (history : Omega -> History) (henv : Measurable env) (hhistory : Measurable history) (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] (hpair : mu.map (fun omega => (env omega, history omega)) = canonicalJointMeasure prior likelihood) : (canonicalPosterior prior likelihood).kernel =ᵐ[mu.map history] ProbabilityTheory.condDistrib env history mu","missing":[],"search":"canonicalposterior_kernel_ae_eq_conddistrib_of_pair_map_eq banditrlproof.posteriorkernel.canonicalposterior_kernel_ae_eq_conddistrib_of_pair_map_eq identify the canonical posterior on any source with the prescribed bayesian environment/history pair law. the proof uses the defining composition-product identity for mathlib's `posterior`. thus the posterior conditional-law equality is produced from a joint-law transport, rather than assumed as a separate bayes-law field. theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_fst_snd","label":"canonicalPosterior_kernel_ae_eq_condDistrib_fst_snd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_fst_snd","description":"On the canonical Bayesian product space itself, the Mathlib posterior is the conditional distribution of the environment coordinate given the history coordinate.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-9e0ce40d0634","parent":"module:BanditRLProof.PosteriorKernel","order":6712,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:247"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem canonicalPosterior_kernel_ae_eq_condDistrib_fst_snd [StandardBorelSpace Env] [Nonempty Env] (prior : Measure Env) [IsProbabilityMeasure prior] (likelihood : ProbabilityTheory.Kernel Env History) [ProbabilityTheory.IsMarkovKernel likelihood] : (canonicalPosterior prior likelihood).kernel =ᵐ[ (canonicalJointMeasure prior likelihood).map Prod.snd] ProbabilityTheory.condDistrib Prod.fst Prod.snd (canonicalJointMeasure prior likelihood)","missing":[],"search":"canonicalposterior_kernel_ae_eq_conddistrib_fst_snd banditrlproof.posteriorkernel.canonicalposterior_kernel_ae_eq_conddistrib_fst_snd on the canonical bayesian product space itself, the mathlib posterior is the conditional distribution of the environment coordinate given the history coordinate. theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.BayesianPosteriorSurface","label":"BayesianPosteriorSurface","kind":"structure","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.BayesianPosteriorSurface","description":"Minimal prior/likelihood/posterior package. This records the Bayesian objects that future Thompson-sampling leaves need to name. It does not assert that `posterior` satisfies Bayes' rule for `prior` and `likelihood`.","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-5d94ddb5bd0c","parent":"module:BanditRLProof.PosteriorKernel","order":6713,"meta":[["Kind","structure"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:268"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure BayesianPosteriorSurface (Env : Type u) (History : Type v) [MeasurableSpace Env] [MeasurableSpace History] where","missing":[],"search":"bayesianposteriorsurface banditrlproof.posteriorkernel.bayesianposteriorsurface minimal prior/likelihood/posterior package. this records the bayesian objects that future thompson-sampling leaves need to name. it does not assert that `posterior` satisfies bayes' rule for `prior` and `likelihood`. structure compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PosteriorKernel.BayesianPosteriorSurface.posterior_isProbabilityMeasure_apply","label":"posterior_isProbabilityMeasure_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PosteriorKernel.BayesianPosteriorSurface.posterior_isProbabilityMeasure_apply","description":"theorem posterior_isProbabilityMeasure_apply (surface : BayesianPosteriorSurface Env' History') (history : History') : IsProbabilityMeasure (surface.posterior.kernel history)","url":"../modules/banditrlproof-posteriorkernel/index.html#decl-80350593ec96","parent":"module:BanditRLProof.PosteriorKernel","order":6714,"meta":[["Kind","theorem"],["Module","BanditRLProof.PosteriorKernel"],["Source","BanditRLProof/PosteriorKernel.lean:292"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem posterior_isProbabilityMeasure_apply (surface : BayesianPosteriorSurface Env' History') (history : History') : IsProbabilityMeasure (surface.posterior.kernel history)","missing":[],"search":"posterior_isprobabilitymeasure_apply banditrlproof.posteriorkernel.bayesianposteriorsurface.posterior_isprobabilitymeasure_apply theorem posterior_isprobabilitymeasure_apply (surface : bayesianposteriorsurface env' history') (history : history') : isprobabilitymeasure (surface.posterior.kernel history) theorem compiled","shard":"modules/a39773eeb1d9043d.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.PowerTailIntegral.source_cutoff_normalization","label":"source_cutoff_normalization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PowerTailIntegral.source_cutoff_normalization","description":"theorem source_cutoff_normalization (C N γ ω : ℝ) (hC : 0<C) (hN : 0<N) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) : (2/ω)/(2/ω-1)*(N*((C*γ^(2/ω))/N)^(1/(2/ω)))= (2*γ/(2-ω))*C^(ω/2)*N^(1-ω/2)","url":"../modules/banditrlproof-powercutoffnormalization/index.html#decl-a04d740ae0cd","parent":"module:BanditRLProof.PowerCutoffNormalization","order":6715,"meta":[["Kind","theorem"],["Module","BanditRLProof.PowerCutoffNormalization"],["Source","BanditRLProof/PowerCutoffNormalization.lean:7"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem source_cutoff_normalization (C N γ ω : ℝ) (hC : 0<C) (hN : 0<N) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) : (2/ω)/(2/ω-1)*(N*((C*γ^(2/ω))/N)^(1/(2/ω)))= (2*γ/(2-ω))*C^(ω/2)*N^(1-ω/2)","missing":[],"search":"source_cutoff_normalization banditrlproof.powertailintegral.source_cutoff_normalization theorem source_cutoff_normalization (c n γ ω : ℝ) (hc : 0<c) (hn : 0<n) (hγ : 0<γ) (hω : 0<ω) (hω1 : ω≤1) : (2/ω)/(2/ω-1)*(n*((c*γ^(2/ω))/n)^(1/(2/ω)))= (2*γ/(2-ω))*c^(ω/2)*n^(1-ω/2) theorem compiled","shard":"modules/fb515e09d31724d7.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.PowerTailIntegral.integral_power_tail_le","label":"integral_power_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PowerTailIntegral.integral_power_tail_le","description":"theorem integral_power_tail_le (a b q : ℝ) (ha : 0<a) (hab : a≤b) (hq : 1<q) : (∫x in a..b, x^(-q))≤a^(1-q)/(q-1)","url":"../modules/banditrlproof-powertailintegral/index.html#decl-b2a5a9e58a8e","parent":"module:BanditRLProof.PowerTailIntegral","order":6716,"meta":[["Kind","theorem"],["Module","BanditRLProof.PowerTailIntegral"],["Source","BanditRLProof/PowerTailIntegral.lean:8"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_power_tail_le (a b q : ℝ) (ha : 0<a) (hab : a≤b) (hq : 1<q) : (∫x in a..b, x^(-q))≤a^(1-q)/(q-1)","missing":[],"search":"integral_power_tail_le banditrlproof.powertailintegral.integral_power_tail_le theorem integral_power_tail_le (a b q : ℝ) (ha : 0<a) (hab : a≤b) (hq : 1<q) : (∫x in a..b, x^(-q))≤a^(1-q)/(q-1) theorem compiled","shard":"modules/8dd7d79ea72bf61c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.PowerTailIntegral.cutoff_balance","label":"cutoff_balance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PowerTailIntegral.cutoff_balance","description":"theorem cutoff_balance (K N q : ℝ) (hK : 0<K) (hN : 0<N) (hq : 0<q) : K*((K/N)^(1/q))^(-q)=N","url":"../modules/banditrlproof-powertailintegral/index.html#decl-6beaa8d199cb","parent":"module:BanditRLProof.PowerTailIntegral","order":6717,"meta":[["Kind","theorem"],["Module","BanditRLProof.PowerTailIntegral"],["Source","BanditRLProof/PowerTailIntegral.lean:24"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cutoff_balance (K N q : ℝ) (hK : 0<K) (hN : 0<N) (hq : 0<q) : K*((K/N)^(1/q))^(-q)=N","missing":[],"search":"cutoff_balance banditrlproof.powertailintegral.cutoff_balance theorem cutoff_balance (k n q : ℝ) (hk : 0<k) (hn : 0<n) (hq : 0<q) : k*((k/n)^(1/q))^(-q)=n theorem compiled","shard":"modules/8dd7d79ea72bf61c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.PowerTailIntegral.cutoff_objective","label":"cutoff_objective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.PowerTailIntegral.cutoff_objective","description":"theorem cutoff_objective (K N q : ℝ) (hK : 0<K) (hN : 0<N) (hq : 1<q) : N*(K/N)^(1/q)+K*((K/N)^(1/q))^(1-q)/(q-1)= q/(q-1)*(N*(K/N)^(1/q))","url":"../modules/banditrlproof-powertailintegral/index.html#decl-1c056b57d447","parent":"module:BanditRLProof.PowerTailIntegral","order":6718,"meta":[["Kind","theorem"],["Module","BanditRLProof.PowerTailIntegral"],["Source","BanditRLProof/PowerTailIntegral.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem cutoff_objective (K N q : ℝ) (hK : 0<K) (hN : 0<N) (hq : 1<q) : N*(K/N)^(1/q)+K*((K/N)^(1/q))^(1-q)/(q-1)= q/(q-1)*(N*(K/N)^(1/q))","missing":[],"search":"cutoff_objective banditrlproof.powertailintegral.cutoff_objective theorem cutoff_objective (k n q : ℝ) (hk : 0<k) (hn : 0<n) (hq : 1<q) : n*(k/n)^(1/q)+k*((k/n)^(1/q))^(1-q)/(q-1)= q/(q-1)*(n*(k/n)^(1/q)) theorem compiled","shard":"modules/8dd7d79ea72bf61c.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le","label":"measure_biUnion_finset_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le","description":"The measure of a finite union of events is bounded by the finite sum of their measures. This is an outer-measure wrapper around `MeasureTheory.measure_biUnion_finset_le`, so it does not require the events to be measurable and does not require `mu` to be a probability measure.","url":"../modules/banditrlproof-probabilityunionbound/index.html#decl-27f0ec5e816e","parent":"module:BanditRLProof.ProbabilityUnionBound","order":6719,"meta":[["Kind","theorem"],["Module","BanditRLProof.ProbabilityUnionBound"],["Source","BanditRLProof/ProbabilityUnionBound.lean:26"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_biUnion_finset_le {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} (mu : Measure Omega) (s : Finset Idx) (E : Idx -> Set Omega) : mu (⋃ i ∈ s, E i) <= s.sum (fun i => mu (E i))","missing":[],"search":"measure_biunion_finset_le banditrlproof.probabilityunionbound.measure_biunion_finset_le the measure of a finite union of events is bounded by the finite sum of their measures. this is an outer-measure wrapper around `measuretheory.measure_biunion_finset_le`, so it does not require the events to be measurable and does not require `mu` to be a probability measure. theorem compiled","shard":"modules/807bc11a7f36bc43.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le_of_uniform","label":"measure_biUnion_finset_le_of_uniform","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le_of_uniform","description":"Equal-share finite-union bound. Every event receives confidence budget `delta / s.card`; nonemptiness of the index set lets the finite sum normalize back to `delta`. As with `measure_biUnion_finset_le`, no event measurability or probability-measure assumption is required.","url":"../modules/banditrlproof-probabilityunionbound/index.html#decl-8a97f517f5db","parent":"module:BanditRLProof.ProbabilityUnionBound","order":6720,"meta":[["Kind","theorem"],["Module","BanditRLProof.ProbabilityUnionBound"],["Source","BanditRLProof/ProbabilityUnionBound.lean:48"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_biUnion_finset_le_of_uniform {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} [DecidableEq Idx] (mu : Measure Omega) (s : Finset Idx) (hs : s.Nonempty) (delta : Real) (E : Idx -> Set Omega) (hE : forall i, i ∈ s -> mu (E i) <= ENNReal.ofReal (delta / (s.card : Real))) : mu (⋃ i ∈ s, E i) <= ENNReal.ofReal delta","missing":[],"search":"measure_biunion_finset_le_of_uniform banditrlproof.probabilityunionbound.measure_biunion_finset_le_of_uniform equal-share finite-union bound. every event receives confidence budget `delta / s.card`; nonemptiness of the index set lets the finite sum normalize back to `delta`. as with `measure_biunion_finset_le`, no event measurability or probability-measure assumption is required. theorem compiled","shard":"modules/807bc11a7f36bc43.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ProbabilityUnionBound.measure_iUnion_fintype_le_sum","label":"measure_iUnion_fintype_le_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ProbabilityUnionBound.measure_iUnion_fintype_le_sum","description":"Fintype specialization of the finite-union probability bound. The index set is `(Finset.univ : Finset Idx)`, matching the finite-arm style used by the bandit proofs.","url":"../modules/banditrlproof-probabilityunionbound/index.html#decl-dea02896dd4a","parent":"module:BanditRLProof.ProbabilityUnionBound","order":6721,"meta":[["Kind","theorem"],["Module","BanditRLProof.ProbabilityUnionBound"],["Source","BanditRLProof/ProbabilityUnionBound.lean:83"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measure_iUnion_fintype_le_sum {Omega : Type u} [MeasurableSpace Omega] {Idx : Type v} [Fintype Idx] (mu : Measure Omega) (E : Idx -> Set Omega) : mu (⋃ i, E i) <= (Finset.univ : Finset Idx).sum (fun i => mu (E i))","missing":[],"search":"measure_iunion_fintype_le_sum banditrlproof.probabilityunionbound.measure_iunion_fintype_le_sum fintype specialization of the finite-union probability bound. the index set is `(finset.univ : finset idx)`, matching the finite-arm style used by the bandit proofs. theorem compiled","shard":"modules/807bc11a7f36bc43.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.finset_sum_pullCount_eq_time","label":"finset_sum_pullCount_eq_time","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.finset_sum_pullCount_eq_time","description":"The pull counts over a finite action space partition the time horizon. This is the deterministic `PULLCOUNT-SUM-TIME` leaf. It consumes the compiled `pullCount` Finset wrapper and Mathlib's finite fiber-cardinality theorem.","url":"../modules/banditrlproof-pullcountdecomposition/index.html#decl-14e4b175577a","parent":"module:BanditRLProof.PullCountDecomposition","order":6722,"meta":[["Kind","theorem"],["Module","BanditRLProof.PullCountDecomposition"],["Source","BanditRLProof/PullCountDecomposition.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finset_sum_pullCount_eq_time : (Finset.univ : Finset Action).sum (fun a : Action => pullCount action a t) = t","missing":[],"search":"finset_sum_pullcount_eq_time banditrlproof.finset_sum_pullcount_eq_time the pull counts over a finite action space partition the time horizon. this is the deterministic `pullcount-sum-time` leaf. it consumes the compiled `pullcount` finset wrapper and mathlib's finite fiber-cardinality theorem. theorem compiled","shard":"modules/0d00fbec8502d108.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.finset_sum_comp_pullCount","label":"finset_sum_comp_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.finset_sum_comp_pullCount","description":"Reindex a sum over action times by arm and the arm's prior pull count. At a time `s`, the value `pullCount action (action s) s` is the zero-based index of that pull among occurrences of the selected arm. This is the local Mathlib-backed counterpart of LML's `sum_comp_pullCount` bookkeeping lemma.","url":"../modules/banditrlproof-pullcountdecomposition/index.html#decl-cf737654ca24","parent":"module:BanditRLProof.PullCountDecomposition","order":6723,"meta":[["Kind","theorem"],["Module","BanditRLProof.PullCountDecomposition"],["Source","BanditRLProof/PullCountDecomposition.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem finset_sum_comp_pullCount {R : Type v} [AddCommMonoid R] (f : Nat -> R) : ∑ s ∈ Finset.range t, f (pullCount action (action s) s) = ∑ a : Action, ∑ j ∈ Finset.range (pullCount action a t), f j","missing":[],"search":"finset_sum_comp_pullcount banditrlproof.finset_sum_comp_pullcount reindex a sum over action times by arm and the arm's prior pull count. at a time `s`, the value `pullcount action (action s) s` is the zero-based index of that pull among occurrences of the selected arm. this is the local mathlib-backed counterpart of lml's `sum_comp_pullcount` bookkeeping lemma. theorem compiled","shard":"modules/0d00fbec8502d108.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.sum_selected_pullCount","label":"sum_selected_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sum_selected_pullCount","description":"Each selected round visits exactly one successive pre-pull count.","url":"../modules/banditrlproof-pullcountreindex/index.html#decl-37dd52afd06b","parent":"module:BanditRLProof.PullCountReindex","order":6724,"meta":[["Kind","theorem"],["Module","BanditRLProof.PullCountReindex"],["Source","BanditRLProof/PullCountReindex.lean:11"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem sum_selected_pullCount {Action : Type*} [DecidableEq Action] (action : ActionTrace Action) (a : Action) (f : ℕ → ℝ) (T : ℕ) : (∑ t ∈ range T, if action t = a then f (pullCount action a t) else 0) = ∑ s ∈ range (pullCount action a T), f s","missing":[],"search":"sum_selected_pullcount banditrlproof.sum_selected_pullcount each selected round visits exactly one successive pre-pull count. theorem compiled","shard":"modules/709c9ec5e35c3a52.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pullCount_le_one_add_eventCount","label":"pullCount_le_one_add_eventCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pullCount_le_one_add_eventCount","description":"Event-count transport retains the first pull, whose pre-pull count is zero. The event premise must be supplied separately by the algorithm analysis.","url":"../modules/banditrlproof-pullcountreindex/index.html#decl-73a1de76024c","parent":"module:BanditRLProof.PullCountReindex","order":6725,"meta":[["Kind","theorem"],["Module","BanditRLProof.PullCountReindex"],["Source","BanditRLProof/PullCountReindex.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pullCount_le_one_add_eventCount {Action : Type*} [DecidableEq Action] (action : ActionTrace Action) (a : Action) (P : ℕ → Prop) [DecidablePred P] (T : ℕ) (hselected : ∀ t < T, action t = a → 0 < pullCount action a t → P (pullCount action a t)) : (pullCount action a T : ℝ) ≤ 1 + ∑ s ∈ range T, if P (s + 1) then (1 : ℝ) else 0","missing":[],"search":"pullcount_le_one_add_eventcount banditrlproof.pullcount_le_one_add_eventcount event-count transport retains the first pull, whose pre-pull count is zero. the event premise must be supplied separately by the algorithm analysis. theorem compiled","shard":"modules/709c9ec5e35c3a52.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.rawCount","label":"rawCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.rawCount","description":"The uncentered real count selected by a visit or transition coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-5f1aeb8af71a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6726,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def rawCount {mdp : MDP State Action} (coordinate : CountCoordinate mdp) {episodes : Nat} (batch : EpisodeBatch mdp episodes) : Real","missing":[],"search":"rawcount banditrlproof.finitehorizonrl.countcoordinate.rawcount the uncentered real count selected by a visit or transition coordinate. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.policyMean","label":"policyMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.policyMean","description":"The iid batch mean of the selected raw count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-2abca78758f9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6727,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def policyMean {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : Real","missing":[],"search":"policymean banditrlproof.finitehorizonrl.countcoordinate.policymean the iid batch mean of the selected raw count. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.measurable_rawCount","label":"measurable_rawCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.measurable_rawCount","description":"Every selected raw count is measurable on the batch space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-c0321548c5e4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6728,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_rawCount {mdp : MDP State Action} (coordinate : CountCoordinate mdp) {episodes : Nat} : Measurable (coordinate.rawCount : EpisodeBatch mdp episodes -> Real)","missing":[],"search":"measurable_rawcount banditrlproof.finitehorizonrl.countcoordinate.measurable_rawcount every selected raw count is measurable on the batch space. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation_eq_rawCount_sub_policyMean","label":"deviation_eq_rawCount_sub_policyMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation_eq_rawCount_sub_policyMean","description":"The existing coordinate deviation is raw count minus policy mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-2e09a78976cf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6729,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:74"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem deviation_eq_rawCount_sub_policyMean {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (batch : EpisodeBatch mdp episodes) : coordinate.deviation policy initialState batch = coordinate.rawCount batch - coordinate.policyMean policy initialState episodes","missing":[],"search":"deviation_eq_rawcount_sub_policymean banditrlproof.finitehorizonrl.countcoordinate.deviation_eq_rawcount_sub_policymean the existing coordinate deviation is raw count minus policy mean. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.integral_rawCount_iidEpisodeBatchMeasure","label":"integral_rawCount_iidEpisodeBatchMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.integral_rawCount_iidEpisodeBatchMeasure","description":"The kernel integral of a selected raw count is its policy mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-a3f97f6769a1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6730,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:86"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_rawCount_iidEpisodeBatchMeasure {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : integral (policy.iidEpisodeBatchMeasure initialState episodes) coordinate.rawCount = coordinate.policyMean policy initialState episodes","missing":[],"search":"integral_rawcount_iidepisodebatchmeasure banditrlproof.finitehorizonrl.countcoordinate.integral_rawcount_iidepisodebatchmeasure the kernel integral of a selected raw count is its policy mean. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation_hasSubgaussianMGF","label":"deviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation_hasSubgaussianMGF","description":"The selected batch-count deviation has the sharp within-batch Bernoulli proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-6820d0fba480","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6731,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:177"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem deviation_hasSubgaussianMGF {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.HasSubgaussianMGF (coordinate.deviation policy initialState) (MarkovPolicy.iidBernoulliVarianceProxy episodes) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"deviation_hassubgaussianmgf banditrlproof.finitehorizonrl.countcoordinate.deviation_hassubgaussianmgf the selected batch-count deviation has the sharp within-batch bernoulli proxy. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateKernelMean","label":"coordinateKernelMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateKernelMean","description":"Predictable mean of the next selected raw count under the history kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-4a1476539622","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6732,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:221"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def coordinateKernelMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (coordinate : CountCoordinate mdp) (history : EpisodeBatchPrefix mdp episodes n) : Real","missing":[],"search":"coordinatekernelmean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinatekernelmean predictable mean of the next selected raw count under the history kernel. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_coordinateKernelMean","label":"measurable_coordinateKernelMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_coordinateKernelMean","description":"The history-kernel raw-count mean is measurable in the finite prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-aed2f63e6ee2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6733,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_coordinateKernelMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (coordinate : CountCoordinate mdp) : Measurable (source.coordinateKernelMean n coordinate)","missing":[],"search":"measurable_coordinatekernelmean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_coordinatekernelmean the history-kernel raw-count mean is measurable in the finite prefix. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateKernelMean_eq_policyMean","label":"coordinateKernelMean_eq_policyMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateKernelMean_eq_policyMean","description":"The predictable kernel mean equals the selected-policy iid batch mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-132ae4ef5f27","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6734,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:241"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem coordinateKernelMean_eq_policyMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (coordinate : CountCoordinate mdp) (history : EpisodeBatchPrefix mdp episodes n) : source.coordinateKernelMean n coordinate history = coordinate.policyMean (source.successorPolicy n history) initialState episodes","missing":[],"search":"coordinatekernelmean_eq_policymean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinatekernelmean_eq_policymean the predictable kernel mean equals the selected-policy iid batch mean. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinatePrefixIncrement","label":"coordinatePrefixIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinatePrefixIncrement","description":"Count increment at a finite prefix. At successor rounds the center is the measurable history-kernel integral, not the potentially nonmeasurable policy selector exposed by the source structure.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-29d23c89f0cf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6735,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def coordinatePrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) : (round : Nat) -> EpisodeBatchPrefix mdp episodes round -> Real | 0, history => coordinate.rawCount (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩) - coordinate.policyMean source.initialPolicy initialState episodes | n + 1, history => coordinate.rawCount (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) - source.coordinateKernelMean n coordinate (Preorder.frestrictLe₂ (π := fun _ : Nat => EpisodeBatch mdp episodes) (Nat.le_succ n) history) omit [Nonempty State] [Nonempty Action] in /-- Every finite-prefix increment is measurable. -/ theorem measurable_coordinatePrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeas…","missing":[],"search":"coordinateprefixincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateprefixincrement count increment at a finite prefix. at successor rounds the center is the measurable history-kernel integral, not the potentially nonmeasurable policy selector exposed by the source structure. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_coordinatePrefixIncrement","label":"measurable_coordinatePrefixIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_coordinatePrefixIncrement","description":"Every finite-prefix increment is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-171b2c4bc108","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6736,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_coordinatePrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (round : Nat) : Measurable (source.coordinatePrefixIncrement coordinate round)","missing":[],"search":"measurable_coordinateprefixincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_coordinateprefixincrement every finite-prefix increment is measurable. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement","label":"coordinateIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement","description":"The adaptive kernel-centered coordinate increment process.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-bfadb01f4e42","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6737,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:300"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def coordinateIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"coordinateincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateincrement the adaptive kernel-centered coordinate increment process. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_stronglyAdapted_piLE","label":"coordinateIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_stronglyAdapted_piLE","description":"The coordinate increment process is adapted to the canonical prefix filtration.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-f03760c9e7f5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6738,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem coordinateIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes)) (source.coordinateIncrement coordinate)","missing":[],"search":"coordinateincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateincrement_stronglyadapted_pile the coordinate increment process is adapted to the canonical prefix filtration. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_rawCount","label":"trajectoryMeasure_condDistrib_rawCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_rawCount","description":"The next raw-count conditional distribution is the mapped batch kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-193c5b4edad2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6739,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:325"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_rawCount {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (coordinate : CountCoordinate mdp) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => coordinate.rawCount (trajectory (n + 1))) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] (source.batchKernel n).map coordinate.rawCount","missing":[],"search":"trajectorymeasure_conddistrib_rawcount banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conddistrib_rawcount the next raw-count conditional distribution is the mapped batch kernel. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_rawCount_eq_batchKernel","label":"condExpKernel_map_rawCount_eq_batchKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_rawCount_eq_batchKernel","description":"Trimmed conditional-expectation-kernel law for the next adaptive raw count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-1f8caebf64f9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6740,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:371"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_rawCount_eq_batchKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (coordinate : CountCoordinate mdp) : Filter.Eventually (fun trajectory : EpisodeBatchTrajectory mdp episodes => Measure.map (fun path : EpisodeBatchTrajectory mdp episodes => coordinate.rawCount (path (n + 1))) (ProbabilityTheory.condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (EpisodeBatchPrefix mdp episodes n)).comap (Preorder.frestrictLe n)) trajectory) = ((source.batchKernel n).map coordinate.rawCount) (Preorder.frestrictLe n trajectory)) (ae (source.trajectoryMeasure.trim (Preorder.measurable_frestrictLe n).comap_le))","missing":[],"search":"condexpkernel_map_rawcount_eq_batchkernel banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.condexpkernel_map_rawcount_eq_batchkernel trimmed conditional-expectation-kernel law for the next adaptive raw count. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_zero_hasSubgaussianMGF","label":"coordinateIncrement_zero_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_zero_hasSubgaussianMGF","description":"Initial adaptive batch-count increment has the iid Bernoulli sum proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-005da84ba63e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6741,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:409"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem coordinateIncrement_zero_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) : ProbabilityTheory.HasSubgaussianMGF (source.coordinateIncrement coordinate 0) (MarkovPolicy.iidBernoulliVarianceProxy episodes) source.trajectoryMeasure","missing":[],"search":"coordinateincrement_zero_hassubgaussianmgf banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateincrement_zero_hassubgaussianmgf initial adaptive batch-count increment has the iid bernoulli sum proxy. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_succ_hasCondSubgaussianMGF","label":"coordinateIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_succ_hasCondSubgaussianMGF","description":"Every successor adaptive count increment is conditionally sub-Gaussian.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-935ea1e2c35a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6742,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:431"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem coordinateIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (coordinate : CountCoordinate mdp) : ProbabilityTheory.HasCondSubgaussianMGF (Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes) n) ((Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes)).le n) (source.coordinateIncrement coordinate (n + 1)) (MarkovPolicy.iidBernoulliVarianceProxy episodes) source.trajectoryMeasure","missing":[],"search":"coordinateincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateincrement_succ_hascondsubgaussianmgf every successor adaptive count increment is conditionally sub-gaussian. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateVarianceProxy","label":"cumulativeCoordinateVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateVarianceProxy","description":"Total within-batch variance proxy for a finite adaptive prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-7489109d3d02","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6743,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:553"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeCoordinateVarianceProxy (episodes rounds : Nat) : NNReal","missing":[],"search":"cumulativecoordinatevarianceproxy banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinatevarianceproxy total within-batch variance proxy for a finite adaptive prefix. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateConfidenceRadius","label":"cumulativeCoordinateConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateConfidenceRadius","description":"Delta-calibrated radius for one adaptive cumulative count coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-4446129f5876","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6744,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:558"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeCoordinateConfidenceRadius (episodes rounds : Nat) (delta : Real) : Real","missing":[],"search":"cumulativecoordinateconfidenceradius banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinateconfidenceradius delta-calibrated radius for one adaptive cumulative count coordinate. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation","label":"cumulativeCoordinateDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation","description":"Sum of kernel-centered increments over the first `rounds` batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-ea7e076eb727","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6745,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:564"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeCoordinateDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativecoordinatedeviation banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinatedeviation sum of kernel-centered increments over the first `rounds` batches. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_cumulativeCoordinateDeviation","label":"measurable_cumulativeCoordinateDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_cumulativeCoordinateDeviation","description":"Every fixed cumulative coordinate deviation is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-7b49669c2067","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6746,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:575"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeCoordinateDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (rounds : Nat) : Measurable (source.cumulativeCoordinateDeviation coordinate rounds)","missing":[],"search":"measurable_cumulativecoordinatedeviation banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_cumulativecoordinatedeviation every fixed cumulative coordinate deviation is measurable. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_cumulativeCoordinateDeviation_abs_tail_le","label":"trajectoryMeasure_cumulativeCoordinateDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_cumulativeCoordinateDeviation_abs_tail_le","description":"Two-sided square-root prefix tail for one adaptive cumulative count coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-01e689273df9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6747,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:589"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeCoordinateDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectoryMeasure {trajectory | cumulativeCoordinateConfidenceRadius episodes rounds delta <= |source.cumulativeCoordinateDeviation coordinate rounds trajectory|} <= ENNReal.ofReal delta","missing":[],"search":"trajectorymeasure_cumulativecoordinatedeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_cumulativecoordinatedeviation_abs_tail_le two-sided square-root prefix tail for one adaptive cumulative count coordinate. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta","label":"cumulativeCountLocalDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta","description":"Equal confidence share for every queried prefix/coordinate pair.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-d37f956c014a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6748,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:630"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeCountLocalDelta (mdp : MDP State Action) (rounds : Nat) (delta : Real) : Real","missing":[],"search":"cumulativecountlocaldelta banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecountlocaldelta equal confidence share for every queried prefix/coordinate pair. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeCountBadEvent","label":"adaptiveCumulativeCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeCountBadEvent","description":"One global bad event covering every cumulative prefix and count coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-0196ae9fba27","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6749,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:636"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"adaptivecumulativecountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.adaptivecumulativecountbadevent one global bad event covering every cumulative prefix and count coordinate. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_adaptiveCumulativeCountBadEvent","label":"measurableSet_adaptiveCumulativeCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_adaptiveCumulativeCountBadEvent","description":"The finite round-coordinate cumulative bad event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-9ca3c6961197","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6750,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:650"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_adaptiveCumulativeCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : MeasurableSet (source.adaptiveCumulativeCountBadEvent rounds delta)","missing":[],"search":"measurableset_adaptivecumulativecountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurableset_adaptivecumulativecountbadevent the finite round-coordinate cumulative bad event is measurable. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountIndex_nonempty","label":"cumulativeCountIndex_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountIndex_nonempty","description":"The finite round-coordinate family is nonempty at positive horizon/rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-f7f709e5c19c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6751,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:662"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCountIndex_nonempty {mdp : MDP State Action} {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) : Nonempty (Fin rounds × CountCoordinate mdp)","missing":[],"search":"cumulativecountindex_nonempty banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecountindex_nonempty the finite round-coordinate family is nonempty at positive horizon/rounds. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta_pos","label":"cumulativeCountLocalDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta_pos","description":"A positive global delta gives a positive local round-coordinate share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-c92ffb213a8b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6752,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:671"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCountLocalDelta_pos {mdp : MDP State Action} {rounds : Nat} (hindex : Nonempty (Fin rounds × CountCoordinate mdp)) {delta : Real} (hdelta : 0 < delta) : 0 < cumulativeCountLocalDelta mdp rounds delta","missing":[],"search":"cumulativecountlocaldelta_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecountlocaldelta_pos a positive global delta gives a positive local round-coordinate share. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta_le_one","label":"cumulativeCountLocalDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta_le_one","description":"A valid global delta gives every nonempty-family local share at most one.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-9c5a37b9a5a2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6753,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:681"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCountLocalDelta_le_one {mdp : MDP State Action} {rounds : Nat} (hindex : Nonempty (Fin rounds × CountCoordinate mdp)) {delta : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : cumulativeCountLocalDelta mdp rounds delta <= 1","missing":[],"search":"cumulativecountlocaldelta_le_one banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecountlocaldelta_le_one a valid global delta gives every nonempty-family local share at most one. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveCumulativeCountBadEvent_le","label":"trajectoryMeasure_adaptiveCumulativeCountBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveCumulativeCountBadEvent_le","description":"The global adaptive cumulative count event has the requested delta budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-fc05c885fa45","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6754,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:692"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_adaptiveCumulativeCountBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectoryMeasure (source.adaptiveCumulativeCountBadEvent rounds delta) <= ENNReal.ofReal delta","missing":[],"search":"trajectorymeasure_adaptivecumulativecountbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_adaptivecumulativecountbadevent_le the global adaptive cumulative count event has the requested delta budget. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation_abs_lt_of_not_mem_badEvent","label":"cumulativeCoordinateDeviation_abs_lt_of_not_mem_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation_abs_lt_of_not_mem_badEvent","description":"Outside the union, every queried cumulative deviation is strictly small.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-713fe86902f5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6755,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:728"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCoordinateDeviation_abs_lt_of_not_mem_badEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) {delta : Real} {trajectory : EpisodeBatchTrajectory mdp episodes} (htrajectory : trajectory ∉ source.adaptiveCumulativeCountBadEvent rounds delta) (round : Fin rounds) (coordinate : CountCoordinate mdp) : |source.cumulativeCoordinateDeviation coordinate (round + 1) trajectory| < cumulativeCoordinateConfidenceRadius episodes (round + 1) (cumulativeCountLocalDelta mdp rounds delta)","missing":[],"search":"cumulativecoordinatedeviation_abs_lt_of_not_mem_badevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinatedeviation_abs_lt_of_not_mem_badevent outside the union, every queried cumulative deviation is strictly small. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateMeanAt","label":"coordinateMeanAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateMeanAt","description":"Predictable center used at one adaptive batch coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-3abd646b9db9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6756,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:744"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def coordinateMeanAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) : Nat -> EpisodeBatchTrajectory mdp episodes -> Real | 0, _trajectory => coordinate.policyMean source.initialPolicy initialState episodes | n + 1, trajectory => source.coordinateKernelMean n coordinate (Preorder.frestrictLe n trajectory) /-- Sum of predictable coordinate means over a finite prefix. -/ noncomputable def cumulativeCoordinateMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"coordinatemeanat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinatemeanat predictable center used at one adaptive batch coordinate. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateMean","label":"cumulativeCoordinateMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateMean","description":"Sum of predictable coordinate means over a finite prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-943bdf056803","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6757,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:757"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeCoordinateMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativecoordinatemean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinatemean sum of predictable coordinate means over a finite prefix. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount","label":"cumulativeCoordinateRawCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount","description":"Sum of uncentered coordinate counts over a finite prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-5a550d6d7366","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6758,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:767"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeCoordinateRawCount {mdp : MDP State Action} {episodes : Nat} (coordinate : CountCoordinate mdp) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativecoordinaterawcount banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinaterawcount sum of uncentered coordinate counts over a finite prefix. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_eq_rawCount_sub_meanAt","label":"coordinateIncrement_eq_rawCount_sub_meanAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_eq_rawCount_sub_meanAt","description":"Each adaptive increment is raw count minus its predictable mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-3ab4e772b5eb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6759,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:775"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem coordinateIncrement_eq_rawCount_sub_meanAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : source.coordinateIncrement coordinate round trajectory = coordinate.rawCount (trajectory round) - source.coordinateMeanAt coordinate round trajectory","missing":[],"search":"coordinateincrement_eq_rawcount_sub_meanat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateincrement_eq_rawcount_sub_meanat each adaptive increment is raw count minus its predictable mean. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation_eq_rawCount_sub_mean","label":"cumulativeCoordinateDeviation_eq_rawCount_sub_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation_eq_rawCount_sub_mean","description":"The cumulative martingale deviation is raw count minus cumulative mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-bcd9a862dd5b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6760,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:791"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCoordinateDeviation_eq_rawCount_sub_mean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (coordinate : CountCoordinate mdp) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : source.cumulativeCoordinateDeviation coordinate rounds trajectory = cumulativeCoordinateRawCount coordinate rounds trajectory - source.cumulativeCoordinateMean coordinate rounds trajectory","missing":[],"search":"cumulativecoordinatedeviation_eq_rawcount_sub_mean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinatedeviation_eq_rawcount_sub_mean the cumulative martingale deviation is raw count minus cumulative mean. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount_visit","label":"cumulativeCoordinateRawCount_visit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount_visit","description":"Cumulative visit raw counts are exactly the summary's visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-fea22858e3ce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6761,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:818"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCoordinateRawCount_visit {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : cumulativeCoordinateRawCount (.visit stage state action) (round + 1) trajectory = ((cumulativeTransitionCountSummaryAt trajectory round).visitCount stage state action : Real)","missing":[],"search":"cumulativecoordinaterawcount_visit banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinaterawcount_visit cumulative visit raw counts are exactly the summary's visit count. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount_transition","label":"cumulativeCoordinateRawCount_transition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount_transition","description":"Cumulative transition raw counts are exactly the summary coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-97cbdc080d5d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6762,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:834"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCoordinateRawCount_transition {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : cumulativeCoordinateRawCount (.transition stage state action nextState) (round + 1) trajectory = (cumulativeTransitionCountSummaryAt trajectory round stage state action nextState : Real)","missing":[],"search":"cumulativecoordinaterawcount_transition banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinaterawcount_transition cumulative transition raw counts are exactly the summary coordinate. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateMeanAt_transition_eq_visit_mul_transition","label":"coordinateMeanAt_transition_eq_visit_mul_transition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateMeanAt_transition_eq_visit_mul_transition","description":"Each transition-count predictable mean factors through its visit mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-476cc95ab846","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6763,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:849"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem coordinateMeanAt_transition_eq_visit_mul_transition {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : source.coordinateMeanAt (.transition stage state action nextState) round trajectory = source.coordinateMeanAt (.visit stage state action) round trajectory * (mdp.transition (state, action)).real {nextState}","missing":[],"search":"coordinatemeanat_transition_eq_visit_mul_transition banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinatemeanat_transition_eq_visit_mul_transition each transition-count predictable mean factors through its visit mean. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateMean_transition_eq_visit_mul_transition","label":"cumulativeCoordinateMean_transition_eq_visit_mul_transition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateMean_transition_eq_visit_mul_transition","description":"Cumulative transition centers factor through the cumulative visit center.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-3d131c2adebb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6764,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:877"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCoordinateMean_transition_eq_visit_mul_transition {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : source.cumulativeCoordinateMean (.transition stage state action nextState) rounds trajectory = source.cumulativeCoordinateMean (.visit stage state action) rounds trajectory * (mdp.transition (state, action)).real {nextState}","missing":[],"search":"cumulativecoordinatemean_transition_eq_visit_mul_transition banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinatemean_transition_eq_visit_mul_transition cumulative transition centers factor through the cumulative visit center. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeEmpiricalTransitionMass_abs_sub_transition_lt","label":"cumulativeEmpiricalTransitionMass_abs_sub_transition_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeEmpiricalTransitionMass_abs_sub_transition_lt","description":"Positive cumulative visits convert the two martingale deviations into a random-denominator empirical-transition singleton bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-c9a669c4a4d7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6765,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:900"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeEmpiricalTransitionMass_abs_sub_transition_lt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Fin rounds) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (radius : Real) (hvisitPos : 0 < (cumulativeTransitionCountSummaryAt trajectory round).visitCount stage state action) (hvisitDeviation : |((cumulativeTransitionCountSummaryAt trajectory round).visitCount stage state action : Real) - source.cumulativeCoordinateMean (.visit stage state action) (round + 1) trajectory| < radius) (htransitionDeviation : |(cumulativeTransitionCountSummaryAt trajectory round stage state action nextState : Real) - source.cumulativeCoordinateMean (.tr…","missing":[],"search":"cumulativeempiricaltransitionmass_abs_sub_transition_lt banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeempiricaltransitionmass_abs_sub_transition_lt positive cumulative visits convert the two martingale deviations into a random-denominator empirical-transition singleton bound. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeTransitionCoordinateRadius","label":"adaptiveCumulativeTransitionCoordinateRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeTransitionCoordinateRadius","description":"Coordinate radius obtained from the cumulative visit denominator. The zero visit branch uses the trivial probability-mass bound; the positive branch uses the paired visit and transition martingale deviations.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-3c704eb7e3c3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6766,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:986"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeTransitionCoordinateRadius {mdp : MDP State Action} {episodes rounds : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Fin rounds) (delta : Real) (stage : Fin mdp.horizon) (state : State) (action : Action) (_nextState : State) : Real","missing":[],"search":"adaptivecumulativetransitioncoordinateradius banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.adaptivecumulativetransitioncoordinateradius coordinate radius obtained from the cumulative visit denominator. the zero visit branch uses the trivial probability-mass bound; the positive branch uses the paired visit and transition martingale deviations. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.AdaptiveCumulativeCountMartingaleCover","label":"AdaptiveCumulativeCountMartingaleCover","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.AdaptiveCumulativeCountMartingaleCover","description":"Regularity contract connecting the statistical coordinate radius to the planner's count radius after multiplication by the recursive value envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-a03904102b80","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6767,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:1005"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def AdaptiveCumulativeCountMartingaleCover {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (countRadius : TransitionCountRadius) (delta rewardBound : Real) : Prop","missing":[],"search":"adaptivecumulativecountmartingalecover banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.adaptivecumulativecountmartingalecover regularity contract connecting the statistical coordinate radius to the planner's count radius after multiplication by the recursive value envelope. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateConfidence_of_not_mem_adaptiveCumulativeCountBadEvent","label":"coordinateConfidence_of_not_mem_adaptiveCumulativeCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateConfidence_of_not_mem_adaptiveCumulativeCountBadEvent","description":"Outside the cumulative count-martingale event, the cumulative empirical plan has the finite-coordinate confidence required by the optimistic Bellman route.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-d22a596aa0a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6768,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:1029"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def coordinateConfidence_of_not_mem_adaptiveCumulativeCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) {delta : Real} {trajectory : EpisodeBatchTrajectory mdp episodes} (htrajectory : trajectory ∉ source.adaptiveCumulativeCountBadEvent rounds delta) (round : Fin rounds) (defaultState : State) (countRadius : TransitionCountRadius) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hcover : forall (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action), (∑ nextState, adaptiveCumulativeTransitionCoordinateRadius trajectory round delta (mdp.decisionStageRemaining remaining hremaining) state action nextState * empiricalFiniteBatchValueEnvelope rew…","missing":[],"search":"coordinateconfidence_of_not_mem_adaptivecumulativecountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.coordinateconfidence_of_not_mem_adaptivecumulativecountbadevent outside the cumulative count-martingale event, the cumulative empirical plan has the finite-coordinate confidence required by the optimistic bellman route. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeCoordinateConfidenceContract_of_martingale","label":"adaptiveCumulativeCoordinateConfidenceContract_of_martingale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeCoordinateConfidenceContract_of_martingale","description":"The cumulative count martingale event and cover produce the reusable global coordinate-confidence contract for cumulative optimistic recommendations.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-646c083cb02f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6769,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:1124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeCoordinateConfidenceContract_of_martingale {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (defaultState : State) (countRadius : TransitionCountRadius) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcover : AdaptiveCumulativeCountMartingaleCover (rounds := rounds) source countRadius delta rewardBound) : AdaptiveCumulativeCoordinateConfidenceContract source defaultState countRadius rounds delta where","missing":[],"search":"adaptivecumulativecoordinateconfidencecontract_of_martingale banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.adaptivecumulativecoordinateconfidencecontract_of_martingale the cumulative count martingale event and cover produce the reusable global coordinate-confidence contract for cumulative optimistic recommendations. definition compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveCumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","label":"trajectoryMeasure_adaptiveCumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveCumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","description":"Generic route endpoint: cumulative count martingales supply the probability producer for optimism and explicit recommended-policy expected regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-e97c650683b1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6770,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:1157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_adaptiveCumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (defaultState : State) (countRadius : TransitionCountRadius) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcover : AdaptiveCumulativeCountMartingaleCover (rounds := rounds) source countRadius delta rewardBound) (radiusEnvelope : Fin rounds -> Real) (hradius : forall trajectory, trajectory ∉ source.adapt…","missing":[],"search":"trajectorymeasure_adaptivecumulativecountmartingale_optimism_and_explicitrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_adaptivecumulativecountmartingale_optimism_and_explicitrecommendedexpectedregret generic route endpoint: cumulative count martingales supply the probability producer for optimism and explicit recommended-policy expected regret. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","description":"Concrete exploratory-source endpoint for the cumulative count-martingale confidence route. The conclusion concerns recommended policies, not behavior or realized regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativecountmartingaleconfidence/index.html#decl-06730fabf8aa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","order":6771,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence.lean:1216"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hcover : AdaptiveEpisodeBatchSource.AdaptiveCumulativeCountMartingaleCover (rounds := rounds) (exploratorySource mdp initialState e…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativecountmartingale_optimism_and_explicitrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativecountmartingale_optimism_and_explicitrecommendedexpectedregret concrete exploratory-source endpoint for the cumulative count-martingale confidence route. the conclusion concerns recommended policies, not behavior or realized regret. theorem compiled","shard":"modules/d537a32c5a68b945.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor_eq_rate_mul_one","label":"exploratoryActionProbabilityFloor_eq_rate_mul_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor_eq_rate_mul_one","description":"theorem exploratoryActionProbabilityFloor_eq_rate_mul_one (explorationRate : NNReal) : exploratoryActionProbabilityFloor Action explorationRate = (explorationRate : Real) * exploratoryActionProbabilityFloor Action 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-1250f91ca622","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6772,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryActionProbabilityFloor_eq_rate_mul_one (explorationRate : NNReal) : exploratoryActionProbabilityFloor Action explorationRate = (explorationRate : Real) * exploratoryActionProbabilityFloor Action 1","missing":[],"search":"exploratoryactionprobabilityfloor_eq_rate_mul_one banditrlproof.finitehorizonrl.exploratoryactionprobabilityfloor_eq_rate_mul_one theorem exploratoryactionprobabilityfloor_eq_rate_mul_one (explorationrate : nnreal) : exploratoryactionprobabilityfloor action explorationrate = (explorationrate : real) * exploratoryactionprobabilityfloor action 1 theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat_eq_rate_pow_mul_one","label":"exploratoryPathStateLowerNat_eq_rate_pow_mul_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat_eq_rate_pow_mul_one","description":"theorem exploratoryPathStateLowerNat_eq_rate_pow_mul_one {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) : forall (stage : Nat) (hstage : stage < mdp.horizon) (state : State), exploratoryPathStateLowerNat support explorationRate stage hstage state = (explorationRate : Real) ^ stage * exploratoryPathStateLowerNat support 1 stage hs…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-46f863f74585","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6773,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPathStateLowerNat_eq_rate_pow_mul_one {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) : forall (stage : Nat) (hstage : stage < mdp.horizon) (state : State), exploratoryPathStateLowerNat support explorationRate stage hstage state = (explorationRate : Real) ^ stage * exploratoryPathStateLowerNat support 1 stage hstage state","missing":[],"search":"exploratorypathstatelowernat_eq_rate_pow_mul_one banditrlproof.finitehorizonrl.exploratorypathstatelowernat_eq_rate_pow_mul_one theorem exploratorypathstatelowernat_eq_rate_pow_mul_one {mdp : mdp state action} {initialstate : measure state} (support : exploratorypathsupport mdp initialstate) (explorationrate : nnreal) : forall (stage : nat) (hstage : stage < mdp.horizon) (state : state), exploratorypathstatelowernat support explorationrate stage hstage state = (explorationrate : real) ^ stage * exploratorypathstatelowernat support 1 stage hstage state theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathVisitLower_eq_rate_pow_mul_one","label":"exploratoryPathVisitLower_eq_rate_pow_mul_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathVisitLower_eq_rate_pow_mul_one","description":"theorem exploratoryPathVisitLower_eq_rate_pow_mul_one {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (stage : Fin mdp.horizon) (state : State) : exploratoryPathStateLower support explorationRate stage state * exploratoryActionProbabilityFloor Action explorationRate = (explorationRate : Real) ^ (stage.val + 1) * (exploratoryPathSt…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-0fa7ecc98320","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6774,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPathVisitLower_eq_rate_pow_mul_one {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (stage : Fin mdp.horizon) (state : State) : exploratoryPathStateLower support explorationRate stage state * exploratoryActionProbabilityFloor Action explorationRate = (explorationRate : Real) ^ (stage.val + 1) * (exploratoryPathStateLower support 1 stage state * exploratoryActionProbabilityFloor Action 1)","missing":[],"search":"exploratorypathvisitlower_eq_rate_pow_mul_one banditrlproof.finitehorizonrl.exploratorypathvisitlower_eq_rate_pow_mul_one theorem exploratorypathvisitlower_eq_rate_pow_mul_one {mdp : mdp state action} {initialstate : measure state} (support : exploratorypathsupport mdp initialstate) (explorationrate : nnreal) (stage : fin mdp.horizon) (state : state) : exploratorypathstatelower support explorationrate stage state * exploratoryactionprobabilityfloor action explorationrate = (explorationrate : real) ^ (stage.val + 1) * (exploratorypathstatelower support 1 stage state * exploratoryactionprobabilityfloor action 1) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor.scale_explorationRate","label":"scale_explorationRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor.scale_explorationRate","description":"theorem ExploratoryPathUniformVisitFloor.scale_explorationRate {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) {baseVisitFloor : Real} (hfloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : ExploratoryPathUniformVisitFloor support explorationRate (baseVisitFloor * (explorat…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-88770050c861","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6775,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ExploratoryPathUniformVisitFloor.scale_explorationRate {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) {baseVisitFloor : Real} (hfloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : ExploratoryPathUniformVisitFloor support explorationRate (baseVisitFloor * (explorationRate : Real) ^ mdp.horizon)","missing":[],"search":"scale_explorationrate banditrlproof.finitehorizonrl.exploratorypathuniformvisitfloor.scale_explorationrate theorem exploratorypathuniformvisitfloor.scale_explorationrate {mdp : mdp state action} {initialstate : measure state} (support : exploratorypathsupport mdp initialstate) {basevisitfloor : real} (hfloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) : exploratorypathuniformvisitfloor support explorationrate (basevisitfloor * (explorationrate : real) ^ mdp.horizon) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScale","label":"decayingExplorationScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScale","description":"def decayingExplorationScale (n : Nat) : Nat","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-e3f8d7fd2231","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6776,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:135"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def decayingExplorationScale (n : Nat) : Nat","missing":[],"search":"decayingexplorationscale banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationscale def decayingexplorationscale (n : nat) : nat definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate","label":"decayingExplorationRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate","description":"noncomputable def decayingExplorationRate (n : Nat) : NNReal","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-45644eeb6dba","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6777,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:137"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationRate (n : Nat) : NNReal","missing":[],"search":"decayingexplorationrate banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrate noncomputable def decayingexplorationrate (n : nat) : nnreal definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRounds","label":"decayingExplorationRounds","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRounds","description":"def decayingExplorationRounds (mdp : MDP State Action) (n : Nat) : Nat","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-36d3b5306ca4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6778,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def decayingExplorationRounds (mdp : MDP State Action) (n : Nat) : Nat","missing":[],"search":"decayingexplorationrounds banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrounds def decayingexplorationrounds (mdp : mdp state action) (n : nat) : nat definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor","label":"decayingExplorationVisitFloor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor","description":"noncomputable def decayingExplorationVisitFloor (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-acbcfa8641d5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6779,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationVisitFloor (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationvisitfloor banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationvisitfloor noncomputable def decayingexplorationvisitfloor (mdp : mdp state action) (basevisitfloor : real) (n : nat) : real definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes","label":"decayingExplorationScheduledEpisodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes","description":"noncomputable def decayingExplorationScheduledEpisodes (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Nat","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-55d2b31a20d3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6780,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationScheduledEpisodes (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Nat","missing":[],"search":"decayingexplorationscheduledepisodes banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationscheduledepisodes noncomputable def decayingexplorationscheduledepisodes (mdp : mdp state action) (basevisitfloor : real) (n : nat) : nat definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRecommendedExpectedRegretBound","label":"decayingExplorationAverageRecommendedExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRecommendedExpectedRegretBound","description":"noncomputable def decayingExplorationAverageRecommendedExpectedRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-4a64292f7a15","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6781,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:153"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageRecommendedExpectedRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationaveragerecommendedexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaveragerecommendedexpectedregretbound noncomputable def decayingexplorationaveragerecommendedexpectedregretbound (mdp : mdp state action) (basevisitfloor : real) (n : nat) : real definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageExploratoryBehaviorExpectedRegretBound","label":"decayingExplorationAverageExploratoryBehaviorExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageExploratoryBehaviorExpectedRegretBound","description":"noncomputable def decayingExplorationAverageExploratoryBehaviorExpectedRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-6b94f58fd276","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6782,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageExploratoryBehaviorExpectedRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationaverageexploratorybehaviorexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaverageexploratorybehaviorexpectedregretbound noncomputable def decayingexplorationaverageexploratorybehaviorexpectedregretbound (mdp : mdp state action) (basevisitfloor : real) (n : nat) : real definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageEnvelope","label":"decayingExplorationAverageEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageEnvelope","description":"noncomputable def decayingExplorationAverageEnvelope (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-46bcb2ab8935","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6783,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageEnvelope (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationaverageenvelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaverageenvelope noncomputable def decayingexplorationaverageenvelope (mdp : mdp state action) (basevisitfloor : real) (n : nat) : real definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScale_pos","label":"decayingExplorationScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScale_pos","description":"theorem decayingExplorationScale_pos (n : Nat) : 0 < decayingExplorationScale n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-b54d204ebc9f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6784,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:175"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationScale_pos (n : Nat) : 0 < decayingExplorationScale n","missing":[],"search":"decayingexplorationscale_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationscale_pos theorem decayingexplorationscale_pos (n : nat) : 0 < decayingexplorationscale n theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate_pos","label":"decayingExplorationRate_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate_pos","description":"theorem decayingExplorationRate_pos (n : Nat) : 0 < decayingExplorationRate n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-d546d68ce784","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6785,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationRate_pos (n : Nat) : 0 < decayingExplorationRate n","missing":[],"search":"decayingexplorationrate_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrate_pos theorem decayingexplorationrate_pos (n : nat) : 0 < decayingexplorationrate n theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate_le_one","label":"decayingExplorationRate_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate_le_one","description":"theorem decayingExplorationRate_le_one (n : Nat) : decayingExplorationRate n <= 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-dc91e2282d96","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6786,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:191"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationRate_le_one (n : Nat) : decayingExplorationRate n <= 1","missing":[],"search":"decayingexplorationrate_le_one banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrate_le_one theorem decayingexplorationrate_le_one (n : nat) : decayingexplorationrate n <= 1 theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRounds_pos","label":"decayingExplorationRounds_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRounds_pos","description":"theorem decayingExplorationRounds_pos (mdp : MDP State Action) (n : Nat) : 0 < decayingExplorationRounds mdp n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-36bacb5d781f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6787,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:201"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationRounds_pos (mdp : MDP State Action) (n : Nat) : 0 < decayingExplorationRounds mdp n","missing":[],"search":"decayingexplorationrounds_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrounds_pos theorem decayingexplorationrounds_pos (mdp : mdp state action) (n : nat) : 0 < decayingexplorationrounds mdp n theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor_pos","label":"decayingExplorationVisitFloor_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor_pos","description":"theorem decayingExplorationVisitFloor_pos (mdp : MDP State Action) {baseVisitFloor : Real} (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 < decayingExplorationVisitFloor mdp baseVisitFloor n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-2ef459d6e333","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6788,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationVisitFloor_pos (mdp : MDP State Action) {baseVisitFloor : Real} (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 < decayingExplorationVisitFloor mdp baseVisitFloor n","missing":[],"search":"decayingexplorationvisitfloor_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationvisitfloor_pos theorem decayingexplorationvisitfloor_pos (mdp : mdp state action) {basevisitfloor : real} (hbasevisitfloor : 0 < basevisitfloor) (n : nat) : 0 < decayingexplorationvisitfloor mdp basevisitfloor n theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor_mul_rounds","label":"decayingExplorationVisitFloor_mul_rounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor_mul_rounds","description":"theorem decayingExplorationVisitFloor_mul_rounds (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : decayingExplorationVisitFloor mdp baseVisitFloor n * (decayingExplorationRounds mdp n : Real) = baseVisitFloor * (decayingExplorationScale n : Real) ^ 4","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-37548fdd547e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6789,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:220"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationVisitFloor_mul_rounds (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : decayingExplorationVisitFloor mdp baseVisitFloor n * (decayingExplorationRounds mdp n : Real) = baseVisitFloor * (decayingExplorationScale n : Real) ^ 4","missing":[],"search":"decayingexplorationvisitfloor_mul_rounds banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationvisitfloor_mul_rounds theorem decayingexplorationvisitfloor_mul_rounds (mdp : mdp state action) (basevisitfloor : real) (n : nat) : decayingexplorationvisitfloor mdp basevisitfloor n * (decayingexplorationrounds mdp n : real) = basevisitfloor * (decayingexplorationscale n : real) ^ 4 theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationUniformVisitFloor","label":"decayingExplorationUniformVisitFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationUniformVisitFloor","description":"theorem decayingExplorationUniformVisitFloor {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) {baseVisitFloor : Real} (hfloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (n : Nat) : ExploratoryPathUniformVisitFloor support (decayingExplorationRate n) (decayingExplorationVisitFloor mdp baseVisitFloor n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-394c96e305ad","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6790,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationUniformVisitFloor {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) {baseVisitFloor : Real} (hfloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (n : Nat) : ExploratoryPathUniformVisitFloor support (decayingExplorationRate n) (decayingExplorationVisitFloor mdp baseVisitFloor n)","missing":[],"search":"decayingexplorationuniformvisitfloor banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationuniformvisitfloor theorem decayingexplorationuniformvisitfloor {mdp : mdp state action} {initialstate : measure state} (support : exploratorypathsupport mdp initialstate) {basevisitfloor : real} (hfloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (n : nat) : exploratorypathuniformvisitfloor support (decayingexplorationrate n) (decayingexplorationvisitfloor mdp basevisitfloor n) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedScheduledAverageEnvelope_decayingExploration_eq","label":"normalizedScheduledAverageEnvelope_decayingExploration_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedScheduledAverageEnvelope_decayingExploration_eq","description":"theorem normalizedScheduledAverageEnvelope_decayingExploration_eq (mdp : MDP State Action) {baseVisitFloor : Real} (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : normalizedCumulativeInverseSqrtScheduledAverageEnvelope mdp (decayingExplorationRounds mdp n) (decayingExplorationVisitFloor mdp baseVisitFloor n) = decayingExplorationAverageEnvelope mdp baseVisitFloor n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-f6c0fbbeaba2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6791,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:265"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedScheduledAverageEnvelope_decayingExploration_eq (mdp : MDP State Action) {baseVisitFloor : Real} (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : normalizedCumulativeInverseSqrtScheduledAverageEnvelope mdp (decayingExplorationRounds mdp n) (decayingExplorationVisitFloor mdp baseVisitFloor n) = decayingExplorationAverageEnvelope mdp baseVisitFloor n","missing":[],"search":"normalizedscheduledaverageenvelope_decayingexploration_eq banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedscheduledaverageenvelope_decayingexploration_eq theorem normalizedscheduledaverageenvelope_decayingexploration_eq (mdp : mdp state action) {basevisitfloor : real} (hbasevisitfloor : 0 < basevisitfloor) (n : nat) : normalizedcumulativeinversesqrtscheduledaverageenvelope mdp (decayingexplorationrounds mdp n) (decayingexplorationvisitfloor mdp basevisitfloor n) = decayingexplorationaverageenvelope mdp basevisitfloor n theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageEnvelope_tendsto_zero","label":"decayingExplorationAverageEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageEnvelope_tendsto_zero","description":"theorem decayingExplorationAverageEnvelope_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) : Tendsto (fun n => decayingExplorationAverageEnvelope mdp baseVisitFloor n) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-d2b470bc7203","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6792,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:295"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationAverageEnvelope_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) : Tendsto (fun n => decayingExplorationAverageEnvelope mdp baseVisitFloor n) atTop (nhds 0)","missing":[],"search":"decayingexplorationaverageenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaverageenvelope_tendsto_zero theorem decayingexplorationaverageenvelope_tendsto_zero (mdp : mdp state action) (basevisitfloor : real) : tendsto (fun n => decayingexplorationaverageenvelope mdp basevisitfloor n) attop (nhds 0) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationBehaviorCharge_tendsto_zero","label":"decayingExplorationBehaviorCharge_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationBehaviorCharge_tendsto_zero","description":"theorem decayingExplorationBehaviorCharge_tendsto_zero (mdp : MDP State Action) : Tendsto (fun n => exploratoryBehaviorRegretCharge mdp (decayingExplorationRate n) 1) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-9566636f859c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6793,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:315"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationBehaviorCharge_tendsto_zero (mdp : MDP State Action) : Tendsto (fun n => exploratoryBehaviorRegretCharge mdp (decayingExplorationRate n) 1) atTop (nhds 0)","missing":[],"search":"decayingexplorationbehaviorcharge_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationbehaviorcharge_tendsto_zero theorem decayingexplorationbehaviorcharge_tendsto_zero (mdp : mdp state action) : tendsto (fun n => exploratorybehaviorregretcharge mdp (decayingexplorationrate n) 1) attop (nhds 0) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageExploratoryBehaviorBound_tendsto_zero","label":"decayingExplorationAverageExploratoryBehaviorBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageExploratoryBehaviorBound_tendsto_zero","description":"theorem decayingExplorationAverageExploratoryBehaviorBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => decayingExplorationAverageExploratoryBehaviorExpectedRegretBound mdp baseVisitFloor n) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-269f731ccdc5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6794,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:334"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationAverageExploratoryBehaviorBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => decayingExplorationAverageExploratoryBehaviorExpectedRegretBound mdp baseVisitFloor n) atTop (nhds 0)","missing":[],"search":"decayingexplorationaverageexploratorybehaviorbound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaverageexploratorybehaviorbound_tendsto_zero theorem decayingexplorationaverageexploratorybehaviorbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (fun n => decayingexplorationaverageexploratorybehaviorexpectedregretbound mdp basevisitfloor n) attop (nhds 0) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationDeltaAndAverageExploratoryBehaviorBound_tendsto_zero","label":"decayingExplorationDeltaAndAverageExploratoryBehaviorBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationDeltaAndAverageExploratoryBehaviorBound_tendsto_zero","description":"theorem decayingExplorationDeltaAndAverageExploratoryBehaviorBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (ENNReal.ofReal (vanishingAverageConfidenceDelta n), decayingExplorationAverageExploratoryBehaviorExpectedRegretBound mdp baseVisitFloor n)) atTop (nhds (0, 0))","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-233b3ded2bbd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6795,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:372"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationDeltaAndAverageExploratoryBehaviorBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (ENNReal.ofReal (vanishingAverageConfidenceDelta n), decayingExplorationAverageExploratoryBehaviorExpectedRegretBound mdp baseVisitFloor n)) atTop (nhds (0, 0))","missing":[],"search":"decayingexplorationdeltaandaverageexploratorybehaviorbound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationdeltaandaverageexploratorybehaviorbound_tendsto_zero theorem decayingexplorationdeltaandaverageexploratorybehaviorbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (fun n => (ennreal.ofreal (vanishingaverageconfidencedelta n), decayingexplorationaverageexploratorybehaviorexpectedregretbound mdp basevisitfloor n)) attop (nhds (0, 0)) theorem compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationAverageExploratoryBehaviorRegretViolationSet","label":"decayingExplorationAverageExploratoryBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationAverageExploratoryBehaviorRegretViolationSet","description":"noncomputable def decayingExplorationAverageExploratoryBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Set (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-55739c21f774","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6796,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:390"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageExploratoryBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Set (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","missing":[],"search":"decayingexplorationaverageexploratorybehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationaverageexploratorybehaviorregretviolationset noncomputable def decayingexplorationaverageexploratorybehaviorregretviolationset (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (defaultstate : state) (basevisitfloor : real) (n : nat) : set (episodebatchtrajectory mdp (adaptiveepisodebatchsource.decayingexplorationscheduledepisodes mdp basevisitfloor n)) definition compiled","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","description":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSp…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationbehaviorconsistency/index.html#decl-7e09173c566d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","order":6797,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency.lean:412"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := Adapti…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageexploratorybehaviorexpectedregret theorem exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageexploratorybehaviorexpectedregret (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (basevisitfloor : real) (n : nat) [standardborelspace (episodebatch mdp (adaptiveepisodebatchsource.decayingexplorationscheduledepisodes mdp basevisitfloor n))] [standardborelspace (episodebatchtrajectory mdp (adaptiveepisodebatchsource.decayingexplorationscheduledepisodes mdp basevisitfloor n))] (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : let rounds := adaptiveepisodebatchsource.decayingexplorationrounds mdp n let delta := adaptiveepisodebatchsource.vanishingaverageconfidencedelta n let exploration…","shard":"modules/c2c0349d2431cbfc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.batchReturnVarianceProxy_coe","label":"batchReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.batchReturnVarianceProxy_coe","description":"The coarse bounded-return proxy is exactly the square of the batch range.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-e1d7bb782807","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6798,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem batchReturnVarianceProxy_coe (mdp : MDP State Action) (episodes : Nat) : ((batchReturnVarianceProxy mdp episodes : NNReal) : Real) = ((episodes : Real) * (mdp.horizon : Real)) ^ 2","missing":[],"search":"batchreturnvarianceproxy_coe banditrlproof.finitehorizonrl.markovpolicy.batchreturnvarianceproxy_coe the coarse bounded-return proxy is exactly the square of the batch range. theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnVarianceProxy_coe","label":"cumulativeSuccessorReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnVarianceProxy_coe","description":"The successor-return proxy contains exactly `rounds` nonzero batch terms.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-75756ef8f33a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6799,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorReturnVarianceProxy_coe (mdp : MDP State Action) (episodes rounds : Nat) : ((cumulativeSuccessorReturnVarianceProxy mdp episodes rounds : NNReal) : Real) = (rounds : Real) * ((episodes : Real) * (mdp.horizon : Real)) ^ 2","missing":[],"search":"cumulativesuccessorreturnvarianceproxy_coe banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativesuccessorreturnvarianceproxy_coe the successor-return proxy contains exactly `rounds` nonzero batch terms. theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius","label":"normalizedSuccessorReturnConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius","description":"Return-deviation radius after normalization by all successor episodes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-2f188f9a8bc3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6800,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedSuccessorReturnConfidenceRadius (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) : Real","missing":[],"search":"normalizedsuccessorreturnconfidenceradius banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedsuccessorreturnconfidenceradius return-deviation radius after normalization by all successor episodes. definition compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_nonneg","label":"normalizedSuccessorReturnConfidenceRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_nonneg","description":"theorem normalizedSuccessorReturnConfidenceRadius_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) : 0 <= normalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-536298740744","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6801,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorReturnConfidenceRadius_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) : 0 <= normalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta","missing":[],"search":"normalizedsuccessorreturnconfidenceradius_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedsuccessorreturnconfidenceradius_nonneg theorem normalizedsuccessorreturnconfidenceradius_nonneg (mdp : mdp state action) (episodes rounds : nat) (delta : real) : 0 <= normalizedsuccessorreturnconfidenceradius mdp episodes rounds delta theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_eq","label":"normalizedSuccessorReturnConfidenceRadius_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_eq","description":"The scheduled batch size cancels exactly from the normalized radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-3cfcf0b720e5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6802,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorReturnConfidenceRadius_eq (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) (hepisodes : 0 < episodes) (hrounds : 0 < rounds) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : normalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta = (mdp.horizon : Real) * Real.sqrt (2 * Real.log (2 / delta) / (rounds : Real))","missing":[],"search":"normalizedsuccessorreturnconfidenceradius_eq banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedsuccessorreturnconfidenceradius_eq the scheduled batch size cancels exactly from the normalized radius. theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationReturnRadiusEnvelope","label":"decayingExplorationReturnRadiusEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationReturnRadiusEnvelope","description":"Elementary deterministic envelope for the normalized return radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-897e947237e4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6803,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationReturnRadiusEnvelope (mdp : MDP State Action) (n : Nat) : Real","missing":[],"search":"decayingexplorationreturnradiusenvelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationreturnradiusenvelope elementary deterministic envelope for the normalized return radius. definition compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","label":"normalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","description":"The decaying schedule dominates the normalized logarithmic return radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-58637a6bea9f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6804,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:150"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope (mdp : MDP State Action) (episodes : Nat) (n : Nat) (hepisodes : 0 < episodes) : normalizedSuccessorReturnConfidenceRadius mdp episodes (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n) <= decayingExplorationReturnRadiusEnvelope mdp n","missing":[],"search":"normalizedsuccessorreturnconfidenceradius_le_decayingenvelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedsuccessorreturnconfidenceradius_le_decayingenvelope the decaying schedule dominates the normalized logarithmic return radius. theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes_pos","label":"decayingExplorationScheduledEpisodes_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes_pos","description":"theorem decayingExplorationScheduledEpisodes_pos (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : 0 < decayingExplorationScheduledEpisodes mdp baseVisitFloor n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-42658f31d5ce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6805,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:208"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationScheduledEpisodes_pos (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : 0 < decayingExplorationScheduledEpisodes mdp baseVisitFloor n","missing":[],"search":"decayingexplorationscheduledepisodes_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationscheduledepisodes_pos theorem decayingexplorationscheduledepisodes_pos (mdp : mdp state action) (basevisitfloor : real) (n : nat) : 0 < decayingexplorationscheduledepisodes mdp basevisitfloor n theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationReturnRadiusEnvelope_tendsto_zero","label":"decayingExplorationReturnRadiusEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationReturnRadiusEnvelope_tendsto_zero","description":"theorem decayingExplorationReturnRadiusEnvelope_tendsto_zero (mdp : MDP State Action) : Tendsto (decayingExplorationReturnRadiusEnvelope mdp) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-41cb6b2a2ec5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6806,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationReturnRadiusEnvelope_tendsto_zero (mdp : MDP State Action) : Tendsto (decayingExplorationReturnRadiusEnvelope mdp) atTop (nhds 0)","missing":[],"search":"decayingexplorationreturnradiusenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationreturnradiusenvelope_tendsto_zero theorem decayingexplorationreturnradiusenvelope_tendsto_zero (mdp : mdp state action) : tendsto (decayingexplorationreturnradiusenvelope mdp) attop (nhds 0) theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationNormalizedReturnRadius_tendsto_zero","label":"decayingExplorationNormalizedReturnRadius_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationNormalizedReturnRadius_tendsto_zero","description":"theorem decayingExplorationNormalizedReturnRadius_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) : Tendsto (fun n => normalizedSuccessorReturnConfidenceRadius mdp (decayingExplorationScheduledEpisodes mdp baseVisitFloor n) (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n)) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-c7534617cb59","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6807,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:236"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationNormalizedReturnRadius_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) : Tendsto (fun n => normalizedSuccessorReturnConfidenceRadius mdp (decayingExplorationScheduledEpisodes mdp baseVisitFloor n) (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n)) atTop (nhds 0)","missing":[],"search":"decayingexplorationnormalizedreturnradius_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationnormalizedreturnradius_tendsto_zero theorem decayingexplorationnormalizedreturnradius_tendsto_zero (mdp : mdp state action) (basevisitfloor : real) : tendsto (fun n => normalizedsuccessorreturnconfidenceradius mdp (decayingexplorationscheduledepisodes mdp basevisitfloor n) (decayingexplorationrounds mdp n) (vanishingaverageconfidencedelta n)) attop (nhds 0) theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound","label":"decayingExplorationAverageRealizedBehaviorRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound","description":"Deterministic realized-behavior certificate at schedule index `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-19f683a843db","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6808,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:255"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageRealizedBehaviorRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationaveragerealizedbehaviorregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaveragerealizedbehaviorregretbound deterministic realized-behavior certificate at schedule index `n`. definition compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget","label":"decayingExplorationRealizedFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget","description":"Count and return deviations each consume one scheduled confidence share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-4eaf7ec63ed8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6809,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:265"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationRealizedFailureBudget (n : Nat) : ENNReal","missing":[],"search":"decayingexplorationrealizedfailurebudget banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrealizedfailurebudget count and return deviations each consume one scheduled confidence share. definition compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound_nonneg","label":"decayingExplorationAverageRealizedBehaviorRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound_nonneg","description":"theorem decayingExplorationAverageRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= decayingExplorationAverageRealizedBehaviorRegretBound mdp baseVisitFloor n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-a71693306466","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6810,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationAverageRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= decayingExplorationAverageRealizedBehaviorRegretBound mdp baseVisitFloor n","missing":[],"search":"decayingexplorationaveragerealizedbehaviorregretbound_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaveragerealizedbehaviorregretbound_nonneg theorem decayingexplorationaveragerealizedbehaviorregretbound_nonneg (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) (n : nat) : 0 <= decayingexplorationaveragerealizedbehaviorregretbound mdp basevisitfloor n theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound_tendsto_zero","label":"decayingExplorationAverageRealizedBehaviorRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound_tendsto_zero","description":"theorem decayingExplorationAverageRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationAverageRealizedBehaviorRegretBound mdp baseVisitFloor) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-510e24a7c4bb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6811,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:290"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationAverageRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationAverageRealizedBehaviorRegretBound mdp baseVisitFloor) atTop (nhds 0)","missing":[],"search":"decayingexplorationaveragerealizedbehaviorregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationaveragerealizedbehaviorregretbound_tendsto_zero theorem decayingexplorationaveragerealizedbehaviorregretbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (decayingexplorationaveragerealizedbehaviorregretbound mdp basevisitfloor) attop (nhds 0) theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget_tendsto_zero","label":"decayingExplorationRealizedFailureBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget_tendsto_zero","description":"theorem decayingExplorationRealizedFailureBudget_tendsto_zero : Tendsto decayingExplorationRealizedFailureBudget atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-d76351572196","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6812,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:305"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationRealizedFailureBudget_tendsto_zero : Tendsto decayingExplorationRealizedFailureBudget atTop (nhds 0)","missing":[],"search":"decayingexplorationrealizedfailurebudget_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrealizedfailurebudget_tendsto_zero theorem decayingexplorationrealizedfailurebudget_tendsto_zero : tendsto decayingexplorationrealizedfailurebudget attop (nhds 0) theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureAndRegretBound_tendsto_zero","label":"decayingExplorationRealizedFailureAndRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureAndRegretBound_tendsto_zero","description":"theorem decayingExplorationRealizedFailureAndRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (decayingExplorationRealizedFailureBudget n, decayingExplorationAverageRealizedBehaviorRegretBound mdp baseVisitFloor n)) atTop (nhds (0, 0))","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-7964c33541ab","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6813,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationRealizedFailureAndRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (decayingExplorationRealizedFailureBudget n, decayingExplorationAverageRealizedBehaviorRegretBound mdp baseVisitFloor n)) atTop (nhds (0, 0))","missing":[],"search":"decayingexplorationrealizedfailureandregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationrealizedfailureandregretbound_tendsto_zero theorem decayingexplorationrealizedfailureandregretbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (fun n => (decayingexplorationrealizedfailurebudget n, decayingexplorationaveragerealizedbehaviorregretbound mdp basevisitfloor n)) attop (nhds (0, 0)) theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationAverageRealizedBehaviorRegretViolationSet","label":"decayingExplorationAverageRealizedBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationAverageRealizedBehaviorRegretViolationSet","description":"Realized-regret violation set for one decaying-exploration window.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-be4ed6c1416f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6814,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:330"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Set (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","missing":[],"search":"decayingexplorationaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationaveragerealizedbehaviorregretviolationset realized-regret violation set for one decaying-exploration window. definition compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","description":"One finite window: the realized violation set is covered by the measurable count/return union, whose tail is the two-share scheduled budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-3d39c2e26b7c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6815,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:361"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpis…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorconsistency banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorconsistency one finite window: the realized violation set is covered by the measurable count/return union, whose tail is the two-share scheduled budget. theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","description":"All finite windows plus the joint scalar limit. The indexed Borel witnesses make the changing sample spaces explicit and prevent a common-space reading.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativedecayingexplorationrealizedbehaviorconsistency/index.html#decl-9d1283bb14db","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","order":6816,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency.lean:482"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < bas…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_decayingexplorationaveragerealizedbehaviorconsistency_allwindows banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_decayingexplorationaveragerealizedbehaviorconsistency_allwindows all finite windows plus the joint scalar limit. the indexed borel witnesses make the changing sample spaces explicit and prevent a common-space reading. theorem compiled","shard":"modules/277247cf8ab0a5f8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius","label":"TransitionCountRadius","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius","description":"A nonnegative transition radius that can only decrease as visits accumulate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-f24ce5d5d905","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6817,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure TransitionCountRadius where","missing":[],"search":"transitioncountradius banditrlproof.finitehorizonrl.transitioncountradius a nonnegative transition radius that can only decrease as visits accumulate. structure compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.linearDecay","label":"linearDecay","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.linearDecay","description":"A finite linear-decay radius, useful as an executable shrinking-radius canary.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-923b53186e97","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6818,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def linearDecay (budget : Nat) : TransitionCountRadius where","missing":[],"search":"lineardecay banditrlproof.finitehorizonrl.transitioncountradius.lineardecay a finite linear-decay radius, useful as an executable shrinking-radius canary. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeTransitionCountSummary","label":"cumulativeTransitionCountSummary","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeTransitionCountSummary","description":"Sum all transition-count coordinates in a nonempty finite batch prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-88ba69ae2b51","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6819,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeTransitionCountSummary {mdp : MDP State Action} {episodes n : Nat} (history : EpisodeBatchPrefix mdp episodes n) : TransitionCountSummary mdp","missing":[],"search":"cumulativetransitioncountsummary banditrlproof.finitehorizonrl.episodebatchprefix.cumulativetransitioncountsummary sum all transition-count coordinates in a nonempty finite batch prefix. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeTransitionCountSummary","label":"measurable_cumulativeTransitionCountSummary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeTransitionCountSummary","description":"The complete cumulative transition-count summary is history measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-e66926cd9620","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6820,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeTransitionCountSummary {mdp : MDP State Action} {episodes n : Nat} : Measurable (cumulativeTransitionCountSummary : EpisodeBatchPrefix mdp episodes n -> TransitionCountSummary mdp)","missing":[],"search":"measurable_cumulativetransitioncountsummary banditrlproof.finitehorizonrl.episodebatchprefix.measurable_cumulativetransitioncountsummary the complete cumulative transition-count summary is history measurable. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeTransitionCountSummary_visitCount","label":"cumulativeTransitionCountSummary_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeTransitionCountSummary_visitCount","description":"Cumulative next-state counts still partition the cumulative visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-a5d093847632","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6821,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:93"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeTransitionCountSummary_visitCount {mdp : MDP State Action} {episodes n : Nat} (history : EpisodeBatchPrefix mdp episodes n) (stage : Fin mdp.horizon) (state : State) (action : Action) : history.cumulativeTransitionCountSummary.visitCount stage state action = ∑ i : Fin (n + 1), (history ⟨i, Finset.mem_Iic.mpr (Nat.le_of_lt_succ i.isLt)⟩).visitCount stage state action","missing":[],"search":"cumulativetransitioncountsummary_visitcount banditrlproof.finitehorizonrl.episodebatchprefix.cumulativetransitioncountsummary_visitcount cumulative next-state counts still partition the cumulative visit count. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt","label":"cumulativeTransitionCountSummaryAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt","description":"Cumulative count summary through trajectory coordinate `round`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-89c5a88ea5bb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6822,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeTransitionCountSummaryAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) : TransitionCountSummary mdp","missing":[],"search":"cumulativetransitioncountsummaryat banditrlproof.finitehorizonrl.cumulativetransitioncountsummaryat cumulative count summary through trajectory coordinate `round`. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt_succ","label":"cumulativeTransitionCountSummaryAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt_succ","description":"Extending the trajectory prefix adds exactly the new batch coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-a3edf2e378bb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6823,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeTransitionCountSummaryAt_succ {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : cumulativeTransitionCountSummaryAt trajectory (round + 1) stage state action nextState = cumulativeTransitionCountSummaryAt trajectory round stage state action nextState + (trajectory (round + 1)).transitionCount stage state action nextState","missing":[],"search":"cumulativetransitioncountsummaryat_succ banditrlproof.finitehorizonrl.cumulativetransitioncountsummaryat_succ extending the trajectory prefix adds exactly the new batch coordinate. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt_visitCount_le_succ","label":"cumulativeTransitionCountSummaryAt_visitCount_le_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt_visitCount_le_succ","description":"Every cumulative state-action visit count is monotone across rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-e2a6865ed22c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6824,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeTransitionCountSummaryAt_visitCount_le_succ {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : (cumulativeTransitionCountSummaryAt trajectory round).visitCount stage state action <= (cumulativeTransitionCountSummaryAt trajectory (round + 1)).visitCount stage state action","missing":[],"search":"cumulativetransitioncountsummaryat_visitcount_le_succ banditrlproof.finitehorizonrl.cumulativetransitioncountsummaryat_visitcount_le_succ every cumulative state-action visit count is monotone across rounds. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.radius_cumulativeVisitCount_succ_le","label":"radius_cumulativeVisitCount_succ_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.radius_cumulativeVisitCount_succ_le","description":"Count-antitone radii shrink along every accumulated state-action row.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-975cc9b014ac","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6825,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:159"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem TransitionCountRadius.radius_cumulativeVisitCount_succ_le {mdp : MDP State Action} {episodes : Nat} (countRadius : TransitionCountRadius) (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : countRadius.radius ((cumulativeTransitionCountSummaryAt trajectory (round + 1)).visitCount stage state action) <= countRadius.radius ((cumulativeTransitionCountSummaryAt trajectory round).visitCount stage state action)","missing":[],"search":"radius_cumulativevisitcount_succ_le banditrlproof.finitehorizonrl.transitioncountradius.radius_cumulativevisitcount_succ_le count-antitone radii shrink along every accumulated state-action row. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan","label":"countRadiusOptimisticPlan","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan","description":"Known-reward empirical plan with a radius selected from cumulative visits.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-3567d12567fa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6826,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:177"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def countRadiusOptimisticPlan (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (countRadius : TransitionCountRadius) : mdp.EstimatedModelPlan where","missing":[],"search":"countradiusoptimisticplan banditrlproof.finitehorizonrl.transitioncountsummary.countradiusoptimisticplan known-reward empirical plan with a radius selected from cumulative visits. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan_upperValueRemaining_abs_le","label":"countRadiusOptimisticPlan_upperValueRemaining_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan_upperValueRemaining_abs_le","description":"Every count-radius optimistic value is controlled by the radius at zero visits.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-12ef967a5776","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6827,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countRadiusOptimisticPlan_upperValueRemaining_abs_le (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (countRadius : TransitionCountRadius) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) : forall (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State), |(summary.countRadiusOptimisticPlan mdp defaultState countRadius).upperValueRemaining remaining hremaining state| <= empiricalFiniteBatchValueEnvelope rewardBound (countRadius.radius 0) remaining","missing":[],"search":"countradiusoptimisticplan_uppervalueremaining_abs_le banditrlproof.finitehorizonrl.transitioncountsummary.countradiusoptimisticplan_uppervalueremaining_abs_le every count-radius optimistic value is controlled by the radius at zero visits. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan_selectedRadiusRemaining","label":"countRadiusOptimisticPlan_selectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan_selectedRadiusRemaining","description":"The cumulative count-radius plan selects exactly its chosen row radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-f6984cf03207","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6828,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:281"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countRadiusOptimisticPlan_selectedRadiusRemaining (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (countRadius : TransitionCountRadius) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (summary.countRadiusOptimisticPlan mdp defaultState countRadius).selectedRadiusRemaining remaining hremaining state = countRadius.radius (summary.visitCount (mdp.decisionStageRemaining remaining hremaining) state ((summary.countRadiusOptimisticPlan mdp defaultState countRadius).optimisticAction (mdp.decisionStageRemaining remaining hremaining) ((summary.countRadiusOptimisticPlan mdp defaultState countRadius).upperValueRemaining remaining (by omega)) state))","missing":[],"search":"countradiusoptimisticplan_selectedradiusremaining banditrlproof.finitehorizonrl.transitioncountsummary.countradiusoptimisticplan_selectedradiusremaining the cumulative count-radius plan selects exactly its chosen row radius. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPolicyTable","label":"countRadiusOptimisticPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPolicyTable","description":"Deterministic optimistic table selected from a cumulative count summary.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-0a93d6788f55","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6829,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:299"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def countRadiusOptimisticPolicyTable (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (countRadius : TransitionCountRadius) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"countradiusoptimisticpolicytable banditrlproof.finitehorizonrl.transitioncountsummary.countradiusoptimisticpolicytable deterministic optimistic table selected from a cumulative count summary. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPolicyTable_toMarkovPolicy","label":"countRadiusOptimisticPolicyTable_toMarkovPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPolicyTable_toMarkovPolicy","description":"The table interpretation is the cumulative count-radius plan's policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-06371b9fb421","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6830,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:308"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countRadiusOptimisticPolicyTable_toMarkovPolicy (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (countRadius : TransitionCountRadius) : (summary.countRadiusOptimisticPolicyTable mdp defaultState countRadius).toMarkovPolicy = (summary.countRadiusOptimisticPlan mdp defaultState countRadius).optimisticPolicy","missing":[],"search":"countradiusoptimisticpolicytable_tomarkovpolicy banditrlproof.finitehorizonrl.transitioncountsummary.countradiusoptimisticpolicytable_tomarkovpolicy the table interpretation is the cumulative count-radius plan's policy. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticPlanAt","label":"adaptiveCumulativeEmpiricalOptimisticPlanAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticPlanAt","description":"Cumulative known-reward optimistic plan recommended after one trajectory prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-bda5420a660d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6831,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:318"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeEmpiricalOptimisticPlanAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (round : Nat) : mdp.EstimatedModelPlan","missing":[],"search":"adaptivecumulativeempiricaloptimisticplanat banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticplanat cumulative known-reward optimistic plan recommended after one trajectory prefix. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticRecommendedExpectedRegret","label":"adaptiveCumulativeEmpiricalOptimisticRecommendedExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticRecommendedExpectedRegret","description":"Sum of expected regrets of all cumulative empirical recommendations.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-b1ddee09c88e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6832,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:327"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeEmpiricalOptimisticRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (rounds : Nat) : Real","missing":[],"search":"adaptivecumulativeempiricaloptimisticrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticrecommendedexpectedregret sum of expected regrets of all cumulative empirical recommendations. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.successorTable","label":"successorTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.successorTable","description":"Cumulative optimistic table selected from every observed batch in a prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-07e732956fa7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6833,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:340"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (countRadius : TransitionCountRadius) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"successortable banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.successortable cumulative optimistic table selected from every observed batch in a prefix. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurable_successorTable","label":"measurable_successorTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurable_successorTable","description":"The cumulative history-to-table selector is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-9f1b1cb547e5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6834,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (countRadius : TransitionCountRadius) (n : Nat) : Measurable (successorTable (mdp := mdp) (episodes := episodes) defaultState countRadius n)","missing":[],"search":"measurable_successortable banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.measurable_successortable the cumulative history-to-table selector is measurable. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource","label":"exploratorySource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource","description":"Exploratory behavior source centered on cumulative empirical recommendations.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-0688ce60a089","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6835,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:362"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratorySource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : AdaptiveEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"exploratorysource banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource exploratory behavior source centered on cumulative empirical recommendations. definition compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorPolicy","label":"exploratorySource_successorPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorPolicy","description":"The source's successor behavior is centered on the cumulative optimistic table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-c68249ad248b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6836,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:387"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).successorPolicy n history = (successorTable defaultState countRadius n history).exploratoryPolicy explorationRate hexplorationRate","missing":[],"search":"exploratorysource_successorpolicy banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_successorpolicy the source's successor behavior is centered on the cumulative optimistic table. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeCoordinateConfidenceContract","label":"AdaptiveCumulativeCoordinateConfidenceContract","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeCoordinateConfidenceContract","description":"One global event contract for cumulative empirical plans. The missing statistical producer must supply coordinate confidence for every cumulative prefix plan outside `badEvent`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-a87f3cab301b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6837,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:407"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure AdaptiveCumulativeCoordinateConfidenceContract {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (defaultState : State) (countRadius : TransitionCountRadius) (rounds : Nat) (delta : Real) where","missing":[],"search":"adaptivecumulativecoordinateconfidencecontract banditrlproof.finitehorizonrl.adaptivecumulativecoordinateconfidencecontract one global event contract for cumulative empirical plans. the missing statistical producer must supply coordinate confidence for every cumulative prefix plan outside `badevent`. structure compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeCoordinateConfidence_optimism_and_recommendedExpectedRegret","label":"adaptiveCumulativeCoordinateConfidence_optimism_and_recommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeCoordinateConfidence_optimism_and_recommendedExpectedRegret","description":"Roundwise cumulative confidence and a selected-radius envelope imply global optimism and the explicit finite recommendation-regret sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-4071b27e181c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6838,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:426"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeCoordinateConfidence_optimism_and_recommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (radiusEnvelope : Fin rounds -> Real) (confidence : forall round : Fin rounds, (adaptiveCumulativeEmpiricalOptimisticPlanAt trajectory defaultState countRadius round).CoordinateConfidence) (hradius : forall (round : Fin rounds) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State), (adaptiveCumulativeEmpiricalOptimisticPlanAt trajectory defaultState countRadius round).selectedRadiusRemaining remaining hremaining state <= radiusEnvelope round) : (forall round : Fin rounds, forall state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= (adaptiveCumulat…","missing":[],"search":"adaptivecumulativecoordinateconfidence_optimism_and_recommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativecoordinateconfidence_optimism_and_recommendedexpectedregret roundwise cumulative confidence and a selected-radius envelope imply global optimism and the explicit finite recommendation-regret sum. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeCoordinateConfidenceContract.trajectoryMeasure_optimism_and_explicitRecommendedExpectedRegret","label":"trajectoryMeasure_optimism_and_explicitRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeCoordinateConfidenceContract.trajectoryMeasure_optimism_and_explicitRecommendedExpectedRegret","description":"Route endpoint: one measurable global cumulative-confidence event yields global optimism and an explicit shrinking-radius recommendation-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeempiricaloptimisticregret/index.html#decl-37b39a969c23","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","order":6839,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret.lean:485"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem AdaptiveCumulativeCoordinateConfidenceContract.trajectoryMeasure_optimism_and_explicitRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (defaultState : State) (countRadius : TransitionCountRadius) (delta : Real) (contract : AdaptiveCumulativeCoordinateConfidenceContract source defaultState countRadius rounds delta) (radiusEnvelope : Fin rounds -> Real) (hradius : forall trajectory, trajectory ∉ contract.badEvent -> forall (round : Fin rounds) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State), (adaptiveCumulativeEmpiricalOptimisticPlanAt trajectory defaultState countRadius round).selectedRadiusRemaining remaining hremaining state <= radiusEnvelope round) : MeasurableSet contract.badEvent /\\ source.t…","missing":[],"search":"trajectorymeasure_optimism_and_explicitrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativecoordinateconfidencecontract.trajectorymeasure_optimism_and_explicitrecommendedexpectedregret route endpoint: one measurable global cumulative-confidence event yields global optimism and an explicit shrinking-radius recommendation-regret bound. theorem compiled","shard":"modules/b000410b6d687ac3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","label":"measurable_realizedSuccessorCumulativeRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","description":"Realized cumulative successor regret is measurable in the finite trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-f1531dcce77b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6840,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.realizedSuccessorCumulativeRegret trajectory rounds)","missing":[],"search":"measurable_realizedsuccessorcumulativeregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_realizedsuccessorcumulativeregret realized cumulative successor regret is measurable in the finite trajectory. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","label":"measurable_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","description":"Realized average successor regret is measurable in the finite trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-3c34e9e3b018","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6841,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.realizedSuccessorAverageRegret trajectory rounds)","missing":[],"search":"measurable_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_realizedsuccessoraverageregret realized average successor regret is measurable in the finite trajectory. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedCumulativeRegret_nonneg","label":"successorExpectedCumulativeRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedCumulativeRegret_nonneg","description":"A finite sum of policy expected regrets is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-4ff8827b0188","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6842,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorExpectedCumulativeRegret_nonneg {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : 0 <= source.successorExpectedCumulativeRegret trajectory rounds","missing":[],"search":"successorexpectedcumulativeregret_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorexpectedcumulativeregret_nonneg a finite sum of policy expected regrets is nonnegative. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedAverageRegret_nonneg","label":"successorExpectedAverageRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedAverageRegret_nonneg","description":"The average of successor policy expected regrets is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-71cddf24b292","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6843,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorExpectedAverageRegret_nonneg {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : 0 <= source.successorExpectedAverageRegret trajectory rounds","missing":[],"search":"successorexpectedaverageregret_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorexpectedaverageregret_nonneg the average of successor policy expected regrets is nonnegative. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","label":"abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","description":"An expected-regret upper bound and a two-sided return-deviation bound control the absolute realized average regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-e76b9aae09f2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6844,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (expectedBound deviationBound : Real) (hexpected : source.successorExpectedAverageRegret trajectory rounds <= expectedBound) (hdeviation : |source.cumulativeSuccessorReturnDeviation rounds trajectory| <= deviationBound) : |source.realizedSuccessorAverageRegret trajectory rounds| <= expectedBound + deviationBound / ((episodes : Real) * (rounds : Real))","missing":[],"search":"abs_realizedsuccessoraverageregret_le_of_expected_le_of_deviation_abs_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.abs_realizedsuccessoraverageregret_le_of_expected_le_of_deviation_abs_le an expected-regret upper bound and a two-sided return-deviation bound control the absolute realized average regret. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageAbsoluteRealizedBehaviorConsistency","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageAbsoluteRealizedBehaviorConsistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageAbsoluteRealizedBehaviorConsistency","description":"The sharp finite-window certificate controls the absolute realized regret, not only its upper tail. This is the finite-window input needed by convergence in probability.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-9f1b8f54dc7e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6845,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageAbsoluteRealizedBehaviorConsistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rou…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationepisodewiseaverageabsoluterealizedbehaviorconsistency banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationepisodewiseaverageabsoluterealizedbehaviorconsistency the sharp finite-window certificate controls the absolute realized regret, not only its upper tail. this is the finite-window input needed by convergence in probability. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.DecayingExplorationEpisodewiseWindowSpace","label":"DecayingExplorationEpisodewiseWindowSpace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.DecayingExplorationEpisodewiseWindowSpace","description":"The dependent sample space containing one complete scheduled experiment per coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-cf48b2c7712b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6846,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:262"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev DecayingExplorationEpisodewiseWindowSpace (mdp : MDP State Action) (baseVisitFloor : Real)","missing":[],"search":"decayingexplorationepisodewisewindowspace banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisewindowspace the dependent sample space containing one complete scheduled experiment per coordinate. abbreviation compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseWindowSource","label":"decayingExplorationEpisodewiseWindowSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseWindowSource","description":"The concrete adaptive source used at schedule coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-ae8a06e31c39","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6847,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseWindowSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : AdaptiveEpisodeBatchSource mdp initialState (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n)","missing":[],"search":"decayingexplorationepisodewisewindowsource banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisewindowsource the concrete adaptive source used at schedule coordinate `n`. definition compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseWindowMeasure","label":"decayingExplorationEpisodewiseWindowMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseWindowMeasure","description":"The scheduled adaptive trajectory law at common-space coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-1f3a18142c9e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6848,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:290"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseWindowMeasure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Measure (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","missing":[],"search":"decayingexplorationepisodewisewindowmeasure banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisewindowmeasure the scheduled adaptive trajectory law at common-space coordinate `n`. definition compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure","label":"decayingExplorationEpisodewiseCommonMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure","description":"Independent-coordinate coupling of all scheduled finite-window trajectory laws.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-83f2b5687b3d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6849,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:315"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseCommonMeasure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) : Measure (DecayingExplorationEpisodewiseWindowSpace mdp baseVisitFloor)","missing":[],"search":"decayingexplorationepisodewisecommonmeasure banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisecommonmeasure independent-coordinate coupling of all scheduled finite-window trajectory laws. definition compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_map_eval","label":"decayingExplorationEpisodewiseCommonMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_map_eval","description":"Every common-space coordinate has exactly its scheduled adaptive law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-fdd9e8a2840f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6850,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:337"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseCommonMeasure_map_eval (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : (decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor).map (fun omega => omega n) = decayingExplorationEpisodewiseWindowMeasure mdp initialState initialTable defaultState baseVisitFloor n","missing":[],"search":"decayingexplorationepisodewisecommonmeasure_map_eval banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisecommonmeasure_map_eval every common-space coordinate has exactly its scheduled adaptive law. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","label":"decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","description":"Scheduled realized successor-average regret as one common-space process.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-8949d3bc9f8e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6851,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) (omega : DecayingExplorationEpisodewiseWindowSpace mdp baseVisitFloor) : Real","missing":[],"search":"decayingexplorationepisodewiserealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiserealizedbehaviorregretprocess scheduled realized successor-average regret as one common-space process. definition compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","label":"measurable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","description":"Every scheduled regret coordinate is a measurable real random variable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-b7184ae58a0e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6852,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:362"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Measurable (decayingExplorationEpisodewiseRealizedBehaviorRegretProcess mdp initialState initialTable defaultState baseVisitFloor n)","missing":[],"search":"measurable_decayingexplorationepisodewiserealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.measurable_decayingexplorationepisodewiserealizedbehaviorregretprocess every scheduled regret coordinate is a measurable real random variable. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonBadEvent","label":"decayingExplorationEpisodewiseCommonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonBadEvent","description":"Pull the sharp count/return union at coordinate `n` to the common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-71030a08847d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6853,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:378"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseCommonBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Set (DecayingExplorationEpisodewiseWindowSpace mdp baseVisitFloor)","missing":[],"search":"decayingexplorationepisodewisecommonbadevent banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisecommonbadevent pull the sharp count/return union at coordinate `n` to the common space. definition compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_badEvent_le","label":"decayingExplorationEpisodewiseCommonMeasure_badEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_badEvent_le","description":"The pulled-back coordinate bad event inherits the exact finite-window budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-8f4292ea73a8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6854,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:393"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseCommonMeasure_badEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : decayingExplorationEpisodewiseCommonMeasure mdp initi…","missing":[],"search":"decayingexplorationepisodewisecommonmeasure_badevent_le banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisecommonmeasure_badevent_le the pulled-back coordinate bad event inherits the exact finite-window budget. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.abs_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","label":"abs_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.abs_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","description":"Outside the pulled-back bad event, coordinate `n` has the sharp absolute bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-99779890257a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6855,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:446"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_le_of_not_mem_badEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) (omega : DecayingExplora…","missing":[],"search":"abs_decayingexplorationepisodewiserealizedbehaviorregretprocess_le_of_not_mem_badevent banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.abs_decayingexplorationepisodewiserealizedbehaviorregretprocess_le_of_not_mem_badevent outside the pulled-back bad event, coordinate `n` has the sharp absolute bound. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","label":"exploratorySource_decayingExplorationEpisodewiseCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","description":"Terminal common-space theorem: measurable regret coordinates, exact scheduled marginals, and convergence in probability of the process to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceconsistency/index.html#decl-dd3f67b5aff1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","order":6856,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency.lean:488"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_decayingExplorationEpisodewiseCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor)…","missing":[],"search":"exploratorysource_decayingexplorationepisodewisecommonmeasure_marginals_and_realizedbehaviorregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_decayingexplorationepisodewisecommonmeasure_marginals_and_realizedbehaviorregret_tendstoinmeasure_zero terminal common-space theorem: measurable regret coordinates, exact scheduled marginals, and convergence in probability of the process to zero. theorem compiled","shard":"modules/b1dfbc5b216b0574.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.abs_cumulativeReward_le_horizon","label":"abs_cumulativeReward_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.abs_cumulativeReward_le_horizon","description":"A deterministic finite trajectory has return bounded by the horizon.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-cfb64ca1645b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6857,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_cumulativeReward_le_horizon (mdp : MDP State Action) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (trajectory : Prod State (StepTrace Action State mdp.horizon)) : |mdp.cumulativeReward trajectory| <= (mdp.horizon : Real)","missing":[],"search":"abs_cumulativereward_le_horizon banditrlproof.finitehorizonrl.mdp.abs_cumulativereward_le_horizon a deterministic finite trajectory has return bounded by the horizon. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_optimalInitialExpectedReturn_le_horizon","label":"abs_optimalInitialExpectedReturn_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_optimalInitialExpectedReturn_le_horizon","description":"The optimal initial expected return inherits the deterministic horizon envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-614513d9bbe1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6858,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:63"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_optimalInitialExpectedReturn_le_horizon (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (hrewardBound : forall state action, |mdp.reward state action| <= 1) : |optimalInitialExpectedReturn mdp initialState| <= (mdp.horizon : Real)","missing":[],"search":"abs_optimalinitialexpectedreturn_le_horizon banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.abs_optimalinitialexpectedreturn_le_horizon the optimal initial expected return inherits the deterministic horizon envelope. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successor_rewardConsistent_ae","label":"trajectoryMeasure_successor_rewardConsistent_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successor_rewardConsistent_ae","description":"Every adaptive successor batch is reward-consistent almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-a8258cfbfb54","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6859,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successor_rewardConsistent_ae {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : ∀ᵐ trajectory ∂source.trajectoryMeasure, (trajectory (n + 1)).RewardConsistent","missing":[],"search":"trajectorymeasure_successor_rewardconsistent_ae banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_successor_rewardconsistent_ae every adaptive successor batch is reward-consistent almost everywhere. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_two_mul_horizon","label":"abs_realizedSuccessorAverageRegret_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_two_mul_horizon","description":"Reward-consistent successor batches give a uniform realized-average envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-5e0a5ca5e479","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6860,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_realizedSuccessorAverageRegret_le_two_mul_horizon {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hconsistent : forall round : Fin rounds, (trajectory ((round : Nat) + 1)).RewardConsistent) : |source.realizedSuccessorAverageRegret trajectory rounds| <= 2 * (mdp.horizon : Real)","missing":[],"search":"abs_realizedsuccessoraverageregret_le_two_mul_horizon banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.abs_realizedsuccessoraverageregret_le_two_mul_horizon reward-consistent successor batches give a uniform realized-average envelope. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_abs_realizedSuccessorAverageRegret_le_two_mul_horizon_ae","label":"trajectoryMeasure_abs_realizedSuccessorAverageRegret_le_two_mul_horizon_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_abs_realizedSuccessorAverageRegret_le_two_mul_horizon_ae","description":"The adaptive realized average has the `2H` envelope almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-71e8512cf30f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6861,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:168"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_abs_realizedSuccessorAverageRegret_le_two_mul_horizon_ae {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ∀ᵐ trajectory ∂source.trajectoryMeasure, |source.realizedSuccessorAverageRegret trajectory rounds| <= 2 * (mdp.horizon : Real)","missing":[],"search":"trajectorymeasure_abs_realizedsuccessoraverageregret_le_two_mul_horizon_ae banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_abs_realizedsuccessoraverageregret_le_two_mul_horizon_ae the adaptive realized average has the `2h` envelope almost everywhere. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","label":"integrable_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","description":"The adaptive realized successor-average regret is integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-af2901124701","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6862,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:187"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : Integrable (fun trajectory => source.realizedSuccessorAverageRegret trajectory rounds) source.trajectoryMeasure","missing":[],"search":"integrable_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.integrable_realizedsuccessoraverageregret the adaptive realized successor-average regret is integrable. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_abs_realizedBehaviorRegretProcess_le_two_mul_horizon_ae","label":"decayingExplorationEpisodewiseCommonMeasure_abs_realizedBehaviorRegretProcess_le_two_mul_horizon_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_abs_realizedBehaviorRegretProcess_le_two_mul_horizon_ae","description":"The common-space regret process has the deterministic `2H` envelope a.e.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-f4721647eb13","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6863,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:208"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseCommonMeasure_abs_realizedBehaviorRegretProcess_le_two_mul_horizon_ae (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : ∀ᵐ omega ∂decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor, |decayingExplorationEpisodewiseRealizedBehaviorRegretProcess mdp initialState initialTable defaultState baseVisitFloor n omega| <= 2 * (mdp.horizon : Real)","missing":[],"search":"decayingexplorationepisodewisecommonmeasure_abs_realizedbehaviorregretprocess_le_two_mul_horizon_ae banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewisecommonmeasure_abs_realizedbehaviorregretprocess_le_two_mul_horizon_ae the common-space regret process has the deterministic `2h` envelope a.e. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.integrable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","label":"integrable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.integrable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","description":"Every common-space regret coordinate is integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-27dabca11518","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6864,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:246"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : Integrable (decayingExplorationEpisodewiseRealizedBehaviorRegretProcess mdp initialState initialTable defaultState baseVisitFloor n) (decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor)","missing":[],"search":"integrable_decayingexplorationepisodewiserealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.integrable_decayingexplorationepisodewiserealizedbehaviorregretprocess every common-space regret coordinate is integrable. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurableSet_decayingExplorationEpisodewiseCommonBadEvent","label":"measurableSet_decayingExplorationEpisodewiseCommonBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurableSet_decayingExplorationEpisodewiseCommonBadEvent","description":"The pulled-back sharp finite-window bad event is measurable on the common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-ded3341128be","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6865,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:268"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_decayingExplorationEpisodewiseCommonBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : MeasurableSet (decayingExplorationEpisodewiseCommo…","missing":[],"search":"measurableset_decayingexplorationepisodewisecommonbadevent banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.measurableset_decayingexplorationepisodewisecommonbadevent the pulled-back sharp finite-window bad event is measurable on the common space. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret","label":"decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret","description":"Expected absolute scheduled realized-behavior regret on the common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-e54eae6da595","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6866,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:313"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregret expected absolute scheduled realized-behavior regret on the common space. definition compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound","label":"decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound","description":"The deterministic good-event radius plus the uniform bad-event contribution.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-68547c0805c4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6867,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:326"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregretbound banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregretbound the deterministic good-event radius plus the uniform bad-event contribution. definition compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_nonneg","label":"decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_nonneg","description":"Expected absolute realized-behavior regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-6a97c91afbf7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6868,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:334"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : 0 <= decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret mdp initialState initialTable defaultState baseVisitFloor n","missing":[],"search":"decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregret_nonneg expected absolute realized-behavior regret is nonnegative. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","label":"decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","description":"The explicit expected-absolute bound is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-978be9fc5f44","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6869,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:345"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound mdp baseVisitFloor n","missing":[],"search":"decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregretbound_nonneg banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregretbound_nonneg the explicit expected-absolute bound is nonnegative. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_le_bound","label":"decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_le_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_le_bound","description":"The common-space expected absolute regret is controlled by the sharp good-event radius plus the uniform `2H` envelope times the failure probability.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-d66b7aa3c4fb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6870,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:360"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_le_bound (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : decayingExplorationEpisodewiseE…","missing":[],"search":"decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregret_le_bound banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregret_le_bound the common-space expected absolute regret is controlled by the sharp good-event radius plus the uniform `2h` envelope times the failure probability. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationRealizedFailureBudget_toReal_tendsto_zero","label":"decayingExplorationRealizedFailureBudget_toReal_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationRealizedFailureBudget_toReal_tendsto_zero","description":"The real-valued failure budget tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-9443f38c7cd3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6871,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:467"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationRealizedFailureBudget_toReal_tendsto_zero : Tendsto (fun n => (AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget n).toReal) atTop (nhds 0)","missing":[],"search":"decayingexplorationrealizedfailurebudget_toreal_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationrealizedfailurebudget_toreal_tendsto_zero the real-valued failure budget tends to zero. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","label":"decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","description":"The explicit expected-absolute envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-cd6aa1629cb7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6872,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:477"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound mdp baseVisitFloor) atTop (nhds 0)","missing":[],"search":"decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseexpectedabsoluterealizedbehaviorregretbound_tendsto_zero the explicit expected-absolute envelope tends to zero. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","label":"exploratorySource_decayingExplorationEpisodewiseCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","description":"Terminal expected-consistency theorem: all coordinates are integrable, every finite expected absolute regret obeys the explicit bound, and those expectations converge to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspaceexpectedconsistency/index.html#decl-369dd28edf0c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","order":6873,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency.lean:495"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_decayingExplorationEpisodewiseCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFlo…","missing":[],"search":"exploratorysource_decayingexplorationepisodewisecommonmeasure_integrable_expectedabsoluterealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_decayingexplorationepisodewisecommonmeasure_integrable_expectedabsoluterealizedbehaviorregret_tendsto_zero terminal expected-consistency theorem: all coordinates are integrable, every finite expected absolute regret obeys the explicit bound, and those expectations converge to zero. theorem compiled","shard":"modules/0a573959ce6556ae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.memLp_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","label":"memLp_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.memLp_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","description":"Every scheduled common-space realized-regret coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-b94945a20a50","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6874,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : MemLp (decayingExplorationEpisodewiseRealizedBehaviorRegretProcess mdp initialState initialTable defaultState baseVisitFloor n) 1 (decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor)","missing":[],"search":"memlp_one_decayingexplorationepisodewiserealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.memlp_one_decayingexplorationepisodewiserealizedbehaviorregretprocess every scheduled common-space realized-regret coordinate belongs to `l1`. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_eq","label":"eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_eq","description":"At exponent one, the `eLpNorm` is exactly the lifted expected absolute regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-df24cef53d61","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6875,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:48"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : eLpNorm (decayingExplorationEpisodewiseRealizedBehaviorRegretProcess mdp initialState initialTable defaultState baseVisitFloor n) 1 (decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor) = ENNReal.ofReal (decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret mdp initialState initialTable defaultState baseVisitFloor n)","missing":[],"search":"elpnorm_one_decayingexplorationepisodewiserealizedbehaviorregretprocess_eq banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.elpnorm_one_decayingexplorationepisodewiserealizedbehaviorregretprocess_eq at exponent one, the `elpnorm` is exactly the lifted expected absolute regret. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_tendsto_zero","label":"eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_tendsto_zero","description":"The exponent-one extended `Lp` norm of the scheduled process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-1ea33958d96b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6876,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => eLpNorm (decayin…","missing":[],"search":"elpnorm_one_decayingexplorationepisodewiserealizedbehaviorregretprocess_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.elpnorm_one_decayingexplorationepisodewiserealizedbehaviorregretprocess_tendsto_zero the exponent-one extended `lp` norm of the scheduled process tends to zero. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","label":"eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","description":"Canonical `L1` convergence form: the norm of the difference from zero tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-f38334e52b96","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6877,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => eLpNorm…","missing":[],"search":"elpnorm_one_decayingexplorationepisodewiserealizedbehaviorregretprocess_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.elpnorm_one_decayingexplorationepisodewiserealizedbehaviorregretprocess_sub_zero_tendsto_zero canonical `l1` convergence form: the norm of the difference from zero tends to zero. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp","label":"decayingExplorationEpisodewiseRealizedBehaviorRegretLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp","description":"The scheduled realized-regret process represented as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-8274e88fb498","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6878,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseRealizedBehaviorRegretLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : Lp Real 1 (decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor)","missing":[],"search":"decayingexplorationepisodewiserealizedbehaviorregretlp banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiserealizedbehaviorregretlp the scheduled realized-regret process represented as an `lp real 1` value. definition compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp_coeFn_ae_eq","label":"decayingExplorationEpisodewiseRealizedBehaviorRegretLp_coeFn_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp_coeFn_ae_eq","description":"The named `Lp` coordinate represents the original process almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-5dc41397f6e9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6879,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseRealizedBehaviorRegretLp_coeFn_ae_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : (decayingExplorationEpisodewiseRealizedBehaviorRegretLp mdp initialState initialTable defaultState baseVisitFloor hrewardBound n : DecayingExplorationEpisodewiseWindowSpace mdp baseVisitFloor -> Real) =ᵐ[ decayingExplorationEpisodewiseCommonMeasure mdp initialState initialTable defaultState baseVisitFloor] decayingExplorationEpisodewiseRealizedBehaviorRegretProcess mdp initialState initialTable defaultState baseVisitFloor n","missing":[],"search":"decayingexplorationepisodewiserealizedbehaviorregretlp_coefn_ae_eq banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiserealizedbehaviorregretlp_coefn_ae_eq the named `lp` coordinate represents the original process almost everywhere. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp_tendsto_zero","label":"decayingExplorationEpisodewiseRealizedBehaviorRegretLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp_tendsto_zero","description":"The named `Lp Real 1` scheduled process converges to zero in the `Lp` topology.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-8aba1eac67fa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6880,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseRealizedBehaviorRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationEpisodewiseRealizedBeha…","missing":[],"search":"decayingexplorationepisodewiserealizedbehaviorregretlp_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiserealizedbehaviorregretlp_tendsto_zero the named `lp real 1` scheduled process converges to zero in the `lp` topology. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","label":"exploratorySource_decayingExplorationEpisodewiseCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","description":"Terminal `L1` theorem: coordinate membership, exact norms, canonical norm convergence, convergence in the `Lp` topology, and the induced convergence in measure all hold on the same common probability space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewisecommonspacel1consistency/index.html#decl-0d76849a7e6a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","order":6881,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency.lean:226"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_decayingExplorationEpisodewiseCommonMeasure_memLp_eLpNorm_L1_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall n, MemLp (decayingE…","missing":[],"search":"exploratorysource_decayingexplorationepisodewisecommonmeasure_memlp_elpnorm_l1_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_decayingexplorationepisodewisecommonmeasure_memlp_elpnorm_l1_tendsto_zero terminal `l1` theorem: coordinate membership, exact norms, canonical norm convergence, convergence in the `lp` topology, and the induced convergence in measure all hold on the same common probability space. theorem compiled","shard":"modules/2f23217ea804e4e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_episode","label":"measurable_episode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_episode","description":"A complete episode row is a measurable coordinate of a finite batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-0b7761d57833","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6882,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_episode {mdp : MDP State Action} {episodes : Nat} (episode : Fin episodes) : Measurable (fun batch : EpisodeBatch mdp episodes => batch episode)","missing":[],"search":"measurable_episode banditrlproof.finitehorizonrl.episodebatch.measurable_episode a complete episode row is a measurable coordinate of a finite batch. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeRowOfTrajectory","label":"iIndepFun_episodeRowOfTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeRowOfTrajectory","description":"Complete generated episode rows are independent product coordinates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-f3a788a0b410","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6883,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_episodeRowOfTrajectory {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.iIndepFun (fun episode trajectories => fun stage => mdp.episodeStepOfTrajectory (trajectories episode) stage) (policy.iidTrajectoryFamilyMeasure initialState episodes)","missing":[],"search":"iindepfun_episoderowoftrajectory banditrlproof.finitehorizonrl.markovpolicy.iindepfun_episoderowoftrajectory complete generated episode rows are independent product coordinates. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_episode","label":"iIndepFun_iidEpisodeBatch_episode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_episode","description":"Complete episode rows remain independent after mapping trajectories to a batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-4551ea161b79","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6884,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:63"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_iidEpisodeBatch_episode {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.iIndepFun (fun episode (batch : EpisodeBatch mdp episodes) => batch episode) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iindepfun_iidepisodebatch_episode banditrlproof.finitehorizonrl.markovpolicy.iindepfun_iidepisodebatch_episode complete episode rows remain independent after mapping trajectories to a batch. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_episodeReturn","label":"iIndepFun_iidEpisodeBatch_episodeReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_episodeReturn","description":"Full episode returns are independent across iid batch coordinates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-f9fca43f8ebe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6885,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_iidEpisodeBatch_episodeReturn {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.iIndepFun (fun episode (batch : EpisodeBatch mdp episodes) => EpisodeBatch.episodeReturn batch episode) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iindepfun_iidepisodebatch_episodereturn banditrlproof.finitehorizonrl.markovpolicy.iindepfun_iidepisodebatch_episodereturn full episode returns are independent across iid batch coordinates. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_episodeReturn_iidEpisodeBatchMeasure","label":"integral_episodeReturn_iidEpisodeBatchMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_episodeReturn_iidEpisodeBatchMeasure","description":"Every episode return in an iid batch has the common trajectory-return mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-e321f14a1740","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6886,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_episodeReturn_iidEpisodeBatchMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) : integral (policy.iidEpisodeBatchMeasure initialState episodes) (fun batch => EpisodeBatch.episodeReturn batch episode) = integral (policy.trajectoryMeasure initialState) mdp.cumulativeReward","missing":[],"search":"integral_episodereturn_iidepisodebatchmeasure banditrlproof.finitehorizonrl.markovpolicy.integral_episodereturn_iidepisodebatchmeasure every episode return in an iid batch has the common trajectory-return mean. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturnVarianceProxy","label":"episodeReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturnVarianceProxy","description":"One full episode's bounded-return Hoeffding proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-1406345417b4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6887,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def episodeReturnVarianceProxy (mdp : MDP State Action) : NNReal","missing":[],"search":"episodereturnvarianceproxy banditrlproof.finitehorizonrl.markovpolicy.episodereturnvarianceproxy one full episode's bounded-return hoeffding proxy. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturn_centered_hasSubgaussianMGF","label":"episodeReturn_centered_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturn_centered_hasSubgaussianMGF","description":"Each bounded centered episode return is sub-Gaussian with proxy `horizon^2`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-65157982a0ec","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6888,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:148"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeReturn_centered_hasSubgaussianMGF {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ProbabilityTheory.HasSubgaussianMGF (fun batch : EpisodeBatch mdp episodes => EpisodeBatch.episodeReturn batch episode - integral (policy.trajectoryMeasure initialState) mdp.cumulativeReward) (episodeReturnVarianceProxy mdp) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"episodereturn_centered_hassubgaussianmgf banditrlproof.finitehorizonrl.markovpolicy.episodereturn_centered_hassubgaussianmgf each bounded centered episode return is sub-gaussian with proxy `horizon^2`. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodewiseBatchReturnVarianceProxy","label":"episodewiseBatchReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodewiseBatchReturnVarianceProxy","description":"Sum of the independent full-episode return proxies in one batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-2bc619def9ff","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6889,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:178"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def episodewiseBatchReturnVarianceProxy (mdp : MDP State Action) (episodes : Nat) : NNReal","missing":[],"search":"episodewisebatchreturnvarianceproxy banditrlproof.finitehorizonrl.markovpolicy.episodewisebatchreturnvarianceproxy sum of the independent full-episode return proxies in one batch. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturnVarianceProxy_coe","label":"episodeReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturnVarianceProxy_coe","description":"theorem episodeReturnVarianceProxy_coe (mdp : MDP State Action) : ((episodeReturnVarianceProxy mdp : NNReal) : Real) = ((mdp.horizon : Nat) : Real) ^ 2","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-32732b602a9d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6890,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:185"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeReturnVarianceProxy_coe (mdp : MDP State Action) : ((episodeReturnVarianceProxy mdp : NNReal) : Real) = ((mdp.horizon : Nat) : Real) ^ 2","missing":[],"search":"episodereturnvarianceproxy_coe banditrlproof.finitehorizonrl.markovpolicy.episodereturnvarianceproxy_coe theorem episodereturnvarianceproxy_coe (mdp : mdp state action) : ((episodereturnvarianceproxy mdp : nnreal) : real) = ((mdp.horizon : nat) : real) ^ 2 theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodewiseBatchReturnVarianceProxy_coe","label":"episodewiseBatchReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodewiseBatchReturnVarianceProxy_coe","description":"theorem episodewiseBatchReturnVarianceProxy_coe (mdp : MDP State Action) (episodes : Nat) : ((episodewiseBatchReturnVarianceProxy mdp episodes : NNReal) : Real) = (episodes : Real) * ((mdp.horizon : Nat) : Real) ^ 2","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-c8c7ec7f223d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6891,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:197"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseBatchReturnVarianceProxy_coe (mdp : MDP State Action) (episodes : Nat) : ((episodewiseBatchReturnVarianceProxy mdp episodes : NNReal) : Real) = (episodes : Real) * ((mdp.horizon : Nat) : Real) ^ 2","missing":[],"search":"episodewisebatchreturnvarianceproxy_coe banditrlproof.finitehorizonrl.markovpolicy.episodewisebatchreturnvarianceproxy_coe theorem episodewisebatchreturnvarianceproxy_coe (mdp : mdp state action) (episodes : nat) : ((episodewisebatchreturnvarianceproxy mdp episodes : nnreal) : real) = (episodes : real) * ((mdp.horizon : nat) : real) ^ 2 theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.totalReturn_centered_episodewise_hasSubgaussianMGF","label":"totalReturn_centered_episodewise_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.totalReturn_centered_episodewise_hasSubgaussianMGF","description":"The centered total batch return has the sharp episodewise proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-d4b8aba6eee7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6892,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:204"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem totalReturn_centered_episodewise_hasSubgaussianMGF {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ProbabilityTheory.HasSubgaussianMGF (fun batch : EpisodeBatch mdp episodes => EpisodeBatch.totalReturn batch - integral (policy.iidEpisodeBatchMeasure initialState episodes) EpisodeBatch.totalReturn) (episodewiseBatchReturnVarianceProxy mdp episodes) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"totalreturn_centered_episodewise_hassubgaussianmgf banditrlproof.finitehorizonrl.markovpolicy.totalreturn_centered_episodewise_hassubgaussianmgf the centered total batch return has the sharp episodewise proxy. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_episodewise_hasCondSubgaussianMGF","label":"successorReturnIncrement_succ_episodewise_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_episodewise_hasCondSubgaussianMGF","description":"The successor batch increment inherits the sharp episodewise batch proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-009ca9178b51","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6893,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:253"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorReturnIncrement_succ_episodewise_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ProbabilityTheory.HasCondSubgaussianMGF (Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes) n) ((Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes)).le n) (source.successorReturnIncrement (n + 1)) (MarkovPolicy.episodewiseBatchReturnVarianceProxy mdp episodes) source.trajectoryMeasure","missing":[],"search":"successorreturnincrement_succ_episodewise_hascondsubgaussianmgf banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnincrement_succ_episodewise_hascondsubgaussianmgf the successor batch increment inherits the sharp episodewise batch proxy. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseCumulativeSuccessorReturnVarianceProxy","label":"episodewiseCumulativeSuccessorReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseCumulativeSuccessorReturnVarianceProxy","description":"Sum of the sharp batch proxies over the successor rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-0de0eb1df665","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6894,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:369"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def episodewiseCumulativeSuccessorReturnVarianceProxy (mdp : MDP State Action) (episodes rounds : Nat) : NNReal","missing":[],"search":"episodewisecumulativesuccessorreturnvarianceproxy banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisecumulativesuccessorreturnvarianceproxy sum of the sharp batch proxies over the successor rounds. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseCumulativeSuccessorReturnVarianceProxy_coe","label":"episodewiseCumulativeSuccessorReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseCumulativeSuccessorReturnVarianceProxy_coe","description":"The sharp cumulative proxy is `rounds * episodes * horizon^2`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-d0392567c615","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6895,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:380"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseCumulativeSuccessorReturnVarianceProxy_coe (mdp : MDP State Action) (episodes rounds : Nat) : ((episodewiseCumulativeSuccessorReturnVarianceProxy mdp episodes rounds : NNReal) : Real) = (rounds : Real) * (episodes : Real) * (mdp.horizon : Real) ^ 2","missing":[],"search":"episodewisecumulativesuccessorreturnvarianceproxy_coe banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisecumulativesuccessorreturnvarianceproxy_coe the sharp cumulative proxy is `rounds * episodes * horizon^2`. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseBatchReturnVarianceProxy_pos","label":"episodewiseBatchReturnVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseBatchReturnVarianceProxy_pos","description":"theorem episodewiseBatchReturnVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) : 0 < MarkovPolicy.episodewiseBatchReturnVarianceProxy mdp episodes","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-e4799faef390","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6896,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:397"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseBatchReturnVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) : 0 < MarkovPolicy.episodewiseBatchReturnVarianceProxy mdp episodes","missing":[],"search":"episodewisebatchreturnvarianceproxy_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisebatchreturnvarianceproxy_pos theorem episodewisebatchreturnvarianceproxy_pos (mdp : mdp state action) (episodes : nat) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) : 0 < markovpolicy.episodewisebatchreturnvarianceproxy mdp episodes theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_episodewiseCumulativeSuccessorReturnDeviation_abs_tail_le","label":"trajectoryMeasure_episodewiseCumulativeSuccessorReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_episodewiseCumulativeSuccessorReturnDeviation_abs_tail_le","description":"Two-sided adaptive return tail with the sharp episodewise proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-d8f43a7a2890","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6897,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:409"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_episodewiseCumulativeSuccessorReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectoryMeasure {trajectory | Concentration.subGaussianSumConfidenceRadius (episodewiseCumulativeSuccessorReturnVarianceProxy mdp episodes rounds) delta <= |source.cumulativeSuccessorReturnDeviation rounds trajectory|} <= ENNReal.ofReal delta","missing":[],"search":"trajectorymeasure_episodewisecumulativesuccessorreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_episodewisecumulativesuccessorreturndeviation_abs_tail_le two-sided adaptive return tail with the sharp episodewise proxy. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseSuccessorReturnDeviationBadEvent","label":"episodewiseSuccessorReturnDeviationBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseSuccessorReturnDeviationBadEvent","description":"Sharp successor-return deviation event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-ea88cc396b37","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6898,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:460"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def episodewiseSuccessorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"episodewisesuccessorreturndeviationbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisesuccessorreturndeviationbadevent sharp successor-return deviation event. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_episodewiseSuccessorReturnDeviationBadEvent","label":"measurableSet_episodewiseSuccessorReturnDeviationBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_episodewiseSuccessorReturnDeviationBadEvent","description":"theorem measurableSet_episodewiseSuccessorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : MeasurableSet (source.episodewiseSuccessorReturnDeviationBadEvent rounds delta)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-d519a5274e27","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6899,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:472"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_episodewiseSuccessorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : MeasurableSet (source.episodewiseSuccessorReturnDeviationBadEvent rounds delta)","missing":[],"search":"measurableset_episodewisesuccessorreturndeviationbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurableset_episodewisesuccessorreturndeviationbadevent theorem measurableset_episodewisesuccessorreturndeviationbadevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (delta : real) : measurableset (source.episodewisesuccessorreturndeviationbadevent rounds delta) theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_episodewiseSuccessorReturnDeviationBadEvent_le","label":"trajectoryMeasure_episodewiseSuccessorReturnDeviationBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_episodewiseSuccessorReturnDeviationBadEvent_le","description":"theorem trajectoryMeasure_episodewiseSuccessorReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes)…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-0c2a06536c73","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6900,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:482"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_episodewiseSuccessorReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectoryMeasure (source.episodewiseSuccessorReturnDeviationBadEvent rounds delta) <= ENNReal.ofReal delta","missing":[],"search":"trajectorymeasure_episodewisesuccessorreturndeviationbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_episodewisesuccessorreturndeviationbadevent_le theorem trajectorymeasure_episodewisesuccessorreturndeviationbadevent_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] [standardborelspace (episodebatchtrajectory mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectorymeasure (source.episodewisesuccessorreturndeviationbadevent rounds delta) <= ennreal.ofreal delta theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_episodewise_transport","label":"trajectoryMeasure_expected_to_realized_successor_average_regret_episodewise_transport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_episodewise_transport","description":"Transport an expected-regret certificate through the sharp return event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-28be22b56d7d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6901,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:500"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_expected_to_realized_successor_average_regret_episodewise_transport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (countBadEvent : Set (EpisodeBatchTrajectory mdp episodes)) (expectedBound : Real) (Good : EpisodeBatchTrajectory mdp episodes -> Prop) (hcountMeasurable : MeasurableSet countBadEvent) (hcountTail : source.trajectoryMeasure countBadEvent <= ENNReal.ofReal delta) (hcountGood : forall trajectory,…","missing":[],"search":"trajectorymeasure_expected_to_realized_successor_average_regret_episodewise_transport banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_expected_to_realized_successor_average_regret_episodewise_transport transport an expected-regret certificate through the sharp return event. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius","label":"episodewiseNormalizedSuccessorReturnConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius","description":"Sharp return-deviation radius after normalization by all sampled episodes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-0ddd6970be16","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6902,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:591"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def episodewiseNormalizedSuccessorReturnConfidenceRadius (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) : Real","missing":[],"search":"episodewisenormalizedsuccessorreturnconfidenceradius banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisenormalizedsuccessorreturnconfidenceradius sharp return-deviation radius after normalization by all sampled episodes. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_nonneg","label":"episodewiseNormalizedSuccessorReturnConfidenceRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_nonneg","description":"theorem episodewiseNormalizedSuccessorReturnConfidenceRadius_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) : 0 <= episodewiseNormalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-3c46ce44e344","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6903,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:601"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseNormalizedSuccessorReturnConfidenceRadius_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) : 0 <= episodewiseNormalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta","missing":[],"search":"episodewisenormalizedsuccessorreturnconfidenceradius_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisenormalizedsuccessorreturnconfidenceradius_nonneg theorem episodewisenormalizedsuccessorreturnconfidenceradius_nonneg (mdp : mdp state action) (episodes rounds : nat) (delta : real) : 0 <= episodewisenormalizedsuccessorreturnconfidenceradius mdp episodes rounds delta theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_eq","label":"episodewiseNormalizedSuccessorReturnConfidenceRadius_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_eq","description":"Exact normalized radius: episode count now improves concentration.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-609240eae9a7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6904,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:614"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseNormalizedSuccessorReturnConfidenceRadius_eq (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) (hepisodes : 0 < episodes) (hrounds : 0 < rounds) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : episodewiseNormalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta = (mdp.horizon : Real) * Real.sqrt (2 * Real.log (2 / delta) / ((episodes : Real) * (rounds : Real)))","missing":[],"search":"episodewisenormalizedsuccessorreturnconfidenceradius_eq banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisenormalizedsuccessorreturnconfidenceradius_eq exact normalized radius: episode count now improves concentration. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_le_normalized","label":"episodewiseNormalizedSuccessorReturnConfidenceRadius_le_normalized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_le_normalized","description":"The sharp radius is no larger than the compiled whole-batch radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-7a20d4d4be60","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6905,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:671"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseNormalizedSuccessorReturnConfidenceRadius_le_normalized (mdp : MDP State Action) (episodes rounds : Nat) (delta : Real) (hepisodes : 0 < episodes) (hrounds : 0 < rounds) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : episodewiseNormalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta <= normalizedSuccessorReturnConfidenceRadius mdp episodes rounds delta","missing":[],"search":"episodewisenormalizedsuccessorreturnconfidenceradius_le_normalized banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisenormalizedsuccessorreturnconfidenceradius_le_normalized the sharp radius is no larger than the compiled whole-batch radius. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","label":"episodewiseNormalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","description":"theorem episodewiseNormalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope (mdp : MDP State Action) (episodes : Nat) (n : Nat) (hepisodes : 0 < episodes) : episodewiseNormalizedSuccessorReturnConfidenceRadius mdp episodes (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n) <= decayingExplorationReturnRadiusEnvelope mdp n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-41f1317ed005","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6906,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:698"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodewiseNormalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope (mdp : MDP State Action) (episodes : Nat) (n : Nat) (hepisodes : 0 < episodes) : episodewiseNormalizedSuccessorReturnConfidenceRadius mdp episodes (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n) <= decayingExplorationReturnRadiusEnvelope mdp n","missing":[],"search":"episodewisenormalizedsuccessorreturnconfidenceradius_le_decayingenvelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.episodewisenormalizedsuccessorreturnconfidenceradius_le_decayingenvelope theorem episodewisenormalizedsuccessorreturnconfidenceradius_le_decayingenvelope (mdp : mdp state action) (episodes : nat) (n : nat) (hepisodes : 0 < episodes) : episodewisenormalizedsuccessorreturnconfidenceradius mdp episodes (decayingexplorationrounds mdp n) (vanishingaverageconfidencedelta n) <= decayingexplorationreturnradiusenvelope mdp n theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseNormalizedReturnRadius_tendsto_zero","label":"decayingExplorationEpisodewiseNormalizedReturnRadius_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseNormalizedReturnRadius_tendsto_zero","description":"theorem decayingExplorationEpisodewiseNormalizedReturnRadius_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) : Tendsto (fun n => episodewiseNormalizedSuccessorReturnConfidenceRadius mdp (decayingExplorationScheduledEpisodes mdp baseVisitFloor n) (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n)) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-0d712a5ea31c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6907,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:717"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseNormalizedReturnRadius_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) : Tendsto (fun n => episodewiseNormalizedSuccessorReturnConfidenceRadius mdp (decayingExplorationScheduledEpisodes mdp baseVisitFloor n) (decayingExplorationRounds mdp n) (vanishingAverageConfidenceDelta n)) atTop (nhds 0)","missing":[],"search":"decayingexplorationepisodewisenormalizedreturnradius_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationepisodewisenormalizedreturnradius_tendsto_zero theorem decayingexplorationepisodewisenormalizedreturnradius_tendsto_zero (mdp : mdp state action) (basevisitfloor : real) : tendsto (fun n => episodewisenormalizedsuccessorreturnconfidenceradius mdp (decayingexplorationscheduledepisodes mdp basevisitfloor n) (decayingexplorationrounds mdp n) (vanishingaverageconfidencedelta n)) attop (nhds 0) theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound","label":"decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound","description":"Expected behavior bound plus the sharp episodewise return radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-ad968a7561dc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6908,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:736"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationepisodewiseaveragerealizedbehaviorregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationepisodewiseaveragerealizedbehaviorregretbound expected behavior bound plus the sharp episodewise return radius. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_nonneg","label":"decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_nonneg","description":"theorem decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound mdp baseVisitFloor n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-39bec56ee96e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6909,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:745"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound mdp baseVisitFloor n","missing":[],"search":"decayingexplorationepisodewiseaveragerealizedbehaviorregretbound_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationepisodewiseaveragerealizedbehaviorregretbound_nonneg theorem decayingexplorationepisodewiseaveragerealizedbehaviorregretbound_nonneg (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) (n : nat) : 0 <= decayingexplorationepisodewiseaveragerealizedbehaviorregretbound mdp basevisitfloor n theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_tendsto_zero","label":"decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_tendsto_zero","description":"theorem decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound mdp baseVisitFloor) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-6ec9dc2da3f8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6910,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:765"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound mdp baseVisitFloor) atTop (nhds 0)","missing":[],"search":"decayingexplorationepisodewiseaveragerealizedbehaviorregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationepisodewiseaveragerealizedbehaviorregretbound_tendsto_zero theorem decayingexplorationepisodewiseaveragerealizedbehaviorregretbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (decayingexplorationepisodewiseaveragerealizedbehaviorregretbound mdp basevisitfloor) attop (nhds 0) theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseRealizedFailureAndRegretBound_tendsto_zero","label":"decayingExplorationEpisodewiseRealizedFailureAndRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseRealizedFailureAndRegretBound_tendsto_zero","description":"theorem decayingExplorationEpisodewiseRealizedFailureAndRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (decayingExplorationRealizedFailureBudget n, decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound mdp baseVisitFloor n)) atTop (nhds (0, 0))","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-090a734c33f5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6911,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:777"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationEpisodewiseRealizedFailureAndRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (decayingExplorationRealizedFailureBudget n, decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound mdp baseVisitFloor n)) atTop (nhds (0, 0))","missing":[],"search":"decayingexplorationepisodewiserealizedfailureandregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.decayingexplorationepisodewiserealizedfailureandregretbound_tendsto_zero theorem decayingexplorationepisodewiserealizedfailureandregretbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (fun n => (decayingexplorationrealizedfailurebudget n, decayingexplorationepisodewiseaveragerealizedbehaviorregretbound mdp basevisitfloor n)) attop (nhds (0, 0)) theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretViolationSet","label":"decayingExplorationEpisodewiseAverageRealizedBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretViolationSet","description":"Realized-regret violation set for the sharp episodewise return certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-2d5fd2f1f103","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6912,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:796"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationEpisodewiseAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Set (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","missing":[],"search":"decayingexplorationepisodewiseaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.decayingexplorationepisodewiseaveragerealizedbehaviorregretviolationset realized-regret violation set for the sharp episodewise return certificate. definition compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency","description":"One scheduled finite window with episodewise return concentration. The realized violation set is covered by the measurable count/return union while the good side retains optimism and the sharp normalized return radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-ab338a7b3759","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6913,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:828"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := A…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationepisodewiseaveragerealizedbehaviorconsistency banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationepisodewiseaveragerealizedbehaviorconsistency one scheduled finite window with episodewise return concentration. the realized violation set is covered by the measurable count/return union while the good side retains optimism and the sharp normalized return radius. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency_allWindows","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency_allWindows","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency_allWindows","description":"All scheduled finite windows with their indexed Borel witnesses, together with the joint scalar limit. The sample spaces may vary with the schedule index.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeepisodewiserealizedbehaviorconsistency/index.html#decl-a1b4d14d50ed","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","order":6914,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency.lean:986"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency_allWindows (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloo…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_decayingexplorationepisodewiseaveragerealizedbehaviorconsistency_allwindows banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_decayingexplorationepisodewiseaveragerealizedbehaviorconsistency_allwindows all scheduled finite windows with their indexed borel witnesses, together with the joint scalar limit. the sample spaces may vary with the schedule index. theorem compiled","shard":"modules/a8dc325512317251.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionPMF_apply","label":"exploratoryActionPMF_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionPMF_apply","description":"theorem exploratoryActionPMF_apply {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (action : Action) : table.exploratoryActionPMF explorationRate hexplorationRate stage state action = (explorationRate : ENNReal) * (Fintype.card Action : ENNReal)⁻¹ + ((1 - explorationRate : NNReal) : EN…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-74c99a1fd93c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6915,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryActionPMF_apply {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (action : Action) : table.exploratoryActionPMF explorationRate hexplorationRate stage state action = (explorationRate : ENNReal) * (Fintype.card Action : ENNReal)⁻¹ + ((1 - explorationRate : NNReal) : ENNReal) * if action = table stage state then 1 else 0","missing":[],"search":"exploratoryactionpmf_apply banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryactionpmf_apply theorem exploratoryactionpmf_apply {mdp : mdp state action} (table : deterministicmarkovpolicytable mdp) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (stage : fin mdp.horizon) (state : state) (action : action) : table.exploratoryactionpmf explorationrate hexplorationrate stage state action = (explorationrate : ennreal) * (fintype.card action : ennreal)⁻¹ + ((1 - explorationrate : nnreal) : ennreal) * if action = table stage state then 1 else 0 theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.integral_exploratoryActionPMF","label":"integral_exploratoryActionPMF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.integral_exploratoryActionPMF","description":"theorem integral_exploratoryActionPMF {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (f : Action -> Real) : (∫ action, f action ∂ (table.exploratoryActionPMF explorationRate hexplorationRate stage state).toMeasure) = (explorationRate : Real) * (∑ action, f action) / (Fintype.card Acti…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-a2d1e5536464","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6916,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_exploratoryActionPMF {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (f : Action -> Real) : (∫ action, f action ∂ (table.exploratoryActionPMF explorationRate hexplorationRate stage state).toMeasure) = (explorationRate : Real) * (∑ action, f action) / (Fintype.card Action : Real) + (1 - (explorationRate : Real)) * f (table stage state)","missing":[],"search":"integral_exploratoryactionpmf banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.integral_exploratoryactionpmf theorem integral_exploratoryactionpmf {mdp : mdp state action} (table : deterministicmarkovpolicytable mdp) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (stage : fin mdp.horizon) (state : state) (f : action -> real) : (∫ action, f action ∂ (table.exploratoryactionpmf explorationrate hexplorationrate stage state).tomeasure) = (explorationrate : real) * (∑ action, f action) / (fintype.card action : real) + (1 - (explorationrate : real)) * f (table stage state) theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.selected_sub_integral_exploratoryActionPMF_le","label":"selected_sub_integral_exploratoryActionPMF_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.selected_sub_integral_exploratoryActionPMF_le","description":"theorem selected_sub_integral_exploratoryActionPMF_le {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (f : Action -> Real) (bound : Real) (hbound : forall action, |f action| <= bound) : f (table stage state) - (∫ action, f action ∂ (table.exploratoryActionPMF explorationRate hexplorati…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-f66a6eed3b8f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6917,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selected_sub_integral_exploratoryActionPMF_le {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (f : Action -> Real) (bound : Real) (hbound : forall action, |f action| <= bound) : f (table stage state) - (∫ action, f action ∂ (table.exploratoryActionPMF explorationRate hexplorationRate stage state).toMeasure) <= 2 * (explorationRate : Real) * bound","missing":[],"search":"selected_sub_integral_exploratoryactionpmf_le banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.selected_sub_integral_exploratoryactionpmf_le theorem selected_sub_integral_exploratoryactionpmf_le {mdp : mdp state action} (table : deterministicmarkovpolicytable mdp) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (stage : fin mdp.horizon) (state : state) (f : action -> real) (bound : real) (hbound : forall action, |f action| <= bound) : f (table stage state) - (∫ action, f action ∂ (table.exploratoryactionpmf explorationrate hexplorationrate stage state).tomeasure) <= 2 * (explorationrate : real) * bound theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_abs_le","label":"bellmanQ_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_abs_le","description":"A bounded continuation gives a bounded one-step action value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-e58ea8db2867","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6918,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:151"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanQ_abs_le (mdp : MDP State Action) (rewardBound continuationBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) {value : State -> Real} (hvalue : forall state, |value state| <= continuationBound) (state : State) (action : Action) : |mdp.bellmanQ value state action| <= rewardBound + continuationBound","missing":[],"search":"bellmanq_abs_le banditrlproof.finitehorizonrl.mdp.bellmanq_abs_le a bounded continuation gives a bounded one-step action value. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue_sub_le_const","label":"transitionValue_sub_le_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionValue_sub_le_const","description":"A pointwise continuation difference bound transports through one transition kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-ed7e9ae05a54","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6919,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:172"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionValue_sub_le_const (mdp : MDP State Action) {left right : State -> Real} {bound : Real} (hbound : forall state, left state - right state <= bound) (state : State) (action : Action) : mdp.transitionValue left state action - mdp.transitionValue right state action <= bound","missing":[],"search":"transitionvalue_sub_le_const banditrlproof.finitehorizonrl.mdp.transitionvalue_sub_le_const a pointwise continuation difference bound transports through one transition kernel. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_abs_le","label":"valueRemaining_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_abs_le","description":"Every bounded-reward policy value is bounded by remaining horizon times the reward bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-873364244743","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6920,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:199"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueRemaining_abs_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : |policy.valueRemaining remaining hremaining state| <= (remaining : Real) * rewardBound","missing":[],"search":"valueremaining_abs_le banditrlproof.finitehorizonrl.markovpolicy.valueremaining_abs_le every bounded-reward policy value is bounded by remaining horizon times the reward bound. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy_bellman","label":"toMarkovPolicy_bellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy_bellman","description":"Bellman evaluation of a deterministic table selects exactly its table action.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-cb09a77e9a95","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6921,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:234"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem toMarkovPolicy_bellman {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : table.toMarkovPolicy.bellman stage value state = mdp.bellmanQ value state (table stage state)","missing":[],"search":"tomarkovpolicy_bellman banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.tomarkovpolicy_bellman bellman evaluation of a deterministic table selects exactly its table action. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_bellman","label":"exploratoryPolicy_bellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_bellman","description":"Bellman evaluation of the exploratory table is its explicit PMF mixture.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-8575f0d282e9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6922,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:247"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPolicy_bellman {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : (table.exploratoryPolicy explorationRate hexplorationRate).bellman stage value state = (explorationRate : Real) * (∑ action, mdp.bellmanQ value state action) / (Fintype.card Action : Real) + (1 - (explorationRate : Real)) * mdp.bellmanQ value state (table stage state)","missing":[],"search":"exploratorypolicy_bellman banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratorypolicy_bellman bellman evaluation of the exploratory table is its explicit pmf mixture. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy_valueRemaining_sub_exploratoryPolicy_valueRemaining_le","label":"toMarkovPolicy_valueRemaining_sub_exploratoryPolicy_valueRemaining_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy_valueRemaining_sub_exploratoryPolicy_valueRemaining_le","description":"Exploration around a deterministic table loses at most the explicit quadratic-horizon charge. The charge is linear in the exploration rate and reward bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-6444e14ac0fe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6923,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem toMarkovPolicy_valueRemaining_sub_exploratoryPolicy_valueRemaining_le {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : table.toMarkovPolicy.valueRemaining remaining hremaining state - (table.exploratoryPolicy explorationRate hexplorationRate).valueRemaining remaining hremaining state <= (explorationRate : Real) * rewardBound * (remaining : Real) * ((remaining + 1 : Nat) : Real)","missing":[],"search":"tomarkovpolicy_valueremaining_sub_exploratorypolicy_valueremaining_le banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.tomarkovpolicy_valueremaining_sub_exploratorypolicy_valueremaining_le exploration around a deterministic table loses at most the explicit quadratic-horizon charge. the charge is linear in the exploration rate and reward bound. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_expectedRegret_le_toMarkovPolicy_expectedRegret_add_charge","label":"exploratoryPolicy_expectedRegret_le_toMarkovPolicy_expectedRegret_add_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_expectedRegret_le_toMarkovPolicy_expectedRegret_add_charge","description":"Exploratory-behavior expected regret is bounded by table regret plus its exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-945a9bc13579","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6924,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:356"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPolicy_expectedRegret_le_toMarkovPolicy_expectedRegret_add_charge {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) : (table.exploratoryPolicy explorationRate hexplorationRate).expectedRegret initialState <= table.toMarkovPolicy.expectedRegret initialState + (explorationRate : Real) * rewardBound * (mdp.horizon : Real) * ((mdp.horizon + 1 : Nat) : Real)","missing":[],"search":"exploratorypolicy_expectedregret_le_tomarkovpolicy_expectedregret_add_charge banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratorypolicy_expectedregret_le_tomarkovpolicy_expectedregret_add_charge exploratory-behavior expected regret is bounded by table regret plus its exploration charge. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryBehaviorRegretCharge","label":"exploratoryBehaviorRegretCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryBehaviorRegretCharge","description":"The per-policy price of uniform exploration over a bounded-reward horizon.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-6be39bff586d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6925,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:438"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryBehaviorRegretCharge (mdp : MDP State Action) (explorationRate : NNReal) (rewardBound : Real) : Real","missing":[],"search":"exploratorybehaviorregretcharge banditrlproof.finitehorizonrl.exploratorybehaviorregretcharge the per-policy price of uniform exploration over a bounded-reward horizon. definition compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret","label":"adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret","description":"Sum of expected regrets of exploratory behaviors centered on cumulative recommendations.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-f9e6c84f0c8e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6926,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:445"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) : Real","missing":[],"search":"adaptivecumulativeempiricaloptimisticexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticexploratorybehaviorexpectedregret sum of expected regrets of exploratory behaviors centered on cumulative recommendations. definition compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret","label":"adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret","description":"Average expected regret of those exploratory behaviors.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-7c162758e7bf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6927,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:458"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) : Real","missing":[],"search":"adaptivecumulativeempiricaloptimisticaverageexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticaverageexploratorybehaviorexpectedregret average expected regret of those exploratory behaviors. definition compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret_le","label":"adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret_le","description":"Cumulative exploratory behavior regret is bounded by recommendation regret plus one charge per round.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-b8bf55541f9b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6928,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:471"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) : adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret (initialState := initialState) trajectory defaultState countRadius explorationRate hexplorationRate rounds <= adaptiveCumulativeEmpiricalOptimisticRecommendedExpectedRegret (initialState := initialState) trajectory defaultState countRadius rounds + (rounds : Real) * exploratoryBehaviorRegretCharge mdp explorationRate rewardBound","missing":[],"search":"adaptivecumulativeempiricaloptimisticexploratorybehaviorexpectedregret_le banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticexploratorybehaviorexpectedregret_le cumulative exploratory behavior regret is bounded by recommendation regret plus one charge per round. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret_le","label":"adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret_le","description":"For a nonempty window, averaging removes the repeated-round factor from the charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-c534d715855b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6929,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:520"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) (hrounds : 0 < rounds) : adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret (initialState := initialState) trajectory defaultState countRadius explorationRate hexplorationRate rounds <= adaptiveCumulativeEmpiricalOptimisticAverageRecommendedExpectedRegret (initialState := initialState) trajectory defaultState countRadius rounds + exploratoryBehaviorRegretCharge mdp exploratio…","missing":[],"search":"adaptivecumulativeempiricaloptimisticaverageexploratorybehaviorexpectedregret_le banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticaverageexploratorybehaviorexpectedregret_le for a nonempty window, averaging removes the repeated-round factor from the charge. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegretBound","label":"vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegretBound","description":"The finite-window behavior-regret certificate adds the explicit exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-4b4c4ac3da1d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6930,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:563"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegretBound (mdp : MDP State Action) (n : Nat) (visitFloor : Real) (explorationRate : NNReal) : Real","missing":[],"search":"vanishingdeltascheduledaverageexploratorybehaviorexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledaverageexploratorybehaviorexpectedregretbound the finite-window behavior-regret certificate adds the explicit exploration charge. definition compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageExploratoryBehaviorBound_tendsto_charge","label":"vanishingDeltaScheduledAverageExploratoryBehaviorBound_tendsto_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageExploratoryBehaviorBound_tendsto_charge","description":"With a fixed exploration rate, the deterministic certificate tends to its exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-4f390a621403","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6931,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:570"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingDeltaScheduledAverageExploratoryBehaviorBound_tendsto_charge (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) (explorationRate : NNReal) : Filter.Tendsto (fun n => vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegretBound mdp n visitFloor explorationRate) Filter.atTop (nhds (exploratoryBehaviorRegretCharge mdp explorationRate 1))","missing":[],"search":"vanishingdeltascheduledaverageexploratorybehaviorbound_tendsto_charge banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledaverageexploratorybehaviorbound_tendsto_charge with a fixed exploration rate, the deterministic certificate tends to its exploration charge. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_policyAt_succ_eq_cumulativeOptimisticExploratoryPolicy","label":"exploratorySource_policyAt_succ_eq_cumulativeOptimisticExploratoryPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_policyAt_succ_eq_cumulativeOptimisticExploratoryPolicy","description":"The regret term indexed by `round` is the source's successor behavior at coordinate `round + 1`; the initial behavior at coordinate zero is intentionally excluded.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-a74be047ad25","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6932,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:594"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_policyAt_succ_eq_cumulativeOptimisticExploratoryPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) : (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).policyAt trajectory (round + 1) = ((cumulativeTransitionCountSummaryAt trajectory round) |>.countRadiusOptimisticPolicyTable mdp defaultState countRadius |>.exploratoryPolicy explorationRate hexplorationRate)","missing":[],"search":"exploratorysource_policyat_succ_eq_cumulativeoptimisticexploratorypolicy banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_policyat_succ_eq_cumulativeoptimisticexploratorypolicy the regret term indexed by `round` is the source's successor behavior at coordinate `round + 1`; the initial behavior at coordinate zero is intentionally excluded. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.vanishingDeltaScheduledAverageExploratoryBehaviorRegretViolationSet","label":"vanishingDeltaScheduledAverageExploratoryBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.vanishingDeltaScheduledAverageExploratoryBehaviorRegretViolationSet","description":"Trajectories whose average exploratory-behavior regret exceeds its charged certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-80effa16b15e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6933,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:609"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def vanishingDeltaScheduledAverageExploratoryBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (n : Nat) (visitFloor : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : Set (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))","missing":[],"search":"vanishingdeltascheduledaverageexploratorybehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.vanishingdeltascheduledaverageexploratorybehaviorregretviolationset trajectories whose average exploratory-behavior regret exceeds its charged certificate. definition compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegret","description":"Finite-window high-probability endpoint for successor exploratory behaviors at source coordinates `1` through `n + 1`; the initial-table behavior at coordinate zero is not charged here. The violation set is contained in the measurable count bad event; outside that event, optimism and the charged average behavior-regret certificate hold simultaneously.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeexploratorybehaviorregret/index.html#decl-cad808dc37a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","order":6934,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret.lean:635"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (n : Nat) (visitFloor : Real) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hvi…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_vanishingdeltascheduledaverageexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_vanishingdeltascheduledaverageexploratorybehaviorexpectedregret finite-window high-probability endpoint for successor exploratory behaviors at source coordinates `1` through `n + 1`; the initial-table behavior at coordinate zero is not charged here. the violation set is contained in the measurable count bad event; outside that event, optimism and the charged average behavior-regret certificate hold simultaneously. theorem compiled","shard":"modules/539fd9fbc8cd3b09.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardSumSummary","label":"RewardSumSummary","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardSumSummary","description":"Reward sums indexed by stage, state, and action.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-6ad7a571771b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6935,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev RewardSumSummary (mdp : MDP State Action)","missing":[],"search":"rewardsumsummary banditrlproof.finitehorizonrl.rewardsumsummary reward sums indexed by stage, state, and action. abbreviation compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState","label":"AdaptiveCumulativeEmpiricalModelState","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState","description":"The cumulative empirical-model state needed by UCB-VI: transition counts and sampled reward sums. State-action visit counts are derived from the transition summary, so they cannot drift away from the transition counts.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-9216028e0515","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6936,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev AdaptiveCumulativeEmpiricalModelState (mdp : MDP State Action)","missing":[],"search":"adaptivecumulativeempiricalmodelstate banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstate the cumulative empirical-model state needed by ucb-vi: transition counts and sampled reward sums. state-action visit counts are derived from the transition summary, so they cannot drift away from the transition counts. abbreviation compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AggregateVisitCountSummary","label":"AggregateVisitCountSummary","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AggregateVisitCountSummary","description":"UCBVI-CH state-action visits aggregated across every stage.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-ce713e668ff8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6937,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev AggregateVisitCountSummary","missing":[],"search":"aggregatevisitcountsummary banditrlproof.finitehorizonrl.aggregatevisitcountsummary ucbvi-ch state-action visits aggregated across every stage. abbreviation compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateVisitCount","label":"aggregateVisitCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateVisitCount","description":"Sum a stage-indexed transition summary into the paper count `N_k(x,a)`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-b8bba90a6c96","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6938,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def aggregateVisitCount {mdp : MDP State Action} (summary : TransitionCountSummary mdp) : AggregateVisitCountSummary (State := State) (Action := Action)","missing":[],"search":"aggregatevisitcount banditrlproof.finitehorizonrl.transitioncountsummary.aggregatevisitcount sum a stage-indexed transition summary into the paper count `n_k(x,a)`. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.measurable_aggregateVisitCount","label":"measurable_aggregateVisitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.measurable_aggregateVisitCount","description":"The complete cross-stage aggregate count view is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-25e74d253bef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6939,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateVisitCount {mdp : MDP State Action} : Measurable (aggregateVisitCount : TransitionCountSummary mdp -> AggregateVisitCountSummary (State := State) (Action := Action))","missing":[],"search":"measurable_aggregatevisitcount banditrlproof.finitehorizonrl.transitioncountsummary.measurable_aggregatevisitcount the complete cross-stage aggregate count view is measurable. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeRewardSumSummary","label":"cumulativeRewardSumSummary","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeRewardSumSummary","description":"Sum every reward coordinate in a nonempty finite generated-batch prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-acf1d74e83a0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6940,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeRewardSumSummary {mdp : MDP State Action} {episodes n : Nat} (history : EpisodeBatchPrefix mdp episodes n) : RewardSumSummary mdp","missing":[],"search":"cumulativerewardsumsummary banditrlproof.finitehorizonrl.episodebatchprefix.cumulativerewardsumsummary sum every reward coordinate in a nonempty finite generated-batch prefix. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_rewardSum_coordinate","label":"measurable_rewardSum_coordinate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_rewardSum_coordinate","description":"A fixed batch reward sum is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-fd43f9a94eb9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6941,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_rewardSum_coordinate {mdp : MDP State Action} {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable fun batch : EpisodeBatch mdp episodes => batch.rewardSum stage state action","missing":[],"search":"measurable_rewardsum_coordinate banditrlproof.finitehorizonrl.episodebatchprefix.measurable_rewardsum_coordinate a fixed batch reward sum is measurable. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeRewardSumSummary","label":"measurable_cumulativeRewardSumSummary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeRewardSumSummary","description":"The complete cumulative reward-sum table is history measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-dda25aaeeda2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6942,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:121"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeRewardSumSummary {mdp : MDP State Action} {episodes n : Nat} : Measurable (cumulativeRewardSumSummary : EpisodeBatchPrefix mdp episodes n -> RewardSumSummary mdp)","missing":[],"search":"measurable_cumulativerewardsumsummary banditrlproof.finitehorizonrl.episodebatchprefix.measurable_cumulativerewardsumsummary the complete cumulative reward-sum table is history measurable. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeEmpiricalModelState","label":"cumulativeEmpiricalModelState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeEmpiricalModelState","description":"Cumulative transition counts and reward sums from exactly the same prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-794c0be270f7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6943,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeEmpiricalModelState {mdp : MDP State Action} {episodes n : Nat} (history : EpisodeBatchPrefix mdp episodes n) : AdaptiveCumulativeEmpiricalModelState mdp","missing":[],"search":"cumulativeempiricalmodelstate banditrlproof.finitehorizonrl.episodebatchprefix.cumulativeempiricalmodelstate cumulative transition counts and reward sums from exactly the same prefix. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeEmpiricalModelState","label":"measurable_cumulativeEmpiricalModelState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeEmpiricalModelState","description":"The paired cumulative empirical-model state is history measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-c34bbf43a8a0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6944,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeEmpiricalModelState {mdp : MDP State Action} {episodes n : Nat} : Measurable (cumulativeEmpiricalModelState : EpisodeBatchPrefix mdp episodes n -> AdaptiveCumulativeEmpiricalModelState mdp)","missing":[],"search":"measurable_cumulativeempiricalmodelstate banditrlproof.finitehorizonrl.episodebatchprefix.measurable_cumulativeempiricalmodelstate the paired cumulative empirical-model state is history measurable. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt","label":"adaptiveCumulativeEmpiricalModelStateAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt","description":"Cumulative empirical-model state through trajectory coordinate `round`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-f48c00538f52","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6945,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def adaptiveCumulativeEmpiricalModelStateAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) : AdaptiveCumulativeEmpiricalModelState mdp","missing":[],"search":"adaptivecumulativeempiricalmodelstateat banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstateat cumulative empirical-model state through trajectory coordinate `round`. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeEmpiricalModelStateAt","label":"measurable_adaptiveCumulativeEmpiricalModelStateAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeEmpiricalModelStateAt","description":"The cumulative empirical-model state at a fixed generated round is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-0b957e2de69c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6946,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_adaptiveCumulativeEmpiricalModelStateAt {mdp : MDP State Action} {episodes : Nat} (round : Nat) : Measurable fun trajectory : EpisodeBatchTrajectory mdp episodes => adaptiveCumulativeEmpiricalModelStateAt trajectory round","missing":[],"search":"measurable_adaptivecumulativeempiricalmodelstateat banditrlproof.finitehorizonrl.measurable_adaptivecumulativeempiricalmodelstateat the cumulative empirical-model state at a fixed generated round is measurable. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_transitionCount_succ","label":"adaptiveCumulativeEmpiricalModelStateAt_transitionCount_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_transitionCount_succ","description":"Extending the prefix adds exactly the new transition-count coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-3971e2ca1e17","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6947,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeEmpiricalModelStateAt_transitionCount_succ {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (adaptiveCumulativeEmpiricalModelStateAt trajectory (round + 1)).1 stage state action nextState = (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1 stage state action nextState + (trajectory (round + 1)).transitionCount stage state action nextState","missing":[],"search":"adaptivecumulativeempiricalmodelstateat_transitioncount_succ banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstateat_transitioncount_succ extending the prefix adds exactly the new transition-count coordinate. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_rewardSum_succ","label":"adaptiveCumulativeEmpiricalModelStateAt_rewardSum_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_rewardSum_succ","description":"Extending the prefix adds exactly the new reward-sum coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-73d54e5ee848","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6948,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:196"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeEmpiricalModelStateAt_rewardSum_succ {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : (adaptiveCumulativeEmpiricalModelStateAt trajectory (round + 1)).2 stage state action = (adaptiveCumulativeEmpiricalModelStateAt trajectory round).2 stage state action + (trajectory (round + 1)).rewardSum stage state action","missing":[],"search":"adaptivecumulativeempiricalmodelstateat_rewardsum_succ banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstateat_rewardsum_succ extending the prefix adds exactly the new reward-sum coordinate. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_visitCount","label":"adaptiveCumulativeEmpiricalModelStateAt_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_visitCount","description":"Visit counts are definitionally derived from the same cumulative state.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-fd9c19ce4957","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6949,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:221"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeEmpiricalModelStateAt_visitCount {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1.visitCount stage state action = ∑ i : Fin (round + 1), (trajectory i).visitCount stage state action","missing":[],"search":"adaptivecumulativeempiricalmodelstateat_visitcount banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstateat_visitcount visit counts are definitionally derived from the same cumulative state. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt","label":"adaptiveCumulativeAggregateVisitCountAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt","description":"The paper count `N_k(x,a)` read from exactly one generated prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-791eb7c5068d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6950,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:235"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def adaptiveCumulativeAggregateVisitCountAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) : AggregateVisitCountSummary (State := State) (Action := Action)","missing":[],"search":"adaptivecumulativeaggregatevisitcountat banditrlproof.finitehorizonrl.adaptivecumulativeaggregatevisitcountat the paper count `n_k(x,a)` read from exactly one generated prefix. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeAggregateVisitCountAt","label":"measurable_adaptiveCumulativeAggregateVisitCountAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeAggregateVisitCountAt","description":"Every fixed-round aggregate state-action count table is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-f81807091aa7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6951,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:243"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_adaptiveCumulativeAggregateVisitCountAt {mdp : MDP State Action} {episodes : Nat} (round : Nat) : Measurable fun trajectory : EpisodeBatchTrajectory mdp episodes => adaptiveCumulativeAggregateVisitCountAt trajectory round","missing":[],"search":"measurable_adaptivecumulativeaggregatevisitcountat banditrlproof.finitehorizonrl.measurable_adaptivecumulativeaggregatevisitcountat every fixed-round aggregate state-action count table is measurable. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt_eq_sum","label":"adaptiveCumulativeAggregateVisitCountAt_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt_eq_sum","description":"The aggregate count is the literal sum over prior episodes and stages.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-6a7c764684db","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6952,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:253"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeAggregateVisitCountAt_eq_sum {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (state : State) (action : Action) : adaptiveCumulativeAggregateVisitCountAt trajectory round state action = ∑ i : Fin (round + 1), ∑ stage : Fin mdp.horizon, (trajectory i).visitCount stage state action","missing":[],"search":"adaptivecumulativeaggregatevisitcountat_eq_sum banditrlproof.finitehorizonrl.adaptivecumulativeaggregatevisitcountat_eq_sum the aggregate count is the literal sum over prior episodes and stages. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt_succ","label":"adaptiveCumulativeAggregateVisitCountAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt_succ","description":"Extending the generated prefix adds exactly one episode's stage visits.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-32fdb5617471","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6953,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeAggregateVisitCountAt_succ {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (state : State) (action : Action) : adaptiveCumulativeAggregateVisitCountAt trajectory (round + 1) state action = adaptiveCumulativeAggregateVisitCountAt trajectory round state action + ∑ stage : Fin mdp.horizon, (trajectory (round + 1)).visitCount stage state action","missing":[],"search":"adaptivecumulativeaggregatevisitcountat_succ banditrlproof.finitehorizonrl.adaptivecumulativeaggregatevisitcountat_succ extending the generated prefix adds exactly one episode's stage visits. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState.empiricalReward","label":"empiricalReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState.empiricalReward","description":"Empirical reward with the explicit zero-count convention.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-37044b362dd4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6954,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:283"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def AdaptiveCumulativeEmpiricalModelState.empiricalReward {mdp : MDP State Action} (model : AdaptiveCumulativeEmpiricalModelState mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) : Real","missing":[],"search":"empiricalreward banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstate.empiricalreward empirical reward with the explicit zero-count convention. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState.empiricalReward_of_visitCount_eq_zero","label":"empiricalReward_of_visitCount_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState.empiricalReward_of_visitCount_eq_zero","description":"theorem AdaptiveCumulativeEmpiricalModelState.empiricalReward_of_visitCount_eq_zero {mdp : MDP State Action} (model : AdaptiveCumulativeEmpiricalModelState mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) (hzero : model.1.visitCount stage state action = 0) : model.empiricalReward stage state action = 0","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-db2f2411ba55","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6955,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:294"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem AdaptiveCumulativeEmpiricalModelState.empiricalReward_of_visitCount_eq_zero {mdp : MDP State Action} (model : AdaptiveCumulativeEmpiricalModelState mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) (hzero : model.1.visitCount stage state action = 0) : model.empiricalReward stage state action = 0","missing":[],"search":"empiricalreward_of_visitcount_eq_zero banditrlproof.finitehorizonrl.adaptivecumulativeempiricalmodelstate.empiricalreward_of_visitcount_eq_zero theorem adaptivecumulativeempiricalmodelstate.empiricalreward_of_visitcount_eq_zero {mdp : mdp state action} (model : adaptivecumulativeempiricalmodelstate mdp) (stage : fin mdp.horizon) (state : state) (action : action) (hzero : model.1.visitcount stage state action = 0) : model.empiricalreward stage state action = 0 theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalSteps","label":"totalSteps","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalSteps","description":"Total number of environment time steps after `episodes` episodes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-9dc652641be6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6956,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:305"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def totalSteps (mdp : MDP State Action) (episodes : Nat) : Nat","missing":[],"search":"totalsteps banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.totalsteps total number of environment time steps after `episodes` episodes. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.confidenceNumerator","label":"confidenceNumerator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.confidenceNumerator","description":"Integer numerator used in the UCBVI-CH logarithmic factor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-ee9811df10f5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6957,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def confidenceNumerator (mdp : MDP State Action) (episodes : Nat) : Nat","missing":[],"search":"confidencenumerator banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.confidencenumerator integer numerator used in the ucbvi-ch logarithmic factor. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor","label":"logFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor","description":"Safe total version of the UCBVI-CH logarithmic factor. The `max 1` only specifies invalid/degenerate inputs and disappears under the positive task contract.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-9e9af9a5f255","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6958,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:319"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def logFactor (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Real","missing":[],"search":"logfactor banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.logfactor safe total version of the ucbvi-ch logarithmic factor. the `max 1` only specifies invalid/degenerate inputs and disappears under the positive task contract. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.one_le_confidenceNumerator_div","label":"one_le_confidenceNumerator_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.one_le_confidenceNumerator_div","description":"theorem one_le_confidenceNumerator_div (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 1 <= ((confidenceNumerator (State := State) (Action := Action) mdp episodes : Nat) : Real) / delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-f23fab3f02ad","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6959,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:327"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_le_confidenceNumerator_div (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 1 <= ((confidenceNumerator (State := State) (Action := Action) mdp episodes : Nat) : Real) / delta","missing":[],"search":"one_le_confidencenumerator_div banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.one_le_confidencenumerator_div theorem one_le_confidencenumerator_div (mdp : mdp state action) (episodes : nat) (delta : real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 1 <= ((confidencenumerator (state := state) (action := action) mdp episodes : nat) : real) / delta theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor_eq_paper","label":"logFactor_eq_paper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor_eq_paper","description":"On the task domain, the safe total logarithm is exactly the paper factor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-d7383c28c334","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6960,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:349"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem logFactor_eq_paper (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : logFactor (State := State) (Action := Action) mdp episodes delta = Real.log (((confidenceNumerator (State := State) (Action := Action) mdp episodes : Nat) : Real) / delta)","missing":[],"search":"logfactor_eq_paper banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.logfactor_eq_paper on the task domain, the safe total logarithm is exactly the paper factor. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor_nonneg","label":"logFactor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor_nonneg","description":"theorem logFactor_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= logFactor (State := State) (Action := Action) mdp episodes delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-209bc248764a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6961,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:365"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem logFactor_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= logFactor (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"logfactor_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.logfactor_nonneg theorem logfactor_nonneg (mdp : mdp state action) (episodes : nat) (delta : real) : 0 <= logfactor (state := state) (action := action) mdp episodes delta theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.scale","label":"scale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.scale","description":"Paper-shaped Hoeffding scale `7 H L`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-da1b07352775","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6962,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:372"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def scale (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Real","missing":[],"search":"scale banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.scale paper-shaped hoeffding scale `7 h l`. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.scale_nonneg","label":"scale_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.scale_nonneg","description":"theorem scale_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= scale (State := State) (Action := Action) mdp episodes delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-82773724e1d8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6963,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:380"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem scale_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= scale (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"scale_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.scale_nonneg theorem scale_nonneg (mdp : mdp state action) (episodes : nat) (delta : real) : 0 <= scale (state := state) (action := action) mdp episodes delta theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius","label":"countRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius","description":"Clipped Hoeffding radius: `H` at zero visits and `min H (7 H L / sqrt N)` afterwards.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-e5436bb12709","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6964,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def countRadius (mdp : MDP State Action) (episodes : Nat) (delta : Real) : TransitionCountRadius","missing":[],"search":"countradius banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.countradius clipped hoeffding radius: `h` at zero visits and `min h (7 h l / sqrt n)` afterwards. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius_zero","label":"countRadius_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius_zero","description":"theorem countRadius_zero (mdp : MDP State Action) (episodes : Nat) (delta : Real) : (countRadius (State := State) (Action := Action) mdp episodes delta).radius 0 = (mdp.horizon : Real)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-0c96fde311b6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6965,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:404"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countRadius_zero (mdp : MDP State Action) (episodes : Nat) (delta : Real) : (countRadius (State := State) (Action := Action) mdp episodes delta).radius 0 = (mdp.horizon : Real)","missing":[],"search":"countradius_zero banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.countradius_zero theorem countradius_zero (mdp : mdp state action) (episodes : nat) (delta : real) : (countradius (state := state) (action := action) mdp episodes delta).radius 0 = (mdp.horizon : real) theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius_of_pos","label":"countRadius_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius_of_pos","description":"theorem countRadius_of_pos (mdp : MDP State Action) (episodes : Nat) (delta : Real) {count : Nat} (hcount : 0 < count) : (countRadius (State := State) (Action := Action) mdp episodes delta).radius count = min (mdp.horizon : Real) (scale (State := State) (Action := Action) mdp episodes delta / Real.sqrt count)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-37d4c3ff5b62","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6966,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:413"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countRadius_of_pos (mdp : MDP State Action) (episodes : Nat) (delta : Real) {count : Nat} (hcount : 0 < count) : (countRadius (State := State) (Action := Action) mdp episodes delta).radius count = min (mdp.horizon : Real) (scale (State := State) (Action := Action) mdp episodes delta / Real.sqrt count)","missing":[],"search":"countradius_of_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.countradius_of_pos theorem countradius_of_pos (mdp : mdp state action) (episodes : nat) (delta : real) {count : nat} (hcount : 0 < count) : (countradius (state := state) (action := action) mdp episodes delta).radius count = min (mdp.horizon : real) (scale (state := state) (action := action) mdp episodes delta / real.sqrt count) theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source","label":"source","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source","description":"One-episode-at-a-time cumulative optimistic source. Coordinate zero follows `initialTable`; coordinate `n+1` is generated by the pure optimistic policy computed from all coordinates through `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-867ef5967fc6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6967,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:430"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def source (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (episodes : Nat) (delta : Real) : AdaptiveEpisodeBatchSource mdp initialState 1 where","missing":[],"search":"source banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.source one-episode-at-a-time cumulative optimistic source. coordinate zero follows `initialtable`; coordinate `n+1` is generated by the pure optimistic policy computed from all coordinates through `n`. definition compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_successorPolicy_eq_optimisticPolicy","label":"source_successorPolicy_eq_optimisticPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_successorPolicy_eq_optimisticPolicy","description":"The next policy is exactly the cumulative empirical optimistic plan.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-60c70f9b4c25","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6968,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:461"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_successorPolicy_eq_optimisticPolicy (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (history : EpisodeBatchPrefix mdp 1 n) : (source mdp initialState initialTable defaultState episodes delta).successorPolicy n history = (history.cumulativeTransitionCountSummary.countRadiusOptimisticPlan mdp defaultState (countRadius (State := State) (Action := Action) mdp episodes delta) ).optimisticPolicy","missing":[],"search":"source_successorpolicy_eq_optimisticpolicy banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.source_successorpolicy_eq_optimisticpolicy the next policy is exactly the cumulative empirical optimistic plan. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_policyAt_succ_eq_optimisticPolicy","label":"source_policyAt_succ_eq_optimisticPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_policyAt_succ_eq_optimisticPolicy","description":"On a complete generated trajectory, the policy at successor episode `n+1` uses exactly the cumulative counts through `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-e07fe88bb6c4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6969,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:480"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_policyAt_succ_eq_optimisticPolicy (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : (source mdp initialState initialTable defaultState episodes delta).policyAt trajectory (n + 1) = (adaptiveCumulativeEmpiricalOptimisticPlanAt trajectory defaultState (countRadius (State := State) (Action := Action) mdp episodes delta) n ).optimisticPolicy","missing":[],"search":"source_policyat_succ_eq_optimisticpolicy banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.source_policyat_succ_eq_optimisticpolicy on a complete generated trajectory, the policy at successor episode `n+1` uses exactly the cumulative counts through `n`. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_policyAt_succ_eq_modelState_optimisticPolicy","label":"source_policyAt_succ_eq_modelState_optimisticPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_policyAt_succ_eq_modelState_optimisticPolicy","description":"The policy/model alignment names the paired cumulative state explicitly.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativehoeffdingucbvi/index.html#decl-7633d3c49bdd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","order":6970,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI.lean:495"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_policyAt_succ_eq_modelState_optimisticPolicy (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : (source mdp initialState initialTable defaultState episodes delta).policyAt trajectory (n + 1) = ((adaptiveCumulativeEmpiricalModelStateAt trajectory n).1 |>.countRadiusOptimisticPlan mdp defaultState (countRadius (State := State) (Action := Action) mdp episodes delta) |>.optimisticPolicy)","missing":[],"search":"source_policyat_succ_eq_modelstate_optimisticpolicy banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.source_policyat_succ_eq_modelstate_optimisticpolicy the policy/model alignment names the paired cumulative state explicitly. theorem compiled","shard":"modules/56041dd13028c84c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold","label":"normalizedCumulativeInverseSqrtScheduledEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold","description":"Real-valued batch-size target that clears both normalized calibration and the one-batch logarithmic visit-mass requirement.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-f8d8b4c87ac8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6971,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtScheduledEpisodeThreshold (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"normalizedcumulativeinversesqrtscheduledepisodethreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledepisodethreshold real-valued batch-size target that clears both normalized calibration and the one-batch logarithmic visit-mass requirement. definition compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes","label":"normalizedCumulativeInverseSqrtScheduledEpisodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes","description":"Explicit positive natural batch size associated with the scheduled target.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-bdff55ab2be5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6972,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtScheduledEpisodes (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : Nat","missing":[],"search":"normalizedcumulativeinversesqrtscheduledepisodes banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledepisodes explicit positive natural batch size associated with the scheduled target. definition compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageEnvelope","label":"normalizedCumulativeInverseSqrtScheduledAverageEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageEnvelope","description":"The inverse-square-root envelope for the scheduled average bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-d5bd11108a17","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6973,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtScheduledAverageEnvelope (mdp : MDP State Action) (rounds : Nat) (visitFloor : Real) : Real","missing":[],"search":"normalizedcumulativeinversesqrtscheduledaverageenvelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledaverageenvelope the inverse-square-root envelope for the scheduled average bound. definition compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_pos","label":"cumulativeInverseSqrtLogFactor_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_pos","description":"The valid global confidence budget makes the shared logarithm strict.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-91517b08f3dc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6974,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeInverseSqrtLogFactor_pos (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < cumulativeInverseSqrtLogFactor mdp rounds delta","missing":[],"search":"cumulativeinversesqrtlogfactor_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtlogfactor_pos the valid global confidence budget makes the shared logarithm strict. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_nonneg","label":"normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_nonneg","description":"The scheduled real-valued target is nonnegative under valid parameters.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-d47e3dc2289e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6975,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_nonneg (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) : 0 <= normalizedCumulativeInverseSqrtScheduledEpisodeThreshold mdp rounds delta visitFloor","missing":[],"search":"normalizedcumulativeinversesqrtscheduledepisodethreshold_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledepisodethreshold_nonneg the scheduled real-valued target is nonnegative under valid parameters. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes_pos","label":"normalizedCumulativeInverseSqrtScheduledEpisodes_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes_pos","description":"Every scheduled batch size is positive.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-8bb81c6c4f19","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6976,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledEpisodes_pos (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : 0 < normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor","missing":[],"search":"normalizedcumulativeinversesqrtscheduledepisodes_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledepisodes_pos every scheduled batch size is positive. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_lt_episodes","label":"normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_lt_episodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_lt_episodes","description":"The real-valued target is strictly below the scheduled natural batch size.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-168b792a5963","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6977,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_lt_episodes (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : normalizedCumulativeInverseSqrtScheduledEpisodeThreshold mdp rounds delta visitFloor < (normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor : Real)","missing":[],"search":"normalizedcumulativeinversesqrtscheduledepisodethreshold_lt_episodes banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledepisodethreshold_lt_episodes the real-valued target is strictly below the scheduled natural batch size. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtEpisodeThreshold_lt_scheduledEpisodes","label":"normalizedCumulativeInverseSqrtEpisodeThreshold_lt_scheduledEpisodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtEpisodeThreshold_lt_scheduledEpisodes","description":"The schedule strictly clears the parent normalized calibration threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-847395167e7c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6978,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:130"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtEpisodeThreshold_lt_scheduledEpisodes (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : normalizedCumulativeInverseSqrtEpisodeThreshold mdp rounds delta visitFloor < (normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor : Real)","missing":[],"search":"normalizedcumulativeinversesqrtepisodethreshold_lt_scheduledepisodes banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtepisodethreshold_lt_scheduledepisodes the schedule strictly clears the parent normalized calibration threshold. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_lt_scheduledEpisodeVisitMass","label":"cumulativeInverseSqrtLogFactor_lt_scheduledEpisodeVisitMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_lt_scheduledEpisodeVisitMass","description":"One scheduled batch has strictly more visit mass than the log factor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-5c9439142518","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6979,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:146"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeInverseSqrtLogFactor_lt_scheduledEpisodeVisitMass (mdp : MDP State Action) (rounds : Nat) (delta : Real) {visitFloor : Real} (hvisitFloor : 0 < visitFloor) : cumulativeInverseSqrtLogFactor mdp rounds delta < (normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor : Real) * visitFloor / 2","missing":[],"search":"cumulativeinversesqrtlogfactor_lt_scheduledepisodevisitmass banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtlogfactor_lt_scheduledepisodevisitmass one scheduled batch has strictly more visit mass than the log factor. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_nonneg","label":"normalizedCumulativeInverseSqrtScheduledAverageBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_nonneg","description":"The normalized scheduled average bound is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-c218453248f7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6980,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:163"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledAverageBound_nonneg (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) : 0 <= normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound mdp (normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor) rounds delta visitFloor","missing":[],"search":"normalizedcumulativeinversesqrtscheduledaveragebound_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledaveragebound_nonneg the normalized scheduled average bound is nonnegative. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_le_envelope","label":"normalizedCumulativeInverseSqrtScheduledAverageBound_le_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_le_envelope","description":"The scheduled average bound is controlled by a pure inverse-square-root rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-0221ffe1ea28","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6981,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:182"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledAverageBound_le_envelope (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) : normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound mdp (normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor) rounds delta visitFloor <= normalizedCumulativeInverseSqrtScheduledAverageEnvelope mdp rounds visitFloor","missing":[],"search":"normalizedcumulativeinversesqrtscheduledaveragebound_le_envelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledaveragebound_le_envelope the scheduled average bound is controlled by a pure inverse-square-root rate. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageEnvelope_tendsto_zero","label":"normalizedCumulativeInverseSqrtScheduledAverageEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageEnvelope_tendsto_zero","description":"The deterministic inverse-square-root envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-035f55dd2b85","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6982,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:297"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledAverageEnvelope_tendsto_zero (mdp : MDP State Action) (visitFloor : Real) : Tendsto (fun n : Nat => normalizedCumulativeInverseSqrtScheduledAverageEnvelope mdp (n + 1) visitFloor) atTop (nhds 0)","missing":[],"search":"normalizedcumulativeinversesqrtscheduledaverageenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledaverageenvelope_tendsto_zero the deterministic inverse-square-root envelope tends to zero. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_tendsto_zero","label":"normalizedCumulativeInverseSqrtScheduledAverageBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_tendsto_zero","description":"The scheduled scalar average recommendation-regret bound tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-3d8e5eda2047","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6983,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:315"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScheduledAverageBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (delta visitFloor : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) : Tendsto (fun n : Nat => normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound mdp (normalizedCumulativeInverseSqrtScheduledEpisodes mdp (n + 1) delta visitFloor) (n + 1) delta visitFloor) atTop (nhds 0)","missing":[],"search":"normalizedcumulativeinversesqrtscheduledaveragebound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscheduledaveragebound_tendsto_zero the scheduled scalar average recommendation-regret bound tends to zero. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_scheduledAverageRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_scheduledAverageRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_scheduledAverageRecommendedExpectedRegret","description":"At every positive finite window, the explicit schedule discharges calibration and preserves the parent's measurable event, delta tail, optimism, and average recommended-policy expected-regret conclusion.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaverageconsistency/index.html#decl-068cfb39c431","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","order":6984,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency.lean:346"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_scheduledAverageRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds : Nat) (delta visitFloor : Real) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes mdp rounds delta visitFloor))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hrewardBound : forall state action, |mdp.reward state ac…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_scheduledaveragerecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_scheduledaveragerecommendedexpectedregret at every positive finite window, the explicit schedule discharges calibration and preserves the parent's measurable event, delta tail, optimism, and average recommended-policy expected-regret conclusion. theorem compiled","shard":"modules/57c1674abe5c367c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeExploratoryEpisodeCount","label":"cumulativeExploratoryEpisodeCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeExploratoryEpisodeCount","description":"Total exploratory episodes in `rounds` batches of size `episodes`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaveragerate/index.html#decl-b90c04a8f919","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","order":6985,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeExploratoryEpisodeCount (episodes rounds : Nat) : Nat","missing":[],"search":"cumulativeexploratoryepisodecount banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeexploratoryepisodecount total exploratory episodes in `rounds` batches of size `episodes`. definition compiled","shard":"modules/68d37f57cb8b326f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound","label":"normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound","description":"The normalized cumulative recommendation-regret bound per recommendation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaveragerate/index.html#decl-929042a8bc9f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","order":6986,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound (mdp : MDP State Action) (episodes rounds : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"normalizedcumulativeinversesqrtaveragerecommendedexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtaveragerecommendedexpectedregretbound the normalized cumulative recommendation-regret bound per recommendation. definition compiled","shard":"modules/68d37f57cb8b326f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound_eq_totalEpisodes","label":"normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound_eq_totalEpisodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound_eq_totalEpisodes","description":"The normalized average bound exposes the square-root rate in the total number of exploratory episodes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaveragerate/index.html#decl-ab9fae3c1c83","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","order":6987,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound_eq_totalEpisodes (mdp : MDP State Action) {episodes rounds : Nat} (hepisodes : 0 < episodes) (hrounds : 0 < rounds) (delta : Real) {visitFloor : Real} (hvisitFloor : 0 < visitFloor) : normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound mdp episodes rounds delta visitFloor = 2 * (mdp.horizon : Real) * min 1 (8 * (Fintype.card State : Real) * (mdp.horizon : Real) * Real.sqrt (cumulativeInverseSqrtLogFactor mdp rounds delta) / Real.sqrt visitFloor / Real.sqrt ((cumulativeExploratoryEpisodeCount episodes rounds : Nat) * visitFloor / 2))","missing":[],"search":"normalizedcumulativeinversesqrtaveragerecommendedexpectedregretbound_eq_totalepisodes banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtaveragerecommendedexpectedregretbound_eq_totalepisodes the normalized average bound exposes the square-root rate in the total number of exploratory episodes. theorem compiled","shard":"modules/68d37f57cb8b326f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageRecommendedExpectedRegret","label":"adaptiveCumulativeEmpiricalOptimisticAverageRecommendedExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageRecommendedExpectedRegret","description":"Average expected regret of the cumulative empirical recommendations.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaveragerate/index.html#decl-a6bc91bafbee","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","order":6988,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate.lean:117"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveCumulativeEmpiricalOptimisticAverageRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (countRadius : TransitionCountRadius) (rounds : Nat) : Real","missing":[],"search":"adaptivecumulativeempiricaloptimisticaveragerecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticaveragerecommendedexpectedregret average expected regret of the cumulative empirical recommendations. definition compiled","shard":"modules/68d37f57cb8b326f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedAverageRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedAverageRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedAverageRecommendedExpectedRegret","description":"Normalized average-recommendation endpoint expressed through all exploratory episodes, with the exact parent event and optimism conclusion unchanged.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtaveragerate/index.html#decl-cec6a6d4459b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","order":6989,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate.lean:133"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedAverageRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (ht…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_normalizedaveragerecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_normalizedaveragerecommendedexpectedregret normalized average-recommendation endpoint expressed through all exploratory episodes, with the exact parent event and optimism conclusion unchanged. theorem compiled","shard":"modules/68d37f57cb8b326f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt","label":"inverseSqrt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt","description":"A concrete count radius: the supplied budget at zero visits and inverse-square root decay after the first visit.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-330c8f7ee6aa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6990,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrt (budget : Real) (hbudget : 0 <= budget) : TransitionCountRadius where","missing":[],"search":"inversesqrt banditrlproof.finitehorizonrl.transitioncountradius.inversesqrt a concrete count radius: the supplied budget at zero visits and inverse-square root decay after the first visit. definition compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt_radius_zero","label":"inverseSqrt_radius_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt_radius_zero","description":"The inverse-square-root radius exposes its supplied zero-count budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-c7110347b98d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6991,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"@[simp] theorem inverseSqrt_radius_zero (budget : Real) (hbudget : 0 <= budget) : (inverseSqrt budget hbudget).radius 0 = budget","missing":[],"search":"inversesqrt_radius_zero banditrlproof.finitehorizonrl.transitioncountradius.inversesqrt_radius_zero the inverse-square-root radius exposes its supplied zero-count budget. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt_radius_of_pos","label":"inverseSqrt_radius_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt_radius_of_pos","description":"Positive counts use the genuine inverse-square-root branch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-44d4d5dd1cd1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6992,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrt_radius_of_pos (budget : Real) (hbudget : 0 <= budget) {count : Nat} (hcount : 0 < count) : (inverseSqrt budget hbudget).radius count = budget / Real.sqrt count","missing":[],"search":"inversesqrt_radius_of_pos banditrlproof.finitehorizonrl.transitioncountradius.inversesqrt_radius_of_pos positive counts use the genuine inverse-square-root branch. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt","label":"cappedInverseSqrt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt","description":"A usable capped inverse-square-root radius. `budget` controls the zero-count value envelope, while `scale` controls the statistical decay after enough visits; the cap preserves antitonicity across the first visit.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-7b93e30b3493","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6993,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cappedInverseSqrt (budget scale : Real) (hbudget : 0 <= budget) (hscale : 0 <= scale) : TransitionCountRadius where","missing":[],"search":"cappedinversesqrt banditrlproof.finitehorizonrl.transitioncountradius.cappedinversesqrt a usable capped inverse-square-root radius. `budget` controls the zero-count value envelope, while `scale` controls the statistical decay after enough visits; the cap preserves antitonicity across the first visit. definition compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt_radius_zero","label":"cappedInverseSqrt_radius_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt_radius_zero","description":"@[simp] theorem cappedInverseSqrt_radius_zero (budget scale : Real) (hbudget : 0 <= budget) (hscale : 0 <= scale) : (cappedInverseSqrt budget scale hbudget hscale).radius 0 = budget","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-a68fdd0df3ed","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6994,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:114"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"@[simp] theorem cappedInverseSqrt_radius_zero (budget scale : Real) (hbudget : 0 <= budget) (hscale : 0 <= scale) : (cappedInverseSqrt budget scale hbudget hscale).radius 0 = budget","missing":[],"search":"cappedinversesqrt_radius_zero banditrlproof.finitehorizonrl.transitioncountradius.cappedinversesqrt_radius_zero @[simp] theorem cappedinversesqrt_radius_zero (budget scale : real) (hbudget : 0 <= budget) (hscale : 0 <= scale) : (cappedinversesqrt budget scale hbudget hscale).radius 0 = budget theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt_radius_of_pos","label":"cappedInverseSqrt_radius_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt_radius_of_pos","description":"theorem cappedInverseSqrt_radius_of_pos (budget scale : Real) (hbudget : 0 <= budget) (hscale : 0 <= scale) {count : Nat} (hcount : 0 < count) : (cappedInverseSqrt budget scale hbudget hscale).radius count = min budget (scale / Real.sqrt count)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-217618c2daff","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6995,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cappedInverseSqrt_radius_of_pos (budget scale : Real) (hbudget : 0 <= budget) (hscale : 0 <= scale) {count : Nat} (hcount : 0 < count) : (cappedInverseSqrt budget scale hbudget hscale).radius count = min budget (scale / Real.sqrt count)","missing":[],"search":"cappedinversesqrt_radius_of_pos banditrlproof.finitehorizonrl.transitioncountradius.cappedinversesqrt_radius_of_pos theorem cappedinversesqrt_radius_of_pos (budget scale : real) (hbudget : 0 <= budget) (hscale : 0 <= scale) {count : nat} (hcount : 0 < count) : (cappedinversesqrt budget scale hbudget hscale).radius count = min budget (scale / real.sqrt count) theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativePathVisitExpectedFloor","label":"cumulativePathVisitExpectedFloor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativePathVisitExpectedFloor","description":"Predictable cumulative visit floor supplied by one path-support batch floor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-41a81dfd9801","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6996,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:131"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativePathVisitExpectedFloor (episodes prefixRounds : Nat) (visitFloor : Real) : Real","missing":[],"search":"cumulativepathvisitexpectedfloor banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativepathvisitexpectedfloor predictable cumulative visit floor supplied by one path-support batch floor. definition compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativePathVisitLowerMargin","label":"cumulativePathVisitLowerMargin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativePathVisitLowerMargin","description":"Predictable visit floor after subtracting the cumulative confidence radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-7813eb757af3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6997,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativePathVisitLowerMargin (mdp : MDP State Action) (episodes rounds : Nat) (delta visitFloor : Real) (round : Fin rounds) : Real","missing":[],"search":"cumulativepathvisitlowermargin banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativepathvisitlowermargin predictable visit floor after subtracting the cumulative confidence radius. definition compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtRadiusEnvelope","label":"cumulativeInverseSqrtRadiusEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtRadiusEnvelope","description":"Deterministic selected-radius envelope obtained from the lower visit margin.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-f7e7194eb799","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6998,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeInverseSqrtRadiusEnvelope (mdp : MDP State Action) (episodes rounds : Nat) (delta visitFloor budget scale : Real) (round : Fin rounds) : Real","missing":[],"search":"cumulativeinversesqrtradiusenvelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtradiusenvelope deterministic selected-radius envelope obtained from the lower visit margin. definition compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.CumulativeInverseSqrtPathCalibration","label":"CumulativeInverseSqrtPathCalibration","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.CumulativeInverseSqrtPathCalibration","description":"Scalar regularity needed to fit all transition-coordinate errors under the inverse-square-root planner radius. It is deterministic and roundwise; path support and the martingale event supply the corresponding realized counts.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-7dc387660d76","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":6999,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:155"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure CumulativeInverseSqrtPathCalibration (mdp : MDP State Action) (episodes rounds : Nat) (delta visitFloor rewardBound budget scale : Real) : Prop where","missing":[],"search":"cumulativeinversesqrtpathcalibration banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtpathcalibration scalar regularity needed to fit all transition-coordinate errors under the inverse-square-root planner radius. it is deterministic and roundwise; path support and the martingale event supply the corresponding realized counts. structure compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_coordinateMeanAt_visit_ge_pathFloor","label":"exploratorySource_coordinateMeanAt_visit_ge_pathFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_coordinateMeanAt_visit_ge_pathFloor","description":"Every adaptive exploratory batch retains the common path-support visit floor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-24a599936cb2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":7000,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:187"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_coordinateMeanAt_visit_ge_pathFloor (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : (episodes : Real) * visitFloor <= (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).coordinateMeanAt (.visit stage state action) round trajectory","missing":[],"search":"exploratorysource_coordinatemeanat_visit_ge_pathfloor banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_coordinatemeanat_visit_ge_pathfloor every adaptive exploratory batch retains the common path-support visit floor. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_cumulativeCoordinateMean_visit_ge_pathFloor","label":"exploratorySource_cumulativeCoordinateMean_visit_ge_pathFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_cumulativeCoordinateMean_visit_ge_pathFloor","description":"The common batch floor sums to a predictable cumulative visit floor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-a8f8141a3871","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":7001,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:222"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_cumulativeCoordinateMean_visit_ge_pathFloor (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (prefixRounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : AdaptiveEpisodeBatchSource.cumulativePathVisitExpectedFloor episodes prefixRounds visitFloor <= (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).cumulativeCoordinateMean (.visit stage s…","missing":[],"search":"exploratorysource_cumulativecoordinatemean_visit_ge_pathfloor banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_cumulativecoordinatemean_visit_ge_pathfloor the common batch floor sums to a predictable cumulative visit floor. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_cumulativePathVisitLowerMargin_lt_visitCount","label":"exploratorySource_cumulativePathVisitLowerMargin_lt_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_cumulativePathVisitLowerMargin_lt_visitCount","description":"Outside the global count event, every cumulative visit count exceeds its lower margin.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-e00c3f853cbe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":7002,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:257"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_cumulativePathVisitLowerMargin_lt_visitCount (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor delta : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) {trajectory : EpisodeBatchTrajectory mdp episodes} (htrajectory : trajectory ∉ (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).adaptiveCumulativeCountBadEvent rounds delta) (round : Fin rounds) (stage : Fin mdp.horizon) (state : State) (action : Action) : AdaptiveEpisodeBatchSource.cumulativePathVisitLo…","missing":[],"search":"exploratorysource_cumulativepathvisitlowermargin_lt_visitcount banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_cumulativepathvisitlowermargin_lt_visitcount outside the global count event, every cumulative visit count exceeds its lower margin. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_adaptiveCumulativeCountMartingaleCover_of_pathSupport_inverseSqrtCalibration","label":"exploratorySource_adaptiveCumulativeCountMartingaleCover_of_pathSupport_inverseSqrtCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_adaptiveCumulativeCountMartingaleCover_of_pathSupport_inverseSqrtCalibration","description":"Path support and the two-scale calibration discharge the full capped inverse-sqrt cover.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-7db8bd13532c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":7003,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_adaptiveCumulativeCountMartingaleCover_of_pathSupport_inverseSqrtCalibration (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor delta : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (rewardBound budget scale : Real) (calibration : AdaptiveEpisodeBatchSource.CumulativeInverseSqrtPathCalibration mdp episodes rounds delta visitFloor rewardBound budget scale) : AdaptiveEpisodeBatchSource.AdaptiveCumulativeCountMartingaleCover (rounds := rounds) (exploratorySource mdp initialState episodes initialTable defaultState (TransitionCountRadius.cappedInverseSqrt budg…","missing":[],"search":"exploratorysource_adaptivecumulativecountmartingalecover_of_pathsupport_inversesqrtcalibration banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_adaptivecumulativecountmartingalecover_of_pathsupport_inversesqrtcalibration path support and the two-scale calibration discharge the full capped inverse-sqrt cover. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.adaptiveCumulativeEmpiricalOptimisticPlanAt_selectedRadiusRemaining_le_inverseSqrtEnvelope","label":"adaptiveCumulativeEmpiricalOptimisticPlanAt_selectedRadiusRemaining_le_inverseSqrtEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.adaptiveCumulativeEmpiricalOptimisticPlanAt_selectedRadiusRemaining_le_inverseSqrtEnvelope","description":"The selected capped inverse-sqrt radius is controlled by the lower-margin envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-bfc6235b4a7f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":7004,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:418"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeEmpiricalOptimisticPlanAt_selectedRadiusRemaining_le_inverseSqrtEnvelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor delta : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (rewardBound budget scale : Real) (calibration : AdaptiveEpisodeBatchSource.CumulativeInverseSqrtPathCalibration mdp episodes rounds delta visitFloor rewardBound budget scale) {trajectory : EpisodeBatchTrajectory mdp episodes} (htrajectory : trajectory ∉ (exploratorySource mdp initialState episodes initialTable defaultState (TransitionCountRadius.cappedInverseSqrt budget scale cal…","missing":[],"search":"adaptivecumulativeempiricaloptimisticplanat_selectedradiusremaining_le_inversesqrtenvelope banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.adaptivecumulativeempiricaloptimisticplanat_selectedradiusremaining_le_inversesqrtenvelope the selected capped inverse-sqrt radius is controlled by the lower-margin envelope. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_explicitRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_explicitRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_explicitRecommendedExpectedRegret","description":"Concrete route endpoint: path support and capped inverse-sqrt calibration produce one measurable cumulative count event, optimism, and a round-indexed finite-sum bound for recommended-policy expected regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtcalibration/index.html#decl-7718f68a1eb9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","order":7005,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration.lean:483"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_explicitRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (rewardBound budget scale : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1)…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_explicitrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_explicitrecommendedexpectedregret concrete route endpoint: path support and capped inverse-sqrt calibration produce one measurable cumulative count event, optimism, and a round-indexed finite-sum bound for recommended-policy expected regret. theorem compiled","shard":"modules/5729df991a0898dd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor","label":"cumulativeInverseSqrtLogFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor","description":"The logarithmic factor shared by every queried cumulative prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-e9dee2a51e4f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7006,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeInverseSqrtLogFactor (mdp : MDP State Action) (rounds : Nat) (delta : Real) : Real","missing":[],"search":"cumulativeinversesqrtlogfactor banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtlogfactor the logarithmic factor shared by every queried cumulative prefix. definition compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCoverCoefficient","label":"cumulativeInverseSqrtCoverCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCoverCoefficient","description":"The deterministic coefficient multiplying each cumulative count radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-fa4a211beee9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7007,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeInverseSqrtCoverCoefficient (mdp : MDP State Action) (rewardBound budget : Real) : Real","missing":[],"search":"cumulativeinversesqrtcovercoefficient banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtcovercoefficient the deterministic coefficient multiplying each cumulative count radius. definition compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold","label":"cumulativeInverseSqrtCalibrationEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold","description":"A sufficient episode threshold for both a half expected-visit margin and the budget branch of the capped transition-radius cover.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-a531f10a242b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7008,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeInverseSqrtCalibrationEpisodeThreshold (mdp : MDP State Action) (rounds : Nat) (delta visitFloor rewardBound budget : Real) : Real","missing":[],"search":"cumulativeinversesqrtcalibrationepisodethreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtcalibrationepisodethreshold a sufficient episode threshold for both a half expected-visit margin and the budget branch of the capped transition-radius cover. definition compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateConfidenceRadius_sq_eq","label":"cumulativeCoordinateConfidenceRadius_sq_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateConfidenceRadius_sq_eq","description":"Exact square of one cumulative coordinate confidence radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-27102768e0d6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7009,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeCoordinateConfidenceRadius_sq_eq (mdp : MDP State Action) {episodes rounds prefixRounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (cumulativeCoordinateConfidenceRadius episodes prefixRounds (cumulativeCountLocalDelta mdp rounds delta)) ^ 2 = (prefixRounds : Real) * (episodes : Real) / 2 * cumulativeInverseSqrtLogFactor mdp rounds delta","missing":[],"search":"cumulativecoordinateconfidenceradius_sq_eq banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativecoordinateconfidenceradius_sq_eq exact square of one cumulative coordinate confidence radius. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_nonneg","label":"cumulativeInverseSqrtLogFactor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_nonneg","description":"The shared logarithmic factor is nonnegative at a valid global delta.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-5cd3f5d94175","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7010,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeInverseSqrtLogFactor_nonneg (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 <= cumulativeInverseSqrtLogFactor mdp rounds delta","missing":[],"search":"cumulativeinversesqrtlogfactor_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtlogfactor_nonneg the shared logarithmic factor is nonnegative at a valid global delta. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.half_cumulativePathVisitExpectedFloor_lt_lowerMargin_of_episodeThreshold","label":"half_cumulativePathVisitExpectedFloor_lt_lowerMargin_of_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.half_cumulativePathVisitExpectedFloor_lt_lowerMargin_of_episodeThreshold","description":"The episode threshold leaves at least half of the predictable expected visits after subtracting the cumulative confidence radius, uniformly over prefixes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-75de1c1bc996","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7011,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem half_cumulativePathVisitExpectedFloor_lt_lowerMargin_of_episodeThreshold (mdp : MDP State Action) {episodes rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) {delta visitFloor rewardBound budget : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hthreshold : cumulativeInverseSqrtCalibrationEpisodeThreshold mdp rounds delta visitFloor rewardBound budget < (episodes : Real)) (round : Fin rounds) : (round + 1 : Real) * (episodes : Real) * visitFloor / 2 < cumulativePathVisitLowerMargin mdp episodes rounds delta visitFloor round","missing":[],"search":"half_cumulativepathvisitexpectedfloor_lt_lowermargin_of_episodethreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.half_cumulativepathvisitexpectedfloor_lt_lowermargin_of_episodethreshold the episode threshold leaves at least half of the predictable expected visits after subtracting the cumulative confidence radius, uniformly over prefixes. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtPathCalibration_of_episodeThreshold","label":"cumulativeInverseSqrtPathCalibration_of_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtPathCalibration_of_episodeThreshold","description":"The explicit episode threshold and scale-square condition construct the full two-scale path calibration; no roundwise cover premise remains.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-62a42f0ca158","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7012,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:175"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeInverseSqrtPathCalibration_of_episodeThreshold (mdp : MDP State Action) {episodes rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) {delta visitFloor rewardBound budget scale : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hrewardBound : 0 <= rewardBound) (hbudget : 0 < budget) (hscale : 0 <= scale) (hscaleCover : (cumulativeInverseSqrtCoverCoefficient mdp rewardBound budget) ^ 2 * cumulativeInverseSqrtLogFactor mdp rounds delta <= scale ^ 2 * visitFloor) (hthreshold : cumulativeInverseSqrtCalibrationEpisodeThreshold mdp rounds delta visitFloor rewardBound budget < (episodes : Real)) : CumulativeInverseSqrtPathCalibration mdp episodes rounds delta visitFloor rewardBound budget scale","missing":[],"search":"cumulativeinversesqrtpathcalibration_of_episodethreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtpathcalibration_of_episodethreshold the explicit episode threshold and scale-square condition construct the full two-scale path calibration; no roundwise cover premise remains. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtEnvelopeSumBound","label":"cumulativeInverseSqrtEnvelopeSumBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtEnvelopeSumBound","description":"Closed-form cap-versus-square-root bound for the complete round sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-87160478ff78","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7013,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:352"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeInverseSqrtEnvelopeSumBound (episodes rounds : Nat) (visitFloor budget scale : Real) : Real","missing":[],"search":"cumulativeinversesqrtenvelopesumbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtenvelopesumbound closed-form cap-versus-square-root bound for the complete round sum. definition compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_cumulativeInverseSqrtRadiusEnvelope_le_explicit","label":"sum_cumulativeInverseSqrtRadiusEnvelope_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_cumulativeInverseSqrtRadiusEnvelope_le_explicit","description":"The round-indexed capped inverse-square-root envelopes sum to the minimum of the linear cap and an explicit square-root-in-rounds bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-03de4f25159c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7014,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:362"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_cumulativeInverseSqrtRadiusEnvelope_le_explicit (mdp : MDP State Action) {episodes rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) {delta visitFloor rewardBound budget scale : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hscale : 0 <= scale) (hthreshold : cumulativeInverseSqrtCalibrationEpisodeThreshold mdp rounds delta visitFloor rewardBound budget < (episodes : Real)) : (∑ round : Fin rounds, cumulativeInverseSqrtRadiusEnvelope mdp episodes rounds delta visitFloor budget scale round) <= cumulativeInverseSqrtEnvelopeSumBound episodes rounds visitFloor budget scale","missing":[],"search":"sum_cumulativeinversesqrtradiusenvelope_le_explicit banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.sum_cumulativeinversesqrtradiusenvelope_le_explicit the round-indexed capped inverse-square-root envelopes sum to the minimum of the linear cap and an explicit square-root-in-rounds bound. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtRecommendedExpectedRegretBound","label":"cumulativeInverseSqrtRecommendedExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtRecommendedExpectedRegretBound","description":"Explicit closed form replacing the terminal's unsimplified round sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-897a545a0076","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7015,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:462"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeInverseSqrtRecommendedExpectedRegretBound (mdp : MDP State Action) (episodes rounds : Nat) (visitFloor budget scale : Real) : Real","missing":[],"search":"cumulativeinversesqrtrecommendedexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtrecommendedexpectedregretbound explicit closed form replacing the terminal's unsimplified round sum. definition compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_horizon_mul_two_cumulativeInverseSqrtRadiusEnvelope_le_explicit","label":"sum_horizon_mul_two_cumulativeInverseSqrtRadiusEnvelope_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_horizon_mul_two_cumulativeInverseSqrtRadiusEnvelope_le_explicit","description":"The terminal's horizon-weighted finite sum is bounded by the closed form.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-e0e8f527b27f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7016,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:469"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_horizon_mul_two_cumulativeInverseSqrtRadiusEnvelope_le_explicit (mdp : MDP State Action) {episodes rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) {delta visitFloor rewardBound budget scale : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hscale : 0 <= scale) (hthreshold : cumulativeInverseSqrtCalibrationEpisodeThreshold mdp rounds delta visitFloor rewardBound budget < (episodes : Real)) : (∑ round : Fin rounds, (mdp.horizon : Real) * (2 * cumulativeInverseSqrtRadiusEnvelope mdp episodes rounds delta visitFloor budget scale round)) <= cumulativeInverseSqrtRecommendedExpectedRegretBound mdp episodes rounds visitFloor budget scale","missing":[],"search":"sum_horizon_mul_two_cumulativeinversesqrtradiusenvelope_le_explicit banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.sum_horizon_mul_two_cumulativeinversesqrtradiusenvelope_le_explicit the terminal's horizon-weighted finite sum is bounded by the closed form. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_closedFormRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_closedFormRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_closedFormRecommendedExpectedRegret","description":"Closed-form route endpoint: deterministic episode and scale inequalities construct the capped calibration and replace the round sum by an explicit minimum of a linear cap and a square-root-in-rounds rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtexplicitrate/index.html#decl-b179170f0c7f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","order":7017,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate.lean:514"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_closedFormRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (rewardBound budget scale : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hbudget : 0 < budget) (hscale : 0 <= scale) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (h…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_closedformrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_closedformrecommendedexpectedregret closed-form route endpoint: deterministic episode and scale inequalities construct the capped calibration and replace the round sum by an explicit minimum of a linear cap and a square-root-in-rounds rate. theorem compiled","shard":"modules/ea9784a14729365e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta","label":"vanishingAverageConfidenceDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta","description":"The confidence budget used at finite window `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-22b880dbe77e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7018,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def vanishingAverageConfidenceDelta (n : Nat) : Real","missing":[],"search":"vanishingaverageconfidencedelta banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingaverageconfidencedelta the confidence budget used at finite window `n`. definition compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes","label":"vanishingDeltaScheduledEpisodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes","description":"The scheduled batch size at finite window `n`, with `n + 1` rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-c9fe125a156a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7019,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def vanishingDeltaScheduledEpisodes (mdp : MDP State Action) (n : Nat) (visitFloor : Real) : Nat","missing":[],"search":"vanishingdeltascheduledepisodes banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledepisodes the scheduled batch size at finite window `n`, with `n + 1` rounds. definition compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageRecommendedExpectedRegretBound","label":"vanishingDeltaScheduledAverageRecommendedExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageRecommendedExpectedRegretBound","description":"The finite-window average recommendation-regret certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-5f45d4b90fc4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7020,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def vanishingDeltaScheduledAverageRecommendedExpectedRegretBound (mdp : MDP State Action) (n : Nat) (visitFloor : Real) : Real","missing":[],"search":"vanishingdeltascheduledaveragerecommendedexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledaveragerecommendedexpectedregretbound the finite-window average recommendation-regret certificate. definition compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_pos","label":"vanishingAverageConfidenceDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_pos","description":"Every confidence budget in the schedule is strictly positive.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-6e9e3d5b9acf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7021,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingAverageConfidenceDelta_pos (n : Nat) : 0 < vanishingAverageConfidenceDelta n","missing":[],"search":"vanishingaverageconfidencedelta_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingaverageconfidencedelta_pos every confidence budget in the schedule is strictly positive. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_le_one","label":"vanishingAverageConfidenceDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_le_one","description":"Every confidence budget in the schedule is at most one.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-79b1a21b9915","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7022,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingAverageConfidenceDelta_le_one (n : Nat) : vanishingAverageConfidenceDelta n <= 1","missing":[],"search":"vanishingaverageconfidencedelta_le_one banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingaverageconfidencedelta_le_one every confidence budget in the schedule is at most one. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_tendsto_zero","label":"vanishingAverageConfidenceDelta_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_tendsto_zero","description":"The real-valued confidence budget tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-c54c27486c40","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7023,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingAverageConfidenceDelta_tendsto_zero : Tendsto vanishingAverageConfidenceDelta atTop (nhds 0)","missing":[],"search":"vanishingaverageconfidencedelta_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingaverageconfidencedelta_tendsto_zero the real-valued confidence budget tends to zero. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_ennreal_tendsto_zero","label":"vanishingAverageConfidenceDelta_ennreal_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_ennreal_tendsto_zero","description":"The `ENNReal` failure budget used by measure bounds tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-c271ce655f73","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7024,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:93"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingAverageConfidenceDelta_ennreal_tendsto_zero : Tendsto (fun n => ENNReal.ofReal (vanishingAverageConfidenceDelta n)) atTop (nhds 0)","missing":[],"search":"vanishingaverageconfidencedelta_ennreal_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingaverageconfidencedelta_ennreal_tendsto_zero the `ennreal` failure budget used by measure bounds tends to zero. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_nonneg","label":"vanishingDeltaScheduledAverageBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_nonneg","description":"The varying-delta finite-window certificate is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-8cc5dc3a7bf4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7025,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingDeltaScheduledAverageBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (n : Nat) {visitFloor : Real} (hvisitFloor : 0 < visitFloor) : 0 <= vanishingDeltaScheduledAverageRecommendedExpectedRegretBound mdp n visitFloor","missing":[],"search":"vanishingdeltascheduledaveragebound_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledaveragebound_nonneg the varying-delta finite-window certificate is nonnegative. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_le_envelope","label":"vanishingDeltaScheduledAverageBound_le_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_le_envelope","description":"The varying-delta certificate is controlled by the same pure rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-66e2eba2febf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7026,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingDeltaScheduledAverageBound_le_envelope (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (n : Nat) {visitFloor : Real} (hvisitFloor : 0 < visitFloor) : vanishingDeltaScheduledAverageRecommendedExpectedRegretBound mdp n visitFloor <= normalizedCumulativeInverseSqrtScheduledAverageEnvelope mdp (n + 1) visitFloor","missing":[],"search":"vanishingdeltascheduledaveragebound_le_envelope banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledaveragebound_le_envelope the varying-delta certificate is controlled by the same pure rate. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_tendsto_zero","label":"vanishingDeltaScheduledAverageBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_tendsto_zero","description":"The varying-delta average recommendation-regret certificate tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-469285b4b376","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7027,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingDeltaScheduledAverageBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) : Tendsto (fun n => vanishingDeltaScheduledAverageRecommendedExpectedRegretBound mdp n visitFloor) atTop (nhds 0)","missing":[],"search":"vanishingdeltascheduledaveragebound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltascheduledaveragebound_tendsto_zero the varying-delta average recommendation-regret certificate tends to zero. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaAndScheduledAverageBound_tendsto_zero","label":"vanishingDeltaAndScheduledAverageBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaAndScheduledAverageBound_tendsto_zero","description":"Failure budget and deterministic certificate jointly tend to `(0, 0)`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-eed263f0766a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7028,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem vanishingDeltaAndScheduledAverageBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) : Tendsto (fun n => (ENNReal.ofReal (vanishingAverageConfidenceDelta n), vanishingDeltaScheduledAverageRecommendedExpectedRegretBound mdp n visitFloor)) atTop (nhds (0, 0))","missing":[],"search":"vanishingdeltaandscheduledaveragebound_tendsto_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.vanishingdeltaandscheduledaveragebound_tendsto_zero failure budget and deterministic certificate jointly tend to `(0, 0)`. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.vanishingDeltaScheduledAverageRegretViolationSet","label":"vanishingDeltaScheduledAverageRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.vanishingDeltaScheduledAverageRegretViolationSet","description":"The trajectories whose average recommendation-regret exceeds the vanishing-delta finite-window certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-595a5122a39b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7029,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def vanishingDeltaScheduledAverageRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (n : Nat) (visitFloor : Real) : Set (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))","missing":[],"search":"vanishingdeltascheduledaverageregretviolationset banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.vanishingdeltascheduledaverageregretviolationset the trajectories whose average recommendation-regret exceeds the vanishing-delta finite-window certificate. definition compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret","description":"At one finite window, the average-regret violation set is contained in the measurable simultaneous count bad event and inherits its vanishing confidence budget as an outer-measure bound. No measurability claim is made for the violation set itself. Outside the measurable bad event, optimism and the explicit average certificate both hold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-adb6d70c92f4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7030,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (n : Nat) (visitFloor : Real) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hvisitFloor…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_vanishingdeltascheduledaveragerecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_vanishingdeltascheduledaveragerecommendedexpectedregret at one finite window, the average-regret violation set is contained in the measurable simultaneous count bad event and inherits its vanishing confidence budget as an outer-measure bound. no measurability claim is made for the violation set itself. outside the measurable bad event, optimism and the explicit average certificate both hold. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret_allWindows","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret_allWindows","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret_allWindows","description":"Dependent family of all finite-window high-probability certificates. The two standard-Borel assumptions are themselves indexed by the changing scheduled sample space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrthighprobabilityaverageconsistency/index.html#decl-95819af6e753","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","order":7031,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret_allWindows (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (visitFloor : Real) (hbatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))) (htrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes mdp n visitFloor))) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hrewardBound : forall state action, |mdp.reward state a…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_vanishingdeltascheduledaveragerecommendedexpectedregret_allwindows banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_vanishingdeltascheduledaveragerecommendedexpectedregret_allwindows dependent family of all finite-window high-probability certificates. the two standard-borel assumptions are themselves indexed by the changing scheduled sample space. theorem compiled","shard":"modules/a45d7177ee4135c4.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale","label":"normalizedCumulativeInverseSqrtScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale","description":"Canonical statistical scale for rewards and zero-count budget bounded by one.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-35a387c19862","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7032,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtScale (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"normalizedcumulativeinversesqrtscale banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscale canonical statistical scale for rewards and zero-count budget bounded by one. definition compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtEpisodeThreshold","label":"normalizedCumulativeInverseSqrtEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtEpisodeThreshold","description":"One sufficient episode threshold for the normalized scale choice.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-61c00197b776","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7033,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtEpisodeThreshold (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"normalizedcumulativeinversesqrtepisodethreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtepisodethreshold one sufficient episode threshold for the normalized scale choice. definition compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtCountRadius","label":"normalizedCumulativeInverseSqrtCountRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtCountRadius","description":"The normalized capped count radius used by the concrete source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-05a5495e5edc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7034,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtCountRadius (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : TransitionCountRadius","missing":[],"search":"normalizedcumulativeinversesqrtcountradius banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtcountradius the normalized capped count radius used by the concrete source. definition compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound","label":"normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound","description":"Closed recommendation-regret bound after fixing budget and scale.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-f283cb4143c7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7035,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound (mdp : MDP State Action) (episodes rounds : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"normalizedcumulativeinversesqrtrecommendedexpectedregretbound banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtrecommendedexpectedregretbound closed recommendation-regret bound after fixing budget and scale. definition compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale_nonneg","label":"normalizedCumulativeInverseSqrtScale_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale_nonneg","description":"The normalized scale is nonnegative without additional scalar premises.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-17ab53cbff00","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7036,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScale_nonneg (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : 0 <= normalizedCumulativeInverseSqrtScale mdp rounds delta visitFloor","missing":[],"search":"normalizedcumulativeinversesqrtscale_nonneg banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscale_nonneg the normalized scale is nonnegative without additional scalar premises. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale_cover","label":"normalizedCumulativeInverseSqrtScale_cover","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale_cover","description":"The normalized scale exactly covers the explicit two-scale coefficient.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-4329cba3fb53","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7037,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:74"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtScale_cover (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) : (cumulativeInverseSqrtCoverCoefficient mdp 1 1) ^ 2 * cumulativeInverseSqrtLogFactor mdp rounds delta <= (normalizedCumulativeInverseSqrtScale mdp rounds delta visitFloor) ^ 2 * visitFloor","missing":[],"search":"normalizedcumulativeinversesqrtscale_cover banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtscale_cover the normalized scale exactly covers the explicit two-scale coefficient. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold_le_normalized","label":"cumulativeInverseSqrtCalibrationEpisodeThreshold_le_normalized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold_le_normalized","description":"The old max threshold is bounded by the single normalized threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-908abd95e86e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7038,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeInverseSqrtCalibrationEpisodeThreshold_le_normalized (mdp : MDP State Action) {rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) : cumulativeInverseSqrtCalibrationEpisodeThreshold mdp rounds delta visitFloor 1 1 <= normalizedCumulativeInverseSqrtEpisodeThreshold mdp rounds delta visitFloor","missing":[],"search":"cumulativeinversesqrtcalibrationepisodethreshold_le_normalized banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtcalibrationepisodethreshold_le_normalized the old max threshold is bounded by the single normalized threshold. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold_lt_of_normalizedThreshold","label":"cumulativeInverseSqrtCalibrationEpisodeThreshold_lt_of_normalizedThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold_lt_of_normalizedThreshold","description":"A strict normalized threshold implies the exact parent threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-79d9167ce283","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7039,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeInverseSqrtCalibrationEpisodeThreshold_lt_of_normalizedThreshold (mdp : MDP State Action) {episodes rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hthreshold : normalizedCumulativeInverseSqrtEpisodeThreshold mdp rounds delta visitFloor < (episodes : Real)) : cumulativeInverseSqrtCalibrationEpisodeThreshold mdp rounds delta visitFloor 1 1 < (episodes : Real)","missing":[],"search":"cumulativeinversesqrtcalibrationepisodethreshold_lt_of_normalizedthreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativeinversesqrtcalibrationepisodethreshold_lt_of_normalizedthreshold a strict normalized threshold implies the exact parent threshold. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtPathCalibration_of_episodeThreshold","label":"normalizedCumulativeInverseSqrtPathCalibration_of_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtPathCalibration_of_episodeThreshold","description":"The single normalized threshold constructs the complete parent calibration.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-8e436d24a3e7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7040,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtPathCalibration_of_episodeThreshold (mdp : MDP State Action) {episodes rounds : Nat} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) {delta visitFloor : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hthreshold : normalizedCumulativeInverseSqrtEpisodeThreshold mdp rounds delta visitFloor < (episodes : Real)) : CumulativeInverseSqrtPathCalibration mdp episodes rounds delta visitFloor 1 1 (normalizedCumulativeInverseSqrtScale mdp rounds delta visitFloor)","missing":[],"search":"normalizedcumulativeinversesqrtpathcalibration_of_episodethreshold banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtpathcalibration_of_episodethreshold the single normalized threshold constructs the complete parent calibration. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound_eq","label":"normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound_eq","description":"Explicit expansion of the normalized recommendation-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-fb5baf599254","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7041,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:185"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound_eq (mdp : MDP State Action) (episodes rounds : Nat) (delta visitFloor : Real) : normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound mdp episodes rounds delta visitFloor = 2 * (mdp.horizon : Real) * min (rounds : Real) (8 * (Fintype.card State : Real) * (mdp.horizon : Real) * Real.sqrt (cumulativeInverseSqrtLogFactor mdp rounds delta) / Real.sqrt visitFloor * Real.sqrt (rounds : Real) / Real.sqrt ((episodes : Real) * visitFloor / 2))","missing":[],"search":"normalizedcumulativeinversesqrtrecommendedexpectedregretbound_eq banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.normalizedcumulativeinversesqrtrecommendedexpectedregretbound_eq explicit expansion of the normalized recommendation-regret bound. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedRecommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedRecommendedExpectedRegret","description":"Normalized deterministic-reward endpoint with one episode threshold and no caller-visible budget or statistical scale.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeinversesqrtnormalizedrate/index.html#decl-4b0edaba14cf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","order":7042,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate.lean:221"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedRecommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes rounds : Nat) [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hthreshol…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_normalizedrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_normalizedrecommendedexpectedregret normalized deterministic-reward endpoint with one episode threshold and no caller-visible budget or statistical scale. theorem compiled","shard":"modules/31c20cbc996b71d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.traceStateAtFrom","label":"traceStateAtFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.traceStateAtFrom","description":"def traceStateAtFrom (state : State) {n : Nat} (trace : StepTrace Action State n) (stage : Fin n) : State","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-db8e8e3a3cce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7043,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def traceStateAtFrom (state : State) {n : Nat} (trace : StepTrace Action State n) (stage : Fin n) : State","missing":[],"search":"tracestateatfrom banditrlproof.finitehorizonrl.mdp.tracestateatfrom def tracestateatfrom (state : state) {n : nat} (trace : steptrace action state n) (stage : fin n) : state definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.traceStateAtFrom_tail","label":"traceStateAtFrom_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.traceStateAtFrom_tail","description":"theorem traceStateAtFrom_tail (state : State) {n : Nat} (trace : StepTrace Action State (n + 1)) (stage : Fin n) : traceStateAtFrom (trace 0).2 (Fin.tail trace) stage = traceStateAtFrom state trace stage.succ","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-3436c51587ef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7044,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem traceStateAtFrom_tail (state : State) {n : Nat} (trace : StepTrace Action State (n + 1)) (stage : Fin n) : traceStateAtFrom (trace 0).2 (Fin.tail trace) stage = traceStateAtFrom state trace stage.succ","missing":[],"search":"tracestateatfrom_tail banditrlproof.finitehorizonrl.mdp.tracestateatfrom_tail theorem tracestateatfrom_tail (state : state) {n : nat} (trace : steptrace action state (n + 1)) (stage : fin n) : tracestateatfrom (trace 0).2 (fin.tail trace) stage = tracestateatfrom state trace stage.succ theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.cumulativeRewardFrom_eq_sum_traceReward","label":"cumulativeRewardFrom_eq_sum_traceReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.cumulativeRewardFrom_eq_sum_traceReward","description":"theorem cumulativeRewardFrom_eq_sum_traceReward (mdp : MDP State Action) (n : Nat) (state : State) (trace : StepTrace Action State n) : mdp.cumulativeRewardFrom n state trace = ∑ stage : Fin n, mdp.reward (traceStateAtFrom state trace stage) (trace stage).1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-9003883c90f4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7045,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeRewardFrom_eq_sum_traceReward (mdp : MDP State Action) (n : Nat) (state : State) (trace : StepTrace Action State n) : mdp.cumulativeRewardFrom n state trace = ∑ stage : Fin n, mdp.reward (traceStateAtFrom state trace stage) (trace stage).1","missing":[],"search":"cumulativerewardfrom_eq_sum_tracereward banditrlproof.finitehorizonrl.mdp.cumulativerewardfrom_eq_sum_tracereward theorem cumulativerewardfrom_eq_sum_tracereward (mdp : mdp state action) (n : nat) (state : state) (trace : steptrace action state n) : mdp.cumulativerewardfrom n state trace = ∑ stage : fin n, mdp.reward (tracestateatfrom state trace stage) (trace stage).1 theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.episodeReturn","label":"episodeReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.episodeReturn","description":"def episodeReturn {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (episode : Fin episodes) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-5319e772e5d8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7046,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:74"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def episodeReturn {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (episode : Fin episodes) : Real","missing":[],"search":"episodereturn banditrlproof.finitehorizonrl.episodebatch.episodereturn def episodereturn {mdp : mdp state action} {episodes : nat} (batch : episodebatch mdp episodes) (episode : fin episodes) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.totalReturn","label":"totalReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.totalReturn","description":"def totalReturn {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-8f7a6afa0f49","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7047,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def totalReturn {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) : Real","missing":[],"search":"totalreturn banditrlproof.finitehorizonrl.episodebatch.totalreturn def totalreturn {mdp : mdp state action} {episodes : nat} (batch : episodebatch mdp episodes) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_episodeReturn","label":"measurable_episodeReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_episodeReturn","description":"theorem measurable_episodeReturn {mdp : MDP State Action} {episodes : Nat} (episode : Fin episodes) : Measurable (fun batch : EpisodeBatch mdp episodes => episodeReturn batch episode)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-2a3f97267725","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7048,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_episodeReturn {mdp : MDP State Action} {episodes : Nat} (episode : Fin episodes) : Measurable (fun batch : EpisodeBatch mdp episodes => episodeReturn batch episode)","missing":[],"search":"measurable_episodereturn banditrlproof.finitehorizonrl.episodebatch.measurable_episodereturn theorem measurable_episodereturn {mdp : mdp state action} {episodes : nat} (episode : fin episodes) : measurable (fun batch : episodebatch mdp episodes => episodereturn batch episode) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_totalReturn","label":"measurable_totalReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_totalReturn","description":"theorem measurable_totalReturn {mdp : MDP State Action} {episodes : Nat} : Measurable (totalReturn : EpisodeBatch mdp episodes -> Real)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-0a14ed9e8f34","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7049,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:90"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_totalReturn {mdp : MDP State Action} {episodes : Nat} : Measurable (totalReturn : EpisodeBatch mdp episodes -> Real)","missing":[],"search":"measurable_totalreturn banditrlproof.finitehorizonrl.episodebatch.measurable_totalreturn theorem measurable_totalreturn {mdp : mdp state action} {episodes : nat} : measurable (totalreturn : episodebatch mdp episodes -> real) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_episodeReturn_le_horizon_of_rewardConsistent","label":"abs_episodeReturn_le_horizon_of_rewardConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_episodeReturn_le_horizon_of_rewardConsistent","description":"theorem abs_episodeReturn_le_horizon_of_rewardConsistent {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (hbatch : batch.RewardConsistent) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (episode : Fin episodes) : |episodeReturn batch episode| <= (mdp.horizon : Real)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-ba6dec9fdc87","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7050,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:95"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_episodeReturn_le_horizon_of_rewardConsistent {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (hbatch : batch.RewardConsistent) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (episode : Fin episodes) : |episodeReturn batch episode| <= (mdp.horizon : Real)","missing":[],"search":"abs_episodereturn_le_horizon_of_rewardconsistent banditrlproof.finitehorizonrl.episodebatch.abs_episodereturn_le_horizon_of_rewardconsistent theorem abs_episodereturn_le_horizon_of_rewardconsistent {mdp : mdp state action} {episodes : nat} (batch : episodebatch mdp episodes) (hbatch : batch.rewardconsistent) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (episode : fin episodes) : |episodereturn batch episode| <= (mdp.horizon : real) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_totalReturn_le_of_rewardConsistent","label":"abs_totalReturn_le_of_rewardConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_totalReturn_le_of_rewardConsistent","description":"theorem abs_totalReturn_le_of_rewardConsistent {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (hbatch : batch.RewardConsistent) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : |totalReturn batch| <= (episodes : Real) * (mdp.horizon : Real)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-26af6d60a9bd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7051,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_totalReturn_le_of_rewardConsistent {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (hbatch : batch.RewardConsistent) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : |totalReturn batch| <= (episodes : Real) * (mdp.horizon : Real)","missing":[],"search":"abs_totalreturn_le_of_rewardconsistent banditrlproof.finitehorizonrl.episodebatch.abs_totalreturn_le_of_rewardconsistent theorem abs_totalreturn_le_of_rewardconsistent {mdp : mdp state action} {episodes : nat} (batch : episodebatch mdp episodes) (hbatch : batch.rewardconsistent) (hrewardbound : forall state action, |mdp.reward state action| <= 1) : |totalreturn batch| <= (episodes : real) * (mdp.horizon : real) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeReturn_episodeBatchOfTrajectories","label":"episodeReturn_episodeBatchOfTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeReturn_episodeBatchOfTrajectories","description":"theorem episodeReturn_episodeBatchOfTrajectories (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (episode : Fin episodes) : EpisodeBatch.episodeReturn (mdp.episodeBatchOfTrajectories episodes trajectories) episode = mdp.cumulativeReward (trajectories episode)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-240e5b580fcc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7052,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:132"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeReturn_episodeBatchOfTrajectories (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (episode : Fin episodes) : EpisodeBatch.episodeReturn (mdp.episodeBatchOfTrajectories episodes trajectories) episode = mdp.cumulativeReward (trajectories episode)","missing":[],"search":"episodereturn_episodebatchoftrajectories banditrlproof.finitehorizonrl.mdp.episodereturn_episodebatchoftrajectories theorem episodereturn_episodebatchoftrajectories (mdp : mdp state action) (episodes : nat) (trajectories : fin episodes -> state × steptrace action state mdp.horizon) (episode : fin episodes) : episodebatch.episodereturn (mdp.episodebatchoftrajectories episodes trajectories) episode = mdp.cumulativereward (trajectories episode) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.totalReturn_episodeBatchOfTrajectories","label":"totalReturn_episodeBatchOfTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.totalReturn_episodeBatchOfTrajectories","description":"theorem totalReturn_episodeBatchOfTrajectories (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) : EpisodeBatch.totalReturn (mdp.episodeBatchOfTrajectories episodes trajectories) = ∑ episode : Fin episodes, mdp.cumulativeReward (trajectories episode)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-41676f114bfd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7053,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem totalReturn_episodeBatchOfTrajectories (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) : EpisodeBatch.totalReturn (mdp.episodeBatchOfTrajectories episodes trajectories) = ∑ episode : Fin episodes, mdp.cumulativeReward (trajectories episode)","missing":[],"search":"totalreturn_episodebatchoftrajectories banditrlproof.finitehorizonrl.mdp.totalreturn_episodebatchoftrajectories theorem totalreturn_episodebatchoftrajectories (mdp : mdp state action) (episodes : nat) (trajectories : fin episodes -> state × steptrace action state mdp.horizon) : episodebatch.totalreturn (mdp.episodebatchoftrajectories episodes trajectories) = ∑ episode : fin episodes, mdp.cumulativereward (trajectories episode) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.batchReturnVarianceProxy","label":"batchReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.batchReturnVarianceProxy","description":"noncomputable def batchReturnVarianceProxy (mdp : MDP State Action) (episodes : Nat) : NNReal","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-578a0272ffdf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7054,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:159"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def batchReturnVarianceProxy (mdp : MDP State Action) (episodes : Nat) : NNReal","missing":[],"search":"batchreturnvarianceproxy banditrlproof.finitehorizonrl.markovpolicy.batchreturnvarianceproxy noncomputable def batchreturnvarianceproxy (mdp : mdp state action) (episodes : nat) : nnreal definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeReward_eval_iidTrajectoryFamilyMeasure","label":"integral_cumulativeReward_eval_iidTrajectoryFamilyMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeReward_eval_iidTrajectoryFamilyMeasure","description":"theorem integral_cumulativeReward_eval_iidTrajectoryFamilyMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) : (∫ trajectories, mdp.cumulativeReward (trajectories episode) ∂policy.iidTrajectoryFamilyMeasure initialState episodes) = ∫ trajectory, mdp.cumulativeReward trajectory ∂policy.trajectoryMeas…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-a63ea118e01d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7055,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_cumulativeReward_eval_iidTrajectoryFamilyMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) : (∫ trajectories, mdp.cumulativeReward (trajectories episode) ∂policy.iidTrajectoryFamilyMeasure initialState episodes) = ∫ trajectory, mdp.cumulativeReward trajectory ∂policy.trajectoryMeasure initialState","missing":[],"search":"integral_cumulativereward_eval_iidtrajectoryfamilymeasure banditrlproof.finitehorizonrl.markovpolicy.integral_cumulativereward_eval_iidtrajectoryfamilymeasure theorem integral_cumulativereward_eval_iidtrajectoryfamilymeasure {mdp : mdp state action} (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] {episodes : nat} (episode : fin episodes) : (∫ trajectories, mdp.cumulativereward (trajectories episode) ∂policy.iidtrajectoryfamilymeasure initialstate episodes) = ∫ trajectory, mdp.cumulativereward trajectory ∂policy.trajectorymeasure initialstate theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_totalReturn_iidEpisodeBatchMeasure","label":"integral_totalReturn_iidEpisodeBatchMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_totalReturn_iidEpisodeBatchMeasure","description":"theorem integral_totalReturn_iidEpisodeBatchMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : integral (policy.iidEpisodeBatchMeasure initialState episodes) EpisodeBatch.totalReturn = (episodes : Real) * integral (policy.trajectoryMeasure initialState) mdp.cumulativeReward","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-b7733d07305b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7056,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_totalReturn_iidEpisodeBatchMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : integral (policy.iidEpisodeBatchMeasure initialState episodes) EpisodeBatch.totalReturn = (episodes : Real) * integral (policy.trajectoryMeasure initialState) mdp.cumulativeReward","missing":[],"search":"integral_totalreturn_iidepisodebatchmeasure banditrlproof.finitehorizonrl.markovpolicy.integral_totalreturn_iidepisodebatchmeasure theorem integral_totalreturn_iidepisodebatchmeasure {mdp : mdp state action} (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) : integral (policy.iidepisodebatchmeasure initialstate episodes) episodebatch.totalreturn = (episodes : real) * integral (policy.trajectorymeasure initialstate) mdp.cumulativereward theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.totalReturn_centered_hasSubgaussianMGF","label":"totalReturn_centered_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.totalReturn_centered_hasSubgaussianMGF","description":"theorem totalReturn_centered_hasSubgaussianMGF {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ProbabilityTheory.HasSubgaussianMGF (fun batch : EpisodeBatch mdp episodes => EpisodeBatch.totalReturn batch - integral (policy.iidEpisodeBatchMeasure initialState…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-883789bcccd3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7057,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem totalReturn_centered_hasSubgaussianMGF {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ProbabilityTheory.HasSubgaussianMGF (fun batch : EpisodeBatch mdp episodes => EpisodeBatch.totalReturn batch - integral (policy.iidEpisodeBatchMeasure initialState episodes) EpisodeBatch.totalReturn) (batchReturnVarianceProxy mdp episodes) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"totalreturn_centered_hassubgaussianmgf banditrlproof.finitehorizonrl.markovpolicy.totalreturn_centered_hassubgaussianmgf theorem totalreturn_centered_hassubgaussianmgf {mdp : mdp state action} (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) (hrewardbound : forall state action, |mdp.reward state action| <= 1) : probabilitytheory.hassubgaussianmgf (fun batch : episodebatch mdp episodes => episodebatch.totalreturn batch - integral (policy.iidepisodebatchmeasure initialstate episodes) episodebatch.totalreturn) (batchreturnvarianceproxy mdp episodes) (policy.iidepisodebatchmeasure initialstate episodes) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnKernelMean","label":"successorReturnKernelMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnKernelMean","description":"noncomputable def successorReturnKernelMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-3589355cda04","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7058,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:244"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorReturnKernelMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : Real","missing":[],"search":"successorreturnkernelmean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnkernelmean noncomputable def successorreturnkernelmean {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (n : nat) (history : episodebatchprefix mdp episodes n) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_successorReturnKernelMean","label":"measurable_successorReturnKernelMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_successorReturnKernelMean","description":"theorem measurable_successorReturnKernelMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : Measurable (source.successorReturnKernelMean n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-de0816fa1484","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7059,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorReturnKernelMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : Measurable (source.successorReturnKernelMean n)","missing":[],"search":"measurable_successorreturnkernelmean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_successorreturnkernelmean theorem measurable_successorreturnkernelmean {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (n : nat) : measurable (source.successorreturnkernelmean n) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnKernelMean_eq_selectedPolicy","label":"successorReturnKernelMean_eq_selectedPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnKernelMean_eq_selectedPolicy","description":"theorem successorReturnKernelMean_eq_selectedPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : source.successorReturnKernelMean n history = (episodes : Real) * integral ((source.successorPolicy n history).trajectoryMeasure initialS…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-f7162f6b9a55","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7060,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:258"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorReturnKernelMean_eq_selectedPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : source.successorReturnKernelMean n history = (episodes : Real) * integral ((source.successorPolicy n history).trajectoryMeasure initialState) mdp.cumulativeReward","missing":[],"search":"successorreturnkernelmean_eq_selectedpolicy banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnkernelmean_eq_selectedpolicy theorem successorreturnkernelmean_eq_selectedpolicy {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (n : nat) (history : episodebatchprefix mdp episodes n) : source.successorreturnkernelmean n history = (episodes : real) * integral ((source.successorpolicy n history).trajectorymeasure initialstate) mdp.cumulativereward theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnPrefixIncrement","label":"successorReturnPrefixIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnPrefixIncrement","description":"noncomputable def successorReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) : (round : Nat) -> EpisodeBatchPrefix mdp episodes round -> Real | 0, _history => 0 | n + 1, history => EpisodeBatch.totalReturn (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) - source.successorRetur…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-7326a67a9773","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7061,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:273"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) : (round : Nat) -> EpisodeBatchPrefix mdp episodes round -> Real | 0, _history => 0 | n + 1, history => EpisodeBatch.totalReturn (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) - source.successorReturnKernelMean n (Preorder.frestrictLe₂ (π := fun _ : Nat => EpisodeBatch mdp episodes) (Nat.le_succ n) history) theorem measurable_successorReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.successorReturnPrefixIncrement round)","missing":[],"search":"successorreturnprefixincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnprefixincrement noncomputable def successorreturnprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) : (round : nat) -> episodebatchprefix mdp episodes round -> real | 0, _history => 0 | n + 1, history => episodebatch.totalreturn (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩) - source.successorreturnkernelmean n (preorder.frestrictle₂ (π := fun _ : nat => episodebatch mdp episodes) (nat.le_succ n) history) theorem measurable_successorreturnprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (round : nat) : measurable (source.successorreturnprefixincrement round) definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_successorReturnPrefixIncrement","label":"measurable_successorReturnPrefixIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_successorReturnPrefixIncrement","description":"theorem measurable_successorReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.successorReturnPrefixIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-9eac5ed84530","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7062,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:287"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.successorReturnPrefixIncrement round)","missing":[],"search":"measurable_successorreturnprefixincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_successorreturnprefixincrement theorem measurable_successorreturnprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (round : nat) : measurable (source.successorreturnprefixincrement round) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement","label":"successorReturnIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement","description":"noncomputable def successorReturnIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-921fd98339df","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7063,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:307"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorReturnIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorreturnincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnincrement noncomputable def successorreturnincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (round : nat) (trajectory : episodebatchtrajectory mdp episodes) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_stronglyAdapted_piLE","label":"successorReturnIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_stronglyAdapted_piLE","description":"theorem successorReturnIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes)) source.successorReturnIncrement","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-1e9e09935e67","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7064,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:315"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorReturnIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes)) source.successorReturnIncrement","missing":[],"search":"successorreturnincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnincrement_stronglyadapted_pile theorem successorreturnincrement_stronglyadapted_pile {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) : stronglyadapted (filtration.pile (x := fun _ : nat => episodebatch mdp episodes)) source.successorreturnincrement theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_totalReturn","label":"trajectoryMeasure_condDistrib_totalReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_totalReturn","description":"theorem trajectoryMeasure_condDistrib_totalReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => EpisodeBatch.totalReturn (trajectory (n + 1))) (…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-059860871bfa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7065,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:327"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_totalReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => EpisodeBatch.totalReturn (trajectory (n + 1))) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] (source.batchKernel n).map EpisodeBatch.totalReturn","missing":[],"search":"trajectorymeasure_conddistrib_totalreturn banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conddistrib_totalreturn theorem trajectorymeasure_conddistrib_totalreturn {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (n : nat) : probabilitytheory.conddistrib (fun trajectory : episodebatchtrajectory mdp episodes => episodebatch.totalreturn (trajectory (n + 1))) (preorder.frestrictle n) source.trajectorymeasure =ᵐ[ source.trajectorymeasure.map (preorder.frestrictle n)] (source.batchkernel n).map episodebatch.totalreturn theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_totalReturn_eq_batchKernel","label":"condExpKernel_map_totalReturn_eq_batchKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_totalReturn_eq_batchKernel","description":"theorem condExpKernel_map_totalReturn_eq_batchKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : Filter.Eventually (fun trajectory : EpisodeBatchTrajectory mdp episodes =…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-60fa718cf4ee","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7066,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:370"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_totalReturn_eq_batchKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : Filter.Eventually (fun trajectory : EpisodeBatchTrajectory mdp episodes => Measure.map (fun path : EpisodeBatchTrajectory mdp episodes => EpisodeBatch.totalReturn (path (n + 1))) (ProbabilityTheory.condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (EpisodeBatchPrefix mdp episodes n)).comap (Preorder.frestrictLe n)) trajectory) = ((source.batchKernel n).map EpisodeBatch.totalReturn) (Preorder.frestrictLe n trajectory)) (ae (source.trajectoryMeasure.trim (Preorder.measurable_frestrictLe n).comap_le))","missing":[],"search":"condexpkernel_map_totalreturn_eq_batchkernel banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.condexpkernel_map_totalreturn_eq_batchkernel theorem condexpkernel_map_totalreturn_eq_batchkernel {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] [standardborelspace (episodebatchtrajectory mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (n : nat) : filter.eventually (fun trajectory : episodebatchtrajectory mdp episodes => measure.map (fun path : episodebatchtrajectory mdp episodes => episodebatch.totalreturn (path (n + 1))) (probabilitytheory.condexpkernel source.trajectorymeasure ((inferinstance : measurablespace (episodebatchprefix mdp episodes n)).comap (preorder.frestrictle n)) trajectory) = ((source.batchkernel n).map episodebatch.totalreturn) (preorder.frestrictle n trajectory)) (ae (source.trajectorymeasure.trim (preorder.measurable_frestrictle n).comap_le)) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_hasCondSubgaussianMGF","label":"successorReturnIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_hasCondSubgaussianMGF","description":"theorem successorReturnIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (hrewardBound : forall state action, |mdp.reward state action| <= 1)…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-18e6b7ac8c41","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7067,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:409"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorReturnIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : ProbabilityTheory.HasCondSubgaussianMGF (Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes) n) ((Filtration.piLE (X := fun _ : Nat => EpisodeBatch mdp episodes)).le n) (source.successorReturnIncrement (n + 1)) (MarkovPolicy.batchReturnVarianceProxy mdp episodes) source.trajectoryMeasure","missing":[],"search":"successorreturnincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnincrement_succ_hascondsubgaussianmgf theorem successorreturnincrement_succ_hascondsubgaussianmgf {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] [standardborelspace (episodebatchtrajectory mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (n : nat) (hrewardbound : forall state action, |mdp.reward state action| <= 1) : probabilitytheory.hascondsubgaussianmgf (filtration.pile (x := fun _ : nat => episodebatch mdp episodes) n) ((filtration.pile (x := fun _ : nat => episodebatch mdp episodes)).le n) (source.successorreturnincrement (n + 1)) (markovpolicy.batchreturnvarianceproxy mdp episodes) source.trajectorymeasure theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnVarianceProxy","label":"cumulativeSuccessorReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnVarianceProxy","description":"noncomputable def cumulativeSuccessorReturnVarianceProxy (mdp : MDP State Action) (episodes rounds : Nat) : NNReal","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-9278de9a4434","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7068,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:523"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSuccessorReturnVarianceProxy (mdp : MDP State Action) (episodes rounds : Nat) : NNReal","missing":[],"search":"cumulativesuccessorreturnvarianceproxy banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativesuccessorreturnvarianceproxy noncomputable def cumulativesuccessorreturnvarianceproxy (mdp : mdp state action) (episodes rounds : nat) : nnreal definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnDeviation","label":"cumulativeSuccessorReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnDeviation","description":"noncomputable def cumulativeSuccessorReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-86e3ea7e3239","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7069,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:530"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSuccessorReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : EpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativesuccessorreturndeviation banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativesuccessorreturndeviation noncomputable def cumulativesuccessorreturndeviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (trajectory : episodebatchtrajectory mdp episodes) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchReturnVarianceProxy_pos","label":"batchReturnVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchReturnVarianceProxy_pos","description":"theorem batchReturnVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) : 0 < MarkovPolicy.batchReturnVarianceProxy mdp episodes","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-f874ec0a7177","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7070,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:537"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem batchReturnVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) : 0 < MarkovPolicy.batchReturnVarianceProxy mdp episodes","missing":[],"search":"batchreturnvarianceproxy_pos banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.batchreturnvarianceproxy_pos theorem batchreturnvarianceproxy_pos (mdp : mdp state action) (episodes : nat) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) : 0 < markovpolicy.batchreturnvarianceproxy mdp episodes theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorReturnDeviation_abs_tail_le","label":"trajectoryMeasure_cumulativeSuccessorReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorReturnDeviation_abs_tail_le","description":"theorem trajectoryMeasure_cumulativeSuccessorReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes)…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-42645d37064f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7071,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:553"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSuccessorReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectoryMeasure {trajectory | Concentration.subGaussianSumConfidenceRadius (cumulativeSuccessorReturnVarianceProxy mdp episodes rounds) delta <= |source.cumulativeSuccessorReturnDeviation rounds trajectory|} <= ENNReal.ofReal delta","missing":[],"search":"trajectorymeasure_cumulativesuccessorreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_cumulativesuccessorreturndeviation_abs_tail_le theorem trajectorymeasure_cumulativesuccessorreturndeviation_abs_tail_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] [standardborelspace (episodebatchtrajectory mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectorymeasure {trajectory | concentration.subgaussiansumconfidenceradius (cumulativesuccessorreturnvarianceproxy mdp episodes rounds) delta <= |source.cumulativesuccessorreturndeviation rounds trajectory|} <= ennreal.ofreal delta theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.optimalInitialExpectedReturn","label":"optimalInitialExpectedReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.optimalInitialExpectedReturn","description":"noncomputable def optimalInitialExpectedReturn (mdp : MDP State Action) (initialState : Measure State) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-0260c4c234a4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7072,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:600"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalInitialExpectedReturn (mdp : MDP State Action) (initialState : Measure State) : Real","missing":[],"search":"optimalinitialexpectedreturn banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.optimalinitialexpectedreturn noncomputable def optimalinitialexpectedreturn (mdp : mdp state action) (initialstate : measure state) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedCumulativeRegret","label":"successorExpectedCumulativeRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedCumulativeRegret","description":"noncomputable def successorExpectedCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-0089ce0e1878","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7073,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:605"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorExpectedCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"successorexpectedcumulativeregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorexpectedcumulativeregret noncomputable def successorexpectedcumulativeregret {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedAverageRegret","label":"successorExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedAverageRegret","description":"noncomputable def successorExpectedAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-f6bf820a3d4e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7074,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:613"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorExpectedAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"successorexpectedaverageregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorexpectedaverageregret noncomputable def successorexpectedaverageregret {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorCumulativeRegret","label":"realizedSuccessorCumulativeRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorCumulativeRegret","description":"noncomputable def realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-bad5a84b0aa5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7075,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:620"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"realizedsuccessorcumulativeregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.realizedsuccessorcumulativeregret noncomputable def realizedsuccessorcumulativeregret {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorAverageRegret","label":"realizedSuccessorAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorAverageRegret","description":"noncomputable def realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-34ccc20c8dfb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7076,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:629"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.realizedsuccessoraverageregret noncomputable def realizedsuccessoraverageregret {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : real definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnDeviationBadEvent","label":"successorReturnDeviationBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnDeviationBadEvent","description":"noncomputable def successorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-1881de4f40ba","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7077,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:637"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"successorreturndeviationbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturndeviationbadevent noncomputable def successorreturndeviationbadevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (delta : real) : set (episodebatchtrajectory mdp episodes) definition compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_cumulativeSuccessorReturnDeviation","label":"measurable_cumulativeSuccessorReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_cumulativeSuccessorReturnDeviation","description":"theorem measurable_cumulativeSuccessorReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (source.cumulativeSuccessorReturnDeviation rounds)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-8ce35bb3aa9d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7078,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:648"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeSuccessorReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (source.cumulativeSuccessorReturnDeviation rounds)","missing":[],"search":"measurable_cumulativesuccessorreturndeviation banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_cumulativesuccessorreturndeviation theorem measurable_cumulativesuccessorreturndeviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) : measurable (source.cumulativesuccessorreturndeviation rounds) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_successorReturnDeviationBadEvent","label":"measurableSet_successorReturnDeviationBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_successorReturnDeviationBadEvent","description":"theorem measurableSet_successorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : MeasurableSet (source.successorReturnDeviationBadEvent rounds delta)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-5c576ad54f2a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7079,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:659"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_successorReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : MeasurableSet (source.successorReturnDeviationBadEvent rounds delta)","missing":[],"search":"measurableset_successorreturndeviationbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurableset_successorreturndeviationbadevent theorem measurableset_successorreturndeviationbadevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (delta : real) : measurableset (source.successorreturndeviationbadevent rounds delta) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_eq_totalReturn_sub_selectedPolicyMean","label":"successorReturnIncrement_succ_eq_totalReturn_sub_selectedPolicyMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_eq_totalReturn_sub_selectedPolicyMean","description":"theorem successorReturnIncrement_succ_eq_totalReturn_sub_selectedPolicyMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (n : Nat) : source.successorReturnIncrement (n + 1) trajectory = EpisodeBatch.totalReturn (trajectory (n + 1)) - (episo…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-3687cbfe8616","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7080,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:668"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorReturnIncrement_succ_eq_totalReturn_sub_selectedPolicyMean {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (n : Nat) : source.successorReturnIncrement (n + 1) trajectory = EpisodeBatch.totalReturn (trajectory (n + 1)) - (episodes : Real) * integral ((source.policyAt trajectory (n + 1)).trajectoryMeasure initialState) mdp.cumulativeReward","missing":[],"search":"successorreturnincrement_succ_eq_totalreturn_sub_selectedpolicymean banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorreturnincrement_succ_eq_totalreturn_sub_selectedpolicymean theorem successorreturnincrement_succ_eq_totalreturn_sub_selectedpolicymean {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (n : nat) : source.successorreturnincrement (n + 1) trajectory = episodebatch.totalreturn (trajectory (n + 1)) - (episodes : real) * integral ((source.policyat trajectory (n + 1)).trajectorymeasure initialstate) mdp.cumulativereward theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnDeviation_eq_fin_sum","label":"cumulativeSuccessorReturnDeviation_eq_fin_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnDeviation_eq_fin_sum","description":"theorem cumulativeSuccessorReturnDeviation_eq_fin_sum {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.cumulativeSuccessorReturnDeviation rounds trajectory = ∑ round : Fin rounds, source.successorReturnIncrement ((round…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-dcb5951d4c65","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7081,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:684"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorReturnDeviation_eq_fin_sum {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.cumulativeSuccessorReturnDeviation rounds trajectory = ∑ round : Fin rounds, source.successorReturnIncrement ((round : Nat) + 1) trajectory","missing":[],"search":"cumulativesuccessorreturndeviation_eq_fin_sum banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.cumulativesuccessorreturndeviation_eq_fin_sum theorem cumulativesuccessorreturndeviation_eq_fin_sum {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : source.cumulativesuccessorreturndeviation rounds trajectory = ∑ round : fin rounds, source.successorreturnincrement ((round : nat) + 1) trajectory theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","label":"realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","description":"theorem realizedSuccessorCumulativeRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.realizedSuccessorCumulativeRegret trajectory rounds = (episodes : Real) * source.successorExpectedCumul…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-3991d7eba902","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7082,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:701"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem realizedSuccessorCumulativeRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.realizedSuccessorCumulativeRegret trajectory rounds = (episodes : Real) * source.successorExpectedCumulativeRegret trajectory rounds - source.cumulativeSuccessorReturnDeviation rounds trajectory","missing":[],"search":"realizedsuccessorcumulativeregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.realizedsuccessorcumulativeregret_eq_expected_sub_deviation theorem realizedsuccessorcumulativeregret_eq_expected_sub_deviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : source.realizedsuccessorcumulativeregret trajectory rounds = (episodes : real) * source.successorexpectedcumulativeregret trajectory rounds - source.cumulativesuccessorreturndeviation rounds trajectory theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","label":"realizedSuccessorAverageRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","description":"theorem realizedSuccessorAverageRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) : source.realizedSuccessorAverageRegret trajectory rounds = sourc…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-189dc6f36165","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7083,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:733"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem realizedSuccessorAverageRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) : source.realizedSuccessorAverageRegret trajectory rounds = source.successorExpectedAverageRegret trajectory rounds - source.cumulativeSuccessorReturnDeviation rounds trajectory / ((episodes : Real) * (rounds : Real))","missing":[],"search":"realizedsuccessoraverageregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.realizedsuccessoraverageregret_eq_expected_sub_deviation theorem realizedsuccessoraverageregret_eq_expected_sub_deviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptiveepisodebatchsource mdp initialstate episodes) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) : source.realizedsuccessoraverageregret trajectory rounds = source.successorexpectedaverageregret trajectory rounds - source.cumulativesuccessorreturndeviation rounds trajectory / ((episodes : real) * (rounds : real)) theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successorReturnDeviationBadEvent_le","label":"trajectoryMeasure_successorReturnDeviationBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successorReturnDeviationBadEvent_le","description":"theorem trajectoryMeasure_successorReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon :…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-a2c20ab6cbbe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7084,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:750"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successorReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectoryMeasure (source.successorReturnDeviationBadEvent rounds delta) <= ENNReal.ofReal delta","missing":[],"search":"trajectorymeasure_successorreturndeviationbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_successorreturndeviationbadevent_le theorem trajectorymeasure_successorreturndeviationbadevent_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] [standardborelspace (episodebatchtrajectory mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectorymeasure (source.successorreturndeviationbadevent rounds delta) <= ennreal.ofreal delta theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport","label":"trajectoryMeasure_expected_to_realized_successor_average_regret_transport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport","description":"theorem trajectoryMeasure_expected_to_realized_successor_average_regret_transport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < e…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-3fc02ea37637","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7085,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:767"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_expected_to_realized_successor_average_regret_transport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (countBadEvent : Set (EpisodeBatchTrajectory mdp episodes)) (expectedBound : Real) (Good : EpisodeBatchTrajectory mdp episodes -> Prop) (hcountMeasurable : MeasurableSet countBadEvent) (hcountTail : source.trajectoryMeasure countBadEvent <= ENNReal.ofReal delta) (hcountGood : forall trajectory, trajectory ∉…","missing":[],"search":"trajectorymeasure_expected_to_realized_successor_average_regret_transport banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_expected_to_realized_successor_average_regret_transport theorem trajectorymeasure_expected_to_realized_successor_average_regret_transport {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (episodebatch mdp episodes)] [standardborelspace (episodebatchtrajectory mdp episodes)] (source : adaptiveepisodebatchsource mdp initialstate episodes) (rounds : nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hhorizon : 0 < mdp.horizon) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (countbadevent : set (episodebatchtrajectory mdp episodes)) (expectedbound : real) (good : episodebatchtrajectory mdp episodes -> prop) (hcountmeasurable : measurableset countbadevent) (hcounttail : source.trajectorymeasure countbadevent <= ennreal.ofreal delta) (hcountgood : forall trajectory, trajectory ∉ countbadevent -> good trajectory /\\ source.successorexpectedaverageregret trajectory rounds <= expectedbound) : let returnbadevent := source.successorreturndeviationbadevent rounds delta let combinedbadevent := countbadevent ∪ returnbadevent measurableset combinedbadevent /\\ source.trajectorym…","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq","label":"exploratorySource_successorExpectedCumulativeRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq","description":"theorem exploratorySource_successorExpectedCumulativeRegret_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat)…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-bf562cd3130f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7086,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:854"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedCumulativeRegret_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).successorExpectedCumulativeRegret trajectory rounds = adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret (initialState := initialState) trajectory defaultState countRadius explorationRate hexplorationRate rounds","missing":[],"search":"exploratorysource_successorexpectedcumulativeregret_eq banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_successorexpectedcumulativeregret_eq theorem exploratorysource_successorexpectedcumulativeregret_eq {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (countradius : transitioncountradius) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : (exploratorysource mdp initialstate episodes initialtable defaultstate countradius explorationrate hexplorationrate).successorexpectedcumulativeregret trajectory rounds = adaptivecumulativeempiricaloptimisticexploratorybehaviorexpectedregret (initialstate := initialstate) trajectory defaultstate countradius explorationrate hexplorationrate rounds theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq","label":"exploratorySource_successorExpectedAverageRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq","description":"theorem exploratorySource_successorExpectedAverageRegret_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) :…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-a9867cbf0d90","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7087,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:873"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedAverageRegret_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : EpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).successorExpectedAverageRegret trajectory rounds = adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret (initialState := initialState) trajectory defaultState countRadius explorationRate hexplorationRate rounds","missing":[],"search":"exploratorysource_successorexpectedaverageregret_eq banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_successorexpectedaverageregret_eq theorem exploratorysource_successorexpectedaverageregret_eq {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (countradius : transitioncountradius) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (trajectory : episodebatchtrajectory mdp episodes) (rounds : nat) : (exploratorysource mdp initialstate episodes initialtable defaultstate countradius explorationrate hexplorationrate).successorexpectedaverageregret trajectory rounds = adaptivecumulativeempiricaloptimisticaverageexploratorybehaviorexpectedregret (initialstate := initialstate) trajectory defaultstate countradius explorationrate hexplorationrate rounds theorem compiled","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorRegret","label":"exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorRegret","description":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (Episod…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativerealizedbehaviorregret/index.html#decl-b7406d6c9012","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","order":7088,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret.lean:890"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpisodeBa…","missing":[],"search":"exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivecumulativeempiricaloptimisticsource.exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorregret theorem exploratorysource_trajectorymeasure_cumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorregret (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (basevisitfloor : real) (n : nat) [standardborelspace (episodebatch mdp (adaptiveepisodebatchsource.decayingexplorationscheduledepisodes mdp basevisitfloor n))] [standardborelspace (episodebatchtrajectory mdp (adaptiveepisodebatchsource.decayingexplorationscheduledepisodes mdp basevisitfloor n))] (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : let rounds := adaptiveepisodebatchsource.decayingexplorationrounds mdp n let delta := adaptiveepisodebatchsource.vanishingaverageconfidencedelta n let explorationrate := adaptiveepisodebatchsourc…","shard":"modules/bece1f714ca529d9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_prod_of_countable_left","label":"measurable_prod_of_countable_left","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_prod_of_countable_left","description":"A function with a countable discrete left coordinate is measurable when every right section is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-bd3a4d3c94ce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7089,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:14"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_prod_of_countable_left {Alpha Beta Gamma : Type*} [MeasurableSpace Alpha] [MeasurableSpace Beta] [MeasurableSpace Gamma] [Countable Alpha] [MeasurableSingletonClass Alpha] (f : Alpha × Beta -> Gamma) (hf : forall a, Measurable (fun b => f (a, b))) : Measurable f","missing":[],"search":"measurable_prod_of_countable_left banditrlproof.measurable_prod_of_countable_left a function with a countable discrete left coordinate is measurable when every right section is measurable. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:MeasureTheory.Measure.compProd_map_dependent","label":"compProd_map_dependent","kind":"lemma","status":"compiled","subtitle":"MeasureTheory.Measure.compProd_map_dependent","description":"Mapping a composition-product by a function which may also read the conditioning coordinate is the composition-product with the correspondingly mapped copy-and-kernel law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-2c1c2f029a5f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7090,"meta":[["Kind","lemma"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"lemma compProd_map_dependent {Alpha : Type u} {Beta : Type v} {Gamma : Type w} [MeasurableSpace Alpha] [MeasurableSpace Beta] [MeasurableSpace Gamma] (mu : Measure Alpha) [SFinite mu] (kappa : ProbabilityTheory.Kernel Alpha Beta) [ProbabilityTheory.IsSFiniteKernel kappa] (f : Alpha × Beta -> Gamma) (hf : Measurable f) : (mu ⊗ₘ kappa).map (fun p => (p.1, f p)) = mu ⊗ₘ ((ProbabilityTheory.Kernel.id ×ₖ kappa).map f)","missing":[],"search":"compprod_map_dependent measuretheory.measure.compprod_map_dependent mapping a composition-product by a function which may also read the conditioning coordinate is the composition-product with the correspondingly mapped copy-and-kernel law. lemma compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_historyBatchStatistic","label":"trajectoryMeasure_condDistrib_historyBatchStatistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_historyBatchStatistic","description":"Conditional law of a statistic which may depend measurably on both the observed prefix and the newly generated batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-f5d47a31ba16","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7091,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_historyBatchStatistic {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (statistic : EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes -> Real) (hstatistic : Measurable statistic) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => statistic (Preorder.frestrictLe n trajectory, trajectory (n + 1))) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] ((ProbabilityTheory.Kernel.id ×ₖ source.batchKernel n).map statistic)","missing":[],"search":"trajectorymeasure_conddistrib_historybatchstatistic banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conddistrib_historybatchstatistic conditional law of a statistic which may depend measurably on both the observed prefix and the newly generated batch. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_historyBatchStatistic_eq_batchKernel","label":"condExpKernel_map_historyBatchStatistic_eq_batchKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_historyBatchStatistic_eq_batchKernel","description":"Trimmed conditional-expectation-kernel version of the preceding dependent statistic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-755c082a909a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7092,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:175"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_historyBatchStatistic_eq_batchKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (statistic : EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes -> Real) (hstatistic : Measurable statistic) : Filter.Eventually (fun trajectory : EpisodeBatchTrajectory mdp episodes => Measure.map (fun path : EpisodeBatchTrajectory mdp episodes => statistic (Preorder.frestrictLe n path, path (n + 1))) (ProbabilityTheory.condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (EpisodeBatchPrefix mdp episodes n)).comap (Preorder.frestrictLe n)) trajectory) = ((ProbabilityTheory.Kernel.id ×ₖ source.batchKernel n).…","missing":[],"search":"condexpkernel_map_historybatchstatistic_eq_batchkernel banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.condexpkernel_map_historybatchstatistic_eq_batchkernel trimmed conditional-expectation-kernel version of the preceding dependent statistic law. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation","label":"bellmanInflation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation","description":"Inflation factor generated by the `z/(32H)` self-bounding term.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-8dc13a5e50ad","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7093,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:222"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellmanInflation (mdp : MDP State Action) : Real","missing":[],"search":"bellmaninflation banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninflation inflation factor generated by the `z/(32h)` self-bounding term. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanWeightCap","label":"bellmanWeightCap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanWeightCap","description":"Uniform cap for every finite-horizon power of `bellmanInflation`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-8ba56b2b5ec4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7094,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:226"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellmanWeightCap : Real","missing":[],"search":"bellmanweightcap banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmanweightcap uniform cap for every finite-horizon power of `bellmaninflation`. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation_one_le","label":"bellmanInflation_one_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation_one_le","description":"theorem bellmanInflation_one_le (mdp : MDP State Action) : 1 <= bellmanInflation mdp","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-06d9b5ce01cb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7095,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:228"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanInflation_one_le (mdp : MDP State Action) : 1 <= bellmanInflation mdp","missing":[],"search":"bellmaninflation_one_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninflation_one_le theorem bellmaninflation_one_le (mdp : mdp state action) : 1 <= bellmaninflation mdp theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation_pow_le_cap","label":"bellmanInflation_pow_le_cap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation_pow_le_cap","description":"theorem bellmanInflation_pow_le_cap (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : bellmanInflation mdp ^ remaining <= bellmanWeightCap","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-0518813617cd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7096,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:235"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanInflation_pow_le_cap (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : bellmanInflation mdp ^ remaining <= bellmanWeightCap","missing":[],"search":"bellmaninflation_pow_le_cap banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninflation_pow_le_cap theorem bellmaninflation_pow_le_cap (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (remaining : nat) (hremaining : remaining <= mdp.horizon) : bellmaninflation mdp ^ remaining <= bellmanweightcap theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanWeight","label":"normalizedBellmanWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanWeight","description":"Normalized chronological recursion weight. Multiplication by the global cap `32/31` later recovers the exact factor `alpha^(stage+1)`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-f32725e8e4b2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7097,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedBellmanWeight (mdp : MDP State Action) (stage : Fin mdp.horizon) : Real","missing":[],"search":"normalizedbellmanweight banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanweight normalized chronological recursion weight. multiplication by the global cap `32/31` later recovers the exact factor `alpha^(stage+1)`. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight","label":"normalizedBellmanChargeWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight","description":"Weight of the local charge at a chronological stage.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-24b8855a512f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7098,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:276"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedBellmanChargeWeight (mdp : MDP State Action) (stage : Fin mdp.horizon) : Real","missing":[],"search":"normalizedbellmanchargeweight banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanchargeweight weight of the local charge at a chronological stage. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight_mem_Icc","label":"normalizedBellmanChargeWeight_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight_mem_Icc","description":"theorem normalizedBellmanChargeWeight_mem_Icc (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (stage : Fin mdp.horizon) : normalizedBellmanChargeWeight mdp stage ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-16feab418cf4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7099,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:280"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedBellmanChargeWeight_mem_Icc (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (stage : Fin mdp.horizon) : normalizedBellmanChargeWeight mdp stage ∈ Set.Icc (0 : Real) 1","missing":[],"search":"normalizedbellmanchargeweight_mem_icc banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanchargeweight_mem_icc theorem normalizedbellmanchargeweight_mem_icc (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (stage : fin mdp.horizon) : normalizedbellmanchargeweight mdp stage ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight_mul_inflation","label":"normalizedBellmanChargeWeight_mul_inflation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight_mul_inflation","description":"theorem normalizedBellmanChargeWeight_mul_inflation (mdp : MDP State Action) (stage : Fin mdp.horizon) : normalizedBellmanChargeWeight mdp stage * bellmanInflation mdp = normalizedBellmanWeight mdp stage","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-c9281103e3b8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7100,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:301"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedBellmanChargeWeight_mul_inflation (mdp : MDP State Action) (stage : Fin mdp.horizon) : normalizedBellmanChargeWeight mdp stage * bellmanInflation mdp = normalizedBellmanWeight mdp stage","missing":[],"search":"normalizedbellmanchargeweight_mul_inflation banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanchargeweight_mul_inflation theorem normalizedbellmanchargeweight_mul_inflation (mdp : mdp state action) (stage : fin mdp.horizon) : normalizedbellmanchargeweight mdp stage * bellmaninflation mdp = normalizedbellmanweight mdp stage theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanWeight_mem_Icc","label":"normalizedBellmanWeight_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanWeight_mem_Icc","description":"theorem normalizedBellmanWeight_mem_Icc (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (stage : Fin mdp.horizon) : normalizedBellmanWeight mdp stage ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-ffb3b6350d8e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7101,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedBellmanWeight_mem_Icc (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (stage : Fin mdp.horizon) : normalizedBellmanWeight mdp stage ∈ Set.Icc (0 : Real) 1","missing":[],"search":"normalizedbellmanweight_mem_icc banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanweight_mem_icc theorem normalizedbellmanweight_mem_icc (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (stage : fin mdp.horizon) : normalizedbellmanweight mdp stage ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedSuccessorGapFeatureOfSummaries","label":"clippedSuccessorGapFeatureOfSummaries","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedSuccessorGapFeatureOfSummaries","description":"The clipped, normalized and chronologically weighted continuation gap used by the predictable martingale. Clipping is inactive on the joint confidence event; normalization keeps the global MGF range exactly `[0,H]`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-dfe6a22a23bf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7102,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedSuccessorGapFeatureOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (summaries : Fin (n + 1) -> TransitionCountSummary mdp) (stage : Fin mdp.horizon) (state : State) : Real","missing":[],"search":"clippedsuccessorgapfeatureofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedsuccessorgapfeatureofsummaries the clipped, normalized and chronologically weighted continuation gap used by the predictable martingale. clipping is inactive on the joint confidence event; normalization keeps the global mgf range exactly `[0,h]`. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedSuccessorGapFeatureOfSummaries_mem_Icc","label":"clippedSuccessorGapFeatureOfSummaries_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedSuccessorGapFeatureOfSummaries_mem_Icc","description":"The martingale feature is globally in `[0,H]`, independently of whether the statistical event holds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-cd78cc38da32","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7103,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:351"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedSuccessorGapFeatureOfSummaries_mem_Icc (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (summaries : Fin (n + 1) -> TransitionCountSummary mdp) (stage : Fin mdp.horizon) (state : State) : clippedSuccessorGapFeatureOfSummaries mdp defaultState episodes delta n summaries stage state ∈ Set.Icc (0 : Real) mdp.horizon","missing":[],"search":"clippedsuccessorgapfeatureofsummaries_mem_icc banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedsuccessorgapfeatureofsummaries_mem_icc the martingale feature is globally in `[0,h]`, independently of whether the statistical event holds. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanStatisticOfSummaries","label":"successorBellmanStatisticOfSummaries","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanStatisticOfSummaries","description":"Compensated batch statistic for one fixed discrete prefix-summary table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-1bb39982ba71","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7104,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:378"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorBellmanStatisticOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (tilt : Real) (input : (Fin (n + 1) -> TransitionCountSummary mdp) × EpisodeBatch mdp 1) : Real","missing":[],"search":"successorbellmanstatisticofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorbellmanstatisticofsummaries compensated batch statistic for one fixed discrete prefix-summary table. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanStatisticOfSummaries","label":"measurable_successorBellmanStatisticOfSummaries","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanStatisticOfSummaries","description":"theorem measurable_successorBellmanStatisticOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (tilt : Real) : Measurable (successorBellmanStatisticOfSummaries mdp defaultState episodes delta n tilt)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-43a1e5556da8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7105,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorBellmanStatisticOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (tilt : Real) : Measurable (successorBellmanStatisticOfSummaries mdp defaultState episodes delta n tilt)","missing":[],"search":"measurable_successorbellmanstatisticofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_successorbellmanstatisticofsummaries theorem measurable_successorbellmanstatisticofsummaries (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (n : nat) (tilt : real) : measurable (successorbellmanstatisticofsummaries mdp defaultstate episodes delta n tilt) theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanHistoryBatchStatistic","label":"successorBellmanHistoryBatchStatistic","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanHistoryBatchStatistic","description":"The same statistic with the prefix summarized exactly as the recurrent planner summarizes it.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-f7e6f933d9d1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7106,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:412"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorBellmanHistoryBatchStatistic (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (tilt : Real) (input : EpisodeBatchPrefix mdp 1 n × EpisodeBatch mdp 1) : Real","missing":[],"search":"successorbellmanhistorybatchstatistic banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorbellmanhistorybatchstatistic the same statistic with the prefix summarized exactly as the recurrent planner summarizes it. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanHistoryBatchStatistic","label":"measurable_successorBellmanHistoryBatchStatistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanHistoryBatchStatistic","description":"theorem measurable_successorBellmanHistoryBatchStatistic (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (tilt : Real) : Measurable (successorBellmanHistoryBatchStatistic mdp defaultState episodes delta n tilt)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-fd7404aa2488","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7107,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:419"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorBellmanHistoryBatchStatistic (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (tilt : Real) : Measurable (successorBellmanHistoryBatchStatistic mdp defaultState episodes delta n tilt)","missing":[],"search":"measurable_successorbellmanhistorybatchstatistic banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_successorbellmanhistorybatchstatistic theorem measurable_successorbellmanhistorybatchstatistic (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (n : nat) (tilt : real) : measurable (successorbellmanhistorybatchstatistic mdp defaultstate episodes delta n tilt) theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessorBellmanInnovation_compensated_hasCondMGFUpperBoundAt","label":"recurrentSuccessorBellmanInnovation_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessorBellmanInnovation_compensated_hasCondMGFUpperBoundAt","description":"A successor episode's compensated, prefix-predictable Bellman innovation has conditional MGF at most one under the literal recurrent batch kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-4955b4f64358","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7108,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:433"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSuccessorBellmanInnovation_compensated_hasCondMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (n : Nat) (tilt : Real) : Concentration.HasCondMGFUpperBoundAt (BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration (mdp := mdp) 1 n) ((BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration (mdp := mdp) 1).le n) (fun trajectory : EpisodeBatchTrajectory mdp 1 => successorBellmanHistoryBatchStatistic mdp defaultState episodes delta n tilt (Preorder.frestrictLe n trajectory, trajectory (n + 1))) 1 0 (recurrentSource mdp initialState defaultState episodes delta |>.trajectoryMeasure)","missing":[],"search":"recurrentsuccessorbellmaninnovation_compensated_hascondmgfupperboundat banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.recurrentsuccessorbellmaninnovation_compensated_hascondmgfupperboundat a successor episode's compensated, prefix-predictable bellman innovation has conditional mgf at most one under the literal recurrent batch kernel. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanInnovationOfSummaries","label":"successorBellmanInnovationOfSummaries","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanInnovationOfSummaries","description":"The un-compensated Bellman innovation of one successor episode, with the policy and clipped gap feature read from one fixed prefix-summary table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-27fc91501648","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7109,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:581"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorBellmanInnovationOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (input : (Fin (n + 1) -> TransitionCountSummary mdp) × EpisodeBatch mdp 1) : Real","missing":[],"search":"successorbellmaninnovationofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorbellmaninnovationofsummaries the un-compensated bellman innovation of one successor episode, with the policy and clipped gap feature read from one fixed prefix-summary table. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanInnovationOfSummaries","label":"measurable_successorBellmanInnovationOfSummaries","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanInnovationOfSummaries","description":"theorem measurable_successorBellmanInnovationOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) : Measurable (successorBellmanInnovationOfSummaries mdp defaultState episodes delta n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-84452b90b253","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7110,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:594"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorBellmanInnovationOfSummaries (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) : Measurable (successorBellmanInnovationOfSummaries mdp defaultState episodes delta n)","missing":[],"search":"measurable_successorbellmaninnovationofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_successorbellmaninnovationofsummaries theorem measurable_successorbellmaninnovationofsummaries (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (n : nat) : measurable (successorbellmaninnovationofsummaries mdp defaultstate episodes delta n) theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanInnovationOfHistoryBatch","label":"successorBellmanInnovationOfHistoryBatch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanInnovationOfHistoryBatch","description":"The successor Bellman innovation with the recurrent prefix summarized in exactly the same way as the generated source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-b017d57b1eed","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7111,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:614"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorBellmanInnovationOfHistoryBatch (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) (input : EpisodeBatchPrefix mdp 1 n × EpisodeBatch mdp 1) : Real","missing":[],"search":"successorbellmaninnovationofhistorybatch banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorbellmaninnovationofhistorybatch the successor bellman innovation with the recurrent prefix summarized in exactly the same way as the generated source. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanInnovationOfHistoryBatch","label":"measurable_successorBellmanInnovationOfHistoryBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanInnovationOfHistoryBatch","description":"theorem measurable_successorBellmanInnovationOfHistoryBatch (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) : Measurable (successorBellmanInnovationOfHistoryBatch mdp defaultState episodes delta n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-65355809939a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7112,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:621"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorBellmanInnovationOfHistoryBatch (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (n : Nat) : Measurable (successorBellmanInnovationOfHistoryBatch mdp defaultState episodes delta n)","missing":[],"search":"measurable_successorbellmaninnovationofhistorybatch banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_successorbellmaninnovationofhistorybatch theorem measurable_successorbellmaninnovationofhistorybatch (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (n : nat) : measurable (successorbellmaninnovationofhistorybatch mdp defaultstate episodes delta n) theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationPrefix","label":"recurrentBellmanInnovationPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationPrefix","description":"Prefix-predictable Bellman innovation. Coordinate zero is deliberately zero: the canonical initial episode is charged separately by its deterministic `H` envelope, while every successor coordinate uses its strict prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-4807e7bd1f61","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7113,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:634"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentBellmanInnovationPrefix (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) : (round : Nat) -> EpisodeBatchPrefix mdp 1 round -> Real | 0, _history => 0 | n + 1, history => successorBellmanInnovationOfHistoryBatch mdp defaultState episodes delta n (Preorder.frestrictLe₂ (π := fun _ : Nat => EpisodeBatch mdp 1) (Nat.le_succ n) history, history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) theorem measurable_recurrentBellmanInnovationPrefix (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (round : Nat) : Measurable (recurrentBellmanInnovationPrefix mdp defaultState episodes delta round)","missing":[],"search":"recurrentbellmaninnovationprefix banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentbellmaninnovationprefix prefix-predictable bellman innovation. coordinate zero is deliberately zero: the canonical initial episode is charged separately by its deterministic `h` envelope, while every successor coordinate uses its strict prefix. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_recurrentBellmanInnovationPrefix","label":"measurable_recurrentBellmanInnovationPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_recurrentBellmanInnovationPrefix","description":"theorem measurable_recurrentBellmanInnovationPrefix (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (round : Nat) : Measurable (recurrentBellmanInnovationPrefix mdp defaultState episodes delta round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-80be6a7f641c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7114,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:646"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_recurrentBellmanInnovationPrefix (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (round : Nat) : Measurable (recurrentBellmanInnovationPrefix mdp defaultState episodes delta round)","missing":[],"search":"measurable_recurrentbellmaninnovationprefix banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_recurrentbellmaninnovationprefix theorem measurable_recurrentbellmaninnovationprefix (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (round : nat) : measurable (recurrentbellmaninnovationprefix mdp defaultstate episodes delta round) theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationProcess","label":"recurrentBellmanInnovationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationProcess","description":"Actual successor-episode Bellman innovation process on the generated recurrent trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-bacc84b13014","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7115,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:666"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentBellmanInnovationProcess (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"recurrentbellmaninnovationprocess banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentbellmaninnovationprocess actual successor-episode bellman innovation process on the generated recurrent trajectory. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationVarianceProcess","label":"recurrentBellmanInnovationVarianceProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationVarianceProcess","description":"Deterministic per-episode variance budget. Coordinate zero is zero because its regret is paid separately; every successor episode has budget `H^3`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-ac18e4d227d5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7116,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:675"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentBellmanInnovationVarianceProcess (mdp : MDP State Action) (round : Nat) (_trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"recurrentbellmaninnovationvarianceprocess banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentbellmaninnovationvarianceprocess deterministic per-episode variance budget. coordinate zero is zero because its regret is paid separately; every successor episode has budget `h^3`. definition compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovation_compensated_stronglyAdapted","label":"recurrentBellmanInnovation_compensated_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovation_compensated_stronglyAdapted","description":"theorem recurrentBellmanInnovation_compensated_stronglyAdapted (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta tilt : Real) : StronglyAdapted (BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration (mdp := mdp) 1) (fun round trajectory => tilt * recurrentBellmanInnovationProcess mdp defaultState episodes delta round trajectory - (tilt ^ 2 / 8) * recurrentBellmanInnovat…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-a3e0c5e6b905","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7117,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:682"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentBellmanInnovation_compensated_stronglyAdapted (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta tilt : Real) : StronglyAdapted (BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration (mdp := mdp) 1) (fun round trajectory => tilt * recurrentBellmanInnovationProcess mdp defaultState episodes delta round trajectory - (tilt ^ 2 / 8) * recurrentBellmanInnovationVarianceProcess mdp round trajectory)","missing":[],"search":"recurrentbellmaninnovation_compensated_stronglyadapted banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentbellmaninnovation_compensated_stronglyadapted theorem recurrentbellmaninnovation_compensated_stronglyadapted (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta tilt : real) : stronglyadapted (banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.batchprefixfiltration (mdp := mdp) 1) (fun round trajectory => tilt * recurrentbellmaninnovationprocess mdp defaultstate episodes delta round trajectory - (tilt ^ 2 / 8) * recurrentbellmaninnovationvarianceprocess mdp round trajectory) theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovation_zero_compensated_hasMGFUpperBoundAt","label":"recurrentBellmanInnovation_zero_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovation_zero_compensated_hasMGFUpperBoundAt","description":"Coordinate zero has the trivial compensated MGF certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-e96f3f95789a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7118,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:717"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentBellmanInnovation_zero_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory : EpisodeBatchTrajectory mdp 1 => tilt * recurrentBellmanInnovationProcess mdp defaultState episodes delta 0 trajectory - (tilt ^ 2 / 8) * recurrentBellmanInnovationVarianceProcess mdp 0 trajectory) 1 0 (recurrentSource mdp initialState defaultState episodes delta |>.trajectoryMeasure)","missing":[],"search":"recurrentbellmaninnovation_zero_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentbellmaninnovation_zero_compensated_hasmgfupperboundat coordinate zero has the trivial compensated mgf certificate. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.trajectoryMeasure_recurrentBellmanInnovation_sum_ge_le","label":"trajectoryMeasure_recurrentBellmanInnovation_sum_ge_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.trajectoryMeasure_recurrentBellmanInnovation_sum_ge_le","description":"Fixed-tilt upper tail for the sum of the generated successor-episode Bellman innovations. The source, prefix policy, and transition law are the same objects as in the recurrent planner.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviadaptivebellmanmartingale/index.html#decl-7d8ebe110ccc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","order":7119,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale.lean:742"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_recurrentBellmanInnovation_sum_ge_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (tilt threshold : Real) (htilt : 0 < tilt) : let source := recurrentSource mdp initialState defaultState episodes delta source.trajectoryMeasure {trajectory | threshold <= (Finset.range episodes).sum (fun round => recurrentBellmanInnovationProcess mdp defaultState episodes delta round trajectory)} <= ENNReal.ofReal (Real.exp (-tilt * threshold + (tilt ^ 2 / 8) * ((episodes : Real) * ((mdp.horizon : Real) * (mdp.horizon : Real) ^ 2))))","missing":[],"search":"trajectorymeasure_recurrentbellmaninnovation_sum_ge_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.trajectorymeasure_recurrentbellmaninnovation_sum_ge_le fixed-tilt upper tail for the sum of the generated successor-episode bellman innovations. the source, prefix policy, and transition law are the same objects as in the recurrent planner. theorem compiled","shard":"modules/2489d370681ce83c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateTransitionCount","label":"aggregateTransitionCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateTransitionCount","description":"The UCBVI-CH numerator `N_k(x,a,y)`, pooled across every decision stage.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-9247f190b375","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7120,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def aggregateTransitionCount {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (state : State) (action : Action) (nextState : State) : Nat","missing":[],"search":"aggregatetransitioncount banditrlproof.finitehorizonrl.transitioncountsummary.aggregatetransitioncount the ucbvi-ch numerator `n_k(x,a,y)`, pooled across every decision stage. definition compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.measurable_aggregateTransitionCount","label":"measurable_aggregateTransitionCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.measurable_aggregateTransitionCount","description":"The complete pooled transition-count table is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-165dc0215059","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7121,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionCount {mdp : MDP State Action} : Measurable fun summary : TransitionCountSummary mdp => fun state action nextState => summary.aggregateTransitionCount state action nextState","missing":[],"search":"measurable_aggregatetransitioncount banditrlproof.finitehorizonrl.transitioncountsummary.measurable_aggregatetransitioncount the complete pooled transition-count table is measurable. theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.sum_aggregateTransitionCount_eq_aggregateVisitCount","label":"sum_aggregateTransitionCount_eq_aggregateVisitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.sum_aggregateTransitionCount_eq_aggregateVisitCount","description":"Pooling transition destinations gives exactly the pooled visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-f7ea56ac1c15","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7122,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_aggregateTransitionCount_eq_aggregateVisitCount {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (state : State) (action : Action) : (∑ nextState : State, summary.aggregateTransitionCount state action nextState) = summary.aggregateVisitCount state action","missing":[],"search":"sum_aggregatetransitioncount_eq_aggregatevisitcount banditrlproof.finitehorizonrl.transitioncountsummary.sum_aggregatetransitioncount_eq_aggregatevisitcount pooling transition destinations gives exactly the pooled visit count. theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF","label":"aggregateEmpiricalTransitionPMF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF","description":"The normalized pooled empirical row, with a Dirac fallback at zero visits.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-a003a1227747","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7123,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateEmpiricalTransitionPMF {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (state : State) (action : Action) : PMF State","missing":[],"search":"aggregateempiricaltransitionpmf banditrlproof.finitehorizonrl.transitioncountsummary.aggregateempiricaltransitionpmf the normalized pooled empirical row, with a dirac fallback at zero visits. definition compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF_of_aggregateVisitCount_eq_zero","label":"aggregateEmpiricalTransitionPMF_of_aggregateVisitCount_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF_of_aggregateVisitCount_eq_zero","description":"theorem aggregateEmpiricalTransitionPMF_of_aggregateVisitCount_eq_zero {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (state : State) (action : Action) (hzero : summary.aggregateVisitCount state action = 0) : summary.aggregateEmpiricalTransitionPMF defaultState state action = PMF.pure defaultState","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-3e8a47c122b9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7124,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateEmpiricalTransitionPMF_of_aggregateVisitCount_eq_zero {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (state : State) (action : Action) (hzero : summary.aggregateVisitCount state action = 0) : summary.aggregateEmpiricalTransitionPMF defaultState state action = PMF.pure defaultState","missing":[],"search":"aggregateempiricaltransitionpmf_of_aggregatevisitcount_eq_zero banditrlproof.finitehorizonrl.transitioncountsummary.aggregateempiricaltransitionpmf_of_aggregatevisitcount_eq_zero theorem aggregateempiricaltransitionpmf_of_aggregatevisitcount_eq_zero {mdp : mdp state action} (summary : transitioncountsummary mdp) (defaultstate : state) (state : state) (action : action) (hzero : summary.aggregatevisitcount state action = 0) : summary.aggregateempiricaltransitionpmf defaultstate state action = pmf.pure defaultstate theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF_apply_of_aggregateVisitCount_pos","label":"aggregateEmpiricalTransitionPMF_apply_of_aggregateVisitCount_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF_apply_of_aggregateVisitCount_pos","description":"theorem aggregateEmpiricalTransitionPMF_apply_of_aggregateVisitCount_pos {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (state : State) (action : Action) (nextState : State) (hpos : 0 < summary.aggregateVisitCount state action) : summary.aggregateEmpiricalTransitionPMF defaultState state action nextState = (summary.aggregateTransitionCount state action nextState : ENNReal) / (…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-9dc881a7508d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7125,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateEmpiricalTransitionPMF_apply_of_aggregateVisitCount_pos {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (state : State) (action : Action) (nextState : State) (hpos : 0 < summary.aggregateVisitCount state action) : summary.aggregateEmpiricalTransitionPMF defaultState state action nextState = (summary.aggregateTransitionCount state action nextState : ENNReal) / (summary.aggregateVisitCount state action : ENNReal)","missing":[],"search":"aggregateempiricaltransitionpmf_apply_of_aggregatevisitcount_pos banditrlproof.finitehorizonrl.transitioncountsummary.aggregateempiricaltransitionpmf_apply_of_aggregatevisitcount_pos theorem aggregateempiricaltransitionpmf_apply_of_aggregatevisitcount_pos {mdp : mdp state action} (summary : transitioncountsummary mdp) (defaultstate : state) (state : state) (action : action) (nextstate : state) (hpos : 0 < summary.aggregatevisitcount state action) : summary.aggregateempiricaltransitionpmf defaultstate state action nextstate = (summary.aggregatetransitioncount state action nextstate : ennreal) / (summary.aggregatevisitcount state action : ennreal) theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel","label":"aggregateEmpiricalTransitionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel","description":"The pooled empirical transition kernel is stage-homogeneous.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-380fce1795d4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7126,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:121"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateEmpiricalTransitionKernel {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) : ProbabilityTheory.Kernel (State × Action) State","missing":[],"search":"aggregateempiricaltransitionkernel banditrlproof.finitehorizonrl.transitioncountsummary.aggregateempiricaltransitionkernel the pooled empirical transition kernel is stage-homogeneous. definition compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel_isMarkov","label":"aggregateEmpiricalTransitionKernel_isMarkov","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel_isMarkov","description":"theorem aggregateEmpiricalTransitionKernel_isMarkov {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) : ProbabilityTheory.IsMarkovKernel (summary.aggregateEmpiricalTransitionKernel defaultState) where","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-3886513c114f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7127,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:128"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateEmpiricalTransitionKernel_isMarkov {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) : ProbabilityTheory.IsMarkovKernel (summary.aggregateEmpiricalTransitionKernel defaultState) where","missing":[],"search":"aggregateempiricaltransitionkernel_ismarkov banditrlproof.finitehorizonrl.transitioncountsummary.aggregateempiricaltransitionkernel_ismarkov theorem aggregateempiricaltransitionkernel_ismarkov {mdp : mdp state action} (summary : transitioncountsummary mdp) (defaultstate : state) : probabilitytheory.ismarkovkernel (summary.aggregateempiricaltransitionkernel defaultstate) where theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt","label":"adaptiveCumulativeAggregateTransitionCountAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt","description":"The pooled transition numerator through one generated trajectory prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-39ade56afc36","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7128,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def adaptiveCumulativeAggregateTransitionCountAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (state : State) (action : Action) (nextState : State) : Nat","missing":[],"search":"adaptivecumulativeaggregatetransitioncountat banditrlproof.finitehorizonrl.adaptivecumulativeaggregatetransitioncountat the pooled transition numerator through one generated trajectory prefix. definition compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeAggregateTransitionCountAt","label":"measurable_adaptiveCumulativeAggregateTransitionCountAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeAggregateTransitionCountAt","description":"theorem measurable_adaptiveCumulativeAggregateTransitionCountAt {mdp : MDP State Action} {episodes : Nat} (round : Nat) (state : State) (action : Action) (nextState : State) : Measurable fun trajectory : EpisodeBatchTrajectory mdp episodes => adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-b277d81b6332","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7129,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:150"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_adaptiveCumulativeAggregateTransitionCountAt {mdp : MDP State Action} {episodes : Nat} (round : Nat) (state : State) (action : Action) (nextState : State) : Measurable fun trajectory : EpisodeBatchTrajectory mdp episodes => adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState","missing":[],"search":"measurable_adaptivecumulativeaggregatetransitioncountat banditrlproof.finitehorizonrl.measurable_adaptivecumulativeaggregatetransitioncountat theorem measurable_adaptivecumulativeaggregatetransitioncountat {mdp : mdp state action} {episodes : nat} (round : nat) (state : state) (action : action) (nextstate : state) : measurable fun trajectory : episodebatchtrajectory mdp episodes => adaptivecumulativeaggregatetransitioncountat trajectory round state action nextstate theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_eq_sum","label":"adaptiveCumulativeAggregateTransitionCountAt_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_eq_sum","description":"The pooled numerator is the literal episode-by-stage generated count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-2922b36f6902","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7130,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:166"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeAggregateTransitionCountAt_eq_sum {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (state : State) (action : Action) (nextState : State) : adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState = ∑ i : Fin (round + 1), ∑ stage : Fin mdp.horizon, (trajectory i).transitionCount stage state action nextState","missing":[],"search":"adaptivecumulativeaggregatetransitioncountat_eq_sum banditrlproof.finitehorizonrl.adaptivecumulativeaggregatetransitioncountat_eq_sum the pooled numerator is the literal episode-by-stage generated count. theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_succ","label":"adaptiveCumulativeAggregateTransitionCountAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_succ","description":"Extending the prefix adds exactly one episode's pooled transition row.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-1e04f69b15f3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7131,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:186"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeAggregateTransitionCountAt_succ {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (state : State) (action : Action) (nextState : State) : adaptiveCumulativeAggregateTransitionCountAt trajectory (round + 1) state action nextState = adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState + ∑ stage : Fin mdp.horizon, (trajectory (round + 1)).transitionCount stage state action nextState","missing":[],"search":"adaptivecumulativeaggregatetransitioncountat_succ banditrlproof.finitehorizonrl.adaptivecumulativeaggregatetransitioncountat_succ extending the prefix adds exactly one episode's pooled transition row. theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","label":"sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","description":"At every generated prefix the pooled numerator row sums to the denominator.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviaggregatetransition/index.html#decl-c5aed9a4bdb0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","order":7132,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition.lean:205"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (round : Nat) (state : State) (action : Action) : (∑ nextState : State, adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState) = adaptiveCumulativeAggregateVisitCountAt trajectory round state action","missing":[],"search":"sum_adaptivecumulativeaggregatetransitioncountat_eq_visitcountat banditrlproof.finitehorizonrl.sum_adaptivecumulativeaggregatetransitioncountat_eq_visitcountat at every generated prefix the pooled numerator row sums to the denominator. theorem compiled","shard":"modules/92a9beb680429b6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryStateAt","label":"measurable_trajectoryStateAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryStateAt","description":"theorem measurable_trajectoryStateAt (mdp : MDP State Action) (stage : Fin mdp.horizon) : Measurable (fun trajectory : State × StepTrace Action State mdp.horizon => mdp.trajectoryStateAt trajectory stage)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvialignment/index.html#decl-cae680952a39","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","order":7133,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAlignment.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_trajectoryStateAt (mdp : MDP State Action) (stage : Fin mdp.horizon) : Measurable (fun trajectory : State × StepTrace Action State mdp.horizon => mdp.trajectoryStateAt trajectory stage)","missing":[],"search":"measurable_trajectorystateat banditrlproof.finitehorizonrl.mdp.measurable_trajectorystateat theorem measurable_trajectorystateat (mdp : mdp state action) (stage : fin mdp.horizon) : measurable (fun trajectory : state × steptrace action state mdp.horizon => mdp.trajectorystateat trajectory stage) theorem compiled","shard":"modules/c24ee449f56274f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.trajectoryMeasure_action_eq_table_ae","label":"trajectoryMeasure_action_eq_table_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.trajectoryMeasure_action_eq_table_ae","description":"A trajectory generated by a deterministic table records exactly that table's action at every chronological stage, almost surely.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvialignment/index.html#decl-5e0397c7eb47","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","order":7134,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAlignment.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_action_eq_table_ae {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : ∀ᵐ trajectory ∂table.toMarkovPolicy.trajectoryMeasure initialState, ∀ stage : Fin mdp.horizon, (mdp.episodeStepOfTrajectory trajectory stage).action = table stage (mdp.trajectoryStateAt trajectory stage)","missing":[],"search":"trajectorymeasure_action_eq_table_ae banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.trajectorymeasure_action_eq_table_ae a trajectory generated by a deterministic table records exactly that table's action at every chronological stage, almost surely. theorem compiled","shard":"modules/c24ee449f56274f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.iidEpisodeBatchMeasure_successorBatchAligned_ae","label":"iidEpisodeBatchMeasure_successorBatchAligned_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.iidEpisodeBatchMeasure_successorBatchAligned_ae","description":"Under the one-episode mapped batch law, every reconstructed state/action record coincides with the genuine generated trajectory and deterministic table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvialignment/index.html#decl-6154911cdda2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","order":7135,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAlignment.lean:102"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_successorBatchAligned_ae (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (table : DeterministicMarkovPolicyTable mdp) (hhorizon : 0 < mdp.horizon) : ∀ᵐ batch ∂table.toMarkovPolicy.iidEpisodeBatchMeasure initialState 1, let reconstructed : State × StepTrace Action State mdp.horizon := (batch.reconstructedInitialState defaultState, batch.reconstructedStepTrace) ∀ stage : Fin mdp.horizon, (batch 0 stage).state = mdp.trajectoryStateAt reconstructed stage ∧ (batch 0 stage).action = table stage (mdp.trajectoryStateAt reconstructed stage)","missing":[],"search":"iidepisodebatchmeasure_successorbatchaligned_ae banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.iidepisodebatchmeasure_successorbatchaligned_ae under the one-episode mapped batch law, every reconstructed state/action record coincides with the genuine generated trajectory and deterministic table. theorem compiled","shard":"modules/c24ee449f56274f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_successorBatchAligned_ae","label":"recurrentSource_trajectoryMeasure_successorBatchAligned_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_successorBatchAligned_ae","description":"Every successor coordinate of the actual recurrent `Kernel.trajMeasure` is aligned with the deterministic table computed from that same trajectory's strict prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvialignment/index.html#decl-feb08581a9eb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","order":7136,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAlignment.lean:195"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_trajectoryMeasure_successorBatchAligned_ae (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (n : Nat) : ∀ᵐ trajectory ∂ (recurrentSource mdp initialState defaultState episodes delta).trajectoryMeasure, SuccessorBatchAligned mdp defaultState episodes delta trajectory n","missing":[],"search":"recurrentsource_trajectorymeasure_successorbatchaligned_ae banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_successorbatchaligned_ae every successor coordinate of the actual recurrent `kernel.trajmeasure` is aligned with the deterministic table computed from that same trajectory's strict prefix. theorem compiled","shard":"modules/c24ee449f56274f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_all_successorBatchAligned_ae","label":"recurrentSource_trajectoryMeasure_all_successorBatchAligned_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_all_successorBatchAligned_ae","description":"Countable conjunction of the same-source alignment certificates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvialignment/index.html#decl-3ca6af11cc4f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","order":7137,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIAlignment.lean:310"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_trajectoryMeasure_all_successorBatchAligned_ae (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) : ∀ᵐ trajectory ∂ (recurrentSource mdp initialState defaultState episodes delta).trajectoryMeasure, ∀ n : Nat, SuccessorBatchAligned mdp defaultState episodes delta trajectory n","missing":[],"search":"recurrentsource_trajectorymeasure_all_successorbatchaligned_ae banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_all_successorbatchaligned_ae countable conjunction of the same-source alignment certificates. theorem compiled","shard":"modules/c24ee449f56274f6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeDeterministicGapInnovationFrom","label":"sampledCumulativeDeterministicGapInnovationFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeDeterministicGapInnovationFrom","description":"Chronological Bellman innovation for a deterministic action table. The sampled action coordinate is deliberately ignored: the table chooses the action from the recursively reconstructed current state. This makes the pathwise recursion canonical on every batch, while its law is still the exact generated deterministic-policy trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-e831b9b6c802","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7138,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeDeterministicGapInnovationFrom (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> StepTrace Action State remaining -> Real | 0, _, _, _ => 0 | remaining + 1, hremaining, state, trace => let stage : Fin mdp.horizon := ⟨mdp.horizon - (remaining + 1), by omega⟩ mdp.transitionValue (feature stage) state (table stage state) - feature stage (trace 0).2 + sampledCumulativeDeterministicGapInnovationFrom mdp table feature remaining (by omega) (trace 0).2 (Fin.tail trace) omit [Nonempty State] [Nonempty Action] in theorem measurable_sampledCumulativeDeterministicGapInnovationFrom (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining <= mdp…","missing":[],"search":"sampledcumulativedeterministicgapinnovationfrom banditrlproof.finitehorizonrl.mdp.sampledcumulativedeterministicgapinnovationfrom chronological bellman innovation for a deterministic action table. the sampled action coordinate is deliberately ignored: the table chooses the action from the recursively reconstructed current state. this makes the pathwise recursion canonical on every batch, while its law is still the exact generated deterministic-policy trajectory law. definition compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeDeterministicGapInnovationFrom","label":"measurable_sampledCumulativeDeterministicGapInnovationFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeDeterministicGapInnovationFrom","description":"theorem measurable_sampledCumulativeDeterministicGapInnovationFrom (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (fun p : State × StepTrace Action State remaining => mdp.sampledCumulativeDeterministicGapInnovationFrom table feature remaining hremaining p.1 p.2)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-a2e4b4e4ece9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7139,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeDeterministicGapInnovationFrom (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (fun p : State × StepTrace Action State remaining => mdp.sampledCumulativeDeterministicGapInnovationFrom table feature remaining hremaining p.1 p.2)","missing":[],"search":"measurable_sampledcumulativedeterministicgapinnovationfrom banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativedeterministicgapinnovationfrom theorem measurable_sampledcumulativedeterministicgapinnovationfrom (mdp : mdp state action) (table : deterministicmarkovpolicytable mdp) (feature : fin mdp.horizon -> state -> real) (remaining : nat) (hremaining : remaining <= mdp.horizon) : measurable (fun p : state × steptrace action state remaining => mdp.sampledcumulativedeterministicgapinnovationfrom table feature remaining hremaining p.1 p.2) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","label":"actionStateKernel_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","description":"theorem actionStateKernel_deterministicGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : State -> Real) (hfeature : forall nextState, feature nextState ∈ Set.Icc (0 : Real) mdp.horizon) (stage : Fin mdp.horizon) (currentState : State) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * (mdp.transitionValue f…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-b69bb36f271e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7140,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_deterministicGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : State -> Real) (hfeature : forall nextState, feature nextState ∈ Set.Icc (0 : Real) mdp.horizon) (stage : Fin mdp.horizon) (currentState : State) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * (mdp.transitionValue feature currentState (table stage currentState) - feature head.2) - tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (table.toMarkovPolicy.actionStateKernel stage currentState)","missing":[],"search":"actionstatekernel_deterministicgapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.actionstatekernel_deterministicgapinnovation_compensated_hasmgfupperboundat theorem actionstatekernel_deterministicgapinnovation_compensated_hasmgfupperboundat (mdp : mdp state action) (table : deterministicmarkovpolicytable mdp) (feature : state -> real) (hfeature : forall nextstate, feature nextstate ∈ set.icc (0 : real) mdp.horizon) (stage : fin mdp.horizon) (currentstate : state) (tilt : real) : concentration.hasmgfupperboundat (fun head : action × state => tilt * (mdp.transitionvalue feature currentstate (table stage currentstate) - feature head.2) - tilt ^ 2 * (mdp.horizon : real) ^ 2 / 8) 1 0 (table.tomarkovpolicy.actionstatekernel stage currentstate) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","label":"trajectoryKernelRemaining_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","description":"Stage-resolved MGF for the canonicalized deterministic trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-c67a22a9fadd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7141,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_deterministicGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) (tilt : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (currentState : State) : Concentration.HasMGFUpperBoundAt (fun trace => tilt * mdp.sampledCumulativeDeterministicGapInnovationFrom table feature remaining hremaining currentState trace - (remaining : Real) * tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (table.toMarkovPolicy.trajectoryKernelRemaining remaining hremaining currentState)","missing":[],"search":"trajectorykernelremaining_deterministicgapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorykernelremaining_deterministicgapinnovation_compensated_hasmgfupperboundat stage-resolved mgf for the canonicalized deterministic trajectory. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","label":"trajectoryMeasure_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","description":"Integrating the start state preserves the canonical deterministic-table stage-resolved MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-2e7b1f3f0e99","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7142,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_deterministicGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory => tilt * mdp.sampledCumulativeDeterministicGapInnovationFrom table feature mdp.horizon le_rfl trajectory.1 trajectory.2 - (mdp.horizon : Real) * tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (table.toMarkovPolicy.trajectoryMeasure initialState)","missing":[],"search":"trajectorymeasure_deterministicgapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorymeasure_deterministicgapinnovation_compensated_hasmgfupperboundat integrating the start state preserves the canonical deterministic-table stage-resolved mgf. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.abs_sampledCumulativeDeterministicGapInnovationFrom_le","label":"abs_sampledCumulativeDeterministicGapInnovationFrom_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.abs_sampledCumulativeDeterministicGapInnovationFrom_le","description":"Pathwise envelope for the canonical deterministic-table innovation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-d8068f81dace","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7143,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:326"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_sampledCumulativeDeterministicGapInnovationFrom_le (mdp : MDP State Action) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) : forall remaining (hremaining : remaining <= mdp.horizon) (currentState : State) (trace : StepTrace Action State remaining), |mdp.sampledCumulativeDeterministicGapInnovationFrom table feature remaining hremaining currentState trace| <= (remaining : Real) * mdp.horizon","missing":[],"search":"abs_sampledcumulativedeterministicgapinnovationfrom_le banditrlproof.finitehorizonrl.mdp.abs_sampledcumulativedeterministicgapinnovationfrom_le pathwise envelope for the canonical deterministic-table innovation. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeTransitionGapInnovationFrom","label":"sampledCumulativeTransitionGapInnovationFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeTransitionGapInnovationFrom","description":"Sum of chronological transition innovations for a stage-indexed bounded continuation table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-da921ba7a3d7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7144,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:385"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeTransitionGapInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> StepTrace Action State remaining -> Real | 0, _, _, _ => 0 | remaining + 1, hremaining, state, trace => let stage : Fin mdp.horizon := ⟨mdp.horizon - (remaining + 1), by omega⟩ mdp.transitionValue (feature stage) state (trace 0).1 - feature stage (trace 0).2 + sampledCumulativeTransitionGapInnovationFrom mdp policy feature remaining (by omega) (trace 0).2 (Fin.tail trace) omit [Nonempty State] [Nonempty Action] in theorem measurable_sampledCumulativeTransitionGapInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (fun p : State × StepTrace…","missing":[],"search":"sampledcumulativetransitiongapinnovationfrom banditrlproof.finitehorizonrl.mdp.sampledcumulativetransitiongapinnovationfrom sum of chronological transition innovations for a stage-indexed bounded continuation table. definition compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeTransitionGapInnovationFrom","label":"measurable_sampledCumulativeTransitionGapInnovationFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeTransitionGapInnovationFrom","description":"theorem measurable_sampledCumulativeTransitionGapInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (fun p : State × StepTrace Action State remaining => mdp.sampledCumulativeTransitionGapInnovationFrom policy feature remaining hremaining p.1 p.2)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-9ef3d60e1895","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7145,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:400"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeTransitionGapInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (fun p : State × StepTrace Action State remaining => mdp.sampledCumulativeTransitionGapInnovationFrom policy feature remaining hremaining p.1 p.2)","missing":[],"search":"measurable_sampledcumulativetransitiongapinnovationfrom banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativetransitiongapinnovationfrom theorem measurable_sampledcumulativetransitiongapinnovationfrom (mdp : mdp state action) (policy : markovpolicy mdp) (feature : fin mdp.horizon -> state -> real) (remaining : nat) (hremaining : remaining <= mdp.horizon) : measurable (fun p : state × steptrace action state remaining => mdp.sampledcumulativetransitiongapinnovationfrom policy feature remaining hremaining p.1 p.2) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.abs_sampledCumulativeTransitionGapInnovationFrom_le","label":"abs_sampledCumulativeTransitionGapInnovationFrom_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.abs_sampledCumulativeTransitionGapInnovationFrom_le","description":"Pathwise envelope matching the sum of `remaining` centered `[0,H]` transition probes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-e245687e6b36","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7146,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:432"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_sampledCumulativeTransitionGapInnovationFrom_le (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) : forall remaining (hremaining : remaining <= mdp.horizon) (currentState : State) (trace : StepTrace Action State remaining), |mdp.sampledCumulativeTransitionGapInnovationFrom policy feature remaining hremaining currentState trace| <= (remaining : Real) * mdp.horizon","missing":[],"search":"abs_sampledcumulativetransitiongapinnovationfrom_le banditrlproof.finitehorizonrl.mdp.abs_sampledcumulativetransitiongapinnovationfrom_le pathwise envelope matching the sum of `remaining` centered `[0,h]` transition probes. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionGapInnovation_compensated_hasMGFUpperBoundAt","label":"actionStateKernel_transitionGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionGapInnovation_compensated_hasMGFUpperBoundAt","description":"One policy transition has the Hoeffding compensated MGF with range `H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-7ad04d6ce328","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7147,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:491"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_transitionGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : State -> Real) (hfeature : forall nextState, feature nextState ∈ Set.Icc (0 : Real) mdp.horizon) (stage : Fin mdp.horizon) (currentState : State) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * (mdp.transitionValue feature currentState head.1 - feature head.2) - tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (policy.actionStateKernel stage currentState)","missing":[],"search":"actionstatekernel_transitiongapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.actionstatekernel_transitiongapinnovation_compensated_hasmgfupperboundat one policy transition has the hoeffding compensated mgf with range `h`. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionGapInnovation_compensated_hasMGFUpperBoundAt","label":"trajectoryKernelRemaining_transitionGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionGapInnovation_compensated_hasMGFUpperBoundAt","description":"Recursive within-episode MGF. Each of the `remaining` transitions pays exactly one `tilt^2 H^2 / 8` compensation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-2f8730971142","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7148,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:581"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_transitionGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) (tilt : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (currentState : State) : Concentration.HasMGFUpperBoundAt (fun trace => tilt * mdp.sampledCumulativeTransitionGapInnovationFrom policy feature remaining hremaining currentState trace - (remaining : Real) * tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (policy.trajectoryKernelRemaining remaining hremaining currentState)","missing":[],"search":"trajectorykernelremaining_transitiongapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorykernelremaining_transitiongapinnovation_compensated_hasmgfupperboundat recursive within-episode mgf. each of the `remaining` transitions pays exactly one `tilt^2 h^2 / 8` compensation. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionGapInnovation_compensated_hasMGFUpperBoundAt","label":"trajectoryMeasure_transitionGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionGapInnovation_compensated_hasMGFUpperBoundAt","description":"Integrating the start state preserves the exact stage-resolved MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-3e0c81aad988","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7149,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:715"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_transitionGapInnovation_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory => tilt * mdp.sampledCumulativeTransitionGapInnovationFrom policy feature mdp.horizon le_rfl trajectory.1 trajectory.2 - (mdp.horizon : Real) * tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (policy.trajectoryMeasure initialState)","missing":[],"search":"trajectorymeasure_transitiongapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorymeasure_transitiongapinnovation_compensated_hasmgfupperboundat integrating the start state preserves the exact stage-resolved mgf. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.reconstructedInitialState","label":"reconstructedInitialState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.reconstructedInitialState","description":"Initial state reconstructed from the first record of a one-episode batch. The default branch specifies the degenerate zero-horizon input only.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-57b0babe2d4e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7150,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:775"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def reconstructedInitialState {mdp : MDP State Action} (defaultState : State) (batch : EpisodeBatch mdp 1) : State","missing":[],"search":"reconstructedinitialstate banditrlproof.finitehorizonrl.episodebatch.reconstructedinitialstate initial state reconstructed from the first record of a one-episode batch. the default branch specifies the degenerate zero-horizon input only. definition compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.reconstructedStepTrace","label":"reconstructedStepTrace","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.reconstructedStepTrace","description":"Action/next-state trace reconstructed from the one-episode records.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-64f7a563054e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7151,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:782"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def reconstructedStepTrace {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) : StepTrace Action State mdp.horizon","missing":[],"search":"reconstructedsteptrace banditrlproof.finitehorizonrl.episodebatch.reconstructedsteptrace action/next-state trace reconstructed from the one-episode records. definition compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_reconstructedInitialState","label":"measurable_reconstructedInitialState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_reconstructedInitialState","description":"theorem measurable_reconstructedInitialState {mdp : MDP State Action} (defaultState : State) : Measurable (reconstructedInitialState (mdp := mdp) defaultState)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-4366eaa6a9fc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7152,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:788"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_reconstructedInitialState {mdp : MDP State Action} (defaultState : State) : Measurable (reconstructedInitialState (mdp := mdp) defaultState)","missing":[],"search":"measurable_reconstructedinitialstate banditrlproof.finitehorizonrl.episodebatch.measurable_reconstructedinitialstate theorem measurable_reconstructedinitialstate {mdp : mdp state action} (defaultstate : state) : measurable (reconstructedinitialstate (mdp := mdp) defaultstate) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_reconstructedStepTrace","label":"measurable_reconstructedStepTrace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_reconstructedStepTrace","description":"theorem measurable_reconstructedStepTrace {mdp : MDP State Action} : Measurable (reconstructedStepTrace (mdp := mdp))","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-3aba78d72802","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7153,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:798"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_reconstructedStepTrace {mdp : MDP State Action} : Measurable (reconstructedStepTrace (mdp := mdp))","missing":[],"search":"measurable_reconstructedsteptrace banditrlproof.finitehorizonrl.episodebatch.measurable_reconstructedsteptrace theorem measurable_reconstructedsteptrace {mdp : mdp state action} : measurable (reconstructedsteptrace (mdp := mdp)) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeDeterministicGapInnovation","label":"cumulativeDeterministicGapInnovation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeDeterministicGapInnovation","description":"Canonical deterministic-table Bellman innovation read from one generated batch. Recorded action fields are intentionally not used.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-6d2b46286742","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7154,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:810"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeDeterministicGapInnovation {mdp : MDP State Action} (defaultState : State) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (batch : EpisodeBatch mdp 1) : Real","missing":[],"search":"cumulativedeterministicgapinnovation banditrlproof.finitehorizonrl.episodebatch.cumulativedeterministicgapinnovation canonical deterministic-table bellman innovation read from one generated batch. recorded action fields are intentionally not used. definition compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_cumulativeDeterministicGapInnovation","label":"measurable_cumulativeDeterministicGapInnovation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_cumulativeDeterministicGapInnovation","description":"theorem measurable_cumulativeDeterministicGapInnovation {mdp : MDP State Action} (defaultState : State) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.cumulativeDeterministicGapInnovation defaultState table feature)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-10381e1a7960","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7155,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:820"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeDeterministicGapInnovation {mdp : MDP State Action} (defaultState : State) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.cumulativeDeterministicGapInnovation defaultState table feature)","missing":[],"search":"measurable_cumulativedeterministicgapinnovation banditrlproof.finitehorizonrl.episodebatch.measurable_cumulativedeterministicgapinnovation theorem measurable_cumulativedeterministicgapinnovation {mdp : mdp state action} (defaultstate : state) (table : deterministicmarkovpolicytable mdp) (feature : fin mdp.horizon -> state -> real) : measurable (fun batch : episodebatch mdp 1 => batch.cumulativedeterministicgapinnovation defaultstate table feature) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeDeterministicGapInnovation_episodeBatchOfTrajectories","label":"cumulativeDeterministicGapInnovation_episodeBatchOfTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeDeterministicGapInnovation_episodeBatchOfTrajectories","description":"theorem cumulativeDeterministicGapInnovation_episodeBatchOfTrajectories (mdp : MDP State Action) (defaultState : State) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (hhorizon : 0 < mdp.horizon) (trajectories : Fin 1 -> State × StepTrace Action State mdp.horizon) : EpisodeBatch.cumulativeDeterministicGapInnovation defaultState table feature (mdp.episodeBatchOfTrajectories…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-1c450051e973","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7156,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:833"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeDeterministicGapInnovation_episodeBatchOfTrajectories (mdp : MDP State Action) (defaultState : State) (table : DeterministicMarkovPolicyTable mdp) (feature : Fin mdp.horizon -> State -> Real) (hhorizon : 0 < mdp.horizon) (trajectories : Fin 1 -> State × StepTrace Action State mdp.horizon) : EpisodeBatch.cumulativeDeterministicGapInnovation defaultState table feature (mdp.episodeBatchOfTrajectories 1 trajectories) = mdp.sampledCumulativeDeterministicGapInnovationFrom table feature mdp.horizon le_rfl (trajectories 0).1 (trajectories 0).2","missing":[],"search":"cumulativedeterministicgapinnovation_episodebatchoftrajectories banditrlproof.finitehorizonrl.episodebatch.cumulativedeterministicgapinnovation_episodebatchoftrajectories theorem cumulativedeterministicgapinnovation_episodebatchoftrajectories (mdp : mdp state action) (defaultstate : state) (table : deterministicmarkovpolicytable mdp) (feature : fin mdp.horizon -> state -> real) (hhorizon : 0 < mdp.horizon) (trajectories : fin 1 -> state × steptrace action state mdp.horizon) : episodebatch.cumulativedeterministicgapinnovation defaultstate table feature (mdp.episodebatchoftrajectories 1 trajectories) = mdp.sampledcumulativedeterministicgapinnovationfrom table feature mdp.horizon le_rfl (trajectories 0).1 (trajectories 0).2 theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeTransitionGapInnovation","label":"cumulativeTransitionGapInnovation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeTransitionGapInnovation","description":"Stage-resolved Bellman innovation read from one generated batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-bfdc13a51c80","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7157,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:859"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeTransitionGapInnovation {mdp : MDP State Action} (defaultState : State) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (batch : EpisodeBatch mdp 1) : Real","missing":[],"search":"cumulativetransitiongapinnovation banditrlproof.finitehorizonrl.episodebatch.cumulativetransitiongapinnovation stage-resolved bellman innovation read from one generated batch. definition compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_cumulativeTransitionGapInnovation","label":"measurable_cumulativeTransitionGapInnovation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_cumulativeTransitionGapInnovation","description":"theorem measurable_cumulativeTransitionGapInnovation {mdp : MDP State Action} (defaultState : State) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.cumulativeTransitionGapInnovation defaultState policy feature)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-e3a63ff78f8d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7158,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:869"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeTransitionGapInnovation {mdp : MDP State Action} (defaultState : State) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.cumulativeTransitionGapInnovation defaultState policy feature)","missing":[],"search":"measurable_cumulativetransitiongapinnovation banditrlproof.finitehorizonrl.episodebatch.measurable_cumulativetransitiongapinnovation theorem measurable_cumulativetransitiongapinnovation {mdp : mdp state action} (defaultstate : state) (policy : markovpolicy mdp) (feature : fin mdp.horizon -> state -> real) : measurable (fun batch : episodebatch mdp 1 => batch.cumulativetransitiongapinnovation defaultstate policy feature) theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeTransitionGapInnovation_episodeBatchOfTrajectories","label":"cumulativeTransitionGapInnovation_episodeBatchOfTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeTransitionGapInnovation_episodeBatchOfTrajectories","description":"Reconstruction is literal on the batch map used by the generated law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-84b8bd3da540","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7159,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:883"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeTransitionGapInnovation_episodeBatchOfTrajectories (mdp : MDP State Action) (defaultState : State) (policy : MarkovPolicy mdp) (feature : Fin mdp.horizon -> State -> Real) (hhorizon : 0 < mdp.horizon) (trajectories : Fin 1 -> State × StepTrace Action State mdp.horizon) : EpisodeBatch.cumulativeTransitionGapInnovation defaultState policy feature (mdp.episodeBatchOfTrajectories 1 trajectories) = mdp.sampledCumulativeTransitionGapInnovationFrom policy feature mdp.horizon le_rfl (trajectories 0).1 (trajectories 0).2","missing":[],"search":"cumulativetransitiongapinnovation_episodebatchoftrajectories banditrlproof.finitehorizonrl.episodebatch.cumulativetransitiongapinnovation_episodebatchoftrajectories reconstruction is literal on the batch map used by the generated law. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchMeasure_one_cumulativeDeterministicGapInnovation_compensated_hasMGFUpperBoundAt","label":"iidEpisodeBatchMeasure_one_cumulativeDeterministicGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchMeasure_one_cumulativeDeterministicGapInnovation_compensated_hasMGFUpperBoundAt","description":"The exact one-episode batch image retains the canonical deterministic-table stage-resolved MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-2d49f21e7f27","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7160,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:915"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_one_cumulativeDeterministicGapInnovation_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (hhorizon : 0 < mdp.horizon) (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun batch : EpisodeBatch mdp 1 => tilt * batch.cumulativeDeterministicGapInnovation defaultState table feature - (mdp.horizon : Real) * tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (table.toMarkovPolicy.iidEpisodeBatchMeasure initialState 1)","missing":[],"search":"iidepisodebatchmeasure_one_cumulativedeterministicgapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.iidepisodebatchmeasure_one_cumulativedeterministicgapinnovation_compensated_hasmgfupperboundat the exact one-episode batch image retains the canonical deterministic-table stage-resolved mgf. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_cumulativeTransitionGapInnovation_compensated_hasMGFUpperBoundAt","label":"iidEpisodeBatchMeasure_one_cumulativeTransitionGapInnovation_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_cumulativeTransitionGapInnovation_compensated_hasMGFUpperBoundAt","description":"The exact one-episode batch image retains the stage-resolved MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibellmaninnovation/index.html#decl-6fdd1415ce39","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","order":7161,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation.lean:994"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_one_cumulativeTransitionGapInnovation_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (hhorizon : 0 < mdp.horizon) (feature : Fin mdp.horizon -> State -> Real) (hfeature : forall stage nextState, feature stage nextState ∈ Set.Icc (0 : Real) mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun batch : EpisodeBatch mdp 1 => tilt * batch.cumulativeTransitionGapInnovation defaultState policy feature - (mdp.horizon : Real) * tilt ^ 2 * (mdp.horizon : Real) ^ 2 / 8) 1 0 (policy.iidEpisodeBatchMeasure initialState 1)","missing":[],"search":"iidepisodebatchmeasure_one_cumulativetransitiongapinnovation_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure_one_cumulativetransitiongapinnovation_compensated_hasmgfupperboundat the exact one-episode batch image retains the stage-resolved mgf. theorem compiled","shard":"modules/9e286f533f4df7e7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionCoordinateVariance","label":"transitionCoordinateVariance","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionCoordinateVariance","description":"True Bernoulli variance of one transition singleton.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-700fb689d498","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7162,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def transitionCoordinateVariance (mdp : MDP State Action) (state : State) (action : Action) (nextState : State) : Real","missing":[],"search":"transitioncoordinatevariance banditrlproof.finitehorizonrl.mdp.transitioncoordinatevariance true bernoulli variance of one transition singleton. definition compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionCoordinateVariance_nonneg","label":"transitionCoordinateVariance_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionCoordinateVariance_nonneg","description":"theorem transitionCoordinateVariance_nonneg (mdp : MDP State Action) (state : State) (action : Action) (nextState : State) : 0 <= mdp.transitionCoordinateVariance state action nextState","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-678729e1d394","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7163,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionCoordinateVariance_nonneg (mdp : MDP State Action) (state : State) (action : Action) (nextState : State) : 0 <= mdp.transitionCoordinateVariance state action nextState","missing":[],"search":"transitioncoordinatevariance_nonneg banditrlproof.finitehorizonrl.mdp.transitioncoordinatevariance_nonneg theorem transitioncoordinatevariance_nonneg (mdp : mdp state action) (state : state) (action : action) (nextstate : state) : 0 <= mdp.transitioncoordinatevariance state action nextstate theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.integral_sq_indicator_sub_transitionMass","label":"integral_sq_indicator_sub_transitionMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.integral_sq_indicator_sub_transitionMass","description":"A centered transition indicator has exact second moment `p(1-p)`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-b3d964f23173","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7164,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_indicator_sub_transitionMass (mdp : MDP State Action) (state : State) (action : Action) (nextState : State) : (∫ y, ((if y = nextState then 1 else 0) - (mdp.transition (state, action)).real {nextState}) ^ 2 ∂mdp.transition (state, action)) = mdp.transitionCoordinateVariance state action nextState","missing":[],"search":"integral_sq_indicator_sub_transitionmass banditrlproof.finitehorizonrl.mdp.integral_sq_indicator_sub_transitionmass a centered transition indicator has exact second moment `p(1-p)`. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","label":"transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","description":"The visited transition singleton has a variance-sensitive compensated MGF. The hard `|tilt| <= 1` contract is the standard Bernstein small-tilt range.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-09530b3f1ce5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7165,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionResidualHead_variance_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (currentState : State) (chosenAction : Action) (tilt : Real) (htilt : |tilt| <= 1) : Concentration.HasMGFUpperBoundAt (fun nextState => tilt * mdp.transitionResidualHead targetState targetAction targetNextState currentState (chosenAction, nextState) - tilt ^ 2 * mdp.transitionCoordinateVariance targetState targetAction targetNextState * mdp.transitionVisitHead targetState targetAction currentState (chosenAction, nextState)) 1 0 (mdp.transition (currentState, chosenAction))","missing":[],"search":"transitionresidualhead_variance_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.transitionresidualhead_variance_compensated_hasmgfupperboundat the visited transition singleton has a variance-sensitive compensated mgf. the hard `|tilt| <= 1` contract is the standard bernstein small-tilt range. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","label":"actionStateKernel_transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","description":"The generated action/transition head preserves the exact coordinate variance compensator.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-56f97c3672b2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7166,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_transitionResidualHead_variance_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (targetState : State) (targetAction : Action) (targetNextState : State) (currentState : State) (stage : Fin mdp.horizon) (tilt : Real) (htilt : |tilt| <= 1) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * mdp.transitionResidualHead targetState targetAction targetNextState currentState head - tilt ^ 2 * mdp.transitionCoordinateVariance targetState targetAction targetNextState * mdp.transitionVisitHead targetState targetAction currentState head) 1 0 (policy.actionStateKernel stage currentState)","missing":[],"search":"actionstatekernel_transitionresidualhead_variance_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.actionstatekernel_transitionresidualhead_variance_compensated_hasmgfupperboundat the generated action/transition head preserves the exact coordinate variance compensator. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionResidual_variance_compensated_hasMGFUpperBoundAt","label":"trajectoryKernelRemaining_transitionResidual_variance_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionResidual_variance_compensated_hasMGFUpperBoundAt","description":"The entire generated episode trace has the exact coordinate-variance compensator, with the compensating count equal to literal visits.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-46c6208b0e23","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7167,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_transitionResidual_variance_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (targetState : State) (targetAction : Action) (targetNextState : State) (tilt : Real) (htilt : |tilt| <= 1) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (currentState : State) : Concentration.HasMGFUpperBoundAt (fun trace => tilt * mdp.transitionResidualFrom targetState targetAction targetNextState remaining currentState trace - tilt ^ 2 * mdp.transitionCoordinateVariance targetState targetAction targetNextState * mdp.transitionVisitFrom targetState targetAction remaining currentState trace) 1 0 (policy.trajectoryKernelRemaining remaining hremaining currentState)","missing":[],"search":"trajectorykernelremaining_transitionresidual_variance_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorykernelremaining_transitionresidual_variance_compensated_hasmgfupperboundat the entire generated episode trace has the exact coordinate-variance compensator, with the compensating count equal to literal visits. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionResidual_variance_compensated_hasMGFUpperBoundAt","label":"trajectoryMeasure_transitionResidual_variance_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionResidual_variance_compensated_hasMGFUpperBoundAt","description":"Integrating the random initial state keeps the exact coordinate-variance compensator on the generated trajectory measure.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-0fd82b98340f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7168,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:412"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_transitionResidual_variance_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (targetState : State) (targetAction : Action) (targetNextState : State) (tilt : Real) (htilt : |tilt| <= 1) : Concentration.HasMGFUpperBoundAt (fun trajectory => tilt * mdp.transitionResidualFrom targetState targetAction targetNextState mdp.horizon trajectory.1 trajectory.2 - tilt ^ 2 * mdp.transitionCoordinateVariance targetState targetAction targetNextState * mdp.transitionVisitFrom targetState targetAction mdp.horizon trajectory.1 trajectory.2) 1 0 (policy.trajectoryMeasure initialState)","missing":[],"search":"trajectorymeasure_transitionresidual_variance_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorymeasure_transitionresidual_variance_compensated_hasmgfupperboundat integrating the random initial state keeps the exact coordinate-variance compensator on the generated trajectory measure. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionResidual_variance_compensated_hasMGFUpperBoundAt","label":"iidEpisodeBatchMeasure_one_aggregateTransitionResidual_variance_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionResidual_variance_compensated_hasMGFUpperBoundAt","description":"The exact one-episode batch image of the generated trace inherits the coordinate-variance MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-bcbaa4f0bb07","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7169,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:481"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_one_aggregateTransitionResidual_variance_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (state : State) (action : Action) (nextState : State) (tilt : Real) (htilt : |tilt| <= 1) : Concentration.HasMGFUpperBoundAt (fun batch : EpisodeBatch mdp 1 => tilt * batch.aggregateTransitionResidual state action nextState - tilt ^ 2 * mdp.transitionCoordinateVariance state action nextState * batch.aggregateVisitReal state action) 1 0 (policy.iidEpisodeBatchMeasure initialState 1)","missing":[],"search":"iidepisodebatchmeasure_one_aggregatetransitionresidual_variance_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure_one_aggregatetransitionresidual_variance_compensated_hasmgfupperboundat the exact one-episode batch image of the generated trace inherits the coordinate-variance mgf. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_succ_variance_compensated_hasCondMGFUpperBoundAt","label":"aggregateTransitionResidual_succ_variance_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_succ_variance_compensated_hasCondMGFUpperBoundAt","description":"Every successor episode in the recurrent generated process inherits the same coordinate-variance compensated conditional MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-468770fb1bbe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7170,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:572"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionResidual_succ_variance_compensated_hasCondMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (n : Nat) (tilt : Real) (htilt : |tilt| <= 1) : Concentration.HasCondMGFUpperBoundAt (mΩ := MeasurableSpace.pi) (batchPrefixFiltration (mdp := mdp) 1 n) ((batchPrefixFiltration (mdp := mdp) 1).le n) (fun trajectory => tilt * source.aggregateTransitionResidualIncrement state action nextState (n + 1) trajectory - tilt ^ 2 * mdp.transitionCoordinateVariance state action nextState * source.aggregateVisitIncrement state action (n + 1) trajectory) 1 0 source.trajectoryMeasure","missing":[],"search":"aggregatetransitionresidual_succ_variance_compensated_hascondmgfupperboundat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidual_succ_variance_compensated_hascondmgfupperboundat every successor episode in the recurrent generated process inherits the same coordinate-variance compensated conditional mgf. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_zero_variance_compensated_hasMGFUpperBoundAt","label":"aggregateTransitionResidual_zero_variance_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_zero_variance_compensated_hasMGFUpperBoundAt","description":"Coordinate zero has the same exact coordinate-variance certificate under the initial marginal of the generated recurrent trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-e05bce8a70d4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7171,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:656"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionResidual_zero_variance_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (tilt : Real) (htilt : |tilt| <= 1) : Concentration.HasMGFUpperBoundAt (fun trajectory : EpisodeBatchTrajectory mdp 1 => tilt * source.aggregateTransitionResidualIncrement state action nextState 0 trajectory - tilt ^ 2 * mdp.transitionCoordinateVariance state action nextState * source.aggregateVisitIncrement state action 0 trajectory) 1 0 source.trajectoryMeasure","missing":[],"search":"aggregatetransitionresidual_zero_variance_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidual_zero_variance_compensated_hasmgfupperboundat coordinate zero has the same exact coordinate-variance certificate under the initial marginal of the generated recurrent trajectory. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","label":"measure_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","description":"Fixed-tilt actual-count upper tail with the true transition-coordinate variance.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-f2b98cea9074","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7172,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:689"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (rounds : Nat) (tilt threshold visitBudget : Real) (htilt : 0 < tilt) (htilt_le : tilt <= 1) : source.trajectoryMeasure {trajectory | threshold <= ∑ i ∈ Finset.range rounds, source.aggregateTransitionResidualIncrement state action nextState i trajectory ∧ (∑ i ∈ Finset.range rounds, source.aggregateVisitIncrement state action i trajectory) <= visitBudget} <= ENNReal.ofReal (Real.exp (-tilt * threshold + tilt ^ 2 * mdp.transitionCoordinateVariance state action nextState * visitBudget))","missing":[],"search":"measure_aggregatetransitionresidualsum_ge_inter_visitsum_le_variance banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measure_aggregatetransitionresidualsum_ge_inter_visitsum_le_variance fixed-tilt actual-count upper tail with the true transition-coordinate variance. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","label":"measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","description":"Two-sided coordinate-variance prefix tail on the same generated law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvibernsteinconfidence/index.html#decl-ea3e71565f3f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","order":7173,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence.lean:744"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (rounds : Nat) (tilt threshold visitBudget : Real) (htilt : 0 < tilt) (htilt_le : tilt <= 1) : source.trajectoryMeasure {trajectory | threshold <= |∑ i ∈ Finset.range rounds, source.aggregateTransitionResidualIncrement state action nextState i trajectory| ∧ (∑ i ∈ Finset.range rounds, source.aggregateVisitIncrement state action i trajectory) <= visitBudget} <= 2 * ENNReal.ofReal (Real.exp (-tilt * threshold + tilt ^ 2 * mdp.transitionCoordinateVariance state action nextState * visitBudget))","missing":[],"search":"measure_abs_aggregatetransitionresidualsum_ge_inter_visitsum_le_variance banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measure_abs_aggregatetransitionresidualsum_ge_inter_visitsum_le_variance two-sided coordinate-variance prefix tail on the same generated law. theorem compiled","shard":"modules/ecc80b678ea37274.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement","label":"generatedBatchAggregateVisitIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement","description":"Aggregate visits contributed by one actual generated batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-d5241406d817","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7174,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def generatedBatchAggregateVisitIncrement {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : Nat","missing":[],"search":"generatedbatchaggregatevisitincrement banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedbatchaggregatevisitincrement aggregate visits contributed by one actual generated batch. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_generatedBatchAggregateVisitIncrement","label":"batchedPrefixCount_generatedBatchAggregateVisitIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_generatedBatchAggregateVisitIncrement","description":"theorem batchedPrefixCount_generatedBatchAggregateVisitIncrement {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (round : Nat) : batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) (round + 1) = adaptiveCumulativeAggregateVisitCountAt trajectory round state action","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-1c24316841df","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7175,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem batchedPrefixCount_generatedBatchAggregateVisitIncrement {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (round : Nat) : batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) (round + 1) = adaptiveCumulativeAggregateVisitCountAt trajectory round state action","missing":[],"search":"batchedprefixcount_generatedbatchaggregatevisitincrement banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batchedprefixcount_generatedbatchaggregatevisitincrement theorem batchedprefixcount_generatedbatchaggregatevisitincrement {mdp : mdp state action} (trajectory : episodebatchtrajectory mdp 1) (state : state) (action : action) (round : nat) : batchedprefixcount (generatedbatchaggregatevisitincrement trajectory state action) (round + 1) = adaptivecumulativeaggregatevisitcountat trajectory round state action theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement_le_horizon","label":"generatedBatchAggregateVisitIncrement_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement_le_horizon","description":"theorem generatedBatchAggregateVisitIncrement_le_horizon {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : generatedBatchAggregateVisitIncrement trajectory state action episode <= mdp.horizon","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-23fc86d91ea0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7176,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem generatedBatchAggregateVisitIncrement_le_horizon {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : generatedBatchAggregateVisitIncrement trajectory state action episode <= mdp.horizon","missing":[],"search":"generatedbatchaggregatevisitincrement_le_horizon banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedbatchaggregatevisitincrement_le_horizon theorem generatedbatchaggregatevisitincrement_le_horizon {mdp : mdp state action} (trajectory : episodebatchtrajectory mdp 1) (state : state) (action : action) (episode : nat) : generatedbatchaggregatevisitincrement trajectory state action episode <= mdp.horizon theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.SuccessorBatchAligned","label":"SuccessorBatchAligned","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.SuccessorBatchAligned","description":"Every stored state/action record is aligned with its reconstructed full trajectory and with the deterministic recurrent choice. This is the exact portion of mapped-batch identity consumed by the regret ledger.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-f9d4f84ebdeb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7177,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def SuccessorBatchAligned (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : Prop","missing":[],"search":"successorbatchaligned banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorbatchaligned every stored state/action record is aligned with its reconstructed full trajectory and with the deterministic recurrent choice. this is the exact portion of mapped-batch identity consumed by the regret ledger. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBatchAligned_step","label":"successorBatchAligned_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBatchAligned_step","description":"On an aligned successor batch, each canonical recursively selected state/action pair is exactly the corresponding stored empirical record.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-979e9ae108b0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7178,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorBatchAligned_step {mdp : MDP State Action} {defaultState : State} {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} {n : Nat} (haligned : SuccessorBatchAligned mdp defaultState episodes delta trajectory n) (stage : Fin mdp.horizon) : let batch := trajectory (n + 1) let reconstructed : State × StepTrace Action State mdp.horizon := (batch.reconstructedInitialState defaultState, batch.reconstructedStepTrace) (batch 0 stage).state = mdp.trajectoryStateAt reconstructed stage ∧ (batch 0 stage).action = successorPolicyTable mdp defaultState episodes delta trajectory n stage (mdp.trajectoryStateAt reconstructed stage)","missing":[],"search":"successorbatchaligned_step banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorbatchaligned_step on an aligned successor batch, each canonical recursively selected state/action pair is exactly the corresponding stored empirical record. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement_eq_indicatorSum","label":"generatedBatchAggregateVisitIncrement_eq_indicatorSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement_eq_indicatorSum","description":"Number of stored visits to one pair is the sum of its stage indicators.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-61527ce6f95f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7179,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem generatedBatchAggregateVisitIncrement_eq_indicatorSum {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : generatedBatchAggregateVisitIncrement trajectory state action episode = ∑ stage : Fin mdp.horizon, if (trajectory episode 0 stage).state = state ∧ (trajectory episode 0 stage).action = action then 1 else 0","missing":[],"search":"generatedbatchaggregatevisitincrement_eq_indicatorsum banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedbatchaggregatevisitincrement_eq_indicatorsum number of stored visits to one pair is the sum of its stage indicators. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedBatchAggregateVisitIncrement_eq_horizon","label":"sum_generatedBatchAggregateVisitIncrement_eq_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedBatchAggregateVisitIncrement_eq_horizon","description":"The sum of all pair increments in one one-trajectory batch is exactly H.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-d64a24ab4f8e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7180,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:109"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_generatedBatchAggregateVisitIncrement_eq_horizon {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (episode : Nat) : ∑ state : State, ∑ action : Action, generatedBatchAggregateVisitIncrement trajectory state action episode = mdp.horizon","missing":[],"search":"sum_generatedbatchaggregatevisitincrement_eq_horizon banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_generatedbatchaggregatevisitincrement_eq_horizon the sum of all pair increments in one one-trajectory batch is exactly h. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairLocalCharge","label":"generatedPairLocalCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairLocalCharge","description":"The clipped UCBVI-CH charge assigned to one state-action pair in one episode, using the literal strict-prefix count of generated batch records.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-88634a78cd57","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7181,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:150"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def generatedPairLocalCharge (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : Real","missing":[],"search":"generatedpairlocalcharge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedpairlocalcharge the clipped ucbvi-ch charge assigned to one state-action pair in one episode, using the literal strict-prefix count of generated batch records. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairLocalCharge_nonneg","label":"generatedPairLocalCharge_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairLocalCharge_nonneg","description":"theorem generatedPairLocalCharge_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : 0 <= generatedPairLocalCharge mdp episodes delta trajectory state action episode","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-1441966bb953","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7182,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:167"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem generatedPairLocalCharge_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : 0 <= generatedPairLocalCharge mdp episodes delta trajectory state action episode","missing":[],"search":"generatedpairlocalcharge_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedpairlocalcharge_nonneg theorem generatedpairlocalcharge_nonneg (mdp : mdp state action) (episodes : nat) (delta : real) (trajectory : episodebatchtrajectory mdp 1) (state : state) (action : action) (episode : nat) : 0 <= generatedpairlocalcharge mdp episodes delta trajectory state action episode theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_eq_generatedPairLocalCharge_of_aligned","label":"successorLocalCharge_eq_generatedPairLocalCharge_of_aligned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_eq_generatedPairLocalCharge_of_aligned","description":"Alignment converts the policy-selected local charge at a recorded stage to the pair charge indexed by that record's exact state and action.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-eeb51e04c093","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7183,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:182"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorLocalCharge_eq_generatedPairLocalCharge_of_aligned {mdp : MDP State Action} {defaultState : State} {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} {n : Nat} (haligned : SuccessorBatchAligned mdp defaultState episodes delta trajectory n) (stage : Fin mdp.horizon) : successorLocalCharge mdp defaultState episodes delta trajectory n (mdp.horizon - (stage.val + 1)) (by omega) (trajectory (n + 1) 0 stage).state = generatedPairLocalCharge mdp episodes delta trajectory (trajectory (n + 1) 0 stage).state (trajectory (n + 1) 0 stage).action (n + 1)","missing":[],"search":"successorlocalcharge_eq_generatedpairlocalcharge_of_aligned banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorlocalcharge_eq_generatedpairlocalcharge_of_aligned alignment converts the policy-selected local charge at a recorded stage to the pair charge indexed by that record's exact state and action. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_successorLocalCharge_eq_sum_pairIncrement_mul_charge_of_aligned","label":"sum_successorLocalCharge_eq_sum_pairIncrement_mul_charge_of_aligned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_successorLocalCharge_eq_sum_pairIncrement_mul_charge_of_aligned","description":"One generated episode's unweighted local charge is exactly the sum of pair charges repeated by the actual state-action visit multiplicities.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-e3afdc1d4787","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7184,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:222"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_successorLocalCharge_eq_sum_pairIncrement_mul_charge_of_aligned {mdp : MDP State Action} {defaultState : State} {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} {n : Nat} (haligned : SuccessorBatchAligned mdp defaultState episodes delta trajectory n) : (∑ stage : Fin mdp.horizon, successorLocalCharge mdp defaultState episodes delta trajectory n (mdp.horizon - (stage.val + 1)) (by omega) (trajectory (n + 1) 0 stage).state) = ∑ state : State, ∑ action : Action, (generatedBatchAggregateVisitIncrement trajectory state action (n + 1) : Real) * generatedPairLocalCharge mdp episodes delta trajectory state action (n + 1)","missing":[],"search":"sum_successorlocalcharge_eq_sum_pairincrement_mul_charge_of_aligned banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_successorlocalcharge_eq_sum_pairincrement_mul_charge_of_aligned one generated episode's unweighted local charge is exactly the sum of pair charges repeated by the actual state-action visit multiplicities. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.remainingGlobalStage","label":"remainingGlobalStage","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.remainingGlobalStage","description":"Chronological global stage corresponding to a coordinate of a remaining suffix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-96f17176f686","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7185,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:275"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def remainingGlobalStage (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (stage : Fin remaining) : Fin mdp.horizon","missing":[],"search":"remainingglobalstage banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.remainingglobalstage chronological global stage corresponding to a coordinate of a remaining suffix. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_eq_of_remaining_eq","label":"successorLocalCharge_eq_of_remaining_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_eq_of_remaining_eq","description":"theorem successorLocalCharge_eq_of_remaining_eq (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) {left right : Nat} (hleft : left + 1 <= mdp.horizon) (hright : right + 1 <= mdp.horizon) (h : left = right) (state : State) : successorLocalCharge mdp defaultState episodes delta trajectory n left hleft state = successorLocalCharge mdp d…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-62128a38441b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7186,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:280"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorLocalCharge_eq_of_remaining_eq (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) {left right : Nat} (hleft : left + 1 <= mdp.horizon) (hright : right + 1 <= mdp.horizon) (h : left = right) (state : State) : successorLocalCharge mdp defaultState episodes delta trajectory n left hleft state = successorLocalCharge mdp defaultState episodes delta trajectory n right hright state","missing":[],"search":"successorlocalcharge_eq_of_remaining_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorlocalcharge_eq_of_remaining_eq theorem successorlocalcharge_eq_of_remaining_eq (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (trajectory : episodebatchtrajectory mdp 1) (n : nat) {left right : nat} (hleft : left + 1 <= mdp.horizon) (hright : right + 1 <= mdp.horizon) (h : left = right) (state : state) : successorlocalcharge mdp defaultstate episodes delta trajectory n left hleft state = successorlocalcharge mdp defaultstate episodes delta trajectory n right hright state theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorStageLocalCharge","label":"successorStageLocalCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorStageLocalCharge","description":"Proof-irrelevant chronological wrapper around the remaining-indexed local charge. Keeping the proof argument out of finite sums avoids dependent rewrites in the charge telescope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-a7b15a429ffa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7187,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:296"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorStageLocalCharge (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) (stage : Fin mdp.horizon) (state : State) : Real","missing":[],"search":"successorstagelocalcharge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorstagelocalcharge proof-irrelevant chronological wrapper around the remaining-indexed local charge. keeping the proof argument out of finite sums avoids dependent rewrites in the charge telescope. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeFrom_eq_sum","label":"successorWeightedChargeFrom_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeFrom_eq_sum","description":"The recursive charge is exactly its chronological finite sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-5f4154a6d639","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7188,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:305"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorWeightedChargeFrom_eq_sum (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : forall (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (trace : StepTrace Action State remaining), successorWeightedChargeFrom mdp defaultState episodes delta trajectory n remaining hremaining state trace = ∑ stage : Fin remaining, normalizedBellmanChargeWeight mdp (remainingGlobalStage mdp remaining hremaining stage) * successorStageLocalCharge mdp defaultState episodes delta trajectory n (remainingGlobalStage mdp remaining hremaining stage) (StepTrace.stateAt state trace stage)","missing":[],"search":"successorweightedchargefrom_eq_sum banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorweightedchargefrom_eq_sum the recursive charge is exactly its chronological finite sum. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge_eq_successorWeightedChargeSum_of_aligned","label":"successorCanonicalWeightedCharge_eq_successorWeightedChargeSum_of_aligned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge_eq_successorWeightedChargeSum_of_aligned","description":"On an aligned generated successor batch, the recursively reconstructed canonical charge is the chronological recorded charge sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-f476b0d5cc7c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7189,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:378"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorCanonicalWeightedCharge_eq_successorWeightedChargeSum_of_aligned {mdp : MDP State Action} {defaultState : State} {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} {n : Nat} (haligned : SuccessorBatchAligned mdp defaultState episodes delta trajectory n) : successorCanonicalWeightedCharge mdp defaultState episodes delta trajectory n = successorWeightedChargeSum mdp defaultState episodes delta trajectory n","missing":[],"search":"successorcanonicalweightedcharge_eq_successorweightedchargesum_of_aligned banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorcanonicalweightedcharge_eq_successorweightedchargesum_of_aligned on an aligned generated successor batch, the recursively reconstructed canonical charge is the chronological recorded charge sum. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge_le_sum_pairIncrement_mul_charge_of_aligned","label":"successorCanonicalWeightedCharge_le_sum_pairIncrement_mul_charge_of_aligned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge_le_sum_pairIncrement_mul_charge_of_aligned","description":"One aligned canonical successor charge is dominated by its unweighted state-action multiplicity sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-664524b71469","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7190,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:407"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorCanonicalWeightedCharge_le_sum_pairIncrement_mul_charge_of_aligned {mdp : MDP State Action} {defaultState : State} {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} {n : Nat} (hhorizon : 0 < mdp.horizon) (haligned : SuccessorBatchAligned mdp defaultState episodes delta trajectory n) : successorCanonicalWeightedCharge mdp defaultState episodes delta trajectory n <= ∑ state : State, ∑ action : Action, (generatedBatchAggregateVisitIncrement trajectory state action (n + 1) : Real) * generatedPairLocalCharge mdp episodes delta trajectory state action (n + 1)","missing":[],"search":"successorcanonicalweightedcharge_le_sum_pairincrement_mul_charge_of_aligned banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorcanonicalweightedcharge_le_sum_pairincrement_mul_charge_of_aligned one aligned canonical successor charge is dominated by its unweighted state-action multiplicity sum. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge","label":"totalGeneratedPairCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge","description":"Total pair charge over the first `episodes` generated batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-a0f07996cb90","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7191,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:437"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def totalGeneratedPairCharge (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"totalgeneratedpaircharge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.totalgeneratedpaircharge total pair charge over the first `episodes` generated batches. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairChargeTerm_nonneg","label":"generatedPairChargeTerm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairChargeTerm_nonneg","description":"theorem generatedPairChargeTerm_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : 0 <= (generatedBatchAggregateVisitIncrement trajectory state action episode : Real) * generatedPairLocalCharge mdp episodes delta trajectory state action episode","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-099669b11d5e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7192,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:445"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem generatedPairChargeTerm_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (episode : Nat) : 0 <= (generatedBatchAggregateVisitIncrement trajectory state action episode : Real) * generatedPairLocalCharge mdp episodes delta trajectory state action episode","missing":[],"search":"generatedpairchargeterm_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedpairchargeterm_nonneg theorem generatedpairchargeterm_nonneg (mdp : mdp state action) (episodes : nat) (delta : real) (trajectory : episodebatchtrajectory mdp 1) (state : state) (action : action) (episode : nat) : 0 <= (generatedbatchaggregatevisitincrement trajectory state action episode : real) * generatedpairlocalcharge mdp episodes delta trajectory state action episode theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_range_succ_shift_le_sum_range","label":"sum_range_succ_shift_le_sum_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_range_succ_shift_le_sum_range","description":"private theorem sum_range_succ_shift_le_sum_range (f : Nat -> Real) (rounds : Nat) (hf : forall n, 0 <= f n) : (Finset.range (rounds - 1)).sum (fun n => f (n + 1)) <= (Finset.range rounds).sum f","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-cd12e1a43158","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7193,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:454"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem sum_range_succ_shift_le_sum_range (f : Nat -> Real) (rounds : Nat) (hf : forall n, 0 <= f n) : (Finset.range (rounds - 1)).sum (fun n => f (n + 1)) <= (Finset.range rounds).sum f","missing":[],"search":"sum_range_succ_shift_le_sum_range banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_range_succ_shift_le_sum_range private theorem sum_range_succ_shift_le_sum_range (f : nat -> real) (rounds : nat) (hf : forall n, 0 <= f n) : (finset.range (rounds - 1)).sum (fun n => f (n + 1)) <= (finset.range rounds).sum f theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_successorCanonicalWeightedCharge_le_totalGeneratedPairCharge","label":"sum_successorCanonicalWeightedCharge_le_totalGeneratedPairCharge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_successorCanonicalWeightedCharge_le_totalGeneratedPairCharge","description":"Summing all aligned successor episodes is dominated by the all-batch pair charge ledger; coordinate zero appears only on the right and is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-a6c1e0acd0dd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7194,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:471"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_successorCanonicalWeightedCharge_le_totalGeneratedPairCharge {mdp : MDP State Action} {defaultState : State} {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (haligned : forall n, n + 1 < episodes -> SuccessorBatchAligned mdp defaultState episodes delta trajectory n) : (Finset.range (episodes - 1)).sum (fun n => successorCanonicalWeightedCharge mdp defaultState episodes delta trajectory n) <= totalGeneratedPairCharge mdp episodes delta trajectory","missing":[],"search":"sum_successorcanonicalweightedcharge_le_totalgeneratedpaircharge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_successorcanonicalweightedcharge_le_totalgeneratedpaircharge summing all aligned successor episodes is dominated by the all-batch pair charge ledger; coordinate zero appears only on the right and is nonnegative. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.min_add_le_add_min","label":"min_add_le_add_min","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.min_add_le_add_min","description":"private theorem min_add_le_add_min {cap x y : Real} (hx : 0 <= x) : min cap (x + y) <= x + min cap y","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-832a926d3b61","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7195,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:512"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem min_add_le_add_min {cap x y : Real} (hx : 0 <= x) : min cap (x + y) <= x + min cap y","missing":[],"search":"min_add_le_add_min banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.min_add_le_add_min private theorem min_add_le_add_min {cap x y : real} (hx : 0 <= x) : min cap (x + y) <= x + min cap y theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.reciprocalChargeThreshold","label":"reciprocalChargeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.reciprocalChargeThreshold","description":"Natural threshold at which the clipped reciprocal correction starts its logarithmic telescope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-f638746fb59f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7196,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:521"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def reciprocalChargeThreshold (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Nat","missing":[],"search":"reciprocalchargethreshold banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.reciprocalchargethreshold natural threshold at which the clipped reciprocal correction starts its logarithmic telescope. definition compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedPairCharge_le_accounting","label":"sum_generatedPairCharge_le_accounting","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedPairCharge_le_accounting","description":"Exact one-pair accounting before the final paper-constant simplification.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-c633cba25584","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7197,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:527"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_generatedPairCharge_le_accounting (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (hhorizon : 0 < mdp.horizon) (hlog : 1 <= logFactor (State := State) (Action := Action) mdp episodes delta) : (Finset.range episodes).sum (fun episode => (generatedBatchAggregateVisitIncrement trajectory state action episode : Real) * generatedPairLocalCharge mdp episodes delta trajectory state action episode) <= (mdp.horizon : Real) * (1 + mdp.horizon) + (9 * (mdp.horizon : Real) * logFactor (State := State) (Action := Action) mdp episodes delta) * (2 * Real.sqrt (batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes) + 2 * mdp.horizon + 2 * mdp.horizon * Real.log ((max 1 (batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes) : Nat)…","missing":[],"search":"sum_generatedpaircharge_le_accounting banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_generatedpaircharge_le_accounting exact one-pair accounting before the final paper-constant simplification. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedPairCharge_le_explicit","label":"sum_generatedPairCharge_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedPairCharge_le_explicit","description":"The accounting ledger fits the explicit per-pair `18/238` budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-8104b995788c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7198,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:654"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_generatedPairCharge_le_explicit (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (hhorizon : 0 < mdp.horizon) (hlog : 1 <= logFactor (State := State) (Action := Action) mdp episodes delta) (hlogCount : Real.log ((max 1 (batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes) : Nat) : Real) <= logFactor (State := State) (Action := Action) mdp episodes delta) : (Finset.range episodes).sum (fun episode => (generatedBatchAggregateVisitIncrement trajectory state action episode : Real) * generatedPairLocalCharge mdp episodes delta trajectory state action episode) <= 18 * (mdp.horizon : Real) * logFactor (State := State) (Action := Action) mdp episodes delta * Real.sqrt (batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes…","missing":[],"search":"sum_generatedpaircharge_le_explicit banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_generatedpaircharge_le_explicit the accounting ledger fits the explicit per-pair `18/238` budget. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.one_le_logFactor","label":"one_le_logFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.one_le_logFactor","description":"On the positive task domain the paper logarithmic factor is at least one. The factor `5` in the confidence numerator is deliberately retained here: it is what makes the statement true uniformly for every `delta <= 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-80cadd429fae","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7199,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:760"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_le_logFactor (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 1 <= logFactor (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"one_le_logfactor banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.one_le_logfactor on the positive task domain the paper logarithmic factor is at least one. the factor `5` in the confidence numerator is deliberately retained here: it is what makes the statement true uniformly for every `delta <= 1`. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_generatedBatchAggregateVisitIncrement_le_totalSteps","label":"batchedPrefixCount_generatedBatchAggregateVisitIncrement_le_totalSteps","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_generatedBatchAggregateVisitIncrement_le_totalSteps","description":"A single state-action pair is visited at most once per stage, hence at most `episodes * H` times in the complete generated ledger.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-3a2431427f79","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7200,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:803"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem batchedPrefixCount_generatedBatchAggregateVisitIncrement_le_totalSteps (mdp : MDP State Action) (episodes : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) : batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes <= totalSteps mdp episodes","missing":[],"search":"batchedprefixcount_generatedbatchaggregatevisitincrement_le_totalsteps banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batchedprefixcount_generatedbatchaggregatevisitincrement_le_totalsteps a single state-action pair is visited at most once per stage, hence at most `episodes * h` times in the complete generated ledger. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.log_max_batchedPrefixCount_le_logFactor","label":"log_max_batchedPrefixCount_le_logFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.log_max_batchedPrefixCount_le_logFactor","description":"The logarithm of every final pair count is controlled by the single paper logarithmic factor used by the generated policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-0f23bc28c9b7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7201,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:822"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem log_max_batchedPrefixCount_le_logFactor (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (state : State) (action : Action) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : Real.log ((max 1 (batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes) : Nat) : Real) <= logFactor (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"log_max_batchedprefixcount_le_logfactor banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.log_max_batchedprefixcount_le_logfactor the logarithm of every final pair count is controlled by the single paper logarithmic factor used by the generated policy. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_batchedPrefixCount_generatedBatchAggregateVisitIncrement_eq_totalSteps","label":"sum_batchedPrefixCount_generatedBatchAggregateVisitIncrement_eq_totalSteps","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_batchedPrefixCount_generatedBatchAggregateVisitIncrement_eq_totalSteps","description":"The aggregate count ledger is exact: summing final counts over every state-action pair gives the number `episodes * H` of generated transitions.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-8a64a70ea873","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7202,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:888"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_batchedPrefixCount_generatedBatchAggregateVisitIncrement_eq_totalSteps (mdp : MDP State Action) (episodes : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : (∑ state : State, ∑ action : Action, batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes) = totalSteps mdp episodes","missing":[],"search":"sum_batchedprefixcount_generatedbatchaggregatevisitincrement_eq_totalsteps banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_batchedprefixcount_generatedbatchaggregatevisitincrement_eq_totalsteps the aggregate count ledger is exact: summing final counts over every state-action pair gives the number `episodes * h` of generated transitions. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_sqrt_batchedPrefixCount_le","label":"sum_sqrt_batchedPrefixCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_sqrt_batchedPrefixCount_le","description":"Cauchy--Schwarz converts the exact visit ledger into the canonical `sqrt(S A K H)` exploration scale.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-9332e987a198","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7203,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:904"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_sqrt_batchedPrefixCount_le (mdp : MDP State Action) (episodes : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : (∑ state : State, ∑ action : Action, Real.sqrt (batchedPrefixCount (generatedBatchAggregateVisitIncrement trajectory state action) episodes)) <= Real.sqrt (totalSteps mdp episodes : Nat) * Real.sqrt ((Fintype.card State * Fintype.card Action : Nat) : Real)","missing":[],"search":"sum_sqrt_batchedprefixcount_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_sqrt_batchedprefixcount_le cauchy--schwarz converts the exact visit ledger into the canonical `sqrt(s a k h)` exploration scale. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","label":"totalGeneratedPairCharge_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","description":"Global deterministic charge budget for every generated batch in the canonical recurrent UCBVI-CH ledger. The two square-root factors are kept separate here so the exact total-count identity remains visible.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvichargesummation/index.html#decl-7e84c6d2d636","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","order":7204,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation.lean:942"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem totalGeneratedPairCharge_le_explicit (mdp : MDP State Action) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : totalGeneratedPairCharge mdp episodes delta trajectory <= 18 * (mdp.horizon : Real) * logFactor (State := State) (Action := Action) mdp episodes delta * (Real.sqrt (totalSteps mdp episodes : Nat) * Real.sqrt ((Fintype.card State * Fintype.card Action : Nat) : Real)) + 238 * (Fintype.card State : Real) ^ 2 * Fintype.card Action * (mdp.horizon : Real) ^ 2 * logFactor (State := State) (Action := Action) mdp episodes delta ^ 2","missing":[],"search":"totalgeneratedpaircharge_le_explicit banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.totalgeneratedpaircharge_le_explicit global deterministic charge budget for every generated batch in the canonical recurrent ucbvi-ch ledger. the two square-root factors are kept separate here so the exact total-count identity remains visible. theorem compiled","shard":"modules/4ac7987599d9953d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.QTable","label":"QTable","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.QTable","description":"A chronological action-value table for all actual decision stages.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-7bee3eb79dba","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7205,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev QTable (mdp : MDP State Action)","missing":[],"search":"qtable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.qtable a chronological action-value table for all actual decision stages. abbreviation compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.initialQTable","label":"initialQTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.initialQTable","description":"The UCBVI-CH initialization `Q_{0,h}(x,a)=H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-de9a15215220","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7206,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def initialQTable (mdp : MDP State Action) : QTable mdp","missing":[],"search":"initialqtable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.initialqtable the ucbvi-ch initialization `q_{0,h}(x,a)=h`. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeSummaryOfSequence","label":"cumulativeSummaryOfSequence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeSummaryOfSequence","description":"Sum a finite sequence of observed transition summaries coordinatewise.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-e62aa1e676a4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7207,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeSummaryOfSequence {mdp : MDP State Action} {n : Nat} (summaries : Fin n -> TransitionCountSummary mdp) : TransitionCountSummary mdp","missing":[],"search":"cumulativesummaryofsequence banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.cumulativesummaryofsequence sum a finite sequence of observed transition summaries coordinatewise. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining","label":"clippedValueRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining","description":"The recursively selected clipped value with a given previous-episode Q table and one pooled empirical transition model. The index is decisions remaining; its successor case uses the chronological stage `H-(remaining+1)`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-a2758cad8ec9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7208,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedValueRemaining {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> Real | 0, _ => fun _state => 0 | remaining + 1, hremaining => let stage := mdp.decisionStageRemaining remaining hremaining let tail := clippedValueRemaining previousQ summary defaultState bonusScale remaining (by omega) let score : State -> Action -> Real := fun state action => let count := summary.aggregateVisitCount state action if count = 0 then (mdp.horizon : Real) else min (previousQ stage state action) (min (mdp.horizon : Real) (mdp.reward state action + (∫ nextState, tail nextState ∂ summary.aggregateEmpiricalTransitionKernel defaultState (state, action)) + bonusScale / Real.sqrt count)) fun state => score state (FiniteRealArgmax.choose (fun action =>…","missing":[],"search":"clippedvalueremaining banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedvalueremaining the recursively selected clipped value with a given previous-episode q table and one pooled empirical transition model. the index is decisions remaining; its successor case uses the chronological stage `h-(remaining+1)`. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining","label":"clippedQRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining","description":"The action-value used at one remaining-horizon coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-74c433b4fd6b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7209,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedQRemaining {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) : Real","missing":[],"search":"clippedqremaining banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqremaining the action-value used at one remaining-horizon coordinate. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable","label":"clippedQTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable","description":"The complete chronological Q table produced by one recurrent update.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-687eded872ac","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7210,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedQTable (mdp : MDP State Action) (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) : QTable mdp","missing":[],"search":"clippedqtable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqtable the complete chronological q table produced by one recurrent update. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_eq_zero","label":"clippedQRemaining_of_aggregateVisitCount_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_eq_zero","description":"theorem clippedQRemaining_of_aggregateVisitCount_eq_zero {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) (hzero : summary.aggregateVisitCount state action = 0) : clippedQRemaining previousQ summary defaultState bonusScale remaining hremain…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-5060c6087d20","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7211,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedQRemaining_of_aggregateVisitCount_eq_zero {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) (hzero : summary.aggregateVisitCount state action = 0) : clippedQRemaining previousQ summary defaultState bonusScale remaining hremaining state action = (mdp.horizon : Real)","missing":[],"search":"clippedqremaining_of_aggregatevisitcount_eq_zero banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqremaining_of_aggregatevisitcount_eq_zero theorem clippedqremaining_of_aggregatevisitcount_eq_zero {mdp : mdp state action} (previousq : qtable mdp) (summary : transitioncountsummary mdp) (defaultstate : state) (bonusscale : real) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) (state : state) (action : action) (hzero : summary.aggregatevisitcount state action = 0) : clippedqremaining previousq summary defaultstate bonusscale remaining hremaining state action = (mdp.horizon : real) theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_pos","label":"clippedQRemaining_of_aggregateVisitCount_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_pos","description":"theorem clippedQRemaining_of_aggregateVisitCount_pos {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) (hpos : 0 < summary.aggregateVisitCount state action) : clippedQRemaining previousQ summary defaultState bonusScale remaining hremaining s…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-94934b6410dd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7212,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedQRemaining_of_aggregateVisitCount_pos {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) (hpos : 0 < summary.aggregateVisitCount state action) : clippedQRemaining previousQ summary defaultState bonusScale remaining hremaining state action = min (previousQ (mdp.decisionStageRemaining remaining hremaining) state action) (min (mdp.horizon : Real) (mdp.reward state action + (∫ nextState, clippedValueRemaining previousQ summary defaultState bonusScale remaining (by omega) nextState ∂ summary.aggregateEmpiricalTransitionKernel defaultState (state, action)) + bonusScale / Real.sqrt (summary.aggregateVisitCount state action)))","missing":[],"search":"clippedqremaining_of_aggregatevisitcount_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqremaining_of_aggregatevisitcount_pos theorem clippedqremaining_of_aggregatevisitcount_pos {mdp : mdp state action} (previousq : qtable mdp) (summary : transitioncountsummary mdp) (defaultstate : state) (bonusscale : real) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) (state : state) (action : action) (hpos : 0 < summary.aggregatevisitcount state action) : clippedqremaining previousq summary defaultstate bonusscale remaining hremaining state action = min (previousq (mdp.decisionstageremaining remaining hremaining) state action) (min (mdp.horizon : real) (mdp.reward state action + (∫ nextstate, clippedvalueremaining previousq summary defaultstate bonusscale remaining (by omega) nextstate ∂ summary.aggregateempiricaltransitionkernel defaultstate (state, action)) + bonusscale / real.sqrt (summary.aggregatevisitcount state action))) theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyTable","label":"clippedPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyTable","description":"Fixed-enumeration argmax table of one recurrent Q update.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-83b01844b704","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7213,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedPolicyTable (mdp : MDP State Action) (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"clippedpolicytable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedpolicytable fixed-enumeration argmax table of one recurrent q update. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_le_selected","label":"clippedQTable_le_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_le_selected","description":"The selected recurrent action maximizes the actual clipped Q table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-b7ef5c7a5595","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7214,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedQTable_le_selected (mdp : MDP State Action) (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (stage : Fin mdp.horizon) (state : State) (action : Action) : clippedQTable mdp previousQ summary defaultState bonusScale stage state action <= clippedQTable mdp previousQ summary defaultState bonusScale stage state (clippedPolicyTable mdp previousQ summary defaultState bonusScale stage state)","missing":[],"search":"clippedqtable_le_selected banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqtable_le_selected the selected recurrent action maximizes the actual clipped q table. theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_clippedPolicyTable","label":"measurable_clippedPolicyTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_clippedPolicyTable","description":"Every recurrent selector is measurable on the finite state space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-5141af96ca1a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7215,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_clippedPolicyTable (mdp : MDP State Action) (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (stage : Fin mdp.horizon) : Measurable (clippedPolicyTable mdp previousQ summary defaultState bonusScale stage)","missing":[],"search":"measurable_clippedpolicytable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_clippedpolicytable every recurrent selector is measurable on the finite state space. theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentQTableOfSummaries","label":"recurrentQTableOfSummaries","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentQTableOfSummaries","description":"Fold observed summaries into the genuine sequence of previous-Q updates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-daf3d866374d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7216,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentQTableOfSummaries (mdp : MDP State Action) (defaultState : State) (bonusScale : Real) : (n : Nat) -> (Fin n -> TransitionCountSummary mdp) -> QTable mdp | 0, _summaries => initialQTable mdp | n + 1, summaries => let previous := recurrentQTableOfSummaries mdp defaultState bonusScale n (fun i => summaries i.castSucc) let cumulative := cumulativeSummaryOfSequence summaries clippedQTable mdp previous cumulative defaultState bonusScale /-- Policy table obtained after the same finite previous-Q fold. -/ noncomputable def recurrentPolicyTableOfSummaries (mdp : MDP State Action) (defaultState : State) (bonusScale : Real) (n : Nat) (summaries : Fin n -> TransitionCountSummary mdp) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"recurrentqtableofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentqtableofsummaries fold observed summaries into the genuine sequence of previous-q updates. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentPolicyTableOfSummaries","label":"recurrentPolicyTableOfSummaries","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentPolicyTableOfSummaries","description":"Policy table obtained after the same finite previous-Q fold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-2db1dcf3f414","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7217,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentPolicyTableOfSummaries (mdp : MDP State Action) (defaultState : State) (bonusScale : Real) (n : Nat) (summaries : Fin n -> TransitionCountSummary mdp) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"recurrentpolicytableofsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentpolicytableofsummaries policy table obtained after the same finite previous-q fold. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.prefixTransitionSummaries","label":"prefixTransitionSummaries","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.prefixTransitionSummaries","description":"Extract exactly the transition summaries stored in one finite prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-ca16333600d4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7218,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:205"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def prefixTransitionSummaries {mdp : MDP State Action} {episodes n : Nat} (history : EpisodeBatchPrefix mdp episodes n) : Fin (n + 1) -> TransitionCountSummary mdp","missing":[],"search":"prefixtransitionsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.prefixtransitionsummaries extract exactly the transition summaries stored in one finite prefix. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_prefixTransitionSummaries","label":"measurable_prefixTransitionSummaries","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_prefixTransitionSummaries","description":"theorem measurable_prefixTransitionSummaries {mdp : MDP State Action} {episodes n : Nat} : Measurable (prefixTransitionSummaries : EpisodeBatchPrefix mdp episodes n -> Fin (n + 1) -> TransitionCountSummary mdp)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-4e182caaae87","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7219,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:214"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_prefixTransitionSummaries {mdp : MDP State Action} {episodes n : Nat} : Measurable (prefixTransitionSummaries : EpisodeBatchPrefix mdp episodes n -> Fin (n + 1) -> TransitionCountSummary mdp)","missing":[],"search":"measurable_prefixtransitionsummaries banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_prefixtransitionsummaries theorem measurable_prefixtransitionsummaries {mdp : mdp state action} {episodes n : nat} : measurable (prefixtransitionsummaries : episodebatchprefix mdp episodes n -> fin (n + 1) -> transitioncountsummary mdp) theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSuccessorTable","label":"recurrentSuccessorTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSuccessorTable","description":"Measurable strict-prefix selector for the recurrent generated source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-2ce7cd610e08","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7220,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:225"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentSuccessorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (bonusScale : Real) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"recurrentsuccessortable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsuccessortable measurable strict-prefix selector for the recurrent generated source. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_recurrentSuccessorTable","label":"measurable_recurrentSuccessorTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_recurrentSuccessorTable","description":"theorem measurable_recurrentSuccessorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (bonusScale : Real) (n : Nat) : Measurable (recurrentSuccessorTable (mdp := mdp) (episodes := episodes) defaultState bonusScale n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-1d217cfe3186","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7221,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:234"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_recurrentSuccessorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (bonusScale : Real) (n : Nat) : Measurable (recurrentSuccessorTable (mdp := mdp) (episodes := episodes) defaultState bonusScale n)","missing":[],"search":"measurable_recurrentsuccessortable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_recurrentsuccessortable theorem measurable_recurrentsuccessortable {mdp : mdp state action} {episodes : nat} (defaultstate : state) (bonusscale : real) (n : nat) : measurable (recurrentsuccessortable (mdp := mdp) (episodes := episodes) defaultstate bonusscale n) theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentInitialTable","label":"recurrentInitialTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentInitialTable","description":"The initial table is the fixed argmax of the all-`H` Q initialization.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-416ca63c663e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7222,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:257"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentInitialTable (mdp : MDP State Action) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"recurrentinitialtable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentinitialtable the initial table is the fixed argmax of the all-`h` q initialization. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource","label":"recurrentSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource","description":"Canonical one-episode recurrent UCBVI source. It has no arbitrary initial policy parameter: coordinate zero executes `recurrentInitialTable`; coordinate `n+1` executes the fold of exactly the observed coordinates `0,...,n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-e03865d46910","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7223,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:268"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def recurrentSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) : AdaptiveEpisodeBatchSource mdp initialState 1 where","missing":[],"search":"recurrentsource banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource canonical one-episode recurrent ucbvi source. it has no arbitrary initial policy parameter: coordinate zero executes `recurrentinitialtable`; coordinate `n+1` executes the fold of exactly the observed coordinates `0,...,n`. definition compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_zero","label":"recurrentSource_policyAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_zero","description":"Coordinate zero uses the all-`H` initialized recurrent policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-5e0e8ae84387","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7224,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_policyAt_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) : (recurrentSource mdp initialState defaultState episodes delta).policyAt trajectory 0 = (recurrentInitialTable mdp).toMarkovPolicy","missing":[],"search":"recurrentsource_policyat_zero banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_policyat_zero coordinate zero uses the all-`h` initialized recurrent policy. theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","label":"recurrentSource_policyAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","description":"Every successor policy is the exact strict-prefix previous-Q fold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviclippedplanner/index.html#decl-3862e4d21689","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","order":7225,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner.lean:304"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_policyAt_succ (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : (recurrentSource mdp initialState defaultState episodes delta).policyAt trajectory (n + 1) = (recurrentSuccessorTable defaultState (scale (State := State) (Action := Action) mdp episodes delta) n (Preorder.frestrictLe n trajectory)).toMarkovPolicy","missing":[],"search":"recurrentsource_policyat_succ banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_policyat_succ every successor policy is the exact strict-prefix previous-q fold. theorem compiled","shard":"modules/255177b2865b2c62.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt","label":"bernsteinTilt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt","description":"Bernstein tilt optimized for a deterministic variance budget, totalized at zero variance.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html#decl-5c0d71cf8a80","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","order":7226,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean:9"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinTilt (logBudget varianceBudget : Real) : Real","missing":[],"search":"bernsteintilt banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteintilt bernstein tilt optimized for a deterministic variance budget, totalized at zero variance. definition compiled","shard":"modules/7af76fc4fbb4e66a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinCoordinateThreshold","label":"bernsteinCoordinateThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinCoordinateThreshold","description":"noncomputable def bernsteinCoordinateThreshold (logBudget varianceBudget : Real) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html#decl-12192815273a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","order":7227,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean:13"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinCoordinateThreshold (logBudget varianceBudget : Real) : Real","missing":[],"search":"bernsteincoordinatethreshold banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteincoordinatethreshold noncomputable def bernsteincoordinatethreshold (logbudget variancebudget : real) : real definition compiled","shard":"modules/7af76fc4fbb4e66a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_pos","label":"bernsteinTilt_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_pos","description":"theorem bernsteinTilt_pos {logBudget varianceBudget : Real} (hlog : 0 < logBudget) (hvariance : 0 <= varianceBudget) : 0 < bernsteinTilt logBudget varianceBudget","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html#decl-b0ece69bb276","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","order":7228,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean:17"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bernsteinTilt_pos {logBudget varianceBudget : Real} (hlog : 0 < logBudget) (hvariance : 0 <= varianceBudget) : 0 < bernsteinTilt logBudget varianceBudget","missing":[],"search":"bernsteintilt_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteintilt_pos theorem bernsteintilt_pos {logbudget variancebudget : real} (hlog : 0 < logbudget) (hvariance : 0 <= variancebudget) : 0 < bernsteintilt logbudget variancebudget theorem compiled","shard":"modules/7af76fc4fbb4e66a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_le_one","label":"bernsteinTilt_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_le_one","description":"theorem bernsteinTilt_le_one (logBudget varianceBudget : Real) : bernsteinTilt logBudget varianceBudget <= 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html#decl-c69c01a473dc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","order":7229,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bernsteinTilt_le_one (logBudget varianceBudget : Real) : bernsteinTilt logBudget varianceBudget <= 1","missing":[],"search":"bernsteintilt_le_one banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteintilt_le_one theorem bernsteintilt_le_one (logbudget variancebudget : real) : bernsteintilt logbudget variancebudget <= 1 theorem compiled","shard":"modules/7af76fc4fbb4e66a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_abs_le_one","label":"bernsteinTilt_abs_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_abs_le_one","description":"theorem bernsteinTilt_abs_le_one {logBudget varianceBudget : Real} (hlog : 0 < logBudget) (hvariance : 0 <= varianceBudget) : |bernsteinTilt logBudget varianceBudget| <= 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html#decl-38bb83fd5023","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","order":7230,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bernsteinTilt_abs_le_one {logBudget varianceBudget : Real} (hlog : 0 < logBudget) (hvariance : 0 <= varianceBudget) : |bernsteinTilt logBudget varianceBudget| <= 1","missing":[],"search":"bernsteintilt_abs_le_one banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteintilt_abs_le_one theorem bernsteintilt_abs_le_one {logbudget variancebudget : real} (hlog : 0 < logbudget) (hvariance : 0 <= variancebudget) : |bernsteintilt logbudget variancebudget| <= 1 theorem compiled","shard":"modules/7af76fc4fbb4e66a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_exponent_le","label":"bernsteinTilt_exponent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_exponent_le","description":"theorem bernsteinTilt_exponent_le {logBudget varianceBudget : Real} (hlog : 0 < logBudget) (hvariance : 0 <= varianceBudget) : -bernsteinTilt logBudget varianceBudget * bernsteinCoordinateThreshold logBudget varianceBudget + bernsteinTilt logBudget varianceBudget ^ 2 * varianceBudget <= -2 * logBudget","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviconfidencetuning/index.html#decl-ff359b46ecbb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","order":7231,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bernsteinTilt_exponent_le {logBudget varianceBudget : Real} (hlog : 0 < logBudget) (hvariance : 0 <= varianceBudget) : -bernsteinTilt logBudget varianceBudget * bernsteinCoordinateThreshold logBudget varianceBudget + bernsteinTilt logBudget varianceBudget ^ 2 * varianceBudget <= -2 * logBudget","missing":[],"search":"bernsteintilt_exponent_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteintilt_exponent_le theorem bernsteintilt_exponent_le {logbudget variancebudget : real} (hlog : 0 < logbudget) (hvariance : 0 <= variancebudget) : -bernsteintilt logbudget variancebudget * bernsteincoordinatethreshold logbudget variancebudget + bernsteintilt logbudget variancebudget ^ 2 * variancebudget <= -2 * logbudget theorem compiled","shard":"modules/7af76fc4fbb4e66a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel_real_singleton_of_pos","label":"aggregateEmpiricalTransitionKernel_real_singleton_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel_real_singleton_of_pos","description":"At a positive pooled count, the singleton mass of the planner's empirical kernel is exactly the pooled numerator divided by the pooled denominator.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicoordinatealignment/index.html#decl-a6946c1cc17a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","order":7232,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateEmpiricalTransitionKernel_real_singleton_of_pos {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (state : State) (action : Action) (nextState : State) (hpos : 0 < summary.aggregateVisitCount state action) : (summary.aggregateEmpiricalTransitionKernel defaultState (state, action)).real {nextState} = (summary.aggregateTransitionCount state action nextState : Real) / (summary.aggregateVisitCount state action : Real)","missing":[],"search":"aggregateempiricaltransitionkernel_real_singleton_of_pos banditrlproof.finitehorizonrl.transitioncountsummary.aggregateempiricaltransitionkernel_real_singleton_of_pos at a positive pooled count, the singleton mass of the planner's empirical kernel is exactly the pooled numerator divided by the pooled denominator. theorem compiled","shard":"modules/7baff7bd1aabc720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinCoordinateThreshold_mul_probability_div_le","label":"bernsteinCoordinateThreshold_mul_probability_div_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinCoordinateThreshold_mul_probability_div_le","description":"theorem bernsteinCoordinateThreshold_mul_probability_div_le {probability logBudget visits : Real} (hprobability : probability ∈ Set.Icc (0 : Real) 1) (hlog : 0 <= logBudget) (hvisits : 0 < visits) : bernsteinCoordinateThreshold logBudget (probability * (1 - probability) * visits) / visits <= 2 * Real.sqrt (2 * logBudget / visits) * Real.sqrt probability + 2 * logBudget / visits","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicoordinatealignment/index.html#decl-8700f74eeba4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","order":7233,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bernsteinCoordinateThreshold_mul_probability_div_le {probability logBudget visits : Real} (hprobability : probability ∈ Set.Icc (0 : Real) 1) (hlog : 0 <= logBudget) (hvisits : 0 < visits) : bernsteinCoordinateThreshold logBudget (probability * (1 - probability) * visits) / visits <= 2 * Real.sqrt (2 * logBudget / visits) * Real.sqrt probability + 2 * logBudget / visits","missing":[],"search":"bernsteincoordinatethreshold_mul_probability_div_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteincoordinatethreshold_mul_probability_div_le theorem bernsteincoordinatethreshold_mul_probability_div_le {probability logbudget visits : real} (hprobability : probability ∈ set.icc (0 : real) 1) (hlog : 0 <= logbudget) (hvisits : 0 < visits) : bernsteincoordinatethreshold logbudget (probability * (1 - probability) * visits) / visits <= 2 * real.sqrt (2 * logbudget / visits) * real.sqrt probability + 2 * logbudget / visits theorem compiled","shard":"modules/7baff7bd1aabc720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.abs_sum_weight_mul_massError_le_transitionValue_div_thirtyTwo_add","label":"abs_sum_weight_mul_massError_le_transitionValue_div_thirtyTwo_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.abs_sum_weight_mul_massError_le_transitionValue_div_thirtyTwo_add","description":"A finite weighted Bernstein coordinate family controls a bounded continuation value. The small `z/(32H)` term is the self-bounding part used by the Bellman recursion; the remaining term is harmonic in the actual count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicoordinatealignment/index.html#decl-1087db46dec4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","order":7234,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment.lean:116"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_sum_weight_mul_massError_le_transitionValue_div_thirtyTwo_add (probability empirical value : State -> Real) (horizon logBudget visits : Real) (hhorizon : 1 <= horizon) (hlog : 0 <= logBudget) (hvisits : 0 < visits) (hprobability : forall state, 0 <= probability state) (hvalue : forall state, value state ∈ Set.Icc (0 : Real) horizon) (hcoordinate : forall state, |empirical state - probability state| <= 2 * Real.sqrt (2 * logBudget / visits) * Real.sqrt (probability state) + 2 * logBudget / visits) : |∑ state : State, value state * (empirical state - probability state)| <= (∑ state : State, value state * probability state) / (32 * horizon) + 66 * Fintype.card State * horizon ^ 2 * logBudget / visits","missing":[],"search":"abs_sum_weight_mul_masserror_le_transitionvalue_div_thirtytwo_add banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.abs_sum_weight_mul_masserror_le_transitionvalue_div_thirtytwo_add a finite weighted bernstein coordinate family controls a bounded continuation value. the small `z/(32h)` term is the self-bounding part used by the bellman recursion; the remaining term is harmonic in the actual count. theorem compiled","shard":"modules/7baff7bd1aabc720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransitionMass_sub_lt_bernstein","label":"abs_empiricalTransitionMass_sub_lt_bernstein","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransitionMass_sub_lt_bernstein","description":"The simultaneous singleton event controls the exact empirical mass error at the actual positive count. No expected count or auxiliary sample appears.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicoordinatealignment/index.html#decl-6b368305403e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","order":7235,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment.lean:286"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_empiricalTransitionMass_sub_lt_bernstein {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {logBudget : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes logBudget) (defaultState : State) (index : BernsteinCoordinateIndex mdp episodes) (hactual : adaptiveCumulativeAggregateVisitCountAt trajectory index.round index.state index.action = index.count + 1) : |((adaptiveCumulativeEmpiricalModelStateAt trajectory index.round).1 |>.aggregateEmpiricalTransitionKernel defaultState (index.state, index.action)).real {index.nextState} - (mdp.transition (index.state, index.action)).real {index.nextState}| < bernsteinCoordinateThreshold logBudget (mdp.transitionCoordinateVariance index.state index.…","missing":[],"search":"abs_empiricaltransitionmass_sub_lt_bernstein banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.abs_empiricaltransitionmass_sub_lt_bernstein the simultaneous singleton event controls the exact empirical mass error at the actual positive count. no expected count or auxiliary sample appears. theorem compiled","shard":"modules/7baff7bd1aabc720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransition_integral_sub_le_transitionValue_div_thirtyTwo_add","label":"abs_empiricalTransition_integral_sub_le_transitionValue_div_thirtyTwo_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransition_integral_sub_le_transitionValue_div_thirtyTwo_add","description":"The exact generated singleton family yields the self-bounding transition value estimate used in the UCBVI recursion.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicoordinatealignment/index.html#decl-a7192ee6d7c5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","order":7236,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment.lean:354"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_empiricalTransition_integral_sub_le_transitionValue_div_thirtyTwo_add {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {logBudget : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes logBudget) (defaultState : State) (round : Fin episodes) (state : State) (action : Action) (count : Fin (episodes * mdp.horizon)) (hactual : adaptiveCumulativeAggregateVisitCountAt trajectory round state action = count + 1) (value : State -> Real) (hvalue : forall nextState, value nextState ∈ Set.Icc (0 : Real) mdp.horizon) (hhorizon : 0 < mdp.horizon) (hlog : 0 <= logBudget) : |(∫ nextState, value nextState ∂TransitionCountSummary.aggregateEmpiricalTransitionKernel (adaptiveCumulativeEmpiricalModelState…","missing":[],"search":"abs_empiricaltransition_integral_sub_le_transitionvalue_div_thirtytwo_add banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.abs_empiricaltransition_integral_sub_le_transitionvalue_div_thirtytwo_add the exact generated singleton family yields the self-bounding transition value estimate used in the ucbvi recursion. theorem compiled","shard":"modules/7baff7bd1aabc720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount","label":"batchedPrefixCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount","description":"Strict-prefix cumulative mass of a batched nonnegative increment stream.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-d2720c031440","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7237,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:18"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def batchedPrefixCount (increment : Nat -> Nat) (round : Nat) : Nat","missing":[],"search":"batchedprefixcount banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batchedprefixcount strict-prefix cumulative mass of a batched nonnegative increment stream. definition compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_zero","label":"batchedPrefixCount_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_zero","description":"@[simp] theorem batchedPrefixCount_zero (increment : Nat -> Nat) : batchedPrefixCount increment 0 = 0","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-65bf73802f13","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7238,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:21"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"@[simp] theorem batchedPrefixCount_zero (increment : Nat -> Nat) : batchedPrefixCount increment 0 = 0","missing":[],"search":"batchedprefixcount_zero banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batchedprefixcount_zero @[simp] theorem batchedprefixcount_zero (increment : nat -> nat) : batchedprefixcount increment 0 = 0 theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_succ","label":"batchedPrefixCount_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_succ","description":"theorem batchedPrefixCount_succ (increment : Nat -> Nat) (round : Nat) : batchedPrefixCount increment (round + 1) = batchedPrefixCount increment round + increment round","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-932da9c4a985","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7239,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem batchedPrefixCount_succ (increment : Nat -> Nat) (round : Nat) : batchedPrefixCount increment (round + 1) = batchedPrefixCount increment round + increment round","missing":[],"search":"batchedprefixcount_succ banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batchedprefixcount_succ theorem batchedprefixcount_succ (increment : nat -> nat) (round : nat) : batchedprefixcount increment (round + 1) = batchedprefixcount increment round + increment round theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_mono","label":"batchedPrefixCount_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_mono","description":"theorem batchedPrefixCount_mono (increment : Nat -> Nat) : Monotone (batchedPrefixCount increment)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-f937970b9c55","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7240,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem batchedPrefixCount_mono (increment : Nat -> Nat) : Monotone (batchedPrefixCount increment)","missing":[],"search":"batchedprefixcount_mono banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batchedprefixcount_mono theorem batchedprefixcount_mono (increment : nat -> nat) : monotone (batchedprefixcount increment) theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_range_forwardDifference","label":"sum_range_forwardDifference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_range_forwardDifference","description":"private theorem sum_range_forwardDifference (f : Nat -> Real) (rounds : Nat) : (∑ round ∈ Finset.range rounds, (f (round + 1) - f round)) = f rounds - f 0","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-f29b082509e7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7241,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem sum_range_forwardDifference (f : Nat -> Real) (rounds : Nat) : (∑ round ∈ Finset.range rounds, (f (round + 1) - f round)) = f rounds - f 0","missing":[],"search":"sum_range_forwarddifference banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_range_forwarddifference private theorem sum_range_forwarddifference (f : nat -> real) (rounds : nat) : (∑ round ∈ finset.range rounds, (f (round + 1) - f round)) = f rounds - f 0 theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_low_batchedIncrement_le","label":"sum_low_batchedIncrement_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_low_batchedIncrement_le","description":"All increments whose strict-prefix mass is below a threshold occupy at most the threshold plus one batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-b4e6b9c57d66","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7242,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_low_batchedIncrement_le (increment : Nat -> Nat) (batch threshold rounds : Nat) (hincrement : forall round, increment round <= batch) : (∑ round ∈ Finset.range rounds, if batchedPrefixCount increment round < threshold then increment round else 0) <= threshold + batch","missing":[],"search":"sum_low_batchedincrement_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_low_batchedincrement_le all increments whose strict-prefix mass is below a threshold occupy at most the threshold plus one batch. theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batched_ratio_le_two_log_increment","label":"batched_ratio_le_two_log_increment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batched_ratio_le_two_log_increment","description":"private theorem batched_ratio_le_two_log_increment {mass increment : Nat} (hmass : 0 < mass) (hincrement : increment <= mass) : (increment : Real) / mass <= 2 * (Real.log (mass + increment : Nat) - Real.log mass)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-502df9aabcd5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7243,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:74"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem batched_ratio_le_two_log_increment {mass increment : Nat} (hmass : 0 < mass) (hincrement : increment <= mass) : (increment : Real) / mass <= 2 * (Real.log (mass + increment : Nat) - Real.log mass)","missing":[],"search":"batched_ratio_le_two_log_increment banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batched_ratio_le_two_log_increment private theorem batched_ratio_le_two_log_increment {mass increment : nat} (hmass : 0 < mass) (hincrement : increment <= mass) : (increment : real) / mass <= 2 * (real.log (mass + increment : nat) - real.log mass) theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_high_batchedRatio_le_log","label":"sum_high_batchedRatio_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_high_batchedRatio_le_log","description":"High-prefix reciprocal masses telescope into one logarithm.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-7459c93df0b7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7244,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_high_batchedRatio_le_log (increment : Nat -> Nat) (batch threshold rounds : Nat) (hbatch : 0 < batch) (hthreshold : batch <= threshold) (hincrement : forall round, increment round <= batch) : (∑ round ∈ Finset.range rounds, if threshold <= batchedPrefixCount increment round then (increment round : Real) / batchedPrefixCount increment round else 0) <= 2 * Real.log (max 1 (batchedPrefixCount increment rounds) : Nat)","missing":[],"search":"sum_high_batchedratio_le_log banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_high_batchedratio_le_log high-prefix reciprocal masses telescope into one logarithm. theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batched_inverseSqrt_step","label":"batched_inverseSqrt_step","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batched_inverseSqrt_step","description":"private theorem batched_inverseSqrt_step {mass increment batch : Nat} (hmass : 0 < mass) (hincrement : increment <= batch) : (increment : Real) / Real.sqrt mass <= 2 * (Real.sqrt (mass + increment : Nat) - Real.sqrt mass) + (batch : Real) * ((increment : Real) / mass)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-b3b75b442c37","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7245,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem batched_inverseSqrt_step {mass increment batch : Nat} (hmass : 0 < mass) (hincrement : increment <= batch) : (increment : Real) / Real.sqrt mass <= 2 * (Real.sqrt (mass + increment : Nat) - Real.sqrt mass) + (batch : Real) * ((increment : Real) / mass)","missing":[],"search":"batched_inversesqrt_step banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.batched_inversesqrt_step private theorem batched_inversesqrt_step {mass increment batch : nat} (hmass : 0 < mass) (hincrement : increment <= batch) : (increment : real) / real.sqrt mass <= 2 * (real.sqrt (mass + increment : nat) - real.sqrt mass) + (batch : real) * ((increment : real) / mass) theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedInvSqrt_le","label":"sum_positive_batchedInvSqrt_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedInvSqrt_le","description":"Batched inverse-square-root masses retain the leading constant two; all within-episode staleness is isolated in one low-count batch and one logarithm.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-d471ee75dc51","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7246,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:218"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_positive_batchedInvSqrt_le (increment : Nat -> Nat) (batch rounds : Nat) (hbatch : 0 < batch) (hincrement : forall round, increment round <= batch) : (∑ round ∈ Finset.range rounds, if batchedPrefixCount increment round = 0 then 0 else (increment round : Real) / Real.sqrt (batchedPrefixCount increment round)) <= 2 * Real.sqrt (batchedPrefixCount increment rounds) + 2 * batch + 2 * batch * Real.log ((max 1 (batchedPrefixCount increment rounds) : Nat) : Real)","missing":[],"search":"sum_positive_batchedinvsqrt_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_positive_batchedinvsqrt_le batched inverse-square-root masses retain the leading constant two; all within-episode staleness is isolated in one low-count batch and one logarithm. theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedMinReciprocal_le","label":"sum_positive_batchedMinReciprocal_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedMinReciprocal_le","description":"A reciprocal correction clipped at one horizon pays the threshold region once (plus one stale batch) and then telescopes logarithmically.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvicounting/index.html#decl-14b2f5ccd670","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","order":7247,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVICounting.lean:396"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_positive_batchedMinReciprocal_le (increment : Nat -> Nat) (batch threshold rounds : Nat) (cap coefficient : Real) (hbatch : 0 < batch) (hthreshold : batch <= threshold) (hincrement : forall round, increment round <= batch) (hcap : 0 <= cap) (hcoefficient : 0 <= coefficient) : (∑ round ∈ Finset.range rounds, if batchedPrefixCount increment round = 0 then 0 else (increment round : Real) * min cap (coefficient / batchedPrefixCount increment round)) <= cap * (threshold + batch) + 2 * coefficient * Real.log ((max 1 (batchedPrefixCount increment rounds) : Nat) : Real)","missing":[],"search":"sum_positive_batchedminreciprocal_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_positive_batchedminreciprocal_le a reciprocal correction clipped at one horizon pays the threshold region once (plus one stale batch) and then telescopes logarithmically. theorem compiled","shard":"modules/9d51a190447b4fc3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_nonneg_of_reward_nonneg","label":"valueRemaining_nonneg_of_reward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_nonneg_of_reward_nonneg","description":"Nonnegative deterministic rewards give nonnegative finite-horizon policy values.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-c059f43d711c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7248,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:22"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueRemaining_nonneg_of_reward_nonneg {mdp : MDP State Action} (policy : MarkovPolicy mdp) (hreward : ∀ state action, 0 <= mdp.reward state action) : ∀ remaining (hremaining : remaining <= mdp.horizon) state, 0 <= policy.valueRemaining remaining hremaining state","missing":[],"search":"valueremaining_nonneg_of_reward_nonneg banditrlproof.finitehorizonrl.markovpolicy.valueremaining_nonneg_of_reward_nonneg nonnegative deterministic rewards give nonnegative finite-horizon policy values. theorem compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedQTable","label":"generatedQTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedQTable","description":"The recurrent Q table used at generated episode coordinate `episode`: exactly the strict prefix `0,...,episode-1` is folded.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-462fd5eaebab","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7249,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def generatedQTable (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (episode : Nat) : QTable mdp","missing":[],"search":"generatedqtable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedqtable the recurrent q table used at generated episode coordinate `episode`: exactly the strict prefix `0,...,episode-1` is folded. definition compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPolicyTable","label":"generatedPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPolicyTable","description":"The deterministic argmax table of that exact generated Q table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-02e96f7d14f0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7250,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def generatedPolicyTable (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (episode : Nat) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"generatedpolicytable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedpolicytable the deterministic argmax table of that exact generated q table. definition compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_eq_generatedPolicyTable","label":"recurrentSource_policyAt_eq_generatedPolicyTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_eq_generatedPolicyTable","description":"The canonical source policy at every coordinate is definitionally the argmax of `generatedQTable` built from its strict prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-5fd7c3bb83a1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7251,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_policyAt_eq_generatedPolicyTable (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (episode : Nat) : (recurrentSource mdp initialState defaultState episodes delta).policyAt trajectory episode = DeterministicMarkovPolicyTable.toMarkovPolicy (generatedPolicyTable mdp defaultState episodes delta trajectory episode)","missing":[],"search":"recurrentsource_policyat_eq_generatedpolicytable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_policyat_eq_generatedpolicytable the canonical source policy at every coordinate is definitionally the argmax of `generatedqtable` built from its strict prefix. theorem compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodeInitialState","label":"generatedEpisodeInitialState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodeInitialState","description":"Totalized initial state of one generated single-episode batch. On the canonical positive-horizon domain it is literally stage zero's state.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-0964b0c71beb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7252,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def generatedEpisodeInitialState (mdp : MDP State Action) (defaultState : State) (batch : EpisodeBatch mdp 1) : State","missing":[],"search":"generatedepisodeinitialstate banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedepisodeinitialstate totalized initial state of one generated single-episode batch. on the canonical positive-horizon domain it is literally stage zero's state. definition compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodePseudoRegret","label":"generatedEpisodePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodePseudoRegret","description":"Pathwise policy-value pseudo-regret of one generated coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-59c291a76763","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7253,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def generatedEpisodePseudoRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (defaultState : State) (trajectory : EpisodeBatchTrajectory mdp 1) (episode : Nat) : Real","missing":[],"search":"generatedepisodepseudoregret banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedepisodepseudoregret pathwise policy-value pseudo-regret of one generated coordinate. definition compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret","label":"cumulativeEpisodePseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret","description":"Raw generated cumulative episode pseudo-regret over exactly coordinates `0,...,K-1`; coordinate zero is included and never hidden.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-fdd1a4a585c7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7254,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeEpisodePseudoRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (defaultState : State) (episodes : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"cumulativeepisodepseudoregret banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.cumulativeepisodepseudoregret raw generated cumulative episode pseudo-regret over exactly coordinates `0,...,k-1`; coordinate zero is included and never hidden. definition compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodePseudoRegret_mem_Icc","label":"generatedEpisodePseudoRegret_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodePseudoRegret_mem_Icc","description":"Every generated policy-value pseudo-regret lies in `[0,H]` under rewards in `[0,1]`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-40a7308a3bcd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7255,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem generatedEpisodePseudoRegret_mem_Icc {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (defaultState : State) (trajectory : EpisodeBatchTrajectory mdp 1) (episode : Nat) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) : generatedEpisodePseudoRegret source defaultState trajectory episode ∈ Set.Icc (0 : Real) mdp.horizon","missing":[],"search":"generatedepisodepseudoregret_mem_icc banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.generatedepisodepseudoregret_mem_icc every generated policy-value pseudo-regret lies in `[0,h]` under rewards in `[0,1]`. theorem compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret_mem_Icc","label":"cumulativeEpisodePseudoRegret_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret_mem_Icc","description":"Consequently the exact `K`-episode raw pseudo-regret lies in `[0,K H]`; this is the envelope later used only on the terminal failure event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviepisoderegret/index.html#decl-e7f78c2fe58b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","order":7256,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeEpisodePseudoRegret_mem_Icc {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (defaultState : State) (episodes : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) : cumulativeEpisodePseudoRegret source defaultState episodes trajectory ∈ Set.Icc (0 : Real) ((episodes : Real) * mdp.horizon)","missing":[],"search":"cumulativeepisodepseudoregret_mem_icc banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.cumulativeepisodepseudoregret_mem_icc consequently the exact `k`-episode raw pseudo-regret lies in `[0,k h]`; this is the envelope later used only on the terminal failure event. theorem compiled","shard":"modules/772b4376bbc41857.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedPolicyTable","label":"measurable_generatedPolicyTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedPolicyTable","description":"theorem measurable_generatedPolicyTable (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (episode : Nat) : Measurable (fun trajectory : EpisodeBatchTrajectory mdp 1 => generatedPolicyTable mdp defaultState episodes delta trajectory episode)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviexpectedregret/index.html#decl-4fadb597da3a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","order":7257,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedPolicyTable (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (episode : Nat) : Measurable (fun trajectory : EpisodeBatchTrajectory mdp 1 => generatedPolicyTable mdp defaultState episodes delta trajectory episode)","missing":[],"search":"measurable_generatedpolicytable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_generatedpolicytable theorem measurable_generatedpolicytable (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (episode : nat) : measurable (fun trajectory : episodebatchtrajectory mdp 1 => generatedpolicytable mdp defaultstate episodes delta trajectory episode) theorem compiled","shard":"modules/6941bfb05ef21651.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedEpisodeInitialState","label":"measurable_generatedEpisodeInitialState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedEpisodeInitialState","description":"theorem measurable_generatedEpisodeInitialState (mdp : MDP State Action) (defaultState : State) : Measurable (generatedEpisodeInitialState mdp defaultState)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviexpectedregret/index.html#decl-d028dfc70195","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","order":7258,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret.lean:48"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedEpisodeInitialState (mdp : MDP State Action) (defaultState : State) : Measurable (generatedEpisodeInitialState mdp defaultState)","missing":[],"search":"measurable_generatedepisodeinitialstate banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_generatedepisodeinitialstate theorem measurable_generatedepisodeinitialstate (mdp : mdp state action) (defaultstate : state) : measurable (generatedepisodeinitialstate mdp defaultstate) theorem compiled","shard":"modules/6941bfb05ef21651.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedEpisodePseudoRegret","label":"measurable_generatedEpisodePseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedEpisodePseudoRegret","description":"theorem measurable_generatedEpisodePseudoRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (episode : Nat) : Measurable (fun trajectory : EpisodeBatchTrajectory mdp 1 => generatedEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState trajectory episode)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviexpectedregret/index.html#decl-bc2ff3dfb4f9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","order":7259,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_generatedEpisodePseudoRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (episode : Nat) : Measurable (fun trajectory : EpisodeBatchTrajectory mdp 1 => generatedEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState trajectory episode)","missing":[],"search":"measurable_generatedepisodepseudoregret banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_generatedepisodepseudoregret theorem measurable_generatedepisodepseudoregret (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (defaultstate : state) (episodes : nat) (delta : real) (episode : nat) : measurable (fun trajectory : episodebatchtrajectory mdp 1 => generatedepisodepseudoregret (recurrentsource mdp initialstate defaultstate episodes delta) defaultstate trajectory episode) theorem compiled","shard":"modules/6941bfb05ef21651.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_cumulativeEpisodePseudoRegret","label":"measurable_cumulativeEpisodePseudoRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_cumulativeEpisodePseudoRegret","description":"theorem measurable_cumulativeEpisodePseudoRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) : Measurable (fun trajectory : EpisodeBatchTrajectory mdp 1 => cumulativeEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState episodes trajectory)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviexpectedregret/index.html#decl-2a71cad187b4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","order":7260,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret.lean:89"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeEpisodePseudoRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) : Measurable (fun trajectory : EpisodeBatchTrajectory mdp 1 => cumulativeEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState episodes trajectory)","missing":[],"search":"measurable_cumulativeepisodepseudoregret banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.measurable_cumulativeepisodepseudoregret theorem measurable_cumulativeepisodepseudoregret (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (defaultstate : state) (episodes : nat) (delta : real) : measurable (fun trajectory : episodebatchtrajectory mdp 1 => cumulativeepisodepseudoregret (recurrentsource mdp initialstate defaultstate episodes delta) defaultstate episodes trajectory) theorem compiled","shard":"modules/6941bfb05ef21651.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure","label":"integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure","description":"Canonical finite-time expected pseudo-regret. The `K H delta` summand is the explicit contribution of the proved terminal failure event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviexpectedregret/index.html#decl-401539f24394","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","order":7261,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","finite-horizon-rl"]],"statement":"theorem integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) : let source := recurrentSource mdp initialState defaultState episodes delta ∫ trajectory, cumulativeEpisodePseudoRegret source defaultState episodes trajectory ∂source.trajectoryMeasure <= canonicalRegretBound (State := State) (Action := Action) mdp episodes delta + (episodes : Real) * mdp.horizon * delta","missing":[],"search":"integral_cumulativeepisodepseudoregret_recurrentsource_le_canonicalregretbound_add_failure banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.integral_cumulativeepisodepseudoregret_recurrentsource_le_canonicalregretbound_add_failure canonical finite-time expected pseudo-regret. the `k h delta` summand is the explicit contribution of the proved terminal failure event. theorem compiled","shard":"modules/6941bfb05ef21651.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":["finite-horizon-rl"]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_eq_of_eq","label":"optimalValueAt_eq_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_eq_of_eq","description":"Transport the proof-indexed chronological optimal value across a stage equality.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-a39e19223dcf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7262,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueAt_eq_of_eq (mdp : MDP State Action) {left right : Nat} (hleft : left <= mdp.horizon) (hright : right <= mdp.horizon) (h : left = right) : mdp.optimalValueAt left hleft = mdp.optimalValueAt right hright","missing":[],"search":"optimalvalueat_eq_of_eq banditrlproof.finitehorizonrl.mdp.optimalvalueat_eq_of_eq transport the proof-indexed chronological optimal value across a stage equality. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_nonneg_of_reward_nonneg","label":"optimalValueAt_nonneg_of_reward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_nonneg_of_reward_nonneg","description":"Nonnegative rewards give a nonnegative optimal finite-horizon value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-fb082a3b8511","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7263,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueAt_nonneg_of_reward_nonneg (mdp : MDP State Action) (hreward : forall state action, 0 <= mdp.reward state action) (stage : Nat) (hstage : stage <= mdp.horizon) (state : State) : 0 <= mdp.optimalValueAt stage hstage state","missing":[],"search":"optimalvalueat_nonneg_of_reward_nonneg banditrlproof.finitehorizonrl.mdp.optimalvalueat_nonneg_of_reward_nonneg nonnegative rewards give a nonnegative optimal finite-horizon value. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.decisionStageRemaining_succ_eq","label":"decisionStageRemaining_succ_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.decisionStageRemaining_succ_eq","description":"theorem decisionStageRemaining_succ_eq (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : (mdp.decisionStageRemaining remaining hremaining : Nat) + 1 = mdp.horizon - remaining","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-25a302ce8c8f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7264,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decisionStageRemaining_succ_eq (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : (mdp.decisionStageRemaining remaining hremaining : Nat) + 1 = mdp.horizon - remaining","missing":[],"search":"decisionstageremaining_succ_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.decisionstageremaining_succ_eq theorem decisionstageremaining_succ_eq (mdp : mdp state action) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) : (mdp.decisionstageremaining remaining hremaining : nat) + 1 = mdp.horizon - remaining theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_eq_of_eq","label":"clippedQRemaining_eq_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_eq_of_eq","description":"Transport the proof-indexed clipped Q surface across equality of the remaining-horizon coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-21d743864013","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7265,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedQRemaining_eq_of_eq {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) {left right : Nat} (hleft : left + 1 <= mdp.horizon) (hright : right + 1 <= mdp.horizon) (h : left = right) : clippedQRemaining previousQ summary defaultState bonusScale left hleft = clippedQRemaining previousQ summary defaultState bonusScale right hright","missing":[],"search":"clippedqremaining_eq_of_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqremaining_eq_of_eq transport the proof-indexed clipped q surface across equality of the remaining-horizon coordinate. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_succ_eq_selectedQ","label":"clippedValueRemaining_succ_eq_selectedQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_succ_eq_selectedQ","description":"The selected value in the backward recursion is the selected entry of the chronological clipped Q table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-a32a770d4821","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7266,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:74"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedValueRemaining_succ_eq_selectedQ {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : clippedValueRemaining previousQ summary defaultState bonusScale (remaining + 1) hremaining state = clippedQRemaining previousQ summary defaultState bonusScale remaining hremaining state (clippedPolicyTable mdp previousQ summary defaultState bonusScale (mdp.decisionStageRemaining remaining hremaining) state)","missing":[],"search":"clippedvalueremaining_succ_eq_selectedq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedvalueremaining_succ_eq_selectedq the selected value in the backward recursion is the selected entry of the chronological clipped q table. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_le_horizon","label":"clippedQTable_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_le_horizon","description":"Every entry of one clipped Q update is at most `H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-455f818689e8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7267,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:117"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedQTable_le_horizon (mdp : MDP State Action) (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (stage : Fin mdp.horizon) (state : State) (action : Action) : clippedQTable mdp previousQ summary defaultState bonusScale stage state action <= (mdp.horizon : Real)","missing":[],"search":"clippedqtable_le_horizon banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqtable_le_horizon every entry of one clipped q update is at most `h`. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_le_horizon","label":"clippedValueRemaining_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_le_horizon","description":"Clipping makes every selected upper value at most `H`, independently of whether the statistical event holds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-21011efd82e8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7268,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:133"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedValueRemaining_le_horizon {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : clippedValueRemaining previousQ summary defaultState bonusScale remaining hremaining state <= (mdp.horizon : Real)","missing":[],"search":"clippedvalueremaining_le_horizon banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedvalueremaining_le_horizon clipping makes every selected upper value at most `h`, independently of whether the statistical event holds. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalValueAt_le_clippedValueRemaining","label":"optimalValueAt_le_clippedValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalValueAt_le_clippedValueRemaining","description":"Pointwise optimal-Q dominance implies that the selected clipped value dominates the optimal value at the matching chronological stage.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-99e50885dd42","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7269,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:152"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueAt_le_clippedValueRemaining {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (hdominates : QDominatesOptimal mdp (clippedQTable mdp previousQ summary defaultState bonusScale)) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : mdp.optimalValueAt (mdp.horizon - remaining) (Nat.sub_le _ _ ) state <= clippedValueRemaining previousQ summary defaultState bonusScale remaining hremaining state","missing":[],"search":"optimalvalueat_le_clippedvalueremaining banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.optimalvalueat_le_clippedvalueremaining pointwise optimal-q dominance implies that the selected clipped value dominates the optimal value at the matching chronological stage. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_sub_policyValueRemaining_le_of_pos","label":"clippedValueRemaining_sub_policyValueRemaining_le_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_sub_policyValueRemaining_le_of_pos","description":"At a positive actual count, the selected clipped value minus the selected policy value propagates through the true transition kernel, plus exactly the empirical-model error and the configured bonus.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-7706be3c1afd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7270,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:216"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedValueRemaining_sub_policyValueRemaining_le_of_pos {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale modelError : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (hpos : 0 < summary.aggregateVisitCount state (clippedPolicyTable mdp previousQ summary defaultState bonusScale (mdp.decisionStageRemaining remaining hremaining) state)) (hmodel : (∫ nextState, clippedValueRemaining previousQ summary defaultState bonusScale remaining (by omega) nextState ∂summary.aggregateEmpiricalTransitionKernel defaultState (state, clippedPolicyTable mdp previousQ summary defaultState bonusScale (mdp.decisionStageRemaining remaining hremaining) state)) - mdp.transitionValue (clippedValueRemaining previousQ summary defaultState bonusScale remaining (by omega)) state (clippedPolicyTable m…","missing":[],"search":"clippedvalueremaining_sub_policyvalueremaining_le_of_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedvalueremaining_sub_policyvalueremaining_le_of_pos at a positive actual count, the selected clipped value minus the selected policy value propagates through the true transition kernel, plus exactly the empirical-model error and the configured bonus. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyGapRemaining","label":"clippedPolicyGapRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyGapRemaining","description":"Bellman gap between one clipped recurrent value surface and the policy selected by that same surface. Naming this exact difference keeps the generated-source recursion readable without hiding either policy identity.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-8411ddddd20c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7271,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:318"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedPolicyGapRemaining {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : Real","missing":[],"search":"clippedpolicygapremaining banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedpolicygapremaining bellman gap between one clipped recurrent value surface and the policy selected by that same surface. naming this exact difference keeps the generated-source recursion readable without hiding either policy identity. definition compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyGapRemaining_le_of_pos","label":"clippedPolicyGapRemaining_le_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyGapRemaining_le_of_pos","description":"The deterministic positive-count recursion, expressed through the named same-policy gap.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-8eb87aeb05a7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7272,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:330"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedPolicyGapRemaining_le_of_pos {mdp : MDP State Action} (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale modelError : Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (hpos : 0 < summary.aggregateVisitCount state (clippedPolicyTable mdp previousQ summary defaultState bonusScale (mdp.decisionStageRemaining remaining hremaining) state)) (hmodel : (∫ nextState, clippedValueRemaining previousQ summary defaultState bonusScale remaining (by omega) nextState ∂summary.aggregateEmpiricalTransitionKernel defaultState (state, clippedPolicyTable mdp previousQ summary defaultState bonusScale (mdp.decisionStageRemaining remaining hremaining) state)) - mdp.transitionValue (clippedValueRemaining previousQ summary defaultState bonusScale remaining (by omega)) state (clippedPolicyTable mdp previousQ summary…","missing":[],"search":"clippedpolicygapremaining_le_of_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedpolicygapremaining_le_of_pos the deterministic positive-count recursion, expressed through the named same-policy gap. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.selectedUpperTransitionModelError_le","label":"selectedUpperTransitionModelError_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.selectedUpperTransitionModelError_le","description":"On the same generated transition event, the model error of the selected clipped continuation is self-bounded by the true propagated policy-value gap. The other two terms are the sharp optimal-tail coordinate and the harmonic Bernstein correction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-9229ba988144","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7273,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:374"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selectedUpperTransitionModelError_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {logBudget : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes logBudget) (defaultState : State) (round : Fin episodes) (previousQ : QTable mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (count : Fin (episodes * mdp.horizon)) (hactual : adaptiveCumulativeAggregateVisitCountAt trajectory round state (clippedPolicyTable mdp previousQ (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1 defaultState (scale (State := State) (Action := Action) mdp episodes delta) (mdp.decisionStageRemaining remaining hremaining) state) = count + 1) (hdominates : QDominatesOptimal…","missing":[],"search":"selecteduppertransitionmodelerror_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.selecteduppertransitionmodelerror_le on the same generated transition event, the model error of the selected clipped continuation is self-bounded by the true propagated policy-value gap. the other two terms are the sharp optimal-tail coordinate and the harmonic bernstein correction. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.selectedPolicyGap_le_of_actual_count","label":"selectedPolicyGap_le_of_actual_count","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.selectedPolicyGap_le_of_actual_count","description":"Complete local UCBVI-CH recursion at an actual positive generated count. The `7HL` planner bonus and the `2HL` sharp transition-value deviation combine to `9HL`; the remaining coordinate term is self-bounded by the propagated same-policy gap plus the explicit harmonic correction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvilocalbellman/index.html#decl-d39082c57d3b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","order":7274,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVILocalBellman.lean:590"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selectedPolicyGap_le_of_actual_count {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (round : Fin episodes) (previousQ : QTable mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (count : Fin (episodes * mdp.horizon)) (hactual : adaptiveCumulativeAggregateVisitCountAt trajectory round state (clippedPolicyTable mdp previousQ (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1 defaultState (scale (State := State) (Action := Action) mdp episodes delta) (mdp.decisionStageRemaining remaining hremaining)…","missing":[],"search":"selectedpolicygap_le_of_actual_count banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.selectedpolicygap_le_of_actual_count complete local ucbvi-ch recursion at an actual positive generated count. the `7hl` planner bonus and the `2hl` sharp transition-value deviation combine to `9hl`; the remaining coordinate term is self-bounded by the propagated same-policy gap plus the explicit harmonic correction. theorem compiled","shard":"modules/6b56a3a95d9a8545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationThreshold","label":"bellmanInnovationThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationThreshold","description":"Bellman-martingale charge used in the frozen `20/250` terminal.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html#decl-a501b10d3c67","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","order":7275,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean:22"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellmanInnovationThreshold (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Real","missing":[],"search":"bellmaninnovationthreshold banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninnovationthreshold bellman-martingale charge used in the frozen `20/250` terminal. definition compiled","shard":"modules/dc864b477258d0b5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt","label":"bellmanInnovationTilt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt","description":"Chernoff tilt optimized for the deterministic `K H^3` variance budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html#decl-78561313ba5e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","order":7276,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellmanInnovationTilt (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Real","missing":[],"search":"bellmaninnovationtilt banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninnovationtilt chernoff tilt optimized for the deterministic `k h^3` variance budget. definition compiled","shard":"modules/dc864b477258d0b5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationThreshold_nonneg","label":"bellmanInnovationThreshold_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationThreshold_nonneg","description":"theorem bellmanInnovationThreshold_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= bellmanInnovationThreshold (State := State) (Action := Action) mdp episodes delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html#decl-961677bdd454","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","order":7277,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanInnovationThreshold_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= bellmanInnovationThreshold (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"bellmaninnovationthreshold_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninnovationthreshold_nonneg theorem bellmaninnovationthreshold_nonneg (mdp : mdp state action) (episodes : nat) (delta : real) : 0 <= bellmaninnovationthreshold (state := state) (action := action) mdp episodes delta theorem compiled","shard":"modules/dc864b477258d0b5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt_pos","label":"bellmanInnovationTilt_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt_pos","description":"theorem bellmanInnovationTilt_pos (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < bellmanInnovationTilt (State := State) (Action := Action) mdp episodes delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html#decl-545d0f92f393","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","order":7278,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanInnovationTilt_pos (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < bellmanInnovationTilt (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"bellmaninnovationtilt_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninnovationtilt_pos theorem bellmaninnovationtilt_pos (mdp : mdp state action) (episodes : nat) (delta : real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : 0 < bellmaninnovationtilt (state := state) (action := action) mdp episodes delta theorem compiled","shard":"modules/dc864b477258d0b5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt_exponent_le_neg_two_logFactor","label":"bellmanInnovationTilt_exponent_le_neg_two_logFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt_exponent_le_neg_two_logFactor","description":"theorem bellmanInnovationTilt_exponent_le_neg_two_logFactor (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : -bellmanInnovationTilt (State := State) (Action := Action) mdp episodes delta * bellmanInnovationThreshold (State := State) (Action := Action) mdp episodes delta + (bellmanInnovationTilt (State…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html#decl-1ba493fccd40","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","order":7279,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanInnovationTilt_exponent_le_neg_two_logFactor (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : -bellmanInnovationTilt (State := State) (Action := Action) mdp episodes delta * bellmanInnovationThreshold (State := State) (Action := Action) mdp episodes delta + (bellmanInnovationTilt (State := State) (Action := Action) mdp episodes delta ^ 2 / 8) * ((episodes : Real) * ((mdp.horizon : Real) * (mdp.horizon : Real) ^ 2)) <= -2 * logFactor (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"bellmaninnovationtilt_exponent_le_neg_two_logfactor banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bellmaninnovationtilt_exponent_le_neg_two_logfactor theorem bellmaninnovationtilt_exponent_le_neg_two_logfactor (mdp : mdp state action) (episodes : nat) (delta : real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : -bellmaninnovationtilt (state := state) (action := action) mdp episodes delta * bellmaninnovationthreshold (state := state) (action := action) mdp episodes delta + (bellmaninnovationtilt (state := state) (action := action) mdp episodes delta ^ 2 / 8) * ((episodes : real) * ((mdp.horizon : real) * (mdp.horizon : real) ^ 2)) <= -2 * logfactor (state := state) (action := action) mdp episodes delta theorem compiled","shard":"modules/dc864b477258d0b5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","label":"recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","description":"The tuned Bellman-innovation tail consumes at most one fifth of `delta`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvimartingaletuning/index.html#decl-42ba8e82c022","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","order":7280,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let source := recurrentSource mdp initialState defaultState episodes delta source.trajectoryMeasure {trajectory | bellmanInnovationThreshold (State := State) (Action := Action) mdp episodes delta <= (Finset.range episodes).sum (fun round => recurrentBellmanInnovationProcess mdp defaultState episodes delta round trajectory)} <= ENNReal.ofReal (delta / 5)","missing":[],"search":"recurrentsource_trajectorymeasure_bellmaninnovation_sum_ge_threshold_le_fifth banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_bellmaninnovation_sum_ge_threshold_le_fifth the tuned bellman-innovation tail consumes at most one fifth of `delta`. theorem compiled","shard":"modules/dc864b477258d0b5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_optimalValue_sub_eq_two_mul_horizon_mul_integral_probe_sub","label":"integral_optimalValue_sub_eq_two_mul_horizon_mul_integral_probe_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_optimalValue_sub_eq_two_mul_horizon_mul_integral_probe_sub","description":"Integrating the normalized optimal-tail probe against two probability measures and subtracting transports exactly back to the `V*` difference.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimaltailalignment/index.html#decl-43b9db17d864","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","order":7281,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment.lean:23"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_optimalValue_sub_eq_two_mul_horizon_mul_integral_probe_sub (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : (∫ nextState, mdp.optimalValueAt (stage + 1) (Nat.succ_le_of_lt stage.isLt) nextState ∂summary.aggregateEmpiricalTransitionKernel defaultState (state, action)) - mdp.transitionValue (mdp.optimalValueAt (stage + 1) (Nat.succ_le_of_lt stage.isLt)) state action = 2 * (mdp.horizon : Real) * ((∫ nextState, mdp.optimalTailProbe stage nextState ∂summary.aggregateEmpiricalTransitionKernel defaultState (state, action)) - mdp.transitionValue (mdp.optimalTailProbe stage) state action)","missing":[],"search":"integral_optimalvalue_sub_eq_two_mul_horizon_mul_integral_probe_sub banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.integral_optimalvalue_sub_eq_two_mul_horizon_mul_integral_probe_sub integrating the normalized optimal-tail probe against two probability measures and subtracting transports exactly back to the `v*` difference. theorem compiled","shard":"modules/b50caa3959474ba5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransition_optimalValue_sub_lt","label":"abs_empiricalTransition_optimalValue_sub_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransition_optimalValue_sub_lt","description":"Outside the proved joint event, the empirical transition kernel consumed by the planner has the sharp `V*` projection error at every peeled positive actual count. No confidence statement is supplied by the caller.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimaltailalignment/index.html#decl-916caaa60326","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","order":7282,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_empiricalTransition_optimalValue_sub_lt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {logBudget : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes logBudget) (defaultState : State) (index : OptimalTailIndex mdp episodes) (hactual : adaptiveCumulativeAggregateVisitCountAt trajectory index.round index.state index.action = index.count + 1) : |(∫ nextState, mdp.optimalValueAt (index.stage + 1) (Nat.succ_le_of_lt index.stage.isLt) nextState ∂TransitionCountSummary.aggregateEmpiricalTransitionKernel (adaptiveCumulativeEmpiricalModelStateAt trajectory index.round).1 defaultState (index.state, index.action)) - mdp.transitionValue (mdp.optimalValueAt (index.stage + 1) (Nat.succ_le_of_lt…","missing":[],"search":"abs_empiricaltransition_optimalvalue_sub_lt banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.abs_empiricaltransition_optimalvalue_sub_lt outside the proved joint event, the empirical transition kernel consumed by the planner has the sharp `v*` projection error at every peeled positive actual count. no confidence statement is supplied by the caller. theorem compiled","shard":"modules/b50caa3959474ba5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalQAt","label":"optimalQAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalQAt","description":"noncomputable def optimalQAt (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-40071fc82a26","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7283,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:20"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalQAt (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) : Real","missing":[],"search":"optimalqat banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.optimalqat noncomputable def optimalqat (mdp : mdp state action) (stage : fin mdp.horizon) (state : state) (action : action) : real definition compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.QDominatesOptimal","label":"QDominatesOptimal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.QDominatesOptimal","description":"def QDominatesOptimal (mdp : MDP State Action) (table : QTable mdp) : Prop","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-af25ea85e92e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7284,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def QDominatesOptimal (mdp : MDP State Action) (table : QTable mdp) : Prop","missing":[],"search":"qdominatesoptimal banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.qdominatesoptimal def qdominatesoptimal (mdp : mdp state action) (table : qtable mdp) : prop definition compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalQAt_le_horizon","label":"optimalQAt_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalQAt_le_horizon","description":"Bounded rewards imply every optimal action value is at most `H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-4a4a60c7219f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7285,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalQAt_le_horizon (mdp : MDP State Action) (hreward : ∀ state action, |mdp.reward state action| <= 1) (stage : Fin mdp.horizon) (state : State) (action : Action) : optimalQAt mdp stage state action <= (mdp.horizon : Real)","missing":[],"search":"optimalqat_le_horizon banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.optimalqat_le_horizon bounded rewards imply every optimal action value is at most `h`. theorem compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.initialQTable_dominatesOptimal","label":"initialQTable_dominatesOptimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.initialQTable_dominatesOptimal","description":"The all-`H` initialization is optimistic.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-e822e296dabc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7286,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem initialQTable_dominatesOptimal (mdp : MDP State Action) (hreward : ∀ state action, |mdp.reward state action| <= 1) : QDominatesOptimal mdp (initialQTable mdp)","missing":[],"search":"initialqtable_dominatesoptimal banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.initialqtable_dominatesoptimal the all-`h` initialization is optimistic. theorem compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.HasOptimalTailConfidence","label":"HasOptimalTailConfidence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.HasOptimalTailConfidence","description":"Statistical premise for one deterministic clipped update. Downstream it is discharged from the same-source finite event; it is kept abstract here so the Bellman induction remains reusable and auditable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-46b3b020ecb2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7287,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def HasOptimalTailConfidence (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) : Prop","missing":[],"search":"hasoptimaltailconfidence banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.hasoptimaltailconfidence statistical premise for one deterministic clipped update. downstream it is discharged from the same-source finite event; it is kept abstract here so the bellman induction remains reusable and auditable. definition compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_dominatesOptimal","label":"clippedQTable_dominatesOptimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_dominatesOptimal","description":"One clipped update preserves pointwise optimal-Q dominance.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-40ef57fabf32","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7288,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedQTable_dominatesOptimal (mdp : MDP State Action) (previousQ : QTable mdp) (summary : TransitionCountSummary mdp) (defaultState : State) (bonusScale : Real) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) (hprevious : QDominatesOptimal mdp previousQ) (hconfidence : HasOptimalTailConfidence mdp summary defaultState bonusScale) : QDominatesOptimal mdp (clippedQTable mdp previousQ summary defaultState bonusScale)","missing":[],"search":"clippedqtable_dominatesoptimal banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.clippedqtable_dominatesoptimal one clipped update preserves pointwise optimal-q dominance. theorem compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.adaptiveCumulativeAggregateVisitCountAt_le","label":"adaptiveCumulativeAggregateVisitCountAt_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.adaptiveCumulativeAggregateVisitCountAt_le","description":"A pooled state-action count through prefix `round` is at most the number of observed episodes times `H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-1582d0eac22d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7289,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:254"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveCumulativeAggregateVisitCountAt_le {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (round : Nat) (state : State) (action : Action) : adaptiveCumulativeAggregateVisitCountAt trajectory round state action <= (round + 1) * mdp.horizon","missing":[],"search":"adaptivecumulativeaggregatevisitcountat_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptivecumulativeaggregatevisitcountat_le a pooled state-action count through prefix `round` is at most the number of observed episodes times `h`. theorem compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.hasOptimalTailConfidence_of_not_mem_simultaneousTransitionFailureEvent","label":"hasOptimalTailConfidence_of_not_mem_simultaneousTransitionFailureEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.hasOptimalTailConfidence_of_not_mem_simultaneousTransitionFailureEvent","description":"The same-source joint event discharges the complete scalar confidence contract for the exact pooled summary at a queried positive prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvioptimism/index.html#decl-dcda20b8b58e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","order":7290,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIOptimism.lean:276"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasOptimalTailConfidence_of_not_mem_simultaneousTransitionFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (round : Fin episodes) : HasOptimalTailConfidence mdp (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1 defaultState (scale (State := State) (Action := Action) mdp episodes delta)","missing":[],"search":"hasoptimaltailconfidence_of_not_mem_simultaneoustransitionfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.hasoptimaltailconfidence_of_not_mem_simultaneoustransitionfailureevent the same-source joint event discharges the complete scalar confidence contract for the exact pooled summary at a queried positive prefix. theorem compiled","shard":"modules/4749049da345191f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.card_bernsteinCoordinateIndex","label":"card_bernsteinCoordinateIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.card_bernsteinCoordinateIndex","description":"theorem card_bernsteinCoordinateIndex (mdp : MDP State Action) (episodes : Nat) : Fintype.card (BernsteinCoordinateIndex mdp episodes) = episodes * Fintype.card State * Fintype.card Action * Fintype.card State * (episodes * mdp.horizon)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-3056873282ac","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7291,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:22"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem card_bernsteinCoordinateIndex (mdp : MDP State Action) (episodes : Nat) : Fintype.card (BernsteinCoordinateIndex mdp episodes) = episodes * Fintype.card State * Fintype.card Action * Fintype.card State * (episodes * mdp.horizon)","missing":[],"search":"card_bernsteincoordinateindex banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.card_bernsteincoordinateindex theorem card_bernsteincoordinateindex (mdp : mdp state action) (episodes : nat) : fintype.card (bernsteincoordinateindex mdp episodes) = episodes * fintype.card state * fintype.card action * fintype.card state * (episodes * mdp.horizon) theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.card_optimalTailIndex","label":"card_optimalTailIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.card_optimalTailIndex","description":"theorem card_optimalTailIndex (mdp : MDP State Action) (episodes : Nat) : Fintype.card (OptimalTailIndex mdp episodes) = episodes * mdp.horizon * Fintype.card State * Fintype.card Action * (episodes * mdp.horizon)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-0e7250a1f4b3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7292,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem card_optimalTailIndex (mdp : MDP State Action) (episodes : Nat) : Fintype.card (OptimalTailIndex mdp episodes) = episodes * mdp.horizon * Fintype.card State * Fintype.card Action * (episodes * mdp.horizon)","missing":[],"search":"card_optimaltailindex banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.card_optimaltailindex theorem card_optimaltailindex (mdp : mdp state action) (episodes : nat) : fintype.card (optimaltailindex mdp episodes) = episodes * mdp.horizon * fintype.card state * fintype.card action * (episodes * mdp.horizon) theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.exp_neg_two_mul_logFactor_eq","label":"exp_neg_two_mul_logFactor_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.exp_neg_two_mul_logFactor_eq","description":"Exact exponential simplification at the paper logarithmic factor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-53fca0d8aa10","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7293,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_two_mul_logFactor_eq (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : Real.exp (-2 * logFactor (State := State) (Action := Action) mdp episodes delta) = (delta / (confidenceNumerator (State := State) (Action := Action) mdp episodes : Nat)) ^ 2","missing":[],"search":"exp_neg_two_mul_logfactor_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.exp_neg_two_mul_logfactor_eq exact exponential simplification at the paper logarithmic factor. theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.exp_neg_two_mul_logFactor_sq_le","label":"exp_neg_two_mul_logFactor_sq_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.exp_neg_two_mul_logFactor_sq_le","description":"The optimal-tail scalar term is no larger than the coordinate term because the task log factor is at least one.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-7d31bfb062cd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7294,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:90"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exp_neg_two_mul_logFactor_sq_le (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : Real.exp (-2 * logFactor (State := State) (Action := Action) mdp episodes delta ^ 2) <= Real.exp (-2 * logFactor (State := State) (Action := Action) mdp episodes delta)","missing":[],"search":"exp_neg_two_mul_logfactor_sq_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.exp_neg_two_mul_logfactor_sq_le the optimal-tail scalar term is no larger than the coordinate term because the task log factor is at least one. theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.ofReal_card_mul_two_mul_exp_le","label":"ofReal_card_mul_two_mul_exp_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.ofReal_card_mul_two_mul_exp_le","description":"private theorem ofReal_card_mul_two_mul_exp_le (card : Nat) (x budget : Real) (hx : 0 <= x) (hbudget : 0 <= budget) (h : (card : Real) * (2 * x) <= budget) : (card : ENNReal) * (2 * ENNReal.ofReal x) <= ENNReal.ofReal budget","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-c91a38433c30","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7295,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem ofReal_card_mul_two_mul_exp_le (card : Nat) (x budget : Real) (hx : 0 <= x) (hbudget : 0 <= budget) (h : (card : Real) * (2 * x) <= budget) : (card : ENNReal) * (2 * ENNReal.ofReal x) <= ENNReal.ofReal budget","missing":[],"search":"ofreal_card_mul_two_mul_exp_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.ofreal_card_mul_two_mul_exp_le private theorem ofreal_card_mul_two_mul_exp_le (card : nat) (x budget : real) (hx : 0 <= x) (hbudget : 0 <= budget) (h : (card : real) * (2 * x) <= budget) : (card : ennreal) * (2 * ennreal.ofreal x) <= ennreal.ofreal budget theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.simultaneousTransitionTailBudget_le_fifth","label":"simultaneousTransitionTailBudget_le_fifth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.simultaneousTransitionTailBudget_le_fifth","description":"The complete coordinate-plus-optimal-tail confidence family consumes at most one fifth of `delta`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-36351d1010a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7296,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:118"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousTransitionTailBudget_le_fifth (mdp : MDP State Action) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (Fintype.card (BernsteinCoordinateIndex mdp episodes) : ENNReal) * (2 * ENNReal.ofReal (Real.exp (-2 * logFactor (State := State) (Action := Action) mdp episodes delta))) + (Fintype.card (OptimalTailIndex mdp episodes) : ENNReal) * (2 * ENNReal.ofReal (Real.exp (-2 * logFactor (State := State) (Action := Action) mdp episodes delta ^ 2))) <= ENNReal.ofReal (delta / 5)","missing":[],"search":"simultaneoustransitiontailbudget_le_fifth banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.simultaneoustransitiontailbudget_le_fifth the complete coordinate-plus-optimal-tail confidence family consumes at most one fifth of `delta`. theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","label":"recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","description":"Specialized same-source confidence event probability.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviprobabilitybudget/index.html#decl-02d4b17fca6b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","order":7297,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget.lean:234"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hreward : ∀ state action, |mdp.reward state action| <= 1) : let source := recurrentSource mdp initialState defaultState episodes delta source.trajectoryMeasure (AdaptiveEpisodeBatchSource.simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) <= ENNReal.ofReal (delta / 5)","missing":[],"search":"recurrentsource_trajectorymeasure_simultaneoustransitionfailureevent_le_fifth banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_simultaneoustransitionfailureevent_le_fifth specialized same-source confidence event probability. theorem compiled","shard":"modules/aaf9cea6994cf22b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeSummaryOfSequence_prefixTransitionSummaries_eq","label":"cumulativeSummaryOfSequence_prefixTransitionSummaries_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeSummaryOfSequence_prefixTransitionSummaries_eq","description":"The planner fold and the statistical state use literally the same prefix transition summary.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvirecurrentoptimism/index.html#decl-90ab3811fce7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","order":7298,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism.lean:23"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSummaryOfSequence_prefixTransitionSummaries_eq {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (round : Nat) : cumulativeSummaryOfSequence (prefixTransitionSummaries (Preorder.frestrictLe round trajectory)) = (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1","missing":[],"search":"cumulativesummaryofsequence_prefixtransitionsummaries_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.cumulativesummaryofsequence_prefixtransitionsummaries_eq the planner fold and the statistical state use literally the same prefix transition summary. theorem compiled","shard":"modules/d7010abdf823bb53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","label":"recurrentQTableOfTrajectory_dominatesOptimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","description":"Every finite Q-table fold through the first `n` actually generated episodes is optimistic on the proved same-source event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvirecurrentoptimism/index.html#decl-a868ac036c84","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","order":7299,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentQTableOfTrajectory_dominatesOptimal {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) : ∀ n, n <= episodes -> QDominatesOptimal mdp (recurrentQTableOfSummaries mdp defaultState (scale (State := State) (Action := Action) mdp episodes delta) n (fun i => (trajectory i).transitionCountSummary))","missing":[],"search":"recurrentqtableoftrajectory_dominatesoptimal banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.recurrentqtableoftrajectory_dominatesoptimal every finite q-table fold through the first `n` actually generated episodes is optimistic on the proved same-source event. theorem compiled","shard":"modules/d7010abdf823bb53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessor_selectedQ_ge_optimalQ","label":"recurrentSuccessor_selectedQ_ge_optimalQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessor_selectedQ_ge_optimalQ","description":"The generated successor episode selects an action whose actual recurrent Q value dominates every optimal action value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvirecurrentoptimism/index.html#decl-c63605a42979","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","order":7300,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSuccessor_selectedQ_ge_optimalQ {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (round : Fin episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : optimalQAt mdp stage state action <= recurrentQTableOfSummaries mdp defaultState (scale (State := State) (Action := Action) mdp episodes delt…","missing":[],"search":"recurrentsuccessor_selectedq_ge_optimalq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.recurrentsuccessor_selectedq_ge_optimalq the generated successor episode selects an action whose actual recurrent q value dominates every optimal action value. theorem compiled","shard":"modules/d7010abdf823bb53.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPreviousQ","label":"successorPreviousQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPreviousQ","description":"Previous Q table for successor episode `n+1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-2d9ac9c072d1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7301,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:22"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorPreviousQ (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : QTable mdp","missing":[],"search":"successorpreviousq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorpreviousq previous q table for successor episode `n+1`. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorSummary","label":"successorSummary","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorSummary","description":"Exact pooled summary available before successor episode `n+1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-fbefddb828b8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7302,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def successorSummary {mdp : MDP State Action} (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : TransitionCountSummary mdp","missing":[],"search":"successorsummary banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorsummary exact pooled summary available before successor episode `n+1`. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyTable","label":"successorPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyTable","description":"Greedy recurrent table actually used by successor episode `n+1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-096918747c36","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7303,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorPolicyTable (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"successorpolicytable banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorpolicytable greedy recurrent table actually used by successor episode `n+1`. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyGapRemaining","label":"successorPolicyGapRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyGapRemaining","description":"Same-policy continuation gap in the successor episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-273960adaec8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7304,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorPolicyGapRemaining (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : Real","missing":[],"search":"successorpolicygapremaining banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorpolicygapremaining same-policy continuation gap in the successor episode. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge","label":"successorLocalCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge","description":"Per-stage deterministic charge. A zero previous count is paid by `H`; positive counts pay the `9HL/sqrt N` bonus/projection term plus the explicit `66 S H^2 L/N` Bernstein correction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-4a76983b1e26","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7305,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:63"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorLocalCharge (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : Real","missing":[],"search":"successorlocalcharge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorlocalcharge per-stage deterministic charge. a zero previous count is paid by `h`; positive counts pay the `9hl/sqrt n` bonus/projection term plus the explicit `66 s h^2 l/n` bernstein correction. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyGapRemaining_mem_Icc","label":"successorPolicyGapRemaining_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyGapRemaining_mem_Icc","description":"theorem successorPolicyGapRemaining_mem_Icc {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= m…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-54325535f0c8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7306,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorPolicyGapRemaining_mem_Icc {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ AdaptiveEpisodeBatchSource.simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (n : Nat) (hn : n + 1 <= episodes) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : successorPolicyGapRemaining mdp defaultState episodes delta trajectory n remaining…","missing":[],"search":"successorpolicygapremaining_mem_icc banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorpolicygapremaining_mem_icc theorem successorpolicygapremaining_mem_icc {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) {episodes : nat} {delta : real} {trajectory : episodebatchtrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardnonneg : forall state action, 0 <= mdp.reward state action) (hrewardone : forall state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ adaptiveepisodebatchsource.simultaneoustransitionfailureevent source episodes (logfactor (state := state) (action := action) mdp episodes delta)) (defaultstate : state) (n : nat) (hn : n + 1 <= episodes) (remaining : nat) (hremaining : remaining <= mdp.horizon) (state : state) : successorpolicygapremaining mdp defaultstate episodes delta trajectory n remaining hremaining state ∈ set.icc (0 : real) mdp.horizon theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_nonneg","label":"successorLocalCharge_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_nonneg","description":"theorem successorLocalCharge_nonneg (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : 0 <= successorLocalCharge mdp defaultState episodes delta trajectory n remaining hremaining state","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-13019eeb68ff","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7307,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorLocalCharge_nonneg (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : 0 <= successorLocalCharge mdp defaultState episodes delta trajectory n remaining hremaining state","missing":[],"search":"successorlocalcharge_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorlocalcharge_nonneg theorem successorlocalcharge_nonneg (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (trajectory : episodebatchtrajectory mdp 1) (n remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) (state : state) : 0 <= successorlocalcharge mdp defaultstate episodes delta trajectory n remaining hremaining state theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight","label":"normalizedBellmanRemainingWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight","description":"Normalized weight at the start of a subproblem with `remaining` decisions.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-e1e7041c3a37","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7308,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedBellmanRemainingWeight (mdp : MDP State Action) (remaining : Nat) : Real","missing":[],"search":"normalizedbellmanremainingweight banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanremainingweight normalized weight at the start of a subproblem with `remaining` decisions. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_zero","label":"normalizedBellmanRemainingWeight_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_zero","description":"theorem normalizedBellmanRemainingWeight_zero (mdp : MDP State Action) : normalizedBellmanRemainingWeight mdp mdp.horizon = 31 / 32","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-9dced3872ac9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7309,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedBellmanRemainingWeight_zero (mdp : MDP State Action) : normalizedBellmanRemainingWeight mdp mdp.horizon = 31 / 32","missing":[],"search":"normalizedbellmanremainingweight_zero banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanremainingweight_zero theorem normalizedbellmanremainingweight_zero (mdp : mdp state action) : normalizedbellmanremainingweight mdp mdp.horizon = 31 / 32 theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_succ_eq_chargeWeight","label":"normalizedBellmanRemainingWeight_succ_eq_chargeWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_succ_eq_chargeWeight","description":"theorem normalizedBellmanRemainingWeight_succ_eq_chargeWeight (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : normalizedBellmanRemainingWeight mdp (remaining + 1) = normalizedBellmanChargeWeight mdp (mdp.decisionStageRemaining remaining hremaining)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-a07424c4677f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7310,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:167"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedBellmanRemainingWeight_succ_eq_chargeWeight (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : normalizedBellmanRemainingWeight mdp (remaining + 1) = normalizedBellmanChargeWeight mdp (mdp.decisionStageRemaining remaining hremaining)","missing":[],"search":"normalizedbellmanremainingweight_succ_eq_chargeweight banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanremainingweight_succ_eq_chargeweight theorem normalizedbellmanremainingweight_succ_eq_chargeweight (mdp : mdp state action) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) : normalizedbellmanremainingweight mdp (remaining + 1) = normalizedbellmanchargeweight mdp (mdp.decisionstageremaining remaining hremaining) theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_eq_stageWeight","label":"normalizedBellmanRemainingWeight_eq_stageWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_eq_stageWeight","description":"theorem normalizedBellmanRemainingWeight_eq_stageWeight (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : normalizedBellmanRemainingWeight mdp remaining = normalizedBellmanWeight mdp (mdp.decisionStageRemaining remaining hremaining)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-c77667c4eb33","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7311,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:176"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedBellmanRemainingWeight_eq_stageWeight (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : normalizedBellmanRemainingWeight mdp remaining = normalizedBellmanWeight mdp (mdp.decisionStageRemaining remaining hremaining)","missing":[],"search":"normalizedbellmanremainingweight_eq_stageweight banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.normalizedbellmanremainingweight_eq_stageweight theorem normalizedbellmanremainingweight_eq_stageweight (mdp : mdp state action) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) : normalizedbellmanremainingweight mdp remaining = normalizedbellmanweight mdp (mdp.decisionstageremaining remaining hremaining) theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeFrom","label":"successorWeightedChargeFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeFrom","description":"Weighted local charges along the state sequence reconstructed recursively from one successor batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-7a80611c0131","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7312,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:191"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorWeightedChargeFrom (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> StepTrace Action State remaining -> Real | 0, _, _, _ => 0 | remaining + 1, hremaining, state, trace => let stage := mdp.decisionStageRemaining remaining hremaining normalizedBellmanChargeWeight mdp stage * successorLocalCharge mdp defaultState episodes delta trajectory n remaining hremaining state + successorWeightedChargeFrom mdp defaultState episodes delta trajectory n remaining (by omega) (trace 0).2 (Fin.tail trace) /-- Full successor-episode charge on the canonical state reconstruction. -/ noncomputable def successorCanonicalWeightedCharge (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchT…","missing":[],"search":"successorweightedchargefrom banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorweightedchargefrom weighted local charges along the state sequence reconstructed recursively from one successor batch. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge","label":"successorCanonicalWeightedCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge","description":"Full successor-episode charge on the canonical state reconstruction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-172c91da5919","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7313,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:207"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorCanonicalWeightedCharge (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : Real","missing":[],"search":"successorcanonicalweightedcharge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorcanonicalweightedcharge full successor-episode charge on the canonical state reconstruction. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeSum","label":"successorWeightedChargeSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeSum","description":"Chronological weighted charge of one successor episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-1c36b0e9d3a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7314,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorWeightedChargeSum (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : Real","missing":[],"search":"successorweightedchargesum banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorweightedchargesum chronological weighted charge of one successor episode. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeSum_nonneg","label":"successorWeightedChargeSum_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeSum_nonneg","description":"theorem successorWeightedChargeSum_nonneg (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : 0 <= successorWeightedChargeSum mdp defaultState episodes delta trajectory n","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-100ef610d68d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7315,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorWeightedChargeSum_nonneg (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : 0 <= successorWeightedChargeSum mdp defaultState episodes delta trajectory n","missing":[],"search":"successorweightedchargesum_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorweightedchargesum_nonneg theorem successorweightedchargesum_nonneg (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (trajectory : episodebatchtrajectory mdp 1) (n : nat) : 0 <= successorweightedchargesum mdp defaultstate episodes delta trajectory n theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedVisitedGap","label":"successorWeightedVisitedGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedVisitedGap","description":"Weighted state-gap transported along one generated successor batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-8ca6d10a7463","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7316,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:240"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorWeightedVisitedGap (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) (stage : Fin mdp.horizon) : Real","missing":[],"search":"successorweightedvisitedgap banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.successorweightedvisitedgap weighted state-gap transported along one generated successor batch. definition compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.successorPolicyGapRemaining_le_inflation_mul_transition_add_charge","label":"successorPolicyGapRemaining_le_inflation_mul_transition_add_charge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.successorPolicyGapRemaining_le_inflation_mul_transition_add_charge","description":"The good same-source transition event discharges the complete local inflated Bellman recursion at every successor episode and state.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-a56d259d5a01","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7317,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:254"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorPolicyGapRemaining_le_inflation_mul_transition_add_charge {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (n : Nat) (hn : n + 1 < episodes) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : successorPolicyGapRemaining mdp defaultState episodes delta trajectory n (re…","missing":[],"search":"successorpolicygapremaining_le_inflation_mul_transition_add_charge banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.successorpolicygapremaining_le_inflation_mul_transition_add_charge the good same-source transition event discharges the complete local inflated bellman recursion at every successor episode and state. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.clippedSuccessorGapFeatureOfSummaries_eq_weight_mul_gap","label":"clippedSuccessorGapFeatureOfSummaries_eq_weight_mul_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.clippedSuccessorGapFeatureOfSummaries_eq_weight_mul_gap","description":"On the joint confidence event, the globally clipped martingale feature is exactly the normalized recurrent policy gap.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-3785c8ea76d4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7318,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:410"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem clippedSuccessorGapFeatureOfSummaries_eq_weight_mul_gap {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (n : Nat) (hn : n + 1 <= episodes) (stage : Fin mdp.horizon) (state : State) : clippedSuccessorGapFeatureOfSummaries mdp defaultState episodes delta n (fun i => (trajectory i).transitionCountSummary) s…","missing":[],"search":"clippedsuccessorgapfeatureofsummaries_eq_weight_mul_gap banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.clippedsuccessorgapfeatureofsummaries_eq_weight_mul_gap on the joint confidence event, the globally clipped martingale feature is exactly the normalized recurrent policy gap. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.normalizedGap_le_weightedChargeFrom_add_deterministicInnovation","label":"normalizedGap_le_weightedChargeFrom_add_deterministicInnovation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.normalizedGap_le_weightedChargeFrom_add_deterministicInnovation","description":"Exact weighted Bellman telescope for one successor episode. The charge, policy table, recursively reconstructed state path, and martingale innovation are all the same objects used by the recurrent generated source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-70970b8eaa47","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7319,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:460"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedGap_le_weightedChargeFrom_add_deterministicInnovation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (n : Nat) (hn : n + 1 < episodes) : forall (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (trace : StepTrace Action State remaining), normalizedBellmanRemainingWeight…","missing":[],"search":"normalizedgap_le_weightedchargefrom_add_deterministicinnovation banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.normalizedgap_le_weightedchargefrom_add_deterministicinnovation exact weighted bellman telescope for one successor episode. the charge, policy table, recursively reconstructed state path, and martingale innovation are all the same objects used by the recurrent generated source. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.successorPolicyGapRemaining_le_cap_mul_charge_add_innovation","label":"successorPolicyGapRemaining_le_cap_mul_charge_add_innovation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.successorPolicyGapRemaining_le_cap_mul_charge_add_innovation","description":"The full-horizon successor policy gap is bounded by the globally capped canonical charge plus the exact successor Bellman innovation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-7e0bcf1a2d83","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7320,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:571"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorPolicyGapRemaining_le_cap_mul_charge_add_innovation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (defaultState : State) (n : Nat) (hn : n + 1 < episodes) : successorPolicyGapRemaining mdp defaultState episodes delta trajectory n mdp.horizon le_rfl ((trajectory (n + 1)).reconstructedInitialState defaultState) <= bel…","missing":[],"search":"successorpolicygapremaining_le_cap_mul_charge_add_innovation banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.successorpolicygapremaining_le_cap_mul_charge_add_innovation the full-horizon successor policy gap is bounded by the globally capped canonical charge plus the exact successor bellman innovation. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedEpisodePseudoRegret_le_successorPolicyGapRemaining","label":"recurrentSource_generatedEpisodePseudoRegret_le_successorPolicyGapRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedEpisodePseudoRegret_le_successorPolicyGapRemaining","description":"Optimism converts the generated successor episode's policy-value pseudo-regret into the same recurrent policy gap used by the telescope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-687264ece725","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7321,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:613"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_generatedEpisodePseudoRegret_le_successorPolicyGapRemaining {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (defaultState : State) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent (recurrentSource mdp initialState defaultState episodes delta) episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (n : Nat) (hn : n + 1 < episodes) : generatedEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState trajectory (n + 1) <= successorPolicyGapR…","missing":[],"search":"recurrentsource_generatedepisodepseudoregret_le_successorpolicygapremaining banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.recurrentsource_generatedepisodepseudoregret_le_successorpolicygapremaining optimism converts the generated successor episode's policy-value pseudo-regret into the same recurrent policy gap used by the telescope. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentBellmanInnovationProcess_succ_eq_canonical","label":"recurrentBellmanInnovationProcess_succ_eq_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentBellmanInnovationProcess_succ_eq_canonical","description":"The named adaptive martingale process is definitionally the direct canonical innovation used in the successor telescope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-8236e9e2c6cd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7322,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:684"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentBellmanInnovationProcess_succ_eq_canonical (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) (n : Nat) : recurrentBellmanInnovationProcess mdp defaultState episodes delta (n + 1) trajectory = mdp.sampledCumulativeDeterministicGapInnovationFrom (successorPolicyTable mdp defaultState episodes delta trajectory n) (clippedSuccessorGapFeatureOfSummaries mdp defaultState episodes delta n (fun i => (trajectory i).transitionCountSummary)) mdp.horizon le_rfl ((trajectory (n + 1)).reconstructedInitialState defaultState) ((trajectory (n + 1)).reconstructedStepTrace)","missing":[],"search":"recurrentbellmaninnovationprocess_succ_eq_canonical banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.recurrentbellmaninnovationprocess_succ_eq_canonical the named adaptive martingale process is definitionally the direct canonical innovation used in the successor telescope. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","label":"recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","description":"End-to-end good-event decomposition for one generated successor episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviregretdecomposition/index.html#decl-b631547f72c5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","order":7323,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition.lean:700"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : forall state action, 0 <= mdp.reward state action) (hrewardOne : forall state action, mdp.reward state action <= 1) (defaultState : State) (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent (recurrentSource mdp initialState defaultState episodes delta) episodes (logFactor (State := State) (Action := Action) mdp episodes delta)) (n : Nat) (hn : n + 1 < episodes) : generatedEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState trajectory (n + 1) <= bellmanWeightCap * (suc…","missing":[],"search":"recurrentsource_generatedsuccessorpseudoregret_le_charge_add_innovation banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.recurrentsource_generatedsuccessorpseudoregret_le_charge_add_innovation end-to-end good-event decomposition for one generated successor episode. theorem compiled","shard":"modules/2bc011b11da9fa3a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead","label":"transitionResidualHead","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead","description":"One transition-coordinate residual on a chronological trace head.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-bc9f4cd10e22","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7324,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionResidualHead (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (currentState : State) (head : Action × State) : Real","missing":[],"search":"transitionresidualhead banditrlproof.finitehorizonrl.mdp.transitionresidualhead one transition-coordinate residual on a chronological trace head. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitHead","label":"transitionVisitHead","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionVisitHead","description":"The realized visit charged by one transition-coordinate residual.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-e671c65f3a08","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7325,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionVisitHead (_mdp : MDP State Action) (targetState : State) (targetAction : Action) (currentState : State) (head : Action × State) : Real","missing":[],"search":"transitionvisithead banditrlproof.finitehorizonrl.mdp.transitionvisithead the realized visit charged by one transition-coordinate residual. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom","label":"transitionResidualFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom","description":"Sum of one fixed transition residual over the remaining generated trace.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-284b2a86bccd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7326,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionResidualFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) : (remaining : Nat) -> State -> StepTrace Action State remaining -> Real | 0, _currentState, _trace => 0 | remaining + 1, currentState, trace => mdp.transitionResidualHead targetState targetAction targetNextState currentState (trace 0) + mdp.transitionResidualFrom targetState targetAction targetNextState remaining (trace 0).2 (Fin.tail trace) /-- Sum of actual visits to one state-action coordinate over the remaining trace. -/ def transitionVisitFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) : (remaining : Nat) -> State -> StepTrace Action State remaining -> Real | 0, _currentState, _trace => 0 | remaining + 1, currentState, trace => mdp.transitionVisitHead targetState targetAction currentState (trace 0) + mdp.transitionVisitFrom targetS…","missing":[],"search":"transitionresidualfrom banditrlproof.finitehorizonrl.mdp.transitionresidualfrom sum of one fixed transition residual over the remaining generated trace. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom","label":"transitionVisitFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom","description":"Sum of actual visits to one state-action coordinate over the remaining trace.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-7b3197f77c2d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7327,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:63"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionVisitFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) : (remaining : Nat) -> State -> StepTrace Action State remaining -> Real | 0, _currentState, _trace => 0 | remaining + 1, currentState, trace => mdp.transitionVisitHead targetState targetAction currentState (trace 0) + mdp.transitionVisitFrom targetState targetAction remaining (trace 0).2 (Fin.tail trace) omit [Nonempty State] [Nonempty Action] in theorem measurable_transitionResidualFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionResidualFrom targetState targetAction targetNextState remaining p.1 p.2)","missing":[],"search":"transitionvisitfrom banditrlproof.finitehorizonrl.mdp.transitionvisitfrom sum of actual visits to one state-action coordinate over the remaining trace. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionResidualFrom","label":"measurable_transitionResidualFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionResidualFrom","description":"theorem measurable_transitionResidualFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionResidualFrom targetState targetAction targetNextState remaining p.1 p.2)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-b04a9f694072","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7328,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionResidualFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionResidualFrom targetState targetAction targetNextState remaining p.1 p.2)","missing":[],"search":"measurable_transitionresidualfrom banditrlproof.finitehorizonrl.mdp.measurable_transitionresidualfrom theorem measurable_transitionresidualfrom (mdp : mdp state action) (targetstate : state) (targetaction : action) (targetnextstate : state) (remaining : nat) : measurable (fun p : state × steptrace action state remaining => mdp.transitionresidualfrom targetstate targetaction targetnextstate remaining p.1 p.2) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionVisitFrom","label":"measurable_transitionVisitFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionVisitFrom","description":"theorem measurable_transitionVisitFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionVisitFrom targetState targetAction remaining p.1 p.2)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-c1cd6d3b2bd4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7329,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionVisitFrom (mdp : MDP State Action) (targetState : State) (targetAction : Action) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionVisitFrom targetState targetAction remaining p.1 p.2)","missing":[],"search":"measurable_transitionvisitfrom banditrlproof.finitehorizonrl.mdp.measurable_transitionvisitfrom theorem measurable_transitionvisitfrom (mdp : mdp state action) (targetstate : state) (targetaction : action) (remaining : nat) : measurable (fun p : state × steptrace action state remaining => mdp.transitionvisitfrom targetstate targetaction remaining p.1 p.2) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead_compensated_hasMGFUpperBoundAt","label":"transitionResidualHead_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead_compensated_hasMGFUpperBoundAt","description":"theorem transitionResidualHead_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (currentState : State) (chosenAction : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun nextState => tilt * mdp.transitionResidualHead targetState targetAction targetNextState currentState (chosenAction, nextState) - (tilt ^ 2 / 8) * mdp.transitio…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-e652aaff8e22","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7330,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:90"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionResidualHead_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (currentState : State) (chosenAction : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun nextState => tilt * mdp.transitionResidualHead targetState targetAction targetNextState currentState (chosenAction, nextState) - (tilt ^ 2 / 8) * mdp.transitionVisitHead targetState targetAction currentState (chosenAction, nextState)) 1 0 (mdp.transition (currentState, chosenAction))","missing":[],"search":"transitionresidualhead_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.transitionresidualhead_compensated_hasmgfupperboundat theorem transitionresidualhead_compensated_hasmgfupperboundat (mdp : mdp state action) (targetstate : state) (targetaction : action) (targetnextstate : state) (currentstate : state) (chosenaction : action) (tilt : real) : concentration.hasmgfupperboundat (fun nextstate => tilt * mdp.transitionresidualhead targetstate targetaction targetnextstate currentstate (chosenaction, nextstate) - (tilt ^ 2 / 8) * mdp.transitionvisithead targetstate targetaction currentstate (chosenaction, nextstate)) 1 0 (mdp.transition (currentstate, chosenaction)) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionResidualHead_compensated_hasMGFUpperBoundAt","label":"actionStateKernel_transitionResidualHead_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionResidualHead_compensated_hasMGFUpperBoundAt","description":"One generated action/transition head pays only its realized visit budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-0c06c285f733","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7331,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_transitionResidualHead_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (targetState : State) (targetAction : Action) (targetNextState : State) (currentState : State) (stage : Fin mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * mdp.transitionResidualHead targetState targetAction targetNextState currentState head - (tilt ^ 2 / 8) * mdp.transitionVisitHead targetState targetAction currentState head) 1 0 (policy.actionStateKernel stage currentState)","missing":[],"search":"actionstatekernel_transitionresidualhead_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.actionstatekernel_transitionresidualhead_compensated_hasmgfupperboundat one generated action/transition head pays only its realized visit budget. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionResidual_compensated_hasMGFUpperBoundAt","label":"trajectoryKernelRemaining_transitionResidual_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionResidual_compensated_hasMGFUpperBoundAt","description":"The whole generated finite trace satisfies the same fixed-tilt exponential bound, with compensator equal to its literal state-action visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-aede22c1cde9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7332,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:242"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_transitionResidual_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (targetState : State) (targetAction : Action) (targetNextState : State) (tilt : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (currentState : State) : Concentration.HasMGFUpperBoundAt (fun trace => tilt * mdp.transitionResidualFrom targetState targetAction targetNextState remaining currentState trace - (tilt ^ 2 / 8) * mdp.transitionVisitFrom targetState targetAction remaining currentState trace) 1 0 (policy.trajectoryKernelRemaining remaining hremaining currentState)","missing":[],"search":"trajectorykernelremaining_transitionresidual_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorykernelremaining_transitionresidual_compensated_hasmgfupperboundat the whole generated finite trace satisfies the same fixed-tilt exponential bound, with compensator equal to its literal state-action visit count. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom_eq_sum","label":"transitionResidualFrom_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom_eq_sum","description":"Residual recursion is exactly the chronological finite sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-319251b1d3f9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7333,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:379"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionResidualFrom_eq_sum (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (remaining : Nat) (currentState : State) (trace : StepTrace Action State remaining) : mdp.transitionResidualFrom targetState targetAction targetNextState remaining currentState trace = ∑ stage : Fin remaining, mdp.transitionResidualHead targetState targetAction targetNextState (StepTrace.stateAt currentState trace stage) (trace stage)","missing":[],"search":"transitionresidualfrom_eq_sum banditrlproof.finitehorizonrl.mdp.transitionresidualfrom_eq_sum residual recursion is exactly the chronological finite sum. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom_eq_sum","label":"transitionVisitFrom_eq_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom_eq_sum","description":"Visit recursion is exactly the chronological finite sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-009f136ba2e2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7334,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:406"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionVisitFrom_eq_sum (mdp : MDP State Action) (targetState : State) (targetAction : Action) (remaining : Nat) (currentState : State) (trace : StepTrace Action State remaining) : mdp.transitionVisitFrom targetState targetAction remaining currentState trace = ∑ stage : Fin remaining, mdp.transitionVisitHead targetState targetAction (StepTrace.stateAt currentState trace stage) (trace stage)","missing":[],"search":"transitionvisitfrom_eq_sum banditrlproof.finitehorizonrl.mdp.transitionvisitfrom_eq_sum visit recursion is exactly the chronological finite sum. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom_eq_aggregateCounts","label":"transitionResidualFrom_eq_aggregateCounts","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom_eq_aggregateCounts","description":"The full-trace residual is the aggregate transition count minus the true singleton mass times the aggregate visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-fa2a3872782e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7335,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:435"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionResidualFrom_eq_aggregateCounts (mdp : MDP State Action) (targetState : State) (targetAction : Action) (targetNextState : State) (trajectory : State × StepTrace Action State mdp.horizon) : mdp.transitionResidualFrom targetState targetAction targetNextState mdp.horizon trajectory.1 trajectory.2 = ((mdp.episodeBatchOfTrajectories 1 (fun _ => trajectory)) |>.transitionCountSummary.aggregateTransitionCount targetState targetAction targetNextState : Real) - ((mdp.episodeBatchOfTrajectories 1 (fun _ => trajectory)) |>.transitionCountSummary.aggregateVisitCount targetState targetAction : Real) * (mdp.transition (targetState, targetAction)).real {targetNextState}","missing":[],"search":"transitionresidualfrom_eq_aggregatecounts banditrlproof.finitehorizonrl.mdp.transitionresidualfrom_eq_aggregatecounts the full-trace residual is the aggregate transition count minus the true singleton mass times the aggregate visit count. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom_eq_aggregateVisitCount","label":"transitionVisitFrom_eq_aggregateVisitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom_eq_aggregateVisitCount","description":"The full-trace compensator is the aggregate visit count of its one-batch image.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-c26cd509d5af","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7336,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:496"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionVisitFrom_eq_aggregateVisitCount (mdp : MDP State Action) (targetState : State) (targetAction : Action) (trajectory : State × StepTrace Action State mdp.horizon) : mdp.transitionVisitFrom targetState targetAction mdp.horizon trajectory.1 trajectory.2 = ((mdp.episodeBatchOfTrajectories 1 (fun _ => trajectory)) |>.transitionCountSummary.aggregateVisitCount targetState targetAction : Real)","missing":[],"search":"transitionvisitfrom_eq_aggregatevisitcount banditrlproof.finitehorizonrl.mdp.transitionvisitfrom_eq_aggregatevisitcount the full-trace compensator is the aggregate visit count of its one-batch image. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionResidual_compensated_hasMGFUpperBoundAt","label":"trajectoryMeasure_transitionResidual_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionResidual_compensated_hasMGFUpperBoundAt","description":"The compensated fixed-tilt witness after integrating the random initial state.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-b61d1a96d516","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7337,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:520"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_transitionResidual_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (targetState : State) (targetAction : Action) (targetNextState : State) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory => tilt * mdp.transitionResidualFrom targetState targetAction targetNextState mdp.horizon trajectory.1 trajectory.2 - (tilt ^ 2 / 8) * mdp.transitionVisitFrom targetState targetAction mdp.horizon trajectory.1 trajectory.2) 1 0 (policy.trajectoryMeasure initialState)","missing":[],"search":"trajectorymeasure_transitionresidual_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorymeasure_transitionresidual_compensated_hasmgfupperboundat the compensated fixed-tilt witness after integrating the random initial state. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateTransitionResidual","label":"aggregateTransitionResidual","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateTransitionResidual","description":"The generated one-batch aggregate transition residual.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-50b81ad7b8c7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7338,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:582"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateTransitionResidual {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (state : State) (action : Action) (nextState : State) : Real","missing":[],"search":"aggregatetransitionresidual banditrlproof.finitehorizonrl.episodebatch.aggregatetransitionresidual the generated one-batch aggregate transition residual. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateVisitReal","label":"aggregateVisitReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateVisitReal","description":"The generated one-batch aggregate state-action visit count as a real.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-dc31033749cd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7339,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:590"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def aggregateVisitReal {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (state : State) (action : Action) : Real","missing":[],"search":"aggregatevisitreal banditrlproof.finitehorizonrl.episodebatch.aggregatevisitreal the generated one-batch aggregate state-action visit count as a real. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateTransitionResidual","label":"measurable_aggregateTransitionResidual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateTransitionResidual","description":"theorem measurable_aggregateTransitionResidual {mdp : MDP State Action} (state : State) (action : Action) (nextState : State) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.aggregateTransitionResidual state action nextState)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-252e48a281b9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7340,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:596"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionResidual {mdp : MDP State Action} (state : State) (action : Action) (nextState : State) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.aggregateTransitionResidual state action nextState)","missing":[],"search":"measurable_aggregatetransitionresidual banditrlproof.finitehorizonrl.episodebatch.measurable_aggregatetransitionresidual theorem measurable_aggregatetransitionresidual {mdp : mdp state action} (state : state) (action : action) (nextstate : state) : measurable (fun batch : episodebatch mdp 1 => batch.aggregatetransitionresidual state action nextstate) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateVisitReal","label":"measurable_aggregateVisitReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateVisitReal","description":"theorem measurable_aggregateVisitReal {mdp : MDP State Action} (state : State) (action : Action) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.aggregateVisitReal state action)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-bde34728db7b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7341,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:630"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateVisitReal {mdp : MDP State Action} (state : State) (action : Action) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.aggregateVisitReal state action)","missing":[],"search":"measurable_aggregatevisitreal banditrlproof.finitehorizonrl.episodebatch.measurable_aggregatevisitreal theorem measurable_aggregatevisitreal {mdp : mdp state action} (state : state) (action : action) : measurable (fun batch : episodebatch mdp 1 => batch.aggregatevisitreal state action) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateVisitCount_le_horizon","label":"aggregateVisitCount_le_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateVisitCount_le_horizon","description":"One generated episode contributes at most one pooled visit per stage.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-9e088d5c2942","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7342,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:646"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateVisitCount_le_horizon {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (state : State) (action : Action) : batch.transitionCountSummary.aggregateVisitCount state action <= mdp.horizon","missing":[],"search":"aggregatevisitcount_le_horizon banditrlproof.finitehorizonrl.episodebatch.aggregatevisitcount_le_horizon one generated episode contributes at most one pooled visit per stage. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_aggregateTransitionResidual_le_two_mul_horizon","label":"abs_aggregateTransitionResidual_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_aggregateTransitionResidual_le_two_mul_horizon","description":"The one-episode transition residual is bounded by twice the horizon.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-5f530b28cf8a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7343,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:679"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_aggregateTransitionResidual_le_two_mul_horizon {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (state : State) (action : Action) (nextState : State) : |batch.aggregateTransitionResidual state action nextState| <= 2 * (mdp.horizon : Real)","missing":[],"search":"abs_aggregatetransitionresidual_le_two_mul_horizon banditrlproof.finitehorizonrl.episodebatch.abs_aggregatetransitionresidual_le_two_mul_horizon the one-episode transition residual is bounded by twice the horizon. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionResidual_compensated_hasMGFUpperBoundAt","label":"iidEpisodeBatchMeasure_one_aggregateTransitionResidual_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionResidual_compensated_hasMGFUpperBoundAt","description":"The exact one-episode generated batch law inherits the trace-level compensated MGF; no independent empirical batch is introduced.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-ab5201640bbe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7344,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:723"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_one_aggregateTransitionResidual_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (state : State) (action : Action) (nextState : State) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun batch : EpisodeBatch mdp 1 => tilt * batch.aggregateTransitionResidual state action nextState - (tilt ^ 2 / 8) * batch.aggregateVisitReal state action) 1 0 (policy.iidEpisodeBatchMeasure initialState 1)","missing":[],"search":"iidepisodebatchmeasure_one_aggregatetransitionresidual_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure_one_aggregatetransitionresidual_compensated_hasmgfupperboundat the exact one-episode generated batch law inherits the trace-level compensated mgf; no independent empirical batch is introduced. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration","label":"batchPrefixFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration","description":"Canonical finite-prefix filtration presented with the actual `frestrictLe` history object consumed by `AdaptiveEpisodeBatchSource`. It is extensionally the usual product filtration, but this presentation keeps conditional kernels definitionally aligned with the generated source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-96c4f19b2b2d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7345,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:811"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def batchPrefixFiltration {mdp : MDP State Action} (episodes : Nat) : @Filtration (EpisodeBatchTrajectory mdp episodes) Nat _ MeasurableSpace.pi where","missing":[],"search":"batchprefixfiltration banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.batchprefixfiltration canonical finite-prefix filtration presented with the actual `frestrictle` history object consumed by `adaptiveepisodebatchsource`. it is extensionally the usual product filtration, but this presentation keeps conditional kernels definitionally aligned with the generated source. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.hasCondMGFUpperBoundAt_of_condExpKernel_map_eq","label":"hasCondMGFUpperBoundAt_of_condExpKernel_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.hasCondMGFUpperBoundAt_of_condExpKernel_map_eq","description":"Transport pointwise fixed-tilt MGF bounds through an identified conditional kernel map. This is the fixed-MGF analogue of the repository's centered sub-Gaussian condExpKernel bridge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-226c2de5a71c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7346,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:833"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasCondMGFUpperBoundAt_of_condExpKernel_map_eq {Omega : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond : MeasurableSpace Omega) (hmcond : mcond <= mOmega) (X : Omega -> Real) (hX : @Measurable Omega Real mOmega inferInstance X) (tilt budget : Real) (target : Omega -> Measure Real) (hintegrable : ∀ s, Integrable (fun omega => Real.exp (s * X omega)) mu) (hmap : Filter.Eventually (fun omega => @Measure.map Omega Real mOmega inferInstance X ((@ProbabilityTheory.condExpKernel Omega mOmega _ mu _ mcond) omega) = target omega) (ae (@MeasureTheory.Measure.trim Omega mcond mOmega mu hmcond))) (htarget : Filter.Eventually (fun omega => Concentration.HasMGFUpperBoundAt id tilt budget (target omega)) (ae (@MeasureTheory.Measure.trim Omega mcond mOmega mu hmcond))) : Concentration.HasCondMGFUpperBoundAt mcond hmcond X tilt…","missing":[],"search":"hascondmgfupperboundat_of_condexpkernel_map_eq banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.hascondmgfupperboundat_of_condexpkernel_map_eq transport pointwise fixed-tilt mgf bounds through an identified conditional kernel map. this is the fixed-mgf analogue of the repository's centered sub-gaussian condexpkernel bridge. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.hasCondMGFUpperBoundAt_congr_measurableSpace","label":"hasCondMGFUpperBoundAt_congr_measurableSpace","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.hasCondMGFUpperBoundAt_congr_measurableSpace","description":"Transport a fixed-tilt conditional MGF certificate across propositionally equal conditioning measurable spaces.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-f178f412f047","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7347,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:873"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasCondMGFUpperBoundAt_congr_measurableSpace {Omega : Type*} [mOmega : MeasurableSpace Omega] [StandardBorelSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (mcond mcond' : MeasurableSpace Omega) (hmcond : mcond <= mOmega) (hmcond' : mcond' <= mOmega) (hspaces : mcond = mcond') (X : Omega -> Real) (tilt budget : Real) (h : Concentration.HasCondMGFUpperBoundAt mcond hmcond X tilt budget mu) : Concentration.HasCondMGFUpperBoundAt mcond' hmcond' X tilt budget mu","missing":[],"search":"hascondmgfupperboundat_congr_measurablespace banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.hascondmgfupperboundat_congr_measurablespace transport a fixed-tilt conditional mgf certificate across propositionally equal conditioning measurable spaces. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidualPrefix","label":"aggregateTransitionResidualPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidualPrefix","description":"Residual of the last batch visible in a finite generated prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-2d7e78af7355","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7348,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:890"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateTransitionResidualPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) (history : EpisodeBatchPrefix mdp 1 round) : Real","missing":[],"search":"aggregatetransitionresidualprefix banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidualprefix residual of the last batch visible in a finite generated prefix. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateVisitPrefix","label":"aggregateVisitPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateVisitPrefix","description":"Visit count of the last batch visible in a finite generated prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-33413d5d0e44","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7349,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:900"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def aggregateVisitPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (round : Nat) (history : EpisodeBatchPrefix mdp 1 round) : Real","missing":[],"search":"aggregatevisitprefix banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatevisitprefix visit count of the last batch visible in a finite generated prefix. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionResidualPrefix","label":"measurable_aggregateTransitionResidualPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionResidualPrefix","description":"theorem measurable_aggregateTransitionResidualPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) : Measurable (source.aggregateTransitionResidualPrefix state action nextState round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-8fa4f467334c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7350,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:909"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionResidualPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) : Measurable (source.aggregateTransitionResidualPrefix state action nextState round)","missing":[],"search":"measurable_aggregatetransitionresidualprefix banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_aggregatetransitionresidualprefix theorem measurable_aggregatetransitionresidualprefix {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (state : state) (action : action) (nextstate : state) (round : nat) : measurable (source.aggregatetransitionresidualprefix state action nextstate round) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateVisitPrefix","label":"measurable_aggregateVisitPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateVisitPrefix","description":"theorem measurable_aggregateVisitPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateVisitPrefix state action round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-cd5815c2df2b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7351,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:922"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateVisitPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateVisitPrefix state action round)","missing":[],"search":"measurable_aggregatevisitprefix banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_aggregatevisitprefix theorem measurable_aggregatevisitprefix {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (state : state) (action : action) (round : nat) : measurable (source.aggregatevisitprefix state action round) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidualIncrement","label":"aggregateTransitionResidualIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidualIncrement","description":"The actual one-episode transition residual at adaptive coordinate `round`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-423cbf3438f1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7352,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:933"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateTransitionResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"aggregatetransitionresidualincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidualincrement the actual one-episode transition residual at adaptive coordinate `round`. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateVisitIncrement","label":"aggregateVisitIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateVisitIncrement","description":"The actual pooled visit count at adaptive coordinate `round`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-43f1421db70a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7353,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:943"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def aggregateVisitIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"aggregatevisitincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatevisitincrement the actual pooled visit count at adaptive coordinate `round`. definition compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionResidualIncrement","label":"measurable_aggregateTransitionResidualIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionResidualIncrement","description":"theorem measurable_aggregateTransitionResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) : Measurable (source.aggregateTransitionResidualIncrement state action nextState round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-96cc2d43fba6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7354,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:953"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) : Measurable (source.aggregateTransitionResidualIncrement state action nextState round)","missing":[],"search":"measurable_aggregatetransitionresidualincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_aggregatetransitionresidualincrement theorem measurable_aggregatetransitionresidualincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (state : state) (action : action) (nextstate : state) (round : nat) : measurable (source.aggregatetransitionresidualincrement state action nextstate round) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateVisitIncrement","label":"measurable_aggregateVisitIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateVisitIncrement","description":"theorem measurable_aggregateVisitIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateVisitIncrement state action round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-ede2487bb5a1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7355,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:964"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateVisitIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateVisitIncrement state action round)","missing":[],"search":"measurable_aggregatevisitincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_aggregatevisitincrement theorem measurable_aggregatevisitincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (state : state) (action : action) (round : nat) : measurable (source.aggregatevisitincrement state action round) theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_compensated_stronglyAdapted_piLE","label":"aggregateTransitionResidual_compensated_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_compensated_stronglyAdapted_piLE","description":"The compensated residual process is adapted to the canonical prefix filtration.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-9668a74ce8ac","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7356,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:975"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionResidual_compensated_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (tilt varianceCoeff : Real) : StronglyAdapted (batchPrefixFiltration (mdp := mdp) 1) (fun round trajectory => tilt * source.aggregateTransitionResidualIncrement state action nextState round trajectory - varianceCoeff * source.aggregateVisitIncrement state action round trajectory)","missing":[],"search":"aggregatetransitionresidual_compensated_stronglyadapted_pile banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidual_compensated_stronglyadapted_pile the compensated residual process is adapted to the canonical prefix filtration. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_batchStatistic","label":"trajectoryMeasure_condDistrib_batchStatistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_batchStatistic","description":"Any measurable real statistic of the next generated episode has exactly the mapped history kernel as its regular conditional law. This reusable bridge is stated for the whole generated batch, so downstream confidence proofs do not replace the recurrent process by an offline batch model.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-8b40ee590b38","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7357,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1008"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_batchStatistic {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (statistic : EpisodeBatch mdp episodes -> Real) (hstatistic : Measurable statistic) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => statistic (trajectory (n + 1))) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] (source.batchKernel n).map statistic","missing":[],"search":"trajectorymeasure_conddistrib_batchstatistic banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conddistrib_batchstatistic any measurable real statistic of the next generated episode has exactly the mapped history kernel as its regular conditional law. this reusable bridge is stated for the whole generated batch, so downstream confidence proofs do not replace the recurrent process by an offline batch model. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_batchStatistic_eq_batchKernel","label":"condExpKernel_map_batchStatistic_eq_batchKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_batchStatistic_eq_batchKernel","description":"Trimmed conditional-expectation-kernel law for an arbitrary measurable real statistic of the next generated episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-6fdec9e6666f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7358,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1055"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_batchStatistic_eq_batchKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [StandardBorelSpace (EpisodeBatchTrajectory mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) (statistic : EpisodeBatch mdp episodes -> Real) (hstatistic : Measurable statistic) : Filter.Eventually (fun trajectory : EpisodeBatchTrajectory mdp episodes => Measure.map (fun path : EpisodeBatchTrajectory mdp episodes => statistic (path (n + 1))) (ProbabilityTheory.condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (EpisodeBatchPrefix mdp episodes n)).comap (Preorder.frestrictLe n)) trajectory) = ((source.batchKernel n).map statistic) (Preorder.frestrictLe n trajectory)) (ae (source.trajectoryMeasure.trim (Preorder.measur…","missing":[],"search":"condexpkernel_map_batchstatistic_eq_batchkernel banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.condexpkernel_map_batchstatistic_eq_batchkernel trimmed conditional-expectation-kernel law for an arbitrary measurable real statistic of the next generated episode. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_exp_mul_aggregateTransitionResidual_compensatedIncrement","label":"integrable_exp_mul_aggregateTransitionResidual_compensatedIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_exp_mul_aggregateTransitionResidual_compensatedIncrement","description":"Every compensated adaptive residual increment is globally exponentially integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-8d3d17521302","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7359,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1095"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_mul_aggregateTransitionResidual_compensatedIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (round : Nat) (tilt varianceCoeff s : Real) : Integrable (fun trajectory : EpisodeBatchTrajectory mdp 1 => Real.exp (s * (tilt * source.aggregateTransitionResidualIncrement state action nextState round trajectory - varianceCoeff * source.aggregateVisitIncrement state action round trajectory))) source.trajectoryMeasure","missing":[],"search":"integrable_exp_mul_aggregatetransitionresidual_compensatedincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.integrable_exp_mul_aggregatetransitionresidual_compensatedincrement every compensated adaptive residual increment is globally exponentially integrable. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_succ_compensated_hasCondMGFUpperBoundAt","label":"aggregateTransitionResidual_succ_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_succ_compensated_hasCondMGFUpperBoundAt","description":"At every successor episode, the actual transition residual has the same visit-charged fixed-tilt conditional MGF under the recurrent source law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-97fcf30b02b4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7360,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionResidual_succ_compensated_hasCondMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (n : Nat) (tilt : Real) : Concentration.HasCondMGFUpperBoundAt (mΩ := MeasurableSpace.pi) (batchPrefixFiltration (mdp := mdp) 1 n) ((batchPrefixFiltration (mdp := mdp) 1).le n) (fun trajectory => tilt * source.aggregateTransitionResidualIncrement state action nextState (n + 1) trajectory - (tilt ^ 2 / 8) * source.aggregateVisitIncrement state action (n + 1) trajectory) 1 0 source.trajectoryMeasure","missing":[],"search":"aggregatetransitionresidual_succ_compensated_hascondmgfupperboundat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidual_succ_compensated_hascondmgfupperboundat at every successor episode, the actual transition residual has the same visit-charged fixed-tilt conditional mgf under the recurrent source law. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_zero_compensated_hasMGFUpperBoundAt","label":"aggregateTransitionResidual_zero_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_zero_compensated_hasMGFUpperBoundAt","description":"Coordinate zero has the same compensated MGF certificate under the exact initial marginal of the generated recurrent trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-7ded9963e037","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7361,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1262"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionResidual_zero_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory : EpisodeBatchTrajectory mdp 1 => tilt * source.aggregateTransitionResidualIncrement state action nextState 0 trajectory - (tilt ^ 2 / 8) * source.aggregateVisitIncrement state action 0 trajectory) 1 0 source.trajectoryMeasure","missing":[],"search":"aggregatetransitionresidual_zero_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionresidual_zero_compensated_hasmgfupperboundat coordinate zero has the same compensated mgf certificate under the exact initial marginal of the generated recurrent trajectory. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateTransitionResidualIncrement_eq_prefixAggregateResidual","label":"sum_aggregateTransitionResidualIncrement_eq_prefixAggregateResidual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateTransitionResidualIncrement_eq_prefixAggregateResidual","description":"Summing the adaptive one-batch residuals through coordinate `round` recovers the exact pooled numerator minus its true transition mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-5529fd4b3a7c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7362,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1295"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_aggregateTransitionResidualIncrement_eq_prefixAggregateResidual {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (trajectory : EpisodeBatchTrajectory mdp 1) (round : Nat) (state : State) (action : Action) (nextState : State) : (∑ i ∈ Finset.range (round + 1), source.aggregateTransitionResidualIncrement state action nextState i trajectory) = (adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState : Real) - (adaptiveCumulativeAggregateVisitCountAt trajectory round state action : Real) * (mdp.transition (state, action)).real {nextState}","missing":[],"search":"sum_aggregatetransitionresidualincrement_eq_prefixaggregateresidual banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.sum_aggregatetransitionresidualincrement_eq_prefixaggregateresidual summing the adaptive one-batch residuals through coordinate `round` recovers the exact pooled numerator minus its true transition mass. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateVisitIncrement_eq_prefixAggregateVisitCount","label":"sum_aggregateVisitIncrement_eq_prefixAggregateVisitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateVisitIncrement_eq_prefixAggregateVisitCount","description":"The compensator sum is literally the aggregate visit denominator used by the recurrent planner at the same generated prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-7b7e998a7551","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7363,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1342"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_aggregateVisitIncrement_eq_prefixAggregateVisitCount {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (trajectory : EpisodeBatchTrajectory mdp 1) (round : Nat) (state : State) (action : Action) : (∑ i ∈ Finset.range (round + 1), source.aggregateVisitIncrement state action i trajectory) = (adaptiveCumulativeAggregateVisitCountAt trajectory round state action : Real)","missing":[],"search":"sum_aggregatevisitincrement_eq_prefixaggregatevisitcount banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.sum_aggregatevisitincrement_eq_prefixaggregatevisitcount the compensator sum is literally the aggregate visit denominator used by the recurrent planner at the same generated prefix. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_aggregateTransitionResidualSum_ge_inter_visitSum_le","label":"measure_aggregateTransitionResidualSum_ge_inter_visitSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_aggregateTransitionResidualSum_ge_inter_visitSum_le","description":"Fixed-tilt, actual-count upper tail for one pooled transition coordinate on a finite prefix of the generated recurrent process. The random count is retained in the event; no expected-occupancy lower bound is substituted.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-44f41ea4e9c8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7364,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1366"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_aggregateTransitionResidualSum_ge_inter_visitSum_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (rounds : Nat) (tilt threshold visitBudget : Real) (htilt : 0 < tilt) : source.trajectoryMeasure {trajectory | threshold <= ∑ i ∈ Finset.range rounds, source.aggregateTransitionResidualIncrement state action nextState i trajectory ∧ (∑ i ∈ Finset.range rounds, source.aggregateVisitIncrement state action i trajectory) <= visitBudget} <= ENNReal.ofReal (Real.exp (-tilt * threshold + (tilt ^ 2 / 8) * visitBudget))","missing":[],"search":"measure_aggregatetransitionresidualsum_ge_inter_visitsum_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measure_aggregatetransitionresidualsum_ge_inter_visitsum_le fixed-tilt, actual-count upper tail for one pooled transition coordinate on a finite prefix of the generated recurrent process. the random count is retained in the event; no expected-occupancy lower bound is substituted. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le","label":"measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le","description":"Two-sided version of the exact-count prefix tail. Both signs are proved from the same generated conditional law; the factor two is only the final finite union.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisamesourceconfidence/index.html#decl-b7b373c66da2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","order":7365,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence.lean:1419"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (state : State) (action : Action) (nextState : State) (rounds : Nat) (tilt threshold visitBudget : Real) (htilt : 0 < tilt) : source.trajectoryMeasure {trajectory | threshold <= |∑ i ∈ Finset.range rounds, source.aggregateTransitionResidualIncrement state action nextState i trajectory| ∧ (∑ i ∈ Finset.range rounds, source.aggregateVisitIncrement state action i trajectory) <= visitBudget} <= 2 * ENNReal.ofReal (Real.exp (-tilt * threshold + (tilt ^ 2 / 8) * visitBudget))","missing":[],"search":"measure_abs_aggregatetransitionresidualsum_ge_inter_visitsum_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measure_abs_aggregatetransitionresidualsum_ge_inter_visitsum_le two-sided version of the exact-count prefix tail. both signs are proved from the same generated conditional law; the factor two is only the final finite union. theorem compiled","shard":"modules/02eed265a80dad25.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.BernsteinCoordinateIndex","label":"BernsteinCoordinateIndex","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.BernsteinCoordinateIndex","description":"One peeled singleton-transition query. `count` encodes the positive actual visit count `count + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-91d935795de4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7366,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure BernsteinCoordinateIndex (mdp : MDP State Action) (episodes : Nat) where","missing":[],"search":"bernsteincoordinateindex banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.bernsteincoordinateindex one peeled singleton-transition query. `count` encodes the positive actual visit count `count + 1`. structure compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.OptimalTailIndex","label":"OptimalTailIndex","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.OptimalTailIndex","description":"One peeled optimal-tail scalar query.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-9134736cebfb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7367,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure OptimalTailIndex (mdp : MDP State Action) (episodes : Nat) where","missing":[],"search":"optimaltailindex banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.optimaltailindex one peeled optimal-tail scalar query. structure compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.bernsteinCoordinateFailureEvent","label":"bernsteinCoordinateFailureEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.bernsteinCoordinateFailureEvent","description":"noncomputable def bernsteinCoordinateFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (logBudget : Real) (index : BernsteinCoordinateIndex mdp episodes) : Set (EpisodeBatchTrajectory mdp 1)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-340fa72e0b7d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7368,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bernsteinCoordinateFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (logBudget : Real) (index : BernsteinCoordinateIndex mdp episodes) : Set (EpisodeBatchTrajectory mdp 1)","missing":[],"search":"bernsteincoordinatefailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.bernsteincoordinatefailureevent noncomputable def bernsteincoordinatefailureevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (logbudget : real) (index : bernsteincoordinateindex mdp episodes) : set (episodebatchtrajectory mdp 1) definition compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.optimalTailFailureEvent","label":"optimalTailFailureEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.optimalTailFailureEvent","description":"noncomputable def optimalTailFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (logBudget : Real) (index : OptimalTailIndex mdp episodes) : Set (EpisodeBatchTrajectory mdp 1)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-000236cd0cd3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7369,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalTailFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (logBudget : Real) (index : OptimalTailIndex mdp episodes) : Set (EpisodeBatchTrajectory mdp 1)","missing":[],"search":"optimaltailfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.optimaltailfailureevent noncomputable def optimaltailfailureevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (logbudget : real) (index : optimaltailindex mdp episodes) : set (episodebatchtrajectory mdp 1) definition compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.simultaneousTransitionFailureEvent","label":"simultaneousTransitionFailureEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.simultaneousTransitionFailureEvent","description":"The complete finite same-source transition confidence failure set.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-801166d556fa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7370,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:89"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousTransitionFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (episodes : Nat) (logBudget : Real) : Set (EpisodeBatchTrajectory mdp 1)","missing":[],"search":"simultaneoustransitionfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.simultaneoustransitionfailureevent the complete finite same-source transition confidence failure set. definition compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_bernsteinCoordinateFailureEvent_le","label":"trajectoryMeasure_bernsteinCoordinateFailureEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_bernsteinCoordinateFailureEvent_le","description":"theorem trajectoryMeasure_bernsteinCoordinateFailureEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (logBudget : Real) (hlog : 0 < logBudget) (index : BernsteinCoordinateIndex mdp episodes) : source.trajectoryMeasur…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-72897535871a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7371,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_bernsteinCoordinateFailureEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (logBudget : Real) (hlog : 0 < logBudget) (index : BernsteinCoordinateIndex mdp episodes) : source.trajectoryMeasure (bernsteinCoordinateFailureEvent source logBudget index) <= 2 * ENNReal.ofReal (Real.exp (-2 * logBudget))","missing":[],"search":"trajectorymeasure_bernsteincoordinatefailureevent_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.trajectorymeasure_bernsteincoordinatefailureevent_le theorem trajectorymeasure_bernsteincoordinatefailureevent_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] [standardborelspace (episodebatch mdp 1)] [standardborelspace (episodebatchtrajectory mdp 1)] (source : adaptiveepisodebatchsource mdp initialstate 1) (logbudget : real) (hlog : 0 < logbudget) (index : bernsteincoordinateindex mdp episodes) : source.trajectorymeasure (bernsteincoordinatefailureevent source logbudget index) <= 2 * ennreal.ofreal (real.exp (-2 * logbudget)) theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt","label":"functionalTilt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt","description":"noncomputable def functionalTilt (logBudget visitBudget : Real) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-7b59ff0e5a78","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7372,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:152"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def functionalTilt (logBudget visitBudget : Real) : Real","missing":[],"search":"functionaltilt banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.functionaltilt noncomputable def functionaltilt (logbudget visitbudget : real) : real definition compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt_pos","label":"functionalTilt_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt_pos","description":"theorem functionalTilt_pos {logBudget visitBudget : Real} (hlog : 0 < logBudget) (hvisit : 0 < visitBudget) : 0 < functionalTilt logBudget visitBudget","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-8e81ea13e320","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7373,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:155"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem functionalTilt_pos {logBudget visitBudget : Real} (hlog : 0 < logBudget) (hvisit : 0 < visitBudget) : 0 < functionalTilt logBudget visitBudget","missing":[],"search":"functionaltilt_pos banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.functionaltilt_pos theorem functionaltilt_pos {logbudget visitbudget : real} (hlog : 0 < logbudget) (hvisit : 0 < visitbudget) : 0 < functionaltilt logbudget visitbudget theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt_exponent_eq","label":"functionalTilt_exponent_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt_exponent_eq","description":"theorem functionalTilt_exponent_eq {logBudget visitBudget : Real} (hvisit : 0 < visitBudget) : -functionalTilt logBudget visitBudget * (logBudget * Real.sqrt visitBudget) + (functionalTilt logBudget visitBudget ^ 2 / 8) * visitBudget = -2 * logBudget ^ 2","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-5b17345f0676","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7374,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem functionalTilt_exponent_eq {logBudget visitBudget : Real} (hvisit : 0 < visitBudget) : -functionalTilt logBudget visitBudget * (logBudget * Real.sqrt visitBudget) + (functionalTilt logBudget visitBudget ^ 2 / 8) * visitBudget = -2 * logBudget ^ 2","missing":[],"search":"functionaltilt_exponent_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.functionaltilt_exponent_eq theorem functionaltilt_exponent_eq {logbudget visitbudget : real} (hvisit : 0 < visitbudget) : -functionaltilt logbudget visitbudget * (logbudget * real.sqrt visitbudget) + (functionaltilt logbudget visitbudget ^ 2 / 8) * visitbudget = -2 * logbudget ^ 2 theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_optimalTailFailureEvent_le","label":"trajectoryMeasure_optimalTailFailureEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_optimalTailFailureEvent_le","description":"theorem trajectoryMeasure_optimalTailFailureEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (hreward : ∀ state action, |mdp.reward state action| <= 1) (logBudget : Real) (hlog : 0 < logBudget) (index : OptimalTailIn…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-ee402c2713de","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7375,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:177"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_optimalTailFailureEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (hreward : ∀ state action, |mdp.reward state action| <= 1) (logBudget : Real) (hlog : 0 < logBudget) (index : OptimalTailIndex mdp episodes) : source.trajectoryMeasure (optimalTailFailureEvent source logBudget index) <= 2 * ENNReal.ofReal (Real.exp (-2 * logBudget ^ 2))","missing":[],"search":"trajectorymeasure_optimaltailfailureevent_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.trajectorymeasure_optimaltailfailureevent_le theorem trajectorymeasure_optimaltailfailureevent_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] [standardborelspace (episodebatch mdp 1)] [standardborelspace (episodebatchtrajectory mdp 1)] (source : adaptiveepisodebatchsource mdp initialstate 1) (hreward : ∀ state action, |mdp.reward state action| <= 1) (logbudget : real) (hlog : 0 < logbudget) (index : optimaltailindex mdp episodes) : source.trajectorymeasure (optimaltailfailureevent source logbudget index) <= 2 * ennreal.ofreal (real.exp (-2 * logbudget ^ 2)) theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_simultaneousTransitionFailureEvent_le","label":"trajectoryMeasure_simultaneousTransitionFailureEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_simultaneousTransitionFailureEvent_le","description":"A transparent finite-union bound. The later UCBVI specialization proves that the displayed right hand side is at most its allotted fraction of `delta`; keeping this intermediate theorem exact makes the event accounting auditable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-70f2990c0d0c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7376,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:214"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_simultaneousTransitionFailureEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (hreward : ∀ state action, |mdp.reward state action| <= 1) (episodes : Nat) (logBudget : Real) (hlog : 0 < logBudget) : source.trajectoryMeasure (simultaneousTransitionFailureEvent source episodes logBudget) <= (Fintype.card (BernsteinCoordinateIndex mdp episodes) : ENNReal) * (2 * ENNReal.ofReal (Real.exp (-2 * logBudget))) + (Fintype.card (OptimalTailIndex mdp episodes) : ENNReal) * (2 * ENNReal.ofReal (Real.exp (-2 * logBudget ^ 2)))","missing":[],"search":"trajectorymeasure_simultaneoustransitionfailureevent_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.trajectorymeasure_simultaneoustransitionfailureevent_le a transparent finite-union bound. the later ucbvi specialization proves that the displayed right hand side is at most its allotted fraction of `delta`; keeping this intermediate theorem exact makes the event accounting auditable. theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_coordinateResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","label":"abs_coordinateResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_coordinateResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","description":"Outside the joint event, every peeled singleton residual is strictly below its variance-sensitive threshold whenever the peel equals the actual count.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-b16a69ffb78e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7377,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_coordinateResidual_lt_of_not_mem_simultaneousTransitionFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {logBudget : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes logBudget) (index : BernsteinCoordinateIndex mdp episodes) (hcount : (∑ i ∈ Finset.range (index.round + 1), source.aggregateVisitIncrement index.state index.action i trajectory) = (index.count + 1 : Nat)) : |∑ i ∈ Finset.range (index.round + 1), source.aggregateTransitionResidualIncrement index.state index.action index.nextState i trajectory| < bernsteinCoordinateThreshold logBudget (mdp.transitionCoordinateVariance index.state index.action index.nextState * (index.count + 1 : Nat))","missing":[],"search":"abs_coordinateresidual_lt_of_not_mem_simultaneoustransitionfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.abs_coordinateresidual_lt_of_not_mem_simultaneoustransitionfailureevent outside the joint event, every peeled singleton residual is strictly below its variance-sensitive threshold whenever the peel equals the actual count. theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_optimalTailResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","label":"abs_optimalTailResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_optimalTailResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","description":"The sharp optimal-tail scalar coordinate is extracted from the same joint event, again at its exact actual-count peel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvisimultaneousconfidence/index.html#decl-ba26f2ddf9fb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","order":7378,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence.lean:287"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_optimalTailResidual_lt_of_not_mem_simultaneousTransitionFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) {episodes : Nat} {logBudget : Real} {trajectory : EpisodeBatchTrajectory mdp 1} (htrajectory : trajectory ∉ simultaneousTransitionFailureEvent source episodes logBudget) (index : OptimalTailIndex mdp episodes) (hcount : (∑ i ∈ Finset.range (index.round + 1), source.aggregateVisitIncrement index.state index.action i trajectory) = (index.count + 1 : Nat)) : |∑ i ∈ Finset.range (index.round + 1), source.aggregateTransitionFunctionalResidualIncrement (mdp.optimalTailProbe index.stage) index.state index.action i trajectory| < logBudget * Real.sqrt (index.count + 1 : Nat)","missing":[],"search":"abs_optimaltailresidual_lt_of_not_mem_simultaneoustransitionfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.adaptiveepisodebatchsource.abs_optimaltailresidual_lt_of_not_mem_simultaneoustransitionfailureevent the sharp optimal-tail scalar coordinate is extracted from the same joint event, again at its exact actual-count peel. theorem compiled","shard":"modules/211bebb9426c7476.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalRegretBound","label":"canonicalRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalRegretBound","description":"Frozen Azar--Osband--Munos-shaped Hoeffding UCBVI-CH bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-8dfb4601e83c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7379,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalRegretBound (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Real","missing":[],"search":"canonicalregretbound banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.canonicalregretbound frozen azar--osband--munos-shaped hoeffding ucbvi-ch bound. definition compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalRegretBound_nonneg","label":"canonicalRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalRegretBound_nonneg","description":"theorem canonicalRegretBound_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= canonicalRegretBound (State := State) (Action := Action) mdp episodes delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-12e67d03e06a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7380,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem canonicalRegretBound_nonneg (mdp : MDP State Action) (episodes : Nat) (delta : Real) : 0 <= canonicalRegretBound (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"canonicalregretbound_nonneg banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.canonicalregretbound_nonneg theorem canonicalregretbound_nonneg (mdp : mdp state action) (episodes : nat) (delta : real) : 0 <= canonicalregretbound (state := state) (action := action) mdp episodes delta theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.alignmentFailureEvent","label":"alignmentFailureEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.alignmentFailureEvent","description":"noncomputable def alignmentFailureEvent (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp 1)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-e4854a5b37f8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7381,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def alignmentFailureEvent (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp 1)","missing":[],"search":"alignmentfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.alignmentfailureevent noncomputable def alignmentfailureevent (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) : set (episodebatchtrajectory mdp 1) definition compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalFailureEvent","label":"canonicalFailureEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalFailureEvent","description":"Complete terminal failure set: same-source confidence, Bellman innovation, and the measure-zero generated-record alignment complement.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-be5f1765aaed","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7382,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalFailureEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp 1)","missing":[],"search":"canonicalfailureevent banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.canonicalfailureevent complete terminal failure set: same-source confidence, bellman innovation, and the measure-zero generated-record alignment complement. definition compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_alignmentFailureEvent_eq_zero","label":"recurrentSource_trajectoryMeasure_alignmentFailureEvent_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_alignmentFailureEvent_eq_zero","description":"theorem recurrentSource_trajectoryMeasure_alignmentFailureEvent_eq_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) : (recurrentSource mdp initialState defaultState episodes delta).trajectoryMeasure (alignmentFailureEvent mdp defaultState episodes delta) = 0","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-3f041388410a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7383,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_trajectoryMeasure_alignmentFailureEvent_eq_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) : (recurrentSource mdp initialState defaultState episodes delta).trajectoryMeasure (alignmentFailureEvent mdp defaultState episodes delta) = 0","missing":[],"search":"recurrentsource_trajectorymeasure_alignmentfailureevent_eq_zero banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_alignmentfailureevent_eq_zero theorem recurrentsource_trajectorymeasure_alignmentfailureevent_eq_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (defaultstate : state) (episodes : nat) (delta : real) (hhorizon : 0 < mdp.horizon) : (recurrentsource mdp initialstate defaultstate episodes delta).trajectorymeasure (alignmentfailureevent mdp defaultstate episodes delta) = 0 theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_canonicalFailureEvent_le","label":"recurrentSource_trajectoryMeasure_canonicalFailureEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_canonicalFailureEvent_le","description":"The complete generated terminal failure set costs at most `delta`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-17b6ff616d43","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7384,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem recurrentSource_trajectoryMeasure_canonicalFailureEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) : (recurrentSource mdp initialState defaultState episodes delta).trajectoryMeasure (canonicalFailureEvent (initialState := initialState) defaultState episodes delta) <= ENNReal.ofReal delta","missing":[],"search":"recurrentsource_trajectorymeasure_canonicalfailureevent_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_canonicalfailureevent_le the complete generated terminal failure set costs at most `delta`. theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_shift_recurrentBellmanInnovation_eq","label":"sum_shift_recurrentBellmanInnovation_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_shift_recurrentBellmanInnovation_eq","description":"private theorem sum_shift_recurrentBellmanInnovation_eq (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) : (Finset.range (episodes - 1)).sum (fun n => recurrentBellmanInnovationProcess mdp defaultState episodes delta (n + 1) trajectory) = (Finset.range episodes).sum (fun round => recurrentBellmanInnovationProcess mdp defaultState episodes del…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-7b9b1f45474a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7385,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private theorem sum_shift_recurrentBellmanInnovation_eq (mdp : MDP State Action) (defaultState : State) (episodes : Nat) (delta : Real) (trajectory : EpisodeBatchTrajectory mdp 1) : (Finset.range (episodes - 1)).sum (fun n => recurrentBellmanInnovationProcess mdp defaultState episodes delta (n + 1) trajectory) = (Finset.range episodes).sum (fun round => recurrentBellmanInnovationProcess mdp defaultState episodes delta round trajectory)","missing":[],"search":"sum_shift_recurrentbellmaninnovation_eq banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sum_shift_recurrentbellmaninnovation_eq private theorem sum_shift_recurrentbellmaninnovation_eq (mdp : mdp state action) (defaultstate : state) (episodes : nat) (delta : real) (trajectory : episodebatchtrajectory mdp 1) : (finset.range (episodes - 1)).sum (fun n => recurrentbellmaninnovationprocess mdp defaultstate episodes delta (n + 1) trajectory) = (finset.range episodes).sum (fun round => recurrentbellmaninnovationprocess mdp defaultstate episodes delta round trajectory) theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sqrt_totalSteps_mul_card_eq_paper","label":"sqrt_totalSteps_mul_card_eq_paper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sqrt_totalSteps_mul_card_eq_paper","description":"theorem sqrt_totalSteps_mul_card_eq_paper (mdp : MDP State Action) (episodes : Nat) : Real.sqrt (totalSteps mdp episodes : Nat) * Real.sqrt ((Fintype.card State * Fintype.card Action : Nat) : Real) = Real.sqrt mdp.horizon * Real.sqrt ((Fintype.card State : Real) * Fintype.card Action * episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-9d21f5240445","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7386,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sqrt_totalSteps_mul_card_eq_paper (mdp : MDP State Action) (episodes : Nat) : Real.sqrt (totalSteps mdp episodes : Nat) * Real.sqrt ((Fintype.card State * Fintype.card Action : Nat) : Real) = Real.sqrt mdp.horizon * Real.sqrt ((Fintype.card State : Real) * Fintype.card Action * episodes)","missing":[],"search":"sqrt_totalsteps_mul_card_eq_paper banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.sqrt_totalsteps_mul_card_eq_paper theorem sqrt_totalsteps_mul_card_eq_paper (mdp : mdp state action) (episodes : nat) : real.sqrt (totalsteps mdp episodes : nat) * real.sqrt ((fintype.card state * fintype.card action : nat) : real) = real.sqrt mdp.horizon * real.sqrt ((fintype.card state : real) * fintype.card action * episodes) theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret_le_canonicalRegretBound_of_not_mem","label":"cumulativeEpisodePseudoRegret_le_canonicalRegretBound_of_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret_le_canonicalRegretBound_of_not_mem","description":"Pathwise generated regret bound outside the proved terminal failure set.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-48b3b0ce4d55","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7387,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeEpisodePseudoRegret_le_canonicalRegretBound_of_not_mem (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) (trajectory : EpisodeBatchTrajectory mdp 1) (htrajectory : trajectory ∉ canonicalFailureEvent (initialState := initialState) defaultState episodes delta) : cumulativeEpisodePseudoRegret (recurrentSource mdp initialState defaultState episodes delta) defaultState episodes trajectory <= canonicalRegretBound (State := State) (Action := Action) mdp episodes delta","missing":[],"search":"cumulativeepisodepseudoregret_le_canonicalregretbound_of_not_mem banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.cumulativeepisodepseudoregret_le_canonicalregretbound_of_not_mem pathwise generated regret bound outside the proved terminal failure set. theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","label":"recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","description":"Canonical generated-policy high-probability regret terminal.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbviterminal/index.html#decl-1968b6ce60d8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","order":7388,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITerminal.lean:403"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","finite-horizon-rl"]],"statement":"theorem recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (defaultState : State) (episodes : Nat) (delta : Real) (hhorizon : 0 < mdp.horizon) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hrewardNonneg : ∀ state action, 0 <= mdp.reward state action) (hrewardOne : ∀ state action, mdp.reward state action <= 1) : let source := recurrentSource mdp initialState defaultState episodes delta source.trajectoryMeasure {trajectory | canonicalRegretBound (State := State) (Action := Action) mdp episodes delta < cumulativeEpisodePseudoRegret source defaultState episodes trajectory} <= ENNReal.ofReal delta","missing":[],"search":"recurrentsource_trajectorymeasure_cumulativeepisodepseudoregret_gt_canonicalregretbound_le banditrlproof.finitehorizonrl.adaptivecumulativehoeffdingucbvi.recurrentsource_trajectorymeasure_cumulativeepisodepseudoregret_gt_canonicalregretbound_le canonical generated-policy high-probability regret terminal. theorem compiled","shard":"modules/774753115e42f47c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":["finite-horizon-rl"]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalTailProbe","label":"optimalTailProbe","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalTailProbe","description":"The normalized optimal continuation-value probe used by the sharp transition-value confidence coordinate. The symmetric shift avoids requiring nonnegativity merely to apply Hoeffding.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-f846744581bf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7389,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalTailProbe (mdp : MDP State Action) (stage : Fin mdp.horizon) (nextState : State) : Real","missing":[],"search":"optimaltailprobe banditrlproof.finitehorizonrl.mdp.optimaltailprobe the normalized optimal continuation-value probe used by the sharp transition-value confidence coordinate. the symmetric shift avoids requiring nonnegativity merely to apply hoeffding. definition compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalTailProbe_mem_Icc","label":"optimalTailProbe_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalTailProbe_mem_Icc","description":"Under rewards in `[-1,1]`, the normalized optimal tail is in `[0,1]`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-cac14a36fcd7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7390,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalTailProbe_mem_Icc (mdp : MDP State Action) (hreward : ∀ state action, |mdp.reward state action| <= 1) (stage : Fin mdp.horizon) (nextState : State) : mdp.optimalTailProbe stage nextState ∈ Set.Icc (0 : Real) 1","missing":[],"search":"optimaltailprobe_mem_icc banditrlproof.finitehorizonrl.mdp.optimaltailprobe_mem_icc under rewards in `[-1,1]`, the normalized optimal tail is in `[0,1]`. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead","label":"transitionFunctionalResidualHead","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead","description":"Linear transition residual for a fixed bounded probe. Writing it as the finite linear combination of singleton residuals makes its relation to the coordinate event explicit.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-561f19cb0af7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7391,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:75"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def transitionFunctionalResidualHead (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (currentState : State) (head : Action × State) : Real","missing":[],"search":"transitionfunctionalresidualhead banditrlproof.finitehorizonrl.mdp.transitionfunctionalresidualhead linear transition residual for a fixed bounded probe. writing it as the finite linear combination of singleton residuals makes its relation to the coordinate event explicit. definition compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom","label":"transitionFunctionalResidualFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom","description":"The same linear probe on a remaining generated trace.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-65e14554b69d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7392,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def transitionFunctionalResidualFrom (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (remaining : Nat) (currentState : State) (trace : StepTrace Action State remaining) : Real","missing":[],"search":"transitionfunctionalresidualfrom banditrlproof.finitehorizonrl.mdp.transitionfunctionalresidualfrom the same linear probe on a remaining generated trace. definition compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionFunctionalResidualFrom","label":"measurable_transitionFunctionalResidualFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionFunctionalResidualFrom","description":"theorem measurable_transitionFunctionalResidualFrom (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionFunctionalResidualFrom feature targetState targetAction remaining p.1 p.2)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-e352783ef22c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7393,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:93"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionFunctionalResidualFrom (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (remaining : Nat) : Measurable (fun p : State × StepTrace Action State remaining => mdp.transitionFunctionalResidualFrom feature targetState targetAction remaining p.1 p.2)","missing":[],"search":"measurable_transitionfunctionalresidualfrom banditrlproof.finitehorizonrl.mdp.measurable_transitionfunctionalresidualfrom theorem measurable_transitionfunctionalresidualfrom (mdp : mdp state action) (feature : state -> real) (targetstate : state) (targetaction : action) (remaining : nat) : measurable (fun p : state × steptrace action state remaining => mdp.transitionfunctionalresidualfrom feature targetstate targetaction remaining p.1 p.2) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue_eq_sum_measureReal","label":"transitionValue_eq_sum_measureReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionValue_eq_sum_measureReal","description":"A finite-state transition integral is its finite singleton-mass sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-133cd7c62f2a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7394,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionValue_eq_sum_measureReal (mdp : MDP State Action) (feature : State -> Real) (state : State) (action : Action) : mdp.transitionValue feature state action = ∑ nextState : State, feature nextState * (mdp.transition (state, action)).real {nextState}","missing":[],"search":"transitionvalue_eq_sum_measurereal banditrlproof.finitehorizonrl.mdp.transitionvalue_eq_sum_measurereal a finite-state transition integral is its finite singleton-mass sum. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead_eq","label":"transitionFunctionalResidualHead_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead_eq","description":"On a visited cell, the linear coordinate residual is exactly the centered bounded feature; away from the cell it is zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-f3ebde17c59b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7395,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionFunctionalResidualHead_eq (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (currentState : State) (head : Action × State) : mdp.transitionFunctionalResidualHead feature targetState targetAction currentState head = if currentState = targetState ∧ head.1 = targetAction then feature head.2 - mdp.transitionValue feature targetState targetAction else 0","missing":[],"search":"transitionfunctionalresidualhead_eq banditrlproof.finitehorizonrl.mdp.transitionfunctionalresidualhead_eq on a visited cell, the linear coordinate residual is exactly the centered bounded feature; away from the cell it is zero. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom_succ","label":"transitionFunctionalResidualFrom_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom_succ","description":"The linear probe obeys the same chronological recursion as every singleton coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-63aef8cf30ae","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7396,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:154"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionFunctionalResidualFrom_succ (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (remaining : Nat) (currentState : State) (trace : StepTrace Action State (remaining + 1)) : mdp.transitionFunctionalResidualFrom feature targetState targetAction (remaining + 1) currentState trace = mdp.transitionFunctionalResidualHead feature targetState targetAction currentState (trace 0) + mdp.transitionFunctionalResidualFrom feature targetState targetAction remaining (trace 0).2 (Fin.tail trace)","missing":[],"search":"transitionfunctionalresidualfrom_succ banditrlproof.finitehorizonrl.mdp.transitionfunctionalresidualfrom_succ the linear probe obeys the same chronological recursion as every singleton coordinate. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","label":"transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","description":"A `[0,1]` probe has the standard Hoeffding compensated MGF on one true transition draw.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-435319c52065","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7397,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:173"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (currentState : State) (chosenAction : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun nextState => tilt * mdp.transitionFunctionalResidualHead feature targetState targetAction currentState (chosenAction, nextState) - (tilt ^ 2 / 8) * mdp.transitionVisitHead targetState targetAction currentState (chosenAction, nextState)) 1 0 (mdp.transition (currentState, chosenAction))","missing":[],"search":"transitionfunctionalresidualhead_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.transitionfunctionalresidualhead_compensated_hasmgfupperboundat a `[0,1]` probe has the standard hoeffding compensated mgf on one true transition draw. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","label":"actionStateKernel_transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","description":"theorem actionStateKernel_transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (currentState : State) (stage : Fin mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * md…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-c064d2f16c1c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7398,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (currentState : State) (stage : Fin mdp.horizon) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun head : Action × State => tilt * mdp.transitionFunctionalResidualHead feature targetState targetAction currentState head - (tilt ^ 2 / 8) * mdp.transitionVisitHead targetState targetAction currentState head) 1 0 (policy.actionStateKernel stage currentState)","missing":[],"search":"actionstatekernel_transitionfunctionalresidualhead_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.actionstatekernel_transitionfunctionalresidualhead_compensated_hasmgfupperboundat theorem actionstatekernel_transitionfunctionalresidualhead_compensated_hasmgfupperboundat (mdp : mdp state action) (policy : markovpolicy mdp) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (targetstate : state) (targetaction : action) (currentstate : state) (stage : fin mdp.horizon) (tilt : real) : concentration.hasmgfupperboundat (fun head : action × state => tilt * mdp.transitionfunctionalresidualhead feature targetstate targetaction currentstate head - (tilt ^ 2 / 8) * mdp.transitionvisithead targetstate targetaction currentstate head) 1 0 (policy.actionstatekernel stage currentstate) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","label":"trajectoryKernelRemaining_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","description":"theorem trajectoryKernelRemaining_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (tilt : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (currentState : State) : Concentration.HasMGFUpperBoundAt (fu…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-f85a051dfa9f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7399,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (tilt : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (currentState : State) : Concentration.HasMGFUpperBoundAt (fun trace => tilt * mdp.transitionFunctionalResidualFrom feature targetState targetAction remaining currentState trace - (tilt ^ 2 / 8) * mdp.transitionVisitFrom targetState targetAction remaining currentState trace) 1 0 (policy.trajectoryKernelRemaining remaining hremaining currentState)","missing":[],"search":"trajectorykernelremaining_transitionfunctionalresidual_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorykernelremaining_transitionfunctionalresidual_compensated_hasmgfupperboundat theorem trajectorykernelremaining_transitionfunctionalresidual_compensated_hasmgfupperboundat (mdp : mdp state action) (policy : markovpolicy mdp) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (targetstate : state) (targetaction : action) (tilt : real) (remaining : nat) (hremaining : remaining <= mdp.horizon) (currentstate : state) : concentration.hasmgfupperboundat (fun trace => tilt * mdp.transitionfunctionalresidualfrom feature targetstate targetaction remaining currentstate trace - (tilt ^ 2 / 8) * mdp.transitionvisitfrom targetstate targetaction remaining currentstate trace) 1 0 (policy.trajectorykernelremaining remaining hremaining currentstate) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","label":"trajectoryMeasure_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","description":"theorem trajectoryMeasure_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory => tilt *…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-f1cbe4b380b3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7400,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:428"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (targetState : State) (targetAction : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory => tilt * mdp.transitionFunctionalResidualFrom feature targetState targetAction mdp.horizon trajectory.1 trajectory.2 - (tilt ^ 2 / 8) * mdp.transitionVisitFrom targetState targetAction mdp.horizon trajectory.1 trajectory.2) 1 0 (policy.trajectoryMeasure initialState)","missing":[],"search":"trajectorymeasure_transitionfunctionalresidual_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.mdp.trajectorymeasure_transitionfunctionalresidual_compensated_hasmgfupperboundat theorem trajectorymeasure_transitionfunctionalresidual_compensated_hasmgfupperboundat (mdp : mdp state action) (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (targetstate : state) (targetaction : action) (tilt : real) : concentration.hasmgfupperboundat (fun trajectory => tilt * mdp.transitionfunctionalresidualfrom feature targetstate targetaction mdp.horizon trajectory.1 trajectory.2 - (tilt ^ 2 / 8) * mdp.transitionvisitfrom targetstate targetaction mdp.horizon trajectory.1 trajectory.2) 1 0 (policy.trajectorymeasure initialstate) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom_eq_aggregateFunctionalResidual","label":"transitionFunctionalResidualFrom_eq_aggregateFunctionalResidual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom_eq_aggregateFunctionalResidual","description":"theorem transitionFunctionalResidualFrom_eq_aggregateFunctionalResidual (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (trajectory : State × StepTrace Action State mdp.horizon) : mdp.transitionFunctionalResidualFrom feature targetState targetAction mdp.horizon trajectory.1 trajectory.2 = ∑ nextState : State, feature nextState * EpisodeBatch.aggregateTransitionResidua…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-592f2177a2c8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7401,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:487"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionFunctionalResidualFrom_eq_aggregateFunctionalResidual (mdp : MDP State Action) (feature : State -> Real) (targetState : State) (targetAction : Action) (trajectory : State × StepTrace Action State mdp.horizon) : mdp.transitionFunctionalResidualFrom feature targetState targetAction mdp.horizon trajectory.1 trajectory.2 = ∑ nextState : State, feature nextState * EpisodeBatch.aggregateTransitionResidual (mdp.episodeBatchOfTrajectories 1 (fun _ => trajectory)) targetState targetAction nextState","missing":[],"search":"transitionfunctionalresidualfrom_eq_aggregatefunctionalresidual banditrlproof.finitehorizonrl.mdp.transitionfunctionalresidualfrom_eq_aggregatefunctionalresidual theorem transitionfunctionalresidualfrom_eq_aggregatefunctionalresidual (mdp : mdp state action) (feature : state -> real) (targetstate : state) (targetaction : action) (trajectory : state × steptrace action state mdp.horizon) : mdp.transitionfunctionalresidualfrom feature targetstate targetaction mdp.horizon trajectory.1 trajectory.2 = ∑ nextstate : state, feature nextstate * episodebatch.aggregatetransitionresidual (mdp.episodebatchoftrajectories 1 (fun _ => trajectory)) targetstate targetaction nextstate theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateTransitionFunctionalResidual","label":"aggregateTransitionFunctionalResidual","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateTransitionFunctionalResidual","description":"A bounded transition functional is another finite coordinate of the same aggregate singleton-residual table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-2772e8063e40","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7402,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:510"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateTransitionFunctionalResidual {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (feature : State -> Real) (state : State) (action : Action) : Real","missing":[],"search":"aggregatetransitionfunctionalresidual banditrlproof.finitehorizonrl.episodebatch.aggregatetransitionfunctionalresidual a bounded transition functional is another finite coordinate of the same aggregate singleton-residual table. definition compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateTransitionFunctionalResidual","label":"measurable_aggregateTransitionFunctionalResidual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateTransitionFunctionalResidual","description":"theorem measurable_aggregateTransitionFunctionalResidual {mdp : MDP State Action} (feature : State -> Real) (state : State) (action : Action) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.aggregateTransitionFunctionalResidual feature state action)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-5cda48658362","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7403,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:517"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionFunctionalResidual {mdp : MDP State Action} (feature : State -> Real) (state : State) (action : Action) : Measurable (fun batch : EpisodeBatch mdp 1 => batch.aggregateTransitionFunctionalResidual feature state action)","missing":[],"search":"measurable_aggregatetransitionfunctionalresidual banditrlproof.finitehorizonrl.episodebatch.measurable_aggregatetransitionfunctionalresidual theorem measurable_aggregatetransitionfunctionalresidual {mdp : mdp state action} (feature : state -> real) (state : state) (action : action) : measurable (fun batch : episodebatch mdp 1 => batch.aggregatetransitionfunctionalresidual feature state action) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_aggregateTransitionFunctionalResidual_le","label":"abs_aggregateTransitionFunctionalResidual_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_aggregateTransitionFunctionalResidual_le","description":"theorem abs_aggregateTransitionFunctionalResidual_le {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) : |batch.aggregateTransitionFunctionalResidual feature state action| <= 2 * Fintype.card State * (mdp.horizon : Real)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-ddeccab4d4ea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7404,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:529"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_aggregateTransitionFunctionalResidual_le {mdp : MDP State Action} (batch : EpisodeBatch mdp 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) : |batch.aggregateTransitionFunctionalResidual feature state action| <= 2 * Fintype.card State * (mdp.horizon : Real)","missing":[],"search":"abs_aggregatetransitionfunctionalresidual_le banditrlproof.finitehorizonrl.episodebatch.abs_aggregatetransitionfunctionalresidual_le theorem abs_aggregatetransitionfunctionalresidual_le {mdp : mdp state action} (batch : episodebatch mdp 1) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (state : state) (action : action) : |batch.aggregatetransitionfunctionalresidual feature state action| <= 2 * fintype.card state * (mdp.horizon : real) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.integral_aggregateEmpiricalTransitionKernel_eq_sum_div","label":"integral_aggregateEmpiricalTransitionKernel_eq_sum_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.integral_aggregateEmpiricalTransitionKernel_eq_sum_div","description":"At a positive pooled count, empirical transition integration is the exact normalized finite count sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-b76781c13eea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7405,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:570"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_aggregateEmpiricalTransitionKernel_eq_sum_div {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (feature : State -> Real) (state : State) (action : Action) (hpos : 0 < summary.aggregateVisitCount state action) : (∫ nextState, feature nextState ∂summary.aggregateEmpiricalTransitionKernel defaultState (state, action)) = (∑ nextState : State, feature nextState * (summary.aggregateTransitionCount state action nextState : Real)) / (summary.aggregateVisitCount state action : Real)","missing":[],"search":"integral_aggregateempiricaltransitionkernel_eq_sum_div banditrlproof.finitehorizonrl.transitioncountsummary.integral_aggregateempiricaltransitionkernel_eq_sum_div at a positive pooled count, empirical transition integration is the exact normalized finite count sum. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionFunctionalResidual_compensated_hasMGFUpperBoundAt","label":"iidEpisodeBatchMeasure_one_aggregateTransitionFunctionalResidual_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionFunctionalResidual_compensated_hasMGFUpperBoundAt","description":"theorem iidEpisodeBatchMeasure_one_aggregateTransitionFunctionalResidual_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun batch : Episod…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-d5bf4b48babc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7406,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:598"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_one_aggregateTransitionFunctionalResidual_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun batch : EpisodeBatch mdp 1 => tilt * batch.aggregateTransitionFunctionalResidual feature state action - (tilt ^ 2 / 8) * batch.aggregateVisitReal state action) 1 0 (policy.iidEpisodeBatchMeasure initialState 1)","missing":[],"search":"iidepisodebatchmeasure_one_aggregatetransitionfunctionalresidual_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure_one_aggregatetransitionfunctionalresidual_compensated_hasmgfupperboundat theorem iidepisodebatchmeasure_one_aggregatetransitionfunctionalresidual_compensated_hasmgfupperboundat {mdp : mdp state action} (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (state : state) (action : action) (tilt : real) : concentration.hasmgfupperboundat (fun batch : episodebatch mdp 1 => tilt * batch.aggregatetransitionfunctionalresidual feature state action - (tilt ^ 2 / 8) * batch.aggregatevisitreal state action) 1 0 (policy.iidepisodebatchmeasure initialstate 1) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidualPrefix","label":"aggregateTransitionFunctionalResidualPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidualPrefix","description":"Fixed-round functional residual read from the actual generated batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-aa22fb4a8bf1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7407,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:688"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateTransitionFunctionalResidualPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) (history : EpisodeBatchPrefix mdp 1 round) : Real","missing":[],"search":"aggregatetransitionfunctionalresidualprefix banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionfunctionalresidualprefix fixed-round functional residual read from the actual generated batch. definition compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidualIncrement","label":"aggregateTransitionFunctionalResidualIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidualIncrement","description":"noncomputable def aggregateTransitionFunctionalResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-96180769aacc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7408,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:697"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def aggregateTransitionFunctionalResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) (trajectory : EpisodeBatchTrajectory mdp 1) : Real","missing":[],"search":"aggregatetransitionfunctionalresidualincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionfunctionalresidualincrement noncomputable def aggregatetransitionfunctionalresidualincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (state : state) (action : action) (round : nat) (trajectory : episodebatchtrajectory mdp 1) : real definition compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionFunctionalResidualPrefix","label":"measurable_aggregateTransitionFunctionalResidualPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionFunctionalResidualPrefix","description":"theorem measurable_aggregateTransitionFunctionalResidualPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateTransitionFunctionalResidualPrefix feature state action round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-551b37535baa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7409,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:707"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionFunctionalResidualPrefix {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateTransitionFunctionalResidualPrefix feature state action round)","missing":[],"search":"measurable_aggregatetransitionfunctionalresidualprefix banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_aggregatetransitionfunctionalresidualprefix theorem measurable_aggregatetransitionfunctionalresidualprefix {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (state : state) (action : action) (round : nat) : measurable (source.aggregatetransitionfunctionalresidualprefix feature state action round) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionFunctionalResidualIncrement","label":"measurable_aggregateTransitionFunctionalResidualIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionFunctionalResidualIncrement","description":"theorem measurable_aggregateTransitionFunctionalResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateTransitionFunctionalResidualIncrement feature state action round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-3d1101e12c26","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7410,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:721"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_aggregateTransitionFunctionalResidualIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (round : Nat) : Measurable (source.aggregateTransitionFunctionalResidualIncrement feature state action round)","missing":[],"search":"measurable_aggregatetransitionfunctionalresidualincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurable_aggregatetransitionfunctionalresidualincrement theorem measurable_aggregatetransitionfunctionalresidualincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (state : state) (action : action) (round : nat) : measurable (source.aggregatetransitionfunctionalresidualincrement feature state action round) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_compensated_stronglyAdapted_piLE","label":"aggregateTransitionFunctionalResidual_compensated_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_compensated_stronglyAdapted_piLE","description":"theorem aggregateTransitionFunctionalResidual_compensated_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (tilt varianceCoeff : Real) : StronglyAdapted (batchPrefixFiltration (mdp := mdp) 1) (fun round trajectory => tilt * source.aggrega…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-1bd2947d796e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7411,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:733"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionFunctionalResidual_compensated_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (state : State) (action : Action) (tilt varianceCoeff : Real) : StronglyAdapted (batchPrefixFiltration (mdp := mdp) 1) (fun round trajectory => tilt * source.aggregateTransitionFunctionalResidualIncrement feature state action round trajectory - varianceCoeff * source.aggregateVisitIncrement state action round trajectory)","missing":[],"search":"aggregatetransitionfunctionalresidual_compensated_stronglyadapted_pile banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionfunctionalresidual_compensated_stronglyadapted_pile theorem aggregatetransitionfunctionalresidual_compensated_stronglyadapted_pile {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (state : state) (action : action) (tilt variancecoeff : real) : stronglyadapted (batchprefixfiltration (mdp := mdp) 1) (fun round trajectory => tilt * source.aggregatetransitionfunctionalresidualincrement feature state action round trajectory - variancecoeff * source.aggregatevisitincrement state action round trajectory) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_exp_mul_aggregateTransitionFunctionalResidual_compensatedIncrement","label":"integrable_exp_mul_aggregateTransitionFunctionalResidual_compensatedIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_exp_mul_aggregateTransitionFunctionalResidual_compensatedIncrement","description":"theorem integrable_exp_mul_aggregateTransitionFunctionalResidual_compensatedIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (round : Nat) (tilt varianceCoeff s : Real) : Integrable…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-92df48b2f3c5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7412,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:758"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_exp_mul_aggregateTransitionFunctionalResidual_compensatedIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (round : Nat) (tilt varianceCoeff s : Real) : Integrable (fun trajectory : EpisodeBatchTrajectory mdp 1 => Real.exp (s * (tilt * source.aggregateTransitionFunctionalResidualIncrement feature state action round trajectory - varianceCoeff * source.aggregateVisitIncrement state action round trajectory))) source.trajectoryMeasure","missing":[],"search":"integrable_exp_mul_aggregatetransitionfunctionalresidual_compensatedincrement banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.integrable_exp_mul_aggregatetransitionfunctionalresidual_compensatedincrement theorem integrable_exp_mul_aggregatetransitionfunctionalresidual_compensatedincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (state : state) (action : action) (round : nat) (tilt variancecoeff s : real) : integrable (fun trajectory : episodebatchtrajectory mdp 1 => real.exp (s * (tilt * source.aggregatetransitionfunctionalresidualincrement feature state action round trajectory - variancecoeff * source.aggregatevisitincrement state action round trajectory))) source.trajectorymeasure theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_succ_compensated_hasCondMGFUpperBoundAt","label":"aggregateTransitionFunctionalResidual_succ_compensated_hasCondMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_succ_compensated_hasCondMGFUpperBoundAt","description":"theorem aggregateTransitionFunctionalResidual_succ_compensated_hasCondMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real)…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-4f49905338aa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7413,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:840"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionFunctionalResidual_succ_compensated_hasCondMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (n : Nat) (tilt : Real) : Concentration.HasCondMGFUpperBoundAt (mΩ := MeasurableSpace.pi) (batchPrefixFiltration (mdp := mdp) 1 n) ((batchPrefixFiltration (mdp := mdp) 1).le n) (fun trajectory => tilt * source.aggregateTransitionFunctionalResidualIncrement feature state action (n + 1) trajectory - (tilt ^ 2 / 8) * source.aggregateVisitIncrement state action (n + 1) trajectory) 1 0 source.trajectoryMeasure","missing":[],"search":"aggregatetransitionfunctionalresidual_succ_compensated_hascondmgfupperboundat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionfunctionalresidual_succ_compensated_hascondmgfupperboundat theorem aggregatetransitionfunctionalresidual_succ_compensated_hascondmgfupperboundat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] [standardborelspace (episodebatch mdp 1)] [standardborelspace (episodebatchtrajectory mdp 1)] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (state : state) (action : action) (n : nat) (tilt : real) : concentration.hascondmgfupperboundat (mω := measurablespace.pi) (batchprefixfiltration (mdp := mdp) 1 n) ((batchprefixfiltration (mdp := mdp) 1).le n) (fun trajectory => tilt * source.aggregatetransitionfunctionalresidualincrement feature state action (n + 1) trajectory - (tilt ^ 2 / 8) * source.aggregatevisitincrement state action (n + 1) trajectory) 1 0 source.trajectorymeasure theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_zero_compensated_hasMGFUpperBoundAt","label":"aggregateTransitionFunctionalResidual_zero_compensated_hasMGFUpperBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_zero_compensated_hasMGFUpperBoundAt","description":"theorem aggregateTransitionFunctionalResidual_zero_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun traject…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-bba580733b10","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7414,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:922"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionFunctionalResidual_zero_compensated_hasMGFUpperBoundAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (tilt : Real) : Concentration.HasMGFUpperBoundAt (fun trajectory : EpisodeBatchTrajectory mdp 1 => tilt * source.aggregateTransitionFunctionalResidualIncrement feature state action 0 trajectory - (tilt ^ 2 / 8) * source.aggregateVisitIncrement state action 0 trajectory) 1 0 source.trajectoryMeasure","missing":[],"search":"aggregatetransitionfunctionalresidual_zero_compensated_hasmgfupperboundat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionfunctionalresidual_zero_compensated_hasmgfupperboundat theorem aggregatetransitionfunctionalresidual_zero_compensated_hasmgfupperboundat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (state : state) (action : action) (tilt : real) : concentration.hasmgfupperboundat (fun trajectory : episodebatchtrajectory mdp 1 => tilt * source.aggregatetransitionfunctionalresidualincrement feature state action 0 trajectory - (tilt ^ 2 / 8) * source.aggregatevisitincrement state action 0 trajectory) 1 0 source.trajectorymeasure theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateTransitionFunctionalResidualIncrement_eq_prefixResidual","label":"sum_aggregateTransitionFunctionalResidualIncrement_eq_prefixResidual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateTransitionFunctionalResidualIncrement_eq_prefixResidual","description":"Summing the generated functional-probe increments is exactly the finite linear combination of the cumulative singleton residuals.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-bb98c56cd62a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7415,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:954"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_aggregateTransitionFunctionalResidualIncrement_eq_prefixResidual {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (trajectory : EpisodeBatchTrajectory mdp 1) (round : Nat) (feature : State -> Real) (state : State) (action : Action) : (∑ i ∈ Finset.range (round + 1), source.aggregateTransitionFunctionalResidualIncrement feature state action i trajectory) = ∑ nextState : State, feature nextState * ((adaptiveCumulativeAggregateTransitionCountAt trajectory round state action nextState : Real) - (adaptiveCumulativeAggregateVisitCountAt trajectory round state action : Real) * (mdp.transition (state, action)).real {nextState})","missing":[],"search":"sum_aggregatetransitionfunctionalresidualincrement_eq_prefixresidual banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.sum_aggregatetransitionfunctionalresidualincrement_eq_prefixresidual summing the generated functional-probe increments is exactly the finite linear combination of the cumulative singleton residuals. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_eq_count_mul_transitionValue_sub","label":"aggregateTransitionFunctionalResidual_eq_count_mul_transitionValue_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_eq_count_mul_transitionValue_sub","description":"At positive actual count, the functional residual divided by that count is exactly `(P_hat-P) feature` for the empirical kernel consumed by the planner.","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-29ec918f14aa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7416,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:986"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem aggregateTransitionFunctionalResidual_eq_count_mul_transitionValue_sub {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (trajectory : EpisodeBatchTrajectory mdp 1) (round : Nat) (defaultState : State) (feature : State -> Real) (state : State) (action : Action) (hpos : 0 < adaptiveCumulativeAggregateVisitCountAt trajectory round state action) : (∑ i ∈ Finset.range (round + 1), source.aggregateTransitionFunctionalResidualIncrement feature state action i trajectory) = (adaptiveCumulativeAggregateVisitCountAt trajectory round state action : Real) * ((∫ nextState, feature nextState ∂TransitionCountSummary.aggregateEmpiricalTransitionKernel (adaptiveCumulativeEmpiricalModelStateAt trajectory round).1 defaultState (state, action)) - mdp.transitionValue feature state action)","missing":[],"search":"aggregatetransitionfunctionalresidual_eq_count_mul_transitionvalue_sub banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.aggregatetransitionfunctionalresidual_eq_count_mul_transitionvalue_sub at positive actual count, the functional residual divided by that count is exactly `(p_hat-p) feature` for the empirical kernel consumed by the planner. theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionFunctionalResidualSum_ge_inter_visitSum_le","label":"measure_abs_aggregateTransitionFunctionalResidualSum_ge_inter_visitSum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionFunctionalResidualSum_ge_inter_visitSum_le","description":"theorem measure_abs_aggregateTransitionFunctionalResidualSum_ge_inter_visitSum_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (…","url":"../modules/banditrlproof-rl-finitehorizonadaptivecumulativeucbvitransitionvalueconfidence/index.html#decl-b25092e13d3d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","order":7417,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence.lean:1056"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_abs_aggregateTransitionFunctionalResidualSum_ge_inter_visitSum_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [StandardBorelSpace (EpisodeBatch mdp 1)] [StandardBorelSpace (EpisodeBatchTrajectory mdp 1)] (source : AdaptiveEpisodeBatchSource mdp initialState 1) (feature : State -> Real) (hfeature : ∀ nextState, feature nextState ∈ Set.Icc (0 : Real) 1) (state : State) (action : Action) (rounds : Nat) (tilt threshold visitBudget : Real) (htilt : 0 < tilt) : source.trajectoryMeasure {trajectory | threshold <= |∑ i ∈ Finset.range rounds, source.aggregateTransitionFunctionalResidualIncrement feature state action i trajectory| ∧ (∑ i ∈ Finset.range rounds, source.aggregateVisitIncrement state action i trajectory) <= visitBudget} <= 2 * ENNReal.ofReal (Real.exp (-tilt * threshold + (tilt ^ 2 / 8) * visitBudget))","missing":[],"search":"measure_abs_aggregatetransitionfunctionalresidualsum_ge_inter_visitsum_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measure_abs_aggregatetransitionfunctionalresidualsum_ge_inter_visitsum_le theorem measure_abs_aggregatetransitionfunctionalresidualsum_ge_inter_visitsum_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] [standardborelspace (episodebatch mdp 1)] [standardborelspace (episodebatchtrajectory mdp 1)] (source : adaptiveepisodebatchsource mdp initialstate 1) (feature : state -> real) (hfeature : ∀ nextstate, feature nextstate ∈ set.icc (0 : real) 1) (state : state) (action : action) (rounds : nat) (tilt threshold visitbudget : real) (htilt : 0 < tilt) : source.trajectorymeasure {trajectory | threshold <= |∑ i ∈ finset.range rounds, source.aggregatetransitionfunctionalresidualincrement feature state action i trajectory| ∧ (∑ i ∈ finset.range rounds, source.aggregatevisitincrement state action i trajectory) <= visitbudget} <= 2 * ennreal.ofreal (real.exp (-tilt * threshold + (tilt ^ 2 / 8) * visitbudget)) theorem compiled","shard":"modules/7a25aa1e1906511c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_visitCount","label":"transitionCountSummary_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_visitCount","description":"Compressing a batch preserves every visit count.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-bae170019437","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7418,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionCountSummary_visitCount {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : batch.transitionCountSummary.visitCount stage state action = batch.visitCount stage state action","missing":[],"search":"transitioncountsummary_visitcount banditrlproof.finitehorizonrl.episodebatch.transitioncountsummary_visitcount compressing a batch preserves every visit count. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_empiricalTransitionPMF","label":"transitionCountSummary_empiricalTransitionPMF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_empiricalTransitionPMF","description":"The summary-normalized PMF is exactly the raw batch empirical PMF.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-69b767fdb45d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7419,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionCountSummary_empiricalTransitionPMF {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : batch.transitionCountSummary.empiricalTransitionPMF defaultState stage state action = batch.empiricalTransitionPMF defaultState stage state action","missing":[],"search":"transitioncountsummary_empiricaltransitionpmf banditrlproof.finitehorizonrl.episodebatch.transitioncountsummary_empiricaltransitionpmf the summary-normalized pmf is exactly the raw batch empirical pmf. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_empiricalTransitionKernel_real_singleton","label":"transitionCountSummary_empiricalTransitionKernel_real_singleton","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_empiricalTransitionKernel_real_singleton","description":"The summary kernel singleton mass is the named raw empirical mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-9601fe7737d6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7420,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionCountSummary_empiricalTransitionKernel_real_singleton {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (batch.transitionCountSummary.empiricalTransitionKernel defaultState stage (state, action)).real {nextState} = batch.empiricalTransitionMass defaultState stage state action nextState","missing":[],"search":"transitioncountsummary_empiricaltransitionkernel_real_singleton banditrlproof.finitehorizonrl.episodebatch.transitioncountsummary_empiricaltransitionkernel_real_singleton the summary kernel singleton mass is the named raw empirical mass. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPlan_upperValueRemaining_abs_le","label":"optimisticPlan_upperValueRemaining_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPlan_upperValueRemaining_abs_le","description":"The known-reward empirical plan has the same explicit linear value envelope as the raw empirical model when the fixed transition bonus is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-e13563f102d0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7421,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimisticPlan_upperValueRemaining_abs_le (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (rewardBound transitionBonus : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) : forall (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State), |(summary.optimisticPlan mdp defaultState transitionBonus).upperValueRemaining remaining hremaining state| <= empiricalFiniteBatchValueEnvelope rewardBound transitionBonus remaining","missing":[],"search":"optimisticplan_uppervalueremaining_abs_le banditrlproof.finitehorizonrl.transitioncountsummary.optimisticplan_uppervalueremaining_abs_le the known-reward empirical plan has the same explicit linear value envelope as the raw empirical model when the fixed transition bonus is nonnegative. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.EmpiricalOptimisticCalibration","label":"EmpiricalOptimisticCalibration","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.EmpiricalOptimisticCalibration","description":"Policy-local statistical calibration needed to make the fixed transition bonus cover all finite-state empirical transition-coordinate errors.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-e2f3c2acbbcf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7422,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure EmpiricalOptimisticCalibration {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta rewardBound transitionBonus : Real) : Prop where","missing":[],"search":"empiricaloptimisticcalibration banditrlproof.finitehorizonrl.markovpolicy.empiricaloptimisticcalibration policy-local statistical calibration needed to make the fixed transition bonus cover all finite-state empirical transition-coordinate errors. structure compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalOptimisticPlanCoordinateConfidence_of_not_mem","label":"empiricalOptimisticPlanCoordinateConfidence_of_not_mem","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalOptimisticPlanCoordinateConfidence_of_not_mem","description":"One good generated batch and one calibration contract produce coordinate confidence for the exact known-reward empirical plan used by the source.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-11d1bb7911f4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7423,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalOptimisticPlanCoordinateConfidence_of_not_mem {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (defaultState : State) (rewardBound transitionBonus : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (calibration : policy.EmpiricalOptimisticCalibration initialState episodes delta rewardBound transitionBonus) : (batch.transitionCountSummary.optimisticPlan mdp defaultState transitionBonus).CoordinateConfidence where","missing":[],"search":"empiricaloptimisticplancoordinateconfidence_of_not_mem banditrlproof.finitehorizonrl.markovpolicy.empiricaloptimisticplancoordinateconfidence_of_not_mem one good generated batch and one calibration contract produce coordinate confidence for the exact known-reward empirical plan used by the source. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionPMF","label":"exploratoryActionPMF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionPMF","description":"Action law that explores uniformly with the supplied probability and otherwise uses the deterministic optimistic table action.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-3c2956067c88","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7424,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:277"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryActionPMF {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) : PMF Action","missing":[],"search":"exploratoryactionpmf banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryactionpmf action law that explores uniformly with the supplied probability and otherwise uses the deterministic optimistic table action. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.explorationRate_mul_inv_card_le_exploratoryActionPMF","label":"explorationRate_mul_inv_card_le_exploratoryActionPMF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.explorationRate_mul_inv_card_le_exploratoryActionPMF","description":"Every action receives at least its uniform-exploration mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-ccc1e3daa1bd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7425,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explorationRate_mul_inv_card_le_exploratoryActionPMF {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (action : Action) : (explorationRate : ENNReal) * (Fintype.card Action : ENNReal)⁻¹ <= table.exploratoryActionPMF explorationRate hexplorationRate stage state action","missing":[],"search":"explorationrate_mul_inv_card_le_exploratoryactionpmf banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.explorationrate_mul_inv_card_le_exploratoryactionpmf every action receives at least its uniform-exploration mass. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy","label":"exploratoryPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy","description":"Markov behavior policy obtained by uniformly exploring around one table.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-1b31e73ca811","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7426,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:302"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryPolicy {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : MarkovPolicy mdp where","missing":[],"search":"exploratorypolicy banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratorypolicy markov behavior policy obtained by uniformly exploring around one table. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDEpisodeBatchKernel","label":"exploratoryIIDEpisodeBatchKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDEpisodeBatchKernel","description":"Iid generated episode-batch kernel indexed by exploratory table policies.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-05071b7ed60c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7427,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:317"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryIIDEpisodeBatchKernel {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : ProbabilityTheory.Kernel (DeterministicMarkovPolicyTable mdp) (EpisodeBatch mdp episodes)","missing":[],"search":"exploratoryiidepisodebatchkernel banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryiidepisodebatchkernel iid generated episode-batch kernel indexed by exploratory table policies. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDEpisodeBatchKernel_apply","label":"exploratoryIIDEpisodeBatchKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDEpisodeBatchKernel_apply","description":"theorem exploratoryIIDEpisodeBatchKernel_apply {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (table : DeterministicMarkovPolicyTable mdp) : exploratoryIIDEpisodeBatchKernel initialState episodes explorationRate hexplorationRate table = (table.exploratoryPolicy explorationRate hexplorati…","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-62b28da716f6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7428,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:329"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryIIDEpisodeBatchKernel_apply {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (table : DeterministicMarkovPolicyTable mdp) : exploratoryIIDEpisodeBatchKernel initialState episodes explorationRate hexplorationRate table = (table.exploratoryPolicy explorationRate hexplorationRate).iidEpisodeBatchMeasure initialState episodes","missing":[],"search":"exploratoryiidepisodebatchkernel_apply banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryiidepisodebatchkernel_apply theorem exploratoryiidepisodebatchkernel_apply {mdp : mdp state action} (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (table : deterministicmarkovpolicytable mdp) : exploratoryiidepisodebatchkernel initialstate episodes explorationrate hexplorationrate table = (table.exploratorypolicy explorationrate hexplorationrate).iidepisodebatchmeasure initialstate episodes theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt","label":"adaptiveEmpiricalOptimisticPlanAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt","description":"The known-reward empirical plan selected from one batch coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-2a26886a684f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7429,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:354"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveEmpiricalOptimisticPlanAt {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (round : Nat) : mdp.EstimatedModelPlan","missing":[],"search":"adaptiveempiricaloptimisticplanat banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticplanat the known-reward empirical plan selected from one batch coordinate. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticRecommendedExpectedRegret","label":"adaptiveEmpiricalOptimisticRecommendedExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticRecommendedExpectedRegret","description":"Sum of expected regrets of the optimistic policies recommended by each batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-47e7f3aa33b3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7430,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:363"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveEmpiricalOptimisticRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (rounds : Nat) : Real","missing":[],"search":"adaptiveempiricaloptimisticrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticrecommendedexpectedregret sum of expected regrets of the optimistic policies recommended by each batch. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticOccupancyRadiusSum","label":"adaptiveEmpiricalOptimisticOccupancyRadiusSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticOccupancyRadiusSum","description":"Sum of the confidence-produced occupancy radius bounds after each batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-b33cd0718564","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7431,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:375"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveEmpiricalOptimisticOccupancyRadiusSum {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (rounds : Nat) : Real","missing":[],"search":"adaptiveempiricaloptimisticoccupancyradiussum banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticoccupancyradiussum sum of the confidence-produced occupancy radius bounds after each batch. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource","label":"exploratorySource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource","description":"Concrete behavior source with uniform action exploration around every latest-batch optimistic table.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-9964633dabb3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7432,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:395"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratorySource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : AdaptiveEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"exploratorysource banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource concrete behavior source with uniform action exploration around every latest-batch optimistic table. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_successorPolicy","label":"exploratorySource_successorPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_successorPolicy","description":"The exploratory successor behavior is centered on the latest optimistic table.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-2ca6672451af","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7433,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:421"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : (exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate).successorPolicy n history = (successorTable defaultState transitionBonus n history).exploratoryPolicy explorationRate hexplorationRate","missing":[],"search":"exploratorysource_successorpolicy banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_successorpolicy the exploratory successor behavior is centered on the latest optimistic table. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurableSet_selectedExploratorySimultaneousCountBadEvent","label":"measurableSet_selectedExploratorySimultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurableSet_selectedExploratorySimultaneousCountBadEvent","description":"Measurability of finite-table-selected exploratory count events.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-821c07b97c76","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7434,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:436"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selectedExploratorySimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} {History : Type*} [MeasurableSpace History] (selector : History -> DeterministicMarkovPolicyTable mdp) (hselector : Measurable selector) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (delta : Real) : MeasurableSet {pair : History × EpisodeBatch mdp episodes | pair.2 ∈ ((selector pair.1).exploratoryPolicy explorationRate hexplorationRate).simultaneousCountBadEvent initialState episodes delta}","missing":[],"search":"measurableset_selectedexploratorysimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.measurableset_selectedexploratorysimultaneouscountbadevent measurability of finite-table-selected exploratory count events. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_measurableSet_successorSimultaneousCountBadEvent","label":"exploratorySource_measurableSet_successorSimultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_measurableSet_successorSimultaneousCountBadEvent","description":"Every selected successor event of the exploratory source is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-128cf20ffdf4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7435,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:471"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_measurableSet_successorSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (delta : Real) (n : Nat) : MeasurableSet (AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent (exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate) rounds delta n)","missing":[],"search":"exploratorysource_measurableset_successorsimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_measurableset_successorsimultaneouscountbadevent every selected successor event of the exploratory source is measurable. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_adaptiveSimultaneousCountConfidence","label":"exploratorySource_trajectoryMeasure_adaptiveSimultaneousCountConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_adaptiveSimultaneousCountConfidence","description":"The exploratory behavior source inherits the adaptive global count event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-3821fe632c25","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7436,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:496"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_adaptiveSimultaneousCountConfidence {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate MeasurableSet (behaviorSource.adaptiveSimultaneousCountBadEvent rounds delta) ∧ behaviorSource.trajectoryMeasure (behaviorSource.adaptiveSimultaneousCountBadEvent rounds delta) <= ENNReal.ofReal delta ∧ forall trajectory, trajectory ∉ behaviorSource.adaptiveSimultaneo…","missing":[],"search":"exploratorysource_trajectorymeasure_adaptivesimultaneouscountconfidence banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_adaptivesimultaneouscountconfidence the exploratory behavior source inherits the adaptive global count event. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_policyAt_succ_eq_adaptiveEmpiricalOptimisticPlanAt","label":"source_policyAt_succ_eq_adaptiveEmpiricalOptimisticPlanAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_policyAt_succ_eq_adaptiveEmpiricalOptimisticPlanAt","description":"The plan from batch `n` is exactly the source policy used at `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-3268ef7feca1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7437,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:531"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_policyAt_succ_eq_adaptiveEmpiricalOptimisticPlanAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (trajectory : EpisodeBatchTrajectory mdp episodes) (n : Nat) : let concreteSource := source mdp initialState episodes initialTable defaultState transitionBonus concreteSource.policyAt trajectory (n + 1) = (adaptiveEmpiricalOptimisticPlanAt (mdp := mdp) (episodes := episodes) trajectory defaultState transitionBonus n).optimisticPolicy","missing":[],"search":"source_policyat_succ_eq_adaptiveempiricaloptimisticplanat banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.source_policyat_succ_eq_adaptiveempiricaloptimisticplanat the plan from batch `n` is exactly the source policy used at `n + 1`. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.policyAt_batch_not_mem_simultaneousCountBadEvent","label":"policyAt_batch_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.policyAt_batch_not_mem_simultaneousCountBadEvent","description":"Outside the adaptive union, every batch avoids its selected local event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-bcb7005ee1f6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7438,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:555"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem policyAt_batch_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) {delta : Real} (trajectory : EpisodeBatchTrajectory mdp episodes) (htrajectory : trajectory ∉ source.adaptiveSimultaneousCountBadEvent rounds delta) (round : Fin rounds) : trajectory round ∉ (source.policyAt trajectory round).simultaneousCountBadEvent initialState episodes (multiBatchLocalDelta rounds delta)","missing":[],"search":"policyat_batch_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.policyat_batch_not_mem_simultaneouscountbadevent outside the adaptive union, every batch avoids its selected local event. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceCalibration","label":"SourceCalibration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceCalibration","description":"Calibration contract selected by the data-generating policy at each round.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-7d6d028909ca","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7439,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:589"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def SourceCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta rewardBound transitionBonus : Real) : Prop","missing":[],"search":"sourcecalibration banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.sourcecalibration calibration contract selected by the data-generating policy at each round. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_coordinateConfidenceAt_of_not_mem","label":"exploratorySource_coordinateConfidenceAt_of_not_mem","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_coordinateConfidenceAt_of_not_mem","description":"Every observed batch plan has coordinate confidence outside the global event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-6b99f28b4bfc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7440,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:603"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratorySource_coordinateConfidenceAt_of_not_mem {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (calibration : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate SourceCalibration behaviorSource rounds delta rewardBound transitionBonus) (trajectory : EpisodeBatchTrajectory mdp episodes) (htrajectory : trajectory ∉ (exploratorySource mdp initialState episodes initialTable defaultState transitionBonus e…","missing":[],"search":"exploratorysource_coordinateconfidenceat_of_not_mem banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_coordinateconfidenceat_of_not_mem every observed batch plan has coordinate confidence outside the global event. definition compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_optimism_and_recommendedExpectedRegret_of_not_mem","label":"exploratorySource_optimism_and_recommendedExpectedRegret_of_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_optimism_and_recommendedExpectedRegret_of_not_mem","description":"Outside the one adaptive event, every batch plan is optimistic and the finite sum of recommended optimistic-policy expected regrets is radius-controlled.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-675f6b797a06","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7441,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:652"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_optimism_and_recommendedExpectedRegret_of_not_mem {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (calibration : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate SourceCalibration behaviorSource rounds delta rewardBound transitionBonus) (trajectory : EpisodeBatchTrajectory mdp episodes) (htrajectory : trajectory ∉ (exploratorySource mdp initialState episodes initialTable defaultState transitionB…","missing":[],"search":"exploratorysource_optimism_and_recommendedexpectedregret_of_not_mem banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_optimism_and_recommendedexpectedregret_of_not_mem outside the one adaptive event, every batch plan is optimistic and the finite sum of recommended optimistic-policy expected regrets is radius-controlled. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","description":"Route endpoint: calibrated all-coordinate confidence, global optimism, and a finite recommended-policy expected-regret sum under one global delta event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticconfidence/index.html#decl-8d9e90f0499e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","order":7442,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticConfidence.lean:710"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (calibration : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate SourceCalibration behaviorSource rounds delta rewardBound transitionBonus) : let behavio…","missing":[],"search":"exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret route endpoint: calibrated all-coordinate confidence, global optimism, and a finite recommended-policy expected-regret sum under one global delta event. theorem compiled","shard":"modules/1b73d33366a0cecb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_const","label":"occupancySumRemaining_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_const","description":"A probability occupancy sum evaluates a constant stage cost exactly.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticoccupancyenvelope/index.html#decl-66ddb8ecba22","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","order":7443,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem MarkovPolicy.occupancySumRemaining_const {mdp : MDP State Action} (policy : MarkovPolicy mdp) (c : Real) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : policy.occupancySumRemaining (fun _remaining _hremaining _state => c) remaining hremaining mu = (remaining : Real) * c","missing":[],"search":"occupancysumremaining_const banditrlproof.finitehorizonrl.markovpolicy.occupancysumremaining_const a probability occupancy sum evaluates a constant stage cost exactly. theorem compiled","shard":"modules/b84f7effecf44232.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt_selectedRadiusRemaining","label":"adaptiveEmpiricalOptimisticPlanAt_selectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt_selectedRadiusRemaining","description":"The concrete known-reward empirical plan selects its fixed transition bonus.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticoccupancyenvelope/index.html#decl-d24b5c59a7ce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","order":7444,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveEmpiricalOptimisticPlanAt_selectedRadiusRemaining {mdp : MDP State Action} {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (round remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (adaptiveEmpiricalOptimisticPlanAt (mdp := mdp) (episodes := episodes) trajectory defaultState transitionBonus round).selectedRadiusRemaining remaining hremaining state = transitionBonus","missing":[],"search":"adaptiveempiricaloptimisticplanat_selectedradiusremaining banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticplanat_selectedradiusremaining the concrete known-reward empirical plan selects its fixed transition bonus. theorem compiled","shard":"modules/b84f7effecf44232.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","label":"adaptiveEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","description":"Each round's selected-radius occupancy term is the horizon times the fixed cost.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticoccupancyenvelope/index.html#decl-68f5e4b08c83","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","order":7445,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (round : Nat) : let plan := adaptiveEmpiricalOptimisticPlanAt (mdp := mdp) (episodes := episodes) trajectory defaultState transitionBonus round plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState = (mdp.horizon : Real) * (2 * transitionBonus)","missing":[],"search":"adaptiveempiricaloptimisticplanat_occupancyselectedradiusremaining_eq banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticplanat_occupancyselectedradiusremaining_eq each round's selected-radius occupancy term is the horizon times the fixed cost. theorem compiled","shard":"modules/b84f7effecf44232.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticOccupancyRadiusSum_eq","label":"adaptiveEmpiricalOptimisticOccupancyRadiusSum_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticOccupancyRadiusSum_eq","description":"The complete adaptive selected-radius occupancy sum has a closed form.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticoccupancyenvelope/index.html#decl-1f99d483f7a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","order":7446,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveEmpiricalOptimisticOccupancyRadiusSum_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (rounds : Nat) : adaptiveEmpiricalOptimisticOccupancyRadiusSum (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaultState transitionBonus rounds = (rounds : Real) * ((mdp.horizon : Real) * (2 * transitionBonus))","missing":[],"search":"adaptiveempiricaloptimisticoccupancyradiussum_eq banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticoccupancyradiussum_eq the complete adaptive selected-radius occupancy sum has a closed form. theorem compiled","shard":"modules/b84f7effecf44232.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_explicitRecommendedExpectedRegret_of_pathSupport_episodeThreshold","label":"exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_explicitRecommendedExpectedRegret_of_pathSupport_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_explicitRecommendedExpectedRegret_of_pathSupport_episodeThreshold","description":"Route endpoint: the path-support episode threshold now yields a fully explicit fixed-bonus bound for the recommended optimistic policies.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticoccupancyenvelope/index.html#decl-a12c96af70fe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","order":7447,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope.lean:122"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_explicitRecommendedExpectedRegret_of_pathSupport_episodeThreshold {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hthreshold : exploratoryPathCalibrationEpisodeThreshold m…","missing":[],"search":"exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_explicitrecommendedexpectedregret_of_pathsupport_episodethreshold banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_explicitrecommendedexpectedregret_of_pathsupport_episodethreshold route endpoint: the path-support episode threshold now yields a fully explicit fixed-bonus bound for the recommended optimistic policies. theorem compiled","shard":"modules/b84f7effecf44232.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary","label":"TransitionCountSummary","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary","description":"All stage/state/action/next-state transition counts from one batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-174af67f2fd7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7448,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev TransitionCountSummary (mdp : MDP State Action)","missing":[],"search":"transitioncountsummary banditrlproof.finitehorizonrl.transitioncountsummary all stage/state/action/next-state transition counts from one batch. abbreviation compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary","label":"transitionCountSummary","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary","description":"Compress an episode batch to the transition counts used by the planner.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-cf5334140af7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7449,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:48"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionCountSummary {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) : TransitionCountSummary mdp","missing":[],"search":"transitioncountsummary banditrlproof.finitehorizonrl.episodebatch.transitioncountsummary compress an episode batch to the transition counts used by the planner. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_transitionCountSummary","label":"measurable_transitionCountSummary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_transitionCountSummary","description":"The complete transition-count summary is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-685487dcb376","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7450,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionCountSummary {mdp : MDP State Action} {episodes : Nat} : Measurable (transitionCountSummary : EpisodeBatch mdp episodes -> TransitionCountSummary mdp)","missing":[],"search":"measurable_transitioncountsummary banditrlproof.finitehorizonrl.episodebatch.measurable_transitioncountsummary the complete transition-count summary is measurable. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.visitCount","label":"visitCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.visitCount","description":"Total visits represented by one state-action transition-count row.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-4ace61bbcf7e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7451,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:75"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def visitCount {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) : Nat","missing":[],"search":"visitcount banditrlproof.finitehorizonrl.transitioncountsummary.visitcount total visits represented by one state-action transition-count row. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionPMF","label":"empiricalTransitionPMF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionPMF","description":"Normalized empirical transition PMF, with an explicit zero-count fallback.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-ec687122f42a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7452,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:81"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalTransitionPMF {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : PMF State","missing":[],"search":"empiricaltransitionpmf banditrlproof.finitehorizonrl.transitioncountsummary.empiricaltransitionpmf normalized empirical transition pmf, with an explicit zero-count fallback. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionKernel","label":"empiricalTransitionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionKernel","description":"The summary-indexed empirical transition kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-67a098fd4a58","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7453,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalTransitionKernel {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel (State × Action) State","missing":[],"search":"empiricaltransitionkernel banditrlproof.finitehorizonrl.transitioncountsummary.empiricaltransitionkernel the summary-indexed empirical transition kernel. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionKernel_isMarkov","label":"empiricalTransitionKernel_isMarkov","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionKernel_isMarkov","description":"theorem empiricalTransitionKernel_isMarkov {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (stage : Fin mdp.horizon) : ProbabilityTheory.IsMarkovKernel (summary.empiricalTransitionKernel defaultState stage) where","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-3a504adad090","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7454,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:131"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionKernel_isMarkov {mdp : MDP State Action} (summary : TransitionCountSummary mdp) (defaultState : State) (stage : Fin mdp.horizon) : ProbabilityTheory.IsMarkovKernel (summary.empiricalTransitionKernel defaultState stage) where","missing":[],"search":"empiricaltransitionkernel_ismarkov banditrlproof.finitehorizonrl.transitioncountsummary.empiricaltransitionkernel_ismarkov theorem empiricaltransitionkernel_ismarkov {mdp : mdp state action} (summary : transitioncountsummary mdp) (defaultstate : state) (stage : fin mdp.horizon) : probabilitytheory.ismarkovkernel (summary.empiricaltransitionkernel defaultstate stage) where theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPlan","label":"optimisticPlan","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPlan","description":"Known-reward empirical-transition plan with zero reward radius and one fixed transition bonus at every coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-de57331c2ee4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7455,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticPlan (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (transitionBonus : Real) : mdp.EstimatedModelPlan where","missing":[],"search":"optimisticplan banditrlproof.finitehorizonrl.transitioncountsummary.optimisticplan known-reward empirical-transition plan with zero reward radius and one fixed transition bonus at every coordinate. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable","label":"DeterministicMarkovPolicyTable","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable","description":"A deterministic action choice at every stage and state.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-bfc068d901c9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7456,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev DeterministicMarkovPolicyTable (mdp : MDP State Action)","missing":[],"search":"deterministicmarkovpolicytable banditrlproof.finitehorizonrl.deterministicmarkovpolicytable a deterministic action choice at every stage and state. abbreviation compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy","label":"toMarkovPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy","description":"Interpret a deterministic action table as a Markov policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-f57d44246f5d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7457,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:168"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def toMarkovPolicy {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) : MarkovPolicy mdp where","missing":[],"search":"tomarkovpolicy banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.tomarkovpolicy interpret a deterministic action table as a markov policy. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPolicyTable","label":"optimisticPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPolicyTable","description":"The deterministic optimistic action table computed from a count summary.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-615ec83cfc07","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7458,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticPolicyTable (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (transitionBonus : Real) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"optimisticpolicytable banditrlproof.finitehorizonrl.transitioncountsummary.optimisticpolicytable the deterministic optimistic action table computed from a count summary. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPolicyTable_toMarkovPolicy","label":"optimisticPolicyTable_toMarkovPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPolicyTable_toMarkovPolicy","description":"The action-table interpretation is exactly the plan's optimistic policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-80a913d0b584","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7459,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimisticPolicyTable_toMarkovPolicy (mdp : MDP State Action) (summary : TransitionCountSummary mdp) (defaultState : State) (transitionBonus : Real) : (summary.optimisticPolicyTable mdp defaultState transitionBonus).toMarkovPolicy = (summary.optimisticPlan mdp defaultState transitionBonus).optimisticPolicy","missing":[],"search":"optimisticpolicytable_tomarkovpolicy banditrlproof.finitehorizonrl.transitioncountsummary.optimisticpolicytable_tomarkovpolicy the action-table interpretation is exactly the plan's optimistic policy. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalOptimisticPolicyTable","label":"empiricalOptimisticPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalOptimisticPolicyTable","description":"The empirical optimistic action table computed from one observed batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-2cdb3d0fc356","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7460,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:201"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalOptimisticPolicyTable {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (transitionBonus : Real) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"empiricaloptimisticpolicytable banditrlproof.finitehorizonrl.episodebatch.empiricaloptimisticpolicytable the empirical optimistic action table computed from one observed batch. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalOptimisticPolicyTable","label":"measurable_empiricalOptimisticPolicyTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalOptimisticPolicyTable","description":"The empirical optimistic table is measurable in the raw batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-e9815c277a2c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7461,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_empiricalOptimisticPolicyTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (transitionBonus : Real) : Measurable fun batch : EpisodeBatch mdp episodes => batch.empiricalOptimisticPolicyTable defaultState transitionBonus","missing":[],"search":"measurable_empiricaloptimisticpolicytable banditrlproof.finitehorizonrl.episodebatch.measurable_empiricaloptimisticpolicytable the empirical optimistic table is measurable in the raw batch. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchKernel","label":"iidEpisodeBatchKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchKernel","description":"Iid generated episode-batch law indexed by a deterministic policy table.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-a125b7c73632","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7462,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:226"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidEpisodeBatchKernel {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.Kernel (DeterministicMarkovPolicyTable mdp) (EpisodeBatch mdp episodes)","missing":[],"search":"iidepisodebatchkernel banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.iidepisodebatchkernel iid generated episode-batch law indexed by a deterministic policy table. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchKernel_apply","label":"iidEpisodeBatchKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchKernel_apply","description":"theorem iidEpisodeBatchKernel_apply {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (table : DeterministicMarkovPolicyTable mdp) : iidEpisodeBatchKernel initialState episodes table = table.toMarkovPolicy.iidEpisodeBatchMeasure initialState episodes","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-a5d6e0540c81","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7463,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:237"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchKernel_apply {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (table : DeterministicMarkovPolicyTable mdp) : iidEpisodeBatchKernel initialState episodes table = table.toMarkovPolicy.iidEpisodeBatchMeasure initialState episodes","missing":[],"search":"iidepisodebatchkernel_apply banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.iidepisodebatchkernel_apply theorem iidepisodebatchkernel_apply {mdp : mdp state action} (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) (table : deterministicmarkovpolicytable mdp) : iidepisodebatchkernel initialstate episodes table = table.tomarkovpolicy.iidepisodebatchmeasure initialstate episodes theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.latestBatch","label":"latestBatch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.latestBatch","description":"The latest observed batch in a finite nonempty prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-c9faacd271ff","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7464,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def latestBatch {mdp : MDP State Action} {episodes n : Nat} (history : EpisodeBatchPrefix mdp episodes n) : EpisodeBatch mdp episodes","missing":[],"search":"latestbatch banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.latestbatch the latest observed batch in a finite nonempty prefix. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurable_latestBatch","label":"measurable_latestBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurable_latestBatch","description":"theorem measurable_latestBatch {mdp : MDP State Action} {episodes n : Nat} : Measurable (latestBatch : EpisodeBatchPrefix mdp episodes n -> EpisodeBatch mdp episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-179b9fb5706b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7465,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:267"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_latestBatch {mdp : MDP State Action} {episodes n : Nat} : Measurable (latestBatch : EpisodeBatchPrefix mdp episodes n -> EpisodeBatch mdp episodes)","missing":[],"search":"measurable_latestbatch banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.measurable_latestbatch theorem measurable_latestbatch {mdp : mdp state action} {episodes n : nat} : measurable (latestbatch : episodebatchprefix mdp episodes n -> episodebatch mdp episodes) theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.successorTable","label":"successorTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.successorTable","description":"Optimistic table selected from the latest batch in a finite prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-443988aac0d2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7466,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:274"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (transitionBonus : Real) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"successortable banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.successortable optimistic table selected from the latest batch in a finite prefix. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurable_successorTable","label":"measurable_successorTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurable_successorTable","description":"theorem measurable_successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (transitionBonus : Real) (n : Nat) : Measurable (successorTable (mdp := mdp) (episodes := episodes) defaultState transitionBonus n)","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-41f76b2ad2fc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7467,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:283"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (transitionBonus : Real) (n : Nat) : Measurable (successorTable (mdp := mdp) (episodes := episodes) defaultState transitionBonus n)","missing":[],"search":"measurable_successortable banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.measurable_successortable theorem measurable_successortable {mdp : mdp state action} {episodes : nat} (defaultstate : state) (transitionbonus : real) (n : nat) : measurable (successortable (mdp := mdp) (episodes := episodes) defaultstate transitionbonus n) theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source","label":"source","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source","description":"Concrete adaptive source: after every batch, recompute the known-reward empirical-transition optimistic table from that batch and sample the next iid batch under the selected deterministic policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-99e30cb0407d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7468,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:297"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def source (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) : AdaptiveEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"source banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.source concrete adaptive source: after every batch, recompute the known-reward empirical-transition optimistic table from that batch and sample the next iid batch under the selected deterministic policy. definition compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_successorPolicy_eq_optimisticPolicy","label":"source_successorPolicy_eq_optimisticPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_successorPolicy_eq_optimisticPolicy","description":"The selected policy is exactly the latest batch's optimistic plan policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-0012005b92a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7469,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_successorPolicy_eq_optimisticPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : (source mdp initialState episodes initialTable defaultState transitionBonus).successorPolicy n history = ((latestBatch history).transitionCountSummary.optimisticPlan mdp defaultState transitionBonus).optimisticPolicy","missing":[],"search":"source_successorpolicy_eq_optimisticpolicy banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.source_successorpolicy_eq_optimisticpolicy the selected policy is exactly the latest batch's optimistic plan policy. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurableSet_selectedSimultaneousCountBadEvent","label":"measurableSet_selectedSimultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurableSet_selectedSimultaneousCountBadEvent","description":"Measurability of a selected count event for any measurable finite policy-table selector. This discharges the regularity premise of the adaptive union route.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-f38a5733a72b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7470,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:338"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selectedSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} {History : Type*} [MeasurableSpace History] (selector : History -> DeterministicMarkovPolicyTable mdp) (hselector : Measurable selector) (delta : Real) : MeasurableSet {pair : History × EpisodeBatch mdp episodes | pair.2 ∈ (selector pair.1).toMarkovPolicy.simultaneousCountBadEvent initialState episodes delta}","missing":[],"search":"measurableset_selectedsimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.measurableset_selectedsimultaneouscountbadevent measurability of a selected count event for any measurable finite policy-table selector. this discharges the regularity premise of the adaptive union route. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_measurableSet_successorSimultaneousCountBadEvent","label":"source_measurableSet_successorSimultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_measurableSet_successorSimultaneousCountBadEvent","description":"Every selected successor count event of the concrete source is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-21cabe9c9ead","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7471,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:365"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_measurableSet_successorSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (rounds : Nat) (delta : Real) (n : Nat) : MeasurableSet (AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent (source mdp initialState episodes initialTable defaultState transitionBonus) rounds delta n)","missing":[],"search":"source_measurableset_successorsimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.source_measurableset_successorsimultaneouscountbadevent every selected successor count event of the concrete source is measurable. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_trajectoryMeasure_condDistrib_eq_empiricalOptimisticPolicyBatchLaw","label":"source_trajectoryMeasure_condDistrib_eq_empiricalOptimisticPolicyBatchLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_trajectoryMeasure_condDistrib_eq_empiricalOptimisticPolicyBatchLaw","description":"The concrete source's next-batch conditional law is its latest empirical optimistic policy law.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-7f75bdc62dbd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7472,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:387"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_trajectoryMeasure_condDistrib_eq_empiricalOptimisticPolicyBatchLaw {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (n : Nat) : let concreteSource := source mdp initialState episodes initialTable defaultState transitionBonus Filter.EventuallyEq (ae (concreteSource.trajectoryMeasure.map (Preorder.frestrictLe n))) (ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => trajectory (n + 1)) (Preorder.frestrictLe n) concreteSource.trajectoryMeasure) (fun history => MarkovPolicy.iidEpisodeBatchMeasure ((latestBatch history).transitionCountSummary.optimisticPlan mdp defaultState transitionBonus).optimisticPolicy initialState episodes)","missing":[],"search":"source_trajectorymeasure_conddistrib_eq_empiricaloptimisticpolicybatchlaw banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.source_trajectorymeasure_conddistrib_eq_empiricaloptimisticpolicybatchlaw the concrete source's next-batch conditional law is its latest empirical optimistic policy law. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_trajectoryMeasure_adaptiveSimultaneousCountConfidence","label":"source_trajectoryMeasure_adaptiveSimultaneousCountConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_trajectoryMeasure_adaptiveSimultaneousCountConfidence","description":"Concrete adaptive count-confidence terminal for the empirical optimistic source. No selected-law or successor-event measurability premise remains.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveempiricaloptimisticsource/index.html#decl-6e8f00872814","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","order":7473,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEmpiricalOptimisticSource.lean:419"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem source_trajectoryMeasure_adaptiveSimultaneousCountConfidence {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let concreteSource := source mdp initialState episodes initialTable defaultState transitionBonus MeasurableSet (concreteSource.adaptiveSimultaneousCountBadEvent rounds delta) ∧ concreteSource.trajectoryMeasure (concreteSource.adaptiveSimultaneousCountBadEvent rounds delta) <= ENNReal.ofReal delta ∧ forall trajectory, trajectory ∉ concreteSource.adaptiveSimultaneousCountBadEvent rounds delta -> forall round : Fin rounds, forall coordinate : CountCoordinate mdp, |coordinate.deviation (c…","missing":[],"search":"source_trajectorymeasure_adaptivesimultaneouscountconfidence banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.source_trajectorymeasure_adaptivesimultaneouscountconfidence concrete adaptive count-confidence terminal for the empirical optimistic source. no selected-law or successor-event measurability premise remains. theorem compiled","shard":"modules/8b1ce42ec8462187.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix","label":"EpisodeBatchPrefix","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix","description":"Finite history through batch coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-6f3d269a0013","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7474,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev EpisodeBatchPrefix (mdp : MDP State Action) (episodes n : Nat)","missing":[],"search":"episodebatchprefix banditrlproof.finitehorizonrl.episodebatchprefix finite history through batch coordinate `n`. abbreviation compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchTrajectory","label":"EpisodeBatchTrajectory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatchTrajectory","description":"Infinite batch trajectory used by the adaptive Ionescu--Tulcea law.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-518ccef2face","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7475,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev EpisodeBatchTrajectory (mdp : MDP State Action) (episodes : Nat)","missing":[],"search":"episodebatchtrajectory banditrlproof.finitehorizonrl.episodebatchtrajectory infinite batch trajectory used by the adaptive ionescu--tulcea law. abbreviation compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource","label":"AdaptiveEpisodeBatchSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource","description":"An adaptive batch source. Coordinate zero is generated by `initialPolicy`. After observing the prefix through coordinate `n`, `successorPolicy n history` selects the policy for coordinate `n + 1`. `batchKernel_eq_iidEpisodeBatchMeasure` is the exact law transport contract: it also certifies that the history-indexed family is a measurable Markov kernel rather than merely a pointwise family of measures.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-c532386359b1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7476,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure AdaptiveEpisodeBatchSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) where","missing":[],"search":"adaptiveepisodebatchsource banditrlproof.finitehorizonrl.adaptiveepisodebatchsource an adaptive batch source. coordinate zero is generated by `initialpolicy`. after observing the prefix through coordinate `n`, `successorpolicy n history` selects the policy for coordinate `n + 1`. `batchkernel_eq_iidepisodebatchmeasure` is the exact law transport contract: it also certifies that the history-indexed family is a measurable markov kernel rather than merely a pointwise family of measures. structure compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure","label":"trajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure","description":"The adaptive infinite batch-trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-ff32c57c6a95","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7477,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def trajectoryMeasure {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) : Measure (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"trajectorymeasure banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure the adaptive infinite batch-trajectory law. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_map_eval_zero","label":"trajectoryMeasure_map_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_map_eval_zero","description":"Coordinate zero has the batch law of the configured initial policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-6ab2749bd702","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7478,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_map_eval_zero {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) : source.trajectoryMeasure.map (Function.eval 0) = source.initialPolicy.iidEpisodeBatchMeasure initialState episodes","missing":[],"search":"trajectorymeasure_map_eval_zero banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_map_eval_zero coordinate zero has the batch law of the configured initial policy. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_prefix_compProd","label":"trajectoryMeasure_prefix_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_prefix_compProd","description":"Every successor prefix/next-batch marginal is the compProd of the prefix law and the configured adaptive batch kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-bd1ad80c5ada","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7479,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:131"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_prefix_compProd {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : source.trajectoryMeasure.map (Preorder.frestrictLe n) ⊗ₘ source.batchKernel n = source.trajectoryMeasure.map (fun trajectory => (Preorder.frestrictLe n trajectory, trajectory (n + 1)))","missing":[],"search":"trajectorymeasure_prefix_compprod banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_prefix_compprod every successor prefix/next-batch marginal is the compprod of the prefix law and the configured adaptive batch kernel. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib","label":"trajectoryMeasure_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib","description":"The next batch conditioned on the full finite prefix has the configured history-dependent Markov kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-8155602b1c26","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7480,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:153"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => trajectory (n + 1)) (Preorder.frestrictLe n) source.trajectoryMeasure =ᶠ[ ae (source.trajectoryMeasure.map (Preorder.frestrictLe n))] source.batchKernel n","missing":[],"search":"trajectorymeasure_conddistrib banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conddistrib the next batch conditioned on the full finite prefix has the configured history-dependent markov kernel. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_eq_iidEpisodeBatchMeasure","label":"trajectoryMeasure_condDistrib_eq_iidEpisodeBatchMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_eq_iidEpisodeBatchMeasure","description":"The next batch conditioned on the finite prefix is exactly the generated iid batch law of the policy selected from that prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-e70de1b21132","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7481,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:176"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_eq_iidEpisodeBatchMeasure {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : EpisodeBatchTrajectory mdp episodes => trajectory (n + 1)) (Preorder.frestrictLe n) source.trajectoryMeasure =ᶠ[ ae (source.trajectoryMeasure.map (Preorder.frestrictLe n))] fun history => (source.successorPolicy n history).iidEpisodeBatchMeasure initialState episodes","missing":[],"search":"trajectorymeasure_conddistrib_eq_iidepisodebatchmeasure banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conddistrib_eq_iidepisodebatchmeasure the next batch conditioned on the finite prefix is exactly the generated iid batch law of the policy selected from that prefix. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialBadEvent","label":"initialBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialBadEvent","description":"Pull an initial-coordinate event back to the adaptive trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-2f7ab438e3b4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7482,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:202"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def initialBadEvent {mdp : MDP State Action} {episodes : Nat} (bad : Set (EpisodeBatch mdp episodes)) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"initialbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.initialbadevent pull an initial-coordinate event back to the adaptive trajectory. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorBadEvent","label":"successorBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorBadEvent","description":"Pull a prefix-dependent successor event back to the adaptive trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-8fef647ed939","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7483,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def successorBadEvent {mdp : MDP State Action} {episodes : Nat} (n : Nat) (bad : Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"successorbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorbadevent pull a prefix-dependent successor event back to the adaptive trajectory. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.roundBadEvent","label":"roundBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.roundBadEvent","description":"Round-indexed adapted bad event, with a separate coordinate-zero event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-a957fb301fa6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7484,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:218"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def roundBadEvent {mdp : MDP State Action} {episodes : Nat} (initialBad : Set (EpisodeBatch mdp episodes)) (successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)) : Nat -> Set (EpisodeBatchTrajectory mdp episodes) | 0 => initialBadEvent initialBad | n + 1 => successorBadEvent n (successorBad n) /-- Union of the first `rounds` adapted bad events. -/ def finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat} (rounds : Nat) (initialBad : Set (EpisodeBatch mdp episodes)) (successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"roundbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.roundbadevent round-indexed adapted bad event, with a separate coordinate-zero event. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.finiteHorizonBadEvent","label":"finiteHorizonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.finiteHorizonBadEvent","description":"Union of the first `rounds` adapted bad events.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-c37c1baed441","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7485,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:228"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat} (rounds : Nat) (initialBad : Set (EpisodeBatch mdp episodes)) (successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"finitehorizonbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.finitehorizonbadevent union of the first `rounds` adapted bad events. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_roundBadEvent","label":"measurableSet_roundBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_roundBadEvent","description":"Measurability of every pulled-back adapted round event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-dfa314e46a29","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7486,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:241"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_roundBadEvent {mdp : MDP State Action} {episodes : Nat} {initialBad : Set (EpisodeBatch mdp episodes)} {successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)} (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, MeasurableSet (successorBad n)) (round : Nat) : MeasurableSet (roundBadEvent initialBad successorBad round)","missing":[],"search":"measurableset_roundbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurableset_roundbadevent measurability of every pulled-back adapted round event. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","label":"measurableSet_finiteHorizonBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","description":"Measurability of the finite adapted bad-event union.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-3e72e3eddba7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7487,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_finiteHorizonBadEvent {mdp : MDP State Action} {episodes rounds : Nat} {initialBad : Set (EpisodeBatch mdp episodes)} {successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)} (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (successorBad n)) : MeasurableSet (finiteHorizonBadEvent rounds initialBad successorBad)","missing":[],"search":"measurableset_finitehorizonbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurableset_finitehorizonbadevent measurability of the finite adapted bad-event union. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_initialBadEvent","label":"trajectoryMeasure_initialBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_initialBadEvent","description":"Exact mass of a pulled-back coordinate-zero event.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-cec7300c244a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7488,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_initialBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) {bad : Set (EpisodeBatch mdp episodes)} (hbad : MeasurableSet bad) : source.trajectoryMeasure (initialBadEvent bad) = source.initialPolicy.iidEpisodeBatchMeasure initialState episodes bad","missing":[],"search":"trajectorymeasure_initialbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_initialbadevent exact mass of a pulled-back coordinate-zero event. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","label":"trajectoryMeasure_successorBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","description":"An adapted successor event inherits a uniform bound on every history fiber.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-c887fa477424","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7489,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:298"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successorBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (n : Nat) {bad : Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)} (hbad : MeasurableSet bad) (budget : ENNReal) (hfiber : forall history, source.batchKernel n history (Prod.mk history ⁻¹' bad) <= budget) : source.trajectoryMeasure (successorBadEvent n bad) <= budget","missing":[],"search":"trajectorymeasure_successorbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_successorbadevent_le an adapted successor event inherits a uniform bound on every history fiber. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_roundBadEvent_le","label":"trajectoryMeasure_roundBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_roundBadEvent_le","description":"Every adapted round event inherits the supplied local budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-6d51c590161a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7490,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:336"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_roundBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) {initialBad : Set (EpisodeBatch mdp episodes)} {successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)} (rounds : Nat) (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (successorBad n)) (budget : ENNReal) (hinitial_le : source.initialPolicy.iidEpisodeBatchMeasure initialState episodes initialBad <= budget) (hsuccessor_le : forall n, n + 1 < rounds -> forall history, source.batchKernel n history (Prod.mk history ⁻¹' successorBad n) <= budget) (round : Fin rounds) : source.trajectoryMeasure (roundBadEvent initialBad successorBad round) <= budget","missing":[],"search":"trajectorymeasure_roundbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_roundbadevent_le every adapted round event inherits the supplied local budget. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_finiteHorizonBadEvent_le","label":"trajectoryMeasure_finiteHorizonBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_finiteHorizonBadEvent_le","description":"Finite-horizon adaptive union bound with equal confidence shares `delta / rounds`. Independence between batches is not assumed.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-0139f6d2546b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7491,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:374"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_finiteHorizonBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (delta : Real) {initialBad : Set (EpisodeBatch mdp episodes)} {successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)} (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (successorBad n)) (hinitial_le : source.initialPolicy.iidEpisodeBatchMeasure initialState episodes initialBad <= ENNReal.ofReal (multiBatchLocalDelta rounds delta)) (hsuccessor_le : forall n, n + 1 < rounds -> forall history, source.batchKernel n history (Prod.mk history ⁻¹' successorBad n) <= ENNReal.ofReal (multiBatchLocalDelta rounds delta)) : source.trajectoryMeasure (finiteHo…","missing":[],"search":"trajectorymeasure_finitehorizonbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_finitehorizonbadevent_le finite-horizon adaptive union bound with equal confidence shares `delta / rounds`. independence between batches is not assumed. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_conditionalLaw_and_finiteHorizonBadEvent_le","label":"trajectoryMeasure_conditionalLaw_and_finiteHorizonBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_conditionalLaw_and_finiteHorizonBadEvent_le","description":"Terminal adaptive law-and-budget package: every successor conditional law is the generated iid batch law of the history-selected policy, and arbitrary measurable adapted local events obey one global finite-horizon delta budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-13efd84d70dd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7492,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:415"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_conditionalLaw_and_finiteHorizonBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (delta : Real) {initialBad : Set (EpisodeBatch mdp episodes)} {successorBad : (n : Nat) -> Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)} (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (successorBad n)) (hinitial_le : source.initialPolicy.iidEpisodeBatchMeasure initialState episodes initialBad <= ENNReal.ofReal (multiBatchLocalDelta rounds delta)) (hsuccessor_le : forall n, n + 1 < rounds -> forall history, source.batchKernel n history (Prod.mk history ⁻¹' successorBad n) <= ENNReal.ofReal (mult…","missing":[],"search":"trajectorymeasure_conditionallaw_and_finitehorizonbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_conditionallaw_and_finitehorizonbadevent_le terminal adaptive law-and-budget package: every successor conditional law is the generated iid batch law of the history-selected policy, and arbitrary measurable adapted local events obey one global finite-horizon delta budget. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialSimultaneousCountBadEvent","label":"initialSimultaneousCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialSimultaneousCountBadEvent","description":"Count-confidence event for the initial policy's coordinate-zero batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-12f26dadc60a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7493,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:455"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def initialSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : Set (EpisodeBatch mdp episodes)","missing":[],"search":"initialsimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.initialsimultaneouscountbadevent count-confidence event for the initial policy's coordinate-zero batch. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent","label":"successorSimultaneousCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent","description":"Prefix/next-batch count event for the policy selected from that prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-a86b4ba545ca","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7494,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:466"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def successorSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) (n : Nat) : Set (EpisodeBatchPrefix mdp episodes n × EpisodeBatch mdp episodes)","missing":[],"search":"successorsimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorsimultaneouscountbadevent prefix/next-batch count event for the policy selected from that prefix. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveSimultaneousCountBadEvent","label":"adaptiveSimultaneousCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveSimultaneousCountBadEvent","description":"The finite-horizon union of initial and history-selected simultaneous count events on the adaptive batch trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-11a807987ce6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7495,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:480"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def adaptiveSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) : Set (EpisodeBatchTrajectory mdp episodes)","missing":[],"search":"adaptivesimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.adaptivesimultaneouscountbadevent the finite-horizon union of initial and history-selected simultaneous count events on the adaptive batch trajectory. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.policyAt","label":"policyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.policyAt","description":"Policy used at a batch coordinate of one adaptive trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-56211f94706e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7496,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:491"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def policyAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (trajectory : EpisodeBatchTrajectory mdp episodes) : Nat -> MarkovPolicy mdp | 0 => source.initialPolicy | n + 1 => source.successorPolicy n (Preorder.frestrictLe n trajectory) omit [Nonempty State] [Nonempty Action] in /-- The initial count event receives the common local confidence share. -/ theorem initialSimultaneousCountBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.initialPolicy.iidEpisodeBatchMeasure initialState episodes (so…","missing":[],"search":"policyat banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.policyat policy used at a batch coordinate of one adaptive trajectory. definition compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialSimultaneousCountBadEvent_le","label":"initialSimultaneousCountBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialSimultaneousCountBadEvent_le","description":"The initial count event receives the common local confidence share.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-5eb18ba99eb4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7497,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:503"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem initialSimultaneousCountBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.initialPolicy.iidEpisodeBatchMeasure initialState episodes (source.initialSimultaneousCountBadEvent rounds delta) <= ENNReal.ofReal (multiBatchLocalDelta rounds delta)","missing":[],"search":"initialsimultaneouscountbadevent_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.initialsimultaneouscountbadevent_le the initial count event receives the common local confidence share. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent_fiber_le","label":"successorSimultaneousCountBadEvent_fiber_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent_fiber_le","description":"Every history fiber of the selected-policy successor count event receives the same local confidence share.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-e76b4e8b2d75","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7498,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:530"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorSimultaneousCountBadEvent_fiber_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (n : Nat) (history : EpisodeBatchPrefix mdp episodes n) : source.batchKernel n history (Prod.mk history ⁻¹' source.successorSimultaneousCountBadEvent rounds delta n) <= ENNReal.ofReal (multiBatchLocalDelta rounds delta)","missing":[],"search":"successorsimultaneouscountbadevent_fiber_le banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.successorsimultaneouscountbadevent_fiber_le every history fiber of the selected-policy successor count event receives the same local confidence share. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_adaptiveSimultaneousCountBadEvent","label":"measurableSet_adaptiveSimultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_adaptiveSimultaneousCountBadEvent","description":"Measurability of the adaptive simultaneous-count union.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-966f2de26709","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7499,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:563"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_adaptiveSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta : Real) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (source.successorSimultaneousCountBadEvent rounds delta n)) : MeasurableSet (source.adaptiveSimultaneousCountBadEvent rounds delta)","missing":[],"search":"measurableset_adaptivesimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.measurableset_adaptivesimultaneouscountbadevent measurability of the adaptive simultaneous-count union. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveSimultaneousCountConfidence","label":"trajectoryMeasure_adaptiveSimultaneousCountConfidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveSimultaneousCountConfidence","description":"Adaptive global-delta simultaneous count confidence. Outside one measurable finite-horizon event, every realized batch satisfies the strict count deviation bound relative to the policy selected from its preceding history.","url":"../modules/banditrlproof-rl-finitehorizonadaptiveepisodebatchlaw/index.html#decl-64296109a223","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","order":7500,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveEpisodeBatchLaw.lean:584"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_adaptiveSimultaneousCountConfidence {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (source.successorSimultaneousCountBadEvent rounds delta n)) : MeasurableSet (source.adaptiveSimultaneousCountBadEvent rounds delta) ∧ source.trajectoryMeasure (source.adaptiveSimultaneousCountBadEvent rounds delta) <= ENNReal.ofReal delta ∧ forall trajectory, trajectory ∉ source.adaptiveSimultaneousCountBadEvent rounds delta -> forall round : Fin rounds, forall coordinate : CountCoordinate mdp, |coordinate.deviation (source.policyAt trajectory round) initialState (trajector…","missing":[],"search":"trajectorymeasure_adaptivesimultaneouscountconfidence banditrlproof.finitehorizonrl.adaptiveepisodebatchsource.trajectorymeasure_adaptivesimultaneouscountconfidence adaptive global-delta simultaneous count confidence. outside one measurable finite-horizon event, every realized batch satisfies the strict count deviation bound relative to the policy selected from its preceding history. theorem compiled","shard":"modules/2913f0f8ee26ce69.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","label":"measurable_realizedSuccessorCumulativeRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","description":"Stochastic realized cumulative successor regret is trajectory-measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-811b7cb82a00","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7501,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.realizedSuccessorCumulativeRegret trajectory rounds)","missing":[],"search":"measurable_realizedsuccessorcumulativeregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_realizedsuccessorcumulativeregret stochastic realized cumulative successor regret is trajectory-measurable. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","label":"measurable_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","description":"Stochastic realized average successor regret is trajectory-measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-c056be380d41","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7502,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.realizedSuccessorAverageRegret trajectory rounds)","missing":[],"search":"measurable_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_realizedsuccessoraverageregret stochastic realized average successor regret is trajectory-measurable. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedCumulativeRegret_nonneg","label":"successorExpectedCumulativeRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedCumulativeRegret_nonneg","description":"A finite sum of selected stochastic-policy expected regrets is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-0fb0f23716e6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7503,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorExpectedCumulativeRegret_nonneg {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : 0 <= source.successorExpectedCumulativeRegret trajectory rounds","missing":[],"search":"successorexpectedcumulativeregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorexpectedcumulativeregret_nonneg a finite sum of selected stochastic-policy expected regrets is nonnegative. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret_nonneg","label":"successorExpectedAverageRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret_nonneg","description":"The average selected stochastic-policy expected regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-25f9f875110d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7504,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorExpectedAverageRegret_nonneg {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : 0 <= source.successorExpectedAverageRegret trajectory rounds","missing":[],"search":"successorexpectedaverageregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorexpectedaverageregret_nonneg the average selected stochastic-policy expected regret is nonnegative. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","label":"abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","description":"An expected-regret upper bound and a two-sided global return-deviation bound control the absolute stochastic realized average regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-96c02c7b0c70","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7505,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (expectedBound deviationBound : Real) (hexpected : source.successorExpectedAverageRegret trajectory rounds <= expectedBound) (hdeviation : |source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory| <= deviationBound) : |source.realizedSuccessorAverageRegret trajectory rounds| <= expectedBound + deviationBound / ((episodes : Real) * (rounds : Real))","missing":[],"search":"abs_realizedsuccessoraverageregret_le_of_expected_le_of_deviation_abs_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.abs_realizedsuccessoraverageregret_le_of_expected_le_of_deviation_abs_le an expected-regret upper bound and a two-sided global return-deviation bound control the absolute stochastic realized average regret. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageAbsoluteRealizedBehaviorConsistency_of_standardBorel","label":"exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageAbsoluteRealizedBehaviorConsistency_of_standardBorel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageAbsoluteRealizedBehaviorConsistency_of_standardBorel","description":"The stochastic finite-window certificate upgraded to absolute realized regret. This is the exact finite-window input consumed by convergence in probability.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-7590f9421037","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7506,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:141"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageAbsoluteRealizedBehaviorConsistency_of_standardBorel (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpisodeBatchSource.decayingExpl…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageabsoluterealizedbehaviorconsistency_of_standardborel banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageabsoluterealizedbehaviorconsistency_of_standardborel the stochastic finite-window certificate upgraded to absolute realized regret. this is the exact finite-window input consumed by convergence in probability. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.DecayingExplorationStochasticWindowSpace","label":"DecayingExplorationStochasticWindowSpace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.DecayingExplorationStochasticWindowSpace","description":"One complete scheduled stochastic experiment at every product coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-23f55e72aebd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7507,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:276"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev DecayingExplorationStochasticWindowSpace (mdp : MDP State Action) (baseVisitFloor : Real)","missing":[],"search":"decayingexplorationstochasticwindowspace banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticwindowspace one complete scheduled stochastic experiment at every product coordinate. abbreviation compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticWindowSource","label":"decayingExplorationStochasticWindowSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticWindowSource","description":"The stochastic adaptive source used at schedule coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-5587cb137751","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7508,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:283"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticWindowSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : AdaptiveStochasticEpisodeBatchSource mdp initialState (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n)","missing":[],"search":"decayingexplorationstochasticwindowsource banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticwindowsource the stochastic adaptive source used at schedule coordinate `n`. definition compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticWindowMeasure","label":"decayingExplorationStochasticWindowMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticWindowMeasure","description":"The scheduled stochastic trajectory law at product coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-3b8942c36797","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7509,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:305"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticWindowMeasure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Measure (StochasticEpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","missing":[],"search":"decayingexplorationstochasticwindowmeasure banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticwindowmeasure the scheduled stochastic trajectory law at product coordinate `n`. definition compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure","label":"decayingExplorationStochasticCommonMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure","description":"Independent product coupling of the complete scheduled stochastic laws.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-ce45d0847764","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7510,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:330"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticCommonMeasure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) : Measure (DecayingExplorationStochasticWindowSpace mdp baseVisitFloor)","missing":[],"search":"decayingexplorationstochasticcommonmeasure banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcommonmeasure independent product coupling of the complete scheduled stochastic laws. definition compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_map_eval","label":"decayingExplorationStochasticCommonMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_map_eval","description":"Each common-space coordinate has exactly its scheduled stochastic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-672bf49c0d64","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7511,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:354"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticCommonMeasure_map_eval (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : (decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor).map (fun omega => omega n) = decayingExplorationStochasticWindowMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor n","missing":[],"search":"decayingexplorationstochasticcommonmeasure_map_eval banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcommonmeasure_map_eval each common-space coordinate has exactly its scheduled stochastic law. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretProcess","label":"decayingExplorationStochasticRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretProcess","description":"Scheduled stochastic realized successor-average regret on the common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-cd18cb94ae28","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7512,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:368"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) (omega : DecayingExplorationStochasticWindowSpace mdp baseVisitFloor) : Real","missing":[],"search":"decayingexplorationstochasticrealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticrealizedbehaviorregretprocess scheduled stochastic realized successor-average regret on the common space. definition compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.measurable_decayingExplorationStochasticRealizedBehaviorRegretProcess","label":"measurable_decayingExplorationStochasticRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.measurable_decayingExplorationStochasticRealizedBehaviorRegretProcess","description":"Every scheduled stochastic regret coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-7f11439919ea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7513,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:381"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_decayingExplorationStochasticRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Measurable (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState baseVisitFloor n)","missing":[],"search":"measurable_decayingexplorationstochasticrealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.measurable_decayingexplorationstochasticrealizedbehaviorregretprocess every scheduled stochastic regret coordinate is measurable. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonBadEvent","label":"decayingExplorationStochasticCommonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonBadEvent","description":"Pull the finite projected-count/global-return union to the common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-c28709e6a4c0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7514,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:399"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticCommonBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (rewardVarianceProxy : NNReal) (n : Nat) : Set (DecayingExplorationStochasticWindowSpace mdp baseVisitFloor)","missing":[],"search":"decayingexplorationstochasticcommonbadevent banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcommonbadevent pull the finite projected-count/global-return union to the common space. definition compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_badEvent_le","label":"decayingExplorationStochasticCommonMeasure_badEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_badEvent_le","description":"The pulled-back stochastic bad event inherits the finite-window budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-6ddd307e3934","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7515,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:429"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticCommonMeasure_badEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor (decayingExplorationStochasticCommonBadEvent mdp initi…","missing":[],"search":"decayingexplorationstochasticcommonmeasure_badevent_le banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcommonmeasure_badevent_le the pulled-back stochastic bad event inherits the finite-window budget. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.abs_decayingExplorationStochasticRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","label":"abs_decayingExplorationStochasticRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.abs_decayingExplorationStochasticRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","description":"Outside the pulled-back bad event, coordinate `n` has the absolute bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-f49a23ffe786","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7516,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:490"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_decayingExplorationStochasticRealizedBehaviorRegretProcess_le_of_not_mem_badEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) (omega : DecayingExplorationStochasticWindowSpace mdp baseVisitFloor) (homega : omega ∉ decayingExplorationStochasticCommonBadEvent mdp ini…","missing":[],"search":"abs_decayingexplorationstochasticrealizedbehaviorregretprocess_le_of_not_mem_badevent banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.abs_decayingexplorationstochasticrealizedbehaviorregretprocess_le_of_not_mem_badevent outside the pulled-back bad event, coordinate `n` has the absolute bound. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","label":"exploratorySource_decayingExplorationStochasticCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","description":"Terminal theorem: exact stochastic schedule marginals and convergence in probability of realized successor-average behavior regret to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-cee96569336f/index.html#decl-667505a24e61","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","order":7517,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency.lean:527"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_decayingExplorationStochasticCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall n, Measurable (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource i…","missing":[],"search":"exploratorysource_decayingexplorationstochasticcommonmeasure_marginals_and_realizedbehaviorregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_decayingexplorationstochasticcommonmeasure_marginals_and_realizedbehaviorregret_tendstoinmeasure_zero terminal theorem: exact stochastic schedule marginals and convergence in probability of realized successor-average behavior regret to zero. theorem compiled","shard":"modules/306fe948ddbfcdcc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_le_two_mul_horizon_of_rewardBound","label":"expectedRegret_le_two_mul_horizon_of_rewardBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_le_two_mul_horizon_of_rewardBound","description":"Mean-reward expected regret has the deterministic `2H` envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-f6c81c98bc38","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7518,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_le_two_mul_horizon_of_rewardBound {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (hrewardBound : forall state action, |mdp.reward state action| <= 1) : policy.expectedRegret initialState <= 2 * (mdp.horizon : Real)","missing":[],"search":"expectedregret_le_two_mul_horizon_of_rewardbound banditrlproof.finitehorizonrl.markovpolicy.expectedregret_le_two_mul_horizon_of_rewardbound mean-reward expected regret has the deterministic `2h` envelope. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret_le_two_mul_horizon","label":"successorExpectedAverageRegret_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret_le_two_mul_horizon","description":"Every selected-policy expected successor average is at most `2H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-60a7c19a42b5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7519,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorExpectedAverageRegret_le_two_mul_horizon {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : source.successorExpectedAverageRegret trajectory rounds <= 2 * (mdp.horizon : Real)","missing":[],"search":"successorexpectedaverageregret_le_two_mul_horizon banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorexpectedaverageregret_le_two_mul_horizon every selected-policy expected successor average is at most `2h`. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","label":"trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","description":"The cumulative globally centered successor deviation inherits one global sub-Gaussian MGF from the strongly adapted conditional increments.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-039b7eea3fe7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7520,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasSubgaussianMGF (source.cumulativeSuccessorGlobalReturnDeviation rounds) (cumulativeSuccessorGlobalReturnVarianceProxy m…","missing":[],"search":"trajectorymeasure_cumulativesuccessorglobalreturndeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_cumulativesuccessorglobalreturndeviation_hassubgaussianmgf the cumulative globally centered successor deviation inherits one global sub-gaussian mgf from the strongly adapted conditional increments. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","label":"integrable_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","description":"Finite-window stochastic realized successor-average regret is integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-c6451112d1cc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7521,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (hrewardBoundOne : forall state action, |mdp.reward state action| <= 1) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) : Integrable (fun trajectory => source.realizedSuccesso…","missing":[],"search":"integrable_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.integrable_realizedsuccessoraverageregret finite-window stochastic realized successor-average regret is integrable. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound","label":"normalizedSuccessorGlobalReturnMGFFirstMomentBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound","description":"The scaled-MGF first-moment contribution after episode/round normalization.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-50c1e830ae0b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7522,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:229"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedSuccessorGlobalReturnMGFFirstMomentBound (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : Real","missing":[],"search":"normalizedsuccessorglobalreturnmgffirstmomentbound banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.normalizedsuccessorglobalreturnmgffirstmomentbound the scaled-mgf first-moment contribution after episode/round normalization. definition compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","label":"normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","description":"The scaled-MGF contribution is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-596049d5d158","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7523,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:243"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : 0 <= normalizedSuccessorGlobalReturnMGFFirstMomentBound mdp episodes rounds rewardBound rewardVarianceProxy","missing":[],"search":"normalizedsuccessorglobalreturnmgffirstmomentbound_nonneg banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.normalizedsuccessorglobalreturnmgffirstmomentbound_nonneg the scaled-mgf contribution is nonnegative. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","label":"normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","description":"Whenever `log (2/delta) >= 1/2`, the scaled-MGF first-moment term is at most `2 * exp(1/2)` times the compiled normalized confidence radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-cb55c9e6febb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7524,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:258"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) (hepisodes : 0 < episodes) (hrounds : 0 < rounds) (hlog : (1 / 2 : Real) <= Real.log (2 / delta)) : normalizedSuccessorGlobalReturnMGFFirstMomentBound mdp episodes rounds rewardBound rewardVarianceProxy <= 2 * Real.exp (1 / 2 : Real) * AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy delta","missing":[],"search":"normalizedsuccessorglobalreturnmgffirstmomentbound_le_confidenceradius banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.normalizedsuccessorglobalreturnmgffirstmomentbound_le_confidenceradius whenever `log (2/delta) >= 1/2`, the scaled-mgf first-moment term is at most `2 * exp(1/2)` times the compiled normalized confidence radius. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.one_half_le_log_two_div_vanishingAverageConfidenceDelta","label":"one_half_le_log_two_div_vanishingAverageConfidenceDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.one_half_le_log_two_div_vanishingAverageConfidenceDelta","description":"The scheduled confidence logarithm is uniformly at least one half.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-57896b845cf5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7525,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:308"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_half_le_log_two_div_vanishingAverageConfidenceDelta (n : Nat) : (1 / 2 : Real) <= Real.log (2 / AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n)","missing":[],"search":"one_half_le_log_two_div_vanishingaverageconfidencedelta banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.one_half_le_log_two_div_vanishingaverageconfidencedelta the scheduled confidence logarithm is uniformly at least one half. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationNormalizedSuccessorGlobalReturnMGFFirstMomentBound_tendsto_zero","label":"decayingExplorationNormalizedSuccessorGlobalReturnMGFFirstMomentBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationNormalizedSuccessorGlobalReturnMGFFirstMomentBound_tendsto_zero","description":"The scheduled scaled-MGF contribution tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-695b8dd5d266","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7526,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationNormalizedSuccessorGlobalReturnMGFFirstMomentBound_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) (rewardBound rewardVarianceProxy : NNReal) : Tendsto (fun n => normalizedSuccessorGlobalReturnMGFFirstMomentBound mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n) (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) rewardBound rewardVarianceProxy) atTop (nhds 0)","missing":[],"search":"decayingexplorationnormalizedsuccessorglobalreturnmgffirstmomentbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationnormalizedsuccessorglobalreturnmgffirstmomentbound_tendsto_zero the scheduled scaled-mgf contribution tends to zero. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCumulativeReturnDeviationProcess","label":"decayingExplorationStochasticCumulativeReturnDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCumulativeReturnDeviationProcess","description":"Scheduled cumulative sampled-return deviation on the stochastic common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-1f8926211b12","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7527,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:374"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticCumulativeReturnDeviationProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : DecayingExplorationStochasticWindowSpace mdp baseVisitFloor -> Real","missing":[],"search":"decayingexplorationstochasticcumulativereturndeviationprocess banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcumulativereturndeviationprocess scheduled cumulative sampled-return deviation on the stochastic common space. definition compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCumulativeReturnDeviationProcess_hasSubgaussianMGF","label":"decayingExplorationStochasticCumulativeReturnDeviationProcess_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCumulativeReturnDeviationProcess_hasSubgaussianMGF","description":"Every common-space cumulative deviation coordinate has its exact MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-f8f0ef47da95","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7528,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:387"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticCumulativeReturnDeviationProcess_hasSubgaussianMGF (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : ProbabilityTheory.HasSubgaussianMGF (decayingExplorationStochasticCumulativeReturnDeviationProcess mdp initialState rewardSource initialTable defaultState baseVisitFloor n) (AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy mdp (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) (AdaptiveEpisod…","missing":[],"search":"decayingexplorationstochasticcumulativereturndeviationprocess_hassubgaussianmgf banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcumulativereturndeviationprocess_hassubgaussianmgf every common-space cumulative deviation coordinate has its exact mgf. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.integrable_decayingExplorationStochasticRealizedBehaviorRegretProcess","label":"integrable_decayingExplorationStochasticRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.integrable_decayingExplorationStochasticRealizedBehaviorRegretProcess","description":"Every scheduled common-space realized-regret coordinate is integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-287a6304849e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7529,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:439"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_decayingExplorationStochasticRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : Integrable (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState baseVisitFloor n) (decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor)","missing":[],"search":"integrable_decayingexplorationstochasticrealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.integrable_decayingexplorationstochasticrealizedbehaviorregretprocess every scheduled common-space realized-regret coordinate is integrable. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonCountBadEvent","label":"decayingExplorationStochasticCommonCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonCountBadEvent","description":"Pull only the projected-count failure event to the stochastic common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-7c8265b121eb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7530,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:481"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticCommonCountBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Set (DecayingExplorationStochasticWindowSpace mdp baseVisitFloor)","missing":[],"search":"decayingexplorationstochasticcommoncountbadevent banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcommoncountbadevent pull only the projected-count failure event to the stochastic common space. definition compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.measurableSet_decayingExplorationStochasticCommonCountBadEvent","label":"measurableSet_decayingExplorationStochasticCommonCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.measurableSet_decayingExplorationStochasticCommonCountBadEvent","description":"The pulled-back projected-count event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-a66441e3c36b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7531,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:505"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_decayingExplorationStochasticCommonCountBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : MeasurableSet (decayingExplorationStochasticCommonCountBadEvent mdp initialState initialTable defaultState baseVisitFloor n)","missing":[],"search":"measurableset_decayingexplorationstochasticcommoncountbadevent banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.measurableset_decayingexplorationstochasticcommoncountbadevent the pulled-back projected-count event is measurable. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_countBadEvent_le","label":"decayingExplorationStochasticCommonMeasure_countBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_countBadEvent_le","description":"The common-space projected-count event consumes one confidence share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-c47cfa188b39","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7532,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:531"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticCommonMeasure_countBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor (decayingExplorationStochasticCommonCountBadEvent mdp initialState initialTable defaultState baseVisitFloor n) <= ENNReal.ofReal (AdaptiveEpisodeBatc…","missing":[],"search":"decayingexplorationstochasticcommonmeasure_countbadevent_le banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticcommonmeasure_countbadevent_le the common-space projected-count event consumes one confidence share. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret","label":"decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret","description":"Expected absolute scheduled stochastic realized-behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-ecd292be24e8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7533,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:567"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"decayingexplorationstochasticexpectedabsoluterealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticexpectedabsoluterealizedbehaviorregret expected absolute scheduled stochastic realized-behavior regret. definition compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound","label":"decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound","description":"Planner radius, count-failure contribution, and normalized stochastic MGF first-moment contribution.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-7724297c98ff","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7534,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:584"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (rewardVarianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"decayingexplorationstochasticexpectedabsoluterealizedbehaviorregretbound banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticexpectedabsoluterealizedbehaviorregretbound planner radius, count-failure contribution, and normalized stochastic mgf first-moment contribution. definition compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_nonneg","label":"decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_nonneg","description":"Expected absolute stochastic realized regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-554952f6b05a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7535,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:598"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (n : Nat) : 0 <= decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState baseVisitFloor n","missing":[],"search":"decayingexplorationstochasticexpectedabsoluterealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticexpectedabsoluterealizedbehaviorregret_nonneg expected absolute stochastic realized regret is nonnegative. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","label":"decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","description":"The explicit stochastic expected-absolute bound is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-3decfae24959","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7536,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:610"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) (n : Nat) : 0 <= decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy n","missing":[],"search":"decayingexplorationstochasticexpectedabsoluterealizedbehaviorregretbound_nonneg banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticexpectedabsoluterealizedbehaviorregretbound_nonneg the explicit stochastic expected-absolute bound is nonnegative. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_le_bound","label":"decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_le_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_le_bound","description":"The common-space expected absolute stochastic realized regret is bounded by the planner radius, one count-failure contribution, and the directly integrated normalized sub-Gaussian deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-fea1f8f153c6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7537,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:644"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_le_bound (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState baseVisitFloor n <= de…","missing":[],"search":"decayingexplorationstochasticexpectedabsoluterealizedbehaviorregret_le_bound banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticexpectedabsoluterealizedbehaviorregret_le_bound the common-space expected absolute stochastic realized regret is bounded by the planner radius, one count-failure contribution, and the directly integrated normalized sub-gaussian deviation. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","label":"decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","description":"The explicit stochastic expected-absolute envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-686ece100c15","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7538,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:850"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) : Tendsto (decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy) atTop (nhds 0)","missing":[],"search":"decayingexplorationstochasticexpectedabsoluterealizedbehaviorregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticexpectedabsoluterealizedbehaviorregretbound_tendsto_zero the explicit stochastic expected-absolute envelope tends to zero. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","label":"exploratorySource_decayingExplorationStochasticCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","description":"Terminal expected-consistency theorem for unbounded sampled rewards: every coordinate is integrable, obeys the explicit three-term bound, and its expected absolute realized regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-8bda54eed8a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7539,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:880"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_decayingExplorationStochasticCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall n, Integrable (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSour…","missing":[],"search":"exploratorysource_decayingexplorationstochasticcommonmeasure_integrable_expectedabsoluterealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_decayingexplorationstochasticcommonmeasure_integrable_expectedabsoluterealizedbehaviorregret_tendsto_zero terminal expected-consistency theorem for unbounded sampled rewards: every coordinate is integrable, obeys the explicit three-term bound, and its expected absolute realized regret tends to zero. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.memLp_one_decayingExplorationStochasticRealizedBehaviorRegretProcess","label":"memLp_one_decayingExplorationStochasticRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.memLp_one_decayingExplorationStochasticRealizedBehaviorRegretProcess","description":"Every scheduled stochastic common-space coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-d4b0715e7267","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7540,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:933"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_decayingExplorationStochasticRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : MemLp (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState baseVisitFloor n) 1 (decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor)","missing":[],"search":"memlp_one_decayingexplorationstochasticrealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.memlp_one_decayingexplorationstochasticrealizedbehaviorregretprocess every scheduled stochastic common-space coordinate belongs to `l1`. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_eq","label":"eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_eq","description":"At exponent one, `eLpNorm` is the lifted expected absolute regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-0e34090d4366","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7541,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:956"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : eLpNorm (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState baseVisitFloor n) 1 (decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor) = ENNReal.ofReal (decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorReg…","missing":[],"search":"elpnorm_one_decayingexplorationstochasticrealizedbehaviorregretprocess_eq banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.elpnorm_one_decayingexplorationstochasticrealizedbehaviorregretprocess_eq at exponent one, `elpnorm` is the lifted expected absolute regret. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_tendsto_zero","label":"eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_tendsto_zero","description":"The exponent-one extended norm of the stochastic process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-25ed6b966770","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7542,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:984"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => eLpNorm (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState baseVis…","missing":[],"search":"elpnorm_one_decayingexplorationstochasticrealizedbehaviorregretprocess_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.elpnorm_one_decayingexplorationstochasticrealizedbehaviorregretprocess_tendsto_zero the exponent-one extended norm of the stochastic process tends to zero. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","label":"eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","description":"Canonical `L1` norm-of-the-difference convergence.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-b517a16e5053","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7543,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:1018"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => eLpNorm (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultStat…","missing":[],"search":"elpnorm_one_decayingexplorationstochasticrealizedbehaviorregretprocess_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.elpnorm_one_decayingexplorationstochasticrealizedbehaviorregretprocess_sub_zero_tendsto_zero canonical `l1` norm-of-the-difference convergence. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp","label":"decayingExplorationStochasticRealizedBehaviorRegretLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp","description":"The stochastic scheduled realized-regret process as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-acc5afb898e0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7544,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:1052"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticRealizedBehaviorRegretLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : Lp Real 1 (decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFloor)","missing":[],"search":"decayingexplorationstochasticrealizedbehaviorregretlp banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticrealizedbehaviorregretlp the stochastic scheduled realized-regret process as an `lp real 1` value. definition compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp_coeFn_ae_eq","label":"decayingExplorationStochasticRealizedBehaviorRegretLp_coeFn_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp_coeFn_ae_eq","description":"The named stochastic `Lp` coordinate represents the original process a.e.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-0717d2bd63d1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7545,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:1073"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticRealizedBehaviorRegretLp_coeFn_ae_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (n : Nat) : (decayingExplorationStochasticRealizedBehaviorRegretLp mdp initialState rewardSource rewardVarianceProxy law initialTable defaultState baseVisitFloor hrewardBound n : DecayingExplorationStochasticWindowSpace mdp baseVisitFloor -> Real) =ᵐ[ decayingExplorationStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState baseVisitFl…","missing":[],"search":"decayingexplorationstochasticrealizedbehaviorregretlp_coefn_ae_eq banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticrealizedbehaviorregretlp_coefn_ae_eq the named stochastic `lp` coordinate represents the original process a.e. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp_tendsto_zero","label":"decayingExplorationStochasticRealizedBehaviorRegretLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp_tendsto_zero","description":"The named stochastic `Lp Real 1` process converges to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-7ea400c20e3e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7546,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:1097"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticRealizedBehaviorRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (decayingExplorationStochasticRealizedBehaviorRegretLp mdp initialState rewardSource rewardVarianceProxy law initialTable defaultState baseVisitFloor hrewardB…","missing":[],"search":"decayingexplorationstochasticrealizedbehaviorregretlp_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationstochasticrealizedbehaviorregretlp_tendsto_zero the named stochastic `lp real 1` process converges to zero. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","label":"exploratorySource_decayingExplorationStochasticCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","description":"Terminal stochastic `L1` theorem: coordinate membership, exact exponent-one norms, `Lp` convergence, and induced convergence in measure hold on the same independent product of complete scheduled experiments.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-c93596c9b64c/index.html#decl-4894a18ba11a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","order":7547,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency.lean:1143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_decayingExplorationStochasticCommonMeasure_memLp_eLpNorm_L1_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall n, MemLp (decayingExplorationStochasticRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState baseVisit…","missing":[],"search":"exploratorysource_decayingexplorationstochasticcommonmeasure_memlp_elpnorm_l1_tendsto_zero banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_decayingexplorationstochasticcommonmeasure_memlp_elpnorm_l1_tendsto_zero terminal stochastic `l1` theorem: coordinate membership, exact exponent-one norms, `lp` convergence, and induced convergence in measure hold on the same independent product of complete scheduled experiments. theorem compiled","shard":"modules/9a92f03a55ac7941.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalReturnDeviationPerEpisodeVarianceProxy","label":"globalReturnDeviationPerEpisodeVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.globalReturnDeviationPerEpisodeVarianceProxy","description":"One-episode proxy underlying the globally centered stochastic batch proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-a8bc6fe7ff49","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7548,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def globalReturnDeviationPerEpisodeVarianceProxy (mdp : MDP State Action) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"globalreturndeviationperepisodevarianceproxy banditrlproof.finitehorizonrl.mdp.globalreturndeviationperepisodevarianceproxy one-episode proxy underlying the globally centered stochastic batch proxy. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidGlobalSampledCumulativeReturnDeviationVarianceProxy_eq","label":"iidGlobalSampledCumulativeReturnDeviationVarianceProxy_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.iidGlobalSampledCumulativeReturnDeviationVarianceProxy_eq","description":"The honest global iid proxy is exactly episode-linear.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-b688e9c8af0d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7549,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidGlobalSampledCumulativeReturnDeviationVarianceProxy_eq (mdp : MDP State Action) (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : mdp.iidGlobalSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy = (episodes : NNReal) * mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy","missing":[],"search":"iidglobalsampledcumulativereturndeviationvarianceproxy_eq banditrlproof.finitehorizonrl.mdp.iidglobalsampledcumulativereturndeviationvarianceproxy_eq the honest global iid proxy is exactly episode-linear. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalReturnDeviationPerEpisodeVarianceProxy_pos","label":"globalReturnDeviationPerEpisodeVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.globalReturnDeviationPerEpisodeVarianceProxy_pos","description":"Positive horizon and positive reward bound make the per-episode proxy positive.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-9723f7d79371","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7550,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem globalReturnDeviationPerEpisodeVarianceProxy_pos (mdp : MDP State Action) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : 0 < rewardBound) (hhorizon : 0 < mdp.horizon) : 0 < mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy","missing":[],"search":"globalreturndeviationperepisodevarianceproxy_pos banditrlproof.finitehorizonrl.mdp.globalreturndeviationperepisodevarianceproxy_pos positive horizon and positive reward bound make the per-episode proxy positive. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_coe","label":"cumulativeSuccessorGlobalReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_coe","description":"The successor proxy is exactly rounds times episodes times one base proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-70bc9e546401","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7551,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:89"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorGlobalReturnVarianceProxy_coe (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : ((cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy : NNReal) : Real) = (rounds : Real) * (episodes : Real) * (mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy : Real)","missing":[],"search":"cumulativesuccessorglobalreturnvarianceproxy_coe banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturnvarianceproxy_coe the successor proxy is exactly rounds times episodes times one base proxy. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius","label":"normalizedSuccessorGlobalReturnConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius","description":"Globally centered stochastic return radius normalized by all successor samples.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-e8e7bba77002","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7552,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedSuccessorGlobalReturnConfidenceRadius (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : Real","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius globally centered stochastic return radius normalized by all successor samples. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_nonneg","label":"normalizedSuccessorGlobalReturnConfidenceRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_nonneg","description":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : 0 <= normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy delta","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-02b1285bed54","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7553,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:114"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_nonneg (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : 0 <= normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy delta","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius_nonneg theorem normalizedsuccessorglobalreturnconfidenceradius_nonneg (mdp : mdp state action) (episodes rounds : nat) (rewardbound rewardvarianceproxy : nnreal) (delta : real) : 0 <= normalizedsuccessorglobalreturnconfidenceradius mdp episodes rounds rewardbound rewardvarianceproxy delta theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_eq","label":"normalizedSuccessorGlobalReturnConfidenceRadius_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_eq","description":"Exact normalized-radius formula after exposing episode and round scaling.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-0eb6a4c7f12a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7554,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:128"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_eq (mdp : MDP State Action) (episodes rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) (hepisodes : 0 < episodes) (hrounds : 0 < rounds) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy delta = Real.sqrt (2 * (mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy : Real) * Real.log (2 / delta) / ((episodes : Real) * (rounds : Real)))","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius_eq exact normalized-radius formula after exposing episode and round scaling. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticReturnRadiusEnvelope","label":"decayingExplorationStochasticReturnRadiusEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticReturnRadiusEnvelope","description":"A simple inverse-scale envelope for the scheduled stochastic radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-e3d32ea790c1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7555,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:187"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticReturnRadiusEnvelope (mdp : MDP State Action) (rewardBound rewardVarianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"decayingexplorationstochasticreturnradiusenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticreturnradiusenvelope a simple inverse-scale envelope for the scheduled stochastic radius. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_le_decayingEnvelope","label":"normalizedSuccessorGlobalReturnConfidenceRadius_le_decayingEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_le_decayingEnvelope","description":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_le_decayingEnvelope (mdp : MDP State Action) (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) (n : Nat) (hepisodes : 0 < episodes) : normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) rewardBound rewardVarianceProxy (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n) <=…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-b4e2c9d5ab2c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7556,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:198"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_le_decayingEnvelope (mdp : MDP State Action) (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) (n : Nat) (hepisodes : 0 < episodes) : normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) rewardBound rewardVarianceProxy (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n) <= decayingExplorationStochasticReturnRadiusEnvelope mdp rewardBound rewardVarianceProxy n","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius_le_decayingenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius_le_decayingenvelope theorem normalizedsuccessorglobalreturnconfidenceradius_le_decayingenvelope (mdp : mdp state action) (episodes : nat) (rewardbound rewardvarianceproxy : nnreal) (n : nat) (hepisodes : 0 < episodes) : normalizedsuccessorglobalreturnconfidenceradius mdp episodes (adaptiveepisodebatchsource.decayingexplorationrounds mdp n) rewardbound rewardvarianceproxy (adaptiveepisodebatchsource.vanishingaverageconfidencedelta n) <= decayingexplorationstochasticreturnradiusenvelope mdp rewardbound rewardvarianceproxy n theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticReturnRadiusEnvelope_tendsto_zero","label":"decayingExplorationStochasticReturnRadiusEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticReturnRadiusEnvelope_tendsto_zero","description":"theorem decayingExplorationStochasticReturnRadiusEnvelope_tendsto_zero (mdp : MDP State Action) (rewardBound rewardVarianceProxy : NNReal) : Tendsto (decayingExplorationStochasticReturnRadiusEnvelope mdp rewardBound rewardVarianceProxy) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-a96613ae5e1c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7557,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:287"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticReturnRadiusEnvelope_tendsto_zero (mdp : MDP State Action) (rewardBound rewardVarianceProxy : NNReal) : Tendsto (decayingExplorationStochasticReturnRadiusEnvelope mdp rewardBound rewardVarianceProxy) atTop (nhds 0)","missing":[],"search":"decayingexplorationstochasticreturnradiusenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticreturnradiusenvelope_tendsto_zero theorem decayingexplorationstochasticreturnradiusenvelope_tendsto_zero (mdp : mdp state action) (rewardbound rewardvarianceproxy : nnreal) : tendsto (decayingexplorationstochasticreturnradiusenvelope mdp rewardbound rewardvarianceproxy) attop (nhds 0) theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationNormalizedSuccessorGlobalReturnRadius_tendsto_zero","label":"decayingExplorationNormalizedSuccessorGlobalReturnRadius_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationNormalizedSuccessorGlobalReturnRadius_tendsto_zero","description":"The exact scheduled stochastic return radius tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-c969cfe50a68","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7558,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:312"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationNormalizedSuccessorGlobalReturnRadius_tendsto_zero (mdp : MDP State Action) (baseVisitFloor : Real) (rewardBound rewardVarianceProxy : NNReal) : Tendsto (fun n => normalizedSuccessorGlobalReturnConfidenceRadius mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n) (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) rewardBound rewardVarianceProxy (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n)) atTop (nhds 0)","missing":[],"search":"decayingexplorationnormalizedsuccessorglobalreturnradius_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationnormalizedsuccessorglobalreturnradius_tendsto_zero the exact scheduled stochastic return radius tends to zero. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound","label":"decayingExplorationStochasticAverageRealizedBehaviorRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound","description":"Stochastic realized-behavior certificate at schedule index `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-db102fbcb3e7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7559,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:342"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticAverageRealizedBehaviorRegretBound (mdp : MDP State Action) (baseVisitFloor : Real) (rewardVarianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"decayingexplorationstochasticaveragerealizedbehaviorregretbound banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticaveragerealizedbehaviorregretbound stochastic realized-behavior certificate at schedule index `n`. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureBudget","label":"decayingExplorationStochasticRealizedFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureBudget","description":"Projected-count and stochastic-return deviations consume one share each.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-d6c0a9b01a57","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7560,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:355"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationStochasticRealizedFailureBudget (n : Nat) : ENNReal","missing":[],"search":"decayingexplorationstochasticrealizedfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticrealizedfailurebudget projected-count and stochastic-return deviations consume one share each. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound_nonneg","label":"decayingExplorationStochasticAverageRealizedBehaviorRegretBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound_nonneg","description":"theorem decayingExplorationStochasticAverageRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) (n : Nat) : 0 <= decayingExplorationStochasticAverageRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy n","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-4b26680fe049","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7561,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:360"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticAverageRealizedBehaviorRegretBound_nonneg (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) (n : Nat) : 0 <= decayingExplorationStochasticAverageRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy n","missing":[],"search":"decayingexplorationstochasticaveragerealizedbehaviorregretbound_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticaveragerealizedbehaviorregretbound_nonneg theorem decayingexplorationstochasticaveragerealizedbehaviorregretbound_nonneg (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) (rewardvarianceproxy : nnreal) (n : nat) : 0 <= decayingexplorationstochasticaveragerealizedbehaviorregretbound mdp basevisitfloor rewardvarianceproxy n theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound_tendsto_zero","label":"decayingExplorationStochasticAverageRealizedBehaviorRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound_tendsto_zero","description":"theorem decayingExplorationStochasticAverageRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) : Tendsto (decayingExplorationStochasticAverageRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-aa27228fef42","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7562,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:383"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticAverageRealizedBehaviorRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) : Tendsto (decayingExplorationStochasticAverageRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy) atTop (nhds 0)","missing":[],"search":"decayingexplorationstochasticaveragerealizedbehaviorregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticaveragerealizedbehaviorregretbound_tendsto_zero theorem decayingexplorationstochasticaveragerealizedbehaviorregretbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) (rewardvarianceproxy : nnreal) : tendsto (decayingexplorationstochasticaveragerealizedbehaviorregretbound mdp basevisitfloor rewardvarianceproxy) attop (nhds 0) theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureBudget_tendsto_zero","label":"decayingExplorationStochasticRealizedFailureBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureBudget_tendsto_zero","description":"theorem decayingExplorationStochasticRealizedFailureBudget_tendsto_zero : Tendsto decayingExplorationStochasticRealizedFailureBudget atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-ffc45c05b625","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7563,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:399"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticRealizedFailureBudget_tendsto_zero : Tendsto decayingExplorationStochasticRealizedFailureBudget atTop (nhds 0)","missing":[],"search":"decayingexplorationstochasticrealizedfailurebudget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticrealizedfailurebudget_tendsto_zero theorem decayingexplorationstochasticrealizedfailurebudget_tendsto_zero : tendsto decayingexplorationstochasticrealizedfailurebudget attop (nhds 0) theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureAndRegretBound_tendsto_zero","label":"decayingExplorationStochasticRealizedFailureAndRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureAndRegretBound_tendsto_zero","description":"theorem decayingExplorationStochasticRealizedFailureAndRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) : Tendsto (fun n => (decayingExplorationStochasticRealizedFailureBudget n, decayingExplorationStochasticAverageRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy n)) atTop (nh…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-6db8d9135129","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7564,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:406"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationStochasticRealizedFailureAndRegretBound_tendsto_zero (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (baseVisitFloor : Real) (hbaseVisitFloor : 0 < baseVisitFloor) (rewardVarianceProxy : NNReal) : Tendsto (fun n => (decayingExplorationStochasticRealizedFailureBudget n, decayingExplorationStochasticAverageRealizedBehaviorRegretBound mdp baseVisitFloor rewardVarianceProxy n)) atTop (nhds (0, 0))","missing":[],"search":"decayingexplorationstochasticrealizedfailureandregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationstochasticrealizedfailureandregretbound_tendsto_zero theorem decayingexplorationstochasticrealizedfailureandregretbound_tendsto_zero (mdp : mdp state action) (hhorizon : 0 < mdp.horizon) (basevisitfloor : real) (hbasevisitfloor : 0 < basevisitfloor) (rewardvarianceproxy : nnreal) : tendsto (fun n => (decayingexplorationstochasticrealizedfailurebudget n, decayingexplorationstochasticaveragerealizedbehaviorregretbound mdp basevisitfloor rewardvarianceproxy n)) attop (nhds (0, 0)) theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource","label":"exploratorySource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource","description":"Stochastic-reward lift of the cumulative empirical-optimistic exploratory source. The cumulative table selector is evaluated only on projected history.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-66ae6cc88ed5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7565,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:429"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratorySource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"exploratorysource banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource stochastic-reward lift of the cumulative empirical-optimistic exploratory source. the cumulative table selector is evaluated only on projected history. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_initialPolicy","label":"exploratorySource_initialPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_initialPolicy","description":"The initial policies of the stochastic lift and deterministic source agree.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-9b486f631c78","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7566,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:483"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_initialPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate).initialPolicy = (AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).initialPolicy","missing":[],"search":"exploratorysource_initialpolicy banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_initialpolicy the initial policies of the stochastic lift and deterministic source agree. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorPolicy","label":"exploratorySource_successorPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorPolicy","description":"Every stochastic successor policy is the deterministic projected policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-e1c779b0fcaa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7567,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:499"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate).successorPolicy n history = (AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).successorPolicy n (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix (mdp := mdp) episodes n history)","missing":[],"search":"exploratorysource_successorpolicy banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_successorpolicy every stochastic successor policy is the deterministic projected policy. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_initialBatch_map_knownRewardEpisodeBatch","label":"exploratorySource_initialBatch_map_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_initialBatch_map_knownRewardEpisodeBatch","description":"The initial stochastic batch maps to the deterministic cumulative fiber.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-df17a82f36fd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7568,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:519"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_initialBatch_map_knownRewardEpisodeBatch {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : (rewardSource.iidStochasticTrajectoryFamilyMeasure (exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate).initialPolicy initialState episodes).map (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch (mdp := mdp) episodes) = (AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).initialPolicy.iidEpisodeBat…","missing":[],"search":"exploratorysource_initialbatch_map_knownrewardepisodebatch banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_initialbatch_map_knownrewardepisodebatch the initial stochastic batch maps to the deterministic cumulative fiber. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_batchKernel_map_knownRewardEpisodeBatch","label":"exploratorySource_batchKernel_map_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_batchKernel_map_knownRewardEpisodeBatch","description":"Every selected stochastic successor batch maps to its deterministic fiber.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-88683e8a4c5c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7569,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:544"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_batchKernel_map_knownRewardEpisodeBatch {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : ((exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate).batchKernel n history).map (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch (mdp := mdp) episodes) = (AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate).batchKernel n (MDP.MeanCom…","missing":[],"search":"exploratorysource_batchkernel_map_knownrewardepisodebatch banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_batchkernel_map_knownrewardepisodebatch every selected stochastic successor batch maps to its deterministic fiber. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","label":"exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","description":"Projected prefix and next-batch joint law of the cumulative source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-8378ecad7551","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7570,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:584"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) : let stochasticSource := exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate let deterministicSource := AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate stochasticSource.trajectoryMeasure.map (fun trajectory => (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix (mdp := m…","missing":[],"search":"exploratorysource_trajectorymeasure_map_projectedprefix_next_eq_compprod banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_map_projectedprefix_next_eq_compprod projected prefix and next-batch joint law of the cumulative source. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_condDistrib_projectedNext","label":"exploratorySource_trajectoryMeasure_condDistrib_projectedNext","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_condDistrib_projectedNext","description":"Conditional projected next-batch law for the cumulative source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-380d10d671ed","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7571,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:694"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_condDistrib_projectedNext {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [Nonempty (EpisodeBatch mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) : let stochasticSource := exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate let deterministicSource := AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate ProbabilityTheory.condDistrib (fun trajectory : Stoc…","missing":[],"search":"exploratorysource_trajectorymeasure_conddistrib_projectednext banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_conddistrib_projectednext conditional projected next-batch law for the cumulative source. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","label":"exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","description":"The complete known-reward projection equals the deterministic source law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-13e0899192ca","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7572,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:738"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [Nonempty (EpisodeBatch mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : let stochasticSource := exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate let deterministicSource := AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState countRadius explorationRate hexplorationRate stochasticSource.trajectoryMeasure.map (MDP.MeanCo…","missing":[],"search":"exploratorysource_trajectorymeasure_map_knownrewardepisodebatchtrajectory banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_map_knownrewardepisodebatchtrajectory the complete known-reward projection equals the deterministic source law. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.projectedAdaptiveCumulativeCountBadEvent","label":"projectedAdaptiveCumulativeCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.projectedAdaptiveCumulativeCountBadEvent","description":"Pullback of the deterministic cumulative count event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-71b7dd42d82c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7573,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:827"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def projectedAdaptiveCumulativeCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (delta : Real) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"projectedadaptivecumulativecountbadevent banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.projectedadaptivecumulativecountbadevent pullback of the deterministic cumulative count event. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_projected","label":"exploratorySource_successorExpectedCumulativeRegret_eq_projected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_projected","description":"Stochastic successor expected regret is the projected cumulative quantity.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-9583e8adf1d9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7574,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:869"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedCumulativeRegret_eq_projected {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate).successorExpectedCumulativeRegret trajectory rounds = adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret (initialState := initialState) (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory (mdp := mdp) episodes trajectory) defaultState countRadiu…","missing":[],"search":"exploratorysource_successorexpectedcumulativeregret_eq_projected banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_successorexpectedcumulativeregret_eq_projected stochastic successor expected regret is the projected cumulative quantity. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_projected","label":"exploratorySource_successorExpectedAverageRegret_eq_projected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_projected","description":"Average stochastic successor expected regret is the projected average.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-457c938cefc9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7575,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:896"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedAverageRegret_eq_projected {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (countRadius : TransitionCountRadius) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState countRadius explorationRate hexplorationRate).successorExpectedAverageRegret trajectory rounds = adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret (initialState := initialState) (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory (mdp := mdp) episodes trajectory) defaultState countRadi…","missing":[],"search":"exploratorysource_successorexpectedaverageregret_eq_projected banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_successorexpectedaverageregret_eq_projected average stochastic successor expected regret is the projected average. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","label":"exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","description":"The projected stochastic source inherits the deterministic decaying count, optimism, and expected exploratory-behavior certificate for one window.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-b481bf929ee3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7576,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:920"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [Nonempty (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 b…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaverageexploratorybehaviorexpectedregret the projected stochastic source inherits the deterministic decaying count, optimism, and expected exploratory-behavior certificate for one window. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationAverageRealizedBehaviorRegretViolationSet","label":"decayingExplorationAverageRealizedBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationAverageRealizedBehaviorRegretViolationSet","description":"Stochastic realized-regret violation set for one scheduled window.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-022a68a1a7f2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7577,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:1032"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def decayingExplorationAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (rewardVarianceProxy : NNReal) (n : Nat) : Set (StochasticEpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))","missing":[],"search":"decayingexplorationaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.decayingexplorationaveragerealizedbehaviorregretviolationset stochastic realized-regret violation set for one scheduled window. definition compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","label":"exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","description":"One stochastic finite window: projected counts and globally centered returns cover the realized-regret violation set under the two-share scheduled budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-35aaa727fdd9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7578,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:1065"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [Nonempty (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [StandardBorelSpace (StochasticEpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))] [Nonempty (StochasticEpisodeBatch mdp (Adaptive…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorconsistency banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorconsistency one stochastic finite window: projected counts and globally centered returns cover the realized-regret violation set under the two-share scheduled budget. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","label":"exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","description":"All scheduled stochastic windows plus the joint scalar limit. The indexed Borel witnesses expose the changing deterministic and stochastic sample spaces.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexplorationconsistency/index.html#decl-cd377c475b69","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","order":7579,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency.lean:1259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (hdetBatchBorel : forall n, StandardBorelSpace (EpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (hdetTrajectoryBorel : forall n, StandardBorelSpace (EpisodeBatchTrajectory mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (hstochasticBatchBorel : forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes mdp baseVisitFloor n))) (hstochasticTrajectoryBorel : forall n, StandardBorelSpace (StochasticEpisodeBatchT…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_decayingexplorationaveragerealizedbehaviorconsistency_allwindows banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_decayingexplorationaveragerealizedbehaviorconsistency_allwindows all scheduled stochastic windows plus the joint scalar limit. the indexed borel witnesses expose the changing deterministic and stochastic sample spaces. theorem compiled","shard":"modules/2a78859d916b17e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency_of_standardBorel","label":"exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency_of_standardBorel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency_of_standardBorel","description":"One scheduled stochastic window with every composite Standard Borel instance inferred from the state and action spaces.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-54b2a9d89968/index.html#decl-3617bbabb4f7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","order":7580,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency_of_standardBorel (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (baseVisitFloor : Real) (n : Nat) [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpisodeBatchSource.decayingExplorationR…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorconsistency_of_standardborel banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_optimism_and_decayingexplorationaveragerealizedbehaviorconsistency_of_standardborel one scheduled stochastic window with every composite standard borel instance inferred from the state and action spaces. theorem compiled","shard":"modules/f37d4f8f41f1fe22.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel","label":"exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel","description":"All scheduled stochastic windows and the joint scalar limit, with no indexed batch or trajectory Standard Borel witnesses supplied by the caller.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardcumulativedecayingexploration-54b2a9d89968/index.html#decl-7cad5df68c21","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","order":7581,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (rewardVarianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (AdaptiveStochasticEpisodeBatchSource.decayingExplorati…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_decayingexplorationaveragerealizedbehaviorconsistency_allwindows_of_standardborel banditrlproof.finitehorizonrl.adaptivecumulativestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedcumulativeinversesqrtpathsupport_decayingexplorationaveragerealizedbehaviorconsistency_allwindows_of_standardborel all scheduled stochastic windows and the joint scalar limit, with no indexed batch or trajectory standard borel witnesses supplied by the caller. theorem compiled","shard":"modules/f37d4f8f41f1fe22.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ProbabilityTheory.measure_compProd_map_prodMap_of_map_eq","label":"measure_compProd_map_prodMap_of_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ProbabilityTheory.measure_compProd_map_prodMap_of_map_eq","description":"Map both coordinates of a measure composition product when the base measure and every kernel fiber have the prescribed mapped laws.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-8ceb7f5dbeec","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7582,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_compProd_map_prodMap_of_map_eq {Alpha Alpha' Beta Beta' : Type*} [MeasurableSpace Alpha] [MeasurableSpace Alpha'] [MeasurableSpace Beta] [MeasurableSpace Beta'] (mu : Measure Alpha) [IsProbabilityMeasure mu] (kappa : ProbabilityTheory.Kernel Alpha Beta) [ProbabilityTheory.IsMarkovKernel kappa] (mu' : Measure Alpha') [IsProbabilityMeasure mu'] (kappa' : ProbabilityTheory.Kernel Alpha' Beta') [ProbabilityTheory.IsMarkovKernel kappa'] (f : Alpha -> Alpha') (hf : Measurable f) (g : Beta -> Beta') (hg : Measurable g) (hmu : mu.map f = mu') (hkappa : forall alpha, (kappa alpha).map g = kappa' (f alpha)) : (mu ⊗ₘ kappa).map (Prod.map f g) = mu' ⊗ₘ kappa'","missing":[],"search":"measure_compprod_map_prodmap_of_map_eq banditrlproof.finitehorizonrl.probabilitytheory.measure_compprod_map_prodmap_of_map_eq map both coordinates of a measure composition product when the base measure and every kernel fiber have the prescribed mapped laws. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch","label":"knownRewardEpisodeBatch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch","description":"Project one stochastic batch to the existing known-mean empirical batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-d5912a72cb58","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7583,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:89"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def knownRewardEpisodeBatch (episodes : Nat) (batch : StochasticEpisodeBatch mdp episodes) : EpisodeBatch mdp episodes","missing":[],"search":"knownrewardepisodebatch banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.knownrewardepisodebatch project one stochastic batch to the existing known-mean empirical batch. definition compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatch","label":"measurable_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatch","description":"The stochastic-batch known-reward projection is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-27d7249a7179","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7584,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_knownRewardEpisodeBatch (episodes : Nat) : Measurable (knownRewardEpisodeBatch (mdp := mdp) episodes)","missing":[],"search":"measurable_knownrewardepisodebatch banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_knownrewardepisodebatch the stochastic-batch known-reward projection is measurable. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix","label":"knownRewardEpisodeBatchPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix","description":"Project every coordinate of a finite stochastic batch prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-91386f732609","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7585,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def knownRewardEpisodeBatchPrefix (episodes n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : EpisodeBatchPrefix mdp episodes n","missing":[],"search":"knownrewardepisodebatchprefix banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.knownrewardepisodebatchprefix project every coordinate of a finite stochastic batch prefix. definition compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchPrefix","label":"measurable_knownRewardEpisodeBatchPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchPrefix","description":"The coordinatewise known-reward prefix projection is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-905cf5840502","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7586,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_knownRewardEpisodeBatchPrefix (episodes n : Nat) : Measurable (knownRewardEpisodeBatchPrefix (mdp := mdp) episodes n)","missing":[],"search":"measurable_knownrewardepisodebatchprefix banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_knownrewardepisodebatchprefix the coordinatewise known-reward prefix projection is measurable. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory","label":"knownRewardEpisodeBatchTrajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory","description":"Project every coordinate of an infinite stochastic batch trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-ac960f0b88d6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7587,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def knownRewardEpisodeBatchTrajectory (episodes : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : EpisodeBatchTrajectory mdp episodes","missing":[],"search":"knownrewardepisodebatchtrajectory banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.knownrewardepisodebatchtrajectory project every coordinate of an infinite stochastic batch trajectory. definition compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchTrajectory","label":"measurable_knownRewardEpisodeBatchTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchTrajectory","description":"The coordinatewise complete known-reward trajectory projection is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-98164c20ac43","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7588,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_knownRewardEpisodeBatchTrajectory (episodes : Nat) : Measurable (knownRewardEpisodeBatchTrajectory (mdp := mdp) episodes)","missing":[],"search":"measurable_knownrewardepisodebatchtrajectory banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_knownrewardepisodebatchtrajectory the coordinatewise complete known-reward trajectory projection is measurable. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix_frestrictLe","label":"knownRewardEpisodeBatchPrefix_frestrictLe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix_frestrictLe","description":"Projection commutes with restricting a complete trajectory to a prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-ead4a5341b90","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7589,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:139"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem knownRewardEpisodeBatchPrefix_frestrictLe (episodes n : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : knownRewardEpisodeBatchPrefix episodes n (Preorder.frestrictLe n trajectory) = Preorder.frestrictLe n (knownRewardEpisodeBatchTrajectory episodes trajectory)","missing":[],"search":"knownrewardepisodebatchprefix_frestrictle banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.knownrewardepisodebatchprefix_frestrictle projection commutes with restricting a complete trajectory to a prefix. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDStochasticEpisodeBatchKernel","label":"exploratoryIIDStochasticEpisodeBatchKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDStochasticEpisodeBatchKernel","description":"Stochastic iid episode-batch law indexed by an exploratory policy table.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-8d02b6cfb1d7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7590,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:150"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryIIDStochasticEpisodeBatchKernel {mdp : MDP State Action} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : ProbabilityTheory.Kernel (DeterministicMarkovPolicyTable mdp) (StochasticEpisodeBatch mdp episodes)","missing":[],"search":"exploratoryiidstochasticepisodebatchkernel banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryiidstochasticepisodebatchkernel stochastic iid episode-batch law indexed by an exploratory policy table. definition compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDStochasticEpisodeBatchKernel_apply","label":"exploratoryIIDStochasticEpisodeBatchKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDStochasticEpisodeBatchKernel_apply","description":"theorem exploratoryIIDStochasticEpisodeBatchKernel_apply {mdp : MDP State Action} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (table : DeterministicMarkovPolicyTable mdp) : exploratoryIIDStochasticEpisodeBatchKernel rewardSource initialState episodes exploration…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-e094d8d0f7e7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7591,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryIIDStochasticEpisodeBatchKernel_apply {mdp : MDP State Action} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (table : DeterministicMarkovPolicyTable mdp) : exploratoryIIDStochasticEpisodeBatchKernel rewardSource initialState episodes explorationRate hexplorationRate table = rewardSource.iidStochasticTrajectoryFamilyMeasure (table.exploratoryPolicy explorationRate hexplorationRate) initialState episodes","missing":[],"search":"exploratoryiidstochasticepisodebatchkernel_apply banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryiidstochasticepisodebatchkernel_apply theorem exploratoryiidstochasticepisodebatchkernel_apply {mdp : mdp state action} (rewardsource : mdp.meancompatiblerewardkernel) (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) (explorationrate : nnreal) (hexplorationrate : explorationrate <= 1) (table : deterministicmarkovpolicytable mdp) : exploratoryiidstochasticepisodebatchkernel rewardsource initialstate episodes explorationrate hexplorationrate table = rewardsource.iidstochastictrajectoryfamilymeasure (table.exploratorypolicy explorationrate hexplorationrate) initialstate episodes theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.measurable_selectedExploratorySampledReturnDeviation","label":"measurable_selectedExploratorySampledReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.measurable_selectedExploratorySampledReturnDeviation","description":"A finite table selector and its batch statistic form a measurable dynamic sampled-return deviation. This discharges the regularity field required by `AdaptiveStochasticEpisodeBatchSource` without constraining sampled rewards.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-c337fb77b363","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7592,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:202"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedExploratorySampledReturnDeviation {mdp : MDP State Action} {episodes : Nat} {History : Type*} [MeasurableSpace History] (selector : History -> DeterministicMarkovPolicyTable mdp) (hselector : Measurable selector) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : Measurable fun pair : History × StochasticEpisodeBatch mdp episodes => mdp.sampledCumulativeReturnDeviationSum ((selector pair.1).exploratoryPolicy explorationRate hexplorationRate) episodes pair.2","missing":[],"search":"measurable_selectedexploratorysampledreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.measurable_selectedexploratorysampledreturndeviation a finite table selector and its batch statistic form a measurable dynamic sampled-return deviation. this discharges the regularity field required by `adaptivestochasticepisodebatchsource` without constraining sampled rewards. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource","label":"exploratorySource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource","description":"Concrete stochastic-reward lift of the deterministic exploratory empirical optimistic source. Every policy update factors through known-reward history.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-8e4fa1b4e602","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7593,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:238"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratorySource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"exploratorysource banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource concrete stochastic-reward lift of the deterministic exploratory empirical optimistic source. every policy update factors through known-reward history. definition compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_initialPolicy","label":"exploratorySource_initialPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_initialPolicy","description":"The stochastic and deterministic initial policies agree definitionally.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-f7f522848dd0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7594,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:291"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_initialPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate).initialPolicy = (AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate).initialPolicy","missing":[],"search":"exploratorysource_initialpolicy banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_initialpolicy the stochastic and deterministic initial policies agree definitionally. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorPolicy","label":"exploratorySource_successorPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorPolicy","description":"The stochastic successor policy is selected from the projected prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-dfa7ba160bb1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7595,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:308"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate).successorPolicy n history = (AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate).successorPolicy n (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix (mdp := mdp) episodes n history)","missing":[],"search":"exploratorysource_successorpolicy banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_successorpolicy the stochastic successor policy is selected from the projected prefix. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_initialBatch_map_knownRewardEpisodeBatch","label":"exploratorySource_initialBatch_map_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_initialBatch_map_knownRewardEpisodeBatch","description":"The initial stochastic batch maps to the deterministic exploratory batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-e6d02e2fbcf2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7596,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:329"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_initialBatch_map_knownRewardEpisodeBatch {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : (rewardSource.iidStochasticTrajectoryFamilyMeasure (exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate).initialPolicy initialState episodes).map (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch (mdp := mdp) episodes) = (AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate).initialPolicy.iidEpisodeBatchMeasure initi…","missing":[],"search":"exploratorysource_initialbatch_map_knownrewardepisodebatch banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_initialbatch_map_knownrewardepisodebatch the initial stochastic batch maps to the deterministic exploratory batch. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_batchKernel_map_knownRewardEpisodeBatch","label":"exploratorySource_batchKernel_map_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_batchKernel_map_knownRewardEpisodeBatch","description":"Every selected stochastic successor batch maps to its deterministic fiber.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-069f710f9234","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7597,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:354"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_batchKernel_map_knownRewardEpisodeBatch {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : ((exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate).batchKernel n history).map (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch (mdp := mdp) episodes) = (AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate).batchKernel n (MDP.MeanCompatibleRewardKe…","missing":[],"search":"exploratorysource_batchkernel_map_knownrewardepisodebatch banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_batchkernel_map_knownrewardepisodebatch every selected stochastic successor batch maps to its deterministic fiber. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","label":"exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","description":"After projecting both coordinates, every stochastic prefix/next-batch joint law is the projected-prefix measure composed with the deterministic source kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-74208248c3d5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7598,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:396"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) : let stochasticSource := exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate let deterministicSource := AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate stochasticSource.trajectoryMeasure.map (fun trajectory => (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix (mdp := mdp) episodes n…","missing":[],"search":"exploratorysource_trajectorymeasure_map_projectedprefix_next_eq_compprod banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_map_projectedprefix_next_eq_compprod after projecting both coordinates, every stochastic prefix/next-batch joint law is the projected-prefix measure composed with the deterministic source kernel. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_condDistrib_projectedNext","label":"exploratorySource_trajectoryMeasure_condDistrib_projectedNext","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_condDistrib_projectedNext","description":"Conditioned on the projected stochastic prefix, the projected next batch has the deterministic exploratory empirical-optimistic source kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-9a03c3c1bc8e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7599,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:504"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_condDistrib_projectedNext {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [Nonempty (EpisodeBatch mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) : let stochasticSource := exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate let deterministicSource := AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate ProbabilityTheory.condDistrib (fun trajectory : StochasticEpisodeBa…","missing":[],"search":"exploratorysource_trajectorymeasure_conddistrib_projectednext banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_conddistrib_projectednext conditioned on the projected stochastic prefix, the projected next batch has the deterministic exploratory empirical-optimistic source kernel. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","label":"exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","description":"The complete known-reward projection of the concrete stochastic adaptive source is exactly the existing deterministic exploratory source trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-aa5756952239","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7600,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:550"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [Nonempty (EpisodeBatch mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : let stochasticSource := exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate let deterministicSource := AdaptiveEmpiricalOptimisticSource.exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate stochasticSource.trajectoryMeasure.map (MDP.MeanCompatibleRewardK…","missing":[],"search":"exploratorysource_trajectorymeasure_map_knownrewardepisodebatchtrajectory banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_map_knownrewardepisodebatchtrajectory the complete known-reward projection of the concrete stochastic adaptive source is exactly the existing deterministic exploratory source trajectory law. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAdaptiveSimultaneousCountBadEvent","label":"projectedAdaptiveSimultaneousCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAdaptiveSimultaneousCountBadEvent","description":"The stochastic count bad event is the inverse image of the deterministic adaptive count event under complete known-reward projection.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-b09205361f91","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7601,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:640"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def projectedAdaptiveSimultaneousCountBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (delta : Real) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"projectedadaptivesimultaneouscountbadevent banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.projectedadaptivesimultaneouscountbadevent the stochastic count bad event is the inverse image of the deterministic adaptive count event under complete known-reward projection. definition compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_recommendedExpectedRegret","label":"exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_recommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_recommendedExpectedRegret","description":"Route endpoint: the concrete stochastic source inherits the deterministic known-mean all-coordinate confidence event, projected optimism, and projected recommended-policy expected-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticprojection/index.html#decl-47ff6b5c1d48","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","order":7602,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection.lean:660"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_recommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (EpisodeBatch mdp episodes)] [Nonempty (EpisodeBatch mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (calibration : let deterministicSource := AdaptiveEmpiricalOptimisticSource.exploratorySource mdp i…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedallcoordinateconfidence_optimism_and_recommendedexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedallcoordinateconfidence_optimism_and_recommendedexpectedregret route endpoint: the concrete stochastic source inherits the deterministic known-mean all-coordinate confidence event, projected optimism, and projected recommended-policy expected-regret bound. theorem compiled","shard":"modules/3fb09df8644b7887.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport_two_delta","label":"trajectoryMeasure_expected_to_realized_successor_average_regret_transport_two_delta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport_two_delta","description":"Generic two-budget version of the expected-to-realized successor-regret transport. The caller's event keeps `countDelta`, while the return event uses `returnDelta`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-c8cd3b2e0e3f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7603,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_expected_to_realized_successor_average_regret_transport_two_delta {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes rewa…","missing":[],"search":"trajectorymeasure_expected_to_realized_successor_average_regret_transport_two_delta banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_expected_to_realized_successor_average_regret_transport_two_delta generic two-budget version of the expected-to-realized successor-regret transport. the caller's event keeps `countdelta`, while the return event uses `returndelta`. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedExploratoryBehaviorExpectedRegret","label":"projectedExploratoryBehaviorExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedExploratoryBehaviorExpectedRegret","description":"Sum of expected regrets of the projected empirical-optimistic exploratory behaviors.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-8abf1b72e3cf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7604,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def projectedExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) : Real","missing":[],"search":"projectedexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.projectedexploratorybehaviorexpectedregret sum of expected regrets of the projected empirical-optimistic exploratory behaviors. definition compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAverageExploratoryBehaviorExpectedRegret","label":"projectedAverageExploratoryBehaviorExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAverageExploratoryBehaviorExpectedRegret","description":"Average expected regret of the projected exploratory behaviors.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-a3b9fd1f8cc3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7605,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:171"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def projectedAverageExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) : Real","missing":[],"search":"projectedaverageexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.projectedaverageexploratorybehaviorexpectedregret average expected regret of the projected exploratory behaviors. definition compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedExploratoryBehaviorExpectedRegret_le","label":"projectedExploratoryBehaviorExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedExploratoryBehaviorExpectedRegret_le","description":"Exploratory behavior regret is recommendation regret plus one charge per round.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-9f486e15bf07","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7606,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:184"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem projectedExploratoryBehaviorExpectedRegret_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) : projectedExploratoryBehaviorExpectedRegret (initialState := initialState) trajectory defaultState transitionBonus explorationRate hexplorationRate rounds <= adaptiveEmpiricalOptimisticRecommendedExpectedRegret (initialState := initialState) trajectory defaultState transitionBonus rounds + (rounds : Real) * exploratoryBehaviorRegretCharge mdp explorationRate rewardBound","missing":[],"search":"projectedexploratorybehaviorexpectedregret_le banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.projectedexploratorybehaviorexpectedregret_le exploratory behavior regret is recommendation regret plus one charge per round. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAverageExploratoryBehaviorExpectedRegret_le","label":"projectedAverageExploratoryBehaviorExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAverageExploratoryBehaviorExpectedRegret_le","description":"Averaging removes the repeated-round exploration factor.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-d1c5f5b1a96d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7607,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:232"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem projectedAverageExploratoryBehaviorExpectedRegret_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : EpisodeBatchTrajectory mdp episodes) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) (hrounds : 0 < rounds) : projectedAverageExploratoryBehaviorExpectedRegret (initialState := initialState) trajectory defaultState transitionBonus explorationRate hexplorationRate rounds <= adaptiveEmpiricalOptimisticRecommendedExpectedRegret (initialState := initialState) trajectory defaultState transitionBonus rounds / (rounds : Real) + exploratoryBehaviorRegretCharge mdp explorationRate rewardBound","missing":[],"search":"projectedaverageexploratorybehaviorexpectedregret_le banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.projectedaverageexploratorybehaviorexpectedregret_le averaging removes the repeated-round exploration factor. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.measurable_selectedExploratoryGlobalReturnDeviation","label":"measurable_selectedExploratoryGlobalReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.measurable_selectedExploratoryGlobalReturnDeviation","description":"Dynamic global-return measurability for any finite exploratory table selector.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-820f348e4993","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7608,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedExploratoryGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} {episodes : Nat} {History : Type*} [MeasurableSpace History] (selector : History -> DeterministicMarkovPolicyTable mdp) (hselector : Measurable selector) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : Measurable fun pair : History × StochasticEpisodeBatch mdp episodes => mdp.globalSampledCumulativeReturnDeviationSum ((selector pair.1).exploratoryPolicy explorationRate hexplorationRate) initialState episodes pair.2","missing":[],"search":"measurable_selectedexploratoryglobalreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.measurable_selectedexploratoryglobalreturndeviation dynamic global-return measurability for any finite exploratory table selector. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_projected","label":"exploratorySource_successorExpectedCumulativeRegret_eq_projected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_projected","description":"The concrete source's successor policies are the projected exploratory policies.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-142c2195a690","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7609,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:330"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedCumulativeRegret_eq_projected {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate).successorExpectedCumulativeRegret trajectory rounds = projectedExploratoryBehaviorExpectedRegret (initialState := initialState) (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory (mdp := mdp) episodes trajectory) defaultState transitionBonus explorationRate hexplorationRat…","missing":[],"search":"exploratorysource_successorexpectedcumulativeregret_eq_projected banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_successorexpectedcumulativeregret_eq_projected the concrete source's successor policies are the projected exploratory policies. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_projected","label":"exploratorySource_successorExpectedAverageRegret_eq_projected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_projected","description":"Average form of the projected successor-policy identity.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-8f577b7423ea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7610,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:357"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedAverageRegret_eq_projected {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState transitionBonus explorationRate hexplorationRate).successorExpectedAverageRegret trajectory rounds = projectedAverageExploratoryBehaviorExpectedRegret (initialState := initialState) (MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory (mdp := mdp) episodes trajectory) defaultState transitionBonus explorationRate hexplorationRa…","missing":[],"search":"exploratorysource_successorexpectedaverageregret_eq_projected banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_successorexpectedaverageregret_eq_projected average form of the projected successor-policy identity. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","label":"exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","description":"Concrete fixed-window route endpoint: projected count confidence and optimism, exploratory behavior charge, and stochastic realized-return concentration hold simultaneously with separate confidence budgets.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardempiricaloptimisticrealizedbehaviorregret/index.html#decl-9f504d38fd79","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","order":7611,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret.lean:382"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (EpisodeBatch mdp episodes)] [Nonempty (EpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (…","missing":[],"search":"exploratorysource_trajectorymeasure_projectedallcoordinateconfidence_optimism_and_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticempiricaloptimisticsource.exploratorysource_trajectorymeasure_projectedallcoordinateconfidence_optimism_and_realizedsuccessoraverageregret concrete fixed-window route endpoint: projected count confidence and optimism, exploratory behavior charge, and stochastic realized-return concentration hold simultaneously with separate confidence budgets. theorem compiled","shard":"modules/732d9425e49f22fc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviation","label":"initialPolicyValueDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviation","description":"The sampled policy-value fluctuation at one complete episode coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-0dce92cddb3a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7612,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def initialPolicyValueDeviation (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (trajectory : State × RewardStepTrace Action State mdp.horizon) : Real","missing":[],"search":"initialpolicyvaluedeviation banditrlproof.finitehorizonrl.mdp.initialpolicyvaluedeviation the sampled policy-value fluctuation at one complete episode coordinate. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviation","label":"measurable_initialPolicyValueDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviation","description":"theorem measurable_initialPolicyValueDeviation (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) : Measurable (mdp.initialPolicyValueDeviation policy initialState)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-eb5e9346a9c8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7613,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_initialPolicyValueDeviation (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) : Measurable (mdp.initialPolicyValueDeviation policy initialState)","missing":[],"search":"measurable_initialpolicyvaluedeviation banditrlproof.finitehorizonrl.mdp.measurable_initialpolicyvaluedeviation theorem measurable_initialpolicyvaluedeviation (mdp : mdp state action) (policy : markovpolicy mdp) (initialstate : measure state) : measurable (mdp.initialpolicyvaluedeviation policy initialstate) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueVarianceProxy","label":"initialPolicyValueVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueVarianceProxy","description":"Hoeffding proxy for the policy value of one sampled initial state.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-9cf6b32898bc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7614,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def initialPolicyValueVarianceProxy (mdp : MDP State Action) (rewardBound : NNReal) : NNReal","missing":[],"search":"initialpolicyvaluevarianceproxy banditrlproof.finitehorizonrl.mdp.initialpolicyvaluevarianceproxy hoeffding proxy for the policy value of one sampled initial state. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_map_fst","label":"stochasticTrajectoryMeasure_map_fst","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_map_fst","description":"The generated complete trajectory has the exact supplied initial marginal.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-368349888bab","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7615,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryMeasure_map_fst (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : (source.stochasticTrajectoryMeasure policy initialState).map Prod.fst = initialState","missing":[],"search":"stochastictrajectorymeasure_map_fst banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure_map_fst the generated complete trajectory has the exact supplied initial marginal. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_initialPolicyValueDeviation_hasSubgaussianMGF","label":"stochasticTrajectoryMeasure_initialPolicyValueDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_initialPolicyValueDeviation_hasSubgaussianMGF","description":"The initial-state policy-value fluctuation is sub-Gaussian.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-aa6eafa57464","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7616,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryMeasure_initialPolicyValueDeviation_hasSubgaussianMGF (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) : ProbabilityTheory.HasSubgaussianMGF (mdp.initialPolicyValueDeviation policy initialState) (mdp.initialPolicyValueVarianceProxy rewardBound) (source.stochasticTrajectoryMeasure policy initialState)","missing":[],"search":"stochastictrajectorymeasure_initialpolicyvaluedeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure_initialpolicyvaluedeviation_hassubgaussianmgf the initial-state policy-value fluctuation is sub-gaussian. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviationAtEpisode","label":"initialPolicyValueDeviationAtEpisode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviationAtEpisode","description":"Initial-state value fluctuation at one coordinate of an iid episode family.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-35fa1a57040b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7617,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def initialPolicyValueDeviationAtEpisode (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) {episodes : Nat} (episode : Fin episodes) (trajectories : StochasticEpisodeBatch mdp episodes) : Real","missing":[],"search":"initialpolicyvaluedeviationatepisode banditrlproof.finitehorizonrl.mdp.initialpolicyvaluedeviationatepisode initial-state value fluctuation at one coordinate of an iid episode family. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviationAtEpisode","label":"measurable_initialPolicyValueDeviationAtEpisode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviationAtEpisode","description":"theorem measurable_initialPolicyValueDeviationAtEpisode (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) {episodes : Nat} (episode : Fin episodes) : Measurable (mdp.initialPolicyValueDeviationAtEpisode policy initialState episode)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-6981b0703e75","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7618,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:118"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_initialPolicyValueDeviationAtEpisode (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) {episodes : Nat} (episode : Fin episodes) : Measurable (mdp.initialPolicyValueDeviationAtEpisode policy initialState episode)","missing":[],"search":"measurable_initialpolicyvaluedeviationatepisode banditrlproof.finitehorizonrl.mdp.measurable_initialpolicyvaluedeviationatepisode theorem measurable_initialpolicyvaluedeviationatepisode (mdp : mdp state action) (policy : markovpolicy mdp) (initialstate : measure state) {episodes : nat} (episode : fin episodes) : measurable (mdp.initialpolicyvaluedeviationatepisode policy initialstate episode) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviationSum","label":"initialPolicyValueDeviationSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviationSum","description":"Sum of initial-state policy-value fluctuations in one iid batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-dfe03a2e0518","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7619,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def initialPolicyValueDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) (trajectories : StochasticEpisodeBatch mdp episodes) : Real","missing":[],"search":"initialpolicyvaluedeviationsum banditrlproof.finitehorizonrl.mdp.initialpolicyvaluedeviationsum sum of initial-state policy-value fluctuations in one iid batch. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviationSum","label":"measurable_initialPolicyValueDeviationSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviationSum","description":"theorem measurable_initialPolicyValueDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) : Measurable (mdp.initialPolicyValueDeviationSum policy initialState episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-beb0f2ee9d5b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7620,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:135"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_initialPolicyValueDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) : Measurable (mdp.initialPolicyValueDeviationSum policy initialState episodes)","missing":[],"search":"measurable_initialpolicyvaluedeviationsum banditrlproof.finitehorizonrl.mdp.measurable_initialpolicyvaluedeviationsum theorem measurable_initialpolicyvaluedeviationsum (mdp : mdp state action) (policy : markovpolicy mdp) (initialstate : measure state) (episodes : nat) : measurable (mdp.initialpolicyvaluedeviationsum policy initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidInitialPolicyValueDeviationVarianceProxy","label":"iidInitialPolicyValueDeviationVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.iidInitialPolicyValueDeviationVarianceProxy","description":"Episode-linear proxy for the initial-state value fluctuation sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-ed8faef82f52","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7621,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidInitialPolicyValueDeviationVarianceProxy (mdp : MDP State Action) (episodes : Nat) (rewardBound : NNReal) : NNReal","missing":[],"search":"iidinitialpolicyvaluedeviationvarianceproxy banditrlproof.finitehorizonrl.mdp.iidinitialpolicyvaluedeviationvarianceproxy episode-linear proxy for the initial-state value fluctuation sum. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardSum","label":"sampledCumulativeRewardSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardSum","description":"Sum of all sampled rewards in one complete stochastic episode batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-1dab9fd1a940","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7622,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:151"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def sampledCumulativeRewardSum (mdp : MDP State Action) (episodes : Nat) (trajectories : StochasticEpisodeBatch mdp episodes) : Real","missing":[],"search":"sampledcumulativerewardsum banditrlproof.finitehorizonrl.mdp.sampledcumulativerewardsum sum of all sampled rewards in one complete stochastic episode batch. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardSum","label":"measurable_sampledCumulativeRewardSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardSum","description":"theorem measurable_sampledCumulativeRewardSum (mdp : MDP State Action) (episodes : Nat) : Measurable (mdp.sampledCumulativeRewardSum episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-79a52bf41b1e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7623,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeRewardSum (mdp : MDP State Action) (episodes : Nat) : Measurable (mdp.sampledCumulativeRewardSum episodes)","missing":[],"search":"measurable_sampledcumulativerewardsum banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativerewardsum theorem measurable_sampledcumulativerewardsum (mdp : mdp state action) (episodes : nat) : measurable (mdp.sampledcumulativerewardsum episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalSampledCumulativeReturnDeviationSum","label":"globalSampledCumulativeReturnDeviationSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.globalSampledCumulativeReturnDeviationSum","description":"One batch's sampled return centered by the selected policy's global initial-law value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-89d9df967726","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7624,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def globalSampledCumulativeReturnDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) (trajectories : StochasticEpisodeBatch mdp episodes) : Real","missing":[],"search":"globalsampledcumulativereturndeviationsum banditrlproof.finitehorizonrl.mdp.globalsampledcumulativereturndeviationsum one batch's sampled return centered by the selected policy's global initial-law value. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalSampledCumulativeReturnDeviationSum_eq","label":"globalSampledCumulativeReturnDeviationSum_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.globalSampledCumulativeReturnDeviationSum_eq","description":"The global batch deviation is actual sampled return minus its policy mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-b2ce2c9dca8f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7625,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:178"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem globalSampledCumulativeReturnDeviationSum_eq (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) (trajectories : StochasticEpisodeBatch mdp episodes) : mdp.globalSampledCumulativeReturnDeviationSum policy initialState episodes trajectories = mdp.sampledCumulativeRewardSum episodes trajectories - (episodes : Real) * integral initialState (policy.valueAt 0 (Nat.zero_le mdp.horizon))","missing":[],"search":"globalsampledcumulativereturndeviationsum_eq banditrlproof.finitehorizonrl.mdp.globalsampledcumulativereturndeviationsum_eq the global batch deviation is actual sampled return minus its policy mean. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_globalSampledCumulativeReturnDeviationSum","label":"measurable_globalSampledCumulativeReturnDeviationSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_globalSampledCumulativeReturnDeviationSum","description":"theorem measurable_globalSampledCumulativeReturnDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) : Measurable (mdp.globalSampledCumulativeReturnDeviationSum policy initialState episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-c68ef5c801c9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7626,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:200"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_globalSampledCumulativeReturnDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) (episodes : Nat) : Measurable (mdp.globalSampledCumulativeReturnDeviationSum policy initialState episodes)","missing":[],"search":"measurable_globalsampledcumulativereturndeviationsum banditrlproof.finitehorizonrl.mdp.measurable_globalsampledcumulativereturndeviationsum theorem measurable_globalsampledcumulativereturndeviationsum (mdp : mdp state action) (policy : markovpolicy mdp) (initialstate : measure state) (episodes : nat) : measurable (mdp.globalsampledcumulativereturndeviationsum policy initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidGlobalSampledCumulativeReturnDeviationVarianceProxy","label":"iidGlobalSampledCumulativeReturnDeviationVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.iidGlobalSampledCumulativeReturnDeviationVarianceProxy","description":"Honest same-space proxy for the two globally centered batch components.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-5fab77ebf199","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7627,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:211"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidGlobalSampledCumulativeReturnDeviationVarianceProxy (mdp : MDP State Action) (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"iidglobalsampledcumulativereturndeviationvarianceproxy banditrlproof.finitehorizonrl.mdp.iidglobalsampledcumulativereturndeviationvarianceproxy honest same-space proxy for the two globally centered batch components. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_initialPolicyValueDeviationAtEpisode","label":"iIndepFun_initialPolicyValueDeviationAtEpisode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_initialPolicyValueDeviationAtEpisode","description":"theorem iIndepFun_initialPolicyValueDeviationAtEpisode (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.initialPolicyValueDeviationAtEpisode policy initialState episode trajectories) (source.iidStochasticTrajectoryFamilyMeasure policy initialState epi…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-59c3da430869","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7628,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:223"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_initialPolicyValueDeviationAtEpisode (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.initialPolicyValueDeviationAtEpisode policy initialState episode trajectories) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iindepfun_initialpolicyvaluedeviationatepisode banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iindepfun_initialpolicyvaluedeviationatepisode theorem iindepfun_initialpolicyvaluedeviationatepisode (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) : probabilitytheory.iindepfun (fun episode trajectories => mdp.initialpolicyvaluedeviationatepisode policy initialstate episode trajectories) (source.iidstochastictrajectoryfamilymeasure policy initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.initialPolicyValueDeviationAtEpisode_hasSubgaussianMGF","label":"initialPolicyValueDeviationAtEpisode_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.initialPolicyValueDeviationAtEpisode_hasSubgaussianMGF","description":"theorem initialPolicyValueDeviationAtEpisode_hasSubgaussianMGF (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) {episodes : Nat} (episode : Fin episodes) : ProbabilityTheory.HasSubgaussianMGF (mdp.initialPolicyValueDeviationAt…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-26268a5b2120","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7629,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:238"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem initialPolicyValueDeviationAtEpisode_hasSubgaussianMGF (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) {episodes : Nat} (episode : Fin episodes) : ProbabilityTheory.HasSubgaussianMGF (mdp.initialPolicyValueDeviationAtEpisode policy initialState episode) (mdp.initialPolicyValueVarianceProxy rewardBound) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"initialpolicyvaluedeviationatepisode_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.initialpolicyvaluedeviationatepisode_hassubgaussianmgf theorem initialpolicyvaluedeviationatepisode_hassubgaussianmgf (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardbound : nnreal) (hrewardbound : ∀ state action, |mdp.reward state action| ≤ (rewardbound : real)) {episodes : nat} (episode : fin episodes) : probabilitytheory.hassubgaussianmgf (mdp.initialpolicyvaluedeviationatepisode policy initialstate episode) (mdp.initialpolicyvaluevarianceproxy rewardbound) (source.iidstochastictrajectoryfamilymeasure policy initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_initialPolicyValueDeviationSum_hasSubgaussianMGF","label":"iidStochasticTrajectoryFamilyMeasure_initialPolicyValueDeviationSum_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_initialPolicyValueDeviationSum_hasSubgaussianMGF","description":"theorem iidStochasticTrajectoryFamilyMeasure_initialPolicyValueDeviationSum_hasSubgaussianMGF (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) : ProbabilityTheory.HasSubgaussianMGF (mdp.initialPolicyValueDevia…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-c0e75edc8184","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7630,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:266"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_initialPolicyValueDeviationSum_hasSubgaussianMGF (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) : ProbabilityTheory.HasSubgaussianMGF (mdp.initialPolicyValueDeviationSum policy initialState episodes) (mdp.iidInitialPolicyValueDeviationVarianceProxy episodes rewardBound) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iidstochastictrajectoryfamilymeasure_initialpolicyvaluedeviationsum_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_initialpolicyvaluedeviationsum_hassubgaussianmgf theorem iidstochastictrajectoryfamilymeasure_initialpolicyvaluedeviationsum_hassubgaussianmgf (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) (rewardbound : nnreal) (hrewardbound : ∀ state action, |mdp.reward state action| ≤ (rewardbound : real)) : probabilitytheory.hassubgaussianmgf (mdp.initialpolicyvaluedeviationsum policy initialstate episodes) (mdp.iidinitialpolicyvaluedeviationvarianceproxy episodes rewardbound) (source.iidstochastictrajectoryfamilymeasure policy initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_globalSampledCumulativeReturnDeviationSum_hasSubgaussianMGF","label":"iidStochasticTrajectoryFamilyMeasure_globalSampledCumulativeReturnDeviationSum_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_globalSampledCumulativeReturnDeviationSum_hasSubgaussianMGF","description":"theorem iidStochasticTrajectoryFamilyMeasure_globalSampledCumulativeReturnDeviationSum_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-25a58df3fee2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7631,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:291"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_globalSampledCumulativeReturnDeviationSum_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasSubgaussianMGF (mdp.globalSampledCumulativeReturnDeviationSum policy initialState episodes) (mdp.iidGlobalSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iidstochastictrajectoryfamilymeasure_globalsampledcumulativereturndeviationsum_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_globalsampledcumulativereturndeviationsum_hassubgaussianmgf theorem iidstochastictrajectoryfamilymeasure_globalsampledcumulativereturndeviationsum_hassubgaussianmgf [standardborelspace state] [standardborelspace action] [nonempty action] (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (episodes : nat) (rewardbound rewardvarianceproxy : nnreal) (hrewardbound : ∀ state action, |mdp.reward state action| ≤ (rewardbound : real)) (law : source.uniformsubgaussianrewardlaw rewardvarianceproxy) : probabilitytheory.hassubgaussianmgf (mdp.globalsampledcumulativereturndeviationsum policy initialstate episodes) (mdp.iidglobalsampledcumulativereturndeviationvarianceproxy episodes rewardbound rewardvarianceproxy) (source.iidstochastictrajectoryfamilymeasure policy initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.GlobalReturnMeasurability","label":"GlobalReturnMeasurability","kind":"typeclass","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.GlobalReturnMeasurability","description":"Explicit measurability of the history-selected globally centered batch return.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-99addd110a7d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7632,"meta":[["Kind","typeclass"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:325"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"class GlobalReturnMeasurability {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : Prop where","missing":[],"search":"globalreturnmeasurability banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.globalreturnmeasurability explicit measurability of the history-selected globally centered batch return. typeclass compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviation","label":"successorGlobalReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviation","description":"Dynamic globally centered return on a prefix/next-batch pair.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-a8aa7d05585b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7633,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:338"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (pair : StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes) : Real","missing":[],"search":"successorglobalreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturndeviation dynamic globally centered return on a prefix/next-batch pair. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviation","label":"measurable_successorGlobalReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviation","description":"theorem measurable_successorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviation n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-948391351370","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7634,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:348"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviation n)","missing":[],"search":"measurable_successorglobalreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_successorglobalreturndeviation theorem measurable_successorglobalreturndeviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) : measurable (source.successorglobalreturndeviation n) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel","label":"successorGlobalReturnDeviationKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel","description":"Selected conditional law of the globally centered next-batch return.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-10de6d86e471","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7635,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:357"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviationKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.Kernel (StochasticEpisodeBatchPrefix mdp episodes n) Real","missing":[],"search":"successorglobalreturndeviationkernel banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturndeviationkernel selected conditional law of the globally centered next-batch return. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel_apply","label":"successorGlobalReturnDeviationKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel_apply","description":"theorem successorGlobalReturnDeviationKernel_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : source.successorGlobalReturnDeviationKernel n history = (source.rewardSource.iidSt…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-328c482f2940","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7636,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:378"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnDeviationKernel_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : source.successorGlobalReturnDeviationKernel n history = (source.rewardSource.iidStochasticTrajectoryFamilyMeasure (source.successorPolicy n history) initialState episodes).map (mdp.globalSampledCumulativeReturnDeviationSum (source.successorPolicy n history) initialState episodes)","missing":[],"search":"successorglobalreturndeviationkernel_apply banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturndeviationkernel_apply theorem successorglobalreturndeviationkernel_apply {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) (history : stochasticepisodebatchprefix mdp episodes n) : source.successorglobalreturndeviationkernel n history = (source.rewardsource.iidstochastictrajectoryfamilymeasure (source.successorpolicy n history) initialstate episodes).map (mdp.globalsampledcumulativereturndeviationsum (source.successorpolicy n history) initialstate episodes) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationAt","label":"successorGlobalReturnDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationAt","description":"Trajectory-level globally centered return at successor coordinate `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-232360b921f3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7637,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:412"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorglobalreturndeviationat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturndeviationat trajectory-level globally centered return at successor coordinate `n + 1`. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviationAt","label":"measurable_successorGlobalReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviationAt","description":"theorem measurable_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviationAt n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-b175b62427c9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7638,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:420"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviationAt n)","missing":[],"search":"measurable_successorglobalreturndeviationat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_successorglobalreturndeviationat theorem measurable_successorglobalreturndeviationat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) : measurable (source.successorglobalreturndeviationat n) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","label":"trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","description":"theorem trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Probabilit…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-45ba24d5fc90","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7639,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:430"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : ProbabilityTheory.condDistrib (source.successorGlobalReturnDeviationAt n) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] source.successorGlobalReturnDeviationKernel n","missing":[],"search":"trajectorymeasure_conddistrib_successorglobalreturndeviationat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_conddistrib_successorglobalreturndeviationat theorem trajectorymeasure_conddistrib_successorglobalreturndeviationat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (stochasticepisodebatch mdp episodes)] [nonempty (stochasticepisodebatch mdp episodes)] (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) : probabilitytheory.conddistrib (source.successorglobalreturndeviationat n) (preorder.frestrictle n) source.trajectorymeasure =ᵐ[ source.trajectorymeasure.map (preorder.frestrictle n)] source.successorglobalreturndeviationkernel n theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorGlobalReturnDeviationAt_eq","label":"condExpKernel_map_successorGlobalReturnDeviationAt_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorGlobalReturnDeviationAt_eq","description":"theorem condExpKernel_map_successorGlobalReturnDeviationAt_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episode…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-1be6ff5ae8a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7640,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:455"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_successorGlobalReturnDeviationAt_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Filter.Eventually (fun trajectory : StochasticEpisodeBatchTrajectory mdp episodes => Measure.map (source.successorGlobalReturnDeviationAt n) (ProbabilityTheory.condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (StochasticEpisodeBatchPrefix mdp episodes n)).comap (Preorder.frestrictLe n)) trajectory) = source.successorGlobalReturnDeviationKernel n (Preorder.frestrictLe n trajectory)) (ae (source.trajector…","missing":[],"search":"condexpkernel_map_successorglobalreturndeviationat_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.condexpkernel_map_successorglobalreturndeviationat_eq theorem condexpkernel_map_successorglobalreturndeviationat_eq {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace (stochasticepisodebatch mdp episodes)] [nonempty (stochasticepisodebatch mdp episodes)] [standardborelspace (stochasticepisodebatchtrajectory mdp episodes)] (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) : filter.eventually (fun trajectory : stochasticepisodebatchtrajectory mdp episodes => measure.map (source.successorglobalreturndeviationat n) (probabilitytheory.condexpkernel source.trajectorymeasure ((inferinstance : measurablespace (stochasticepisodebatchprefix mdp episodes n)).comap (preorder.frestrictle n)) trajectory) = source.successorglobalreturndeviationkernel n (preorder.frestrictle n trajectory)) (ae (source.trajectorymeasure.trim (preorder.measurable_frestrictle n).comap_le)) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnPrefixIncrement","label":"successorGlobalReturnPrefixIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnPrefixIncrement","description":"Prefix process that deliberately leaves coordinate zero uncharged.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-2cb75f5120af","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7641,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:485"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : (round : Nat) → StochasticEpisodeBatchPrefix mdp episodes round → Real | 0, _history => 0 | n + 1, history => source.successorGlobalReturnDeviation n (Preorder.frestrictLe₂ (π := fun _ : Nat => StochasticEpisodeBatch mdp episodes) (Nat.le_succ n) history, history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) theorem measurable_successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (round : Nat) : Measurable (source.successorGlobalReturnPrefixIncrement round)","missing":[],"search":"successorglobalreturnprefixincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturnprefixincrement prefix process that deliberately leaves coordinate zero uncharged. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnPrefixIncrement","label":"measurable_successorGlobalReturnPrefixIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnPrefixIncrement","description":"theorem measurable_successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (round : Nat) : Measurable (source.successorGlobalReturnPrefixIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-0033693fb9e9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7642,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:498"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (round : Nat) : Measurable (source.successorGlobalReturnPrefixIncrement round)","missing":[],"search":"measurable_successorglobalreturnprefixincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_successorglobalreturnprefixincrement theorem measurable_successorglobalreturnprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (round : nat) : measurable (source.successorglobalreturnprefixincrement round) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement","label":"successorGlobalReturnIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement","description":"noncomputable def successorGlobalReturnIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-8bb36e479d6e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7643,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:517"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorglobalreturnincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturnincrement noncomputable def successorglobalreturnincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) (round : nat) (trajectory : stochasticepisodebatchtrajectory mdp episodes) : real definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_stronglyAdapted_piLE","label":"successorGlobalReturnIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_stronglyAdapted_piLE","description":"theorem successorGlobalReturnIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp episodes)) source.successorGlobalReturnIncrement","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-12505419d056","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7644,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:526"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp episodes)) source.successorGlobalReturnIncrement","missing":[],"search":"successorglobalreturnincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturnincrement_stronglyadapted_pile theorem successorglobalreturnincrement_stronglyadapted_pile {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] : stronglyadapted (filtration.pile (x := fun _ : nat => stochasticepisodebatch mdp episodes)) source.successorglobalreturnincrement theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","label":"successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","description":"theorem successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episo…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-5cd60ceb2a19","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7645,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:540"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasCondSubgaussianMGF (Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp episodes) n) ((Filtration.piLE (X := fun _ : Na…","missing":[],"search":"successorglobalreturnincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturnincrement_succ_hascondsubgaussianmgf theorem successorglobalreturnincrement_succ_hascondsubgaussianmgf {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace state] [standardborelspace action] [nonempty action] [standardborelspace (stochasticepisodebatch mdp episodes)] [nonempty (stochasticepisodebatch mdp episodes)] [standardborelspace (stochasticepisodebatchtrajectory mdp episodes)] (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) (rewardbound rewardvarianceproxy : nnreal) (hrewardbound : ∀ state action, |mdp.reward state action| ≤ (rewardbound : real)) (law : source.rewardsource.uniformsubgaussianrewardlaw rewardvarianceproxy) : probabilitytheory.hascondsubgaussianmgf (filtration.pile (x := fun _ : nat => stochasticepisodebatch mdp episodes) n) ((filtration.pile (x := fun _ : nat => stochasticepisodebatch mdp episodes)).le n) (source.successorglobalreturnincrement (n + 1)) (mdp.iidglobalsampledcumulativereturndeviationvarianceproxy episodes rewardbound rewardvarianceproxy) source.trajectorymeasure theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy","label":"cumulativeSuccessorGlobalReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy","description":"Zero plus `rounds` successor proxies.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-952ee2fecfbc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7646,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:648"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSuccessorGlobalReturnVarianceProxy (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"cumulativesuccessorglobalreturnvarianceproxy banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturnvarianceproxy zero plus `rounds` successor proxies. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_eq","label":"cumulativeSuccessorGlobalReturnVarianceProxy_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_eq","description":"theorem cumulativeSuccessorGlobalReturnVarianceProxy_eq (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy = (rounds : NNReal) * mdp.iidGlobalSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-2a671dcfbd30","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7647,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:658"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorGlobalReturnVarianceProxy_eq (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy = (rounds : NNReal) * mdp.iidGlobalSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy","missing":[],"search":"cumulativesuccessorglobalreturnvarianceproxy_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturnvarianceproxy_eq theorem cumulativesuccessorglobalreturnvarianceproxy_eq (mdp : mdp state action) (rounds episodes : nat) (rewardbound rewardvarianceproxy : nnreal) : cumulativesuccessorglobalreturnvarianceproxy mdp rounds episodes rewardbound rewardvarianceproxy = (rounds : nnreal) * mdp.iidglobalsampledcumulativereturndeviationvarianceproxy episodes rewardbound rewardvarianceproxy theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation","label":"cumulativeSuccessorGlobalReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation","description":"Cumulative globally centered deviation over successor coordinates only.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-671c7bd964be","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7648,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:670"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSuccessorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativesuccessorglobalreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturndeviation cumulative globally centered deviation over successor coordinates only. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","label":"trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","description":"theorem trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTraject…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-e254b300f39f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7649,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:679"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy : NNReal) : Real)) (del…","missing":[],"search":"trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le theorem trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace state] [standardborelspace action] [nonempty action] [standardborelspace (stochasticepisodebatch mdp episodes)] [nonempty (stochasticepisodebatch mdp episodes)] [standardborelspace (stochasticepisodebatchtrajectory mdp episodes)] (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) (rewardbound rewardvarianceproxy : nnreal) (hrewardbound : ∀ state action, |mdp.reward state action| ≤ (rewardbound : real)) (law : source.rewardsource.uniformsubgaussianrewardlaw rewardvarianceproxy) (htotal : 0 < ((cumulativesuccessorglobalreturnvarianceproxy mdp rounds episodes rewardbound rewardvarianceproxy : nnreal) : real)) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) : source.trajectorymeasure {trajectory | concentration.subgaussiansumconfidenceradius (cumulativesuccessorglobalreturnvarianceproxy mdp rounds episodes rewardbound rewardvarianceproxy) delta ≤ |source.cumulativesuccessorglobalretur…","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorPolicyAt","label":"successorPolicyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorPolicyAt","description":"The selected successor policy after observing coordinates through `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-91d3479fc761","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7650,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:740"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorPolicyAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) : MarkovPolicy mdp","missing":[],"search":"successorpolicyat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorpolicyat the selected successor policy after observing coordinates through `n`. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn","label":"optimalInitialExpectedReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn","description":"Optimal expected return under the supplied initial-state law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-4ce3f70d27e3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7651,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:749"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalInitialExpectedReturn [Nonempty Action] (mdp : MDP State Action) (initialState : Measure State) : Real","missing":[],"search":"optimalinitialexpectedreturn banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.optimalinitialexpectedreturn optimal expected return under the supplied initial-state law. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedCumulativeRegret","label":"successorExpectedCumulativeRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedCumulativeRegret","description":"Sum of selected successor-policy expected regrets.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-46c65da3c875","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7652,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:756"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorExpectedCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [Nonempty Action] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"successorexpectedcumulativeregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorexpectedcumulativeregret sum of selected successor-policy expected regrets. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret","label":"successorExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret","description":"Average selected successor-policy expected regret per adaptive round.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-ed419924f3fa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7653,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:766"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorExpectedAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [Nonempty Action] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"successorexpectedaverageregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorexpectedaverageregret average selected successor-policy expected regret per adaptive round. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret","label":"realizedSuccessorCumulativeRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret","description":"Realized regret of all sampled successor batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-baf0265dd922","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7654,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:775"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [Nonempty Action] {episodes : Nat} (_source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"realizedsuccessorcumulativeregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.realizedsuccessorcumulativeregret realized regret of all sampled successor batches. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret","label":"realizedSuccessorAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret","description":"Realized successor regret averaged over sampled episodes and rounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-53c88992e2bc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7655,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:787"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [Nonempty Action] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.realizedsuccessoraverageregret realized successor regret averaged over sampled episodes and rounds. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_eq","label":"successorGlobalReturnIncrement_succ_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_eq","description":"The adaptive successor increment is actual batch return minus policy mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-fecce1a14b47","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7656,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:797"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnIncrement_succ_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) : source.successorGlobalReturnIncrement (n + 1) trajectory = mdp.sampledCumulativeRewardSum episodes (trajectory (n + 1)) - (episodes : Real) * integral initialState ((source.successorPolicyAt trajectory n).valueAt 0 (Nat.zero_le mdp.horizon))","missing":[],"search":"successorglobalreturnincrement_succ_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturnincrement_succ_eq the adaptive successor increment is actual batch return minus policy mean. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","label":"cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","description":"The cumulative deviation is the finite sum over successor coordinates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-672e9bd4fb43","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7657,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:816"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory = ∑ round : Fin rounds, source.successorGlobalReturnIncrement ((round : Nat) + 1) trajectory","missing":[],"search":"cumulativesuccessorglobalreturndeviation_eq_fin_sum banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturndeviation_eq_fin_sum the cumulative deviation is the finite sum over successor coordinates. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","label":"realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","description":"Exact realized equals expected minus globally centered return deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-fce1fb212b5a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7658,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:836"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem realizedSuccessorCumulativeRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [Nonempty Action] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.realizedSuccessorCumulativeRegret trajectory rounds = (episodes : Real) * source.successorExpectedCumulativeRegret trajectory rounds - source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory","missing":[],"search":"realizedsuccessorcumulativeregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.realizedsuccessorcumulativeregret_eq_expected_sub_deviation exact realized equals expected minus globally centered return deviation. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","label":"realizedSuccessorAverageRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","description":"Exact averaged form of the stochastic realized-regret decomposition.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-9f48c8936553","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7659,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:875"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem realizedSuccessorAverageRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] [Nonempty Action] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) : source.realizedSuccessorAverageRegret trajectory rounds = source.successorExpectedAverageRegret trajectory rounds - source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory / ((episodes : Real) * (rounds : Real))","missing":[],"search":"realizedsuccessoraverageregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.realizedsuccessoraverageregret_eq_expected_sub_deviation exact averaged form of the stochastic realized-regret decomposition. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationBadEvent","label":"successorGlobalReturnDeviationBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationBadEvent","description":"Return-deviation event used by fixed-window stochastic regret transport.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-adcbc4f6228f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7660,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:894"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"successorglobalreturndeviationbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorglobalreturndeviationbadevent return-deviation event used by fixed-window stochastic regret transport. definition compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_cumulativeSuccessorGlobalReturnDeviation","label":"measurable_cumulativeSuccessorGlobalReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_cumulativeSuccessorGlobalReturnDeviation","description":"theorem measurable_cumulativeSuccessorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) : Measurable (source.cumulativeSuccessorGlobalReturnDeviation rounds)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-c85370712272","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7661,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:906"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeSuccessorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) : Measurable (source.cumulativeSuccessorGlobalReturnDeviation rounds)","missing":[],"search":"measurable_cumulativesuccessorglobalreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_cumulativesuccessorglobalreturndeviation theorem measurable_cumulativesuccessorglobalreturndeviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) : measurable (source.cumulativesuccessorglobalreturndeviation rounds) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_successorGlobalReturnDeviationBadEvent","label":"measurableSet_successorGlobalReturnDeviationBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_successorGlobalReturnDeviationBadEvent","description":"theorem measurableSet_successorGlobalReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : MeasurableSet (source.successorGlobalReturnDeviationBadEvent roun…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-1c25ed41a582","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7662,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:917"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_successorGlobalReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : MeasurableSet (source.successorGlobalReturnDeviationBadEvent rounds rewardBound rewardVarianceProxy delta)","missing":[],"search":"measurableset_successorglobalreturndeviationbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurableset_successorglobalreturndeviationbadevent theorem measurableset_successorglobalreturndeviationbadevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) (rewardbound rewardvarianceproxy : nnreal) (delta : real) : measurableset (source.successorglobalreturndeviationbadevent rounds rewardbound rewardvarianceproxy delta) theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","label":"trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","description":"theorem trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp epi…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-567788bc8bc5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7663,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:929"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy : NNReal) : Real)) (delta : Real)…","missing":[],"search":"trajectorymeasure_successorglobalreturndeviationbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_successorglobalreturndeviationbadevent_le theorem trajectorymeasure_successorglobalreturndeviationbadevent_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} [standardborelspace state] [standardborelspace action] [nonempty action] [standardborelspace (stochasticepisodebatch mdp episodes)] [nonempty (stochasticepisodebatch mdp episodes)] [standardborelspace (stochasticepisodebatchtrajectory mdp episodes)] (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) (rewardbound rewardvarianceproxy : nnreal) (hrewardbound : ∀ state action, |mdp.reward state action| ≤ (rewardbound : real)) (law : source.rewardsource.uniformsubgaussianrewardlaw rewardvarianceproxy) (htotal : 0 < ((cumulativesuccessorglobalreturnvarianceproxy mdp rounds episodes rewardbound rewardvarianceproxy : nnreal) : real)) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) : source.trajectorymeasure (source.successorglobalreturndeviationbadevent rounds rewardbound rewardvarianceproxy delta) ≤ ennreal.ofreal delta theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport","label":"trajectoryMeasure_expected_to_realized_successor_average_regret_transport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport","description":"Combine a caller-supplied count/optimism event with the stochastic return event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardrealizedbehaviorregret/index.html#decl-513c106da022","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","order":7664,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret.lean:959"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_expected_to_realized_successor_average_regret_transport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((cumulativeSuccessorGlobalReturnVarianceProxy mdp rounds episodes re…","missing":[],"search":"trajectorymeasure_expected_to_realized_successor_average_regret_transport banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_expected_to_realized_successor_average_regret_transport combine a caller-supplied count/optimism event with the stochastic return event. theorem compiled","shard":"modules/7eb45ee6543142db.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.policyAt","label":"policyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.policyAt","description":"Policy whose iid stochastic law generated a batch coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-48b267c774ff","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7665,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def policyAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Nat -> MarkovPolicy mdp | 0 => source.initialPolicy | n + 1 => source.successorPolicy n (Preorder.frestrictLe n trajectory) /-- Pull an initial stochastic-batch event back to the adaptive trajectory. -/ def initialBadEvent {mdp : MDP State Action} {episodes : Nat} (bad : Set (StochasticEpisodeBatch mdp episodes)) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"policyat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.policyat policy whose iid stochastic law generated a batch coordinate. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialBadEvent","label":"initialBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialBadEvent","description":"Pull an initial stochastic-batch event back to the adaptive trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-85cbba8423d6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7666,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def initialBadEvent {mdp : MDP State Action} {episodes : Nat} (bad : Set (StochasticEpisodeBatch mdp episodes)) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"initialbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.initialbadevent pull an initial stochastic-batch event back to the adaptive trajectory. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorBadEvent","label":"successorBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorBadEvent","description":"Pull a prefix-dependent successor stochastic-batch event back to the trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-9db7e99f28dc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7667,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:48"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def successorBadEvent {mdp : MDP State Action} {episodes : Nat} (n : Nat) (bad : Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"successorbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorbadevent pull a prefix-dependent successor stochastic-batch event back to the trajectory. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.roundBadEvent","label":"roundBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.roundBadEvent","description":"Round-indexed adapted stochastic-batch event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-0a52accb7cd5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7668,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def roundBadEvent {mdp : MDP State Action} {episodes : Nat} (initialBad : Set (StochasticEpisodeBatch mdp episodes)) (successorBad : (n : Nat) -> Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)) : Nat -> Set (StochasticEpisodeBatchTrajectory mdp episodes) | 0 => initialBadEvent initialBad | n + 1 => successorBadEvent n (successorBad n) /-- Union of the first `rounds` adapted stochastic-batch events. -/ def finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat} (rounds : Nat) (initialBad : Set (StochasticEpisodeBatch mdp episodes)) (successorBad : (n : Nat) -> Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"roundbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.roundbadevent round-indexed adapted stochastic-batch event. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent","label":"finiteHorizonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent","description":"Union of the first `rounds` adapted stochastic-batch events.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-3c2e47ca8298","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7669,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat} (rounds : Nat) (initialBad : Set (StochasticEpisodeBatch mdp episodes)) (successorBad : (n : Nat) -> Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"finitehorizonbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.finitehorizonbadevent union of the first `rounds` adapted stochastic-batch events. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","label":"measurableSet_finiteHorizonBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","description":"Measurability of the finite adapted stochastic-batch union.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-d9261f523958","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7670,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_finiteHorizonBadEvent {mdp : MDP State Action} {episodes rounds : Nat} {initialBad : Set (StochasticEpisodeBatch mdp episodes)} {successorBad : (n : Nat) -> Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)} (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (successorBad n)) : MeasurableSet (finiteHorizonBadEvent rounds initialBad successorBad)","missing":[],"search":"measurableset_finitehorizonbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurableset_finitehorizonbadevent measurability of the finite adapted stochastic-batch union. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_initialBadEvent","label":"trajectoryMeasure_initialBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_initialBadEvent","description":"Exact mass of a pulled-back initial stochastic-batch event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-51ee25b5a51d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7671,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_initialBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) {bad : Set (StochasticEpisodeBatch mdp episodes)} (hbad : MeasurableSet bad) : source.trajectoryMeasure (initialBadEvent bad) = source.rewardSource.iidStochasticTrajectoryFamilyMeasure source.initialPolicy initialState episodes bad","missing":[],"search":"trajectorymeasure_initialbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_initialbadevent exact mass of a pulled-back initial stochastic-batch event. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","label":"trajectoryMeasure_successorBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","description":"A successor event inherits a uniform bound on every history fiber.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-d4b5e435c972","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7672,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successorBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) {bad : Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)} (hbad : MeasurableSet bad) (budget : ENNReal) (hfiber : forall history, source.batchKernel n history (Prod.mk history ⁻¹' bad) <= budget) : source.trajectoryMeasure (successorBadEvent n bad) <= budget","missing":[],"search":"trajectorymeasure_successorbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_successorbadevent_le a successor event inherits a uniform bound on every history fiber. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent","label":"initialAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent","description":"Initial selected count-and-reward empirical-model event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-75eb1632531e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7673,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def initialAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) : Set (StochasticEpisodeBatch mdp episodes)","missing":[],"search":"initialallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.initialallcoordinateempiricalmodelbadevent initial selected count-and-reward empirical-model event. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent","label":"successorAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent","description":"Prefix-selected successor count-and-reward empirical-model event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-59da0b0fd1d5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7674,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:173"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) (n : Nat) : Set (StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes)","missing":[],"search":"successorallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorallcoordinateempiricalmodelbadevent prefix-selected successor count-and-reward empirical-model event. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.adaptiveAllCoordinateEmpiricalModelBadEvent","label":"adaptiveAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.adaptiveAllCoordinateEmpiricalModelBadEvent","description":"Finite-round selected count-and-reward empirical-model event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-c76f75c145d6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7675,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) : Set (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"adaptiveallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.adaptiveallcoordinateempiricalmodelbadevent finite-round selected count-and-reward empirical-model event. definition compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.multiBatchLocalDelta_pos_of_pos","label":"multiBatchLocalDelta_pos_of_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.multiBatchLocalDelta_pos_of_pos","description":"A positive global share gives a positive finite-round local share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-6c8f41a0a837","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7676,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:202"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem multiBatchLocalDelta_pos_of_pos {rounds : Nat} (hrounds : 0 < rounds) {delta : Real} (hdelta : 0 < delta) : 0 < multiBatchLocalDelta rounds delta","missing":[],"search":"multibatchlocaldelta_pos_of_pos banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.multibatchlocaldelta_pos_of_pos a positive global share gives a positive finite-round local share. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.multiBatchLocalDelta_le_one_of_le_one","label":"multiBatchLocalDelta_le_one_of_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.multiBatchLocalDelta_le_one_of_le_one","description":"A global share at most one gives every finite-round local share at most one.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-4a192a88ffce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7677,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem multiBatchLocalDelta_le_one_of_le_one {rounds : Nat} (hrounds : 0 < rounds) {delta : Real} (hdelta_le_one : delta <= 1) : multiBatchLocalDelta rounds delta <= 1","missing":[],"search":"multibatchlocaldelta_le_one_of_le_one banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.multibatchlocaldelta_le_one_of_le_one a global share at most one gives every finite-round local share at most one. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_initialAllCoordinateEmpiricalModelBadEvent","label":"measurableSet_initialAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_initialAllCoordinateEmpiricalModelBadEvent","description":"The initial selected empirical-model event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-8034810bc677","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7678,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:220"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_initialAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) : MeasurableSet (source.initialAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta)","missing":[],"search":"measurableset_initialallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurableset_initialallcoordinateempiricalmodelbadevent the initial selected empirical-model event is measurable. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","label":"measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","description":"Measurability of every selected successor event closes the global event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-cd1817ab8fa6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7679,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:234"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (source.successorAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta n)) : MeasurableSet (source.adaptiveAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta)","missing":[],"search":"measurableset_adaptiveallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurableset_adaptiveallcoordinateempiricalmodelbadevent measurability of every selected successor event closes the global event. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent_le","label":"initialAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent_le","description":"The initial selected empirical-model event receives its two local shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-4b9210ac6c00","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7680,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem initialAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) : (source.rewardSource.iidStochasticTrajectoryFamilyMeasure source.initialPolicy initialState episodes) (source.initialAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta re…","missing":[],"search":"initialallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.initialallcoordinateempiricalmodelbadevent_le the initial selected empirical-model event receives its two local shares. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent_fiber_le","label":"successorAllCoordinateEmpiricalModelBadEvent_fiber_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent_fiber_le","description":"Every selected successor fiber receives the same two local shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-3bad3214e0be","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7681,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:280"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorAllCoordinateEmpiricalModelBadEvent_fiber_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : source.batchKernel n history (Prod.mk history ⁻¹' source.successorAllCoordinateEmpiricalModelBadEvent rounds vari…","missing":[],"search":"successorallcoordinateempiricalmodelbadevent_fiber_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorallcoordinateempiricalmodelbadevent_fiber_le every selected successor fiber receives the same two local shares. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.ofReal_multiBatchLocalDelta_add","label":"ofReal_multiBatchLocalDelta_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.ofReal_multiBatchLocalDelta_add","description":"The two local shares are the local share of their sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-b763af453753","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7682,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:319"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ofReal_multiBatchLocalDelta_add (rounds : Nat) {countDelta rewardDelta : Real} (hcountDelta : 0 <= countDelta) (hrewardDelta : 0 <= rewardDelta) : ENNReal.ofReal (multiBatchLocalDelta rounds countDelta) + ENNReal.ofReal (multiBatchLocalDelta rounds rewardDelta) = ENNReal.ofReal (multiBatchLocalDelta rounds (countDelta + rewardDelta))","missing":[],"search":"ofreal_multibatchlocaldelta_add banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.ofreal_multibatchlocaldelta_add the two local shares are the local share of their sum. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","label":"trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","description":"Finite-round adaptive count-and-reward confidence under exact selected-policy iid fibers. The two global shares remain separate in the terminal bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-bda047aed7c1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7683,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:348"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (source.successorAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta n)) : source.traj…","missing":[],"search":"trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le finite-round adaptive count-and-reward confidence under exact selected-policy iid fibers. the two global shares remain separate in the terminal bound. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","label":"policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","description":"Outside the adaptive event, each batch avoids its generating policy's event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-5ec426778ff5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7684,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:425"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (varianceProxy : NNReal) (countDelta rewardDelta : Real) (htrajectory : trajectory ∉ source.adaptiveAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta) (round : Fin rounds) : trajectory round ∉ source.rewardSource.stochasticAllCoordinateEmpiricalModelBadEvent (source.policyAt trajectory round) initialState episodes varianceProxy (multiBatchLocalDelta rounds countDelta) (multiBatchLocalDelta rounds rewardDelta)","missing":[],"search":"policyat_batch_not_mem_allcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.policyat_batch_not_mem_allcoordinateempiricalmodelbadevent outside the adaptive event, each batch avoids its generating policy's event. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selectedExploratoryStochasticAllCoordinateEmpiricalModelBadEvent","label":"measurableSet_selectedExploratoryStochasticAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selectedExploratoryStochasticAllCoordinateEmpiricalModelBadEvent","description":"Selected exploratory count-and-reward events are measurable after pulling the stochastic batch back to its raw sampled `EpisodeBatch`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-03733caaed6a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7685,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:466"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selectedExploratoryStochasticAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} {History : Type*} [MeasurableSpace History] (rewardSource : mdp.MeanCompatibleRewardKernel) (selector : History -> DeterministicMarkovPolicyTable mdp) (hselector : Measurable selector) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (varianceProxy : NNReal) (countDelta rewardDelta : Real) : MeasurableSet {pair : History × StochasticEpisodeBatch mdp episodes | pair.2 ∈ rewardSource.stochasticAllCoordinateEmpiricalModelBadEvent ((selector pair.1).exploratoryPolicy explorationRate hexplorationRate) initialState episodes varianceProxy countDelta rewardDelta}","missing":[],"search":"measurableset_selectedexploratorystochasticallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selectedexploratorystochasticallcoordinateempiricalmodelbadevent selected exploratory count-and-reward events are measurable after pulling the stochastic batch back to its raw sampled `episodebatch`. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","label":"exploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","description":"Every successor event of the concrete sampled optimistic source is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-2ce78cf24ea6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7686,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:521"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) (n : Nat) : let source := exploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate MeasurableSet (source.successorAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta n)","missing":[],"search":"exploratorysource_measurableset_successorallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_measurableset_successorallcoordinateempiricalmodelbadevent every successor event of the concrete sampled optimistic source is measurable. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","label":"exploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","description":"The concrete sampled optimistic source inherits the finite-round selected count-and-reward event without any caller-supplied measurability premise.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-1090ec47dc2e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7687,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:557"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) :…","missing":[],"search":"exploratorysource_trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le the concrete sampled optimistic source inherits the finite-round selected count-and-reward event without any caller-supplied measurability premise. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","description":"Finite-round sampled-reward adaptive optimism with explicit exploratory path support calibration. Every model is built from the actual observed rewards; outside one measurable event, every round is optimistic and its recommended policy satisfies the compiled selected-radius expected-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticconfidence/index.html#decl-3ff6f8b080d1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","order":7688,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence.lean:608"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_explicitcalibration finite-round sampled-reward adaptive optimism with explicit exploratory path support calibration. every model is built from the actual observed rewards; outside one measurable event, every round is optimistic and its recommended policy satisfies the compiled selected-radius expected-regret bound. theorem compiled","shard":"modules/e52a060cdefebe8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.sampledEmpiricalOptimisticPolicyTable_toMarkovPolicy","label":"sampledEmpiricalOptimisticPolicyTable_toMarkovPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.sampledEmpiricalOptimisticPolicyTable_toMarkovPolicy","description":"The sampled optimistic action table represents the sampled plan's policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-0d81ec0bbfb0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7689,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledEmpiricalOptimisticPolicyTable_toMarkovPolicy {mdp : MDP State Action} {episodes : Nat} (batch : StochasticEpisodeBatch mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) : (batch.sampledEmpiricalOptimisticPolicyTable defaultState rewardBudget transitionBudget).toMarkovPolicy = (mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes (mdp.sampledEpisodeBatchOfStochasticTrajectories episodes batch) defaultState rewardBudget transitionBudget).plan.optimisticPolicy","missing":[],"search":"sampledempiricaloptimisticpolicytable_tomarkovpolicy banditrlproof.finitehorizonrl.stochasticepisodebatch.sampledempiricaloptimisticpolicytable_tomarkovpolicy the sampled optimistic action table represents the sampled plan's policy. theorem compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret","label":"adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret","description":"Expected regret of the actual successor exploratory policies selected from the sampled plans in a finite window.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-c0a6b092f655","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7690,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) : Real","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticsuccessorexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsuccessorexploratorybehaviorexpectedregret expected regret of the actual successor exploratory policies selected from the sampled plans in a finite window. definition compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorPolicy_frestrictLe_eq_sampledPlanExploratoryPolicy","label":"exploratorySource_successorPolicy_frestrictLe_eq_sampledPlanExploratoryPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorPolicy_frestrictLe_eq_sampledPlanExploratoryPolicy","description":"The policy selected from sampled coordinate `n` generates successor coordinate `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-8ddabf872c57","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7691,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorPolicy_frestrictLe_eq_sampledPlanExploratoryPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : let source := exploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate source.successorPolicy n (Preorder.frestrictLe n trajectory) = ((trajectory n).sampledEmpiricalOptimisticPolicyTable defaultState rewardBudget transitionBudget).exploratoryPolicy explorationRate hexplorationRate","missing":[],"search":"exploratorysource_successorpolicy_frestrictle_eq_sampledplanexploratorypolicy banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_successorpolicy_frestrictle_eq_sampledplanexploratorypolicy the policy selected from sampled coordinate `n` generates successor coordinate `n + 1`. theorem compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_cumulativeSuccessorPolicyExpectedRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","label":"exploratorySource_cumulativeSuccessorPolicyExpectedRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_cumulativeSuccessorPolicyExpectedRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","description":"The named sampled-plan sum is the actual source successor-policy regret sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-9c2b6c87f4c5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7692,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_cumulativeSuccessorPolicyExpectedRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : let source := exploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate (∑ round : Fin rounds, (source.successorPolicy (round : Nat) (Preorder.frestrictLe (round : Nat) trajectory)).expectedRegret initialState) = adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBe…","missing":[],"search":"exploratorysource_cumulativesuccessorpolicyexpectedregret_eq_sampledplanexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_cumulativesuccessorpolicyexpectedregret_eq_sampledplanexploratorybehaviorexpectedregret the named sampled-plan sum is the actual source successor-policy regret sum. theorem compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret_le","label":"adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret_le","description":"Cumulative successor exploratory-policy regret is recommendation regret plus one explicit exploration charge per selected sampled plan.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-39dae787fef6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7693,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) : adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaultState rewardBudget transitionBudget explorationRate hexplorationRate rounds <= adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret (mdp := mdp) (initialState := initialState) (episodes := episo…","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticsuccessorexploratorybehaviorexpectedregret_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsuccessorexploratorybehaviorexpectedregret_le cumulative successor exploratory-policy regret is recommendation regret plus one explicit exploration charge per selected sampled plan. theorem compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret","label":"adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret","description":"On sampled-model confidence, cumulative successor behavior regret is bounded by the occupancy-radius certificate plus the explicit exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-88239fb84352","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7694,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:167"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hconfidence : (forall round : Fin rounds, forall state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= (adaptiveStochasticSampledEmpiricalOptimisticPlanAt trajectory defaultState rewardBudget transitionBudget round).upperValueRemaining mdp.horizon le_rfl state) ∧ adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret (mdp := mdp) (initialState :=…","missing":[],"search":"adaptivestochasticsampledempiricaloptimistic_optimism_and_cumulativesuccessorexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimistic_optimism_and_cumulativesuccessorexploratorybehaviorexpectedregret on sampled-model confidence, cumulative successor behavior regret is bounded by the occupancy-radius certificate plus the explicit exploration charge. theorem compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_explicitCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_explicitCalibration","description":"Finite-round actual-sampled confidence, optimism, and cumulative expected regret of the selected successor exploratory policies.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-297b24f7f88e/index.html#decl-72a06186b702","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","order":7695,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret.lean:229"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_explicitCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_l…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_cumulativesuccessorexploratorybehaviorexpectedregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_cumulativesuccessorexploratorybehaviorexpectedregret_of_pathsupport_explicitcalibration finite-round actual-sampled confidence, optimism, and cumulative expected regret of the selected successor exploratory policies. theorem compiled","shard":"modules/3d1c93c4686a0400.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt","label":"adaptiveStochasticSampledEmpiricalOptimisticPlanAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt","description":"Empirical optimistic plan built from one actual sampled-reward batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-48200eab5c41/index.html#decl-87911ed50a44","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","order":7696,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveStochasticSampledEmpiricalOptimisticPlanAt {mdp : MDP State Action} {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (round : Nat) : mdp.EstimatedModelPlan","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticplanat banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticplanat empirical optimistic plan built from one actual sampled-reward batch. definition compiled","shard":"modules/9ae36733e5c8ca42.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret","label":"adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret","description":"Sum of the recommended policies' expected regrets over the sampled window.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-48200eab5c41/index.html#decl-5211e7a7f0c2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","order":7697,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (rounds : Nat) : Real","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticrecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticrecommendedexpectedregret sum of the recommended policies' expected regrets over the sampled window. definition compiled","shard":"modules/9ae36733e5c8ca42.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum","label":"adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum","description":"Sum of the occupancy selected-radius bounds over the sampled window.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-48200eab5c41/index.html#decl-a6917d1bf488","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","order":7698,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (rounds : Nat) : Real","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticoccupancyradiussum banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticoccupancyradiussum sum of the occupancy selected-radius bounds over the sampled window. definition compiled","shard":"modules/9ae36733e5c8ca42.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeRecommendedExpectedRegret","label":"adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeRecommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeRecommendedExpectedRegret","description":"Finite summation of pointwise sampled-model optimism and regret bounds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-48200eab5c41/index.html#decl-666853d45d48","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","order":7699,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeRecommendedExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes rounds : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (hround : forall round : Fin rounds, let plan := adaptiveStochasticSampledEmpiricalOptimisticPlanAt trajectory defaultState rewardBudget transitionBudget round (forall state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= plan.upperValueRemaining mdp.horizon le_rfl state) ∧ plan.optimisticPolicy.expectedRegret initialState <= plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState) : (forall round : Fin rounds, forall state, mdp.op…","missing":[],"search":"adaptivestochasticsampledempiricaloptimistic_optimism_and_cumulativerecommendedexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimistic_optimism_and_cumulativerecommendedexpectedregret finite summation of pointwise sampled-model optimism and regret bounds. theorem compiled","shard":"modules/9ae36733e5c8ca42.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeRecommendedExpectedRegret_of_pathSupport_explicitCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeRecommendedExpectedRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeRecommendedExpectedRegret_of_pathSupport_explicitCalibration","description":"Finite-round actual-sampled-model confidence, optimism, and cumulative recommended-policy expected regret under explicit path-support calibration.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticcum-48200eab5c41/index.html#decl-3b95fb7eb7a7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","order":7700,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeRecommendedExpectedRegret_of_pathSupport_explicitCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDel…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_cumulativerecommendedexpectedregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_cumulativerecommendedexpectedregret_of_pathsupport_explicitcalibration finite-round actual-sampled-model confidence, optimism, and cumulative recommended-policy expected regret under explicit path-support calibration. theorem compiled","shard":"modules/9ae36733e5c8ca42.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt_selectedRadiusRemaining","label":"adaptiveStochasticSampledEmpiricalOptimisticPlanAt_selectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt_selectedRadiusRemaining","description":"An actual sampled empirical plan selects its two fixed model budgets.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html#decl-df741638b35a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","order":7701,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimisticPlanAt_selectedRadiusRemaining {mdp : MDP State Action} {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (round remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (adaptiveStochasticSampledEmpiricalOptimisticPlanAt trajectory defaultState rewardBudget transitionBudget round).selectedRadiusRemaining remaining hremaining state = rewardBudget + transitionBudget","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticplanat_selectedradiusremaining banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticplanat_selectedradiusremaining an actual sampled empirical plan selects its two fixed model budgets. theorem compiled","shard":"modules/a75a4e61326820fd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","label":"adaptiveStochasticSampledEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","description":"One sampled plan's selected-radius occupancy term has a closed form.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html#decl-879b0c8c843d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","order":7702,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (round : Nat) : let plan := adaptiveStochasticSampledEmpiricalOptimisticPlanAt trajectory defaultState rewardBudget transitionBudget round plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState = (mdp.horizon : Real) * (2 * (rewardBudget + transitionBudget))","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticplanat_occupancyselectedradiusremaining_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticplanat_occupancyselectedradiusremaining_eq one sampled plan's selected-radius occupancy term has a closed form. theorem compiled","shard":"modules/a75a4e61326820fd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum_eq","label":"adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum_eq","description":"The complete actual-sampled occupancy-radius sum is deterministic.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html#decl-839538e703cc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","order":7703,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) (rounds : Nat) : adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaultState rewardBudget transitionBudget rounds = (rounds : Real) * ((mdp.horizon : Real) * (2 * (rewardBudget + transitionBudget)))","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticoccupancyradiussum_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticoccupancyradiussum_eq the complete actual-sampled occupancy-radius sum is deterministic. theorem compiled","shard":"modules/a75a4e61326820fd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticExplicitBudgetAverageBound","label":"adaptiveStochasticSampledEmpiricalOptimisticExplicitBudgetAverageBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticExplicitBudgetAverageBound","description":"Planning part of the calibrated realized average-regret certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html#decl-95f7bf578134","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","order":7704,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean:93"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveStochasticSampledEmpiricalOptimisticExplicitBudgetAverageBound (mdp : MDP State Action) (explorationRate : NNReal) (rewardBound rewardBudget : Real) : Real","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticexplicitbudgetaveragebound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticexplicitbudgetaveragebound planning part of the calibrated realized average-regret certificate. definition compiled","shard":"modules/a75a4e61326820fd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_explicitBudgetAverageBound","label":"adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_explicitBudgetAverageBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_explicitBudgetAverageBound","description":"Uniform-floor budgets close the averaged occupancy and exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html#decl-e5deb517c3ee","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","order":7705,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean:102"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_explicitBudgetAverageBound {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (rewardBudget rewardBound : Real) (explorationRate : NNReal) (rounds : Nat) (hrounds : 0 < rounds) : (adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaultState rewardBudget (uniformFloorStochasticTransitionBudget rewardBound rewardBudget) rounds + (rounds : Real) * exploratoryBehaviorRegretCharge mdp explorationRate rewardBound) / (rounds : Real) = adaptiveStochasticSampledEmpiricalOptimisticExplicitBudgetAverageBound mdp explorationRate rewardBound rewardBudget","missing":[],"search":"adaptivestochasticsampledempiricaloptimistic_occupancyandchargeaverage_eq_explicitbudgetaveragebound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimistic_occupancyandchargeaverage_eq_explicitbudgetaveragebound uniform-floor budgets close the averaged occupancy and exploration charge. theorem compiled","shard":"modules/a75a4e61326820fd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_explicitBudgetRealizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_explicitBudgetRealizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_explicitBudgetRealizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","description":"Actual sampled-model confidence and globally centered return concentration with a deterministic explicit-budget planning envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticexp-3d57f4149b98/index.html#decl-ab6b2a76ae7e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","order":7706,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret.lean:132"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_explicitBudgetRealizedSuccessorAverageRegret_of_pathSupport_explicitCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (hmodelTotal : 0 < ((((episodes : NNReal) * varianceProxy : NNR…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_explicitbudgetrealizedsuccessoraverageregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_explicitbudgetrealizedsuccessoraverageregret_of_pathsupport_explicitcalibration actual sampled-model confidence and globally centered return concentration with a deterministic explicit-budget planning envelope. theorem compiled","shard":"modules/a75a4e61326820fd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","label":"exploratorySource_successorExpectedCumulativeRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","description":"The generic source cumulative expected regret is the named sampled-plan sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticrea-28c98ba81ea7/index.html#decl-db36229ab602","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","order":7707,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedCumulativeRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate).successorExpectedCumulativeRegret trajectory rounds = adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaul…","missing":[],"search":"exploratorysource_successorexpectedcumulativeregret_eq_sampledplanexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_successorexpectedcumulativeregret_eq_sampledplanexploratorybehaviorexpectedregret the generic source cumulative expected regret is the named sampled-plan sum. theorem compiled","shard":"modules/49b43e9e188d456d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","label":"exploratorySource_successorExpectedAverageRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","description":"Average form of the exact sampled-plan successor-policy identity.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticrea-28c98ba81ea7/index.html#decl-af4ab554775b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","order":7708,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret.lean:75"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_successorExpectedAverageRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : (exploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate).successorExpectedAverageRegret trajectory rounds = adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaultState…","missing":[],"search":"exploratorysource_successorexpectedaverageregret_eq_sampledplanexploratorybehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_successorexpectedaverageregret_eq_sampledplanexploratorybehaviorexpectedregret average form of the exact sampled-plan successor-policy identity. theorem compiled","shard":"modules/49b43e9e188d456d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","description":"Concrete three-share fixed-window endpoint for actual sampled-model planning and realized successor behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticrea-28c98ba81ea7/index.html#decl-f33865c95b3e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","order":7709,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret_of_pathSupport_explicitCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (hmodelTotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real)))…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_realizedsuccessoraverageregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_realizedsuccessoraverageregret_of_pathsupport_explicitcalibration concrete three-share fixed-window endpoint for actual sampled-model planning and realized successor behavior regret. theorem compiled","shard":"modules/49b43e9e188d456d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyAt_expectedRegret_le_rateAt_of_not_mem_modelRoundBadEvent","label":"selfConsistentScheduledCausalSource_successorPolicyAt_expectedRegret_le_rateAt_of_not_mem_modelRoundBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyAt_expectedRegret_le_rateAt_of_not_mem_modelRoundBadEvent","description":"A model-good coordinate bounds the actual successor policy by the causal rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-20fa414d9371/index.html#decl-7ac215b31b57","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","order":7710,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_successorPolicyAt_expectedRegret_le_rateAt_of_not_mem_modelRoundBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBa…","missing":[],"search":"selfconsistentscheduledcausalsource_successorpolicyat_expectedregret_le_rateat_of_not_mem_modelroundbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_successorpolicyat_expectedregret_le_rateat_of_not_mem_modelroundbadevent a model-good coordinate bounds the actual successor policy by the causal rate. theorem compiled","shard":"modules/675fd8e299ce3279.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","label":"selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","description":"Expected regret of the actual exploratory successor policy at one coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-20fa414d9371/index.html#decl-64c6fa0fb71a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","order":7711,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency.lean:152"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (t : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess expected regret of the actual exploratory successor policy at one coordinate. definition compiled","shard":"modules/675fd8e299ce3279.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_nonneg","label":"selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_nonneg","description":"Every actual successor-policy expected-regret coordinate is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-20fa414d9371/index.html#decl-44fa26c442c6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","order":7712,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency.lean:169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (t : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s)) : 0 <= selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_nonneg every actual successor-policy expected-regret coordinate is nonnegative. theorem compiled","shard":"modules/675fd8e299ce3279.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoAlmostEverywhere_zero","description":"Actual successor-policy expected regret converges to zero almost surely.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-20fa414d9371/index.html#decl-b79b204586d9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","order":7713,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianc…","missing":[],"search":"selfconsistentscheduledcausalsource_successorpolicyexpectedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_successorpolicyexpectedregret_tendstoalmosteverywhere_zero actual successor-policy expected regret converges to zero almost surely. theorem compiled","shard":"modules/675fd8e299ce3279.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_behaviorExpected_and_realizedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_behaviorExpected_and_realizedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_behaviorExpected_and_realizedRegret_tendstoAlmostEverywhere_zero","description":"Eventual optimism and expected/realized behavior consistency hold jointly a.e.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-20fa414d9371/index.html#decl-f6475b1fac11","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","order":7714,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency.lean:257"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_behaviorExpected_and_realizedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let episodes := fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistent…","missing":[],"search":"selfconsistentscheduledcausalsource_eventually_modeloptimistic_and_behaviorexpected_and_realizedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_eventually_modeloptimistic_and_behaviorexpected_and_realizedregret_tendstoalmosteverywhere_zero eventual optimism and expected/realized behavior consistency hold jointly a.e. theorem compiled","shard":"modules/675fd8e299ce3279.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.summable_exp_neg_sqrt_natCast","label":"summable_exp_neg_sqrt_natCast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.summable_exp_neg_sqrt_natCast","description":"The stretched-exponential sequence `exp (-sqrt n)` is summable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-0458ff7023c6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7715,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:20"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_exp_neg_sqrt_natCast : Summable (fun n : Nat => Real.exp (-Real.sqrt (n : Real)))","missing":[],"search":"summable_exp_neg_sqrt_natcast banditrlproof.summable_exp_neg_sqrt_natcast the stretched-exponential sequence `exp (-sqrt n)` is summable. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.natCast_le_successorEpisodeMass","label":"natCast_le_successorEpisodeMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.natCast_le_successorEpisodeMass","description":"Positive episode batches make successor mass dominate the round index.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-3a5a7c32e2ae","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7716,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem natCast_le_successorEpisodeMass (episodes : Nat -> Nat) (hepisodes : forall t, 0 < episodes t) (rounds : Nat) : (rounds : Real) <= successorEpisodeMass episodes rounds","missing":[],"search":"natcast_le_successorepisodemass banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.natcast_le_successorepisodemass positive episode batches make successor mass dominate the round index. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.summable_exp_neg_sqrt_successorEpisodeMass","label":"summable_exp_neg_sqrt_successorEpisodeMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.summable_exp_neg_sqrt_successorEpisodeMass","description":"A mass-adapted `exp (-sqrt mass)` schedule is summable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-fa0ac1750301","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7717,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_exp_neg_sqrt_successorEpisodeMass (episodes : Nat -> Nat) (hepisodes : forall t, 0 < episodes t) : Summable (fun rounds => Real.exp (-Real.sqrt (successorEpisodeMass episodes rounds)))","missing":[],"search":"summable_exp_neg_sqrt_successorepisodemass banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.summable_exp_neg_sqrt_successorepisodemass a mass-adapted `exp (-sqrt mass)` schedule is summable. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledCausalVanishingReturnDelta","label":"summable_selfConsistentScheduledCausalVanishingReturnDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledCausalVanishingReturnDelta","description":"The natural-causal mass-adapted return failure schedule is summable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-84b2ee255992","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7718,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:116"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_selfConsistentScheduledCausalVanishingReturnDelta (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Summable (selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor)","missing":[],"search":"summable_selfconsistentscheduledcausalvanishingreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_selfconsistentscheduledcausalvanishingreturndelta the natural-causal mass-adapted return failure schedule is summable. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalVanishingReturnFailureBudget_ne_top","label":"tsum_selfConsistentScheduledCausalVanishingReturnFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalVanishingReturnFailureBudget_ne_top","description":"The return-failure ENNReal budget is finite.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-4ff85c5fc056","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7719,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:135"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_selfConsistentScheduledCausalVanishingReturnFailureBudget_ne_top (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : (∑' rounds, ENNReal.ofReal (selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor rounds)) ≠ ∞","missing":[],"search":"tsum_selfconsistentscheduledcausalvanishingreturnfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_selfconsistentscheduledcausalvanishingreturnfailurebudget_ne_top the return-failure ennreal budget is finite. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnBadEvent","label":"selfConsistentScheduledCausalVanishingReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnBadEvent","description":"The actual return-deviation event charged by the summable mass schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-28e879be7473","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7720,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalVanishingReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentscheduledcausalvanishingreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturnbadevent the actual return-deviation event charged by the summable mass schedule. definition compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalVanishingReturnBadEvent","label":"measurableSet_selfConsistentScheduledCausalVanishingReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalVanishingReturnBadEvent","description":"Each event in the summable return schedule is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-d3bd0162ba18","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7721,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledCausalVanishingReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : MeasurableSet (selfConsistentScheduledCausalVanishingReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurableset_selfconsistentscheduledcausalvanishingreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentscheduledcausalvanishingreturnbadevent each event in the summable return schedule is measurable. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_vanishingReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_vanishingReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_vanishingReturnBadEvent_le","description":"Every return event is bounded by its mass-adapted failure share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-261420b3c86b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7722,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:186"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_vanishingReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (rounds : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure (selfConsistentScheduledCausalVanishingReturnBadEvent mdp initialState rewardSource initialTable defaultState varian…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_vanishingreturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_vanishingreturnbadevent_le every return event is bounded by its mass-adapted failure share. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalModelRoundBadEvent_measure_ne_top","label":"tsum_selfConsistentScheduledCausalModelRoundBadEvent_measure_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalModelRoundBadEvent_measure_ne_top","description":"The actual coordinate-model event measures have finite total mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-9efe3fcefa7f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7723,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:255"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_selfConsistentScheduledCausalModelRoundBadEvent_measure_ne_top (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (∑' t, source.trajectoryMeasure (selfConsistentScheduledCausalModelRoundBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t)) ≠ ∞","missing":[],"search":"tsum_selfconsistentscheduledcausalmodelroundbadevent_measure_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_selfconsistentscheduledcausalmodelroundbadevent_measure_ne_top the actual coordinate-model event measures have finite total mass. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalVanishingReturnBadEvent_measure_ne_top","label":"tsum_selfConsistentScheduledCausalVanishingReturnBadEvent_measure_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalVanishingReturnBadEvent_measure_ne_top","description":"The actual return-event measures have finite total mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-5d1b4c8fa52d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7724,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_selfConsistentScheduledCausalVanishingReturnBadEvent_measure_ne_top (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (∑' rounds, source.trajectoryMeasure (selfConsistentScheduledCausalVanishingReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy bas…","missing":[],"search":"tsum_selfconsistentscheduledcausalvanishingreturnbadevent_measure_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_selfconsistentscheduledcausalvanishingreturnbadevent_measure_ne_top the actual return-event measures have finite total mass. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_selfConsistentScheduledCausalModel_and_returnBadEvents","label":"ae_eventually_not_mem_selfConsistentScheduledCausalModel_and_returnBadEvents","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_selfConsistentScheduledCausalModel_and_returnBadEvents","description":"Almost every natural-causal trajectory is eventually model-good and return-good.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-e3f190f461f2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7725,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_not_mem_selfConsistentScheduledCausalModel_and_returnBadEvents (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor ∀ᵐ trajectory ∂source.trajectoryMeasure, (∀ᶠ t in atTop, trajectory ∉ selfConsistentScheduledCausalModelRoundBadEvent mdp initialState rewardSource initialTable…","missing":[],"search":"ae_eventually_not_mem_selfconsistentscheduledcausalmodel_and_returnbadevents banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_not_mem_selfconsistentscheduledcausalmodel_and_returnbadevents almost every natural-causal trajectory is eventually model-good and return-good. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_abs_selfConsistentScheduledNaturalCausalRealizedRegret_le_envelope","label":"ae_eventually_abs_selfConsistentScheduledNaturalCausalRealizedRegret_le_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_abs_selfConsistentScheduledNaturalCausalRealizedRegret_le_envelope","description":"Almost every trajectory eventually obeys one fixed-burn-in regret envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-e882fc133cbe","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7726,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:356"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_abs_selfConsistentScheduledNaturalCausalRealizedRegret_le_envelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVis…","missing":[],"search":"ae_eventually_abs_selfconsistentschedulednaturalcausalrealizedregret_le_envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_abs_selfconsistentschedulednaturalcausalrealizedregret_le_envelope almost every trajectory eventually obeys one fixed-burn-in regret envelope. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_realizedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_realizedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_realizedRegret_tendstoAlmostEverywhere_zero","description":"Natural-causal realized successor-average regret converges almost surely.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-14b5db8afbb1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7727,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:431"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_realizedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisi…","missing":[],"search":"selfconsistentscheduledcausalsource_realizedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_realizedregret_tendstoalmosteverywhere_zero natural-causal realized successor-average regret converges almost surely. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_realizedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_realizedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_realizedRegret_tendstoAlmostEverywhere_zero","description":"Almost surely, late empirical models are optimistic while realized regret vanishes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-872db4cc4093/index.html#decl-2378be9d9799","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","order":7728,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency.lean:492"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_realizedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let episodes := fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp…","missing":[],"search":"selfconsistentscheduledcausalsource_eventually_modeloptimistic_and_realizedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_eventually_modeloptimistic_and_realizedregret_tendstoalmosteverywhere_zero almost surely, late empirical models are optimistic while realized regret vanishes. theorem compiled","shard":"modules/14c1bcee766d279d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.integral_le_add_const_mul_measureReal_of_le_on_compl","label":"integral_le_add_const_mul_measureReal_of_le_on_compl","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.integral_le_add_const_mul_measureReal_of_le_on_compl","description":"Integrate a local bound off one bad event and a global bound on it.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-cb3971df576f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7729,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_le_add_const_mul_measureReal_of_le_on_compl {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (f : Omega -> Real) (bad : Set Omega) (hbad : MeasurableSet bad) (hf : Integrable f mu) (rate envelope : Real) (hrate : 0 <= rate) (hglobal : forall omega, f omega <= envelope) (hgood : forall omega, omega ∉ bad -> f omega <= rate) : integral mu f <= rate + envelope * (mu bad).toReal","missing":[],"search":"integral_le_add_const_mul_measurereal_of_le_on_compl banditrlproof.finitehorizonrl.integral_le_add_const_mul_measurereal_of_le_on_compl integrate a local bound off one bad event and a global bound on it. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt_nonneg","label":"selfConsistentScheduledCausalPlanningRateAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt_nonneg","description":"The named causal planning envelope is pointwise nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-2c4ac602e1b3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7730,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:84"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalPlanningRateAt_nonneg (mdp : MDP State Action) (t : Nat) : 0 <= selfConsistentScheduledCausalPlanningRateAt mdp t","missing":[],"search":"selfconsistentscheduledcausalplanningrateat_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalplanningrateat_nonneg the named causal planning envelope is pointwise nonnegative. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalCoordinateModelFailureBudget_toReal_eq","label":"selfConsistentScheduledCausalCoordinateModelFailureBudget_toReal_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalCoordinateModelFailureBudget_toReal_eq","description":"The two coordinate model-confidence shares have real mass `2 * delta_t`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-4c524b803a5a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7731,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalCoordinateModelFailureBudget_toReal_eq (mdp : MDP State Action) (t : Nat) : (selfConsistentScheduledCausalCoordinateModelFailureBudget mdp t).toReal = 2 * AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp t","missing":[],"search":"selfconsistentscheduledcausalcoordinatemodelfailurebudget_toreal_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalcoordinatemodelfailurebudget_toreal_eq the two coordinate model-confidence shares have real mass `2 * delta_t`. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt","label":"selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt","description":"Planning rate plus the one-event `2H` expectation fallback.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-1f0f8a89e68f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7732,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt (mdp : MDP State Action) (t : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat planning rate plus the one-event `2h` expectation fallback. definition compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq","label":"selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq","description":"Closed form of the finite-coordinate integrated behavior envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-c01c84e082c5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7733,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq (mdp : MDP State Action) (t : Nat) : selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp t = (mdp.horizon : Real) * (2 * (1 / (AdaptiveEpisodeBatchSource.decayingExplorationScale t : Real) ^ 2 + (12 * (Fintype.card State : Real) * (mdp.horizon : Real)) / (AdaptiveEpisodeBatchSource.decayingExplorationScale t : Real) ^ 2)) + exploratoryBehaviorRegretCharge mdp (AdaptiveEpisodeBatchSource.decayingExplorationRate (t + 1)) 1 + (4 * (mdp.horizon : Real)) / (((t + 2 : Nat) : Real) ^ (mdp.horizon + 5))","missing":[],"search":"selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_eq closed form of the finite-coordinate integrated behavior envelope. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_nonneg","label":"selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_nonneg","description":"The explicit integrated behavior envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-3b4bc7a372ec","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7734,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:146"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_nonneg (mdp : MDP State Action) (t : Nat) : 0 <= selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp t","missing":[],"search":"selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_nonneg the explicit integrated behavior envelope is nonnegative. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_le_rateAt","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_le_rateAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_le_rateAt","description":"Expected absolute behavior regret obeys the explicit finite-coordinate rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-25d6f9424c00","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7735,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:156"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_le_rateAt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (t : Nat) : selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret mdp initialState rewardSource initialTable defaultStat…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret_le_rateat banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret_le_rateat expected absolute behavior regret obeys the explicit finite-coordinate rate. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_tendsto_zero","label":"selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_tendsto_zero","description":"The explicit integrated behavior envelope vanishes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-0e9a47f69dfa","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7736,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:249"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_tendsto_zero the explicit integrated behavior envelope vanishes. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero_of_explicit_rate","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero_of_explicit_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero_of_explicit_rate","description":"Quantitative squeeze proof of expected-absolute convergence.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-fa5fe3581630","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7737,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:264"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero_of_explicit_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret mdp initialState rewardSource initi…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret_tendsto_zero_of_explicit_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret_tendsto_zero_of_explicit_rate quantitative squeeze proof of expected-absolute convergence. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpectedRegret_explicitIntegratedRate","label":"selfConsistentScheduledCausalSource_behaviorExpectedRegret_explicitIntegratedRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpectedRegret_explicitIntegratedRate","description":"The finite-coordinate rate, its limit, and the induced expectation limit.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-550a9c120089/index.html#decl-47baadb459b9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","order":7738,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate.lean:295"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_behaviorExpectedRegret_explicitIntegratedRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall t, selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret mdp initialState rewardSource initialTable defau…","missing":[],"search":"selfconsistentscheduledcausalsource_behaviorexpectedregret_explicitintegratedrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_behaviorexpectedregret_explicitintegratedrate the finite-coordinate rate, its limit, and the induced expectation limit. theorem compiled","shard":"modules/c4ecfc2d77946946.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","description":"Natural-prefix cumulative expected regret of the actual successor behavior.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-df01d09a9f1e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7739,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat)","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess natural-prefix cumulative expected regret of the actual successor behavior. definition compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret","description":"Expectation of the natural-prefix cumulative behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-79ce14d0ea8d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7740,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret expectation of the natural-prefix cumulative behavior regret. definition compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate","label":"selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate","description":"Finite-prefix sum of the explicit integrated coordinate rates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-cf4382017f41","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7741,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate finite-prefix sum of the explicit integrated coordinate rates. definition compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret","description":"Natural-prefix average behavior expected regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-310962b0b312","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7742,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret natural-prefix average behavior expected regret. definition compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate","label":"selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate","description":"Cesaro average of the explicit integrated coordinate rates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-37d89bac40c1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7743,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate cesaro average of the explicit integrated coordinate rates. definition compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_nonneg","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_nonneg","description":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory) : 0 <= selfConsistentScheduledNatura…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-60f020646085","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7744,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory) : 0 <= selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_nonneg theorem selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_nonneg (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) (trajectory) : 0 <= selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds trajectory theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","label":"integrable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","description":"theorem integrable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state actio…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-068822e59299","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7745,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : Integrable (selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess theorem integrable_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (rounds : nat) : integrable (selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds) (selfconsistentscheduledcausalsource mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor).trajectorymeasure theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_eq_sum_expectedAbsolute","label":"selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_eq_sum_expectedAbsolute","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_eq_sum_expectedAbsolute","description":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_eq_sum_expectedAbsolute (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-cc0ce364303c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7746,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_eq_sum_expectedAbsolute (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds = (Finset.range rounds).sum fun t => selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_eq_sum_expectedabsolute banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_eq_sum_expectedabsolute theorem selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_eq_sum_expectedabsolute (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (rounds : nat) : selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds = (finset.range rounds).sum fun t => selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor t theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_nonneg","description":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedCumul…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-0f1283d1607f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7747,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:178"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_nonneg theorem selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_nonneg (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) : 0 <= selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","label":"selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","description":"theorem selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_nonneg (mdp : MDP State Action) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-ce7e82c678bb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7748,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:197"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_nonneg (mdp : MDP State Action) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_nonneg theorem selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_nonneg (mdp : mdp state action) (rounds : nat) : 0 <= selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_rate","label":"selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_rate","description":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-3a9bcbf4da85","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7749,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:206"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret mdp initialState rewardSource initialTable defa…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_le_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_le_rate theorem selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_le_rate (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (rounds : nat) : selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds <= selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_rate","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_rate","description":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : D…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-8b6d63cb6942","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7750,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:236"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret mdp initialState rewardSource initialTable defaultSta…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_le_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_le_rate theorem selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_le_rate (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (rounds : nat) : selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds <= selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_eq_natWeightedAverage","label":"selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_eq_natWeightedAverage","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_eq_natWeightedAverage","description":"theorem selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_eq_natWeightedAverage (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate mdp rounds = natWeightedAverage (fun _ => 1) (selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp) rounds","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-b28f8cad80c1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7751,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:268"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_eq_natWeightedAverage (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate mdp rounds = natWeightedAverage (fun _ => 1) (selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp) rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate_eq_natweightedaverage banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate_eq_natweightedaverage theorem selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate_eq_natweightedaverage (mdp : mdp state action) (rounds : nat) : selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate mdp rounds = natweightedaverage (fun _ => 1) (selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat mdp) rounds theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","label":"selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","description":"theorem selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate mdp) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-0440406cce87","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7752,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate_tendsto_zero theorem selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate_tendsto_zero (mdp : mdp state action) : tendsto (selfconsistentschedulednaturalcausalaverageintegratedbehaviorexpectedregretrate mdp) attop (nhds 0) theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_nonneg","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_nonneg","description":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalAverageBehaviorE…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-cca4fb21c94d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7753,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:300"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_nonneg theorem selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_nonneg (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) : 0 <= selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero","description":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTabl…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-b75793519549","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7754,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:317"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret mdp initialState rewardSource initialTable defaultStat…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_tendsto_zero theorem selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor) attop (nhds 0) theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitRate","label":"selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitRate","description":"Same-source finite-prefix cumulative and Cesaro-average behavior-regret route.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-249451db3f9e/index.html#decl-e2f1392fb1f0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","order":7755,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState variance…","missing":[],"search":"selfconsistentscheduledcausalsource_cumulative_and_averagebehaviorexpectedregret_explicitrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_cumulative_and_averagebehaviorexpectedregret_explicitrate same-source finite-prefix cumulative and cesaro-average behavior-regret route. theorem compiled","shard":"modules/248397c259af1e3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.measurable_exploratoryPolicy_expectedRegret_comp","label":"measurable_exploratoryPolicy_expectedRegret_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.measurable_exploratoryPolicy_expectedRegret_comp","description":"Expected regret after a measurable finite table selection is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-9d6602db783e/index.html#decl-d79967acf59c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","order":7756,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_exploratoryPolicy_expectedRegret_comp {Omega : Type w} [MeasurableSpace Omega] {mdp : MDP State Action} (initialState : Measure State) [IsProbabilityMeasure initialState] (selector : Omega -> DeterministicMarkovPolicyTable mdp) (hselector : Measurable selector) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : Measurable fun omega => ((selector omega).exploratoryPolicy explorationRate hexplorationRate).expectedRegret initialState","missing":[],"search":"measurable_exploratorypolicy_expectedregret_comp banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.measurable_exploratorypolicy_expectedregret_comp expected regret after a measurable finite table selection is measurable. theorem compiled","shard":"modules/f62a2dca35fd9200.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","description":"Every actual causal successor-policy expected-regret coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-9d6602db783e/index.html#decl-7fc17436b237","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","order":7757,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency.lean:81"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (t : Nat) : Measurable (selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess every actual causal successor-policy expected-regret coordinate is measurable. theorem compiled","shard":"modules/f62a2dca35fd9200.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoInMeasure_zero","label":"selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoInMeasure_zero","description":"Actual successor-policy expected regret converges in measure to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-9d6602db783e/index.html#decl-8ef7098148be","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","order":7758,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency.lean:121"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy…","missing":[],"search":"selfconsistentscheduledcausalsource_successorpolicyexpectedregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_successorpolicyexpectedregret_tendstoinmeasure_zero actual successor-policy expected regret converges in measure to zero. theorem compiled","shard":"modules/f62a2dca35fd9200.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_tendstoInMeasure_zero","label":"selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_tendstoInMeasure_zero","description":"Actual behavior expected and realized regret both converge in measure.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-9d6602db783e/index.html#decl-441cbfc062ba","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","order":7759,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency.lean:166"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState variance…","missing":[],"search":"selfconsistentscheduledcausalsource_behaviorexpected_and_realizedregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_behaviorexpected_and_realizedregret_tendstoinmeasure_zero actual behavior expected and realized regret both converge in measure. theorem compiled","shard":"modules/f62a2dca35fd9200.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_le_two_mul_horizon","label":"selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_le_two_mul_horizon","description":"The actual successor behavior's expected regret has the global `2H` envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-75075e5a4264","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7760,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_le_two_mul_horizon (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (t : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s)) : selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t trajectory <= 2 * (mdp.horizon : Real)","missing":[],"search":"selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_le_two_mul_horizon banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_le_two_mul_horizon the actual successor behavior's expected regret has the global `2h` envelope. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","label":"integrable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","description":"Every behavior expected-regret coordinate is integrable on the causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-9cfb5bfd929c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7761,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (t : Nat) : Integrable (selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t) (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess every behavior expected-regret coordinate is integrable on the causal source. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret","description":"Expected absolute behavior regret at one natural-causal coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-734f05d85dea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7762,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (t : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret expected absolute behavior regret at one natural-causal coordinate. definition compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero","description":"Expected absolute behavior regret converges to zero on the same causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-8a781bfe1f05","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7763,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret mdp initialState rewardSource initialTable defaultSt…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutebehaviorregret_tendsto_zero expected absolute behavior regret converges to zero on the same causal source. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","label":"memLp_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","description":"Every coordinate of the behavior expected-regret process belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-bca36a5c8012","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7764,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:193"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (t : Nat) : MemLp (selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t) 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess every coordinate of the behavior expected-regret process belongs to `l1`. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_eq","description":"The exponent-one extended norm is the lifted expected absolute behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-d52a5f09aef8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7765,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:216"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (t : Nat) : eLpNorm (selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t) 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure = ENNReal.ofReal (selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret mdp initialState rewardSource initialTable defaultState vari…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_eq the exponent-one extended norm is the lifted expected absolute behavior regret. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_tendsto_zero","description":"The exponent-one extended norm of the behavior process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-9734928013e2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7766,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:244"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun t => eLpNorm (selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess mdp initia…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_tendsto_zero the exponent-one extended norm of the behavior process tends to zero. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_sub_zero_tendsto_zero","description":"Canonical exponent-one norm-of-the-difference convergence.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-3034e93d546e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7767,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:279"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun t => eLpNorm (selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess m…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalsuccessorpolicyexpectedregretprocess_sub_zero_tendsto_zero canonical exponent-one norm-of-the-difference convergence. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp","label":"selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp","description":"The behavior expected-regret process as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-4b205b5c1a63","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7768,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:314"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (t : Nat) : Lp Real 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"selfconsistentschedulednaturalcausalbehaviorexpectedregretlp banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalbehaviorexpectedregretlp the behavior expected-regret process as an `lp real 1` value. definition compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_coeFn_ae_eq","label":"selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_coeFn_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_coeFn_ae_eq","description":"The named `Lp` coordinate represents the behavior expected-regret process a.e.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-c68ba7d3ad9f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7769,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:334"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_coeFn_ae_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (t : Nat) : (selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor hrewardBound t : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s) -> Real) =ᵐ[ (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajecto…","missing":[],"search":"selfconsistentschedulednaturalcausalbehaviorexpectedregretlp_coefn_ae_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalbehaviorexpectedregretlp_coefn_ae_eq the named `lp` coordinate represents the behavior expected-regret process a.e. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_tendsto_zero","label":"selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_tendsto_zero","description":"The named behavior expected-regret `Lp Real 1` process converges to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-ff61168f1c43","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7770,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:361"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp mdp initialState rewardSource initialTable defaultState variance…","missing":[],"search":"selfconsistentschedulednaturalcausalbehaviorexpectedregretlp_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalbehaviorexpectedregretlp_tendsto_zero the named behavior expected-regret `lp real 1` process converges to zero. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpectedRegret_memLp_eLpNorm_L1_tendsto_zero","label":"selfConsistentScheduledCausalSource_behaviorExpectedRegret_memLp_eLpNorm_L1_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpectedRegret_memLp_eLpNorm_L1_tendsto_zero","description":"Full behavior expected-regret `L1` terminal on the genuine causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-15e23bee38d1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7771,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:408"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_behaviorExpectedRegret_memLp_eLpNorm_L1_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy…","missing":[],"search":"selfconsistentscheduledcausalsource_behaviorexpectedregret_memlp_elpnorm_l1_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_behaviorexpectedregret_memlp_elpnorm_l1_tendsto_zero full behavior expected-regret `l1` terminal on the genuine causal source. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_L1_tendsto_zero","label":"selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_L1_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_L1_tendsto_zero","description":"Actual behavior expected and realized regret converge jointly in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-f5e78083b78d/index.html#decl-c40d18ed282e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","order":7772,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency.lean:479"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_L1_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy…","missing":[],"search":"selfconsistentscheduledcausalsource_behaviorexpected_and_realizedregret_l1_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_behaviorexpected_and_realizedregret_l1_tendsto_zero actual behavior expected and realized regret converge jointly in `l1`. theorem compiled","shard":"modules/3121aa04f88d38e9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","label":"measurable_realizedSuccessorCumulativeRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","description":"Realized successor cumulative regret is measurable on the causal space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-f07eaf0fc0a5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7773,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.realizedSuccessorCumulativeRegret trajectory rounds)","missing":[],"search":"measurable_realizedsuccessorcumulativeregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_realizedsuccessorcumulativeregret realized successor cumulative regret is measurable on the causal space. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","label":"measurable_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","description":"Realized successor average regret is measurable on the causal space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-9e4e3f384c05","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7774,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.realizedSuccessorAverageRegret trajectory rounds)","missing":[],"search":"measurable_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_realizedsuccessoraverageregret realized successor average regret is measurable on the causal space. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedCumulativeRegret_nonneg","label":"successorWeightedExpectedCumulativeRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedCumulativeRegret_nonneg","description":"A positive-weight sum of selected-policy expected regrets is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-66a88c5c7199","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7775,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorWeightedExpectedCumulativeRegret_nonneg {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : 0 <= source.successorWeightedExpectedCumulativeRegret trajectory rounds","missing":[],"search":"successorweightedexpectedcumulativeregret_nonneg banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorweightedexpectedcumulativeregret_nonneg a positive-weight sum of selected-policy expected regrets is nonnegative. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret_nonneg","label":"successorWeightedExpectedAverageRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret_nonneg","description":"The positive-weight successor expected average regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-3e4f0cf34ab5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7776,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorWeightedExpectedAverageRegret_nonneg {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : 0 <= source.successorWeightedExpectedAverageRegret trajectory rounds","missing":[],"search":"successorweightedexpectedaverageregret_nonneg banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorweightedexpectedaverageregret_nonneg the positive-weight successor expected average regret is nonnegative. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","label":"abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","description":"An expected-regret upper bound and a two-sided global return bound control the absolute realized regret under the exact heterogeneous successor mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-5b094edc8373","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7777,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hmass : 0 < successorEpisodeMass episodes rounds) (expectedBound deviationBound : Real) (hexpected : source.successorWeightedExpectedAverageRegret trajectory rounds <= expectedBound) (hdeviation : |source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory| <= deviationBound) : |source.realizedSuccessorAverageRegret trajectory rounds| <= expectedBound + deviationBound / successorEpisodeMass episodes rounds","missing":[],"search":"abs_realizedsuccessoraverageregret_le_of_expected_le_of_deviation_abs_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.abs_realizedsuccessoraverageregret_le_of_expected_le_of_deviation_abs_le an expected-regret upper bound and a two-sided global return bound control the absolute realized regret under the exact heterogeneous successor mass. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_eq_inv_pow","label":"selfConsistentScheduledLocalDelta_eq_inv_pow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_eq_inv_pow","description":"The local count and reward share is an explicit shifted p-series term.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-8a663427f0de","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7778,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledLocalDelta_eq_inv_pow (mdp : MDP State Action) (t : Nat) : AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp t = 1 / (((t + 2 : Nat) : Real) ^ (mdp.horizon + 5))","missing":[],"search":"selfconsistentscheduledlocaldelta_eq_inv_pow banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledlocaldelta_eq_inv_pow the local count and reward share is an explicit shifted p-series term. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledLocalDelta","label":"summable_selfConsistentScheduledLocalDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledLocalDelta","description":"Coordinatewise causal model confidence shares are summable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-5826f18b84d1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7779,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:181"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_selfConsistentScheduledLocalDelta (mdp : MDP State Action) : Summable (AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp)","missing":[],"search":"summable_selfconsistentscheduledlocaldelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_selfconsistentscheduledlocaldelta coordinatewise causal model confidence shares are summable. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalCoordinateModelFailureBudget","label":"selfConsistentScheduledCausalCoordinateModelFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalCoordinateModelFailureBudget","description":"Two model-confidence shares are charged at every causal coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-24b9e67ad6e6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7780,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:199"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalCoordinateModelFailureBudget (mdp : MDP State Action) (t : Nat) : ENNReal","missing":[],"search":"selfconsistentscheduledcausalcoordinatemodelfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalcoordinatemodelfailurebudget two model-confidence shares are charged at every causal coordinate. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget","label":"selfConsistentScheduledCausalTailModelFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget","description":"Infinite model-confidence budget after deleting a finite burn-in prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-aca005b4e67e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7781,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalTailModelFailureBudget (mdp : MDP State Action) (burnin : Nat) : ENNReal","missing":[],"search":"selfconsistentscheduledcausaltailmodelfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelfailurebudget infinite model-confidence budget after deleting a finite burn-in prefix. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalCoordinateModelFailureBudget_ne_top","label":"tsum_selfConsistentScheduledCausalCoordinateModelFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalCoordinateModelFailureBudget_ne_top","description":"The full coordinatewise model-confidence budget is finite.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-afd707d61217","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7782,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:218"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_selfConsistentScheduledCausalCoordinateModelFailureBudget_ne_top (mdp : MDP State Action) : ∑' t, selfConsistentScheduledCausalCoordinateModelFailureBudget mdp t ≠ ∞","missing":[],"search":"tsum_selfconsistentscheduledcausalcoordinatemodelfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_selfconsistentscheduledcausalcoordinatemodelfailurebudget_ne_top the full coordinatewise model-confidence budget is finite. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_tendsto_zero","label":"selfConsistentScheduledCausalTailModelFailureBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_tendsto_zero","description":"Deleting a growing finite prefix makes the infinite model tail vanish.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-08fd65a29d3d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7783,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:242"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalTailModelFailureBudget_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledCausalTailModelFailureBudget mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausaltailmodelfailurebudget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelfailurebudget_tendsto_zero deleting a growing finite prefix makes the infinite model tail vanish. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelRoundBadEvent","label":"selfConsistentScheduledCausalModelRoundBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelRoundBadEvent","description":"The actual sampled-model event at one coordinate of the causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-3a6606540095","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7784,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:252"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalModelRoundBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (t : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s))","missing":[],"search":"selfconsistentscheduledcausalmodelroundbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelroundbadevent the actual sampled-model event at one coordinate of the causal source. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelBadEvent","label":"selfConsistentScheduledCausalTailModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelBadEvent","description":"Union of all actual sampled-model failures after a finite burn-in.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-a119e4927b1b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7785,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:274"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalTailModelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s))","missing":[],"search":"selfconsistentscheduledcausaltailmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelbadevent union of all actual sampled-model failures after a finite burn-in. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalModelRoundBadEvent","label":"measurableSet_selfConsistentScheduledCausalModelRoundBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalModelRoundBadEvent","description":"Every concrete causal coordinate model event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-2ccb99526e71","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7786,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:290"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledCausalModelRoundBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (t : Nat) : MeasurableSet (selfConsistentScheduledCausalModelRoundBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t)","missing":[],"search":"measurableset_selfconsistentscheduledcausalmodelroundbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentscheduledcausalmodelroundbadevent every concrete causal coordinate model event is measurable. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalTailModelBadEvent","label":"measurableSet_selfConsistentScheduledCausalTailModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalTailModelBadEvent","description":"The infinite tail event is measurable by countable union.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-c189dcd3bb5a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7787,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledCausalTailModelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin : Nat) : MeasurableSet (selfConsistentScheduledCausalTailModelBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor burnin)","missing":[],"search":"measurableset_selfconsistentscheduledcausaltailmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentscheduledcausaltailmodelbadevent the infinite tail event is measurable by countable union. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_modelRoundBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_modelRoundBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_modelRoundBadEvent_le","description":"One causal coordinate receives its exact pair of local confidence shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-3bde798768e9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7788,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_modelRoundBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (t : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure (selfConsistentScheduledCausalModelRoundBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t) <= selfConsistentScheduledCausalCoordinateModelFailureBudget mdp t","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_modelroundbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_modelroundbadevent_le one causal coordinate receives its exact pair of local confidence shares. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_tailModelBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_tailModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_tailModelBadEvent_le","description":"The actual infinite tail event is controlled by the summable tail budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-e5aece55cff8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7789,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:433"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_tailModelBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (burnin : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure (selfConsistentScheduledCausalTailModelBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor burnin) <= selfConsistentScheduledCausalTailModelFailureBudget mdp burnin","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_tailmodelbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_tailmodelbadevent_le the actual infinite tail event is controlled by the summable tail budget. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_tail","label":"not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_tail","description":"Outside the tail event, every coordinate after burn-in is model-good.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-f6b5636f5ac3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7790,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:473"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_tail (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin t : Nat) (hburnin : burnin <= t) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s)) (htrajectory : trajectory ∉ selfConsistentScheduledCausalTailModelBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor burnin) : trajectory ∉ selfConsistentScheduledCausalModelRoundBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t","missing":[],"search":"not_mem_selfconsistentscheduledcausalmodelroundbadevent_of_not_mem_tail banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.not_mem_selfconsistentscheduledcausalmodelroundbadevent_of_not_mem_tail outside the tail event, every coordinate after burn-in is model-good. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_coordinateConfidence_of_not_mem_modelRoundBadEvent","label":"selfConsistentScheduledCausalSource_coordinateConfidence_of_not_mem_modelRoundBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_coordinateConfidence_of_not_mem_modelRoundBadEvent","description":"One model-good causal coordinate yields optimism and recommended regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-cd3a83d9fed6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7791,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:502"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_coordinateConfidence_of_not_mem_modelRoundBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsiste…","missing":[],"search":"selfconsistentscheduledcausalsource_coordinateconfidence_of_not_mem_modelroundbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_coordinateconfidence_of_not_mem_modelroundbadevent one model-good causal coordinate yields optimism and recommended regret. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalLocalPlanningBound_le_rateAt","label":"selfConsistentScheduledCausalLocalPlanningBound_le_rateAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalLocalPlanningBound_le_rateAt","description":"The exact local planning budget is bounded by the named vanishing rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-83e3a47cc87e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7792,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:670"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalLocalPlanningBound_le_rateAt (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (t : Nat) : let rewardBudget := AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget mdp varianceProxy baseVisitFloor let transitionBudget := AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget mdp varianceProxy baseVisitFloor (mdp.horizon : Real) * (2 * (rewardBudget t + transitionBudget t)) + exploratoryBehaviorRegretCharge mdp (AdaptiveEpisodeBatchSource.decayingExplorationRate (t + 1)) 1 <= selfConsistentScheduledCausalPlanningRateAt mdp t","missing":[],"search":"selfconsistentscheduledcausallocalplanningbound_le_rateat banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausallocalplanningbound_le_rateat the exact local planning budget is bounded by the named vanishing rate. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope","label":"selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope","description":"Burn-in regret plus the full weighted causal planning-rate average.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-5c78883939a9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7793,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:707"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalburninexpectedregretrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninexpectedregretrateenvelope burn-in regret plus the full weighted causal planning-rate average. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_tendsto_zero","label":"selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_tendsto_zero","description":"For fixed burn-in, early regret is diluted and the tail planning rate vanishes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-11d39111b538","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7794,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:725"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin : Nat) : Tendsto (selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope mdp varianceProxy baseVisitFloor burnin) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalburninexpectedregretrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninexpectedregretrateenvelope_tendsto_zero for fixed burn-in, early regret is diluted and the tail planning rate vanishes. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_weightedExpectedSuccessorAverageRegret_le_burninEnvelope","label":"selfConsistentScheduledCausalSource_weightedExpectedSuccessorAverageRegret_le_burninEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_weightedExpectedSuccessorAverageRegret_le_burninEnvelope","description":"Tail model confidence yields a burn-in-diluted weighted expected regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-c0c37c89e7b0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7795,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:759"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_weightedExpectedSuccessorAverageRegret_le_burninEnvelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (hrounds : 0 < rounds) (trajectory : HeterogeneousStochasticEpisod…","missing":[],"search":"selfconsistentscheduledcausalsource_weightedexpectedsuccessoraverageregret_le_burninenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_weightedexpectedsuccessoraverageregret_le_burninenvelope tail model confidence yields a burn-in-diluted weighted expected regret. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta","label":"selfConsistentScheduledCausalVanishingReturnDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta","description":"Return confidence share adapted to the actual causal successor mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-5e3a8a21c8a4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7796,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:988"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalVanishingReturnDelta (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalvanishingreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturndelta return confidence share adapted to the actual causal successor mass. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_pos","label":"selfConsistentScheduledCausalVanishingReturnDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_pos","description":"theorem selfConsistentScheduledCausalVanishingReturnDelta_pos (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 < selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor rounds","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-1990a33cddcb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7797,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1002"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalVanishingReturnDelta_pos (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 < selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentscheduledcausalvanishingreturndelta_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturndelta_pos theorem selfconsistentscheduledcausalvanishingreturndelta_pos (mdp : mdp state action) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) : 0 < selfconsistentscheduledcausalvanishingreturndelta mdp varianceproxy basevisitfloor rounds theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_le_one","label":"selfConsistentScheduledCausalVanishingReturnDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_le_one","description":"theorem selfConsistentScheduledCausalVanishingReturnDelta_le_one (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor rounds <= 1","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-b299142101c5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7798,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1013"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalVanishingReturnDelta_le_one (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor rounds <= 1","missing":[],"search":"selfconsistentscheduledcausalvanishingreturndelta_le_one banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturndelta_le_one theorem selfconsistentscheduledcausalvanishingreturndelta_le_one (mdp : mdp state action) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) : selfconsistentscheduledcausalvanishingreturndelta mdp varianceproxy basevisitfloor rounds <= 1 theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_tendsto_zero","label":"selfConsistentScheduledCausalVanishingReturnDelta_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_tendsto_zero","description":"The mass-adapted return failure share tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-ba0ac8d2b62d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7799,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1026"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalVanishingReturnDelta_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalvanishingreturndelta_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturndelta_tendsto_zero the mass-adapted return failure share tends to zero. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnRateEnvelope","label":"selfConsistentScheduledCausalVanishingReturnRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnRateEnvelope","description":"Explicit normalized return radius for the mass-adapted confidence share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-c7f28ac6fd05","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7800,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1069"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalVanishingReturnRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalvanishingreturnrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturnrateenvelope explicit normalized return radius for the mass-adapted confidence share. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnConfidenceRadius_vanishingDelta_eq","label":"normalizedSuccessorGlobalReturnConfidenceRadius_vanishingDelta_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnConfidenceRadius_vanishingDelta_eq","description":"The mass-adapted normalized return radius has the explicit envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-82af0118f993","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7801,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1089"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_vanishingDelta_eq (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (hrounds : 0 < rounds) : let episodes := fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds 1 varianceProxy (selfConsistentScheduledCausalVanishingReturnDelta mdp varianceProxy baseVisitFloor rounds) = selfConsistentScheduledCausalVanishingReturnRateEnvelope mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius_vanishingdelta_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.normalizedsuccessorglobalreturnconfidenceradius_vanishingdelta_eq the mass-adapted normalized return radius has the explicit envelope. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnRateEnvelope_tendsto_zero","label":"selfConsistentScheduledCausalVanishingReturnRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnRateEnvelope_tendsto_zero","description":"The mass-adapted normalized return radius tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-114d6f5a697f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7802,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalVanishingReturnRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalVanishingReturnRateEnvelope mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalvanishingreturnrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalvanishingreturnrateenvelope_tendsto_zero the mass-adapted normalized return radius tends to zero. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelReturnBadEvent","label":"selfConsistentScheduledCausalTailModelReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelReturnBadEvent","description":"Tail model event combined with the mass-adapted two-sided return event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-e70d6dbf0f3e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7803,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1198"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalTailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentscheduledcausaltailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelreturnbadevent tail model event combined with the mass-adapted two-sided return event. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelReturnFailureBudget","label":"selfConsistentScheduledCausalTailModelReturnFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelReturnFailureBudget","description":"Exact tail-model plus mass-adapted return failure budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-f3d377053b34","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7804,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1218"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalTailModelReturnFailureBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : ENNReal","missing":[],"search":"selfconsistentscheduledcausaltailmodelreturnfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelreturnfailurebudget exact tail-model plus mass-adapted return failure budget. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope","label":"selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope","description":"Burn-in expected-regret envelope plus the mass-adapted return radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-c410c20278de","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7805,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalburninrealizedregretrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninrealizedregretrateenvelope burn-in expected-regret envelope plus the mass-adapted return radius. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope_tendsto_zero","label":"selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope_tendsto_zero","description":"For fixed burn-in, the complete deterministic realized envelope vanishes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-34fc1fb1d634","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7806,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1239"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin : Nat) : Tendsto (selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope mdp varianceProxy baseVisitFloor burnin) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalburninrealizedregretrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninrealizedregretrateenvelope_tendsto_zero for fixed burn-in, the complete deterministic realized envelope vanishes. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_tail_optimism_and_absoluteRealizedRegret","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_tail_optimism_and_absoluteRealizedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_tail_optimism_and_absoluteRealizedRegret","description":"Tail optimism and absolute realized regret on the natural causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-45a408cf66da","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7807,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1258"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_tail_optimism_and_absoluteRealizedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (hrounds : 0 < rounds) : let episodes := fun t => AdaptiveStocha…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_tail_optimism_and_absoluterealizedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_tail_optimism_and_absoluterealizedregret tail optimism and absolute realized regret on the natural causal source. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretProcess","label":"selfConsistentScheduledNaturalCausalRealizedRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretProcess","description":"Realized successor-average regret process on one natural causal trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-aba031c58788","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7808,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1462"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRealizedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedregretprocess realized successor-average regret process on one natural causal trajectory. definition compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","description":"Every coordinate of the natural causal regret process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-5facc54a7a48","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7809,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1478"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalRealizedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalRealizedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalrealizedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalrealizedregretprocess every coordinate of the natural causal regret process is measurable. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_realizedRegret_tendstoInMeasure_zero","label":"selfConsistentScheduledCausalSource_realizedRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_realizedRegret_tendstoInMeasure_zero","description":"Natural causal realized successor-average regret converges in probability.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-334d05996e51/index.html#decl-51ab705849c3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","order":7810,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency.lean:1501"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_realizedRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor…","missing":[],"search":"selfconsistentscheduledcausalsource_realizedregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_realizedregret_tendstoinmeasure_zero natural causal realized successor-average regret converges in probability. theorem compiled","shard":"modules/f810818bfcf103b6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret_le_two_mul_horizon","label":"successorWeightedExpectedAverageRegret_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret_le_two_mul_horizon","description":"Every heterogeneous weighted selected-policy average is at most `2H`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-98c45e4c45d2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7811,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorWeightedExpectedAverageRegret_le_two_mul_horizon {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hmass : 0 < successorEpisodeMass episodes rounds) (hrewardBound : forall state action, |mdp.reward state action| <= 1) : source.successorWeightedExpectedAverageRegret trajectory rounds <= 2 * (mdp.horizon : Real)","missing":[],"search":"successorweightedexpectedaverageregret_le_two_mul_horizon banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorweightedexpectedaverageregret_le_two_mul_horizon every heterogeneous weighted selected-policy average is at most `2h`. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","label":"trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","description":"The heterogeneous successor deviation has one global sub-Gaussian MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-15575fbafc5d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7812,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:94"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBo…","missing":[],"search":"trajectorymeasure_cumulativesuccessorglobalreturndeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_cumulativesuccessorglobalreturndeviation_hassubgaussianmgf the heterogeneous successor deviation has one global sub-gaussian mgf. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","label":"integrable_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","description":"Finite-prefix heterogeneous realized successor-average regret is integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-ae27cc832435","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7813,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : forall t, 0 < episodes t) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward…","missing":[],"search":"integrable_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.integrable_realizedsuccessoraverageregret finite-prefix heterogeneous realized successor-average regret is integrable. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound","label":"normalizedSuccessorGlobalReturnMGFFirstMomentBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound","description":"Direct first-moment envelope for the normalized heterogeneous deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-bfd6d8749438","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7814,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:216"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedSuccessorGlobalReturnMGFFirstMomentBound (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : Real","missing":[],"search":"normalizedsuccessorglobalreturnmgffirstmomentbound banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnmgffirstmomentbound direct first-moment envelope for the normalized heterogeneous deviation. definition compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","label":"normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","description":"The normalized heterogeneous MGF first-moment envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-6ab93f77bc29","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7815,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:230"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : 0 <= normalizedSuccessorGlobalReturnMGFFirstMomentBound mdp episodes rounds rewardBound rewardVarianceProxy","missing":[],"search":"normalizedsuccessorglobalreturnmgffirstmomentbound_nonneg banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnmgffirstmomentbound_nonneg the normalized heterogeneous mgf first-moment envelope is nonnegative. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","label":"normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","description":"A sufficiently wide normalized confidence radius dominates the MGF mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-bb69f58d310c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7816,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:244"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) (hmass : 0 < successorEpisodeMass episodes rounds) (hlog : (1 / 2 : Real) <= Real.log (2 / delta)) : normalizedSuccessorGlobalReturnMGFFirstMomentBound mdp episodes rounds rewardBound rewardVarianceProxy <= 2 * Real.exp (1 / 2 : Real) * normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy delta","missing":[],"search":"normalizedsuccessorglobalreturnmgffirstmomentbound_le_confidenceradius banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnmgffirstmomentbound_le_confidenceradius a sufficiently wide normalized confidence radius dominates the mgf mean. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_ne_top","label":"selfConsistentScheduledCausalTailModelFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_ne_top","description":"Every post-burn-in model budget is finite.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-7da635e9d085","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7817,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:298"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalTailModelFailureBudget_ne_top (mdp : MDP State Action) (burnin : Nat) : selfConsistentScheduledCausalTailModelFailureBudget mdp burnin ≠ ⊤","missing":[],"search":"selfconsistentscheduledcausaltailmodelfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelfailurebudget_ne_top every post-burn-in model budget is finite. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_toReal_tendsto_zero","label":"selfConsistentScheduledCausalTailModelFailureBudget_toReal_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_toReal_tendsto_zero","description":"The real-valued post-burn-in model tail also tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-22613f06506a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7818,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:317"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalTailModelFailureBudget_toReal_tendsto_zero (mdp : MDP State Action) : Tendsto (fun burnin => (selfConsistentScheduledCausalTailModelFailureBudget mdp burnin).toReal) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausaltailmodelfailurebudget_toreal_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausaltailmodelfailurebudget_toreal_tendsto_zero the real-valued post-burn-in model tail also tends to zero. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound","label":"selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound","description":"Scheduled direct-MGF contribution under actual successor mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-2ae5e7630e2f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7819,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:328"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalnormalizedreturnmgffirstmomentbound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalnormalizedreturnmgffirstmomentbound scheduled direct-mgf contribution under actual successor mass. definition compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_nonneg","label":"selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_nonneg","description":"The scheduled direct-MGF contribution is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-52a146e779ab","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7820,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:341"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentscheduledcausalnormalizedreturnmgffirstmomentbound_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalnormalizedreturnmgffirstmomentbound_nonneg the scheduled direct-mgf contribution is nonnegative. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_tendsto_zero","label":"selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_tendsto_zero","description":"The scheduled direct-MGF first-moment contribution tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-55192174408f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7821,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:353"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalnormalizedreturnmgffirstmomentbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalnormalizedreturnmgffirstmomentbound_tendsto_zero the scheduled direct-mgf first-moment contribution tends to zero. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_nonneg","label":"selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_nonneg","description":"The fixed-burn-in expected-regret envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-edb1a5c3a6ef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7822,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:421"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : 0 <= selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope mdp varianceProxy baseVisitFloor burnin rounds","missing":[],"search":"selfconsistentscheduledcausalburninexpectedregretrateenvelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninexpectedregretrateenvelope_nonneg the fixed-burn-in expected-regret envelope is nonnegative. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret","description":"Expected absolute regret of one natural causal prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-acda0e96ff8a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7823,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:452"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret expected absolute regret of one natural causal prefix. definition compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope","label":"selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope","description":"Two-parameter direct `L1` envelope after a fixed model burn-in.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-5eff567d1299","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7824,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:467"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalburninexpectedabsoluteregretl1envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninexpectedabsoluteregretl1envelope two-parameter direct `l1` envelope after a fixed model burn-in. definition compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope_nonneg","label":"selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope_nonneg","description":"The direct natural-causal `L1` envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-5aca727494ad","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7825,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:481"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) : 0 <= selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope mdp varianceProxy baseVisitFloor burnin rounds","missing":[],"search":"selfconsistentscheduledcausalburninexpectedabsoluteregretl1envelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalburninexpectedabsoluteregretl1envelope_nonneg the direct natural-causal `l1` envelope is nonnegative. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","label":"integrable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","description":"Every coordinate of the natural causal realized-regret process is integrable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-3e9f7aaa7363","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7826,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:496"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalRealizedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : Integrable (selfConsistentScheduledNaturalCausalRealizedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalrealizedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalrealizedregretprocess every coordinate of the natural causal realized-regret process is integrable. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_nonneg","description":"Expected absolute natural-causal regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-995f3958e54a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7827,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:547"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret_nonneg expected absolute natural-causal regret is nonnegative. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_le_burninL1Envelope","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_le_burninL1Envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_le_burninL1Envelope","description":"The expected absolute natural-causal regret is controlled by the good-event planning envelope, the bounded model-tail overflow, and the directly integrated global-return MGF contribution.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-d2121259f230","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7828,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:566"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_le_burninL1Envelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (hrounds : 0 < rounds) : selfConsistentScheduledNaturalCausalExpectedAbs…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret_le_burninl1envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret_le_burninl1envelope the expected absolute natural-causal regret is controlled by the good-event planning envelope, the bounded model-tail overflow, and the directly integrated global-return mgf contribution. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_tendsto_zero","description":"Expected absolute regret on the natural causal prefixes tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-aa45b578b5cc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7829,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:755"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret mdp initialState rewardSource initialTable defaultSt…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluterealizedregret_tendsto_zero expected absolute regret on the natural causal prefixes tends to zero. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess","label":"memLp_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess","description":"Every coordinate of the natural causal process belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-d8f50c10f1b1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7830,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:860"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : MemLp (selfConsistentScheduledNaturalCausalRealizedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalrealizedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalrealizedregretprocess every coordinate of the natural causal process belongs to `l1`. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_eq","description":"At exponent one, `eLpNorm` is the lifted expected absolute regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-953d52e484bf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7831,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:884"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : eLpNorm (selfConsistentScheduledNaturalCausalRealizedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure = ENNReal.ofReal (selfConsistentScheduledNatura…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalrealizedregretprocess_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalrealizedregretprocess_eq at exponent one, `elpnorm` is the lifted expected absolute regret. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_tendsto_zero","description":"The exponent-one extended norm on the causal prefixes tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-2f8c986c7b87","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7832,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:914"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun rounds => eLpNorm (selfConsistentScheduledNaturalCausalRealizedRegretProcess mdp initialState rewardSource initi…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalrealizedregretprocess_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalrealizedregretprocess_tendsto_zero the exponent-one extended norm on the causal prefixes tends to zero. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_sub_zero_tendsto_zero","description":"Canonical natural-causal `L1` norm-of-the-difference convergence.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-3839d17f9f62","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7833,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:949"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun rounds => eLpNorm (selfConsistentScheduledNaturalCausalRealizedRegretProcess mdp initialState rewardSou…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalrealizedregretprocess_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalrealizedregretprocess_sub_zero_tendsto_zero canonical natural-causal `l1` norm-of-the-difference convergence. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp","label":"selfConsistentScheduledNaturalCausalRealizedRegretLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp","description":"The natural causal realized-regret process as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-cb43bb8f94bb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7834,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:984"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRealizedRegretLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : Lp Real 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedregretlp banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedregretlp the natural causal realized-regret process as an `lp real 1` value. definition compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp_coeFn_ae_eq","label":"selfConsistentScheduledNaturalCausalRealizedRegretLp_coeFn_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp_coeFn_ae_eq","description":"The named natural-causal `Lp` coordinate represents the process a.e.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-bf7fa743b677","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7835,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:1005"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRealizedRegretLp_coeFn_ae_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : (selfConsistentScheduledNaturalCausalRealizedRegretLp mdp initialState rewardSource varianceProxy law initialTable defaultState baseVisitFloor hrewardBound rounds : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real) =ᵐ[ (selfConsistent…","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedregretlp_coefn_ae_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedregretlp_coefn_ae_eq the named natural-causal `lp` coordinate represents the process a.e. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp_tendsto_zero","label":"selfConsistentScheduledNaturalCausalRealizedRegretLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp_tendsto_zero","description":"The named natural-causal `Lp Real 1` process converges to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-11c3852afefb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7836,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:1034"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRealizedRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalRealizedRegretLp mdp initialState rewardSource varianceProxy law initialTable defaultState baseVi…","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedregretlp_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedregretlp_tendsto_zero the named natural-causal `lp real 1` process converges to zero. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalRealizedRegret_memLp_eLpNorm_L1_tendsto_zero","label":"selfConsistentScheduledCausalSource_naturalRealizedRegret_memLp_eLpNorm_L1_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalRealizedRegret_memLp_eLpNorm_L1_tendsto_zero","description":"Terminal natural-causal `L1` theorem on one heterogeneous dependent trajectory measure: exact exponent-one norms, `Lp` convergence, and convergence in measure.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a3aa468b7915/index.html#decl-d32cf0526653","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","order":7837,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency.lean:1084"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_naturalRealizedRegret_memLp_eLpNorm_L1_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy…","missing":[],"search":"selfconsistentscheduledcausalsource_naturalrealizedregret_memlp_elpnorm_l1_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_naturalrealizedregret_memlp_elpnorm_l1_tendsto_zero terminal natural-causal `l1` theorem on one heterogeneous dependent trajectory measure: exact exponent-one norms, `lp` convergence, and convergence in measure. theorem compiled","shard":"modules/8de342c11ea19258.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.natWeightedAverage","label":"natWeightedAverage","kind":"definition","status":"compiled","subtitle":"BanditRLProof.natWeightedAverage","description":"A finite average with positive natural-number weights.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-494c11b2b2d6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7838,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def natWeightedAverage (weight : Nat -> Nat) (value : Nat -> Real) (rounds : Nat) : Real","missing":[],"search":"natweightedaverage banditrlproof.natweightedaverage a finite average with positive natural-number weights. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.tendsto_natWeightedAverage_zero","label":"tendsto_natWeightedAverage_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tendsto_natWeightedAverage_zero","description":"Positive natural weights preserve a zero limit under finite weighted averaging.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-cde73c5dc5a7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7839,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tendsto_natWeightedAverage_zero (weight : Nat -> Nat) (hweight : forall t, 0 < weight t) (value : Nat -> Real) (hvalue : Tendsto value atTop (nhds 0)) : Tendsto (natWeightedAverage weight value) atTop (nhds 0)","missing":[],"search":"tendsto_natweightedaverage_zero banditrlproof.tendsto_natweightedaverage_zero positive natural weights preserve a zero limit under finite weighted averaging. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_eq_sum_range","label":"successorEpisodeMass_eq_sum_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_eq_sum_range","description":"The successor episode mass is the corresponding range sum.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-6163b4b8d867","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7840,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorEpisodeMass_eq_sum_range (episodes : Nat -> Nat) (rounds : Nat) : successorEpisodeMass episodes rounds = (Finset.range rounds).sum (fun t => (episodes (t + 1) : Real))","missing":[],"search":"successorepisodemass_eq_sum_range banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorepisodemass_eq_sum_range the successor episode mass is the corresponding range sum. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_tendsto_atTop","label":"successorEpisodeMass_tendsto_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_tendsto_atTop","description":"Positive coordinate batch sizes make successor mass diverge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-1e3bd234f976","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7841,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:93"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorEpisodeMass_tendsto_atTop (episodes : Nat -> Nat) (hepisodes : forall t, 0 < episodes t) : Tendsto (successorEpisodeMass episodes) atTop atTop","missing":[],"search":"successorepisodemass_tendsto_attop banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorepisodemass_tendsto_attop positive coordinate batch sizes make successor mass diverge. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_coe","label":"cumulativeSuccessorGlobalReturnVarianceProxy_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_coe","description":"The heterogeneous global return proxy is exactly linear in successor mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-f42ed510f895","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7842,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorGlobalReturnVarianceProxy_coe (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : ((cumulativeSuccessorGlobalReturnVarianceProxy mdp episodes rounds rewardBound rewardVarianceProxy : NNReal) : Real) = successorEpisodeMass episodes rounds * (mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy : Real)","missing":[],"search":"cumulativesuccessorglobalreturnvarianceproxy_coe banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturnvarianceproxy_coe the heterogeneous global return proxy is exactly linear in successor mass. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius","label":"normalizedSuccessorGlobalReturnConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius","description":"Globally centered successor-return radius divided by actual successor mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-0a1870523f77","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7843,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def normalizedSuccessorGlobalReturnConfidenceRadius (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : Real","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius globally centered successor-return radius divided by actual successor mass. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_eq","label":"normalizedSuccessorGlobalReturnConfidenceRadius_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_eq","description":"Exact square-root formula for the heterogeneous normalized return radius.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-117df1abc323","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7844,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:139"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_eq (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) (hmass : 0 < successorEpisodeMass episodes rounds) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy delta = Real.sqrt (2 * (mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy : Real) * Real.log (2 / delta) / successorEpisodeMass episodes rounds)","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius_eq banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius_eq exact square-root formula for the heterogeneous normalized return radius. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.fixedHalfSuccessorGlobalReturnRateEnvelope","label":"fixedHalfSuccessorGlobalReturnRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.fixedHalfSuccessorGlobalReturnRateEnvelope","description":"Fixed-half return envelope on the actual heterogeneous successor mass.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-bff10f2b4577","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7845,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:198"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def fixedHalfSuccessorGlobalReturnRateEnvelope (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : Real","missing":[],"search":"fixedhalfsuccessorglobalreturnrateenvelope banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.fixedhalfsuccessorglobalreturnrateenvelope fixed-half return envelope on the actual heterogeneous successor mass. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_half_eq","label":"normalizedSuccessorGlobalReturnConfidenceRadius_half_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_half_eq","description":"At confidence share one half, the normalized radius is the fixed-half envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-d3f9aaed4524","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7846,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem normalizedSuccessorGlobalReturnConfidenceRadius_half_eq (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hmass : 0 < successorEpisodeMass episodes rounds) : normalizedSuccessorGlobalReturnConfidenceRadius mdp episodes rounds rewardBound rewardVarianceProxy (1 / 2) = fixedHalfSuccessorGlobalReturnRateEnvelope mdp episodes rounds rewardBound rewardVarianceProxy","missing":[],"search":"normalizedsuccessorglobalreturnconfidenceradius_half_eq banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.normalizedsuccessorglobalreturnconfidenceradius_half_eq at confidence share one half, the normalized radius is the fixed-half envelope. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.fixedHalfSuccessorGlobalReturnRateEnvelope_tendsto_zero","label":"fixedHalfSuccessorGlobalReturnRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.fixedHalfSuccessorGlobalReturnRateEnvelope_tendsto_zero","description":"The fixed-half heterogeneous return envelope vanishes with prefix length.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-a49d0ccc2723","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7847,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem fixedHalfSuccessorGlobalReturnRateEnvelope_tendsto_zero (mdp : MDP State Action) (episodes : Nat -> Nat) (hepisodes : forall t, 0 < episodes t) (rewardBound rewardVarianceProxy : NNReal) : Tendsto (fun rounds => fixedHalfSuccessorGlobalReturnRateEnvelope mdp episodes rounds rewardBound rewardVarianceProxy) atTop (nhds 0)","missing":[],"search":"fixedhalfsuccessorglobalreturnrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.fixedhalfsuccessorglobalreturnrateenvelope_tendsto_zero the fixed-half heterogeneous return envelope vanishes with prefix length. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt","label":"selfConsistentScheduledCausalPlanningRateAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt","description":"Coordinatewise planning rate for the genuinely causal successor batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-3ef51db7900e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7848,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:250"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalPlanningRateAt (mdp : MDP State Action) (t : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalplanningrateat banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalplanningrateat coordinatewise planning rate for the genuinely causal successor batch. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt_tendsto_zero","label":"selfConsistentScheduledCausalPlanningRateAt_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt_tendsto_zero","description":"The causal coordinatewise planning rate vanishes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-4b61f8af4186","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7849,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:264"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalPlanningRateAt_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledCausalPlanningRateAt mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalplanningrateat_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalplanningrateat_tendsto_zero the causal coordinatewise planning rate vanishes. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalWeightedPlanningRateEnvelope","label":"selfConsistentScheduledCausalWeightedPlanningRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalWeightedPlanningRateEnvelope","description":"Positive scheduled successor weights applied to the coordinatewise rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-ea15874a08ae","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7850,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:302"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalWeightedPlanningRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalweightedplanningrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalweightedplanningrateenvelope positive scheduled successor weights applied to the coordinatewise rate. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningAverageBound_le_rateEnvelope","label":"selfConsistentScheduledCausalSuccessorPlanningAverageBound_le_rateEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningAverageBound_le_rateEnvelope","description":"The exact causal planning average is controlled by the weighted rate envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-c1f723cf589c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7851,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSuccessorPlanningAverageBound_le_rateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledCausalSuccessorPlanningAverageBound mdp varianceProxy baseVisitFloor rounds <= selfConsistentScheduledCausalWeightedPlanningRateEnvelope mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentscheduledcausalsuccessorplanningaveragebound_le_rateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsuccessorplanningaveragebound_le_rateenvelope the exact causal planning average is controlled by the weighted rate envelope. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalWeightedPlanningRateEnvelope_tendsto_zero","label":"selfConsistentScheduledCausalWeightedPlanningRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalWeightedPlanningRateEnvelope_tendsto_zero","description":"The scheduled weighted causal planning envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-cc742db8a2cf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7852,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalWeightedPlanningRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalWeightedPlanningRateEnvelope mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalweightedplanningrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalweightedplanningrateenvelope_tendsto_zero the scheduled weighted causal planning envelope tends to zero. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalReturnRateEnvelope","label":"selfConsistentScheduledCausalReturnRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalReturnRateEnvelope","description":"Fixed-half global return envelope for the scheduled causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-d0d8baef710f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7853,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:406"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalReturnRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalreturnrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalreturnrateenvelope fixed-half global return envelope for the scheduled causal source. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalReturnRateEnvelope_tendsto_zero","label":"selfConsistentScheduledCausalReturnRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalReturnRateEnvelope_tendsto_zero","description":"The scheduled causal return envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-5cba7020b017","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7854,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:419"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalReturnRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalReturnRateEnvelope mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalreturnrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalreturnrateenvelope_tendsto_zero the scheduled causal return envelope tends to zero. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope","label":"selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope","description":"Full deterministic realized-regret envelope on the causal successor prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-2a3570b2e23c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7855,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:434"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope full deterministic realized-regret envelope on the causal successor prefix. definition compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","label":"selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","description":"The complete deterministic causal realized-regret envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-ff9697a666af","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7856,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:446"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope_tendsto_zero the complete deterministic causal realized-regret envelope tends to zero. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret_le_explicitRateEnvelope","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret_le_explicitRateEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret_le_explicitRateEnvelope","description":"The complete deterministic causal realized-regret envelope tends to zero. -/ theorem selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope mdp varianceProxy baseVisitFloor) atTop (nhds 0) := by unfold selfConsistentScheduledCausalR…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-6eb799af80ec/index.html#decl-677caaa5c425","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","order":7857,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate.lean:465"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret_le_explicitRateEnvelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) : let episodes := fun t => AdaptiveStochasticEpiso…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_optimism_and_realizedsuccessoraverageregret_le_explicitrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_optimism_and_realizedsuccessoraverageregret_le_explicitrateenvelope the complete deterministic causal realized-regret envelope tends to zero. -/ theorem selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope_tendsto_zero (mdp : mdp state action) (varianceproxy : nnreal) (basevisitfloor : real) : tendsto (selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope mdp varianceproxy basevisitfloor) attop (nhds 0) := by unfold selfconsistentscheduledcausalrealizedsuccessoraverageregretrateenvelope simpa using (selfconsistentscheduledcausalweightedplanningrateenvelope_tendsto_zero mdp varianceproxy basevisitfloor).add (selfconsistentscheduledcausalreturnrateenvelope_tendsto_zero mdp varianceproxy basevisitfloor) /- the deterministic regret envelope vanishes, but the exact finite-prefix model failure budget below accumulates the early coordinate events. this theorem therefore deliberately stops at a finite-prefix certificate. theorem compiled","shard":"modules/08096d6e6305c654.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.policyAt","label":"policyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.policyAt","description":"Policy whose exact iid law generated coordinate `t`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-109a4f75e57a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7858,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def policyAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Nat -> MarkovPolicy mdp | 0 => source.initialPolicy | n + 1 => source.successorPolicy n (Preorder.frestrictLe n trajectory) /-- Pull an initial heterogeneous-batch event back to the causal trajectory. -/ def initialBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} (bad : Set (StochasticEpisodeBatch mdp (episodes 0))) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"policyat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.policyat policy whose exact iid law generated coordinate `t`. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialBadEvent","label":"initialBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialBadEvent","description":"Pull an initial heterogeneous-batch event back to the causal trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-30199423255e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7859,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def initialBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} (bad : Set (StochasticEpisodeBatch mdp (episodes 0))) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"initialbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.initialbadevent pull an initial heterogeneous-batch event back to the causal trajectory. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorBadEvent","label":"successorBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorBadEvent","description":"Pull a prefix-dependent successor event back to the causal trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-3dbe5e3bbfdb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7860,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def successorBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} (n : Nat) (bad : Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"successorbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorbadevent pull a prefix-dependent successor event back to the causal trajectory. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.roundBadEvent","label":"roundBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.roundBadEvent","description":"Coordinate-indexed adapted heterogeneous-batch event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-17283e2b9150","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7861,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def roundBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} (initialBad : Set (StochasticEpisodeBatch mdp (episodes 0))) (successorBad : (n : Nat) -> Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))) : Nat -> Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) | 0 => initialBadEvent initialBad | n + 1 => successorBadEvent n (successorBad n) /-- Union of the first `rounds` heterogeneous adapted events. -/ def finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} (rounds : Nat) (initialBad : Set (StochasticEpisodeBatch mdp (episodes 0))) (successorBad : (n : Nat) -> Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"roundbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.roundbadevent coordinate-indexed adapted heterogeneous-batch event. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent","label":"finiteHorizonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent","description":"Union of the first `rounds` heterogeneous adapted events.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-2571fa007bbc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7862,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} (rounds : Nat) (initialBad : Set (StochasticEpisodeBatch mdp (episodes 0))) (successorBad : (n : Nat) -> Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"finitehorizonbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.finitehorizonbadevent union of the first `rounds` heterogeneous adapted events. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","label":"measurableSet_finiteHorizonBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","description":"Measurability of the finite heterogeneous adapted union.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-14737942c632","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7863,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_finiteHorizonBadEvent {mdp : MDP State Action} {episodes : Nat -> Nat} {rounds : Nat} {initialBad : Set (StochasticEpisodeBatch mdp (episodes 0))} {successorBad : (n : Nat) -> Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))} (hinitial : MeasurableSet initialBad) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (successorBad n)) : MeasurableSet (finiteHorizonBadEvent rounds initialBad successorBad)","missing":[],"search":"measurableset_finitehorizonbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurableset_finitehorizonbadevent measurability of the finite heterogeneous adapted union. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_initialBadEvent","label":"trajectoryMeasure_initialBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_initialBadEvent","description":"Exact mass of a pulled-back initial heterogeneous-batch event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-4b1bc3b9bec8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7864,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_initialBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) {bad : Set (StochasticEpisodeBatch mdp (episodes 0))} (hbad : MeasurableSet bad) : source.trajectoryMeasure (initialBadEvent bad) = source.rewardSource.iidStochasticTrajectoryFamilyMeasure source.initialPolicy initialState (episodes 0) bad","missing":[],"search":"trajectorymeasure_initialbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_initialbadevent exact mass of a pulled-back initial heterogeneous-batch event. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","label":"trajectoryMeasure_successorBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","description":"A successor event inherits a uniform bound on every selected history fiber.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-7bf6a3daa4cd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7865,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successorBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) {bad : Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))} (hbad : MeasurableSet bad) (budget : ENNReal) (hfiber : forall history, source.batchKernel n history (Prod.mk history ⁻¹' bad) <= budget) : source.trajectoryMeasure (successorBadEvent n bad) <= budget","missing":[],"search":"trajectorymeasure_successorbadevent_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_successorbadevent_le a successor event inherits a uniform bound on every selected history fiber. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent","label":"initialAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent","description":"Initial selected empirical-model event with coordinate-specific shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-2d89b9e08749","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7866,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def initialAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) : Set (StochasticEpisodeBatch mdp (episodes 0))","missing":[],"search":"initialallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.initialallcoordinateempiricalmodelbadevent initial selected empirical-model event with coordinate-specific shares. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent","label":"successorAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent","description":"Prefix-selected successor empirical-model event with local shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-bb3510ccf662","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7867,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:192"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) (n : Nat) : Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))","missing":[],"search":"successorallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorallcoordinateempiricalmodelbadevent prefix-selected successor empirical-model event with local shares. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.adaptiveAllCoordinateEmpiricalModelBadEvent","label":"adaptiveAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.adaptiveAllCoordinateEmpiricalModelBadEvent","description":"First-`rounds` selected count-and-reward empirical-model event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-0ea4d1ee75eb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7868,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:206"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"adaptiveallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.adaptiveallcoordinateempiricalmodelbadevent first-`rounds` selected count-and-reward empirical-model event. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeModelFailureBudget","label":"cumulativeModelFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeModelFailureBudget","description":"Exact finite sum of coordinate-specific count and reward failure shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-a57466fe6487","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7869,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:220"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeModelFailureBudget (rounds : Nat) (countDelta rewardDelta : Nat -> Real) : ENNReal","missing":[],"search":"cumulativemodelfailurebudget banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativemodelfailurebudget exact finite sum of coordinate-specific count and reward failure shares. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_initialAllCoordinateEmpiricalModelBadEvent","label":"measurableSet_initialAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_initialAllCoordinateEmpiricalModelBadEvent","description":"The initial coordinate event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-6555fa8468f6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7870,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:226"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_initialAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) : MeasurableSet (source.initialAllCoordinateEmpiricalModelBadEvent varianceProxy countDelta rewardDelta)","missing":[],"search":"measurableset_initialallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurableset_initialallcoordinateempiricalmodelbadevent the initial coordinate event is measurable. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","label":"measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","description":"Selected successor measurability closes the finite heterogeneous event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-d2254bd98196","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7871,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:240"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) (hsuccessor : forall n, n + 1 < rounds -> MeasurableSet (source.successorAllCoordinateEmpiricalModelBadEvent varianceProxy countDelta rewardDelta n)) : MeasurableSet (source.adaptiveAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta)","missing":[],"search":"measurableset_adaptiveallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurableset_adaptiveallcoordinateempiricalmodelbadevent selected successor measurability closes the finite heterogeneous event. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent_le","label":"initialAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent_le","description":"The initial selected iid fiber receives its coordinate shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-a73282ca837b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7872,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:257"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem initialAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (hepisodes : 0 < episodes 0) (varianceProxy : NNReal) (law : source.rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes 0 : NNReal) * varianceProxy : NNReal) : Real))) (countDelta rewardDelta : Nat -> Real) (hcountDelta : 0 < countDelta 0) (hcountDelta_le_one : countDelta 0 <= 1) (hrewardDelta : 0 < rewardDelta 0) (hrewardDelta_le_one : rewardDelta 0 <= 1) : (source.rewardSource.iidStochasticTrajectoryFamilyMeasure source.initialPolicy initialState (episodes 0)) (source.initialAllCoordinateEmpiricalModelBadEvent varianceProxy countDelta rewardDelta)…","missing":[],"search":"initialallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.initialallcoordinateempiricalmodelbadevent_le the initial selected iid fiber receives its coordinate shares. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent_fiber_le","label":"successorAllCoordinateEmpiricalModelBadEvent_fiber_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent_fiber_le","description":"Every selected successor iid fiber receives its coordinate shares.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-75f0fe8dda0b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7873,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorAllCoordinateEmpiricalModelBadEvent_fiber_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (varianceProxy : NNReal) (law : source.rewardSource.UniformSubgaussianRewardLaw varianceProxy) (countDelta rewardDelta : Nat -> Real) (n : Nat) (hepisodes : 0 < episodes (n + 1)) (htotal : 0 < ((((episodes (n + 1) : NNReal) * varianceProxy : NNReal) : Real))) (hcountDelta : 0 < countDelta (n + 1)) (hcountDelta_le_one : countDelta (n + 1) <= 1) (hrewardDelta : 0 < rewardDelta (n + 1)) (hrewardDelta_le_one : rewardDelta (n + 1) <= 1) (history : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) : source.batchKernel n history (Prod.mk history ⁻¹' source.successorAllCoo…","missing":[],"search":"successorallcoordinateempiricalmodelbadevent_fiber_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorallcoordinateempiricalmodelbadevent_fiber_le every selected successor iid fiber receives its coordinate shares. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","label":"trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","description":"Every selected successor iid fiber receives its coordinate shares. -/ theorem successorAllCoordinateEmpiricalModelBadEvent_fiber_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (varianceProxy : NNReal) (law…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-3420174e40df","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7874,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:323"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (hepisodes : forall t, 0 < episodes t) (varianceProxy : NNReal) (law : source.rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : forall t, 0 < ((((episodes t : NNReal) * varianceProxy : NNReal) : Real))) (countDelta rewardDelta : Nat -> Real) (hcountDelta : forall t, t < rounds -> 0 < countDelta t) (hcountDelta_le_one : forall t, t < rounds -> countDelta t <= 1) (hrewardDelta : forall t, t < rounds -> 0 < rewardDelta t) (hrewardDelta_le_one : forall t, t < rounds -> rewardDelta t <= 1) (hsuccessor : forall n, n + 1 < rounds -> Measu…","missing":[],"search":"trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le every selected successor iid fiber receives its coordinate shares. -/ theorem successorallcoordinateempiricalmodelbadevent_fiber_le [standardborelspace state] [standardborelspace action] {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) (varianceproxy : nnreal) (law : source.rewardsource.uniformsubgaussianrewardlaw varianceproxy) (countdelta rewarddelta : nat -> real) (n : nat) (hepisodes : 0 < episodes (n + 1)) (htotal : 0 < ((((episodes (n + 1) : nnreal) * varianceproxy : nnreal) : real))) (hcountdelta : 0 < countdelta (n + 1)) (hcountdelta_le_one : countdelta (n + 1) <= 1) (hrewarddelta : 0 < rewarddelta (n + 1)) (hrewarddelta_le_one : rewarddelta (n + 1) <= 1) (history : heterogeneousstochasticepisodebatchprefix mdp episodes n) : source.batchkernel n history (prod.mk history ⁻¹' source.successorallcoordinateempiricalmodelbadevent varianceproxy countdelta rewarddelta n) <= ennreal.ofreal (countdelta (n + 1)) + ennreal.ofreal (rewarddelta (n + 1)) := by rw [source.batchkernel_eq_iidstochastictrajectoryfamilymeasure] ch…","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","label":"policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","description":"Outside the finite event, every batch avoids its generating-policy event.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-3041cdc1f00b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7875,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:403"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) (htrajectory : trajectory ∉ source.adaptiveAllCoordinateEmpiricalModelBadEvent rounds varianceProxy countDelta rewardDelta) (round : Fin rounds) : trajectory round ∉ source.rewardSource.stochasticAllCoordinateEmpiricalModelBadEvent (source.policyAt trajectory round) initialState (episodes round) varianceProxy (countDelta round) (rewardDelta round)","missing":[],"search":"policyat_batch_not_mem_allcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.policyat_batch_not_mem_allcoordinateempiricalmodelbadevent outside the finite event, every batch avoids its generating-policy event. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","label":"heterogeneousExploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","description":"Concrete heterogeneous successor events are measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-6c2d357284ce","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7876,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:442"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem heterogeneousExploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (explorationRate : Nat -> NNReal) (hexplorationRate : forall t, explorationRate t <= 1) (varianceProxy : NNReal) (countDelta rewardDelta : Nat -> Real) (n : Nat) : let source := heterogeneousExploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate MeasurableSet (source.successorAllCoordinateEmpiricalModelBadEvent varianceProxy countDelta rewardDelta n)","missing":[],"search":"heterogeneousexploratorysource_measurableset_successorallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneousexploratorysource_measurableset_successorallcoordinateempiricalmodelbadevent concrete heterogeneous successor events are measurable. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","label":"heterogeneousExploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","description":"Concrete heterogeneous source event and its finite failure-budget bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-51a8c4c61239","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7877,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:477"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem heterogeneousExploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (explorationRate : Nat -> NNReal) (hexplorationRate : forall t, explorationRate t <= 1) (rounds : Nat) (hepisodes : forall t, 0 < episodes t) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : forall t, 0 < ((((episodes t : NNReal) * varianceProxy : NNReal) : Real))) (countDelta rewardDelta : Nat -> Real) (hcountDelta : forall t, t < rounds -> 0 < countDelta t) (hcountDelta_le_one : forall t, t < rounds -…","missing":[],"search":"heterogeneousexploratorysource_trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneousexploratorysource_trajectorymeasure_adaptiveallcoordinateempiricalmodelbadevent_le concrete heterogeneous source event and its finite failure-budget bound. theorem compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget","label":"selfConsistentScheduledCausalModelFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget","description":"Model-confidence failure budget of the first causal schedule coordinates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-85b6aa3f5f8b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7878,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:526"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalModelFailureBudget (mdp : MDP State Action) (rounds : Nat) : ENNReal","missing":[],"search":"selfconsistentscheduledcausalmodelfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelfailurebudget model-confidence failure budget of the first causal schedule coordinates. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelBadEvent","label":"selfConsistentScheduledCausalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelBadEvent","description":"Named actual-sampled model event on the self-consistent causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-172b734b0233","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7879,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:534"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalModelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentscheduledcausalmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelbadevent named actual-sampled model event on the self-consistent causal source. definition compiled","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","description":"Named actual-sampled model event on the self-consistent causal source. -/ noncomputable def selfConsistentScheduledCausalModelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Set…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-977e98383175/index.html#decl-6375c7fbda5b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","order":7880,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence.lean:556"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : let episodes := fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistent…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret named actual-sampled model event on the self-consistent causal source. -/ noncomputable def selfconsistentscheduledcausalmodelbadevent (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) : set (heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) := let source := selfconsistentscheduledcausalsource mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor source.adaptiveallcoordinateempiricalmodelbadevent rounds varianceproxy (adaptivestochasticepisodebatchsource.selfconsistentscheduledlocaldelta mdp) (adaptivestochasticepisodebatchsource.selfconsistentscheduledlocaldelta mdp) /- the genuine causal schedule receives coordinatewise sampled-model confidence. every model is computed from the actual batch at tha…","shard":"modules/3e6e674281b0bfa7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorPolicyAt","label":"successorPolicyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorPolicyAt","description":"The selected policy that generates successor coordinate `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-527be072c54e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7881,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorPolicyAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) : MarkovPolicy mdp","missing":[],"search":"successorpolicyat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorpolicyat the selected policy that generates successor coordinate `n + 1`. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass","label":"successorEpisodeMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass","description":"Total number of complete episodes in successor coordinates `1..rounds`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-cb4035554bd5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7882,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorEpisodeMass (episodes : Nat -> Nat) (rounds : Nat) : Real","missing":[],"search":"successorepisodemass banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorepisodemass total number of complete episodes in successor coordinates `1..rounds`. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_pos","label":"successorEpisodeMass_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_pos","description":"theorem successorEpisodeMass_pos (episodes : Nat -> Nat) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : forall t, 0 < episodes t) : 0 < successorEpisodeMass episodes rounds","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-03a4632146ea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7883,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorEpisodeMass_pos (episodes : Nat -> Nat) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : forall t, 0 < episodes t) : 0 < successorEpisodeMass episodes rounds","missing":[],"search":"successorepisodemass_pos banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorepisodemass_pos theorem successorepisodemass_pos (episodes : nat -> nat) (rounds : nat) (hrounds : 0 < rounds) (hepisodes : forall t, 0 < episodes t) : 0 < successorepisodemass episodes rounds theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedCumulativeRegret","label":"successorWeightedExpectedCumulativeRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedCumulativeRegret","description":"Batch-size-weighted expected regret of the selected successor policies.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-9e6baf0d3402","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7884,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorWeightedExpectedCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"successorweightedexpectedcumulativeregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorweightedexpectedcumulativeregret batch-size-weighted expected regret of the selected successor policies. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret","label":"successorWeightedExpectedAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret","description":"Weighted successor expected regret per sampled successor episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-3eee1fc59ad3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7885,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorWeightedExpectedAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"successorweightedexpectedaverageregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorweightedexpectedaverageregret weighted successor expected regret per sampled successor episode. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret","label":"realizedSuccessorCumulativeRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret","description":"Realized regret of every sampled successor batch, with its actual size.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-db11bac5079a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7886,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def realizedSuccessorCumulativeRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (_source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"realizedsuccessorcumulativeregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.realizedsuccessorcumulativeregret realized regret of every sampled successor batch, with its actual size. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret","label":"realizedSuccessorAverageRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret","description":"Realized successor regret per actual sampled successor episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-11eb29963c37","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7887,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def realizedSuccessorAverageRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"realizedsuccessoraverageregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.realizedsuccessoraverageregret realized successor regret per actual sampled successor episode. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_eq","label":"successorGlobalReturnIncrement_succ_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_eq","description":"A successor increment is actual return minus its selected policy mean.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-e0321d06dc89","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7888,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:138"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnIncrement_succ_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) : source.successorGlobalReturnIncrement (n + 1) trajectory = mdp.sampledCumulativeRewardSum (episodes (n + 1)) (trajectory (n + 1)) - (episodes (n + 1) : Real) * integral initialState ((source.successorPolicyAt trajectory n).valueAt 0 (Nat.zero_le mdp.horizon))","missing":[],"search":"successorglobalreturnincrement_succ_eq banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturnincrement_succ_eq a successor increment is actual return minus its selected policy mean. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","label":"cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","description":"The global deviation is exactly the finite sum of successor increments.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-0b31b62e5a5b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7889,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:163"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory = ∑ round : Fin rounds, source.successorGlobalReturnIncrement ((round : Nat) + 1) trajectory","missing":[],"search":"cumulativesuccessorglobalreturndeviation_eq_fin_sum banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturndeviation_eq_fin_sum the global deviation is exactly the finite sum of successor increments. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","label":"realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","description":"Exact weighted realized-equals-expected-minus-deviation identity.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-a178b29adea9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7890,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem realizedSuccessorCumulativeRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.realizedSuccessorCumulativeRegret trajectory rounds = source.successorWeightedExpectedCumulativeRegret trajectory rounds - source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory","missing":[],"search":"realizedsuccessorcumulativeregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.realizedsuccessorcumulativeregret_eq_expected_sub_deviation exact weighted realized-equals-expected-minus-deviation identity. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","label":"realizedSuccessorAverageRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","description":"Exact weighted average form of the realized-regret decomposition.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-7e3348619053","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7891,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:235"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem realizedSuccessorAverageRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hmass : 0 < successorEpisodeMass episodes rounds) : source.realizedSuccessorAverageRegret trajectory rounds = source.successorWeightedExpectedAverageRegret trajectory rounds - source.cumulativeSuccessorGlobalReturnDeviation rounds trajectory / successorEpisodeMass episodes rounds","missing":[],"search":"realizedsuccessoraverageregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.realizedsuccessoraverageregret_eq_expected_sub_deviation exact weighted average form of the realized-regret decomposition. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationBadEvent","label":"successorGlobalReturnDeviationBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationBadEvent","description":"Named successor-only global-return event for heterogeneous batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-c7257dc46202","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7892,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:254"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"successorglobalreturndeviationbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturndeviationbadevent named successor-only global-return event for heterogeneous batches. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_cumulativeSuccessorGlobalReturnDeviation","label":"measurable_cumulativeSuccessorGlobalReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_cumulativeSuccessorGlobalReturnDeviation","description":"theorem measurable_cumulativeSuccessorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) : Measurable (source.cumulativeSuccessorGlobalReturnDeviation rounds)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-241b75db7b37","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7893,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeSuccessorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) : Measurable (source.cumulativeSuccessorGlobalReturnDeviation rounds)","missing":[],"search":"measurable_cumulativesuccessorglobalreturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_cumulativesuccessorglobalreturndeviation theorem measurable_cumulativesuccessorglobalreturndeviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) : measurable (source.cumulativesuccessorglobalreturndeviation rounds) theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_successorGlobalReturnDeviationBadEvent","label":"measurableSet_successorGlobalReturnDeviationBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_successorGlobalReturnDeviationBadEvent","description":"theorem measurableSet_successorGlobalReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : MeasurableSet (source.successorGlobalReturnDe…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-7f2e1c4f6367","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7894,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:286"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_successorGlobalReturnDeviationBadEvent {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (delta : Real) : MeasurableSet (source.successorGlobalReturnDeviationBadEvent rounds rewardBound rewardVarianceProxy delta)","missing":[],"search":"measurableset_successorglobalreturndeviationbadevent banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurableset_successorglobalreturndeviationbadevent theorem measurableset_successorglobalreturndeviationbadevent {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) (rewardbound rewardvarianceproxy : nnreal) (delta : real) : measurableset (source.successorglobalreturndeviationbadevent rounds rewardbound rewardvarianceproxy delta) theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","label":"trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","description":"theorem trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n,…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-a92353202d10","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7895,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:300"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law…","missing":[],"search":"trajectorymeasure_successorglobalreturndeviationbadevent_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_successorglobalreturndeviationbadevent_le theorem trajectorymeasure_successorglobalreturndeviationbadevent_le {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} [standardborelspace state] [standardborelspace action] [forall n, standardborelspace (heterogeneousstochasticepisodebatchprefix mdp episodes n)] [forall n, nonempty (heterogeneousstochasticepisodebatchprefix mdp episodes n)] [forall n, standardborelspace (stochasticepisodebatch mdp (episodes n))] [forall n, nonempty (stochasticepisodebatch mdp (episodes n))] [standardborelspace (heterogeneousstochasticepisodebatchtrajectory mdp episodes)] (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) (rewardbound rewardvarianceproxy : nnreal) (hrewardbound : forall state action, |mdp.reward state action| <= (rewardbound : real)) (law : source.rewardsource.uniformsubgaussianrewardlaw rewardvarianceproxy) (htotal : 0 < ((cumulativesuccessorglobalreturnvarianceproxy mdp episodes rounds rewardbound rewardvarianceproxy : nnreal) : real)) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : source.trajectorymeasure (sou…","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_weightedExpected_to_realized_successor_average_regret_transport","label":"trajectoryMeasure_weightedExpected_to_realized_successor_average_regret_transport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_weightedExpected_to_realized_successor_average_regret_transport","description":"theorem trajectoryMeasure_weightedExpected_to_realized_successor_average_regret_transport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp e…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-03683bf51bf7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7896,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:340"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_weightedExpected_to_realized_successor_average_regret_transport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : forall t, 0 < episodes t) (rewardBound rewardVarianceProxy : NNReal) (hreward…","missing":[],"search":"trajectorymeasure_weightedexpected_to_realized_successor_average_regret_transport banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_weightedexpected_to_realized_successor_average_regret_transport theorem trajectorymeasure_weightedexpected_to_realized_successor_average_regret_transport {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} [standardborelspace state] [standardborelspace action] [forall n, standardborelspace (heterogeneousstochasticepisodebatchprefix mdp episodes n)] [forall n, nonempty (heterogeneousstochasticepisodebatchprefix mdp episodes n)] [forall n, standardborelspace (stochasticepisodebatch mdp (episodes n))] [forall n, nonempty (stochasticepisodebatch mdp (episodes n))] [standardborelspace (heterogeneousstochasticepisodebatchtrajectory mdp episodes)] (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (rounds : nat) (hrounds : 0 < rounds) (hepisodes : forall t, 0 < episodes t) (rewardbound rewardvarianceproxy : nnreal) (hrewardbound : forall state action, |mdp.reward state action| <= (rewardbound : real)) (law : source.rewardsource.uniformsubgaussianrewardlaw rewardvarianceproxy) (htotal : 0 < ((cumulativesuccessorglobalreturnvarianceproxy mdp episodes rounds rewardbound rewardv…","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel_occupancySelectedRadiusRemaining_eq","label":"stochasticAllCoordinateEmpiricalFiniteBatchModel_occupancySelectedRadiusRemaining_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel_occupancySelectedRadiusRemaining_eq","description":"A sampled stochastic empirical model has the two supplied fixed radii.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-3dfb70062640","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7897,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:458"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel_occupancySelectedRadiusRemaining_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) : let model := mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget model.plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * model.plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState = (mdp.horizon : Real) * (2 * (rewardBudget + transitionBudget))","missing":[],"search":"stochasticallcoordinateempiricalfinitebatchmodel_occupancyselectedradiusremaining_eq banditrlproof.finitehorizonrl.mdp.stochasticallcoordinateempiricalfinitebatchmodel_occupancyselectedradiusremaining_eq a sampled stochastic empirical model has the two supplied fixed radii. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_successorPolicyAt_eq_sampledPlanExploratoryPolicy","label":"heterogeneousExploratorySource_successorPolicyAt_eq_sampledPlanExploratoryPolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_successorPolicyAt_eq_sampledPlanExploratoryPolicy","description":"The sampled table at coordinate `n` is exactly the next source policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-d059918c2344","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7898,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:491"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem heterogeneousExploratorySource_successorPolicyAt_eq_sampledPlanExploratoryPolicy {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (explorationRate : Nat -> NNReal) (hexplorationRate : forall t, explorationRate t <= 1) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) : let source := heterogeneousExploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate source.successorPolicyAt trajectory n = ((trajectory n).sampledEmpiricalOptimisticPolicyTable defaultState (rewardBudget n) (transitionBudget n)).exploratoryPolicy (explorationRa…","missing":[],"search":"heterogeneousexploratorysource_successorpolicyat_eq_sampledplanexploratorypolicy banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneousexploratorysource_successorpolicyat_eq_sampledplanexploratorypolicy the sampled table at coordinate `n` is exactly the next source policy. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_successorPolicyAt_expectedRegret_le","label":"heterogeneousExploratorySource_successorPolicyAt_expectedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_successorPolicyAt_expectedRegret_le","description":"One selected causal exploratory policy pays only its local charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-b6616a93a70c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7899,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:511"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem heterogeneousExploratorySource_successorPolicyAt_expectedRegret_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (explorationRate : Nat -> NNReal) (hexplorationRate : forall t, explorationRate t <= 1) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) (rewardBound : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) : let source := heterogeneousExploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate let model := mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel (episodes n) (mdp.sampledE…","missing":[],"search":"heterogeneousexploratorysource_successorpolicyat_expectedregret_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneousexploratorysource_successorpolicyat_expectedregret_le one selected causal exploratory policy pays only its local charge. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningCumulativeBound","label":"selfConsistentScheduledCausalSuccessorPlanningCumulativeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningCumulativeBound","description":"Weighted planning envelope for all causal successor batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-2f450393fc17","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7900,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:550"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalSuccessorPlanningCumulativeBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalsuccessorplanningcumulativebound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsuccessorplanningcumulativebound weighted planning envelope for all causal successor batches. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningAverageBound","label":"selfConsistentScheduledCausalSuccessorPlanningAverageBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningAverageBound","description":"The same planning envelope per actual sampled successor episode.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-a90db9234322","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7901,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:571"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalSuccessorPlanningAverageBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentscheduledcausalsuccessorplanningaveragebound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsuccessorplanningaveragebound the same planning envelope per actual sampled successor episode. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_weightedExpectedSuccessorAverageRegret","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_weightedExpectedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_weightedExpectedSuccessorAverageRegret","description":"The same planning envelope per actual sampled successor episode. -/ noncomputable def selfConsistentScheduledCausalSuccessorPlanningAverageBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real := selfConsistentScheduledCausalSuccessorPlanningCumulativeBound mdp varianceProxy baseVisitFloor rounds / HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-7cc10e778df6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7902,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:587"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_weightedExpectedSuccessorAverageRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) : let episodes := fun t => AdaptiveStochasticEpisodeBatchSource.se…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_optimism_and_weightedexpectedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_optimism_and_weightedexpectedsuccessoraverageregret the same planning envelope per actual sampled successor episode. -/ noncomputable def selfconsistentscheduledcausalsuccessorplanningaveragebound (mdp : mdp state action) (varianceproxy : nnreal) (basevisitfloor : real) (rounds : nat) : real := selfconsistentscheduledcausalsuccessorplanningcumulativebound mdp varianceproxy basevisitfloor rounds / heterogeneousadaptivestochasticepisodebatchsource.successorepisodemass (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) rounds /- the sampled-model event controls the actual exploratory successor policies. every coordinate keeps its own model budgets, next-coordinate exploration rate, and successor batch-size weight. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorReturnBadEvent","label":"selfConsistentScheduledCausalSuccessorReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorReturnBadEvent","description":"Named successor-only return event for the self-consistent causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-cc945147495f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7903,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:754"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalSuccessorReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentscheduledcausalsuccessorreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsuccessorreturnbadevent named successor-only return event for the self-consistent causal source. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelReturnBadEvent","label":"selfConsistentScheduledCausalModelReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelReturnBadEvent","description":"Union of the causal sampled-model and globally centered return events.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-c55e0dca357f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7904,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:771"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentscheduledcausalmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelreturnbadevent union of the causal sampled-model and globally centered return events. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelReturnFailureBudget","label":"selfConsistentScheduledCausalModelReturnFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelReturnFailureBudget","description":"Exact finite-prefix model-plus-return failure budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-8e58b226dc31","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7905,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:789"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalModelReturnFailureBudget (mdp : MDP State Action) (rounds : Nat) (returnDelta : Real) : ENNReal","missing":[],"search":"selfconsistentscheduledcausalmodelreturnfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelreturnfailurebudget exact finite-prefix model-plus-return failure budget. definition compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret","description":"Exact finite-prefix model-plus-return failure budget. -/ noncomputable def selfConsistentScheduledCausalModelReturnFailureBudget (mdp : MDP State Action) (rounds : Nat) (returnDelta : Real) : ENNReal := selfConsistentScheduledCausalModelFailureBudget mdp rounds + ENNReal.ofReal returnDelta /- End-to-end causal finite-prefix theorem. Off one named event, every actual sampled model is optimistic and the weighted reali…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-27084604259d/index.html#decl-1cbfe68218cc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","order":7906,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret.lean:800"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDelta_le_one…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_optimism_and_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_optimism_and_realizedsuccessoraverageregret exact finite-prefix model-plus-return failure budget. -/ noncomputable def selfconsistentscheduledcausalmodelreturnfailurebudget (mdp : mdp state action) (rounds : nat) (returndelta : real) : ennreal := selfconsistentscheduledcausalmodelfailurebudget mdp rounds + ennreal.ofreal returndelta /- end-to-end causal finite-prefix theorem. off one named event, every actual sampled model is optimistic and the weighted realized successor-average behavior regret is controlled by the coordinatewise schedule plus the heterogeneous globally centered return radius. theorem compiled","shard":"modules/f4d04e51161ffd03.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviation","label":"successorSampledReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviation","description":"Dynamic next-coordinate sampled-return deviation on a prefix/batch pair.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-d7947d01c6c7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7907,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorSampledReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (pair : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1))) : Real","missing":[],"search":"successorsampledreturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorsampledreturndeviation dynamic next-coordinate sampled-return deviation on a prefix/batch pair. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorDeviationKernel","label":"successorDeviationKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorDeviationKernel","description":"Conditional kernel of the dynamic heterogeneous successor deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-fe206cc9b55f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7908,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorDeviationKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Kernel (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) Real","missing":[],"search":"successordeviationkernel banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successordeviationkernel conditional kernel of the dynamic heterogeneous successor deviation. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorDeviationKernel_apply","label":"successorDeviationKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorDeviationKernel_apply","description":"Every dynamic deviation fiber is the selected policy's iid statistic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-7e8044f567e9","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7909,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorDeviationKernel_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (history : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) : source.successorDeviationKernel n history = (source.rewardSource.iidStochasticTrajectoryFamilyMeasure (source.successorPolicy n history) initialState (episodes (n + 1))).map (mdp.sampledCumulativeReturnDeviationSum (source.successorPolicy n history) (episodes (n + 1)))","missing":[],"search":"successordeviationkernel_apply banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successordeviationkernel_apply every dynamic deviation fiber is the selected policy's iid statistic law. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviationAt","label":"successorSampledReturnDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviationAt","description":"Dynamic deviation evaluated at successor trajectory coordinate `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-a74c57d25bf7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7910,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorsampledreturndeviationat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorsampledreturndeviationat dynamic deviation evaluated at successor trajectory coordinate `n + 1`. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorSampledReturnDeviationAt","label":"measurable_successorSampledReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorSampledReturnDeviationAt","description":"theorem measurable_successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Measurable (source.successorSampledReturnDeviationAt n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-a4147ade0e07","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7911,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:137"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Measurable (source.successorSampledReturnDeviationAt n)","missing":[],"search":"measurable_successorsampledreturndeviationat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_successorsampledreturndeviationat theorem measurable_successorsampledreturndeviationat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) (n : nat) : measurable (source.successorsampledreturndeviationat n) theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","label":"trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","description":"Conditional law of the heterogeneous dynamic successor deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-d310d46436fd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7912,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:151"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] : condDistrib (source.successorSampledReturnDeviationAt n) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] source.successorDeviationKernel n","missing":[],"search":"trajectorymeasure_conddistrib_successorsampledreturndeviationat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_conddistrib_successorsampledreturndeviationat conditional law of the heterogeneous dynamic successor deviation. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorSampledReturnDeviationAt_eq","label":"condExpKernel_map_successorSampledReturnDeviationAt_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorSampledReturnDeviationAt_eq","description":"Trimmed conditional-expectation kernel form of the same dynamic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-492015ae3adb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7913,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_successorSampledReturnDeviationAt_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] : Filter.Eventually (fun trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes => Measure.map (source.successorSampledReturnDeviationAt n) (condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (HeterogeneousStochasticEpisodeB…","missing":[],"search":"condexpkernel_map_successorsampledreturndeviationat_eq banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.condexpkernel_map_successorsampledreturndeviationat_eq trimmed conditional-expectation kernel form of the same dynamic law. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationPrefixIncrement","label":"sampledReturnDeviationPrefixIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationPrefixIncrement","description":"Prefix-level heterogeneous increment, including coordinate zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-896dc4315766","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7914,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:220"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : (round : Nat) -> HeterogeneousStochasticEpisodeBatchPrefix mdp episodes round -> Real | 0, history => mdp.sampledCumulativeReturnDeviationSum source.initialPolicy (episodes 0) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩) | n + 1, history => source.successorSampledReturnDeviation n (Preorder.frestrictLe₂ (π := fun k : Nat => StochasticEpisodeBatch mdp (episodes k)) (Nat.le_succ n) history, history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) omit [DecidableEq State] [DecidableEq Action] [MeasurableSingletonClass State] [MeasurableSingletonClass Action] [Nonempty State] [Nonempty Action] in theorem measurable_sampledReturnDeviationPrefixIncremen…","missing":[],"search":"sampledreturndeviationprefixincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.sampledreturndeviationprefixincrement prefix-level heterogeneous increment, including coordinate zero. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationPrefixIncrement","label":"measurable_sampledReturnDeviationPrefixIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationPrefixIncrement","description":"theorem measurable_sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationPrefixIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-8f8d7c1155ea","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7915,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:240"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationPrefixIncrement round)","missing":[],"search":"measurable_sampledreturndeviationprefixincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_sampledreturndeviationprefixincrement theorem measurable_sampledreturndeviationprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) (round : nat) : measurable (source.sampledreturndeviationprefixincrement round) theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement","label":"sampledReturnDeviationIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement","description":"Adapted heterogeneous sampled-return increment on the full trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-6e82c9479c6d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7916,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:263"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"sampledreturndeviationincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.sampledreturndeviationincrement adapted heterogeneous sampled-return increment on the full trajectory. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationIncrement","label":"measurable_sampledReturnDeviationIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationIncrement","description":"theorem measurable_sampledReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-8e28f80d3f72","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7917,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:276"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationIncrement round)","missing":[],"search":"measurable_sampledreturndeviationincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_sampledreturndeviationincrement theorem measurable_sampledreturndeviationincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) (round : nat) : measurable (source.sampledreturndeviationincrement round) theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_stronglyAdapted_piLE","label":"sampledReturnDeviationIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_stronglyAdapted_piLE","description":"theorem sampledReturnDeviationIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : StronglyAdapted (Filtration.piLE (X := fun n : Nat => StochasticEpisodeBatch mdp (episodes n))) source.sampledReturnDeviationIncrement","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-97e268011c9e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7918,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledReturnDeviationIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : StronglyAdapted (Filtration.piLE (X := fun n : Nat => StochasticEpisodeBatch mdp (episodes n))) source.sampledReturnDeviationIncrement","missing":[],"search":"sampledreturndeviationincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.sampledreturndeviationincrement_stronglyadapted_pile theorem sampledreturndeviationincrement_stronglyadapted_pile {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) : stronglyadapted (filtration.pile (x := fun n : nat => stochasticepisodebatch mdp (episodes n))) source.sampledreturndeviationincrement theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","label":"sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","description":"Coordinate zero inherits the iid batch MGF at `episodes 0`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-4dce41807ac6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7919,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:306"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledReturnDeviationIncrement_zero_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes 0))] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) : HasSubgaussianMGF (source.sampledReturnDeviationIncrement 0) (mdp.iidSampledCumulativeReturnDeviationVarianceProxy (episodes 0) rewardBound rewardVarianceProxy) source.trajectoryMeasure","missing":[],"search":"sampledreturndeviationincrement_zero_hassubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.sampledreturndeviationincrement_zero_hassubgaussianmgf coordinate zero inherits the iid batch mgf at `episodes 0`. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","label":"sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","description":"Coordinate `n + 1` is conditionally sub-Gaussian at its own batch size.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-e97e56e10ee7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7920,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:339"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProx…","missing":[],"search":"sampledreturndeviationincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.sampledreturndeviationincrement_succ_hascondsubgaussianmgf coordinate `n + 1` is conditionally sub-gaussian at its own batch size. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy","label":"cumulativeSampledReturnDeviationVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy","description":"Sum of coordinate-specific sampled-return variance proxies.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-3ed34e8a95ef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7921,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:457"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSampledReturnDeviationVarianceProxy (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"cumulativesampledreturndeviationvarianceproxy banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesampledreturndeviationvarianceproxy sum of coordinate-specific sampled-return variance proxies. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy_pos","label":"cumulativeSampledReturnDeviationVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy_pos","description":"The heterogeneous total proxy is positive under a positive batch schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-29a048468d56","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7922,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:468"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSampledReturnDeviationVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrounds : 0 < rounds) (hepisodes : forall n, 0 < episodes n) (hhorizon : 0 < mdp.horizon) (hrewardVarianceProxy : 0 < rewardVarianceProxy) : 0 < ((cumulativeSampledReturnDeviationVarianceProxy mdp episodes rounds rewardBound rewardVarianceProxy : NNReal) : Real)","missing":[],"search":"cumulativesampledreturndeviationvarianceproxy_pos banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesampledreturndeviationvarianceproxy_pos the heterogeneous total proxy is positive under a positive batch schedule. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviation","label":"cumulativeSampledReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviation","description":"Cumulative heterogeneous sampled-return deviation through `rounds`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-2732823d8bb2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7923,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:502"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSampledReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativesampledreturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesampledreturndeviation cumulative heterogeneous sampled-return deviation through `rounds`. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","label":"trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","description":"Fixed-round two-sided tail on the heterogeneous causal law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-15405bbedcfc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7924,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:515"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSource.UniformSu…","missing":[],"search":"trajectorymeasure_cumulativesampledreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_cumulativesampledreturndeviation_abs_tail_le fixed-round two-sided tail on the heterogeneous causal law. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.GlobalReturnMeasurability","label":"GlobalReturnMeasurability","kind":"typeclass","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.GlobalReturnMeasurability","description":"Measurability contract for history-selected globally centered returns.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-38a907bdb41d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7925,"meta":[["Kind","typeclass"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:577"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"class GlobalReturnMeasurability {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : Prop where","missing":[],"search":"globalreturnmeasurability banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.globalreturnmeasurability measurability contract for history-selected globally centered returns. typeclass compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviation","label":"successorGlobalReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviation","description":"Dynamic globally centered return on a heterogeneous prefix/batch pair.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-18bf0b668078","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7926,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:591"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (pair : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1))) : Real","missing":[],"search":"successorglobalreturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturndeviation dynamic globally centered return on a heterogeneous prefix/batch pair. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviation","label":"measurable_successorGlobalReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviation","description":"theorem measurable_successorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviation n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-b8ed9edabc6b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7927,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:604"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviation n)","missing":[],"search":"measurable_successorglobalreturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_successorglobalreturndeviation theorem measurable_successorglobalreturndeviation {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) : measurable (source.successorglobalreturndeviation n) theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel","label":"successorGlobalReturnDeviationKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel","description":"Selected conditional kernel of the globally centered successor return.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-786d32725723","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7928,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:613"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviationKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Kernel (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) Real","missing":[],"search":"successorglobalreturndeviationkernel banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturndeviationkernel selected conditional kernel of the globally centered successor return. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel_apply","label":"successorGlobalReturnDeviationKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel_apply","description":"Every global-return kernel fiber is the exact selected iid statistic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-1aa1d360ceda","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7929,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:636"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnDeviationKernel_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) (history : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) : source.successorGlobalReturnDeviationKernel n history = (source.rewardSource.iidStochasticTrajectoryFamilyMeasure (source.successorPolicy n history) initialState (episodes (n + 1))).map (mdp.globalSampledCumulativeReturnDeviationSum (source.successorPolicy n history) initialState (episodes (n + 1)))","missing":[],"search":"successorglobalreturndeviationkernel_apply banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturndeviationkernel_apply every global-return kernel fiber is the exact selected iid statistic law. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationAt","label":"successorGlobalReturnDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationAt","description":"Globally centered return evaluated at successor coordinate `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-712062a7e55e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7930,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:675"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorglobalreturndeviationat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturndeviationat globally centered return evaluated at successor coordinate `n + 1`. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviationAt","label":"measurable_successorGlobalReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviationAt","description":"theorem measurable_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviationAt n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-4e73a7c8ec17","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7931,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:688"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) : Measurable (source.successorGlobalReturnDeviationAt n)","missing":[],"search":"measurable_successorglobalreturndeviationat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_successorglobalreturndeviationat theorem measurable_successorglobalreturndeviationat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (n : nat) : measurable (source.successorglobalreturndeviationat n) theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","label":"trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","description":"Dynamic conditional law of the globally centered successor return.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-83f7fc930023","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7932,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:702"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] : condDistrib (source.successorGlobalReturnDeviationAt n) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] source.successorGlobalReturnDeviationKernel n","missing":[],"search":"trajectorymeasure_conddistrib_successorglobalreturndeviationat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_conddistrib_successorglobalreturndeviationat dynamic conditional law of the globally centered successor return. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorGlobalReturnDeviationAt_eq","label":"condExpKernel_map_successorGlobalReturnDeviationAt_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorGlobalReturnDeviationAt_eq","description":"Trimmed conditional-expectation kernel form of the global-return law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-62d4a1662b1a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7933,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:733"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_successorGlobalReturnDeviationAt_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] : Filter.Eventually (fun trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes => Measure.map (source.successorGlobalReturnDeviationAt n) (condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace…","missing":[],"search":"condexpkernel_map_successorglobalreturndeviationat_eq banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.condexpkernel_map_successorglobalreturndeviationat_eq trimmed conditional-expectation kernel form of the global-return law. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnPrefixIncrement","label":"successorGlobalReturnPrefixIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnPrefixIncrement","description":"Prefix process with coordinate zero uncharged and successors globally centered.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-4569c82843cc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7934,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:769"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : (round : Nat) -> HeterogeneousStochasticEpisodeBatchPrefix mdp episodes round -> Real | 0, _history => 0 | n + 1, history => source.successorGlobalReturnDeviation n (Preorder.frestrictLe₂ (π := fun k : Nat => StochasticEpisodeBatch mdp (episodes k)) (Nat.le_succ n) history, history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) omit [DecidableEq State] [DecidableEq Action] [MeasurableSingletonClass State] [MeasurableSingletonClass Action] [Nonempty State] [Nonempty Action] in theorem measurable_successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Na…","missing":[],"search":"successorglobalreturnprefixincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturnprefixincrement prefix process with coordinate zero uncharged and successors globally centered. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnPrefixIncrement","label":"measurable_successorGlobalReturnPrefixIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnPrefixIncrement","description":"theorem measurable_successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (round : Nat) : Measurable (source.successorGlobalReturnPrefixIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-42670baa0810","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7935,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:787"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorGlobalReturnPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (round : Nat) : Measurable (source.successorGlobalReturnPrefixIncrement round)","missing":[],"search":"measurable_successorglobalreturnprefixincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_successorglobalreturnprefixincrement theorem measurable_successorglobalreturnprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] (round : nat) : measurable (source.successorglobalreturnprefixincrement round) theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement","label":"successorGlobalReturnIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement","description":"noncomputable def successorGlobalReturnIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-7ceac0439b1a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7936,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:808"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorGlobalReturnIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorglobalreturnincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturnincrement noncomputable def successorglobalreturnincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) (round : nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp episodes) : real definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_stronglyAdapted_piLE","label":"successorGlobalReturnIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_stronglyAdapted_piLE","description":"theorem successorGlobalReturnIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] : StronglyAdapted (Filtration.piLE (X := fun n : Nat => StochasticEpisodeBatch mdp (episodes n))) source.successorGlobalR…","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-b361e72cce0b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7937,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:821"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] : StronglyAdapted (Filtration.piLE (X := fun n : Nat => StochasticEpisodeBatch mdp (episodes n))) source.successorGlobalReturnIncrement","missing":[],"search":"successorglobalreturnincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturnincrement_stronglyadapted_pile theorem successorglobalreturnincrement_stronglyadapted_pile {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) [source.globalreturnmeasurability] : stronglyadapted (filtration.pile (x := fun n : nat => stochasticepisodebatch mdp (episodes n))) source.successorglobalreturnincrement theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","label":"successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","description":"Every successor global-return coordinate has its selected conditional MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-801bf3009064","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7938,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:838"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSource.UniformSubga…","missing":[],"search":"successorglobalreturnincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.successorglobalreturnincrement_succ_hascondsubgaussianmgf every successor global-return coordinate has its selected conditional mgf. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy","label":"cumulativeSuccessorGlobalReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy","description":"Zero at coordinate zero plus heterogeneous global proxies at successors.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-0847e05c991d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7939,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:957"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSuccessorGlobalReturnVarianceProxy (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"cumulativesuccessorglobalreturnvarianceproxy banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturnvarianceproxy zero at coordinate zero plus heterogeneous global proxies at successors. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_pos","label":"cumulativeSuccessorGlobalReturnVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_pos","description":"Positive successor count and coordinate proxy imply a positive total proxy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-a65cdf3d5c3d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7940,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:971"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSuccessorGlobalReturnVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrounds : 0 < rounds) (hepisodes : forall n, 0 < episodes n) (hhorizon : 0 < mdp.horizon) (hrewardVarianceProxy : 0 < rewardVarianceProxy) : 0 < ((cumulativeSuccessorGlobalReturnVarianceProxy mdp episodes rounds rewardBound rewardVarianceProxy : NNReal) : Real)","missing":[],"search":"cumulativesuccessorglobalreturnvarianceproxy_pos banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturnvarianceproxy_pos positive successor count and coordinate proxy imply a positive total proxy. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation","label":"cumulativeSuccessorGlobalReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation","description":"Cumulative globally centered deviation over successor coordinates only.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-441761d54021","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7941,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:1016"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSuccessorGlobalReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativesuccessorglobalreturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.cumulativesuccessorglobalreturndeviation cumulative globally centered deviation over successor coordinates only. definition compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","label":"trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","description":"Fixed-round successor-only globally centered two-sided tail.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-5454e09a4c09","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7942,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:1029"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound :…","missing":[],"search":"trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le fixed-round successor-only globally centered two-sided tail. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","description":"Concrete self-consistent schedule wrapper for the heterogeneous tail.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-6550352dcc44","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7943,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:1119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (hrounds : 0 < rounds) (hhorizon : 0 < mdp.horizon) (hrewardVarianceProxy : 0 < rewardVarianceProxy) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTab…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_cumulativesampledreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_cumulativesampledreturndeviation_abs_tail_le concrete self-consistent schedule wrapper for the heterogeneous tail. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","description":"Self-consistent successor-only globally centered heterogeneous tail.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-bc480c75f165/index.html#decl-373130f47d95","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","order":7944,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration.lean:1166"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (hrounds : 0 < rounds) (hhorizon : 0 < mdp.horizon) (hrewardVarianceProxy : 0 < rewardVarianceProxy) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource in…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_cumulativesuccessorglobalreturndeviation_abs_tail_le self-consistent successor-only globally centered heterogeneous tail. theorem compiled","shard":"modules/fc32965a64f21720.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousStochasticEpisodeBatchTrajectory","label":"HeterogeneousStochasticEpisodeBatchTrajectory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousStochasticEpisodeBatchTrajectory","description":"A trajectory whose complete stochastic batch size may vary by coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-1773a74ba990","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7945,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev HeterogeneousStochasticEpisodeBatchTrajectory (mdp : MDP State Action) (episodes : Nat -> Nat)","missing":[],"search":"heterogeneousstochasticepisodebatchtrajectory banditrlproof.finitehorizonrl.heterogeneousstochasticepisodebatchtrajectory a trajectory whose complete stochastic batch size may vary by coordinate. abbreviation compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousStochasticEpisodeBatchPrefix","label":"HeterogeneousStochasticEpisodeBatchPrefix","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousStochasticEpisodeBatchPrefix","description":"A finite dependent batch history through coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-799618d6e363","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7946,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev HeterogeneousStochasticEpisodeBatchPrefix (mdp : MDP State Action) (episodes : Nat -> Nat) (n : Nat)","missing":[],"search":"heterogeneousstochasticepisodebatchprefix banditrlproof.finitehorizonrl.heterogeneousstochasticepisodebatchprefix a finite dependent batch history through coordinate `n`. abbreviation compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource","label":"HeterogeneousAdaptiveStochasticEpisodeBatchSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource","description":"An adaptive source of stochastic episode batches with coordinate-dependent batch sizes.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-5a392d59ae77","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7947,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure HeterogeneousAdaptiveStochasticEpisodeBatchSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat -> Nat) where","missing":[],"search":"heterogeneousadaptivestochasticepisodebatchsource banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource an adaptive source of stochastic episode batches with coordinate-dependent batch sizes. structure compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure","label":"trajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure","description":"The one-process dependent Ionescu-Tulcea trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-e440321659f8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7948,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def trajectoryMeasure {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : Measure (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"trajectorymeasure banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure the one-process dependent ionescu-tulcea trajectory law. definition compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_eval_zero","label":"trajectoryMeasure_map_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_eval_zero","description":"Coordinate zero has the configured initial stochastic batch law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-5d57450b3ffc","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7949,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_map_eval_zero {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : source.trajectoryMeasure.map (Function.eval 0) = source.rewardSource.iidStochasticTrajectoryFamilyMeasure source.initialPolicy initialState (episodes 0)","missing":[],"search":"trajectorymeasure_map_eval_zero banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_map_eval_zero coordinate zero has the configured initial stochastic batch law. theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_prefix_compProd","label":"trajectoryMeasure_prefix_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_prefix_compProd","description":"Every prefix/next marginal has the configured causal `compProd` law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-ff191bcb3295","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7950,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_prefix_compProd {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : source.trajectoryMeasure.map (Preorder.frestrictLe n) ⊗ₘ source.batchKernel n = source.trajectoryMeasure.map (fun trajectory => (Preorder.frestrictLe n trajectory, trajectory (n + 1)))","missing":[],"search":"trajectorymeasure_prefix_compprod banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_prefix_compprod every prefix/next marginal has the configured causal `compprod` law. theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_nextBatch","label":"trajectoryMeasure_condDistrib_nextBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_nextBatch","description":"The regular conditional law of coordinate `n + 1` is its source kernel.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-645bcd42981d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7951,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:164"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_nextBatch {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] : condDistrib (fun trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes => trajectory (n + 1)) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] source.batchKernel n","missing":[],"search":"trajectorymeasure_conddistrib_nextbatch banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_conddistrib_nextbatch the regular conditional law of coordinate `n + 1` is its source kernel. theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_prefix_projective","label":"trajectoryMeasure_map_prefix_projective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_prefix_projective","description":"Finite marginals of the causal law form a projective family.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-8e1c3906795f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7952,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_map_prefix_projective {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) {m n : Nat} (hmn : m <= n) : (source.trajectoryMeasure.map (Preorder.frestrictLe n)).map (Preorder.frestrictLe₂ (π := fun k : Nat => StochasticEpisodeBatch mdp (episodes k)) hmn) = source.trajectoryMeasure.map (Preorder.frestrictLe m)","missing":[],"search":"trajectorymeasure_map_prefix_projective banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_map_prefix_projective finite marginals of the causal law form a projective family. theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousLatestBatch","label":"heterogeneousLatestBatch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousLatestBatch","description":"The latest batch in a dependent finite history.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-cc0b7fe05a9c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7953,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def heterogeneousLatestBatch {mdp : MDP State Action} {episodes : Nat -> Nat} {n : Nat} (history : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) : StochasticEpisodeBatch mdp (episodes n)","missing":[],"search":"heterogeneouslatestbatch banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneouslatestbatch the latest batch in a dependent finite history. definition compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_heterogeneousLatestBatch","label":"measurable_heterogeneousLatestBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_heterogeneousLatestBatch","description":"theorem measurable_heterogeneousLatestBatch {mdp : MDP State Action} {episodes : Nat -> Nat} {n : Nat} : Measurable (heterogeneousLatestBatch : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n -> StochasticEpisodeBatch mdp (episodes n))","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-5c65617e80d5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7954,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:218"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_heterogeneousLatestBatch {mdp : MDP State Action} {episodes : Nat -> Nat} {n : Nat} : Measurable (heterogeneousLatestBatch : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n -> StochasticEpisodeBatch mdp (episodes n))","missing":[],"search":"measurable_heterogeneouslatestbatch banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_heterogeneouslatestbatch theorem measurable_heterogeneouslatestbatch {mdp : mdp state action} {episodes : nat -> nat} {n : nat} : measurable (heterogeneouslatestbatch : heterogeneousstochasticepisodebatchprefix mdp episodes n -> stochasticepisodebatch mdp (episodes n)) theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousSuccessorTable","label":"heterogeneousSuccessorTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousSuccessorTable","description":"The actual-sampled optimistic table selected from dependent history.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-582d4162bda4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7955,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:228"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def heterogeneousSuccessorTable {mdp : MDP State Action} {episodes : Nat -> Nat} (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (n : Nat) (history : HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"heterogeneoussuccessortable banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneoussuccessortable the actual-sampled optimistic table selected from dependent history. definition compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_heterogeneousSuccessorTable","label":"measurable_heterogeneousSuccessorTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_heterogeneousSuccessorTable","description":"theorem measurable_heterogeneousSuccessorTable {mdp : MDP State Action} {episodes : Nat -> Nat} (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (n : Nat) : Measurable (heterogeneousSuccessorTable (mdp := mdp) (episodes := episodes) defaultState rewardBudget transitionBudget n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-22a04813ee29","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7956,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:237"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_heterogeneousSuccessorTable {mdp : MDP State Action} {episodes : Nat -> Nat} (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (n : Nat) : Measurable (heterogeneousSuccessorTable (mdp := mdp) (episodes := episodes) defaultState rewardBudget transitionBudget n)","missing":[],"search":"measurable_heterogeneoussuccessortable banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_heterogeneoussuccessortable theorem measurable_heterogeneoussuccessortable {mdp : mdp state action} {episodes : nat -> nat} (defaultstate : state) (rewardbudget transitionbudget : nat -> real) (n : nat) : measurable (heterogeneoussuccessortable (mdp := mdp) (episodes := episodes) defaultstate rewardbudget transitionbudget n) theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource","label":"heterogeneousExploratorySource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource","description":"Actual-sampled exploratory source with coordinate-dependent batch sizes, budgets, and exploration rates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-6d83389e7ad5","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7957,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:252"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def heterogeneousExploratorySource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat -> Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Nat -> Real) (explorationRate : Nat -> NNReal) (hexplorationRate : forall n, explorationRate n <= 1) : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"heterogeneousexploratorysource banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.heterogeneousexploratorysource actual-sampled exploratory source with coordinate-dependent batch sizes, budgets, and exploration rates. definition compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource","label":"selfConsistentScheduledCausalSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource","description":"The genuinely causal round-varying self-consistent scheduled source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-55b8cf615c71","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7958,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:292"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledCausalSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState (fun n => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n)","missing":[],"search":"selfconsistentscheduledcausalsource banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource the genuinely causal round-varying self-consistent scheduled source. definition compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_exactLaws_and_projective","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_exactLaws_and_projective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_exactLaws_and_projective","description":"Exact conditional and projective laws of the self-consistent causal source.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-1d5becc5aea0/index.html#decl-78ca158d77df","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","order":7959,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource.lean:323"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_exactLaws_and_projective (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) : let episodes := fun n => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure.map (Function.eval 0) = rewardSource.iidStochasticTrajectoryFamilyMeasure (initialTable.exploratoryPolicy (AdaptiveEpisodeBatchSource.decayingExplorationRate 0) (AdaptiveEpisodeBatchSource.decayingExplo…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_exactlaws_and_projective banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_exactlaws_and_projective exact conditional and projective laws of the self-consistent causal source. theorem compiled","shard":"modules/d1a06f4cd913d098.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_absoluteRealizedSuccessorAverageRegret","label":"exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_absoluteRealizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_absoluteRealizedSuccessorAverageRegret","description":"Actual-sampled optimism and an absolute realized-regret explicit rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-e24bc637b8a4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7960,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_absoluteRealizedSuccessorAverageRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (n : Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpisodeBatchSource.decayingE…","missing":[],"search":"exploratorysource_trajectorymeasure_selfconsistentscheduledexplicitrate_allcoordinateconfidence_optimism_and_absoluterealizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_selfconsistentscheduledexplicitrate_allcoordinateconfidence_optimism_and_absoluterealizedsuccessoraverageregret actual-sampled optimism and an absolute realized-regret explicit rate. theorem compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.SelfConsistentScheduledStochasticWindowSpace","label":"SelfConsistentScheduledStochasticWindowSpace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.SelfConsistentScheduledStochasticWindowSpace","description":"One complete actual-sampled self-consistent experiment at each coordinate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-c1a89a65834c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7961,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:290"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev SelfConsistentScheduledStochasticWindowSpace (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real)","missing":[],"search":"selfconsistentscheduledstochasticwindowspace banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticwindowspace one complete actual-sampled self-consistent experiment at each coordinate. abbreviation compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticWindowSource","label":"selfConsistentScheduledStochasticWindowSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticWindowSource","description":"The actual-sampled self-consistent source at schedule coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-a2f52a3f19b4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7962,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:298"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledStochasticWindowSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : AdaptiveStochasticEpisodeBatchSource mdp initialState (AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n)","missing":[],"search":"selfconsistentscheduledstochasticwindowsource banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticwindowsource the actual-sampled self-consistent source at schedule coordinate `n`. definition compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticWindowMeasure","label":"selfConsistentScheduledStochasticWindowMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticWindowMeasure","description":"The actual-sampled self-consistent trajectory law at coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-0ac858967326","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7963,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledStochasticWindowMeasure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Measure (StochasticEpisodeBatchTrajectory mdp (AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n))","missing":[],"search":"selfconsistentscheduledstochasticwindowmeasure banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticwindowmeasure the actual-sampled self-consistent trajectory law at coordinate `n`. definition compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure","label":"selfConsistentScheduledStochasticCommonMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure","description":"Independent product coupling of the complete self-consistent window laws.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-1971430057e4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7964,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:347"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledStochasticCommonMeasure (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) : Measure (SelfConsistentScheduledStochasticWindowSpace mdp varianceProxy baseVisitFloor)","missing":[],"search":"selfconsistentscheduledstochasticcommonmeasure banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticcommonmeasure independent product coupling of the complete self-consistent window laws. definition compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure_map_eval","label":"selfConsistentScheduledStochasticCommonMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure_map_eval","description":"Each common-space coordinate has exactly its scheduled trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-c4d0a20d065a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7965,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:374"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledStochasticCommonMeasure_map_eval (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : (selfConsistentScheduledStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).map (fun omega => omega n) = selfConsistentScheduledStochasticWindowMeasure mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n","missing":[],"search":"selfconsistentscheduledstochasticcommonmeasure_map_eval banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticcommonmeasure_map_eval each common-space coordinate has exactly its scheduled trajectory law. theorem compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticRealizedRegretProcess","label":"selfConsistentScheduledStochasticRealizedRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticRealizedRegretProcess","description":"Scheduled actual-sampled realized successor-average regret.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-d2ac9dcc4dcd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7966,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:390"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledStochasticRealizedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) (omega : SelfConsistentScheduledStochasticWindowSpace mdp varianceProxy baseVisitFloor) : Real","missing":[],"search":"selfconsistentscheduledstochasticrealizedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticrealizedregretprocess scheduled actual-sampled realized successor-average regret. definition compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledStochasticRealizedRegretProcess","label":"measurable_selfConsistentScheduledStochasticRealizedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledStochasticRealizedRegretProcess","description":"Every coordinate of the scheduled actual-sampled regret is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-78b6e6b7f194","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7967,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:405"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledStochasticRealizedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Measurable (selfConsistentScheduledStochasticRealizedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n)","missing":[],"search":"measurable_selfconsistentscheduledstochasticrealizedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentscheduledstochasticrealizedregretprocess every coordinate of the scheduled actual-sampled regret is measurable. theorem compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonBadEvent","label":"selfConsistentScheduledStochasticCommonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonBadEvent","description":"Pull the model/global-return union at coordinate `n` to the common space.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-15a9b810def4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7968,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:424"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledStochasticCommonBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Set (SelfConsistentScheduledStochasticWindowSpace mdp varianceProxy baseVisitFloor)","missing":[],"search":"selfconsistentscheduledstochasticcommonbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticcommonbadevent pull the model/global-return union at coordinate `n` to the common space. definition compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure_badEvent_le","label":"selfConsistentScheduledStochasticCommonMeasure_badEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure_badEvent_le","description":"The pulled-back bad event inherits the explicit finite-window failure rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-b9811f818d2f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7969,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:444"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledStochasticCommonMeasure_badEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledStochasticCommonMeasure mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfCo…","missing":[],"search":"selfconsistentscheduledstochasticcommonmeasure_badevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledstochasticcommonmeasure_badevent_le the pulled-back bad event inherits the explicit finite-window failure rate. theorem compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.abs_selfConsistentScheduledStochasticRealizedRegretProcess_le_of_not_mem_badEvent","label":"abs_selfConsistentScheduledStochasticRealizedRegretProcess_le_of_not_mem_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.abs_selfConsistentScheduledStochasticRealizedRegretProcess_le_of_not_mem_badEvent","description":"Outside the pulled-back event, coordinate `n` has the absolute rate bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-7817086f7e81","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7970,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:497"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_selfConsistentScheduledStochasticRealizedRegretProcess_le_of_not_mem_badEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) (omega : SelfConsistentScheduledStochasticWindowSpace mdp varianceProxy baseVisitFloor) (homega : omega ∉ selfConsiste…","missing":[],"search":"abs_selfconsistentscheduledstochasticrealizedregretprocess_le_of_not_mem_badevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.abs_selfconsistentscheduledstochasticrealizedregretprocess_le_of_not_mem_badevent outside the pulled-back event, coordinate `n` has the absolute rate bound. theorem compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_selfConsistentScheduledStochasticCommonMeasure_marginals_and_realizedRegret_tendstoInMeasure_zero","label":"exploratorySource_selfConsistentScheduledStochasticCommonMeasure_marginals_and_realizedRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_selfConsistentScheduledStochasticCommonMeasure_marginals_and_realizedRegret_tendstoInMeasure_zero","description":"Terminal theorem: exact scheduled marginals and convergence in probability of actual-sampled realized successor-average regret to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-e77cad3a946e/index.html#decl-283f52f88276","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","order":7971,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency.lean:536"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_selfConsistentScheduledStochasticCommonMeasure_marginals_and_realizedRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall n, Measurable (selfConsistentScheduledStochasticRealizedRegretProcess mdp initialSta…","missing":[],"search":"exploratorysource_selfconsistentscheduledstochasticcommonmeasure_marginals_and_realizedregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_selfconsistentscheduledstochasticcommonmeasure_marginals_and_realizedregret_tendstoinmeasure_zero terminal theorem: exact scheduled marginals and convergence in probability of actual-sampled realized successor-average regret to zero. theorem compiled","shard":"modules/2c234f58b610d060.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudgetRateEnvelope","label":"selfConsistentScheduledTransitionBudgetRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudgetRateEnvelope","description":"Three times the compiled contraction envelope controls the fixed-point budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-5e758176c03e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7972,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledTransitionBudgetRateEnvelope (mdp : MDP State Action) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledtransitionbudgetrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitionbudgetrateenvelope three times the compiled contraction envelope controls the fixed-point budget. definition compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudgetRateEnvelope_eq","label":"selfConsistentScheduledTransitionBudgetRateEnvelope_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudgetRateEnvelope_eq","description":"The explicit transition-budget envelope is `12 * |State| * horizon / scale^2`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-caa3af9ab9f2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7973,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionBudgetRateEnvelope_eq (mdp : MDP State Action) (n : Nat) : selfConsistentScheduledTransitionBudgetRateEnvelope mdp n = (12 * (Fintype.card State : Real) * (mdp.horizon : Real)) / (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) ^ 2","missing":[],"search":"selfconsistentscheduledtransitionbudgetrateenvelope_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitionbudgetrateenvelope_eq the explicit transition-budget envelope is `12 * |state| * horizon / scale^2`. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_le_rateEnvelope","label":"selfConsistentScheduledTransitionBudget_le_rateEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_le_rateEnvelope","description":"The exact fixed-point transition budget has an explicit scale-squared rate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-5a52f2c48f30","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7974,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:62"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionBudget_le_rateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledTransitionBudget mdp varianceProxy baseVisitFloor n <= selfConsistentScheduledTransitionBudgetRateEnvelope mdp n","missing":[],"search":"selfconsistentscheduledtransitionbudget_le_rateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitionbudget_le_rateenvelope the exact fixed-point transition budget has an explicit scale-squared rate. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretRateEnvelope","label":"selfConsistentScheduledPlanningAverageRegretRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretRateEnvelope","description":"Explicit planning envelope: scale-squared model error plus the exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-95e4273be926","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7975,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledPlanningAverageRegretRateEnvelope (mdp : MDP State Action) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledplanningaverageregretrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledplanningaverageregretrateenvelope explicit planning envelope: scale-squared model error plus the exploration charge. definition compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound_le_rateEnvelope","label":"selfConsistentScheduledPlanningAverageRegretBound_le_rateEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound_le_rateEnvelope","description":"The scheduled planning certificate is bounded by the explicit rate envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-312b8d6b4c28","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7976,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:137"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledPlanningAverageRegretBound_le_rateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledPlanningAverageRegretBound mdp varianceProxy baseVisitFloor n <= selfConsistentScheduledPlanningAverageRegretRateEnvelope mdp n","missing":[],"search":"selfconsistentscheduledplanningaverageregretbound_le_rateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledplanningaverageregretbound_le_rateenvelope the scheduled planning certificate is bounded by the explicit rate envelope. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope","label":"selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope","description":"Closed full realized-rate envelope, including the global return fluctuation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-4fd573d4433b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7977,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:171"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledrealizedsuccessoraverageregretrateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedsuccessoraverageregretrateenvelope closed full realized-rate envelope, including the global return fluctuation. definition compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound_le_rateEnvelope","label":"selfConsistentScheduledRealizedSuccessorAverageRegretBound_le_rateEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound_le_rateEnvelope","description":"The original realized-regret certificate is bounded by the closed rate envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-75e9365d2a19","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7978,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:177"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRealizedSuccessorAverageRegretBound_le_rateEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledRealizedSuccessorAverageRegretBound mdp varianceProxy baseVisitFloor n <= selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope mdp varianceProxy n","missing":[],"search":"selfconsistentscheduledrealizedsuccessoraverageregretbound_le_rateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedsuccessoraverageregretbound_le_rateenvelope the original realized-regret certificate is bounded by the closed rate envelope. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureRateEnvelope","label":"selfConsistentScheduledRealizedFailureRateEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureRateEnvelope","description":"The three confidence shares equal the explicit `3 / (n + 2)` envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-fda135b0eecf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7979,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:196"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledRealizedFailureRateEnvelope (n : Nat) : ENNReal","missing":[],"search":"selfconsistentscheduledrealizedfailurerateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedfailurerateenvelope the three confidence shares equal the explicit `3 / (n + 2)` envelope. definition compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget_eq_rateEnvelope","label":"selfConsistentScheduledRealizedFailureBudget_eq_rateEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget_eq_rateEnvelope","description":"The old three-share budget is exactly the explicit failure-rate envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-d6d8c4632e41","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7980,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:207"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRealizedFailureBudget_eq_rateEnvelope (n : Nat) : selfConsistentScheduledRealizedFailureBudget n = selfConsistentScheduledRealizedFailureRateEnvelope n","missing":[],"search":"selfconsistentscheduledrealizedfailurebudget_eq_rateenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedfailurebudget_eq_rateenvelope the old three-share budget is exactly the explicit failure-rate envelope. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretRateEnvelope_tendsto_zero","label":"selfConsistentScheduledPlanningAverageRegretRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretRateEnvelope_tendsto_zero","description":"The explicit planning-rate envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-308546a32072","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7981,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:226"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledPlanningAverageRegretRateEnvelope_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledPlanningAverageRegretRateEnvelope mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledplanningaverageregretrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledplanningaverageregretrateenvelope_tendsto_zero the explicit planning-rate envelope tends to zero. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","label":"selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","description":"The full explicit realized-regret envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-f099ae7ed13c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7982,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:257"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope mdp varianceProxy) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledrealizedsuccessoraverageregretrateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedsuccessoraverageregretrateenvelope_tendsto_zero the full explicit realized-regret envelope tends to zero. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureRateEnvelope_tendsto_zero","label":"selfConsistentScheduledRealizedFailureRateEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureRateEnvelope_tendsto_zero","description":"The explicit `3 / (n + 2)` failure envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-b3a19a4f3324","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7983,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:275"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRealizedFailureRateEnvelope_tendsto_zero : Tendsto selfConsistentScheduledRealizedFailureRateEnvelope atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledrealizedfailurerateenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedfailurerateenvelope_tendsto_zero the explicit `3 / (n + 2)` failure envelope tends to zero. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledExplicitRateEnvelopes_tendsto_zero","label":"selfConsistentScheduledExplicitRateEnvelopes_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledExplicitRateEnvelopes_tendsto_zero","description":"Explicit failure and realized-regret rates vanish jointly.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-4d7a09831b66","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7984,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledExplicitRateEnvelopes_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (fun n => (selfConsistentScheduledRealizedFailureRateEnvelope n, selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope mdp varianceProxy n)) atTop (nhds (0, 0))","missing":[],"search":"selfconsistentscheduledexplicitrateenvelopes_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledexplicitrateenvelopes_tendsto_zero explicit failure and realized-regret rates vanish jointly. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","label":"exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","description":"Actual-sampled optimism and realized regret with explicit finite-window rates.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-7de2fbb569a4/index.html#decl-398836e63600","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","order":7985,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate.lean:312"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (n : Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpisodeBatchSource.decayingExplorati…","missing":[],"search":"exploratorysource_trajectorymeasure_selfconsistentscheduledexplicitrate_allcoordinateconfidence_optimism_and_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_selfconsistentscheduledexplicitrate_allcoordinateconfidence_optimism_and_realizedsuccessoraverageregret actual-sampled optimism and realized regret with explicit finite-window rates. theorem compiled","shard":"modules/872311f9b2035ad5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_selfConsistentCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_selfConsistentCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_selfConsistentCalibration","description":"Finite-round actual-sampled confidence under the exact self-consistent transition budget. Every selected model is built from its observed batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a072b4ceb8e1/index.html#decl-0019087fea3d","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","order":7986,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_selfConsistentCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_selfconsistentcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_selfconsistentcalibration finite-round actual-sampled confidence under the exact self-consistent transition budget. every selected model is built from its observed batch. theorem compiled","shard":"modules/62f926b73e242815.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_selfConsistentCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_selfConsistentCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_selfConsistentCalibration","description":"The self-consistent per-round certificate sums to the successor exploratory behavior expected-regret bound, including the explicit exploration charge.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a072b4ceb8e1/index.html#decl-1e3b9a2386c3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","order":7987,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret.lean:190"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_selfConsistentCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardD…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_cumulativesuccessorexploratorybehaviorexpectedregret_of_pathsupport_selfconsistentcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_cumulativesuccessorexploratorybehaviorexpectedregret_of_pathsupport_selfconsistentcalibration the self-consistent per-round certificate sums to the successor exploratory behavior expected-regret bound, including the explicit exploration charge. theorem compiled","shard":"modules/62f926b73e242815.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSelfConsistentBudgetAverageBound","label":"adaptiveStochasticSampledEmpiricalOptimisticSelfConsistentBudgetAverageBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSelfConsistentBudgetAverageBound","description":"Planning part of the self-consistent realized average-regret certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a072b4ceb8e1/index.html#decl-10d31667dfdd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","order":7988,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret.lean:304"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def adaptiveStochasticSampledEmpiricalOptimisticSelfConsistentBudgetAverageBound (mdp : MDP State Action) (episodes : Nat) (countDelta visitFloor : Real) (explorationRate : NNReal) (rewardBound rewardBudget : Real) : Real","missing":[],"search":"adaptivestochasticsampledempiricaloptimisticselfconsistentbudgetaveragebound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticselfconsistentbudgetaveragebound planning part of the self-consistent realized average-regret certificate. definition compiled","shard":"modules/62f926b73e242815.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_selfConsistentBudgetAverageBound","label":"adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_selfConsistentBudgetAverageBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_selfConsistentBudgetAverageBound","description":"The exact occupancy sum closes to the self-consistent average bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a072b4ceb8e1/index.html#decl-5c20c9d6b760","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","order":7989,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret.lean:316"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_selfConsistentBudgetAverageBound {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) (defaultState : State) (countDelta visitFloor rewardBound rewardBudget : Real) (explorationRate : NNReal) (rounds : Nat) (hrounds : 0 < rounds) : (adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum (mdp := mdp) (initialState := initialState) (episodes := episodes) trajectory defaultState rewardBudget (uniformFloorStochasticSelfConsistentTransitionBudget mdp episodes countDelta visitFloor rewardBound rewardBudget) rounds + (rounds : Real) * exploratoryBehaviorRegretCharge mdp explorationRate rewardBound) / (rounds : Real) = adaptiveStochasticSampledEmpiricalOptimisticSelfConsistentBudgetAverageBo…","missing":[],"search":"adaptivestochasticsampledempiricaloptimistic_occupancyandchargeaverage_eq_selfconsistentbudgetaveragebound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimistic_occupancyandchargeaverage_eq_selfconsistentbudgetaveragebound the exact occupancy sum closes to the self-consistent average bound. theorem compiled","shard":"modules/62f926b73e242815.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_selfConsistentBudgetRealizedSuccessorAverageRegret_of_pathSupport_selfConsistentCalibration","label":"exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_selfConsistentBudgetRealizedSuccessorAverageRegret_of_pathSupport_selfConsistentCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_selfConsistentBudgetRealizedSuccessorAverageRegret_of_pathSupport_selfConsistentCalibration","description":"Actual sampled-model confidence and globally centered realized successor regret under the shrinking self-consistent transition budget.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-a072b4ceb8e1/index.html#decl-75053d164bf6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","order":7990,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret.lean:346"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_selfConsistentBudgetRealizedSuccessorAverageRegret_of_pathSupport_selfConsistentCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (hmodelTotal : 0 < ((((episodes : NNReal) * varianc…","missing":[],"search":"exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_selfconsistentbudgetrealizedsuccessoraverageregret_of_pathsupport_selfconsistentcalibration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_finiteround_allcoordinateconfidence_optimism_and_selfconsistentbudgetrealizedsuccessoraverageregret_of_pathsupport_selfconsistentcalibration actual sampled-model confidence and globally centered realized successor regret under the shrinking self-consistent transition budget. theorem compiled","shard":"modules/62f926b73e242815.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentCountShrinkEpisodeThreshold","label":"selfConsistentCountShrinkEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.selfConsistentCountShrinkEpisodeThreshold","description":"Explicit episode threshold making the normalized count radius smaller than `scale⁻²`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-b0488a63e312","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7991,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentCountShrinkEpisodeThreshold (mdp : MDP State Action) (rounds : Nat) (delta visitFloor scale : Real) : Real","missing":[],"search":"selfconsistentcountshrinkepisodethreshold banditrlproof.finitehorizonrl.selfconsistentcountshrinkepisodethreshold explicit episode threshold making the normalized count radius smaller than `scale⁻²`. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentRewardShrinkEpisodeThreshold","label":"selfConsistentRewardShrinkEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.selfConsistentRewardShrinkEpisodeThreshold","description":"Explicit episode threshold making the uniform sampled-reward radius smaller than `scale⁻²`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-b7149093320f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7992,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentRewardShrinkEpisodeThreshold (mdp : MDP State Action) (rounds : Nat) (varianceProxy : NNReal) (delta visitFloor scale : Real) : Real","missing":[],"search":"selfconsistentrewardshrinkepisodethreshold banditrlproof.finitehorizonrl.selfconsistentrewardshrinkepisodethreshold explicit episode threshold making the uniform sampled-reward radius smaller than `scale⁻²`. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_scale_sq_of_threshold","label":"simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_scale_sq_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_scale_sq_of_threshold","description":"The count threshold gives a scale-squared normalized count-radius bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-0c6b9aec2699","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7993,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_scale_sq_of_threshold (mdp : MDP State Action) {rounds episodes : Nat} {delta visitFloor scale : Real} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hscale : 0 < scale) (hthreshold : selfConsistentCountShrinkEpisodeThreshold mdp rounds delta visitFloor scale < (episodes : Real)) : simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFloor / scale ^ 2","missing":[],"search":"simultaneouscountconfidenceradius_lt_episodes_mul_visitfloor_div_scale_sq_of_threshold banditrlproof.finitehorizonrl.simultaneouscountconfidenceradius_lt_episodes_mul_visitfloor_div_scale_sq_of_threshold the count threshold gives a scale-squared normalized count-radius bound. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousRewardSumConfidenceRadius_lt_episodes_mul_visitFloor_div_two_scale_sq_of_threshold","label":"simultaneousRewardSumConfidenceRadius_lt_episodes_mul_visitFloor_div_two_scale_sq_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.simultaneousRewardSumConfidenceRadius_lt_episodes_mul_visitFloor_div_two_scale_sq_of_threshold","description":"The reward threshold gives a half-scale-squared sampled-reward-sum bound.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-c095b9e9a52b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7994,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:117"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousRewardSumConfidenceRadius_lt_episodes_mul_visitFloor_div_two_scale_sq_of_threshold (mdp : MDP State Action) {rounds episodes : Nat} (varianceProxy : NNReal) {delta visitFloor scale : Real} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hscale : 0 < scale) (hthreshold : selfConsistentRewardShrinkEpisodeThreshold mdp rounds varianceProxy delta visitFloor scale < (episodes : Real)) : MDP.MeanCompatibleRewardKernel.simultaneousRewardSumConfidenceRadius mdp episodes varianceProxy (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFloor / (2 * scale ^ 2)","missing":[],"search":"simultaneousrewardsumconfidenceradius_lt_episodes_mul_visitfloor_div_two_scale_sq_of_threshold banditrlproof.finitehorizonrl.simultaneousrewardsumconfidenceradius_lt_episodes_mul_visitfloor_div_two_scale_sq_of_threshold the reward threshold gives a half-scale-squared sampled-reward-sum bound. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodeThreshold","label":"selfConsistentScheduledEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodeThreshold","description":"Maximum of calibration, count-shrinkage, and reward-shrinkage thresholds.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-381ff7398505","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7995,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:193"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledEpisodeThreshold (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledepisodethreshold banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodethreshold maximum of calibration, count-shrinkage, and reward-shrinkage thresholds. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes","label":"selfConsistentScheduledEpisodes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes","description":"Positive natural episode count one step above the explicit maximum threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-a79eec1015c3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7996,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:218"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledEpisodes (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Nat","missing":[],"search":"selfconsistentscheduledepisodes banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes positive natural episode count one step above the explicit maximum threshold. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes_pos","label":"selfConsistentScheduledEpisodes_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes_pos","description":"theorem selfConsistentScheduledEpisodes_pos (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : 0 < selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-0282d4df4e1b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7997,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledEpisodes_pos (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : 0 < selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n","missing":[],"search":"selfconsistentscheduledepisodes_pos banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes_pos theorem selfconsistentscheduledepisodes_pos (mdp : mdp state action) (varianceproxy : nnreal) (basevisitfloor : real) (n : nat) : 0 < selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor n theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodeThreshold_lt_episodes","label":"selfConsistentScheduledEpisodeThreshold_lt_episodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodeThreshold_lt_episodes","description":"theorem selfConsistentScheduledEpisodeThreshold_lt_episodes (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : selfConsistentScheduledEpisodeThreshold mdp varianceProxy baseVisitFloor n < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-fa6af5a1a26c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7998,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:237"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledEpisodeThreshold_lt_episodes (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : selfConsistentScheduledEpisodeThreshold mdp varianceProxy baseVisitFloor n < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real)","missing":[],"search":"selfconsistentscheduledepisodethreshold_lt_episodes banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodethreshold_lt_episodes theorem selfconsistentscheduledepisodethreshold_lt_episodes (mdp : mdp state action) (varianceproxy : nnreal) (basevisitfloor : real) (n : nat) : selfconsistentscheduledepisodethreshold mdp varianceproxy basevisitfloor n < (selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor n : real) theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.exploratoryPathCalibrationEpisodeThreshold_lt_selfConsistentScheduledEpisodes","label":"exploratoryPathCalibrationEpisodeThreshold_lt_selfConsistentScheduledEpisodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.exploratoryPathCalibrationEpisodeThreshold_lt_selfConsistentScheduledEpisodes","description":"The scheduled episode count strictly exceeds the path-calibration threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-77e0b2c83bd7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":7999,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:254"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPathCalibrationEpisodeThreshold_lt_selfConsistentScheduledEpisodes (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : exploratoryPathCalibrationEpisodeThreshold mdp (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n) (AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor mdp baseVisitFloor n) < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real)","missing":[],"search":"exploratorypathcalibrationepisodethreshold_lt_selfconsistentscheduledepisodes banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.exploratorypathcalibrationepisodethreshold_lt_selfconsistentscheduledepisodes the scheduled episode count strictly exceeds the path-calibration threshold. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentCountShrinkEpisodeThreshold_lt_scheduledEpisodes","label":"selfConsistentCountShrinkEpisodeThreshold_lt_scheduledEpisodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentCountShrinkEpisodeThreshold_lt_scheduledEpisodes","description":"The scheduled episode count strictly exceeds the count-shrink threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-d014ebf0cd16","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8000,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentCountShrinkEpisodeThreshold_lt_scheduledEpisodes (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : selfConsistentCountShrinkEpisodeThreshold mdp (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n) (AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor mdp baseVisitFloor n) (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real)","missing":[],"search":"selfconsistentcountshrinkepisodethreshold_lt_scheduledepisodes banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentcountshrinkepisodethreshold_lt_scheduledepisodes the scheduled episode count strictly exceeds the count-shrink threshold. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentRewardShrinkEpisodeThreshold_lt_scheduledEpisodes","label":"selfConsistentRewardShrinkEpisodeThreshold_lt_scheduledEpisodes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentRewardShrinkEpisodeThreshold_lt_scheduledEpisodes","description":"The scheduled episode count strictly exceeds the sampled-reward shrink threshold.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-6249d1dde34e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8001,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:284"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentRewardShrinkEpisodeThreshold_lt_scheduledEpisodes (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : selfConsistentRewardShrinkEpisodeThreshold mdp (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) varianceProxy (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n) (AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor mdp baseVisitFloor n) (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real)","missing":[],"search":"selfconsistentrewardshrinkepisodethreshold_lt_scheduledepisodes banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentrewardshrinkepisodethreshold_lt_scheduledepisodes the scheduled episode count strictly exceeds the sampled-reward shrink threshold. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduled_countMargin_and_halfContraction","label":"selfConsistentScheduled_countMargin_and_halfContraction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduled_countMargin_and_halfContraction","description":"The explicit schedule discharges the strict count margin and half contraction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-6649cc573bd8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8002,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:301"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduled_countMargin_and_halfContraction (mdp : MDP State Action) (witnessState : State) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : simultaneousCountConfidenceRadius mdp (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n) (multiBatchLocalDelta (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n)) < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real) * AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor mdp baseVisitFloor n /\\ uniformFloorStochasticTransitionContraction mdp (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n) (multiBatchLocalDelta (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) (AdaptiveEpisodeBatchSource.vanishingAverag…","missing":[],"search":"selfconsistentscheduled_countmargin_and_halfcontraction banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduled_countmargin_and_halfcontraction the explicit schedule discharges the strict count margin and half contraction. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta","label":"selfConsistentScheduledLocalDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta","description":"Shared per-round confidence share for the self-consistent schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-a9349eb97bdf","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8003,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledLocalDelta (mdp : MDP State Action) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledlocaldelta banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledlocaldelta shared per-round confidence share for the self-consistent schedule. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget","label":"selfConsistentScheduledRewardBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget","description":"Actual sampled-reward coordinate budget under the explicit schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-989c0212e53a","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8004,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:340"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledRewardBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledrewardbudget banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrewardbudget actual sampled-reward coordinate budget under the explicit schedule. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction","label":"selfConsistentScheduledTransitionContraction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction","description":"Actual transition contraction under the explicit schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-4605167e95fd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8005,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:351"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledTransitionContraction (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledtransitioncontraction banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontraction actual transition contraction under the explicit schedule. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget","label":"selfConsistentScheduledTransitionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget","description":"Exact fixed-point transition budget under the explicit schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-2a54792d01b8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8006,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:361"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledTransitionBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledtransitionbudget banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitionbudget exact fixed-point transition budget under the explicit schedule. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledCountRadius_lt_mass_div_scale_sq","label":"selfConsistentScheduledCountRadius_lt_mass_div_scale_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledCountRadius_lt_mass_div_scale_sq","description":"Scheduled count confidence is smaller than visit mass divided by `scale^2`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-f9a1c7e4f11e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8007,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:372"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCountRadius_lt_mass_div_scale_sq (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : simultaneousCountConfidenceRadius mdp (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n) (selfConsistentScheduledLocalDelta mdp n) < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real) * AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor mdp baseVisitFloor n / (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) ^ 2","missing":[],"search":"selfconsistentscheduledcountradius_lt_mass_div_scale_sq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledcountradius_lt_mass_div_scale_sq scheduled count confidence is smaller than visit mass divided by `scale^2`. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardSumRadius_lt_mass_div_two_scale_sq","label":"selfConsistentScheduledRewardSumRadius_lt_mass_div_two_scale_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardSumRadius_lt_mass_div_two_scale_sq","description":"Scheduled reward-sum confidence is smaller than visit mass divided by `2*scale^2`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-945b211f68ef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8008,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:398"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRewardSumRadius_lt_mass_div_two_scale_sq (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : MDP.MeanCompatibleRewardKernel.simultaneousRewardSumConfidenceRadius mdp (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n) varianceProxy (selfConsistentScheduledLocalDelta mdp n) < (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n : Real) * AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor mdp baseVisitFloor n / (2 * (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) ^ 2)","missing":[],"search":"selfconsistentscheduledrewardsumradius_lt_mass_div_two_scale_sq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrewardsumradius_lt_mass_div_two_scale_sq scheduled reward-sum confidence is smaller than visit mass divided by `2*scale^2`. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContractionEnvelope","label":"selfConsistentScheduledTransitionContractionEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContractionEnvelope","description":"A simple deterministic envelope for the scheduled transition contraction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-2f431a2989e6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8009,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:424"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledTransitionContractionEnvelope (mdp : MDP State Action) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledtransitioncontractionenvelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontractionenvelope a simple deterministic envelope for the scheduled transition contraction. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_nonneg","label":"selfConsistentScheduledRewardBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_nonneg","description":"The scheduled sampled-reward budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-3f919a7e0d25","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8010,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:430"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRewardBudget_nonneg (mdp : MDP State Action) (witnessState : State) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= selfConsistentScheduledRewardBudget mdp varianceProxy baseVisitFloor n","missing":[],"search":"selfconsistentscheduledrewardbudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrewardbudget_nonneg the scheduled sampled-reward budget is nonnegative. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_lt_inv_scale_sq","label":"selfConsistentScheduledRewardBudget_lt_inv_scale_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_lt_inv_scale_sq","description":"The scheduled sampled-reward budget is strictly below `scale^-2`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-485360b2add4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8011,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:441"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRewardBudget_lt_inv_scale_sq (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledRewardBudget mdp varianceProxy baseVisitFloor n < 1 / (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) ^ 2","missing":[],"search":"selfconsistentscheduledrewardbudget_lt_inv_scale_sq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrewardbudget_lt_inv_scale_sq the scheduled sampled-reward budget is strictly below `scale^-2`. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_nonneg","label":"selfConsistentScheduledTransitionContraction_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_nonneg","description":"The scheduled transition contraction is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-20c08275dc72","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8012,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:499"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionContraction_nonneg (mdp : MDP State Action) (witnessState : State) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= selfConsistentScheduledTransitionContraction mdp varianceProxy baseVisitFloor n","missing":[],"search":"selfconsistentscheduledtransitioncontraction_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontraction_nonneg the scheduled transition contraction is nonnegative. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_lt_envelope","label":"selfConsistentScheduledTransitionContraction_lt_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_lt_envelope","description":"The scheduled transition contraction is bounded by its `scale^-2` envelope.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-d3fd429ac2b7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8013,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:511"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionContraction_lt_envelope (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledTransitionContraction mdp varianceProxy baseVisitFloor n < selfConsistentScheduledTransitionContractionEnvelope mdp n","missing":[],"search":"selfconsistentscheduledtransitioncontraction_lt_envelope banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontraction_lt_envelope the scheduled transition contraction is bounded by its `scale^-2` envelope. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_lt_one","label":"selfConsistentScheduledTransitionContraction_lt_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_lt_one","description":"The calibration half bound makes the scheduled contraction strictly smaller than one.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-9dc9c84239d2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8014,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:586"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionContraction_lt_one (mdp : MDP State Action) (witnessState : State) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : selfConsistentScheduledTransitionContraction mdp varianceProxy baseVisitFloor n < 1","missing":[],"search":"selfconsistentscheduledtransitioncontraction_lt_one banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontraction_lt_one the calibration half bound makes the scheduled contraction strictly smaller than one. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_nonneg","label":"selfConsistentScheduledTransitionBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_nonneg","description":"The exact scheduled transition budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-77753c25edb8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8015,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:601"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionBudget_nonneg (mdp : MDP State Action) (witnessState : State) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : 0 <= selfConsistentScheduledTransitionBudget mdp varianceProxy baseVisitFloor n","missing":[],"search":"selfconsistentscheduledtransitionbudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitionbudget_nonneg the exact scheduled transition budget is nonnegative. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationScale_sq_tendsto_atTop","label":"decayingExplorationScale_sq_tendsto_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationScale_sq_tendsto_atTop","description":"The square of the decaying-exploration scale tends to infinity.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-c0f6b3ca1e53","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8016,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:621"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem decayingExplorationScale_sq_tendsto_atTop : Tendsto (fun n : Nat => (AdaptiveEpisodeBatchSource.decayingExplorationScale n : Real) ^ 2) atTop atTop","missing":[],"search":"decayingexplorationscale_sq_tendsto_attop banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.decayingexplorationscale_sq_tendsto_attop the square of the decaying-exploration scale tends to infinity. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContractionEnvelope_tendsto_zero","label":"selfConsistentScheduledTransitionContractionEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContractionEnvelope_tendsto_zero","description":"The deterministic contraction envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-fa3aba062ec4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8017,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:638"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionContractionEnvelope_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledTransitionContractionEnvelope mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledtransitioncontractionenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontractionenvelope_tendsto_zero the deterministic contraction envelope tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_tendsto_zero","label":"selfConsistentScheduledRewardBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_tendsto_zero","description":"The actual scheduled sampled-reward budget tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-703bce4b6f3b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8018,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:646"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRewardBudget_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledRewardBudget mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledrewardbudget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrewardbudget_tendsto_zero the actual scheduled sampled-reward budget tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_tendsto_zero","label":"selfConsistentScheduledTransitionContraction_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_tendsto_zero","description":"The actual scheduled transition contraction tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-2b3de120e0a3","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8019,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:664"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionContraction_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledTransitionContraction mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledtransitioncontraction_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitioncontraction_tendsto_zero the actual scheduled transition contraction tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_tendsto_zero","label":"selfConsistentScheduledTransitionBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_tendsto_zero","description":"The exact fixed-point transition budget tends to zero with its contraction.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-6a8c76418c3b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8020,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:683"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledTransitionBudget_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledTransitionBudget mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledtransitionbudget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledtransitionbudget_tendsto_zero the exact fixed-point transition budget tends to zero with its contraction. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound","label":"selfConsistentScheduledPlanningAverageRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound","description":"Planning part of the scheduled self-consistent realized certificate.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-f8c2a44cab7f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8021,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:719"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledPlanningAverageRegretBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledplanningaverageregretbound banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledplanningaverageregretbound planning part of the scheduled self-consistent realized certificate. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound","label":"selfConsistentScheduledRealizedSuccessorAverageRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound","description":"Full realized successor-average regret bound under the explicit schedule.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-c9b68ee5cf41","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8022,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:731"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledRealizedSuccessorAverageRegretBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"selfconsistentscheduledrealizedsuccessoraverageregretbound banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedsuccessoraverageregretbound full realized successor-average regret bound under the explicit schedule. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget","label":"selfConsistentScheduledRealizedFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget","description":"Count, reward, and globally centered return events each consume one share.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-acf0936e959e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8023,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:743"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledRealizedFailureBudget (n : Nat) : ENNReal","missing":[],"search":"selfconsistentscheduledrealizedfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedfailurebudget count, reward, and globally centered return events each consume one share. definition compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound_tendsto_zero","label":"selfConsistentScheduledPlanningAverageRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound_tendsto_zero","description":"The scheduled planning average-regret bound tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-41a2b5e6db00","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8024,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:749"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledPlanningAverageRegretBound_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledPlanningAverageRegretBound mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledplanningaverageregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledplanningaverageregretbound_tendsto_zero the scheduled planning average-regret bound tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledNormalizedSuccessorGlobalReturnRadius_tendsto_zero","label":"selfConsistentScheduledNormalizedSuccessorGlobalReturnRadius_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledNormalizedSuccessorGlobalReturnRadius_tendsto_zero","description":"The scheduled globally centered return radius tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-e196870fc5da","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8025,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:784"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNormalizedSuccessorGlobalReturnRadius_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (fun n => normalizedSuccessorGlobalReturnConfidenceRadius mdp (selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor n) (AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp n) 1 varianceProxy (AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta n)) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednormalizedsuccessorglobalreturnradius_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentschedulednormalizedsuccessorglobalreturnradius_tendsto_zero the scheduled globally centered return radius tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound_tendsto_zero","label":"selfConsistentScheduledRealizedSuccessorAverageRegretBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound_tendsto_zero","description":"The full scheduled realized successor-average bound tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-ad008f34df87","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8026,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:806"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRealizedSuccessorAverageRegretBound_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledRealizedSuccessorAverageRegretBound mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledrealizedsuccessoraverageregretbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedsuccessoraverageregretbound_tendsto_zero the full scheduled realized successor-average bound tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget_tendsto_zero","label":"selfConsistentScheduledRealizedFailureBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget_tendsto_zero","description":"The three-share failure budget tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-fbb296711eb7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8027,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:825"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledRealizedFailureBudget_tendsto_zero : Tendsto selfConsistentScheduledRealizedFailureBudget atTop (nhds 0)","missing":[],"search":"selfconsistentscheduledrealizedfailurebudget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledrealizedfailurebudget_tendsto_zero the three-share failure budget tends to zero. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledFailureAndRealizedBound_tendsto_zero","label":"selfConsistentScheduledFailureAndRealizedBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledFailureAndRealizedBound_tendsto_zero","description":"Failure probability and realized-regret certificates vanish together.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-fc7de5c45acb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8028,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:834"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledFailureAndRealizedBound_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) {baseVisitFloor : Real} (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun n => (selfConsistentScheduledRealizedFailureBudget n, selfConsistentScheduledRealizedSuccessorAverageRegretBound mdp varianceProxy baseVisitFloor n)) atTop (nhds (0, 0))","missing":[],"search":"selfconsistentscheduledfailureandrealizedbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.selfconsistentscheduledfailureandrealizedbound_tendsto_zero failure probability and realized-regret certificates vanish together. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduled_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","label":"exploratorySource_trajectoryMeasure_selfConsistentScheduled_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduled_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","description":"At every schedule index, actual sampled count/reward confidence, optimism, and globally centered realized successor-average regret hold outside three finite bad-event shares. The episode and trajectory spaces may change with `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsel-83f00d79874e/index.html#decl-63f82b40bbed","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","order":8029,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule.lean:858"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_selfConsistentScheduled_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (baseVisitFloor : Real) (n : Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let rounds := AdaptiveEpisodeBatchSource.decayingExplorationRounds mdp…","missing":[],"search":"exploratorysource_trajectorymeasure_selfconsistentscheduled_allcoordinateconfidence_optimism_and_realizedsuccessoraverageregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_trajectorymeasure_selfconsistentscheduled_allcoordinateconfidence_optimism_and_realizedsuccessoraverageregret at every schedule index, actual sampled count/reward confidence, optimism, and globally centered realized successor-average regret hold outside three finite bad-event shares. the episode and trajectory spaces may change with `n`. theorem compiled","shard":"modules/9aab168693a59424.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_rewardSum","label":"measurable_rewardSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_rewardSum","description":"A fixed sampled-reward sum is measurable on raw episode batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-064917c2e104","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8030,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_rewardSum {mdp : MDP State Action} {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable fun batch : EpisodeBatch mdp episodes => batch.rewardSum stage state action","missing":[],"search":"measurable_rewardsum banditrlproof.finitehorizonrl.episodebatch.measurable_rewardsum a fixed sampled-reward sum is measurable on raw episode batches. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalReward","label":"measurable_empiricalReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalReward","description":"A fixed empirical sampled-reward mean is measurable on raw batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-f87658eb6a16","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8031,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:62"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_empiricalReward {mdp : MDP State Action} {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable fun batch : EpisodeBatch mdp episodes => batch.empiricalReward stage state action","missing":[],"search":"measurable_empiricalreward banditrlproof.finitehorizonrl.episodebatch.measurable_empiricalreward a fixed empirical sampled-reward mean is measurable on raw batches. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.stateKernel","label":"stateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.TransitionCountSummary.stateKernel","description":"A fixed empirical next-state law as a kernel in the count summary.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-75123a4f1762","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8032,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stateKernel {mdp : MDP State Action} (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : ProbabilityTheory.Kernel (TransitionCountSummary mdp) State","missing":[],"search":"statekernel banditrlproof.finitehorizonrl.transitioncountsummary.statekernel a fixed empirical next-state law as a kernel in the count summary. definition compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionStateKernel","label":"empiricalTransitionStateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionStateKernel","description":"The raw-batch empirical next-state law as a Markov kernel in the batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-7aa0cde5e0a7","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8033,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:102"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalTransitionStateKernel {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : ProbabilityTheory.Kernel (EpisodeBatch mdp episodes) State","missing":[],"search":"empiricaltransitionstatekernel banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionstatekernel the raw-batch empirical next-state law as a markov kernel in the batch. definition compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionStateKernel_apply","label":"empiricalTransitionStateKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionStateKernel_apply","description":"theorem empiricalTransitionStateKernel_apply {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : empiricalTransitionStateKernel (episodes := episodes) defaultState stage state action batch = batch.empiricalTransitionKernel defaultState stage (state, action)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-3bcae36c2dc0","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8034,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:122"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionStateKernel_apply {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : empiricalTransitionStateKernel (episodes := episodes) defaultState stage state action batch = batch.empiricalTransitionKernel defaultState stage (state, action)","missing":[],"search":"empiricaltransitionstatekernel_apply banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionstatekernel_apply theorem empiricaltransitionstatekernel_apply {mdp : mdp state action} {episodes : nat} (batch : episodebatch mdp episodes) (defaultstate : state) (stage : fin mdp.horizon) (state : state) (action : action) : empiricaltransitionstatekernel (episodes := episodes) defaultstate stage state action batch = batch.empiricaltransitionkernel defaultstate stage (state, action) theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_uncurry_of_forall_measurable","label":"measurable_uncurry_of_forall_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_uncurry_of_forall_measurable","description":"Finite-state coordinatewise measurability yields joint measurability.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-3b6dab7a135b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8035,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_uncurry_of_forall_measurable {Omega : Type*} [MeasurableSpace Omega] (value : Omega -> State -> Real) (hvalue : forall nextState : State, Measurable fun omega => value omega nextState) : Measurable (Function.uncurry value)","missing":[],"search":"measurable_uncurry_of_forall_measurable banditrlproof.finitehorizonrl.episodebatch.measurable_uncurry_of_forall_measurable finite-state coordinatewise measurability yields joint measurability. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalTransitionValue","label":"measurable_empiricalTransitionValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalTransitionValue","description":"The empirical transition integral is measurable when every continuation-value coordinate is measurable in the same raw batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-8bf3984ccb6f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8036,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:163"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_empiricalTransitionValue {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (value : EpisodeBatch mdp episodes -> State -> Real) (hvalue : forall nextState : State, Measurable fun batch => value batch nextState) : Measurable fun batch : EpisodeBatch mdp episodes => ∫ nextState, value batch nextState ∂batch.empiricalTransitionKernel defaultState stage (state, action)","missing":[],"search":"measurable_empiricaltransitionvalue banditrlproof.finitehorizonrl.episodebatch.measurable_empiricaltransitionvalue the empirical transition integral is measurable when every continuation-value coordinate is measurable in the same raw batch. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticQ","label":"measurable_stochasticAllCoordinateEmpiricalOptimisticQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticQ","description":"Every sampled empirical optimistic action-value coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-2222f0eb4d83","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8037,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:187"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stochasticAllCoordinateEmpiricalOptimisticQ (mdp : MDP State Action) (episodes : Nat) (defaultState : State) (rewardBudget transitionBudget : Real) (stage : Fin mdp.horizon) (state : State) (action : Action) (value : EpisodeBatch mdp episodes -> State -> Real) (hvalue : forall nextState : State, Measurable fun batch => value batch nextState) : Measurable fun batch : EpisodeBatch mdp episodes => let model := mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget model.plan.optimisticQ stage (value batch) state action","missing":[],"search":"measurable_stochasticallcoordinateempiricaloptimisticq banditrlproof.finitehorizonrl.mdp.measurable_stochasticallcoordinateempiricaloptimisticq every sampled empirical optimistic action-value coordinate is measurable. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalUpperValueRemaining","label":"measurable_stochasticAllCoordinateEmpiricalUpperValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalUpperValueRemaining","description":"The sampled empirical optimistic value recursion is batch-measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-b7af06a885e4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8038,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stochasticAllCoordinateEmpiricalUpperValueRemaining (mdp : MDP State Action) (episodes : Nat) (defaultState : State) (rewardBudget transitionBudget : Real) : forall (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State), Measurable fun batch : EpisodeBatch mdp episodes => let model := mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget model.plan.upperValueRemaining remaining hremaining state | 0, _hremaining, state => by simp [EstimatedModelPlan.upperValueRemaining] | remaining + 1, hremaining, state => by let stage := mdp.decisionStageRemaining remaining hremaining let value := fun batch : EpisodeBatch mdp episodes => let model := mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget model.plan.upperValueRemaining remaining (by omega…","missing":[],"search":"measurable_stochasticallcoordinateempiricaluppervalueremaining banditrlproof.finitehorizonrl.mdp.measurable_stochasticallcoordinateempiricaluppervalueremaining the sampled empirical optimistic value recursion is batch-measurable. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticActionAt","label":"measurable_stochasticAllCoordinateEmpiricalOptimisticActionAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticActionAt","description":"Every chronological sampled empirical optimistic action is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-36ed16cb7828","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8039,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:256"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stochasticAllCoordinateEmpiricalOptimisticActionAt (mdp : MDP State Action) (episodes : Nat) (defaultState : State) (rewardBudget transitionBudget : Real) (stage : Fin mdp.horizon) (state : State) : Measurable fun batch : EpisodeBatch mdp episodes => let model := mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget model.plan.optimisticActionAt stage state","missing":[],"search":"measurable_stochasticallcoordinateempiricaloptimisticactionat banditrlproof.finitehorizonrl.mdp.measurable_stochasticallcoordinateempiricaloptimisticactionat every chronological sampled empirical optimistic action is measurable. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticPolicyTable","label":"measurable_stochasticAllCoordinateEmpiricalOptimisticPolicyTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticPolicyTable","description":"The complete sampled empirical optimistic policy table is measurable.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-a0baea56a483","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8040,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:291"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stochasticAllCoordinateEmpiricalOptimisticPolicyTable (mdp : MDP State Action) (episodes : Nat) (defaultState : State) (rewardBudget transitionBudget : Real) : Measurable fun batch : EpisodeBatch mdp episodes => fun stage state => (mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget).plan |>.optimisticActionAt stage state","missing":[],"search":"measurable_stochasticallcoordinateempiricaloptimisticpolicytable banditrlproof.finitehorizonrl.mdp.measurable_stochasticallcoordinateempiricaloptimisticpolicytable the complete sampled empirical optimistic policy table is measurable. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.sampledEmpiricalOptimisticPolicyTable","label":"sampledEmpiricalOptimisticPolicyTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.sampledEmpiricalOptimisticPolicyTable","description":"Optimistic table computed from the actual sampled rewards in one batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-303d5eda5156","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8041,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledEmpiricalOptimisticPolicyTable {mdp : MDP State Action} {episodes : Nat} (batch : StochasticEpisodeBatch mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"sampledempiricaloptimisticpolicytable banditrlproof.finitehorizonrl.stochasticepisodebatch.sampledempiricaloptimisticpolicytable optimistic table computed from the actual sampled rewards in one batch. definition compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.measurable_sampledEmpiricalOptimisticPolicyTable","label":"measurable_sampledEmpiricalOptimisticPolicyTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.measurable_sampledEmpiricalOptimisticPolicyTable","description":"The sampled-reward optimistic table is measurable in the stochastic batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-8ab373bf00bb","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8042,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledEmpiricalOptimisticPolicyTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (rewardBudget transitionBudget : Real) : Measurable fun batch : StochasticEpisodeBatch mdp episodes => batch.sampledEmpiricalOptimisticPolicyTable defaultState rewardBudget transitionBudget","missing":[],"search":"measurable_sampledempiricaloptimisticpolicytable banditrlproof.finitehorizonrl.stochasticepisodebatch.measurable_sampledempiricaloptimisticpolicytable the sampled-reward optimistic table is measurable in the stochastic batch. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.latestBatch","label":"latestBatch","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.latestBatch","description":"The latest complete stochastic batch in a finite nonempty prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-4129bdc98a91","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8043,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:336"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def latestBatch {mdp : MDP State Action} {episodes n : Nat} (history : StochasticEpisodeBatchPrefix mdp episodes n) : StochasticEpisodeBatch mdp episodes","missing":[],"search":"latestbatch banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.latestbatch the latest complete stochastic batch in a finite nonempty prefix. definition compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_latestBatch","label":"measurable_latestBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_latestBatch","description":"theorem measurable_latestBatch {mdp : MDP State Action} {episodes n : Nat} : Measurable (latestBatch : StochasticEpisodeBatchPrefix mdp episodes n -> StochasticEpisodeBatch mdp episodes)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-35f79ae01fca","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8044,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:345"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_latestBatch {mdp : MDP State Action} {episodes n : Nat} : Measurable (latestBatch : StochasticEpisodeBatchPrefix mdp episodes n -> StochasticEpisodeBatch mdp episodes)","missing":[],"search":"measurable_latestbatch banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_latestbatch theorem measurable_latestbatch {mdp : mdp state action} {episodes n : nat} : measurable (latestbatch : stochasticepisodebatchprefix mdp episodes n -> stochasticepisodebatch mdp episodes) theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.successorTable","label":"successorTable","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.successorTable","description":"Sampled-reward optimistic table selected from the latest stochastic batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-5c01d61479ef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8045,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:354"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (rewardBudget transitionBudget : Real) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : DeterministicMarkovPolicyTable mdp","missing":[],"search":"successortable banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.successortable sampled-reward optimistic table selected from the latest stochastic batch. definition compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_successorTable","label":"measurable_successorTable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_successorTable","description":"theorem measurable_successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (rewardBudget transitionBudget : Real) (n : Nat) : Measurable (successorTable (mdp := mdp) (episodes := episodes) defaultState rewardBudget transitionBudget n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-a37814bd8d34","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8046,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:362"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorTable {mdp : MDP State Action} {episodes : Nat} (defaultState : State) (rewardBudget transitionBudget : Real) (n : Nat) : Measurable (successorTable (mdp := mdp) (episodes := episodes) defaultState rewardBudget transitionBudget n)","missing":[],"search":"measurable_successortable banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_successortable theorem measurable_successortable {mdp : mdp state action} {episodes : nat} (defaultstate : state) (rewardbudget transitionbudget : real) (n : nat) : measurable (successortable (mdp := mdp) (episodes := episodes) defaultstate rewardbudget transitionbudget n) theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource","label":"exploratorySource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource","description":"Adaptive exploratory source whose successor policy uses actual sampled rewards.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-0f041c9a5516","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8047,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:374"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratorySource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes where","missing":[],"search":"exploratorysource banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource adaptive exploratory source whose successor policy uses actual sampled rewards. definition compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_batchKernel_eq_selectedPolicy_iidLaw","label":"exploratorySource_batchKernel_eq_selectedPolicy_iidLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_batchKernel_eq_selectedPolicy_iidLaw","description":"Every history fiber has the exact iid stochastic law of its selected policy.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardsampledempiricaloptimisticsource/index.html#decl-af0c737c2acd","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","order":8048,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource.lean:408"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_batchKernel_eq_selectedPolicy_iidLaw {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBudget transitionBudget : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : let source := exploratorySource mdp initialState episodes rewardSource initialTable defaultState rewardBudget transitionBudget explorationRate hexplorationRate source.batchKernel n history = rewardSource.iidStochasticTrajectoryFamilyMeasure ((successorTable defaultState rewardBudget transitionBudget n history) |>.exploratoryPolicy explorationRate hexplorationRate) initialState episodes","missing":[],"search":"exploratorysource_batchkernel_eq_selectedpolicy_iidlaw banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exploratorysource_batchkernel_eq_selectedpolicy_iidlaw every history fiber has the exact iid stochastic law of its selected policy. theorem compiled","shard":"modules/13af6ace591b20c3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:ProbabilityTheory.Kernel.retainedInputKernel","label":"retainedInputKernel","kind":"definition","status":"compiled","subtitle":"ProbabilityTheory.Kernel.retainedInputKernel","description":"Pair a kernel output with the input at which the kernel is evaluated.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-b87783ed6c4c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8049,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def retainedInputKernel (kernel : ProbabilityTheory.Kernel Input Output) : ProbabilityTheory.Kernel Input (Input × Output)","missing":[],"search":"retainedinputkernel probabilitytheory.kernel.retainedinputkernel pair a kernel output with the input at which the kernel is evaluated. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:ProbabilityTheory.Kernel.retainedInputKernel_apply","label":"retainedInputKernel_apply","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.Kernel.retainedInputKernel_apply","description":"theorem retainedInputKernel_apply (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (input : Input) : retainedInputKernel kernel input = (kernel input).map (Prod.mk input)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-9d2e32881e65","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8050,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem retainedInputKernel_apply (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (input : Input) : retainedInputKernel kernel input = (kernel input).map (Prod.mk input)","missing":[],"search":"retainedinputkernel_apply probabilitytheory.kernel.retainedinputkernel_apply theorem retainedinputkernel_apply (kernel : probabilitytheory.kernel input output) [probabilitytheory.ismarkovkernel kernel] (input : input) : retainedinputkernel kernel input = (kernel input).map (prod.mk input) theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:ProbabilityTheory.Kernel.condDistrib_pair_ae_eq_retainedInputKernel_of_pair_map_eq_compProd","label":"condDistrib_pair_ae_eq_retainedInputKernel_of_pair_map_eq_compProd","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.Kernel.condDistrib_pair_ae_eq_retainedInputKernel_of_pair_map_eq_compProd","description":"Recover the conditional law of the retained input/output pair.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-43987d677232","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8051,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_pair_ae_eq_retainedInputKernel_of_pair_map_eq_compProd {Sample : Type w} [MeasurableSpace Sample] [StandardBorelSpace (Input × Output)] [Nonempty (Input × Output)] (mu : Measure Sample) [IsFiniteMeasure mu] (condition : Sample → Input) (next : Sample → Output) (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (hcondition : Measurable condition) (hnext : Measurable next) (hpair : mu.map (fun sample => (condition sample, next sample)) = (mu.map condition).compProd kernel) : ProbabilityTheory.condDistrib (fun sample => (condition sample, next sample)) condition mu =ᵐ[ mu.map condition] retainedInputKernel kernel","missing":[],"search":"conddistrib_pair_ae_eq_retainedinputkernel_of_pair_map_eq_compprod probabilitytheory.kernel.conddistrib_pair_ae_eq_retainedinputkernel_of_pair_map_eq_compprod recover the conditional law of the retained input/output pair. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:ProbabilityTheory.Kernel.condDistrib_dynamic_map_ae_eq_of_pair_map_eq_compProd","label":"condDistrib_dynamic_map_ae_eq_of_pair_map_eq_compProd","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.Kernel.condDistrib_dynamic_map_ae_eq_of_pair_map_eq_compProd","description":"Map a statistic which depends jointly on the retained input and output.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-d734815606f4","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8052,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condDistrib_dynamic_map_ae_eq_of_pair_map_eq_compProd {Sample : Type w} [MeasurableSpace Sample] [StandardBorelSpace (Input × Output)] [Nonempty (Input × Output)] {Result : Type*} [MeasurableSpace Result] [StandardBorelSpace Result] [Nonempty Result] (mu : Measure Sample) [IsFiniteMeasure mu] (condition : Sample → Input) (next : Sample → Output) (kernel : ProbabilityTheory.Kernel Input Output) [ProbabilityTheory.IsMarkovKernel kernel] (g : Input × Output → Result) (hcondition : Measurable condition) (hnext : Measurable next) (hg : Measurable g) (hpair : mu.map (fun sample => (condition sample, next sample)) = (mu.map condition).compProd kernel) : ProbabilityTheory.condDistrib (fun sample => g (condition sample, next sample)) condition mu =ᵐ[ mu.map condition] (retainedInputKernel kernel).map g","missing":[],"search":"conddistrib_dynamic_map_ae_eq_of_pair_map_eq_compprod probabilitytheory.kernel.conddistrib_dynamic_map_ae_eq_of_pair_map_eq_compprod map a statistic which depends jointly on the retained input and output. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch","label":"StochasticEpisodeBatch","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch","description":"A complete stochastic-reward episode batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-a83150fab73e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8053,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:168"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev StochasticEpisodeBatch (mdp : MDP State Action) (episodes : Nat)","missing":[],"search":"stochasticepisodebatch banditrlproof.finitehorizonrl.stochasticepisodebatch a complete stochastic-reward episode batch. abbreviation compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatchPrefix","label":"StochasticEpisodeBatchPrefix","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatchPrefix","description":"Finite history through stochastic batch coordinate `n`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-8add2f6d5cef","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8054,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:173"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev StochasticEpisodeBatchPrefix (mdp : MDP State Action) (episodes n : Nat)","missing":[],"search":"stochasticepisodebatchprefix banditrlproof.finitehorizonrl.stochasticepisodebatchprefix finite history through stochastic batch coordinate `n`. abbreviation compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatchTrajectory","label":"StochasticEpisodeBatchTrajectory","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatchTrajectory","description":"Infinite stochastic episode-batch trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-a544a3c0a0e8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8055,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:178"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev StochasticEpisodeBatchTrajectory (mdp : MDP State Action) (episodes : Nat)","missing":[],"search":"stochasticepisodebatchtrajectory banditrlproof.finitehorizonrl.stochasticepisodebatchtrajectory infinite stochastic episode-batch trajectory. abbreviation compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource","label":"AdaptiveStochasticEpisodeBatchSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource","description":"An adaptive source of complete stochastic-reward episode batches. The dynamic measurability field is the only additional policy-selection regularity needed by this route. The exact batch-kernel equality supplies the pointwise conditionally iid law selected by each observed prefix.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-1a97250e78c2","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8056,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure AdaptiveStochasticEpisodeBatchSource (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) where","missing":[],"search":"adaptivestochasticepisodebatchsource banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource an adaptive source of complete stochastic-reward episode batches. the dynamic measurability field is the only additional policy-selection regularity needed by this route. the exact batch-kernel equality supplies the pointwise conditionally iid law selected by each observed prefix. structure compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure","label":"trajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure","description":"The adaptive infinite stochastic batch-trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-e8c9b4729768","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8057,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:224"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def trajectoryMeasure {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : Measure (StochasticEpisodeBatchTrajectory mdp episodes)","missing":[],"search":"trajectorymeasure banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure the adaptive infinite stochastic batch-trajectory law. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_eval_zero","label":"trajectoryMeasure_map_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_eval_zero","description":"Coordinate zero has the configured initial stochastic batch law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-ecd97391ce2c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8058,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:244"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_map_eval_zero {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : source.trajectoryMeasure.map (Function.eval 0) = source.rewardSource.iidStochasticTrajectoryFamilyMeasure source.initialPolicy initialState episodes","missing":[],"search":"trajectorymeasure_map_eval_zero banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_map_eval_zero coordinate zero has the configured initial stochastic batch law. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_prefix_compProd","label":"trajectoryMeasure_prefix_compProd","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_prefix_compProd","description":"The prefix/next-batch marginal has the configured compProd law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-c85742fc900f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8059,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_prefix_compProd {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : source.trajectoryMeasure.map (Preorder.frestrictLe n) ⊗ₘ source.batchKernel n = source.trajectoryMeasure.map (fun trajectory => (Preorder.frestrictLe n trajectory, trajectory (n + 1)))","missing":[],"search":"trajectorymeasure_prefix_compprod banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_prefix_compprod the prefix/next-batch marginal has the configured compprod law. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviation","label":"successorSampledReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviation","description":"The dynamic next-round sampled-return statistic on prefix/batch pairs.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-bf6e6273f47e","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8060,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorSampledReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (pair : StochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp episodes) : Real","missing":[],"search":"successorsampledreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorsampledreturndeviation the dynamic next-round sampled-return statistic on prefix/batch pairs. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorDeviationKernel","label":"successorDeviationKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorDeviationKernel","description":"Conditional kernel of the dynamic next-round deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-f4f052c0bc27","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8061,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:289"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorDeviationKernel {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.Kernel (StochasticEpisodeBatchPrefix mdp episodes n) Real","missing":[],"search":"successordeviationkernel banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successordeviationkernel conditional kernel of the dynamic next-round deviation. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorDeviationKernel_apply","label":"successorDeviationKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorDeviationKernel_apply","description":"Each deviation-kernel fiber is the selected policy's iid statistic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-2919f73cc2b1","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8062,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:310"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem successorDeviationKernel_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (history : StochasticEpisodeBatchPrefix mdp episodes n) : source.successorDeviationKernel n history = (source.rewardSource.iidStochasticTrajectoryFamilyMeasure (source.successorPolicy n history) initialState episodes).map (mdp.sampledCumulativeReturnDeviationSum (source.successorPolicy n history) episodes)","missing":[],"search":"successordeviationkernel_apply banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successordeviationkernel_apply each deviation-kernel fiber is the selected policy's iid statistic law. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviationAt","label":"successorSampledReturnDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviationAt","description":"The trajectory-level dynamic deviation at successor coordinate `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-393ac417549c","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8063,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:343"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"successorsampledreturndeviationat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.successorsampledreturndeviationat the trajectory-level dynamic deviation at successor coordinate `n + 1`. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorSampledReturnDeviationAt","label":"measurable_successorSampledReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorSampledReturnDeviationAt","description":"theorem measurable_successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Measurable (source.successorSampledReturnDeviationAt n)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-371fa9f84477","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8064,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:351"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Measurable (source.successorSampledReturnDeviationAt n)","missing":[],"search":"measurable_successorsampledreturndeviationat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_successorsampledreturndeviationat theorem measurable_successorsampledreturndeviationat {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) (n : nat) : measurable (source.successorsampledreturndeviationat n) theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","label":"trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","description":"The conditional law of the dynamic successor deviation.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-28c8e721c0c8","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8065,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:361"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : ProbabilityTheory.condDistrib (source.successorSampledReturnDeviationAt n) (Preorder.frestrictLe n) source.trajectoryMeasure =ᵐ[ source.trajectoryMeasure.map (Preorder.frestrictLe n)] source.successorDeviationKernel n","missing":[],"search":"trajectorymeasure_conddistrib_successorsampledreturndeviationat banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_conddistrib_successorsampledreturndeviationat the conditional law of the dynamic successor deviation. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorSampledReturnDeviationAt_eq","label":"condExpKernel_map_successorSampledReturnDeviationAt_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorSampledReturnDeviationAt_eq","description":"The trimmed conditional-expectation kernel has the same dynamic law.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-2b1e70b52b25","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8066,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:387"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem condExpKernel_map_successorSampledReturnDeviationAt_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : Filter.Eventually (fun trajectory : StochasticEpisodeBatchTrajectory mdp episodes => Measure.map (source.successorSampledReturnDeviationAt n) (ProbabilityTheory.condExpKernel source.trajectoryMeasure ((inferInstance : MeasurableSpace (StochasticEpisodeBatchPrefix mdp episodes n)).comap (Preorder.frestrictLe n)) trajectory) = source.successorDeviationKernel n (Preorder.frestrictLe n trajectory)) (ae (source.trajectoryMeasure.trim (Preorder.measurable_frestrictL…","missing":[],"search":"condexpkernel_map_successorsampledreturndeviationat_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.condexpkernel_map_successorsampledreturndeviationat_eq the trimmed conditional-expectation kernel has the same dynamic law. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationPrefixIncrement","label":"sampledReturnDeviationPrefixIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationPrefixIncrement","description":"Prefix-level increment, including the genuine initial stochastic batch.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-1ec1a8410c58","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8067,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:417"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : (round : Nat) → StochasticEpisodeBatchPrefix mdp episodes round → Real | 0, history => mdp.sampledCumulativeReturnDeviationSum source.initialPolicy episodes (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩) | n + 1, history => source.successorSampledReturnDeviation n (Preorder.frestrictLe₂ (π := fun _ : Nat => StochasticEpisodeBatch mdp episodes) (Nat.le_succ n) history, history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) theorem measurable_sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round :…","missing":[],"search":"sampledreturndeviationprefixincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.sampledreturndeviationprefixincrement prefix-level increment, including the genuine initial stochastic batch. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationPrefixIncrement","label":"measurable_sampledReturnDeviationPrefixIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationPrefixIncrement","description":"theorem measurable_sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationPrefixIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-864d4b9d9b01","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8068,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:432"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledReturnDeviationPrefixIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationPrefixIncrement round)","missing":[],"search":"measurable_sampledreturndeviationprefixincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_sampledreturndeviationprefixincrement theorem measurable_sampledreturndeviationprefixincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) (round : nat) : measurable (source.sampledreturndeviationprefixincrement round) theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement","label":"sampledReturnDeviationIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement","description":"Adapted sampled-return deviation increment on the full trajectory.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-4114daf1bb5b","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8069,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:455"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"sampledreturndeviationincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.sampledreturndeviationincrement adapted sampled-return deviation increment on the full trajectory. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationIncrement","label":"measurable_sampledReturnDeviationIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationIncrement","description":"theorem measurable_sampledReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationIncrement round)","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-f31ed59a5b65","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8070,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:464"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) : Measurable (source.sampledReturnDeviationIncrement round)","missing":[],"search":"measurable_sampledreturndeviationincrement banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.measurable_sampledreturndeviationincrement theorem measurable_sampledreturndeviationincrement {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) (round : nat) : measurable (source.sampledreturndeviationincrement round) theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_stronglyAdapted_piLE","label":"sampledReturnDeviationIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_stronglyAdapted_piLE","description":"theorem sampledReturnDeviationIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp episodes)) source.sampledReturnDeviationIncrement","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-923900ca57df","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8071,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:473"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledReturnDeviationIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : StronglyAdapted (Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp episodes)) source.sampledReturnDeviationIncrement","missing":[],"search":"sampledreturndeviationincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.sampledreturndeviationincrement_stronglyadapted_pile theorem sampledreturndeviationincrement_stronglyadapted_pile {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat} (source : adaptivestochasticepisodebatchsource mdp initialstate episodes) : stronglyadapted (filtration.pile (x := fun _ : nat => stochasticepisodebatch mdp episodes)) source.sampledreturndeviationincrement theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","label":"sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","description":"The initial adaptive coordinate inherits the iid batch MGF.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-cd29dc8b62a6","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8072,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:487"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledReturnDeviationIncrement_zero_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasSubgaussianMGF (source.sampledReturnDeviationIncrement 0) (mdp.iidSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy) source.trajectoryMeasure","missing":[],"search":"sampledreturndeviationincrement_zero_hassubgaussianmgf banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.sampledreturndeviationincrement_zero_hassubgaussianmgf the initial adaptive coordinate inherits the iid batch mgf. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","label":"sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","description":"Every successor adaptive coordinate is conditionally sub-Gaussian.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-b7b7cd291762","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8073,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:520"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasCondSubgaussianMGF (Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp episodes) n) ((Filtration.piLE (X := fun _ : Nat => StochasticEpisodeBatch mdp ep…","missing":[],"search":"sampledreturndeviationincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.sampledreturndeviationincrement_succ_hascondsubgaussianmgf every successor adaptive coordinate is conditionally sub-gaussian. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy","label":"cumulativeSampledReturnDeviationVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy","description":"Sum of per-round proxies over `rounds` adaptive stochastic batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-e7fdf7b21def","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8074,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:626"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSampledReturnDeviationVarianceProxy (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"cumulativesampledreturndeviationvarianceproxy banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesampledreturndeviationvarianceproxy sum of per-round proxies over `rounds` adaptive stochastic batches. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy_eq","label":"cumulativeSampledReturnDeviationVarianceProxy_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy_eq","description":"theorem cumulativeSampledReturnDeviationVarianceProxy_eq (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : cumulativeSampledReturnDeviationVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy = (rounds : NNReal) * mdp.iidSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-add846167c6f","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8075,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:633"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem cumulativeSampledReturnDeviationVarianceProxy_eq (mdp : MDP State Action) (rounds episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : cumulativeSampledReturnDeviationVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy = (rounds : NNReal) * mdp.iidSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy","missing":[],"search":"cumulativesampledreturndeviationvarianceproxy_eq banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesampledreturndeviationvarianceproxy_eq theorem cumulativesampledreturndeviationvarianceproxy_eq (mdp : mdp state action) (rounds episodes : nat) (rewardbound rewardvarianceproxy : nnreal) : cumulativesampledreturndeviationvarianceproxy mdp rounds episodes rewardbound rewardvarianceproxy = (rounds : nnreal) * mdp.iidsampledcumulativereturndeviationvarianceproxy episodes rewardbound rewardvarianceproxy theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviation","label":"cumulativeSampledReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviation","description":"Cumulative sampled-return deviation over `rounds` adaptive batches.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-bf87808d2f70","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8076,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:644"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeSampledReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : StochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"cumulativesampledreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.cumulativesampledreturndeviation cumulative sampled-return deviation over `rounds` adaptive batches. definition compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","label":"trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","description":"Fixed-round two-sided tail for adaptive stochastic sampled returns.","url":"../modules/banditrlproof-rl-finitehorizonadaptivestochasticrewardtotalreturnconcentration/index.html#decl-b2ea08bf7269","parent":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","order":8077,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration.lean:654"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] [StandardBorelSpace (StochasticEpisodeBatch mdp episodes)] [Nonempty (StochasticEpisodeBatch mdp episodes)] [StandardBorelSpace (StochasticEpisodeBatchTrajectory mdp episodes)] (source : AdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (law : source.rewardSource.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((cumulativeSampledReturnDeviationVarianceProxy mdp rounds episodes rewardBound rewardVarianceProxy : NNReal) : Real)) (delta : Real) (hdelta : 0 < delta) (hdelta_le…","missing":[],"search":"trajectorymeasure_cumulativesampledreturndeviation_abs_tail_le banditrlproof.finitehorizonrl.adaptivestochasticepisodebatchsource.trajectorymeasure_cumulativesampledreturndeviation_abs_tail_le fixed-round two-sided tail for adaptive stochastic sampled returns. theorem compiled","shard":"modules/84480fd4aa05d31d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.abs_integral_sub_integral_le_sum_coordinateRadius_mul_envelope","label":"abs_integral_sub_integral_le_sum_coordinateRadius_mul_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.abs_integral_sub_integral_le_sum_coordinateRadius_mul_envelope","description":"On a finite measurable space, coordinate bounds on singleton masses control the expectation error of every pointwise-enveloped real value.","url":"../modules/banditrlproof-rl-finitehorizoncoordinatemodelconfidence/index.html#decl-499731b37f8f","parent":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","order":8078,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonCoordinateModelConfidence.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem abs_integral_sub_integral_le_sum_coordinateRadius_mul_envelope (estimated trueMeasure : Measure State) [IsFiniteMeasure estimated] [IsFiniteMeasure trueMeasure] (value coordinateRadius : State -> Real) (envelope : Real) (hcoordinate : forall state, |estimated.real {state} - trueMeasure.real {state}| <= coordinateRadius state) (hvalue : forall state, |value state| <= envelope) : |(∫ state, value state ∂estimated) - ∫ state, value state ∂trueMeasure| <= ∑ state, coordinateRadius state * envelope","missing":[],"search":"abs_integral_sub_integral_le_sum_coordinateradius_mul_envelope banditrlproof.finitehorizonrl.abs_integral_sub_integral_le_sum_coordinateradius_mul_envelope on a finite measurable space, coordinate bounds on singleton masses control the expectation error of every pointwise-enveloped real value. theorem compiled","shard":"modules/e4a4d6678bda70c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence","label":"CoordinateConfidence","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence","description":"Finite-state coordinate confidence sufficient for the recursive optimistic model route. The transition radius may be any upper bound on the displayed coordinate sum, so later concentration producers can choose their own radii.","url":"../modules/banditrlproof-rl-finitehorizoncoordinatemodelconfidence/index.html#decl-e7852548e846","parent":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","order":8079,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonCoordinateModelConfidence.lean:81"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure CoordinateConfidence {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) where","missing":[],"search":"coordinateconfidence banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.coordinateconfidence finite-state coordinate confidence sufficient for the recursive optimistic model route. the transition radius may be any upper bound on the displayed coordinate sum, so later concentration producers can choose their own radii. structure compiled","shard":"modules/e4a4d6678bda70c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.transitionError_le_radius","label":"transitionError_le_radius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.transitionError_le_radius","description":"Coordinate transition confidence implies the recursive Bellman error bound.","url":"../modules/banditrlproof-rl-finitehorizoncoordinatemodelconfidence/index.html#decl-00c0488eb2aa","parent":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","order":8080,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonCoordinateModelConfidence.lean:116"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem CoordinateConfidence.transitionError_le_radius {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.CoordinateConfidence) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) : |plan.transitionValue (mdp.decisionStageRemaining remaining hremaining) (plan.upperValueRemaining remaining (by omega)) state action - mdp.transitionValue (plan.upperValueRemaining remaining (by omega)) state action| <= plan.transitionRadius (mdp.decisionStageRemaining remaining hremaining) state action","missing":[],"search":"transitionerror_le_radius banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.coordinateconfidence.transitionerror_le_radius coordinate transition confidence implies the recursive bellman error bound. theorem compiled","shard":"modules/e4a4d6678bda70c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.toConfidence","label":"toConfidence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.toConfidence","description":"Package coordinate confidence as the existing estimated-model confidence.","url":"../modules/banditrlproof-rl-finitehorizoncoordinatemodelconfidence/index.html#decl-84eb05a6712a","parent":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","order":8081,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonCoordinateModelConfidence.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def CoordinateConfidence.toConfidence {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.CoordinateConfidence) : plan.Confidence where","missing":[],"search":"toconfidence banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.coordinateconfidence.toconfidence package coordinate confidence as the existing estimated-model confidence. definition compiled","shard":"modules/e4a4d6678bda70c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","label":"optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","description":"Route endpoint: coordinate model confidence gives global optimism and the compiled selected-radius single-episode expected-regret bound.","url":"../modules/banditrlproof-rl-finitehorizoncoordinatemodelconfidence/index.html#decl-51db79b460b3","parent":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","order":8082,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonCoordinateModelConfidence.lean:154"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem CoordinateConfidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.CoordinateConfidence) (initialState : Measure State) [IsProbabilityMeasure initialState] : (forall state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= plan.upperValueRemaining mdp.horizon le_rfl state) /\\ plan.optimisticPolicy.expectedRegret initialState <= plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState","missing":[],"search":"optimism_and_expectedregret_le_two_occupancyselectedradiusremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.coordinateconfidence.optimism_and_expectedregret_le_two_occupancyselectedradiusremaining route endpoint: coordinate model confidence gives global optimism and the compiled selected-radius single-episode expected-regret bound. theorem compiled","shard":"modules/e4a4d6678bda70c7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep","label":"EpisodeStep","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep","description":"One recorded finite-horizon transition with its observed reward.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-c6f547f10477","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8083,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure EpisodeStep (State : Type u) (Action : Type v) where","missing":[],"search":"episodestep banditrlproof.finitehorizonrl.episodestep one recorded finite-horizon transition with its observed reward. structure compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_state","label":"measurable_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_state","description":"theorem measurable_state : Measurable (fun step : EpisodeStep State Action => step.state)","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-82b5c83c6d49","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8084,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_state : Measurable (fun step : EpisodeStep State Action => step.state)","missing":[],"search":"measurable_state banditrlproof.finitehorizonrl.episodestep.measurable_state theorem measurable_state : measurable (fun step : episodestep state action => step.state) theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_action","label":"measurable_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_action","description":"theorem measurable_action : Measurable (fun step : EpisodeStep State Action => step.action)","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-0b3ea4382c00","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8085,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_action : Measurable (fun step : EpisodeStep State Action => step.action)","missing":[],"search":"measurable_action banditrlproof.finitehorizonrl.episodestep.measurable_action theorem measurable_action : measurable (fun step : episodestep state action => step.action) theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_reward","label":"measurable_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_reward","description":"theorem measurable_reward : Measurable (fun step : EpisodeStep State Action => step.reward)","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-88fa5aafeb5f","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8086,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_reward : Measurable (fun step : EpisodeStep State Action => step.reward)","missing":[],"search":"measurable_reward banditrlproof.finitehorizonrl.episodestep.measurable_reward theorem measurable_reward : measurable (fun step : episodestep state action => step.reward) theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_nextState","label":"measurable_nextState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_nextState","description":"theorem measurable_nextState : Measurable (fun step : EpisodeStep State Action => step.nextState)","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-86b05465fc21","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8087,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_nextState : Measurable (fun step : EpisodeStep State Action => step.nextState)","missing":[],"search":"measurable_nextstate banditrlproof.finitehorizonrl.episodestep.measurable_nextstate theorem measurable_nextstate : measurable (fun step : episodestep state action => step.nextstate) theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch","label":"EpisodeBatch","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch","description":"A finite table with one record at each valid stage of every episode. This raw type does not enforce cross-stage state continuity or identify the records with samples from the MDP trajectory law; those are downstream probabilistic laws.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-c9b119d6ce9a","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8088,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev EpisodeBatch (mdp : MDP State Action) (episodes : Nat)","missing":[],"search":"episodebatch banditrlproof.finitehorizonrl.episodebatch a finite table with one record at each valid stage of every episode. this raw type does not enforce cross-stage state continuity or identify the records with samples from the mdp trajectory law; those are downstream probabilistic laws. abbreviation compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.visitCount","label":"visitCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.visitCount","description":"Number of batch episodes visiting a state-action pair at one stage.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-bfed9d456214","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8089,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def visitCount {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : Nat","missing":[],"search":"visitcount banditrlproof.finitehorizonrl.episodebatch.visitcount number of batch episodes visiting a state-action pair at one stage. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.rewardSum","label":"rewardSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.rewardSum","description":"Sum of rewards recorded at a state-action pair and stage.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-5c225cc56b62","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8090,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def rewardSum {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : Real","missing":[],"search":"rewardsum banditrlproof.finitehorizonrl.episodebatch.rewardsum sum of rewards recorded at a state-action pair and stage. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalReward","label":"empiricalReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalReward","description":"Empirical reward mean, with the conventional zero value at zero visits.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-e5b833e8e577","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8091,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:118"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalReward {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : Real","missing":[],"search":"empiricalreward banditrlproof.finitehorizonrl.episodebatch.empiricalreward empirical reward mean, with the conventional zero value at zero visits. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCount","label":"transitionCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCount","description":"Number of matching transitions to a fixed next state.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-1b72681fea08","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8092,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionCount {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : Nat","missing":[],"search":"transitioncount banditrlproof.finitehorizonrl.episodebatch.transitioncount number of matching transitions to a fixed next state. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.sum_transitionCount_eq_visitCount","label":"sum_transitionCount_eq_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.sum_transitionCount_eq_visitCount","description":"Next-state transition counts partition the state-action visit count.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-8b604f5c5697","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8093,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:135"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_transitionCount_eq_visitCount {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : ∑ nextState, batch.transitionCount stage state action nextState = batch.visitCount stage state action","missing":[],"search":"sum_transitioncount_eq_visitcount banditrlproof.finitehorizonrl.episodebatch.sum_transitioncount_eq_visitcount next-state transition counts partition the state-action visit count. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF","label":"empiricalTransitionPMF","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF","description":"Empirical next-state PMF with an explicit zero-visit fallback state.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-9fe1c536a0b4","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8094,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:153"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalTransitionPMF {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : PMF State","missing":[],"search":"empiricaltransitionpmf banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionpmf empirical next-state pmf with an explicit zero-visit fallback state. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionMass","label":"empiricalTransitionMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionMass","description":"Real singleton mass of the empirical transition PMF.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-dfc87a5e6057","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8095,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalTransitionMass {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : Real","missing":[],"search":"empiricaltransitionmass banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionmass real singleton mass of the empirical transition pmf. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF_eq_pure_of_visitCount_eq_zero","label":"empiricalTransitionPMF_eq_pure_of_visitCount_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF_eq_pure_of_visitCount_eq_zero","description":"Zero-visit state-action pairs use the declared fallback distribution.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-ddefe6fd8e24","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8096,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:204"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionPMF_eq_pure_of_visitCount_eq_zero {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (hzero : batch.visitCount stage state action = 0) : batch.empiricalTransitionPMF defaultState stage state action = PMF.pure defaultState","missing":[],"search":"empiricaltransitionpmf_eq_pure_of_visitcount_eq_zero banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionpmf_eq_pure_of_visitcount_eq_zero zero-visit state-action pairs use the declared fallback distribution. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF_apply_of_visitCount_ne_zero","label":"empiricalTransitionPMF_apply_of_visitCount_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF_apply_of_visitCount_ne_zero","description":"At a positive visit count, empirical PMF mass is normalized count.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-77da9b52cdd3","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8097,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:216"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionPMF_apply_of_visitCount_ne_zero {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (hvisit : batch.visitCount stage state action ≠ 0) : batch.empiricalTransitionPMF defaultState stage state action nextState = (batch.transitionCount stage state action nextState : ENNReal) / (batch.visitCount stage state action : ENNReal)","missing":[],"search":"empiricaltransitionpmf_apply_of_visitcount_ne_zero banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionpmf_apply_of_visitcount_ne_zero at a positive visit count, empirical pmf mass is normalized count. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionMass_eq_div_of_visitCount_ne_zero","label":"empiricalTransitionMass_eq_div_of_visitCount_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionMass_eq_div_of_visitCount_ne_zero","description":"Real empirical singleton mass is the usual count divided by visit count.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-c34e36d21fd9","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8098,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:230"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionMass_eq_div_of_visitCount_ne_zero {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (hvisit : batch.visitCount stage state action ≠ 0) : batch.empiricalTransitionMass defaultState stage state action nextState = (batch.transitionCount stage state action nextState : Real) / (batch.visitCount stage state action : Real)","missing":[],"search":"empiricaltransitionmass_eq_div_of_visitcount_ne_zero banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionmass_eq_div_of_visitcount_ne_zero real empirical singleton mass is the usual count divided by visit count. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel","label":"empiricalTransitionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel","description":"State-action indexed empirical PMFs form a measurable finite-state kernel.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-6943db18b3e4","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8099,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:245"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def empiricalTransitionKernel {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel (State × Action) State","missing":[],"search":"empiricaltransitionkernel banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionkernel state-action indexed empirical pmfs form a measurable finite-state kernel. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_apply","label":"empiricalTransitionKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_apply","description":"theorem empiricalTransitionKernel_apply {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : batch.empiricalTransitionKernel defaultState stage (state, action) = (batch.empiricalTransitionPMF defaultState stage state action).toMeasure","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-103f13c500de","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8100,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:255"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionKernel_apply {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) : batch.empiricalTransitionKernel defaultState stage (state, action) = (batch.empiricalTransitionPMF defaultState stage state action).toMeasure","missing":[],"search":"empiricaltransitionkernel_apply banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionkernel_apply theorem empiricaltransitionkernel_apply {mdp : mdp state action} {episodes : nat} (batch : episodebatch mdp episodes) (defaultstate : state) (stage : fin mdp.horizon) (state : state) (action : action) : batch.empiricaltransitionkernel defaultstate stage (state, action) = (batch.empiricaltransitionpmf defaultstate stage state action).tomeasure theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_isMarkov","label":"empiricalTransitionKernel_isMarkov","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_isMarkov","description":"Every section of the empirical transition kernel is a probability law.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-6b3f86628cee","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8101,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:265"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionKernel_isMarkov {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) : ProbabilityTheory.IsMarkovKernel (batch.empiricalTransitionKernel defaultState stage) where","missing":[],"search":"empiricaltransitionkernel_ismarkov banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionkernel_ismarkov every section of the empirical transition kernel is a probability law. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_real_singleton","label":"empiricalTransitionKernel_real_singleton","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_real_singleton","description":"Kernel singleton mass agrees with the named empirical transition mass.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-79c4c059d6ae","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8102,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionKernel_real_singleton {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (defaultState : State) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (batch.empiricalTransitionKernel defaultState stage (state, action)).real {nextState} = batch.empiricalTransitionMass defaultState stage state action nextState","missing":[],"search":"empiricaltransitionkernel_real_singleton banditrlproof.finitehorizonrl.episodebatch.empiricaltransitionkernel_real_singleton kernel singleton mass agrees with the named empirical transition mass. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel","label":"FiniteBatchModel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel","description":"A finite batch together with reward and transition radii. The empirical reward and transition model are derived from `batch`; zero transition counts use the explicit `defaultState` fallback rather than an implicit arbitrary law.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-6412b3b044b9","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8103,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:300"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure FiniteBatchModel (mdp : MDP State Action) (episodes : Nat) where","missing":[],"search":"finitebatchmodel banditrlproof.finitehorizonrl.mdp.finitebatchmodel a finite batch together with reward and transition radii. the empirical reward and transition model are derived from `batch`; zero transition counts use the explicit `defaultstate` fallback rather than an implicit arbitrary law. structure compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.plan","label":"plan","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.plan","description":"The estimated-model plan canonically generated by a finite episode batch.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-c28f02277566","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8104,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def plan {mdp : MDP State Action} {episodes : Nat} (model : FiniteBatchModel mdp episodes) : EstimatedModelPlan mdp where","missing":[],"search":"plan banditrlproof.finitehorizonrl.mdp.finitebatchmodel.plan the estimated-model plan canonically generated by a finite episode batch. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence","label":"Confidence","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence","description":"Raw finite-batch confidence contracts. Reward errors and singleton-frequency errors are stated directly on the empirical statistics, while the envelope and radius-cover fields discharge the value-dependent coordinate transport.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-f22d4a053f0b","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8105,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:327"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure Confidence {mdp : MDP State Action} {episodes : Nat} (model : FiniteBatchModel mdp episodes) where","missing":[],"search":"confidence banditrlproof.finitehorizonrl.mdp.finitebatchmodel.confidence raw finite-batch confidence contracts. reward errors and singleton-frequency errors are stated directly on the empirical statistics, while the envelope and radius-cover fields discharge the value-dependent coordinate transport. structure compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence.toCoordinateConfidence","label":"toCoordinateConfidence","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence.toCoordinateConfidence","description":"Raw empirical-statistic confidence gives coordinate model confidence.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-b88d2721f921","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8106,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:362"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def Confidence.toCoordinateConfidence {mdp : MDP State Action} {episodes : Nat} {model : FiniteBatchModel mdp episodes} (confidence : model.Confidence) : model.plan.CoordinateConfidence where","missing":[],"search":"tocoordinateconfidence banditrlproof.finitehorizonrl.mdp.finitebatchmodel.confidence.tocoordinateconfidence raw empirical-statistic confidence gives coordinate model confidence. definition compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","label":"optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","description":"Route endpoint: finite-batch reward and singleton-frequency confidence imply global optimism and the compiled selected-radius expected-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonempiricalmodel/index.html#decl-f6beafdf2e94","parent":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","order":8107,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEmpiricalModel"],["Source","BanditRLProof/RL/FiniteHorizonEmpiricalModel.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining {mdp : MDP State Action} {episodes : Nat} {model : FiniteBatchModel mdp episodes} (confidence : model.Confidence) (initialState : Measure State) [IsProbabilityMeasure initialState] : (forall state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= model.plan.upperValueRemaining mdp.horizon le_rfl state) /\\ model.plan.optimisticPolicy.expectedRegret initialState <= model.plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * model.plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState","missing":[],"search":"optimism_and_expectedregret_le_two_occupancyselectedradiusremaining banditrlproof.finitehorizonrl.mdp.finitebatchmodel.confidence.optimism_and_expectedregret_le_two_occupancyselectedradiusremaining route endpoint: finite-batch reward and singleton-frequency confidence imply global optimism and the compiled selected-radius expected-regret bound. theorem compiled","shard":"modules/7783b9dec15003e8.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.toProdEquiv","label":"toProdEquiv","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.toProdEquiv","description":"Product coordinates underlying the measurable structure on `EpisodeStep`.","url":"../modules/banditrlproof-rl-finitehorizonepisodebatchstandardborel/index.html#decl-e69217e35658","parent":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","order":8108,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel"],["Source","BanditRLProof/RL/FiniteHorizonEpisodeBatchStandardBorel.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def toProdEquiv : EpisodeStep State Action ≃ State × Action × Real × State where","missing":[],"search":"toprodequiv banditrlproof.finitehorizonrl.episodestep.toprodequiv product coordinates underlying the measurable structure on `episodestep`. definition compiled","shard":"modules/2c82e1f72ed73243.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.decisionStageRemaining","label":"decisionStageRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.decisionStageRemaining","description":"Chronological stage corresponding to a successor remaining-horizon index.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-99e0fe4f8940","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8109,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def decisionStageRemaining (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : Fin mdp.horizon","missing":[],"search":"decisionstageremaining banditrlproof.finitehorizonrl.mdp.decisionstageremaining chronological stage corresponding to a successor remaining-horizon index. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan","label":"EstimatedModelPlan","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan","description":"A stage-indexed estimated MDP together with separate reward and transition confidence radii. Confidence itself is a downstream proposition because the transition error is evaluated on the recursively generated upper value.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-5cfc50290a4c","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8110,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure EstimatedModelPlan (mdp : MDP State Action) where","missing":[],"search":"estimatedmodelplan banditrlproof.finitehorizonrl.mdp.estimatedmodelplan a stage-indexed estimated mdp together with separate reward and transition confidence radii. confidence itself is a downstream proposition because the transition error is evaluated on the recursively generated upper value. structure compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.exactModelPlan","label":"exactModelPlan","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.exactModelPlan","description":"The true MDP with zero radii is the canonical exact estimated plan.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-3e6b86b08be6","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8111,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def exactModelPlan (mdp : MDP State Action) : EstimatedModelPlan mdp where","missing":[],"search":"exactmodelplan banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.exactmodelplan the true mdp with zero radii is the canonical exact estimated plan. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.transitionValue","label":"transitionValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.transitionValue","description":"Estimated next-state expectation of a continuation value.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-44163d853748","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8112,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def transitionValue {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) (action : Action) : Real","missing":[],"search":"transitionvalue banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.transitionvalue estimated next-state expectation of a continuation value. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.bellmanQ","label":"bellmanQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.bellmanQ","description":"Estimated one-step reward plus continuation value.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-adaff1620fb8","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8113,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellmanQ {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) (action : Action) : Real","missing":[],"search":"bellmanq banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.bellmanq estimated one-step reward plus continuation value. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticQ","label":"optimisticQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticQ","description":"Estimated Bellman action value enlarged by both confidence radii.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-2332244fca8f","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8114,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:91"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticQ {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) (action : Action) : Real","missing":[],"search":"optimisticq banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticq estimated bellman action value enlarged by both confidence radii. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_transitionValue","label":"measurable_transitionValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_transitionValue","description":"The estimated transition-value surface is measurable in the state-action pair.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-a677d4bdec88","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8115,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionValue {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) {value : State -> Real} (hvalue : Measurable value) : Measurable (Function.uncurry (plan.transitionValue stage value))","missing":[],"search":"measurable_transitionvalue banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.measurable_transitionvalue the estimated transition-value surface is measurable in the state-action pair. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticQ","label":"measurable_optimisticQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticQ","description":"The optimistic estimated action-value surface is measurable.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-23c70dddaff6","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8116,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimisticQ {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) {value : State -> Real} (hvalue : Measurable value) : Measurable (Function.uncurry (plan.optimisticQ stage value))","missing":[],"search":"measurable_optimisticq banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.measurable_optimisticq the optimistic estimated action-value surface is measurable. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticAction","label":"optimisticAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticAction","description":"A finite action maximizing the estimated optimistic action value.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-d3677e66124f","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8117,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:122"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticAction {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : Action","missing":[],"search":"optimisticaction banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticaction a finite action maximizing the estimated optimistic action value. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticQ_le_optimisticAction","label":"optimisticQ_le_optimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticQ_le_optimisticAction","description":"Every action is bounded by the selected optimistic estimated action value.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-69599bf1e38b","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8118,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:130"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimisticQ_le_optimisticAction {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) (action : Action) : plan.optimisticQ stage value state action <= plan.optimisticQ stage value state (plan.optimisticAction stage value state)","missing":[],"search":"optimisticq_le_optimisticaction banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticq_le_optimisticaction every action is bounded by the selected optimistic estimated action value. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticBellman","label":"optimisticBellman","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticBellman","description":"Pointwise maximum of the estimated optimistic action values.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-9b162d9ed9ca","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8119,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:141"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticBellman {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : Real","missing":[],"search":"optimisticbellman banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticbellman pointwise maximum of the estimated optimistic action values. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticAction","label":"measurable_optimisticAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticAction","description":"The finite-state estimated optimistic selector is measurable.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-8ead73b3dc50","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8120,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:148"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimisticAction {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) : Measurable (plan.optimisticAction stage value)","missing":[],"search":"measurable_optimisticaction banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.measurable_optimisticaction the finite-state estimated optimistic selector is measurable. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.upperValueRemaining","label":"upperValueRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.upperValueRemaining","description":"Recursive estimated optimistic value indexed by decisions remaining.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-54a87a8b467a","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8121,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:155"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def upperValueRemaining {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> Real | 0, _ => fun _ => 0 | remaining + 1, hremaining => plan.optimisticBellman (mdp.decisionStageRemaining remaining hremaining) (plan.upperValueRemaining remaining (by omega)) omit [MeasurableSingletonClass State] in /-- Transport the dependent upper-value recursion across equal remaining horizons. -/ theorem upperValueRemaining_eq_of_eq {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) {left right : Nat} (hleft : left <= mdp.horizon) (hright : right <= mdp.horizon) (h : left = right) : plan.upperValueRemaining left hleft = plan.upperValueRemaining right hright","missing":[],"search":"uppervalueremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.uppervalueremaining recursive estimated optimistic value indexed by decisions remaining. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.upperValueRemaining_eq_of_eq","label":"upperValueRemaining_eq_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.upperValueRemaining_eq_of_eq","description":"Transport the dependent upper-value recursion across equal remaining horizons.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-1c19af1169ad","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8122,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem upperValueRemaining_eq_of_eq {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) {left right : Nat} (hleft : left <= mdp.horizon) (hright : right <= mdp.horizon) (h : left = right) : plan.upperValueRemaining left hleft = plan.upperValueRemaining right hright","missing":[],"search":"uppervalueremaining_eq_of_eq banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.uppervalueremaining_eq_of_eq transport the dependent upper-value recursion across equal remaining horizons. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_upperValueRemaining","label":"measurable_upperValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_upperValueRemaining","description":"Every recursively generated upper-value surface is measurable.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-bfe888623ebb","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8123,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:175"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_upperValueRemaining {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (plan.upperValueRemaining remaining hremaining)","missing":[],"search":"measurable_uppervalueremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.measurable_uppervalueremaining every recursively generated upper-value surface is measurable. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence","label":"Confidence","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence","description":"Two-sided model confidence on the recursive upper-value route. The transition contract is deliberately not quantified over arbitrary unbounded values.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-633bd08ba5f2","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8124,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:185"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure Confidence {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) : Prop where","missing":[],"search":"confidence banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence two-sided model confidence on the recursive upper-value route. the transition contract is deliberately not quantified over arbitrary unbounded values. structure compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.exactModelPlan_confidence","label":"exactModelPlan_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.exactModelPlan_confidence","description":"The canonical exact estimated plan satisfies confidence by reflexivity.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-915d2d164276","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8125,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:203"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exactModelPlan_confidence (mdp : MDP State Action) : (exactModelPlan mdp).Confidence","missing":[],"search":"exactmodelplan_confidence banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.exactmodelplan_confidence the canonical exact estimated plan satisfies confidence by reflexivity. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.trueBellmanQ_le_optimisticQ","label":"trueBellmanQ_le_optimisticQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.trueBellmanQ_le_optimisticQ","description":"Two-sided confidence makes every true action value optimistic.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-24b305aefbc9","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8126,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:213"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.trueBellmanQ_le_optimisticQ {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.Confidence) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action) : mdp.bellmanQ (plan.upperValueRemaining remaining (by omega)) state action <= plan.optimisticQ (mdp.decisionStageRemaining remaining hremaining) (plan.upperValueRemaining remaining (by omega)) state action","missing":[],"search":"truebellmanq_le_optimisticq banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence.truebellmanq_le_optimisticq two-sided confidence makes every true action value optimistic. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.certificate","label":"certificate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.certificate","description":"The recursive estimated optimistic values form a true Bellman certificate.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-8c6c8945afc6","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8127,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def certificate {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (confidence : plan.Confidence) : mdp.OptimisticBellmanCertificate where","missing":[],"search":"certificate banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.certificate the recursive estimated optimistic values form a true bellman certificate. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.optimalValueRemaining_le_upperValueRemaining","label":"optimalValueRemaining_le_upperValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.optimalValueRemaining_le_upperValueRemaining","description":"Estimated-model confidence implies pointwise global optimism.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-252318e27cdf","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8128,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:258"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.optimalValueRemaining_le_upperValueRemaining {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.Confidence) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : mdp.optimalValueRemaining remaining hremaining state <= plan.upperValueRemaining remaining hremaining state","missing":[],"search":"optimalvalueremaining_le_uppervalueremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence.optimalvalueremaining_le_uppervalueremaining estimated-model confidence implies pointwise global optimism. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticActionAt","label":"optimisticActionAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticActionAt","description":"Estimated optimistic action at a chronological stage.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-d46e5eade4c5","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8129,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:268"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticActionAt {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) : State -> Action","missing":[],"search":"optimisticactionat banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticactionat estimated optimistic action at a chronological stage. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticActionAt","label":"measurable_optimisticActionAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticActionAt","description":"The chronological estimated optimistic selector is measurable.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-a71a97b28173","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8130,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:276"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimisticActionAt {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) : Measurable (plan.optimisticActionAt stage)","missing":[],"search":"measurable_optimisticactionat banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.measurable_optimisticactionat the chronological estimated optimistic selector is measurable. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticPolicy","label":"optimisticPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticPolicy","description":"Deterministic Markov policy greedy for each estimated optimistic backup.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-5496367e3709","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8131,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:283"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimisticPolicy {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) : MarkovPolicy mdp where","missing":[],"search":"optimisticpolicy banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticpolicy deterministic markov policy greedy for each estimated optimistic backup. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticActionAt_decisionStageRemaining","label":"optimisticActionAt_decisionStageRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticActionAt_decisionStageRemaining","description":"Remaining-horizon and chronological selectors agree exactly.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-6508c2621789","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8132,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:296"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimisticActionAt_decisionStageRemaining {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : plan.optimisticActionAt (mdp.decisionStageRemaining remaining hremaining) = plan.optimisticAction (mdp.decisionStageRemaining remaining hremaining) (plan.upperValueRemaining remaining (by omega))","missing":[],"search":"optimisticactionat_decisionstageremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticactionat_decisionstageremaining remaining-horizon and chronological selectors agree exactly. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticPolicy_bellman_eq_bellmanQ","label":"optimisticPolicy_bellman_eq_bellmanQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticPolicy_bellman_eq_bellmanQ","description":"The deterministic estimated-greedy policy Bellman integral selects its action.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-f7a7ff8588a3","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8133,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimisticPolicy_bellman_eq_bellmanQ {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : plan.optimisticPolicy.bellman stage value state = mdp.bellmanQ value state (plan.optimisticActionAt stage state)","missing":[],"search":"optimisticpolicy_bellman_eq_bellmanq banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.optimisticpolicy_bellman_eq_bellmanq the deterministic estimated-greedy policy bellman integral selects its action. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.selectedRadiusRemaining","label":"selectedRadiusRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.selectedRadiusRemaining","description":"Reward-plus-transition radius selected by the estimated optimistic policy.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-db960ff853b1","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8134,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:324"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedRadiusRemaining {mdp : MDP State Action} (plan : EstimatedModelPlan mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : Real","missing":[],"search":"selectedradiusremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.selectedradiusremaining reward-plus-transition radius selected by the estimated optimistic policy. definition compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.selectedRadiusRemaining_nonneg","label":"selectedRadiusRemaining_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.selectedRadiusRemaining_nonneg","description":"Every selected radius is nonnegative under the two-sided confidence contract.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-16954de7cf37","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8135,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:336"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.selectedRadiusRemaining_nonneg {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.Confidence) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : 0 <= plan.selectedRadiusRemaining remaining hremaining state","missing":[],"search":"selectedradiusremaining_nonneg banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence.selectedradiusremaining_nonneg every selected radius is nonnegative under the two-sided confidence contract. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.policyBellmanResidual_le_two_selectedRadiusRemaining","label":"policyBellmanResidual_le_two_selectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.policyBellmanResidual_le_two_selectedRadiusRemaining","description":"The estimated-greedy policy residual is at most twice its selected model confidence radius. One side of each absolute-error bound produced optimism; the other side controls the remaining selected-action overestimate.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-b022999abb8f","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8136,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:357"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.policyBellmanResidual_le_two_selectedRadiusRemaining {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.Confidence) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (plan.certificate confidence).policyBellmanResidual plan.optimisticPolicy remaining hremaining state <= 2 * plan.selectedRadiusRemaining remaining hremaining state","missing":[],"search":"policybellmanresidual_le_two_selectedradiusremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence.policybellmanresidual_le_two_selectedradiusremaining the estimated-greedy policy residual is at most twice its selected model confidence radius. one side of each absolute-error bound produced optimism; the other side controls the remaining selected-action overestimate. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.expectedRegret_le_two_occupancySelectedRadiusRemaining","label":"expectedRegret_le_two_occupancySelectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.expectedRegret_le_two_occupancySelectedRadiusRemaining","description":"Full residual chain for the estimated-greedy policy under a probability initial law.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-145182db48c7","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8137,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.expectedRegret_le_two_occupancySelectedRadiusRemaining {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.Confidence) (initialState : Measure State) [IsProbabilityMeasure initialState] : 0 <= (plan.certificate confidence).residualOccupancyRemaining plan.optimisticPolicy mdp.horizon le_rfl initialState /\\ plan.optimisticPolicy.expectedRegret initialState <= (plan.certificate confidence).residualOccupancyRemaining plan.optimisticPolicy mdp.horizon le_rfl initialState /\\ (plan.certificate confidence).residualOccupancyRemaining plan.optimisticPolicy mdp.horizon le_rfl initialState <= plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState /\\ plan.optimisticPolicy.expectedRegret initialState <= plan.optimisticPolicy.occupancySumRemain…","missing":[],"search":"expectedregret_le_two_occupancyselectedradiusremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence.expectedregret_le_two_occupancyselectedradiusremaining full residual chain for the estimated-greedy policy under a probability initial law. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","label":"optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","description":"Route endpoint: estimated-model confidence simultaneously gives global optimism and the selected-radius single-episode expected-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonestimatedmodelcertificate/index.html#decl-e3f74d0b8df5","parent":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","order":8138,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate"],["Source","BanditRLProof/RL/FiniteHorizonEstimatedModelCertificate.lean:425"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining {mdp : MDP State Action} {plan : EstimatedModelPlan mdp} (confidence : plan.Confidence) (initialState : Measure State) [IsProbabilityMeasure initialState] : (forall state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= plan.upperValueRemaining mdp.horizon le_rfl state) /\\ plan.optimisticPolicy.expectedRegret initialState <= plan.optimisticPolicy.occupancySumRemaining (fun remaining hremaining state => 2 * plan.selectedRadiusRemaining remaining hremaining state) mdp.horizon le_rfl initialState","missing":[],"search":"optimism_and_expectedregret_le_two_occupancyselectedradiusremaining banditrlproof.finitehorizonrl.mdp.estimatedmodelplan.confidence.optimism_and_expectedregret_le_two_occupancyselectedradiusremaining route endpoint: estimated-model confidence simultaneously gives global optimism and the selected-radius single-episode expected-regret bound. theorem compiled","shard":"modules/e0b34e8febac75ca.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathCalibrationDimensionFactor","label":"exploratoryPathCalibrationDimensionFactor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathCalibrationDimensionFactor","description":"Dimension factor sufficient to turn a radius margin into the half contraction.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-6726e4ef36f6","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8139,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryPathCalibrationDimensionFactor (mdp : MDP State Action) : Real","missing":[],"search":"exploratorypathcalibrationdimensionfactor banditrlproof.finitehorizonrl.exploratorypathcalibrationdimensionfactor dimension factor sufficient to turn a radius margin into the half contraction. definition compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.one_le_exploratoryPathCalibrationDimensionFactor","label":"one_le_exploratoryPathCalibrationDimensionFactor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.one_le_exploratoryPathCalibrationDimensionFactor","description":"The calibration dimension factor is at least one.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-103a8082a875","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8140,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_le_exploratoryPathCalibrationDimensionFactor (mdp : MDP State Action) : 1 <= exploratoryPathCalibrationDimensionFactor mdp","missing":[],"search":"one_le_exploratorypathcalibrationdimensionfactor banditrlproof.finitehorizonrl.one_le_exploratorypathcalibrationdimensionfactor the calibration dimension factor is at least one. theorem compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathCalibrationEpisodeThreshold","label":"exploratoryPathCalibrationEpisodeThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathCalibrationEpisodeThreshold","description":"Closed-form Real episode threshold for the local simultaneous-count budget. Its denominator is positive when `visitFloor` is positive.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-ad8017e96aff","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8141,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryPathCalibrationEpisodeThreshold (mdp : MDP State Action) (rounds : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"exploratorypathcalibrationepisodethreshold banditrlproof.finitehorizonrl.exploratorypathcalibrationepisodethreshold closed-form real episode threshold for the local simultaneous-count budget. its denominator is positive when `visitfloor` is positive. definition compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_dimensionFactor_of_threshold","label":"simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_dimensionFactor_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_dimensionFactor_of_threshold","description":"Above the explicit episode threshold, the simultaneous count radius is less than the common expected-count floor divided by the calibration dimension factor.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-3c322027c076","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8142,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_dimensionFactor_of_threshold (mdp : MDP State Action) (witnessState : State) {rounds episodes : Nat} {delta visitFloor : Real} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hthreshold : exploratoryPathCalibrationEpisodeThreshold mdp rounds delta visitFloor < (episodes : Real)) : simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFloor / exploratoryPathCalibrationDimensionFactor mdp","missing":[],"search":"simultaneouscountconfidenceradius_lt_episodes_mul_visitfloor_div_dimensionfactor_of_threshold banditrlproof.finitehorizonrl.simultaneouscountconfidenceradius_lt_episodes_mul_visitfloor_div_dimensionfactor_of_threshold above the explicit episode threshold, the simultaneous count radius is less than the common expected-count floor divided by the calibration dimension factor. theorem compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.episodeThreshold_countMargin_and_halfContraction","label":"episodeThreshold_countMargin_and_halfContraction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.episodeThreshold_countMargin_and_halfContraction","description":"The episode threshold simultaneously discharges the strict count margin and the half-contraction premise of the explicit path-support calibration route.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-4a4b552bd206","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8143,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:137"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeThreshold_countMargin_and_halfContraction (mdp : MDP State Action) (witnessState : State) {rounds episodes : Nat} {delta visitFloor : Real} (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (hvisitFloor : 0 < visitFloor) (hthreshold : exploratoryPathCalibrationEpisodeThreshold mdp rounds delta visitFloor < (episodes : Real)) : simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFloor /\\ (Fintype.card State : Real) * uniformFloorTransitionCoordinateRadius mdp episodes (multiBatchLocalDelta rounds delta) visitFloor * (mdp.horizon : Real) <= 1 / 2","missing":[],"search":"episodethreshold_countmargin_and_halfcontraction banditrlproof.finitehorizonrl.episodethreshold_countmargin_and_halfcontraction the episode threshold simultaneously discharges the strict count margin and the half-contraction premise of the explicit path-support calibration route. theorem compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceTransitionBonusCover_of_pathSupport_episodeThreshold","label":"exploratorySource_sourceTransitionBonusCover_of_pathSupport_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceTransitionBonusCover_of_pathSupport_episodeThreshold","description":"The closed-form episode threshold constructs the source-wide transition cover.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-c396fc995993","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8144,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:195"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceTransitionBonusCover_of_pathSupport_episodeThreshold {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hthreshold : exploratoryPathCalibrationEpisodeThreshold mdp rounds delta visitFloor < (episodes : Real)) (hrewardBound_nonneg : 0 <= rewardBound) : let behaviorSource := exploratorySource mdp initialState ep…","missing":[],"search":"exploratorysource_sourcetransitionbonuscover_of_pathsupport_episodethreshold banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcetransitionbonuscover_of_pathsupport_episodethreshold the closed-form episode threshold constructs the source-wide transition cover. theorem compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport_episodeThreshold","label":"exploratorySource_sourceCalibration_of_pathSupport_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport_episodeThreshold","description":"The closed-form episode threshold constructs the full source calibration.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-d49aeac4f4b3","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8145,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:222"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceCalibration_of_pathSupport_episodeThreshold {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hthreshold : exploratoryPathCalibrationEpisodeThreshold mdp rounds delta visitFloor < (episodes : Real)) (hrewardBound_nonneg : 0 <= rewardBound) : let behaviorSource := exploratorySource mdp initialState episodes in…","missing":[],"search":"exploratorysource_sourcecalibration_of_pathsupport_episodethreshold banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcecalibration_of_pathsupport_episodethreshold the closed-form episode threshold constructs the full source calibration. theorem compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_episodeThreshold","label":"exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_episodeThreshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_episodeThreshold","description":"Route endpoint: path support and one explicit episode threshold imply global adaptive count confidence, roundwise optimism, and recommended expected regret.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportepisodethreshold/index.html#decl-adcc10383f51","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","order":8146,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportEpisodeThreshold.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_episodeThreshold {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) (hhorizon : 0 < mdp.horizon) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hvisitFloor : 0 < visitFloor) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hthreshold : exploratoryPathCalibrationEpisodeThreshold mdp round…","missing":[],"search":"exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_episodethreshold banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_episodethreshold route endpoint: path support and one explicit episode threshold imply global adaptive count confidence, roundwise optimism, and recommended expected regret. theorem compiled","shard":"modules/b2cf52ffdd58f3de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorTransitionCoordinateRadius","label":"uniformFloorTransitionCoordinateRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorTransitionCoordinateRadius","description":"Uniform transition-coordinate radius obtained from one common expected-count floor. The denominator is useful when the count radius is strictly smaller than `episodes * visitFloor`.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-e685ae1b5ba6","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8147,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformFloorTransitionCoordinateRadius (mdp : MDP State Action) (episodes : Nat) (delta visitFloor : Real) : Real","missing":[],"search":"uniformfloortransitioncoordinateradius banditrlproof.finitehorizonrl.uniformfloortransitioncoordinateradius uniform transition-coordinate radius obtained from one common expected-count floor. the denominator is useful when the count radius is strictly smaller than `episodes * visitfloor`. definition compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorTransitionCoordinateRadius_nonneg","label":"uniformFloorTransitionCoordinateRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorTransitionCoordinateRadius_nonneg","description":"A positive common denominator makes the uniform coordinate radius nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-5ecbc79aba7e","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8148,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformFloorTransitionCoordinateRadius_nonneg {mdp : MDP State Action} {episodes : Nat} {delta visitFloor : Real} (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < (episodes : Real) * visitFloor) : 0 <= uniformFloorTransitionCoordinateRadius mdp episodes delta visitFloor","missing":[],"search":"uniformfloortransitioncoordinateradius_nonneg banditrlproof.finitehorizonrl.uniformfloortransitioncoordinateradius_nonneg a positive common denominator makes the uniform coordinate radius nonnegative. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCountTransitionCoordinateRadius_le_uniformFloor","label":"expectedCountTransitionCoordinateRadius_le_uniformFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCountTransitionCoordinateRadius_le_uniformFloor","description":"A common expected-count floor bounds every deterministic transition-coordinate radius by the corresponding uniform-denominator radius.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-14342edef504","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8149,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedCountTransitionCoordinateRadius_le_uniformFloor {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta visitFloor : Real) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < (episodes : Real) * visitFloor) (hcountFloor : forall coordinate : VisitCoordinate mdp, (episodes : Real) * visitFloor <= coordinate.expectedCount policy initialState episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.expectedCountTransitionCoordinateRadius initialState episodes delta stage state action nextState <= uniformFloorTransitionCoordinateRadius mdp episodes delta visitFloor","missing":[],"search":"expectedcounttransitioncoordinateradius_le_uniformfloor banditrlproof.finitehorizonrl.markovpolicy.expectedcounttransitioncoordinateradius_le_uniformfloor a common expected-count floor bounds every deterministic transition-coordinate radius by the corresponding uniform-denominator radius. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor","label":"ExploratoryPathUniformVisitFloor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor","description":"One state-action floor shared by every stage and target state on the path certificate.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-904e948fa617","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8150,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:94"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def ExploratoryPathUniformVisitFloor {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (visitFloor : Real) : Prop","missing":[],"search":"exploratorypathuniformvisitfloor banditrlproof.finitehorizonrl.exploratorypathuniformvisitfloor one state-action floor shared by every stage and target state on the path certificate. definition compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor.exploratoryStateCountMargin","label":"exploratoryStateCountMargin","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor.exploratoryStateCountMargin","description":"A strict scalar count inequality discharges the full exploratory margin.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-e37aaae258f0","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8151,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:107"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ExploratoryPathUniformVisitFloor.exploratoryStateCountMargin {mdp : MDP State Action} {initialState : Measure State} {episodes : Nat} {delta : Real} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < (episodes : Real) * visitFloor) : ExploratoryStateCountMargin mdp episodes delta explorationRate (exploratoryPathStateLower support explorationRate)","missing":[],"search":"exploratorystatecountmargin banditrlproof.finitehorizonrl.exploratorypathuniformvisitfloor.exploratorystatecountmargin a strict scalar count inequality discharges the full exploratory margin. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.uniformVisitFloor_expectedCount_le","label":"uniformVisitFloor_expectedCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.uniformVisitFloor_expectedCount_le","description":"Every exploratory table inherits the common expected-count floor.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-3425b6c8503d","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8152,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformVisitFloor_expectedCount_le {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (coordinate : VisitCoordinate mdp) : (episodes : Real) * visitFloor <= coordinate.expectedCount (table.exploratoryPolicy explorationRate hexplorationRate) initialState episodes","missing":[],"search":"uniformvisitfloor_expectedcount_le banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.uniformvisitfloor_expectedcount_le every exploratory table inherits the common expected-count floor. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.transitionBonusCover_rewardBound_of_uniformExpectedCountFloor","label":"transitionBonusCover_rewardBound_of_uniformExpectedCountFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.transitionBonusCover_rewardBound_of_uniformExpectedCountFloor","description":"If the uniform transition coefficient is at most one half, `rewardBound` covers every transition-radius/value-envelope sum when it is used as the transition bonus.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-5b7b568d252f","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8153,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:164"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionBonusCover_rewardBound_of_uniformExpectedCountFloor {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta visitFloor rewardBound : Real) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < (episodes : Real) * visitFloor) (hcountFloor : forall coordinate : VisitCoordinate mdp, (episodes : Real) * visitFloor <= coordinate.expectedCount policy initialState episodes) (hrewardBound_nonneg : 0 <= rewardBound) (hcontraction : (Fintype.card State : Real) * uniformFloorTransitionCoordinateRadius mdp episodes delta visitFloor * (mdp.horizon : Real) <= 1 / 2) : policy.TransitionBonusCover initialState episodes delta rewardBound rewardBound","missing":[],"search":"transitionbonuscover_rewardbound_of_uniformexpectedcountfloor banditrlproof.finitehorizonrl.markovpolicy.transitionbonuscover_rewardbound_of_uniformexpectedcountfloor if the uniform transition coefficient is at most one half, `rewardbound` covers every transition-radius/value-envelope sum when it is used as the transition bonus. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceTransitionBonusCover_of_pathSupport_explicitCalibration","label":"exploratorySource_sourceTransitionBonusCover_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceTransitionBonusCover_of_pathSupport_explicitCalibration","description":"The scalar path-support conditions construct the source-wide bonus cover.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-af88cdebd6e8","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8154,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:225"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceTransitionBonusCover_of_pathSupport_explicitCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hmargin : simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFloor) (hrewardBound_nonneg : 0 <= rewardBound) (hcontraction : (Fintype.card State : Real) * uniformFloorTransitionCoordinateRadius mdp episodes (multiBatchLocalDelta rounds delta) visitFloor * (mdp.horizon : Real) <= 1 / 2) : let behaviorSour…","missing":[],"search":"exploratorysource_sourcetransitionbonuscover_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcetransitionbonuscover_of_pathsupport_explicitcalibration the scalar path-support conditions construct the source-wide bonus cover. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport_explicitCalibration","label":"exploratorySource_sourceCalibration_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport_explicitCalibration","description":"Explicit path support and scalar rate conditions construct `SourceCalibration`.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-f8ad26daf63b","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8155,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceCalibration_of_pathSupport_explicitCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hmargin : simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFloor) (hrewardBound_nonneg : 0 <= rewardBound) (hcontraction : (Fintype.card State : Real) * uniformFloorTransitionCoordinateRadius mdp episodes (multiBatchLocalDelta rounds delta) visitFloor * (mdp.horizon : Real) <= 1 / 2) : let behaviorSource := exp…","missing":[],"search":"exploratorysource_sourcecalibration_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcecalibration_of_pathsupport_explicitcalibration explicit path support and scalar rate conditions construct `sourcecalibration`. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","label":"exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","description":"Route endpoint: path support plus explicit scalar count and half-contraction conditions yield the adaptive global confidence, optimism, and recommended expected-regret theorem with `transitionBonus = rewardBound`.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportexplicitcalibration/index.html#decl-0c2e75c5443e","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","order":8156,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportExplicitCalibration.lean:306"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (support : ExploratoryPathSupport mdp initialState) (visitFloor : Real) (hfloor : ExploratoryPathUniformVisitFloor support explorationRate visitFloor) (hmargin : simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < (episodes : Real) * visitFl…","missing":[],"search":"exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport_explicitcalibration route endpoint: path support plus explicit scalar count and half-contraction conditions yield the adaptive global confidence, optimism, and recommended expected-regret theorem with `transitionbonus = rewardbound`. theorem compiled","shard":"modules/eb62366eb7df2f7f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryStateAt_zero","label":"trajectoryStateAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryStateAt_zero","description":"theorem trajectoryStateAt_zero (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (hhorizon : 0 < mdp.horizon) : mdp.trajectoryStateAt trajectory ⟨0, hhorizon⟩ = trajectory.1","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-5b4a186732c9","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8157,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryStateAt_zero (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (hhorizon : 0 < mdp.horizon) : mdp.trajectoryStateAt trajectory ⟨0, hhorizon⟩ = trajectory.1","missing":[],"search":"trajectorystateat_zero banditrlproof.finitehorizonrl.mdp.trajectorystateat_zero theorem trajectorystateat_zero (mdp : mdp state action) (trajectory : state × steptrace action state mdp.horizon) (hhorizon : 0 < mdp.horizon) : mdp.trajectorystateat trajectory ⟨0, hhorizon⟩ = trajectory.1 theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeStepOfTrajectory_nextState_eq_trajectoryStateAt_succ","label":"episodeStepOfTrajectory_nextState_eq_trajectoryStateAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeStepOfTrajectory_nextState_eq_trajectoryStateAt_succ","description":"theorem episodeStepOfTrajectory_nextState_eq_trajectoryStateAt_succ (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (stage : Nat) (hstage : stage + 1 < mdp.horizon) : (mdp.episodeStepOfTrajectory trajectory ⟨stage, by omega⟩).nextState = mdp.trajectoryStateAt trajectory ⟨stage + 1, hstage⟩","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-06b62eb5a416","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8158,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeStepOfTrajectory_nextState_eq_trajectoryStateAt_succ (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (stage : Nat) (hstage : stage + 1 < mdp.horizon) : (mdp.episodeStepOfTrajectory trajectory ⟨stage, by omega⟩).nextState = mdp.trajectoryStateAt trajectory ⟨stage + 1, hstage⟩","missing":[],"search":"episodestepoftrajectory_nextstate_eq_trajectorystateat_succ banditrlproof.finitehorizonrl.mdp.episodestepoftrajectory_nextstate_eq_trajectorystateat_succ theorem episodestepoftrajectory_nextstate_eq_trajectorystateat_succ (mdp : mdp state action) (trajectory : state × steptrace action state mdp.horizon) (stage : nat) (hstage : stage + 1 < mdp.horizon) : (mdp.episodestepoftrajectory trajectory ⟨stage, by omega⟩).nextstate = mdp.trajectorystateat trajectory ⟨stage + 1, hstage⟩ theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability_zero","label":"stageStateProbability_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability_zero","description":"The generated state mass at stage zero is the initial singleton mass.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-5e36b28669d7","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8159,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageStateProbability_zero {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (hhorizon : 0 < mdp.horizon) (state : State) : policy.stageStateProbability initialState ⟨0, hhorizon⟩ state = initialState.real {state}","missing":[],"search":"stagestateprobability_zero banditrlproof.finitehorizonrl.markovpolicy.stagestateprobability_zero the generated state mass at stage zero is the initial singleton mass. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_le_stageStateProbability_succ","label":"stageTransitionJointProbability_le_stageStateProbability_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_le_stageStateProbability_succ","description":"A selected transition event is contained in its successor state event.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-76ecb5502b23","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8160,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageTransitionJointProbability_le_stageStateProbability_succ {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Nat) (hstage : stage + 1 < mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.stageTransitionJointProbability initialState ⟨stage, by omega⟩ state action nextState <= policy.stageStateProbability initialState ⟨stage + 1, hstage⟩ nextState","missing":[],"search":"stagetransitionjointprobability_le_stagestateprobability_succ banditrlproof.finitehorizonrl.markovpolicy.stagetransitionjointprobability_le_stagestateprobability_succ a selected transition event is contained in its successor state event. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathSupport","label":"ExploratoryPathSupport","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ExploratoryPathSupport","description":"One explicit predecessor path for each successor-stage state, together with nonnegative singleton-mass floors along those paths.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-d7718caf9342","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8161,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure ExploratoryPathSupport (mdp : MDP State Action) (initialState : Measure State) where","missing":[],"search":"exploratorypathsupport banditrlproof.finitehorizonrl.exploratorypathsupport one explicit predecessor path for each successor-stage state, together with nonnegative singleton-mass floors along those paths. structure compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat","label":"exploratoryPathStateLowerNat","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat","description":"Recursive state floor along the selected predecessor paths.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-8982b8e3b03a","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8162,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:121"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryPathStateLowerNat {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) : (stage : Nat) -> stage < mdp.horizon -> State -> Real | 0, _hstage, state => support.initialFloor state | stage + 1, hstage, state => exploratoryPathStateLowerNat support explorationRate stage (by omega) (support.predecessorState ⟨stage + 1, hstage⟩ state) * exploratoryActionProbabilityFloor Action explorationRate * support.transitionFloor ⟨stage + 1, hstage⟩ state /-- The selected-path lower envelope indexed by valid chronological stages. -/ noncomputable def exploratoryPathStateLower {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (stage : Fin mdp.horizon) (state : State) : Real","missing":[],"search":"exploratorypathstatelowernat banditrlproof.finitehorizonrl.exploratorypathstatelowernat recursive state floor along the selected predecessor paths. definition compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLower","label":"exploratoryPathStateLower","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathStateLower","description":"The selected-path lower envelope indexed by valid chronological stages.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-2615c94e77f0","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8163,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:134"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryPathStateLower {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (stage : Fin mdp.horizon) (state : State) : Real","missing":[],"search":"exploratorypathstatelower banditrlproof.finitehorizonrl.exploratorypathstatelower the selected-path lower envelope indexed by valid chronological stages. definition compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat_nonneg","label":"exploratoryPathStateLowerNat_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat_nonneg","description":"Every recursively constructed path-support floor is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-706fb7f55f99","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8164,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPathStateLowerNat_nonneg {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) : forall (stage : Nat) (hstage : stage < mdp.horizon) (state : State), 0 <= exploratoryPathStateLowerNat support explorationRate stage hstage state","missing":[],"search":"exploratorypathstatelowernat_nonneg banditrlproof.finitehorizonrl.exploratorypathstatelowernat_nonneg every recursively constructed path-support floor is nonnegative. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLower_nonneg","label":"exploratoryPathStateLower_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryPathStateLower_nonneg","description":"The Fin-indexed selected-path envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-d8db0ea8830c","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8165,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPathStateLower_nonneg {mdp : MDP State Action} {initialState : Measure State} (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (stage : Fin mdp.horizon) (state : State) : 0 <= exploratoryPathStateLower support explorationRate stage state","missing":[],"search":"exploratorypathstatelower_nonneg banditrlproof.finitehorizonrl.exploratorypathstatelower_nonneg the fin-indexed selected-path envelope is nonnegative. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPathStateLower_le_stageStateProbability","label":"exploratoryPathStateLower_le_stageStateProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPathStateLower_le_stageStateProbability","description":"Every exploratory table policy dominates the same explicit path-support state envelope, independently of the table's deterministic center.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-1023166fb6e4","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8166,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPathStateLower_le_stageStateProbability {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (support : ExploratoryPathSupport mdp initialState) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) : exploratoryPathStateLower support explorationRate stage state <= (table.exploratoryPolicy explorationRate hexplorationRate).stageStateProbability initialState stage state","missing":[],"search":"exploratorypathstatelower_le_stagestateprobability banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratorypathstatelower_le_stagestateprobability every exploratory table policy dominates the same explicit path-support state envelope, independently of the table's deterministic center. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceStateReachability_of_pathSupport","label":"exploratorySource_sourceStateReachability_of_pathSupport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceStateReachability_of_pathSupport","description":"Every policy selected by the exploratory source shares the path-support floor.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-59322a5628cc","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8167,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceStateReachability_of_pathSupport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (support : ExploratoryPathSupport mdp initialState) : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate SourceStateReachability behaviorSource rounds (exploratoryPathStateLower support explorationRate)","missing":[],"search":"exploratorysource_sourcestatereachability_of_pathsupport banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcestatereachability_of_pathsupport every policy selected by the exploratory source shares the path-support floor. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport","label":"exploratorySource_sourceCalibration_of_pathSupport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport","description":"Explicit path support constructs the exact adaptive source calibration.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-208e5de40baa","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8168,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceCalibration_of_pathSupport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (support : ExploratoryPathSupport mdp initialState) (hmargin : ExploratoryStateCountMargin mdp episodes (multiBatchLocalDelta rounds delta) explorationRate (exploratoryPathStateLower support explorationRate)) (hcover : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate SourceTransitionBonusCover behaviorSource rounds (multiBatchLocalDelta rounds delta) rewardBound transitionBonus) : let behaviorSource := exploratorySource mdp initialSt…","missing":[],"search":"exploratorysource_sourcecalibration_of_pathsupport banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcecalibration_of_pathsupport explicit path support constructs the exact adaptive source calibration. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport","label":"exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport","description":"Route endpoint: explicit initial and transition path support replaces the abstract source state-reachability premise in the adaptive global theorem.","url":"../modules/banditrlproof-rl-finitehorizonexploratorypathsupportreachability/index.html#decl-e152b20e0997","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","order":8169,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryPathSupportReachability.lean:327"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (support : ExploratoryPathSupport mdp initialState) (hmargin : ExploratoryStateCountMargin mdp episodes (multiBatchLocalDelta rounds delta) explorationRate (exploratoryPathStateLower support explorationRate)) (hcover : let behavi…","missing":[],"search":"exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_pathsupport route endpoint: explicit initial and transition path support replaces the abstract source state-reachability premise in the adaptive global theorem. theorem compiled","shard":"modules/cb6c725009331d24.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor","label":"exploratoryActionProbabilityFloor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor","description":"Real probability floor contributed by uniform exploration.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-e864df2b6439","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8170,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def exploratoryActionProbabilityFloor (Action : Type v) [Fintype Action] (explorationRate : NNReal) : Real","missing":[],"search":"exploratoryactionprobabilityfloor banditrlproof.finitehorizonrl.exploratoryactionprobabilityfloor real probability floor contributed by uniform exploration. definition compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor_nonneg","label":"exploratoryActionProbabilityFloor_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor_nonneg","description":"theorem exploratoryActionProbabilityFloor_nonneg (explorationRate : NNReal) : 0 <= exploratoryActionProbabilityFloor Action explorationRate","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-bf43b44a6a1f","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8171,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryActionProbabilityFloor_nonneg (explorationRate : NNReal) : 0 <= exploratoryActionProbabilityFloor Action explorationRate","missing":[],"search":"exploratoryactionprobabilityfloor_nonneg banditrlproof.finitehorizonrl.exploratoryactionprobabilityfloor_nonneg theorem exploratoryactionprobabilityfloor_nonneg (explorationrate : nnreal) : 0 <= exploratoryactionprobabilityfloor action explorationrate theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability_nonneg","label":"stageStateProbability_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability_nonneg","description":"Every generated stage-state probability is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-23955dbc62bc","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8172,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:48"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageStateProbability_nonneg {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) : 0 <= policy.stageStateProbability initialState stage state","missing":[],"search":"stagestateprobability_nonneg banditrlproof.finitehorizonrl.markovpolicy.stagestateprobability_nonneg every generated stage-state probability is nonnegative. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.TransitionBonusCover","label":"TransitionBonusCover","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.TransitionBonusCover","description":"The transition-coordinate cover field retained by empirical calibration.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-0032aeada375","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8173,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def TransitionBonusCover {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta rewardBound transitionBonus : Real) : Prop","missing":[],"search":"transitionbonuscover banditrlproof.finitehorizonrl.markovpolicy.transitionbonuscover the transition-coordinate cover field retained by empirical calibration. definition compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionProbabilityFloor_le_exploratoryActionPMF_toReal","label":"exploratoryActionProbabilityFloor_le_exploratoryActionPMF_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionProbabilityFloor_le_exploratoryActionPMF_toReal","description":"The ENNReal exploratory PMF floor as a Real inequality.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-bc52119ab836","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8174,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:81"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryActionProbabilityFloor_le_exploratoryActionPMF_toReal {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (action : Action) : exploratoryActionProbabilityFloor Action explorationRate <= (table.exploratoryActionPMF explorationRate hexplorationRate stage state action).toReal","missing":[],"search":"exploratoryactionprobabilityfloor_le_exploratoryactionpmf_toreal banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryactionprobabilityfloor_le_exploratoryactionpmf_toreal the ennreal exploratory pmf floor as a real inequality. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionProbabilityFloor_le_actionKernel_real","label":"exploratoryActionProbabilityFloor_le_actionKernel_real","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionProbabilityFloor_le_actionKernel_real","description":"The exploratory Markov kernel inherits the Real singleton floor.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-3da770b67eb6","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8175,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryActionProbabilityFloor_le_actionKernel_real {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stage : Fin mdp.horizon) (state : State) (action : Action) : exploratoryActionProbabilityFloor Action explorationRate <= ((table.exploratoryPolicy explorationRate hexplorationRate).actionKernel stage state {action}).toReal","missing":[],"search":"exploratoryactionprobabilityfloor_le_actionkernel_real banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratoryactionprobabilityfloor_le_actionkernel_real the exploratory markov kernel inherits the real singleton floor. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.stateLower_mul_exploratoryActionProbabilityFloor_le_stageVisitProbability","label":"stateLower_mul_exploratoryActionProbabilityFloor_le_stageVisitProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.stateLower_mul_exploratoryActionProbabilityFloor_le_stageVisitProbability","description":"State reachability times the exploratory action floor lower-bounds visits.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-4a56a9cdc9d8","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8176,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateLower_mul_exploratoryActionProbabilityFloor_le_stageVisitProbability {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stateLower : Fin mdp.horizon -> State -> Real) (hstateLower : forall stage state, stateLower stage state <= (table.exploratoryPolicy explorationRate hexplorationRate).stageStateProbability initialState stage state) (stage : Fin mdp.horizon) (state : State) (action : Action) : stateLower stage state * exploratoryActionProbabilityFloor Action explorationRate <= (table.exploratoryPolicy explorationRate hexplorationRate).stageVisitProbability initialState stage state action","missing":[],"search":"statelower_mul_exploratoryactionprobabilityfloor_le_stagevisitprobability banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.statelower_mul_exploratoryactionprobabilityfloor_le_stagevisitprobability state reachability times the exploratory action floor lower-bounds visits. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.stateLower_expectedCount_le","label":"stateLower_expectedCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.stateLower_expectedCount_le","description":"The state/action floor also lower-bounds the genuine expected visit count.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-ea7b4f0b8b59","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8177,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateLower_expectedCount_le {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stateLower : Fin mdp.horizon -> State -> Real) (hstateLower : forall stage state, stateLower stage state <= (table.exploratoryPolicy explorationRate hexplorationRate).stageStateProbability initialState stage state) (coordinate : VisitCoordinate mdp) : (episodes : Real) * (stateLower coordinate.stage coordinate.state * exploratoryActionProbabilityFloor Action explorationRate) <= coordinate.expectedCount (table.exploratoryPolicy explorationRate hexplorationRate) initialState episodes","missing":[],"search":"statelower_expectedcount_le banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.statelower_expectedcount_le the state/action floor also lower-bounds the genuine expected visit count. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryStateCountMargin","label":"ExploratoryStateCountMargin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ExploratoryStateCountMargin","description":"Strict count margin implied by a state lower envelope and exploration.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-cc8a7a480d6c","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8178,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def ExploratoryStateCountMargin (mdp : MDP State Action) (episodes : Nat) (delta : Real) (explorationRate : NNReal) (stateLower : Fin mdp.horizon -> State -> Real) : Prop","missing":[],"search":"exploratorystatecountmargin banditrlproof.finitehorizonrl.exploratorystatecountmargin strict count margin implied by a state lower envelope and exploration. definition compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalOptimisticCalibration_exploratoryPolicy_of_stateReachability","label":"empiricalOptimisticCalibration_exploratoryPolicy_of_stateReachability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalOptimisticCalibration_exploratoryPolicy_of_stateReachability","description":"Reachability plus the unchanged cover constructs policy-local calibration.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-ac40f023c937","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8179,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def empiricalOptimisticCalibration_exploratoryPolicy_of_stateReachability {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta rewardBound transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (stateLower : Fin mdp.horizon -> State -> Real) (hstateLower : forall stage state, stateLower stage state <= (table.exploratoryPolicy explorationRate hexplorationRate).stageStateProbability initialState stage state) (hmargin : ExploratoryStateCountMargin mdp episodes delta explorationRate stateLower) (hcover : (table.exploratoryPolicy explorationRate hexplorationRate).TransitionBonusCover initialState episodes delta rewardBound transitionBonus) : (table.exploratoryPolicy explorationRate hexplorationRate).EmpiricalOptimisticCalibration initialState episo…","missing":[],"search":"empiricaloptimisticcalibration_exploratorypolicy_of_statereachability banditrlproof.finitehorizonrl.markovpolicy.empiricaloptimisticcalibration_exploratorypolicy_of_statereachability reachability plus the unchanged cover constructs policy-local calibration. definition compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceStateReachability","label":"SourceStateReachability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceStateReachability","description":"State-only reachability envelope for every batch-generating source policy.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-cb395f13bf32","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8180,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:216"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def SourceStateReachability {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (stateLower : Fin mdp.horizon -> State -> Real) : Prop","missing":[],"search":"sourcestatereachability banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.sourcestatereachability state-only reachability envelope for every batch-generating source policy. definition compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceTransitionBonusCover","label":"SourceTransitionBonusCover","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceTransitionBonusCover","description":"Transition-bonus cover for every batch-generating source policy.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-9ded2214ff55","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8181,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def SourceTransitionBonusCover {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (source : AdaptiveEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (delta rewardBound transitionBonus : Real) : Prop","missing":[],"search":"sourcetransitionbonuscover banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.sourcetransitionbonuscover transition-bonus cover for every batch-generating source policy. definition compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_stateReachability","label":"exploratorySource_sourceCalibration_of_stateReachability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_stateReachability","description":"State reachability and exploration discharge every expected-count margin in the exact adaptive calibration contract; the transition-bonus cover is preserved.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-75c1b262e21f","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8182,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:248"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_sourceCalibration_of_stateReachability {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus delta : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (rounds : Nat) (stateLower : Fin mdp.horizon -> State -> Real) (hmargin : ExploratoryStateCountMargin mdp episodes (multiBatchLocalDelta rounds delta) explorationRate stateLower) (hreachability : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorationRate SourceStateReachability behaviorSource rounds stateLower) (hcover : let behaviorSource := exploratorySource mdp initialState episodes initialTable defaultState transitionBonus explorationRate hexplorat…","missing":[],"search":"exploratorysource_sourcecalibration_of_statereachability banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_sourcecalibration_of_statereachability state reachability and exploration discharge every expected-count margin in the exact adaptive calibration contract; the transition-bonus cover is preserved. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_stateReachability","label":"exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_stateReachability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_stateReachability","description":"Route endpoint: state reachability plus uniform exploration discharges the calibration premise of the adaptive confidence, optimism, and recommended expected-regret theorem.","url":"../modules/banditrlproof-rl-finitehorizonexploratoryreachabilitycalibration/index.html#decl-023d9d31fa1f","parent":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","order":8183,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration"],["Source","BanditRLProof/RL/FiniteHorizonExploratoryReachabilityCalibration.lean:294"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_stateReachability {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat} (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (rewardBound transitionBonus : Real) (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (htransitionBonus_nonneg : 0 <= transitionBonus) (rounds : Nat) (hrounds : 0 < rounds) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (stateLower : Fin mdp.horizon -> State -> Real) (hmargin : ExploratoryStateCountMargin mdp episodes (multiBatchLocalDelta rounds delta) explorationRate stateLower) (hreachability : let behaviorSource := exploratorySource md…","missing":[],"search":"exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_statereachability banditrlproof.finitehorizonrl.adaptiveempiricaloptimisticsource.exploratorysource_trajectorymeasure_allcoordinateconfidence_optimism_and_recommendedexpectedregret_of_statereachability route endpoint: state reachability plus uniform exploration discharges the calibration premise of the adaptive confidence, optimism, and recommended expected-regret theorem. theorem compiled","shard":"modules/305f505d9aa83d08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.empiricalFiniteBatchValueEnvelope","label":"empiricalFiniteBatchValueEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.empiricalFiniteBatchValueEnvelope","description":"Explicit noncircular envelope for a fixed reward and transition budget.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-c1c0b0e99619","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8184,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def empiricalFiniteBatchValueEnvelope (rewardBound transitionBudget : Real) (remaining : Nat) : Real","missing":[],"search":"empiricalfinitebatchvalueenvelope banditrlproof.finitehorizonrl.empiricalfinitebatchvalueenvelope explicit noncircular envelope for a fixed reward and transition budget. definition compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.allCoordinateEmpiricalFiniteBatchModel","label":"allCoordinateEmpiricalFiniteBatchModel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.allCoordinateEmpiricalFiniteBatchModel","description":"Canonical empirical model for the all-coordinate route: exact generated rewards use radius zero, while every transition coordinate shares one fixed budget.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-3cc533207b1a","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8185,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def allCoordinateEmpiricalFiniteBatchModel (mdp : MDP State Action) (episodes : Nat) (batch : EpisodeBatch mdp episodes) (defaultState : State) (transitionBudget : Real) : FiniteBatchModel mdp episodes where","missing":[],"search":"allcoordinateempiricalfinitebatchmodel banditrlproof.finitehorizonrl.mdp.allcoordinateempiricalfinitebatchmodel canonical empirical model for the all-coordinate route: exact generated rewards use radius zero, while every transition coordinate shares one fixed budget. definition compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCountTransitionCoordinateRadius","label":"expectedCountTransitionCoordinateRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCountTransitionCoordinateRadius","description":"Deterministic lower-margin coordinate radius based on the genuine visit mean.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-604314fbceb1","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8186,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedCountTransitionCoordinateRadius {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta : Real) (stage : Fin mdp.horizon) (state : State) (action : Action) (_nextState : State) : Real","missing":[],"search":"expectedcounttransitioncoordinateradius banditrlproof.finitehorizonrl.markovpolicy.expectedcounttransitioncoordinateradius deterministic lower-margin coordinate radius based on the genuine visit mean. definition compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCount_sub_radius_lt_count_of_not_mem_simultaneousCountBadEvent","label":"expectedCount_sub_radius_lt_count_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCount_sub_radius_lt_count_of_not_mem_simultaneousCountBadEvent","description":"Outside the simultaneous event, realized count exceeds its deterministic lower margin.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-36648693fdfa","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8187,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedCount_sub_radius_lt_count_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (coordinate : VisitCoordinate mdp) : coordinate.expectedCount policy initialState episodes - simultaneousCountConfidenceRadius mdp episodes delta < (coordinate.count batch : Real)","missing":[],"search":"expectedcount_sub_radius_lt_count_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.expectedcount_sub_radius_lt_count_of_not_mem_simultaneouscountbadevent outside the simultaneous event, realized count exceeds its deterministic lower margin. theorem compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalTransitionMass_abs_sub_transition_le_expectedCountRadius_of_not_mem","label":"empiricalTransitionMass_abs_sub_transition_le_expectedCountRadius_of_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalTransitionMass_abs_sub_transition_le_expectedCountRadius_of_not_mem","description":"The random-denominator transition error is bounded by the deterministic expected-count lower-margin radius.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-af203de18209","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8188,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:94"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionMass_abs_sub_transition_le_expectedCountRadius_of_not_mem {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (defaultState : State) (coordinate : VisitCoordinate mdp) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) (nextState : State) : |batch.empiricalTransitionMass defaultState coordinate.stage coordinate.state coordinate.action nextState - (mdp.transition (coordinate.state, coordinate.action)).real {nextState}| ≤ policy.expectedCountTransitionCoordinateRadius initialState episodes delta coordinate.stage coordinate.state coordinate.action nextState","missing":[],"search":"empiricaltransitionmass_abs_sub_transition_le_expectedcountradius_of_not_mem banditrlproof.finitehorizonrl.markovpolicy.empiricaltransitionmass_abs_sub_transition_le_expectedcountradius_of_not_mem the random-denominator transition error is bounded by the deterministic expected-count lower-margin radius. theorem compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalReward_eq_of_not_mem_simultaneousCountBadEvent","label":"empiricalReward_eq_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalReward_eq_of_not_mem_simultaneousCountBadEvent","description":"Full-coordinate margins make every reward-consistent empirical reward exact.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-1653c4422db7","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8189,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalReward_eq_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (hreward : batch.RewardConsistent) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : batch.empiricalReward stage state action = mdp.reward state action","missing":[],"search":"empiricalreward_eq_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.empiricalreward_eq_of_not_mem_simultaneouscountbadevent full-coordinate margins make every reward-consistent empirical reward exact. theorem compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.AllCoordinateConfidence.upperValueRemaining_abs_le","label":"upperValueRemaining_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.AllCoordinateConfidence.upperValueRemaining_abs_le","description":"Exact empirical rewards and fixed nonnegative budgets give the explicit linear absolute envelope for every recursive optimistic value.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-8ef4d090a9b4","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8190,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:171"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem upperValueRemaining_abs_le (hrewardExact : ∀ stage state action, batch.empiricalReward stage state action = mdp.reward state action) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ rewardBound) (htransitionBudget_nonneg : 0 ≤ transitionBudget) : ∀ (remaining : Nat) (hremaining : remaining ≤ mdp.horizon) (state : State), |(mdp.allCoordinateEmpiricalFiniteBatchModel episodes batch defaultState transitionBudget).plan.upperValueRemaining remaining hremaining state| ≤ empiricalFiniteBatchValueEnvelope rewardBound transitionBudget remaining","missing":[],"search":"uppervalueremaining_abs_le banditrlproof.finitehorizonrl.markovpolicy.allcoordinateconfidence.uppervalueremaining_abs_le exact empirical rewards and fixed nonnegative budgets give the explicit linear absolute envelope for every recursive optimistic value. theorem compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.allCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","label":"allCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.allCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","description":"Pathwise producer: full genuine occupancy margins, deterministic radius cover, and reward consistency construct the complete raw finite-batch confidence object.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-e27191dd2f6e","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8191,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def allCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (hreward : batch.RewardConsistent) (defaultState : State) (rewardBound transitionBudget : Real) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ rewardBound) (htransitionBudget_nonneg : 0 ≤ transitionBudget) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) (hcover : ∀ (remaining : Nat) (hremaining : remaining + 1 ≤ mdp.horizon) (state : State) (action : Action), (∑ nextState, policy.expectedCountTransitionCoordinateRadius initialS…","missing":[],"search":"allcoordinateempiricalfinitebatchmodelconfidence_of_not_mem banditrlproof.finitehorizonrl.markovpolicy.allcoordinateempiricalfinitebatchmodelconfidence_of_not_mem pathwise producer: full genuine occupancy margins, deterministic radius cover, and reward consistency construct the complete raw finite-batch confidence object. definition compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_allCoordinate_finiteBatchModel_confidence","label":"iidEpisodeBatch_allCoordinate_finiteBatchModel_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_allCoordinate_finiteBatchModel_confidence","description":"Mapped-iid confidence endpoint with the unchanged simultaneous-event failure budget. No additional reward or confidence event is introduced.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-91a1f0aa0702","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8192,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:315"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_allCoordinate_finiteBatchModel_confidence {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) (defaultState : State) (rewardBound transitionBudget : Real) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ rewardBound) (htransitionBudget_nonneg : 0 ≤ transitionBudget) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) (hcover : ∀ (remaining : Nat) (hremaining : remaining + 1 ≤ mdp.horizon) (state : State) (action : Action), (∑ nextState, policy.expectedCountTransitionCoordinateRadius initialState episodes delta (mdp.decisionStageRemaining remaining hremaining) state action next…","missing":[],"search":"iidepisodebatch_allcoordinate_finitebatchmodel_confidence banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_allcoordinate_finitebatchmodel_confidence mapped-iid confidence endpoint with the unchanged simultaneous-event failure budget. no additional reward or confidence event is introduced. theorem compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_allCoordinate_optimism_and_expectedRegret","label":"iidEpisodeBatch_allCoordinate_optimism_and_expectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_allCoordinate_optimism_and_expectedRegret","description":"The produced confidence object immediately yields global optimism and the existing selected-radius single-episode expected-regret bound almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizoniidallcoordinatefinitebatchconfidence/index.html#decl-6276ae57224e","parent":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","order":8193,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDAllCoordinateFiniteBatchConfidence.lean:362"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_allCoordinate_optimism_and_expectedRegret {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) (defaultState : State) (rewardBound transitionBudget : Real) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ rewardBound) (htransitionBudget_nonneg : 0 ≤ transitionBudget) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) (hcover : ∀ (remaining : Nat) (hremaining : remaining + 1 ≤ mdp.horizon) (state : State) (action : Action), (∑ nextState, policy.expectedCountTransitionCoordinateRadius initialState episodes delta (mdp.decisionStageRemaining remaining hremaining) state action next…","missing":[],"search":"iidepisodebatch_allcoordinate_optimism_and_expectedregret banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_allcoordinate_optimism_and_expectedregret the produced confidence object immediately yields global optimism and the existing selected-radius single-episode expected-regret bound almost everywhere. theorem compiled","shard":"modules/b60cbc42af79d2ee.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.visitIndicator","label":"visitIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.visitIndicator","description":"Real indicator that one record visits a fixed state-action coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-9192cdb9ac89","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8194,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def visitIndicator (state : State) (action : Action) (step : EpisodeStep State Action) : Real","missing":[],"search":"visitindicator banditrlproof.finitehorizonrl.episodestep.visitindicator real indicator that one record visits a fixed state-action coordinate. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.transitionIndicator","label":"transitionIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.transitionIndicator","description":"Real indicator that one record realizes a fixed transition coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-6599b742dd99","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8195,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def transitionIndicator (state : State) (action : Action) (nextState : State) (step : EpisodeStep State Action) : Real","missing":[],"search":"transitionindicator banditrlproof.finitehorizonrl.episodestep.transitionindicator real indicator that one record realizes a fixed transition coordinate. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_visitIndicator","label":"measurable_visitIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_visitIndicator","description":"A fixed visit indicator is measurable on empirical records.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-63606b24e17b","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8196,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_visitIndicator (state : State) (action : Action) : Measurable (visitIndicator state action)","missing":[],"search":"measurable_visitindicator banditrlproof.finitehorizonrl.episodestep.measurable_visitindicator a fixed visit indicator is measurable on empirical records. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_transitionIndicator","label":"measurable_transitionIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_transitionIndicator","description":"A fixed transition indicator is measurable on empirical records.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-72593320fd8d","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8197,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionIndicator (state : State) (action : Action) (nextState : State) : Measurable (transitionIndicator state action nextState)","missing":[],"search":"measurable_transitionindicator banditrlproof.finitehorizonrl.episodestep.measurable_transitionindicator a fixed transition indicator is measurable on empirical records. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.visitIndicator_mem_Icc","label":"visitIndicator_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.visitIndicator_mem_Icc","description":"Every visit indicator lies in the Hoeffding interval `[0,1]`.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-d85e4ff2e5e5","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8198,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem visitIndicator_mem_Icc (state : State) (action : Action) (step : EpisodeStep State Action) : visitIndicator state action step ∈ Set.Icc (0 : Real) 1","missing":[],"search":"visitindicator_mem_icc banditrlproof.finitehorizonrl.episodestep.visitindicator_mem_icc every visit indicator lies in the hoeffding interval `[0,1]`. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.transitionIndicator_mem_Icc","label":"transitionIndicator_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeStep.transitionIndicator_mem_Icc","description":"Every transition indicator lies in the Hoeffding interval `[0,1]`.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-4a08c835c0b4","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8199,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionIndicator_mem_Icc (state : State) (action : Action) (nextState : State) (step : EpisodeStep State Action) : transitionIndicator state action nextState step ∈ Set.Icc (0 : Real) 1","missing":[],"search":"transitionindicator_mem_icc banditrlproof.finitehorizonrl.episodestep.transitionindicator_mem_icc every transition indicator lies in the hoeffding interval `[0,1]`. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability","label":"stageVisitProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability","description":"Genuine single-trajectory mean of a fixed stage/state/action visit indicator.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-043e3e6cbb71","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8200,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:91"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stageVisitProbability {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) : Real","missing":[],"search":"stagevisitprobability banditrlproof.finitehorizonrl.markovpolicy.stagevisitprobability genuine single-trajectory mean of a fixed stage/state/action visit indicator. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability","label":"stageTransitionJointProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability","description":"Genuine single-trajectory joint probability of a fixed `(state, action, nextState)` coordinate. This is not a conditional transition probability given the current state and action.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-8ac75d92519f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8201,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stageTransitionJointProbability {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : Real","missing":[],"search":"stagetransitionjointprobability banditrlproof.finitehorizonrl.markovpolicy.stagetransitionjointprobability genuine single-trajectory joint probability of a fixed `(state, action, nextstate)` coordinate. this is not a conditional transition probability given the current state and action. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_eq_measureReal","label":"stageVisitProbability_eq_measureReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_eq_measureReal","description":"A stage visit mean is the real mass of its measurable trajectory event.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-88e819c7c06f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8202,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageVisitProbability_eq_measureReal {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) : policy.stageVisitProbability initialState stage state action = (policy.trajectoryMeasure initialState).real {trajectory | (mdp.episodeStepOfTrajectory trajectory stage).state = state /\\ (mdp.episodeStepOfTrajectory trajectory stage).action = action}","missing":[],"search":"stagevisitprobability_eq_measurereal banditrlproof.finitehorizonrl.markovpolicy.stagevisitprobability_eq_measurereal a stage visit mean is the real mass of its measurable trajectory event. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_eq_measureReal","label":"stageTransitionJointProbability_eq_measureReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_eq_measureReal","description":"A stage joint-transition mean is the real mass of its measurable event.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-66113a717034","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8203,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageTransitionJointProbability_eq_measureReal {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.stageTransitionJointProbability initialState stage state action nextState = (policy.trajectoryMeasure initialState).real {trajectory | (mdp.episodeStepOfTrajectory trajectory stage).state = state /\\ (mdp.episodeStepOfTrajectory trajectory stage).action = action /\\ (mdp.episodeStepOfTrajectory trajectory stage).nextState = nextState}","missing":[],"search":"stagetransitionjointprobability_eq_measurereal banditrlproof.finitehorizonrl.markovpolicy.stagetransitionjointprobability_eq_measurereal a stage joint-transition mean is the real mass of its measurable event. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_mem_Icc","label":"stageVisitProbability_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_mem_Icc","description":"Every genuine stage visit probability lies in `[0,1]`.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-7139a87413d3","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8204,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:171"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageVisitProbability_mem_Icc {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) : policy.stageVisitProbability initialState stage state action ∈ Set.Icc (0 : Real) 1","missing":[],"search":"stagevisitprobability_mem_icc banditrlproof.finitehorizonrl.markovpolicy.stagevisitprobability_mem_icc every genuine stage visit probability lies in `[0,1]`. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_mem_Icc","label":"stageTransitionJointProbability_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_mem_Icc","description":"Every genuine stage joint-transition probability lies in `[0,1]`.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-121e285b4c6d","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8205,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:181"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageTransitionJointProbability_mem_Icc {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.stageTransitionJointProbability initialState stage state action nextState ∈ Set.Icc (0 : Real) 1","missing":[],"search":"stagetransitionjointprobability_mem_icc banditrlproof.finitehorizonrl.markovpolicy.stagetransitionjointprobability_mem_icc every genuine stage joint-transition probability lies in `[0,1]`. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_le_stageVisitProbability","label":"stageTransitionJointProbability_le_stageVisitProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_le_stageVisitProbability","description":"A fixed joint-transition probability is bounded by its visit probability.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-13809aed20b2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8206,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageTransitionJointProbability_le_stageVisitProbability {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.stageTransitionJointProbability initialState stage state action nextState ≤ policy.stageVisitProbability initialState stage state action","missing":[],"search":"stagetransitionjointprobability_le_stagevisitprobability banditrlproof.finitehorizonrl.markovpolicy.stagetransitionjointprobability_le_stagevisitprobability a fixed joint-transition probability is bounded by its visit probability. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_visitIndicator_iidEpisodeBatchMeasure_eval","label":"integral_visitIndicator_iidEpisodeBatchMeasure_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_visitIndicator_iidEpisodeBatchMeasure_eval","description":"Every mapped episode coordinate has the common genuine visit-indicator mean.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-b20a7aeb0834","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8207,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_visitIndicator_iidEpisodeBatchMeasure_eval {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : integral (policy.iidEpisodeBatchMeasure initialState episodes) (fun batch => EpisodeStep.visitIndicator state action (batch episode stage)) = policy.stageVisitProbability initialState stage state action","missing":[],"search":"integral_visitindicator_iidepisodebatchmeasure_eval banditrlproof.finitehorizonrl.markovpolicy.integral_visitindicator_iidepisodebatchmeasure_eval every mapped episode coordinate has the common genuine visit-indicator mean. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_transitionIndicator_iidEpisodeBatchMeasure_eval","label":"integral_transitionIndicator_iidEpisodeBatchMeasure_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_transitionIndicator_iidEpisodeBatchMeasure_eval","description":"Every mapped episode coordinate has the common genuine transition-indicator mean.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-0a92b1081ed2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8208,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:245"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_transitionIndicator_iidEpisodeBatchMeasure_eval {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : integral (policy.iidEpisodeBatchMeasure initialState episodes) (fun batch => EpisodeStep.transitionIndicator state action nextState (batch episode stage)) = policy.stageTransitionJointProbability initialState stage state action nextState","missing":[],"search":"integral_transitionindicator_iidepisodebatchmeasure_eval banditrlproof.finitehorizonrl.markovpolicy.integral_transitionindicator_iidepisodebatchmeasure_eval every mapped episode coordinate has the common genuine transition-indicator mean. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredVisitIndicator","label":"centeredVisitIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredVisitIndicator","description":"Centered visit indicator for one episode coordinate of a mapped batch.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-bdd5527d3554","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8209,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:281"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredVisitIndicator {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) (episode : Fin episodes) (batch : EpisodeBatch mdp episodes) : Real","missing":[],"search":"centeredvisitindicator banditrlproof.finitehorizonrl.markovpolicy.centeredvisitindicator centered visit indicator for one episode coordinate of a mapped batch. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredTransitionIndicator","label":"centeredTransitionIndicator","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredTransitionIndicator","description":"Centered transition indicator for one episode coordinate of a mapped batch.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-e0db0c5d9e26","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8210,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:290"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def centeredTransitionIndicator {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (episode : Fin episodes) (batch : EpisodeBatch mdp episodes) : Real","missing":[],"search":"centeredtransitionindicator banditrlproof.finitehorizonrl.markovpolicy.centeredtransitionindicator centered transition indicator for one episode coordinate of a mapped batch. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_visitIndicator_eq_cast_visitCount","label":"sum_visitIndicator_eq_cast_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_visitIndicator_eq_cast_visitCount","description":"The named visit count is the Real sum of mapped record indicators.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-3ddeea83fcdc","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8211,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:302"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_visitIndicator_eq_cast_visitCount {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : (∑ episode : Fin episodes, EpisodeStep.visitIndicator state action (batch episode stage)) = (batch.visitCount stage state action : Real)","missing":[],"search":"sum_visitindicator_eq_cast_visitcount banditrlproof.finitehorizonrl.markovpolicy.sum_visitindicator_eq_cast_visitcount the named visit count is the real sum of mapped record indicators. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_transitionIndicator_eq_cast_transitionCount","label":"sum_transitionIndicator_eq_cast_transitionCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_transitionIndicator_eq_cast_transitionCount","description":"The named transition count is the Real sum of mapped record indicators.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-3bca2ac096f2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8212,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:314"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_transitionIndicator_eq_cast_transitionCount {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (∑ episode : Fin episodes, EpisodeStep.transitionIndicator state action nextState (batch episode stage)) = (batch.transitionCount stage state action nextState : Real)","missing":[],"search":"sum_transitionindicator_eq_cast_transitioncount banditrlproof.finitehorizonrl.markovpolicy.sum_transitionindicator_eq_cast_transitioncount the named transition count is the real sum of mapped record indicators. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_centeredVisitIndicator_eq_cast_visitCount_sub","label":"sum_centeredVisitIndicator_eq_cast_visitCount_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_centeredVisitIndicator_eq_cast_visitCount_sub","description":"Centered visit-indicator sums are exactly count minus episode-count times mean.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-59b1aef657d7","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8213,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:327"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_centeredVisitIndicator_eq_cast_visitCount_sub {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) : (∑ episode : Fin episodes, policy.centeredVisitIndicator initialState stage state action episode batch) = (batch.visitCount stage state action : Real) - (episodes : Real) * policy.stageVisitProbability initialState stage state action","missing":[],"search":"sum_centeredvisitindicator_eq_cast_visitcount_sub banditrlproof.finitehorizonrl.markovpolicy.sum_centeredvisitindicator_eq_cast_visitcount_sub centered visit-indicator sums are exactly count minus episode-count times mean. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_centeredTransitionIndicator_eq_cast_transitionCount_sub","label":"sum_centeredTransitionIndicator_eq_cast_transitionCount_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_centeredTransitionIndicator_eq_cast_transitionCount_sub","description":"Centered transition-indicator sums are count minus episode-count times mean.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-e1939ce3582a","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8214,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_centeredTransitionIndicator_eq_cast_transitionCount_sub {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (batch : EpisodeBatch mdp episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (∑ episode : Fin episodes, policy.centeredTransitionIndicator initialState stage state action nextState episode batch) = (batch.transitionCount stage state action nextState : Real) - (episodes : Real) * policy.stageTransitionJointProbability initialState stage state action nextState","missing":[],"search":"sum_centeredtransitionindicator_eq_cast_transitioncount_sub banditrlproof.finitehorizonrl.markovpolicy.sum_centeredtransitionindicator_eq_cast_transitioncount_sub centered transition-indicator sums are count minus episode-count times mean. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_cast_visitCount","label":"measurable_cast_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_cast_visitCount","description":"A fixed visit count, coerced to `Real`, is measurable on episode batches.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-bef38bb030b0","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8215,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:376"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cast_visitCount {mdp : MDP State Action} {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable fun batch : EpisodeBatch mdp episodes => (batch.visitCount stage state action : Real)","missing":[],"search":"measurable_cast_visitcount banditrlproof.finitehorizonrl.markovpolicy.measurable_cast_visitcount a fixed visit count, coerced to `real`, is measurable on episode batches. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_cast_transitionCount","label":"measurable_cast_transitionCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_cast_transitionCount","description":"A fixed transition count, coerced to `Real`, is measurable on episode batches.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-b4d7d89f5433","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8216,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:391"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cast_transitionCount {mdp : MDP State Action} {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : Measurable fun batch : EpisodeBatch mdp episodes => (batch.transitionCount stage state action nextState : Real)","missing":[],"search":"measurable_cast_transitioncount banditrlproof.finitehorizonrl.markovpolicy.measurable_cast_transitioncount a fixed transition count, coerced to `real`, is measurable on episode batches. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_visitCountDeviation","label":"measurable_visitCountDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_visitCountDeviation","description":"The fixed visit-count deviation from its genuine mean is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-0714b75b8047","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8217,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:408"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_visitCountDeviation {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable fun batch : EpisodeBatch mdp episodes => (batch.visitCount stage state action : Real) - (episodes : Real) * policy.stageVisitProbability initialState stage state action","missing":[],"search":"measurable_visitcountdeviation banditrlproof.finitehorizonrl.markovpolicy.measurable_visitcountdeviation the fixed visit-count deviation from its genuine mean is measurable. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_transitionCountDeviation","label":"measurable_transitionCountDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_transitionCountDeviation","description":"The fixed joint-transition-count deviation from its genuine mean is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-7ac37d75fe3e","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8218,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:421"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionCountDeviation {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : Measurable fun batch : EpisodeBatch mdp episodes => (batch.transitionCount stage state action nextState : Real) - (episodes : Real) * policy.stageTransitionJointProbability initialState stage state action nextState","missing":[],"search":"measurable_transitioncountdeviation banditrlproof.finitehorizonrl.markovpolicy.measurable_transitioncountdeviation the fixed joint-transition-count deviation from its genuine mean is measurable. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy","label":"iidBernoulliVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy","description":"Total Hoeffding variance proxy for a finite iid family of `[0,1]` indicators.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-fa5f99eee207","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8219,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:433"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidBernoulliVarianceProxy (episodes : Nat) : NNReal","missing":[],"search":"iidbernoullivarianceproxy banditrlproof.finitehorizonrl.markovpolicy.iidbernoullivarianceproxy total hoeffding variance proxy for a finite iid family of `[0,1]` indicators. definition compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy_eq","label":"iidBernoulliVarianceProxy_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy_eq","description":"The iid Bernoulli Hoeffding proxy is exactly one quarter per episode.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-f2eddc9c13de","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8220,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:437"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidBernoulliVarianceProxy_eq (episodes : Nat) : iidBernoulliVarianceProxy episodes = (episodes : NNReal) / 4","missing":[],"search":"iidbernoullivarianceproxy_eq banditrlproof.finitehorizonrl.markovpolicy.iidbernoullivarianceproxy_eq the iid bernoulli hoeffding proxy is exactly one quarter per episode. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy_pos","label":"iidBernoulliVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy_pos","description":"A positive episode count gives a positive total Bernoulli variance proxy.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-975f0b1fde82","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8221,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:443"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidBernoulliVarianceProxy_pos {episodes : Nat} (hepisodes : 0 < episodes) : 0 < ((iidBernoulliVarianceProxy episodes : NNReal) : Real)","missing":[],"search":"iidbernoullivarianceproxy_pos banditrlproof.finitehorizonrl.markovpolicy.iidbernoullivarianceproxy_pos a positive episode count gives a positive total bernoulli variance proxy. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredVisitIndicator_hasSubgaussianMGF","label":"centeredVisitIndicator_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredVisitIndicator_hasSubgaussianMGF","description":"One centered visit coordinate has the `[0,1]` Hoeffding MGF proxy.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-fd7b85131ea1","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8222,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:457"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem centeredVisitIndicator_hasSubgaussianMGF {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) (episode : Fin episodes) : ProbabilityTheory.HasSubgaussianMGF (policy.centeredVisitIndicator initialState stage state action episode) (Concentration.intervalVarianceProxy 0 1) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"centeredvisitindicator_hassubgaussianmgf banditrlproof.finitehorizonrl.markovpolicy.centeredvisitindicator_hassubgaussianmgf one centered visit coordinate has the `[0,1]` hoeffding mgf proxy. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredTransitionIndicator_hasSubgaussianMGF","label":"centeredTransitionIndicator_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredTransitionIndicator_hasSubgaussianMGF","description":"One centered transition coordinate has the `[0,1]` Hoeffding MGF proxy.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-ae2b01d339cf","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8223,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:488"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem centeredTransitionIndicator_hasSubgaussianMGF {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (episode : Fin episodes) : ProbabilityTheory.HasSubgaussianMGF (policy.centeredTransitionIndicator initialState stage state action nextState episode) (Concentration.intervalVarianceProxy 0 1) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"centeredtransitionindicator_hassubgaussianmgf banditrlproof.finitehorizonrl.markovpolicy.centeredtransitionindicator_hassubgaussianmgf one centered transition coordinate has the `[0,1]` hoeffding mgf proxy. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_centeredVisitIndicator","label":"iIndepFun_centeredVisitIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_centeredVisitIndicator","description":"Centered visit coordinates remain independent under the mapped batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-11cfdd80a5f4","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8224,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:522"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_centeredVisitIndicator {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : ProbabilityTheory.iIndepFun (policy.centeredVisitIndicator initialState stage state action) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iindepfun_centeredvisitindicator banditrlproof.finitehorizonrl.markovpolicy.iindepfun_centeredvisitindicator centered visit coordinates remain independent under the mapped batch law. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_centeredTransitionIndicator","label":"iIndepFun_centeredTransitionIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_centeredTransitionIndicator","description":"Centered transition coordinates remain independent under the mapped batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-b3c213800a22","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8225,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:538"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_centeredTransitionIndicator {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : ProbabilityTheory.iIndepFun (policy.centeredTransitionIndicator initialState stage state action nextState) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iindepfun_centeredtransitionindicator banditrlproof.finitehorizonrl.markovpolicy.iindepfun_centeredtransitionindicator centered transition coordinates remain independent under the mapped batch law. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_visitCountBadEvent","label":"measurableSet_visitCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_visitCountBadEvent","description":"The fixed visit-count two-sided bad event is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-648ad5a7de91","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8226,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:555"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_visitCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (delta : Real) : MeasurableSet {batch : EpisodeBatch mdp episodes | Concentration.subGaussianSumConfidenceRadius (iidBernoulliVarianceProxy episodes) delta ≤ |(batch.visitCount stage state action : Real) - (episodes : Real) * policy.stageVisitProbability initialState stage state action|}","missing":[],"search":"measurableset_visitcountbadevent banditrlproof.finitehorizonrl.markovpolicy.measurableset_visitcountbadevent the fixed visit-count two-sided bad event is measurable. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_transitionCountBadEvent","label":"measurableSet_transitionCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_transitionCountBadEvent","description":"The fixed joint-transition-count two-sided bad event is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-8eb5eba551c2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8227,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:572"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_transitionCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (delta : Real) : MeasurableSet {batch : EpisodeBatch mdp episodes | Concentration.subGaussianSumConfidenceRadius (iidBernoulliVarianceProxy episodes) delta ≤ |(batch.transitionCount stage state action nextState : Real) - (episodes : Real) * policy.stageTransitionJointProbability initialState stage state action nextState|}","missing":[],"search":"measurableset_transitioncountbadevent banditrlproof.finitehorizonrl.markovpolicy.measurableset_transitioncountbadevent the fixed joint-transition-count two-sided bad event is measurable. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_visitCount_abs_tail_le","label":"iidEpisodeBatch_visitCount_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_visitCount_abs_tail_le","description":"Two-sided delta confidence tail for one fixed visit-count coordinate under the mapped fixed-policy iid episode-batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-5af28045ac56","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8228,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:594"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_visitCount_abs_tail_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (policy.iidEpisodeBatchMeasure initialState episodes) {batch | Concentration.subGaussianSumConfidenceRadius (iidBernoulliVarianceProxy episodes) delta <= |(batch.visitCount stage state action : Real) - (episodes : Real) * policy.stageVisitProbability initialState stage state action|} <= ENNReal.ofReal delta","missing":[],"search":"iidepisodebatch_visitcount_abs_tail_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_visitcount_abs_tail_le two-sided delta confidence tail for one fixed visit-count coordinate under the mapped fixed-policy iid episode-batch law. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_transitionCount_abs_tail_le","label":"iidEpisodeBatch_transitionCount_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_transitionCount_abs_tail_le","description":"Two-sided delta confidence tail for one fixed transition-count coordinate under the mapped fixed-policy iid episode-batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-f9910d8d3ff7","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8229,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:629"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_transitionCount_abs_tail_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (policy.iidEpisodeBatchMeasure initialState episodes) {batch | Concentration.subGaussianSumConfidenceRadius (iidBernoulliVarianceProxy episodes) delta <= |(batch.transitionCount stage state action nextState : Real) - (episodes : Real) * policy.stageTransitionJointProbability initialState stage state action nextState|} <= ENNReal.ofReal delta","missing":[],"search":"iidepisodebatch_transitioncount_abs_tail_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_transitioncount_abs_tail_le two-sided delta confidence tail for one fixed transition-count coordinate under the mapped fixed-policy iid episode-batch law. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_visit_and_transition_count_abs_tail_le","label":"iidEpisodeBatch_visit_and_transition_count_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_visit_and_transition_count_abs_tail_le","description":"Route endpoint: both named fixed-coordinate count tails are available under the same mapped iid episode-batch law. This conjunction does not union the two bad events or spend a shared failure budget.","url":"../modules/banditrlproof-rl-finitehorizoniidcountconcentration/index.html#decl-3939e2f09361","parent":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","order":8230,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDCountConcentration"],["Source","BanditRLProof/RL/FiniteHorizonIIDCountConcentration.lean:667"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_visit_and_transition_count_abs_tail_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (policy.iidEpisodeBatchMeasure initialState episodes) {batch | Concentration.subGaussianSumConfidenceRadius (iidBernoulliVarianceProxy episodes) delta <= |(batch.visitCount stage state action : Real) - (episodes : Real) * policy.stageVisitProbability initialState stage state action|} <= ENNReal.ofReal delta /\\ (policy.iidEpisodeBatchMeasure initialState episodes) {batch | Concentration.subGaussianSumConfidenceRadius (iidBernoulliVarianceProxy episodes) delta <= |(batch.transitionCount stage state action nextState : Real…","missing":[],"search":"iidepisodebatch_visit_and_transition_count_abs_tail_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_visit_and_transition_count_abs_tail_le route endpoint: both named fixed-coordinate count tails are available under the same mapped iid episode-batch law. this conjunction does not union the two bad events or spend a shared failure budget. theorem compiled","shard":"modules/f6528463e67d2f76.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalTransitionMass_abs_sub_transition_lt_of_not_mem_simultaneousCountBadEvent","label":"empiricalTransitionMass_abs_sub_transition_lt_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalTransitionMass_abs_sub_transition_lt_of_not_mem_simultaneousCountBadEvent","description":"A single eligible coordinate inherits a strict empirical transition-mass bound from the simultaneous visit and joint-transition count deviations.","url":"../modules/banditrlproof-rl-finitehorizoniideligibleempiricaltransitionconfidence/index.html#decl-e1c484b474bd","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","order":8231,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleEmpiricalTransitionConfidence.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalTransitionMass_abs_sub_transition_lt_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (defaultState : State) (coordinate : VisitCoordinate mdp) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) (nextState : State) : |batch.empiricalTransitionMass defaultState coordinate.stage coordinate.state coordinate.action nextState - (mdp.transition (coordinate.state, coordinate.action)).real {nextState}| < 2 * simultaneousCountConfidenceRadius mdp episodes delta / (coordinate.count batch : Real)","missing":[],"search":"empiricaltransitionmass_abs_sub_transition_lt_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.empiricaltransitionmass_abs_sub_transition_lt_of_not_mem_simultaneouscountbadevent a single eligible coordinate inherits a strict empirical transition-mass bound from the simultaneous visit and joint-transition count deviations. theorem compiled","shard":"modules/cc8ccf653922594c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_empiricalTransitionMass_confidence","label":"iidEpisodeBatch_eligible_empiricalTransitionMass_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_empiricalTransitionMass_confidence","description":"Route endpoint: the existing measurable simultaneous-count event has global delta mass, and outside it all eligible empirical transition singleton masses obey the positive-random-denominator confidence bound simultaneously.","url":"../modules/banditrlproof-rl-finitehorizoniideligibleempiricaltransitionconfidence/index.html#decl-45badb4f4a4f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","order":8232,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleEmpiricalTransitionConfidence.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_eligible_empiricalTransitionMass_confidence {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) (defaultState : State) (eligible : Finset (VisitCoordinate mdp)) (hmargin : ∀ coordinate ∈ eligible, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : MeasurableSet (policy.simultaneousCountBadEvent initialState episodes delta) ∧ (policy.iidEpisodeBatchMeasure initialState episodes) (policy.simultaneousCountBadEvent initialState episodes delta) ≤ ENNReal.ofReal delta ∧ ∀ batch ∉ policy.simultaneousCountBadEvent initialState episodes delta, ∀ coordinate ∈ eligible, ∀ nextState, |batch.empiricalTransitionMass defaultState coordinate.stag…","missing":[],"search":"iidepisodebatch_eligible_empiricaltransitionmass_confidence banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_eligible_empiricaltransitionmass_confidence route endpoint: the existing measurable simultaneous-count event has global delta mass, and outside it all eligible empirical transition singleton masses obey the positive-random-denominator confidence bound simultaneously. theorem compiled","shard":"modules/cc8ccf653922594c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate","label":"VisitCoordinate","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.VisitCoordinate","description":"One stage/state/action coordinate whose empirical visit count may be used as a denominator.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-498524c4aab2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8233,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure VisitCoordinate (mdp : MDP State Action) where","missing":[],"search":"visitcoordinate banditrlproof.finitehorizonrl.visitcoordinate one stage/state/action coordinate whose empirical visit count may be used as a denominator. structure compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.count","label":"count","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.VisitCoordinate.count","description":"Realized visit count selected by a visit coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-3d6d5c0f28d4","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8234,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def count {mdp : MDP State Action} (coordinate : VisitCoordinate mdp) {episodes : Nat} (batch : EpisodeBatch mdp episodes) : Nat","missing":[],"search":"count banditrlproof.finitehorizonrl.visitcoordinate.count realized visit count selected by a visit coordinate. definition compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.expectedCount","label":"expectedCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.VisitCoordinate.expectedCount","description":"Genuine expected visit count under the fixed-policy single-episode law.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-586cb9c5cbc9","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8235,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedCount {mdp : MDP State Action} (coordinate : VisitCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : Real","missing":[],"search":"expectedcount banditrlproof.finitehorizonrl.visitcoordinate.expectedcount genuine expected visit count under the fixed-policy single-episode law. definition compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.zeroCountEvent","label":"zeroCountEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.VisitCoordinate.zeroCountEvent","description":"Event that the selected visit coordinate has zero realized count.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-0385a84e6104","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8236,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def zeroCountEvent {mdp : MDP State Action} (coordinate : VisitCoordinate mdp) (episodes : Nat) : Set (EpisodeBatch mdp episodes)","missing":[],"search":"zerocountevent banditrlproof.finitehorizonrl.visitcoordinate.zerocountevent event that the selected visit coordinate has zero realized count. definition compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.measurableSet_zeroCountEvent","label":"measurableSet_zeroCountEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.VisitCoordinate.measurableSet_zeroCountEvent","description":"A selected zero visit-count event is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-9527790e4991","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8237,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_zeroCountEvent {mdp : MDP State Action} (coordinate : VisitCoordinate mdp) (episodes : Nat) : MeasurableSet (coordinate.zeroCountEvent episodes)","missing":[],"search":"measurableset_zerocountevent banditrlproof.finitehorizonrl.visitcoordinate.measurableset_zerocountevent a selected zero visit-count event is measurable. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.eligibleZeroVisitCountEvent","label":"eligibleZeroVisitCountEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.eligibleZeroVisitCountEvent","description":"Union of zero-count events over a finite caller-selected coordinate set.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-e18eb0e43d94","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8238,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def eligibleZeroVisitCountEvent {mdp : MDP State Action} (episodes : Nat) (eligible : Finset (VisitCoordinate mdp)) : Set (EpisodeBatch mdp episodes)","missing":[],"search":"eligiblezerovisitcountevent banditrlproof.finitehorizonrl.eligiblezerovisitcountevent union of zero-count events over a finite caller-selected coordinate set. definition compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.measurableSet_eligibleZeroVisitCountEvent","label":"measurableSet_eligibleZeroVisitCountEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.measurableSet_eligibleZeroVisitCountEvent","description":"The finite eligible zero-count union is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-f9ff1ddd1009","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8239,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_eligibleZeroVisitCountEvent {mdp : MDP State Action} (episodes : Nat) (eligible : Finset (VisitCoordinate mdp)) : MeasurableSet (eligibleZeroVisitCountEvent episodes eligible)","missing":[],"search":"measurableset_eligiblezerovisitcountevent banditrlproof.finitehorizonrl.measurableset_eligiblezerovisitcountevent the finite eligible zero-count union is measurable. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.visitCoordinate_count_pos_of_not_mem_eligibleZeroVisitCountEvent","label":"visitCoordinate_count_pos_of_not_mem_eligibleZeroVisitCountEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.visitCoordinate_count_pos_of_not_mem_eligibleZeroVisitCountEvent","description":"Outside the eligible zero-count union, every selected Nat count is positive.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-dc28bf5513fa","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8240,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem visitCoordinate_count_pos_of_not_mem_eligibleZeroVisitCountEvent {mdp : MDP State Action} {episodes : Nat} (eligible : Finset (VisitCoordinate mdp)) (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ eligibleZeroVisitCountEvent episodes eligible) (coordinate : VisitCoordinate mdp) (hcoordinate : coordinate ∈ eligible) : 0 < coordinate.count batch","missing":[],"search":"visitcoordinate_count_pos_of_not_mem_eligiblezerovisitcountevent banditrlproof.finitehorizonrl.visitcoordinate_count_pos_of_not_mem_eligiblezerovisitcountevent outside the eligible zero-count union, every selected nat count is positive. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.visitCoordinate_count_pos_of_not_mem_simultaneousCountBadEvent","label":"visitCoordinate_count_pos_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.visitCoordinate_count_pos_of_not_mem_simultaneousCountBadEvent","description":"A strict expected-count margin turns the simultaneous deviation bound into a positive realized visit count.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-bffce07ab3eb","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8241,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem visitCoordinate_count_pos_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (coordinate : VisitCoordinate mdp) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : 0 < coordinate.count batch","missing":[],"search":"visitcoordinate_count_pos_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.visitcoordinate_count_pos_of_not_mem_simultaneouscountbadevent a strict expected-count margin turns the simultaneous deviation bound into a positive realized visit count. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.eligibleZeroVisitCountEvent_subset_simultaneousCountBadEvent","label":"eligibleZeroVisitCountEvent_subset_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.eligibleZeroVisitCountEvent_subset_simultaneousCountBadEvent","description":"Under the eligible margins, every eligible zero-count outcome lies in the already-budgeted simultaneous bad event.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-93e88b8c2689","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8242,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:154"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eligibleZeroVisitCountEvent_subset_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (eligible : Finset (VisitCoordinate mdp)) (hmargin : ∀ coordinate ∈ eligible, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : eligibleZeroVisitCountEvent episodes eligible ⊆ policy.simultaneousCountBadEvent initialState episodes delta","missing":[],"search":"eligiblezerovisitcountevent_subset_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.eligiblezerovisitcountevent_subset_simultaneouscountbadevent under the eligible margins, every eligible zero-count outcome lies in the already-budgeted simultaneous bad event. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligibleZeroVisitCountEvent_le","label":"iidEpisodeBatch_eligibleZeroVisitCountEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligibleZeroVisitCountEvent_le","description":"The eligible zero-count union inherits the simultaneous global-delta tail.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-0537eb1ae33b","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8243,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_eligibleZeroVisitCountEvent_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) (eligible : Finset (VisitCoordinate mdp)) (hmargin : ∀ coordinate ∈ eligible, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : (policy.iidEpisodeBatchMeasure initialState episodes) (eligibleZeroVisitCountEvent episodes eligible) ≤ ENNReal.ofReal delta","missing":[],"search":"iidepisodebatch_eligiblezerovisitcountevent_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_eligiblezerovisitcountevent_le the eligible zero-count union inherits the simultaneous global-delta tail. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_visit_count_positivity","label":"iidEpisodeBatch_eligible_visit_count_positivity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_visit_count_positivity","description":"Route endpoint: eligible zero counts have global-delta probability, and every eligible denominator is positive outside that exact zero-count event. The named subset theorem separately embeds this event in the compiled simultaneous bad event.","url":"../modules/banditrlproof-rl-finitehorizoniideligiblevisitcountpositivity/index.html#decl-3940d786f03d","parent":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","order":8244,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity"],["Source","BanditRLProof/RL/FiniteHorizonIIDEligibleVisitCountPositivity.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_eligible_visit_count_positivity {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) (eligible : Finset (VisitCoordinate mdp)) (hmargin : ∀ coordinate ∈ eligible, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : MeasurableSet (eligibleZeroVisitCountEvent episodes eligible) ∧ (policy.iidEpisodeBatchMeasure initialState episodes) (eligibleZeroVisitCountEvent episodes eligible) ≤ ENNReal.ofReal delta ∧ ∀ batch ∉ eligibleZeroVisitCountEvent episodes eligible, ∀ coordinate ∈ eligible, 0 < coordinate.count batch","missing":[],"search":"iidepisodebatch_eligible_visit_count_positivity banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_eligible_visit_count_positivity route endpoint: eligible zero counts have global-delta probability, and every eligible denominator is positive outside that exact zero-count event. the named subset theorem separately embeds this event in the compiled simultaneous bad event. theorem compiled","shard":"modules/d0660da0b0b6fa3b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.RewardConsistent","label":"RewardConsistent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.RewardConsistent","description":"Every recorded reward agrees with the MDP reward at its recorded state and action.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-0f855f21b383","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8245,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:29"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def RewardConsistent {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) : Prop","missing":[],"search":"rewardconsistent banditrlproof.finitehorizonrl.episodebatch.rewardconsistent every recorded reward agrees with the mdp reward at its recorded state and action. definition compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurableSet_rewardConsistent","label":"measurableSet_rewardConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurableSet_rewardConsistent","description":"Reward consistency is a measurable property of finite episode batches.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-2bdcb47270ea","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8246,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_rewardConsistent {mdp : MDP State Action} {episodes : Nat} : MeasurableSet {batch : EpisodeBatch mdp episodes | batch.RewardConsistent}","missing":[],"search":"measurableset_rewardconsistent banditrlproof.finitehorizonrl.episodebatch.measurableset_rewardconsistent reward consistency is a measurable property of finite episode batches. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.rewardSum_eq_visitCount_mul_reward_of_rewardConsistent","label":"rewardSum_eq_visitCount_mul_reward_of_rewardConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.rewardSum_eq_visitCount_mul_reward_of_rewardConsistent","description":"In a reward-consistent batch, the reward sum is visit count times the true reward.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-775474fd3676","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8247,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem rewardSum_eq_visitCount_mul_reward_of_rewardConsistent [DecidableEq State] [DecidableEq Action] {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (hbatch : batch.RewardConsistent) (stage : Fin mdp.horizon) (state : State) (action : Action) : batch.rewardSum stage state action = (batch.visitCount stage state action : Real) * mdp.reward state action","missing":[],"search":"rewardsum_eq_visitcount_mul_reward_of_rewardconsistent banditrlproof.finitehorizonrl.episodebatch.rewardsum_eq_visitcount_mul_reward_of_rewardconsistent in a reward-consistent batch, the reward sum is visit count times the true reward. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalReward_eq_reward_of_rewardConsistent","label":"empiricalReward_eq_reward_of_rewardConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalReward_eq_reward_of_rewardConsistent","description":"Positive visits cancel the denominator, so empirical reward is exact.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-6b4c2ca0c84e","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8248,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalReward_eq_reward_of_rewardConsistent [DecidableEq State] [DecidableEq Action] {mdp : MDP State Action} {episodes : Nat} (batch : EpisodeBatch mdp episodes) (hbatch : batch.RewardConsistent) (stage : Fin mdp.horizon) (state : State) (action : Action) (hcount : batch.visitCount stage state action ≠ 0) : batch.empiricalReward stage state action = mdp.reward state action","missing":[],"search":"empiricalreward_eq_reward_of_rewardconsistent banditrlproof.finitehorizonrl.episodebatch.empiricalreward_eq_reward_of_rewardconsistent positive visits cancel the denominator, so empirical reward is exact. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardConsistent","label":"episodeBatchOfTrajectories_rewardConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardConsistent","description":"Every finite episode batch extracted from genuine trajectories is reward-consistent.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-4a452b06f53c","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8249,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:95"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_rewardConsistent (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) : (mdp.episodeBatchOfTrajectories episodes trajectories).RewardConsistent","missing":[],"search":"episodebatchoftrajectories_rewardconsistent banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_rewardconsistent every finite episode batch extracted from genuine trajectories is reward-consistent. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardSum_eq_visitCount_mul_reward","label":"episodeBatchOfTrajectories_rewardSum_eq_visitCount_mul_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardSum_eq_visitCount_mul_reward","description":"Generated reward sums are exactly visit counts times deterministic MDP rewards.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-3ee004462bd3","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8250,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_rewardSum_eq_visitCount_mul_reward [DecidableEq State] [DecidableEq Action] (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) : (mdp.episodeBatchOfTrajectories episodes trajectories).rewardSum stage state action = ((mdp.episodeBatchOfTrajectories episodes trajectories).visitCount stage state action : Real) * mdp.reward state action","missing":[],"search":"episodebatchoftrajectories_rewardsum_eq_visitcount_mul_reward banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_rewardsum_eq_visitcount_mul_reward generated reward sums are exactly visit counts times deterministic mdp rewards. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_empiricalReward_eq_reward","label":"episodeBatchOfTrajectories_empiricalReward_eq_reward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_empiricalReward_eq_reward","description":"A generated empirical reward is exact whenever its visit count is nonzero.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-1242974f26eb","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8251,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_empiricalReward_eq_reward [DecidableEq State] [DecidableEq Action] (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) (hcount : (mdp.episodeBatchOfTrajectories episodes trajectories).visitCount stage state action ≠ 0) : (mdp.episodeBatchOfTrajectories episodes trajectories).empiricalReward stage state action = mdp.reward state action","missing":[],"search":"episodebatchoftrajectories_empiricalreward_eq_reward banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_empiricalreward_eq_reward a generated empirical reward is exact whenever its visit count is nonzero. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_rewardConsistent_ae","label":"iidEpisodeBatchMeasure_rewardConsistent_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_rewardConsistent_ae","description":"The mapped iid episode-batch law is supported a.e. on reward-consistent records.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-f130585b2f34","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8252,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_rewardConsistent_ae {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ∀ᵐ batch ∂policy.iidEpisodeBatchMeasure initialState episodes, batch.RewardConsistent","missing":[],"search":"iidepisodebatchmeasure_rewardconsistent_ae banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure_rewardconsistent_ae the mapped iid episode-batch law is supported a.e. on reward-consistent records. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalReward_eq_and_transition_lt_of_not_mem_simultaneousCountBadEvent","label":"empiricalReward_eq_and_transition_lt_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalReward_eq_and_transition_lt_of_not_mem_simultaneousCountBadEvent","description":"On the simultaneous good event, reward consistency gives exact empirical reward and the existing eligible margin gives every next-state transition bound.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-7e9aed92e879","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8253,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem empiricalReward_eq_and_transition_lt_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (hreward : batch.RewardConsistent) (defaultState : State) (coordinate : VisitCoordinate mdp) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : batch.empiricalReward coordinate.stage coordinate.state coordinate.action = mdp.reward coordinate.state coordinate.action ∧ ∀ nextState, |batch.empiricalTransitionMass defaultState coordinate.stage coordinate.state coordinate.action nextState - (mdp.transition (coordinate.state, coordinate.action)).real {nextState}| < 2 * simultane…","missing":[],"search":"empiricalreward_eq_and_transition_lt_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.empiricalreward_eq_and_transition_lt_of_not_mem_simultaneouscountbadevent on the simultaneous good event, reward consistency gives exact empirical reward and the existing eligible margin gives every next-state transition bound. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeBatchOfTrajectories_empiricalReward_eq_and_transition_lt","label":"episodeBatchOfTrajectories_empiricalReward_eq_and_transition_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeBatchOfTrajectories_empiricalReward_eq_and_transition_lt","description":"Generated trajectory batches satisfy the reward side of the same good-event endpoint.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-3e85861b19cc","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8254,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:191"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_empiricalReward_eq_and_transition_lt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (hbatch : mdp.episodeBatchOfTrajectories episodes trajectories ∉ policy.simultaneousCountBadEvent initialState episodes delta) (defaultState : State) (coordinate : VisitCoordinate mdp) (hmargin : simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : (mdp.episodeBatchOfTrajectories episodes trajectories).empiricalReward coordinate.stage coordinate.state coordinate.action = mdp.reward coordinate.state coordinate.action ∧ ∀ nextState, |(mdp.episodeBatchOfTrajectories episodes trajectories).empiricalTransitionMass defaultState coordinate.s…","missing":[],"search":"episodebatchoftrajectories_empiricalreward_eq_and_transition_lt banditrlproof.finitehorizonrl.markovpolicy.episodebatchoftrajectories_empiricalreward_eq_and_transition_lt generated trajectory batches satisfy the reward side of the same good-event endpoint. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_empiricalReward_exact_and_transition_confidence","label":"iidEpisodeBatch_eligible_empiricalReward_exact_and_transition_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_empiricalReward_exact_and_transition_confidence","description":"Route endpoint: the existing global-delta event simultaneously controls every eligible transition coordinate, while reward-consistent batches have zero reward error at those same positive-count coordinates.","url":"../modules/banditrlproof-rl-finitehorizoniidgeneratedempiricalrewardexactness/index.html#decl-2b0d9d229adf","parent":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","order":8255,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness"],["Source","BanditRLProof/RL/FiniteHorizonIIDGeneratedEmpiricalRewardExactness.lean:223"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_eligible_empiricalReward_exact_and_transition_confidence {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) (defaultState : State) (eligible : Finset (VisitCoordinate mdp)) (hmargin : ∀ coordinate ∈ eligible, simultaneousCountConfidenceRadius mdp episodes delta < coordinate.expectedCount policy initialState episodes) : MeasurableSet (policy.simultaneousCountBadEvent initialState episodes delta) ∧ (policy.iidEpisodeBatchMeasure initialState episodes) (policy.simultaneousCountBadEvent initialState episodes delta) ≤ ENNReal.ofReal delta ∧ ∀ᵐ batch ∂policy.iidEpisodeBatchMeasure initialState episodes, batch ∉ policy.simultaneousCountBadEvent initialState episodes delta -> ∀ coordinate ∈ eligib…","missing":[],"search":"iidepisodebatch_eligible_empiricalreward_exact_and_transition_confidence banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_eligible_empiricalreward_exact_and_transition_confidence route endpoint: the existing global-delta event simultaneously controls every eligible transition coordinate, while reward-consistent batches have zero reward error at those same positive-count coordinates. theorem compiled","shard":"modules/8fb2f0caabd20af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.multiBatchLocalDelta","label":"multiBatchLocalDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.multiBatchLocalDelta","description":"Equal confidence share assigned to every product-batch coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-f08eadfb5305","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8256,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def multiBatchLocalDelta (rounds : Nat) (delta : Real) : Real","missing":[],"search":"multibatchlocaldelta banditrlproof.finitehorizonrl.multibatchlocaldelta equal confidence share assigned to every product-batch coordinate. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure","label":"iidEpisodeBatchFamilyMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure","description":"Finite product of one fixed-policy iid episode-batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-7f6aedcba957","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8257,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidEpisodeBatchFamilyMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds episodes : Nat) : Measure (Fin rounds -> EpisodeBatch mdp episodes)","missing":[],"search":"iidepisodebatchfamilymeasure banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamilymeasure finite product of one fixed-policy iid episode-batch law. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_map_eval","label":"iidEpisodeBatchFamilyMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_map_eval","description":"Every product coordinate has the compiled single-batch marginal law.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-4c48432aa1d5","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8258,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchFamilyMeasure_map_eval {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds episodes : Nat) (round : Fin rounds) : (policy.iidEpisodeBatchFamilyMeasure initialState rounds episodes).map (Function.eval round) = policy.iidEpisodeBatchMeasure initialState episodes","missing":[],"search":"iidepisodebatchfamilymeasure_map_eval banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamilymeasure_map_eval every product coordinate has the compiled single-batch marginal law. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchRoundBadEvent","label":"multiBatchRoundBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchRoundBadEvent","description":"Pullback of one local simultaneous-count event to a product coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-5ae1ea57c112","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8259,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def multiBatchRoundBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds episodes : Nat) (delta : Real) (round : Fin rounds) : Set (Fin rounds -> EpisodeBatch mdp episodes)","missing":[],"search":"multibatchroundbadevent banditrlproof.finitehorizonrl.markovpolicy.multibatchroundbadevent pullback of one local simultaneous-count event to a product coordinate. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchSimultaneousCountBadEvent","label":"multiBatchSimultaneousCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchSimultaneousCountBadEvent","description":"Union of the local simultaneous-count events over all product batches.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-1e131f5bc85d","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8260,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:75"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def multiBatchSimultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds episodes : Nat) (delta : Real) : Set (Fin rounds -> EpisodeBatch mdp episodes)","missing":[],"search":"multibatchsimultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.multibatchsimultaneouscountbadevent union of the local simultaneous-count events over all product batches. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_multiBatchSimultaneousCountBadEvent","label":"measurableSet_multiBatchSimultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_multiBatchSimultaneousCountBadEvent","description":"The pulled-back finite union is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-b30e27901710","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8261,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_multiBatchSimultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds episodes : Nat) (delta : Real) : MeasurableSet (policy.multiBatchSimultaneousCountBadEvent initialState rounds episodes delta)","missing":[],"search":"measurableset_multibatchsimultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.measurableset_multibatchsimultaneouscountbadevent the pulled-back finite union is measurable. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_roundBadEvent_le","label":"iidEpisodeBatchFamilyMeasure_roundBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_roundBadEvent_le","description":"Each pulled-back local event has its equal-share probability bound.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-4c96e36aa8f6","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8262,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchFamilyMeasure_roundBadEvent_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds : Nat) (hrounds : 0 < rounds) (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (round : Fin rounds) : (policy.iidEpisodeBatchFamilyMeasure initialState rounds episodes) (policy.multiBatchRoundBadEvent initialState rounds episodes delta round) <= ENNReal.ofReal (multiBatchLocalDelta rounds delta)","missing":[],"search":"iidepisodebatchfamilymeasure_roundbadevent_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamilymeasure_roundbadevent_le each pulled-back local event has its equal-share probability bound. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_multiBatchSimultaneousCountBadEvent_le","label":"iidEpisodeBatchFamilyMeasure_multiBatchSimultaneousCountBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_multiBatchSimultaneousCountBadEvent_le","description":"Equal-share union over product coordinates retains the global delta.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-8613e27f7a26","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8263,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:133"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchFamilyMeasure_multiBatchSimultaneousCountBadEvent_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds : Nat) (hrounds : 0 < rounds) (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (policy.iidEpisodeBatchFamilyMeasure initialState rounds episodes) (policy.multiBatchSimultaneousCountBadEvent initialState rounds episodes delta) <= ENNReal.ofReal delta","missing":[],"search":"iidepisodebatchfamilymeasure_multibatchsimultaneouscountbadevent_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamilymeasure_multibatchsimultaneouscountbadevent_le equal-share union over product coordinates retains the global delta. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_rewardConsistent_ae","label":"iidEpisodeBatchFamilyMeasure_rewardConsistent_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_rewardConsistent_ae","description":"Every product batch is reward-consistent almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-2ee6ed222990","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8264,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchFamilyMeasure_rewardConsistent_ae {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds episodes : Nat) : ∀ᵐ batches ∂policy.iidEpisodeBatchFamilyMeasure initialState rounds episodes, ∀ round, (batches round).RewardConsistent","missing":[],"search":"iidepisodebatchfamilymeasure_rewardconsistent_ae banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamilymeasure_rewardconsistent_ae every product batch is reward-consistent almost everywhere. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchEmpiricalModelAt","label":"multiBatchEmpiricalModelAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchEmpiricalModelAt","description":"Batch-specific canonical empirical model at one product coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-98655d77c41c","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8265,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def multiBatchEmpiricalModelAt {mdp : MDP State Action} (_policy : MarkovPolicy mdp) {rounds episodes : Nat} (batches : Fin rounds -> EpisodeBatch mdp episodes) (defaultState : State) (transitionBudget : Real) (round : Fin rounds) : MDP.FiniteBatchModel mdp episodes","missing":[],"search":"multibatchempiricalmodelat banditrlproof.finitehorizonrl.markovpolicy.multibatchempiricalmodelat batch-specific canonical empirical model at one product coordinate. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchCumulativeExpectedRegret","label":"multiBatchCumulativeExpectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchCumulativeExpectedRegret","description":"Sum of batch-specific optimistic-policy expected regrets.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-62b657e745e9","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8266,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def multiBatchCumulativeExpectedRegret {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) {rounds episodes : Nat} (batches : Fin rounds -> EpisodeBatch mdp episodes) (defaultState : State) (transitionBudget : Real) : Real","missing":[],"search":"multibatchcumulativeexpectedregret banditrlproof.finitehorizonrl.markovpolicy.multibatchcumulativeexpectedregret sum of batch-specific optimistic-policy expected regrets. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchCumulativeSelectedRadiusOccupancy","label":"multiBatchCumulativeSelectedRadiusOccupancy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchCumulativeSelectedRadiusOccupancy","description":"Sum of the batch-specific selected-radius occupancy bounds.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-d835cb27a4cd","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8267,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:192"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def multiBatchCumulativeSelectedRadiusOccupancy {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) {rounds episodes : Nat} (batches : Fin rounds -> EpisodeBatch mdp episodes) (defaultState : State) (transitionBudget : Real) : Real","missing":[],"search":"multibatchcumulativeselectedradiusoccupancy banditrlproof.finitehorizonrl.markovpolicy.multibatchcumulativeselectedradiusoccupancy sum of the batch-specific selected-radius occupancy bounds. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.allCoordinateConfidenceFamily_of_not_mem","label":"allCoordinateConfidenceFamily_of_not_mem","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.allCoordinateConfidenceFamily_of_not_mem","description":"Pathwise confidence-family producer outside the finite pulled-back bad-event union, assuming the generated reward-consistency support at every coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-fcc941ffe94f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8268,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def allCoordinateConfidenceFamily_of_not_mem {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {rounds episodes : Nat} {delta : Real} (batches : Fin rounds -> EpisodeBatch mdp episodes) (hbatches : batches ∉ policy.multiBatchSimultaneousCountBadEvent initialState rounds episodes delta) (hreward : ∀ round, (batches round).RewardConsistent) (defaultState : State) (rewardBound transitionBudget : Real) (hrewardBound : ∀ state action, |mdp.reward state action| <= rewardBound) (htransitionBudget_nonneg : 0 <= transitionBudget) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < coordinate.expectedCount policy initialState episodes) (hcover : ∀ (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : A…","missing":[],"search":"allcoordinateconfidencefamily_of_not_mem banditrlproof.finitehorizonrl.markovpolicy.allcoordinateconfidencefamily_of_not_mem pathwise confidence-family producer outside the finite pulled-back bad-event union, assuming the generated reward-consistency support at every coordinate. definition compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.confidenceFamily_optimism_and_cumulativeExpectedRegret","label":"confidenceFamily_optimism_and_cumulativeExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.confidenceFamily_optimism_and_cumulativeExpectedRegret","description":"Finite sums preserve all roundwise optimism and expected-regret bounds.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-d3a1211b7028","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8269,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem confidenceFamily_optimism_and_cumulativeExpectedRegret {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {rounds episodes : Nat} (batches : Fin rounds -> EpisodeBatch mdp episodes) (defaultState : State) (transitionBudget : Real) (confidence : ∀ round, (policy.multiBatchEmpiricalModelAt batches defaultState transitionBudget round).Confidence) : (∀ round state, mdp.optimalValueRemaining mdp.horizon le_rfl state <= (policy.multiBatchEmpiricalModelAt batches defaultState transitionBudget round).plan.upperValueRemaining mdp.horizon le_rfl state) ∧ policy.multiBatchCumulativeExpectedRegret initialState batches defaultState transitionBudget <= policy.multiBatchCumulativeSelectedRadiusOccupancy initialState batches defaultState transitionBudget","missing":[],"search":"confidencefamily_optimism_and_cumulativeexpectedregret banditrlproof.finitehorizonrl.markovpolicy.confidencefamily_optimism_and_cumulativeexpectedregret finite sums preserve all roundwise optimism and expected-regret bounds. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamily_allCoordinate_finiteBatchModel_confidence","label":"iidEpisodeBatchFamily_allCoordinate_finiteBatchModel_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamily_allCoordinate_finiteBatchModel_confidence","description":"Mapped finite-product confidence endpoint: one global-delta event and an a.e. family of confidence witnesses, with no measurable witness selection claim.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-efc6b70c4f39","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8270,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:286"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchFamily_allCoordinate_finiteBatchModel_confidence {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds : Nat) (hrounds : 0 < rounds) (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (defaultState : State) (rewardBound transitionBudget : Real) (hrewardBound : ∀ state action, |mdp.reward state action| <= rewardBound) (htransitionBudget_nonneg : 0 <= transitionBudget) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < coordinate.expectedCount policy initialState episodes) (hcover : ∀ (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action), (∑ nextState, policy.expectedCountTransitionCoordinateRadius initialState epis…","missing":[],"search":"iidepisodebatchfamily_allcoordinate_finitebatchmodel_confidence banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamily_allcoordinate_finitebatchmodel_confidence mapped finite-product confidence endpoint: one global-delta event and an a.e. family of confidence witnesses, with no measurable witness selection claim. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamily_allCoordinate_optimism_and_cumulativeExpectedRegret","label":"iidEpisodeBatchFamily_allCoordinate_optimism_and_cumulativeExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamily_allCoordinate_optimism_and_cumulativeExpectedRegret","description":"Route endpoint: the same finite-product event simultaneously yields optimism for every batch model and the cumulative finite sum of expected-regret bounds.","url":"../modules/banditrlproof-rl-finitehorizoniidmultibatchcumulativeconfidenceregret/index.html#decl-dff57e46b78e","parent":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","order":8271,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret"],["Source","BanditRLProof/RL/FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret.lean:339"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchFamily_allCoordinate_optimism_and_cumulativeExpectedRegret {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rounds : Nat) (hrounds : 0 < rounds) (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) (defaultState : State) (rewardBound transitionBudget : Real) (hrewardBound : ∀ state action, |mdp.reward state action| <= rewardBound) (htransitionBudget_nonneg : 0 <= transitionBudget) (hmargin : ∀ coordinate : VisitCoordinate mdp, simultaneousCountConfidenceRadius mdp episodes (multiBatchLocalDelta rounds delta) < coordinate.expectedCount policy initialState episodes) (hcover : ∀ (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (action : Action), (∑ nextState, policy.expectedCountTransitionCoordinateRadius initial…","missing":[],"search":"iidepisodebatchfamily_allcoordinate_optimism_and_cumulativeexpectedregret banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchfamily_allcoordinate_optimism_and_cumulativeexpectedregret route endpoint: the same finite-product event simultaneously yields optimism for every batch model and the cumulative finite sum of expected-regret bounds. theorem compiled","shard":"modules/7b98124166f88ccc.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate","label":"CountCoordinate","kind":"inductive type","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate","description":"Finite index of all visit and joint-transition count coordinates of an MDP.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-4d8a7535c8a8","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8272,"meta":[["Kind","inductive type"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"inductive CountCoordinate (mdp : MDP State Action) where","missing":[],"search":"countcoordinate banditrlproof.finitehorizonrl.countcoordinate finite index of all visit and joint-transition count coordinates of an mdp. inductive type compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.equivVisitSumTransition","label":"equivVisitSumTransition","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.equivVisitSumTransition","description":"Explicit finite-sum presentation of the two coordinate families.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-190be2077736","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8273,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def equivVisitSumTransition (mdp : MDP State Action) : CountCoordinate mdp ≃ (Fin mdp.horizon × State × Action) ⊕ (Fin mdp.horizon × State × Action × State) where","missing":[],"search":"equivvisitsumtransition banditrlproof.finitehorizonrl.countcoordinate.equivvisitsumtransition explicit finite-sum presentation of the two coordinate families. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.countCoordinateCard","label":"countCoordinateCard","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.countCoordinateCard","description":"Number of visit and joint-transition coordinates in the simultaneous family.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-d55ba88db85f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8274,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:63"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def countCoordinateCard (mdp : MDP State Action) : Nat","missing":[],"search":"countcoordinatecard banditrlproof.finitehorizonrl.countcoordinatecard number of visit and joint-transition coordinates in the simultaneous family. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.countCoordinateCard_eq","label":"countCoordinateCard_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.countCoordinateCard_eq","description":"The simultaneous family contains `H*S*A` visits and `H*S*A*S` transitions.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-1300a717915a","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8275,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:70"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countCoordinateCard_eq (mdp : MDP State Action) : countCoordinateCard mdp = mdp.horizon * Fintype.card State * Fintype.card Action + mdp.horizon * Fintype.card State * Fintype.card Action * Fintype.card State","missing":[],"search":"countcoordinatecard_eq banditrlproof.finitehorizonrl.countcoordinatecard_eq the simultaneous family contains `h*s*a` visits and `h*s*a*s` transitions. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountDelta","label":"simultaneousCountDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.simultaneousCountDelta","description":"Equal confidence share assigned to each count coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-7c90be3dd802","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8276,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousCountDelta (mdp : MDP State Action) (delta : Real) : Real","missing":[],"search":"simultaneouscountdelta banditrlproof.finitehorizonrl.simultaneouscountdelta equal confidence share assigned to each count coordinate. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius","label":"simultaneousCountConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius","description":"Common count radius after allocating the global confidence budget equally.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-af78e4f1b242","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8277,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousCountConfidenceRadius (mdp : MDP State Action) (episodes : Nat) (delta : Real) : Real","missing":[],"search":"simultaneouscountconfidenceradius banditrlproof.finitehorizonrl.simultaneouscountconfidenceradius common count radius after allocating the global confidence budget equally. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation","label":"deviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation","description":"Real count deviation selected by a visit or joint-transition coordinate.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-98498d1c0f2d","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8278,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:94"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def deviation {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (batch : EpisodeBatch mdp episodes) : Real","missing":[],"search":"deviation banditrlproof.finitehorizonrl.countcoordinate.deviation real count deviation selected by a visit or joint-transition coordinate. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.measurable_deviation","label":"measurable_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.measurable_deviation","description":"Every selected count deviation is measurable on the mapped batch space.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-deb448cfcd3f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8279,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_deviation {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} : Measurable (coordinate.deviation policy initialState : EpisodeBatch mdp episodes → Real)","missing":[],"search":"measurable_deviation banditrlproof.finitehorizonrl.countcoordinate.measurable_deviation every selected count deviation is measurable on the mapped batch space. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.badEvent","label":"badEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.badEvent","description":"Two-sided bad event for one selected coordinate at a supplied delta.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-d75533ae1eb3","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8280,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def badEvent {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (coordinateDelta : Real) : Set (EpisodeBatch mdp episodes)","missing":[],"search":"badevent banditrlproof.finitehorizonrl.countcoordinate.badevent two-sided bad event for one selected coordinate at a supplied delta. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.measurableSet_badEvent","label":"measurableSet_badEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.CountCoordinate.measurableSet_badEvent","description":"Every selected fixed-coordinate bad event is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-afa8339598f5","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8281,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:138"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_badEvent {mdp : MDP State Action} (coordinate : CountCoordinate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (coordinateDelta : Real) : MeasurableSet (coordinate.badEvent policy initialState episodes coordinateDelta)","missing":[],"search":"measurableset_badevent banditrlproof.finitehorizonrl.countcoordinate.measurableset_badevent every selected fixed-coordinate bad event is measurable. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measure_countCoordinate_badEvent_le","label":"measure_countCoordinate_badEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measure_countCoordinate_badEvent_le","description":"The compiled marginal tail dispatches over the finite coordinate type.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-1365e436309e","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8282,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:153"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measure_countCoordinate_badEvent_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (coordinate : CountCoordinate mdp) (coordinateDelta : Real) (hdelta : 0 < coordinateDelta) (hdelta_le_one : coordinateDelta ≤ 1) : (policy.iidEpisodeBatchMeasure initialState episodes) (coordinate.badEvent policy initialState episodes coordinateDelta) ≤ ENNReal.ofReal coordinateDelta","missing":[],"search":"measure_countcoordinate_badevent_le banditrlproof.finitehorizonrl.markovpolicy.measure_countcoordinate_badevent_le the compiled marginal tail dispatches over the finite coordinate type. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountBadEvent","label":"simultaneousCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountBadEvent","description":"Union of every visit and joint-transition count bad event.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-e4116fe5f9e3","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8283,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:176"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta : Real) : Set (EpisodeBatch mdp episodes)","missing":[],"search":"simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.simultaneouscountbadevent union of every visit and joint-transition count bad event. definition compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_simultaneousCountBadEvent","label":"measurableSet_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_simultaneousCountBadEvent","description":"The simultaneous count bad event is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-b005ef76fad8","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8284,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:186"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta : Real) : MeasurableSet (policy.simultaneousCountBadEvent initialState episodes delta)","missing":[],"search":"measurableset_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.measurableset_simultaneouscountbadevent the simultaneous count bad event is measurable. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountDelta_pos","label":"simultaneousCountDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountDelta_pos","description":"A nonempty coordinate family receives a positive equal delta share.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-d1529a0c90fc","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8285,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:200"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousCountDelta_pos {mdp : MDP State Action} (hcoordinate : Nonempty (CountCoordinate mdp)) {delta : Real} (hdelta : 0 < delta) : 0 < simultaneousCountDelta mdp delta","missing":[],"search":"simultaneouscountdelta_pos banditrlproof.finitehorizonrl.markovpolicy.simultaneouscountdelta_pos a nonempty coordinate family receives a positive equal delta share. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountDelta_le_one","label":"simultaneousCountDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountDelta_le_one","description":"A global delta at most one gives every nonempty-family share at most one.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-d176e472da21","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8286,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:212"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousCountDelta_le_one {mdp : MDP State Action} (hcoordinate : Nonempty (CountCoordinate mdp)) {delta : Real} (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) : simultaneousCountDelta mdp delta ≤ 1","missing":[],"search":"simultaneouscountdelta_le_one banditrlproof.finitehorizonrl.markovpolicy.simultaneouscountdelta_le_one a global delta at most one gives every nonempty-family share at most one. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_simultaneousCountBadEvent_le","label":"iidEpisodeBatch_simultaneousCountBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_simultaneousCountBadEvent_le","description":"All finite visit and joint-transition count deviations share one global delta budget. When the coordinate family is empty (in particular at horizon zero), the bad union is empty and no positive-horizon premise is needed.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-5936421e7416","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8287,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:228"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_simultaneousCountBadEvent_le {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) : (policy.iidEpisodeBatchMeasure initialState episodes) (policy.simultaneousCountBadEvent initialState episodes delta) ≤ ENNReal.ofReal delta","missing":[],"search":"iidepisodebatch_simultaneouscountbadevent_le banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_simultaneouscountbadevent_le all finite visit and joint-transition count deviations share one global delta budget. when the coordinate family is empty (in particular at horizon zero), the bad union is empty and no positive-horizon premise is needed. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.countCoordinate_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","label":"countCoordinate_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.countCoordinate_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","description":"Outside the simultaneous union, every indexed deviation is below its radius.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-ae2bb08a3117","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8288,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem countCoordinate_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (coordinate : CountCoordinate mdp) : |coordinate.deviation policy initialState batch| < simultaneousCountConfidenceRadius mdp episodes delta","missing":[],"search":"countcoordinate_abs_deviation_lt_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.countcoordinate_abs_deviation_lt_of_not_mem_simultaneouscountbadevent outside the simultaneous union, every indexed deviation is below its radius. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.visitCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","label":"visitCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.visitCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","description":"Visit-count specialization of the simultaneous good-side bound.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-50c492e27260","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8289,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:290"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem visitCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (stage : Fin mdp.horizon) (state : State) (action : Action) : |(batch.visitCount stage state action : Real) - (episodes : Real) * policy.stageVisitProbability initialState stage state action| < simultaneousCountConfidenceRadius mdp episodes delta","missing":[],"search":"visitcount_abs_deviation_lt_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.visitcount_abs_deviation_lt_of_not_mem_simultaneouscountbadevent visit-count specialization of the simultaneous good-side bound. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.transitionCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","label":"transitionCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.transitionCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","description":"Joint-transition-count specialization of the simultaneous good-side bound.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-0e3b8943993b","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8290,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:307"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {delta : Real} (batch : EpisodeBatch mdp episodes) (hbatch : batch ∉ policy.simultaneousCountBadEvent initialState episodes delta) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : |(batch.transitionCount stage state action nextState : Real) - (episodes : Real) * policy.stageTransitionJointProbability initialState stage state action nextState| < simultaneousCountConfidenceRadius mdp episodes delta","missing":[],"search":"transitioncount_abs_deviation_lt_of_not_mem_simultaneouscountbadevent banditrlproof.finitehorizonrl.markovpolicy.transitioncount_abs_deviation_lt_of_not_mem_simultaneouscountbadevent joint-transition-count specialization of the simultaneous good-side bound. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_simultaneous_count_confidence","label":"iidEpisodeBatch_simultaneous_count_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_simultaneous_count_confidence","description":"Route endpoint: one global-delta bad union and all coordinatewise good-side bounds under the same mapped fixed-policy iid episode-batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidsimultaneouscountconfidence/index.html#decl-2d9725779bd2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","order":8291,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence"],["Source","BanditRLProof/RL/FiniteHorizonIIDSimultaneousCountConfidence.lean:328"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_simultaneous_count_confidence {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) : (policy.iidEpisodeBatchMeasure initialState episodes) (policy.simultaneousCountBadEvent initialState episodes delta) ≤ ENNReal.ofReal delta ∧ ∀ batch ∉ policy.simultaneousCountBadEvent initialState episodes delta, ∀ coordinate : CountCoordinate mdp, |coordinate.deviation policy initialState batch| < simultaneousCountConfidenceRadius mdp episodes delta","missing":[],"search":"iidepisodebatch_simultaneous_count_confidence banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_simultaneous_count_confidence route endpoint: one global-delta bad union and all coordinatewise good-side bounds under the same mapped fixed-policy iid episode-batch law. theorem compiled","shard":"modules/0115ce4b57aba6bb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryStateAt","label":"trajectoryStateAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryStateAt","description":"Current state immediately before a recorded trajectory action.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-30ee24a710b2","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8292,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def trajectoryStateAt (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : State","missing":[],"search":"trajectorystateat banditrlproof.finitehorizonrl.mdp.trajectorystateat current state immediately before a recorded trajectory action. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeStepOfTrajectory","label":"episodeStepOfTrajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeStepOfTrajectory","description":"A full generated trajectory viewed as one empirical record at a stage.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-5da81f3505a0","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8293,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def episodeStepOfTrajectory (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : EpisodeStep State Action where","missing":[],"search":"episodestepoftrajectory banditrlproof.finitehorizonrl.mdp.episodestepoftrajectory a full generated trajectory viewed as one empirical record at a stage. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories","label":"episodeBatchOfTrajectories","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories","description":"A finite family of full trajectories mapped to empirical episode records.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-2c4382396408","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8294,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def episodeBatchOfTrajectories (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) : EpisodeBatch mdp episodes","missing":[],"search":"episodebatchoftrajectories banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories a finite family of full trajectories mapped to empirical episode records. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryVisitContribution","label":"trajectoryVisitContribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryVisitContribution","description":"One trajectory's contribution to a stage/state/action visit count.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-c098a474add1","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8295,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def trajectoryVisitContribution (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) (trajectory : State × StepTrace Action State mdp.horizon) : Nat","missing":[],"search":"trajectoryvisitcontribution banditrlproof.finitehorizonrl.mdp.trajectoryvisitcontribution one trajectory's contribution to a stage/state/action visit count. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryRewardContribution","label":"trajectoryRewardContribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryRewardContribution","description":"One trajectory's contribution to a stage/state/action reward sum.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-b05a6d19534f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8296,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def trajectoryRewardContribution (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) (trajectory : State × StepTrace Action State mdp.horizon) : Real","missing":[],"search":"trajectoryrewardcontribution banditrlproof.finitehorizonrl.mdp.trajectoryrewardcontribution one trajectory's contribution to a stage/state/action reward sum. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryTransitionContribution","label":"trajectoryTransitionContribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.trajectoryTransitionContribution","description":"One trajectory's contribution to a stage/state/action/next-state count.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-6ad88a44089b","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8297,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def trajectoryTransitionContribution (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) (trajectory : State × StepTrace Action State mdp.horizon) : Nat","missing":[],"search":"trajectorytransitioncontribution banditrlproof.finitehorizonrl.mdp.trajectorytransitioncontribution one trajectory's contribution to a stage/state/action/next-state count. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_episodeStepOfTrajectory","label":"measurable_episodeStepOfTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_episodeStepOfTrajectory","description":"The stage record extracted from a finite trajectory is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-3cf57252c158","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8298,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:84"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_episodeStepOfTrajectory (mdp : MDP State Action) (stage : Fin mdp.horizon) : Measurable (fun trajectory => mdp.episodeStepOfTrajectory trajectory stage)","missing":[],"search":"measurable_episodestepoftrajectory banditrlproof.finitehorizonrl.mdp.measurable_episodestepoftrajectory the stage record extracted from a finite trajectory is measurable. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_episodeBatchOfTrajectories","label":"measurable_episodeBatchOfTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_episodeBatchOfTrajectories","description":"Mapping a finite trajectory family to its empirical batch is measurable.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-f0030b075efb","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8299,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_episodeBatchOfTrajectories (mdp : MDP State Action) (episodes : Nat) : Measurable (mdp.episodeBatchOfTrajectories episodes)","missing":[],"search":"measurable_episodebatchoftrajectories banditrlproof.finitehorizonrl.mdp.measurable_episodebatchoftrajectories mapping a finite trajectory family to its empirical batch is measurable. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_apply","label":"episodeBatchOfTrajectories_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_apply","description":"theorem episodeBatchOfTrajectories_apply (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (episode : Fin episodes) (stage : Fin mdp.horizon) : mdp.episodeBatchOfTrajectories episodes trajectories episode stage = mdp.episodeStepOfTrajectory (trajectories episode) stage","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-253e4f72faae","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8300,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_apply (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (episode : Fin episodes) (stage : Fin mdp.horizon) : mdp.episodeBatchOfTrajectories episodes trajectories episode stage = mdp.episodeStepOfTrajectory (trajectories episode) stage","missing":[],"search":"episodebatchoftrajectories_apply banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_apply theorem episodebatchoftrajectories_apply (mdp : mdp state action) (episodes : nat) (trajectories : fin episodes -> state × steptrace action state mdp.horizon) (episode : fin episodes) (stage : fin mdp.horizon) : mdp.episodebatchoftrajectories episodes trajectories episode stage = mdp.episodestepoftrajectory (trajectories episode) stage theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_visitCount","label":"episodeBatchOfTrajectories_visitCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_visitCount","description":"Extracted batch visits are exactly the sum of trajectory contributions.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-dd313aa4282f","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8301,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_visitCount (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) : (mdp.episodeBatchOfTrajectories episodes trajectories).visitCount stage state action = ∑ episode, mdp.trajectoryVisitContribution stage state action (trajectories episode)","missing":[],"search":"episodebatchoftrajectories_visitcount banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_visitcount extracted batch visits are exactly the sum of trajectory contributions. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardSum","label":"episodeBatchOfTrajectories_rewardSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardSum","description":"Extracted batch rewards are exactly the sum of trajectory contributions.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-6fd897ccd0a3","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8302,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:128"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_rewardSum (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) : (mdp.episodeBatchOfTrajectories episodes trajectories).rewardSum stage state action = ∑ episode, mdp.trajectoryRewardContribution stage state action (trajectories episode)","missing":[],"search":"episodebatchoftrajectories_rewardsum banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_rewardsum extracted batch rewards are exactly the sum of trajectory contributions. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_transitionCount","label":"episodeBatchOfTrajectories_transitionCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_transitionCount","description":"Extracted transition counts are exactly trajectory-indicator sums.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-0485620f8931","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8303,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem episodeBatchOfTrajectories_transitionCount (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (mdp.episodeBatchOfTrajectories episodes trajectories).transitionCount stage state action nextState = ∑ episode, mdp.trajectoryTransitionContribution stage state action nextState (trajectories episode)","missing":[],"search":"episodebatchoftrajectories_transitioncount banditrlproof.finitehorizonrl.mdp.episodebatchoftrajectories_transitioncount extracted transition counts are exactly trajectory-indicator sums. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryVisitContribution","label":"measurable_trajectoryVisitContribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryVisitContribution","description":"A fixed visit contribution is measurable on the generated trajectory.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-9937fcba5509","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8304,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_trajectoryVisitContribution (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable (mdp.trajectoryVisitContribution stage state action)","missing":[],"search":"measurable_trajectoryvisitcontribution banditrlproof.finitehorizonrl.mdp.measurable_trajectoryvisitcontribution a fixed visit contribution is measurable on the generated trajectory. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryRewardContribution","label":"measurable_trajectoryRewardContribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryRewardContribution","description":"A fixed reward contribution is measurable on the generated trajectory.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-df801da002b6","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8305,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_trajectoryRewardContribution (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) : Measurable (mdp.trajectoryRewardContribution stage state action)","missing":[],"search":"measurable_trajectoryrewardcontribution banditrlproof.finitehorizonrl.mdp.measurable_trajectoryrewardcontribution a fixed reward contribution is measurable on the generated trajectory. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryTransitionContribution","label":"measurable_trajectoryTransitionContribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryTransitionContribution","description":"A fixed transition contribution is measurable on the generated trajectory.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-7b9accb5e1f1","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8306,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:172"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_trajectoryTransitionContribution (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : Measurable (mdp.trajectoryTransitionContribution stage state action nextState)","missing":[],"search":"measurable_trajectorytransitioncontribution banditrlproof.finitehorizonrl.mdp.measurable_trajectorytransitioncontribution a fixed transition contribution is measurable on the generated trajectory. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidTrajectoryFamilyMeasure","label":"iidTrajectoryFamilyMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidTrajectoryFamilyMeasure","description":"Finite iid product of the generated single-episode trajectory law.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-a7b383dc0448","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8307,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:184"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidTrajectoryFamilyMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : Measure (Fin episodes -> State × StepTrace Action State mdp.horizon)","missing":[],"search":"iidtrajectoryfamilymeasure banditrlproof.finitehorizonrl.markovpolicy.iidtrajectoryfamilymeasure finite iid product of the generated single-episode trajectory law. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure","label":"iidEpisodeBatchMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure","description":"Pushforward law of the empirical batch extracted from iid trajectories.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-9514155aaafe","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8308,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:201"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidEpisodeBatchMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : Measure (EpisodeBatch mdp episodes)","missing":[],"search":"iidepisodebatchmeasure banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure pushforward law of the empirical batch extracted from iid trajectories. definition compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidTrajectoryFamilyMeasure_map_eval","label":"iidTrajectoryFamilyMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidTrajectoryFamilyMeasure_map_eval","description":"Every product coordinate has the generated single-episode trajectory law.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-e4d8849a0b38","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8309,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:221"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidTrajectoryFamilyMeasure_map_eval {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) : (policy.iidTrajectoryFamilyMeasure initialState episodes).map (Function.eval episode) = policy.trajectoryMeasure initialState","missing":[],"search":"iidtrajectoryfamilymeasure_map_eval banditrlproof.finitehorizonrl.markovpolicy.iidtrajectoryfamilymeasure_map_eval every product coordinate has the generated single-episode trajectory law. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_map_eval","label":"iidEpisodeBatchMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_map_eval","description":"Each episode/stage coordinate of the mapped batch has the corresponding pushforward of the genuine generated single-trajectory law.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-7b2707fc3927","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8310,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:239"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatchMeasure_map_eval {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) (stage : Fin mdp.horizon) : (policy.iidEpisodeBatchMeasure initialState episodes).map (fun batch => batch episode stage) = (policy.trajectoryMeasure initialState).map (fun trajectory => mdp.episodeStepOfTrajectory trajectory stage)","missing":[],"search":"iidepisodebatchmeasure_map_eval banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatchmeasure_map_eval each episode/stage coordinate of the mapped batch has the corresponding pushforward of the genuine generated single-trajectory law. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeStepOfTrajectory","label":"iIndepFun_episodeStepOfTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeStepOfTrajectory","description":"Stage records from distinct product coordinates are independent.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-0d240687d7a1","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8311,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:273"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_episodeStepOfTrajectory {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.episodeStepOfTrajectory (trajectories episode) stage) (policy.iidTrajectoryFamilyMeasure initialState episodes)","missing":[],"search":"iindepfun_episodestepoftrajectory banditrlproof.finitehorizonrl.markovpolicy.iindepfun_episodestepoftrajectory stage records from distinct product coordinates are independent. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_eval","label":"iIndepFun_iidEpisodeBatch_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_eval","description":"Fixed-stage record coordinates are independent under the mapped batch law.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-bd210332f027","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8312,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_iidEpisodeBatch_eval {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) : ProbabilityTheory.iIndepFun (fun episode (batch : EpisodeBatch mdp episodes) => batch episode stage) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iindepfun_iidepisodebatch_eval banditrlproof.finitehorizonrl.markovpolicy.iindepfun_iidepisodebatch_eval fixed-stage record coordinates are independent under the mapped batch law. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_statistic","label":"iIndepFun_iidEpisodeBatch_statistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_statistic","description":"Measurable fixed-stage batch-record statistics are independent by episode.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-875b985c9dc5","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8313,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:342"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_iidEpisodeBatch_statistic {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) {Target : Type w} [MeasurableSpace Target] (statistic : EpisodeStep State Action -> Target) (hstatistic : Measurable statistic) : ProbabilityTheory.iIndepFun (fun episode (batch : EpisodeBatch mdp episodes) => statistic (batch episode stage)) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iindepfun_iidepisodebatch_statistic banditrlproof.finitehorizonrl.markovpolicy.iindepfun_iidepisodebatch_statistic measurable fixed-stage batch-record statistics are independent by episode. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeStatistic","label":"iIndepFun_episodeStatistic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeStatistic","description":"Any measurable statistic of a fixed-stage record remains independent by episode.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-350678eb0592","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8314,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:360"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_episodeStatistic {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) {Target : Type w} [MeasurableSpace Target] (statistic : EpisodeStep State Action -> Target) (hstatistic : Measurable statistic) : ProbabilityTheory.iIndepFun (fun episode trajectories => statistic (mdp.episodeStepOfTrajectory (trajectories episode) stage)) (policy.iidTrajectoryFamilyMeasure initialState episodes)","missing":[],"search":"iindepfun_episodestatistic banditrlproof.finitehorizonrl.markovpolicy.iindepfun_episodestatistic any measurable statistic of a fixed-stage record remains independent by episode. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryVisitContribution","label":"iIndepFun_trajectoryVisitContribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryVisitContribution","description":"Visit-count summands are independent across iid episode trajectories.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-6d90feb540a0","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8315,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:377"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_trajectoryVisitContribution {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.trajectoryVisitContribution stage state action (trajectories episode)) (policy.iidTrajectoryFamilyMeasure initialState episodes)","missing":[],"search":"iindepfun_trajectoryvisitcontribution banditrlproof.finitehorizonrl.markovpolicy.iindepfun_trajectoryvisitcontribution visit-count summands are independent across iid episode trajectories. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryRewardContribution","label":"iIndepFun_trajectoryRewardContribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryRewardContribution","description":"Reward-sum summands are independent across iid episode trajectories.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-d69c1f66fa98","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8316,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:393"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_trajectoryRewardContribution {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.trajectoryRewardContribution stage state action (trajectories episode)) (policy.iidTrajectoryFamilyMeasure initialState episodes)","missing":[],"search":"iindepfun_trajectoryrewardcontribution banditrlproof.finitehorizonrl.markovpolicy.iindepfun_trajectoryrewardcontribution reward-sum summands are independent across iid episode trajectories. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryTransitionContribution","label":"iIndepFun_trajectoryTransitionContribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryTransitionContribution","description":"Transition-count summands are independent across iid episode trajectories.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-8c35f261a29a","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8317,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:409"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_trajectoryTransitionContribution {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.trajectoryTransitionContribution stage state action nextState (trajectories episode)) (policy.iidTrajectoryFamilyMeasure initialState episodes)","missing":[],"search":"iindepfun_trajectorytransitioncontribution banditrlproof.finitehorizonrl.markovpolicy.iindepfun_trajectorytransitioncontribution transition-count summands are independent across iid episode trajectories. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_stepLaw_and_independence","label":"iidEpisodeBatch_stepLaw_and_independence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_stepLaw_and_independence","description":"Route endpoint: generated batch coordinates have the correct marginal law and are independent across episodes at every fixed stage.","url":"../modules/banditrlproof-rl-finitehorizoniidtrajectorybatch/index.html#decl-75c12639d7bb","parent":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","order":8318,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch"],["Source","BanditRLProof/RL/FiniteHorizonIIDTrajectoryBatch.lean:430"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidEpisodeBatch_stepLaw_and_independence {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) : (forall episode : Fin episodes, (policy.iidEpisodeBatchMeasure initialState episodes).map (fun batch => batch episode stage) = (policy.trajectoryMeasure initialState).map (fun trajectory => mdp.episodeStepOfTrajectory trajectory stage)) /\\ ProbabilityTheory.iIndepFun (fun episode (batch : EpisodeBatch mdp episodes) => batch episode stage) (policy.iidEpisodeBatchMeasure initialState episodes)","missing":[],"search":"iidepisodebatch_steplaw_and_independence banditrlproof.finitehorizonrl.markovpolicy.iidepisodebatch_steplaw_and_independence route endpoint: generated batch coordinates have the correct marginal law and are independent across episodes at every fixed stage. theorem compiled","shard":"modules/67ee00221d433aae.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP","label":"MDP","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP","description":"A finite-state, finite-action, finite-horizon MDP backed by a Mathlib Markov transition kernel. The reward is allowed to depend on the current state and action; stochastic rewards can be added later through a separate reward kernel.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html#decl-9d750d4e9d4a","parent":"module:BanditRLProof.RL.FiniteHorizonMDP","order":8319,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonMDP"],["Source","BanditRLProof/RL/FiniteHorizonMDP.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure MDP (State : Type u) (Action : Type v) [MeasurableSpace State] [MeasurableSpace Action] [Fintype State] [Fintype Action] where","missing":[],"search":"mdp banditrlproof.finitehorizonrl.mdp a finite-state, finite-action, finite-horizon mdp backed by a mathlib markov transition kernel. the reward is allowed to depend on the current state and action; stochastic rewards can be added later through a separate reward kernel. structure compiled","shard":"modules/1b81b75c6f8471fa.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue","label":"transitionValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionValue","description":"Expected continuation value after taking `action` in `state`.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html#decl-578ef7d39b2d","parent":"module:BanditRLProof.RL.FiniteHorizonMDP","order":8320,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonMDP"],["Source","BanditRLProof/RL/FiniteHorizonMDP.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def transitionValue (mdp : MDP State Action) (value : State → Real) (state : State) (action : Action) : Real","missing":[],"search":"transitionvalue banditrlproof.finitehorizonrl.mdp.transitionvalue expected continuation value after taking `action` in `state`. definition compiled","shard":"modules/1b81b75c6f8471fa.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ","label":"bellmanQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.bellmanQ","description":"One-step Bellman action value `r(s,a) + E[V(S') | s,a]`.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html#decl-45c54a0a0951","parent":"module:BanditRLProof.RL.FiniteHorizonMDP","order":8321,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonMDP"],["Source","BanditRLProof/RL/FiniteHorizonMDP.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellmanQ (mdp : MDP State Action) (value : State → Real) (state : State) (action : Action) : Real","missing":[],"search":"bellmanq banditrlproof.finitehorizonrl.mdp.bellmanq one-step bellman action value `r(s,a) + e[v(s') | s,a]`. definition compiled","shard":"modules/1b81b75c6f8471fa.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionValue","label":"measurable_transitionValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionValue","description":"The continuation-value surface is measurable in the state-action pair.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html#decl-d2e6ce40b5b1","parent":"module:BanditRLProof.RL.FiniteHorizonMDP","order":8322,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonMDP"],["Source","BanditRLProof/RL/FiniteHorizonMDP.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_transitionValue (mdp : MDP State Action) {value : State → Real} (hvalue : Measurable value) : Measurable (Function.uncurry (mdp.transitionValue value))","missing":[],"search":"measurable_transitionvalue banditrlproof.finitehorizonrl.mdp.measurable_transitionvalue the continuation-value surface is measurable in the state-action pair. theorem compiled","shard":"modules/1b81b75c6f8471fa.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_bellmanQ","label":"measurable_bellmanQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_bellmanQ","description":"The one-step Bellman action value is measurable in the state-action pair.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html#decl-8d207f49f09a","parent":"module:BanditRLProof.RL.FiniteHorizonMDP","order":8323,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonMDP"],["Source","BanditRLProof/RL/FiniteHorizonMDP.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_bellmanQ (mdp : MDP State Action) {value : State → Real} (hvalue : Measurable value) : Measurable (Function.uncurry (mdp.bellmanQ value))","missing":[],"search":"measurable_bellmanq banditrlproof.finitehorizonrl.mdp.measurable_bellmanq the one-step bellman action value is measurable in the state-action pair. theorem compiled","shard":"modules/1b81b75c6f8471fa.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_zero","label":"bellmanQ_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_zero","description":"With zero continuation value, the Bellman action value is the reward.","url":"../modules/banditrlproof-rl-finitehorizonmdp/index.html#decl-c3fe66340768","parent":"module:BanditRLProof.RL.FiniteHorizonMDP","order":8324,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonMDP"],["Source","BanditRLProof/RL/FiniteHorizonMDP.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanQ_zero (mdp : MDP State Action) (state : State) (action : Action) : mdp.bellmanQ (fun _ => 0) state action = mdp.reward state action","missing":[],"search":"bellmanq_zero banditrlproof.finitehorizonrl.mdp.bellmanq_zero with zero continuation value, the bellman action value is the reward. theorem compiled","shard":"modules/1b81b75c6f8471fa.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","description":"Pathwise equal-round average of successor-policy expected regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-eab445b0a940","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8325,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess pathwise equal-round average of successor-policy expected regret. definition compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnDeviationProcess","label":"selfConsistentScheduledNaturalCausalAverageReturnDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnDeviationProcess","description":"Cumulative normalized successor-return deviation divided by round count.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-2f72a96e214c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8326,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageReturnDeviationProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturndeviationprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturndeviationprocess cumulative normalized successor-return deviation divided by round count. definition compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta","label":"naturalAllPrefixReturnDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta","description":"Summable confidence share for the return event at prefix `n + 1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-014752e1fe2b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8327,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAllPrefixReturnDelta (n : Nat) : Real","missing":[],"search":"naturalallprefixreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixreturndelta summable confidence share for the return event at prefix `n + 1`. definition compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta_pos","label":"naturalAllPrefixReturnDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta_pos","description":"theorem naturalAllPrefixReturnDelta_pos (n : Nat) : 0 < naturalAllPrefixReturnDelta n","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-90f10804e397","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8328,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAllPrefixReturnDelta_pos (n : Nat) : 0 < naturalAllPrefixReturnDelta n","missing":[],"search":"naturalallprefixreturndelta_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixreturndelta_pos theorem naturalallprefixreturndelta_pos (n : nat) : 0 < naturalallprefixreturndelta n theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta_le_one","label":"naturalAllPrefixReturnDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta_le_one","description":"theorem naturalAllPrefixReturnDelta_le_one (n : Nat) : naturalAllPrefixReturnDelta n <= 1","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-2e8c2da70858","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8329,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:84"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAllPrefixReturnDelta_le_one (n : Nat) : naturalAllPrefixReturnDelta n <= 1","missing":[],"search":"naturalallprefixreturndelta_le_one banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixreturndelta_le_one theorem naturalallprefixreturndelta_le_one (n : nat) : naturalallprefixreturndelta n <= 1 theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_naturalAllPrefixReturnDelta","label":"summable_naturalAllPrefixReturnDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_naturalAllPrefixReturnDelta","description":"theorem summable_naturalAllPrefixReturnDelta : Summable naturalAllPrefixReturnDelta","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-a06f4571c7bd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8330,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:93"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_naturalAllPrefixReturnDelta : Summable naturalAllPrefixReturnDelta","missing":[],"search":"summable_naturalallprefixreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_naturalallprefixreturndelta theorem summable_naturalallprefixreturndelta : summable naturalallprefixreturndelta theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius","label":"naturalAllPrefixAverageReturnConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius","description":"Fixed-prefix return radius after division by the positive prefix length.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-4ca6112cd559","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8331,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAllPrefixAverageReturnConfidenceRadius (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"naturalallprefixaveragereturnconfidenceradius banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixaveragereturnconfidenceradius fixed-prefix return radius after division by the positive prefix length. definition compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope","label":"naturalAllPrefixAverageReturnConfidenceEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope","description":"Deterministic square-root envelope for the normalized all-prefix radius.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-519216ecd0ed","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8332,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:114"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAllPrefixAverageReturnConfidenceEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"naturalallprefixaveragereturnconfidenceenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixaveragereturnconfidenceenvelope deterministic square-root envelope for the normalized all-prefix radius. definition compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.log_two_div_naturalAllPrefixReturnDelta_le","label":"log_two_div_naturalAllPrefixReturnDelta_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.log_two_div_naturalAllPrefixReturnDelta_le","description":"The inverse-square confidence share contributes at most three shifted logs.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-178d248adccc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8333,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem log_two_div_naturalAllPrefixReturnDelta_le (n : Nat) : Real.log (2 / naturalAllPrefixReturnDelta n) <= 3 * (1 + Real.log ((n + 2 : Nat) : Real))","missing":[],"search":"log_two_div_naturalallprefixreturndelta_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.log_two_div_naturalallprefixreturndelta_le the inverse-square confidence share contributes at most three shifted logs. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope_nonneg","label":"naturalAllPrefixAverageReturnConfidenceEnvelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope_nonneg","description":"theorem naturalAllPrefixAverageReturnConfidenceEnvelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : 0 <= naturalAllPrefixAverageReturnConfidenceEnvelope mdp varianceProxy n","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-91e2f1fdf132","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8334,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:152"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAllPrefixAverageReturnConfidenceEnvelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : 0 <= naturalAllPrefixAverageReturnConfidenceEnvelope mdp varianceProxy n","missing":[],"search":"naturalallprefixaveragereturnconfidenceenvelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixaveragereturnconfidenceenvelope_nonneg theorem naturalallprefixaveragereturnconfidenceenvelope_nonneg (mdp : mdp state action) (varianceproxy : nnreal) (n : nat) : 0 <= naturalallprefixaveragereturnconfidenceenvelope mdp varianceproxy n theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius_le_envelope","label":"naturalAllPrefixAverageReturnConfidenceRadius_le_envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius_le_envelope","description":"The normalized fixed-prefix return radius is below its deterministic envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-1c9688e30aaa","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8335,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAllPrefixAverageReturnConfidenceRadius_le_envelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : naturalAllPrefixAverageReturnConfidenceRadius mdp varianceProxy baseVisitFloor n <= naturalAllPrefixAverageReturnConfidenceEnvelope mdp varianceProxy n","missing":[],"search":"naturalallprefixaveragereturnconfidenceradius_le_envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixaveragereturnconfidenceradius_le_envelope the normalized fixed-prefix return radius is below its deterministic envelope. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope_tendsto_zero","label":"naturalAllPrefixAverageReturnConfidenceEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope_tendsto_zero","description":"The deterministic all-prefix return-radius envelope vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-d6ebe9f6f62a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8336,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:257"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAllPrefixAverageReturnConfidenceEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (naturalAllPrefixAverageReturnConfidenceEnvelope mdp varianceProxy) atTop (nhds 0)","missing":[],"search":"naturalallprefixaveragereturnconfidenceenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixaveragereturnconfidenceenvelope_tendsto_zero the deterministic all-prefix return-radius envelope vanishes. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius_tendsto_zero","label":"naturalAllPrefixAverageReturnConfidenceRadius_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius_tendsto_zero","description":"The normalized all-prefix return confidence radius vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-2340f65c1541","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8337,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAllPrefixAverageReturnConfidenceRadius_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (naturalAllPrefixAverageReturnConfidenceRadius mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"naturalallprefixaveragereturnconfidenceradius_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixaveragereturnconfidenceradius_tendsto_zero the normalized all-prefix return confidence radius vanishes. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnBadEvent","label":"naturalAllPrefixReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnBadEvent","description":"Return bad event at positive prefix `n + 1` and inverse-square share.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-ac1226fb6573","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8338,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:339"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAllPrefixReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"naturalallprefixreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.naturalallprefixreturnbadevent return bad event at positive prefix `n + 1` and inverse-square share. definition compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_naturalAllPrefixReturnBadEvent","label":"measurableSet_naturalAllPrefixReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_naturalAllPrefixReturnBadEvent","description":"The all-prefix shifted return event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-73e4fb772dee","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8339,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:355"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_naturalAllPrefixReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : MeasurableSet (naturalAllPrefixReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n)","missing":[],"search":"measurableset_naturalallprefixreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_naturalallprefixreturnbadevent the all-prefix shifted return event is measurable. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalAllPrefixReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_naturalAllPrefixReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalAllPrefixReturnBadEvent_le","description":"Each shifted return event has probability at most its inverse-square share.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-8bf7ef2b9771","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8340,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:371"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_naturalAllPrefixReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (n : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure (naturalAllPrefixReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n) <= ENNReal.ofReal (naturalAllPr…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_naturalallprefixreturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_naturalallprefixreturnbadevent_le each shifted return event has probability at most its inverse-square share. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_naturalAllPrefixReturnBadEvent_measure_ne_top","label":"tsum_naturalAllPrefixReturnBadEvent_measure_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_naturalAllPrefixReturnBadEvent_measure_ne_top","description":"The shifted return-event probabilities have finite total mass.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-2c9cefe2630f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8341,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:397"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_naturalAllPrefixReturnBadEvent_measure_ne_top (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (∑' n, source.trajectoryMeasure (naturalAllPrefixReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n)) ≠ ∞","missing":[],"search":"tsum_naturalallprefixreturnbadevent_measure_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_naturalallprefixreturnbadevent_measure_ne_top the shifted return-event probabilities have finite total mass. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_naturalAllPrefixReturnBadEvent","label":"ae_eventually_not_mem_naturalAllPrefixReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_naturalAllPrefixReturnBadEvent","description":"Almost every trajectory eventually avoids all shifted return bad events.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-0642c564368c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8342,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:422"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_not_mem_naturalAllPrefixReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor ∀ᵐ trajectory ∂source.trajectoryMeasure, ∀ᶠ n in atTop, trajectory ∉ naturalAllPrefixReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n","missing":[],"search":"ae_eventually_not_mem_naturalallprefixreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_not_mem_naturalallprefixreturnbadevent almost every trajectory eventually avoids all shifted return bad events. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_naturalAverageBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","description":"The pathwise behavior expected-regret Cesaro average tends to zero a.e.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-f2f43da86a0a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8343,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:447"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_naturalAverageBehaviorExpectedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState…","missing":[],"search":"selfconsistentscheduledcausalsource_naturalaveragebehaviorexpectedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_naturalaveragebehaviorexpectedregret_tendstoalmosteverywhere_zero the pathwise behavior expected-regret cesaro average tends to zero a.e. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageReturnDeviation_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_naturalAverageReturnDeviation_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageReturnDeviation_tendstoAlmostEverywhere_zero","description":"The equal-round normalized return deviation vanishes on almost every trajectory.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-7c1179a9e5b6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8344,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:483"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_naturalAverageReturnDeviation_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor ∀ᵐ trajectory ∂source.trajectoryMeasure, Tendsto (fun rounds => selfConsistentScheduledNaturalCausalAverageReturnDeviationProcess mdp initialState rewardSource initialTable defaul…","missing":[],"search":"selfconsistentscheduledcausalsource_naturalaveragereturndeviation_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_naturalaveragereturndeviation_tendstoalmosteverywhere_zero the equal-round normalized return deviation vanishes on almost every trajectory. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","description":"All-prefix natural average realized behavior regret converges almost surely.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsureconsistency/index.html#decl-78135017c395","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","order":8345,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency.lean:557"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState…","missing":[],"search":"selfconsistentscheduledcausalsource_naturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_naturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero all-prefix natural average realized behavior regret converges almost surely. theorem compiled","shard":"modules/39bb2d69667e182e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","label":"explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","description":"Shifted exponent-three envelope for the scheduled behavior-regret L1 term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-b88e3553fae0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8346,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope (mdp : MDP State Action) (n : Nat) : Real","missing":[],"search":"explicitpolynomialprefixaveragebehaviorregretl1summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragebehaviorregretl1summableenvelope shifted exponent-three envelope for the scheduled behavior-regret l1 term. definition compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnL1SummableEnvelope","label":"explicitPolynomialPrefixAverageReturnL1SummableEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnL1SummableEnvelope","description":"Shifted exponent-two envelope for the scheduled normalized-return L1 term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-d5c1da0af1ec","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8347,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageReturnL1SummableEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"explicitpolynomialprefixaveragereturnl1summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragereturnl1summableenvelope shifted exponent-two envelope for the scheduled normalized-return l1 term. definition compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","description":"Summable deterministic L1 envelope on the fourth-power prefix schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-65972e958b9b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8348,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : Real","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretl1summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretl1summableenvelope summable deterministic l1 envelope on the fourth-power prefix schedule. definition compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_explicitRounds_le_summableEnvelope","label":"selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_explicitRounds_le_summableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_explicitRounds_le_summableEnvelope","description":"The scheduled logarithmic average is dominated by a shifted exponent-three term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-ef48fa8db988","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8349,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_explicitRounds_le_summableEnvelope (mdp : MDP State Action) (n : Nat) : selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate mdp (explicitHighProbabilityRounds n) <= explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope mdp n","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_explicitrounds_le_summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_explicitrounds_le_summableenvelope the scheduled logarithmic average is dominated by a shifted exponent-three term. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_explicitRounds_le_summableEnvelope","label":"selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_explicitRounds_le_summableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_explicitRounds_le_summableEnvelope","description":"The scheduled return first moment is dominated by a shifted exponent-two term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-71d863c01be5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8350,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_explicitRounds_le_summableEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound mdp varianceProxy baseVisitFloor (explicitHighProbabilityRounds n) <= explicitPolynomialPrefixAverageReturnL1SummableEnvelope mdp varianceProxy n","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_explicitrounds_le_summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_explicitrounds_le_summableenvelope the scheduled return first moment is dominated by a shifted exponent-two term. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope_nonneg","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope_nonneg","description":"The explicit scheduled L1 envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-1ee5cde6656b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8351,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:131"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (n : Nat) : 0 <= explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope mdp varianceProxy n","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretl1summableenvelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretl1summableenvelope_nonneg the explicit scheduled l1 envelope is nonnegative. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_explicitRounds_le_summableEnvelope","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_explicitRounds_le_summableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_explicitRounds_le_summableEnvelope","description":"The all-prefix L1 envelope is pointwise controlled on fourth-power prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-ae8845827125","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8352,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:150"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_explicitRounds_le_summableEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope mdp varianceProxy baseVisitFloor (explicitHighProbabilityRounds n) <= explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope mdp varianceProxy n","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_explicitrounds_le_summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_explicitrounds_le_summableenvelope the all-prefix l1 envelope is pointwise controlled on fourth-power prefixes. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","label":"summable_explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","description":"The shifted exponent-three behavior envelope is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-18e3c8ae9318","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8353,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope (mdp : MDP State Action) : Summable (explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope mdp)","missing":[],"search":"summable_explicitpolynomialprefixaveragebehaviorregretl1summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_explicitpolynomialprefixaveragebehaviorregretl1summableenvelope the shifted exponent-three behavior envelope is summable. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageReturnL1SummableEnvelope","label":"summable_explicitPolynomialPrefixAverageReturnL1SummableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageReturnL1SummableEnvelope","description":"The shifted exponent-two return envelope is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-4c8f8986072a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8354,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_explicitPolynomialPrefixAverageReturnL1SummableEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) : Summable (explicitPolynomialPrefixAverageReturnL1SummableEnvelope mdp varianceProxy)","missing":[],"search":"summable_explicitpolynomialprefixaveragereturnl1summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_explicitpolynomialprefixaveragereturnl1summableenvelope the shifted exponent-two return envelope is summable. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","label":"summable_explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","description":"The deterministic fourth-power scheduled L1 envelope is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-8219f5723185","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8355,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:211"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) : Summable (explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope mdp varianceProxy)","missing":[],"search":"summable_explicitpolynomialprefixaveragerealizedbehaviorregretl1summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_explicitpolynomialprefixaveragerealizedbehaviorregretl1summableenvelope the deterministic fourth-power scheduled l1 envelope is summable. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","label":"explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","description":"Expected absolute exact average realized regret on the fourth-power schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-111d36691a86","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8356,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:224"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret expected absolute exact average realized regret on the fourth-power schedule. definition compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","label":"explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","description":"The scheduled expected absolute process is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-4780af8f5cce","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8357,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:237"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : 0 <= explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n","missing":[],"search":"explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_nonneg the scheduled expected absolute process is nonnegative. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_le_summableEnvelope","label":"explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_le_summableEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_le_summableEnvelope","description":"The scheduled expected absolute process is bounded by the summable envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-2fd036084c1b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8358,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:253"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_le_summableEnvelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialState rewardSource initialT…","missing":[],"search":"explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_le_summableenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_le_summableenvelope the scheduled expected absolute process is bounded by the summable envelope. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","label":"summable_explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","description":"Scheduled expected absolute exact average realized regret is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-5e9098ee39ed","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8359,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Summable (explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaul…","missing":[],"search":"summable_explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret scheduled expected absolute exact average realized regret is summable. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixReciprocalThreshold","label":"explicitPolynomialPrefixReciprocalThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixReciprocalThreshold","description":"Reciprocal thresholds used to turn fixed-threshold Borel-Cantelli into convergence.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-6c3555ee6aaa","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8360,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:314"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixReciprocalThreshold (k : Nat) : Real","missing":[],"search":"explicitpolynomialprefixreciprocalthreshold banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixreciprocalthreshold reciprocal thresholds used to turn fixed-threshold borel-cantelli into convergence. definition compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixReciprocalThreshold_pos","label":"explicitPolynomialPrefixReciprocalThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixReciprocalThreshold_pos","description":"theorem explicitPolynomialPrefixReciprocalThreshold_pos (k : Nat) : 0 < explicitPolynomialPrefixReciprocalThreshold k","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-fbed1f36492f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8361,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:317"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixReciprocalThreshold_pos (k : Nat) : 0 < explicitPolynomialPrefixReciprocalThreshold k","missing":[],"search":"explicitpolynomialprefixreciprocalthreshold_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixreciprocalthreshold_pos theorem explicitpolynomialprefixreciprocalthreshold_pos (k : nat) : 0 < explicitpolynomialprefixreciprocalthreshold k theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_expectedAbsolute_div","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_expectedAbsolute_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_expectedAbsolute_div","description":"Markov's inequality for one scheduled distance-from-zero violation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-4e1d14fbbd55","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8362,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:323"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_expectedAbsolute_div (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor epsilon : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hepsilon : 0 < epsilon) (n : Nat) : explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon n <= ENNReal.ofReal (explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialState reward…","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_le_expectedabsolute_div banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_le_expectedabsolute_div markov's inequality for one scheduled distance-from-zero violation. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_ne_top","label":"tsum_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_ne_top","description":"Fixed positive scheduled violation probabilities have finite total mass.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-8ead6740008a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8363,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:399"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_ne_top (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : (∑' n, explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViol…","missing":[],"search":"tsum_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_ne_top fixed positive scheduled violation probabilities have finite total mass. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_explicitPolynomialPrefixAverageRealizedBehaviorRegretReciprocalDistanceViolationSet","label":"ae_eventually_not_mem_explicitPolynomialPrefixAverageRealizedBehaviorRegretReciprocalDistanceViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_explicitPolynomialPrefixAverageRealizedBehaviorRegretReciprocalDistanceViolationSet","description":"Almost every trajectory eventually avoids every reciprocal scheduled violation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-1b5a1c016d25","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8364,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:435"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_not_mem_explicitPolynomialPrefixAverageRealizedBehaviorRegretReciprocalDistanceViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultSt…","missing":[],"search":"ae_eventually_not_mem_explicitpolynomialprefixaveragerealizedbehaviorregretreciprocaldistanceviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_not_mem_explicitpolynomialprefixaveragerealizedbehaviorregretreciprocaldistanceviolationset almost every trajectory eventually avoids every reciprocal scheduled violation. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","description":"The exact equal-round average process converges a.e. on fourth-power prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-a1e32379b8b9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8365,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:470"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseV…","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero the exact equal-round average process converges a.e. on fourth-power prefixes. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","description":"Scheduled L1 summability and almost-sure consistency on the common source.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretalmostsuree-90d339330735/index.html#decl-903bda9dd326","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","order":8366,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule.lean:523"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTa…","missing":[],"search":"selfconsistentscheduledcausalsource_explicitpolynomialprefixaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_explicitpolynomialprefixaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero scheduled l1 summability and almost-sure consistency on the common source. theorem compiled","shard":"modules/c727419ba9ecf20b.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_hasSubgaussianMGF","label":"trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_hasSubgaussianMGF","description":"The natural normalized successor-return sum has one global sub-Gaussian MGF.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-e3c327435044","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8367,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (…","missing":[],"search":"trajectorymeasure_naturalcumulativesuccessoraveragereturndeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_naturalcumulativesuccessoraveragereturndeviation_hassubgaussianmgf the natural normalized successor-return sum has one global sub-gaussian mgf. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_hasSubgaussianMGF","label":"selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_hasSubgaussianMGF","description":"The self-consistent natural cumulative normalized-return deviation has a global MGF.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-77d29b0eb5a1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8368,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:91"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_hasSubgaussianMGF (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : HasSubgaussianMGF (selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) (selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy mdp varianceProxy baseVisitFloor rounds) (selfConsistentScheduledCausalSource mdp initia…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativereturndeviationprocess_hassubgaussianmgf banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativereturndeviationprocess_hassubgaussianmgf the self-consistent natural cumulative normalized-return deviation has a global mgf. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound","label":"selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound","description":"First-moment envelope for the round-normalized cumulative return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-f8db2aea2811","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8369,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound first-moment envelope for the round-normalized cumulative return deviation. definition compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope","label":"selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope","description":"Coarser inverse-square-root envelope obtained from the linear proxy bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-00bc9ed59b00","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8370,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturninversesqrtenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturninversesqrtenvelope coarser inverse-square-root envelope obtained from the linear proxy bound. definition compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_nonneg","label":"selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_nonneg","description":"The normalized-return first-moment envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-232528d99292","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8371,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_nonneg the normalized-return first-moment envelope is nonnegative. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_le","label":"integral_abs_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_le","description":"The MGF controls the absolute first moment of the cumulative return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-c14947a1e6e0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8372,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:156"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : integral (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure (fun trajectory => |selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds traject…","missing":[],"search":"integral_abs_selfconsistentschedulednaturalcausalcumulativereturndeviationprocess_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_abs_selfconsistentschedulednaturalcausalcumulativereturndeviationprocess_le the mgf controls the absolute first moment of the cumulative return deviation. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_le_inverseSqrtEnvelope","label":"selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_le_inverseSqrtEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_le_inverseSqrtEnvelope","description":"The normalized first-moment envelope is bounded by an inverse square root.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-865d6d9c465b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8373,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_le_inverseSqrtEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (hrounds : 0 < rounds) : selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound mdp varianceProxy baseVisitFloor rounds <= selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope mdp varianceProxy rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_le_inversesqrtenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_le_inversesqrtenvelope the normalized first-moment envelope is bounded by an inverse square root. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope_tendsto_zero","label":"selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope_tendsto_zero","description":"The deterministic inverse-square-root return envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-5f8bffabeb22","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8374,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:235"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope mdp varianceProxy) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturninversesqrtenvelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturninversesqrtenvelope_tendsto_zero the deterministic inverse-square-root return envelope tends to zero. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_tendsto_zero","label":"selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_tendsto_zero","description":"The exact normalized-return first-moment envelope tends to zero on all prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-4ea8fdbe6621","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8375,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragereturnfirstmomentbound_tendsto_zero the exact normalized-return first-moment envelope tends to zero on all prefixes. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","label":"integrable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","description":"The cumulative normalized-return deviation is integrable on every prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-2f245590f421","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8376,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : Integrable (selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalcumulativereturndeviationprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalcumulativereturndeviationprocess the cumulative normalized-return deviation is integrable on every prefix. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","label":"integrable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","description":"The equal-round-weighted natural average realized behavior regret is integrable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-fdc3d0cb0a86","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8377,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:292"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : Integrable (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess the equal-round-weighted natural average realized behavior regret is integrable. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret","description":"Expected absolute equal-round-weighted natural average realized behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-bb2ca7bf3626","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8378,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:340"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret expected absolute equal-round-weighted natural average realized behavior regret. definition compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope","description":"Deterministic all-prefix L1 envelope for the exact average process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-61f5e1144e8f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8379,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:356"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope deterministic all-prefix l1 envelope for the exact average process. definition compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_nonneg","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_nonneg","description":"The deterministic L1 envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-777b895f68de","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8380,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:368"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_nonneg the deterministic l1 envelope is nonnegative. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_tendsto_zero","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_tendsto_zero","description":"The deterministic all-prefix L1 envelope tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-47c3f2ff58f1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8381,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:383"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_tendsto_zero the deterministic all-prefix l1 envelope tends to zero. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","description":"Expected absolute average realized regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-bf691ad41f6c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8382,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:397"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret_nonneg expected absolute average realized regret is nonnegative. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_le_L1Envelope","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_le_L1Envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_le_L1Envelope","description":"The expected absolute exact average process is bounded by the all-prefix L1 envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-68880cdf9ff0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8383,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:411"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_le_L1Envelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialStat…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret_le_l1envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret_le_l1envelope the expected absolute exact average process is bounded by the all-prefix l1 envelope. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_tendsto_zero","description":"Expected absolute exact average realized behavior regret tends to zero on all prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-aac0c9cef272","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8384,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:537"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret mdp initialState rewar…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteaveragerealizedbehaviorregret_tendsto_zero expected absolute exact average realized behavior regret tends to zero on all prefixes. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","label":"memLp_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","description":"Every exact equal-round average realized-regret coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-7bb9b5db1e60","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8385,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:570"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : MemLp (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess every exact equal-round average realized-regret coordinate belongs to `l1`. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq","description":"At exponent one, `eLpNorm` is the lifted expected absolute exact average regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-3bae7ac3f17c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8386,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:595"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : eLpNorm (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure = ENNReal.ofReal…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_eq at exponent one, `elpnorm` is the lifted expected absolute exact average regret. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_tendsto_zero","description":"The exponent-one extended norm of the exact average process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-5e838f0407b9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8387,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:625"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun rounds => eLpNorm (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp i…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_tendsto_zero the exponent-one extended norm of the exact average process tends to zero. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","description":"The exponent-one norm of the difference from zero tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-9c7e79108513","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8388,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:660"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (fun rounds => eLpNorm (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProc…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_sub_zero_tendsto_zero the exponent-one norm of the difference from zero tends to zero. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp","description":"The exact equal-round average realized behavior regret as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-0abac35331ff","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8389,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:695"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : Lp Real 1 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp the exact equal-round average realized behavior regret as an `lp real 1` value. definition compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_coeFn_ae_eq","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_coeFn_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_coeFn_ae_eq","description":"The named `Lp` coordinate represents the exact average process a.e.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-3d37604099ae","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8390,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:717"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_coeFn_ae_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp mdp initialState rewardSource varianceProxy law initialTable defaultState baseVisitFloor hrewardBound rounds : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp_coefn_ae_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp_coefn_ae_eq the named `lp` coordinate represents the exact average process a.e. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_tendsto_zero","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_tendsto_zero","description":"The named exact average `Lp Real 1` process converges to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-8fe20041fc36","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8391,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:746"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp mdp initialState rewardSource varianceProxy law in…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp_tendsto_zero the named exact average `lp real 1` process converges to zero. theorem compiled","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_allPrefix_L1_tendsto_zero","label":"selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_allPrefix_L1_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_allPrefix_L1_tendsto_zero","description":"The named exact average `Lp Real 1` process converges to zero. -/ theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law :…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalaveragerealizedbehaviorregretl1consistency/index.html#decl-1f7841eb8d06","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","order":8392,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency.lean:797"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_allPrefix_L1_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState var…","missing":[],"search":"selfconsistentscheduledcausalsource_naturalaveragerealizedbehaviorregret_allprefix_l1_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_naturalaveragerealizedbehaviorregret_allprefix_l1_tendsto_zero the named exact average `lp real 1` process converges to zero. -/ theorem selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlp mdp initialstate rewardsource varianceproxy law initialtable defaultstate basevisitfloor hrewardbound) attop (nhds 0) := by let source := selfconsistentscheduledcausalsource mdp initialstate rewardsource i…","shard":"modules/b977209f4618234f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate","label":"selfConsistentScheduledNaturalCausalCumulativePlanningRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate","description":"Natural-prefix sum of the causal planning-rate coordinates.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-207043870a1b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8393,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativePlanningRate (mdp : MDP State Action) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeplanningrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeplanningrate natural-prefix sum of the causal planning-rate coordinates. definition compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_integrated","label":"selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_integrated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_integrated","description":"The pathwise planning sum is dominated by the integrated finite-prefix sum.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-63144ec0f241","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8394,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_integrated (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativePlanningRate mdp rounds <= selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeplanningrate_le_integrated banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeplanningrate_le_integrated the pathwise planning sum is dominated by the integrated finite-prefix sum. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_logarithmic","label":"selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_logarithmic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_logarithmic","description":"The natural-prefix planning sum has the existing explicit logarithmic envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-d7eb9c2523e2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8395,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_logarithmic (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativePlanningRate mdp rounds <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeplanningrate_le_logarithmic banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeplanningrate_le_logarithmic the natural-prefix planning sum has the existing explicit logarithmic envelope. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_eq_fin_sum","label":"selfConsistentScheduledCausalModelFailureBudget_eq_fin_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_eq_fin_sum","description":"The accumulated prefix budget is exactly the finite sum of both model shares.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-454dcee6609c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8396,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalModelFailureBudget_eq_fin_sum (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledCausalModelFailureBudget mdp rounds = ∑ round : Fin rounds, (ENNReal.ofReal (AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp round) + ENNReal.ofReal (AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp round))","missing":[],"search":"selfconsistentscheduledcausalmodelfailurebudget_eq_fin_sum banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelfailurebudget_eq_fin_sum the accumulated prefix budget is exactly the finite sum of both model shares. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_prefix","label":"not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_prefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_prefix","description":"Avoiding the prefix event implies avoiding every coordinate event in it.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-b110b62366fd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8397,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_prefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s)) {rounds t : Nat} (ht : t < rounds) (hprefix : trajectory ∉ selfConsistentScheduledCausalModelBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) : trajectory ∉ selfConsistentScheduledCausalModelRoundBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor t","missing":[],"search":"not_mem_selfconsistentscheduledcausalmodelroundbadevent_of_not_mem_prefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.not_mem_selfconsistentscheduledcausalmodelroundbadevent_of_not_mem_prefix avoiding the prefix event implies avoiding every coordinate event in it. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","description":"The random natural-prefix cumulative behavior expected-regret process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-0d081b09ca50","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8398,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:138"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess the random natural-prefix cumulative behavior expected-regret process is measurable. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_planning_of_not_mem_modelBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_planning_of_not_mem_modelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_planning_of_not_mem_modelBadEvent","description":"Off the finite-prefix model event, the actual process obeys the planning sum.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-199fd1c8c27d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8399,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:156"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_planning_of_not_mem_modelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => Adaptive…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_planning_of_not_mem_modelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_planning_of_not_mem_modelbadevent off the finite-prefix model event, the actual process obeys the planning sum. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_logarithmic_of_not_mem_modelBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_logarithmic_of_not_mem_modelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_logarithmic_of_not_mem_modelBadEvent","description":"Off the finite-prefix model event, the actual process obeys the log envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-b143836739f4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8400,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:197"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_logarithmic_of_not_mem_modelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => Adapt…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_logarithmic_of_not_mem_modelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_logarithmic_of_not_mem_modelbadevent off the finite-prefix model event, the actual process obeys the log envelope. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","description":"One-sided fixed-prefix violation event for the random cumulative process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-56a1cb059cef","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8401,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:232"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun s => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor s))","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretlogarithmicviolationset one-sided fixed-prefix violation event for the random cumulative process. definition compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","description":"The fixed-prefix logarithmic violation event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-7e37f62b6354","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8402,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:251"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : MeasurableSet (selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretlogarithmicviolationset the fixed-prefix logarithmic violation event is measurable. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet_subset_modelBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet_subset_modelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet_subset_modelBadEvent","description":"Every logarithmic violation lies in the actual finite-prefix model event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-97f3f0f8d469","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8403,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet_subset_modelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicV…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretlogarithmicviolationset_subset_modelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretlogarithmicviolationset_subset_modelbadevent every logarithmic violation lies in the actual finite-prefix model event. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeBehaviorExpectedRegretLogarithmicViolationSet_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeBehaviorExpectedRegretLogarithmicViolationSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeBehaviorExpectedRegretLogarithmicViolationSet_le","description":"The one-sided logarithmic violation probability uses the exact prefix budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-65e8e80cec05","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8404,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:299"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeBehaviorExpectedRegretLogarithmicViolationSet_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_cumulativebehaviorexpectedregretlogarithmicviolationset_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_cumulativebehaviorexpectedregretlogarithmicviolationset_le the one-sided logarithmic violation probability uses the exact prefix budget. theorem compiled","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeBehaviorExpectedRegret","label":"selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeBehaviorExpectedRegret","description":"The one-sided logarithmic violation probability uses the exact prefix budget. -/ theorem selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeBehaviorExpectedRegretLogarithmicViolationSet_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNRea…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregrethighprobabilitylograte/index.html#decl-144e12e9ea96","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","order":8405,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initial…","missing":[],"search":"selfconsistentscheduledcausalsource_fixedprefixhighprobabilitylogarithmiccumulativebehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_fixedprefixhighprobabilitylogarithmiccumulativebehaviorexpectedregret the one-sided logarithmic violation probability uses the exact prefix budget. -/ theorem selfconsistentscheduledcausalsource_trajectorymeasure_cumulativebehaviorexpectedregretlogarithmicviolationset_le (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (rounds : nat) : let source := selfconsistentscheduledcausalsource mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor source.trajectorymeasure (selfconsistentschedulednaturalcausalcumulativ…","shard":"modules/58bf0ff46185918e.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.sum_range_one_div_natCast_add_two_sq_le_one","label":"sum_range_one_div_natCast_add_two_sq_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sum_range_one_div_natCast_add_two_sq_le_one","description":"theorem sum_range_one_div_natCast_add_two_sq_le_one (rounds : Nat) : (Finset.range rounds).sum (fun t => 1 / (((t + 2 : Nat) : Real) ^ 2)) <= 1","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-6048a7e9589c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8406,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_range_one_div_natCast_add_two_sq_le_one (rounds : Nat) : (Finset.range rounds).sum (fun t => 1 / (((t + 2 : Nat) : Real) ^ 2)) <= 1","missing":[],"search":"sum_range_one_div_natcast_add_two_sq_le_one banditrlproof.sum_range_one_div_natcast_add_two_sq_le_one theorem sum_range_one_div_natcast_add_two_sq_le_one (rounds : nat) : (finset.range rounds).sum (fun t => 1 / (((t + 2 : nat) : real) ^ 2)) <= 1 theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.sum_range_one_div_natCast_add_two_pow_le_one","label":"sum_range_one_div_natCast_add_two_pow_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sum_range_one_div_natCast_add_two_pow_le_one","description":"theorem sum_range_one_div_natCast_add_two_pow_le_one (rounds exponent : Nat) (hexponent : 2 <= exponent) : (Finset.range rounds).sum (fun t => 1 / (((t + 2 : Nat) : Real) ^ exponent)) <= 1","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-9bb62a52be65","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8407,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_range_one_div_natCast_add_two_pow_le_one (rounds exponent : Nat) (hexponent : 2 <= exponent) : (Finset.range rounds).sum (fun t => 1 / (((t + 2 : Nat) : Real) ^ exponent)) <= 1","missing":[],"search":"sum_range_one_div_natcast_add_two_pow_le_one banditrlproof.sum_range_one_div_natcast_add_two_pow_le_one theorem sum_range_one_div_natcast_add_two_pow_le_one (rounds exponent : nat) (hexponent : 2 <= exponent) : (finset.range rounds).sum (fun t => 1 / (((t + 2 : nat) : real) ^ exponent)) <= 1 theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.sum_range_one_div_natCast_add_three_le_one_add_log","label":"sum_range_one_div_natCast_add_three_le_one_add_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sum_range_one_div_natCast_add_three_le_one_add_log","description":"theorem sum_range_one_div_natCast_add_three_le_one_add_log (rounds : Nat) : (Finset.range rounds).sum (fun t => 1 / (((t + 3 : Nat) : Real))) <= 1 + Real.log (rounds : Real)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-023c2835bc2c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8408,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_range_one_div_natCast_add_three_le_one_add_log (rounds : Nat) : (Finset.range rounds).sum (fun t => 1 / (((t + 3 : Nat) : Real))) <= 1 + Real.log (rounds : Real)","missing":[],"search":"sum_range_one_div_natcast_add_three_le_one_add_log banditrlproof.sum_range_one_div_natcast_add_three_le_one_add_log theorem sum_range_one_div_natcast_add_three_le_one_add_log (rounds : nat) : (finset.range rounds).sum (fun t => 1 / (((t + 3 : nat) : real))) <= 1 + real.log (rounds : real) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.one_le_one_add_log_natCast","label":"one_le_one_add_log_natCast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.one_le_one_add_log_natCast","description":"theorem one_le_one_add_log_natCast (rounds : Nat) : (1 : Real) <= 1 + Real.log (rounds : Real)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-90f0f1abbb7b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8409,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_le_one_add_log_natCast (rounds : Nat) : (1 : Real) <= 1 + Real.log (rounds : Real)","missing":[],"search":"one_le_one_add_log_natcast banditrlproof.one_le_one_add_log_natcast theorem one_le_one_add_log_natcast (rounds : nat) : (1 : real) <= 1 + real.log (rounds : real) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.tendsto_one_add_log_natCast_div_natCast_zero","label":"tendsto_one_add_log_natCast_div_natCast_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tendsto_one_add_log_natCast_div_natCast_zero","description":"theorem tendsto_one_add_log_natCast_div_natCast_zero : Filter.Tendsto (fun rounds : Nat => (1 + Real.log (rounds : Real)) / (rounds : Real)) Filter.atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-9409dfc84927","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8410,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:114"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tendsto_one_add_log_natCast_div_natCast_zero : Filter.Tendsto (fun rounds : Nat => (1 + Real.log (rounds : Real)) / (rounds : Real)) Filter.atTop (nhds 0)","missing":[],"search":"tendsto_one_add_log_natcast_div_natcast_zero banditrlproof.tendsto_one_add_log_natcast_div_natcast_zero theorem tendsto_one_add_log_natcast_div_natcast_zero : filter.tendsto (fun rounds : nat => (1 + real.log (rounds : real)) / (rounds : real)) filter.attop (nhds 0) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient","label":"selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient","description":"noncomputable def selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient (mdp : MDP State Action) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-c55cdb6fa591","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8411,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient (mdp : MDP State Action) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesquareratecoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesquareratecoefficient noncomputable def selfconsistentschedulednaturalcausalinversesquareratecoefficient (mdp : mdp state action) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient","label":"selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient","description":"noncomputable def selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient (mdp : MDP State Action) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-090317f044df","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8412,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:154"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient (mdp : MDP State Action) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient noncomputable def selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient (mdp : mdp state action) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalHighPowerRateCoefficient","label":"selfConsistentScheduledNaturalCausalHighPowerRateCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalHighPowerRateCoefficient","description":"noncomputable def selfConsistentScheduledNaturalCausalHighPowerRateCoefficient (mdp : MDP State Action) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-b13664361222","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8413,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalHighPowerRateCoefficient (mdp : MDP State Action) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalhighpowerratecoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalhighpowerratecoefficient noncomputable def selfconsistentschedulednaturalcausalhighpowerratecoefficient (mdp : mdp state action) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient","label":"selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient","description":"noncomputable def selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient (mdp : MDP State Action) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-9d1ecce74b17","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8414,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient (mdp : MDP State Action) : Real","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicratecoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicratecoefficient noncomputable def selfconsistentschedulednaturalcausallogarithmicratecoefficient (mdp : mdp state action) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate","label":"selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate","description":"noncomputable def selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-b94d6f2e0d7a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8415,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:168"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate noncomputable def selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate (mdp : mdp state action) (rounds : nat) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate","label":"selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate","description":"noncomputable def selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-ed05c248362a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8416,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:175"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate noncomputable def selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate (mdp : mdp state action) (rounds : nat) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate","label":"selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate","description":"noncomputable def selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-19fc1f39d35d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8417,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate (mdp : MDP State Action) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate noncomputable def selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate (mdp : mdp state action) (rounds : nat) : real definition compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq_threeTerm","label":"selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq_threeTerm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq_threeTerm","description":"theorem selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq_threeTerm (mdp : MDP State Action) (t : Nat) : selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp t = selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient mdp * (1 / (((t + 2 : Nat) : Real) ^ 2)) + selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient mdp * (1 / (((t + 3 : Na…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-61936debe28b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8418,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq_threeTerm (mdp : MDP State Action) (t : Nat) : selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt mdp t = selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient mdp * (1 / (((t + 2 : Nat) : Real) ^ 2)) + selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient mdp * (1 / (((t + 3 : Nat) : Real))) + selfConsistentScheduledNaturalCausalHighPowerRateCoefficient mdp * (1 / (((t + 2 : Nat) : Real) ^ (mdp.horizon + 5)))","missing":[],"search":"selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_eq_threeterm banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_eq_threeterm theorem selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat_eq_threeterm (mdp : mdp state action) (t : nat) : selfconsistentschedulednaturalcausalintegratedbehaviorexpectedregretrateat mdp t = selfconsistentschedulednaturalcausalinversesquareratecoefficient mdp * (1 / (((t + 2 : nat) : real) ^ 2)) + selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient mdp * (1 / (((t + 3 : nat) : real))) + selfconsistentschedulednaturalcausalhighpowerratecoefficient mdp * (1 / (((t + 2 : nat) : real) ^ (mdp.horizon + 5))) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient_nonneg","description":"theorem selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient mdp","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-e8d2ab82f208","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8419,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:212"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient mdp","missing":[],"search":"selfconsistentschedulednaturalcausalinversesquareratecoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesquareratecoefficient_nonneg theorem selfconsistentschedulednaturalcausalinversesquareratecoefficient_nonneg (mdp : mdp state action) : 0 <= selfconsistentschedulednaturalcausalinversesquareratecoefficient mdp theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient_nonneg","label":"selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient_nonneg","description":"theorem selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient mdp","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-9e9b5477ad5d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8420,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:221"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient mdp","missing":[],"search":"selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient_nonneg theorem selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient_nonneg (mdp : mdp state action) : 0 <= selfconsistentschedulednaturalcausalexplorationharmonicratecoefficient mdp theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalHighPowerRateCoefficient_nonneg","label":"selfConsistentScheduledNaturalCausalHighPowerRateCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalHighPowerRateCoefficient_nonneg","description":"theorem selfConsistentScheduledNaturalCausalHighPowerRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalHighPowerRateCoefficient mdp","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-fbbb0d2c9830","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8421,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:230"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalHighPowerRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalHighPowerRateCoefficient mdp","missing":[],"search":"selfconsistentschedulednaturalcausalhighpowerratecoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalhighpowerratecoefficient_nonneg theorem selfconsistentschedulednaturalcausalhighpowerratecoefficient_nonneg (mdp : mdp state action) : 0 <= selfconsistentschedulednaturalcausalhighpowerratecoefficient mdp theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient_nonneg","label":"selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient_nonneg","description":"theorem selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient mdp","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-95676faf0331","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8422,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:239"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient_nonneg (mdp : MDP State Action) : 0 <= selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient mdp","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicratecoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicratecoefficient_nonneg theorem selfconsistentschedulednaturalcausallogarithmicratecoefficient_nonneg (mdp : mdp state action) : 0 <= selfconsistentschedulednaturalcausallogarithmicratecoefficient mdp theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_refined","label":"selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_refined","description":"theorem selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_refined (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds <= selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-e00c1b664142","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8423,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:252"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_refined (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds <= selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_le_refined banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_le_refined theorem selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_le_refined (mdp : mdp state action) (rounds : nat) : selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate mdp rounds <= selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","label":"selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","description":"theorem selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-a5a92f7b5cbd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8424,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:297"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate_le_logarithmic banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate_le_logarithmic theorem selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate_le_logarithmic (mdp : mdp state action) (rounds : nat) : selfconsistentschedulednaturalcausalrefinedcumulativeintegratedbehaviorexpectedregretrate mdp rounds <= selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","label":"selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","description":"theorem selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-b4c27f68f42d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8425,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:325"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_le_logarithmic banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_le_logarithmic theorem selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate_le_logarithmic (mdp : mdp state action) (rounds : nat) : selfconsistentschedulednaturalcausalcumulativeintegratedbehaviorexpectedregretrate mdp rounds <= selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","label":"selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","description":"theorem selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_nonneg (mdp : MDP State Action) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-58b2d7827250","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8426,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:339"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_nonneg (mdp : MDP State Action) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate_nonneg theorem selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate_nonneg (mdp : mdp state action) (rounds : nat) : 0 <= selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_nonneg","label":"selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_nonneg","description":"theorem selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_nonneg (mdp : MDP State Action) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate mdp rounds","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-effc186ea516","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8427,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:352"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_nonneg (mdp : MDP State Action) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate mdp rounds","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_nonneg theorem selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_nonneg (mdp : mdp state action) (rounds : nat) : 0 <= selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","label":"selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","description":"theorem selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate mdp) atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-118f88795c2f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8428,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:365"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero (mdp : MDP State Action) : Tendsto (selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate mdp) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_tendsto_zero theorem selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_tendsto_zero (mdp : mdp state action) : tendsto (selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate mdp) attop (nhds 0) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_isBigO","label":"selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_isBigO","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_isBigO","description":"theorem selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_isBigO (mdp : MDP State Action) : (selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp) =O[atTop] (fun rounds : Nat => 1 + Real.log (rounds : Real))","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-609e89f30d5e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8429,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:382"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_isBigO (mdp : MDP State Action) : (selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate mdp) =O[atTop] (fun rounds : Nat => 1 + Real.log (rounds : Real))","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate_isbigo banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate_isbigo theorem selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate_isbigo (mdp : mdp state action) : (selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate mdp) =o[attop] (fun rounds : nat => 1 + real.log (rounds : real)) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_logarithmic","label":"selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_logarithmic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_logarithmic","description":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_logarithmic (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initia…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-253be876f18d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8430,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_logarithmic (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret mdp initialState rewardSource initialTab…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_le_logarithmic banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_le_logarithmic theorem selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_le_logarithmic (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (rounds : nat) : selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds <= selfconsistentschedulednaturalcausallogarithmiccumulativeintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_isBigO_one_add_log","label":"selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_isBigO_one_add_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_isBigO_one_add_log","description":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_isBigO_one_add_log (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (in…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-459cc3e55c17","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8431,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:419"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_isBigO_one_add_log (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret mdp initialState rewardSource initialTable default…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_isbigo_one_add_log banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_isbigo_one_add_log theorem selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret_isbigo_one_add_log (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : (selfconsistentschedulednaturalcausalexpectedcumulativebehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor) =o[attop] (fun rounds : nat => 1 + real.log (rounds : real)) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_logarithmic","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_logarithmic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_logarithmic","description":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_logarithmic (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTa…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-db2b2b504591","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8432,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:473"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_logarithmic (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) : selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret mdp initialState rewardSource initialTable def…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_le_logarithmic banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_le_logarithmic theorem selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_le_logarithmic (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (rounds : nat) : selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor rounds <= selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate mdp rounds theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_isBigO_log_div_natCast","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_isBigO_log_div_natCast","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_isBigO_log_div_natCast","description":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_isBigO_log_div_natCast (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (i…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-197485d7ff54","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8433,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:502"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_isBigO_log_div_natCast (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret mdp initialState rewardSource initialTable defaultSt…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_isbigo_log_div_natcast banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_isbigo_log_div_natcast theorem selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_isbigo_log_div_natcast (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : (selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor) =o[attop] (fun rounds : nat => (1 + real.log (rounds : real)) / (rounds : real)) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero_of_logarithmicRate","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero_of_logarithmicRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero_of_logarithmicRate","description":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero_of_logarithmicRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw variance…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-0e9cd0ed6435","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8434,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:567"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero_of_logarithmicRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret mdp initialState rewardSource initi…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_tendsto_zero_of_logarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_tendsto_zero_of_logarithmicrate theorem selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret_tendsto_zero_of_logarithmicrate (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (selfconsistentschedulednaturalcausalaveragebehaviorexpectedregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor) attop (nhds 0) theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitLogarithmicRate","label":"selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitLogarithmicRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitLogarithmicRate","description":"Same-source explicit logarithmic cumulative and `log(n) / n` average route.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalbehaviorexpectedregretlograte/index.html#decl-f7ecfc638d2f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","order":8435,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate.lean:600"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitLogarithmicRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (forall rounds, selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate mdp…","missing":[],"search":"selfconsistentscheduledcausalsource_cumulative_and_averagebehaviorexpectedregret_explicitlogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_cumulative_and_averagebehaviorexpectedregret_explicitlogarithmicrate same-source explicit logarithmic cumulative and `log(n) / n` average route. theorem compiled","shard":"modules/2c1d7f021910a352.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_rounds_mul_two_horizon","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_rounds_mul_two_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_rounds_mul_two_horizon","description":"The cumulative behavior expected-regret process has its deterministic finite-prefix envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-1ba3b5b67772","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8436,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_rounds_mul_two_horizon (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory <= (rounds : Real) * (2 * (mdp.horizon : Real))","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_rounds_mul_two_horizon banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_rounds_mul_two_horizon the cumulative behavior expected-regret process has its deterministic finite-prefix envelope. theorem compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope","description":"Deterministic envelope for the second moment at one positive prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-93e078b9651e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8437,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretsecondmomentenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretsecondmomentenvelope deterministic envelope for the second moment at one positive prefix. definition compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope_nonneg","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope_nonneg","description":"The deterministic coordinate second-moment envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-4b23dba8f66f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8438,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope mdp varianceProxy baseVisitFloor rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretsecondmomentenvelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretsecondmomentenvelope_nonneg the deterministic coordinate second-moment envelope is nonnegative. theorem compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_secondMomentEnvelope","label":"integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_secondMomentEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_secondMomentEnvelope","description":"At every positive deterministic prefix, the exact average realized behavior-regret second moment is bounded by the deterministic envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-bca3be2e9757","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8439,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_secondMomentEnvelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) (hrounds : 0 < rounds) : integral (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure (fun trajectory => selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaul…","missing":[],"search":"integral_sq_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_secondmomentenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_sq_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_secondmomentenvelope at every positive deterministic prefix, the exact average realized behavior-regret second moment is bounded by the deterministic envelope. theorem compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget","label":"selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget","description":"A finite deterministic second-moment budget covering every positive prefix up to `maxRounds`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-6d600c8361a3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8440,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:246"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingexplicitsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingexplicitsecondmomentbudget a finite deterministic second-moment budget covering every positive prefix up to `maxrounds`. definition compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget_nonneg","label":"selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget_nonneg","description":"The deterministic finite-prefix second-moment budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-2aee83103756","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8441,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:254"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget mdp varianceProxy baseVisitFloor maxRounds","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingexplicitsecondmomentbudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingexplicitsecondmomentbudget_nonneg the deterministic finite-prefix second-moment budget is nonnegative. theorem compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_le_explicitBudget","label":"selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_le_explicitBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_le_explicitBudget","description":"A positive bounded stopping time selects one coordinate from the finite prefix sum, so its exact second moment is bounded by the deterministic budget. No optional-stopping theorem is used.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-60328e984fad","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8442,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:267"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_le_explicitBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (htau : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProx…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment_le_explicitbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment_le_explicitbudget a positive bounded stopping time selects one coordinate from the finite prefix sum, so its exact second moment is bounded by the deterministic budget. no optional-stopping theorem is used. theorem compiled","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","description":"A positive bounded stopping time selects one coordinate from the finite prefix sum, so its exact second moment is bounded by the deterministic budget. No optional-stopping theorem is used. -/ theorem selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_le_explicitBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace Stat…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitdeterministic-320ae32c6b88/index.html#decl-363e8b9cc983","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","order":8443,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret.lean:369"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpiso…","missing":[],"search":"selfconsistentscheduledcausalsource_boundedstoppingtimeexplicitdeterministicmomentexpectedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_boundedstoppingtimeexplicitdeterministicmomentexpectedaveragerealizedbehaviorregret a positive bounded stopping time selects one coordinate from the finite prefix sum, so its exact second moment is bounded by the deterministic budget. no optional-stopping theorem is used. -/ theorem selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment_le_explicitbudget (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (basevisitfloor : real) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (tau : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (htau : isstoppingtime (selfconsistentschedulednaturalcausaltrajectoryfiltration mdp initialstate rewardsource initialtable defaultstate…","shard":"modules/ad54d36e4a57dcfe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","label":"memLp_two_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","description":"The cumulative behavior expected-regret process belongs to `L2` under its deterministic finite-prefix envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-c591e5d651ef","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8444,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_two_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : MemLp (selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) 2 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"memlp_two_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_two_selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess the cumulative behavior expected-regret process belongs to `l2` under its deterministic finite-prefix envelope. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","label":"memLp_two_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","description":"Every deterministic-prefix average realized behavior-regret coordinate is in `L2`: the behavior component is bounded and the centered return component is sub-Gaussian.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-4d822577ef90","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8445,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:84"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_two_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : MemLp (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds) 2 (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure","missing":[],"search":"memlp_two_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_two_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess every deterministic-prefix average realized behavior-regret coordinate is in `l2`: the behavior component is bounded and the centered return component is sub-gaussian. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","label":"memLp_two_selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","description":"Mathlib bounded-stopping transport gives `L2` for the exact stopped average realized behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-823483452c0d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8446,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_two_selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (htau : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) ta…","missing":[],"search":"memlp_two_selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_two_selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret mathlib bounded-stopping transport gives `l2` for the exact stopped average realized behavior regret. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_nonneg","label":"selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_nonneg","description":"Every deterministic logarithmic average-rate coordinate is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-a9faf2a22446","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8447,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:170"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : 0 <= selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate mdp varianceProxy baseVisitFloor rounds returnDelta","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate_nonneg every deterministic logarithmic average-rate coordinate is nonnegative. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget","label":"selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget","description":"Deterministic finite sum which dominates the logarithmic rate selected by any positive stopping time bounded by `maxRounds`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-63eb68218264","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8448,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:186"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingexplicitexpectedratebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingexplicitexpectedratebudget deterministic finite sum which dominates the logarithmic rate selected by any positive stopping time bounded by `maxrounds`. definition compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget_nonneg","label":"selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget_nonneg","description":"The finite expected-rate budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-20c7bf139afc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8449,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:196"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) : 0 <= selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget mdp varianceProxy baseVisitFloor maxRounds","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingexplicitexpectedratebudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingexplicitexpectedratebudget_nonneg the finite expected-rate budget is nonnegative. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_le_explicitExpectedRateBudget","label":"selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_le_explicitExpectedRateBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_le_explicitExpectedRateBudget","description":"The stopped logarithmic rate is charged to the finite positive-prefix rate budget without assuming endpoint monotonicity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-38669181a31d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8450,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:208"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_le_explicitExpectedRateBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (maxRounds : Nat) (htau_pos : forall trajectory, (1 : WithTop Nat) <= tau trajectory) (htau_le : forall trajectory, tau trajectory <= (maxRounds : WithTop Nat)) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloo…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate_le_explicitexpectedratebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate_le_explicitexpectedratebudget the stopped logarithmic rate is charged to the finite positive-prefix rate budget without assuming endpoint monotonicity. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment","label":"selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment","description":"Exact stopped second moment on the generated causal trajectory measure.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-a91ee6367357","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8451,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:246"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment exact stopped second moment on the generated causal trajectory measure. definition compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_nonneg","label":"selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_nonneg","description":"The exact stopped second moment is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-3df697b6284e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8452,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:266"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) : 0 <= selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor tau","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment_nonneg the exact stopped second moment is nonnegative. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","description":"The exact stopped second moment is nonnegative. -/ theorem selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitexpectedavera-7693804df910/index.html#decl-a2e40d31842e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","order":8453,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfC…","missing":[],"search":"selfconsistentscheduledcausalsource_boundedstoppingtimeexplicitexpectedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_boundedstoppingtimeexplicitexpectedaveragerealizedbehaviorregret the exact stopped second moment is nonnegative. -/ theorem selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment_nonneg (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (tau : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) : 0 <= selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor tau := by unfold selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregretsecondmoment exact integral_nonneg fun _ => sq_nonneg _ /- terminal expected-regret route. the bad-event contribution is charged through the exact stopped second moment; no optional-stopping theorem is used. theorem compiled","shard":"modules/0dccafeeffcf5dbe.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.sum_range_one_div_natCast_add_two_pow_le_one_div_sixteen","label":"sum_range_one_div_natCast_add_two_pow_le_one_div_sixteen","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.sum_range_one_div_natCast_add_two_pow_le_one_div_sixteen","description":"A shifted inverse-power prefix with exponent at least six is at most one sixteenth. This refines the existing inverse-square comparison.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequarterg-2d26031403a4/index.html#decl-1d62424cc18b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","order":8454,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret.lean:23"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sum_range_one_div_natCast_add_two_pow_le_one_div_sixteen (rounds exponent : Nat) (hexponent : 6 <= exponent) : (Finset.range rounds).sum (fun t => 1 / (((t + 2 : Nat) : Real) ^ exponent)) <= 1 / 16","missing":[],"search":"sum_range_one_div_natcast_add_two_pow_le_one_div_sixteen banditrlproof.sum_range_one_div_natcast_add_two_pow_le_one_div_sixteen a shifted inverse-power prefix with exponent at least six is at most one sixteenth. this refines the existing inverse-square comparison. theorem compiled","shard":"modules/058aae79cf868369.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_eq_ofReal_two_mul_sum","label":"selfConsistentScheduledCausalModelFailureBudget_eq_ofReal_two_mul_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_eq_ofReal_two_mul_sum","description":"The finite scheduled model budget is the ENNReal image of twice the real sum of its local confidence schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequarterg-2d26031403a4/index.html#decl-7e97be051fba","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","order":8455,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalModelFailureBudget_eq_ofReal_two_mul_sum (mdp : MDP State Action) (rounds : Nat) : selfConsistentScheduledCausalModelFailureBudget mdp rounds = ENNReal.ofReal (2 * (Finset.range rounds).sum fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp t)","missing":[],"search":"selfconsistentscheduledcausalmodelfailurebudget_eq_ofreal_two_mul_sum banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelfailurebudget_eq_ofreal_two_mul_sum the finite scheduled model budget is the ennreal image of twice the real sum of its local confidence schedule. theorem compiled","shard":"modules/058aae79cf868369.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_le_one_eighth","label":"selfConsistentScheduledCausalModelFailureBudget_le_one_eighth","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_le_one_eighth","description":"Every finite prefix of the existing self-consistent model schedule spends at most one eighth of probability mass across both model-confidence shares.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequarterg-2d26031403a4/index.html#decl-dfb4bbd3ebd7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","order":8456,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret.lean:156"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalModelFailureBudget_le_one_eighth (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (rounds : Nat) : selfConsistentScheduledCausalModelFailureBudget mdp rounds <= ENNReal.ofReal (1 / 8 : Real)","missing":[],"search":"selfconsistentscheduledcausalmodelfailurebudget_le_one_eighth banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelfailurebudget_le_one_eighth every finite prefix of the existing self-consistent model schedule spends at most one eighth of probability mass across both model-confidence shares. theorem compiled","shard":"modules/058aae79cf868369.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le_one_quarter","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le_one_quarter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le_one_quarter","description":"With return budget one eighth, the joint horizon-level model event and all positive-prefix return events have probability at most one quarter. Its measurable complement has real probability at least three quarters.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequarterg-2d26031403a4/index.html#decl-c4ee5024792d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","order":8457,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le_one_quarter (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (maxRounds : Nat) (hmaxRounds : 0 < maxRounds) : let source := selfConsistentScheduledCausalSource m…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingsinglemodelreturnbadevent_le_one_quarter banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingsinglemodelreturnbadevent_le_one_quarter with return budget one eighth, the joint horizon-level model event and all positive-prefix return events have probability at most one quarter. its measurable complement has real probability at least three quarters. theorem compiled","shard":"modules/058aae79cf868369.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","description":"With return budget one eighth, the joint horizon-level model event and all positive-prefix return events have probability at most one quarter. Its measurable complement has real probability at least three quarters. -/ theorem selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le_one_quarter (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initi…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimeexplicitthreequarterg-2d26031403a4/index.html#decl-685d7142c482","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","order":8458,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatc…","missing":[],"search":"selfconsistentscheduledcausalsource_boundedstoppingtimeexplicitthreequartergoodeventaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_boundedstoppingtimeexplicitthreequartergoodeventaveragerealizedbehaviorregret with return budget one eighth, the joint horizon-level model event and all positive-prefix return events have probability at most one quarter. its measurable complement has real probability at least three quarters. -/ theorem selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingsinglemodelreturnbadevent_le_one_quarter (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (maxrounds : nat) (hmaxrounds : 0 < maxrounds) : let source := selfconsis…","shard":"modules/058aae79cf868369.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.one_le_untopA_and_untopA_le_of_withTop_bounds","label":"one_le_untopA_and_untopA_le_of_withTop_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.one_le_untopA_and_untopA_le_of_withTop_bounds","description":"A positive `WithTop Nat` time bounded by a finite horizon has a positive finite `untopA` value in the same deterministic range.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-dc63e50573da","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8459,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_le_untopA_and_untopA_le_of_withTop_bounds {Omega : Type*} (tau : Omega -> WithTop Nat) (maxRounds : Nat) (htau_pos : forall omega, (1 : WithTop Nat) <= tau omega) (htau_le : forall omega, tau omega <= (maxRounds : WithTop Nat)) (omega : Omega) : 1 <= (tau omega).untopA /\\ (tau omega).untopA <= maxRounds","missing":[],"search":"one_le_untopa_and_untopa_le_of_withtop_bounds banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.one_le_untopa_and_untopa_le_of_withtop_bounds a positive `withtop nat` time bounded by a finite horizon has a positive finite `untopa` value in the same deterministic range. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_stronglyAdapted","label":"selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_stronglyAdapted","description":"The prefix-scheduled deterministic logarithmic rate is strongly adapted to the natural trajectory filtration.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-29521670a7af","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8460,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_stronglyAdapted (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) : StronglyAdapted (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (fun rounds (_trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) => selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate mdp varianceProxy baseVisitFloor rounds (returnDeltaAt rounds))","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate_stronglyadapted banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate_stronglyadapted the prefix-scheduled deterministic logarithmic rate is strongly adapted to the natural trajectory filtration. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","label":"selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","description":"Exact natural average realized behavior regret evaluated at one stopping time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-ba36057eb14c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8461,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret exact natural average realized behavior regret evaluated at one stopping time. definition compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret_apply","label":"selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => A…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-fcd2d6195f34","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8462,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor tau trajectory = sel…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret_apply theorem selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (tau : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) : selfconsistentschedulednaturalcausalstoppedaveragerealizedbehaviorregret mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor tau trajectory = selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor (tau trajectory).untopa trajectory theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate","label":"selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate","description":"The scheduled logarithmic average rate evaluated at the same stopping time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-7016751bd830","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8463,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:130"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate the scheduled logarithmic average rate evaluated at the same stopping time. definition compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_apply","label":"selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) (tau : HeterogeneousStochasticEpisode…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-3332647f8bdd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8464,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:156"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate mdp initialState rewardSource initialTable defaultState varianceProxy bas…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate_apply theorem selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (returndeltaat : nat -> real) (tau : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) : selfconsistentschedulednaturalcausalstoppedrealizedaveragelogarithmicrate mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor returndeltaat tau trajectory = selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate mdp varianceproxy basevisitfloor (tau trajectory).untopa (returndeltaat (tau trajectory).untopa) theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","label":"selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","description":"One-sided violation of the stopped logarithmic average-rate certificate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-606baa5cff3e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8465,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset one-sided violation of the stopped logarithmic average-rate certificate. definition compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","description":"A bounded stopping-time violation is measurable at the deterministic natural-filtration bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-f64fd0809492","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8466,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:205"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (htau : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) tau) (maxRounds : Nat) (htau_le : forall trajectory, tau trajectory <= (maxRounds : WithTop Nat)) : MeasurableSet[ selfConsistentSc…","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset a bounded stopping-time violation is measurable at the deterministic natural-filtration bound. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","label":"selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","description":"Finite union of all positive fixed-prefix average-regret violations through the deterministic stopping-time bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-3a7481481f33","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8467,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:244"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) (returnDeltaAt : Nat -> Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalpositiveprefixaveragerealizedbehaviorregretviolationwindow banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpositiveprefixaveragerealizedbehaviorregretviolationwindow finite union of all positive fixed-prefix average-regret violations through the deterministic stopping-time bound. definition compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","label":"measurableSet_selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","description":"The finite positive-prefix violation window is ambient measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-00d9f3b49641","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8468,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:262"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) (returnDeltaAt : Nat -> Real) : MeasurableSet (selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor maxRounds returnDeltaAt)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalpositiveprefixaveragerealizedbehaviorregretviolationwindow banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalpositiveprefixaveragerealizedbehaviorregretviolationwindow the finite positive-prefix violation window is ambient measurable. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_positivePrefixWindow","label":"selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_positivePrefixWindow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_positivePrefixWindow","description":"Every positive bounded stopped violation occurs at one fixed prefix in the finite window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-37e01733080b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8469,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:281"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_positivePrefixWindow (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (returnDeltaAt : Nat -> Real) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (maxRounds : Nat) (htau_pos : forall trajectory, (1 : WithTop Nat) <= tau trajectory) (htau_le : forall trajectory, tau trajectory <= (maxRounds : WithTop Nat)) : selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet mdp initialState rewardS…","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset_subset_positiveprefixwindow banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset_subset_positiveprefixwindow every positive bounded stopped violation occurs at one fixed prefix in the finite window. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_positivePrefixAverageRealizedBehaviorRegretViolationWindow_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_positivePrefixAverageRealizedBehaviorRegretViolationWindow_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_positivePrefixAverageRealizedBehaviorRegretViolationWindow_le","description":"The positive-prefix violation window has the exact finite sum of the per-prefix model and return failure budgets.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-cfc0960ff748","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8470,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:316"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_positivePrefixAverageRealizedBehaviorRegretViolationWindow_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (maxRounds : Nat) (returnDeltaAt : Nat -> Real) (hreturnDeltaAt : forall rounds, rounds ∈ Fins…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_positiveprefixaveragerealizedbehaviorregretviolationwindow_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_positiveprefixaveragerealizedbehaviorregretviolationwindow_le the positive-prefix violation window has the exact finite sum of the per-prefix model and return failure budgets. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_le","description":"The stopped violation inherits the same explicit finite-window budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-3bdcf32e88ac","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8471,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:390"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisode…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingtimeaveragerealizedbehaviorregretviolationset_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingtimeaveragerealizedbehaviorregretviolationset_le the stopped violation inherits the same explicit finite-window budget. theorem compiled","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_boundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","description":"The stopped violation inherits the same explicit finite-window budget. -/ theorem selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal)…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimehighprobabilityaverag-1c90c83f63f3/index.html#decl-4cc63e3163b6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","order":8472,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret.lean:446"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_boundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfCo…","missing":[],"search":"selfconsistentscheduledcausalsource_boundedstoppingtimehighprobabilityaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_boundedstoppingtimehighprobabilityaveragerealizedbehaviorregret the stopped violation inherits the same explicit finite-window budget. -/ theorem selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingtimeaveragerealizedbehaviorregretviolationset_le (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (tau : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (maxrounds : nat) (htau_pos : forall trajectory,…","shard":"modules/0f03556de4324f19.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent_mono","label":"finiteHorizonBadEvent_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent_mono","description":"A finite-horizon heterogeneous bad-event union is monotone in its outer horizon when the underlying coordinate event family is unchanged.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-f38b407f4ad2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8473,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem finiteHorizonBadEvent_mono {mdp : MDP State Action} {episodes : Nat -> Nat} {rounds maxRounds : Nat} {initialBad : Set (StochasticEpisodeBatch mdp (episodes 0))} {successorBad : (n : Nat) -> Set (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n × StochasticEpisodeBatch mdp (episodes (n + 1)))} (hrounds : rounds <= maxRounds) : finiteHorizonBadEvent rounds initialBad successorBad ⊆ finiteHorizonBadEvent maxRounds initialBad successorBad","missing":[],"search":"finitehorizonbadevent_mono banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.finitehorizonbadevent_mono a finite-horizon heterogeneous bad-event union is monotone in its outer horizon when the underlying coordinate event family is unchanged. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelBadEvent_mono","label":"selfConsistentScheduledCausalModelBadEvent_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelBadEvent_mono","description":"The actual self-consistent model-confidence event at a prefix is contained in the corresponding event at every larger deterministic horizon.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-cdd2d2c589e2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8474,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalModelBadEvent_mono (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) {rounds maxRounds : Nat} (hrounds : rounds <= maxRounds) : selfConsistentScheduledCausalModelBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds ⊆ selfConsistentScheduledCausalModelBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor maxRounds","missing":[],"search":"selfconsistentscheduledcausalmodelbadevent_mono banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalmodelbadevent_mono the actual self-consistent model-confidence event at a prefix is contained in the corresponding event at every larger deterministic horizon. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare","label":"selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare","description":"Equal allocation of one global return confidence budget over the possible positive stopped prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-a577878432fe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8475,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:74"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare (maxRounds : Nat) (returnDelta : Real) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingequalreturnshare banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingequalreturnshare equal allocation of one global return confidence budget over the possible positive stopped prefixes. definition compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixIndex_nonempty","label":"selfConsistentScheduledNaturalCausalPositivePrefixIndex_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixIndex_nonempty","description":"A positive deterministic horizon has at least one positive prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-db8e2d6c33e8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8476,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalPositivePrefixIndex_nonempty (maxRounds : Nat) (hmaxRounds : 0 < maxRounds) : (Finset.Icc 1 maxRounds).Nonempty","missing":[],"search":"selfconsistentschedulednaturalcausalpositiveprefixindex_nonempty banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpositiveprefixindex_nonempty a positive deterministic horizon has at least one positive prefix. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare_spec","label":"selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare_spec","description":"The equal return share is positive and at most one.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-9ccb3612a7bf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8477,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare_spec (maxRounds : Nat) (hmaxRounds : 0 < maxRounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDelta_le_one : returnDelta <= 1) : 0 < selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare maxRounds returnDelta ∧ selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare maxRounds returnDelta <= 1","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingequalreturnshare_spec banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingequalreturnshare_spec the equal return share is positive and at most one. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","label":"selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","description":"Finite window containing only the return-deviation events at the possible positive stopped prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-ba0857cd7cbe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8478,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingreturnbadeventwindow banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingreturnbadeventwindow finite window containing only the return-deviation events at the possible positive stopped prefixes. definition compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","label":"measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","description":"The finite return-only window is ambient measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-339f3f36f5c7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8479,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) (returnDelta : Real) : MeasurableSet (selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor maxRounds returnDelta)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalboundedstoppingreturnbadeventwindow banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalboundedstoppingreturnbadeventwindow the finite return-only window is ambient measurable. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingReturnBadEventWindow_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingReturnBadEventWindow_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingReturnBadEventWindow_le","description":"Equal allocation bounds the finite return-deviation window by the one global return confidence budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-2af7ff232b90","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8480,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingReturnBadEventWindow_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (maxRounds : Nat) (hmaxRounds : 0 < maxRounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDelta_le_one : returnDelta <= 1) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure (selfConsisten…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingreturnbadeventwindow_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingreturnbadeventwindow_le equal allocation bounds the finite return-deviation window by the one global return confidence budget. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingSingleModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalBoundedStoppingSingleModelReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingSingleModelReturnBadEvent","description":"One horizon model-confidence event joined with the finite return-only window for all possible positive stopped prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-c82d9b524485","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8481,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:208"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedStoppingSingleModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (maxRounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingsinglemodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingsinglemodelreturnbadevent one horizon model-confidence event joined with the finite return-only window for all possible positive stopped prefixes. definition compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow_subset_boundedStoppingSingleModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow_subset_boundedStoppingSingleModelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow_subset_boundedStoppingSingleModelReturnBadEvent","description":"Every positive fixed-prefix average-regret violation with the equal return share lies in the single horizon model event or the return-only window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-66a29a82b70d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8482,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow_subset_boundedStoppingSingleModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (maxRounds : Nat) (returnDelta : Real) : selfConsistentScheduledNat…","missing":[],"search":"selfconsistentschedulednaturalcausalpositiveprefixaveragerealizedbehaviorregretviolationwindow_subset_boundedstoppingsinglemodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpositiveprefixaveragerealizedbehaviorregretviolationwindow_subset_boundedstoppingsinglemodelreturnbadevent every positive fixed-prefix average-regret violation with the equal return share lies in the single horizon model event or the return-only window. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le","description":"The single-model return event is measurable and is charged by one horizon model budget plus the one global return budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-1b44c8d5bc10","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8483,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (maxRounds : Nat) (hmaxRounds : 0 < maxRounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDel…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingsinglemodelreturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_boundedstoppingsinglemodelreturnbadevent_le the single-model return event is measurable and is charged by one horizon model budget plus the one global return budget. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_singleModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_singleModelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_singleModelReturnBadEvent","description":"The stopped violation is contained in the sharper single-model return event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-43c3ba2212bb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8484,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:344"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_singleModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStoch…","missing":[],"search":"selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset_subset_singlemodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset_subset_singlemodelreturnbadevent the stopped violation is contained in the sharper single-model return event. theorem compiled","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_boundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","description":"The stopped violation is contained in the sharper single-model return event. -/ theorem selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_singleModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varia…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedstoppingtimesinglemodeleventhighp-1a0d3cf891a7/index.html#decl-63d54b7f14cf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","order":8485,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret.lean:392"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_boundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (tau : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBa…","missing":[],"search":"selfconsistentscheduledcausalsource_boundedstoppingtimesinglemodeleventhighprobabilityaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_boundedstoppingtimesinglemodeleventhighprobabilityaveragerealizedbehaviorregret the stopped violation is contained in the sharper single-model return event. -/ theorem selfconsistentschedulednaturalcausalboundedstoppingtimeaveragerealizedbehaviorregretviolationset_subset_singlemodelreturnbadevent (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (tau : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat)…","shard":"modules/a6eaf19b50e052d0.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_window_offset_untopA_eq_of_withTop_bounds","label":"exists_window_offset_untopA_eq_of_withTop_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_window_offset_untopA_eq_of_withTop_bounds","description":"A finite `WithTop Nat` window selects one explicit natural offset.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-53db3135731e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8486,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exists_window_offset_untopA_eq_of_withTop_bounds {Omega : Type*} (tau : Omega -> WithTop Nat) (start window : Nat) (htau_lower : forall omega, (start : WithTop Nat) <= tau omega) (htau_upper : forall omega, tau omega <= ((start + window : Nat) : WithTop Nat)) (omega : Omega) : exists offset, offset ∈ Finset.range (window + 1) /\\ (tau omega).untopA = start + offset","missing":[],"search":"exists_window_offset_untopa_eq_of_withtop_bounds banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exists_window_offset_untopa_eq_of_withtop_bounds a finite `withtop nat` window selects one explicit natural offset. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget","label":"selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget","description":"Sum of the coordinate L1 envelopes over one fixed-width shifted window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-a8e78aa912a9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8487,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (window scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalboundedwindowstoppingl1budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedwindowstoppingl1budget sum of the coordinate l1 envelopes over one fixed-width shifted window. definition compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_nonneg","label":"selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_nonneg","description":"The fixed-window L1 budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-de812204011e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8488,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:70"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (window scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget mdp varianceProxy baseVisitFloor window scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalboundedwindowstoppingl1budget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedwindowstoppingl1budget_nonneg the fixed-window l1 budget is nonnegative. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_tendsto_zero","label":"selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_tendsto_zero","description":"Every fixed-width shifted finite sum of coordinate L1 envelopes vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-083ae486434c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8489,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (window : Nat) : Tendsto (selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget mdp varianceProxy baseVisitFloor window) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalboundedwindowstoppingl1budget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalboundedwindowstoppingl1budget_tendsto_zero every fixed-width shifted finite sum of coordinate l1 envelopes vanishes. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret","description":"Expected absolute value of the fixed-window stopped average process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-0d116862a20e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8490,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret expected absolute value of the fixed-window stopped average process. definition compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret","description":"Every pointwise bounded stopped average process belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-da25d95e9308","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8491,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:128"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret every pointwise bounded stopped average process belongs to `l1`. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_nonneg","description":"Expected absolute stopped regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-fe841fb5a4ab","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8492,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:166"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret_nonneg expected absolute stopped regret is nonnegative. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_le_budget","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_le_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_le_budget","description":"The selected stopped coordinate is bounded by the shifted finite L1 budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-b32014dd3e5d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8493,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_le_budget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStoc…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret_le_budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret_le_budget the selected stopped coordinate is bounded by the shifted finite l1 budget. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","description":"Expected absolute stopped average regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-e981f26576d1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8494,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:306"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveS…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteboundedwindowstoppingaveragerealizedbehaviorregret_tendsto_zero expected absolute stopped average regret tends to zero. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_eq","description":"At exponent one, the stopped norm is its lifted expected absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-e40aeae60db0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8495,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:357"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardS…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret_eq at exponent one, the stopped norm is its lifted expected absolute value. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","description":"The stopped process converges to zero in the exponent-one extended norm.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-6983387b731b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8496,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:402"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => Adap…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero the stopped process converges to zero in the exponent-one extended norm. theorem compiled","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","label":"selfConsistentScheduledCausalSource_boundedWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","description":"The stopped process converges to zero in the exponent-one extended norm. -/ theorem eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NN…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalboundedwindowstoppingtimel1averagerealiz-38a651046106/index.html#decl-81575a834af2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8497,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:489"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_boundedWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStoc…","missing":[],"search":"selfconsistentscheduledcausalsource_boundedwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_boundedwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency the stopped process converges to zero in the exponent-one extended norm. -/ theorem elpnorm_one_selfconsistentschedulednaturalcausalboundedwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat)…","shard":"modules/39e5c563a3541cad.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat","description":"The first natural prefix in the double-linear raw window whose observed average regret is at most the deterministic threshold, with right-end fallback.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-19451714214b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8498,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Nat","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat the first natural prefix in the double-linear raw window whose observed average regret is at most the deterministic threshold, with right-end fallback. definition compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","description":"The `WithTop Nat` surface consumed by Mathlib stopped-process APIs.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-b6999f054a37","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8499,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix the `withtop nat` surface consumed by mathlib stopped-process apis. definition compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_of_first_hit","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_of_first_hit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_of_first_hit","description":"A specified first threshold hit is exactly the Mathlib finite hitting time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-d8015889affc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8500,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_of_first_hit (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex candidate : Nat) (trajectory) (hcandidateLower : explicitHighProbabilityRounds scheduleIndex <= candidate) (hcandidateUpper : candidate <= explicitHighProbabilityRounds scheduleIndex + (2 * scheduleIndex + 1)) (hhit : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor candidate trajectory <= threshold scheduleIndex) (hfirst : forall earlier, explicitHighProbabilityRounds sche…","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_eq_of_first_hit banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_eq_of_first_hit a specified first threshold hit is exactly the mathlib finite hitting time. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_base_of_le","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_base_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_base_of_le","description":"If the base observation crosses the threshold, first passage is immediate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-3737967bde28","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8501,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_base_of_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) (hthreshold : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (explicitHighProbabilityRounds scheduleIndex) trajectory <= threshold scheduleIndex) : selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory = e…","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_eq_base_of_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_eq_base_of_le if the base observation crosses the threshold, first passage is immediate. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_right_of_forall_lt","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_right_of_forall_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_right_of_forall_lt","description":"Without an earlier crossing, the capped first-passage rule returns the right endpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-10a8cde47ccf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8502,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:164"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_right_of_forall_lt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) (hbefore : forall earlier, explicitHighProbabilityRounds scheduleIndex <= earlier -> earlier < explicitHighProbabilityRounds scheduleIndex + (2 * scheduleIndex + 1) -> threshold scheduleIndex < selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor earlier trajectory) : selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat…","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_eq_right_of_forall_lt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_eq_right_of_forall_lt without an earlier crossing, the capped first-passage rule returns the right endpoint. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_before_gt","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_before_gt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_before_gt","description":"Every strict pre-stopping prefix in the window is above the threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-14a9c017ef1c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8503,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_before_gt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex earlier : Nat) (trajectory) (hlower : explicitHighProbabilityRounds scheduleIndex <= earlier) (hearlier : earlier < selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory) : threshold scheduleIndex < selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProx…","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_before_gt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_before_gt every strict pre-stopping prefix in the window is above the threshold. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_le_threshold_of_lt_right","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_le_threshold_of_lt_right","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_le_threshold_of_lt_right","description":"A strict stop before the cap is an actual threshold hit.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-3c2f7b0ae5ef","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8504,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:233"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_le_threshold_of_lt_right (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) (hstopping : selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory < explicitHighProbabilityRounds scheduleIndex + (2 * scheduleIndex + 1)) : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (self…","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_le_threshold_of_lt_right banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassageprefixnat_le_threshold_of_lt_right a strict stop before the cap is an actual threshold hit. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_isStoppingTime","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_isStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_isStoppingTime","description":"The capped first-passage rule is an exact natural-filtration stopping time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-ac8e573fd6a2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8505,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_isStoppingTime (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex)","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_isstoppingtime banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_isstoppingtime the capped first-passage rule is an exact natural-filtration stopping time. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_lower","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_lower","description":"The capped first-passage rule never stops before its fourth-power base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-38cef162b7fc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8506,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_lower (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) : (explicitHighProbabilityRounds scheduleIndex : WithTop Nat) <= selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_lower banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_lower the capped first-passage rule never stops before its fourth-power base. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_upper","label":"selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_upper","description":"The capped first-passage rule never exceeds its double-linear endpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-30836058b881","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8507,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_upper (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory <= (explicitHighProbabilityRounds scheduleIndex + (2 * scheduleIndex + 1) : WithTop Nat)","missing":[],"search":"selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_upper banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_upper the capped first-passage rule never exceeds its double-linear endpoint. theorem compiled","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cappedDoubleLinearRawWindowFirstPassageStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","label":"selfConsistentScheduledCausalSource_cappedDoubleLinearRawWindowFirstPassageStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cappedDoubleLinearRawWindowFirstPassageStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","description":"The capped first-passage rule never exceeds its double-linear endpoint. -/ theorem selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_upper (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal)…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalcappeddoublelinearrawwindowfirstpassages-782d55d77dac/index.html#decl-d2e88635399f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8508,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:351"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_cappedDoubleLinearRawWindowFirstPassageStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (threshold : Nat -> Real) : let stoppingPrefix := selfConsistentSchedul…","missing":[],"search":"selfconsistentscheduledcausalsource_cappeddoublelinearrawwindowfirstpassagestoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_cappeddoublelinearrawwindowfirstpassagestoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency the capped first-passage rule never exceeds its double-linear endpoint. -/ theorem selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix_upper (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (threshold : nat -> real) (scheduleindex : nat) (trajectory) : selfconsistentschedulednaturalcausalcappeddoublelinearrawwindowfirstpassagestoppingprefix mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor threshold scheduleindex trajectory <= (explicithighprobabilityrounds scheduleindex + (2 * scheduleindex + 1) : withtop nat) := by have hbaseright : explicithighprobabilityrounds scheduleindex <= explicithighprobabilityrounds scheduleindex + (2 * scheduleindex + 1) := nat.le_add_right _ _ have hbound := (measuretheory.hittingbtwn_mem_icc (u := selfconsistent…","shard":"modules/d01efa4903b1d7de.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_mono","label":"explicitHighProbabilityRounds_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_mono","description":"The explicit fourth-power prefix grid is monotone in its index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-35c8742ff302","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8509,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_mono {left right : Nat} (hle : left <= right) : explicitHighProbabilityRounds left <= explicitHighProbabilityRounds right","missing":[],"search":"explicithighprobabilityrounds_mono banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_mono the explicit fourth-power prefix grid is monotone in its index. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.scheduleIndex_le_explicitHighProbabilityRounds_add","label":"scheduleIndex_le_explicitHighProbabilityRounds_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.scheduleIndex_le_explicitHighProbabilityRounds_add","description":"Every explicit fourth-power prefix at index `n + offset` dominates `n`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-fb1b709b21ea","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8510,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem scheduleIndex_le_explicitHighProbabilityRounds_add (scheduleIndex offset : Nat) : scheduleIndex <= explicitHighProbabilityRounds (scheduleIndex + offset)","missing":[],"search":"scheduleindex_le_explicithighprobabilityrounds_add banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.scheduleindex_le_explicithighprobabilityrounds_add every explicit fourth-power prefix at index `n + offset` dominates `n`. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_growingWindowGrid_offset_untopA_eq","label":"exists_growingWindowGrid_offset_untopA_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_growingWindowGrid_offset_untopA_eq","description":"A grid-valued stopping prefix selects one finite window offset after `untopA`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-58dd8a1c1b38","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8511,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exists_growingWindowGrid_offset_untopA_eq {Omega : Type*} (stoppingPrefix : Nat -> Omega -> WithTop Nat) (windowAt : Nat -> Nat) (hgrid : forall scheduleIndex trajectory, exists offset, offset ∈ Finset.range (windowAt scheduleIndex + 1) /\\ stoppingPrefix scheduleIndex trajectory = (explicitHighProbabilityRounds (scheduleIndex + offset) : WithTop Nat)) (scheduleIndex : Nat) (trajectory : Omega) : exists offset, offset ∈ Finset.range (windowAt scheduleIndex + 1) /\\ (stoppingPrefix scheduleIndex trajectory).untopA = explicitHighProbabilityRounds (scheduleIndex + offset)","missing":[],"search":"exists_growingwindowgrid_offset_untopa_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exists_growingwindowgrid_offset_untopa_eq a grid-valued stopping prefix selects one finite window offset after `untopa`. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.growingWindowGrid_stoppingPrefix_le","label":"growingWindowGrid_stoppingPrefix_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.growingWindowGrid_stoppingPrefix_le","description":"Every grid-valued stopping prefix has a pointwise finite upper bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-01ee18d970af","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8512,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem growingWindowGrid_stoppingPrefix_le {Omega : Type*} (stoppingPrefix : Nat -> Omega -> WithTop Nat) (windowAt : Nat -> Nat) (hgrid : forall scheduleIndex trajectory, exists offset, offset ∈ Finset.range (windowAt scheduleIndex + 1) /\\ stoppingPrefix scheduleIndex trajectory = (explicitHighProbabilityRounds (scheduleIndex + offset) : WithTop Nat)) (scheduleIndex : Nat) (trajectory : Omega) : stoppingPrefix scheduleIndex trajectory <= (explicitHighProbabilityRounds (scheduleIndex + windowAt scheduleIndex) : WithTop Nat)","missing":[],"search":"growingwindowgrid_stoppingprefix_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.growingwindowgrid_stoppingprefix_le every grid-valued stopping prefix has a pointwise finite upper bound. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget","label":"selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget","description":"Finite L1 budget for a window of fourth-power grid candidates.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-3542bb982909","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8513,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:90"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget (mdp : MDP State Action) (varianceProxy : NNReal) (windowAt : Nat -> Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget finite l1 budget for a window of fourth-power grid candidates. definition compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail","label":"selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail","description":"Infinite shifted tail controlling every finite grid window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-5b60a44c11e5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8514,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1tail banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1tail infinite shifted tail controlling every finite grid window. definition compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_nonneg","label":"selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_nonneg","description":"Every finite growing-window grid budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-090c33f87efe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8515,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (windowAt : Nat -> Nat) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget mdp varianceProxy windowAt scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget_nonneg every finite growing-window grid budget is nonnegative. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_le_tail","label":"selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_le_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_le_tail","description":"Every finite grid-window budget is bounded by the full shifted tail.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-8a89c3eaa879","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8516,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_le_tail (mdp : MDP State Action) (varianceProxy : NNReal) (windowAt : Nat -> Nat) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget mdp varianceProxy windowAt scheduleIndex <= selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail mdp varianceProxy scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget_le_tail banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget_le_tail every finite grid-window budget is bounded by the full shifted tail. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail_tendsto_zero","label":"selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail_tendsto_zero","description":"The infinite shifted grid-envelope tail tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-44b468ae77d8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8517,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:151"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail mdp varianceProxy) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1tail_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1tail_tendsto_zero the infinite shifted grid-envelope tail tends to zero. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_tendsto_zero","label":"selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_tendsto_zero","description":"Any finite grid-window budget vanishes, even when its width grows.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-004094949661","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8518,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (windowAt : Nat -> Nat) : Tendsto (selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget mdp varianceProxy windowAt) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalgrowingwindowgridstoppingl1budget_tendsto_zero any finite grid-window budget vanishes, even when its width grows. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret","description":"Expected absolute value of the growing-window grid-stopped process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-915f67ca788c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8519,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret expected absolute value of the growing-window grid-stopped process. definition compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret","description":"Every grid-window stopped coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-fd36d6295559","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8520,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:206"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSo…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret every grid-window stopped coordinate belongs to `l1`. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_nonneg","description":"Expected absolute growing-window grid-stopped regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-0ea679a79fd8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8521,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:248"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret_nonneg expected absolute growing-window grid-stopped regret is nonnegative. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_le_budget","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_le_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_le_budget","description":"The selected grid coordinate is bounded by the finite candidate budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-1286215a4442","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8522,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_le_budget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => Adaptive…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret_le_budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret_le_budget the selected grid coordinate is bounded by the finite candidate budget. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_tendsto_zero","description":"Expected absolute growing-window grid-stopped regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-44cc98f5a72c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8523,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:410"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => Adapt…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutegrowingwindowgridstoppingaveragerealizedbehaviorregret_tendsto_zero expected absolute growing-window grid-stopped regret tends to zero. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_eq","description":"At exponent one, the stopped norm is its lifted expected absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-3601b593b6d4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8524,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:458"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rew…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret_eq at exponent one, the stopped norm is its lifted expected absolute value. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","description":"The grid-stopped process converges to zero in the exponent-one norm.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-773649316ef6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8525,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:505"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t =>…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero the grid-stopped process converges to zero in the exponent-one norm. theorem compiled","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_growingWindowGridStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","label":"selfConsistentScheduledCausalSource_growingWindowGridStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_growingWindowGridStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","description":"The grid-stopped process converges to zero in the exponent-one norm. -/ theorem eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NN…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalgrowingwindowgridstoppingtimel1averagere-d9988871a3d1/index.html#decl-02679e18a79d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8526,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:588"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_growingWindowGridStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => Adaptive…","missing":[],"search":"selfconsistentscheduledcausalsource_growingwindowgridstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_growingwindowgridstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency the grid-stopped process converges to zero in the exponent-one norm. -/ theorem elpnorm_one_selfconsistentschedulednaturalcausalgrowingwindowgridstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> witht…","shard":"modules/5d2fc239ecf72c08.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold","label":"selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold","description":"Inverse-square-root threshold `1/sqrt(n+1)` for first passage.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-d37041bbbe0b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8527,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtfirstpassagethreshold banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtfirstpassagethreshold inverse-square-root threshold `1/sqrt(n+1)` for first passage. definition compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_pos","label":"selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_pos","description":"Every inverse-square-root first-passage threshold is positive.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-2ebb03ffca22","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8528,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_pos (scheduleIndex : Nat) : 0 < selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtfirstpassagethreshold_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtfirstpassagethreshold_pos every inverse-square-root first-passage threshold is positive. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_tendsto_zero","description":"The inverse-square-root first-passage threshold tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-8baa18d3c39a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8529,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_tendsto_zero : Tendsto selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtfirstpassagethreshold_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtfirstpassagethreshold_tendsto_zero the inverse-square-root first-passage threshold tends to zero. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","description":"Capped first passage at the inverse-square-root threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-1442151f6b72","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8530,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix capped first passage at the inverse-square-root threshold. definition compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","description":"Event that inverse-square-root first passage advances beyond its base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-00a71df465f4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8531,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset event that inverse-square-root first passage advances beyond its base. definition compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","description":"Probability that inverse-square-root first passage advances past its base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-a6f14621eacd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8532,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : ENNReal","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability probability that inverse-square-root first passage advances past its base. definition compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","description":"Markov rate obtained by dividing the scheduled L1 envelope by the inverse-square-root threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-44e141316f2e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8533,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate markov rate obtained by dividing the scheduled l1 envelope by the inverse-square-root threshold. definition compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","description":"Delay is exactly strict one-sided threshold violation at the base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-541b38871211","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8534,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex = {trajectory | selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex < explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory}","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_eq delay is exactly strict one-sided threshold violation at the base. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","label":"measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","description":"The inverse-square-root delay event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-151c3852edac","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8535,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:215"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : MeasurableSet (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset the inverse-square-root delay event is measurable. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","description":"Delayed inverse-square-root first passage is contained in the scheduled distance violation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-b930e6c2c233","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8536,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:236"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex ⊆ explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex) scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_subset_distanceviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_subset_distanceviolationset delayed inverse-square-root first passage is contained in the scheduled distance violation. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_div_pow_three_eq_inverse_rpow_five_halves","label":"sqrt_div_pow_three_eq_inverse_rpow_five_halves","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_div_pow_three_eq_inverse_rpow_five_halves","description":"private lemma sqrt_div_pow_three_eq_inverse_rpow_five_halves (s : Real) (hs : 0 < s) : Real.sqrt s / s ^ 3 = 1 / s ^ (5 / 2 : Real)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-250f58ddb71d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8537,"meta":[["Kind","lemma"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private lemma sqrt_div_pow_three_eq_inverse_rpow_five_halves (s : Real) (hs : 0 < s) : Real.sqrt s / s ^ 3 = 1 / s ^ (5 / 2 : Real)","missing":[],"search":"sqrt_div_pow_three_eq_inverse_rpow_five_halves banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.sqrt_div_pow_three_eq_inverse_rpow_five_halves private lemma sqrt_div_pow_three_eq_inverse_rpow_five_halves (s : real) (hs : 0 < s) : real.sqrt s / s ^ 3 = 1 / s ^ (5 / 2 : real) lemma compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_div_pow_two_eq_inverse_rpow_three_halves","label":"sqrt_div_pow_two_eq_inverse_rpow_three_halves","kind":"lemma","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_div_pow_two_eq_inverse_rpow_three_halves","description":"private lemma sqrt_div_pow_two_eq_inverse_rpow_three_halves (s : Real) (hs : 0 < s) : Real.sqrt s / s ^ 2 = 1 / s ^ (3 / 2 : Real)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-c125c27bf184","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8538,"meta":[["Kind","lemma"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:279"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"private lemma sqrt_div_pow_two_eq_inverse_rpow_three_halves (s : Real) (hs : 0 < s) : Real.sqrt s / s ^ 2 = 1 / s ^ (3 / 2 : Real)","missing":[],"search":"sqrt_div_pow_two_eq_inverse_rpow_three_halves banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.sqrt_div_pow_two_eq_inverse_rpow_three_halves private lemma sqrt_div_pow_two_eq_inverse_rpow_three_halves (s : real) (hs : 0 < s) : real.sqrt s / s ^ 2 = 1 / s ^ (3 / 2 : real) lemma compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","description":"The inverse-square-root delay rate is an inverse-`5/2` behavior term plus an inverse-`3/2` return term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-ef4cbf15f617","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8539,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:292"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate mdp varianceProxy scheduleIndex = 4 * selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient mdp / (explicitHighProbabilityScale scheduleIndex : Real) ^ (5 / 2 : Real) + (2 * Real.sqrt (mdp.globalReturnDeviationPerEpisodeVarianceProxy 1 varianceProxy : Real) * Real.exp (1 / 2 : Real)) / (explicitHighProbabilityScale scheduleIndex : Real) ^ (3 / 2 : Real)","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_eq the inverse-square-root delay rate is an inverse-`5/2` behavior term plus an inverse-`3/2` return term. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_inverseSqrtFirstPassageThreshold_le_delayRate","label":"explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_inverseSqrtFirstPassageThreshold_le_delayRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_inverseSqrtFirstPassageThreshold_le_delayRate","description":"The expected absolute base process divided by the inverse-square-root threshold is bounded by the explicit delay rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-0896ef8f45be","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8540,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:341"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_inverseSqrtFirstPassageThreshold_le_delayRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorReg…","missing":[],"search":"explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_div_inversesqrtfirstpassagethreshold_le_delayrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_div_inversesqrtfirstpassagethreshold_le_delayrate the expected absolute base process divided by the inverse-square-root threshold is bounded by the explicit delay rate. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","label":"summable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","description":"The inverse-square-root delay rate is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-38f829128414","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8541,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:378"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate (mdp : MDP State Action) (varianceProxy : NNReal) : Summable (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate mdp varianceProxy)","missing":[],"search":"summable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate the inverse-square-root delay rate is summable. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","description":"Summability implies that the inverse-square-root delay rate tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-596b49d8073a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8542,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:437"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate mdp varianceProxy) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_tendsto_zero summability implies that the inverse-square-root delay rate tends to zero. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","description":"Markov control of inverse-square-root first-passage delay.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-032ba2f619ea","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8543,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:447"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDo…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_le_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_le_rate markov control of inverse-square-root first-passage delay. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_ne_top","label":"tsum_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_ne_top","description":"Inverse-square-root delay probabilities have finite total ENNReal mass.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-4865d33dfb9b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8544,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:519"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_ne_top (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (∑' scheduleIndex, selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedD…","missing":[],"search":"tsum_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_ne_top inverse-square-root delay probabilities have finite total ennreal mass. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","label":"ae_eventually_not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","description":"Almost every trajectory eventually avoids every inverse-square-root delay event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-93515a41c6ee","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8545,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:550"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource…","missing":[],"search":"ae_eventually_not_mem_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_not_mem_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedset almost every trajectory eventually avoids every inverse-square-root delay event. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_eq_base","label":"ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_eq_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_eq_base","description":"Almost surely, the inverse-square-root first-passage rule eventually stops exactly at the fourth-power base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-a39310e7841f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8546,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:585"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_eq_base (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSou…","missing":[],"search":"ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix_eq_base banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix_eq_base almost surely, the inverse-square-root first-passage rule eventually stops exactly at the fourth-power base. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","description":"The summably bounded inverse-square-root delay probability tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-525b067f91df","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8547,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:633"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLine…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_tendsto_zero the summably bounded inverse-square-root delay probability tends to zero. theorem compiled","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassage_summableDelay_eventuallyImmediateStopping_and_L1_consistency","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassage_summableDelay_eventuallyImmediateStopping_and_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassage_summableDelay_eventuallyImmediateStopping_and_L1_consistency","description":"The summably bounded inverse-square-root delay probability tends to zero. -/ theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (variancePro…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappeddoublelinearra-3b95c6b754eb/index.html#decl-bcd92a75adc2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","order":8548,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency.lean:667"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassage_summableDelay_eventuallyImmediateStopping_and_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let threshold := selfConsistentScheduledNaturalCaus…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappeddoublelinearrawwindowfirstpassage_summabledelay_eventuallyimmediatestopping_and_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappeddoublelinearrawwindowfirstpassage_summabledelay_eventuallyimmediatestopping_and_l1_consistency the summably bounded inverse-square-root delay probability tends to zero. -/ theorem selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagedelay…","shard":"modules/141cb8fa34505117.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_tendsto_zero_of_memLp_one_of_eLpNorm_tendsto_zero","label":"integral_tendsto_zero_of_memLp_one_of_eLpNorm_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_tendsto_zero_of_memLp_one_of_eLpNorm_tendsto_zero","description":"Exponent-one norm convergence to zero implies convergence of signed Bochner integrals to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html#decl-abb04f1fd99f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","order":8549,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_tendsto_zero_of_memLp_one_of_eLpNorm_tendsto_zero {Omega : Type w} [MeasurableSpace Omega] {mu : Measure Omega} {f : Nat -> Omega -> Real} (hf : forall n, MemLp (f n) 1 mu) (hnorm : Tendsto (fun n => eLpNorm (f n) 1 mu) atTop (nhds 0)) : Tendsto (fun n => integral mu (f n)) atTop (nhds 0)","missing":[],"search":"integral_tendsto_zero_of_memlp_one_of_elpnorm_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_tendsto_zero_of_memlp_one_of_elpnorm_tendsto_zero exponent-one norm convergence to zero implies convergence of signed bochner integrals to zero. theorem compiled","shard":"modules/8ced8c9611756ef1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegralDifference_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegralDifference_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegralDifference_tendsto_zero","description":"The signed integral of the exact uncapped-minus-capped truncation error tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html#decl-f0b587ccb79b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","order":8550,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegralDifference_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp ini…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretintegraldifference_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretintegraldifference_tendsto_zero the signed integral of the exact uncapped-minus-capped truncation error tends to zero. theorem compiled","shard":"modules/8ced8c9611756ef1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGap_tendsto_zero","description":"The signed difference between the uncapped and capped stopped-process expectations tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html#decl-03eaee6d00ea","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","order":8551,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialSta…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedgap_tendsto_zero the signed difference between the uncapped and capped stopped-process expectations tends to zero. theorem compiled","shard":"modules/8ced8c9611756ef1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGapAbs_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGapAbs_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGapAbs_tendsto_zero","description":"The absolute gap between the uncapped and capped stopped-process expectations tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html#decl-9c22b8401e63","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","order":8552,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean:233"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGapAbs_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initial…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedgapabs_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedgapabs_tendsto_zero the absolute gap between the uncapped and capped stopped-process expectations tends to zero. theorem compiled","shard":"modules/8ced8c9611756ef1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","description":"The capped stopped-process signed expectation tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html#decl-cce006731a5b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","order":8553,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean:280"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialT…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_tendsto_zero the capped stopped-process signed expectation tends to zero. theorem compiled","shard":"modules/8ced8c9611756ef1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_expected_truncation_replacement","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_expected_truncation_replacement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_expected_truncation_replacement","description":"The capped stopped-process signed expectation tends to zero. -/ theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvaria…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-ce44cae20bc0/index.html#decl-39f8f338833b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","order":8554,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement.lean:343"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_expected_truncation_replacement (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp i…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedaveragerealizedbehaviorregret_expected_truncation_replacement banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedaveragerealizedbehaviorregret_expected_truncation_replacement the capped stopped-process signed expectation tends to zero. -/ theorem selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : let source := selfconsistentscheduledcausalsource mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor le…","shard":"modules/8ced8c9611756ef1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppingPrefixes_eq_base_of_not_mem_delayedSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppingPrefixes_eq_base_of_not_mem_delayedSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppingPrefixes_eq_base_of_not_mem_delayedSet","description":"Outside the capped delayed set, the capped and uncapped inverse-square-root first-passage rules both stop at the common deterministic base, so their stopped regret coordinates agree pointwise.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-97d79a22a164","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8555,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppingPrefixes_eq_base_of_not_mem_delayedSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : let cappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let uncappedStoppingPrefix := selfConsistentScheduledNaturalCausalInve…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppingprefixes_eq_base_of_not_mem_delayedset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppingprefixes_eq_base_of_not_mem_delayedset outside the capped delayed set, the capped and uncapped inverse-square-root first-passage rules both stop at the common deterministic base, so their stopped regret coordinates agree pointwise. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_eq_unboundedHittingAfter","label":"ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_eq_unboundedHittingAfter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_eq_unboundedHittingAfter","description":"Almost every trajectory eventually sees exact equality between the capped and uncapped stopped processes. This follows from summable avoidance of the capped delayed set and the pointwise support theorem above.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-ef191885c36c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8556,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:120"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_eq_unboundedHittingAfter (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rew…","missing":[],"search":"ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregret_eq_unboundedhittingafter banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregret_eq_unboundedhittingafter almost every trajectory eventually sees exact equality between the capped and uncapped stopped processes. this follows from summable avoidance of the capped delayed set and the pointwise support theorem above. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret","description":"Every coordinate of the exact capped inverse-square-root stopped process belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-13406b0f9a06","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8557,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:172"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSour…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregret every coordinate of the exact capped inverse-square-root stopped process belongs to `l1`. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_tendsto_zero","description":"The exponent-one norm of the exact capped inverse-square-root stopped process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-984135d94448","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8558,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:213"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource init…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregret_tendsto_zero the exponent-one norm of the exact capped inverse-square-root stopped process tends to zero. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference","description":"Each uncapped-minus-capped stopped-process coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-527ce73ae634","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8559,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:274"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSou…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference each uncapped-minus-capped stopped-process coordinate belongs to `l1`. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendsto_zero","description":"The `L1` norm of the uncapped-minus-capped truncation error tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-782f28850e71","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8560,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:322"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference_tendsto_zero the `l1` norm of the uncapped-minus-capped truncation error tends to zero. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp","description":"The uncapped-minus-capped stopped truncation error as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-595a2adc5a96","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8561,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:439"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : Lp Real 1 (selfConsistentScheduledCausalSour…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifferencelp banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifferencelp the uncapped-minus-capped stopped truncation error as an `lp real 1` value. definition compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp_tendsto_zero","description":"The named capped/uncapped truncation error converges to zero in `Lp Real 1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-c963bd3723f9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8562,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:481"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThresho…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifferencelp_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifferencelp_tendsto_zero the named capped/uncapped truncation error converges to zero in `lp real 1`. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendstoInMeasure_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendstoInMeasure_zero","description":"The raw truncation-error representatives converge to zero in measure.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-5b792e21bf2d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8563,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:564"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp in…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference_tendstoinmeasure_zero the raw truncation-error representatives converge to zero in measure. theorem compiled","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_L1_truncation_equivalence","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_L1_truncation_equivalence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_L1_truncation_equivalence","description":"The raw truncation-error representatives converge to zero in measure. -/ theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewar…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-1c417b09b148/index.html#decl-b352491f78b6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","order":8564,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence.lean:627"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_L1_truncation_equivalence (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initial…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedaveragerealizedbehaviorregret_l1_truncation_equivalence banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedaveragerealizedbehaviorregret_l1_truncation_equivalence the raw truncation-error representatives converge to zero in measure. -/ theorem selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedaveragerealizedbehaviorregretdifference_tendstoinmeasure_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : let source := selfconsistentscheduledcausalsource mdp initialstate rewardsource initialtable defaultstate va…","shard":"modules/eb8cdeca410f3171.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageSampledReturn","label":"naturalSuccessorBatchAverageSampledReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageSampledReturn","description":"Observed sample mean in successor batch `t + 1`. Lean's division on `Real` is total, so a zero-size successor batch gives zero. The self-consistent scheduled source used below has positive batch sizes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-3005e5382272","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8565,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalSuccessorBatchAverageSampledReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (_source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (t : Nat) : Real","missing":[],"search":"naturalsuccessorbatchaveragesampledreturn banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessorbatchaveragesampledreturn observed sample mean in successor batch `t + 1`. lean's division on `real` is total, so a zero-size successor batch gives zero. the self-consistent scheduled source used below has positive batch sizes. definition compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn","label":"naturalAverageSampledReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn","description":"Average of observed successor-batch sample means over a natural prefix. At the empty prefix it uses the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-705cadd51e29","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8566,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAverageSampledReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"naturalaveragesampledreturn banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragesampledreturn average of observed successor-batch sample means over a natural prefix. at the empty prefix it uses the optimal initial expected return. definition compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn_eq_optimal_sub_naturalAverageRealizedBehaviorRegret","label":"naturalAverageSampledReturn_eq_optimal_sub_naturalAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn_eq_optimal_sub_naturalAverageRealizedBehaviorRegret","description":"At every prefix, observed average return is optimal value minus realized regret. The empty-prefix convention makes the identity valid at zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-93e8da6ca7c1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8567,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAverageSampledReturn_eq_optimal_sub_naturalAverageRealizedBehaviorRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.naturalAverageSampledReturn trajectory rounds = AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn mdp initialState - source.naturalAverageRealizedBehaviorRegret trajectory rounds","missing":[],"search":"naturalaveragesampledreturn_eq_optimal_sub_naturalaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragesampledreturn_eq_optimal_sub_naturalaveragerealizedbehaviorregret at every prefix, observed average return is optimal value minus realized regret. the empty-prefix convention makes the identity valid at zero. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageSampledReturn","label":"measurable_naturalAverageSampledReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageSampledReturn","description":"A fixed-prefix average sampled-return coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-1c3e0ecaa5a7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8568,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalAverageSampledReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.naturalAverageSampledReturn trajectory rounds)","missing":[],"search":"measurable_naturalaveragesampledreturn banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalaveragesampledreturn a fixed-prefix average sampled-return coordinate is measurable. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageSampledReturn_naturalTrajectoryFiltration","label":"measurable_naturalAverageSampledReturn_naturalTrajectoryFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageSampledReturn_naturalTrajectoryFiltration","description":"A fixed-prefix average sampled return is measurable at its natural filtration level.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-6d2b5b912b13","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8569,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:121"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalAverageSampledReturn_naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : @Measurable (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) Real (source.naturalTrajectoryFiltration rounds) inferInstance (fun trajectory => source.naturalAverageSampledReturn trajectory rounds)","missing":[],"search":"measurable_naturalaveragesampledreturn_naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalaveragesampledreturn_naturaltrajectoryfiltration a fixed-prefix average sampled return is measurable at its natural filtration level. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn_stronglyAdapted_naturalTrajectoryFiltration","label":"naturalAverageSampledReturn_stronglyAdapted_naturalTrajectoryFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn_stronglyAdapted_naturalTrajectoryFiltration","description":"The natural average sampled-return process is strongly adapted.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-8cbc4383fe63","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8570,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:148"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAverageSampledReturn_stronglyAdapted_naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : StronglyAdapted source.naturalTrajectoryFiltration (fun rounds trajectory => source.naturalAverageSampledReturn trajectory rounds)","missing":[],"search":"naturalaveragesampledreturn_stronglyadapted_naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragesampledreturn_stronglyadapted_naturaltrajectoryfiltration the natural average sampled-return process is strongly adapted. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess","label":"selfConsistentScheduledNaturalCausalAverageSampledReturnProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess","description":"Natural-prefix average of the observed successor-batch sample means for the self-consistent causal source.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-bc6efed1ed34","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8571,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:167"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageSampledReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragesampledreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragesampledreturnprocess natural-prefix average of the observed successor-batch sample means for the self-consistent causal source. definition compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_eq_optimal_sub_realized","label":"selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_eq_optimal_sub_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_eq_optimal_sub_realized","description":"The project-specific sampled-return process is exactly the complement of the average realized-regret process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-83ab4ac7246d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8572,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:184"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_eq_optimal_sub_realized (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory) : selfConsistentScheduledNaturalCausalAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory = AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn mdp initialState - selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalaveragesampledreturnprocess_eq_optimal_sub_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragesampledreturnprocess_eq_optimal_sub_realized the project-specific sampled-return process is exactly the complement of the average realized-regret process. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_stronglyAdapted","label":"selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_stronglyAdapted","description":"The self-consistent sampled-return process is strongly adapted to the exact natural trajectory filtration.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-9ce75b51bc93","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8573,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_stronglyAdapted (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) : StronglyAdapted (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (selfConsistentScheduledNaturalCausalAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor)","missing":[],"search":"selfconsistentschedulednaturalcausalaveragesampledreturnprocess_stronglyadapted banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragesampledreturnprocess_stronglyadapted the self-consistent sampled-return process is strongly adapted to the exact natural trajectory filtration. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","description":"Average sampled return evaluated at a `WithTop Nat` stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-7b1a1a5d35db","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8574,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:229"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess average sampled return evaluated at a `withtop nat` stopping prefix. definition compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_apply","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTraje…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-07d61b4ea3f5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8575,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:253"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex trajectory = selfConsistentScheduledNaturalCausalAverageSampledReturnProcess mdp initialState rewardSource initi…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_apply theorem selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (scheduleindex : nat) (trajectory) : selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor stoppingprefix scheduleindex trajectory = selfconsistentschedulednaturalcausalaveragesampledreturnprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor (stoppingprefix scheduleindex trajectory).untopa trajectory theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","description":"Stopping preserves the exact sampled-return/realized-regret complement.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-2c079b2c9e45","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8576,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:275"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex trajectory = AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn mdp initialStat…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_eq_optimal_sub_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_eq_optimal_sub_realized stopping preserves the exact sampled-return/realized-regret complement. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","label":"measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","description":"Mathlib stopped-value measurability for the sampled-return process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-5c5c41a62dd4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8577,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:304"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (hstopping : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (stoppingPrefix scheduleIndex)) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess mdp initialState r…","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess mathlib stopped-value measurability for the sampled-return process. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_of_integrable_realized","label":"integrable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_of_integrable_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_of_integrable_realized","description":"Integrability of stopped realized regret transfers to stopped sampled return under any finite measure.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-6c2966d8dc21","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8578,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:341"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_of_integrable_realized {Omega : Type w} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (optimal : Real) (sampledReturn realizedRegret : Omega -> Real) (hpoint : forall omega, sampledReturn omega = optimal - realizedRegret omega) (hrealized : Integrable realizedRegret mu) : Integrable sampledReturn mu","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_of_integrable_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_of_integrable_realized integrability of stopped realized regret transfers to stopped sampled return under any finite measure. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","label":"integral_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","description":"The corresponding Bochner expectation is the optimal constant minus the expected realized regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-e7c017c79c44","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8579,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:355"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized {Omega : Type w} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (optimal : Real) (sampledReturn realizedRegret : Omega -> Real) (hpoint : forall omega, sampledReturn omega = optimal - realizedRegret omega) (hrealized : Integrable realizedRegret mu) : integral mu sampledReturn = optimal - integral mu realizedRegret","missing":[],"search":"integral_selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_eq_optimal_sub_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnprocess_eq_optimal_sub_realized the corresponding bochner expectation is the optimal constant minus the expected realized regret. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","label":"integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","description":"The capped inverse-sqrt first-passage sampled-return coordinate is integrable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-5c243a2729b4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8580,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:370"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initi…","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturn the capped inverse-sqrt first-passage sampled-return coordinate is integrable. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","label":"integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","description":"The genuine uncapped `hittingAfter` sampled-return coordinate is integrable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-b1d1ff744c79","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8581,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:431"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rew…","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturn the genuine uncapped `hittingafter` sampled-return coordinate is integrable. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","description":"For the capped prefix, expected sampled return is exactly optimal value minus expected realized regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-61c719e99ab8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8582,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:492"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialSta…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnintegral_eq_optimal_sub_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnintegral_eq_optimal_sub_realized for the capped prefix, expected sampled return is exactly optimal value minus expected realized regret. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","description":"For the uncapped prefix, expected sampled return is exactly optimal value minus expected realized regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-6ccd28135315","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8583,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:548"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnintegral_eq_optimal_sub_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnintegral_eq_optimal_sub_realized for the uncapped prefix, expected sampled return is exactly optimal value minus expected realized regret. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_tendsto_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_tendsto_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_tendsto_optimal","description":"Expected sampled return at the capped prefix converges to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-e98e1754e138","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8584,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:604"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_tendsto_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable d…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnintegral_tendsto_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnintegral_tendsto_optimal expected sampled return at the capped prefix converges to the optimal initial expected return. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_tendsto_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_tendsto_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_tendsto_optimal","description":"Expected sampled return at the genuine uncapped `hittingAfter` prefix converges to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-ad1f3dc568a7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8585,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:671"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_tendsto_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnintegral_tendsto_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnintegral_tendsto_optimal expected sampled return at the genuine uncapped `hittingafter` prefix converges to the optimal initial expected return. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageSampledReturn_expected_optimality","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageSampledReturn_expected_optimality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageSampledReturn_expected_optimality","description":"Terminal sampled-return semantics and expected-optimality package for the capped approximation and the genuine uncapped stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-cc5096f8e8ad/index.html#decl-04586854cea2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","order":8586,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality.lean:738"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageSampledReturn_expected_optimality (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSou…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedaveragesampledreturn_expected_optimality banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedaveragesampledreturn_expected_optimality terminal sampled-return semantics and expected-optimality package for the capped approximation and the genuine uncapped stopping prefix. theorem compiled","shard":"modules/5983195049ff0445.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGap_tendsto_zero","description":"The signed expectation gap between the uncapped and capped stopped behavior expected-regret coordinates tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-970415c3ecf9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8587,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewa…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretexpectedgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretexpectedgap_tendsto_zero the signed expectation gap between the uncapped and capped stopped behavior expected-regret coordinates tends to zero. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGapAbs_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGapAbs_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGapAbs_tendsto_zero","description":"The absolute expectation gap between the uncapped and capped stopped behavior expected-regret coordinates tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-c0a98b1873d6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8588,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGapAbs_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState r…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretexpectedgapabs_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretexpectedgapabs_tendsto_zero the absolute expectation gap between the uncapped and capped stopped behavior expected-regret coordinates tends to zero. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGap_tendsto_zero","description":"The signed expectation gap between the uncapped and capped stopped return-deviation coordinates tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-182ce2bd12b9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8589,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSourc…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationexpectedgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationexpectedgap_tendsto_zero the signed expectation gap between the uncapped and capped stopped return-deviation coordinates tends to zero. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGapAbs_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGapAbs_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGapAbs_tendsto_zero","description":"The absolute expectation gap between the uncapped and capped stopped return-deviation coordinates tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-bbb0f21f39cb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8590,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:270"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGapAbs_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSo…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationexpectedgapabs_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationexpectedgapabs_tendsto_zero the absolute expectation gap between the uncapped and capped stopped return-deviation coordinates tends to zero. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","description":"Exact expected decomposition for the capped first-passage prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-560a9d3543e0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8591,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:317"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCa…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_eq_behaviorexpected_sub_returndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_eq_behaviorexpected_sub_returndeviation exact expected decomposition for the capped first-passage prefix. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","description":"Exact expected decomposition for the genuine uncapped `hittingAfter` prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-4b447f33bb9c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8592,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:399"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsis…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_eq_behaviorexpected_sub_returndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_eq_behaviorexpected_sub_returndeviation exact expected decomposition for the genuine uncapped `hittingafter` prefix. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_expected_truncation_replacement","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_expected_truncation_replacement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_expected_truncation_replacement","description":"Terminal expected semantic replacement package for the capped and genuine uncapped first-passage prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-b315bfb9255e/index.html#decl-0871e1336271","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","order":8593,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement.lean:481"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_expected_truncation_replacement (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausa…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedbehaviorexpectedregret_and_returndeviation_expected_truncation_replacement banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedbehaviorexpectedregret_and_returndeviation_expected_truncation_replacement terminal expected semantic replacement package for the capped and genuine uncapped first-passage prefixes. theorem compiled","shard":"modules/a25d67f6ee687193.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_atTop","label":"ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_atTop","description":"The capped inverse-square-root first-passage prefix diverges after `WithTop.untopA`, almost surely.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-354ca23ef796","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8594,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_atTop (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardS…","missing":[],"search":"ae_tendsto_untopa_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix_attop banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_tendsto_untopa_selfconsistentschedulednaturalcausalinversesqrtthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix_attop the capped inverse-square-root first-passage prefix diverges after `withtop.untopa`, almost surely. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegret_and_returnDeviation_eq","label":"ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegret_and_returnDeviation_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegret_and_returnDeviation_eq","description":"Capped and uncapped stopped policy-value and return-deviation coordinates are eventually exactly equal, almost surely.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-941fc3f6868e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8595,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegret_and_returnDeviation_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp init…","missing":[],"search":"ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregret_and_returndeviation_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregret_and_returndeviation_eq capped and uncapped stopped policy-value and return-deviation coordinates are eventually exactly equal, almost surely. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","description":"The capped stopped behavior expected-regret process converges almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-238383cfff3e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8596,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initia…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedbehaviorexpectedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedbehaviorexpectedregret_tendstoalmosteverywhere_zero the capped stopped behavior expected-regret process converges almost everywhere. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","description":"Every capped stopped behavior expected-regret coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-97d811d1c3dd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8597,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:185"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregret every capped stopped behavior expected-regret coordinate is measurable. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","description":"Every capped stopped behavior expected-regret coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-2de7b72e895a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8598,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:215"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let behaviorProcess := selfConsistentScheduledNaturalCausalStoppingTimeAverageBe…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregret every capped stopped behavior expected-regret coordinate belongs to `l1`. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret_tendsto_zero","description":"The exponent-one norm of the capped stopped behavior expected regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-514e1e60a10c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8599,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:256"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTabl…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregret_tendsto_zero the exponent-one norm of the capped stopped behavior expected regret tends to zero. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegretIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegretIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegretIntegral_tendsto_zero","description":"Signed expectation of the capped stopped behavior expected regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-2ad5eaadbf23","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8600,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:373"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegretIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable de…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregretintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedbehaviorexpectedregretintegral_tendsto_zero signed expectation of the capped stopped behavior expected regret tends to zero. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation","description":"Every capped stopped return-deviation coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-262c5122e765","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8601,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:414"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTabl…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedreturndeviation every capped stopped return-deviation coordinate belongs to `l1`. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation_tendsto_zero","description":"Exponent-one norm of the capped stopped return deviation tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-be151e5193b5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8602,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:490"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defau…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedreturndeviation_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedreturndeviation_tendsto_zero exponent-one norm of the capped stopped return deviation tends to zero. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviationIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviationIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviationIntegral_tendsto_zero","description":"Signed expectation of the capped stopped return deviation tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-b29c41a84519","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8603,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:606"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviationIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultSt…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedreturndeviationintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedreturndeviationintegral_tendsto_zero signed expectation of the capped stopped return deviation tends to zero. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference","description":"The uncapped-minus-capped stopped behavior expected-regret coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-34f89b6894ac","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8604,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:649"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let cappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let uncappedStoppingPrefix := selfConsisten…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretdifference banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretdifference the uncapped-minus-capped stopped behavior expected-regret coordinate belongs to `l1`. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference_tendsto_zero","description":"Capped and uncapped stopped behavior expected regret are asymptotically equivalent in exponent-one norm.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-cb8fec9148f5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8605,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:690"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initia…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretdifference_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretdifference_tendsto_zero capped and uncapped stopped behavior expected regret are asymptotically equivalent in exponent-one norm. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifferenceIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifferenceIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifferenceIntegral_tendsto_zero","description":"Signed integral of the stopped behavior truncation difference tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-e349807cc677","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8606,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:800"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifferenceIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialSta…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretdifferenceintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedbehaviorexpectedregretdifferenceintegral_tendsto_zero signed integral of the stopped behavior truncation difference tends to zero. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_eq_behaviorExpectedRegretDifference_sub_realizedBehaviorRegretDifference","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_eq_behaviorExpectedRegretDifference_sub_realizedBehaviorRegretDifference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_eq_behaviorExpectedRegretDifference_sub_realizedBehaviorRegretDifference","description":"The semantic return-difference is exactly the behavior-value difference minus the realized-regret difference.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-90809dfa4422","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8607,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:850"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_eq_behaviorExpectedRegretDifference_sub_realizedBehaviorRegretDifference (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : let cappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let uncappe…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifference_eq_behaviorexpectedregretdifference_sub_realizedbehaviorregretdifference banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifference_eq_behaviorexpectedregretdifference_sub_realizedbehaviorregretdifference the semantic return-difference is exactly the behavior-value difference minus the realized-regret difference. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference","description":"The uncapped-minus-capped stopped return-deviation coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-8349065cbb8e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8608,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:921"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initia…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifference banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifference the uncapped-minus-capped stopped return-deviation coordinate belongs to `l1`. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_tendsto_zero","description":"Capped and uncapped stopped return deviation are asymptotically equivalent in exponent-one norm.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-ed5286319906","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8609,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:970"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifference_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifference_tendsto_zero capped and uncapped stopped return deviation are asymptotically equivalent in exponent-one norm. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifferenceIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifferenceIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifferenceIntegral_tendsto_zero","description":"Signed integral of the stopped return-deviation truncation difference tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-7225de7e2d15","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8610,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:1118"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifferenceIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewa…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifferenceintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturndeviationdifferenceintegral_tendsto_zero signed integral of the stopped return-deviation truncation difference tends to zero. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_truncation_equivalence","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_truncation_equivalence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_truncation_equivalence","description":"Terminal componentwise semantic and `L1` truncation-equivalence package for the capped and genuine uncapped first-passage prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-fee5c7b763ca/index.html#decl-8b8607084f44","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","order":8611,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence.lean:1169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_truncation_equivalence (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSourc…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedbehaviorexpectedregret_and_returndeviation_l1_truncation_equivalence banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedbehaviorexpectedregret_and_returndeviation_l1_truncation_equivalence terminal componentwise semantic and `l1` truncation-equivalence package for the capped and genuine uncapped first-passage prefixes. theorem compiled","shard":"modules/b9d5504243a1df3c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","description":"For the capped prefix, the expected realized-minus-policy-value gap is exactly the negative expected return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-db5be44ca073","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8612,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source :=…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_sub_behaviorexpectedregretintegral_eq_neg_returndeviationintegral banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretintegral_sub_behaviorexpectedregretintegral_eq_neg_returndeviationintegral for the capped prefix, the expected realized-minus-policy-value gap is exactly the negative expected return deviation. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","description":"For the genuine uncapped `hittingAfter` prefix, the expected realized-minus-policy-value gap is exactly the negative expected return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-4dc5ba353436","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8613,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat)…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_sub_behaviorexpectedregretintegral_eq_neg_returndeviationintegral banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_sub_behaviorexpectedregretintegral_eq_neg_returndeviationintegral for the genuine uncapped `hittingafter` prefix, the expected realized-minus-policy-value gap is exactly the negative expected return deviation. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","description":"The capped expected realized-minus-policy-value gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-bd4b10a5e1c2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8614,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardS…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgap_tendsto_zero the capped expected realized-minus-policy-value gap tends to zero. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","description":"The absolute capped expected realized-minus-policy-value gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-75c690fd3c64","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8615,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:203"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewa…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgapabs_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgapabs_tendsto_zero the absolute capped expected realized-minus-policy-value gap tends to zero. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","description":"The genuine uncapped expected realized-minus-policy-value gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-0ad42d2be8bd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8616,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:247"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initi…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgap_tendsto_zero the genuine uncapped expected realized-minus-policy-value gap tends to zero. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","description":"The absolute genuine uncapped expected realized-minus-policy-value gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-4f5287397aea","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8617,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:326"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp in…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgapabs_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretpolicyvalueexpectedgapabs_tendsto_zero the absolute genuine uncapped expected realized-minus-policy-value gap tends to zero. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedRealizedBehaviorRegret_and_policyValue_expected_consistency","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedRealizedBehaviorRegret_and_policyValue_expected_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedRealizedBehaviorRegret_and_policyValue_expected_consistency","description":"Terminal expected-consistency square for capped and genuine uncapped first-passage prefixes. The horizontal edges replace capped expectations by uncapped expectations, while the vertical edges replace successor-policy value gaps by realized regret expectations.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-46d57f4e4800/index.html#decl-dc7ec2bca187","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","order":8618,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency.lean:372"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedRealizedBehaviorRegret_and_policyValue_expected_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp ini…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedrealizedbehaviorregret_and_policyvalue_expected_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedrealizedbehaviorregret_and_policyvalue_expected_consistency terminal expected-consistency square for capped and genuine uncapped first-passage prefixes. the horizontal edges replace capped expectations by uncapped expectations, while the vertical edges replace successor-policy value gaps by realized regret expectations. theorem compiled","shard":"modules/f1d931ec59766ce7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.probReal_compl_gt_one_sub_of_measure_lt_ofReal","label":"probReal_compl_gt_one_sub_of_measure_lt_ofReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.probReal_compl_gt_one_sub_of_measure_lt_ofReal","description":"A strict `ENNReal` upper bound on a measurable bad event gives the corresponding strict real-probability lower bound on its complement.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-75508fd13d99","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8619,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem probReal_compl_gt_one_sub_of_measure_lt_ofReal {Omega : Type w} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] {event : Set Omega} (hevent : MeasurableSet event) {delta : Real} (hdelta : 0 < delta) (htail : mu event < ENNReal.ofReal delta) : 1 - delta < mu.real event.compl","missing":[],"search":"probreal_compl_gt_one_sub_of_measure_lt_ofreal banditrlproof.finitehorizonrl.probreal_compl_gt_one_sub_of_measure_lt_ofreal a strict `ennreal` upper bound on a measurable bad event gives the corresponding strict real-probability lower bound on its complement. theorem compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","description":"The simultaneous six-coordinate stopped-return good event is exactly the complement of the existing joint violation event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-9204347ac407","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8620,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset the simultaneous six-coordinate stopped-return good event is exactly the complement of the existing joint violation event. definition compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","label":"measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","description":"The named six-coordinate good event is measurable at every schedule index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-28b7112634af","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8621,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : MeasurableSet (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon scheduleIndex)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset the named six-coordinate good event is measurable at every schedule index. theorem compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_iff","label":"mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_iff","description":"Membership in the named good event is equivalent to all six literal stopped-return errors being strictly below the same threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-6cfc2d7e82b9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8622,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:91"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_iff (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : let cappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let uncappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSq…","missing":[],"search":"mem_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset_iff banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.mem_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset_iff membership in the named good event is equivalent to all six literal stopped-return errors being strictly below the same threshold. theorem compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_probReal_gt_one_sub","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_probReal_gt_one_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_probReal_gt_one_sub","description":"A strict mass bound for the named bad event yields the corresponding real probability lower bound for the named good event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-d5697d5c9d8f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8623,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_probReal_gt_one_sub (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon delta : Real) (scheduleIndex : Nat) (hdelta : 0 < delta) (htail : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon scheduleIndex < ENNReal.ofReal delta) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor 1 - delta < source.trajectoryMeasure.real…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset_probreal_gt_one_sub banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset_probreal_gt_one_sub a strict mass bound for the named bad event yields the corresponding real probability lower bound for the named good event. theorem compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_tailStart_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","label":"exists_tailStart_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_tailStart_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","description":"The qualitative eventual confidence theorem has one deterministic natural tail start that controls every later schedule index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-6c9d6254db3b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8624,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:202"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exists_tailStart_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon delta : Real) (hepsilon : 0 < epsilon) (hdelta : 0 < delta)…","missing":[],"search":"exists_tailstart_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability_lt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exists_tailstart_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability_lt the qualitative eventual confidence theorem has one deterministic natural tail start that controls every later schedule index. theorem compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_deterministicTailHighProbability_optimality","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_deterministicTailHighProbability_optimality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_deterministicTailHighProbability_optimality","description":"Terminal deterministic-tail confidence certificate. One noncomputable cutoff works for every later schedule index, and the named good event combines real probability mass with the six literal stopped-return error bounds.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-54cef3e28ac5/index.html#decl-1d2230dedb36","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","order":8625,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_deterministicTailHighProbability_optimality (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsist…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_deterministictailhighprobability_optimality banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_deterministictailhighprobability_optimality terminal deterministic-tail confidence certificate. one noncomputable cutoff works for every later schedule index, and the named good event combines real probability mass with the six literal stopped-return error bounds. theorem compiled","shard":"modules/c3ed91af080dc8bd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","description":"The capped stopped realized-regret process inherits the genuine uncapped almost-sure limit because the two processes are eventually equal almost surely.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-19cd53ebbace","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8626,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero the capped stopped realized-regret process inherits the genuine uncapped almost-sure limit because the two processes are eventually equal almost surely. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_eq_behaviorExpected_sub_realized","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_eq_behaviorExpected_sub_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_eq_behaviorExpected_sub_realized","description":"The stopped return deviation is behavior expected regret minus realized behavior regret, pointwise at any common stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-67bf5791d9b0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8627,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_eq_behaviorExpected_sub_realized (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex = fun trajectory => selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedReg…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess_eq_behaviorexpected_sub_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess_eq_behaviorexpected_sub_realized the stopped return deviation is behavior expected regret minus realized behavior regret, pointwise at any common stopping prefix. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn_tendstoInMeasure_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn_tendstoInMeasure_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn_tendstoInMeasure_optimal","description":"The literal capped stopped sampled return converges in measure to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-79a5b8cd2a22","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8628,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn_tendstoInMeasure_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturn_tendstoinmeasure_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturn_tendstoinmeasure_optimal the literal capped stopped sampled return converges in measure to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn_tendstoInMeasure_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn_tendstoInMeasure_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn_tendstoInMeasure_optimal","description":"The literal uncapped stopped sampled return converges in measure to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-a23f152e0662","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8629,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn_tendstoInMeasure_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSourc…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturn_tendstoinmeasure_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturn_tendstoinmeasure_optimal the literal uncapped stopped sampled return converges in measure to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","description":"The literal capped stopped expected return of the selected successor policies converges in measure to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-4a93b9b3772c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8630,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:207"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSour…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturn_tendstoinmeasure_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturn_tendstoinmeasure_optimal the literal capped stopped expected return of the selected successor policies converges in measure to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","description":"The literal uncapped stopped expected return of the selected successor policies converges in measure to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-755b697b2e61","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8631,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:252"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialS…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturn_tendstoinmeasure_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturn_tendstoinmeasure_optimal the literal uncapped stopped expected return of the selected successor policies converges in measure to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","description":"The capped same-prefix sampled/policy return gap converges in measure to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-34f7c68f869c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8632,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:297"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initia…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap_tendstoinmeasure_zero the capped same-prefix sampled/policy return gap converges in measure to zero. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","description":"The uncapped same-prefix sampled/policy return gap converges in measure to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-b0e34e11f641","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8633,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:344"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewa…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap_tendstoinmeasure_zero the uncapped same-prefix sampled/policy return gap converges in measure to zero. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","description":"The literal capped stopped sampled return converges almost surely to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-6989066a3b8c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8634,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:391"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initi…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedaveragesampledreturn_tendstoalmosteverywhere_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedaveragesampledreturn_tendstoalmosteverywhere_optimal the literal capped stopped sampled return converges almost surely to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","description":"The literal uncapped stopped sampled return converges almost surely to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-61b828c28d02","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8635,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:454"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rew…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragesampledreturn_tendstoalmosteverywhere_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragesampledreturn_tendstoalmosteverywhere_optimal the literal uncapped stopped sampled return converges almost surely to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","description":"The literal capped stopped expected return of the selected successor policies converges almost surely to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-7e9fa7222075","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8636,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:517"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState re…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedaveragesuccessorpolicyexpectedreturn_tendstoalmosteverywhere_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedaveragesuccessorpolicyexpectedreturn_tendstoalmosteverywhere_optimal the literal capped stopped expected return of the selected successor policies converges almost surely to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","description":"The literal uncapped stopped expected return of the selected successor policies converges almost surely to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-d917a56f12bf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8637,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:580"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragesuccessorpolicyexpectedreturn_tendstoalmosteverywhere_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragesuccessorpolicyexpectedreturn_tendstoalmosteverywhere_optimal the literal uncapped stopped expected return of the selected successor policies converges almost surely to the optimal initial expected return. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","description":"The capped same-prefix sampled/policy return gap converges almost surely to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-87408f8c1b50","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8638,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:643"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSourc…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedsampledpolicyexpectedreturngap_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcapped_stoppedsampledpolicyexpectedreturngap_tendstoalmosteverywhere_zero the capped same-prefix sampled/policy return gap converges almost surely to zero. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","description":"The uncapped same-prefix sampled/policy return gap converges almost surely to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-9eb38c8ad69e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8639,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:720"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialSt…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedsampledpolicyexpectedreturngap_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedsampledpolicyexpectedreturngap_tendstoalmosteverywhere_zero the uncapped same-prefix sampled/policy return gap converges almost surely to zero. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_inMeasure_and_almostSure_optimality","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_inMeasure_and_almostSure_optimality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_inMeasure_and_almostSure_optimality","description":"Terminal convergence-mode package for literal stopped sampled return, the literal expected return of the actually selected successor policies, and their same-prefix gap at the capped and genuine uncapped stopping prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-096e546f61ce/index.html#decl-bf9e0157603a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","order":8640,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality.lean:798"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_inMeasure_and_almostSure_optimality (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentSched…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_inmeasure_and_almostsure_optimality banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_inmeasure_and_almostsure_optimality terminal convergence-mode package for literal stopped sampled return, the literal expected return of the actually selected successor policies, and their same-prefix gap at the capped and genuine uncapped stopping prefixes. theorem compiled","shard":"modules/8b5632841b59a829.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_fun_neg","label":"eLpNorm_fun_neg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_fun_neg","description":"Eta-expanded negation has the same `eLpNorm`. This small wrapper keeps the return-process transport proofs independent of the syntactic representation of pointwise negation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-dc4aef3f2a2a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8641,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_fun_neg {Omega : Type*} [MeasurableSpace Omega] (f : Omega -> Real) (p : ENNReal) (mu : Measure Omega) : eLpNorm (fun omega => -f omega) p mu = eLpNorm f p mu","missing":[],"search":"elpnorm_fun_neg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_fun_neg eta-expanded negation has the same `elpnorm`. this small wrapper keeps the return-process transport proofs independent of the syntactic representation of pointwise negation. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess","description":"Sampled return centered at the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-c409d3405022","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8642,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnoptimalityerrorprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnoptimalityerrorprocess sampled return centered at the optimal initial expected return. definition compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess","description":"Literal successor-policy expected return centered at the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-424a4d1aa84f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8643,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnoptimalityerrorprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnoptimalityerrorprocess literal successor-policy expected return centered at the optimal initial expected return. definition compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess","description":"Same-prefix sampled return minus literal successor-policy expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-a7f064b077c0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8644,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledpolicyexpectedreturngapprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledpolicyexpectedreturngapprocess same-prefix sampled return minus literal successor-policy expected return. definition compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess_eq_neg_realized","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess_eq_neg_realized","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess_eq_neg_realized","description":"The stopped sampled-return optimality error is the negative stopped realized behavior regret, pointwise and before integration.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-120a52e86427","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8645,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess_eq_neg_realized (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex = fun trajectory => -selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedB…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnoptimalityerrorprocess_eq_neg_realized banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturnoptimalityerrorprocess_eq_neg_realized the stopped sampled-return optimality error is the negative stopped realized behavior regret, pointwise and before integration. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess_eq_neg_behaviorExpected","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess_eq_neg_behaviorExpected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess_eq_neg_behaviorExpected","description":"The stopped literal successor-policy return optimality error is the negative stopped behavior expected regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-8294ca0f4ae5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8646,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:155"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess_eq_neg_behaviorExpected (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex = fun trajectory => -selfConsistentScheduledN…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnoptimalityerrorprocess_eq_neg_behaviorexpected banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnoptimalityerrorprocess_eq_neg_behaviorexpected the stopped literal successor-policy return optimality error is the negative stopped behavior expected regret. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess_eq_returnDeviation","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess_eq_returnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess_eq_returnDeviation","description":"The stopped sampled/policy return gap is exactly the stopped return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-2c8658f5be8b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8647,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess_eq_returnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex = selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProces…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledpolicyexpectedreturngapprocess_eq_returndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledpolicyexpectedreturngapprocess_eq_returndeviation the stopped sampled/policy return gap is exactly the stopped return deviation. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError","description":"Every capped stopped sampled-return optimality error belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-80c0f3b65a38","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8648,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSourc…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledreturnoptimalityerror banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledreturnoptimalityerror every capped stopped sampled-return optimality error belongs to `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError_tendsto_zero","description":"The capped sampled-return optimality error converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-40057ec76a12","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8649,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:246"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initi…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledreturnoptimalityerror_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledreturnoptimalityerror_tendsto_zero the capped sampled-return optimality error converges to zero in `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError","description":"Every uncapped `hittingAfter` stopped sampled-return optimality error belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-9a3a1b95617c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8650,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:286"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialSt…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledreturnoptimalityerror banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledreturnoptimalityerror every uncapped `hittingafter` stopped sampled-return optimality error belongs to `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError_tendsto_zero","description":"The genuine uncapped stopped sampled-return optimality error converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-45636f6e71d4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8651,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:323"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rew…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledreturnoptimalityerror_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledreturnoptimalityerror_tendsto_zero the genuine uncapped stopped sampled-return optimality error converges to zero in `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError","description":"Every capped stopped literal successor-policy return optimality error belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-ef0b747bee1d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8652,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:363"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let errorProcess := selfConsistentScheduledNaturalCausalSt…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsuccessorpolicyexpectedreturnoptimalityerror banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsuccessorpolicyexpectedreturnoptimalityerror every capped stopped literal successor-policy return optimality error belongs to `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","description":"The capped literal successor-policy return optimality error converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-80887ca7ef33","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8653,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:393"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState re…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsuccessorpolicyexpectedreturnoptimalityerror_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsuccessorpolicyexpectedreturnoptimalityerror_tendsto_zero the capped literal successor-policy return optimality error converges to zero in `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError","description":"Every genuine uncapped stopped literal successor-policy return optimality error belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-a34c89be49ce","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8654,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:433"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let errorProcess := selfConsistentScheduledNaturalCausalStopp…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsuccessorpolicyexpectedreturnoptimalityerror banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsuccessorpolicyexpectedreturnoptimalityerror every genuine uncapped stopped literal successor-policy return optimality error belongs to `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","description":"The genuine uncapped literal successor-policy return optimality error converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-faec39e977df","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8655,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:463"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsuccessorpolicyexpectedreturnoptimalityerror_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsuccessorpolicyexpectedreturnoptimalityerror_tendsto_zero the genuine uncapped literal successor-policy return optimality error converges to zero in `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","description":"Every capped stopped sampled/policy return gap belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-eb101794d961","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8656,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:502"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSou…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap every capped stopped sampled/policy return gap belongs to `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendsto_zero","description":"The capped stopped sampled/policy return gap converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-c1ede8fbf503","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8657,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:538"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource ini…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap_tendsto_zero the capped stopped sampled/policy return gap converges to zero in `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","description":"Every genuine uncapped stopped sampled/policy return gap belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-8b91d561a823","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8658,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:576"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initial…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap every genuine uncapped stopped sampled/policy return gap belongs to `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendsto_zero","description":"The genuine uncapped stopped sampled/policy return gap converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-7b52a42d6118","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8659,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:613"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState r…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap_tendsto_zero the genuine uncapped stopped sampled/policy return gap converges to zero in `l1`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_L1_optimality","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_L1_optimality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_L1_optimality","description":"Terminal true-`L1` package for stopped sampled return, the literal return of the actually selected successor policies, and their same-prefix gap at both the capped first-passage approximation and genuine uncapped `hittingAfter`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-0130b328725b/index.html#decl-c6b98916566c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","order":8660,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality.lean:653"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_L1_optimality (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp i…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_l1_optimality banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_l1_optimality terminal true-`l1` package for stopped sampled return, the literal return of the actually selected successor policies, and their same-prefix gap at both the capped first-passage approximation and genuine uncapped `hittingafter`. theorem compiled","shard":"modules/071a18b7f72511e3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","description":"Every capped stopped sampled-return coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-92879534b1c0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8661,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturn every capped stopped sampled-return coordinate is measurable. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","description":"Every genuine uncapped stopped sampled-return coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-bf45b088604c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8662,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturn every genuine uncapped stopped sampled-return coordinate is measurable. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","description":"Every capped stopped sampled-minus-policy-return gap is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-34d771c1fca8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8663,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:86"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedsampledpolicyexpectedreturngap every capped stopped sampled-minus-policy-return gap is measurable. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","description":"Every genuine uncapped stopped sampled-minus-policy-return gap is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-754f9314e28d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8664,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedsampledpolicyexpectedreturngap every genuine uncapped stopped sampled-minus-policy-return gap is measurable. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","description":"At one schedule index, at least one of the six literal capped/uncapped stopped-return errors is at least `epsilon`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-0e37fc46dcf6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8665,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:139"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset at one schedule index, at least one of the six literal capped/uncapped stopped-return errors is at least `epsilon`. definition compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","description":"The six-way stopped-return violation event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-08d54aabf6d1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8666,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:198"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : MeasurableSet (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon scheduleIndex)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset the six-way stopped-return violation event is measurable. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_iff","label":"not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_iff","description":"Outside the joint violation event, all six literal errors are strictly smaller than the same accuracy threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-67f767c1c52d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8667,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:250"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_iff (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : let cappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let uncappedStoppingPrefix := selfConsistentScheduledNaturalCausal…","missing":[],"search":"not_mem_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset_iff banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.not_mem_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset_iff outside the joint violation event, all six literal errors are strictly smaller than the same accuracy threshold. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability","description":"Probability of the six-way simultaneous stopped-return violation event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-95d85b808732","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8668,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:313"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : ENNReal","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability probability of the six-way simultaneous stopped-return violation event. definition compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_tendsto_zero","description":"The simultaneous six-way violation probability tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-f04d62bca9a0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8669,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:328"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : Tendsto (selfConsistentSchedule…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability_tendsto_zero the simultaneous six-way violation probability tends to zero. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","label":"eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","description":"Every positive accuracy and confidence budget is eventually met simultaneously by all six stopped-return coordinates.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-b90d052b9196","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8670,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:481"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon delta : Real) (hepsilon : 0 < epsilon) (hdelta : 0 < delta) : ∀ᶠ…","missing":[],"search":"eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability_lt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationprobability_lt every positive accuracy and confidence budget is eventually met simultaneously by all six stopped-return coordinates. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_simultaneousHighProbability_optimality","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_simultaneousHighProbability_optimality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_simultaneousHighProbability_optimality","description":"Terminal simultaneous high-probability certificate for literal stopped sampled return, actual successor-policy return, and their same-prefix gap at the capped and genuine uncapped inverse-square-root first-passage prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-2bfa90f5b3a0/index.html#decl-15ee000d4075","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","order":8671,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality.lean:512"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_simultaneousHighProbability_optimality (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let violationSet := selfConsis…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_simultaneoushighprobability_optimality banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_simultaneoushighprobability_optimality terminal simultaneous high-probability certificate for literal stopped sampled return, actual successor-policy return, and their same-prefix gap at the capped and genuine uncapped inverse-square-root first-passage prefixes. theorem compiled","shard":"modules/2427fcc33fef519d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorPolicyExpectedReturn","label":"naturalSuccessorPolicyExpectedReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorPolicyExpectedReturn","description":"Literal expected cumulative reward of the successor policy selected from the dependent prefix through `t`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-4a251322e486","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8672,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalSuccessorPolicyExpectedReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (t : Nat) : Real","missing":[],"search":"naturalsuccessorpolicyexpectedreturn banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessorpolicyexpectedreturn literal expected cumulative reward of the successor policy selected from the dependent prefix through `t`. definition compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","label":"naturalSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","description":"The literal successor-policy return is optimal value minus that policy's expected regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-7786de25cacd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8673,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (t : Nat) : source.naturalSuccessorPolicyExpectedReturn trajectory t = AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn mdp initialState - (source.successorPolicyAt trajectory t).expectedRegret initialState","missing":[],"search":"naturalsuccessorpolicyexpectedreturn_eq_optimal_sub_expectedregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessorpolicyexpectedreturn_eq_optimal_sub_expectedregret the literal successor-policy return is optimal value minus that policy's expected regret. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSuccessorPolicyExpectedReturn","label":"naturalAverageSuccessorPolicyExpectedReturn","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSuccessorPolicyExpectedReturn","description":"Equal-round average of literal successor-policy expected returns. The empty prefix uses the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-79b58fc7cdbf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8674,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAverageSuccessorPolicyExpectedReturn {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"naturalaveragesuccessorpolicyexpectedreturn banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragesuccessorpolicyexpectedreturn equal-round average of literal successor-policy expected returns. the empty prefix uses the optimal initial expected return. definition compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","label":"naturalAverageSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","description":"At every prefix, the literal policy-return average is optimal value minus the average successor-policy expected regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-b2f4407a3dbc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8675,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAverageSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.naturalAverageSuccessorPolicyExpectedReturn trajectory rounds = AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn mdp initialState - (∑ t ∈ Finset.range rounds, (source.successorPolicyAt trajectory t).expectedRegret initialState) / (rounds : Real)","missing":[],"search":"naturalaveragesuccessorpolicyexpectedreturn_eq_optimal_sub_expectedregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragesuccessorpolicyexpectedreturn_eq_optimal_sub_expectedregret at every prefix, the literal policy-return average is optimal value minus the average successor-policy expected regret. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","label":"selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","description":"Natural-prefix average of literal expected returns of the actual successor policies selected by the self-consistent causal source.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-54b1ec25cd0f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8676,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:131"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess natural-prefix average of literal expected returns of the actual successor policies selected by the self-consistent causal source. definition compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","label":"selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","description":"The literal prefix policy return is the complement of the existing average behavior expected-regret process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-6f6f0478da37","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8677,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:150"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory) : selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory = AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn mdp initialState - selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess_eq_optimal_sub_behaviorexpected banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess_eq_optimal_sub_behaviorexpected the literal prefix policy return is the complement of the existing average behavior expected-regret process. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","label":"measurable_selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","description":"Every deterministic-prefix literal policy-return coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-6e294401c5a7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8678,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess every deterministic-prefix literal policy-return coordinate is measurable. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","description":"Literal average successor-policy expected return evaluated at a `WithTop Nat` stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-b5cef1c33fab","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8679,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:213"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess literal average successor-policy expected return evaluated at a `withtop nat` stopping prefix. definition compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_apply","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticE…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-d8e7940266f6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8680,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:238"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex trajectory = selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedR…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess_apply theorem selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (scheduleindex : nat) (trajectory) : selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor stoppingprefix scheduleindex trajectory = selfconsistentschedulednaturalcausalaveragesuccessorpolicyexpectedreturnprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor (stoppingprefix scheduleindex trajectory).untopa trajectory theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","description":"Stopping preserves the exact literal policy-return/behavior-regret complement.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-527170c43fb6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8681,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:263"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex trajectory = AdaptiveStochasticEpisodeBatchSource.opti…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess_eq_optimal_sub_behaviorexpected banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess_eq_optimal_sub_behaviorexpected stopping preserves the exact literal policy-return/behavior-regret complement. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","label":"measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","description":"A stopping time gives a measurable stopped literal policy-return coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-4f6f01f76d7e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8682,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:295"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (hstopping : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (stoppingPrefix scheduleIndex)) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpected…","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragesuccessorpolicyexpectedreturnprocess a stopping time gives a measurable stopped literal policy-return coordinate. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_stoppedSuccessorPolicyExpectedReturn_of_integrable_behavior","label":"integrable_stoppedSuccessorPolicyExpectedReturn_of_integrable_behavior","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_stoppedSuccessorPolicyExpectedReturn_of_integrable_behavior","description":"Integrability transfers from a stopped behavior expected-regret process to its literal policy-return complement.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-f7684cee5e3f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8683,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:340"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_stoppedSuccessorPolicyExpectedReturn_of_integrable_behavior {Omega : Type w} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (optimal : Real) (policyReturn behaviorRegret : Omega -> Real) (hpoint : forall omega, policyReturn omega = optimal - behaviorRegret omega) (hbehavior : Integrable behaviorRegret mu) : Integrable policyReturn mu","missing":[],"search":"integrable_stoppedsuccessorpolicyexpectedreturn_of_integrable_behavior banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_stoppedsuccessorpolicyexpectedreturn_of_integrable_behavior integrability transfers from a stopped behavior expected-regret process to its literal policy-return complement. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_stoppedSuccessorPolicyExpectedReturn_eq_optimal_sub_behavior","label":"integral_stoppedSuccessorPolicyExpectedReturn_eq_optimal_sub_behavior","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_stoppedSuccessorPolicyExpectedReturn_eq_optimal_sub_behavior","description":"The expectation of a stopped literal policy return is the optimal constant minus the expected behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-fdafff927ccd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8684,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:354"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_stoppedSuccessorPolicyExpectedReturn_eq_optimal_sub_behavior {Omega : Type w} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (optimal : Real) (policyReturn behaviorRegret : Omega -> Real) (hpoint : forall omega, policyReturn omega = optimal - behaviorRegret omega) (hbehavior : Integrable behaviorRegret mu) : integral mu policyReturn = optimal - integral mu behaviorRegret","missing":[],"search":"integral_stoppedsuccessorpolicyexpectedreturn_eq_optimal_sub_behavior banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_stoppedsuccessorpolicyexpectedreturn_eq_optimal_sub_behavior the expectation of a stopped literal policy return is the optimal constant minus the expected behavior regret. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturn_sub_successorPolicyExpectedReturn_eq_returnDeviation","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturn_sub_successorPolicyExpectedReturn_eq_returnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturn_sub_successorPolicyExpectedReturn_eq_returnDeviation","description":"At any common stopping prefix, observed sampled return minus the literal successor-policy expected return is exactly the normalized return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-921f2f08f0f4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8685,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:369"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturn_sub_successorPolicyExpectedReturn_eq_returnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex trajectory - selfConsistentScheduledNaturalCausalStoppingTimeAverageSucc…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturn_sub_successorpolicyexpectedreturn_eq_returndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragesampledreturn_sub_successorpolicyexpectedreturn_eq_returndeviation at any common stopping prefix, observed sampled return minus the literal successor-policy expected return is exactly the normalized return deviation. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","description":"Every capped stopped literal successor-policy expected-return coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-1ef7db90c1e3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8686,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:400"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturn every capped stopped literal successor-policy expected-return coordinate is measurable. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","description":"Every genuine uncapped stopped literal successor-policy expected-return coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-3f0cf0e2bde2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8687,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:430"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturn every genuine uncapped stopped literal successor-policy expected-return coordinate is measurable. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","label":"integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","description":"The capped stopped literal successor-policy expected return is integrable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-174af84dc05a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8688,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:456"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let policyReturn := selfConsistentScheduledNaturalCausalStoppingT…","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturn the capped stopped literal successor-policy expected return is integrable. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","label":"integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","description":"The genuine uncapped stopped literal successor-policy expected return is integrable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-583be539e52b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8689,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:511"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let policyReturn := selfConsistentScheduledNaturalCausalStoppingTime…","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturn banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturn the genuine uncapped stopped literal successor-policy expected return is integrable. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","description":"For the capped prefix, expected literal policy return is exactly optimal value minus expected behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-38dffeb13516","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8690,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:566"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let policyReturn := selfConsistentSc…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturnintegral_eq_optimal_sub_behaviorexpected banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturnintegral_eq_optimal_sub_behaviorexpected for the capped prefix, expected literal policy return is exactly optimal value minus expected behavior regret. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","description":"For the genuine uncapped prefix, expected literal policy return is exactly optimal value minus expected behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-c74fa1e5c61c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8691,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:616"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let policyReturn := selfConsistentSched…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturnintegral_eq_optimal_sub_behaviorexpected banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturnintegral_eq_optimal_sub_behaviorexpected for the genuine uncapped prefix, expected literal policy return is exactly optimal value minus expected behavior regret. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","description":"At the capped prefix, the expected sampled/policy-return gap is exactly the expected return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-3bcbe031574f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8692,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:666"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfC…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnintegral_sub_successorpolicyexpectedreturnintegral_eq_returndeviationintegral banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnintegral_sub_successorpolicyexpectedreturnintegral_eq_returndeviationintegral at the capped prefix, the expected sampled/policy-return gap is exactly the expected return deviation. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","description":"At the genuine uncapped prefix, the expected sampled/policy-return gap is exactly the expected return deviation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-04f0f34e1f4b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8693,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:725"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnintegral_sub_successorpolicyexpectedreturnintegral_eq_returndeviationintegral banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnintegral_sub_successorpolicyexpectedreturnintegral_eq_returndeviationintegral at the genuine uncapped prefix, the expected sampled/policy-return gap is exactly the expected return deviation. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","description":"The capped expected sampled/policy-return gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-807b08309e75","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8694,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:783"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState reward…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnsuccessorpolicyexpectedreturngap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnsuccessorpolicyexpectedreturngap_tendsto_zero the capped expected sampled/policy-return gap tends to zero. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","description":"The genuine uncapped expected sampled/policy-return gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-2ee9e6903b18","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8695,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:859"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp init…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnsuccessorpolicyexpectedreturngap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnsuccessorpolicyexpectedreturngap_tendsto_zero the genuine uncapped expected sampled/policy-return gap tends to zero. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","description":"The absolute capped expected sampled/policy-return gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-86f5f7c89a3a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8696,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:935"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rew…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnsuccessorpolicyexpectedreturnabsgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesampledreturnsuccessorpolicyexpectedreturnabsgap_tendsto_zero the absolute capped expected sampled/policy-return gap tends to zero. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","description":"The absolute genuine uncapped expected sampled/policy-return gap tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-d6e18be86991","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8697,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:976"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp i…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnsuccessorpolicyexpectedreturnabsgap_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesampledreturnsuccessorpolicyexpectedreturnabsgap_tendsto_zero the absolute genuine uncapped expected sampled/policy-return gap tends to zero. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","description":"Expected literal successor-policy return at the capped prefix converges to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-5796d29528f4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8698,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:1017"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSourc…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturnintegral_tendsto_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedstoppedaveragesuccessorpolicyexpectedreturnintegral_tendsto_optimal expected literal successor-policy return at the capped prefix converges to the optimal initial expected return. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","description":"Expected literal successor-policy return at the genuine uncapped prefix converges to the optimal initial expected return.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-6d4ce03631bd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8699,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:1084"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialSt…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturnintegral_tendsto_optimal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragesuccessorpolicyexpectedreturnintegral_tendsto_optimal expected literal successor-policy return at the genuine uncapped prefix converges to the optimal initial expected return. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_consistency","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_consistency","description":"Terminal literal policy-return semantics and stopped expected-consistency package for the capped approximation and genuine uncapped stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdcappedunboundedhitti-70effcc712a4/index.html#decl-208d36495dd3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","order":8700,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency.lean:1151"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp ini…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_consistency terminal literal policy-return semantics and stopped expected-consistency package for the capped approximation and genuine uncapped stopping prefix. theorem compiled","shard":"modules/a7011fb34dc315d6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","description":"Uncapped first passage below the inverse-square-root threshold after the fourth-power scheduled base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-9822b4408f41","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8701,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix uncapped first passage below the inverse-square-root threshold after the fourth-power scheduled base. definition compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_lower","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_lower","description":"Uncapped first passage cannot precede its scheduled base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-285c511b22fe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8702,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_lower (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : (explicitHighProbabilityRounds scheduleIndex : WithTop Nat) <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_lower banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_lower uncapped first passage cannot precede its scheduled base. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base_of_process_le","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base_of_process_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base_of_process_le","description":"If the process already lies below threshold at the base, uncapped first passage stops exactly at the base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-f9af30b72e29","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8703,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base_of_process_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) (hprocess : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (explicitHighProbabilityRounds scheduleIndex) trajectory <= selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleInde…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_eq_base_of_process_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_eq_base_of_process_le if the process already lies below threshold at the base, uncapped first passage stops exactly at the base. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_untopA_unboundedHittingAfter_le_threshold","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_untopA_unboundedHittingAfter_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_untopA_unboundedHittingAfter_le_threshold","description":"A finite uncapped first passage really lands in the inverse-square-root lower interval.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-29f9ff03ee4a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8704,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:109"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_untopA_unboundedHittingAfter_le_threshold (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) (hfinite : selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory ≠ ⊤) : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_untopa_unboundedhittingafter_le_threshold banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_untopa_unboundedhittingafter_le_threshold a finite uncapped first passage really lands in the inverse-square-root lower interval. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_isStoppingTime","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_isStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_isStoppingTime","description":"Every schedule-indexed uncapped inverse-square-root first passage is a stopping time for the exact natural causal filtration.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-0791ab9d64a8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8705,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:166"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_isStoppingTime (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex)","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_isstoppingtime banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_isstoppingtime every schedule-indexed uncapped inverse-square-root first passage is a stopping time for the exact natural causal filtration. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","label":"ae_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","description":"For every fixed schedule index, all-prefix almost-sure convergence forces the uncapped inverse-square-root first passage to be finite almost surely.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-287cef9d279d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8706,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource init…","missing":[],"search":"ae_ne_top_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_ne_top_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix for every fixed schedule index, all-prefix almost-sure convergence forces the uncapped inverse-square-root first passage to be finite almost surely. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_all_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","label":"ae_all_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_all_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","description":"Countability places all fixed-index a.e. finiteness statements on one common almost-sure set.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-57185b8c15c0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8707,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_all_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultSt…","missing":[],"search":"ae_all_ne_top_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_all_ne_top_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix countability places all fixed-index a.e. finiteness statements on one common almost-sure set. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base","label":"ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base","description":"Summable inverse-square-root delay forces the genuine uncapped first passage to hit immediately at the base for all sufficiently large schedule indices, almost surely.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-f236144572cd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8708,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable d…","missing":[],"search":"ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_eq_base banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_eventually_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_eq_base summable inverse-square-root delay forces the genuine uncapped first passage to hit immediately at the base for all sufficiently large schedule indices, almost surely. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_atTop","label":"ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_atTop","description":"Eventual exact-base equality makes the uncapped stopped prefixes diverge after applying Mathlib's `WithTop.untopA`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-31fb1a79ec35","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8709,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_atTop (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable…","missing":[],"search":"ae_tendsto_untopa_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_attop banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.ae_tendsto_untopa_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingprefix_attop eventual exact-base equality makes the uncapped stopped prefixes diverge after applying mathlib's `withtop.untopa`. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoAlmostEverywhere_zero","description":"The genuine uncapped inverse-square-root first-passage stopped process is measurable and converges almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-83acd93ab525","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8710,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:371"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initia…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedprocess_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedprocess_tendstoalmosteverywhere_zero the genuine uncapped inverse-square-root first-passage stopped process is measurable and converges almost everywhere. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoInMeasure_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoInMeasure_zero","description":"Almost-everywhere convergence of the measurable uncapped stopped process implies convergence in measure under the generated probability law.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-ac64859c6952","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8711,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:427"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedprocess_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedprocess_tendstoinmeasure_zero almost-everywhere convergence of the measurable uncapped stopped process implies convergence in measure under the generated probability law. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_aeFinite_eventuallyImmediateStopping_and_inMeasure_consistency","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_aeFinite_eventuallyImmediateStopping_and_inMeasure_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_aeFinite_eventuallyImmediateStopping_and_inMeasure_consistency","description":"Complete uncapped inverse-square-root first-passage a.e.-finiteness and stopped-process consistency package.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-654dbee51b26/index.html#decl-e7d50ce3296d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","order":8712,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency.lean:465"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_aeFinite_eventuallyImmediateStopping_and_inMeasure_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let threshold := selfConsistentScheduledNaturalCausalInverseSqrtFir…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_aefinite_eventuallyimmediatestopping_and_inmeasure_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_aefinite_eventuallyimmediatestopping_and_inmeasure_consistency complete uncapped inverse-square-root first-passage a.e.-finiteness and stopped-process consistency package. theorem compiled","shard":"modules/dccd43c981ca7bb6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale","label":"inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale","description":"Real degree-four comparison scale for the threshold schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-11b01364e489","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8713,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale (scheduleIndex : Nat) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdegreefourscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdegreefourscale real degree-four comparison scale for the threshold schedule. definition compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale_nonneg","label":"inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale_nonneg","description":"theorem inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale_nonneg (scheduleIndex : Nat) : 0 <= inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale scheduleIndex","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-c40cc8a0d62e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8714,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale_nonneg (scheduleIndex : Nat) : 0 <= inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdegreefourscale_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdegreefourscale_nonneg theorem inversesqrtthresholdunboundedhittingafterdegreefourscale_nonneg (scheduleindex : nat) : 0 <= inversesqrtthresholdunboundedhittingafterdegreefourscale scheduleindex theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","label":"sqrt_inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","description":"The square root of the degree-eight scale is exactly the degree-four scale, with no asymptotic slack.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-9f357f488628","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8715,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sqrt_inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale (scheduleIndex : Nat) : Real.sqrt (inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex) = inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale scheduleIndex","missing":[],"search":"sqrt_inversesqrtthresholdunboundedhittingafterdegreeeightscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.sqrt_inversesqrtthresholdunboundedhittingafterdegreeeightscale the square root of the degree-eight scale is exactly the degree-four scale, with no asymptotic slack. theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeFour","label":"sqrt_inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeFour","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeFour","description":"Taking square roots of the accepted degree-eight polynomial moment budget yields a degree-four asymptotic envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-3ad34be130c5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8716,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sqrt_inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeFour (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => Real.sqrt (inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex)) =O[atTop] inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale","missing":[],"search":"sqrt_inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_isbigo_degreefour banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.sqrt_inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_isbigo_degreefour taking square roots of the accepted degree-eight polynomial moment budget yields a degree-four asymptotic envelope. theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget","description":"Cauchy--Schwarz polynomial absolute-first-moment budget. Its only varying factor is the square root of the accepted stopping-round second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-74265e6c8d1a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8717,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget cauchy--schwarz polynomial absolute-first-moment budget. its only varying factor is the square root of the accepted stopping-round second-moment budget. definition compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_nonneg","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-9b2903a0da9d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8718,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget_nonneg theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget_nonneg (mdp : mdp state action) (varianceproxy : nnreal) (scheduleindex : nat) : 0 <= selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget mdp varianceproxy scheduleindex theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_isBigO_degreeFour","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_isBigO_degreeFour","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_isBigO_degreeFour","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_isBigO_degreeFour (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex) =O[atTop] inverseSqrtThresho…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-976534e7c802","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8719,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:122"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_isBigO_degreeFour (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex) =O[atTop] inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget_isbigo_degreefour banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget_isbigo_degreefour theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget_isbigo_degreefour (mdp : mdp state action) (varianceproxy : nnreal) : (fun scheduleindex : nat => selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingaftercauchyschwarzpolynomialabsolutefirstmomentbudget mdp varianceproxy scheduleindex) =o[attop] inversesqrtthresholdunboundedhittingafterdegreefourscale theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_cauchySchwarzPolynomialBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_cauchySchwarzPolynomialBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_cauchySchwarzPolynomialBudget","description":"For each fixed threshold index, Cauchy--Schwarz controls the actual stopped average realized behavior regret by the square root of the accepted polynomial stopping-round second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-6dc146eabd24","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8720,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:145"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_cauchySchwarzPolynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_cauchyschwarzpolynomialbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_cauchyschwarzpolynomialbudget for each fixed threshold index, cauchy--schwarz controls the actual stopped average realized behavior regret by the square root of the accepted polynomial stopping-round second-moment budget. theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_cauchySchwarzPolynomialBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_cauchySchwarzPolynomialBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_cauchySchwarzPolynomialBudget","description":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_cauchySchwarzPolynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varia…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-01125f501f13","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8721,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:248"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_cauchySchwarzPolynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : selfConsistentSchedule…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_le_cauchyschwarzpolynomialbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_le_cauchyschwarzpolynomialbudget theorem selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_le_cauchyschwarzpolynomialbudget (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (scheduleindex : nat) : selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret…","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeFour","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeFour","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeFour","description":"With all model/source parameters fixed, the actual expected absolute stopped regret grows at most at the Cauchy--Schwarz degree-four rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-dc3af8200013/index.html#decl-0a872534464a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","order":8722,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeFour (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUn…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_isbigo_degreefour banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_isbigo_degreefour with all model/source parameters fixed, the actual expected absolute stopped regret grows at most at the cauchy--schwarz degree-four rate. theorem compiled","shard":"modules/5256cac1c538aadf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tendsto_integral_abs_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","label":"tendsto_integral_abs_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tendsto_integral_abs_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","description":"A uniformly integrable Real family has vanishing restricted L1 mass on a measurable sequence of events whose measures tend to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-b0637306eb5b/index.html#decl-15d2cf2bf643","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","order":8723,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tendsto_integral_abs_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero {Omega : Type w} [MeasurableSpace Omega] {mu : Measure Omega} {f : Nat -> Omega -> Real} {event : Nat -> Set Omega} (hui : UniformIntegrable f 1 mu) (hevent : forall n, MeasurableSet (event n)) (hmeasure : Tendsto (fun n => mu (event n)) atTop (nhds 0)) : Tendsto (fun n => integral (mu.restrict (event n)) (fun omega => |f n omega|)) atTop (nhds 0)","missing":[],"search":"tendsto_integral_abs_restrict_of_uniformintegrable_one_of_measure_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tendsto_integral_abs_restrict_of_uniformintegrable_one_of_measure_tendsto_zero a uniformly integrable real family has vanishing restricted l1 mass on a measurable sequence of events whose measures tend to zero. theorem compiled","shard":"modules/ca4515a29f24abbf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tendsto_abs_integral_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","label":"tendsto_abs_integral_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tendsto_abs_integral_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","description":"The signed restricted integrals also vanish, by domination with the restricted integral of the absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-b0637306eb5b/index.html#decl-2c5f1eac787e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","order":8724,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tendsto_abs_integral_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero {Omega : Type w} [MeasurableSpace Omega] {mu : Measure Omega} {f : Nat -> Omega -> Real} {event : Nat -> Set Omega} (hui : UniformIntegrable f 1 mu) (hevent : forall n, MeasurableSet (event n)) (hmeasure : Tendsto (fun n => mu (event n)) atTop (nhds 0)) : Tendsto (fun n => |integral (mu.restrict (event n)) (f n)|) atTop (nhds 0)","missing":[],"search":"tendsto_abs_integral_restrict_of_uniformintegrable_one_of_measure_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tendsto_abs_integral_restrict_of_uniformintegrable_one_of_measure_tendsto_zero the signed restricted integrals also vanish, by domination with the restricted integral of the absolute value. theorem compiled","shard":"modules/ca4515a29f24abbf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedEvent_expectedContribution_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedEvent_expectedContribution_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedEvent_expectedContribution_tendsto_zero","description":"The existing inverse-square-root delayed event has vanishing expected contribution for the exact uncapped `hittingAfter` stopped average realized behavior-regret process. The event comes from the capped first-passage route, but the integrated process and stopping prefix remain genuinely uncapped.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-b0637306eb5b/index.html#decl-9775e1d39d7d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","order":8725,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedEvent_expectedContribution_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource ini…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_delayedevent_expectedcontribution_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_delayedevent_expectedcontribution_tendsto_zero the existing inverse-square-root delayed event has vanishing expected contribution for the exact uncapped `hittingafter` stopped average realized behavior-regret process. the event comes from the capped first-passage route, but the integrated process and stopping prefix remain genuinely uncapped. theorem compiled","shard":"modules/ca4515a29f24abbf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart","description":"Expected positive part of the exact stopped average realized behavior regret as a function of the threshold schedule index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-d9d3109ce9f7/index.html#decl-3d2dfe2c3a6f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","order":8726,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedpositivepart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedpositivepart expected positive part of the exact stopped average realized behavior regret as a function of the threshold schedule index. definition compiled","shard":"modules/e656c5924aebfa74.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart_nonneg","description":"The expected positive part of the exact stopped process is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-d9d3109ce9f7/index.html#decl-939c37335f07","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","order":8727,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedpositivepart_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedpositivepart_nonneg the expected positive part of the exact stopped process is nonnegative. theorem compiled","shard":"modules/e656c5924aebfa74.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretPositivePart_integrable_and_integral_le_threshold","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretPositivePart_integrable_and_integral_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretPositivePart_integrable_and_integral_le_threshold","description":"At every fixed threshold index, the positive part of the exact stopped average realized behavior regret is integrable and its expectation is bounded by the inverse-square-root hit threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-d9d3109ce9f7/index.html#decl-4ef7bc795622","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","order":8728,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretPositivePart_integrable_and_integral_le_threshold (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfCons…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretpositivepart_integrable_and_integral_le_threshold banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretpositivepart_integrable_and_integral_le_threshold at every fixed threshold index, the positive part of the exact stopped average realized behavior regret is integrable and its expectation is bounded by the inverse-square-root hit threshold. theorem compiled","shard":"modules/e656c5924aebfa74.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedPositivePart_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedPositivePart_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedPositivePart_tendsto_zero","description":"The expected positive part of the exact stopped average realized behavior regret tends to zero with the inverse-square-root threshold schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-d9d3109ce9f7/index.html#decl-f6ebd2c0acde","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","order":8729,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedPositivePart_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThre…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedpositivepart_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedpositivepart_tendsto_zero the expected positive part of the exact stopped average realized behavior regret tends to zero with the inverse-square-root threshold schedule. theorem compiled","shard":"modules/e656c5924aebfa74.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient","description":"Coefficient of a reciprocal-linear envelope for the exact scheduled realized-regret rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-6015b10498c0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8730,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient (mdp : MDP State Action) (varianceProxy : NNReal) : Real","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretratelinearcoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretratelinearcoefficient coefficient of a reciprocal-linear envelope for the exact scheduled realized-regret rate. definition compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient_nonneg","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient_nonneg","description":"The reciprocal-linear rate coefficient is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-251301793efe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8731,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient mdp varianceProxy","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretratelinearcoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretratelinearcoefficient_nonneg the reciprocal-linear rate coefficient is nonnegative. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_le_linearEnvelope","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_le_linearEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_le_linearEnvelope","description":"The exact scheduled realized-regret rate is controlled by a reciprocal linear envelope in the fourth-power schedule scale.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-24cedf33cb4f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8732,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_le_linearEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : explicitPolynomialPrefixAverageRealizedBehaviorRegretRate mdp varianceProxy baseVisitFloor n <= explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient mdp varianceProxy / (explicitHighProbabilityScale n : Real)","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretrate_le_linearenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretrate_le_linearenvelope the exact scheduled realized-regret rate is controlled by a reciprocal linear envelope in the fourth-power schedule scale. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart","description":"Explicit checkpoint index that clears the fixed inverse-square-root threshold for every later scheduled regret-rate checkpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-bcf9f272d938","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8733,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Nat","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicittailstart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicittailstart explicit checkpoint index that clears the fixed inverse-square-root threshold for every later scheduled regret-rate checkpoint. definition compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_spec","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_spec","description":"The explicit ceiling witness is beyond the threshold index and validates the exact scheduled rate comparison at every later checkpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-acd07b72fb7b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8734,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_spec (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart mdp varianceProxy scheduleIndex /\\ forall n, inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart mdp varianceProxy scheduleIndex <= n -> explicitPolynomialPrefixAverageRealizedBehaviorRegretRate mdp varianceProxy baseVisitFloor n <= selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicittailstart_spec banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicittailstart_spec the explicit ceiling witness is beyond the threshold index and validates the exact scheduled rate comparison at every later checkpoint. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart_le_explicitTailStart","label":"inverseSqrtThresholdUnboundedHittingAfterTailStart_le_explicitTailStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart_le_explicitTailStart","description":"The accepted least eventual witness is no larger than the explicit ceiling-based tail start.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-bd1cd64fe859","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8735,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:245"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterTailStart_le_explicitTailStart (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : inverseSqrtThresholdUnboundedHittingAfterTailStart mdp varianceProxy baseVisitFloor scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingaftertailstart_le_explicittailstart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingaftertailstart_le_explicittailstart the accepted least eventual witness is no larger than the explicit ceiling-based tail start. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget","description":"ENNReal second-moment budget obtained by replacing the canonical tail start with its explicit ceiling-based upper bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-3320d857e496","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8736,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:262"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : ENNReal","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentennrealbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentennrealbudget ennreal second-moment budget obtained by replacing the canonical tail start with its explicit ceiling-based upper bound. definition compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_ne_top","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_ne_top","description":"The explicit-start ENNReal budget is finite under the same horizon-five contract as the accepted canonical budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-e24a6fe5336a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8737,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_ne_top (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) (hhorizon : 4 < mdp.horizon) : inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget mdp varianceProxy scheduleIndex ≠ ∞","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentennrealbudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentennrealbudget_ne_top the explicit-start ennreal budget is finite under the same horizon-five contract as the accepted canonical budget. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget","description":"Real-valued form of the explicit-start stopping-round second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-6117d01c2576","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8738,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentbudget real-valued form of the explicit-start stopping-round second-moment budget. definition compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitStoppingRoundSecondMomentENNRealBudgetAt_mono","label":"explicitStoppingRoundSecondMomentENNRealBudgetAt_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitStoppingRoundSecondMomentENNRealBudgetAt_mono","description":"The checkpoint-square plus fixed weighted failure series is monotone in its deterministic checkpoint start.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-574d46445252","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8739,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:305"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitStoppingRoundSecondMomentENNRealBudgetAt_mono (mdp : MDP State Action) {left right : Nat} (h : left <= right) : (((explicitHighProbabilityRounds left + 1) ^ 2 : Nat) : ENNReal) + ∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * explicitPolynomialPrefixTailModelReturnFailureBudget mdp n <= (((explicitHighProbabilityRounds right + 1) ^ 2 : Nat) : ENNReal) + ∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * explicitPolynomialPrefixTailModelReturnFailureBudget mdp n","missing":[],"search":"explicitstoppingroundsecondmomentennrealbudgetat_mono banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitstoppingroundsecondmomentennrealbudgetat_mono the checkpoint-square plus fixed weighted failure series is monotone in its deterministic checkpoint start. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_le_explicit","label":"inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_le_explicit","description":"Replacing the least eventual witness by the explicit tail start can only increase the deterministic ENNReal second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-7a50fd604666","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8740,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:322"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_le_explicit (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget mdp varianceProxy baseVisitFloor scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentennrealbudget_le_explicit banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentennrealbudget_le_explicit replacing the least eventual witness by the explicit tail start can only increase the deterministic ennreal second-moment budget. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget_le_explicit","label":"inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget_le_explicit","description":"Real-valued canonical second-moment budget is bounded by the explicit tail-start budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-b40be85cc8d8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8741,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:339"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget_le_explicit (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (hhorizon : 4 < mdp.horizon) : inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget mdp varianceProxy baseVisitFloor scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentbudget_le_explicit banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentbudget_le_explicit real-valued canonical second-moment budget is bounded by the explicit tail-start budget. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_explicitTailStartBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_explicitTailStartBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_explicitTailStartBudget","description":"The actual successor stopping-round second moment is bounded by the explicit ceiling-based deterministic budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-37d783f19e4a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8742,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:359"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_explicitTailStartBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp i…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_le_explicittailstartbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_le_explicittailstartbudget the actual successor stopping-round second moment is bounded by the explicit ceiling-based deterministic budget. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","description":"Explicit-tail-start absolute-first-moment budget. Its public parameters contain no canonical `Nat.find` witness and no unevaluated random integral.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-9479b793e17c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8743,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:397"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterexplicittailstartdeterministicstoppingroundsecondmomentabsolutefirstmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterexplicittailstartdeterministicstoppingroundsecondmomentabsolutefirstmomentbudget explicit-tail-start absolute-first-moment budget. its public parameters contain no canonical `nat.find` witness and no unevaluated random integral. definition compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_explicitTailStartDeterministicStoppingRoundSecondMomentBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_explicitTailStartDeterministicStoppingRoundSecondMomentBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_explicitTailStartDeterministicStoppingRoundSecondMomentBudget","description":"For each fixed threshold index, the exact stopped average realized behavior regret is integrable and its absolute first moment is controlled by the explicit ceiling-based deterministic second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-54a382704d14/index.html#decl-cbe86c365f5c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","order":8744,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound.lean:413"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_explicitTailStartDeterministicStoppingRoundSecondMomentBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (s…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_explicittailstartdeterministicstoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_explicittailstartdeterministicstoppingroundsecondmomentbudget for each fixed threshold index, the exact stopped average realized behavior regret is integrable and its absolute first moment is controlled by the explicit ceiling-based deterministic second-moment budget. theorem compiled","shard":"modules/077479f4659a43a5.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope","description":"A round-independent second-moment envelope for every deterministic exact natural-causal average realized behavior-regret coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-bc1d0a453aca","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8745,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope (mdp : MDP State Action) (varianceProxy : NNReal) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretuniformsecondmomentenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretuniformsecondmomentenvelope a round-independent second-moment envelope for every deterministic exact natural-causal average realized behavior-regret coordinate. definition compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope_nonneg","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope_nonneg","description":"The uniform deterministic-coordinate second-moment envelope is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-bf653e7d7869","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8746,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope mdp varianceProxy","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretuniformsecondmomentenvelope_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretuniformsecondmomentenvelope_nonneg the uniform deterministic-coordinate second-moment envelope is nonnegative. theorem compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingFiberAbsoluteFirstMomentBudget","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingFiberAbsoluteFirstMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingFiberAbsoluteFirstMomentBudget","description":"The explicit fixed-index stopping-fiber budget used to control the absolute first moment at an uncapped inverse-sqrt hitting time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-1630f0f5cdbb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8747,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingFiberAbsoluteFirstMomentBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingfiberabsolutefirstmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingfiberabsolutefirstmomentbudget the explicit fixed-index stopping-fiber budget used to control the absolute first moment at an uncapped inverse-sqrt hitting time. definition compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentAbsoluteFirstMomentBudget","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentAbsoluteFirstMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentAbsoluteFirstMomentBudget","description":"The fixed-index absolute first-moment budget after eliminating the stopping-fiber sum in favor of the actual stopping-round second moment and the universal inverse-square series.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-49d171367162","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8748,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:82"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentAbsoluteFirstMomentBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentabsolutefirstmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentabsolutefirstmomentbudget the fixed-index absolute first-moment budget after eliminating the stopping-fiber sum in favor of the actual stopping-round second moment and the universal inverse-square series. definition compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","description":"Deterministic fixed-index absolute first-moment budget obtained by replacing the actual stopping-round second moment with the canonical checkpoint/failure-series upper bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-916bb54fcca7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8749,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdeterministicstoppingroundsecondmomentabsolutefirstmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdeterministicstoppingroundsecondmomentabsolutefirstmomentbudget deterministic fixed-index absolute first-moment budget obtained by replacing the actual stopping-round second moment with the canonical checkpoint/failure-series upper bound. definition compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_uniformSecondMomentEnvelope","label":"integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_uniformSecondMomentEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_uniformSecondMomentEnvelope","description":"Every deterministic coordinate has second moment bounded by one constant, including the zero-round coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-52734271287c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8750,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_uniformSecondMomentEnvelope (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) : integral (selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor).trajectoryMeasure (fun trajectory => selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceP…","missing":[],"search":"integral_sq_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_uniformsecondmomentenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_sq_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_uniformsecondmomentenvelope every deterministic coordinate has second moment bounded by one constant, including the zero-round coordinate. theorem compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","description":"For every fixed threshold index, the exact average realized behavior regret stopped at the genuine uncapped inverse-sqrt first passage is integrable and its expectation is at most the hit threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-2fb8d3f442f8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8751,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentSchedu…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_le_threshold banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_le_threshold for every fixed threshold index, the exact average realized behavior regret stopped at the genuine uncapped inverse-sqrt first passage is integrable and its expectation is at most the hit threshold. theorem compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingFiberBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingFiberBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingFiberBudget","description":"For each fixed threshold index, the exact stopped average realized behavior regret has an explicit absolute first-moment bound given by the summable square-root masses of the genuine stopping fibers. The bound is not uniform in the threshold index and does not use optional stopping.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-848daf0e43ff","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8752,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:330"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingFiberBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfCo…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_stoppingfiberbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_stoppingfiberbudget for each fixed threshold index, the exact stopped average realized behavior regret has an explicit absolute first-moment bound given by the summable square-root masses of the genuine stopping fibers. the bound is not uniform in the threshold index and does not use optional stopping. theorem compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingRoundSecondMomentBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingRoundSecondMomentBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingRoundSecondMomentBudget","description":"For each fixed threshold index, the exact stopped average realized behavior regret is integrable and its absolute first moment is controlled by the actual stopping-round second moment plus the universal inverse-square series. This is not a schedule-index-uniform or asymptotic estimate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-d90632155cc0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8753,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:426"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingRoundSecondMomentBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let sour…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_stoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_stoppingroundsecondmomentbudget for each fixed threshold index, the exact stopped average realized behavior regret is integrable and its absolute first moment is controlled by the actual stopping-round second moment plus the universal inverse-square series. this is not a schedule-index-uniform or asymptotic estimate. theorem compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_deterministicStoppingRoundSecondMomentBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_deterministicStoppingRoundSecondMomentBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_deterministicStoppingRoundSecondMomentBudget","description":"For each fixed threshold index, the exact stopped average realized behavior regret is integrable and its absolute first moment is controlled by a fully deterministic checkpoint/failure-series budget. The endpoint contains no unevaluated stopping-time integral.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-3752d32798f7/index.html#decl-4cb1eaa38e21","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","order":8754,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound.lean:522"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_deterministicStoppingRoundSecondMomentBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Na…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_deterministicstoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_deterministicstoppingroundsecondmomentbudget for each fixed threshold index, the exact stopped average realized behavior regret is integrable and its absolute first moment is controlled by a fully deterministic checkpoint/failure-series budget. the endpoint contains no unevaluated stopping-time integral. theorem compiled","shard":"modules/e5f6b467a20c8791.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityQuarticBlockWeight","label":"explicitHighProbabilityQuarticBlockWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityQuarticBlockWeight","description":"Cubic envelope for the width of one consecutive fourth-power checkpoint block.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-aac086a964d1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8755,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def explicitHighProbabilityQuarticBlockWeight (n : Nat) : Nat","missing":[],"search":"explicithighprobabilityquarticblockweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityquarticblockweight cubic envelope for the width of one consecutive fourth-power checkpoint block. definition compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_succ_sub_le_quarticBlockWeight","label":"explicitHighProbabilityRounds_succ_sub_le_quarticBlockWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_succ_sub_le_quarticBlockWeight","description":"A consecutive fourth-power checkpoint gap is bounded by the cubic block weight.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-ff3526d5efc2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8756,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_succ_sub_le_quarticBlockWeight (n : Nat) : explicitHighProbabilityRounds (n + 1) - explicitHighProbabilityRounds n <= explicitHighProbabilityQuarticBlockWeight n","missing":[],"search":"explicithighprobabilityrounds_succ_sub_le_quarticblockweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_succ_sub_le_quarticblockweight a consecutive fourth-power checkpoint gap is bounded by the cubic block weight. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_le_inv_pow_six","label":"selfConsistentScheduledLocalDelta_le_inv_pow_six","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_le_inv_pow_six","description":"Positive horizon makes every local confidence share no larger than a shifted inverse sixth power.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-cc995aa99c7d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8757,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledLocalDelta_le_inv_pow_six (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (t : Nat) : AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp t <= 1 / (((t + 2 : Nat) : Real) ^ 6)","missing":[],"search":"selfconsistentscheduledlocaldelta_le_inv_pow_six banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledlocaldelta_le_inv_pow_six positive horizon makes every local confidence share no larger than a shifted inverse sixth power. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedInverseCubePairEnvelope","label":"quarticBlockShiftedInverseCubePairEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedInverseCubePairEnvelope","description":"Shifted inverse-cube envelope used after the cubic block width cancels three powers.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-4454f0db2343","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8758,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:75"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def quarticBlockShiftedInverseCubePairEnvelope (p : Nat × Nat) : ENNReal","missing":[],"search":"quarticblockshiftedinversecubepairenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticblockshiftedinversecubepairenvelope shifted inverse-cube envelope used after the cubic block width cancels three powers. definition compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockShiftedInverseCubePairEnvelope_ne_top","label":"tsum_quarticBlockShiftedInverseCubePairEnvelope_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockShiftedInverseCubePairEnvelope_ne_top","description":"The shifted inverse-cube envelope is summable over checkpoint/tail-offset pairs.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-4c15896f6875","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8759,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:84"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticBlockShiftedInverseCubePairEnvelope_ne_top : Ne (∑' p : Nat × Nat, quarticBlockShiftedInverseCubePairEnvelope p) ∞","missing":[],"search":"tsum_quarticblockshiftedinversecubepairenvelope_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticblockshiftedinversecubepairenvelope_ne_top the shifted inverse-cube envelope is summable over checkpoint/tail-offset pairs. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedCoordinateModelFailureCharge","label":"quarticBlockShiftedCoordinateModelFailureCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedCoordinateModelFailureCharge","description":"One checkpoint block weight times one shifted coordinate model-failure charge.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-7ab44da1a43c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8760,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def quarticBlockShiftedCoordinateModelFailureCharge (mdp : MDP State Action) (p : Nat × Nat) : ENNReal","missing":[],"search":"quarticblockshiftedcoordinatemodelfailurecharge banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticblockshiftedcoordinatemodelfailurecharge one checkpoint block weight times one shifted coordinate model-failure charge. definition compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","label":"quarticBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","description":"Positive horizon makes the actual weighted coordinate charge fit the inverse-cube pair envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-d27e113bdb60","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8761,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem quarticBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) (p : Nat × Nat) : quarticBlockShiftedCoordinateModelFailureCharge mdp p <= quarticBlockShiftedInverseCubePairEnvelope p","missing":[],"search":"quarticblockshiftedcoordinatemodelfailurecharge_le_pairenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticblockshiftedcoordinatemodelfailurecharge_le_pairenvelope positive horizon makes the actual weighted coordinate charge fit the inverse-cube pair envelope. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockShiftedCoordinateModelFailureCharge_ne_top","label":"tsum_quarticBlockShiftedCoordinateModelFailureCharge_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockShiftedCoordinateModelFailureCharge_ne_top","description":"The actual weighted shifted coordinate charges have finite total ENNReal mass.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-95353cdb54ee","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8762,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticBlockShiftedCoordinateModelFailureCharge_ne_top (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) : Ne (∑' p : Nat × Nat, quarticBlockShiftedCoordinateModelFailureCharge mdp p) ∞","missing":[],"search":"tsum_quarticblockshiftedcoordinatemodelfailurecharge_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticblockshiftedcoordinatemodelfailurecharge_ne_top the actual weighted shifted coordinate charges have finite total ennreal mass. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","label":"quarticBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","description":"A model tail after `n+1`, charged by one quartic block weight, is the shifted coordinate row.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-2e26ca29dfc4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8763,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:273"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem quarticBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge (mdp : MDP State Action) (n : Nat) : (explicitHighProbabilityQuarticBlockWeight n : ENNReal) * selfConsistentScheduledCausalTailModelFailureBudget mdp (n + 1) = ∑' j : Nat, quarticBlockShiftedCoordinateModelFailureCharge mdp (n, j)","missing":[],"search":"quarticblockweight_mul_tailmodelfailurebudget_eq_tsum_shiftedcoordinatecharge banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticblockweight_mul_tailmodelfailurebudget_eq_tsum_shiftedcoordinatecharge a model tail after `n+1`, charged by one quartic block weight, is the shifted coordinate row. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_tailModelFailureBudget_ne_top","label":"tsum_quarticBlockWeight_mul_tailModelFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_tailModelFailureBudget_ne_top","description":"The fourth-power block weights are summable against the exact infinite model-tail budgets.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-7b6d12442a67","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8764,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:299"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticBlockWeight_mul_tailModelFailureBudget_ne_top (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) : Ne (∑' n : Nat, (explicitHighProbabilityQuarticBlockWeight n : ENNReal) * selfConsistentScheduledCausalTailModelFailureBudget mdp (n + 1)) ∞","missing":[],"search":"tsum_quarticblockweight_mul_tailmodelfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticblockweight_mul_tailmodelfailurebudget_ne_top the fourth-power block weights are summable against the exact infinite model-tail budgets. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta","label":"summable_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta","description":"A cubic fourth-power block weight times the explicit exponential return share is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-31e7a00cd046","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8765,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:321"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta : Summable (fun n : Nat => ((explicitHighProbabilityQuarticBlockWeight n : Nat) : Real) * explicitHighProbabilityReturnDelta n)","missing":[],"search":"summable_quarticblockweight_mul_explicithighprobabilityreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_quarticblockweight_mul_explicithighprobabilityreturndelta a cubic fourth-power block weight times the explicit exponential return share is summable. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","label":"tsum_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","description":"The exact exponential return shares have finite total mass after quartic block charging.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-b6f6e2cae23e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8766,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:358"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top : Ne (∑' n : Nat, (explicitHighProbabilityQuarticBlockWeight n : ENNReal) * ENNReal.ofReal (explicitHighProbabilityReturnDelta n)) ∞","missing":[],"search":"tsum_quarticblockweight_mul_explicithighprobabilityreturndelta_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticblockweight_mul_explicithighprobabilityreturndelta_ne_top the exact exponential return shares have finite total mass after quartic block charging. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","label":"tsum_quarticBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","description":"The exact checkpoint violation budgets remain summable after paying every quartic block width.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-132b5ff392ec","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8767,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:377"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top (mdp : MDP State Action) (hhorizon : 0 < mdp.horizon) : Ne (∑' n : Nat, (explicitHighProbabilityQuarticBlockWeight n : ENNReal) * explicitPolynomialPrefixTailModelReturnFailureBudget mdp n) ∞","missing":[],"search":"tsum_quarticblockweight_mul_explicitpolynomialprefixtailmodelreturnfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticblockweight_mul_explicitpolynomialprefixtailmodelreturnfailurebudget_ne_top the exact checkpoint violation budgets remain summable after paying every quartic block width. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_add_le_add_sum_Ico_quarticBlockWeight","label":"explicitHighProbabilityRounds_add_le_add_sum_Ico_quarticBlockWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_add_le_add_sum_Ico_quarticBlockWeight","description":"Cubic block weights dominate every finite telescope of the fourth-power checkpoint grid.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-c178ba2dc9b0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8768,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:397"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_add_le_add_sum_Ico_quarticBlockWeight (start width : Nat) : explicitHighProbabilityRounds (start + width) <= explicitHighProbabilityRounds start + (Finset.Ico start (start + width)).sum explicitHighProbabilityQuarticBlockWeight","missing":[],"search":"explicithighprobabilityrounds_add_le_add_sum_ico_quarticblockweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_add_le_add_sum_ico_quarticblockweight cubic block weights dominate every finite telescope of the fourth-power checkpoint grid. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.natCast_succ_le_checkpoint_add_tsum_quarticBlockWeight_of_delayed","label":"natCast_succ_le_checkpoint_add_tsum_quarticBlockWeight_of_delayed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.natCast_succ_le_checkpoint_add_tsum_quarticBlockWeight_of_delayed","description":"A finite natural time is paid for by an initial checkpoint plus all preceding delayed blocks.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-c2bb800da7cd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8769,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:440"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem natCast_succ_le_checkpoint_add_tsum_quarticBlockWeight_of_delayed (start time : Nat) : (time + 1 : ENNReal) <= (explicitHighProbabilityRounds start + 1 : ENNReal) + ∑' n : Nat, if start <= n ∧ explicitHighProbabilityRounds n < time then (explicitHighProbabilityQuarticBlockWeight n : ENNReal) else 0","missing":[],"search":"natcast_succ_le_checkpoint_add_tsum_quarticblockweight_of_delayed banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.natcast_succ_le_checkpoint_add_tsum_quarticblockweight_of_delayed a finite natural time is paid for by an initial checkpoint plus all preceding delayed blocks. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_untopA_add_one_of_eventually_quarticCheckpointTail","label":"integrable_untopA_add_one_of_eventually_quarticCheckpointTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_untopA_add_one_of_eventually_quarticCheckpointTail","description":"Eventually summable fourth-power checkpoint crossing tails imply a finite first moment.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-30b1b1f990ed","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8770,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:539"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_untopA_add_one_of_eventually_quarticCheckpointTail {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (budget : Nat -> ENNReal) (hbudget : Ne (∑' n : Nat, (explicitHighProbabilityQuarticBlockWeight n : ENNReal) * budget n) ∞) (htail : ∀ᶠ n : Nat in atTop, mu {omega | (explicitHighProbabilityRounds n : WithTop Nat) < tau omega} <= budget n) : Integrable (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) mu","missing":[],"search":"integrable_untopa_add_one_of_eventually_quarticcheckpointtail banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_untopa_add_one_of_eventually_quarticcheckpointtail eventually summable fourth-power checkpoint crossing tails imply a finite first moment. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","description":"Checkpoint trajectories whose fixed-index uncapped first passage has not yet occurred.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-7d44bf38d42c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8771,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:646"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex checkpointIndex : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedcheckpointset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedcheckpointset checkpoint trajectories whose fixed-index uncapped first passage has not yet occurred. definition compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","label":"measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","description":"Every fixed delayed-checkpoint event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-3d2ea914d6bd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8772,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:665"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex checkpointIndex : Nat) : MeasurableSet (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex checkpointIndex)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedcheckpointset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedcheckpointset every fixed delayed-checkpoint event is measurable. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_subset_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_subset_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_subset_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","description":"Once a checkpoint lies beyond the fixed base and its deterministic rate is below the fixed threshold, every delayed first passage violates the compiled checkpoint regret bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-4e68a8589071","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8773,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:687"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_subset_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) {scheduleIndex checkpointIndex : Nat} (hindex : scheduleIndex <= checkpointIndex) (hrate : explicitPolynomialPrefixAverageRealizedBehaviorRegretRate mdp varianceProxy baseVisitFloor checkpointIndex <= selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet mdp initialState rewardSource initialTable defaultState varianceProxy b…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedcheckpointset_subset_explicitpolynomialprefixaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedcheckpointset_subset_explicitpolynomialprefixaveragerealizedbehaviorregretviolationset once a checkpoint lies beyond the fixed base and its deterministic rate is below the fixed threshold, every delayed first passage violates the compiled checkpoint regret bound. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le","label":"eventually_selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le","description":"For each fixed threshold index, checkpoint crossing tails are eventually bounded by the explicit summable burn-in-tail/model-return budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-57eb46e916e6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8774,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:741"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eventually_selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp in…","missing":[],"search":"eventually_selfconsistentscheduledcausalsource_trajectorymeasure_inversesqrtthresholdunboundedhittingafterdelayedcheckpointset_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.eventually_selfconsistentscheduledcausalsource_trajectorymeasure_inversesqrtthresholdunboundedhittingafterdelayedcheckpointset_le for each fixed threshold index, checkpoint crossing tails are eventually bounded by the explicit summable burn-in-tail/model-return budget. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integrableFiniteStoppingTime","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integrableFiniteStoppingTime","description":"Each fixed-index genuine uncapped inverse-square-root first passage is finite almost surely and has an integrable random horizon.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-84e1e76d235b/index.html#decl-7684c39a4cb4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","order":8775,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime.lean:791"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integrableFiniteStoppingTime (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integrablefinitestoppingtime banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integrablefinitestoppingtime each fixed-index genuine uncapped inverse-square-root first passage is finite almost surely and has an integrable random horizon. theorem compiled","shard":"modules/f0bba0fa4b54c7a3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret_succ_eq","label":"naturalAverageRealizedBehaviorRegret_succ_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret_succ_eq","description":"Exact one-step recursion for the natural round-average realized regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-698516ceb787","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8776,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAverageRealizedBehaviorRegret_succ_eq {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (n : Nat) : source.naturalAverageRealizedBehaviorRegret trajectory (n + 1) = ((n : Real) * source.naturalAverageRealizedBehaviorRegret trajectory n + source.naturalSuccessorBatchAverageRealizedRegret trajectory n) / ((n + 1 : Nat) : Real)","missing":[],"search":"naturalaveragerealizedbehaviorregret_succ_eq banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragerealizedbehaviorregret_succ_eq exact one-step recursion for the natural round-average realized regret. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalSuccessorAverageReturnDeviationIncrement_succ_hasSubgaussianMGF","label":"trajectoryMeasure_naturalSuccessorAverageReturnDeviationIncrement_succ_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalSuccessorAverageReturnDeviationIncrement_succ_hasSubgaussianMGF","description":"A conditional successor-average return MGF also gives its unconditional sub-Gaussian MGF on the trajectory measure.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-d9d742e1f433","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8777,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:63"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_naturalSuccessorAverageReturnDeviationIncrement_succ_hasSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : s…","missing":[],"search":"trajectorymeasure_naturalsuccessoraveragereturndeviationincrement_succ_hassubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_naturalsuccessoraveragereturndeviationincrement_succ_hassubgaussianmgf a conditional successor-average return mgf also gives its unconditional sub-gaussian mgf on the trajectory measure. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnVarianceProxyAt_succ_le_globalReturnDeviationPerEpisodeVarianceProxy","label":"naturalSuccessorAverageReturnVarianceProxyAt_succ_le_globalReturnDeviationPerEpisodeVarianceProxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnVarianceProxyAt_succ_le_globalReturnDeviationPerEpisodeVarianceProxy","description":"A positive successor batch contributes at most one one-episode return variance proxy after normalization by its batch size.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-db3f8f98ec5f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8778,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalSuccessorAverageReturnVarianceProxyAt_succ_le_globalReturnDeviationPerEpisodeVarianceProxy (mdp : MDP State Action) (episodes : Nat -> Nat) (n : Nat) (rewardBound rewardVarianceProxy : NNReal) (hepisodes : 0 < episodes (n + 1)) : naturalSuccessorAverageReturnVarianceProxyAt mdp episodes (n + 1) rewardBound rewardVarianceProxy <= mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy","missing":[],"search":"naturalsuccessoraveragereturnvarianceproxyat_succ_le_globalreturndeviationperepisodevarianceproxy banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessoraveragereturnvarianceproxyat_succ_le_globalreturndeviationperepisodevarianceproxy a positive successor batch contributes at most one one-episode return variance proxy after normalization by its batch size. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegretSecondMomentEnvelope","label":"naturalSuccessorBatchAverageRealizedRegretSecondMomentEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegretSecondMomentEnvelope","description":"Uniform deterministic second-moment envelope for one successor-batch average realized-regret coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-19d0b62ca21a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8779,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalSuccessorBatchAverageRealizedRegretSecondMomentEnvelope (mdp : MDP State Action) (rewardVarianceProxy : NNReal) : Real","missing":[],"search":"naturalsuccessorbatchaveragerealizedregretsecondmomentenvelope banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessorbatchaveragerealizedregretsecondmomentenvelope uniform deterministic second-moment envelope for one successor-batch average realized-regret coordinate. definition compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.memLp_two_naturalSuccessorBatchAverageRealizedRegret","label":"memLp_two_naturalSuccessorBatchAverageRealizedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.memLp_two_naturalSuccessorBatchAverageRealizedRegret","description":"Every positive-count successor-batch average realized-regret coordinate belongs to `L2`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-c8ecfe9e6040","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8780,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_two_naturalSuccessorBatchAverageRealizedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (law : source.rewardSource.UniformSubgaussianRewardLaw re…","missing":[],"search":"memlp_two_naturalsuccessorbatchaveragerealizedregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.memlp_two_naturalsuccessorbatchaveragerealizedregret every positive-count successor-batch average realized-regret coordinate belongs to `l2`. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.integral_sq_naturalSuccessorBatchAverageRealizedRegret_le_secondMomentEnvelope","label":"integral_sq_naturalSuccessorBatchAverageRealizedRegret_le_secondMomentEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.integral_sq_naturalSuccessorBatchAverageRealizedRegret_le_secondMomentEnvelope","description":"The second moment of every positive-count successor-batch average realized-regret coordinate is bounded by the uniform envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-b54c07502b2c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8781,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:214"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sq_naturalSuccessorBatchAverageRealizedRegret_le_secondMomentEnvelope {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (law : source.rewardSource.Unif…","missing":[],"search":"integral_sq_naturalsuccessorbatchaveragerealizedregret_le_secondmomentenvelope banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.integral_sq_naturalsuccessorbatchaveragerealizedregret_le_secondmomentenvelope the second moment of every positive-count successor-batch average realized-regret coordinate is bounded by the uniform envelope. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.max_neg_natWeightedAverage_succ_le_abs_increment_div","label":"max_neg_natWeightedAverage_succ_le_abs_increment_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.max_neg_natWeightedAverage_succ_le_abs_increment_div","description":"If the previous average is nonnegative, one new coordinate can create at most its absolute value divided by the new sample count as negative overshoot.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-12eb0d4fea78","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8782,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:334"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem max_neg_natWeightedAverage_succ_le_abs_increment_div (n : Nat) (previous increment : Real) (hprevious : 0 <= previous) : max (-(((n : Real) * previous + increment) / ((n + 1 : Nat) : Real))) 0 <= |increment| / ((n + 1 : Nat) : Real)","missing":[],"search":"max_neg_natweightedaverage_succ_le_abs_increment_div banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.max_neg_natweightedaverage_succ_le_abs_increment_div if the previous average is nonnegative, one new coordinate can create at most its absolute value divided by the new sample count as negative overshoot. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.hittingAfter_predecessor_gt_of_untopA_gt_base","label":"hittingAfter_predecessor_gt_of_untopA_gt_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.hittingAfter_predecessor_gt_of_untopA_gt_base","description":"A delayed finite `hittingAfter` has a predecessor outside the target lower interval.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-c92d4034e2b5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8783,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:353"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hittingAfter_predecessor_gt_of_untopA_gt_base {Omega : Type*} (process : Nat -> Omega -> Real) (threshold : Real) (base : Nat) (omega : Omega) (hfinite : MeasureTheory.hittingAfter process (Set.Iic threshold) base omega ≠ ⊤) (hdelayed : base < (MeasureTheory.hittingAfter process (Set.Iic threshold) base omega).untopA) : threshold < process ((MeasureTheory.hittingAfter process (Set.Iic threshold) base omega).untopA - 1) omega","missing":[],"search":"hittingafter_predecessor_gt_of_untopa_gt_base banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.hittingafter_predecessor_gt_of_untopa_gt_base a delayed finite `hittingafter` has a predecessor outside the target lower interval. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight","label":"inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight","description":"Reciprocal overshoot weight, active only strictly after the scheduled base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-5daa50a3b087","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8784,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:383"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight (base n : Nat) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight reciprocal overshoot weight, active only strictly after the scheduled base. definition compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.max_neg_hittingAfter_untopA_le_base_abs_add_abs_stoppedValue_delayedReciprocalIncrement","label":"max_neg_hittingAfter_untopA_le_base_abs_add_abs_stoppedValue_delayedReciprocalIncrement","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.max_neg_hittingAfter_untopA_le_base_abs_add_abs_stoppedValue_delayedReciprocalIncrement","description":"At a finite first hit of a positive lower threshold, the negative part is bounded by the base absolute value plus the reciprocal-weighted final increment.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-934bef0652d7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8785,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:390"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem max_neg_hittingAfter_untopA_le_base_abs_add_abs_stoppedValue_delayedReciprocalIncrement {Omega : Type*} (average increment : Nat -> Omega -> Real) (threshold : Real) (hthreshold : 0 < threshold) (base : Nat) (omega : Omega) (hrecursion : forall n, average (n + 1) omega = ((n : Real) * average n omega + increment n omega) / ((n + 1 : Nat) : Real)) (hfinite : MeasureTheory.hittingAfter average (Set.Iic threshold) base omega ≠ ⊤) : max (-average (MeasureTheory.hittingAfter average (Set.Iic threshold) base omega).untopA omega) 0 <= |average base omega| + |stoppedValue (fun hit trajectory => inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight base hit * increment (hit - 1) trajectory) (MeasureTheory.hittingAfter average (Set.Iic threshold) base) omega|","missing":[],"search":"max_neg_hittingafter_untopa_le_base_abs_add_abs_stoppedvalue_delayedreciprocalincrement banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.max_neg_hittingafter_untopa_le_base_abs_add_abs_stoppedvalue_delayedreciprocalincrement at a finite first hit of a positive lower threshold, the negative part is bounded by the base absolute value plus the reciprocal-weighted final increment. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq","label":"summable_inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq","description":"The square of the delayed reciprocal weight is dominated by the classical inverse-square series.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-f529a7490492","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8786,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:487"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq (base : Nat) : Summable (fun n => inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight base n ^ 2)","missing":[],"search":"summable_inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight_sq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight_sq the square of the delayed reciprocal weight is dominated by the classical inverse-square series. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_tendsto_zero","label":"inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_tendsto_zero","description":"The squared reciprocal tail after a growing base tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-d16254e1f56f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8787,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:508"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_tendsto_zero : Tendsto (fun base => ∑' n : Nat, inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight base n ^ 2) atTop (nhds 0)","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight_sq_tsum_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight_sq_tsum_tendsto_zero the squared reciprocal tail after a growing base tends to zero. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_sqrt_tendsto_zero","label":"inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_sqrt_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_sqrt_tendsto_zero","description":"The square root of the squared reciprocal tail also vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-c9cec2f2ce71","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8788,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:537"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_sqrt_tendsto_zero : Tendsto (fun base => Real.sqrt (∑' n : Nat, inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight base n ^ 2)) atTop (nhds 0)","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight_sq_tsum_sqrt_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdelayedreciprocalweight_sq_tsum_sqrt_tendsto_zero the square root of the squared reciprocal tail also vanishes. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegret","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegret","description":"Reciprocal-weighted final successor-batch coordinate at the uncapped inverse-square-root hitting time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-c91e82abdf3b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8789,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:549"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedreciprocalsuccessorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedreciprocalsuccessorregret reciprocal-weighted final successor-batch coordinate at the uncapped inverse-square-root hitting time. definition compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfter_stoppedNegativePart_le_baseAbsolute_add_delayedReciprocalSuccessorRegretAbsolute","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfter_stoppedNegativePart_le_baseAbsolute_add_delayedReciprocalSuccessorRegretAbsolute","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfter_stoppedNegativePart_le_baseAbsolute_add_delayedReciprocalSuccessorRegretAbsolute","description":"Pathwise negative-part decomposition at a finite uncapped hit: the base absolute average pays for an immediate hit, while the reciprocal-weighted successor coordinate pays for a delayed hit.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-2ecd6529cbcb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8790,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:577"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfter_stoppedNegativePart_le_baseAbsolute_add_delayedReciprocalSuccessorRegretAbsolute (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) (hfinite : selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory ≠ ⊤) : let stoppedProcess := selfConsistentScheduledN…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafter_stoppednegativepart_le_baseabsolute_add_delayedreciprocalsuccessorregretabsolute banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafter_stoppednegativepart_le_baseabsolute_add_delayedreciprocalsuccessorregretabsolute pathwise negative-part decomposition at a finite uncapped hit: the base absolute average pays for an immediate hit, while the reciprocal-weighted successor coordinate pays for a delayed hit. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegret_integrable_and_integral_abs_le","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegret_integrable_and_integral_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegret_integrable_and_integral_abs_le","description":"The reciprocal-weighted final successor coordinate is integrable and its absolute first moment is controlled by the square root of the reciprocal square tail.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-c81141a371c8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8791,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:648"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegret_integrable_and_integral_abs_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let delayedRegret := selfConsistentScheduled…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_delayedreciprocalsuccessorregret_integrable_and_integral_abs_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_delayedreciprocalsuccessorregret_integrable_and_integral_abs_le the reciprocal-weighted final successor coordinate is integrable and its absolute first moment is controlled by the square root of the reciprocal square tail. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegretExpectedAbsolute","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegretExpectedAbsolute","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegretExpectedAbsolute","description":"Expected absolute reciprocal-weighted final successor coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-594301656149","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8792,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:757"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegretExpectedAbsolute (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedreciprocalsuccessorregretexpectedabsolute banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterdelayedreciprocalsuccessorregretexpectedabsolute expected absolute reciprocal-weighted final successor coordinate. definition compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegretExpectedAbsolute_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegretExpectedAbsolute_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegretExpectedAbsolute_tendsto_zero","description":"The expected absolute reciprocal-weighted overshoot coordinate vanishes as the scheduled base tends to infinity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-171e5cd1aadb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8793,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:774"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegretExpectedAbsolute_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnb…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_delayedreciprocalsuccessorregretexpectedabsolute_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_delayedreciprocalsuccessorregretexpectedabsolute_tendsto_zero the expected absolute reciprocal-weighted overshoot coordinate vanishes as the scheduled base tends to infinity. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedNegativePart","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedNegativePart","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedNegativePart","description":"Expected negative part of the exact stopped average realized behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-8cbc734f8033","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8794,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:827"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedNegativePart (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectednegativepart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectednegativepart expected negative part of the exact stopped average realized behavior regret. definition compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedNegativePart_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedNegativePart_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedNegativePart_tendsto_zero","description":"The expected negative part vanishes: immediate hits are paid by the summable base-prefix L1 term, and delayed hits by the reciprocal-weighted L2 overshoot term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-6254a0d4b950","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8795,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:851"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedNegativePart_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThre…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectednegativepart_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectednegativepart_tendsto_zero the expected negative part vanishes: immediate hits are paid by the summable base-prefix l1 term, and delayed hits by the reciprocal-weighted l2 overshoot term. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_tendsto_zero","description":"The exact stopped average realized behavior regret at the uncapped inverse-square-root `hittingAfter` converges to zero in expected absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-8050754bad64/index.html#decl-829316d8f8c8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","order":8796,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency.lean:1004"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThreshol…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_tendsto_zero the exact stopped average realized behavior regret at the uncapped inverse-square-root `hittingafter` converges to zero in expected absolute value. theorem compiled","shard":"modules/060c2dc5d2b73eb3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","description":"Every coordinate of the exact uncapped stopped process belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-91366a884ffb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8797,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialS…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret every coordinate of the exact uncapped stopped process belongs to `l1`. theorem compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_eq","description":"At exponent one, the extended norm of the exact uncapped stopped process is the lifted expected absolute regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-e60e5efbc6ba","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8798,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp ini…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret_eq at exponent one, the extended norm of the exact uncapped stopped process is the lifted expected absolute regret. theorem compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_tendsto_zero","description":"The exponent-one extended norm of the exact uncapped stopped process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-03bbd3da61ba","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8799,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState re…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret_tendsto_zero the exponent-one extended norm of the exact uncapped stopped process tends to zero. theorem compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","description":"The exponent-one extended norm of the difference from zero tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-38dbef1b7b39","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8800,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:154"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initia…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret_sub_zero_tendsto_zero the exponent-one extended norm of the difference from zero tends to zero. theorem compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp","description":"The exact uncapped stopped average realized behavior regret as an `Lp Real 1` value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-438e01ac5e8f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8801,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:197"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : Lp Real 1 (selfConsistentScheduledCausalSource mdp initialSt…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp the exact uncapped stopped average realized behavior regret as an `lp real 1` value. definition compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_coeFn_ae_eq","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_coeFn_ae_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_coeFn_ae_eq","description":"The named `Lp` coordinate represents the exact uncapped stopped process almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-b5afc90e49df","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8802,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_coeFn_ae_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : Filter.EventuallyEq (ae (selfConsistentScheduledCausalSour…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp_coefn_ae_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp_coefn_ae_eq the named `lp` coordinate represents the exact uncapped stopped process almost everywhere. theorem compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_tendsto_zero","description":"The named exact uncapped stopped `Lp Real 1` process converges to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-44cf17b6d7f7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8803,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:272"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHitti…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp_tendsto_zero the named exact uncapped stopped `lp real 1` process converges to zero. theorem compiled","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_memLp_eLpNorm_Lp_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_memLp_eLpNorm_Lp_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_memLp_eLpNorm_Lp_tendsto_zero","description":"The named exact uncapped stopped `Lp Real 1` process converges to zero. -/ theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (variancePro…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-e22699d055e9/index.html#decl-abb952cf1183","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","order":8804,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency.lean:330"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_memLp_eLpNorm_Lp_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialSt…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_memlp_elpnorm_lp_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_memlp_elpnorm_lp_tendsto_zero the named exact uncapped stopped `lp real 1` process converges to zero. -/ theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretlp mdp init…","shard":"modules/f6ab9e44e83da294.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","label":"inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","description":"Real degree-eight comparison scale for the threshold schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-819a9010ad9b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8805,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale (scheduleIndex : Nat) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdegreeeightscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdegreeeightscale real degree-eight comparison scale for the threshold schedule. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_nonneg","label":"inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_nonneg","description":"theorem inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_nonneg (scheduleIndex : Nat) : 0 <= inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-33fe8018c136","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8806,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_nonneg (scheduleIndex : Nat) : 0 <= inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdegreeeightscale_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdegreeeightscale_nonneg theorem inversesqrtthresholdunboundedhittingafterdegreeeightscale_nonneg (scheduleindex : nat) : 0 <= inversesqrtthresholdunboundedhittingafterdegreeeightscale scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_one_le","label":"inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_one_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_one_le","description":"theorem inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_one_le (scheduleIndex : Nat) : 1 <= inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-6867b77423e1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8807,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_one_le (scheduleIndex : Nat) : 1 <= inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterdegreeeightscale_one_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterdegreeeightscale_one_le theorem inversesqrtthresholdunboundedhittingafterdegreeeightscale_one_le (scheduleindex : nat) : 1 <= inversesqrtthresholdunboundedhittingafterdegreeeightscale scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointAsymptoticCoefficient","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointAsymptoticCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointAsymptoticCoefficient","description":"Fixed natural coefficient for the degree-eight checkpoint envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-d6ecc0f1dc48","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8808,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointAsymptoticCoefficient (mdp : MDP State Action) (varianceProxy : NNReal) : Nat","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialcheckpointasymptoticcoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialcheckpointasymptoticcoefficient fixed natural coefficient for the degree-eight checkpoint envelope. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare_le_asymptoticCoefficient_mul_scale_pow_eight","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare_le_asymptoticCoefficient_mul_scale_pow_eight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare_le_asymptoticCoefficient_mul_scale_pow_eight","description":"The explicit checkpoint polynomial is bounded by one fixed coefficient times the degree-eight scale.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-3126f00a3f87","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8809,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare_le_asymptoticCoefficient_mul_scale_pow_eight (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare mdp varianceProxy scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointAsymptoticCoefficient mdp varianceProxy * explicitHighProbabilityScale scheduleIndex ^ 8","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialcheckpointsquare_le_asymptoticcoefficient_mul_scale_pow_eight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialcheckpointsquare_le_asymptoticcoefficient_mul_scale_pow_eight the explicit checkpoint polynomial is bounded by one fixed coefficient times the degree-eight scale. theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient","description":"Fixed real coefficient for the polynomial stopping-round second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-55fe1a37a37c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8810,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient (mdp : MDP State Action) (varianceProxy : NNReal) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentasymptoticcoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentasymptoticcoefficient fixed real coefficient for the polynomial stopping-round second-moment budget. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient_nonneg","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient_nonneg","description":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient mdp varianceProxy","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-ff8fafbcc267","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8811,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:119"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient mdp varianceProxy","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentasymptoticcoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentasymptoticcoefficient_nonneg theorem inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentasymptoticcoefficient_nonneg (mdp : mdp state action) (varianceproxy : nnreal) : 0 <= inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentasymptoticcoefficient mdp varianceproxy theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","description":"Pointwise degree-eight bound for the real deterministic moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-bc54d7a80888","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8812,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:134"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient mdp varianceProxy * inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_le_asymptoticcoefficient_mul_degreeeightscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_le_asymptoticcoefficient_mul_degreeeightscale pointwise degree-eight bound for the real deterministic moment budget. theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_nonneg","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_nonneg","description":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : 0 <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-bb5cf14fb724","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8813,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:219"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : 0 <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_nonneg theorem inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_nonneg (mdp : mdp state action) (varianceproxy : nnreal) (scheduleindex : nat) : 0 <= inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget mdp varianceproxy scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeEight","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeEight","description":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeEight (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex) =O[atTop] inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-08364e923d4e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8814,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:234"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeEight (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex) =O[atTop] inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_isbigo_degreeeight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_isbigo_degreeeight theorem inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget_isbigo_degreeeight (mdp : mdp state action) (varianceproxy : nnreal) : (fun scheduleindex : nat => inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget mdp varianceproxy scheduleindex) =o[attop] inversesqrtthresholdunboundedhittingafterdegreeeightscale theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment","description":"Actual successor stopping-round second moment as a function of the threshold schedule index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-28bae7bd62b0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8815,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment actual successor stopping-round second moment as a function of the threshold schedule index. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment_nonneg","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : 0 <= selfCons…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-b170262ec72f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8816,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:277"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment_nonneg theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment_nonneg (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (scheduleindex : nat) : 0 <= selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppingRoundSecondMoment_le_polynomialBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppingRoundSecondMoment_le_polynomialBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppingRoundSecondMoment_le_polynomialBudget","description":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppingRoundSecondMoment_le_polynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSub…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-94b378fc2a8c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8817,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:295"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppingRoundSecondMoment_le_polynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboun…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppingroundsecondmoment_le_polynomialbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppingroundsecondmoment_le_polynomialbudget theorem selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppingroundsecondmoment_le_polynomialbudget (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (scheduleindex : nat) : selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor scheduleindex <= inversesqrtthresholdunbounde…","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_isBigO_degreeEight","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_isBigO_degreeEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_isBigO_degreeEight","description":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_isBigO_degreeEight (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubg…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-047d77707b4e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8818,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:323"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_isBigO_degreeEight (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppin…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_isbigo_degreeeight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_isbigo_degreeeight theorem selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_isbigo_degreeeight (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : (selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppingroundsecondmoment mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor) =o[attop] inversesqrtthresholdunboundedhittingafterdegreeeightscale…","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant","label":"inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant","description":"Universal shifted inverse-square series appearing in the stopped-value absolute-first-moment bridge.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-880237e39fc9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8819,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:363"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingaftershiftedinversesquareseriesconstant banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingaftershiftedinversesquareseriesconstant universal shifted inverse-square series appearing in the stopped-value absolute-first-moment bridge. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant_nonneg","label":"inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant_nonneg","description":"theorem inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant_nonneg : 0 <= inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-d36e060036db","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8820,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:371"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant_nonneg : 0 <= inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant","missing":[],"search":"inversesqrtthresholdunboundedhittingaftershiftedinversesquareseriesconstant_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingaftershiftedinversesquareseriesconstant_nonneg theorem inversesqrtthresholdunboundedhittingaftershiftedinversesquareseriesconstant_nonneg : 0 <= inversesqrtthresholdunboundedhittingaftershiftedinversesquareseriesconstant theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient","description":"Fixed coefficient for the polynomial absolute-first-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-90eebfc0c512","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8821,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:379"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient (mdp : MDP State Action) (varianceProxy : NNReal) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient fixed coefficient for the polynomial absolute-first-moment budget. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient_nonneg","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient mdp varianceProxy","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-f4301915d9d3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8822,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:393"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient mdp varianceProxy","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient_nonneg theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient_nonneg (mdp : mdp state action) (varianceproxy : nnreal) : 0 <= selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient mdp varianceproxy theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget mdp variancePro…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-8793935a0e0c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8823,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:411"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient mdp varianceProxy * inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_le_asymptoticcoefficient_mul_degreeeightscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_le_asymptoticcoefficient_mul_degreeeightscale theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_le_asymptoticcoefficient_mul_degreeeightscale (mdp : mdp state action) (varianceproxy : nnreal) (scheduleindex : nat) : selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget mdp varianceproxy scheduleindex <= selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialabsolutefirstmomentasymptoticcoefficient mdp varianceproxy * inversesqrtthresholdunboundedhittingafterdegreeeightscale scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_nonneg","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-c7bbbe71c1df","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8824,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:515"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_nonneg theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_nonneg (mdp : mdp state action) (varianceproxy : nnreal) (scheduleindex : nat) : 0 <= selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget mdp varianceproxy scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_isBigO_degreeEight","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_isBigO_degreeEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_isBigO_degreeEight","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_isBigO_degreeEight (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex) =O[…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-19c0fed7e265","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8825,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:542"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_isBigO_degreeEight (mdp : MDP State Action) (varianceProxy : NNReal) : (fun scheduleIndex : Nat => selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget mdp varianceProxy scheduleIndex) =O[atTop] inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_isbigo_degreeeight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_isbigo_degreeeight theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget_isbigo_degreeeight (mdp : mdp state action) (varianceproxy : nnreal) : (fun scheduleindex : nat => selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget mdp varianceproxy scheduleindex) =o[attop] inversesqrtthresholdunboundedhittingafterdegreeeightscale theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute","description":"Expected absolute stopped average realized behavior regret as a function of the threshold schedule index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-9d95a1d76fca","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8826,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:568"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute expected absolute stopped average realized behavior regret as a function of the threshold schedule index. definition compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute_nonneg","description":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleI…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-180887b35dac","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8827,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:589"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute_nonneg theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute_nonneg (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (scheduleindex : nat) : 0 <= selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor scheduleindex theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_polynomialBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_polynomialBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_polynomialBudget","description":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_polynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (la…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-1ceb90a17b23","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8828,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:607"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_polynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausa…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_le_polynomialbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_le_polynomialbudget theorem selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_le_polynomialbudget (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (scheduleindex : nat) : selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretexpectedabsolute mdp initialstate rewar…","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeEight","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeEight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeEight","description":"With all model/source parameters fixed, the actual expected absolute stopped regret grows at most at the compiled degree-eight rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-ceed8fe354ff/index.html#decl-03e61f81073a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","order":8829,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics.lean:637"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeEight (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : (selfConsistentScheduledNaturalCausalInverseSqrtThresholdU…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_isbigo_degreeeight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregretexpectedabsolute_isbigo_degreeeight with all model/source parameters fixed, the actual expected absolute stopped regret grows at most at the compiled degree-eight rate. theorem compiled","shard":"modules/0b1376067a4a4888.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialScaleCoefficient","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialScaleCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialScaleCoefficient","description":"Natural model coefficient used by the polynomial tail-start envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-df655b683fa7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8830,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterPolynomialScaleCoefficient (mdp : MDP State Action) (varianceProxy : NNReal) : Nat","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialscalecoefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialscalecoefficient natural model coefficient used by the polynomial tail-start envelope. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_succ_le_polynomialScale","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_succ_le_polynomialScale","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_succ_le_polynomialScale","description":"The explicit ceiling start grows at most linearly in the schedule scale.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-e498e7722d04","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8831,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_succ_le_polynomialScale (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart mdp varianceProxy scheduleIndex + 1 <= inverseSqrtThresholdUnboundedHittingAfterPolynomialScaleCoefficient mdp varianceProxy * explicitHighProbabilityScale scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicittailstart_succ_le_polynomialscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicittailstart_succ_le_polynomialscale the explicit ceiling start grows at most linearly in the schedule scale. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare","description":"Explicit degree-eight natural checkpoint-square envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-adfcf636538c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8832,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Nat","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialcheckpointsquare banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialcheckpointsquare explicit degree-eight natural checkpoint-square envelope. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitCheckpointSquare_le_polynomial","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitCheckpointSquare_le_polynomial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitCheckpointSquare_le_polynomial","description":"The exact explicit-start checkpoint square is bounded by the polynomial envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-e9b23a52f6fe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8833,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:117"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterExplicitCheckpointSquare_le_polynomial (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : (explicitHighProbabilityRounds (inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart mdp varianceProxy scheduleIndex) + 1) ^ 2 <= inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicitcheckpointsquare_le_polynomial banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicitcheckpointsquare_le_polynomial the exact explicit-start checkpoint square is bounded by the polynomial envelope. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant","label":"inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant","description":"The weighted model/return failure contribution to the stopping-round second moment. It depends only on the MDP.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-a3de85b9c4f9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8834,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:138"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant (mdp : MDP State Action) : ENNReal","missing":[],"search":"inversesqrtthresholdunboundedhittingafterweightedfailuresecondmomentennrealconstant banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterweightedfailuresecondmomentennrealconstant the weighted model/return failure contribution to the stopping-round second moment. it depends only on the mdp. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant_ne_top","label":"inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant_ne_top","description":"The weighted failure constant is finite under the horizon-five contract.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-1be5ec890c42","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8835,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant_ne_top (mdp : MDP State Action) (hhorizon : 4 < mdp.horizon) : inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant mdp ≠ ∞","missing":[],"search":"inversesqrtthresholdunboundedhittingafterweightedfailuresecondmomentennrealconstant_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterweightedfailuresecondmomentennrealconstant_ne_top the weighted failure constant is finite under the horizon-five contract. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentConstant","label":"inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentConstant","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentConstant","description":"Real-valued form of the weighted model/return failure constant.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-186364163310","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8836,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentConstant (mdp : MDP State Action) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterweightedfailuresecondmomentconstant banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterweightedfailuresecondmomentconstant real-valued form of the weighted model/return failure constant. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget","description":"ENNReal polynomial stopping-round second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-ea29304e5ed0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8837,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:167"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : ENNReal","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentennrealbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentennrealbudget ennreal polynomial stopping-round second-moment budget. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_ne_top","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_ne_top","description":"The polynomial ENNReal budget is finite under the same horizon contract.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-f31ddb66116c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8838,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:180"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_ne_top (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) (hhorizon : 4 < mdp.horizon) : inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget mdp varianceProxy scheduleIndex ≠ ∞","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentennrealbudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentennrealbudget_ne_top the polynomial ennreal budget is finite under the same horizon contract. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget","description":"Real polynomial stopping-round second-moment budget, displayed as a degree-eight checkpoint term plus the named MDP failure constant.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-70dbb030a28f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8839,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:195"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentbudget real polynomial stopping-round second-moment budget, displayed as a degree-eight checkpoint term plus the named mdp failure constant. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_toReal","label":"inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_toReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_toReal","description":"Taking `toReal` of the polynomial ENNReal budget gives its displayed real polynomial-plus-constant form.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-9402242b6331","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8840,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_toReal (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) (hhorizon : 4 < mdp.horizon) : (inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget mdp varianceProxy scheduleIndex).toReal = inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentennrealbudget_toreal banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentennrealbudget_toreal taking `toreal` of the polynomial ennreal budget gives its displayed real polynomial-plus-constant form. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_le_polynomial","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_le_polynomial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_le_polynomial","description":"The ceiling-start ENNReal budget is bounded by the polynomial budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-8d6efb9d8034","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8841,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:230"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_le_polynomial (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget mdp varianceProxy scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentennrealbudget_le_polynomial banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentennrealbudget_le_polynomial the ceiling-start ennreal budget is bounded by the polynomial budget. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget_le_polynomial","label":"inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget_le_polynomial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget_le_polynomial","description":"The real ceiling-start second-moment budget is bounded by its displayed polynomial-plus-constant envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-58f504cadcf7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8842,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:253"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget_le_polynomial (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) (hhorizon : 4 < mdp.horizon) : inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget mdp varianceProxy scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentbudget_le_polynomial banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterexplicitstoppingroundsecondmomentbudget_le_polynomial the real ceiling-start second-moment budget is bounded by its displayed polynomial-plus-constant envelope. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_polynomialBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_polynomialBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_polynomialBudget","description":"The actual successor stopping-round second moment is bounded by the polynomial-plus-model-constant budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-4f423a7346c4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8843,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:281"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_polynomialBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialS…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_le_polynomialbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_le_polynomialbudget the actual successor stopping-round second moment is bounded by the polynomial-plus-model-constant budget. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget","description":"Polynomial-envelope absolute-first-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-6a5a709b84a2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8844,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:318"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterpolynomialstoppingroundsecondmomentabsolutefirstmomentbudget polynomial-envelope absolute-first-moment budget. definition compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_polynomialStoppingRoundSecondMomentBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_polynomialStoppingRoundSecondMomentBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_polynomialStoppingRoundSecondMomentBudget","description":"For each fixed threshold index, the stopped average realized behavior regret has a polynomial-plus-model-constant absolute first-moment bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-44f9271b353d/index.html#decl-b57a47e96a19","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","order":8845,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_polynomialStoppingRoundSecondMomentBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat)…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_polynomialstoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_integrable_and_integral_abs_le_polynomialstoppingroundsecondmomentbudget for each fixed threshold index, the stopped average realized behavior regret has a polynomial-plus-model-constant absolute first-moment bound. theorem compiled","shard":"modules/f2b75c0aecc1158c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityQuarticSquareBlockWeight","label":"explicitHighProbabilityQuarticSquareBlockWeight","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityQuarticSquareBlockWeight","description":"Seventh-degree envelope for one consecutive squared fourth-power checkpoint block.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-625951fb08e4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8846,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def explicitHighProbabilityQuarticSquareBlockWeight (n : Nat) : Nat","missing":[],"search":"explicithighprobabilityquarticsquareblockweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityquarticsquareblockweight seventh-degree envelope for one consecutive squared fourth-power checkpoint block. definition compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_succ_square_sub_le_quarticSquareBlockWeight","label":"explicitHighProbabilityRounds_succ_square_sub_le_quarticSquareBlockWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_succ_square_sub_le_quarticSquareBlockWeight","description":"A consecutive squared fourth-power checkpoint gap is bounded by the seventh-degree weight.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-c44938cdaab0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8847,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_succ_square_sub_le_quarticSquareBlockWeight (n : Nat) : (explicitHighProbabilityRounds (n + 1) + 1) ^ 2 - (explicitHighProbabilityRounds n + 1) ^ 2 <= explicitHighProbabilityQuarticSquareBlockWeight n","missing":[],"search":"explicithighprobabilityrounds_succ_square_sub_le_quarticsquareblockweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_succ_square_sub_le_quarticsquareblockweight a consecutive squared fourth-power checkpoint gap is bounded by the seventh-degree weight. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_le_inv_pow_ten","label":"selfConsistentScheduledLocalDelta_le_inv_pow_ten","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_le_inv_pow_ten","description":"Horizon at least five makes every local confidence share no larger than a shifted inverse tenth power.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-ab3fb7d6a9ad","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8848,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledLocalDelta_le_inv_pow_ten (mdp : MDP State Action) (hhorizon : 4 < mdp.horizon) (t : Nat) : AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta mdp t <= 1 / (((t + 2 : Nat) : Real) ^ 10)","missing":[],"search":"selfconsistentscheduledlocaldelta_le_inv_pow_ten banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledlocaldelta_le_inv_pow_ten horizon at least five makes every local confidence share no larger than a shifted inverse tenth power. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedInverseCubePairEnvelope","label":"quarticSquareBlockShiftedInverseCubePairEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedInverseCubePairEnvelope","description":"Inverse-cube pair envelope after a seventh-degree weight cancels seven inverse powers.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-f7b6aedb2af4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8849,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def quarticSquareBlockShiftedInverseCubePairEnvelope (p : Nat × Nat) : ENNReal","missing":[],"search":"quarticsquareblockshiftedinversecubepairenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticsquareblockshiftedinversecubepairenvelope inverse-cube pair envelope after a seventh-degree weight cancels seven inverse powers. definition compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockShiftedInverseCubePairEnvelope_ne_top","label":"tsum_quarticSquareBlockShiftedInverseCubePairEnvelope_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockShiftedInverseCubePairEnvelope_ne_top","description":"The square-block shifted pair envelope has finite total ENNReal mass.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-319070557916","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8850,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticSquareBlockShiftedInverseCubePairEnvelope_ne_top : Ne (∑' p : Nat × Nat, quarticSquareBlockShiftedInverseCubePairEnvelope p) ∞","missing":[],"search":"tsum_quarticsquareblockshiftedinversecubepairenvelope_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticsquareblockshiftedinversecubepairenvelope_ne_top the square-block shifted pair envelope has finite total ennreal mass. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedCoordinateModelFailureCharge","label":"quarticSquareBlockShiftedCoordinateModelFailureCharge","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedCoordinateModelFailureCharge","description":"One squared-checkpoint block weight times one shifted coordinate model-failure charge.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-485629302f6b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8851,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def quarticSquareBlockShiftedCoordinateModelFailureCharge (mdp : MDP State Action) (p : Nat × Nat) : ENNReal","missing":[],"search":"quarticsquareblockshiftedcoordinatemodelfailurecharge banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticsquareblockshiftedcoordinatemodelfailurecharge one squared-checkpoint block weight times one shifted coordinate model-failure charge. definition compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","label":"quarticSquareBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","description":"The inverse-tenth local share puts every seventh-weighted coordinate charge below the pair envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-1253d0414db6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8852,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem quarticSquareBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope (mdp : MDP State Action) (hhorizon : 4 < mdp.horizon) (p : Nat × Nat) : quarticSquareBlockShiftedCoordinateModelFailureCharge mdp p <= quarticSquareBlockShiftedInverseCubePairEnvelope p","missing":[],"search":"quarticsquareblockshiftedcoordinatemodelfailurecharge_le_pairenvelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticsquareblockshiftedcoordinatemodelfailurecharge_le_pairenvelope the inverse-tenth local share puts every seventh-weighted coordinate charge below the pair envelope. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockShiftedCoordinateModelFailureCharge_ne_top","label":"tsum_quarticSquareBlockShiftedCoordinateModelFailureCharge_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockShiftedCoordinateModelFailureCharge_ne_top","description":"The seventh-weighted shifted coordinate charges have finite total ENNReal mass.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-3b15919093e4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8853,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:170"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticSquareBlockShiftedCoordinateModelFailureCharge_ne_top (mdp : MDP State Action) (hhorizon : 4 < mdp.horizon) : Ne (∑' p : Nat × Nat, quarticSquareBlockShiftedCoordinateModelFailureCharge mdp p) ∞","missing":[],"search":"tsum_quarticsquareblockshiftedcoordinatemodelfailurecharge_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticsquareblockshiftedcoordinatemodelfailurecharge_ne_top the seventh-weighted shifted coordinate charges have finite total ennreal mass. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","label":"quarticSquareBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","description":"A model tail after `n+1`, charged by one squared-checkpoint block, is the shifted coordinate row.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-e56003251685","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8854,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:184"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem quarticSquareBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge (mdp : MDP State Action) (n : Nat) : (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * selfConsistentScheduledCausalTailModelFailureBudget mdp (n + 1) = ∑' j : Nat, quarticSquareBlockShiftedCoordinateModelFailureCharge mdp (n, j)","missing":[],"search":"quarticsquareblockweight_mul_tailmodelfailurebudget_eq_tsum_shiftedcoordinatecharge banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.quarticsquareblockweight_mul_tailmodelfailurebudget_eq_tsum_shiftedcoordinatecharge a model tail after `n+1`, charged by one squared-checkpoint block, is the shifted coordinate row. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_tailModelFailureBudget_ne_top","label":"tsum_quarticSquareBlockWeight_mul_tailModelFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_tailModelFailureBudget_ne_top","description":"Squared-checkpoint block weights are summable against the exact infinite model-tail budgets.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-ff67ae15481c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8855,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticSquareBlockWeight_mul_tailModelFailureBudget_ne_top (mdp : MDP State Action) (hhorizon : 4 < mdp.horizon) : Ne (∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * selfConsistentScheduledCausalTailModelFailureBudget mdp (n + 1)) ∞","missing":[],"search":"tsum_quarticsquareblockweight_mul_tailmodelfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticsquareblockweight_mul_tailmodelfailurebudget_ne_top squared-checkpoint block weights are summable against the exact infinite model-tail budgets. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta","label":"summable_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta","description":"A seventh-degree squared-checkpoint weight times the exponential return share is summable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-e2f9eaaaed28","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8856,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:232"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem summable_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta : Summable (fun n : Nat => ((explicitHighProbabilityQuarticSquareBlockWeight n : Nat) : Real) * explicitHighProbabilityReturnDelta n)","missing":[],"search":"summable_quarticsquareblockweight_mul_explicithighprobabilityreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.summable_quarticsquareblockweight_mul_explicithighprobabilityreturndelta a seventh-degree squared-checkpoint weight times the exponential return share is summable. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","label":"tsum_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","description":"The exact return shares have finite total mass after squared-checkpoint charging.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-36f21ed6a577","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8857,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top : Ne (∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * ENNReal.ofReal (explicitHighProbabilityReturnDelta n)) ∞","missing":[],"search":"tsum_quarticsquareblockweight_mul_explicithighprobabilityreturndelta_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticsquareblockweight_mul_explicithighprobabilityreturndelta_ne_top the exact return shares have finite total mass after squared-checkpoint charging. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","label":"tsum_quarticSquareBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","description":"The exact checkpoint violation budgets remain summable after squared-checkpoint charging.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-b7b137a6cc34","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8858,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:288"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tsum_quarticSquareBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top (mdp : MDP State Action) (hhorizon : 4 < mdp.horizon) : Ne (∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * explicitPolynomialPrefixTailModelReturnFailureBudget mdp n) ∞","missing":[],"search":"tsum_quarticsquareblockweight_mul_explicitpolynomialprefixtailmodelreturnfailurebudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.tsum_quarticsquareblockweight_mul_explicitpolynomialprefixtailmodelreturnfailurebudget_ne_top the exact checkpoint violation budgets remain summable after squared-checkpoint charging. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_add_square_le_add_sum_Ico_quarticSquareBlockWeight","label":"explicitHighProbabilityRounds_add_square_le_add_sum_Ico_quarticSquareBlockWeight","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_add_square_le_add_sum_Ico_quarticSquareBlockWeight","description":"Seventh-degree weights dominate every finite telescope of squared checkpoint values.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-387a3297ec20","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8859,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_add_square_le_add_sum_Ico_quarticSquareBlockWeight (start width : Nat) : (explicitHighProbabilityRounds (start + width) + 1) ^ 2 <= (explicitHighProbabilityRounds start + 1) ^ 2 + (Finset.Ico start (start + width)).sum explicitHighProbabilityQuarticSquareBlockWeight","missing":[],"search":"explicithighprobabilityrounds_add_square_le_add_sum_ico_quarticsquareblockweight banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_add_square_le_add_sum_ico_quarticsquareblockweight seventh-degree weights dominate every finite telescope of squared checkpoint values. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.natCast_succ_square_le_checkpoint_square_add_tsum_quarticSquareBlockWeight_of_delayed","label":"natCast_succ_square_le_checkpoint_square_add_tsum_quarticSquareBlockWeight_of_delayed","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.natCast_succ_square_le_checkpoint_square_add_tsum_quarticSquareBlockWeight_of_delayed","description":"The square of a finite natural time is paid for by an initial squared checkpoint plus preceding delayed blocks.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-f9dce88b1d20","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8860,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:358"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem natCast_succ_square_le_checkpoint_square_add_tsum_quarticSquareBlockWeight_of_delayed (start time : Nat) : (((time + 1) ^ 2 : Nat) : ENNReal) <= (((explicitHighProbabilityRounds start + 1) ^ 2 : Nat) : ENNReal) + ∑' n : Nat, if start <= n ∧ explicitHighProbabilityRounds n < time then (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) else 0","missing":[],"search":"natcast_succ_square_le_checkpoint_square_add_tsum_quarticsquareblockweight_of_delayed banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.natcast_succ_square_le_checkpoint_square_add_tsum_quarticsquareblockweight_of_delayed the square of a finite natural time is paid for by an initial squared checkpoint plus preceding delayed blocks. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_inverseSqrtThresholdUnboundedHittingAfterTailStart","label":"exists_inverseSqrtThresholdUnboundedHittingAfterTailStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_inverseSqrtThresholdUnboundedHittingAfterTailStart","description":"A deterministic checkpoint after which the scheduled regret envelope is below the fixed inverse-square-root first-passage threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-2180854ceeb3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8861,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:468"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exists_inverseSqrtThresholdUnboundedHittingAfterTailStart (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : ∃ start : Nat, scheduleIndex <= start ∧ ∀ n, start <= n -> explicitPolynomialPrefixAverageRealizedBehaviorRegretRate mdp varianceProxy baseVisitFloor n <= selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex","missing":[],"search":"exists_inversesqrtthresholdunboundedhittingaftertailstart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exists_inversesqrtthresholdunboundedhittingaftertailstart a deterministic checkpoint after which the scheduled regret envelope is below the fixed inverse-square-root first-passage threshold. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart","label":"inverseSqrtThresholdUnboundedHittingAfterTailStart","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart","description":"Canonical deterministic witness for the eventual delayed-checkpoint tail bound at a fixed inverse-square-root threshold index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-30a271499e3d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8862,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:490"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterTailStart (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Nat","missing":[],"search":"inversesqrtthresholdunboundedhittingaftertailstart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingaftertailstart canonical deterministic witness for the eventual delayed-checkpoint tail bound at a fixed inverse-square-root threshold index. definition compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart_spec","label":"inverseSqrtThresholdUnboundedHittingAfterTailStart_spec","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart_spec","description":"The canonical tail start is beyond the threshold index and validates the scheduled regret-rate comparison at every later checkpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-19461ed70cea","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8863,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:500"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterTailStart_spec (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : scheduleIndex <= inverseSqrtThresholdUnboundedHittingAfterTailStart mdp varianceProxy baseVisitFloor scheduleIndex ∧ ∀ n, inverseSqrtThresholdUnboundedHittingAfterTailStart mdp varianceProxy baseVisitFloor scheduleIndex <= n -> explicitPolynomialPrefixAverageRealizedBehaviorRegretRate mdp varianceProxy baseVisitFloor n <= selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold scheduleIndex","missing":[],"search":"inversesqrtthresholdunboundedhittingaftertailstart_spec banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingaftertailstart_spec the canonical tail start is beyond the threshold index and validates the scheduled regret-rate comparison at every later checkpoint. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget","label":"inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget","description":"Explicit deterministic ENNReal budget for the second moment of the successor round count at the uncapped inverse-square-root hitting time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-bd39a16af19d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8864,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:520"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : ENNReal","missing":[],"search":"inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentennrealbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentennrealbudget explicit deterministic ennreal budget for the second moment of the successor round count at the uncapped inverse-square-root hitting time. definition compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_ne_top","label":"inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_ne_top","description":"The deterministic stopping-round second-moment budget is finite whenever the horizon supplies the inverse-tenth confidence exponent.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-986d6bc4330e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8865,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:533"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_ne_top (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (hhorizon : 4 < mdp.horizon) : inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget mdp varianceProxy baseVisitFloor scheduleIndex ≠ ∞","missing":[],"search":"inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentennrealbudget_ne_top banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentennrealbudget_ne_top the deterministic stopping-round second-moment budget is finite whenever the horizon supplies the inverse-tenth confidence exponent. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget","label":"inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget","description":"Real-valued form of the deterministic stopping-round second-moment budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-19fe02ea0b01","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8866,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:549"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.inversesqrtthresholdunboundedhittingafterstoppingroundsecondmomentbudget real-valued form of the deterministic stopping-round second-moment budget. definition compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.lintegral_sq_untopA_add_one_le_quarticSquareCheckpointBudget","label":"lintegral_sq_untopA_add_one_le_quarticSquareCheckpointBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.lintegral_sq_untopA_add_one_le_quarticSquareCheckpointBudget","description":"A pointwise delayed-checkpoint tail budget gives an explicit ENNReal upper bound for the successor stopping-round second moment.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-a8b3740fd058","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8867,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:561"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem lintegral_sq_untopA_add_one_le_quarticSquareCheckpointBudget {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (budget : Nat -> ENNReal) (start : Nat) (htailStart : ∀ n, start <= n -> mu {omega | (explicitHighProbabilityRounds n : WithTop Nat) < tau omega} <= budget n) : ∫⁻ omega, ENNReal.ofReal (((((tau omega).untopA + 1 : Nat) : Real)) ^ 2) ∂mu <= (((explicitHighProbabilityRounds start + 1) ^ 2 : Nat) : ENNReal) + ∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * budget n","missing":[],"search":"lintegral_sq_untopa_add_one_le_quarticsquarecheckpointbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.lintegral_sq_untopa_add_one_le_quarticsquarecheckpointbudget a pointwise delayed-checkpoint tail budget gives an explicit ennreal upper bound for the successor stopping-round second moment. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_untopA_add_one_of_eventually_quarticSquareCheckpointTail","label":"memLp_two_untopA_add_one_of_eventually_quarticSquareCheckpointTail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_untopA_add_one_of_eventually_quarticSquareCheckpointTail","description":"Eventually summable squared fourth-power checkpoint crossing tails imply a finite second moment.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-0b92dccb41c5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8868,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:660"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_two_untopA_add_one_of_eventually_quarticSquareCheckpointTail {Omega : Type*} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (budget : Nat -> ENNReal) (hbudget : Ne (∑' n : Nat, (explicitHighProbabilityQuarticSquareBlockWeight n : ENNReal) * budget n) ∞) (htail : ∀ᶠ n : Nat in atTop, mu {omega | (explicitHighProbabilityRounds n : WithTop Nat) < tau omega} <= budget n) : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu","missing":[],"search":"memlp_two_untopa_add_one_of_eventually_quarticsquarecheckpointtail banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_two_untopa_add_one_of_eventually_quarticsquarecheckpointtail eventually summable squared fourth-power checkpoint crossing tails imply a finite second moment. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le_of_tailStart","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le_of_tailStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le_of_tailStart","description":"From the canonical tail start onward, every delayed uncapped first-passage checkpoint has the exact compiled model/return failure budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-dee9ac27d84a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8869,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:773"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le_of_tailStart (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex checkpointIndex : Nat) (hcheckpointIndex : inverseSqrtThreshold…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_inversesqrtthresholdunboundedhittingafterdelayedcheckpointset_le_of_tailstart banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_inversesqrtthresholdunboundedhittingafterdelayedcheckpointset_le_of_tailstart from the canonical tail start onward, every delayed uncapped first-passage checkpoint has the exact compiled model/return failure budget. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_lintegral_stoppingRound_sq_le_ENNRealBudget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_lintegral_stoppingRound_sq_le_ENNRealBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_lintegral_stoppingRound_sq_le_ENNRealBudget","description":"The actual successor stopping-round square has an explicit deterministic ENNReal upper bound at every fixed inverse-square-root threshold index.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-dd9210f2ef35","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8870,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:826"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_lintegral_stoppingRound_sq_le_ENNRealBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialSta…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_lintegral_stoppinground_sq_le_ennrealbudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_lintegral_stoppinground_sq_le_ennrealbudget the actual successor stopping-round square has an explicit deterministic ennreal upper bound at every fixed inverse-square-root threshold index. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime","description":"Each fixed-index genuine uncapped inverse-square-root first passage has a finite second moment when the finite-horizon confidence exponent is at least ten.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-5d537b98d971","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8871,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:876"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState reward…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_squareintegrablefinitestoppingtime banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_squareintegrablefinitestoppingtime each fixed-index genuine uncapped inverse-square-root first passage has a finite second moment when the finite-horizon confidence exponent is at least ten. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_budget","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_budget","description":"The Bochner second moment of the actual successor stopping-round count is bounded by the canonical deterministic checkpoint/failure budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6ad0e95a3a25/index.html#decl-bc915aec614d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","order":8872,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime.lean:928"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_budget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewar…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_le_budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_integral_stoppinground_sq_le_budget the bochner second moment of the actual successor stopping-round count is bounded by the canonical deterministic checkpoint/failure budget. theorem compiled","shard":"modules/5c53f54a74e9eb34.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","description":"Average successor-policy expected regret evaluated at a stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-5a79f563c130","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8873,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess average successor-policy expected regret evaluated at a stopping prefix. definition compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_apply","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeB…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-5b0615f76e32","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8874,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTabl…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_apply theorem selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (scheduleindex : nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) : selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor stoppingprefix scheduleindex trajectory = selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisi…","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess","description":"Average normalized return deviation evaluated at a stopping prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-283301725d1b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8875,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess average normalized return deviation evaluated at a stopping prefix. definition compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_apply","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTra…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-8b9863723042","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8876,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess mdp initialState rewardSource initialTable defaultState…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess_apply theorem selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (scheduleindex : nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) : selfconsistentschedulednaturalcausalstoppingtimeaveragereturndeviationprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor stoppingprefix scheduleindex trajectory = selfconsistentschedulednaturalcausalaveragereturndeviationprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor (stoppingprefix scheduleinde…","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","description":"Every deterministic-prefix average behavior expected-regret coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-fbb1c191c456","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8877,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess every deterministic-prefix average behavior expected-regret coordinate is measurable. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_nonneg","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_nonneg","description":"Every deterministic-prefix average behavior expected regret is nonnegative, including the zero-prefix convention.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-a2aba99189c7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8878,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : 0 <= selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess_nonneg every deterministic-prefix average behavior expected regret is nonnegative, including the zero-prefix convention. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","label":"selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","description":"The deterministic-prefix average behavior expected regret has the global policy-value envelope `2H`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-49eec55e9d8e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8879,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:185"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_le_two_mul_horizon (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory <= 2 * (mdp.horizon : Real)","missing":[],"search":"selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess_le_two_mul_horizon banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragebehaviorexpectedregretprocess_le_two_mul_horizon the deterministic-prefix average behavior expected regret has the global policy-value envelope `2h`. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","description":"A measurable stopping prefix gives a measurable stopped behavior expected-regret coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-b0d071e85dd2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8880,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (hstopping : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (stoppingPrefix scheduleIndex)) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess…","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess a measurable stopping prefix gives a measurable stopped behavior expected-regret coordinate. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_nonneg","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_nonneg","description":"Stopping preserves nonnegativity of the pathwise behavior expected-regret average.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-88de0d2db360","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8881,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:255"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : 0 <= selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initi…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_nonneg stopping preserves nonnegativity of the pathwise behavior expected-regret average. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","description":"Stopping preserves the deterministic `2H` policy-value envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-23f09a3927c8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8882,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:283"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_le_two_mul_horizon (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStopping…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_le_two_mul_horizon banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragebehaviorexpectedregretprocess_le_two_mul_horizon stopping preserves the deterministic `2h` policy-value envelope. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_eq_behaviorExpected_sub_returnDeviation","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_eq_behaviorExpected_sub_returnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_eq_behaviorExpected_sub_returnDeviation","description":"The stopped realized process is exactly stopped behavior expected regret minus stopped normalized return deviation at the same prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-b846ecbe0c5d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8883,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:313"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_eq_behaviorExpected_sub_returnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess mdp ini…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess_eq_behaviorexpected_sub_returndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess_eq_behaviorexpected_sub_returndeviation the stopped realized process is exactly stopped behavior expected regret minus stopped normalized return deviation at the same prefix. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","description":"The exact uncapped stopped behavior expected-regret process converges almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-51f018f602f8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8884,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:350"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewa…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedbehaviorexpectedregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedbehaviorexpectedregret_tendstoalmosteverywhere_zero the exact uncapped stopped behavior expected-regret process converges almost everywhere. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","description":"Every exact uncapped stopped behavior expected-regret coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-1c0dcfbb4092","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8885,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:396"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret every exact uncapped stopped behavior expected-regret coordinate is measurable. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","label":"integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","description":"Every exact uncapped stopped behavior expected-regret coordinate is integrable by the global `2H` envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-6b5ebf8e254f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8886,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:422"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : Integrable (selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex) (selfConsistentScheduledCausalSource mdp…","missing":[],"search":"integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret every exact uncapped stopped behavior expected-regret coordinate is integrable by the global `2h` envelope. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","description":"Every exact uncapped stopped behavior expected-regret coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-9749aa75c285","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8887,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:468"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : MemLp (selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) scheduleIndex) 1 (selfConsistentScheduledCausalSource mdp ini…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret every exact uncapped stopped behavior expected-regret coordinate belongs to `l1`. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","label":"integral_abs_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","description":"Expected absolute stopped behavior expected regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-c46515dd2ebd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8888,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:496"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSo…","missing":[],"search":"integral_abs_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_abs_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret_tendsto_zero expected absolute stopped behavior expected regret tends to zero. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_eq","description":"Exponent-one norm of the exact stopped behavior expected-regret process is the lifted expected absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-e1b9102dfbad","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8889,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:591"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let stoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let behaviorProcess := selfConsistentScheduledNaturalCausalStoppingTimeAverage…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret_eq exponent-one norm of the exact stopped behavior expected-regret process is the lifted expected absolute value. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","description":"The exact stopped behavior expected-regret process converges to zero in exponent-one norm.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-183155bbc63d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8890,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:624"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSou…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregret_tendsto_zero the exact stopped behavior expected-regret process converges to zero in exponent-one norm. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretIntegral_tendsto_zero","description":"Signed expectation of the exact stopped behavior expected-regret process tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-00220d1d738a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8891,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:667"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregretintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedbehaviorexpectedregretintegral_tendsto_zero signed expectation of the exact stopped behavior expected-regret process tends to zero. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation","label":"memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation","description":"The exact uncapped stopped return-deviation coordinates belong to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-1c0d7bb97718","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8892,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:708"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSou…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedreturndeviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedreturndeviation the exact uncapped stopped return-deviation coordinates belong to `l1`. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation_tendsto_zero","description":"Exponent-one norm of the exact stopped return deviation tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-96f8f3ef3a93","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8893,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:784"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource ini…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedreturndeviation_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedreturndeviation_tendsto_zero exponent-one norm of the exact stopped return deviation tends to zero. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviationIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviationIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviationIntegral_tendsto_zero","description":"Signed expectation of the stopped return deviation tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-b7e431cdb5f0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8894,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:896"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviationIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initial…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedreturndeviationintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedreturndeviationintegral_tendsto_zero signed expectation of the stopped return deviation tends to zero. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_consistency","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_consistency","description":"Terminal policy-value semantic and `L1` package at genuine uncapped `hittingAfter`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-6c68733bace1/index.html#decl-e60c52f85c40","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","order":8895,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency.lean:939"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialStat…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedbehaviorexpectedregret_and_returndeviation_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedbehaviorexpectedregret_and_returndeviation_l1_consistency terminal policy-value semantic and `l1` package at genuine uncapped `hittingafter`. theorem compiled","shard":"modules/327c3e55ce790481.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_restrict_le_of_uniformIntegrable_one","label":"integral_abs_restrict_le_of_uniformIntegrable_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_restrict_le_of_uniformIntegrable_one","description":"Probability-theory uniform integrability at exponent one gives uniform absolute continuity of the expected norm over measurable events. This is a thin real-valued wrapper around Mathlib's `UnifIntegrable` epsilon-delta interface.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-f72757b422c4/index.html#decl-5769a1fc60c6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","order":8896,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_restrict_le_of_uniformIntegrable_one {Omega : Type w} [MeasurableSpace Omega] {mu : Measure Omega} {f : Nat -> Omega -> Real} (hui : UniformIntegrable f 1 mu) : forall epsilon : Real, 0 < epsilon -> exists delta : Real, 0 < delta /\\ forall i event, MeasurableSet event -> mu event <= ENNReal.ofReal delta -> integral (mu.restrict event) (fun omega => |f i omega|) <= epsilon","missing":[],"search":"integral_abs_restrict_le_of_uniformintegrable_one banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.integral_abs_restrict_le_of_uniformintegrable_one probability-theory uniform integrability at exponent one gives uniform absolute continuity of the expected norm over measurable events. this is a thin real-valued wrapper around mathlib's `unifintegrable` epsilon-delta interface. theorem compiled","shard":"modules/96497046a61986e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformAbsoluteContinuity","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformAbsoluteContinuity","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformAbsoluteContinuity","description":"Uniform absolute continuity of the exact uncapped stopped average realized behavior-regret family. One delta works for every schedule index and controls both the absolute signed set integral and the set integral of the absolute stopped regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-f72757b422c4/index.html#decl-3f58d4e40464","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","order":8897,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformAbsoluteContinuity (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_uniformabsolutecontinuity banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_uniformabsolutecontinuity uniform absolute continuity of the exact uncapped stopped average realized behavior-regret family. one delta works for every schedule index and controls both the absolute signed set integral and the set integral of the absolute stopped regret. theorem compiled","shard":"modules/96497046a61986e1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.uniformIntegrable_one_of_memLp_and_tendsto_eLpNorm_sub_zero","label":"uniformIntegrable_one_of_memLp_and_tendsto_eLpNorm_sub_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.uniformIntegrable_one_of_memLp_and_tendsto_eLpNorm_sub_zero","description":"L1 convergence to zero gives probability-theory uniform integrability. The `UnifIntegrable` component is Mathlib's Lp convergence theorem. The additional uniform L1 bound is obtained from boundedness of the convergent sequence in `Lp` and transported back through `MemLp.toLp`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-22fe344b7372/index.html#decl-2d983355b704","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","order":8898,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformIntegrable_one_of_memLp_and_tendsto_eLpNorm_sub_zero {Omega : Type w} [MeasurableSpace Omega] {mu : Measure Omega} {f : Nat -> Omega -> Real} (hf : forall n, MemLp (f n) 1 mu) (hfg : Tendsto (fun n => eLpNorm (f n - (fun _ => 0)) 1 mu) atTop (nhds 0)) : UniformIntegrable f 1 mu","missing":[],"search":"uniformintegrable_one_of_memlp_and_tendsto_elpnorm_sub_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.uniformintegrable_one_of_memlp_and_tendsto_elpnorm_sub_zero l1 convergence to zero gives probability-theory uniform integrability. the `unifintegrable` component is mathlib's lp convergence theorem. the additional uniform l1 bound is obtained from boundedness of the convergent sequence in `lp` and transported back through `memlp.tolp`. theorem compiled","shard":"modules/34fff1c8285532b7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.uniformIntegrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","label":"uniformIntegrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.uniformIntegrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","description":"The exact uncapped stopped average realized behavior-regret family is uniformly integrable in Mathlib's probability-theory sense at exponent one.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-22fe344b7372/index.html#decl-6a3188b1b761","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","order":8899,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformIntegrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSou…","missing":[],"search":"uniformintegrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.uniformintegrable_selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregret the exact uncapped stopped average realized behavior-regret family is uniformly integrable in mathlib's probability-theory sense at exponent one. theorem compiled","shard":"modules/34fff1c8285532b7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","description":"The signed expectation of the exact uncapped stopped process tends to zero. This is continuity of the Bochner integral under L1 convergence, not an optional-stopping identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-22fe344b7372/index.html#decl-49447ae0d196","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","order":8900,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency.lean:128"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState reward…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_tendsto_zero the signed expectation of the exact uncapped stopped process tends to zero. this is continuity of the bochner integral under l1 convergence, not an optional-stopping identity. theorem compiled","shard":"modules/34fff1c8285532b7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformIntegrable_and_integral_tendsto_zero","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformIntegrable_and_integral_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformIntegrable_and_integral_tendsto_zero","description":"The signed expectation of the exact uncapped stopped process tends to zero. This is continuity of the Bochner integral under L1 convergence, not an optional-stopping identity. -/ theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [Stan…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalinversesqrtthresholdunboundedhittingafte-22fe344b7372/index.html#decl-3b695dd30705","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","order":8901,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency.lean:193"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformIntegrable_and_integral_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_uniformintegrable_and_integral_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdunboundedhittingafter_stoppedaveragerealizedbehaviorregret_uniformintegrable_and_integral_tendsto_zero the signed expectation of the exact uncapped stopped process tends to zero. this is continuity of the bochner integral under l1 convergence, not an optional-stopping identity. -/ theorem selfconsistentschedulednaturalcausalinversesqrtthresholdunboundedhittingafterstoppedaveragerealizedbehaviorregretintegral_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloo…","shard":"modules/34fff1c8285532b7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.one_add_log_le_two_mul_sqrt","label":"one_add_log_le_two_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.one_add_log_le_two_mul_sqrt","description":"A positive real number satisfies `1 + log x <= 2 * sqrt x`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-5519c3ea9bef","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8902,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem one_add_log_le_two_mul_sqrt {x : Real} (hx : 0 < x) : 1 + Real.log x <= 2 * Real.sqrt x","missing":[],"search":"one_add_log_le_two_mul_sqrt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.one_add_log_le_two_mul_sqrt a positive real number satisfies `1 + log x <= 2 * sqrt x`. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_le_inverseSqrt","label":"selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_le_inverseSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_le_inverseSqrt","description":"The compiled logarithmic average rate admits an inverse-square-root bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-f6e0be92c997","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8903,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_le_inverseSqrt (mdp : MDP State Action) (rounds : Nat) (hrounds : 0 < rounds) : selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate mdp rounds <= 2 * selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient mdp / Real.sqrt (rounds : Real)","missing":[],"search":"selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_le_inversesqrt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausallogarithmicaverageintegratedbehaviorexpectedregretrate_le_inversesqrt the compiled logarithmic average rate admits an inverse-square-root bound. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Coefficient","label":"selfConsistentScheduledNaturalCausalRawWindowL1Coefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Coefficient","description":"Coefficient of the common inverse-square-root all-prefix L1 envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-04b0f68081e6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8904,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRawWindowL1Coefficient (mdp : MDP State Action) (varianceProxy : NNReal) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalrawwindowl1coefficient banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrawwindowl1coefficient coefficient of the common inverse-square-root all-prefix l1 envelope. definition compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope","label":"selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope","description":"Common inverse-square-root envelope used on every raw candidate prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-da7b0f708198","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8905,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope (mdp : MDP State Action) (varianceProxy : NNReal) (rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalrawwindowinversesqrtl1envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrawwindowinversesqrtl1envelope common inverse-square-root envelope used on every raw candidate prefix. definition compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Coefficient_nonneg","label":"selfConsistentScheduledNaturalCausalRawWindowL1Coefficient_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Coefficient_nonneg","description":"The common raw-window L1 coefficient is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-cf2374b15df1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8906,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:97"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRawWindowL1Coefficient_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) : 0 <= selfConsistentScheduledNaturalCausalRawWindowL1Coefficient mdp varianceProxy","missing":[],"search":"selfconsistentschedulednaturalcausalrawwindowl1coefficient_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrawwindowl1coefficient_nonneg the common raw-window l1 coefficient is nonnegative. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_le_rawWindowInverseSqrtL1Envelope","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_le_rawWindowInverseSqrtL1Envelope","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_le_rawWindowInverseSqrtL1Envelope","description":"The exact all-prefix L1 envelope is bounded by the common inverse square root.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-fe6714cd2ad5","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8907,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_le_rawWindowInverseSqrtL1Envelope (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (hrounds : 0 < rounds) : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope mdp varianceProxy baseVisitFloor rounds <= selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope mdp varianceProxy rounds","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_le_rawwindowinversesqrtl1envelope banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretl1envelope_le_rawwindowinversesqrtl1envelope the exact all-prefix l1 envelope is bounded by the common inverse square root. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_sq_le_sqrt_rounds_add","label":"explicitHighProbabilityScale_sq_le_sqrt_rounds_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_sq_le_sqrt_rounds_add","description":"A raw prefix after the fourth-power base has square root at least `(n+1)^2`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-13c9c5681b14","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8908,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:149"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityScale_sq_le_sqrt_rounds_add (scheduleIndex offset : Nat) : (explicitHighProbabilityScale scheduleIndex : Real) ^ 2 <= Real.sqrt (explicitHighProbabilityRounds scheduleIndex + offset : Nat)","missing":[],"search":"explicithighprobabilityscale_sq_le_sqrt_rounds_add banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityscale_sq_le_sqrt_rounds_add a raw prefix after the fourth-power base has square root at least `(n+1)^2`. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_inverseSquare","label":"selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_inverseSquare","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_inverseSquare","description":"Every candidate in the raw window costs at most one inverse-square term.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-6fd69140e4e4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8909,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:170"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_inverseSquare (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex offset : Nat) : selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope mdp varianceProxy (explicitHighProbabilityRounds scheduleIndex + offset) <= selfConsistentScheduledNaturalCausalRawWindowL1Coefficient mdp varianceProxy / (explicitHighProbabilityScale scheduleIndex : Real) ^ 2","missing":[],"search":"selfconsistentschedulednaturalcausalrawwindowinversesqrtl1envelope_add_le_inversesquare banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrawwindowinversesqrtl1envelope_add_le_inversesquare every candidate in the raw window costs at most one inverse-square term. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget","label":"selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget","description":"L1 budget over all raw prefixes from `(n+1)^4` through `(n+1)^4+n`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-4acf8306ed03","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8910,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget l1 budget over all raw prefixes from `(n+1)^4` through `(n+1)^4+n`. definition compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_nonneg","label":"selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_nonneg","description":"The finite raw-window L1 budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-fcc79665c0a7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8911,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:201"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget mdp varianceProxy baseVisitFloor scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget_nonneg the finite raw-window l1 budget is nonnegative. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_le_rate","label":"selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_le_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_le_rate","description":"The `n+1` raw candidates have total budget at most `D/(n+1)`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-73c44dd42a0c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8912,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_le_rate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget mdp varianceProxy baseVisitFloor scheduleIndex <= selfConsistentScheduledNaturalCausalRawWindowL1Coefficient mdp varianceProxy / (explicitHighProbabilityScale scheduleIndex : Real)","missing":[],"search":"selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget_le_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget_le_rate the `n+1` raw candidates have total budget at most `d/(n+1)`. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Rate_tendsto_zero","label":"selfConsistentScheduledNaturalCausalRawWindowL1Rate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Rate_tendsto_zero","description":"The explicit `D/(n+1)` raw-window budget rate tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-f446814276da","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8913,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:256"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRawWindowL1Rate_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (fun scheduleIndex => selfConsistentScheduledNaturalCausalRawWindowL1Coefficient mdp varianceProxy / (explicitHighProbabilityScale scheduleIndex : Real)) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalrawwindowl1rate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrawwindowl1rate_tendsto_zero the explicit `d/(n+1)` raw-window budget rate tends to zero. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_tendsto_zero","label":"selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_tendsto_zero","description":"The finite growing raw-window L1 budget tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-2f452f5fa3aa","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8914,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingl1budget_tendsto_zero the finite growing raw-window l1 budget tends to zero. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_polynomialBaseGrowingRawWindow_offset_untopA_eq","label":"exists_polynomialBaseGrowingRawWindow_offset_untopA_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_polynomialBaseGrowingRawWindow_offset_untopA_eq","description":"The WithTop bounds select one raw natural prefix in the growing window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-5b36188a51b7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8915,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:289"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exists_polynomialBaseGrowingRawWindow_offset_untopA_eq {Omega : Type*} (stoppingPrefix : Nat -> Omega -> WithTop Nat) (hstoppingLower : forall scheduleIndex trajectory, (explicitHighProbabilityRounds scheduleIndex : WithTop Nat) <= stoppingPrefix scheduleIndex trajectory) (hstoppingUpper : forall scheduleIndex trajectory, stoppingPrefix scheduleIndex trajectory <= (explicitHighProbabilityRounds scheduleIndex + scheduleIndex : WithTop Nat)) (scheduleIndex : Nat) (trajectory : Omega) : exists offset, offset ∈ Finset.range (scheduleIndex + 1) /\\ (stoppingPrefix scheduleIndex trajectory).untopA = explicitHighProbabilityRounds scheduleIndex + offset","missing":[],"search":"exists_polynomialbasegrowingrawwindow_offset_untopa_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exists_polynomialbasegrowingrawwindow_offset_untopa_eq the withtop bounds select one raw natural prefix in the growing window. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","description":"Expected absolute value of the polynomial-base raw-window stopped process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-ca578a87a325","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8916,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:308"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret expected absolute value of the polynomial-base raw-window stopped process. definition compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","description":"Every polynomial-base raw-window stopped coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-16aeafa2631f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8917,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:331"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialS…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret every polynomial-base raw-window stopped coordinate belongs to `l1`. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","description":"Expected absolute polynomial-base raw-window stopped regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-cee07ecaa553","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8918,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:371"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_nonneg expected absolute polynomial-base raw-window stopped regret is nonnegative. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","label":"selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","description":"The selected raw coordinate is bounded by the finite candidate L1 budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-a2e411d4f570","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8919,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:393"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_le_budget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_le_budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_le_budget the selected raw coordinate is bounded by the finite candidate l1 budget. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","description":"Expected absolute raw-window stopped regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-8aab7f93b831","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8920,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:526"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (f…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsolutepolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_tendsto_zero expected absolute raw-window stopped regret tends to zero. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_eq","description":"At exponent one, the raw-window stopped norm is its expected absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-8f186f94ff0f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8921,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:577"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp ini…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_eq at exponent one, the raw-window stopped norm is its expected absolute value. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","description":"The raw-window stopped process converges to zero in the exponent-one norm.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-3121ff81305a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8922,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:623"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory m…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero the raw-window stopped process converges to zero in the exponent-one norm. theorem compiled","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialBaseGrowingRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","label":"selfConsistentScheduledCausalSource_explicitPolynomialBaseGrowingRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialBaseGrowingRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","description":"The raw-window stopped process converges to zero in the exponent-one norm. -/ theorem eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel)…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalpolynomialbasegrowingrawwindowstoppingti-446f774eda58/index.html#decl-09516ce57f34","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8923,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:708"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_explicitPolynomialBaseGrowingRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory m…","missing":[],"search":"selfconsistentscheduledcausalsource_explicitpolynomialbasegrowingrawwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_explicitpolynomialbasegrowingrawwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency the raw-window stopped process converges to zero in the exponent-one norm. -/ theorem elpnorm_one_selfconsistentschedulednaturalcausalpolynomialbasegrowingrawwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistents…","shard":"modules/3da26b2e8ecfe6ac.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_apply_randomNat","label":"measurable_apply_randomNat","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_apply_randomNat","description":"A countable random coordinate of a measurable process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-75f09cefa8ae","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8924,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_apply_randomNat {Omega : Type w} {Beta : Type x} [MeasurableSpace Omega] [MeasurableSpace Beta] (process : Nat -> Omega -> Beta) (randomIndex : Omega -> Nat) (hprocess : forall n, Measurable (process n)) (hrandomIndex : Measurable randomIndex) : Measurable (fun omega => process (randomIndex omega) omega)","missing":[],"search":"measurable_apply_randomnat banditrlproof.measurable_apply_randomnat a countable random coordinate of a measurable process is measurable. theorem compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.ae_tendsto_apply_randomPrefix","label":"ae_tendsto_apply_randomPrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ae_tendsto_apply_randomPrefix","description":"An almost-everywhere limit survives an almost-everywhere diverging random prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-35d8357fccf9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8925,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:41"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem ae_tendsto_apply_randomPrefix {Omega : Type w} {Beta : Type x} [MeasurableSpace Omega] [TopologicalSpace Beta] {mu : Measure Omega} {process : Nat -> Omega -> Beta} {randomPrefix : Nat -> Omega -> Nat} {z : Beta} (hprocess : ∀ᵐ omega ∂mu, Tendsto (fun n => process n omega) atTop (nhds z)) (hrandomPrefix : ∀ᵐ omega ∂mu, Tendsto (fun n => randomPrefix n omega) atTop atTop) : ∀ᵐ omega ∂mu, Tendsto (fun n => process (randomPrefix n omega) omega) atTop (nhds z)","missing":[],"search":"ae_tendsto_apply_randomprefix banditrlproof.ae_tendsto_apply_randomprefix an almost-everywhere limit survives an almost-everywhere diverging random prefix. theorem compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.tendsto_randomPrefix_atTop_of_nat_le","label":"tendsto_randomPrefix_atTop_of_nat_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tendsto_randomPrefix_atTop_of_nat_le","description":"A pointwise deterministic lower envelope forces a Nat-valued random prefix to diverge.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-f9ac5edf40cf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8926,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem tendsto_randomPrefix_atTop_of_nat_le {Omega : Type w} (randomPrefix : Nat -> Omega -> Nat) (hlower : forall n omega, n <= randomPrefix n omega) (omega : Omega) : Tendsto (fun n => randomPrefix n omega) atTop atTop","missing":[],"search":"tendsto_randomprefix_attop_of_nat_le banditrlproof.tendsto_randomprefix_attop_of_nat_le a pointwise deterministic lower envelope forces a nat-valued random prefix to diverge. theorem compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","label":"selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","description":"The exact natural average realized behavior regret at a random prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-c1b7cd6123e0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8927,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (randomPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess the exact natural average realized behavior regret at a random prefix. definition compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess_apply","label":"selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess_apply","description":"theorem selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (randomPrefix : Nat -> HeterogeneousStochasticEpisodeBat…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-4ac0ec5877b0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8928,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (randomPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultS…","missing":[],"search":"selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess_apply theorem selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (randomprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> nat) (scheduleindex : nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) : selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor randomprefix scheduleindex trajectory = selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor (rand…","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","description":"Every coordinate of the exact random-prefix process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-fff307229af9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8929,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:127"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (randomPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Nat) (scheduleIndex : Nat) (hrandomPrefix : Measurable (randomPrefix scheduleIndex)) : Measurable (selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor randomPrefix scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalrandomprefixaveragerealizedbehaviorregretprocess every coordinate of the exact random-prefix process is measurable. theorem compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","description":"The exact natural average realized behavior regret remains almost-surely consistent at every measurable random prefix which diverges almost everywhere.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-c97fe787ddc6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8930,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (randomPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStoch…","missing":[],"search":"selfconsistentscheduledcausalsource_randomprefixnaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_randomprefixnaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero the exact natural average realized behavior regret remains almost-surely consistent at every measurable random prefix which diverges almost everywhere. theorem compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","label":"selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","description":"A pointwise lower envelope is a practical sufficient random-prefix contract.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrandomprefixaveragerealizedbehaviorregre-1304969ada20/index.html#decl-594dae5947a7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","order":8931,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency.lean:221"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (randomPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => Ada…","missing":[],"search":"selfconsistentscheduledcausalsource_randomprefixnaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero_of_nat_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_randomprefixnaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero_of_nat_le a pointwise lower envelope is a practical sufficient random-prefix contract. theorem compiled","shard":"modules/f878fac0ec6429cf.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget","label":"selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget","description":"Sum of all coordinate L1 envelopes in a parameterized raw window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-e5b8d50e5c58","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8932,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (baseRounds windowWidth : Nat -> Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget sum of all coordinate l1 envelopes in a parameterized raw window. definition compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate","label":"selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate","description":"Candidate-count times inverse-square-root rate for a raw window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-d5e98ea2db8b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8933,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate (mdp : MDP State Action) (varianceProxy : NNReal) (baseRounds windowWidth : Nat -> Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1rate candidate-count times inverse-square-root rate for a raw window. definition compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_baseRounds_le_sqrt_add","label":"sqrt_baseRounds_le_sqrt_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_baseRounds_le_sqrt_add","description":"Adding a natural offset cannot decrease the square root of a base prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-35c84dd91670","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8934,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sqrt_baseRounds_le_sqrt_add (baseRounds offset : Nat) : Real.sqrt (baseRounds : Real) <= Real.sqrt (baseRounds + offset : Nat)","missing":[],"search":"sqrt_baserounds_le_sqrt_add banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.sqrt_baserounds_le_sqrt_add adding a natural offset cannot decrease the square root of a base prefix. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_base","label":"selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_base","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_base","description":"Every candidate envelope is controlled by the inverse square root of its base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-78d97638db67","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8935,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_base (mdp : MDP State Action) (varianceProxy : NNReal) (baseRounds offset : Nat) (hbaseRounds : 0 < baseRounds) : selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope mdp varianceProxy (baseRounds + offset) <= selfConsistentScheduledNaturalCausalRawWindowL1Coefficient mdp varianceProxy / Real.sqrt (baseRounds : Real)","missing":[],"search":"selfconsistentschedulednaturalcausalrawwindowinversesqrtl1envelope_add_le_base banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrawwindowinversesqrtl1envelope_add_le_base every candidate envelope is controlled by the inverse square root of its base. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_nonneg","label":"selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_nonneg","description":"A parameterized finite raw-window L1 budget is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-39fe80d8d856","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8936,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_nonneg (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (baseRounds windowWidth : Nat -> Nat) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget mdp varianceProxy baseVisitFloor baseRounds windowWidth scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget_nonneg a parameterized finite raw-window l1 budget is nonnegative. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_le_rate","label":"selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_le_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_le_rate","description":"The candidate sum is bounded by candidate count times the base envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-5821c7564196","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8937,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_le_rate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (baseRounds windowWidth : Nat -> Nat) (hbaseRounds : forall scheduleIndex, 0 < baseRounds scheduleIndex) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget mdp varianceProxy baseVisitFloor baseRounds windowWidth scheduleIndex <= selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate mdp varianceProxy baseRounds windowWidth scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget_le_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget_le_rate the candidate sum is bounded by candidate count times the base envelope. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate_tendsto_zero","label":"selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate_tendsto_zero","description":"A vanishing candidate-count/base-square-root ratio gives a vanishing rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-89619a885f72","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8938,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseRounds windowWidth : Nat -> Nat) (hcandidateRate : Tendsto (fun scheduleIndex => (((windowWidth scheduleIndex + 1 : Nat) : Real) / Real.sqrt (baseRounds scheduleIndex : Real))) atTop (nhds 0)) : Tendsto (selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate mdp varianceProxy baseRounds windowWidth) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1rate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1rate_tendsto_zero a vanishing candidate-count/base-square-root ratio gives a vanishing rate. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_tendsto_zero","label":"selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_tendsto_zero","description":"The parameterized raw-window L1 budget tends to zero under the rate contract.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-9c054b41b70d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8939,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (baseRounds windowWidth : Nat -> Nat) (hbaseRounds : forall scheduleIndex, 0 < baseRounds scheduleIndex) (hcandidateRate : Tendsto (fun scheduleIndex => (((windowWidth scheduleIndex + 1 : Nat) : Real) / Real.sqrt (baseRounds scheduleIndex : Real))) atTop (nhds 0)) : Tendsto (selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget mdp varianceProxy baseVisitFloor baseRounds windowWidth) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingl1budget_tendsto_zero the parameterized raw-window l1 budget tends to zero under the rate contract. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_rateControlledRawWindow_offset_untopA_eq","label":"exists_rateControlledRawWindow_offset_untopA_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_rateControlledRawWindow_offset_untopA_eq","description":"WithTop bounds select one candidate in a parameterized raw window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-45112ecbc3d0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8940,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:189"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exists_rateControlledRawWindow_offset_untopA_eq {Omega : Type*} (stoppingPrefix : Nat -> Omega -> WithTop Nat) (baseRounds windowWidth : Nat -> Nat) (hstoppingLower : forall scheduleIndex trajectory, (baseRounds scheduleIndex : WithTop Nat) <= stoppingPrefix scheduleIndex trajectory) (hstoppingUpper : forall scheduleIndex trajectory, stoppingPrefix scheduleIndex trajectory <= (baseRounds scheduleIndex + windowWidth scheduleIndex : WithTop Nat)) (scheduleIndex : Nat) (trajectory : Omega) : exists offset, offset ∈ Finset.range (windowWidth scheduleIndex + 1) /\\ (stoppingPrefix scheduleIndex trajectory).untopA = baseRounds scheduleIndex + offset","missing":[],"search":"exists_ratecontrolledrawwindow_offset_untopa_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.exists_ratecontrolledrawwindow_offset_untopa_eq withtop bounds select one candidate in a parameterized raw window. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRawWindowCandidateRate_tendsto_zero","label":"explicitHighProbabilityRawWindowCandidateRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRawWindowCandidateRate_tendsto_zero","description":"The fourth-power base with width `n` satisfies the generic rate contract.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-9c7c458843b2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8941,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:211"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRawWindowCandidateRate_tendsto_zero : Tendsto (fun scheduleIndex => (((scheduleIndex + 1 : Nat) : Real) / Real.sqrt (explicitHighProbabilityRounds scheduleIndex : Real))) atTop (nhds 0)","missing":[],"search":"explicithighprobabilityrawwindowcandidaterate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrawwindowcandidaterate_tendsto_zero the fourth-power base with width `n` satisfies the generic rate contract. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityDoubleLinearRawWindowCandidateRate_tendsto_zero","label":"explicitHighProbabilityDoubleLinearRawWindowCandidateRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityDoubleLinearRawWindowCandidateRate_tendsto_zero","description":"The fourth-power base also supports the strictly wider raw width `2*n+1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-8af32251da15","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8942,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:247"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityDoubleLinearRawWindowCandidateRate_tendsto_zero : Tendsto (fun scheduleIndex => ((((2 * scheduleIndex + 1) + 1 : Nat) : Real) / Real.sqrt (explicitHighProbabilityRounds scheduleIndex : Real))) atTop (nhds 0)","missing":[],"search":"explicithighprobabilitydoublelinearrawwindowcandidaterate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilitydoublelinearrawwindowcandidaterate_tendsto_zero the fourth-power base also supports the strictly wider raw width `2*n+1`. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","description":"Expected absolute value of a rate-controlled raw-window stopped process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-de9d040c1277","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8943,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret expected absolute value of a rate-controlled raw-window stopped process. definition compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","label":"memLp_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","description":"Every rate-controlled raw-window stopped coordinate belongs to `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-8d15b8f5302d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8944,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:305"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem memLp_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (baseRounds windowWidth : Nat -> Nat) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCausalTr…","missing":[],"search":"memlp_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.memlp_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret every rate-controlled raw-window stopped coordinate belongs to `l1`. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","description":"Expected absolute rate-controlled stopped regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-931214b32f93","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8945,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:345"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : 0 <= selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor stoppingPrefix scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_nonneg expected absolute rate-controlled stopped regret is nonnegative. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","description":"The selected raw coordinate is bounded by the full finite candidate budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-aa86ea06d738","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8946,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:367"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_le_budget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (baseRounds windowWidth : Nat -> Nat) (stoppingPrefix : Nat -> HeterogeneousStochasticE…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_le_budget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_le_budget the selected raw coordinate is bounded by the full finite candidate budget. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","label":"selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","description":"Expected absolute rate-controlled stopped regret tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-1dc27458c6c1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8947,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:493"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (baseRounds windowWidth : Nat -> Nat) (hbaseRounds : forall scheduleIndex, 0 < baseR…","missing":[],"search":"selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalexpectedabsoluteratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_tendsto_zero expected absolute rate-controlled stopped regret tends to zero. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_eq","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_eq","description":"At exponent one, the stopped norm is its expected absolute value.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-e54bcf31cd64","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8948,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:551"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (baseRounds windowWidth : Nat -> Nat) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (hstopping : forall scheduleIndex, IsStoppingTime (selfConsistentScheduledNaturalCau…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_eq at exponent one, the stopped norm is its expected absolute value. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","label":"eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","description":"The rate-controlled stopped process converges to zero in `L1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-3ea447aa4ed1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8949,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:597"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (baseRounds windowWidth : Nat -> Nat) (hbaseRounds : forall scheduleIndex, 0 <…","missing":[],"search":"elpnorm_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.elpnorm_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero the rate-controlled stopped process converges to zero in `l1`. theorem compiled","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_rateControlledRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","label":"selfConsistentScheduledCausalSource_rateControlledRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_rateControlledRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","description":"The rate-controlled stopped process converges to zero in `L1`. -/ theorem eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NN…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalratecontrolledrawwindowstoppingtimel1ave-142f68002d96/index.html#decl-6829a096bf73","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":8950,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:689"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_rateControlledRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (baseRounds windowWidth : Nat -> Nat) (hbaseRounds : forall scheduleIndex, 0 < baseRoun…","missing":[],"search":"selfconsistentscheduledcausalsource_ratecontrolledrawwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_ratecontrolledrawwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency the rate-controlled stopped process converges to zero in `l1`. -/ theorem elpnorm_one_selfconsistentschedulednaturalcausalratecontrolledrawwindowstoppingaveragerealizedbehaviorregret_sub_zero_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (baserounds windowwidth : nat -> nat) (hbaserounds : forall scheduleindex, 0 < baserounds scheduleindex) (hcandidaterate : tendsto (fun scheduleindex => (((windowwidth scheduleindex + 1…","shard":"modules/9f887de0e0743805.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeBehaviorExpectedRegretLogarithmicRate","label":"selfConsistentScheduledNaturalCausalBurninCumulativeBehaviorExpectedRegretLogarithmicRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeBehaviorExpectedRegretLogarithmicRate","description":"Uniform burn-in charge plus the compiled full logarithmic planning sum.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-6bdbd73420c3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8951,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninCumulativeBehaviorExpectedRegretLogarithmicRate (mdp : MDP State Action) (burnin rounds : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalburnincumulativebehaviorexpectedregretlogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburnincumulativebehaviorexpectedregretlogarithmicrate uniform burn-in charge plus the compiled full logarithmic planning sum. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelBadEvent","description":"Outside the infinite model tail, only the first `burnin` natural rounds need the uniform `2 * horizon` charge. Every later round uses its actual coordinate model certificate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-7ca43b84808d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8952,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (trajectory : HeterogeneousStoch…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_burnin_logarithmic_of_not_mem_tailmodelbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativebehaviorexpectedregretprocess_le_burnin_logarithmic_of_not_mem_tailmodelbadevent outside the infinite model tail, only the first `burnin` natural rounds need the uniform `2 * horizon` charge. every later round uses its actual coordinate model certificate. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninTailModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalBurninTailModelReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninTailModelReturnBadEvent","description":"Infinite model tail after `burnin`, union the fixed-prefix normalized return deviation event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-6a0abd7e7bfc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8953,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:183"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninTailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalburnintailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburnintailmodelreturnbadevent infinite model tail after `burnin`, union the fixed-prefix normalized return deviation event. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninTailModelReturnFailureBudget","label":"selfConsistentScheduledNaturalCausalBurninTailModelReturnFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninTailModelReturnFailureBudget","description":"Exact infinite-tail model share plus caller-supplied return share.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-afd938ebdf00","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8954,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:201"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninTailModelReturnFailureBudget (mdp : MDP State Action) (burnin : Nat) (returnDelta : Real) : ENNReal","missing":[],"search":"selfconsistentschedulednaturalcausalburnintailmodelreturnfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburnintailmodelreturnfailurebudget exact infinite-tail model share plus caller-supplied return share. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalBurninTailModelReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_naturalBurninTailModelReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalBurninTailModelReturnBadEvent_le","description":"The joint burn-in tail event is measurable and has its exact union budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-1306f29e2d58","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8955,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:208"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_naturalBurninTailModelReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (burnin rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDelta_le_one : returnDelta <= 1) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_naturalburnintailmodelreturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_naturalburnintailmodelreturnbadevent_le the joint burn-in tail event is measurable and has its exact union budget. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninRealizedCumulativeLogarithmicRate","label":"selfConsistentScheduledNaturalCausalBurninRealizedCumulativeLogarithmicRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninRealizedCumulativeLogarithmicRate","description":"Burn-in expected-regret envelope plus normalized return radius.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-490d4a84aaf9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8956,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:281"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninRealizedCumulativeLogarithmicRate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalburninrealizedcumulativelogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburninrealizedcumulativelogarithmicrate burn-in expected-regret envelope plus normalized return radius. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninRealizedAverageLogarithmicRate","label":"selfConsistentScheduledNaturalCausalBurninRealizedAverageLogarithmicRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninRealizedAverageLogarithmicRate","description":"Positive-round average form of the burn-in realized rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-c2db6bc043fb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8957,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninRealizedAverageLogarithmicRate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalburninrealizedaveragelogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburninrealizedaveragelogarithmicrate positive-round average form of the burn-in realized rate. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","description":"Outside the tail-model/return union, cumulative successor-batch-average realized behavior regret obeys the burn-in logarithmic envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-bf05517cd91c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8958,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:306"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (returnDelta : Real) (traj…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess_le_burnin_logarithmic_of_not_mem_tailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess_le_burnin_logarithmic_of_not_mem_tailmodelreturnbadevent outside the tail-model/return union, cumulative successor-batch-average realized behavior regret obeys the burn-in logarithmic envelope. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","description":"Joint-good paths also obey the positive-round average burn-in rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-dc6e58f8cd9a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8959,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:405"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (hrounds : 0 < rounds) (retur…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_burnin_logarithmic_of_not_mem_tailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_burnin_logarithmic_of_not_mem_tailmodelreturnbadevent joint-good paths also obey the positive-round average burn-in rate. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","label":"selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","description":"One-sided cumulative burn-in realized-regret violation set.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-e787e2ca51aa","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8960,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:450"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalburnincumulativerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburnincumulativerealizedbehaviorregretlogarithmicviolationset one-sided cumulative burn-in realized-regret violation set. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","label":"selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","description":"One-sided average burn-in realized-regret violation set.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-8da93e8e65b6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8961,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:470"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset one-sided average burn-in realized-regret violation set. definition compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","description":"The cumulative burn-in violation set is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-1c6297746f8f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8962,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:490"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : MeasurableSet (selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor burnin rounds returnDelta)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalburnincumulativerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalburnincumulativerealizedbehaviorregretlogarithmicviolationset the cumulative burn-in violation set is measurable. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","description":"The average burn-in violation set is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-76037523b053","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8963,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:510"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (burnin rounds : Nat) (returnDelta : Real) : MeasurableSet (selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor burnin rounds returnDelta)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset the average burn-in violation set is measurable. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","description":"Every cumulative burn-in violation belongs to the joint tail event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-e3de9fb25a85","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8964,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:530"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (returnDelta : Real) : selfCon…","missing":[],"search":"selfconsistentschedulednaturalcausalburnincumulativerealizedbehaviorregretlogarithmicviolationset_subset_tailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburnincumulativerealizedbehaviorregretlogarithmicviolationset_subset_tailmodelreturnbadevent every cumulative burn-in violation belongs to the joint tail event. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","description":"Every average burn-in violation belongs to the joint tail event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-75994bd2b976","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8965,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:564"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (hrounds : 0 < rounds) (returnDel…","missing":[],"search":"selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset_subset_tailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset_subset_tailmodelreturnbadevent every average burn-in violation belongs to the joint tail event. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_burninTailHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_burninTailHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_burninTailHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","description":"Terminal burn-in tail high-probability logarithmic cumulative and average realized behavior-regret certificate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityburninlograte/index.html#decl-e59bb5b89942","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","order":8966,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate.lean:601"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_burninTailHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (burnin rounds : Nat) (hburnin : burnin <= rounds) (hrounds : 0 < rounds) (returnDelta : Real) (hr…","missing":[],"search":"selfconsistentscheduledcausalsource_burnintailhighprobabilitylogarithmiccumulativeaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_burnintailhighprobabilitylogarithmiccumulativeaveragerealizedbehaviorregret terminal burn-in tail high-probability logarithmic cumulative and average realized behavior-regret certificate. theorem compiled","shard":"modules/a63cc67926ea20ec.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy_le_rounds_mul","label":"naturalCumulativeSuccessorAverageReturnVarianceProxy_le_rounds_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy_le_rounds_mul","description":"Each positive-count normalized successor coordinate contributes at most one copy of the one-episode global return proxy.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-848e8ed7bbfc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8967,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalCumulativeSuccessorAverageReturnVarianceProxy_le_rounds_mul (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hepisodes : forall n, 0 < episodes n) : naturalCumulativeSuccessorAverageReturnVarianceProxy mdp episodes rounds rewardBound rewardVarianceProxy <= (rounds : NNReal) * mdp.globalReturnDeviationPerEpisodeVarianceProxy rewardBound rewardVarianceProxy","missing":[],"search":"naturalcumulativesuccessoraveragereturnvarianceproxy_le_rounds_mul banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativesuccessoraveragereturnvarianceproxy_le_rounds_mul each positive-count normalized successor coordinate contributes at most one copy of the one-episode global return proxy. theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale","label":"explicitHighProbabilityScale","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale","description":"Positive scale used by the explicit prefix schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-5dd88fa424ec","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8968,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def explicitHighProbabilityScale (n : Nat) : Nat","missing":[],"search":"explicithighprobabilityscale banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityscale positive scale used by the explicit prefix schedule. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityBurnin","label":"explicitHighProbabilityBurnin","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityBurnin","description":"The model-tail burn-in is linear in the schedule scale.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-9935fd96c97b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8969,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:95"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def explicitHighProbabilityBurnin (n : Nat) : Nat","missing":[],"search":"explicithighprobabilityburnin banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityburnin the model-tail burn-in is linear in the schedule scale. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds","label":"explicitHighProbabilityRounds","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds","description":"Natural successor prefixes are sampled along a fourth-power subsequence.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-287ab5991bc0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8970,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def explicitHighProbabilityRounds (n : Nat) : Nat","missing":[],"search":"explicithighprobabilityrounds banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds natural successor prefixes are sampled along a fourth-power subsequence. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta","label":"explicitHighProbabilityReturnDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta","description":"Exponentially vanishing fixed-prefix return failure share.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-65e025bcc3db","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8971,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:103"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitHighProbabilityReturnDelta (n : Nat) : Real","missing":[],"search":"explicithighprobabilityreturndelta banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityreturndelta exponentially vanishing fixed-prefix return failure share. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_pos","label":"explicitHighProbabilityScale_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_pos","description":"theorem explicitHighProbabilityScale_pos (n : Nat) : 0 < explicitHighProbabilityScale n","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-67d620612c27","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8972,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityScale_pos (n : Nat) : 0 < explicitHighProbabilityScale n","missing":[],"search":"explicithighprobabilityscale_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityscale_pos theorem explicithighprobabilityscale_pos (n : nat) : 0 < explicithighprobabilityscale n theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityBurnin_le_rounds","label":"explicitHighProbabilityBurnin_le_rounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityBurnin_le_rounds","description":"theorem explicitHighProbabilityBurnin_le_rounds (n : Nat) : explicitHighProbabilityBurnin n <= explicitHighProbabilityRounds n","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-7bd7f308ec02","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8973,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:110"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityBurnin_le_rounds (n : Nat) : explicitHighProbabilityBurnin n <= explicitHighProbabilityRounds n","missing":[],"search":"explicithighprobabilityburnin_le_rounds banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityburnin_le_rounds theorem explicithighprobabilityburnin_le_rounds (n : nat) : explicithighprobabilityburnin n <= explicithighprobabilityrounds n theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_pos","label":"explicitHighProbabilityRounds_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_pos","description":"theorem explicitHighProbabilityRounds_pos (n : Nat) : 0 < explicitHighProbabilityRounds n","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-6313b5087175","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8974,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_pos (n : Nat) : 0 < explicitHighProbabilityRounds n","missing":[],"search":"explicithighprobabilityrounds_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_pos theorem explicithighprobabilityrounds_pos (n : nat) : 0 < explicithighprobabilityrounds n theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_pos","label":"explicitHighProbabilityReturnDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_pos","description":"theorem explicitHighProbabilityReturnDelta_pos (n : Nat) : 0 < explicitHighProbabilityReturnDelta n","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-6c967b6d02db","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8975,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:120"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityReturnDelta_pos (n : Nat) : 0 < explicitHighProbabilityReturnDelta n","missing":[],"search":"explicithighprobabilityreturndelta_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityreturndelta_pos theorem explicithighprobabilityreturndelta_pos (n : nat) : 0 < explicithighprobabilityreturndelta n theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_le_one","label":"explicitHighProbabilityReturnDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_le_one","description":"theorem explicitHighProbabilityReturnDelta_le_one (n : Nat) : explicitHighProbabilityReturnDelta n <= 1","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-128b6bd19b04","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8976,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityReturnDelta_le_one (n : Nat) : explicitHighProbabilityReturnDelta n <= 1","missing":[],"search":"explicithighprobabilityreturndelta_le_one banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityreturndelta_le_one theorem explicithighprobabilityreturndelta_le_one (n : nat) : explicithighprobabilityreturndelta n <= 1 theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_tendsto_atTop","label":"explicitHighProbabilityScale_tendsto_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_tendsto_atTop","description":"theorem explicitHighProbabilityScale_tendsto_atTop : Tendsto explicitHighProbabilityScale atTop atTop","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-48fa54b259c2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8977,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:129"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityScale_tendsto_atTop : Tendsto explicitHighProbabilityScale atTop atTop","missing":[],"search":"explicithighprobabilityscale_tendsto_attop banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityscale_tendsto_attop theorem explicithighprobabilityscale_tendsto_attop : tendsto explicithighprobabilityscale attop attop theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_real_tendsto_atTop","label":"explicitHighProbabilityScale_real_tendsto_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_real_tendsto_atTop","description":"theorem explicitHighProbabilityScale_real_tendsto_atTop : Tendsto (fun n => (explicitHighProbabilityScale n : Real)) atTop atTop","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-c1235941a2ee","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8978,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:134"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityScale_real_tendsto_atTop : Tendsto (fun n => (explicitHighProbabilityScale n : Real)) atTop atTop","missing":[],"search":"explicithighprobabilityscale_real_tendsto_attop banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityscale_real_tendsto_attop theorem explicithighprobabilityscale_real_tendsto_attop : tendsto (fun n => (explicithighprobabilityscale n : real)) attop attop theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_tendsto_atTop","label":"explicitHighProbabilityRounds_tendsto_atTop","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_tendsto_atTop","description":"theorem explicitHighProbabilityRounds_tendsto_atTop : Tendsto explicitHighProbabilityRounds atTop atTop","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-5e282f8e9848","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8979,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:139"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityRounds_tendsto_atTop : Tendsto explicitHighProbabilityRounds atTop atTop","missing":[],"search":"explicithighprobabilityrounds_tendsto_attop banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityrounds_tendsto_attop theorem explicithighprobabilityrounds_tendsto_attop : tendsto explicithighprobabilityrounds attop attop theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_tendsto_zero","label":"explicitHighProbabilityReturnDelta_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_tendsto_zero","description":"theorem explicitHighProbabilityReturnDelta_tendsto_zero : Tendsto explicitHighProbabilityReturnDelta atTop (nhds 0)","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-dd8ed50ec2c3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8980,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitHighProbabilityReturnDelta_tendsto_zero : Tendsto explicitHighProbabilityReturnDelta atTop (nhds 0)","missing":[],"search":"explicithighprobabilityreturndelta_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicithighprobabilityreturndelta_tendsto_zero theorem explicithighprobabilityreturndelta_tendsto_zero : tendsto explicithighprobabilityreturndelta attop (nhds 0) theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy_le_rounds_mul","label":"selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy_le_rounds_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy_le_rounds_mul","description":"Self-consistent specialization of the generic own-count proxy bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-4744c12a046b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8981,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:155"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy_le_rounds_mul (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy mdp varianceProxy baseVisitFloor rounds <= (rounds : NNReal) * mdp.globalReturnDeviationPerEpisodeVarianceProxy 1 varianceProxy","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativereturnvarianceproxy_le_rounds_mul banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativereturnvarianceproxy_le_rounds_mul self-consistent specialization of the generic own-count proxy bound. theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius","label":"explicitPolynomialPrefixAverageReturnRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius","description":"Return-confidence contribution after division by the scheduled prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-2432173c7de8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8982,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:174"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageReturnRadius (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"explicitpolynomialprefixaveragereturnradius banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragereturnradius return-confidence contribution after division by the scheduled prefix. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius_le","label":"explicitPolynomialPrefixAverageReturnRadius_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius_le","description":"A coarse `O(1 / scale)` envelope for the scheduled average radius.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-bffa2469ee8e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8983,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:187"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageReturnRadius_le (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : explicitPolynomialPrefixAverageReturnRadius mdp varianceProxy baseVisitFloor n <= 2 * (((mdp.globalReturnDeviationPerEpisodeVarianceProxy 1 varianceProxy : NNReal) : Real) + 1) / (explicitHighProbabilityScale n : Real)","missing":[],"search":"explicitpolynomialprefixaveragereturnradius_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragereturnradius_le a coarse `o(1 / scale)` envelope for the scheduled average radius. theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius_tendsto_zero","label":"explicitPolynomialPrefixAverageReturnRadius_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius_tendsto_zero","description":"The explicit fixed-prefix return confidence contribution vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-fc8c0502cde4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8984,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageReturnRadius_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (explicitPolynomialPrefixAverageReturnRadius mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"explicitpolynomialprefixaveragereturnradius_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragereturnradius_tendsto_zero the explicit fixed-prefix return confidence contribution vanishes. theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnFailureBudget","label":"explicitPolynomialPrefixTailModelReturnFailureBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnFailureBudget","description":"Exact tail-model plus return-share budget along the explicit schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-7fb507b529ed","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8985,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:329"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixTailModelReturnFailureBudget (mdp : MDP State Action) (n : Nat) : ENNReal","missing":[],"search":"explicitpolynomialprefixtailmodelreturnfailurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixtailmodelreturnfailurebudget exact tail-model plus return-share budget along the explicit schedule. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate","description":"Scheduled positive-prefix average realized behavior-regret envelope.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-dcba22243b5b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8986,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:336"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretRate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Real","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretrate scheduled positive-prefix average realized behavior-regret envelope. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnFailureBudget_tendsto_zero","label":"explicitPolynomialPrefixTailModelReturnFailureBudget_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnFailureBudget_tendsto_zero","description":"The exact union-bound failure budget vanishes along the schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-2ab8850233ad","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8987,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:349"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixTailModelReturnFailureBudget_tendsto_zero (mdp : MDP State Action) : Tendsto (explicitPolynomialPrefixTailModelReturnFailureBudget mdp) atTop (nhds 0)","missing":[],"search":"explicitpolynomialprefixtailmodelreturnfailurebudget_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixtailmodelreturnfailurebudget_tendsto_zero the exact union-bound failure budget vanishes along the schedule. theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_tendsto_zero","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_tendsto_zero","description":"The full scheduled average realized-regret envelope vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-71b2c17fd9a4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8988,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:374"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) : Tendsto (explicitPolynomialPrefixAverageRealizedBehaviorRegretRate mdp varianceProxy baseVisitFloor) atTop (nhds 0)","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretrate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretrate_tendsto_zero the full scheduled average realized-regret envelope vanishes. theorem compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnBadEvent","label":"explicitPolynomialPrefixTailModelReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnBadEvent","description":"Scheduled model-tail/normalized-return union event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-ee36a880d44e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8989,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:439"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixTailModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"explicitpolynomialprefixtailmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixtailmodelreturnbadevent scheduled model-tail/normalized-return union event. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","description":"Scheduled one-sided average realized behavior-regret violation set.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-f895d7145382","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8990,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:457"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretviolationset scheduled one-sided average realized behavior-regret violation set. definition compiled","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixHighProbabilityAverageRealizedBehaviorRegretConsistency","label":"selfConsistentScheduledCausalSource_explicitPolynomialPrefixHighProbabilityAverageRealizedBehaviorRegretConsistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixHighProbabilityAverageRealizedBehaviorRegretConsistency","description":"Scheduled one-sided average realized behavior-regret violation set. -/ noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real)…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilityexplicitschedule/index.html#decl-2f3f6bef3c02","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","order":8991,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule.lean:479"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_explicitPolynomialPrefixHighProbabilityAverageRealizedBehaviorRegretConsistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable…","missing":[],"search":"selfconsistentscheduledcausalsource_explicitpolynomialprefixhighprobabilityaveragerealizedbehaviorregretconsistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_explicitpolynomialprefixhighprobabilityaveragerealizedbehaviorregretconsistency scheduled one-sided average realized behavior-regret violation set. -/ noncomputable def explicitpolynomialprefixaveragerealizedbehaviorregretviolationset (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (n : nat) : set (heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) := selfconsistentschedulednaturalcausalburninaveragerealizedbehaviorregretlogarithmicviolationset mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor (explicithighprobabilityburnin n) (explicithighprobabilityrounds n) (explicithighprobabilityreturndelta n) /- terminal scheduled-prefix high-probability average realized behavior-regret consistency certificate. it packages every fixed-prefix event and pathwise certificate together with the two vanish…","shard":"modules/bfa07784d0928c6f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement","label":"naturalSuccessorAverageReturnDeviationIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement","description":"Successor return deviation divided by the actual successor batch size. Coordinate zero is a dummy zero so coordinate `n + 1` remains conditioned on the prefix filtration at `n`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-bc45cca8197d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8992,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalSuccessorAverageReturnDeviationIncrement {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (round : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"naturalsuccessoraveragereturndeviationincrement banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessoraveragereturndeviationincrement successor return deviation divided by the actual successor batch size. coordinate zero is a dummy zero so coordinate `n + 1` remains conditioned on the prefix filtration at `n`. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnVarianceProxyAt","label":"naturalSuccessorAverageReturnVarianceProxyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnVarianceProxyAt","description":"Exact square-scaled conditional proxy for the normalized increment.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-0f6bde49bdf0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8993,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalSuccessorAverageReturnVarianceProxyAt (mdp : MDP State Action) (episodes : Nat -> Nat) (round : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"naturalsuccessoraveragereturnvarianceproxyat banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessoraveragereturnvarianceproxyat exact square-scaled conditional proxy for the normalized increment. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy","label":"naturalCumulativeSuccessorAverageReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy","description":"Total proxy for normalized successor sample-average deviations.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-9a325a089277","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8994,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalCumulativeSuccessorAverageReturnVarianceProxy (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"naturalcumulativesuccessoraveragereturnvarianceproxy banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativesuccessoraveragereturnvarianceproxy total proxy for normalized successor sample-average deviations. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnDeviation","label":"naturalCumulativeSuccessorAverageReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnDeviation","description":"Cumulative normalized deviation over natural successor rounds.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-42a9cb11ba03","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8995,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalCumulativeSuccessorAverageReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) : Real","missing":[],"search":"naturalcumulativesuccessoraveragereturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativesuccessoraveragereturndeviation cumulative normalized deviation over natural successor rounds. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement_stronglyAdapted_piLE","label":"naturalSuccessorAverageReturnDeviationIncrement_stronglyAdapted_piLE","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement_stronglyAdapted_piLE","description":"The normalized successor return process is strongly adapted to `piLE`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-8b9313b55754","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8996,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:101"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalSuccessorAverageReturnDeviationIncrement_stronglyAdapted_piLE {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] : StronglyAdapted (Filtration.piLE (X := fun n : Nat => StochasticEpisodeBatch mdp (episodes n))) source.naturalSuccessorAverageReturnDeviationIncrement","missing":[],"search":"naturalsuccessoraveragereturndeviationincrement_stronglyadapted_pile banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessoraveragereturndeviationincrement_stronglyadapted_pile the normalized successor return process is strongly adapted to `pile`. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeSuccessorAverageReturnDeviation","label":"measurable_naturalCumulativeSuccessorAverageReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeSuccessorAverageReturnDeviation","description":"The normalized cumulative return deviation is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-a0e977810182","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8997,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:129"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalCumulativeSuccessorAverageReturnDeviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) : Measurable (source.naturalCumulativeSuccessorAverageReturnDeviation rounds)","missing":[],"search":"measurable_naturalcumulativesuccessoraveragereturndeviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalcumulativesuccessoraveragereturndeviation the normalized cumulative return deviation is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement_succ_hasCondSubgaussianMGF","label":"naturalSuccessorAverageReturnDeviationIncrement_succ_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement_succ_hasCondSubgaussianMGF","description":"Scalar transport of the selected conditional successor-return MGF.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-0470a00792f6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8998,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:146"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalSuccessorAverageReturnDeviationIncrement_succ_hasCondSubgaussianMGF {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (n : Nat) [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [StandardBorelSpace (StochasticEpisodeBatch mdp (episodes (n + 1)))] [Nonempty (StochasticEpisodeBatch mdp (episodes (n + 1)))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.rewardSo…","missing":[],"search":"naturalsuccessoraveragereturndeviationincrement_succ_hascondsubgaussianmgf banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessoraveragereturndeviationincrement_succ_hascondsubgaussianmgf scalar transport of the selected conditional successor-return mgf. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy_pos","label":"naturalCumulativeSuccessorAverageReturnVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy_pos","description":"A positive prefix has positive normalized total return proxy.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-d75928654241","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":8999,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:185"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalCumulativeSuccessorAverageReturnVarianceProxy_pos (mdp : MDP State Action) (episodes : Nat -> Nat) (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrounds : 0 < rounds) (hepisodes : forall n, 0 < episodes n) (hrewardBound_pos : 0 < rewardBound) (hhorizon : 0 < mdp.horizon) : 0 < ((naturalCumulativeSuccessorAverageReturnVarianceProxy mdp episodes rounds rewardBound rewardVarianceProxy : NNReal) : Real)","missing":[],"search":"naturalcumulativesuccessoraveragereturnvarianceproxy_pos banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativesuccessoraveragereturnvarianceproxy_pos a positive prefix has positive normalized total return proxy. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_abs_tail_le","label":"trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_abs_tail_le","description":"Fixed-prefix two-sided tail for normalized successor sample-average returns.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-6d4d5c8f28b7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9000,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_abs_tail_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} [StandardBorelSpace State] [StandardBorelSpace Action] [forall n, StandardBorelSpace (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, Nonempty (HeterogeneousStochasticEpisodeBatchPrefix mdp episodes n)] [forall n, StandardBorelSpace (StochasticEpisodeBatch mdp (episodes n))] [forall n, Nonempty (StochasticEpisodeBatch mdp (episodes n))] [StandardBorelSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)] (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) [source.GlobalReturnMeasurability] (rounds : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (reward…","missing":[],"search":"trajectorymeasure_naturalcumulativesuccessoraveragereturndeviation_abs_tail_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.trajectorymeasure_naturalcumulativesuccessoraveragereturndeviation_abs_tail_le fixed-prefix two-sided tail for normalized successor sample-average returns. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnDeviation_eq_sum_range","label":"naturalCumulativeSuccessorAverageReturnDeviation_eq_sum_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnDeviation_eq_sum_range","description":"The dummy-zero process is exactly the natural successor-round sum.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-53df64d6e7df","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9001,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:295"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalCumulativeSuccessorAverageReturnDeviation_eq_sum_range {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : source.naturalCumulativeSuccessorAverageReturnDeviation rounds trajectory = ∑ t ∈ Finset.range rounds, ((episodes (t + 1) : Real)⁻¹) * source.successorGlobalReturnIncrement (t + 1) trajectory","missing":[],"search":"naturalcumulativesuccessoraveragereturndeviation_eq_sum_range banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativesuccessoraveragereturndeviation_eq_sum_range the dummy-zero process is exactly the natural successor-round sum. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegret","label":"naturalSuccessorBatchAverageRealizedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegret","description":"Realized regret of the sample average in successor batch `t + 1`.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-2cacb28e5198","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9002,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalSuccessorBatchAverageRealizedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (_source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (t : Nat) : Real","missing":[],"search":"naturalsuccessorbatchaveragerealizedregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessorbatchaveragerealizedregret realized regret of the sample average in successor batch `t + 1`. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeRealizedBehaviorRegret","label":"naturalCumulativeRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeRealizedBehaviorRegret","description":"Natural-prefix cumulative successor-batch-average realized regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-3fc47ae2443d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9003,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:324"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalCumulativeRealizedBehaviorRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"naturalcumulativerealizedbehaviorregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativerealizedbehaviorregret natural-prefix cumulative successor-batch-average realized regret. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret","label":"naturalAverageRealizedBehaviorRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret","description":"Natural-prefix round-average realized behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-5bf0e13fa572","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9004,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:335"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def naturalAverageRealizedBehaviorRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) : Real","missing":[],"search":"naturalaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragerealizedbehaviorregret natural-prefix round-average realized behavior regret. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalSuccessorBatchAverageRealizedRegret","label":"measurable_naturalSuccessorBatchAverageRealizedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalSuccessorBatchAverageRealizedRegret","description":"A successor-batch-average realized-regret coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-81827b7a62f6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9005,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:349"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalSuccessorBatchAverageRealizedRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (t : Nat) : Measurable (fun trajectory => source.naturalSuccessorBatchAverageRealizedRegret trajectory t)","missing":[],"search":"measurable_naturalsuccessorbatchaveragerealizedregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalsuccessorbatchaveragerealizedregret a successor-batch-average realized-regret coordinate is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeRealizedBehaviorRegret","label":"measurable_naturalCumulativeRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeRealizedBehaviorRegret","description":"The natural cumulative realized behavior-regret process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-a70907c8d142","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9006,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:365"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalCumulativeRealizedBehaviorRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.naturalCumulativeRealizedBehaviorRegret trajectory rounds)","missing":[],"search":"measurable_naturalcumulativerealizedbehaviorregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalcumulativerealizedbehaviorregret the natural cumulative realized behavior-regret process is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageRealizedBehaviorRegret","label":"measurable_naturalAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageRealizedBehaviorRegret","description":"The natural round-average realized behavior-regret process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-90946c22c35d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9007,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:380"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalAverageRealizedBehaviorRegret {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : Measurable (fun trajectory => source.naturalAverageRealizedBehaviorRegret trajectory rounds)","missing":[],"search":"measurable_naturalaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalaveragerealizedbehaviorregret the natural round-average realized behavior-regret process is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegret_eq_expected_sub_deviation","label":"naturalSuccessorBatchAverageRealizedRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegret_eq_expected_sub_deviation","description":"Exact one-round batch-average realized/expected/deviation identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-e1450f6dc5cc","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9008,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:395"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalSuccessorBatchAverageRealizedRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (t : Nat) (hepisodes : 0 < episodes (t + 1)) : source.naturalSuccessorBatchAverageRealizedRegret trajectory t = (source.successorPolicyAt trajectory t).expectedRegret initialState - source.naturalSuccessorAverageReturnDeviationIncrement (t + 1) trajectory","missing":[],"search":"naturalsuccessorbatchaveragerealizedregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalsuccessorbatchaveragerealizedregret_eq_expected_sub_deviation exact one-round batch-average realized/expected/deviation identity. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeRealizedBehaviorRegret_eq_expected_sub_deviation","label":"naturalCumulativeRealizedBehaviorRegret_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeRealizedBehaviorRegret_eq_expected_sub_deviation","description":"Exact natural-prefix cumulative expected-minus-normalized-deviation identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-5ceab7c47840","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9009,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:421"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalCumulativeRealizedBehaviorRegret_eq_expected_sub_deviation {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (rounds : Nat) (hepisodes : forall n, 0 < episodes n) : source.naturalCumulativeRealizedBehaviorRegret trajectory rounds = (∑ t ∈ Finset.range rounds, (source.successorPolicyAt trajectory t).expectedRegret initialState) - source.naturalCumulativeSuccessorAverageReturnDeviation rounds trajectory","missing":[],"search":"naturalcumulativerealizedbehaviorregret_eq_expected_sub_deviation banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalcumulativerealizedbehaviorregret_eq_expected_sub_deviation exact natural-prefix cumulative expected-minus-normalized-deviation identity. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy","label":"selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy","description":"Self-consistent normalized successor-return proxy for a natural prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-51c49d3ec664","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9010,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:459"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : NNReal","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativereturnvarianceproxy banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativereturnvarianceproxy self-consistent normalized successor-return proxy for a natural prefix. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","label":"selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","description":"Normalized successor-return deviation on the self-consistent causal source.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-cfa4a3282afe","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9011,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:470"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativereturndeviationprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativereturndeviationprocess normalized successor-return deviation on the self-consistent causal source. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","label":"selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","description":"Natural cumulative successor-batch-average realized behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-11a3218d8199","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9012,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:486"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess natural cumulative successor-batch-average realized behavior regret. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","description":"Natural round-average successor-batch-average realized behavior regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-cc0f728ca9c9","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9013,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:503"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess natural round-average successor-batch-average realized behavior regret. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","description":"Fixed-prefix two-sided normalized return-deviation event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-46724858025c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9014,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:519"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativereturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativereturnbadevent fixed-prefix two-sided normalized return-deviation event. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","label":"measurable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","description":"The self-consistent normalized return-deviation process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-577bf596ed52","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9015,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:539"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalcumulativereturndeviationprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalcumulativereturndeviationprocess the self-consistent normalized return-deviation process is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","description":"The natural cumulative batch-average realized-regret process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-d4bd0ebb0cdb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9016,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:560"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess the natural cumulative batch-average realized-regret process is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","description":"The natural round-average batch-average realized-regret process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-3612a3202112","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9017,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:577"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) : Measurable (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess the natural round-average batch-average realized-regret process is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","label":"measurableSet_selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","description":"The fixed-prefix normalized return event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-3ef5dc8d26da","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9018,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:594"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : MeasurableSet (selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds returnDelta)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalcumulativereturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalcumulativereturnbadevent the fixed-prefix normalized return event is measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_eq_expected_sub_deviation","label":"selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_eq_expected_sub_deviation","description":"Exact natural-prefix cumulative realized/expected/deviation identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-b4c9e9ac3ee0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9019,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:612"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_eq_expected_sub_deviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory = selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy b…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess_eq_expected_sub_deviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess_eq_expected_sub_deviation exact natural-prefix cumulative realized/expected/deviation identity. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq_expected_sub_deviation","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq_expected_sub_deviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq_expected_sub_deviation","description":"Exact natural-prefix round-average realized/expected/deviation identity.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-e47d011e7480","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9020,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:650"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq_expected_sub_deviation (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds trajectory = (selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVi…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_eq_expected_sub_deviation banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_eq_expected_sub_deviation exact natural-prefix round-average realized/expected/deviation identity. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalCumulativeReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_naturalCumulativeReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalCumulativeReturnBadEvent_le","description":"The self-consistent normalized return event has its exact return share.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-450cc7e72363","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9021,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:679"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_naturalCumulativeReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (baseVisitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDelta_le_one : returnDelta <= 1) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor source.trajectoryMeasure (selfConsistentScheduledNat…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_naturalcumulativereturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_naturalcumulativereturnbadevent_le the self-consistent normalized return event has its exact return share. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedCumulativeLogarithmicRate","label":"selfConsistentScheduledNaturalCausalRealizedCumulativeLogarithmicRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedCumulativeLogarithmicRate","description":"Logarithmic model rate plus the normalized successor-return radius.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-d0eb7e222949","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9022,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:728"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRealizedCumulativeLogarithmicRate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedcumulativelogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedcumulativelogarithmicrate logarithmic model rate plus the normalized successor-return radius. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate","label":"selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate","description":"Round-average form of the fixed-prefix realized logarithmic rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-9ea7ba0065e8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9023,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:738"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate (mdp : MDP State Action) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalrealizedaveragelogarithmicrate round-average form of the fixed-prefix realized logarithmic rate. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalModelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalModelReturnBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalModelReturnBadEvent","description":"Union of the actual finite-prefix model event and normalized return event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-5f50583716d3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9024,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:745"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalModelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalmodelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalmodelreturnbadevent union of the actual finite-prefix model event and normalized return event. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","label":"selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","description":"One-sided cumulative realized-regret violation set.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-b8c92e5c1a82","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9025,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:763"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretlogarithmicviolationset one-sided cumulative realized-regret violation set. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","description":"One-sided round-average realized-regret violation set.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-578ffb81f2d1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9026,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:782"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset one-sided round-average realized-regret violation set. definition compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","description":"The cumulative realized-regret violation set is Borel measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-f27fcb224cb8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9027,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:801"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : MeasurableSet (selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds returnDelta)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretlogarithmicviolationset the cumulative realized-regret violation set is borel measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","label":"measurableSet_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","description":"The average realized-regret violation set is Borel measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-7ccf08e8c23b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9028,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:819"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (rounds : Nat) (returnDelta : Real) : MeasurableSet (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor rounds returnDelta)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset the average realized-regret violation set is borel measurable. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalModelReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_naturalModelReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalModelReturnBadEvent_le","description":"The model/return union is measurable and obeys the sum of its two shares.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-d07155bf9829","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9029,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:837"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_naturalModelReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hreturnDelta_le_one : returnDelta…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_naturalmodelreturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_naturalmodelreturnbadevent_le the model/return union is measurable and obeys the sum of its two shares. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","description":"Outside the joint event, cumulative realized regret obeys the explicit rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-d46ef1753113","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9030,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:900"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (returnDelta : Real) (trajectory : HeterogeneousStochasticEpisodeBatchTra…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess_le_logarithmic_of_not_mem_modelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretprocess_le_logarithmic_of_not_mem_modelreturnbadevent outside the joint event, cumulative realized regret obeys the explicit rate. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","description":"Outside the joint event, average realized regret obeys the divided rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-9a6cda64c8a1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9031,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:991"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) (trajectory : HeterogeneousStoch…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_logarithmic_of_not_mem_modelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_le_logarithmic_of_not_mem_modelreturnbadevent outside the joint event, average realized regret obeys the divided rate. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","description":"Every cumulative realized-regret violation lies in the joint event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-30227a83a039","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9032,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:1032"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (returnDelta : Real) : selfConsistentScheduledNaturalCausalCumulativeRealize…","missing":[],"search":"selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretlogarithmicviolationset_subset_modelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalcumulativerealizedbehaviorregretlogarithmicviolationset_subset_modelreturnbadevent every cumulative realized-regret violation lies in the joint event. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","description":"Every average realized-regret violation lies in the joint event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-a181b85d6eec","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9033,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:1063"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) : selfConsistentScheduledNaturalCau…","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset_subset_modelreturnbadevent banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset_subset_modelreturnbadevent every average realized-regret violation lies in the joint event. theorem compiled","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","label":"selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","description":"Every average realized-regret violation lies in the joint event. -/ theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvaria…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregrethighprobabilitylograte/index.html#decl-a1ad3e062d22","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","order":9034,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate.lean:1100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (rounds : Nat) (hrounds : 0 < rounds) (returnDelta : Real) (hreturnDelta : 0 < returnDelta) (hret…","missing":[],"search":"selfconsistentscheduledcausalsource_fixedprefixhighprobabilitylogarithmiccumulativeaveragerealizedbehaviorregret banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_fixedprefixhighprobabilitylogarithmiccumulativeaveragerealizedbehaviorregret every average realized-regret violation lies in the joint event. -/ theorem selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset_subset_modelreturnbadevent (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (rounds : nat) (hrounds : 0 < rounds) (returndelta : real) : selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretlogarithmicviolationset mdp initialstate rewardsource initialtable defaultstate varianceproxy…","shard":"modules/4bec63ff105ca405.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","description":"The equal-round-weighted natural realized-regret process on the explicit schedule.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-c8bc3bbc30cb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9035,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretprocess the equal-round-weighted natural realized-regret process on the explicit schedule. definition compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","description":"Fixed-threshold distance-from-zero violation for the scheduled process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-b8b11beea4b6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9036,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:52"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (n : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset fixed-threshold distance-from-zero violation for the scheduled process. definition compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability","description":"Trajectory probability of the scheduled distance violation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-e54f4dbe4c53","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9037,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (n : Nat) : ENNReal","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability trajectory probability of the scheduled distance violation. definition compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","label":"measurable_explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","description":"Every scheduled coordinate of the equal-round-weighted process is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-7cd170435c23","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9038,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) : Measurable (explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n)","missing":[],"search":"measurable_explicitpolynomialprefixaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_explicitpolynomialprefixaveragerealizedbehaviorregretprocess every scheduled coordinate of the equal-round-weighted process is measurable. theorem compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","label":"measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","description":"Every fixed-threshold scheduled distance violation is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-0bdc280285b8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9039,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (n : Nat) : MeasurableSet (explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon n)","missing":[],"search":"measurableset_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset every fixed-threshold scheduled distance violation is measurable. theorem compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess_neg_averageReturnRadius_lt_of_not_mem_event","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess_neg_averageReturnRadius_lt_of_not_mem_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess_neg_averageReturnRadius_lt_of_not_mem_event","description":"Every fixed-threshold scheduled distance violation is measurable. -/ theorem measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloo…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-62eec66d7bec","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9040,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:130"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess_neg_averageReturnRadius_lt_of_not_mem_event (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (n : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) (htrajectory : trajectory ∉ explicitPolynomialPrefixTailModelReturnBadEvent mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n) : -explicitPolynomialPrefixAverageReturnRadius mdp varianceProxy baseVisitFloor n < explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess mdp initialState reward…","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretprocess_neg_averagereturnradius_lt_of_not_mem_event banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretprocess_neg_averagereturnradius_lt_of_not_mem_event every fixed-threshold scheduled distance violation is measurable. -/ theorem measurableset_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor epsilon : real) (n : nat) : measurableset (explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor epsilon n) := by unfold explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset exact measurableset_le measurable_const (measurable.dist (measurable_explicitpolynomialprefixaveragerealizedbehaviorregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor n) measurable_const) /- outside the scheduled union event, the return deviation is strictly smaller than its confidence radius. nonnegative behavior expected regret therefo…","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixTailModelReturnBadEvent_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixTailModelReturnBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixTailModelReturnBadEvent_le","description":"Direct projection of the scheduled parent-event probability bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-01cc20f9bad4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9041,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:207"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixTailModelReturnBadEvent_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_explicitpolynomialprefixtailmodelreturnbadevent_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_explicitpolynomialprefixtailmodelreturnbadevent_le direct projection of the scheduled parent-event probability bound. theorem compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet_subset_event","label":"eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet_subset_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet_subset_event","description":"Eventually every fixed positive distance violation lies in the parent event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-074da192d6c6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9042,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:239"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet_subset_event (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : ∀ᶠ n in atTop, explicitPolynomialPrefixAverageRealizedBehaviorRegret…","missing":[],"search":"eventually_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset_subset_event banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.eventually_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationset_subset_event eventually every fixed positive distance violation lies in the parent event. theorem compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_failureBudget","label":"eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_failureBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_failureBudget","description":"Eventually every fixed distance probability is bounded by the exact budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-14e8a392bf17","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9043,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:306"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_failureBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : ∀ᶠ n in atTop, explicitPolynomialPrefixAverageRealizedBe…","missing":[],"search":"eventually_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_le_failurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.eventually_explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_le_failurebudget eventually every fixed distance probability is bounded by the exact budget. theorem compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_tendsto_zero","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_tendsto_zero","description":"The scheduled distance-violation probability vanishes at every positive threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-b9becf125a9f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9044,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:342"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : Tendsto (explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceV…","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_tendsto_zero the scheduled distance-violation probability vanishes at every positive threshold. theorem compiled","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoInMeasure_zero","label":"selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoInMeasure_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoInMeasure_zero","description":"The scheduled distance-violation probability vanishes at every positive threshold. -/ theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvariance…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretinmeasureexplicitschedule/index.html#decl-156e18ce7a2f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","order":9045,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule.lean:375"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoInMeasure_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSource initialTable def…","missing":[],"search":"selfconsistentscheduledcausalsource_explicitpolynomialprefixaveragerealizedbehaviorregret_tendstoinmeasure_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_explicitpolynomialprefixaveragerealizedbehaviorregret_tendstoinmeasure_zero the scheduled distance-violation probability vanishes at every positive threshold. -/ theorem explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) (epsilon : real) (hepsilon : 0 < epsilon) : tendsto (explicitpolynomialprefixaveragerealizedbehaviorregretdistanceviolationprobability mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor epsi…","shard":"modules/5bf9e4aedf8353f7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","description":"Fixed-threshold upper-tail event for the scheduled average realized regret.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-37c524b8c301","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9046,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (n : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretuppertailset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretuppertailset fixed-threshold upper-tail event for the scheduled average realized regret. definition compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability","description":"Trajectory probability of the fixed-threshold scheduled upper tail.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-36f0418ae7a0","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9047,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (n : Nat) : ENNReal","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretuppertailprobability banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretuppertailprobability trajectory probability of the fixed-threshold scheduled upper tail. definition compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","label":"measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","description":"Every fixed-threshold scheduled upper-tail event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-b463f782d852","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9048,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (n : Nat) : MeasurableSet (explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon n)","missing":[],"search":"measurableset_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailset every fixed-threshold scheduled upper-tail event is measurable. theorem compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet_subset_violationSet","label":"eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet_subset_violationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet_subset_violationSet","description":"Once the deterministic envelope is below a positive fixed threshold, the fixed-threshold upper tail is contained in the compiled envelope violation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-ef3fe5c77e9b","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9049,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet_subset_violationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (hepsilon : 0 < epsilon) : ∀ᶠ n in atTop, explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon n ⊆ explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor n","missing":[],"search":"eventually_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailset_subset_violationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.eventually_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailset_subset_violationset once the deterministic envelope is below a positive fixed threshold, the fixed-threshold upper tail is contained in the compiled envelope violation. theorem compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet_le","label":"selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet_le","description":"Direct projection of the compiled scheduled violation probability bound.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-56b717004c7a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9050,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (n : Nat) : let source := selfConsistentScheduledCausalSource mdp initialState rewardSo…","missing":[],"search":"selfconsistentscheduledcausalsource_trajectorymeasure_explicitpolynomialprefixaveragerealizedbehaviorregretviolationset_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_trajectorymeasure_explicitpolynomialprefixaveragerealizedbehaviorregretviolationset_le direct projection of the compiled scheduled violation probability bound. theorem compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_le_failureBudget","label":"eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_le_failureBudget","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_le_failureBudget","description":"Eventually every fixed positive upper tail is bounded by the exact budget.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-91f9675420cf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9051,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_le_failureBudget (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : ∀ᶠ n in atTop, explicitPolynomialPrefixAverageRealizedBehaviorRe…","missing":[],"search":"eventually_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailprobability_le_failurebudget banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.eventually_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailprobability_le_failurebudget eventually every fixed positive upper tail is bounded by the exact budget. theorem compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_tendsto_zero","label":"explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_tendsto_zero","description":"The fixed-positive-threshold scheduled upper-tail probability vanishes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-f00b97c591c1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9052,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:192"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (epsilon : Real) (hepsilon : 0 < epsilon) : Tendsto (explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbabil…","missing":[],"search":"explicitpolynomialprefixaveragerealizedbehaviorregretuppertailprobability_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixaveragerealizedbehaviorregretuppertailprobability_tendsto_zero the fixed-positive-threshold scheduled upper-tail probability vanishes. theorem compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailInProbability","label":"selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailInProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailInProbability","description":"All fixed positive one-sided upper-tail probabilities vanish on the same explicit fourth-power prefix subsequence.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalrealizedbehaviorregretuppertailinprobability/index.html#decl-1e43e0bb0a0a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","order":9053,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability.lean:225"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailInProbability (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : forall epsilon, 0 < epsilon -> (forall n, MeasurableSet (explicitPolynomialPrefixAverageRealized…","missing":[],"search":"selfconsistentscheduledcausalsource_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailinprobability banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_explicitpolynomialprefixaveragerealizedbehaviorregretuppertailinprobability all fixed positive one-sided upper-tail probabilities vanish on the same explicit fourth-power prefix subsequence. theorem compiled","shard":"modules/f4bd7cceaf19f542.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold","label":"selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold","description":"Reciprocal threshold `1/(n+1)` for the scheduled first-passage scan.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-f63b32a63113","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9054,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:38"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalfirstpassagethreshold banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalfirstpassagethreshold reciprocal threshold `1/(n+1)` for the scheduled first-passage scan. definition compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_pos","label":"selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_pos","description":"Every reciprocal first-passage threshold is strictly positive.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-f966a6bfaa32","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9055,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_pos (scheduleIndex : Nat) : 0 < selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalfirstpassagethreshold_pos banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalfirstpassagethreshold_pos every reciprocal first-passage threshold is strictly positive. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_tendsto_zero","label":"selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_tendsto_zero","description":"The reciprocal first-passage threshold tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-fa2ae3e6b04f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9056,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_tendsto_zero : Tendsto selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalfirstpassagethreshold_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalfirstpassagethreshold_tendsto_zero the reciprocal first-passage threshold tends to zero. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","description":"The capped first-passage stopping rule at the reciprocal threshold.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-b4674b02473f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9057,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagestoppingprefix the capped first-passage stopping rule at the reciprocal threshold. definition compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","description":"Event that reciprocal-threshold first passage does not stop at its base.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-c6ee60580479","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9058,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset event that reciprocal-threshold first passage does not stop at its base. definition compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","description":"Probability of failing to stop at the first prefix in the scan window.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-8e2c38944ee6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9059,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:104"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : ENNReal","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability probability of failing to stop at the first prefix in the scan window. definition compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","description":"Deterministic Markov rate for reciprocal-threshold first-passage delay.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-862f6f0d3fad","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9060,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:120"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : Real","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayrate deterministic markov rate for reciprocal-threshold first-passage delay. definition compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","description":"Delay is exactly strict one-sided threshold violation at the base prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-f492d91be131","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9061,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:130"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex = {trajectory | selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold scheduleIndex < explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory}","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_eq delay is exactly strict one-sided threshold violation at the base prefix. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","label":"measurableSet_selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","description":"The reciprocal-threshold delay event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-19c6f5151773","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9062,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:222"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : MeasurableSet (selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset the reciprocal-threshold delay event is measurable. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","description":"Delayed first passage is contained in the scheduled distance violation.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-f536f25809a3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9063,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:242"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex ⊆ explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold scheduleIndex) scheduleIndex","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_subset_distanceviolationset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedset_subset_distanceviolationset delayed first passage is contained in the scheduled distance violation. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","description":"Dividing the scheduled L1 envelope by the reciprocal threshold exposes an inverse-square plus inverse-linear delay rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-fab2f99142a2","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9064,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq (mdp : MDP State Action) (varianceProxy : NNReal) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate mdp varianceProxy scheduleIndex = 4 * selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient mdp / (explicitHighProbabilityScale scheduleIndex : Real) ^ 2 + (2 * Real.sqrt (mdp.globalReturnDeviationPerEpisodeVarianceProxy 1 varianceProxy : Real) * Real.exp (1 / 2 : Real)) / (explicitHighProbabilityScale scheduleIndex : Real)","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_eq banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_eq dividing the scheduled l1 envelope by the reciprocal threshold exposes an inverse-square plus inverse-linear delay rate. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_reciprocalFirstPassageThreshold_le_delayRate","label":"explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_reciprocalFirstPassageThreshold_le_delayRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_reciprocalFirstPassageThreshold_le_delayRate","description":"The scheduled expected absolute base process divided by the reciprocal threshold is bounded by the explicit delay rate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-3481377b60eb","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9065,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:307"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_reciprocalFirstPassageThreshold_le_delayRate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegr…","missing":[],"search":"explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_div_reciprocalfirstpassagethreshold_le_delayrate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.explicitpolynomialprefixexpectedabsoluteaveragerealizedbehaviorregret_div_reciprocalfirstpassagethreshold_le_delayrate the scheduled expected absolute base process divided by the reciprocal threshold is bounded by the explicit delay rate. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","description":"The reciprocal-threshold delay rate tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-1616ef40f838","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9066,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:344"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero (mdp : MDP State Action) (varianceProxy : NNReal) : Tendsto (selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate mdp varianceProxy) atTop (nhds 0)","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayrate_tendsto_zero the reciprocal-threshold delay rate tends to zero. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","description":"Markov control of the reciprocal-threshold first-passage delay event.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-cc22328547da","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9067,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:394"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoub…","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_le_rate banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_le_rate markov control of the reciprocal-threshold first-passage delay event. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","label":"selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","description":"The probability of delaying beyond the first prefix tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-0aa6e2d0146e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9068,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:466"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : Tendsto (selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinear…","missing":[],"search":"selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_tendsto_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_tendsto_zero the probability of delaying beyond the first prefix tends to zero. theorem compiled","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_reciprocalThresholdCappedDoubleLinearRawWindowFirstPassage_vanishingDelayProbability_and_L1_consistency","label":"selfConsistentScheduledCausalSource_reciprocalThresholdCappedDoubleLinearRawWindowFirstPassage_vanishingDelayProbability_and_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_reciprocalThresholdCappedDoubleLinearRawWindowFirstPassage_vanishingDelayProbability_and_L1_consistency","description":"The probability of delaying beyond the first prefix tends to zero. -/ theorem selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNR…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalreciprocalthresholdcappeddoublelinearraw-aba06a37d450/index.html#decl-48d2d9ac644a","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","order":9069,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency.lean:500"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_reciprocalThresholdCappedDoubleLinearRawWindowFirstPassage_vanishingDelayProbability_and_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let threshold := selfConsistentScheduledNaturalCausalReciprocalFirst…","missing":[],"search":"selfconsistentscheduledcausalsource_reciprocalthresholdcappeddoublelinearrawwindowfirstpassage_vanishingdelayprobability_and_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_reciprocalthresholdcappeddoublelinearrawwindowfirstpassage_vanishingdelayprobability_and_l1_consistency the probability of delaying beyond the first prefix tends to zero. -/ theorem selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability_tendsto_zero (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] [standardborelspace state] [standardborelspace action] (rewardsource : mdp.meancompatiblerewardkernel) (varianceproxy : nnreal) (hvarianceproxy : 0 < varianceproxy) (law : rewardsource.uniformsubgaussianrewardlaw varianceproxy) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (support : exploratorypathsupport mdp initialstate) (basevisitfloor : real) (hbasefloor : exploratorypathuniformvisitfloor support 1 basevisitfloor) (hrewardbound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbasevisitfloor : 0 < basevisitfloor) : tendsto (selfconsistentschedulednaturalcausalreciprocalthresholdcappeddoublelinearrawwindowfirstpassagedelayedprobability mdp initialstate rewardsource…","shard":"modules/bceb71c80a576ce6.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalTrajectoryFiltration","label":"naturalTrajectoryFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalTrajectoryFiltration","description":"The dependent coordinate filtration through the current batch coordinate.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-d0a8a9c2d931","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9070,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (_source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : Filtration Nat (inferInstance : MeasurableSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes))","missing":[],"search":"naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturaltrajectoryfiltration the dependent coordinate filtration through the current batch coordinate. definition compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalTrajectoryFiltration_apply","label":"naturalTrajectoryFiltration_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalTrajectoryFiltration_apply","description":"theorem naturalTrajectoryFiltration_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : (source.naturalTrajectoryFiltration n : MeasurableSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)) = Filtration.piLE (X := fun n : Nat => Stoch…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-ed14d16c0a0c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9071,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalTrajectoryFiltration_apply {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (n : Nat) : (source.naturalTrajectoryFiltration n : MeasurableSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes)) = Filtration.piLE (X := fun n : Nat => StochasticEpisodeBatch mdp (episodes n)) n","missing":[],"search":"naturaltrajectoryfiltration_apply banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturaltrajectoryfiltration_apply theorem naturaltrajectoryfiltration_apply {mdp : mdp state action} {initialstate : measure state} [isprobabilitymeasure initialstate] {episodes : nat -> nat} (source : heterogeneousadaptivestochasticepisodebatchsource mdp initialstate episodes) (n : nat) : (source.naturaltrajectoryfiltration n : measurablespace (heterogeneousstochasticepisodebatchtrajectory mdp episodes)) = filtration.pile (x := fun n : nat => stochasticepisodebatch mdp (episodes n)) n theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalTrajectory_coordinate_of_le","label":"measurable_naturalTrajectory_coordinate_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalTrajectory_coordinate_of_le","description":"A trajectory coordinate is measurable at every later natural-filtration level.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-0ebce0910adf","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9072,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:70"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalTrajectory_coordinate_of_le {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) {coordinate horizon : Nat} (hcoordinate : coordinate <= horizon) : @Measurable (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) (StochasticEpisodeBatch mdp (episodes coordinate)) (source.naturalTrajectoryFiltration horizon) inferInstance (fun trajectory => trajectory coordinate)","missing":[],"search":"measurable_naturaltrajectory_coordinate_of_le banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturaltrajectory_coordinate_of_le a trajectory coordinate is measurable at every later natural-filtration level. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalSuccessorBatchAverageRealizedRegret_naturalTrajectoryFiltration","label":"measurable_naturalSuccessorBatchAverageRealizedRegret_naturalTrajectoryFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalSuccessorBatchAverageRealizedRegret_naturalTrajectoryFiltration","description":"One successor-batch realized-regret coordinate is measurable at its batch time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-a1b3132f0835","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9073,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalSuccessorBatchAverageRealizedRegret_naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (t : Nat) : @Measurable (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) Real (source.naturalTrajectoryFiltration (t + 1)) inferInstance (fun trajectory => source.naturalSuccessorBatchAverageRealizedRegret trajectory t)","missing":[],"search":"measurable_naturalsuccessorbatchaveragerealizedregret_naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalsuccessorbatchaveragerealizedregret_naturaltrajectoryfiltration one successor-batch realized-regret coordinate is measurable at its batch time. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeRealizedBehaviorRegret_naturalTrajectoryFiltration","label":"measurable_naturalCumulativeRealizedBehaviorRegret_naturalTrajectoryFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeRealizedBehaviorRegret_naturalTrajectoryFiltration","description":"The cumulative realized-regret prefix is measurable at its prefix level.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-81baddbf884c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9074,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalCumulativeRealizedBehaviorRegret_naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : @Measurable (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) Real (source.naturalTrajectoryFiltration rounds) inferInstance (fun trajectory => source.naturalCumulativeRealizedBehaviorRegret trajectory rounds)","missing":[],"search":"measurable_naturalcumulativerealizedbehaviorregret_naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalcumulativerealizedbehaviorregret_naturaltrajectoryfiltration the cumulative realized-regret prefix is measurable at its prefix level. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageRealizedBehaviorRegret_naturalTrajectoryFiltration","label":"measurable_naturalAverageRealizedBehaviorRegret_naturalTrajectoryFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageRealizedBehaviorRegret_naturalTrajectoryFiltration","description":"The exact average realized-regret prefix is measurable at its prefix level.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-83e8bb9d9030","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9075,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_naturalAverageRealizedBehaviorRegret_naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) (rounds : Nat) : @Measurable (HeterogeneousStochasticEpisodeBatchTrajectory mdp episodes) Real (source.naturalTrajectoryFiltration rounds) inferInstance (fun trajectory => source.naturalAverageRealizedBehaviorRegret trajectory rounds)","missing":[],"search":"measurable_naturalaveragerealizedbehaviorregret_naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.measurable_naturalaveragerealizedbehaviorregret_naturaltrajectoryfiltration the exact average realized-regret prefix is measurable at its prefix level. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret_stronglyAdapted_naturalTrajectoryFiltration","label":"naturalAverageRealizedBehaviorRegret_stronglyAdapted_naturalTrajectoryFiltration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret_stronglyAdapted_naturalTrajectoryFiltration","description":"The exact natural average realized-regret process is strongly adapted.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-57d13240c5b3","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9076,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:166"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem naturalAverageRealizedBehaviorRegret_stronglyAdapted_naturalTrajectoryFiltration {mdp : MDP State Action} {initialState : Measure State} [IsProbabilityMeasure initialState] {episodes : Nat -> Nat} (source : HeterogeneousAdaptiveStochasticEpisodeBatchSource mdp initialState episodes) : StronglyAdapted source.naturalTrajectoryFiltration (fun rounds trajectory => source.naturalAverageRealizedBehaviorRegret trajectory rounds)","missing":[],"search":"naturalaveragerealizedbehaviorregret_stronglyadapted_naturaltrajectoryfiltration banditrlproof.finitehorizonrl.heterogeneousadaptivestochasticepisodebatchsource.naturalaveragerealizedbehaviorregret_stronglyadapted_naturaltrajectoryfiltration the exact natural average realized-regret process is strongly adapted. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalTrajectoryFiltration","label":"selfConsistentScheduledNaturalCausalTrajectoryFiltration","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalTrajectoryFiltration","description":"Natural filtration of the self-consistent heterogeneous causal source.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-c427ab102f87","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9077,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:184"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalTrajectoryFiltration (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) : Filtration Nat (inferInstance : MeasurableSpace (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)))","missing":[],"search":"selfconsistentschedulednaturalcausaltrajectoryfiltration banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausaltrajectoryfiltration natural filtration of the self-consistent heterogeneous causal source. definition compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_stronglyAdapted","label":"selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_stronglyAdapted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_stronglyAdapted","description":"The exact self-consistent natural average process is strongly adapted.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-75f03e39c8bd","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9078,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:201"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_stronglyAdapted (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) : StronglyAdapted (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor)","missing":[],"search":"selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_stronglyadapted banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess_stronglyadapted the exact self-consistent natural average process is strongly adapted. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","description":"The exact average realized-regret process evaluated at a stopping time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-a76be95f9c6c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9079,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:222"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> Real","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess the exact average realized-regret process evaluated at a stopping time. definition compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_apply","label":"selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_apply","description":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeB…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-16445e501c6c","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9080,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:246"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_apply (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTabl…","missing":[],"search":"selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess_apply banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess_apply theorem selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess_apply (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (stoppingprefix : nat -> heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t) -> withtop nat) (scheduleindex : nat) (trajectory : heterogeneousstochasticepisodebatchtrajectory mdp (fun t => adaptivestochasticepisodebatchsource.selfconsistentscheduledepisodes mdp varianceproxy basevisitfloor t)) : selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor stoppingprefix scheduleindex trajectory = selfconsistentschedulednaturalcausalaveragerealizedbehaviorregretprocess mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisi…","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","label":"measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","description":"Mathlib stopped-value measurability for the exact adapted process.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-03eab34e6c4d","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9081,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:272"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat) (scheduleIndex : Nat) (hstopping : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (stoppingPrefix scheduleIndex)) : Measurable (selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess…","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalstoppingtimeaveragerealizedbehaviorregretprocess mathlib stopped-value measurability for the exact adapted process. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","label":"selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","description":"Diverging stopping times preserve almost-sure natural average consistency.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-7fb3c482c4e1","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9082,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:313"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveSto…","missing":[],"search":"selfconsistentscheduledcausalsource_stoppingtimenaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_stoppingtimenaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero diverging stopping times preserve almost-sure natural average consistency. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","label":"selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","description":"The practical lower envelope `n <= tau_n` forces the stopped prefixes to diverge.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalstoppingtimeaveragerealizedbehaviorregre-fb3c2e78d95f/index.html#decl-b8fc1f48fa90","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","order":9083,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency.lean:385"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (stoppingPrefix : Nat -> HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => A…","missing":[],"search":"selfconsistentscheduledcausalsource_stoppingtimenaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero_of_nat_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_stoppingtimenaturalaveragerealizedbehaviorregret_tendstoalmosteverywhere_zero_of_nat_le the practical lower envelope `n <= tau_n` forces the stopped prefixes to diverge. theorem compiled","shard":"modules/c0efccc76ef87329.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","description":"Event on which the observed base-prefix average regret triggers early stopping.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-dc9915563b5f","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9084,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : Set (HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t))","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowearlystopset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowearlystopset event on which the observed base-prefix average regret triggers early stopping. definition compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix","description":"Stop at the observed fourth-power prefix, or wait exactly `2*n+1` more prefixes.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-1f226ed763f6","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9085,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) -> WithTop Nat","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix stop at the observed fourth-power prefix, or wait exactly `2*n+1` more prefixes. definition compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_base_of_le","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_base_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_base_of_le","description":"The threshold-success branch stops at the observed base prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-8348cb7c8ce8","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9086,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:80"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_base_of_le (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) (hthreshold : selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (explicitHighProbabilityRounds scheduleIndex) trajectory <= threshold scheduleIndex) : selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex traj…","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_eq_base_of_le banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_eq_base_of_le the threshold-success branch stops at the observed base prefix. theorem compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_right_of_lt","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_right_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_right_of_lt","description":"The threshold-failure branch waits to the right endpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-d4a831b95e19","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9087,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_right_of_lt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) (hthreshold : threshold scheduleIndex < selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (explicitHighProbabilityRounds scheduleIndex) trajectory) : selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex traj…","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_eq_right_of_lt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_eq_right_of_lt the threshold-failure branch waits to the right endpoint. theorem compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","label":"measurableSet_selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","description":"The early-stop event is known at the fourth-power base prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-da382f44132e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9088,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : MeasurableSet[ selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor (explicitHighProbabilityRounds scheduleIndex)] (selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex)","missing":[],"search":"measurableset_selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowearlystopset banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurableset_selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowearlystopset the early-stop event is known at the fourth-power base prefix. theorem compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_isStoppingTime","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_isStoppingTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_isStoppingTime","description":"The threshold-triggered two-endpoint rule is an exact natural-filtration stopping time.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-9de83938eaea","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9089,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:173"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_isStoppingTime (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) : IsStoppingTime (selfConsistentScheduledNaturalCausalTrajectoryFiltration mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor) (selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex)","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_isstoppingtime banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_isstoppingtime the threshold-triggered two-endpoint rule is an exact natural-filtration stopping time. theorem compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_lower","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_lower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_lower","description":"The threshold rule never stops before its observed base prefix.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-84285769791e","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9090,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:210"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_lower (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) : (explicitHighProbabilityRounds scheduleIndex : WithTop Nat) <= selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_lower banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_lower the threshold rule never stops before its observed base prefix. theorem compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_upper","label":"selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_upper","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_upper","description":"The threshold rule never exceeds the double-linear right endpoint.","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-c8460b2be3e4","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9091,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:233"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_upper (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (threshold : Nat -> Real) (scheduleIndex : Nat) (trajectory) : selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor threshold scheduleIndex trajectory <= (explicitHighProbabilityRounds scheduleIndex + (2 * scheduleIndex + 1) : WithTop Nat)","missing":[],"search":"selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_upper banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_upper the threshold rule never exceeds the double-linear right endpoint. theorem compiled","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_thresholdTriggeredDoubleLinearRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","label":"selfConsistentScheduledCausalSource_thresholdTriggeredDoubleLinearRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_thresholdTriggeredDoubleLinearRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","description":"The threshold rule never exceeds the double-linear right endpoint. -/ theorem selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_upper (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (base…","url":"../modules/banditrlproof-rl-finitehorizonnaturalcausalthresholdtriggereddoublelinearrawwindows-396ec6121572/index.html#decl-17ae0aa1cae7","parent":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","order":9092,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency"],["Source","BanditRLProof/RL/FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_thresholdTriggeredDoubleLinearRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 0 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) (threshold : Nat -> Real) : let stoppingPrefix := selfConsistentSchedul…","missing":[],"search":"selfconsistentscheduledcausalsource_thresholdtriggereddoublelinearrawwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_thresholdtriggereddoublelinearrawwindowstoppingtimenaturalaveragerealizedbehaviorregret_l1_consistency the threshold rule never exceeds the double-linear right endpoint. -/ theorem selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix_upper (mdp : mdp state action) (initialstate : measure state) [isprobabilitymeasure initialstate] (rewardsource : mdp.meancompatiblerewardkernel) (initialtable : deterministicmarkovpolicytable mdp) (defaultstate : state) (varianceproxy : nnreal) (basevisitfloor : real) (threshold : nat -> real) (scheduleindex : nat) (trajectory) : selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix mdp initialstate rewardsource initialtable defaultstate varianceproxy basevisitfloor threshold scheduleindex trajectory <= (explicithighprobabilityrounds scheduleindex + (2 * scheduleindex + 1) : withtop nat) := by classical unfold selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowstoppingprefix by_cases htrajectory : trajectory ∈ selfconsistentschedulednaturalcausalthresholdtriggereddoublelinearrawwindowearlystopset mdp…","shard":"modules/51d238693033e960.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stateOccupancy","label":"stateOccupancy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stateOccupancy","description":"State law at a chronological stage under a Markov policy.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-8a2ca737d6e6","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9093,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stateOccupancy {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) : (stage : Nat) -> stage <= mdp.horizon -> Measure State | 0, _ => initialState | stage + 1, hstage => policy.inducedStateKernel ⟨stage, by omega⟩ ∘ₘ policy.stateOccupancy initialState stage (by omega) /-- Every chronological state occupancy is a probability measure. -/ instance instStateOccupancyIsProbabilityMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Nat) (hstage : stage <= mdp.horizon) : IsProbabilityMeasure (policy.stateOccupancy initialState stage hstage)","missing":[],"search":"stateoccupancy banditrlproof.finitehorizonrl.markovpolicy.stateoccupancy state law at a chronological stage under a markov policy. definition compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap","label":"policyBellmanGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap","description":"Expected one-step optimality gap after averaging the optimal action-value gap over the policy action kernel at a chronological stage.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-c6801664822a","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9094,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def policyBellmanGap {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) : Real","missing":[],"search":"policybellmangap banditrlproof.finitehorizonrl.markovpolicy.policybellmangap expected one-step optimality gap after averaging the optimal action-value gap over the policy action kernel at a chronological stage. definition compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_policyBellmanGap","label":"measurable_policyBellmanGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_policyBellmanGap","description":"The policy-averaged one-step Bellman optimality gap is measurable.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-54d3be6ef1dc","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9095,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_policyBellmanGap {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : Measurable (policy.policyBellmanGap stage)","missing":[],"search":"measurable_policybellmangap banditrlproof.finitehorizonrl.markovpolicy.measurable_policybellmangap the policy-averaged one-step bellman optimality gap is measurable. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap_nonneg","label":"policyBellmanGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap_nonneg","description":"Every policy-averaged one-step Bellman optimality gap is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-f24635f3a1af","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9096,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem policyBellmanGap_nonneg {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) : 0 <= policy.policyBellmanGap stage state","missing":[],"search":"policybellmangap_nonneg banditrlproof.finitehorizonrl.markovpolicy.policybellmangap_nonneg every policy-averaged one-step bellman optimality gap is nonnegative. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap_stageOfRemaining","label":"policyBellmanGap_stageOfRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap_stageOfRemaining","description":"Reindex a chronological policy Bellman gap by the number of decisions remaining.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-fd89c6eda94b","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9097,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem policyBellmanGap_stageOfRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : policy.policyBellmanGap ⟨mdp.horizon - (remaining + 1), by omega⟩ state = mdp.optimalValueRemaining (remaining + 1) hremaining state - policy.bellman ⟨mdp.horizon - (remaining + 1), by omega⟩ (mdp.optimalValueRemaining remaining (by omega)) state","missing":[],"search":"policybellmangap_stageofremaining banditrlproof.finitehorizonrl.markovpolicy.policybellmangap_stageofremaining reindex a chronological policy bellman gap by the number of decisions remaining. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_inducedStateKernel_eq_integral_transitionValue","label":"integral_inducedStateKernel_eq_integral_transitionValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_inducedStateKernel_eq_integral_transitionValue","description":"Integrating a continuation value against the induced state kernel averages its transition expectation over the policy action kernel.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-bc4ffd6ae35e","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9098,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:113"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_inducedStateKernel_eq_integral_transitionValue {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : (∫ nextState, value nextState ∂policy.inducedStateKernel stage state) = ∫ action, mdp.transitionValue value state action ∂policy.actionKernel stage state","missing":[],"search":"integral_inducedstatekernel_eq_integral_transitionvalue banditrlproof.finitehorizonrl.markovpolicy.integral_inducedstatekernel_eq_integral_transitionvalue integrating a continuation value against the induced state kernel averages its transition expectation over the policy action kernel. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_sub_bellman_eq_integral_inducedStateKernel_sub","label":"bellman_sub_bellman_eq_integral_inducedStateKernel_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_sub_bellman_eq_integral_inducedStateKernel_sub","description":"Changing only the continuation value in a policy Bellman step is integration against the induced next-state kernel.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-8674efeb6897","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9099,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:134"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellman_sub_bellman_eq_integral_inducedStateKernel_sub {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (left right : State -> Real) (state : State) : policy.bellman stage left state - policy.bellman stage right state = ∫ nextState, left nextState - right nextState ∂policy.inducedStateKernel stage state","missing":[],"search":"bellman_sub_bellman_eq_integral_inducedstatekernel_sub banditrlproof.finitehorizonrl.markovpolicy.bellman_sub_bellman_eq_integral_inducedstatekernel_sub changing only the continuation value in a policy bellman step is integration against the induced next-state kernel. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_sub_comp_inducedStateKernel_eq_integral_bellman_sub","label":"integral_sub_comp_inducedStateKernel_eq_integral_bellman_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_sub_comp_inducedStateKernel_eq_integral_bellman_sub","description":"Push an integrated continuation-value difference through one policy-induced state step.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-b5aeb9156227","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9100,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:186"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sub_comp_inducedStateKernel_eq_integral_bellman_sub {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (left right : State -> Real) (mu : Measure State) [IsProbabilityMeasure mu] : (∫ nextState, left nextState - right nextState ∂policy.inducedStateKernel stage ∘ₘ mu) = ∫ state, policy.bellman stage left state - policy.bellman stage right state ∂mu","missing":[],"search":"integral_sub_comp_inducedstatekernel_eq_integral_bellman_sub banditrlproof.finitehorizonrl.markovpolicy.integral_sub_comp_inducedstatekernel_eq_integral_bellman_sub push an integrated continuation-value difference through one policy-induced state step. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining","label":"occupancyGapRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining","description":"Finite occupancy sum of policy Bellman gaps, indexed by decisions remaining. Each recursive call advances the state law by the policy-induced kernel.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-7e7ced4856bd","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9101,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:214"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def occupancyGapRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) : (remaining : Nat) -> remaining <= mdp.horizon -> Measure State -> Real | 0, _, _ => 0 | remaining + 1, hremaining, mu => let stage : Fin mdp.horizon := ⟨mdp.horizon - (remaining + 1), by omega⟩ (∫ state, mdp.optimalValueRemaining (remaining + 1) hremaining state - policy.bellman stage (mdp.optimalValueRemaining remaining (by omega)) state ∂mu) + policy.occupancyGapRemaining remaining (by omega) (policy.inducedStateKernel stage ∘ₘ mu) omit [MeasurableSingletonClass State] in /-- The successor recursion is an integral of the chronological policy Bellman gap plus the remaining occupancy gaps under the next-state law. -/ theorem occupancyGapRemaining_succ {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (mu : Measure State) : policy…","missing":[],"search":"occupancygapremaining banditrlproof.finitehorizonrl.markovpolicy.occupancygapremaining finite occupancy sum of policy bellman gaps, indexed by decisions remaining. each recursive call advances the state law by the policy-induced kernel. definition compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_succ","label":"occupancyGapRemaining_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_succ","description":"The successor recursion is an integral of the chronological policy Bellman gap plus the remaining occupancy gaps under the next-state law.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-103dd9e9d029","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9102,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:231"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem occupancyGapRemaining_succ {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (mu : Measure State) : policy.occupancyGapRemaining (remaining + 1) hremaining mu = (∫ state, policy.policyBellmanGap ⟨mdp.horizon - (remaining + 1), by omega⟩ state ∂mu) + policy.occupancyGapRemaining remaining (by omega) (policy.inducedStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ ∘ₘ mu)","missing":[],"search":"occupancygapremaining_succ banditrlproof.finitehorizonrl.markovpolicy.occupancygapremaining_succ the successor recursion is an integral of the chronological policy bellman gap plus the remaining occupancy gaps under the next-state law. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_eq_integral_optimalValueRemaining_sub_valueRemaining","label":"occupancyGapRemaining_eq_integral_optimalValueRemaining_sub_valueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_eq_integral_optimalValueRemaining_sub_valueRemaining","description":"The recursive occupancy-gap sum is exactly the integrated optimal-policy value gap.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-24e6a16a2887","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9103,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:248"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem occupancyGapRemaining_eq_integral_optimalValueRemaining_sub_valueRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : policy.occupancyGapRemaining remaining hremaining mu = ∫ state, mdp.optimalValueRemaining remaining hremaining state - policy.valueRemaining remaining hremaining state ∂mu","missing":[],"search":"occupancygapremaining_eq_integral_optimalvalueremaining_sub_valueremaining banditrlproof.finitehorizonrl.markovpolicy.occupancygapremaining_eq_integral_optimalvalueremaining_sub_valueremaining the recursive occupancy-gap sum is exactly the integrated optimal-policy value gap. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_le_optimalValueRemaining","label":"valueRemaining_le_optimalValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_le_optimalValueRemaining","description":"Backward policy value is bounded by backward optimal value.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-a0ef76509563","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9104,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:291"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueRemaining_le_optimalValueRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : policy.valueRemaining remaining hremaining state <= mdp.optimalValueRemaining remaining hremaining state","missing":[],"search":"valueremaining_le_optimalvalueremaining banditrlproof.finitehorizonrl.markovpolicy.valueremaining_le_optimalvalueremaining backward policy value is bounded by backward optimal value. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_nonneg","label":"occupancyGapRemaining_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_nonneg","description":"The finite occupancy Bellman-gap sum is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-c5971c24f4bb","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9105,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:311"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem occupancyGapRemaining_nonneg {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : 0 <= policy.occupancyGapRemaining remaining hremaining mu","missing":[],"search":"occupancygapremaining_nonneg banditrlproof.finitehorizonrl.markovpolicy.occupancygapremaining_nonneg the finite occupancy bellman-gap sum is nonnegative. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret","label":"expectedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret","description":"Expected finite-horizon regret: optimal initial value minus generated trajectory reward.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-c87c298d4826","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9106,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:323"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedRegret {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) : Real","missing":[],"search":"expectedregret banditrlproof.finitehorizonrl.markovpolicy.expectedregret expected finite-horizon regret: optimal initial value minus generated trajectory reward. definition compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_integral_optimalValueAt_sub_valueAt","label":"expectedRegret_eq_integral_optimalValueAt_sub_valueAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_integral_optimalValueAt_sub_valueAt","description":"Expected trajectory regret is the initial-law integral of the optimal-policy value gap.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-21b7f1d47b39","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9107,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:331"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_eq_integral_optimalValueAt_sub_valueAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : policy.expectedRegret initialState = ∫ state, mdp.optimalValueAt 0 (Nat.zero_le mdp.horizon) state - policy.valueAt 0 (Nat.zero_le mdp.horizon) state ∂initialState","missing":[],"search":"expectedregret_eq_integral_optimalvalueat_sub_valueat banditrlproof.finitehorizonrl.markovpolicy.expectedregret_eq_integral_optimalvalueat_sub_valueat expected trajectory regret is the initial-law integral of the optimal-policy value gap. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_occupancyGapRemaining","label":"expectedRegret_eq_occupancyGapRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_occupancyGapRemaining","description":"Expected trajectory regret equals the full recursively accumulated occupancy Bellman gap.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-187ce64f0edb","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9108,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:351"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_eq_occupancyGapRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : policy.expectedRegret initialState = policy.occupancyGapRemaining mdp.horizon le_rfl initialState","missing":[],"search":"expectedregret_eq_occupancygapremaining banditrlproof.finitehorizonrl.markovpolicy.expectedregret_eq_occupancygapremaining expected trajectory regret equals the full recursively accumulated occupancy bellman gap. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_nonneg","label":"expectedRegret_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_nonneg","description":"Expected finite-horizon regret is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-f46992a9519e","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9109,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:361"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_nonneg {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : 0 <= policy.expectedRegret initialState","missing":[],"search":"expectedregret_nonneg banditrlproof.finitehorizonrl.markovpolicy.expectedregret_nonneg expected finite-horizon regret is nonnegative. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_expectedRegret_eq_zero","label":"optimalPolicy_expectedRegret_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_expectedRegret_eq_zero","description":"The measurable greedy policy has zero expected finite-horizon regret.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-4cc718fc0981","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9110,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:373"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalPolicy_expectedRegret_eq_zero (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] : mdp.optimalPolicy.expectedRegret initialState = 0","missing":[],"search":"optimalpolicy_expectedregret_eq_zero banditrlproof.finitehorizonrl.mdp.optimalpolicy_expectedregret_eq_zero the measurable greedy policy has zero expected finite-horizon regret. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.expectedRegret_eq_occupancyGap_nonneg_and_optimalPolicy_zero","label":"expectedRegret_eq_occupancyGap_nonneg_and_optimalPolicy_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.expectedRegret_eq_occupancyGap_nonneg_and_optimalPolicy_zero","description":"Route endpoint: expected trajectory regret is exactly the finite occupancy Bellman-gap sum, is nonnegative for every Markov policy, and vanishes for the compiled measurable greedy optimal policy.","url":"../modules/banditrlproof-rl-finitehorizonoccupancyregret/index.html#decl-3d67c5f57576","parent":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","order":9111,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOccupancyRegret"],["Source","BanditRLProof/RL/FiniteHorizonOccupancyRegret.lean:386"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_eq_occupancyGap_nonneg_and_optimalPolicy_zero (mdp : MDP State Action) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : policy.expectedRegret initialState = policy.occupancyGapRemaining mdp.horizon le_rfl initialState /\\ 0 <= policy.expectedRegret initialState /\\ mdp.optimalPolicy.expectedRegret initialState = 0","missing":[],"search":"expectedregret_eq_occupancygap_nonneg_and_optimalpolicy_zero banditrlproof.finitehorizonrl.mdp.expectedregret_eq_occupancygap_nonneg_and_optimalpolicy_zero route endpoint: expected trajectory regret is exactly the finite occupancy bellman-gap sum, is nonnegative for every markov policy, and vanishes for the compiled measurable greedy optimal policy. theorem compiled","shard":"modules/682a8686065935b9.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalAction","label":"optimalAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalAction","description":"A finite action maximizing the one-step Bellman action value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-53a2b966352f","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9112,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalAction (mdp : MDP State Action) (value : State -> Real) (state : State) : Action","missing":[],"search":"optimalaction banditrlproof.finitehorizonrl.mdp.optimalaction a finite action maximizing the one-step bellman action value. definition compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_le_optimalAction","label":"bellmanQ_le_optimalAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_le_optimalAction","description":"The selected finite action dominates every Bellman action value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-bf69aaaf7443","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9113,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanQ_le_optimalAction (mdp : MDP State Action) (value : State -> Real) (state : State) (action : Action) : mdp.bellmanQ value state action <= mdp.bellmanQ value state (mdp.optimalAction value state)","missing":[],"search":"bellmanq_le_optimalaction banditrlproof.finitehorizonrl.mdp.bellmanq_le_optimalaction the selected finite action dominates every bellman action value. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalAction","label":"measurable_optimalAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalAction","description":"The finite-state maximizing selector is measurable.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-ce2481f57d3f","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9114,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimalAction (mdp : MDP State Action) (value : State -> Real) : Measurable (mdp.optimalAction value)","missing":[],"search":"measurable_optimalaction banditrlproof.finitehorizonrl.mdp.measurable_optimalaction the finite-state maximizing selector is measurable. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalBellman","label":"optimalBellman","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalBellman","description":"Pointwise finite-action Bellman maximum.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-adc163196623","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9115,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalBellman (mdp : MDP State Action) (value : State -> Real) (state : State) : Real","missing":[],"search":"optimalbellman banditrlproof.finitehorizonrl.mdp.optimalbellman pointwise finite-action bellman maximum. definition compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_le_optimalBellman","label":"bellmanQ_le_optimalBellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_le_optimalBellman","description":"Every action value is bounded by the finite-action Bellman maximum.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-30955c900622","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9116,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanQ_le_optimalBellman (mdp : MDP State Action) (value : State -> Real) (state : State) (action : Action) : mdp.bellmanQ value state action <= mdp.optimalBellman value state","missing":[],"search":"bellmanq_le_optimalbellman banditrlproof.finitehorizonrl.mdp.bellmanq_le_optimalbellman every action value is bounded by the finite-action bellman maximum. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalBellman","label":"measurable_optimalBellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalBellman","description":"The optimal Bellman value is measurable on the finite discrete state space.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-d74d4b895ad2","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9117,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:62"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimalBellman (mdp : MDP State Action) (value : State -> Real) : Measurable (mdp.optimalBellman value)","missing":[],"search":"measurable_optimalbellman banditrlproof.finitehorizonrl.mdp.measurable_optimalbellman the optimal bellman value is measurable on the finite discrete state space. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue_mono","label":"transitionValue_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.transitionValue_mono","description":"Transition expectation is monotone in the continuation value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-5d3f826275ac","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9118,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem transitionValue_mono (mdp : MDP State Action) {left right : State -> Real} (hle : forall state, left state <= right state) (state : State) (action : Action) : mdp.transitionValue left state action <= mdp.transitionValue right state action","missing":[],"search":"transitionvalue_mono banditrlproof.finitehorizonrl.mdp.transitionvalue_mono transition expectation is monotone in the continuation value. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_mono","label":"bellmanQ_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_mono","description":"Bellman action values are monotone in their continuation value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-921eb2cc42b6","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9119,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:83"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellmanQ_mono (mdp : MDP State Action) {left right : State -> Real} (hle : forall state, left state <= right state) (state : State) (action : Action) : mdp.bellmanQ left state action <= mdp.bellmanQ right state action","missing":[],"search":"bellmanq_mono banditrlproof.finitehorizonrl.mdp.bellmanq_mono bellman action values are monotone in their continuation value. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueRemaining","label":"optimalValueRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueRemaining","description":"Backward optimal value indexed by the number of decisions remaining.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-cea16a725af4","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9120,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalValueRemaining (mdp : MDP State Action) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> Real | 0, _ => fun _ => 0 | remaining + 1, hremaining => mdp.optimalBellman (mdp.optimalValueRemaining remaining (by omega)) /-- Every backward optimal-value surface is measurable. -/ theorem measurable_optimalValueRemaining (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (mdp.optimalValueRemaining remaining hremaining)","missing":[],"search":"optimalvalueremaining banditrlproof.finitehorizonrl.mdp.optimalvalueremaining backward optimal value indexed by the number of decisions remaining. definition compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalValueRemaining","label":"measurable_optimalValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalValueRemaining","description":"Every backward optimal-value surface is measurable.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-1f0ab3ed43f5","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9121,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:100"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimalValueRemaining (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (mdp.optimalValueRemaining remaining hremaining)","missing":[],"search":"measurable_optimalvalueremaining banditrlproof.finitehorizonrl.mdp.measurable_optimalvalueremaining every backward optimal-value surface is measurable. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt","label":"optimalValueAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt","description":"Optimal value at chronological stage `stage <= horizon`.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-000dce5f6606","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9122,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:107"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalValueAt (mdp : MDP State Action) (stage : Nat) (_hstage : stage <= mdp.horizon) : State -> Real","missing":[],"search":"optimalvalueat banditrlproof.finitehorizonrl.mdp.optimalvalueat optimal value at chronological stage `stage <= horizon`. definition compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalValueAt","label":"measurable_optimalValueAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalValueAt","description":"Every chronological optimal-value surface is measurable.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-26bb24d333fe","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9123,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_optimalValueAt (mdp : MDP State Action) (stage : Nat) (hstage : stage <= mdp.horizon) : Measurable (mdp.optimalValueAt stage hstage)","missing":[],"search":"measurable_optimalvalueat banditrlproof.finitehorizonrl.mdp.measurable_optimalvalueat every chronological optimal-value surface is measurable. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueRemaining_eq_of_eq","label":"optimalValueRemaining_eq_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueRemaining_eq_of_eq","description":"Transport the dependent optimal-value recursion across equal remaining horizons.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-a451d753e8fd","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9124,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:120"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueRemaining_eq_of_eq (mdp : MDP State Action) {left right : Nat} (hleft : left <= mdp.horizon) (hright : right <= mdp.horizon) (h : left = right) : mdp.optimalValueRemaining left hleft = mdp.optimalValueRemaining right hright","missing":[],"search":"optimalvalueremaining_eq_of_eq banditrlproof.finitehorizonrl.mdp.optimalvalueremaining_eq_of_eq transport the dependent optimal-value recursion across equal remaining horizons. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_horizon","label":"optimalValueAt_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_horizon","description":"The optimal value is zero at the terminal stage.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-9664eb323ee1","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9125,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:132"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueAt_horizon (mdp : MDP State Action) : mdp.optimalValueAt mdp.horizon le_rfl = fun _ => 0","missing":[],"search":"optimalvalueat_horizon banditrlproof.finitehorizonrl.mdp.optimalvalueat_horizon the optimal value is zero at the terminal stage. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_bellman","label":"optimalValueAt_bellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_bellman","description":"Finite-action Bellman recursion for the chronological optimal value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-6a71ee30b1b8","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9126,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:144"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueAt_bellman (mdp : MDP State Action) (stage : Nat) (hstage : stage < mdp.horizon) : mdp.optimalValueAt stage (Nat.le_of_lt hstage) = mdp.optimalBellman (mdp.optimalValueAt (stage + 1) (by omega))","missing":[],"search":"optimalvalueat_bellman banditrlproof.finitehorizonrl.mdp.optimalvalueat_bellman finite-action bellman recursion for the chronological optimal value. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_mono","label":"bellman_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_mono","description":"Policy Bellman expectation is monotone in the continuation value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-ca291d70a4ab","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9127,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:172"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellman_mono {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) {left right : State -> Real} (hle : forall state, left state <= right state) (state : State) : policy.bellman stage left state <= policy.bellman stage right state","missing":[],"search":"bellman_mono banditrlproof.finitehorizonrl.markovpolicy.bellman_mono policy bellman expectation is monotone in the continuation value. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_le_optimalBellman","label":"bellman_le_optimalBellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_le_optimalBellman","description":"Every policy Bellman expectation is bounded by the finite-action maximum.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-31cb01d2c0cc","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9128,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:192"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem bellman_le_optimalBellman {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : policy.bellman stage value state <= mdp.optimalBellman value state","missing":[],"search":"bellman_le_optimalbellman banditrlproof.finitehorizonrl.markovpolicy.bellman_le_optimalbellman every policy bellman expectation is bounded by the finite-action maximum. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_le_optimalValueAt","label":"valueAt_le_optimalValueAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_le_optimalValueAt","description":"Every Markov policy value is pointwise bounded by the optimal value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-1fae62e6a936","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9129,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:213"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueAt_le_optimalValueAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Nat) (hstage : stage <= mdp.horizon) (state : State) : policy.valueAt stage hstage state <= mdp.optimalValueAt stage hstage state","missing":[],"search":"valueat_le_optimalvalueat banditrlproof.finitehorizonrl.markovpolicy.valueat_le_optimalvalueat every markov policy value is pointwise bounded by the optimal value. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy","label":"optimalPolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy","description":"Greedy deterministic Markov policy for the backward optimal value.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-f72113608ee1","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9130,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:241"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalPolicy (mdp : MDP State Action) : MarkovPolicy mdp where","missing":[],"search":"optimalpolicy banditrlproof.finitehorizonrl.mdp.optimalpolicy greedy deterministic markov policy for the backward optimal value. definition compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_bellman_eq_optimalBellman","label":"optimalPolicy_bellman_eq_optimalBellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_bellman_eq_optimalBellman","description":"One greedy-policy Bellman step is exactly the finite-action maximum.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-fde7477eeebc","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9131,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:252"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalPolicy_bellman_eq_optimalBellman (mdp : MDP State Action) (stage : Fin mdp.horizon) (state : State) : mdp.optimalPolicy.bellman stage (mdp.optimalValueAt (stage + 1) (Nat.succ_le_of_lt stage.isLt)) state = mdp.optimalBellman (mdp.optimalValueAt (stage + 1) (Nat.succ_le_of_lt stage.isLt)) state","missing":[],"search":"optimalpolicy_bellman_eq_optimalbellman banditrlproof.finitehorizonrl.mdp.optimalpolicy_bellman_eq_optimalbellman one greedy-policy bellman step is exactly the finite-action maximum. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_valueAt_eq_optimalValueAt","label":"optimalPolicy_valueAt_eq_optimalValueAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_valueAt_eq_optimalValueAt","description":"The measurable greedy deterministic policy attains the optimal value at every stage.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-4e01b69b67cb","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9132,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:269"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalPolicy_valueAt_eq_optimalValueAt (mdp : MDP State Action) (stage : Nat) (hstage : stage <= mdp.horizon) : mdp.optimalPolicy.valueAt stage hstage = mdp.optimalValueAt stage hstage","missing":[],"search":"optimalpolicy_valueat_eq_optimalvalueat banditrlproof.finitehorizonrl.mdp.optimalpolicy_valueat_eq_optimalvalueat the measurable greedy deterministic policy attains the optimal value at every stage. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_dominates_and_is_attained","label":"optimalValueAt_dominates_and_is_attained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_dominates_and_is_attained","description":"Route endpoint: the backward Bellman value dominates every Markov policy and is attained by the measurable greedy deterministic policy.","url":"../modules/banditrlproof-rl-finitehorizonoptimality/index.html#decl-4d9e36ea67e6","parent":"module:BanditRLProof.RL.FiniteHorizonOptimality","order":9133,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimality"],["Source","BanditRLProof/RL/FiniteHorizonOptimality.lean:298"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueAt_dominates_and_is_attained (mdp : MDP State Action) : (forall (policy : MarkovPolicy mdp) (stage : Nat) (hstage : stage <= mdp.horizon) (state : State), policy.valueAt stage hstage state <= mdp.optimalValueAt stage hstage state) /\\ (exists policy : MarkovPolicy mdp, forall (stage : Nat) (hstage : stage <= mdp.horizon), policy.valueAt stage hstage = mdp.optimalValueAt stage hstage)","missing":[],"search":"optimalvalueat_dominates_and_is_attained banditrlproof.finitehorizonrl.mdp.optimalvalueat_dominates_and_is_attained route endpoint: the backward bellman value dominates every markov policy and is attained by the measurable greedy deterministic policy. theorem compiled","shard":"modules/08a6cdf0c7e8de8d.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalBellman_mono","label":"optimalBellman_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalBellman_mono","description":"The finite-action optimal Bellman operator is monotone in its continuation value.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-e21504e90167","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9134,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:31"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalBellman_mono (mdp : MDP State Action) {left right : State -> Real} (hle : forall state, left state <= right state) (state : State) : mdp.optimalBellman left state <= mdp.optimalBellman right state","missing":[],"search":"optimalbellman_mono banditrlproof.finitehorizonrl.mdp.optimalbellman_mono the finite-action optimal bellman operator is monotone in its continuation value. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate","label":"OptimisticBellmanCertificate","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate","description":"An upper-value plan whose terminal value is zero and whose successor surface dominates one true optimal Bellman backup of its tail surface.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-99f66e3c4d82","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9135,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:47"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure OptimisticBellmanCertificate (mdp : MDP State Action) where","missing":[],"search":"optimisticbellmancertificate banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate an upper-value plan whose terminal value is zero and whose successor surface dominates one true optimal bellman backup of its tail surface. structure compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalBellmanCertificate","label":"optimalBellmanCertificate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.optimalBellmanCertificate","description":"The true optimal value itself is the canonical exact optimistic certificate.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-afe1d336bd21","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9136,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:60"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def optimalBellmanCertificate (mdp : MDP State Action) : OptimisticBellmanCertificate mdp where","missing":[],"search":"optimalbellmancertificate banditrlproof.finitehorizonrl.mdp.optimalbellmancertificate the true optimal value itself is the canonical exact optimistic certificate. definition compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.measurable_upperValueRemaining","label":"measurable_upperValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.measurable_upperValueRemaining","description":"Every upper-value surface is measurable on the finite discrete state space.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-1040aecab783","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9137,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_upperValueRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (certificate.upperValueRemaining remaining hremaining)","missing":[],"search":"measurable_uppervalueremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.measurable_uppervalueremaining every upper-value surface is measurable on the finite discrete state space. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.optimalValueRemaining_le_upperValueRemaining","label":"optimalValueRemaining_le_upperValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.optimalValueRemaining_le_upperValueRemaining","description":"Local Bellman optimism implies the upper values dominate the true optimal values.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-aa5cda262d2f","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9138,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalValueRemaining_le_upperValueRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : mdp.optimalValueRemaining remaining hremaining state <= certificate.upperValueRemaining remaining hremaining state","missing":[],"search":"optimalvalueremaining_le_uppervalueremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.optimalvalueremaining_le_uppervalueremaining local bellman optimism implies the upper values dominate the true optimal values. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining","label":"occupancySumRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining","description":"Recursive occupancy sum for a remaining-horizon-indexed stage cost. At a successor step, cost index `remaining` is chronological stage `horizon - (remaining + 1)`.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-b72f0045dbfd","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9139,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def occupancySumRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (cost : (remaining : Nat) -> remaining + 1 <= mdp.horizon -> State -> Real) : (remaining : Nat) -> remaining <= mdp.horizon -> Measure State -> Real | 0, _, _ => 0 | remaining + 1, hremaining, mu => let stage : Fin mdp.horizon := ⟨mdp.horizon - (remaining + 1), by omega⟩ (∫ state, cost remaining hremaining state ∂mu) + policy.occupancySumRemaining cost remaining (by omega) (policy.inducedStateKernel stage ∘ₘ mu) omit [MeasurableSingletonClass State] [Nonempty Action] in /-- Successor equation for a generic recursive occupancy sum. -/ theorem occupancySumRemaining_succ {mdp : MDP State Action} (policy : MarkovPolicy mdp) (cost : (remaining : Nat) -> remaining + 1 <= mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (mu : Measure State) : policy.occupancyS…","missing":[],"search":"occupancysumremaining banditrlproof.finitehorizonrl.markovpolicy.occupancysumremaining recursive occupancy sum for a remaining-horizon-indexed stage cost. at a successor step, cost index `remaining` is chronological stage `horizon - (remaining + 1)`. definition compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_succ","label":"occupancySumRemaining_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_succ","description":"Successor equation for a generic recursive occupancy sum.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-7c5c09b84808","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9140,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:120"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem occupancySumRemaining_succ {mdp : MDP State Action} (policy : MarkovPolicy mdp) (cost : (remaining : Nat) -> remaining + 1 <= mdp.horizon -> State -> Real) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (mu : Measure State) : policy.occupancySumRemaining cost (remaining + 1) hremaining mu = (∫ state, cost remaining hremaining state ∂mu) + policy.occupancySumRemaining cost remaining (by omega) (policy.inducedStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ ∘ₘ mu)","missing":[],"search":"occupancysumremaining_succ banditrlproof.finitehorizonrl.markovpolicy.occupancysumremaining_succ successor equation for a generic recursive occupancy sum. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_mono","label":"occupancySumRemaining_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_mono","description":"Pointwise domination of stage costs lifts to domination of their occupancy sums.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-c5feeaf673a4","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9141,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:135"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem occupancySumRemaining_mono {mdp : MDP State Action} (policy : MarkovPolicy mdp) {left right : (remaining : Nat) -> remaining + 1 <= mdp.horizon -> State -> Real} (hle : forall (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State), left remaining hremaining state <= right remaining hremaining state) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : policy.occupancySumRemaining left remaining hremaining mu <= policy.occupancySumRemaining right remaining hremaining mu","missing":[],"search":"occupancysumremaining_mono banditrlproof.finitehorizonrl.markovpolicy.occupancysumremaining_mono pointwise domination of stage costs lifts to domination of their occupancy sums. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.policyBellmanResidual","label":"policyBellmanResidual","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.policyBellmanResidual","description":"Policy Bellman residual of an optimistic upper-value plan.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-6546dc373ead","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9142,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:169"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def policyBellmanResidual {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : Real","missing":[],"search":"policybellmanresidual banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.policybellmanresidual policy bellman residual of an optimistic upper-value plan. definition compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.measurable_policyBellmanResidual","label":"measurable_policyBellmanResidual","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.measurable_policyBellmanResidual","description":"The optimistic policy Bellman residual is measurable.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-150567bca838","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9143,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:178"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_policyBellmanResidual {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) : Measurable (certificate.policyBellmanResidual policy remaining hremaining)","missing":[],"search":"measurable_policybellmanresidual banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.measurable_policybellmanresidual the optimistic policy bellman residual is measurable. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.policyBellmanResidual_nonneg","label":"policyBellmanResidual_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.policyBellmanResidual_nonneg","description":"Every policy Bellman residual of an optimistic certificate is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-58b3d6f873fc","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9144,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:188"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem policyBellmanResidual_nonneg {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : 0 <= certificate.policyBellmanResidual policy remaining hremaining state","missing":[],"search":"policybellmanresidual_nonneg banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.policybellmanresidual_nonneg every policy bellman residual of an optimistic certificate is nonnegative. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining","label":"residualOccupancyRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining","description":"Recursive true-occupancy sum of the certificate's policy Bellman residuals.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-ac03c5f77219","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9145,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:203"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def residualOccupancyRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) : Real","missing":[],"search":"residualoccupancyremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.residualoccupancyremaining recursive true-occupancy sum of the certificate's policy bellman residuals. definition compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_eq_integral_upperValueRemaining_sub_valueRemaining","label":"residualOccupancyRemaining_eq_integral_upperValueRemaining_sub_valueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_eq_integral_upperValueRemaining_sub_valueRemaining","description":"The residual occupancy sum is exactly the integrated difference between the certificate upper value and the supplied policy value.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-b8df39e05087","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9146,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:214"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem residualOccupancyRemaining_eq_integral_upperValueRemaining_sub_valueRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : certificate.residualOccupancyRemaining policy remaining hremaining mu = ∫ state, certificate.upperValueRemaining remaining hremaining state - policy.valueRemaining remaining hremaining state ∂mu","missing":[],"search":"residualoccupancyremaining_eq_integral_uppervalueremaining_sub_valueremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.residualoccupancyremaining_eq_integral_uppervalueremaining_sub_valueremaining the residual occupancy sum is exactly the integrated difference between the certificate upper value and the supplied policy value. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.optimalBellmanCertificate_residualOccupancyRemaining_eq_occupancyGapRemaining","label":"optimalBellmanCertificate_residualOccupancyRemaining_eq_occupancyGapRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.optimalBellmanCertificate_residualOccupancyRemaining_eq_occupancyGapRemaining","description":"For the canonical exact certificate, residual occupancy is the previously compiled optimality gap occupancy sum.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-0dea7f44d4a6","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9147,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:266"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem optimalBellmanCertificate_residualOccupancyRemaining_eq_occupancyGapRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : mdp.optimalBellmanCertificate.residualOccupancyRemaining policy remaining hremaining mu = policy.occupancyGapRemaining remaining hremaining mu","missing":[],"search":"optimalbellmancertificate_residualoccupancyremaining_eq_occupancygapremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.optimalbellmancertificate_residualoccupancyremaining_eq_occupancygapremaining for the canonical exact certificate, residual occupancy is the previously compiled optimality gap occupancy sum. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_nonneg","label":"residualOccupancyRemaining_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_nonneg","description":"The residual occupancy sum is nonnegative.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-8712299d54a1","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9148,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:278"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem residualOccupancyRemaining_nonneg {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : 0 <= certificate.residualOccupancyRemaining policy remaining hremaining mu","missing":[],"search":"residualoccupancyremaining_nonneg banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.residualoccupancyremaining_nonneg the residual occupancy sum is nonnegative. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.expectedRegret_le_residualOccupancyRemaining","label":"expectedRegret_le_residualOccupancyRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.expectedRegret_le_residualOccupancyRemaining","description":"Single-episode expected regret is bounded by the optimistic residual occupancy sum.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-94108d9c9b3f","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9149,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:294"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_le_residualOccupancyRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : policy.expectedRegret initialState <= certificate.residualOccupancyRemaining policy mdp.horizon le_rfl initialState","missing":[],"search":"expectedregret_le_residualoccupancyremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.expectedregret_le_residualoccupancyremaining single-episode expected regret is bounded by the optimistic residual occupancy sum. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_le_occupancySumRemaining","label":"residualOccupancyRemaining_le_occupancySumRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_le_occupancySumRemaining","description":"A pointwise bonus bound on Bellman residuals bounds the residual occupancy sum.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-f9c03b524798","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9150,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:313"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem residualOccupancyRemaining_le_occupancySumRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (bonus : (remaining : Nat) -> remaining + 1 <= mdp.horizon -> State -> Real) (hbonus : forall (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State), certificate.policyBellmanResidual policy remaining hremaining state <= bonus remaining hremaining state) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (mu : Measure State) [IsProbabilityMeasure mu] : certificate.residualOccupancyRemaining policy remaining hremaining mu <= policy.occupancySumRemaining bonus remaining hremaining mu","missing":[],"search":"residualoccupancyremaining_le_occupancysumremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.residualoccupancyremaining_le_occupancysumremaining a pointwise bonus bound on bellman residuals bounds the residual occupancy sum. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.expectedRegret_le_residual_le_occupancyBonusRemaining","label":"expectedRegret_le_residual_le_occupancyBonusRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.expectedRegret_le_residual_le_occupancyBonusRemaining","description":"Route endpoint: local true-Bellman optimism induces global value optimism; single-episode expected regret is bounded by the true-occupancy residual sum; and any pointwise bonus dominating those residuals bounds regret.","url":"../modules/banditrlproof-rl-finitehorizonoptimisticcertificate/index.html#decl-0cda61a26836","parent":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","order":9151,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonOptimisticCertificate"],["Source","BanditRLProof/RL/FiniteHorizonOptimisticCertificate.lean:334"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedRegret_le_residual_le_occupancyBonusRemaining {mdp : MDP State Action} (certificate : OptimisticBellmanCertificate mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (bonus : (remaining : Nat) -> remaining + 1 <= mdp.horizon -> State -> Real) (hbonus : forall (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State), certificate.policyBellmanResidual policy remaining hremaining state <= bonus remaining hremaining state) : 0 <= certificate.residualOccupancyRemaining policy mdp.horizon le_rfl initialState /\\ policy.expectedRegret initialState <= certificate.residualOccupancyRemaining policy mdp.horizon le_rfl initialState /\\ certificate.residualOccupancyRemaining policy mdp.horizon le_rfl initialState <= policy.occupancySumRemaining bonus mdp.horizon le_rfl initialState /\\ policy.expectedRegret initialSta…","missing":[],"search":"expectedregret_le_residual_le_occupancybonusremaining banditrlproof.finitehorizonrl.mdp.optimisticbellmancertificate.expectedregret_le_residual_le_occupancybonusremaining route endpoint: local true-bellman optimism induces global value optimism; single-episode expected regret is bounded by the true-occupancy residual sum; and any pointwise bonus dominating those residuals bounds regret. theorem compiled","shard":"modules/25f18bdd51c05545.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy","label":"MarkovPolicy","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy","description":"A Markov action kernel for every decision stage before `mdp.horizon`.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-dd1f07a67f34","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9152,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure MarkovPolicy (mdp : MDP State Action) where","missing":[],"search":"markovpolicy banditrlproof.finitehorizonrl.markovpolicy a markov action kernel for every decision stage before `mdp.horizon`. structure compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.inducedStateKernel","label":"inducedStateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.inducedStateKernel","description":"State transition kernel induced by sampling the policy action and then the MDP transition.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-3cee3e7a3dde","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9153,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def inducedStateKernel {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel State State","missing":[],"search":"inducedstatekernel banditrlproof.finitehorizonrl.markovpolicy.inducedstatekernel state transition kernel induced by sampling the policy action and then the mdp transition. definition compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman","label":"bellman","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman","description":"Bellman expectation under one stage of the supplied Markov policy.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-17ec8bb4d067","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9154,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def bellman {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : Real","missing":[],"search":"bellman banditrlproof.finitehorizonrl.markovpolicy.bellman bellman expectation under one stage of the supplied markov policy. definition compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_bellman","label":"measurable_bellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_bellman","description":"A measurable continuation value gives a measurable policy Bellman value.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-ed93a8c994cc","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9155,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:65"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_bellman {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) {value : State -> Real} (hvalue : Measurable value) : Measurable (policy.bellman stage value)","missing":[],"search":"measurable_bellman banditrlproof.finitehorizonrl.markovpolicy.measurable_bellman a measurable continuation value gives a measurable policy bellman value. theorem compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining","label":"valueRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining","description":"Backward policy value indexed by the number of decisions remaining. The first kernel used at `remaining` is chronological stage `horizon - remaining`.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-9b6d116e3c17","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9156,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def valueRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> Real | 0, _ => fun _ => 0 | remaining + 1, hremaining => policy.bellman ⟨mdp.horizon - (remaining + 1), by omega⟩ (policy.valueRemaining remaining (by omega)) /-- The backward policy value is measurable for every valid remaining horizon. -/ theorem measurable_valueRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (policy.valueRemaining remaining hremaining)","missing":[],"search":"valueremaining banditrlproof.finitehorizonrl.markovpolicy.valueremaining backward policy value indexed by the number of decisions remaining. the first kernel used at `remaining` is chronological stage `horizon - remaining`. definition compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_valueRemaining","label":"measurable_valueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_valueRemaining","description":"The backward policy value is measurable for every valid remaining horizon.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-6bbd77b06661","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9157,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_valueRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (policy.valueRemaining remaining hremaining)","missing":[],"search":"measurable_valueremaining banditrlproof.finitehorizonrl.markovpolicy.measurable_valueremaining the backward policy value is measurable for every valid remaining horizon. theorem compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt","label":"valueAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt","description":"Policy value at a chronological stage `stage <= horizon`.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-5902fe75f174","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9158,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def valueAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Nat) (_hstage : stage <= mdp.horizon) : State -> Real","missing":[],"search":"valueat banditrlproof.finitehorizonrl.markovpolicy.valueat policy value at a chronological stage `stage <= horizon`. definition compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_valueAt","label":"measurable_valueAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_valueAt","description":"Every chronological policy-value surface is measurable.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-be7b020cdf38","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9159,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:105"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_valueAt {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Nat) (hstage : stage <= mdp.horizon) : Measurable (policy.valueAt stage hstage)","missing":[],"search":"measurable_valueat banditrlproof.finitehorizonrl.markovpolicy.measurable_valueat every chronological policy-value surface is measurable. theorem compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_eq_of_eq","label":"valueRemaining_eq_of_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_eq_of_eq","description":"Transport `valueRemaining` across equality of the remaining horizon.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-24dffea87129","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9160,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueRemaining_eq_of_eq {mdp : MDP State Action} (policy : MarkovPolicy mdp) {left right : Nat} (hleft : left <= mdp.horizon) (hright : right <= mdp.horizon) (h : left = right) : policy.valueRemaining left hleft = policy.valueRemaining right hright","missing":[],"search":"valueremaining_eq_of_eq banditrlproof.finitehorizonrl.markovpolicy.valueremaining_eq_of_eq transport `valueremaining` across equality of the remaining horizon. theorem compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_horizon","label":"valueAt_horizon","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_horizon","description":"The finite-horizon policy value is zero at the terminal stage.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-1c1473f1b7af","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9161,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueAt_horizon {mdp : MDP State Action} (policy : MarkovPolicy mdp) : policy.valueAt mdp.horizon le_rfl = fun _ => 0","missing":[],"search":"valueat_horizon banditrlproof.finitehorizonrl.markovpolicy.valueat_horizon the finite-horizon policy value is zero at the terminal stage. theorem compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_bellman","label":"valueAt_bellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_bellman","description":"Finite-horizon policy evaluation satisfies the Bellman recursion at every decision stage.","url":"../modules/banditrlproof-rl-finitehorizonpolicy/index.html#decl-fdf7c1a5f79e","parent":"module:BanditRLProof.RL.FiniteHorizonPolicy","order":9162,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonPolicy"],["Source","BanditRLProof/RL/FiniteHorizonPolicy.lean:138"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueAt_bellman {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Nat) (hstage : stage < mdp.horizon) : policy.valueAt stage (Nat.le_of_lt hstage) = policy.bellman ⟨stage, hstage⟩ (policy.valueAt (stage + 1) (by omega))","missing":[],"search":"valueat_bellman banditrlproof.finitehorizonrl.markovpolicy.valueat_bellman finite-horizon policy evaluation satisfies the bellman recursion at every decision stage. theorem compiled","shard":"modules/273e54deb76050f3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt","label":"stateAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateAt","description":"def stateAt : {n : Nat} -> State -> StepTrace Action State n -> Fin n -> State | 0, _initial, _trace, coordinate => Fin.elim0 coordinate | _n + 1, initial, trace, coordinate => Fin.cases initial (fun previous => (trace previous.castSucc).2) coordinate @[simp] theorem stateAt_zero (initial : State) (head : Action × State) (tail : StepTrace Action State n) : stateAt initial (Fin.cons head tail) (0 : Fin (n + 1)) = ini…","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-7fabec363816","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9163,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:24"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def stateAt : {n : Nat} -> State -> StepTrace Action State n -> Fin n -> State | 0, _initial, _trace, coordinate => Fin.elim0 coordinate | _n + 1, initial, trace, coordinate => Fin.cases initial (fun previous => (trace previous.castSucc).2) coordinate @[simp] theorem stateAt_zero (initial : State) (head : Action × State) (tail : StepTrace Action State n) : stateAt initial (Fin.cons head tail) (0 : Fin (n + 1)) = initial","missing":[],"search":"stateat banditrlproof.finitehorizonrl.steptrace.stateat def stateat : {n : nat} -> state -> steptrace action state n -> fin n -> state | 0, _initial, _trace, coordinate => fin.elim0 coordinate | _n + 1, initial, trace, coordinate => fin.cases initial (fun previous => (trace previous.castsucc).2) coordinate @[simp] theorem stateat_zero (initial : state) (head : action × state) (tail : steptrace action state n) : stateat initial (fin.cons head tail) (0 : fin (n + 1)) = initial definition compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_zero","label":"stateAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_zero","description":"theorem stateAt_zero (initial : State) (head : Action × State) (tail : StepTrace Action State n) : stateAt initial (Fin.cons head tail) (0 : Fin (n + 1)) = initial","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-f058d296e67b","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9164,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateAt_zero (initial : State) (head : Action × State) (tail : StepTrace Action State n) : stateAt initial (Fin.cons head tail) (0 : Fin (n + 1)) = initial","missing":[],"search":"stateat_zero banditrlproof.finitehorizonrl.steptrace.stateat_zero theorem stateat_zero (initial : state) (head : action × state) (tail : steptrace action state n) : stateat initial (fin.cons head tail) (0 : fin (n + 1)) = initial theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_zero_apply","label":"stateAt_zero_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_zero_apply","description":"theorem stateAt_zero_apply (initial : State) (trace : StepTrace Action State (n + 1)) : stateAt initial trace (0 : Fin (n + 1)) = initial","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-9776a36ba363","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9165,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateAt_zero_apply (initial : State) (trace : StepTrace Action State (n + 1)) : stateAt initial trace (0 : Fin (n + 1)) = initial","missing":[],"search":"stateat_zero_apply banditrlproof.finitehorizonrl.steptrace.stateat_zero_apply theorem stateat_zero_apply (initial : state) (trace : steptrace action state (n + 1)) : stateat initial trace (0 : fin (n + 1)) = initial theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_succ","label":"stateAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_succ","description":"theorem stateAt_succ (initial : State) (head : Action × State) (tail : StepTrace Action State n) (coordinate : Fin n) : stateAt initial (Fin.cons head tail) coordinate.succ = stateAt head.2 tail coordinate","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-7a3a5312149b","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9166,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:42"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateAt_succ (initial : State) (head : Action × State) (tail : StepTrace Action State n) (coordinate : Fin n) : stateAt initial (Fin.cons head tail) coordinate.succ = stateAt head.2 tail coordinate","missing":[],"search":"stateat_succ banditrlproof.finitehorizonrl.steptrace.stateat_succ theorem stateat_succ (initial : state) (head : action × state) (tail : steptrace action state n) (coordinate : fin n) : stateat initial (fin.cons head tail) coordinate.succ = stateat head.2 tail coordinate theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_eq_if","label":"stateAt_eq_if","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_eq_if","description":"theorem stateAt_eq_if {n : Nat} (initial : State) (trace : StepTrace Action State n) (coordinate : Fin n) : stateAt initial trace coordinate = if _hzero : coordinate.val = 0 then initial else (trace ⟨coordinate.val - 1, by omega⟩).2","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-781bb2256aae","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9167,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateAt_eq_if {n : Nat} (initial : State) (trace : StepTrace Action State n) (coordinate : Fin n) : stateAt initial trace coordinate = if _hzero : coordinate.val = 0 then initial else (trace ⟨coordinate.val - 1, by omega⟩).2","missing":[],"search":"stateat_eq_if banditrlproof.finitehorizonrl.steptrace.stateat_eq_if theorem stateat_eq_if {n : nat} (initial : state) (trace : steptrace action state n) (coordinate : fin n) : stateat initial trace coordinate = if _hzero : coordinate.val = 0 then initial else (trace ⟨coordinate.val - 1, by omega⟩).2 theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionAt","label":"stateActionAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateActionAt","description":"def stateActionAt {n : Nat} (initial : State) (trace : StepTrace Action State n) (coordinate : Fin n) : State × Action","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-d714de5fccdc","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9168,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def stateActionAt {n : Nat} (initial : State) (trace : StepTrace Action State n) (coordinate : Fin n) : State × Action","missing":[],"search":"stateactionat banditrlproof.finitehorizonrl.steptrace.stateactionat def stateactionat {n : nat} (initial : state) (trace : steptrace action state n) (coordinate : fin n) : state × action definition compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionNextAt","label":"stateActionNextAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateActionNextAt","description":"def stateActionNextAt {n : Nat} (initial : State) (trace : StepTrace Action State n) (coordinate : Fin n) : (State × Action) × State","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-dd8ff380b5f3","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9169,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:71"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def stateActionNextAt {n : Nat} (initial : State) (trace : StepTrace Action State n) (coordinate : Fin n) : (State × Action) × State","missing":[],"search":"stateactionnextat banditrlproof.finitehorizonrl.steptrace.stateactionnextat def stateactionnextat {n : nat} (initial : state) (trace : steptrace action state n) (coordinate : fin n) : (state × action) × state definition compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionAt_cons_succ","label":"stateActionAt_cons_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateActionAt_cons_succ","description":"theorem stateActionAt_cons_succ (initial : State) (head : Action × State) (tail : StepTrace Action State n) (coordinate : Fin n) : stateActionAt initial (Fin.cons head tail) coordinate.succ = stateActionAt head.2 tail coordinate","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-b3fc4d485b15","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9170,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:77"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateActionAt_cons_succ (initial : State) (head : Action × State) (tail : StepTrace Action State n) (coordinate : Fin n) : stateActionAt initial (Fin.cons head tail) coordinate.succ = stateActionAt head.2 tail coordinate","missing":[],"search":"stateactionat_cons_succ banditrlproof.finitehorizonrl.steptrace.stateactionat_cons_succ theorem stateactionat_cons_succ (initial : state) (head : action × state) (tail : steptrace action state n) (coordinate : fin n) : stateactionat initial (fin.cons head tail) coordinate.succ = stateactionat head.2 tail coordinate theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionNextAt_cons_succ","label":"stateActionNextAt_cons_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.stateActionNextAt_cons_succ","description":"theorem stateActionNextAt_cons_succ (initial : State) (head : Action × State) (tail : StepTrace Action State n) (coordinate : Fin n) : stateActionNextAt initial (Fin.cons head tail) coordinate.succ = stateActionNextAt head.2 tail coordinate","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-abd45f046e8f","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9171,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:84"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateActionNextAt_cons_succ (initial : State) (head : Action × State) (tail : StepTrace Action State n) (coordinate : Fin n) : stateActionNextAt initial (Fin.cons head tail) coordinate.succ = stateActionNextAt head.2 tail coordinate","missing":[],"search":"stateactionnextat_cons_succ banditrlproof.finitehorizonrl.steptrace.stateactionnextat_cons_succ theorem stateactionnextat_cons_succ (initial : state) (head : action × state) (tail : steptrace action state n) (coordinate : fin n) : stateactionnextat initial (fin.cons head tail) coordinate.succ = stateactionnextat head.2 tail coordinate theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_stateActionAt","label":"measurable_stateActionAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.measurable_stateActionAt","description":"theorem measurable_stateActionAt {n : Nat} (initial : State) (coordinate : Fin n) : Measurable (fun trace : StepTrace Action State n => stateActionAt initial trace coordinate)","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-62b70d8ff802","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9172,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:94"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stateActionAt {n : Nat} (initial : State) (coordinate : Fin n) : Measurable (fun trace : StepTrace Action State n => stateActionAt initial trace coordinate)","missing":[],"search":"measurable_stateactionat banditrlproof.finitehorizonrl.steptrace.measurable_stateactionat theorem measurable_stateactionat {n : nat} (initial : state) (coordinate : fin n) : measurable (fun trace : steptrace action state n => stateactionat initial trace coordinate) theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_stateActionNextAt","label":"measurable_stateActionNextAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.measurable_stateActionNextAt","description":"theorem measurable_stateActionNextAt {n : Nat} (initial : State) (coordinate : Fin n) : Measurable (fun trace : StepTrace Action State n => stateActionNextAt initial trace coordinate)","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-1792dd6b9371","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9173,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:99"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stateActionNextAt {n : Nat} (initial : State) (coordinate : Fin n) : Measurable (fun trace : StepTrace Action State n => stateActionNextAt initial trace coordinate)","missing":[],"search":"measurable_stateactionnextat banditrlproof.finitehorizonrl.steptrace.measurable_stateactionnextat theorem measurable_stateactionnextat {n : nat} (initial : state) (coordinate : fin n) : measurable (fun trace : steptrace action state n => stateactionnextat initial trace coordinate) theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stepTrace_stateAt_eq_trajectoryStateAt","label":"stepTrace_stateAt_eq_trajectoryStateAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.stepTrace_stateAt_eq_trajectoryStateAt","description":"theorem stepTrace_stateAt_eq_trajectoryStateAt (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : StepTrace.stateAt trajectory.1 trajectory.2 stage = mdp.trajectoryStateAt trajectory stage","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-a968921a0622","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9174,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stepTrace_stateAt_eq_trajectoryStateAt (mdp : MDP State Action) (trajectory : State × StepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : StepTrace.stateAt trajectory.1 trajectory.2 stage = mdp.trajectoryStateAt trajectory stage","missing":[],"search":"steptrace_stateat_eq_trajectorystateat banditrlproof.finitehorizonrl.mdp.steptrace_stateat_eq_trajectorystateat theorem steptrace_stateat_eq_trajectorystateat (mdp : mdp state action) (trajectory : state × steptrace action state mdp.horizon) (stage : fin mdp.horizon) : steptrace.stateat trajectory.1 trajectory.2 stage = mdp.trajectorystateat trajectory stage theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel_apply_singleton","label":"actionStateKernel_apply_singleton","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel_apply_singleton","description":"theorem actionStateKernel_apply_singleton {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.actionStateKernel stage state {(action, nextState)} = policy.actionKernel stage state {action} * mdp.transition (state, action) {nextState}","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-8324fa8762f6","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9175,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:128"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_apply_singleton {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.actionStateKernel stage state {(action, nextState)} = policy.actionKernel stage state {action} * mdp.transition (state, action) {nextState}","missing":[],"search":"actionstatekernel_apply_singleton banditrlproof.finitehorizonrl.markovpolicy.actionstatekernel_apply_singleton theorem actionstatekernel_apply_singleton {mdp : mdp state action} (policy : markovpolicy mdp) (stage : fin mdp.horizon) (state : state) (action : action) (nextstate : state) : policy.actionstatekernel stage state {(action, nextstate)} = policy.actionkernel stage state {action} * mdp.transition (state, action) {nextstate} theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel_apply_actionSet","label":"actionStateKernel_apply_actionSet","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel_apply_actionSet","description":"theorem actionStateKernel_apply_actionSet {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) : policy.actionStateKernel stage state ({action} ×ˢ Set.univ) = policy.actionKernel stage state {action}","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-baedc3f8ff33","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9176,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_apply_actionSet {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) : policy.actionStateKernel stage state ({action} ×ˢ Set.univ) = policy.actionKernel stage state {action}","missing":[],"search":"actionstatekernel_apply_actionset banditrlproof.finitehorizonrl.markovpolicy.actionstatekernel_apply_actionset theorem actionstatekernel_apply_actionset {mdp : mdp state action} (policy : markovpolicy mdp) (stage : fin mdp.horizon) (state : state) (action : action) : policy.actionstatekernel stage state ({action} ×ˢ set.univ) = policy.actionkernel stage state {action} theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_map_head","label":"trajectoryKernelRemaining_map_head","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_map_head","description":"theorem trajectoryKernelRemaining_map_head {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (initial : State) : (policy.trajectoryKernelRemaining (remaining + 1) hremaining initial).map (fun trace => trace 0) = policy.actionStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ initial","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-c1629bd25adb","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9177,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:154"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_map_head {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (initial : State) : (policy.trajectoryKernelRemaining (remaining + 1) hremaining initial).map (fun trace => trace 0) = policy.actionStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ initial","missing":[],"search":"trajectorykernelremaining_map_head banditrlproof.finitehorizonrl.markovpolicy.trajectorykernelremaining_map_head theorem trajectorykernelremaining_map_head {mdp : mdp state action} (policy : markovpolicy mdp) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) (initial : state) : (policy.trajectorykernelremaining (remaining + 1) hremaining initial).map (fun trace => trace 0) = policy.actionstatekernel ⟨mdp.horizon - (remaining + 1), by omega⟩ initial theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_transitionEvent_eq_visitEvent_mul","label":"trajectoryKernelRemaining_transitionEvent_eq_visitEvent_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_transitionEvent_eq_visitEvent_mul","description":"theorem trajectoryKernelRemaining_transitionEvent_eq_visitEvent_mul {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (initial state : State) (action : Action) (nextState : State) (coordinate : Fin remaining) : (policy.trajectoryKernelRemaining remaining hremaining initial) {trace | StepTrace.stateActionNextAt initial trace coordinate = ((state, action), n…","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-8f66254f52f6","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9178,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:170"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_transitionEvent_eq_visitEvent_mul {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (initial state : State) (action : Action) (nextState : State) (coordinate : Fin remaining) : (policy.trajectoryKernelRemaining remaining hremaining initial) {trace | StepTrace.stateActionNextAt initial trace coordinate = ((state, action), nextState)} = (policy.trajectoryKernelRemaining remaining hremaining initial) {trace | StepTrace.stateActionAt initial trace coordinate = (state, action)} * mdp.transition (state, action) {nextState}","missing":[],"search":"trajectorykernelremaining_transitionevent_eq_visitevent_mul banditrlproof.finitehorizonrl.markovpolicy.trajectorykernelremaining_transitionevent_eq_visitevent_mul theorem trajectorykernelremaining_transitionevent_eq_visitevent_mul {mdp : mdp state action} (policy : markovpolicy mdp) (remaining : nat) (hremaining : remaining <= mdp.horizon) (initial state : state) (action : action) (nextstate : state) (coordinate : fin remaining) : (policy.trajectorykernelremaining remaining hremaining initial) {trace | steptrace.stateactionnextat initial trace coordinate = ((state, action), nextstate)} = (policy.trajectorykernelremaining remaining hremaining initial) {trace | steptrace.stateactionat initial trace coordinate = (state, action)} * mdp.transition (state, action) {nextstate} theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure_transitionEvent_eq_visitEvent_mul","label":"trajectoryMeasure_transitionEvent_eq_visitEvent_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure_transitionEvent_eq_visitEvent_mul","description":"theorem trajectoryMeasure_transitionEvent_eq_visitEvent_mul {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.trajectoryMeasure initialState {trajectory | StepTrace.stateActionNextAt trajectory.1 trajectory.2 stage = ((state, action), nextState)} = policy.traj…","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-497bea009701","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9179,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:260"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_transitionEvent_eq_visitEvent_mul {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.trajectoryMeasure initialState {trajectory | StepTrace.stateActionNextAt trajectory.1 trajectory.2 stage = ((state, action), nextState)} = policy.trajectoryMeasure initialState {trajectory | StepTrace.stateActionAt trajectory.1 trajectory.2 stage = (state, action)} * mdp.transition (state, action) {nextState}","missing":[],"search":"trajectorymeasure_transitionevent_eq_visitevent_mul banditrlproof.finitehorizonrl.markovpolicy.trajectorymeasure_transitionevent_eq_visitevent_mul theorem trajectorymeasure_transitionevent_eq_visitevent_mul {mdp : mdp state action} (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (stage : fin mdp.horizon) (state : state) (action : action) (nextstate : state) : policy.trajectorymeasure initialstate {trajectory | steptrace.stateactionnextat trajectory.1 trajectory.2 stage = ((state, action), nextstate)} = policy.trajectorymeasure initialstate {trajectory | steptrace.stateactionat trajectory.1 trajectory.2 stage = (state, action)} * mdp.transition (state, action) {nextstate} theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_eq_stageVisitProbability_mul_transition","label":"stageTransitionJointProbability_eq_stageVisitProbability_mul_transition","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_eq_stageVisitProbability_mul_transition","description":"theorem stageTransitionJointProbability_eq_stageVisitProbability_mul_transition [DecidableEq State] [DecidableEq Action] {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.stageTransitionJointProbability initialState stage state action nextState = policy.stageV…","url":"../modules/banditrlproof-rl-finitehorizonstagetransitionjointfactorization/index.html#decl-0513b2720c3d","parent":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","order":9180,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageTransitionJointFactorization.lean:291"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageTransitionJointProbability_eq_stageVisitProbability_mul_transition [DecidableEq State] [DecidableEq Action] {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : policy.stageTransitionJointProbability initialState stage state action nextState = policy.stageVisitProbability initialState stage state action * (mdp.transition (state, action) {nextState}).toReal","missing":[],"search":"stagetransitionjointprobability_eq_stagevisitprobability_mul_transition banditrlproof.finitehorizonrl.markovpolicy.stagetransitionjointprobability_eq_stagevisitprobability_mul_transition theorem stagetransitionjointprobability_eq_stagevisitprobability_mul_transition [decidableeq state] [decidableeq action] {mdp : mdp state action} (policy : markovpolicy mdp) (initialstate : measure state) [isprobabilitymeasure initialstate] (stage : fin mdp.horizon) (state : state) (action : action) (nextstate : state) : policy.stagetransitionjointprobability initialstate stage state action nextstate = policy.stagevisitprobability initialstate stage state action * (mdp.transition (state, action) {nextstate}).toreal theorem compiled","shard":"modules/b8553f8c07d7c892.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate","label":"stageOfRemainingCoordinate","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate","description":"Chronological MDP stage represented by one coordinate of a remaining trace.","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-9c702443738a","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9181,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def stageOfRemainingCoordinate (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (coordinate : Fin remaining) : Fin mdp.horizon","missing":[],"search":"stageofremainingcoordinate banditrlproof.finitehorizonrl.mdp.stageofremainingcoordinate chronological mdp stage represented by one coordinate of a remaining trace. definition compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate_succ","label":"stageOfRemainingCoordinate_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate_succ","description":"theorem stageOfRemainingCoordinate_succ (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (coordinate : Fin remaining) : mdp.stageOfRemainingCoordinate (remaining + 1) hremaining coordinate.succ = mdp.stageOfRemainingCoordinate remaining (by omega) coordinate","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-f5d703a5fbe5","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9182,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageOfRemainingCoordinate_succ (mdp : MDP State Action) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (coordinate : Fin remaining) : mdp.stageOfRemainingCoordinate (remaining + 1) hremaining coordinate.succ = mdp.stageOfRemainingCoordinate remaining (by omega) coordinate","missing":[],"search":"stageofremainingcoordinate_succ banditrlproof.finitehorizonrl.mdp.stageofremainingcoordinate_succ theorem stageofremainingcoordinate_succ (mdp : mdp state action) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) (coordinate : fin remaining) : mdp.stageofremainingcoordinate (remaining + 1) hremaining coordinate.succ = mdp.stageofremainingcoordinate remaining (by omega) coordinate theorem compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate_full","label":"stageOfRemainingCoordinate_full","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate_full","description":"theorem stageOfRemainingCoordinate_full (mdp : MDP State Action) (stage : Fin mdp.horizon) : mdp.stageOfRemainingCoordinate mdp.horizon le_rfl stage = stage","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-f3b202d47f29","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9183,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:48"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageOfRemainingCoordinate_full (mdp : MDP State Action) (stage : Fin mdp.horizon) : mdp.stageOfRemainingCoordinate mdp.horizon le_rfl stage = stage","missing":[],"search":"stageofremainingcoordinate_full banditrlproof.finitehorizonrl.mdp.stageofremainingcoordinate_full theorem stageofremainingcoordinate_full (mdp : mdp state action) (stage : fin mdp.horizon) : mdp.stageofremainingcoordinate mdp.horizon le_rfl stage = stage theorem compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_visitEvent_eq_stateEvent_mul_action","label":"trajectoryKernelRemaining_visitEvent_eq_stateEvent_mul_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_visitEvent_eq_stateEvent_mul_action","description":"Inside a remaining generated trace, a state/action visit factors into the state event and the action singleton selected by the chronological policy kernel.","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-5e4cb3a8ca79","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9184,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryKernelRemaining_visitEvent_eq_stateEvent_mul_action {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (initial state : State) (action : Action) (coordinate : Fin remaining) : (policy.trajectoryKernelRemaining remaining hremaining initial) {trace | StepTrace.stateActionAt initial trace coordinate = (state, action)} = (policy.trajectoryKernelRemaining remaining hremaining initial) {trace | StepTrace.stateAt initial trace coordinate = state} * policy.actionKernel (mdp.stageOfRemainingCoordinate remaining hremaining coordinate) state {action}","missing":[],"search":"trajectorykernelremaining_visitevent_eq_stateevent_mul_action banditrlproof.finitehorizonrl.markovpolicy.trajectorykernelremaining_visitevent_eq_stateevent_mul_action inside a remaining generated trace, a state/action visit factors into the state event and the action singleton selected by the chronological policy kernel. theorem compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure_visitEvent_eq_stateEvent_mul_action","label":"trajectoryMeasure_visitEvent_eq_stateEvent_mul_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure_visitEvent_eq_stateEvent_mul_action","description":"The full trajectory visit event factors into its state event and action mass.","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-e73c467f4722","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9185,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:140"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem trajectoryMeasure_visitEvent_eq_stateEvent_mul_action {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) : policy.trajectoryMeasure initialState {trajectory | (mdp.episodeStepOfTrajectory trajectory stage).state = state /\\ (mdp.episodeStepOfTrajectory trajectory stage).action = action} = policy.trajectoryMeasure initialState {trajectory | mdp.trajectoryStateAt trajectory stage = state} * policy.actionKernel stage state {action}","missing":[],"search":"trajectorymeasure_visitevent_eq_stateevent_mul_action banditrlproof.finitehorizonrl.markovpolicy.trajectorymeasure_visitevent_eq_stateevent_mul_action the full trajectory visit event factors into its state event and action mass. theorem compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability","label":"stageStateProbability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability","description":"Genuine state probability at one chronological stage of a generated trajectory.","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-0ad03600f56f","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9186,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:186"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stageStateProbability {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) : Real","missing":[],"search":"stagestateprobability banditrlproof.finitehorizonrl.markovpolicy.stagestateprobability genuine state probability at one chronological stage of a generated trajectory. definition compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_eq_stageStateProbability_mul_action","label":"stageVisitProbability_eq_stageStateProbability_mul_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_eq_stageStateProbability_mul_action","description":"A generated state/action visit probability is state mass times action mass.","url":"../modules/banditrlproof-rl-finitehorizonstagevisitfactorization/index.html#decl-5030e790eb73","parent":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","order":9187,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStageVisitFactorization"],["Source","BanditRLProof/RL/FiniteHorizonStageVisitFactorization.lean:194"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stageVisitProbability_eq_stageStateProbability_mul_action {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) : policy.stageVisitProbability initialState stage state action = policy.stageStateProbability initialState stage state * (policy.actionKernel stage state {action}).toReal","missing":[],"search":"stagevisitprobability_eq_stagestateprobability_mul_action banditrlproof.finitehorizonrl.markovpolicy.stagevisitprobability_eq_stagestateprobability_mul_action a generated state/action visit probability is state mass times action mass. theorem compiled","shard":"modules/6bc2deb294183cbd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel","label":"MeanCompatibleRewardKernel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel","description":"A stochastic Real reward kernel whose selected rewards are integrable and have the mean stored in the finite-horizon MDP reward field.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-ef020fc0c501","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9188,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure MeanCompatibleRewardKernel (mdp : MDP State Action) where","missing":[],"search":"meancompatiblerewardkernel banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel a stochastic real reward kernel whose selected rewards are integrable and have the mean stored in the finite-horizon mdp reward field. structure compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardNextStateKernel","label":"rewardNextStateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardNextStateKernel","description":"The conditionally independent joint law of reward and next state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-bcb482ca4396","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9189,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:54"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardNextStateKernel (source : MeanCompatibleRewardKernel mdp) : ProbabilityTheory.Kernel (Prod State Action) (Prod Real State)","missing":[],"search":"rewardnextstatekernel banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewardnextstatekernel the conditionally independent joint law of reward and next state. definition compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellmanQ","label":"stochasticBellmanQ","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellmanQ","description":"Expected sampled reward plus continuation value under the joint kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-e22b011bc339","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9190,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:66"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticBellmanQ (source : MeanCompatibleRewardKernel mdp) (value : State -> Real) (state : State) (action : Action) : Real","missing":[],"search":"stochasticbellmanq banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticbellmanq expected sampled reward plus continuation value under the joint kernel. definition compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_reward_add_value","label":"integrable_reward_add_value","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_reward_add_value","description":"The sampled one-step return is integrable for measurable continuation values.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-dca22d1a7ffd","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9191,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_reward_add_value (source : MeanCompatibleRewardKernel mdp) {value : State -> Real} (hvalue : Measurable value) (state : State) (action : Action) : Integrable (fun pair : Prod Real State => pair.1 + value pair.2) (source.rewardNextStateKernel (state, action))","missing":[],"search":"integrable_reward_add_value banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.integrable_reward_add_value the sampled one-step return is integrable for measurable continuation values. theorem compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellmanQ_eq_bellmanQ","label":"stochasticBellmanQ_eq_bellmanQ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellmanQ_eq_bellmanQ","description":"Sampling a mean-compatible reward preserves the existing Bellman action value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-51b52235fb3b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9192,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticBellmanQ_eq_bellmanQ (source : MeanCompatibleRewardKernel mdp) {value : State -> Real} (hvalue : Measurable value) (state : State) (action : Action) : source.stochasticBellmanQ value state action = mdp.bellmanQ value state action","missing":[],"search":"stochasticbellmanq_eq_bellmanq banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticbellmanq_eq_bellmanq sampling a mean-compatible reward preserves the existing bellman action value. theorem compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellman","label":"stochasticBellman","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellman","description":"Policy expectation formed from the sampled stochastic Bellman action value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-324cf49267a1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9193,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticBellman (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (value : State -> Real) (state : State) : Real","missing":[],"search":"stochasticbellman banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticbellman policy expectation formed from the sampled stochastic bellman action value. definition compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellman_eq_bellman","label":"stochasticBellman_eq_bellman","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellman_eq_bellman","description":"The stochastic policy Bellman operator equals the existing mean operator.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-73a89750cef6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9194,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:134"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticBellman_eq_bellman (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) {value : State -> Real} (hvalue : Measurable value) : source.stochasticBellman policy stage value = policy.bellman stage value","missing":[],"search":"stochasticbellman_eq_bellman banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticbellman_eq_bellman the stochastic policy bellman operator equals the existing mean operator. theorem compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueRemaining","label":"stochasticValueRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueRemaining","description":"Backward policy value computed with sampled stochastic reward laws.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-526736f98345","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9195,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticValueRemaining (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) : (remaining : Nat) -> remaining <= mdp.horizon -> State -> Real | 0, _ => fun _ => 0 | remaining + 1, hremaining => source.stochasticBellman policy ⟨mdp.horizon - (remaining + 1), by omega⟩ (source.stochasticValueRemaining policy remaining (by omega)) /-- Stochastic backward policy evaluation equals mean-reward evaluation. -/ theorem stochasticValueRemaining_eq_valueRemaining (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : source.stochasticValueRemaining policy remaining hremaining = policy.valueRemaining remaining hremaining","missing":[],"search":"stochasticvalueremaining banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticvalueremaining backward policy value computed with sampled stochastic reward laws. definition compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueRemaining_eq_valueRemaining","label":"stochasticValueRemaining_eq_valueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueRemaining_eq_valueRemaining","description":"Stochastic backward policy evaluation equals mean-reward evaluation.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-38b8199287ca","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9196,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticValueRemaining_eq_valueRemaining (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : source.stochasticValueRemaining policy remaining hremaining = policy.valueRemaining remaining hremaining","missing":[],"search":"stochasticvalueremaining_eq_valueremaining banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticvalueremaining_eq_valueremaining stochastic backward policy evaluation equals mean-reward evaluation. theorem compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueAt","label":"stochasticValueAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueAt","description":"Stochastic policy value at a chronological stage.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-b52553bc9b7b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9197,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:176"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticValueAt (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Nat) (_hstage : stage <= mdp.horizon) : State -> Real","missing":[],"search":"stochasticvalueat banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticvalueat stochastic policy value at a chronological stage. definition compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueAt_eq_valueAt","label":"stochasticValueAt_eq_valueAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueAt_eq_valueAt","description":"Every chronological stochastic value equals the existing policy value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-380f37e1f05d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9198,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:184"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticValueAt_eq_valueAt (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Nat) (hstage : stage <= mdp.horizon) : source.stochasticValueAt policy stage hstage = policy.valueAt stage hstage","missing":[],"search":"stochasticvalueat_eq_valueat banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticvalueat_eq_valueat every chronological stochastic value equals the existing policy value. theorem compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.deterministic","label":"deterministic","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.deterministic","description":"The deterministic MDP reward, viewed as a kernel, is mean-compatible.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-a7551e8a063d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9199,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:193"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def deterministic (mdp : MDP State Action) : MeanCompatibleRewardKernel mdp where","missing":[],"search":"deterministic banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.deterministic the deterministic mdp reward, viewed as a kernel, is mean-compatible. definition compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.meanPlanningTransport","label":"meanPlanningTransport","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.meanPlanningTransport","description":"Terminal mean-planning transport: the product kernel is Markov and all sampled Bellman and backward values agree with the existing deterministic-mean model.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellman/index.html#decl-3bda84a2ed4e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","order":9200,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellman"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellman.lean:213"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem meanPlanningTransport (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) : ProbabilityTheory.IsMarkovKernel source.rewardNextStateKernel ∧ (forall (value : State -> Real), Measurable value -> forall state action, source.stochasticBellmanQ value state action = mdp.bellmanQ value state action) ∧ (forall (stage : Fin mdp.horizon) (value : State -> Real), Measurable value -> source.stochasticBellman policy stage value = policy.bellman stage value) ∧ (forall (remaining : Nat) (hremaining : remaining <= mdp.horizon), source.stochasticValueRemaining policy remaining hremaining = policy.valueRemaining remaining hremaining) ∧ (forall (stage : Nat) (hstage : stage <= mdp.horizon), source.stochasticValueAt policy stage hstage = policy.valueAt stage hstage)","missing":[],"search":"meanplanningtransport banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.meanplanningtransport terminal mean-planning transport: the product kernel is markov and all sampled bellman and backward values agree with the existing deterministic-mean model. theorem compiled","shard":"modules/30774244479759d1.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.intervalVarianceProxy_neg_coe","label":"intervalVarianceProxy_neg_coe","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.intervalVarianceProxy_neg_coe","description":"A symmetric interval with NNReal radius has variance proxy equal to its square.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-e95f6a6d927a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9201,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:20"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem intervalVarianceProxy_neg_coe (bound : NNReal) : intervalVarianceProxy (-(bound : Real)) (bound : Real) = bound ^ 2","missing":[],"search":"intervalvarianceproxy_neg_coe banditrlproof.concentration.intervalvarianceproxy_neg_coe a symmetric interval with nnreal radius has variance proxy equal to its square. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationStepVarianceProxy","label":"meanBellmanInnovationStepVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.meanBellmanInnovationStepVarianceProxy","description":"Hoeffding proxy for the mean Bellman return with `remaining` decisions.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-067cc9fac9a1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9202,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def meanBellmanInnovationStepVarianceProxy (rewardBound : NNReal) (remaining : Nat) : NNReal","missing":[],"search":"meanbellmaninnovationstepvarianceproxy banditrlproof.finitehorizonrl.meanbellmaninnovationstepvarianceproxy hoeffding proxy for the mean bellman return with `remaining` decisions. definition compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy","label":"meanBellmanInnovationVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy","description":"Sum of the stage-dependent mean Bellman innovation proxies.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-cc09b528a84e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9203,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def meanBellmanInnovationVarianceProxy (rewardBound : NNReal) (remaining : Nat) : NNReal","missing":[],"search":"meanbellmaninnovationvarianceproxy banditrlproof.finitehorizonrl.meanbellmaninnovationvarianceproxy sum of the stage-dependent mean bellman innovation proxies. definition compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_zero","label":"meanBellmanInnovationVarianceProxy_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_zero","description":"theorem meanBellmanInnovationVarianceProxy_zero (rewardBound : NNReal) : meanBellmanInnovationVarianceProxy rewardBound 0 = 0","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-d190b4ba54f1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9204,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem meanBellmanInnovationVarianceProxy_zero (rewardBound : NNReal) : meanBellmanInnovationVarianceProxy rewardBound 0 = 0","missing":[],"search":"meanbellmaninnovationvarianceproxy_zero banditrlproof.finitehorizonrl.meanbellmaninnovationvarianceproxy_zero theorem meanbellmaninnovationvarianceproxy_zero (rewardbound : nnreal) : meanbellmaninnovationvarianceproxy rewardbound 0 = 0 theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_succ","label":"meanBellmanInnovationVarianceProxy_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_succ","description":"theorem meanBellmanInnovationVarianceProxy_succ (rewardBound : NNReal) (remaining : Nat) : meanBellmanInnovationVarianceProxy rewardBound (remaining + 1) = meanBellmanInnovationStepVarianceProxy rewardBound (remaining + 1) + meanBellmanInnovationVarianceProxy rewardBound remaining","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-c58c96a39d37","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9205,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:49"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem meanBellmanInnovationVarianceProxy_succ (rewardBound : NNReal) (remaining : Nat) : meanBellmanInnovationVarianceProxy rewardBound (remaining + 1) = meanBellmanInnovationStepVarianceProxy rewardBound (remaining + 1) + meanBellmanInnovationVarianceProxy rewardBound remaining","missing":[],"search":"meanbellmaninnovationvarianceproxy_succ banditrlproof.finitehorizonrl.meanbellmaninnovationvarianceproxy_succ theorem meanbellmaninnovationvarianceproxy_succ (rewardbound : nnreal) (remaining : nat) : meanbellmaninnovationvarianceproxy rewardbound (remaining + 1) = meanbellmaninnovationstepvarianceproxy rewardbound (remaining + 1) + meanbellmaninnovationvarianceproxy rewardbound remaining theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_pos","label":"meanBellmanInnovationVarianceProxy_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_pos","description":"theorem meanBellmanInnovationVarianceProxy_pos {rewardBound : NNReal} {remaining : Nat} (hrewardBound : 0 < rewardBound) (hremaining : 0 < remaining) : 0 < meanBellmanInnovationVarianceProxy rewardBound remaining","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-55f7c0e67710","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9206,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem meanBellmanInnovationVarianceProxy_pos {rewardBound : NNReal} {remaining : Nat} (hrewardBound : 0 < rewardBound) (hremaining : 0 < remaining) : 0 < meanBellmanInnovationVarianceProxy rewardBound remaining","missing":[],"search":"meanbellmaninnovationvarianceproxy_pos banditrlproof.finitehorizonrl.meanbellmaninnovationvarianceproxy_pos theorem meanbellmaninnovationvarianceproxy_pos {rewardbound : nnreal} {remaining : nat} (hrewardbound : 0 < rewardbound) (hremaining : 0 < remaining) : 0 < meanbellmaninnovationvarianceproxy rewardbound remaining theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_abs_le_of_rewardBound","label":"valueRemaining_abs_le_of_rewardBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_abs_le_of_rewardBound","description":"Every bounded-mean-reward policy value lies in its remaining reward envelope.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-c9047a55e55c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9207,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:70"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem valueRemaining_abs_le_of_rewardBound {mdp : MDP State Action} (policy : MarkovPolicy mdp) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (remaining : Nat) (hremaining : remaining ≤ mdp.horizon) (state : State) : |policy.valueRemaining remaining hremaining state| ≤ (remaining : Real) * (rewardBound : Real)","missing":[],"search":"valueremaining_abs_le_of_rewardbound banditrlproof.finitehorizonrl.markovpolicy.valueremaining_abs_le_of_rewardbound every bounded-mean-reward policy value lies in its remaining reward envelope. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeMeanBellmanInnovationFrom","label":"sampledCumulativeMeanBellmanInnovationFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeMeanBellmanInnovationFrom","description":"Sum of policy-action and transition innovations in the mean Bellman return.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-83e92659f87e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9208,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:122"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeMeanBellmanInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) : (remaining : Nat) → remaining ≤ mdp.horizon → State → RewardStepTrace Action State remaining → Real | 0, _, _, _ => 0 | remaining + 1, hremaining, state, trace => (mdp.reward state (trace 0).1 + policy.valueRemaining remaining (by omega) (trace 0).2.2 - policy.valueRemaining (remaining + 1) hremaining state) + sampledCumulativeMeanBellmanInnovationFrom mdp policy remaining (by omega) (trace 0).2.2 (Fin.tail trace) /-- The recursive mean Bellman innovation is jointly measurable in start state and trace. -/ theorem measurable_sampledCumulativeMeanBellmanInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining ≤ mdp.horizon) : Measurable (fun p : State × RewardStepTrace Action State remaining => mdp.sampledCumulativeMeanBel…","missing":[],"search":"sampledcumulativemeanbellmaninnovationfrom banditrlproof.finitehorizonrl.mdp.sampledcumulativemeanbellmaninnovationfrom sum of policy-action and transition innovations in the mean bellman return. definition compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeMeanBellmanInnovationFrom","label":"measurable_sampledCumulativeMeanBellmanInnovationFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeMeanBellmanInnovationFrom","description":"The recursive mean Bellman innovation is jointly measurable in start state and trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-60b63720b076","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9209,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:135"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeMeanBellmanInnovationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining ≤ mdp.horizon) : Measurable (fun p : State × RewardStepTrace Action State remaining => mdp.sampledCumulativeMeanBellmanInnovationFrom policy remaining hremaining p.1 p.2)","missing":[],"search":"measurable_sampledcumulativemeanbellmaninnovationfrom banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativemeanbellmaninnovationfrom the recursive mean bellman innovation is jointly measurable in start state and trace. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_map_dropReward","label":"actionRewardStateKernel_map_dropReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_map_dropReward","description":"Dropping the sampled reward recovers the ordinary action/next-state kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-5fb7f021bf70","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9210,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:173"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardStateKernel_map_dropReward (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : (source.actionRewardStateKernel policy stage).map (fun head : Action × (Real × State) => (head.1, head.2.2)) = policy.actionStateKernel stage","missing":[],"search":"actionrewardstatekernel_map_dropreward banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardstatekernel_map_dropreward dropping the sampled reward recovers the ordinary action/next-state kernel. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_meanBellmanReturn_actionStateKernel_eq_valueRemaining","label":"integral_meanBellmanReturn_actionStateKernel_eq_valueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_meanBellmanReturn_actionStateKernel_eq_valueRemaining","description":"The one-step mean Bellman return integrates to the recursive policy value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-263e69057647","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9211,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:204"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_meanBellmanReturn_actionStateKernel_eq_valueRemaining (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 ≤ mdp.horizon) (state : State) : ∫ head : Action × State, (mdp.reward state head.1 + policy.valueRemaining remaining (by omega) head.2) ∂ policy.actionStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ state = policy.valueRemaining (remaining + 1) hremaining state","missing":[],"search":"integral_meanbellmanreturn_actionstatekernel_eq_valueremaining banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.integral_meanbellmanreturn_actionstatekernel_eq_valueremaining the one-step mean bellman return integrates to the recursive policy value. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","label":"actionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","description":"A one-step mean Bellman innovation is sub-Gaussian at its stage envelope.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-4b762c581384","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9212,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:244"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionStateKernel_meanBellmanInnovation_hasSubgaussianMGF (policy : MarkovPolicy mdp) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (remaining : Nat) (hremaining : remaining + 1 ≤ mdp.horizon) (state : State) : ProbabilityTheory.HasSubgaussianMGF (fun head : Action × State => mdp.reward state head.1 + policy.valueRemaining remaining (by omega) head.2 - policy.valueRemaining (remaining + 1) hremaining state) (meanBellmanInnovationStepVarianceProxy rewardBound (remaining + 1)) (policy.actionStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ state)","missing":[],"search":"actionstatekernel_meanbellmaninnovation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionstatekernel_meanbellmaninnovation_hassubgaussianmgf a one-step mean bellman innovation is sub-gaussian at its stage envelope. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_meanBellmanInnovation_hasSubgaussianMGF","label":"actionRewardStateKernel_meanBellmanInnovation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_meanBellmanInnovation_hasSubgaussianMGF","description":"The reward-bearing head has the same mean Bellman innovation MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-143b18d05204","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9213,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:296"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardStateKernel_meanBellmanInnovation_hasSubgaussianMGF (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (remaining : Nat) (hremaining : remaining + 1 ≤ mdp.horizon) (state : State) : ProbabilityTheory.HasSubgaussianMGF (fun head : Action × (Real × State) => mdp.reward state head.1 + policy.valueRemaining remaining (by omega) head.2.2 - policy.valueRemaining (remaining + 1) hremaining state) (meanBellmanInnovationStepVarianceProxy rewardBound (remaining + 1)) (source.actionRewardStateKernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state)","missing":[],"search":"actionrewardstatekernel_meanbellmaninnovation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardstatekernel_meanbellmaninnovation_hassubgaussianmgf the reward-bearing head has the same mean bellman innovation mgf. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_hasSubgaussianMGF","label":"stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_hasSubgaussianMGF","description":"The cumulative mean Bellman innovation has the sum of its stage proxies.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-90806b475913","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9214,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:361"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (remaining : Nat) (hremaining : remaining ≤ mdp.horizon) (state : State) : ProbabilityTheory.HasSubgaussianMGF (mdp.sampledCumulativeMeanBellmanInnovationFrom policy remaining hremaining state) (meanBellmanInnovationVarianceProxy rewardBound remaining) (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state)","missing":[],"search":"stochastictrajectorykernelremaining_sampledcumulativemeanbellmaninnovationfrom_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_sampledcumulativemeanbellmaninnovationfrom_hassubgaussianmgf the cumulative mean bellman innovation has the sum of its stage proxies. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_abs_tail_le","label":"stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_abs_tail_le","description":"Fixed-horizon two-sided tail for the cumulative mean Bellman innovation.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardbellmaninnovationconcentration/index.html#decl-c838fd1158fc","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","order":9215,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardBellmanInnovationConcentration.lean:594"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (rewardBound : NNReal) (hrewardBound : ∀ state action, |mdp.reward state action| ≤ (rewardBound : Real)) (remaining : Nat) (hremaining : remaining ≤ mdp.horizon) (state : State) (htotal : 0 < ((meanBellmanInnovationVarianceProxy rewardBound remaining : NNReal) : Real)) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta ≤ 1) : (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state) {trace | Concentration.subGaussianSumConfidenceRadius (meanBellmanInnovationVarianceProxy rewardBound remaining) delta ≤ |mdp.sampledCumulativeMeanBellmanInnovationFrom policy remaining hremaining state trace|} ≤ ENNReal.ofReal delta","missing":[],"search":"stochastictrajectorykernelremaining_sampledcumulativemeanbellmaninnovationfrom_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_sampledcumulativemeanbellmaninnovationfrom_abs_tail_le fixed-horizon two-sided tail for the cumulative mean bellman innovation. theorem compiled","shard":"modules/7e9a1796f28c42cb.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.headRewardMean","label":"headRewardMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.headRewardMean","description":"The selected mean reward at the first coordinate of a positive trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-2796aede074d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9216,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def headRewardMean (mdp : MDP State Action) (state : State) (remaining : Nat) : RewardStepTrace Action State (remaining + 1) -> Real","missing":[],"search":"headrewardmean banditrlproof.finitehorizonrl.mdp.headrewardmean the selected mean reward at the first coordinate of a positive trace. definition compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardMean","label":"measurable_headRewardMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardMean","description":"theorem measurable_headRewardMean (mdp : MDP State Action) (state : State) (remaining : Nat) : Measurable (mdp.headRewardMean state remaining)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-7b876a050403","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9217,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:37"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_headRewardMean (mdp : MDP State Action) (state : State) (remaining : Nat) : Measurable (mdp.headRewardMean state remaining)","missing":[],"search":"measurable_headrewardmean banditrlproof.finitehorizonrl.mdp.measurable_headrewardmean theorem measurable_headrewardmean (mdp : mdp state action) (state : state) (remaining : nat) : measurable (mdp.headrewardmean state remaining) theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardMean_comap_headAction","label":"measurable_headRewardMean_comap_headAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardMean_comap_headAction","description":"The selected head mean is measurable in the sigma-algebra generated by the head action.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-09433bd44c14","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9218,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_headRewardMean_comap_headAction (mdp : MDP State Action) (state : State) (remaining : Nat) : @Measurable (RewardStepTrace Action State (remaining + 1)) Real (MeasurableSpace.comap (RewardStepTrace.headAction (Action := Action) (State := State) remaining) inferInstance) inferInstance (mdp.headRewardMean state remaining)","missing":[],"search":"measurable_headrewardmean_comap_headaction banditrlproof.finitehorizonrl.mdp.measurable_headrewardmean_comap_headaction the selected head mean is measurable in the sigma-algebra generated by the head action. theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.headRewardDeviation","label":"headRewardDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.headRewardDeviation","description":"Actual first sampled reward centered by the mean of its selected action law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-a71725743070","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9219,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def headRewardDeviation (mdp : MDP State Action) (state : State) (remaining : Nat) : RewardStepTrace Action State (remaining + 1) -> Real","missing":[],"search":"headrewarddeviation banditrlproof.finitehorizonrl.mdp.headrewarddeviation actual first sampled reward centered by the mean of its selected action law. definition compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardDeviation","label":"measurable_headRewardDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardDeviation","description":"theorem measurable_headRewardDeviation (mdp : MDP State Action) (state : State) (remaining : Nat) : Measurable (mdp.headRewardDeviation state remaining)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-13d9ccbb49a3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9220,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_headRewardDeviation (mdp : MDP State Action) (state : State) (remaining : Nat) : Measurable (mdp.headRewardDeviation state remaining)","missing":[],"search":"measurable_headrewarddeviation banditrlproof.finitehorizonrl.mdp.measurable_headrewarddeviation theorem measurable_headrewarddeviation (mdp : mdp state action) (state : state) (remaining : nat) : measurable (mdp.headrewarddeviation state remaining) theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.UniformSubgaussianRewardLaw","label":"UniformSubgaussianRewardLaw","kind":"structure","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.UniformSubgaussianRewardLaw","description":"Every selected reward law is sub-Gaussian around its stored MDP mean with one common proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-c5f33b39f2d5","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9221,"meta":[["Kind","structure"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:78"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"structure UniformSubgaussianRewardLaw (source : MeanCompatibleRewardKernel mdp) (varianceProxy : NNReal) : Prop where","missing":[],"search":"uniformsubgaussianrewardlaw banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.uniformsubgaussianrewardlaw every selected reward law is sub-gaussian around its stored mdp mean with one common proxy. structure compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.uniformSubgaussianRewardLaw_of_mem_Icc","label":"uniformSubgaussianRewardLaw_of_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.uniformSubgaussianRewardLaw_of_mem_Icc","description":"A common selected-reward interval constructs a uniform Hoeffding proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-e5bb32d99424","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9222,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformSubgaussianRewardLaw_of_mem_Icc (source : MeanCompatibleRewardKernel mdp) (lo hi : Real) (hbound : forall state action, ∀ᵐ reward ∂ source.rewardKernel.kernel (state, action), reward ∈ Set.Icc lo hi) : source.UniformSubgaussianRewardLaw (Concentration.intervalVarianceProxy lo hi)","missing":[],"search":"uniformsubgaussianrewardlaw_of_mem_icc banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.uniformsubgaussianrewardlaw_of_mem_icc a common selected-reward interval constructs a uniform hoeffding proxy. theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_hasCondSubgaussianMGF","label":"stochasticTrajectoryKernelRemaining_headRewardDeviation_hasCondSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_hasCondSubgaussianMGF","description":"The first generated reward deviation is conditionally sub-Gaussian given its sampled action.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-b07da5d3d6d9","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9223,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_headRewardDeviation_hasCondSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasCondSubgaussianMGF (MeasurableSpace.comap (RewardStepTrace.headAction (Action := Action) (State := State) remaining) inferInstance) (RewardStepTrace.measurable_headAction remaining).comap_le (mdp.headRewardDeviation state remaining) varianceProxy (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state)","missing":[],"search":"stochastictrajectorykernelremaining_headrewarddeviation_hascondsubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_headrewarddeviation_hascondsubgaussianmgf the first generated reward deviation is conditionally sub-gaussian given its sampled action. theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_hasSubgaussianMGF","label":"stochasticTrajectoryKernelRemaining_headRewardDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_hasSubgaussianMGF","description":"The conditional head deviation law also gives its unconditional sub-Gaussian MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-33ae35d60e51","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9224,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:162"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_headRewardDeviation_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (mdp.headRewardDeviation state remaining) varianceProxy (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state)","missing":[],"search":"stochastictrajectorykernelremaining_headrewarddeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_headrewarddeviation_hassubgaussianmgf the conditional head deviation law also gives its unconditional sub-gaussian mgf. theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_abs_tail_le","label":"stochasticTrajectoryKernelRemaining_headRewardDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_abs_tail_le","description":"One-step two-sided delta tail for the generated head reward deviation.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconcentration/index.html#decl-6a2759d595d6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","order":9225,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConcentration.lean:200"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_headRewardDeviation_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (hvariance : 0 < (varianceProxy : Real)) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state) {trace | Concentration.subGaussianSumConfidenceRadius varianceProxy delta <= |mdp.headRewardDeviation state remaining trace|} <= ENNReal.ofReal delta","missing":[],"search":"stochastictrajectorykernelremaining_headrewarddeviation_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_headrewarddeviation_abs_tail_le one-step two-sided delta tail for the generated head reward deviation. theorem compiled","shard":"modules/b5c0fe62bbf0bf28.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.headAction","label":"headAction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.headAction","description":"The first sampled action of a positive reward-bearing trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-da78720e9c44","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9226,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def headAction (remaining : Nat) : RewardStepTrace Action State (remaining + 1) -> Action","missing":[],"search":"headaction banditrlproof.finitehorizonrl.rewardsteptrace.headaction the first sampled action of a positive reward-bearing trace. definition compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headAction","label":"measurable_headAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headAction","description":"theorem measurable_headAction (remaining : Nat) : Measurable (headAction (Action := Action) (State := State) remaining)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-59f17cd91209","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9227,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:35"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_headAction (remaining : Nat) : Measurable (headAction (Action := Action) (State := State) remaining)","missing":[],"search":"measurable_headaction banditrlproof.finitehorizonrl.rewardsteptrace.measurable_headaction theorem measurable_headaction (remaining : nat) : measurable (headaction (action := action) (state := state) remaining) theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.selectedRewardKernelAt","label":"selectedRewardKernelAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.selectedRewardKernelAt","description":"The reward kernel selected after freezing the current state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-ff6198fb1ed1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9228,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:46"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selectedRewardKernelAt (source : MeanCompatibleRewardKernel mdp) (state : State) : ProbabilityTheory.Kernel Action Real","missing":[],"search":"selectedrewardkernelat banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.selectedrewardkernelat the reward kernel selected after freezing the current state. definition compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_eq_compProd_selectedRewardKernelAt","label":"actionRewardKernel_eq_compProd_selectedRewardKernelAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_eq_compProd_selectedRewardKernelAt","description":"The one-step action/reward marginal is the policy law composed with the selected reward law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-c02ac5a727a3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9229,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:59"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardKernel_eq_compProd_selectedRewardKernelAt (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) : source.actionRewardKernel policy stage state = policy.actionKernel stage state ⊗ₘ source.selectedRewardKernelAt state","missing":[],"search":"actionrewardkernel_eq_compprod_selectedrewardkernelat banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardkernel_eq_compprod_selectedrewardkernelat the one-step action/reward marginal is the policy law composed with the selected reward law. theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_map_fst","label":"actionRewardKernel_map_fst","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_map_fst","description":"The action marginal of the one-step action/reward law is the policy action law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-c0aad3e0070d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9230,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:73"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardKernel_map_fst (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) : (source.actionRewardKernel policy stage state).map Prod.fst = policy.actionKernel stage state","missing":[],"search":"actionrewardkernel_map_fst banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardkernel_map_fst the action marginal of the one-step action/reward law is the policy action law. theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headAction","label":"stochasticTrajectoryKernelRemaining_map_headAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headAction","description":"Mapping a generated trace to its first action recovers the policy action law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-d185f2264b29","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9231,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:86"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_map_headAction (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state).map (RewardStepTrace.headAction (Action := Action) (State := State) remaining) = policy.actionKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ state","missing":[],"search":"stochastictrajectorykernelremaining_map_headaction banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_map_headaction mapping a generated trace to its first action recovers the policy action law. theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_condDistrib_reward_given_action","label":"actionRewardKernel_condDistrib_reward_given_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_condDistrib_reward_given_action","description":"Under the one-step joint law, reward conditioned on action is the selected reward kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-d5c4827bd443","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9232,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:115"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardKernel_condDistrib_reward_given_action (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) : ProbabilityTheory.condDistrib Prod.snd Prod.fst (source.actionRewardKernel policy stage state) =ᵐ[ (source.actionRewardKernel policy stage state).map Prod.fst] source.selectedRewardKernelAt state","missing":[],"search":"actionrewardkernel_conddistrib_reward_given_action banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardkernel_conddistrib_reward_given_action under the one-step joint law, reward conditioned on action is the selected reward kernel. theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_condDistrib_headReward_given_headAction","label":"stochasticTrajectoryKernelRemaining_condDistrib_headReward_given_headAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_condDistrib_headReward_given_headAction","description":"On a generated positive trace, reward conditioned on the sampled head action has its selected law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-4258a46329b8","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9233,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:136"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_condDistrib_headReward_given_headAction (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : ProbabilityTheory.condDistrib (RewardStepTrace.headReward (Action := Action) (State := State) remaining) (RewardStepTrace.headAction (Action := Action) (State := State) remaining) (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state) =ᵐ[ (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state).map (RewardStepTrace.headAction (Action := Action) (State := State) remaining)] source.selectedRewardKernelAt state","missing":[],"search":"stochastictrajectorykernelremaining_conddistrib_headreward_given_headaction banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_conddistrib_headreward_given_headaction on a generated positive trace, reward conditioned on the sampled head action has its selected law. theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_condExpKernel_map_headReward_given_headAction","label":"stochasticTrajectoryKernelRemaining_condExpKernel_map_headReward_given_headAction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_condExpKernel_map_headReward_given_headAction","description":"The generated head conditional law on `Real` is also the mapped conditional expectation kernel on the sigma-algebra generated by the sampled head action.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardconditionallaw/index.html#decl-5491d84565eb","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","order":9234,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardConditionalLaw.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_condExpKernel_map_headReward_given_headAction [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : ∀ᵐ trace ∂ (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state).trim (RewardStepTrace.measurable_headAction remaining).comap_le, Measure.map (RewardStepTrace.headReward (Action := Action) (State := State) remaining) (ProbabilityTheory.condExpKernel (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state) (MeasurableSpace.comap (RewardStepTrace.headAction (Action := Action) (State := State) remaining) inferInstance) trace) = source.selectedRewardKernelAt state (RewardStepTrace.headAction (Action := Action) (State := State…","missing":[],"search":"stochastictrajectorykernelremaining_condexpkernel_map_headreward_given_headaction banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_condexpkernel_map_headreward_given_headaction the generated head conditional law on `real` is also the mapped conditional expectation kernel on the sigma-algebra generated by the sampled head action. theorem compiled","shard":"modules/2520e3477eb2999f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.kernel_hasSubgaussianMGF_of_ae","label":"kernel_hasSubgaussianMGF_of_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.kernel_hasSubgaussianMGF_of_ae","description":"theorem kernel_hasSubgaussianMGF_of_ae {Alpha : Type u} {Omega : Type v} [MeasurableSpace Alpha] [MeasurableSpace Omega] (mu : Measure Alpha) [IsFiniteMeasure mu] (kernel : ProbabilityTheory.Kernel Alpha Omega) (X : Omega -> Real) (varianceProxy : NNReal) (hX : Measurable X) (hfiber : ∀ᵐ alpha ∂mu, ProbabilityTheory.HasSubgaussianMGF X varianceProxy (kernel alpha)) : ProbabilityTheory.Kernel.HasSubgaussianMGF X vari…","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-1fe1cb899fd1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9235,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:21"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem kernel_hasSubgaussianMGF_of_ae {Alpha : Type u} {Omega : Type v} [MeasurableSpace Alpha] [MeasurableSpace Omega] (mu : Measure Alpha) [IsFiniteMeasure mu] (kernel : ProbabilityTheory.Kernel Alpha Omega) (X : Omega -> Real) (varianceProxy : NNReal) (hX : Measurable X) (hfiber : ∀ᵐ alpha ∂mu, ProbabilityTheory.HasSubgaussianMGF X varianceProxy (kernel alpha)) : ProbabilityTheory.Kernel.HasSubgaussianMGF X varianceProxy kernel mu","missing":[],"search":"kernel_hassubgaussianmgf_of_ae banditrlproof.concentration.kernel_hassubgaussianmgf_of_ae theorem kernel_hassubgaussianmgf_of_ae {alpha : type u} {omega : type v} [measurablespace alpha] [measurablespace omega] (mu : measure alpha) [isfinitemeasure mu] (kernel : probabilitytheory.kernel alpha omega) (x : omega -> real) (varianceproxy : nnreal) (hx : measurable x) (hfiber : ∀ᵐ alpha ∂mu, probabilitytheory.hassubgaussianmgf x varianceproxy (kernel alpha)) : probabilitytheory.kernel.hassubgaussianmgf x varianceproxy kernel mu theorem compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_apply_of_kernel_dirac","label":"hasSubgaussianMGF_apply_of_kernel_dirac","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasSubgaussianMGF_apply_of_kernel_dirac","description":"theorem hasSubgaussianMGF_apply_of_kernel_dirac {Alpha : Type u} {Omega : Type v} [MeasurableSpace Alpha] [MeasurableSpace Omega] [MeasurableSingletonClass Alpha] (kernel : ProbabilityTheory.Kernel Alpha Omega) (alpha : Alpha) (X : Omega -> Real) (varianceProxy : NNReal) (hkernel : ProbabilityTheory.Kernel.HasSubgaussianMGF X varianceProxy kernel (Measure.dirac alpha)) : ProbabilityTheory.HasSubgaussianMGF X varianc…","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-8eea1e5e5a6c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9236,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:57"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_apply_of_kernel_dirac {Alpha : Type u} {Omega : Type v} [MeasurableSpace Alpha] [MeasurableSpace Omega] [MeasurableSingletonClass Alpha] (kernel : ProbabilityTheory.Kernel Alpha Omega) (alpha : Alpha) (X : Omega -> Real) (varianceProxy : NNReal) (hkernel : ProbabilityTheory.Kernel.HasSubgaussianMGF X varianceProxy kernel (Measure.dirac alpha)) : ProbabilityTheory.HasSubgaussianMGF X varianceProxy (kernel alpha)","missing":[],"search":"hassubgaussianmgf_apply_of_kernel_dirac banditrlproof.concentration.hassubgaussianmgf_apply_of_kernel_dirac theorem hassubgaussianmgf_apply_of_kernel_dirac {alpha : type u} {omega : type v} [measurablespace alpha] [measurablespace omega] [measurablesingletonclass alpha] (kernel : probabilitytheory.kernel alpha omega) (alpha : alpha) (x : omega -> real) (varianceproxy : nnreal) (hkernel : probabilitytheory.kernel.hassubgaussianmgf x varianceproxy kernel (measure.dirac alpha)) : probabilitytheory.hassubgaussianmgf x varianceproxy (kernel alpha) theorem compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardDeviationFrom","label":"sampledCumulativeRewardDeviationFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardDeviationFrom","description":"def sampledCumulativeRewardDeviationFrom (mdp : MDP State Action) : (remaining : Nat) -> State -> RewardStepTrace Action State remaining -> Real | 0, _, _ => 0 | remaining + 1, state, trace => (trace 0).2.1 - mdp.reward state (trace 0).1 + sampledCumulativeRewardDeviationFrom mdp remaining (trace 0).2.2 (Fin.tail trace) theorem measurable_sampledCumulativeRewardDeviationFrom (mdp : MDP State Action) (remaining : Nat…","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-c1762fc0aa59","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9237,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:88"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def sampledCumulativeRewardDeviationFrom (mdp : MDP State Action) : (remaining : Nat) -> State -> RewardStepTrace Action State remaining -> Real | 0, _, _ => 0 | remaining + 1, state, trace => (trace 0).2.1 - mdp.reward state (trace 0).1 + sampledCumulativeRewardDeviationFrom mdp remaining (trace 0).2.2 (Fin.tail trace) theorem measurable_sampledCumulativeRewardDeviationFrom (mdp : MDP State Action) (remaining : Nat) : Measurable (fun p : State × RewardStepTrace Action State remaining => mdp.sampledCumulativeRewardDeviationFrom remaining p.1 p.2)","missing":[],"search":"sampledcumulativerewarddeviationfrom banditrlproof.finitehorizonrl.mdp.sampledcumulativerewarddeviationfrom def sampledcumulativerewarddeviationfrom (mdp : mdp state action) : (remaining : nat) -> state -> rewardsteptrace action state remaining -> real | 0, _, _ => 0 | remaining + 1, state, trace => (trace 0).2.1 - mdp.reward state (trace 0).1 + sampledcumulativerewarddeviationfrom mdp remaining (trace 0).2.2 (fin.tail trace) theorem measurable_sampledcumulativerewarddeviationfrom (mdp : mdp state action) (remaining : nat) : measurable (fun p : state × rewardsteptrace action state remaining => mdp.sampledcumulativerewarddeviationfrom remaining p.1 p.2) definition compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardDeviationFrom","label":"measurable_sampledCumulativeRewardDeviationFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardDeviationFrom","description":"theorem measurable_sampledCumulativeRewardDeviationFrom (mdp : MDP State Action) (remaining : Nat) : Measurable (fun p : State × RewardStepTrace Action State remaining => mdp.sampledCumulativeRewardDeviationFrom remaining p.1 p.2)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-5dee7f745b51","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9238,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeRewardDeviationFrom (mdp : MDP State Action) (remaining : Nat) : Measurable (fun p : State × RewardStepTrace Action State remaining => mdp.sampledCumulativeRewardDeviationFrom remaining p.1 p.2)","missing":[],"search":"measurable_sampledcumulativerewarddeviationfrom banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativerewarddeviationfrom theorem measurable_sampledcumulativerewarddeviationfrom (mdp : mdp state action) (remaining : nat) : measurable (fun p : state × rewardsteptrace action state remaining => mdp.sampledcumulativerewarddeviationfrom remaining p.1 p.2) theorem compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_rewardDeviation_hasSubgaussianMGF","label":"actionRewardStateKernel_rewardDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_rewardDeviation_hasSubgaussianMGF","description":"theorem actionRewardStateKernel_rewardDeviation_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (fun head :…","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-e5c1a814db1d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9239,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:123"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardStateKernel_rewardDeviation_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (fun head : Action × (Real × State) => head.2.1 - mdp.reward state head.1) varianceProxy (source.actionRewardStateKernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state)","missing":[],"search":"actionrewardstatekernel_rewarddeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardstatekernel_rewarddeviation_hassubgaussianmgf theorem actionrewardstatekernel_rewarddeviation_hassubgaussianmgf [standardborelspace state] [standardborelspace action] [nonempty action] (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (remaining : nat) (hremaining : remaining + 1 <= mdp.horizon) (state : state) (varianceproxy : nnreal) (law : source.uniformsubgaussianrewardlaw varianceproxy) : probabilitytheory.hassubgaussianmgf (fun head : action × (real × state) => head.2.1 - mdp.reward state head.1) varianceproxy (source.actionrewardstatekernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state) theorem compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_hasSubgaussianMGF","label":"stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_hasSubgaussianMGF","description":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.H…","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-2f5368b1729d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9240,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (mdp.sampledCumulativeRewardDeviationFrom remaining state) ((remaining : NNReal) * varianceProxy) (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state)","missing":[],"search":"stochastictrajectorykernelremaining_sampledcumulativerewarddeviationfrom_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_sampledcumulativerewarddeviationfrom_hassubgaussianmgf theorem stochastictrajectorykernelremaining_sampledcumulativerewarddeviationfrom_hassubgaussianmgf [standardborelspace state] [standardborelspace action] [nonempty action] (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (remaining : nat) (hremaining : remaining <= mdp.horizon) (state : state) (varianceproxy : nnreal) (law : source.uniformsubgaussianrewardlaw varianceproxy) : probabilitytheory.hassubgaussianmgf (mdp.sampledcumulativerewarddeviationfrom remaining state) ((remaining : nnreal) * varianceproxy) (source.stochastictrajectorykernelremaining policy remaining hremaining state) theorem compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_abs_tail_le","label":"stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_abs_tail_le","description":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((remaining…","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardcumulativeconcentration/index.html#decl-b7e12a528f07","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","order":9241,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardCumulativeConcentration.lean:344"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((remaining : NNReal) * varianceProxy : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state) {trace | Concentration.subGaussianSumConfidenceRadius ((remaining : NNReal) * varianceProxy) delta <= |mdp.sampledCumulativeRewardDeviationFrom remaining state trace|} <= ENNReal.ofReal delta","missing":[],"search":"stochastictrajectorykernelremaining_sampledcumulativerewarddeviationfrom_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_sampledcumulativerewarddeviationfrom_abs_tail_le theorem stochastictrajectorykernelremaining_sampledcumulativerewarddeviationfrom_abs_tail_le [standardborelspace state] [standardborelspace action] [nonempty action] (source : meancompatiblerewardkernel mdp) (policy : markovpolicy mdp) (remaining : nat) (hremaining : remaining <= mdp.horizon) (state : state) (varianceproxy : nnreal) (law : source.uniformsubgaussianrewardlaw varianceproxy) (htotal : 0 < ((((remaining : nnreal) * varianceproxy : nnreal) : real))) (delta : real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.stochastictrajectorykernelremaining policy remaining hremaining state) {trace | concentration.subgaussiansumconfidenceradius ((remaining : nnreal) * varianceproxy) delta <= |mdp.sampledcumulativerewarddeviationfrom remaining state trace|} <= ennreal.ofreal delta theorem compiled","shard":"modules/da549a6056d78ebd.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.eraseRewards","label":"eraseRewards","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.eraseRewards","description":"Discard sampled rewards while retaining every action and next state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-356c1adf6242","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9242,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def eraseRewards (remaining : Nat) : RewardStepTrace Action State remaining -> StepTrace Action State remaining","missing":[],"search":"eraserewards banditrlproof.finitehorizonrl.rewardsteptrace.eraserewards discard sampled rewards while retaining every action and next state. definition compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_eraseRewards","label":"measurable_eraseRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_eraseRewards","description":"Coordinatewise reward erasure is measurable on the finite Pi-space.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-32f99fdd5934","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9243,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_eraseRewards (remaining : Nat) : Measurable (eraseRewards (Action := Action) (State := State) remaining)","missing":[],"search":"measurable_eraserewards banditrlproof.finitehorizonrl.rewardsteptrace.measurable_eraserewards coordinatewise reward erasure is measurable on the finite pi-space. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.eraseRewards_cons","label":"eraseRewards_cons","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.eraseRewards_cons","description":"Reward erasure commutes with prepending one generated coordinate.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-2bdb1cac3389","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9244,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:51"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem eraseRewards_cons (remaining : Nat) (head : Action × (Real × State)) (tail : RewardStepTrace Action State remaining) : eraseRewards (Action := Action) (State := State) (remaining + 1) (@Fin.cons remaining (fun _ => Action × (Real × State)) head tail) = @Fin.cons remaining (fun _ => Action × State) (head.1, head.2.2) (eraseRewards (Action := Action) (State := State) remaining tail)","missing":[],"search":"eraserewards_cons banditrlproof.finitehorizonrl.rewardsteptrace.eraserewards_cons reward erasure commutes with prepending one generated coordinate. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.ProbabilityTheory.compProd_map_prodMap_of_map_eq","label":"compProd_map_prodMap_of_map_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.ProbabilityTheory.compProd_map_prodMap_of_map_eq","description":"Map both outputs of a composition-product kernel when the mapped head kernel and every mapped tail fiber agree with prescribed target kernels.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-9ea3fd3e64ce","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9245,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:70"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem compProd_map_prodMap_of_map_eq {Alpha Beta Beta' Gamma Gamma' : Type*} [MeasurableSpace Alpha] [MeasurableSpace Beta] [MeasurableSpace Beta'] [MeasurableSpace Gamma] [MeasurableSpace Gamma'] (kappa : ProbabilityTheory.Kernel Alpha Beta) [ProbabilityTheory.IsMarkovKernel kappa] (eta : ProbabilityTheory.Kernel (Alpha × Beta) Gamma) [ProbabilityTheory.IsMarkovKernel eta] (kappa' : ProbabilityTheory.Kernel Alpha Beta') [ProbabilityTheory.IsMarkovKernel kappa'] (eta' : ProbabilityTheory.Kernel (Alpha × Beta') Gamma') [ProbabilityTheory.IsMarkovKernel eta'] (f : Beta -> Beta') (hf : Measurable f) (g : Gamma -> Gamma') (hg : Measurable g) (hkappa : kappa.map f = kappa') (heta : forall alpha beta, (eta (alpha, beta)).map g = eta' (alpha, f beta)) : (kappa.compProd eta).map (Prod.map f g) = kappa'.compProd eta'","missing":[],"search":"compprod_map_prodmap_of_map_eq banditrlproof.finitehorizonrl.probabilitytheory.compprod_map_prodmap_of_map_eq map both outputs of a composition-product kernel when the mapped head kernel and every mapped tail fiber agree with prescribed target kernels. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_eraseRewards","label":"stochasticTrajectoryKernelRemaining_map_eraseRewards","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_eraseRewards","description":"Dropping all sampled rewards from a generated finite stochastic trajectory recovers the ordinary action/next-state trajectory kernel exactly.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-933e3c3d6ad6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9246,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:130"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_map_eraseRewards (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : (source.stochasticTrajectoryKernelRemaining policy remaining hremaining).map (RewardStepTrace.eraseRewards (Action := Action) (State := State) remaining) = policy.trajectoryKernelRemaining remaining hremaining","missing":[],"search":"stochastictrajectorykernelremaining_map_eraserewards banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_map_eraserewards dropping all sampled rewards from a generated finite stochastic trajectory recovers the ordinary action/next-state trajectory kernel exactly. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.eraseTrajectory","label":"eraseTrajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.eraseTrajectory","description":"Erase sampled rewards from a full stochastic trajectory, retaining its initial state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-04ffe4bc3c9c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9247,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:235"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def eraseTrajectory (trajectory : State × RewardStepTrace Action State mdp.horizon) : State × StepTrace Action State mdp.horizon","missing":[],"search":"erasetrajectory banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.erasetrajectory erase sampled rewards from a full stochastic trajectory, retaining its initial state. definition compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_eraseTrajectory","label":"measurable_eraseTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_eraseTrajectory","description":"Full-trajectory reward erasure is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-faa2835fdf9a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9248,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:241"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_eraseTrajectory : Measurable (eraseTrajectory (mdp := mdp) (Action := Action))","missing":[],"search":"measurable_erasetrajectory banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_erasetrajectory full-trajectory reward erasure is measurable. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_map_eraseTrajectory","label":"stochasticTrajectoryMeasure_map_eraseTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_map_eraseTrajectory","description":"The full stochastic trajectory law maps exactly to the ordinary trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-16deda35c5f2","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9249,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:247"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryMeasure_map_eraseTrajectory (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : (source.stochasticTrajectoryMeasure policy initialState).map (eraseTrajectory (mdp := mdp) (Action := Action)) = policy.trajectoryMeasure initialState","missing":[],"search":"stochastictrajectorymeasure_map_erasetrajectory banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure_map_erasetrajectory the full stochastic trajectory law maps exactly to the ordinary trajectory law. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.eraseTrajectoryFamily","label":"eraseTrajectoryFamily","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.eraseTrajectoryFamily","description":"Erase sampled rewards coordinatewise from a finite iid trajectory family.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-4680d0645f5d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9250,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:267"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def eraseTrajectoryFamily (episodes : Nat) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : Fin episodes -> State × StepTrace Action State mdp.horizon","missing":[],"search":"erasetrajectoryfamily banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.erasetrajectoryfamily erase sampled rewards coordinatewise from a finite iid trajectory family. definition compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_eraseTrajectoryFamily","label":"measurable_eraseTrajectoryFamily","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_eraseTrajectoryFamily","description":"Finite-family reward erasure is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-4b6b11ea834e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9251,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:274"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_eraseTrajectoryFamily (episodes : Nat) : Measurable (eraseTrajectoryFamily (mdp := mdp) (Action := Action) episodes)","missing":[],"search":"measurable_erasetrajectoryfamily banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_erasetrajectoryfamily finite-family reward erasure is measurable. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_eraseTrajectoryFamily","label":"iidStochasticTrajectoryFamilyMeasure_map_eraseTrajectoryFamily","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_eraseTrajectoryFamily","description":"Finite iid stochastic trajectories map to the deterministic iid family law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-52355593f48d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9252,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_map_eraseTrajectoryFamily (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes).map (eraseTrajectoryFamily (mdp := mdp) (Action := Action) episodes) = policy.iidTrajectoryFamilyMeasure initialState episodes","missing":[],"search":"iidstochastictrajectoryfamilymeasure_map_erasetrajectoryfamily banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_map_erasetrajectoryfamily finite iid stochastic trajectories map to the deterministic iid family law. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchOfStochasticTrajectories","label":"knownRewardEpisodeBatchOfStochasticTrajectories","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchOfStochasticTrajectories","description":"Convert a stochastic trajectory family to the existing empirical batch after erasing rewards. Batch rewards are the known means `mdp.reward`.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-0867db264f3d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9253,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:312"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def knownRewardEpisodeBatchOfStochasticTrajectories (episodes : Nat) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : EpisodeBatch mdp episodes","missing":[],"search":"knownrewardepisodebatchofstochastictrajectories banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.knownrewardepisodebatchofstochastictrajectories convert a stochastic trajectory family to the existing empirical batch after erasing rewards. batch rewards are the known means `mdp.reward`. definition compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchOfStochasticTrajectories","label":"measurable_knownRewardEpisodeBatchOfStochasticTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchOfStochasticTrajectories","description":"The known-reward stochastic-family batch projection is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-31664bbd6828","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9254,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:322"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_knownRewardEpisodeBatchOfStochasticTrajectories (episodes : Nat) : Measurable (knownRewardEpisodeBatchOfStochasticTrajectories (mdp := mdp) episodes)","missing":[],"search":"measurable_knownrewardepisodebatchofstochastictrajectories banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_knownrewardepisodebatchofstochastictrajectories the known-reward stochastic-family batch projection is measurable. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_knownRewardEpisodeBatch_eq_iidEpisodeBatchMeasure","label":"iidStochasticTrajectoryFamilyMeasure_map_knownRewardEpisodeBatch_eq_iidEpisodeBatchMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_knownRewardEpisodeBatch_eq_iidEpisodeBatchMeasure","description":"The known-reward batch extracted from iid stochastic trajectories has exactly the existing deterministic iid episode-batch law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewarderasurelaw/index.html#decl-ba187a5779bf","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","order":9255,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardErasureLaw.lean:335"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_map_knownRewardEpisodeBatch_eq_iidEpisodeBatchMeasure (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes).map (knownRewardEpisodeBatchOfStochasticTrajectories (mdp := mdp) episodes) = policy.iidEpisodeBatchMeasure initialState episodes","missing":[],"search":"iidstochastictrajectoryfamilymeasure_map_knownrewardepisodebatch_eq_iidepisodebatchmeasure banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_map_knownrewardepisodebatch_eq_iidepisodebatchmeasure the known-reward batch extracted from iid stochastic trajectories has exactly the existing deterministic iid episode-batch law. theorem compiled","shard":"modules/92f38ad159165af3.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.rewardStepTrace_stateAt_eq_trajectoryStateAt_eraseTrajectory","label":"rewardStepTrace_stateAt_eq_trajectoryStateAt_eraseTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.rewardStepTrace_stateAt_eq_trajectoryStateAt_eraseTrajectory","description":"Reward erasure preserves the state immediately preceding every coordinate.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-663ca5ef774e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9256,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:30"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem rewardStepTrace_stateAt_eq_trajectoryStateAt_eraseTrajectory (mdp : MDP State Action) (trajectory : State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : RewardStepTrace.stateAt trajectory.1 trajectory.2 stage = mdp.trajectoryStateAt (MeanCompatibleRewardKernel.eraseTrajectory (mdp := mdp) trajectory) stage","missing":[],"search":"rewardsteptrace_stateat_eq_trajectorystateat_erasetrajectory banditrlproof.finitehorizonrl.mdp.rewardsteptrace_stateat_eq_trajectorystateat_erasetrajectory reward erasure preserves the state immediately preceding every coordinate. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_state_eq_knownRewardEpisodeStep","label":"sampledEpisodeStep_state_eq_knownRewardEpisodeStep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_state_eq_knownRewardEpisodeStep","description":"Sampled and known-reward projections have identical stage states.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-ac6557544d74","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9257,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledEpisodeStep_state_eq_knownRewardEpisodeStep (mdp : MDP State Action) (trajectory : State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : (mdp.sampledEpisodeStepOfStochasticTrajectory trajectory stage).state = (mdp.episodeStepOfTrajectory (MeanCompatibleRewardKernel.eraseTrajectory (mdp := mdp) trajectory) stage).state","missing":[],"search":"sampledepisodestep_state_eq_knownrewardepisodestep banditrlproof.finitehorizonrl.mdp.sampledepisodestep_state_eq_knownrewardepisodestep sampled and known-reward projections have identical stage states. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_action_eq_knownRewardEpisodeStep","label":"sampledEpisodeStep_action_eq_knownRewardEpisodeStep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_action_eq_knownRewardEpisodeStep","description":"Sampled and known-reward projections have identical stage actions.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-0a757967883e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9258,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:56"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledEpisodeStep_action_eq_knownRewardEpisodeStep (mdp : MDP State Action) (trajectory : State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : (mdp.sampledEpisodeStepOfStochasticTrajectory trajectory stage).action = (mdp.episodeStepOfTrajectory (MeanCompatibleRewardKernel.eraseTrajectory (mdp := mdp) trajectory) stage).action","missing":[],"search":"sampledepisodestep_action_eq_knownrewardepisodestep banditrlproof.finitehorizonrl.mdp.sampledepisodestep_action_eq_knownrewardepisodestep sampled and known-reward projections have identical stage actions. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_nextState_eq_knownRewardEpisodeStep","label":"sampledEpisodeStep_nextState_eq_knownRewardEpisodeStep","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_nextState_eq_knownRewardEpisodeStep","description":"Sampled and known-reward projections have identical next states.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-2c6c48d69904","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9259,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:68"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledEpisodeStep_nextState_eq_knownRewardEpisodeStep (mdp : MDP State Action) (trajectory : State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : (mdp.sampledEpisodeStepOfStochasticTrajectory trajectory stage).nextState = (mdp.episodeStepOfTrajectory (MeanCompatibleRewardKernel.eraseTrajectory (mdp := mdp) trajectory) stage).nextState","missing":[],"search":"sampledepisodestep_nextstate_eq_knownrewardepisodestep banditrlproof.finitehorizonrl.mdp.sampledepisodestep_nextstate_eq_knownrewardepisodestep sampled and known-reward projections have identical next states. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatch_visitCount_eq_knownRewardEpisodeBatch","label":"sampledEpisodeBatch_visitCount_eq_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatch_visitCount_eq_knownRewardEpisodeBatch","description":"Actual sampled-reward and known-reward projections have the same visit counts.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-49bca2426810","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9260,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledEpisodeBatch_visitCount_eq_knownRewardEpisodeBatch (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) : (mdp.sampledEpisodeBatchOfStochasticTrajectories episodes trajectories).visitCount stage state action = (MeanCompatibleRewardKernel.knownRewardEpisodeBatchOfStochasticTrajectories (mdp := mdp) episodes trajectories).visitCount stage state action","missing":[],"search":"sampledepisodebatch_visitcount_eq_knownrewardepisodebatch banditrlproof.finitehorizonrl.mdp.sampledepisodebatch_visitcount_eq_knownrewardepisodebatch actual sampled-reward and known-reward projections have the same visit counts. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatch_transitionCount_eq_knownRewardEpisodeBatch","label":"sampledEpisodeBatch_transitionCount_eq_knownRewardEpisodeBatch","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatch_transitionCount_eq_knownRewardEpisodeBatch","description":"Actual sampled-reward and known-reward projections have the same transition counts.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-12229fcb17f6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9261,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledEpisodeBatch_transitionCount_eq_knownRewardEpisodeBatch (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) (nextState : State) : (mdp.sampledEpisodeBatchOfStochasticTrajectories episodes trajectories).transitionCount stage state action nextState = (MeanCompatibleRewardKernel.knownRewardEpisodeBatchOfStochasticTrajectories (mdp := mdp) episodes trajectories).transitionCount stage state action nextState","missing":[],"search":"sampledepisodebatch_transitioncount_eq_knownrewardepisodebatch banditrlproof.finitehorizonrl.mdp.sampledepisodebatch_transitioncount_eq_knownrewardepisodebatch actual sampled-reward and known-reward projections have the same transition counts. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticSampledBatchCountBadEvent","label":"stochasticSampledBatchCountBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticSampledBatchCountBadEvent","description":"Count event pulled back along the actual sampled-reward batch projection. Its set is source-independent, while the receiver aligns it with the source's iid stochastic trajectory measure and reward event used by the combined route.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-f1093059fbeb","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9262,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:161"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticSampledBatchCountBadEvent (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta : Real) : Set (Fin episodes -> State × RewardStepTrace Action State mdp.horizon)","missing":[],"search":"stochasticsampledbatchcountbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticsampledbatchcountbadevent count event pulled back along the actual sampled-reward batch projection. its set is source-independent, while the receiver aligns it with the source's iid stochastic trajectory measure and reward event used by the combined route. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.mem_stochasticSampledBatchCountBadEvent_iff_knownReward","label":"mem_stochasticSampledBatchCountBadEvent_iff_knownReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.mem_stochasticSampledBatchCountBadEvent_iff_knownReward","description":"Membership in the sampled and known-reward pullbacks of the count event agrees.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-9be345b81763","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9263,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:171"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem mem_stochasticSampledBatchCountBadEvent_iff_knownReward (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta : Real) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : trajectories ∈ source.stochasticSampledBatchCountBadEvent policy initialState episodes delta ↔ MeanCompatibleRewardKernel.knownRewardEpisodeBatchOfStochasticTrajectories (mdp := mdp) episodes trajectories ∈ policy.simultaneousCountBadEvent initialState episodes delta","missing":[],"search":"mem_stochasticsampledbatchcountbadevent_iff_knownreward banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.mem_stochasticsampledbatchcountbadevent_iff_knownreward membership in the sampled and known-reward pullbacks of the count event agrees. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_stochasticSampledBatchCountBadEvent","label":"measurableSet_stochasticSampledBatchCountBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_stochasticSampledBatchCountBadEvent","description":"The pulled-back sampled-batch count event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-9ca9a86d024f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9264,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:211"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_stochasticSampledBatchCountBadEvent (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (delta : Real) : MeasurableSet (source.stochasticSampledBatchCountBadEvent policy initialState episodes delta)","missing":[],"search":"measurableset_stochasticsampledbatchcountbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurableset_stochasticsampledbatchcountbadevent the pulled-back sampled-batch count event is measurable. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_stochasticSampledBatchCountBadEvent_le","label":"iidStochasticTrajectoryFamilyMeasure_stochasticSampledBatchCountBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_stochasticSampledBatchCountBadEvent_le","description":"The actual sampled-batch count event inherits the compiled count failure share.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-71211bd87c89","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9265,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:223"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_stochasticSampledBatchCountBadEvent_le (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes) (source.stochasticSampledBatchCountBadEvent policy initialState episodes delta) <= ENNReal.ofReal delta","missing":[],"search":"iidstochastictrajectoryfamilymeasure_stochasticsampledbatchcountbadevent_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_stochasticsampledbatchcountbadevent_le the actual sampled-batch count event inherits the compiled count failure share. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta","label":"simultaneousRewardDelta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta","description":"Equal reward confidence share for every stage/state/action coordinate.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-d22ebb556910","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9266,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:259"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousRewardDelta (mdp : MDP State Action) (delta : Real) : Real","missing":[],"search":"simultaneousrewarddelta banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.simultaneousrewarddelta equal reward confidence share for every stage/state/action coordinate. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardSumConfidenceRadius","label":"simultaneousRewardSumConfidenceRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardSumConfidenceRadius","description":"Deterministic fixed-coordinate reward-sum confidence radius.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-a48d34f888dd","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9267,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:264"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousRewardSumConfidenceRadius (mdp : MDP State Action) (episodes : Nat) (varianceProxy : NNReal) (delta : Real) : Real","missing":[],"search":"simultaneousrewardsumconfidenceradius banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.simultaneousrewardsumconfidenceradius deterministic fixed-coordinate reward-sum confidence radius. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardCoordinateBadEvent","label":"rewardCoordinateBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardCoordinateBadEvent","description":"One visit coordinate's sampled-reward deviation event.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-01d9fd683065","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9268,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:272"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardCoordinateBadEvent (source : MeanCompatibleRewardKernel mdp) (episodes : Nat) (varianceProxy : NNReal) (delta : Real) (coordinate : VisitCoordinate mdp) : Set (Fin episodes -> State × RewardStepTrace Action State mdp.horizon)","missing":[],"search":"rewardcoordinatebadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewardcoordinatebadevent one visit coordinate's sampled-reward deviation event. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardBadEvent","label":"simultaneousRewardBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardBadEvent","description":"Union of every stage/state/action sampled-reward deviation event.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-8e9a7e44402e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9269,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:284"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def simultaneousRewardBadEvent (source : MeanCompatibleRewardKernel mdp) (episodes : Nat) (varianceProxy : NNReal) (delta : Real) : Set (Fin episodes -> State × RewardStepTrace Action State mdp.horizon)","missing":[],"search":"simultaneousrewardbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.simultaneousrewardbadevent union of every stage/state/action sampled-reward deviation event. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_rewardCoordinateBadEvent","label":"measurableSet_rewardCoordinateBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_rewardCoordinateBadEvent","description":"Every fixed reward-coordinate event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-65bdb9a4b1a5","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9270,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:292"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_rewardCoordinateBadEvent (source : MeanCompatibleRewardKernel mdp) (episodes : Nat) (varianceProxy : NNReal) (delta : Real) (coordinate : VisitCoordinate mdp) : MeasurableSet (source.rewardCoordinateBadEvent episodes varianceProxy delta coordinate)","missing":[],"search":"measurableset_rewardcoordinatebadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurableset_rewardcoordinatebadevent every fixed reward-coordinate event is measurable. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_simultaneousRewardBadEvent","label":"measurableSet_simultaneousRewardBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_simultaneousRewardBadEvent","description":"The all-coordinate reward union is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-b1c139e22b59","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9271,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:309"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_simultaneousRewardBadEvent (source : MeanCompatibleRewardKernel mdp) (episodes : Nat) (varianceProxy : NNReal) (delta : Real) : MeasurableSet (source.simultaneousRewardBadEvent episodes varianceProxy delta)","missing":[],"search":"measurableset_simultaneousrewardbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurableset_simultaneousrewardbadevent the all-coordinate reward union is measurable. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta_pos","label":"simultaneousRewardDelta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta_pos","description":"A nonempty reward-coordinate family receives a positive equal share.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-2dccc4296c7f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9272,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousRewardDelta_pos (mdp : MDP State Action) (hcoordinate : Nonempty (VisitCoordinate mdp)) {delta : Real} (hdelta : 0 < delta) : 0 < simultaneousRewardDelta mdp delta","missing":[],"search":"simultaneousrewarddelta_pos banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.simultaneousrewarddelta_pos a nonempty reward-coordinate family receives a positive equal share. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta_le_one","label":"simultaneousRewardDelta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta_le_one","description":"A global reward share at most one gives each coordinate a share at most one.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-d26ccc42ea83","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9273,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:329"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem simultaneousRewardDelta_le_one (mdp : MDP State Action) (hcoordinate : Nonempty (VisitCoordinate mdp)) {delta : Real} (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : simultaneousRewardDelta mdp delta <= 1","missing":[],"search":"simultaneousrewarddelta_le_one banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.simultaneousrewarddelta_le_one a global reward share at most one gives each coordinate a share at most one. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_simultaneousRewardBadEvent_le","label":"iidStochasticTrajectoryFamilyMeasure_simultaneousRewardBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_simultaneousRewardBadEvent_le","description":"The finite all-coordinate sampled-reward union consumes only its reward share.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-e7e1110723b2","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9274,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:340"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_simultaneousRewardBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes) (source.simultaneousRewardBadEvent episodes varianceProxy delta) <= ENNReal.ofReal delta","missing":[],"search":"iidstochastictrajectoryfamilymeasure_simultaneousrewardbadevent_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_simultaneousrewardbadevent_le the finite all-coordinate sampled-reward union consumes only its reward share. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviation_sum_abs_lt_of_not_mem_simultaneousRewardBadEvent","label":"maskedRewardDeviation_sum_abs_lt_of_not_mem_simultaneousRewardBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviation_sum_abs_lt_of_not_mem_simultaneousRewardBadEvent","description":"Outside the reward union, every masked reward sum is strictly inside its radius.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-6f2fb4ad7fb2","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9275,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:390"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem maskedRewardDeviation_sum_abs_lt_of_not_mem_simultaneousRewardBadEvent (source : MeanCompatibleRewardKernel mdp) {episodes : Nat} {varianceProxy : NNReal} {delta : Real} (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) (htrajectories : trajectories ∉ source.simultaneousRewardBadEvent episodes varianceProxy delta) (coordinate : VisitCoordinate mdp) : |∑ episode : Fin episodes, source.maskedRewardDeviationAtEpisode coordinate.stage coordinate.state coordinate.action episode trajectories| < simultaneousRewardSumConfidenceRadius mdp episodes varianceProxy delta","missing":[],"search":"maskedrewarddeviation_sum_abs_lt_of_not_mem_simultaneousrewardbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.maskedrewarddeviation_sum_abs_lt_of_not_mem_simultaneousrewardbadevent outside the reward union, every masked reward sum is strictly inside its radius. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviation_sum_eq_sampledBatch_rewardSum_sub","label":"maskedRewardDeviation_sum_eq_sampledBatch_rewardSum_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviation_sum_eq_sampledBatch_rewardSum_sub","description":"The masked iid reward sum is exactly reward sum minus visit count times mean.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-28f5b082bf95","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9276,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:408"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem maskedRewardDeviation_sum_eq_sampledBatch_rewardSum_sub (source : MeanCompatibleRewardKernel mdp) {episodes : Nat} (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) (state : State) (action : Action) : (∑ episode : Fin episodes, source.maskedRewardDeviationAtEpisode stage state action episode trajectories) = (mdp.sampledEpisodeBatchOfStochasticTrajectories episodes trajectories).rewardSum stage state action - ((mdp.sampledEpisodeBatchOfStochasticTrajectories episodes trajectories).visitCount stage state action : Real) * mdp.reward state action","missing":[],"search":"maskedrewarddeviation_sum_eq_sampledbatch_rewardsum_sub banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.maskedrewarddeviation_sum_eq_sampledbatch_rewardsum_sub the masked iid reward sum is exactly reward sum minus visit count times mean. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.expectedCountRewardCoordinateRadius","label":"expectedCountRewardCoordinateRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.expectedCountRewardCoordinateRadius","description":"Deterministic reward-mean radius based on the genuine lower count margin. Its formula is source-independent; the receiver keeps it adjacent to the source-indexed reward event and empirical-reward consumer.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-fcdc767bbad6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9277,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:443"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def expectedCountRewardCoordinateRadius (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) (coordinate : VisitCoordinate mdp) : Real","missing":[],"search":"expectedcountrewardcoordinateradius banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.expectedcountrewardcoordinateradius deterministic reward-mean radius based on the genuine lower count margin. its formula is source-independent; the receiver keeps it adjacent to the source-indexed reward event and empirical-reward consumer. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledBatch_empiricalReward_abs_sub_le_expectedCountRewardCoordinateRadius","label":"sampledBatch_empiricalReward_abs_sub_le_expectedCountRewardCoordinateRadius","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledBatch_empiricalReward_abs_sub_le_expectedCountRewardCoordinateRadius","description":"Outside both events, a sampled empirical reward obeys its deterministic radius.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-3b641098b01e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9278,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:454"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledBatch_empiricalReward_abs_sub_le_expectedCountRewardCoordinateRadius (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {varianceProxy : NNReal} {countDelta rewardDelta : Real} (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) (hcount : trajectories ∉ source.stochasticSampledBatchCountBadEvent policy initialState episodes countDelta) (hreward : trajectories ∉ source.simultaneousRewardBadEvent episodes varianceProxy rewardDelta) (coordinate : VisitCoordinate mdp) (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < coordinate.expectedCount policy initialState episodes) : let batch := mdp.sampledEpisodeBatchOfStochasticTrajectories episodes trajectories |batch.empiricalReward coordinate.stage coordinate.state coordinate.act…","missing":[],"search":"sampledbatch_empiricalreward_abs_sub_le_expectedcountrewardcoordinateradius banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.sampledbatch_empiricalreward_abs_sub_le_expectedcountrewardcoordinateradius outside both events, a sampled empirical reward obeys its deterministic radius. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.stochasticEmpiricalFiniteBatchValueEnvelope","label":"stochasticEmpiricalFiniteBatchValueEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.stochasticEmpiricalFiniteBatchValueEnvelope","description":"Linear value envelope with one empirical-reward error and one reward bonus.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-13bd76798c3b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9279,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:539"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def stochasticEmpiricalFiniteBatchValueEnvelope (rewardBound rewardBudget transitionBudget : Real) (remaining : Nat) : Real","missing":[],"search":"stochasticempiricalfinitebatchvalueenvelope banditrlproof.finitehorizonrl.stochasticempiricalfinitebatchvalueenvelope linear value envelope with one empirical-reward error and one reward bonus. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel","label":"stochasticAllCoordinateEmpiricalFiniteBatchModel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel","description":"Canonical empirical model retaining sampled rewards and fixed reward/transition budgets.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-ba330f6e3532","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9280,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:548"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticAllCoordinateEmpiricalFiniteBatchModel (mdp : MDP State Action) (episodes : Nat) (batch : EpisodeBatch mdp episodes) (defaultState : State) (rewardBudget transitionBudget : Real) : FiniteBatchModel mdp episodes where","missing":[],"search":"stochasticallcoordinateempiricalfinitebatchmodel banditrlproof.finitehorizonrl.mdp.stochasticallcoordinateempiricalfinitebatchmodel canonical empirical model retaining sampled rewards and fixed reward/transition budgets. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.StochasticAllCoordinateConfidence.upperValueRemaining_abs_le","label":"upperValueRemaining_abs_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.StochasticAllCoordinateConfidence.upperValueRemaining_abs_le","description":"Reward error and fixed budgets give a noncircular linear optimistic-value envelope.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-38b46b248def","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9281,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:565"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem upperValueRemaining_abs_le (hrewardError : forall stage state action, |batch.empiricalReward stage state action - mdp.reward state action| <= rewardBudget) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hrewardBudget_nonneg : 0 <= rewardBudget) (htransitionBudget_nonneg : 0 <= transitionBudget) : forall (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State), |(mdp.stochasticAllCoordinateEmpiricalFiniteBatchModel episodes batch defaultState rewardBudget transitionBudget).plan.upperValueRemaining remaining hremaining state| <= stochasticEmpiricalFiniteBatchValueEnvelope rewardBound rewardBudget transitionBudget remaining","missing":[],"search":"uppervalueremaining_abs_le banditrlproof.finitehorizonrl.mdp.stochasticallcoordinateconfidence.uppervalueremaining_abs_le reward error and fixed budgets give a noncircular linear optimistic-value envelope. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticAllCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","label":"stochasticAllCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticAllCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","description":"Pathwise stochastic sampled-batch producer for the complete confidence object.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-185c7c9d9385","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9282,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:680"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticAllCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem {mdp : MDP State Action} (source : mdp.MeanCompatibleRewardKernel) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} {varianceProxy : NNReal} {countDelta rewardDelta : Real} (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) (hcount : trajectories ∉ source.stochasticSampledBatchCountBadEvent policy initialState episodes countDelta) (hreward : trajectories ∉ source.simultaneousRewardBadEvent episodes varianceProxy rewardDelta) (defaultState : State) (rewardBound rewardBudget transitionBudget : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hrewardBudget_nonneg : 0 <= rewardBudget) (htransitionBudget_nonneg : 0 <= transitionBudget) (hmargin : forall coordinate : VisitCoord…","missing":[],"search":"stochasticallcoordinateempiricalfinitebatchmodelconfidence_of_not_mem banditrlproof.finitehorizonrl.markovpolicy.stochasticallcoordinateempiricalfinitebatchmodelconfidence_of_not_mem pathwise stochastic sampled-batch producer for the complete confidence object. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticAllCoordinateEmpiricalModelBadEvent","label":"stochasticAllCoordinateEmpiricalModelBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticAllCoordinateEmpiricalModelBadEvent","description":"The one bad event used by the stochastic empirical-model route: the pulled-back count event or one of the sampled-reward coordinate events.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-0dac2ff6ca16","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9283,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:787"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticAllCoordinateEmpiricalModelBadEvent (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) : Set (Fin episodes -> State × RewardStepTrace Action State mdp.horizon)","missing":[],"search":"stochasticallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochasticallcoordinateempiricalmodelbadevent the one bad event used by the stochastic empirical-model route: the pulled-back count event or one of the sampled-reward coordinate events. definition compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_stochasticAllCoordinateEmpiricalModelBadEvent","label":"measurableSet_stochasticAllCoordinateEmpiricalModelBadEvent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_stochasticAllCoordinateEmpiricalModelBadEvent","description":"The combined count-and-reward empirical-model event is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-9062bed77cbb","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9284,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:799"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurableSet_stochasticAllCoordinateEmpiricalModelBadEvent (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta : Real) : MeasurableSet (source.stochasticAllCoordinateEmpiricalModelBadEvent policy initialState episodes varianceProxy countDelta rewardDelta)","missing":[],"search":"measurableset_stochasticallcoordinateempiricalmodelbadevent banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurableset_stochasticallcoordinateempiricalmodelbadevent the combined count-and-reward empirical-model event is measurable. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_stochasticAllCoordinateEmpiricalModelBadEvent_le","label":"iidStochasticTrajectoryFamilyMeasure_stochasticAllCoordinateEmpiricalModelBadEvent_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_stochasticAllCoordinateEmpiricalModelBadEvent_le","description":"The two separately calibrated failure shares add under the combined event.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-ff700953475b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9285,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:814"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_stochasticAllCoordinateEmpiricalModelBadEvent_le [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes) (source.stochasticAllCoordinateEmpiricalModelBadEvent policy initialState episodes varianceProxy countDelta rewardDelta) <= ENNReal.ofReal countDelta +…","missing":[],"search":"iidstochastictrajectoryfamilymeasure_stochasticallcoordinateempiricalmodelbadevent_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_stochasticallcoordinateempiricalmodelbadevent_le the two separately calibrated failure shares add under the combined event. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_finiteBatchModel_confidence","label":"iidStochasticTrajectoryFamilyMeasure_allCoordinate_finiteBatchModel_confidence","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_finiteBatchModel_confidence","description":"The iid stochastic-reward trajectory law produces a measurable all-coordinate confidence event for the empirical model built from the actual sampled rewards.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-53f04e26b81d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9286,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:848"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_allCoordinate_finiteBatchModel_confidence [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} (source : mdp.MeanCompatibleRewardKernel) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (defaultState : State) (rewardBound rewardBudget transitionBudget : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hrewardBudget_nonneg : 0 <= rewardBudget) (htransitionBud…","missing":[],"search":"iidstochastictrajectoryfamilymeasure_allcoordinate_finitebatchmodel_confidence banditrlproof.finitehorizonrl.markovpolicy.iidstochastictrajectoryfamilymeasure_allcoordinate_finitebatchmodel_confidence the iid stochastic-reward trajectory law produces a measurable all-coordinate confidence event for the empirical model built from the actual sampled rewards. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret","label":"iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret","description":"Outside the compiled stochastic empirical-model event, the sampled model is globally optimistic and its recommended policy satisfies the existing selected-radius expected-regret bound.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidallcoordinateempiricalmodelconfidence/index.html#decl-f1ac38e13130","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","order":9287,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence.lean:923"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} (source : mdp.MeanCompatibleRewardKernel) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (defaultState : State) (rewardBound rewardBudget transitionBudget : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hrewardBudget_nonneg : 0 <= rewardBudget) (htransitionBud…","missing":[],"search":"iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret banditrlproof.finitehorizonrl.markovpolicy.iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret outside the compiled stochastic empirical-model event, the sampled model is globally optimistic and its recommended policy satisfies the existing selected-radius expected-regret bound. theorem compiled","shard":"modules/e4faea21c4e46cf7.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_compProd_of_forall","label":"hasSubgaussianMGF_compProd_of_forall","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasSubgaussianMGF_compProd_of_forall","description":"A uniform fiberwise sub-Gaussian proxy survives an arbitrary probability mixture.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-aab49341be1b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9288,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:26"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_compProd_of_forall {Index : Type u} {Omega : Type v} [MeasurableSpace Index] [MeasurableSpace Omega] (mu : Measure Index) [IsProbabilityMeasure mu] (kappa : ProbabilityTheory.Kernel Index Omega) [ProbabilityTheory.IsMarkovKernel kappa] (X : Index × Omega -> Real) (hX : Measurable X) (c : NNReal) (hfiber : forall index, ProbabilityTheory.HasSubgaussianMGF (fun omega => X (index, omega)) c (kappa index)) : ProbabilityTheory.HasSubgaussianMGF X c (mu.compProd kappa)","missing":[],"search":"hassubgaussianmgf_compprod_of_forall banditrlproof.concentration.hassubgaussianmgf_compprod_of_forall a uniform fiberwise sub-gaussian proxy survives an arbitrary probability mixture. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_zero_of_proxy","label":"hasSubgaussianMGF_zero_of_proxy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasSubgaussianMGF_zero_of_proxy","description":"The zero random variable admits every nonnegative sub-Gaussian proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-3cbf78d9dd1d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9289,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_zero_of_proxy {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsProbabilityMeasure mu] (c : NNReal) : ProbabilityTheory.HasSubgaussianMGF (fun _ : Omega => 0) c mu","missing":[],"search":"hassubgaussianmgf_zero_of_proxy banditrlproof.concentration.hassubgaussianmgf_zero_of_proxy the zero random variable admits every nonnegative sub-gaussian proxy. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt","label":"stateAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt","description":"State immediately before a reward-bearing trajectory coordinate.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-5c50903182db","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9290,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:133"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def stateAt (initialState : State) {remaining : Nat} (trace : RewardStepTrace Action State remaining) (coordinate : Fin remaining) : State","missing":[],"search":"stateat banditrlproof.finitehorizonrl.rewardsteptrace.stateat state immediately before a reward-bearing trajectory coordinate. definition compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_stateAt","label":"measurable_stateAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_stateAt","description":"The pre-coordinate state is measurable as a function of the full trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-6cf4bf9cd23f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9291,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:143"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stateAt (initialState : State) {remaining : Nat} (coordinate : Fin remaining) : Measurable (fun trace : RewardStepTrace Action State remaining => stateAt initialState trace coordinate)","missing":[],"search":"measurable_stateat banditrlproof.finitehorizonrl.rewardsteptrace.measurable_stateat the pre-coordinate state is measurable as a function of the full trace. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt_cons_zero","label":"stateAt_cons_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt_cons_zero","description":"theorem stateAt_cons_zero (initialState : State) (remaining : Nat) (head : Action × (Real × State)) (tail : RewardStepTrace Action State remaining) : stateAt initialState (@Fin.cons remaining (fun _ => Action × (Real × State)) head tail) 0 = initialState","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-8664df69181c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9292,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:156"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateAt_cons_zero (initialState : State) (remaining : Nat) (head : Action × (Real × State)) (tail : RewardStepTrace Action State remaining) : stateAt initialState (@Fin.cons remaining (fun _ => Action × (Real × State)) head tail) 0 = initialState","missing":[],"search":"stateat_cons_zero banditrlproof.finitehorizonrl.rewardsteptrace.stateat_cons_zero theorem stateat_cons_zero (initialstate : state) (remaining : nat) (head : action × (real × state)) (tail : rewardsteptrace action state remaining) : stateat initialstate (@fin.cons remaining (fun _ => action × (real × state)) head tail) 0 = initialstate theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt_cons_succ","label":"stateAt_cons_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt_cons_succ","description":"theorem stateAt_cons_succ (initialState : State) (remaining : Nat) (head : Action × (Real × State)) (tail : RewardStepTrace Action State remaining) (coordinate : Fin remaining) : stateAt initialState (@Fin.cons remaining (fun _ => Action × (Real × State)) head tail) coordinate.succ = stateAt head.2.2 tail coordinate","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-03246b36376f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9293,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:168"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stateAt_cons_succ (initialState : State) (remaining : Nat) (head : Action × (Real × State)) (tail : RewardStepTrace Action State remaining) (coordinate : Fin remaining) : stateAt initialState (@Fin.cons remaining (fun _ => Action × (Real × State)) head tail) coordinate.succ = stateAt head.2.2 tail coordinate","missing":[],"search":"stateat_cons_succ banditrlproof.finitehorizonrl.rewardsteptrace.stateat_cons_succ theorem stateat_cons_succ (initialstate : state) (remaining : nat) (head : action × (real × state)) (tail : rewardsteptrace action state remaining) (coordinate : fin remaining) : stateat initialstate (@fin.cons remaining (fun _ => action × (real × state)) head tail) coordinate.succ = stateat head.2.2 tail coordinate theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_stateAt_prod","label":"measurable_stateAt_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_stateAt_prod","description":"The pre-coordinate state is measurable when the initial state is also an input.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-5d9bc2f3b15b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9294,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:203"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_stateAt_prod {remaining : Nat} (coordinate : Fin remaining) : Measurable (fun p : State × RewardStepTrace Action State remaining => stateAt p.1 p.2 coordinate)","missing":[],"search":"measurable_stateat_prod banditrlproof.finitehorizonrl.rewardsteptrace.measurable_stateat_prod the pre-coordinate state is measurable when the initial state is also an input. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.maskedRewardDeviationAt","label":"maskedRewardDeviationAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.maskedRewardDeviationAt","description":"A sampled reward centered at one target coordinate and masked by its visit event.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-a9d30d00eb54","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9295,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def maskedRewardDeviationAt (mdp : MDP State Action) (initialState : State) {remaining : Nat} (trace : RewardStepTrace Action State remaining) (coordinate : Fin remaining) (state : State) (action : Action) : Real","missing":[],"search":"maskedrewarddeviationat banditrlproof.finitehorizonrl.mdp.maskedrewarddeviationat a sampled reward centered at one target coordinate and masked by its visit event. definition compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_maskedRewardDeviationAt","label":"measurable_maskedRewardDeviationAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_maskedRewardDeviationAt","description":"A fixed visit-masked reward deviation is measurable on reward-bearing traces.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-1ed053fbfb7f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9296,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:227"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_maskedRewardDeviationAt (mdp : MDP State Action) (initialState : State) {remaining : Nat} (coordinate : Fin remaining) (state : State) (action : Action) : Measurable (fun trace => mdp.maskedRewardDeviationAt initialState trace coordinate state action)","missing":[],"search":"measurable_maskedrewarddeviationat banditrlproof.finitehorizonrl.mdp.measurable_maskedrewarddeviationat a fixed visit-masked reward deviation is measurable on reward-bearing traces. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_maskedRewardDeviationAt_trajectory","label":"measurable_maskedRewardDeviationAt_trajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_maskedRewardDeviationAt_trajectory","description":"The masked deviation is measurable when the initial state is part of the input.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-79e330771ba1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9297,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:243"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_maskedRewardDeviationAt_trajectory (mdp : MDP State Action) {remaining : Nat} (coordinate : Fin remaining) (state : State) (action : Action) : Measurable (fun trajectory : State × RewardStepTrace Action State remaining => mdp.maskedRewardDeviationAt trajectory.1 trajectory.2 coordinate state action)","missing":[],"search":"measurable_maskedrewarddeviationat_trajectory banditrlproof.finitehorizonrl.mdp.measurable_maskedrewarddeviationat_trajectory the masked deviation is measurable when the initial state is part of the input. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStepOfStochasticTrajectory","label":"sampledEpisodeStepOfStochasticTrajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStepOfStochasticTrajectory","description":"One sampled reward-bearing trajectory converted to an empirical record.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-18ee4665164c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9298,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:262"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def sampledEpisodeStepOfStochasticTrajectory (mdp : MDP State Action) (trajectory : State × RewardStepTrace Action State mdp.horizon) (stage : Fin mdp.horizon) : EpisodeStep State Action where","missing":[],"search":"sampledepisodestepofstochastictrajectory banditrlproof.finitehorizonrl.mdp.sampledepisodestepofstochastictrajectory one sampled reward-bearing trajectory converted to an empirical record. definition compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledEpisodeStepOfStochasticTrajectory","label":"measurable_sampledEpisodeStepOfStochasticTrajectory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledEpisodeStepOfStochasticTrajectory","description":"Extracting a sampled empirical record from a stochastic trajectory is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-4696801fb21e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9299,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:271"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledEpisodeStepOfStochasticTrajectory (mdp : MDP State Action) (stage : Fin mdp.horizon) : Measurable (fun trajectory => mdp.sampledEpisodeStepOfStochasticTrajectory trajectory stage)","missing":[],"search":"measurable_sampledepisodestepofstochastictrajectory banditrlproof.finitehorizonrl.mdp.measurable_sampledepisodestepofstochastictrajectory extracting a sampled empirical record from a stochastic trajectory is measurable. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatchOfStochasticTrajectories","label":"sampledEpisodeBatchOfStochasticTrajectories","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatchOfStochasticTrajectories","description":"A finite iid family of stochastic trajectories mapped to sampled-reward records.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-c2abb5fb332a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9300,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def sampledEpisodeBatchOfStochasticTrajectories (mdp : MDP State Action) (episodes : Nat) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : EpisodeBatch mdp episodes","missing":[],"search":"sampledepisodebatchofstochastictrajectories banditrlproof.finitehorizonrl.mdp.sampledepisodebatchofstochastictrajectories a finite iid family of stochastic trajectories mapped to sampled-reward records. definition compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledEpisodeBatchOfStochasticTrajectories","label":"measurable_sampledEpisodeBatchOfStochasticTrajectories","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledEpisodeBatchOfStochasticTrajectories","description":"Mapping a finite stochastic trajectory family to sampled episode records is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-9ca9fd1fdc9c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9301,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:302"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledEpisodeBatchOfStochasticTrajectories (mdp : MDP State Action) (episodes : Nat) : Measurable (mdp.sampledEpisodeBatchOfStochasticTrajectories episodes)","missing":[],"search":"measurable_sampledepisodebatchofstochastictrajectories banditrlproof.finitehorizonrl.mdp.measurable_sampledepisodebatchofstochastictrajectories mapping a finite stochastic trajectory family to sampled episode records is measurable. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_maskedRewardDeviation_hasSubgaussianMGF","label":"actionRewardStateKernel_maskedRewardDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_maskedRewardDeviation_hasSubgaussianMGF","description":"A visit-masked one-step reward deviation retains the common reward proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-8acc11b863d9","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9302,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:319"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardStateKernel_maskedRewardDeviation_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (initialState state : State) (action : Action) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (fun head : Action × (Real × State) => if initialState = state /\\ head.1 = action then head.2.1 - mdp.reward state action else 0) varianceProxy (source.actionRewardStateKernel policy stage initialState)","missing":[],"search":"actionrewardstatekernel_maskedrewarddeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardstatekernel_maskedrewarddeviation_hassubgaussianmgf a visit-masked one-step reward deviation retains the common reward proxy. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_maskedRewardDeviationAt_hasSubgaussianMGF","label":"stochasticTrajectoryKernelRemaining_maskedRewardDeviationAt_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_maskedRewardDeviationAt_hasSubgaussianMGF","description":"Every coordinate of a generated reward-bearing trace has the masked reward MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-0337d990d6f5","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9303,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:371"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_maskedRewardDeviationAt_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (initialState : State) (coordinate : Fin remaining) (state : State) (action : Action) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (fun trace => mdp.maskedRewardDeviationAt initialState trace coordinate state action) varianceProxy (source.stochasticTrajectoryKernelRemaining policy remaining hremaining initialState)","missing":[],"search":"stochastictrajectorykernelremaining_maskedrewarddeviationat_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_maskedrewarddeviationat_hassubgaussianmgf every coordinate of a generated reward-bearing trace has the masked reward mgf. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_maskedRewardDeviationAt_hasSubgaussianMGF","label":"stochasticTrajectoryMeasure_maskedRewardDeviationAt_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_maskedRewardDeviationAt_hasSubgaussianMGF","description":"The same fixed coordinate MGF holds after mixing over the initial-state law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-e0f031f5f6f8","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9304,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:454"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryMeasure_maskedRewardDeviationAt_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (fun trajectory : State × RewardStepTrace Action State mdp.horizon => mdp.maskedRewardDeviationAt trajectory.1 trajectory.2 stage state action) varianceProxy (source.stochasticTrajectoryMeasure policy initialState)","missing":[],"search":"stochastictrajectorymeasure_maskedrewarddeviationat_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure_maskedrewarddeviationat_hassubgaussianmgf the same fixed coordinate mgf holds after mixing over the initial-state law. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviationAtEpisode","label":"maskedRewardDeviationAtEpisode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviationAtEpisode","description":"One iid episode's fixed-coordinate masked reward deviation. The receiver keeps this structural coordinate on the same source-indexed API as its law theorems; the pointwise value itself depends only on the sampled trajectory and the MDP.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-163c345f347a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9305,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:482"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def maskedRewardDeviationAtEpisode (source : MeanCompatibleRewardKernel mdp) (stage : Fin mdp.horizon) (state : State) (action : Action) {episodes : Nat} (episode : Fin episodes) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : Real","missing":[],"search":"maskedrewarddeviationatepisode banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.maskedrewarddeviationatepisode one iid episode's fixed-coordinate masked reward deviation. the receiver keeps this structural coordinate on the same source-indexed api as its law theorems; the pointwise value itself depends only on the sampled trajectory and the mdp. definition compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_maskedRewardDeviationAtEpisode","label":"iIndepFun_maskedRewardDeviationAtEpisode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_maskedRewardDeviationAtEpisode","description":"Fixed-coordinate masked reward deviations are independent across iid episodes.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-be784a9f4766","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9306,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:491"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_maskedRewardDeviationAtEpisode (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) : ProbabilityTheory.iIndepFun (source.maskedRewardDeviationAtEpisode stage state action) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iindepfun_maskedrewarddeviationatepisode banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iindepfun_maskedrewarddeviationatepisode fixed-coordinate masked reward deviations are independent across iid episodes. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviationAtEpisode_hasSubgaussianMGF","label":"maskedRewardDeviationAtEpisode_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviationAtEpisode_hasSubgaussianMGF","description":"Every iid episode coordinate inherits the complete-trajectory masked reward MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-311b97ef8bd0","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9307,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:508"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem maskedRewardDeviationAtEpisode_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (stage : Fin mdp.horizon) (state : State) (action : Action) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) {episodes : Nat} (episode : Fin episodes) : ProbabilityTheory.HasSubgaussianMGF (source.maskedRewardDeviationAtEpisode stage state action episode) varianceProxy (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"maskedrewarddeviationatepisode_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.maskedrewarddeviationatepisode_hassubgaussianmgf every iid episode coordinate inherits the complete-trajectory masked reward mgf. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_hasSubgaussianMGF","label":"iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_hasSubgaussianMGF","description":"The iid fixed-coordinate reward-deviation sum has the episode-linear proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-2f089309e38f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9308,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:537"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) : ProbabilityTheory.HasSubgaussianMGF (fun trajectories => ∑ episode : Fin episodes, source.maskedRewardDeviationAtEpisode stage state action episode trajectories) ((episodes : NNReal) * varianceProxy) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iidstochastictrajectoryfamilymeasure_maskedrewarddeviation_sum_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_maskedrewarddeviation_sum_hassubgaussianmgf the iid fixed-coordinate reward-deviation sum has the episode-linear proxy. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_abs_tail_le","label":"iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_abs_tail_le","description":"Fixed-coordinate two-sided reward-sum tail under the iid stochastic family law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidempiricalrewardconfidence/index.html#decl-e19c036c6bd6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","order":9309,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence.lean:561"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (stage : Fin mdp.horizon) (state : State) (action : Action) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes) {trajectories | Concentration.subGaussianSumConfidenceRadius ((episodes : NNReal) * varianceProxy) delta <= |∑ episode : Fin episodes, source.maskedRewardDeviationAtEpisode stage state action episode trajectories|} <= ENNReal.ofReal delta","missing":[],"search":"iidstochastictrajectoryfamilymeasure_maskedrewarddeviation_sum_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_maskedrewarddeviation_sum_abs_tail_le fixed-coordinate two-sided reward-sum tail under the iid stochastic family law. theorem compiled","shard":"modules/4e005ff9f4704783.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticRewardCoordinateRadius","label":"uniformFloorStochasticRewardCoordinateRadius","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticRewardCoordinateRadius","description":"Uniform reward-mean radius obtained from one common expected-count floor.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-b72419c54b63","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9310,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformFloorStochasticRewardCoordinateRadius (mdp : MDP State Action) (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta visitFloor : Real) : Real","missing":[],"search":"uniformfloorstochasticrewardcoordinateradius banditrlproof.finitehorizonrl.uniformfloorstochasticrewardcoordinateradius uniform reward-mean radius obtained from one common expected-count floor. definition compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticRewardCoordinateRadius_nonneg","label":"uniformFloorStochasticRewardCoordinateRadius_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticRewardCoordinateRadius_nonneg","description":"The uniform reward radius is nonnegative under the strict count margin.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-43682e2fe3a4","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9311,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:40"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformFloorStochasticRewardCoordinateRadius_nonneg {mdp : MDP State Action} {episodes : Nat} {varianceProxy : NNReal} {countDelta rewardDelta visitFloor : Real} (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < (episodes : Real) * visitFloor) : 0 <= uniformFloorStochasticRewardCoordinateRadius mdp episodes varianceProxy countDelta rewardDelta visitFloor","missing":[],"search":"uniformfloorstochasticrewardcoordinateradius_nonneg banditrlproof.finitehorizonrl.uniformfloorstochasticrewardcoordinateradius_nonneg the uniform reward radius is nonnegative under the strict count margin. theorem compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionBudget","label":"uniformFloorStochasticTransitionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionBudget","description":"The explicit transition budget paired with the uniform reward budget.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-da1235131b1d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9312,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def uniformFloorStochasticTransitionBudget (rewardBound rewardBudget : Real) : Real","missing":[],"search":"uniformfloorstochastictransitionbudget banditrlproof.finitehorizonrl.uniformfloorstochastictransitionbudget the explicit transition budget paired with the uniform reward budget. definition compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.expectedCountRewardCoordinateRadius_le_uniformFloor","label":"expectedCountRewardCoordinateRadius_le_uniformFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.expectedCountRewardCoordinateRadius_le_uniformFloor","description":"A common expected-count floor dominates every coordinate reward radius.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-735a43fed78a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9313,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:64"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem expectedCountRewardCoordinateRadius_le_uniformFloor (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta visitFloor : Real) (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < (episodes : Real) * visitFloor) (hcountFloor : forall coordinate : VisitCoordinate mdp, (episodes : Real) * visitFloor <= coordinate.expectedCount policy initialState episodes) (coordinate : VisitCoordinate mdp) : source.expectedCountRewardCoordinateRadius policy initialState episodes varianceProxy countDelta rewardDelta coordinate <= uniformFloorStochasticRewardCoordinateRadius mdp episodes varianceProxy countDelta rewardDelta visitFloor","missing":[],"search":"expectedcountrewardcoordinateradius_le_uniformfloor banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.expectedcountrewardcoordinateradius_le_uniformfloor a common expected-count floor dominates every coordinate reward radius. theorem compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticTransitionCover_of_uniformExpectedCountFloor","label":"stochasticTransitionCover_of_uniformExpectedCountFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticTransitionCover_of_uniformExpectedCountFloor","description":"The common count floor and half-contraction condition cover every stochastic transition-radius/value-envelope sum with the explicit transition budget.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-96e7a41753e7","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9314,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTransitionCover_of_uniformExpectedCountFloor {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta visitFloor rewardBound : Real) (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < (episodes : Real) * visitFloor) (hcountFloor : forall coordinate : VisitCoordinate mdp, (episodes : Real) * visitFloor <= coordinate.expectedCount policy initialState episodes) (hrewardBound_nonneg : 0 <= rewardBound) (hcontraction : (Fintype.card State : Real) * uniformFloorTransitionCoordinateRadius mdp episodes countDelta visitFloor * (mdp.horizon : Real) <= 1 / 2) : let rewardBudget := uniformFloorStochasticRewardCoordinateRadius mdp episodes varianceProxy countDelta rewardDelta visitFloor let transitionBudget := uniformFloorStochasticTra…","missing":[],"search":"stochastictransitioncover_of_uniformexpectedcountfloor banditrlproof.finitehorizonrl.markovpolicy.stochastictransitioncover_of_uniformexpectedcountfloor the common count floor and half-contraction condition cover every stochastic transition-radius/value-envelope sum with the explicit transition budget. theorem compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor","label":"iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor","description":"Fixed-policy route endpoint with every coordinate margin and cover produced by one common expected-count floor and one scalar half-contraction condition.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-38ece30b9263","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9315,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:211"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} (source : mdp.MeanCompatibleRewardKernel) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (defaultState : State) (rewardBound visitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hmargin : simultaneousCountConfidenceRadius mdp…","missing":[],"search":"iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_uniformexpectedcountfloor banditrlproof.finitehorizonrl.markovpolicy.iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_uniformexpectedcountfloor fixed-policy route endpoint with every coordinate margin and cover produced by one common expected-count floor and one scalar half-contraction condition. theorem compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_explicitCalibration","label":"exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_explicitCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_explicitCalibration","description":"Practical endpoint: exploratory path support constructs the common count floor, then the explicit stochastic calibration yields confidence and recommended expected regret for the exploratory policy's sampled-reward empirical model.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidexplicitcalibration/index.html#decl-0f195594da36","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","order":9316,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDExplicitCalibration.lean:307"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_explicitCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (source : mdp.MeanCompatibleRewardKernel) (initialState : Measure State) [IsProbabilityMeasure initialState] (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (defaultState : State) (rewardBound : Real) (hrewardBound : forall state…","missing":[],"search":"exploratorypolicy_iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_pathsupport_explicitcalibration banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratorypolicy_iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_pathsupport_explicitcalibration practical endpoint: exploratory path support constructs the common count floor, then the explicit stochastic calibration yields confidence and recommended expected regret for the exploratory policy's sampled-reward empirical model. theorem compiled","shard":"modules/86e110363a4d0d40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget","label":"selfConsistentTransitionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget","description":"Exact nonnegative fixed-point budget for `q * (base + budget) <= budget`.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-68f28c72df03","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9317,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:20"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentTransitionBudget (q base : Real) : Real","missing":[],"search":"selfconsistenttransitionbudget banditrlproof.finitehorizonrl.selfconsistenttransitionbudget exact nonnegative fixed-point budget for `q * (base + budget) <= budget`. definition compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget_nonneg","label":"selfConsistentTransitionBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget_nonneg","description":"The fixed-point budget is nonnegative when `0 <= q < 1` and `base >= 0`.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-a9b02d6f14b5","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9318,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentTransitionBudget_nonneg {q base : Real} (hq_nonneg : 0 <= q) (hq : q < 1) (hbase : 0 <= base) : 0 <= selfConsistentTransitionBudget q base","missing":[],"search":"selfconsistenttransitionbudget_nonneg banditrlproof.finitehorizonrl.selfconsistenttransitionbudget_nonneg the fixed-point budget is nonnegative when `0 <= q < 1` and `base >= 0`. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget_fixedPoint","label":"selfConsistentTransitionBudget_fixedPoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget_fixedPoint","description":"The chosen budget solves the transition-envelope fixed point exactly.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-23847b944a46","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9319,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentTransitionBudget_fixedPoint {q base : Real} (hq : q < 1) : q * (base + selfConsistentTransitionBudget q base) = selfConsistentTransitionBudget q base","missing":[],"search":"selfconsistenttransitionbudget_fixedpoint banditrlproof.finitehorizonrl.selfconsistenttransitionbudget_fixedpoint the chosen budget solves the transition-envelope fixed point exactly. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionContraction","label":"uniformFloorStochasticTransitionContraction","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionContraction","description":"The exact common-floor transition contraction factor.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-890e8c30d8be","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9320,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:50"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformFloorStochasticTransitionContraction (mdp : MDP State Action) (episodes : Nat) (countDelta visitFloor : Real) : Real","missing":[],"search":"uniformfloorstochastictransitioncontraction banditrlproof.finitehorizonrl.uniformfloorstochastictransitioncontraction the exact common-floor transition contraction factor. definition compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionContraction_nonneg","label":"uniformFloorStochasticTransitionContraction_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionContraction_nonneg","description":"The common-floor contraction factor is nonnegative under the count margin.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-89bca1bbdf4f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9321,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformFloorStochasticTransitionContraction_nonneg {mdp : MDP State Action} {episodes : Nat} {countDelta visitFloor : Real} (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < (episodes : Real) * visitFloor) : 0 <= uniformFloorStochasticTransitionContraction mdp episodes countDelta visitFloor","missing":[],"search":"uniformfloorstochastictransitioncontraction_nonneg banditrlproof.finitehorizonrl.uniformfloorstochastictransitioncontraction_nonneg the common-floor contraction factor is nonnegative under the count margin. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget","label":"uniformFloorStochasticSelfConsistentTransitionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget","description":"Shrinking transition budget obtained from the exact contraction factor.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-f6cee6eea7a1","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9322,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:75"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformFloorStochasticSelfConsistentTransitionBudget (mdp : MDP State Action) (episodes : Nat) (countDelta visitFloor rewardBound rewardBudget : Real) : Real","missing":[],"search":"uniformfloorstochasticselfconsistenttransitionbudget banditrlproof.finitehorizonrl.uniformfloorstochasticselfconsistenttransitionbudget shrinking transition budget obtained from the exact contraction factor. definition compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget_nonneg","label":"uniformFloorStochasticSelfConsistentTransitionBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget_nonneg","description":"The shrinking transition budget is nonnegative under `q < 1`.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-696a91caee95","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9323,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:87"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformFloorStochasticSelfConsistentTransitionBudget_nonneg {mdp : MDP State Action} {episodes : Nat} {countDelta visitFloor rewardBound rewardBudget : Real} (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < (episodes : Real) * visitFloor) (hrewardBound_nonneg : 0 <= rewardBound) (hrewardBudget_nonneg : 0 <= rewardBudget) (hq : uniformFloorStochasticTransitionContraction mdp episodes countDelta visitFloor < 1) : 0 <= uniformFloorStochasticSelfConsistentTransitionBudget mdp episodes countDelta visitFloor rewardBound rewardBudget","missing":[],"search":"uniformfloorstochasticselfconsistenttransitionbudget_nonneg banditrlproof.finitehorizonrl.uniformfloorstochasticselfconsistenttransitionbudget_nonneg the shrinking transition budget is nonnegative under `q < 1`. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget_fixedPoint","label":"uniformFloorStochasticSelfConsistentTransitionBudget_fixedPoint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget_fixedPoint","description":"The common-floor shrinking budget satisfies the exact envelope identity.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-9e88326894f6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9324,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:108"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem uniformFloorStochasticSelfConsistentTransitionBudget_fixedPoint {mdp : MDP State Action} {episodes : Nat} {countDelta visitFloor rewardBound rewardBudget : Real} (hq : uniformFloorStochasticTransitionContraction mdp episodes countDelta visitFloor < 1) : uniformFloorStochasticTransitionContraction mdp episodes countDelta visitFloor * (rewardBound + 2 * rewardBudget + uniformFloorStochasticSelfConsistentTransitionBudget mdp episodes countDelta visitFloor rewardBound rewardBudget) = uniformFloorStochasticSelfConsistentTransitionBudget mdp episodes countDelta visitFloor rewardBound rewardBudget","missing":[],"search":"uniformfloorstochasticselfconsistenttransitionbudget_fixedpoint banditrlproof.finitehorizonrl.uniformfloorstochasticselfconsistenttransitionbudget_fixedpoint the common-floor shrinking budget satisfies the exact envelope identity. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticTransitionCover_of_uniformExpectedCountFloor_selfConsistent","label":"stochasticTransitionCover_of_uniformExpectedCountFloor_selfConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticTransitionCover_of_uniformExpectedCountFloor_selfConsistent","description":"The exact `q < 1` fixed point covers every transition-radius/value-envelope sum with a budget that shrinks as `q` tends to zero.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-19518f675d58","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9325,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:131"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTransitionCover_of_uniformExpectedCountFloor_selfConsistent {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (varianceProxy : NNReal) (countDelta rewardDelta visitFloor rewardBound : Real) (hmargin : simultaneousCountConfidenceRadius mdp episodes countDelta < (episodes : Real) * visitFloor) (hcountFloor : forall coordinate : VisitCoordinate mdp, (episodes : Real) * visitFloor <= coordinate.expectedCount policy initialState episodes) (hrewardBound_nonneg : 0 <= rewardBound) (hq : uniformFloorStochasticTransitionContraction mdp episodes countDelta visitFloor < 1) : let rewardBudget := uniformFloorStochasticRewardCoordinateRadius mdp episodes varianceProxy countDelta rewardDelta visitFloor let transitionBudget := uniformFloorStochasticSelfConsistentTransitionBudget mdp episodes countDe…","missing":[],"search":"stochastictransitioncover_of_uniformexpectedcountfloor_selfconsistent banditrlproof.finitehorizonrl.markovpolicy.stochastictransitioncover_of_uniformexpectedcountfloor_selfconsistent the exact `q < 1` fixed point covers every transition-radius/value-envelope sum with a budget that shrinks as `q` tends to zero. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor_selfConsistent","label":"iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor_selfConsistent","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor_selfConsistent","description":"Fixed-policy all-coordinate confidence under the shrinking transition budget.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-0d4cdeb28075","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9326,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:242"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor_selfConsistent [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} (source : mdp.MeanCompatibleRewardKernel) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (defaultState : State) (rewardBound visitFloor : Real) (hrewardBound : forall state action, |mdp.reward state action| <= rewardBound) (hmargin : simultaneousCountConfi…","missing":[],"search":"iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_uniformexpectedcountfloor_selfconsistent banditrlproof.finitehorizonrl.markovpolicy.iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_uniformexpectedcountfloor_selfconsistent fixed-policy all-coordinate confidence under the shrinking transition budget. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_selfConsistentCalibration","label":"exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_selfConsistentCalibration","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_selfConsistentCalibration","description":"Exploratory path support feeds the shrinking fixed-policy calibration.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidselfconsistentcalibration/index.html#decl-a274a121c0ea","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","order":9327,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDSelfConsistentCalibration.lean:333"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_selfConsistentCalibration [StandardBorelSpace State] [StandardBorelSpace Action] {mdp : MDP State Action} (table : DeterministicMarkovPolicyTable mdp) (source : mdp.MeanCompatibleRewardKernel) (initialState : Measure State) [IsProbabilityMeasure initialState] (explorationRate : NNReal) (hexplorationRate : explorationRate <= 1) (episodes : Nat) (hepisodes : 0 < episodes) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (htotal : 0 < ((((episodes : NNReal) * varianceProxy : NNReal) : Real))) (countDelta : Real) (hcountDelta : 0 < countDelta) (hcountDelta_le_one : countDelta <= 1) (rewardDelta : Real) (hrewardDelta : 0 < rewardDelta) (hrewardDelta_le_one : rewardDelta <= 1) (defaultState : State) (rewardBound : Real) (hrewardBound : forall…","missing":[],"search":"exploratorypolicy_iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_pathsupport_selfconsistentcalibration banditrlproof.finitehorizonrl.deterministicmarkovpolicytable.exploratorypolicy_iidstochastictrajectoryfamilymeasure_allcoordinate_optimism_and_expectedregret_of_pathsupport_selfconsistentcalibration exploratory path support feeds the shrinking fixed-policy calibration. theorem compiled","shard":"modules/a2dd38ed46a3af40.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationAtEpisode","label":"sampledCumulativeReturnDeviationAtEpisode","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationAtEpisode","description":"The sampled-return deviation in one coordinate of a finite episode family.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-c86a25b3a5f7","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9328,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:28"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeReturnDeviationAtEpisode (mdp : MDP State Action) (policy : MarkovPolicy mdp) {episodes : Nat} (episode : Fin episodes) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : Real","missing":[],"search":"sampledcumulativereturndeviationatepisode banditrlproof.finitehorizonrl.mdp.sampledcumulativereturndeviationatepisode the sampled-return deviation in one coordinate of a finite episode family. definition compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationAtEpisode","label":"measurable_sampledCumulativeReturnDeviationAtEpisode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationAtEpisode","description":"One episode-coordinate deviation is measurable on the finite product space.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-dfc5cb0d60a9","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9329,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:36"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeReturnDeviationAtEpisode (mdp : MDP State Action) (policy : MarkovPolicy mdp) {episodes : Nat} (episode : Fin episodes) : Measurable (mdp.sampledCumulativeReturnDeviationAtEpisode policy episode)","missing":[],"search":"measurable_sampledcumulativereturndeviationatepisode banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativereturndeviationatepisode one episode-coordinate deviation is measurable on the finite product space. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationSum","label":"sampledCumulativeReturnDeviationSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationSum","description":"Sum of sampled-return deviations over complete iid episode coordinates.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-4f3e5dc44c4d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9330,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeReturnDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (episodes : Nat) (trajectories : Fin episodes -> State × RewardStepTrace Action State mdp.horizon) : Real","missing":[],"search":"sampledcumulativereturndeviationsum banditrlproof.finitehorizonrl.mdp.sampledcumulativereturndeviationsum sum of sampled-return deviations over complete iid episode coordinates. definition compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationSum","label":"measurable_sampledCumulativeReturnDeviationSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationSum","description":"The finite-episode deviation sum is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-4c5417ff408b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9331,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeReturnDeviationSum (mdp : MDP State Action) (policy : MarkovPolicy mdp) (episodes : Nat) : Measurable (mdp.sampledCumulativeReturnDeviationSum policy episodes)","missing":[],"search":"measurable_sampledcumulativereturndeviationsum banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativereturndeviationsum the finite-episode deviation sum is measurable. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidSampledCumulativeReturnDeviationVarianceProxy","label":"iidSampledCumulativeReturnDeviationVarianceProxy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.iidSampledCumulativeReturnDeviationVarianceProxy","description":"Episode-linear variance proxy for the iid sampled-return deviation sum.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-957e9ad106ea","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9332,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:61"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidSampledCumulativeReturnDeviationVarianceProxy (mdp : MDP State Action) (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) : NNReal","missing":[],"search":"iidsampledcumulativereturndeviationvarianceproxy banditrlproof.finitehorizonrl.mdp.iidsampledcumulativereturndeviationvarianceproxy episode-linear variance proxy for the iid sampled-return deviation sum. definition compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure","label":"iidStochasticTrajectoryFamilyMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure","description":"Finite iid product of the complete reward-bearing stochastic trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-120d6e38ffed","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9333,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:76"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def iidStochasticTrajectoryFamilyMeasure (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : Measure (Fin episodes -> State × RewardStepTrace Action State mdp.horizon)","missing":[],"search":"iidstochastictrajectoryfamilymeasure banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure finite iid product of the complete reward-bearing stochastic trajectory law. definition compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_eval","label":"iidStochasticTrajectoryFamilyMeasure_map_eval","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_eval","description":"Every iid product coordinate has the exact complete stochastic trajectory law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-dea21a283d49","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9334,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:96"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_map_eval (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] {episodes : Nat} (episode : Fin episodes) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes).map (Function.eval episode) = source.stochasticTrajectoryMeasure policy initialState","missing":[],"search":"iidstochastictrajectoryfamilymeasure_map_eval banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_map_eval every iid product coordinate has the exact complete stochastic trajectory law. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_sampledCumulativeReturnDeviationAtEpisode","label":"iIndepFun_sampledCumulativeReturnDeviationAtEpisode","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_sampledCumulativeReturnDeviationAtEpisode","description":"Complete sampled-return deviations are independent across iid episodes.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-abf98099822d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9335,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:111"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iIndepFun_sampledCumulativeReturnDeviationAtEpisode (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) : ProbabilityTheory.iIndepFun (fun episode trajectories => mdp.sampledCumulativeReturnDeviationAtEpisode policy episode trajectories) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iindepfun_sampledcumulativereturndeviationatepisode banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iindepfun_sampledcumulativereturndeviationatepisode complete sampled-return deviations are independent across iid episodes. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledCumulativeReturnDeviationAtEpisode_hasSubgaussianMGF","label":"sampledCumulativeReturnDeviationAtEpisode_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledCumulativeReturnDeviationAtEpisode_hasSubgaussianMGF","description":"Each iid episode coordinate inherits the compiled initial-law MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-0d3c3dc844b3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9336,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:126"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledCumulativeReturnDeviationAtEpisode_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) {episodes : Nat} (episode : Fin episodes) : ProbabilityTheory.HasSubgaussianMGF (mdp.sampledCumulativeReturnDeviationAtEpisode policy episode) ((mdp.horizon : NNReal) * rewardVarianceProxy + meanBellmanInnovationVarianceProxy rewardBound mdp.horizon) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"sampledcumulativereturndeviationatepisode_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.sampledcumulativereturndeviationatepisode_hassubgaussianmgf each iid episode coordinate inherits the compiled initial-law mgf. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_hasSubgaussianMGF","label":"iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_hasSubgaussianMGF","description":"The finite iid episode deviation sum has the episode-linear proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-62eae120026b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9337,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:160"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasSubgaussianMGF (mdp.sampledCumulativeReturnDeviationSum policy episodes) (mdp.iidSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy) (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes)","missing":[],"search":"iidstochastictrajectoryfamilymeasure_sampledcumulativereturndeviationsum_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_sampledcumulativereturndeviationsum_hassubgaussianmgf the finite iid episode deviation sum has the episode-linear proxy. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_abs_tail_le","label":"iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_abs_tail_le","description":"Fixed-sample two-sided delta tail for the iid episode deviation sum.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardiidtotalreturnconcentration/index.html#decl-01cd130d4e91","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","order":9338,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardIIDTotalReturnConcentration.lean:193"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (episodes : Nat) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((mdp.iidSampledCumulativeReturnDeviationVarianceProxy episodes rewardBound rewardVarianceProxy : NNReal) : Real)) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.iidStochasticTrajectoryFamilyMeasure policy initialState episodes) {trajectories | Concentration.subGaussianSumConfidenceRadius (mdp.iidSampledCumulativeReturnDeviationVarianceProxy…","missing":[],"search":"iidstochastictrajectoryfamilymeasure_sampledcumulativereturndeviationsum_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.iidstochastictrajectoryfamilymeasure_sampledcumulativereturndeviationsum_abs_tail_le fixed-sample two-sided delta tail for the iid episode deviation sum. theorem compiled","shard":"modules/4e222b16e8fb2ba2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_compProd_of_forall_fintype","label":"hasSubgaussianMGF_compProd_of_forall_fintype","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Concentration.hasSubgaussianMGF_compProd_of_forall_fintype","description":"A common sub-Gaussian proxy on every fiber of a Markov kernel is preserved by mixing the fibers with a probability measure on a finite index type.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardinitiallawtotalreturnconcentration/index.html#decl-a48a70d4f163","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","order":9339,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem hasSubgaussianMGF_compProd_of_forall_fintype {Index Omega : Type*} [MeasurableSpace Index] [MeasurableSpace Omega] [Fintype Index] (mu : Measure Index) [IsProbabilityMeasure mu] (kappa : ProbabilityTheory.Kernel Index Omega) [ProbabilityTheory.IsMarkovKernel kappa] (X : Index × Omega -> Real) (hX : Measurable X) (c : NNReal) (hfiber : forall index, ProbabilityTheory.HasSubgaussianMGF (fun omega => X (index, omega)) c (kappa index)) : ProbabilityTheory.HasSubgaussianMGF X c (mu.compProd kappa)","missing":[],"search":"hassubgaussianmgf_compprod_of_forall_fintype banditrlproof.concentration.hassubgaussianmgf_compprod_of_forall_fintype a common sub-gaussian proxy on every fiber of a markov kernel is preserved by mixing the fibers with a probability measure on a finite index type. theorem compiled","shard":"modules/340d176746b0eace.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviation","label":"sampledCumulativeReturnDeviation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviation","description":"Sampled cumulative return centered by the policy value at the trajectory's own initial state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardinitiallawtotalreturnconcentration/index.html#decl-9de9e4c55d05","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","order":9340,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration.lean:91"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeReturnDeviation (mdp : MDP State Action) (policy : MarkovPolicy mdp) (trajectory : State × RewardStepTrace Action State mdp.horizon) : Real","missing":[],"search":"sampledcumulativereturndeviation banditrlproof.finitehorizonrl.mdp.sampledcumulativereturndeviation sampled cumulative return centered by the policy value at the trajectory's own initial state. definition compiled","shard":"modules/340d176746b0eace.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviation","label":"measurable_sampledCumulativeReturnDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviation","description":"The full initial-state-dependent sampled-return deviation is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardinitiallawtotalreturnconcentration/index.html#decl-431a43d362d7","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","order":9341,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeReturnDeviation (mdp : MDP State Action) (policy : MarkovPolicy mdp) : Measurable (mdp.sampledCumulativeReturnDeviation policy)","missing":[],"search":"measurable_sampledcumulativereturndeviation banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativereturndeviation the full initial-state-dependent sampled-return deviation is measurable. theorem compiled","shard":"modules/340d176746b0eace.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_hasSubgaussianMGF","label":"stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_hasSubgaussianMGF","description":"Initial-law sampled-return MGF bound. The centering is state dependent, so the common statewise proxy passes unchanged through the finite initial-state mix.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardinitiallawtotalreturnconcentration/index.html#decl-3d5ebd74c71a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","order":9342,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration.lean:112"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) : ProbabilityTheory.HasSubgaussianMGF (mdp.sampledCumulativeReturnDeviation policy) ((mdp.horizon : NNReal) * rewardVarianceProxy + meanBellmanInnovationVarianceProxy rewardBound mdp.horizon) (source.stochasticTrajectoryMeasure policy initialState)","missing":[],"search":"stochastictrajectorymeasure_sampledcumulativereturndeviation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure_sampledcumulativereturndeviation_hassubgaussianmgf initial-law sampled-return mgf bound. the centering is state dependent, so the common statewise proxy passes unchanged through the finite initial-state mix. theorem compiled","shard":"modules/340d176746b0eace.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_abs_tail_le","label":"stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_abs_tail_le","description":"Fixed-horizon two-sided delta tail under an arbitrary finite initial-state law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardinitiallawtotalreturnconcentration/index.html#decl-d87dad40e6b9","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","order":9343,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) (htotal : 0 < ((((mdp.horizon : NNReal) * rewardVarianceProxy + meanBellmanInnovationVarianceProxy rewardBound mdp.horizon : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.stochasticTrajectoryMeasure policy initialState) {trajectory | Concentration.subGaussianSumConfidenceRadius ((mdp.horizon : NNReal) * rewardVarianceProxy + meanBellmanInnovationVarianceProxy rewar…","missing":[],"search":"stochastictrajectorymeasure_sampledcumulativereturndeviation_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure_sampledcumulativereturndeviation_abs_tail_le fixed-horizon two-sided delta tail under an arbitrary finite initial-state law. theorem compiled","shard":"modules/340d176746b0eace.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.head","label":"head","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.head","description":"The first sampled action, reward, and next state of a positive trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-b82a0a89f25c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9344,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def head (remaining : Nat) : RewardStepTrace Action State (remaining + 1) -> Prod Action (Prod Real State)","missing":[],"search":"head banditrlproof.finitehorizonrl.rewardsteptrace.head the first sampled action, reward, and next state of a positive trace. definition compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_head","label":"measurable_head","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_head","description":"theorem measurable_head (remaining : Nat) : Measurable (head (Action := Action) (State := State) remaining)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-4a4f8606dd76","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9345,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_head (remaining : Nat) : Measurable (head (Action := Action) (State := State) remaining)","missing":[],"search":"measurable_head banditrlproof.finitehorizonrl.rewardsteptrace.measurable_head theorem measurable_head (remaining : nat) : measurable (head (action := action) (state := state) remaining) theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.headActionReward","label":"headActionReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.headActionReward","description":"The first sampled action/reward pair, with next state discarded.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-9f818314414d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9346,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def headActionReward (remaining : Nat) : RewardStepTrace Action State (remaining + 1) -> Prod Action Real","missing":[],"search":"headactionreward banditrlproof.finitehorizonrl.rewardsteptrace.headactionreward the first sampled action/reward pair, with next state discarded. definition compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headActionReward","label":"measurable_headActionReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headActionReward","description":"theorem measurable_headActionReward (remaining : Nat) : Measurable (headActionReward (Action := Action) (State := State) remaining)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-9f6a6637065b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9347,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:44"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_headActionReward (remaining : Nat) : Measurable (headActionReward (Action := Action) (State := State) remaining)","missing":[],"search":"measurable_headactionreward banditrlproof.finitehorizonrl.rewardsteptrace.measurable_headactionreward theorem measurable_headactionreward (remaining : nat) : measurable (headactionreward (action := action) (state := state) remaining) theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.headReward","label":"headReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.headReward","description":"The first actual sampled Real reward of a positive trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-2fa374f1a484","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9348,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def headReward (remaining : Nat) : RewardStepTrace Action State (remaining + 1) -> Real","missing":[],"search":"headreward banditrlproof.finitehorizonrl.rewardsteptrace.headreward the first actual sampled real reward of a positive trace. definition compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headReward","label":"measurable_headReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headReward","description":"theorem measurable_headReward (remaining : Nat) : Measurable (headReward (Action := Action) (State := State) remaining)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-2a4ae4fbf560","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9349,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:58"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_headReward (remaining : Nat) : Measurable (headReward (Action := Action) (State := State) remaining)","missing":[],"search":"measurable_headreward banditrlproof.finitehorizonrl.rewardsteptrace.measurable_headreward theorem measurable_headreward (remaining : nat) : measurable (headreward (action := action) (state := state) remaining) theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel","label":"actionRewardKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel","description":"The action/reward marginal of one generated stochastic MDP step.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-bf1344d40e3e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9350,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:69"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def actionRewardKernel (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel State (Prod Action Real)","missing":[],"search":"actionrewardkernel banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardkernel the action/reward marginal of one generated stochastic mdp step. definition compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardMarginalKernel","label":"rewardMarginalKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardMarginalKernel","description":"The reward-only marginal after mixing the selected laws over policy actions.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-f9903a5f00b0","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9351,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:86"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardMarginalKernel (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel State Real","missing":[],"search":"rewardmarginalkernel banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewardmarginalkernel the reward-only marginal after mixing the selected laws over policy actions. definition compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_head","label":"stochasticTrajectoryKernelRemaining_map_head","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_head","description":"The first generated coordinate has exactly the compiled one-step law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-70352bdd129e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9352,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:102"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_map_head (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state).map (RewardStepTrace.head (Action := Action) (State := State) remaining) = source.actionRewardStateKernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state","missing":[],"search":"stochastictrajectorykernelremaining_map_head banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_map_head the first generated coordinate has exactly the compiled one-step law. theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headActionReward","label":"stochasticTrajectoryKernelRemaining_map_headActionReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headActionReward","description":"Mapping a generated trace to its first action/reward pair gives the joint marginal.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-cd8cea22f1e3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9353,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:124"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_map_headActionReward (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state).map (RewardStepTrace.headActionReward (Action := Action) (State := State) remaining) = source.actionRewardKernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state","missing":[],"search":"stochastictrajectorykernelremaining_map_headactionreward banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_map_headactionreward mapping a generated trace to its first action/reward pair gives the joint marginal. theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headReward","label":"stochasticTrajectoryKernelRemaining_map_headReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headReward","description":"Mapping a generated trace to its first reward gives the policy reward mixture.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-7e9db9dbe6ea","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9354,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_map_headReward (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state).map (RewardStepTrace.headReward (Action := Action) (State := State) remaining) = source.rewardMarginalKernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state","missing":[],"search":"stochastictrajectorykernelremaining_map_headreward banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_map_headreward mapping a generated trace to its first reward gives the policy reward mixture. theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_apply_prod","label":"actionRewardKernel_apply_prod","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_apply_prod","description":"Exact action/reward rectangle law for a policy-mixed stochastic step.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-850be0a0bd63","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9355,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:187"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardKernel_apply_prod (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) {actionSet : Set Action} {rewardSet : Set Real} (hactionSet : MeasurableSet actionSet) (hrewardSet : MeasurableSet rewardSet) : source.actionRewardKernel policy stage state (actionSet ×ˢ rewardSet) = ∫⁻ action in actionSet, source.rewardKernel.kernel (state, action) rewardSet ∂(policy.actionKernel stage state)","missing":[],"search":"actionrewardkernel_apply_prod banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardkernel_apply_prod exact action/reward rectangle law for a policy-mixed stochastic step. theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardMarginalKernel_apply","label":"rewardMarginalKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardMarginalKernel_apply","description":"Reward-event probability is the selected reward law mixed over policy actions.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-29c0bbc9339c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9356,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:217"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem rewardMarginalKernel_apply (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) (state : State) {rewardSet : Set Real} (hrewardSet : MeasurableSet rewardSet) : source.rewardMarginalKernel policy stage state rewardSet = ∫⁻ action, source.rewardKernel.kernel (state, action) rewardSet ∂(policy.actionKernel stage state)","missing":[],"search":"rewardmarginalkernel_apply banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewardmarginalkernel_apply reward-event probability is the selected reward law mixed over policy actions. theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headMarginalFactorization","label":"stochasticTrajectoryKernelRemaining_headMarginalFactorization","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headMarginalFactorization","description":"Route endpoint: the generated first reward event has the exact randomized policy mixture law, while the joint action/reward rectangle retains the selected-law factorization.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardmarginal/index.html#decl-cd2e28ef6ced","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","order":9357,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardMarginal.lean:243"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_headMarginalFactorization (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) {actionSet : Set Action} {rewardSet : Set Real} (hactionSet : MeasurableSet actionSet) (hrewardSet : MeasurableSet rewardSet) : (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state) ((RewardStepTrace.headActionReward (Action := Action) (State := State) remaining) ⁻¹' (actionSet ×ˢ rewardSet)) = ∫⁻ action in actionSet, source.rewardKernel.kernel (state, action) rewardSet ∂(policy.actionKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ state) ∧ (source.stochasticTrajectoryKernelRemaining policy (remaining + 1) hremaining state) ((RewardStepTrace.headReward (Action := Action) (State := State) remaining) ⁻¹' rewardSet) = ∫⁻ action, source.rewardKe…","missing":[],"search":"stochastictrajectorykernelremaining_headmarginalfactorization banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_headmarginalfactorization route endpoint: the generated first reward event has the exact randomized policy mixture law, while the joint action/reward rectangle retains the selected-law factorization. theorem compiled","shard":"modules/6ecdd848334c317a.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:ProbabilityTheory.Kernel.compProd_prodMkRight_eq_prod","label":"compProd_prodMkRight_eq_prod","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.Kernel.compProd_prodMkRight_eq_prod","description":"Sampling from `kappa` and then from a kernel that ignores the sampled value is the same as sampling from the product kernel at the original input.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-5fa3c260ce37","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9358,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:27"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem compProd_prodMkRight_eq_prod (kappa : Kernel Alpha Beta) (eta : Kernel Alpha Gamma) [IsFiniteKernel kappa] [IsFiniteKernel eta] : kappa ⊗ₖ prodMkRight Beta eta = kappa ×ₖ eta","missing":[],"search":"compprod_prodmkright_eq_prod probabilitytheory.kernel.compprod_prodmkright_eq_prod sampling from `kappa` and then from a kernel that ignores the sampled value is the same as sampling from the product kernel at the original input. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:ProbabilityTheory.Kernel.map_compProd_prodMk","label":"map_compProd_prodMk","kind":"theorem","status":"compiled","subtitle":"ProbabilityTheory.Kernel.map_compProd_prodMk","description":"Push a first-coordinate-dependent map through the second stage of `compProd`.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-2ef08fde6851","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9359,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:39"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem map_compProd_prodMk {Delta : Type*} [MeasurableSpace Delta] (kappa : Kernel Alpha Beta) (eta : Kernel (Alpha × Beta) Gamma) (eta' : Kernel (Alpha × Beta) Delta) [IsSFiniteKernel kappa] [IsSFiniteKernel eta] [IsSFiniteKernel eta'] (g : Beta -> Gamma -> Delta) (hg : Measurable g.uncurry) (hmap : forall a b, (eta (a, b)).map (g b) = eta' (a, b)) : (kappa ⊗ₖ eta).map (fun p => (p.1, g p.1 p.2)) = kappa ⊗ₖ eta'","missing":[],"search":"map_compprod_prodmk probabilitytheory.kernel.map_compprod_prodmk push a first-coordinate-dependent map through the second stage of `compprod`. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.retainedActionStateKernel","label":"retainedActionStateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.retainedActionStateKernel","description":"Retain the current state together with one sampled action/next-state pair.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-67d7cc1f6d30","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9360,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:79"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def retainedActionStateKernel {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel State (State × (Action × State))","missing":[],"search":"retainedactionstatekernel banditrlproof.finitehorizonrl.markovpolicy.retainedactionstatekernel retain the current state together with one sampled action/next-state pair. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationFrom","label":"sampledCumulativeReturnDeviationFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationFrom","description":"Actual sampled cumulative return centered at the recursive policy value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-a0c7e0bdd7ad","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9361,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:98"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledCumulativeReturnDeviationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (trace : RewardStepTrace Action State remaining) : Real","missing":[],"search":"sampledcumulativereturndeviationfrom banditrlproof.finitehorizonrl.mdp.sampledcumulativereturndeviationfrom actual sampled cumulative return centered at the recursive policy value. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationFrom","label":"measurable_sampledCumulativeReturnDeviationFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationFrom","description":"The sampled-return deviation is jointly measurable in start state and trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-7197022f7495","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9362,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:106"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeReturnDeviationFrom (mdp : MDP State Action) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) : Measurable (fun p : State × RewardStepTrace Action State remaining => mdp.sampledCumulativeReturnDeviationFrom policy remaining hremaining p.1 p.2)","missing":[],"search":"measurable_sampledcumulativereturndeviationfrom banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativereturndeviationfrom the sampled-return deviation is jointly measurable in start state and trace. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationFrom_eq_rewardDeviation_add_meanBellmanInnovation","label":"sampledCumulativeReturnDeviationFrom_eq_rewardDeviation_add_meanBellmanInnovation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationFrom_eq_rewardDeviation_add_meanBellmanInnovation","description":"The centered sampled return splits pathwise into selected-reward noise and the mean Bellman innovation. No probabilistic independence is used in this identity.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-43cd0ad6c42c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9363,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:120"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem sampledCumulativeReturnDeviationFrom_eq_rewardDeviation_add_meanBellmanInnovation (mdp : MDP State Action) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (trace : RewardStepTrace Action State remaining) : mdp.sampledCumulativeReturnDeviationFrom policy remaining hremaining state trace = mdp.sampledCumulativeRewardDeviationFrom remaining state trace + mdp.sampledCumulativeMeanBellmanInnovationFrom policy remaining hremaining state trace","missing":[],"search":"sampledcumulativereturndeviationfrom_eq_rewarddeviation_add_meanbellmaninnovation banditrlproof.finitehorizonrl.mdp.sampledcumulativereturndeviationfrom_eq_rewarddeviation_add_meanbellmaninnovation the centered sampled return splits pathwise into selected-reward noise and the mean bellman innovation. no probabilistic independence is used in this identity. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationAfterActionState","label":"rewardDeviationAfterActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationAfterActionState","description":"Center a sampled reward using the retained state/action coordinates.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-c1301333d62d","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9364,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:147"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def rewardDeviationAfterActionState (mdp : MDP State Action) : ((State × (Action × State)) × Real) -> Real","missing":[],"search":"rewarddeviationafteractionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewarddeviationafteractionstate center a sampled reward using the retained state/action coordinates. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_rewardDeviationAfterActionState","label":"measurable_rewardDeviationAfterActionState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_rewardDeviationAfterActionState","description":"theorem measurable_rewardDeviationAfterActionState (mdp : MDP State Action) : Measurable (rewardDeviationAfterActionState mdp)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-a060f0f92508","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9365,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:151"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_rewardDeviationAfterActionState (mdp : MDP State Action) : Measurable (rewardDeviationAfterActionState mdp)","missing":[],"search":"measurable_rewarddeviationafteractionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_rewarddeviationafteractionstate theorem measurable_rewarddeviationafteractionstate (mdp : mdp state action) : measurable (rewarddeviationafteractionstate mdp) theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rawRewardKernelAfterActionState","label":"rawRewardKernelAfterActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rawRewardKernelAfterActionState","description":"Raw reward sampled from the retained current-state/action coordinates.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-23502f45684f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9366,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:158"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rawRewardKernelAfterActionState (source : MeanCompatibleRewardKernel mdp) : ProbabilityTheory.Kernel (State × (Action × State)) Real","missing":[],"search":"rawrewardkernelafteractionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rawrewardkernelafteractionstate raw reward sampled from the retained current-state/action coordinates. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState","label":"rewardDeviationKernelAfterActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState","description":"Reward residual sampled after retaining the current state and an already sampled action/next-state pair.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-b74e7e3e9f4c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9367,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:175"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardDeviationKernelAfterActionState (source : MeanCompatibleRewardKernel mdp) : ProbabilityTheory.Kernel (State × (Action × State)) Real","missing":[],"search":"rewarddeviationkernelafteractionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewarddeviationkernelafteractionstate reward residual sampled after retaining the current state and an already sampled action/next-state pair. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState_apply","label":"rewardDeviationKernelAfterActionState_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState_apply","description":"Pointwise law of the retained-state centered reward kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-7b0d2cd7fb3f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9368,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:191"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem rewardDeviationKernelAfterActionState_apply (source : MeanCompatibleRewardKernel mdp) (p : State × (Action × State)) : source.rewardDeviationKernelAfterActionState p = (source.rewardKernel.kernel (p.1, p.2.1)).map (fun reward => reward - mdp.reward p.1 p.2.1)","missing":[],"search":"rewarddeviationkernelafteractionstate_apply banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewarddeviationkernelafteractionstate_apply pointwise law of the retained-state centered reward kernel. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState_hasSubgaussianMGF","label":"rewardDeviationKernelAfterActionState_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState_hasSubgaussianMGF","description":"The retained-state reward residual kernel inherits the uniform reward MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-885227c42c87","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9369,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:209"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem rewardDeviationKernelAfterActionState_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] (source : MeanCompatibleRewardKernel mdp) (varianceProxy : NNReal) (law : source.UniformSubgaussianRewardLaw varianceProxy) (base : Measure (State × (Action × State))) [IsFiniteMeasure base] : ProbabilityTheory.Kernel.HasSubgaussianMGF id varianceProxy source.rewardDeviationKernelAfterActionState base","missing":[],"search":"rewarddeviationkernelafteractionstate_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewarddeviationkernelafteractionstate_hassubgaussianmgf the retained-state reward residual kernel inherits the uniform reward mgf. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rawRewardKernelAfterRetainedActionState","label":"rawRewardKernelAfterRetainedActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rawRewardKernelAfterRetainedActionState","description":"Ignore the outer duplicated state and sample the raw retained-coordinate reward.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-6d60e82a0750","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9370,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:236"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rawRewardKernelAfterRetainedActionState (source : MeanCompatibleRewardKernel mdp) : ProbabilityTheory.Kernel (State × (State × (Action × State))) Real","missing":[],"search":"rawrewardkernelafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rawrewardkernelafterretainedactionstate ignore the outer duplicated state and sample the raw retained-coordinate reward. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterRetainedActionState","label":"rewardDeviationKernelAfterRetainedActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterRetainedActionState","description":"Ignore the outer duplicated state and sample the centered reward residual.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-0f66844d54f0","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9371,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:243"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def rewardDeviationKernelAfterRetainedActionState (source : MeanCompatibleRewardKernel mdp) : ProbabilityTheory.Kernel (State × (State × (Action × State))) Real","missing":[],"search":"rewarddeviationkernelafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewarddeviationkernelafterretainedactionstate ignore the outer duplicated state and sample the centered reward residual. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.uncenterRewardAfterRetainedActionState","label":"uncenterRewardAfterRetainedActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.uncenterRewardAfterRetainedActionState","description":"Add the retained MDP mean back to a centered reward residual.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-462b7a7e0271","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9372,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:264"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def uncenterRewardAfterRetainedActionState (mdp : MDP State Action) : State × (Action × State) -> Real -> Real","missing":[],"search":"uncenterrewardafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.uncenterrewardafterretainedactionstate add the retained mdp mean back to a centered reward residual. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_uncenterRewardAfterRetainedActionState","label":"measurable_uncenterRewardAfterRetainedActionState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_uncenterRewardAfterRetainedActionState","description":"theorem measurable_uncenterRewardAfterRetainedActionState (mdp : MDP State Action) : Measurable (uncenterRewardAfterRetainedActionState mdp).uncurry","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-7a750a9703f2","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9373,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:268"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_uncenterRewardAfterRetainedActionState (mdp : MDP State Action) : Measurable (uncenterRewardAfterRetainedActionState mdp).uncurry","missing":[],"search":"measurable_uncenterrewardafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_uncenterrewardafterretainedactionstate theorem measurable_uncenterrewardafterretainedactionstate (mdp : mdp state action) : measurable (uncenterrewardafterretainedactionstate mdp).uncurry theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterRetainedActionState_map_uncenter","label":"rewardDeviationKernelAfterRetainedActionState_map_uncenter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterRetainedActionState_map_uncenter","description":"Centering and then restoring the retained MDP mean recovers the raw reward law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-d3ddf20977ae","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9374,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:276"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem rewardDeviationKernelAfterRetainedActionState_map_uncenter (source : MeanCompatibleRewardKernel mdp) (outer : State) (retained : State × (Action × State)) : (source.rewardDeviationKernelAfterRetainedActionState (outer, retained)).map (uncenterRewardAfterRetainedActionState mdp retained) = source.rawRewardKernelAfterRetainedActionState (outer, retained)","missing":[],"search":"rewarddeviationkernelafterretainedactionstate_map_uncenter banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.rewarddeviationkernelafterretainedactionstate_map_uncenter centering and then restoring the retained mdp mean recovers the raw reward law. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rewardDeviation_map_uncenter","label":"retainedActionStateKernel_compProd_rewardDeviation_map_uncenter","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rewardDeviation_map_uncenter","description":"Restore raw rewards throughout the retained action-state product kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-d858c8f46198","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9375,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:300"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem retainedActionStateKernel_compProd_rewardDeviation_map_uncenter (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ((policy.retainedActionStateKernel stage) ⊗ₖ source.rewardDeviationKernelAfterRetainedActionState).map (fun p => (p.1, uncenterRewardAfterRetainedActionState mdp p.1 p.2)) = (policy.retainedActionStateKernel stage) ⊗ₖ source.rawRewardKernelAfterRetainedActionState","missing":[],"search":"retainedactionstatekernel_compprod_rewarddeviation_map_uncenter banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.retainedactionstatekernel_compprod_rewarddeviation_map_uncenter restore raw rewards throughout the retained action-state product kernel. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.assembleRawRewardAfterRetainedActionState","label":"assembleRawRewardAfterRetainedActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.assembleRawRewardAfterRetainedActionState","description":"Arrange retained current/action/next-state coordinates with a raw reward.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-b930f551397b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9376,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:320"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def assembleRawRewardAfterRetainedActionState : ((State × (Action × State)) × Real) -> Action × (Real × State)","missing":[],"search":"assemblerawrewardafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.assemblerawrewardafterretainedactionstate arrange retained current/action/next-state coordinates with a raw reward. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_assembleRawRewardAfterRetainedActionState","label":"measurable_assembleRawRewardAfterRetainedActionState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_assembleRawRewardAfterRetainedActionState","description":"theorem measurable_assembleRawRewardAfterRetainedActionState : Measurable (assembleRawRewardAfterRetainedActionState (State := State) (Action := Action))","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-1de1eff8e921","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9377,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:326"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_assembleRawRewardAfterRetainedActionState : Measurable (assembleRawRewardAfterRetainedActionState (State := State) (Action := Action))","missing":[],"search":"measurable_assemblerawrewardafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_assemblerawrewardafterretainedactionstate theorem measurable_assemblerawrewardafterretainedactionstate : measurable (assemblerawrewardafterretainedactionstate (state := state) (action := action)) theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rawReward_map_assemble","label":"retainedActionStateKernel_compProd_rawReward_map_assemble","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rawReward_map_assemble","description":"The retained raw-reward construction is exactly the existing head kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-1eaaedfbb5a6","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9378,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:334"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem retainedActionStateKernel_compProd_rawReward_map_assemble (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ((policy.retainedActionStateKernel stage) ⊗ₖ source.rawRewardKernelAfterRetainedActionState).map (assembleRawRewardAfterRetainedActionState (State := State) (Action := Action)) = source.actionRewardStateKernel policy stage","missing":[],"search":"retainedactionstatekernel_compprod_rawreward_map_assemble banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.retainedactionstatekernel_compprod_rawreward_map_assemble the retained raw-reward construction is exactly the existing head kernel. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.assembleRewardDeviationAfterRetainedActionState","label":"assembleRewardDeviationAfterRetainedActionState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.assembleRewardDeviationAfterRetainedActionState","description":"Assemble a reward-bearing head directly from a centered residual.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-50a420db5c7e","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9379,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:408"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def assembleRewardDeviationAfterRetainedActionState (mdp : MDP State Action) : ((State × (Action × State)) × Real) -> Action × (Real × State)","missing":[],"search":"assemblerewarddeviationafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.assemblerewarddeviationafterretainedactionstate assemble a reward-bearing head directly from a centered residual. definition compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_assembleRewardDeviationAfterRetainedActionState","label":"measurable_assembleRewardDeviationAfterRetainedActionState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_assembleRewardDeviationAfterRetainedActionState","description":"theorem measurable_assembleRewardDeviationAfterRetainedActionState (mdp : MDP State Action) : Measurable (assembleRewardDeviationAfterRetainedActionState mdp)","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-bf61dba7619f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9380,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:414"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_assembleRewardDeviationAfterRetainedActionState (mdp : MDP State Action) : Measurable (assembleRewardDeviationAfterRetainedActionState mdp)","missing":[],"search":"measurable_assemblerewarddeviationafterretainedactionstate banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_assemblerewarddeviationafterretainedactionstate theorem measurable_assemblerewarddeviationafterretainedactionstate (mdp : mdp state action) : measurable (assemblerewarddeviationafterretainedactionstate mdp) theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rewardDeviation_map_assemble","label":"retainedActionStateKernel_compProd_rewardDeviation_map_assemble","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rewardDeviation_map_assemble","description":"The centered residual construction maps exactly to the existing head kernel.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-ab960f2c2969","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9381,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:423"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem retainedActionStateKernel_compProd_rewardDeviation_map_assemble (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ((policy.retainedActionStateKernel stage) ⊗ₖ source.rewardDeviationKernelAfterRetainedActionState).map (assembleRewardDeviationAfterRetainedActionState mdp) = source.actionRewardStateKernel policy stage","missing":[],"search":"retainedactionstatekernel_compprod_rewarddeviation_map_assemble banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.retainedactionstatekernel_compprod_rewarddeviation_map_assemble the centered residual construction maps exactly to the existing head kernel. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","label":"retainedActionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","description":"Retaining the current state preserves the one-step Bellman innovation MGF.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-b52b3f0a4817","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9382,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:451"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem retainedActionStateKernel_meanBellmanInnovation_hasSubgaussianMGF (policy : MarkovPolicy mdp) (rewardBound : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : ProbabilityTheory.HasSubgaussianMGF (fun retained : State × (Action × State) => mdp.reward retained.1 retained.2.1 + policy.valueRemaining remaining (by omega) retained.2.2 - policy.valueRemaining (remaining + 1) hremaining state) (meanBellmanInnovationStepVarianceProxy rewardBound (remaining + 1)) (policy.retainedActionStateKernel ⟨mdp.horizon - (remaining + 1), by omega⟩ state)","missing":[],"search":"retainedactionstatekernel_meanbellmaninnovation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.retainedactionstatekernel_meanbellmaninnovation_hassubgaussianmgf retaining the current state preserves the one-step bellman innovation mgf. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_sampledReturnBellmanInnovation_hasSubgaussianMGF","label":"actionRewardStateKernel_sampledReturnBellmanInnovation_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_sampledReturnBellmanInnovation_hasSubgaussianMGF","description":"One sampled reward plus the sampled next policy value is sub-Gaussian around the current policy value with the sum of reward and Bellman innovation proxies.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-adf0bccb14e4","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9383,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:513"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem actionRewardStateKernel_sampledReturnBellmanInnovation_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) (remaining : Nat) (hremaining : remaining + 1 <= mdp.horizon) (state : State) : ProbabilityTheory.HasSubgaussianMGF (fun head : Action × (Real × State) => head.2.1 + policy.valueRemaining remaining (by omega) head.2.2 - policy.valueRemaining (remaining + 1) hremaining state) (meanBellmanInnovationStepVarianceProxy rewardBound (remaining + 1) + rewardVarianceProxy) (source.actionRewardStateKernel policy ⟨mdp.horizon - (remaining + 1), by omega⟩ state)","missing":[],"search":"actionrewardstatekernel_sampledreturnbellmaninnovation_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardstatekernel_sampledreturnbellmaninnovation_hassubgaussianmgf one sampled reward plus the sampled next policy value is sub-gaussian around the current policy value with the sum of reward and bellman innovation proxies. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_hasSubgaussianMGF","label":"stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_hasSubgaussianMGF","description":"The full sampled return is sub-Gaussian around the recursive policy value, with the reward-noise proxy plus the mean Bellman innovation proxy.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-53f713c1182f","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9384,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:622"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_hasSubgaussianMGF [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : ProbabilityTheory.HasSubgaussianMGF (mdp.sampledCumulativeReturnDeviationFrom policy remaining hremaining state) ((remaining : NNReal) * rewardVarianceProxy + meanBellmanInnovationVarianceProxy rewardBound remaining) (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state)","missing":[],"search":"stochastictrajectorykernelremaining_sampledcumulativereturndeviationfrom_hassubgaussianmgf banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_sampledcumulativereturndeviationfrom_hassubgaussianmgf the full sampled return is sub-gaussian around the recursive policy value, with the reward-noise proxy plus the mean bellman innovation proxy. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_abs_tail_le","label":"stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_abs_tail_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_abs_tail_le","description":"Fixed-horizon two-sided delta tail for sampled return around policy value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtotalreturnconcentration/index.html#decl-8194f7a631f3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","order":9385,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTotalReturnConcentration.lean:877"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_abs_tail_le [StandardBorelSpace State] [StandardBorelSpace Action] [Nonempty Action] (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (rewardBound rewardVarianceProxy : NNReal) (hrewardBound : forall state action, |mdp.reward state action| <= (rewardBound : Real)) (law : source.UniformSubgaussianRewardLaw rewardVarianceProxy) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) (htotal : 0 < ((((remaining : NNReal) * rewardVarianceProxy + meanBellmanInnovationVarianceProxy rewardBound remaining : NNReal) : Real))) (delta : Real) (hdelta : 0 < delta) (hdelta_le_one : delta <= 1) : (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state) {trace | Concentration.subGaussianSumConfidenceRadius ((remaining : NNReal) * rewardVarianceProxy + meanBellma…","missing":[],"search":"stochastictrajectorykernelremaining_sampledcumulativereturndeviationfrom_abs_tail_le banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining_sampledcumulativereturndeviationfrom_abs_tail_le fixed-horizon two-sided delta tail for sampled return around policy value. theorem compiled","shard":"modules/c1517f0e64986a8f.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace","label":"RewardStepTrace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace","description":"A finite trace of sampled action, reward, and resulting next state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-a8001828c153","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9386,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:26"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev RewardStepTrace (Action : Type v) (State : Type u) (n : Nat)","missing":[],"search":"rewardsteptrace banditrlproof.finitehorizonrl.rewardsteptrace a finite trace of sampled action, reward, and resulting next state. abbreviation compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_cons","label":"measurable_cons","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_cons","description":"Prepending a sampled action-reward-state coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-82d8ca9af4ee","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9387,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:33"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cons (n : Nat) : Measurable (fun p : Prod (Prod Action (Prod Real State)) (RewardStepTrace Action State n) => @Fin.cons n (fun _ => Prod Action (Prod Real State)) p.1 p.2)","missing":[],"search":"measurable_cons banditrlproof.finitehorizonrl.rewardsteptrace.measurable_cons prepending a sampled action-reward-state coordinate is measurable. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_tail","label":"measurable_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_tail","description":"Removing the first sampled coordinate is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-2f0e2e2daff3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9388,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:45"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_tail (n : Nat) : Measurable (fun trace : RewardStepTrace Action State (n + 1) => Fin.tail trace)","missing":[],"search":"measurable_tail banditrlproof.finitehorizonrl.rewardsteptrace.measurable_tail removing the first sampled coordinate is measurable. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.integrable_of_fintype_aestronglyMeasurable","label":"integrable_of_fintype_aestronglyMeasurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.integrable_of_fintype_aestronglyMeasurable","description":"An a.e. strongly measurable Real function on a finite type is integrable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-5e695a1109ea","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9389,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:55"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_of_fintype_aestronglyMeasurable {Omega : Type*} [Fintype Omega] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (f : Omega -> Real) (hf : AEStronglyMeasurable f mu) : Integrable f mu","missing":[],"search":"integrable_of_fintype_aestronglymeasurable banditrlproof.finitehorizonrl.integrable_of_fintype_aestronglymeasurable an a.e. strongly measurable real function on a finite type is integrable. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel","label":"actionRewardStateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel","description":"One policy action followed by its sampled reward and next state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-e3d5200cf084","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9390,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:72"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def actionRewardStateKernel (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel State (Prod Action (Prod Real State))","missing":[],"search":"actionrewardstatekernel banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.actionrewardstatekernel one policy action followed by its sampled reward and next state. definition compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining","label":"stochasticTrajectoryKernelRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining","description":"Kernel of the next `remaining` sampled action-reward-state coordinates. The first chronological stage is `mdp.horizon - remaining`.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-7569f3b8994c","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9391,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:92"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticTrajectoryKernelRemaining (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) : (remaining : Nat) -> remaining <= mdp.horizon -> ProbabilityTheory.Kernel State (RewardStepTrace Action State remaining) | 0, _ => ProbabilityTheory.Kernel.deterministic (fun _ => fun i => Fin.elim0 i) measurable_const | remaining + 1, hremaining => let stage : Fin mdp.horizon := ⟨mdp.horizon - (remaining + 1), by omega⟩ let tailKernel : ProbabilityTheory.Kernel (Prod State (Prod Action (Prod Real State))) (RewardStepTrace Action State remaining)","missing":[],"search":"stochastictrajectorykernelremaining banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorykernelremaining kernel of the next `remaining` sampled action-reward-state coordinates. the first chronological stage is `mdp.horizon - remaining`. definition compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardFrom","label":"sampledCumulativeRewardFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardFrom","description":"Sum of the actual sampled rewards in a reward-bearing finite trace.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-69c38097f65a","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9392,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:134"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def sampledCumulativeRewardFrom : (remaining : Nat) -> RewardStepTrace Action State remaining -> Real | 0, _ => 0 | remaining + 1, trace => (trace 0).2.1 + sampledCumulativeRewardFrom remaining (Fin.tail trace) omit [Fintype State] [Fintype Action] in /-- The sampled finite cumulative reward is measurable. -/ theorem measurable_sampledCumulativeRewardFrom (remaining : Nat) : Measurable (sampledCumulativeRewardFrom (Action := Action) (State := State) remaining)","missing":[],"search":"sampledcumulativerewardfrom banditrlproof.finitehorizonrl.mdp.sampledcumulativerewardfrom sum of the actual sampled rewards in a reward-bearing finite trace. definition compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardFrom","label":"measurable_sampledCumulativeRewardFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardFrom","description":"The sampled finite cumulative reward is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-923ebd4ef0a3","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9393,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:142"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeRewardFrom (remaining : Nat) : Measurable (sampledCumulativeRewardFrom (Action := Action) (State := State) remaining)","missing":[],"search":"measurable_sampledcumulativerewardfrom banditrlproof.finitehorizonrl.mdp.measurable_sampledcumulativerewardfrom the sampled finite cumulative reward is measurable. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining","label":"integrable_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining","description":"The sampled cumulative reward is `L1` under every statewise stochastic trajectory law. No boundedness or second-moment hypothesis is used.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-3aba284101da","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9394,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:165"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : Integrable (MDP.sampledCumulativeRewardFrom (Action := Action) (State := State) remaining) (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state)","missing":[],"search":"integrable_sampledcumulativerewardfrom_stochastictrajectorykernelremaining banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.integrable_sampledcumulativerewardfrom_stochastictrajectorykernelremaining the sampled cumulative reward is `l1` under every statewise stochastic trajectory law. no boundedness or second-moment hypothesis is used. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining_eq_stochasticValueRemaining","label":"integral_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining_eq_stochasticValueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining_eq_stochasticValueRemaining","description":"Statewise stochastic trajectory identity: expected sampled cumulative reward equals the independently defined stochastic backward policy value.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-2fe67564d8b7","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9395,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:312"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining_eq_stochasticValueRemaining (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : integral (source.stochasticTrajectoryKernelRemaining policy remaining hremaining state) (MDP.sampledCumulativeRewardFrom (Action := Action) (State := State) remaining) = source.stochasticValueRemaining policy remaining hremaining state","missing":[],"search":"integral_sampledcumulativerewardfrom_stochastictrajectorykernelremaining_eq_stochasticvalueremaining banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.integral_sampledcumulativerewardfrom_stochastictrajectorykernelremaining_eq_stochasticvalueremaining statewise stochastic trajectory identity: expected sampled cumulative reward equals the independently defined stochastic backward policy value. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure","label":"stochasticTrajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure","description":"Full reward-bearing trajectory law, including the initial state.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-b7e28b4c8dfa","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9396,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:400"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def stochasticTrajectoryMeasure (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) : Measure (Prod State (RewardStepTrace Action State mdp.horizon))","missing":[],"search":"stochastictrajectorymeasure banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.stochastictrajectorymeasure full reward-bearing trajectory law, including the initial state. definition compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledCumulativeReward","label":"sampledCumulativeReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledCumulativeReward","description":"Sampled cumulative reward on the full stochastic trajectory.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-1cb622eeee0b","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9397,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:416"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def sampledCumulativeReward (trajectory : Prod State (RewardStepTrace Action State mdp.horizon)) : Real","missing":[],"search":"sampledcumulativereward banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.sampledcumulativereward sampled cumulative reward on the full stochastic trajectory. definition compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_sampledCumulativeReward","label":"measurable_sampledCumulativeReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_sampledCumulativeReward","description":"The full sampled cumulative reward is measurable.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-92439f588f45","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9398,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:421"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledCumulativeReward : Measurable (sampledCumulativeReward (mdp := mdp))","missing":[],"search":"measurable_sampledcumulativereward banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.measurable_sampledcumulativereward the full sampled cumulative reward is measurable. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_sampledCumulativeReward_stochasticTrajectoryMeasure","label":"integrable_sampledCumulativeReward_stochasticTrajectoryMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_sampledCumulativeReward_stochasticTrajectoryMeasure","description":"The full sampled cumulative reward is integrable under its generated law.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-ecac064a0f44","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9399,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:426"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledCumulativeReward_stochasticTrajectoryMeasure (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : Integrable (sampledCumulativeReward (mdp := mdp)) (source.stochasticTrajectoryMeasure policy initialState)","missing":[],"search":"integrable_sampledcumulativereward_stochastictrajectorymeasure banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.integrable_sampledcumulativereward_stochastictrajectorymeasure the full sampled cumulative reward is integrable under its generated law. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_sampledCumulativeReward_stochasticTrajectoryMeasure_eq_integral_valueAt_zero","label":"integral_sampledCumulativeReward_stochasticTrajectoryMeasure_eq_integral_valueAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_sampledCumulativeReward_stochasticTrajectoryMeasure_eq_integral_valueAt_zero","description":"Route endpoint: expected sampled cumulative reward equals the stochastic and existing mean policy values at chronological stage zero.","url":"../modules/banditrlproof-rl-finitehorizonstochasticrewardtrajectory/index.html#decl-9b6ed4a372ef","parent":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","order":9400,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonStochasticRewardTrajectory.lean:455"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledCumulativeReward_stochasticTrajectoryMeasure_eq_integral_valueAt_zero (source : MeanCompatibleRewardKernel mdp) (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : integral (source.stochasticTrajectoryMeasure policy initialState) (sampledCumulativeReward (mdp := mdp)) = integral initialState (source.stochasticValueAt policy 0 (Nat.zero_le mdp.horizon)) ∧ integral (source.stochasticTrajectoryMeasure policy initialState) (sampledCumulativeReward (mdp := mdp)) = integral initialState (policy.valueAt 0 (Nat.zero_le mdp.horizon))","missing":[],"search":"integral_sampledcumulativereward_stochastictrajectorymeasure_eq_integral_valueat_zero banditrlproof.finitehorizonrl.mdp.meancompatiblerewardkernel.integral_sampledcumulativereward_stochastictrajectorymeasure_eq_integral_valueat_zero route endpoint: expected sampled cumulative reward equals the stochastic and existing mean policy values at chronological stage zero. theorem compiled","shard":"modules/876b9bd0cd8904e2.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace","label":"StepTrace","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace","description":"The finite sequence of sampled actions and their resulting next states.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-9001bd339ba9","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9401,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:25"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"abbrev StepTrace (Action : Type v) (State : Type u) (n : Nat)","missing":[],"search":"steptrace banditrlproof.finitehorizonrl.steptrace the finite sequence of sampled actions and their resulting next states. abbreviation compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_cons","label":"measurable_cons","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.measurable_cons","description":"Prepending one action-state step is measurable for the product Pi-space.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-ce94123cc6ea","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9402,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:32"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cons (n : Nat) : Measurable (fun p : Prod (Prod Action State) (StepTrace Action State n) => @Fin.cons n (fun _ => Prod Action State) p.1 p.2)","missing":[],"search":"measurable_cons banditrlproof.finitehorizonrl.steptrace.measurable_cons prepending one action-state step is measurable for the product pi-space. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_tail","label":"measurable_tail","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.StepTrace.measurable_tail","description":"Removing the first action-state step is measurable.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-ec8ebead27e8","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9403,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:43"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_tail (n : Nat) : Measurable (fun trace : StepTrace Action State (n + 1) => Fin.tail trace)","missing":[],"search":"measurable_tail banditrlproof.finitehorizonrl.steptrace.measurable_tail removing the first action-state step is measurable. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.integrable_of_fintype","label":"integrable_of_fintype","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.integrable_of_fintype","description":"A measurable Real function on a finite type is integrable under every finite measure.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-84b93ebebd95","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9404,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:53"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_of_fintype {Omega : Type*} [Fintype Omega] [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (f : Omega -> Real) (hf : Measurable f) : Integrable f mu","missing":[],"search":"integrable_of_fintype banditrlproof.finitehorizonrl.integrable_of_fintype a measurable real function on a finite type is integrable under every finite measure. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel","label":"actionStateKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel","description":"One chronological policy step, retaining both the action and next state.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-364994def77b","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9405,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:67"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def actionStateKernel {mdp : MDP State Action} (policy : MarkovPolicy mdp) (stage : Fin mdp.horizon) : ProbabilityTheory.Kernel State (Prod Action State)","missing":[],"search":"actionstatekernel banditrlproof.finitehorizonrl.markovpolicy.actionstatekernel one chronological policy step, retaining both the action and next state. definition compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining","label":"trajectoryKernelRemaining","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining","description":"Kernel of the next `remaining` action-state steps. Its first chronological stage is `mdp.horizon - remaining`.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-1568412fc314","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9406,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:85"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def trajectoryKernelRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) : (remaining : Nat) -> remaining <= mdp.horizon -> ProbabilityTheory.Kernel State (StepTrace Action State remaining) | 0, _ => ProbabilityTheory.Kernel.deterministic (fun _ => fun i => Fin.elim0 i) measurable_const | remaining + 1, hremaining => let stage : Fin mdp.horizon := ⟨mdp.horizon - (remaining + 1), by omega⟩ let tailKernel : ProbabilityTheory.Kernel (Prod State (Prod Action State)) (StepTrace Action State remaining)","missing":[],"search":"trajectorykernelremaining banditrlproof.finitehorizonrl.markovpolicy.trajectorykernelremaining kernel of the next `remaining` action-state steps. its first chronological stage is `mdp.horizon - remaining`. definition compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.cumulativeRewardFrom","label":"cumulativeRewardFrom","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.cumulativeRewardFrom","description":"Total reward of a finite trace, starting from the supplied current state.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-4a771adf6654","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9407,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:125"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeRewardFrom (mdp : MDP State Action) : (remaining : Nat) -> State -> StepTrace Action State remaining -> Real | 0, _, _ => 0 | remaining + 1, state, trace => mdp.reward state (trace 0).1 + mdp.cumulativeRewardFrom remaining (trace 0).2 (Fin.tail trace) /-- The finite cumulative reward is measurable jointly in its initial state and trace. -/ theorem measurable_cumulativeRewardFrom (mdp : MDP State Action) (remaining : Nat) : Measurable (fun p : Prod State (StepTrace Action State remaining) => mdp.cumulativeRewardFrom remaining p.1 p.2)","missing":[],"search":"cumulativerewardfrom banditrlproof.finitehorizonrl.mdp.cumulativerewardfrom total reward of a finite trace, starting from the supplied current state. definition compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_cumulativeRewardFrom","label":"measurable_cumulativeRewardFrom","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_cumulativeRewardFrom","description":"The finite cumulative reward is measurable jointly in its initial state and trace.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-e8d800e4fa26","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9408,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:133"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeRewardFrom (mdp : MDP State Action) (remaining : Nat) : Measurable (fun p : Prod State (StepTrace Action State remaining) => mdp.cumulativeRewardFrom remaining p.1 p.2)","missing":[],"search":"measurable_cumulativerewardfrom banditrlproof.finitehorizonrl.mdp.measurable_cumulativerewardfrom the finite cumulative reward is measurable jointly in its initial state and trace. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integrable_cumulativeRewardFrom_trajectoryKernelRemaining","label":"integrable_cumulativeRewardFrom_trajectoryKernelRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integrable_cumulativeRewardFrom_trajectoryKernelRemaining","description":"Every statewise finite-trace cumulative reward is integrable under its trajectory law.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-ef5acc3eaa77","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9409,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:157"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_cumulativeRewardFrom_trajectoryKernelRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : Integrable (mdp.cumulativeRewardFrom remaining state) (policy.trajectoryKernelRemaining remaining hremaining state)","missing":[],"search":"integrable_cumulativerewardfrom_trajectorykernelremaining banditrlproof.finitehorizonrl.markovpolicy.integrable_cumulativerewardfrom_trajectorykernelremaining every statewise finite-trace cumulative reward is integrable under its trajectory law. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeRewardFrom_trajectoryKernelRemaining_eq_valueRemaining","label":"integral_cumulativeRewardFrom_trajectoryKernelRemaining_eq_valueRemaining","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeRewardFrom_trajectoryKernelRemaining_eq_valueRemaining","description":"Statewise policy-evaluation identity: integrating the finite generated return over the remaining trajectory gives the backward policy value.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-7b2632c6c1d6","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9410,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:171"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_cumulativeRewardFrom_trajectoryKernelRemaining_eq_valueRemaining {mdp : MDP State Action} (policy : MarkovPolicy mdp) (remaining : Nat) (hremaining : remaining <= mdp.horizon) (state : State) : (∫ trace, mdp.cumulativeRewardFrom remaining state trace ∂policy.trajectoryKernelRemaining remaining hremaining state) = policy.valueRemaining remaining hremaining state","missing":[],"search":"integral_cumulativerewardfrom_trajectorykernelremaining_eq_valueremaining banditrlproof.finitehorizonrl.markovpolicy.integral_cumulativerewardfrom_trajectorykernelremaining_eq_valueremaining statewise policy-evaluation identity: integrating the finite generated return over the remaining trajectory gives the backward policy value. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure","label":"trajectoryMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure","description":"Joint law of the initial state and all finite action-state steps.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-8ca8abb31559","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9411,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:263"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def trajectoryMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) : Measure (Prod State (StepTrace Action State mdp.horizon))","missing":[],"search":"trajectorymeasure banditrlproof.finitehorizonrl.markovpolicy.trajectorymeasure joint law of the initial state and all finite action-state steps. definition compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.cumulativeReward","label":"cumulativeReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.cumulativeReward","description":"Cumulative reward on the full finite policy trajectory.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-7a01d25ccd70","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9412,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:282"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"def cumulativeReward (mdp : MDP State Action) (trajectory : Prod State (StepTrace Action State mdp.horizon)) : Real","missing":[],"search":"cumulativereward banditrlproof.finitehorizonrl.mdp.cumulativereward cumulative reward on the full finite policy trajectory. definition compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_cumulativeReward","label":"measurable_cumulativeReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MDP.measurable_cumulativeReward","description":"The full finite-trajectory cumulative reward is measurable.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-6774d697a37e","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9413,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:287"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_cumulativeReward (mdp : MDP State Action) : Measurable mdp.cumulativeReward","missing":[],"search":"measurable_cumulativereward banditrlproof.finitehorizonrl.mdp.measurable_cumulativereward the full finite-trajectory cumulative reward is measurable. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integrable_cumulativeReward_trajectoryMeasure","label":"integrable_cumulativeReward_trajectoryMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integrable_cumulativeReward_trajectoryMeasure","description":"The cumulative reward is automatically integrable under the finite trajectory law.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-def6628394b3","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9414,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:296"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integrable_cumulativeReward_trajectoryMeasure {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : Integrable mdp.cumulativeReward (policy.trajectoryMeasure initialState)","missing":[],"search":"integrable_cumulativereward_trajectorymeasure banditrlproof.finitehorizonrl.markovpolicy.integrable_cumulativereward_trajectorymeasure the cumulative reward is automatically integrable under the finite trajectory law. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeReward_trajectoryMeasure_eq_integral_valueAt_zero","label":"integral_cumulativeReward_trajectoryMeasure_eq_integral_valueAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeReward_trajectoryMeasure_eq_integral_valueAt_zero","description":"Route endpoint: expected cumulative reward under the generated finite policy trajectory equals the initial-state expectation of the stage-zero policy value.","url":"../modules/banditrlproof-rl-finitehorizontrajectory/index.html#decl-644cd5a8b792","parent":"module:BanditRLProof.RL.FiniteHorizonTrajectory","order":9415,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.FiniteHorizonTrajectory"],["Source","BanditRLProof/RL/FiniteHorizonTrajectory.lean:307"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem integral_cumulativeReward_trajectoryMeasure_eq_integral_valueAt_zero {mdp : MDP State Action} (policy : MarkovPolicy mdp) (initialState : Measure State) [IsProbabilityMeasure initialState] : (∫ trajectory, mdp.cumulativeReward trajectory ∂policy.trajectoryMeasure initialState) = ∫ state, policy.valueAt 0 (Nat.zero_le mdp.horizon) state ∂initialState","missing":[],"search":"integral_cumulativereward_trajectorymeasure_eq_integral_valueat_zero banditrlproof.finitehorizonrl.markovpolicy.integral_cumulativereward_trajectorymeasure_eq_integral_valueat_zero route endpoint: expected cumulative reward under the generated finite policy trajectory equals the initial-state expectation of the stage-zero policy value. theorem compiled","shard":"modules/1a1bc34fd7cc0065.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","description":"The maximum of the six literal capped/uncapped stopped-return errors at one schedule index.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-b9ac50013ddb","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9416,"meta":[["Kind","definition"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:34"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"noncomputable def selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t) → Real","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror the maximum of the six literal capped/uncapped stopped-return errors at one schedule index. definition compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_nonneg","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_nonneg","description":"The scalar joint error is nonnegative.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-423d3a2e67ce","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9417,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:95"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_nonneg (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : 0 ≤ selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror_nonneg banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror_nonneg the scalar joint error is nonnegative. theorem compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","label":"measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","description":"Every schedule-index coordinate of the scalar joint error is measurable.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-de2856a00419","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9418,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:116"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor : Real) (scheduleIndex : Nat) : Measurable (selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex)","missing":[],"search":"measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.measurable_selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror every schedule-index coordinate of the scalar joint error is measurable. theorem compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_lt_iff","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_lt_iff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_lt_iff","description":"A strict scalar joint-error bound is exactly the conjunction of the same six literal stopped-return bounds.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-1a0d912102ee","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9419,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:179"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_lt_iff (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) (trajectory : HeterogeneousStochasticEpisodeBatchTrajectory mdp (fun t => AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes mdp varianceProxy baseVisitFloor t)) : let cappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor let uncappedStoppingPrefix := selfConsistentScheduledNaturalCausalInverseSqrtT…","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror_lt_iff banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointerror_lt_iff a strict scalar joint-error bound is exactly the conjunction of the same six literal stopped-return bounds. theorem compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_eq_jointError_ge","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_eq_jointError_ge","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_eq_jointError_ge","description":"The accepted named weak-bad event is exactly the scalar joint-error weak superlevel set.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-838ba4d8813d","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9420,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:242"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_eq_jointError_ge (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon scheduleIndex = {trajectory | epsilon ≤ selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory}","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset_eq_jointerror_ge banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointviolationset_eq_jointerror_ge the accepted named weak-bad event is exactly the scalar joint-error weak superlevel set. theorem compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_eq_jointError_lt","label":"selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_eq_jointError_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_eq_jointError_lt","description":"The accepted named strict-good event is exactly the scalar joint-error strict sublevel set.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-2762b2879637","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9421,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:266"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_eq_jointError_lt (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] (rewardSource : mdp.MeanCompatibleRewardKernel) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (varianceProxy : NNReal) (baseVisitFloor epsilon : Real) (scheduleIndex : Nat) : selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor epsilon scheduleIndex = {trajectory | selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError mdp initialState rewardSource initialTable defaultState varianceProxy baseVisitFloor scheduleIndex trajectory < epsilon}","missing":[],"search":"selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset_eq_jointerror_lt banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentschedulednaturalcausalinversesqrtthresholdcappedunboundedhittingafterstoppedreturnjointgoodset_eq_jointerror_lt the accepted named strict-good event is exactly the scalar joint-error strict sublevel set. theorem compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_jointError_deterministicTailHighProbability_optimality","label":"selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_jointError_deterministicTailHighProbability_optimality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_jointError_deterministicTailHighProbability_optimality","description":"Terminal scalar deterministic-tail confidence certificate. One noncomputable cutoff works for every later schedule index.","url":"../modules/banditrlproof-rl-stoppedreturnjointerrordeterministictailhighprobability/index.html#decl-4991a606617a","parent":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","order":9422,"meta":[["Kind","theorem"],["Module","BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability"],["Source","BanditRLProof/RL/StoppedReturnJointErrorDeterministicTailHighProbability.lean:293"],["Chapter","Finite-horizon RL"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:finite-horizon-rl"],["Indexed settings","None registered"]],"statement":"theorem selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_jointError_deterministicTailHighProbability_optimality (mdp : MDP State Action) (initialState : Measure State) [IsProbabilityMeasure initialState] [StandardBorelSpace State] [StandardBorelSpace Action] (rewardSource : mdp.MeanCompatibleRewardKernel) (varianceProxy : NNReal) (hvarianceProxy : 0 < varianceProxy) (law : rewardSource.UniformSubgaussianRewardLaw varianceProxy) (initialTable : DeterministicMarkovPolicyTable mdp) (defaultState : State) (support : ExploratoryPathSupport mdp initialState) (baseVisitFloor : Real) (hbaseFloor : ExploratoryPathUniformVisitFloor support 1 baseVisitFloor) (hrewardBound : forall state action, |mdp.reward state action| <= 1) (hhorizon : 4 < mdp.horizon) (hbaseVisitFloor : 0 < baseVisitFloor) : let source :=…","missing":[],"search":"selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_jointerror_deterministictailhighprobability_optimality banditrlproof.finitehorizonrl.adaptivestochasticsampledempiricaloptimisticsource.selfconsistentscheduledcausalsource_inversesqrtthresholdcappedunboundedhittingafter_stoppedsampledreturn_and_successorpolicyexpectedreturn_jointerror_deterministictailhighprobability_optimality terminal scalar deterministic-tail confidence certificate. one noncomputable cutoff works for every later schedule index. theorem compiled","shard":"modules/34c63cc214810d21.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:finite-horizon-rl"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_rat_div_const","label":"measurable_rat_div_const","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_rat_div_const","description":"Division by a fixed rational is measurable on a measurable singleton Rat space. This is the `RAT-MEASURABLE-DIV-CONST-OF-MEASURABLE-SINGLETON` wrapper. It is kept separate from the ETC empirical-mean theorem so later leaves can decide whether to consume it or keep an explicit division contract.","url":"../modules/banditrlproof-ratmeasurability/index.html#decl-653dd2abc521","parent":"module:BanditRLProof.RatMeasurability","order":9423,"meta":[["Kind","theorem"],["Module","BanditRLProof.RatMeasurability"],["Source","BanditRLProof/RatMeasurability.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem measurable_rat_div_const [MeasurableSpace Rat] [MeasurableSingletonClass Rat] (c : Rat) : Measurable (fun x : Rat => x / c)","missing":[],"search":"measurable_rat_div_const banditrlproof.measurable_rat_div_const division by a fixed rational is measurable on a measurable singleton rat space. this is the `rat-measurable-div-const-of-measurable-singleton` wrapper. it is kept separate from the etc empirical-mean theorem so later leaves can decide whether to consume it or keep an explicit division contract. theorem compiled","shard":"modules/ec5dd601bec2a62a.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realKernelMean","label":"realKernelMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.realKernelMean","description":"Identity-integral mean of a Real-valued arm kernel.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-b271e30a0a43","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9424,"meta":[["Kind","definition"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:20"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def realKernelMean {K : Nat} (nu : ProbabilityTheory.Kernel (Fin K) Real) (a : Fin K) : Real","missing":[],"search":"realkernelmean banditrlproof.realkernelmean identity-integral mean of a real-valued arm kernel. definition compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realKernelGap","label":"realKernelGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.realKernelGap","description":"Gap of a Real-valued arm kernel from its supremum arm mean.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-157a951fd45b","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9425,"meta":[["Kind","definition"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def realKernelGap {K : Nat} (nu : ProbabilityTheory.Kernel (Fin K) Real) (a : Fin K) : Real","missing":[],"search":"realkernelgap banditrlproof.realkernelgap gap of a real-valued arm kernel from its supremum arm mean. definition compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realKernelRegret","label":"realKernelRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.realKernelRegret","description":"Finite-horizon regret of an action trace against a Real-valued arm kernel.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-bab64ce3559c","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9426,"meta":[["Kind","definition"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:30"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def realKernelRegret {K : Nat} (nu : ProbabilityTheory.Kernel (Fin K) Real) (action : ActionTrace (Fin K)) (n : Nat) : Real","missing":[],"search":"realkernelregret banditrlproof.realkernelregret finite-horizon regret of an action trace against a real-valued arm kernel. definition compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realKernelGap_nonneg","label":"realKernelGap_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.realKernelGap_nonneg","description":"Every kernel arm gap is nonnegative when the finite arm type is nonempty.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-1fbd2866f93f","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9427,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:36"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem realKernelGap_nonneg {K : Nat} [Nonempty (Fin K)] (nu : ProbabilityTheory.Kernel (Fin K) Real) (a : Fin K) : 0 <= realKernelGap nu a","missing":[],"search":"realkernelgap_nonneg banditrlproof.realkernelgap_nonneg every kernel arm gap is nonnegative when the finite arm type is nonempty. theorem compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realKernelRegret_eq_finset_sum_gap","label":"realKernelRegret_eq_finset_sum_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.realKernelRegret_eq_finset_sum_gap","description":"Kernel regret is the finite time-indexed sum of selected kernel gaps.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-cb155c367e94","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9428,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:44"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem realKernelRegret_eq_finset_sum_gap {K : Nat} (nu : ProbabilityTheory.Kernel (Fin K) Real) (action : ActionTrace (Fin K)) (n : Nat) : realKernelRegret nu action n = (Finset.range n).sum (fun t => realKernelGap nu (action t))","missing":[],"search":"realkernelregret_eq_finset_sum_gap banditrlproof.realkernelregret_eq_finset_sum_gap kernel regret is the finite time-indexed sum of selected kernel gaps. theorem compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realKernelRegret_eq_sum_gap_mul_pullCount","label":"realKernelRegret_eq_sum_gap_mul_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.realKernelRegret_eq_sum_gap_mul_pullCount","description":"Kernel regret is the gap-weighted finite sum of arm pull counts.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-dc6bcbb99b5d","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9429,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:54"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem realKernelRegret_eq_sum_gap_mul_pullCount {K : Nat} (nu : ProbabilityTheory.Kernel (Fin K) Real) (action : ActionTrace (Fin K)) (n : Nat) : realKernelRegret nu action n = (Finset.univ : Finset (Fin K)).sum (fun a => realKernelGap nu a * (pullCount action a n : Real))","missing":[],"search":"realkernelregret_eq_sum_gap_mul_pullcount banditrlproof.realkernelregret_eq_sum_gap_mul_pullcount kernel regret is the gap-weighted finite sum of arm pull counts. theorem compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integrable_realKernelRegret_of_integrable_pullCount","label":"integrable_realKernelRegret_of_integrable_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integrable_realKernelRegret_of_integrable_pullCount","description":"Pull-count integrability implies integrability of Real kernel regret.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-729b409e3b78","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9430,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:65"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_realKernelRegret_of_integrable_pullCount {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : ProbabilityTheory.Kernel (Fin K) Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) (hcount : forall a : Fin K, Integrable (fun omega => (pullCount (action omega) a n : Real)) mu) : Integrable (fun omega => realKernelRegret nu (action omega) n) mu","missing":[],"search":"integrable_realkernelregret_of_integrable_pullcount banditrlproof.integrable_realkernelregret_of_integrable_pullcount pull-count integrability implies integrability of real kernel regret. theorem compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integral_realKernelRegret_eq_sum_gap_mul_integral_pullCount","label":"integral_realKernelRegret_eq_sum_gap_mul_integral_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_realKernelRegret_eq_sum_gap_mul_integral_pullCount","description":"The Bochner expectation of Real kernel regret is the gap-weighted sum of expected arm pull counts. This is the kernel-facing bookkeeping endpoint for the exact ETC route. It does not assume a Markov kernel, probability measure, identity integrability, algorithm/environment law, sub-Gaussian proxy, or argmax semantics. Those contracts are required only by downstream statistical theorems.","url":"../modules/banditrlproof-realkernelregretpullcount/index.html#decl-d1814fd6f5cb","parent":"module:BanditRLProof.RealKernelRegretPullCount","order":9431,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealKernelRegretPullCount"],["Source","BanditRLProof/RealKernelRegretPullCount.lean:88"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_realKernelRegret_eq_sum_gap_mul_integral_pullCount {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (nu : ProbabilityTheory.Kernel (Fin K) Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) (hcount : forall a : Fin K, Integrable (fun omega => (pullCount (action omega) a n : Real)) mu) : integral mu (fun omega => realKernelRegret nu (action omega) n) = (Finset.univ : Finset (Fin K)).sum (fun a => realKernelGap nu a * integral mu (fun omega => (pullCount (action omega) a n : Real)))","missing":[],"search":"integral_realkernelregret_eq_sum_gap_mul_integral_pullcount banditrlproof.integral_realkernelregret_eq_sum_gap_mul_integral_pullcount the bochner expectation of real kernel regret is the gap-weighted sum of expected arm pull counts. this is the kernel-facing bookkeeping endpoint for the exact etc route. it does not assume a markov kernel, probability measure, identity integrability, algorithm/environment law, sub-gaussian proxy, or argmax semantics. those contracts are required only by downstream statistical theorems. theorem compiled","shard":"modules/bd78f0f47c67d1d3.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realMeanGap","label":"realMeanGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.realMeanGap","description":"Gap from the supremum finite-arm mean, matching the scalar semantics of LML's bandit gap.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html#decl-6c07421c09a3","parent":"module:BanditRLProof.RealMeanRegretPullCount","order":9432,"meta":[["Kind","definition"],["Module","BanditRLProof.RealMeanRegretPullCount"],["Source","BanditRLProof/RealMeanRegretPullCount.lean:22"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def realMeanGap {K : Nat} (mean : Fin K -> Real) (a : Fin K) : Real","missing":[],"search":"realmeangap banditrlproof.realmeangap gap from the supremum finite-arm mean, matching the scalar semantics of lml's bandit gap. definition compiled","shard":"modules/a8b2b232507c2380.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realMeanRegret","label":"realMeanRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.realMeanRegret","description":"Real pseudo-regret written directly from arm means over a finite horizon.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html#decl-5bf6511c4880","parent":"module:BanditRLProof.RealMeanRegretPullCount","order":9433,"meta":[["Kind","definition"],["Module","BanditRLProof.RealMeanRegretPullCount"],["Source","BanditRLProof/RealMeanRegretPullCount.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def realMeanRegret {K : Nat} (mean : Fin K -> Real) (action : ActionTrace (Fin K)) (n : Nat) : Real","missing":[],"search":"realmeanregret banditrlproof.realmeanregret real pseudo-regret written directly from arm means over a finite horizon. definition compiled","shard":"modules/a8b2b232507c2380.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realMeanRegret_eq_finset_sum_gap","label":"realMeanRegret_eq_finset_sum_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.realMeanRegret_eq_finset_sum_gap","description":"Real mean regret is the time-indexed finite sum of selected arm gaps.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html#decl-a4a02dd8c96b","parent":"module:BanditRLProof.RealMeanRegretPullCount","order":9434,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealMeanRegretPullCount"],["Source","BanditRLProof/RealMeanRegretPullCount.lean:32"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem realMeanRegret_eq_finset_sum_gap {K : Nat} (mean : Fin K -> Real) (action : ActionTrace (Fin K)) (n : Nat) : realMeanRegret mean action n = (Finset.range n).sum (fun t => realMeanGap mean (action t))","missing":[],"search":"realmeanregret_eq_finset_sum_gap banditrlproof.realmeanregret_eq_finset_sum_gap real mean regret is the time-indexed finite sum of selected arm gaps. theorem compiled","shard":"modules/a8b2b232507c2380.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.realMeanRegret_eq_sum_gap_mul_pullCount","label":"realMeanRegret_eq_sum_gap_mul_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.realMeanRegret_eq_sum_gap_mul_pullCount","description":"Real mean regret decomposes into each arm gap times its finite-horizon pull count. This is the deterministic half of `REAL-MEAN-REGRET-PULLCOUNT`. The definition uses the same supremum-minus-mean gap as the exact LML theorem card; no rational model, kernel law, measurability, or concentration assumption appears here.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html#decl-e2efe312da40","parent":"module:BanditRLProof.RealMeanRegretPullCount","order":9435,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealMeanRegretPullCount"],["Source","BanditRLProof/RealMeanRegretPullCount.lean:48"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem realMeanRegret_eq_sum_gap_mul_pullCount {K : Nat} (mean : Fin K -> Real) (action : ActionTrace (Fin K)) (n : Nat) : realMeanRegret mean action n = (Finset.univ : Finset (Fin K)).sum (fun a => realMeanGap mean a * (pullCount action a n : Real))","missing":[],"search":"realmeanregret_eq_sum_gap_mul_pullcount banditrlproof.realmeanregret_eq_sum_gap_mul_pullcount real mean regret decomposes into each arm gap times its finite-horizon pull count. this is the deterministic half of `real-mean-regret-pullcount`. the definition uses the same supremum-minus-mean gap as the exact lml theorem card; no rational model, kernel law, measurability, or concentration assumption appears here. theorem compiled","shard":"modules/a8b2b232507c2380.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integrable_realMeanRegret_of_integrable_pullCount","label":"integrable_realMeanRegret_of_integrable_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integrable_realMeanRegret_of_integrable_pullCount","description":"Pull-count integrability implies integrability of Real mean regret.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html#decl-1ec394b3bc92","parent":"module:BanditRLProof.RealMeanRegretPullCount","order":9436,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealMeanRegretPullCount"],["Source","BanditRLProof/RealMeanRegretPullCount.lean:70"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integrable_realMeanRegret_of_integrable_pullCount {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (mean : Fin K -> Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) (hcount : forall a : Fin K, Integrable (fun omega => (pullCount (action omega) a n : Real)) mu) : Integrable (fun omega => realMeanRegret mean (action omega) n) mu","missing":[],"search":"integrable_realmeanregret_of_integrable_pullcount banditrlproof.integrable_realmeanregret_of_integrable_pullcount pull-count integrability implies integrability of real mean regret. theorem compiled","shard":"modules/a8b2b232507c2380.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.integral_realMeanRegret_eq_sum_gap_mul_integral_pullCount","label":"integral_realMeanRegret_eq_sum_gap_mul_integral_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_realMeanRegret_eq_sum_gap_mul_integral_pullCount","description":"The Bochner expectation of Real mean regret is the gap-weighted sum of expected pull counts. This is the expectation half of `REAL-MEAN-REGRET-PULLCOUNT`. Callers supply only pull-count integrability; probability-space, policy, reward-law, and sub-Gaussian contracts remain outside this bookkeeping leaf.","url":"../modules/banditrlproof-realmeanregretpullcount/index.html#decl-0ee98fffd2a3","parent":"module:BanditRLProof.RealMeanRegretPullCount","order":9437,"meta":[["Kind","theorem"],["Module","BanditRLProof.RealMeanRegretPullCount"],["Source","BanditRLProof/RealMeanRegretPullCount.lean:106"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem integral_realMeanRegret_eq_sum_gap_mul_integral_pullCount {Omega : Type u} {K : Nat} [MeasurableSpace Omega] (mu : Measure Omega) (mean : Fin K -> Real) (action : Omega -> ActionTrace (Fin K)) (n : Nat) (hcount : forall a : Fin K, Integrable (fun omega => (pullCount (action omega) a n : Real)) mu) : integral mu (fun omega => realMeanRegret mean (action omega) n) = (Finset.univ : Finset (Fin K)).sum (fun a => realMeanGap mean a * integral mu (fun omega => (pullCount (action omega) a n : Real)))","missing":[],"search":"integral_realmeanregret_eq_sum_gap_mul_integral_pullcount banditrlproof.integral_realmeanregret_eq_sum_gap_mul_integral_pullcount the bochner expectation of real mean regret is the gap-weighted sum of expected pull counts. this is the expectation half of `real-mean-regret-pullcount`. callers supply only pull-count integrability; probability-space, policy, reward-law, and sub-gaussian contracts remain outside this bookkeeping leaf. theorem compiled","shard":"modules/a8b2b232507c2380.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret","label":"pseudoRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.pseudoRegret","description":"Pseudo-regret accumulated from the model gaps along an action trace.","url":"../modules/banditrlproof-regret/index.html#decl-48c4dd14183f","parent":"module:BanditRLProof.Regret","order":9438,"meta":[["Kind","definition"],["Module","BanditRLProof.Regret"],["Source","BanditRLProof/Regret.lean:16"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"noncomputable def pseudoRegret (model : FiniteBanditModel K) (action : Nat → Fin K) : Nat → Rat | 0 => 0 | t + 1 => pseudoRegret model action t + model.gap (action t) @[simp] theorem pseudoRegret_zero (model : FiniteBanditModel K) (action : Nat → Fin K) : pseudoRegret model action 0 = 0","missing":[],"search":"pseudoregret banditrlproof.pseudoregret pseudo-regret accumulated from the model gaps along an action trace. definition compiled","shard":"modules/094980085c9c35a8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_zero","label":"pseudoRegret_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_zero","description":"@[simp] theorem pseudoRegret_zero (model : FiniteBanditModel K) (action : Nat → Fin K) : pseudoRegret model action 0 = 0","url":"../modules/banditrlproof-regret/index.html#decl-22f47ed99a0b","parent":"module:BanditRLProof.Regret","order":9439,"meta":[["Kind","theorem"],["Module","BanditRLProof.Regret"],["Source","BanditRLProof/Regret.lean:21"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pseudoRegret_zero (model : FiniteBanditModel K) (action : Nat → Fin K) : pseudoRegret model action 0 = 0","missing":[],"search":"pseudoregret_zero banditrlproof.pseudoregret_zero @[simp] theorem pseudoregret_zero (model : finitebanditmodel k) (action : nat → fin k) : pseudoregret model action 0 = 0 theorem compiled","shard":"modules/094980085c9c35a8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_succ","label":"pseudoRegret_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_succ","description":"@[simp] theorem pseudoRegret_succ (model : FiniteBanditModel K) (action : Nat → Fin K) (t : Nat) : pseudoRegret model action (t + 1) = pseudoRegret model action t + model.gap (action t)","url":"../modules/banditrlproof-regret/index.html#decl-df07a91295b1","parent":"module:BanditRLProof.Regret","order":9440,"meta":[["Kind","theorem"],["Module","BanditRLProof.Regret"],["Source","BanditRLProof/Regret.lean:25"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"@[simp] theorem pseudoRegret_succ (model : FiniteBanditModel K) (action : Nat → Fin K) (t : Nat) : pseudoRegret model action (t + 1) = pseudoRegret model action t + model.gap (action t)","missing":[],"search":"pseudoregret_succ banditrlproof.pseudoregret_succ @[simp] theorem pseudoregret_succ (model : finitebanditmodel k) (action : nat → fin k) (t : nat) : pseudoregret model action (t + 1) = pseudoregret model action t + model.gap (action t) theorem compiled","shard":"modules/094980085c9c35a8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.RegretBoundCard","label":"RegretBoundCard","kind":"structure","status":"compiled","subtitle":"BanditRLProof.RegretBoundCard","description":"A reusable record for theorem-card style regret bounds.","url":"../modules/banditrlproof-regret/index.html#decl-3c5bc53f2596","parent":"module:BanditRLProof.Regret","order":9441,"meta":[["Kind","structure"],["Module","BanditRLProof.Regret"],["Source","BanditRLProof/Regret.lean:31"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure RegretBoundCard where","missing":[],"search":"regretboundcard banditrlproof.regretboundcard a reusable record for theorem-card style regret bounds. structure compiled","shard":"modules/094980085c9c35a8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.RegretObligation","label":"RegretObligation","kind":"structure","status":"compiled","subtitle":"BanditRLProof.RegretObligation","description":"A proof obligation attached to a regret theorem.","url":"../modules/banditrlproof-regret/index.html#decl-afe45d42f78c","parent":"module:BanditRLProof.Regret","order":9442,"meta":[["Kind","structure"],["Module","BanditRLProof.Regret"],["Source","BanditRLProof/Regret.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"structure RegretObligation where","missing":[],"search":"regretobligation banditrlproof.regretobligation a proof obligation attached to a regret theorem. structure compiled","shard":"modules/094980085c9c35a8.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_le_finset_sum_gap_mul_count_bound","label":"pseudoRegret_le_finset_sum_gap_mul_count_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_le_finset_sum_gap_mul_count_bound","description":"If every arm's pull count is bounded by `B`, then pseudo-regret is bounded by the corresponding gap-weighted count budget. This is the `REGRET-COUNT-BOUND` deterministic scaffold. It consumes only the compiled regret decomposition and model-derived gap nonnegativity.","url":"../modules/banditrlproof-regretcountbounds/index.html#decl-5edac519788c","parent":"module:BanditRLProof.RegretCountBounds","order":9443,"meta":[["Kind","theorem"],["Module","BanditRLProof.RegretCountBounds"],["Source","BanditRLProof/RegretCountBounds.lean:26"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_le_finset_sum_gap_mul_count_bound {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n : Nat) (B : Fin K -> Rat) (hB : forall a : Fin K, ((pullCount action a n : Nat) : Rat) <= B a) : pseudoRegret model action n <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a * B a)","missing":[],"search":"pseudoregret_le_finset_sum_gap_mul_count_bound banditrlproof.pseudoregret_le_finset_sum_gap_mul_count_bound if every arm's pull count is bounded by `b`, then pseudo-regret is bounded by the corresponding gap-weighted count budget. this is the `regret-count-bound` deterministic scaffold. it consumes only the compiled regret decomposition and model-derived gap nonnegativity. theorem compiled","shard":"modules/a21ec45143e9825b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_le_finset_sum_gap_mul_nat_count_bound","label":"pseudoRegret_le_finset_sum_gap_mul_nat_count_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_le_finset_sum_gap_mul_nat_count_bound","description":"Nat-valued per-arm pull-count bounds imply the corresponding gap-weighted pseudo-regret bound after casting the count budget to `Rat`. This is the `REGRET-NAT-COUNT-BOUND` adapter. It is algorithm-neutral and keeps ETC/UCB-specific count facts out of this file.","url":"../modules/banditrlproof-regretcountbounds/index.html#decl-14cd0bd59d47","parent":"module:BanditRLProof.RegretCountBounds","order":9444,"meta":[["Kind","theorem"],["Module","BanditRLProof.RegretCountBounds"],["Source","BanditRLProof/RegretCountBounds.lean:52"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_le_finset_sum_gap_mul_nat_count_bound {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n : Nat) (B : Fin K -> Nat) (hB : forall a : Fin K, pullCount action a n <= B a) : pseudoRegret model action n <= (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a * (((B a : Nat) : Rat)))","missing":[],"search":"pseudoregret_le_finset_sum_gap_mul_nat_count_bound banditrlproof.pseudoregret_le_finset_sum_gap_mul_nat_count_bound nat-valued per-arm pull-count bounds imply the corresponding gap-weighted pseudo-regret bound after casting the count budget to `rat`. this is the `regret-nat-count-bound` adapter. it is algorithm-neutral and keeps etc/ucb-specific count facts out of this file. theorem compiled","shard":"modules/a21ec45143e9825b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_le_sum_gap_mul_uniform_nat_count_bound","label":"pseudoRegret_le_sum_gap_mul_uniform_nat_count_bound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_le_sum_gap_mul_uniform_nat_count_bound","description":"A uniform Nat-valued pull-count bound implies pseudo-regret is bounded by the sum of model gaps times that uniform count budget. This is the `REGRET-UNIFORM-NAT-COUNT-BOUND` adapter. It is still algorithm-neutral and does not prove any ETC/UCB-specific count fact.","url":"../modules/banditrlproof-regretcountbounds/index.html#decl-c462b6029c77","parent":"module:BanditRLProof.RegretCountBounds","order":9445,"meta":[["Kind","theorem"],["Module","BanditRLProof.RegretCountBounds"],["Source","BanditRLProof/RegretCountBounds.lean:80"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_le_sum_gap_mul_uniform_nat_count_bound {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n B : Nat) (hB : forall a : Fin K, pullCount action a n <= B) : pseudoRegret model action n <= ((Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a)) * (((B : Nat) : Rat))","missing":[],"search":"pseudoregret_le_sum_gap_mul_uniform_nat_count_bound banditrlproof.pseudoregret_le_sum_gap_mul_uniform_nat_count_bound a uniform nat-valued pull-count bound implies pseudo-regret is bounded by the sum of model gaps times that uniform count budget. this is the `regret-uniform-nat-count-bound` adapter. it is still algorithm-neutral and does not prove any etc/ucb-specific count fact. theorem compiled","shard":"modules/a21ec45143e9825b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","label":"pseudoRegret_eq_finset_sum_gap_mul_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","description":"Pseudo-regret decomposes into an arm-indexed sum of each arm gap multiplied by its pull count. This is the deterministic `REGRET-PULLCOUNT` bridge. It consumes the compiled `Finset.range` wrappers instead of reopening the recursive definitions of `pseudoRegret` or `pullCount`.","url":"../modules/banditrlproof-regretdecomposition/index.html#decl-c049d1b466b0","parent":"module:BanditRLProof.RegretDecomposition","order":9446,"meta":[["Kind","theorem"],["Module","BanditRLProof.RegretDecomposition"],["Source","BanditRLProof/RegretDecomposition.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem pseudoRegret_eq_finset_sum_gap_mul_pullCount : pseudoRegret model action t = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => model.gap a * (pullCount action a t : Rat))","missing":[],"search":"pseudoregret_eq_finset_sum_gap_mul_pullcount banditrlproof.pseudoregret_eq_finset_sum_gap_mul_pullcount pseudo-regret decomposes into an arm-indexed sum of each arm gap multiplied by its pull count. this is the deterministic `regret-pullcount` bridge. it consumes the compiled `finset.range` wrappers instead of reopening the recursive definitions of `pseudoregret` or `pullcount`. theorem compiled","shard":"modules/b064adda19239008.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.MarkovRewardKernel","label":"MarkovRewardKernel","kind":"structure","status":"compiled","subtitle":"BanditRLProof.RewardKernel.MarkovRewardKernel","description":"A reward distribution indexed by an arm/context object. The underlying object is Mathlib's `ProbabilityTheory.Kernel`; the local wrapper gives later bandit proof leaves a stable project-level name for the regularity contract.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-0afd0979ceac","parent":"module:BanditRLProof.RewardKernel","order":9447,"meta":[["Kind","structure"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:31"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure MarkovRewardKernel (Index : Type u) (Reward : Type v) [MeasurableSpace Index] [MeasurableSpace Reward] where","missing":[],"search":"markovrewardkernel banditrlproof.rewardkernel.markovrewardkernel a reward distribution indexed by an arm/context object. the underlying object is mathlib's `probabilitytheory.kernel`; the local wrapper gives later bandit proof leaves a stable project-level name for the regularity contract. structure compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.ofKernel","label":"ofKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.ofKernel","description":"Build the local reward-kernel contract from a Mathlib Markov kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-eaf6012df9e5","parent":"module:BanditRLProof.RewardKernel","order":9448,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:46"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def ofKernel (kernel : ProbabilityTheory.Kernel Index Reward) (hkernel : ProbabilityTheory.IsMarkovKernel kernel) : MarkovRewardKernel Index Reward where","missing":[],"search":"ofkernel banditrlproof.rewardkernel.ofkernel build the local reward-kernel contract from a mathlib markov kernel. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_kernel","label":"measurable_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_kernel","description":"The reward kernel is measurable as a map into measures.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-b93488a5540f","parent":"module:BanditRLProof.RewardKernel","order":9449,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:54"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_kernel (rewardKernel : MarkovRewardKernel Index Reward) : Measurable rewardKernel.kernel","missing":[],"search":"measurable_kernel banditrlproof.rewardkernel.measurable_kernel the reward kernel is measurable as a map into measures. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_apply_of_measurable_index","label":"measurable_apply_of_measurable_index","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_apply_of_measurable_index","description":"A measurable random index selects a measurable random reward measure.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ab6b56db1cee","parent":"module:BanditRLProof.RewardKernel","order":9450,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:60"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_apply_of_measurable_index {Omega : Type w} [MeasurableSpace Omega] (rewardKernel : MarkovRewardKernel Index Reward) (index : Omega -> Index) (hindex : Measurable index) : Measurable (fun omega : Omega => rewardKernel.kernel (index omega))","missing":[],"search":"measurable_apply_of_measurable_index banditrlproof.rewardkernel.measurable_apply_of_measurable_index a measurable random index selects a measurable random reward measure. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_eventProbability_of_measurable_index","label":"measurable_eventProbability_of_measurable_index","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_eventProbability_of_measurable_index","description":"For every measurable reward event, the selected event probability is a measurable scalar function of the random index.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-a78a3ad19969","parent":"module:BanditRLProof.RewardKernel","order":9451,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:72"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_eventProbability_of_measurable_index {Omega : Type w} [MeasurableSpace Omega] (rewardKernel : MarkovRewardKernel Index Reward) (index : Omega -> Index) (hindex : Measurable index) {event : Set Reward} (hevent : MeasurableSet event) : Measurable (fun omega : Omega => rewardKernel.kernel (index omega) event)","missing":[],"search":"measurable_eventprobability_of_measurable_index banditrlproof.rewardkernel.measurable_eventprobability_of_measurable_index for every measurable reward event, the selected event probability is a measurable scalar function of the random index. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isProbabilityMeasure_apply","label":"isProbabilityMeasure_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isProbabilityMeasure_apply","description":"Every measure selected by a reward kernel is a probability measure.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-984ae6b32cd4","parent":"module:BanditRLProof.RewardKernel","order":9452,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:85"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isProbabilityMeasure_apply (rewardKernel : MarkovRewardKernel Index Reward) (index : Index) : IsProbabilityMeasure (rewardKernel.kernel index)","missing":[],"search":"isprobabilitymeasure_apply banditrlproof.rewardkernel.isprobabilitymeasure_apply every measure selected by a reward kernel is a probability measure. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.apply_univ","label":"apply_univ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.apply_univ","description":"theorem apply_univ (rewardKernel : MarkovRewardKernel Index Reward) (index : Index) : rewardKernel.kernel index Set.univ = 1","url":"../modules/banditrlproof-rewardkernel/index.html#decl-c7ecae82adef","parent":"module:BanditRLProof.RewardKernel","order":9453,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:94"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem apply_univ (rewardKernel : MarkovRewardKernel Index Reward) (index : Index) : rewardKernel.kernel index Set.univ = 1","missing":[],"search":"apply_univ banditrlproof.rewardkernel.apply_univ theorem apply_univ (rewardkernel : markovrewardkernel index reward) (index : index) : rewardkernel.kernel index set.univ = 1 theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.const","label":"const","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.const","description":"Constant probability reward law as a reward-kernel contract.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-2c114d5d89dc","parent":"module:BanditRLProof.RewardKernel","order":9454,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:103"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def const (mu : Measure Reward) (hmu : IsProbabilityMeasure mu) : MarkovRewardKernel Index Reward where","missing":[],"search":"const banditrlproof.rewardkernel.const constant probability reward law as a reward-kernel contract. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.deterministic","label":"deterministic","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.deterministic","description":"Deterministic measurable reward law as a reward-kernel contract.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-8d4aecd8562e","parent":"module:BanditRLProof.RewardKernel","order":9455,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:113"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def deterministic (reward : Index -> Reward) (hreward : Measurable reward) : MarkovRewardKernel Index Reward where","missing":[],"search":"deterministic banditrlproof.rewardkernel.deterministic deterministic measurable reward law as a reward-kernel contract. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.selectedMeasure","label":"selectedMeasure","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.selectedMeasure","description":"Select the reward measure associated with a context/action pair. This is only the one-step reward-law lookup; it is not a trajectory-law construction.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-068fe7137c9d","parent":"module:BanditRLProof.RewardKernel","order":9456,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:127"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def selectedMeasure (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (context : Context) (action : Action) : Measure Reward","missing":[],"search":"selectedmeasure banditrlproof.rewardkernel.selectedmeasure select the reward measure associated with a context/action pair. this is only the one-step reward-law lookup; it is not a trajectory-law construction. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.contextIndependentOfActionLaws","label":"contextIndependentOfActionLaws","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.contextIndependentOfActionLaws","description":"Build a context-independent reward kernel from action-indexed probability laws. Countability and measurable singletons make the action-to-measure map measurable; pulling it back along `Prod.snd` leaves the context unrestricted.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-3c8aca3601b0","parent":"module:BanditRLProof.RewardKernel","order":9457,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:138"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def contextIndependentOfActionLaws [MeasurableSingletonClass Action] [Countable Action] (actionLaw : Action -> Measure Reward) (hprob : forall action, IsProbabilityMeasure (actionLaw action)) : MarkovRewardKernel (Context × Action) Reward where","missing":[],"search":"contextindependentofactionlaws banditrlproof.rewardkernel.contextindependentofactionlaws build a context-independent reward kernel from action-indexed probability laws. countability and measurable singletons make the action-to-measure map measurable; pulling it back along `prod.snd` leaves the context unrestricted. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.selectedMeasure_contextIndependentOfActionLaws","label":"selectedMeasure_contextIndependentOfActionLaws","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.selectedMeasure_contextIndependentOfActionLaws","description":"A context-independent reward kernel selects the original action law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e21b50518cba","parent":"module:BanditRLProof.RewardKernel","order":9458,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:153"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem selectedMeasure_contextIndependentOfActionLaws [MeasurableSingletonClass Action] [Countable Action] (actionLaw : Action -> Measure Reward) (hprob : forall action, IsProbabilityMeasure (actionLaw action)) (context : Context) (action : Action) : selectedMeasure (contextIndependentOfActionLaws actionLaw hprob) context action = actionLaw action","missing":[],"search":"selectedmeasure_contextindependentofactionlaws banditrlproof.rewardkernel.selectedmeasure_contextindependentofactionlaws a context-independent reward kernel selects the original action law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isProbabilityMeasure_selectedMeasure","label":"isProbabilityMeasure_selectedMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isProbabilityMeasure_selectedMeasure","description":"A context/action-selected reward measure is a probability measure.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-5df0321a27e7","parent":"module:BanditRLProof.RewardKernel","order":9459,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:164"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isProbabilityMeasure_selectedMeasure (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (context : Context) (action : Action) : IsProbabilityMeasure (selectedMeasure rewardKernel context action)","missing":[],"search":"isprobabilitymeasure_selectedmeasure banditrlproof.rewardkernel.isprobabilitymeasure_selectedmeasure a context/action-selected reward measure is a probability measure. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.selectedMeasure_univ","label":"selectedMeasure_univ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.selectedMeasure_univ","description":"theorem selectedMeasure_univ (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (context : Context) (action : Action) : selectedMeasure rewardKernel context action Set.univ = 1","url":"../modules/banditrlproof-rewardkernel/index.html#decl-d22f615af9eb","parent":"module:BanditRLProof.RewardKernel","order":9460,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:173"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem selectedMeasure_univ (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (context : Context) (action : Action) : selectedMeasure rewardKernel context action Set.univ = 1","missing":[],"search":"selectedmeasure_univ banditrlproof.rewardkernel.selectedmeasure_univ theorem selectedmeasure_univ (rewardkernel : markovrewardkernel (context × action) reward) (context : context) (action : action) : selectedmeasure rewardkernel context action set.univ = 1 theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_selectedMeasure_of_measurable","label":"measurable_selectedMeasure_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_selectedMeasure_of_measurable","description":"Measurable context and action random variables select a measurable random reward measure from the context/action reward kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-872a1ab80561","parent":"module:BanditRLProof.RewardKernel","order":9461,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:183"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedMeasure_of_measurable {Omega : Type w} [MeasurableSpace Omega] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (context : Omega -> Context) (action : Omega -> Action) (hcontext : Measurable context) (haction : Measurable action) : Measurable (fun omega : Omega => selectedMeasure rewardKernel (context omega) (action omega))","missing":[],"search":"measurable_selectedmeasure_of_measurable banditrlproof.rewardkernel.measurable_selectedmeasure_of_measurable measurable context and action random variables select a measurable random reward measure from the context/action reward kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_selectedEventProbability_of_measurable","label":"measurable_selectedEventProbability_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_selectedEventProbability_of_measurable","description":"For a measurable reward event, measurable context and action random variables select a measurable event-probability scalar from the reward kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-d2b600512756","parent":"module:BanditRLProof.RewardKernel","order":9462,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:202"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedEventProbability_of_measurable {Omega : Type w} [MeasurableSpace Omega] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (context : Omega -> Context) (action : Omega -> Action) (hcontext : Measurable context) (haction : Measurable action) {event : Set Reward} (hevent : MeasurableSet event) : Measurable (fun omega : Omega => selectedMeasure rewardKernel (context omega) (action omega) event)","missing":[],"search":"measurable_selectedeventprobability_of_measurable banditrlproof.rewardkernel.measurable_selectedeventprobability_of_measurable for a measurable reward event, measurable context and action random variables select a measurable event-probability scalar from the reward kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_selectedMeasure_of_policy_state","label":"measurable_selectedMeasure_of_policy_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_selectedMeasure_of_policy_state","description":"A measurable policy applied to a measurable state can be used as the action coordinate of a context/action reward kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e82dda6e7d6a","parent":"module:BanditRLProof.RewardKernel","order":9463,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:224"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedMeasure_of_policy_state {Omega : Type w} {State : Type u} [MeasurableSpace Omega] [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (context : Omega -> Context) (state : Omega -> State) (hcontext : Measurable context) (hstate : Measurable state) : Measurable (fun omega : Omega => selectedMeasure rewardKernel (context omega) (policy.action (state omega)))","missing":[],"search":"measurable_selectedmeasure_of_policy_state banditrlproof.rewardkernel.measurable_selectedmeasure_of_policy_state a measurable policy applied to a measurable state can be used as the action coordinate of a context/action reward kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_selectedEventProbability_of_policy_state","label":"measurable_selectedEventProbability_of_policy_state","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_selectedEventProbability_of_policy_state","description":"Event-probability version of `measurable_selectedMeasure_of_policy_state`.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-fc49de708d46","parent":"module:BanditRLProof.RewardKernel","order":9464,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:247"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_selectedEventProbability_of_policy_state {Omega : Type w} {State : Type u} [MeasurableSpace Omega] [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (context : Omega -> Context) (state : Omega -> State) (hcontext : Measurable context) (hstate : Measurable state) {event : Set Reward} (hevent : MeasurableSet event) : Measurable (fun omega : Omega => selectedMeasure rewardKernel (context omega) (policy.action (state omega)) event)","missing":[],"search":"measurable_selectedeventprobability_of_policy_state banditrlproof.rewardkernel.measurable_selectedeventprobability_of_policy_state event-probability version of `measurable_selectedmeasure_of_policy_state`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.policyContextStateIndex","label":"policyContextStateIndex","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.policyContextStateIndex","description":"The deterministic index map that turns a context/state pair into the context/action pair selected by a measurable policy.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ac879b13d34e","parent":"module:BanditRLProof.RewardKernel","order":9465,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:273"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def policyContextStateIndex {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) : Context × State -> Context × Action","missing":[],"search":"policycontextstateindex banditrlproof.rewardkernel.policycontextstateindex the deterministic index map that turns a context/state pair into the context/action pair selected by a measurable policy. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_policyContextStateIndex","label":"measurable_policyContextStateIndex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_policyContextStateIndex","description":"The policy-induced context/state index map is measurable.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-f03e1a8a1c9f","parent":"module:BanditRLProof.RewardKernel","order":9466,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:280"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_policyContextStateIndex {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) : Measurable (policyContextStateIndex (Context := Context) policy)","missing":[],"search":"measurable_policycontextstateindex banditrlproof.rewardkernel.measurable_policycontextstateindex the policy-induced context/state index map is measurable. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicy","label":"composePolicy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicy","description":"Compose a measurable policy with a context/action reward kernel to obtain the one-step reward kernel indexed by context/state pairs. This is a one-step `KERNEL-POLICY-BIND` precursor. It composes the policy map with reward-kernel lookup, but it does not yet build a finite-horizon or infinite-horizon trajectory law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-81c0e0f40486","parent":"module:BanditRLProof.RewardKernel","order":9467,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:294"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def composePolicy {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) : MarkovRewardKernel (Context × State) Reward where","missing":[],"search":"composepolicy banditrlproof.rewardkernel.composepolicy compose a measurable policy with a context/action reward kernel to obtain the one-step reward kernel indexed by context/state pairs. this is a one-step `kernel-policy-bind` precursor. it composes the policy map with reward-kernel lookup, but it does not yet build a finite-horizon or infinite-horizon trajectory law. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicy_kernel_apply","label":"composePolicy_kernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicy_kernel_apply","description":"theorem composePolicy_kernel_apply {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) : (composePolicy rewardKernel policy).kernel pair = selectedMeasure rewardKernel pair.1 (policy.action pair.2)","url":"../modules/banditrlproof-rewardkernel/index.html#decl-fd068e86c36a","parent":"module:BanditRLProof.RewardKernel","order":9468,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:315"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicy_kernel_apply {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) : (composePolicy rewardKernel policy).kernel pair = selectedMeasure rewardKernel pair.1 (policy.action pair.2)","missing":[],"search":"composepolicy_kernel_apply banditrlproof.rewardkernel.composepolicy_kernel_apply theorem composepolicy_kernel_apply {state : type u} [measurablespace state] (rewardkernel : markovrewardkernel (context × action) reward) (policy : policy.measurablepolicy state action) (pair : context × state) : (composepolicy rewardkernel policy).kernel pair = selectedmeasure rewardkernel pair.1 (policy.action pair.2) theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_composePolicy","label":"isMarkovKernel_composePolicy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_composePolicy","description":"The composed policy/reward kernel is a Mathlib Markov kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ed3989887476","parent":"module:BanditRLProof.RewardKernel","order":9469,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:325"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_composePolicy {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) : ProbabilityTheory.IsMarkovKernel (composePolicy rewardKernel policy).kernel","missing":[],"search":"ismarkovkernel_composepolicy banditrlproof.rewardkernel.ismarkovkernel_composepolicy the composed policy/reward kernel is a mathlib markov kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_composePolicy_eventProbability","label":"measurable_composePolicy_eventProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_composePolicy_eventProbability","description":"For a measurable reward event, the event probability under the composed policy/reward kernel is measurable in the context/state pair.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-5d1ed421dd9f","parent":"module:BanditRLProof.RewardKernel","order":9470,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:337"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_composePolicy_eventProbability {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) {event : Set Reward} (hevent : MeasurableSet event) : Measurable (fun pair : Context × State => (composePolicy rewardKernel policy).kernel pair event)","missing":[],"search":"measurable_composepolicy_eventprobability banditrlproof.rewardkernel.measurable_composepolicy_eventprobability for a measurable reward event, the event probability under the composed policy/reward kernel is measurable in the context/state pair. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.policyActionOfContextState","label":"policyActionOfContextState","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.policyActionOfContextState","description":"The policy-selected action as a deterministic map from context/state pairs. The context coordinate is carried so this map has the same source as the composed context/state reward kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-60bf932c87d8","parent":"module:BanditRLProof.RewardKernel","order":9471,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:358"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def policyActionOfContextState {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) : Context × State -> Action","missing":[],"search":"policyactionofcontextstate banditrlproof.rewardkernel.policyactionofcontextstate the policy-selected action as a deterministic map from context/state pairs. the context coordinate is carried so this map has the same source as the composed context/state reward kernel. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_policyActionOfContextState","label":"measurable_policyActionOfContextState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_policyActionOfContextState","description":"The policy-selected action map on context/state pairs is measurable.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ee5c4204f9b0","parent":"module:BanditRLProof.RewardKernel","order":9472,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:365"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_policyActionOfContextState {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) : Measurable (policyActionOfContextState (Context := Context) policy)","missing":[],"search":"measurable_policyactionofcontextstate banditrlproof.rewardkernel.measurable_policyactionofcontextstate the policy-selected action map on context/state pairs is measurable. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.policyActionKernel","label":"policyActionKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.policyActionKernel","description":"The deterministic action kernel induced by a measurable policy.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-b9fe05bf8dce","parent":"module:BanditRLProof.RewardKernel","order":9473,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:372"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def policyActionKernel {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) : ProbabilityTheory.Kernel (Context × State) Action","missing":[],"search":"policyactionkernel banditrlproof.rewardkernel.policyactionkernel the deterministic action kernel induced by a measurable policy. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.policyActionKernel_apply","label":"policyActionKernel_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.policyActionKernel_apply","description":"theorem policyActionKernel_apply {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) : policyActionKernel (Context := Context) policy pair = Measure.dirac (policy.action pair.2)","url":"../modules/banditrlproof-rewardkernel/index.html#decl-fe7281795202","parent":"module:BanditRLProof.RewardKernel","order":9474,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:381"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem policyActionKernel_apply {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) : policyActionKernel (Context := Context) policy pair = Measure.dirac (policy.action pair.2)","missing":[],"search":"policyactionkernel_apply banditrlproof.rewardkernel.policyactionkernel_apply theorem policyactionkernel_apply {state : type u} [measurablespace state] (policy : policy.measurablepolicy state action) (pair : context × state) : policyactionkernel (context := context) policy pair = measure.dirac (policy.action pair.2) theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_policyActionKernel","label":"isMarkovKernel_policyActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_policyActionKernel","description":"The policy action kernel is a Mathlib Markov kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-dd1a9f67ede7","parent":"module:BanditRLProof.RewardKernel","order":9475,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:390"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_policyActionKernel {State : Type u} [MeasurableSpace State] (policy : Policy.MeasurablePolicy State Action) : ProbabilityTheory.IsMarkovKernel (policyActionKernel (Context := Context) policy)","missing":[],"search":"ismarkovkernel_policyactionkernel banditrlproof.rewardkernel.ismarkovkernel_policyactionkernel the policy action kernel is a mathlib markov kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward","label":"composePolicyActionReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicyActionReward","description":"One-step action/reward kernel induced by a deterministic measurable policy and a context/action reward kernel. The output is the pair `(action, reward)`: the action coordinate is the policy-selected deterministic action, and the reward coordinate is drawn from the reward law selected by that same action. This is still a one-step construction; finite-prefix assembly is provided separately below.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-4ca595f393be","parent":"module:BanditRLProof.RewardKernel","order":9476,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:407"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def composePolicyActionReward {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) : MarkovRewardKernel (Context × State) (Action × Reward) where","missing":[],"search":"composepolicyactionreward banditrlproof.rewardkernel.composepolicyactionreward one-step action/reward kernel induced by a deterministic measurable policy and a context/action reward kernel. the output is the pair `(action, reward)`: the action coordinate is the policy-selected deterministic action, and the reward coordinate is drawn from the reward law selected by that same action. this is still a one-step construction; finite-prefix assembly is provided separately below. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_kernel","label":"composePolicyActionReward_kernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicyActionReward_kernel","description":"theorem composePolicyActionReward_kernel {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) : (composePolicyActionReward rewardKernel policy).kernel = ProbabilityTheory.Kernel.prod (policyActionKernel (Context := Context) policy) (composePolicy rewardKernel policy).kernel","url":"../modules/banditrlproof-rewardkernel/index.html#decl-27b2af75b637","parent":"module:BanditRLProof.RewardKernel","order":9477,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:428"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicyActionReward_kernel {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) : (composePolicyActionReward rewardKernel policy).kernel = ProbabilityTheory.Kernel.prod (policyActionKernel (Context := Context) policy) (composePolicy rewardKernel policy).kernel","missing":[],"search":"composepolicyactionreward_kernel banditrlproof.rewardkernel.composepolicyactionreward_kernel theorem composepolicyactionreward_kernel {state : type u} [measurablespace state] (rewardkernel : markovrewardkernel (context × action) reward) (policy : policy.measurablepolicy state action) : (composepolicyactionreward rewardkernel policy).kernel = probabilitytheory.kernel.prod (policyactionkernel (context := context) policy) (composepolicy rewardkernel policy).kernel theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_composePolicyActionReward","label":"isMarkovKernel_composePolicyActionReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_composePolicyActionReward","description":"The one-step action/reward policy/reward kernel is a Mathlib Markov kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e4a4b6bfd9cd","parent":"module:BanditRLProof.RewardKernel","order":9478,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:439"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_composePolicyActionReward {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) : ProbabilityTheory.IsMarkovKernel (composePolicyActionReward rewardKernel policy).kernel","missing":[],"search":"ismarkovkernel_composepolicyactionreward banditrlproof.rewardkernel.ismarkovkernel_composepolicyactionreward the one-step action/reward policy/reward kernel is a mathlib markov kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_composePolicyActionReward_eventProbability","label":"measurable_composePolicyActionReward_eventProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_composePolicyActionReward_eventProbability","description":"Event-probability measurability for the one-step action/reward kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-4bf89adba9ec","parent":"module:BanditRLProof.RewardKernel","order":9479,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:450"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_composePolicyActionReward_eventProbability {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) {event : Set (Action × Reward)} (hevent : MeasurableSet event) : Measurable (fun pair : Context × State => (composePolicyActionReward rewardKernel policy).kernel pair event)","missing":[],"search":"measurable_composepolicyactionreward_eventprobability banditrlproof.rewardkernel.measurable_composepolicyactionreward_eventprobability event-probability measurability for the one-step action/reward kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_reward_event","label":"composePolicyActionReward_reward_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicyActionReward_reward_event","description":"The reward marginal of the one-step action/reward kernel is exactly the policy-selected reward law. This is a kernel-level law-transfer wrapper: it identifies the second coordinate of the `(Action × Reward)` step kernel, but it does not identify a global conditional expectation kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-83a75d07fb61","parent":"module:BanditRLProof.RewardKernel","order":9480,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:474"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicyActionReward_reward_event {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) {event : Set Reward} (hevent : MeasurableSet event) : (composePolicyActionReward rewardKernel policy).kernel pair (Prod.snd ⁻¹' event) = selectedMeasure rewardKernel pair.1 (policy.action pair.2) event","missing":[],"search":"composepolicyactionreward_reward_event banditrlproof.rewardkernel.composepolicyactionreward_reward_event the reward marginal of the one-step action/reward kernel is exactly the policy-selected reward law. this is a kernel-level law-transfer wrapper: it identifies the second coordinate of the `(action × reward)` step kernel, but it does not identify a global conditional expectation kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_reward_map","label":"composePolicyActionReward_reward_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicyActionReward_reward_map","description":"Measure-level reward marginal of the one-step action/reward kernel. This is the pushforward version of `composePolicyActionReward_reward_event`. It exposes the exact shape needed by conditional-kernel law identification: mapping the `(Action × Reward)` one-step law through `Prod.snd` recovers the policy-selected reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-7090f68fd360","parent":"module:BanditRLProof.RewardKernel","order":9481,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:522"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicyActionReward_reward_map {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) : Measure.map Prod.snd ((composePolicyActionReward rewardKernel policy).kernel pair) = selectedMeasure rewardKernel pair.1 (policy.action pair.2)","missing":[],"search":"composepolicyactionreward_reward_map banditrlproof.rewardkernel.composepolicyactionreward_reward_map measure-level reward marginal of the one-step action/reward kernel. this is the pushforward version of `composepolicyactionreward_reward_event`. it exposes the exact shape needed by conditional-kernel law identification: mapping the `(action × reward)` one-step law through `prod.snd` recovers the policy-selected reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_kernel_apply_eq_map_prod_mk","label":"composePolicyActionReward_kernel_apply_eq_map_prod_mk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicyActionReward_kernel_apply_eq_map_prod_mk","description":"Pointwise measure shape of the one-step action/reward policy/reward kernel. The action coordinate is deterministic, so the full `(Action × Reward)` law is the selected reward law pushed through `Prod.mk` with the policy action fixed.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-3edb8c38b516","parent":"module:BanditRLProof.RewardKernel","order":9482,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:541"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicyActionReward_kernel_apply_eq_map_prod_mk {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Policy.MeasurablePolicy State Action) (pair : Context × State) : (composePolicyActionReward rewardKernel policy).kernel pair = Measure.map (Prod.mk (policy.action pair.2)) (selectedMeasure rewardKernel pair.1 (policy.action pair.2))","missing":[],"search":"composepolicyactionreward_kernel_apply_eq_map_prod_mk banditrlproof.rewardkernel.composepolicyactionreward_kernel_apply_eq_map_prod_mk pointwise measure shape of the one-step action/reward policy/reward kernel. the action coordinate is deterministic, so the full `(action × reward)` law is the selected reward law pushed through `prod.mk` with the policy action fixed. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_action_map","label":"composePolicyActionReward_action_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicyActionReward_action_map","description":"Measure-level action marginal of the one-step action/reward kernel. The action coordinate is deterministic: mapping the `(Action × Reward)` law through `Prod.fst` recovers the Dirac measure at the policy-selected action.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-bb0519f9549c","parent":"module:BanditRLProof.RewardKernel","order":9483,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:574"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicyActionReward_action_map {State : Type u} [MeasurableSpace State] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Policy.MeasurablePolicy State Action) (pair : Prod Context State) : Measure.map Prod.fst ((composePolicyActionReward rewardKernel policy).kernel pair) = Measure.dirac (policy.action pair.2)","missing":[],"search":"composepolicyactionreward_action_map banditrlproof.rewardkernel.composepolicyactionreward_action_map measure-level action marginal of the one-step action/reward kernel. the action coordinate is deterministic: mapping the `(action × reward)` law through `prod.fst` recovers the dirac measure at the policy-selected action. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.CenteredRewardKernelLaw","label":"CenteredRewardKernelLaw","kind":"structure","status":"compiled","subtitle":"BanditRLProof.RewardKernel.CenteredRewardKernelLaw","description":"Pointwise centered-reward law contract for a context/action reward kernel. For every context/action index, the selected reward law has a centered reward with zero integral and a sub-Gaussian MGF. The contract is deliberately one-step: the transfer theorems below move these facts through policy composition and history-indexed step kernels, but they do not identify a global trajectory conditional expectation.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-0067206525e7","parent":"module:BanditRLProof.RewardKernel","order":9484,"meta":[["Kind","structure"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:630"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"structure CenteredRewardKernelLaw {Context : Type x} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) where","missing":[],"search":"centeredrewardkernellaw banditrlproof.rewardkernel.centeredrewardkernellaw pointwise centered-reward law contract for a context/action reward kernel. for every context/action index, the selected reward law has a centered reward with zero integral and a sub-gaussian mgf. the contract is deliberately one-step: the transfer theorems below move these facts through policy composition and history-indexed step kernels, but they do not identify a global trajectory conditional expectation. structure compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicy_centeredReward_integrable","label":"composePolicy_centeredReward_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicy_centeredReward_integrable","description":"A policy-composed one-step reward kernel inherits centered-reward integrability from the underlying context/action reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-0fdffc9fd932","parent":"module:BanditRLProof.RewardKernel","order":9485,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:659"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicy_centeredReward_integrable {Context : Type x} {State : Type u} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (policy : Policy.MeasurablePolicy State Action) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : CenteredRewardKernelLaw rewardKernel mean varianceProxy) (pair : Context × State) : MeasureTheory.Integrable (fun reward : Rat => (((reward - mean pair.1 (policy.action pair.2) : Rat) : Real))) ((composePolicy rewardKernel policy).kernel pair)","missing":[],"search":"composepolicy_centeredreward_integrable banditrlproof.rewardkernel.composepolicy_centeredreward_integrable a policy-composed one-step reward kernel inherits centered-reward integrability from the underlying context/action reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicy_centeredReward_integral_eq_zero","label":"composePolicy_centeredReward_integral_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicy_centeredReward_integral_eq_zero","description":"A policy-composed one-step reward kernel inherits the centered-reward zero-integral fact from the underlying context/action reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-270b69dbebb8","parent":"module:BanditRLProof.RewardKernel","order":9486,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:681"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicy_centeredReward_integral_eq_zero {Context : Type x} {State : Type u} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (policy : Policy.MeasurablePolicy State Action) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : CenteredRewardKernelLaw rewardKernel mean varianceProxy) (pair : Context × State) : MeasureTheory.integral ((composePolicy rewardKernel policy).kernel pair) (fun reward : Rat => (((reward - mean pair.1 (policy.action pair.2) : Rat) : Real))) = 0","missing":[],"search":"composepolicy_centeredreward_integral_eq_zero banditrlproof.rewardkernel.composepolicy_centeredreward_integral_eq_zero a policy-composed one-step reward kernel inherits the centered-reward zero-integral fact from the underlying context/action reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.composePolicy_centeredReward_hasSubgaussianMGF","label":"composePolicy_centeredReward_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.composePolicy_centeredReward_hasSubgaussianMGF","description":"A policy-composed one-step reward kernel inherits the centered-reward sub-Gaussian MGF witness from the underlying context/action reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e01e6d672763","parent":"module:BanditRLProof.RewardKernel","order":9487,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:702"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem composePolicy_centeredReward_hasSubgaussianMGF {Context : Type x} {State : Type u} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (policy : Policy.MeasurablePolicy State Action) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : CenteredRewardKernelLaw rewardKernel mean varianceProxy) (pair : Context × State) : ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => (((reward - mean pair.1 (policy.action pair.2) : Rat) : Real))) (varianceProxy pair.1 (policy.action pair.2)) ((composePolicy rewardKernel policy).kernel pair)","missing":[],"search":"composepolicy_centeredreward_hassubgaussianmgf banditrlproof.rewardkernel.composepolicy_centeredreward_hassubgaussianmgf a policy-composed one-step reward kernel inherits the centered-reward sub-gaussian mgf witness from the underlying context/action reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.historyStepRewardKernel","label":"historyStepRewardKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.historyStepRewardKernel","description":"The one-step reward kernel selected from a finite reward history. This is an Ionescu-Tulcea-facing precursor for `KERNEL-POLICY-BIND`: for each time `n`, a measurable context extractor and a measurable policy-state extractor turn the finite reward history `Π i : Finset.Iic n, Reward` into the context/state pair consumed by the one-step policy/reward kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e78c198be1ad","parent":"module:BanditRLProof.RewardKernel","order":9488,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:729"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepRewardKernel {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : MarkovRewardKernel ((i : Finset.Iic n) -> Reward) Reward where","missing":[],"search":"historysteprewardkernel banditrlproof.rewardkernel.historysteprewardkernel the one-step reward kernel selected from a finite reward history. this is an ionescu-tulcea-facing precursor for `kernel-policy-bind`: for each time `n`, a measurable context extractor and a measurable policy-state extractor turn the finite reward history `π i : finset.iic n, reward` into the context/state pair consumed by the one-step policy/reward kernel. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily","label":"historyStepKernelFamily","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.historyStepKernelFamily","description":"The Mathlib kernel family consumed by `ProbabilityTheory.Kernel.partialTraj` for the constant reward-coordinate type family.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-b21b1ed69375","parent":"module:BanditRLProof.RewardKernel","order":9489,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:760"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"def historyStepKernelFamily {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> Reward) Reward","missing":[],"search":"historystepkernelfamily banditrlproof.rewardkernel.historystepkernelfamily the mathlib kernel family consumed by `probabilitytheory.kernel.partialtraj` for the constant reward-coordinate type family. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_apply","label":"historyStepKernelFamily_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.historyStepKernelFamily_apply","description":"theorem historyStepKernelFamily_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i…","url":"../modules/banditrlproof-rewardkernel/index.html#decl-9aba8de178dc","parent":"module:BanditRLProof.RewardKernel","order":9490,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:777"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Reward) : historyStepKernelFamily rewardKernel policy context state hcontext hstate n history = selectedMeasure rewardKernel (context n history) ((policy n).action (state n history))","missing":[],"search":"historystepkernelfamily_apply banditrlproof.rewardkernel.historystepkernelfamily_apply theorem historystepkernelfamily_apply {context : type x} {state : type u} {action : type y} {reward : type v} [measurablespace context] [measurablespace state] [measurablespace action] [measurablespace reward] (rewardkernel : markovrewardkernel (context × action) reward) (policy : nat -> policy.measurablepolicy state action) (context : (n : nat) -> ((i : finset.iic n) -> reward) -> context) (state : (n : nat) -> ((i : finset.iic n) -> reward) -> state) (hcontext : forall n : nat, measurable (context n)) (hstate : forall n : nat, measurable (state n)) (n : nat) (history : (i : finset.iic n) -> reward) : historystepkernelfamily rewardkernel policy context state hcontext hstate n history = selectedmeasure rewardkernel (context n history) ((policy n).action (state n history)) theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_historyStepKernelFamily","label":"isMarkovKernel_historyStepKernelFamily","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_historyStepKernelFamily","description":"Every kernel in `historyStepKernelFamily` is a Mathlib Markov kernel.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-8b5a32247854","parent":"module:BanditRLProof.RewardKernel","order":9491,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:795"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_historyStepKernelFamily {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) : forall n : Nat, ProbabilityTheory.IsMarkovKernel (historyStepKernelFamily rewardKernel policy context state hcontext hstate n)","missing":[],"search":"ismarkovkernel_historystepkernelfamily banditrlproof.rewardkernel.ismarkovkernel_historystepkernelfamily every kernel in `historystepkernelfamily` is a mathlib markov kernel. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_historyStepKernelFamily_eventProbability","label":"measurable_historyStepKernelFamily_eventProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_historyStepKernelFamily_eventProbability","description":"For any measurable reward event, the event probability selected by one member of the history-indexed kernel family is measurable in the finite reward history.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-83e2f8c50d44","parent":"module:BanditRLProof.RewardKernel","order":9492,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:834"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_historyStepKernelFamily_eventProbability {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) {event : Set Reward} (hevent : MeasurableSet event) : Measurable (fun history : (i : Finset.Iic n) -> Reward => historyStepKernelFamily rewardKernel policy context state hcontext hstate n history event)","missing":[],"search":"measurable_historystepkernelfamily_eventprobability banditrlproof.rewardkernel.measurable_historystepkernelfamily_eventprobability for any measurable reward event, the event probability selected by one member of the history-indexed kernel family is measurable in the finite reward history. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_integrable","label":"historyStepKernelFamily_centeredReward_integrable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_integrable","description":"The history-indexed one-step reward kernel inherits centered-reward integrability from the underlying context/action reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-50714d32eb01","parent":"module:BanditRLProof.RewardKernel","order":9493,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:860"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_integrable {Context : Type x} {State : Type u} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : CenteredRewardKernelLaw rewardKernel mean varianceProxy) (n : Nat) (history : (i : Finset.Iic n) -> Rat) : MeasureTheory.Integrable (fun reward : Rat => (((reward - mean (context n history) ((policy n).action (state n history)) : Rat) : Real))) (historyStepKernelFamily rewardKernel poli…","missing":[],"search":"historystepkernelfamily_centeredreward_integrable banditrlproof.rewardkernel.historystepkernelfamily_centeredreward_integrable the history-indexed one-step reward kernel inherits centered-reward integrability from the underlying context/action reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_integral_eq_zero","label":"historyStepKernelFamily_centeredReward_integral_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_integral_eq_zero","description":"The history-indexed one-step reward kernel inherits the centered-reward zero-integral fact from the underlying context/action reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-55c90c8478a9","parent":"module:BanditRLProof.RewardKernel","order":9494,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:889"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_integral_eq_zero {Context : Type x} {State : Type u} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : CenteredRewardKernelLaw rewardKernel mean varianceProxy) (n : Nat) (history : (i : Finset.Iic n) -> Rat) : MeasureTheory.integral (historyStepKernelFamily rewardKernel policy context state hcontext hstate n history) (fun reward : Rat => (((reward - mean (context n history) ((polic…","missing":[],"search":"historystepkernelfamily_centeredreward_integral_eq_zero banditrlproof.rewardkernel.historystepkernelfamily_centeredreward_integral_eq_zero the history-indexed one-step reward kernel inherits the centered-reward zero-integral fact from the underlying context/action reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_hasSubgaussianMGF","label":"historyStepKernelFamily_centeredReward_hasSubgaussianMGF","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_hasSubgaussianMGF","description":"The history-indexed one-step reward kernel inherits the centered-reward sub-Gaussian MGF witness from the underlying context/action reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-12c5c3beb4aa","parent":"module:BanditRLProof.RewardKernel","order":9495,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:919"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem historyStepKernelFamily_centeredReward_hasSubgaussianMGF {Context : Type x} {State : Type u} {Action : Type y} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] (rewardKernel : MarkovRewardKernel (Context × Action) Rat) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Rat) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (mean : Context -> Action -> Rat) (varianceProxy : Context -> Action -> NNReal) (law : CenteredRewardKernelLaw rewardKernel mean varianceProxy) (n : Nat) (history : (i : Finset.Iic n) -> Rat) : ProbabilityTheory.HasSubgaussianMGF (fun reward : Rat => (((reward - mean (context n history) ((policy n).action (state n history)) : Rat) : Real))) (varianceProxy (context…","missing":[],"search":"historystepkernelfamily_centeredreward_hassubgaussianmgf banditrlproof.rewardkernel.historystepkernelfamily_centeredreward_hassubgaussianmgf the history-indexed one-step reward kernel inherits the centered-reward sub-gaussian mgf witness from the underlying context/action reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.partialTrajectoryKernel","label":"partialTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.partialTrajectoryKernel","description":"Finite-prefix trajectory kernel obtained by feeding the history-indexed policy/reward step-kernel family to Mathlib's `partialTraj` construction. This gives the finite-prefix Ionescu-Tulcea assembly surface for reward histories only. It is not yet the final bandit trajectory law with action trace, conditional reward-law transfer, or adaptive regret theorem.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-de77bf9c82ad","parent":"module:BanditRLProof.RewardKernel","order":9496,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:955"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def partialTrajectoryKernel {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (a b : Nat) : ProbabilityTheory.Kernel ((i : Finset.Iic a) -> Reward) ((i : Finset.Iic b) -> Reward)","missing":[],"search":"partialtrajectorykernel banditrlproof.rewardkernel.partialtrajectorykernel finite-prefix trajectory kernel obtained by feeding the history-indexed policy/reward step-kernel family to mathlib's `partialtraj` construction. this gives the finite-prefix ionescu-tulcea assembly surface for reward histories only. it is not yet the final bandit trajectory law with action trace, conditional reward-law transfer, or adaptive regret theorem. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_partialTrajectoryKernel","label":"isMarkovKernel_partialTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_partialTrajectoryKernel","description":"The finite-prefix trajectory kernel assembled by `partialTraj` is Markov.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-d7f1c60b2038","parent":"module:BanditRLProof.RewardKernel","order":9497,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:976"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_partialTrajectoryKernel {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (a b : Nat) : ProbabilityTheory.IsMarkovKernel (partialTrajectoryKernel rewardKernel policy context state hcontext hstate a b)","missing":[],"search":"ismarkovkernel_partialtrajectorykernel banditrlproof.rewardkernel.ismarkovkernel_partialtrajectorykernel the finite-prefix trajectory kernel assembled by `partialtraj` is markov. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_partialTrajectoryKernel_eventProbability","label":"measurable_partialTrajectoryKernel_eventProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_partialTrajectoryKernel_eventProbability","description":"Event-probability measurability for the finite-prefix trajectory kernel assembled by Mathlib's `partialTraj`.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ae4ba8af07b3","parent":"module:BanditRLProof.RewardKernel","order":9498,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1004"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_partialTrajectoryKernel_eventProbability {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (a b : Nat) {event : Set ((i : Finset.Iic b) -> Reward)} (hevent : MeasurableSet event) : Measurable (fun history : (i : Finset.Iic a) -> Reward => partialTrajectoryKernel rewardKernel policy context state hcontext hstate a b history event)","missing":[],"search":"measurable_partialtrajectorykernel_eventprobability banditrlproof.rewardkernel.measurable_partialtrajectorykernel_eventprobability event-probability measurability for the finite-prefix trajectory kernel assembled by mathlib's `partialtraj`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.partialTrajectoryKernel_succ_next_map","label":"partialTrajectoryKernel_succ_next_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.partialTrajectoryKernel_succ_next_map","description":"For a one-step extension, the reward-only partial trajectory kernel has the configured history-step reward kernel as its next-coordinate marginal. This is a local wrapper around Mathlib's `Kernel.map_partialTraj_succ_self`. It is a trajectory-kernel fact, not a conditional-expectation-kernel identification.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-07256a5457cc","parent":"module:BanditRLProof.RewardKernel","order":9499,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1036"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem partialTrajectoryKernel_succ_next_map {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : (partialTrajectoryKernel rewardKernel policy context state hcontext hstate n (n + 1)).map (fun history : (i : Finset.Iic (n + 1)) -> Reward => history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) = historyStepKernelFamily rewardKernel policy context state hcontext hstate n","missing":[],"search":"partialtrajectorykernel_succ_next_map banditrlproof.rewardkernel.partialtrajectorykernel_succ_next_map for a one-step extension, the reward-only partial trajectory kernel has the configured history-step reward kernel as its next-coordinate marginal. this is a local wrapper around mathlib's `kernel.map_partialtraj_succ_self`. it is a trajectory-kernel fact, not a conditional-expectation-kernel identification. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.partialTrajectoryKernel_succ_next_map_apply","label":"partialTrajectoryKernel_succ_next_map_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.partialTrajectoryKernel_succ_next_map_apply","description":"Pointwise measure form of `partialTrajectoryKernel_succ_next_map`.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-7443fe8ff56e","parent":"module:BanditRLProof.RewardKernel","order":9500,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1068"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem partialTrajectoryKernel_succ_next_map_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Reward) : Measure.map (fun extended : (i : Finset.Iic (n + 1)) -> Reward => extended ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) (partialTrajectoryKernel rewardKernel policy context state hcontext hstate n (n + 1) history) = historyStepKernelFamily rewardKernel policy context state hcontext h…","missing":[],"search":"partialtrajectorykernel_succ_next_map_apply banditrlproof.rewardkernel.partialtrajectorykernel_succ_next_map_apply pointwise measure form of `partialtrajectorykernel_succ_next_map`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernel","label":"actionRewardHistoryStepKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernel","description":"The one-step action/reward kernel selected from a finite action/reward pair history. This is the action/reward analogue of `historyStepRewardKernel`: the state seen by the policy may depend on the finite prefix of previously emitted `(Action × Reward)` pairs, and the next emitted object is again an `(Action × Reward)` pair.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-f423f2b2daa8","parent":"module:BanditRLProof.RewardKernel","order":9501,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1107"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def actionRewardHistoryStepKernel {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : MarkovRewardKernel ((i : Finset.Iic n) -> Action × Reward) (Action × Reward) where","missing":[],"search":"actionrewardhistorystepkernel banditrlproof.rewardkernel.actionrewardhistorystepkernel the one-step action/reward kernel selected from a finite action/reward pair history. this is the action/reward analogue of `historysteprewardkernel`: the state seen by the policy may depend on the finite prefix of previously emitted `(action × reward)` pairs, and the next emitted object is again an `(action × reward)` pair. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily","label":"actionRewardHistoryStepKernelFamily","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily","description":"The Mathlib kernel family consumed by `partialTraj` for constant `(Action × Reward)` trajectory coordinates.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e54659d44747","parent":"module:BanditRLProof.RewardKernel","order":9502,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1144"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def actionRewardHistoryStepKernelFamily {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> Action × Reward) (Action × Reward)","missing":[],"search":"actionrewardhistorystepkernelfamily banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily the mathlib kernel family consumed by `partialtraj` for constant `(action × reward)` trajectory coordinates. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_apply","label":"actionRewardHistoryStepKernelFamily_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_apply","description":"theorem actionRewardHistoryStepKernelFamily_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (sta…","url":"../modules/banditrlproof-rewardkernel/index.html#decl-9da179d3847b","parent":"module:BanditRLProof.RewardKernel","order":9503,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1165"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Action × Reward) : actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n history = (composePolicyActionReward rewardKernel (policy n)).kernel (context n history, state n history)","missing":[],"search":"actionrewardhistorystepkernelfamily_apply banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_apply theorem actionrewardhistorystepkernelfamily_apply {context : type x} {state : type u} {action : type y} {reward : type v} [measurablespace context] [measurablespace state] [measurablespace action] [measurablespace reward] (rewardkernel : markovrewardkernel (context × action) reward) (policy : nat -> policy.measurablepolicy state action) (context : (n : nat) -> ((i : finset.iic n) -> action × reward) -> context) (state : (n : nat) -> ((i : finset.iic n) -> action × reward) -> state) (hcontext : forall n : nat, measurable (context n)) (hstate : forall n : nat, measurable (state n)) (n : nat) (history : (i : finset.iic n) -> action × reward) : actionrewardhistorystepkernelfamily rewardkernel policy context state hcontext hstate n history = (composepolicyactionreward rewardkernel (policy n)).kernel (context n history, state n history) theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_actionRewardHistoryStepKernelFamily","label":"isMarkovKernel_actionRewardHistoryStepKernelFamily","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_actionRewardHistoryStepKernelFamily","description":"Every kernel in `actionRewardHistoryStepKernelFamily` is Markov.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-35d69fd5da87","parent":"module:BanditRLProof.RewardKernel","order":9504,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1186"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_actionRewardHistoryStepKernelFamily {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) : forall n : Nat, ProbabilityTheory.IsMarkovKernel (actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n)","missing":[],"search":"ismarkovkernel_actionrewardhistorystepkernelfamily banditrlproof.rewardkernel.ismarkovkernel_actionrewardhistorystepkernelfamily every kernel in `actionrewardhistorystepkernelfamily` is markov. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_actionRewardHistoryStepKernelFamily_eventProbability","label":"measurable_actionRewardHistoryStepKernelFamily_eventProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_actionRewardHistoryStepKernelFamily_eventProbability","description":"For any measurable action/reward-pair event, the event probability selected by one member of the action/reward history-indexed kernel family is measurable in the finite action/reward pair history.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ccd9fd5f19a5","parent":"module:BanditRLProof.RewardKernel","order":9505,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1232"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_actionRewardHistoryStepKernelFamily_eventProbability {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) {event : Set (Action × Reward)} (hevent : MeasurableSet event) : Measurable (fun history : (i : Finset.Iic n) -> Action × Reward => actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n history event)","missing":[],"search":"measurable_actionrewardhistorystepkernelfamily_eventprobability banditrlproof.rewardkernel.measurable_actionrewardhistorystepkernelfamily_eventprobability for any measurable action/reward-pair event, the event probability selected by one member of the action/reward history-indexed kernel family is measurable in the finite action/reward pair history. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_event","label":"actionRewardHistoryStepKernelFamily_reward_event","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_event","description":"The reward marginal of a history-indexed one-step action/reward kernel is the reward law selected by the same finite history.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-27066aa70cef","parent":"module:BanditRLProof.RewardKernel","order":9506,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1263"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_reward_event {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Action × Reward) {event : Set Reward} (hevent : MeasurableSet event) : actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n history (Prod.snd ⁻¹' event) = selectedMeasure rewardKernel (context n history) ((policy n).action (sta…","missing":[],"search":"actionrewardhistorystepkernelfamily_reward_event banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_reward_event the reward marginal of a history-indexed one-step action/reward kernel is the reward law selected by the same finite history. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_map","label":"actionRewardHistoryStepKernelFamily_reward_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_map","description":"Measure-level reward marginal of a history-indexed one-step action/reward kernel. This is the pushforward version of `actionRewardHistoryStepKernelFamily_reward_event`, and is the local `RewardKernel` side of the future `condExpKernel` reward-coordinate map identification.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-f81871a0ae3c","parent":"module:BanditRLProof.RewardKernel","order":9507,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1297"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_reward_map {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Action × Reward) : Measure.map Prod.snd (actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n history) = selectedMeasure rewardKernel (context n history) ((policy n).action (state n history))","missing":[],"search":"actionrewardhistorystepkernelfamily_reward_map banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_reward_map measure-level reward marginal of a history-indexed one-step action/reward kernel. this is the pushforward version of `actionrewardhistorystepkernelfamily_reward_event`, and is the local `rewardkernel` side of the future `condexpkernel` reward-coordinate map identification. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_action_map","label":"actionRewardHistoryStepKernelFamily_action_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_action_map","description":"Measure-level action marginal of a history-indexed one-step action/reward kernel. For a fixed finite history, the next action is deterministic and equals the action chosen by the time-`n` policy on the state selected from that history.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-a404078f2512","parent":"module:BanditRLProof.RewardKernel","order":9508,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1329"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_action_map {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Prod Action Reward) : Measure.map Prod.fst (actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n history) = Measure.dirac ((policy n).action (state n history))","missing":[],"search":"actionrewardhistorystepkernelfamily_action_map banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_action_map measure-level action marginal of a history-indexed one-step action/reward kernel. for a fixed finite history, the next action is deterministic and equals the action chosen by the time-`n` policy on the state selected from that history. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_apply_eq_map_prod_mk","label":"actionRewardHistoryStepKernelFamily_apply_eq_map_prod_mk","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_apply_eq_map_prod_mk","description":"Pointwise measure shape of a history-indexed action/reward step kernel. For a fixed finite pair history, the next action is the policy-selected action and the reward coordinate is drawn from the corresponding selected reward law.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-f6cdec26b0b4","parent":"module:BanditRLProof.RewardKernel","order":9509,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1358"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_apply_eq_map_prod_mk {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Action × Reward) : actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hstate n history = Measure.map (Prod.mk ((policy n).action (state n history))) (selectedMeasure rewardKernel (context n history) ((policy n).action (state n…","missing":[],"search":"actionrewardhistorystepkernelfamily_apply_eq_map_prod_mk banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_apply_eq_map_prod_mk pointwise measure shape of a history-indexed action/reward step kernel. for a fixed finite pair history, the next action is the policy-selected action and the reward coordinate is drawn from the corresponding selected reward law. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel","label":"actionRewardPartialTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel","description":"Finite-prefix action/reward trajectory kernel obtained by feeding the history-indexed action/reward step-kernel family to Mathlib's `partialTraj`. This is the compiled finite-prefix action/reward trajectory-law surface for `KERNEL-POLICY-BIND`. It still does not prove conditional reward-law transfer, posterior kernels, or final adaptive regret theorems.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-6da81c136f94","parent":"module:BanditRLProof.RewardKernel","order":9510,"meta":[["Kind","definition"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1391"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"noncomputable def actionRewardPartialTrajectoryKernel {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (a b : Nat) : ProbabilityTheory.Kernel ((i : Finset.Iic a) -> Action × Reward) ((i : Finset.Iic b) -> Action × Reward)","missing":[],"search":"actionrewardpartialtrajectorykernel banditrlproof.rewardkernel.actionrewardpartialtrajectorykernel finite-prefix action/reward trajectory kernel obtained by feeding the history-indexed action/reward step-kernel family to mathlib's `partialtraj`. this is the compiled finite-prefix action/reward trajectory-law surface for `kernel-policy-bind`. it still does not prove conditional reward-law transfer, posterior kernels, or final adaptive regret theorems. definition compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_actionRewardPartialTrajectoryKernel","label":"isMarkovKernel_actionRewardPartialTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.isMarkovKernel_actionRewardPartialTrajectoryKernel","description":"The finite-prefix action/reward trajectory kernel is Markov.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-9803b9c970d3","parent":"module:BanditRLProof.RewardKernel","order":9511,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1415"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem isMarkovKernel_actionRewardPartialTrajectoryKernel {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (a b : Nat) : ProbabilityTheory.IsMarkovKernel (actionRewardPartialTrajectoryKernel rewardKernel policy context state hcontext hstate a b)","missing":[],"search":"ismarkovkernel_actionrewardpartialtrajectorykernel banditrlproof.rewardkernel.ismarkovkernel_actionrewardpartialtrajectorykernel the finite-prefix action/reward trajectory kernel is markov. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.measurable_actionRewardPartialTrajectoryKernel_eventProbability","label":"measurable_actionRewardPartialTrajectoryKernel_eventProbability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.measurable_actionRewardPartialTrajectoryKernel_eventProbability","description":"Event-probability measurability for the finite-prefix action/reward trajectory kernel assembled by Mathlib's `partialTraj`.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-46638c1e70c0","parent":"module:BanditRLProof.RewardKernel","order":9512,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1446"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_actionRewardPartialTrajectoryKernel_eventProbability {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (a b : Nat) {event : Set ((i : Finset.Iic b) -> Action × Reward)} (hevent : MeasurableSet event) : Measurable (fun history : (i : Finset.Iic a) -> Action × Reward => actionRewardPartialTrajectoryKernel rewardKernel policy context state hcontext hstate a b history event)","missing":[],"search":"measurable_actionrewardpartialtrajectorykernel_eventprobability banditrlproof.rewardkernel.measurable_actionrewardpartialtrajectorykernel_eventprobability event-probability measurability for the finite-prefix action/reward trajectory kernel assembled by mathlib's `partialtraj`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_next_map","label":"actionRewardPartialTrajectoryKernel_succ_next_map","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_next_map","description":"For a one-step extension, the action/reward partial trajectory kernel has the configured action/reward history-step kernel as its next-coordinate marginal. This is the action/reward-pair version of `partialTrajectoryKernel_succ_next_map` and exposes the exact Mathlib `partialTraj` marginal used by future `condExpKernel` pair-law identification.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-c6a03e33f7b8","parent":"module:BanditRLProof.RewardKernel","order":9513,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1480"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_succ_next_map {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : (actionRewardPartialTrajectoryKernel rewardKernel policy context state hcontext hstate n (n + 1)).map (fun history : (i : Finset.Iic (n + 1)) -> Action × Reward => history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) = actionRewardHistoryStepKernelFamily rewardKernel policy context state hcontext hst…","missing":[],"search":"actionrewardpartialtrajectorykernel_succ_next_map banditrlproof.rewardkernel.actionrewardpartialtrajectorykernel_succ_next_map for a one-step extension, the action/reward partial trajectory kernel has the configured action/reward history-step kernel as its next-coordinate marginal. this is the action/reward-pair version of `partialtrajectorykernel_succ_next_map` and exposes the exact mathlib `partialtraj` marginal used by future `condexpkernel` pair-law identification. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_next_map_apply","label":"actionRewardPartialTrajectoryKernel_succ_next_map_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_next_map_apply","description":"Pointwise measure form of `actionRewardPartialTrajectoryKernel_succ_next_map`.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ecda05e4e723","parent":"module:BanditRLProof.RewardKernel","order":9514,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1518"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_succ_next_map_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Action × Reward) : Measure.map (fun extended : (i : Finset.Iic (n + 1)) -> Action × Reward => extended ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩) (actionRewardPartialTrajectoryKernel rewardKernel policy context state hcontext hstate n (n + 1) history) = actionRe…","missing":[],"search":"actionrewardpartialtrajectorykernel_succ_next_map_apply banditrlproof.rewardkernel.actionrewardpartialtrajectorykernel_succ_next_map_apply pointwise measure form of `actionrewardpartialtrajectorykernel_succ_next_map`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_extend_map_apply","label":"actionRewardPartialTrajectoryKernel_succ_extend_map_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_extend_map_apply","description":"Pointwise one-step extension form of the action/reward `partialTraj` kernel. For one transition, the finite-prefix trajectory kernel is the history-step action/reward kernel pushed through the deterministic operation that appends the sampled next pair to the old finite pair prefix.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-ce5b84e96d48","parent":"module:BanditRLProof.RewardKernel","order":9515,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1557"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardPartialTrajectoryKernel_succ_extend_map_apply {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] (rewardKernel : MarkovRewardKernel (Context × Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Action × Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) (history : (i : Finset.Iic n) -> Action × Reward) : actionRewardPartialTrajectoryKernel rewardKernel policy context state hcontext hstate n (n + 1) history = Measure.map (fun next : Action × Reward => History.extendPairHistorySucc history next) (actionRewardHistoryStepKernelFamily rewa…","missing":[],"search":"actionrewardpartialtrajectorykernel_succ_extend_map_apply banditrlproof.rewardkernel.actionrewardpartialtrajectorykernel_succ_extend_map_apply pointwise one-step extension form of the action/reward `partialtraj` kernel. for one transition, the finite-prefix trajectory kernel is the history-step action/reward kernel pushed through the deterministic operation that appends the sampled next pair to the old finite pair prefix. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure","label":"actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure","description":"Canonical regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. This is the Mathlib-backed trajectory-measure side of the open `COND-EXPECT-REWARD` law-identification route. It does not identify an arbitrary ambient `condExpKernel`; it records that, on Mathlib's canonical `trajMeasure`, conditioning the next action/reward pair on the finite prefix recove…","url":"../modules/banditrlproof-rewardkernel/index.html#decl-67c7a8f99a7a","parent":"module:BanditRLProof.RewardKernel","order":9516,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1640"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [Nonempty (Prod Action Reward)] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : ProbabilityTheory.condDistrib (fun trajectory : (t : Nat) -> Prod Action Reward => trajectory (n + 1)) (Preorder.frestric…","missing":[],"search":"actionrewardhistorystepkernelfamily_conddistrib_trajmeasure banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_conddistrib_trajmeasure canonical regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. this is the mathlib-backed trajectory-measure side of the open `cond-expect-reward` law-identification route. it does not identify an arbitrary ambient `condexpkernel`; it records that, on mathlib's canonical `trajmeasure`, conditioning the next action/reward pair on the finite prefix recovers the configured `actionrewardhistorystepkernelfamily`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_condDistrib_trajMeasure","label":"actionRewardHistoryStepKernelFamily_reward_condDistrib_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_condDistrib_trajMeasure","description":"Canonical reward-marginal regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. This is the `Prod.snd` projection of `actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure`: on Mathlib's canonical `trajMeasure`, conditioning the next reward coordinate on the finite action/reward prefix recovers the reward marginal of the configured `actionRewardHis…","url":"../modules/banditrlproof-rewardkernel/index.html#decl-e525a7aa1cc7","parent":"module:BanditRLProof.RewardKernel","order":9517,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1691"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_reward_condDistrib_trajMeasure {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Reward] [Nonempty (Prod Action Reward)] [Nonempty Reward] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : Filter.EventuallyEq (MeasureTheory.ae ((ProbabilityTheory.Kernel.tra…","missing":[],"search":"actionrewardhistorystepkernelfamily_reward_conddistrib_trajmeasure banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_reward_conddistrib_trajmeasure canonical reward-marginal regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. this is the `prod.snd` projection of `actionrewardhistorystepkernelfamily_conddistrib_trajmeasure`: on mathlib's canonical `trajmeasure`, conditioning the next reward coordinate on the finite action/reward prefix recovers the reward marginal of the configured `actionrewardhistorystepkernelfamily`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_action_condDistrib_trajMeasure","label":"actionRewardHistoryStepKernelFamily_action_condDistrib_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_action_condDistrib_trajMeasure","description":"Canonical action-marginal regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. This is the `Prod.fst` projection of `actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure`: on Mathlib's canonical `trajMeasure`, conditioning the next action coordinate on the finite action/reward prefix recovers the action marginal of the configured `actionRewardHis…","url":"../modules/banditrlproof-rewardkernel/index.html#decl-fa9e0fd359ba","parent":"module:BanditRLProof.RewardKernel","order":9518,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1789"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_action_condDistrib_trajMeasure {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Action] [Nonempty (Prod Action Reward)] [Nonempty Action] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : Filter.EventuallyEq (MeasureTheory.ae ((ProbabilityTheory.Kernel.tra…","missing":[],"search":"actionrewardhistorystepkernelfamily_action_conddistrib_trajmeasure banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_action_conddistrib_trajmeasure canonical action-marginal regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. this is the `prod.fst` projection of `actionrewardhistorystepkernelfamily_conddistrib_trajmeasure`: on mathlib's canonical `trajmeasure`, conditioning the next action coordinate on the finite action/reward prefix recovers the action marginal of the configured `actionrewardhistorystepkernelfamily`. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_selectedAction_condDistrib_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedAction_condDistrib_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_selectedAction_condDistrib_trajMeasure","description":"Canonical selected-action regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. This rewrites the action marginal from `actionRewardHistoryStepKernelFamily_action_condDistrib_trajMeasure` into the Dirac law at the action selected by the frozen finite-history policy state.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-f4779f766263","parent":"module:BanditRLProof.RewardKernel","order":9519,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1885"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedAction_condDistrib_trajMeasure {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Action] [Nonempty (Prod Action Reward)] [Nonempty Action] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : Filter.EventuallyEq (MeasureTheory.ae ((ProbabilityTheory.Ke…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedaction_conddistrib_trajmeasure banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_selectedaction_conddistrib_trajmeasure canonical selected-action regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. this rewrites the action marginal from `actionrewardhistorystepkernelfamily_action_conddistrib_trajmeasure` into the dirac law at the action selected by the frozen finite-history policy state. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_selectedMeasure_condDistrib_trajMeasure","label":"actionRewardHistoryStepKernelFamily_selectedMeasure_condDistrib_trajMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_selectedMeasure_condDistrib_trajMeasure","description":"Canonical selected-reward regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. This rewrites the reward marginal from `actionRewardHistoryStepKernelFamily_reward_condDistrib_trajMeasure` into the selected context/action reward measure at the frozen finite pair history.","url":"../modules/banditrlproof-rewardkernel/index.html#decl-885988c0fb34","parent":"module:BanditRLProof.RewardKernel","order":9520,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardKernel"],["Source","BanditRLProof/RewardKernel.lean:1941"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem actionRewardHistoryStepKernelFamily_selectedMeasure_condDistrib_trajMeasure {Context : Type x} {State : Type u} {Action : Type y} {Reward : Type v} [MeasurableSpace Context] [MeasurableSpace State] [MeasurableSpace Action] [MeasurableSpace Reward] [StandardBorelSpace (Prod Action Reward)] [StandardBorelSpace Reward] [Nonempty (Prod Action Reward)] [Nonempty Reward] (mu0 : Measure (Prod Action Reward)) [MeasureTheory.IsProbabilityMeasure mu0] (rewardKernel : MarkovRewardKernel (Prod Context Action) Reward) (policy : Nat -> Policy.MeasurablePolicy State Action) (context : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> Context) (state : (n : Nat) -> ((i : Finset.Iic n) -> Prod Action Reward) -> State) (hcontext : forall n : Nat, Measurable (context n)) (hstate : forall n : Nat, Measurable (state n)) (n : Nat) : Filter.EventuallyEq (MeasureTheory.ae ((ProbabilityTheory.K…","missing":[],"search":"actionrewardhistorystepkernelfamily_selectedmeasure_conddistrib_trajmeasure banditrlproof.rewardkernel.actionrewardhistorystepkernelfamily_selectedmeasure_conddistrib_trajmeasure canonical selected-reward regular conditional distribution for the action/reward trajectory generated by the local history-step kernel family. this rewrites the reward marginal from `actionrewardhistorystepkernelfamily_reward_conddistrib_trajmeasure` into the selected context/action reward measure at the frozen finite pair history. theorem compiled","shard":"modules/6baeb7b115f3b6c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.trajMeasure_map_eval_zero","label":"trajMeasure_map_eval_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.trajMeasure_map_eval_zero","description":"The zeroth coordinate of an Ionescu-Tulcea trajectory has the supplied initial law. This is a project-level wrapper over `trajMeasure`, `traj_map_frestrictLe`, and `partialTraj_self`. It is general enough to be tracked as a Mathlib candidate.","url":"../modules/banditrlproof-rewardtracelaw/index.html#decl-9d4f9b19b0b8","parent":"module:BanditRLProof.RewardTraceLaw","order":9521,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardTraceLaw"],["Source","BanditRLProof/RewardTraceLaw.lean:26"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem trajMeasure_map_eval_zero {X : Nat -> Type*} [forall n, MeasurableSpace (X n)] (mu0 : Measure (X 0)) [IsProbabilityMeasure mu0] (kernel : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> X i) (X (n + 1))) [forall n, ProbabilityTheory.IsMarkovKernel (kernel n)] : Measure.map (fun trajectory : ((n : Nat) -> X n) => trajectory 0) (ProbabilityTheory.Kernel.trajMeasure mu0 kernel) = mu0","missing":[],"search":"trajmeasure_map_eval_zero banditrlproof.rewardkernel.trajmeasure_map_eval_zero the zeroth coordinate of an ionescu-tulcea trajectory has the supplied initial law. this is a project-level wrapper over `trajmeasure`, `traj_map_frestrictle`, and `partialtraj_self`. it is general enough to be tracked as a mathlib candidate. theorem compiled","shard":"modules/e81e7eb551d58a10.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.rewardTrace_prefix_map_eq_trajMeasure_of_condDistrib","label":"rewardTrace_prefix_map_eq_trajMeasure_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.rewardTrace_prefix_map_eq_trajMeasure_of_condDistrib","description":"Finite reward-prefix law uniqueness from the initial marginal and successor conditional distributions. Only the conditional laws before `n` are required. The proof turns each conditional-distribution identity into a joint prefix/next-reward law with `condDistrib_ae_eq_iff_measure_eq_compProd`, then matches the corresponding Ionescu-Tulcea recurrence.","url":"../modules/banditrlproof-rewardtracelaw/index.html#decl-d7f0394aecc3","parent":"module:BanditRLProof.RewardTraceLaw","order":9522,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardTraceLaw"],["Source","BanditRLProof/RewardTraceLaw.lean:70"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem rewardTrace_prefix_map_eq_trajMeasure_of_condDistrib {Omega Reward : Type*} [MeasurableSpace Omega] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Reward) [IsProbabilityMeasure mu0] (reward : Omega -> RewardTrace Reward) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (kernel : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> Reward) Reward) [forall n, ProbabilityTheory.IsMarkovKernel (kernel n)] (hzero : Measure.map (fun omega : Omega => reward omega 0) mu = mu0) (n : Nat) (hcond : forall i : Nat, i < n -> ProbabilityTheory.condDistrib (fun omega : Omega => reward omega (i + 1)) (fun omega : Omega => History.finiteRewardHistoryOfTrace (reward omega) i) mu =ᵐ[ mu.map (fun omega : Omega => History.finiteRewardHistoryOfTrace (reward omega) i)] kernel i) : Measu…","missing":[],"search":"rewardtrace_prefix_map_eq_trajmeasure_of_conddistrib banditrlproof.rewardkernel.rewardtrace_prefix_map_eq_trajmeasure_of_conddistrib finite reward-prefix law uniqueness from the initial marginal and successor conditional distributions. only the conditional laws before `n` are required. the proof turns each conditional-distribution identity into a joint prefix/next-reward law with `conddistrib_ae_eq_iff_measure_eq_compprod`, then matches the corresponding ionescu-tulcea recurrence. theorem compiled","shard":"modules/e81e7eb551d58a10.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.rewardTrace_map_eq_trajMeasure_of_condDistrib","label":"rewardTrace_map_eq_trajMeasure_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.rewardTrace_map_eq_trajMeasure_of_condDistrib","description":"The complete trace law is uniquely determined by its initial marginal and all successor conditional distributions. The finite-prefix theorem above handles every `Finset.Iic n`. Any finite set of time coordinates embeds measurably into one such prefix, so the external law and the Ionescu-Tulcea law have the same finite-dimensional marginals. Mathlib's projective-limit uniqueness then identifies the full measures.","url":"../modules/banditrlproof-rewardtracelaw/index.html#decl-eae851423463","parent":"module:BanditRLProof.RewardTraceLaw","order":9523,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardTraceLaw"],["Source","BanditRLProof/RewardTraceLaw.lean:267"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem rewardTrace_map_eq_trajMeasure_of_condDistrib {Omega Reward : Type*} [MeasurableSpace Omega] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (mu0 : Measure Reward) [IsProbabilityMeasure mu0] (reward : Omega -> RewardTrace Reward) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (kernel : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> Reward) Reward) [forall n, ProbabilityTheory.IsMarkovKernel (kernel n)] (hzero : Measure.map (fun omega : Omega => reward omega 0) mu = mu0) (hcond : forall i : Nat, ProbabilityTheory.condDistrib (fun omega : Omega => reward omega (i + 1)) (fun omega : Omega => History.finiteRewardHistoryOfTrace (reward omega) i) mu =ᵐ[mu.map (fun omega : Omega => History.finiteRewardHistoryOfTrace (reward omega) i)] kernel i) : Measure.map reward mu = Probabil…","missing":[],"search":"rewardtrace_map_eq_trajmeasure_of_conddistrib banditrlproof.rewardkernel.rewardtrace_map_eq_trajmeasure_of_conddistrib the complete trace law is uniquely determined by its initial marginal and all successor conditional distributions. the finite-prefix theorem above handles every `finset.iic n`. any finite set of time coordinates embeds measurably into one such prefix, so the external law and the ionescu-tulcea law have the same finite-dimensional marginals. mathlib's projective-limit uniqueness then identifies the full measures. theorem compiled","shard":"modules/e81e7eb551d58a10.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.RewardKernel.identDistrib_rewardTrace_of_common_condDistrib","label":"identDistrib_rewardTrace_of_common_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.RewardKernel.identDistrib_rewardTrace_of_common_condDistrib","description":"Two complete traces are identically distributed when they share an initial marginal and the same successor conditional-distribution kernels.","url":"../modules/banditrlproof-rewardtracelaw/index.html#decl-c1a0c224e7db","parent":"module:BanditRLProof.RewardTraceLaw","order":9524,"meta":[["Kind","theorem"],["Module","BanditRLProof.RewardTraceLaw"],["Source","BanditRLProof/RewardTraceLaw.lean:347"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem identDistrib_rewardTrace_of_common_condDistrib {Omega Xi Reward : Type*} [MeasurableSpace Omega] [MeasurableSpace Xi] [MeasurableSpace Reward] [StandardBorelSpace Reward] [Nonempty Reward] (mu : Measure Omega) [IsFiniteMeasure mu] (mu' : Measure Xi) [IsFiniteMeasure mu'] (mu0 : Measure Reward) [IsProbabilityMeasure mu0] (reward : Omega -> RewardTrace Reward) (reward' : Xi -> RewardTrace Reward) (hreward : forall t : Nat, Measurable (fun omega : Omega => reward omega t)) (hreward' : forall t : Nat, Measurable (fun xi : Xi => reward' xi t)) (kernel : (n : Nat) -> ProbabilityTheory.Kernel ((i : Finset.Iic n) -> Reward) Reward) [forall n, ProbabilityTheory.IsMarkovKernel (kernel n)] (hzero : Measure.map (fun omega : Omega => reward omega 0) mu = mu0) (hzero' : Measure.map (fun xi : Xi => reward' xi 0) mu' = mu0) (hcond : forall i : Nat, ProbabilityTheory.condDistrib (fun omega : Ome…","missing":[],"search":"identdistrib_rewardtrace_of_common_conddistrib banditrlproof.rewardkernel.identdistrib_rewardtrace_of_common_conddistrib two complete traces are identically distributed when they share an initial marginal and the same successor conditional-distribution kernels. theorem compiled","shard":"modules/e81e7eb551d58a10.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.ENNReal.ofReal_finset_sum_mul_natCast_of_nonneg","label":"ofReal_finset_sum_mul_natCast_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ENNReal.ofReal_finset_sum_mul_natCast_of_nonneg","description":"For a finite sum of nonnegative real weights times natural counts, `ofReal` commutes with the weighted sum and turns the counts into `ENNReal` casts. This is the `OFREAL-FINSET-WEIGHTED-NAT-FAITHFULNESS` scalar leaf. It is a faithfulness lemma under explicit pointwise nonnegativity of the real weights; it is not an expectation theorem.","url":"../modules/banditrlproof-scalarennreal/index.html#decl-8a1706a42670","parent":"module:BanditRLProof.ScalarENNReal","order":9525,"meta":[["Kind","theorem"],["Module","BanditRLProof.ScalarENNReal"],["Source","BanditRLProof/ScalarENNReal.lean:27"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ofReal_finset_sum_mul_natCast_of_nonneg {ι : Type u} (s : Finset ι) (gap : ι -> Real) (count : ι -> Nat) (hgap : forall i : ι, i ∈ s -> 0 <= gap i) : ENNReal.ofReal (s.sum (fun i : ι => gap i * ((count i : Nat) : Real))) = s.sum (fun i : ι => ENNReal.ofReal (gap i) * ((count i : Nat) : ENNReal))","missing":[],"search":"ofreal_finset_sum_mul_natcast_of_nonneg banditrlproof.ennreal.ofreal_finset_sum_mul_natcast_of_nonneg for a finite sum of nonnegative real weights times natural counts, `ofreal` commutes with the weighted sum and turns the counts into `ennreal` casts. this is the `ofreal-finset-weighted-nat-faithfulness` scalar leaf. it is a faithfulness lemma under explicit pointwise nonnegativity of the real weights; it is not an expectation theorem. theorem compiled","shard":"modules/ffe67e47b1342d9b.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.real_pseudoRegret_eq_univ_sum_model_gap_mul_natCast_pullCount","label":"real_pseudoRegret_eq_univ_sum_model_gap_mul_natCast_pullCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.real_pseudoRegret_eq_univ_sum_model_gap_mul_natCast_pullCount","description":"private theorem real_pseudoRegret_eq_univ_sum_model_gap_mul_natCast_pullCount {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n : Nat) : (((pseudoRegret model action n : Rat) : Real)) = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => (((model.gap a : Rat) : Real) * (((pullCount action a n : Nat) : Real))))","url":"../modules/banditrlproof-scalarpseudoregret/index.html#decl-e8e508ebf215","parent":"module:BanditRLProof.ScalarPseudoRegret","order":9526,"meta":[["Kind","theorem"],["Module","BanditRLProof.ScalarPseudoRegret"],["Source","BanditRLProof/ScalarPseudoRegret.lean:17"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"private theorem real_pseudoRegret_eq_univ_sum_model_gap_mul_natCast_pullCount {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (n : Nat) : (((pseudoRegret model action n : Rat) : Real)) = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => (((model.gap a : Rat) : Real) * (((pullCount action a n : Nat) : Real))))","missing":[],"search":"real_pseudoregret_eq_univ_sum_model_gap_mul_natcast_pullcount banditrlproof.real_pseudoregret_eq_univ_sum_model_gap_mul_natcast_pullcount private theorem real_pseudoregret_eq_univ_sum_model_gap_mul_natcast_pullcount {k : nat} (model : finitebanditmodel k) (action : actiontrace (fin k)) (n : nat) : (((pseudoregret model action n : rat) : real)) = (finset.univ : finset (fin k)).sum (fun a : fin k => (((model.gap a : rat) : real) * (((pullcount action a n : nat) : real)))) theorem compiled","shard":"modules/04f2599e28c848fe.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.ENNReal.ofReal_pseudoRegret_eq_univ_sum_model_gap_ofReal_mul_natCast_pullCount_of_nonneg","label":"ofReal_pseudoRegret_eq_univ_sum_model_gap_ofReal_mul_natCast_pullCount_of_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.ENNReal.ofReal_pseudoRegret_eq_univ_sum_model_gap_ofReal_mul_natCast_pullCount_of_nonneg","description":"Pointwise pseudo-regret is faithfully represented by the `ENNReal.ofReal` weighted pull-count expression when all model gaps are explicitly nonnegative after casting from `Rat` to `Real`. This is the `OFREAL-PSEUDOREGRET-PULLCOUNT-FAITHFULNESS` scalar/model bridge. It is not an expectation theorem and does not introduce measures, filtrations, kernels, or concentration assumptions.","url":"../modules/banditrlproof-scalarpseudoregret/index.html#decl-92652c00e308","parent":"module:BanditRLProof.ScalarPseudoRegret","order":9527,"meta":[["Kind","theorem"],["Module","BanditRLProof.ScalarPseudoRegret"],["Source","BanditRLProof/ScalarPseudoRegret.lean:41"],["Chapter","Foundations"],["Used in books","bandit"],["Reading references","teaching:foundations"],["Indexed settings","None registered"]],"statement":"theorem ofReal_pseudoRegret_eq_univ_sum_model_gap_ofReal_mul_natCast_pullCount_of_nonneg {K : Nat} (model : FiniteBanditModel K) (action : ActionTrace (Fin K)) (hgap : forall a : Fin K, 0 <= (((model.gap a : Rat) : Real))) (n : Nat) : ENNReal.ofReal (((pseudoRegret model action n : Rat) : Real)) = (Finset.univ : Finset (Fin K)).sum (fun a : Fin K => ENNReal.ofReal (((model.gap a : Rat) : Real)) * ((pullCount action a n : Nat) : ENNReal))","missing":[],"search":"ofreal_pseudoregret_eq_univ_sum_model_gap_ofreal_mul_natcast_pullcount_of_nonneg banditrlproof.ennreal.ofreal_pseudoregret_eq_univ_sum_model_gap_ofreal_mul_natcast_pullcount_of_nonneg pointwise pseudo-regret is faithfully represented by the `ennreal.ofreal` weighted pull-count expression when all model gaps are explicitly nonnegative after casting from `rat` to `real`. this is the `ofreal-pseudoregret-pullcount-faithfulness` scalar/model bridge. it is not an expectation theorem and does not introduce measures, filtrations, kernels, or concentration assumptions. theorem compiled","shard":"modules/04f2599e28c848fe.json","books":["bandit"],"chapters":["teaching:foundations"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialProcess","label":"halfTsallisPotentialProcess","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialProcess","description":"A time-indexed paper-normalized half-Tsallis potential.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-3303f86637d5","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9528,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisPotentialProcess {Action : Type u} (arms : Finset Action) (eta : Real) (score probability : Nat -> Action -> Real) (t : Nat) : Real","missing":[],"search":"halftsallispotentialprocess banditrlproof.tsallis.halftsallispotentialprocess a time-indexed paper-normalized half-tsallis potential. definition compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisPotentialStability_eq_linearLoss_sum_add_terminal_sub_initial","label":"sum_halfTsallisPotentialStability_eq_linearLoss_sum_add_terminal_sub_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisPotentialStability_eq_linearLoss_sum_add_terminal_sub_initial","description":"Exact fixed-learning-rate finite-horizon telescope for the candidate potential expression. When every probability is the corresponding certified minimizer, this is the constrained-potential telescope.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-7c19c58c2c19","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9529,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisPotentialStability_eq_linearLoss_sum_add_terminal_sub_initial {Action : Type u} (arms : Finset Action) (eta : Real) (score probability estimate : Nat -> Action -> Real) (T : Nat) (hscoreSucc : forall t, t < T -> score (t + 1) = fun action => score t action + estimate t action) : (Finset.range T).sum (fun t => halfTsallisPotentialStability arms eta (score t) (probability t) (estimate t) (probability (t + 1))) = (Finset.range T).sum (fun t => FTRL.linearLoss arms (probability t) (estimate t)) + halfTsallisPotentialProcess arms eta score probability T - halfTsallisPotentialProcess arms eta score probability 0","missing":[],"search":"sum_halftsallispotentialstability_eq_linearloss_sum_add_terminal_sub_initial banditrlproof.tsallis.sum_halftsallispotentialstability_eq_linearloss_sum_add_terminal_sub_initial exact fixed-learning-rate finite-horizon telescope for the candidate potential expression. when every probability is the corresponding certified minimizer, this is the constrained-potential telescope. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.importanceWeightedPotentialStabilityScore","label":"importanceWeightedPotentialStabilityScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.importanceWeightedPotentialStabilityScore","description":"The realized one-round conjugate-potential score on a history/action pair.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-272344b44f7d","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9530,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:94"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def importanceWeightedPotentialStabilityScore {History : Type u} {Action : Type v} (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (sample : History × Action) : Real","missing":[],"search":"importanceweightedpotentialstabilityscore banditrlproof.tsallis.importanceweightedpotentialstabilityscore the realized one-round conjugate-potential score on a history/action pair. definition compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedPotentialStabilityBound","label":"refinedPotentialStabilityBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedPotentialStabilityBound","description":"The paper-shaped one-round conditional expectation budget.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-d80634a893d0","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9531,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:107"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def refinedPotentialStabilityBound {History : Type u} {Action : Type v} (arms : Finset Action) (eta : Real) (prob : History -> Action -> Real) (history : History) : Real","missing":[],"search":"refinedpotentialstabilitybound banditrlproof.tsallis.refinedpotentialstabilitybound the paper-shaped one-round conditional expectation budget. definition compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_linearLoss_of_coordinatewise","label":"measurable_linearLoss_of_coordinatewise","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_linearLoss_of_coordinatewise","description":"Coordinatewise measurability closes a finite linear-loss sum.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-f4ce10fd7c4a","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9532,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:116"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_linearLoss_of_coordinatewise {History : Type u} {Action : Type v} [MeasurableSpace History] (arms : Finset Action) (prob loss : History -> Action -> Real) (hprob : forall action, action ∈ arms -> Measurable (fun history => prob history action)) (hloss : forall action, action ∈ arms -> Measurable (fun history => loss history action)) : Measurable (fun history => FTRL.linearLoss arms (prob history) (loss history))","missing":[],"search":"measurable_linearloss_of_coordinatewise banditrlproof.tsallis.measurable_linearloss_of_coordinatewise coordinatewise measurability closes a finite linear-loss sum. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_halfTsallisPotentialValue","label":"measurable_halfTsallisPotentialValue","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_halfTsallisPotentialValue","description":"The paper-normalized potential is measurable from supported coordinate measurability of its score and simplex candidate.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-7420ff29e24f","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9533,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:131"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_halfTsallisPotentialValue {History : Type u} {Action : Type v} [MeasurableSpace History] (arms : Finset Action) (eta : Real) (score probability : History -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun history => score history action)) (hprobability : forall action, action ∈ arms -> Measurable (fun history => probability history action)) : Measurable (fun history => halfTsallisPotentialValue arms eta (score history) (probability history))","missing":[],"search":"measurable_halftsallispotentialvalue banditrlproof.tsallis.measurable_halftsallispotentialvalue the paper-normalized potential is measurable from supported coordinate measurability of its score and simplex candidate. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_importanceWeightedPotentialStabilityScore","label":"measurable_importanceWeightedPotentialStabilityScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_importanceWeightedPotentialStabilityScore","description":"The conjugate-potential score is measurable from supported current score, probability, loss, and updated-selector coordinates.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-ee7d324102ef","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9534,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:171"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_importanceWeightedPotentialStabilityScore {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (hscore : forall candidate, candidate ∈ arms -> Measurable (fun history => score history candidate)) (hprob : forall candidate, candidate ∈ arms -> Measurable (fun history => prob history candidate)) (hloss : forall candidate, candidate ∈ arms -> Measurable (fun history => loss history candidate)) (hnext : forall candidate, candidate ∈ arms -> Measurable (fun sample : History × Action => next sample.1 sample.2 candidate)) : Measurable (importanceWeightedPotentialStabilityScore arms eta score prob loss next)","missing":[],"search":"measurable_importanceweightedpotentialstabilityscore banditrlproof.tsallis.measurable_importanceweightedpotentialstabilityscore the conjugate-potential score is measurable from supported current score, probability, loss, and updated-selector coordinates. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_refinedPotentialStabilityBound","label":"measurable_refinedPotentialStabilityBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_refinedPotentialStabilityBound","description":"The refined one-round budget is measurable from supported probability coordinates.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-3cb8d0a2a762","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9535,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:245"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_refinedPotentialStabilityBound {History : Type u} {Action : Type v} [MeasurableSpace History] (arms : Finset Action) (eta : Real) (prob : History -> Action -> Real) (hprob : forall action, action ∈ arms -> Measurable (fun history => prob history action)) : Measurable (refinedPotentialStabilityBound arms eta prob)","missing":[],"search":"measurable_refinedpotentialstabilitybound banditrlproof.tsallis.measurable_refinedpotentialstabilitybound the refined one-round budget is measurable from supported probability coordinates. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_nonneg_of_minimizers","label":"halfTsallisPotentialStability_nonneg_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialStability_nonneg_of_minimizers","description":"A true minimizer-to-minimizer potential step is nonnegative.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-98ee4ffbaaad","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9536,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:259"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialStability_nonneg_of_minimizers {Action : Type u} (arms : Finset Action) (eta : Real) (score probability estimate next : Action -> Real) (heta : 0 < eta) (hprobabilityMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score probability) (hnextMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + estimate action) next) : 0 <= halfTsallisPotentialStability arms eta score probability estimate next","missing":[],"search":"halftsallispotentialstability_nonneg_of_minimizers banditrlproof.tsallis.halftsallispotentialstability_nonneg_of_minimizers a true minimizer-to-minimizer potential step is nonnegative. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_refinedPotentialStabilityBound_of_finiteSimplex","label":"integrable_refinedPotentialStabilityBound_of_finiteSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_refinedPotentialStabilityBound_of_finiteSimplex","description":"The refined budget is uniformly integrable under a finite history measure.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-5c1c35e02cc4","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9537,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:308"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_refinedPotentialStabilityBound_of_finiteSimplex {History : Type u} {Action : Type v} [MeasurableSpace History] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (eta : Real) (prob : History -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (heta : 0 < eta) : Integrable (refinedPotentialStabilityBound arms eta prob) historyMu","missing":[],"search":"integrable_refinedpotentialstabilitybound_of_finitesimplex banditrlproof.tsallis.integrable_refinedpotentialstabilitybound_of_finitesimplex the refined budget is uniformly integrable under a finite history measure. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel","label":"integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel","description":"Measurability plus exact minimizer certificates give integrability of the potential score under the finite sampling kernel, without a probability floor.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-9c9c8aabcec5","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9538,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:361"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (hprobMin : forall history, FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (score history) (prob history)) (hnextMin : forall history chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun candidate => score history candidate + Exp3.importanceWei…","missing":[],"search":"integrable_importanceweightedpotentialstabilityscore_finiteactionkernel banditrlproof.tsallis.integrable_importanceweightedpotentialstabilityscore_finiteactionkernel measurability plus exact minimizer certificates give integrability of the potential score under the finite sampling kernel, without a probability floor. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_score_comp_history_action_of_condDistrib_generic","label":"integrable_score_comp_history_action_of_condDistrib_generic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_score_comp_history_action_of_condDistrib_generic","description":"Product-law integrability pulls back to a realized history/action pair when the kernel is an identified conditional action law.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-866241bafda2","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9539,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:494"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_score_comp_history_action_of_condDistrib_generic {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (score : History × Action -> Real) (policy : Kernel History Action) [IsMarkovKernel policy] (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (hIntegrable : Integrable score (mu.map history ⊗ₘ policy)) : Integrable (fun omega => score (history omega, action omega)) mu","missing":[],"search":"integrable_score_comp_history_action_of_conddistrib_generic banditrlproof.tsallis.integrable_score_comp_history_action_of_conddistrib_generic product-law integrability pulls back to a realized history/action pair when the kernel is an identified conditional action law. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_importanceWeightedPotentialStabilityScore_le_integral_refinedBound_of_condDistrib_of_minimizers","label":"integral_importanceWeightedPotentialStabilityScore_le_integral_refinedBound_of_condDistrib_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_importanceWeightedPotentialStabilityScore_le_integral_refinedBound_of_condDistrib_of_minimizers","description":"An identified finite conditional action law transports the deterministic ordinary-IW conjugate-potential theorem to a one-round integral inequality.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-0800a4a0dbbd","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9540,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:521"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_importanceWeightedPotentialStabilityScore_le_integral_refinedBound_of_condDistrib_of_minimizers {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => Exp3.finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (hprobMin : forall…","missing":[],"search":"integral_importanceweightedpotentialstabilityscore_le_integral_refinedbound_of_conddistrib_of_minimizers banditrlproof.tsallis.integral_importanceweightedpotentialstabilityscore_le_integral_refinedbound_of_conddistrib_of_minimizers an identified finite conditional action law transports the deterministic ordinary-iw conjugate-potential theorem to a one-round integral inequality. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_importanceWeightedPotentialStabilityScore_le_integral_sum_refinedBound_of_condDistrib_of_minimizers","label":"integral_sum_importanceWeightedPotentialStabilityScore_le_integral_sum_refinedBound_of_condDistrib_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_importanceWeightedPotentialStabilityScore_le_integral_sum_refinedBound_of_condDistrib_of_minimizers","description":"Finite-horizon expected conjugate-potential stability under identified conditional action laws and exact current/update minimizer certificates.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-3f93a378631a","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9541,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:619"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_importanceWeightedPotentialStabilityScore_le_integral_sum_refinedBound_of_condDistrib_of_minimizers {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (horizon : Nat) (arms : Finset Action) (eta : Real) (history : Nat -> Omega -> History) (action : Nat -> Omega -> Action) (score prob loss : Nat -> History -> Action -> Real) (next : Nat -> History -> Action -> Action -> Real) (policy : Nat -> Kernel History Action) (hmarkov : forall t, IsMarkovKernel (policy t)) (hhistory : forall t, Measurable (history t)) (haction : forall t, Measurable (action t)) (hpolicy : forall t, policy t =ᵐ[mu.map (history t)] fun h => Exp3.finiteActionMeasure arms (prob t h…","missing":[],"search":"integral_sum_importanceweightedpotentialstabilityscore_le_integral_sum_refinedbound_of_conddistrib_of_minimizers banditrlproof.tsallis.integral_sum_importanceweightedpotentialstabilityscore_le_integral_sum_refinedbound_of_conddistrib_of_minimizers finite-horizon expected conjugate-potential stability under identified conditional action laws and exact current/update minimizer certificates. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryActionPotentialStabilityAt","label":"sampledHalfTsallisHistoryActionPotentialStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryActionPotentialStabilityAt","description":"The canonical generated one-round potential score on a visible environment/prefix and sampled successor action.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-0de9cbb869ed","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9542,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:739"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : (Env × History.FinitePairHistory Action Real n) × Action -> Real","missing":[],"search":"sampledhalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.sampledhalftsallishistoryactionpotentialstabilityat the canonical generated one-round potential score on a visible environment/prefix and sampled successor action. definition compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisRefinedPotentialStabilityBoundAt","label":"sampledHalfTsallisRefinedPotentialStabilityBoundAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisRefinedPotentialStabilityBoundAt","description":"The generated visible-prefix refined potential budget.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-1a9c4414827a","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9543,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:752"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisRefinedPotentialStabilityBoundAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Real","missing":[],"search":"sampledhalftsallisrefinedpotentialstabilityboundat banditrlproof.tsallis.sampledhalftsallisrefinedpotentialstabilityboundat the generated visible-prefix refined potential budget. definition compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryActionPotentialStabilityAt","label":"measurable_sampledHalfTsallisHistoryActionPotentialStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryActionPotentialStabilityAt","description":"The generated canonical potential score is measurable.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-4cea29fc6ff0","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9544,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:760"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Measurable (sampledHalfTsallisHistoryActionPotentialStabilityAt arms harms eta loss n)","missing":[],"search":"measurable_sampledhalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.measurable_sampledhalftsallishistoryactionpotentialstabilityat the generated canonical potential score is measurable. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisRefinedPotentialStabilityBoundAt","label":"integrable_sampledHalfTsallisRefinedPotentialStabilityBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledHalfTsallisRefinedPotentialStabilityBoundAt","description":"The generated refined potential budget is automatically integrable.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-7689512d20e3","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9545,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:787"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledHalfTsallisRefinedPotentialStabilityBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (n : Nat) (historyMu : Measure (Env × History.FinitePairHistory Action Real n)) [IsFiniteMeasure historyMu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) : Integrable (sampledHalfTsallisRefinedPotentialStabilityBoundAt (Env := Env) arms harms eta n) historyMu","missing":[],"search":"integrable_sampledhalftsallisrefinedpotentialstabilityboundat banditrlproof.tsallis.integrable_sampledhalftsallisrefinedpotentialstabilityboundat the generated refined potential budget is automatically integrable. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisHistoryActionPotentialStabilityAt","label":"integrable_sampledHalfTsallisHistoryActionPotentialStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledHalfTsallisHistoryActionPotentialStabilityAt","description":"The generated canonical potential score is automatically integrable under the visible-history/action product law.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-17b2f55d3ed0","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9546,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:814"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : let selector := canonicalHalfTsallisGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (sampledHalfTsallisHistoryActionPotentialStabilityAt arms harms eta loss n) (mu.map (sampledHalfTsallisHistoryAt n) ⊗ₘ sampledHalfTsallisPolicyAt (Env := Env) arms harms eta selector.finiteHistory n)","missing":[],"search":"integrable_sampledhalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.integrable_sampledhalftsallishistoryactionpotentialstabilityat the generated canonical potential score is automatically integrable under the visible-history/action product law. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisSuccessorPotentialStabilityAt","label":"sampledHalfTsallisSuccessorPotentialStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisSuccessorPotentialStabilityAt","description":"The actual generated successor potential step, written with the next prefix's current selector.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-9eecd8b416e4","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9547,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:870"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisSuccessorPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallissuccessorpotentialstabilityat banditrlproof.tsallis.sampledhalftsallissuccessorpotentialstabilityat the actual generated successor potential step, written with the next prefix's current selector. definition compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorPotentialStability_le_integral_sum_refinedPotentialStabilityBound_canonical","label":"integral_sum_sampledHalfTsallisSuccessorPotentialStability_le_integral_sum_refinedPotentialStabilityBound_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorPotentialStability_le_integral_sum_refinedPotentialStabilityBound_canonical","description":"Generated canonical finite-horizon successor conjugate-potential stability. The trajectory action law, score recursion, selector measurability, score integrability, and refined-budget integrability are all discharged internally.","url":"../modules/banditrlproof-tsallisconjugatepotentialfinitehorizon/index.html#decl-7cdad013c584","parent":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","order":9548,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialFiniteHorizon"],["Source","BanditRLProof/TsallisConjugatePotentialFiniteHorizon.lean:894"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledHalfTsallisSuccessorPotentialStability_le_integral_sum_refinedPotentialStabilityBound_canonical {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun n => sampledHalfTsallisSuccessorPotentialStabilityAt arms harms eta loss n sample)) <= integral mu (f…","missing":[],"search":"integral_sum_sampledhalftsallissuccessorpotentialstability_le_integral_sum_refinedpotentialstabilitybound_canonical banditrlproof.tsallis.integral_sum_sampledhalftsallissuccessorpotentialstability_le_integral_sum_refinedpotentialstabilitybound_canonical generated canonical finite-horizon successor conjugate-potential stability. the trajectory action law, score recursion, selector measurability, score integrability, and refined-budget integrability are all discharged internally. theorem compiled","shard":"modules/95b2d9c1b7c75723.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue","label":"halfTsallisPotentialValue","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialValue","description":"The paper-normalized negative regularized-objective value. When `probability` is a certified minimizer and `eta > 0`, this is the constrained half-Tsallis potential under the local-to-paper learning-rate translation.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-8e099cb3cadb","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9549,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisPotentialValue {Action : Type u} (arms : Finset Action) (eta : Real) (score probability : Action -> Real) : Real","missing":[],"search":"halftsallispotentialvalue banditrlproof.tsallis.halftsallispotentialvalue the paper-normalized negative regularized-objective value. when `probability` is a certified minimizer and `eta > 0`, this is the constrained half-tsallis potential under the local-to-paper learning-rate translation. definition compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability","label":"halfTsallisPotentialStability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialStability","description":"The one-step candidate-potential expression before taking a sampling average. The low-level feasible-next bridge treats `next` only as a candidate point; the algorithm-facing theorems require both current and next minimizer certificates.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-bd859f7b7541","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9550,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisPotentialStability {Action : Type u} (arms : Finset Action) (eta : Real) (score probability estimate next : Action -> Real) : Real","missing":[],"search":"halftsallispotentialstability banditrlproof.tsallis.halftsallispotentialstability the one-step candidate-potential expression before taking a sampling average. the low-level feasible-next bridge treats `next` only as a candidate point; the algorithm-facing theorems require both current and next minimizer certificates. definition compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement","label":"halfTsallisConjugateCoordinateIncrement","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement","description":"The coordinate increment obtained from the explicit unconstrained conjugate. The shift is `estimate action - baseline`.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-bcfe785df2ff","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9551,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:62"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisConjugateCoordinateIncrement (eta probability shift : Real) : Real","missing":[],"search":"halftsallisconjugatecoordinateincrement banditrlproof.tsallis.halftsallisconjugatecoordinateincrement the coordinate increment obtained from the explicit unconstrained conjugate. the shift is `estimate action - baseline`. definition compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisConjugatePotentialUpper","label":"halfTsallisConjugatePotentialUpper","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisConjugatePotentialUpper","description":"Finite sum of explicit conjugate coordinate increments.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-ffab631c9494","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9552,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:70"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisConjugatePotentialUpper {Action : Type u} (arms : Finset Action) (eta : Real) (probability estimate : Action -> Real) (baseline : Real) : Real","missing":[],"search":"halftsallisconjugatepotentialupper banditrlproof.tsallis.halftsallisconjugatepotentialupper finite sum of explicit conjugate coordinate increments. definition compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.one_add_eta_mul_shift_mul_sqrt_pos","label":"one_add_eta_mul_shift_mul_sqrt_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.one_add_eta_mul_shift_mul_sqrt_pos","description":"theorem one_add_eta_mul_shift_mul_sqrt_pos {eta probability shift : Real} (_heta : 0 < eta) (hprobability : 0 < probability) (hdomain : -1 <= 2 * eta * shift * Real.sqrt probability) : 0 < 1 + eta * shift * Real.sqrt probability","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-a3a32ddb492b","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9553,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:78"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem one_add_eta_mul_shift_mul_sqrt_pos {eta probability shift : Real} (_heta : 0 < eta) (hprobability : 0 < probability) (hdomain : -1 <= 2 * eta * shift * Real.sqrt probability) : 0 < 1 + eta * shift * Real.sqrt probability","missing":[],"search":"one_add_eta_mul_shift_mul_sqrt_pos banditrlproof.tsallis.one_add_eta_mul_shift_mul_sqrt_pos theorem one_add_eta_mul_shift_mul_sqrt_pos {eta probability shift : real} (_heta : 0 < eta) (hprobability : 0 < probability) (hdomain : -1 <= 2 * eta * shift * real.sqrt probability) : 0 < 1 + eta * shift * real.sqrt probability theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement_eq","label":"halfTsallisConjugateCoordinateIncrement_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement_eq","description":"Exact rational form of the explicit conjugate coordinate increment.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-61d7d8dba76a","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9554,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:87"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisConjugateCoordinateIncrement_eq {eta probability shift : Real} (heta : 0 < eta) (hprobability : 0 < probability) (hdomain : -1 <= 2 * eta * shift * Real.sqrt probability) : halfTsallisConjugateCoordinateIncrement eta probability shift = eta * Real.sqrt probability * probability * shift ^ 2 / (1 + eta * shift * Real.sqrt probability)","missing":[],"search":"halftsallisconjugatecoordinateincrement_eq banditrlproof.tsallis.halftsallisconjugatecoordinateincrement_eq exact rational form of the explicit conjugate coordinate increment. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement_le","label":"halfTsallisConjugateCoordinateIncrement_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement_le","description":"Scalar translated Lemma 19 bound. The positive cubic term is active only when the shifted estimate is negative.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-2bd2d5e6e423","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9555,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisConjugateCoordinateIncrement_le {eta probability shift : Real} (heta : 0 < eta) (hprobability : 0 < probability) (hdomain : -1 <= 2 * eta * shift * Real.sqrt probability) : halfTsallisConjugateCoordinateIncrement eta probability shift <= eta * Real.sqrt probability * probability * shift ^ 2 + 2 * eta ^ 2 * probability ^ 2 * (max (-shift) 0) ^ 3","missing":[],"search":"halftsallisconjugatecoordinateincrement_le banditrlproof.tsallis.halftsallisconjugatecoordinateincrement_le scalar translated lemma 19 bound. the positive cubic term is active only when the shifted estimate is negative. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallis_fenchelCoordinate_le","label":"halfTsallis_fenchelCoordinate_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallis_fenchelCoordinate_le","description":"Coordinatewise Fenchel upper bound for a nonnegative competitor weight.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-e40ad8e80657","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9556,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:184"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallis_fenchelCoordinate_le {eta probability competitor shift : Real} (heta : 0 < eta) (hprobability : 0 < probability) (hcompetitor : 0 <= competitor) (hdomain : -1 <= 2 * eta * shift * Real.sqrt probability) : -competitor / (eta * Real.sqrt probability) - competitor * shift + 2 * Real.sqrt competitor / eta <= Real.sqrt probability / (eta * (1 + eta * shift * Real.sqrt probability))","missing":[],"search":"halftsallis_fenchelcoordinate_le banditrlproof.tsallis.halftsallis_fenchelcoordinate_le coordinatewise fenchel upper bound for a nonnegative competitor weight. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisConjugatePotentialUpper_le_shiftedMoments","label":"halfTsallisConjugatePotentialUpper_le_shiftedMoments","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisConjugatePotentialUpper_le_shiftedMoments","description":"The explicit conjugate finite sum is bounded by shifted quadratic/cubic moments.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-047ea0306fa0","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9557,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:224"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisConjugatePotentialUpper_le_shiftedMoments {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (probability estimate : Action -> Real) (baseline : Real) (heta : 0 < eta) (hprobability : forall action, action ∈ arms -> 0 < probability action) (hdomain : forall action, action ∈ arms -> -1 <= 2 * eta * (estimate action - baseline) * Real.sqrt (probability action)) : halfTsallisConjugatePotentialUpper arms eta probability estimate baseline <= eta * arms.sum (fun action => Real.sqrt (probability action) * probability action * (estimate action - baseline) ^ 2) + 2 * eta ^ 2 * arms.sum (fun action => probability action ^ 2 * (max (baseline - estimate action) 0) ^ 3)","missing":[],"search":"halftsallisconjugatepotentialupper_le_shiftedmoments banditrlproof.tsallis.halftsallisconjugatepotentialupper_le_shiftedmoments the explicit conjugate finite sum is bounded by shifted quadratic/cubic moments. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.importanceWeightedLoss_sub_selectedLoss_conjugate_domain","label":"importanceWeightedLoss_sub_selectedLoss_conjugate_domain","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.importanceWeightedLoss_sub_selectedLoss_conjugate_domain","description":"Ordinary importance-weighted estimates satisfy the translated conjugate domain.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-f405d0ef54cd","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9558,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:275"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem importanceWeightedLoss_sub_selectedLoss_conjugate_domain {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (probability loss : Action -> Real) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (hprobability : FTRL.finiteSimplex arms probability) (hprobabilityPos : forall action, action ∈ arms -> 0 < probability action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) {chosen action : Action} (hchosen : chosen ∈ arms) (haction : action ∈ arms) : -1 <= 2 * eta * (Exp3.importanceWeightedLoss probability loss chosen action - loss chosen) * Real.sqrt (probability action)","missing":[],"search":"importanceweightedloss_sub_selectedloss_conjugate_domain banditrlproof.tsallis.importanceweightedloss_sub_selectedloss_conjugate_domain ordinary importance-weighted estimates satisfy the translated conjugate domain. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_sub_linearLoss_eq_sum_sub_mul","label":"linearLoss_sub_linearLoss_eq_sum_sub_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_sub_linearLoss_eq_sum_sub_mul","description":"theorem linearLoss_sub_linearLoss_eq_sum_sub_mul {Action : Type u} (arms : Finset Action) (p q score : Action -> Real) : FTRL.linearLoss arms p score - FTRL.linearLoss arms q score = arms.sum (fun action => (p action - q action) * score action)","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-cd5c0914e223","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9559,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:329"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_sub_linearLoss_eq_sum_sub_mul {Action : Type u} (arms : Finset Action) (p q score : Action -> Real) : FTRL.linearLoss arms p score - FTRL.linearLoss arms q score = arms.sum (fun action => (p action - q action) * score action)","missing":[],"search":"linearloss_sub_linearloss_eq_sum_sub_mul banditrlproof.tsallis.linearloss_sub_linearloss_eq_sum_sub_mul theorem linearloss_sub_linearloss_eq_sum_sub_mul {action : type u} (arms : finset action) (p q score : action -> real) : ftrl.linearloss arms p score - ftrl.linearloss arms q score = arms.sum (fun action => (p action - q action) * score action) theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sub_mul_sub_baseline_eq_sum_sub_mul","label":"sum_sub_mul_sub_baseline_eq_sum_sub_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sub_mul_sub_baseline_eq_sum_sub_mul","description":"theorem sum_sub_mul_sub_baseline_eq_sum_sub_mul {Action : Type u} (arms : Finset Action) (p q value : Action -> Real) (baseline : Real) (hp : arms.sum p = 1) (hq : arms.sum q = 1) : arms.sum (fun action => (p action - q action) * (value action - baseline)) = arms.sum (fun action => (p action - q action) * value action)","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-0bf53785d6fb","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9560,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:340"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sub_mul_sub_baseline_eq_sum_sub_mul {Action : Type u} (arms : Finset Action) (p q value : Action -> Real) (baseline : Real) (hp : arms.sum p = 1) (hq : arms.sum q = 1) : arms.sum (fun action => (p action - q action) * (value action - baseline)) = arms.sum (fun action => (p action - q action) * value action)","missing":[],"search":"sum_sub_mul_sub_baseline_eq_sum_sub_mul banditrlproof.tsallis.sum_sub_mul_sub_baseline_eq_sum_sub_mul theorem sum_sub_mul_sub_baseline_eq_sum_sub_mul {action : type u} (arms : finset action) (p q value : action -> real) (baseline : real) (hp : arms.sum p = 1) (hq : arms.sum q = 1) : arms.sum (fun action => (p action - q action) * (value action - baseline)) = arms.sum (fun action => (p action - q action) * value action) theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_sub_linearLoss_score_eq_sum_div_sqrt_of_stationary","label":"linearLoss_sub_linearLoss_score_eq_sum_div_sqrt_of_stationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_sub_linearLoss_score_eq_sum_div_sqrt_of_stationary","description":"Stationarity removes the score and its common simplex multiplier.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-0391c6856787","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9561,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:360"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_sub_linearLoss_score_eq_sum_div_sqrt_of_stationary {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p q : Action -> Real) (multiplier : Real) (heta : 0 < eta) (hp : FTRL.finiteSimplex arms p) (hq : FTRL.finiteSimplex arms q) (hpPos : forall action, action ∈ arms -> 0 < p action) (hstationary : HalfTsallisInteriorStationary arms eta score p multiplier) : FTRL.linearLoss arms p score - FTRL.linearLoss arms q score = arms.sum (fun action => (p action - q action) / (eta * Real.sqrt (p action)))","missing":[],"search":"linearloss_sub_linearloss_score_eq_sum_div_sqrt_of_stationary banditrlproof.tsallis.linearloss_sub_linearloss_score_eq_sum_div_sqrt_of_stationary stationarity removes the score and its common simplex multiplier. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_le_conjugatePotentialUpper_of_feasible","label":"halfTsallisPotentialStability_le_conjugatePotentialUpper_of_feasible","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialStability_le_conjugatePotentialUpper_of_feasible","description":"The feasible-next candidate-potential expression is bounded by the explicit unconstrained conjugate sum. Only the current point needs stationarity. This is an algebraic bridge, not an actual constrained-potential theorem unless the caller separately certifies that `next` minimizes the updated objective.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-af468c63afd1","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9562,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:426"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialStability_le_conjugatePotentialUpper_of_feasible {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score probability estimate next : Action -> Real) (baseline multiplier : Real) (heta : 0 < eta) (hprobability : FTRL.finiteSimplex arms probability) (hnext : FTRL.finiteSimplex arms next) (hprobabilityPos : forall action, action ∈ arms -> 0 < probability action) (hstationary : HalfTsallisInteriorStationary arms eta score probability multiplier) (hdomain : forall action, action ∈ arms -> -1 <= 2 * eta * (estimate action - baseline) * Real.sqrt (probability action)) : halfTsallisPotentialStability arms eta score probability estimate next <= halfTsallisConjugatePotentialUpper arms eta probability estimate baseline","missing":[],"search":"halftsallispotentialstability_le_conjugatepotentialupper_of_feasible banditrlproof.tsallis.halftsallispotentialstability_le_conjugatepotentialupper_of_feasible the feasible-next candidate-potential expression is bounded by the explicit unconstrained conjugate sum. only the current point needs stationarity. this is an algebraic bridge, not an actual constrained-potential theorem unless the caller separately certifies that `next` minimizes the updated objective. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_importanceWeightedLoss_le_shiftedMoments_of_minimizers","label":"halfTsallisPotentialStability_importanceWeightedLoss_le_shiftedMoments_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialStability_importanceWeightedLoss_le_shiftedMoments_of_minimizers","description":"Paper-faithful one-step conjugate-potential stability for an ordinary importance-weighted sampled update.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-ca7b65eb15f9","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9563,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:539"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialStability_importanceWeightedLoss_le_shiftedMoments_of_minimizers {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score probability loss next : Action -> Real) (chosen : Action) (hchosen : chosen ∈ arms) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (hprobabilityMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score probability) (hnextMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss probability loss chosen action) next) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : halfTsallisPotentialStability arms eta score probability (Exp3.importanceWeightedLoss probability loss chosen) next <= eta * shiftedHalfPowerImportanceWeighte…","missing":[],"search":"halftsallispotentialstability_importanceweightedloss_le_shiftedmoments_of_minimizers banditrlproof.tsallis.halftsallispotentialstability_importanceweightedloss_le_shiftedmoments_of_minimizers paper-faithful one-step conjugate-potential stability for an ordinary importance-weighted sampled update. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_refined_of_minimizers","label":"sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_refined_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_refined_of_minimizers","description":"Sampled-action averaged ordinary-IW conjugate-potential stability. This is the deterministic finite-action form of the refined coefficient used downstream by the self-bounding route.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-a00f14a22be2","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9564,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:603"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_refined_of_minimizers {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score probability loss : Action -> Real) (next : Action -> Action -> Real) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (hprobabilityMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score probability) (hnextMin : forall chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss probability loss chosen action) (next chosen)) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => probability chosen * halfTsallisPotentialStability arms eta score probability (Exp3.importanceWeighted…","missing":[],"search":"sum_prob_mul_halftsallispotentialstability_importanceweightedloss_le_refined_of_minimizers banditrlproof.tsallis.sum_prob_mul_halftsallispotentialstability_importanceweightedloss_le_refined_of_minimizers sampled-action averaged ordinary-iw conjugate-potential stability. this is the deterministic finite-action form of the refined coefficient used downstream by the self-bounding route. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisMinimizer_mul_potentialStability_le_refined","label":"sum_halfTsallisMinimizer_mul_potentialStability_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisMinimizer_mul_potentialStability_le_refined","description":"Canonical-selector wrapper with all minimizer certificates discharged.","url":"../modules/banditrlproof-tsallisconjugatepotentialstability/index.html#decl-f23cda949fde","parent":"module:BanditRLProof.TsallisConjugatePotentialStability","order":9565,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConjugatePotentialStability"],["Source","BanditRLProof/TsallisConjugatePotentialStability.lean:692"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisMinimizer_mul_potentialStability_le_refined {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : Action -> Real) (heta : 0 < eta) (heta_le : eta <= 1 / 2) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : let probability := halfTsallisMinimizer arms harms eta score let next := fun chosen => halfTsallisUpdatedMinimizer arms harms eta score loss chosen arms.sum (fun chosen => probability chosen * halfTsallisPotentialStability arms eta score probability (Exp3.importanceWeightedLoss probability loss chosen) (next chosen)) <= eta * arms.sum (fun action => Real.sqrt (probability action) * (1 - probability action)) + 2 * eta ^ 2","missing":[],"search":"sum_halftsallisminimizer_mul_potentialstability_le_refined banditrlproof.tsallis.sum_halftsallisminimizer_mul_potentialstability_le_refined canonical-selector wrapper with all minimizer certificates discharged. theorem compiled","shard":"modules/36d26052d5725552.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linear_sub_quadratic_le_sq_div_four","label":"linear_sub_quadratic_le_sq_div_four","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linear_sub_quadratic_le_sq_div_four","description":"A downward quadratic is bounded by its unconstrained vertex value.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-a768aad0c8b6","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9566,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:18"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linear_sub_quadratic_le_sq_div_four (b x c : Real) (hc : 0 < c) : b * x - c * x ^ 2 ≤ b ^ 2 / (4 * c)","missing":[],"search":"linear_sub_quadratic_le_sq_div_four banditrlproof.tsallis.linear_sub_quadratic_le_sq_div_four a downward quadratic is bounded by its unconstrained vertex value. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_linear_sub_quadratic_le_unconstrained","label":"sum_linear_sub_quadratic_le_unconstrained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_linear_sub_quadratic_le_unconstrained","description":"Coordinatewise completion of squares gives the unconstrained finite-sum branch of the paper's quadratic optimization.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-4cd4621c3fef","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9567,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:26"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_linear_sub_quadratic_le_unconstrained {Index : Type u} [DecidableEq Index] (indices : Finset Index) (b : Real) (c x : Index → Real) (hc : ∀ i ∈ indices, 0 < c i) : indices.sum (fun i => b * x i - c i * x i ^ 2) ≤ b ^ 2 / 4 * indices.sum (fun i => 1 / c i)","missing":[],"search":"sum_linear_sub_quadratic_le_unconstrained banditrlproof.tsallis.sum_linear_sub_quadratic_le_unconstrained coordinatewise completion of squares gives the unconstrained finite-sum branch of the paper's quadratic optimization. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_linear_sub_quadratic_le_of_sum_le","label":"sum_linear_sub_quadratic_le_of_sum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_linear_sub_quadratic_le_of_sum_le","description":"If the unconstrained vertex lies beyond a finite `sum x ≤ M` constraint, shift the common linear coefficient and apply completion of squares to obtain the active-constraint branch.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-5e7f9da75638","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9568,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_linear_sub_quadratic_le_of_sum_le {Index : Type u} [DecidableEq Index] (indices : Finset Index) (b M : Real) (c x : Index → Real) (hc : ∀ i ∈ indices, 0 < c i) (hreciprocal : 0 < indices.sum (fun i => 1 / c i)) (hxsum : indices.sum x ≤ M) (hthreshold : 2 * M ≤ b * indices.sum (fun i => 1 / c i)) : indices.sum (fun i => b * x i - c i * x i ^ 2) ≤ b * M - M ^ 2 / indices.sum (fun i => 1 / c i)","missing":[],"search":"sum_linear_sub_quadratic_le_of_sum_le banditrlproof.tsallis.sum_linear_sub_quadratic_le_of_sum_le if the unconstrained vertex lies beyond a finite `sum x ≤ m` constraint, shift the common linear coefficient and apply completion of squares to obtain the active-constraint branch. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_inv_pos_of_nonempty","label":"sum_inv_pos_of_nonempty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_inv_pos_of_nonempty","description":"Positive coefficients on a nonempty finite set have a positive reciprocal sum.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-21af5d866ae7","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9569,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:88"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_inv_pos_of_nonempty {Index : Type u} [DecidableEq Index] (indices : Finset Index) (hindices : indices.Nonempty) (c : Index → Real) (hc : ∀ i ∈ indices, 0 < c i) : 0 < indices.sum (fun i => 1 / c i)","missing":[],"search":"sum_inv_pos_of_nonempty banditrlproof.tsallis.sum_inv_pos_of_nonempty positive coefficients on a nonempty finite set have a positive reciprocal sum. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_le_sqrt_card","label":"sum_erase_sqrt_le_sqrt_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_erase_sqrt_le_sqrt_card","description":"A finite simplex point has suboptimal square-root mass at most the square root of the number of suboptimal coordinates.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-1c1f7a66ac75","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9570,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:100"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_erase_sqrt_le_sqrt_card {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability : Action → Real) (hprobability : FTRL.finiteSimplex arms probability) : (arms.erase best).sum (fun action => Real.sqrt (probability action)) ≤ Real.sqrt ((arms.erase best).card : Real)","missing":[],"search":"sum_erase_sqrt_le_sqrt_card banditrlproof.tsallis.sum_erase_sqrt_le_sqrt_card a finite simplex point has suboptimal square-root mass at most the square root of the number of suboptimal coordinates. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_probability_sub_gap_le_unconstrained","label":"sum_erase_sqrt_probability_sub_gap_le_unconstrained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_erase_sqrt_probability_sub_gap_le_unconstrained","description":"The unconstrained one-round branch after substituting `x action = sqrt (probability action)` and `c action = lambda * gap action`.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-25fa597d5b56","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9571,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:154"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_erase_sqrt_probability_sub_gap_le_unconstrained {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (probability gap : Action → Real) (hprobability : FTRL.finiteSimplex arms probability) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (b lambda : Real) (hlambda : 0 < lambda) : (arms.erase best).sum (fun action => b * Real.sqrt (probability action) - lambda * gap action * probability action) ≤ b ^ 2 / 4 * (arms.erase best).sum (fun action => 1 / (lambda * gap action))","missing":[],"search":"sum_erase_sqrt_probability_sub_gap_le_unconstrained banditrlproof.tsallis.sum_erase_sqrt_probability_sub_gap_le_unconstrained the unconstrained one-round branch after substituting `x action = sqrt (probability action)` and `c action = lambda * gap action`. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_probability_sub_gap_le_of_threshold","label":"sum_erase_sqrt_probability_sub_gap_le_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_erase_sqrt_probability_sub_gap_le_of_threshold","description":"On the active-constraint branch, finite-simplex square-root mass sharpens the unconstrained coordinatewise bound by the common mass constraint.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-48aec5999b9d","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9572,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:190"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_erase_sqrt_probability_sub_gap_le_of_threshold {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability gap : Action → Real) (hprobability : FTRL.finiteSimplex arms probability) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (b lambda : Real) (hlambda : 0 < lambda) (hthreshold : 2 * Real.sqrt ((arms.erase best).card : Real) ≤ b * (arms.erase best).sum (fun action => 1 / (lambda * gap action))) : (arms.erase best).sum (fun action => b * Real.sqrt (probability action) - lambda * gap action * probability action) ≤ b * Real.sqrt ((arms.erase best).card : Real) - (arms.erase best).card / (arms.erase best).sum (fun action => 1 / (lambda * gap action))","missing":[],"search":"sum_erase_sqrt_probability_sub_gap_le_of_threshold banditrlproof.tsallis.sum_erase_sqrt_probability_sub_gap_le_of_threshold on the active-constraint branch, finite-simplex square-root mass sharpens the unconstrained coordinatewise bound by the common mass constraint. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbability_sum_le_unconstrained","label":"sampledScheduledHalfTsallisExpectedProbability_sum_le_unconstrained","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbability_sum_le_unconstrained","description":"Generated expected-probability specialization of the unconstrained one-round quadratic branch. All measure and Jensen obligations are discharged by the existing finite-simplex expectation theorem.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-d04f3e3e76d5","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9573,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:251"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisExpectedProbability_sum_le_unconstrained {Env : Type u} {Action : Type*} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : MeasureTheory.Measure (Env × ((k : Nat) → Action × Real))) [MeasureTheory.IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (t : Nat) {best : Action} (gap : Action → Real) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (b lambda : Real) (hlambda : 0 < lambda) : (arms.erase best).sum (fun action => b * Real.sqrt (sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action) - lambda * gap action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action) ≤ b ^ 2 / 4 * (arms.erase best).sum (fun action => 1 / (lambda * gap action))","missing":[],"search":"sampledscheduledhalftsallisexpectedprobability_sum_le_unconstrained banditrlproof.tsallis.sampledscheduledhalftsallisexpectedprobability_sum_le_unconstrained generated expected-probability specialization of the unconstrained one-round quadratic branch. all measure and jensen obligations are discharged by the existing finite-simplex expectation theorem. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbability_sum_le_of_threshold","label":"sampledScheduledHalfTsallisExpectedProbability_sum_le_of_threshold","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbability_sum_le_of_threshold","description":"Generated expected-probability specialization of the active-constraint one-round quadratic branch. This is the direct consumer for the later time-threshold split in the improved self-bounding route.","url":"../modules/banditrlproof-tsallisconstrainedquadraticoptimization/index.html#decl-d80cfb2a963d","parent":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","order":9574,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisConstrainedQuadraticOptimization"],["Source","BanditRLProof/TsallisConstrainedQuadraticOptimization.lean:282"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisExpectedProbability_sum_le_of_threshold {Env : Type u} {Action : Type*} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : MeasureTheory.Measure (Env × ((k : Nat) → Action × Real))) [MeasureTheory.IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (t : Nat) {best : Action} (hbest : best ∈ arms) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (b lambda : Real) (hlambda : 0 < lambda) (hthreshold : 2 * Real.sqrt ((arms.erase best).card : Real) ≤ b * (arms.erase best).sum (fun action => 1 / (lambda * gap action))) : (arms.erase best).sum (fun action => b * Real.sqrt (sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action) - lambda * gap action * sampledScheduledHalfTsallis…","missing":[],"search":"sampledscheduledhalftsallisexpectedprobability_sum_le_of_threshold banditrlproof.tsallis.sampledscheduledhalftsallisexpectedprobability_sum_le_of_threshold generated expected-probability specialization of the active-constraint one-round quadratic branch. this is the direct consumer for the later time-threshold split in the improved self-bounding route. theorem compiled","shard":"modules/0092d0a097e173b2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.importanceWeightedStabilityScore","label":"importanceWeightedStabilityScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.importanceWeightedStabilityScore","description":"The one-round FTRL stability score generated by an importance-weighted loss estimate.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-36e1795ff09f","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9575,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def importanceWeightedStabilityScore {History : Type u} {Action : Type v} (arms : Finset Action) (prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (sample : History × Action) : Real","missing":[],"search":"importanceweightedstabilityscore banditrlproof.tsallis.importanceweightedstabilityscore the one-round ftrl stability score generated by an importance-weighted loss estimate. definition compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfPowerStabilityBound","label":"halfPowerStabilityBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfPowerStabilityBound","description":"The pointwise half-Tsallis upper bound for the sampling-law averaged one-round stability score.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-794e53fc0ba8","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9576,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfPowerStabilityBound {History : Type u} {Action : Type v} (arms : Finset Action) (eta : Real) (prob : History -> Action -> Real) (history : History) : Real","missing":[],"search":"halfpowerstabilitybound banditrlproof.tsallis.halfpowerstabilitybound the pointwise half-tsallis upper bound for the sampling-law averaged one-round stability score. definition compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_condDistrib_of_minimizers","label":"integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_condDistrib_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_condDistrib_of_minimizers","description":"An identified finite conditional action law transports the pointwise half-Tsallis minimizer stability theorem to a one-round integral inequality. The score and bound integrability hypotheses are the exact analytic boundary needed by Bochner integral monotonicity. No measurability of a particular minimizer selector is inferred here.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-4d8f58cd9f27","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9577,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:54"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_condDistrib_of_minimizers {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => Exp3.finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (heta : 0 < eta) (hprobMin : forall h, FTRL.IsRegularizedM…","missing":[],"search":"integral_importanceweightedstabilityscore_le_integral_halfpowerstabilitybound_of_conddistrib_of_minimizers banditrlproof.tsallis.integral_importanceweightedstabilityscore_le_integral_halfpowerstabilitybound_of_conddistrib_of_minimizers an identified finite conditional action law transports the pointwise half-tsallis minimizer stability theorem to a one-round integral inequality. the score and bound integrability hypotheses are the exact analytic boundary needed by bochner integral monotonicity. no measurability of a particular minimizer selector is inferred here. theorem compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.actionProcess_integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_minimizers","label":"actionProcess_integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.actionProcess_integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_minimizers","description":"Generated finite-action process specialization of the conditional half-Tsallis stability transport.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-73fc74608f37","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9578,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:143"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_minimizers {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (heta : 0 < eta) (hprobMin : forall h, FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (score h) (prob h)) (hnextMin : forall h chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun candidate => score h candi…","missing":[],"search":"actionprocess_integral_importanceweightedstabilityscore_le_integral_halfpowerstabilitybound_of_minimizers banditrlproof.tsallis.actionprocess_integral_importanceweightedstabilityscore_le_integral_halfpowerstabilitybound_of_minimizers generated finite-action process specialization of the conditional half-tsallis stability transport. theorem compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisHistoryMinimizer","label":"halfTsallisHistoryMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisHistoryMinimizer","description":"The fixed half-Tsallis minimizer viewed as a history-indexed action distribution.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-618cef511ac1","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9579,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:198"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisHistoryMinimizer {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) : History -> Action -> Real","missing":[],"search":"halftsallishistoryminimizer banditrlproof.tsallis.halftsallishistoryminimizer the fixed half-tsallis minimizer viewed as a history-indexed action distribution. definition compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisHistoryUpdatedMinimizer","label":"halfTsallisHistoryUpdatedMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisHistoryUpdatedMinimizer","description":"The fixed importance-weighted half-Tsallis update viewed as a history/action-indexed family.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-d734db019b00","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9580,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:207"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisHistoryUpdatedMinimizer {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : History -> Action -> Real) : History -> Action -> Action -> Real","missing":[],"search":"halftsallishistoryupdatedminimizer banditrlproof.tsallis.halftsallishistoryupdatedminimizer the fixed importance-weighted half-tsallis update viewed as a history/action-indexed family. definition compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurableFiniteActionDistribution_halfTsallisHistoryMinimizer","label":"measurableFiniteActionDistribution_halfTsallisHistoryMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurableFiniteActionDistribution_halfTsallisHistoryMinimizer","description":"Coordinate measurability is the only missing contract for turning the fixed history-indexed minimizer into the project's finite-action kernel source. Distribution feasibility follows from the minimizer certificate.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-1dc7010ecf81","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9581,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:219"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def measurableFiniteActionDistribution_halfTsallisHistoryMinimizer {History : Type u} {Action : Type v} [MeasurableSpace History] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : History -> Action -> Real) (hmeasurable : forall action, action ∈ arms -> Measurable (fun history => halfTsallisHistoryMinimizer arms harms eta score history action)) : Exp3.MeasurableFiniteActionDistribution arms (halfTsallisHistoryMinimizer arms harms eta score) where","missing":[],"search":"measurablefiniteactiondistribution_halftsallishistoryminimizer banditrlproof.tsallis.measurablefiniteactiondistribution_halftsallishistoryminimizer coordinate measurability is the only missing contract for turning the fixed history-indexed minimizer into the project's finite-action kernel source. distribution feasibility follows from the minimizer certificate. definition compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.actionProcess_integral_halfTsallisHistoryStability_le_integral_halfPowerStabilityBound","label":"actionProcess_integral_halfTsallisHistoryStability_le_integral_halfPowerStabilityBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.actionProcess_integral_halfTsallisHistoryStability_le_integral_halfPowerStabilityBound","description":"Canonical generated-action one-round stability endpoint. The minimizer certificates and finite-distribution laws are internal. The remaining caller contracts expose exactly what is not yet proved for the `Classical.choose` selector: coordinate measurability and regularity of the updated stability score.","url":"../modules/banditrlproof-tsallisftrlconditionalstability/index.html#decl-ce3a96d75370","parent":"module:BanditRLProof.TsallisFTRLConditionalStability","order":9582,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLConditionalStability"],["Source","BanditRLProof/TsallisFTRLConditionalStability.lean:245"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem actionProcess_integral_halfTsallisHistoryStability_le_integral_halfPowerStabilityBound {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : History -> Action -> Real) (hmeasurable : forall action, action ∈ arms -> Measurable (fun history => halfTsallisHistoryMinimizer arms harms eta score history action)) (heta : 0 < eta) (hloss : forall history candidate, candidate ∈ arms -> 0 <= loss history candidate ∧ loss history candidate <= 1) (hscore : Measurable (importanceWeightedStabilityScore arms (halfTsallisHistoryMinimizer arms harms eta score) loss (halfTsallisHistoryUpdatedMinimizer arms harms eta score loss))) (hIn…","missing":[],"search":"actionprocess_integral_halftsallishistorystability_le_integral_halfpowerstabilitybound banditrlproof.tsallis.actionprocess_integral_halftsallishistorystability_le_integral_halfpowerstabilitybound canonical generated-action one-round stability endpoint. the minimizer certificates and finite-distribution laws are internal. the remaining caller contracts expose exactly what is not yet proved for the `classical.choose` selector: coordinate measurability and regularity of the updated stability score. theorem compiled","shard":"modules/6f0056731c8550fd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisProbabilityAtTime","label":"sampledHalfTsallisProbabilityAtTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisProbabilityAtTime","description":"The pure half-Tsallis sampling probability at an actual trajectory time.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-9c1492318a7c","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9583,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Nat -> Env × ((k : Nat) -> Action × Real) -> Action -> Real | 0, _sample => initialHalfTsallisDistribution arms harms eta | n + 1, sample => sampledHalfTsallisHistoryDistribution arms harms eta n (Preorder.frestrictLe n sample.2) /-- The observed-scalar importance-weighted loss vector at an actual time. -/ noncomputable def sampledHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledhalftsallisprobabilityattime banditrlproof.tsallis.sampledhalftsallisprobabilityattime the pure half-tsallis sampling probability at an actual trajectory time. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedLossAt","label":"sampledHalfTsallisObservedEstimatedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedLossAt","description":"The observed-scalar importance-weighted loss vector at an actual time.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-21cc776dcf90","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9584,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:30"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledhalftsallisobservedestimatedlossat banditrlproof.tsallis.sampledhalftsallisobservedestimatedlossat the observed-scalar importance-weighted loss vector at an actual time. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.cumulativeLoss_sampledHalfTsallisObservedEstimatedLossAt_succ","label":"cumulativeLoss_sampledHalfTsallisObservedEstimatedLossAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.cumulativeLoss_sampledHalfTsallisObservedEstimatedLossAt_succ","description":"The canonical cumulative selector generated by the observed estimator is the actual pure half-Tsallis probability at the same time.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-188032f87603","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9585,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLoss_sampledHalfTsallisObservedEstimatedLossAt_succ {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (sample : Env × ((k : Nat) -> Action × Real)) (n : Nat) : FTRL.cumulativeLoss (fun t => sampledHalfTsallisObservedEstimatedLossAt arms harms eta t sample) (n + 1) = sampledHalfTsallisHistoryScore arms harms eta n (Preorder.frestrictLe n sample.2)","missing":[],"search":"cumulativeloss_sampledhalftsallisobservedestimatedlossat_succ banditrlproof.tsallis.cumulativeloss_sampledhalftsallisobservedestimatedlossat_succ the canonical cumulative selector generated by the observed estimator is the actual pure half-tsallis probability at the same time. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_observedEstimatedLoss_eq_probabilityAtTime","label":"halfTsallisCumulativeMinimizer_observedEstimatedLoss_eq_probabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_observedEstimatedLoss_eq_probabilityAtTime","description":"theorem halfTsallisCumulativeMinimizer_observedEstimatedLoss_eq_probabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) : halfTsallisCumulativeMinimizer arms harms eta (fun s => sampledHalfTsallisObservedEstimatedLossAt arms harms eta s sample) t = sampledHalfTsallisProbabilityAtTime ar…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-cbf4d8923f47","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9586,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:69"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisCumulativeMinimizer_observedEstimatedLoss_eq_probabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) : halfTsallisCumulativeMinimizer arms harms eta (fun s => sampledHalfTsallisObservedEstimatedLossAt arms harms eta s sample) t = sampledHalfTsallisProbabilityAtTime arms harms eta t sample","missing":[],"search":"halftsalliscumulativeminimizer_observedestimatedloss_eq_probabilityattime banditrlproof.tsallis.halftsalliscumulativeminimizer_observedestimatedloss_eq_probabilityattime theorem halftsalliscumulativeminimizer_observedestimatedloss_eq_probabilityattime {env : type u} {action : type v} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (sample : env × ((k : nat) -> action × real)) (t : nat) : halftsalliscumulativeminimizer arms harms eta (fun s => sampledhalftsallisobservedestimatedlossat arms harms eta s sample) t = sampledhalftsallisprobabilityattime arms harms eta t sample theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialStability","label":"sampledHalfTsallisInitialStability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisInitialStability","description":"The time-zero current-minus-updated stability term.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-0048852c0097","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9587,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:86"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisInitialStability {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallisinitialstability banditrlproof.tsallis.sampledhalftsallisinitialstability the time-zero current-minus-updated stability term. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedSuccessorStabilityAt","label":"sampledHalfTsallisObservedSuccessorStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisObservedSuccessorStabilityAt","description":"The pathwise successor stability term formed from the scalar reward stored in the generated trajectory.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-921f9222de5c","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9588,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:100"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisObservedSuccessorStabilityAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallisobservedsuccessorstabilityat banditrlproof.tsallis.sampledhalftsallisobservedsuccessorstabilityat the pathwise successor stability term formed from the scalar reward stored in the generated trajectory. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialUpdatedAt","label":"sampledHalfTsallisInitialUpdatedAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisInitialUpdatedAt","description":"Canonical time-zero update after sampling the initial action.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-7ca9ba12f417","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9589,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:114"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisInitialUpdatedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : Env -> Action -> Action -> Real","missing":[],"search":"sampledhalftsallisinitialupdatedat banditrlproof.tsallis.sampledhalftsallisinitialupdatedat canonical time-zero update after sampling the initial action. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialHistoryActionStability","label":"sampledHalfTsallisInitialHistoryActionStability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisInitialHistoryActionStability","description":"Canonical predictable stability score for the initial action.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-6a0e89ae80d3","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9590,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:124"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisInitialHistoryActionStability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : Env × Action -> Real","missing":[],"search":"sampledhalftsallisinitialhistoryactionstability banditrlproof.tsallis.sampledhalftsallisinitialhistoryactionstability canonical predictable stability score for the initial action. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_observedEstimated_stability_eq_initial_add_successor","label":"sum_observedEstimated_stability_eq_initial_add_successor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_observedEstimated_stability_eq_initial_add_successor","description":"The deterministic FTRL stability sum splits into the initial term and the generated successor terms.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-024f731b02e6","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9591,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:136"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_observedEstimated_stability_eq_initial_add_successor {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (sample : Env × ((k : Nat) -> Action × Real)) (horizon : Nat) : let estimatedLoss := fun t => sampledHalfTsallisObservedEstimatedLossAt arms harms eta t sample let p := halfTsallisCumulativeMinimizer arms harms eta estimatedLoss (Finset.range (horizon + 1)).sum (fun t => FTRL.linearLoss arms (p t) (estimatedLoss t) - FTRL.linearLoss arms (p (t + 1)) (estimatedLoss t)) = sampledHalfTsallisInitialStability arms harms eta sample + (Finset.range horizon).sum (fun n => sampledHalfTsallisObservedSuccessorStabilityAt arms harms eta n sample)","missing":[],"search":"sum_observedestimated_stability_eq_initial_add_successor banditrlproof.tsallis.sum_observedestimated_stability_eq_initial_add_successor the deterministic ftrl stability sum splits into the initial term and the generated successor terms. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallis_observedEstimatedRegret_le","label":"sampledHalfTsallis_observedEstimatedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallis_observedEstimatedRegret_le","description":"Pathwise estimated-loss regret decomposition for time zero followed by `horizon` generated successor rounds.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-f5e3aa2e3493","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9592,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:166"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallis_observedEstimatedRegret_le {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (sample : Env × ((k : Nat) -> Action × Real)) (horizon : Nat) : (Finset.range (horizon + 1)).sum (fun t => FTRL.linearLoss arms (sampledHalfTsallisProbabilityAtTime arms harms eta t sample) (sampledHalfTsallisObservedEstimatedLossAt arms harms eta t sample) - FTRL.linearLoss arms q (sampledHalfTsallisObservedEstimatedLossAt arms harms eta t sample)) <= sampledHalfTsallisInitialStability arms harms eta sample + (Finset.range horizon).sum (fun n => sampledHalfTsallisObservedSuccessorStabilityAt arms harms eta n sample) + ((powerSum arms (1 / 2 : Real) (initialHalfTsallisDistribution arms harms eta) - powerSum arms (1 / 2 : Real) q) / (1 - (1 / 2 : Real))…","missing":[],"search":"sampledhalftsallis_observedestimatedregret_le banditrlproof.tsallis.sampledhalftsallis_observedestimatedregret_le pathwise estimated-loss regret decomposition for time zero followed by `horizon` generated successor rounds. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedLossAt","label":"sampledHalfTsallisPredictableEstimatedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedLossAt","description":"The predictable importance-weighted loss vector at an actual trajectory time, using the same probability as the observed recursive update.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-4459f8b42a0f","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9593,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:202"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledhalftsallispredictableestimatedlossat banditrlproof.tsallis.sampledhalftsallispredictableestimatedlossat the predictable importance-weighted loss vector at an actual trajectory time, using the same probability as the observed recursive update. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedRegret","label":"sampledHalfTsallisObservedEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedRegret","description":"Estimated regret against a fixed comparator distribution, including time zero and `horizon` successor rounds.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-8c367f9a17c0","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9594,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:215"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisObservedEstimatedRegret {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallisobservedestimatedregret banditrlproof.tsallis.sampledhalftsallisobservedestimatedregret estimated regret against a fixed comparator distribution, including time zero and `horizon` successor rounds. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEnvironmentRegret","label":"sampledHalfTsallisPredictableEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableEnvironmentRegret","description":"Predictable environment regret for the generated pure half-Tsallis probabilities against a fixed comparator distribution.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-29d08cc7c62b","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9595,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:231"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisPredictableEnvironmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallispredictableenvironmentregret banditrlproof.tsallis.sampledhalftsallispredictableenvironmentregret predictable environment regret for the generated pure half-tsallis probabilities against a fixed comparator distribution. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedRegret","label":"sampledHalfTsallisPredictableEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedRegret","description":"The same finite-horizon estimated regret after replacing stored rewards by their predictable loss-vector coordinates.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-fd5040ac2906","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9596,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:246"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisPredictableEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallispredictableestimatedregret banditrlproof.tsallis.sampledhalftsallispredictableestimatedregret the same finite-horizon estimated regret after replacing stored rewards by their predictable loss-vector coordinates. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisProbabilityAtTime","label":"measurable_sampledHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisProbabilityAtTime","description":"theorem measurable_sampledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledHalfTsallisProbabilityAtTime arms…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-155a92865a28","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9597,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:262"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledHalfTsallisProbabilityAtTime arms harms eta t sample candidate)","missing":[],"search":"measurable_sampledhalftsallisprobabilityattime banditrlproof.tsallis.measurable_sampledhalftsallisprobabilityattime theorem measurable_sampledhalftsallisprobabilityattime {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledhalftsallisprobabilityattime arms harms eta t sample candidate) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisObservedEstimatedLossAt","label":"measurable_sampledHalfTsallisObservedEstimatedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisObservedEstimatedLossAt","description":"theorem measurable_sampledHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledHalfTsallisObservedEstimated…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-983054c2bebc","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9598,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:287"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledHalfTsallisObservedEstimatedLossAt arms harms eta t sample candidate)","missing":[],"search":"measurable_sampledhalftsallisobservedestimatedlossat banditrlproof.tsallis.measurable_sampledhalftsallisobservedestimatedlossat theorem measurable_sampledhalftsallisobservedestimatedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledhalftsallisobservedestimatedlossat arms harms eta t sample candidate) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisPredictableEstimatedLossAt","label":"measurable_sampledHalfTsallisPredictableEstimatedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisPredictableEstimatedLossAt","description":"theorem measurable_sampledHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Act…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-dd5ba582fa1b","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9599,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:306"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledHalfTsallisPredictableEstimatedLossAt arms harms eta loss t sample candidate)","missing":[],"search":"measurable_sampledhalftsallispredictableestimatedlossat banditrlproof.tsallis.measurable_sampledhalftsallispredictableestimatedlossat theorem measurable_sampledhalftsallispredictableestimatedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : exp3.predictablelossvector env action) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledhalftsallispredictableestimatedlossat arms harms eta loss t sample candidate) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","label":"sampledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","description":"Deterministic predictable feedback identifies the observed estimator with the corresponding predictable estimator almost surely at every actual time.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-c54b2eb2bb92","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9600,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:329"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment (fun sample => sampledHalfTsallisObservedEstimatedLossAt arms harms eta t sample) =ᵐ[mu] (fun sample => sampledHalfTsallisPredictableEstimatedLossAt arms harms eta loss t sample)","missing":[],"search":"sampledhalftsallisobservedestimatedlossat_eq_predictable_ae banditrlproof.tsallis.sampledhalftsallisobservedestimatedlossat_eq_predictable_ae deterministic predictable feedback identifies the observed estimator with the corresponding predictable estimator almost surely at every actual time. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisInitialUpdatedAt","label":"measurable_sampledHalfTsallisInitialUpdatedAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisInitialUpdatedAt","description":"theorem measurable_sampledHalfTsallisInitialUpdatedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × Action => sampledHalfTsallisInitialUp…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-07f9aff6d150","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9601,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:376"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisInitialUpdatedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × Action => sampledHalfTsallisInitialUpdatedAt arms harms eta loss sample.1 sample.2 candidate)","missing":[],"search":"measurable_sampledhalftsallisinitialupdatedat banditrlproof.tsallis.measurable_sampledhalftsallisinitialupdatedat theorem measurable_sampledhalftsallisinitialupdatedat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : exp3.predictablelossvector env action) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × action => sampledhalftsallisinitialupdatedat arms harms eta loss sample.1 sample.2 candidate) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisInitialHistoryActionStability","label":"measurable_sampledHalfTsallisInitialHistoryActionStability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisInitialHistoryActionStability","description":"theorem measurable_sampledHalfTsallisInitialHistoryActionStability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : Measurable (sampledHalfTsallisInitialHistoryActionStability arms harms eta loss)","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-a533e2834ae4","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9602,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:411"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisInitialHistoryActionStability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : Measurable (sampledHalfTsallisInitialHistoryActionStability arms harms eta loss)","missing":[],"search":"measurable_sampledhalftsallisinitialhistoryactionstability banditrlproof.tsallis.measurable_sampledhalftsallisinitialhistoryactionstability theorem measurable_sampledhalftsallisinitialhistoryactionstability {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (loss : exp3.predictablelossvector env action) : measurable (sampledhalftsallisinitialhistoryactionstability arms harms eta loss) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialStability_eq_historyAction_ae","label":"sampledHalfTsallisInitialStability_eq_historyAction_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisInitialStability_eq_historyAction_ae","description":"The stored-reward initial stability term agrees almost surely with the canonical predictable updated-minimizer score.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-6153a4eaf0f8","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9603,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:433"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisInitialStability_eq_historyAction_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment sampledHalfTsallisInitialStability arms harms eta =ᵐ[mu] (fun sample => sampledHalfTsallisInitialHistoryActionStability arms harms eta loss (sample.1, (sample.2 0).1))","missing":[],"search":"sampledhalftsallisinitialstability_eq_historyaction_ae banditrlproof.tsallis.sampledhalftsallisinitialstability_eq_historyaction_ae the stored-reward initial stability term agrees almost surely with the canonical predictable updated-minimizer score. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedSuccessorStabilityAt_eq_predictable_ae","label":"sampledHalfTsallisObservedSuccessorStabilityAt_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisObservedSuccessorStabilityAt_eq_predictable_ae","description":"Observed and predictable successor stability agree almost surely.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-6562708cb7fc","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9604,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:482"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisObservedSuccessorStabilityAt_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment sampledHalfTsallisObservedSuccessorStabilityAt arms harms eta n =ᵐ[mu] sampledHalfTsallisSuccessorStabilityAt arms harms eta loss n","missing":[],"search":"sampledhalftsallisobservedsuccessorstabilityat_eq_predictable_ae banditrlproof.tsallis.sampledhalftsallisobservedsuccessorstabilityat_eq_predictable_ae observed and predictable successor stability agree almost surely. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_mixedImportanceWeightedLoss_of_coordinates","label":"measurable_mixedImportanceWeightedLoss_of_coordinates","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_mixedImportanceWeightedLoss_of_coordinates","description":"theorem measurable_mixedImportanceWeightedLoss_of_coordinates {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (hprob : forall candidate, candidate ∈ arms -> Measurable (fun history => prob history candidate)) (hloss : forall candidate, candidate ∈ arms -> Measu…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-b1c890ad493a","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9605,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:509"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_mixedImportanceWeightedLoss_of_coordinates {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (hprob : forall candidate, candidate ∈ arms -> Measurable (fun history => prob history candidate)) (hloss : forall candidate, candidate ∈ arms -> Measurable (fun history => loss history candidate)) : Measurable (fun sample : History × Action => Exp3.mixedImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2)","missing":[],"search":"measurable_mixedimportanceweightedloss_of_coordinates banditrlproof.tsallis.measurable_mixedimportanceweightedloss_of_coordinates theorem measurable_mixedimportanceweightedloss_of_coordinates {history : type u} {action : type v} [measurablespace history] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (prob loss : history -> action -> real) (hprob : forall candidate, candidate ∈ arms -> measurable (fun history => prob history candidate)) (hloss : forall candidate, candidate ∈ arms -> measurable (fun history => loss history candidate)) : measurable (fun sample : history × action => exp3.mixedimportanceweightedloss arms (prob sample.1) (loss sample.1) sample.2) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_weightedImportanceWeightedLoss_of_coordinates","label":"measurable_weightedImportanceWeightedLoss_of_coordinates","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_weightedImportanceWeightedLoss_of_coordinates","description":"theorem measurable_weightedImportanceWeightedLoss_of_coordinates {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob weight loss : History -> Action -> Real) (hprob : forall candidate, candidate ∈ arms -> Measurable (fun history => prob history candidate)) (hweight : forall candidate, candidate ∈ a…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-52911f1274f2","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9606,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:536"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_weightedImportanceWeightedLoss_of_coordinates {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob weight loss : History -> Action -> Real) (hprob : forall candidate, candidate ∈ arms -> Measurable (fun history => prob history candidate)) (hweight : forall candidate, candidate ∈ arms -> Measurable (fun history => weight history candidate)) (hloss : forall candidate, candidate ∈ arms -> Measurable (fun history => loss history candidate)) : Measurable (fun sample : History × Action => Exp3.weightedImportanceWeightedLoss arms (prob sample.1) (weight sample.1) (loss sample.1) sample.2)","missing":[],"search":"measurable_weightedimportanceweightedloss_of_coordinates banditrlproof.tsallis.measurable_weightedimportanceweightedloss_of_coordinates theorem measurable_weightedimportanceweightedloss_of_coordinates {history : type u} {action : type v} [measurablespace history] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (prob weight loss : history -> action -> real) (hprob : forall candidate, candidate ∈ arms -> measurable (fun history => prob history candidate)) (hweight : forall candidate, candidate ∈ arms -> measurable (fun history => weight history candidate)) (hloss : forall candidate, candidate ∈ arms -> measurable (fun history => loss history candidate)) : measurable (fun sample : history × action => exp3.weightedimportanceweightedloss arms (prob sample.1) (weight sample.1) (loss sample.1) sample.2) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_mixedImportanceWeightedLoss_finiteActionKernel","label":"integrable_mixedImportanceWeightedLoss_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_mixedImportanceWeightedLoss_finiteActionKernel","description":"Mixed importance-weighted loss is integrable under its finite sampling kernel without a uniform probability floor.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-21d599b25dc3","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9607,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:567"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_mixedImportanceWeightedLoss_finiteActionKernel {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (hprobPos : forall history candidate, candidate ∈ arms -> 0 < prob history candidate) (hloss : forall history candidate, candidate ∈ arms -> 0 <= loss history candidate ∧ loss history candidate <= 1) (hscore : Measurable (fun sample : History × Action => Exp3.mixedImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2)) : Integrable (fun sample : History × Action => Exp3.mixedImportanceWeightedLoss arms (prob sample.1) (loss sample.1) sample.2) (historyMu ⊗ₘ Exp3.finiteActionKernel arms pr…","missing":[],"search":"integrable_mixedimportanceweightedloss_finiteactionkernel banditrlproof.tsallis.integrable_mixedimportanceweightedloss_finiteactionkernel mixed importance-weighted loss is integrable under its finite sampling kernel without a uniform probability floor. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_weightedImportanceWeightedLoss_finiteActionKernel","label":"integrable_weightedImportanceWeightedLoss_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_weightedImportanceWeightedLoss_finiteActionKernel","description":"A fixed simplex comparator weighting of the importance-weighted loss is integrable under the sampling kernel, again without a uniform floor.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-d79b58284b4a","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9608,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:658"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_weightedImportanceWeightedLoss_finiteActionKernel {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (hprobPos : forall history candidate, candidate ∈ arms -> 0 < prob history candidate) (hloss : forall history candidate, candidate ∈ arms -> 0 <= loss history candidate ∧ loss history candidate <= 1) (hscore : Measurable (fun sample : History × Action => Exp3.weightedImportanceWeightedLoss arms (prob sample.1) q (loss sample.1) sample.2)) : Integrable (fun sample : History × Action => Exp3.weightedImportanceWeightedLoss arms (prob sample.1) q (los…","missing":[],"search":"integrable_weightedimportanceweightedloss_finiteactionkernel banditrlproof.tsallis.integrable_weightedimportanceweightedloss_finiteactionkernel a fixed simplex comparator weighting of the importance-weighted loss is integrable under the sampling kernel, again without a uniform floor. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.initialHalfTsallisEnvironmentDistributionSource","label":"initialHalfTsallisEnvironmentDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.initialHalfTsallisEnvironmentDistributionSource","description":"Constant environment-indexed source for the initial half-Tsallis law.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-c7f7b61f44c1","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9609,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:749"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def initialHalfTsallisEnvironmentDistributionSource {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Exp3.MeasurableFiniteActionDistribution arms (fun _env : Env => initialHalfTsallisDistribution arms harms eta) where","missing":[],"search":"initialhalftsallisenvironmentdistributionsource banditrlproof.tsallis.initialhalftsallisenvironmentdistributionsource constant environment-indexed source for the initial half-tsallis law. definition compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_score_comp_history_action_of_condDistrib","label":"integrable_score_comp_history_action_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_score_comp_history_action_of_condDistrib","description":"Product-law integrability of an arbitrary score transports to the actual history/action composition under an identified conditional action law.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-28f15d9c0727","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9610,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:763"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_score_comp_history_action_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (policy : Kernel History Action) [IsMarkovKernel policy] (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (score : History × Action -> Real) (hIntegrable : Integrable score (mu.map history ⊗ₘ policy)) : Integrable (fun omega => score (history omega, action omega)) mu","missing":[],"search":"integrable_score_comp_history_action_of_conddistrib banditrlproof.tsallis.integrable_score_comp_history_action_of_conddistrib product-law integrability of an arbitrary score transports to the actual history/action composition under an identified conditional action law. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_mixed_weightedImportanceWeightedLoss_of_condDistrib","label":"integrable_mixed_weightedImportanceWeightedLoss_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_mixed_weightedImportanceWeightedLoss_of_condDistrib","description":"theorem integrable_mixed_weightedImportanceWeightedLoss_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega…","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-51a0a5256c99","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9611,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:786"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_mixed_weightedImportanceWeightedLoss_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (hcond : condDistrib action history mu =ᵐ[mu.map history] Exp3.finiteActionKernel arms prob source) (hprobPos : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (hlossMeas : forall candidate, candidate ∈ arms -> Measurable (fun h => loss h candidat…","missing":[],"search":"integrable_mixed_weightedimportanceweightedloss_of_conddistrib banditrlproof.tsallis.integrable_mixed_weightedimportanceweightedloss_of_conddistrib theorem integrable_mixed_weightedimportanceweightedloss_of_conddistrib {omega : type u} {history : type v} {action : type*} [measurablespace omega] [measurablespace history] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (mu : measure omega) [isfinitemeasure mu] (history : omega -> history) (hhistory : measurable history) (action : omega -> action) (haction : measurable action) (arms : finset action) (prob loss : history -> action -> real) (q : action -> real) (hq : ftrl.finitesimplex arms q) (source : exp3.measurablefiniteactiondistribution arms prob) (hcond : conddistrib action history mu =ᵐ[mu.map history] exp3.finiteactionkernel arms prob source) (hprobpos : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (hlossmeas : forall candidate, candidate ∈ arms -> measurable (fun h => loss h candidate)) (hloss : forall h candidate, candidate ∈ arms -> 0 <= loss h candidate ∧ loss h candidate <= 1) : integrable (fun omega => exp3.mixedimportanceweightedloss arms (prob (history omega)) (loss (history omega)) (action omega)) mu ∧ integrable (fun omega => exp3.weightedimportanceweightedloss arms (prob (history omega)) q (loss (history omega)) (actio…","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_mixed_weightedImportanceWeightedLoss_eq_predictable","label":"integral_mixed_weightedImportanceWeightedLoss_eq_predictable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_mixed_weightedImportanceWeightedLoss_eq_predictable","description":"Conditional first moments for a positive finite sampling law, with no uniform probability floor.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-623ea9133e12","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9612,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:833"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_mixed_weightedImportanceWeightedLoss_eq_predictable {Omega : Type u} {History : Type v} {Action : Type*} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (hcond : condDistrib action history mu =ᵐ[mu.map history] Exp3.finiteActionKernel arms prob source) (hprobPos : forall h candidate, candidate ∈ arms -> 0 < prob h candidate) (hlossMeas : forall candidate, candidate ∈ arms -> Measurable (fun h => loss h candidate)…","missing":[],"search":"integral_mixed_weightedimportanceweightedloss_eq_predictable banditrlproof.tsallis.integral_mixed_weightedimportanceweightedloss_eq_predictable conditional first moments for a positive finite sampling law, with no uniform probability floor. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedLossAt_first_moments","label":"sampledHalfTsallisPredictableEstimatedLossAt_first_moments","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedLossAt_first_moments","description":"At every actual time, both the sampling-distribution mixed estimator and the fixed-comparator weighted estimator have the corresponding predictable first moment under the generated half-Tsallis trajectory.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-e9e715e5b8f5","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9613,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:937"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisPredictableEstimatedLossAt_first_moments {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (t : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment (Integrable (fun sample => FTRL.linearLoss arms (sampledHalfTsallisProbabilityAtTime arms harms eta t sample) (sampledHalfTsallisPredictableEstimatedLossAt arms harms eta loss t sample)) mu ∧ Integrable (fun sample => FTRL.lin…","missing":[],"search":"sampledhalftsallispredictableestimatedlossat_first_moments banditrlproof.tsallis.sampledhalftsallispredictableestimatedlossat_first_moments at every actual time, both the sampling-distribution mixed estimator and the fixed-comparator weighted estimator have the corresponding predictable first moment under the generated half-tsallis trajectory. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledHalfTsallisProbabilityAtTime","label":"finiteSimplex_sampledHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_sampledHalfTsallisProbabilityAtTime","description":"theorem finiteSimplex_sampledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : FTRL.finiteSimplex arms (sampledHalfTsallisProbabilityAtTime arms harms eta t sample)","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-eb00ababff86","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9614,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1105"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_sampledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : FTRL.finiteSimplex arms (sampledHalfTsallisProbabilityAtTime arms harms eta t sample)","missing":[],"search":"finitesimplex_sampledhalftsallisprobabilityattime banditrlproof.tsallis.finitesimplex_sampledhalftsallisprobabilityattime theorem finitesimplex_sampledhalftsallisprobabilityattime {env : type u} {action : type v} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (t : nat) (sample : env × ((k : nat) -> action × real)) : ftrl.finitesimplex arms (sampledhalftsallisprobabilityattime arms harms eta t sample) theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisPredictableLinearLossAt","label":"integrable_sampledHalfTsallisPredictableLinearLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledHalfTsallisPredictableLinearLossAt","description":"Current mixed predictable loss and fixed-comparator predictable loss are integrable at every actual time.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-55e08ef64be2","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9615,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1125"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledHalfTsallisPredictableLinearLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (t : Nat) : Integrable (fun sample => FTRL.linearLoss arms (sampledHalfTsallisProbabilityAtTime arms harms eta t sample) (Exp3.predictableLossAt loss t sample)) mu ∧ Integrable (fun sample => FTRL.linearLoss arms q (Exp3.predictableLossAt loss t sample)) mu","missing":[],"search":"integrable_sampledhalftsallispredictablelinearlossat banditrlproof.tsallis.integrable_sampledhalftsallispredictablelinearlossat current mixed predictable loss and fixed-comparator predictable loss are integrable at every actual time. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","label":"integral_sampledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","description":"Finite-horizon predictable estimated regret is integrable and has exactly the same integral as predictable environment regret.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-9de71452a4a6","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9616,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1211"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (horizon : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment Integrable (sampledHalfTsallisPredictableEstimatedRegret arms harms eta loss q horizon) mu ∧ Integrable (sampledHalfTsallisPredictableEnvironmentRegret arms harms eta loss q horizon) mu ∧ integral mu (sam…","missing":[],"search":"integral_sampledhalftsallispredictableestimatedregret_eq_environmentregret banditrlproof.tsallis.integral_sampledhalftsallispredictableestimatedregret_eq_environmentregret finite-horizon predictable estimated regret is integrable and has exactly the same integral as predictable environment regret. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedRegret_eq_predictable_ae","label":"sampledHalfTsallisObservedEstimatedRegret_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedRegret_eq_predictable_ae","description":"The observed estimated-regret functional agrees almost surely with its predictable-estimator version over every finite horizon.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-d186031fdf89","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9617,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1303"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisObservedEstimatedRegret_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment sampledHalfTsallisObservedEstimatedRegret arms harms eta q horizon =ᵐ[mu] sampledHalfTsallisPredictableEstimatedRegret arms harms eta loss q horizon","missing":[],"search":"sampledhalftsallisobservedestimatedregret_eq_predictable_ae banditrlproof.tsallis.sampledhalftsallisobservedestimatedregret_eq_predictable_ae the observed estimated-regret functional agrees almost surely with its predictable-estimator version over every finite horizon. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisObservedEstimatedRegret_eq_environmentRegret","label":"integral_sampledHalfTsallisObservedEstimatedRegret_eq_environmentRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledHalfTsallisObservedEstimatedRegret_eq_environmentRegret","description":"Observed importance-weighted regret is integrable and has exactly the predictable environment-regret integral.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-7a8c2a4c5c60","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9618,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1343"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledHalfTsallisObservedEstimatedRegret_eq_environmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (horizon : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment Integrable (sampledHalfTsallisObservedEstimatedRegret arms harms eta q horizon) mu ∧ Integrable (sampledHalfTsallisPredictableEnvironmentRegret arms harms eta loss q horizon) mu ∧ integral mu (sampledHalfTsa…","missing":[],"search":"integral_sampledhalftsallisobservedestimatedregret_eq_environmentregret banditrlproof.tsallis.integral_sampledhalftsallisobservedestimatedregret_eq_environmentregret observed importance-weighted regret is integrable and has exactly the predictable environment-regret integral. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisInitialStability_le_halfPower","label":"integral_sampledHalfTsallisInitialStability_le_halfPower","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledHalfTsallisInitialStability_le_halfPower","description":"The time-zero observed stability has the same half-power bound as every successor round. A probability prior keeps the constant bound unscaled.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-63760964bffd","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9619,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1387"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledHalfTsallisInitialStability_le_halfPower {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment Integrable (sampledHalfTsallisInitialStability arms harms eta) mu ∧ integral mu (sampledHalfTsallisInitialStability arms harms eta) <= 2 * eta * powerSum arms (1 / 2 : Real) (initialHalfTsallisDistribution arms harms eta)","missing":[],"search":"integral_sampledhalftsallisinitialstability_le_halfpower banditrlproof.tsallis.integral_sampledhalftsallisinitialstability_le_halfpower the time-zero observed stability has the same half-power bound as every successor round. a probability prior keeps the constant bound unscaled. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisSuccessorStabilitiesAt_canonical","label":"integrable_sampledHalfTsallisSuccessorStabilitiesAt_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledHalfTsallisSuccessorStabilitiesAt_canonical","description":"Predictable and observed successor stability are both integrable under the canonical generated half-Tsallis trajectory.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-9f75a775f4d6","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9620,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1529"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledHalfTsallisSuccessorStabilitiesAt_canonical {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment Integrable (sampledHalfTsallisSuccessorStabilityAt arms harms eta loss n) mu ∧ Integrable (sampledHalfTsallisObservedSuccessorStabilityAt arms harms eta n) mu","missing":[],"search":"integrable_sampledhalftsallissuccessorstabilitiesat_canonical banditrlproof.tsallis.integrable_sampledhalftsallissuccessorstabilitiesat_canonical predictable and observed successor stability are both integrable under the canonical generated half-tsallis trajectory. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_le","label":"integral_sampledHalfTsallisPredictableEnvironmentRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_le","description":"Complete estimated-to-environment regret theorem for the generated pure half-Tsallis policy. The horizon contains time zero plus `horizon` successor rounds; self-bounding and learning-rate tuning remain downstream.","url":"../modules/banditrlproof-tsallisftrlestimatedenvironmentregret/index.html#decl-d6f1e228a2a1","parent":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","order":9621,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret"],["Source","BanditRLProof/TsallisFTRLEstimatedEnvironmentRegret.lean:1625"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledHalfTsallisPredictableEnvironmentRegret_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (horizon : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment integral mu (sampledHalfTsallisPredictableEnvironmentRegret arms harms eta loss q horizon) <= 2 * eta * powerSum arms (1 / 2 : Real) (initialHalfTsallisDistribution arms harms eta) + integral mu (fu…","missing":[],"search":"integral_sampledhalftsallispredictableenvironmentregret_le banditrlproof.tsallis.integral_sampledhalftsallispredictableenvironmentregret_le complete estimated-to-environment regret theorem for the generated pure half-tsallis policy. the horizon contains time zero plus `horizon` successor rounds; self-bounding and learning-rate tuning remain downstream. theorem compiled","shard":"modules/4184d01947728c69.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedStabilityScore_comp_history_action_of_condDistrib","label":"integrable_importanceWeightedStabilityScore_comp_history_action_of_condDistrib","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_importanceWeightedStabilityScore_comp_history_action_of_condDistrib","description":"Product-law integrability transports back to the realized history/action pair when the kernel is the conditional action law.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html#decl-28a800717119","parent":"module:BanditRLProof.TsallisFTRLExpectedStability","order":9622,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLExpectedStability"],["Source","BanditRLProof/TsallisFTRLExpectedStability.lean:29"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_importanceWeightedStabilityScore_comp_history_action_of_condDistrib {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [StandardBorelSpace Action] [Nonempty Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (policy : Kernel History Action) [IsMarkovKernel policy] (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (hIntegrable : Integrable (importanceWeightedStabilityScore arms prob loss next) (mu.map history ⊗ₘ policy)) : Integrable (fun omega => importanceWeightedStabilityScore arms prob loss next (history omega, action omega)) mu","missing":[],"search":"integrable_importanceweightedstabilityscore_comp_history_action_of_conddistrib banditrlproof.tsallis.integrable_importanceweightedstabilityscore_comp_history_action_of_conddistrib product-law integrability transports back to the realized history/action pair when the kernel is the conditional action law. theorem compiled","shard":"modules/80ea1cc48bfe0748.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_halfPowerStabilityBound_comp_history","label":"integrable_halfPowerStabilityBound_comp_history","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_halfPowerStabilityBound_comp_history","description":"Integrability under a history marginal transports back along the history map.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html#decl-3ae05b75d28c","parent":"module:BanditRLProof.TsallisFTRLExpectedStability","order":9623,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLExpectedStability"],["Source","BanditRLProof/TsallisFTRLExpectedStability.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_halfPowerStabilityBound_comp_history {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] (mu : Measure Omega) (history : Omega -> History) (hhistory : Measurable history) (arms : Finset Action) (eta : Real) (prob : History -> Action -> Real) (hIntegrable : Integrable (halfPowerStabilityBound arms eta prob) (mu.map history)) : Integrable (fun omega => halfPowerStabilityBound arms eta prob (history omega)) mu","missing":[],"search":"integrable_halfpowerstabilitybound_comp_history banditrlproof.tsallis.integrable_halfpowerstabilitybound_comp_history integrability under a history marginal transports back along the history map. theorem compiled","shard":"modules/80ea1cc48bfe0748.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_importanceWeightedStabilityScore_le_integral_sum_halfPowerStabilityBound_of_condDistrib_of_minimizers","label":"integral_sum_importanceWeightedStabilityScore_le_integral_sum_halfPowerStabilityBound_of_condDistrib_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_importanceWeightedStabilityScore_le_integral_sum_halfPowerStabilityBound_of_condDistrib_of_minimizers","description":"Expected finite-horizon half-Tsallis stability under identified conditional action laws and explicit current/update minimizer certificates. All rounds live on one ambient measure `mu`. This is the theorem-level bridge from the one-round sampling-law average to the expected finite stability sum.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html#decl-9ad3a21cc69d","parent":"module:BanditRLProof.TsallisFTRLExpectedStability","order":9624,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLExpectedStability"],["Source","BanditRLProof/TsallisFTRLExpectedStability.lean:80"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_importanceWeightedStabilityScore_le_integral_sum_halfPowerStabilityBound_of_condDistrib_of_minimizers {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (horizon : Nat) (arms : Finset Action) (eta : Real) (history : Nat -> Omega -> History) (action : Nat -> Omega -> Action) (score prob loss : Nat -> History -> Action -> Real) (next : Nat -> History -> Action -> Action -> Real) (policy : Nat -> Kernel History Action) (hmarkov : forall t, IsMarkovKernel (policy t)) (hhistory : forall t, Measurable (history t)) (haction : forall t, Measurable (action t)) (hpolicy : forall t, policy t =ᵐ[mu.map (history t)] fun h => Exp3.finiteActionMeasure arms (prob t…","missing":[],"search":"integral_sum_importanceweightedstabilityscore_le_integral_sum_halfpowerstabilitybound_of_conddistrib_of_minimizers banditrlproof.tsallis.integral_sum_importanceweightedstabilityscore_le_integral_sum_halfpowerstabilitybound_of_conddistrib_of_minimizers expected finite-horizon half-tsallis stability under identified conditional action laws and explicit current/update minimizer certificates. all rounds live on one ambient measure `mu`. this is the theorem-level bridge from the one-round sampling-law average to the expected finite stability sum. theorem compiled","shard":"modules/80ea1cc48bfe0748.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_halfTsallisHistoryStability_le_integral_sum_halfPowerStabilityBound","label":"integral_sum_halfTsallisHistoryStability_le_integral_sum_halfPowerStabilityBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_halfTsallisHistoryStability_le_integral_sum_halfPowerStabilityBound","description":"Canonical half-Tsallis finite-horizon expected stability theorem. The current and updated minimizer certificates are internal; selector and score regularity remain explicit.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html#decl-dbadbd9ecbc1","parent":"module:BanditRLProof.TsallisFTRLExpectedStability","order":9625,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLExpectedStability"],["Source","BanditRLProof/TsallisFTRLExpectedStability.lean:190"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_halfTsallisHistoryStability_le_integral_sum_halfPowerStabilityBound {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (history : Nat -> Omega -> History) (action : Nat -> Omega -> Action) (score loss : Nat -> History -> Action -> Real) (policy : Nat -> Kernel History Action) (hmarkov : forall t, IsMarkovKernel (policy t)) (hhistory : forall t, Measurable (history t)) (haction : forall t, Measurable (action t)) (hpolicy : forall t, policy t =ᵐ[mu.map (history t)] fun h => Exp3.finiteActionMeasure arms (halfTsallisHistoryMinimizer arms harms eta (score t) h)) (hcond : forall…","missing":[],"search":"integral_sum_halftsallishistorystability_le_integral_sum_halfpowerstabilitybound banditrlproof.tsallis.integral_sum_halftsallishistorystability_le_integral_sum_halfpowerstabilitybound canonical half-tsallis finite-horizon expected stability theorem. the current and updated minimizer certificates are internal; selector and score regularity remain explicit. theorem compiled","shard":"modules/80ea1cc48bfe0748.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisSuccessorStabilityScore","label":"halfTsallisSuccessorStabilityScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisSuccessorStabilityScore","description":"The realized FTRL stability term using the next round's current half-Tsallis selector.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html#decl-3eebb6100026","parent":"module:BanditRLProof.TsallisFTRLExpectedStability","order":9626,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLExpectedStability"],["Source","BanditRLProof/TsallisFTRLExpectedStability.lean:254"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisSuccessorStabilityScore {Omega : Type u} {History : Type v} {Action : Type w} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (history : Nat -> Omega -> History) (action : Nat -> Omega -> Action) (score loss : Nat -> History -> Action -> Real) (t : Nat) (omega : Omega) : Real","missing":[],"search":"halftsallissuccessorstabilityscore banditrlproof.tsallis.halftsallissuccessorstabilityscore the realized ftrl stability term using the next round's current half-tsallis selector. definition compiled","shard":"modules/80ea1cc48bfe0748.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_halfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_score_succ","label":"integral_sum_halfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_score_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_halfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_score_succ","description":"Expected finite-horizon bound for the actual successor stability sum. The score recursion identifies the next round's current selector with the importance-weighted updated selector from the current round. This closes the action-dependent successor alignment without claiming it pathwise absent the explicit recursion contract.","url":"../modules/banditrlproof-tsallisftrlexpectedstability/index.html#decl-17d93609f98a","parent":"module:BanditRLProof.TsallisFTRLExpectedStability","order":9627,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLExpectedStability"],["Source","BanditRLProof/TsallisFTRLExpectedStability.lean:279"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_halfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_score_succ {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (history : Nat -> Omega -> History) (action : Nat -> Omega -> Action) (score loss : Nat -> History -> Action -> Real) (policy : Nat -> Kernel History Action) (hmarkov : forall t, IsMarkovKernel (policy t)) (hhistory : forall t, Measurable (history t)) (haction : forall t, Measurable (action t)) (hpolicy : forall t, policy t =ᵐ[mu.map (history t)] fun h => Exp3.finiteActionMeasure arms (halfTsallisHistoryMinimizer arms harms eta (score t) h))…","missing":[],"search":"integral_sum_halftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_of_score_succ banditrlproof.tsallis.integral_sum_halftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_of_score_succ expected finite-horizon bound for the actual successor stability sum. the score recursion identifies the next round's current selector with the importance-weighted updated selector from the current round. this closes the action-dependent successor alignment without claiming it pathwise absent the explicit recursion contract. theorem compiled","shard":"modules/80ea1cc48bfe0748.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer","label":"halfTsallisCumulativeMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer","description":"The fixed half-Tsallis minimizer for losses accumulated before round `t`.","url":"../modules/banditrlproof-tsallisftrlfinitehorizonselection/index.html#decl-7605b163ccc6","parent":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","order":9628,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLFiniteHorizonSelection"],["Source","BanditRLProof/TsallisFTRLFiniteHorizonSelection.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisCumulativeMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : Action -> Real","missing":[],"search":"halftsalliscumulativeminimizer banditrlproof.tsallis.halftsalliscumulativeminimizer the fixed half-tsallis minimizer for losses accumulated before round `t`. definition compiled","shard":"modules/2a195d5c65ce1238.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_isRegularizedMinimizer","label":"halfTsallisCumulativeMinimizer_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_isRegularizedMinimizer","description":"Every canonical cumulative selector carries the minimizer certificate required by the finite-horizon FTRL decomposition.","url":"../modules/banditrlproof-tsallisftrlfinitehorizonselection/index.html#decl-6f121754e252","parent":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","order":9629,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLFiniteHorizonSelection"],["Source","BanditRLProof/TsallisFTRLFiniteHorizonSelection.lean:30"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisCumulativeMinimizer_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (FTRL.cumulativeLoss loss t) (halfTsallisCumulativeMinimizer arms harms eta loss t)","missing":[],"search":"halftsalliscumulativeminimizer_isregularizedminimizer banditrlproof.tsallis.halftsalliscumulativeminimizer_isregularizedminimizer every canonical cumulative selector carries the minimizer certificate required by the finite-horizon ftrl decomposition. theorem compiled","shard":"modules/2a195d5c65ce1238.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_succ","label":"halfTsallisCumulativeMinimizer_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_succ","description":"The selector at `t + 1` is the fixed minimizer after appending round `t`'s loss vector.","url":"../modules/banditrlproof-tsallisftrlfinitehorizonselection/index.html#decl-61a074c935f8","parent":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","order":9630,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLFiniteHorizonSelection"],["Source","BanditRLProof/TsallisFTRLFiniteHorizonSelection.lean:43"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisCumulativeMinimizer_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (t : Nat) : halfTsallisCumulativeMinimizer arms harms eta loss (t + 1) = halfTsallisMinimizer arms harms eta (fun action => FTRL.cumulativeLoss loss t action + loss t action)","missing":[],"search":"halftsalliscumulativeminimizer_succ banditrlproof.tsallis.halftsalliscumulativeminimizer_succ the selector at `t + 1` is the fixed minimizer after appending round `t`'s loss vector. theorem compiled","shard":"modules/2a195d5c65ce1238.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_succ_eq_updated","label":"halfTsallisCumulativeMinimizer_succ_eq_updated","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_succ_eq_updated","description":"If the realized round loss is the importance-weighted estimator generated from the current selector, the successor selector is exactly the canonical one-step updated minimizer.","url":"../modules/banditrlproof-tsallisftrlfinitehorizonselection/index.html#decl-b8c3c59135f6","parent":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","order":9631,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLFiniteHorizonSelection"],["Source","BanditRLProof/TsallisFTRLFiniteHorizonSelection.lean:55"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisCumulativeMinimizer_succ_eq_updated {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (rawLoss : Action -> Real) (chosen : Action) (t : Nat) (hloss : loss t = Exp3.importanceWeightedLoss (halfTsallisCumulativeMinimizer arms harms eta loss t) rawLoss chosen) : halfTsallisCumulativeMinimizer arms harms eta loss (t + 1) = halfTsallisUpdatedMinimizer arms harms eta (FTRL.cumulativeLoss loss t) rawLoss chosen","missing":[],"search":"halftsalliscumulativeminimizer_succ_eq_updated banditrlproof.tsallis.halftsalliscumulativeminimizer_succ_eq_updated if the realized round loss is the importance-weighted estimator generated from the current selector, the successor selector is exactly the canonical one-step updated minimizer. theorem compiled","shard":"modules/2a195d5c65ce1238.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty_half_canonical","label":"cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty_half_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty_half_canonical","description":"Finite-horizon half-Tsallis FTRL decomposition with the cumulative minimizer sequence selected internally. The remaining first term on the right is the pathwise stability sum.","url":"../modules/banditrlproof-tsallisftrlfinitehorizonselection/index.html#decl-be4d5d42cab5","parent":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","order":9632,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLFiniteHorizonSelection"],["Source","BanditRLProof/TsallisFTRLFiniteHorizonSelection.lean:73"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty_half_canonical {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Nat -> Action -> Real) (q : Action -> Real) (T : Nat) (heta : 0 < eta) (hq : FTRL.finiteSimplex arms q) : let p := halfTsallisCumulativeMinimizer arms harms eta loss (Finset.range T).sum (fun t => FTRL.linearLoss arms (p t) (loss t) - FTRL.linearLoss arms q (loss t)) <= (Finset.range T).sum (fun t => FTRL.linearLoss arms (p t) (loss t) - FTRL.linearLoss arms (p (t + 1)) (loss t)) + ((powerSum arms (1 / 2 : Real) (p 0) - powerSum arms (1 / 2 : Real) q) / (1 - (1 / 2 : Real))) / eta","missing":[],"search":"cumulativelinearloss_sub_comparator_le_stability_add_powersumpenalty_half_canonical banditrlproof.tsallis.cumulativelinearloss_sub_comparator_le_stability_add_powersumpenalty_half_canonical finite-horizon half-tsallis ftrl decomposition with the cumulative minimizer sequence selected internally. the remaining first term on the right is the pathwise stability sum. theorem compiled","shard":"modules/2a195d5c65ce1238.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HalfTsallisGeneratedSelectorMeasurability","label":"HalfTsallisGeneratedSelectorMeasurability","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.HalfTsallisGeneratedSelectorMeasurability","description":"Coordinate regularity for the generated current and sampled-action updated half-Tsallis selectors. The finite-history component constructs the policy; the second component covers the environment/prefix/action parameter space of the updated canonical minimizer.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-ecfa855d9a6a","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9633,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure HalfTsallisGeneratedSelectorMeasurability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : Prop where","missing":[],"search":"halftsallisgeneratedselectormeasurability banditrlproof.tsallis.halftsallisgeneratedselectormeasurability coordinate regularity for the generated current and sampled-action updated half-tsallis selectors. the finite-history component constructs the policy; the second component covers the environment/prefix/action parameter space of the updated canonical minimizer. structure compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisFiniteHistorySelectorMeasurability","label":"canonicalHalfTsallisFiniteHistorySelectorMeasurability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.canonicalHalfTsallisFiniteHistorySelectorMeasurability","description":"The canonical half-Tsallis selector itself satisfies the finite-history coordinate measurability contract.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-8956122273f7","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9634,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHalfTsallisFiniteHistorySelectorMeasurability {Action : Type v} [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta where","missing":[],"search":"canonicalhalftsallisfinitehistoryselectormeasurability banditrlproof.tsallis.canonicalhalftsallisfinitehistoryselectormeasurability the canonical half-tsallis selector itself satisfies the finite-history coordinate measurability contract. definition compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_importanceWeightedStabilityScore","label":"measurable_importanceWeightedStabilityScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_importanceWeightedStabilityScore","description":"A finite sum of current-minus-updated importance-weighted linear losses is measurable once all current, loss, and updated coordinates are measurable.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-52075fe1063a","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9635,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:52"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_importanceWeightedStabilityScore {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (hprob : forall candidate, candidate ∈ arms -> Measurable (fun history => prob history candidate)) (hloss : forall candidate, candidate ∈ arms -> Measurable (fun history => loss history candidate)) (hnext : forall candidate, candidate ∈ arms -> Measurable (fun sample : History × Action => next sample.1 sample.2 candidate)) : Measurable (importanceWeightedStabilityScore arms prob loss next)","missing":[],"search":"measurable_importanceweightedstabilityscore banditrlproof.tsallis.measurable_importanceweightedstabilityscore a finite sum of current-minus-updated importance-weighted linear losses is measurable once all current, loss, and updated coordinates are measurable. theorem compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisPredictableLossAt","label":"measurable_sampledHalfTsallisPredictableLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisPredictableLossAt","description":"Every supported coordinate of the generated predictable loss vector is measurable on the environment/prefix parameter space.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-344456cd8517","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9636,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:96"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (candidate : Action) : Measurable (fun input : Env × History.FinitePairHistory Action Real n => sampledHalfTsallisPredictableLossAt loss n input candidate)","missing":[],"search":"measurable_sampledhalftsallispredictablelossat banditrlproof.tsallis.measurable_sampledhalftsallispredictablelossat every supported coordinate of the generated predictable loss vector is measurable on the environment/prefix parameter space. theorem compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisUpdatedAt_canonical","label":"measurable_sampledHalfTsallisUpdatedAt_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisUpdatedAt_canonical","description":"Every supported coordinate of the canonical sampled-action update is measurable without an external selector assumption.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-341f1eab8b27","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9637,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:110"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisUpdatedAt_canonical {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : (Env × History.FinitePairHistory Action Real n) × Action => sampledHalfTsallisUpdatedAt arms harms eta loss n sample.1 sample.2 candidate)","missing":[],"search":"measurable_sampledhalftsallisupdatedat_canonical banditrlproof.tsallis.measurable_sampledhalftsallisupdatedat_canonical every supported coordinate of the canonical sampled-action update is measurable without an external selector assumption. theorem compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisGeneratedSelectorMeasurability","label":"canonicalHalfTsallisGeneratedSelectorMeasurability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.canonicalHalfTsallisGeneratedSelectorMeasurability","description":"The canonical current and one-step updated minimizers satisfy the full generated-selector regularity contract; callers no longer need to assume measurability of either `Classical.choose` surface.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-dee7ffd41bef","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9638,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:186"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHalfTsallisGeneratedSelectorMeasurability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) : HalfTsallisGeneratedSelectorMeasurability arms harms eta loss where","missing":[],"search":"canonicalhalftsallisgeneratedselectormeasurability banditrlproof.tsallis.canonicalhalftsallisgeneratedselectormeasurability the canonical current and one-step updated minimizers satisfy the full generated-selector regularity contract; callers no longer need to assume measurability of either `classical.choose` surface. definition compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryActionStabilityAt","label":"measurable_sampledHalfTsallisHistoryActionStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryActionStabilityAt","description":"The generated one-round current-minus-updated stability score is measurable under the generated selector coordinate contract.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-eaedf808024b","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9639,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:200"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisHistoryActionStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (selector : HalfTsallisGeneratedSelectorMeasurability arms harms eta loss) (n : Nat) : Measurable (sampledHalfTsallisHistoryActionStabilityAt arms harms eta loss n)","missing":[],"search":"measurable_sampledhalftsallishistoryactionstabilityat banditrlproof.tsallis.measurable_sampledhalftsallishistoryactionstabilityat the generated one-round current-minus-updated stability score is measurable under the generated selector coordinate contract. theorem compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_selector","label":"integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_selector","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_selector","description":"Generated finite-horizon actual-successor stability with all score and integrability regularity derived from one generated-selector contract.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-d4056a423081","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9640,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:223"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_selector {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (loss : Exp3.PredictableLossVector Env Action) (selector : HalfTsallisGeneratedSelectorMeasurability arms harms eta loss) : let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun n => sampledHalfTsallisSuccessorStabilityAt arms harms eta loss n sample)) <= integral mu (fun sample => (Finset.range horizon).sum (fun n => sampledHal…","missing":[],"search":"integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_of_selector banditrlproof.tsallis.integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_of_selector generated finite-horizon actual-successor stability with all score and integrability regularity derived from one generated-selector contract. theorem compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_canonical","label":"integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_canonical","description":"Generated finite-horizon actual-successor stability for the canonical half-Tsallis policy, with selector measurability proved internally.","url":"../modules/banditrlproof-tsallisftrlgeneratedmeasurability/index.html#decl-30ecdb1a4a2b","parent":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","order":9641,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedMeasurability"],["Source","BanditRLProof/TsallisFTRLGeneratedMeasurability.lean:250"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_canonical {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun n => sampledHalfTsallisSuccessorStabilityAt arms harms eta loss n sample)) <= integral mu (fun sample => (Finset.range horizon).sum (fun n =>…","missing":[],"search":"integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_canonical banditrlproof.tsallis.integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_canonical generated finite-horizon actual-successor stability for the canonical half-tsallis policy, with selector measurability proved internally. theorem compiled","shard":"modules/f8c727f70ad58870.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_apply_le_one","label":"finiteSimplex_apply_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_apply_le_one","description":"Every supported coordinate of a finite simplex point is at most one.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-0876032190ed","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9642,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_apply_le_one {Action : Type u} {arms : Finset Action} {p : Action -> Real} (hp : FTRL.finiteSimplex arms p) {action : Action} (haction : action ∈ arms) : p action <= 1","missing":[],"search":"finitesimplex_apply_le_one banditrlproof.tsallis.finitesimplex_apply_le_one every supported coordinate of a finite simplex point is at most one. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.prob_mul_abs_importanceWeightedStabilityScore_le","label":"prob_mul_abs_importanceWeightedStabilityScore_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.prob_mul_abs_importanceWeightedStabilityScore_le","description":"Sampling mass cancels the inverse probability in one realized stability score, leaving a bound by the current and updated selected coordinates.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-057995ee3b83","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9643,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem prob_mul_abs_importanceWeightedStabilityScore_le {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (history : History) (chosen : Action) (hchosen : chosen ∈ arms) (hnextSimplex : FTRL.finiteSimplex arms (next history chosen)) (hprobPos : 0 < prob history chosen) (hloss : 0 <= loss history chosen ∧ loss history chosen <= 1) : prob history chosen * |importanceWeightedStabilityScore arms prob loss next (history, chosen)| <= prob history chosen + next history chosen chosen","missing":[],"search":"prob_mul_abs_importanceweightedstabilityscore_le banditrlproof.tsallis.prob_mul_abs_importanceweightedstabilityscore_le sampling mass cancels the inverse probability in one realized stability score, leaving a bound by the current and updated selected coordinates. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_finiteActionMeasure","label":"integrable_finiteActionMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_finiteActionMeasure","description":"Every real-valued function is integrable under a finite action law.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-0f72c5dcd9f7","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9644,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:99"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_finiteActionMeasure {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] (arms : Finset Action) (prob : Action -> Real) (f : Action -> Real) : Integrable f (Exp3.finiteActionMeasure arms prob)","missing":[],"search":"integrable_finiteactionmeasure banditrlproof.tsallis.integrable_finiteactionmeasure every real-valued function is integrable under a finite action law. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedStabilityScore_finiteActionKernel","label":"integrable_importanceWeightedStabilityScore_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_importanceWeightedStabilityScore_finiteActionKernel","description":"A measurable one-round stability score is automatically integrable under its finite sampling kernel. No uniform lower probability floor is needed.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-c1af29d198aa","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9645,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:110"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_importanceWeightedStabilityScore_finiteActionKernel {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (hprobPos : forall history action, action ∈ arms -> 0 < prob history action) (hnextSimplex : forall history chosen, chosen ∈ arms -> FTRL.finiteSimplex arms (next history chosen)) (hloss : forall history action, action ∈ arms -> 0 <= loss history action ∧ loss history action <= 1) (hscore : Measurable (importanceWeightedStabilityScore arms prob loss next)) : Integrable (importanceWeightedStabilityScore arms prob loss next) (historyMu ⊗ₘ Exp3.finiteAction…","missing":[],"search":"integrable_importanceweightedstabilityscore_finiteactionkernel banditrlproof.tsallis.integrable_importanceweightedstabilityscore_finiteactionkernel a measurable one-round stability score is automatically integrable under its finite sampling kernel. no uniform lower probability floor is needed. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_halfPowerStabilityBound_of_finiteSimplex","label":"integrable_halfPowerStabilityBound_of_finiteSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_halfPowerStabilityBound_of_finiteSimplex","description":"The half-power stability budget is uniformly bounded and hence integrable under every finite history measure.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-5e675dd7a637","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9646,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:195"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_halfPowerStabilityBound_of_finiteSimplex {History : Type u} {Action : Type v} [MeasurableSpace History] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (eta : Real) (prob : History -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) : Integrable (halfPowerStabilityBound arms eta prob) historyMu","missing":[],"search":"integrable_halfpowerstabilitybound_of_finitesimplex banditrlproof.tsallis.integrable_halfpowerstabilitybound_of_finitesimplex the half-power stability budget is uniformly bounded and hence integrable under every finite history measure. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPolicyAt_eq_finiteActionKernel","label":"sampledHalfTsallisPolicyAt_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPolicyAt_eq_finiteActionKernel","description":"The environment-lifted generated policy is exactly the finite-action kernel carried by its measurable distribution source.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-723d0fd44f56","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9647,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:238"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisPolicyAt_eq_finiteActionKernel {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : sampledHalfTsallisPolicyAt (Env := Env) arms harms eta selector n = Exp3.finiteActionKernel arms (sampledHalfTsallisProbabilityAt (Env := Env) arms harms eta n) (sampledHalfTsallisEnvironmentHistoryDistributionSource (Env := Env) arms harms eta selector n)","missing":[],"search":"sampledhalftsallispolicyat_eq_finiteactionkernel banditrlproof.tsallis.sampledhalftsallispolicyat_eq_finiteactionkernel the environment-lifted generated policy is exactly the finite-action kernel carried by its measurable distribution source. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisHistoryActionStabilityAt","label":"integrable_sampledHalfTsallisHistoryActionStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledHalfTsallisHistoryActionStabilityAt","description":"Measurability of the generated stability score now suffices for its integrability under the generated history/action product law.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-17accea74c9d","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9648,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:259"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledHalfTsallisHistoryActionStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (hscore : Measurable (sampledHalfTsallisHistoryActionStabilityAt arms harms eta loss n)) : Integrable (sampledHalfTsallisHistoryActionStabilityAt arms harms eta loss n) (((prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment).map (sampledHalfTsallisHistoryAt n)) ⊗ₘ sampledHalfTsallisPolicyAt (Env := Env) arms harms eta selector n)","missing":[],"search":"integrable_sampledhalftsallishistoryactionstabilityat banditrlproof.tsallis.integrable_sampledhalftsallishistoryactionstabilityat measurability of the generated stability score now suffices for its integrability under the generated history/action product law. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisHalfPowerBoundAt","label":"integrable_sampledHalfTsallisHalfPowerBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledHalfTsallisHalfPowerBoundAt","description":"The generated half-power budget is automatically integrable under every generated history marginal.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-25b579f5cb1a","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9649,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:308"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledHalfTsallisHalfPowerBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Integrable (sampledHalfTsallisHalfPowerBoundAt (Env := Env) arms harms eta n) ((prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment).map (sampledHalfTsallisHistoryAt n))","missing":[],"search":"integrable_sampledhalftsallishalfpowerboundat banditrlproof.tsallis.integrable_sampledhalftsallishalfpowerboundat the generated half-power budget is automatically integrable under every generated history marginal. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_measurable","label":"integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_measurable","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_measurable","description":"Generated finite-horizon actual-successor stability with both integrability contracts discharged. Stability-score measurability remains the exact selector/update regularity boundary.","url":"../modules/banditrlproof-tsallisftrlgeneratedregularity/index.html#decl-d911d966e961","parent":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","order":9650,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLGeneratedRegularity"],["Source","BanditRLProof/TsallisFTRLGeneratedRegularity.lean:335"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_measurable {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (loss : Exp3.PredictableLossVector Env Action) (hscore : forall n, Measurable (sampledHalfTsallisHistoryActionStabilityAt arms harms eta loss n)) : let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun n => sampledHalfTsallisSuccessorStabilityAt arms harms eta loss n…","missing":[],"search":"integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_of_measurable banditrlproof.tsallis.integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound_of_measurable generated finite-horizon actual-successor stability with both integrability contracts discharged. stability-score measurability remains the exact selector/update regularity boundary. theorem compiled","shard":"modules/fd9ae68b9a1257ea.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_simplexPairShift_of_eq_zero_of_le","label":"finiteSimplex_simplexPairShift_of_eq_zero_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_simplexPairShift_of_eq_zero_of_le","description":"A one-sided transfer from a positive donor to a zero coordinate stays in the finite simplex.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-5b5b6bb60519","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9651,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_simplexPairShift_of_eq_zero_of_le {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (t : Real) (hij : i ≠ j) (hp : FTRL.finiteSimplex arms p) (hi : i ∈ arms) (hj : j ∈ arms) (hpi : p i = 0) (ht : 0 <= t) (htj : t <= p j) : FTRL.finiteSimplex arms (simplexPairShift p i j t)","missing":[],"search":"finitesimplex_simplexpairshift_of_eq_zero_of_le banditrlproof.tsallis.finitesimplex_simplexpairshift_of_eq_zero_of_le a one-sided transfer from a positive donor to a zero coordinate stays in the finite simplex. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_simplexPairShift_eq","label":"linearLoss_simplexPairShift_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_simplexPairShift_eq","description":"Exact linear-loss change under a two-coordinate simplex transfer.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-1ae3fe110b8b","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9652,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_simplexPairShift_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p score : Action -> Real) (i j : Action) (t : Real) (hi : i ∈ arms) (hj : j ∈ arms) : FTRL.linearLoss arms (simplexPairShift p i j t) score = FTRL.linearLoss arms p score + t * (score i - score j)","missing":[],"search":"linearloss_simplexpairshift_eq banditrlproof.tsallis.linearloss_simplexpairshift_eq exact linear-loss change under a two-coordinate simplex transfer. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sqrt_simplexPairShift_eq_of_eq_zero","label":"sum_sqrt_simplexPairShift_eq_of_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sqrt_simplexPairShift_eq_of_eq_zero","description":"Exact square-root-sum change when mass enters a zero coordinate.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-b13b3d57ef66","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9653,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:55"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sqrt_simplexPairShift_eq_of_eq_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (t : Real) (hij : i ≠ j) (hi : i ∈ arms) (hj : j ∈ arms) (hpi : p i = 0) : arms.sum (fun a => Real.sqrt (simplexPairShift p i j t a)) = arms.sum (fun a => Real.sqrt (p a)) + Real.sqrt t + Real.sqrt (p j - t) - Real.sqrt (p j)","missing":[],"search":"sum_sqrt_simplexpairshift_eq_of_eq_zero banditrlproof.tsallis.sum_sqrt_simplexpairshift_eq_of_eq_zero exact square-root-sum change when mass enters a zero coordinate. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_sub_sqrt_sub_le_div_sqrt","label":"sqrt_sub_sqrt_sub_le_div_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_sub_sqrt_sub_le_div_sqrt","description":"Removing mass `t` from a positive coordinate loses at most `t / sqrt r` of square-root mass.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-e9b78cabccf1","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9654,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:90"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sqrt_sub_sqrt_sub_le_div_sqrt {r t : Real} (hr : 0 < r) (ht : 0 <= t) (htr : t <= r) : Real.sqrt r - Real.sqrt (r - t) <= t / Real.sqrt r","missing":[],"search":"sqrt_sub_sqrt_sub_le_div_sqrt banditrlproof.tsallis.sqrt_sub_sqrt_sub_le_div_sqrt removing mass `t` from a positive coordinate loses at most `t / sqrt r` of square-root mass. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_transfer_strictly_improves_half_objective","label":"exists_transfer_strictly_improves_half_objective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_transfer_strictly_improves_half_objective","description":"At a zero coordinate, the square-root gain dominates any fixed linear slope for a sufficiently small positive transfer from a positive donor.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-376b88a0e9a8","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9655,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:104"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_transfer_strictly_improves_half_objective (r d : Real) (hr : 0 < r) : exists t, 0 < t ∧ t < r ∧ d * t < 2 * (Real.sqrt t + Real.sqrt (r - t) - Real.sqrt r)","missing":[],"search":"exists_transfer_strictly_improves_half_objective banditrlproof.tsallis.exists_transfer_strictly_improves_half_objective at a zero coordinate, the square-root gain dominates any fixed linear slope for a sufficiently small positive transfer from a positive donor. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_pos","label":"isRegularizedMinimizer_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.isRegularizedMinimizer_pos","description":"Every supported coordinate of a half-Tsallis finite-simplex minimizer is strictly positive. No sign condition on `eta` or boundedness condition on the finite score is needed.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-44abcb0ef467","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9656,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:168"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem isRegularizedMinimizer_pos {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) : forall action, action ∈ arms -> 0 < p action","missing":[],"search":"isregularizedminimizer_pos banditrlproof.tsallis.isregularizedminimizer_pos every supported coordinate of a half-tsallis finite-simplex minimizer is strictly positive. no sign condition on `eta` or boundedness condition on the finite score is needed. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer_auto","label":"exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer_auto","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer_auto","description":"A half-Tsallis simplex minimizer automatically supplies the interior stationarity certificate used by the one-step stability route.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-545bb6062821","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9657,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:212"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer_auto {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) : exists multiplier, HalfTsallisInteriorStationary arms eta score p multiplier","missing":[],"search":"exists_halftsallisinteriorstationary_of_isregularizedminimizer_auto banditrlproof.tsallis.exists_halftsallisinteriorstationary_of_isregularizedminimizer_auto a half-tsallis simplex minimizer automatically supplies the interior stationarity certificate used by the one-step stability route. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_minimizers","label":"sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_minimizers","description":"Sampling-law half-Tsallis stability directly from current and chosen-update minimizer certificates. Strict positivity and common multipliers are derived internally.","url":"../modules/banditrlproof-tsallisftrlinteriority/index.html#decl-8fd3fb93e44c","parent":"module:BanditRLProof.TsallisFTRLInteriority","order":9658,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLInteriority"],["Source","BanditRLProof/TsallisFTRLInteriority.lean:227"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_minimizers {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score prob loss : Action -> Real) (next : Action -> Action -> Real) (heta : 0 < eta) (hprobMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score prob) (hnextMin : forall chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss prob loss chosen action) (next chosen)) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => prob chosen * (FTRL.linearLoss arms prob (Exp3.importanceWeightedLoss prob loss chosen) - FTRL.linearLoss arms (next chosen) (Exp3.importanceWeightedLoss prob…","missing":[],"search":"sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_powersum_half_of_minimizers banditrlproof.tsallis.sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_powersum_half_of_minimizers sampling-law half-tsallis stability directly from current and chosen-update minimizer certificates. strict positivity and common multipliers are derived internally. theorem compiled","shard":"modules/0cc655f0d890710a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.extendFiniteWeights","label":"extendFiniteWeights","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.extendFiniteWeights","description":"Extend finite-subtype weights by zero outside the explicit arm set.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-8ec7beaf81c1","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9659,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def extendFiniteWeights {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) : Action -> Real","missing":[],"search":"extendfiniteweights banditrlproof.tsallis.extendfiniteweights extend finite-subtype weights by zero outside the explicit arm set. definition compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.extendFiniteWeights_apply_of_mem","label":"extendFiniteWeights_apply_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.extendFiniteWeights_apply_of_mem","description":"theorem extendFiniteWeights_apply_of_mem {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) {action : Action} (haction : action ∈ arms) : extendFiniteWeights arms p action = p ⟨action, haction⟩","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-d4d7334ed35b","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9660,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem extendFiniteWeights_apply_of_mem {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) {action : Action} (haction : action ∈ arms) : extendFiniteWeights arms p action = p ⟨action, haction⟩","missing":[],"search":"extendfiniteweights_apply_of_mem banditrlproof.tsallis.extendfiniteweights_apply_of_mem theorem extendfiniteweights_apply_of_mem {action : type u} [decidableeq action] (arms : finset action) (p : ↥arms -> real) {action : action} (haction : action ∈ arms) : extendfiniteweights arms p action = p ⟨action, haction⟩ theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.extendFiniteWeights_apply_of_not_mem","label":"extendFiniteWeights_apply_of_not_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.extendFiniteWeights_apply_of_not_mem","description":"theorem extendFiniteWeights_apply_of_not_mem {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) {action : Action} (haction : action ∉ arms) : extendFiniteWeights arms p action = 0","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-970a0bd4b7dd","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9661,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem extendFiniteWeights_apply_of_not_mem {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) {action : Action} (haction : action ∉ arms) : extendFiniteWeights arms p action = 0","missing":[],"search":"extendfiniteweights_apply_of_not_mem banditrlproof.tsallis.extendfiniteweights_apply_of_not_mem theorem extendfiniteweights_apply_of_not_mem {action : type u} [decidableeq action] (arms : finset action) (p : ↥arms -> real) {action : action} (haction : action ∉ arms) : extendfiniteweights arms p action = 0 theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_extendFiniteWeights","label":"sum_extendFiniteWeights","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_extendFiniteWeights","description":"theorem sum_extendFiniteWeights {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) : arms.sum (extendFiniteWeights arms p) = ∑ action : ↥arms, p action","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-958d77615401","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9662,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_extendFiniteWeights {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) : arms.sum (extendFiniteWeights arms p) = ∑ action : ↥arms, p action","missing":[],"search":"sum_extendfiniteweights banditrlproof.tsallis.sum_extendfiniteweights theorem sum_extendfiniteweights {action : type u} [decidableeq action] (arms : finset action) (p : ↥arms -> real) : arms.sum (extendfiniteweights arms p) = ∑ action : ↥arms, p action theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_extendFiniteWeights","label":"finiteSimplex_extendFiniteWeights","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_extendFiniteWeights","description":"theorem finiteSimplex_extendFiniteWeights {Action : Type u} [DecidableEq Action] (arms : Finset Action) {p : ↥arms -> Real} (hp : p ∈ stdSimplex Real ↥arms) : FTRL.finiteSimplex arms (extendFiniteWeights arms p)","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-b35c57e872e6","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9663,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_extendFiniteWeights {Action : Type u} [DecidableEq Action] (arms : Finset Action) {p : ↥arms -> Real} (hp : p ∈ stdSimplex Real ↥arms) : FTRL.finiteSimplex arms (extendFiniteWeights arms p)","missing":[],"search":"finitesimplex_extendfiniteweights banditrlproof.tsallis.finitesimplex_extendfiniteweights theorem finitesimplex_extendfiniteweights {action : type u} [decidableeq action] (arms : finset action) {p : ↥arms -> real} (hp : p ∈ stdsimplex real ↥arms) : ftrl.finitesimplex arms (extendfiniteweights arms p) theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.restrict_mem_stdSimplex_of_finiteSimplex","label":"restrict_mem_stdSimplex_of_finiteSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.restrict_mem_stdSimplex_of_finiteSimplex","description":"theorem restrict_mem_stdSimplex_of_finiteSimplex {Action : Type u} [DecidableEq Action] (arms : Finset Action) {p : Action -> Real} (hp : FTRL.finiteSimplex arms p) : Finset.restrict arms p ∈ stdSimplex Real ↥arms","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-864938d8f5b8","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9664,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem restrict_mem_stdSimplex_of_finiteSimplex {Action : Type u} [DecidableEq Action] (arms : Finset Action) {p : Action -> Real} (hp : FTRL.finiteSimplex arms p) : Finset.restrict arms p ∈ stdSimplex Real ↥arms","missing":[],"search":"restrict_mem_stdsimplex_of_finitesimplex banditrlproof.tsallis.restrict_mem_stdsimplex_of_finitesimplex theorem restrict_mem_stdsimplex_of_finitesimplex {action : type u} [decidableeq action] (arms : finset action) {p : action -> real} (hp : ftrl.finitesimplex arms p) : finset.restrict arms p ∈ stdsimplex real ↥arms theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_extendFiniteWeights_eq","label":"linearLoss_extendFiniteWeights_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_extendFiniteWeights_eq","description":"theorem linearLoss_extendFiniteWeights_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) (score : Action -> Real) : FTRL.linearLoss arms (extendFiniteWeights arms p) score = FTRL.linearLoss Finset.univ p (fun action => score action)","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-64f3755b7a59","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9665,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:70"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_extendFiniteWeights_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) (score : Action -> Real) : FTRL.linearLoss arms (extendFiniteWeights arms p) score = FTRL.linearLoss Finset.univ p (fun action => score action)","missing":[],"search":"linearloss_extendfiniteweights_eq banditrlproof.tsallis.linearloss_extendfiniteweights_eq theorem linearloss_extendfiniteweights_eq {action : type u} [decidableeq action] (arms : finset action) (p : ↥arms -> real) (score : action -> real) : ftrl.linearloss arms (extendfiniteweights arms p) score = ftrl.linearloss finset.univ p (fun action => score action) theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_restrict_eq","label":"linearLoss_restrict_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_restrict_eq","description":"theorem linearLoss_restrict_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p score : Action -> Real) : FTRL.linearLoss Finset.univ (Finset.restrict arms p) (fun action => score action) = FTRL.linearLoss arms p score","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-732156e83888","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9666,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:79"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_restrict_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p score : Action -> Real) : FTRL.linearLoss Finset.univ (Finset.restrict arms p) (fun action => score action) = FTRL.linearLoss arms p score","missing":[],"search":"linearloss_restrict_eq banditrlproof.tsallis.linearloss_restrict_eq theorem linearloss_restrict_eq {action : type u} [decidableeq action] (arms : finset action) (p score : action -> real) : ftrl.linearloss finset.univ (finset.restrict arms p) (fun action => score action) = ftrl.linearloss arms p score theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sqrt_extendFiniteWeights_eq","label":"sum_sqrt_extendFiniteWeights_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sqrt_extendFiniteWeights_eq","description":"theorem sum_sqrt_extendFiniteWeights_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) : arms.sum (fun action => Real.sqrt (extendFiniteWeights arms p action)) = ∑ action : ↥arms, Real.sqrt (p action)","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-17cdcddde995","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9667,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:89"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sqrt_extendFiniteWeights_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : ↥arms -> Real) : arms.sum (fun action => Real.sqrt (extendFiniteWeights arms p action)) = ∑ action : ↥arms, Real.sqrt (p action)","missing":[],"search":"sum_sqrt_extendfiniteweights_eq banditrlproof.tsallis.sum_sqrt_extendfiniteweights_eq theorem sum_sqrt_extendfiniteweights_eq {action : type u} [decidableeq action] (arms : finset action) (p : ↥arms -> real) : arms.sum (fun action => real.sqrt (extendfiniteweights arms p action)) = ∑ action : ↥arms, real.sqrt (p action) theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sqrt_restrict_eq","label":"sum_sqrt_restrict_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sqrt_restrict_eq","description":"theorem sum_sqrt_restrict_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) : (Finset.univ : Finset ↥arms).sum (fun action => Real.sqrt (Finset.restrict arms p action)) = arms.sum (fun action => Real.sqrt (p action))","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-3b74ead53a17","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9668,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:97"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sqrt_restrict_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) : (Finset.univ : Finset ↥arms).sum (fun action => Real.sqrt (Finset.restrict arms p action)) = arms.sum (fun action => Real.sqrt (p action))","missing":[],"search":"sum_sqrt_restrict_eq banditrlproof.tsallis.sum_sqrt_restrict_eq theorem sum_sqrt_restrict_eq {action : type u} [decidableeq action] (arms : finset action) (p : action -> real) : (finset.univ : finset ↥arms).sum (fun action => real.sqrt (finset.restrict arms p action)) = arms.sum (fun action => real.sqrt (p action)) theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.regularizedObjective_half_extendFiniteWeights_eq","label":"regularizedObjective_half_extendFiniteWeights_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.regularizedObjective_half_extendFiniteWeights_eq","description":"theorem regularizedObjective_half_extendFiniteWeights_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score : Action -> Real) (p : ↥arms -> Real) : FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (extendFiniteWeights arms p) = FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) (fun action : ↥arms => score ac…","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-df5616df08f7","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9669,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:106"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regularizedObjective_half_extendFiniteWeights_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score : Action -> Real) (p : ↥arms -> Real) : FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (extendFiniteWeights arms p) = FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) (fun action : ↥arms => score action) p","missing":[],"search":"regularizedobjective_half_extendfiniteweights_eq banditrlproof.tsallis.regularizedobjective_half_extendfiniteweights_eq theorem regularizedobjective_half_extendfiniteweights_eq {action : type u} [decidableeq action] (arms : finset action) (eta : real) (score : action -> real) (p : ↥arms -> real) : ftrl.regularizedobjective arms eta (negentropyregularizer arms (1 / 2 : real)) score (extendfiniteweights arms p) = ftrl.regularizedobjective arms.attach eta (negentropyregularizer arms.attach (1 / 2 : real)) (fun action : ↥arms => score action) p theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.regularizedObjective_half_restrict_eq","label":"regularizedObjective_half_restrict_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.regularizedObjective_half_restrict_eq","description":"theorem regularizedObjective_half_restrict_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) : FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) (fun action : ↥arms => score action) (Finset.restrict arms p) = FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-ce3b454ef4c5","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9670,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:120"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regularizedObjective_half_restrict_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) : FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) (fun action : ↥arms => score action) (Finset.restrict arms p) = FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p","missing":[],"search":"regularizedobjective_half_restrict_eq banditrlproof.tsallis.regularizedobjective_half_restrict_eq theorem regularizedobjective_half_restrict_eq {action : type u} [decidableeq action] (arms : finset action) (eta : real) (score p : action -> real) : ftrl.regularizedobjective arms.attach eta (negentropyregularizer arms.attach (1 / 2 : real)) (fun action : ↥arms => score action) (finset.restrict arms p) = ftrl.regularizedobjective arms eta (negentropyregularizer arms (1 / 2 : real)) score p theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.continuous_regularizedObjective_half_univ","label":"continuous_regularizedObjective_half_univ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.continuous_regularizedObjective_half_univ","description":"The half-Tsallis objective is continuous on the finite subtype function space.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-731be202ecaa","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9671,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:135"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem continuous_regularizedObjective_half_univ {Action : Type u} [Fintype Action] (eta : Real) (score : Action -> Real) : Continuous (fun p : Action -> Real => FTRL.regularizedObjective Finset.univ eta (negEntropyRegularizer Finset.univ (1 / 2 : Real)) score p)","missing":[],"search":"continuous_regularizedobjective_half_univ banditrlproof.tsallis.continuous_regularizedobjective_half_univ the half-tsallis objective is continuous on the finite subtype function space. theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_isRegularizedMinimizer_half","label":"exists_isRegularizedMinimizer_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_isRegularizedMinimizer_half","description":"Every nonempty explicit finite arm set admits a half-Tsallis regularized minimizer. The learning rate and finite score coordinates may be arbitrary real numbers.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-640393aa10ad","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9672,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:147"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_isRegularizedMinimizer_half {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) : exists p, FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p","missing":[],"search":"exists_isregularizedminimizer_half banditrlproof.tsallis.exists_isregularizedminimizer_half every nonempty explicit finite arm set admits a half-tsallis regularized minimizer. the learning rate and finite score coordinates may be arbitrary real numbers. theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer","label":"halfTsallisMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisMinimizer","description":"A fixed choice of half-Tsallis minimizer on a nonempty explicit arm set.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-fdb0f47d23ff","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9673,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:181"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) : Action -> Real","missing":[],"search":"halftsallisminimizer banditrlproof.tsallis.halftsallisminimizer a fixed choice of half-tsallis minimizer on a nonempty explicit arm set. definition compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer_isRegularizedMinimizer","label":"halfTsallisMinimizer_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisMinimizer_isRegularizedMinimizer","description":"theorem halfTsallisMinimizer_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (halfTsallisMinimizer arms harms eta score)","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-941cdc4dac78","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9674,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:187"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisMinimizer_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (halfTsallisMinimizer arms harms eta score)","missing":[],"search":"halftsallisminimizer_isregularizedminimizer banditrlproof.tsallis.halftsallisminimizer_isregularizedminimizer theorem halftsallisminimizer_isregularizedminimizer {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (score : action -> real) : ftrl.isregularizedminimizer (ftrl.finitesimplex arms) arms eta (negentropyregularizer arms (1 / 2 : real)) score (halftsallisminimizer arms harms eta score) theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisUpdatedMinimizer","label":"halfTsallisUpdatedMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisUpdatedMinimizer","description":"Canonical half-Tsallis update after observing one importance-weighted loss.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-ebd2b54c3e15","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9675,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:197"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisUpdatedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : Action -> Real) (chosen : Action) : Action -> Real","missing":[],"search":"halftsallisupdatedminimizer banditrlproof.tsallis.halftsallisupdatedminimizer canonical half-tsallis update after observing one importance-weighted loss. definition compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisUpdatedMinimizer_isRegularizedMinimizer","label":"halfTsallisUpdatedMinimizer_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisUpdatedMinimizer_isRegularizedMinimizer","description":"theorem halfTsallisUpdatedMinimizer_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : Action -> Real) (chosen : Action) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss (halfTsallisMinimizer arms harms eta score) lo…","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-239561e37633","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9676,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:206"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisUpdatedMinimizer_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : Action -> Real) (chosen : Action) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss (halfTsallisMinimizer arms harms eta score) loss chosen action) (halfTsallisUpdatedMinimizer arms harms eta score loss chosen)","missing":[],"search":"halftsallisupdatedminimizer_isregularizedminimizer banditrlproof.tsallis.halftsallisupdatedminimizer_isregularizedminimizer theorem halftsallisupdatedminimizer_isregularizedminimizer {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (score loss : action -> real) (chosen : action) : ftrl.isregularizedminimizer (ftrl.finitesimplex arms) arms eta (negentropyregularizer arms (1 / 2 : real)) (fun action => score action + exp3.importanceweightedloss (halftsallisminimizer arms harms eta score) loss chosen action) (halftsallisupdatedminimizer arms harms eta score loss chosen) theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisMinimizer_mul_linearLoss_sub_updated_le_powerSum_half","label":"sum_halfTsallisMinimizer_mul_linearLoss_sub_updated_le_powerSum_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisMinimizer_mul_linearLoss_sub_updated_le_powerSum_half","description":"The sampling-law one-step stability endpoint with both current and updated minimizers selected internally.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-5251322995e5","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9677,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:223"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisMinimizer_mul_linearLoss_sub_updated_le_powerSum_half {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score loss : Action -> Real) (heta : 0 < eta) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : let prob := halfTsallisMinimizer arms harms eta score let next := fun chosen => halfTsallisUpdatedMinimizer arms harms eta score loss chosen arms.sum (fun chosen => prob chosen * (FTRL.linearLoss arms prob (Exp3.importanceWeightedLoss prob loss chosen) - FTRL.linearLoss arms (next chosen) (Exp3.importanceWeightedLoss prob loss chosen))) <= 2 * eta * powerSum arms (1 / 2 : Real) prob","missing":[],"search":"sum_halftsallisminimizer_mul_linearloss_sub_updated_le_powersum_half banditrlproof.tsallis.sum_halftsallisminimizer_mul_linearloss_sub_updated_le_powersum_half the sampling-law one-step stability endpoint with both current and updated minimizers selected internally. theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_minimizer","label":"exists_halfTsallisInteriorStationary_minimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_minimizer","description":"A nonempty finite arm set admits a strictly positive half-Tsallis minimizer together with its common stationarity multiplier.","url":"../modules/banditrlproof-tsallisftrlminimizerexistence/index.html#decl-26cc569af67b","parent":"module:BanditRLProof.TsallisFTRLMinimizerExistence","order":9678,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerExistence"],["Source","BanditRLProof/TsallisFTRLMinimizerExistence.lean:255"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_halfTsallisInteriorStationary_minimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) : exists p multiplier, FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p ∧ (forall action, action ∈ arms -> 0 < p action) ∧ HalfTsallisInteriorStationary arms eta score p multiplier","missing":[],"search":"exists_halftsallisinteriorstationary_minimizer banditrlproof.tsallis.exists_halftsallisinteriorstationary_minimizer a nonempty finite arm set admits a strictly positive half-tsallis minimizer together with its common stationarity multiplier. theorem compiled","shard":"modules/c7437c4d4a66ce50.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer","label":"restrictedHalfTsallisMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer","description":"The canonical project minimizer, restricted to the explicit finite arm subtype and parameterized by a score vector on that subtype.","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-e9ef766960bb","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9679,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def restrictedHalfTsallisMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : ↥arms -> Real) : ↥arms -> Real","missing":[],"search":"restrictedhalftsallisminimizer banditrlproof.tsallis.restrictedhalftsallisminimizer the canonical project minimizer, restricted to the explicit finite arm subtype and parameterized by a score vector on that subtype. definition compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_mem_stdSimplex","label":"restrictedHalfTsallisMinimizer_mem_stdSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_mem_stdSimplex","description":"theorem restrictedHalfTsallisMinimizer_mem_stdSimplex {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : ↥arms -> Real) : restrictedHalfTsallisMinimizer arms harms eta score ∈ stdSimplex Real ↥arms","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-cd0127f71af6","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9680,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem restrictedHalfTsallisMinimizer_mem_stdSimplex {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : ↥arms -> Real) : restrictedHalfTsallisMinimizer arms harms eta score ∈ stdSimplex Real ↥arms","missing":[],"search":"restrictedhalftsallisminimizer_mem_stdsimplex banditrlproof.tsallis.restrictedhalftsallisminimizer_mem_stdsimplex theorem restrictedhalftsallisminimizer_mem_stdsimplex {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (score : ↥arms -> real) : restrictedhalftsallisminimizer arms harms eta score ∈ stdsimplex real ↥arms theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_isMinOn","label":"restrictedHalfTsallisMinimizer_isMinOn","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_isMinOn","description":"theorem restrictedHalfTsallisMinimizer_isMinOn {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : ↥arms -> Real) : IsMinOn (fun weights : ↥arms -> Real => FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) score weights) (stdSimplex Real ↥arms) (restrictedHalfTsallisMinimizer arms harms eta score)","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-4451ab994139","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9681,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem restrictedHalfTsallisMinimizer_isMinOn {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : ↥arms -> Real) : IsMinOn (fun weights : ↥arms -> Real => FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) score weights) (stdSimplex Real ↥arms) (restrictedHalfTsallisMinimizer arms harms eta score)","missing":[],"search":"restrictedhalftsallisminimizer_isminon banditrlproof.tsallis.restrictedhalftsallisminimizer_isminon theorem restrictedhalftsallisminimizer_isminon {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (score : ↥arms -> real) : isminon (fun weights : ↥arms -> real => ftrl.regularizedobjective arms.attach eta (negentropyregularizer arms.attach (1 / 2 : real)) score weights) (stdsimplex real ↥arms) (restrictedhalftsallisminimizer arms harms eta score) theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.continuous_regularizedObjective_half_restricted_joint","label":"continuous_regularizedObjective_half_restricted_joint","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.continuous_regularizedObjective_half_restricted_joint","description":"The restricted half-Tsallis objective is jointly continuous in the finite score vector and simplex vector.","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-ad7c7720224a","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9682,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:67"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem continuous_regularizedObjective_half_restricted_joint {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) : Continuous (fun pair : (↥arms -> Real) × (↥arms -> Real) => FTRL.regularizedObjective arms.attach eta (negEntropyRegularizer arms.attach (1 / 2 : Real)) pair.1 pair.2)","missing":[],"search":"continuous_regularizedobjective_half_restricted_joint banditrlproof.tsallis.continuous_regularizedobjective_half_restricted_joint the restricted half-tsallis objective is jointly continuous in the finite score vector and simplex vector. theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.continuous_restrictedHalfTsallisMinimizer","label":"continuous_restrictedHalfTsallisMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.continuous_restrictedHalfTsallisMinimizer","description":"The finite-coordinate canonical half-Tsallis minimizer is continuous in the finite score vector.","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-a7be2f6e0a46","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9683,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:80"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem continuous_restrictedHalfTsallisMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Continuous (restrictedHalfTsallisMinimizer arms harms eta)","missing":[],"search":"continuous_restrictedhalftsallisminimizer banditrlproof.tsallis.continuous_restrictedhalftsallisminimizer the finite-coordinate canonical half-tsallis minimizer is continuous in the finite score vector. theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer_eq_on_arms_of_score_eq","label":"halfTsallisMinimizer_eq_on_arms_of_score_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisMinimizer_eq_on_arms_of_score_eq","description":"Changing score coordinates outside the explicit arm set does not change the canonical half-Tsallis minimizer on supported coordinates.","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-1e5069d59013","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9684,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:140"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisMinimizer_eq_on_arms_of_score_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score₁ score₂ : Action -> Real) (hscore : forall action, action ∈ arms -> score₁ action = score₂ action) : forall action, action ∈ arms -> halfTsallisMinimizer arms harms eta score₁ action = halfTsallisMinimizer arms harms eta score₂ action","missing":[],"search":"halftsallisminimizer_eq_on_arms_of_score_eq banditrlproof.tsallis.halftsallisminimizer_eq_on_arms_of_score_eq changing score coordinates outside the explicit arm set does not change the canonical half-tsallis minimizer on supported coordinates. theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_restrict_score_apply","label":"restrictedHalfTsallisMinimizer_restrict_score_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_restrict_score_apply","description":"theorem restrictedHalfTsallisMinimizer_restrict_score_apply {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) (action : ↥arms) : restrictedHalfTsallisMinimizer arms harms eta (Finset.restrict arms score) action = halfTsallisMinimizer arms harms eta score action","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-e99747b2a4f4","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9685,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:175"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem restrictedHalfTsallisMinimizer_restrict_score_apply {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Action -> Real) (action : ↥arms) : restrictedHalfTsallisMinimizer arms harms eta (Finset.restrict arms score) action = halfTsallisMinimizer arms harms eta score action","missing":[],"search":"restrictedhalftsallisminimizer_restrict_score_apply banditrlproof.tsallis.restrictedhalftsallisminimizer_restrict_score_apply theorem restrictedhalftsallisminimizer_restrict_score_apply {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (score : action -> real) (action : ↥arms) : restrictedhalftsallisminimizer arms harms eta (finset.restrict arms score) action = halftsallisminimizer arms harms eta score action theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_halfTsallisMinimizer_comp","label":"measurable_halfTsallisMinimizer_comp","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_halfTsallisMinimizer_comp","description":"Coordinatewise measurability of supported scores implies coordinatewise measurability of the existing canonical project minimizer.","url":"../modules/banditrlproof-tsallisftrlminimizermeasurability/index.html#decl-a06bead22a50","parent":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","order":9686,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerMeasurability"],["Source","BanditRLProof/TsallisFTRLMinimizerMeasurability.lean:191"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_halfTsallisMinimizer_comp {Omega : Type*} {Action : Type u} [MeasurableSpace Omega] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score : Omega -> Action -> Real) (hscore : forall action, action ∈ arms -> Measurable (fun omega => score omega action)) (action : Action) (haction : action ∈ arms) : Measurable (fun omega => halfTsallisMinimizer arms harms eta (score omega) action)","missing":[],"search":"measurable_halftsallisminimizer_comp banditrlproof.tsallis.measurable_halftsallisminimizer_comp coordinatewise measurability of supported scores implies coordinatewise measurability of the existing canonical project minimizer. theorem compiled","shard":"modules/0a891c6c0a977227.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.strictConcaveOn_sum_sqrt_stdSimplex","label":"strictConcaveOn_sum_sqrt_stdSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.strictConcaveOn_sum_sqrt_stdSimplex","description":"The finite sum of square roots is strictly concave on every nonempty standard simplex.","url":"../modules/banditrlproof-tsallisftrlminimizeruniqueness/index.html#decl-ea83ee30c9a1","parent":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","order":9687,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerUniqueness"],["Source","BanditRLProof/TsallisFTRLMinimizerUniqueness.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem strictConcaveOn_sum_sqrt_stdSimplex {Action : Type u} [Fintype Action] [Nonempty Action] : StrictConcaveOn Real (stdSimplex Real Action) (fun p : Action -> Real => (Finset.univ : Finset Action).sum (fun action => Real.sqrt (p action)))","missing":[],"search":"strictconcaveon_sum_sqrt_stdsimplex banditrlproof.tsallis.strictconcaveon_sum_sqrt_stdsimplex the finite sum of square roots is strictly concave on every nonempty standard simplex. theorem compiled","shard":"modules/b19721492d6eb203.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.strictConvexOn_regularizedObjective_half_stdSimplex","label":"strictConvexOn_regularizedObjective_half_stdSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.strictConvexOn_regularizedObjective_half_stdSimplex","description":"On a nonempty finite standard simplex, the half-Tsallis regularized objective is strictly convex in the probability vector.","url":"../modules/banditrlproof-tsallisftrlminimizeruniqueness/index.html#decl-3903e7d63ebb","parent":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","order":9688,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerUniqueness"],["Source","BanditRLProof/TsallisFTRLMinimizerUniqueness.lean:53"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem strictConvexOn_regularizedObjective_half_stdSimplex {Action : Type u} [Fintype Action] [Nonempty Action] (eta : Real) (score : Action -> Real) : StrictConvexOn Real (stdSimplex Real Action) (fun p => FTRL.regularizedObjective Finset.univ eta (negEntropyRegularizer Finset.univ (1 / 2 : Real)) score p)","missing":[],"search":"strictconvexon_regularizedobjective_half_stdsimplex banditrlproof.tsallis.strictconvexon_regularizedobjective_half_stdsimplex on a nonempty finite standard simplex, the half-tsallis regularized objective is strictly convex in the probability vector. theorem compiled","shard":"modules/b19721492d6eb203.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_half_eq_on_arms","label":"isRegularizedMinimizer_half_eq_on_arms","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.isRegularizedMinimizer_half_eq_on_arms","description":"Two half-Tsallis minimizers have identical coordinates on the explicit arm set. Coordinates outside `arms` are intentionally not constrained by the project finite-simplex predicate or objective.","url":"../modules/banditrlproof-tsallisftrlminimizeruniqueness/index.html#decl-f68427321216","parent":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","order":9689,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerUniqueness"],["Source","BanditRLProof/TsallisFTRLMinimizerUniqueness.lean:96"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem isRegularizedMinimizer_half_eq_on_arms {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score p q : Action -> Real) (hp : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hq : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score q) : forall action, action ∈ arms -> p action = q action","missing":[],"search":"isregularizedminimizer_half_eq_on_arms banditrlproof.tsallis.isregularizedminimizer_half_eq_on_arms two half-tsallis minimizers have identical coordinates on the explicit arm set. coordinates outside `arms` are intentionally not constrained by the project finite-simplex predicate or objective. theorem compiled","shard":"modules/b19721492d6eb203.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer_eq_on_arms","label":"halfTsallisMinimizer_eq_on_arms","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisMinimizer_eq_on_arms","description":"The canonical selected minimizer agrees on every supported coordinate with any other half-Tsallis minimizer certificate.","url":"../modules/banditrlproof-tsallisftrlminimizeruniqueness/index.html#decl-358abac650e9","parent":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","order":9690,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLMinimizerUniqueness"],["Source","BanditRLProof/TsallisFTRLMinimizerUniqueness.lean:146"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisMinimizer_eq_on_arms {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (score p : Action -> Real) (hp : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) : forall action, action ∈ arms -> halfTsallisMinimizer arms harms eta score action = p action","missing":[],"search":"halftsallisminimizer_eq_on_arms banditrlproof.tsallis.halftsallisminimizer_eq_on_arms the canonical selected minimizer agrees on every supported coordinate with any other half-tsallis minimizer certificate. theorem compiled","shard":"modules/b19721492d6eb203.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HalfTsallisInteriorStationary","label":"HalfTsallisInteriorStationary","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HalfTsallisInteriorStationary","description":"Interior first-order stationarity for the half-Tsallis regularizer. For `alpha = 1 / 2`, the coordinate derivative of the local negative Tsallis entropy is `-p_a^(-1/2)`. The common multiplier records the simplex equality constraint; positivity and normalization are kept as separate theorem inputs.","url":"../modules/banditrlproof-tsallisftrlonestepstability/index.html#decl-dc3a8a8aad22","parent":"module:BanditRLProof.TsallisFTRLOneStepStability","order":9691,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLOneStepStability"],["Source","BanditRLProof/TsallisFTRLOneStepStability.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HalfTsallisInteriorStationary {Action : Type u} (arms : Finset Action) (eta : Real) (score p : Action -> Real) (multiplier : Real) : Prop","missing":[],"search":"halftsallisinteriorstationary banditrlproof.tsallis.halftsallisinteriorstationary interior first-order stationarity for the half-tsallis regularizer. for `alpha = 1 / 2`, the coordinate derivative of the local negative tsallis entropy is `-p_a^(-1/2)`. the common multiplier records the simplex equality constraint; positivity and normalization are kept as separate theorem inputs. definition compiled","shard":"modules/8302fed7a3be334d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisInteriorStationary_rpow_sub_rpow_eq","label":"halfTsallisInteriorStationary_rpow_sub_rpow_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisInteriorStationary_rpow_sub_rpow_eq","description":"Subtracting two half-Tsallis stationarity equations isolates the update.","url":"../modules/banditrlproof-tsallisftrlonestepstability/index.html#decl-c74f6c9c0aeb","parent":"module:BanditRLProof.TsallisFTRLOneStepStability","order":9692,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLOneStepStability"],["Source","BanditRLProof/TsallisFTRLOneStepStability.lean:48"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisInteriorStationary_rpow_sub_rpow_eq {Action : Type u} (arms : Finset Action) (eta : Real) (score increment p q : Action -> Real) (multiplier nextMultiplier : Real) (hp : HalfTsallisInteriorStationary arms eta score p multiplier) (hq : HalfTsallisInteriorStationary arms eta (fun action => score action + increment action) q nextMultiplier) {action : Action} (haction : action ∈ arms) : (q action) ^ (-(1 / 2 : Real)) - (p action) ^ (-(1 / 2 : Real)) = eta * increment action - (nextMultiplier - multiplier)","missing":[],"search":"halftsallisinteriorstationary_rpow_sub_rpow_eq banditrlproof.tsallis.halftsallisinteriorstationary_rpow_sub_rpow_eq subtracting two half-tsallis stationarity equations isolates the update. theorem compiled","shard":"modules/8302fed7a3be334d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub","label":"sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub","description":"Scalar half-Tsallis curvature inequality on the positive cone. This is the one-dimensional inequality that converts a negative-half-power gradient displacement into a displacement of probability mass.","url":"../modules/banditrlproof-tsallisftrlonestepstability/index.html#decl-763d06faf514","parent":"module:BanditRLProof.TsallisFTRLOneStepStability","order":9693,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLOneStepStability"],["Source","BanditRLProof/TsallisFTRLOneStepStability.lean:73"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub {p q : Real} (hp : 0 < p) (hq : 0 < q) (hqp : q <= p) : p - q <= 2 * p ^ (3 / 2 : Real) * (q ^ (-(1 / 2 : Real)) - p ^ (-(1 / 2 : Real)))","missing":[],"search":"sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub banditrlproof.tsallis.sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub scalar half-tsallis curvature inequality on the positive cone. this is the one-dimensional inequality that converts a negative-half-power gradient displacement into a displacement of probability mass. theorem compiled","shard":"modules/8302fed7a3be334d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_sub_next_importanceWeightedLoss_le","label":"linearLoss_sub_next_importanceWeightedLoss_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_sub_next_importanceWeightedLoss_le","description":"Pathwise half-Tsallis FTRL stability for one importance-weighted observation. The current and updated distributions are normalized and strictly positive on `arms`. Their stationarity certificates force the multiplier displacement to lie between zero and the selected coordinate update; the scalar curvature lemma then controls the one-step linear-loss difference.","url":"../modules/banditrlproof-tsallisftrlonestepstability/index.html#decl-388669a046f2","parent":"module:BanditRLProof.TsallisFTRLOneStepStability","order":9694,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLOneStepStability"],["Source","BanditRLProof/TsallisFTRLOneStepStability.lean:134"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_sub_next_importanceWeightedLoss_le {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score prob next loss : Action -> Real) (chosen : Action) (multiplier nextMultiplier : Real) (hchosen : chosen ∈ arms) (heta : 0 < eta) (hprobSimplex : FTRL.finiteSimplex arms prob) (hnextSimplex : FTRL.finiteSimplex arms next) (hprobPos : forall action, action ∈ arms -> 0 < prob action) (hnextPos : forall action, action ∈ arms -> 0 < next action) (hlossNonneg : forall action, action ∈ arms -> 0 <= loss action) (hprobStationary : HalfTsallisInteriorStationary arms eta score prob multiplier) (hnextStationary : HalfTsallisInteriorStationary arms eta (fun action => score action + Exp3.importanceWeightedLoss prob loss chosen action) next nextMultiplier) : FTRL.linearLoss arms prob (Exp3.importanceWeightedLoss prob loss chosen) - FTRL.linearLoss arms next (Exp3.imp…","missing":[],"search":"linearloss_sub_next_importanceweightedloss_le banditrlproof.tsallis.linearloss_sub_next_importanceweightedloss_le pathwise half-tsallis ftrl stability for one importance-weighted observation. the current and updated distributions are normalized and strictly positive on `arms`. their stationarity certificates force the multiplier displacement to lie between zero and the selected coordinate update; the scalar curvature lemma then controls the one-step linear-loss difference. theorem compiled","shard":"modules/8302fed7a3be334d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half","label":"sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half","description":"Sampling-law finite-sum half-Tsallis stability bound. For every possible sampled action, `next chosen` carries its own updated stationarity certificate. Averaging the pathwise FTRL stability terms with the current simplex masses is bounded by `2 * eta` times the half-power sum.","url":"../modules/banditrlproof-tsallisftrlonestepstability/index.html#decl-bf602c2e4d17","parent":"module:BanditRLProof.TsallisFTRLOneStepStability","order":9695,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLOneStepStability"],["Source","BanditRLProof/TsallisFTRLOneStepStability.lean:312"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score prob loss : Action -> Real) (next : Action -> Action -> Real) (multiplier : Real) (nextMultiplier : Action -> Real) (heta : 0 < eta) (hprobSimplex : FTRL.finiteSimplex arms prob) (hprobPos : forall action, action ∈ arms -> 0 < prob action) (hnextSimplex : forall chosen, chosen ∈ arms -> FTRL.finiteSimplex arms (next chosen)) (hnextPos : forall chosen, chosen ∈ arms -> forall action, action ∈ arms -> 0 < next chosen action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) (hprobStationary : HalfTsallisInteriorStationary arms eta score prob multiplier) (hnextStationary : forall chosen, chosen ∈ arms -> HalfTsallisInteriorStationary arms eta (fun action => score action + Exp3.importanceWeightedLoss pr…","missing":[],"search":"sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_powersum_half banditrlproof.tsallis.sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_powersum_half sampling-law finite-sum half-tsallis stability bound. for every possible sampled action, `next chosen` carries its own updated stationarity certificate. averaging the pathwise ftrl stability terms with the current simplex masses is bounded by `2 * eta` times the half-power sum. theorem compiled","shard":"modules/8302fed7a3be334d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HalfTsallisFiniteHistorySelectorMeasurability","label":"HalfTsallisFiniteHistorySelectorMeasurability","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.HalfTsallisFiniteHistorySelectorMeasurability","description":"Reusable regularity boundary for the noncomputable half-Tsallis selector. It asks only that measurable finite-history score coordinates produce measurable selected probability coordinates.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-9a211170949d","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9696,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:26"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure HalfTsallisFiniteHistorySelectorMeasurability {Action : Type u} [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Prop where","missing":[],"search":"halftsallisfinitehistoryselectormeasurability banditrlproof.tsallis.halftsallisfinitehistoryselectormeasurability reusable regularity boundary for the noncomputable half-tsallis selector. it asks only that measurable finite-history score coordinates produce measurable selected probability coordinates. structure compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.initialHalfTsallisDistribution","label":"initialHalfTsallisDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.initialHalfTsallisDistribution","description":"Initial pure half-Tsallis law, before any observed pair.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-bfe54502afe8","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9697,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def initialHalfTsallisDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Action -> Real","missing":[],"search":"initialhalftsallisdistribution banditrlproof.tsallis.initialhalftsallisdistribution initial pure half-tsallis law, before any observed pair. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteActionDistribution_initialHalfTsallisDistribution","label":"finiteActionDistribution_initialHalfTsallisDistribution","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteActionDistribution_initialHalfTsallisDistribution","description":"theorem finiteActionDistribution_initialHalfTsallisDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Exp3.FiniteActionDistribution arms (initialHalfTsallisDistribution arms harms eta)","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-568d24c1fd02","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9698,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:46"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteActionDistribution_initialHalfTsallisDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : Exp3.FiniteActionDistribution arms (initialHalfTsallisDistribution arms harms eta)","missing":[],"search":"finiteactiondistribution_initialhalftsallisdistribution banditrlproof.tsallis.finiteactiondistribution_initialhalftsallisdistribution theorem finiteactiondistribution_initialhalftsallisdistribution {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) : exp3.finiteactiondistribution arms (initialhalftsallisdistribution arms harms eta) theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore","label":"sampledHalfTsallisHistoryScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore","description":"Cumulative importance-weighted score through an inclusive observed pair history, using the pure half-Tsallis law generated by the previous score.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-93dee1b30e61","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9699,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHistoryScore {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) : (n : Nat) -> History.FinitePairHistory Action Real n -> Action -> Real | 0, history, action => Exp3.importanceWeightedLoss (initialHalfTsallisDistribution arms harms eta) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action | n + 1, history, action => let previous := Exp3.previousPairHistory history sampledHalfTsallisHistoryScore arms harms eta n previous action + Exp3.importanceWeightedLoss (halfTsallisMinimizer arms harms eta (sampledHalfTsallisHistoryScore arms harms eta n previous)) (fun _ => (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 action @[simp] theorem sampledHalfTsallisHistoryScore_zero {Action : Type u} [DecidableEq Action] (ar…","missing":[],"search":"sampledhalftsallishistoryscore banditrlproof.tsallis.sampledhalftsallishistoryscore cumulative importance-weighted score through an inclusive observed pair history, using the pure half-tsallis law generated by the previous score. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore_zero","label":"sampledHalfTsallisHistoryScore_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore_zero","description":"theorem sampledHalfTsallisHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (history : History.FinitePairHistory Action Real 0) (action : Action) : sampledHalfTsallisHistoryScore arms harms eta 0 history action = Exp3.importanceWeightedLoss (initialHalfTsallisDistribution arms harms eta) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history…","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-70c56abacbc8","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9700,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:80"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (history : History.FinitePairHistory Action Real 0) (action : Action) : sampledHalfTsallisHistoryScore arms harms eta 0 history action = Exp3.importanceWeightedLoss (initialHalfTsallisDistribution arms harms eta) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action","missing":[],"search":"sampledhalftsallishistoryscore_zero banditrlproof.tsallis.sampledhalftsallishistoryscore_zero theorem sampledhalftsallishistoryscore_zero {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (history : history.finitepairhistory action real 0) (action : action) : sampledhalftsallishistoryscore arms harms eta 0 history action = exp3.importanceweightedloss (initialhalftsallisdistribution arms harms eta) (fun _ => (history ⟨0, finset.mem_iic.mpr le_rfl⟩).2) (history ⟨0, finset.mem_iic.mpr le_rfl⟩).1 action theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore_succ","label":"sampledHalfTsallisHistoryScore_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore_succ","description":"theorem sampledHalfTsallisHistoryScore_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (action : Action) : sampledHalfTsallisHistoryScore arms harms eta (n + 1) history action = sampledHalfTsallisHistoryScore arms harms eta n (Exp3.previousPairHistory history) action + Exp3.importanceWeightedLo…","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-e5f3b3b0c994","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9701,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:92"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisHistoryScore_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (action : Action) : sampledHalfTsallisHistoryScore arms harms eta (n + 1) history action = sampledHalfTsallisHistoryScore arms harms eta n (Exp3.previousPairHistory history) action + Exp3.importanceWeightedLoss (halfTsallisMinimizer arms harms eta (sampledHalfTsallisHistoryScore arms harms eta n (Exp3.previousPairHistory history))) (fun _ => (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 action","missing":[],"search":"sampledhalftsallishistoryscore_succ banditrlproof.tsallis.sampledhalftsallishistoryscore_succ theorem sampledhalftsallishistoryscore_succ {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (n : nat) (history : history.finitepairhistory action real (n + 1)) (action : action) : sampledhalftsallishistoryscore arms harms eta (n + 1) history action = sampledhalftsallishistoryscore arms harms eta n (exp3.previouspairhistory history) action + exp3.importanceweightedloss (halftsallisminimizer arms harms eta (sampledhalftsallishistoryscore arms harms eta n (exp3.previouspairhistory history))) (fun _ => (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).2) (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).1 action theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryScore","label":"measurable_sampledHalfTsallisHistoryScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryScore","description":"The generic selector contract makes the recursively accumulated score measurable at every supported action coordinate.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-1a0ecef50c03","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9702,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:110"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledHalfTsallisHistoryScore {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) : forall n action, action ∈ arms -> Measurable (fun history : History.FinitePairHistory Action Real n => sampledHalfTsallisHistoryScore arms harms eta n history action)","missing":[],"search":"measurable_sampledhalftsallishistoryscore banditrlproof.tsallis.measurable_sampledhalftsallishistoryscore the generic selector contract makes the recursively accumulated score measurable at every supported action coordinate. theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryDistribution","label":"sampledHalfTsallisHistoryDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryDistribution","description":"Pure half-Tsallis probabilities generated by the sampled score.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-9d68e845a0c6","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9703,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:189"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHistoryDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) : History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledhalftsallishistorydistribution banditrlproof.tsallis.sampledhalftsallishistorydistribution pure half-tsallis probabilities generated by the sampled score. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryDistributionSource","label":"sampledHalfTsallisHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryDistributionSource","description":"Measurable finite-action source for the recursively generated policy.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-55f2af35fcff","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9704,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:197"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHistoryDistributionSource {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Exp3.MeasurableFiniteActionDistribution arms (sampledHalfTsallisHistoryDistribution arms harms eta n)","missing":[],"search":"sampledhalftsallishistorydistributionsource banditrlproof.tsallis.sampledhalftsallishistorydistributionsource measurable finite-action source for the recursively generated policy. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryAlgorithm","label":"sampledHalfTsallisHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryAlgorithm","description":"Stochastic finite-history algorithm generated by the recursive pure half-Tsallis score.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-e6d70a55ac2d","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9705,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:218"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHistoryAlgorithm {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) : Thompson.HistoryAlgorithm Action Real where","missing":[],"search":"sampledhalftsallishistoryalgorithm banditrlproof.tsallis.sampledhalftsallishistoryalgorithm stochastic finite-history algorithm generated by the recursive pure half-tsallis score. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryAlgorithm_policy","label":"sampledHalfTsallisHistoryAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryAlgorithm_policy","description":"theorem sampledHalfTsallisHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : (sampledHalfTsallisHistoryAlgorithm arms harms eta selector).policy n = Exp3.finiteActionKernel arms (sampledHalfTsallisHisto…","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-4e21ea59cb10","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9706,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:242"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : (sampledHalfTsallisHistoryAlgorithm arms harms eta selector).policy n = Exp3.finiteActionKernel arms (sampledHalfTsallisHistoryDistribution arms harms eta n) (sampledHalfTsallisHistoryDistributionSource arms harms eta selector n)","missing":[],"search":"sampledhalftsallishistoryalgorithm_policy banditrlproof.tsallis.sampledhalftsallishistoryalgorithm_policy theorem sampledhalftsallishistoryalgorithm_policy {action : type u} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : real) (selector : halftsallisfinitehistoryselectormeasurability arms harms eta) (n : nat) : (sampledhalftsallishistoryalgorithm arms harms eta selector).policy n = exp3.finiteactionkernel arms (sampledhalftsallishistorydistribution arms harms eta n) (sampledhalftsallishistorydistributionsource arms harms eta selector n) theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryKernel","label":"sampledHalfTsallisTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryKernel","description":"Complete environment-indexed recursive pure half-Tsallis trajectory.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-ecc8040f4706","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9707,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:257"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisTrajectoryKernel {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] [Nonempty Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) : Kernel Env ((n : Nat) -> Action × Real)","missing":[],"search":"sampledhalftsallistrajectorykernel banditrlproof.tsallis.sampledhalftsallistrajectorykernel complete environment-indexed recursive pure half-tsallis trajectory. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryMeasure_condDistrib_action","label":"sampledHalfTsallisTrajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryMeasure_condDistrib_action","description":"Every successor action has the recursive pure half-Tsallis finite-action law conditional on its visible pair-history prefix.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-f7e136e278f5","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9708,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:287"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisTrajectoryMeasure_condDistrib_action {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector environment) =ᵐ[ (prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector environment).map (fun sample => Preorder.frestrictLe n sample.2)] Exp3.finiteAc…","missing":[],"search":"sampledhalftsallistrajectorymeasure_conddistrib_action banditrlproof.tsallis.sampledhalftsallistrajectorymeasure_conddistrib_action every successor action has the recursive pure half-tsallis finite-action law conditional on its visible pair-history prefix. theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","label":"sampledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","description":"The same conditional law after retaining the environment in the visible history. The policy kernel is comapped along the pair-history projection.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-96bab5d4d080","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9709,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:319"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2)) (prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector environment) =ᵐ[ (prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector environment)…","missing":[],"search":"sampledhalftsallistrajectorymeasure_conddistrib_action_given_environment banditrlproof.tsallis.sampledhalftsallistrajectorymeasure_conddistrib_action_given_environment the same conditional law after retaining the environment in the visible history. the policy kernel is comapped along the pair-history projection. theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryAt","label":"sampledHalfTsallisHistoryAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryAt","description":"Visible environment/pair-history state before successor action `n + 1`.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-785b37adc707","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9710,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:363"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledHalfTsallisHistoryAt {Env : Type v} {Action : Type u} (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Env × History.FinitePairHistory Action Real n","missing":[],"search":"sampledhalftsallishistoryat banditrlproof.tsallis.sampledhalftsallishistoryat visible environment/pair-history state before successor action `n + 1`. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisActionAt","label":"sampledHalfTsallisActionAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisActionAt","description":"Successor action sampled after the visible prefix through `n`.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-d293cb9603fb","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9711,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:370"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledHalfTsallisActionAt {Env : Type v} {Action : Type u} (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action","missing":[],"search":"sampledhalftsallisactionat banditrlproof.tsallis.sampledhalftsallisactionat successor action sampled after the visible prefix through `n`. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisScoreAt","label":"sampledHalfTsallisScoreAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisScoreAt","description":"Recursive half-Tsallis score on an environment/prefix state.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-8f27b37a5376","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9712,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:376"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisScoreAt {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledhalftsallisscoreat banditrlproof.tsallis.sampledhalftsallisscoreat recursive half-tsallis score on an environment/prefix state. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableLossAt","label":"sampledHalfTsallisPredictableLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableLossAt","description":"Predictable successor loss vector on an environment/prefix state.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-a1eca90114e6","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9713,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:383"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledHalfTsallisPredictableLossAt {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledhalftsallispredictablelossat banditrlproof.tsallis.sampledhalftsallispredictablelossat predictable successor loss vector on an environment/prefix state. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisProbabilityAt","label":"sampledHalfTsallisProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisProbabilityAt","description":"Pure half-Tsallis probability on an environment/prefix state.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-b7c8fc7ab8f1","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9714,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:391"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisProbabilityAt {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledhalftsallisprobabilityat banditrlproof.tsallis.sampledhalftsallisprobabilityat pure half-tsallis probability on an environment/prefix state. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisUpdatedAt","label":"sampledHalfTsallisUpdatedAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisUpdatedAt","description":"Canonical sampled-action update on an environment/prefix state.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-8bbada0374e2","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9715,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:399"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisUpdatedAt {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Action -> Real","missing":[],"search":"sampledhalftsallisupdatedat banditrlproof.tsallis.sampledhalftsallisupdatedat canonical sampled-action update on an environment/prefix state. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryActionStabilityAt","label":"sampledHalfTsallisHistoryActionStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHistoryActionStabilityAt","description":"The one-round stability score on visible history/action pairs.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-d857dd38ee52","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9716,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:410"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHistoryActionStabilityAt {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : (Env × History.FinitePairHistory Action Real n) × Action -> Real","missing":[],"search":"sampledhalftsallishistoryactionstabilityat banditrlproof.tsallis.sampledhalftsallishistoryactionstabilityat the one-round stability score on visible history/action pairs. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHalfPowerBoundAt","label":"sampledHalfTsallisHalfPowerBoundAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisHalfPowerBoundAt","description":"The roundwise half-power budget on a visible history.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-31ea5ceac562","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9717,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:422"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisHalfPowerBoundAt {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Real","missing":[],"search":"sampledhalftsallishalfpowerboundat banditrlproof.tsallis.sampledhalftsallishalfpowerboundat the roundwise half-power budget on a visible history. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisSuccessorStabilityAt","label":"sampledHalfTsallisSuccessorStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisSuccessorStabilityAt","description":"The actual displayed stability term with the next generated prefix's current selector, rather than the sampled-action update notation.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-2d69b7414aea","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9718,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:431"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisSuccessorStabilityAt {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallissuccessorstabilityat banditrlproof.tsallis.sampledhalftsallissuccessorstabilityat the actual displayed stability term with the next generated prefix's current selector, rather than the sampled-action update notation. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisEnvironmentHistoryDistributionSource","label":"sampledHalfTsallisEnvironmentHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisEnvironmentHistoryDistributionSource","description":"Environment-lifted measurable probability source at one prefix level.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-82c153ea1fe4","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9719,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:449"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisEnvironmentHistoryDistributionSource {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Exp3.MeasurableFiniteActionDistribution arms (sampledHalfTsallisProbabilityAt (Env := Env) arms harms eta n)","missing":[],"search":"sampledhalftsallisenvironmenthistorydistributionsource banditrlproof.tsallis.sampledhalftsallisenvironmenthistorydistributionsource environment-lifted measurable probability source at one prefix level. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPolicyAt","label":"sampledHalfTsallisPolicyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPolicyAt","description":"Algorithm policy comapped to the environment/prefix state.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-de84ca10359e","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9720,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:468"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisPolicyAt {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Kernel (Env × History.FinitePairHistory Action Real n) Action","missing":[],"search":"sampledhalftsallispolicyat banditrlproof.tsallis.sampledhalftsallispolicyat algorithm policy comapped to the environment/prefix state. definition compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisScoreAt_succ_ae","label":"sampledHalfTsallisScoreAt_succ_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisScoreAt_succ_ae","description":"The recursive score on generated prefixes follows the predictable importance-weighted update almost surely.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-9c4c2b0cfb53","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9721,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:497"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisScoreAt_succ_ae {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment (fun sample => sampledHalfTsallisHistoryScore arms harms eta (n + 1) (Preorder.frestrictLe (n + 1) sample.2)) =ᵐ[mu] (fun sample candidate => sampledHalfTsallisHistoryScore arms harms eta n (Preorder.frestrictLe n sample.2) candidate + Exp3.importanceWeightedLoss (sampledHalfTsallisHistoryDistribution arms harms eta…","missing":[],"search":"sampledhalftsallisscoreat_succ_ae banditrlproof.tsallis.sampledhalftsallisscoreat_succ_ae the recursive score on generated prefixes follows the predictable importance-weighted update almost surely. theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound","label":"integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound","description":"Generated predictable-trajectory finite-horizon half-Tsallis stability. The trajectory construction discharges the policy-kernel, conditional-law, and successor-score-recursion obligations. The remaining explicit inputs are the canonical selector measurability contract and regularity of the updated stability score under each generated history/action product law.","url":"../modules/banditrlproof-tsallisftrlrecursivetrajectory/index.html#decl-a84071e69d51","parent":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","order":9722,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRecursiveTrajectory"],["Source","BanditRLProof/TsallisFTRLRecursiveTrajectory.lean:550"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (selector : HalfTsallisFiniteHistorySelectorMeasurability arms harms eta) (loss : Exp3.PredictableLossVector Env Action) (hscore : forall n, Measurable (sampledHalfTsallisHistoryActionStabilityAt arms harms eta loss n)) (hIntegrable : forall n, Integrable (sampledHalfTsallisHistoryActionStabilityAt arms harms eta loss n) ((prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment).map (sampledHalfTsallisHistoryAt n) ⊗ₘ sample…","missing":[],"search":"integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound banditrlproof.tsallis.integral_sum_sampledhalftsallissuccessorstability_le_integral_sum_halfpowerstabilitybound generated predictable-trajectory finite-horizon half-tsallis stability. the trajectory construction discharges the policy-kernel, conditional-law, and successor-score-recursion obligations. the remaining explicit inputs are the canonical selector measurability contract and regularity of the updated stability score under each generated history/action product law. theorem compiled","shard":"modules/eca159a1335906e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.cumulativeLoss","label":"cumulativeLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.FTRL.cumulativeLoss","description":"Coordinatewise cumulative loss before round `t`.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-16f13e3f5287","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9723,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def cumulativeLoss {Action : Type u} (loss : Nat -> Action -> Real) (t : Nat) : Action -> Real","missing":[],"search":"cumulativeloss banditrlproof.ftrl.cumulativeloss coordinatewise cumulative loss before round `t`. definition compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.cumulativeLoss_zero","label":"cumulativeLoss_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.cumulativeLoss_zero","description":"@[simp] theorem cumulativeLoss_zero {Action : Type u} (loss : Nat -> Action -> Real) : cumulativeLoss loss 0 = 0","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-a0809663f722","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9724,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"@[simp] theorem cumulativeLoss_zero {Action : Type u} (loss : Nat -> Action -> Real) : cumulativeLoss loss 0 = 0","missing":[],"search":"cumulativeloss_zero banditrlproof.ftrl.cumulativeloss_zero @[simp] theorem cumulativeloss_zero {action : type u} (loss : nat -> action -> real) : cumulativeloss loss 0 = 0 theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.cumulativeLoss_succ","label":"cumulativeLoss_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.cumulativeLoss_succ","description":"Adding one round appends its loss vector to the cumulative loss.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-af809be4c99c","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9725,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:34"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLoss_succ {Action : Type u} (loss : Nat -> Action -> Real) (t : Nat) : cumulativeLoss loss (t + 1) = fun action => cumulativeLoss loss t action + loss t action","missing":[],"search":"cumulativeloss_succ banditrlproof.ftrl.cumulativeloss_succ adding one round appends its loss vector to the cumulative loss. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.linearLoss_add_right","label":"linearLoss_add_right","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.linearLoss_add_right","description":"Finite-action linear loss is additive in its loss-vector argument.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-0ac5763fc50f","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9726,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:42"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_add_right {Action : Type u} (arms : Finset Action) (p x y : Action -> Real) : linearLoss arms p (fun action => x action + y action) = linearLoss arms p x + linearLoss arms p y","missing":[],"search":"linearloss_add_right banditrlproof.ftrl.linearloss_add_right finite-action linear loss is additive in its loss-vector argument. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.linearLoss_cumulativeLoss","label":"linearLoss_cumulativeLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.linearLoss_cumulativeLoss","description":"A linear loss against the cumulative vector is the sum of round losses.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-6da7fb484d85","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9727,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_cumulativeLoss {Action : Type u} (arms : Finset Action) (p : Action -> Real) (loss : Nat -> Action -> Real) (T : Nat) : linearLoss arms p (cumulativeLoss loss T) = (Finset.range T).sum (fun t => linearLoss arms p (loss t))","missing":[],"search":"linearloss_cumulativeloss banditrlproof.ftrl.linearloss_cumulativeloss a linear loss against the cumulative vector is the sum of round losses. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.regularizedObjective_cumulativeLoss_succ","label":"regularizedObjective_cumulativeLoss_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.regularizedObjective_cumulativeLoss_succ","description":"The cumulative regularized objective has the expected successor update.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-33dd9f004059","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9728,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regularizedObjective_cumulativeLoss_succ {Action : Type u} (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Nat -> Action -> Real) (p : Action -> Real) (t : Nat) : regularizedObjective arms eta regularizer (cumulativeLoss loss (t + 1)) p = regularizedObjective arms eta regularizer (cumulativeLoss loss t) p + eta * linearLoss arms p (loss t)","missing":[],"search":"regularizedobjective_cumulativeloss_succ banditrlproof.ftrl.regularizedobjective_cumulativeloss_succ the cumulative regularized objective has the expected successor update. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.eta_mul_sum_next_linearLoss_add_regularizer_zero_le_objective","label":"eta_mul_sum_next_linearLoss_add_regularizer_zero_le_objective","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.eta_mul_sum_next_linearLoss_add_regularizer_zero_le_objective","description":"Scaled regularized be-the-leader inequality for cumulative-loss minimizers. The point `p t` minimizes the objective built from losses before round `t`. Consequently the shifted choices `p (t + 1)` can be compared to the terminal cumulative objective while retaining the initial regularizer value.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-58ad0dfc64a2","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9729,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:76"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem eta_mul_sum_next_linearLoss_add_regularizer_zero_le_objective {Action : Type u} (feasible : (Action -> Real) -> Prop) (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Nat -> Action -> Real) (p : Nat -> Action -> Real) (T : Nat) (hp : forall t, t <= T -> IsRegularizedMinimizer feasible arms eta regularizer (cumulativeLoss loss t) (p t)) : eta * (Finset.range T).sum (fun t => linearLoss arms (p (t + 1)) (loss t)) + regularizer (p 0) <= regularizedObjective arms eta regularizer (cumulativeLoss loss T) (p T)","missing":[],"search":"eta_mul_sum_next_linearloss_add_regularizer_zero_le_objective banditrlproof.ftrl.eta_mul_sum_next_linearloss_add_regularizer_zero_le_objective scaled regularized be-the-leader inequality for cumulative-loss minimizers. the point `p t` minimizes the objective built from losses before round `t`. consequently the shifted choices `p (t + 1)` can be compared to the terminal cumulative objective while retaining the initial regularizer value. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.sum_next_linearLoss_sub_comparator_le_regularizer_penalty","label":"sum_next_linearLoss_sub_comparator_le_regularizer_penalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.sum_next_linearLoss_sub_comparator_le_regularizer_penalty","description":"Regularized be-the-leader bound against any feasible comparator.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-d8eccd460829","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9730,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:134"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_next_linearLoss_sub_comparator_le_regularizer_penalty {Action : Type u} (feasible : (Action -> Real) -> Prop) (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Nat -> Action -> Real) (p : Nat -> Action -> Real) (q : Action -> Real) (T : Nat) (heta : 0 < eta) (hp : forall t, t <= T -> IsRegularizedMinimizer feasible arms eta regularizer (cumulativeLoss loss t) (p t)) (hq : feasible q) : (Finset.range T).sum (fun t => linearLoss arms (p (t + 1)) (loss t) - linearLoss arms q (loss t)) <= (regularizer q - regularizer (p 0)) / eta","missing":[],"search":"sum_next_linearloss_sub_comparator_le_regularizer_penalty banditrlproof.ftrl.sum_next_linearloss_sub_comparator_le_regularizer_penalty regularized be-the-leader bound against any feasible comparator. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.cumulativeLinearLoss_sub_comparator_le_stability_add_penalty","label":"cumulativeLinearLoss_sub_comparator_le_stability_add_penalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.cumulativeLinearLoss_sub_comparator_le_stability_add_penalty","description":"Finite-horizon FTRL stability/penalty regret decomposition. The first sum on the right is the stability term. The second term is the regularizer penalty. No convexity or minimizer-existence theorem is hidden: all cumulative minimizer certificates are explicit inputs.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-71b4f6d2e7c2","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9731,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:181"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLinearLoss_sub_comparator_le_stability_add_penalty {Action : Type u} (feasible : (Action -> Real) -> Prop) (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Nat -> Action -> Real) (p : Nat -> Action -> Real) (q : Action -> Real) (T : Nat) (heta : 0 < eta) (hp : forall t, t <= T -> IsRegularizedMinimizer feasible arms eta regularizer (cumulativeLoss loss t) (p t)) (hq : feasible q) : (Finset.range T).sum (fun t => linearLoss arms (p t) (loss t) - linearLoss arms q (loss t)) <= (Finset.range T).sum (fun t => linearLoss arms (p t) (loss t) - linearLoss arms (p (t + 1)) (loss t)) + (regularizer q - regularizer (p 0)) / eta","missing":[],"search":"cumulativelinearloss_sub_comparator_le_stability_add_penalty banditrlproof.ftrl.cumulativelinearloss_sub_comparator_le_stability_add_penalty finite-horizon ftrl stability/penalty regret decomposition. the first sum on the right is the stability term. the second term is the regularizer penalty. no convexity or minimizer-existence theorem is hidden: all cumulative minimizer certificates are explicit inputs. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.FTRL.cumulativeLinearLoss_sub_comparator_le_stability_add_penalty_simplex","label":"cumulativeLinearLoss_sub_comparator_le_stability_add_penalty_simplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.FTRL.cumulativeLinearLoss_sub_comparator_le_stability_add_penalty_simplex","description":"Finite-simplex specialization of the FTRL decomposition.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-589ebba56453","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9732,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:227"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLinearLoss_sub_comparator_le_stability_add_penalty_simplex {Action : Type u} (arms : Finset Action) (eta : Real) (regularizer : (Action -> Real) -> Real) (loss : Nat -> Action -> Real) (p : Nat -> Action -> Real) (q : Action -> Real) (T : Nat) (heta : 0 < eta) (hp : forall t, t <= T -> IsRegularizedMinimizer (finiteSimplex arms) arms eta regularizer (cumulativeLoss loss t) (p t)) (hq : finiteSimplex arms q) : (Finset.range T).sum (fun t => linearLoss arms (p t) (loss t) - linearLoss arms q (loss t)) <= (Finset.range T).sum (fun t => linearLoss arms (p t) (loss t) - linearLoss arms (p (t + 1)) (loss t)) + (regularizer q - regularizer (p 0)) / eta","missing":[],"search":"cumulativelinearloss_sub_comparator_le_stability_add_penalty_simplex banditrlproof.ftrl.cumulativelinearloss_sub_comparator_le_stability_add_penalty_simplex finite-simplex specialization of the ftrl decomposition. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.negEntropyRegularizer_sub_eq_powerSum_sub_div","label":"negEntropyRegularizer_sub_eq_powerSum_sub_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.negEntropyRegularizer_sub_eq_powerSum_sub_div","description":"Difference of negative Tsallis entropies as a power-sum difference.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-a1b7bdc1a918","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9733,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:254"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem negEntropyRegularizer_sub_eq_powerSum_sub_div {Action : Type u} (arms : Finset Action) (alpha : Real) (p q : Action -> Real) (halpha : Ne alpha 1) : negEntropyRegularizer arms alpha q - negEntropyRegularizer arms alpha p = (powerSum arms alpha p - powerSum arms alpha q) / (1 - alpha)","missing":[],"search":"negentropyregularizer_sub_eq_powersum_sub_div banditrlproof.tsallis.negentropyregularizer_sub_eq_powersum_sub_div difference of negative tsallis entropies as a power-sum difference. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty","label":"cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty","description":"Finite-horizon Tsallis-FTRL stability/penalty regret decomposition. The penalty is explicit in the finite power sums. The remaining algorithmic obligation is to bound the stability sum for the chosen Tsallis exponent and loss estimator, then supply minimizer existence and any stochastic contracts.","url":"../modules/banditrlproof-tsallisftrlregret/index.html#decl-3c8ca355a6f7","parent":"module:BanditRLProof.TsallisFTRLRegret","order":9734,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLRegret"],["Source","BanditRLProof/TsallisFTRLRegret.lean:273"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty {Action : Type u} (arms : Finset Action) (alpha eta : Real) (loss : Nat -> Action -> Real) (p : Nat -> Action -> Real) (q : Action -> Real) (T : Nat) (halpha : Ne alpha 1) (heta : 0 < eta) (hp : forall t, t <= T -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms alpha) (FTRL.cumulativeLoss loss t) (p t)) (hq : FTRL.finiteSimplex arms q) : (Finset.range T).sum (fun t => FTRL.linearLoss arms (p t) (loss t) - FTRL.linearLoss arms q (loss t)) <= (Finset.range T).sum (fun t => FTRL.linearLoss arms (p t) (loss t) - FTRL.linearLoss arms (p (t + 1)) (loss t)) + ((powerSum arms alpha (p 0) - powerSum arms alpha q) / (1 - alpha)) / eta","missing":[],"search":"cumulativelinearloss_sub_comparator_le_stability_add_powersumpenalty banditrlproof.tsallis.cumulativelinearloss_sub_comparator_le_stability_add_powersumpenalty finite-horizon tsallis-ftrl stability/penalty regret decomposition. the penalty is explicit in the finite power sums. the remaining algorithmic obligation is to bound the stability sum for the chosen tsallis exponent and loss estimator, then supply minimizer existence and any stochastic contracts. theorem compiled","shard":"modules/2ffee97809724358.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.pairDirection","label":"pairDirection","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.pairDirection","description":"The zero-sum direction that transfers mass from `j` to `i`.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-3bf5ade5cfa4","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9735,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def pairDirection {Action : Type u} [DecidableEq Action] (i j : Action) (a : Action) : Real","missing":[],"search":"pairdirection banditrlproof.tsallis.pairdirection the zero-sum direction that transfers mass from `j` to `i`. definition compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.simplexPairShift","label":"simplexPairShift","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.simplexPairShift","description":"Transfer scalar mass `t` from coordinate `j` to coordinate `i`.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-a7ea7e106913","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9736,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:36"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def simplexPairShift {Action : Type u} [DecidableEq Action] (p : Action -> Real) (i j : Action) (t : Real) : Action -> Real","missing":[],"search":"simplexpairshift banditrlproof.tsallis.simplexpairshift transfer scalar mass `t` from coordinate `j` to coordinate `i`. definition compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.simplexPairShift_zero","label":"simplexPairShift_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.simplexPairShift_zero","description":"theorem simplexPairShift_zero {Action : Type u} [DecidableEq Action] (p : Action -> Real) (i j : Action) : simplexPairShift p i j 0 = p","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-7c92696e4dbd","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9737,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem simplexPairShift_zero {Action : Type u} [DecidableEq Action] (p : Action -> Real) (i j : Action) : simplexPairShift p i j 0 = p","missing":[],"search":"simplexpairshift_zero banditrlproof.tsallis.simplexpairshift_zero theorem simplexpairshift_zero {action : type u} [decidableeq action] (p : action -> real) (i j : action) : simplexpairshift p i j 0 = p theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_pairDirection_mul","label":"sum_pairDirection_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_pairDirection_mul","description":"theorem sum_pairDirection_mul {Action : Type u} [DecidableEq Action] (arms : Finset Action) (i j : Action) (f : Action -> Real) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (fun a => pairDirection i j a * f a) = f i - f j","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-2a78daad9540","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9738,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_pairDirection_mul {Action : Type u} [DecidableEq Action] (arms : Finset Action) (i j : Action) (f : Action -> Real) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (fun a => pairDirection i j a * f a) = f i - f j","missing":[],"search":"sum_pairdirection_mul banditrlproof.tsallis.sum_pairdirection_mul theorem sum_pairdirection_mul {action : type u} [decidableeq action] (arms : finset action) (i j : action) (f : action -> real) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (fun a => pairdirection i j a * f a) = f i - f j theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_simplexPairShift","label":"sum_simplexPairShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_simplexPairShift","description":"theorem sum_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (t : Real) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (simplexPairShift p i j t) = arms.sum p","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-c2096a2009be","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9739,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:54"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (t : Real) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (simplexPairShift p i j t) = arms.sum p","missing":[],"search":"sum_simplexpairshift banditrlproof.tsallis.sum_simplexpairshift theorem sum_simplexpairshift {action : type u} [decidableeq action] (arms : finset action) (p : action -> real) (i j : action) (t : real) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (simplexpairshift p i j t) = arms.sum p theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_simplexPairShift_of_abs_lt_min","label":"finiteSimplex_simplexPairShift_of_abs_lt_min","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_simplexPairShift_of_abs_lt_min","description":"theorem finiteSimplex_simplexPairShift_of_abs_lt_min {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (t : Real) (hij : i ≠ j) (hp : FTRL.finiteSimplex arms p) (hi : i ∈ arms) (hj : j ∈ arms) (ht : |t| < min (p i) (p j)) : FTRL.finiteSimplex arms (simplexPairShift p i j t)","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-e2dfbe92c8f5","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9740,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:64"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_simplexPairShift_of_abs_lt_min {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (t : Real) (hij : i ≠ j) (hp : FTRL.finiteSimplex arms p) (hi : i ∈ arms) (hj : j ∈ arms) (ht : |t| < min (p i) (p j)) : FTRL.finiteSimplex arms (simplexPairShift p i j t)","missing":[],"search":"finitesimplex_simplexpairshift_of_abs_lt_min banditrlproof.tsallis.finitesimplex_simplexpairshift_of_abs_lt_min theorem finitesimplex_simplexpairshift_of_abs_lt_min {action : type u} [decidableeq action] (arms : finset action) (p : action -> real) (i j : action) (t : real) (hij : i ≠ j) (hp : ftrl.finitesimplex arms p) (hi : i ∈ arms) (hj : j ∈ arms) (ht : |t| < min (p i) (p j)) : ftrl.finitesimplex arms (simplexpairshift p i j t) theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.regularizedObjective_half_eq","label":"regularizedObjective_half_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.regularizedObjective_half_eq","description":"The half-Tsallis objective is a linear term minus twice the square-root sum.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-65cac88b7688","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9741,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:87"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regularizedObjective_half_eq {Action : Type u} (arms : Finset Action) (eta : Real) (score p : Action -> Real) : FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p = eta * FTRL.linearLoss arms p score - 2 * arms.sum (fun a => Real.sqrt (p a)) + 2","missing":[],"search":"regularizedobjective_half_eq banditrlproof.tsallis.regularizedobjective_half_eq the half-tsallis objective is a linear term minus twice the square-root sum. theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasDerivAt_simplexPairShift","label":"hasDerivAt_simplexPairShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasDerivAt_simplexPairShift","description":"theorem hasDerivAt_simplexPairShift {Action : Type u} [DecidableEq Action] (p : Action -> Real) (i j a : Action) : HasDerivAt (fun t => simplexPairShift p i j t a) (pairDirection i j a) 0","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-da7f56ee09ed","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9742,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:98"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_simplexPairShift {Action : Type u} [DecidableEq Action] (p : Action -> Real) (i j a : Action) : HasDerivAt (fun t => simplexPairShift p i j t a) (pairDirection i j a) 0","missing":[],"search":"hasderivat_simplexpairshift banditrlproof.tsallis.hasderivat_simplexpairshift theorem hasderivat_simplexpairshift {action : type u} [decidableeq action] (p : action -> real) (i j a : action) : hasderivat (fun t => simplexpairshift p i j t a) (pairdirection i j a) 0 theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasDerivAt_linearLoss_simplexPairShift","label":"hasDerivAt_linearLoss_simplexPairShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasDerivAt_linearLoss_simplexPairShift","description":"theorem hasDerivAt_linearLoss_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p score : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) : HasDerivAt (fun t => FTRL.linearLoss arms (simplexPairShift p i j t) score) (score i - score j) 0","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-5219ef9a911f","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9743,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:107"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_linearLoss_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p score : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) : HasDerivAt (fun t => FTRL.linearLoss arms (simplexPairShift p i j t) score) (score i - score j) 0","missing":[],"search":"hasderivat_linearloss_simplexpairshift banditrlproof.tsallis.hasderivat_linearloss_simplexpairshift theorem hasderivat_linearloss_simplexpairshift {action : type u} [decidableeq action] (arms : finset action) (p score : action -> real) (i j : action) (hi : i ∈ arms) (hj : j ∈ arms) : hasderivat (fun t => ftrl.linearloss arms (simplexpairshift p i j t) score) (score i - score j) 0 theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_pairDirection_div_two_sqrt","label":"sum_pairDirection_div_two_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_pairDirection_div_two_sqrt","description":"theorem sum_pairDirection_div_two_sqrt {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (fun a => pairDirection i j a / (2 * Real.sqrt (p a))) = 1 / (2 * Real.sqrt (p i)) - 1 / (2 * Real.sqrt (p j))","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-8af352afb488","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9744,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:123"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_pairDirection_div_two_sqrt {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (fun a => pairDirection i j a / (2 * Real.sqrt (p a))) = 1 / (2 * Real.sqrt (p i)) - 1 / (2 * Real.sqrt (p j))","missing":[],"search":"sum_pairdirection_div_two_sqrt banditrlproof.tsallis.sum_pairdirection_div_two_sqrt theorem sum_pairdirection_div_two_sqrt {action : type u} [decidableeq action] (arms : finset action) (p : action -> real) (i j : action) (hi : i ∈ arms) (hj : j ∈ arms) : arms.sum (fun a => pairdirection i j a / (2 * real.sqrt (p a))) = 1 / (2 * real.sqrt (p i)) - 1 / (2 * real.sqrt (p j)) theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasDerivAt_sum_sqrt_simplexPairShift","label":"hasDerivAt_sum_sqrt_simplexPairShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasDerivAt_sum_sqrt_simplexPairShift","description":"theorem hasDerivAt_sum_sqrt_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) (hpPos : forall a, a ∈ arms -> 0 < p a) : HasDerivAt (fun t => arms.sum (fun a => Real.sqrt (simplexPairShift p i j t a))) (1 / (2 * Real.sqrt (p i)) - 1 / (2 * Real.sqrt (p j))) 0","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-0c12e41f8fa7","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9745,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:142"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_sum_sqrt_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (p : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) (hpPos : forall a, a ∈ arms -> 0 < p a) : HasDerivAt (fun t => arms.sum (fun a => Real.sqrt (simplexPairShift p i j t a))) (1 / (2 * Real.sqrt (p i)) - 1 / (2 * Real.sqrt (p j))) 0","missing":[],"search":"hasderivat_sum_sqrt_simplexpairshift banditrlproof.tsallis.hasderivat_sum_sqrt_simplexpairshift theorem hasderivat_sum_sqrt_simplexpairshift {action : type u} [decidableeq action] (arms : finset action) (p : action -> real) (i j : action) (hi : i ∈ arms) (hj : j ∈ arms) (hppos : forall a, a ∈ arms -> 0 < p a) : hasderivat (fun t => arms.sum (fun a => real.sqrt (simplexpairshift p i j t a))) (1 / (2 * real.sqrt (p i)) - 1 / (2 * real.sqrt (p j))) 0 theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasDerivAt_regularizedObjective_half_simplexPairShift","label":"hasDerivAt_regularizedObjective_half_simplexPairShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasDerivAt_regularizedObjective_half_simplexPairShift","description":"theorem hasDerivAt_regularizedObjective_half_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) (hpPos : forall a, a ∈ arms -> 0 < p a) : HasDerivAt (fun t => FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (simplexPairShift p i j t)) (eta * (score i - score j) - (1…","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-71828adc0975","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9746,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:160"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasDerivAt_regularizedObjective_half_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (i j : Action) (hi : i ∈ arms) (hj : j ∈ arms) (hpPos : forall a, a ∈ arms -> 0 < p a) : HasDerivAt (fun t => FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (simplexPairShift p i j t)) (eta * (score i - score j) - (1 / Real.sqrt (p i) - 1 / Real.sqrt (p j))) 0","missing":[],"search":"hasderivat_regularizedobjective_half_simplexpairshift banditrlproof.tsallis.hasderivat_regularizedobjective_half_simplexpairshift theorem hasderivat_regularizedobjective_half_simplexpairshift {action : type u} [decidableeq action] (arms : finset action) (eta : real) (score p : action -> real) (i j : action) (hi : i ∈ arms) (hj : j ∈ arms) (hppos : forall a, a ∈ arms -> 0 < p a) : hasderivat (fun t => ftrl.regularizedobjective arms eta (negentropyregularizer arms (1 / 2 : real)) score (simplexpairshift p i j t)) (eta * (score i - score j) - (1 / real.sqrt (p i) - 1 / real.sqrt (p j))) 0 theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.isLocalMin_regularizedObjective_half_simplexPairShift","label":"isLocalMin_regularizedObjective_half_simplexPairShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.isLocalMin_regularizedObjective_half_simplexPairShift","description":"theorem isLocalMin_regularizedObjective_half_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (i j : Action) (hij : i ≠ j) (hi : i ∈ arms) (hj : j ∈ arms) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hpPos : forall a, a ∈ arms -> 0 < p a) : IsLocalMin (fun t => FTRL.r…","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-5992409a90de","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9747,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:192"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem isLocalMin_regularizedObjective_half_simplexPairShift {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (i j : Action) (hij : i ≠ j) (hi : i ∈ arms) (hj : j ∈ arms) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hpPos : forall a, a ∈ arms -> 0 < p a) : IsLocalMin (fun t => FTRL.regularizedObjective arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score (simplexPairShift p i j t)) 0","missing":[],"search":"islocalmin_regularizedobjective_half_simplexpairshift banditrlproof.tsallis.islocalmin_regularizedobjective_half_simplexpairshift theorem islocalmin_regularizedobjective_half_simplexpairshift {action : type u} [decidableeq action] (arms : finset action) (eta : real) (score p : action -> real) (i j : action) (hij : i ≠ j) (hi : i ∈ arms) (hj : j ∈ arms) (hpmin : ftrl.isregularizedminimizer (ftrl.finitesimplex arms) arms eta (negentropyregularizer arms (1 / 2 : real)) score p) (hppos : forall a, a ∈ arms -> 0 < p a) : islocalmin (fun t => ftrl.regularizedobjective arms eta (negentropyregularizer arms (1 / 2 : real)) score (simplexpairshift p i j t)) 0 theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallis_pairwise_stationary_of_isRegularizedMinimizer","label":"halfTsallis_pairwise_stationary_of_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallis_pairwise_stationary_of_isRegularizedMinimizer","description":"theorem halfTsallis_pairwise_stationary_of_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hpPos : forall a, a ∈ arms -> 0 < p a) {i j : Action} (hi : i ∈ arms) (hj : j ∈ arms) : eta * score i - (p i) ^ (-(1 / 2 : Re…","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-55ddd25b7bf4","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9748,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:220"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallis_pairwise_stationary_of_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hpPos : forall a, a ∈ arms -> 0 < p a) {i j : Action} (hi : i ∈ arms) (hj : j ∈ arms) : eta * score i - (p i) ^ (-(1 / 2 : Real)) = eta * score j - (p j) ^ (-(1 / 2 : Real))","missing":[],"search":"halftsallis_pairwise_stationary_of_isregularizedminimizer banditrlproof.tsallis.halftsallis_pairwise_stationary_of_isregularizedminimizer theorem halftsallis_pairwise_stationary_of_isregularizedminimizer {action : type u} [decidableeq action] (arms : finset action) (eta : real) (score p : action -> real) (hpmin : ftrl.isregularizedminimizer (ftrl.finitesimplex arms) arms eta (negentropyregularizer arms (1 / 2 : real)) score p) (hppos : forall a, a ∈ arms -> 0 < p a) {i j : action} (hi : i ∈ arms) (hj : j ∈ arms) : eta * score i - (p i) ^ (-(1 / 2 : real)) = eta * score j - (p j) ^ (-(1 / 2 : real)) theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer","label":"exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer","description":"An explicitly interior half-Tsallis simplex minimizer admits the stationarity certificate consumed by the one-step stability theorem.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-f688f3093683","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9749,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:252"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (hpMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hpPos : forall a, a ∈ arms -> 0 < p a) : exists multiplier, HalfTsallisInteriorStationary arms eta score p multiplier","missing":[],"search":"exists_halftsallisinteriorstationary_of_isregularizedminimizer banditrlproof.tsallis.exists_halftsallisinteriorstationary_of_isregularizedminimizer an explicitly interior half-tsallis simplex minimizer admits the stationarity certificate consumed by the one-step stability theorem. theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.two_mul_sqrt_sub_sqrt_le_sub_div_sqrt","label":"two_mul_sqrt_sub_sqrt_le_sub_div_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.two_mul_sqrt_sub_sqrt_le_sub_div_sqrt","description":"The supporting-line inequality for the square root at a positive point.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-648c90c49cdb","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9750,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:272"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem two_mul_sqrt_sub_sqrt_le_sub_div_sqrt {p q : Real} (hp : 0 < p) (hq : 0 <= q) : 2 * (Real.sqrt q - Real.sqrt p) <= (q - p) / Real.sqrt p","missing":[],"search":"two_mul_sqrt_sub_sqrt_le_sub_div_sqrt banditrlproof.tsallis.two_mul_sqrt_sub_sqrt_le_sub_div_sqrt the supporting-line inequality for the square root at a positive point. theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_of_halfTsallisInteriorStationary","label":"isRegularizedMinimizer_of_halfTsallisInteriorStationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.isRegularizedMinimizer_of_halfTsallisInteriorStationary","description":"A positive simplex point satisfying half-Tsallis stationarity globally minimizes the corresponding regularized objective on the finite simplex.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-abf496010e46","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9751,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:287"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem isRegularizedMinimizer_of_halfTsallisInteriorStationary {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (multiplier : Real) (hp : FTRL.finiteSimplex arms p) (hpPos : forall a, a ∈ arms -> 0 < p a) (hstationary : HalfTsallisInteriorStationary arms eta score p multiplier) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p","missing":[],"search":"isregularizedminimizer_of_halftsallisinteriorstationary banditrlproof.tsallis.isregularizedminimizer_of_halftsallisinteriorstationary a positive simplex point satisfying half-tsallis stationarity globally minimizes the corresponding regularized objective on the finite simplex. theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_iff_exists_halfTsallisInteriorStationary","label":"isRegularizedMinimizer_iff_exists_halfTsallisInteriorStationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.isRegularizedMinimizer_iff_exists_halfTsallisInteriorStationary","description":"Interior half-Tsallis stationarity is equivalent to simplex minimality.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-57e3b908c9dd","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9752,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:361"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem isRegularizedMinimizer_iff_exists_halfTsallisInteriorStationary {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score p : Action -> Real) (hp : FTRL.finiteSimplex arms p) (hpPos : forall a, a ∈ arms -> 0 < p a) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p ↔ exists multiplier, HalfTsallisInteriorStationary arms eta score p multiplier","missing":[],"search":"isregularizedminimizer_iff_exists_halftsallisinteriorstationary banditrlproof.tsallis.isregularizedminimizer_iff_exists_halftsallisinteriorstationary interior half-tsallis stationarity is equivalent to simplex minimality. theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_positive_minimizers","label":"sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_positive_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_positive_minimizers","description":"Sampling-law half-Tsallis stability directly from concrete regularized-minimizer certificates. Common multipliers are constructed internally from interiority.","url":"../modules/banditrlproof-tsallisftrlstationarity/index.html#decl-574f6dfaf879","parent":"module:BanditRLProof.TsallisFTRLStationarity","order":9753,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFTRLStationarity"],["Source","BanditRLProof/TsallisFTRLStationarity.lean:382"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_positive_minimizers {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score prob loss : Action -> Real) (next : Action -> Action -> Real) (heta : 0 < eta) (hprobMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score prob) (hprobPos : forall action, action ∈ arms -> 0 < prob action) (hnextMin : forall chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss prob loss chosen action) (next chosen)) (hnextPos : forall chosen, chosen ∈ arms -> forall action, action ∈ arms -> 0 < next chosen action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun…","missing":[],"search":"sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_powersum_half_of_positive_minimizers banditrlproof.tsallis.sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_powersum_half_of_positive_minimizers sampling-law half-tsallis stability directly from concrete regularized-minimizer certificates. common multipliers are constructed internally from interiority. theorem compiled","shard":"modules/d144522ae1500cb1.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.armDependentSuboptimalRewardBoostSource","label":"armDependentSuboptimalRewardBoostSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.armDependentSuboptimalRewardBoostSource","description":"A stationary but arm-dependent corruption process that leaves the best arm unchanged and adds a prescribed nonnegative boost to every other arm.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html#decl-7da1e86cc0c7","parent":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","order":9754,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean:11"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def armDependentSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (boost : Fin K -> Real) (hboost : forall arm, 0 <= boost arm) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"armdependentsuboptimalrewardboostsource banditrlproof.tsallis.armdependentsuboptimalrewardboostsource a stationary but arm-dependent corruption process that leaves the best arm unchanged and adds a prescribed nonnegative boost to every other arm. definition compiled","shard":"modules/68d103113a899997.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_armDependentSuboptimalRewardBoostSource","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_armDependentSuboptimalRewardBoostSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_armDependentSuboptimalRewardBoostSource","description":"The arm-dependent suboptimal boost has exact deterministic envelope budget `(T+1) * sum_(a != best) boost(a)`.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html#decl-9aae03515020","parent":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","order":9755,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_armDependentSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Fin K -> Real) (hboost : forall arm, 0 <= boost arm) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (armDependentSuboptimalRewardBoostSource model boost hboost) = (((horizon + 1 : Nat) : Real)) * ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum boost","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_armdependentsuboptimalrewardboostsource banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_armdependentsuboptimalrewardboostsource the arm-dependent suboptimal boost has exact deterministic envelope budget `(t+1) * sum_(a != best) boost(a)`. theorem compiled","shard":"modules/68d103113a899997.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDArmDependentSuboptimalBoostRefinedRegime","label":"finiteArmIIDArmDependentSuboptimalBoostRefinedRegime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDArmDependentSuboptimalBoostRefinedRegime","description":"The coefficient-aware refined regime for an arm-dependent boost budget.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html#decl-2eae7468df3b","parent":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","order":9756,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean:65"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDArmDependentSuboptimalBoostRefinedRegime {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Fin K -> Real) : Prop","missing":[],"search":"finitearmiidarmdependentsuboptimalboostrefinedregime banditrlproof.tsallis.finitearmiidarmdependentsuboptimalboostrefinedregime the coefficient-aware refined regime for an arm-dependent boost budget. definition compiled","shard":"modules/68d103113a899997.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_armDependentSuboptimalRewardBoostSource_of_refinedRegime","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_armDependentSuboptimalRewardBoostSource_of_refinedRegime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_armDependentSuboptimalRewardBoostSource_of_refinedRegime","description":"The named arm-dependent refined regime supplies the existing model-facing compact corruption window after the exact budget rewrite.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html#decl-fee0316675fe","parent":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","order":9757,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean:85"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_armDependentSuboptimalRewardBoostSource_of_refinedRegime {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Fin K -> Real) (hboost : forall arm, 0 <= boost arm) (hregime : finiteArmIIDArmDependentSuboptimalBoostRefinedRegime model horizon boost) : finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow model horizon (armDependentSuboptimalRewardBoostSource model boost hboost)","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow_armdependentsuboptimalrewardboostsource_of_refinedregime banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow_armdependentsuboptimalrewardboostsource_of_refinedregime the named arm-dependent refined regime supplies the existing model-facing compact corruption window after the exact budget rewrite. theorem compiled","shard":"modules/68d103113a899997.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDArmDependentSuboptimalBoostAllRegimeBound","label":"finiteArmIIDArmDependentSuboptimalBoostAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDArmDependentSuboptimalBoostAllRegimeBound","description":"Total regret bound for arm-dependent boosts: the refined local expression inside its named regime and the logarithmic additive-budget expression on the complement.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html#decl-e9999763e574","parent":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","order":9758,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean:99"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDArmDependentSuboptimalBoostAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Fin K -> Real) : Real","missing":[],"search":"finitearmiidarmdependentsuboptimalboostallregimebound banditrlproof.tsallis.finitearmiidarmdependentsuboptimalboostallregimebound total regret bound for arm-dependent boosts: the refined local expression inside its named regime and the logarithmic additive-budget expression on the complement. definition compiled","shard":"modules/68d103113a899997.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDArmDependentSuboptimalBoostRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDArmDependentSuboptimalBoostRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDArmDependentSuboptimalBoostRewardLawRegret_le_allRegimes","description":"Scheduled half-Tsallis regret for every nonnegative arm-dependent suboptimal reward boost and every finite horizon. The result automatically selects the refined or logarithmic branch and requires no caller window proof.","url":"../modules/banditrlproof-tsallisfinitearmiidarmdependentsuboptimalboostregret/index.html#decl-56b685b99035","parent":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","order":9759,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret.lean:123"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDArmDependentSuboptimalBoostRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (boost : Fin K -> Real) (hboost : forall arm, 0 <= boost arm) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidarmdependentsuboptimalboostrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidarmdependentsuboptimalboostrewardlawregret_le_allregimes scheduled half-tsallis regret for every nonnegative arm-dependent suboptimal reward boost and every finite horizon. the result automatically selects the refined or logarithmic branch and requires no caller window proof. theorem compiled","shard":"modules/68d103113a899997.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReal","label":"clippedUnitReal","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReal","description":"Projection of a real value to the unit interval.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-b0003f64c504","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9760,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedUnitReal (value : Real) : Real","missing":[],"search":"clippedunitreal banditrlproof.tsallis.clippedunitreal projection of a real value to the unit interval. definition compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReal_mem_Icc","label":"clippedUnitReal_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReal_mem_Icc","description":"theorem clippedUnitReal_mem_Icc (value : Real) : clippedUnitReal value ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-c42b38e687f2","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9761,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:26"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem clippedUnitReal_mem_Icc (value : Real) : clippedUnitReal value ∈ Set.Icc (0 : Real) 1","missing":[],"search":"clippedunitreal_mem_icc banditrlproof.tsallis.clippedunitreal_mem_icc theorem clippedunitreal_mem_icc (value : real) : clippedunitreal value ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReal_eq_of_mem_Icc","label":"clippedUnitReal_eq_of_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReal_eq_of_mem_Icc","description":"theorem clippedUnitReal_eq_of_mem_Icc (value : Real) (hvalue : value ∈ Set.Icc (0 : Real) 1) : clippedUnitReal value = value","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-630d97a510bb","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9762,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:30"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem clippedUnitReal_eq_of_mem_Icc (value : Real) (hvalue : value ∈ Set.Icc (0 : Real) 1) : clippedUnitReal value = value","missing":[],"search":"clippedunitreal_eq_of_mem_icc banditrlproof.tsallis.clippedunitreal_eq_of_mem_icc theorem clippedunitreal_eq_of_mem_icc (value : real) (hvalue : value ∈ set.icc (0 : real) 1) : clippedunitreal value = value theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_clippedUnitReal_add_sub_self_le","label":"abs_clippedUnitReal_add_sub_self_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_clippedUnitReal_add_sub_self_le","description":"theorem abs_clippedUnitReal_add_sub_self_le (value shift : Real) (hvalue : value ∈ Set.Icc (0 : Real) 1) : |clippedUnitReal (value + shift) - value| <= |shift|","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-8275f5f6e65d","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9763,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:36"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_clippedUnitReal_add_sub_self_le (value shift : Real) (hvalue : value ∈ Set.Icc (0 : Real) 1) : |clippedUnitReal (value + shift) - value| <= |shift|","missing":[],"search":"abs_clippedunitreal_add_sub_self_le banditrlproof.tsallis.abs_clippedunitreal_add_sub_self_le theorem abs_clippedunitreal_add_sub_self_le (value shift : real) (hvalue : value ∈ set.icc (0 : real) 1) : |clippedunitreal (value + shift) - value| <= |shift| theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedReward","label":"finiteArmIIDStationaryCorruptedReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedReward","description":"Reward after a stationary arm-dependent shift and unit-interval clipping.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-aa0d10eab797","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9764,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:51"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDStationaryCorruptedReward {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : Real","missing":[],"search":"finitearmiidstationarycorruptedreward banditrlproof.tsallis.finitearmiidstationarycorruptedreward reward after a stationary arm-dependent shift and unit-interval clipping. definition compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss","label":"finiteArmIIDStationaryCorruptedRewardVectorLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss","description":"Selected loss associated with the stationary corrupted reward vector.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-998efd00e88c","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9765,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:56"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDStationaryCorruptedRewardVectorLoss {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : Real","missing":[],"search":"finitearmiidstationarycorruptedrewardvectorloss banditrlproof.tsallis.finitearmiidstationarycorruptedrewardvectorloss selected loss associated with the stationary corrupted reward vector. definition compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDStationaryCorruptedRewardVectorLoss","label":"measurable_finiteArmIIDStationaryCorruptedRewardVectorLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_finiteArmIIDStationaryCorruptedRewardVectorLoss","description":"theorem measurable_finiteArmIIDStationaryCorruptedRewardVectorLoss {K : Nat} (rewardShift : Fin K -> Real) : Measurable (fun input : (Fin K -> Rat) × Fin K => finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift input.1 input.2)","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-56b6970d21d3","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9766,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:60"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteArmIIDStationaryCorruptedRewardVectorLoss {K : Nat} (rewardShift : Fin K -> Real) : Measurable (fun input : (Fin K -> Rat) × Fin K => finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift input.1 input.2)","missing":[],"search":"measurable_finitearmiidstationarycorruptedrewardvectorloss banditrlproof.tsallis.measurable_finitearmiidstationarycorruptedrewardvectorloss theorem measurable_finitearmiidstationarycorruptedrewardvectorloss {k : nat} (rewardshift : fin k -> real) : measurable (fun input : (fin k -> rat) × fin k => finitearmiidstationarycorruptedrewardvectorloss rewardshift input.1 input.2) theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss_nonneg","label":"finiteArmIIDStationaryCorruptedRewardVectorLoss_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss_nonneg","description":"theorem finiteArmIIDStationaryCorruptedRewardVectorLoss_nonneg {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : 0 <= finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state arm","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-a35063e6bf4d","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9767,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:67"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDStationaryCorruptedRewardVectorLoss_nonneg {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : 0 <= finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state arm","missing":[],"search":"finitearmiidstationarycorruptedrewardvectorloss_nonneg banditrlproof.tsallis.finitearmiidstationarycorruptedrewardvectorloss_nonneg theorem finitearmiidstationarycorruptedrewardvectorloss_nonneg {k : nat} (rewardshift : fin k -> real) (state : fin k -> rat) (arm : fin k) : 0 <= finitearmiidstationarycorruptedrewardvectorloss rewardshift state arm theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss_le_one","label":"finiteArmIIDStationaryCorruptedRewardVectorLoss_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss_le_one","description":"theorem finiteArmIIDStationaryCorruptedRewardVectorLoss_le_one {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state arm <= 1","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-1564fb72c08d","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9768,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:77"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDStationaryCorruptedRewardVectorLoss_le_one {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state arm <= 1","missing":[],"search":"finitearmiidstationarycorruptedrewardvectorloss_le_one banditrlproof.tsallis.finitearmiidstationarycorruptedrewardvectorloss_le_one theorem finitearmiidstationarycorruptedrewardvectorloss_le_one {k : nat} (rewardshift : fin k -> real) (state : fin k -> rat) (arm : fin k) : finitearmiidstationarycorruptedrewardvectorloss rewardshift state arm <= 1 theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_finiteArmIIDStationaryCorruptedReward_sub_base_le","label":"abs_finiteArmIIDStationaryCorruptedReward_sub_base_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_finiteArmIIDStationaryCorruptedReward_sub_base_le","description":"theorem abs_finiteArmIIDStationaryCorruptedReward_sub_base_le {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : |finiteArmIIDStationaryCorruptedReward rewardShift state arm - clippedUnitReward (state arm)| <= |rewardShift arm|","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-07d90a64d796","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9769,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:86"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_finiteArmIIDStationaryCorruptedReward_sub_base_le {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (arm : Fin K) : |finiteArmIIDStationaryCorruptedReward rewardShift state arm - clippedUnitReward (state arm)| <= |rewardShift arm|","missing":[],"search":"abs_finitearmiidstationarycorruptedreward_sub_base_le banditrlproof.tsallis.abs_finitearmiidstationarycorruptedreward_sub_base_le theorem abs_finitearmiidstationarycorruptedreward_sub_base_le {k : nat} (rewardshift : fin k -> real) (state : fin k -> rat) (arm : fin k) : |finitearmiidstationarycorruptedreward rewardshift state arm - clippedunitreward (state arm)| <= |rewardshift arm| theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_stationaryCorruptedLossDiff_sub_baseLossDiff_le","label":"abs_stationaryCorruptedLossDiff_sub_baseLossDiff_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_stationaryCorruptedLossDiff_sub_baseLossDiff_le","description":"theorem abs_stationaryCorruptedLossDiff_sub_baseLossDiff_le {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (best arm : Fin K) : |(finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state arm - finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state best) - (finiteArmIIDRewardVectorLoss state arm - finiteArmIIDRewardVectorLoss state best)| <= |rewardShift arm| + |rewardShift best|","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-16fc6cabd579","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9770,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:94"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_stationaryCorruptedLossDiff_sub_baseLossDiff_le {K : Nat} (rewardShift : Fin K -> Real) (state : Fin K -> Rat) (best arm : Fin K) : |(finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state arm - finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift state best) - (finiteArmIIDRewardVectorLoss state arm - finiteArmIIDRewardVectorLoss state best)| <= |rewardShift arm| + |rewardShift best|","missing":[],"search":"abs_stationarycorruptedlossdiff_sub_baselossdiff_le banditrlproof.tsallis.abs_stationarycorruptedlossdiff_sub_baselossdiff_le theorem abs_stationarycorruptedlossdiff_sub_baselossdiff_le {k : nat} (rewardshift : fin k -> real) (state : fin k -> rat) (best arm : fin k) : |(finitearmiidstationarycorruptedrewardvectorloss rewardshift state arm - finitearmiidstationarycorruptedrewardvectorloss rewardshift state best) - (finitearmiidrewardvectorloss state arm - finitearmiidrewardvectorloss state best)| <= |rewardshift arm| + |rewardshift best| theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_iidLossStateDiff","label":"integrable_iidLossStateDiff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_iidLossStateDiff","description":"theorem integrable_iidLossStateDiff {LossState Action : Type*} [MeasurableSpace LossState] [MeasurableSpace Action] (law : Measure LossState) [IsFiniteMeasure law] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : forall state action, 0 <= value state action) (hvalue_le_one : forall state action, value state action <= 1) (best arm :…","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-533d14d5a18c","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9771,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_iidLossStateDiff {LossState Action : Type*} [MeasurableSpace LossState] [MeasurableSpace Action] (law : Measure LossState) [IsFiniteMeasure law] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : forall state action, 0 <= value state action) (hvalue_le_one : forall state action, value state action <= 1) (best arm : Action) : Integrable (fun state => value state arm - value state best) law","missing":[],"search":"integrable_iidlossstatediff banditrlproof.tsallis.integrable_iidlossstatediff theorem integrable_iidlossstatediff {lossstate action : type*} [measurablespace lossstate] [measurablespace action] (law : measure lossstate) [isfinitemeasure law] (value : lossstate -> action -> real) (hvalue : measurable (fun input : lossstate × action => value input.1 input.2)) (hvalue_nonneg : forall state action, 0 <= value state action) (hvalue_le_one : forall state action, value state action <= 1) (best arm : action) : integrable (fun state => value state arm - value state best) law theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_iidLossStateMeanGap_stationaryCorrupted_sub_modelGap_le","label":"abs_iidLossStateMeanGap_stationaryCorrupted_sub_modelGap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_iidLossStateMeanGap_stationaryCorrupted_sub_modelGap_le","description":"The actual stationary-corrupted IID loss gap remains within the two affected arm shifts of the uncorrupted finite-bandit model gap.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-2f24daea7df4","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9772,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:149"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_iidLossStateMeanGap_stationaryCorrupted_sub_modelGap_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (rewardShift : Fin K -> Real) (arm : Fin K) : |iidLossStateMeanGap (finiteArmIIDRewardVectorLaw armLaw) (finiteArmIIDStationaryCorruptedRewardVectorLoss rewardShift) model.bestArm arm - ((model.gap arm : Rat) : Real)| <= |rewardShift arm| + |rewardShift model.bestArm|","missing":[],"search":"abs_iidlossstatemeangap_stationarycorrupted_sub_modelgap_le banditrlproof.tsallis.abs_iidlossstatemeangap_stationarycorrupted_sub_modelgap_le the actual stationary-corrupted iid loss gap remains within the two affected arm shifts of the uncorrupted finite-bandit model gap. theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.scheduledGapDeviationBudget","label":"scheduledGapDeviationBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.scheduledGapDeviationBudget","description":"Total baseline-gap perturbation allowance through the inclusive horizon.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-646d415acc8f","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9773,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:205"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def scheduledGapDeviationBudget {Action : Type*} [DecidableEq Action] (arms : Finset Action) (best : Action) (horizon : Nat) (deviation : Action -> Real) : Real","missing":[],"search":"scheduledgapdeviationbudget banditrlproof.tsallis.scheduledgapdeviationbudget total baseline-gap perturbation allowance through the inclusive horizon. definition compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.scheduledGapDeviationBudget_eq","label":"scheduledGapDeviationBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.scheduledGapDeviationBudget_eq","description":"theorem scheduledGapDeviationBudget_eq {Action : Type*} [DecidableEq Action] (arms : Finset Action) (best : Action) (horizon : Nat) (deviation : Action -> Real) : scheduledGapDeviationBudget arms best horizon deviation = (((horizon + 1 : Nat) : Real)) * (arms.erase best).sum deviation","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-dd9c457836a7","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9774,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:212"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem scheduledGapDeviationBudget_eq {Action : Type*} [DecidableEq Action] (arms : Finset Action) (best : Action) (horizon : Nat) (deviation : Action -> Real) : scheduledGapDeviationBudget arms best horizon deviation = (((horizon + 1 : Nat) : Real)) * (arms.erase best).sum deviation","missing":[],"search":"scheduledgapdeviationbudget_eq banditrlproof.tsallis.scheduledgapdeviationbudget_eq theorem scheduledgapdeviationbudget_eq {action : type*} [decidableeq action] (arms : finset action) (best : action) (horizon : nat) (deviation : action -> real) : scheduledgapdeviationbudget arms best horizon deviation = (((horizon + 1 : nat) : real)) * (arms.erase best).sum deviation theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_perturbedExpectedGapLaw","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_perturbedExpectedGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_perturbedExpectedGapLaw","description":"An expected law for perturbed gaps yields the baseline-gap self-bound with the accumulated coordinatewise gap-deviation budget.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-cad216d3b7fa","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9775,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:222"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_perturbedExpectedGapLaw {Env Action : Type*} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (baseGap actualGap deviation : Action -> Real) (hactualGapLaw : HasScheduledExpectedGapLaw mu arms harms eta loss best actualGap horizon) (hdeviation : forall action, action ∈ arms.erase best -> |actualGap action - baseGap action| <= deviation action) : (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => baseGap action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_perturbedexpectedgaplaw banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_perturbedexpectedgaplaw an expected law for perturbed gaps yields the baseline-gap self-bound with the accumulated coordinatewise gap-deviation budget. theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget","label":"finiteArmIIDStationaryRewardCorruptionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget","description":"Explicit corruption budget for a stationary arm-dependent reward shift.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-0ba018d70ed3","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9776,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:301"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDStationaryRewardCorruptionBudget {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Fin K -> Real) : Real","missing":[],"search":"finitearmiidstationaryrewardcorruptionbudget banditrlproof.tsallis.finitearmiidstationaryrewardcorruptionbudget explicit corruption budget for a stationary arm-dependent reward shift. definition compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget_eq","label":"finiteArmIIDStationaryRewardCorruptionBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget_eq","description":"theorem finiteArmIIDStationaryRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Fin K -> Real) : finiteArmIIDStationaryRewardCorruptionBudget model horizon rewardShift = (((horizon + 1 : Nat) : Real)) * ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => |rewardShift arm| + |rewardShift model.bestArm|)","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-365bb175f1ea","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9777,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:307"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDStationaryRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Fin K -> Real) : finiteArmIIDStationaryRewardCorruptionBudget model horizon rewardShift = (((horizon + 1 : Nat) : Real)) * ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => |rewardShift arm| + |rewardShift model.bestArm|)","missing":[],"search":"finitearmiidstationaryrewardcorruptionbudget_eq banditrlproof.tsallis.finitearmiidstationaryrewardcorruptionbudget_eq theorem finitearmiidstationaryrewardcorruptionbudget_eq {k : nat} (model : finitebanditmodel k) (horizon : nat) (rewardshift : fin k -> real) : finitearmiidstationaryrewardcorruptionbudget model horizon rewardshift = (((horizon + 1 : nat) : real)) * ((finset.univ : finset (fin k)).erase model.bestarm).sum (fun arm => |rewardshift arm| + |rewardshift model.bestarm|) theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget_zero","label":"finiteArmIIDStationaryRewardCorruptionBudget_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget_zero","description":"theorem finiteArmIIDStationaryRewardCorruptionBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIIDStationaryRewardCorruptionBudget model horizon (fun _ => 0) = 0","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-27b65c01a1a5","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9778,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:317"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDStationaryRewardCorruptionBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIIDStationaryRewardCorruptionBudget model horizon (fun _ => 0) = 0","missing":[],"search":"finitearmiidstationaryrewardcorruptionbudget_zero banditrlproof.tsallis.finitearmiidstationaryrewardcorruptionbudget_zero theorem finitearmiidstationaryrewardcorruptionbudget_zero {k : nat} (model : finitebanditmodel k) (horizon : nat) : finitearmiidstationaryrewardcorruptionbudget model horizon (fun _ => 0) = 0 theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDStationaryCorruptedRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDStationaryCorruptedRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDStationaryCorruptedRewardLawRegret_le_log","description":"Scheduled half-Tsallis logarithmic regret for an IID finite-arm reward model with a fixed clipped reward shift. The additive corruption term is derived from `rewardShift`; it is not a free theorem parameter.","url":"../modules/banditrlproof-tsallisfinitearmiidcorruptedrewardlaw/index.html#decl-69e56c13a3fd","parent":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","order":9779,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDCorruptedRewardLaw.lean:327"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDStationaryCorruptedRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (rewardShift : Fin K -> Real) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidstationarycorruptedrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidstationarycorruptedrewardlawregret_le_log scheduled half-tsallis logarithmic regret for an iid finite-arm reward model with a fixed clipped reward shift. the additive corruption term is derived from `rewardshift`; it is not a free theorem parameter. theorem compiled","shard":"modules/a8370caa8eaa971c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHistoryAdaptiveRewardShiftSource","label":"FiniteArmIIDHistoryAdaptiveRewardShiftSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHistoryAdaptiveRewardShiftSource","description":"A predictable reward-shift process and its deterministic envelope.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-4a3d21796e6e","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9780,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure FiniteArmIIDHistoryAdaptiveRewardShiftSource (K : Nat) where","missing":[],"search":"finitearmiidhistoryadaptiverewardshiftsource banditrlproof.tsallis.finitearmiidhistoryadaptiverewardshiftsource a predictable reward-shift process and its deterministic envelope. structure compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_clippedUnitReal","label":"measurable_clippedUnitReal","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_clippedUnitReal","description":"Projection to the real unit interval is measurable.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-a09ad7e03a9a","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9781,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_clippedUnitReal : Measurable clippedUnitReal","missing":[],"search":"measurable_clippedunitreal banditrlproof.tsallis.measurable_clippedunitreal projection to the real unit interval is measurable. theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_initial","label":"measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_initial","description":"theorem measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_initial {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : Measurable (fun input : (Nat -> (Fin K -> Rat)) × Fin K => finiteArmIIDStationaryCorruptedRewardVectorLoss source.initial (input.1 0) input.2)","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-a89347029d52","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9782,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_initial {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : Measurable (fun input : (Nat -> (Fin K -> Rat)) × Fin K => finiteArmIIDStationaryCorruptedRewardVectorLoss source.initial (input.1 0) input.2)","missing":[],"search":"measurable_finitearmiidhistoryadaptivecorruptedrewardloss_initial banditrlproof.tsallis.measurable_finitearmiidhistoryadaptivecorruptedrewardloss_initial theorem measurable_finitearmiidhistoryadaptivecorruptedrewardloss_initial {k : nat} (source : finitearmiidhistoryadaptiverewardshiftsource k) : measurable (fun input : (nat -> (fin k -> rat)) × fin k => finitearmiidstationarycorruptedrewardvectorloss source.initial (input.1 0) input.2) theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_successor","label":"measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_successor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_successor","description":"theorem measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_successor {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (n : Nat) : Measurable (fun input : (Nat -> (Fin K -> Rat)) × (History.FinitePairHistory (Fin K) Real n × Fin K) => finiteArmIIDStationaryCorruptedRewardVectorLoss (source.successor n input.2.1) (input.1 (n + 1)) input.2.2)","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-87f5b9e40cd2","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9783,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:66"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_successor {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (n : Nat) : Measurable (fun input : (Nat -> (Fin K -> Rat)) × (History.FinitePairHistory (Fin K) Real n × Fin K) => finiteArmIIDStationaryCorruptedRewardVectorLoss (source.successor n input.2.1) (input.1 (n + 1)) input.2.2)","missing":[],"search":"measurable_finitearmiidhistoryadaptivecorruptedrewardloss_successor banditrlproof.tsallis.measurable_finitearmiidhistoryadaptivecorruptedrewardloss_successor theorem measurable_finitearmiidhistoryadaptivecorruptedrewardloss_successor {k : nat} (source : finitearmiidhistoryadaptiverewardshiftsource k) (n : nat) : measurable (fun input : (nat -> (fin k -> rat)) × (history.finitepairhistory (fin k) real n × fin k) => finitearmiidstationarycorruptedrewardvectorloss (source.successor n input.2.1) (input.1 (n + 1)) input.2.2) theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveCorruptedRewardLoss","label":"finiteArmIIDHistoryAdaptiveCorruptedRewardLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveCorruptedRewardLoss","description":"Predictable clipped loss generated by a history-adaptive reward shift.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-3b43e9d37150","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9784,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:99"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveCorruptedRewardLoss {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : Exp3.PredictableLossVector (Nat -> (Fin K -> Rat)) (Fin K) where","missing":[],"search":"finitearmiidhistoryadaptivecorruptedrewardloss banditrlproof.tsallis.finitearmiidhistoryadaptivecorruptedrewardloss predictable clipped loss generated by a history-adaptive reward shift. definition compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_hasIIDStateCoordinateLocality","label":"finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_hasIIDStateCoordinateLocality","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_hasIIDStateCoordinateLocality","description":"The concrete history-adaptive loss reads no future IID state coordinate.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-bae7daf635e6","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9785,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_hasIIDStateCoordinateLocality {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : HasIIDStateCoordinateLocality (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source)","missing":[],"search":"finitearmiidhistoryadaptivecorruptedrewardloss_hasiidstatecoordinatelocality banditrlproof.tsallis.finitearmiidhistoryadaptivecorruptedrewardloss_hasiidstatecoordinatelocality the concrete history-adaptive loss reads no future iid state coordinate. theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_zero","label":"predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_zero","description":"theorem predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_zero {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (sample : (Nat -> (Fin K -> Rat)) × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) 0 sample arm = finiteArmIIDStationaryCorruptedRewardVectorLoss source.initial (sample.1 0) arm","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-54c7eed46c2a","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9786,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:139"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_zero {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (sample : (Nat -> (Fin K -> Rat)) × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) 0 sample arm = finiteArmIIDStationaryCorruptedRewardVectorLoss source.initial (sample.1 0) arm","missing":[],"search":"predictablelossat_finitearmiidhistoryadaptivecorruptedrewardloss_zero banditrlproof.tsallis.predictablelossat_finitearmiidhistoryadaptivecorruptedrewardloss_zero theorem predictablelossat_finitearmiidhistoryadaptivecorruptedrewardloss_zero {k : nat} (source : finitearmiidhistoryadaptiverewardshiftsource k) (sample : (nat -> (fin k -> rat)) × ((k : nat) -> fin k × real)) (arm : fin k) : exp3.predictablelossat (finitearmiidhistoryadaptivecorruptedrewardloss source) 0 sample arm = finitearmiidstationarycorruptedrewardvectorloss source.initial (sample.1 0) arm theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_succ","label":"predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_succ","description":"theorem predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_succ {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (n : Nat) (sample : (Nat -> (Fin K -> Rat)) × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) (n + 1) sample arm = finiteArmIIDStationaryCorruptedRewardVectorLoss (source.successor n (Preorder.fres…","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-6a47854e0a4f","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9787,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:150"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_succ {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (n : Nat) (sample : (Nat -> (Fin K -> Rat)) × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) (n + 1) sample arm = finiteArmIIDStationaryCorruptedRewardVectorLoss (source.successor n (Preorder.frestrictLe n sample.2)) (sample.1 (n + 1)) arm","missing":[],"search":"predictablelossat_finitearmiidhistoryadaptivecorruptedrewardloss_succ banditrlproof.tsallis.predictablelossat_finitearmiidhistoryadaptivecorruptedrewardloss_succ theorem predictablelossat_finitearmiidhistoryadaptivecorruptedrewardloss_succ {k : nat} (source : finitearmiidhistoryadaptiverewardshiftsource k) (n : nat) (sample : (nat -> (fin k -> rat)) × ((k : nat) -> fin k × real)) (arm : fin k) : exp3.predictablelossat (finitearmiidhistoryadaptivecorruptedrewardloss source) (n + 1) sample arm = finitearmiidstationarycorruptedrewardvectorloss (source.successor n (preorder.frestrictle n sample.2)) (sample.1 (n + 1)) arm theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le","label":"abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le","description":"Pointwise actual/reference gap deviation under the source envelope.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-ef3a31bb0d16","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9788,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:163"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (t : Nat) (sample : (Nat -> (Fin K -> Rat)) × ((k : Nat) -> Fin K × Real)) (best arm : Fin K) : |(Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) t sample arm - Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) t sample best) - (Exp3.predictableLossAt (iidLossStatePredictableLossVector finiteArmIIDRewardVectorLoss measurable_finiteArmIIDRewardVectorLoss finiteArmIIDRewardVectorLoss_nonneg finiteArmIIDRewardVectorLoss_le_one) t sample arm - Exp3.predictableLossAt (iidLossStatePredictableLossVector finiteArmIIDRewardVectorLoss measurable_finiteArmIIDRewardVectorLoss finiteArmIIDRewardVectorLoss_nonneg finiteArmIIDRewardVectorLoss_le_one) t sample best)| <= source.envelope t arm + source.e…","missing":[],"search":"abs_historyadaptivecorruptedpredictablelossdiff_sub_baselossdiff_le banditrlproof.tsallis.abs_historyadaptivecorruptedpredictablelossdiff_sub_baselossdiff_le pointwise actual/reference gap deviation under the source envelope. theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget","description":"Explicit deterministic envelope budget for predictable corruption.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-f4da3a6bdd71","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9789,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:217"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveRewardCorruptionBudget {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : Real","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget explicit deterministic envelope budget for predictable corruption. definition compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_eq","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_eq","description":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon source = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => source.envelope t arm + source.envelope t model.bestArm))","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-96bda8673518","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9790,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:224"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon source = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => source.envelope t arm + source.envelope t model.bestArm))","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_eq banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_eq theorem finitearmiidhistoryadaptiverewardcorruptionbudget_eq {k : nat} (model : finitebanditmodel k) (horizon : nat) (source : finitearmiidhistoryadaptiverewardshiftsource k) : finitearmiidhistoryadaptiverewardcorruptionbudget model horizon source = (finset.range (horizon + 1)).sum (fun t => ((finset.univ : finset (fin k)).erase model.bestarm).sum (fun arm => source.envelope t arm + source.envelope t model.bestarm)) theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource","label":"zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource","description":"The uncorrupted predictable shift source.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-98d44f07d933","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9791,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:235"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource (K : Nat) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"zerofinitearmiidhistoryadaptiverewardshiftsource banditrlproof.tsallis.zerofinitearmiidhistoryadaptiverewardshiftsource the uncorrupted predictable shift source. definition compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_zero","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_zero","description":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource K) = 0","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-c20f771776cc","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9792,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:246"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource K) = 0","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_zero banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_zero theorem finitearmiidhistoryadaptiverewardcorruptionbudget_zero {k : nat} (model : finitebanditmodel k) (horizon : nat) : finitearmiidhistoryadaptiverewardcorruptionbudget model horizon (zerofinitearmiidhistoryadaptiverewardshiftsource k) = 0 theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveTrajectoryKernel","label":"hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveTrajectoryKernel","description":"The actual history-adaptive canonical trajectory has the IID-prefix factorization needed to reuse the uncorrupted reference gap law.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-c3211d910382","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9793,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:256"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveTrajectoryKernel {K : Nat} [Nonempty (Fin K)] (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (arms : Finset (Fin K)) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (horizon : Nat) : HasScheduledIIDPrefixKernelFactorization (sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source).environment) horizon","missing":[],"search":"hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallisfinitearmiidhistoryadaptivetrajectorykernel banditrlproof.tsallis.hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallisfinitearmiidhistoryadaptivetrajectorykernel the actual history-adaptive canonical trajectory has the iid-prefix factorization needed to reuse the uncorrupted reference gap law. theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_log","description":"Scheduled half-Tsallis logarithmic regret under a measurable predictable history-adaptive reward shift. The additive allowance is the deterministic envelope budget packaged by `source`, not a free theorem parameter.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptivecorruptedrewardlaw/index.html#decl-2749fc06ca4e","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","order":9794,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw.lean:278"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlawregret_le_log scheduled half-tsallis logarithmic regret under a measurable predictable history-adaptive reward shift. the additive allowance is the deterministic envelope budget packaged by `source`, not a free theorem parameter. theorem compiled","shard":"modules/41dd857f9f56104b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardShiftAt","label":"finiteArmIIDHistoryAdaptiveRewardShiftAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardShiftAt","description":"The reward shift selected from the finite observed pair history available before round `t`.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-80fd877bf959","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9795,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:12"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveRewardShiftAt {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (t : Nat) (trajectory : (k : Nat) -> Fin K × Real) (arm : Fin K) : Real","missing":[],"search":"finitearmiidhistoryadaptiverewardshiftat banditrlproof.tsallis.finitearmiidhistoryadaptiverewardshiftat the reward shift selected from the finite observed pair history available before round `t`. definition compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveRewardShiftAt","label":"measurable_finiteArmIIDHistoryAdaptiveRewardShiftAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveRewardShiftAt","description":"The realized predictable shift of a fixed arm is measurable on the full trajectory space.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-141613ad9b0c","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9796,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteArmIIDHistoryAdaptiveRewardShiftAt {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (t : Nat) (arm : Fin K) : Measurable (fun trajectory : (k : Nat) -> Fin K × Real => finiteArmIIDHistoryAdaptiveRewardShiftAt source t trajectory arm)","missing":[],"search":"measurable_finitearmiidhistoryadaptiverewardshiftat banditrlproof.tsallis.measurable_finitearmiidhistoryadaptiverewardshiftat the realized predictable shift of a fixed arm is measurable on the full trajectory space. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_finiteArmIIDHistoryAdaptiveRewardShiftAt_le","label":"abs_finiteArmIIDHistoryAdaptiveRewardShiftAt_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_finiteArmIIDHistoryAdaptiveRewardShiftAt_le","description":"The realized shift retains the deterministic envelope supplied by the history-adaptive source.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-f51ba312dbbc","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9797,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_finiteArmIIDHistoryAdaptiveRewardShiftAt_le {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (t : Nat) (trajectory : (k : Nat) -> Fin K × Real) (arm : Fin K) : |finiteArmIIDHistoryAdaptiveRewardShiftAt source t trajectory arm| <= source.envelope t arm","missing":[],"search":"abs_finitearmiidhistoryadaptiverewardshiftat_le banditrlproof.tsallis.abs_finitearmiidhistoryadaptiverewardshiftat_le the realized shift retains the deterministic envelope supplied by the history-adaptive source. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le_actualShift","label":"abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le_actualShift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le_actualShift","description":"The actual/reference predictable loss-gap deviation is controlled by the two shifts realized on the observed history, before replacing them by their deterministic envelopes.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-4dd78432eb5d","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9798,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:52"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le_actualShift {K : Nat} (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (t : Nat) (sample : (Nat -> (Fin K -> Rat)) × ((k : Nat) -> Fin K × Real)) (best arm : Fin K) : |(Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) t sample arm - Exp3.predictableLossAt (finiteArmIIDHistoryAdaptiveCorruptedRewardLoss source) t sample best) - (Exp3.predictableLossAt (iidLossStatePredictableLossVector finiteArmIIDRewardVectorLoss measurable_finiteArmIIDRewardVectorLoss finiteArmIIDRewardVectorLoss_nonneg finiteArmIIDRewardVectorLoss_le_one) t sample arm - Exp3.predictableLossAt (iidLossStatePredictableLossVector finiteArmIIDRewardVectorLoss measurable_finiteArmIIDRewardVectorLoss finiteArmIIDRewardVectorLoss_nonneg finiteArmIIDRewardVectorLoss_le_one) t sample best)| <= |finiteArmIIDHistory…","missing":[],"search":"abs_historyadaptivecorruptedpredictablelossdiff_sub_baselossdiff_le_actualshift banditrlproof.tsallis.abs_historyadaptivecorruptedpredictablelossdiff_sub_baselossdiff_le_actualshift the actual/reference predictable loss-gap deviation is controlled by the two shifts realized on the observed history, before replacing them by their deterministic envelopes. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbability_mul_historyAdaptiveRewardShiftDeviation","label":"integrable_sampledScheduledHalfTsallisProbability_mul_historyAdaptiveRewardShiftDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbability_mul_historyAdaptiveRewardShiftDeviation","description":"Probability-weighted realized history-adaptive deviation is integrable under every finite trajectory measure.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-a0cd8b97b964","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9799,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:95"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisProbability_mul_historyAdaptiveRewardShiftDeviation {Env : Type*} {K : Nat} [MeasurableSpace Env] (mu : Measure (Env × ((k : Nat) -> Fin K × Real))) [IsFiniteMeasure mu] (arms : Finset (Fin K)) (harms : arms.Nonempty) (eta : Nat -> Real) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (t : Nat) (best arm : Fin K) (harm : arm ∈ arms) : Integrable (fun sample => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample arm * (|finiteArmIIDHistoryAdaptiveRewardShiftAt source t sample.2 arm| + |finiteArmIIDHistoryAdaptiveRewardShiftAt source t sample.2 best|)) mu","missing":[],"search":"integrable_sampledscheduledhalftsallisprobability_mul_historyadaptiverewardshiftdeviation banditrlproof.tsallis.integrable_sampledscheduledhalftsallisprobability_mul_historyadaptiverewardshiftdeviation probability-weighted realized history-adaptive deviation is integrable under every finite trajectory measure. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget","label":"finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget","description":"Expected corruption weighted by the generated policy's conditional selection probability for each affected suboptimal arm, using the realized history-adaptive shifts.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-6217ec446c00","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9800,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:158"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (mu : Measure ((Nat -> Fin K -> Rat) × ((k : Nat) -> Fin K × Real))) : Real","missing":[],"search":"finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget expected corruption weighted by the generated policy's conditional selection probability for each affected suboptimal arm, using the realized history-adaptive shifts. definition compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_eq","label":"finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_eq","description":"theorem finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (mu : Measure ((Nat -> Fin K -> Rat) × ((k : Nat) -> Fin K × Real))) : finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget model horizon source mu = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase…","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-7e6d72e08833","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9801,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:175"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (mu : Measure ((Nat -> Fin K -> Rat) × ((k : Nat) -> Fin K × Real))) : finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget model horizon source mu = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => integral mu (fun sample => sampledScheduledHalfTsallisProbabilityAtTime (Finset.univ : Finset (Fin K)) ⟨model.bestArm, Finset.mem_univ model.bestArm⟩ sampledScheduledHalfTsallisSqrtSchedule t sample arm * (|finiteArmIIDHistoryAdaptiveRewardShiftAt source t sample.2 arm| + |finiteArmIIDHistoryAdaptiveRewardShiftAt source t sample.2 model.bestArm|))))","missing":[],"search":"finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_eq banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_eq theorem finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_eq {k : nat} (model : finitebanditmodel k) (horizon : nat) (source : finitearmiidhistoryadaptiverewardshiftsource k) (mu : measure ((nat -> fin k -> rat) × ((k : nat) -> fin k × real))) : finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget model horizon source mu = (finset.range (horizon + 1)).sum (fun t => ((finset.univ : finset (fin k)).erase model.bestarm).sum (fun arm => integral mu (fun sample => sampledscheduledhalftsallisprobabilityattime (finset.univ : finset (fin k)) ⟨model.bestarm, finset.mem_univ model.bestarm⟩ sampledscheduledhalftsallissqrtschedule t sample arm * (|finitearmiidhistoryadaptiverewardshiftat source t sample.2 arm| + |finitearmiidhistoryadaptiverewardshiftat source t sample.2 model.bestarm|)))) theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_nonneg","label":"finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_nonneg","description":"The policy-weighted realized corruption budget is nonnegative.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-7fb2c909b2f3","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9802,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:196"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_nonneg {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (mu : Measure ((Nat -> Fin K -> Rat) × ((k : Nat) -> Fin K × Real))) : 0 <= finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget model horizon source mu","missing":[],"search":"finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_nonneg banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_nonneg the policy-weighted realized corruption budget is nonnegative. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_le","label":"finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_le","description":"The expected realized budget never exceeds the source's deterministic all-round envelope budget.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-54b1f83fc17f","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9803,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:219"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_le {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (mu : Measure ((Nat -> Fin K -> Rat) × ((k : Nat) -> Fin K × Real))) [IsProbabilityMeasure mu] : finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget model horizon source mu <= finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon source","missing":[],"search":"finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_le banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedrewardcorruptionbudget_le the expected realized budget never exceeds the source's deterministic all-round envelope budget. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw_hasSelfBounding","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw_hasSelfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw_hasSelfBounding","description":"The finite-arm IID history-adaptive model satisfies self-bounding with the exact policy-weighted realized corruption budget.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-70859c1acb32","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9804,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:287"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw_hasSelfBounding {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw_hasselfbounding banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw_hasselfbounding the finite-arm iid history-adaptive model satisfies self-bounding with the exact policy-weighted realized corruption budget. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_log","description":"The square-root schedule gives logarithmic regret with the exact expected realized history-adaptive corruption budget.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-70b7552f3c52","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9805,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:418"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlawregret_le_log the square-root schedule gives logarithmic regret with the exact expected realized history-adaptive corruption budget. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","label":"finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","description":"The generated-law specialization of the expected realized corruption budget, packaged without exposing the trajectory measure to callers.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-6fb120b2e919","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9806,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:517"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (horizon : Nat) : Real","missing":[],"search":"finitearmiidhistoryadaptiveexpectedrewardcorruptionbudgetforlaw banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedrewardcorruptionbudgetforlaw the generated-law specialization of the expected realized corruption budget, packaged without exposing the trajectory measure to callers. definition compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRefinedCorruptionWindow","label":"finiteArmIIDHistoryAdaptiveExpectedRefinedCorruptionWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRefinedCorruptionWindow","description":"The coefficient-aware refined window evaluated at the generated policy's expected realized history-adaptive corruption.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-630c6ff7cc74","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9807,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:538"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveExpectedRefinedCorruptionWindow {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (horizon : Nat) : Prop","missing":[],"search":"finitearmiidhistoryadaptiveexpectedrefinedcorruptionwindow banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedrefinedcorruptionwindow the coefficient-aware refined window evaluated at the generated policy's expected realized history-adaptive corruption. definition compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","description":"Refined local regret using the policy-weighted realized corruption budget. The compact window discharges the low-level scalar tuning inequalities.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-ad07d968005b","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9808,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:553"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) (hwindow : finiteArmIIDHistoryAdaptiveExpectedRefinedCorruptionWindow model armLaw source horizon…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlawregret_le_refinedlocalexplicit_of_window banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlawregret_le_refinedlocalexplicit_of_window refined local regret using the policy-weighted realized corruption budget. the compact window discharges the low-level scalar tuning inequalities. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound","label":"finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound","description":"Total regret envelope using the refined expected-corruption expression when there is a suboptimal arm and the compact window holds, and the logarithmic expected-corruption expression otherwise.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-746aa34a2d89","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9809,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:686"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (horizon : Nat) : Real","missing":[],"search":"finitearmiidhistoryadaptiveexpectedcorruptionallregimebound banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedcorruptionallregimebound total regret envelope using the refined expected-corruption expression when there is a suboptimal arm and the compact window holds, and the logarithmic expected-corruption expression otherwise. definition compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound_fin_one","label":"finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound_fin_one","description":"With one arm there is no suboptimal coordinate and the expected corruption all-regimes envelope reduces to the logarithmic base term.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-3ac210987f44","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9810,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:714"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Fin 1 -> Measure Rat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource 1) (horizon : Nat) : finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound model armLaw source horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmiidhistoryadaptiveexpectedcorruptionallregimebound_fin_one banditrlproof.tsallis.finitearmiidhistoryadaptiveexpectedcorruptionallregimebound_fin_one with one arm there is no suboptimal coordinate and the expected corruption all-regimes envelope reduces to the logarithmic base term. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","description":"Scheduled half-Tsallis regret for every finite-arm IID history-adaptive reward-shift source and finite horizon, using the policy-weighted realized corruption budget in both automatically selected regimes.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-21a2c2927a87","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9811,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:729"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","adversarial-bobw"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlawregret_le_allregimes scheduled half-tsallis regret for every finite-arm iid history-adaptive reward-shift source and finite horizon, using the policy-weighted realized corruption budget in both automatically selected regimes. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":["adversarial-bobw"]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostExpectedCorruptionRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostExpectedCorruptionRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostExpectedCorruptionRewardLawRegret_le_allRegimes","description":"Measurable history-arm-gated suboptimal boosts inherit the all-regimes bound with gate-open corruption weighted by conditional selection probability, rather than the full deterministic boost schedule.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-d7d5d386cb50","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","order":9812,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw.lean:792"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostExpectedCorruptionRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (initialGate : Set (Fin K)) (gate : (n : Nat) -> Set (History.FinitePairHistory (Fin K) Real n × Fin K)) (hgate : forall n, MeasurableSet (gate n)) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostexpectedcorruptionrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostexpectedcorruptionrewardlawregret_le_allregimes measurable history-arm-gated suboptimal boosts inherit the all-regimes bound with gate-open corruption weighted by conditional selection probability, rather than the full deterministic boost schedule. theorem compiled","shard":"modules/0ea12c1f9e104637.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw_hasSelfBounding","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw_hasSelfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw_hasSelfBounding","description":"The history-adaptive finite-arm corruption model supplies the terminal self-bounding contract consumed by the refined square-root schedule route.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw/index.html#decl-7a756d8866dc","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","order":9813,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw.lean:13"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw_hasSelfBounding {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlaw_hasselfbounding banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlaw_hasselfbounding the history-adaptive finite-arm corruption model supplies the terminal self-bounding contract consumed by the refined square-root schedule route. theorem compiled","shard":"modules/1e549752339dbd02.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit","description":"Refined local square-root corruption regret for the concrete finite-arm IID history-adaptive reward-shift model. The corruption scalar is the source's deterministic envelope budget, rather than a free theorem parameter.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw/index.html#decl-d761a8d51034","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","order":9814,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw.lean:136"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) (hcorruption : 0 < finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon source) (hscalarLower : 2 <= (2 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * (((horizon + 1…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlawregret_le_refinedlocalexplicit banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlawregret_le_refinedlocalexplicit refined local square-root corruption regret for the concrete finite-arm iid history-adaptive reward-shift model. the corruption scalar is the source's deterministic envelope budget, rather than a free theorem parameter. theorem compiled","shard":"modules/1e549752339dbd02.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow","description":"The coefficient-aware refined corruption window specialized to the finite-arm IID history-adaptive reward-shift model.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw/index.html#decl-205db58b94ab","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","order":9815,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw.lean:251"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) : Prop","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow the coefficient-aware refined corruption window specialized to the finite-arm iid history-adaptive reward-shift model. definition compiled","shard":"modules/1e549752339dbd02.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","description":"A model-facing refined theorem with the low-level scalar inequalities replaced by a compact corruption window. Unit-bounded positive model gaps imply that the reciprocal-gap sum dominates the number of suboptimal arms.","url":"../modules/banditrlproof-tsallisfinitearmiidhistoryadaptiverefinedcorruptedrewardlaw/index.html#decl-3e55a4ac37fd","parent":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","order":9816,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw.lean:264"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHistoryAdaptiveRewardShiftSource K) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) (hwindow : finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow model horizon source) : letI : Nonempty (Fi…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlawregret_le_refinedlocalexplicit_of_window banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhistoryadaptivecorruptedrewardlawregret_le_refinedlocalexplicit_of_window a model-facing refined theorem with the low-level scalar inequalities replaced by a compact corruption window. unit-bounded positive model gaps imply that the reciprocal-gap sum dominates the number of suboptimal arms. theorem compiled","shard":"modules/1e549752339dbd02.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource","label":"FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource","description":"Predictable reward shifts and deterministic envelope witnesses needed only for rounds `0, ..., horizon`. Successor data are indexed by `Fin horizon`, so no measurability or boundedness contract is requested after the final round.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-8c33aec01c21","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9817,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource (K horizon : Nat) where","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftsource banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource predictable reward shifts and deterministic envelope witnesses needed only for rounds `0, ..., horizon`. successor data are indexed by `fin horizon`, so no measurability or boundedness contract is requested after the final round. structure compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor","description":"Successor shift obtained by extending a horizon-local source by zero.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-804c3c855762","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9818,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:37"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : Real","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftsuccessor banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsuccessor successor shift obtained by extending a horizon-local source by zero. definition compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope","description":"Deterministic envelope obtained by extending a horizon-local envelope by zero after its final successor round.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-1fa7153c8961","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9819,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftenvelope banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftenvelope deterministic envelope obtained by extending a horizon-local envelope by zero after its final successor round. definition compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_lt","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_lt","description":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor source n history arm = source.successor ⟨n, hn⟩ history arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-11eadb39a3e0","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9820,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor source n history arm = source.successor ⟨n, hn⟩ history arm","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftsuccessor_of_lt banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsuccessor_of_lt theorem finitearmiidhorizonhistoryadaptiverewardshiftsuccessor_of_lt {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : n < horizon) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : finitearmiidhorizonhistoryadaptiverewardshiftsuccessor source n history arm = source.successor ⟨n, hn⟩ history arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_le","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_le","description":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor source n history arm = 0","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-1e6b5909b2a0","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9821,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:69"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor source n history arm = 0","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftsuccessor_of_le banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsuccessor_of_le theorem finitearmiidhorizonhistoryadaptiverewardshiftsuccessor_of_le {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : horizon <= n) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : finitearmiidhorizonhistoryadaptiverewardshiftsuccessor source n history arm = 0 theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_zero","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_zero","description":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_zero {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope source 0 arm = source.initialEnvelope arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-d2bfeb0f85c9","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9822,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:80"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_zero {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope source 0 arm = source.initialEnvelope arm","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftenvelope_zero banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftenvelope_zero theorem finitearmiidhorizonhistoryadaptiverewardshiftenvelope_zero {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (arm : fin k) : finitearmiidhorizonhistoryadaptiverewardshiftenvelope source 0 arm = source.initialenvelope arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_lt","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_lt","description":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope source (Nat.succ n) arm = source.successorEnvelope ⟨n, hn⟩ arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-99ebcd3e50d3","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9823,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:89"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope source (Nat.succ n) arm = source.successorEnvelope ⟨n, hn⟩ arm","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftenvelope_succ_of_lt banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftenvelope_succ_of_lt theorem finitearmiidhorizonhistoryadaptiverewardshiftenvelope_succ_of_lt {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : n < horizon) (arm : fin k) : finitearmiidhorizonhistoryadaptiverewardshiftenvelope source (nat.succ n) arm = source.successorenvelope ⟨n, hn⟩ arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_le","label":"finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_le","description":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope source (Nat.succ n) arm = 0","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-45179d625212","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9824,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:101"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (arm : Fin K) : finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope source (Nat.succ n) arm = 0","missing":[],"search":"finitearmiidhorizonhistoryadaptiverewardshiftenvelope_succ_of_le banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftenvelope_succ_of_le theorem finitearmiidhorizonhistoryadaptiverewardshiftenvelope_succ_of_le {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : horizon <= n) (arm : fin k) : finitearmiidhorizonhistoryadaptiverewardshiftenvelope source (nat.succ n) arm = 0 theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime","label":"toAllTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime","description":"Zero extension of a horizon-local source to the all-time source interface. The supplied simp lemmas show that the extension is unchanged on every round used by the target horizon; it introduces no post-horizon regularity obligation.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-53d2aa8b02cb","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9825,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:115"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"toalltime banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime zero extension of a horizon-local source to the all-time source interface. the supplied simp lemmas show that the extension is unchanged on every round used by the target horizon; it introduces no post-horizon regularity obligation. definition compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_initial","label":"toAllTime_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_initial","description":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_initial {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (arm : Fin K) : source.toAllTime.initial arm = source.initial arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-cc1f03939a3f","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9826,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:163"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_initial {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (arm : Fin K) : source.toAllTime.initial arm = source.initial arm","missing":[],"search":"toalltime_initial banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_initial theorem finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_initial {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (arm : fin k) : source.toalltime.initial arm = source.initial arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_lt","label":"toAllTime_successor_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_lt","description":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : source.toAllTime.successor n history arm = source.successor ⟨n, hn⟩ history arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-e7efd6682448","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9827,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:170"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : source.toAllTime.successor n history arm = source.successor ⟨n, hn⟩ history arm","missing":[],"search":"toalltime_successor_of_lt banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_successor_of_lt theorem finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_successor_of_lt {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : n < horizon) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : source.toalltime.successor n history arm = source.successor ⟨n, hn⟩ history arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_le","label":"toAllTime_successor_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_le","description":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : source.toAllTime.successor n history arm = 0","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-0365d2f410f1","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9828,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:181"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : source.toAllTime.successor n history arm = 0","missing":[],"search":"toalltime_successor_of_le banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_successor_of_le theorem finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_successor_of_le {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : horizon <= n) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : source.toalltime.successor n history arm = 0 theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_zero","label":"toAllTime_envelope_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_zero","description":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_zero {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (arm : Fin K) : source.toAllTime.envelope 0 arm = source.initialEnvelope arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-3f33a54c7eb6","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9829,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:191"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_zero {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (arm : Fin K) : source.toAllTime.envelope 0 arm = source.initialEnvelope arm","missing":[],"search":"toalltime_envelope_zero banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_envelope_zero theorem finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_envelope_zero {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (arm : fin k) : source.toalltime.envelope 0 arm = source.initialenvelope arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_lt","label":"toAllTime_envelope_succ_of_lt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_lt","description":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (arm : Fin K) : source.toAllTime.envelope (Nat.succ n) arm = source.successorEnvelope ⟨n, hn⟩ arm","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-26d86af1f0b9","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9830,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:199"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_lt {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : n < horizon) (arm : Fin K) : source.toAllTime.envelope (Nat.succ n) arm = source.successorEnvelope ⟨n, hn⟩ arm","missing":[],"search":"toalltime_envelope_succ_of_lt banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_envelope_succ_of_lt theorem finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_envelope_succ_of_lt {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : n < horizon) (arm : fin k) : source.toalltime.envelope (nat.succ n) arm = source.successorenvelope ⟨n, hn⟩ arm theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_le","label":"toAllTime_envelope_succ_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_le","description":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (arm : Fin K) : source.toAllTime.envelope (Nat.succ n) arm = 0","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-2709dd1a0655","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9831,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:209"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_le {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (n : Nat) (hn : horizon <= n) (arm : Fin K) : source.toAllTime.envelope (Nat.succ n) arm = 0","missing":[],"search":"toalltime_envelope_succ_of_le banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_envelope_succ_of_le theorem finitearmiidhorizonhistoryadaptiverewardshiftsource.toalltime_envelope_succ_of_le {k horizon : nat} (source : finitearmiidhorizonhistoryadaptiverewardshiftsource k horizon) (n : nat) (hn : horizon <= n) (arm : fin k) : source.toalltime.envelope (nat.succ n) arm = 0 theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveCorruptedRewardLoss","label":"finiteArmIIDHorizonHistoryAdaptiveCorruptedRewardLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveCorruptedRewardLoss","description":"Predictable clipped loss attached to the zero extension of a horizon-local reward-shift source.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-1efcce07cf8f","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9832,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:219"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHorizonHistoryAdaptiveCorruptedRewardLoss {K horizon : Nat} (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon)","missing":[],"search":"finitearmiidhorizonhistoryadaptivecorruptedrewardloss banditrlproof.tsallis.finitearmiidhorizonhistoryadaptivecorruptedrewardloss predictable clipped loss attached to the zero extension of a horizon-local reward-shift source. definition compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","label":"finiteArmIIDHorizonHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","description":"Exact generated-policy expected-corruption budget of a horizon-local source. Only rounds through `horizon` occur in the finite sum.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-7b02a118c6e7","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9833,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:226"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHorizonHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw {K horizon : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) : Real","missing":[],"search":"finitearmiidhorizonhistoryadaptiveexpectedrewardcorruptionbudgetforlaw banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiveexpectedrewardcorruptionbudgetforlaw exact generated-policy expected-corruption budget of a horizon-local source. only rounds through `horizon` occur in the finite sum. definition compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveExpectedCorruptionAllRegimeBound","label":"finiteArmIIDHorizonHistoryAdaptiveExpectedCorruptionAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveExpectedCorruptionAllRegimeBound","description":"All-regimes expected-corruption envelope for a horizon-local source.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-c544a2baf264","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9834,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:235"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDHorizonHistoryAdaptiveExpectedCorruptionAllRegimeBound {K horizon : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) : Real","missing":[],"search":"finitearmiidhorizonhistoryadaptiveexpectedcorruptionallregimebound banditrlproof.tsallis.finitearmiidhorizonhistoryadaptiveexpectedcorruptionallregimebound all-regimes expected-corruption envelope for a horizon-local source. definition compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","description":"Scheduled half-Tsallis regret under a finite-arm IID reward law and a history-adaptive reward-shift source whose regularity contract stops at the target horizon. The conclusion uses the exact generated-policy expected corruption and internally selects the refined or logarithmic branch.","url":"../modules/banditrlproof-tsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlaw/index.html#decl-daa384722f53","parent":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","order":9835,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw.lean:247"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes {K horizon : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (source : FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource K horizon) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidhorizonhistoryadaptiveexpectedcorruptedrewardlawregret_le_allregimes scheduled half-tsallis regret under a finite-arm iid reward law and a history-adaptive reward-shift source whose regularity contract stops at the target horizon. the conclusion uses the exact generated-policy expected corruption and internally selects the refined or logarithmic branch. theorem compiled","shard":"modules/3a1cebcc93d6b123.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurableHistoryArmGatedSuboptimalRewardBoostSource","label":"measurableHistoryArmGatedSuboptimalRewardBoostSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurableHistoryArmGatedSuboptimalRewardBoostSource","description":"A history-adaptive corruption source with an arbitrary initial arm gate and arbitrary measurable finite-pair-history-and-arm successor gates. The best arm is never shifted.","url":"../modules/banditrlproof-tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret/index.html#decl-ffbd85f78077","parent":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","order":9836,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret.lean:12"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def measurableHistoryArmGatedSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (initialGate : Set (Fin K)) (gate : (n : Nat) -> Set (History.FinitePairHistory (Fin K) Real n × Fin K)) (hgate : forall n, MeasurableSet (gate n)) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"measurablehistoryarmgatedsuboptimalrewardboostsource banditrlproof.tsallis.measurablehistoryarmgatedsuboptimalrewardboostsource a history-adaptive corruption source with an arbitrary initial arm gate and arbitrary measurable finite-pair-history-and-arm successor gates. the best arm is never shifted. definition compiled","shard":"modules/b5361afee43a7f6c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_measurableHistoryArmGatedSuboptimalRewardBoostSource","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_measurableHistoryArmGatedSuboptimalRewardBoostSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_measurableHistoryArmGatedSuboptimalRewardBoostSource","description":"The measurable history-arm gate does not enlarge the deterministic envelope, whose budget is exactly the underlying time-varying schedule.","url":"../modules/banditrlproof-tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret/index.html#decl-b4652d451889","parent":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","order":9837,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_measurableHistoryArmGatedSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (initialGate : Set (Fin K)) (gate : (n : Nat) -> Set (History.FinitePairHistory (Fin K) Real n × Fin K)) (hgate : forall n, MeasurableSet (gate n)) (horizon : Nat) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (measurableHistoryArmGatedSuboptimalRewardBoostSource model initialGate gate hgate boost hboost) = finiteArmIIDTimeVaryingSuboptimalBoostBudget model horizon boost","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_measurablehistoryarmgatedsuboptimalrewardboostsource banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_measurablehistoryarmgatedsuboptimalrewardboostsource the measurable history-arm gate does not enlarge the deterministic envelope, whose budget is exactly the underlying time-varying schedule. theorem compiled","shard":"modules/b5361afee43a7f6c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_measurableHistoryArmGatedSuboptimalRewardBoostSource_of_refinedRegime","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_measurableHistoryArmGatedSuboptimalRewardBoostSource_of_refinedRegime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_measurableHistoryArmGatedSuboptimalRewardBoostSource_of_refinedRegime","description":"The named time-varying refined regime supplies the compact window for an arbitrary measurable history-arm-gated source with the same envelope.","url":"../modules/banditrlproof-tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret/index.html#decl-1cd8614147fa","parent":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","order":9838,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret.lean:84"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_measurableHistoryArmGatedSuboptimalRewardBoostSource_of_refinedRegime {K : Nat} (model : FiniteBanditModel K) (initialGate : Set (Fin K)) (gate : (n : Nat) -> Set (History.FinitePairHistory (Fin K) Real n × Fin K)) (hgate : forall n, MeasurableSet (gate n)) (horizon : Nat) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hregime : finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime model horizon boost) : finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow model horizon (measurableHistoryArmGatedSuboptimalRewardBoostSource model initialGate gate hgate boost hboost)","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow_measurablehistoryarmgatedsuboptimalrewardboostsource_of_refinedregime banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow_measurablehistoryarmgatedsuboptimalrewardboostsource_of_refinedregime the named time-varying refined regime supplies the compact window for an arbitrary measurable history-arm-gated source with the same envelope. theorem compiled","shard":"modules/b5361afee43a7f6c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRewardLawRegret_le_allRegimes","description":"Scheduled half-Tsallis regret for every initial arm gate and measurable predictable successor gate on the complete finite pair history and candidate arm. The theorem covers every finite horizon and selects the refined or logarithmic branch internally.","url":"../modules/banditrlproof-tsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostregret/index.html#decl-61c98d2344f0","parent":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","order":9839,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret.lean:105"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (initialGate : Set (Fin K)) (gate : (n : Nat) -> Set (History.FinitePairHistory (Fin K) Real n × Fin K)) (hgate : forall n, MeasurableSet (gate n)) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm ->…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidmeasurablehistoryarmgatedsuboptimalboostrewardlawregret_le_allregimes scheduled half-tsallis regret for every initial arm gate and measurable predictable successor gate on the complete finite pair history and candidate arm. the theorem covers every finite horizon and selects the refined or logarithmic branch internally. theorem compiled","shard":"modules/b5361afee43a7f6c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.previousActionGatedSuboptimalRewardBoostSource","label":"previousActionGatedSuboptimalRewardBoostSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.previousActionGatedSuboptimalRewardBoostSource","description":"A concrete history-adaptive corruption source. At successor time `n+1`, the suboptimal-arm boost is active exactly when the action observed at time `n` equals `triggerArm`. The best arm is never shifted.","url":"../modules/banditrlproof-tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret/index.html#decl-a90adfb30436","parent":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","order":9840,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret.lean:12"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def previousActionGatedSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (triggerArm : Fin K) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"previousactiongatedsuboptimalrewardboostsource banditrlproof.tsallis.previousactiongatedsuboptimalrewardboostsource a concrete history-adaptive corruption source. at successor time `n+1`, the suboptimal-arm boost is active exactly when the action observed at time `n` equals `triggerarm`. the best arm is never shifted. definition compiled","shard":"modules/a79a1b34f8c9e3b5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_previousActionGatedSuboptimalRewardBoostSource","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_previousActionGatedSuboptimalRewardBoostSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_previousActionGatedSuboptimalRewardBoostSource","description":"The deterministic envelope of the previous-action-gated source has the same exact double finite-sum budget as its ungated time-varying schedule.","url":"../modules/banditrlproof-tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret/index.html#decl-5951df8e1c7d","parent":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","order":9841,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret.lean:62"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_previousActionGatedSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (triggerArm : Fin K) (horizon : Nat) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (previousActionGatedSuboptimalRewardBoostSource model triggerArm boost hboost) = finiteArmIIDTimeVaryingSuboptimalBoostBudget model horizon boost","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_previousactiongatedsuboptimalrewardboostsource banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_previousactiongatedsuboptimalrewardboostsource the deterministic envelope of the previous-action-gated source has the same exact double finite-sum budget as its ungated time-varying schedule. theorem compiled","shard":"modules/a79a1b34f8c9e3b5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_previousActionGatedSuboptimalRewardBoostSource_of_refinedRegime","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_previousActionGatedSuboptimalRewardBoostSource_of_refinedRegime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_previousActionGatedSuboptimalRewardBoostSource_of_refinedRegime","description":"The time-varying named regime supplies the compact refined window for the history-adaptive previous-action-gated source.","url":"../modules/banditrlproof-tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret/index.html#decl-e348270e43f1","parent":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","order":9842,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret.lean:81"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_previousActionGatedSuboptimalRewardBoostSource_of_refinedRegime {K : Nat} (model : FiniteBanditModel K) (triggerArm : Fin K) (horizon : Nat) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hregime : finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime model horizon boost) : finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow model horizon (previousActionGatedSuboptimalRewardBoostSource model triggerArm boost hboost)","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow_previousactiongatedsuboptimalrewardboostsource_of_refinedregime banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow_previousactiongatedsuboptimalrewardboostsource_of_refinedregime the time-varying named regime supplies the compact refined window for the history-adaptive previous-action-gated source. theorem compiled","shard":"modules/a79a1b34f8c9e3b5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRewardLawRegret_le_allRegimes","description":"Scheduled half-Tsallis regret for a concrete source whose successor boost depends on the previous sampled action. The theorem covers every finite horizon and internally selects the refined or logarithmic branch.","url":"../modules/banditrlproof-tsallisfinitearmiidpreviousactiongatedsuboptimalboostregret/index.html#decl-c440086ad6a5","parent":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","order":9843,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret.lean:97"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (triggerArm : Fin K) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidpreviousactiongatedsuboptimalboostrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidpreviousactiongatedsuboptimalboostrewardlawregret_le_allregimes scheduled half-tsallis regret for a concrete source whose successor boost depends on the previous sampled action. the theorem covers every finite horizon and internally selects the refined or logarithmic branch. theorem compiled","shard":"modules/a79a1b34f8c9e3b5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReward","label":"clippedUnitReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReward","description":"Pointwise projection of a rational reward into the real unit interval.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-f366fd2c279d","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9844,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def clippedUnitReward (reward : Rat) : Real","missing":[],"search":"clippedunitreward banditrlproof.tsallis.clippedunitreward pointwise projection of a rational reward into the real unit interval. definition compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReward_nonneg","label":"clippedUnitReward_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReward_nonneg","description":"theorem clippedUnitReward_nonneg (reward : Rat) : 0 <= clippedUnitReward reward","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-49c5b2ed3ada","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9845,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem clippedUnitReward_nonneg (reward : Rat) : 0 <= clippedUnitReward reward","missing":[],"search":"clippedunitreward_nonneg banditrlproof.tsallis.clippedunitreward_nonneg theorem clippedunitreward_nonneg (reward : rat) : 0 <= clippedunitreward reward theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReward_le_one","label":"clippedUnitReward_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReward_le_one","description":"theorem clippedUnitReward_le_one (reward : Rat) : clippedUnitReward reward <= 1","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-080f63bfbd8f","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9846,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem clippedUnitReward_le_one (reward : Rat) : clippedUnitReward reward <= 1","missing":[],"search":"clippedunitreward_le_one banditrlproof.tsallis.clippedunitreward_le_one theorem clippedunitreward_le_one (reward : rat) : clippedunitreward reward <= 1 theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.clippedUnitReward_eq_of_mem_Icc","label":"clippedUnitReward_eq_of_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.clippedUnitReward_eq_of_mem_Icc","description":"theorem clippedUnitReward_eq_of_mem_Icc (reward : Rat) (hreward : ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) : clippedUnitReward reward = ((reward : Rat) : Real)","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-d36e90bc88fe","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9847,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem clippedUnitReward_eq_of_mem_Icc (reward : Rat) (hreward : ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) : clippedUnitReward reward = ((reward : Rat) : Real)","missing":[],"search":"clippedunitreward_eq_of_mem_icc banditrlproof.tsallis.clippedunitreward_eq_of_mem_icc theorem clippedunitreward_eq_of_mem_icc (reward : rat) (hreward : ((reward : rat) : real) ∈ set.icc (0 : real) 1) : clippedunitreward reward = ((reward : rat) : real) theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_clippedUnitReward","label":"measurable_clippedUnitReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_clippedUnitReward","description":"theorem measurable_clippedUnitReward : Measurable clippedUnitReward","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-8fd785c80b4f","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9848,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:36"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_clippedUnitReward : Measurable clippedUnitReward","missing":[],"search":"measurable_clippedunitreward banditrlproof.tsallis.measurable_clippedunitreward theorem measurable_clippedunitreward : measurable clippedunitreward theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLaw","label":"finiteArmIIDRewardVectorLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDRewardVectorLaw","description":"The independent one-round reward vector induced by one law per arm.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-7096e1afb2f1","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9849,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDRewardVectorLaw {K : Nat} (armLaw : Fin K -> Measure Rat) : Measure (Fin K -> Rat)","missing":[],"search":"finitearmiidrewardvectorlaw banditrlproof.tsallis.finitearmiidrewardvectorlaw the independent one-round reward vector induced by one law per arm. definition compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss","label":"finiteArmIIDRewardVectorLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss","description":"Convert a sampled reward vector into the selected arm's clipped loss.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-8a2dbe3c7f8f","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9850,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:45"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDRewardVectorLoss {K : Nat} (state : Fin K -> Rat) (arm : Fin K) : Real","missing":[],"search":"finitearmiidrewardvectorloss banditrlproof.tsallis.finitearmiidrewardvectorloss convert a sampled reward vector into the selected arm's clipped loss. definition compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDRewardVectorLoss","label":"measurable_finiteArmIIDRewardVectorLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_finiteArmIIDRewardVectorLoss","description":"theorem measurable_finiteArmIIDRewardVectorLoss {K : Nat} : Measurable (fun input : (Fin K -> Rat) × Fin K => finiteArmIIDRewardVectorLoss input.1 input.2)","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-cdd04f8e5333","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9851,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteArmIIDRewardVectorLoss {K : Nat} : Measurable (fun input : (Fin K -> Rat) × Fin K => finiteArmIIDRewardVectorLoss input.1 input.2)","missing":[],"search":"measurable_finitearmiidrewardvectorloss banditrlproof.tsallis.measurable_finitearmiidrewardvectorloss theorem measurable_finitearmiidrewardvectorloss {k : nat} : measurable (fun input : (fin k -> rat) × fin k => finitearmiidrewardvectorloss input.1 input.2) theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss_nonneg","label":"finiteArmIIDRewardVectorLoss_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss_nonneg","description":"theorem finiteArmIIDRewardVectorLoss_nonneg {K : Nat} (state : Fin K -> Rat) (arm : Fin K) : 0 <= finiteArmIIDRewardVectorLoss state arm","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-df645454744b","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9852,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:54"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDRewardVectorLoss_nonneg {K : Nat} (state : Fin K -> Rat) (arm : Fin K) : 0 <= finiteArmIIDRewardVectorLoss state arm","missing":[],"search":"finitearmiidrewardvectorloss_nonneg banditrlproof.tsallis.finitearmiidrewardvectorloss_nonneg theorem finitearmiidrewardvectorloss_nonneg {k : nat} (state : fin k -> rat) (arm : fin k) : 0 <= finitearmiidrewardvectorloss state arm theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss_le_one","label":"finiteArmIIDRewardVectorLoss_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss_le_one","description":"theorem finiteArmIIDRewardVectorLoss_le_one {K : Nat} (state : Fin K -> Rat) (arm : Fin K) : finiteArmIIDRewardVectorLoss state arm <= 1","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-33b9f98b5696","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9853,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:60"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDRewardVectorLoss_le_one {K : Nat} (state : Fin K -> Rat) (arm : Fin K) : finiteArmIIDRewardVectorLoss state arm <= 1","missing":[],"search":"finitearmiidrewardvectorloss_le_one banditrlproof.tsallis.finitearmiidrewardvectorloss_le_one theorem finitearmiidrewardvectorloss_le_one {k : nat} (state : fin k -> rat) (arm : fin k) : finitearmiidrewardvectorloss state arm <= 1 theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_clippedUnitReward","label":"integrable_clippedUnitReward","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_clippedUnitReward","description":"theorem integrable_clippedUnitReward (mu : Measure Rat) [IsFiniteMeasure mu] : Integrable clippedUnitReward mu","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-a3cca64f1f19","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9854,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:66"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_clippedUnitReward (mu : Measure Rat) [IsFiniteMeasure mu] : Integrable clippedUnitReward mu","missing":[],"search":"integrable_clippedunitreward banditrlproof.tsallis.integrable_clippedunitreward theorem integrable_clippedunitreward (mu : measure rat) [isfinitemeasure mu] : integrable clippedunitreward mu theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_finiteArmIIDRewardVectorLaw_clippedUnitReward_eq_mean","label":"integral_finiteArmIIDRewardVectorLaw_clippedUnitReward_eq_mean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_finiteArmIIDRewardVectorLaw_clippedUnitReward_eq_mean","description":"Under an almost-sure unit-interval contract, the product-coordinate clipped reward has exactly the supplied finite-bandit mean.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-d3c7dc136280","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9855,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:76"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_finiteArmIIDRewardVectorLaw_clippedUnitReward_eq_mean {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (arm : Fin K) : integral (finiteArmIIDRewardVectorLaw armLaw) (fun state => clippedUnitReward (state arm)) = ((model.mean arm : Rat) : Real)","missing":[],"search":"integral_finitearmiidrewardvectorlaw_clippedunitreward_eq_mean banditrlproof.tsallis.integral_finitearmiidrewardvectorlaw_clippedunitreward_eq_mean under an almost-sure unit-interval contract, the product-coordinate clipped reward has exactly the supplied finite-bandit mean. theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.iidLossStateMeanGap_finiteArmIIDRewardVectorLoss_eq_gap","label":"iidLossStateMeanGap_finiteArmIIDRewardVectorLoss_eq_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.iidLossStateMeanGap_finiteArmIIDRewardVectorLoss_eq_gap","description":"The one-round IID loss-state mean gap is exactly the rational finite-bandit model gap after coercion to the reals.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-657b3659db97","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9856,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:104"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem iidLossStateMeanGap_finiteArmIIDRewardVectorLoss_eq_gap {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (arm : Fin K) : iidLossStateMeanGap (finiteArmIIDRewardVectorLaw armLaw) finiteArmIIDRewardVectorLoss model.bestArm arm = ((model.gap arm : Rat) : Real)","missing":[],"search":"iidlossstatemeangap_finitearmiidrewardvectorloss_eq_gap banditrlproof.tsallis.iidlossstatemeangap_finitearmiidrewardvectorloss_eq_gap the one-round iid loss-state mean gap is exactly the rational finite-bandit model gap after coercion to the reals. theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","description":"A finite collection of bounded rational reward laws supplies the concrete IID stochastic model for the generated scheduled half-Tsallis logarithmic regret theorem. The latent one-round law is the independent product of arm laws, while the observed feedback is the selected clipped loss.","url":"../modules/banditrlproof-tsallisfinitearmiidrewardlaw/index.html#decl-7e0452498d9b","parent":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","order":9857,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDRewardLaw.lean:160"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","adversarial-bobw"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) (corruption : Real) (hcorruption : 0 <= corruption) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidrewardlawregret_le_log a finite collection of bounded rational reward laws supplies the concrete iid stochastic model for the generated scheduled half-tsallis logarithmic regret theorem. the latent one-round law is the independent product of arm laws, while the observed feedback is the selected clipped loss. theorem compiled","shard":"modules/ffb1c4fe34ae905a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":["adversarial-bobw"]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedReward","label":"finiteArmIIDTimeVaryingCorruptedReward","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedReward","description":"Reward after a deterministic time-indexed arm shift and clipping.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-ca3778fd39be","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9858,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:19"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDTimeVaryingCorruptedReward {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) (state : Fin K -> Rat) (arm : Fin K) : Real","missing":[],"search":"finitearmiidtimevaryingcorruptedreward banditrlproof.tsallis.finitearmiidtimevaryingcorruptedreward reward after a deterministic time-indexed arm shift and clipping. definition compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","label":"finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","description":"Loss vector induced by the time-indexed clipped reward.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-d05df7dc37e6","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9859,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDTimeVaryingCorruptedRewardVectorLoss {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) (state : Fin K -> Rat) (arm : Fin K) : Real","missing":[],"search":"finitearmiidtimevaryingcorruptedrewardvectorloss banditrlproof.tsallis.finitearmiidtimevaryingcorruptedrewardvectorloss loss vector induced by the time-indexed clipped reward. definition compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","label":"measurable_finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","description":"theorem measurable_finiteArmIIDTimeVaryingCorruptedRewardVectorLoss {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) : Measurable (fun input : (Fin K -> Rat) × Fin K => finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift t input.1 input.2)","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-79eca89ccf13","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9860,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:30"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_finiteArmIIDTimeVaryingCorruptedRewardVectorLoss {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) : Measurable (fun input : (Fin K -> Rat) × Fin K => finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift t input.1 input.2)","missing":[],"search":"measurable_finitearmiidtimevaryingcorruptedrewardvectorloss banditrlproof.tsallis.measurable_finitearmiidtimevaryingcorruptedrewardvectorloss theorem measurable_finitearmiidtimevaryingcorruptedrewardvectorloss {k : nat} (rewardshift : nat -> fin k -> real) (t : nat) : measurable (fun input : (fin k -> rat) × fin k => finitearmiidtimevaryingcorruptedrewardvectorloss rewardshift t input.1 input.2) theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_nonneg","label":"finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_nonneg","description":"theorem finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_nonneg {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) (state : Fin K -> Rat) (arm : Fin K) : 0 <= finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift t state arm","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-3502bb0d3104","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9861,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_nonneg {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) (state : Fin K -> Rat) (arm : Fin K) : 0 <= finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift t state arm","missing":[],"search":"finitearmiidtimevaryingcorruptedrewardvectorloss_nonneg banditrlproof.tsallis.finitearmiidtimevaryingcorruptedrewardvectorloss_nonneg theorem finitearmiidtimevaryingcorruptedrewardvectorloss_nonneg {k : nat} (rewardshift : nat -> fin k -> real) (t : nat) (state : fin k -> rat) (arm : fin k) : 0 <= finitearmiidtimevaryingcorruptedrewardvectorloss rewardshift t state arm theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_le_one","label":"finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_le_one","description":"theorem finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_le_one {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) (state : Fin K -> Rat) (arm : Fin K) : finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift t state arm <= 1","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-c3b8baf22301","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9862,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:52"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_le_one {K : Nat} (rewardShift : Nat -> Fin K -> Real) (t : Nat) (state : Fin K -> Rat) (arm : Fin K) : finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift t state arm <= 1","missing":[],"search":"finitearmiidtimevaryingcorruptedrewardvectorloss_le_one banditrlproof.tsallis.finitearmiidtimevaryingcorruptedrewardvectorloss_le_one theorem finitearmiidtimevaryingcorruptedrewardvectorloss_le_one {k : nat} (rewardshift : nat -> fin k -> real) (t : nat) (state : fin k -> rat) (arm : fin k) : finitearmiidtimevaryingcorruptedrewardvectorloss rewardshift t state arm <= 1 theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_iidLossStateTimeVaryingMeanGap_corrupted_sub_modelGap_le","label":"abs_iidLossStateTimeVaryingMeanGap_corrupted_sub_modelGap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_iidLossStateTimeVaryingMeanGap_corrupted_sub_modelGap_le","description":"The actual mean gap at round `t` differs from the baseline model gap by at most the two affected arm shifts.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-7ab1279bae1f","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9863,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:65"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_iidLossStateTimeVaryingMeanGap_corrupted_sub_modelGap_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : ∀ arm, IsProbabilityMeasure (armLaw arm)) (hbound : ∀ arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : ∀ arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (rewardShift : Nat -> Fin K -> Real) (t : Nat) (arm : Fin K) : |iidLossStateTimeVaryingMeanGap (finiteArmIIDRewardVectorLaw armLaw) (finiteArmIIDTimeVaryingCorruptedRewardVectorLoss rewardShift) t model.bestArm arm - ((model.gap arm : Rat) : Real)| <= |rewardShift t arm| + |rewardShift t model.bestArm|","missing":[],"search":"abs_iidlossstatetimevaryingmeangap_corrupted_sub_modelgap_le banditrlproof.tsallis.abs_iidlossstatetimevaryingmeangap_corrupted_sub_modelgap_le the actual mean gap at round `t` differs from the baseline model gap by at most the two affected arm shifts. theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget","label":"finiteArmIIDTimeVaryingRewardCorruptionBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget","description":"Explicit accumulated budget of a time-indexed oblivious reward-shift schedule through the inclusive horizon.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-d1d8d44c7f1c","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9864,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:89"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDTimeVaryingRewardCorruptionBudget {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmiidtimevaryingrewardcorruptionbudget banditrlproof.tsallis.finitearmiidtimevaryingrewardcorruptionbudget explicit accumulated budget of a time-indexed oblivious reward-shift schedule through the inclusive horizon. definition compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_eq","label":"finiteArmIIDTimeVaryingRewardCorruptionBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_eq","description":"theorem finiteArmIIDTimeVaryingRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Nat -> Fin K -> Real) : finiteArmIIDTimeVaryingRewardCorruptionBudget model horizon rewardShift = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => |rewardShift t arm| + |rewardShift t model.bestArm|))","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-f7839b4e68b6","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9865,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:96"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDTimeVaryingRewardCorruptionBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Nat -> Fin K -> Real) : finiteArmIIDTimeVaryingRewardCorruptionBudget model horizon rewardShift = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => |rewardShift t arm| + |rewardShift t model.bestArm|))","missing":[],"search":"finitearmiidtimevaryingrewardcorruptionbudget_eq banditrlproof.tsallis.finitearmiidtimevaryingrewardcorruptionbudget_eq theorem finitearmiidtimevaryingrewardcorruptionbudget_eq {k : nat} (model : finitebanditmodel k) (horizon : nat) (rewardshift : nat -> fin k -> real) : finitearmiidtimevaryingrewardcorruptionbudget model horizon rewardshift = (finset.range (horizon + 1)).sum (fun t => ((finset.univ : finset (fin k)).erase model.bestarm).sum (fun arm => |rewardshift t arm| + |rewardshift t model.bestarm|)) theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_zero","label":"finiteArmIIDTimeVaryingRewardCorruptionBudget_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_zero","description":"theorem finiteArmIIDTimeVaryingRewardCorruptionBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIIDTimeVaryingRewardCorruptionBudget model horizon (fun _ _ => 0) = 0","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-107d2a678755","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9866,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:107"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDTimeVaryingRewardCorruptionBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIIDTimeVaryingRewardCorruptionBudget model horizon (fun _ _ => 0) = 0","missing":[],"search":"finitearmiidtimevaryingrewardcorruptionbudget_zero banditrlproof.tsallis.finitearmiidtimevaryingrewardcorruptionbudget_zero theorem finitearmiidtimevaryingrewardcorruptionbudget_zero {k : nat} (model : finitebanditmodel k) (horizon : nat) : finitearmiidtimevaryingrewardcorruptionbudget model horizon (fun _ _ => 0) = 0 theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_const_eq_stationary","label":"finiteArmIIDTimeVaryingRewardCorruptionBudget_const_eq_stationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_const_eq_stationary","description":"A constant shift schedule recovers the stationary corruption budget.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-5e111d72b9af","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9867,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:115"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDTimeVaryingRewardCorruptionBudget_const_eq_stationary {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (rewardShift : Fin K -> Real) : finiteArmIIDTimeVaryingRewardCorruptionBudget model horizon (fun _ => rewardShift) = finiteArmIIDStationaryRewardCorruptionBudget model horizon rewardShift","missing":[],"search":"finitearmiidtimevaryingrewardcorruptionbudget_const_eq_stationary banditrlproof.tsallis.finitearmiidtimevaryingrewardcorruptionbudget_const_eq_stationary a constant shift schedule recovers the stationary corruption budget. theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingCorruptedRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingCorruptedRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingCorruptedRewardLawRegret_le_log","description":"Scheduled half-Tsallis logarithmic regret under an explicit deterministic time-indexed reward-shift schedule. The additive allowance is derived from the schedule and is not a free corruption parameter.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingcorruptedrewardlaw/index.html#decl-c7bb6ba29a7e","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","order":9868,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw.lean:130"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingCorruptedRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : ∀ arm, IsProbabilityMeasure (armLaw arm)) (hbound : ∀ arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : ∀ arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (rewardShift : Nat -> Fin K -> Real) (hgapPos : ∀ arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidtimevaryingcorruptedrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidtimevaryingcorruptedrewardlawregret_le_log scheduled half-tsallis logarithmic regret under an explicit deterministic time-indexed reward-shift schedule. the additive allowance is derived from the schedule and is not a free corruption parameter. theorem compiled","shard":"modules/d6d5939c9663cbce.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.timeVaryingSuboptimalRewardBoostSource","label":"timeVaryingSuboptimalRewardBoostSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.timeVaryingSuboptimalRewardBoostSource","description":"A deterministic time-and-arm-dependent corruption process that leaves the best arm unchanged and adds a prescribed nonnegative boost to every other arm.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-04dbcee8ceda","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9869,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:11"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def timeVaryingSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"timevaryingsuboptimalrewardboostsource banditrlproof.tsallis.timevaryingsuboptimalrewardboostsource a deterministic time-and-arm-dependent corruption process that leaves the best arm unchanged and adds a prescribed nonnegative boost to every other arm. definition compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostBudget","label":"finiteArmIIDTimeVaryingSuboptimalBoostBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostBudget","description":"Exact deterministic corruption mass of a time-varying suboptimal-arm boost through the inclusive horizon.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-83beda728ac8","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9870,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDTimeVaryingSuboptimalBoostBudget {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmiidtimevaryingsuboptimalboostbudget banditrlproof.tsallis.finitearmiidtimevaryingsuboptimalboostbudget exact deterministic corruption mass of a time-varying suboptimal-arm boost through the inclusive horizon. definition compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_timeVaryingSuboptimalRewardBoostSource","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_timeVaryingSuboptimalRewardBoostSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_timeVaryingSuboptimalRewardBoostSource","description":"The source envelope budget is exactly the finite time-and-arm boost sum.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-84036fe3c64e","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9871,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:42"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_timeVaryingSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (timeVaryingSuboptimalRewardBoostSource model boost hboost) = finiteArmIIDTimeVaryingSuboptimalBoostBudget model horizon boost","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_timevaryingsuboptimalrewardboostsource banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_timevaryingsuboptimalrewardboostsource the source envelope budget is exactly the finite time-and-arm boost sum. theorem compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime","label":"finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime","description":"The coefficient-aware refined regime for a time-varying suboptimal-arm boost budget.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-e00f5e813b6e","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9872,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:59"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Nat -> Fin K -> Real) : Prop","missing":[],"search":"finitearmiidtimevaryingsuboptimalboostrefinedregime banditrlproof.tsallis.finitearmiidtimevaryingsuboptimalboostrefinedregime the coefficient-aware refined regime for a time-varying suboptimal-arm boost budget. definition compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_timeVaryingSuboptimalRewardBoostSource_of_refinedRegime","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_timeVaryingSuboptimalRewardBoostSource_of_refinedRegime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_timeVaryingSuboptimalRewardBoostSource_of_refinedRegime","description":"The named time-varying refined regime supplies the existing model-facing compact corruption window after the exact budget rewrite.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-9e071ead275c","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9873,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:79"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_timeVaryingSuboptimalRewardBoostSource_of_refinedRegime {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hregime : finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime model horizon boost) : finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow model horizon (timeVaryingSuboptimalRewardBoostSource model boost hboost)","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow_timevaryingsuboptimalrewardboostsource_of_refinedregime banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow_timevaryingsuboptimalrewardboostsource_of_refinedregime the named time-varying refined regime supplies the existing model-facing compact corruption window after the exact budget rewrite. theorem compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostAllRegimeBound","label":"finiteArmIIDTimeVaryingSuboptimalBoostAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostAllRegimeBound","description":"Total regret bound for deterministic time-varying boosts: the refined local expression inside its named regime and the logarithmic additive-budget expression on the complement.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-147d39bf0f66","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9874,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:93"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDTimeVaryingSuboptimalBoostAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (boost : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmiidtimevaryingsuboptimalboostallregimebound banditrlproof.tsallis.finitearmiidtimevaryingsuboptimalboostallregimebound total regret bound for deterministic time-varying boosts: the refined local expression inside its named regime and the logarithmic additive-budget expression on the complement. definition compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingSuboptimalBoostRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingSuboptimalBoostRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingSuboptimalBoostRewardLawRegret_le_allRegimes","description":"Scheduled half-Tsallis regret for every deterministic nonnegative time-varying suboptimal reward boost and every finite horizon. The theorem selects the refined or logarithmic branch and requires no caller window proof.","url":"../modules/banditrlproof-tsallisfinitearmiidtimevaryingsuboptimalboostregret/index.html#decl-04a8dfb3a29a","parent":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","order":9875,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret.lean:117"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingSuboptimalBoostRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (boost : Nat -> Fin K -> Real) (hboost : forall t arm, 0 <= boost t arm) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiidtimevaryingsuboptimalboostrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiidtimevaryingsuboptimalboostrewardlawregret_le_allregimes scheduled half-tsallis regret for every deterministic nonnegative time-varying suboptimal reward boost and every finite horizon. the theorem selects the refined or logarithmic branch and requires no caller window proof. theorem compiled","shard":"modules/cc17252d1756c662.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.uniformSuboptimalRewardBoostSource","label":"uniformSuboptimalRewardBoostSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.uniformSuboptimalRewardBoostSource","description":"A concrete corruption process that leaves the best arm unchanged and adds the same nonnegative reward boost to every other arm at every round.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-a8bf023c18ed","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9876,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:11"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def uniformSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (epsilon : Real) (hepsilon : 0 <= epsilon) : FiniteArmIIDHistoryAdaptiveRewardShiftSource K where","missing":[],"search":"uniformsuboptimalrewardboostsource banditrlproof.tsallis.uniformsuboptimalrewardboostsource a concrete corruption process that leaves the best arm unchanged and adds the same nonnegative reward boost to every other arm at every round. definition compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_uniformSuboptimalRewardBoostSource","label":"finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_uniformSuboptimalRewardBoostSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_uniformSuboptimalRewardBoostSource","description":"The uniform suboptimal-arm boost has the exact deterministic envelope budget `(T+1) * (# suboptimal arms) * epsilon`.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-57872359a82a","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9877,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_uniformSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (epsilon : Real) (hepsilon : 0 <= epsilon) : finiteArmIIDHistoryAdaptiveRewardCorruptionBudget model horizon (uniformSuboptimalRewardBoostSource model epsilon hepsilon) = (((horizon + 1 : Nat) : Real)) * (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * epsilon","missing":[],"search":"finitearmiidhistoryadaptiverewardcorruptionbudget_uniformsuboptimalrewardboostsource banditrlproof.tsallis.finitearmiidhistoryadaptiverewardcorruptionbudget_uniformsuboptimalrewardboostsource the uniform suboptimal-arm boost has the exact deterministic envelope budget `(t+1) * (# suboptimal arms) * epsilon`. theorem compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDUniformSuboptimalBoostRefinedRegime","label":"finiteArmIIDUniformSuboptimalBoostRefinedRegime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDUniformSuboptimalBoostRefinedRegime","description":"The explicit scalar regime in which the coefficient-aware refined theorem is used for the uniform suboptimal-arm boost.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-dd659fd491cb","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9878,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:73"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDUniformSuboptimalBoostRefinedRegime {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (epsilon : Real) : Prop","missing":[],"search":"finitearmiiduniformsuboptimalboostrefinedregime banditrlproof.tsallis.finitearmiiduniformsuboptimalboostrefinedregime the explicit scalar regime in which the coefficient-aware refined theorem is used for the uniform suboptimal-arm boost. definition compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource","description":"Natural scalar conditions for the uniform suboptimal-arm boost imply the coefficient-aware refined corruption window.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-37156805c58c","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9879,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:91"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (epsilon : Real) (hepsilon : 0 <= epsilon) (hhorizon : 25 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => 1 / ((model.gap arm : Rat) : Real))) ^ 2 <= (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * (((horizon + 1 : Nat) : Real))) (hepsilonGap : epsilon * ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => 1 / ((model.gap arm : Rat) : Real)) <= 1) (hcorruptionLower : 25 * ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => 1 / ((model.gap arm : Rat) : Real)) * (Real.log ((2 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * (((horizon + 1 : Nat) : Real))) / (25 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => 1 / ((model…","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow_uniformsuboptimalrewardboostsource banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow_uniformsuboptimalrewardboostsource natural scalar conditions for the uniform suboptimal-arm boost imply the coefficient-aware refined corruption window. theorem compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource_of_refinedRegime","label":"finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource_of_refinedRegime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource_of_refinedRegime","description":"The named uniform-boost refined regime supplies the model-facing compact corruption window without exposing its three clauses separately.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-388986333f9a","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9880,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:149"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource_of_refinedRegime {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (epsilon : Real) (hepsilon : 0 <= epsilon) (hregime : finiteArmIIDUniformSuboptimalBoostRefinedRegime model horizon epsilon) : finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow model horizon (uniformSuboptimalRewardBoostSource model epsilon hepsilon)","missing":[],"search":"finitearmiidhistoryadaptiverefinedcorruptionwindow_uniformsuboptimalrewardboostsource_of_refinedregime banditrlproof.tsallis.finitearmiidhistoryadaptiverefinedcorruptionwindow_uniformsuboptimalrewardboostsource_of_refinedregime the named uniform-boost refined regime supplies the model-facing compact corruption window without exposing its three clauses separately. theorem compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIIDUniformSuboptimalBoostAllRegimeBound","label":"finiteArmIIDUniformSuboptimalBoostAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIIDUniformSuboptimalBoostAllRegimeBound","description":"A total explicit bound: use the refined square-root branch inside its coefficient-aware window and the logarithmic additive-budget branch everywhere else. In particular, the fallback covers zero and small corruption.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-77d5bbd095c2","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9881,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:164"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIIDUniformSuboptimalBoostAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (epsilon : Real) : Real","missing":[],"search":"finitearmiiduniformsuboptimalboostallregimebound banditrlproof.tsallis.finitearmiiduniformsuboptimalboostallregimebound a total explicit bound: use the refined square-root branch inside its coefficient-aware window and the logarithmic additive-budget branch everywhere else. in particular, the fallback covers zero and small corruption. definition compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_refinedLocalExplicit","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_refinedLocalExplicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_refinedLocalExplicit","description":"Refined local regret for the concrete corruption process that uniformly boosts every suboptimal arm by `epsilon` at every round. The corruption budget is exposed explicitly rather than through the abstract source envelope.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-28243c874566","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9882,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:186"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_refinedLocalExplicit {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (epsilon : Real) (hepsilon : 0 <= epsilon) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) (hhorizon : 25 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => 1 / ((model.gap arm : Rat) : Real))) ^ 2 <= (…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiiduniformsuboptimalboostrewardlawregret_le_refinedlocalexplicit banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiiduniformsuboptimalboostrewardlawregret_le_refinedlocalexplicit refined local regret for the concrete corruption process that uniformly boosts every suboptimal arm by `epsilon` at every round. the corruption budget is exposed explicitly rather than through the abstract source envelope. theorem compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_allRegimes","description":"Uniform suboptimal-arm boost regret for every nonnegative `epsilon` and finite horizon. The theorem selects the refined local branch when its named regime holds and otherwise falls back to the compiled logarithmic theorem.","url":"../modules/banditrlproof-tsallisfinitearmiiduniformsuboptimalboostrefinedregret/index.html#decl-0169c2b735a4","parent":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","order":9883,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret.lean:275"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Fin K -> Measure Rat) (hprob : forall arm, IsProbabilityMeasure (armLaw arm)) (hbound : forall arm, ∀ᵐ reward ∂armLaw arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall arm, integral (armLaw arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (epsilon : Real) (hepsilon : 0 <= epsilon) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmiiduniformsuboptimalboostrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmiiduniformsuboptimalboostrewardlawregret_le_allregimes uniform suboptimal-arm boost regret for every nonnegative `epsilon` and finite horizon. the theorem selects the refined local branch when its named regime holds and otherwise falls back to the compiled logarithmic theorem. theorem compiled","shard":"modules/f6fddcb46ffd96bb.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanAllRegimeBound","label":"finiteArmIndependentDriftingMeanAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanAllRegimeBound","description":"Total explicit fixed-comparator bound for independent nonidentical reward laws with drifting means. The refined expression is used exactly inside its coefficient-aware window; the logarithmic explicit-budget expression is used on the complement.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanallregimes/index.html#decl-19c74b967af4","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","order":9884,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanAllRegimes.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentDriftingMeanAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmindependentdriftingmeanallregimebound banditrlproof.tsallis.finitearmindependentdriftingmeanallregimebound total explicit fixed-comparator bound for independent nonidentical reward laws with drifting means. the refined expression is used exactly inside its coefficient-aware window; the logarithmic explicit-budget expression is used on the complement. definition compiled","shard":"modules/51c9062786bf444c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanAllRegimeBound_fin_one","label":"finiteArmIndependentDriftingMeanAllRegimeBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanAllRegimeBound_fin_one","description":"With one arm there is no suboptimal coordinate, so the all-regimes envelope reduces to the logarithmic base term.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanallregimes/index.html#decl-44ad62354182","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","order":9885,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanAllRegimes.lean:46"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentDriftingMeanAllRegimeBound_fin_one (model : FiniteBanditModel 1) (horizon : Nat) (meanDeviation : Nat -> Fin 1 -> Real) : finiteArmIndependentDriftingMeanAllRegimeBound model horizon meanDeviation = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmindependentdriftingmeanallregimebound_fin_one banditrlproof.tsallis.finitearmindependentdriftingmeanallregimebound_fin_one with one arm there is no suboptimal coordinate, so the all-regimes envelope reduces to the logarithmic base term. theorem compiled","shard":"modules/51c9062786bf444c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_allRegimes","description":"Generated scheduled half-Tsallis regret against the fixed baseline comparator for every deterministic mean-deviation envelope and finite horizon. The theorem automatically selects the compact-window refined branch or the logarithmic fallback and requires no caller window proof. This is not dynamic regret.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanallregimes/index.html#decl-2314af934a44","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","order":9886,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanAllRegimes.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlawregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlawregret_le_allregimes generated scheduled half-tsallis regret against the fixed baseline comparator for every deterministic mean-deviation envelope and finite horizon. the theorem automatically selects the compact-window refined branch or the logarithmic fallback and requires no caller window proof. this is not dynamic regret. theorem compiled","shard":"modules/51c9062786bf444c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret","label":"sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret","description":"Predictable environment regret against a deterministic comparator that may change with the round.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-3c033d66da0e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9887,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (comparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret banditrlproof.tsallis.sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret predictable environment regret against a deterministic comparator that may change with the round. definition compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","label":"sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","description":"Moving-comparator regret is fixed-comparator regret plus the cumulative loss advantage of the moving comparator over the fixed arm.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-bf7c0126a0f0","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9888,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (comparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta loss comparator horizon sample = sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon sample + (Finset.range (horizon + 1)).sum (fun t => Exp3.predictableLossAt loss t sample best - Exp3.predictableLossAt loss t sample (comparator t))","missing":[],"search":"sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_eq_fixed_add banditrlproof.tsallis.sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_eq_fixed_add moving-comparator regret is fixed-comparator regret plus the cumulative loss advantage of the moving comparator over the fixed arm. theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentBestArmAt","label":"finiteArmIndependentBestArmAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentBestArmAt","description":"An actual-mean maximizing arm at round `t`. The finite action space is nonempty because it comes from a `FiniteBanditModel`.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-3ce8c062f25a","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9889,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:66"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentBestArmAt {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) : Fin K","missing":[],"search":"finitearmindependentbestarmat banditrlproof.tsallis.finitearmindependentbestarmat an actual-mean maximizing arm at round `t`. the finite action space is nonempty because it comes from a `finitebanditmodel`. definition compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_le_bestArmAt","label":"finiteArmIndependentRewardMean_le_bestArmAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardMean_le_bestArmAt","description":"The selected dynamic arm maximizes the actual roundwise reward mean.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-3145ec624fec","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9890,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:75"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentRewardMean_le_bestArmAt {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : finiteArmIndependentRewardMean armLaw t arm <= finiteArmIndependentRewardMean armLaw t (finiteArmIndependentBestArmAt model armLaw t)","missing":[],"search":"finitearmindependentrewardmean_le_bestarmat banditrlproof.tsallis.finitearmindependentrewardmean_le_bestarmat the selected dynamic arm maximizes the actual roundwise reward mean. theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty","label":"finiteArmIndependentDynamicComparatorPenalty","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty","description":"Extra envelope budget needed to move from the fixed model comparator to the actual-mean maximizing arm at every included round.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-4cb0e48f1a5e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9891,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:90"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentDynamicComparatorPenalty {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmindependentdynamiccomparatorpenalty banditrlproof.tsallis.finitearmindependentdynamiccomparatorpenalty extra envelope budget needed to move from the fixed model comparator to the actual-mean maximizing arm at every included round. definition compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty_fin_one","label":"finiteArmIndependentDynamicComparatorPenalty_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty_fin_one","description":"With one arm the dynamic comparator cannot move, so its extra penalty is zero.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-c46eb4993dac","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9892,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:104"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentDynamicComparatorPenalty_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) (meanDeviation : Nat -> Fin 1 -> Real) : finiteArmIndependentDynamicComparatorPenalty model armLaw horizon meanDeviation = 0","missing":[],"search":"finitearmindependentdynamiccomparatorpenalty_fin_one banditrlproof.tsallis.finitearmindependentdynamiccomparatorpenalty_fin_one with one arm the dynamic comparator cannot move, so its extra penalty is zero. theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_sub_bestArm_le_meanDeviation","label":"finiteArmIndependentRewardMean_sub_bestArm_le_meanDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardMean_sub_bestArm_le_meanDeviation","description":"The actual-mean advantage of any arm over the fixed baseline best arm is controlled by the two corresponding mean-deviation envelopes.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-e63c3e4ab4a2","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9893,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:117"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentRewardMean_sub_bestArm_le_meanDeviation {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (t : Nat) (arm : Fin K) : finiteArmIndependentRewardMean armLaw t arm - finiteArmIndependentRewardMean armLaw t model.bestArm <= meanDeviation t arm + meanDeviation t model.bestArm","missing":[],"search":"finitearmindependentrewardmean_sub_bestarm_le_meandeviation banditrlproof.tsallis.finitearmindependentrewardmean_sub_bestarm_le_meandeviation the actual-mean advantage of any arm over the fixed baseline best arm is controlled by the two corresponding mean-deviation envelopes. theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentMovingComparatorRewardLawRegret_eq_fixed_add_meanAdvantage","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentMovingComparatorRewardLawRegret_eq_fixed_add_meanAdvantage","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentMovingComparatorRewardLawRegret_eq_fixed_add_meanAdvantage","description":"For the concrete independent nonidentical generated law, integrated moving-comparator regret is fixed-baseline regret plus the exact cumulative actual-mean advantage of the comparator.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-e14395285ca2","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9894,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:155"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentMovingComparatorRewardLawRegret_eq_fixed_add_meanAdvantage {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (comparator : Nat -> Fin K) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentmovingcomparatorrewardlawregret_eq_fixed_add_meanadvantage banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentmovingcomparatorrewardlawregret_eq_fixed_add_meanadvantage for the concrete independent nonidentical generated law, integrated moving-comparator regret is fixed-baseline regret plus the exact cumulative actual-mean advantage of the comparator. theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanDynamicAllRegimeBound","label":"finiteArmIndependentDriftingMeanDynamicAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanDynamicAllRegimeBound","description":"Dynamic all-regimes bound: fixed-baseline all-regimes regret plus the explicit envelope cost of following the actual-mean maximizing arm.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-029bc0d0874b","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9895,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:308"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentDriftingMeanDynamicAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmindependentdriftingmeandynamicallregimebound banditrlproof.tsallis.finitearmindependentdriftingmeandynamicallregimebound dynamic all-regimes bound: fixed-baseline all-regimes regret plus the explicit envelope cost of following the actual-mean maximizing arm. definition compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanDynamicAllRegimeBound_fin_one","label":"finiteArmIndependentDriftingMeanDynamicAllRegimeBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanDynamicAllRegimeBound_fin_one","description":"theorem finiteArmIndependentDriftingMeanDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) (meanDeviation : Nat -> Fin 1 -> Real) : finiteArmIndependentDriftingMeanDynamicAllRegimeBound model armLaw horizon meanDeviation = 1 + Real.log (((horizon + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-9c22a7494b45","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9896,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:318"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentDriftingMeanDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) (meanDeviation : Nat -> Fin 1 -> Real) : finiteArmIndependentDriftingMeanDynamicAllRegimeBound model armLaw horizon meanDeviation = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmindependentdriftingmeandynamicallregimebound_fin_one banditrlproof.tsallis.finitearmindependentdriftingmeandynamicallregimebound_fin_one theorem finitearmindependentdriftingmeandynamicallregimebound_fin_one (model : finitebanditmodel 1) (armlaw : nat -> fin 1 -> measure rat) (horizon : nat) (meandeviation : nat -> fin 1 -> real) : finitearmindependentdriftingmeandynamicallregimebound model armlaw horizon meandeviation = 1 + real.log (((horizon + 1 : nat) : real)) theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes","description":"Generated scheduled half-Tsallis dynamic regret against the arm with the largest actual reward mean at each round. No caller supplies the dynamic comparator, a refined-window proof, or a nonempty-suboptimal-arm proof.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeandynamicregret/index.html#decl-995fe6ce3b52","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","order":9897,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanDynamicRegret.lean:330"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","delayed-nonstationary"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeandynamicregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeandynamicregret_le_allregimes generated scheduled half-tsallis dynamic regret against the arm with the largest actual reward mean at each round. no caller supplies the dynamic comparator, a refined-window proof, or a nonempty-suboptimal-arm proof. theorem compiled","shard":"modules/28ce4068a99827b0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanRefinedCorruptionWindow","label":"finiteArmIndependentDriftingMeanRefinedCorruptionWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanRefinedCorruptionWindow","description":"The coefficient-aware refined corruption window specialized to the explicit mean-deviation budget.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrefinedregret/index.html#decl-21ef0bcf7b29","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","order":9898,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRefinedRegret.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentDriftingMeanRefinedCorruptionWindow {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : Prop","missing":[],"search":"finitearmindependentdriftingmeanrefinedcorruptionwindow banditrlproof.tsallis.finitearmindependentdriftingmeanrefinedcorruptionwindow the coefficient-aware refined corruption window specialized to the explicit mean-deviation budget. definition compiled","shard":"modules/55162eabfb656bf3.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_refinedLocalExplicit_of_window","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_refinedLocalExplicit_of_window","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_refinedLocalExplicit_of_window","description":"Refined local square-root regret against the fixed baseline comparator for independent, nonidentical reward laws whose means drift within an explicit armwise envelope. This is not dynamic regret.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrefinedregret/index.html#decl-f476f5c141ff","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","order":9899,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRefinedRegret.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_refinedLocalExplicit_of_window {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (hsuboptimal : ((Finset.univ : Finset (Fin K)).erase model.bestArm).Nonempty) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) (hwindow : finiteArmIndependentDriftingMeanRefinedCorruptionWindow model horizon meanDeviation) : let…","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlawregret_le_refinedlocalexplicit_of_window banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlawregret_le_refinedlocalexplicit_of_window refined local square-root regret against the fixed baseline comparator for independent, nonidentical reward laws whose means drift within an explicit armwise envelope. this is not dynamic regret. theorem compiled","shard":"modules/55162eabfb656bf3.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean","label":"finiteArmIndependentRewardMean","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardMean","description":"The actual mean reward of arm `arm` at round `t`.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-4d2bb200f38d","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9900,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:18"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentRewardMean {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"finitearmindependentrewardmean banditrlproof.tsallis.finitearmindependentrewardmean the actual mean reward of arm `arm` at round `t`. definition compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_mean_sub","label":"independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_mean_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_mean_sub","description":"Under the unit-support contract, the roundwise product-law loss gap is the difference between the actual mean rewards of the best and selected arms.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-2d57941d6bfd","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9901,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_mean_sub {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (best arm : Fin K) : independentLossStateTimeVaryingMeanGap (finiteArmIndependentRewardVectorLaw armLaw) (fun _ => finiteArmIIDRewardVectorLoss) t best arm = finiteArmIndependentRewardMean armLaw t best - finiteArmIndependentRewardMean armLaw t arm","missing":[],"search":"independentlossstatetimevaryingmeangap_finitearmindependentrewardvectorloss_eq_mean_sub banditrlproof.tsallis.independentlossstatetimevaryingmeangap_finitearmindependentrewardvectorloss_eq_mean_sub under the unit-support contract, the roundwise product-law loss gap is the difference between the actual mean rewards of the best and selected arms. theorem compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_independentLossStateTimeVaryingMeanGap_sub_modelGap_le","label":"abs_independentLossStateTimeVaryingMeanGap_sub_modelGap_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_independentLossStateTimeVaryingMeanGap_sub_modelGap_le","description":"The actual loss gap stays within the sum of the two supplied arm-mean deviation envelopes from the fixed model gap.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-c5f3d4860500","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9902,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:81"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_independentLossStateTimeVaryingMeanGap_sub_modelGap_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (t : Nat) (arm : Fin K) (harm : arm ≠ model.bestArm) : |independentLossStateTimeVaryingMeanGap (finiteArmIndependentRewardVectorLaw armLaw) (fun _ => finiteArmIIDRewardVectorLoss) t model.bestArm arm - ((model.gap arm : Rat) : Real)| <= meanDeviation t arm + meanDeviation t model.bestArm","missing":[],"search":"abs_independentlossstatetimevaryingmeangap_sub_modelgap_le banditrlproof.tsallis.abs_independentlossstatetimevaryingmeangap_sub_modelgap_le the actual loss gap stays within the sum of the two supplied arm-mean deviation envelopes from the fixed model gap. theorem compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget","label":"finiteArmIndependentMeanDeviationBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget","description":"Explicit accumulated gap-deviation budget induced by armwise drifting means through the inclusive horizon.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-817344452f5e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9903,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentMeanDeviationBudget {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : Real","missing":[],"search":"finitearmindependentmeandeviationbudget banditrlproof.tsallis.finitearmindependentmeandeviationbudget explicit accumulated gap-deviation budget induced by armwise drifting means through the inclusive horizon. definition compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_eq","label":"finiteArmIndependentMeanDeviationBudget_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_eq","description":"theorem finiteArmIndependentMeanDeviationBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : finiteArmIndependentMeanDeviationBudget model horizon meanDeviation = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => meanDeviation t arm + meanDeviation t model.bestArm))","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-a62ce5f94e76","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9904,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:134"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentMeanDeviationBudget_eq {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) (meanDeviation : Nat -> Fin K -> Real) : finiteArmIndependentMeanDeviationBudget model horizon meanDeviation = (Finset.range (horizon + 1)).sum (fun t => ((Finset.univ : Finset (Fin K)).erase model.bestArm).sum (fun arm => meanDeviation t arm + meanDeviation t model.bestArm))","missing":[],"search":"finitearmindependentmeandeviationbudget_eq banditrlproof.tsallis.finitearmindependentmeandeviationbudget_eq theorem finitearmindependentmeandeviationbudget_eq {k : nat} (model : finitebanditmodel k) (horizon : nat) (meandeviation : nat -> fin k -> real) : finitearmindependentmeandeviationbudget model horizon meandeviation = (finset.range (horizon + 1)).sum (fun t => ((finset.univ : finset (fin k)).erase model.bestarm).sum (fun arm => meandeviation t arm + meandeviation t model.bestarm)) theorem compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_zero","label":"finiteArmIndependentMeanDeviationBudget_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_zero","description":"theorem finiteArmIndependentMeanDeviationBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIndependentMeanDeviationBudget model horizon (fun _ _ => 0) = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-31f50339e81a","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9905,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:145"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentMeanDeviationBudget_zero {K : Nat} (model : FiniteBanditModel K) (horizon : Nat) : finiteArmIndependentMeanDeviationBudget model horizon (fun _ _ => 0) = 0","missing":[],"search":"finitearmindependentmeandeviationbudget_zero banditrlproof.tsallis.finitearmindependentmeandeviationbudget_zero theorem finitearmindependentmeandeviationbudget_zero {k : nat} (model : finitebanditmodel k) (horizon : nat) : finitearmindependentmeandeviationbudget model horizon (fun _ _ => 0) = 0 theorem compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLaw_hasSelfBounding","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLaw_hasSelfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLaw_hasSelfBounding","description":"Independent nonidentical reward laws with an armwise mean-deviation envelope supply the terminal self-bound consumed by both logarithmic and refined square-root schedule routes.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-e5e74d4373fc","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9906,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:155"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLaw_hasSelfBounding {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlaw_hasselfbounding banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlaw_hasselfbounding independent nonidentical reward laws with an armwise mean-deviation envelope supply the terminal self-bound consumed by both logarithmic and refined square-root schedule routes. theorem compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_log","description":"Generated scheduled half-Tsallis logarithmic regret against the fixed baseline comparator `model.bestArm` for independent, nonidentical finite-arm reward laws whose means drift within an explicit coordinatewise envelope. This is a static-comparator theorem, not dynamic regret against each round's best actual mean.","url":"../modules/banditrlproof-tsallisfinitearmindependentdriftingmeanrewardlaw/index.html#decl-342348bd8d36","parent":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","order":9907,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentDriftingMeanRewardLaw.lean:277"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (meanDeviation : Nat -> Fin K -> Real) (hmeanDeviation : forall t arm, |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= meanDeviation t arm) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentdriftingmeanrewardlawregret_le_log generated scheduled half-tsallis logarithmic regret against the fixed baseline comparator `model.bestarm` for independent, nonidentical finite-arm reward laws whose means drift within an explicit coordinatewise envelope. this is a static-comparator theorem, not dynamic regret against each round's best actual mean. theorem compiled","shard":"modules/3309325c46f3735e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_nonneg","label":"finiteArmIndependentCumulativeGlobalMeanSwitchCount_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_nonneg","description":"The global population-mean switch count is nonnegative.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-6ab97a8ce376","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9908,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:19"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeGlobalMeanSwitchCount_nonneg {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) : 0 <= finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t","missing":[],"search":"finitearmindependentcumulativeglobalmeanswitchcount_nonneg banditrlproof.tsallis.finitearmindependentcumulativeglobalmeanswitchcount_nonneg the global population-mean switch count is nonnegative. theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_mono","label":"finiteArmIndependentCumulativeGlobalMeanSwitchCount_mono","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_mono","description":"Enlarging the prefix cannot decrease the global switch count.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-3aca4dba19c2","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9909,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:29"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeGlobalMeanSwitchCount_mono {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) {t u : Nat} (htu : t <= u) : finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t <= finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw u","missing":[],"search":"finitearmindependentcumulativeglobalmeanswitchcount_mono banditrlproof.tsallis.finitearmindependentcumulativeglobalmeanswitchcount_mono enlarging the prefix cannot decrease the global switch count. theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_globalMeanSwitchCount_le","label":"finiteArmIndependentMeanDeviationBudget_globalMeanSwitchCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_globalMeanSwitchCount_le","description":"The fixed-comparator mean-deviation budget at the global prefix envelope is bounded by the terminal count times horizon mass and suboptimal-arm cardinality.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-45c14ac0c48c","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9910,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:42"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentMeanDeviationBudget_globalMeanSwitchCount_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : finiteArmIndependentMeanDeviationBudget model horizon (fun t _ => finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t) <= 2 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * (((horizon + 1 : Nat) : Real)) * finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw horizon","missing":[],"search":"finitearmindependentmeandeviationbudget_globalmeanswitchcount_le banditrlproof.tsallis.finitearmindependentmeandeviationbudget_globalmeanswitchcount_le the fixed-comparator mean-deviation budget at the global prefix envelope is bounded by the terminal count times horizon mass and suboptimal-arm cardinality. theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty_globalMeanSwitchCount_le","label":"finiteArmIndependentDynamicComparatorPenalty_globalMeanSwitchCount_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty_globalMeanSwitchCount_le","description":"The moving-comparator penalty at the global prefix envelope is bounded by the same terminal-count expression as the fixed-comparator deviation budget. The erased-arm cardinality makes the bound exactly zero for `Fin 1`.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-81c98f608505","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9911,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:94"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentDynamicComparatorPenalty_globalMeanSwitchCount_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : finiteArmIndependentDynamicComparatorPenalty model armLaw horizon (fun t _ => finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t) <= 2 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * (((horizon + 1 : Nat) : Real)) * finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw horizon","missing":[],"search":"finitearmindependentdynamiccomparatorpenalty_globalmeanswitchcount_le banditrlproof.tsallis.finitearmindependentdynamiccomparatorpenalty_globalmeanswitchcount_le the moving-comparator penalty at the global prefix envelope is bounded by the same terminal-count expression as the fixed-comparator deviation budget. the erased-arm cardinality makes the bound exactly zero for `fin 1`. theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCount_totalBudget_le","label":"finiteArmIndependentGlobalMeanSwitchCount_totalBudget_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCount_totalBudget_le","description":"One terminal global switch count controls both the fixed-comparator mean-deviation budget and the moving-comparator penalty.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-dcdbb4cf1fff","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9912,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:165"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentGlobalMeanSwitchCount_totalBudget_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : finiteArmIndependentMeanDeviationBudget model horizon (fun t _ => finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t) + finiteArmIndependentDynamicComparatorPenalty model armLaw horizon (fun t _ => finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t) <= 4 * (((Finset.univ : Finset (Fin K)).erase model.bestArm).card : Real) * (((horizon + 1 : Nat) : Real)) * finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw horizon","missing":[],"search":"finitearmindependentglobalmeanswitchcount_totalbudget_le banditrlproof.tsallis.finitearmindependentglobalmeanswitchcount_totalbudget_le one terminal global switch count controls both the fixed-comparator mean-deviation budget and the moving-comparator penalty. theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound","label":"finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound","description":"Explicit logarithmic dynamic-regret bound with a single terminal global population-mean switch count.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-edd4dc6029d5","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9913,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:189"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : Real","missing":[],"search":"finitearmindependentglobalmeanswitchcounthorizoncompressedlogdynamicbound banditrlproof.tsallis.finitearmindependentglobalmeanswitchcounthorizoncompressedlogdynamicbound explicit logarithmic dynamic-regret bound with a single terminal global population-mean switch count. definition compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound_fin_one","label":"finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound_fin_one","description":"theorem finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-9212560d3331","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9914,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:202"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmindependentglobalmeanswitchcounthorizoncompressedlogdynamicbound_fin_one banditrlproof.tsallis.finitearmindependentglobalmeanswitchcounthorizoncompressedlogdynamicbound_fin_one theorem finitearmindependentglobalmeanswitchcounthorizoncompressedlogdynamicbound_fin_one (model : finitebanditmodel 1) (armlaw : nat -> fin 1 -> measure rat) (horizon : nat) : finitearmindependentglobalmeanswitchcounthorizoncompressedlogdynamicbound model armlaw horizon = 1 + real.log (((horizon + 1 : nat) : real)) theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountHorizonCompressedDynamicRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountHorizonCompressedDynamicRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountHorizonCompressedDynamicRegret_le_log","description":"Generated expected predictable-environment dynamic regret with all time-indexed population-mean switch envelopes compressed into the terminal global count. This is a linear horizon-level compression, not a minimax switch-rate theorem.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountcompresseddynamicregret/index.html#decl-42fa27646f5f","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","order":9915,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret.lean:218"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountHorizonCompressedDynamicRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentglobalmeanswitchcounthorizoncompresseddynamicregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentglobalmeanswitchcounthorizoncompresseddynamicregret_le_log generated expected predictable-environment dynamic regret with all time-indexed population-mean switch envelopes compressed into the terminal global count. this is a linear horizon-level compression, not a minimax switch-rate theorem. theorem compiled","shard":"modules/d1f98a4d6221487f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount","label":"finiteArmIndependentCumulativeGlobalMeanSwitchCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount","description":"Real-valued prefix count of rounds before `t` at which at least one arm's population mean changes.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-cd0356eff39e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9916,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:19"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentCumulativeGlobalMeanSwitchCount {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) : Real","missing":[],"search":"finitearmindependentcumulativeglobalmeanswitchcount banditrlproof.tsallis.finitearmindependentcumulativeglobalmeanswitchcount real-valued prefix count of rounds before `t` at which at least one arm's population mean changes. definition compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_zero","label":"finiteArmIndependentCumulativeGlobalMeanSwitchCount_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_zero","description":"theorem finiteArmIndependentCumulativeGlobalMeanSwitchCount_zero {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) : finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw 0 = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-9d447babb326","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9917,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeGlobalMeanSwitchCount_zero {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) : finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw 0 = 0","missing":[],"search":"finitearmindependentcumulativeglobalmeanswitchcount_zero banditrlproof.tsallis.finitearmindependentcumulativeglobalmeanswitchcount_zero theorem finitearmindependentcumulativeglobalmeanswitchcount_zero {k : nat} (armlaw : nat -> fin k -> measure rat) : finitearmindependentcumulativeglobalmeanswitchcount armlaw 0 = 0 theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_card","label":"finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_card","description":"The global real-valued count is the coercion of the filtered cardinality of rounds with at least one population-mean change.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-910c1b850b3d","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9918,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_card {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) : finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t = ((((Finset.range t).filter (fun s => ∃ arm : Fin K, finiteArmIndependentRewardMean armLaw (s + 1) arm ≠ finiteArmIndependentRewardMean armLaw s arm)).card : Nat) : Real)","missing":[],"search":"finitearmindependentcumulativeglobalmeanswitchcount_eq_card banditrlproof.tsallis.finitearmindependentcumulativeglobalmeanswitchcount_eq_card the global real-valued count is the coercion of the filtered cardinality of rounds with at least one population-mean change. theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_le_globalMeanSwitchCount","label":"finiteArmIndependentCumulativeMeanSwitchCount_le_globalMeanSwitchCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_le_globalMeanSwitchCount","description":"Every armwise prefix switch count is bounded by the global count.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-e0b3977d74a5","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9919,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeMeanSwitchCount_le_globalMeanSwitchCount {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : finiteArmIndependentCumulativeMeanSwitchCount armLaw t arm <= finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t","missing":[],"search":"finitearmindependentcumulativemeanswitchcount_le_globalmeanswitchcount banditrlproof.tsallis.finitearmindependentcumulativemeanswitchcount_le_globalmeanswitchcount every armwise prefix switch count is bounded by the global count. theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_le_globalMeanSwitchCount","label":"finiteArmIndependentCumulativeMeanPathVariation_le_globalMeanSwitchCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_le_globalMeanSwitchCount","description":"Unit-supported cumulative path variation of any arm is bounded by the single global population-mean switch count.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-e8c93f659cc7","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9920,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:76"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeMeanPathVariation_le_globalMeanSwitchCount {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (arm : Fin K) : finiteArmIndependentCumulativeMeanPathVariation armLaw t arm <= finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t","missing":[],"search":"finitearmindependentcumulativemeanpathvariation_le_globalmeanswitchcount banditrlproof.tsallis.finitearmindependentcumulativemeanpathvariation_le_globalmeanswitchcount unit-supported cumulative path variation of any arm is bounded by the single global population-mean switch count. theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_globalMeanSwitchCount","label":"abs_finiteArmIndependentRewardMean_sub_model_le_globalMeanSwitchCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_globalMeanSwitchCount","description":"Initial model matching turns the global prefix switch count into the all-time deviation envelope required by the dynamic-regret theorem.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-e6f120875178","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9921,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:92"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_finiteArmIndependentRewardMean_sub_model_le_globalMeanSwitchCount {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (t : Nat) (arm : Fin K) : |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw t","missing":[],"search":"abs_finitearmindependentrewardmean_sub_model_le_globalmeanswitchcount banditrlproof.tsallis.abs_finitearmindependentrewardmean_sub_model_le_globalmeanswitchcount initial model matching turns the global prefix switch count into the all-time deviation envelope required by the dynamic-regret theorem. theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound","label":"finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound","description":"Dynamic all-regimes bound specialized to the exact global prefix population-mean switch count.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-ba3079cdc6d4","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9922,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:113"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : Real","missing":[],"search":"finitearmindependentglobalmeanswitchcountdynamicallregimebound banditrlproof.tsallis.finitearmindependentglobalmeanswitchcountdynamicallregimebound dynamic all-regimes bound specialized to the exact global prefix population-mean switch count. definition compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound_fin_one","label":"finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound_fin_one","description":"theorem finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-87774fee2d12","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9923,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:122"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmindependentglobalmeanswitchcountdynamicallregimebound_fin_one banditrlproof.tsallis.finitearmindependentglobalmeanswitchcountdynamicallregimebound_fin_one theorem finitearmindependentglobalmeanswitchcountdynamicallregimebound_fin_one (model : finitebanditmodel 1) (armlaw : nat -> fin 1 -> measure rat) (horizon : nat) : finitearmindependentglobalmeanswitchcountdynamicallregimebound model armlaw horizon = 1 + real.log (((horizon + 1 : nat) : real)) theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret_le_allRegimes","description":"Generated expected predictable-environment dynamic regret with the all-time deviation envelope derived from one global prefix count of population-mean change-points. No caller supplies a comparator, variation family, armwise switch budget, or global switch budget. This is not a minimax or horizon-compressed switch-rate theorem.","url":"../modules/banditrlproof-tsallisfinitearmindependentglobalmeanswitchcountdynamicregret/index.html#decl-fb60fa893a84","parent":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","order":9924,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret.lean:135"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentglobalmeanswitchcountdynamicregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentglobalmeanswitchcountdynamicregret_le_allregimes generated expected predictable-environment dynamic regret with the all-time deviation envelope derived from one global prefix count of population-mean change-points. no caller supplies a comparator, variation family, armwise switch budget, or global switch budget. this is not a minimax or horizon-compressed switch-rate theorem. theorem compiled","shard":"modules/ece891261c1097c5.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_mem_Icc","label":"finiteArmIndependentRewardMean_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardMean_mem_Icc","description":"A bounded probability reward law has its population mean in the unit interval.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-8b7f66c7f967","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9925,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentRewardMean_mem_Icc {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (arm : Fin K) : finiteArmIndependentRewardMean armLaw t arm ∈ Set.Icc (0 : Real) 1","missing":[],"search":"finitearmindependentrewardmean_mem_icc banditrlproof.tsallis.finitearmindependentrewardmean_mem_icc a bounded probability reward law has its population mean in the unit interval. theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount","label":"finiteArmIndependentCumulativeMeanSwitchCount","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount","description":"Real-valued prefix count of the rounds before `t` at which one arm's population mean changes.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-83882c0f8e8f","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9926,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentCumulativeMeanSwitchCount {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"finitearmindependentcumulativemeanswitchcount banditrlproof.tsallis.finitearmindependentcumulativemeanswitchcount real-valued prefix count of the rounds before `t` at which one arm's population mean changes. definition compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_zero","label":"finiteArmIndependentCumulativeMeanSwitchCount_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_zero","description":"theorem finiteArmIndependentCumulativeMeanSwitchCount_zero {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (arm : Fin K) : finiteArmIndependentCumulativeMeanSwitchCount armLaw 0 arm = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-ff208b4033a0","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9927,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeMeanSwitchCount_zero {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (arm : Fin K) : finiteArmIndependentCumulativeMeanSwitchCount armLaw 0 arm = 0","missing":[],"search":"finitearmindependentcumulativemeanswitchcount_zero banditrlproof.tsallis.finitearmindependentcumulativemeanswitchcount_zero theorem finitearmindependentcumulativemeanswitchcount_zero {k : nat} (armlaw : nat -> fin k -> measure rat) (arm : fin k) : finitearmindependentcumulativemeanswitchcount armlaw 0 arm = 0 theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_eq_card","label":"finiteArmIndependentCumulativeMeanSwitchCount_eq_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_eq_card","description":"The real-valued switch count is exactly the coercion of the filtered cardinality of nonzero consecutive population-mean changes.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-6f767e2ae5b8","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9928,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:68"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeMeanSwitchCount_eq_card {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : finiteArmIndependentCumulativeMeanSwitchCount armLaw t arm = ((((Finset.range t).filter (fun s => finiteArmIndependentRewardMean armLaw (s + 1) arm ≠ finiteArmIndependentRewardMean armLaw s arm)).card : Nat) : Real)","missing":[],"search":"finitearmindependentcumulativemeanswitchcount_eq_card banditrlproof.tsallis.finitearmindependentcumulativemeanswitchcount_eq_card the real-valued switch count is exactly the coercion of the filtered cardinality of nonzero consecutive population-mean changes. theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_le_switchCount","label":"finiteArmIndependentCumulativeMeanPathVariation_le_switchCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_le_switchCount","description":"Unit-supported reward laws make every nonzero population-mean jump at most one, so cumulative mean path variation is bounded by switch count.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-0ed2000edd7d","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9929,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:81"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeMeanPathVariation_le_switchCount {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (arm : Fin K) : finiteArmIndependentCumulativeMeanPathVariation armLaw t arm <= finiteArmIndependentCumulativeMeanSwitchCount armLaw t arm","missing":[],"search":"finitearmindependentcumulativemeanpathvariation_le_switchcount banditrlproof.tsallis.finitearmindependentcumulativemeanpathvariation_le_switchcount unit-supported reward laws make every nonzero population-mean jump at most one, so cumulative mean path variation is bounded by switch count. theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_switchCount","label":"abs_finiteArmIndependentRewardMean_sub_model_le_switchCount","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_switchCount","description":"Initial model matching turns the armwise prefix switch count into the all-time deviation envelope required by the dynamic-regret theorem.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-a8d52966218b","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9930,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:110"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_finiteArmIndependentRewardMean_sub_model_le_switchCount {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (t : Nat) (arm : Fin K) : |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= finiteArmIndependentCumulativeMeanSwitchCount armLaw t arm","missing":[],"search":"abs_finitearmindependentrewardmean_sub_model_le_switchcount banditrlproof.tsallis.abs_finitearmindependentrewardmean_sub_model_le_switchcount initial model matching turns the armwise prefix switch count into the all-time deviation envelope required by the dynamic-regret theorem. theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound","label":"finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound","description":"Dynamic all-regimes bound specialized to the exact armwise prefix population-mean switch count.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-8f47bc003465","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9931,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:131"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : Real","missing":[],"search":"finitearmindependentmeanswitchcountdynamicallregimebound banditrlproof.tsallis.finitearmindependentmeanswitchcountdynamicallregimebound dynamic all-regimes bound specialized to the exact armwise prefix population-mean switch count. definition compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound_fin_one","label":"finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound_fin_one","description":"theorem finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-32e1130fff59","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9932,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:139"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmindependentmeanswitchcountdynamicallregimebound_fin_one banditrlproof.tsallis.finitearmindependentmeanswitchcountdynamicallregimebound_fin_one theorem finitearmindependentmeanswitchcountdynamicallregimebound_fin_one (model : finitebanditmodel 1) (armlaw : nat -> fin 1 -> measure rat) (horizon : nat) : finitearmindependentmeanswitchcountdynamicallregimebound model armlaw horizon = 1 + real.log (((horizon + 1 : nat) : real)) theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentMeanSwitchCountDynamicRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentMeanSwitchCountDynamicRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentMeanSwitchCountDynamicRegret_le_allRegimes","description":"Generated expected predictable-environment dynamic regret with the all-time deviation envelope derived from armwise population-mean switch counts. No caller supplies a comparator, variation family, or switch budget. This is an exact prefix-envelope specialization, not a minimax change-point or horizon-compressed standard nonstationary rate.","url":"../modules/banditrlproof-tsallisfinitearmindependentmeanswitchcountdynamicregret/index.html#decl-04cd5b2e97ab","parent":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","order":9933,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret.lean:152"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentMeanSwitchCountDynamicRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentmeanswitchcountdynamicregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentmeanswitchcountdynamicregret_le_allregimes generated expected predictable-environment dynamic regret with the all-time deviation envelope derived from armwise population-mean switch counts. no caller supplies a comparator, variation family, or switch budget. this is an exact prefix-envelope specialization, not a minimax change-point or horizon-compressed standard nonstationary rate. theorem compiled","shard":"modules/35ce12786adfe8b9.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation","label":"finiteArmIndependentCumulativeMeanPathVariation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation","description":"Cumulative absolute variation of one arm's actual reward mean before round `t`.","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-9b530036bbe7","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9934,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentCumulativeMeanPathVariation {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : Real","missing":[],"search":"finitearmindependentcumulativemeanpathvariation banditrlproof.tsallis.finitearmindependentcumulativemeanpathvariation cumulative absolute variation of one arm's actual reward mean before round `t`. definition compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_zero","label":"finiteArmIndependentCumulativeMeanPathVariation_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_zero","description":"theorem finiteArmIndependentCumulativeMeanPathVariation_zero {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (arm : Fin K) : finiteArmIndependentCumulativeMeanPathVariation armLaw 0 arm = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-ee334003513e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9935,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeMeanPathVariation_zero {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (arm : Fin K) : finiteArmIndependentCumulativeMeanPathVariation armLaw 0 arm = 0","missing":[],"search":"finitearmindependentcumulativemeanpathvariation_zero banditrlproof.tsallis.finitearmindependentcumulativemeanpathvariation_zero theorem finitearmindependentcumulativemeanpathvariation_zero {k : nat} (armlaw : nat -> fin k -> measure rat) (arm : fin k) : finitearmindependentcumulativemeanpathvariation armlaw 0 arm = 0 theorem compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_zero_le_cumulativeMeanPathVariation","label":"abs_finiteArmIndependentRewardMean_sub_zero_le_cumulativeMeanPathVariation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_zero_le_cumulativeMeanPathVariation","description":"theorem abs_finiteArmIndependentRewardMean_sub_zero_le_cumulativeMeanPathVariation {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : |finiteArmIndependentRewardMean armLaw t arm - finiteArmIndependentRewardMean armLaw 0 arm| <= finiteArmIndependentCumulativeMeanPathVariation armLaw t arm","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-2a5eb372680a","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9936,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:32"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_finiteArmIndependentRewardMean_sub_zero_le_cumulativeMeanPathVariation {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : |finiteArmIndependentRewardMean armLaw t arm - finiteArmIndependentRewardMean armLaw 0 arm| <= finiteArmIndependentCumulativeMeanPathVariation armLaw t arm","missing":[],"search":"abs_finitearmindependentrewardmean_sub_zero_le_cumulativemeanpathvariation banditrlproof.tsallis.abs_finitearmindependentrewardmean_sub_zero_le_cumulativemeanpathvariation theorem abs_finitearmindependentrewardmean_sub_zero_le_cumulativemeanpathvariation {k : nat} (armlaw : nat -> fin k -> measure rat) (t : nat) (arm : fin k) : |finitearmindependentrewardmean armlaw t arm - finitearmindependentrewardmean armlaw 0 arm| <= finitearmindependentcumulativemeanpathvariation armlaw t arm theorem compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_cumulativeMeanPathVariation","label":"abs_finiteArmIndependentRewardMean_sub_model_le_cumulativeMeanPathVariation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_cumulativeMeanPathVariation","description":"Initial mean matching turns cumulative actual-mean path variation into the all-time model-deviation envelope required by the drifting-mean theorem.","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-5cab6fb07bca","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9937,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:72"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem abs_finiteArmIndependentRewardMean_sub_model_le_cumulativeMeanPathVariation {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (t : Nat) (arm : Fin K) : |finiteArmIndependentRewardMean armLaw t arm - ((model.mean arm : Rat) : Real)| <= finiteArmIndependentCumulativeMeanPathVariation armLaw t arm","missing":[],"search":"abs_finitearmindependentrewardmean_sub_model_le_cumulativemeanpathvariation banditrlproof.tsallis.abs_finitearmindependentrewardmean_sub_model_le_cumulativemeanpathvariation initial mean matching turns cumulative actual-mean path variation into the all-time model-deviation envelope required by the drifting-mean theorem. theorem compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentPathVariationDynamicAllRegimeBound","label":"finiteArmIndependentPathVariationDynamicAllRegimeBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentPathVariationDynamicAllRegimeBound","description":"The compiled dynamic all-regimes bound specialized to the law-derived cumulative population-mean path-variation envelope.","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-1b9e4a826873","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9938,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:91"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentPathVariationDynamicAllRegimeBound {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : Real","missing":[],"search":"finitearmindependentpathvariationdynamicallregimebound banditrlproof.tsallis.finitearmindependentpathvariationdynamicallregimebound the compiled dynamic all-regimes bound specialized to the law-derived cumulative population-mean path-variation envelope. definition compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentPathVariationDynamicAllRegimeBound_fin_one","label":"finiteArmIndependentPathVariationDynamicAllRegimeBound_fin_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentPathVariationDynamicAllRegimeBound_fin_one","description":"theorem finiteArmIndependentPathVariationDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentPathVariationDynamicAllRegimeBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-5c819416325a","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9939,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:99"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentPathVariationDynamicAllRegimeBound_fin_one (model : FiniteBanditModel 1) (armLaw : Nat -> Fin 1 -> Measure Rat) (horizon : Nat) : finiteArmIndependentPathVariationDynamicAllRegimeBound model armLaw horizon = 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"finitearmindependentpathvariationdynamicallregimebound_fin_one banditrlproof.tsallis.finitearmindependentpathvariationdynamicallregimebound_fin_one theorem finitearmindependentpathvariationdynamicallregimebound_fin_one (model : finitebanditmodel 1) (armlaw : nat -> fin 1 -> measure rat) (horizon : nat) : finitearmindependentpathvariationdynamicallregimebound model armlaw horizon = 1 + real.log (((horizon + 1 : nat) : real)) theorem compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentPathVariationDynamicRegret_le_allRegimes","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentPathVariationDynamicRegret_le_allRegimes","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentPathVariationDynamicRegret_le_allRegimes","description":"Generated expected predictable-environment dynamic regret with the all-time deviation envelope derived from the reward laws' population-mean path variation. No caller supplies a deviation envelope or the moving comparator. This retains one cumulative prefix envelope at every included time; it is not a horizon-compressed or minimax-sharp standard `V_T` bound.","url":"../modules/banditrlproof-tsallisfinitearmindependentpathvariationdynamicregret/index.html#decl-0d72cc30694f","parent":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","order":9940,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret"],["Source","BanditRLProof/TsallisFiniteArmIndependentPathVariationDynamicRegret.lean:112"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentPathVariationDynamicRegret_le_allRegimes {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hinitialMean : forall arm, finiteArmIndependentRewardMean armLaw 0 arm = ((model.mean arm : Rat) : Real)) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (hgapLeOne : forall arm, arm ≠ model.bestArm -> ((model.gap arm : Rat) : Real) <= 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentpathvariationdynamicregret_le_allregimes banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentpathvariationdynamicregret_le_allregimes generated expected predictable-environment dynamic regret with the all-time deviation envelope derived from the reward laws' population-mean path variation. no caller supplies a deviation envelope or the moving comparator. this retains one cumulative prefix envelope at every included time; it is not a horizon-compressed or minimax-sharp standard `v_t` bound. theorem compiled","shard":"modules/5aa0257adad200c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardVectorLaw","label":"finiteArmIndependentRewardVectorLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardVectorLaw","description":"The independent finite-arm reward-vector law used at round `t`.","url":"../modules/banditrlproof-tsallisfinitearmindependentrewardlaw/index.html#decl-4b7af391ed72","parent":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","order":9941,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentRewardLaw.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentRewardVectorLaw {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) : Measure (Fin K -> Rat)","missing":[],"search":"finitearmindependentrewardvectorlaw banditrlproof.tsallis.finitearmindependentrewardvectorlaw the independent finite-arm reward-vector law used at round `t`. definition compiled","shard":"modules/4555088178dcc4f4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_gap","label":"independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_gap","description":"Every roundwise loss gap has the fixed finite-bandit model gap when the possibly time-varying arm laws preserve the model means.","url":"../modules/banditrlproof-tsallisfinitearmindependentrewardlaw/index.html#decl-eb9baac3b577","parent":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","order":9942,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentRewardLaw.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_gap {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall t arm, integral (armLaw t arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (t : Nat) (arm : Fin K) : independentLossStateTimeVaryingMeanGap (finiteArmIndependentRewardVectorLaw armLaw) (fun _ => finiteArmIIDRewardVectorLoss) t model.bestArm arm = ((model.gap arm : Rat) : Real)","missing":[],"search":"independentlossstatetimevaryingmeangap_finitearmindependentrewardvectorloss_eq_gap banditrlproof.tsallis.independentlossstatetimevaryingmeangap_finitearmindependentrewardvectorloss_eq_gap every roundwise loss gap has the fixed finite-bandit model gap when the possibly time-varying arm laws preserve the model means. theorem compiled","shard":"modules/4555088178dcc4f4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledIndependentMeanGapLaw_of_finiteArmIndependentRewardVectorLaw","label":"hasScheduledIndependentMeanGapLaw_of_finiteArmIndependentRewardVectorLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledIndependentMeanGapLaw_of_finiteArmIndependentRewardVectorLaw","description":"Independent roundwise product reward vectors with fixed arm means supply the fixed model-gap law required by the scheduled self-bounding route.","url":"../modules/banditrlproof-tsallisfinitearmindependentrewardlaw/index.html#decl-59ebe32bc312","parent":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","order":9943,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentRewardLaw.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledIndependentMeanGapLaw_of_finiteArmIndependentRewardVectorLaw {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall t arm, integral (armLaw t arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (horizon : Nat) (trajectoryKernel : Kernel (Nat -> Fin K -> Rat) ((k : Nat) -> Fin K × Real)) [IsMarkovKernel trajectoryKernel] (hfactor : HasScheduledIIDPrefixKernelFactorization trajectoryKernel horizon) : let law := finiteArmIndependentRewardVectorLaw armLaw let value := fun _ : Nat => finiteArmIIDRewardVectorLoss let prior := Measure.infinitePi law let mu := prior ⊗ₘ trajectoryKernel HasScheduledIndependentMeanGapLaw mu (Finset.univ : Finset (Fin…","missing":[],"search":"hasscheduledindependentmeangaplaw_of_finitearmindependentrewardvectorlaw banditrlproof.tsallis.hasscheduledindependentmeangaplaw_of_finitearmindependentrewardvectorlaw independent roundwise product reward vectors with fixed arm means supply the fixed model-gap law required by the scheduled self-bounding route. theorem compiled","shard":"modules/4555088178dcc4f4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentRewardLawRegret_le_log","label":"integral_sampledScheduledHalfTsallisFiniteArmIndependentRewardLawRegret_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentRewardLawRegret_le_log","description":"A time-varying collection of bounded rational arm-reward laws with fixed finite-bandit means supplies a concrete independent, nonidentically distributed stochastic model for the generated scheduled half-Tsallis logarithmic regret theorem.","url":"../modules/banditrlproof-tsallisfinitearmindependentrewardlaw/index.html#decl-5971274cf4f4","parent":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","order":9944,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentRewardLaw"],["Source","BanditRLProof/TsallisFiniteArmIndependentRewardLaw.lean:103"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteArmIndependentRewardLawRegret_le_log {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hmean : forall t arm, integral (armLaw t arm) (fun reward : Rat => ((reward : Rat) : Real)) = ((model.mean arm : Rat) : Real)) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) (corruption : Real) (hcorruption : 0 <= corruption) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitearmindependentrewardlawregret_le_log banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitearmindependentrewardlawregret_le_log a time-varying collection of bounded rational arm-reward laws with fixed finite-bandit means supplies a concrete independent, nonidentically distributed stochastic model for the generated scheduled half-tsallis logarithmic regret theorem. theorem compiled","shard":"modules/4555088178dcc4f4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionModel","label":"finiteArmIndependentSingleSwitchObstructionModel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionModel","description":"Baseline two-arm model used by the single-switch obstruction.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-1a3c968120cd","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9945,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def finiteArmIndependentSingleSwitchObstructionModel : FiniteBanditModel 2 where","missing":[],"search":"finitearmindependentsingleswitchobstructionmodel banditrlproof.tsallis.finitearmindependentsingleswitchobstructionmodel baseline two-arm model used by the single-switch obstruction. definition compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionModel_bestArm","label":"finiteArmIndependentSingleSwitchObstructionModel_bestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionModel_bestArm","description":"theorem finiteArmIndependentSingleSwitchObstructionModel_bestArm : finiteArmIndependentSingleSwitchObstructionModel.bestArm = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-6f96af0962a4","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9946,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionModel_bestArm : finiteArmIndependentSingleSwitchObstructionModel.bestArm = 0","missing":[],"search":"finitearmindependentsingleswitchobstructionmodel_bestarm banditrlproof.tsallis.finitearmindependentsingleswitchobstructionmodel_bestarm theorem finitearmindependentsingleswitchobstructionmodel_bestarm : finitearmindependentsingleswitchobstructionmodel.bestarm = 0 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw","label":"finiteArmIndependentSingleSwitchObstructionLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw","description":"Arm zero stays at mean `1/2`; arm one moves from `1/4` at round zero to `3/4` forever after round zero.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-31450bec417a","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9947,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:32"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentSingleSwitchObstructionLaw (t : Nat) (arm : Fin 2) : Measure Rat","missing":[],"search":"finitearmindependentsingleswitchobstructionlaw banditrlproof.tsallis.finitearmindependentsingleswitchobstructionlaw arm zero stays at mean `1/2`; arm one moves from `1/4` at round zero to `3/4` forever after round zero. definition compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw_isProbabilityMeasure","label":"finiteArmIndependentSingleSwitchObstructionLaw_isProbabilityMeasure","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw_isProbabilityMeasure","description":"theorem finiteArmIndependentSingleSwitchObstructionLaw_isProbabilityMeasure (t : Nat) (arm : Fin 2) : IsProbabilityMeasure (finiteArmIndependentSingleSwitchObstructionLaw t arm)","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-ed228895024a","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9948,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionLaw_isProbabilityMeasure (t : Nat) (arm : Fin 2) : IsProbabilityMeasure (finiteArmIndependentSingleSwitchObstructionLaw t arm)","missing":[],"search":"finitearmindependentsingleswitchobstructionlaw_isprobabilitymeasure banditrlproof.tsallis.finitearmindependentsingleswitchobstructionlaw_isprobabilitymeasure theorem finitearmindependentsingleswitchobstructionlaw_isprobabilitymeasure (t : nat) (arm : fin 2) : isprobabilitymeasure (finitearmindependentsingleswitchobstructionlaw t arm) theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw_mem_Icc_ae","label":"finiteArmIndependentSingleSwitchObstructionLaw_mem_Icc_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw_mem_Icc_ae","description":"theorem finiteArmIndependentSingleSwitchObstructionLaw_mem_Icc_ae (t : Nat) (arm : Fin 2) : ∀ᵐ reward ∂finiteArmIndependentSingleSwitchObstructionLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-90e50a29c1e4","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9949,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:48"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionLaw_mem_Icc_ae (t : Nat) (arm : Fin 2) : ∀ᵐ reward ∂finiteArmIndependentSingleSwitchObstructionLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1","missing":[],"search":"finitearmindependentsingleswitchobstructionlaw_mem_icc_ae banditrlproof.tsallis.finitearmindependentsingleswitchobstructionlaw_mem_icc_ae theorem finitearmindependentsingleswitchobstructionlaw_mem_icc_ae (t : nat) (arm : fin 2) : ∀ᵐ reward ∂finitearmindependentsingleswitchobstructionlaw t arm, ((reward : rat) : real) ∈ set.icc (0 : real) 1 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_zero","label":"finiteArmIndependentSingleSwitchObstructionMean_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_zero","description":"theorem finiteArmIndependentSingleSwitchObstructionMean_zero (t : Nat) : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw t 0 = 1 / 2","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-d088c3549682","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9950,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionMean_zero (t : Nat) : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw t 0 = 1 / 2","missing":[],"search":"finitearmindependentsingleswitchobstructionmean_zero banditrlproof.tsallis.finitearmindependentsingleswitchobstructionmean_zero theorem finitearmindependentsingleswitchobstructionmean_zero (t : nat) : finitearmindependentrewardmean finitearmindependentsingleswitchobstructionlaw t 0 = 1 / 2 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_one_zero","label":"finiteArmIndependentSingleSwitchObstructionMean_one_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_one_zero","description":"theorem finiteArmIndependentSingleSwitchObstructionMean_one_zero : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw 0 1 = 1 / 4","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-fc917bd951f4","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9951,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:67"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionMean_one_zero : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw 0 1 = 1 / 4","missing":[],"search":"finitearmindependentsingleswitchobstructionmean_one_zero banditrlproof.tsallis.finitearmindependentsingleswitchobstructionmean_one_zero theorem finitearmindependentsingleswitchobstructionmean_one_zero : finitearmindependentrewardmean finitearmindependentsingleswitchobstructionlaw 0 1 = 1 / 4 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_one_succ","label":"finiteArmIndependentSingleSwitchObstructionMean_one_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_one_succ","description":"theorem finiteArmIndependentSingleSwitchObstructionMean_one_succ (t : Nat) : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw (t + 1) 1 = 3 / 4","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-7bd1405548fc","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9952,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:75"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionMean_one_succ (t : Nat) : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw (t + 1) 1 = 3 / 4","missing":[],"search":"finitearmindependentsingleswitchobstructionmean_one_succ banditrlproof.tsallis.finitearmindependentsingleswitchobstructionmean_one_succ theorem finitearmindependentsingleswitchobstructionmean_one_succ (t : nat) : finitearmindependentrewardmean finitearmindependentsingleswitchobstructionlaw (t + 1) 1 = 3 / 4 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionInitialMean","label":"finiteArmIndependentSingleSwitchObstructionInitialMean","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionInitialMean","description":"theorem finiteArmIndependentSingleSwitchObstructionInitialMean (arm : Fin 2) : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw 0 arm = ((finiteArmIndependentSingleSwitchObstructionModel.mean arm : Rat) : Real)","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-5e64ac17dfc5","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9953,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:83"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionInitialMean (arm : Fin 2) : finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw 0 arm = ((finiteArmIndependentSingleSwitchObstructionModel.mean arm : Rat) : Real)","missing":[],"search":"finitearmindependentsingleswitchobstructioninitialmean banditrlproof.tsallis.finitearmindependentsingleswitchobstructioninitialmean theorem finitearmindependentsingleswitchobstructioninitialmean (arm : fin 2) : finitearmindependentrewardmean finitearmindependentsingleswitchobstructionlaw 0 arm = ((finitearmindependentsingleswitchobstructionmodel.mean arm : rat) : real) theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGap_pos","label":"finiteArmIndependentSingleSwitchObstructionGap_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGap_pos","description":"theorem finiteArmIndependentSingleSwitchObstructionGap_pos (arm : Fin 2) (harm : arm ≠ finiteArmIndependentSingleSwitchObstructionModel.bestArm) : 0 < ((finiteArmIndependentSingleSwitchObstructionModel.gap arm : Rat) : Real)","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-87e8ae4cb327","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9954,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:92"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionGap_pos (arm : Fin 2) (harm : arm ≠ finiteArmIndependentSingleSwitchObstructionModel.bestArm) : 0 < ((finiteArmIndependentSingleSwitchObstructionModel.gap arm : Rat) : Real)","missing":[],"search":"finitearmindependentsingleswitchobstructiongap_pos banditrlproof.tsallis.finitearmindependentsingleswitchobstructiongap_pos theorem finitearmindependentsingleswitchobstructiongap_pos (arm : fin 2) (harm : arm ≠ finitearmindependentsingleswitchobstructionmodel.bestarm) : 0 < ((finitearmindependentsingleswitchobstructionmodel.gap arm : rat) : real) theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGap_le_one","label":"finiteArmIndependentSingleSwitchObstructionGap_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGap_le_one","description":"theorem finiteArmIndependentSingleSwitchObstructionGap_le_one (arm : Fin 2) (harm : arm ≠ finiteArmIndependentSingleSwitchObstructionModel.bestArm) : ((finiteArmIndependentSingleSwitchObstructionModel.gap arm : Rat) : Real) <= 1","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-0bd4e7425638","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9955,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:106"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionGap_le_one (arm : Fin 2) (harm : arm ≠ finiteArmIndependentSingleSwitchObstructionModel.bestArm) : ((finiteArmIndependentSingleSwitchObstructionModel.gap arm : Rat) : Real) <= 1","missing":[],"search":"finitearmindependentsingleswitchobstructiongap_le_one banditrlproof.tsallis.finitearmindependentsingleswitchobstructiongap_le_one theorem finitearmindependentsingleswitchobstructiongap_le_one (arm : fin 2) (harm : arm ≠ finitearmindependentsingleswitchobstructionmodel.bestarm) : ((finitearmindependentsingleswitchobstructionmodel.gap arm : rat) : real) <= 1 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionBestArmAt_zero","label":"finiteArmIndependentSingleSwitchObstructionBestArmAt_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionBestArmAt_zero","description":"theorem finiteArmIndependentSingleSwitchObstructionBestArmAt_zero : finiteArmIndependentBestArmAt finiteArmIndependentSingleSwitchObstructionModel finiteArmIndependentSingleSwitchObstructionLaw 0 = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-4790c85adcaf","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9956,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:120"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionBestArmAt_zero : finiteArmIndependentBestArmAt finiteArmIndependentSingleSwitchObstructionModel finiteArmIndependentSingleSwitchObstructionLaw 0 = 0","missing":[],"search":"finitearmindependentsingleswitchobstructionbestarmat_zero banditrlproof.tsallis.finitearmindependentsingleswitchobstructionbestarmat_zero theorem finitearmindependentsingleswitchobstructionbestarmat_zero : finitearmindependentbestarmat finitearmindependentsingleswitchobstructionmodel finitearmindependentsingleswitchobstructionlaw 0 = 0 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionBestArmAt_succ","label":"finiteArmIndependentSingleSwitchObstructionBestArmAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionBestArmAt_succ","description":"theorem finiteArmIndependentSingleSwitchObstructionBestArmAt_succ (t : Nat) : finiteArmIndependentBestArmAt finiteArmIndependentSingleSwitchObstructionModel finiteArmIndependentSingleSwitchObstructionLaw (t + 1) = 1","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-be57af29b777","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9957,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:151"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionBestArmAt_succ (t : Nat) : finiteArmIndependentBestArmAt finiteArmIndependentSingleSwitchObstructionModel finiteArmIndependentSingleSwitchObstructionLaw (t + 1) = 1","missing":[],"search":"finitearmindependentsingleswitchobstructionbestarmat_succ banditrlproof.tsallis.finitearmindependentsingleswitchobstructionbestarmat_succ theorem finitearmindependentsingleswitchobstructionbestarmat_succ (t : nat) : finitearmindependentbestarmat finitearmindependentsingleswitchobstructionmodel finitearmindependentsingleswitchobstructionlaw (t + 1) = 1 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionIndicator","label":"finiteArmIndependentSingleSwitchObstructionIndicator","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionIndicator","description":"private theorem finiteArmIndependentSingleSwitchObstructionIndicator (s : Nat) : (if ∃ arm : Fin 2, finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw (s + 1) arm ≠ finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw s arm then (1 : Real) else 0) = if s = 0 then 1 else 0","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-cd3f4eb810c2","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9958,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:183"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem finiteArmIndependentSingleSwitchObstructionIndicator (s : Nat) : (if ∃ arm : Fin 2, finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw (s + 1) arm ≠ finiteArmIndependentRewardMean finiteArmIndependentSingleSwitchObstructionLaw s arm then (1 : Real) else 0) = if s = 0 then 1 else 0","missing":[],"search":"finitearmindependentsingleswitchobstructionindicator banditrlproof.tsallis.finitearmindependentsingleswitchobstructionindicator private theorem finitearmindependentsingleswitchobstructionindicator (s : nat) : (if ∃ arm : fin 2, finitearmindependentrewardmean finitearmindependentsingleswitchobstructionlaw (s + 1) arm ≠ finitearmindependentrewardmean finitearmindependentsingleswitchobstructionlaw s arm then (1 : real) else 0) = if s = 0 then 1 else 0 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGlobalCount_succ","label":"finiteArmIndependentSingleSwitchObstructionGlobalCount_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGlobalCount_succ","description":"Every positive prefix contains exactly the one change at `0 -> 1`.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-2e68a7842350","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9959,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:219"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchObstructionGlobalCount_succ (t : Nat) : finiteArmIndependentCumulativeGlobalMeanSwitchCount finiteArmIndependentSingleSwitchObstructionLaw (t + 1) = 1","missing":[],"search":"finitearmindependentsingleswitchobstructionglobalcount_succ banditrlproof.tsallis.finitearmindependentsingleswitchobstructionglobalcount_succ every positive prefix contains exactly the one change at `0 -> 1`. theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage","label":"finiteArmIndependentSingleSwitchComparatorAdvantage","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage","description":"Exact actual-mean advantage charged by the moving-comparator decomposition on the single-switch law.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-613a92e55a9e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9960,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:230"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentSingleSwitchComparatorAdvantage (horizon : Nat) : Real","missing":[],"search":"finitearmindependentsingleswitchcomparatoradvantage banditrlproof.tsallis.finitearmindependentsingleswitchcomparatoradvantage exact actual-mean advantage charged by the moving-comparator decomposition on the single-switch law. definition compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_zero","label":"finiteArmIndependentSingleSwitchComparatorAdvantage_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_zero","description":"theorem finiteArmIndependentSingleSwitchComparatorAdvantage_zero : finiteArmIndependentSingleSwitchComparatorAdvantage 0 = 0","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-c4f7ba99863c","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9961,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:243"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchComparatorAdvantage_zero : finiteArmIndependentSingleSwitchComparatorAdvantage 0 = 0","missing":[],"search":"finitearmindependentsingleswitchcomparatoradvantage_zero banditrlproof.tsallis.finitearmindependentsingleswitchcomparatoradvantage_zero theorem finitearmindependentsingleswitchcomparatoradvantage_zero : finitearmindependentsingleswitchcomparatoradvantage 0 = 0 theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_eq","label":"finiteArmIndependentSingleSwitchComparatorAdvantage_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_eq","description":"One permanent switch makes the exact comparator advantage equal to `horizon / 4`, despite the global switch count being one.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-c4af74d02c1b","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9962,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:250"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchComparatorAdvantage_eq (horizon : Nat) : finiteArmIndependentSingleSwitchComparatorAdvantage horizon = (horizon : Real) / 4","missing":[],"search":"finitearmindependentsingleswitchcomparatoradvantage_eq banditrlproof.tsallis.finitearmindependentsingleswitchcomparatoradvantage_eq one permanent switch makes the exact comparator advantage equal to `horizon / 4`, despite the global switch count being one. theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchDynamicComparatorPenalty_eq","label":"finiteArmIndependentSingleSwitchDynamicComparatorPenalty_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchDynamicComparatorPenalty_eq","description":"The current repeated-prefix envelope penalty is exactly `2 * horizon` on the same one-switch law.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-3c2f440f09dc","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9963,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:267"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchDynamicComparatorPenalty_eq (horizon : Nat) : finiteArmIndependentDynamicComparatorPenalty finiteArmIndependentSingleSwitchObstructionModel finiteArmIndependentSingleSwitchObstructionLaw horizon (fun t _ => finiteArmIndependentCumulativeGlobalMeanSwitchCount finiteArmIndependentSingleSwitchObstructionLaw t) = 2 * (horizon : Real)","missing":[],"search":"finitearmindependentsingleswitchdynamiccomparatorpenalty_eq banditrlproof.tsallis.finitearmindependentsingleswitchdynamiccomparatorpenalty_eq the current repeated-prefix envelope penalty is exactly `2 * horizon` on the same one-switch law. theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_gt_nat_mul_sqrt","label":"finiteArmIndependentSingleSwitchComparatorAdvantage_gt_nat_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_gt_nat_mul_sqrt","description":"For every natural coefficient, a one-switch horizon exists where the exact separately charged comparator advantage exceeds that coefficient times `sqrt(horizon)`. This rules out obtaining a uniform square-root switch-rate by bounding this term independently.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-cbafb7e7f12e","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9964,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:292"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchComparatorAdvantage_gt_nat_mul_sqrt (coefficient : Nat) : let horizon := (4 * coefficient + 1) ^ 2 finiteArmIndependentCumulativeGlobalMeanSwitchCount finiteArmIndependentSingleSwitchObstructionLaw horizon = 1 ∧ (coefficient : Real) * Real.sqrt (horizon : Real) < finiteArmIndependentSingleSwitchComparatorAdvantage horizon","missing":[],"search":"finitearmindependentsingleswitchcomparatoradvantage_gt_nat_mul_sqrt banditrlproof.tsallis.finitearmindependentsingleswitchcomparatoradvantage_gt_nat_mul_sqrt for every natural coefficient, a one-switch horizon exists where the exact separately charged comparator advantage exceeds that coefficient times `sqrt(horizon)`. this rules out obtaining a uniform square-root switch-rate by bounding this term independently. theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchDynamicComparatorRouteObstruction","label":"finiteArmIndependentSingleSwitchDynamicComparatorRouteObstruction","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchDynamicComparatorRouteObstruction","description":"Compiled blocker certificate for the current dynamic-comparator proof route. It does not assert a regret lower bound: a sharper proof could exploit cancellation between fixed-comparator regret and comparator advantage.","url":"../modules/banditrlproof-tsallisfinitearmindependentsingleswitchcomparatorobstruction/index.html#decl-5fa74518c1c2","parent":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","order":9965,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction"],["Source","BanditRLProof/TsallisFiniteArmIndependentSingleSwitchComparatorObstruction.lean:325"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentSingleSwitchDynamicComparatorRouteObstruction (horizon : Nat) (horizon_pos : 0 < horizon) : finiteArmIndependentCumulativeGlobalMeanSwitchCount finiteArmIndependentSingleSwitchObstructionLaw horizon = 1 ∧ finiteArmIndependentSingleSwitchComparatorAdvantage horizon = (horizon : Real) / 4 ∧ finiteArmIndependentDynamicComparatorPenalty finiteArmIndependentSingleSwitchObstructionModel finiteArmIndependentSingleSwitchObstructionLaw horizon (fun t _ => finiteArmIndependentCumulativeGlobalMeanSwitchCount finiteArmIndependentSingleSwitchObstructionLaw t) = 2 * (horizon : Real)","missing":[],"search":"finitearmindependentsingleswitchdynamiccomparatorrouteobstruction banditrlproof.tsallis.finitearmindependentsingleswitchdynamiccomparatorrouteobstruction compiled blocker certificate for the current dynamic-comparator proof route. it does not assert a regret lower bound: a sharper proof could exploit cancellation between fixed-comparator regret and comparator advantage. theorem compiled","shard":"modules/a6a975af03ac1544.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteBanditMeanLoss","label":"finiteBanditMeanLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteBanditMeanLoss","description":"The stationary predictable loss family obtained from bounded arm means.","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html#decl-729346cf1080","parent":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","order":9966,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisFiniteBanditMeanLoss"],["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean:26"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteBanditMeanLoss {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) : PredictableLossVector Env (Fin K) where","missing":[],"search":"finitebanditmeanloss banditrlproof.exp3.finitebanditmeanloss the stationary predictable loss family obtained from bounded arm means. definition compiled","shard":"modules/d5a6aca3bc377c2a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteBanditMeanLoss_initial","label":"finiteBanditMeanLoss_initial","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteBanditMeanLoss_initial","description":"theorem finiteBanditMeanLoss_initial {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (env : Env) (arm : Fin K) : (finiteBanditMeanLoss (Env := Env) model hmean).initial env arm = 1 - ((model.mean arm : Rat) : Real)","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html#decl-fe1085764a3f","parent":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","order":9967,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteBanditMeanLoss"],["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean:53"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteBanditMeanLoss_initial {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (env : Env) (arm : Fin K) : (finiteBanditMeanLoss (Env := Env) model hmean).initial env arm = 1 - ((model.mean arm : Rat) : Real)","missing":[],"search":"finitebanditmeanloss_initial banditrlproof.exp3.finitebanditmeanloss_initial theorem finitebanditmeanloss_initial {k : nat} {env : type u} [measurablespace env] (model : finitebanditmodel k) (hmean : forall arm, ((model.mean arm : rat) : real) ∈ set.icc (0 : real) 1) (env : env) (arm : fin k) : (finitebanditmeanloss (env := env) model hmean).initial env arm = 1 - ((model.mean arm : rat) : real) theorem compiled","shard":"modules/d5a6aca3bc377c2a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.finiteBanditMeanLoss_successor","label":"finiteBanditMeanLoss_successor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.finiteBanditMeanLoss_successor","description":"theorem finiteBanditMeanLoss_successor {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : (finiteBanditMeanLoss (Env := Env) model hmean).successor n env history arm = 1 - ((model.mean arm : Rat) : Real)","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html#decl-b67abb1a6840","parent":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","order":9968,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteBanditMeanLoss"],["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean:64"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteBanditMeanLoss_successor {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (n : Nat) (env : Env) (history : History.FinitePairHistory (Fin K) Real n) (arm : Fin K) : (finiteBanditMeanLoss (Env := Env) model hmean).successor n env history arm = 1 - ((model.mean arm : Rat) : Real)","missing":[],"search":"finitebanditmeanloss_successor banditrlproof.exp3.finitebanditmeanloss_successor theorem finitebanditmeanloss_successor {k : nat} {env : type u} [measurablespace env] (model : finitebanditmodel k) (hmean : forall arm, ((model.mean arm : rat) : real) ∈ set.icc (0 : real) 1) (n : nat) (env : env) (history : history.finitepairhistory (fin k) real n) (arm : fin k) : (finitebanditmeanloss (env := env) model hmean).successor n env history arm = 1 - ((model.mean arm : rat) : real) theorem compiled","shard":"modules/d5a6aca3bc377c2a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.predictableLossAt_finiteBanditMeanLoss","label":"predictableLossAt_finiteBanditMeanLoss","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.predictableLossAt_finiteBanditMeanLoss","description":"theorem predictableLossAt_finiteBanditMeanLoss {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (sample : Env × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : predictableLossAt (finiteBanditMeanLoss (Env := Env) model hmean) t sample arm = 1 - ((model.mean arm : Rat) : Real)","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html#decl-8885749dd4e7","parent":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","order":9969,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteBanditMeanLoss"],["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean:77"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_finiteBanditMeanLoss {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (sample : Env × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : predictableLossAt (finiteBanditMeanLoss (Env := Env) model hmean) t sample arm = 1 - ((model.mean arm : Rat) : Real)","missing":[],"search":"predictablelossat_finitebanditmeanloss banditrlproof.exp3.predictablelossat_finitebanditmeanloss theorem predictablelossat_finitebanditmeanloss {k : nat} {env : type u} [measurablespace env] (model : finitebanditmodel k) (hmean : forall arm, ((model.mean arm : rat) : real) ∈ set.icc (0 : real) 1) (t : nat) (sample : env × ((k : nat) -> fin k × real)) (arm : fin k) : predictablelossat (finitebanditmeanloss (env := env) model hmean) t sample arm = 1 - ((model.mean arm : rat) : real) theorem compiled","shard":"modules/d5a6aca3bc377c2a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Exp3.predictableLossAt_finiteBanditMeanLoss_sub_bestArm_eq_gap","label":"predictableLossAt_finiteBanditMeanLoss_sub_bestArm_eq_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Exp3.predictableLossAt_finiteBanditMeanLoss_sub_bestArm_eq_gap","description":"Mean-loss differences against the selected best arm are model gaps.","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html#decl-7186fb0a2c60","parent":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","order":9970,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteBanditMeanLoss"],["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean:90"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_finiteBanditMeanLoss_sub_bestArm_eq_gap {K : Nat} {Env : Type u} [MeasurableSpace Env] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (sample : Env × ((k : Nat) -> Fin K × Real)) (arm : Fin K) : predictableLossAt (finiteBanditMeanLoss (Env := Env) model hmean) t sample arm - predictableLossAt (finiteBanditMeanLoss (Env := Env) model hmean) t sample model.bestArm = ((model.gap arm : Rat) : Real)","missing":[],"search":"predictablelossat_finitebanditmeanloss_sub_bestarm_eq_gap banditrlproof.exp3.predictablelossat_finitebanditmeanloss_sub_bestarm_eq_gap mean-loss differences against the selected best arm are model gaps. theorem compiled","shard":"modules/d5a6aca3bc377c2a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteBanditMeanLossRegret_le_log_fixedGap","label":"integral_sampledScheduledHalfTsallisFiniteBanditMeanLossRegret_le_log_fixedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteBanditMeanLossRegret_le_log_fixedGap","description":"The square-root scheduled half-Tsallis algorithm has logarithmic fixed-gap regret on the deterministic mean-loss environment of a bounded finite-bandit model with strictly positive non-best gaps.","url":"../modules/banditrlproof-tsallisfinitebanditmeanloss/index.html#decl-8955a87f0720","parent":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","order":9971,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisFiniteBanditMeanLoss"],["Source","BanditRLProof/TsallisFiniteBanditMeanLoss.lean:115"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisFiniteBanditMeanLossRegret_le_log_fixedGap {K : Nat} {Env : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] (prior : Measure Env) [IsProbabilityMeasure prior] (model : FiniteBanditModel K) (hmean : forall arm, ((model.mean arm : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (hgapPos : forall arm, arm ≠ model.bestArm -> 0 < ((model.gap arm : Rat) : Real)) (horizon : Nat) (corruption : Real) (hcorruption : 0 <= corruption) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledscheduledhalftsallisfinitebanditmeanlossregret_le_log_fixedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallisfinitebanditmeanlossregret_le_log_fixedgap the square-root scheduled half-tsallis algorithm has logarithmic fixed-gap regret on the deterministic mean-loss environment of a bounded finite-bandit model with strictly positive non-best gaps. theorem compiled","shard":"modules/d5a6aca3bc377c2a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.powerWeightedSquaredImportanceWeightedLoss","label":"powerWeightedSquaredImportanceWeightedLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.powerWeightedSquaredImportanceWeightedLoss","description":"The inverse-Hessian-weighted square of one sampled loss estimate.","url":"../modules/banditrlproof-tsallisimportanceweightedmoment/index.html#decl-31af464ba76c","parent":"module:BanditRLProof.TsallisImportanceWeightedMoment","order":9972,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisImportanceWeightedMoment"],["Source","BanditRLProof/TsallisImportanceWeightedMoment.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def powerWeightedSquaredImportanceWeightedLoss {Action : Type u} (arms : Finset Action) (alpha : Real) (prob loss : Action -> Real) (chosen : Action) : Real","missing":[],"search":"powerweightedsquaredimportanceweightedloss banditrlproof.tsallis.powerweightedsquaredimportanceweightedloss the inverse-hessian-weighted square of one sampled loss estimate. definition compiled","shard":"modules/f8ebe049226fec43.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.powerWeightedSquaredImportanceWeightedLoss_eq_selected","label":"powerWeightedSquaredImportanceWeightedLoss_eq_selected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.powerWeightedSquaredImportanceWeightedLoss_eq_selected","description":"Pathwise power-moment identity: only the sampled coordinate remains.","url":"../modules/banditrlproof-tsallisimportanceweightedmoment/index.html#decl-2d7ba9a9453f","parent":"module:BanditRLProof.TsallisImportanceWeightedMoment","order":9973,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisImportanceWeightedMoment"],["Source","BanditRLProof/TsallisImportanceWeightedMoment.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem powerWeightedSquaredImportanceWeightedLoss_eq_selected {Action : Type u} [DecidableEq Action] (arms : Finset Action) (alpha : Real) (prob loss : Action -> Real) (chosen : Action) (hchosen : chosen ∈ arms) (hprob : 0 < prob chosen) : powerWeightedSquaredImportanceWeightedLoss arms alpha prob loss chosen = (loss chosen) ^ 2 * (prob chosen) ^ (-alpha)","missing":[],"search":"powerweightedsquaredimportanceweightedloss_eq_selected banditrlproof.tsallis.powerweightedsquaredimportanceweightedloss_eq_selected pathwise power-moment identity: only the sampled coordinate remains. theorem compiled","shard":"modules/f8ebe049226fec43.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_eq","label":"sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_eq","description":"The sampling-mass-weighted finite sum equals the power-weighted loss square.","url":"../modules/banditrlproof-tsallisimportanceweightedmoment/index.html#decl-53415df4e79c","parent":"module:BanditRLProof.TsallisImportanceWeightedMoment","order":9974,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisImportanceWeightedMoment"],["Source","BanditRLProof/TsallisImportanceWeightedMoment.lean:64"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (alpha : Real) (prob loss : Action -> Real) (hprob : forall action, action ∈ arms -> 0 < prob action) : arms.sum (fun chosen => prob chosen * powerWeightedSquaredImportanceWeightedLoss arms alpha prob loss chosen) = arms.sum (fun action => (loss action) ^ 2 * (prob action) ^ (1 - alpha))","missing":[],"search":"sum_prob_mul_powerweightedsquaredimportanceweightedloss_eq banditrlproof.tsallis.sum_prob_mul_powerweightedsquaredimportanceweightedloss_eq the sampling-mass-weighted finite sum equals the power-weighted loss square. theorem compiled","shard":"modules/f8ebe049226fec43.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_le_powerSum","label":"sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_le_powerSum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_le_powerSum","description":"For losses in `[0,1]`, the Tsallis importance-weighted power moment is bounded by the finite power sum with exponent `1 - alpha`.","url":"../modules/banditrlproof-tsallisimportanceweightedmoment/index.html#decl-54bc4c25f8ad","parent":"module:BanditRLProof.TsallisImportanceWeightedMoment","order":9975,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisImportanceWeightedMoment"],["Source","BanditRLProof/TsallisImportanceWeightedMoment.lean:87"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_le_powerSum {Action : Type u} [DecidableEq Action] (arms : Finset Action) (alpha : Real) (prob loss : Action -> Real) (hprob : forall action, action ∈ arms -> 0 < prob action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => prob chosen * powerWeightedSquaredImportanceWeightedLoss arms alpha prob loss chosen) <= powerSum arms (1 - alpha) prob","missing":[],"search":"sum_prob_mul_powerweightedsquaredimportanceweightedloss_le_powersum banditrlproof.tsallis.sum_prob_mul_powerweightedsquaredimportanceweightedloss_le_powersum for losses in `[0,1]`, the tsallis importance-weighted power moment is bounded by the finite power sum with exponent `1 - alpha`. theorem compiled","shard":"modules/f8ebe049226fec43.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartEpochRounds","label":"oracleRestartEpochRounds","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartEpochRounds","description":"Included rounds assigned to one epoch.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html#decl-0c298db8b4a3","parent":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","order":9976,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleRestartEpochRounds {Epoch : Type w} [DecidableEq Epoch] (epochOf : Nat -> Epoch) (horizon : Nat) (epoch : Epoch) : Finset Nat","missing":[],"search":"oraclerestartepochrounds banditrlproof.tsallis.oraclerestartepochrounds included rounds assigned to one epoch. definition compiled","shard":"modules/05d4c2590f7ed16f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret","label":"sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret","description":"Predictable fixed-comparator regret contributed by one epoch fiber.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html#decl-3cba3dae5f0a","parent":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","order":9977,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret {Env : Type u} {Action : Type v} {Epoch : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] [DecidableEq Epoch] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (epochOf : Nat -> Epoch) (epochComparator : Epoch -> Action) (horizon : Nat) (epoch : Epoch) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallispredictableoraclerestartepochregret banditrlproof.tsallis.sampledscheduledhalftsallispredictableoraclerestartepochregret predictable fixed-comparator regret contributed by one epoch fiber. definition compiled","shard":"modules/05d4c2590f7ed16f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_oracleRestartEpochRegret","label":"sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_oracleRestartEpochRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_oracleRestartEpochRegret","description":"Moving-comparator regret against an epochwise constant comparator is exactly the sum of its epoch-fiber fixed-comparator regrets.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html#decl-ceff4338bb0a","parent":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","order":9978,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean:46"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_oracleRestartEpochRegret {Env : Type u} {Action : Type v} {Epoch : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] [DecidableEq Epoch] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (epochs : Finset Epoch) (epochOf : Nat -> Epoch) (epochComparator : Epoch -> Action) (horizon : Nat) (hEpochOf : ∀ t ∈ Finset.range (horizon + 1), epochOf t ∈ epochs) (sample : Env × ((k : Nat) -> Action × Real)) : sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta loss (fun t => epochComparator (epochOf t)) horizon sample = epochs.sum (fun epoch => sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret arms harms eta loss epochOf epochComparator horizon epoch sample)","missing":[],"search":"sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_eq_sum_oraclerestartepochregret banditrlproof.tsallis.sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_eq_sum_oraclerestartepochregret moving-comparator regret against an epochwise constant comparator is exactly the sum of its epoch-fiber fixed-comparator regrets. theorem compiled","shard":"modules/05d4c2590f7ed16f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sqrt_oracleRestartEpochRounds_card_le","label":"sum_sqrt_oracleRestartEpochRounds_card_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sqrt_oracleRestartEpochRounds_card_le","description":"Cauchy--Schwarz bound for the square roots of epoch-fiber lengths.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html#decl-cfd444eb2ffd","parent":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","order":9979,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean:75"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sqrt_oracleRestartEpochRounds_card_le {Epoch : Type w} [DecidableEq Epoch] (epochs : Finset Epoch) (epochOf : Nat -> Epoch) (horizon : Nat) (hEpochOf : ∀ t ∈ Finset.range (horizon + 1), epochOf t ∈ epochs) : epochs.sum (fun epoch => Real.sqrt ((oracleRestartEpochRounds epochOf horizon epoch).card : Real)) <= Real.sqrt (epochs.card : Real) * Real.sqrt (((horizon + 1 : Nat) : Real))","missing":[],"search":"sum_sqrt_oraclerestartepochrounds_card_le banditrlproof.tsallis.sum_sqrt_oraclerestartepochrounds_card_le cauchy--schwarz bound for the square roots of epoch-fiber lengths. theorem compiled","shard":"modules/05d4c2590f7ed16f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSqrt","label":"sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSqrt","description":"Oracle-restart assembly: epoch-local square-root fixed-comparator certificates yield a global square-root moving-comparator bound.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html#decl-d0969a8f7ed4","parent":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","order":9980,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean:109"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSqrt {Env : Type u} {Action : Type v} {Epoch : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] [DecidableEq Epoch] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (epochs : Finset Epoch) (epochOf : Nat -> Epoch) (epochComparator : Epoch -> Action) (horizon : Nat) (hEpochOf : ∀ t ∈ Finset.range (horizon + 1), epochOf t ∈ epochs) (coefficient : Real) (hcoefficient : 0 <= coefficient) (sample : Env × ((k : Nat) -> Action × Real)) (hEpochRegret : ∀ epoch ∈ epochs, sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret arms harms eta loss epochOf epochComparator horizon epoch sample <= coefficient * Real.sqrt ((oracleRestartEpochRounds epochOf horizon epoch).card : Real)) : sampledScheduledHalfTsa…","missing":[],"search":"sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_le_oraclerestartsqrt banditrlproof.tsallis.sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_le_oraclerestartsqrt oracle-restart assembly: epoch-local square-root fixed-comparator certificates yield a global square-root moving-comparator bound. theorem compiled","shard":"modules/05d4c2590f7ed16f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSwitchCountSqrt","label":"sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSwitchCountSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSwitchCountSqrt","description":"Switch-count-facing oracle-restart assembly. When the epoch partition has at most one more epoch than switches, the global moving-comparator bound has the standard `sqrt((switches + 1) * (horizon + 1))` product form.","url":"../modules/banditrlproof-tsallisoraclerestartdynamicregret/index.html#decl-a7276c24d02d","parent":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","order":9981,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartDynamicRegret.lean:162"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSwitchCountSqrt {Env : Type u} {Action : Type v} {Epoch : Type w} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] [DecidableEq Epoch] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (epochs : Finset Epoch) (epochOf : Nat -> Epoch) (epochComparator : Epoch -> Action) (horizon switches : Nat) (hEpochOf : ∀ t ∈ Finset.range (horizon + 1), epochOf t ∈ epochs) (hEpochCard : epochs.card <= switches + 1) (coefficient : Real) (hcoefficient : 0 <= coefficient) (sample : Env × ((k : Nat) -> Action × Real)) (hEpochRegret : ∀ epoch ∈ epochs, sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret arms harms eta loss epochOf epochComparator horizon epoch sample <= coefficient * Real.sqrt ((oracleRestartEpochRounds…","missing":[],"search":"sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_le_oraclerestartswitchcountsqrt banditrlproof.tsallis.sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret_le_oraclerestartswitchcountsqrt switch-count-facing oracle-restart assembly. when the epoch partition has at most one more epoch than switches, the global moving-comparator bound has the standard `sqrt((switches + 1) * (horizon + 1))` product form. theorem compiled","shard":"modules/05d4c2590f7ed16f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","label":"sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","description":"Predictable importance-weighted loss using the generated restart probability at the same actual trajectory time.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-8e85e43b5d3a","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9982,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallispredictableestimatedlossat banditrlproof.tsallis.sampledoraclerestarthalftsallispredictableestimatedlossat predictable importance-weighted loss using the generated restart probability at the same actual trajectory time. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt","label":"sampledOracleRestartHalfTsallisObservedEstimatedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt","description":"Stored-reward importance-weighted loss under the generated restart probability at the same actual trajectory time.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-86f9ef113ef5","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9983,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:36"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallisobservedestimatedlossat banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedestimatedlossat stored-reward importance-weighted loss under the generated restart probability at the same actual trajectory time. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAt","label":"sampledOracleRestartHalfTsallisProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAt","description":"Restarted successor probability on an environment/global-prefix state.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-aaceb4901ce9","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9984,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:48"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisProbabilityAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityat banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityat restarted successor probability on an environment/global-prefix state. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisEnvironmentHistoryDistributionSource","label":"sampledOracleRestartHalfTsallisEnvironmentHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisEnvironmentHistoryDistributionSource","description":"Environment-lifted measurable source for one restarted successor law.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-6aa46a0bc495","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9985,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisEnvironmentHistoryDistributionSource {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : Exp3.MeasurableFiniteActionDistribution arms (sampledOracleRestartHalfTsallisProbabilityAt (Env := Env) arms harms eta schedule n)","missing":[],"search":"sampledoraclerestarthalftsallisenvironmenthistorydistributionsource banditrlproof.tsallis.sampledoraclerestarthalftsallisenvironmenthistorydistributionsource environment-lifted measurable source for one restarted successor law. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPolicyAt","label":"sampledOracleRestartHalfTsallisPolicyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPolicyAt","description":"The restarted policy comapped to the environment/global-prefix state.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-9d687c36b21a","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9986,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:78"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPolicyAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : Kernel (Env × History.FinitePairHistory Action Real n) Action","missing":[],"search":"sampledoraclerestarthalftsallispolicyat banditrlproof.tsallis.sampledoraclerestarthalftsallispolicyat the restarted policy comapped to the environment/global-prefix state. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPolicyAt_eq_finiteActionKernel","label":"sampledOracleRestartHalfTsallisPolicyAt_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPolicyAt_eq_finiteActionKernel","description":"theorem sampledOracleRestartHalfTsallisPolicyAt_eq_finiteActionKernel {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : sampledOracleRestartHalfTsallisPolicyAt (Env := Env) arms harms eta schedule n = Exp3.finiteActionKe…","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-68145629cb2d","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9987,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:103"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPolicyAt_eq_finiteActionKernel {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : sampledOracleRestartHalfTsallisPolicyAt (Env := Env) arms harms eta schedule n = Exp3.finiteActionKernel arms (sampledOracleRestartHalfTsallisProbabilityAt (Env := Env) arms harms eta schedule n) (sampledOracleRestartHalfTsallisEnvironmentHistoryDistributionSource (Env := Env) arms harms eta schedule n)","missing":[],"search":"sampledoraclerestarthalftsallispolicyat_eq_finiteactionkernel banditrlproof.tsallis.sampledoraclerestarthalftsallispolicyat_eq_finiteactionkernel theorem sampledoraclerestarthalftsallispolicyat_eq_finiteactionkernel {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (n : nat) : sampledoraclerestarthalftsallispolicyat (env := env) arms harms eta schedule n = exp3.finiteactionkernel arms (sampledoraclerestarthalftsallisprobabilityat (env := env) arms harms eta schedule n) (sampledoraclerestarthalftsallisenvironmenthistorydistributionsource (env := env) arms harms eta schedule n) theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisProbabilityAtTime","label":"measurable_sampledOracleRestartHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisProbabilityAtTime","description":"theorem measurable_sampledOracleRestartHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Acti…","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-1438ed7cde6f","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9988,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:123"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledOracleRestartHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta schedule t sample candidate)","missing":[],"search":"measurable_sampledoraclerestarthalftsallisprobabilityattime banditrlproof.tsallis.measurable_sampledoraclerestarthalftsallisprobabilityattime theorem measurable_sampledoraclerestarthalftsallisprobabilityattime {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledoraclerestarthalftsallisprobabilityattime arms harms eta schedule t sample candidate) theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","label":"measurable_sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","description":"theorem measurable_sampledOracleRestartHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ a…","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-1d0f3b511f5a","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9989,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:143"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledOracleRestartHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledOracleRestartHalfTsallisPredictableEstimatedLossAt arms harms eta schedule loss t sample candidate)","missing":[],"search":"measurable_sampledoraclerestarthalftsallispredictableestimatedlossat banditrlproof.tsallis.measurable_sampledoraclerestarthalftsallispredictableestimatedlossat theorem measurable_sampledoraclerestarthalftsallispredictableestimatedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (loss : exp3.predictablelossvector env action) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledoraclerestarthalftsallispredictableestimatedlossat arms harms eta schedule loss t sample candidate) theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","label":"sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","description":"Deterministic predictable feedback identifies each stored-reward restart estimator with its predictable counterpart almost surely.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-23fe6d1996ee","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9990,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:167"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment (fun sample => sampledOracleRestartHalfTsallisObservedEstimatedLossAt arms harms eta schedule t sample) =ᵐ[mu] (fun sample => sampledOracleRestartHalfTsallisPredictableEstimatedLossAt arms harms eta schedule loss t sample)","missing":[],"search":"sampledoraclerestarthalftsallisobservedestimatedlossat_eq_predictable_ae banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedestimatedlossat_eq_predictable_ae deterministic predictable feedback identifies each stored-reward restart estimator with its predictable counterpart almost surely. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEstimatedLossAt_first_moments","label":"sampledOracleRestartHalfTsallisPredictableEstimatedLossAt_first_moments","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEstimatedLossAt_first_moments","description":"At each generated restart time, mixed and comparator-weighted importance-weighted estimators are integrable and have the corresponding predictable environment first moments.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-aff3e70f1171","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9991,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:217"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableEstimatedLossAt_first_moments {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (t : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment (Integrable (fun sample => FTRL.linearLoss arms (sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta schedule t sample) (sampledOracleRestartHalfTsallisPredictableEstimatedLossAt arms harms eta schedule loss t sample)) mu ∧ Integrab…","missing":[],"search":"sampledoraclerestarthalftsallispredictableestimatedlossat_first_moments banditrlproof.tsallis.sampledoraclerestarthalftsallispredictableestimatedlossat_first_moments at each generated restart time, mixed and comparator-weighted importance-weighted estimators are integrable and have the corresponding predictable environment first moments. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledOracleRestartHalfTsallisProbabilityAtTime","label":"finiteSimplex_sampledOracleRestartHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_sampledOracleRestartHalfTsallisProbabilityAtTime","description":"theorem finiteSimplex_sampledOracleRestartHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : FTRL.finiteSimplex arms (sampledOracleRestartHalfTsallisProbabilityAtTime a…","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-3966d4fcf55e","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9992,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:391"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_sampledOracleRestartHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : FTRL.finiteSimplex arms (sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta schedule t sample)","missing":[],"search":"finitesimplex_sampledoraclerestarthalftsallisprobabilityattime banditrlproof.tsallis.finitesimplex_sampledoraclerestarthalftsallisprobabilityattime theorem finitesimplex_sampledoraclerestarthalftsallisprobabilityattime {env : type u} {action : type v} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (t : nat) (sample : env × ((k : nat) -> action × real)) : ftrl.finitesimplex arms (sampledoraclerestarthalftsallisprobabilityattime arms harms eta schedule t sample) theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisPredictableLinearLossAt","label":"integrable_sampledOracleRestartHalfTsallisPredictableLinearLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisPredictableLinearLossAt","description":"Generated restart mixed predictable loss and any fixed simplex comparator loss are integrable at every actual time.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-c7073c8b67de","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9993,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:420"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledOracleRestartHalfTsallisPredictableLinearLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (t : Nat) : Integrable (fun sample => FTRL.linearLoss arms (sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta schedule t sample) (Exp3.predictableLossAt loss t sample)) mu ∧ Integrable (fun sample => FTRL.linearLoss arms q (Exp3.predictableLossAt loss t sample)) mu","missing":[],"search":"integrable_sampledoraclerestarthalftsallispredictablelinearlossat banditrlproof.tsallis.integrable_sampledoraclerestarthalftsallispredictablelinearlossat generated restart mixed predictable loss and any fixed simplex comparator loss are integrable at every actual time. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret","label":"sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret","description":"Estimated regret contributed by one actual schedule epoch.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-4ea1ac6cf768","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9994,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:518"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon epoch : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallispredictablescheduleepochestimatedregret banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablescheduleepochestimatedregret estimated regret contributed by one actual schedule epoch. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret","label":"sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret","description":"Stored-reward estimated regret contributed by one actual schedule epoch.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-9eb28e90639c","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9995,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:538"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (q : Action -> Real) (horizon epoch : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret stored-reward estimated regret contributed by one actual schedule epoch. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret","label":"sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret","description":"Environment regret against a fixed comparator distribution on one actual schedule epoch.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-762437d77e67","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9996,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:556"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon epoch : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallispredictablescheduleepochenvironmentregret banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablescheduleepochenvironmentregret environment regret against a fixed comparator distribution on one actual schedule epoch. definition compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_eq_environmentRegret","label":"integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_eq_environmentRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_eq_environmentRegret","description":"Epoch-local estimated regret is integrable and has exactly the same integral as fixed-comparator environment regret on the actual generated restart law.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-8d63950be1d7","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9997,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:575"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_eq_environmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (horizon epoch : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment Integrable (sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret arms harms eta schedule loss q horizon epoch) mu ∧ Integrable (sampledOracleRestartHalfTsallisPredictableScheduleEpochEn…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablescheduleepochestimatedregret_eq_environmentregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablescheduleepochestimatedregret_eq_environmentregret epoch-local estimated regret is integrable and has exactly the same integral as fixed-comparator environment regret on the actual generated restart law. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_predictable_ae","label":"sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_predictable_ae","description":"Stored-reward and predictable-estimator regret agree almost surely on every actual schedule epoch fiber.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-e131c28b23f1","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9998,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:677"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon epoch : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret arms harms eta schedule q horizon epoch =ᵐ[mu] sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret arms harms eta schedule loss q horizon epoch","missing":[],"search":"sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_eq_predictable_ae banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_eq_predictable_ae stored-reward and predictable-estimator regret agree almost surely on every actual schedule epoch fiber. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret_pointMass_eq","label":"sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret_pointMass_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret_pointMass_eq","description":"A point-mass comparator on one schedule fiber is the existing epoch-comparator environment-regret surface.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-c24dd8b98c7d","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":9999,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:716"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret_pointMass_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (epoch : Nat) (hcomparator : epochComparator epoch ∈ arms) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret arms harms eta schedule loss (pointMass (epochComparator epoch)) horizon epoch sample = sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret arms harms eta schedule loss epochComparator horizon epoch sample","missing":[],"search":"sampledoraclerestarthalftsallispredictablescheduleepochenvironmentregret_pointmass_eq banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablescheduleepochenvironmentregret_pointmass_eq a point-mass comparator on one schedule fiber is the existing epoch-comparator environment-regret surface. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","label":"integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","description":"Fixed-arm epoch estimated regret transports exactly to the existing epoch environment-regret integral on the generated restart law.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-a7db27382be7","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10000,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:743"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_pointMass_eq_epochRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (epoch : Nat) (hcomparator : epochComparator epoch ∈ arms) (horizon : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment Integrable (sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret arms harms eta schedule loss (pointMass (epochComparator epoch)) horizon epoch) mu…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablescheduleepochestimatedregret_pointmass_eq_epochregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablescheduleepochestimatedregret_pointmass_eq_epochregret fixed-arm epoch estimated regret transports exactly to the existing epoch environment-regret integral on the generated restart law. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","label":"integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","description":"Stored-reward fixed-arm epoch estimated regret is integrable and has exactly the existing epoch environment-regret integral.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-cc2dc24ae499","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10001,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:794"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_eq_epochRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (epoch : Nat) (hcomparator : epochComparator epoch ∈ arms) (horizon : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment Integrable (sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret arms harms eta schedule (pointMass (epochComparator epoch)) horizon epoch) mu ∧ Integrabl…","missing":[],"search":"integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_eq_epochregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_eq_epochregret stored-reward fixed-arm epoch estimated regret is integrable and has exactly the existing epoch environment-regret integral. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret_le_of_estimatedRegret_le","label":"integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret_le_of_estimatedRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret_le_of_estimatedRegret_le","description":"Any expected estimated-regret certificate on an actual schedule epoch transports to the corresponding fixed-arm environment-regret certificate.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-17e51e610da8","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10002,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:846"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret_le_of_estimatedRegret_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (epoch : Nat) (hcomparator : epochComparator epoch ∈ arms) (horizon : Nat) (bound : Real) (hEstimated : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment integral mu (sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret arms harms eta schedule loss (pointMass (epochComparator epoch))…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablescheduleepochregret_le_of_estimatedregret_le banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablescheduleepochregret_le_of_estimatedregret_le any expected estimated-regret certificate on an actual schedule epoch transports to the corresponding fixed-arm environment-regret certificate. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochEstimatedRegret","label":"integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochEstimatedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochEstimatedRegret","description":"Epoch-local expected estimated-regret certificates assemble into an expected moving-comparator square-root bound on the actual generated restart trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-a346c6664d58","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10003,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:883"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon : Nat) (hcomparator : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, epochComparator epoch ∈ arms) (coefficient : Real) (hcoefficient : 0 <= coefficient) (hEstimated : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, integral (prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedul…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_schedulesqrt_of_epochestimatedregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_schedulesqrt_of_epochestimatedregret epoch-local expected estimated-regret certificates assemble into an expected moving-comparator square-root bound on the actual generated restart trajectory. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochEstimatedRegret","label":"integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochEstimatedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochEstimatedRegret","description":"Switch-count-facing expected restart bound under an explicit cardinality contract and epoch-local estimated-regret certificates.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-7adbe9a8db07","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10004,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:1029"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon switches : Nat) (hcomparator : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, epochComparator epoch ∈ arms) (hEpochCard : (oracleRestartScheduleEpochs schedule horizon).card <= switches + 1) (coefficient : Real) (hcoefficient : 0 <= coefficient) (hEstimated : ∀ epoch ∈ oracleRestartScheduleEpochs sche…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_scheduleswitchcountsqrt_of_epochestimatedregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_scheduleswitchcountsqrt_of_epochestimatedregret switch-count-facing expected restart bound under an explicit cardinality contract and epoch-local estimated-regret certificates. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochObservedEstimatedRegret","label":"integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochObservedEstimatedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochObservedEstimatedRegret","description":"Stored-reward epoch certificates assemble into the same expected moving-comparator square-root bound on the actual generated restart trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-af5195e55b0e","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10005,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:1099"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochObservedEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon : Nat) (hcomparator : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, epochComparator epoch ∈ arms) (coefficient : Real) (hcoefficient : 0 <= coefficient) (hObserved : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, integral (prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_schedulesqrt_of_epochobservedestimatedregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_schedulesqrt_of_epochobservedestimatedregret stored-reward epoch certificates assemble into the same expected moving-comparator square-root bound on the actual generated restart trajectory. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochObservedEstimatedRegret","label":"integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochObservedEstimatedRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochObservedEstimatedRegret","description":"Stored-reward epoch certificates imply the switch-count-facing expected restart bound under the explicit schedule-cardinality contract.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedregret/index.html#decl-0f7cbfdfc12e","parent":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","order":10006,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedRegret"],["Source","BanditRLProof/TsallisOracleRestartExpectedRegret.lean:1170"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochObservedEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon switches : Nat) (hcomparator : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, epochComparator epoch ∈ arms) (hEpochCard : (oracleRestartScheduleEpochs schedule horizon).card <= switches + 1) (coefficient : Real) (hcoefficient : 0 <= coefficient) (hObserved : ∀ epoch ∈ oracleRestartScheduleEpoc…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_scheduleswitchcountsqrt_of_epochobservedestimatedregret banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_scheduleswitchcountsqrt_of_epochobservedestimatedregret stored-reward epoch certificates imply the switch-count-facing expected restart bound under the explicit schedule-cardinality contract. theorem compiled","shard":"modules/f2bb346c779f5a7f.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartLocalTime","label":"oracleRestartLocalTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartLocalTime","description":"Local time of an actual round in its restart epoch.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-013571e21397","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10007,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleRestartLocalTime (schedule : OracleRestartSchedule) (t : Nat) : Nat","missing":[],"search":"oraclerestartlocaltime banditrlproof.tsallis.oraclerestartlocaltime local time of an actual round in its restart epoch. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAt","label":"sampledOracleRestartHalfTsallisHistoryAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAt","description":"Environment and visible global prefix before actual action `n + 1`.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-9c5dbfa72e35","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10008,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:26"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledOracleRestartHalfTsallisHistoryAt {Env : Type u} {Action : Type v} (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Env × History.FinitePairHistory Action Real n","missing":[],"search":"sampledoraclerestarthalftsallishistoryat banditrlproof.tsallis.sampledoraclerestarthalftsallishistoryat environment and visible global prefix before actual action `n + 1`. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisActionAt","label":"sampledOracleRestartHalfTsallisActionAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisActionAt","description":"Actual successor action after the visible global prefix through `n`.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-1e7c79a6ca78","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10009,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledOracleRestartHalfTsallisActionAt {Env : Type u} {Action : Type v} (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action","missing":[],"search":"sampledoraclerestarthalftsallisactionat banditrlproof.tsallis.sampledoraclerestarthalftsallisactionat actual successor action after the visible global prefix through `n`. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.frestrictLe_oracleRestartShiftedTrajectory_eq_localPairHistory","label":"frestrictLe_oracleRestartShiftedTrajectory_eq_localPairHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.frestrictLe_oracleRestartShiftedTrajectory_eq_localPairHistory","description":"Restricting a shifted trajectory to its local predecessor prefix is the same finite history as reindexing the corresponding global prefix.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-ac033fe5d1c9","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10010,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem frestrictLe_oracleRestartShiftedTrajectory_eq_localPairHistory {Env : Type u} {Action : Type v} (start n : Nat) (hstart : start <= n) (sample : Env × ((k : Nat) -> Action × Real)) : Preorder.frestrictLe (n - start) (oracleRestartShiftedTrajectory start sample).2 = oracleRestartLocalPairHistory start n hstart (Preorder.frestrictLe n sample.2)","missing":[],"search":"frestrictle_oraclerestartshiftedtrajectory_eq_localpairhistory banditrlproof.tsallis.frestrictle_oraclerestartshiftedtrajectory_eq_localpairhistory restricting a shifted trajectory to its local predecessor prefix is the same finite history as reindexing the corresponding global prefix. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisScoreAt","label":"sampledOracleRestartHalfTsallisScoreAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisScoreAt","description":"Restart-local cumulative score before actual action `n + 1`.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-2ff373a46406","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10011,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:52"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisScoreAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallisscoreat banditrlproof.tsallis.sampledoraclerestarthalftsallisscoreat restart-local cumulative score before actual action `n + 1`. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableLossAt","label":"sampledOracleRestartHalfTsallisPredictableLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableLossAt","description":"Predictable loss at actual successor time `n + 1`.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-35c8ba016423","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10012,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:68"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledOracleRestartHalfTsallisPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallispredictablelossat banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablelossat predictable loss at actual successor time `n + 1`. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisUpdatedAt","label":"sampledOracleRestartHalfTsallisUpdatedAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisUpdatedAt","description":"Same-local-rate update after actual action `n + 1`.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-6ecaf2fedfbb","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10013,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:76"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisUpdatedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallisupdatedat banditrlproof.tsallis.sampledoraclerestarthalftsallisupdatedat same-local-rate update after actual action `n + 1`. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","label":"sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","description":"Restart-local predictable history/action potential-stability score.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-09447ad59776","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10014,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:97"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : (Env × History.FinitePairHistory Action Real n) × Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.sampledoraclerestarthalftsallishistoryactionpotentialstabilityat restart-local predictable history/action potential-stability score. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","label":"sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","description":"Refined restart-local one-round stability budget.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-9d0373bc3ac7","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10015,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:115"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Real","missing":[],"search":"sampledoraclerestarthalftsallisrefinedpotentialstabilityboundat banditrlproof.tsallis.sampledoraclerestarthalftsallisrefinedpotentialstabilityboundat refined restart-local one-round stability budget. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor","label":"sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor","description":"Restart-local potential stability at actual successor time `n + 1`, written with the stored reward before the predictable-law rewrite.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-58b65c105a69","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10016,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:128"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallisobservedpotentialstabilityatsuccessor banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedpotentialstabilityatsuccessor restart-local potential stability at actual successor time `n + 1`, written with the stored reward before the predictable-law rewrite. definition compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor_eq_shifted","label":"sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor_eq_shifted","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor_eq_shifted","description":"The stored-reward restart-local potential term is exactly the scheduled potential term on the path shifted to the current epoch.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-4ce5304d8ba7","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10017,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:158"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor_eq_shifted {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor arms harms eta schedule n sample = sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory (schedule.start (n + 1)) sample) (oracleRestartLocalTime schedule (n + 1))","missing":[],"search":"sampledoraclerestarthalftsallisobservedpotentialstabilityatsuccessor_eq_shifted banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedpotentialstabilityatsuccessor_eq_shifted the stored-reward restart-local potential term is exactly the scheduled potential term on the path shifted to the current epoch. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_eq_historyAction_ae","label":"sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_eq_historyAction_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_eq_historyAction_ae","description":"On the global generated restart law, the shifted stored-reward local stability term agrees almost surely with the predictable restart-local history/action score.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-5b5bf087d977","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10018,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:238"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_eq_historyAction_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory (schedule.start (n + 1)) sample) (oracleRestartLocalTime schedule (n + 1))) =ᵐ[mu] (fun sample => sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt arms…","missing":[],"search":"sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_eq_historyaction_ae banditrlproof.tsallis.sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_eq_historyaction_ae on the global generated restart law, the shifted stored-reward local stability term agrees almost surely with the predictable restart-local history/action score. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisScoreAt","label":"measurable_sampledOracleRestartHalfTsallisScoreAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisScoreAt","description":"Supported coordinates of the restart-local cumulative score are measurable on the visible global prefix.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-786fbae6ee33","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10019,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:304"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledOracleRestartHalfTsallisScoreAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun input : Env × History.FinitePairHistory Action Real n => sampledOracleRestartHalfTsallisScoreAt arms harms eta schedule n input candidate)","missing":[],"search":"measurable_sampledoraclerestarthalftsallisscoreat banditrlproof.tsallis.measurable_sampledoraclerestarthalftsallisscoreat supported coordinates of the restart-local cumulative score are measurable on the visible global prefix. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisUpdatedAt","label":"measurable_sampledOracleRestartHalfTsallisUpdatedAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisUpdatedAt","description":"Supported coordinates of the restart-local same-rate update are measurable.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-d0566c146e11","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10020,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:333"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledOracleRestartHalfTsallisUpdatedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : (Env × History.FinitePairHistory Action Real n) × Action => sampledOracleRestartHalfTsallisUpdatedAt arms harms eta schedule loss n sample.1 sample.2 candidate)","missing":[],"search":"measurable_sampledoraclerestarthalftsallisupdatedat banditrlproof.tsallis.measurable_sampledoraclerestarthalftsallisupdatedat supported coordinates of the restart-local same-rate update are measurable. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","label":"measurable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","description":"The restart-local predictable history/action potential score is measurable.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-d6ab2b2567e3","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10021,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:404"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Measurable (sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt arms harms eta schedule loss n)","missing":[],"search":"measurable_sampledoraclerestarthalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.measurable_sampledoraclerestarthalftsallishistoryactionpotentialstabilityat the restart-local predictable history/action potential score is measurable. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","label":"integrable_sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","description":"The refined restart-local budget is integrable under every finite visible history law when the local rate is positive.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-ceebe40e4245","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10022,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:434"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (n : Nat) (historyMu : Measure (Env × History.FinitePairHistory Action Real n)) [IsFiniteMeasure historyMu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (heta : 0 < eta (oracleRestartLocalTime schedule (n + 1))) : Integrable (sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt (Env := Env) arms harms eta schedule n) historyMu","missing":[],"search":"integrable_sampledoraclerestarthalftsallisrefinedpotentialstabilityboundat banditrlproof.tsallis.integrable_sampledoraclerestarthalftsallisrefinedpotentialstabilityboundat the refined restart-local budget is integrable under every finite visible history law when the local rate is positive. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAt_isRegularizedMinimizer","label":"sampledOracleRestartHalfTsallisProbabilityAt_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAt_isRegularizedMinimizer","description":"The actual restart probability is the regularized minimizer of the restart-local pre-action score at the local learning rate.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-941df18dca0c","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10023,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:460"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisProbabilityAt_isRegularizedMinimizer {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (input : Env × History.FinitePairHistory Action Real n) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms (eta (oracleRestartLocalTime schedule (n + 1))) (negEntropyRegularizer arms (1 / 2 : Real)) (sampledOracleRestartHalfTsallisScoreAt arms harms eta schedule n input) (sampledOracleRestartHalfTsallisProbabilityAt arms harms eta schedule n input)","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityat_isregularizedminimizer banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityat_isregularizedminimizer the actual restart probability is the regularized minimizer of the restart-local pre-action score at the local learning rate. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisUpdatedAt_isRegularizedMinimizer","label":"sampledOracleRestartHalfTsallisUpdatedAt_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisUpdatedAt_isRegularizedMinimizer","description":"The same-local-rate restart update is the regularized minimizer after the predictable ordinary-IW increment.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-43316e045110","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10024,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:507"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisUpdatedAt_isRegularizedMinimizer {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (input : Env × History.FinitePairHistory Action Real n) (chosen : Action) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms (eta (oracleRestartLocalTime schedule (n + 1))) (negEntropyRegularizer arms (1 / 2 : Real)) (fun candidate => sampledOracleRestartHalfTsallisScoreAt arms harms eta schedule n input candidate + Exp3.importanceWeightedLoss (sampledOracleRestartHalfTsallisProbabilityAt arms harms eta schedule n input) (sampledOracleRestartHalfTsallisPredictableLossAt loss n input) chosen candidate) (sampledOracleRestartHalfTsallisUpdatedAt arms har…","missing":[],"search":"sampledoraclerestarthalftsallisupdatedat_isregularizedminimizer banditrlproof.tsallis.sampledoraclerestarthalftsallisupdatedat_isregularizedminimizer the same-local-rate restart update is the regularized minimizer after the predictable ordinary-iw increment. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","label":"integrable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","description":"The restart-local predictable potential score is automatically integrable under its visible-history/action finite kernel.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-519fcf1b5d6c","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10025,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:534"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (oracleRestartLocalTime schedule (n + 1))) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment Integrable (sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt arms harms eta schedule loss n) (mu.map (sampledOracleRestartHalfTsallisHistoryAt n) ⊗ₘ sampledOracleRestartHalfTsallisPolicyAt (Env := Env) arms harms eta sc…","missing":[],"search":"integrable_sampledoraclerestarthalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.integrable_sampledoraclerestarthalftsallishistoryactionpotentialstabilityat the restart-local predictable potential score is automatically integrable under its visible-history/action finite kernel. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt_le_one","label":"integral_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt_le_one","description":"Under the single generated restart law, the predictable restart-local successor stability score is integrable and has coarse expected budget one.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-814a55233a39","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10026,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:585"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt_le_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (oracleRestartLocalTime schedule (n + 1))) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment let term := fun sample => sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt arms harms eta schedule loss n (sampledOracleRestartHalfTsallisHistoryAt n sample, sampledOracleRestartHalfTsallisActionAt n sample) Integr…","missing":[],"search":"integral_sampledoraclerestarthalftsallishistoryactionpotentialstabilityat_le_one banditrlproof.tsallis.integral_sampledoraclerestarthalftsallishistoryactionpotentialstabilityat_le_one under the single generated restart law, the predictable restart-local successor stability score is integrable and has coarse expected budget one. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_one","label":"integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_one","description":"The actual shifted restart-local successor stability term is integrable and has coarse expected budget one under the single global restart law.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-8f526c51de54","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10027,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:696"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (oracleRestartLocalTime schedule (n + 1))) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment let term := fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory (schedule.start (n + 1)) sample) (oracleRestartLocalTime schedule (n + 1)) Integrable term mu ∧ integr…","missing":[],"search":"integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_le_one banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_le_one the actual shifted restart-local successor stability term is integrable and has coarse expected budget one under the single global restart law. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_refined","label":"integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_refined","description":"At local rates at most one half, the actual shifted restart-local successor stability term has the refined expected conjugate-potential bound under the single global restart law.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-933dd55a649e","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10028,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:753"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_refined {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (oracleRestartLocalTime schedule (n + 1))) (heta_le : eta (oracleRestartLocalTime schedule (n + 1)) <= 1 / 2) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment let term := fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory (schedule.start (n + 1)) sample…","missing":[],"search":"integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_le_refined banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_le_refined at local rates at most one half, the actual shifted restart-local successor stability term has the refined expected conjugate-potential bound under the single global restart law. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisInitialPotentialStabilityAtTime_le_one","label":"integral_sampledOracleRestartHalfTsallisInitialPotentialStabilityAtTime_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisInitialPotentialStabilityAtTime_le_one","description":"At global time zero, the restart process has the canonical initial half-Tsallis action law and the usual coarse expected stability budget.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-f698b49faee6","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10029,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:897"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisInitialPotentialStabilityAtTime_le_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (heta : 0 < eta 0) (loss : Exp3.PredictableLossVector Env Action) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory 0 sample) 0) mu ∧ integral mu (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory 0 sample) 0)…","missing":[],"search":"integral_sampledoraclerestarthalftsallisinitialpotentialstabilityattime_le_one banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisinitialpotentialstabilityattime_le_one at global time zero, the restart process has the canonical initial half-tsallis action law and the usual coarse expected stability budget. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtTime_le_integral_one","label":"integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtTime_le_integral_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtTime_le_integral_one","description":"Every actual restart time, including global time zero and later restart boundaries, has the mass-scaled coarse local stability budget under the one global generated law.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-315722f80f20","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10030,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:1073"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtTime_le_integral_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (heta : 0 < eta (oracleRestartLocalTime schedule t)) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment let term := fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta (oracleRestartShiftedTrajectory (schedule.start t) sample) (oracleRestartLocalTime schedule t) Integrable term mu ∧ integral mu term <=…","missing":[],"search":"integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityattime_le_integral_one banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityattime_le_integral_one every actual restart time, including global time zero and later restart boundaries, has the mass-scaled coarse local stability budget under the one global generated law. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_le_card","label":"integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_le_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_le_card","description":"On a deterministic contiguous restart epoch, the complete shifted local stability prefix is integrable and its expectation is at most the prefix cardinality. The probability assumption turns the finite-measure one-round budget into the literal constant `1`.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-9f484133efa6","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10031,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:1109"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_le_card {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epoch localHorizon : Nat) (hstart : ∀ localTime, localTime ≤ localHorizon -> schedule.start (epoch + localTime) = epoch) (heta : ∀ localTime, localTime ≤ localHorizon -> 0 < eta localTime) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment let stabilitySum := fun sample => (Finset.range (localHorizon + 1)).sum (fun localTime => samp…","missing":[],"search":"integral_sum_sampledoraclerestarthalftsallisshiftedpotentialstabilityatlocalprefix_le_card banditrlproof.tsallis.integral_sum_sampledoraclerestarthalftsallisshiftedpotentialstabilityatlocalprefix_le_card on a deterministic contiguous restart epoch, the complete shifted local stability prefix is integrable and its expectation is at most the prefix cardinality. the probability assumption turns the finite-measure one-round budget into the literal constant `1`. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_card_add_penalty_of_epochRounds_eq","label":"integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_card_add_penalty_of_epochRounds_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_card_add_penalty_of_epochRounds_eq","description":"A contiguous actual restart epoch inherits an expected observed estimated-regret certificate from the one global generated law. This coarse endpoint is linear in the epoch cardinality; obtaining the target square-root certificate still requires summing and tuning the refined one-round bounds.","url":"../modules/banditrlproof-tsallisoraclerestartexpectedstability/index.html#decl-5b08552da672","parent":"module:BanditRLProof.TsallisOracleRestartExpectedStability","order":10032,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartExpectedStability"],["Source","BanditRLProof/TsallisOracleRestartExpectedStability.lean:1188"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_card_add_penalty_of_epochRounds_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (horizon epoch localHorizon : Nat) {best : Action} (hbest : best ∈ arms) (hRounds : oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime)) (heta : ∀ localTime, localTime ≤ localHorizon -> 0 < eta localTime) (hetaMono : ∀ localTime, localTime < localHorizon -> eta (localTime + 1) ≤…","missing":[],"search":"integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_le_card_add_penalty_of_epochrounds_eq banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_le_card_add_penalty_of_epochrounds_eq a contiguous actual restart epoch inherits an expected observed estimated-regret certificate from the one global generated law. this coarse endpoint is linear in the epoch cardinality; obtaining the target square-root certificate still requires summing and tuning the refined one-round bounds. theorem compiled","shard":"modules/5fe8711e9fdf9945.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSqrt","label":"integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSqrt","description":"The generated square-root-schedule epoch certificate assembles into the schedule-cardinality-facing expected moving-comparator regret bound.","url":"../modules/banditrlproof-tsallisoraclerestartgenerateddynamicregret/index.html#decl-1abfee1efa48","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","order":10033,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartGeneratedDynamicRegret.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon : Nat) (hcomparator : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, epochComparator epoch ∈ arms) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule schedule loss.environment Integrable (sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret arms h…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_sqrtschedule_le_schedulesqrt banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_sqrtschedule_le_schedulesqrt the generated square-root-schedule epoch certificate assembles into the schedule-cardinality-facing expected moving-comparator regret bound. theorem compiled","shard":"modules/163d2cea7bd3fa4a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSwitchCountSqrt","label":"integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSwitchCountSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSwitchCountSqrt","description":"Under an explicit schedule-epoch count contract, the same generated certificate yields the switch-count-facing expected dynamic-regret bound.","url":"../modules/banditrlproof-tsallisoraclerestartgenerateddynamicregret/index.html#decl-ec4297f66200","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","order":10034,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret"],["Source","BanditRLProof/TsallisOracleRestartGeneratedDynamicRegret.lean:67"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSwitchCountSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon switches : Nat) (hcomparator : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, epochComparator epoch ∈ arms) (hEpochCard : (oracleRestartScheduleEpochs schedule horizon).card <= switches + 1) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule schedule loss.env…","missing":[],"search":"integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_sqrtschedule_le_scheduleswitchcountsqrt banditrlproof.tsallis.integral_sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_sqrtschedule_le_scheduleswitchcountsqrt under an explicit schedule-epoch count contract, the same generated certificate yields the switch-count-facing expected dynamic-regret bound. theorem compiled","shard":"modules/163d2cea7bd3fa4a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule","label":"OracleRestartSchedule","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.OracleRestartSchedule","description":"Start time of the current epoch, with either continuation or a fresh restart at every successor time.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-cf75e3fb4479","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10035,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure OracleRestartSchedule where","missing":[],"search":"oraclerestartschedule banditrlproof.tsallis.oraclerestartschedule start time of the current epoch, with either continuation or a fresh restart at every successor time. structure compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleNeverRestartSchedule","label":"oracleNeverRestartSchedule","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleNeverRestartSchedule","description":"One epoch containing the whole trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-c12e13e75952","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10036,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleNeverRestartSchedule : OracleRestartSchedule where","missing":[],"search":"oracleneverrestartschedule banditrlproof.tsallis.oracleneverrestartschedule one epoch containing the whole trajectory. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartEveryRoundSchedule","label":"oracleRestartEveryRoundSchedule","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartEveryRoundSchedule","description":"A fresh one-round epoch at every actual time.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-3668c632a97a","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10037,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:34"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleRestartEveryRoundSchedule : OracleRestartSchedule where","missing":[],"search":"oraclerestarteveryroundschedule banditrlproof.tsallis.oraclerestarteveryroundschedule a fresh one-round epoch at every actual time. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.start_succ_le_of_ne","label":"start_succ_le_of_ne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.OracleRestartSchedule.start_succ_le_of_ne","description":"Away from a restart boundary, the current epoch starts no later than the last observed round.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-e152b53177bd","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10038,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:42"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem OracleRestartSchedule.start_succ_le_of_ne (schedule : OracleRestartSchedule) (n : Nat) (hboundary : schedule.start (n + 1) ≠ n + 1) : schedule.start (n + 1) <= n","missing":[],"search":"start_succ_le_of_ne banditrlproof.tsallis.oraclerestartschedule.start_succ_le_of_ne away from a restart boundary, the current epoch starts no later than the last observed round. theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartLocalPairHistory","label":"oracleRestartLocalPairHistory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartLocalPairHistory","description":"Reindex the inclusive global history segment `start..n` as a local history through `n-start`.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-8a136cdc3259","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10039,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:51"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleRestartLocalPairHistory {Action Reward : Type*} (start n : Nat) (hstart : start <= n) (history : History.FinitePairHistory Action Reward n) : History.FinitePairHistory Action Reward (n - start)","missing":[],"search":"oraclerestartlocalpairhistory banditrlproof.tsallis.oraclerestartlocalpairhistory reindex the inclusive global history segment `start..n` as a local history through `n-start`. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartLocalPairHistory_zero","label":"oracleRestartLocalPairHistory_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartLocalPairHistory_zero","description":"theorem oracleRestartLocalPairHistory_zero {Action Reward : Type*} (n : Nat) (history : History.FinitePairHistory Action Reward n) : oracleRestartLocalPairHistory 0 n (Nat.zero_le n) history = history","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-bb7fd4ea06da","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10040,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartLocalPairHistory_zero {Action Reward : Type*} (n : Nat) (history : History.FinitePairHistory Action Reward n) : oracleRestartLocalPairHistory 0 n (Nat.zero_le n) history = history","missing":[],"search":"oraclerestartlocalpairhistory_zero banditrlproof.tsallis.oraclerestartlocalpairhistory_zero theorem oraclerestartlocalpairhistory_zero {action reward : type*} (n : nat) (history : history.finitepairhistory action reward n) : oraclerestartlocalpairhistory 0 n (nat.zero_le n) history = history theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_oracleRestartLocalPairHistory","label":"measurable_oracleRestartLocalPairHistory","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_oracleRestartLocalPairHistory","description":"Epoch-suffix reindexing is measurable coordinatewise.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-7bec124263aa","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10041,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:69"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_oracleRestartLocalPairHistory {Action Reward : Type*} [MeasurableSpace Action] [MeasurableSpace Reward] (start n : Nat) (hstart : start <= n) : Measurable (oracleRestartLocalPairHistory (Action := Action) (Reward := Reward) start n hstart)","missing":[],"search":"measurable_oraclerestartlocalpairhistory banditrlproof.tsallis.measurable_oraclerestartlocalpairhistory epoch-suffix reindexing is measurable coordinatewise. theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution","label":"sampledOracleRestartHalfTsallisHistoryDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution","description":"Restarted successor action distribution. Boundary times use a fresh initial law; continuation times use only the current epoch's local suffix.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-a3810364a540","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10042,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:86"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisHistoryDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledoraclerestarthalftsallishistorydistribution banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistribution restarted successor action distribution. boundary times use a fresh initial law; continuation times use only the current epoch's local suffix. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_of_boundary","label":"sampledOracleRestartHalfTsallisHistoryDistribution_of_boundary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_of_boundary","description":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_of_boundary {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (hboundary : schedule.start (n + 1) = n + 1) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n history = initialHalf…","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-b85bdc92f196","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10043,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:102"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_of_boundary {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (hboundary : schedule.start (n + 1) = n + 1) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n history = initialHalfTsallisDistribution arms harms (eta 0)","missing":[],"search":"sampledoraclerestarthalftsallishistorydistribution_of_boundary banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistribution_of_boundary theorem sampledoraclerestarthalftsallishistorydistribution_of_boundary {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (n : nat) (hboundary : schedule.start (n + 1) = n + 1) (history : history.finitepairhistory action real n) : sampledoraclerestarthalftsallishistorydistribution arms harms eta schedule n history = initialhalftsallisdistribution arms harms (eta 0) theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_of_continuation","label":"sampledOracleRestartHalfTsallisHistoryDistribution_of_continuation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_of_continuation","description":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_of_continuation {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (hboundary : schedule.start (n + 1) ≠ n + 1) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n history = sampled…","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-d1e1ead05fb7","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10044,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:114"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_of_continuation {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (hboundary : schedule.start (n + 1) ≠ n + 1) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n history = sampledScheduledHalfTsallisHistoryDistribution arms harms eta (n - schedule.start (n + 1)) (oracleRestartLocalPairHistory (schedule.start (n + 1)) n (schedule.start_succ_le_of_ne n hboundary) history)","missing":[],"search":"sampledoraclerestarthalftsallishistorydistribution_of_continuation banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistribution_of_continuation theorem sampledoraclerestarthalftsallishistorydistribution_of_continuation {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (n : nat) (hboundary : schedule.start (n + 1) ≠ n + 1) (history : history.finitepairhistory action real n) : sampledoraclerestarthalftsallishistorydistribution arms harms eta schedule n history = sampledscheduledhalftsallishistorydistribution arms harms eta (n - schedule.start (n + 1)) (oraclerestartlocalpairhistory (schedule.start (n + 1)) n (schedule.start_succ_le_of_ne n hboundary) history) theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_pos","label":"sampledOracleRestartHalfTsallisHistoryDistribution_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_pos","description":"Every arm in the finite action set keeps strictly positive probability under either the boundary reset or the continued local scheduled policy.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-50f2bfa9f22f","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10045,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:131"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_pos {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) (history : History.FinitePairHistory Action Real n) (candidate : Action) (hcandidate : candidate ∈ arms) : 0 < sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n history candidate","missing":[],"search":"sampledoraclerestarthalftsallishistorydistribution_pos banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistribution_pos every arm in the finite action set keeps strictly positive probability under either the boundary reset or the continued local scheduled policy. theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_neverRestart","label":"sampledOracleRestartHalfTsallisHistoryDistribution_neverRestart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_neverRestart","description":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_neverRestart {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta oracleNeverRestartSchedule n history = sampledScheduledHalfTsallisHistoryDistribution arms harms eta n history","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-521a2e54c366","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10046,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:166"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_neverRestart {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta oracleNeverRestartSchedule n history = sampledScheduledHalfTsallisHistoryDistribution arms harms eta n history","missing":[],"search":"sampledoraclerestarthalftsallishistorydistribution_neverrestart banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistribution_neverrestart theorem sampledoraclerestarthalftsallishistorydistribution_neverrestart {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (n : nat) (history : history.finitepairhistory action real n) : sampledoraclerestarthalftsallishistorydistribution arms harms eta oracleneverrestartschedule n history = sampledscheduledhalftsallishistorydistribution arms harms eta n history theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_restartEveryRound","label":"sampledOracleRestartHalfTsallisHistoryDistribution_restartEveryRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_restartEveryRound","description":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_restartEveryRound {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta oracleRestartEveryRoundSchedule n history = initialHalfTsallisDistribution arms harms (eta 0)","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-1d2d972c9c8a","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10047,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:179"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisHistoryDistribution_restartEveryRound {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (history : History.FinitePairHistory Action Real n) : sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta oracleRestartEveryRoundSchedule n history = initialHalfTsallisDistribution arms harms (eta 0)","missing":[],"search":"sampledoraclerestarthalftsallishistorydistribution_restarteveryround banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistribution_restarteveryround theorem sampledoraclerestarthalftsallishistorydistribution_restarteveryround {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (n : nat) (history : history.finitepairhistory action real n) : sampledoraclerestarthalftsallishistorydistribution arms harms eta oraclerestarteveryroundschedule n history = initialhalftsallisdistribution arms harms (eta 0) theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistributionSource","label":"sampledOracleRestartHalfTsallisHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistributionSource","description":"Measurable finite-action source for every restarted successor policy.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-a74f87220f7c","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10048,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:190"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisHistoryDistributionSource {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : Exp3.MeasurableFiniteActionDistribution arms (sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n)","missing":[],"search":"sampledoraclerestarthalftsallishistorydistributionsource banditrlproof.tsallis.sampledoraclerestarthalftsallishistorydistributionsource measurable finite-action source for every restarted successor policy. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAlgorithm","label":"sampledOracleRestartHalfTsallisHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAlgorithm","description":"Stochastic finite-history algorithm whose score resets at every scheduled epoch boundary.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-86066d62a6eb","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10049,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:232"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisHistoryAlgorithm {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) : Thompson.HistoryAlgorithm Action Real where","missing":[],"search":"sampledoraclerestarthalftsallishistoryalgorithm banditrlproof.tsallis.sampledoraclerestarthalftsallishistoryalgorithm stochastic finite-history algorithm whose score resets at every scheduled epoch boundary. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAlgorithm_policy","label":"sampledOracleRestartHalfTsallisHistoryAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAlgorithm_policy","description":"theorem sampledOracleRestartHalfTsallisHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : (sampledOracleRestartHalfTsallisHistoryAlgorithm arms harms eta schedule).policy n = Exp3.finiteActionKernel arms (sampledOracleRestartHalfTsall…","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-5210f3601925","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10050,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:259"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (n : Nat) : (sampledOracleRestartHalfTsallisHistoryAlgorithm arms harms eta schedule).policy n = Exp3.finiteActionKernel arms (sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n) (sampledOracleRestartHalfTsallisHistoryDistributionSource arms harms eta schedule n)","missing":[],"search":"sampledoraclerestarthalftsallishistoryalgorithm_policy banditrlproof.tsallis.sampledoraclerestarthalftsallishistoryalgorithm_policy theorem sampledoraclerestarthalftsallishistoryalgorithm_policy {action : type u} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (schedule : oraclerestartschedule) (n : nat) : (sampledoraclerestarthalftsallishistoryalgorithm arms harms eta schedule).policy n = exp3.finiteactionkernel arms (sampledoraclerestarthalftsallishistorydistribution arms harms eta schedule n) (sampledoraclerestarthalftsallishistorydistributionsource arms harms eta schedule n) theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime","label":"sampledOracleRestartHalfTsallisProbabilityAtTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime","description":"Restarted sampling probabilities at actual trajectory times.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-c4b9284ba076","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10051,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:275"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisProbabilityAtTime {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) : Nat -> Env × ((k : Nat) -> Action × Real) -> Action -> Real | 0, _sample => initialHalfTsallisDistribution arms harms (eta 0) | n + 1, sample => sampledOracleRestartHalfTsallisHistoryDistribution arms harms eta schedule n (Preorder.frestrictLe n sample.2) @[simp] theorem sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta oracleNeverRestartSchedule t sample = sampledScheduledHalfTsallisProbabilityAtTime arms harms eta…","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityattime banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityattime restarted sampling probabilities at actual trajectory times. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart","label":"sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart","description":"theorem sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta oracleNeverRestartSchedule t sample = sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-1879c3e0a9c7","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10052,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:286"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta oracleNeverRestartSchedule t sample = sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityattime_neverrestart banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityattime_neverrestart theorem sampledoraclerestarthalftsallisprobabilityattime_neverrestart {env : type v} {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (t : nat) (sample : env × ((k : nat) -> action × real)) : sampledoraclerestarthalftsallisprobabilityattime arms harms eta oracleneverrestartschedule t sample = sampledscheduledhalftsallisprobabilityattime arms harms eta t sample theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_restartEveryRound","label":"sampledOracleRestartHalfTsallisProbabilityAtTime_restartEveryRound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_restartEveryRound","description":"theorem sampledOracleRestartHalfTsallisProbabilityAtTime_restartEveryRound {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta oracleRestartEveryRoundSchedule t sample = initialHalfTsallisDistribution arms harms (eta 0)","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-16436f67cdef","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10053,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:301"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisProbabilityAtTime_restartEveryRound {Env : Type v} {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta oracleRestartEveryRoundSchedule t sample = initialHalfTsallisDistribution arms harms (eta 0)","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityattime_restarteveryround banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityattime_restarteveryround theorem sampledoraclerestarthalftsallisprobabilityattime_restarteveryround {env : type v} {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (t : nat) (sample : env × ((k : nat) -> action × real)) : sampledoraclerestarthalftsallisprobabilityattime arms harms eta oraclerestarteveryroundschedule t sample = initialhalftsallisdistribution arms harms (eta 0) theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryKernel","label":"sampledOracleRestartHalfTsallisTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryKernel","description":"Full environment-indexed trajectory generated by the restarted policy.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-e3f9bf2371a1","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10054,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:314"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisTrajectoryKernel {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] [Nonempty Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) : Kernel Env ((n : Nat) -> Action × Real)","missing":[],"search":"sampledoraclerestarthalftsallistrajectorykernel banditrlproof.tsallis.sampledoraclerestarthalftsallistrajectorykernel full environment-indexed trajectory generated by the restarted policy. definition compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action","label":"sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action","description":"The generated successor action has the restarted finite-action law conditional on the complete visible global history.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-ce77cecf9528","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10055,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:343"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule environment) =ᵐ[ (prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule environment).map (fun sample => Preorder.frestrictLe n sample.2)] Exp3.f…","missing":[],"search":"sampledoraclerestarthalftsallistrajectorymeasure_conddistrib_action banditrlproof.tsallis.sampledoraclerestarthalftsallistrajectorymeasure_conddistrib_action the generated successor action has the restarted finite-action law conditional on the complete visible global history. theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","label":"sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","description":"The restarted successor action law after retaining both the environment and the visible global prefix. This is the conditioning surface needed by predictable-loss first-moment transport.","url":"../modules/banditrlproof-tsallisoraclerestartgeneratedtrajectory/index.html#decl-2ae59016d64f","parent":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","order":10056,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGeneratedTrajectory"],["Source","BanditRLProof/TsallisOracleRestartGeneratedTrajectory.lean:378"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2)) (prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule environment) =ᵐ[ (prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule envir…","missing":[],"search":"sampledoraclerestarthalftsallistrajectorymeasure_conddistrib_action_given_environment banditrlproof.tsallis.sampledoraclerestarthalftsallistrajectorymeasure_conddistrib_action_given_environment the restarted successor action law after retaining both the environment and the visible global prefix. this is the conditioning surface needed by predictable-loss first-moment transport. theorem compiled","shard":"modules/eb8cb529c6ce11ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartStart","label":"oracleChangePointRestartStart","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleChangePointRestartStart","description":"Start of the most recent epoch when a restart occurs after every `change` boundary. A true value at `s` starts the new epoch at `s + 1`.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-b98a595c6af0","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10057,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleChangePointRestartStart (change : Nat -> Prop) [DecidablePred change] : Nat -> Nat | 0 => 0 | n + 1 => if change n then n + 1 else oracleChangePointRestartStart change n /-- A valid restart schedule generated by an arbitrary decidable boundary predicate. -/ def oracleChangePointRestartSchedule (change : Nat -> Prop) [DecidablePred change] : OracleRestartSchedule where","missing":[],"search":"oraclechangepointrestartstart banditrlproof.tsallis.oraclechangepointrestartstart start of the most recent epoch when a restart occurs after every `change` boundary. a true value at `s` starts the new epoch at `s + 1`. definition compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartSchedule","label":"oracleChangePointRestartSchedule","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleChangePointRestartSchedule","description":"A valid restart schedule generated by an arbitrary decidable boundary predicate.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-e8c5886363ab","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10058,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:29"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleChangePointRestartSchedule (change : Nat -> Prop) [DecidablePred change] : OracleRestartSchedule where","missing":[],"search":"oraclechangepointrestartschedule banditrlproof.tsallis.oraclechangepointrestartschedule a valid restart schedule generated by an arbitrary decidable boundary predicate. definition compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartSchedule_start_zero","label":"oracleChangePointRestartSchedule_start_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleChangePointRestartSchedule_start_zero","description":"theorem oracleChangePointRestartSchedule_start_zero (change : Nat -> Prop) [DecidablePred change] : (oracleChangePointRestartSchedule change).start 0 = 0","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-fec94c76cf4c","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10059,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleChangePointRestartSchedule_start_zero (change : Nat -> Prop) [DecidablePred change] : (oracleChangePointRestartSchedule change).start 0 = 0","missing":[],"search":"oraclechangepointrestartschedule_start_zero banditrlproof.tsallis.oraclechangepointrestartschedule_start_zero theorem oraclechangepointrestartschedule_start_zero (change : nat -> prop) [decidablepred change] : (oraclechangepointrestartschedule change).start 0 = 0 theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartSchedule_start_succ","label":"oracleChangePointRestartSchedule_start_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleChangePointRestartSchedule_start_succ","description":"theorem oracleChangePointRestartSchedule_start_succ (change : Nat -> Prop) [DecidablePred change] (n : Nat) : (oracleChangePointRestartSchedule change).start (n + 1) = if change n then n + 1 else (oracleChangePointRestartSchedule change).start n","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-880a8a7989be","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10060,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:52"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleChangePointRestartSchedule_start_succ (change : Nat -> Prop) [DecidablePred change] (n : Nat) : (oracleChangePointRestartSchedule change).start (n + 1) = if change n then n + 1 else (oracleChangePointRestartSchedule change).start n","missing":[],"search":"oraclechangepointrestartschedule_start_succ banditrlproof.tsallis.oraclechangepointrestartschedule_start_succ theorem oraclechangepointrestartschedule_start_succ (change : nat -> prop) [decidablepred change] (n : nat) : (oraclechangepointrestartschedule change).start (n + 1) = if change n then n + 1 else (oraclechangepointrestartschedule change).start n theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_oracleChangePointRestartSchedule","label":"oracleRestartScheduleEpochs_oracleChangePointRestartSchedule","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartScheduleEpochs_oracleChangePointRestartSchedule","description":"The registered epochs are zero together with every successor of a change point before the inclusive horizon.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-8ce0d7c0ca86","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10061,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:60"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartScheduleEpochs_oracleChangePointRestartSchedule (change : Nat -> Prop) [DecidablePred change] (horizon : Nat) : oracleRestartScheduleEpochs (oracleChangePointRestartSchedule change) horizon = insert 0 (((Finset.range horizon).filter change).image (fun s => s + 1))","missing":[],"search":"oraclerestartscheduleepochs_oraclechangepointrestartschedule banditrlproof.tsallis.oraclerestartscheduleepochs_oraclechangepointrestartschedule the registered epochs are zero together with every successor of a change point before the inclusive horizon. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_oracleChangePointRestartSchedule","label":"oracleRestartScheduleEpochs_card_oracleChangePointRestartSchedule","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_oracleChangePointRestartSchedule","description":"A change-point restart schedule has exactly one more visited epoch than the number of change boundaries before the inclusive horizon.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-64109f3a9ba1","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10062,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:99"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartScheduleEpochs_card_oracleChangePointRestartSchedule (change : Nat -> Prop) [DecidablePred change] (horizon : Nat) : (oracleRestartScheduleEpochs (oracleChangePointRestartSchedule change) horizon).card = ((Finset.range horizon).filter change).card + 1","missing":[],"search":"oraclerestartscheduleepochs_card_oraclechangepointrestartschedule banditrlproof.tsallis.oraclerestartscheduleepochs_card_oraclechangepointrestartschedule a change-point restart schedule has exactly one more visited epoch than the number of change boundaries before the inclusive horizon. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeAt","label":"finiteArmIndependentGlobalMeanChangeAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeAt","description":"A boundary has a global population-mean change when at least one arm changes mean between its two adjacent reward laws.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-f8b5e5c13c9e","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10063,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:113"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def finiteArmIndependentGlobalMeanChangeAt {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (s : Nat) : Prop","missing":[],"search":"finitearmindependentglobalmeanchangeat banditrlproof.tsallis.finitearmindependentglobalmeanchangeat a boundary has a global population-mean change when at least one arm changes mean between its two adjacent reward laws. definition compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeTimes","label":"finiteArmIndependentGlobalMeanChangeTimes","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeTimes","description":"Global population-mean change boundaries strictly before `horizon`.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-e9b184bf444d","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10064,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:120"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentGlobalMeanChangeTimes {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : Finset Nat","missing":[],"search":"finitearmindependentglobalmeanchangetimes banditrlproof.tsallis.finitearmindependentglobalmeanchangetimes global population-mean change boundaries strictly before `horizon`. definition compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeRestartSchedule","label":"finiteArmIndependentGlobalMeanChangeRestartSchedule","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeRestartSchedule","description":"Population-mean oracle schedule: restart immediately after every global mean-change boundary.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-67af44122415","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10065,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:128"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def finiteArmIndependentGlobalMeanChangeRestartSchedule {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) : OracleRestartSchedule","missing":[],"search":"finitearmindependentglobalmeanchangerestartschedule banditrlproof.tsallis.finitearmindependentglobalmeanchangerestartschedule population-mean oracle schedule: restart immediately after every global mean-change boundary. definition compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_finiteArmIndependentGlobalMeanChangeRestartSchedule","label":"oracleRestartScheduleEpochs_card_finiteArmIndependentGlobalMeanChangeRestartSchedule","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_finiteArmIndependentGlobalMeanChangeRestartSchedule","description":"The population-mean restart schedule has exactly one initial epoch plus one epoch for every global mean-change boundary.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-33380777d65e","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10066,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:137"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartScheduleEpochs_card_finiteArmIndependentGlobalMeanChangeRestartSchedule {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : (oracleRestartScheduleEpochs (finiteArmIndependentGlobalMeanChangeRestartSchedule armLaw) horizon).card = (finiteArmIndependentGlobalMeanChangeTimes armLaw horizon).card + 1","missing":[],"search":"oraclerestartscheduleepochs_card_finitearmindependentglobalmeanchangerestartschedule banditrlproof.tsallis.oraclerestartscheduleepochs_card_finitearmindependentglobalmeanchangerestartschedule the population-mean restart schedule has exactly one initial epoch plus one epoch for every global mean-change boundary. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_changeTimes_card","label":"finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_changeTimes_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_changeTimes_card","description":"The existing real-valued global switch count is the coercion of the named change-time finset cardinality.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-1249d703d21f","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10067,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:153"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_changeTimes_card {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw horizon = ((finiteArmIndependentGlobalMeanChangeTimes armLaw horizon).card : Real)","missing":[],"search":"finitearmindependentcumulativeglobalmeanswitchcount_eq_changetimes_card banditrlproof.tsallis.finitearmindependentcumulativeglobalmeanswitchcount_eq_changetimes_card the existing real-valued global switch count is the coercion of the named change-time finset cardinality. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_cast_finiteArmIndependentGlobalMeanChangeRestartSchedule","label":"oracleRestartScheduleEpochs_card_cast_finiteArmIndependentGlobalMeanChangeRestartSchedule","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_cast_finiteArmIndependentGlobalMeanChangeRestartSchedule","description":"In real-valued form, the exact epoch count is the existing cumulative global population-mean switch count plus one.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-765eefcbe1dc","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10068,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:171"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartScheduleEpochs_card_cast_finiteArmIndependentGlobalMeanChangeRestartSchedule {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (horizon : Nat) : ((oracleRestartScheduleEpochs (finiteArmIndependentGlobalMeanChangeRestartSchedule armLaw) horizon).card : Real) = finiteArmIndependentCumulativeGlobalMeanSwitchCount armLaw horizon + 1","missing":[],"search":"oraclerestartscheduleepochs_card_cast_finitearmindependentglobalmeanchangerestartschedule banditrlproof.tsallis.oraclerestartscheduleepochs_card_cast_finitearmindependentglobalmeanchangerestartschedule in real-valued form, the exact epoch count is the existing cumulative global population-mean switch count plus one. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_eq_globalMeanChangeRestartStart","label":"finiteArmIndependentRewardMean_eq_globalMeanChangeRestartStart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardMean_eq_globalMeanChangeRestartStart","description":"Since every global mean change starts a new epoch, each arm's population mean is constant from the current epoch start through the current round.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-23f87a436505","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10069,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:186"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentRewardMean_eq_globalMeanChangeRestartStart {K : Nat} (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : finiteArmIndependentRewardMean armLaw t arm = finiteArmIndependentRewardMean armLaw ((finiteArmIndependentGlobalMeanChangeRestartSchedule armLaw).start t) arm","missing":[],"search":"finitearmindependentrewardmean_eq_globalmeanchangerestartstart banditrlproof.tsallis.finitearmindependentrewardmean_eq_globalmeanchangerestartstart since every global mean change starts a new epoch, each arm's population mean is constant from the current epoch start through the current round. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_le_globalMeanChangeRestartBestArm","label":"finiteArmIndependentRewardMean_le_globalMeanChangeRestartBestArm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteArmIndependentRewardMean_le_globalMeanChangeRestartBestArm","description":"The arm maximizing population mean at the current epoch start remains a population-mean maximizer at every round in that epoch.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-764f7b81bc0d","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10070,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:218"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteArmIndependentRewardMean_le_globalMeanChangeRestartBestArm {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (t : Nat) (arm : Fin K) : finiteArmIndependentRewardMean armLaw t arm <= finiteArmIndependentRewardMean armLaw t (finiteArmIndependentBestArmAt model armLaw ((finiteArmIndependentGlobalMeanChangeRestartSchedule armLaw).start t))","missing":[],"search":"finitearmindependentrewardmean_le_globalmeanchangerestartbestarm banditrlproof.tsallis.finitearmindependentrewardmean_le_globalmeanchangerestartbestarm the arm maximizing population mean at the current epoch start remains a population-mean maximizer at every round in that epoch. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_globalMeanChangeRestartBestArm_nonneg","label":"independentLossStateTimeVaryingMeanGap_globalMeanChangeRestartBestArm_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_globalMeanChangeRestartBestArm_nonneg","description":"Under unit support, the raw-population-mean oracle comparator is also optimal for the clipped loss used by the generated trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-63457efd9355","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10071,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:243"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem independentLossStateTimeVaryingMeanGap_globalMeanChangeRestartBestArm_nonneg {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : forall t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : forall t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (t : Nat) (arm : Fin K) : 0 <= independentLossStateTimeVaryingMeanGap (finiteArmIndependentRewardVectorLaw armLaw) (fun _ => finiteArmIIDRewardVectorLoss) t (finiteArmIndependentBestArmAt model armLaw ((finiteArmIndependentGlobalMeanChangeRestartSchedule armLaw).start t)) arm","missing":[],"search":"independentlossstatetimevaryingmeangap_globalmeanchangerestartbestarm_nonneg banditrlproof.tsallis.independentlossstatetimevaryingmeangap_globalmeanchangerestartbestarm_nonneg under unit support, the raw-population-mean oracle comparator is also optimal for the clipped loss used by the generated trajectory. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le","label":"integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le","description":"The concrete independent reward law, restarted at every global population-mean change, has generated expected dynamic regret bounded by the square root of its exact population-mean switch count. The comparator selected at each epoch start remains optimal for the clipped loss throughout that epoch under the almost-sure unit-support contract.","url":"../modules/banditrlproof-tsallisoraclerestartglobalmeanswitchcount/index.html#decl-6678d5fce9a1","parent":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","order":10072,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount"],["Source","BanditRLProof/TsallisOracleRestartGlobalMeanSwitchCount.lean:272"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","delayed-nonstationary"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le {K : Nat} (model : FiniteBanditModel K) (armLaw : Nat -> Fin K -> Measure Rat) (hprob : ∀ t arm, IsProbabilityMeasure (armLaw t arm)) (hbound : ∀ t arm, ∀ᵐ reward ∂armLaw t arm, ((reward : Rat) : Real) ∈ Set.Icc (0 : Real) 1) (horizon : Nat) : letI : Nonempty (Fin K)","missing":[],"search":"integral_sampledoraclerestarthalftsallisfinitearmindependentglobalmeanchangedynamicregret_le banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisfinitearmindependentglobalmeanchangedynamicregret_le the concrete independent reward law, restarted at every global population-mean change, has generated expected dynamic regret bounded by the square root of its exact population-mean switch count. the comparator selected at each epoch start remains optimal for the clipped loss throughout that epoch under the almost-sure unit-support contract. theorem compiled","shard":"modules/e59ee124d04434ff.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":["delayed-nonstationary"]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEnvironmentRegret","label":"sampledOracleRestartHalfTsallisPredictableEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEnvironmentRegret","description":"Predictable environment regret of the generated restart probabilities against a fixed comparator distribution through the inclusive horizon.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-9362f5a05f77","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10073,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPredictableEnvironmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallispredictableenvironmentregret banditrlproof.tsallis.sampledoraclerestarthalftsallispredictableenvironmentregret predictable environment regret of the generated restart probabilities against a fixed comparator distribution through the inclusive horizon. definition compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret","label":"sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret","description":"Predictable environment regret of the generated restart probabilities against a deterministic comparator arm that may change with the round.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-3482968c9c9d","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10074,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:39"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (comparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret predictable environment regret of the generated restart probabilities against a deterministic comparator arm that may change with the round. definition compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEnvironmentRegret_neverRestart","label":"sampledOracleRestartHalfTsallisPredictableEnvironmentRegret_neverRestart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEnvironmentRegret_neverRestart","description":"theorem sampledOracleRestartHalfTsallisPredictableEnvironmentRegret_neverRestart {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisPredict…","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-e95016ca2c0f","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10075,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:55"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableEnvironmentRegret_neverRestart {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisPredictableEnvironmentRegret arms harms eta oracleNeverRestartSchedule loss q horizon sample = sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms eta loss q horizon sample","missing":[],"search":"sampledoraclerestarthalftsallispredictableenvironmentregret_neverrestart banditrlproof.tsallis.sampledoraclerestarthalftsallispredictableenvironmentregret_neverrestart theorem sampledoraclerestarthalftsallispredictableenvironmentregret_neverrestart {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (loss : exp3.predictablelossvector env action) (q : action -> real) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : sampledoraclerestarthalftsallispredictableenvironmentregret arms harms eta oracleneverrestartschedule loss q horizon sample = sampledscheduledhalftsallispredictableenvironmentregret arms harms eta loss q horizon sample theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_neverRestart","label":"sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_neverRestart","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_neverRestart","description":"theorem sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_neverRestart {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (comparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleR…","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-8354f0c29b7a","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10076,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:70"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_neverRestart {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (comparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta oracleNeverRestartSchedule loss comparator horizon sample = sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta loss comparator horizon sample","missing":[],"search":"sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_neverrestart banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_neverrestart theorem sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_neverrestart {env : type u} {action : type v} [measurablespace env] [measurablespace action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (loss : exp3.predictablelossvector env action) (comparator : nat -> action) (horizon : nat) (sample : env × ((k : nat) -> action × real)) : sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret arms harms eta oracleneverrestartschedule loss comparator horizon sample = sampledscheduledhalftsallispredictablemovingcomparatorenvironmentregret arms harms eta loss comparator horizon sample theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","label":"sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","description":"Restarted moving-comparator regret is fixed point-mass regret plus the cumulative advantage of the moving comparator over the fixed arm.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-e0c4d014e330","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10077,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:88"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (comparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta schedule loss comparator horizon sample = sampledOracleRestartHalfTsallisPredictableEnvironmentRegret arms harms eta schedule loss (pointMass best) horizon sample + (Finset.range (horizon + 1)).sum (fun t => Exp3.predictableLossAt loss t sample best - Exp3.predictableLossAt loss t sample (comparator t))","missing":[],"search":"sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_eq_fixed_add banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_eq_fixed_add restarted moving-comparator regret is fixed point-mass regret plus the cumulative advantage of the moving comparator over the fixed arm. theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs","label":"oracleRestartScheduleEpochs","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartScheduleEpochs","description":"Epoch ids actually visited by the restart schedule through the inclusive horizon. The epoch assignment is exactly `schedule.start`.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-595e99978b40","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10078,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:115"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleRestartScheduleEpochs (schedule : OracleRestartSchedule) (horizon : Nat) : Finset Nat","missing":[],"search":"oraclerestartscheduleepochs banditrlproof.tsallis.oraclerestartscheduleepochs epoch ids actually visited by the restart schedule through the inclusive horizon. the epoch assignment is exactly `schedule.start`. definition compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartSchedule_start_mem_epochs","label":"oracleRestartSchedule_start_mem_epochs","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartSchedule_start_mem_epochs","description":"theorem oracleRestartSchedule_start_mem_epochs (schedule : OracleRestartSchedule) (horizon t : Nat) (ht : t ∈ Finset.range (horizon + 1)) : schedule.start t ∈ oracleRestartScheduleEpochs schedule horizon","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-309731baf4a6","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10079,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:119"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartSchedule_start_mem_epochs (schedule : OracleRestartSchedule) (horizon t : Nat) (ht : t ∈ Finset.range (horizon + 1)) : schedule.start t ∈ oracleRestartScheduleEpochs schedule horizon","missing":[],"search":"oraclerestartschedule_start_mem_epochs banditrlproof.tsallis.oraclerestartschedule_start_mem_epochs theorem oraclerestartschedule_start_mem_epochs (schedule : oraclerestartschedule) (horizon t : nat) (ht : t ∈ finset.range (horizon + 1)) : schedule.start t ∈ oraclerestartscheduleepochs schedule horizon theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret","label":"sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret","description":"Restart-specific predictable regret contributed by one schedule epoch.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-aacd9b447b5c","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10080,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon epoch : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallispredictablescheduleepochregret banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablescheduleepochregret restart-specific predictable regret contributed by one schedule epoch. definition compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_scheduleEpochRegret","label":"sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_scheduleEpochRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_scheduleEpochRegret","description":"Schedule-aligned moving regret is exactly the sum of its epoch-fiber regrets, with no independent `epochOf` compatibility premise.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-fa99e4e50d42","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10081,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:144"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_scheduleEpochRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta schedule loss (fun t => epochComparator (schedule.start t)) horizon sample = (oracleRestartScheduleEpochs schedule horizon).sum (fun epoch => sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret arms harms eta schedule loss epochComparator horizon epoch sample)","missing":[],"search":"sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_eq_sum_scheduleepochregret banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_eq_sum_scheduleepochregret schedule-aligned moving regret is exactly the sum of its epoch-fiber regrets, with no independent `epochof` compatibility premise. theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt","label":"sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt","description":"Schedule-aligned epoch certificates assemble into a global square-root bound for the generated restart regret surface.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-397fdfff568c","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10082,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:175"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon : Nat) (coefficient : Real) (hcoefficient : 0 <= coefficient) (sample : Env × ((k : Nat) -> Action × Real)) (hEpochRegret : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret arms harms eta schedule loss epochComparator horizon epoch sample <= coefficient * Real.sqrt ((oracleRestartEpochRounds schedule.start horizon epoch).card : Real)) : sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret arms harms eta schedul…","missing":[],"search":"sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_schedulesqrt banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_schedulesqrt schedule-aligned epoch certificates assemble into a global square-root bound for the generated restart regret surface. theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt","label":"sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt","description":"Switch-count-facing schedule assembly under an explicit cardinality contract on the epochs actually visited by `schedule.start`.","url":"../modules/banditrlproof-tsallisoraclerestartpredictableregret/index.html#decl-743d3ebd1b76","parent":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","order":10083,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartPredictableRegret"],["Source","BanditRLProof/TsallisOracleRestartPredictableRegret.lean:233"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epochComparator : Nat -> Action) (horizon switches : Nat) (hEpochCard : (oracleRestartScheduleEpochs schedule horizon).card <= switches + 1) (coefficient : Real) (hcoefficient : 0 <= coefficient) (sample : Env × ((k : Nat) -> Action × Real)) (hEpochRegret : ∀ epoch ∈ oracleRestartScheduleEpochs schedule horizon, sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret arms harms eta schedule loss epochComparator horizon epoch sample <= coefficient * Real.sqrt ((oracleRestartEpochRounds schedule.start horizon epoch).card : Real…","missing":[],"search":"sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_scheduleswitchcountsqrt banditrlproof.tsallis.sampledoraclerestarthalftsallispredictablemovingcomparatorenvironmentregret_le_scheduleswitchcountsqrt switch-count-facing schedule assembly under an explicit cardinality contract on the epochs actually visited by `schedule.start`. theorem compiled","shard":"modules/8feacfdbb9c9dfb6.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedPotentialStabilityBound_le_two_mul_eta_mul_sqrt_erase_card","label":"refinedPotentialStabilityBound_le_two_mul_eta_mul_sqrt_erase_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedPotentialStabilityBound_le_two_mul_eta_mul_sqrt_erase_card","description":"The refined all-arm half-Tsallis budget is bounded by the square-root mass of the arms other than any distinguished supported arm.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-911da883c507","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10084,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedPotentialStabilityBound_le_two_mul_eta_mul_sqrt_erase_card {History : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (eta : Real) (heta : 0 <= eta) (probability : History -> Action -> Real) (history : History) (hprobability : FTRL.finiteSimplex arms (probability history)) : refinedPotentialStabilityBound arms eta probability history <= 2 * eta * Real.sqrt ((arms.erase best).card : Real) + 2 * eta ^ 2","missing":[],"search":"refinedpotentialstabilitybound_le_two_mul_eta_mul_sqrt_erase_card banditrlproof.tsallis.refinedpotentialstabilitybound_le_two_mul_eta_mul_sqrt_erase_card the refined all-arm half-tsallis budget is bounded by the square-root mass of the arms other than any distinguished supported arm. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_two_mul_eta_mul_sqrt_erase_card","label":"integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_two_mul_eta_mul_sqrt_erase_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_two_mul_eta_mul_sqrt_erase_card","description":"One actual restart-local successor has the deterministic refined square-root-cardinality budget under the single global generated law.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-0516cd1e2d22","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10085,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:73"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_two_mul_eta_mul_sqrt_erase_card {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) {best : Action} (hbest : best ∈ arms) (heta : 0 < eta (oracleRestartLocalTime schedule (n + 1))) (heta_le : eta (oracleRestartLocalTime schedule (n + 1)) <= 1 / 2) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms eta schedule loss.environment let term := fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms…","missing":[],"search":"integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_le_two_mul_eta_mul_sqrt_erase_card banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisshiftedpotentialstabilityatsuccessor_le_two_mul_eta_mul_sqrt_erase_card one actual restart-local successor has the deterministic refined square-root-cardinality budget under the single global generated law. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.one_div_natSucc_le_one_div_sqrt_natSucc","label":"one_div_natSucc_le_one_div_sqrt_natSucc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.one_div_natSucc_le_one_div_sqrt_natSucc","description":"theorem one_div_natSucc_le_one_div_sqrt_natSucc (t : Nat) : 1 / (((t + 1 : Nat) : Real)) <= 1 / Real.sqrt (((t + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-a53d6271c347","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10086,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:155"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem one_div_natSucc_le_one_div_sqrt_natSucc (t : Nat) : 1 / (((t + 1 : Nat) : Real)) <= 1 / Real.sqrt (((t + 1 : Nat) : Real))","missing":[],"search":"one_div_natsucc_le_one_div_sqrt_natsucc banditrlproof.tsallis.one_div_natsucc_le_one_div_sqrt_natsucc theorem one_div_natsucc_le_one_div_sqrt_natsucc (t : nat) : 1 / (((t + 1 : nat) : real)) <= 1 / real.sqrt (((t + 1 : nat) : real)) theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_two_mul","label":"sampledScheduledHalfTsallisSqrtSchedule_two_mul","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_two_mul","description":"theorem sampledScheduledHalfTsallisSqrtSchedule_two_mul (t : Nat) : 2 * sampledScheduledHalfTsallisSqrtSchedule t = 1 / Real.sqrt (((t + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-a371bd18b36e","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10087,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:172"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSqrtSchedule_two_mul (t : Nat) : 2 * sampledScheduledHalfTsallisSqrtSchedule t = 1 / Real.sqrt (((t + 1 : Nat) : Real))","missing":[],"search":"sampledscheduledhalftsallissqrtschedule_two_mul banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule_two_mul theorem sampledscheduledhalftsallissqrtschedule_two_mul (t : nat) : 2 * sampledscheduledhalftsallissqrtschedule t = 1 / real.sqrt (((t + 1 : nat) : real)) theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_two_mul_sq","label":"sampledScheduledHalfTsallisSqrtSchedule_two_mul_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_two_mul_sq","description":"theorem sampledScheduledHalfTsallisSqrtSchedule_two_mul_sq (t : Nat) : 2 * sampledScheduledHalfTsallisSqrtSchedule t ^ 2 = (1 / 2 : Real) * (1 / (((t + 1 : Nat) : Real)))","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-a7a8d7c2d11f","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10088,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:183"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSqrtSchedule_two_mul_sq (t : Nat) : 2 * sampledScheduledHalfTsallisSqrtSchedule t ^ 2 = (1 / 2 : Real) * (1 / (((t + 1 : Nat) : Real)))","missing":[],"search":"sampledscheduledhalftsallissqrtschedule_two_mul_sq banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule_two_mul_sq theorem sampledscheduledhalftsallissqrtschedule_two_mul_sq (t : nat) : 2 * sampledscheduledhalftsallissqrtschedule t ^ 2 = (1 / 2 : real) * (1 / (((t + 1 : nat) : real))) theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.one_div_sampledScheduledHalfTsallisSqrtSchedule","label":"one_div_sampledScheduledHalfTsallisSqrtSchedule","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.one_div_sampledScheduledHalfTsallisSqrtSchedule","description":"theorem one_div_sampledScheduledHalfTsallisSqrtSchedule (t : Nat) : 1 / sampledScheduledHalfTsallisSqrtSchedule t = 2 * Real.sqrt (((t + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-dd31de4ef18b","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10089,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:190"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem one_div_sampledScheduledHalfTsallisSqrtSchedule (t : Nat) : 1 / sampledScheduledHalfTsallisSqrtSchedule t = 2 * Real.sqrt (((t + 1 : Nat) : Real))","missing":[],"search":"one_div_sampledscheduledhalftsallissqrtschedule banditrlproof.tsallis.one_div_sampledscheduledhalftsallissqrtschedule theorem one_div_sampledscheduledhalftsallissqrtschedule (t : nat) : 1 / sampledscheduledhalftsallissqrtschedule t = 2 * real.sqrt (((t + 1 : nat) : real)) theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_range_sampledScheduledHalfTsallisSqrtSchedule_refinedBudget_le_three_mul_sqrt","label":"sum_range_sampledScheduledHalfTsallisSqrtSchedule_refinedBudget_le_three_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_range_sampledScheduledHalfTsallisSqrtSchedule_refinedBudget_le_three_mul_sqrt","description":"The deterministic refined budgets of the square-root schedule have an inclusive prefix bound of order `sqrt(card * time)`.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-62e2cd2b18d1","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10090,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:203"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_range_sampledScheduledHalfTsallisSqrtSchedule_refinedBudget_le_three_mul_sqrt {Action : Type u} (arms : Finset Action) (harms : arms.Nonempty) (n : Nat) : (Finset.range n).sum (fun t => 2 * sampledScheduledHalfTsallisSqrtSchedule t * Real.sqrt (arms.card : Real) + 2 * sampledScheduledHalfTsallisSqrtSchedule t ^ 2) <= 3 * Real.sqrt (arms.card : Real) * Real.sqrt (n : Real)","missing":[],"search":"sum_range_sampledscheduledhalftsallissqrtschedule_refinedbudget_le_three_mul_sqrt banditrlproof.tsallis.sum_range_sampledscheduledhalftsallissqrtschedule_refinedbudget_le_three_mul_sqrt the deterministic refined budgets of the square-root schedule have an inclusive prefix bound of order `sqrt(card * time)`. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_sqrtSchedule_le_four_mul_sqrt","label":"integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_sqrtSchedule_le_four_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_sqrtSchedule_le_four_mul_sqrt","description":"A deterministic contiguous restart epoch has an expected shifted stability prefix bounded by `4 * sqrt(K) * sqrt(epoch length)`.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-33e2a8a307f6","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10091,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:277"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_sqrtSchedule_le_four_mul_sqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (epoch localHorizon : Nat) {best : Action} (hbest : best ∈ arms) (hstart : ∀ localTime, localTime <= localHorizon -> schedule.start (epoch + localTime) = epoch) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule schedule loss.environment let stabilitySum := fun sample => (Finset.range (localHorizon + 1)).sum (fun localTim…","missing":[],"search":"integral_sum_sampledoraclerestarthalftsallisshiftedpotentialstabilityatlocalprefix_sqrtschedule_le_four_mul_sqrt banditrlproof.tsallis.integral_sum_sampledoraclerestarthalftsallisshiftedpotentialstabilityatlocalprefix_sqrtschedule_le_four_mul_sqrt a deterministic contiguous restart epoch has an expected shifted stability prefix bounded by `4 * sqrt(k) * sqrt(epoch length)`. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.initialHalfTsallisPotentialMass_sqrtSchedule_pointMassPenalty_le_four_mul_sqrt","label":"initialHalfTsallisPotentialMass_sqrtSchedule_pointMassPenalty_le_four_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.initialHalfTsallisPotentialMass_sqrtSchedule_pointMassPenalty_le_four_mul_sqrt","description":"The terminal point-mass penalty of one square-root-scheduled epoch is bounded by `4 * sqrt(K) * sqrt(epoch length)`.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-514e5edfd72a","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10092,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:476"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem initialHalfTsallisPotentialMass_sqrtSchedule_pointMassPenalty_le_four_mul_sqrt {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) {best : Action} (hbest : best ∈ arms) (localHorizon : Nat) : halfTsallisPotentialMass arms (initialHalfTsallisDistribution arms harms (sampledScheduledHalfTsallisSqrtSchedule 0)) / sampledScheduledHalfTsallisSqrtSchedule localHorizon - 1 / sampledScheduledHalfTsallisSqrtSchedule localHorizon <= 4 * Real.sqrt (arms.card : Real) * Real.sqrt ((localHorizon + 1 : Nat) : Real)","missing":[],"search":"initialhalftsallispotentialmass_sqrtschedule_pointmasspenalty_le_four_mul_sqrt banditrlproof.tsallis.initialhalftsallispotentialmass_sqrtschedule_pointmasspenalty_le_four_mul_sqrt the terminal point-mass penalty of one square-root-scheduled epoch is bounded by `4 * sqrt(k) * sqrt(epoch length)`. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt","label":"integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt","description":"An actual contiguous restart epoch run with the square-root schedule has the `C * sqrt(epoch length)` observed estimated-regret certificate required by the restart dynamic-regret assembly, with `C = 8 * sqrt(K)`.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-152e5764e391","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10093,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:558"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (horizon epoch localHorizon : Nat) {best : Action} (hbest : best ∈ arms) (hRounds : oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime)) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule schedule loss.environment let observed := sampled…","missing":[],"search":"integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_sqrtschedule_le_eight_mul_sqrt banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_sqrtschedule_le_eight_mul_sqrt an actual contiguous restart epoch run with the square-root schedule has the `c * sqrt(epoch length)` observed estimated-regret certificate required by the restart dynamic-regret assembly, with `c = 8 * sqrt(k)`. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card","label":"integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card","description":"Cardinality-shaped form of the tuned epoch certificate.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-0d3e4291b719","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10094,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:683"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (horizon epoch localHorizon : Nat) {best : Action} (hbest : best ∈ arms) (hRounds : oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime)) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule schedule loss.environment let observed := sa…","missing":[],"search":"integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_sqrtschedule_le_eight_mul_sqrt_card banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_sqrtschedule_le_eight_mul_sqrt_card cardinality-shaped form of the tuned epoch certificate. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card_of_mem","label":"integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card_of_mem","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card_of_mem","description":"Every actual epoch of the restart schedule exposes the cardinality-shaped certificate directly, with no caller-supplied local-horizon witness.","url":"../modules/banditrlproof-tsallisoraclerestartrefinedstabilitytuning/index.html#decl-7505c9c9d162","parent":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","order":10095,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartRefinedStabilityTuning"],["Source","BanditRLProof/TsallisOracleRestartRefinedStabilityTuning.lean:725"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card_of_mem {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (schedule : OracleRestartSchedule) (loss : Exp3.PredictableLossVector Env Action) (horizon epoch : Nat) {best : Action} (hbest : best ∈ arms) (hepoch : epoch ∈ oracleRestartScheduleEpochs schedule horizon) : let mu := prior ⊗ₘ sampledOracleRestartHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule schedule loss.environment let observed := sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret (Env := Env) arms har…","missing":[],"search":"integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_sqrtschedule_le_eight_mul_sqrt_card_of_mem banditrlproof.tsallis.integral_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_sqrtschedule_le_eight_mul_sqrt_card_of_mem every actual epoch of the restart schedule exposes the cardinality-shaped certificate directly, with no caller-supplied local-horizon witness. theorem compiled","shard":"modules/accbe1b237996ad8.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartShiftedTrajectory","label":"oracleRestartShiftedTrajectory","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartShiftedTrajectory","description":"Shift a generated trajectory so that `start` becomes local time zero.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-4c1158766fed","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10096,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:18"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def oracleRestartShiftedTrajectory {Env : Type u} {Action : Type v} (start : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Env × ((k : Nat) -> Action × Real)","missing":[],"search":"oraclerestartshiftedtrajectory banditrlproof.tsallis.oraclerestartshiftedtrajectory shift a generated trajectory so that `start` becomes local time zero. definition compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartShiftedTrajectory_fst","label":"oracleRestartShiftedTrajectory_fst","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartShiftedTrajectory_fst","description":"theorem oracleRestartShiftedTrajectory_fst {Env : Type u} {Action : Type v} (start : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (oracleRestartShiftedTrajectory start sample).1 = sample.1","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-bd6d4a208a9c","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10097,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartShiftedTrajectory_fst {Env : Type u} {Action : Type v} (start : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : (oracleRestartShiftedTrajectory start sample).1 = sample.1","missing":[],"search":"oraclerestartshiftedtrajectory_fst banditrlproof.tsallis.oraclerestartshiftedtrajectory_fst theorem oraclerestartshiftedtrajectory_fst {env : type u} {action : type v} (start : nat) (sample : env × ((k : nat) -> action × real)) : (oraclerestartshiftedtrajectory start sample).1 = sample.1 theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.oracleRestartShiftedTrajectory_snd_apply","label":"oracleRestartShiftedTrajectory_snd_apply","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.oracleRestartShiftedTrajectory_snd_apply","description":"theorem oracleRestartShiftedTrajectory_snd_apply {Env : Type u} {Action : Type v} (start : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (localTime : Nat) : (oracleRestartShiftedTrajectory start sample).2 localTime = sample.2 (start + localTime)","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-f852d09d34aa","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10098,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:32"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem oracleRestartShiftedTrajectory_snd_apply {Env : Type u} {Action : Type v} (start : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (localTime : Nat) : (oracleRestartShiftedTrajectory start sample).2 localTime = sample.2 (start + localTime)","missing":[],"search":"oraclerestartshiftedtrajectory_snd_apply banditrlproof.tsallis.oraclerestartshiftedtrajectory_snd_apply theorem oraclerestartshiftedtrajectory_snd_apply {env : type u} {action : type v} (start : nat) (sample : env × ((k : nat) -> action × real)) (localtime : nat) : (oraclerestartshiftedtrajectory start sample).2 localtime = sample.2 (start + localtime) theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.monotone_start","label":"monotone_start","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.OracleRestartSchedule.monotone_start","description":"Restart epoch starts are monotone in actual time.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-1b016c1ed322","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10099,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem OracleRestartSchedule.monotone_start (schedule : OracleRestartSchedule) : Monotone schedule.start","missing":[],"search":"monotone_start banditrlproof.tsallis.oraclerestartschedule.monotone_start restart epoch starts are monotone in actual time. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.start_start","label":"start_start","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.OracleRestartSchedule.start_start","description":"Every epoch identifier visited by the schedule is a fixed point of `schedule.start`.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-ba98c60763e7","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10100,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:54"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem OracleRestartSchedule.start_start (schedule : OracleRestartSchedule) (t : Nat) : schedule.start (schedule.start t) = schedule.start t","missing":[],"search":"start_start banditrlproof.tsallis.oraclerestartschedule.start_start every epoch identifier visited by the schedule is a fixed point of `schedule.start`. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.start_eq_of_between","label":"start_eq_of_between","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.OracleRestartSchedule.start_eq_of_between","description":"A restart schedule cannot leave an epoch and later return to it.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-c12fea599f45","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10101,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:65"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem OracleRestartSchedule.start_eq_of_between (schedule : OracleRestartSchedule) {epoch localTime t : Nat} (ht : schedule.start t = epoch) (hepoch : epoch <= localTime) (hlocalTime : localTime <= t) : schedule.start localTime = epoch","missing":[],"search":"start_eq_of_between banditrlproof.tsallis.oraclerestartschedule.start_eq_of_between a restart schedule cannot leave an epoch and later return to it. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_eq_scheduled_shift","label":"sampledOracleRestartHalfTsallisProbabilityAtTime_eq_scheduled_shift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_eq_scheduled_shift","description":"At every actual time, the restarted probability is exactly the scheduled probability at local time `t - schedule.start t` on the shifted trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-368176454b82","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10102,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:79"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisProbabilityAtTime_eq_scheduled_shift {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta schedule t sample = sampledScheduledHalfTsallisProbabilityAtTime arms harms eta (t - schedule.start t) (oracleRestartShiftedTrajectory (schedule.start t) sample)","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityattime_eq_scheduled_shift banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityattime_eq_scheduled_shift at every actual time, the restarted probability is exactly the scheduled probability at local time `t - schedule.start t` on the shifted trajectory. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_scheduled_shift","label":"sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_scheduled_shift","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_scheduled_shift","description":"The stored-reward restart estimator is the scheduled stored-reward estimator at the current epoch's local time.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-e4a3c359ce78","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10103,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_scheduled_shift {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledOracleRestartHalfTsallisObservedEstimatedLossAt arms harms eta schedule t sample = sampledScheduledHalfTsallisObservedEstimatedLossAt arms harms eta (t - schedule.start t) (oracleRestartShiftedTrajectory (schedule.start t) sample)","missing":[],"search":"sampledoraclerestarthalftsallisobservedestimatedlossat_eq_scheduled_shift banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedestimatedlossat_eq_scheduled_shift the stored-reward restart estimator is the scheduled stored-reward estimator at the current epoch's local time. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_add_eq_scheduled_of_start_eq","label":"sampledOracleRestartHalfTsallisProbabilityAtTime_add_eq_scheduled_of_start_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_add_eq_scheduled_of_start_eq","description":"On a fixed epoch fiber, the restarted probability is the scheduled local probability on the trajectory shifted by that epoch.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-437430ff71b2","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10104,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:146"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisProbabilityAtTime_add_eq_scheduled_of_start_eq {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (epoch localTime : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (hstart : schedule.start (epoch + localTime) = epoch) : sampledOracleRestartHalfTsallisProbabilityAtTime arms harms eta schedule (epoch + localTime) sample = sampledScheduledHalfTsallisProbabilityAtTime arms harms eta localTime (oracleRestartShiftedTrajectory epoch sample)","missing":[],"search":"sampledoraclerestarthalftsallisprobabilityattime_add_eq_scheduled_of_start_eq banditrlproof.tsallis.sampledoraclerestarthalftsallisprobabilityattime_add_eq_scheduled_of_start_eq on a fixed epoch fiber, the restarted probability is the scheduled local probability on the trajectory shifted by that epoch. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_add_eq_scheduled_of_start_eq","label":"sampledOracleRestartHalfTsallisObservedEstimatedLossAt_add_eq_scheduled_of_start_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_add_eq_scheduled_of_start_eq","description":"On a fixed epoch fiber, the stored-reward restarted estimator is the scheduled local estimator on the shifted trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-793d29b540d4","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10105,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:163"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedEstimatedLossAt_add_eq_scheduled_of_start_eq {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (epoch localTime : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (hstart : schedule.start (epoch + localTime) = epoch) : sampledOracleRestartHalfTsallisObservedEstimatedLossAt arms harms eta schedule (epoch + localTime) sample = sampledScheduledHalfTsallisObservedEstimatedLossAt arms harms eta localTime (oracleRestartShiftedTrajectory epoch sample)","missing":[],"search":"sampledoraclerestarthalftsallisobservedestimatedlossat_add_eq_scheduled_of_start_eq banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedestimatedlossat_add_eq_scheduled_of_start_eq on a fixed epoch fiber, the stored-reward restarted estimator is the scheduled local estimator on the shifted trajectory. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret","label":"sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret","description":"Stored-reward estimated regret on an explicit inclusive prefix of one restart epoch, indexed by local times `0, ..., localHorizon`.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-0eb03d63659d","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10106,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:180"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (q : Action -> Real) (epoch localHorizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledoraclerestarthalftsallislocalprefixobservedestimatedregret banditrlproof.tsallis.sampledoraclerestarthalftsallislocalprefixobservedestimatedregret stored-reward estimated regret on an explicit inclusive prefix of one restart epoch, indexed by local times `0, ..., localhorizon`. definition compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_eq_scheduled","label":"sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_eq_scheduled","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_eq_scheduled","description":"A contiguous restart-epoch prefix is definitionally the existing scheduled estimated regret on the shifted trajectory.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-c7ff7155a294","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10107,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:199"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_eq_scheduled {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (q : Action -> Real) (epoch localHorizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (hstart : forall localTime, localTime <= localHorizon -> schedule.start (epoch + localTime) = epoch) : sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret arms harms eta schedule q epoch localHorizon sample = sampledScheduledHalfTsallisEstimatedRegret arms harms eta q localHorizon (oracleRestartShiftedTrajectory epoch sample)","missing":[],"search":"sampledoraclerestarthalftsallislocalprefixobservedestimatedregret_eq_scheduled banditrlproof.tsallis.sampledoraclerestarthalftsallislocalprefixobservedestimatedregret_eq_scheduled a contiguous restart-epoch prefix is definitionally the existing scheduled estimated regret on the shifted trajectory. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_pointMass_le_stability_add_penalty","label":"sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_pointMass_le_stability_add_penalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_pointMass_le_stability_add_penalty","description":"Pathwise FTRL certificate for any explicit contiguous prefix of one restart epoch.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-b0f890aa720e","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10108,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:229"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_pointMass_le_stability_add_penalty {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (epoch localHorizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) {best : Action} (hbest : best ∈ arms) (hstart : forall localTime, localTime <= localHorizon -> schedule.start (epoch + localTime) = epoch) (heta : forall localTime, localTime <= localHorizon -> 0 < eta localTime) (hetaMono : forall localTime, localTime < localHorizon -> eta (localTime + 1) <= eta localTime) : sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret arms harms eta schedule (pointMass best) epoch localHorizon sample <= (Finset.range (localHorizon + 1)).sum (fun localTime => sampledScheduledHalfTsallisPotentialStabilityAtTime arms h…","missing":[],"search":"sampledoraclerestarthalftsallislocalprefixobservedestimatedregret_pointmass_le_stability_add_penalty banditrlproof.tsallis.sampledoraclerestarthalftsallislocalprefixobservedestimatedregret_pointmass_le_stability_add_penalty pathwise ftrl certificate for any explicit contiguous prefix of one restart epoch. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_oracleRestartEpochRounds_eq_image_range","label":"exists_oracleRestartEpochRounds_eq_image_range","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_oracleRestartEpochRounds_eq_image_range","description":"Every visited restart epoch fiber is a nonempty contiguous range starting at its epoch identifier.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-503d9de4192c","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10109,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:266"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_oracleRestartEpochRounds_eq_image_range (schedule : OracleRestartSchedule) (horizon epoch : Nat) (hepoch : epoch ∈ oracleRestartScheduleEpochs schedule horizon) : ∃ localHorizon, oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime)","missing":[],"search":"exists_oraclerestartepochrounds_eq_image_range banditrlproof.tsallis.exists_oraclerestartepochrounds_eq_image_range every visited restart epoch fiber is a nonempty contiguous range starting at its epoch identifier. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_localPrefix_of_epochRounds_eq","label":"sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_localPrefix_of_epochRounds_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_localPrefix_of_epochRounds_eq","description":"An explicit range representation of an actual epoch fiber identifies its stored-reward regret with the local-prefix surface.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-2aa64cfb6559","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10110,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:335"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_localPrefix_of_epochRounds_eq {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (q : Action -> Real) (horizon epoch localHorizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) (hRounds : oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime)) : sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret arms harms eta schedule q horizon epoch sample = sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret arms harms eta schedule q epoch localHorizon sample","missing":[],"search":"sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_eq_localprefix_of_epochrounds_eq banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_eq_localprefix_of_epochrounds_eq an explicit range representation of an actual epoch fiber identifies its stored-reward regret with the local-prefix surface. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty_of_epochRounds_eq","label":"sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty_of_epochRounds_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty_of_epochRounds_eq","description":"Pathwise FTRL certificate for an actual epoch once its finite fiber is presented as a contiguous local-time range.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-e279b6511758","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10111,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:358"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty_of_epochRounds_eq {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (horizon epoch localHorizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) {best : Action} (hbest : best ∈ arms) (hRounds : oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime)) (heta : forall localTime, localTime <= localHorizon -> 0 < eta localTime) (hetaMono : forall localTime, localTime < localHorizon -> eta (localTime + 1) <= eta localTime) : sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret arms harms eta schedule (pointMass best) horizon epoch sample <= (Finset.range (localHorizon + 1)).sum (fun lo…","missing":[],"search":"sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_le_stability_add_penalty_of_epochrounds_eq banditrlproof.tsallis.sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_le_stability_add_penalty_of_epochrounds_eq pathwise ftrl certificate for an actual epoch once its finite fiber is presented as a contiguous local-time range. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty","label":"exists_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty","description":"Every visited epoch admits a local horizon whose cardinality is the actual fiber cardinality and whose stored-reward regret satisfies the pathwise FTRL stability-plus-penalty certificate.","url":"../modules/banditrlproof-tsallisoraclerestartscorealignment/index.html#decl-4bb7a48a371c","parent":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","order":10112,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisOracleRestartScoreAlignment"],["Source","BanditRLProof/TsallisOracleRestartScoreAlignment.lean:412"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (schedule : OracleRestartSchedule) (horizon epoch : Nat) (hepoch : epoch ∈ oracleRestartScheduleEpochs schedule horizon) (sample : Env × ((k : Nat) -> Action × Real)) {best : Action} (hbest : best ∈ arms) (heta : forall localTime, 0 < eta localTime) (hetaMono : forall localTime, eta (localTime + 1) <= eta localTime) : ∃ localHorizon, oracleRestartEpochRounds schedule.start horizon epoch = (Finset.range (localHorizon + 1)).image (fun localTime => epoch + localTime) ∧ (oracleRestartEpochRounds schedule.start horizon epoch).card = localHorizon + 1 ∧ sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret arms harms eta schedule (pointMass best…","missing":[],"search":"exists_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_le_stability_add_penalty banditrlproof.tsallis.exists_sampledoraclerestarthalftsallisobservedscheduleepochestimatedregret_pointmass_le_stability_add_penalty every visited epoch admits a local horizon whose cardinality is the actual fiber cardinality and whose stored-reward regret satisfies the pathwise ftrl stability-plus-penalty certificate. theorem compiled","shard":"modules/109822b415d3fc83.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.CounterAction","label":"CounterAction","kind":"abbreviation","status":"compiled","subtitle":"BanditRLProof.Tsallis.CounterAction","description":"private abbrev CounterAction","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-2d6cea402d45","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10113,"meta":[["Kind","abbreviation"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private abbrev CounterAction","missing":[],"search":"counteraction banditrlproof.tsallis.counteraction private abbrev counteraction abbreviation compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterArms","label":"counterArms","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterArms","description":"private noncomputable def counterArms : Finset CounterAction","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-9a14e6a30225","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10114,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterArms : Finset CounterAction","missing":[],"search":"counterarms banditrlproof.tsallis.counterarms private noncomputable def counterarms : finset counteraction definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterEta","label":"counterEta","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterEta","description":"private noncomputable def counterEta : Real","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-9cd7e9e586a1","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10115,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:26"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterEta : Real","missing":[],"search":"countereta banditrlproof.tsallis.countereta private noncomputable def countereta : real definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterProb","label":"counterProb","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterProb","description":"private noncomputable def counterProb : CounterAction -> Real","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-3074fbb4a056","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10116,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:28"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterProb : CounterAction -> Real","missing":[],"search":"counterprob banditrlproof.tsallis.counterprob private noncomputable def counterprob : counteraction -> real definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterScore","label":"counterScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterScore","description":"private noncomputable def counterScore : CounterAction -> Real","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-6ee50b85f75f","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10117,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterScore : CounterAction -> Real","missing":[],"search":"counterscore banditrlproof.tsallis.counterscore private noncomputable def counterscore : counteraction -> real definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterLoss","label":"counterLoss","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterLoss","description":"private noncomputable def counterLoss : CounterAction -> Real","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-7ae3417af063","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10118,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:34"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterLoss : CounterAction -> Real","missing":[],"search":"counterloss banditrlproof.tsallis.counterloss private noncomputable def counterloss : counteraction -> real definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterNext","label":"counterNext","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterNext","description":"private noncomputable def counterNext (chosen action : CounterAction) : Real","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-50d567f3dcb8","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10119,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterNext (chosen action : CounterAction) : Real","missing":[],"search":"counternext banditrlproof.tsallis.counternext private noncomputable def counternext (chosen action : counteraction) : real definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterNextMultiplier","label":"counterNextMultiplier","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterNextMultiplier","description":"private noncomputable def counterNextMultiplier (chosen : CounterAction) : Real","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-7482bbc14702","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10120,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private noncomputable def counterNextMultiplier (chosen : CounterAction) : Real","missing":[],"search":"counternextmultiplier banditrlproof.tsallis.counternextmultiplier private noncomputable def counternextmultiplier (chosen : counteraction) : real definition compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_49","label":"sqrt_49","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_49","description":"private theorem sqrt_49 : Real.sqrt (49 : Real) = 7","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-f7d111038082","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10121,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:51"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_49 : Real.sqrt (49 : Real) = 7","missing":[],"search":"sqrt_49 banditrlproof.tsallis.sqrt_49 private theorem sqrt_49 : real.sqrt (49 : real) = 7 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_576","label":"sqrt_576","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_576","description":"private theorem sqrt_576 : Real.sqrt (576 : Real) = 24","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-afc7894cd018","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10122,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:54"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_576 : Real.sqrt (576 : Real) = 24","missing":[],"search":"sqrt_576 banditrlproof.tsallis.sqrt_576 private theorem sqrt_576 : real.sqrt (576 : real) = 24 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_625","label":"sqrt_625","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_625","description":"private theorem sqrt_625 : Real.sqrt (625 : Real) = 25","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-1da2613bdf59","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10123,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:57"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_625 : Real.sqrt (625 : Real) = 25","missing":[],"search":"sqrt_625 banditrlproof.tsallis.sqrt_625 private theorem sqrt_625 : real.sqrt (625 : real) = 25 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_24778200568643041","label":"sqrt_24778200568643041","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_24778200568643041","description":"private theorem sqrt_24778200568643041 : Real.sqrt (24778200568643041 : Real) = 157410929","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-f3f6749e327e","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10124,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:60"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_24778200568643041 : Real.sqrt (24778200568643041 : Real) = 157410929","missing":[],"search":"sqrt_24778200568643041 banditrlproof.tsallis.sqrt_24778200568643041 private theorem sqrt_24778200568643041 : real.sqrt (24778200568643041 : real) = 157410929 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_1813828968643041","label":"sqrt_1813828968643041","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_1813828968643041","description":"private theorem sqrt_1813828968643041 : Real.sqrt (1813828968643041 : Real) = 42589071","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-cfa0c008c290","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10125,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:65"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_1813828968643041 : Real.sqrt (1813828968643041 : Real) = 42589071","missing":[],"search":"sqrt_1813828968643041 banditrlproof.tsallis.sqrt_1813828968643041 private theorem sqrt_1813828968643041 : real.sqrt (1813828968643041 : real) = 42589071 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_22964371600000000","label":"sqrt_22964371600000000","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_22964371600000000","description":"private theorem sqrt_22964371600000000 : Real.sqrt (22964371600000000 : Real) = 151540000","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-b3a653de4313","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10126,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:70"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_22964371600000000 : Real.sqrt (22964371600000000 : Real) = 151540000","missing":[],"search":"sqrt_22964371600000000 banditrlproof.tsallis.sqrt_22964371600000000 private theorem sqrt_22964371600000000 : real.sqrt (22964371600000000 : real) = 151540000 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_1524122302720081","label":"sqrt_1524122302720081","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_1524122302720081","description":"private theorem sqrt_1524122302720081 : Real.sqrt (1524122302720081 : Real) = 39040009","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-91488c46980c","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10127,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:75"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_1524122302720081 : Real.sqrt (1524122302720081 : Real) = 39040009","missing":[],"search":"sqrt_1524122302720081 banditrlproof.tsallis.sqrt_1524122302720081 private theorem sqrt_1524122302720081 : real.sqrt (1524122302720081 : real) = 39040009 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_120121402720081","label":"sqrt_120121402720081","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_120121402720081","description":"private theorem sqrt_120121402720081 : Real.sqrt (120121402720081 : Real) = 10959991","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-47b9ff33ea0e","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10128,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:80"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_120121402720081 : Real.sqrt (120121402720081 : Real) = 10959991","missing":[],"search":"sqrt_120121402720081 banditrlproof.tsallis.sqrt_120121402720081 private theorem sqrt_120121402720081 : real.sqrt (120121402720081 : real) = 10959991 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_1404000900000000","label":"sqrt_1404000900000000","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_1404000900000000","description":"private theorem sqrt_1404000900000000 : Real.sqrt (1404000900000000 : Real) = 37470000","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-f67e7de8eb5f","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10129,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:85"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_1404000900000000 : Real.sqrt (1404000900000000 : Real) = 37470000","missing":[],"search":"sqrt_1404000900000000 banditrlproof.tsallis.sqrt_1404000900000000 private theorem sqrt_1404000900000000 : real.sqrt (1404000900000000 : real) = 37470000 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterEta_pos","label":"counterEta_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterEta_pos","description":"private theorem counterEta_pos : 0 < counterEta","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-299cf9ea28a9","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10130,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:90"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterEta_pos : 0 < counterEta","missing":[],"search":"countereta_pos banditrlproof.tsallis.countereta_pos private theorem countereta_pos : 0 < countereta theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterEta_le_one","label":"counterEta_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterEta_le_one","description":"private theorem counterEta_le_one : counterEta <= 1","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-eef6d39838ec","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10131,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:93"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterEta_le_one : counterEta <= 1","missing":[],"search":"countereta_le_one banditrlproof.tsallis.countereta_le_one private theorem countereta_le_one : countereta <= 1 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterProb_simplex","label":"counterProb_simplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterProb_simplex","description":"private theorem counterProb_simplex : FTRL.finiteSimplex counterArms counterProb","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-8749bb9262ce","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10132,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:96"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterProb_simplex : FTRL.finiteSimplex counterArms counterProb","missing":[],"search":"counterprob_simplex banditrlproof.tsallis.counterprob_simplex private theorem counterprob_simplex : ftrl.finitesimplex counterarms counterprob theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterProb_pos","label":"counterProb_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterProb_pos","description":"private theorem counterProb_pos (action : CounterAction) (haction : action ∈ counterArms) : 0 < counterProb action","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-f222c6ecc083","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10133,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:103"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterProb_pos (action : CounterAction) (haction : action ∈ counterArms) : 0 < counterProb action","missing":[],"search":"counterprob_pos banditrlproof.tsallis.counterprob_pos private theorem counterprob_pos (action : counteraction) (haction : action ∈ counterarms) : 0 < counterprob action theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterLoss_mem_Icc","label":"counterLoss_mem_Icc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterLoss_mem_Icc","description":"private theorem counterLoss_mem_Icc (action : CounterAction) (haction : action ∈ counterArms) : 0 <= counterLoss action ∧ counterLoss action <= 1","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-2aeafec8d5bc","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10134,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:108"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterLoss_mem_Icc (action : CounterAction) (haction : action ∈ counterArms) : 0 <= counterLoss action ∧ counterLoss action <= 1","missing":[],"search":"counterloss_mem_icc banditrlproof.tsallis.counterloss_mem_icc private theorem counterloss_mem_icc (action : counteraction) (haction : action ∈ counterarms) : 0 <= counterloss action ∧ counterloss action <= 1 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterNext_simplex","label":"counterNext_simplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterNext_simplex","description":"private theorem counterNext_simplex (chosen : CounterAction) (hchosen : chosen ∈ counterArms) : FTRL.finiteSimplex counterArms (counterNext chosen)","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-0044979c5ad1","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10135,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:113"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterNext_simplex (chosen : CounterAction) (hchosen : chosen ∈ counterArms) : FTRL.finiteSimplex counterArms (counterNext chosen)","missing":[],"search":"counternext_simplex banditrlproof.tsallis.counternext_simplex private theorem counternext_simplex (chosen : counteraction) (hchosen : chosen ∈ counterarms) : ftrl.finitesimplex counterarms (counternext chosen) theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterNext_pos","label":"counterNext_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterNext_pos","description":"private theorem counterNext_pos (chosen : CounterAction) (hchosen : chosen ∈ counterArms) (action : CounterAction) (haction : action ∈ counterArms) : 0 < counterNext chosen action","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-be4f035140db","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10136,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:122"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterNext_pos (chosen : CounterAction) (hchosen : chosen ∈ counterArms) (action : CounterAction) (haction : action ∈ counterArms) : 0 < counterNext chosen action","missing":[],"search":"counternext_pos banditrlproof.tsallis.counternext_pos private theorem counternext_pos (chosen : counteraction) (hchosen : chosen ∈ counterarms) (action : counteraction) (haction : action ∈ counterarms) : 0 < counternext chosen action theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.rpow_neg_half_eq_inv_sqrt","label":"rpow_neg_half_eq_inv_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.rpow_neg_half_eq_inv_sqrt","description":"private theorem rpow_neg_half_eq_inv_sqrt {x : Real} (hx : 0 < x) : x ^ (-(1 / 2 : Real)) = 1 / Real.sqrt x","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-187eaf24c509","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10137,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:128"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem rpow_neg_half_eq_inv_sqrt {x : Real} (hx : 0 < x) : x ^ (-(1 / 2 : Real)) = 1 / Real.sqrt x","missing":[],"search":"rpow_neg_half_eq_inv_sqrt banditrlproof.tsallis.rpow_neg_half_eq_inv_sqrt private theorem rpow_neg_half_eq_inv_sqrt {x : real} (hx : 0 < x) : x ^ (-(1 / 2 : real)) = 1 / real.sqrt x theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterProb_stationary","label":"counterProb_stationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterProb_stationary","description":"private theorem counterProb_stationary : HalfTsallisInteriorStationary counterArms counterEta counterScore counterProb 0","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-a611b00b2e35","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10138,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:133"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterProb_stationary : HalfTsallisInteriorStationary counterArms counterEta counterScore counterProb 0","missing":[],"search":"counterprob_stationary banditrlproof.tsallis.counterprob_stationary private theorem counterprob_stationary : halftsallisinteriorstationary counterarms countereta counterscore counterprob 0 theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterNext_stationary","label":"counterNext_stationary","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterNext_stationary","description":"private theorem counterNext_stationary (chosen : CounterAction) (hchosen : chosen ∈ counterArms) : HalfTsallisInteriorStationary counterArms counterEta (fun action => counterScore action + Exp3.importanceWeightedLoss counterProb counterLoss chosen action) (counterNext chosen) (counterNextMultiplier chosen)","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-1340dd694754","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10139,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:142"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterNext_stationary (chosen : CounterAction) (hchosen : chosen ∈ counterArms) : HalfTsallisInteriorStationary counterArms counterEta (fun action => counterScore action + Exp3.importanceWeightedLoss counterProb counterLoss chosen action) (counterNext chosen) (counterNextMultiplier chosen)","missing":[],"search":"counternext_stationary banditrlproof.tsallis.counternext_stationary private theorem counternext_stationary (chosen : counteraction) (hchosen : chosen ∈ counterarms) : halftsallisinteriorstationary counterarms countereta (fun action => counterscore action + exp3.importanceweightedloss counterprob counterloss chosen action) (counternext chosen) (counternextmultiplier chosen) theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterProb_minimizer","label":"counterProb_minimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterProb_minimizer","description":"private theorem counterProb_minimizer : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex counterArms) counterArms counterEta (negEntropyRegularizer counterArms (1 / 2 : Real)) counterScore counterProb","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-879bad42375c","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10140,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:159"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterProb_minimizer : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex counterArms) counterArms counterEta (negEntropyRegularizer counterArms (1 / 2 : Real)) counterScore counterProb","missing":[],"search":"counterprob_minimizer banditrlproof.tsallis.counterprob_minimizer private theorem counterprob_minimizer : ftrl.isregularizedminimizer (ftrl.finitesimplex counterarms) counterarms countereta (negentropyregularizer counterarms (1 / 2 : real)) counterscore counterprob theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.counterNext_minimizer","label":"counterNext_minimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.counterNext_minimizer","description":"private theorem counterNext_minimizer (chosen : CounterAction) (hchosen : chosen ∈ counterArms) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex counterArms) counterArms counterEta (negEntropyRegularizer counterArms (1 / 2 : Real)) (fun action => counterScore action + Exp3.importanceWeightedLoss counterProb counterLoss chosen action) (counterNext chosen)","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-cdbbb82943ff","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10141,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:167"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem counterNext_minimizer (chosen : CounterAction) (hchosen : chosen ∈ counterArms) : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex counterArms) counterArms counterEta (negEntropyRegularizer counterArms (1 / 2 : Real)) (fun action => counterScore action + Exp3.importanceWeightedLoss counterProb counterLoss chosen action) (counterNext chosen)","missing":[],"search":"counternext_minimizer banditrlproof.tsallis.counternext_minimizer private theorem counternext_minimizer (chosen : counteraction) (hchosen : chosen ∈ counterarms) : ftrl.isregularizedminimizer (ftrl.finitesimplex counterarms) counterarms countereta (negentropyregularizer counterarms (1 / 2 : real)) (fun action => counterscore action + exp3.importanceweightedloss counterprob counterloss chosen action) (counternext chosen) theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_minimizer_counterexample_to_refinedAveragedStability","label":"exists_minimizer_counterexample_to_refinedAveragedStability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_minimizer_counterexample_to_refinedAveragedStability","description":"Even with strict positive simplex minimizers and `[0,1]` losses, the current sampled-action average of `<p - p_next, hatLoss>` can exceed the locally scaled paper coefficient `eta * sum sqrt(p) * (1-p) + 2 * eta^2`.","url":"../modules/banditrlproof-tsallisrefinedaveragedstabilityobstruction/index.html#decl-eae32f91aceb","parent":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","order":10142,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedAveragedStabilityObstruction"],["Source","BanditRLProof/TsallisRefinedAveragedStabilityObstruction.lean:188"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_minimizer_counterexample_to_refinedAveragedStability : ∃ (eta : Real) (score prob loss : Fin 2 -> Real) (next : Fin 2 -> Fin 2 -> Real), 0 < eta ∧ eta <= 1 ∧ FTRL.finiteSimplex Finset.univ prob ∧ (∀ action ∈ (Finset.univ : Finset (Fin 2)), 0 < prob action) ∧ (∀ action ∈ (Finset.univ : Finset (Fin 2)), 0 <= loss action ∧ loss action <= 1) ∧ FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex Finset.univ) Finset.univ eta (negEntropyRegularizer Finset.univ (1 / 2 : Real)) score prob ∧ (∀ chosen ∈ (Finset.univ : Finset (Fin 2)), ∀ action ∈ (Finset.univ : Finset (Fin 2)), 0 < next chosen action) ∧ (∀ chosen ∈ (Finset.univ : Finset (Fin 2)), FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex Finset.univ) Finset.univ eta (negEntropyRegularizer Finset.univ (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss prob loss chosen action) (next chosen)) ∧ eta * (Finset.un…","missing":[],"search":"exists_minimizer_counterexample_to_refinedaveragedstability banditrlproof.tsallis.exists_minimizer_counterexample_to_refinedaveragedstability even with strict positive simplex minimizers and `[0,1]` losses, the current sampled-action average of `<p - p_next, hatloss>` can exceed the locally scaled paper coefficient `eta * sum sqrt(p) * (1-p) + 2 * eta^2`. theorem compiled","shard":"modules/0d8cb34db93ae3ba.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.shiftedHalfPowerImportanceWeightedMoment","label":"shiftedHalfPowerImportanceWeightedMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.shiftedHalfPowerImportanceWeightedMoment","description":"Inverse-half-Tsallis-Hessian quadratic moment after subtracting a baseline.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-d230424ad39b","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10143,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def shiftedHalfPowerImportanceWeightedMoment {Action : Type u} (arms : Finset Action) (prob loss : Action -> Real) (chosen : Action) : Real","missing":[],"search":"shiftedhalfpowerimportanceweightedmoment banditrlproof.tsallis.shiftedhalfpowerimportanceweightedmoment inverse-half-tsallis-hessian quadratic moment after subtracting a baseline. definition compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.shiftedPositiveCubicImportanceWeightedMoment","label":"shiftedPositiveCubicImportanceWeightedMoment","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.shiftedPositiveCubicImportanceWeightedMoment","description":"Positive cubic remainder used by the refined Taylor bound.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-184da0295f54","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10144,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:36"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def shiftedPositiveCubicImportanceWeightedMoment {Action : Type u} (arms : Finset Action) (prob loss : Action -> Real) (chosen : Action) : Real","missing":[],"search":"shiftedpositivecubicimportanceweightedmoment banditrlproof.tsallis.shiftedpositivecubicimportanceweightedmoment positive cubic remainder used by the refined taylor bound. definition compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_mul_sum_erase_eq_sum_mul_one_sub","label":"sum_mul_sum_erase_eq_sum_mul_one_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_mul_sum_erase_eq_sum_mul_one_sub","description":"Reindex a weighted sum over complements using simplex normalization.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-c4f1f256c8e7","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10145,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:46"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_mul_sum_erase_eq_sum_mul_one_sub {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob weight : Action -> Real) (hsum : arms.sum prob = 1) : arms.sum (fun chosen => prob chosen * (arms.erase chosen).sum weight) = arms.sum (fun action => weight action * (1 - prob action))","missing":[],"search":"sum_mul_sum_erase_eq_sum_mul_one_sub banditrlproof.tsallis.sum_mul_sum_erase_eq_sum_mul_one_sub reindex a weighted sum over complements using simplex normalization. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.prob_mul_shiftedHalfPowerImportanceWeightedMoment_eq","label":"prob_mul_shiftedHalfPowerImportanceWeightedMoment_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.prob_mul_shiftedHalfPowerImportanceWeightedMoment_eq","description":"Exact sampled-action expansion of the shifted quadratic moment.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-debb2deb48ea","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10146,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:89"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem prob_mul_shiftedHalfPowerImportanceWeightedMoment_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) {chosen : Action} (hchosen : chosen ∈ arms) (hprob : 0 < prob chosen) : prob chosen * shiftedHalfPowerImportanceWeightedMoment arms prob loss chosen = (loss chosen) ^ 2 * (Real.sqrt (prob chosen) * (1 - prob chosen) ^ 2 + prob chosen * (arms.erase chosen).sum (fun action => Real.sqrt (prob action) * prob action))","missing":[],"search":"prob_mul_shiftedhalfpowerimportanceweightedmoment_eq banditrlproof.tsallis.prob_mul_shiftedhalfpowerimportanceweightedmoment_eq exact sampled-action expansion of the shifted quadratic moment. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_shiftedHalfPowerImportanceWeightedMoment_le","label":"sum_prob_mul_shiftedHalfPowerImportanceWeightedMoment_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_shiftedHalfPowerImportanceWeightedMoment_le","description":"The sampled shifted quadratic IW moment is bounded by the refined all-arm half-power mass.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-0f4316a0beb4","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10147,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:136"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_shiftedHalfPowerImportanceWeightedMoment_le {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hprobability : FTRL.finiteSimplex arms prob) (hprob : forall action, action ∈ arms -> 0 < prob action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => prob chosen * shiftedHalfPowerImportanceWeightedMoment arms prob loss chosen) <= arms.sum (fun action => Real.sqrt (prob action) * (1 - prob action))","missing":[],"search":"sum_prob_mul_shiftedhalfpowerimportanceweightedmoment_le banditrlproof.tsallis.sum_prob_mul_shiftedhalfpowerimportanceweightedmoment_le the sampled shifted quadratic iw moment is bounded by the refined all-arm half-power mass. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.prob_mul_shiftedPositiveCubicImportanceWeightedMoment_eq","label":"prob_mul_shiftedPositiveCubicImportanceWeightedMoment_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.prob_mul_shiftedPositiveCubicImportanceWeightedMoment_eq","description":"Exact sampled-action expansion of the positive cubic remainder.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-a140ba017811","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10148,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:195"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem prob_mul_shiftedPositiveCubicImportanceWeightedMoment_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) {chosen : Action} (hchosen : chosen ∈ arms) (hprob : 0 < prob chosen) (hprobLeOne : prob chosen <= 1) (hlossNonneg : 0 <= loss chosen) : prob chosen * shiftedPositiveCubicImportanceWeightedMoment arms prob loss chosen = prob chosen * (loss chosen) ^ 3 * (arms.erase chosen).sum (fun action => (prob action) ^ 2)","missing":[],"search":"prob_mul_shiftedpositivecubicimportanceweightedmoment_eq banditrlproof.tsallis.prob_mul_shiftedpositivecubicimportanceweightedmoment_eq exact sampled-action expansion of the positive cubic remainder. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_shiftedPositiveCubicImportanceWeightedMoment_le_one","label":"sum_prob_mul_shiftedPositiveCubicImportanceWeightedMoment_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_shiftedPositiveCubicImportanceWeightedMoment_le_one","description":"The sampled positive cubic IW remainder is at most one.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-378d7121f5c3","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10149,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:226"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_shiftedPositiveCubicImportanceWeightedMoment_le_one {Action : Type u} [DecidableEq Action] (arms : Finset Action) (prob loss : Action -> Real) (hprobability : FTRL.finiteSimplex arms prob) (hprob : forall action, action ∈ arms -> 0 < prob action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => prob chosen * shiftedPositiveCubicImportanceWeightedMoment arms prob loss chosen) <= 1","missing":[],"search":"sum_prob_mul_shiftedpositivecubicimportanceweightedmoment_le_one banditrlproof.tsallis.sum_prob_mul_shiftedpositivecubicimportanceweightedmoment_le_one the sampled positive cubic iw remainder is at most one. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_stability_le_refinedHalfPower_add_square","label":"sum_prob_mul_stability_le_refinedHalfPower_add_square","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_stability_le_refinedHalfPower_add_square","description":"Paper-shaped finite-sum consumer for a shifted Taylor/Hessian stability bound. The remaining hypothesis is deliberately pointwise: a producer must compare its instantaneous stability quantity with the shifted quadratic and positive cubic moments. This theorem performs all sampled-action averaging.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-9739d46f3016","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10150,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:300"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_stability_le_refinedHalfPower_add_square {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (prob loss stability : Action -> Real) (heta : 0 <= eta) (hprobability : FTRL.finiteSimplex arms prob) (hprob : forall action, action ∈ arms -> 0 < prob action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) (hpointwise : forall chosen, chosen ∈ arms -> stability chosen <= eta / 2 * shiftedHalfPowerImportanceWeightedMoment arms prob loss chosen + eta ^ 2 / 2 * shiftedPositiveCubicImportanceWeightedMoment arms prob loss chosen) : arms.sum (fun chosen => prob chosen * stability chosen) <= eta / 2 * arms.sum (fun action => Real.sqrt (prob action) * (1 - prob action)) + eta ^ 2 / 2","missing":[],"search":"sum_prob_mul_stability_le_refinedhalfpower_add_square banditrlproof.tsallis.sum_prob_mul_stability_le_refinedhalfpower_add_square paper-shaped finite-sum consumer for a shifted taylor/hessian stability bound. the remaining hypothesis is deliberately pointwise: a producer must compare its instantaneous stability quantity with the shifted quadratic and positive cubic moments. this theorem performs all sampled-action averaging. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_refined_of_shiftedTaylor","label":"sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_refined_of_shiftedTaylor","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_refined_of_shiftedTaylor","description":"Exact current-FTRL-expression wrapper around the shifted-moment consumer. The sole unresolved premise is the deterministic shifted Taylor/Hessian comparison for each sampled-action update.","url":"../modules/banditrlproof-tsallisrefinedimportanceweightedmoment/index.html#decl-56d44022f17b","parent":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","order":10151,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedImportanceWeightedMoment"],["Source","BanditRLProof/TsallisRefinedImportanceWeightedMoment.lean:399"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_refined_of_shiftedTaylor {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (prob loss : Action -> Real) (next : Action -> Action -> Real) (heta : 0 <= eta) (hprobability : FTRL.finiteSimplex arms prob) (hprob : forall action, action ∈ arms -> 0 < prob action) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) (hshiftedTaylor : forall chosen, chosen ∈ arms -> FTRL.linearLoss arms prob (Exp3.importanceWeightedLoss prob loss chosen) - FTRL.linearLoss arms (next chosen) (Exp3.importanceWeightedLoss prob loss chosen) <= eta / 2 * shiftedHalfPowerImportanceWeightedMoment arms prob loss chosen + eta ^ 2 / 2 * shiftedPositiveCubicImportanceWeightedMoment arms prob loss chosen) : arms.sum (fun chosen => prob chosen * (FTRL.linearLoss arms prob (Exp3.importanceWeightedLoss prob l…","missing":[],"search":"sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_refined_of_shiftedtaylor banditrlproof.tsallis.sum_prob_mul_linearloss_sub_next_importanceweightedloss_le_refined_of_shiftedtaylor exact current-ftrl-expression wrapper around the shifted-moment consumer. the sole unresolved premise is the deterministic shifted taylor/hessian comparison for each sampled-action update. theorem compiled","shard":"modules/277aeb8ba81dc22e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.probability_le_sqrt","label":"probability_le_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.probability_le_sqrt","description":"theorem probability_le_sqrt (probability : Real) (hprobability : 0 <= probability) (hprobability_le_one : probability <= 1) : probability <= Real.sqrt probability","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html#decl-5a8eb0f37e9a","parent":"module:BanditRLProof.TsallisRefinedSuboptimalStability","order":10152,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedSuboptimalStability"],["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem probability_le_sqrt (probability : Real) (hprobability : 0 <= probability) (hprobability_le_one : probability <= 1) : probability <= Real.sqrt probability","missing":[],"search":"probability_le_sqrt banditrlproof.tsallis.probability_le_sqrt theorem probability_le_sqrt (probability : real) (hprobability : 0 <= probability) (hprobability_le_one : probability <= 1) : probability <= real.sqrt probability theorem compiled","shard":"modules/08461fc43d8decb4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_mul_one_sub_le_sqrt","label":"sqrt_mul_one_sub_le_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_mul_one_sub_le_sqrt","description":"theorem sqrt_mul_one_sub_le_sqrt (probability : Real) (hprobability : 0 <= probability) : Real.sqrt probability * (1 - probability) <= Real.sqrt probability","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html#decl-b74535939a1b","parent":"module:BanditRLProof.TsallisRefinedSuboptimalStability","order":10153,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedSuboptimalStability"],["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sqrt_mul_one_sub_le_sqrt (probability : Real) (hprobability : 0 <= probability) : Real.sqrt probability * (1 - probability) <= Real.sqrt probability","missing":[],"search":"sqrt_mul_one_sub_le_sqrt banditrlproof.tsallis.sqrt_mul_one_sub_le_sqrt theorem sqrt_mul_one_sub_le_sqrt (probability : real) (hprobability : 0 <= probability) : real.sqrt probability * (1 - probability) <= real.sqrt probability theorem compiled","shard":"modules/08461fc43d8decb4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_mul_one_sub_le_one_sub","label":"sqrt_mul_one_sub_le_one_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_mul_one_sub_le_one_sub","description":"theorem sqrt_mul_one_sub_le_one_sub (probability : Real) (hprobability_le_one : probability <= 1) : Real.sqrt probability * (1 - probability) <= 1 - probability","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html#decl-7f00d87d42cb","parent":"module:BanditRLProof.TsallisRefinedSuboptimalStability","order":10154,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedSuboptimalStability"],["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sqrt_mul_one_sub_le_one_sub (probability : Real) (hprobability_le_one : probability <= 1) : Real.sqrt probability * (1 - probability) <= 1 - probability","missing":[],"search":"sqrt_mul_one_sub_le_one_sub banditrlproof.tsallis.sqrt_mul_one_sub_le_one_sub theorem sqrt_mul_one_sub_le_one_sub (probability : real) (hprobability_le_one : probability <= 1) : real.sqrt probability * (1 - probability) <= 1 - probability theorem compiled","shard":"modules/08461fc43d8decb4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.one_sub_eq_sum_erase","label":"one_sub_eq_sum_erase","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.one_sub_eq_sum_erase","description":"theorem one_sub_eq_sum_erase {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability : Action -> Real) (hprobability : FTRL.finiteSimplex arms probability) : 1 - probability best = (arms.erase best).sum probability","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html#decl-6c3d912dff82","parent":"module:BanditRLProof.TsallisRefinedSuboptimalStability","order":10155,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedSuboptimalStability"],["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem one_sub_eq_sum_erase {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability : Action -> Real) (hprobability : FTRL.finiteSimplex arms probability) : 1 - probability best = (arms.erase best).sum probability","missing":[],"search":"one_sub_eq_sum_erase banditrlproof.tsallis.one_sub_eq_sum_erase theorem one_sub_eq_sum_erase {action : type u} [decidableeq action] (arms : finset action) {best : action} (hbest : best ∈ arms) (probability : action -> real) (hprobability : ftrl.finitesimplex arms probability) : 1 - probability best = (arms.erase best).sum probability theorem compiled","shard":"modules/08461fc43d8decb4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt","label":"sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt","description":"Eliminate one distinguished arm from the paper's refined half-power stability budget.","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html#decl-c50054ba034e","parent":"module:BanditRLProof.TsallisRefinedSuboptimalStability","order":10156,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedSuboptimalStability"],["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean:60"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability : Action -> Real) (hprobability : FTRL.finiteSimplex arms probability) : arms.sum (fun action => Real.sqrt (probability action) * (1 - probability action)) <= 2 * (arms.erase best).sum (fun action => Real.sqrt (probability action))","missing":[],"search":"sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt banditrlproof.tsallis.sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt eliminate one distinguished arm from the paper's refined half-power stability budget. theorem compiled","shard":"modules/08461fc43d8decb4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.regret_le_of_refinedHalfPowerSelfBounding","label":"regret_le_of_refinedHalfPowerSelfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.regret_le_of_refinedHalfPowerSelfBounding","description":"A finite time-by-suboptimal-arm consumer for the paper-shaped refined stability budget. The remaining algorithmic obligation is exactly `hupper`.","url":"../modules/banditrlproof-tsallisrefinedsuboptimalstability/index.html#decl-61ecf035497c","parent":"module:BanditRLProof.TsallisRefinedSuboptimalStability","order":10157,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRefinedSuboptimalStability"],["Source","BanditRLProof/TsallisRefinedSuboptimalStability.lean:100"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regret_le_of_refinedHalfPowerSelfBounding {Time : Type u} {Action : Type v} [DecidableEq Time] [DecidableEq Action] (times : Finset Time) (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability : Time -> Action -> Real) (coefficient : Time -> Real) (gap : Action -> Real) (regret base corruption : Real) (hprobability : ∀ time ∈ times, FTRL.finiteSimplex arms (probability time)) (hcoefficient : ∀ time ∈ times, 0 <= coefficient time) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (hselfBounding : times.sum (fun time => (arms.erase best).sum (fun action => gap action * probability time action)) - corruption <= regret) (hupper : regret <= base + times.sum (fun time => coefficient time * arms.sum (fun action => Real.sqrt (probability time action) * (1 - probability time action)))) : regret <= 2 * base + (times.product (arms.erase best)).sum (fun index => (2 * co…","missing":[],"search":"regret_le_of_refinedhalfpowerselfbounding banditrlproof.tsallis.regret_le_of_refinedhalfpowerselfbounding a finite time-by-suboptimal-arm consumer for the paper-shaped refined stability budget. the remaining algorithmic obligation is exactly `hupper`. theorem compiled","shard":"modules/08461fc43d8decb4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.powerSum","label":"powerSum","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.powerSum","description":"Finite Tsallis power sum `sum_a p_a^alpha`.","url":"../modules/banditrlproof-tsallisregularizer/index.html#decl-2f226c198b60","parent":"module:BanditRLProof.TsallisRegularizer","order":10158,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRegularizer"],["Source","BanditRLProof/TsallisRegularizer.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def powerSum {Action : Type u} (arms : Finset Action) (alpha : Real) (p : Action -> Real) : Real","missing":[],"search":"powersum banditrlproof.tsallis.powersum finite tsallis power sum `sum_a p_a^alpha`. definition compiled","shard":"modules/bef5147b4a2f36c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.entropy","label":"entropy","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.entropy","description":"Tsallis entropy on a finite action set, with denominator left explicit.","url":"../modules/banditrlproof-tsallisregularizer/index.html#decl-5dc99abd5a82","parent":"module:BanditRLProof.TsallisRegularizer","order":10159,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRegularizer"],["Source","BanditRLProof/TsallisRegularizer.lean:29"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def entropy {Action : Type u} (arms : Finset Action) (alpha : Real) (p : Action -> Real) : Real","missing":[],"search":"entropy banditrlproof.tsallis.entropy tsallis entropy on a finite action set, with denominator left explicit. definition compiled","shard":"modules/bef5147b4a2f36c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.negEntropyRegularizer","label":"negEntropyRegularizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.negEntropyRegularizer","description":"Negative Tsallis entropy as the FTRL regularizer convention.","url":"../modules/banditrlproof-tsallisregularizer/index.html#decl-d852e7cbd2ac","parent":"module:BanditRLProof.TsallisRegularizer","order":10160,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisRegularizer"],["Source","BanditRLProof/TsallisRegularizer.lean:34"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def negEntropyRegularizer {Action : Type u} (arms : Finset Action) (alpha : Real) (p : Action -> Real) : Real","missing":[],"search":"negentropyregularizer banditrlproof.tsallis.negentropyregularizer negative tsallis entropy as the ftrl regularizer convention. definition compiled","shard":"modules/bef5147b4a2f36c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.one_sub_exponent_ne_zero","label":"one_sub_exponent_ne_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.one_sub_exponent_ne_zero","description":"The Tsallis denominator is nonzero when `alpha != 1`.","url":"../modules/banditrlproof-tsallisregularizer/index.html#decl-543ea2a4b7ea","parent":"module:BanditRLProof.TsallisRegularizer","order":10161,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRegularizer"],["Source","BanditRLProof/TsallisRegularizer.lean:39"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem one_sub_exponent_ne_zero {alpha : Real} (halpha : alpha ≠ 1) : 1 - alpha ≠ 0","missing":[],"search":"one_sub_exponent_ne_zero banditrlproof.tsallis.one_sub_exponent_ne_zero the tsallis denominator is nonzero when `alpha != 1`. theorem compiled","shard":"modules/bef5147b4a2f36c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.powerSum_nonneg_of_finiteSimplex","label":"powerSum_nonneg_of_finiteSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.powerSum_nonneg_of_finiteSimplex","description":"The Tsallis power sum is nonnegative on the finite simplex.","url":"../modules/banditrlproof-tsallisregularizer/index.html#decl-82e84bb7b6d0","parent":"module:BanditRLProof.TsallisRegularizer","order":10162,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRegularizer"],["Source","BanditRLProof/TsallisRegularizer.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem powerSum_nonneg_of_finiteSimplex {Action : Type u} (arms : Finset Action) (alpha : Real) (p : Action -> Real) (hp : FTRL.finiteSimplex arms p) : 0 <= powerSum arms alpha p","missing":[],"search":"powersum_nonneg_of_finitesimplex banditrlproof.tsallis.powersum_nonneg_of_finitesimplex the tsallis power sum is nonnegative on the finite simplex. theorem compiled","shard":"modules/bef5147b4a2f36c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.negEntropyRegularizer_wellDefined_on_finiteSimplex","label":"negEntropyRegularizer_wellDefined_on_finiteSimplex","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.negEntropyRegularizer_wellDefined_on_finiteSimplex","description":"Well-definedness package for the finite-simplex Tsallis regularizer. The two facts exposed here are the local obligations needed before later Tsallis/FTRL leaves can use `Real.rpow` algebra and division by `1 - alpha`.","url":"../modules/banditrlproof-tsallisregularizer/index.html#decl-f6fd85a31bb2","parent":"module:BanditRLProof.TsallisRegularizer","order":10163,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisRegularizer"],["Source","BanditRLProof/TsallisRegularizer.lean:60"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem negEntropyRegularizer_wellDefined_on_finiteSimplex {Action : Type u} (arms : Finset Action) (alpha : Real) (p : Action -> Real) (hp : FTRL.finiteSimplex arms p) (halpha : alpha ≠ 1) : 0 <= powerSum arms alpha p ∧ 1 - alpha ≠ 0 ∧ negEntropyRegularizer arms alpha p = - ((powerSum arms alpha p - 1) / (1 - alpha))","missing":[],"search":"negentropyregularizer_welldefined_on_finitesimplex banditrlproof.tsallis.negentropyregularizer_welldefined_on_finitesimplex well-definedness package for the finite-simplex tsallis regularizer. the two facts exposed here are the local obligations needed before later tsallis/ftrl leaves can use `real.rpow` algebra and division by `1 - alpha`. theorem compiled","shard":"modules/bef5147b4a2f36c0.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_le_linearLoss_sub_of_minimizers","label":"halfTsallisPotentialStability_le_linearLoss_sub_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialStability_le_linearLoss_sub_of_minimizers","description":"Comparing the old objective at its minimizer with the updated minimizer reduces conjugate-potential stability to a difference of linear losses.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-7e7d13384c06","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10164,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialStability_le_linearLoss_sub_of_minimizers {Action : Type u} (arms : Finset Action) (eta : Real) (score probability estimate next : Action -> Real) (heta : 0 < eta) (hprobabilityMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score probability) (hnextMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + estimate action) next) : halfTsallisPotentialStability arms eta score probability estimate next <= FTRL.linearLoss arms probability estimate - FTRL.linearLoss arms next estimate","missing":[],"search":"halftsallispotentialstability_le_linearloss_sub_of_minimizers banditrlproof.tsallis.halftsallispotentialstability_le_linearloss_sub_of_minimizers comparing the old objective at its minimizer with the updated minimizer reduces conjugate-potential stability to a difference of linear losses. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","label":"halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","description":"An ordinary importance-weighted minimizer step is at most one, with no upper bound on the positive learning rate.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-2bd70397c59b","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10165,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:72"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score probability loss next : Action -> Real) (chosen : Action) (hchosen : chosen ∈ arms) (heta : 0 < eta) (hprobabilityMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score probability) (hnextMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss probability loss chosen action) next) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : halfTsallisPotentialStability arms eta score probability (Exp3.importanceWeightedLoss probability loss chosen) next <= 1","missing":[],"search":"halftsallispotentialstability_importanceweightedloss_le_one_of_minimizers banditrlproof.tsallis.halftsallispotentialstability_importanceweightedloss_le_one_of_minimizers an ordinary importance-weighted minimizer step is at most one, with no upper bound on the positive learning rate. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","label":"sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","description":"Averaging the arbitrary-rate pointwise bound under the current simplex keeps the coarse one-round budget equal to one.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-243683af4f08","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10166,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:113"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers {Action : Type u} [DecidableEq Action] (arms : Finset Action) (eta : Real) (score probability loss : Action -> Real) (next : Action -> Action -> Real) (heta : 0 < eta) (hprobabilityMin : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score probability) (hnextMin : forall chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun action => score action + Exp3.importanceWeightedLoss probability loss chosen action) (next chosen)) (hloss : forall action, action ∈ arms -> 0 <= loss action ∧ loss action <= 1) : arms.sum (fun chosen => probability chosen * halfTsallisPotentialStability arms eta score probability (Exp3.importanceWeightedLoss probability loss chosen)…","missing":[],"search":"sum_prob_mul_halftsallispotentialstability_importanceweightedloss_le_one_of_minimizers banditrlproof.tsallis.sum_prob_mul_halftsallispotentialstability_importanceweightedloss_le_one_of_minimizers averaging the arbitrary-rate pointwise bound under the current simplex keeps the coarse one-round budget equal to one. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel_coarse","label":"integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel_coarse","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel_coarse","description":"The finite-action product-law score is integrable for every positive rate. The absolute-value argument uses nonnegativity of true minimizer steps and the coarse averaged bound, so no probability floor or upper rate bound appears.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-759562296d3b","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10167,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:151"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel_coarse {History : Type u} {Action : Type v} [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (historyMu : Measure History) [IsFiniteMeasure historyMu] (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (source : Exp3.MeasurableFiniteActionDistribution arms prob) (heta : 0 < eta) (hprobMin : forall history, FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (score history) (prob history)) (hnextMin : forall history chosen, chosen ∈ arms -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun candidate => score history candidate + Exp3.importanceWeightedLoss (prob hi…","missing":[],"search":"integrable_importanceweightedpotentialstabilityscore_finiteactionkernel_coarse banditrlproof.tsallis.integrable_importanceweightedpotentialstabilityscore_finiteactionkernel_coarse the finite-action product-law score is integrable for every positive rate. the absolute-value argument uses nonnegativity of true minimizer steps and the coarse averaged bound, so no probability floor or upper rate bound appears. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_importanceWeightedPotentialStabilityScore_le_integral_one_of_condDistrib_of_minimizers","label":"integral_importanceWeightedPotentialStabilityScore_le_integral_one_of_condDistrib_of_minimizers","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_importanceWeightedPotentialStabilityScore_le_integral_one_of_condDistrib_of_minimizers","description":"An identified finite conditional action law transports the arbitrary-rate ordinary-IW bound to a one-round integral inequality with constant budget.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-84b125c350a6","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10168,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:245"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_importanceWeightedPotentialStabilityScore_le_integral_one_of_condDistrib_of_minimizers {Omega : Type u} {History : Type v} {Action : Type w} [MeasurableSpace Omega] [MeasurableSpace History] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (mu : Measure Omega) [IsFiniteMeasure mu] (history : Omega -> History) (hhistory : Measurable history) (action : Omega -> Action) (haction : Measurable action) (arms : Finset Action) (eta : Real) (score prob loss : History -> Action -> Real) (next : History -> Action -> Action -> Real) (policy : Kernel History Action) [IsMarkovKernel policy] (hpolicy : policy =ᵐ[mu.map history] fun h => Exp3.finiteActionMeasure arms (prob h)) (hcond : condDistrib action history mu =ᵐ[mu.map history] policy) (heta : 0 < eta) (hprobMin : forall h, FTRL.IsRegularizedMinimizer (F…","missing":[],"search":"integral_importanceweightedpotentialstabilityscore_le_integral_one_of_conddistrib_of_minimizers banditrlproof.tsallis.integral_importanceweightedpotentialstabilityscore_le_integral_one_of_conddistrib_of_minimizers an identified finite conditional action law transports the arbitrary-rate ordinary-iw bound to a one-round integral inequality with constant budget. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_one","label":"integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_one","description":"A generated scheduled successor term is integrable and has coarse expected budget one at every positive local rate.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-843b844c9623","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10169,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:335"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (n + 1)) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample (n + 1)) mu ∧ integral mu (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms…","missing":[],"search":"integral_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_one banditrlproof.tsallis.integral_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_one a generated scheduled successor term is integrable and has coarse expected budget one at every positive local rate. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","label":"integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","description":"One generated successor round exposes the refined budget directly, rather than only through a sum whose rates are all assumed small.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-5a2cd6232cae","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10170,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:497"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (n + 1)) (heta_le : eta (n + 1) <= 1 / 2) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample (n + 1)) mu ∧ Integrable (fun sample => sampledScheduledHalfT…","missing":[],"search":"integral_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_refined banditrlproof.tsallis.integral_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_refined one generated successor round exposes the refined budget directly, rather than only through a sum whose rates are all assumed small. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_one","label":"integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_one","description":"The generated scheduled time-zero term has the same arbitrary-rate coarse budget as successor terms.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-a43fdf3909e8","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10171,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:658"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_one {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : 0 < eta 0) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample 0) mu ∧ integral mu (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample 0) <= i…","missing":[],"search":"integral_sampledscheduledhalftsallisinitialpotentialstabilityattime_le_one banditrlproof.tsallis.integral_sampledscheduledhalftsallisinitialpotentialstabilityattime_le_one the generated scheduled time-zero term has the same arbitrary-rate coarse budget as successor terms. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSuccessorAllRatePotentialStabilityBoundAt","label":"sampledScheduledHalfTsallisSuccessorAllRatePotentialStabilityBoundAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSuccessorAllRatePotentialStabilityBoundAt","description":"Piecewise successor budget: use the refined expression at small local rates and the coarse constant otherwise.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-0b2a6794e9ed","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10172,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:794"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisSuccessorAllRatePotentialStabilityBoundAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (input : Env × History.FinitePairHistory Action Real n) : Real","missing":[],"search":"sampledscheduledhalftsallissuccessorallratepotentialstabilityboundat banditrlproof.tsallis.sampledscheduledhalftsallissuccessorallratepotentialstabilityboundat piecewise successor budget: use the refined expression at small local rates and the coarse constant otherwise. definition compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound","label":"sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound","description":"Piecewise initial budget with the same local-rate threshold.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-127dfd339df0","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10173,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:804"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (env : Env) : Real","missing":[],"search":"sampledscheduledhalftsallisinitialallratepotentialstabilitybound banditrlproof.tsallis.sampledscheduledhalftsallisinitialallratepotentialstabilitybound piecewise initial budget with the same local-rate threshold. definition compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime","label":"sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime","description":"Piecewise budget indexed by the actual scheduled time.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-b7ef486072b6","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10174,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:814"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) : Nat -> Real | 0 => sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound arms harms eta sample.1 | n + 1 => sampledScheduledHalfTsallisSuccessorAllRatePotentialStabilityBoundAt arms harms eta n (sampledScheduledHalfTsallisHistoryAt n sample) /-- A successor term and its piecewise all-rate budget are integrable, and the expected term is bounded by that budget. -/ theorem integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_allRateBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty A…","missing":[],"search":"sampledscheduledhalftsallisallratepotentialstabilityboundattime banditrlproof.tsallis.sampledscheduledhalftsallisallratepotentialstabilityboundattime piecewise budget indexed by the actual scheduled time. definition compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_allRateBound","label":"integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_allRateBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_allRateBound","description":"A successor term and its piecewise all-rate budget are integrable, and the expected term is bounded by that budget.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-c5b82e2c66eb","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10175,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:826"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_allRateBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (n + 1)) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample (n + 1)) mu ∧ Integrable (fun sample => sampledScheduledHalfTsallisSuccessorAllRatePotent…","missing":[],"search":"integral_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_allratebound banditrlproof.tsallis.integral_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_allratebound a successor term and its piecewise all-rate budget are integrable, and the expected term is bounded by that budget. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_allRateBound","label":"integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_allRateBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_allRateBound","description":"The initial term and its piecewise all-rate budget satisfy the analogous integrability and expectation contract.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-46dbd5801c47","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10176,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:897"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_allRateBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : 0 < eta 0) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample 0) mu ∧ Integrable (fun sample => sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound arms har…","missing":[],"search":"integral_sampledscheduledhalftsallisinitialpotentialstabilityattime_le_allratebound banditrlproof.tsallis.integral_sampledscheduledhalftsallisinitialpotentialstabilityattime_le_allratebound the initial term and its piecewise all-rate budget satisfy the analogous integrability and expectation contract. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","label":"integral_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","description":"Every actual scheduled time has an integrable stability term and an integrable piecewise all-rate budget, with the corresponding expectation inequality.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-77c277cb1168","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10177,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:973"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (heta : 0 < eta t) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample t) mu ∧ Integrable (fun sample => sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime arms h…","missing":[],"search":"integral_sampledscheduledhalftsallispotentialstabilityattime_le_allratebound banditrlproof.tsallis.integral_sampledscheduledhalftsallispotentialstabilityattime_le_allratebound every actual scheduled time has an integrable stability term and an integrable piecewise all-rate budget, with the corresponding expectation inequality. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_allRate","label":"integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_allRate","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_allRate","description":"Under positivity of every included local rate, both the exact full scheduled stability sum and its piecewise all-rate budget are integrable.","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-a6c029a3d9f2","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10178,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:1013"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_allRate {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : forall t, t <= horizon -> 0 < eta t) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => (Finset.range (horizon + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample t)) mu ∧ Integrable (fun sample…","missing":[],"search":"integrable_sum_sampledscheduledhalftsallispotentialstabilityattime_allrate banditrlproof.tsallis.integrable_sum_sampledscheduledhalftsallispotentialstabilityattime_allrate under positivity of every included local rate, both the exact full scheduled stability sum and its piecewise all-rate budget are integrable. theorem compiled","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","label":"integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","description":"Under positivity of every included local rate, both the exact full scheduled stability sum and its piecewise all-rate budget are integrable. -/ theorem integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_allRate {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [Decida…","url":"../modules/banditrlproof-tsallisscheduledallrateexpectedstability/index.html#decl-2f3dc44dba91","parent":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","order":10179,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllRateExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllRateExpectedStability.lean:1063"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : forall t, t <= horizon -> 0 < eta t) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range (horizon + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample t)) <= integral mu (fun…","missing":[],"search":"integral_sum_sampledscheduledhalftsallispotentialstabilityattime_le_allratebound banditrlproof.tsallis.integral_sum_sampledscheduledhalftsallispotentialstabilityattime_le_allratebound under positivity of every included local rate, both the exact full scheduled stability sum and its piecewise all-rate budget are integrable. -/ theorem integrable_sum_sampledscheduledhalftsallispotentialstabilityattime_allrate {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (horizon : nat) (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (heta : forall t, t <= horizon -> 0 < eta t) (loss : exp3.predictablelossvector env action) : let selector := canonicalhalftsallisschedulegeneratedselectormeasurability arms harms eta loss let mu := prior ⊗ₘ sampledscheduledhalftsallistrajectorykernel arms harms eta selector.finitehistory loss.environment integrable (fun sample => (finset.range (horizon + 1)).sum (fun t => sampledscheduledhalftsallispotentialstabilityattime arms harms eta sample t)) mu ∧ integrable (fun sample => (finset.range (horizon + 1)).sum (fun t => sampledscheduledhalftsallisallratepotentialstabilityboundattime arms harms eta sample t)) mu := by dsimp only let selector :=…","shard":"modules/6f27333bec391692.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","label":"integrable_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","description":"One actual scheduled successor potential term is integrable under the canonical generated trajectory.","url":"../modules/banditrlproof-tsallisscheduledalltimesexpectedstability/index.html#decl-e2d6ec205f61","parent":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","order":10180,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllTimesExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllTimesExpectedStability.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (n + 1)) (heta_le : eta (n + 1) <= 1 / 2) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample (n + 1)) mu","missing":[],"search":"integrable_sampledscheduledhalftsallissuccessorpotentialstabilityattime banditrlproof.tsallis.integrable_sampledscheduledhalftsallissuccessorpotentialstabilityattime one actual scheduled successor potential term is integrable under the canonical generated trajectory. theorem compiled","shard":"modules/6066298328b3c736.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","label":"integrable_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","description":"The complete scheduled successor stability sum is integrable.","url":"../modules/banditrlproof-tsallisscheduledalltimesexpectedstability/index.html#decl-328e5b62ba23","parent":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","order":10181,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllTimesExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllTimesExpectedStability.lean:98"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : forall n, n < horizon -> 0 < eta (n + 1)) (heta_le : forall n, n < horizon -> eta (n + 1) <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => (Finset.range horizon).sum (fun n => sampledScheduledHalfTsallisPotentialStabilityAt…","missing":[],"search":"integrable_sum_sampledscheduledhalftsallissuccessorpotentialstabilityattime banditrlproof.tsallis.integrable_sum_sampledscheduledhalftsallissuccessorpotentialstabilityattime the complete scheduled successor stability sum is integrable. theorem compiled","shard":"modules/6066298328b3c736.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime","label":"integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime","description":"The full scheduled stability sum from time zero through `horizon` is integrable under the canonical generated trajectory.","url":"../modules/banditrlproof-tsallisscheduledalltimesexpectedstability/index.html#decl-fd9525409c73","parent":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","order":10182,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllTimesExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllTimesExpectedStability.lean:135"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => (Finset.range (horizon + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialStabilityAtTime arms har…","missing":[],"search":"integrable_sum_sampledscheduledhalftsallispotentialstabilityattime banditrlproof.tsallis.integrable_sum_sampledscheduledhalftsallispotentialstabilityattime the full scheduled stability sum from time zero through `horizon` is integrable under the canonical generated trajectory. theorem compiled","shard":"modules/6066298328b3c736.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_refined","label":"integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_refined","description":"The full scheduled stability sum from time zero through `horizon` is integrable under the canonical generated trajectory. -/ theorem integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Meas…","url":"../modules/banditrlproof-tsallisscheduledalltimesexpectedstability/index.html#decl-88eb4b71a36d","parent":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","order":10183,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledAllTimesExpectedStability"],["Source","BanditRLProof/TsallisScheduledAllTimesExpectedStability.lean:179"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_refined {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range (horizon + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialStabilityAtTim…","missing":[],"search":"integral_sum_sampledscheduledhalftsallispotentialstabilityattime_le_refined banditrlproof.tsallis.integral_sum_sampledscheduledhalftsallispotentialstabilityattime_le_refined the full scheduled stability sum from time zero through `horizon` is integrable under the canonical generated trajectory. -/ theorem integrable_sum_sampledscheduledhalftsallispotentialstabilityattime {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (horizon : nat) (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (loss : exp3.predictablelossvector env action) : let selector := canonicalhalftsallisschedulegeneratedselectormeasurability arms harms eta loss let mu := prior ⊗ₘ sampledscheduledhalftsallistrajectorykernel arms harms eta selector.finitehistory loss.environment integrable (fun sample => (finset.range (horizon + 1)).sum (fun t => sampledscheduledhalftsallispotentialstabilityattime arms harms eta sample t)) mu := by dsimp only let selector := canonicalhalftsallisschedulegeneratedselectormeasurability arms harms eta loss let mu := prior ⊗ₘ sampledscheduledhalftsallistrajectorykernel arms…","shard":"modules/6066298328b3c736.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPastSigma","label":"sampledScheduledHalfTsallisPastSigma","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPastSigma","description":"Information in the generated action/reward trace strictly before the scheduled action at time `t`. At time zero there is no trace information; at time `n + 1` this is the sigma-algebra generated by the prefix through `n`.","url":"../modules/banditrlproof-tsallisscheduledconditionalmeangap/index.html#decl-c3462177101e","parent":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","order":10184,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledConditionalMeanGap"],["Source","BanditRLProof/TsallisScheduledConditionalMeanGap.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"@[reducible] def sampledScheduledHalfTsallisPastSigma {Env : Type u} {Action : Type v} [MeasurableSpace Action] (t : Nat) : MeasurableSpace (Env × ((k : Nat) -> Action × Real))","missing":[],"search":"sampledscheduledhalftsallispastsigma banditrlproof.tsallis.sampledscheduledhalftsallispastsigma information in the generated action/reward trace strictly before the scheduled action at time `t`. at time zero there is no trace information; at time `n + 1` this is the sigma-algebra generated by the prefix through `n`. definition compiled","shard":"modules/8bdf146310c99785.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPastSigma_le","label":"sampledScheduledHalfTsallisPastSigma_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPastSigma_le","description":"The scheduled past sigma-algebra is a sub-sigma-algebra of the ambient trajectory sigma-algebra.","url":"../modules/banditrlproof-tsallisscheduledconditionalmeangap/index.html#decl-e5a36710b22a","parent":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","order":10185,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledConditionalMeanGap"],["Source","BanditRLProof/TsallisScheduledConditionalMeanGap.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPastSigma_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (t : Nat) : sampledScheduledHalfTsallisPastSigma (Env := Env) (Action := Action) t <= (inferInstance : MeasurableSpace (Env × ((k : Nat) -> Action × Real)))","missing":[],"search":"sampledscheduledhalftsallispastsigma_le banditrlproof.tsallis.sampledscheduledhalftsallispastsigma_le the scheduled past sigma-algebra is a sub-sigma-algebra of the ambient trajectory sigma-algebra. theorem compiled","shard":"modules/8bdf146310c99785.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisProbabilityAtTime_pastSigma","label":"measurable_sampledScheduledHalfTsallisProbabilityAtTime_pastSigma","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisProbabilityAtTime_pastSigma","description":"Every scheduled action-probability coordinate is measurable using only the trace prefix available before that action is sampled.","url":"../modules/banditrlproof-tsallisscheduledconditionalmeangap/index.html#decl-ac155a022a3b","parent":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","order":10186,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledConditionalMeanGap"],["Source","BanditRLProof/TsallisScheduledConditionalMeanGap.lean:48"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisProbabilityAtTime_pastSigma {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : @Measurable (Env × ((k : Nat) -> Action × Real)) Real (sampledScheduledHalfTsallisPastSigma (Env := Env) (Action := Action) t) inferInstance (fun sample => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample candidate)","missing":[],"search":"measurable_sampledscheduledhalftsallisprobabilityattime_pastsigma banditrlproof.tsallis.measurable_sampledscheduledhalftsallisprobabilityattime_pastsigma every scheduled action-probability coordinate is measurable using only the trace prefix available before that action is sampled. theorem compiled","shard":"modules/8bdf146310c99785.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledConditionalMeanGapLaw","label":"HasScheduledConditionalMeanGapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledConditionalMeanGapLaw","description":"Coordinatewise conditional-mean gap law relative to the information available before the scheduled action at each time.","url":"../modules/banditrlproof-tsallisscheduledconditionalmeangap/index.html#decl-2a2c84623433","parent":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","order":10187,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledConditionalMeanGap"],["Source","BanditRLProof/TsallisScheduledConditionalMeanGap.lean:79"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledConditionalMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Action -> Real) (horizon : Nat) : Prop","missing":[],"search":"hasscheduledconditionalmeangaplaw banditrlproof.tsallis.hasscheduledconditionalmeangaplaw coordinatewise conditional-mean gap law relative to the information available before the scheduled action at each time. definition compiled","shard":"modules/8bdf146310c99785.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledExpectedGapLaw_of_conditionalMeanGapLaw","label":"hasScheduledExpectedGapLaw_of_conditionalMeanGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledExpectedGapLaw_of_conditionalMeanGapLaw","description":"A constant conditional loss-gap law implies the expected-gap law consumed by scheduled stochastic self-bounding. The proof uses only Mathlib's pull-out property and preservation of the integral by conditional expectation.","url":"../modules/banditrlproof-tsallisscheduledconditionalmeangap/index.html#decl-e0d64fca875f","parent":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","order":10188,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledConditionalMeanGap"],["Source","BanditRLProof/TsallisScheduledConditionalMeanGap.lean:98"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledExpectedGapLaw_of_conditionalMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Action -> Real) (horizon : Nat) (hgap : HasScheduledConditionalMeanGapLaw mu arms loss best gap horizon) : HasScheduledExpectedGapLaw mu arms harms eta loss best gap horizon","missing":[],"search":"hasscheduledexpectedgaplaw_of_conditionalmeangaplaw banditrlproof.tsallis.hasscheduledexpectedgaplaw_of_conditionalmeangaplaw a constant conditional loss-gap law implies the expected-gap law consumed by scheduled stochastic self-bounding. the proof uses only mathlib's pull-out property and preservation of the integral by conditional expectation. theorem compiled","shard":"modules/8bdf146310c99785.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledExpectedGapLaw","label":"HasScheduledExpectedGapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledExpectedGapLaw","description":"The per-time, per-suboptimal-arm first-moment law needed by stochastic self-bounding. It is deliberately finer than the final summed regret law so that later conditional-expectation producers can discharge it one coordinate at a time.","url":"../modules/banditrlproof-tsallisscheduledexpectedgapselfbounding/index.html#decl-c28a7abae6e0","parent":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","order":10189,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledExpectedGapSelfBounding.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledExpectedGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Action -> Real) (horizon : Nat) : Prop","missing":[],"search":"hasscheduledexpectedgaplaw banditrlproof.tsallis.hasscheduledexpectedgaplaw the per-time, per-suboptimal-arm first-moment law needed by stochastic self-bounding. it is deliberately finer than the final summed regret law so that later conditional-expectation producers can discharge it one coordinate at a time. definition compiled","shard":"modules/7b012c9da4988fc4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbability_mul_predictableLossDiffAt","label":"integrable_sampledScheduledHalfTsallisProbability_mul_predictableLossDiffAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbability_mul_predictableLossDiffAt","description":"One probability-weighted predictable loss difference is integrable.","url":"../modules/banditrlproof-tsallisscheduledexpectedgapselfbounding/index.html#decl-5af42dcafb2c","parent":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","order":10190,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledExpectedGapSelfBounding.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisProbability_mul_predictableLossDiffAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best action : Action) (haction : action ∈ arms) (t : Nat) : Integrable (fun sample => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample action * (Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best)) mu","missing":[],"search":"integrable_sampledscheduledhalftsallisprobability_mul_predictablelossdiffat banditrlproof.tsallis.integrable_sampledscheduledhalftsallisprobability_mul_predictablelossdiffat one probability-weighted predictable loss difference is integrable. theorem compiled","shard":"modules/7b012c9da4988fc4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_weightedLossGapMass","label":"sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_weightedLossGapMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_weightedLossGapMass","description":"Pathwise scheduled regret against a best-arm point mass is the finite sum of probability-weighted predictable loss differences over suboptimal arms.","url":"../modules/banditrlproof-tsallisscheduledexpectedgapselfbounding/index.html#decl-fd4ea5e00865","parent":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","order":10191,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledExpectedGapSelfBounding.lean:77"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_weightedLossGapMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon sample = (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample action * (Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best)))","missing":[],"search":"sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_weightedlossgapmass banditrlproof.tsallis.sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_weightedlossgapmass pathwise scheduled regret against a best-arm point mass is the finite sum of probability-weighted predictable loss differences over suboptimal arms. theorem compiled","shard":"modules/7b012c9da4988fc4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass_of_expectedGapLaw","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass_of_expectedGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass_of_expectedGapLaw","description":"A coordinatewise expected-gap law identifies integrated scheduled regret with the expected suboptimal-arm gap mass.","url":"../modules/banditrlproof-tsallisscheduledexpectedgapselfbounding/index.html#decl-62ec24a917d2","parent":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","order":10192,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledExpectedGapSelfBounding.lean:109"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass_of_expectedGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgapLaw : HasScheduledExpectedGapLaw mu arms harms eta loss best gap horizon) : integral mu (sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon) = (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => gap action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t actio…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimalexpectedgapmass_of_expectedgaplaw banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimalexpectedgapmass_of_expectedgaplaw a coordinatewise expected-gap law identifies integrated scheduled regret with the expected suboptimal-arm gap mass. theorem compiled","shard":"modules/7b012c9da4988fc4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_expectedGapLaw","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_expectedGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_expectedGapLaw","description":"A nonnegative corruption allowance turns the expected-gap identity into the self-bounding premise consumed by completion of squares.","url":"../modules/banditrlproof-tsallisscheduledexpectedgapselfbounding/index.html#decl-51329ace417a","parent":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","order":10193,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledExpectedGapSelfBounding.lean:159"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_expectedGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgapLaw : HasScheduledExpectedGapLaw mu arms harms eta loss best gap horizon) (corruption : Real) (hcorruption : 0 <= corruption) : (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => gap action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action)) - corruption <= integral mu (sampledScheduledHalfTsallisPredictableEnvironmentRegret…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_expectedgaplaw banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_expectedgaplaw a nonnegative corruption allowance turns the expected-gap identity into the self-bounding premise consumed by completion of squares. theorem compiled","shard":"modules/7b012c9da4988fc4.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedLossAt","label":"sampledScheduledHalfTsallisPredictableEstimatedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedLossAt","description":"Predictable importance-weighted loss using the scheduled probability at the same actual trajectory time.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-1e5a9c2a217d","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10194,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledscheduledhalftsallispredictableestimatedlossat banditrlproof.tsallis.sampledscheduledhalftsallispredictableestimatedlossat predictable importance-weighted loss using the scheduled probability at the same actual trajectory time. definition compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret","label":"sampledScheduledHalfTsallisPredictableEnvironmentRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret","description":"Predictable environment regret of the scheduled generated probabilities against a fixed comparator through the inclusive terminal time.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-c9326518de57","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10195,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPredictableEnvironmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallispredictableenvironmentregret banditrlproof.tsallis.sampledscheduledhalftsallispredictableenvironmentregret predictable environment regret of the scheduled generated probabilities against a fixed comparator through the inclusive terminal time. definition compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedRegret","label":"sampledScheduledHalfTsallisPredictableEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedRegret","description":"Finite-horizon scheduled regret after replacing stored rewards by the predictable importance-weighted loss vectors.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-bb0503bf280a","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10196,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:51"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPredictableEstimatedRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallispredictableestimatedregret banditrlproof.tsallis.sampledscheduledhalftsallispredictableestimatedregret finite-horizon scheduled regret after replacing stored rewards by the predictable importance-weighted loss vectors. definition compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisProbabilityAtTime","label":"measurable_sampledScheduledHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisProbabilityAtTime","description":"theorem measurable_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledScheduledHalfTsall…","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-e143385d2e8f","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10197,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:68"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample candidate)","missing":[],"search":"measurable_sampledscheduledhalftsallisprobabilityattime banditrlproof.tsallis.measurable_sampledscheduledhalftsallisprobabilityattime theorem measurable_sampledscheduledhalftsallisprobabilityattime {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledscheduledhalftsallisprobabilityattime arms harms eta t sample candidate) theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisPredictableEstimatedLossAt","label":"measurable_sampledScheduledHalfTsallisPredictableEstimatedLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisPredictableEstimatedLossAt","description":"theorem measurable_sampledScheduledHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × (…","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-5f41dbd0e9e5","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10198,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:93"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisPredictableEstimatedLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : Env × ((k : Nat) -> Action × Real) => sampledScheduledHalfTsallisPredictableEstimatedLossAt arms harms eta loss t sample candidate)","missing":[],"search":"measurable_sampledscheduledhalftsallispredictableestimatedlossat banditrlproof.tsallis.measurable_sampledscheduledhalftsallispredictableestimatedlossat theorem measurable_sampledscheduledhalftsallispredictableestimatedlossat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (loss : exp3.predictablelossvector env action) (t : nat) (candidate : action) (hcandidate : candidate ∈ arms) : measurable (fun sample : env × ((k : nat) -> action × real) => sampledscheduledhalftsallispredictableestimatedlossat arms harms eta loss t sample candidate) theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","label":"sampledScheduledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","description":"Deterministic predictable feedback identifies each stored-reward scheduled estimator with its predictable counterpart almost surely.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-11dbdc25e29c","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10199,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:116"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (t : Nat) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment (fun sample => sampledScheduledHalfTsallisObservedEstimatedLossAt arms harms eta t sample) =ᵐ[mu] (fun sample => sampledScheduledHalfTsallisPredictableEstimatedLossAt arms harms eta loss t sample)","missing":[],"search":"sampledscheduledhalftsallisobservedestimatedlossat_eq_predictable_ae banditrlproof.tsallis.sampledscheduledhalftsallisobservedestimatedlossat_eq_predictable_ae deterministic predictable feedback identifies each stored-reward scheduled estimator with its predictable counterpart almost surely. theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedLossAt_first_moments","label":"sampledScheduledHalfTsallisPredictableEstimatedLossAt_first_moments","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedLossAt_first_moments","description":"At each scheduled time, the mixed and comparator-weighted predictable IW estimators are integrable and have their corresponding environment first moments.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-89286e370a12","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10200,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:170"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableEstimatedLossAt_first_moments {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (t : Nat) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment (Integrable (fun sample => FTRL.linearLoss arms (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample) (sampledScheduledHalfTsallisPredictableEstimatedLossAt arms…","missing":[],"search":"sampledscheduledhalftsallispredictableestimatedlossat_first_moments banditrlproof.tsallis.sampledscheduledhalftsallispredictableestimatedlossat_first_moments at each scheduled time, the mixed and comparator-weighted predictable iw estimators are integrable and have their corresponding environment first moments. theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledScheduledHalfTsallisProbabilityAtTime","label":"finiteSimplex_sampledScheduledHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_sampledScheduledHalfTsallisProbabilityAtTime","description":"theorem finiteSimplex_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : FTRL.finiteSimplex arms (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample)","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-724ffe8e5b7f","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10201,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:347"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : FTRL.finiteSimplex arms (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample)","missing":[],"search":"finitesimplex_sampledscheduledhalftsallisprobabilityattime banditrlproof.tsallis.finitesimplex_sampledscheduledhalftsallisprobabilityattime theorem finitesimplex_sampledscheduledhalftsallisprobabilityattime {env : type u} {action : type v} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (t : nat) (sample : env × ((k : nat) -> action × real)) : ftrl.finitesimplex arms (sampledscheduledhalftsallisprobabilityattime arms harms eta t sample) theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisPredictableLinearLossAt","label":"integrable_sampledScheduledHalfTsallisPredictableLinearLossAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisPredictableLinearLossAt","description":"Current scheduled mixed predictable loss and a fixed-comparator predictable loss are integrable at every actual time.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-fb74c79178ef","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10202,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:369"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisPredictableLinearLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (t : Nat) : Integrable (fun sample => FTRL.linearLoss arms (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample) (Exp3.predictableLossAt loss t sample)) mu ∧ Integrable (fun sample => FTRL.linearLoss arms q (Exp3.predictableLossAt loss t sample)) mu","missing":[],"search":"integrable_sampledscheduledhalftsallispredictablelinearlossat banditrlproof.tsallis.integrable_sampledscheduledhalftsallispredictablelinearlossat current scheduled mixed predictable loss and a fixed-comparator predictable loss are integrable at every actual time. theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","label":"integral_sampledScheduledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","description":"Predictable scheduled estimated regret is integrable and has exactly the same finite-horizon integral as predictable environment regret.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-b20a895d260d","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10203,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:459"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (horizon : Nat) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (sampledScheduledHalfTsallisPredictableEstimatedRegret arms harms eta loss q horizon) mu ∧ Integrable (sampledScheduledHalfTsallisPredictableEnvi…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableestimatedregret_eq_environmentregret banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableestimatedregret_eq_environmentregret predictable scheduled estimated regret is integrable and has exactly the same finite-horizon integral as predictable environment regret. theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_eq_predictable_ae","label":"sampledScheduledHalfTsallisEstimatedRegret_eq_predictable_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_eq_predictable_ae","description":"Observed scheduled estimated regret agrees almost surely with the predictable-estimator version over every finite horizon.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-8475b2c09afe","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10204,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:555"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisEstimatedRegret_eq_predictable_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (horizon : Nat) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment sampledScheduledHalfTsallisEstimatedRegret arms harms eta q horizon =ᵐ[mu] sampledScheduledHalfTsallisPredictableEstimatedRegret arms harms eta loss q horizon","missing":[],"search":"sampledscheduledhalftsallisestimatedregret_eq_predictable_ae banditrlproof.tsallis.sampledscheduledhalftsallisestimatedregret_eq_predictable_ae observed scheduled estimated regret agrees almost surely with the predictable-estimator version over every finite horizon. theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret","label":"integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret","description":"Observed scheduled IW regret is integrable and has exactly the predictable environment-regret integral.","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-005c6f36f96b","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10205,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:597"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (q : Action -> Real) (hq : FTRL.finiteSimplex arms q) (horizon : Nat) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (sampledScheduledHalfTsallisEstimatedRegret arms harms eta q horizon) mu ∧ Integrable (sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms et…","missing":[],"search":"integral_sampledscheduledhalftsallisestimatedregret_eq_environmentregret banditrlproof.tsallis.integral_sampledscheduledhalftsallisestimatedregret_eq_environmentregret observed scheduled iw regret is integrable and has exactly the predictable environment-regret integral. theorem compiled","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","description":"Observed scheduled IW regret is integrable and has exactly the predictable environment-regret integral. -/ theorem integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [Is…","url":"../modules/banditrlproof-tsallisscheduledexpectedregret/index.html#decl-b37ae4da4e92","parent":"module:BanditRLProof.TsallisScheduledExpectedRegret","order":10206,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedRegret"],["Source","BanditRLProof/TsallisScheduledExpectedRegret.lean:648"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (sampledScheduledHalfTsallisPredictableEnvir…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_allratebound banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_allratebound observed scheduled iw regret is integrable and has exactly the predictable environment-regret integral. -/ theorem integral_sampledscheduledhalftsallisestimatedregret_eq_environmentregret {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (loss : exp3.predictablelossvector env action) (q : action -> real) (hq : ftrl.finitesimplex arms q) (horizon : nat) : let selector := canonicalhalftsallisschedulegeneratedselectormeasurability arms harms eta loss let mu := prior ⊗ₘ sampledscheduledhalftsallistrajectorykernel arms harms eta selector.finitehistory loss.environment integrable (sampledscheduledhalftsallisestimatedregret arms harms eta q horizon) mu ∧ integrable (sampledscheduledhalftsallispredictableenvironmentregret arms harms eta loss q horizon) mu ∧ integral mu (sampledscheduledhalftsallisestimatedregret arms harms eta q horizon) = integral mu (sampledscheduledhalftsallispredictableenvironmentregret arms harms et…","shard":"modules/afeded321c420efe.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAt","label":"sampledScheduledHalfTsallisHistoryAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAt","description":"Visible environment/pair-history state before scheduled action `n + 1`.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-49be878d0a68","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10207,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledScheduledHalfTsallisHistoryAt {Env : Type u} {Action : Type v} (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Env × History.FinitePairHistory Action Real n","missing":[],"search":"sampledscheduledhalftsallishistoryat banditrlproof.tsallis.sampledscheduledhalftsallishistoryat visible environment/pair-history state before scheduled action `n + 1`. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisActionAt","label":"sampledScheduledHalfTsallisActionAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisActionAt","description":"Scheduled successor action after the visible prefix through `n`.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-2d1da8516f31","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10208,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:32"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledScheduledHalfTsallisActionAt {Env : Type u} {Action : Type v} (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action","missing":[],"search":"sampledscheduledhalftsallisactionat banditrlproof.tsallis.sampledscheduledhalftsallisactionat scheduled successor action after the visible prefix through `n`. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisScoreAt","label":"sampledScheduledHalfTsallisScoreAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisScoreAt","description":"Scheduled recursive score on an environment/prefix state.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-35bfea91f529","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10209,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisScoreAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledscheduledhalftsallisscoreat banditrlproof.tsallis.sampledscheduledhalftsallisscoreat scheduled recursive score on an environment/prefix state. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableLossAt","label":"sampledScheduledHalfTsallisPredictableLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableLossAt","description":"Predictable successor loss on a scheduled environment/prefix state.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-eadff0799c2a","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10210,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:47"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def sampledScheduledHalfTsallisPredictableLossAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledscheduledhalftsallispredictablelossat banditrlproof.tsallis.sampledscheduledhalftsallispredictablelossat predictable successor loss on a scheduled environment/prefix state. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisProbabilityAt","label":"sampledScheduledHalfTsallisProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisProbabilityAt","description":"Scheduled successor probability on an environment/prefix state.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-9ec7525fb945","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10211,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:55"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisProbabilityAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledscheduledhalftsallisprobabilityat banditrlproof.tsallis.sampledscheduledhalftsallisprobabilityat scheduled successor probability on an environment/prefix state. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisUpdatedAt","label":"sampledScheduledHalfTsallisUpdatedAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisUpdatedAt","description":"Same-rate canonical update after sampling scheduled action `n + 1`.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-e7ac6ddeceea","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10212,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:64"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisUpdatedAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Action -> Action -> Real","missing":[],"search":"sampledscheduledhalftsallisupdatedat banditrlproof.tsallis.sampledscheduledhalftsallisupdatedat same-rate canonical update after sampling scheduled action `n + 1`. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEnvironmentHistoryDistributionSource","label":"sampledScheduledHalfTsallisEnvironmentHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisEnvironmentHistoryDistributionSource","description":"Environment-lifted measurable scheduled probability source.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-c58128e8df56","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10213,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:77"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisEnvironmentHistoryDistributionSource {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Exp3.MeasurableFiniteActionDistribution arms (sampledScheduledHalfTsallisProbabilityAt (Env := Env) arms harms eta n)","missing":[],"search":"sampledscheduledhalftsallisenvironmenthistorydistributionsource banditrlproof.tsallis.sampledscheduledhalftsallisenvironmenthistorydistributionsource environment-lifted measurable scheduled probability source. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPolicyAt","label":"sampledScheduledHalfTsallisPolicyAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPolicyAt","description":"The scheduled policy comapped to the environment/prefix state.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-1a10cfebd2b6","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10214,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:97"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPolicyAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Kernel (Env × History.FinitePairHistory Action Real n) Action","missing":[],"search":"sampledscheduledhalftsallispolicyat banditrlproof.tsallis.sampledscheduledhalftsallispolicyat the scheduled policy comapped to the environment/prefix state. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPolicyAt_eq_finiteActionKernel","label":"sampledScheduledHalfTsallisPolicyAt_eq_finiteActionKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPolicyAt_eq_finiteActionKernel","description":"The environment-lifted scheduled policy is the corresponding finite action kernel.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-d911b258dd96","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10215,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPolicyAt_eq_finiteActionKernel {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : sampledScheduledHalfTsallisPolicyAt (Env := Env) arms harms eta selector n = Exp3.finiteActionKernel arms (sampledScheduledHalfTsallisProbabilityAt (Env := Env) arms harms eta n) (sampledScheduledHalfTsallisEnvironmentHistoryDistributionSource (Env := Env) arms harms eta selector n)","missing":[],"search":"sampledscheduledhalftsallispolicyat_eq_finiteactionkernel banditrlproof.tsallis.sampledscheduledhalftsallispolicyat_eq_finiteactionkernel the environment-lifted scheduled policy is the corresponding finite action kernel. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HalfTsallisScheduleGeneratedSelectorMeasurability","label":"HalfTsallisScheduleGeneratedSelectorMeasurability","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.HalfTsallisScheduleGeneratedSelectorMeasurability","description":"Coordinate regularity for the scheduled current and same-rate updated selectors.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-eba02b238a95","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10216,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:149"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure HalfTsallisScheduleGeneratedSelectorMeasurability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) : Prop where","missing":[],"search":"halftsallisschedulegeneratedselectormeasurability banditrlproof.tsallis.halftsallisschedulegeneratedselectormeasurability coordinate regularity for the scheduled current and same-rate updated selectors. structure compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisUpdatedAt_canonical","label":"measurable_sampledScheduledHalfTsallisUpdatedAt_canonical","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisUpdatedAt_canonical","description":"Every supported coordinate of the scheduled same-rate update is measurable.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-0cf74065e654","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10217,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:165"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisUpdatedAt_canonical {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (candidate : Action) (hcandidate : candidate ∈ arms) : Measurable (fun sample : (Env × History.FinitePairHistory Action Real n) × Action => sampledScheduledHalfTsallisUpdatedAt arms harms eta loss n sample.1 sample.2 candidate)","missing":[],"search":"measurable_sampledscheduledhalftsallisupdatedat_canonical banditrlproof.tsallis.measurable_sampledscheduledhalftsallisupdatedat_canonical every supported coordinate of the scheduled same-rate update is measurable. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisScheduleGeneratedSelectorMeasurability","label":"canonicalHalfTsallisScheduleGeneratedSelectorMeasurability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.canonicalHalfTsallisScheduleGeneratedSelectorMeasurability","description":"The canonical scheduled selector satisfies current and updated coordinate measurability at every successor round.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-47c6605731a7","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10218,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:236"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHalfTsallisScheduleGeneratedSelectorMeasurability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) : HalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss where","missing":[],"search":"canonicalhalftsallisschedulegeneratedselectormeasurability banditrlproof.tsallis.canonicalhalftsallisschedulegeneratedselectormeasurability the canonical scheduled selector satisfies current and updated coordinate measurability at every successor round. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","label":"sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","description":"Scheduled one-round conjugate-potential score on a visible prefix and sampled successor action.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-cba8b2b18e9e","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10219,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:253"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : (Env × History.FinitePairHistory Action Real n) × Action -> Real","missing":[],"search":"sampledscheduledhalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.sampledscheduledhalftsallishistoryactionpotentialstabilityat scheduled one-round conjugate-potential score on a visible prefix and sampled successor action. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","label":"sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","description":"Refined one-round budget for scheduled successor action `n + 1`.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-6ec702f4a2ff","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10220,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:266"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) : Env × History.FinitePairHistory Action Real n -> Real","missing":[],"search":"sampledscheduledhalftsallisrefinedpotentialstabilityboundat banditrlproof.tsallis.sampledscheduledhalftsallisrefinedpotentialstabilityboundat refined one-round budget for scheduled successor action `n + 1`. definition compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","label":"measurable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","description":"The canonical scheduled successor potential score is measurable.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-f24cadda6803","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10221,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:275"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : Measurable (sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt arms harms eta loss n)","missing":[],"search":"measurable_sampledscheduledhalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.measurable_sampledscheduledhalftsallishistoryactionpotentialstabilityat the canonical scheduled successor potential score is measurable. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","label":"integrable_sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","description":"The scheduled refined budget is automatically integrable under any finite visible-history law when its local rate is positive.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-49fe75062cd1","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10222,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:303"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (n : Nat) (historyMu : Measure (Env × History.FinitePairHistory Action Real n)) [IsFiniteMeasure historyMu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : 0 < eta (n + 1)) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) : Integrable (sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt (Env := Env) arms harms eta n) historyMu","missing":[],"search":"integrable_sampledscheduledhalftsallisrefinedpotentialstabilityboundat banditrlproof.tsallis.integrable_sampledscheduledhalftsallisrefinedpotentialstabilityboundat the scheduled refined budget is automatically integrable under any finite visible-history law when its local rate is positive. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","label":"integrable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","description":"The canonical scheduled successor potential score is automatically integrable under its visible-history/action product law.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-455819a04edd","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10223,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:327"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) (heta : 0 < eta (n + 1)) (heta_le : eta (n + 1) <= 1 / 2) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt arms harms eta loss n) (mu.map (sampledScheduledHalfTsallisHistoryAt n) ⊗ₘ sampledScheduledHalfTs…","missing":[],"search":"integrable_sampledscheduledhalftsallishistoryactionpotentialstabilityat banditrlproof.tsallis.integrable_sampledscheduledhalftsallishistoryactionpotentialstabilityat the canonical scheduled successor potential score is automatically integrable under its visible-history/action product law. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime_succ_eq_historyAction_ae","label":"sampledScheduledHalfTsallisPotentialStabilityAtTime_succ_eq_historyAction_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime_succ_eq_historyAction_ae","description":"The actual scheduled successor potential term agrees almost surely with the canonical visible-history/action potential score.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-8d1a3ffe17d3","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10224,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:384"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPotentialStabilityAtTime_succ_eq_historyAction_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (n : Nat) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample (n + 1)) =ᵐ[mu] (fun sample => sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt arms harms eta loss n (sampledScheduledHal…","missing":[],"search":"sampledscheduledhalftsallispotentialstabilityattime_succ_eq_historyaction_ae banditrlproof.tsallis.sampledscheduledhalftsallispotentialstabilityattime_succ_eq_historyaction_ae the actual scheduled successor potential term agrees almost surely with the canonical visible-history/action potential score. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","label":"integral_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","description":"Expected finite-horizon refined stability for all generated successor rounds. The left side is exactly the same-rate stability sum consumed by the pathwise scheduled score/penalty decomposition, restricted to actual times `n + 1`. Time zero and rates above `1 / 2` are deliberately not claimed.","url":"../modules/banditrlproof-tsallisscheduledexpectedstability/index.html#decl-9dede6f4f22c","parent":"module:BanditRLProof.TsallisScheduledExpectedStability","order":10225,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledExpectedStability"],["Source","BanditRLProof/TsallisScheduledExpectedStability.lean:479"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : forall n, n < horizon -> 0 < eta (n + 1)) (heta_le : forall n, n < horizon -> eta (n + 1) <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range horizon).sum (fun n => sampledScheduledHalfTsallisPotentialS…","missing":[],"search":"integral_sum_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_refined banditrlproof.tsallis.integral_sum_sampledscheduledhalftsallissuccessorpotentialstabilityattime_le_refined expected finite-horizon refined stability for all generated successor rounds. the left side is exactly the same-rate stability sum consumed by the pathwise scheduled score/penalty decomposition, restricted to actual times `n + 1`. time zero and rates above `1 / 2` are deliberately not claimed. theorem compiled","shard":"modules/3bfc0e062b30519e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableSuboptimalGapMass","label":"sampledScheduledHalfTsallisPredictableSuboptimalGapMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableSuboptimalGapMass","description":"Pathwise suboptimal-arm gap mass of the scheduled generated laws.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html#decl-54b8666cc1f0","parent":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","order":10226,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledFixedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPredictableSuboptimalGapMass {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (best : Action) (gap : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallispredictablesuboptimalgapmass banditrlproof.tsallis.sampledscheduledhalftsallispredictablesuboptimalgapmass pathwise suboptimal-arm gap mass of the scheduled generated laws. definition compiled","shard":"modules/2efdc5dff2ff34e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalGapMass","label":"sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalGapMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalGapMass","description":"An exact predictable fixed-gap law identifies scheduled environment regret against the best-arm point mass with suboptimal-arm gap mass pathwise.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html#decl-fc5d434d48c9","parent":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","order":10227,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledFixedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalGapMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) (sample : Env × ((k : Nat) -> Action × Real)) : sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon sample = sampledScheduledHalfTsallisPredictableSuboptimalGapMass arms harms eta best gap horizon sample","missing":[],"search":"sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimalgapmass banditrlproof.tsallis.sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimalgapmass an exact predictable fixed-gap law identifies scheduled environment regret against the best-arm point mass with suboptimal-arm gap mass pathwise. theorem compiled","shard":"modules/2efdc5dff2ff34e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableSuboptimalGapMass_eq","label":"integral_sampledScheduledHalfTsallisPredictableSuboptimalGapMass_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableSuboptimalGapMass_eq","description":"Integrating scheduled pathwise suboptimal gap mass gives the deterministic time-by-arm sum of gaps times expected scheduled action probabilities.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html#decl-c43a1091d532","parent":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","order":10228,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledFixedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean:67"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableSuboptimalGapMass_eq {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (best : Action) (gap : Action -> Real) (horizon : Nat) : integral mu (sampledScheduledHalfTsallisPredictableSuboptimalGapMass arms harms eta best gap horizon) = (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => gap action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action))","missing":[],"search":"integral_sampledscheduledhalftsallispredictablesuboptimalgapmass_eq banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictablesuboptimalgapmass_eq integrating scheduled pathwise suboptimal gap mass gives the deterministic time-by-arm sum of gaps times expected scheduled action probabilities. theorem compiled","shard":"modules/2efdc5dff2ff34e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass","description":"The exact predictable fixed-gap law identifies the integrated scheduled environment regret with the expected suboptimal-arm gap mass.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html#decl-831b87cd7d5f","parent":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","order":10229,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledFixedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean:110"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finite…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimalexpectedgapmass banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimalexpectedgapmass the exact predictable fixed-gap law identifies the integrated scheduled environment regret with the expected suboptimal-arm gap mass. theorem compiled","shard":"modules/2efdc5dff2ff34e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_fixedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_fixedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_fixedGap","description":"Any nonnegative corruption allowance turns the exact fixed-gap identity into the explicit self-bounding inequality required by completion of squares.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html#decl-d11b2ccec8f5","parent":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","order":10230,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledFixedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean:160"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_fixedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) (corruption : Real) (hcorruption : 0 <= corruption) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajec…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_fixedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_fixedgap any nonnegative corruption allowance turns the exact fixed-gap identity into the explicit self-bounding inequality required by completion of squares. theorem compiled","shard":"modules/2efdc5dff2ff34e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_fixedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_fixedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_fixedGap","description":"The generated scheduled regret bound with an explicit self-bounding premise becomes automatic under an exact predictable fixed-gap law.","url":"../modules/banditrlproof-tsallisscheduledfixedgapselfbounding/index.html#decl-cf91b88af0c0","parent":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","order":10231,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledFixedGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledFixedGapSelfBounding.lean:194"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_fixedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) (gap : Action -> Real) (hgapPos : forall action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sam…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_of_fixedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_of_fixedgap the generated scheduled regret bound with an explicit self-bounding premise becomes automatic under an exact predictable fixed-gap law. theorem compiled","shard":"modules/2efdc5dff2ff34e7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasIIDStateCoordinateLocality","label":"HasIIDStateCoordinateLocality","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasIIDStateCoordinateLocality","description":"A predictable loss family reads state coordinate `0` initially and state coordinate `n + 1` after a history of length `n + 1`. The successor value may otherwise depend arbitrarily and measurably on that pre-action history.","url":"../modules/banditrlproof-tsallisschedulediidhistoryadaptive/index.html#decl-d9318536541b","parent":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","order":10232,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisScheduledIIDHistoryAdaptive"],["Source","BanditRLProof/TsallisScheduledIIDHistoryAdaptive.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure HasIIDStateCoordinateLocality {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (loss : Exp3.PredictableLossVector (Nat -> LossState) Action) : Prop where","missing":[],"search":"hasiidstatecoordinatelocality banditrlproof.tsallis.hasiidstatecoordinatelocality a predictable loss family reads state coordinate `0` initially and state coordinate `n + 1` after a history of length `n + 1`. the successor value may otherwise depend arbitrarily and measurably on that pre-action history. structure compiled","shard":"modules/373736b152e0a15b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidStateCoordinateLocality_prefix_eq","label":"sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidStateCoordinateLocality_prefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidStateCoordinateLocality_prefix_eq","description":"Equal IID state prefixes generate equal visible trajectory prefixes for a history-adaptive coordinate-local predictable loss family.","url":"../modules/banditrlproof-tsallisschedulediidhistoryadaptive/index.html#decl-df44e99914ba","parent":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","order":10233,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDHistoryAdaptive"],["Source","BanditRLProof/TsallisScheduledIIDHistoryAdaptive.lean:36"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidStateCoordinateLocality_prefix_eq {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (loss : Exp3.PredictableLossVector (Nat -> LossState) Action) (hlocal : HasIIDStateCoordinateLocality loss) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (environment1 environment2 : Nat -> LossState) (n : Nat) (henvironment : Preorder.frestrictLe n environment1 = Preorder.frestrictLe n environment2) : (sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment environment1).map (Preorder.frestrictLe n) = (sampledScheduledHalfTsall…","missing":[],"search":"sampledscheduledhalftsallistrajectorykernel_map_frestrictle_eq_of_iidstatecoordinatelocality_prefix_eq banditrlproof.tsallis.sampledscheduledhalftsallistrajectorykernel_map_frestrictle_eq_of_iidstatecoordinatelocality_prefix_eq equal iid state prefixes generate equal visible trajectory prefixes for a history-adaptive coordinate-local predictable loss family. theorem compiled","shard":"modules/373736b152e0a15b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDHistoryAdaptivePrefixKernel","label":"sampledScheduledHalfTsallisIIDHistoryAdaptivePrefixKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDHistoryAdaptivePrefixKernel","description":"Finite visible-prefix kernel induced by a coordinate-local predictable loss family on an IID state trace.","url":"../modules/banditrlproof-tsallisschedulediidhistoryadaptive/index.html#decl-5921923c9dcb","parent":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","order":10234,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDHistoryAdaptive"],["Source","BanditRLProof/TsallisScheduledIIDHistoryAdaptive.lean:92"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisIIDHistoryAdaptivePrefixKernel {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] [DecidableEq Action] (fallback : LossState) (loss : Exp3.PredictableLossVector (Nat -> LossState) Action) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Kernel ((i : Finset.Iic n) -> LossState) (History.FinitePairHistory Action Real n)","missing":[],"search":"sampledscheduledhalftsallisiidhistoryadaptiveprefixkernel banditrlproof.tsallis.sampledscheduledhalftsallisiidhistoryadaptiveprefixkernel finite visible-prefix kernel induced by a coordinate-local predictable loss family on an iid state trace. definition compiled","shard":"modules/373736b152e0a15b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisHistoryAdaptiveTrajectoryKernel","label":"hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisHistoryAdaptiveTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisHistoryAdaptiveTrajectoryKernel","description":"The canonical scheduled trajectory of a history-adaptive coordinate-local loss family factors through every finite IID state prefix.","url":"../modules/banditrlproof-tsallisschedulediidhistoryadaptive/index.html#decl-3f1594616c44","parent":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","order":10235,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDHistoryAdaptive"],["Source","BanditRLProof/TsallisScheduledIIDHistoryAdaptive.lean:134"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisHistoryAdaptiveTrajectoryKernel {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (fallback : LossState) (loss : Exp3.PredictableLossVector (Nat -> LossState) Action) (hlocal : HasIIDStateCoordinateLocality loss) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (horizon : Nat) : HasScheduledIIDPrefixKernelFactorization (sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment) horizon","missing":[],"search":"hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallishistoryadaptivetrajectorykernel banditrlproof.tsallis.hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallishistoryadaptivetrajectorykernel the canonical scheduled trajectory of a history-adaptive coordinate-local loss family factors through every finite iid state prefix. theorem compiled","shard":"modules/373736b152e0a15b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.iidLossStatePredictableLossVector","label":"iidLossStatePredictableLossVector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.iidLossStatePredictableLossVector","description":"A predictable loss vector obtained by reading one fresh loss state at each time. Joint measurability of `value` is the sole evaluation regularity contract; pointwise bounds make this a valid `[0,1]` loss process.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-0d1bc8d11aef","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10236,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def iidLossStatePredictableLossVector {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) : Exp3.PredictableLossVector (Nat -> LossState) Action where","missing":[],"search":"iidlossstatepredictablelossvector banditrlproof.tsallis.iidlossstatepredictablelossvector a predictable loss vector obtained by reading one fresh loss state at each time. joint measurability of `value` is the sole evaluation regularity contract; pointwise bounds make this a valid `[0,1]` loss process. definition compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.predictableLossAt_iidLossStatePredictableLossVector","label":"predictableLossAt_iidLossStatePredictableLossVector","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.predictableLossAt_iidLossStatePredictableLossVector","description":"theorem predictableLossAt_iidLossStatePredictableLossVector {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (t : Nat) (sample : (Nat -> LossS…","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-f8bbe3404bcd","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10237,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:61"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_iidLossStatePredictableLossVector {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (t : Nat) (sample : (Nat -> LossState) × ((k : Nat) -> Action × Real)) (action : Action) : Exp3.predictableLossAt (iidLossStatePredictableLossVector value hvalue hvalue_nonneg hvalue_le_one) t sample action = value (sample.1 t) action","missing":[],"search":"predictablelossat_iidlossstatepredictablelossvector banditrlproof.tsallis.predictablelossat_iidlossstatepredictablelossvector theorem predictablelossat_iidlossstatepredictablelossvector {lossstate : type u} {action : type v} [measurablespace lossstate] [measurablespace action] (value : lossstate -> action -> real) (hvalue : measurable (fun input : lossstate × action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (t : nat) (sample : (nat -> lossstate) × ((k : nat) -> action × real)) (action : action) : exp3.predictablelossat (iidlossstatepredictablelossvector value hvalue hvalue_nonneg hvalue_le_one) t sample action = value (sample.1 t) action theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.extendLossStatePrefix","label":"extendLossStatePrefix","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.extendLossStatePrefix","description":"Extend a finite loss-state prefix to an infinite stream by a fixed fallback state after the prefix endpoint.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-674eb6e4414d","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10238,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:80"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def extendLossStatePrefix {LossState : Type u} (fallback : LossState) (n : Nat) (statePrefix : (i : Finset.Iic n) -> LossState) : Nat -> LossState","missing":[],"search":"extendlossstateprefix banditrlproof.tsallis.extendlossstateprefix extend a finite loss-state prefix to an infinite stream by a fixed fallback state after the prefix endpoint. definition compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_extendLossStatePrefix","label":"measurable_extendLossStatePrefix","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_extendLossStatePrefix","description":"theorem measurable_extendLossStatePrefix {LossState : Type u} [MeasurableSpace LossState] (fallback : LossState) (n : Nat) : Measurable (extendLossStatePrefix fallback n)","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-46be79e4b7fa","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10239,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:86"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_extendLossStatePrefix {LossState : Type u} [MeasurableSpace LossState] (fallback : LossState) (n : Nat) : Measurable (extendLossStatePrefix fallback n)","missing":[],"search":"measurable_extendlossstateprefix banditrlproof.tsallis.measurable_extendlossstateprefix theorem measurable_extendlossstateprefix {lossstate : type u} [measurablespace lossstate] (fallback : lossstate) (n : nat) : measurable (extendlossstateprefix fallback n) theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.extendLossStatePrefix_apply_of_le","label":"extendLossStatePrefix_apply_of_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.extendLossStatePrefix_apply_of_le","description":"theorem extendLossStatePrefix_apply_of_le {LossState : Type u} (fallback : LossState) (n t : Nat) (ht : t <= n) (statePrefix : (i : Finset.Iic n) -> LossState) : extendLossStatePrefix fallback n statePrefix t = statePrefix ⟨t, Finset.mem_Iic.mpr ht⟩","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-b8ff05ea9919","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10240,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:98"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem extendLossStatePrefix_apply_of_le {LossState : Type u} (fallback : LossState) (n t : Nat) (ht : t <= n) (statePrefix : (i : Finset.Iic n) -> LossState) : extendLossStatePrefix fallback n statePrefix t = statePrefix ⟨t, Finset.mem_Iic.mpr ht⟩","missing":[],"search":"extendlossstateprefix_apply_of_le banditrlproof.tsallis.extendlossstateprefix_apply_of_le theorem extendlossstateprefix_apply_of_le {lossstate : type u} (fallback : lossstate) (n t : nat) (ht : t <= n) (stateprefix : (i : finset.iic n) -> lossstate) : extendlossstateprefix fallback n stateprefix t = stateprefix ⟨t, finset.mem_iic.mpr ht⟩ theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidLossState_prefix_eq","label":"sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidLossState_prefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidLossState_prefix_eq","description":"The visible scheduled trajectory prefix generated from an IID loss-state stream is unchanged when the stream is replaced by any other stream with the same finite loss-state prefix.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-39f6456359fd","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10241,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:108"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidLossState_prefix_eq {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (environment₁ environment₂ : Nat -> LossState) (n : Nat) (henvironment : Preorder.frestrictLe n environment₁ = Preorder.frestrictLe n environment₂) : (sampledScheduledHalfTsallisTrajectoryKernel…","missing":[],"search":"sampledscheduledhalftsallistrajectorykernel_map_frestrictle_eq_of_iidlossstate_prefix_eq banditrlproof.tsallis.sampledscheduledhalftsallistrajectorykernel_map_frestrictle_eq_of_iidlossstate_prefix_eq the visible scheduled trajectory prefix generated from an iid loss-state stream is unchanged when the stream is replaced by any other stream with the same finite loss-state prefix. theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDPrefixKernel","label":"sampledScheduledHalfTsallisIIDPrefixKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDPrefixKernel","description":"Finite pair-prefix kernel obtained by extending the supplied loss-state prefix and running the canonical scheduled trajectory.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-06c8a23eb5af","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10242,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:171"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisIIDPrefixKernel {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] [DecidableEq Action] (fallback : LossState) (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Kernel ((i : Finset.Iic n) -> LossState) (History.FinitePairHistory Action Real n)","missing":[],"search":"sampledscheduledhalftsallisiidprefixkernel banditrlproof.tsallis.sampledscheduledhalftsallisiidprefixkernel finite pair-prefix kernel obtained by extending the supplied loss-state prefix and running the canonical scheduled trajectory. definition compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.iidLossStateMeanGap","label":"iidLossStateMeanGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.iidLossStateMeanGap","description":"The stationary mean loss gap induced by one coordinate law.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-1d7d560d059e","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10243,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:225"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def iidLossStateMeanGap {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] (law : Measure LossState) (value : LossState -> Action -> Real) (best action : Action) : Real","missing":[],"search":"iidlossstatemeangap banditrlproof.tsallis.iidlossstatemeangap the stationary mean loss gap induced by one coordinate law. definition compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledIIDPrefixKernelFactorization","label":"HasScheduledIIDPrefixKernelFactorization","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledIIDPrefixKernelFactorization","description":"Exact remaining trajectory-law obligation for the IID route. At every scheduled successor time, the generated visible pair prefix is a Markov-kernel extension of the corresponding finite loss-state prefix.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-ec7f10aceeba","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10244,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:235"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledIIDPrefixKernelFactorization {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (trajectoryKernel : Kernel (Nat -> LossState) ((k : Nat) -> Action × Real)) (horizon : Nat) : Prop","missing":[],"search":"hasschedulediidprefixkernelfactorization banditrlproof.tsallis.hasschedulediidprefixkernelfactorization exact remaining trajectory-law obligation for the iid route. at every scheduled successor time, the generated visible pair prefix is a markov-kernel extension of the corresponding finite loss-state prefix. definition compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTrajectoryKernel","label":"hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTrajectoryKernel","description":"The canonical scheduled half-Tsallis trajectory automatically factors through every finite IID loss-state prefix.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-9214c926af37","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10245,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:252"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTrajectoryKernel {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (fallback : LossState) (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (horizon : Nat) : HasScheduledIIDPrefixKernelFactorization (sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector (iidLossStatePredictableLossVector value hvalue…","missing":[],"search":"hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallistrajectorykernel banditrlproof.tsallis.hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallistrajectorykernel the canonical scheduled half-tsallis trajectory automatically factors through every finite iid loss-state prefix. theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledIndependentMeanGapLaw_of_iidLossState","label":"hasScheduledIndependentMeanGapLaw_of_iidLossState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledIndependentMeanGapLaw_of_iidLossState","description":"Infinite-product loss states plus finite-prefix factorization produce the independence and global-mean contract consumed by scheduled self-bounding.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-13d417b353e6","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10246,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:324"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledIndependentMeanGapLaw_of_iidLossState {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (law : Measure LossState) [IsProbabilityMeasure law] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (arms : Finset Action) (best : Action) (horizon : Nat) (trajectoryKernel : Kernel (Nat -> LossState) ((k : Nat) -> Action × Real)) [IsMarkovKernel trajectoryKernel] (hfactor : HasScheduledIIDPrefixKernelFactorization trajectoryKernel horizon) : let prior := Measure.infinitePi (fun _ : Nat => law) let mu := prior ⊗ₘ trajectoryKer…","missing":[],"search":"hasscheduledindependentmeangaplaw_of_iidlossstate banditrlproof.tsallis.hasscheduledindependentmeangaplaw_of_iidlossstate infinite-product loss states plus finite-prefix factorization produce the independence and global-mean contract consumed by scheduled self-bounding. theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_iidLossState","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_iidLossState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_iidLossState","description":"The canonical scheduled half-Tsallis trajectory reaches the explicit logarithmic bound for an IID loss-state environment. Its finite-prefix factorization is constructed internally.","url":"../modules/banditrlproof-tsallisschedulediidmeangap/index.html#decl-d501dca34b13","parent":"module:BanditRLProof.TsallisScheduledIIDMeanGap","order":10247,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDMeanGap.lean:498"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_iidLossState {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (law : Measure LossState) [IsProbabilityMeasure law] (value : LossState -> Action -> Real) (hvalue : Measurable (fun input : LossState × Action => value input.1 input.2)) (hvalue_nonneg : ∀ state action, 0 <= value state action) (hvalue_le_one : ∀ state action, value state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < iidLossStateMeanGap law value best action) (corruption : Real) (hcorruption : 0 <= corruption) : let prior := Measure.…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_iidlossstate banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_iidlossstate the canonical scheduled half-tsallis trajectory reaches the explicit logarithmic bound for an iid loss-state environment. its finite-prefix factorization is constructed internally. theorem compiled","shard":"modules/20709c4534b72170.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.iidTimeVaryingLossStatePredictableLossVector","label":"iidTimeVaryingLossStatePredictableLossVector","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.iidTimeVaryingLossStatePredictableLossVector","description":"A predictable loss vector whose fresh-state evaluator may vary by round.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-79fdcb45a6f9","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10248,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def iidTimeVaryingLossStatePredictableLossVector {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) : Exp3.PredictableLossVector (Nat -> LossState) Action where","missing":[],"search":"iidtimevaryinglossstatepredictablelossvector banditrlproof.tsallis.iidtimevaryinglossstatepredictablelossvector a predictable loss vector whose fresh-state evaluator may vary by round. definition compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.predictableLossAt_iidTimeVaryingLossStatePredictableLossVector","label":"predictableLossAt_iidTimeVaryingLossStatePredictableLossVector","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.predictableLossAt_iidTimeVaryingLossStatePredictableLossVector","description":"theorem predictableLossAt_iidTimeVaryingLossStatePredictableLossVector {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1)…","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-0f8a9378d870","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10249,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem predictableLossAt_iidTimeVaryingLossStatePredictableLossVector {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (t : Nat) (sample : (Nat -> LossState) × ((k : Nat) -> Action × Real)) (action : Action) : Exp3.predictableLossAt (iidTimeVaryingLossStatePredictableLossVector value hvalue hvalue_nonneg hvalue_le_one) t sample action = value t (sample.1 t) action","missing":[],"search":"predictablelossat_iidtimevaryinglossstatepredictablelossvector banditrlproof.tsallis.predictablelossat_iidtimevaryinglossstatepredictablelossvector theorem predictablelossat_iidtimevaryinglossstatepredictablelossvector {lossstate : type u} {action : type v} [measurablespace lossstate] [measurablespace action] (value : nat -> lossstate -> action -> real) (hvalue : ∀ t, measurable (fun input : lossstate × action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (t : nat) (sample : (nat -> lossstate) × ((k : nat) -> action × real)) (action : action) : exp3.predictablelossat (iidtimevaryinglossstatepredictablelossvector value hvalue hvalue_nonneg hvalue_le_one) t sample action = value t (sample.1 t) action theorem compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidTimeVaryingLossState_prefix_eq","label":"sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidTimeVaryingLossState_prefix_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidTimeVaryingLossState_prefix_eq","description":"Equal state prefixes generate equal visible trajectory prefixes for a time-varying fresh-state evaluator.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-877190fec48d","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10250,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:77"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidTimeVaryingLossState_prefix_eq {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (environment₁ environment₂ : Nat -> LossState) (n : Nat) (henvironment : Preorder.frestrictLe n environment₁ = Preorder.frestrictLe n environment₂) : (sampledSche…","missing":[],"search":"sampledscheduledhalftsallistrajectorykernel_map_frestrictle_eq_of_iidtimevaryinglossstate_prefix_eq banditrlproof.tsallis.sampledscheduledhalftsallistrajectorykernel_map_frestrictle_eq_of_iidtimevaryinglossstate_prefix_eq equal state prefixes generate equal visible trajectory prefixes for a time-varying fresh-state evaluator. theorem compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDTimeVaryingPrefixKernel","label":"sampledScheduledHalfTsallisIIDTimeVaryingPrefixKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDTimeVaryingPrefixKernel","description":"Finite visible-prefix kernel for a time-varying fresh-state evaluator.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-3ab49a6f4b3d","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10251,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:141"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisIIDTimeVaryingPrefixKernel {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [Nonempty Action] [DecidableEq Action] (fallback : LossState) (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Kernel ((i : Finset.Iic n) -> LossState) (History.FinitePairHistory Action Real n)","missing":[],"search":"sampledscheduledhalftsallisiidtimevaryingprefixkernel banditrlproof.tsallis.sampledscheduledhalftsallisiidtimevaryingprefixkernel finite visible-prefix kernel for a time-varying fresh-state evaluator. definition compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTimeVaryingTrajectoryKernel","label":"hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTimeVaryingTrajectoryKernel","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTimeVaryingTrajectoryKernel","description":"The canonical trajectory factors through every finite IID state prefix even when the deterministic evaluator varies with time.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-b116a36b3d8e","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10252,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:195"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTimeVaryingTrajectoryKernel {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (fallback : LossState) (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (horizon : Nat) : HasScheduledIIDPrefixKernelFactorization (sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector (iidTimeVarying…","missing":[],"search":"hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallistimevaryingtrajectorykernel banditrlproof.tsallis.hasschedulediidprefixkernelfactorization_sampledscheduledhalftsallistimevaryingtrajectorykernel the canonical trajectory factors through every finite iid state prefix even when the deterministic evaluator varies with time. theorem compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap","label":"independentLossStateTimeVaryingMeanGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap","description":"Mean loss gap of the round-`t` evaluator under its coordinate law.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-6bc646e44d99","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10253,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:267"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def independentLossStateTimeVaryingMeanGap {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] (law : Nat -> Measure LossState) (value : Nat -> LossState -> Action -> Real) (t : Nat) (best action : Action) : Real","missing":[],"search":"independentlossstatetimevaryingmeangap banditrlproof.tsallis.independentlossstatetimevaryingmeangap mean loss gap of the round-`t` evaluator under its coordinate law. definition compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.iidLossStateTimeVaryingMeanGap","label":"iidLossStateTimeVaryingMeanGap","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.iidLossStateTimeVaryingMeanGap","description":"Mean loss gap of the round-`t` evaluator under the common coordinate law.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-1c17abb9d964","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10254,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:276"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def iidLossStateTimeVaryingMeanGap {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] (law : Measure LossState) (value : Nat -> LossState -> Action -> Real) (t : Nat) (best action : Action) : Real","missing":[],"search":"iidlossstatetimevaryingmeangap banditrlproof.tsallis.iidlossstatetimevaryingmeangap mean loss gap of the round-`t` evaluator under the common coordinate law. definition compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingIndependentMeanGapLaw_of_independentLossState","label":"hasScheduledTimeVaryingIndependentMeanGapLaw_of_independentLossState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledTimeVaryingIndependentMeanGapLaw_of_independentLossState","description":"Independent, potentially nonidentically distributed coordinates plus prefix factorization produce the time-varying independence and global-mean contract.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-c78ecce5c6fc","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10255,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:286"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledTimeVaryingIndependentMeanGapLaw_of_independentLossState {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (law : Nat -> Measure LossState) [∀ t, IsProbabilityMeasure (law t)] (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (arms : Finset Action) (best : Action) (horizon : Nat) (trajectoryKernel : Kernel (Nat -> LossState) ((k : Nat) -> Action × Real)) [IsMarkovKernel trajectoryKernel] (hfactor : HasScheduledIIDPrefixKernelFactorization trajectoryKernel horizon) : let prior := Measure.infinit…","missing":[],"search":"hasscheduledtimevaryingindependentmeangaplaw_of_independentlossstate banditrlproof.tsallis.hasscheduledtimevaryingindependentmeangaplaw_of_independentlossstate independent, potentially nonidentically distributed coordinates plus prefix factorization produce the time-varying independence and global-mean contract. theorem compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingIndependentMeanGapLaw_of_iidLossState","label":"hasScheduledTimeVaryingIndependentMeanGapLaw_of_iidLossState","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledTimeVaryingIndependentMeanGapLaw_of_iidLossState","description":"IID coordinates are the constant-law specialization of the independent coordinate producer.","url":"../modules/banditrlproof-tsallisschedulediidtimevaryingmeangap/index.html#decl-c0e1dc36044e","parent":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","order":10256,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap"],["Source","BanditRLProof/TsallisScheduledIIDTimeVaryingMeanGap.lean:464"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledTimeVaryingIndependentMeanGapLaw_of_iidLossState {LossState : Type u} {Action : Type v} [MeasurableSpace LossState] [StandardBorelSpace LossState] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (law : Measure LossState) [IsProbabilityMeasure law] (value : Nat -> LossState -> Action -> Real) (hvalue : ∀ t, Measurable (fun input : LossState × Action => value t input.1 input.2)) (hvalue_nonneg : ∀ t state action, 0 <= value t state action) (hvalue_le_one : ∀ t state action, value t state action <= 1) (arms : Finset Action) (best : Action) (horizon : Nat) (trajectoryKernel : Kernel (Nat -> LossState) ((k : Nat) -> Action × Real)) [IsMarkovKernel trajectoryKernel] (hfactor : HasScheduledIIDPrefixKernelFactorization trajectoryKernel horizon) : let prior := Measure.infinitePi (fun _ : Nat => law)…","missing":[],"search":"hasscheduledtimevaryingindependentmeangaplaw_of_iidlossstate banditrlproof.tsallis.hasscheduledtimevaryingindependentmeangaplaw_of_iidlossstate iid coordinates are the constant-law specialization of the independent coordinate producer. theorem compiled","shard":"modules/fbe3fe5a6b58802d.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledIndependentMeanGapLaw","label":"HasScheduledIndependentMeanGapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledIndependentMeanGapLaw","description":"For every included time and suboptimal arm, the predictable loss difference is independent of the pre-action trace sigma-algebra and has global mean equal to the arm gap.","url":"../modules/banditrlproof-tsallisscheduledindependentmeangap/index.html#decl-934ac4a8f1f0","parent":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","order":10257,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledIndependentMeanGap"],["Source","BanditRLProof/TsallisScheduledIndependentMeanGap.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledIndependentMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Action -> Real) (horizon : Nat) : Prop","missing":[],"search":"hasscheduledindependentmeangaplaw banditrlproof.tsallis.hasscheduledindependentmeangaplaw for every included time and suboptimal arm, the predictable loss difference is independent of the pre-action trace sigma-algebra and has global mean equal to the arm gap. definition compiled","shard":"modules/203d1924167df739.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledConditionalMeanGapLaw_of_independentMeanGapLaw","label":"hasScheduledConditionalMeanGapLaw_of_independentMeanGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledConditionalMeanGapLaw_of_independentMeanGapLaw","description":"Independence from the pre-action trace and the correct global mean imply the scheduled conditional-mean gap law.","url":"../modules/banditrlproof-tsallisscheduledindependentmeangap/index.html#decl-9749063259df","parent":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","order":10258,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIndependentMeanGap"],["Source","BanditRLProof/TsallisScheduledIndependentMeanGap.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledConditionalMeanGapLaw_of_independentMeanGapLaw {Env : Type u} {Action : Type v} [mEnv : MeasurableSpace Env] [mAction : MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Action -> Real) (horizon : Nat) (hgap : HasScheduledIndependentMeanGapLaw mu arms loss best gap horizon) : HasScheduledConditionalMeanGapLaw mu arms loss best gap horizon","missing":[],"search":"hasscheduledconditionalmeangaplaw_of_independentmeangaplaw banditrlproof.tsallis.hasscheduledconditionalmeangaplaw_of_independentmeangaplaw independence from the pre-action trace and the correct global mean imply the scheduled conditional-mean gap law. theorem compiled","shard":"modules/203d1924167df739.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledExpectedGapLaw_of_independentMeanGapLaw","label":"hasScheduledExpectedGapLaw_of_independentMeanGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledExpectedGapLaw_of_independentMeanGapLaw","description":"The independence-plus-mean contract also directly produces the coordinatewise first-moment law used by scheduled self-bounding.","url":"../modules/banditrlproof-tsallisscheduledindependentmeangap/index.html#decl-4f5777600f67","parent":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","order":10259,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIndependentMeanGap"],["Source","BanditRLProof/TsallisScheduledIndependentMeanGap.lean:98"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledExpectedGapLaw_of_independentMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Action -> Real) (horizon : Nat) (hgap : HasScheduledIndependentMeanGapLaw mu arms loss best gap horizon) : HasScheduledExpectedGapLaw mu arms harms eta loss best gap horizon","missing":[],"search":"hasscheduledexpectedgaplaw_of_independentmeangaplaw banditrlproof.tsallis.hasscheduledexpectedgaplaw_of_independentmeangaplaw the independence-plus-mean contract also directly produces the coordinatewise first-moment law used by scheduled self-bounding. theorem compiled","shard":"modules/203d1924167df739.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_independentMeanGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_independentMeanGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_independentMeanGap","description":"Under coordinatewise independence from the pre-action trace and the correct global loss-gap means, the generated square-root schedule satisfies the explicit logarithmic regret bound.","url":"../modules/banditrlproof-tsallisscheduledindependentmeangap/index.html#decl-1359224421a4","parent":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","order":10260,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledIndependentMeanGap"],["Source","BanditRLProof/TsallisScheduledIndependentMeanGap.lean:120"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_independentMeanGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampledScheduledHalfTsallisSqrtSchedule loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule selector.finiteHistory loss.…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_independentmeangap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_independentmeangap under coordinatewise independence from the pre-action trace and the correct global loss-gap means, the generated square-root schedule satisfies the explicit logarithmic regret bound. theorem compiled","shard":"modules/203d1924167df739.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","label":"sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","description":"Canonical time-zero conjugate-potential score on the environment and sampled initial action.","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html#decl-aa1a618653f6","parent":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","order":10261,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledInitialExpectedStability"],["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisInitialHistoryActionPotentialStability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) : Env × Action -> Real","missing":[],"search":"sampledscheduledhalftsallisinitialhistoryactionpotentialstability banditrlproof.tsallis.sampledscheduledhalftsallisinitialhistoryactionpotentialstability canonical time-zero conjugate-potential score on the environment and sampled initial action. definition compiled","shard":"modules/a01c4cca68684c12.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","label":"sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","description":"Refined time-zero budget under the canonical initial half-Tsallis law.","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html#decl-975a6aeaab75","parent":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","order":10262,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledInitialExpectedStability"],["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean:35"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) : Env -> Real","missing":[],"search":"sampledscheduledhalftsallisinitialrefinedpotentialstabilitybound banditrlproof.tsallis.sampledscheduledhalftsallisinitialrefinedpotentialstabilitybound refined time-zero budget under the canonical initial half-tsallis law. definition compiled","shard":"modules/a01c4cca68684c12.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","label":"measurable_sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","description":"The canonical time-zero history/action potential score is measurable.","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html#decl-83dad65974f9","parent":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","order":10263,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledInitialExpectedStability"],["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean:44"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisInitialHistoryActionPotentialStability {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) : Measurable (sampledScheduledHalfTsallisInitialHistoryActionPotentialStability arms harms eta loss)","missing":[],"search":"measurable_sampledscheduledhalftsallisinitialhistoryactionpotentialstability banditrlproof.tsallis.measurable_sampledscheduledhalftsallisinitialhistoryactionpotentialstability the canonical time-zero history/action potential score is measurable. theorem compiled","shard":"modules/a01c4cca68684c12.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","label":"integrable_sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","description":"The time-zero refined budget is integrable under any finite environment law when the initial rate is positive.","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html#decl-2b0ebabbb984","parent":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","order":10264,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledInitialExpectedStability"],["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean:69"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : 0 < eta 0) : Integrable (sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound (Env := Env) arms harms eta) prior","missing":[],"search":"integrable_sampledscheduledhalftsallisinitialrefinedpotentialstabilitybound banditrlproof.tsallis.integrable_sampledscheduledhalftsallisinitialrefinedpotentialstabilitybound the time-zero refined budget is integrable under any finite environment law when the initial rate is positive. theorem compiled","shard":"modules/a01c4cca68684c12.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime_zero_eq_initial_ae","label":"sampledScheduledHalfTsallisPotentialStabilityAtTime_zero_eq_initial_ae","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime_zero_eq_initial_ae","description":"The actual scheduled time-zero potential term agrees almost surely with the canonical environment/action score.","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html#decl-9932e970d3d1","parent":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","order":10265,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledInitialExpectedStability"],["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean:89"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisPotentialStabilityAtTime_zero_eq_initial_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample 0) =ᵐ[mu] (fun sample => sampledScheduledHalfTsallisInitialHistoryActionPotentialStability arms harms eta loss (sample.1, (sample.2 0).1))","missing":[],"search":"sampledscheduledhalftsallispotentialstabilityattime_zero_eq_initial_ae banditrlproof.tsallis.sampledscheduledhalftsallispotentialstabilityattime_zero_eq_initial_ae the actual scheduled time-zero potential term agrees almost surely with the canonical environment/action score. theorem compiled","shard":"modules/a01c4cca68684c12.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_refined","label":"integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_refined","description":"The actual scheduled time-zero potential term agrees almost surely with the canonical environment/action score. -/ theorem sampledScheduledHalfTsallisPotentialStabilityAtTime_zero_eq_initial_ae {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure En…","url":"../modules/banditrlproof-tsallisscheduledinitialexpectedstability/index.html#decl-7adf7d456d1a","parent":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","order":10266,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledInitialExpectedStability"],["Source","BanditRLProof/TsallisScheduledInitialExpectedStability.lean:164"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_refined {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (heta : 0 < eta 0) (heta_le : eta 0 <= 1 / 2) (loss : Exp3.PredictableLossVector Env Action) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment Integrable (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample 0) mu ∧ integral mu (fun sample => sampledScheduledHalfTsallisPotentialStabilityAtTim…","missing":[],"search":"integral_sampledscheduledhalftsallisinitialpotentialstabilityattime_le_refined banditrlproof.tsallis.integral_sampledscheduledhalftsallisinitialpotentialstabilityattime_le_refined the actual scheduled time-zero potential term agrees almost surely with the canonical environment/action score. -/ theorem sampledscheduledhalftsallispotentialstabilityattime_zero_eq_initial_ae {env : type u} {action : type v} [measurablespace env] [standardborelspace env] [measurablespace action] [measurablesingletonclass action] [standardborelspace action] [nonempty action] [decidableeq action] (prior : measure env) [isfinitemeasure prior] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (loss : exp3.predictablelossvector env action) : let selector := canonicalhalftsallisschedulegeneratedselectormeasurability arms harms eta loss let mu := prior ⊗ₘ sampledscheduledhalftsallistrajectorykernel arms harms eta selector.finitehistory loss.environment (fun sample => sampledscheduledhalftsallispotentialstabilityattime arms harms eta sample 0) =ᵐ[mu] (fun sample => sampledscheduledhalftsallisinitialhistoryactionpotentialstability arms harms eta loss (sample.1, (sample.2 0).1)) := by dsimp only let selector := canonicalhalftsallisschedulegeneratedselectormeasurability arms harms eta loss let algorithm := sampledscheduledhalftsallishistoryalgorithm arms harms eta selector.finitehistory let m…","shard":"modules/a01c4cca68684c12.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HalfTsallisScheduleFiniteHistorySelectorMeasurability","label":"HalfTsallisScheduleFiniteHistorySelectorMeasurability","kind":"structure","status":"compiled","subtitle":"BanditRLProof.Tsallis.HalfTsallisScheduleFiniteHistorySelectorMeasurability","description":"Roundwise selector regularity for a deterministic learning-rate schedule.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-346bf71ffa02","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10267,"meta":[["Kind","structure"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"structure HalfTsallisScheduleFiniteHistorySelectorMeasurability {Action : Type u} [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) : Prop where","missing":[],"search":"halftsallisschedulefinitehistoryselectormeasurability banditrlproof.tsallis.halftsallisschedulefinitehistoryselectormeasurability roundwise selector regularity for a deterministic learning-rate schedule. structure compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisScheduleFiniteHistorySelectorMeasurability","label":"canonicalHalfTsallisScheduleFiniteHistorySelectorMeasurability","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.canonicalHalfTsallisScheduleFiniteHistorySelectorMeasurability","description":"The canonical selected minimizer satisfies the scheduled regularity contract at every round.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-4a26073fb1de","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10268,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def canonicalHalfTsallisScheduleFiniteHistorySelectorMeasurability {Action : Type u} [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta where","missing":[],"search":"canonicalhalftsallisschedulefinitehistoryselectormeasurability banditrlproof.tsallis.canonicalhalftsallisschedulefinitehistoryselectormeasurability the canonical selected minimizer satisfies the scheduled regularity contract at every round. definition compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore","label":"sampledScheduledHalfTsallisHistoryScore","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore","description":"Cumulative importance-weighted score through an inclusive history under the scheduled policies.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-8f80d7ed1e27","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10269,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:42"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisHistoryScore {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) : (n : Nat) -> History.FinitePairHistory Action Real n -> Action -> Real | 0, history, action => Exp3.importanceWeightedLoss (initialHalfTsallisDistribution arms harms (eta 0)) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action | n + 1, history, action => let previous := Exp3.previousPairHistory history sampledScheduledHalfTsallisHistoryScore arms harms eta n previous action + Exp3.importanceWeightedLoss (halfTsallisMinimizer arms harms (eta (n + 1)) (sampledScheduledHalfTsallisHistoryScore arms harms eta n previous)) (fun _ => (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 action @[simp] theorem sampledScheduledHalfTsallisHis…","missing":[],"search":"sampledscheduledhalftsallishistoryscore banditrlproof.tsallis.sampledscheduledhalftsallishistoryscore cumulative importance-weighted score through an inclusive history under the scheduled policies. definition compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore_zero","label":"sampledScheduledHalfTsallisHistoryScore_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore_zero","description":"theorem sampledScheduledHalfTsallisHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (history : History.FinitePairHistory Action Real 0) (action : Action) : sampledScheduledHalfTsallisHistoryScore arms harms eta 0 history action = Exp3.importanceWeightedLoss (initialHalfTsallisDistribution arms harms (eta 0)) (fun _ => (history ⟨0, Finset.mem_…","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-974b0317c9f1","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10270,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:62"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisHistoryScore_zero {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (history : History.FinitePairHistory Action Real 0) (action : Action) : sampledScheduledHalfTsallisHistoryScore arms harms eta 0 history action = Exp3.importanceWeightedLoss (initialHalfTsallisDistribution arms harms (eta 0)) (fun _ => (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨0, Finset.mem_Iic.mpr le_rfl⟩).1 action","missing":[],"search":"sampledscheduledhalftsallishistoryscore_zero banditrlproof.tsallis.sampledscheduledhalftsallishistoryscore_zero theorem sampledscheduledhalftsallishistoryscore_zero {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (history : history.finitepairhistory action real 0) (action : action) : sampledscheduledhalftsallishistoryscore arms harms eta 0 history action = exp3.importanceweightedloss (initialhalftsallisdistribution arms harms (eta 0)) (fun _ => (history ⟨0, finset.mem_iic.mpr le_rfl⟩).2) (history ⟨0, finset.mem_iic.mpr le_rfl⟩).1 action theorem compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore_succ","label":"sampledScheduledHalfTsallisHistoryScore_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore_succ","description":"theorem sampledScheduledHalfTsallisHistoryScore_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (action : Action) : sampledScheduledHalfTsallisHistoryScore arms harms eta (n + 1) history action = sampledScheduledHalfTsallisHistoryScore arms harms eta n (Exp3.previousPairHistory history)…","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-3316586097f8","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10271,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:74"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisHistoryScore_succ {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) (history : History.FinitePairHistory Action Real (n + 1)) (action : Action) : sampledScheduledHalfTsallisHistoryScore arms harms eta (n + 1) history action = sampledScheduledHalfTsallisHistoryScore arms harms eta n (Exp3.previousPairHistory history) action + Exp3.importanceWeightedLoss (halfTsallisMinimizer arms harms (eta (n + 1)) (sampledScheduledHalfTsallisHistoryScore arms harms eta n (Exp3.previousPairHistory history))) (fun _ => (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).2) (history ⟨n + 1, Finset.mem_Iic.mpr le_rfl⟩).1 action","missing":[],"search":"sampledscheduledhalftsallishistoryscore_succ banditrlproof.tsallis.sampledscheduledhalftsallishistoryscore_succ theorem sampledscheduledhalftsallishistoryscore_succ {action : type u} [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (n : nat) (history : history.finitepairhistory action real (n + 1)) (action : action) : sampledscheduledhalftsallishistoryscore arms harms eta (n + 1) history action = sampledscheduledhalftsallishistoryscore arms harms eta n (exp3.previouspairhistory history) action + exp3.importanceweightedloss (halftsallisminimizer arms harms (eta (n + 1)) (sampledscheduledhalftsallishistoryscore arms harms eta n (exp3.previouspairhistory history))) (fun _ => (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).2) (history ⟨n + 1, finset.mem_iic.mpr le_rfl⟩).1 action theorem compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisHistoryScore","label":"measurable_sampledScheduledHalfTsallisHistoryScore","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisHistoryScore","description":"Scheduled selector regularity makes every supported score coordinate measurable.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-d64f48ffe843","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10272,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:93"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem measurable_sampledScheduledHalfTsallisHistoryScore {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) : forall n action, action ∈ arms -> Measurable (fun history : History.FinitePairHistory Action Real n => sampledScheduledHalfTsallisHistoryScore arms harms eta n history action)","missing":[],"search":"measurable_sampledscheduledhalftsallishistoryscore banditrlproof.tsallis.measurable_sampledscheduledhalftsallishistoryscore scheduled selector regularity makes every supported score coordinate measurable. theorem compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryDistribution","label":"sampledScheduledHalfTsallisHistoryDistribution","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryDistribution","description":"Scheduled successor probabilities after the visible prefix through `n`.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-62b65eefa1aa","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10273,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:173"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisHistoryDistribution {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (n : Nat) : History.FinitePairHistory Action Real n -> Action -> Real","missing":[],"search":"sampledscheduledhalftsallishistorydistribution banditrlproof.tsallis.sampledscheduledhalftsallishistorydistribution scheduled successor probabilities after the visible prefix through `n`. definition compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryDistributionSource","label":"sampledScheduledHalfTsallisHistoryDistributionSource","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryDistributionSource","description":"Measurable finite-action source for one scheduled successor policy.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-3d10dbc2abaf","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10274,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:182"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisHistoryDistributionSource {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : Exp3.MeasurableFiniteActionDistribution arms (sampledScheduledHalfTsallisHistoryDistribution arms harms eta n)","missing":[],"search":"sampledscheduledhalftsallishistorydistributionsource banditrlproof.tsallis.sampledscheduledhalftsallishistorydistributionsource measurable finite-action source for one scheduled successor policy. definition compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAlgorithm","label":"sampledScheduledHalfTsallisHistoryAlgorithm","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAlgorithm","description":"Stochastic finite-history algorithm generated by the scheduled recursive score.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-1cd99b247d6b","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10275,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:205"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisHistoryAlgorithm {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) : Thompson.HistoryAlgorithm Action Real where","missing":[],"search":"sampledscheduledhalftsallishistoryalgorithm banditrlproof.tsallis.sampledscheduledhalftsallishistoryalgorithm stochastic finite-history algorithm generated by the scheduled recursive score. definition compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAlgorithm_policy","label":"sampledScheduledHalfTsallisHistoryAlgorithm_policy","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAlgorithm_policy","description":"theorem sampledScheduledHalfTsallisHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : (sampledScheduledHalfTsallisHistoryAlgorithm arms harms eta selector).policy n = Exp3.finiteActionKer…","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-c52ae4b87a4d","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10276,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:229"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisHistoryAlgorithm_policy {Action : Type u} [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (n : Nat) : (sampledScheduledHalfTsallisHistoryAlgorithm arms harms eta selector).policy n = Exp3.finiteActionKernel arms (sampledScheduledHalfTsallisHistoryDistribution arms harms eta n) (sampledScheduledHalfTsallisHistoryDistributionSource arms harms eta selector n)","missing":[],"search":"sampledscheduledhalftsallishistoryalgorithm_policy banditrlproof.tsallis.sampledscheduledhalftsallishistoryalgorithm_policy theorem sampledscheduledhalftsallishistoryalgorithm_policy {action : type u} [measurablespace action] [measurablesingletonclass action] [decidableeq action] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (selector : halftsallisschedulefinitehistoryselectormeasurability arms harms eta) (n : nat) : (sampledscheduledhalftsallishistoryalgorithm arms harms eta selector).policy n = exp3.finiteactionkernel arms (sampledscheduledhalftsallishistorydistribution arms harms eta n) (sampledscheduledhalftsallishistorydistributionsource arms harms eta selector n) theorem compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel","label":"sampledScheduledHalfTsallisTrajectoryKernel","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel","description":"Complete environment-indexed scheduled half-Tsallis trajectory.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-8791deff4975","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10277,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:245"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisTrajectoryKernel {Env : Type v} {Action : Type u} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] [Nonempty Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) : Kernel Env ((n : Nat) -> Action × Real)","missing":[],"search":"sampledscheduledhalftsallistrajectorykernel banditrlproof.tsallis.sampledscheduledhalftsallistrajectorykernel complete environment-indexed scheduled half-tsallis trajectory. definition compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action","label":"sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action","description":"The successor action has its scheduled finite-action law conditional on the visible pair-history prefix.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-f0e06ba25939","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10278,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:276"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample => Preorder.frestrictLe n sample.2) (prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector environment) =ᵐ[ (prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector environment).map (fun sample => Preor…","missing":[],"search":"sampledscheduledhalftsallistrajectorymeasure_conddistrib_action banditrlproof.tsallis.sampledscheduledhalftsallistrajectorymeasure_conddistrib_action the successor action has its scheduled finite-action law conditional on the visible pair-history prefix. theorem compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","label":"sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","description":"The scheduled conditional law after retaining the environment in the visible history.","url":"../modules/banditrlproof-tsallisscheduledrecursivetrajectory/index.html#decl-4c2478f39c1e","parent":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","order":10279,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRecursiveTrajectory"],["Source","BanditRLProof/TsallisScheduledRecursiveTrajectory.lean:310"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment {Env : Type v} {Action : Type u} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsFiniteMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (selector : HalfTsallisScheduleFiniteHistorySelectorMeasurability arms harms eta) (environment : Thompson.MeasurableHistoryEnvironment Env Action Real) (n : Nat) : condDistrib (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.2 (n + 1)).1) (fun sample : Env × ((k : Nat) -> Action × Real) => (sample.1, Preorder.frestrictLe n sample.2)) (prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector environment) =ᵐ[ (prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryK…","missing":[],"search":"sampledscheduledhalftsallistrajectorymeasure_conddistrib_action_given_environment banditrlproof.tsallis.sampledscheduledhalftsallistrajectorymeasure_conddistrib_action_given_environment the scheduled conditional law after retaining the environment in the visible history. theorem compiled","shard":"modules/8634a6ad55dda625.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw_of_expectedDeviation","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw_of_expectedDeviation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw_of_expectedDeviation","description":"A reference expected-gap law plus an integrable sample-dependent predictable perturbation yields a self-bound with the exact expected weighted deviation allowance.","url":"../modules/banditrlproof-tsallisscheduledreferencegapexpecteddeviationselfbounding/index.html#decl-5233c6c5e7e9","parent":"module:BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","order":10280,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding"],["Source","BanditRLProof/TsallisScheduledReferenceGapExpectedDeviationSelfBounding.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw_of_expectedDeviation {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (actualLoss referenceLoss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (baseGap : Action -> Real) (deviation : Nat -> (Env × ((k : Nat) -> Action × Real)) -> Action -> Real) (horizon : Nat) (hreferenceGapLaw : HasScheduledExpectedGapLaw mu arms harms eta referenceLoss best baseGap horizon) (hdeviation_integrable : forall t, t <= horizon -> forall action, action ∈ arms.erase best -> Integrable (fun sample => sampledScheduledHalfTsallisProbabilit…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_referenceexpectedgaplaw_of_expecteddeviation banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_referenceexpectedgaplaw_of_expecteddeviation a reference expected-gap law plus an integrable sample-dependent predictable perturbation yields a self-bound with the exact expected weighted deviation allowance. theorem compiled","shard":"modules/eca4fc0ad7196908.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw","description":"A reference expected-gap law plus a pointwise predictable perturbation bound yields the fixed-baseline self-bound for the actual generated regret.","url":"../modules/banditrlproof-tsallisscheduledreferencegapselfbounding/index.html#decl-ccf9b2538625","parent":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","order":10281,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledReferenceGapSelfBounding"],["Source","BanditRLProof/TsallisScheduledReferenceGapSelfBounding.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (actualLoss referenceLoss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (baseGap : Action -> Real) (deviation : Nat -> Action -> Real) (horizon : Nat) (hreferenceGapLaw : HasScheduledExpectedGapLaw mu arms harms eta referenceLoss best baseGap horizon) (hdeviation : ∀ t, t <= horizon -> ∀ sample action, action ∈ arms.erase best -> |(Exp3.predictableLossAt actualLoss t sample action - Exp3.predictableLossAt actualLoss t sample best) - (Exp3.predictableLossAt reference…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_referenceexpectedgaplaw banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_referenceexpectedgaplaw a reference expected-gap law plus a pointwise predictable perturbation bound yields the fixed-baseline self-bound for the actual generated regret. theorem compiled","shard":"modules/963e29109b7e5e4c.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_sub_self_le_half_one_sub","label":"sqrt_sub_self_le_half_one_sub","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_sub_self_le_half_one_sub","description":"The concavity tangent at one bounds the excess of `sqrt x` over `x`.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-56d4b7e18cad","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10282,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sqrt_sub_self_le_half_one_sub (x : Real) (hx : 0 <= x) : Real.sqrt x - x <= (1 - x) / 2","missing":[],"search":"sqrt_sub_self_le_half_one_sub banditrlproof.tsallis.sqrt_sub_self_le_half_one_sub the concavity tangent at one bounds the excess of `sqrt x` over `x`. theorem compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_sub_one_le_two_mul_sum_erase_refined","label":"halfTsallisPotentialMass_sub_one_le_two_mul_sum_erase_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialMass_sub_one_le_two_mul_sum_erase_refined","description":"Above the point-mass baseline, half-Tsallis potential mass is controlled by the paper's refined suboptimal-arm mass.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-b8d68b973824","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10283,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:29"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialMass_sub_one_le_two_mul_sum_erase_refined {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (probability : Action -> Real) (hprobability : FTRL.finiteSimplex arms probability) : halfTsallisPotentialMass arms probability - 1 <= 2 * (arms.erase best).sum (fun action => Real.sqrt (probability action) - probability action / 2)","missing":[],"search":"halftsallispotentialmass_sub_one_le_two_mul_sum_erase_refined banditrlproof.tsallis.halftsallispotentialmass_sub_one_le_two_mul_sum_erase_refined above the point-mass baseline, half-tsallis potential mass is controlled by the paper's refined suboptimal-arm mass. theorem compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le_refined","label":"sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le_refined","description":"Generated pathwise point-mass penalty with every reciprocal-rate increment retained and every potential mass replaced by its refined suboptimal-arm upper bound.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-7c1b0e10726c","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10284,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:57"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le_refined {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) {best : Action} (hbest : best ∈ arms) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hetaMono : forall t, t < n -> eta (t + 1) <= eta t) : (Finset.range (n + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialPenaltyAtTime arms harms eta (pointMass best) sample t) <= 2 / eta 0 * (arms.erase best).sum (fun action => Real.sqrt (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta 0 sample action) - sampledScheduledHalfTsallisProbabilityAtTime arms harms eta 0 sample action / 2) + (Finset.range n).sum (fun t => 2 * (1 / eta (t + 1) - 1 / eta t) * (arms.erase best).sum (fun action => Real.sqrt (sampledScheduledHalfTsallisProbabilityAtTime…","missing":[],"search":"sum_sampledscheduledhalftsallispotentialpenalty_pointmass_le_refined banditrlproof.tsallis.sum_sampledscheduledhalftsallispotentialpenalty_pointmass_le_refined generated pathwise point-mass penalty with every reciprocal-rate increment retained and every potential mass replaced by its refined suboptimal-arm upper bound. theorem compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedSuboptimalMassAt","label":"sampledScheduledHalfTsallisRefinedSuboptimalMassAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedSuboptimalMassAt","description":"Refined suboptimal-arm mass on one generated trajectory sample.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-65cf7be583b6","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10285,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:196"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisRefinedSuboptimalMassAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (best : Action) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallisrefinedsuboptimalmassat banditrlproof.tsallis.sampledscheduledhalftsallisrefinedsuboptimalmassat refined suboptimal-arm mass on one generated trajectory sample. definition compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt","label":"sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt","description":"Deterministic Jensen target for one refined suboptimal-arm mass.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-7be8f57e6882","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10286,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:208"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (best : Action) (t : Nat) : Real","missing":[],"search":"sampledscheduledhalftsallisexpectedrefinedsuboptimalmassat banditrlproof.tsallis.sampledscheduledhalftsallisexpectedrefinedsuboptimalmassat deterministic jensen target for one refined suboptimal-arm mass. definition compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisRefinedSuboptimalMassAt","label":"integrable_sampledScheduledHalfTsallisRefinedSuboptimalMassAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisRefinedSuboptimalMassAt","description":"theorem integrable_sampledScheduledHalfTsallisRefinedSuboptimalMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) {best : Action} (t : Nat) : Integrable (sampledScheduledHalfTsallisRefined…","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-05f79a7f05e3","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10287,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:220"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisRefinedSuboptimalMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) {best : Action} (t : Nat) : Integrable (sampledScheduledHalfTsallisRefinedSuboptimalMassAt arms harms eta best t) mu","missing":[],"search":"integrable_sampledscheduledhalftsallisrefinedsuboptimalmassat banditrlproof.tsallis.integrable_sampledscheduledhalftsallisrefinedsuboptimalmassat theorem integrable_sampledscheduledhalftsallisrefinedsuboptimalmassat {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) -> action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) {best : action} (t : nat) : integrable (sampledscheduledhalftsallisrefinedsuboptimalmassat arms harms eta best t) mu theorem compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisRefinedSuboptimalMassAt_le_expected","label":"integral_sampledScheduledHalfTsallisRefinedSuboptimalMassAt_le_expected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisRefinedSuboptimalMassAt_le_expected","description":"Jensen transport for the complete refined suboptimal-arm mass at one scheduled time.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-d137a6e52907","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10288,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:247"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisRefinedSuboptimalMassAt_le_expected {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) {best : Action} (t : Nat) : integral mu (sampledScheduledHalfTsallisRefinedSuboptimalMassAt arms harms eta best t) <= sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt mu arms harms eta best t","missing":[],"search":"integral_sampledscheduledhalftsallisrefinedsuboptimalmassat_le_expected banditrlproof.tsallis.integral_sampledscheduledhalftsallisrefinedsuboptimalmassat_le_expected jensen transport for the complete refined suboptimal-arm mass at one scheduled time. theorem compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedExpectedPenalty","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedExpectedPenalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedExpectedPenalty","description":"The generated predictable environment regret retains the refined time-varying penalty after expectation. Unlike the coarse endpoint theorem, no terminal potential mass divided by the final learning rate remains.","url":"../modules/banditrlproof-tsallisscheduledrefinedexpectedpenalty/index.html#decl-610a61c63410","parent":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","order":10289,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedExpectedPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedExpectedPenalty.lean:292"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedExpectedPenalty {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (sampledScheduledHalfTsallisPredic…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedexpectedpenalty banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedexpectedpenalty the generated predictable environment regret retains the refined time-varying penalty after expectation. unlike the coarse endpoint theorem, no terminal potential mass divided by the final learning rate remains. theorem compiled","shard":"modules/f4fa09eb862ed7a2.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedSuboptimalSqrtMassAt","label":"sampledScheduledHalfTsallisExpectedSuboptimalSqrtMassAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedSuboptimalSqrtMassAt","description":"Sum of square roots of expected probabilities over suboptimal arms.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-e0ffe99f58cb","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10290,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisExpectedSuboptimalSqrtMassAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (best : Action) (t : Nat) : Real","missing":[],"search":"sampledscheduledhalftsallisexpectedsuboptimalsqrtmassat banditrlproof.tsallis.sampledscheduledhalftsallisexpectedsuboptimalsqrtmassat sum of square roots of expected probabilities over suboptimal arms. definition compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt_le_sqrtMass","label":"sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt_le_sqrtMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt_le_sqrtMass","description":"Dropping the nonpositive linear correction weakens the expected refined mass to the square-root mass used by completion of squares.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-17382391d0ee","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10291,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt_le_sqrtMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) {best : Action} (t : Nat) : sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt mu arms harms eta best t <= sampledScheduledHalfTsallisExpectedSuboptimalSqrtMassAt mu arms harms eta best t","missing":[],"search":"sampledscheduledhalftsallisexpectedrefinedsuboptimalmassat_le_sqrtmass banditrlproof.tsallis.sampledscheduledhalftsallisexpectedrefinedsuboptimalmassat_le_sqrtmass dropping the nonpositive linear correction weakens the expected refined mass to the square-root mass used by completion of squares. theorem compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient","label":"sampledScheduledHalfTsallisRefinedCoefficient","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient","description":"Combined coefficient multiplying `sqrt (E[p_t(a)])` after adding refined stability and reciprocal-rate penalty terms.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-c01ee8b27bcc","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10292,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:58"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisRefinedCoefficient (eta : Nat -> Real) : Nat -> Real | 0 => 2 * eta 0 + 2 / eta 0 | t + 1 => 2 * eta (t + 1) + 2 * (1 / eta (t + 1) - 1 / eta t) /-- Generated predictable environment regret is bounded by the deterministic quadratic-rate budget plus one refined square-root mass per actual time. -/ theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <…","missing":[],"search":"sampledscheduledhalftsallisrefinedcoefficient banditrlproof.tsallis.sampledscheduledhalftsallisrefinedcoefficient combined coefficient multiplying `sqrt (e[p_t(a)])` after adding refined stability and reciprocal-rate penalty terms. definition compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty","description":"Generated predictable environment regret is bounded by the deterministic quadratic-rate budget plus one refined square-root mass per actual time.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-6c2edb740080","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10293,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:66"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.envi…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty generated predictable environment regret is bounded by the deterministic quadratic-rate budget plus one refined square-root mass per actual time. theorem compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_selfBounding","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_selfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_selfBounding","description":"The combined refined scheduled upper bound consumes any matching expected self-bounding law and yields a squared-coefficient-over-gap theorem.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-d1d1f9e63102","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10294,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:220"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_selfBounding {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) (gap : Action -> Real) (hgapPos : forall action, action ∈ arms.erase best -> 0 < gap action) (corruption : Real) (hselfBounding : let selector := canonicalHalfTsallisScheduleGeneratedSelector…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty_of_selfbounding banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty_of_selfbounding the combined refined scheduled upper bound consumes any matching expected self-bounding law and yields a squared-coefficient-over-gap theorem. theorem compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_fixedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_fixedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_fixedGap","description":"Under an exact predictable fixed-gap law, the combined refined scheduled upper bound automatically yields a squared-coefficient-over-gap theorem.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-8b0b6ace33ef","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10295,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:316"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_fixedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) (gap : Action -> Real) (hgapPos : forall action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.pred…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty_of_fixedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty_of_fixedgap under an exact predictable fixed-gap law, the combined refined scheduled upper bound automatically yields a squared-coefficient-over-gap theorem. theorem compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_expectedGapLaw","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_expectedGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_expectedGapLaw","description":"A coordinatewise expected-gap law supplies the refined scheduled squared-coefficient-over-gap theorem without a samplewise fixed-gap premise.","url":"../modules/banditrlproof-tsallisscheduledrefinedstabilitypenalty/index.html#decl-10fa336c8cb1","parent":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","order":10296,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledRefinedStabilityPenalty"],["Source","BanditRLProof/TsallisScheduledRefinedStabilityPenalty.lean:357"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_expectedGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) (gap : Action -> Real) (hgapPos : forall action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty_of_expectedgaplaw banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedstabilitypenalty_of_expectedgaplaw a coordinatewise expected-gap law supplies the refined scheduled squared-coefficient-over-gap theorem without a samplewise fixed-gap premise. theorem compiled","shard":"modules/46556ada78c750ef.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisProbabilityAtTime","label":"sampledScheduledHalfTsallisProbabilityAtTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisProbabilityAtTime","description":"The pure scheduled half-Tsallis sampling probability at an actual time.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-d0e81f5f1e1b","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10297,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) : Nat -> Env × ((k : Nat) -> Action × Real) -> Action -> Real | 0, _sample => initialHalfTsallisDistribution arms harms (eta 0) | n + 1, sample => sampledScheduledHalfTsallisHistoryDistribution arms harms eta n (Preorder.frestrictLe n sample.2) /-- The actual observed-scalar IW loss vector under the scheduled policy. -/ noncomputable def sampledScheduledHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledscheduledhalftsallisprobabilityattime banditrlproof.tsallis.sampledscheduledhalftsallisprobabilityattime the pure scheduled half-tsallis sampling probability at an actual time. definition compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisObservedEstimatedLossAt","label":"sampledScheduledHalfTsallisObservedEstimatedLossAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisObservedEstimatedLossAt","description":"The actual observed-scalar IW loss vector under the scheduled policy.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-dbdf1cfbc197","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10298,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisObservedEstimatedLossAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Action -> Real","missing":[],"search":"sampledscheduledhalftsallisobservedestimatedlossat banditrlproof.tsallis.sampledscheduledhalftsallisobservedestimatedlossat the actual observed-scalar iw loss vector under the scheduled policy. definition compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.cumulativeLoss_sampledScheduledHalfTsallisObservedEstimatedLossAt_succ","label":"cumulativeLoss_sampledScheduledHalfTsallisObservedEstimatedLossAt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.cumulativeLoss_sampledScheduledHalfTsallisObservedEstimatedLossAt_succ","description":"The cumulative observed IW loss through time `n` is exactly the inclusive scheduled finite-history score at level `n`.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-af000626cad7","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10299,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:44"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem cumulativeLoss_sampledScheduledHalfTsallisObservedEstimatedLossAt_succ {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) (n : Nat) : FTRL.cumulativeLoss (fun t => sampledScheduledHalfTsallisObservedEstimatedLossAt arms harms eta t sample) (n + 1) = sampledScheduledHalfTsallisHistoryScore arms harms eta n (Preorder.frestrictLe n sample.2)","missing":[],"search":"cumulativeloss_sampledscheduledhalftsallisobservedestimatedlossat_succ banditrlproof.tsallis.cumulativeloss_sampledscheduledhalftsallisobservedestimatedlossat_succ the cumulative observed iw loss through time `n` is exactly the inclusive scheduled finite-history score at level `n`. theorem compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisScheduledMinimizer_observedEstimatedLoss_eq_probabilityAtTime","label":"halfTsallisScheduledMinimizer_observedEstimatedLoss_eq_probabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisScheduledMinimizer_observedEstimatedLoss_eq_probabilityAtTime","description":"The canonical scheduled cumulative selector is the actual generated sampling probability at the same time.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-d55577f1edcf","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10300,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:75"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisScheduledMinimizer_observedEstimatedLoss_eq_probabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) : halfTsallisScheduledMinimizer arms harms eta (fun s => sampledScheduledHalfTsallisObservedEstimatedLossAt arms harms eta s sample) t = sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample","missing":[],"search":"halftsallisscheduledminimizer_observedestimatedloss_eq_probabilityattime banditrlproof.tsallis.halftsallisscheduledminimizer_observedestimatedloss_eq_probabilityattime the canonical scheduled cumulative selector is the actual generated sampling probability at the same time. theorem compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSameRateNextAt","label":"sampledScheduledHalfTsallisSameRateNextAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSameRateNextAt","description":"The same-rate auxiliary minimizer after appending the actual observed IW loss at time `t`.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-df108bcf4315","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10301,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:99"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisSameRateNextAt {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) : Action -> Real","missing":[],"search":"sampledscheduledhalftsallissameratenextat banditrlproof.tsallis.sampledscheduledhalftsallissameratenextat the same-rate auxiliary minimizer after appending the actual observed iw loss at time `t`. definition compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime","label":"sampledScheduledHalfTsallisPotentialStabilityAtTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime","description":"The same-rate conjugate-potential stability term at an actual trajectory time.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-d79b05a9bf3e","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10302,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:110"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPotentialStabilityAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) : Real","missing":[],"search":"sampledscheduledhalftsallispotentialstabilityattime banditrlproof.tsallis.sampledscheduledhalftsallispotentialstabilityattime the same-rate conjugate-potential stability term at an actual trajectory time. definition compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialPenaltyAtTime","label":"sampledScheduledHalfTsallisPotentialPenaltyAtTime","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialPenaltyAtTime","description":"The scheduled potential-change penalty at an actual trajectory time.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-bdfbcabd5405","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10303,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:126"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisPotentialPenaltyAtTime {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (q : Action -> Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) : Real","missing":[],"search":"sampledscheduledhalftsallispotentialpenaltyattime banditrlproof.tsallis.sampledscheduledhalftsallispotentialpenaltyattime the scheduled potential-change penalty at an actual trajectory time. definition compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret","label":"sampledScheduledHalfTsallisEstimatedRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret","description":"Pathwise estimated regret against a finite-simplex comparator through the inclusive terminal time `n`.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-dcba75ff6a3e","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10304,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:146"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisEstimatedRegret {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (q : Action -> Real) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledscheduledhalftsallisestimatedregret banditrlproof.tsallis.sampledscheduledhalftsallisestimatedregret pathwise estimated regret against a finite-simplex comparator through the inclusive terminal time `n`. definition compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_eq_stability_add_penalty","label":"sampledScheduledHalfTsallisEstimatedRegret_eq_stability_add_penalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_eq_stability_add_penalty","description":"Exact pathwise decomposition of generated scheduled estimated regret into same-rate potential stability plus the learning-rate-change penalty.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-2ba60770536a","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10305,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:163"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisEstimatedRegret_eq_stability_add_penalty {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (q : Action -> Real) (n : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : sampledScheduledHalfTsallisEstimatedRegret arms harms eta q n sample = (Finset.range (n + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample t) + (Finset.range (n + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialPenaltyAtTime arms harms eta q sample t)","missing":[],"search":"sampledscheduledhalftsallisestimatedregret_eq_stability_add_penalty banditrlproof.tsallis.sampledscheduledhalftsallisestimatedregret_eq_stability_add_penalty exact pathwise decomposition of generated scheduled estimated regret into same-rate potential stability plus the learning-rate-change penalty. theorem compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le","label":"sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le","description":"Pathwise generated-trajectory specialization of the scheduled point-mass penalty theorem, retaining the explicit terminal `-1 / eta n` contribution.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-336c842a2dc3","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10306,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:188"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) {best : Action} (hbest : best ∈ arms) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hetaMono : forall t, t < n -> eta (t + 1) <= eta t) : (Finset.range (n + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialPenaltyAtTime arms harms eta (pointMass best) sample t) <= halfTsallisPotentialMass arms (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta 0 sample) / eta n - 1 / eta n","missing":[],"search":"sum_sampledscheduledhalftsallispotentialpenalty_pointmass_le banditrlproof.tsallis.sum_sampledscheduledhalftsallispotentialpenalty_pointmass_le pathwise generated-trajectory specialization of the scheduled point-mass penalty theorem, retaining the explicit terminal `-1 / eta n` contribution. theorem compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_pointMass_le_stability_add_penalty","label":"sampledScheduledHalfTsallisEstimatedRegret_pointMass_le_stability_add_penalty","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_pointMass_le_stability_add_penalty","description":"Generated pathwise best-arm estimated regret is bounded by the scheduled same-rate stability sum plus the explicit initial-minus-terminal penalty.","url":"../modules/banditrlproof-tsallisscheduledscorealignment/index.html#decl-f490bfe55388","parent":"module:BanditRLProof.TsallisScheduledScoreAlignment","order":10307,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledScoreAlignment"],["Source","BanditRLProof/TsallisScheduledScoreAlignment.lean:213"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisEstimatedRegret_pointMass_le_stability_add_penalty {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) {best : Action} (hbest : best ∈ arms) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hetaMono : forall t, t < n -> eta (t + 1) <= eta t) : sampledScheduledHalfTsallisEstimatedRegret arms harms eta (pointMass best) n sample <= (Finset.range (n + 1)).sum (fun t => sampledScheduledHalfTsallisPotentialStabilityAtTime arms harms eta sample t) + halfTsallisPotentialMass arms (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta 0 sample) / eta n - 1 / eta n","missing":[],"search":"sampledscheduledhalftsallisestimatedregret_pointmass_le_stability_add_penalty banditrlproof.tsallis.sampledscheduledhalftsallisestimatedregret_pointmass_le_stability_add_penalty generated pathwise best-arm estimated regret is bounded by the scheduled same-rate stability sum plus the explicit initial-minus-terminal penalty. theorem compiled","shard":"modules/b2f112affca17175.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.regret_le_selfBoundingInterpolation","label":"regret_le_selfBoundingInterpolation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.regret_le_selfBoundingInterpolation","description":"Algebraic lambda interpolation between an upper regret bound and a terminal self-bounding lower estimate.","url":"../modules/banditrlproof-tsallisscheduledselfboundinginterpolation/index.html#decl-e381dfad82c7","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","order":10308,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingInterpolation"],["Source","BanditRLProof/TsallisScheduledSelfBoundingInterpolation.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regret_le_selfBoundingInterpolation (regret upper gapMass corruption lambda : Real) (hlambda : lambda ∈ Set.Icc (0 : Real) 1) (hupper : regret ≤ upper) (hselfBounding : gapMass - corruption ≤ regret) : regret ≤ (1 + lambda) * upper - lambda * gapMass + lambda * corruption","missing":[],"search":"regret_le_selfboundinginterpolation banditrlproof.tsallis.regret_le_selfboundinginterpolation algebraic lambda interpolation between an upper regret bound and a terminal self-bounding lower estimate. theorem compiled","shard":"modules/80ebd323c80dd90b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingInterpolation","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingInterpolation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingInterpolation","description":"The generated scheduled half-Tsallis upper estimate and a terminal self-bounding constraint imply the paper-facing lambda-interpolated bound. The remaining route is a finite-dimensional optimization over the expected action probabilities and `lambda`; kernel, integral, and Jensen obligations have already been discharged here.","url":"../modules/banditrlproof-tsallisscheduledselfboundinginterpolation/index.html#decl-348cf8d16080","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","order":10309,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingInterpolation"],["Source","BanditRLProof/TsallisScheduledSelfBoundingInterpolation.lean:46"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingInterpolation {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : ∀ t, t ≤ horizon → 0 < eta t) (heta_le : ∀ t, t ≤ horizon → eta t ≤ 1 / 2) (hetaMono : ∀ t, t < horizon → eta (t + 1) ≤ eta t) (gap : Action → Real) (corruption lambda : Real) (hlambda : lambda ∈ Set.Icc (0 : Real) 1) (hselfBounding : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sample…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_selfboundinginterpolation banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_selfboundinginterpolation the generated scheduled half-tsallis upper estimate and a terminal self-bounding constraint imply the paper-facing lambda-interpolated bound. the remaining route is a finite-dimensional optimization over the expected action probabilities and `lambda`; kernel, integral, and jensen obligations have already been discharged here. theorem compiled","shard":"modules/80ebd323c80dd90b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisExpectedProbability_le_quadraticFilterSplit","label":"sum_sampledScheduledHalfTsallisExpectedProbability_le_quadraticFilterSplit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisExpectedProbability_le_quadraticFilterSplit","description":"Split a finite collection of generated expected-probability quadratic terms at the exact active-mass threshold. The active times use the simplex mass constraint; all other times use coordinatewise completion of squares.","url":"../modules/banditrlproof-tsallisscheduledselfboundingoptimization/index.html#decl-a4ca4ce8422b","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","order":10310,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingOptimization"],["Source","BanditRLProof/TsallisScheduledSelfBoundingOptimization.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_sampledScheduledHalfTsallisExpectedProbability_le_quadraticFilterSplit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (times : Finset Nat) {best : Action} (hbest : best ∈ arms) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (b : Nat → Real) (lambda : Real) (hlambda : 0 < lambda) : times.sum (fun t => (arms.erase best).sum (fun action => b t * Real.sqrt (sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action) - lambda * gap action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action)) ≤ (times.filter fun t => 2 * Real.sqrt ((arms.erase…","missing":[],"search":"sum_sampledscheduledhalftsallisexpectedprobability_le_quadraticfiltersplit banditrlproof.tsallis.sum_sampledscheduledhalftsallisexpectedprobability_le_quadraticfiltersplit split a finite collection of generated expected-probability quadratic terms at the exact active-mass threshold. the active times use the simplex mass constraint; all other times use coordinatewise completion of squares. theorem compiled","shard":"modules/348c2c11d8f79600.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_range_sampledScheduledHalfTsallisExpectedProbability_le_quadraticPrefixSplit","label":"sum_range_sampledScheduledHalfTsallisExpectedProbability_le_quadraticPrefixSplit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_range_sampledScheduledHalfTsallisExpectedProbability_le_quadraticPrefixSplit","description":"Prefix/suffix form of the one-round split. A caller supplies a cutoff and proves the active threshold only on its prefix; the suffix always admits the unconstrained branch.","url":"../modules/banditrlproof-tsallisscheduledselfboundingoptimization/index.html#decl-13e850aa3721","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","order":10311,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingOptimization"],["Source","BanditRLProof/TsallisScheduledSelfBoundingOptimization.lean:116"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_range_sampledScheduledHalfTsallisExpectedProbability_le_quadraticPrefixSplit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) → Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (n cutoff : Nat) (hcutoff : cutoff ≤ n) {best : Action} (hbest : best ∈ arms) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (b : Nat → Real) (lambda : Real) (hlambda : 0 < lambda) (hthreshold : ∀ t < cutoff, 2 * Real.sqrt ((arms.erase best).card : Real) ≤ b t * (arms.erase best).sum (fun action => 1 / (lambda * gap action))) : (Finset.range n).sum (fun t => (arms.erase best).sum (fun action => b t * Real.sqrt (sampledScheduledHalfTsallisExpectedProbabilityA…","missing":[],"search":"sum_range_sampledscheduledhalftsallisexpectedprobability_le_quadraticprefixsplit banditrlproof.tsallis.sum_range_sampledscheduledhalftsallisexpectedprobability_le_quadraticprefixsplit prefix/suffix form of the one-round split. a caller supplies a cutoff and proves the active threshold only on its prefix; the suffix always admits the unconstrained branch. theorem compiled","shard":"modules/348c2c11d8f79600.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingQuadraticSplit","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingQuadraticSplit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingQuadraticSplit","description":"The generated scheduled regret theorem after terminal self-bounding and the exact finite-time quadratic branch split. All probabilistic, conditional law, Jensen, and finite-simplex obligations have been discharged; the remaining right-hand side is deterministic schedule algebra.","url":"../modules/banditrlproof-tsallisscheduledselfboundingoptimization/index.html#decl-f187cc0d4e0d","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","order":10312,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingOptimization"],["Source","BanditRLProof/TsallisScheduledSelfBoundingOptimization.lean:194"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingQuadraticSplit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : ∀ t, t ≤ horizon → 0 < eta t) (heta_le : ∀ t, t ≤ horizon → eta t ≤ 1 / 2) (hetaMono : ∀ t, t < horizon → eta (t + 1) ≤ eta t) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (hselfBounding : let selector := canonica…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_selfboundingquadraticsplit banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_selfboundingquadraticsplit the generated scheduled regret theorem after terminal self-bounding and the exact finite-time quadratic branch split. all probabilistic, conditional law, jensen, and finite-simplex obligations have been discharged; the remaining right-hand side is deterministic schedule algebra. theorem compiled","shard":"modules/348c2c11d8f79600.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSum_of_refinedCoefficient_le","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSum_of_refinedCoefficient_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSum_of_refinedCoefficient_le","description":"Refined generated interpolation with any deterministic coefficient envelope `b`. This is the common pre-split surface for filter and prefix quadratic consumers.","url":"../modules/banditrlproof-tsallisscheduledselfboundingoptimization/index.html#decl-5bd2093ad331","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","order":10313,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingOptimization"],["Source","BanditRLProof/TsallisScheduledSelfBoundingOptimization.lean:363"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSum_of_refinedCoefficient_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : ∀ t, t ≤ horizon → 0 < eta t) (heta_le : ∀ t, t ≤ horizon → eta t ≤ 1 / 2) (hetaMono : ∀ t, t < horizon → eta (t + 1) ≤ eta t) (gap : Action → Real) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (b : Nat → Real) (hb : ∀ t, t ≤ horizon → (1 + lambda) * sampledScheduledHalfTsallisRefinedCoefficient eta…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedselfboundingquadraticsum_of_refinedcoefficient_le banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedselfboundingquadraticsum_of_refinedcoefficient_le refined generated interpolation with any deterministic coefficient envelope `b`. this is the common pre-split surface for filter and prefix quadratic consumers. theorem compiled","shard":"modules/348c2c11d8f79600.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSplit_of_refinedCoefficient_le","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSplit_of_refinedCoefficient_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSplit_of_refinedCoefficient_le","description":"Refined generated quadratic split for the theorem route. Unlike the coarse scheduled interpolation above, this theorem consumes the uncollapsed stability-penalty coefficient, so its deterministic base contains only `sum_t 2 * eta_t^2` and no terminal `1 / eta_T` potential term.","url":"../modules/banditrlproof-tsallisscheduledselfboundingoptimization/index.html#decl-f13f6169f6b0","parent":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","order":10314,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSelfBoundingOptimization"],["Source","BanditRLProof/TsallisScheduledSelfBoundingOptimization.lean:514"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSplit_of_refinedCoefficient_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat → Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : ∀ t, t ≤ horizon → 0 < eta t) (heta_le : ∀ t, t ≤ horizon → eta t ≤ 1 / 2) (hetaMono : ∀ t, t < horizon → eta (t + 1) ≤ eta t) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (b : Nat…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedselfboundingquadraticsplit_of_refinedcoefficient_le banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedselfboundingquadraticsplit_of_refinedcoefficient_le refined generated quadratic split for the theorem route. unlike the coarse scheduled interpolation above, this theorem consumes the uncollapsed stability-penalty coefficient, so its deterministic base contains only `sum_t 2 * eta_t^2` and no terminal `1 / eta_t` potential term. theorem compiled","shard":"modules/348c2c11d8f79600.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbabilityAt","label":"sampledScheduledHalfTsallisExpectedProbabilityAt","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbabilityAt","description":"Expected probability of one action under an arbitrary trajectory law.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-e644d71c5af8","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10315,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisExpectedProbabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (action : Action) : Real","missing":[],"search":"sampledscheduledhalftsallisexpectedprobabilityat banditrlproof.tsallis.sampledscheduledhalftsallisexpectedprobabilityat expected probability of one action under an arbitrary trajectory law. definition compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbabilityAtTime","label":"integrable_sampledScheduledHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbabilityAtTime","description":"theorem integrable_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (action : Action) (haction : action ∈ arms) : Integrable (fun sample =…","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-c5edf03f7a79","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10316,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (action : Action) (haction : action ∈ arms) : Integrable (fun sample => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample action) mu","missing":[],"search":"integrable_sampledscheduledhalftsallisprobabilityattime banditrlproof.tsallis.integrable_sampledscheduledhalftsallisprobabilityattime theorem integrable_sampledscheduledhalftsallisprobabilityattime {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) -> action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (t : nat) (action : action) (haction : action ∈ arms) : integrable (fun sample => sampledscheduledhalftsallisprobabilityattime arms harms eta t sample action) mu theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integrable_sqrt_sampledScheduledHalfTsallisProbabilityAtTime","label":"integrable_sqrt_sampledScheduledHalfTsallisProbabilityAtTime","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integrable_sqrt_sampledScheduledHalfTsallisProbabilityAtTime","description":"theorem integrable_sqrt_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (action : Action) (haction : action ∈ arms) : Integrable (fun sam…","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-f0659a02a9c6","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10317,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:53"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integrable_sqrt_sampledScheduledHalfTsallisProbabilityAtTime {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsFiniteMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (action : Action) (haction : action ∈ arms) : Integrable (fun sample => Real.sqrt (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample action)) mu","missing":[],"search":"integrable_sqrt_sampledscheduledhalftsallisprobabilityattime banditrlproof.tsallis.integrable_sqrt_sampledscheduledhalftsallisprobabilityattime theorem integrable_sqrt_sampledscheduledhalftsallisprobabilityattime {env : type u} {action : type v} [measurablespace env] [measurablespace action] [measurablesingletonclass action] [decidableeq action] (mu : measure (env × ((k : nat) -> action × real))) [isfinitemeasure mu] (arms : finset action) (harms : arms.nonempty) (eta : nat -> real) (t : nat) (action : action) (haction : action ∈ arms) : integrable (fun sample => real.sqrt (sampledscheduledhalftsallisprobabilityattime arms harms eta t sample action)) mu theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledScheduledHalfTsallisExpectedProbabilityAt","label":"finiteSimplex_sampledScheduledHalfTsallisExpectedProbabilityAt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_sampledScheduledHalfTsallisExpectedProbabilityAt","description":"Integrating a generated finite-simplex law under a probability measure again gives a finite-simplex law.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-455dd8c2d63e","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10318,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:78"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_sampledScheduledHalfTsallisExpectedProbabilityAt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) : FTRL.finiteSimplex arms (sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t)","missing":[],"search":"finitesimplex_sampledscheduledhalftsallisexpectedprobabilityat banditrlproof.tsallis.finitesimplex_sampledscheduledhalftsallisexpectedprobabilityat integrating a generated finite-simplex law under a probability measure again gives a finite-simplex law. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sqrt_sampledScheduledHalfTsallisProbabilityAtTime_le_sqrt_expected","label":"integral_sqrt_sampledScheduledHalfTsallisProbabilityAtTime_le_sqrt_expected","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sqrt_sampledScheduledHalfTsallisProbabilityAtTime_le_sqrt_expected","description":"Jensen transport for one scheduled action coordinate.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-40885ca0d6e8","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10319,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:115"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sqrt_sampledScheduledHalfTsallisProbabilityAtTime_le_sqrt_expected {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (t : Nat) (action : Action) (haction : action ∈ arms) : integral mu (fun sample => Real.sqrt (sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample action)) <= Real.sqrt (sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harms eta t action)","missing":[],"search":"integral_sqrt_sampledscheduledhalftsallisprobabilityattime_le_sqrt_expected banditrlproof.tsallis.integral_sqrt_sampledscheduledhalftsallisprobabilityattime_le_sqrt_expected jensen transport for one scheduled action coordinate. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_eq_refined","label":"sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_eq_refined","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_eq_refined","description":"On the small-rate branch, the all-rate actual-time budget is the refined half-power budget of the actual scheduled probability.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-a74efcd1cecd","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10320,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:149"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_eq_refined {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (sample : Env × ((k : Nat) -> Action × Real)) (t : Nat) (heta_le : eta t <= 1 / 2) : sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime arms harms eta sample t = refinedPotentialStabilityBound arms (eta t) (fun sample action => sampledScheduledHalfTsallisProbabilityAtTime arms harms eta t sample action) sample","missing":[],"search":"sampledscheduledhalftsallisallratepotentialstabilityboundattime_eq_refined banditrlproof.tsallis.sampledscheduledhalftsallisallratepotentialstabilityboundattime_eq_refined on the small-rate branch, the all-rate actual-time budget is the refined half-power budget of the actual scheduled probability. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","label":"integral_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","description":"One integrated small-rate budget is bounded by suboptimal-arm square roots of expected scheduled probabilities.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-25c99298d654","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10321,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:184"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (t : Nat) (heta : 0 < eta t) (heta_le : eta t <= 1 / 2) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime arms harms eta sample…","missing":[],"search":"integral_sampledscheduledhalftsallisallratepotentialstabilityboundattime_le_suboptimalexpectedsqrt banditrlproof.tsallis.integral_sampledscheduledhalftsallisallratepotentialstabilityboundattime_le_suboptimalexpectedsqrt one integrated small-rate budget is bounded by suboptimal-arm square roots of expected scheduled probabilities. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","label":"integral_sum_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","description":"The complete integrated all-rate budget, on its small-rate branch, is bounded by a deterministic time/suboptimal-arm sum of square roots of expected scheduled probabilities.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-9be880d754f5","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10322,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:288"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sum_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (horizon : Nat) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.environment integral mu (fun sample => (Finset.range (horizon…","missing":[],"search":"integral_sum_sampledscheduledhalftsallisallratepotentialstabilityboundattime_le_suboptimalexpectedsqrt banditrlproof.tsallis.integral_sum_sampledscheduledhalftsallisallratepotentialstabilityboundattime_le_suboptimalexpectedsqrt the complete integrated all-rate budget, on its small-rate branch, is bounded by a deterministic time/suboptimal-arm sum of square roots of expected scheduled probabilities. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_suboptimalExpectedSqrt","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_suboptimalExpectedSqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_suboptimalExpectedSqrt","description":"Generated scheduled predictable environment regret against the best-arm point mass has the deterministic suboptimal-arm upper needed by the self-bounding completion-of-squares step.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-6d76c05e928b","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10323,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:354"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_suboptimalExpectedSqrt {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms eta loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms eta selector.finiteHistory loss.envir…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_suboptimalexpectedsqrt banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_suboptimalexpectedsqrt generated scheduled predictable environment regret against the best-arm point mass has the deterministic suboptimal-arm upper needed by the self-bounding completion-of-squares step. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_selfBounding","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_selfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_selfBounding","description":"The generated scheduled upper bound feeds the abstract self-bounding completion-of-squares theorem without any remaining pathwise or Jensen obligation.","url":"../modules/banditrlproof-tsallisscheduledsuboptimalexpectedbound/index.html#decl-8868b5c97884","parent":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","order":10324,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledSuboptimalExpectedBound"],["Source","BanditRLProof/TsallisScheduledSuboptimalExpectedBound.lean:399"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_selfBounding {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (heta : forall t, t <= horizon -> 0 < eta t) (heta_le : forall t, t <= horizon -> eta t <= 1 / 2) (hetaMono : forall t, t < horizon -> eta (t + 1) <= eta t) (gap : Action -> Real) (hgap : forall action, action ∈ arms.erase best -> 0 < gap action) (corruption : Real) (hselfBounding : (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => gap acti…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_of_selfbounding banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_of_selfbounding the generated scheduled upper bound feeds the abstract self-bounding completion-of-squares theorem without any remaining pathwise or jensen obligation. theorem compiled","shard":"modules/e53088856a6c064e.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledTimeVaryingExpectedGapLaw","label":"HasScheduledTimeVaryingExpectedGapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledTimeVaryingExpectedGapLaw","description":"Per-time first-moment law for probability-weighted predictable loss gaps.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-5c4055ae21de","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10325,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:22"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledTimeVaryingExpectedGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Nat -> Action -> Real) (horizon : Nat) : Prop","missing":[],"search":"hasscheduledtimevaryingexpectedgaplaw banditrlproof.tsallis.hasscheduledtimevaryingexpectedgaplaw per-time first-moment law for probability-weighted predictable loss gaps. definition compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledTimeVaryingConditionalMeanGapLaw","label":"HasScheduledTimeVaryingConditionalMeanGapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledTimeVaryingConditionalMeanGapLaw","description":"Per-time conditional mean law relative to the information available before the scheduled action.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-798aa5525ec1","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10326,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledTimeVaryingConditionalMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Nat -> Action -> Real) (horizon : Nat) : Prop","missing":[],"search":"hasscheduledtimevaryingconditionalmeangaplaw banditrlproof.tsallis.hasscheduledtimevaryingconditionalmeangaplaw per-time conditional mean law relative to the information available before the scheduled action. definition compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasScheduledTimeVaryingIndependentMeanGapLaw","label":"HasScheduledTimeVaryingIndependentMeanGapLaw","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasScheduledTimeVaryingIndependentMeanGapLaw","description":"Per-time independence and global-mean law for predictable loss gaps.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-78be49de4758","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10327,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:57"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasScheduledTimeVaryingIndependentMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) (arms : Finset Action) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Nat -> Action -> Real) (horizon : Nat) : Prop","missing":[],"search":"hasscheduledtimevaryingindependentmeangaplaw banditrlproof.tsallis.hasscheduledtimevaryingindependentmeangaplaw per-time independence and global-mean law for predictable loss gaps. definition compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingConditionalMeanGapLaw_of_independentMeanGapLaw","label":"hasScheduledTimeVaryingConditionalMeanGapLaw_of_independentMeanGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledTimeVaryingConditionalMeanGapLaw_of_independentMeanGapLaw","description":"Independence from the scheduled past identifies each time-varying conditional mean.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-714bbb8944dd","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10328,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:74"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledTimeVaryingConditionalMeanGapLaw_of_independentMeanGapLaw {Env : Type u} {Action : Type v} [mEnv : MeasurableSpace Env] [mAction : MeasurableSpace Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Nat -> Action -> Real) (horizon : Nat) (hgap : HasScheduledTimeVaryingIndependentMeanGapLaw mu arms loss best gap horizon) : HasScheduledTimeVaryingConditionalMeanGapLaw mu arms loss best gap horizon","missing":[],"search":"hasscheduledtimevaryingconditionalmeangaplaw_of_independentmeangaplaw banditrlproof.tsallis.hasscheduledtimevaryingconditionalmeangaplaw_of_independentmeangaplaw independence from the scheduled past identifies each time-varying conditional mean. theorem compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingExpectedGapLaw_of_conditionalMeanGapLaw","label":"hasScheduledTimeVaryingExpectedGapLaw_of_conditionalMeanGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledTimeVaryingExpectedGapLaw_of_conditionalMeanGapLaw","description":"A time-varying conditional mean law yields the corresponding weighted expected-gap law by pulling the scheduled probability through `condExp`.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-173a8a023bcc","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10329,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:123"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledTimeVaryingExpectedGapLaw_of_conditionalMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Nat -> Action -> Real) (horizon : Nat) (hgap : HasScheduledTimeVaryingConditionalMeanGapLaw mu arms loss best gap horizon) : HasScheduledTimeVaryingExpectedGapLaw mu arms harms eta loss best gap horizon","missing":[],"search":"hasscheduledtimevaryingexpectedgaplaw_of_conditionalmeangaplaw banditrlproof.tsallis.hasscheduledtimevaryingexpectedgaplaw_of_conditionalmeangaplaw a time-varying conditional mean law yields the corresponding weighted expected-gap law by pulling the scheduled probability through `condexp`. theorem compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingExpectedGapLaw_of_independentMeanGapLaw","label":"hasScheduledTimeVaryingExpectedGapLaw_of_independentMeanGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.hasScheduledTimeVaryingExpectedGapLaw_of_independentMeanGapLaw","description":"The independent time-varying mean contract directly feeds the weighted expected-gap law.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-2ca31f118942","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10330,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:201"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem hasScheduledTimeVaryingExpectedGapLaw_of_independentMeanGapLaw {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) (best : Action) (gap : Nat -> Action -> Real) (horizon : Nat) (hgap : HasScheduledTimeVaryingIndependentMeanGapLaw mu arms loss best gap horizon) : HasScheduledTimeVaryingExpectedGapLaw mu arms harms eta loss best gap horizon","missing":[],"search":"hasscheduledtimevaryingexpectedgaplaw_of_independentmeangaplaw banditrlproof.tsallis.hasscheduledtimevaryingexpectedgaplaw_of_independentmeangaplaw the independent time-varying mean contract directly feeds the weighted expected-gap law. theorem compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalTimeVaryingExpectedGapMass","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalTimeVaryingExpectedGapMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalTimeVaryingExpectedGapMass","description":"A time-varying expected-gap law identifies integrated regret with the time-by-arm actual gap mass.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-080cb4bc20d9","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10331,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:222"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalTimeVaryingExpectedGapMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Nat -> Action -> Real) (horizon : Nat) (hgapLaw : HasScheduledTimeVaryingExpectedGapLaw mu arms harms eta loss best gap horizon) : integral mu (sampledScheduledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon) = (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => gap t action * sampledScheduledHalfTsallisExpectedProbabilityAt mu arms harm…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimaltimevaryingexpectedgapmass banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_eq_suboptimaltimevaryingexpectedgapmass a time-varying expected-gap law identifies integrated regret with the time-by-arm actual gap mass. theorem compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.scheduledTimeVaryingGapDeviationBudget","label":"scheduledTimeVaryingGapDeviationBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.scheduledTimeVaryingGapDeviationBudget","description":"Accumulated coordinatewise deviation from fixed baseline gaps.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-a6f7bf027889","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10332,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:271"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def scheduledTimeVaryingGapDeviationBudget {Action : Type*} [DecidableEq Action] (arms : Finset Action) (best : Action) (horizon : Nat) (deviation : Nat -> Action -> Real) : Real","missing":[],"search":"scheduledtimevaryinggapdeviationbudget banditrlproof.tsallis.scheduledtimevaryinggapdeviationbudget accumulated coordinatewise deviation from fixed baseline gaps. definition compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_timeVaryingPerturbedExpectedGapLaw","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_timeVaryingPerturbedExpectedGapLaw","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_timeVaryingPerturbedExpectedGapLaw","description":"A time-varying actual-gap law within a known coordinatewise distance of fixed baseline gaps yields the self-bound consumed by schedule tuning.","url":"../modules/banditrlproof-tsallisscheduledtimevaryingexpectedgap/index.html#decl-4b7de965a77b","parent":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","order":10333,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisScheduledTimeVaryingExpectedGap"],["Source","BanditRLProof/TsallisScheduledTimeVaryingExpectedGap.lean:280"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_timeVaryingPerturbedExpectedGapLaw {Env Action : Type*} [MeasurableSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [DecidableEq Action] (mu : Measure (Env × ((k : Nat) -> Action × Real))) [IsProbabilityMeasure mu] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (baseGap : Action -> Real) (actualGap deviation : Nat -> Action -> Real) (hactualGapLaw : HasScheduledTimeVaryingExpectedGapLaw mu arms harms eta loss best actualGap horizon) (hdeviation : ∀ t, t <= horizon -> ∀ action, action ∈ arms.erase best -> |actualGap t action - baseGap action| <= deviation t action) : (Finset.range (horizon + 1)).sum (fun t => (arms.erase best).sum (fun action => baseGap acti…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_timevaryingperturbedexpectedgaplaw banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_hasselfbounding_of_timevaryingperturbedexpectedgaplaw a time-varying actual-gap law within a known coordinatewise distance of fixed baseline gaps yields the self-bound consumed by schedule tuning. theorem compiled","shard":"modules/4962101e2a47ba19.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.pointMass","label":"pointMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.pointMass","description":"Unit mass at one action, used as the fixed optimal-arm comparator.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-be646b7a223f","parent":"module:BanditRLProof.TsallisSelfBounding","order":10334,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:27"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def pointMass {Action : Type u} [DecidableEq Action] (best : Action) : Action -> Real","missing":[],"search":"pointmass banditrlproof.tsallis.pointmass unit mass at one action, used as the fixed optimal-arm comparator. definition compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.finiteSimplex_pointMass","label":"finiteSimplex_pointMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.finiteSimplex_pointMass","description":"theorem finiteSimplex_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) : FTRL.finiteSimplex arms (pointMass best)","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-454daf0a5532","parent":"module:BanditRLProof.TsallisSelfBounding","order":10335,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem finiteSimplex_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) : FTRL.finiteSimplex arms (pointMass best)","missing":[],"search":"finitesimplex_pointmass banditrlproof.tsallis.finitesimplex_pointmass theorem finitesimplex_pointmass {action : type u} [decidableeq action] (arms : finset action) {best : action} (hbest : best ∈ arms) : ftrl.finitesimplex arms (pointmass best) theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_pointMass","label":"linearLoss_pointMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_pointMass","description":"theorem linearLoss_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (loss : Action -> Real) : FTRL.linearLoss arms (pointMass best) loss = loss best","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-a87a0f7464e2","parent":"module:BanditRLProof.TsallisSelfBounding","order":10336,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:40"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (loss : Action -> Real) : FTRL.linearLoss arms (pointMass best) loss = loss best","missing":[],"search":"linearloss_pointmass banditrlproof.tsallis.linearloss_pointmass theorem linearloss_pointmass {action : type u} [decidableeq action] (arms : finset action) {best : action} (hbest : best ∈ arms) (loss : action -> real) : ftrl.linearloss arms (pointmass best) loss = loss best theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.linearLoss_sub_pointMass_eq_gapMass","label":"linearLoss_sub_pointMass_eq_gapMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.linearLoss_sub_pointMass_eq_gapMass","description":"One-round regret against a point mass is the probability-weighted gap.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-1025765b3be3","parent":"module:BanditRLProof.TsallisSelfBounding","order":10337,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:49"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem linearLoss_sub_pointMass_eq_gapMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) (prob loss gap : Action -> Real) (hprob : FTRL.finiteSimplex arms prob) (hgap : ∀ action ∈ arms, loss action - loss best = gap action) : FTRL.linearLoss arms prob loss - FTRL.linearLoss arms (pointMass best) loss = arms.sum (fun action => prob action * gap action)","missing":[],"search":"linearloss_sub_pointmass_eq_gapmass banditrlproof.tsallis.linearloss_sub_pointmass_eq_gapmass one-round regret against a point mass is the probability-weighted gap. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableGapMass","label":"sampledHalfTsallisPredictableGapMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableGapMass","description":"Gap mass accumulated by the generated actual-time sampling laws.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-b266302229ce","parent":"module:BanditRLProof.TsallisSelfBounding","order":10338,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:78"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledHalfTsallisPredictableGapMass {Env : Type u} {Action : Type v} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (gap : Action -> Real) (horizon : Nat) (sample : Env × ((k : Nat) -> Action × Real)) : Real","missing":[],"search":"sampledhalftsallispredictablegapmass banditrlproof.tsallis.sampledhalftsallispredictablegapmass gap mass accumulated by the generated actual-time sampling laws. definition compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.HasSelfBoundingRegret","label":"HasSelfBoundingRegret","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.HasSelfBoundingRegret","description":"Paper-facing scalar form of a `(Delta, C, T)` self-bounding constraint.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-14bdd368d7af","parent":"module:BanditRLProof.TsallisSelfBounding","order":10339,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:89"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def HasSelfBoundingRegret (expectedGapMass regret corruption : Real) : Prop","missing":[],"search":"hasselfboundingregret banditrlproof.tsallis.hasselfboundingregret paper-facing scalar form of a `(delta, c, t)` self-bounding constraint. definition compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_gapMass","label":"sampledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_gapMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_gapMass","description":"A predictable fixed-gap law identifies environment regret pathwise.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-51eb87a7e485","parent":"module:BanditRLProof.TsallisSelfBounding","order":10340,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:94"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_gapMass {Env : Type u} {Action : Type v} [MeasurableSpace Env] [MeasurableSpace Action] [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgap : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) (sample : Env × ((k : Nat) -> Action × Real)) : sampledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon sample = sampledHalfTsallisPredictableGapMass arms harms eta gap horizon sample","missing":[],"search":"sampledhalftsallispredictableenvironmentregret_pointmass_eq_gapmass banditrlproof.tsallis.sampledhalftsallispredictableenvironmentregret_pointmass_eq_gapmass a predictable fixed-gap law identifies environment regret pathwise. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding","label":"integral_sampledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding","description":"The fixed-gap predictable environment satisfies the integrated self-bounding condition, with any nonnegative corruption allowance.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-5107a44d0908","parent":"module:BanditRLProof.TsallisSelfBounding","order":10341,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:123"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (gap : Action -> Real) (horizon : Nat) (hgap : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) (corruption : Real) (hcorruption : 0 ≤ corruption) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environmen…","missing":[],"search":"integral_sampledhalftsallispredictableenvironmentregret_hasselfbounding banditrlproof.tsallis.integral_sampledhalftsallispredictableenvironmentregret_hasselfbounding the fixed-gap predictable environment satisfies the integrated self-bounding condition, with any nonnegative corruption allowance. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_pointMass_le","label":"integral_sampledHalfTsallisPredictableEnvironmentRegret_pointMass_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_pointMass_le","description":"Actual generated environment-regret upper bound specialized to the optimal-arm point-mass comparator.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-ff638402e345","parent":"module:BanditRLProof.TsallisSelfBounding","order":10342,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:170"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledHalfTsallisPredictableEnvironmentRegret_pointMass_le {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (eta : Real) (heta : 0 < eta) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) : let selector := canonicalHalfTsallisFiniteHistorySelectorMeasurability arms harms eta let mu := prior ⊗ₘ sampledHalfTsallisTrajectoryKernel arms harms eta selector loss.environment integral mu (sampledHalfTsallisPredictableEnvironmentRegret arms harms eta loss (pointMass best) horizon) <= 2 * eta * powerSum arms (1 / 2 : Real) (initialHalfTsallisDistribution arms harms eta) + integr…","missing":[],"search":"integral_sampledhalftsallispredictableenvironmentregret_pointmass_le banditrlproof.tsallis.integral_sampledhalftsallispredictableenvironmentregret_pointmass_le actual generated environment-regret upper bound specialized to the optimal-arm point-mass comparator. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap","label":"two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap","description":"Coordinate completion of squares used by the self-bounding conversion.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-92865423ccd1","parent":"module:BanditRLProof.TsallisSelfBounding","order":10343,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:200"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap (coeff probability gap : Real) (hprobability : 0 ≤ probability) (hgap : 0 < gap) : 2 * coeff * Real.sqrt probability - gap * probability ≤ coeff ^ 2 / gap","missing":[],"search":"two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap banditrlproof.tsallis.two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap coordinate completion of squares used by the self-bounding conversion. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption","label":"regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption","description":"Finite `(Delta, C, T)` self-bounding conversion after a refined upper bound has removed every zero-gap coordinate.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-d7be7be7576f","parent":"module:BanditRLProof.TsallisSelfBounding","order":10344,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:221"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption {Index : Type u} (indices : Finset Index) (probability coeff gap : Index -> Real) (regret base corruption : Real) (hprobability : ∀ index ∈ indices, 0 ≤ probability index) (hgap : ∀ index ∈ indices, 0 < gap index) (hselfBounding : indices.sum (fun index => gap index * probability index) - corruption ≤ regret) (hupper : regret ≤ base + indices.sum (fun index => coeff index * Real.sqrt (probability index))) : regret ≤ 2 * base + indices.sum (fun index => coeff index ^ 2 / gap index) + corruption","missing":[],"search":"regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption banditrlproof.tsallis.regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption finite `(delta, c, t)` self-bounding conversion after a refined upper bound has removed every zero-gap coordinate. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.powerSum_pointMass_half","label":"powerSum_pointMass_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.powerSum_pointMass_half","description":"theorem powerSum_pointMass_half {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) : powerSum arms (1 / 2 : Real) (pointMass best) = 1","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-a16574ca0540","parent":"module:BanditRLProof.TsallisSelfBounding","order":10345,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:268"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem powerSum_pointMass_half {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) : powerSum arms (1 / 2 : Real) (pointMass best) = 1","missing":[],"search":"powersum_pointmass_half banditrlproof.tsallis.powersum_pointmass_half theorem powersum_pointmass_half {action : type u} [decidableeq action] (arms : finset action) {best : action} (hbest : best ∈ arms) : powersum arms (1 / 2 : real) (pointmass best) = 1 theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_pointMass","label":"sum_erase_sqrt_pointMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_erase_sqrt_pointMass","description":"theorem sum_erase_sqrt_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) (best : Action) : (arms.erase best).sum (fun action => Real.sqrt (pointMass best action)) = 0","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-3734daea07de","parent":"module:BanditRLProof.TsallisSelfBounding","order":10346,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:275"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_erase_sqrt_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) (best : Action) : (arms.erase best).sum (fun action => Real.sqrt (pointMass best action)) = 0","missing":[],"search":"sum_erase_sqrt_pointmass banditrlproof.tsallis.sum_erase_sqrt_pointmass theorem sum_erase_sqrt_pointmass {action : type u} [decidableeq action] (arms : finset action) (best : action) : (arms.erase best).sum (fun action => real.sqrt (pointmass best action)) = 0 theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.not_forall_powerSum_half_le_sum_erase_sqrt","label":"not_forall_powerSum_half_le_sum_erase_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.not_forall_powerSum_half_le_sum_erase_sqrt","description":"The existing total half-power budget cannot uniformly be replaced by a suboptimal-arm square-root budget. A refined `(1-p)` factor is genuinely needed before the self-bounding conversion can consume the trajectory bound.","url":"../modules/banditrlproof-tsallisselfbounding/index.html#decl-f826e94b19ba","parent":"module:BanditRLProof.TsallisSelfBounding","order":10347,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBounding"],["Source","BanditRLProof/TsallisSelfBounding.lean:288"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem not_forall_powerSum_half_le_sum_erase_sqrt {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) : ¬ ∀ p : Action -> Real, FTRL.finiteSimplex arms p -> powerSum arms (1 / 2 : Real) p ≤ (arms.erase best).sum (fun action => Real.sqrt (p action))","missing":[],"search":"not_forall_powersum_half_le_sum_erase_sqrt banditrlproof.tsallis.not_forall_powersum_half_le_sum_erase_sqrt the existing total half-power budget cannot uniformly be replaced by a suboptimal-arm square-root budget. a refined `(1-p)` factor is genuinely needed before the self-bounding conversion can consume the trajectory bound. theorem compiled","shard":"modules/d1fde7e2599ad19b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.selfBoundingBetaEquation","label":"selfBoundingBetaEquation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.selfBoundingBetaEquation","description":"The scalar equation used to tune the refined self-bounding parameter.","url":"../modules/banditrlproof-tsallisselfboundingbetaroot/index.html#decl-c96e01f1b96e","parent":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","order":10348,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSelfBoundingBetaRoot"],["Source","BanditRLProof/TsallisSelfBoundingBetaRoot.lean:19"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def selfBoundingBetaEquation (scale reciprocalGap corruption beta : Real) : Real","missing":[],"search":"selfboundingbetaequation banditrlproof.tsallis.selfboundingbetaequation the scalar equation used to tune the refined self-bounding parameter. definition compiled","shard":"modules/40098b95879e0116.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.continuousOn_selfBoundingBetaEquation","label":"continuousOn_selfBoundingBetaEquation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.continuousOn_selfBoundingBetaEquation","description":"The beta equation is continuous on every interval bounded below by one.","url":"../modules/banditrlproof-tsallisselfboundingbetaroot/index.html#decl-29aab2d58685","parent":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","order":10349,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBoundingBetaRoot"],["Source","BanditRLProof/TsallisSelfBoundingBetaRoot.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem continuousOn_selfBoundingBetaEquation (scale reciprocalGap corruption upper : Real) : ContinuousOn (selfBoundingBetaEquation scale reciprocalGap corruption) (Icc 1 upper)","missing":[],"search":"continuouson_selfboundingbetaequation banditrlproof.tsallis.continuouson_selfboundingbetaequation the beta equation is continuous on every interval bounded below by one. theorem compiled","shard":"modules/40098b95879e0116.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_selfBoundingBetaEquation_eq_zero","label":"exists_selfBoundingBetaEquation_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_selfBoundingBetaEquation_eq_zero","description":"Under the corruption window used by the refined Tsallis-INF analysis, the scalar beta equation has a zero between one and `scale / reciprocalGap ^ 2`. This is the intermediate-value certificate that precedes any quantitative Lambert-W estimate.","url":"../modules/banditrlproof-tsallisselfboundingbetaroot/index.html#decl-e3afbc4379f9","parent":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","order":10350,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSelfBoundingBetaRoot"],["Source","BanditRLProof/TsallisSelfBoundingBetaRoot.lean:41"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_selfBoundingBetaEquation_eq_zero (scale reciprocalGap corruption : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hupper : 1 <= scale / reciprocalGap ^ 2) (hcorruptionUpper : corruption * reciprocalGap <= scale) (hcorruptionLower : reciprocalGap * (Real.log (scale / reciprocalGap ^ 2) + 1) <= corruption) : exists beta, beta ∈ Icc 1 (scale / reciprocalGap ^ 2) ∧ selfBoundingBetaEquation scale reciprocalGap corruption beta = 0","missing":[],"search":"exists_selfboundingbetaequation_eq_zero banditrlproof.tsallis.exists_selfboundingbetaequation_eq_zero under the corruption window used by the refined tsallis-inf analysis, the scalar beta equation has a zero between one and `scale / reciprocalgap ^ 2`. this is the intermediate-value certificate that precedes any quantitative lambert-w estimate. theorem compiled","shard":"modules/40098b95879e0116.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule","label":"sampledScheduledHalfTsallisSqrtSchedule","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule","description":"The concrete small-rate schedule used by the fixed-gap route.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-4af39deeb22b","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10351,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:20"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisSqrtSchedule (t : Nat) : Real","missing":[],"search":"sampledscheduledhalftsallissqrtschedule banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule the concrete small-rate schedule used by the fixed-gap route. definition compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget","label":"sampledScheduledHalfTsallisHarmonicBudget","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget","description":"The finite harmonic budget through the inclusive horizon.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-98f4282e39c4","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10352,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:24"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisHarmonicBudget (horizon : Nat) : Real","missing":[],"search":"sampledscheduledhalftsallisharmonicbudget banditrlproof.tsallis.sampledscheduledhalftsallisharmonicbudget the finite harmonic budget through the inclusive horizon. definition compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget_eq_harmonic","label":"sampledScheduledHalfTsallisHarmonicBudget_eq_harmonic","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget_eq_harmonic","description":"The local real-valued budget is the real cast of Mathlib's rational harmonic number.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-702e4f6f2957","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10353,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:31"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisHarmonicBudget_eq_harmonic (horizon : Nat) : sampledScheduledHalfTsallisHarmonicBudget horizon = (harmonic (horizon + 1) : Real)","missing":[],"search":"sampledscheduledhalftsallisharmonicbudget_eq_harmonic banditrlproof.tsallis.sampledscheduledhalftsallisharmonicbudget_eq_harmonic the local real-valued budget is the real cast of mathlib's rational harmonic number. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget_le_one_add_log","label":"sampledScheduledHalfTsallisHarmonicBudget_le_one_add_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget_le_one_add_log","description":"Mathlib's finite harmonic estimate gives the explicit logarithmic budget.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-4388df83f1ea","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10354,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:43"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisHarmonicBudget_le_one_add_log (horizon : Nat) : sampledScheduledHalfTsallisHarmonicBudget horizon <= 1 + Real.log (((horizon + 1 : Nat) : Real))","missing":[],"search":"sampledscheduledhalftsallisharmonicbudget_le_one_add_log banditrlproof.tsallis.sampledscheduledhalftsallisharmonicbudget_le_one_add_log mathlib's finite harmonic estimate gives the explicit logarithmic budget. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_pos","label":"sampledScheduledHalfTsallisSqrtSchedule_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_pos","description":"theorem sampledScheduledHalfTsallisSqrtSchedule_pos (t : Nat) : 0 < sampledScheduledHalfTsallisSqrtSchedule t","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-8241be38be2c","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10355,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:50"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSqrtSchedule_pos (t : Nat) : 0 < sampledScheduledHalfTsallisSqrtSchedule t","missing":[],"search":"sampledscheduledhalftsallissqrtschedule_pos banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule_pos theorem sampledscheduledhalftsallissqrtschedule_pos (t : nat) : 0 < sampledscheduledhalftsallissqrtschedule t theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_le_half","label":"sampledScheduledHalfTsallisSqrtSchedule_le_half","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_le_half","description":"theorem sampledScheduledHalfTsallisSqrtSchedule_le_half (t : Nat) : sampledScheduledHalfTsallisSqrtSchedule t <= 1 / 2","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-9bc2e5350d4f","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10356,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:55"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSqrtSchedule_le_half (t : Nat) : sampledScheduledHalfTsallisSqrtSchedule t <= 1 / 2","missing":[],"search":"sampledscheduledhalftsallissqrtschedule_le_half banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule_le_half theorem sampledscheduledhalftsallissqrtschedule_le_half (t : nat) : sampledscheduledhalftsallissqrtschedule t <= 1 / 2 theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_succ_le","label":"sampledScheduledHalfTsallisSqrtSchedule_succ_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_succ_le","description":"theorem sampledScheduledHalfTsallisSqrtSchedule_succ_le (t : Nat) : sampledScheduledHalfTsallisSqrtSchedule (t + 1) <= sampledScheduledHalfTsallisSqrtSchedule t","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-8e3d74aaf536","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10357,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:65"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSqrtSchedule_succ_le (t : Nat) : sampledScheduledHalfTsallisSqrtSchedule (t + 1) <= sampledScheduledHalfTsallisSqrtSchedule t","missing":[],"search":"sampledscheduledhalftsallissqrtschedule_succ_le banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule_succ_le theorem sampledscheduledhalftsallissqrtschedule_succ_le (t : nat) : sampledscheduledhalftsallissqrtschedule (t + 1) <= sampledscheduledhalftsallissqrtschedule t theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_four_mul_sq","label":"sampledScheduledHalfTsallisSqrtSchedule_four_mul_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_four_mul_sq","description":"theorem sampledScheduledHalfTsallisSqrtSchedule_four_mul_sq (t : Nat) : 4 * (sampledScheduledHalfTsallisSqrtSchedule t) ^ 2 = 1 / (((t + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-58620c6c1438","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10358,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:79"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSqrtSchedule_four_mul_sq (t : Nat) : 4 * (sampledScheduledHalfTsallisSqrtSchedule t) ^ 2 = 1 / (((t + 1 : Nat) : Real))","missing":[],"search":"sampledscheduledhalftsallissqrtschedule_four_mul_sq banditrlproof.tsallis.sampledscheduledhalftsallissqrtschedule_four_mul_sq theorem sampledscheduledhalftsallissqrtschedule_four_mul_sq (t : nat) : 4 * (sampledscheduledhalftsallissqrtschedule t) ^ 2 = 1 / (((t + 1 : nat) : real)) theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sqrt_succ_sub_sqrt_le_one_div_sqrt_succ","label":"sqrt_succ_sub_sqrt_le_one_div_sqrt_succ","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sqrt_succ_sub_sqrt_le_one_div_sqrt_succ","description":"private theorem sqrt_succ_sub_sqrt_le_one_div_sqrt_succ (t : Nat) : Real.sqrt (((t + 2 : Nat) : Real)) - Real.sqrt (((t + 1 : Nat) : Real)) <= 1 / Real.sqrt (((t + 2 : Nat) : Real))","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-e56e53cd6192","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10359,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:92"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"private theorem sqrt_succ_sub_sqrt_le_one_div_sqrt_succ (t : Nat) : Real.sqrt (((t + 2 : Nat) : Real)) - Real.sqrt (((t + 1 : Nat) : Real)) <= 1 / Real.sqrt (((t + 2 : Nat) : Real))","missing":[],"search":"sqrt_succ_sub_sqrt_le_one_div_sqrt_succ banditrlproof.tsallis.sqrt_succ_sub_sqrt_le_one_div_sqrt_succ private theorem sqrt_succ_sub_sqrt_le_one_div_sqrt_succ (t : nat) : real.sqrt (((t + 2 : nat) : real)) - real.sqrt (((t + 1 : nat) : real)) <= 1 / real.sqrt (((t + 2 : nat) : real)) theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_nonneg","label":"sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_nonneg","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_nonneg","description":"theorem sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_nonneg (t : Nat) : 0 <= sampledScheduledHalfTsallisRefinedCoefficient sampledScheduledHalfTsallisSqrtSchedule t","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-b4d62c17e6f4","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10360,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:116"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_nonneg (t : Nat) : 0 <= sampledScheduledHalfTsallisRefinedCoefficient sampledScheduledHalfTsallisSqrtSchedule t","missing":[],"search":"sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_nonneg banditrlproof.tsallis.sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_nonneg theorem sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_nonneg (t : nat) : 0 <= sampledscheduledhalftsallisrefinedcoefficient sampledscheduledhalftsallissqrtschedule t theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_le","label":"sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_le","description":"theorem sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_le (t : Nat) : sampledScheduledHalfTsallisRefinedCoefficient sampledScheduledHalfTsallisSqrtSchedule t <= 5 / Real.sqrt (((t + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-7316ed6d7822","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10361,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:132"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_le (t : Nat) : sampledScheduledHalfTsallisRefinedCoefficient sampledScheduledHalfTsallisSqrtSchedule t <= 5 / Real.sqrt (((t + 1 : Nat) : Real))","missing":[],"search":"sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_le banditrlproof.tsallis.sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_le theorem sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_le (t : nat) : sampledscheduledhalftsallisrefinedcoefficient sampledscheduledhalftsallissqrtschedule t <= 5 / real.sqrt (((t + 1 : nat) : real)) theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_sq_le","label":"sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_sq_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_sq_le","description":"theorem sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_sq_le (t : Nat) : (sampledScheduledHalfTsallisRefinedCoefficient sampledScheduledHalfTsallisSqrtSchedule t) ^ 2 <= 25 / (((t + 1 : Nat) : Real))","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-fc6ae386cf71","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10362,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:156"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_sq_le (t : Nat) : (sampledScheduledHalfTsallisRefinedCoefficient sampledScheduledHalfTsallisSqrtSchedule t) ^ 2 <= 25 / (((t + 1 : Nat) : Real))","missing":[],"search":"sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_sq_le banditrlproof.tsallis.sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_sq_le theorem sampledscheduledhalftsallisrefinedcoefficient_sqrtschedule_sq_le (t : nat) : (sampledscheduledhalftsallisrefinedcoefficient sampledscheduledhalftsallissqrtschedule t) ^ 2 <= 25 / (((t + 1 : nat) : real)) theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_of_selfBounding","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_of_selfBounding","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_of_selfBounding","description":"The square-root schedule closes any matching self-bounding route with a finite harmonic budget and an explicit reciprocal-gap factor.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-5e36e61b8e7e","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10363,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:179"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_of_selfBounding {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < gap action) (corruption : Real) (hselfBounding : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampledScheduledHalfTsallisSqrtSchedule loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule selector.…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_of_selfbounding banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_of_selfbounding the square-root schedule closes any matching self-bounding route with a finite harmonic budget and an explicit reciprocal-gap factor. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_fixedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_fixedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_fixedGap","description":"The square-root schedule closes the exact fixed-gap route with a finite harmonic budget and an explicit reciprocal-gap factor.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-530bb0a3a070","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10364,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:304"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_fixedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) (corruption : Real) (hcorruption : 0 <= corruption) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampledSc…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_fixedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_fixedgap the square-root schedule closes the exact fixed-gap route with a finite harmonic budget and an explicit reciprocal-gap factor. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_expectedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_expectedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_expectedGap","description":"The square-root schedule also closes a coordinatewise expected-gap law; no samplewise fixed-gap identity is required.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-ecfee37e0d5b","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10365,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:341"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_expectedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampledScheduledHalfTsallisSqrtSchedule loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule selector.finiteHistory loss.environment…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_expectedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_expectedgap the square-root schedule also closes a coordinatewise expected-gap law; no samplewise fixed-gap identity is required. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_fixedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_fixedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_fixedGap","description":"The concrete square-root schedule has an explicit logarithmic fixed-gap regret bound, obtained from Mathlib's finite harmonic estimate.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-4abf4c0f7a1a","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10366,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:385"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_fixedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : ∀ t sample action, action ∈ arms -> Exp3.predictableLossAt loss t sample action - Exp3.predictableLossAt loss t sample best = gap action) (corruption : Real) (hcorruption : 0 <= corruption) : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampl…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_fixedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_fixedgap the concrete square-root schedule has an explicit logarithmic fixed-gap regret bound, obtained from mathlib's finite harmonic estimate. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_expectedGap","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_expectedGap","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_expectedGap","description":"Mathlib's harmonic estimate turns the coordinatewise expected-gap route into the same explicit logarithmic square-root-schedule bound.","url":"../modules/banditrlproof-tsallissqrtschedulefixedgap/index.html#decl-def8ff8952f7","parent":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","order":10367,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleFixedGap"],["Source","BanditRLProof/TsallisSqrtScheduleFixedGap.lean:457"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_expectedGap {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hgapPos : ∀ action, action ∈ arms.erase best -> 0 < gap action) (hgapLaw : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampledScheduledHalfTsallisSqrtSchedule loss let mu := prior ⊗ₘ sampledScheduledHalfTsallisTrajectoryKernel arms harms sampledScheduledHalfTsallisSqrtSchedule selector.finiteHistory loss.environ…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_expectedgap banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_log_expectedgap mathlib's harmonic estimate turns the coordinatewise expected-gap route into the same explicit logarithmic square-root-schedule bound. theorem compiled","shard":"modules/ce68acf4a8761837.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_range_one_div_sqrt_natSucc_le_two_sqrt","label":"sum_range_one_div_sqrt_natSucc_le_two_sqrt","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_range_one_div_sqrt_natSucc_le_two_sqrt","description":"Integral-comparison bound for the shifted inverse-square-root prefix.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-dd18429a71ea","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10368,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:21"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_range_one_div_sqrt_natSucc_le_two_sqrt (n : Nat) : (Finset.range n).sum (fun t => 1 / Real.sqrt (((t + 1 : Nat) : Real))) ≤ 2 * Real.sqrt (n : Real)","missing":[],"search":"sum_range_one_div_sqrt_natsucc_le_two_sqrt banditrlproof.tsallis.sum_range_one_div_sqrt_natsucc_le_two_sqrt integral-comparison bound for the shifted inverse-square-root prefix. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_Ico_one_div_natSucc_le_log_div","label":"sum_Ico_one_div_natSucc_le_log_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_Ico_one_div_natSucc_le_log_div","description":"Harmonic tail bound with the exact logarithmic ratio needed by the self-bounding threshold split.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-f374e0e54b25","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10369,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:56"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_Ico_one_div_natSucc_le_log_div (m n : Nat) (hm : 0 < m) (hmn : m ≤ n) : (Finset.Ico m n).sum (fun t => 1 / (((t + 1 : Nat) : Real))) ≤ Real.log ((n : Real) / (m : Real))","missing":[],"search":"sum_ico_one_div_natsucc_le_log_div banditrlproof.tsallis.sum_ico_one_div_natsucc_le_log_div harmonic tail bound with the exact logarithmic ratio needed by the self-bounding threshold split. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_range_sqrtSchedule_activeBranch_le_closedForm","label":"sum_range_sqrtSchedule_activeBranch_le_closedForm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_range_sqrtSchedule_activeBranch_le_closedForm","description":"The active-prefix branch of the square-root schedule has a closed-form inverse-square-root bound.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-a01028a43966","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10370,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:83"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_range_sqrtSchedule_activeBranch_le_closedForm (cutoff : Nat) (amplitude sqrtCard card reciprocalGap : Real) (hamplitude : 0 <= amplitude) (hsqrtCard : 0 <= sqrtCard) : (Finset.range cutoff).sum (fun t => amplitude / Real.sqrt (((t + 1 : Nat) : Real)) * sqrtCard - card / reciprocalGap) <= 2 * amplitude * sqrtCard * Real.sqrt (cutoff : Real) - (cutoff : Real) * card / reciprocalGap","missing":[],"search":"sum_range_sqrtschedule_activebranch_le_closedform banditrlproof.tsallis.sum_range_sqrtschedule_activebranch_le_closedform the active-prefix branch of the square-root schedule has a closed-form inverse-square-root bound. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_Ico_sqrtSchedule_unconstrainedBranch_le_log","label":"sum_Ico_sqrtSchedule_unconstrainedBranch_le_log","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_Ico_sqrtSchedule_unconstrainedBranch_le_log","description":"The unconstrained tail branch of the square-root schedule is controlled by the logarithmic harmonic tail.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-4bc5d54e9a7f","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10371,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:117"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_Ico_sqrtSchedule_unconstrainedBranch_le_log (m n : Nat) (hm : 0 < m) (hmn : m <= n) (amplitude reciprocalGap : Real) (hreciprocalGap : 0 <= reciprocalGap) : (Finset.Ico m n).sum (fun t => (amplitude / Real.sqrt (((t + 1 : Nat) : Real))) ^ 2 / 4 * reciprocalGap) <= (amplitude ^ 2 / 4 * reciprocalGap) * Real.log ((n : Real) / (m : Real))","missing":[],"search":"sum_ico_sqrtschedule_unconstrainedbranch_le_log banditrlproof.tsallis.sum_ico_sqrtschedule_unconstrainedbranch_le_log the unconstrained tail branch of the square-root schedule is controlled by the logarithmic harmonic tail. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticSplit","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticSplit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticSplit","description":"The refined terminal self-bounding route under the concrete square-root schedule, reduced to two explicit filtered scalar sums.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-77702da734ee","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10372,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:151"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticSplit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (hselfBounding : let selector := canonicalHalfTsallisScheduleGeneratedSelectorMeasurability arms harms sampledScheduledHalfTsallisSqrtSchedule loss let mu := prior ⊗ₘ sampledS…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingquadraticsplit banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingquadraticsplit the refined terminal self-bounding route under the concrete square-root schedule, reduced to two explicit filtered scalar sums. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticPrefixSplit","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticPrefixSplit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticPrefixSplit","description":"Prefix/suffix specialization with a single scalar cutoff certificate.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-9321e9c01968","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10373,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:261"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticPrefixSplit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon cutoff : Nat) (hcutoff : cutoff ≤ horizon + 1) (gap : Action → Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (hcutoffThreshold : 2 * Real.sqrt ((arms.erase best).card : Real) * Real.sqrt (cutoff : Real) ≤ 5 * (1 + lambda) * (arms.erase be…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingquadraticprefixsplit banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingquadraticprefixsplit prefix/suffix specialization with a single scalar cutoff certificate. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticClosedForm","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticClosedForm","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticClosedForm","description":"Closed-form generated-regret endpoint after choosing a positive cutoff. The only remaining scalar optimization is the choice of `cutoff` and `lambda`.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingoptimization/index.html#decl-409260db1126","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","order":10374,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingOptimization.lean:410"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticClosedForm {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon cutoff : Nat) (hcutoffPos : 0 < cutoff) (hcutoff : cutoff <= horizon + 1) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (hcutoffThreshold : 2 * Real.sqrt ((arms.erase best).card : Real) * Real.sqrt (cutoff : Real) <= 5 * (…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingquadraticclosedform banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingquadraticclosedform closed-form generated-regret endpoint after choosing a positive cutoff. the only remaining scalar optimization is the choice of `cutoff` and `lambda`. theorem compiled","shard":"modules/011d9f7321b4e3dd.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalBetaEquation","label":"refinedLocalBetaEquation","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalBetaEquation","description":"The beta equation produced by the local coefficient-five envelope.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-f7e14565a887","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10375,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:19"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def refinedLocalBetaEquation (scale reciprocalGap corruption beta : Real) : Real","missing":[],"search":"refinedlocalbetaequation banditrlproof.tsallis.refinedlocalbetaequation the beta equation produced by the local coefficient-five envelope. definition compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.continuousOn_refinedLocalBetaEquation","label":"continuousOn_refinedLocalBetaEquation","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.continuousOn_refinedLocalBetaEquation","description":"The local beta equation is continuous on every interval bounded below by one.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-6a84b4ecb75c","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10376,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem continuousOn_refinedLocalBetaEquation (scale reciprocalGap corruption upper : Real) : ContinuousOn (refinedLocalBetaEquation scale reciprocalGap corruption) (Icc 1 upper)","missing":[],"search":"continuouson_refinedlocalbetaequation banditrlproof.tsallis.continuouson_refinedlocalbetaequation the local beta equation is continuous on every interval bounded below by one. theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sq_sqrt_sub_one_le_sub_log_sub_one","label":"sq_sqrt_sub_one_le_sub_log_sub_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sq_sqrt_sub_one_le_sub_log_sub_one","description":"The elementary inequality behind the Lambert-free beta estimate.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-1f9bdd62bf00","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10377,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:39"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sq_sqrt_sub_one_le_sub_log_sub_one (weight : Real) (hweight : 1 <= weight) : (Real.sqrt weight - 1) ^ 2 <= weight - Real.log weight - 1","missing":[],"search":"sq_sqrt_sub_one_le_sub_log_sub_one banditrlproof.tsallis.sq_sqrt_sub_one_le_sub_log_sub_one the elementary inequality behind the lambert-free beta estimate. theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalBetaWeight_bounds_of_eq_zero","label":"refinedLocalBetaWeight_bounds_of_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalBetaWeight_bounds_of_eq_zero","description":"A root of the coefficient-five equation admits the same kind of elementary upper estimate as the paper's Lambert-W expression, with the offset correction appearing as `+1` under the outer square root.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-7e4b1ce85c98","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10378,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:56"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalBetaWeight_bounds_of_eq_zero (scale reciprocalGap corruption beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hcorruption : 0 < corruption) (hbeta : 1 <= beta) (hcorruptionUpper : corruption * reciprocalGap <= scale) (hroot : refinedLocalBetaEquation scale reciprocalGap corruption beta = 0) : let weight := corruption * reciprocalGap / scale * beta 1 <= weight ∧ weight <= (1 + Real.sqrt (Real.log (scale / (corruption * reciprocalGap)) + 1)) ^ 2","missing":[],"search":"refinedlocalbetaweight_bounds_of_eq_zero banditrlproof.tsallis.refinedlocalbetaweight_bounds_of_eq_zero a root of the coefficient-five equation admits the same kind of elementary upper estimate as the paper's lambert-w expression, with the offset correction appearing as `+1` under the outer square root. theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_refinedLocalBetaEquation_eq_zero","label":"exists_refinedLocalBetaEquation_eq_zero","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_refinedLocalBetaEquation_eq_zero","description":"Under the coefficient-aware corruption window, the corrected beta equation has a root in the interval that corresponds to `alpha` between the local horizon threshold and one.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-fa593730f5cc","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10379,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:125"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_refinedLocalBetaEquation_eq_zero (scale reciprocalGap corruption : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hupper : 2 <= scale / (25 * reciprocalGap ^ 2)) (hcorruptionUpper : 2 * (corruption * reciprocalGap) <= scale) (hcorruptionLower : 25 * reciprocalGap * (Real.log (scale / (25 * reciprocalGap ^ 2)) + 2) <= corruption) : exists beta, beta ∈ Icc 2 (scale / (25 * reciprocalGap ^ 2)) ∧ refinedLocalBetaEquation scale reciprocalGap corruption beta = 0","missing":[],"search":"exists_refinedlocalbetaequation_eq_zero banditrlproof.tsallis.exists_refinedlocalbetaequation_eq_zero under the coefficient-aware corruption window, the corrected beta equation has a root in the interval that corresponds to `alpha` between the local horizon threshold and one. theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_refinedLocalBetaEquation_eq_zero_and_weight_bounds","label":"exists_refinedLocalBetaEquation_eq_zero_and_weight_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_refinedLocalBetaEquation_eq_zero_and_weight_bounds","description":"The coefficient-aware corruption window yields a beta root together with the elementary quantitative weight estimate needed by the refined regret route.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-b5380737531d","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10380,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:178"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_refinedLocalBetaEquation_eq_zero_and_weight_bounds (scale reciprocalGap corruption : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hcorruption : 0 < corruption) (hupper : 2 <= scale / (25 * reciprocalGap ^ 2)) (hcorruptionUpper : 2 * (corruption * reciprocalGap) <= scale) (hcorruptionLower : 25 * reciprocalGap * (Real.log (scale / (25 * reciprocalGap ^ 2)) + 2) <= corruption) : exists beta, beta ∈ Icc 2 (scale / (25 * reciprocalGap ^ 2)) ∧ refinedLocalBetaEquation scale reciprocalGap corruption beta = 0 ∧ let weight := corruption * reciprocalGap / scale * beta 1 <= weight ∧ weight <= (1 + Real.sqrt (Real.log (scale / (corruption * reciprocalGap)) + 1)) ^ 2","missing":[],"search":"exists_refinedlocalbetaequation_eq_zero_and_weight_bounds banditrlproof.tsallis.exists_refinedlocalbetaequation_eq_zero_and_weight_bounds the coefficient-aware corruption window yields a beta root together with the elementary quantitative weight estimate needed by the refined regret route. theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha","label":"refinedLocalAlpha","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalAlpha","description":"The coefficient-aware change of variables from beta to alpha.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-906122e8877f","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10381,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:208"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def refinedLocalAlpha (scale reciprocalGap beta : Real) : Real","missing":[],"search":"refinedlocalalpha banditrlproof.tsallis.refinedlocalalpha the coefficient-aware change of variables from beta to alpha. definition compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalLambda","label":"refinedLocalLambda","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalLambda","description":"The inverse of `alpha = 2 * lambda / (1 + lambda)`.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-8d0043734215","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10382,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:213"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def refinedLocalLambda (scale reciprocalGap beta : Real) : Real","missing":[],"search":"refinedlocallambda banditrlproof.tsallis.refinedlocallambda the inverse of `alpha = 2 * lambda / (1 + lambda)`. definition compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha_sq","label":"refinedLocalAlpha_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalAlpha_sq","description":"theorem refinedLocalAlpha_sq (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hbeta : 0 <= beta) : refinedLocalAlpha scale reciprocalGap beta ^ 2 = 25 * reciprocalGap ^ 2 * beta / scale","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-5c6cb818a339","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10383,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:218"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalAlpha_sq (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hbeta : 0 <= beta) : refinedLocalAlpha scale reciprocalGap beta ^ 2 = 25 * reciprocalGap ^ 2 * beta / scale","missing":[],"search":"refinedlocalalpha_sq banditrlproof.tsallis.refinedlocalalpha_sq theorem refinedlocalalpha_sq (scale reciprocalgap beta : real) (hscale : 0 < scale) (hbeta : 0 <= beta) : refinedlocalalpha scale reciprocalgap beta ^ 2 = 25 * reciprocalgap ^ 2 * beta / scale theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha_pos","label":"refinedLocalAlpha_pos","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalAlpha_pos","description":"theorem refinedLocalAlpha_pos (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hbeta : 0 < beta) : 0 < refinedLocalAlpha scale reciprocalGap beta","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-4d833b86cc99","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10384,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:227"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalAlpha_pos (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hbeta : 0 < beta) : 0 < refinedLocalAlpha scale reciprocalGap beta","missing":[],"search":"refinedlocalalpha_pos banditrlproof.tsallis.refinedlocalalpha_pos theorem refinedlocalalpha_pos (scale reciprocalgap beta : real) (hscale : 0 < scale) (hreciprocalgap : 0 < reciprocalgap) (hbeta : 0 < beta) : 0 < refinedlocalalpha scale reciprocalgap beta theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha_le_one","label":"refinedLocalAlpha_le_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalAlpha_le_one","description":"theorem refinedLocalAlpha_le_one (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hbeta : beta <= scale / (25 * reciprocalGap ^ 2)) : refinedLocalAlpha scale reciprocalGap beta <= 1","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-b939926277af","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10385,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:236"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalAlpha_le_one (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hbeta : beta <= scale / (25 * reciprocalGap ^ 2)) : refinedLocalAlpha scale reciprocalGap beta <= 1","missing":[],"search":"refinedlocalalpha_le_one banditrlproof.tsallis.refinedlocalalpha_le_one theorem refinedlocalalpha_le_one (scale reciprocalgap beta : real) (hscale : 0 < scale) (hreciprocalgap : 0 < reciprocalgap) (hbeta : beta <= scale / (25 * reciprocalgap ^ 2)) : refinedlocalalpha scale reciprocalgap beta <= 1 theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalLambda_mem_Ioc","label":"refinedLocalLambda_mem_Ioc","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalLambda_mem_Ioc","description":"theorem refinedLocalLambda_mem_Ioc (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hbetaPos : 0 < beta) (hbetaUpper : beta <= scale / (25 * reciprocalGap ^ 2)) : refinedLocalLambda scale reciprocalGap beta ∈ Set.Ioc (0 : Real) 1","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-a08ecaa72d70","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10386,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:249"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalLambda_mem_Ioc (scale reciprocalGap beta : Real) (hscale : 0 < scale) (hreciprocalGap : 0 < reciprocalGap) (hbetaPos : 0 < beta) (hbetaUpper : beta <= scale / (25 * reciprocalGap ^ 2)) : refinedLocalLambda scale reciprocalGap beta ∈ Set.Ioc (0 : Real) 1","missing":[],"search":"refinedlocallambda_mem_ioc banditrlproof.tsallis.refinedlocallambda_mem_ioc theorem refinedlocallambda_mem_ioc (scale reciprocalgap beta : real) (hscale : 0 < scale) (hreciprocalgap : 0 < reciprocalgap) (hbetapos : 0 < beta) (hbetaupper : beta <= scale / (25 * reciprocalgap ^ 2)) : refinedlocallambda scale reciprocalgap beta ∈ set.ioc (0 : real) 1 theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.one_add_refinedLocalLambda_eq","label":"one_add_refinedLocalLambda_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.one_add_refinedLocalLambda_eq","description":"theorem one_add_refinedLocalLambda_eq (scale reciprocalGap beta : Real) (halpha : refinedLocalAlpha scale reciprocalGap beta < 2) : 1 + refinedLocalLambda scale reciprocalGap beta = 2 / (2 - refinedLocalAlpha scale reciprocalGap beta)","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-787d611c8ccb","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10387,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:269"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem one_add_refinedLocalLambda_eq (scale reciprocalGap beta : Real) (halpha : refinedLocalAlpha scale reciprocalGap beta < 2) : 1 + refinedLocalLambda scale reciprocalGap beta = 2 / (2 - refinedLocalAlpha scale reciprocalGap beta)","missing":[],"search":"one_add_refinedlocallambda_eq banditrlproof.tsallis.one_add_refinedlocallambda_eq theorem one_add_refinedlocallambda_eq (scale reciprocalgap beta : real) (halpha : refinedlocalalpha scale reciprocalgap beta < 2) : 1 + refinedlocallambda scale reciprocalgap beta = 2 / (2 - refinedlocalalpha scale reciprocalgap beta) theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.two_mul_refinedLocalLambda_div_one_add_eq_alpha","label":"two_mul_refinedLocalLambda_div_one_add_eq_alpha","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.two_mul_refinedLocalLambda_div_one_add_eq_alpha","description":"theorem two_mul_refinedLocalLambda_div_one_add_eq_alpha (scale reciprocalGap beta : Real) (halpha : refinedLocalAlpha scale reciprocalGap beta < 2) : 2 * refinedLocalLambda scale reciprocalGap beta / (1 + refinedLocalLambda scale reciprocalGap beta) = refinedLocalAlpha scale reciprocalGap beta","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedscalar/index.html#decl-2f37439e9a25","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","order":10388,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedScalar.lean:279"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem two_mul_refinedLocalLambda_div_one_add_eq_alpha (scale reciprocalGap beta : Real) (halpha : refinedLocalAlpha scale reciprocalGap beta < 2) : 2 * refinedLocalLambda scale reciprocalGap beta / (1 + refinedLocalLambda scale reciprocalGap beta) = refinedLocalAlpha scale reciprocalGap beta","missing":[],"search":"two_mul_refinedlocallambda_div_one_add_eq_alpha banditrlproof.tsallis.two_mul_refinedlocallambda_div_one_add_eq_alpha theorem two_mul_refinedlocallambda_div_one_add_eq_alpha (scale reciprocalgap beta : real) (halpha : refinedlocalalpha scale reciprocalgap beta < 2) : 2 * refinedlocallambda scale reciprocalgap beta / (1 + refinedlocallambda scale reciprocalgap beta) = refinedlocalalpha scale reciprocalgap beta theorem compiled","shard":"modules/fb68ef46c66a1727.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_one_div_lambda_mul_eq_div","label":"sum_one_div_lambda_mul_eq_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_one_div_lambda_mul_eq_div","description":"theorem sum_one_div_lambda_mul_eq_div {Action : Type u} [DecidableEq Action] (actions : Finset Action) (gap : Action -> Real) (lambda : Real) (hlambda : lambda ≠ 0) (hgap : ∀ action ∈ actions, gap action ≠ 0) : actions.sum (fun action => 1 / (lambda * gap action)) = actions.sum (fun action => 1 / gap action) / lambda","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html#decl-1c4159099196","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","order":10389,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean:18"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_one_div_lambda_mul_eq_div {Action : Type u} [DecidableEq Action] (actions : Finset Action) (gap : Action -> Real) (lambda : Real) (hlambda : lambda ≠ 0) (hgap : ∀ action ∈ actions, gap action ≠ 0) : actions.sum (fun action => 1 / (lambda * gap action)) = actions.sum (fun action => 1 / gap action) / lambda","missing":[],"search":"sum_one_div_lambda_mul_eq_div banditrlproof.tsallis.sum_one_div_lambda_mul_eq_div theorem sum_one_div_lambda_mul_eq_div {action : type u} [decidableeq action] (actions : finset action) (gap : action -> real) (lambda : real) (hlambda : lambda ≠ 0) (hgap : ∀ action ∈ actions, gap action ≠ 0) : actions.sum (fun action => 1 / (lambda * gap action)) = actions.sum (fun action => 1 / gap action) / lambda theorem compiled","shard":"modules/84246877a8b9e43a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingThreshold_refinedLocalLambda_eq","label":"sampledScheduledHalfTsallisSelfBoundingThreshold_refinedLocalLambda_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingThreshold_refinedLocalLambda_eq","description":"With the coefficient-aware alpha/lambda change of variables, the actual continuous threshold used by the generated theorem is exactly `2 * (horizon + 1) / beta`.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html#decl-34f0d7e19170","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","order":10390,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean:37"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSelfBoundingThreshold_refinedLocalLambda_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (horizon : Nat) (beta : Real) (hbetaLower : 2 <= beta) (hbetaUpper : beta <= (2 * ((arms.erase best).card : Real) * (((horizon + 1 : Nat) : Real))) / (25 * ((arms.erase best).sum (fun action => 1 / gap action)) ^ 2)) : sampledScheduledHalfTsallisSelfBoundingThreshold arms best gap (refinedLocalLambda (2 * ((arms.erase best).card : Real) * (((horizon + 1 : Nat) : Real))) ((arms.erase best).sum (fun action => 1 / gap action)) beta) = (2 * (((horizon + 1 : Nat) : Real))) / beta","missing":[],"search":"sampledscheduledhalftsallisselfboundingthreshold_refinedlocallambda_eq banditrlproof.tsallis.sampledscheduledhalftsallisselfboundingthreshold_refinedlocallambda_eq with the coefficient-aware alpha/lambda change of variables, the actual continuous threshold used by the generated theorem is exactly `2 * (horizon + 1) / beta`. theorem compiled","shard":"modules/84246877a8b9e43a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalTunedRegretBound","label":"refinedLocalTunedRegretBound","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalTunedRegretBound","description":"The generated-regret scalar bound after substituting the coefficient-aware beta-dependent learning-rate multiplier.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html#decl-4a8f6217c04e","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","order":10391,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean:146"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def refinedLocalTunedRegretBound (scale horizonMass reciprocalGap corruption beta : Real) : Real","missing":[],"search":"refinedlocaltunedregretbound banditrlproof.tsallis.refinedlocaltunedregretbound the generated-regret scalar bound after substituting the coefficient-aware beta-dependent learning-rate multiplier. definition compiled","shard":"modules/84246877a8b9e43a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalTunedRegretBound_le_explicit","label":"refinedLocalTunedRegretBound_le_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalTunedRegretBound_le_explicit","description":"The Lambert-free explicit estimate for the local tuned scalar bound.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html#decl-d66bce02bc2d","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","order":10392,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean:155"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalTunedRegretBound_le_explicit (scale horizonMass reciprocalGap corruption beta : Real) (hscale : 0 < scale) (hhorizonMass : 1 <= horizonMass) (hreciprocalGap : 0 < reciprocalGap) (hcorruption : 0 < corruption) (hbeta : 1 <= beta) (hbetaUpper : beta <= scale / (25 * reciprocalGap ^ 2)) (hcorruptionUpper : corruption * reciprocalGap <= scale) (hroot : refinedLocalBetaEquation scale reciprocalGap corruption beta = 0) : refinedLocalTunedRegretBound scale horizonMass reciprocalGap corruption beta <= 1 + Real.log horizonMass + 10 * Real.sqrt (corruption * reciprocalGap) * (2 + Real.sqrt (Real.log (scale / (corruption * reciprocalGap)) + 1))","missing":[],"search":"refinedlocaltunedregretbound_le_explicit banditrlproof.tsallis.refinedlocaltunedregretbound_le_explicit the lambert-free explicit estimate for the local tuned scalar bound. theorem compiled","shard":"modules/84246877a8b9e43a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.exists_integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalTuned","label":"exists_integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalTuned","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.exists_integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalTuned","description":"The coefficient-aware scalar root now drives the actual generated theorem: it constructs the beta-dependent lambda, discharges the floor-threshold window, and rewrites the logarithmic tail as `log beta`.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html#decl-b755c9cba979","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","order":10393,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean:348"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem exists_integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalTuned {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption : Real) (hcorruption : 0 < corruption) (hscalarLower : 2 <= (2 * ((arms.erase best).card : Real) * (((horizon + 1 : Nat) : Real))) / (25 * ((arms.erase best).sum (fun action => 1 / gap action)) ^ 2)) (hscalarThresholdOne : (2 * ((arms.erase best).card :…","missing":[],"search":"exists_integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedlocaltuned banditrlproof.tsallis.exists_integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedlocaltuned the coefficient-aware scalar root now drives the actual generated theorem: it constructs the beta-dependent lambda, discharges the floor-threshold window, and rewrites the logarithmic tail as `log beta`. theorem compiled","shard":"modules/84246877a8b9e43a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalExplicit","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalExplicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalExplicit","description":"The coefficient-aware generated route with the auxiliary root eliminated. The constants reflect the local floor theorem's amplitude `5 * (1 + lambda)`; this is intentionally not presented as the paper's sharper constant.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedtuning/index.html#decl-ca4e9ab32735","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","order":10394,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedTuning.lean:527"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalExplicit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption : Real) (hcorruption : 0 < corruption) (hscalarLower : 2 <= (2 * ((arms.erase best).card : Real) * (((horizon + 1 : Nat) : Real))) / (25 * ((arms.erase best).sum (fun action => 1 / gap action)) ^ 2)) (hscalarThresholdOne : (2 * ((arms.erase best).card : Rea…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedlocalexplicit banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_refinedlocalexplicit the coefficient-aware generated route with the auxiliary root eliminated. the constants reflect the local floor theorem's amplitude `5 * (1 + lambda)`; this is intentionally not presented as the paper's sharper constant. theorem compiled","shard":"modules/84246877a8b9e43a.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.RefinedLocalCorruptionWindow","label":"RefinedLocalCorruptionWindow","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.RefinedLocalCorruptionWindow","description":"A compact sufficient window for the coefficient-aware refined optimizer. Here `armCount` is the number of suboptimal arms, `horizonMass = T + 1`, `reciprocalGap` is the sum of inverse gaps, and `corruption` is the self-bound allowance.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedwindow/index.html#decl-2c125eae0ac5","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","order":10395,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedWindow.lean:13"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"def RefinedLocalCorruptionWindow (armCount horizonMass reciprocalGap corruption : Real) : Prop","missing":[],"search":"refinedlocalcorruptionwindow banditrlproof.tsallis.refinedlocalcorruptionwindow a compact sufficient window for the coefficient-aware refined optimizer. here `armcount` is the number of suboptimal arms, `horizonmass = t + 1`, `reciprocalgap` is the sum of inverse gaps, and `corruption` is the self-bound allowance. definition compiled","shard":"modules/962b76b0d7bfa6f7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.refinedLocalCorruptionWindow_scalar_bounds","label":"refinedLocalCorruptionWindow_scalar_bounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.refinedLocalCorruptionWindow_scalar_bounds","description":"The compact window plus `armCount <= reciprocalGap` supplies all scalar contracts used by the generated refined theorem.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingrefinedwindow/index.html#decl-e97300eb9668","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","order":10396,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingRefinedWindow.lean:25"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem refinedLocalCorruptionWindow_scalar_bounds (armCount horizonMass reciprocalGap corruption : Real) (harmCount : 1 <= armCount) (hhorizonMass : 0 < horizonMass) (hreciprocalGap : 0 < reciprocalGap) (hcountGap : armCount <= reciprocalGap) (hwindow : RefinedLocalCorruptionWindow armCount horizonMass reciprocalGap corruption) : 2 <= (2 * armCount * horizonMass) / (25 * reciprocalGap ^ 2) ∧ (2 * armCount * horizonMass) / (25 * reciprocalGap ^ 2) <= 2 * horizonMass ∧ 2 * (corruption * reciprocalGap) <= 2 * armCount * horizonMass ∧ 25 * reciprocalGap * (Real.log ((2 * armCount * horizonMass) / (25 * reciprocalGap ^ 2)) + 2) <= corruption ∧ 0 < corruption","missing":[],"search":"refinedlocalcorruptionwindow_scalar_bounds banditrlproof.tsallis.refinedlocalcorruptionwindow_scalar_bounds the compact window plus `armcount <= reciprocalgap` supplies all scalar contracts used by the generated refined theorem. theorem compiled","shard":"modules/962b76b0d7bfa6f7.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.natFloor_positive_and_half_le","label":"natFloor_positive_and_half_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.natFloor_positive_and_half_le","description":"A positive real threshold and its natural floor differ by at most a factor of two once the threshold is at least one.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-5fa5876144b4","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10397,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:23"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem natFloor_positive_and_half_le (q : Real) (hq : 1 <= q) : 0 < ⌊q⌋₊ ∧ ((⌊q⌋₊ : Nat) : Real) <= q ∧ q <= 2 * ((⌊q⌋₊ : Nat) : Real)","missing":[],"search":"natfloor_positive_and_half_le banditrlproof.tsallis.natfloor_positive_and_half_le a positive real threshold and its natural floor differ by at most a factor of two once the threshold is at least one. theorem compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingThreshold","label":"sampledScheduledHalfTsallisSelfBoundingThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingThreshold","description":"The continuous threshold whose floor is used for the active-prefix split.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-c57c27786d57","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10398,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:38"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisSelfBoundingThreshold {Action : Type u} [DecidableEq Action] (arms : Finset Action) (best : Action) (gap : Action -> Real) (lambda : Real) : Real","missing":[],"search":"sampledscheduledhalftsallisselfboundingthreshold banditrlproof.tsallis.sampledscheduledhalftsallisselfboundingthreshold the continuous threshold whose floor is used for the active-prefix split. definition compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingFloorCutoff","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingFloorCutoff","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingFloorCutoff","description":"In the large-horizon branch, flooring the continuous threshold removes the explicit cutoff from the refined generated-regret theorem.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-e390442812bc","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10399,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:50"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingFloorCutoff {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption lambda : Real) (hlambda : lambda ∈ Set.Ioc (0 : Real) 1) (hthresholdOne : 1 <= sampledScheduledHalfTsallisSelfBoundingThreshold arms best gap lambda) (hthresholdHorizon : sampledScheduledHalfTsallisSelfBoundingThreshold arms best gap…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingfloorcutoff banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_refinedselfboundingfloorcutoff in the large-horizon branch, flooring the continuous threshold removes the explicit cutoff from the refined generated-regret theorem. theorem compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingOneThreshold","label":"sampledScheduledHalfTsallisSelfBoundingOneThreshold","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingOneThreshold","description":"The `lambda = 1` continuous threshold, written with the ordinary reciprocal-gap mass.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-b31aad05b843","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10400,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:235"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def sampledScheduledHalfTsallisSelfBoundingOneThreshold {Action : Type u} [DecidableEq Action] (arms : Finset Action) (best : Action) (gap : Action -> Real) : Real","missing":[],"search":"sampledscheduledhalftsallisselfboundingonethreshold banditrlproof.tsallis.sampledscheduledhalftsallisselfboundingonethreshold the `lambda = 1` continuous threshold, written with the ordinary reciprocal-gap mass. definition compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingOneThreshold_eq","label":"sampledScheduledHalfTsallisSelfBoundingOneThreshold_eq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingOneThreshold_eq","description":"The `lambda = 1` threshold has the expected reciprocal-gap/card form.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-bf172022e0d9","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10401,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:243"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sampledScheduledHalfTsallisSelfBoundingOneThreshold_eq {Action : Type u} [DecidableEq Action] (arms : Finset Action) (best : Action) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) : sampledScheduledHalfTsallisSelfBoundingOneThreshold arms best gap = 25 * ((arms.erase best).sum (fun action => 1 / gap action)) ^ 2 / ((arms.erase best).card : Real)","missing":[],"search":"sampledscheduledhalftsallisselfboundingonethreshold_eq banditrlproof.tsallis.sampledscheduledhalftsallisselfboundingonethreshold_eq the `lambda = 1` threshold has the expected reciprocal-gap/card form. theorem compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne","description":"A concrete large-horizon self-bounding theorem with `lambda = 1`. Unlike the refined corruption endpoint, this theorem needs no scalar optimization beyond checking that its continuous threshold lies in the horizon.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-a7506e29e8f9","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10402,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:266"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption : Real) (hthresholdOne : 1 <= sampledScheduledHalfTsallisSelfBoundingOneThreshold arms best gap) (hthresholdHorizon : sampledScheduledHalfTsallisSelfBoundingOneThreshold arms best gap <= ((horizon + 1 : Nat) : Real)) (hselfBounding : let selector :=…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_selfboundingone banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_selfboundingone a concrete large-horizon self-bounding theorem with `lambda = 1`. unlike the refined corruption endpoint, this theorem needs no scalar optimization beyond checking that its continuous threshold lies in the horizon. theorem compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne_explicit","label":"integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne_explicit","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne_explicit","description":"Fully explicit `lambda = 1` large-horizon bound, with the continuous threshold rewritten as `25 * S^2 / (K-1)`.","url":"../modules/banditrlproof-tsallissqrtscheduleselfboundingtuning/index.html#decl-eeb1c62726e5","parent":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","order":10403,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning"],["Source","BanditRLProof/TsallisSqrtScheduleSelfBoundingTuning.lean:347"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne_explicit {Env : Type u} {Action : Type v} [MeasurableSpace Env] [StandardBorelSpace Env] [MeasurableSpace Action] [MeasurableSingletonClass Action] [StandardBorelSpace Action] [Nonempty Action] [DecidableEq Action] (prior : Measure Env) [IsProbabilityMeasure prior] (arms : Finset Action) (harms : arms.Nonempty) (loss : Exp3.PredictableLossVector Env Action) {best : Action} (hbest : best ∈ arms) (horizon : Nat) (gap : Action -> Real) (hsuboptimal : (arms.erase best).Nonempty) (hgap : ∀ action ∈ arms.erase best, 0 < gap action) (corruption : Real) (hthresholdOne : 1 <= 25 * ((arms.erase best).sum (fun action => 1 / gap action)) ^ 2 / ((arms.erase best).card : Real)) (hthresholdHorizon : 25 * ((arms.erase best).sum (fun action => 1 / gap action)) ^ 2 / ((arms.erase best).card…","missing":[],"search":"integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_selfboundingone_explicit banditrlproof.tsallis.integral_sampledscheduledhalftsallispredictableenvironmentregret_pointmass_le_sqrtschedule_selfboundingone_explicit fully explicit `lambda = 1` large-horizon bound, with the continuous threshold rewritten as `25 * s^2 / (k-1)`. theorem compiled","shard":"modules/5744f46d69559543.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass","label":"halfTsallisPotentialMass","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialMass","description":"The paper-normalized regularizer mass carried by the local potential.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-83fb36900ab3","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10404,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:28"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisPotentialMass {Action : Type u} (arms : Finset Action) (p : Action -> Real) : Real","missing":[],"search":"halftsallispotentialmass banditrlproof.tsallis.halftsallispotentialmass the paper-normalized regularizer mass carried by the local potential. definition compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_eq_two_mul_powerSum_sub_one","label":"halfTsallisPotentialMass_eq_two_mul_powerSum_sub_one","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialMass_eq_two_mul_powerSum_sub_one","description":"The regularizer mass is `2 * sum sqrt(p_a) - 1`.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-2055973b460f","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10405,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:33"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialMass_eq_two_mul_powerSum_sub_one {Action : Type u} (arms : Finset Action) (p : Action -> Real) : halfTsallisPotentialMass arms p = 2 * powerSum arms (1 / 2 : Real) p - 1","missing":[],"search":"halftsallispotentialmass_eq_two_mul_powersum_sub_one banditrlproof.tsallis.halftsallispotentialmass_eq_two_mul_powersum_sub_one the regularizer mass is `2 * sum sqrt(p_a) - 1`. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_pointMass","label":"halfTsallisPotentialMass_pointMass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialMass_pointMass","description":"A supported point mass has paper-normalized potential mass one.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-1ddbb19672ac","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10406,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:42"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialMass_pointMass {Action : Type u} [DecidableEq Action] (arms : Finset Action) {best : Action} (hbest : best ∈ arms) : halfTsallisPotentialMass arms (pointMass best) = 1","missing":[],"search":"halftsallispotentialmass_pointmass banditrlproof.tsallis.halftsallispotentialmass_pointmass a supported point mass has paper-normalized potential mass one. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue_eq_neg_linearLoss_add_mass_div","label":"halfTsallisPotentialValue_eq_neg_linearLoss_add_mass_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialValue_eq_neg_linearLoss_add_mass_div","description":"The potential is linear loss with the regularizer mass divided by `eta`.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-4b51090a1595","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10407,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:51"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialValue_eq_neg_linearLoss_add_mass_div {Action : Type u} (arms : Finset Action) (eta : Real) (score p : Action -> Real) (heta : eta ≠ 0) : halfTsallisPotentialValue arms eta score p = -FTRL.linearLoss arms p score + halfTsallisPotentialMass arms p / eta","missing":[],"search":"halftsallispotentialvalue_eq_neg_linearloss_add_mass_div banditrlproof.tsallis.halftsallispotentialvalue_eq_neg_linearloss_add_mass_div the potential is linear loss with the regularizer mass divided by `eta`. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue_le_of_isRegularizedMinimizer","label":"halfTsallisPotentialValue_le_of_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialValue_le_of_isRegularizedMinimizer","description":"A regularized-objective minimizer maximizes the corresponding potential.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-44b0a270cc62","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10408,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:63"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialValue_le_of_isRegularizedMinimizer {Action : Type u} (arms : Finset Action) (eta : Real) (score p q : Action -> Real) (heta : 0 < eta) (hp : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) score p) (hq : FTRL.finiteSimplex arms q) : halfTsallisPotentialValue arms eta score q <= halfTsallisPotentialValue arms eta score p","missing":[],"search":"halftsallispotentialvalue_le_of_isregularizedminimizer banditrlproof.tsallis.halftsallispotentialvalue_le_of_isregularizedminimizer a regularized-objective minimizer maximizes the corresponding potential. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue_new_sub_old_le_rateChange_mul_mass","label":"halfTsallisPotentialValue_new_sub_old_le_rateChange_mul_mass","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialValue_new_sub_old_le_rateChange_mul_mass","description":"Changing the learning rate at a fixed score costs the reciprocal-rate increment times the mass of the new minimizer.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-3719b0c33982","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10409,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:85"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialValue_new_sub_old_le_rateChange_mul_mass {Action : Type u} (arms : Finset Action) (etaOld etaNew : Real) (score pOld pNew : Action -> Real) (hetaOld : 0 < etaOld) (hetaNew : 0 < etaNew) (hpOld : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms etaOld (negEntropyRegularizer arms (1 / 2 : Real)) score pOld) (hpNew : FTRL.finiteSimplex arms pNew) : halfTsallisPotentialValue arms etaNew score pNew - halfTsallisPotentialValue arms etaOld score pOld <= (1 / etaNew - 1 / etaOld) * halfTsallisPotentialMass arms pNew","missing":[],"search":"halftsallispotentialvalue_new_sub_old_le_ratechange_mul_mass banditrlproof.tsallis.halftsallispotentialvalue_new_sub_old_le_ratechange_mul_mass changing the learning rate at a fixed score costs the reciprocal-rate increment times the mass of the new minimizer. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_le_of_zero_isRegularizedMinimizer","label":"halfTsallisPotentialMass_le_of_zero_isRegularizedMinimizer","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisPotentialMass_le_of_zero_isRegularizedMinimizer","description":"A zero-score minimizer maximizes the half-Tsallis regularizer mass.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-d314ba7f3133","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10410,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:114"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem halfTsallisPotentialMass_le_of_zero_isRegularizedMinimizer {Action : Type u} (arms : Finset Action) (eta : Real) (p q : Action -> Real) (heta : 0 < eta) (hp : FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms eta (negEntropyRegularizer arms (1 / 2 : Real)) (fun _ => 0) p) (hq : FTRL.finiteSimplex arms q) : halfTsallisPotentialMass arms q <= halfTsallisPotentialMass arms p","missing":[],"search":"halftsallispotentialmass_le_of_zero_isregularizedminimizer banditrlproof.tsallis.halftsallispotentialmass_le_of_zero_isregularizedminimizer a zero-score minimizer maximizes the half-tsallis regularizer mass. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_range_succ_sub_eq_first_sub_last_add_cross","label":"sum_range_succ_sub_eq_first_sub_last_add_cross","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_range_succ_sub_eq_first_sub_last_add_cross","description":"Algebraic telescope with different left and right endpoint processes.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-ae10e6ccff85","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10411,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:133"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_range_succ_sub_eq_first_sub_last_add_cross (A B : Nat -> Real) (n : Nat) : (Finset.range (n + 1)).sum (fun t => A t - B t) = A 0 - B n + (Finset.range n).sum (fun t => A (t + 1) - B t)","missing":[],"search":"sum_range_succ_sub_eq_first_sub_last_add_cross banditrlproof.tsallis.sum_range_succ_sub_eq_first_sub_last_add_cross algebraic telescope with different left and right endpoint processes. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisScheduledMinimizer","label":"halfTsallisScheduledMinimizer","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisScheduledMinimizer","description":"Canonical scheduled minimizer before round `t`.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-776f88eaf1ac","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10412,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:145"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisScheduledMinimizer {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Nat -> Action -> Real) (t : Nat) : Action -> Real","missing":[],"search":"halftsallisscheduledminimizer banditrlproof.tsallis.halftsallisscheduledminimizer canonical scheduled minimizer before round `t`. definition compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.halfTsallisScheduledSameRateNext","label":"halfTsallisScheduledSameRateNext","kind":"definition","status":"compiled","subtitle":"BanditRLProof.Tsallis.halfTsallisScheduledSameRateNext","description":"Canonical same-rate auxiliary minimizer after appending round `t`.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-acdd5aed3ae4","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10413,"meta":[["Kind","definition"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:153"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"noncomputable def halfTsallisScheduledSameRateNext {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Nat -> Action -> Real) (t : Nat) : Action -> Real","missing":[],"search":"halftsallisscheduledsameratenext banditrlproof.tsallis.halftsallisscheduledsameratenext canonical same-rate auxiliary minimizer after appending round `t`. definition compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisScheduledPotentialPenalty_le","label":"sum_halfTsallisScheduledPotentialPenalty_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisScheduledPotentialPenalty_le","description":"Deterministic time-varying potential penalty with its terminal comparator contribution left explicit.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-82babbf2ab93","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10414,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:162"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisScheduledPotentialPenalty_le {Action : Type u} (arms : Finset Action) (eta : Nat -> Real) (loss : Nat -> Action -> Real) (current sameRateNext : Nat -> Action -> Real) (q : Action -> Real) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hcurrent : forall t, t <= n -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms (eta t) (negEntropyRegularizer arms (1 / 2 : Real)) (FTRL.cumulativeLoss loss t) (current t)) (hnext : forall t, t <= n -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms (eta t) (negEntropyRegularizer arms (1 / 2 : Real)) (FTRL.cumulativeLoss loss (t + 1)) (sameRateNext t)) (hq : FTRL.finiteSimplex arms q) : (Finset.range (n + 1)).sum (fun t => halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss t) (current t) - halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss (t + 1)) (sameRateNext t) - FTRL.linearLo…","missing":[],"search":"sum_halftsallisscheduledpotentialpenalty_le banditrlproof.tsallis.sum_halftsallisscheduledpotentialpenalty_le deterministic time-varying potential penalty with its terminal comparator contribution left explicit. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisScheduledPotentialPenalty_le_initial_sub_comparator_div","label":"sum_halfTsallisScheduledPotentialPenalty_le_initial_sub_comparator_div","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisScheduledPotentialPenalty_le_initial_sub_comparator_div","description":"Under a nonincreasing positive schedule, every rate-change mass is bounded by the initial zero-score mass, so the reciprocal-rate increments telescope.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-eae7d1bcf209","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10415,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:282"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisScheduledPotentialPenalty_le_initial_sub_comparator_div {Action : Type u} (arms : Finset Action) (eta : Nat -> Real) (loss : Nat -> Action -> Real) (current sameRateNext : Nat -> Action -> Real) (q : Action -> Real) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hetaMono : forall t, t < n -> eta (t + 1) <= eta t) (hcurrent : forall t, t <= n -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms (eta t) (negEntropyRegularizer arms (1 / 2 : Real)) (FTRL.cumulativeLoss loss t) (current t)) (hnext : forall t, t <= n -> FTRL.IsRegularizedMinimizer (FTRL.finiteSimplex arms) arms (eta t) (negEntropyRegularizer arms (1 / 2 : Real)) (FTRL.cumulativeLoss loss (t + 1)) (sameRateNext t)) (hq : FTRL.finiteSimplex arms q) : (Finset.range (n + 1)).sum (fun t => halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss t) (current t) - halfTsallisPotentialValue…","missing":[],"search":"sum_halftsallisscheduledpotentialpenalty_le_initial_sub_comparator_div banditrlproof.tsallis.sum_halftsallisscheduledpotentialpenalty_le_initial_sub_comparator_div under a nonincreasing positive schedule, every rate-change mass is bounded by the initial zero-score mass, so the reciprocal-rate increments telescope. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisCanonicalScheduledPotentialPenalty_le","label":"sum_halfTsallisCanonicalScheduledPotentialPenalty_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisCanonicalScheduledPotentialPenalty_le","description":"Canonical minimizer endpoint for the deterministic scheduled penalty.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-69ca8db8a52a","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10416,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:370"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisCanonicalScheduledPotentialPenalty_le {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Nat -> Action -> Real) (q : Action -> Real) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hetaMono : forall t, t < n -> eta (t + 1) <= eta t) (hq : FTRL.finiteSimplex arms q) : (Finset.range (n + 1)).sum (fun t => halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss t) (halfTsallisScheduledMinimizer arms harms eta loss t) - halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss (t + 1)) (halfTsallisScheduledSameRateNext arms harms eta loss t) - FTRL.linearLoss arms q (loss t)) <= (halfTsallisPotentialMass arms (halfTsallisScheduledMinimizer arms harms eta loss 0) - halfTsallisPotentialMass arms q) / eta n","missing":[],"search":"sum_halftsalliscanonicalscheduledpotentialpenalty_le banditrlproof.tsallis.sum_halftsalliscanonicalscheduledpotentialpenalty_le canonical minimizer endpoint for the deterministic scheduled penalty. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.Tsallis.sum_halfTsallisCanonicalScheduledPotentialPenalty_pointMass_le","label":"sum_halfTsallisCanonicalScheduledPotentialPenalty_pointMass_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.Tsallis.sum_halfTsallisCanonicalScheduledPotentialPenalty_pointMass_le","description":"Best-arm specialization retaining the explicit terminal `-1 / eta n` contribution encoded by the point-mass comparator.","url":"../modules/banditrlproof-tsallistimevaryingpenalty/index.html#decl-29f61978d6a3","parent":"module:BanditRLProof.TsallisTimeVaryingPenalty","order":10417,"meta":[["Kind","theorem"],["Module","BanditRLProof.TsallisTimeVaryingPenalty"],["Source","BanditRLProof/TsallisTimeVaryingPenalty.lean:404"],["Chapter","Tsallis-FTRL"],["Used in books","bandit, online-learning"],["Reading references","teaching:tsallis"],["Indexed settings","None registered"]],"statement":"theorem sum_halfTsallisCanonicalScheduledPotentialPenalty_pointMass_le {Action : Type u} [DecidableEq Action] (arms : Finset Action) (harms : arms.Nonempty) (eta : Nat -> Real) (loss : Nat -> Action -> Real) {best : Action} (hbest : best ∈ arms) (n : Nat) (heta : forall t, t <= n -> 0 < eta t) (hetaMono : forall t, t < n -> eta (t + 1) <= eta t) : (Finset.range (n + 1)).sum (fun t => halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss t) (halfTsallisScheduledMinimizer arms harms eta loss t) - halfTsallisPotentialValue arms (eta t) (FTRL.cumulativeLoss loss (t + 1)) (halfTsallisScheduledSameRateNext arms harms eta loss t) - FTRL.linearLoss arms (pointMass best) (loss t)) <= halfTsallisPotentialMass arms (halfTsallisScheduledMinimizer arms harms eta loss 0) / eta n - 1 / eta n","missing":[],"search":"sum_halftsalliscanonicalscheduledpotentialpenalty_pointmass_le banditrlproof.tsallis.sum_halftsalliscanonicalscheduledpotentialpenalty_pointmass_le best-arm specialization retaining the explicit terminal `-1 / eta n` contribution encoded by the point-mass comparator. theorem compiled","shard":"modules/c0d31421123d3a7b.json","books":["bandit","online-learning"],"chapters":["teaching:tsallis"],"settings":[]},{"id":"declaration:BanditRLProof.UCBSummability.finiteHorizonBadEvent","label":"finiteHorizonBadEvent","kind":"definition","status":"compiled","subtitle":"BanditRLProof.UCBSummability.finiteHorizonBadEvent","description":"The union of arm-time bad events over all finite arms and times `< T`.","url":"../modules/banditrlproof-ucbsummability/index.html#decl-f9a4f4e96f84","parent":"module:BanditRLProof.UCBSummability","order":10418,"meta":[["Kind","definition"],["Module","BanditRLProof.UCBSummability"],["Source","BanditRLProof/UCBSummability.lean:19"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"def finiteHorizonBadEvent {Omega : Type u} {Arm : Type v} (bad : Arm -> Nat -> Set Omega) (T : Nat) : Set Omega","missing":[],"search":"finitehorizonbadevent banditrlproof.ucbsummability.finitehorizonbadevent the union of arm-time bad events over all finite arms and times `< t`. definition compiled","shard":"modules/dde1574221a3812c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCBSummability.measure_finiteHorizonBadEvent_le_sum","label":"measure_finiteHorizonBadEvent_le_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCBSummability.measure_finiteHorizonBadEvent_le_sum","description":"Finite-horizon union bound for a UCB-style arm-time bad-event family. No event measurability is required: this is an outer-measure bound inherited from Mathlib's finite-union measure inequality.","url":"../modules/banditrlproof-ucbsummability/index.html#decl-d3d0e2905dc9","parent":"module:BanditRLProof.UCBSummability","order":10419,"meta":[["Kind","theorem"],["Module","BanditRLProof.UCBSummability"],["Source","BanditRLProof/UCBSummability.lean:30"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonBadEvent_le_sum {Omega : Type u} [MeasurableSpace Omega] {Arm : Type v} [Fintype Arm] (mu : Measure Omega) (bad : Arm -> Nat -> Set Omega) (T : Nat) : mu (finiteHorizonBadEvent bad T) <= (Finset.univ : Finset Arm).sum (fun a => (Finset.range T).sum (fun t => mu (bad a t)))","missing":[],"search":"measure_finitehorizonbadevent_le_sum banditrlproof.ucbsummability.measure_finitehorizonbadevent_le_sum finite-horizon union bound for a ucb-style arm-time bad-event family. no event measurability is required: this is an outer-measure bound inherited from mathlib's finite-union measure inequality. theorem compiled","shard":"modules/dde1574221a3812c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.UCBSummability.measure_finiteHorizonBadEvent_le_tail_sum","label":"measure_finiteHorizonBadEvent_le_tail_sum","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.UCBSummability.measure_finiteHorizonBadEvent_le_tail_sum","description":"Tail-bound consumer for finite-horizon UCB bad events. The hypothesis `htail` is the per-arm/per-time concentration result; this wrapper only assembles those local bounds into the finite bad-event sum.","url":"../modules/banditrlproof-ucbsummability/index.html#decl-9a11ec3f2a00","parent":"module:BanditRLProof.UCBSummability","order":10420,"meta":[["Kind","theorem"],["Module","BanditRLProof.UCBSummability"],["Source","BanditRLProof/UCBSummability.lean:58"],["Chapter","UCB"],["Used in books","bandit, reinforcement-learning"],["Reading references","teaching:ucb"],["Indexed settings","None registered"]],"statement":"theorem measure_finiteHorizonBadEvent_le_tail_sum {Omega : Type u} [MeasurableSpace Omega] {Arm : Type v} [Fintype Arm] (mu : Measure Omega) (bad : Arm -> Nat -> Set Omega) (tail : Arm -> Nat -> ENNReal) (T : Nat) (htail : forall a t, t < T -> mu (bad a t) <= tail a t) : mu (finiteHorizonBadEvent bad T) <= (Finset.univ : Finset Arm).sum (fun a => (Finset.range T).sum (fun t => tail a t))","missing":[],"search":"measure_finitehorizonbadevent_le_tail_sum banditrlproof.ucbsummability.measure_finitehorizonbadevent_le_tail_sum tail-bound consumer for finite-horizon ucb bad events. the hypothesis `htail` is the per-arm/per-time concentration result; this wrapper only assembles those local bounds into the finite bad-event sum. theorem compiled","shard":"modules/dde1574221a3812c.json","books":["bandit","reinforcement-learning"],"chapters":["teaching:ucb"],"settings":[]},{"id":"declaration:BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberMeasure_eq_lintegral_rounds_sq","label":"tsum_natSuccSquare_mul_stoppingFiberMeasure_eq_lintegral_rounds_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberMeasure_eq_lintegral_rounds_sq","description":"The squared successor-round count is the countable sum of its weighted equality fibers. This is an equality in `ENNReal`, so no integrability assumption is needed.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-0d689a030a01","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10421,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:26"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_natSuccSquare_mul_stoppingFiberMeasure_eq_lintegral_rounds_sq {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) : (∑' n : Nat, ((((n + 1) ^ 2 : Nat) : ENNReal) * mu {omega | tau omega = (n : WithTop Nat)})) = ∫⁻ omega, ENNReal.ofReal (((((tau omega).untopA + 1 : Nat) : Real)) ^ 2) ∂mu","missing":[],"search":"tsum_natsuccsquare_mul_stoppingfibermeasure_eq_lintegral_rounds_sq banditrlproof.tsum_natsuccsquare_mul_stoppingfibermeasure_eq_lintegral_rounds_sq the squared successor-round count is the countable sum of its weighted equality fibers. this is an equality in `ennreal`, so no integrability assumption is needed. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberMeasure_ne_top","label":"tsum_natSuccSquare_mul_stoppingFiberMeasure_ne_top","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberMeasure_ne_top","description":"The equality fibers of an a.e.-finite L2 stopping time have finite total mass after weighting by the squared successor index.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-279d269da080","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10422,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:89"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_natSuccSquare_mul_stoppingFiberMeasure_ne_top {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) : Ne (∑' n : Nat, ((((n + 1) ^ 2 : Nat) : ENNReal) * mu {omega | tau omega = (n : WithTop Nat)})) ∞","missing":[],"search":"tsum_natsuccsquare_mul_stoppingfibermeasure_ne_top banditrlproof.tsum_natsuccsquare_mul_stoppingfibermeasure_ne_top the equality fibers of an a.e.-finite l2 stopping time have finite total mass after weighting by the squared successor index. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberRealMeasure_eq_integral_rounds_sq","label":"tsum_natSuccSquare_mul_stoppingFiberRealMeasure_eq_integral_rounds_sq","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberRealMeasure_eq_integral_rounds_sq","description":"The real weighted fiber sum is exactly the second moment of the successor-round count.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-1839cf459d65","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10423,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:112"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_natSuccSquare_mul_stoppingFiberRealMeasure_eq_integral_rounds_sq {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) : (∑' n : Nat, (((n + 1 : Nat) : Real) ^ 2) * mu.real {omega | tau omega = (n : WithTop Nat)}) = integral mu (fun omega => ((((tau omega).untopA + 1 : Nat) : Real)) ^ 2)","missing":[],"search":"tsum_natsuccsquare_mul_stoppingfiberrealmeasure_eq_integral_rounds_sq banditrlproof.tsum_natsuccsquare_mul_stoppingfiberrealmeasure_eq_integral_rounds_sq the real weighted fiber sum is exactly the second moment of the successor-round count. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.summable_sqrt_stoppingFiberRealMeasure_of_memLp_two","label":"summable_sqrt_stoppingFiberRealMeasure_of_memLp_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.summable_sqrt_stoppingFiberRealMeasure_of_memLp_two","description":"The square roots of the real equality-fiber masses are summable for an a.e.-finite stopping time whose successor round count belongs to `L2`.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-22a3e6aa46a8","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10424,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:153"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem summable_sqrt_stoppingFiberRealMeasure_of_memLp_two {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) : Summable (fun n : Nat => Real.sqrt (mu.real {omega | tau omega = (n : WithTop Nat)}))","missing":[],"search":"summable_sqrt_stoppingfiberrealmeasure_of_memlp_two banditrlproof.summable_sqrt_stoppingfiberrealmeasure_of_memlp_two the square roots of the real equality-fiber masses are summable for an a.e.-finite stopping time whose successor round count belongs to `l2`. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.tsum_sqrt_stoppingFiberRealMeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natSuccSquare_of_memLp_two","label":"tsum_sqrt_stoppingFiberRealMeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natSuccSquare_of_memLp_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tsum_sqrt_stoppingFiberRealMeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natSuccSquare_of_memLp_two","description":"The square-root fiber-mass sum is quantitatively controlled by one half of the stopping-round second moment plus the universal inverse-square series. This is a fixed-stopping-time estimate, not a uniform family bound.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-57dea8a354fa","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10425,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:226"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_sqrt_stoppingFiberRealMeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natSuccSquare_of_memLp_two {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) : (∑' n : Nat, Real.sqrt (mu.real {omega | tau omega = (n : WithTop Nat)})) <= (1 / 2 : Real) * (integral mu (fun omega => ((((tau omega).untopA + 1 : Nat) : Real)) ^ 2) + ∑' n : Nat, 1 / (((n + 1 : Nat) : Real) ^ 2))","missing":[],"search":"tsum_sqrt_stoppingfiberrealmeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natsuccsquare_of_memlp_two banditrlproof.tsum_sqrt_stoppingfiberrealmeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natsuccsquare_of_memlp_two the square-root fiber-mass sum is quantitatively controlled by one half of the stopping-round second moment plus the universal inverse-square series. this is a fixed-stopping-time estimate, not a uniform family bound. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.tsum_sqrt_stoppingFiberRealMeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two","label":"tsum_sqrt_stoppingFiberRealMeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.tsum_sqrt_stoppingFiberRealMeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two","description":"Cauchy--Schwarz controls the square-root stopping-fiber masses by the actual successor-round second moment and the shifted inverse-square series. This is a fixed-stopping-time estimate, not a uniform family bound.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-c37b508c96a4","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10426,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:336"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem tsum_sqrt_stoppingFiberRealMeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) : (∑' n : Nat, Real.sqrt (mu.real {omega | tau omega = (n : WithTop Nat)})) <= Real.sqrt (integral mu (fun omega => ((((tau omega).untopA + 1 : Nat) : Real)) ^ 2)) * Real.sqrt (∑' n : Nat, 1 / (((n + 1 : Nat) : Real) ^ 2))","missing":[],"search":"tsum_sqrt_stoppingfiberrealmeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natsuccsquare_of_memlp_two banditrlproof.tsum_sqrt_stoppingfiberrealmeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natsuccsquare_of_memlp_two cauchy--schwarz controls the square-root stopping-fiber masses by the actual successor-round second moment and the shifted inverse-square series. this is a fixed-stopping-time estimate, not a uniform family bound. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.integrable_stoppedValue_of_uniform_secondMoment_of_memLp_two_rounds","label":"integrable_stoppedValue_of_uniform_secondMoment_of_memLp_two_rounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integrable_stoppedValue_of_uniform_secondMoment_of_memLp_two_rounds","description":"Uniform deterministic-coordinate second moments and an L2 finite stopping time make the corresponding unbounded stopped value integrable. The proof is a countable equality-fiber decomposition; it does not use optional stopping.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-08cd8375ffdb","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10427,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:437"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_stoppedValue_of_uniform_secondMoment_of_memLp_two_rounds {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) (process : Nat -> Omega -> Real) (hstoppedMeasurable : Measurable (stoppedValue process tau)) (secondMomentEnvelope : Real) (hprocessMemLp : ∀ n, MemLp (process n) 2 mu) (hprocessSecondMoment : ∀ n, integral mu (fun omega => process n omega ^ 2) <= secondMomentEnvelope) : Integrable (stoppedValue process tau) mu","missing":[],"search":"integrable_stoppedvalue_of_uniform_secondmoment_of_memlp_two_rounds banditrlproof.integrable_stoppedvalue_of_uniform_secondmoment_of_memlp_two_rounds uniform deterministic-coordinate second moments and an l2 finite stopping time make the corresponding unbounded stopped value integrable. the proof is a countable equality-fiber decomposition; it does not use optional stopping. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_tsum_sqrt_stoppingFiberRealMeasure_of_memLp_two_rounds","label":"integral_abs_stoppedValue_le_uniformSecondMoment_mul_tsum_sqrt_stoppingFiberRealMeasure_of_memLp_two_rounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_tsum_sqrt_stoppingFiberRealMeasure_of_memLp_two_rounds","description":"A quantitative version of the stopping-fiber transport: the absolute first moment of the stopped value is bounded by the uniform coordinate L2 envelope times the sum of square roots of the stopping-fiber masses. This is a fixed-stopping-time bound, not an index-uniform rate or optional stopping.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-baa20f916624","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10428,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:554"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_stoppedValue_le_uniformSecondMoment_mul_tsum_sqrt_stoppingFiberRealMeasure_of_memLp_two_rounds {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) (process : Nat -> Omega -> Real) (secondMomentEnvelope : Real) (hprocessMemLp : ∀ n, MemLp (process n) 2 mu) (hprocessSecondMoment : ∀ n, integral mu (fun omega => process n omega ^ 2) <= secondMomentEnvelope) : integral mu (fun omega => |stoppedValue process tau omega|) <= Real.sqrt secondMomentEnvelope * ∑' n : Nat, Real.sqrt (mu.real {omega | tau omega = (n : WithTop Nat)})","missing":[],"search":"integral_abs_stoppedvalue_le_uniformsecondmoment_mul_tsum_sqrt_stoppingfiberrealmeasure_of_memlp_two_rounds banditrlproof.integral_abs_stoppedvalue_le_uniformsecondmoment_mul_tsum_sqrt_stoppingfiberrealmeasure_of_memlp_two_rounds a quantitative version of the stopping-fiber transport: the absolute first moment of the stopped value is bounded by the uniform coordinate l2 envelope times the sum of square roots of the stopping-fiber masses. this is a fixed-stopping-time bound, not an index-uniform rate or optional stopping. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two_rounds","label":"integral_abs_stoppedValue_le_uniformSecondMoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two_rounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two_rounds","description":"The stopped-value first moment inherits the Cauchy--Schwarz stopping-fiber bound. The estimate is for one fixed stopping time and does not invoke optional stopping.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-944b35aba9b7","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10429,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:670"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_stoppedValue_le_uniformSecondMoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two_rounds {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) (process : Nat -> Omega -> Real) (secondMomentEnvelope : Real) (hprocessMemLp : ∀ n, MemLp (process n) 2 mu) (hprocessSecondMoment : ∀ n, integral mu (fun omega => process n omega ^ 2) <= secondMomentEnvelope) : integral mu (fun omega => |stoppedValue process tau omega|) <= Real.sqrt secondMomentEnvelope * (Real.sqrt (integral mu (fun omega => ((((tau omega).untopA + 1 : Nat) : Real)) ^ 2)) * Real.sqrt (∑' n : Nat, 1 / (((n + 1 : Nat) : Real) ^ 2)))","missing":[],"search":"integral_abs_stoppedvalue_le_uniformsecondmoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natsuccsquare_of_memlp_two_rounds banditrlproof.integral_abs_stoppedvalue_le_uniformsecondmoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natsuccsquare_of_memlp_two_rounds the stopped-value first moment inherits the cauchy--schwarz stopping-fiber bound. the estimate is for one fixed stopping time and does not invoke optional stopping. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_half_roundSecondMoment_add_inverseSquareTsum_of_memLp_two_rounds","label":"integral_abs_stoppedValue_le_uniformSecondMoment_mul_half_roundSecondMoment_add_inverseSquareTsum_of_memLp_two_rounds","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_half_roundSecondMoment_add_inverseSquareTsum_of_memLp_two_rounds","description":"The stopping-fiber absolute first-moment estimate with its fiber sum eliminated in favor of the actual successor-round second moment and the universal inverse-square series.","url":"../modules/banditrlproof-unboundedstoppingtimel2coordinateintegrability/index.html#decl-854b4332f455","parent":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","order":10430,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeL2CoordinateIntegrability.lean:717"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integral_abs_stoppedValue_le_uniformSecondMoment_mul_half_roundSecondMoment_add_inverseSquareTsum_of_memLp_two_rounds {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (hrounds : MemLp (fun omega => (((tau omega).untopA + 1 : Nat) : Real)) 2 mu) (process : Nat -> Omega -> Real) (secondMomentEnvelope : Real) (hprocessMemLp : ∀ n, MemLp (process n) 2 mu) (hprocessSecondMoment : ∀ n, integral mu (fun omega => process n omega ^ 2) <= secondMomentEnvelope) : integral mu (fun omega => |stoppedValue process tau omega|) <= Real.sqrt secondMomentEnvelope * ((1 / 2 : Real) * (integral mu (fun omega => ((((tau omega).untopA + 1 : Nat) : Real)) ^ 2) + ∑' n : Nat, 1 / (((n + 1 : Nat) : Real) ^ 2)))","missing":[],"search":"integral_abs_stoppedvalue_le_uniformsecondmoment_mul_half_roundsecondmoment_add_inversesquaretsum_of_memlp_two_rounds banditrlproof.integral_abs_stoppedvalue_le_uniformsecondmoment_mul_half_roundsecondmoment_add_inversesquaretsum_of_memlp_two_rounds the stopping-fiber absolute first-moment estimate with its fiber sum eliminated in favor of the actual successor-round second moment and the universal inverse-square series. theorem compiled","shard":"modules/b61978be8cbce3c1.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.measurable_stoppedValue_of_measurable_coordinates","label":"measurable_stoppedValue_of_measurable_coordinates","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.measurable_stoppedValue_of_measurable_coordinates","description":"A stopped value is measurable when the stopping index and every deterministic coordinate are measurable. The proof decomposes the dynamic evaluation into countably many natural-number fibers.","url":"../modules/banditrlproof-unboundedstoppingtimeweightedl2coordinateintegrability/index.html#decl-a6937f3d7f25","parent":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","order":10431,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeWeightedL2CoordinateIntegrability.lean:24"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem measurable_stoppedValue_of_measurable_coordinates {Omega : Type u} [MeasurableSpace Omega] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (process : Nat -> Omega -> Real) (hprocess : forall n, Measurable (process n)) : Measurable (stoppedValue process tau)","missing":[],"search":"measurable_stoppedvalue_of_measurable_coordinates banditrlproof.measurable_stoppedvalue_of_measurable_coordinates a stopped value is measurable when the stopping index and every deterministic coordinate are measurable. the proof decomposes the dynamic evaluation into countably many natural-number fibers. theorem compiled","shard":"modules/ddbe6f306a28ebff.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.summable_abs_weight_mul_sqrt_stoppingFiberRealMeasure_and_tsum_le","label":"summable_abs_weight_mul_sqrt_stoppingFiberRealMeasure_and_tsum_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.summable_abs_weight_mul_sqrt_stoppingFiberRealMeasure_and_tsum_le","description":"Square-summable deterministic weights are summable against the square roots of the real masses of measurable stopping fibers.","url":"../modules/banditrlproof-unboundedstoppingtimeweightedl2coordinateintegrability/index.html#decl-f2f715d2f565","parent":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","order":10432,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeWeightedL2CoordinateIntegrability.lean:53"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem summable_abs_weight_mul_sqrt_stoppingFiberRealMeasure_and_tsum_le {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (weight : Nat -> Real) (hweightSq : Summable (fun n => weight n ^ 2)) : Summable (fun n => |weight n| * Real.sqrt (mu.real {omega | tau omega = (n : WithTop Nat)})) /\\ (∑' n : Nat, |weight n| * Real.sqrt (mu.real {omega | tau omega = (n : WithTop Nat)})) <= Real.sqrt (∑' n : Nat, weight n ^ 2) * Real.sqrt (mu.real Set.univ)","missing":[],"search":"summable_abs_weight_mul_sqrt_stoppingfiberrealmeasure_and_tsum_le banditrlproof.summable_abs_weight_mul_sqrt_stoppingfiberrealmeasure_and_tsum_le square-summable deterministic weights are summable against the square roots of the real masses of measurable stopping fibers. theorem compiled","shard":"modules/ddbe6f306a28ebff.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"declaration:BanditRLProof.integrable_and_integral_abs_stoppedValue_weight_mul_le","label":"integrable_and_integral_abs_stoppedValue_weight_mul_le","kind":"theorem","status":"compiled","subtitle":"BanditRLProof.integrable_and_integral_abs_stoppedValue_weight_mul_le","description":"Uniform deterministic-coordinate second moments and a square-summable deterministic weight make the weighted stopped value integrable and bound its absolute first moment.","url":"../modules/banditrlproof-unboundedstoppingtimeweightedl2coordinateintegrability/index.html#decl-2a4d1c665cca","parent":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","order":10433,"meta":[["Kind","theorem"],["Module","BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability"],["Source","BanditRLProof/UnboundedStoppingTimeWeightedL2CoordinateIntegrability.lean:130"],["Chapter","Probability layer"],["Used in books","bandit, reinforcement-learning, online-learning"],["Reading references","teaching:probability"],["Indexed settings","None registered"]],"statement":"theorem integrable_and_integral_abs_stoppedValue_weight_mul_le {Omega : Type u} [MeasurableSpace Omega] (mu : Measure Omega) [IsFiniteMeasure mu] (tau : Omega -> WithTop Nat) (htau : Measurable tau) (hfinite : ∀ᵐ omega ∂mu, tau omega ≠ ⊤) (process : Nat -> Omega -> Real) (weight : Nat -> Real) (hweightSq : Summable (fun n => weight n ^ 2)) (hstoppedMeasurable : Measurable (stoppedValue (fun n omega => weight n * process n omega) tau)) (secondMomentEnvelope : Real) (hprocessMemLp : forall n, MemLp (process n) 2 mu) (hprocessSecondMoment : forall n, integral mu (fun omega => process n omega ^ 2) <= secondMomentEnvelope) : Integrable (stoppedValue (fun n omega => weight n * process n omega) tau) mu /\\ integral mu (fun omega => |stoppedValue (fun n omega => weight n * process n omega) tau omega|) <= Real.sqrt secondMomentEnvelope * (Real.sqrt (∑' n : Nat, weight n ^ 2) * Real.sqrt (mu.real…","missing":[],"search":"integrable_and_integral_abs_stoppedvalue_weight_mul_le banditrlproof.integrable_and_integral_abs_stoppedvalue_weight_mul_le uniform deterministic-coordinate second moments and a square-summable deterministic weight make the weighted stopped value integrable and bound its absolute first moment. theorem compiled","shard":"modules/ddbe6f306a28ebff.json","books":["bandit","reinforcement-learning","online-learning"],"chapters":["teaching:probability"],"settings":[]},{"id":"milestone:CAUSAL-PARALLEL-ACTUAL-REGRET","label":"Repaired parallel causal allocation and actual expected simple regret","kind":"mathematical milestone","status":"compiled","subtitle":"CAUSAL-PARALLEL-ACTUAL-REGRET","description":"For N>=2 known independent binary roots, the constructed strict rarity r and normalized allocation have actual design cost<=2r, including q=0/1. Actual graph-law transport gives expected regret <=(2sqrt(2)+7)sqrt(2r log(2TK)/T)+1/T for the repaired design and the attained optimal design, each using its own true cost in the threshold. Local fair/deterministic witnesses have exact costs 2/4. Independent source correct…","url":"../implementation-map/index.html#causal-parallel-actual-regret","parent":"group:milestones","order":0,"meta":[["Book Map chapter","ucb"],["Lean declarations","4"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"repaired parallel causal allocation and actual expected simple regret causal-parallel-actual-regret for n>=2 known independent binary roots, the constructed strict rarity r and normalized allocation have actual design cost<=2r, including q=0/1. actual graph-law transport gives expected regret <=(2sqrt(2)+7)sqrt(2r log(2tk)/t)+1/t for the repaired design and the attained optimal design, each using its own true cost in the threshold. local fair/deterministic witnesses have exact costs 2/4. independent source correction and semantic review accepted; source screening and iclr evidence remain open. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","label":"Causal actual-sampling expected simple regret on a common finite alphabet","kind":"mathematical milestone","status":"compiled","subtitle":"CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","description":"Actual intervention samples and fixed-order empirical recommendation satisfy R_T <= (2sqrt(2)+7)sqrt(m log(2TK)/T)+1/T and R_T<=1; explicit rate constant 3sqrt(2)+7. Internally derived confidence, binary reward readout and covered allocation. Independently reviewed common-alphabet scope; the separately reviewed native heterogeneous bridge is now available. Parallel allocation and boundary witnesses are now reviewed…","url":"../implementation-map/index.html#causal-common-alphabet-expected-regret","parent":"group:milestones","order":1,"meta":[["Book Map chapter","ucb"],["Lean declarations","4"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"causal actual-sampling expected simple regret on a common finite alphabet causal-common-alphabet-expected-regret actual intervention samples and fixed-order empirical recommendation satisfy r_t <= (2sqrt(2)+7)sqrt(m log(2tk)/t)+1/t and r_t<=1; explicit rate constant 3sqrt(2)+7. internally derived confidence, binary reward readout and covered allocation. independently reviewed common-alphabet scope; the separately reviewed native heterogeneous bridge is now available. parallel allocation and boundary witnesses are now reviewed separately. remaining source screening and iclr evidence remain open. the frozen noisy diagnostic supplement proves concentrated cost 8/3, exact biases, conditional/interventional distinction and tuned one-round expected regret 1/5. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HOO-REWARD-FAMILY-RATE","label":"HOO full expected regret for arbitrary reward families","kind":"mathematical milestone","status":"compiled","subtitle":"HOO-REWARD-FAMILY-RATE","description":"The same causal HOO trajectory with arbitrary arm-indexed probability reward laws satisfies the all-horizon source-repaired dimension rate, without global reward-kernel measurability. Actual and pseudo-regret expectations coincide. Fixed choices, log(max(N,2)) and other source repairs remain explicit; no whole-topic or efficiency claim.","url":"../implementation-map/index.html#hoo-reward-family-rate","parent":"group:milestones","order":2,"meta":[["Book Map chapter","ucb"],["Lean declarations","3"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"hoo full expected regret for arbitrary reward families hoo-reward-family-rate the same causal hoo trajectory with arbitrary arm-indexed probability reward laws satisfies the all-horizon source-repaired dimension rate, without global reward-kernel measurability. actual and pseudo-regret expectations coincide. fixed choices, log(max(n,2)) and other source repairs remain explicit; no whole-topic or efficiency claim. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HEAVY-TAIL-GENALTI-FINITE-SUPREMUM","label":"Finite raw-moment regret supremum and actual-law adapter","kind":"mathematical milestone","status":"compiled","subtitle":"HEAVY-TAIL-GENALTI-FINITE-SUPREMUM","description":"At each finite horizon, actual raw-moment arm laws and arbitrary probability trace laws yield normalized expected pseudo-regret at most2T; the EReal supremum is not positive infinity. The actual-process pushforward preserves the integral. A two-arm+1/-1 witness attains2T for epsilon1. This obstructs Genalti2024 Eq5 on positive scales, not asymptotic nonadaptivity or a fixed-algorithm lower bound.","url":"../implementation-map/index.html#heavy-tail-genalti-finite-supremum","parent":"group:milestones","order":3,"meta":[["Book Map chapter","ucb"],["Lean declarations","3"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"finite raw-moment regret supremum and actual-law adapter heavy-tail-genalti-finite-supremum at each finite horizon, actual raw-moment arm laws and arbitrary probability trace laws yield normalized expected pseudo-regret at most2t; the ereal supremum is not positive infinity. the actual-process pushforward preserves the integral. a two-arm+1/-1 witness attains2t for epsilon1. this obstructs genalti2024 eq5 on positive scales, not asymptotic nonadaptivity or a fixed-algorithm lower bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HEAVY-TAIL-PRINTED-REGRET-COUNTEREXAMPLE","label":"Finite counterexample to the printed source-policy regret coefficient","kind":"mathematical milestone","status":"compiled","subtitle":"HEAVY-TAIL-PRINTED-REGRET-COUNTEREXAMPLE","description":"For the unchanged source-radius4/r^-2 policy with its permissible deterministic ties, valid Dirac0/-1 arm laws and epsilon=u=1, expected pseudo-regret at T=2^50 exceeds32logT+5. The actual finite count contradiction, raw-second-moment witness, product-law expectation bridge and positive-gap-sum negation compile. This rejects that literal printed coefficient, not every paper estimator or logarithmic regret order.","url":"../implementation-map/index.html#heavy-tail-printed-regret-counterexample","parent":"group:milestones","order":4,"meta":[["Book Map chapter","ucb"],["Lean declarations","2"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"finite counterexample to the printed source-policy regret coefficient heavy-tail-printed-regret-counterexample for the unchanged source-radius4/r^-2 policy with its permissible deterministic ties, valid dirac0/-1 arm laws and epsilon=u=1, expected pseudo-regret at t=2^50 exceeds32logt+5. the actual finite count contradiction, raw-second-moment witness, product-law expectation bridge and positive-gap-sum negation compile. this rejects that literal printed coefficient, not every paper estimator or logarithmic regret order. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HEAVY-TAIL-CLIPPED-CORRUPTION-TRANSFER","label":"Clipped-prefix confidence under consumed-sample corruption","kind":"mathematical milestone","status":"compiled","subtitle":"HEAVY-TAIL-CLIPPED-CORRUPTION-TRANSFER","description":"Independent clean coordinates with raw p moments produce clipped-mean confidence at any outcome-dependent positive count N<=t, enlarged by C/N for a pathwise consumed-prefix corruption budget. Nonmeasurable count/adversary events use finite outer measure. Explicit estimator adaptation on one action trace, not a corruption-robust policy regret theorem.","url":"../implementation-map/index.html#heavy-tail-clipped-corruption-transfer","parent":"group:milestones","order":5,"meta":[["Book Map chapter","ucb"],["Lean declarations","3"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"clipped-prefix confidence under consumed-sample corruption heavy-tail-clipped-corruption-transfer independent clean coordinates with raw p moments produce clipped-mean confidence at any outcome-dependent positive count n<=t, enlarged by c/n for a pathwise consumed-prefix corruption budget. nonmeasurable count/adversary events use finite outer measure. explicit estimator adaptation on one action trace, not a corruption-robust policy regret theorem. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HEAVY-TAIL-CORRECTED-SOURCE-REGRET","label":"Corrected expected regret for source-parameter robust UCB","kind":"mathematical milestone","status":"compiled","subtitle":"HEAVY-TAIL-CORRECTED-SOURCE-REGRET","description":"For the unchanged radius4/r^-2 causal policy, raw p moments imply expected pseudo-regret at most sum over positive gaps of Delta*(A+5), A=2log(max(T,1))/(Delta/(8u^(1/p)))^(p/epsilon). This explicit corrected coefficient is128u/Delta at epsilon1; the printed32 coefficient now has a complete finite Lean counterexample for the instantiated source tie rule.","url":"../implementation-map/index.html#heavy-tail-corrected-source-regret","parent":"group:milestones","order":6,"meta":[["Book Map chapter","ucb"],["Lean declarations","2"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"corrected expected regret for source-parameter robust ucb heavy-tail-corrected-source-regret for the unchanged radius4/r^-2 causal policy, raw p moments imply expected pseudo-regret at most sum over positive gaps of delta*(a+5), a=2log(max(t,1))/(delta/(8u^(1/p)))^(p/epsilon). this explicit corrected coefficient is128u/delta at epsilon1; the printed32 coefficient now has a complete finite lean counterexample for the instantiated source tie rule. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HEAVY-TAIL-SOURCE-SCHEDULE-CONFIDENCE","label":"Source-schedule causal robust-UCB confidence budgets","kind":"mathematical milestone","status":"compiled","subtitle":"HEAVY-TAIL-SOURCE-SCHEDULE-CONFIDENCE","description":"The actual history-based robust UCB with source radius4 and paper-round confidence parameter r^(-2) has each per-arm signed deviation probability at most t exp(-5L_t/4), L_t=2log(t+1), after initialization. Each signed finite-time event-probability sum is at most2. This does not establish the printed regret coefficient.","url":"../implementation-map/index.html#heavy-tail-source-schedule-confidence","parent":"group:milestones","order":7,"meta":[["Book Map chapter","ucb"],["Lean declarations","2"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"source-schedule causal robust-ucb confidence budgets heavy-tail-source-schedule-confidence the actual history-based robust ucb with source radius4 and paper-round confidence parameter r^(-2) has each per-arm signed deviation probability at most t exp(-5l_t/4), l_t=2log(t+1), after initialization. each signed finite-time event-probability sum is at most2. this does not establish the printed regret coefficient. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","label":"Sample-index truncated confidence with source radius four","kind":"mathematical milestone","status":"compiled","subtitle":"HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","description":"For independent common-mean samples with bounded raw (1+epsilon)-moments, each closed signed deviation of the unchanged sample-index truncated mean at radius 4 u^(1/(1+epsilon)) (L/n)^(epsilon/(1+epsilon)) has probability at most exp(-5L/4), for L>0. This strengthens the source delta bound at L=log(1/delta); The actual source policy and explicitly corrected regret now have separate accepted results; the literal prin…","url":"../implementation-map/index.html#heavy-tail-source-confidence-four","parent":"group:milestones","order":8,"meta":[["Book Map chapter","probability"],["Lean declarations","4"],["Prerequisite milestones","4"]],"statement":"","missing":[],"search":"sample-index truncated confidence with source radius four heavy-tail-source-confidence-four for independent common-mean samples with bounded raw (1+epsilon)-moments, each closed signed deviation of the unchanged sample-index truncated mean at radius 4 u^(1/(1+epsilon)) (l/n)^(epsilon/(1+epsilon)) has probability at most exp(-5l/4), for l>0. this strengthens the source delta bound at l=log(1/delta); the actual source policy and explicitly corrected regret now have separate accepted results; the literal printed coefficient is refuted separately. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:FOUNDATION-REGRET-DECOMPOSITION","label":"Finite-arm pseudo-regret decomposition","kind":"mathematical milestone","status":"compiled","subtitle":"FOUNDATION-REGRET-DECOMPOSITION","description":"Finite-horizon pseudo-regret equals the sum over arms of each gap multiplied by its pull count.","url":"../implementation-map/index.html#foundation-regret-decomposition","parent":"group:milestones","order":9,"meta":[["Book Map chapter","foundations"],["Lean declarations","1"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"finite-arm pseudo-regret decomposition foundation-regret-decomposition finite-horizon pseudo-regret equals the sum over arms of each gap multiplied by its pull count. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:PROBABILITY-GENERATED-COND-MGF","label":"Generated-history conditional sub-Gaussian reward law","kind":"mathematical milestone","status":"compiled","subtitle":"PROBABILITY-GENERATED-COND-MGF","description":"A selected successor reward on the canonical generated trajectory inherits the centered conditional MGF bound supplied by its step kernel.","url":"../implementation-map/index.html#probability-generated-cond-mgf","parent":"group:milestones","order":10,"meta":[["Book Map chapter","probability"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"generated-history conditional sub-gaussian reward law probability-generated-cond-mgf a selected successor reward on the canonical generated trajectory inherits the centered conditional mgf bound supplied by its step kernel. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:PROBABILITY-FINTYPE-GEOMETRIC-ALL-TIME-UNION","label":"Finite-index geometric all-time confidence union","kind":"mathematical milestone","status":"compiled","subtitle":"PROBABILITY-FINTYPE-GEOMETRIC-ALL-TIME-UNION","description":"If every index in a nonempty finite family receives an equal geometric confidence share at every time, the outer measure of any time-index failure is at most the total confidence budget.","url":"../implementation-map/index.html#probability-fintype-geometric-all-time-union","parent":"group:milestones","order":11,"meta":[["Book Map chapter","probability"],["Lean declarations","2"],["Prerequisite milestones","2"]],"statement":"","missing":["Each ETC, UCB, or RL consumer must still supply its per-time, per-index tail events and model-specific law assumptions."],"search":"finite-index geometric all-time confidence union probability-fintype-geometric-all-time-union if every index in a nonempty finite family receives an equal geometric confidence share at every time, the outer measure of any time-index failure is at most the total confidence budget. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-GEOMETRIC-ALL-TIME","label":"Generated finite-arm empirical-mean all-time confidence","kind":"mathematical milestone","status":"compiled","subtitle":"PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-GEOMETRIC-ALL-TIME","description":"On one canonical generated action/reward trajectory, every finite arm and every positive successor horizon obeys the existing random-pull-count empirical-mean radius outside an event of outer measure at most the total geometric confidence budget.","url":"../implementation-map/index.html#probability-generated-fintype-empirical-mean-geometric-all-time","parent":"group:milestones","order":12,"meta":[["Book Map chapter","probability"],["Lean declarations","2"],["Prerequisite milestones","2"]],"statement":"","missing":["The ordinary-UCB chapter uses its exact finite-arm/time confidence producer. A fixed-policy anytime UCB consumer for this stronger geometric event remains a separate extension."],"search":"generated finite-arm empirical-mean all-time confidence probability-generated-fintype-empirical-mean-geometric-all-time on one canonical generated action/reward trajectory, every finite arm and every positive successor horizon obeys the existing random-pull-count empirical-mean radius outside an event of outer measure at most the total geometric confidence budget. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","label":"Generated finite-arm telescoping all-time empirical-mean confidence","kind":"mathematical milestone","status":"compiled","subtitle":"PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","description":"On one canonical generated action/reward trajectory, every finite arm and positive successor horizon obeys the existing random-pull-count empirical-mean radius at share delta/((n+1)(n+2))/|A| outside an event of outer measure at most delta.","url":"../implementation-map/index.html#probability-generated-fintype-empirical-mean-telescoping-all-time","parent":"group:milestones","order":13,"meta":[["Book Map chapter","probability"],["Lean declarations","4"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"generated finite-arm telescoping all-time empirical-mean confidence probability-generated-fintype-empirical-mean-telescoping-all-time on one canonical generated action/reward trajectory, every finite arm and positive successor horizon obeys the existing random-pull-count empirical-mean radius at share delta/((n+1)(n+2))/|a| outside an event of outer measure at most delta. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","label":"Fixed-policy telescoping anytime UCB confidence and regret","kind":"mathematical milestone","status":"compiled","subtitle":"UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","description":"One horizon-free scheduled UCB policy uses the telescoping per-round confidence share on its own generated finite history. On the same canonical action/reward trajectory measure, one all-time confidence event controls every positive-gap arm at every finite horizon, and each finite horizon has an explicit expected pseudo-regret bound.","url":"../implementation-map/index.html#ucb-fixed-policy-telescoping-anytime-regret","parent":"group:milestones","order":14,"meta":[["Book Map chapter","ucb"],["Lean declarations","5"],["Prerequisite milestones","2"]],"statement":"","missing":["The finite-time expectation keeps the explicit T times delta failure contribution; fixed-delta expected-average consistency is not claimed.","The compiled bounded KL-UCB extension is mapped separately; literal pinned-LML identity remains cross-toolchain work."],"search":"fixed-policy telescoping anytime ucb confidence and regret ucb-fixed-policy-telescoping-anytime-regret one horizon-free scheduled ucb policy uses the telescoping per-round confidence share on its own generated finite history. on the same canonical action/reward trajectory measure, one all-time confidence event controls every positive-gap arm at every finite horizon, and each finite horizon has an explicit expected pseudo-regret bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:PROBABILITY-ETC-UCB-ROUTE-SURFACE","label":"Probability interfaces used by canonical ETC and ordinary UCB","kind":"mathematical milestone","status":"compiled","subtitle":"PROBABILITY-ETC-UCB-ROUTE-SURFACE","description":"The generated-history law, conditional sub-Gaussian MGF, and finite-arm/time empirical-mean event actually required by the scoped ETC and horizon-indexed UCB routes compile through one external Book Map canary; the reusable countable adapter is exposed separately.","url":"../implementation-map/index.html#probability-etc-ucb-route-surface","parent":"group:milestones","order":15,"meta":[["Book Map chapter","probability"],["Lean declarations","3"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"probability interfaces used by canonical etc and ordinary ucb probability-etc-ucb-route-surface the generated-history law, conditional sub-gaussian mgf, and finite-arm/time empirical-mean event actually required by the scoped etc and horizon-indexed ucb routes compile through one external book map canary; the reusable countable adapter is exposed separately. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:ETC-CANONICAL-SUBGAUSSIAN-REGRET","label":"Canonical sub-Gaussian ETC expected regret","kind":"mathematical milestone","status":"compiled","subtitle":"ETC-CANONICAL-SUBGAUSSIAN-REGRET","description":"The generated ETC policy under finite-arm sub-Gaussian reward laws satisfies the explicit exploration-plus-wrong-commit expected pseudo-regret bound.","url":"../implementation-map/index.html#etc-canonical-subgaussian-regret","parent":"group:milestones","order":16,"meta":[["Book Map chapter","etc"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"canonical sub-gaussian etc expected regret etc-canonical-subgaussian-regret the generated etc policy under finite-arm sub-gaussian reward laws satisfies the explicit exploration-plus-wrong-commit expected pseudo-regret bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:ETC-CANONICAL-RAT-LEAST-TIE","label":"Canonical Rat ETC least-encoded tie rule","kind":"mathematical milestone","status":"compiled","subtitle":"ETC-CANONICAL-RAT-LEAST-TIE","description":"The Rat commit oracle used by the measurable generated ETC policy is Mathlib's first-occurrence argmax on Fin K, hence chooses the least encoded arm among tied maxima.","url":"../implementation-map/index.html#etc-canonical-rat-least-tie","parent":"group:milestones","order":17,"meta":[["Book Map chapter","etc"],["Lean declarations","2"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"canonical rat etc least-encoded tie rule etc-canonical-rat-least-tie the rat commit oracle used by the measurable generated etc policy is mathlib's first-occurrence argmax on fin k, hence chooses the least encoded arm among tied maxima. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:ETC-LML-PORT","label":"Local ETC endpoint aligned with the LML theorem card","kind":"mathematical milestone","status":"compiled","subtitle":"ETC-LML-PORT","description":"The local canonical ETC chapter route compiles; direct identity with the pinned LML declaration remains a separate cross-toolchain theorem-card route and is not claimed here.","url":"../implementation-map/index.html#etc-lml-port","parent":"group:milestones","order":18,"meta":[["Book Map chapter","etc"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":["Direct imported-LML symbol and toolchain identity remains cross-toolchain work."],"search":"local etc endpoint aligned with the lml theorem card etc-lml-port the local canonical etc chapter route compiles; direct identity with the pinned lml declaration remains a separate cross-toolchain theorem-card route and is not claimed here. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:UCB-FINITE-ARM-SUBGAUSSIAN","label":"Finite-arm sub-Gaussian UCB gap-sum bound","kind":"mathematical milestone","status":"compiled","subtitle":"UCB-FINITE-ARM-SUBGAUSSIAN","description":"The generated selected-policy UCB action satisfies a textbook-shaped Real pseudo-regret gap-sum bound, including the zero-proxy case.","url":"../implementation-map/index.html#ucb-finite-arm-subgaussian","parent":"group:milestones","order":19,"meta":[["Book Map chapter","ucb"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"finite-arm sub-gaussian ucb gap-sum bound ucb-finite-arm-subgaussian the generated selected-policy ucb action satisfies a textbook-shaped real pseudo-regret gap-sum bound, including the zero-proxy case. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:UCB-EXPECTED-AVERAGE-CONSISTENCY","label":"Finite-arm UCB expected-average consistency","kind":"mathematical milestone","status":"compiled","subtitle":"UCB-EXPECTED-AVERAGE-CONSISTENCY","description":"For the compiled finite-arm sub-Gaussian source, expected pseudo-regret divided by the horizon tends to zero.","url":"../implementation-map/index.html#ucb-expected-average-consistency","parent":"group:milestones","order":20,"meta":[["Book Map chapter","ucb"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"finite-arm ucb expected-average consistency ucb-expected-average-consistency for the compiled finite-arm sub-gaussian source, expected pseudo-regret divided by the horizon tends to zero. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:UCB-HORIZON-INDEXED-CANONICAL-CHAIN","label":"Canonical horizon-indexed UCB confidence-to-consistency chain","kind":"mathematical milestone","status":"compiled","subtitle":"UCB-HORIZON-INDEXED-CANONICAL-CHAIN","description":"For each explicit horizon and confidence budget, one generated pair-trajectory family carries the finite-arm/time confidence event through large-gap and pull-count bounds to finite-arm expected pseudo-regret; the scheduled family has vanishing expected average regret.","url":"../implementation-map/index.html#ucb-horizon-indexed-canonical-chain","parent":"group:milestones","order":21,"meta":[["Book Map chapter","ucb"],["Lean declarations","4"],["Prerequisite milestones","1"]],"statement":"","missing":["This is a horizon-indexed policy family, not a single fixed-policy anytime UCB theorem."],"search":"canonical horizon-indexed ucb confidence-to-consistency chain ucb-horizon-indexed-canonical-chain for each explicit horizon and confidence budget, one generated pair-trajectory family carries the finite-arm/time confidence event through large-gap and pull-count bounds to finite-arm expected pseudo-regret; the scheduled family has vanishing expected average regret. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:UCB-LML-PORT","label":"Pinned LML UCB theorem-card port","kind":"mathematical milestone","status":"partial","subtitle":"UCB-LML-PORT","description":"Local ordinary-UCB and Real arm-stream one-policy results compile; only literal identity with the pinned upstream declaration remains a separate theorem-card/cross-toolchain gate.","url":"../implementation-map/index.html#ucb-lml-port","parent":"group:milestones","order":22,"meta":[["Book Map chapter","ucb"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":["Direct imported-LML symbol and toolchain identity remain cross-toolchain work; the local expected pull-count ledger is already compiled.","Do not treat the upstream theorem card as a local proof term."],"search":"pinned lml ucb theorem-card port ucb-lml-port local ordinary-ucb and real arm-stream one-policy results compile; only literal identity with the pinned upstream declaration remains a separate theorem-card/cross-toolchain gate. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:LML-DIRECT-TOOLCHAIN-IDENTITY","label":"Direct LeanMachineLearning toolchain identity","kind":"mathematical milestone","status":"blocked","subtitle":"LML-DIRECT-TOOLCHAIN-IDENTITY","description":"ABRL has compiled local ETC/UCB theorems shaped against the pinned LeanMachineLearning cards, but it does not import or consume the actual upstream Bandits.ETC.regret_le or Bandits.UCB.regret_le declarations.","url":"../implementation-map/index.html#lml-direct-toolchain-identity","parent":"group:milestones","order":23,"meta":[["Book Map chapter","frontier"],["Lean declarations","0"],["Prerequisite milestones","0"]],"statement":"","missing":["Reconcile ABRL's Lean 4.29.1 and Mathlib v4.29.1 environment with the recorded LML seed's newer Lean/Mathlib toolchain in an isolated migration build.","Add a pinned LML dependency and compile the real LeanMachineLearning.Online.Bandit.Algorithms.ETC and UCB imports.","Consume the actual upstream symbols in ABRL wrapper theorems without copied or shadow declarations.","Pass the complete Lean, test, license, notice, attribution, and website gates on the unified toolchain."],"search":"direct leanmachinelearning toolchain identity lml-direct-toolchain-identity abrl has compiled local etc/ucb theorems shaped against the pinned leanmachinelearning cards, but it does not import or consume the actual upstream bandits.etc.regret_le or bandits.ucb.regret_le declarations. mathematical milestone blocked","shard":"views/milestones.json"},{"id":"milestone:OFUL-ELLIPTICAL-POTENTIAL","label":"Logarithmic elliptical-potential inequality","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-ELLIPTICAL-POTENTIAL","description":"Clipped inverse-Gram quadratic widths are bounded by a dimension-scaled log-determinant growth term.","url":"../implementation-map/index.html#oful-elliptical-potential","parent":"group:milestones","order":24,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"logarithmic elliptical-potential inequality oful-elliptical-potential clipped inverse-gram quadratic widths are bounded by a dimension-scaled log-determinant growth term. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-SELF-NORMALIZED-RIDGE-CONFIDENCE","label":"Conditional-MGF to ridge confidence ellipsoid","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-SELF-NORMALIZED-RIDGE-CONFIDENCE","description":"The finite-dimensional conditional-MGF and Gaussian-mixture route controls the ridge-estimation error in the regularized matrix norm.","url":"../implementation-map/index.html#oful-self-normalized-ridge-confidence","parent":"group:milestones","order":25,"meta":[["Book Map chapter","oful"],["Lean declarations","2"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"conditional-mgf to ridge confidence ellipsoid oful-self-normalized-ridge-confidence the finite-dimensional conditional-mgf and gaussian-mixture route controls the ridge-estimation error in the regularized matrix norm. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-MEASURABLE-GENERATED-POLICY","label":"Measurable horizon-free optimistic policy","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-MEASURABLE-GENERATED-POLICY","description":"A strict-fold finite-action selector turns ridge estimates and scheduled confidence radii into one measurable history algorithm without a terminal horizon parameter.","url":"../implementation-map/index.html#oful-measurable-generated-policy","parent":"group:milestones","order":26,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"measurable horizon-free optimistic policy oful-measurable-generated-policy a strict-fold finite-action selector turns ridge estimates and scheduled confidence radii into one measurable history algorithm without a terminal horizon parameter. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-ALL-TIME-CONFIDENCE","label":"One-policy all-time OFUL confidence","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-ALL-TIME-CONFIDENCE","description":"A telescoping confidence-budget schedule controls one countable failure event for the generated scalar-ridge policy at every deterministic horizon.","url":"../implementation-map/index.html#oful-all-time-confidence","parent":"group:milestones","order":27,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"one-policy all-time oful confidence oful-all-time-confidence a telescoping confidence-budget schedule controls one countable failure event for the generated scalar-ridge policy at every deterministic horizon. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-ALL-HORIZON-HIGH-PROBABILITY-REGRET","label":"One-policy all-horizon OFUL pseudo-regret","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-ALL-HORIZON-HIGH-PROBABILITY-REGRET","description":"On the same canonical trajectory generated by the horizon-free telescoping policy, one outer-measure event controls the explicit pseudo-regret bound for every finite horizon.","url":"../implementation-map/index.html#oful-all-horizon-high-probability-regret","parent":"group:milestones","order":28,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"one-policy all-horizon oful pseudo-regret oful-all-horizon-high-probability-regret on the same canonical trajectory generated by the horizon-free telescoping policy, one outer-measure event controls the explicit pseudo-regret bound for every finite horizon. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-EXPECTED-AVERAGE-CONSISTENCY","label":"Fixed-model OFUL expected-average consistency","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-EXPECTED-AVERAGE-CONSISTENCY","description":"For the separate horizon-indexed fixed-model policy family, the canonical expected pseudo-regret bound is little-o of the horizon, so the corresponding expected average tends to zero.","url":"../implementation-map/index.html#oful-expected-average-consistency","parent":"group:milestones","order":29,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"fixed-model oful expected-average consistency oful-expected-average-consistency for the separate horizon-indexed fixed-model policy family, the canonical expected pseudo-regret bound is little-o of the horizon, so the corresponding expected average tends to zero. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-BOUNDED-STOPPING-EXPECTED-REGRET","label":"Bounded stopping-time OFUL expected regret","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-BOUNDED-STOPPING-EXPECTED-REGRET","description":"For the same horizon-free telescoping generated policy, any stopping time bounded by a deterministic horizon has nonnegative expected stopped pseudo-regret bounded by the endpoint budget plus the explicit delta-weighted envelope.","url":"../implementation-map/index.html#oful-bounded-stopping-expected-regret","parent":"group:milestones","order":30,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"bounded stopping-time oful expected regret oful-bounded-stopping-expected-regret for the same horizon-free telescoping generated policy, any stopping time bounded by a deterministic horizon has nonnegative expected stopped pseudo-regret bounded by the endpoint budget plus the explicit delta-weighted envelope. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:OFUL-UNBOUNDED-STOPPING-EXPECTED-REGRET","label":"Square-integrable random-horizon OFUL expected regret","kind":"mathematical milestone","status":"compiled","subtitle":"OFUL-UNBOUNDED-STOPPING-EXPECTED-REGRET","description":"Under an explicit square-integrable finite stopping-time contract, the stopped high-probability pseudo-regret is integrable and receives a second-moment-controlled expectation bound.","url":"../implementation-map/index.html#oful-unbounded-stopping-expected-regret","parent":"group:milestones","order":31,"meta":[["Book Map chapter","oful"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"square-integrable random-horizon oful expected regret oful-unbounded-stopping-expected-regret under an explicit square-integrable finite stopping-time contract, the stopped high-probability pseudo-regret is integrable and receives a second-moment-controlled expectation bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-POSTERIOR-KERNEL","label":"Posterior kernel equals the conditional environment law","kind":"mathematical milestone","status":"compiled","subtitle":"THOMPSON-POSTERIOR-KERNEL","description":"When the observed environment-history pair has the canonical prior-likelihood joint law, the canonical posterior kernel is almost everywhere the conditional distribution of the environment given history.","url":"../implementation-map/index.html#thompson-posterior-kernel","parent":"group:milestones","order":32,"meta":[["Book Map chapter","thompson"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"posterior kernel equals the conditional environment law thompson-posterior-kernel when the observed environment-history pair has the canonical prior-likelihood joint law, the canonical posterior kernel is almost everywhere the conditional distribution of the environment given history. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-CANONICAL-SAMPLER","label":"Canonical one-step probability matching","kind":"mathematical milestone","status":"compiled","subtitle":"THOMPSON-CANONICAL-SAMPLER","description":"Sampling an environment from the canonical posterior and applying the measurable best-action selector gives the same conditional action law as the posterior best action.","url":"../implementation-map/index.html#thompson-canonical-sampler","parent":"group:milestones","order":33,"meta":[["Book Map chapter","thompson"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"canonical one-step probability matching thompson-canonical-sampler sampling an environment from the canonical posterior and applying the measurable best-action selector gives the same conditional action law as the posterior best action. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-RECURSIVE-PROBABILITY-MATCHING","label":"Probability matching on the actual recursive trajectory","kind":"mathematical milestone","status":"compiled","subtitle":"THOMPSON-RECURSIVE-PROBABILITY-MATCHING","description":"For every round, the successor action on the recursively generated Thompson trajectory conditioned on its own finite history has the posterior-best-action conditional law.","url":"../implementation-map/index.html#thompson-recursive-probability-matching","parent":"group:milestones","order":34,"meta":[["Book Map chapter","thompson"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"probability matching on the actual recursive trajectory thompson-recursive-probability-matching for every round, the successor action on the recursively generated thompson trajectory conditioned on its own finite history has the posterior-best-action conditional law. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-BAYES-CLIPPED-DECOMPOSITION","label":"Bayesian regret and clipped-UCB decomposition","kind":"mathematical milestone","status":"compiled","subtitle":"THOMPSON-BAYES-CLIPPED-DECOMPOSITION","description":"Probability matching transports finite-history scores so generated comparator-relative mean regret splits exactly into selector and selected-action clipped-score terms; it becomes Bayesian regret when the selector is mean-optimal.","url":"../implementation-map/index.html#thompson-bayes-clipped-decomposition","parent":"group:milestones","order":35,"meta":[["Book Map chapter","thompson"],["Lean declarations","2"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"bayesian regret and clipped-ucb decomposition thompson-bayes-clipped-decomposition probability matching transports finite-history scores so generated comparator-relative mean regret splits exactly into selector and selected-action clipped-score terms; it becomes bayesian regret when the selector is mean-optimal. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-LATENT-STREAM-SUPPORT","label":"Generated rewards align with the stationary latent arm stream","kind":"mathematical milestone","status":"compiled","subtitle":"THOMPSON-LATENT-STREAM-SUPPORT","description":"The actual recursive trajectory reward coordinates agree almost everywhere with the next-unused-coordinate reward read from the selected arm's latent stream.","url":"../implementation-map/index.html#thompson-latent-stream-support","parent":"group:milestones","order":36,"meta":[["Book Map chapter","thompson"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"generated rewards align with the stationary latent arm stream thompson-latent-stream-support the actual recursive trajectory reward coordinates agree almost everywhere with the next-unused-coordinate reward read from the selected arm's latent stream. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-STATIONARY-REGRET","label":"Stationary latent-arm-stream Thompson regret","kind":"mathematical milestone","status":"compiled","subtitle":"THOMPSON-STATIONARY-REGRET","description":"Under the explicit pointwise mean-optimal selector, stationary Markov, bounded-mean, and centered sub-Gaussian contracts, the canonical generated Thompson trajectory satisfies E[R_n^Bayes] <= (2K+1)(u-l)+8 sqrt(sigma^2 K n log n).","url":"../implementation-map/index.html#thompson-stationary-regret","parent":"group:milestones","order":37,"meta":[["Book Map chapter","thompson"],["Lean declarations","2"],["Prerequisite milestones","4"]],"statement":"","missing":[],"search":"stationary latent-arm-stream thompson regret thompson-stationary-regret under the explicit pointwise mean-optimal selector, stationary markov, bounded-mean, and centered sub-gaussian contracts, the canonical generated thompson trajectory satisfies e[r_n^bayes] <= (2k+1)(u-l)+8 sqrt(sigma^2 k n log n). mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:THOMPSON-GENERAL-PORT","label":"General Thompson/LML Bayesian port","kind":"mathematical milestone","status":"partial","subtitle":"THOMPSON-GENERAL-PORT","description":"The stationary local endpoint is complete, but posterior-law producers outside that model and exact upstream compatibility remain separate obligations.","url":"../implementation-map/index.html#thompson-general-port","parent":"group:milestones","order":38,"meta":[["Book Map chapter","thompson"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":["Add posterior-law producers for broader models.","Close the exact upstream compatibility gate."],"search":"general thompson/lml bayesian port thompson-general-port the stationary local endpoint is complete, but posterior-law producers outside that model and exact upstream compatibility remain separate obligations. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:EXP3-EXPECTED-REGRET","label":"Tuned expected EXP3 regret","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-EXPECTED-REGRET","description":"The generated predictable EXP3 process satisfies an explicit square-root expected-regret bound.","url":"../implementation-map/index.html#exp3-expected-regret","parent":"group:milestones","order":39,"meta":[["Book Map chapter","exp3"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"tuned expected exp3 regret exp3-expected-regret the generated predictable exp3 process satisfies an explicit square-root expected-regret bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:EXP3-BEST-ARM-REALIZED-HIGH-PROBABILITY","label":"Per-horizon best-arm realized high-probability EXP3","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-BEST-ARM-REALIZED-HIGH-PROBABILITY","description":"For each supplied positive horizon, finite comparator aggregation gives the generated horizon-tuned EXP3 law a best-supported-arm realized-regret tail; changing the horizon changes the parameters and law.","url":"../implementation-map/index.html#exp3-best-arm-realized-high-probability","parent":"group:milestones","order":40,"meta":[["Book Map chapter","exp3"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":["This is not one horizon-free policy with a simultaneous confidence event over all horizons."],"search":"per-horizon best-arm realized high-probability exp3 exp3-best-arm-realized-high-probability for each supplied positive horizon, finite comparator aggregation gives the generated horizon-tuned exp3 law a best-supported-arm realized-regret tail; changing the horizon changes the parameters and law. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:CONCENTRATION-COUNTABLE-SCHEDULED-QUADRATIC-TAIL","label":"Countable scheduled quadratic fixed-MGF tail","kind":"mathematical milestone","status":"compiled","subtitle":"CONCENTRATION-COUNTABLE-SCHEDULED-QUADRATIC-TAIL","description":"A countable union of indexwise deviation-and-variance events is controlled by the sum of their confidence shares and hence by a caller-supplied total budget.","url":"../implementation-map/index.html#concentration-countable-scheduled-quadratic-tail","parent":"group:milestones","order":41,"meta":[["Book Map chapter","probability"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"countable scheduled quadratic fixed-mgf tail concentration-countable-scheduled-quadratic-tail a countable union of indexwise deviation-and-variance events is controlled by the sum of their confidence shares and hence by a caller-supplied total budget. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:EXP3-PREDICTABLE-VARIANCE-GEOMETRIC-ALL-TIME","label":"All-positive-prefix EXP3 predictable-variance tail","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-PREDICTABLE-VARIANCE-GEOMETRIC-ALL-TIME","description":"On one generated EXP3 trajectory law, geometric confidence shares control the joint selected-loss deviation and predictable-variance failures over every positive prefix by one outer budget.","url":"../implementation-map/index.html#exp3-predictable-variance-geometric-all-time","parent":"group:milestones","order":42,"meta":[["Book Map chapter","exp3"],["Lean declarations","1"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"all-positive-prefix exp3 predictable-variance tail exp3-predictable-variance-geometric-all-time on one generated exp3 trajectory law, geometric confidence shares control the joint selected-loss deviation and predictable-variance failures over every positive prefix by one outer budget. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:EXP3-REALIZED-DEVIATION-GEOMETRIC-ALL-TIME","label":"All-positive-prefix EXP3 realized-deviation tail","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-REALIZED-DEVIATION-GEOMETRIC-ALL-TIME","description":"On one fixed generated EXP3 process, the realized selected loss minus its predictable counterpart stays below its geometrically scheduled radius at every positive prefix outside one failure event of outer mass at most delta.","url":"../implementation-map/index.html#exp3-realized-deviation-geometric-all-time","parent":"group:milestones","order":43,"meta":[["Book Map chapter","exp3"],["Lean declarations","1"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"all-positive-prefix exp3 realized-deviation tail exp3-realized-deviation-geometric-all-time on one fixed generated exp3 process, the realized selected loss minus its predictable counterpart stays below its geometrically scheduled radius at every positive prefix outside one failure event of outer mass at most delta. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:EXP3-PREDICTABLE-REGRET-GEOMETRIC-ALL-TIME","label":"All-positive-prefix EXP3 predictable-regret tail","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-PREDICTABLE-REGRET-GEOMETRIC-ALL-TIME","description":"For one fixed generated EXP3 process and one supported comparator, every positive-prefix predictable-regret failure is covered by a single geometrically budgeted event.","url":"../implementation-map/index.html#exp3-predictable-regret-geometric-all-time","parent":"group:milestones","order":44,"meta":[["Book Map chapter","exp3"],["Lean declarations","1"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"all-positive-prefix exp3 predictable-regret tail exp3-predictable-regret-geometric-all-time for one fixed generated exp3 process and one supported comparator, every positive-prefix predictable-regret failure is covered by a single geometrically budgeted event. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:EXP3-REALIZED-REGRET-GEOMETRIC-ALL-TIME","label":"All-positive-prefix EXP3 realized-regret tail","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-REALIZED-REGRET-GEOMETRIC-ALL-TIME","description":"For one generated EXP3 process and one supported comparator, the realized selected-loss regret at every positive prefix is controlled by the sum of the predictable-regret and realized-deviation schedules outside one event of outer mass at most delta.","url":"../implementation-map/index.html#exp3-realized-regret-geometric-all-time","parent":"group:milestones","order":45,"meta":[["Book Map chapter","exp3"],["Lean declarations","2"],["Prerequisite milestones","2"]],"statement":"","missing":["The process parameters eta and gamma and the comparator are fixed across prefixes.","The theorem is not a horizon-varying tuned sublinear all-time guarantee, a best-arm minimum, or an ideal EXP3.P theorem."],"search":"all-positive-prefix exp3 realized-regret tail exp3-realized-regret-geometric-all-time for one generated exp3 process and one supported comparator, the realized selected-loss regret at every positive prefix is controlled by the sum of the predictable-regret and realized-deviation schedules outside one event of outer mass at most delta. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:EXP3-SPARSE-ALL-HORIZON","label":"Sparse-loss all-horizon high-probability EXP3","kind":"mathematical milestone","status":"compiled","subtitle":"EXP3-SPARSE-ALL-HORIZON","description":"The best-arm realized-regret tail is controlled with the supplied sparsity-failure probability left explicit.","url":"../implementation-map/index.html#exp3-sparse-all-horizon","parent":"group:milestones","order":46,"meta":[["Book Map chapter","exp3"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"sparse-loss all-horizon high-probability exp3 exp3-sparse-all-horizon the best-arm realized-regret tail is controlled with the supplied sparsity-failure probability left explicit. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TSALLIS-IID-LOG","label":"Finite-arm IID half-Tsallis logarithmic regret","kind":"mathematical milestone","status":"compiled","subtitle":"TSALLIS-IID-LOG","description":"IID probability arm laws with exact model means and positive non-best gaps yield a logarithmic reciprocal-gap regret bound.","url":"../implementation-map/index.html#tsallis-iid-log","parent":"group:milestones","order":47,"meta":[["Book Map chapter","tsallis"],["Lean declarations","1"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"finite-arm iid half-tsallis logarithmic regret tsallis-iid-log iid probability arm laws with exact model means and positive non-best gaps yield a logarithmic reciprocal-gap regret bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TSALLIS-HISTORY-ADAPTIVE-CORRUPTION","label":"History-adaptive expected-corruption all-regimes bound","kind":"mathematical milestone","status":"compiled","subtitle":"TSALLIS-HISTORY-ADAPTIVE-CORRUPTION","description":"A measurable predictable corruption model receives an internally selected refined or logarithmic regret bound.","url":"../implementation-map/index.html#tsallis-history-adaptive-corruption","parent":"group:milestones","order":48,"meta":[["Book Map chapter","tsallis"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"history-adaptive expected-corruption all-regimes bound tsallis-history-adaptive-corruption a measurable predictable corruption model receives an internally selected refined or logarithmic regret bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TSALLIS-DYNAMIC-REGRET","label":"Nonidentical drifting-mean dynamic regret","kind":"mathematical milestone","status":"compiled","subtitle":"TSALLIS-DYNAMIC-REGRET","description":"Predictable-environment regret to the actual moving best arm is bounded by the fixed-comparator route plus an explicit mean-drift penalty.","url":"../implementation-map/index.html#tsallis-dynamic-regret","parent":"group:milestones","order":49,"meta":[["Book Map chapter","tsallis"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"nonidentical drifting-mean dynamic regret tsallis-dynamic-regret predictable-environment regret to the actual moving best arm is bounded by the fixed-comparator route plus an explicit mean-drift penalty. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TSALLIS-ORACLE-RESTART-GENERATED","label":"Generated oracle-restart switch-count dynamic regret","kind":"mathematical milestone","status":"compiled","subtitle":"TSALLIS-ORACLE-RESTART-GENERATED","description":"A change-point schedule built from population-mean switches generates a single restart trajectory whose expected dynamic regret is bounded by the square-root switch-count rate under the route's support assumptions.","url":"../implementation-map/index.html#tsallis-oracle-restart-generated","parent":"group:milestones","order":50,"meta":[["Book Map chapter","tsallis"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"generated oracle-restart switch-count dynamic regret tsallis-oracle-restart-generated a change-point schedule built from population-mean switches generates a single restart trajectory whose expected dynamic regret is bounded by the square-root switch-count rate under the route's support assumptions. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-FINITE-MDP-BELLMAN","label":"Finite-horizon MDP and Bellman interface","kind":"mathematical milestone","status":"compiled","subtitle":"RL-FINITE-MDP-BELLMAN","description":"Finite-horizon MDP data, measurable Markov policies, recursive value functions, optimal Bellman operators, and an attaining optimal policy are formalized locally.","url":"../implementation-map/index.html#rl-finite-mdp-bellman","parent":"group:milestones","order":51,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","2"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"finite-horizon mdp and bellman interface rl-finite-mdp-bellman finite-horizon mdp data, measurable markov policies, recursive value functions, optimal bellman operators, and an attaining optimal policy are formalized locally. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-OCCUPANCY-REGRET","label":"Expected regret as an occupancy Bellman gap","kind":"mathematical milestone","status":"compiled","subtitle":"RL-OCCUPANCY-REGRET","description":"A Markov policy's expected regret equals the occupancy-weighted Bellman optimality gap and is nonnegative; the optimal policy has zero regret.","url":"../implementation-map/index.html#rl-occupancy-regret","parent":"group:milestones","order":52,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","2"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"expected regret as an occupancy bellman gap rl-occupancy-regret a markov policy's expected regret equals the occupancy-weighted bellman optimality gap and is nonnegative; the optimal policy has zero regret. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-ADAPTIVE-REALIZED-CONSISTENCY","label":"Adaptive realized behavior-regret consistency","kind":"mathematical milestone","status":"compiled","subtitle":"RL-ADAPTIVE-REALIZED-CONSISTENCY","description":"With the explicit path-support, bounded-mean, sub-Gaussian, scheduling, and Standard Borel contracts, both the failure budget and average realized behavior-regret envelope tend to zero across windows.","url":"../implementation-map/index.html#rl-adaptive-realized-consistency","parent":"group:milestones","order":53,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"adaptive realized behavior-regret consistency rl-adaptive-realized-consistency with the explicit path-support, bounded-mean, sub-gaussian, scheduling, and standard borel contracts, both the failure budget and average realized behavior-regret envelope tend to zero across windows. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-UNBOUNDED-HITTINGAFTER-L2","label":"Inverse-sqrt hittingAfter is a square-integrable finite stopping time","kind":"mathematical milestone","status":"compiled","subtitle":"RL-UNBOUNDED-HITTINGAFTER-L2","description":"For each fixed threshold index and horizon greater than four, the genuine uncapped Mathlib hittingAfter has an L2 round count under the exact generated causal source.","url":"../implementation-map/index.html#rl-unbounded-hittingafter-l2","parent":"group:milestones","order":54,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","1"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"inverse-sqrt hittingafter is a square-integrable finite stopping time rl-unbounded-hittingafter-l2 for each fixed threshold index and horizon greater than four, the genuine uncapped mathlib hittingafter has an l2 round count under the exact generated causal source. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-UNBOUNDED-HITTINGAFTER-EXPECTED-UPPER-BOUND","label":"Stopped realized behavior regret is integrable and below its hit threshold in expectation","kind":"mathematical milestone","status":"compiled","subtitle":"RL-UNBOUNDED-HITTINGAFTER-EXPECTED-UPPER-BOUND","description":"For every fixed inverse-sqrt threshold index and horizon greater than four, the exact average realized behavior-regret process stopped at the uncapped hit is integrable and its integral is at most that threshold.","url":"../implementation-map/index.html#rl-unbounded-hittingafter-expected-upper-bound","parent":"group:milestones","order":55,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","1"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"stopped realized behavior regret is integrable and below its hit threshold in expectation rl-unbounded-hittingafter-expected-upper-bound for every fixed inverse-sqrt threshold index and horizon greater than four, the exact average realized behavior-regret process stopped at the uncapped hit is integrable and its integral is at most that threshold. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","label":"Generated adaptive cumulative Hoeffding UCBVI-CH chain","kind":"mathematical milestone","status":"compiled","subtitle":"RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","description":"One generated adaptive process now carries exact aggregate transition numerators and visit denominators, a previous-Q clipped recurrent planner, a strict-prefix measurable policy, joint same-source singleton-Bernstein and optimal-tail confidence, Bellman optimism, raw generated episode pseudo-regret, actual-count charge summation, and a generated-filtration Bellman martingale.","url":"../implementation-map/index.html#rl-ucbvi-hoeffding-generated-foundation","parent":"group:milestones","order":56,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","8"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"generated adaptive cumulative hoeffding ucbvi-ch chain rl-ucbvi-hoeffding-generated-foundation one generated adaptive process now carries exact aggregate transition numerators and visit denominators, a previous-q clipped recurrent planner, a strict-prefix measurable policy, joint same-source singleton-bernstein and optimal-tail confidence, bellman optimism, raw generated episode pseudo-regret, actual-count charge summation, and a generated-filtration bellman martingale. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:RL-UCBVI-HOEFFDING-CANONICAL-TERMINALS","label":"Canonical known-reward Hoeffding UCBVI-CH terminals","kind":"mathematical milestone","status":"compiled","subtitle":"RL-UCBVI-HOEFFDING-CANONICAL-TERMINALS","description":"On the recurrent source's own trajectory measure, the probability that raw K-episode policy-value pseudo-regret exceeds 20 H sqrt(H) L sqrt(S A K) + 250 H^2 S^2 A L^2 is at most delta. The matching integrable expectation is at most that bound plus K H delta.","url":"../implementation-map/index.html#rl-ucbvi-hoeffding-canonical-terminals","parent":"group:milestones","order":57,"meta":[["Book Map chapter","finite-horizon-rl"],["Lean declarations","2"],["Prerequisite milestones","4"]],"statement":"","missing":["Bernstein/variance-aware minimax UCB-VI leading-rate theorem.","Stochastic-reward and realized sampled-return UCBVI terminals.","Posterior-sampling, model-free, and continuous-space RL extensions."],"search":"canonical known-reward hoeffding ucbvi-ch terminals rl-ucbvi-hoeffding-canonical-terminals on the recurrent source's own trajectory measure, the probability that raw k-episode policy-value pseudo-regret exceeds 20 h sqrt(h) l sqrt(s a k) + 250 h^2 s^2 a l^2 is at most delta. the matching integrable expectation is at most that bound plus k h delta. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:BWK-STOPPING-FOUNDATION","label":"Budget-exhaustion stopping time","kind":"mathematical milestone","status":"compiled","subtitle":"BWK-STOPPING-FOUNDATION","description":"An adapted natural-valued spending process reaches a fixed budget at a stopping time.","url":"../implementation-map/index.html#bwk-stopping-foundation","parent":"group:milestones","order":58,"meta":[["Book Map chapter","frontier"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"budget-exhaustion stopping time bwk-stopping-foundation an adapted natural-valued spending process reaches a fixed budget at a stopping time. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:BWK-FINAL-REGRET","label":"Bandits-with-knapsacks regret theorem","kind":"mathematical milestone","status":"blocked","subtitle":"BWK-FINAL-REGRET","description":"Stopping-time and several positive-cost budget adapters compile, but there is no full resource process, feasibility invariant, primal-dual comparison, and BwK regret theorem.","url":"../implementation-map/index.html#bwk-final-regret","parent":"group:milestones","order":59,"meta":[["Book Map chapter","frontier"],["Lean declarations","1"],["Prerequisite milestones","0"]],"statement":"","missing":["Resource-consumption and feasibility model.","Primal-dual comparison.","Final resource-constrained regret assembly."],"search":"bandits-with-knapsacks regret theorem bwk-final-regret stopping-time and several positive-cost budget adapters compile, but there is no full resource process, feasibility invariant, primal-dual comparison, and bwk regret theorem. mathematical milestone blocked","shard":"views/milestones.json"},{"id":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","label":"Generated bounded-reward KL-UCB confidence and regret","kind":"mathematical milestone","status":"compiled","subtitle":"KL-UCB-BOUNDED-GENERATED-REGRET","description":"One horizon-free measurable policy uses the actual Bernoulli-KL confidence-set supremum on its generated reward history. On the same canonical trajectory measure, a telescoping all-time confidence event controls every positive-gap arm at every finite horizon and yields a conservative finite-time expected pseudo-regret bound.","url":"../implementation-map/index.html#kl-ucb-bounded-generated-regret","parent":"group:milestones","order":60,"meta":[["Book Map chapter","ucb"],["Lean declarations","8"],["Prerequisite milestones","2"]],"statement":"","missing":["The finite-time conservative route assumes AE [0,1] rewards and means in [margin,1-margin], and retains T times delta.","Sharp KL-Chernoff concentration, the Garivier-Cappe leading constant, and asymptotic optimality are not claimed."],"search":"generated bounded-reward kl-ucb confidence and regret kl-ucb-bounded-generated-regret one horizon-free measurable policy uses the actual bernoulli-kl confidence-set supremum on its generated reward history. on the same canonical trajectory measure, a telescoping all-time confidence event controls every positive-gap arm at every finite horizon and yields a conservative finite-time expected pseudo-regret bound. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","label":"Chapter 13 lower-bound semantic, optimality, and deterministic spine","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","description":"The Chapter 13 conversion window compiles explicit ENNReal worst-case/minimax semantics, fixed-class minimax-optimality without assuming general attainment, a source-shaped least-explored alternative-arm theorem from the exact expected pull budget, and quantitative deterministic two-environment regret algebra whose cross-law pull discrepancy remains a visible error premise.","url":"../implementation-map/index.html#textbook-part-iv-ch13-basic-ideas-lean-spine","parent":"group:milestones","order":61,"meta":[["Book Map chapter","foundations"],["Lean declarations","15"],["Prerequisite milestones","0"]],"statement":"","missing":["The Chapter 14 event-testing foundation and Chapter 15 same-policy adaptive-history KL identity compile.","The separate Chapter 15 consumer closes Theorem 13.1 with c=1/54 and GaussianHypothesisTesting closes exact Eq. (13.1); the broader-class fixed-horizon MOSS near-minimax consequence also compiles. Whole-chapter integration, review and deployment gates remain pending."],"search":"chapter 13 lower-bound semantic, optimality, and deterministic spine textbook-part-iv-ch13-basic-ideas-lean-spine the chapter 13 conversion window compiles explicit ennreal worst-case/minimax semantics, fixed-class minimax-optimality without assuming general attainment, a source-shaped least-explored alternative-arm theorem from the exact expected pull budget, and quantitative deterministic two-environment regret algebra whose cross-law pull discrepancy remains a visible error premise. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","label":"Chapter 13 exact Gaussian testing bounds and Chernoff companion","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","description":"The canonical iid Gaussian empirical mean has law N(mu,1/n). Both midpoint error events and their maximum-risk Chernoff upper bound compile. Separately, the exact lower and upper Mills-ratio integrals of Eq. (13.4) give the printed Eq. (13.1) bounds for the zero-mean error, with denominator constants 16 and 32/pi, for n>0 and Delta>0.","url":"../implementation-map/index.html#textbook-part-iv-ch13-gaussian-testing-companion","parent":"group:milestones","order":62,"meta":[["Book Map chapter","foundations"],["Lean declarations","23"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"chapter 13 exact gaussian testing bounds and chernoff companion textbook-part-iv-ch13-gaussian-testing-companion the canonical iid gaussian empirical mean has law n(mu,1/n). both midpoint error events and their maximum-risk chernoff upper bound compile. separately, the exact lower and upper mills-ratio integrals of eq. (13.4) give the printed eq. (13.1) bounds for the zero-mean error, with denominator constants 16 and 32/pi, for n>0 and delta>0. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","label":"Chapter 14 coding, KL data-processing, and Bretagnolle–Huber spine","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","description":"The frozen Chapter 14 required body is compiled: recursive Huffman optimality and its entropy sandwich, exact-real arithmetic block coding and universal converse, finite/partition/common-density KL, the source affinity/overlap route, and positive-variance Gaussian testing. Independent review, PR #106, main run 33959196451, Pages and live acceptance passed; the model qualifications below remain part of the claim boun…","url":"../implementation-map/index.html#textbook-part-iv-ch14-information-theory-lean-spine","parent":"group:milestones","order":63,"meta":[["Book Map chapter","foundations"],["Lean declarations","58"],["Prerequisite milestones","2"]],"statement":"","missing":["Required-body independent review, full main gate and public desktop/mobile acceptance passed. Optional Notes/Bibliographic Remarks/Exercises are not claimed complete; full Exercise 14.10 is an additional compiled result.","Nonempty singleton codewords are a local model convention. Uniform fixed-length optimality holds for power-of-two cardinalities; a compiled ternary counterexample refutes the arbitrary-cardinality reading.","Arithmetic coding is exact-real and classical with constant support/escape overhead, not executable finite-precision code. Cross-entropy differences are unrounded; finite KL iff absolute continuity is finite-alphabet-only.","The adaptive same-policy bandit-history divergence decomposition is a Chapter 15 target, not a Chapter 14 theorem."],"search":"chapter 14 coding, kl data-processing, and bretagnolle–huber spine textbook-part-iv-ch14-information-theory-lean-spine the frozen chapter 14 required body is compiled: recursive huffman optimality and its entropy sandwich, exact-real arithmetic block coding and universal converse, finite/partition/common-density kl, the source affinity/overlap route, and positive-variance gaussian testing. independent review, pr #106, main run 33959196451, pages and live acceptance passed; the model qualifications below remain part of the claim boundary. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","label":"Chapter 15 unit-Gaussian likelihood-ratio and KL dependency slice","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","description":"For real means mu and nu, the project constructs the unit-variance Gaussian arm laws, identifies their log Radon–Nikodym derivative under the first law, proves its integrability, and derives the exact extended-real identity D(N(mu,1),N(nu,1))=(mu-nu)^2/2. The changed-arm cost, source gap, exact information exponent one half, and unit-cube gap upper bound also compile.","url":"../implementation-map/index.html#textbook-part-iv-ch15-gaussian-kl-dependency-slice","parent":"group:milestones","order":64,"meta":[["Book Map chapter","foundations"],["Lean declarations","11"],["Prerequisite milestones","2"]],"statement":"","missing":["This milestone is intentionally arm-local; the separate compiled Theorem 15.2 milestone consumes it together with Lemma 15.1 and Chapter 14 event testing."],"search":"chapter 15 unit-gaussian likelihood-ratio and kl dependency slice textbook-part-iv-ch15-gaussian-kl-dependency-slice for real means mu and nu, the project constructs the unit-variance gaussian arm laws, identifies their log radon–nikodym derivative under the first law, proves its integrability, and derives the exact extended-real identity d(n(mu,1),n(nu,1))=(mu-nu)^2/2. the changed-arm cost, source gap, exact information exponent one half, and unit-cube gap upper bound also compile. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","label":"Chapter 15 same-policy adaptive-history KL decomposition","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","description":"For finite arms, a countably generated reward space, arbitrary stationary Markov arm laws, and one common randomized history policy, the directed KL between the two canonical finite history laws equals the finite sum of first-law realized expected pull counts times the corresponding directed arm KL. The proof includes singular fibres and infinite arm KL through an extended-real conditional-kernel chain rule.","url":"../implementation-map/index.html#textbook-part-iv-ch15-same-policy-history-kl-decomposition","parent":"group:milestones","order":65,"meta":[["Book Map chapter","foundations"],["Lean declarations","9"],["Prerequisite milestones","1"]],"statement":"","missing":["The local inclusive-round index lastRound represents lastRound+1 observations; later source consumers must preserve that convention explicitly.","The compiled Theorem 15.2 milestone consumes this identity, but later Chapter 16–17 terminals remain separate and are not consequences of Lemma 15.1 alone."],"search":"chapter 15 same-policy adaptive-history kl decomposition textbook-part-iv-ch15-same-policy-history-kl-decomposition for finite arms, a countably generated reward space, arbitrary stationary markov arm laws, and one common randomized history policy, the directed kl between the two canonical finite history laws equals the finite sum of first-law realized expected pull counts times the corresponding directed arm kl. the proof includes singular fibres and infinite arm kl through an extended-real conditional-kernel chain rule. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH15-EX15-7-DATA-PROCESSING-LEAF","label":"Chapter 15 Exercise 15.7 measurable-observation KL leaf","kind":"mathematical milestone","status":"partial","subtitle":"TEXTBOOK-PART-IV-CH15-EX15-7-DATA-PROCESSING-LEAF","description":"KL divergence between finite measures cannot increase after an arbitrary measurable observation. Combined with the compiled deterministic-horizon Lemma 15.1, any measurable statistic of the complete finite bandit history is bounded by the first-law expected pull-count-weighted arm information. The bounded stopping-time history bound required by Exercise 15.7 remains open.","url":"../implementation-map/index.html#textbook-part-iv-ch15-ex15-7-data-processing-leaf","parent":"group:milestones","order":66,"meta":[["Book Map chapter","foundations"],["Lean declarations","2"],["Prerequisite milestones","1"]],"statement":"","missing":["Construct a measurable bounded stopped-history law and prove that its KL charges arm information only through the realized stopping time.","Factor an arbitrary F_tau-measurable random element through the stopped history before applying data processing."],"search":"chapter 15 exercise 15.7 measurable-observation kl leaf textbook-part-iv-ch15-ex15-7-data-processing-leaf kl divergence between finite measures cannot increase after an arbitrary measurable observation. combined with the compiled deterministic-horizon lemma 15.1, any measurable statistic of the complete finite bandit history is bounded by the first-law expected pull-count-weighted arm information. the bounded stopping-time history bound required by exercise 15.7 remains open. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","label":"Chapter 15 finite-armed Gaussian minimax lower bound","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","description":"For every possibly randomized nonanticipating finite-history policy, every k greater than one, and every horizon n at least k-1, the library constructs a unit-cube mean vector for a k-armed unit-variance Gaussian bandit whose expected pseudo-regret is at least sqrt((k-1)n)/27. The proof uses the actual canonical history law, least-explored arm, source event T_0(n) at most n/2, Lemma 15.1 in the first-environment KL…","url":"../implementation-map/index.html#textbook-part-iv-ch15-gaussian-minimax-lower-bound","parent":"group:milestones","order":67,"meta":[["Book Map chapter","foundations"],["Lean declarations","12"],["Prerequisite milestones","4"]],"statement":"","missing":["The Bernoulli refinement in the Chapter 15 notes and Exercises 15.1–15.6 and 15.8 are outside this compiled main-theorem milestone.","Exercise 15.7 has a separate compiled data-processing dependency, but its stopped-history inequality remains open.","Chapter 16 and Chapter 17 lower-bound terminals remain separate consumers and are not promoted by this result."],"search":"chapter 15 finite-armed gaussian minimax lower bound textbook-part-iv-ch15-gaussian-minimax-lower-bound for every possibly randomized nonanticipating finite-history policy, every k greater than one, and every horizon n at least k-1, the library constructs a unit-cube mean vector for a k-armed unit-variance gaussian bandit whose expected pseudo-regret is at least sqrt((k-1)n)/27. the proof uses the actual canonical history law, least-explored arm, source event t_0(n) at most n/2, lemma 15.1 in the first-environment kl direction, bretagnolle–huber, the exact source gap, and an explicit worst-case/minimax lattice layer. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","label":"Chapter 16 exact Gaussian d_inf and one-arm information dependency slice","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","description":"Twenty named Chapter 16 dependencies compile: Definition 16.1's consistency interface and power/log-growth consequences; the extended-real d_inf interface and exact unit-Gaussian Table 16.1 row; the same-policy one-arm history-KL identity; the measurable majority event and Bretagnolle–Huber information inequality; and the finite-KL scalar regret/log assembly. The separate compiled event-to-regret record now consumes…","url":"../implementation-map/index.html#textbook-part-iv-ch16-consistency-dinf-dependency-slice","parent":"group:milestones","order":68,"meta":[["Book Map chapter","foundations"],["Lean declarations","20"],["Prerequisite milestones","4"]],"statement":"","missing":["This historical dependency record predates the now-compiled finite-mean producer, Lemma 16.3, and Theorem 16.4.","Theorem 16.2 still needs its per-arm information-to-liminf extraction, including the zero, finite, and infinite d_inf branches."],"search":"chapter 16 exact gaussian d_inf and one-arm information dependency slice textbook-part-iv-ch16-consistency-dinf-dependency-slice twenty named chapter 16 dependencies compile: definition 16.1's consistency interface and power/log-growth consequences; the extended-real d_inf interface and exact unit-gaussian table 16.1 row; the same-policy one-arm history-kl identity; the measurable majority event and bretagnolle–huber information inequality; and the finite-kl scalar regret/log assembly. the separate compiled event-to-regret record now consumes this layer without changing the scope of these twenty declarations. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","label":"Chapter 16 majority-event expected-pseudo-regret producers","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","description":"Fifteen named declarations define canonical gap-times-pull-count pseudo-regret, charge the majority event and its complement with exact half-horizon factors, and combine those charges with the one-arm KL/Bretagnolle–Huber layer. The later finite-mean producer now consumes this historical layer to prove Lemma 16.3 and Theorem 16.4.","url":"../implementation-map/index.html#textbook-part-iv-ch16-event-regret-producers","parent":"group:milestones","order":69,"meta":[["Book Map chapter","foundations"],["Lean declarations","15"],["Prerequisite milestones","4"]],"statement":"","missing":["This historical producer record stops below the now-compiled finite-mean environment consumer.","Theorem 16.2 now compiles with every extended-real d_inf branch and finite-count Fatou."],"search":"chapter 16 majority-event expected-pseudo-regret producers textbook-part-iv-ch16-event-regret-producers fifteen named declarations define canonical gap-times-pull-count pseudo-regret, charge the majority event and its complement with exact half-horizon factors, and combine those charges with the one-arm kl/bretagnolle–huber layer. the later finite-mean producer now consumes this historical layer to prove lemma 16.3 and theorem 16.4. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH16-SOURCE-TERMINALS","label":"Chapter 16 instance-dependent asymptotic and finite-time source terminals","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH16-SOURCE-TERMINALS","description":"The source-frozen endpoints are Theorem 16.2's unstructured-class liminf regret constant, Lemma 16.3's one-arm finite-time expected-pull inequality, and Theorem 16.4's unit-Gaussian local-envelope positive-part lower bound. All three source terminals now compile exactly, including Theorem 16.2 with every extended-real information branch.","url":"../implementation-map/index.html#textbook-part-iv-ch16-source-terminals","parent":"group:milestones","order":70,"meta":[["Book Map chapter","foundations"],["Lean declarations","3"],["Prerequisite milestones","10"]],"statement":"","missing":[],"search":"chapter 16 instance-dependent asymptotic and finite-time source terminals textbook-part-iv-ch16-source-terminals the source-frozen endpoints are theorem 16.2's unstructured-class liminf regret constant, lemma 16.3's one-arm finite-time expected-pull inequality, and theorem 16.4's unit-gaussian local-envelope positive-part lower bound. all three source terminals now compile exactly, including theorem 16.2 with every extended-real information branch. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","label":"Chapter 17 stochastic tail terminals and adversarial pathwise core","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","description":"The project compiles Theorem 17.1 with its full gap-at-most-one Gaussian source-class premise and a unit-cube witness, Corollaries 17.2–17.3 with exact quantifiers and constants, Claim 17.5's first-moment witness, the correlated clipped shared-noise path, construction-level Eq. (17.8), event subtraction, and quarter-horizon algebra.","url":"../implementation-map/index.html#textbook-part-iv-ch17-first-moment-and-tail-dependency-slice","parent":"group:milestones","order":71,"meta":[["Book Map chapter","foundations"],["Lean declarations","16"],["Prerequisite milestones","1"]],"statement":"","missing":[],"search":"chapter 17 stochastic tail terminals and adversarial pathwise core textbook-part-iv-ch17-first-moment-and-tail-dependency-slice the project compiles theorem 17.1 with its full gap-at-most-one gaussian source-class premise and a unit-cube witness, corollaries 17.2–17.3 with exact quantifiers and constants, claim 17.5's first-moment witness, the correlated clipped shared-noise path, construction-level eq. (17.8), event subtraction, and quarter-horizon algebra. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","label":"Chapter 17 stochastic and corrected adversarial high-probability terminals","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","description":"Theorem 17.1, Corollaries 17.2–17.3, Claims 17.5–17.7, Eq. (17.8), same-policy shared-noise coupling, deterministic matrix extraction, and Theorem 17.4 pass the full local Lean/Tests/harness gate. User-approved corrections: Claim 17.6 uses T_i ≤ n/2; Theorem 17.4 uses 0 < δ ≤ 1/32, c=1/160, C=64 and the strict random-regret CDF tail. No proof of the false printed statements or new deployment is claimed.","url":"../implementation-map/index.html#textbook-part-iv-ch17-source-terminals","parent":"group:milestones","order":72,"meta":[["Book Map chapter","foundations"],["Lean declarations","11"],["Prerequisite milestones","5"]],"statement":"","missing":[],"search":"chapter 17 stochastic and corrected adversarial high-probability terminals textbook-part-iv-ch17-source-terminals theorem 17.1, corollaries 17.2–17.3, claims 17.5–17.7, eq. (17.8), same-policy shared-noise coupling, deterministic matrix extraction, and theorem 17.4 pass the full local lean/tests/harness gate. user-approved corrections: claim 17.6 uses t_i ≤ n/2; theorem 17.4 uses 0 < δ ≤ 1/32, c=1/160, c=64 and the strict random-regret cdf tail. no proof of the false printed statements or new deployment is claimed. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:TEXTBOOK-PART-IV-THEOREM-13-1-GAUSSIAN-MINIMAX","label":"Theorem 13.1 Gaussian finite-arm minimax lower bound","kind":"mathematical milestone","status":"compiled","subtitle":"TEXTBOOK-PART-IV-THEOREM-13-1-GAUSSIAN-MINIMAX","description":"For k-armed unit-variance Gaussian bandits with mean vectors in the unit cube, the source's universal-constant minimax lower bound of order sqrt(k n) now compiles locally. The exact Chapter 15 theorem yields the explicit admissible constant c=1/54 whenever k is greater than one and n is at least k.","url":"../implementation-map/index.html#textbook-part-iv-theorem-13-1-gaussian-minimax","parent":"group:milestones","order":73,"meta":[["Book Map chapter","foundations"],["Lean declarations","1"],["Prerequisite milestones","3"]],"statement":"","missing":["Exact Eq. (13.1) is separately compiled in GaussianHypothesisTesting; the broader-class fixed-horizon MOSS near-minimax consequence compiles in SubgaussianMinimax, with publication gates pending.","Notes 13.2 and Exercises 13.1–13.2 are optional and are not implied by the theorem."],"search":"theorem 13.1 gaussian finite-arm minimax lower bound textbook-part-iv-theorem-13-1-gaussian-minimax for k-armed unit-variance gaussian bandits with mean vectors in the unit cube, the source's universal-constant minimax lower bound of order sqrt(k n) now compiles locally. the exact chapter 15 theorem yields the explicit admissible constant c=1/54 whenever k is greater than one and n is at least k. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","label":"Source-faithful delayed-feedback accounting","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-FEEDBACK-SOURCE-ACCOUNTING","description":"For the frozen NeurIPS 2025 delayed best-of-both-worlds source, the library formalizes strict pre-action observability, outstanding rounds, their exact finite partition, paper-style end-of-round missing-feedback counts, and the explicit one-based indexing bridge. This is deterministic accounting, not a regret theorem.","url":"../implementation-map/index.html#delayed-feedback-source-accounting","parent":"group:milestones","order":74,"meta":[["Book Map chapter","frontier"],["Lean declarations","17"],["Prerequisite milestones","0"]],"statement":"","missing":[],"search":"source-faithful delayed-feedback accounting delayed-feedback-source-accounting for the frozen neurips 2025 delayed best-of-both-worlds source, the library formalizes strict pre-action observability, outstanding rounds, their exact finite partition, paper-style end-of-round missing-feedback counts, and the explicit one-based indexing bridge. this is deterministic accounting, not a regret theorem. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","label":"Causal action-time view and new-feedback processing","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-FEEDBACK-CAUSAL-PROCESSING","description":"The learner view exposes past actions and exactly the losses available before the next action. Observation-equivalent hidden worlds are indistinguishable to any typed causal decision rule, and the set-level update processes precisely the newly observed rounds. A separate downstream layer now constructs a one-round measure-valued rule; a measurable recursive policy kernel and the paper's ordered state updates remain…","url":"../implementation-map/index.html#delayed-feedback-causal-processing","parent":"group:milestones","order":75,"meta":[["Book Map chapter","frontier"],["Lean declarations","20"],["Prerequisite milestones","2"]],"statement":"","missing":["A measurable stochastic policy kernel and recursively generated action law depending only on this view.","The source's simultaneous-arrival order and BSC/EAP state-transition invariants."],"search":"causal action-time view and new-feedback processing delayed-feedback-causal-processing the learner view exposes past actions and exactly the losses available before the next action. observation-equivalent hidden worlds are indistinguishable to any typed causal decision rule, and the set-level update processes precisely the newly observed rounds. a separate downstream layer now constructs a one-round measure-valued rule; a measurable recursive policy kernel and the paper's ordered state updates remain open. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","label":"Delayed SAPO active-arm allocation leaf","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-ACTIVE-ALLOCATION","description":"Algorithm 5's equal residual allocation over a nonempty active set is normalized and coordinatewise nonnegative whenever the inactive coordinates are nonnegative and carry at most unit mass. No claim is made that the unformalized EAP state maintains those premises.","url":"../implementation-map/index.html#delayed-sapo-active-allocation","parent":"group:milestones","order":76,"meta":[["Book Map chapter","frontier"],["Lean declarations","8"],["Prerequisite milestones","2"]],"statement":"","missing":["A source-faithful EAP state and proof that every update preserves nonnegative inactive mass at most one.","A measurable sampling kernel using this vector on the recursively generated delayed-feedback history; the one-round probability measure is compiled downstream."],"search":"delayed sapo active-arm allocation leaf delayed-sapo-active-allocation algorithm 5's equal residual allocation over a nonempty active set is normalized and coordinatewise nonnegative whenever the inactive coordinates are nonnegative and carry at most unit mass. no claim is made that the unformalized eap state maintains those premises. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","label":"Optimal-arm survival and causal one-round action law","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-ELIMINATION-ACTION-LAW","description":"Algorithm 5 lines 7–8 compile as a source-exact elimination snapshot. The deterministic core of Lemma D.9 proves that an explicit optimal-arm-survival certificate closes the nonempty-active premise for line 15. With EAP's still-explicit nonnegativity and mass premises, the resulting vector induces a genuine one-round probability measure, and every allocation rule typed on the causal view returns the same measure in…","url":"../implementation-map/index.html#delayed-sapo-elimination-action-law","parent":"group:milestones","order":77,"meta":[["Book Map chapter","frontier"],["Lean declarations","19"],["Prerequisite milestones","4"]],"statement":"","missing":["The Definition-D.1 count, phase, error, and delay clauses, their probability bound, and full recursive source Lemma D.9.","EAP preservation of its inactive-probability premises.","Coordinate measurability, a Markov kernel over generated histories, and recursive delayed trajectory generation."],"search":"optimal-arm survival and causal one-round action law delayed-sapo-elimination-action-law algorithm 5 lines 7–8 compile as a source-exact elimination snapshot. the deterministic core of lemma d.9 proves that an explicit optimal-arm-survival certificate closes the nonempty-active premise for line 15. with eap's still-explicit nonnegativity and mass premises, the resulting vector induces a genuine one-round probability measure, and every allocation rule typed on the causal view returns the same measure in observation-equivalent hidden worlds. the source-shaped good-event projection that constructs the certificate is compiled separately; this milestone is not the full lemma, a measurable history kernel, or a regret endpoint. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","label":"Source-shaped good-event projection for optimal-arm survival","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","description":"A source-shaped confidence snapshot defines ucbStar as the finite-arm infimum of the two upper-confidence surfaces used in the paper. Its elimination projection of Definition D.1 derives muStar <= ucbStar rather than assuming that inequality as an independent certificate, constructs the deterministic Lemma-D.9 survival certificate, and transports any externally supplied complement-good-event probability bound to an…","url":"../implementation-map/index.html#delayed-sapo-good-event-d9-projection","parent":"group:milestones","order":78,"meta":[["Book Map chapter","frontier"],["Lean declarations","11"],["Prerequisite milestones","2"]],"statement":"","missing":["The full Definition D.1 count, phase, error, and delay clauses and their measurable simultaneous event.","The D.2–D.7 concentration/counting lemmas that produce the six component probability bounds.","Persistence across the recursive Delayed SAPO state machine and the stochastic/adversarial regret endpoints."],"search":"source-shaped good-event projection for optimal-arm survival delayed-sapo-good-event-d9-projection a source-shaped confidence snapshot defines ucbstar as the finite-arm infimum of the two upper-confidence surfaces used in the paper. its elimination projection of definition d.1 derives mustar <= ucbstar rather than assuming that inequality as an independent certificate, constructs the deterministic lemma-d.9 survival certificate, and transports any externally supplied complement-good-event probability bound to an optimal-arm-elimination bound. corollary d.8's six-event union assembly is compiled separately; the full definition-d.1 event and the d.2–d.7 component concentration/counting producers remain open. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","label":"Corollary-D.8 union assembly to D.9 survival","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-D8-D9-ASSEMBLY","description":"Six explicitly named failure components represent the clauses discharged by source Lemmas D.2–D.7. The corrected Corollary-D.8 layer uses the source-exact three 2/T budgets and three 1/T budgets, proves their exact 9/T sum, and transports a recorded full-event-to-elimination projection through the compiled D.9 consumer to bound optimal-arm elimination. The six concentration/counting bounds and the semantic projectio…","url":"../implementation-map/index.html#delayed-sapo-d8-d9-assembly","parent":"group:milestones","order":79,"meta":[["Book Map chapter","frontier"],["Lean declarations","14"],["Prerequisite milestones","2"]],"statement":"","missing":["Source-faithful random variables, events, and proofs of Lemmas D.2–D.7 on one generated Delayed SAPO law.","A proved projection from the complete Definition-D.1 event to the compiled elimination slice.","Recursive optimal-arm persistence and either paper-level regret endpoint."],"search":"corollary-d.8 union assembly to d.9 survival delayed-sapo-d8-d9-assembly six explicitly named failure components represent the clauses discharged by source lemmas d.2–d.7. the corrected corollary-d.8 layer uses the source-exact three 2/t budgets and three 1/t budgets, proves their exact 9/t sum, and transports a recorded full-event-to-elimination projection through the compiled d.9 consumer to bound optimal-arm elimination. the six concentration/counting bounds and the semantic projection are hypotheses, not claimed source theorems. an earlier 1/t^2 encoding was removed as source drift. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","label":"Lemma-D.10/D.12 width-direction diagnostic and conditional same-snapshot skeleton","kind":"mathematical milestone","status":"partial","subtitle":"DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","description":"Lean proves that the source inverse-square-root empirical width is antitone in a positive pull count and gives both a normalized count-one/count-four witness and a literal T=4 witness against the printed reverse transport. It also proves the exact source-shaped small-count implication count <= 192 log T -> 1 <= 10 width. A branched active-arm consumer uses current-UCB and the optimal-to-later factor-three edge only…","url":"../implementation-map/index.html#delayed-sapo-d10-d12-gap-ordering-audit","parent":"group:milestones","order":80,"meta":[["Book Map chapter","frontier"],["Lean declarations","19"],["Prerequisite milestones","0"]],"statement":"","missing":["A measurable generated Delayed-SAPO trajectory, its Algorithm-5 transition-and-invariant-to-summary producer, and the D.4 simultaneous probability producer for the two count inequalities consumed by the compiled trace-summary adapter.","A source amendment or author clarification for the intended printed D.10 prefix-to-elimination width step; the compiled conditional skeleton bypasses rather than validates that step.","Unconditional source Lemmas D.10/D.12, main-text Lemma 4.2, Theorem 4.1, and either regret endpoint."],"search":"lemma-d.10/d.12 width-direction diagnostic and conditional same-snapshot skeleton delayed-sapo-d10-d12-gap-ordering-audit lean proves that the source inverse-square-root empirical width is antitone in a positive pull count and gives both a normalized count-one/count-four witness and a literal t=4 witness against the printed reverse transport. it also proves the exact source-shaped small-count implication count <= 192 log t -> 1 <= 10 width. a branched active-arm consumer uses current-ucb and the optimal-to-later factor-three edge only in the large-count branch, while the small-count branch uses bounded means plus an explicit source-width/count certificate. a separate processed-prefix producer now derives these algebraic inputs and the later-to-earlier factor-ten edge from a source-time ledger and d.1 count certificate. the earlier explicit four-edge consumer remains for comparison. these 19 diagnostic declarations, even together with that producer, do not verify or refute source lemmas d.10/d.12, main-text lemma 4.2, or theorem 4.1. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","label":"Definition-D.1 active-count to same-prefix width producer","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","description":"An ordered source-time ledger records each processed round's chosen arm, active set, and line-15 allocation, together with the prior-prefix empirical-UCB vector. For arms active throughout the prefix, Lean derives equal cumulative pull mass directly from Algorithm 5's equal active allocation. The exact D.1 lower and upper count clauses then imply n_j >= n_i/4 - 6 log T. Splitting at 192 log T produces the large-bran…","url":"../implementation-map/index.html#delayed-sapo-d1-active-count-width-producer","parent":"group:milestones","order":81,"meta":[["Book Map chapter","frontier"],["Lean declarations","16"],["Prerequisite milestones","3"]],"statement":"","missing":["A measurable randomized Algorithm-5 trajectory and numerical BSC/EAP transition that supplies the explicit fields consumed by the compiled structural processing step.","Lemma D.4's simultaneous probability bound for the D.1 count clause on the generated trajectory.","An ordered elimination-snapshot wrapper and the remaining D.13–D.21 regret chain."],"search":"definition-d.1 active-count to same-prefix width producer delayed-sapo-d1-active-count-width-producer an ordered source-time ledger records each processed round's chosen arm, active set, and line-15 allocation, together with the prior-prefix empirical-ucb vector. for arms active throughout the prefix, lean derives equal cumulative pull mass directly from algorithm 5's equal active allocation. the exact d.1 lower and upper count clauses then imply n_j >= n_i/4 - 6 log t. splitting at 192 log t produces the large-branch factor-three width comparison and the branch-free, certificate-level same-prefix factor-ten comparison; the recursive ucb definition supplies the current-ucb edge. these producers remove the manually supplied branch and pair-width premises from the repaired same-snapshot factor-20 gap theorem. a separate compiled trace-summary adapter now constructs this certificate from distinct, strictly available source-indexed data and the explicit d.4 count clause; producing that summary from algorithm 5 and proving the clause with probability 1 - 2/t remain open. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","label":"Processed trace summary to source-time ledger","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","description":"A deterministic trace summary records distinct source indices for processed feedback items without requiring chronological source order and carries the strict source availability condition s + d_s < t. Its ledger reads the chosen action, source-round active set, and line-15 allocation at each source index rather than reconstructing them from later data. The intra-round current active set is kept separate from the an…","url":"../implementation-map/index.html#delayed-sapo-processed-trace-summary-adapter","parent":"group:milestones","order":82,"meta":[["Book Map chapter","frontier"],["Lean declarations","9"],["Prerequisite milestones","2"]],"statement":"","missing":["A measurable causal randomized Delayed-SAPO kernel and numerical BSC/EAP transition that supply the explicit round state and confidence surfaces consumed by the compiled structural processing step.","Lemma D.4's simultaneous 2/T probability bound for the two count inequalities over the source's processed-prefix family.","The switch branch, an unconditional generated-trajectory D.12 / main-text Lemma 4.2 theorem, and the remaining source regret chain."],"search":"processed trace summary to source-time ledger delayed-sapo-processed-trace-summary-adapter a deterministic trace summary records distinct source indices for processed feedback items without requiring chronological source order and carries the strict source availability condition s + d_s < t. its ledger reads the chosen action, source-round active set, and line-15 allocation at each source index rather than reconstructing them from later data. the intra-round current active set is kept separate from the antitone source-round trace, and current-to-source containment is an explicit summary invariant rather than a generated result. the projected confidence snapshot defines the source inverse-square-root width and recursive empirical ucb directly. given only the two displayed d.4 count inequalities, lean constructs the existing processed-prefix certificate and reaches the conditional same-snapshot factor-20 gap consumer. this is a deterministic adapter, not an algorithm-5 transition/invariant producer, d.4 probability theorem, or full delayed-sapo trajectory. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","label":"One ordered no-switch Algorithm-5 processing step","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","description":"A 15-declaration structural transition refines one iteration of Algorithm 5 lines 3–4 and 7–8. It keeps the paper sequence S as a duplicate-free list, accepts an arbitrary newly observed source in B(t) minus S, and appends it without imposing chronological order. Lean derives source-index injectivity, the exact strict availability condition s + d_s < t, and current-to-source active-set containment before constructin…","url":"../implementation-map/index.html#delayed-sapo-ordered-no-switch-process-one","parent":"group:milestones","order":83,"meta":[["Book Map chapter","frontier"],["Lean declarations","15"],["Prerequisite milestones","3"]],"statement":"","missing":["The numerical BSC/EAP update that constructs the empirical and importance-weighted confidence surfaces read by this structural step.","Round finalization, the switch branch, and a measurable randomized recursive Delayed-SAPO trajectory.","Lemma D.4's simultaneous probability bound, the generated-trajectory version of the multi-snapshot elimination argument, and every paper regret endpoint."],"search":"one ordered no-switch algorithm-5 processing step delayed-sapo-ordered-no-switch-process-one a 15-declaration structural transition refines one iteration of algorithm 5 lines 3–4 and 7–8. it keeps the paper sequence s as a duplicate-free list, accepts an arbitrary newly observed source in b(t) minus s, and appends it without imposing chronological order. lean derives source-index injectivity, the exact strict availability condition s + d_s < t, and current-to-source active-set containment before constructing the line-7 processed trace summary. the line-8 successor uses the exact remainingactive set and preserves the round-start active-set invariant, so another arbitrary new arrival can be processed. a focused canary processes source round one after source round three. this is a deterministic no-switch structural step: its numerical confidence surfaces are explicit inputs, and it does not implement bsc, eap, a generated trajectory, d.4, a switch path, or a regret endpoint. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","label":"Ordered no-switch trace closes the temporal elimination premise","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","description":"A 12-declaration deterministic trace layer composes exact no-switch processing steps with explicit exhausted-round advances. Lean proves that active sets can only shrink along the reflexive-transitive trace, then derives that an arm eliminated by a later step was still present in the earlier line-8 remaining set. This discharges the previously caller-supplied temporal premise of the repaired factor-20 consumer. The…","url":"../implementation-map/index.html#delayed-sapo-ordered-no-switch-trace-ordering","parent":"group:milestones","order":84,"meta":[["Book Map chapter","frontier"],["Lean declarations","12"],["Prerequisite milestones","3"]],"statement":"","missing":["Numerical BSC/EAP production and a measurable randomized Delayed-SAPO trajectory that generate the structural states and confidence surfaces.","Lemma D.4's simultaneous 2/T probability bound and the Definition-D.1 good-event producer.","The switch branch, an unconditional source Lemma D.12 / main-text Lemma 4.2 theorem, and both regret endpoints."],"search":"ordered no-switch trace closes the temporal elimination premise delayed-sapo-ordered-no-switch-trace-ordering a 12-declaration deterministic trace layer composes exact no-switch processing steps with explicit exhausted-round advances. lean proves that active sets can only shrink along the reflexive-transitive trace, then derives that an arm eliminated by a later step was still present in the earlier line-8 remaining set. this discharges the previously caller-supplied temporal premise of the repaired factor-20 consumer. the final theorem still consumes the earlier snapshot's explicit d.4 count clause and elimination-good projection; it is not a generated stochastic trajectory, a proof of d.4, or a source-paper regret theorem. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","label":"Appendix D.11 nonnegative-gap half-set core","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","description":"Lean promotes the Markov-style counting step on exactly the nonnegative domain used by stochastic loss gaps: for any finite nonnegative gap family, at most half the arms lie strictly above twice the average. Empty and zero-average families are covered explicitly, and optimal-arm minimality produces nonnegative source loss gaps. The unrestricted real-valued formulation is not promoted; a signed regression canary guar…","url":"../implementation-map/index.html#delayed-sapo-d11-nonnegative-gap-half-set","parent":"group:milestones","order":85,"meta":[["Book Map chapter","frontier"],["Lean declarations","6"],["Prerequisite milestones","1"]],"statement":"","missing":["A source-faithful Lemma D.13 statement resolving the printed witness/index mismatch and half-active-set convention.","The generated stochastic Delayed-SAPO process and its regret endpoint."],"search":"appendix d.11 nonnegative-gap half-set core delayed-sapo-d11-nonnegative-gap-half-set lean promotes the markov-style counting step on exactly the nonnegative domain used by stochastic loss gaps: for any finite nonnegative gap family, at most half the arms lie strictly above twice the average. empty and zero-average families are covered explicitly, and optimal-arm minimality produces nonnegative source loss gaps. the unrestricted real-valued formulation is not promoted; a signed regression canary guards the premise boundary without claiming a source correction. this six-declaration leaf closes only the deterministic d.11 counting fact; it does not prove d.13 or a stochastic regret endpoint. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","label":"Algorithm 5 line-10 eliminated-arm initialization","kind":"mathematical milestone","status":"compiled","subtitle":"DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","description":"A 31-declaration producer constructs the literal state assigned to each arm selected by Algorithm 5 line 7: the elimination round, frozen processed order and empirical mean, p_i^1 = 1/(2K) + n_i(S)/(2T), surrogate gap 8 width_i(S), real-valued first phase target 1280/(p_i^1 Delta-tilde_i^2), zero error count, phase one, and empty sample/confidence sets. The bank update changes only the current eliminated set, preser…","url":"../implementation-map/index.html#delayed-sapo-line10-eliminated-arm-initialization","parent":"group:milestones","order":86,"meta":[["Book Map chapter","frontier"],["Lean declarations","31"],["Prerequisite milestones","3"]],"statement":"","missing":["The source EAP phase transition, integer stopping rule, sample/confidence-set evolution, and recursive probability bank.","BSC, the measurable generated Delayed-SAPO trajectory, D.4's simultaneous probability theorem, and both regret endpoints."],"search":"algorithm 5 line-10 eliminated-arm initialization delayed-sapo-line10-eliminated-arm-initialization a 31-declaration producer constructs the literal state assigned to each arm selected by algorithm 5 line 7: the elimination round, frozen processed order and empirical mean, p_i^1 = 1/(2k) + n_i(s)/(2t), surrogate gap 8 width_i(s), real-valued first phase target 1280/(p_i^1 delta-tilde_i^2), zero error count, phase one, and empty sample/confidence sets. the bank update changes only the current eliminated set, preserves surviving arms, and proves positive probability, surrogate gap, and phase target under the explicit nontrivial-horizon boundary 1 < t. a concrete two-arm canary exercises both the eliminated and surviving branches. this initializes eap data but does not execute eap or bsc, construct a randomized trajectory, prove d.4, or reach either paper endpoint. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","label":"Source-frozen delayed best-of-both-worlds endpoint audit","kind":"mathematical milestone","status":"partial","subtitle":"NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","description":"The exact NeurIPS 2025 source is hash-frozen, a same-algorithm multi-regime contract compiles, and 197 named source-audit declarations compile across accounting, causality, processing, allocation, elimination, source-scoped good-event projection/union assembly, the one-round action law, the 19-declaration D.10/D.12 diagnostic layer, a 16-declaration processed-prefix count-to-width producer, a 9-declaration processed…","url":"../implementation-map/index.html#neurips-2025-delayed-bobw-central-endpoints","parent":"group:milestones","order":87,"meta":[["Book Map chapter","frontier"],["Lean declarations","5"],["Prerequisite milestones","19"]],"statement":"","missing":["The complete Definition-D.1 event, its D.2–D.7 component probability producers, Delayed SAPO BSC/EAP phase transitions beyond the line-10 initializer, switching rule, and ordered update semantics.","A measurable causal randomized sampling kernel and recursively generated delayed-feedback trajectory law; only the one-round measure-valued rule now compiles.","A measurable generated-state trajectory and numerical BSC/EAP producer for the structural steps, the D.4 probability proof, a clarification or amendment of the printed D.10 transport, and an unconditional generated-trajectory D.12 / main-text Lemma 4.2 bridge.","The stochastic-instance and oblivious-adversarial regret endpoints for the same algorithm identity."],"search":"source-frozen delayed best-of-both-worlds endpoint audit neurips-2025-delayed-bobw-central-endpoints the exact neurips 2025 source is hash-frozen, a same-algorithm multi-regime contract compiles, and 197 named source-audit declarations compile across accounting, causality, processing, allocation, elimination, source-scoped good-event projection/union assembly, the one-round action law, the 19-declaration d.10/d.12 diagnostic layer, a 16-declaration processed-prefix count-to-width producer, a 9-declaration processed-trace-summary adapter, a 15-declaration ordered no-switch structural transition, a 12-declaration ordered trace, a six-declaration nonnegative-domain d.11 counting core, and a 31-declaration algorithm-5 line-10 initializer. the trace proves active-set monotonicity and the temporal factor-20 premise; the new numerical producer initializes only the first eliminated-arm eap state. banditrllib still does not execute eap/bsc, generate a measurable randomized trajectory, prove d.4's simultaneous 2/t probability bound, cover the switch branch, resolve d.13, or reach a regret endpoint. it therefore does not promote source lemmas d.4, d.10/d.12, d.13, main-text lemma 4.2, or theorem 4.1, and it does not yet prove the d.2–d.7 component bounds on a generated trajectory or either regret endpoint. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","label":"Source-frozen succinct lower-bound geometry audit","kind":"mathematical milestone","status":"partial","subtitle":"NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","description":"The exact Zeng–Honorio NeurIPS 2025 source is hash-frozen, and 54 named local declarations compile the nonempty symmetric unit-atom system, the literal succinct-support correlation contract, the source-shaped real-valued Q and R definitions, Definitions 3.1–3.3, and Lemmas 3.1–3.4. For the same vector, the strict-support route turns local R equality into unit correlations and applies finite Bessel to prove that a st…","url":"../implementation-map/index.html#neurips-2025-succinct-lower-bound-geometry-audit","parent":"group:milestones","order":88,"meta":[["Book Map chapter","frontier"],["Lean declarations","54"],["Prerequisite milestones","0"]],"statement":"","missing":["A source-faithful global repair for R: a spanning/nondegeneracy premise, an extended-real codomain, or restriction to the atom-generated span or quotient.","A source-faithful use of the repaired global R interface in the global primal/dual statements from Lemmas 3.5–3.6.","Assumption 3.7's grouped-support geometry, Theorem 3.8's action/parameter construction, its same-policy history-information chain, and both regret regimes."],"search":"source-frozen succinct lower-bound geometry audit neurips-2025-succinct-lower-bound-geometry-audit the exact zeng–honorio neurips 2025 source is hash-frozen, and 54 named local declarations compile the nonempty symmetric unit-atom system, the literal succinct-support correlation contract, the source-shaped real-valued q and r definitions, definitions 3.1–3.3, and lemmas 3.1–3.4. for the same vector, the strict-support route turns local r equality into unit correlations and applies finite bessel to prove that a strict representation uses no more atoms than any succinct representation; the two-direction consumer makes strict representation size unique. lean also proves a separate diagnostic: if a nonzero ambient vector is orthogonal to every atom, then the candidate set defining the paper's global real-valued r is unbounded. this exposes a regularity/codomain obligation without claiming that the paper is incorrect. the global lemmas 3.5–3.6, assumption 3.7, theorem 3.8, and every stochastic-bandit regret endpoint remain uncompiled. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","label":"Source-frozen stochastic-gradient-bandit Theorem-1 endpoint and Theorem-4 contract audit","kind":"mathematical milestone","status":"partial","subtitle":"NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","description":"The exact Baudry–Johnson–Vary–Pike-Burke–Rebeschini NeurIPS 2025 source is hash-frozen. Two hundred twenty-three named declarations comprise a frozen 215-declaration, twelve-layer Theorem-1 stack in the exact 26+18+18+14+4+10+3+25+19+9+37+32 split plus a separate eight-declaration Appendix-E/Theorem-4 source-contract audit. The first stack closes twoArmFixedIIDDirac_theoremOne: the source Theorem 1 for bounded two-a…","url":"../implementation-map/index.html#neurips-2025-stochastic-gradient-bandit-mechanism-audit","parent":"group:milestones","order":89,"meta":[["Book Map chapter","frontier"],["Lean declarations","223"],["Prerequisite milestones","0"]],"statement":"","missing":["The source-faithful two-arm learning-rate regimes and regret endpoints in Theorems 2–3.","Theorem 4 still requires the source-faithful general-K generated SGB process, its uniform buffered-event and survival-probability producer, the stopped-supermartingale/Doob route, and the final regret assembly; only the finite source-contract consumers compile.","Any broader non-Dirac environment-mixture or non-source extension must remain separate from the exact fixed-IID/Dirac Theorem-1 endpoint."],"search":"source-frozen stochastic-gradient-bandit theorem-1 endpoint and theorem-4 contract audit neurips-2025-stochastic-gradient-bandit-mechanism-audit the exact baudry–johnson–vary–pike-burke–rebeschini neurips 2025 source is hash-frozen. two hundred twenty-three named declarations comprise a frozen 215-declaration, twelve-layer theorem-1 stack in the exact 26+18+18+14+4+10+3+25+19+9+37+32 split plus a separate eight-declaration appendix-e/theorem-4 source-contract audit. the first stack closes twoarmfixediiddirac_theoremone: the source theorem 1 for bounded two-arm fixed-iid laws with exact arm means, a dirac environment prior, actual generated sampled pseudo-regret, 0 < delta < 1, eta > 0, eta c_eta < delta, and t = tailhorizon + 1. the separate gate proves only the positive equation-(22) drift margin, an audited finite survival-event lower bound under explicit premises, and finite geometric transient-phase bounds. it exposes an unresolved step-4 conditioning/direction mismatch but does not construct the general-k generated process, uniform buffered-event producer, stopped supermartingale/doob route, or theorem 4. theorems 2–4 remain open. mathematical milestone partial","shard":"views/milestones.json"},{"id":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","label":"Prospectively frozen SGB Corollary-1 and Theorem-2 follow-on","kind":"mathematical milestone","status":"partial","subtitle":"NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","description":"This follow-on preserves the historical 223-declaration SGB mechanism audit and adds 138 counted audit-slice declarations, for an exact 361 = 223 + 23 + 25 + 26 + 7 + 8 + 13 + 28 + 8 inventory. The counted layers compile Corollary 1, deterministic starvation and terminal-count consumers, chronological nth-pull infrastructure, latent product/readout, deferred-decisions prefix factorization, action/readout interfaces,…","url":"../implementation-map/index.html#neurips-2025-sgb-phase-transition-followon","parent":"group:milestones","order":90,"meta":[["Book Map chapter","frontier"],["Lean declarations","138"],["Prerequisite milestones","0"]],"statement":"","missing":["A bridge from the compiled terminal-count-below event to a fixed-cutoff starvation trigger/event; the exact probability split and missing-pull-to-terminal-count inclusion do not supply occurrence-conditioned IID.","The stopped-prefix future-cylinder law needed to prove conditional no-return probability at least one half; one-step fixed-history action kernels do not supply that law.","The Rademacher/binomial anti-concentration and ballot-prefix phase producer, followed by the finite-to-asymptotic polynomial-regret assembly.","The frozen terminal twoArmRademacherDirac_theoremTwo_polynomialRegret; Corollary 1 and the compiled nth-pull, latent product/readout, finite-prefix factorization, branch-locality, one-step freshness, full native-law, masked selected-block, exact phase-event, phase dichotomy, missing-pull-to-terminal-count, and deterministic starvation layers are not terminal evidence."],"search":"prospectively frozen sgb corollary-1 and theorem-2 follow-on neurips-2025-sgb-phase-transition-followon this follow-on preserves the historical 223-declaration sgb mechanism audit and adds 138 counted audit-slice declarations, for an exact 361 = 223 + 23 + 25 + 26 + 7 + 8 + 13 + 28 + 8 inventory. the counted layers compile corollary 1, deterministic starvation and terminal-count consumers, chronological nth-pull infrastructure, latent product/readout, deferred-decisions prefix factorization, action/readout interfaces, count-capped branch locality, and deterministic-time selected-reward freshness. a separate ten-declaration module identifies the complete visible/native trajectory law. the selected-block module now has 36 declarations: eight transport finite pull-time/reward blocks to an exact masked latent-coupling law, fourteen define and transport the exact finite appendix-c `s0/s1` event, ten split its pure latent probability into the generated all-present event plus an explicit missing-pull event, and four connect that branch to the generated finite-horizon terminal-count-below event and its nonnegative-gap expected sampled pseudo-regret consumer through the exact visible marginal. this transport does not prove positive missing-pull mass, a product or selected-iid theorem, or a fixed-cutoff starvation trigger; future/no-return, ballot probability, asymptotic assembly, and t…","shard":"views/milestones.json"},{"id":"milestone:TARGET-DRIFT-V2-CONTROLLED-EVALUATION","label":"Balanced target-drift controlled evaluation","kind":"mathematical milestone","status":"planned","subtitle":"TARGET-DRIFT-V2-CONTROLLED-EVALUATION","description":"Version 2 reuses 30 frozen source cases while balancing 75 source-faithful and 75 injected-drift target-replicate triplets across compile-only, source-aware blueprint, and full ABRL conditions. Both variants use one matched field/value template, with a frozen text-only leakage diagnostic. The result-free protocol and component-tested code specify a pre-audit common workspace, opaque identifiers, content-addressed se…","url":"../implementation-map/index.html#target-drift-v2-controlled-evaluation","parent":"group:milestones","order":91,"meta":[["Book Map chapter","frontier"],["Lean declarations","0"],["Prerequisite milestones","0"]],"statement":"","missing":["Freeze the provider, operator-attested immutable model version, exact Codex CLI provider-client bytes/version, auth-only runtime boundary, reasoning effort, service tier, sampling semantics, dated cache-read/cache-write/input/output token prices, replicate semantics, and token/tool/build/time/cost budgets; hash-seal the implemented missing-run policy and completion-ledger builder, keep the adapter to one CLI invocation, and explicitly freeze or disclose the provider-client-internal retry boundary.","Publish and freeze the final production checker image from the reviewed ephemeral candidate, freeze the provider image and commands, and reseal all seven bound checker probes under that final published image; the result-free candidate run passed the same seven probes but is not the final production seal.","Extend the result-free-only root PID-1/control/credential boundary candidate with a separately reviewed real-execution action, freeze its auth/visibility/network/tool/active-budget boundaries, publish and re-pull the final digest, and rerun all bound probes; only then run the implemented, hash-bound real-provider/real-sandbox one-case-by-three-condition smoke lane, whose operator-only plan and checker state permanently force result_eligible=false and exclude it from the 450, blind grading, and inferential analysis.","Complete the frozen-model source-absent wording control and independent blind wording review; freeze hash-verified source paths, grader identities, the sealed pack digest, and the resulting internal-pack and grader-only-export digests.","Complete the separately pinned and still-unrun LeanFlow external-system calibration: first freeze its adapter, schedule, fairness boundary, graders, and analysis, then pass its result-ineligible smoke, and only then execute the planned 30-run descriptive comparison without pooling it into the 450-run primary estimand.","Execute and neutrally check all 450 matched runs, complete independent grading, and analyze the frozen target-level endpoints."],"search":"balanced target-drift controlled evaluation target-drift-v2-controlled-evaluation version 2 reuses 30 frozen source cases while balancing 75 source-faithful and 75 injected-drift target-replicate triplets across compile-only, source-aware blueprint, and full abrl conditions. both variants use one matched field/value template, with a frozen text-only leakage diagnostic. the result-free protocol and component-tested code specify a pre-audit common workspace, opaque identifiers, content-addressed sealing, workflow-artifact presence/hash records, an atomic operator grading pack, a physically separate exact-positive-allowlist grader export, and source/target-aware analysis with multiplicity-controlled secondary endpoints. the grader-only export contains only normalized packets, the frozen prompt and rubric, a response template, and a digest manifest; it excludes the operator mapping, completion ledger, semantic labels, and execution/workflow metadata. before export or assembly, the internal pack is reconstructed byte for byte from the sealed run manifest, current complete ledger, and checked-run evidence, including its deterministic grade mapping. before inference, the analyzer reruns the hash-matched sealed assembler from the two grader responses and adjudication and requires an exact byte match with the supplied grade ledger. these are component-tested result-free controls; no gr…","shard":"views/milestones.json"},{"id":"milestone:CUCB-FULL-REPAIRED-RATES","label":"Full triggered CUCB approximation regret with explicit analysis repairs","kind":"mathematical milestone","status":"compiled","subtitle":"CUCB-FULL-REPAIRED-RATES","description":"One actual causal CUCB trajectory satisfies the refined gap-integral Theorem1 and both polynomial Theorem2 branches, with exact constants and H=1. Observed-law compatibility, inverse range and normalized analysis charge are explicit; signed alpha-beta benchmark and oracle-failure credit are retained. Independent semantic acceptance is scoped to the frozen finite-subset model, not all applications or topic completion.","url":"../implementation-map/index.html#cucb-full-repaired-rates","parent":"group:milestones","order":92,"meta":[["Book Map chapter","ucb"],["Lean declarations","3"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"full triggered cucb approximation regret with explicit analysis repairs cucb-full-repaired-rates one actual causal cucb trajectory satisfies the refined gap-integral theorem1 and both polynomial theorem2 branches, with exact constants and h=1. observed-law compatibility, inverse range and normalized analysis charge are explicit; signed alpha-beta benchmark and oracle-failure credit are retained. independent semantic acceptance is scoped to the frozen finite-subset model, not all applications or topic completion. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","label":"Native heterogeneous finite-node causal expected simple regret","kind":"mathematical milestone","status":"compiled","subtitle":"CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","description":"Reviewed native heterogeneous finite-node intervention sampling and expected simple regret, with binary reward readout. Parallel allocation and witnesses are now separately reviewed; source screening and ICLR evidence remain open, and the topic is incomplete. Actual joint and full sample-law pushforwards preserve estimator, ties, exact design cost and expectation. The final theorem constructs its finite codec intern…","url":"../implementation-map/index.html#causal-native-heterogeneous-expected-regret","parent":"group:milestones","order":93,"meta":[["Book Map chapter","ucb"],["Lean declarations","4"],["Prerequisite milestones","2"]],"statement":"","missing":[],"search":"native heterogeneous finite-node causal expected simple regret causal-native-heterogeneous-expected-regret reviewed native heterogeneous finite-node intervention sampling and expected simple regret, with binary reward readout. parallel allocation and witnesses are now separately reviewed; source screening and iclr evidence remain open, and the topic is incomplete. actual joint and full sample-law pushforwards preserve estimator, ties, exact design cost and expectation. the final theorem constructs its finite codec internally and retains the repaired coefficient and 1/t residual. the frozen noisy diagnostic supplement proves concentrated cost 8/3, exact biases, conditional/interventional distinction and tuned one-round expected regret 1/5. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:MULTI-AGENT-STATIC-COORDINATION-OCCUPATION","label":"Static Musical Chairs actual coordination and finite expected occupation","kind":"mathematical milestone","status":"compiled","subtitle":"MULTI-AGENT-STATIC-COORDINATION-OCCUPATION","description":"Actual finite coordination law, local collision-bit update, fixed-arm invariants, derived1/(4n) hazard and marginal survival give finite expected unfixed occupation<=4n². Candidate correctness is an intermediate input. The2U charge and coordination mean-regret bound are now separately reviewed; the complete unknown-N learner is now separately source-reviewed and being integrated; topic incomplete.","url":"../implementation-map/index.html#multi-agent-static-coordination-occupation","parent":"group:milestones","order":94,"meta":[["Book Map chapter","ucb"],["Lean declarations","3"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"static musical chairs actual coordination and finite expected occupation multi-agent-static-coordination-occupation actual finite coordination law, local collision-bit update, fixed-arm invariants, derived1/(4n) hazard and marginal survival give finite expected unfixed occupation<=4n². candidate correctness is an intermediate input. the2u charge and coordination mean-regret bound are now separately reviewed; the complete unknown-n learner is now separately source-reviewed and being integrated; topic incomplete. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:MULTI-AGENT-STATIC-COORDINATION-REGRET","label":"Static Musical Chairs actual finite expected coordination mean regret","kind":"mathematical milestone","status":"compiled","subtitle":"MULTI-AGENT-STATIC-COORDINATION-REGRET","description":"Actual real comparator deficit<=2U and support nonnegativity give finite expected common-set mean regret<=8n^2 through the constructed state/draw law. Exact real expectation conversion avoids positive-part substitution. This component takes common S; the separately reviewed full learner produces it from local exploration and transports the conditional law. Topic incomplete.","url":"../implementation-map/index.html#multi-agent-static-coordination-regret","parent":"group:milestones","order":95,"meta":[["Book Map chapter","ucb"],["Lean declarations","2"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"static musical chairs actual finite expected coordination mean regret multi-agent-static-coordination-regret actual real comparator deficit<=2u and support nonnegativity give finite expected common-set mean regret<=8n^2 through the constructed state/draw law. exact real expectation conversion avoids positive-part substitution. this component takes common s; the separately reviewed full learner produces it from local exploration and transports the conditional law. topic incomplete. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"milestone:MULTI-AGENT-STATIC-LEARNER-REGRET","label":"Static unknown-N learner: actual expected visible regret","kind":"mathematical milestone","status":"compiled","subtitle":"MULTI-AGENT-STATIC-LEARNER-REGRET","description":"Shared production chain passed local acceptance: primitive bounded reward laws and actual local estimation/ranking/coordination imply allH expected pseudo and unconditional visible regret<=min(nH,nS+8n^2+delta*nH). Repaired all-player budget explicit. Full noisy public-root canary and shared acceptance passed. Bounded recent-source selection and descriptive case accepted, with zero direct outside-case ABRL proof/val…","url":"../implementation-map/index.html#multi-agent-static-learner-regret","parent":"group:milestones","order":96,"meta":[["Book Map chapter","ucb"],["Lean declarations","3"],["Prerequisite milestones","3"]],"statement":"","missing":[],"search":"static unknown-n learner: actual expected visible regret multi-agent-static-learner-regret shared production chain passed local acceptance: primitive bounded reward laws and actual local estimation/ranking/coordination imply allh expected pseudo and unconditional visible regret<=min(nh,ns+8n^2+delta*nh). repaired all-player budget explicit. full noisy public-root canary and shared acceptance passed. bounded recent-source selection and descriptive case accepted, with zero direct outside-case abrl proof/value references. all-topic controlled evaluation and manuscript integration remain open. mathematical milestone compiled","shard":"views/milestones.json"},{"id":"spine:13","label":"Chapter 13 · Lower Bounds: Basic Ideas","kind":"source chapter","status":"compiled","subtitle":"Theorem 13.1 (source statement; proof deferred to Chapter 15)","description":"Theorem 13.1 compiles through Chapter 15 with c=1/54. Chapter 13 also compiles fixed-class minimax-optimality, the canonical iid Gaussian empirical-mean law, midpoint error events, the Chernoff companion, and both exact Mills-ratio bounds of Eq. (13.4) rescaled to the printed Eq. (13.1). The broader 1-subgaussian class with gaps in [0,1] now has a compiled fixed-horizon MOSS upper bound and constant-factor near-mini…","url":"../textbook-spine/chapter-13-basic-ideas/index.html","parent":"group:textbook-spine","order":0,"meta":[["Whole-chapter status","compiled"],["Source theorem","compiled"],["Lean correspondences","31"],["Open gaps","2"]],"statement":"","missing":["The frozen main-text contract is complete, including the broader finite-arm 1-subgaussian near-minimax consequence. Integrated proof, rendered export, structured review, PR, main, Pages and live acceptance are recorded for b38630c.","Optional: Notes 13.2 and Exercises 13.1–13.2 are not formalized and do not block the chapter contract."],"search":"chapter 13 · lower bounds: basic ideas theorem 13.1 (source statement; proof deferred to chapter 15) theorem 13.1 compiles through chapter 15 with c=1/54. chapter 13 also compiles fixed-class minimax-optimality, the canonical iid gaussian empirical-mean law, midpoint error events, the chernoff companion, and both exact mills-ratio bounds of eq. (13.4) rescaled to the printed eq. (13.1). the broader 1-subgaussian class with gaps in [0,1] now has a compiled fixed-horizon moss upper bound and constant-factor near-minimax theorem. the frozen main-text contract is complete: pr #105, authoritative-main checks, pages deployment and live desktop/mobile acceptance pass for b38630c. notes and exercises remain optional and unformalized. source chapter compiled","shard":"views/spine.json"},{"id":"spine:14","label":"Chapter 14 · Foundations of Information Theory","kind":"source chapter","status":"compiled","subtitle":"Theorem 14.2 / Eq. (14.7) (Bretagnolle–Huber)","description":"The frozen required body is compiled: Huffman optimality, exact-real arithmetic block coding and converse, finite/partition/common-density KL, the source affinity/overlap route, and Gaussian testing. Independent review, PR #106, main run 33959196451, Pages and live desktop/mobile acceptance passed. The singleton and uniform-code qualifications below are part of the accepted boundary; optional Notes/Exercises are not…","url":"../textbook-spine/chapter-14-information-theory/index.html","parent":"group:textbook-spine","order":1,"meta":[["Whole-chapter status","compiled"],["Source theorem","compiled"],["Lean correspondences","46"],["Open gaps","4"]],"statement":"","missing":["The frozen required-body contract passed independent review, main compilation and live publication acceptance; optional Notes/Bibliographic Remarks/Exercises are outside that contract. Full Exercise 14.10 is an additional compiled result.","Model qualification: singleton codewords are nonempty. Uniform fixed-length optimality is proved for power-of-two cardinalities, not arbitrary cardinalities: the ternary prefix code 0, 10, 11 has mean length 5/3 rather than 2.","Arithmetic coding is a classical exact-real construction with constant support/escape overhead, not an executable finite-precision encoder. Cross-entropy differences are unrounded; finite KL iff absolute continuity is finite-alphabet-only.","Adaptive-bandit history likelihood ratios and KL decomposition belong to Chapter 15; the scoped finite-arm same-policy identity now compiles there."],"search":"chapter 14 · foundations of information theory theorem 14.2 / eq. (14.7) (bretagnolle–huber) the frozen required body is compiled: huffman optimality, exact-real arithmetic block coding and converse, finite/partition/common-density kl, the source affinity/overlap route, and gaussian testing. independent review, pr #106, main run 33959196451, pages and live desktop/mobile acceptance passed. the singleton and uniform-code qualifications below are part of the accepted boundary; optional notes/exercises are not claimed complete. source chapter compiled","shard":"views/spine.json"},{"id":"spine:15","label":"Chapter 15 · Minimax Lower Bounds","kind":"source chapter","status":"compiled","subtitle":"Lemma 15.1 / Eq. (15.1) and Theorem 15.2","description":"The frozen required body (§15.1–15.2) is compiled: Lemma 15.1 and Theorem 15.2 use one arbitrary randomized HistoryAlgorithm, the canonical finite-history law, exact unit-Gaussian construction and 1/27 constant, with worst-case and minimax consequences. Optional Exercise 15.7 remains partial: measurable-map KL contraction reuses Chapter 14's trim API, and the fixed-horizon observation corollary compiles; stopped-his…","url":"../textbook-spine/chapter-15-minimax-lower-bounds/index.html","parent":"group:textbook-spine","order":2,"meta":[["Whole-chapter status","compiled"],["Source theorem","compiled"],["Lean correspondences","26"],["Open gaps","2"]],"statement":"","missing":["The Bernoulli refinement discussed in the §15.3 Notes and Exercises 15.1–15.6 and 15.8 are outside the current compiled slice.","Exercise 15.7 is not complete: generic data processing compiles, but the stopped-history information bound and F_tau-measurable factorization remain planned."],"search":"chapter 15 · minimax lower bounds lemma 15.1 / eq. (15.1) and theorem 15.2 the frozen required body (§15.1–15.2) is compiled: lemma 15.1 and theorem 15.2 use one arbitrary randomized historyalgorithm, the canonical finite-history law, exact unit-gaussian construction and 1/27 constant, with worst-case and minimax consequences. optional exercise 15.7 remains partial: measurable-map kl contraction reuses chapter 14's trim api, and the fixed-horizon observation corollary compiles; stopped-history information and f_tau factorization remain open. notes and other exercises are outside the required-body completion contract. source chapter compiled","shard":"views/spine.json"},{"id":"spine:16","label":"Chapter 16 · Instance-Dependent Lower Bounds","kind":"source chapter","status":"compiled","subtitle":"Definition 16.1, Theorem 16.2, Lemma 16.3, and Theorem 16.4","description":"Definition 16.1, Theorem 16.2, Lemma 16.3, and Theorem 16.4 compile with the source quantifiers, information branches, constants, and positive-part placement.","url":"../textbook-spine/chapter-16-instance-dependent/index.html","parent":"group:textbook-spine","order":3,"meta":[["Whole-chapter status","compiled"],["Source theorem","compiled"],["Lean correspondences","28"],["Open gaps","0"]],"statement":"","missing":[],"search":"chapter 16 · instance-dependent lower bounds definition 16.1, theorem 16.2, lemma 16.3, and theorem 16.4 definition 16.1, theorem 16.2, lemma 16.3, and theorem 16.4 compile with the source quantifiers, information branches, constants, and positive-part placement. source chapter compiled","shard":"views/spine.json"},{"id":"spine:17","label":"Chapter 17 · High-Probability Lower Bounds","kind":"source chapter","status":"compiled","subtitle":"Theorem 17.1 (stochastic high-probability lower bound)","description":"All Chapter 17 body endpoints pass the full Lean/Tests/harness gate, including same-policy hard-law coupling and deterministic matrix extraction; integrated into main via PR #101. Approved corrections: Claim 17.6 uses T_i ≤ n/2; Theorem 17.4 uses 0 < δ ≤ 1/32 with c=1/160, C=64 and a strict CDF tail. This is corrected-chapter closure, not a proof of the unchanged printed statements or every optional exercise.","url":"../textbook-spine/chapter-17-high-probability/index.html","parent":"group:textbook-spine","order":4,"meta":[["Whole-chapter status","compiled"],["Source theorem","compiled"],["Lean correspondences","16"],["Open gaps","0"]],"statement":"","missing":[],"search":"chapter 17 · high-probability lower bounds theorem 17.1 (stochastic high-probability lower bound) all chapter 17 body endpoints pass the full lean/tests/harness gate, including same-policy hard-law coupling and deterministic matrix extraction; integrated into main via pr #101. approved corrections: claim 17.6 uses t_i ≤ n/2; theorem 17.4 uses 0 < δ ≤ 1/32 with c=1/160, c=64 and a strict cdf tail. this is corrected-chapter closure, not a proof of the unchanged printed statements or every optional exercise. source chapter compiled","shard":"views/spine.json"}],"edges":[{"source":"library:banditrl","target":"group:book-map","relation":"contains"},{"source":"library:banditrl","target":"group:textbook-spine","relation":"contains"},{"source":"library:banditrl","target":"group:milestones","relation":"contains"},{"source":"library:banditrl","target":"group:proof-laboratory","relation":"contains"},{"source":"group:book-map","target":"chapter:foundations","relation":"contains"},{"source":"group:book-map","target":"chapter:probability","relation":"contains"},{"source":"group:book-map","target":"chapter:etc","relation":"contains"},{"source":"group:book-map","target":"chapter:ucb","relation":"contains"},{"source":"group:book-map","target":"chapter:oful","relation":"contains"},{"source":"group:book-map","target":"chapter:thompson","relation":"contains"},{"source":"group:book-map","target":"chapter:exp3","relation":"contains"},{"source":"group:book-map","target":"chapter:tsallis","relation":"contains"},{"source":"group:book-map","target":"chapter:finite-horizon-rl","relation":"contains"},{"source":"group:book-map","target":"chapter:frontier","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.ArmStreamPolicy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBActualReward","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBCharge","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBChargedConditional","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBChargedMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBConcentration","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBFiniteExample","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBGapCutoff","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBGapInverse","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBHistory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBNiceEvent","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBObservationMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBRegretTail","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBRewardKernel","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBRoundMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBSourceModel","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBThreshold","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBThresholdTail","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBTrajectory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBUnderCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalAllocation","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalAllocationRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalExpectedRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalHeterogeneous","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalImportance","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalImportanceTransport","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalMarginalLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalOrderedLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalParallelDesign","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalParallelLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalParallelRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalRecommendation","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalSampleMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalSampling","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.CausalTuning","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETC","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCCountLemmas","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCMeasurability","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCRegretLemmas","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCTrace","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","relation":"contains"},{"source":"chapter:etc","target":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOActualRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOConcentration","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOConditionalMGF","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOODepthOptimization","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOExpectedRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOExpectedVisits","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOHistory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOIndexConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOMeasurable","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOPathComparison","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOPrefix","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOORate","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOORegretAlgebra","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOORegretPartition","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOORewardFamily","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOSelectionTail","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOTrajectory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HOOTree","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailHistory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.HeavyTailUCB","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.KLUCBBernoulli","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSS","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSConditionalReward","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSConstants","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSHistory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSOccupancy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSOptimism","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSPeeling","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSRewardBranch","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSStream","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsCollision","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsExploration","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsRanking","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsRealized","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Algorithms.MusicalChairsReward","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.Thompson","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","relation":"contains"},{"source":"chapter:thompson","target":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCB","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamSource","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmStreamTail","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Automation","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.BoundedRewardKernelLaw","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.BudgetStoppingTime","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationCappedOccupancy","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationConditionalMGF","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationConfidenceSchedule","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationDyadicExponential","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationFixedMGF","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationGaussianOccupancy","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationIndexOccupancy","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationMartingaleMaximal","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationQuadraticMaximal","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationQuadraticScheduled","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationSubGaussian","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationTailIntegration","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConcentrationVariance","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalExpectationReward","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalRewardFoundation","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalRewardLawSource","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Core","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.CurvatureNoiseGapGeometry","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.Accounting","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.ActionLaw","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.CausalView","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.Elimination","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.Processing","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ActionProcess","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3BernsteinAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3BernsteinExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3BernsteinTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3BestArm","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ComparatorBernstein","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ComparatorConfidence","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ConditionalMoments","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ExpectedRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ExplorationBias","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3HedgeRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3HighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ImportanceWeighted","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernstein","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareConfidence","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3Potential","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PredictableAdversary","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PredictableHedge","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PredictableIntegration","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PredictableMoments","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PredictableRegretAllTime","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PureBernstein","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3PureConfidence","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedConcentration","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedConfidence","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedDeviationAllTime","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedDeviationTail","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedPredictableVariance","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedRegret","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RealizedRegretAllTime","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3RecursiveTrajectory","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3SampledHedge","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3SampledHistoryScore","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3ScoreRegularity","relation":"contains"},{"source":"chapter:exp3","target":"module:BanditRLProof.Exp3UniformRegret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationBochnerSums","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationFiniteBanditBounds","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationFoundation","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationPullCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationPullCountBounds","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationRegretPullCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationSums","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationWeightedPullCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ExpectationWeightedPullCountBounds","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.FTRLOneStep","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.FiniteArmRewardKernelLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.FiniteBanditModelInvariants","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.FiniteContextVarianceProxy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.FiniteGapCutoff","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.FiniteGapLayerCake","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.FiniteRealArgmax","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOCantorModel","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOCantorRate","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOODimension","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOGeometry","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOLevels","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOModel","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOOptimalBranch","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOPacking","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOPartition","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HOOTailSum","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailArmLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailClippedConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailClippedMoments","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailClippedScheduled","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailClippedTransfer","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailClipping","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailFixedTilt","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailGapThreshold","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailPowerSum","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailScheduledConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailSourceConfidence","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailSourceGap","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailSourceSchedule","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailTailSum","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailTruncation","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailTuning","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.HeavyTailUnshiftedMGF","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.HistoryFiltration","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.IndependenceFoundation","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.IntegrabilitySums","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.KernelIndependentExtension","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.KernelTrajectoryPrefix","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LeafLemmas","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.Literature","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.AffinityKL","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.BanditHistoryKL","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.BasicIdeas","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.BlockEntropy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.CodingEntropyBound","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.CommonDensityKL","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.CommonDomination","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.CrossEntropy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.DyadicAddresses","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.FinitePartitionKL","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.FixedLengthCoding","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.GaussianMinimax","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.GaussianTesting","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.HighProbability","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.HuffmanConstruction","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.HuffmanStep","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.InformationTheory","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.InstanceDependent","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.Minimax","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.PrefixCodePruning","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.ShannonLengths","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.LowerBounds.UniformCoding","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MartingaleDifference","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.MathlibWrappers","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasurableLocalQuantities","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasurablePullCount","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasurablePullCountCast","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasurableRegret","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasurableSums","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasureFoundation","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.MeasureL2Indicator","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULAllTimeConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULConfidenceEllipsoid","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULEllipticalPotential","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULEllipticalPotentialFoundation","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULExpectedRegretAsymptotics","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULExpectedRegretConsistency","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULExpectedRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULFiniteActionOptimism","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULFiniteHorizonScoreGram","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGaussianCovarianceMixture","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGaussianEvaluatedMixture","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGaussianMixture","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGaussianMixtureMeasurability","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGaussianSpectralMixture","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULHighProbabilityRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULInitialRoundGap","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULMeasurableRecursiveSelection","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULNormalizedRadiusWidth","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScalarRegularizationBias","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledAllTimeConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULSelectedWidthSummation","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULSelfNormalizedConfidence","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULSelfNormalizedMarkov","relation":"contains"},{"source":"chapter:oful","target":"module:BanditRLProof.OFULUniformTimeConfidence","relation":"contains"},{"source":"chapter:frontier","target":"module:BanditRLProof.OpenProblems","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.PolicyMeasurability","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.PosteriorKernel","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.PowerCutoffNormalization","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.PowerTailIntegral","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.ProbabilityUnionBound","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.PullCountDecomposition","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.PullCountReindex","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonMDP","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonOptimality","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonPolicy","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.FiniteHorizonTrajectory","relation":"contains"},{"source":"chapter:finite-horizon-rl","target":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.RatMeasurability","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.RealKernelRegretPullCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.RealMeanRegretPullCount","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.Regret","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.RegretCountBounds","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.RegretDecomposition","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.RewardKernel","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.RewardTraceLaw","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ScalarENNReal","relation":"contains"},{"source":"chapter:foundations","target":"module:BanditRLProof.ScalarPseudoRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisConjugatePotentialStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLConditionalStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLExpectedStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLInteriority","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLMinimizerExistence","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLOneStepStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFTRLStationarity","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisImportanceWeightedMoment","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartExpectedStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisRefinedSuboptimalStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisRegularizer","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledExpectedRegret","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledExpectedStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledIIDMeanGap","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledScoreAlignment","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSelfBounding","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","relation":"contains"},{"source":"chapter:tsallis","target":"module:BanditRLProof.TsallisTimeVaryingPenalty","relation":"contains"},{"source":"chapter:ucb","target":"module:BanditRLProof.UCBSummability","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","relation":"contains"},{"source":"chapter:probability","target":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.action_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.action_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.history_eq_trace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.measurable_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.measurable_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"declaration:BanditRLProof.ArmStreamPolicy.history_ucb","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_actual_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.actual_reward_expectation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.cumulative_actual_reward_expectation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_eq_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_actual_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_eq_gap_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.choose","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.choose_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.choose_sufficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_le_time","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_mono_step","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_causal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.charge_before_feedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.chargedObservations","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.chargedObservations_le_observationCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.chargedObservations_le_counters","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_le_observations_of_always_triggered","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_choose","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_counters","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_path_charge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.chargeValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.successValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.compensation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.exp_compensation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.compensation_abs_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.sum_chargeValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.sum_successValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_compensation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.integrable_exp_compensation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_counters_piLE","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.compensation_adapted","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.charged_successor_condMGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.charged_initial_MGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.charged_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"declaration:BanditRLProof.CUCB.ChargeData.observation_below_charged_half","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.prefixActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.prefixCounters","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.prefixCounters_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_prefixActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_prefixCounters","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.cucb_condExp_chargedTriggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"declaration:BanditRLProof.CUCB.ChargeData.cucb_initial_chargedTriggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"declaration:BanditRLProof.CUCB.ChargeData.chargedTriggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"declaration:BanditRLProof.CUCB.ChargeData.chargedTriggerFactor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"declaration:BanditRLProof.CUCB.ChargeData.chargedTriggerFactor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"declaration:BanditRLProof.CUCB.ChargeData.measurable_chargedTriggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"declaration:BanditRLProof.CUCB.ChargeData.integrable_chargedTriggerFactor_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"declaration:BanditRLProof.CUCB.ChargeData.roundKernel_chargedTriggerFactor_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.pathNoise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.pathCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.measurable_round_piLE","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.path_compensated_adapted","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.path_noise_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.path_noise_count_tail_optimized","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.path_negative_noise_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.path_negative_noise_count_tail_optimized","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.sum_pathCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.sum_pathNoise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.observed_sum_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"declaration:BanditRLProof.CUCB.observed_sum_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.observedNoise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.observedIndicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.observedCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.measurable_observedCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.observedCompensated_abs_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.integrable_exp_observedCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.exp_observedCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.cucb_condExp_observedCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.cucb_successor_condMGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"declaration:BanditRLProof.CUCB.cucb_initial_MGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConfidence","target":"declaration:BanditRLProof.CUCB.observationCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConfidence","target":"declaration:BanditRLProof.CUCB.pathDeviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConfidence","target":"declaration:BanditRLProof.CUCB.path_deviation_slice","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBConfidence","target":"declaration:BanditRLProof.CUCB.path_deviation_confidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"declaration:BanditRLProof.CUCB.cucbTrajectory_ae_round_property","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"declaration:BanditRLProof.CUCB.ChargeData.triggerSupport","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"declaration:BanditRLProof.CUCB.ChargeData.measurableSet_triggerSupport","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"declaration:BanditRLProof.CUCB.ChargeData.environment_ae_observed_of_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"declaration:BanditRLProof.CUCB.ChargeData.roundKernel_ae_triggerSupport_of_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_le_observations_ae_of_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.trueInput","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.expectedReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.expectedReward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerProbability_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerProbability_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.mem_triggerActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.triggerActions_nonempty","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.minTrigger_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.chargeData","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.chargeData_trigger_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.deterministic_counter_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.charged_observation_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"declaration:BanditRLProof.CUCB.FeedbackModel.nice_event_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","target":"declaration:BanditRLProof.CUCB.SourceModel.finite_power_sum_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeCount_power_sum_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.fullFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.fullEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_observed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_compatible","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_reward_integrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_reward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.fullModel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_expected_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.fullSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_minTrigger","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.full_globalMinTrigger","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.deterministic_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.Sample","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.sampleLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.bit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.outcome","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.selected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.feedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.environment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.law","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.sampleLaw_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.environment_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.law_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.environment_observed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.observation_compatible","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.reward_integrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.reward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.model","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.mean_value","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.expected_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.law_one_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.noisy_each_arm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.selected_size","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.choose","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.measurable_choose","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.choose_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.product_smooth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.source","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.trigger_value","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.minTrigger_value","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.globalMinTrigger_value","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.true_score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.true_optimum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.max_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.choose_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.initial_action_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.probabilistic_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.randomizedSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.randomized_action_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.randomized_probabilistic_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.noBadSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.no_bad_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.first_action_law","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.regret_one_positive","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"declaration:BanditRLProof.CUCB.FiniteExample.refined_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"declaration:BanditRLProof.CUCB.SourceModel.card_underChargeTimes_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"declaration:BanditRLProof.CUCB.SourceModel.sum_card_underChargeTimes_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight_le_cutoff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"declaration:BanditRLProof.CUCB.SourceModel.sum_underSampledGap_le_cutoff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_gap_cutoff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_large_cutoff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.thresholdCoefficient_antitone","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.gapDomain","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_unique","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_strictMono","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_antitone","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_intervalIntegrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.UnitOutcome","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.Feedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observationCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observationSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.empiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.upperIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observation_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observationCount_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observationSum_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observationSum_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.observationSum_le_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.empiricalMean_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.empiricalMean_unobserved","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.empiricalMean_observed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.empiricalMean_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.upperIndex_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.oracleInput","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.statistics_causal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.oracleInput_causal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.oracleInput_visible","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_observationCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_observationCount_real","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_observationSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_empiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_upperIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"declaration:BanditRLProof.CUCB.measurable_oracleInput","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","target":"declaration:BanditRLProof.CUCB.confidenceRadius_lt_half","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","target":"declaration:BanditRLProof.CUCB.SourceModel.not_bad_of_nice_and_sufficient_observations","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.confidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.measurable_confidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.count_mul_radius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.empirical_bad_implies_deviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.upperIndex_of_confidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.empirical_bad_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.round_confidence_arithmetic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.empirical_bad_probability_source","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.NiceEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.measurableSet_niceEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"declaration:BanditRLProof.CUCB.niceEvent_complement_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.marginalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.centeredFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.measurable_centeredFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.marginal_subgaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.integrable_centeredFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.integral_centeredFactor_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.observedSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.measurableSet_observedSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.ObservationCompatible","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.observed_centeredFactor_integrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.observed_centeredFactor_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.observedFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.observedFactor_eq_piecewise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.observedFactor_eq_exp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.integrable_observedFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"declaration:BanditRLProof.CUCB.integral_observedFactor_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"declaration:BanditRLProof.CUCB.SourceModel.continuous_score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"declaration:BanditRLProof.CUCB.SourceModel.continuous_optimum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"declaration:BanditRLProof.CUCB.SourceModel.measurable_joint_score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"declaration:BanditRLProof.CUCB.SourceModel.oracleSuccess","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"declaration:BanditRLProof.CUCB.SourceModel.measurableSet_oracleSuccess","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"declaration:BanditRLProof.CUCB.SourceModel.measurableSet_path_oracleSuccess","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.successIndicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.measurable_successIndicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.successIndicator_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_successIndicator_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.roundKernel_success_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.condExp_oracle_success","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.initial_oracle_success","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.expected_oracle_success","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"declaration:BanditRLProof.CUCB.SourceModel.oracle_failure_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.threshold_integral_le_power","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_threshold_integral_deterministic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_threshold_integral_probabilistic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_cutoff_regret_deterministic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.polynomial_cutoff_regret_probabilistic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.base_arm_count_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_linear_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_one_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.theorem_two_deterministic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.theorem_two_probabilistic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_polynomial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_polynomial_square","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseAt_polynomial_reciprocal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_deterministic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_probabilistic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"declaration:BanditRLProof.CUCB.SourceModel.gapThreshold_polynomial_upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.expected_underSampledGap_le_refined","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.theorem_one_refined_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.armRefinedTerm_eq_zero_of_no_bad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_nonpos_of_no_bad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.underSampledGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.sufficientIndicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.maxPositiveGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.underSampledGap_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.sufficientIndicator_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.gap_decomposition","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.measurable_actual_underSampledGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.measurableSet_sufficientSuccessfulCharge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_actual_underSampledGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_sufficientIndicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.expected_gap_decomposition","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_underSampled_add_sufficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretTail","target":"declaration:BanditRLProof.CUCB.SourceModel.expected_sufficientIndicator_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretTail","target":"declaration:BanditRLProof.CUCB.SourceModel.sum_inverse_square_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretTail","target":"declaration:BanditRLProof.CUCB.SourceModel.sum_expected_sufficientIndicator_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretTail","target":"declaration:BanditRLProof.CUCB.SourceModel.approximationRegret_le_underSampled_add_source_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_trueScore_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.environment_norm_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_round_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.round_reward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integral_round_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integral_round_norm_reward_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integral_round_reward_eq_score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integrable_joint_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"declaration:BanditRLProof.CUCB.SourceModel.integral_joint_reward_eq_score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.marginalMean_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.centeredFactor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.centeredFactor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.observedFactor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.observedFactor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.measurable_observedFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.integrable_observedFactor_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"declaration:BanditRLProof.CUCB.roundKernel_observedFactor_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.scoreOptimum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.score_le_optimum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.sourceGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.maxPositiveGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.gap_le_maxPositiveGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.bad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.inverseGap_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.chargeData","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.chargeData_sufficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.optimum_monotone","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.maxPositiveGap_eq_zero_of_no_bad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"declaration:BanditRLProof.CUCB.SourceModel.counters_zero_of_no_bad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"declaration:BanditRLProof.CUCB.SourceModel.SufficientSuccessfulCharge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_subset","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability_of_all_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"declaration:BanditRLProof.CUCB.SourceModel.sufficient_successful_charge_probability_source","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.thresholdCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.samplingThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.thresholdCoefficient_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.samplingThreshold_deterministic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.samplingThreshold_probabilistic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.normalizedCharge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.normalizedCharge_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.normalizedCharge_sufficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.exists_normalized_charge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"declaration:BanditRLProof.CUCB.exists_threshold_charge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.probabilistic_threshold_crossing","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.FeedbackModel.arms_nonempty","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.FeedbackModel.globalMinTrigger_eq_one_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_fixed_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_slice","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.TriggerShortfall","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_probability_of_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_union_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"declaration:BanditRLProof.CUCB.SourceModel.trigger_shortfall_union_probability_of_all_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.Round","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.Input","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.emptyFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.initialInput","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.oracleInput_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.roundKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.roundKernel_rectangle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.roundKernel_action_law","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.feedbackExtension","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.measurable_feedbackExtension","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.oracleInput_feedbackExtension","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.cucbStepKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.cucbTrajectory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.cucbStepKernel_apply_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.cucbTrajectory_prefix_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.cucbTrajectory_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"declaration:BanditRLProof.CUCB.cucbTrajectory_initial_law","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.triggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.measurable_triggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.triggerFactor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.triggerFactor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.triggerFactor_piecewise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.integrable_triggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.integral_triggerFactor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"declaration:BanditRLProof.CUCB.integral_triggerFactor_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_monotone","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_strict_of_charge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.ChargeData.counters_injOn_charges","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.ChargeData.card_charges_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeGapTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.card_underChargeGapTail_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.badActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes_mem_badActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeTimes_empty_of_no_bad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.minBadGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.maxBadGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"declaration:BanditRLProof.CUCB.SourceModel.badGap_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight_le_refined","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.armRefinedTerm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.underChargeWeight_le_armRefinedTerm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.underSampledGap_eq_sum_arms","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.sum_underSampledGap_eq_weights","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"declaration:BanditRLProof.CUCB.SourceModel.sum_underSampledGap_le_refined","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.secondMoment_ge_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.uniform_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.uniform_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.uniform_ratio_le_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.uniform_secondMoment_le_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.designCost","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.secondMoment_le_designCost","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.designCost_ge_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.uniform_designCost_le_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.convexLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.convexLaw_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.inverse_convex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.convexLaw_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.secondMoment_convex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.mixture_convexLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.designCost_convex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"declaration:BanditRLProof.Causal.design_sublevel_mass_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocationRegret","target":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_uniform","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalAllocationRegret","target":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_optimal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.estimateCenter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_centered_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_signed_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_simultaneous_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"declaration:BanditRLProof.Causal.GraphModel.sampleEstimate_source_confidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalExpectedRegret","target":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_le_tuned","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalExpectedRegret","target":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalExpectedRegret","target":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_source_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalExpectedRegret","target":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_explicit_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeTables","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeTables.prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.nodeJoint","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeAssignment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.decodeAssignment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.decode_encodeAssignment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeTables","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeAssignment_snoc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.joint_encodeTables","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.decode_joint_encodeTables","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.productNodeCodec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.nodeIntervene","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeTables_intervene","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.joint_intervene","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.encodeGraph","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.intervention_joint_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeCodec.encoded_joint_valid","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.nodeHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.ParentConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.parentConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.encodeParent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.decodeParent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.decode_encodeParent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.parentLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"declaration:BanditRLProof.Causal.NodeGraphModel.parentLaw_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","target":"declaration:BanditRLProof.Causal.nodeJoint_snoc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","target":"declaration:BanditRLProof.Causal.nodeJoint_factorization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","target":"declaration:BanditRLProof.Causal.nodeJoint_normalized","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","target":"declaration:BanditRLProof.Causal.nodeIntervention_factorization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.rewardMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.rewardMean_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleRecommendation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.simpleRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleRecommendation_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.simpleRegret_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.covers_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_source_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_explicit_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_uniform","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_optimal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.roundLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeRound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeCodec.encodeSamples","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.roundLaw_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleLaw_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleWeightedBit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleEstimate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleWeightedBit_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.sampleEstimate_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"declaration:BanditRLProof.Causal.NodeGraphModel.designCost_encoded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.mass_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.mass_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.sum_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.mixture","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.mixture_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.Covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.ratio","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.covered_cancel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.importance_identity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.positive_allocation_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.ratio_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.secondMoment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.ratio_second_moment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.truncationBias","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.truncatedMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.truncatedMean_add_bias","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.truncationBias_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.truncationBias_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.pairedLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.mixture_pairedLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.pairedLaw_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.weightedBit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.weightedBit_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.weightedBit_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.weightedBit_second_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"declaration:BanditRLProof.Causal.GraphModel.mixture_parent_joint","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.mass_map_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.ratio_map_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.mixture_map","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.covers_map_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.weightedBit_map_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.secondMoment_map_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"declaration:BanditRLProof.Causal.designCost_map_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.joint_map_init","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.take","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.Tables.restrict","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.joint_map_take","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.ParentConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.parentConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.parentTable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.table_eq_parentTable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.joint_map_last_pair","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.last_parent_joint","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.joint_map_node_pair","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.joint_map_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.parentLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.parentLaw_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.intervention_parent_joint","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.paired_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"declaration:BanditRLProof.Causal.GraphModel.intervention_parent_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.allocationOfWeights","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.allocationOfWeights_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.mass_mem_simplex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.mixtureWeight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.coordinateCost","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.coordinateCost_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.safeAllocations","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.continuous_mixtureWeight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.safeAllocations_compact","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.sublevel_mem_safe","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.uniform_mem_safe","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.safe_denominator_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.coordinateCost_continuousOn","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.safe_allocation_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.exists_optimal_allocation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.optimalAllocation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.optimalAllocation_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.optimalAllocation_cost_le_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.optimalAllocation_minimizes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.secondMoment_self","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"declaration:BanditRLProof.Causal.constant_designCost","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.Tables","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.Tables.prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.joint","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.intervene","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.intervene_none","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.intervene_at","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.joint_snoc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.joint_factorization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.joint_normalized","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.GraphModel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.GraphModel.doModel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.doModel_factorization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"declaration:BanditRLProof.Causal.intervention_incompatible_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rareIndices","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.exists_rarity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rarity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_minimal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.valueProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.valueProbability_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rarity_reciprocal_le_half","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rareActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rareActions_card_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rareWeight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.atomicTotal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.rareWeight_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.atomicTotal_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocationWeight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocationWeight_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocationWeight_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_empty_ge_half","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.covers_of_mass_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.ratio_le_of_mass_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"declaration:BanditRLProof.Causal.secondMoment_le_of_mass_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.mixture_mass_ge_component","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.rootTable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.rootTable_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.rootAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.rootLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.rootLaw_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.rootLaw_atom_mass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.atomic_observational_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocated_mass_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_cost_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"declaration:BanditRLProof.Causal.ParallelParameters.optimal_cost_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graphAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graphAction_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph_parentLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph_parentConfig_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph_allocation_covers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph_designCost","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph_allocation_cost_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.graph_optimal_cost_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"declaration:BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel_optimal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.maximizingActions","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.maximizingActions_nonempty","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.orderedArgmax","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.score_le_orderedArgmax","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.orderedArgmax_le_of_maximal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.rewardMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.rewardMean_eq_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.rewardMean_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.sampleRecommendation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.simpleRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.simpleRegret_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"declaration:BanditRLProof.Causal.GraphModel.simpleRegret_le_on_confidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampleMGF","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_mgf","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampleMGF","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_sum_mgf","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampleMGF","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_signed_sum_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.roundLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.roundLaw_observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.sampleLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.sampleLaw_observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.integral_sample_observation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_independent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"declaration:BanditRLProof.Causal.GraphModel.sampleWeightedBit_second_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceLog","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceLog_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceLog_union_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceThreshold_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceTilt_admissible","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceTilt_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceTilt_exponent_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceThreshold_scale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceThreshold_bias_scale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceRegret_scale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceLog_half_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"declaration:BanditRLProof.Causal.sourceResidual_le_scale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.Spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.exploreArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.exploreArm_val","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.exploreArm_eq_of_mod_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.exploreArm_eq_iff_mod_eq_val","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.exploreArm_add_K","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.CommitOracle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"declaration:BanditRLProof.ETC.obligationNames","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.score_le_foldl_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.argmax_cons_eq_some_foldl_rat_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.argmaxCommitOracle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.argmaxCommitOracle_argmax_finRange","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.argmaxCommitOracle_encode_le_of_score_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.argmaxCommitOracle_choose_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.argmaxCommitOracle_eq_arm_subset_empMean_ge_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_eq_arm_le_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm_of_argmaxOracle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_pastReward_iSup_infinitePi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_historyFiltrationSucc_infinitePi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.boundedRewardTraceSource_infinitePi_actionWithCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_infinitePi_bounded_actionMean_canonicalTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_infinitePi_bounded_actionMean_condSubGaussian_canonicalTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_infinitePi_bounded_actionMean_condSubGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.BoundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_integrable_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_integral_eq_zero_of_boundedRewardTraceSource_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.measurable_centeredReward_actionWithCommit_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.stronglyAdapted_centeredReward_actionWithCommit_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_actionWithCommit_integrable_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_actionWithCommit_succMartingaleDifferencePrefix_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_hasSubgaussianMGF_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_boundedRewardTraceSource_condSubGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.centeredRewardCondSubGaussianWitnesses_of_boundedRewardTraceSource_canonicalTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_boundedRewardTraceSource_condSubGaussian_canonicalTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource_condSubGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_boundedRewardTraceSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredRewardBoundVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredReward_integrable_of_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredReward_integral_eq_zero_of_integral_eq_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredReward_integral_eq_zero_of_mem_Icc_integral_eq_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredReward_hasSubgaussianMGF_of_mem_Icc_integral_eq_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_bounded_centered","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_centeredReward_subG","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_action_bounded_centered","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","target":"declaration:BanditRLProof.ETC.centeredDiffSubGaussianTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","target":"declaration:BanditRLProof.ETC.centeredDiffSubGaussianWitnesses_of_indep_subG","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiff_indep_subG","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","target":"declaration:BanditRLProof.ETC.iIndepFun_centeredPairwiseRewardDiff_of_iIndepFun_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiffVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail_of_reward_iIndepFun_centeredReward_subG","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.CenteredDiffSubGaussianWitnesses","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiffSubGaussianWitnesses","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.measurable_centeredPairwiseRewardDiff_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.stronglyAdapted_centeredPairwiseRewardDiff_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredReward_condExp_eq_zero_of_indep","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredReward_condExp_historyFiltrationSucc_eq_zero_of_indep","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_pastReward_iSup_of_iIndepFun_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.historyFiltrationSucc_actionWithCommit_le_pastReward_iSup","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.indep_centeredReward_succ_historyFiltrationSucc_of_iIndepFun_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_indep","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredReward_succ_condExp_historyFiltrationSucc_eq_zero_of_iIndepFun_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.hasCondSubgaussianMGF_of_indep_comap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredReward_succ_hasCondSubgaussianMGF_historyFiltrationSucc_of_iIndepFun_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasSubgaussianMGF_of_action_miss","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_miss","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_miss","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_arm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_action_eq_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_arm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_action_eq_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_of_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff_hasCondSubgaussianMGF_historyFiltrationSucc_of_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.CenteredRewardCondSubGaussianWitnesses","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.CenteredDiffCondSubGaussianWitnesses","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.centeredDiffCondSubGaussianWitnesses_of_centeredRewardCondSubGaussianWitnesses","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centeredDiffCondSubGaussianWitnesses","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_exploreArm_K_eq_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_exploreArm_add_K_eq_add_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_exploreArm_mul_K_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_exploreArm_explorationPulls_mul_K_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"declaration:BanditRLProof.ETC.empMeanAtExploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"declaration:BanditRLProof.ETC.empMeanAtExploration_eq_of_eq_on_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"declaration:BanditRLProof.ETC.empMeanAtExploration_completeRewardTrace_eq_of_explorationHorizon_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"declaration:BanditRLProof.ETC.empMeanAtExploration_eq_sumRewards_div_explorationPulls","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"declaration:BanditRLProof.ETC.empMeanAtExploration_le_iff_sumRewards_le_of_explorationPulls_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"declaration:BanditRLProof.ETC.empMeanAtExploration_ge_best_event_subset_sumRewards_tail_event_of_imp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","target":"declaration:BanditRLProof.ETC.measurable_sumRewards_actionWithCommit_exploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","target":"declaration:BanditRLProof.ETC.measurable_empMeanAtExploration_of_measurable_div_const","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","target":"declaration:BanditRLProof.ETC.measurable_empMeanAtExploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","target":"declaration:BanditRLProof.ETC.measurable_empMeanAtExploration_coordinates","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"declaration:BanditRLProof.ETC.sum_centeredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"declaration:BanditRLProof.ETC.centeredPairwiseGapThreshold_eq_explorationPulls_mul_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"declaration:BanditRLProof.ETC.canonicalSubGaussianArmPairwiseTailReal_eq_exp_neg_explorationPulls_mul_gap_sq_div_four_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_exp_neg_explorationPulls_mul_gap_sq_div_four_mul_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_explorationArgmaxAction_le_exploration_add_remaining_mul_exp_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"declaration:BanditRLProof.ETC.integrable_real_pullCount_actionWithCommit_choice_of_measurable_commit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_suffix_mul_commit_prob","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_eq_exploration_add_remaining_mul_commit_prob","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_actionWithCommit_choice_le_exploration_add_remaining_mul_of_commit_prob_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integrable_real_pseudoRegret_actionWithCommit_choice_of_measurable_commit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_sum_gap_mul_commit_prob","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_actionWithCommit_choice_le_exploration_add_suffix_badGap_prob","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.RewardKernel.pair_map_eq_compProd_of_map_eq_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.RewardKernel.condDistrib_pair_ae_eq_compProd_of_split","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.RewardKernel.condDistrib_ae_eq_const_of_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.RewardKernel.map_eq_of_condDistrib_ae_eq_const","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.RewardKernel.condDistrib_ae_eq_const_of_ae_eq_selected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.finiteArmCenteredRewardKernelLaw_of_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.finiteArmBoundedCenteredRewardKernelLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_boundedArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardProcess_sum_tail_ennreal_of_boundedArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_boundedArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredRewardCondSubGaussianWitnesses_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_boundedArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_pairwiseEmpMeanTailContract_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_prob_commit_eq_arm_le_pairwiseTail_of_boundedArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_prob_wrongCommit_le_pairwiseTailSum_of_boundedArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalBoundedArmWrongCommitTailBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalBoundedArmWrongCommitTailBudgetReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalBoundedArmMaxGapIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalBoundedArmPairwiseTailReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalBoundedArmPerArmIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalSubGaussianArmPairwiseTailReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.canonicalSubGaussianArmPerArmIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_canonicalSubGaussianArmPairwiseTailReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_eq_arm_le_canonicalBoundedArmPairwiseTailReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.real_measure_explorationArgmaxCommit_ne_bestArm_le_canonicalBoundedArmWrongCommitTailBudgetReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxPrefixRegretReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.measurable_explorationArgmaxPrefixRegretReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxPrefixRegretReal_finiteRewardHistoryOfTrace_generated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_eq_of_explorationPrefix_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_explorationPrefix_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_explorationPrefix_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_stepKernel_apply_eq_exploreArmLaw_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_initial_map_eq_explorationArm_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionRewardHistory_explorationArm_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmMaxGapIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalBoundedArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_reward_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal_of_actionDependent_actionRewardHistory_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.measurable_explorationArgmaxHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedAction_eq_explorationArgmaxAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedAction_eq_actionWithCommit_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxGeneratedActionPartialTrajectoryPairLawSource_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"declaration:BanditRLProof.ETC.explorationArgmaxHistory_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductArgmaxCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductArgmaxAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductMaxGapLintegralRegretBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductWrongCommitTailBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductBadGapLintegralRegretBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductSumGapLintegralRegretBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductWrongCommitTailBudgetReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductBadGapIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductSumGapIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.fixedProductMaxGapIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.explorationArgmaxCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.explorationArgmaxAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.explorationMaxGapIntegralRegretBoundReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapLintegralRegretBound_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapLintegralRegretBound_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.real_measure_fixedProductArgmaxCommit_ne_bestArm_le_fixedProductWrongCommitTailBudgetReal_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_badGap_prob_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductBadGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductSumGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxAction_le_explorationMaxGapIntegralRegretBoundReal_of_infinitePi_bounded_exploreMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_sumGap_prob_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_argmaxCommitOracle_actionWithCommit_le_exploration_add_suffix_maxGap_prob_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"declaration:BanditRLProof.ETC.lintegral_ofReal_pseudoRegret_fixedProductArgmaxAction_le_fixedProductMaxGapLintegralRegretBound_of_infinitePi_bounded_actionMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurableSet_commitArm_ne_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurable_empMeanVector_of_forall_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurable_commitOracle_choose_of_measurable_empMeanVector","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurable_commitOracle_choose_of_forall_measurable_empMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurableSet_commitOracle_ne_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurableSet_commitOracle_ne_bestArm_of_forall_measurable_empMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurableSet_empMean_ge_empMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.measurableSet_exists_ne_bestArm_empMean_ge_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.wrong_commit_subset_exists_empMean_ge_bestArm_of_commitOracle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_wrong_mean_events_of_subset","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_exists_ne_bestArm_empMean_ge_bestArm_le_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_wrong_mean_events","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_sum_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_sum_nonbest_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitArm_ne_bestArm_le_filtered_sum_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_filtered_sum_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"declaration:BanditRLProof.ETC.prob_commitOracle_ne_bestArm_le_sum_nonbest_pairwise_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_centered_subGaussian_event_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","target":"declaration:BanditRLProof.ETC.pairwiseEmpMeanTailContract_of_subGaussian_event_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","target":"declaration:BanditRLProof.ETC.PairwiseEmpMeanTailContract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_eq_arm_le_pairwise_tail_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_pairwise_tail_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.ratArmLawRealKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.ratArmLawRealKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.isMarkovKernel_ratArmLawRealKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.realKernelMean_ratArmLawRealKernel_eq_integral_cast","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.realKernelMean_ratArmLawRealKernel_eq_modelMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.ciSup_modelMean_cast_eq_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.realKernelGap_ratArmLawRealKernel_eq_modelGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_explorationArgmaxAction_le_exact_sum_of_armLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.argmax_cons_eq_some_foldl_real_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realArgmaxCommit_argmax_finRange","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realArgmaxCommit_encode_le_of_score_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.RealEncodedArgmaxCandidate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.exists_realEncodedArgmaxCandidate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmaxIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmaxIndex_candidate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_encode_eq_index","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_encode_le_of_isMax","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.realLeastEncodedArgmax_eq_realArgmaxCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.eventually_realExplorationArgmaxAction_eq_of_roundRobin_leastEncodedCommit_persist","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_leastEncodedCommit_persist","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realEmpMeanAtExploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realEmpMeanAtExploration_eq_sumRewards_div_explorationPulls","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.measurable_realEmpMeanAtExploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.real_score_le_foldl_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realArgmaxCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realArgmaxCommit_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realArgmaxCommit_const","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.measurable_selected_real_score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.measurable_foldl_real_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.measurable_realArgmaxCommit_of_forall_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.realExplorationArgmaxAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.measurable_realExplorationArgmaxCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_eq_exploration_add_remaining_mul_commit_prob","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_le_exploration_add_remaining_mul_of_commit_prob_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistoryPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistorySumRewards","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistoryEmpMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistoryPullCount_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistorySumRewards_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistoryEmpMean_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.pullCount_eq_of_eq_on_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.sumRewards_eq_of_action_eq_on_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.realHistoryEmpMean_exploration_eq_realEmpMeanAtExploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib_of_historyLeastEncodedCommit_persist","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realKernelBestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realKernelMean_le_realKernelBestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.ciSup_realKernelMean_eq_realKernelBestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realKernelGap_eq_realKernelBestArm_sub","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realCenteredPairwiseRewardDiff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realCenteredPairwiseGapThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realCenteredPairwiseRewardDiffVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.real_selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.real_meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.real_sumRewards_le_imp_centered_pairwise_sum_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit_eq_arm_event_subset_centeredPairwise_sum_event","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.iIndepFun_realCenteredPairwiseRewardDiff_of_iIndepFun_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.realCenteredPairwiseRewardDiff_hasSubgaussianMGF_of_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.sum_realCenteredPairwiseRewardDiffVarianceProxy_const_eq_two_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.real_measure_realExplorationArgmaxCommit_eq_arm_le_exp_of_infinitePi_kernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.integral_real_pullCount_realExplorationArgmaxAction_le_exp_of_infinitePi_kernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_infinitePi_kernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","target":"declaration:BanditRLProof.ETC.RealStationaryETCSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","target":"declaration:BanditRLProof.ETC.regret_le_of_realStationaryETCSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationRewardPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realRewardTraceOfExplorationPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationRewardPrefix_realRewardTraceOfExplorationPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.measurable_realExplorationRewardPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.measurable_realRewardTraceOfExplorationPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.measurable_realExplorationRewardPrefix_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.sumRewards_eq_of_eq_on_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realEmpMeanAtExploration_eq_of_eq_on_exploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit_eq_of_eq_on_exploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationArgmaxCommit_realRewardTraceOf_prefix_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationArgmaxAction_realRewardTraceOf_prefix_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.measurable_realKernelRegret_of_forall_measurable_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realKernelRegretOfExplorationPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.measurable_realKernelRegretOfExplorationPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realKernelRegret_realExplorationArgmaxAction_eq_prefixFunctional","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realKernelRegret_eq_of_action_eq_on_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.real_trajMeasure_const_eq_infinitePi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationPrefixOfFiniteRewardHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.measurable_realExplorationPrefixOfFiniteRewardHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.realExplorationPrefixOfFiniteRewardHistory_of_trace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_realExplorationArgmaxAction_le_exact_sum_of_prefixLaw_eq_infinitePi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_prefixLaw_eq_infinitePi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_initial_map_eq_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","target":"declaration:BanditRLProof.ETC.integral_realKernelRegret_externalAction_le_exact_sum_of_actionDependent_actionRewardHistory_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_exploreArm_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_le_sum_gap_mul_explorationPulls","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_suffix_count_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_add_suffix_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_eq_of_commitArm_eq_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_of_commitArm_eq_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_explorationPulls_mul_K_add_le_sum_gap_mul_explorationPulls_add_suffix_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"declaration:BanditRLProof.ETC.centeredPairwiseRewardDiff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"declaration:BanditRLProof.ETC.centeredPairwiseGapThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"declaration:BanditRLProof.ETC.selectedSubMean_sum_eq_sumRewards_sub_pullCount_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"declaration:BanditRLProof.ETC.meanSubSelected_sum_eq_pullCount_mul_sub_sumRewards","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"declaration:BanditRLProof.ETC.sumRewards_le_imp_centered_pairwise_sum_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"declaration:BanditRLProof.ETC.empMeanAtExploration_ge_best_event_subset_centered_pairwise_sum_event","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTrace","target":"declaration:BanditRLProof.ETC.actionWithCommit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTrace","target":"declaration:BanditRLProof.ETC.actionWithCommit_eq_exploreArm_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTrace","target":"declaration:BanditRLProof.ETC.actionWithCommit_eq_commitArm_of_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTrace","target":"declaration:BanditRLProof.ETC.actionWithCommit_eq_bestArm_of_commitArm_eq_bestArm_of_explorationPulls_mul_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_eq_pullCount_exploreArm_of_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.ratCast_pullCount_actionWithCommit_explorationPulls_mul_K_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_succ_eq_add_if_commitArm_of_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_of_ne","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"declaration:BanditRLProof.ETC.pullCount_actionWithCommit_explorationPulls_mul_K_add_eq_commitArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","target":"declaration:BanditRLProof.ETC.prob_argmaxCommitOracle_ne_bestArm_le_filtered_sum_centeredDiffSubGaussianTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","target":"declaration:BanditRLProof.ETC.pseudoRegret_actionWithCommit_choice_le_sum_gap_mul_explorationPulls_add_suffix_badGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"declaration:BanditRLProof.HOO.trajectory_reward_bounded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"declaration:BanditRLProof.HOO.integrable_trajectory_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"declaration:BanditRLProof.HOO.integral_trajectory_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"declaration:BanditRLProof.HOO.Covering.integral_actual_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.regionNoise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.regionCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.action_prefixExtension_at","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.measurable_action_piLE","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.measurable_coordinate_piLE","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_compensated_adapted","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_compensated_successor","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_compensated_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_noise_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.sum_regionCount_eq_visits","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.sum_regionNoise_eq_rewardSum_sub_means","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_negative_noise_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_noise_visits_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"declaration:BanditRLProof.HOO.region_negative_noise_visits_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.nodeMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.bounded_node_subgaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.bounded_node_fixedMGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.selectedNode","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.regionIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.regionSelected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.region_step_fixedMGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.region_step_compensated_integral_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.measurable_selectedNode","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.measurable_regionSelected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.measurable_regionIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.regionCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.measurable_regionCompensated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.integrable_regionCompensated_exp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.trajectory_region_condExp_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"declaration:BanditRLProof.HOO.trajectory_region_condMGF","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConfidence","target":"declaration:BanditRLProof.HOO.regionDeviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConfidence","target":"declaration:BanditRLProof.HOO.visits_history_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConfidence","target":"declaration:BanditRLProof.HOO.region_deviation_slice","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOConfidence","target":"declaration:BanditRLProof.HOO.region_deviation_confidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOODepthOptimization","target":"declaration:BanditRLProof.HOO.balance_identities","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOODepthOptimization","target":"declaration:BanditRLProof.HOO.exists_regret_depth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOODepthOptimization","target":"declaration:BanditRLProof.HOO.log_horizon_pos_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedRegret","target":"declaration:BanditRLProof.HOO.RegularCovering.integrable_actual_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedRegret","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_regret_partition_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedRegret","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_regret_dimension_sums","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.visitTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.measurable_visitTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.visitTrace_zero_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.visitTrace_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.lintegral_visits_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.visitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.visitThreshold_controls","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.RegularCovering.poor_region_lintegral_visits","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.integrable_visits","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"declaration:BanditRLProof.HOO.RegularCovering.poor_region_expected_visits","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.Observations","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.expanded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.visits","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.rewardSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.next","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.step","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.step_length","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.history_length","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.expanded_step","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.next_not_expanded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.next_ne_root","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.next_not_previously_played","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.visits_step","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.rewardSum_step","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.history_causal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.action_causal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.expanded_card_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.depthBound_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.action_depth_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.history_eq_ofFn","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.action_ne_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"declaration:BanditRLProof.HOO.action_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.regional_means_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.regional_means_upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.count_mul_width","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.upper_le_implies_lower_deviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.upper_underestimate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.width_le_half_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.upper_ge_implies_upper_deviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.upper_overestimate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.Covering.nodeMean_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.RegularCovering.optimal_descendant_mean_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.Covering.descendant_mean_upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.RegularCovering.optimal_upper_underestimate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"declaration:BanditRLProof.HOO.RegularCovering.poor_upper_overestimate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.measurable_backward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.measurable_walk","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.measurable_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.visits_ofFn","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.expanded_ofFn","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.measurable_next_ofFn","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"declaration:BanditRLProof.HOO.measurable_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPathComparison","target":"declaration:BanditRLProof.HOO.Covering.branch_root_optimistic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPathComparison","target":"declaration:BanditRLProof.HOO.Covering.selected_underestimate_implies_branch_underestimate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPathComparison","target":"declaration:BanditRLProof.HOO.Covering.history_selected_underestimate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"declaration:BanditRLProof.HOO.proper_prefix_walk_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"declaration:BanditRLProof.HOO.proper_prefix_select_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"declaration:BanditRLProof.HOO.expanded_history_prefix_closed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"declaration:BanditRLProof.HOO.bValue_le_upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"declaration:BanditRLProof.HOO.bValue_le_prefix_walk","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"declaration:BanditRLProof.HOO.visits_pos_mem_expanded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORate","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"declaration:BanditRLProof.HOO.regretSumConstant","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"declaration:BanditRLProof.HOO.regretSumConstant_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"declaration:BanditRLProof.HOO.nat_pow_rpow","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"declaration:BanditRLProof.HOO.regret_level_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"declaration:BanditRLProof.HOO.regret_sums_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"declaration:BanditRLProof.HOO.action_singleton_count_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.deep_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.shallow_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.bad_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.pathwise_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.boundary_expected_visits","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.Covering.familyNodeLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.Covering.familyNodeLaw_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.Covering.familyNodeLaw_eq_nodeLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.RegularCovering.discrete","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.familyDiscreteKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate_family","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret_family","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"declaration:BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate_family","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOSelectionTail","target":"declaration:BanditRLProof.HOO.RegularCovering.poor_region_selection_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.prefixExtension","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.measurable_prefixExtension","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.action_prefixExtension","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.stepKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.trajectory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.stepKernel_apply_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.trajectory_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.trajectory_prefix_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"declaration:BanditRLProof.HOO.trajectory_initial_law","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.Node","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.child","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.child_length","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.prefix_child","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.depthBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.length_le_depthBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.not_mem_of_depth_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.backward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.backward_not_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.backward_stable_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.bValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.bValue_not_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.bValue_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.preferred","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.preferred_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.walk","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.walk_not_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.prefix_walk","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.walk_length_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.select_not_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.select_length_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"declaration:BanditRLProof.HOO.bValue_le_walk","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robustMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robustMean_latent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robust_pullCount_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robustMean_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robustIndex_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robustMean_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robust_selected_gap_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robust_selected_small_radius_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robust_initial_count_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"declaration:BanditRLProof.HeavyTail.robust_large_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"declaration:BanditRLProof.HeavyTail.lintegral_pullCount_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"declaration:BanditRLProof.HeavyTail.robust_lintegral_count_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"declaration:BanditRLProof.HeavyTail.robust_integrable_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"declaration:BanditRLProof.HeavyTail.robust_integral_count_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.historyAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.historyReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.measurable_historyAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.measurable_historyReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.historyTruncatedMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.measurable_historyTruncatedMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.history_count_trace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.historyTruncatedMean_trace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.truncated_observed_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"declaration:BanditRLProof.HeavyTail.historyTruncatedMean_latent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegret","target":"declaration:BanditRLProof.HeavyTail.integrable_id_of_raw_moment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegret","target":"declaration:BanditRLProof.HeavyTail.robust_expected_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegret","target":"declaration:BanditRLProof.HeavyTail.robust_regret_integrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.mean_abs_le_raw_scale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.trace_regret_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.expected_normalized_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalizedValues","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalizedValues_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_ne_top","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.process_value_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extremeKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_moment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_best","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.extreme_trace_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.cap_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_two_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.selected_gap_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_selected_small_radius_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_initial_count_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_large_count_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.dropped","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.truncate_neg_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.dropped_filter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.sum_truncate_neg_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.dropped_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.suboptimal_index_gt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.optimal_index_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.deterministicStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.trace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.mean_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.mean_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.radius_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.source_index_sub_gt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.source_index_best_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.forced_suboptimal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.horizonHeight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.log_two_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.height_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.late_log_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.late_best_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.finite_count_obstruction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_raw_moment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.ae_deterministic","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_best_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.regret_eq_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.expected_regret_eq_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.printed_coefficient_counterexample","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.printed_gap_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.literal_printed_bound_false","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_lintegral_count_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integrable_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.sampleThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.historyIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.nextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.measurable_historyIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.measurable_nextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.measurable_robustAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustAction_initialization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustAction_maximizes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_latent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_pullCount_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustIndex_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.arm_adaptive_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.arm_adaptive_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_expected_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_regret_integrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.confidenceLog","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.confidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.historyIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.nextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.measurable_historyIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.measurable_nextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.robustAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.robustReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.measurable_robustAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.robustAction_initialization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"declaration:BanditRLProof.HeavyTail.robustAction_maximizes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.IsBernoulliParameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLCore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_of_not_left","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_of_not_right","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_zero_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_one_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_right_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_right_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_right_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_top_right_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_zero_left_of_interior","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_one_left_of_interior","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_self","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_eq_klFun","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_eq_of_interior","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLExpanded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_eq_expanded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.hasDerivAt_bernoulliKLExpanded_right","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.half_sq_sub_le_bernoulliKLCore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_le_sq_div","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.ennnreal_half_sq_sub_le_bernoulliKL","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_le_of_sq_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.bernoulliKL_self","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.continuousAt_bernoulliKLCore_right","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.confidenceSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.index","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.confidenceSet_bddAbove","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.mem_confidenceSet_self","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.mem_confidenceSet_of_natCast_mul_core_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.natCast_mul_half_sq_sub_le_budget_of_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.confidenceSet_nonempty","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.index_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.index_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.index_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.mem_confidenceSet_zero_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.index_zero_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.le_index_of_mem_confidenceSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"declaration:BanditRLProof.KLUCB.exists_mem_confidenceSet_of_lt_index","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedBudgetAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedBudgetAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedIndexAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.historyIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measurable_historyIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.historyNextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.historyIndex_le_nextArm_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.pairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.historyState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measurable_historyState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.historyPolicy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedAction_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.pairHistory_eq_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.historyIndex_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedIndexAt_le_selected_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedAction_succ_eq_initializationArm_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.successorArmPullCount_generatedAction_K_add_one_eq_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.successorArmPullCount_generatedAction_pos_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.K_le_of_generatedAction_selected_and_count_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.armMean_mem_confidenceSet_of_abs_lt_radius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.armMean_le_generatedIndexAt_of_abs_lt_radius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.margin_mul_gap_div_eight_le_radius_of_selected_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.pullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.pullCount_le_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.actionRewardHistoryStepKernelFamily_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedKLAllTimeBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedKLAllTimeBadEvent_subset_absBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.sumRewards_nonneg_of_mem_Icc_zero_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.sumRewards_le_pullCount_of_mem_Icc_zero_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.successorArmEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedEmpiricalMean_isBernoulliParameter_of_rewards_mem_Icc'","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measure_pullCount_gt_threshold_le_of_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.lintegral_pullCount_le_of_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.generatedRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measurable_generatedRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.pullCount_generatedRegretAction_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.allHorizonPullCount_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.actionRewardTrajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measure_allTimeBadEvent_le_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"declaration:BanditRLProof.KLUCB.measurable_generatedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.logPlus","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.radius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.index","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.logPlus_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.radius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.radius_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.radius_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.action_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.action_initial_arm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.action_index_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"declaration:BanditRLProof.MOSS.selected_index_gt_mean_add_half_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalHistory_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalHistory_empiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalAction_succ_eq_historyAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalAction_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.measurable_canonicalAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.measurable_canonicalReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.measurable_canonicalHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalHistory_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"declaration:BanditRLProof.MOSS.canonicalHistory_eq_of_eq_consumed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"declaration:BanditRLProof.MOSS.centeredRewardTable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"declaration:BanditRLProof.MOSS.integral_canonicalReward_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"declaration:BanditRLProof.MOSS.mean_add_centeredRewardTable_average","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"declaration:BanditRLProof.MOSS.streamTrace_pullCount_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"declaration:BanditRLProof.MOSS.canonicalReward_action_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSConditionalReward","target":"declaration:BanditRLProof.MOSS.map_condition_reward_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSConditionalReward","target":"declaration:BanditRLProof.MOSS.canonicalReward_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSConstants","target":"declaration:BanditRLProof.MOSS.log_sixtyFour_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSConstants","target":"declaration:BanditRLProof.MOSS.largeGap_constant_fifteen","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSConstants","target":"declaration:BanditRLProof.MOSS.largeGap_scaled_constant_fifteen","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.streamMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.fixedLogExceedanceCount_eq_fixedRadiusCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.integrable_indexExceedanceCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.integral_indexExceedanceCount_le_sharp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.integral_indexExceedanceCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.gap_mul_integral_indexExceedanceCount_le_sharp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"declaration:BanditRLProof.MOSS.gap_mul_integral_indexExceedanceCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","target":"declaration:BanditRLProof.MOSS.integral_streamTrace_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"declaration:BanditRLProof.MOSS.historyAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"declaration:BanditRLProof.MOSS.measurable_historyAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"declaration:BanditRLProof.MOSS.historyAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"declaration:BanditRLProof.MOSS.historyAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"declaration:BanditRLProof.MOSS.historyAction_initialization","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"declaration:BanditRLProof.MOSS.historyAction_index_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","target":"declaration:BanditRLProof.MOSS.canonical_initialPair_map","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","target":"declaration:BanditRLProof.MOSS.canonical_action_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","target":"declaration:BanditRLProof.MOSS.canonical_historySequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","target":"declaration:BanditRLProof.MOSS.map_canonicalHistory_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","target":"declaration:BanditRLProof.finiteHistoryPullCountENNReal_trace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","target":"declaration:BanditRLProof.MOSS.canonicalHistory_gapRegret_toReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","target":"declaration:BanditRLProof.MOSS.canonicalGapExpectedRegret_eq_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","target":"declaration:BanditRLProof.MOSS.canonicalGapExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.fixedLogRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.sampleRadius_le_fixedLogRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.indexExceedanceCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.pullCount_le_one_add_indexExceedanceCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.fixedLogExceedanceCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.smallSampleCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.indexExceedanceCount_le_small_add_fixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.smallSampleCount_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.smallSampleCount_le_inv_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"declaration:BanditRLProof.MOSS.indexExceedanceCount_le_inv_sq_add_fixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.centeredIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.optimismDeficit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.optimismDeficit_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.le_optimismDeficit_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.stronglyMeasurable_centeredIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.integrable_centeredIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.stronglyMeasurable_optimismDeficit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.integrable_optimismDeficit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.measure_optimismDeficit_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.integral_optimismDeficit_eq_integral_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.integral_optimismDeficit_le_two_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"declaration:BanditRLProof.MOSS.twice_horizon_mul_integral_optimismDeficit_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.logPlus_mono","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.exp_neg_logPlus_inv_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.peelingBarrier","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.blockBarrier","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.blockBarrier_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.blockBarrier_le_peelingBarrier","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.exp_neg_blockBarrier_sq_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.peelingSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.blockBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.measure_blockBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.scaledBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.measure_scaledBadEvent_le_fifteen","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.meanBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.meanBadEvent_subset_scaledBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"declaration:BanditRLProof.MOSS.measure_meanBadEvent_le_fifteen","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRegret","target":"declaration:BanditRLProof.MOSS.streamTrace_gapSum_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRegret","target":"declaration:BanditRLProof.MOSS.streamTrace_realMeanRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRegret","target":"declaration:BanditRLProof.MOSS.integral_largeGapCountSum_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.conditionCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.conditionBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.measurable_conditionCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.measurableSet_conditionBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.measurable_canonicalNextCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.canonicalReward_succ_eq_coordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.rebuilt_mem_conditionBranch_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.map_rebuilt_restrict_conditionBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"declaration:BanditRLProof.MOSS.map_condition_reward_restrict_branch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.neg_optimismDeficit_le_centeredIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.radius_eq_streamRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.streamEmpirical","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.pullCount_le_of_stream_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.streamCounts","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.streamTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.pullCount_streamTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.streamTrace_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"declaration:BanditRLProof.MOSS.streamTrace_pullCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","target":"declaration:BanditRLProof.MOSS.measurable_streamMean_at_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","target":"declaration:BanditRLProof.MOSS.measurable_action_of_state","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","target":"declaration:BanditRLProof.MOSS.measurable_streamCounts","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","target":"declaration:BanditRLProof.MOSS.measurable_streamTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","target":"declaration:BanditRLProof.MOSS.integrable_streamTrace_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalCondition","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalNextCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalNextCoordinate_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalHistory_eq_of_complement_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalNextCoordinate_eq_iff_insert","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalConditionWithout","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.canonicalCondition_eq_without","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.measurable_canonicalCondition","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.measurable_canonicalConditionWithout","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.indepFun_coordinate_canonicalConditionWithout","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"declaration:BanditRLProof.MOSS.map_canonicalConditionWithout_coordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_toMeasure_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.eventIndicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.eventIndicator_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_indicator_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_indicator_subGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_indicator_independent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.eventCount_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_centered_sum_subGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_centered_sum_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.FinitePMF.iid_frequency_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.collisionProbReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.collision_indicator_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.localCollisionRate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.localCollisionRate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.collisionBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.collisionBadEvent_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.allCollisionBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.allCollisionBadEvent_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.collision_exploration_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.allCollisionBadEvent_le_half_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate_eq_compl","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationRewardLaw_fst_event","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.statisticsBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.statisticsBadEvent_le_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_eq_compl","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationLength","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationLength_mean_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationLength_collision_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationLength_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_at_explorationLength","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.State","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.CollisionFree","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.step","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.DistinctFixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.action_of_fixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.step_preserves_fixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.step_fixed_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.new_fixed_collisionFree","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.step_distinct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.initial_distinct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.trajectory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.trajectory_distinct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.action_local","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.jointDraw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.transition","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.stateLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.transition_distinct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.stateLaw_distinct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.localUpdate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.collisionBit","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.step_localUpdate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.jointDraw_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.FixedWithin","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.step_within","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.stateLaw_within","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.jointDraw_rectangle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.jointDraw_rectangle_product","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.fixationWindow","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.fixationWindow_fixes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.transition_fixation_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.isolationWindow","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.isolationWindow_subset","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.jointDraw_isolation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.transition_common_fixation_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.exists_unoccupied","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_hazard","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.fixedPlayers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.hitFixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.safeFixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.fixed_action_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.hitFixed_partner","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.hitFixed_card_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.safeFixed_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.roundMeanReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.safeFixed_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.safeFixed_arms_subset","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.safeFixed_reward_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_le_twice_unfixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.expected_round_charge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_toReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_real_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.quarter_le_avoidance_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.quarter_le_avoidance","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.real_uniform_hazard_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.uniform_hazard_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_hazard_quarter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.event_add_compl","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.transition_fixed_no_return","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.unfixed_survival","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.quarter_rate_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.unfixed_survival_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.total_unfixed_occupation_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.unfixedCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.unfixedCount_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.expected_unfixed_eq_prob_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"declaration:BanditRLProof.MusicalChairs.expected_unfixed_occupation_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.iid","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.iid_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.iid_product_expectation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.eventCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.pow_eventCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.indicator_weight_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.FinitePMF.iid_count_pgf","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.explorationDraw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.explorationLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.observes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.observes_rectangle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.exploration_observes_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.exploration_collisionFree_split","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.exploration_collisionFree_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.exploration_collision_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.observationCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.collisionCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.observationCount_pgf","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.collisionCount_pgf","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.exploration_observes_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.observationCount_laplace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.ExplorationFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.explorationFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localObservationCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localCollisionCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localRewardSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localEmpiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localObservationCount_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localCollisionCount_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"declaration:BanditRLProof.MusicalChairs.localRewardSum_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.rankedArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.rankIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.rankedArm_rankIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.rankIndex_rankedArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.mem_topArms_iff_rankIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.ranked_score_antitone","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.boundaryGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.boundaryGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.boundaryGap_separates","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_orderStatistic_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.measurableSet_actualCandidate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.CandidateConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.measurable_learnedConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.configDrawPathKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_time_independent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.commonConfig","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedConfig_eq_common_on_good","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_on_good","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_fst_event","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_rectangle_on_good","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_good_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.trajectory_eq_of_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.extendedCoordinationDraws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.continuationAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.continuationAction_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.continuationAction_fixed_stays","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_positive","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_normalized_rectangle","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.ordered_set_sum_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.topArms_sum_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.collisionFreePlayers","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.earnedArms","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.collisionFree_action_injective","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.earnedArms_card_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.earnedArms_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.globalRoundRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.globalRoundRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.globalRoundRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_global","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerAction_explore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerAction_coordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.explorationPrefixRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerRegret_split","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.explorationPrefixRegret_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerRegret_bounds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerRegret_le_exploration_add","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerAction_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.integrable_learnerRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.integrable_pathRegret_projection","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerRegret_good_integral_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.learnerRegret_expected_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.source_conditional_learnerRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_coarse","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.coordination_residual_source_constant","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_published_residual","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_inverse_horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"declaration:BanditRLProof.MusicalChairs.source_conditional_learnerRegret_published_residual","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.FinitePMF.iid_sum_snoc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.FinitePMF.iid_snoc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.trajectory_snoc_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.trajectory_snoc_last","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_trajectory_stateLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.FinitePMF.iid_take","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.FinitePMF.sum_map","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.trajectory_take","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_trajectory_marginal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_last_state_draw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_round_state_draw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.FinitePMF.sum_map_real","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.FinitePMF.sum_bind_real","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.pathCoordinationRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_round_regret_expectation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_path_regret_expectation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_path_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_state_marginal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.learnedDrawPathKernel_round_marginal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_restricted_path","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_good_regret_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_conditional_coordination_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.iid_continuationAction_marginal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"declaration:BanditRLProof.MusicalChairs.source_conditional_coordination_regret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.base_gamma_upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.base_neg_gamma_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.avoidanceBase_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.avoidance_mass_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.collision_accuracy_sandwich","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.collision_accuracy_survival_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.population_inverse_margin","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.population_inverse_round","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationInverse","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.localPopulationEstimate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.localCollisionCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.noncollision_fraction_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationInverse_ge_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate_all_collision","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate_zero_collision","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate_correct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.localPopulationEstimate_correct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate_regular_cast","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.collisionCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.localPopulationEstimate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationRecovered","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.allCollisionAccurate_subset_populationRecovered","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.populationRecovered_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect_eq_inter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.measurableSet_explorationEstimatesCorrect","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.explorationStatisticsAccurate_subset_estimatesCorrect","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.scoreOrder","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.rankedArms","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.topArms","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.topArms_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.topArms_before_unselected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.topArms_eq_of_order_separated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.topArms_eq_of_strict_separation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.topArms_eq_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.populationEstimate_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.localCandidateSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_nonempty","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.empirical_strict_separation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_correct","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.localCandidateSet_eq_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.measurableSet_localScoreOrder","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_eq_comparisons","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.measurableSet_explorationGoodEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.explorationEstimatesCorrect_subset_goodEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.trueTopArms","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.trueTopArms_eq_of_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.explorationGoodEvent_trueTop_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.armMean_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.separating_gap_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"declaration:BanditRLProof.MusicalChairs.source_exploration_good_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_integrable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.scheduleRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.integrable_scheduleRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.scheduleRealizedRegret_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_scheduleRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.integrable_randomScheduleRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.randomScheduleRegret_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.explorationRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.continuationRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.exploration_schedule_pseudo","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.continuation_schedule_pseudo","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.integrable_explorationRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.explorationRealizedRegret_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.FullSample","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.completeLearnerLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.completeLearnerLaw_first_preserving","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.explorationContinuationLaw_first_preserving","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.completeLearnerLaw_exploration_preserving","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.integral_preserving_pullback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_continuationSchedule","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.integrable_realizedLearnerRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.continuationRealizedRegret_mean_joint","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret_expected_eq_pseudo","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.source_expected_realizedLearnerRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.continuationFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_explore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_coordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_arm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_collision","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_completed_exploration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnedConfig_from_feedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.continuationFeedback_local_update","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.learnerFeedback_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.finite_sum_split_at","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret_eq_feedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_feedback_mk","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_explorationFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_continuationFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_learnerFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_learnerTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.visibleLearnerLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.visibleRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_feedback_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.measurable_visibleRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.visibleRegret_pullback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.integrable_visibleRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.visibleRegret_expected_eq_realized","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"declaration:BanditRLProof.MusicalChairs.source_expected_visibleRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.rewardLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.armMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_preserving","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_bounded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.reward_coordinate_subGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.reward_time_independent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.selected_sum_subGaussian","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.selected_sum_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.selectedMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.selected_centered_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.selectedMean_tail_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.selectedMean_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.observedTimes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.localEmpiricalMean_eq_selectedMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.localEmpiricalMean_fixed_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.measurable_selectedMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.measurable_jointEmpiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.explorationRewardLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.meanBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.measurableSet_meanBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.meanBadEvent_mixture","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.meanBadEvent_count_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.half_le_one_sub_exp_neg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.count_mixture_exp_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.explorationProbReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.explorationProbReal_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.explorationProbReal_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.explorationProbReal_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.ofReal_explorationProbReal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.meanBadEvent_exponential_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.allMeanBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.allMeanBadEvent_exponential_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.mean_exploration_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.allMeanBadEvent_le_half_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.allMeanAccurate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.allMeanAccurate_eq_compl","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"declaration:BanditRLProof.MusicalChairs.allMeanAccurate_probability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxDenominator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxDenominator_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceIncrement_eq_indicator","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sum_sourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.policyValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gradientCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_eq_bestMean_sub_policyValue","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.gapExpectedIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapExpectedIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_ge_minGap_mul_failureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.gapExpectedIncrement_best_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_le_maxGap_mul_failureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.failureMass_eq_successFailure_add_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_exp_actionReward_le_sourceEqEight_of_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight_of_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_exp_actionReward_le_sourceEqEight_of_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmActionGap_le_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_le_gap_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSampledPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_gap_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceTheoremOne_margin_of_two_mul_eta_sourceC_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceTheoremOne_constant_le_inv_eta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneRate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneRate_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_horizon_eq_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_rate_eq_log","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_inv_eta_le_inv_log_two_mul_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_log_argument_le_horizon_pow_four","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_log_term_le_two_mul_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_gap_mul_horizon_le_exp_constant_mul_rate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneAbsoluteConstant","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_piecewise_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne_piecewise","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.two_mul_abs_pow_div_factorial_add_two_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceC","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_terms_summable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_mono","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_le_exp_two_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo_terms_summable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.exp_eq_one_add_self_add_expTailTwo","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo_le_of_abs_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.sq_div_two_mul_sourceC_abs_div_two","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.exp_mul_le_sourceEqEight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_ge","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourFiniteGeometricPhaseMass_le_inv","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourFiniteTransientMass_le_inv","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_fixedArmFinitePrefix_eq_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_dirac_eq_map_trajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.StochasticGradientBandit.stationaryRewardKernelAt_twoArmFixedIIDRewardKernel_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_latentCoordinate_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.stationaryRewardHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.historyStepKernel_stationaryRewardHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextPair_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleInitialPair_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.trajMeasure_map_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.frestrictLe_succ_eq_extendPairHistorySucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.nativeStationaryTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_map_frestrictLe_eq_native","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_frestrictLe_eq_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.UCB.extendArmStreamFinitePrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.UCB.measurable_extendArmStreamFinitePrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.UCB.extendArmStreamFinitePrefix_apply_of_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_prefix_next_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamFeedback_eq_of_withoutCoordinate_eq_of_selectedCoordinate_ne","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_eq_of_withoutCoordinate_eq_of_target_count_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamNextActionNeSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamInitialSafeArmSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamNextActionNeSet","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_restrict_nextActionNe_eq_of_withoutCoordinate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamVisiblePrefixNextAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamVisibleNextReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountCap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.singletonPairHistory_preimage_latentArmStreamPrefixCountCap_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCapLocality_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountLt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountLt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountEq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountEq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.realHistoryPullCount_extendPairHistorySucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.mem_latentArmStreamPrefixCountCap_extendPairHistorySucc_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCap_of_extendPairHistorySucc_mem","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.selectedCoordinate_ne_of_extendPairHistorySucc_mem_prefixCountCap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamSuccessorCountCap_preimage","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamSuccessorCountCapSection","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamSuccessorCountCapSection","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_restrict_successorCountCap_eq_of_withoutCoordinate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.LatentArmStreamVisiblePrefixNextActionBranchLocality","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality_of_prefixCountCapLocality","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod_of_locality","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamSelectedCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_branch_eq_prod","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_joint_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_condDistrib_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_joint_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixGeneratedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixGeneratedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixOptimalPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount_environmentPrefix_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInclusiveOptimalPullCountProcess","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmInclusiveOptimalPullCountProcess","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.isStoppingTime_twoArmNthOptimalPullTime","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullTime","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_eq_top_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_succ_of_nthOptimalPullTime_eq_top","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq_of_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_action_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmGeneratedReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_of_time_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmSuccessProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullSuccessProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_eq_of_time_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmHistoryEnvironment_ext","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullTimeRewardBlock","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullTimeRewardBlock","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmLatentMaskedOptimalPullBlock","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullTimeRewardBlock_eq_latentMasked_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNativeOptimalPullTimeRewardBlock_map_eq_latentMasked","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_snd_eq_nativeStationary","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_visible_eq_generated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPhaseOnePrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmAppendixCPhaseOnePrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCRewardPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCAllPullsPresent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCAllPullsPresent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCObservedPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCObservedPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCLatentPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCLatentPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock_preimage_appendixCObservedPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCGeneratedPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCGeneratedPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCPureLatentRewardEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCMissingPullLatentPhaseEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.mem_twoArmAppendixCMissingPullLatentPhaseEvent_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_probability_le_countBelow","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent_eq_union_phase_missing","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.disjoint_twoArmAppendixCLatentPhaseEvent_missing","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_pi","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_purePhaseEvent_eq_phase_add_missing","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_le_one_div_two_mul_nat_of_exp_two_mul_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_one_div_two_mul_nat_of_exp_parameter_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_le_one_div_two_mul_nat_of_time_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_exp_two_mul_parameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability_le_one_div_two_mul_horizon_of_parameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCGeneratedPhaseEvent_exists_lastPullTime","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneTriggerEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmGeneratedAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmTerminalOptimalPullCountEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmOptimalPullCountBelowEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_eq_iUnion_terminalCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneTriggerEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneStarvationEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.finTwo_eq_zero_or_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_suboptimalPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_horizon_sub_of_optimalPullCount_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent_sampledPseudoRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.mem_twoArmStepOneStarvationEvent_of_lowProbability_noFurtherOptimalPull","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_sampledPseudoRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret_of_finiteMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_charge_mul_probability_le_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_softmaxProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_sourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_historyParameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxFiniteActionDistribution","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.historySoftmaxDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.historyAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.historyAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_sourceIncrement_eq_expectedSourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_sourceIncrement_eq_expectedSourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_expectedSourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action_zero_given_environment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_isMarkov","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_initialFeedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDReward_aestronglyMeasurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zeroInitialization_finTwo","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardRecurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseRecurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardSuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseSuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardRecurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseRecurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.TwoArmBoundedFixedMeanEnvironmentContract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmEnvironmentPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNextPair","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmEnvironmentPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNextPair","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma_mono","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixFiltration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardTrajectorySuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseTrajectorySuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_inverseSuccessor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_zero_abs_le_one_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_succ_abs_le_one_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_abs_le_one_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_prefix_rewards_abs_le_one_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_le_abs_reward_of_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_softmax_le_abs_reward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmInitialPairKernel_sourceIncrement_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.abs_historyParameter_zeroInitialization_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessorPotential_eq_exp_historyParameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessorPotential_eq_exp_historyParameter","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardTrajectorySuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseTrajectorySuccessorPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_le_recurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_sum_eq_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_zeroInitialization_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_sum_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_one_eq_neg_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_one_eq_one_sub_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.finTwo_one_eq_neg_zero_of_sum_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.exp_two_mul_zero_mul_one_sub_softmaxProbability_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.exp_neg_two_mul_zero_mul_softmaxProbability_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_exp_two_mul_failure_eq_success","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_exp_two_mul_zero_eq_odds","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardQ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseQ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardQ_mul_reward_eq_sourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseQ_mul_reward_eq_sourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardEqEightRemainder_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseEqEightRemainder_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le_add_success_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le_sub_failure_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectorySourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectorySourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectoryParameterZero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectorySourceIncrement_eq_successFailure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialSourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialSourceIncrement","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialSourceIncrement_eq_quarter_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_eq_successFailureSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessFailureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessFailureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFailureMass_eq_successFailure_add_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessProbability_sq_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_le_source_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_source_log_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmActionGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmActionGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialActionGap_eq_half","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessorActionGap_eq_failureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_eq_generated","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_sourceTheoremOne","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInversePotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFailureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectoryParameterZero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInversePotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessProbability","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmFailureMass","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInversePotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessProbability_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmFailureMass_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessor_eq_nextPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessor_eq_nextPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardRecurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseRecurrenceBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardUnconditionalRecurrence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseUnconditionalRecurrence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmScalarForwardIterate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmScalarInverseTelescope","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqTelescope","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqSum_le_initial_div","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialForwardPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialInversePotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialForwardPotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialInversePotential","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardPotential_zero_eq_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInversePotential_zero_eq_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_zero_kernel_eq_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInversePotential_zero_kernel_eq_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardInitialUnconditionalRecurrence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseInitialUnconditionalRecurrence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration_from_source_initial","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PriorSketch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.obligationNames","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.bestAction_measurable_of_countable_env","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.ofCountableEnv","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.ofPosteriorMap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_eq_posterior_map","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_apply_eq_posteriorBest_map","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.PosteriorActionIdentityLedger.actionKernel_apply_singleton_eq_posteriorBest_preimage","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.BayesianPosteriorActionSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_bayesianPosteriorActionSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_posteriorMap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_bayesianPairMap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"declaration:BanditRLProof.Thompson.condDistrib_action_ae_eq_bestAction_of_canonicalPriorLikelihood","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.compProd_withDensity_left","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.comp_withDensity_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.map_swap_withDensity_snd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.compProd_eq_compProd_withDensity_snd_of_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.AlgorithmDensityPosteriorSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.algorithmDensityPosteriorSource_of_condDistrib_history_withDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.referencePosterior_ae_eq_condDistrib_of_algorithmDensitySource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_algorithmDensitySource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_condDistrib_history_withDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.HistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.HistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.historyStepKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.HistoryAlgorithmAbsolutelyContinuous","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.singletonPairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.pairHistoryPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.pairHistoryLast","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.measurable_singletonPairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.measurable_pairHistoryPrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.measurable_pairHistoryLast","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.pairHistoryPrefix_extendPairHistorySucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.pairHistoryLast_extendPairHistorySucc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.historyDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.measurable_historyDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.map_withDensity_comp","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.compProd_withDensity_withDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.kernel_withDensity_rnDeriv_eq_of_absolutelyContinuous","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.kernel_compProd_withDensity_left","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.historyStepKernel_eq_withDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.IsHistoryAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.isHistoryAlgorithmEnvironmentSequence_of_split","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.nextPairJointLaw_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.finitePairHistory_map_eq_withDensity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.ConditionalHistoryAlgorithmDensitySource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.condDistrib_finitePairHistory_eq_withDensity_of_conditionalProcessSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalProcessSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.ConditionalHistoryAlgorithmEnvironmentSplitSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.ConditionalHistoryAlgorithmDensitySplitSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmDensitySource_of_split","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.condDistrib_finitePairHistory_eq_withDensity_of_conditionalSplitSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_conditionalSplitSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.HistoryActionScore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.HistoryActionScore.atTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.HistoryActionScore.atBestTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.HistoryActionScore.measurable_atTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.HistoryActionScore.measurable_atBestTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.integral_comp_eq_of_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.integral_historyAction_eq_of_condDistrib_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_action_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryMeasure_map_action_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_map_action_zero_eq_bestAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.trajectoryHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.trajectoryBestHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_integral_historyScore_eq_bestAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.IsOptimalMeanSelector","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.trajectoryBayesMeanRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_historyScore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalActionKernelOnPair","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerEnv","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerEnv_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerHistory_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSamplerAction_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.map_compProd_comap_snd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSampler_env_history_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSampler_history_action_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_actionKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_bestAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectoryReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_canonicalHistoryTrajectoryAction_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_canonicalHistoryTrajectoryReward_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectory_initialPair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectory_step_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.initialAction_map_eq_of_historyAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.initialFeedback_condDistrib_of_historyAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.historyStepKernel_map_fst","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.policy_condDistrib_of_historyAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.feedback_condDistrib_of_historyAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.HistoryAlgorithmEnvironmentSplitSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.HistoryAlgorithmEnvironmentSplitSource.toSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSplitSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalHistoryAlgorithmEnvironmentSequence_of_split","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.kernelWithInput","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.kernelWithInput_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.condDistrib_id_fst_compProd_ae_eq_kernelWithInput","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.environmentTrajectoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.environmentTrajectoryReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_environmentTrajectoryAction_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_environmentTrajectoryReward_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.historyAlgorithmEnvironmentSequence_of_measure_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.mappedCanonicalHistoryAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.kernelWithInputHistoryAlgorithmEnvironmentSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmEnvironmentSequence_of_canonicalTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmEnvironmentSplitSource_of_canonicalTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.conditionalHistoryAlgorithmDensitySplitSource_of_canonicalTrajectoryKernels","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_canonicalTrajectoryKernels","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCB","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCBHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCB_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCB_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.finset_sum_sqrt_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.finset_sum_one_div_sqrt_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.sum_clippedUCB_action_sub_mean_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCBHistory_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.measurable_clippedUCBHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.measurable_uncurry_clippedUCBHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCBHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCBHistory_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCBHistoryScore_atTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.clippedUCBHistoryScore_atBestTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.integrable_of_measurable_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.integrable_trajectoryHistoryScore_clippedUCB","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.integrable_trajectoryBestHistoryScore_clippedUCB","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.integrable_trajectoryMean_bestAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.integrable_trajectoryMean_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.MeasurableHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.MeasurableHistoryEnvironment.at","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableEnvironmentInitialPairKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableEnvironmentHistoryStepKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableEnvironmentInitialPairKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableEnvironmentHistoryStepKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_measurableTrajectoryPrefixEnvironmentHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.retainEnvironmentKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.retainEnvironmentKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.retainEnvironmentKernel_map_snd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentStepKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentStepKernel_succ_apply_map_snd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableEnvironmentInitialStatePrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_measurableEnvironmentInitialStatePrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableEnvironmentPairTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurable_measurableEnvironmentPairTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment_initialStatePrefix","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.measurableTrajectoryPrefixEnvironment_ae_eq_of_traj","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_prefix_next_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_condDistrib_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical_of_step_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_apply_eq_canonical","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment_stepCondDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_measurableEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.uniformActionMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.absolutelyContinuous_uniformActionMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.uniformHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.historyAlgorithmAbsolutelyContinuous_uniform","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.trajectoryMixture_map_history_action_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.trajectoryMixture_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_history_action_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.finitePairReferencePosterior_ae_eq_condDistrib_of_conditionalProcessSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_initialAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.referencePosteriorHistoryAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.referencePosterior","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.referencePosterior_kernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.referenceActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerEnv","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerEnv_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerHistory_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySamplerAction_measurable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.map_compProd_comap_history","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySampler_base_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySampler_history_action_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySampler_condDistrib_action_ae_eq_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.policySampler_condDistrib_env_ae_eq_of_base","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.referencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"declaration:BanditRLProof.Thompson.finitePairReferencePolicySampler_condDistrib_action_ae_eq_bestAction_of_posterior_invariance","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.UnitArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.uniformUnitArmStreamMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryRewardSampler","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_uncurry_stationaryRewardSampler","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryRewardSampler_map_volume","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.rewardStreamOfUnitArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_rewardStreamOfUnitArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryRewardKernelAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryRewardKernelAt_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryArmStreamKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryArmStreamKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_initialFeedback","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_at_initialFeedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryMeasurableHistoryEnvironment_at_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamInitialReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamInitialReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamNextReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamNextReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_initialFeedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamNextReward_fixed","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_at_initialFeedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamMeasurableHistoryEnvironment_at_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_zero_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_succ_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measurable_rewardFromArmStream_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measure_pos_and_sumRewards_sub_pullCount_mul_ge_le_of_armStream_identDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measure_pos_and_pullCount_mul_sub_sumRewards_ge_le_of_armStream_identDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le_of_canonicalArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le_of_canonicalArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamTrajectoryAction_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamTrajectoryReward_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_pullCount_selectedArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_sumRewards_selectedArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_realEmpiricalMean_selectedArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.identDistrib_fst_latentArmStreamTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryReward_eq_rewardFromArmStream_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_sumRewards_sub_pullCount_mul_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pullCount_mul_sub_sumRewards_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pos_and_sumRewards_sub_pullCount_mul_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_pos_and_pullCount_mul_sub_sumRewards_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_sumRewards_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_sumRewards_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryKernel_pos_and_sumRewards_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryAction_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measurable_stationaryLatentArmStreamTrajectoryReward_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_sumRewards_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_sumRewards_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_compProd_le_of_forall_kernel_apply_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_pos_and_sumRewards_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_sq_div_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.sum_clippedCountWidthThreshold_tail_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_biUnion_clippedCountWidthThreshold_le_mul_sub_armPrefixSum_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_le_mul_mean_sub_sumRewards","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.clippedCountWidthThreshold_le_sumRewards_sub_mul_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_realEmpiricalMean_add_width_le_mean_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamPrior","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_map_prodAssoc_symm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_realEmpiricalMean_add_width_le_mean_le_nat_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_mean_le_realEmpiricalMean_sub_width_le_nat_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_selectedArm_realEmpiricalMean_add_width_le_mean_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_mean_bestAction_sub_clippedUCB_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_biUnion_clippedCountWidthThreshold_le_armPrefixSum_sub_mul_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_mean_le_realEmpiricalMean_sub_width_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.measure_latentArmStreamTrajectory_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_exists_arm_exists_mean_le_realEmpiricalMean_sub_width_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_clippedUCB_action_sub_mean_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le_of_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.Spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.IndexState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.score","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.score_eq_empiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScore","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScore_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.score_le_foldl_select","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.scoreArgmax","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.scoreArgmax_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxAction_score_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxAction_score_max_of_selected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.meanGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.meanGap_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_confidenceScore_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.not_two_radius_lt_meanGap_of_confidenceScore_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.upperConfidenceBad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lowerConfidenceBad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measurableSet_upperConfidenceBad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measurableSet_lowerConfidenceBad","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measurableSet_confidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.not_upperConfidenceBad_of_not_confidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.not_lowerConfidenceBad_of_not_confidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_not_confidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_confidenceBadEvent_le_sum_upper_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceBadEventAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measurableSet_confidenceBadEventAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.finiteHorizonConfidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.not_confidenceBadEventAt_of_not_finiteHorizonConfidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_not_finiteHorizonConfidenceBadEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.mem_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap_of_score_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.scoreMaxEvent_subset_finiteHorizonConfidenceBadEvent_of_two_radius_lt_meanGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_sum_upper_lower","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.upperConfidenceBad_subset_absDeviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lowerConfidenceBad_subset_absDeviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_absDeviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_absDeviation","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_absDeviation_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.chebyshevAbsDeviationTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_absDeviation_le_chebyshev_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_chebyshev_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail_le_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianBudgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianBudgetRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianBudgetRadius_sq_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail_budgetRadius_le_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_budgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_budgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_oneSided_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_exp_neg_budget_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_budgetRadius_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.exp_neg_log_eq_inv","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianLogBudgetRadius_sq_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianOneSidedDeviationTail_logBudgetRadius_le_inv_scale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_upperConfidenceBad_le_subGaussian_logBudgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_lowerConfidenceBad_le_subGaussian_logBudgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_logBudgetRadius_inv_scale_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianConstantLogBudgetRadius_sq_domination","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.constant_invScale_double_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_constantLogBudgetRadius_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.constant_invScale_double_sum_le_of_real","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.textbookDeltaScale","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.textbookDeltaScale_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.textbookDeltaScale_total_inv_budget_eq_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.constant_invScale_double_sum_textbookDeltaScale_le_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_textbookDeltaRadius_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_scoreMaxEvent_le_subGaussian_textbookDeltaRadius_delta_of_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedEvent_subset_scoreMaxEvent_of_action_score_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedEventOn_subset_finiteHorizonConfidenceBadEvent_of_action_score_max","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargeGapEvent_le_subGaussian_textbookDeltaRadius_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_subGaussian_textbookDeltaRadius_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.sum_measure_confidenceScoreArgmax_selectedLargeGapEventOn_le_card_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_selectedLargeGapCountOn_le_card_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_horizon_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_free_or_delta_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_freeBudget_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.freeTimes_indicator_sum_le_card","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedSmallPullCount_sum_eq_min_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedSmallPullCount_sum_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedSmallPullCount_indicator_sum_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_selectedSmallPullCount_indicator_sum_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedPullCount_sum_eq_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedPullCount_indicator_sum_eq_natCast_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.selectedPullCount_indicator_sum_eq_selectedSmall_add_selectedLargePullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.natCast_pullCount_le_threshold_add_selectedLargePullCount_indicator_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measurableSet_selectedLargePullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_selectedLargePullCount_indicator_sum_eq_sum_measure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_subGaussian_textbookDeltaRadius_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.sum_measure_confidenceScoreArgmax_selectedLargePullCountEvent_le_horizon_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_selectedLargePullCount_indicator_sum_le_horizon_mul_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_threshold_add_horizon_delta_of_selectedLargePullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_freeCard_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.mem_subGaussianTextbookDeltaRadiusChargedTimes_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.mem_subGaussianTextbookDeltaRadiusFreeTimes_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes_of_not_free","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusChargedTimes_gap_large","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusFreeCard_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadiusFreeTimes_card_le_threshold_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusThreshold_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_large_gap_of_lt_half_meanGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusHalfGapThreshold_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_sq_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_eight_mul_lt_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusEightProxyLogThreshold_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_lt_gap_sq_div","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusProxyThreshold_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_lt_half_meanGap_of_proxy_le_variance_div_count","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountThreshold_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianTextbookDeltaRadius_count_large_of_threshold_lt_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountLowerBound_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusRecursiveSampleCount_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmax_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_historyAction_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_generatedActionTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.identityActionPolicy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxGeneratedState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.confidenceScoreArgmaxGeneratedTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.lintegral_confidenceScoreArgmaxGeneratedTrace_pullCount_le_textbookDeltaRadiusSampleCountSource_add_horizon_delta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.subGaussianAbsDeviationTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_absDeviation_le_subGaussian_tail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.measure_finiteHorizonConfidenceBadEvent_le_subGaussian_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"declaration:BanditRLProof.UCB.obligationNames","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamPSeriesTerm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamPSeriesTailBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamPSeriesTerm_summable","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.indexTail_four_eq_coe_armStreamPSeriesTerm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.constSum_four_le_armStreamPSeriesTailBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.constSum_four_toReal_le_armStreamPSeriesTailBound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamAsymptoticModelCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamAsymptoticModelCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.lml_sum_four_le_armStreamAsymptoticModelCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamExpectedRegret_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamExpectedRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"declaration:BanditRLProof.UCB.armStreamExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamWithoutCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamInsertCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamNextCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamCoordinateOfHistoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistoryActionCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamSelectedRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamNextCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistoryActionFromWithout","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamWithoutCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamInsertCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamWithoutCoordinate_insertCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamHistoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamNextCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamCoordinateOfHistoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurableSet_armStreamHistoryActionCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurableSet_armStreamNextCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamNextCoordinate_eq_coordinateOfHistoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.pairwise_disjoint_armStreamHistoryActionCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.pairwise_disjoint_armStreamNextCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.iUnion_armStreamHistoryActionCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.iUnion_armStreamNextCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measure_eq_sum_restrict_armStreamHistoryActionCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamMeasure_eq_sum_restrict_nextCoordinateBranch","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.measurable_armStreamHistoryActionFromWithout","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamNextCoordinate_fst_eq_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamReward_succ_eq_nextCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistory_eq_of_eq_below_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistory_eq_of_withoutCoordinate_eq_of_pullCount_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamNextCoordinate_eq_iff_insertCoordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistoryAction_eq_fromWithout_of_nextCoordinate_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.iIndepFun_armStreamMeasure_coordinate","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.indepFun_armStreamMeasure_coordinate_without","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.indepFun_armStreamMeasure_coordinate_historyActionFromWithout","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_historyActionFromWithout_coordinate_eq_prod","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamHistoryActionFromWithout_mem_coordinateBranch_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.map_historyActionFromWithout_restrict_coordinateBranch_eq_historyAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_historyAction_reward_restrict_branch_eq_prod","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_historyAction_reward_succ_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamReward_succ_condDistrib_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment_feedback_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamMeasure_condDistrib_coordinate_given_without","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.armStreamReward_zero_condDistrib_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment_initialFeedback_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.armStreamAction_eq_initializationArm_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.pullCount_armStreamAction_K_eq_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.pullCount_armStreamAction_pos_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.K_lt_of_one_lt_pullCount_armStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.armStreamAction_eq_realIndexAction_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.realIndex_le_realIndex_armStreamAction_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.meanGap_le_two_realWidth_of_selected","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.pullCount_le_eight_scale_log_div_gap_sq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.realPullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.pullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.indexTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.constSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.selectedLargePullCountEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.lowerIndexFailure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.upperIndexFailure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.selectedLargePullCountEvent_subset_lower_union_upper","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.measure_selectedLargePullCountEvent_le_two_mul_indexTail","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.lintegral_selectedLargePullCount_indicator_sum_le_two_mul_constSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.lintegral_natCast_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.indexTail_ne_top","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.constSum_ne_top","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integrable_real_pullCount_armStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_real_pullCount_armStreamAction_le_threshold_add_two_mul_constSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.pullThreshold_cast_le_realPullThreshold_add_two","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_real_pullCount_armStreamAction_le_realThreshold_add_two_add_two_mul_constSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_sum_gap_mul_realThreshold_add_two_add_two_mul_constSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_lml_sum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.measurable_armStreamActionTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.measurable_realKernelRegret_actionTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.identDistrib_action_armStreamAction_of_identDistrib_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.identDistrib_action_of_identDistrib_actionRewardTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_identDistrib_actionRewardTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_common_actionReward_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.identDistrib_actionRewardTrace_of_condDistrib_eq_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.identDistrib_actionRewardTrace_of_split_condDistrib_eq_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_split_condDistrib_eq_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"declaration:BanditRLProof.UCB.integral_realKernelRegret_externalAction_le_lml_sum_of_condDistrib_eq_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.finiteArmRealRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.finiteArmRealRewardKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.finiteArmRealRewardKernel_isMarkov","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamFiniteArmSubgaussianExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamBoundedFiniteArmExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.armStreamArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.initializationArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.initializationArm_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.realHistoryNextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.measurable_realHistoryNextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.measurable_armRewardStream_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamAction_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamAction_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamHistory_eq_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.measurable_armStreamHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.measurable_armStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.measurable_armStreamReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamAction_succ_eq_realHistoryNextArm_actualHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamAction_succ_eq_realHistoryIndexAction_of_not_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamUCBFixedArmPrefixSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.armStreamMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_armStreamUCB_mem_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"declaration:BanditRLProof.UCB.rewardFromArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"declaration:BanditRLProof.UCB.sumRewards_rewardFromArmStream_eq_armPrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"declaration:BanditRLProof.UCB.fixedArmPrefixSourceOfArmStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_rewardFromArmStream_mem_le_identDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"declaration:BanditRLProof.UCB.canonicalFixedArmPrefixSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_rewardFromCanonicalArmStream_mem_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.armStreamMeasure_map_coord","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.iIndepFun_armStreamMeasure_coord_sub","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.iIndepFun_armStreamMeasure_sub_coord","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.hasSubgaussianMGF_armStreamMeasure_coord_sub","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.hasSubgaussianMGF_armStreamMeasure_sub_coord","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.sum_coord_sub_eq_armPrefixSum_sub","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.sum_sub_coord_eq_mul_sub_armPrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_armPrefixSum_sub_mul_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_mul_sub_armPrefixSum_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.armPrefixEmpiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.armPrefixAverageConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_armPrefixAverageConfidenceRadius_le_abs_empiricalMean_sub","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.upperDeviationPairs","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.lowerDeviationPairs","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.positiveUpperDeviationPairs","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.positiveLowerDeviationPairs","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.mem_fst_image_upperDeviationPairs","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.mem_fst_image_lowerDeviationPairs","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.mem_fst_image_positiveUpperDeviationPairs_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.mem_fst_image_positiveLowerDeviationPairs_iff","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_sumRewards_sub_pullCount_mul_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_pullCount_mul_sub_sumRewards_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_pos_and_sumRewards_sub_pullCount_mul_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_pos_and_pullCount_mul_sub_sumRewards_ge_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.countWidthThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.countWidthThreshold_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.countWidthThreshold_le_mul_mean_sub_sumRewards_of_empiricalMean_add_width_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.countWidthThreshold_le_sumRewards_sub_mul_mean_of_mean_le_empiricalMean_sub_width","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.countWidthThreshold_sq_div_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.positiveCountFilter_eq_Icc","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.sum_countWidthThreshold_tail_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.natCast_mul_exp_neg_log_le_inv_rpow_sub_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean_log_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth_log_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_realEmpiricalMean_add_realWidth_le_mean_rpow_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"declaration:BanditRLProof.UCB.measure_mean_le_realEmpiricalMean_sub_realWidth_rpow_bound","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorArmwiseBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_boundedFiniteArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_armwiseBoundedFiniteArmLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedPseudoRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorBoundedFiniteArmExpectedAveragePseudoRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorEmpiricalMeanAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRadiusAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorIndexAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorInitializedScoreMaxSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorLargeGapEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorInitializedScoreMaxSource.meanGap_le_two_radius_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_centeredKernel_of_variance_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorRewardMapLaw","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"declaration:BanditRLProof.UCB.integrable_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.SelectedPolicySuccessorFiniteHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.completeFinitePairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.completeFinitePairHistoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.completeFinitePairHistoryReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.completeFinitePairHistory_finitePairHistoryOfTrace_apply_of_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryNextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex_le_nextArm_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPairHistory_eq_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.sumRewards_eq_of_forall_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmPullCount_completeFinitePairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmRewardSum_completeFinitePairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorHistoryIndex_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBAction_succ_eq_initializationArm_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_K_add_one_eq_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_pos_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBInitializedScoreMaxSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.exists_selected_with_threshold_le_prior_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRealPullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPullThreshold_contracts","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.two_mul_successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_global_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanFiniteArmTimePeelingRadius_lt_gap_of_explicitPullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.K_le_of_selectedPolicySuccessorGeneratedUCBAction_selected_and_count_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_largeGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_of_global_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_of_largeGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_threshold_le_ennreal_delta_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorGeneratedUCBAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.lintegral_natCast_le_threshold_add_bound_mul_of_measure_gt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_largeGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_threshold_add_horizon_mul_delta_of_global_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_largeGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.modelMeanGap_bestArm_eq_realGap","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorPullThreshold_cast_le_realThreshold_add_two","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.gap_mul_selectedPolicySuccessorPullThreshold_cast_le_textbookGapBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.ofReal_gap_mul_selectedPolicySuccessorPullThreshold_le_textbookGapBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.sum_gap_mul_explicitThreshold_add_failure_le_textbookGapSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorGeneratedUCBRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.pullCount_selectedPolicySuccessorGeneratedUCBRegretAction_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_le_sum_gap_mul_bound_of_positiveGap_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRewardStepKernelFamily","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"declaration:BanditRLProof.UCB.isMarkovKernel_selectedPolicySuccessorRewardStepKernelFamily","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorRewardTrajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBSelectedRewardLawSource_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCB_reward_map_eq_selected_policy_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"declaration:BanditRLProof.UCB.measure_actionRewardHistoryStepKernelFamily_selectedPolicySuccessorLargeGapEvent_le_ennreal_delta_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"declaration:BanditRLProof.UCB.lintegral_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_le_explicitPullThreshold_add_horizon_mul_delta_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_explicitThresholdSum_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticDelta","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.horizon_mul_selectedPolicySuccessorAsymptoticDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmTimeLogBudget_asymptoticDelta_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticGapCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapBudget_add_failure_asymptoticDelta_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticModelCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTextbookGapSum_asymptoticDelta_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorAsymptoticModelCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.one_add_log_natCast_succ_isBigO_log_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.log_natCast_succ_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedPseudoRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorActionRewardTrajMeasureExpectedAveragePseudoRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","target":"declaration:BanditRLProof.UCB.actionRewardTrajectorySuccessorAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorGeneratedUCBRegretAction_ae_eq_actionRewardTrajectorySuccessorAction_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_actionRewardTrajectorySuccessorAction_le_textbookGapSum_actionRewardTrajMeasure_centeredKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentBoundedRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_contextDependentSubgaussianRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_trajMeasure_finiteContextDependentSubgaussianRewardKernel_without_proxy_positivity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.finiteArmSubgaussianInitialActionRewardMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.finiteArmSubgaussianInitialActionRewardMeasure_isProbabilityMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.ArmRewardStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.armPrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.measurable_armPrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.FixedArmPrefixSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.FixedArmPrefixSource.measurable_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.FixedArmPrefixSource.measurable_armPrefixSum","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"declaration:BanditRLProof.UCB.measure_pullCount_prod_sumRewards_mem_le_of_fixedArmPrefixSource_identDistrib","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingRadiusAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingRadiusAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingIndexAt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingHistoryIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryNextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex_le_nextArm_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingHistoryState","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction_succ","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory_eq_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryIndex_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingIndexAt_le_generatedAction_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBAction_succ_eq_initializationArm_of_lt","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_K_add_one_eq_one","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.successorArmPullCount_selectedPolicySuccessorTelescopingGeneratedUCBAction_pos_of_K_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.K_le_of_selectedPolicySuccessorTelescopingGeneratedUCBAction_selected_and_count_pos","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingGeneratedUCBAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingLogBudget_mono","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.successorArmEmpiricalMeanTelescopingPeelingRadius_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingRealPullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPullThreshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPullThreshold_contracts","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.two_mul_successorArmEmpiricalMeanTelescopingPeelingRadius_lt_gap_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_meanGap_le_two_radius_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_pullCount_le_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.actionRewardHistoryStepKernelFamily_selectedPolicySuccessorTelescoping_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_of_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.lintegral_selectedPolicySuccessorTelescoping_pullCount_le_of_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measurable_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.pullCount_selectedPolicySuccessorTelescopingGeneratedUCBRegretAction_eq","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_of_allTimeConfidence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_allHorizonPullCount_of_not_badEvent","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingActionRewardTrajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_allTimeBadEvent_le_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_pullCount_gt_threshold_le_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realEmpiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realWidth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryWidth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realIndexAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryIndexAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realHistoryPullCount","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realHistorySumRewards","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realHistoryEmpMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realHistoryWidth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realHistoryIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realHistoryIndexAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realEmpiricalMean","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realWidth","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realIndex","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realIndexAction_spec","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.measurable_realIndexAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryEmpiricalMean_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryWidth_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryIndex_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"declaration:BanditRLProof.UCB.realHistoryIndexAction_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","target":"declaration:BanditRLProof.UCB.RealStationaryUCBSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","target":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","target":"declaration:BanditRLProof.UCB.identDistrib_actionRewardTrace_of_realStationaryUCBSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","target":"declaration:BanditRLProof.UCB.regret_le_of_realStationaryUCBSequence","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.finiteArmCanonicalKernelTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_canonicalKernelTrajectory","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalRealUCBHistorySelector","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalRealUCBPolicyKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.measurable_canonicalRealUCBHistorySelector","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalRealUCBPolicyKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm_initialAction_eq_dirac","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalArmStreamHistoryAlgorithm_policy_ae_eq_explicitPolicyKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectory_finitePairHistory_map_eq_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_condDistrib_ae_eq_explicitPolicyKernel","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_zero_ae_eq_initializationArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_succ_ae_eq_realHistoryNextArm_all","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryAction_follows_realHistoryNextArm_ae","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryExpectedRegret_eq_armStreamExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmModelCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"declaration:BanditRLProof.UCB.realStationaryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.condDistrib_comp_measurePreserving","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.measurePreservingArmStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.measurePreservingArmStreamReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_comp_measurePreserving_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.productNoiseArmStreamAction","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.productNoiseArmStreamReward","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.realStationaryUCBSequence_productNoise_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.finiteArmProductNoiseMeasure","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.productNoiseArmwiseBoundedFiniteArmExpectedRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"declaration:BanditRLProof.UCB.productNoiseArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectory_historyAction_map_eq_armStream","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryReward_zero_condDistrib_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryReward_succ_condDistrib_ae_eq_nu","relation":"contains"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","target":"declaration:BanditRLProof.UCB.canonicalKernelTrajectoryArmwiseBoundedFiniteArmExpectedAverageRegret_tendsto_zero_and_explicitPolicy_and_selectedRewardLaws","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.HarnessProfile","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.AgentRole","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.TaskKind","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.TaskStatus","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.ArtifactSpec","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.AcceptanceGate","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.HarnessTask","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.defaultLeanGate","relation":"contains"},{"source":"module:BanditRLProof.Automation","target":"declaration:BanditRLProof.defaultHarnessProfile","relation":"contains"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.intervalVarianceProxy_pos_of_lt","relation":"contains"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"declaration:BanditRLProof.RewardKernel.centeredRewardKernelLaw_of_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"declaration:BanditRLProof.RewardKernel.boundedCenteredRewardKernelLaw","relation":"contains"},{"source":"module:BanditRLProof.BudgetStoppingTime","target":"declaration:BanditRLProof.Budget.budgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.BudgetStoppingTime","target":"declaration:BanditRLProof.Budget.isStoppingTime_budgetExhaustionTime_of_adapted","relation":"contains"},{"source":"module:BanditRLProof.BudgetStoppingTime","target":"declaration:BanditRLProof.Budget.measurableSet_budgetExhaustionTime_le_of_adapted","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationCappedOccupancy","target":"declaration:BanditRLProof.Concentration.cappedOccupancyTail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationCappedOccupancy","target":"declaration:BanditRLProof.Concentration.cappedOccupancyTail_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationCappedOccupancy","target":"declaration:BanditRLProof.Concentration.cappedOccupancyTail_antitone","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationCappedOccupancy","target":"declaration:BanditRLProof.Concentration.integral_cappedOccupancyTail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationCappedOccupancy","target":"declaration:BanditRLProof.Concentration.sum_le_occupancy_bound_sharp","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConditionalMGF","target":"declaration:BanditRLProof.Concentration.hasCondMGFUpperBoundAt_of_condExp_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.geometricConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.geometricConfidenceShare_pos","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.tsum_ofReal_geometricConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.telescopingConfidenceWeight","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.telescopingConfidenceWeight_eq_sub","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.sum_range_telescopingConfidenceWeight","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.telescopingConfidenceWeight_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.hasSum_telescopingConfidenceWeight","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.telescopingConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.telescopingConfidenceShare_eq_div","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.telescopingConfidenceShare_pos","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"declaration:BanditRLProof.Concentration.tsum_ofReal_telescopingConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"declaration:BanditRLProof.Concentration.mul_exp_neg_le_exp_difference","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"declaration:BanditRLProof.Concentration.sum_dyadic_mul_exp_neg_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"declaration:BanditRLProof.Concentration.sum_dyadic_mul_exp_neg_le_three_div_two","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"declaration:BanditRLProof.Concentration.tsum_dyadic_mul_exp_neg_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"declaration:BanditRLProof.Concentration.sum_moss_peeling_exponential_le_twelve","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"declaration:BanditRLProof.Concentration.tsum_moss_peeling_exponential_le_fifteen","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","target":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_tsum_of_uniform","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","target":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","target":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.ae_integrable_exp_mul","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.ae_forall_integrable_exp_mul","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.congr","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.congr_iff","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.memLp_exp_mul","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.zero_kernel","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.zero_measure","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.prodMkLeft_compProd","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.integrable_exp_add_compProd","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.add_compProd","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.Kernel.HasMGFUpperBoundAt.add_comp","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.hasMGFUpperBoundAt_iff_kernel","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.of_map","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.id_map_iff","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.compensated","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.trim","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.add_of_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.measure_ge_le_exp_add","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.HasMGFUpperBoundAt.sum_of_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.measure_sum_ge_le_of_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"declaration:BanditRLProof.Concentration.measure_sum_ge_inter_sum_le_of_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.integral_mul_exp_neg_mul_sq_Ioi","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.integral_transformed_occupancy_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.occupancyTail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.occupancyTail_antitoneOn","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.occupancySubstitution","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.occupancySubstitution_image","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.occupancySubstitution_injOn","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.hasDerivAt_occupancySubstitution","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.integral_occupancyTail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.integrableOn_occupancyTail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.sum_occupancyTail_shift_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"declaration:BanditRLProof.Concentration.sum_le_occupancy_bound","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.fixedRadiusMeanEvent","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.measure_fixedRadiusMeanEvent_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.sum_measureReal_fixedRadiusMeanEvent_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.measurableSet_fixedRadiusMeanEvent","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.fixedRadiusCount","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.integrable_fixedRadiusCount","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.integral_fixedRadiusCount_le","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"declaration:BanditRLProof.Concentration.integral_fixedRadiusCount_le_sharp","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"declaration:BanditRLProof.Concentration.submartingale_exp_of_martingale","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"declaration:BanditRLProof.Concentration.measure_exists_le_martingale_ge_le_exp","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"declaration:BanditRLProof.Concentration.measure_exists_le_martingale_ge_le_subgaussian","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"declaration:BanditRLProof.Concentration.measure_exists_le_independent_partialSum_ge_le_subgaussian","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"declaration:BanditRLProof.Concentration.exists_tilt_quadratic_fixedMGF_exponent_le_neg","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"declaration:BanditRLProof.Concentration.quadraticFixedMGFRadius","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"declaration:BanditRLProof.Concentration.measure_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticMaximal","target":"declaration:BanditRLProof.Concentration.quadraticFixedMGFMaximalRadius","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticMaximal","target":"declaration:BanditRLProof.Concentration.measure_biUnion_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticScheduled","target":"declaration:BanditRLProof.Concentration.quadraticFixedMGFScheduledRadius","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticScheduled","target":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_tsum_of_fixedTilt_quadratic_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationQuadraticScheduled","target":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.of_measurableSpace_eq","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.integrable","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.indicator","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.indicator_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.intervalVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.boundedCentered_hasSubgaussianMGF_of_mem_Icc_integral_eq","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.integral_abs_le_two_mul_sqrt_mul_exp_half_of_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.integral_sq_le_four_mul_proxy_mul_exp_half_of_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussian_sum_tail_of_iIndepFun","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussian_sum_tail_ennreal_of_iIndepFun","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_tail_of_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_tail_ennreal_of_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussian_sum_abs_tail_ennreal_of_iIndepFun","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_abs_tail_ennreal_of_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussianSumConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussianSumConfidenceRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussianSumConfidenceRadius_sq","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.two_mul_exp_neg_subGaussianSumConfidenceRadius_sq_div_eq_delta","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussian_sum_abs_tail_ennreal_delta_of_iIndepFun","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_sum_abs_tail_ennreal_delta_of_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussianAverageConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussianAverageConfidenceRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.measure_average_abs_tail_le_of_measure_sum_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.measure_randomCount_average_abs_tail_le_of_measure_sum_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.measure_positive_randomCount_event_le_sum_exactCount","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.measure_positive_randomCount_event_le_of_exactCount_uniform","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussian_average_abs_tail_ennreal_delta_of_iIndepFun","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_average_abs_tail_ennreal_delta_of_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_indicator_sum_tail_predictableVariance_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.subGaussianPredictableVarianceRadius","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"declaration:BanditRLProof.Concentration.condSubGaussian_indicator_sum_abs_tail_predictableVariance_delta","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationTailIntegration","target":"declaration:BanditRLProof.Concentration.integral_positive_tail_le_two_sqrt","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationVariance","target":"declaration:BanditRLProof.Concentration.variance_chebyshev_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationVariance","target":"declaration:BanditRLProof.Concentration.evariance_chebyshev_tail","relation":"contains"},{"source":"module:BanditRLProof.ConcentrationVariance","target":"declaration:BanditRLProof.Concentration.variance_sum_of_pairwise_indep","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_integral_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.hasSubgaussianMGF_mono_varianceProxy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_deterministic_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_dirac_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_countable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_countable_trim","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_map_eq_of_condDistrib_ae_eq_real_trim","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.eventuallyEq_const_of_map_eq_dirac","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.pair_map_eq_map_prod_mk_of_action_ae_eq_const_reward_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_pair_map_eq_map_prod_mk_of_action_ae_reward_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_trim","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_identDistrib_trajMeasure_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_of_identDistrib_trajMeasure_trim","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_action_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_actionMarginal_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedAction_condExpKernel_ae_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_reward_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_selectedMeasure_rewardHistoryOfTrace_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_condExpKernel_map_trajMeasure_of_selectedAction_ae_selectedMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_of_condExpKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_centered_of_condExpKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_congr_measurableSpace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_condExpKernel_integral_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_integral_eq_historyStepKernel_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_integral_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExp_eq_zero_of_condExpKernel_map_eq_historyStepKernel_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.hasCondSubgaussianMGF_of_condExpKernel_map_eq_historyStepKernel_centeredReward","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_event_real_eq_indicator_of_measurableSet","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.condExpKernel_ae_eq_const_of_countable_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_prefix_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_finitePairHistoryOfTrace_condExpKernel_map_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_of_measurable_of_policy_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_of_pairHistory_measurable_of_action_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_action_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.action_condExpKernel_ae_eq_policy_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_policy_of_action_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.random_pair_condExpKernel_map_eq_actual_action_of_generatedActionTraceSucc_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.pair_condExpKernel_map_eq_frozen_actual_action_of_generatedActionTraceSucc_random_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_of_coordinate_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finiteRewardHistory_condExpKernel_frozen_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_of_coordinate_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_condExpKernel_frozen_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_ae_eq_extend_of_pairHistory_frozen","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_condExpKernel_ae_eq_extend_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.finitePairHistory_succ_condExpKernel_map_eq_extend_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_frozenPast_ae_of_history_frozen","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_of_action_ae_eq_policy_reward_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_of_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_actual_action_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_pair_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardFoundation","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure_of_condSubgaussian","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardFoundation","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_conditionalRewardFoundation_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.eventually_ae_trim_of_eq_measurableSpace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_measurable","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.comap_finitePairHistoryOfTrace_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyFiltrationSucc_generatedActionFromRewardHistory_eq_comap_finiteRewardHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_succ_measurable_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_selectedMeasure_condExpKernel_map_trajMeasure_generatedActionFromRewardHistory_finitePairHistoryOfTrace_trim","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionPartialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_partialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_action_ae_eq_policy_reward_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_finitePairHistory_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionSelectedRewardFinitePairHistoryLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_comap_trim_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_identDistrib_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionSelectedRewardFinitePairHistoryLawSource_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_selectedRewardFinitePairHistoryLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_condExp_eq_zero_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_centeredRewardSuccProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionFromRewardHistory_armMaskedCenteredRewardSuccProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredRewardSuccProcess_average_tail_ennreal_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_comap_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_comap_trim_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_selectedRewardFinitePairHistoryLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_comap_trim_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_definitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_definitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_definitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_partialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_generatedActionActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.reward_condExpKernel_map_eq_selected_policy_of_generatedActionRandomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_definitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.GeneratedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_comap_trim_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_partialTrajectoryPairLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_comap_trim_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardPartialTrajectoryKernel_extend_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_actionRewardHistoryStepKernelFamily_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.rawReward_succ_aemeasurable_of_measurable_reward","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.selectedMean_succ_aemeasurable_of_measurable_mean","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.selectedMean_succ_bound_of_mean_range_bound","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.rawReward_succ_bound_of_reward_range_bound","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_historyStepKernelFamily_condExpKernel_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_of_coordinate_measurable_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_projected_of_context_state_measurable_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_action_ae_eq_policy_reward_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_selected_policy_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawMeanBoundedSource_of_rawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeanBoundedSource_of_rawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeasurableMeanBoundedSource_of_rawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource_of_rawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_rawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.integrable_exp_mul_of_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_exp_of_generatedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_boundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_randomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionSelectedRewardFinitePairHistoryLawSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_reward_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_actionRewardPartialTrajectoryKernel_extend_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionTraceSucc_random_pair_map_eq_actual_action_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionActualRewardMapSource_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalMapSource_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairMapSource_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalMapSource_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairBoundedCenteredSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawBoundMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_aemeasurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_measurable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_bound_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_integrable_exp_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_definitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairMapSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionPartialTrajectoryPairLawSource_of_historyVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionDefinitionalActualRewardMapSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionActualRewardMapSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairBoundedCenteredSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairCenteredSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalCenteredSource_of_uniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_pair_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_centered_meas","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_variance_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_uniform_variance_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_history_variance_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_historyVarianceSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_partialTrajectoryPairLawSource_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_selectedRewardFinitePairHistoryLawSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_via_selectedRewardFinitePairHistoryLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_via_selectedRewardFinitePairHistoryLawSource_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_comap_trim_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardPartialTrajectoryKernel_extend_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_actionRewardHistoryStepKernelFamily_pair_map_eq_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_condExp_eq_zero_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_generatedActionDefinitionalActualRewardMapSource","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_generatedActionDefinitionalActualRewardMapSource_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeBoundedSource_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalMapSource_rawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeUniformVarianceBoundedSource_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeUniformVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_actual_action","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_actual_action_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.generatedActionRandomPairDefinitionalRawRangeMeasurableMeanRangeHistoryVarianceBoundedSource_of_reward_map_eq_selected_policy","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded_of_varianceCeiling_le","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_sum_tail_ennreal_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmPullCount","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmPullCount_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmRewardSum","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_eq_successorArmRewardSum_sub_pullCount_mul","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedVarianceSuccProcess_sum_eq_mul_successorArmPullCount","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.armMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanExactCountRadius","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanPeelingRadius","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeConfidenceShare","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimePeelingRadius","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFiniteArmTimeBadEvent","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMean_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_average_abs_tail_ennreal_delta_of_reward_map_eq_selected_policy_definitionalRawRangeMeasurableMeanRangeHistoryVarianceBounded","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeGeometricAllTimeBadEvent","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardPartialTrajectoryKernel_map_eq_historyFiltrationSucc_finitePairHistoryOfTrace_of_pair_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredRewardSuccProcess_stronglyAdapted_historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib_of_ae_variance","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.centeredReward_succ_hasCondSubgaussianMGF_of_pair_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_of_pair_condDistrib_on_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_centeredRewardSuccProcess_sum_tail_ennreal_trajMeasure_on_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_action_succ_ae_eq_policy_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_policyArmMaskedCenteredRewardSuccProcess_sum_abs_tail_predictableVariance_ennreal_delta_trajMeasure_on_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_sampledArmMaskedCenteredRewardSuccProcess_sum_abs_tail_successorPullCount_ennreal_delta_trajMeasure_on_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_exact_pullCount_ennreal_delta_trajMeasure_on_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","target":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent","relation":"contains"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.ActionTrace","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.RewardTrace","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.pullCount","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.pullCount_zero","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.pullCount_succ","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.sumRewards","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.sumRewards_zero","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.sumRewards_succ","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.FiniteBanditModel","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.FiniteBanditModel.bestArm","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.FiniteBanditModel.bestMean","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.FiniteBanditModel.gap","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.FiniteBanditModel.gap_bestArm","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.PolicySketch","relation":"contains"},{"source":"module:BanditRLProof.Core","target":"declaration:BanditRLProof.CertificateStatus","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.IsSimplexTangent","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.tangentPairing","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.tangentPairing_add_const_of_isSimplexTangent","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedCenter","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.sum_weight_mul_sub_weightedCenter_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition_of_centered","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_center_le","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_eq_center_iff","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_decomposition","relation":"contains"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_le_two","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.observedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.outstandingAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.observedBefore_disjoint_outstandingAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.observedBefore_union_outstandingAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.card_observedBefore_add_card_outstandingAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.outstandingCount","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.maxOutstandingBeforeThrough","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.outstandingCount_le_round","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.outstandingCount_le_maxOutstandingBeforeThrough","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.oneBasedDelayShift","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperMissingAtEnd","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperMissingAtEnd_eq_outstandingAt_oneBasedDelayShift","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_eq_outstandingCount_oneBasedDelayShift","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_le_round","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperSigmaMaxThrough","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_le_paperSigmaMaxThrough","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability_nonnegative","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.sum_probability_eq_one","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.finiteActionDistribution","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure_isProbabilityMeasure","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_isProbabilityMeasure","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_eq_of_observation_equivalent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.inactiveArms","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.activeEqualShare","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_active","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_inactive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.activeEqualShare_nonneg","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_nonneg","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"declaration:BanditRLProof.DelayedFeedback.sum_delayedSAPOProbability_eq_one","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.ActionTimeView","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.ActionTimeView.ext","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.CausalDecisionRule","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_lt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_not_lt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_mem","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_outstanding_loss_hidden","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_eq_of_observation_equivalent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"declaration:BanditRLProof.DelayedFeedback.causalDecision_eq_of_observation_equivalent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOInitialEliminatedProbability","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOInitialPhaseTarget","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_pos","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmBank","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.ActiveArmsUninitialized","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_some_iff","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_none_iff","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_mem","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.mem_eliminated_of_initializeNewlyEliminated_ne_prior","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_eq_prior_of_mem_remainingActive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.prior_eq_none_of_mem_eliminated","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.remainingActive_uninitialized_after_initialize","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_arm","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationRound","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationProcessedOrder","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_errorCount","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseIndex","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseSamples","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_processedAtProbabilityLevel","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_pos","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_le_one","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_pos","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_pos","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_spec_of_mem","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.eliminated","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_eliminated_iff","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_remainingActive_iff","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.OptimalArmSurvivalCertificate","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.optimal_mem_remainingActive_of_certificate","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive_nonempty_of_certificate","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.sum_delayedSAPOProbability_after_elimination_eq_one","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","target":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","target":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","target":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","target":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim_iff_shared_fields","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","target":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim_iff_shared_fields","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.processedOrder_toFinset_eq_observedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActionRound","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_processedOrder","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralStep","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralReachable","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralStep","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralReachable","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.mem_earlierRemainingActive_of_laterEliminated","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.source_le_roundStart","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.currentActive_subset_activeAtSourceRound","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.sourceRound_not_mem","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_nodup","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_available","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedSource_le_roundStart","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.currentActive_subset_extendedSourceActive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.toPreEliminationSummary","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.line8RemainingActive_subset_currentActive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_currentActive_subset_before","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_preserves_roundStart","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.processedPullCount","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass_eq_of_active_throughout","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_nonneg","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_one","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_three_of_count_le_eight_mul","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.expectedPullMass_eq_of_mem_active","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.quarter_count_sub_six_log_le_count_of_mem_active","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.eighth_count_le_count_of_large_count","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_three_of_large_count","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_ten_of_mem_active","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.ucbStar_le_empiricalMean_add_width","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.activeArmGapBranch","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countCertificate","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.newlyObservedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.observedBefore_mono","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.processed_disjoint_newlyObservedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.processed_union_newlyObservedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.processAllNew","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.processAllNew_eq_observedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.previousObservedBefore_subset_current","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.processAllNew_from_previous_eq_current","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"declaration:BanditRLProof.DelayedFeedback.outstandingAt_disjoint_newlyObservedBefore","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefix","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalWidthAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalUpperAt","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toConfidenceSnapshot","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.D4CountClause","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.currentActive_subset_activeAt_sourceIndex","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefixCountCertificate","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_traceSummary","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"declaration:BanditRLProof.DelayedFeedback.finiteAverageGap","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"declaration:BanditRLProof.DelayedFeedback.aboveTwiceAverageGap","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"declaration:BanditRLProof.DelayedFeedback.two_mul_card_aboveTwiceAverageGap_le","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"declaration:BanditRLProof.DelayedFeedback.sourceStochasticLossGap","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"declaration:BanditRLProof.DelayedFeedback.sourceStochasticLossGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"declaration:BanditRLProof.DelayedFeedback.two_mul_card_sourceStochasticLossGap_aboveTwiceAverage_le","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_antitone","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_of_count_le_96_mul_scale","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_two_log_of_small_count","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_one","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_four","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_one_le_four","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_horizon_four_one_le_four","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.eight_mul_empiricalWidth_lt_gap_of_mem_eliminated","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive_of_large_or_small_count","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap_le_gap","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_two_mul_surrogateGap","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.d12_gap_ordering_chain","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_twenty_mul_gap_of_eliminationPrefixIndex_le","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.sourceUcbStar","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.EliminationGoodEvent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalMean_le_ucbStar_of_eliminationGoodEvent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalArmSurvivalCertificate_of_eliminationGoodEvent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimal_mem_remainingActive_of_eliminationGoodEvent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalSurvivalEventSet","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet_subset_optimalSurvivalEventSet","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le_of_goodEvent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventComponent","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.componentFailure","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.failureSet","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet_compl","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_sum","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.linearFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.doubleLinearFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceComponentFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sum_sourceComponentFailureBudget_eq_nine_div","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_nine_div","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_eliminationGoodEventSet_compl_le_nine_div","relation":"contains"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_optimalSurvivalEventSet_compl_le_nine_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.MeasurableFiniteActionDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.finiteActionKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcessMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcessHistory","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcessAction","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcessHistory_measurable","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcessAction_measurable","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcess_history_map_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcess_condDistrib_action_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.finiteActionKernel_ae_eq_finiteActionMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcess_integral_importanceWeightedLoss_eq_integral_loss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.predictableLossAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_realizedLoss_le_one_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_finiteHorizon_realizedLoss_le_one_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedRegret_le_horizon_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_trivialRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.bernsteinLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.bernsteinAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinArmEntropyExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinConfidenceExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinRealizedExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.rpow_inv_three_le_half_of_eight_mul_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.sqrt_div_le_half_of_four_mul_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.numerator_le_cube_mul_of_rpow_inv_three_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.numerator_le_sq_mul_of_sqrt_div_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitBernsteinRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinHighProbabilityRegret_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinRealizedHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinConfidenceRadius_le_three_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.realizedDeviationRadius_le_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_sq_mul","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinHighProbabilityLearningRate_le_sq_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinEntropyBudget_le_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinUnscaledSquareBudget_le_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.bernsteinHedgeBudget_le_three_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.log_one_div_third_eq_log_three_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinRealizedHighProbabilityRegretBudget_le_eleven_mul","relation":"contains"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedBernsteinRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3BestArm","target":"declaration:BanditRLProof.Exp3.sampledPredictableBestArmCumulativeLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3BestArm","target":"declaration:BanditRLProof.Exp3.threshold_le_sampledPredictableRealizedLoss_sub_bestArmCumulativeLoss_iff","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Concentration.exp_le_one_add_self_add_sq_of_abs_le_one","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Concentration.exists_tilt_fixedMGF_exponent_le_neg","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_comparatorEstimatorDeviation_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_comparatorEstimatorDeviation_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.finiteActionComparatorEstimator_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.comparatorEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorBernsteinConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_bernstein_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.comparatorEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableComparatorEstimatorDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryObservedComparatorEstimatorDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorDeviationProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviationProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorDeviationProxy_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledComparatorEstimatorVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedComparatorEstimatorDeviation_sum_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.FiniteActionDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.finiteActionMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.finiteActionMeasure_isProbabilityMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.integral_finiteActionMeasure_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.integral_historyAction_eq_integral_sum_of_condDistrib_ae_eq_finiteActionMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.integral_importanceWeightedLoss_eq_integral_loss_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.integral_weightedImportanceWeightedLoss_eq_integral_weightedLoss_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"declaration:BanditRLProof.Exp3.integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.eventually_const_le_natCast","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.eventually_const_le_natCast_pow_three","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.eventually_doubleVarianceProbabilisticSparseLossLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold_eq_explicit_of_largeHorizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure_of_largeHorizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_largeHorizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le_of_largeHorizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"declaration:BanditRLProof.Exp3.eventually_sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.expectedRegretBudget_le_four_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedExplorationRate_sq_mul_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedExplorationRate_mul_eq_sqrt_mul","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.tunedPredictableTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExplorationBias","target":"declaration:BanditRLProof.Exp3.distribution_le_sampledTrajectoryProbabilityAt_div_one_sub_gamma","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExplorationBias","target":"declaration:BanditRLProof.Exp3.mixedSquaredLoss_sampledTrajectoryObservedLoss_le_inv_one_sub_gamma","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExplorationBias","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedLoss_le_pure_add_gamma","relation":"contains"},{"source":"module:BanditRLProof.Exp3ExplorationBias","target":"declaration:BanditRLProof.Exp3.sampledTrajectory_finiteHorizon_explorationBias_secondMoment","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.cumulativeLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.weight","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.totalWeight","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.distribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.mixedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.mixedSquaredLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.cumulativeLoss_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.weight_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.weight_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.totalWeight_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.distribution_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.distribution_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.sum_distribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.totalWeight_zero","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.exp_neg_le_one_sub_add_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.exp_neg_mul_le_one_sub_add_sq_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.exp_neg_mul_le_one_sub_add_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.totalWeight_succ_div_eq_sum_distribution_mul_exp","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.totalWeight_succ_div_le_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.totalWeight_succ_div_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.log_totalWeight_succ_sub_le_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.log_totalWeight_succ_sub_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.mixedSquaredLoss_le_one","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt_le_inv_explorationFloor","relation":"contains"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_mem_unitInterval_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_observedMixedSquared_sum_le_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.importanceWeightedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.mixedImportanceWeightedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.weightedImportanceWeightedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.importanceWeightedLoss_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_importanceWeightedLoss_eq_loss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.mixedImportanceWeightedLoss_eq_selectedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_mixedImportanceWeightedLoss_eq_mixedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_weightedImportanceWeightedLoss_eq_weightedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss_eq_selectedLoss_sq_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_mixedSquaredImportanceWeightedLoss_eq_sum_loss_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_mixedSquaredImportanceWeightedLoss_le_card","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredImportanceWeightedLoss_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_card_div_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimator_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_bernstein_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredBernsteinVarianceCoefficient","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredBernsteinConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_bernstein_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinSquareHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.bernsteinSquareLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.bernsteinSquareAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.bernsteinSquareBestArmAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredBernsteinVarianceCoefficient_eq_card_sq_div_gamma","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_mul_sampledMixedSquaredBernsteinConfidenceRadius_le_three_mul_sq_mul_horizon_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareRealizedExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareRealizedTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedBernsteinSquareRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitBernsteinSquareRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_bernsteinSquareRealizedHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityScale_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityLearningRate_sq_mul_scale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.bernsteinSquareRealizedTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableBernsteinSquareRealizedHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedBernsteinSquareRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.mixedSquaredImportanceWeightedLoss_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredDeviationProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredDeviationProxy_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredAt_le_card","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredSum_le_card_mul","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquared_sum_tail_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledObservedMixedSquaredSum_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_exponential","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonExponentialSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredVarianceProxy_coe_eq_card_div_two_gamma_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_mul_sampledMixedSquaredConfidenceRadius_le_sq_mul_horizon_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBalancedSqrt_le_two_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRealizedExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedExponentialSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareArmExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareMixedExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.rpow_inv_six_le_half_of_sixtyfour_mul_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.numerator_le_pow_six_mul_of_rpow_inv_six_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitExponentialSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_exponentialSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityScale_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityLearningRate_sq_mul_scale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.exponentialSquareBernsteinRealizedTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableExponentialSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedExponentialSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.measurable_mixedSquaredEstimatorCenteredSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_le_card_div_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.integral_sq_mixedSquaredEstimatorDeviation_finiteActionMeasure_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorDeviation_condExpKernel_map_eq_finiteActionMeasure_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.integral_sq_mixedSquaredEstimatorDeviation_condExpKernel_eq_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_zero_condExpKernel_integral_sq_eq_variance","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_succ_condExpKernel_integral_sq_eq_variance","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviationProcess_condExpKernel_integral_sq_eq_varianceProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableMixedSquaredVarianceAt_filtration","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess_isPredictable","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVariance_sum_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_joint_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_mixedSquaredEstimatorDeviation_le_inv_floor_mul_sum_loss_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimatorCenteredSecondMoment_le_inv_floor_mul_sum_loss_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableMixedSquaredVarianceAt_le_inv_floor_mul_lossSquaredAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossSquaredSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossSquaredSum_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareLossEnergyRealizedMarkovHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableDoubleVarianceRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_doublePredictableVarianceRealizedHighProbabilityRegret_tail_joint","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_joint_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.measurable_sampledPredictableMixedSquaredVarianceSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableMixedSquaredVarianceSum_gt_le_lintegral_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedHighProbabilityRegret_tail_of_lintegral_variance_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareRealizedMarkovHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareRealizedMarkovHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossDoubleVarianceRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_smallLossDoublePredictableVarianceRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.predictableLossAt_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredAt_le_lossMassAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossSquaredSum_le_lossMassSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_inv_floor_mul_lossMassSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_of_lossMassSum_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_off_bad_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_predictableVariance_of_lossMassSum_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossHighProbabilityRegret_tail_joint","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_off_bad_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint_of_lossMassSum_le_or_mem","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedHighProbabilityRegret_tail_joint","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSmallLossRealizedMarkovHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableSparseRealizedVarianceBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceSum_le_sparseRealizedVarianceBudget_or_mem_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_doubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRealizedExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.boundedRealizedLargeHorizon_of_doubleVarianceLargeHorizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.doubleVarianceProbabilisticSparseLossClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseRealizedPredictableVarianceRadius_le_three_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossDoubleVarianceRealizedTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableDoubleVarianceProbabilisticSparseLossRealizedHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedDoubleVarianceProbabilisticSparseLossRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_ae_sparsity","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_one_div_fifth_eq_log_five_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_mul_sparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceArmExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceMarkovExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceConfidenceExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.rpow_inv_five_le_half_of_thirtytwo_mul_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.numerator_le_pow_five_mul_of_rpow_inv_five_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossSupport","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassAt_le_supportCard","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_sparsity_mul_horizon_of_sample","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_sparsity_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta_of_ae_sparsity","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableSparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_or_mem_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableLossMassSum_le_card_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableGlobalVarianceMeanBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceLIntegral_le_globalLossMass","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_mul_probabilisticSparseLossPredictableVarianceRadius_le_three_mul_sq_mul_horizon_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBalancedSqrt_le_two_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceMarkovExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceBudget_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityScale_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.probabilisticSparseLossPredictableVarianceRealizedMarkovTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPredictableVarianceRealizedMarkovRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceBudget_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityScale_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityLearningRate_sq_mul_scale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceHighProbabilityHedgeBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sparseLossPredictableVarianceRealizedMarkovTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareSparseLossRealizedMarkovHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedSparseLossPredictableVarianceRealizedMarkovRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableSparsePathwiseVarianceBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredVarianceSum_le_sparsePathwiseVarianceBudget_or_mem_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"declaration:BanditRLProof.Exp3.sampledPredictable_predictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossBestArmAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_single_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonProbabilisticSparseLossPathwiseVarianceBestArmRealizedRegret_tail_of_sparsityFailure_le_single_charge","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableSparsePathwiseVarianceBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_mul_pathwiseVarianceProbabilisticSparseLossRadius_le_three_mul_sq_mul_horizon_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossBalancedSqrt_le_two_mul_gamma_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossMixedExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityScale_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityLearningRate_sq_mul_scale","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossHighProbabilityHedgeBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.pathwiseVarianceProbabilisticSparseLossRealizedTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableVarianceSquareProbabilisticSparseLossRealizedPathwiseVarianceHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_off_sparsityFailure","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedProbabilisticSparseLossPathwiseVarianceRealizedRegret_tail_of_sparsityFailure_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.abs_mixedSquaredEstimatorDeviation_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_hasMGFUpperBoundAt_variance","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.finiteActionMixedSquaredEstimator_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.mixedSquaredEstimator_compensated_hasCondMGFUpperBoundAt_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensated_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensated_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredCompensatedProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledMixedSquaredPredictableVarianceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableMixedSquaredDeviation_sum_tail_predictableVariance_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.potential","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.updatedWeight","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.updatedPotential","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.updatedPotential_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.updatedWeight_nonneg_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.updatedPotential_nonneg_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.updatedPotential_sub_potential_eq_sum_weight_mul_exp_sub_one","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.sum_range_forward_difference","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.potentialProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3Potential","target":"declaration:BanditRLProof.Exp3Potential.potentialProcess_telescope_sum_range","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.PredictableLossVector","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.PredictableLossVector.environment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.PredictableLossVector.environment_initialFeedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.PredictableLossVector.environment_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.PredictableLossVector.initial_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.PredictableLossVector.successor_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.trajectoryMixture_condDistrib_action_given_environment_history","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryMeasure_condDistrib_action_given_environment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableHedge","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_nonneg_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableHedge","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_finiteHorizon_reward_nonneg_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableHedge","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_hedge_regret_le_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableHedge","target":"declaration:BanditRLProof.Exp3.sampledPredictableScoreHedge_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureProbabilitySourceAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryExploredPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureObservedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPurePredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableLossAt_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryPurePredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryExploredPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryExploredPredictableLossAt_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryExploredPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureObservedLossAt_ae_eq_weightedPredictable","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryPureObservedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictablePureObservedInitial_integral_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictablePureObservedSuccessor_integral_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictablePureObservedAt_integral_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictablePureObserved_finiteHorizon_integral_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_hedge_exploredSecondMoment_le_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictable_integral_pureHedge_le_exploredSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictable_integral_exploredLoss_le_pure_add_gamma","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictableObserved_finiteHorizon_secondMoment_integral_le_card_mul","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.trajectoryMixture_map_environment_history_output_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_prefix_next_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_nextPair_given_environment_prefix","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledEnvironmentHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableSuccessorLossRegularity","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledInitialEnvironmentDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableInitialLossRegularity","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_map_environment_action_zero","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalMeasurableEnvironmentTrajectoryMeasure_condDistrib_action_zero_given_environment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalPredictableTrajectoryMeasure_reward_zero_eq_initialLoss_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedInitial_first_second_moment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.canonicalPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryMeasure_reward_eq_successorLoss_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableSuccessorLoss_first_second_moment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedSuccessor_first_second_moment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.predictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.observedImportanceWeightedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilitySourceAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableTrajectoryLossRegularityAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.measurable_predictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.measurable_observedImportanceWeightedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.measurable_observedMixedSquaredImportanceWeightedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.integrable_predictableImportanceWeightedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.integrable_predictableMixedSquaredImportanceWeightedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.integrable_predictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.integrable_predictableLossSqSumAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.observedAt_eq_predictableAt_ae","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.integrable_observedAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedAt_first_second_moment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"declaration:BanditRLProof.Exp3.sampledPredictableObserved_finiteHorizon_first_second_moment","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableRegretAllTime","target":"declaration:BanditRLProof.Exp3.sampledPredictableRegretGeometricAllTimeBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableRegretAllTime","target":"declaration:BanditRLProof.Exp3.sampledPredictableRegretGeometricAllTimeFailureSet","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableRegretAllTime","target":"declaration:BanditRLProof.Exp3.mem_sampledPredictableRegretGeometricAllTimeFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.Exp3PredictableRegretAllTime","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.weightedImportanceWeightedLoss_eq_selected","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sum_prob_mul_sq_weightedEstimatorMeanMinusRaw_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.finiteActionWeightedEstimatorMeanMinusRaw_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.weightedEstimatorMeanMinusRaw_hasCondMGFUpperBoundAt_of_condDistrib_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableMinusWeightedAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusWeighted_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusWeighted_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedBernsteinConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_bernstein_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.weightedEstimator_hasCondSubgaussianMGF_of_condDistrib_ae_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryWeightedPurePredictableDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPureObservedDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledWeightedPurePredictableDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledWeightedPurePredictableDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviationConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPureObservedDeviation_sum_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPurePredictableMinusObservedAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObservedProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"declaration:BanditRLProof.Exp3.sampledPurePredictableMinusObserved_sum_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinLargeHorizonCondition","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinAllHorizonRegretThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonRandomSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.log_one_div_fourth_eq_log_four_div","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedExplicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedTunedThreshold_le_explicitThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_gammaCharacterizedRandomSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinConfidenceExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_eq_raw","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRawExplorationRate_le_half_of_horizon_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinClippedExplorationRate_contracts","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_explicitRandomSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinRealizedHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityLearningRate_sq_mul","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.randomSquareHighProbabilityHedgeBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.randomSquareBernsteinRealizedTunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinRealizedHighProbabilityRegretBudget_le_tunedThreshold","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"declaration:BanditRLProof.Exp3.sampledPredictable_tunedRandomSquareBernsteinRealizedRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledObservedMixedSquaredSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.observedMixedSquaredImportanceWeightedLossAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledObservedMixedSquaredSum_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.measurable_sampledObservedMixedSquaredSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.integrable_sampledPredictableObservedMixedSquaredSum","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableObservedMixedSquared_sum_tail_markov","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRandomSquareBernsteinHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_randomSquareBernsteinHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"declaration:BanditRLProof.Exp3.condExpKernel_map_eq_finiteActionMeasure_of_condDistrib_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"declaration:BanditRLProof.Exp3.sampledTrajectorySelectedDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectorySelectedDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"declaration:BanditRLProof.Exp3.intervalVarianceProxy_zero_one_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationConfidenceRadius_sq_domination","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_exp_neg_budget","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceLinearBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceLinearBudget_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedDeviationGeometricAllTimeRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedDeviationGeometricAllTimeFailureSet","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.mem_sampledRealizedDeviationGeometricAllTimeFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationAllTimeFailureSet_linearBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_zero_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableDeviationFiltration","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableDeviationFiltration_zero","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableDeviationFiltration_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProxy","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationProxy_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_ennreal","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedHighProbabilityRegretBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedHighProbabilityRegret_tail_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedHighProbabilityRegret_tail_total_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.measurable_selectedLossCenteredSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_one","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_lossMass","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.finiteActionSelectedLossDeviation_hasMGFUpperBoundAt_variance","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.finiteActionSelectedLossDeviation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableRealizedVarianceAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_le_lossMassAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryPredictableRealizedVarianceAt_le_one","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.selectedLossDeviation_compensated_hasCondMGFUpperBoundAt_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedLossCompensated_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableSelectedLossCompensated_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensated_zero_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensated_succ_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectoryPredictableRealizedVarianceAt_filtration","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess_isPredictable","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVarianceProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_lossMass","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceGeometricRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviationAllTimeFailureSet","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","target":"declaration:BanditRLProof.Exp3.mem_sampledPredictableRealizedDeviationAllTimeFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","target":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceMaximalRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_prefix_max_tail_predictableVariance_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedCompensatedProcess_sum_range_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_fixedTilt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledRealizedPredictableVarianceRadius","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_delta","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedLossAt_ae_eq_selectedPredictable","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.measurable_sampledTrajectorySelectedPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledTrajectorySelectedPredictableLossAt_mem_unitInterval","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectorySelectedPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.integrable_sampledTrajectoryRealizedLossAt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedInitial_integral_eq_explored","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedSuccessor_integral_eq_explored","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedAt_integral_eq_explored","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictableRealized_finiteHorizon_integral_eq_explored","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le_four_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeBudget","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeFailureSet","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"declaration:BanditRLProof.Exp3.mem_sampledRealizedRegretGeometricAllTimeFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"declaration:BanditRLProof.Exp3.sampledRealizedRegretGeometricAllTimeFailureSet_subset","relation":"contains"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.historyWeight","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.historyTotalWeight","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.normalizedHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.historyTotalWeight_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.normalizedHistoryDistribution_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.sum_normalizedHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredHistoryDistribution_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.sum_exploredHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredHistoryDistribution_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.explorationFloor_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.finiteActionDistribution_exploredHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.measurable_historyWeight","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.measurable_historyTotalWeight","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.measurable_normalizedHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.measurable_exploredHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.MeasurableFiniteHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.normalizedHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.initialExploredDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.finiteActionDistribution_initialExploredDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredHistoryAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"declaration:BanditRLProof.Exp3.exploredTrajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryObservedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.previousPairHistory_frestrictLe","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledHistoryScore_frestrictLe_eq_cumulativeLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.distribution_sampledTrajectoryObservedLoss_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilityAt_eq_mix_distribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryProbabilityAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledTrajectoryObservedLoss_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledTrajectory_hedge_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"declaration:BanditRLProof.Exp3.sampledHistoryScore_hedge_regret_le","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.previousPairHistory","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.measurable_previousPairHistory","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.measurable_observedImportanceWeightedLoss","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledHistoryScore_zero","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledHistoryScore_succ","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.measurable_sampledHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.measurableFiniteHistoryScore_sampledHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledHistoryDistribution_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedHistoryAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"declaration:BanditRLProof.Exp3.sampledImportanceWeightedTrajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.BoundedMeasurableLossWithProbabilityFloor","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.BoundedMeasurableLossWithProbabilityFloor.prob_pos","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.measurable_importanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.measurable_mixedImportanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.measurable_weightedImportanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.measurable_mixedSquaredImportanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.norm_importanceWeightedLoss_score_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.norm_mixedImportanceWeightedLoss_score_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.norm_weightedImportanceWeightedLoss_score_le_inv_floor","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.norm_mixedSquaredImportanceWeightedLoss_score_le_inv_floor_sq","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_importanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_mixedImportanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_weightedImportanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_mixedSquaredImportanceWeightedLoss_score","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_importanceWeightedLoss_selected_of_isFiniteMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_weightedImportanceWeightedLoss_selected_of_isFiniteMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.integrable_mixedSquaredImportanceWeightedLoss_selected_of_isFiniteMeasure","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.actionProcess_integral_importanceWeightedLoss_eq_integral_loss_of_regularity","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedImportanceWeightedLoss_eq_integral_mixedLoss_of_regularity","relation":"contains"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"declaration:BanditRLProof.Exp3.actionProcess_integral_mixedSquaredImportanceWeightedLoss_eq_integral_sum_loss_sq_of_regularity","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_exploredExpectedRegret_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_realizedExpectedRegret_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.clippedExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.clippedLearningRate","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.clippedExplorationRate_nonneg","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.clippedExplorationRate_le_half","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.clippedExplorationRate_eq_tuned","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.clippedPredictableTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"declaration:BanditRLProof.Exp3.sampledPredictable_clippedRealizedExpectedRegret_le_min","relation":"contains"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"declaration:BanditRLProof.ExpectationBochnerSums.integral_finset_sum","relation":"contains"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"declaration:BanditRLProof.ExpectationBochnerSums.integral_univ_sum","relation":"contains"},{"source":"module:BanditRLProof.ExpectationFiniteBanditBounds","target":"declaration:BanditRLProof.lintegral_univ_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","relation":"contains"},{"source":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","target":"declaration:BanditRLProof.lintegral_univ_sum_model_gap_ofReal_mul_natCast_pullCount_le_sum_model_gap_ofReal_mul_time","relation":"contains"},{"source":"module:BanditRLProof.ExpectationFoundation","target":"declaration:BanditRLProof.lintegral_actionTrace_eval_eq_indicator_one","relation":"contains"},{"source":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","target":"declaration:BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","target":"declaration:BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time_of_rat_gap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","target":"declaration:BanditRLProof.lintegral_ofReal_pseudoRegret_le_sum_model_gap_ofReal_mul_time","relation":"contains"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"declaration:BanditRLProof.ennreal_natCast_pullCount_eq_finset_range_indicator_one","relation":"contains"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"declaration:BanditRLProof.lintegral_natCast_pullCount_eq_sum_measure_actionTrace_eval_eq","relation":"contains"},{"source":"module:BanditRLProof.ExpectationPullCountBounds","target":"declaration:BanditRLProof.lintegral_natCast_pullCount_le_time","relation":"contains"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"declaration:BanditRLProof.real_pseudoRegret_eq_univ_sum_gap_mul_natCast_pullCount","relation":"contains"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"declaration:BanditRLProof.integrable_real_pullCount_of_measurable_action","relation":"contains"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"declaration:BanditRLProof.integrable_real_pseudoRegret_of_integrable_pullCount","relation":"contains"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"declaration:BanditRLProof.integral_real_pseudoRegret_eq_sum_gap_mul_integral_pullCount","relation":"contains"},{"source":"module:BanditRLProof.ExpectationSums","target":"declaration:BanditRLProof.lintegral_finset_sum_actionTrace_eval_eq_indicator_one","relation":"contains"},{"source":"module:BanditRLProof.ExpectationWeightedPullCount","target":"declaration:BanditRLProof.lintegral_finset_sum_gap_mul_natCast_pullCount_eq","relation":"contains"},{"source":"module:BanditRLProof.ExpectationWeightedPullCountBounds","target":"declaration:BanditRLProof.lintegral_finset_sum_gap_mul_natCast_pullCount_le_sum_gap_mul_time","relation":"contains"},{"source":"module:BanditRLProof.FTRLOneStep","target":"declaration:BanditRLProof.FTRL.linearLoss","relation":"contains"},{"source":"module:BanditRLProof.FTRLOneStep","target":"declaration:BanditRLProof.FTRL.finiteSimplex","relation":"contains"},{"source":"module:BanditRLProof.FTRLOneStep","target":"declaration:BanditRLProof.FTRL.regularizedObjective","relation":"contains"},{"source":"module:BanditRLProof.FTRLOneStep","target":"declaration:BanditRLProof.FTRL.IsRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.FTRLOneStep","target":"declaration:BanditRLProof.FTRL.linearLoss_sub_le_regularizer_sub_div_of_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.FTRLOneStep","target":"declaration:BanditRLProof.FTRL.linearLoss_sub_le_regularizer_sub_div_of_simplex_minimizer","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.finiteArmIntervalVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.intervalVarianceProxy_le_finiteArmIntervalVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.finiteArmIntervalVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.finiteArmVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteArmVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.finiteArmVarianceProxy_pos_of_exists","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.finiteArmPositiveVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteArmPositiveVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.Concentration.finiteArmPositiveVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.RewardKernel.contextIndependentCenteredRewardKernelLaw_of_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.RewardKernel.contextIndependentBoundedCenteredRewardKernelLaw","relation":"contains"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"declaration:BanditRLProof.RewardKernel.contextIndependentArmwiseBoundedCenteredRewardKernelLaw","relation":"contains"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"declaration:BanditRLProof.FiniteBanditModel.mean_le_foldl_select","relation":"contains"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"declaration:BanditRLProof.FiniteBanditModel.mean_le_bestArm_mean","relation":"contains"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"declaration:BanditRLProof.FiniteBanditModel.gap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"declaration:BanditRLProof.FiniteBanditModel.maxGap","relation":"contains"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"declaration:BanditRLProof.FiniteBanditModel.gap_le_maxGap","relation":"contains"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"declaration:BanditRLProof.FiniteBanditModel.maxGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"declaration:BanditRLProof.Concentration.finiteContextArmVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteContextArmVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"declaration:BanditRLProof.Concentration.finiteContextArmVarianceProxy_pos_of_exists","relation":"contains"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"declaration:BanditRLProof.Concentration.finiteContextArmPositiveVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"declaration:BanditRLProof.Concentration.varianceProxy_le_finiteContextArmPositiveVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"declaration:BanditRLProof.Concentration.finiteContextArmPositiveVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.FiniteGapCutoff","target":"declaration:BanditRLProof.FiniteGapLayerCake.sum_le_cutoff_integral","relation":"contains"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"declaration:BanditRLProof.FiniteGapLayerCake.intervalIntegrable_step","relation":"contains"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"declaration:BanditRLProof.FiniteGapLayerCake.integral_step","relation":"contains"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"declaration:BanditRLProof.FiniteGapLayerCake.intervalIntegrable_card","relation":"contains"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"declaration:BanditRLProof.FiniteGapLayerCake.sum_eq_layerCake","relation":"contains"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"declaration:BanditRLProof.FiniteGapLayerCake.sum_le_refined_integral","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.score_le_foldl_select","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.chooseFin","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.score_le_chooseFin","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.measurable_selected_score","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.measurable_foldl_select","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.measurable_chooseFin_of_forall_measurable","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.choose","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.score_le_choose","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.measurable_choose_of_forall_measurable","relation":"contains"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"declaration:BanditRLProof.FiniteRealArgmax.measurable_selected_score_of_forall_measurable","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.Arm","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.center","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.region","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.ell","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.center_child_of_lt","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.center_child_last","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.mem_child_iff","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.region_children","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.region_measurable","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.region_diameter","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.ball_subset_region","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.region_disjoint","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.covering","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.mean","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.mean_le_best","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.first_coordinate_distance","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.weaklyLipschitz","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.bitReward","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.bitKernel","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.law","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.law_bounded","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.law_zero_mass","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.law_not_dirac","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.law_mean","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.global_sup","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.poorNode","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.poor_region_mean","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.poor_sup","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorModel","target":"declaration:BanditRLProof.HOO.CantorModel.expected_poor_visits","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorRate","target":"declaration:BanditRLProof.HOO.CantorModel.packing_le_two_div","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorRate","target":"declaration:BanditRLProof.HOO.CantorModel.dimension_le_two","relation":"contains"},{"source":"module:BanditRLProof.HOOCantorRate","target":"declaration:BanditRLProof.HOO.CantorModel.expected_actual_rate","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalPacking","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.packingExponent","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.packingExponent_zero","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.packingExponent_positive","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalityDimension","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalityDimension_nonneg","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.packing_le_rpow_of_exponent_lt","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.RegularCovering.eventually_nearOptimalPacking_le","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.RegularCovering.uniform_nearOptimalPacking_le","relation":"contains"},{"source":"module:BanditRLProof.HOODimension","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalNodes_power_bound","relation":"contains"},{"source":"module:BanditRLProof.HOOGeometry","target":"declaration:BanditRLProof.HOO.WeaklyLipschitz","relation":"contains"},{"source":"module:BanditRLProof.HOOGeometry","target":"declaration:BanditRLProof.HOO.regionSup","relation":"contains"},{"source":"module:BanditRLProof.HOOGeometry","target":"declaration:BanditRLProof.HOO.region_gap_le","relation":"contains"},{"source":"module:BanditRLProof.HOOGeometry","target":"declaration:BanditRLProof.HOO.near_optimal_region","relation":"contains"},{"source":"module:BanditRLProof.HOOLevels","target":"declaration:BanditRLProof.HOO.nodesAtDepth","relation":"contains"},{"source":"module:BanditRLProof.HOOLevels","target":"declaration:BanditRLProof.HOO.mem_nodesAtDepth","relation":"contains"},{"source":"module:BanditRLProof.HOOLevels","target":"declaration:BanditRLProof.HOO.card_nodesAtDepth","relation":"contains"},{"source":"module:BanditRLProof.HOOLevels","target":"declaration:BanditRLProof.HOO.Covering.exists_region_at_depth","relation":"contains"},{"source":"module:BanditRLProof.HOOLevels","target":"declaration:BanditRLProof.HOO.RegularCovering.disjoint_ball_family_card_le","relation":"contains"},{"source":"module:BanditRLProof.HOOLevels","target":"declaration:BanditRLProof.HOO.RegularCovering.exists_finite_packing_bound","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.child_subset","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.append_subset","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.descendant_subset","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.representative","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.representative_mem","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.RegularCovering","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.RegularCovering.region_near_optimal","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.nodeLaw","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.arm","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.arm_mem","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.measurable_arm","relation":"contains"},{"source":"module:BanditRLProof.HOOModel","target":"declaration:BanditRLProof.HOO.Covering.stepKernel_actual_arm","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.regionSup_children","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.optimalChild","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.optimalChild_sup","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.optimalPath","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.optimalPath_length","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.optimalPath_sup","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.optimalPath_prefix","relation":"contains"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"declaration:BanditRLProof.HOO.Covering.exists_first_unexpanded","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.ContainedBallPacking","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.packingSizes","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.zero_mem_packingSizes","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.packingSizes_bddAbove","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber_attained","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber_mono","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.containedPacking_card_le","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimalNodes_card_le_packing","relation":"contains"},{"source":"module:BanditRLProof.HOOPacking","target":"declaration:BanditRLProof.HOO.RegularCovering.packingNumber_le_ambient","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.mem_nearOptimalNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.root_nearOptimal","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.nearOptimal_prefix","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.boundaryNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.mem_boundaryNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.boundaryNodes_card_le","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.boundaryNodes_poor","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.node_partition_cover","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.descendant_nearOptimal_gap","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.descendant_boundary_gap","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.prefix_of_prefixes_length","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.deepGoodNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.shallowGoodNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.badSubtreeNodes","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.deep_shallow_disjoint","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.deep_bad_disjoint","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.shallow_bad_disjoint","relation":"contains"},{"source":"module:BanditRLProof.HOOPartition","target":"declaration:BanditRLProof.HOO.RegularCovering.actual_regret_partition","relation":"contains"},{"source":"module:BanditRLProof.HOOTailSum","target":"declaration:BanditRLProof.HOO.selectionFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.HOOTailSum","target":"declaration:BanditRLProof.HOO.selection_failure_exp_eq","relation":"contains"},{"source":"module:BanditRLProof.HOOTailSum","target":"declaration:BanditRLProof.HOO.selection_failure_le_telescope","relation":"contains"},{"source":"module:BanditRLProof.HOOTailSum","target":"declaration:BanditRLProof.HOO.selection_failure_sum_le_three","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailArmLaw","target":"declaration:BanditRLProof.HeavyTail.arm_coordinate_integral","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailArmLaw","target":"declaration:BanditRLProof.HeavyTail.arm_coordinate_integrable","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailArmLaw","target":"declaration:BanditRLProof.HeavyTail.arm_adaptive_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedConfidence","target":"declaration:BanditRLProof.HeavyTail.clipped_centered_mgf","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedConfidence","target":"declaration:BanditRLProof.HeavyTail.clipped_sum_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedConfidence","target":"declaration:BanditRLProof.HeavyTail.clipped_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.clip_eq_self","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.abs_clip_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.abs_clip_le_abs","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.abs_sub_clip_le_truncate","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.abs_sub_clip_moment_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.sq_clip_moment_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.measurable_clip","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.integral_clip_bias_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"declaration:BanditRLProof.HeavyTail.integral_sq_clip_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedScheduled","target":"declaration:BanditRLProof.HeavyTail.scheduled_clipped_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedScheduled","target":"declaration:BanditRLProof.HeavyTail.scheduled_adaptive_clipped_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedTransfer","target":"declaration:BanditRLProof.HeavyTail.integrable_of_raw_moment","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedTransfer","target":"declaration:BanditRLProof.HeavyTail.adaptive_corrupted_clipped_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedTransfer","target":"declaration:BanditRLProof.HeavyTail.observed_corrupted_clipped_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClippedTransfer","target":"declaration:BanditRLProof.HeavyTail.arm_corrupted_clipped_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.clip","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.abs_clip_sub_clip_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.clip_corruption_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.prefixMean","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.clipped_prefix_corruption_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.clipped_observed_prefix","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.corrupted_clipped_estimator_error_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.actual_clipped_corruption_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"declaration:BanditRLProof.HeavyTail.truncate_not_unit_lipschitz","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.exists_variance_tilt","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.neg_mgf","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.fixed_mgf_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.truncated_sum_abs_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.sum_mean_tail_of_centered","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.truncated_sum_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"declaration:BanditRLProof.HeavyTail.truncated_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"declaration:BanditRLProof.HeavyTail.bounded_centered_mgf","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"declaration:BanditRLProof.HeavyTail.bounded_centering_mgf","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"declaration:BanditRLProof.HeavyTail.truncated_centered_mgf","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"declaration:BanditRLProof.HeavyTail.independent_sum_mgf","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"declaration:BanditRLProof.HeavyTail.truncated_sum_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailGapThreshold","target":"declaration:BanditRLProof.HeavyTail.gapThreshold","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailGapThreshold","target":"declaration:BanditRLProof.HeavyTail.gapThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailGapThreshold","target":"declaration:BanditRLProof.HeavyTail.confidenceLog_mono","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailGapThreshold","target":"declaration:BanditRLProof.HeavyTail.twice_radius_lt_gap","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailPowerSum","target":"declaration:BanditRLProof.HeavyTail.rpow_increment_lower","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailPowerSum","target":"declaration:BanditRLProof.HeavyTail.sum_shifted_rpow_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailScheduledConfidence","target":"declaration:BanditRLProof.HeavyTail.scheduled_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailScheduledConfidence","target":"declaration:BanditRLProof.HeavyTail.scheduled_adaptive_mean_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.sourceTruncationThreshold","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.sourceThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.sourceThreshold_le_terminal","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.sourceThreshold_bias_average","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.sourceThreshold_variance_sum","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_centered_sum_upper_tail_sharp","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_centered_sum_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log_sharp","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.truncate_neg","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_lower_tail_log_sharp","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.horizonLog","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapBudget","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapThreshold","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.horizonLog_nonneg","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.horizonLog_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.gapThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.sourceLog_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"declaration:BanditRLProof.HeavyTail.SourcePolicy.twice_radius_le_gap","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.sourceConfidenceLog","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.sourceConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.sourceConfidenceLog_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.source_adaptive_mean_upper_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.source_adaptive_mean_lower_tail","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.inverse_sqrt_step","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.source_schedule_exp_eq","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.source_schedule_tail_le_telescope","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"declaration:BanditRLProof.HeavyTail.source_schedule_tail_sum_le_two","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTailSum","target":"declaration:BanditRLProof.HeavyTail.scheduled_exp_eq","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTailSum","target":"declaration:BanditRLProof.HeavyTail.cubic_tail_le_telescope","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTailSum","target":"declaration:BanditRLProof.HeavyTail.reciprocal_telescope","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTailSum","target":"declaration:BanditRLProof.HeavyTail.scheduled_tail_sum_le_two","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.truncate","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.abs_truncate_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.measurable_truncate","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.abs_sub_truncate_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.sq_truncate_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.integral_truncate_bias_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.integral_sq_truncate_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.transformed_observed_prefix","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"declaration:BanditRLProof.HeavyTail.estimator_error_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.power_threshold_bias_term","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.power_threshold_bias_sum","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold_factor","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold_bias_sum","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.power_scale_bias","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.power_bias_normalization","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold_bias_average","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.threshold_scale_identity","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.threshold_variance_identity","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.confidenceLog_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold_le_terminal","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.sampleThreshold_variance_sum","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"declaration:BanditRLProof.HeavyTail.tuned_radius_le","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailUnshiftedMGF","target":"declaration:BanditRLProof.HeavyTail.exp_le_one_add_self_add_three_quarters_sq","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailUnshiftedMGF","target":"declaration:BanditRLProof.HeavyTail.bounded_centering_mgf_unshifted_sharp","relation":"contains"},{"source":"module:BanditRLProof.HeavyTailUnshiftedMGF","target":"declaration:BanditRLProof.HeavyTail.bounded_centering_mgf_unshifted","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.FiniteActionHistory","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.FiniteRewardHistory","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.FinitePairHistory","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.FiniteHistory","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteActionHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteRewardHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.completeRewardTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.extendPairHistorySucc","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteActionHistoryOfTrace_apply","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteRewardHistoryOfTrace_apply","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.completeRewardTrace_finiteRewardHistoryOfTrace_apply_of_le","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteHistoryOfTrace_fst","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finiteHistoryOfTrace_snd","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finitePairHistoryOfTrace_apply","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.extendPairHistorySucc_apply_of_le","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.extendPairHistorySucc_apply_succ","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.finitePairHistoryOfTrace_succ","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.pairHistoryRewardProjection","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.pairHistoryRewardProjection_apply","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.pairHistoryRewardProjection_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteActionHistory_eval","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteRewardHistory_eval","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteHistory_action_eval","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteHistory_reward_eval","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_pairHistoryRewardProjection","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_extendPairHistorySucc","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteActionHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteRewardHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finiteHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyGenerators","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyGenerators_mono","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyMeasurableSpace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyMeasurableSpace_mono","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyMeasurableSpace_le","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltration","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltration_apply","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltrationSucc","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltrationSucc_apply","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurableSet_action_mem_historyFiltration","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_action_mem_historyFiltration_of_lt","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurableSet_reward_mem_historyFiltration","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_reward_mem_historyFiltration_of_lt","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.measurable_finitePairHistoryOfTrace_mem_historyFiltration_of_lt","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltration_succ_eq_comap_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltrationSucc_eq_comap_finitePairHistoryOfTrace","relation":"contains"},{"source":"module:BanditRLProof.HistoryFiltration","target":"declaration:BanditRLProof.History.historyFiltrationSucc_eq_of_action_eq_on_prefix","relation":"contains"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"declaration:BanditRLProof.IndependenceFoundation.iIndepFun_infinitePi_coord","relation":"contains"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"declaration:BanditRLProof.IndependenceFoundation.iIndepFun_rewardTrace_infinitePi","relation":"contains"},{"source":"module:BanditRLProof.IntegrabilitySums","target":"declaration:BanditRLProof.IntegrabilitySums.integrable_finset_sum","relation":"contains"},{"source":"module:BanditRLProof.IntegrabilitySums","target":"declaration:BanditRLProof.IntegrabilitySums.integrable_univ_sum","relation":"contains"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"declaration:BanditRLProof.Measure.compProd_restrict_prod","relation":"contains"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"declaration:BanditRLProof.Measure.compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","relation":"contains"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"declaration:BanditRLProof.Measure.map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","relation":"contains"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"declaration:BanditRLProof.IndepFun.comp_of_map","relation":"contains"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"declaration:BanditRLProof.indepFun_fst_snd_compProd_comap_of_indepFun","relation":"contains"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"declaration:BanditRLProof.map_snd_x_compProd_comap_eq_prod_map_of_indepFun","relation":"contains"},{"source":"module:BanditRLProof.KernelTrajectoryPrefix","target":"declaration:BanditRLProof.KernelTrajectoryPrefix.partialTraj_zero_congr","relation":"contains"},{"source":"module:BanditRLProof.KernelTrajectoryPrefix","target":"declaration:BanditRLProof.KernelTrajectoryPrefix.trajMeasure_map_frestrictLe_congr","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_one","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_succ_of_eq","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_succ_of_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_le_succ","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_succ_le_succ","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_lt_of_forall_succ_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_mono","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_le_time","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_add_le","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_le_add","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_eq_zero_of_forall_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_eq_time_of_forall_eq","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_pos_of_eq_before","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_eq_of_forall_lt","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_const_self","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_const_of_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_add_eq_of_forall_ne_between","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_add_eq_add_of_forall_eq_between","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pullCount_eq_list_filter_length","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_succ_of_eq","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_succ_of_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_eq_zero_of_forall_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_const_of_ne","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_add_eq_of_forall_ne_between","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_eq_list_range_foldl","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.sumRewards_eq_list_range_filter_foldl","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.FiniteBanditModel.bestMean_eq_mean_bestArm","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.FiniteBanditModel.gap_of_ne_bestArm","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_one","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_succ_of_bestArm","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_succ_of_gap_zero","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_eq_zero_of_forall_bestArm","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_eq_zero_of_forall_gap_zero","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_const_bestArm","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_const_of_gap_zero","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_add_eq_of_forall_bestArm_between","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_add_eq_of_forall_gap_zero_between","relation":"contains"},{"source":"module:BanditRLProof.LeafLemmas","target":"declaration:BanditRLProof.pseudoRegret_eq_list_range_foldl","relation":"contains"},{"source":"module:BanditRLProof.Literature","target":"declaration:BanditRLProof.UpstreamRef","relation":"contains"},{"source":"module:BanditRLProof.Literature","target":"declaration:BanditRLProof.lmlRef","relation":"contains"},{"source":"module:BanditRLProof.Literature","target":"declaration:BanditRLProof.lmlBanditDeclarationCards","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.mul_exp_neg_half_log_eq_sqrt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.integrable_sqrt_rnDeriv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.integrable_exp_neg_half_llr","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.integral_exp_neg_half_llr_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.exp_neg_half_integral_llr_le_rnAffinity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.rnAffinity_eq_commonDensityAffinity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.exp_neg_half_integral_llr_le_commonDensityAffinity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_le_half_commonDensityAffinity_sq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.sourceBlockList","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.sourceBlockList_length","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.sourceBlockList_injective","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.sourceBlockList_mass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.exists_arithmeticBlockSupport","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_expected_length_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_payload_interval","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_rate_sandwich","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_rate_tendsto_entropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticOffset","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticOffset_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticOffset_add_le_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticInterval","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_width","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticOffset_separated","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_head_separated","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_separated","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.arithmeticInterval_interior_unique","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"declaration:BanditRLProof.LowerBounds.exists_grid_cell_inside","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"declaration:BanditRLProof.LowerBounds.arithmeticAddress_prefix_forces_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"declaration:BanditRLProof.LowerBounds.exists_arithmeticPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"declaration:BanditRLProof.LowerBounds.arithmeticLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"declaration:BanditRLProof.LowerBounds.arithmeticLength_width_budget","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"declaration:BanditRLProof.LowerBounds.arithmeticLength_le_information_add_two","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"declaration:BanditRLProof.LowerBounds.supportTaggedWord","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"declaration:BanditRLProof.LowerBounds.supportTaggedWord_prefixFree","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.extendZeroMass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_extendZeroMass_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"declaration:BanditRLProof.LowerBounds.exists_zeroSafe_arithmeticCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","target":"declaration:BanditRLProof.LowerBounds.klDiv_map_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","target":"declaration:BanditRLProof.LowerBounds.klDiv_observedBanditHistory_le_expectedPulls_sum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.pairHistoryZeroMeasurableEquiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.pairHistorySuccMeasurableEquiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.pairHistoryZeroMeasurableEquiv_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.pairHistorySuccMeasurableEquiv_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment_initialFeedback","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.stationaryBanditHistoryEnvironment_feedback_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalPolicyArmMass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalExpectedPullCountThrough_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.lintegral_lintegral_fin_eq_sum_armMass_mul","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryPullCountENNReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_pairHistoryZeroMeasurableEquiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_pairHistorySuccMeasurableEquiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.lintegral_historyStepKernel_armIndicator_eq_policy_mass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_zero_general","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_succ_general","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_expectedPullCount_mul_armKL_general","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_realizedExpectedPullCount_mul_armKL","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.worstCaseExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal.mem_policyClass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal.eq_minimaxExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.expectedRegret_le_worstCaseExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret_le_worstCaseExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.le_minimaxExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.exists_alternative_le_average","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.alternativeExpectedPullBudget_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.exists_leastExploredAlternative","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.baseEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.changedEnvironmentRegretLowerBound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.entropy_product_term","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteEntropy_prod","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteEntropy_prod_probability","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_prod_probability","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.SourceBlock","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.sourceBlockMass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.sourceBlockMass_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.sum_sourceBlockMass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_sourceBlockMass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.exists_sourceBlock_code_rate_sandwich","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.sourceBlock_code_rate_lower_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.exists_sourceBlock_code_family_tendsto_entropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"declaration:BanditRLProof.LowerBounds.sourceBlock_code_family_limit_ge_entropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CodingEntropyBound","target":"declaration:BanditRLProof.LowerBounds.entropy_term_le_codeLength_remainder","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CodingEntropyBound","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_le_expectedCodeLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"declaration:BanditRLProof.LowerBounds.llr_ae_eq_log_commonDensity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"declaration:BanditRLProof.LowerBounds.integrable_commonDensity_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_of_integrable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_eq_if","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_klFun","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.memLp_sqrt_of_integrable_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.integral_sqrt_mul_sq_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.commonDensityAffinity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.integrable_commonDensityAffinity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.half_commonDensityAffinity_sq_le_overlap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.commonDensityComparisonEvent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.measurableSet_commonDensityComparisonEvent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.integrable_min_commonDensity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_eq_testingError","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_le_testingError","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_le_commonDensityOverlap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDomination","target":"declaration:BanditRLProof.LowerBounds.exists_commonFiniteDominatingMeasure","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDomination","target":"declaration:BanditRLProof.LowerBounds.exists_commonSigmaFiniteDominatingMeasure","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CommonDomination","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_lt_top_iff_ac","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_kernelRN_ae","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_kernelRN","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_ae","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_map_measurableEquiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_finiteAction_compProd_eq_lintegral_armKL","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_finiteAction_compProd_eq_sum_mass_mul_armKL","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CrossEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteCrossEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CrossEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteCrossEntropy_sub_entropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CrossEntropy","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_crossEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.CrossEntropy","target":"declaration:BanditRLProof.LowerBounds.entropyTerm_tendsto_zero_right","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.binaryAddressValue","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.binaryAddressValue_lt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.exists_binaryAddress","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.binaryAddressValue_append","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.dyadicAddressLower","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.dyadicAddressUpper","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.dyadicAddress_width","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.dyadicAddress_nonempty","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.dyadicAddress_append_contained","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.dyadicAddress_prefix_contained","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"declaration:BanditRLProof.LowerBounds.exists_dyadicAddress_inside","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.absolutelyContinuous_iff_atom_support","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.rnDeriv_mul_atom","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.rnDeriv_atom_eq_div","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_atom_support_mismatch","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_klFun","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_sum_log","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_if","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_top_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.totalMass_klFun_le_relativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.sum_relativeEntropy_restrict_fibers","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_map_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_le_relativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_map_le_finitePartitionRelativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_fin_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_map_eq_if","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.exists_binary_map_relativeEntropy_eq_top_of_event","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_top_of_not_absolutelyContinuous","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy_of_not_absolutelyContinuous","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","target":"declaration:BanditRLProof.LowerBounds.exists_fin_encoding_of_finite_range","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","target":"declaration:BanditRLProof.LowerBounds.exists_fin_observation_densityApproximation","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_mono","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","target":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FixedLengthCoding","target":"declaration:BanditRLProof.LowerBounds.exists_fixedLengthPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FixedLengthCoding","target":"declaration:BanditRLProof.LowerBounds.exists_ceilingLogPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.FixedLengthCoding","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_fixedLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance_pos","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianIIDObservationLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianCoordinateAverage","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianIIDSumLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianIIDSampleMeanLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_zero_error_event","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_gap_error_event","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_id_gaussianReal_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_gap_sub_id_gaussianReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_Ici_le_exp_neg_sq_div_two_variance","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_gap_Iio_half_le_exp_neg_sq_div_two_variance","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_le_exp","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability_le_exp","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_mills_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_source_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison_denominator_pos","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.hasDerivAt_gaussianMillsComparison","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison_lower_derivative_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsComparison_pos","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.tendsto_gaussianMillsComparison","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMills_lower_integral","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMills_sign_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMills_sign_threshold","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_factor","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_nonneg_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsErrorDerivative_source_nonneg_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsError","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.hasDerivAt_gaussianMillsError","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsError_source_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.tendsto_gaussianMillsError","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMillsError_source_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussian_integral_split","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMills_upper_integral","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_tail_integral","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_half_tail_integral","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_half_mills_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_standardized_tail","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_mills_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"declaration:BanditRLProof.LowerBounds.gaussianMills_expression_rescale","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal_ne_top","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.sum_finiteHistoryPullCountENNReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.sum_finiteHistoryPullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountReal_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryPullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.ofReal_mul_probReal_le_lintegral_of_event","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.sixteen_div_twentySeven_le_exp_neg_half","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryGaussianPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret_eq_sum_expectedPulls","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPullCountReal_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.sum_gaussianExpectedPullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseMean","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseMean_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseMean_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_selected","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedMean_other","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxChangedEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_ne_top","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_toReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxBaseSmallPullEvent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.measurableSet_gaussianMinimaxBaseSmallPullEvent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_base_toReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGaussianPseudoRegret_changed_toReal_lower","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.base_event_forces_gaussianPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.changed_complement_forces_gaussianPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.base_event_probability_lower_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.changed_complement_probability_lower_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianKernel_base_changed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianMinimax_base_changed_history","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_leastExploredAlternative","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_historyKL_le_half","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.horizon_mul_gaussianMinimaxGap_eq_half_sqrt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.exists_unitGaussianBandit_expectedPseudoRegret_ge_sqrt_div_twentySeven","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.exists_unitGaussianBanditEnvironment_expectedPseudoRegret_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_same_variance","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_same_variance_ae","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_same_variance","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_same_variance","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.three_fifths_le_exp_neg_half","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.gaussian_testing_error_lower_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.gaussian_testing_error_three_tenths","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"declaration:BanditRLProof.LowerBounds.gaussian_testing_max_error_three_twentieths","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.tailAtLeast","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHighProbabilityThreshold","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegretReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.GapOneGaussianBanditEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gapOneGaussianExpectedPseudoRegretReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gapOneGaussianRandomPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment.toGapOne","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gapOneGaussianExpectedPseudoRegretReal_toGapOne","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gapOneGaussianRandomPseudoRegret_toGapOne","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityGap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_gaussianRandomPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integrable_gaussianRandomPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret_toReal_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integral_gaussianRandomPseudoRegret_eq_expected","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integral_le_threshold_add_bound_mul_tailMass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegretReal_base_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.horizon_mul_sqrt_div_eq_sqrt_mul","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.sqrt_mul_mul_sqrt_div_eq_alternatives","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.horizon_mul_stochasticHighProbabilityGap_div_two","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticHighProbability_informationExponent_le_log","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1_of_four_mul_delta_lt_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1_unitCube","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticMinimax_sourceTerm_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold_at_minimax_scale","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold_le_quarter_root","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.minimax_expected_scale_identity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integral_exp_neg_rpow_inv_le_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integral_le_scale_of_all_rpow_log_tail","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_corollary17_2","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.noUniformGaussianRandomPseudoRegretTail_corollary17_3","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.exists_tailMass_ge_of_integral_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.exists_cdfTail_ge_of_integral_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measureReal_diff_ge_delta","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression_ge_quarter","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.randomRegret_ge_quarter_of_clippingDecomposition","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.clipUnitReward","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.clipUnitReward_mono","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.clipUnitReward_eq_self","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHardShift","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedArmLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClippedArmMap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_reward_marginal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipHistory","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClipHistory","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipHistoryAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipHistory_pullCount","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipHistory_pullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialUnclippedKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedKernel_eq_map","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipped_initialPairLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipped_historyStepLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipped_prefixStepLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_eq_map","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistoryLaw_pullSmall","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialUnclippedKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_common_scale","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.klDiv_adversarialUnclippedKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.klDiv_adversarialUnclipped_base_changed_history","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClaim17_6Gap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClaim17_6Gap_information_calibration","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.sum_adversarialUnclipped_expectedPulls","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedHistory_pull_le_half_claim17_6","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_distinguished","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_le_two_mul","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHardShift_le_gap_of_ne","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward_distinguished_mono","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward_gap_of_not_clipped","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialPullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippingCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipIndicator","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClipIndicator","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClipIndicator_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.gaussianReal_tenth_abs_quarter_le_eighth","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integral_adversarialClipIndicator_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClippingCount_tail_claim17_7","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialBoundaryClippingCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialBoundaryClippingCountReal_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialBoundaryClippingCount_tail_claim17_7","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialComparatorRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialComparatorRegret_le_randomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialComparatorRegret_ge_eq17_8","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_eq17_8","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.clipUnitReward_eq_self_of_ne_endpoints","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_boundary_eq17_8","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullClippedReward","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullHardShift_separation","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullClippedReward_best","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullRandomRegret_ge_boundary_eq17_8","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount_tail_claim17_7","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.AdversarialRewardTable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableInitialFeedback","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableNextFeedback","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableStepKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableStepKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableHistoryKernel_prefix_congr","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullRewardTable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialFullRewardTable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialFullRewardTable_at","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_noise_marginal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel_update_future","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_split","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryKernel_split_future","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialCenteredNoiseLaw_split","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw_full_reward_marginal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialTableHistoryKernel_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialTableHistoryKernel_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialFreshNoise_step","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_succ_slice","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_succ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.lintegral_adversarialNoiseHistoryKernel_eq_clipped","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_pull_le_half_claim17_6","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialClippingCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_clipping_tail","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_good_event","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHistoryActions","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHistoryActions_pullCountENNReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHistoryActions_pullCountReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHistory_randomRegret_ge_quarter","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_randomRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.exists_kernel_section_mass_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialRandomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialJointRandomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.exists_adversarialTable_randomRegret_tail","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialConfidence_log_calibration","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialClaim17_6Gap_tenth_sq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialHorizon_calibration","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialThreshold_calibration","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.exists_adversarialTable_randomRegret_gt_theorem17_4","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableRandomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTableCDF","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.measurable_adversarialTableRandomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.integrable_adversarialTableRandomRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialTable_strictTail_eq_one_sub_CDF","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_theorem17_4","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_relabel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.relabel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.HuffmanRemainder","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.huffmanSplitEquiv","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.huffmanSplitEquiv_false","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.huffmanSplitEquiv_true","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.huffman_merged_card_lt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"declaration:BanditRLProof.LowerBounds.exists_two_least_weights","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"declaration:BanditRLProof.LowerBounds.oneBitCode_optimal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"declaration:BanditRLProof.LowerBounds.emptyRemainderRoot","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"declaration:BanditRLProof.LowerBounds.huffmanOptimalCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"declaration:BanditRLProof.LowerBounds.huffmanCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"declaration:BanditRLProof.LowerBounds.huffmanCode_optimal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"declaration:BanditRLProof.LowerBounds.huffmanCode_entropy_sandwich","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanStep","target":"declaration:BanditRLProof.LowerBounds.exists_oriented_sibling_code","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.HuffmanStep","target":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.expand_least_weights","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.uniquelyDecodable_range","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.codebook","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.coe_codebook","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.kraft_inequality","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.discreteEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.discreteEntropy_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_absolutelyContinuous_of_integrable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_probability_absolutelyContinuous_of_integrable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_not_absolutelyContinuous","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_zero_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.rnDeriv_restrict_restrict","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_restrict_add_compl","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bernoulliKLCore_event_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.mul_sqrt_div_eq_sqrt_mul","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.exp_neg_half_bernoulliKLCore_le_affinity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.half_binaryAffinity_sq_le_eventError","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuberCore","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentPolicyOver","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.gap_bestArm","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMeanIncrease","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmChangedMargin","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.mean_eq_of_armLaw_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.InUnstructuredClass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_law","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_unique","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.withImprovedArm_mem","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMeanIncrease_sub_gap_eq_changedMargin","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMeanChange_produces_gap_contract","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.gap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.gap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean_mean","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment.toFiniteMean_gap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_changed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_other","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_uniqueBest","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedMean_mem_localBox","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_uniqueBest","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_sameArmLaw","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_meanIncrease","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_changedMargin","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL_toReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.InChapter16GaussianLocalClass","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.inChapter16GaussianLocalClass_self","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.inChapter16GaussianLocalClass_changed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.add","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_add_le_rpow","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_pull_div_log_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.liminf_pull_div_log_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.divergenceInfimum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_exists_alternative_lt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment.exists_confusingEnvironment_lt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_eq_top_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_le_perturbed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMajorityPullEvent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.measurableSet_oneArmMajorityPullEvent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryGapPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_ne_top","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_toReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough_general","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_ne_top","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_ne_top","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal_eq_sum_expectedPulls","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.unitVarianceGaussianExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMajority_forces_gapPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_forces_gapPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMajority_probability_charge_le_expectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_mul_eq_exp","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.exp_testing_bound_of_majority_regret_bounds","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_of_exp_testing_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.gapPseudoRegret_add_pos_of_only_arm_changed","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedPullCount","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedPullCount_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteMeanExpectedRegret_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.finiteMeanNormalizedRegret_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.consistentRegret_liminf_expectedPull_div_log_ge_of_alternative","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedPull_div_log_ge_inv_dInf","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedRegret_div_log_ge","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPullCount_ge_finiteTimeInstanceDependent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.chapter16Gaussian_finiteTime_log_identity","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianArm","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianBandit","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_gaussianPDFReal_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianArm_zero_two_mul","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_sq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_le_half","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.binaryWords","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.card_binaryWords","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.mem_binaryWords_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.binaryExtensions","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.card_binaryExtensions","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.mem_binaryExtensions_iff","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.binaryExtensions_disjoint_of_incomparable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_binaryWord_avoiding_prefixes","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.binary_level_mul_kraft_weight","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_binaryWord_of_kraft_lt_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_prefixFree_insert_of_kraft_lt_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_prefix_encoding_of_kraft_le_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_prefix_encoding_of_kraft_lt_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_binaryPrefixCode_of_kraft_le_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_binaryPrefixCode_of_kraft_lt_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_prefixCode_of_uniquelyDecodable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"declaration:BanditRLProof.LowerBounds.exists_binaryPrefixCode_entropy_sandwich","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.relabel","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_swap","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_swap_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.length_antitone","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.entropy_sandwich","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.one_le_expectedCodeLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.singletonPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"declaration:BanditRLProof.LowerBounds.singletonPrefixCode_optimal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_swap_le_allow_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","target":"declaration:BanditRLProof.LowerBounds.exists_no_worse_least_weight_siblings","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.deepest_parent_incomparable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.replaceWord","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.pruneDeepest","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_replaceWord","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_pruneDeepest","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_pruneDeepest_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.totalCodeLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.totalCodeLength_pruneDeepest","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.exists_minimal_totalCodeLength_competitor","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.exists_competitor_with_deepest_siblings","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"declaration:BanditRLProof.LowerBounds.exists_no_worse_deepest_sibling_pair","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.extended_prefix_parent_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.siblingExpandedWord","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.siblingExpandedWord_prefixFree","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.expandSibling","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_expandSibling","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.sibling_parent_not_prefix_other","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.siblingContractedWord","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.siblingContractedWord_prefixFree","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.contractSibling","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_contractSibling","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.binaryRootPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"declaration:BanditRLProof.LowerBounds.binaryRootPrefixCode_optimal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_eq_lintegral_condExp","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_map_eq_trim_of_absolutelyContinuous","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_trim_of_density_measurable","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"declaration:BanditRLProof.LowerBounds.densityApproximationFiltration","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"declaration:BanditRLProof.LowerBounds.measurable_density_iSup_approximationFiltration","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_densityApproximation_trim","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_triangle_counterexample","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","target":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_asymmetry","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.shannonLength","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.shannonLength_pos","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.shannonLength_kraft_weight_lt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.shannonLength_le_information_add_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.weighted_shannonLength_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.sum_weighted_shannonLength_le_entropy_add_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.sum_positive_shannon_weights_lt_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"declaration:BanditRLProof.LowerBounds.exists_lengths_kraft_lt_one_entropy_bound","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.UnitSubgaussianBanditEnvironment","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment.toSubgaussian","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.subgaussianExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.subgaussianWorstCaseExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.subgaussianMinimaxExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.subgaussianExpectedPseudoRegret_gaussian","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimax_le_subgaussianMinimax","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.moss_subgaussianExpectedPseudoRegret_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.subgaussianMinimax_sandwich","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"declaration:BanditRLProof.LowerBounds.moss_nearMinimax","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQSet_bddAbove","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.le_sourceQ_of_mem","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_le_norm","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_zero","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.abs_inner_le_sourceQ_of_mem","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_eq_zero_of_atom_orthogonal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceR","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceRSet_not_bddAbove_of_nonzero_atom_orthogonal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.correlationSum_le_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_basis_basis","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.orthonormal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_le_maxAbsCoefficient","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.exists_abs_eq_maxAbsCoefficient","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportCombination","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom_mem","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_basis","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_signedSupportAtom","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_coefficientSign","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_coefficientSign","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul_self","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportSignCombination","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_coefficientSign","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportSignCombination","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.norm_sq_supportSignCombination","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_supportSignCombination","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_mul_sourceQ","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_of_sourceQ_le_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceRSet_bddAbove","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceR_supportCombination_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctAt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsStrictlySuccinctAt","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceR_eq_sumAbs","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.inner_supportSignCombination_eq_sumAbs","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceQ_supportSignCombination_eq_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.norm_sq_supportSignCombination_eq_size","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sumAbs_eq","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.abs_inner_strictBasis_supportSignCombination_eq_one","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.strictSize_le","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation.abs_coefficient_pos","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.strictlySuccinctSize_unique","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"declaration:BanditRLProof.LowerBounds.uniformPowerTwo_entropy","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"declaration:BanditRLProof.LowerBounds.fixedLength_uniformPowerTwo_optimal","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"declaration:BanditRLProof.LowerBounds.ternaryPrefixWord","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"declaration:BanditRLProof.LowerBounds.ternaryPrefixCode","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"declaration:BanditRLProof.LowerBounds.ternaryPrefixCode_uniform_length","relation":"contains"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"declaration:BanditRLProof.LowerBounds.uniform_three_fixedLength_not_optimal","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.toPrefix","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.stronglyAdapted'","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.integrable'","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifference.condExp_succ_ae_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.stronglyAdapted'","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.integrable_of_lt","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.SuccMartingaleDifferencePrefix.condExp_succ_ae_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.centeredRewardProcess","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.succMartingaleDifference_centeredRewardProcess_of_condExp","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.succMartingaleDifferencePrefix_centeredRewardProcess_of_condExp","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.partialSumsSucc","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.martingale_partialSumsSucc_of_succMartingaleDifference","relation":"contains"},{"source":"module:BanditRLProof.MartingaleDifference","target":"declaration:BanditRLProof.MartingaleDiff.martingale_partialSumsSucc_centeredRewardProcess_of_condExp","relation":"contains"},{"source":"module:BanditRLProof.MathlibWrappers","target":"declaration:BanditRLProof.pullCount_eq_finset_filter_card","relation":"contains"},{"source":"module:BanditRLProof.MathlibWrappers","target":"declaration:BanditRLProof.sumRewards_eq_finset_filter_sum","relation":"contains"},{"source":"module:BanditRLProof.MathlibWrappers","target":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum","relation":"contains"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"declaration:BanditRLProof.sumRewards_eq_finset_range_indicator_reward","relation":"contains"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"declaration:BanditRLProof.measurable_sumRewards","relation":"contains"},{"source":"module:BanditRLProof.MeasurablePullCount","target":"declaration:BanditRLProof.measurable_pullCount","relation":"contains"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"declaration:BanditRLProof.measurable_natCast_pullCount","relation":"contains"},{"source":"module:BanditRLProof.MeasurableRegret","target":"declaration:BanditRLProof.measurable_pseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.MeasurableSums","target":"declaration:BanditRLProof.measurable_finset_sum_indicator_reward","relation":"contains"},{"source":"module:BanditRLProof.MeasureFoundation","target":"declaration:BanditRLProof.measurableSet_actionTrace_eval_eq","relation":"contains"},{"source":"module:BanditRLProof.MeasureFoundation","target":"declaration:BanditRLProof.measurable_actionTrace_eval_eq_indicator_const","relation":"contains"},{"source":"module:BanditRLProof.MeasureFoundation","target":"declaration:BanditRLProof.measurable_actionTrace_eval_eq_indicator_reward","relation":"contains"},{"source":"module:BanditRLProof.MeasureL2Indicator","target":"declaration:BanditRLProof.integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeScheduledScalarRidgeConfidenceFailureSet","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.not_mem_allTimeScheduledScalarRidgeConfidenceFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_allTimeScheduledScalarRidgeConfidenceFailureSet_le_tsum","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight_eq_sub","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.sum_range_allTimeTelescopingWeight","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingWeight_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.hasSum_allTimeTelescopingWeight","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_eq_div","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.tsum_ofReal_allTimeTelescopingDelta","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingScalarRidgeConfidenceFailureSet","relation":"contains"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_matrix_det_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_matrix_adjugate_apply_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_matrix_nonsingInv_apply_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_matrix_mulVec_apply_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_dotProduct_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_confidenceWidth_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_optimisticScore_of_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonFeatureGram_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonResponseVector_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonRidgeEstimate_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonScalarConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryObservedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryObservedResponse","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryObservedFeature_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryObservedResponse","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeEstimate","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeDesign","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeRadius","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryFixedActionFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticScore","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScalarRidgeOptimisticScore","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction_score_max","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeSelectedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_observedFeature_succ_ae_eq_finiteHistoryScalarRidgeSelectedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.matrixNorm","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.matrixNorm_add_le","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.matrixNorm_sub_le","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.matrixNorm_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonResponseVector","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonRidgeEstimate","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonConfidenceThreshold","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.one_le_det_add_posSemidef_div_det","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonConfidenceThreshold_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonResponseVector_eq_featureGram_mulVec_add_noiseScore","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.posDef_nonsingInv_mulVec_mulVec","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.posDef_mulVec_nonsingInv_mulVec","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.matrixNorm_nonsingInv_mulVec_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonRidgeEstimate_sub_eq_inverseScore_sub_inverseBias","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.finiteHorizonRidgeEstimate_error_matrixNorm_le_score_add_bias","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.selfNormalizedQuadratic_gt_of_ridgeEstimate_error_matrixNorm_gt_confidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.rankOneGram_eq_replicateCol_mul_replicateRow","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.rankOneGram_isHermitian","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_one_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_rankOne_update_factor_eq_one_add_dotProduct_inv_mulVec","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_scalar_identity","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_scalar_identity_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.isUnit_det_scalar_identity","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_scalar_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.quadraticForm","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.quadraticForm_add","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.dotProduct_mulVec_eq_quadraticForm","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.quadraticForm_scalar_identity","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_sq_pos_of_exists_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.exists_coord_ne_zero_of_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.featureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.featureGram_isHermitian","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_eq_scalar_add_featureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_isHermitian","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prefixFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prefixFeatureGram_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prefixFeatureGram_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prefixFeatureGram_isHermitian","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_isHermitian","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_scalar_identity","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_prefixFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_regularizedPrefixFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_regularizedPrefixFeatureGram_le","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.finset_prod_le_pow_sum_div_card_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prod_univ_le_pow_sum_div_card_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_posSemidef_le_pow_trace_div_card","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_bound_average","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.rankOneGram_quadraticForm_eq_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.rankOneGram_posSemidef","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.rankOneGram_quadraticForm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.featureGram_quadraticForm_eq_sum_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prefixFeatureGram_quadraticForm_eq_sum_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_eq_sum_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_eq_sum_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.featureGram_quadraticForm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.prefixFeatureGram_quadraticForm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_quadraticForm_pos_of_pos_lambda","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_quadraticForm_pos_of_pos_lambda","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_div_card","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_pow_trace_bound_average_of_pos_lambda","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_quadratic_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_det_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_det_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_det_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_det_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.isUnit_det_regularizedFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.isUnit_det_regularizedPrefixFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedFeatureGram_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_add_rankOneGram_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_add_rankOneGram_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedFeatureGram_add_rankOneGram_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_add_rankOneGram_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedFeatureGram_rankOne_update_factor_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_rankOne_update_factor_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.log_det_regularizedFeatureGram_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_add_rankOneGram","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.log_det_regularizedFeatureGram_add_rankOneGram_sub","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_add_rankOneGram_sub","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_succ_sub","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_forward_difference","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_log_update_factor_eq_log_det_ratio","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.div_one_add_self_le_log_one_add","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.self_le_two_log_one_add_of_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.one_le_two_log_two","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.min_one_le_two_log_one_add","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_one_le_two_sum_log_one_add","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_one_le_two_of_sum_log_one_add_le","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_le_two_sum_log_one_add_of_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_le_two_of_sum_log_one_add_le_of_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_ratio","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_log_regularizedPrefixFeatureGram_update_factor_eq_log_det_sub_base","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_sub_base","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_sub_base","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_sub_base_of_pos_lambda","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_sub_base_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_log_det_upper","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_log_det_upper_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.log_det_regularizedPrefixFeatureGram_sub_base_le_of_det_le_mul_exp","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_det_mul_exp_upper","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_det_mul_exp_upper_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp_dim_scaled","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp_log","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_average_exp_exponent_dim_cancel","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.trace_average_pow_le_lambda_pow_mul_exp","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_of_trace_average_bound","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_of_trace_average_bound","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_of_trace_average_bound_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_dim_scaled","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_dim_scaled","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_dim_scaled_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.det_regularizedPrefixFeatureGram_le_mul_exp_trace_average_log","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_log","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"declaration:BanditRLProof.OFUL.sum_range_prefix_update_le_two_trace_average_log_of_update_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULEllipticalPotentialFoundation","target":"declaration:BanditRLProof.OFUL.standardLogDeterminantAndEllipticalPotential","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGap","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.IsOptimalLinearArm","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.measurable_canonicalHistoryTrajectorySumRangeAllGap","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.abs_linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.abs_canonicalHistoryTrajectorySumRangeAllGap_le_envelope","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.integrable_canonicalHistoryTrajectorySumRangeAllGap","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.measurableSet_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_real_measure","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectorySumRangeAllFixedComparatorGap_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standard_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardExpectedRegretDelta","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardExpectedRegretDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardExpectedRegretDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_mul_standardExpectedRegretDelta","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_standardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.standardExpectedRegretLogBudget_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.sqrt_log_succ_isBigO_log_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.standardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_isBigO_sqrt_mul_log","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.sqrt_mul_log_succ_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.standardExpectedPseudoRegretBound_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedPseudoRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.standardExpectedRegretLogBudget","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.standardExpectedRegretAlgorithmDelta_eq_inv_sq","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_standardExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.standardExpectedPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound_add_initial_eq_standardExpectedPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.integral_canonicalHistoryTrajectoryPseudoRegret_nonneg_and_le_explicitStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.confidenceWidth","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.linearValue","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.optimisticScore","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteActionArgmax","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteActionArgmax_spec","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.abs_dotProduct_le_matrixNorm_mul_confidenceWidth","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.linearValue_le_optimisticScore_of_matrixNorm_sub_le","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.optimisticScore_le_linearValue_add_two_mul_bonus_of_matrixNorm_sub_le","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteActionOptimisticChoice","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteActionOptimisticChoice_score_max","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.linearValue_sub_finiteActionOptimisticChoice_le_two_mul_bonus","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarGram","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarOptimisticAction_gap_le","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarOptimismViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizonScalarOptimismViolationSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.finiteHorizonNoiseScore","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.varianceWeightedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.inner_finiteHorizonNoiseScore_eq_sum_projection_mul_noise","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram_quadraticForm_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram_posSemidef","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonVarianceGram_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonNoiseScore","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.compensatedScore_eq_inner_sub_varianceGram","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.finiteHorizonScoreVarianceGram_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.integral_exp_inner_finiteHorizonScore_sub_varianceGram_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.integral_gaussianQuadraticExponential_finiteHorizon_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_prod_multivariateGaussian_zero_inv_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.inner_toEuclideanCLM_selfAdjoint","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.inner_toEuclideanCLM_congruence","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_transformed","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.sqrt_inv_congruence_factorization","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.det_one_add_sqrt_inv_congruence_eq_ratio","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.sqrt_inv_congruence_inverse","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.dotProduct_sqrt_inv_congruence_inverse","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv_detRatio","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"declaration:BanditRLProof.OFUL.integrable_exp_inner_sub_quadratic_multivariateGaussian_zero_inv","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_multivariateGaussian_zero_inv","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"declaration:BanditRLProof.OFUL.finiteHorizonDetRatioInvGramExponential","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_finiteHorizon_multivariateGaussian_zero_inv","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"declaration:BanditRLProof.OFUL.measurable_finiteHorizonDetRatioInvGramExponential","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"declaration:BanditRLProof.OFUL.lintegral_finiteHorizon_detRatio_invGramExponential_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_linear_sub_quadratic_gaussianReal_zero_one","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_sum_linear_sub_diagonal_quadratic_pi_gaussianReal_det","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.gaussianQuadraticExponential","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.measurable_gaussianQuadraticExponential_dot","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.measurable_gaussianQuadraticExponential","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.gaussianQuadraticExponentialENNReal","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.measurable_gaussianQuadraticExponentialENNReal","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_prod","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"declaration:BanditRLProof.OFUL.lintegral_gaussianQuadraticExponentialENNReal_prod_multivariateGaussian_zero_inv","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.inner_sum_smul_orthonormalBasis","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.inner_sum_smul_orthonormalBasis_left","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.inner_toEuclideanCLM_sum_eigenvectorBasis","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_stdGaussian_eigenvalues","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.det_one_add_posSemidef_eq_prod_eigenvalues","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.toEuclideanCLM_one_add_posSemidef_inv_eigenvectorBasis","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.dotProduct_one_add_posSemidef_inv_mulVec_eq_sum_eigenvalues","relation":"contains"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"declaration:BanditRLProof.OFUL.integral_exp_inner_sub_quadratic_stdGaussian_det","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.linearValue_sub_selected_le_two_mul_bonus_of_score_max","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeOptimisticAction_gap_le","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryResponse","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHistoryObservedFeature_finitePairHistoryOfTrace_of_le","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHistoryObservedResponse_finitePairHistoryOfTrace_of_le","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHorizonFeatureGram_finitePairHistoryOfTrace_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHorizonResponseVector_finitePairHistoryOfTrace_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeDesign_finitePairHistoryOfTrace_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeEstimate_finitePairHistoryOfTrace_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidgeRadius_finitePairHistoryOfTrace_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_gap_le_on_uniformConfidence","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryBeforeFiltration_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScalarRidgeOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_zero_ae_eq_initialArm","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_predictableFeature_all","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableResidual","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.measurable_canonicalHistoryTrajectoryResponse_before_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.finiteActionProjectionBound","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.finiteActionProjectionBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.predictableFeature_projection_le_finiteActionProjectionBound","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.CanonicalPredictableScalarRidgeResidualLaw","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.mem_scalarRidgeConfidenceFailureAt_iff_of_feature_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.mem_finiteHorizonUniformScalarRidgeConfidenceFailureSet_iff_of_feature_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le_of_predictableResidualLaw","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_predictableResidualLaw","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardSelectedWidthBudget","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardScalarRadiusWidthBound","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarGram_eq_regularizedPrefixFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_mono","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_le_standardUpper","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.standardSelectedWidthBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_subset","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","target":"declaration:BanditRLProof.OFUL.CanonicalScalarRidgeConfidenceSource","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectory_uniformScalarRidgeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeSuccGapViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.standardHighProbabilityRegretLogBudget","relation":"contains"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_highProbabilityRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.standardHighProbabilityPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound_eq_standardHighProbabilityPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_map_eq_historyAlgorithmInitialFeedback","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condDistrib_eq_historyStepReward","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_historyStepReward_comap","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_historyAlgorithmInitialFeedback_unitComap","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.finiteHistoryScalarRidge_historyStepKernel_map_snd","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_map_eq_scalarRidgeInitialFeedback","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condDistrib_eq_scalarRidgeStepReward","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_reward_succ_condExpKernel_map_eq_scalarRidgeStepReward_comap","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialReward_condExpKernel_map_eq_scalarRidgeInitialFeedback_unitComap","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.CanonicalLinearSubgaussianEnvironmentLaw","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.canonicalPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.abs_linearValue_le_euclideanLength_mul_euclideanLength","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.abs_linearValue_le_parameterFeatureBound","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.linearValue_sub_linearValue_le_two_mul_parameterFeatureBound","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.standardScalarInitialGapBound","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_initialGap_le_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapBound","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_ae_le_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeAllGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticScore","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAction_score_max","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.finiteHistoryOptimisticSelectedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_candidateFeature_succ_ae_eq_selectedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"declaration:BanditRLProof.OFUL.regularizedPrefixFeatureGram_inv_quadratic_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"declaration:BanditRLProof.OFUL.confidenceWidth_regularizedPrefixFeatureGram_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_confidenceWidth_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_radius_mul_width_le_standard_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"declaration:BanditRLProof.OFUL.measure_canonicalHistoryTrajectorySumRangeSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"declaration:BanditRLProof.OFUL.euclideanLength","relation":"contains"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"declaration:BanditRLProof.OFUL.scalarIdentity_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"declaration:BanditRLProof.OFUL.matrixNorm_nonsingInv_scalar_mulVec_le_sqrt_mul_euclideanLength","relation":"contains"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"declaration:BanditRLProof.OFUL.finiteHorizon_scalarRegularizationBias_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizonScalarRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.squareIntegrableFiniteStoppingTime_of_bounded_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_le_sq_of_bounded_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_reachHorizon_ae_of_spent_reach","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.AlignedWindowPositiveActionCostAE","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_iff_forall_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.canonicalActionCostProcess_alignedWindowPositive_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.budget_le_cumulativeActionCost_budget_mul_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_aeAlignedWindowPositiveActionCostBudgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.telescopingStandardScalarAllRoundGapBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_initialGap_le_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_subset_succ_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllHorizonAllRoundGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_antitone","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.standardScalarConfidenceRadiusUpper_antitone_delta","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper_of_indices","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.finiteHorizonScalarConfidenceRadius_telescoping_le_standardUpper","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.telescopingStandardScalarRadiusWidthBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_sum_range_succ_telescoping_radius_mul_width_le_standard_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_radius_mul_width_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_sum_range_succ_gap_le_standard_ae_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_subset_confidenceFailure_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllHorizonSuccGapStandardViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.allTimeTelescopingDelta_eq_outerBudget_div_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.standardHighProbabilityRegretLogBudget_outerBudget_eq_telescoping","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.standardHighProbabilityPseudoRegretBound_outerBudget_eq_telescoping","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingStandardScalarAllRoundGapBound_eq_telescopingHighProbabilityPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretAllHorizonViolationSet_eq_standard","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeRadius","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticScore","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticScore_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScheduledScalarRidgeOptimisticScore","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAction_score_max","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeOptimisticAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_initialAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidge_historyStepKernel_map_snd","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHistoryScheduledScalarRidgeSelectedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_observedFeature_succ_ae_eq_scheduledSelectedFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryScheduledScalarRidgeOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableFeature_stronglyMeasurable","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectory_action_zero_ae_eq_initialArm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_scheduledPredictableFeature_all","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableResidual","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.scheduledCanonicalHistoryTrajectoryPredictableResidual_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.scheduledPredictableFeature_projection_le_finiteActionProjectionBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.CanonicalScheduledPredictableScalarRidgeResidualLaw","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalScheduledPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.mem_allTimeTelescopingScalarRidgeConfidenceFailureSet_iff_of_feature_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_action_succ_ae_eq_scheduledAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectory_action_succ_gap_le_of_not_mem_confidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_subset_confidenceFailure_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectoryAllTimeSuccGapViolationSet_le_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedIndexSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.card_blockStartForcedIndexSet_le_div_add_one","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_card_mul","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_le_div_add_one_mul_two_mul_parameterFeatureBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableFeature_stronglyMeasurable","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonical_action_zero_ae_eq_initialArm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryFeature_ae_eq_deterministicHistoryPredictableFeature_all","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableResidual","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistoryCanonicalPredictableResidual_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistoryPredictableFeature_projection_le_finiteActionProjectionBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.deterministicHistory_historyStepKernel_map_snd","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.CanonicalDeterministicHistoryPredictableScalarRidgeResidualLaw","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalDeterministicHistoryPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_deterministicHistoryCanonical_allTimeTelescopingScalarRidgeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalBlockStartForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_blockStartForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalHistoryTrajectory_initialGap_le_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.Thompson.deterministicHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.Thompson.deterministicHistoryAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.Thompson.canonicalHistoryTrajectory_action_succ_ae_eq_deterministicHistorySelector","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_mul_window","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryBlockStartForcedTelescopingScalarRidgeAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_blockStart_succ_ae_eq_forcedAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_finiteHistoryBlockStartForcedTelescopingScalarRidgeAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalTrajectoryMeasure_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegretViolationSet_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.blockStartForcedHorizonIndexedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"declaration:BanditRLProof.OFUL.blockStartForcedIndexSet_horizon_self_eq_singleton","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_horizon_self","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret_horizon_self_le_two_mul_parameterFeatureBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegretHorizonWindowViolationSet_subset","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"declaration:BanditRLProof.OFUL.blockStartForcedCanonicalStandardHighProbabilityPseudoRegret_horizonWindow_nonneg_and_finiteHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_of_blockStartTelescopingActionCostPositive","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.alignedWindowPositiveActionCostAE_of_blockStartForcedTelescopingAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_blockStartForcedPositiveActionCostBudgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_of_mod_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.finiteHistoryBlockStartForcedTelescopingScalarRidgeAction_eq_forcedAction_of_mod_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_mod_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.blockStartForcedSuccessorPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.blockStartForcedActionSuccessorPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorRadiusWidthCharge","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_gap_le_of_eq_telescopingAction_of_not_mem_confidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_blockStartForced_add_optimistic","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.blockStartForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.blockStartOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_blockStartForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_forcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_mono","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_mono","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.standardScalarAllRoundGapEnvelope_mono","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_envelope","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_boundedStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_endpoint_add_envelope_mul_real_measure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedConfidenceLog_isBigO_log_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedRegretLogBudget_isBigO_log","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound_isBigO_sqrt_mul_log","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_isBigO_sqrt_mul_log","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","target":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedAverageStoppedPseudoRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedRegretLogBudget","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_standardExpectedRegretDelta","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingStandardExpectedPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_add_initial_standardExpectedRegretDelta_eq","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_telescopingStandardExpectedBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalTelescopingStandardExpectedStoppedPseudoRegret_nonneg_and_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryAllRoundFiltration","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectoryAllRoundFiltration_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.measurable_canonicalHistoryTrajectory_coordinate_allRound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.boundedTrajectoryTime_ne_top","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.coe_untopA_boundedTrajectoryTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.measurable_canonicalStandardHighProbabilityPseudoRegret_allRound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_stronglyAdapted_allRound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_stronglyAdapted_allRound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_subset_allHorizon","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_boundedStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_boundedStoppingTime_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.squareIntegrableFiniteStoppingTime_of_bounded","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_le_sq_of_bounded","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_budget_of_spent_budget","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_reachHorizon_of_spent_reach","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_reachHorizon","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_reachHorizon","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_reachHorizonRoundsSq_add_initialGap_mul_reachHorizonRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_reachedBy","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.block_le_cumulativeSpent_mul_of_alignedWindowPositive","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.budget_le_cumulativeSpent_budget_mul_of_alignedWindowPositive","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetWindowRoundsSq_add_initialGap_mul_budgetWindowRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativeAlignedWindowPositiveCostBudgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeSpent","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeSpent_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeSpent_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.adapted_cumulativeSpent","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeSpent_unitGrowth_of_one_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_cumulativePositiveCostBudgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.canonicalActionCostProcess","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.adapted_canonicalActionCostProcess","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.canonicalActionCostProcess_one_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeActionCost","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeActionCost_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.adapted_cumulativeActionCost","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.cumulativeActionCost_unitGrowth_of_one_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_positiveActionCostBudgetExhaustionTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorRadiusWidthCharge","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorPseudoRegret_le_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForced_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_le_initial_add_powerOfTwoForcedActionCharge_add_radiusWidthCharge_ae_of_not_mem_allTimeConfidenceFailure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.canonicalPowerOfTwoForcedPredictableScalarRidgeResidualLaw_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_powerOfTwoForcedCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorRadiusWidthCharge_le_telescopingStandardScalarRadiusWidthBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalHistoryTrajectory_initialGap_le_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_le_forcedActionCharge_add_explicitBound_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretAllHorizonViolationSet_subset_confidenceFailure_ae","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_of_not_mem_allHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegretConsistencyFailureSet_subset_allHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_tendsto_zero_off_violation_and_consistencyFailure_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isLittleO_natCast_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityAveragePseudoRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityAveragePseudoRegret_le_averageBound_of_not_mem_allHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_averageBudget_tendsto_zero_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_anti_delta","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_isBigO_log_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.natCast_log2_add_one_isBigO_log_succ","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedScalarHighProbabilityPseudoRegretBound_isBigO_sqrt_mul_log","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalAsymptoticHighProbabilityPseudoRegretAllHorizonViolationSet_eq_scalar","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_asymptoticRate_nonneg_and_allHorizon_tail_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.measurable_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAlgorithm_policy_apply","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_pow_sub_one","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_pow_ae_eq_forcedAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.isPowerOfTwoForcedIndex","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.isPowerOfTwoForcedIndex_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedIndexSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.mem_powerOfTwoForcedIndexSet_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedIndexSet_zero","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.zero_mem_powerOfTwoForcedIndexSet_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"declaration:BanditRLProof.OFUL.card_powerOfTwoForcedIndexSet_le_log2_add_one","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_of_not_forced","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.finiteHistoryPowerOfTwoForcedTelescopingScalarRidgeAction_eq_forcedAction","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_finiteHistoryTelescopingScalarRidgeOptimisticAction_of_not_powerOfTwoForced","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalHistoryTrajectory_action_succ_ae_eq_forcedAction_of_powerOfTwoForced","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedSuccessorPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.powerOfTwoOptimisticSuccessorPseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_eq_initial_add_powerOfTwoForced_add_optimistic","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedSuccessorPseudoRegret_ae_eq_forcedActionCharge","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"declaration:BanditRLProof.OFUL.canonicalStandardHighProbabilityPseudoRegret_ae_eq_initial_add_powerOfTwoForcedAction_add_optimistic","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_card_mul","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedActionSuccessorPseudoRegret_le_log2_add_one_mul_two_mul_parameterFeatureBound","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegretScalarAllHorizonViolationSet_subset","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"declaration:BanditRLProof.OFUL.powerOfTwoForcedCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_scalarAllHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.IntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.abs_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_randomHorizonEnvelope","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.measurable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_stoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.measurable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_stoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integrable_standardScalarAllRoundGapEnvelope_at_stoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integrable_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_of_integrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.measurableSet_telescopingCanonicalExplicitHighProbabilityPseudoRegretStoppedViolationSet_of_stoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_integral_badIndicator_randomHorizonEnvelope_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.standardScalarLogDetBudget_le_rounds_mul_div","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretQuadraticCoefficient","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretQuadraticCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityRegretLogBudget_le_rounds_sq_mul","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.telescopingHighProbabilityPseudoRegretBound_le_rounds_sq_mul_coefficient","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.integrable_stoppedValue_telescopingHighProbabilityPseudoRegretBound_of_squareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime_automaticBudgetIntegrability","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","target":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","target":"declaration:BanditRLProof.OFUL.stoppingTimeRoundSecondMoment_nonneg","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.SquareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.SquareIntegrableFiniteStoppingTime.toIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.integral_indicator_le_sqrt_secondMoment_mul_sqrt_real_measure","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.integral_badIndicator_standardScalarAllRoundGapEnvelope_at_stoppingTime_le","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_integral_stoppedBudget_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_telescopingHighProbabilityPseudoRegretBound_le_quadraticCoefficient_mul_roundSecondMoment_of_squareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_roundSecondMoment_add_initialGap_mul_sqrt_roundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.index_le_spent_of_unitGrowth","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.budgetExhaustionTime_le_budget_of_unitGrowth","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.squareIntegrableFiniteStoppingTime_budgetExhaustionTime_of_unitGrowth","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.stoppingTimeRoundSecondMoment_budgetExhaustionTime_le_of_unitGrowth","relation":"contains"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"declaration:BanditRLProof.Budget.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_budgetRoundsSq_add_initialGap_mul_budgetRounds_mul_sqrt_delta_and_stoppedViolation_measure_le_of_budgetExhaustionTime_of_unitGrowth","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.clippedConfidenceWidth","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.clippedConfidenceWidth_eq_min_one_confidenceWidth","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.sum_range_clippedConfidenceWidth_le_sqrt_mul_sqrt_log","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.sum_range_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.sum_range_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.sum_range_selectedAction_min_one_confidenceWidth_le_sqrt_mul_sqrt_log","relation":"contains"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"declaration:BanditRLProof.OFUL.sum_range_selectedAction_confidenceWidth_le_sqrt_mul_sqrt_log_of_width_le_one","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.integrable_exp_mul_predictable_mul_compensated_of_abs_le","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"declaration:ProbabilityTheory.HasCondSubgaussianMGF.predictable_mul_compensated_hasCondMGFUpperBoundAt_of_abs_le","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"declaration:BanditRLProof.OFUL.fixedDirectionCompensatedScore_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.constantSquaredVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.finiteHorizonFeatureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.finiteHorizonVarianceGram_constantSquared_eq_smul_featureGram","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.finiteHorizonFeatureGram_posSemidef","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.sq_smul_posDef","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.det_sq_smul_div_det_sq_smul","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.dotProduct_sq_smul_inv_mulVec","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.measure_inv_le_finiteHorizonDetRatioInvGramExponential_le","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.measure_invOfReal_le_finiteHorizonDetRatioInvGramExponential_le","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.invOfReal_le_gaussianDetRatioExponential_of_invGramQuadratic_gt","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizon_invGramQuadratic_gt_two_log_detRatio_div_le","relation":"contains"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizon_selfNormalizedQuadratic_gt_two_mul_sq_mul_log_detRatio_div_le","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.scalarRidgeConfidenceFailureAt","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHorizonScheduledScalarRidgeConfidenceFailureSet","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.not_mem_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_iff","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizonScheduledScalarRidgeConfidenceFailureSet_le_sum","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.finiteHorizonUniformScalarRidgeConfidenceFailureSet","relation":"contains"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizonUniformScalarRidgeConfidenceFailureSet_le","relation":"contains"},{"source":"module:BanditRLProof.OpenProblems","target":"declaration:BanditRLProof.ProblemArea","relation":"contains"},{"source":"module:BanditRLProof.OpenProblems","target":"declaration:BanditRLProof.OpenProblem","relation":"contains"},{"source":"module:BanditRLProof.OpenProblems","target":"declaration:BanditRLProof.seedOpenProblems","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.MeasurablePolicy","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_action_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_action_mem_filtration_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_action_mem_historyFiltration_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.generatedActionTrace","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_generatedActionTrace_eval_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.generatedActionTraceSucc","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.generatedActionTraceSucc_zero","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.generatedActionTraceSucc_succ","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.generatedActionTraceSucc_succ_eq","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_generatedActionTraceSucc_eval_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_generatedActionTraceSucc_succ_mem_filtration_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_generatedActionTrace_eval_mem_filtration_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"declaration:BanditRLProof.Policy.measurable_generatedActionTrace_eval_mem_historyFiltration_of_measurable_state","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.MarkovPosteriorKernel","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.ofKernel","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.ofMeasureSelector","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.ofMeasureSelector_apply","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.ofCountableHistorySelector","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.ofCountableHistorySelector_apply","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.measurable_kernel","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.measurable_apply_of_measurable_history","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.measurable_eventProbability_of_measurable_history","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.isProbabilityMeasure_apply","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.apply_univ","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.canonicalJointMeasure","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_fst_snd","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.BayesianPosteriorSurface","relation":"contains"},{"source":"module:BanditRLProof.PosteriorKernel","target":"declaration:BanditRLProof.PosteriorKernel.BayesianPosteriorSurface.posterior_isProbabilityMeasure_apply","relation":"contains"},{"source":"module:BanditRLProof.PowerCutoffNormalization","target":"declaration:BanditRLProof.PowerTailIntegral.source_cutoff_normalization","relation":"contains"},{"source":"module:BanditRLProof.PowerTailIntegral","target":"declaration:BanditRLProof.PowerTailIntegral.integral_power_tail_le","relation":"contains"},{"source":"module:BanditRLProof.PowerTailIntegral","target":"declaration:BanditRLProof.PowerTailIntegral.cutoff_balance","relation":"contains"},{"source":"module:BanditRLProof.PowerTailIntegral","target":"declaration:BanditRLProof.PowerTailIntegral.cutoff_objective","relation":"contains"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"declaration:BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le","relation":"contains"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"declaration:BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le_of_uniform","relation":"contains"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"declaration:BanditRLProof.ProbabilityUnionBound.measure_iUnion_fintype_le_sum","relation":"contains"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"declaration:BanditRLProof.finset_sum_pullCount_eq_time","relation":"contains"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"declaration:BanditRLProof.finset_sum_comp_pullCount","relation":"contains"},{"source":"module:BanditRLProof.PullCountReindex","target":"declaration:BanditRLProof.sum_selected_pullCount","relation":"contains"},{"source":"module:BanditRLProof.PullCountReindex","target":"declaration:BanditRLProof.pullCount_le_one_add_eventCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.rawCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.policyMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.measurable_rawCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation_eq_rawCount_sub_policyMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.integral_rawCount_iidEpisodeBatchMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateKernelMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_coordinateKernelMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateKernelMean_eq_policyMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinatePrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_coordinatePrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_rawCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_rawCount_eq_batchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_zero_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_cumulativeCoordinateDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_cumulativeCoordinateDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_adaptiveCumulativeCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountIndex_nonempty","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCountLocalDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveCumulativeCountBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation_abs_lt_of_not_mem_badEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateMeanAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateIncrement_eq_rawCount_sub_meanAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateDeviation_eq_rawCount_sub_mean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount_visit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateRawCount_transition","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateMeanAt_transition_eq_visit_mul_transition","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateMean_transition_eq_visit_mul_transition","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeEmpiricalTransitionMass_abs_sub_transition_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeTransitionCoordinateRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.AdaptiveCumulativeCountMartingaleCover","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.coordinateConfidence_of_not_mem_adaptiveCumulativeCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveCumulativeCoordinateConfidenceContract_of_martingale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveCumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeCountMartingale_optimism_and_explicitRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor_eq_rate_mul_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat_eq_rate_pow_mul_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathVisitLower_eq_rate_pow_mul_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor.scale_explorationRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRounds","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRecommendedExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageExploratoryBehaviorExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScale_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRate_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRounds_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationVisitFloor_mul_rounds","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationUniformVisitFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedScheduledAverageEnvelope_decayingExploration_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationBehaviorCharge_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageExploratoryBehaviorBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationDeltaAndAverageExploratoryBehaviorBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationAverageExploratoryBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.batchReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationReturnRadiusEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationScheduledEpisodes_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationReturnRadiusEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationNormalizedReturnRadius_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationAverageRealizedBehaviorRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationRealizedFailureAndRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.linearDecay","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeTransitionCountSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeTransitionCountSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeTransitionCountSummary_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.cumulativeTransitionCountSummaryAt_visitCount_le_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.radius_cumulativeVisitCount_succ_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan_upperValueRemaining_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPlan_selectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.countRadiusOptimisticPolicyTable_toMarkovPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticPlanAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.successorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurable_successorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeCoordinateConfidenceContract","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeCoordinateConfidence_optimism_and_recommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeCoordinateConfidenceContract.trajectoryMeasure_optimism_and_explicitRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedCumulativeRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedAverageRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageAbsoluteRealizedBehaviorConsistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.DecayingExplorationEpisodewiseWindowSpace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseWindowSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseWindowMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_badEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.abs_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.abs_cumulativeReward_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_optimalInitialExpectedReturn_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successor_rewardConsistent_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_abs_realizedSuccessorAverageRegret_le_two_mul_horizon_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseCommonMeasure_abs_realizedBehaviorRegretProcess_le_two_mul_horizon_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.integrable_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.measurableSet_decayingExplorationEpisodewiseCommonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegret_le_bound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationRealizedFailureBudget_toReal_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.memLp_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationEpisodewiseRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp_coeFn_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseRealizedBehaviorRegretLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_decayingExplorationEpisodewiseCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_episode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeRowOfTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_episode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_episodeReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_episodeReturn_iidEpisodeBatchMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturn_centered_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodewiseBatchReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodewiseBatchReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.totalReturn_centered_episodewise_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_episodewise_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseCumulativeSuccessorReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseCumulativeSuccessorReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseBatchReturnVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_episodewiseCumulativeSuccessorReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseSuccessorReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_episodewiseSuccessorReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_episodewiseSuccessorReturnDeviationBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_episodewise_transport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_le_normalized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.episodewiseNormalizedSuccessorReturnConfidenceRadius_le_decayingEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseNormalizedReturnRadius_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.decayingExplorationEpisodewiseRealizedFailureAndRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.decayingExplorationEpisodewiseAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_decayingExplorationEpisodewiseAverageRealizedBehaviorConsistency_allWindows","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionPMF_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.integral_exploratoryActionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.selected_sub_integral_exploratoryActionPMF_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue_sub_le_const","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy_bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy_valueRemaining_sub_exploratoryPolicy_valueRemaining_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_expectedRegret_le_toMarkovPolicy_expectedRegret_add_charge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryBehaviorRegretCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticExploratoryBehaviorExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageExploratoryBehaviorExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageExploratoryBehaviorBound_tendsto_charge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_policyAt_succ_eq_cumulativeOptimisticExploratoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.vanishingDeltaScheduledAverageExploratoryBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardSumSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AggregateVisitCountSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateVisitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.measurable_aggregateVisitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeRewardSumSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_rewardSum_coordinate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeRewardSumSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.cumulativeEmpiricalModelState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix.measurable_cumulativeEmpiricalModelState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeEmpiricalModelStateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_transitionCount_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_rewardSum_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalModelStateAt_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeAggregateVisitCountAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateVisitCountAt_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState.empiricalReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalModelState.empiricalReward_of_visitCount_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalSteps","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.confidenceNumerator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.one_le_confidenceNumerator_div","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor_eq_paper","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.logFactor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.scale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.scale_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.countRadius_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_successorPolicy_eq_optimisticPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_policyAt_succ_eq_optimisticPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.source_policyAt_succ_eq_modelState_optimisticPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodes_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledEpisodeThreshold_lt_episodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtEpisodeThreshold_lt_scheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_lt_scheduledEpisodeVisitMass","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_le_envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScheduledAverageBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_scheduledAverageRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeExploratoryEpisodeCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtAverageRecommendedExpectedRegretBound_eq_totalEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeEmpiricalOptimisticAverageRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedAverageRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt_radius_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.inverseSqrt_radius_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt_radius_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountRadius.cappedInverseSqrt_radius_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativePathVisitExpectedFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativePathVisitLowerMargin","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtRadiusEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.CumulativeInverseSqrtPathCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_coordinateMeanAt_visit_ge_pathFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_cumulativeCoordinateMean_visit_ge_pathFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_cumulativePathVisitLowerMargin_lt_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_adaptiveCumulativeCountMartingaleCover_of_pathSupport_inverseSqrtCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.adaptiveCumulativeEmpiricalOptimisticPlanAt_selectedRadiusRemaining_le_inverseSqrtEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_explicitRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCoverCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeCoordinateConfidenceRadius_sq_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtLogFactor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.half_cumulativePathVisitExpectedFloor_lt_lowerMargin_of_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtPathCalibration_of_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtEnvelopeSumBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_cumulativeInverseSqrtRadiusEnvelope_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtRecommendedExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_horizon_mul_two_cumulativeInverseSqrtRadiusEnvelope_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_closedFormRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageRecommendedExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingAverageConfidenceDelta_ennreal_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_le_envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaScheduledAverageBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.vanishingDeltaAndScheduledAverageBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.vanishingDeltaScheduledAverageRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_vanishingDeltaScheduledAverageRecommendedExpectedRegret_allWindows","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtCountRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtScale_cover","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold_le_normalized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeInverseSqrtCalibrationEpisodeThreshold_lt_of_normalizedThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtPathCalibration_of_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.normalizedCumulativeInverseSqrtRecommendedExpectedRegretBound_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_normalizedRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.traceStateAtFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.traceStateAtFrom_tail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.cumulativeRewardFrom_eq_sum_traceReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.episodeReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.totalReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_episodeReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_totalReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_episodeReturn_le_horizon_of_rewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_totalReturn_le_of_rewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeReturn_episodeBatchOfTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.totalReturn_episodeBatchOfTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.batchReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeReward_eval_iidTrajectoryFamilyMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_totalReturn_iidEpisodeBatchMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.totalReturn_centered_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnKernelMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_successorReturnKernelMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnKernelMean_eq_selectedPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_successorReturnPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_totalReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_totalReturn_eq_batchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchReturnVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.optimalInitialExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_cumulativeSuccessorReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_successorReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorReturnIncrement_succ_eq_totalReturn_sub_selectedPolicyMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.cumulativeSuccessorReturnDeviation_eq_fin_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successorReturnDeviationBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_cumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.measurable_prod_of_countable_left","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:MeasureTheory.Measure.compProd_map_dependent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_historyBatchStatistic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_historyBatchStatistic_eq_batchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanWeightCap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation_one_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInflation_pow_le_cap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanChargeWeight_mul_inflation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanWeight_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedSuccessorGapFeatureOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedSuccessorGapFeatureOfSummaries_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanStatisticOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanStatisticOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanHistoryBatchStatistic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanHistoryBatchStatistic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessorBellmanInnovation_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanInnovationOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanInnovationOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBellmanInnovationOfHistoryBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_successorBellmanInnovationOfHistoryBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_recurrentBellmanInnovationPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovationVarianceProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovation_compensated_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentBellmanInnovation_zero_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.trajectoryMeasure_recurrentBellmanInnovation_sum_ge_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateTransitionCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.measurable_aggregateTransitionCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.sum_aggregateTransitionCount_eq_aggregateVisitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF_of_aggregateVisitCount_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionPMF_apply_of_aggregateVisitCount_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel_isMarkov","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.measurable_adaptiveCumulativeAggregateTransitionCountAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"declaration:BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryStateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.trajectoryMeasure_action_eq_table_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.iidEpisodeBatchMeasure_successorBatchAligned_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_successorBatchAligned_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_all_successorBatchAligned_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeDeterministicGapInnovationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeDeterministicGapInnovationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_deterministicGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.abs_sampledCumulativeDeterministicGapInnovationFrom_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeTransitionGapInnovationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeTransitionGapInnovationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.abs_sampledCumulativeTransitionGapInnovationFrom_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.reconstructedInitialState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.reconstructedStepTrace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_reconstructedInitialState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_reconstructedStepTrace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeDeterministicGapInnovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_cumulativeDeterministicGapInnovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeDeterministicGapInnovation_episodeBatchOfTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeTransitionGapInnovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_cumulativeTransitionGapInnovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.cumulativeTransitionGapInnovation_episodeBatchOfTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchMeasure_one_cumulativeDeterministicGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_cumulativeTransitionGapInnovation_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionCoordinateVariance","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionCoordinateVariance_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.integral_sq_indicator_sub_transitionMass","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionResidualHead_variance_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionResidual_variance_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionResidual_variance_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionResidual_variance_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_succ_variance_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_zero_variance_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le_variance","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_generatedBatchAggregateVisitIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.SuccessorBatchAligned","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorBatchAligned_step","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedBatchAggregateVisitIncrement_eq_indicatorSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedBatchAggregateVisitIncrement_eq_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairLocalCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairLocalCharge_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_eq_generatedPairLocalCharge_of_aligned","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_successorLocalCharge_eq_sum_pairIncrement_mul_charge_of_aligned","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.remainingGlobalStage","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_eq_of_remaining_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorStageLocalCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeFrom_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge_eq_successorWeightedChargeSum_of_aligned","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge_le_sum_pairIncrement_mul_charge_of_aligned","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPairChargeTerm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_range_succ_shift_le_sum_range","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_successorCanonicalWeightedCharge_le_totalGeneratedPairCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.min_add_le_add_min","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.reciprocalChargeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedPairCharge_le_accounting","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_generatedPairCharge_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.one_le_logFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_generatedBatchAggregateVisitIncrement_le_totalSteps","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.log_max_batchedPrefixCount_le_logFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_batchedPrefixCount_generatedBatchAggregateVisitIncrement_eq_totalSteps","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_sqrt_batchedPrefixCount_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.QTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.initialQTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeSummaryOfSequence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_le_selected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_clippedPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentQTableOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentPolicyTableOfSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.prefixTransitionSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_prefixTransitionSummaries","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSuccessorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_recurrentSuccessorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentInitialTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinCoordinateThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_abs_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinTilt_exponent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.aggregateEmpiricalTransitionKernel_real_singleton_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bernsteinCoordinateThreshold_mul_probability_div_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.abs_sum_weight_mul_massError_le_transitionValue_div_thirtyTwo_add","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransitionMass_sub_lt_bernstein","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransition_integral_sub_le_transitionValue_div_thirtyTwo_add","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batchedPrefixCount_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_range_forwardDifference","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_low_batchedIncrement_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batched_ratio_le_two_log_increment","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_high_batchedRatio_le_log","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.batched_inverseSqrt_step","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedInvSqrt_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedMinReciprocal_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_nonneg_of_reward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedQTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_eq_generatedPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodeInitialState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.generatedEpisodePseudoRegret_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedEpisodeInitialState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_generatedEpisodePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_cumulativeEpisodePseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_eq_of_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_nonneg_of_reward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.decisionStageRemaining_succ_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_eq_of_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_succ_eq_selectedQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalValueAt_le_clippedValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedValueRemaining_sub_policyValueRemaining_le_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyGapRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedPolicyGapRemaining_le_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.selectedUpperTransitionModelError_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.selectedPolicyGap_le_of_actual_count","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationThreshold_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.bellmanInnovationTilt_exponent_le_neg_two_logFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_optimalValue_sub_eq_two_mul_horizon_mul_integral_probe_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_empiricalTransition_optimalValue_sub_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalQAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.QDominatesOptimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.optimalQAt_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.initialQTable_dominatesOptimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.HasOptimalTailConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_dominatesOptimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.adaptiveCumulativeAggregateVisitCountAt_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.hasOptimalTailConfidence_of_not_mem_simultaneousTransitionFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.card_bernsteinCoordinateIndex","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.card_optimalTailIndex","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.exp_neg_two_mul_logFactor_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.exp_neg_two_mul_logFactor_sq_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.ofReal_card_mul_two_mul_exp_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.simultaneousTransitionTailBudget_le_fifth","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeSummaryOfSequence_prefixTransitionSummaries_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessor_selectedQ_ge_optimalQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPreviousQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyGapRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorPolicyGapRemaining_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorLocalCharge_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_succ_eq_chargeWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.normalizedBellmanRemainingWeight_eq_stageWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorCanonicalWeightedCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedChargeSum_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.successorWeightedVisitedGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.successorPolicyGapRemaining_le_inflation_mul_transition_add_charge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.clippedSuccessorGapFeatureOfSummaries_eq_weight_mul_gap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.normalizedGap_le_weightedChargeFrom_add_deterministicInnovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.successorPolicyGapRemaining_le_cap_mul_charge_add_innovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedEpisodePseudoRegret_le_successorPolicyGapRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentBellmanInnovationProcess_succ_eq_canonical","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitHead","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionResidualFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionVisitFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualHead_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionResidualHead_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionResidual_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom_eq_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionResidualFrom_eq_aggregateCounts","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionVisitFrom_eq_aggregateVisitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionResidual_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateTransitionResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateVisitReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateTransitionResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateVisitReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateVisitCount_le_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_aggregateTransitionResidual_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionResidual_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.batchPrefixFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.hasCondMGFUpperBoundAt_of_condExpKernel_map_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.hasCondMGFUpperBoundAt_congr_measurableSpace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidualPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateVisitPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionResidualPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateVisitPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidualIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateVisitIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionResidualIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateVisitIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_compensated_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_batchStatistic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.condExpKernel_map_batchStatistic_eq_batchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_exp_mul_aggregateTransitionResidual_compensatedIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_succ_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionResidual_zero_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateTransitionResidualIncrement_eq_prefixAggregateResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateVisitIncrement_eq_prefixAggregateVisitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_aggregateTransitionResidualSum_ge_inter_visitSum_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionResidualSum_ge_inter_visitSum_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.BernsteinCoordinateIndex","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.OptimalTailIndex","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.bernsteinCoordinateFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.optimalTailFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.simultaneousTransitionFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_bernsteinCoordinateFailureEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.functionalTilt_exponent_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_optimalTailFailureEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_simultaneousTransitionFailureEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_coordinateResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.abs_optimalTailResidual_lt_of_not_mem_simultaneousTransitionFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.alignmentFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.canonicalFailureEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_alignmentFailureEvent_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_canonicalFailureEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_shift_recurrentBellmanInnovation_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sqrt_totalSteps_mul_card_eq_paper","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.cumulativeEpisodePseudoRegret_le_canonicalRegretBound_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalTailProbe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalTailProbe_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionFunctionalResidualFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue_eq_sum_measureReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.actionStateKernel_transitionFunctionalResidualHead_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryKernelRemaining_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryMeasure_transitionFunctionalResidual_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionFunctionalResidualFrom_eq_aggregateFunctionalResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.aggregateTransitionFunctionalResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_aggregateTransitionFunctionalResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.abs_aggregateTransitionFunctionalResidual_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.integral_aggregateEmpiricalTransitionKernel_eq_sum_div","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_one_aggregateTransitionFunctionalResidual_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidualPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidualIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionFunctionalResidualPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurable_aggregateTransitionFunctionalResidualIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_compensated_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.integrable_exp_mul_aggregateTransitionFunctionalResidual_compensatedIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_succ_compensated_hasCondMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_zero_compensated_hasMGFUpperBoundAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.sum_aggregateTransitionFunctionalResidualIncrement_eq_prefixResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.aggregateTransitionFunctionalResidual_eq_count_mul_transitionValue_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measure_abs_aggregateTransitionFunctionalResidualSum_ge_inter_visitSum_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_empiricalTransitionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary_empiricalTransitionKernel_real_singleton","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPlan_upperValueRemaining_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.EmpiricalOptimisticCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalOptimisticPlanCoordinateConfidence_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.explorationRate_mul_inv_card_le_exploratoryActionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDEpisodeBatchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDEpisodeBatchKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticOccupancyRadiusSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_successorPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurableSet_selectedExploratorySimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_measurableSet_successorSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_adaptiveSimultaneousCountConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_policyAt_succ_eq_adaptiveEmpiricalOptimisticPlanAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.policyAt_batch_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_coordinateConfidenceAt_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_optimism_and_recommendedExpectedRegret_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_const","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt_selectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveEmpiricalOptimisticOccupancyRadiusSum_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_explicitRecommendedExpectedRegret_of_pathSupport_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCountSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_transitionCountSummary","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.empiricalTransitionKernel_isMarkov","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPlan","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.toMarkovPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.optimisticPolicyTable_toMarkovPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalOptimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalOptimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.iidEpisodeBatchKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.latestBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurable_latestBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.successorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurable_successorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_successorPolicy_eq_optimisticPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.measurableSet_selectedSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_measurableSet_successorSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_trajectoryMeasure_condDistrib_eq_empiricalOptimisticPolicyBatchLaw","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.source_trajectoryMeasure_adaptiveSimultaneousCountConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_map_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_prefix_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_condDistrib_eq_iidEpisodeBatchMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.roundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_roundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_initialBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_roundBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_finiteHorizonBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_conditionalLaw_and_finiteHorizonBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.adaptiveSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.policyAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.initialSimultaneousCountBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.successorSimultaneousCountBadEvent_fiber_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.measurableSet_adaptiveSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEpisodeBatchSource.trajectoryMeasure_adaptiveSimultaneousCountConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedCumulativeRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageAbsoluteRealizedBehaviorConsistency_of_standardBorel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.DecayingExplorationStochasticWindowSpace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticWindowSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticWindowMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.measurable_decayingExplorationStochasticRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_badEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.abs_decayingExplorationStochasticRealizedBehaviorRegretProcess_le_of_not_mem_badEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_marginals_and_realizedBehaviorRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_le_two_mul_horizon_of_rewardBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.one_half_le_log_two_div_vanishingAverageConfidenceDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationNormalizedSuccessorGlobalReturnMGFFirstMomentBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCumulativeReturnDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCumulativeReturnDeviationProcess_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.integrable_decayingExplorationStochasticRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.measurableSet_decayingExplorationStochasticCommonCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticCommonMeasure_countBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegret_le_bound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticExpectedAbsoluteRealizedBehaviorRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_integrable_expectedAbsoluteRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.memLp_one_decayingExplorationStochasticRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.eLpNorm_one_decayingExplorationStochasticRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp_coeFn_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationStochasticRealizedBehaviorRegretLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_decayingExplorationStochasticCommonMeasure_memLp_eLpNorm_L1_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalReturnDeviationPerEpisodeVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidGlobalSampledCumulativeReturnDeviationVarianceProxy_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalReturnDeviationPerEpisodeVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticReturnRadiusEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_le_decayingEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticReturnRadiusEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationNormalizedSuccessorGlobalReturnRadius_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticAverageRealizedBehaviorRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationStochasticRealizedFailureAndRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_initialPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_initialBatch_map_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_batchKernel_map_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_condDistrib_projectedNext","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.projectedAdaptiveCumulativeCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_projected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_projected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.decayingExplorationAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_optimism_and_decayingExplorationAverageRealizedBehaviorConsistency_of_standardBorel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.ProbabilityTheory.measure_compProd_map_prodMap_of_map_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchPrefix_frestrictLe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDStochasticEpisodeBatchKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryIIDStochasticEpisodeBatchKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.measurable_selectedExploratorySampledReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_initialPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_initialBatch_map_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_batchKernel_map_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_projectedPrefix_next_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_condDistrib_projectedNext","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_map_knownRewardEpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAdaptiveSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_recommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport_two_delta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAverageExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedExploratoryBehaviorExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.projectedAverageExploratoryBehaviorExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.measurable_selectedExploratoryGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_projected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_projected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedAllCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_map_fst","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_initialPolicyValueDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.initialPolicyValueDeviationSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_initialPolicyValueDeviationSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidInitialPolicyValueDeviationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalSampledCumulativeReturnDeviationSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.globalSampledCumulativeReturnDeviationSum_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_globalSampledCumulativeReturnDeviationSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidGlobalSampledCumulativeReturnDeviationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_initialPolicyValueDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.initialPolicyValueDeviationAtEpisode_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_initialPolicyValueDeviationSum_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_globalSampledCumulativeReturnDeviationSum_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.GlobalReturnMeasurability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorGlobalReturnDeviationAt_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorPolicyAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.optimalInitialExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_cumulativeSuccessorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_successorGlobalReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_expected_to_realized_successor_average_regret_transport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.policyAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.roundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_initialBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.adaptiveAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.multiBatchLocalDelta_pos_of_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.multiBatchLocalDelta_le_one_of_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_initialAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent_fiber_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.ofReal_multiBatchLocalDelta_add","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selectedExploratoryStochasticAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.sampledEmpiricalOptimisticPolicyTable_toMarkovPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorPolicy_frestrictLe_eq_sampledPlanExploratoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_cumulativeSuccessorPolicyExpectedRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSuccessorExploratoryBehaviorExpectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_optimism_and_cumulativeRecommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeRecommendedExpectedRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt_selectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticPlanAt_occupancySelectedRadiusRemaining_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticOccupancyRadiusSum_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticExplicitBudgetAverageBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_explicitBudgetAverageBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_explicitBudgetRealizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorExpectedCumulativeRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_successorExpectedAverageRegret_eq_sampledPlanExploratoryBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyAt_expectedRegret_le_rateAt_of_not_mem_modelRoundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_behaviorExpected_and_realizedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.summable_exp_neg_sqrt_natCast","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.natCast_le_successorEpisodeMass","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.summable_exp_neg_sqrt_successorEpisodeMass","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledCausalVanishingReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalVanishingReturnFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalVanishingReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_vanishingReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalModelRoundBadEvent_measure_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalVanishingReturnBadEvent_measure_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_selfConsistentScheduledCausalModel_and_returnBadEvents","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_abs_selfConsistentScheduledNaturalCausalRealizedRegret_le_envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_realizedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_eventually_modelOptimistic_and_realizedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.integral_le_add_const_mul_measureReal_of_le_on_compl","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalCoordinateModelFailureBudget_toReal_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_le_rateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero_of_explicit_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpectedRegret_explicitIntegratedRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_eq_sum_expectedAbsolute","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_eq_natWeightedAverage","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.measurable_exploratoryPolicy_expectedRegret_comp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_successorPolicyExpectedRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalSuccessorPolicyExpectedRegretProcess_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_coeFn_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBehaviorExpectedRegretLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpectedRegret_memLp_eLpNorm_L1_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_behaviorExpected_and_realizedRegret_L1_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedCumulativeRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.abs_realizedSuccessorAverageRegret_le_of_expected_le_of_deviation_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_eq_inv_pow","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledLocalDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalCoordinateModelFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledCausalCoordinateModelFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelRoundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalModelRoundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledCausalTailModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_modelRoundBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_tailModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_tail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_coordinateConfidence_of_not_mem_modelRoundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalLocalPlanningBound_le_rateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_weightedExpectedSuccessorAverageRegret_le_burninEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnDelta_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.normalizedSuccessorGlobalReturnConfidenceRadius_vanishingDelta_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalVanishingReturnRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelReturnFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninRealizedRegretRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_tail_optimism_and_absoluteRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_realizedRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.integrable_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnMGFFirstMomentBound_le_confidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalTailModelFailureBudget_toReal_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalNormalizedReturnMGFFirstMomentBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedRegretRateEnvelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalBurninExpectedAbsoluteRegretL1Envelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalRealizedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_le_burninL1Envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRealizedRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRealizedRegretProcess_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp_coeFn_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedRegretLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalRealizedRegret_memLp_eLpNorm_L1_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.natWeightedAverage","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.tendsto_natWeightedAverage_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_eq_sum_range","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_tendsto_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.fixedHalfSuccessorGlobalReturnRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.normalizedSuccessorGlobalReturnConfidenceRadius_half_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.fixedHalfSuccessorGlobalReturnRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalPlanningRateAt_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalWeightedPlanningRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningAverageBound_le_rateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalWeightedPlanningRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalReturnRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalReturnRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret_le_explicitRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.policyAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.roundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_initialBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.adaptiveAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeModelFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_initialAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_adaptiveAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.initialAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorAllCoordinateEmpiricalModelBadEvent_fiber_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.policyAt_batch_not_mem_allCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_measurableSet_successorAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_trajectoryMeasure_adaptiveAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorPolicyAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorEpisodeMass_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorWeightedExpectedAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation_eq_fin_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorCumulativeRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.realizedSuccessorAverageRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_cumulativeSuccessorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurableSet_successorGlobalReturnDeviationBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_successorGlobalReturnDeviationBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_weightedExpected_to_realized_successor_average_regret_transport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel_occupancySelectedRadiusRemaining_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_successorPolicyAt_eq_sampledPlanExploratoryPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource_successorPolicyAt_expectedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningCumulativeBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorPlanningAverageBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_weightedExpectedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSuccessorReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelReturnFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_optimism_and_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorDeviationKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorDeviationKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorSampledReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorSampledReturnDeviationAt_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.GlobalReturnMeasurability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorGlobalReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorGlobalReturnDeviationAt_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_successorGlobalReturnPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.successorGlobalReturnIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.cumulativeSuccessorGlobalReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeSuccessorGlobalReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousStochasticEpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousStochasticEpisodeBatchPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_prefix_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_nextBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_prefix_projective","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousLatestBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_heterogeneousLatestBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousSuccessorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_heterogeneousSuccessorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.heterogeneousExploratorySource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_exactLaws_and_projective","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_absoluteRealizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.SelfConsistentScheduledStochasticWindowSpace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticWindowSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticWindowMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticRealizedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledStochasticRealizedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledStochasticCommonMeasure_badEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.abs_selfConsistentScheduledStochasticRealizedRegretProcess_le_of_not_mem_badEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_selfConsistentScheduledStochasticCommonMeasure_marginals_and_realizedRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudgetRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudgetRateEnvelope_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_le_rateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound_le_rateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound_le_rateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureRateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget_eq_rateEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureRateEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledExplicitRateEnvelopes_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduledExplicitRate_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_selfConsistentCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_cumulativeSuccessorExploratoryBehaviorExpectedRegret_of_pathSupport_selfConsistentCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimisticSelfConsistentBudgetAverageBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveStochasticSampledEmpiricalOptimistic_occupancyAndChargeAverage_eq_selfConsistentBudgetAverageBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_finiteRound_allCoordinateConfidence_optimism_and_selfConsistentBudgetRealizedSuccessorAverageRegret_of_pathSupport_selfConsistentCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentCountShrinkEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentRewardShrinkEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_scale_sq_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousRewardSumConfidenceRadius_lt_episodes_mul_visitFloor_div_two_scale_sq_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodes_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledEpisodeThreshold_lt_episodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.exploratoryPathCalibrationEpisodeThreshold_lt_selfConsistentScheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentCountShrinkEpisodeThreshold_lt_scheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentRewardShrinkEpisodeThreshold_lt_scheduledEpisodes","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduled_countMargin_and_halfContraction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledLocalDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledCountRadius_lt_mass_div_scale_sq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardSumRadius_lt_mass_div_two_scale_sq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContractionEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_lt_inv_scale_sq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_lt_envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_lt_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.decayingExplorationScale_sq_tendsto_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContractionEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRewardBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionContraction_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledTransitionBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledPlanningAverageRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledNormalizedSuccessorGlobalReturnRadius_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedSuccessorAverageRegretBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledRealizedFailureBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.selfConsistentScheduledFailureAndRealizedBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_selfConsistentScheduled_allCoordinateConfidence_optimism_and_realizedSuccessorAverageRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_rewardSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.stateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionStateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionStateKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_uncurry_of_forall_measurable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurable_empiricalTransitionValue","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalUpperValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticActionAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_stochasticAllCoordinateEmpiricalOptimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.sampledEmpiricalOptimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch.measurable_sampledEmpiricalOptimisticPolicyTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.latestBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_latestBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.successorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_successorTable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exploratorySource_batchKernel_eq_selectedPolicy_iidLaw","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:ProbabilityTheory.Kernel.retainedInputKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:ProbabilityTheory.Kernel.retainedInputKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:ProbabilityTheory.Kernel.condDistrib_pair_ae_eq_retainedInputKernel_of_pair_map_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:ProbabilityTheory.Kernel.condDistrib_dynamic_map_ae_eq_of_pair_map_eq_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatchPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.StochasticEpisodeBatchTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_map_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_prefix_compProd","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorDeviationKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorDeviationKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.successorSampledReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_successorSampledReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_condDistrib_successorSampledReturnDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.condExpKernel_map_successorSampledReturnDeviationAt_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationPrefixIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.measurable_sampledReturnDeviationIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_zero_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.sampledReturnDeviationIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviationVarianceProxy_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.cumulativeSampledReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_cumulativeSampledReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.abs_integral_sub_integral_le_sum_coordinateRadius_mul_envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.transitionError_le_radius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.toConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.CoordinateConfidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_state","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_action","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_reward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_nextState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.rewardSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.transitionCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.sum_transitionCount_eq_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionMass","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF_eq_pure_of_visitCount_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionPMF_apply_of_visitCount_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionMass_eq_div_of_visitCount_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_isMarkov","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalTransitionKernel_real_singleton","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.plan","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence.toCoordinateConfidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.FiniteBatchModel.Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.toProdEquiv","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.decisionStageRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.exactModelPlan","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.transitionValue","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.bellmanQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_transitionValue","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticQ_le_optimisticAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.upperValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.upperValueRemaining_eq_of_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_upperValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.exactModelPlan_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.trueBellmanQ_le_optimisticQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.certificate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.optimalValueRemaining_le_upperValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticActionAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.measurable_optimisticActionAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticActionAt_decisionStageRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.optimisticPolicy_bellman_eq_bellmanQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.selectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.selectedRadiusRemaining_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.policyBellmanResidual_le_two_selectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.expectedRegret_le_two_occupancySelectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.EstimatedModelPlan.Confidence.optimism_and_expectedRegret_le_two_occupancySelectedRadiusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathCalibrationDimensionFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.one_le_exploratoryPathCalibrationDimensionFactor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathCalibrationEpisodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius_lt_episodes_mul_visitFloor_div_dimensionFactor_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.episodeThreshold_countMargin_and_halfContraction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceTransitionBonusCover_of_pathSupport_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_episodeThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorTransitionCoordinateRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorTransitionCoordinateRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCountTransitionCoordinateRadius_le_uniformFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathUniformVisitFloor.exploratoryStateCountMargin","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.uniformVisitFloor_expectedCount_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.transitionBonusCover_rewardBound_of_uniformExpectedCountFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceTransitionBonusCover_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryStateAt_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeStepOfTrajectory_nextState_eq_trajectoryStateAt_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_le_stageStateProbability_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryPathSupport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLower","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLowerNat_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryPathStateLower_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPathStateLower_le_stageStateProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceStateReachability_of_pathSupport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_pathSupport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_pathSupport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.exploratoryActionProbabilityFloor_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.TransitionBonusCover","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionProbabilityFloor_le_exploratoryActionPMF_toReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryActionProbabilityFloor_le_actionKernel_real","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.stateLower_mul_exploratoryActionProbabilityFloor_le_stageVisitProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.stateLower_expectedCount_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.ExploratoryStateCountMargin","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalOptimisticCalibration_exploratoryPolicy_of_stateReachability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceStateReachability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.SourceTransitionBonusCover","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_sourceCalibration_of_stateReachability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_allCoordinateConfidence_optimism_and_recommendedExpectedRegret_of_stateReachability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.empiricalFiniteBatchValueEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.allCoordinateEmpiricalFiniteBatchModel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCountTransitionCoordinateRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedCount_sub_radius_lt_count_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalTransitionMass_abs_sub_transition_le_expectedCountRadius_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalReward_eq_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.AllCoordinateConfidence.upperValueRemaining_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.allCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_allCoordinate_finiteBatchModel_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_allCoordinate_optimism_and_expectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.visitIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.transitionIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_visitIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.measurable_transitionIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.visitIndicator_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeStep.transitionIndicator_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_eq_measureReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_eq_measureReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_le_stageVisitProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_visitIndicator_iidEpisodeBatchMeasure_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_transitionIndicator_iidEpisodeBatchMeasure_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredVisitIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredTransitionIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_visitIndicator_eq_cast_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_transitionIndicator_eq_cast_transitionCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_centeredVisitIndicator_eq_cast_visitCount_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.sum_centeredTransitionIndicator_eq_cast_transitionCount_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_cast_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_cast_transitionCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_visitCountDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_transitionCountDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidBernoulliVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredVisitIndicator_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.centeredTransitionIndicator_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_centeredVisitIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_centeredTransitionIndicator","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_visitCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_transitionCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_visitCount_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_transitionCount_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_visit_and_transition_count_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalTransitionMass_abs_sub_transition_lt_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_empiricalTransitionMass_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.count","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.expectedCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.zeroCountEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.VisitCoordinate.measurableSet_zeroCountEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.eligibleZeroVisitCountEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.measurableSet_eligibleZeroVisitCountEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.visitCoordinate_count_pos_of_not_mem_eligibleZeroVisitCountEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.visitCoordinate_count_pos_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.eligibleZeroVisitCountEvent_subset_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligibleZeroVisitCountEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_visit_count_positivity","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.RewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.measurableSet_rewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.rewardSum_eq_visitCount_mul_reward_of_rewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.EpisodeBatch.empiricalReward_eq_reward_of_rewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardSum_eq_visitCount_mul_reward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_empiricalReward_eq_reward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_rewardConsistent_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.empiricalReward_eq_and_transition_lt_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.episodeBatchOfTrajectories_empiricalReward_eq_and_transition_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_eligible_empiricalReward_exact_and_transition_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.multiBatchLocalDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchRoundBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_multiBatchSimultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_roundBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_multiBatchSimultaneousCountBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamilyMeasure_rewardConsistent_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchEmpiricalModelAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchCumulativeExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.multiBatchCumulativeSelectedRadiusOccupancy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.allCoordinateConfidenceFamily_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.confidenceFamily_optimism_and_cumulativeExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamily_allCoordinate_finiteBatchModel_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchFamily_allCoordinate_optimism_and_cumulativeExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.equivVisitSumTransition","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.countCoordinateCard","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.countCoordinateCard_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.simultaneousCountConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.measurable_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.badEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.CountCoordinate.measurableSet_badEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measure_countCoordinate_badEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurableSet_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.simultaneousCountDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_simultaneousCountBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.countCoordinate_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.visitCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.transitionCount_abs_deviation_lt_of_not_mem_simultaneousCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_simultaneous_count_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryStateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeStepOfTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryVisitContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryRewardContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.trajectoryTransitionContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_episodeStepOfTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_episodeBatchOfTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_visitCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_rewardSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.episodeBatchOfTrajectories_transitionCount","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryVisitContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryRewardContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_trajectoryTransitionContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidTrajectoryFamilyMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidTrajectoryFamilyMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatchMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeStepOfTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_iidEpisodeBatch_statistic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_episodeStatistic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryVisitContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryRewardContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iIndepFun_trajectoryTransitionContribution","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidEpisodeBatch_stepLaw_and_independence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_transitionValue","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_bellmanQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_naturalAllPrefixReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.log_two_div_naturalAllPrefixReturnDelta_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius_le_envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixAverageReturnConfidenceRadius_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.naturalAllPrefixReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_naturalAllPrefixReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalAllPrefixReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_naturalAllPrefixReturnBadEvent_measure_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_naturalAllPrefixReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageReturnDeviation_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnL1SummableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_explicitRounds_le_summableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_explicitRounds_le_summableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_explicitRounds_le_summableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageBehaviorRegretL1SummableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageReturnL1SummableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixAverageRealizedBehaviorRegretL1SummableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_le_summableEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixReciprocalThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixReciprocalThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_expectedAbsolute_div","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_explicitPolynomialPrefixAverageRealizedBehaviorRegretReciprocalDistanceViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_le_inverseSqrtEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnInverseSqrtEnvelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageReturnFirstMomentBound_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_le_L1Envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_coeFn_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_naturalAverageRealizedBehaviorRegret_allPrefix_L1_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_integrated","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativePlanningRate_le_logarithmic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_eq_fin_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledCausalModelRoundBadEvent_of_not_mem_prefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_planning_of_not_mem_modelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_logarithmic_of_not_mem_modelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretLogarithmicViolationSet_subset_modelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_cumulativeBehaviorExpectedRegretLogarithmicViolationSet_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.sum_range_one_div_natCast_add_two_sq_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.sum_range_one_div_natCast_add_two_pow_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.sum_range_one_div_natCast_add_three_le_one_add_log","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.one_le_one_add_log_natCast","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.tendsto_one_add_log_natCast_div_natCast_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalHighPowerRateCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalIntegratedBehaviorExpectedRegretRateAt_eq_threeTerm","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSquareRateCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExplorationHarmonicRateCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalHighPowerRateCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicRateCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_refined","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRefinedCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeIntegratedBehaviorExpectedRegretRate_le_logarithmic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicCumulativeIntegratedBehaviorExpectedRegretRate_isBigO","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_le_logarithmic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedCumulativeBehaviorRegret_isBigO_one_add_log","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_le_logarithmic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_isBigO_log_div_natCast","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegret_tendsto_zero_of_logarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cumulative_and_averageBehaviorExpectedRegret_explicitLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_rounds_mul_two_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretSecondMomentEnvelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_secondMomentEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitSecondMomentBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_le_explicitBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingExplicitExpectedRateBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_le_explicitExpectedRateBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegretSecondMoment_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.sum_range_one_div_natCast_add_two_pow_le_one_div_sixteen","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_eq_ofReal_two_mul_sum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelFailureBudget_le_one_eighth","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le_one_quarter","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.one_le_untopA_and_untopA_le_of_withTop_bounds","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedAverageRealizedBehaviorRegret_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppedRealizedAverageLogarithmicRate_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_positivePrefixWindow","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_positivePrefixAverageRealizedBehaviorRegretViolationWindow_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.finiteHorizonBadEvent_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalModelBadEvent_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixIndex_nonempty","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingEqualReturnShare_spec","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBoundedStoppingReturnBadEventWindow","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingReturnBadEventWindow_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingSingleModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPositivePrefixAverageRealizedBehaviorRegretViolationWindow_subset_boundedStoppingSingleModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_boundedStoppingSingleModelReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedStoppingTimeAverageRealizedBehaviorRegretViolationSet_subset_singleModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_window_offset_untopA_eq_of_withTop_bounds","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBoundedWindowStoppingL1Budget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_le_budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteBoundedWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalBoundedWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_boundedWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_of_first_hit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_base_of_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_eq_right_of_forall_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_before_gt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassagePrefixNat_le_threshold_of_lt_right","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_isStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_lower","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_upper","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_cappedDoubleLinearRawWindowFirstPassageStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.scheduleIndex_le_explicitHighProbabilityRounds_add","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_growingWindowGrid_offset_untopA_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.growingWindowGrid_stoppingPrefix_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_le_tail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Tail_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingL1Budget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_le_budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteGrowingWindowGridStoppingAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalGrowingWindowGridStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_growingWindowGridStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtFirstPassageThreshold_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_div_pow_three_eq_inverse_rpow_five_halves","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_div_pow_two_eq_inverse_rpow_three_halves","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_inverseSqrtFirstPassageThreshold_le_delayRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_eq_base","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassage_summableDelay_eventuallyImmediateStopping_and_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_tendsto_zero_of_memLp_one_of_eLpNorm_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegralDifference_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedGapAbs_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_expected_truncation_replacement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppingPrefixes_eq_base_of_not_mem_delayedSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_eq_unboundedHittingAfter","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifferenceLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretDifference_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_L1_truncation_equivalence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn_eq_optimal_sub_naturalAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageSampledReturn_naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSampledReturn_stronglyAdapted_naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_eq_optimal_sub_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSampledReturnProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_of_integrable_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnProcess_eq_optimal_sub_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_eq_optimal_sub_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_tendsto_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_tendsto_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedAverageSampledReturn_expected_optimality","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretExpectedGapAbs_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationExpectedGapAbs_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_eq_behaviorExpected_sub_returnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_expected_truncation_replacement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegret_and_returnDeviation_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedBehaviorExpectedRegretIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviation_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedReturnDeviationIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifference_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretDifferenceIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_eq_behaviorExpectedRegretDifference_sub_realizedBehaviorRegretDifference","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifference_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnDeviationDifferenceIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_truncation_equivalence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_sub_behaviorExpectedRegretIntegral_eq_neg_returnDeviationIntegral","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretPolicyValueExpectedGapAbs_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedRealizedBehaviorRegret_and_policyValue_expected_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.probReal_compl_gt_one_sub_of_measure_lt_ofReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_iff","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_probReal_gt_one_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_tailStart_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_deterministicTailHighProbability_optimality","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_eq_behaviorExpected_sub_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn_tendstoInMeasure_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn_tendstoInMeasure_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn_tendstoInMeasure_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSampledReturn_tendstoAlmostEverywhere_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageSuccessorPolicyExpectedReturn_tendstoAlmostEverywhere_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCapped_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedSampledPolicyExpectedReturnGap_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_inMeasure_and_almostSure_optimality","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_fun_neg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturnOptimalityErrorProcess_eq_neg_realized","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnOptimalityErrorProcess_eq_neg_behaviorExpected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledPolicyExpectedReturnGapProcess_eq_returnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledReturnOptimalityError_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledReturnOptimalityError_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSuccessorPolicyExpectedReturnOptimalityError_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_L1_optimality","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedSampledPolicyExpectedReturnGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedSampledPolicyExpectedReturnGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.not_mem_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_iff","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationProbability_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_simultaneousHighProbability_optimality","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorPolicyExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSuccessorPolicyExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageSuccessorPolicyExpectedReturn_eq_optimal_sub_expectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageSuccessorPolicyExpectedReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess_eq_optimal_sub_behaviorExpected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageSuccessorPolicyExpectedReturnProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_stoppedSuccessorPolicyExpectedReturn_of_integrable_behavior","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_stoppedSuccessorPolicyExpectedReturn_eq_optimal_sub_behavior","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageSampledReturn_sub_successorPolicyExpectedReturn_eq_returnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturn","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_eq_optimal_sub_behaviorExpected","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnIntegral_sub_successorPolicyExpectedReturnIntegral_eq_returnDeviationIntegral","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSampledReturnSuccessorPolicyExpectedReturnAbsGap_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageSuccessorPolicyExpectedReturnIntegral_tendsto_optimal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_lower","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base_of_process_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_untopA_unboundedHittingAfter_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_isStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_all_ne_top_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_eventually_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_eq_base","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.ae_tendsto_untopA_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingPrefix_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedProcess_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_aeFinite_eventuallyImmediateStopping_and_inMeasure_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeFourScale_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeFour","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzPolynomialAbsoluteFirstMomentBudget_isBigO_degreeFour","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_cauchySchwarzPolynomialBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_cauchySchwarzPolynomialBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeFour","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tendsto_integral_abs_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tendsto_abs_integral_restrict_of_uniformIntegrable_one_of_measure_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedEvent_expectedContribution_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedPositivePart_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretPositivePart_integrable_and_integral_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedPositivePart_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRateLinearCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_le_linearEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_spec","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart_le_explicitTailStart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitStoppingRoundSecondMomentENNRealBudgetAt_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_explicitTailStartBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_explicitTailStartDeterministicStoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretUniformSecondMomentEnvelope_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingFiberAbsoluteFirstMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentAbsoluteFirstMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDeterministicStoppingRoundSecondMomentAbsoluteFirstMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_sq_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_uniformSecondMomentEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingFiberBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_stoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_deterministicStoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityQuarticBlockWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_succ_sub_le_quarticBlockWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_le_inv_pow_six","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedInverseCubePairEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockShiftedInverseCubePairEnvelope_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedCoordinateModelFailureCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockShiftedCoordinateModelFailureCharge_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_tailModelFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_add_le_add_sum_Ico_quarticBlockWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.natCast_succ_le_checkpoint_add_tsum_quarticBlockWeight_of_delayed","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_untopA_add_one_of_eventually_quarticCheckpointTail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_subset_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret_succ_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalSuccessorAverageReturnDeviationIncrement_succ_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnVarianceProxyAt_succ_le_globalReturnDeviationPerEpisodeVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegretSecondMomentEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.memLp_two_naturalSuccessorBatchAverageRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.integral_sq_naturalSuccessorBatchAverageRealizedRegret_le_secondMomentEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.max_neg_natWeightedAverage_succ_le_abs_increment_div","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.hittingAfter_predecessor_gt_of_untopA_gt_base","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.max_neg_hittingAfter_untopA_le_base_abs_add_abs_stoppedValue_delayedReciprocalIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalWeight_sq_tsum_sqrt_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfter_stoppedNegativePart_le_baseAbsolute_add_delayedReciprocalSuccessorRegretAbsolute","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegret_integrable_and_integral_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedReciprocalSuccessorRegretExpectedAbsolute","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_delayedReciprocalSuccessorRegretExpectedAbsolute_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedNegativePart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedNegativePart_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_coeFn_ae_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretLp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_memLp_eLpNorm_Lp_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterDegreeEightScale_one_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointAsymptoticCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare_le_asymptoticCoefficient_mul_scale_pow_eight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAsymptoticCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget_isBigO_degreeEight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMoment_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppingRoundSecondMoment_le_polynomialBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_isBigO_degreeEight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterShiftedInverseSquareSeriesConstant_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialAbsoluteFirstMomentAsymptoticCoefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_le_asymptoticCoefficient_mul_degreeEightScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget_isBigO_degreeEight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretExpectedAbsolute_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_le_polynomialBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegretExpectedAbsolute_isBigO_degreeEight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialScaleCoefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitTailStart_succ_le_polynomialScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialCheckpointSquare","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitCheckpointSquare_le_polynomial","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentENNRealConstant_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterWeightedFailureSecondMomentConstant","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentENNRealBudget_toReal","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentENNRealBudget_le_polynomial","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterExplicitStoppingRoundSecondMomentBudget_le_polynomial","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_polynomialBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialStoppingRoundSecondMomentAbsoluteFirstMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_abs_le_polynomialStoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityQuarticSquareBlockWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_succ_square_sub_le_quarticSquareBlockWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledLocalDelta_le_inv_pow_ten","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedInverseCubePairEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockShiftedInverseCubePairEnvelope_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedCoordinateModelFailureCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockShiftedCoordinateModelFailureCharge_le_pairEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockShiftedCoordinateModelFailureCharge_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.quarticSquareBlockWeight_mul_tailModelFailureBudget_eq_tsum_shiftedCoordinateCharge","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_tailModelFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.summable_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_explicitHighProbabilityReturnDelta_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.tsum_quarticSquareBlockWeight_mul_explicitPolynomialPrefixTailModelReturnFailureBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_add_square_le_add_sum_Ico_quarticSquareBlockWeight","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.natCast_succ_square_le_checkpoint_square_add_tsum_quarticSquareBlockWeight_of_delayed","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_inverseSqrtThresholdUnboundedHittingAfterTailStart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterTailStart_spec","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentENNRealBudget_ne_top","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.inverseSqrtThresholdUnboundedHittingAfterStoppingRoundSecondMomentBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.lintegral_sq_untopA_add_one_le_quarticSquareCheckpointBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_two_untopA_add_one_of_eventually_quarticSquareCheckpointTail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_inverseSqrtThresholdUnboundedHittingAfterDelayedCheckpointSet_le_of_tailStart","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_lintegral_stoppingRound_sq_le_ENNRealBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_integral_stoppingRound_sq_le_budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageReturnDeviationProcess_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageBehaviorExpectedRegretProcess_le_two_mul_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_eq_behaviorExpected_sub_returnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviation_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedReturnDeviationIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedBehaviorExpectedRegret_and_returnDeviation_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.integral_abs_restrict_le_of_uniformIntegrable_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformAbsoluteContinuity","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.uniformIntegrable_one_of_memLp_and_tendsto_eLpNorm_sub_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.uniformIntegrable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedAverageRealizedBehaviorRegretIntegral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_uniformIntegrable_and_integral_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.one_add_log_le_two_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalLogarithmicAverageIntegratedBehaviorExpectedRegretRate_le_inverseSqrt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Coefficient","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Coefficient_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretL1Envelope_le_rawWindowInverseSqrtL1Envelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_sq_le_sqrt_rounds_add","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_inverseSquare","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_le_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowL1Rate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingL1Budget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_polynomialBaseGrowingRawWindow_offset_untopA_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsolutePolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalPolynomialBaseGrowingRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialBaseGrowingRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.measurable_apply_randomNat","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.ae_tendsto_apply_randomPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.tendsto_randomPrefix_atTop_of_nat_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalRandomPrefixAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_randomPrefixNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.sqrt_baseRounds_le_sqrt_add","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRawWindowInverseSqrtL1Envelope_add_le_base","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_le_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Rate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingL1Budget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.exists_rateControlledRawWindow_offset_untopA_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRawWindowCandidateRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityDoubleLinearRawWindowCandidateRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.memLp_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_le_budget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalExpectedAbsoluteRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eLpNorm_one_selfConsistentScheduledNaturalCausalRateControlledRawWindowStoppingAverageRealizedBehaviorRegret_sub_zero_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_rateControlledRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeBehaviorExpectedRegretLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeBehaviorExpectedRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninTailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninTailModelReturnFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalBurninTailModelReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninRealizedCumulativeLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninRealizedAverageLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_burnin_logarithmic_of_not_mem_tailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalBurninAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_tailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_burninTailHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy_le_rounds_mul","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityBurnin","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityBurnin_le_rounds","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_tendsto_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityScale_real_tendsto_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityRounds_tendsto_atTop","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitHighProbabilityReturnDelta_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy_le_rounds_mul","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageReturnRadius_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnFailureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnFailureBudget_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixTailModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixHighProbabilityAverageRealizedBehaviorRegretConsistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnVarianceProxyAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement_stronglyAdapted_piLE","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeSuccessorAverageReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorAverageReturnDeviationIncrement_succ_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.trajectoryMeasure_naturalCumulativeSuccessorAverageReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeSuccessorAverageReturnDeviation_eq_sum_range","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalSuccessorBatchAverageRealizedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalSuccessorBatchAverageRealizedRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalCumulativeRealizedBehaviorRegret_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeReturnDeviationProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_eq_expected_sub_deviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalCumulativeReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedCumulativeLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalRealizedAverageLogarithmicRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalModelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_naturalModelReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_le_logarithmic_of_not_mem_modelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalCumulativeRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretLogarithmicViolationSet_subset_modelReturnBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_fixedPrefixHighProbabilityLogarithmicCumulativeAverageRealizedBehaviorRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretProcess_neg_averageReturnRadius_lt_of_not_mem_event","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixTailModelReturnBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationSet_subset_event","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_le_failureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretDistanceViolationProbability_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegret_tendstoInMeasure_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailSet_subset_violationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_trajectoryMeasure_explicitPolynomialPrefixAverageRealizedBehaviorRegretViolationSet_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.eventually_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_le_failureBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailProbability_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_explicitPolynomialPrefixAverageRealizedBehaviorRegretUpperTailInProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalFirstPassageThreshold_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedSet_subset_distanceViolationSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.explicitPolynomialPrefixExpectedAbsoluteAverageRealizedBehaviorRegret_div_reciprocalFirstPassageThreshold_le_delayRate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayRate_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_le_rate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageDelayedProbability_tendsto_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_reciprocalThresholdCappedDoubleLinearRawWindowFirstPassage_vanishingDelayProbability_and_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalTrajectoryFiltration_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalTrajectory_coordinate_of_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalSuccessorBatchAverageRealizedRegret_naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalCumulativeRealizedBehaviorRegret_naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.measurable_naturalAverageRealizedBehaviorRegret_naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.HeterogeneousAdaptiveStochasticEpisodeBatchSource.naturalAverageRealizedBehaviorRegret_stronglyAdapted_naturalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalTrajectoryFiltration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalAverageRealizedBehaviorRegretProcess_stronglyAdapted","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalStoppingTimeAverageRealizedBehaviorRegretProcess","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_stoppingTimeNaturalAverageRealizedBehaviorRegret_tendstoAlmostEverywhere_zero_of_nat_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_base_of_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_eq_right_of_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurableSet_selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowEarlyStopSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_isStoppingTime","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_lower","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingPrefix_upper","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_thresholdTriggeredDoubleLinearRawWindowStoppingTimeNaturalAverageRealizedBehaviorRegret_L1_consistency","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stateOccupancy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_policyBellmanGap","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.policyBellmanGap_stageOfRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_inducedStateKernel_eq_integral_transitionValue","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_sub_bellman_eq_integral_inducedStateKernel_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_sub_comp_inducedStateKernel_eq_integral_bellman_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_eq_integral_optimalValueRemaining_sub_valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_le_optimalValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancyGapRemaining_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_integral_optimalValueAt_sub_valueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_occupancyGapRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_expectedRegret_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.expectedRegret_eq_occupancyGap_nonneg_and_optimalPolicy_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_le_optimalAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_le_optimalBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.transitionValue_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.bellmanQ_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_optimalValueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueRemaining_eq_of_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman_le_optimalBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_le_optimalValueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_bellman_eq_optimalBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalPolicy_valueAt_eq_optimalValueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_dominates_and_is_attained","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalBellman_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalBellmanCertificate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.measurable_upperValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.optimalValueRemaining_le_upperValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.occupancySumRemaining_mono","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.policyBellmanResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.measurable_policyBellmanResidual","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.policyBellmanResidual_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_eq_integral_upperValueRemaining_sub_valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.optimalBellmanCertificate_residualOccupancyRemaining_eq_occupancyGapRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.expectedRegret_le_residualOccupancyRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.residualOccupancyRemaining_le_occupancySumRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.OptimisticBellmanCertificate.expectedRegret_le_residual_le_occupancyBonusRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.inducedStateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.measurable_valueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_eq_of_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_horizon","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueAt_bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_zero_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateAt_eq_if","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionNextAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionAt_cons_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.stateActionNextAt_cons_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_stateActionAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_stateActionNextAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stepTrace_stateAt_eq_trajectoryStateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel_apply_singleton","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel_apply_actionSet","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_map_head","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_transitionEvent_eq_visitEvent_mul","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure_transitionEvent_eq_visitEvent_mul","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageTransitionJointProbability_eq_stageVisitProbability_mul_transition","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stageOfRemainingCoordinate_full","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining_visitEvent_eq_stateEvent_mul_action","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure_visitEvent_eq_stateEvent_mul_action","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageStateProbability","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stageVisitProbability_eq_stageStateProbability_mul_action","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardNextStateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellmanQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_reward_add_value","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellmanQ_eq_bellmanQ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticBellman_eq_bellman","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueRemaining_eq_valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticValueAt_eq_valueAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.deterministic","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.meanPlanningTransport","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.Concentration.intervalVarianceProxy_neg_coe","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationStepVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.meanBellmanInnovationVarianceProxy_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.valueRemaining_abs_le_of_rewardBound","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeMeanBellmanInnovationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeMeanBellmanInnovationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_map_dropReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_meanBellmanReturn_actionStateKernel_eq_valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_meanBellmanInnovation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeMeanBellmanInnovationFrom_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.headRewardMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardMean","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardMean_comap_headAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.headRewardDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_headRewardDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.UniformSubgaussianRewardLaw","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.uniformSubgaussianRewardLaw_of_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_hasCondSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headRewardDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.headAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.selectedRewardKernelAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_eq_compProd_selectedRewardKernelAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_map_fst","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_condDistrib_reward_given_action","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_condDistrib_headReward_given_headAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_condExpKernel_map_headReward_given_headAction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.Concentration.kernel_hasSubgaussianMGF_of_ae","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_apply_of_kernel_dirac","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardDeviationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardDeviationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_rewardDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeRewardDeviationFrom_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.eraseRewards","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_eraseRewards","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.eraseRewards_cons","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.ProbabilityTheory.compProd_map_prodMap_of_map_eq","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_eraseRewards","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.eraseTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_eraseTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_map_eraseTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.eraseTrajectoryFamily","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_eraseTrajectoryFamily","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_eraseTrajectoryFamily","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.knownRewardEpisodeBatchOfStochasticTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_knownRewardEpisodeBatchOfStochasticTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_knownRewardEpisodeBatch_eq_iidEpisodeBatchMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.rewardStepTrace_stateAt_eq_trajectoryStateAt_eraseTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_state_eq_knownRewardEpisodeStep","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_action_eq_knownRewardEpisodeStep","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStep_nextState_eq_knownRewardEpisodeStep","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatch_visitCount_eq_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatch_transitionCount_eq_knownRewardEpisodeBatch","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticSampledBatchCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.mem_stochasticSampledBatchCountBadEvent_iff_knownReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_stochasticSampledBatchCountBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_stochasticSampledBatchCountBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardSumConfidenceRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardCoordinateBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_rewardCoordinateBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_simultaneousRewardBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta_pos","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.simultaneousRewardDelta_le_one","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_simultaneousRewardBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviation_sum_abs_lt_of_not_mem_simultaneousRewardBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviation_sum_eq_sampledBatch_rewardSum_sub","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.expectedCountRewardCoordinateRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledBatch_empiricalReward_abs_sub_le_expectedCountRewardCoordinateRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.stochasticEmpiricalFiniteBatchValueEnvelope","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.stochasticAllCoordinateEmpiricalFiniteBatchModel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.StochasticAllCoordinateConfidence.upperValueRemaining_abs_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticAllCoordinateEmpiricalFiniteBatchModelConfidence_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurableSet_stochasticAllCoordinateEmpiricalModelBadEvent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_stochasticAllCoordinateEmpiricalModelBadEvent_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_finiteBatchModel_confidence","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_compProd_of_forall","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_zero_of_proxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_stateAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt_cons_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.stateAt_cons_succ","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_stateAt_prod","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.maskedRewardDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_maskedRewardDeviationAt","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_maskedRewardDeviationAt_trajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeStepOfStochasticTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledEpisodeStepOfStochasticTrajectory","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledEpisodeBatchOfStochasticTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledEpisodeBatchOfStochasticTrajectories","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_maskedRewardDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_maskedRewardDeviationAt_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_maskedRewardDeviationAt_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_maskedRewardDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.maskedRewardDeviationAtEpisode_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_maskedRewardDeviation_sum_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticRewardCoordinateRadius","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticRewardCoordinateRadius_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.expectedCountRewardCoordinateRadius_le_uniformFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticTransitionCover_of_uniformExpectedCountFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_explicitCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.selfConsistentTransitionBudget_fixedPoint","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionContraction","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticTransitionContraction_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.uniformFloorStochasticSelfConsistentTransitionBudget_fixedPoint","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.stochasticTransitionCover_of_uniformExpectedCountFloor_selfConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_uniformExpectedCountFloor_selfConsistent","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"declaration:BanditRLProof.FiniteHorizonRL.DeterministicMarkovPolicyTable.exploratoryPolicy_iidStochasticTrajectoryFamilyMeasure_allCoordinate_optimism_and_expectedRegret_of_pathSupport_selfConsistentCalibration","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationSum","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.iidSampledCumulativeReturnDeviationVarianceProxy","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_map_eval","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iIndepFun_sampledCumulativeReturnDeviationAtEpisode","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledCumulativeReturnDeviationAtEpisode_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.iidStochasticTrajectoryFamilyMeasure_sampledCumulativeReturnDeviationSum_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"declaration:BanditRLProof.Concentration.hasSubgaussianMGF_compProd_of_forall_fintype","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure_sampledCumulativeReturnDeviation_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.head","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_head","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.headActionReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headActionReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.headReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_headReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardMarginalKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_head","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headActionReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_map_headReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardKernel_apply_prod","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardMarginalKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_headMarginalFactorization","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:ProbabilityTheory.Kernel.compProd_prodMkRight_eq_prod","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:ProbabilityTheory.Kernel.map_compProd_prodMk","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.retainedActionStateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeReturnDeviationFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeReturnDeviationFrom_eq_rewardDeviation_add_meanBellmanInnovation","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationAfterActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_rewardDeviationAfterActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rawRewardKernelAfterActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState_apply","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterActionState_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rawRewardKernelAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.uncenterRewardAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_uncenterRewardAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.rewardDeviationKernelAfterRetainedActionState_map_uncenter","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rewardDeviation_map_uncenter","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.assembleRawRewardAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_assembleRawRewardAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rawReward_map_assemble","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.assembleRewardDeviationAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_assembleRewardDeviationAfterRetainedActionState","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_compProd_rewardDeviation_map_assemble","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.retainedActionStateKernel_meanBellmanInnovation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel_sampledReturnBellmanInnovation_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining_sampledCumulativeReturnDeviationFrom_abs_tail_le","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_cons","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.RewardStepTrace.measurable_tail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.integrable_of_fintype_aestronglyMeasurable","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.actionRewardStateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryKernelRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.sampledCumulativeRewardFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_sampledCumulativeRewardFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_sampledCumulativeRewardFrom_stochasticTrajectoryKernelRemaining_eq_stochasticValueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.stochasticTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.sampledCumulativeReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.measurable_sampledCumulativeReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integrable_sampledCumulativeReward_stochasticTrajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.MeanCompatibleRewardKernel.integral_sampledCumulativeReward_stochasticTrajectoryMeasure_eq_integral_valueAt_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_cons","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.StepTrace.measurable_tail","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.integrable_of_fintype","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.actionStateKernel","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryKernelRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.cumulativeRewardFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_cumulativeRewardFrom","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integrable_cumulativeRewardFrom_trajectoryKernelRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeRewardFrom_trajectoryKernelRemaining_eq_valueRemaining","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.trajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.cumulativeReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MDP.measurable_cumulativeReward","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integrable_cumulativeReward_trajectoryMeasure","relation":"contains"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.integral_cumulativeReward_trajectoryMeasure_eq_integral_valueAt_zero","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.measurable_selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointError_lt_iff","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointViolationSet_eq_jointError_ge","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedReturnJointGoodSet_eq_jointError_lt","relation":"contains"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdCappedUnboundedHittingAfter_stoppedSampledReturn_and_successorPolicyExpectedReturn_jointError_deterministicTailHighProbability_optimality","relation":"contains"},{"source":"module:BanditRLProof.RatMeasurability","target":"declaration:BanditRLProof.measurable_rat_div_const","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.realKernelMean","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.realKernelGap","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.realKernelRegret","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.realKernelGap_nonneg","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.realKernelRegret_eq_finset_sum_gap","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.realKernelRegret_eq_sum_gap_mul_pullCount","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.integrable_realKernelRegret_of_integrable_pullCount","relation":"contains"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"declaration:BanditRLProof.integral_realKernelRegret_eq_sum_gap_mul_integral_pullCount","relation":"contains"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"declaration:BanditRLProof.realMeanGap","relation":"contains"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"declaration:BanditRLProof.realMeanRegret","relation":"contains"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"declaration:BanditRLProof.realMeanRegret_eq_finset_sum_gap","relation":"contains"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"declaration:BanditRLProof.realMeanRegret_eq_sum_gap_mul_pullCount","relation":"contains"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"declaration:BanditRLProof.integrable_realMeanRegret_of_integrable_pullCount","relation":"contains"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"declaration:BanditRLProof.integral_realMeanRegret_eq_sum_gap_mul_integral_pullCount","relation":"contains"},{"source":"module:BanditRLProof.Regret","target":"declaration:BanditRLProof.pseudoRegret","relation":"contains"},{"source":"module:BanditRLProof.Regret","target":"declaration:BanditRLProof.pseudoRegret_zero","relation":"contains"},{"source":"module:BanditRLProof.Regret","target":"declaration:BanditRLProof.pseudoRegret_succ","relation":"contains"},{"source":"module:BanditRLProof.Regret","target":"declaration:BanditRLProof.RegretBoundCard","relation":"contains"},{"source":"module:BanditRLProof.Regret","target":"declaration:BanditRLProof.RegretObligation","relation":"contains"},{"source":"module:BanditRLProof.RegretCountBounds","target":"declaration:BanditRLProof.pseudoRegret_le_finset_sum_gap_mul_count_bound","relation":"contains"},{"source":"module:BanditRLProof.RegretCountBounds","target":"declaration:BanditRLProof.pseudoRegret_le_finset_sum_gap_mul_nat_count_bound","relation":"contains"},{"source":"module:BanditRLProof.RegretCountBounds","target":"declaration:BanditRLProof.pseudoRegret_le_sum_gap_mul_uniform_nat_count_bound","relation":"contains"},{"source":"module:BanditRLProof.RegretDecomposition","target":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.MarkovRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.ofKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_kernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_apply_of_measurable_index","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_eventProbability_of_measurable_index","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isProbabilityMeasure_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.apply_univ","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.const","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.deterministic","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.selectedMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.contextIndependentOfActionLaws","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.selectedMeasure_contextIndependentOfActionLaws","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isProbabilityMeasure_selectedMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.selectedMeasure_univ","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_selectedMeasure_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_selectedEventProbability_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_selectedMeasure_of_policy_state","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_selectedEventProbability_of_policy_state","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.policyContextStateIndex","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_policyContextStateIndex","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicy","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicy_kernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_composePolicy","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_composePolicy_eventProbability","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.policyActionOfContextState","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_policyActionOfContextState","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.policyActionKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.policyActionKernel_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_policyActionKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_kernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_composePolicyActionReward","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_composePolicyActionReward_eventProbability","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_reward_event","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_reward_map","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_kernel_apply_eq_map_prod_mk","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicyActionReward_action_map","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.CenteredRewardKernelLaw","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicy_centeredReward_integrable","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicy_centeredReward_integral_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.composePolicy_centeredReward_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.historyStepRewardKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_historyStepKernelFamily","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_historyStepKernelFamily_eventProbability","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_integrable","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_integral_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.historyStepKernelFamily_centeredReward_hasSubgaussianMGF","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.partialTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_partialTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_partialTrajectoryKernel_eventProbability","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.partialTrajectoryKernel_succ_next_map","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.partialTrajectoryKernel_succ_next_map_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_actionRewardHistoryStepKernelFamily","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_actionRewardHistoryStepKernelFamily_eventProbability","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_event","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_map","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_action_map","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_apply_eq_map_prod_mk","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.isMarkovKernel_actionRewardPartialTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.measurable_actionRewardPartialTrajectoryKernel_eventProbability","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_next_map","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_next_map_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardPartialTrajectoryKernel_succ_extend_map_apply","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_condDistrib_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_reward_condDistrib_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_action_condDistrib_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_selectedAction_condDistrib_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardKernel","target":"declaration:BanditRLProof.RewardKernel.actionRewardHistoryStepKernelFamily_selectedMeasure_condDistrib_trajMeasure","relation":"contains"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"declaration:BanditRLProof.RewardKernel.trajMeasure_map_eval_zero","relation":"contains"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"declaration:BanditRLProof.RewardKernel.rewardTrace_prefix_map_eq_trajMeasure_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"declaration:BanditRLProof.RewardKernel.rewardTrace_map_eq_trajMeasure_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"declaration:BanditRLProof.RewardKernel.identDistrib_rewardTrace_of_common_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.ScalarENNReal","target":"declaration:BanditRLProof.ENNReal.ofReal_finset_sum_mul_natCast_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.ScalarPseudoRegret","target":"declaration:BanditRLProof.real_pseudoRegret_eq_univ_sum_model_gap_mul_natCast_pullCount","relation":"contains"},{"source":"module:BanditRLProof.ScalarPseudoRegret","target":"declaration:BanditRLProof.ENNReal.ofReal_pseudoRegret_eq_univ_sum_model_gap_ofReal_mul_natCast_pullCount_of_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialProcess","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisPotentialStability_eq_linearLoss_sum_add_terminal_sub_initial","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.importanceWeightedPotentialStabilityScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.refinedPotentialStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.measurable_linearLoss_of_coordinatewise","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.measurable_halfTsallisPotentialValue","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.measurable_importanceWeightedPotentialStabilityScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.measurable_refinedPotentialStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_nonneg_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integrable_refinedPotentialStabilityBound_of_finiteSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integrable_score_comp_history_action_of_condDistrib_generic","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integral_importanceWeightedPotentialStabilityScore_le_integral_refinedBound_of_condDistrib_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integral_sum_importanceWeightedPotentialStabilityScore_le_integral_sum_refinedBound_of_condDistrib_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisRefinedPotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisRefinedPotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisSuccessorPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorPotentialStability_le_integral_sum_refinedPotentialStabilityBound_canonical","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisConjugatePotentialUpper","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.one_add_eta_mul_shift_mul_sqrt_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisConjugateCoordinateIncrement_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallis_fenchelCoordinate_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisConjugatePotentialUpper_le_shiftedMoments","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.importanceWeightedLoss_sub_selectedLoss_conjugate_domain","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.linearLoss_sub_linearLoss_eq_sum_sub_mul","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.sum_sub_mul_sub_baseline_eq_sum_sub_mul","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.linearLoss_sub_linearLoss_score_eq_sum_div_sqrt_of_stationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_le_conjugatePotentialUpper_of_feasible","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_importanceWeightedLoss_le_shiftedMoments_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_refined_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisMinimizer_mul_potentialStability_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.linear_sub_quadratic_le_sq_div_four","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sum_linear_sub_quadratic_le_unconstrained","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sum_linear_sub_quadratic_le_of_sum_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sum_inv_pos_of_nonempty","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_le_sqrt_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_probability_sub_gap_le_unconstrained","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_probability_sub_gap_le_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbability_sum_le_unconstrained","relation":"contains"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbability_sum_le_of_threshold","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.importanceWeightedStabilityScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.halfPowerStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_condDistrib_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.actionProcess_integral_importanceWeightedStabilityScore_le_integral_halfPowerStabilityBound_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisHistoryMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisHistoryUpdatedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.measurableFiniteActionDistribution_halfTsallisHistoryMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"declaration:BanditRLProof.Tsallis.actionProcess_integral_halfTsallisHistoryStability_le_integral_halfPowerStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.cumulativeLoss_sampledHalfTsallisObservedEstimatedLossAt_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_observedEstimatedLoss_eq_probabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedSuccessorStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialUpdatedAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialHistoryActionStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sum_observedEstimated_stability_eq_initial_add_successor","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallis_observedEstimatedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisObservedEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisPredictableEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisInitialUpdatedAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisInitialHistoryActionStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisInitialStability_eq_historyAction_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedSuccessorStabilityAt_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_mixedImportanceWeightedLoss_of_coordinates","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.measurable_weightedImportanceWeightedLoss_of_coordinates","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integrable_mixedImportanceWeightedLoss_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integrable_weightedImportanceWeightedLoss_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.initialHalfTsallisEnvironmentDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integrable_score_comp_history_action_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integrable_mixed_weightedImportanceWeightedLoss_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integral_mixed_weightedImportanceWeightedLoss_eq_predictable","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEstimatedLossAt_first_moments","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisPredictableLinearLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisObservedEstimatedRegret_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisObservedEstimatedRegret_eq_environmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisInitialStability_le_halfPower","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisSuccessorStabilitiesAt_canonical","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedStabilityScore_comp_history_action_of_condDistrib","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_halfPowerStabilityBound_comp_history","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_importanceWeightedStabilityScore_le_integral_sum_halfPowerStabilityBound_of_condDistrib_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_halfTsallisHistoryStability_le_integral_sum_halfPowerStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisSuccessorStabilityScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_halfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_score_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"declaration:BanditRLProof.Tsallis.halfTsallisCumulativeMinimizer_succ_eq_updated","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"declaration:BanditRLProof.Tsallis.cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty_half_canonical","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.HalfTsallisGeneratedSelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisFiniteHistorySelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.measurable_importanceWeightedStabilityScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisUpdatedAt_canonical","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisGeneratedSelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryActionStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_selector","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_canonical","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_apply_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.prob_mul_abs_importanceWeightedStabilityScore_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.integrable_finiteActionMeasure","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedStabilityScore_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.integrable_halfPowerStabilityBound_of_finiteSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPolicyAt_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisHistoryActionStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.integrable_sampledHalfTsallisHalfPowerBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound_of_measurable","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_simplexPairShift_of_eq_zero_of_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.linearLoss_simplexPairShift_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.sum_sqrt_simplexPairShift_eq_of_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.sqrt_sub_sqrt_sub_le_div_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.exists_transfer_strictly_improves_half_objective","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer_auto","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.extendFiniteWeights","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.extendFiniteWeights_apply_of_mem","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.extendFiniteWeights_apply_of_not_mem","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.sum_extendFiniteWeights","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_extendFiniteWeights","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.restrict_mem_stdSimplex_of_finiteSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.linearLoss_extendFiniteWeights_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.linearLoss_restrict_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.sum_sqrt_extendFiniteWeights_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.sum_sqrt_restrict_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.regularizedObjective_half_extendFiniteWeights_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.regularizedObjective_half_restrict_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.continuous_regularizedObjective_half_univ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.exists_isRegularizedMinimizer_half","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.halfTsallisUpdatedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.halfTsallisUpdatedMinimizer_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisMinimizer_mul_linearLoss_sub_updated_le_powerSum_half","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"declaration:BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_minimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_mem_stdSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_isMinOn","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.continuous_regularizedObjective_half_restricted_joint","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.continuous_restrictedHalfTsallisMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer_eq_on_arms_of_score_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.restrictedHalfTsallisMinimizer_restrict_score_apply","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"declaration:BanditRLProof.Tsallis.measurable_halfTsallisMinimizer_comp","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","target":"declaration:BanditRLProof.Tsallis.strictConcaveOn_sum_sqrt_stdSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","target":"declaration:BanditRLProof.Tsallis.strictConvexOn_regularizedObjective_half_stdSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","target":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_half_eq_on_arms","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","target":"declaration:BanditRLProof.Tsallis.halfTsallisMinimizer_eq_on_arms","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"declaration:BanditRLProof.Tsallis.HalfTsallisInteriorStationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisInteriorStationary_rpow_sub_rpow_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"declaration:BanditRLProof.Tsallis.sub_le_two_mul_rpow_three_halves_mul_neg_half_rpow_sub","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"declaration:BanditRLProof.Tsallis.linearLoss_sub_next_importanceWeightedLoss_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.HalfTsallisFiniteHistorySelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.initialHalfTsallisDistribution","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.finiteActionDistribution_initialHalfTsallisDistribution","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryScore_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.measurable_sampledHalfTsallisHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisActionAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisScoreAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisUpdatedAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHistoryActionStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisHalfPowerBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisSuccessorStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisEnvironmentHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPolicyAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisScoreAt_succ_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledHalfTsallisSuccessorStability_le_integral_sum_halfPowerStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.cumulativeLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.cumulativeLoss_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.cumulativeLoss_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.linearLoss_add_right","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.linearLoss_cumulativeLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.regularizedObjective_cumulativeLoss_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.eta_mul_sum_next_linearLoss_add_regularizer_zero_le_objective","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.sum_next_linearLoss_sub_comparator_le_regularizer_penalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.cumulativeLinearLoss_sub_comparator_le_stability_add_penalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.FTRL.cumulativeLinearLoss_sub_comparator_le_stability_add_penalty_simplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.Tsallis.negEntropyRegularizer_sub_eq_powerSum_sub_div","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"declaration:BanditRLProof.Tsallis.cumulativeLinearLoss_sub_comparator_le_stability_add_powerSumPenalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.pairDirection","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.simplexPairShift_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.sum_pairDirection_mul","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.sum_simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_simplexPairShift_of_abs_lt_min","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.regularizedObjective_half_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.hasDerivAt_simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.hasDerivAt_linearLoss_simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.sum_pairDirection_div_two_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.hasDerivAt_sum_sqrt_simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.hasDerivAt_regularizedObjective_half_simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.isLocalMin_regularizedObjective_half_simplexPairShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.halfTsallis_pairwise_stationary_of_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.exists_halfTsallisInteriorStationary_of_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.two_mul_sqrt_sub_sqrt_le_sub_div_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_of_halfTsallisInteriorStationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.isRegularizedMinimizer_iff_exists_halfTsallisInteriorStationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_powerSum_half_of_positive_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.armDependentSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_armDependentSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDArmDependentSuboptimalBoostRefinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_armDependentSuboptimalRewardBoostSource_of_refinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDArmDependentSuboptimalBoostAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDArmDependentSuboptimalBoostRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReal","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReal_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReal_eq_of_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_clippedUnitReal_add_sub_self_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedReward","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDStationaryCorruptedRewardVectorLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryCorruptedRewardVectorLoss_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_finiteArmIIDStationaryCorruptedReward_sub_base_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_stationaryCorruptedLossDiff_sub_baseLossDiff_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integrable_iidLossStateDiff","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_iidLossStateMeanGap_stationaryCorrupted_sub_modelGap_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.scheduledGapDeviationBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.scheduledGapDeviationBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_perturbedExpectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDStationaryRewardCorruptionBudget_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDStationaryCorruptedRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHistoryAdaptiveRewardShiftSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_clippedUnitReal","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_initial","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_successor","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveCorruptedRewardLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_hasIIDStateCoordinateLocality","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.predictableLossAt_finiteArmIIDHistoryAdaptiveCorruptedRewardLoss_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.zeroFiniteArmIIDHistoryAdaptiveRewardShiftSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardShiftAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDHistoryAdaptiveRewardShiftAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_finiteArmIIDHistoryAdaptiveRewardShiftAt_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_historyAdaptiveCorruptedPredictableLossDiff_sub_baseLossDiff_le_actualShift","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbability_mul_historyAdaptiveRewardShiftDeviation","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudget_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw_hasSelfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedRefinedCorruptionWindow","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveExpectedCorruptionAllRegimeBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostExpectedCorruptionRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw_hasSelfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLawRegret_le_refinedLocalExplicit_of_window","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_lt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftSuccessor_of_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_lt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveRewardShiftEnvelope_succ_of_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_initial","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_lt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_successor_of_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_lt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.FiniteArmIIDHorizonHistoryAdaptiveRewardShiftSource.toAllTime_envelope_succ_of_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveCorruptedRewardLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveExpectedRewardCorruptionBudgetForLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHorizonHistoryAdaptiveExpectedCorruptionAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.measurableHistoryArmGatedSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_measurableHistoryArmGatedSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_measurableHistoryArmGatedSuboptimalRewardBoostSource_of_refinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.previousActionGatedSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_previousActionGatedSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_previousActionGatedSuboptimalRewardBoostSource_of_refinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReward","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReward_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReward_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.clippedUnitReward_eq_of_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_clippedUnitReward","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDRewardVectorLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDRewardVectorLoss_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.integrable_clippedUnitReward","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_finiteArmIIDRewardVectorLaw_clippedUnitReward_eq_mean","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.iidLossStateMeanGap_finiteArmIIDRewardVectorLoss_eq_gap","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedReward","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.measurable_finiteArmIIDTimeVaryingCorruptedRewardVectorLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingCorruptedRewardVectorLoss_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_iidLossStateTimeVaryingMeanGap_corrupted_sub_modelGap_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingRewardCorruptionBudget_const_eq_stationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingCorruptedRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.timeVaryingSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_timeVaryingSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostRefinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_timeVaryingSuboptimalRewardBoostSource_of_refinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDTimeVaryingSuboptimalBoostAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDTimeVaryingSuboptimalBoostRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.uniformSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRewardCorruptionBudget_uniformSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDUniformSuboptimalBoostRefinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDHistoryAdaptiveRefinedCorruptionWindow_uniformSuboptimalRewardBoostSource_of_refinedRegime","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIIDUniformSuboptimalBoostAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_refinedLocalExplicit","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDUniformSuboptimalBoostRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanAllRegimeBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentBestArmAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_le_bestArmAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_sub_bestArm_le_meanDeviation","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentMovingComparatorRewardLawRegret_eq_fixed_add_meanAdvantage","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanDynamicAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanDynamicAllRegimeBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDriftingMeanRefinedCorruptionWindow","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_refinedLocalExplicit_of_window","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_mean_sub","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.abs_independentLossStateTimeVaryingMeanGap_sub_modelGap_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLaw_hasSelfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_mono","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanDeviationBudget_globalMeanSwitchCount_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentDynamicComparatorPenalty_globalMeanSwitchCount_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCount_totalBudget_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountHorizonCompressedLogDynamicBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountHorizonCompressedDynamicRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_le_globalMeanSwitchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_le_globalMeanSwitchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_globalMeanSwitchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanSwitchCountDynamicAllRegimeBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanSwitchCount_eq_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_le_switchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_switchCount","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentMeanSwitchCountDynamicAllRegimeBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentMeanSwitchCountDynamicRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeMeanPathVariation_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_zero_le_cumulativeMeanPathVariation","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.abs_finiteArmIndependentRewardMean_sub_model_le_cumulativeMeanPathVariation","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentPathVariationDynamicAllRegimeBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentPathVariationDynamicAllRegimeBound_fin_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentPathVariationDynamicRegret_le_allRegimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardVectorLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","target":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_finiteArmIndependentRewardVectorLoss_eq_gap","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","target":"declaration:BanditRLProof.Tsallis.hasScheduledIndependentMeanGapLaw_of_finiteArmIndependentRewardVectorLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentRewardLawRegret_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionModel","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionModel_bestArm","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw_isProbabilityMeasure","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionLaw_mem_Icc_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_one_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionMean_one_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionInitialMean","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGap_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGap_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionBestArmAt_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionBestArmAt_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionIndicator","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchObstructionGlobalCount_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchDynamicComparatorPenalty_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchComparatorAdvantage_gt_nat_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentSingleSwitchDynamicComparatorRouteObstruction","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"declaration:BanditRLProof.Exp3.finiteBanditMeanLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"declaration:BanditRLProof.Exp3.finiteBanditMeanLoss_initial","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"declaration:BanditRLProof.Exp3.finiteBanditMeanLoss_successor","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"declaration:BanditRLProof.Exp3.predictableLossAt_finiteBanditMeanLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"declaration:BanditRLProof.Exp3.predictableLossAt_finiteBanditMeanLoss_sub_bestArm_eq_gap","relation":"contains"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteBanditMeanLossRegret_le_log_fixedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.powerWeightedSquaredImportanceWeightedLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.powerWeightedSquaredImportanceWeightedLoss_eq_selected","relation":"contains"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_powerWeightedSquaredImportanceWeightedLoss_le_powerSum","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"declaration:BanditRLProof.Tsallis.oracleRestartEpochRounds","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableOracleRestartEpochRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_oracleRestartEpochRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sum_sqrt_oracleRestartEpochRounds_card_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSwitchCountSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisEnvironmentHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPolicyAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPolicyAt_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisPredictableEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEstimatedLossAt_first_moments","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledOracleRestartHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisPredictableLinearLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_eq_environmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochEnvironmentRegret_pointMass_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_eq_epochRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret_le_of_estimatedRegret_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt_of_epochObservedEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt_of_epochObservedEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.oracleRestartLocalTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisActionAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.frestrictLe_oracleRestartShiftedTrajectory_eq_localPairHistory","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisScoreAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisUpdatedAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedPotentialStabilityAtSuccessor_eq_shifted","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_eq_historyAction_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisScoreAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisUpdatedAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisRefinedPotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAt_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisUpdatedAt_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisHistoryActionPotentialStabilityAt_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisInitialPotentialStabilityAtTime_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtTime_le_integral_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_le_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_card_add_penalty_of_epochRounds_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_sqrtSchedule_le_scheduleSwitchCountSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.oracleNeverRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.oracleRestartEveryRoundSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.start_succ_le_of_ne","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.oracleRestartLocalPairHistory","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.oracleRestartLocalPairHistory_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.measurable_oracleRestartLocalPairHistory","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_of_boundary","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_of_continuation","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_neverRestart","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistribution_restartEveryRound","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisHistoryAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_neverRestart","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_restartEveryRound","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartStart","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartSchedule_start_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleChangePointRestartSchedule_start_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_oracleChangePointRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_oracleChangePointRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeTimes","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentGlobalMeanChangeRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_finiteArmIndependentGlobalMeanChangeRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentCumulativeGlobalMeanSwitchCount_eq_changeTimes_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs_card_cast_finiteArmIndependentGlobalMeanChangeRestartSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_eq_globalMeanChangeRestartStart","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.finiteArmIndependentRewardMean_le_globalMeanChangeRestartBestArm","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap_globalMeanChangeRestartBestArm_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableEnvironmentRegret_neverRestart","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_neverRestart","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_fixed_add","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.oracleRestartScheduleEpochs","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.oracleRestartSchedule_start_mem_epochs","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableScheduleEpochRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_eq_sum_scheduleEpochRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_scheduleSwitchCountSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.refinedPotentialStabilityBound_le_two_mul_eta_mul_sqrt_erase_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtSuccessor_le_two_mul_eta_mul_sqrt_erase_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.one_div_natSucc_le_one_div_sqrt_natSucc","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_two_mul","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_two_mul_sq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.one_div_sampledScheduledHalfTsallisSqrtSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.sum_range_sampledScheduledHalfTsallisSqrtSchedule_refinedBudget_le_three_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledOracleRestartHalfTsallisShiftedPotentialStabilityAtLocalPrefix_sqrtSchedule_le_four_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.initialHalfTsallisPotentialMass_sqrtSchedule_pointMassPenalty_le_four_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_sqrtSchedule_le_eight_mul_sqrt_card_of_mem","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.oracleRestartShiftedTrajectory","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.oracleRestartShiftedTrajectory_fst","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.oracleRestartShiftedTrajectory_snd_apply","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.monotone_start","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.start_start","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.OracleRestartSchedule.start_eq_of_between","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_eq_scheduled_shift","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_eq_scheduled_shift","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisProbabilityAtTime_add_eq_scheduled_of_start_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedEstimatedLossAt_add_eq_scheduled_of_start_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_eq_scheduled","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisLocalPrefixObservedEstimatedRegret_pointMass_le_stability_add_penalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.exists_oracleRestartEpochRounds_eq_image_range","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_eq_localPrefix_of_epochRounds_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty_of_epochRounds_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"declaration:BanditRLProof.Tsallis.exists_sampledOracleRestartHalfTsallisObservedScheduleEpochEstimatedRegret_pointMass_le_stability_add_penalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.CounterAction","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterArms","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterEta","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterProb","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterLoss","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterNext","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterNextMultiplier","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_49","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_576","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_625","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_24778200568643041","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_1813828968643041","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_22964371600000000","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_1524122302720081","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_120121402720081","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.sqrt_1404000900000000","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterEta_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterEta_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterProb_simplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterProb_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterLoss_mem_Icc","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterNext_simplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterNext_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.rpow_neg_half_eq_inv_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterProb_stationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterNext_stationary","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterProb_minimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.counterNext_minimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"declaration:BanditRLProof.Tsallis.exists_minimizer_counterexample_to_refinedAveragedStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.shiftedHalfPowerImportanceWeightedMoment","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.shiftedPositiveCubicImportanceWeightedMoment","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_mul_sum_erase_eq_sum_mul_one_sub","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.prob_mul_shiftedHalfPowerImportanceWeightedMoment_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_shiftedHalfPowerImportanceWeightedMoment_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.prob_mul_shiftedPositiveCubicImportanceWeightedMoment_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_shiftedPositiveCubicImportanceWeightedMoment_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_stability_le_refinedHalfPower_add_square","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_linearLoss_sub_next_importanceWeightedLoss_le_refined_of_shiftedTaylor","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"declaration:BanditRLProof.Tsallis.probability_le_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"declaration:BanditRLProof.Tsallis.sqrt_mul_one_sub_le_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"declaration:BanditRLProof.Tsallis.sqrt_mul_one_sub_le_one_sub","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"declaration:BanditRLProof.Tsallis.one_sub_eq_sum_erase","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"declaration:BanditRLProof.Tsallis.sum_sqrt_mul_one_sub_le_two_mul_sum_erase_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"declaration:BanditRLProof.Tsallis.regret_le_of_refinedHalfPowerSelfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"declaration:BanditRLProof.Tsallis.powerSum","relation":"contains"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"declaration:BanditRLProof.Tsallis.entropy","relation":"contains"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"declaration:BanditRLProof.Tsallis.negEntropyRegularizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"declaration:BanditRLProof.Tsallis.one_sub_exponent_ne_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"declaration:BanditRLProof.Tsallis.powerSum_nonneg_of_finiteSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"declaration:BanditRLProof.Tsallis.negEntropyRegularizer_wellDefined_on_finiteSimplex","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_le_linearLoss_sub_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.sum_prob_mul_halfTsallisPotentialStability_importanceWeightedLoss_le_one_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_importanceWeightedPotentialStabilityScore_finiteActionKernel_coarse","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_importanceWeightedPotentialStabilityScore_le_integral_one_of_condDistrib_of_minimizers","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSuccessorAllRatePotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialAllRatePotentialStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_allRateBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_allRateBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_allRate","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPastSigma","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPastSigma_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisProbabilityAtTime_pastSigma","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"declaration:BanditRLProof.Tsallis.HasScheduledConditionalMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledExpectedGapLaw_of_conditionalMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.HasScheduledExpectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbability_mul_predictableLossDiffAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_weightedLossGapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass_of_expectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_expectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisPredictableEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisObservedEstimatedLossAt_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEstimatedLossAt_first_moments","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledScheduledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisPredictableLinearLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEstimatedRegret_eq_environmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_eq_predictable_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisActionAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisScoreAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisUpdatedAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEnvironmentHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPolicyAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPolicyAt_eq_finiteActionKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.HalfTsallisScheduleGeneratedSelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisUpdatedAt_canonical","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisScheduleGeneratedSelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisRefinedPotentialStabilityBoundAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisHistoryActionPotentialStabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime_succ_eq_historyAction_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisSuccessorPotentialStabilityAtTime_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableSuboptimalGapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalGapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableSuboptimalGapMass_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalExpectedGapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_fixedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_fixedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","target":"declaration:BanditRLProof.Tsallis.HasIIDStateCoordinateLocality","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidStateCoordinateLocality_prefix_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDHistoryAdaptivePrefixKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","target":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisHistoryAdaptiveTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.iidLossStatePredictableLossVector","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.predictableLossAt_iidLossStatePredictableLossVector","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.extendLossStatePrefix","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.measurable_extendLossStatePrefix","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.extendLossStatePrefix_apply_of_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidLossState_prefix_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDPrefixKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.iidLossStateMeanGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.HasScheduledIIDPrefixKernelFactorization","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledIndependentMeanGapLaw_of_iidLossState","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_iidLossState","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.iidTimeVaryingLossStatePredictableLossVector","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.predictableLossAt_iidTimeVaryingLossStatePredictableLossVector","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel_map_frestrictLe_eq_of_iidTimeVaryingLossState_prefix_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisIIDTimeVaryingPrefixKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledIIDPrefixKernelFactorization_sampledScheduledHalfTsallisTimeVaryingTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.independentLossStateTimeVaryingMeanGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.iidLossStateTimeVaryingMeanGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingIndependentMeanGapLaw_of_independentLossState","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingIndependentMeanGapLaw_of_iidLossState","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"declaration:BanditRLProof.Tsallis.HasScheduledIndependentMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledConditionalMeanGapLaw_of_independentMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledExpectedGapLaw_of_independentMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_independentMeanGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisInitialHistoryActionPotentialStability","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisInitialRefinedPotentialStabilityBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime_zero_eq_initial_ae","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisInitialPotentialStabilityAtTime_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.HalfTsallisScheduleFiniteHistorySelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.canonicalHalfTsallisScheduleFiniteHistorySelectorMeasurability","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryScore_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.measurable_sampledScheduledHalfTsallisHistoryScore","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryDistribution","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryDistributionSource","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAlgorithm","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHistoryAlgorithm_policy","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryKernel","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisTrajectoryMeasure_condDistrib_action_given_environment","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw_of_expectedDeviation","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_referenceExpectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.sqrt_sub_self_le_half_one_sub","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_sub_one_le_two_mul_sum_erase_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedSuboptimalMassAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisRefinedSuboptimalMassAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisRefinedSuboptimalMassAt_le_expected","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedExpectedPenalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedSuboptimalSqrtMassAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedRefinedSuboptimalMassAt_le_sqrtMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_selfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_fixedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedStabilityPenalty_of_expectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisObservedEstimatedLossAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.cumulativeLoss_sampledScheduledHalfTsallisObservedEstimatedLossAt_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.halfTsallisScheduledMinimizer_observedEstimatedLoss_eq_probabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSameRateNextAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialStabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPotentialPenaltyAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_eq_stability_add_penalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisPotentialPenalty_pointMass_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_pointMass_le_stability_add_penalty","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","target":"declaration:BanditRLProof.Tsallis.regret_le_selfBoundingInterpolation","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingInterpolation","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.sum_sampledScheduledHalfTsallisExpectedProbability_le_quadraticFilterSplit","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.sum_range_sampledScheduledHalfTsallisExpectedProbability_le_quadraticPrefixSplit","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_selfBoundingQuadraticSplit","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSum_of_refinedCoefficient_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedSelfBoundingQuadraticSplit_of_refinedCoefficient_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisExpectedProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integrable_sampledScheduledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integrable_sqrt_sampledScheduledHalfTsallisProbabilityAtTime","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_sampledScheduledHalfTsallisExpectedProbabilityAt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integral_sqrt_sampledScheduledHalfTsallisProbabilityAtTime_le_sqrt_expected","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_eq_refined","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisAllRatePotentialStabilityBoundAtTime_le_suboptimalExpectedSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_suboptimalExpectedSqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_of_selfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.HasScheduledTimeVaryingExpectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.HasScheduledTimeVaryingConditionalMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.HasScheduledTimeVaryingIndependentMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingConditionalMeanGapLaw_of_independentMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingExpectedGapLaw_of_conditionalMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.hasScheduledTimeVaryingExpectedGapLaw_of_independentMeanGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_suboptimalTimeVaryingExpectedGapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.scheduledTimeVaryingGapDeviationBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding_of_timeVaryingPerturbedExpectedGapLaw","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.pointMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.finiteSimplex_pointMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.linearLoss_pointMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.linearLoss_sub_pointMass_eq_gapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableGapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.HasSelfBoundingRegret","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.sampledHalfTsallisPredictableEnvironmentRegret_pointMass_eq_gapMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_hasSelfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.integral_sampledHalfTsallisPredictableEnvironmentRegret_pointMass_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.two_mul_coeff_mul_sqrt_sub_gap_mul_le_sq_div_gap","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.regret_le_two_mul_base_add_sum_sq_div_gap_add_corruption","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.powerSum_pointMass_half","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.sum_erase_sqrt_pointMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"declaration:BanditRLProof.Tsallis.not_forall_powerSum_half_le_sum_erase_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","target":"declaration:BanditRLProof.Tsallis.selfBoundingBetaEquation","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","target":"declaration:BanditRLProof.Tsallis.continuousOn_selfBoundingBetaEquation","relation":"contains"},{"source":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","target":"declaration:BanditRLProof.Tsallis.exists_selfBoundingBetaEquation_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget_eq_harmonic","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisHarmonicBudget_le_one_add_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_le_half","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_succ_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSqrtSchedule_four_mul_sq","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sqrt_succ_sub_sqrt_le_one_div_sqrt_succ","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_nonneg","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisRefinedCoefficient_sqrtSchedule_sq_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_of_selfBounding","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_fixedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_expectedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_fixedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_expectedGap","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.sum_range_one_div_sqrt_natSucc_le_two_sqrt","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.sum_Ico_one_div_natSucc_le_log_div","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.sum_range_sqrtSchedule_activeBranch_le_closedForm","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.sum_Ico_sqrtSchedule_unconstrainedBranch_le_log","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticSplit","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticPrefixSplit","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingQuadraticClosedForm","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalBetaEquation","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.continuousOn_refinedLocalBetaEquation","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.sq_sqrt_sub_one_le_sub_log_sub_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalBetaWeight_bounds_of_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.exists_refinedLocalBetaEquation_eq_zero","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.exists_refinedLocalBetaEquation_eq_zero_and_weight_bounds","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalLambda","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha_sq","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha_pos","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalAlpha_le_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.refinedLocalLambda_mem_Ioc","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.one_add_refinedLocalLambda_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"declaration:BanditRLProof.Tsallis.two_mul_refinedLocalLambda_div_one_add_eq_alpha","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"declaration:BanditRLProof.Tsallis.sum_one_div_lambda_mul_eq_div","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingThreshold_refinedLocalLambda_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"declaration:BanditRLProof.Tsallis.refinedLocalTunedRegretBound","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"declaration:BanditRLProof.Tsallis.refinedLocalTunedRegretBound_le_explicit","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"declaration:BanditRLProof.Tsallis.exists_integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalTuned","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_refinedLocalExplicit","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","target":"declaration:BanditRLProof.Tsallis.RefinedLocalCorruptionWindow","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","target":"declaration:BanditRLProof.Tsallis.refinedLocalCorruptionWindow_scalar_bounds","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.natFloor_positive_and_half_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingThreshold","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_refinedSelfBoundingFloorCutoff","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingOneThreshold","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisSelfBoundingOneThreshold_eq","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne","relation":"contains"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_selfBoundingOne_explicit","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_eq_two_mul_powerSum_sub_one","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_pointMass","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue_eq_neg_linearLoss_add_mass_div","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue_le_of_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialValue_new_sub_old_le_rateChange_mul_mass","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisPotentialMass_le_of_zero_isRegularizedMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.sum_range_succ_sub_eq_first_sub_last_add_cross","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisScheduledMinimizer","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.halfTsallisScheduledSameRateNext","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisScheduledPotentialPenalty_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisScheduledPotentialPenalty_le_initial_sub_comparator_div","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisCanonicalScheduledPotentialPenalty_le","relation":"contains"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"declaration:BanditRLProof.Tsallis.sum_halfTsallisCanonicalScheduledPotentialPenalty_pointMass_le","relation":"contains"},{"source":"module:BanditRLProof.UCBSummability","target":"declaration:BanditRLProof.UCBSummability.finiteHorizonBadEvent","relation":"contains"},{"source":"module:BanditRLProof.UCBSummability","target":"declaration:BanditRLProof.UCBSummability.measure_finiteHorizonBadEvent_le_sum","relation":"contains"},{"source":"module:BanditRLProof.UCBSummability","target":"declaration:BanditRLProof.UCBSummability.measure_finiteHorizonBadEvent_le_tail_sum","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberMeasure_eq_lintegral_rounds_sq","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberMeasure_ne_top","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.tsum_natSuccSquare_mul_stoppingFiberRealMeasure_eq_integral_rounds_sq","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.summable_sqrt_stoppingFiberRealMeasure_of_memLp_two","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.tsum_sqrt_stoppingFiberRealMeasure_le_half_mul_integral_rounds_sq_add_tsum_inverse_natSuccSquare_of_memLp_two","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.tsum_sqrt_stoppingFiberRealMeasure_le_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.integrable_stoppedValue_of_uniform_secondMoment_of_memLp_two_rounds","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_tsum_sqrt_stoppingFiberRealMeasure_of_memLp_two_rounds","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_sqrt_integral_rounds_sq_mul_sqrt_tsum_inverse_natSuccSquare_of_memLp_two_rounds","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"declaration:BanditRLProof.integral_abs_stoppedValue_le_uniformSecondMoment_mul_half_roundSecondMoment_add_inverseSquareTsum_of_memLp_two_rounds","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","target":"declaration:BanditRLProof.measurable_stoppedValue_of_measurable_coordinates","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","target":"declaration:BanditRLProof.summable_abs_weight_mul_sqrt_stoppingFiberRealMeasure_and_tsum_le","relation":"contains"},{"source":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","target":"declaration:BanditRLProof.integrable_and_integral_abs_stoppedValue_weight_mul_le","relation":"contains"},{"source":"group:milestones","target":"milestone:CAUSAL-PARALLEL-ACTUAL-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:HOO-REWARD-FAMILY-RATE","relation":"contains"},{"source":"group:milestones","target":"milestone:HEAVY-TAIL-GENALTI-FINITE-SUPREMUM","relation":"contains"},{"source":"group:milestones","target":"milestone:HEAVY-TAIL-PRINTED-REGRET-COUNTEREXAMPLE","relation":"contains"},{"source":"group:milestones","target":"milestone:HEAVY-TAIL-CLIPPED-CORRUPTION-TRANSFER","relation":"contains"},{"source":"group:milestones","target":"milestone:HEAVY-TAIL-CORRECTED-SOURCE-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:HEAVY-TAIL-SOURCE-SCHEDULE-CONFIDENCE","relation":"contains"},{"source":"group:milestones","target":"milestone:HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","relation":"contains"},{"source":"group:milestones","target":"milestone:FOUNDATION-REGRET-DECOMPOSITION","relation":"contains"},{"source":"group:milestones","target":"milestone:PROBABILITY-GENERATED-COND-MGF","relation":"contains"},{"source":"group:milestones","target":"milestone:PROBABILITY-FINTYPE-GEOMETRIC-ALL-TIME-UNION","relation":"contains"},{"source":"group:milestones","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-GEOMETRIC-ALL-TIME","relation":"contains"},{"source":"group:milestones","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","relation":"contains"},{"source":"group:milestones","target":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:PROBABILITY-ETC-UCB-ROUTE-SURFACE","relation":"contains"},{"source":"group:milestones","target":"milestone:ETC-CANONICAL-SUBGAUSSIAN-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:ETC-CANONICAL-RAT-LEAST-TIE","relation":"contains"},{"source":"group:milestones","target":"milestone:ETC-LML-PORT","relation":"contains"},{"source":"group:milestones","target":"milestone:UCB-FINITE-ARM-SUBGAUSSIAN","relation":"contains"},{"source":"group:milestones","target":"milestone:UCB-EXPECTED-AVERAGE-CONSISTENCY","relation":"contains"},{"source":"group:milestones","target":"milestone:UCB-HORIZON-INDEXED-CANONICAL-CHAIN","relation":"contains"},{"source":"group:milestones","target":"milestone:UCB-LML-PORT","relation":"contains"},{"source":"group:milestones","target":"milestone:LML-DIRECT-TOOLCHAIN-IDENTITY","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-ELLIPTICAL-POTENTIAL","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-SELF-NORMALIZED-RIDGE-CONFIDENCE","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-MEASURABLE-GENERATED-POLICY","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-ALL-TIME-CONFIDENCE","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-ALL-HORIZON-HIGH-PROBABILITY-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-EXPECTED-AVERAGE-CONSISTENCY","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-BOUNDED-STOPPING-EXPECTED-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:OFUL-UNBOUNDED-STOPPING-EXPECTED-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-POSTERIOR-KERNEL","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-CANONICAL-SAMPLER","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-RECURSIVE-PROBABILITY-MATCHING","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-BAYES-CLIPPED-DECOMPOSITION","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-LATENT-STREAM-SUPPORT","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-STATIONARY-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:THOMPSON-GENERAL-PORT","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-EXPECTED-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-BEST-ARM-REALIZED-HIGH-PROBABILITY","relation":"contains"},{"source":"group:milestones","target":"milestone:CONCENTRATION-COUNTABLE-SCHEDULED-QUADRATIC-TAIL","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-PREDICTABLE-VARIANCE-GEOMETRIC-ALL-TIME","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-REALIZED-DEVIATION-GEOMETRIC-ALL-TIME","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-PREDICTABLE-REGRET-GEOMETRIC-ALL-TIME","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-REALIZED-REGRET-GEOMETRIC-ALL-TIME","relation":"contains"},{"source":"group:milestones","target":"milestone:EXP3-SPARSE-ALL-HORIZON","relation":"contains"},{"source":"group:milestones","target":"milestone:TSALLIS-IID-LOG","relation":"contains"},{"source":"group:milestones","target":"milestone:TSALLIS-HISTORY-ADAPTIVE-CORRUPTION","relation":"contains"},{"source":"group:milestones","target":"milestone:TSALLIS-DYNAMIC-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:TSALLIS-ORACLE-RESTART-GENERATED","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-FINITE-MDP-BELLMAN","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-OCCUPANCY-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-ADAPTIVE-REALIZED-CONSISTENCY","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-UNBOUNDED-HITTINGAFTER-L2","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-UNBOUNDED-HITTINGAFTER-EXPECTED-UPPER-BOUND","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"contains"},{"source":"group:milestones","target":"milestone:RL-UCBVI-HOEFFDING-CANONICAL-TERMINALS","relation":"contains"},{"source":"group:milestones","target":"milestone:BWK-STOPPING-FOUNDATION","relation":"contains"},{"source":"group:milestones","target":"milestone:BWK-FINAL-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH15-EX15-7-DATA-PROCESSING-LEAF","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH16-SOURCE-TERMINALS","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"contains"},{"source":"group:milestones","target":"milestone:TEXTBOOK-PART-IV-THEOREM-13-1-GAUSSIAN-MINIMAX","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"contains"},{"source":"group:milestones","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"contains"},{"source":"group:milestones","target":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","relation":"contains"},{"source":"group:milestones","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"contains"},{"source":"group:milestones","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"contains"},{"source":"group:milestones","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"contains"},{"source":"group:milestones","target":"milestone:TARGET-DRIFT-V2-CONTROLLED-EVALUATION","relation":"contains"},{"source":"group:milestones","target":"milestone:CUCB-FULL-REPAIRED-RATES","relation":"contains"},{"source":"group:milestones","target":"milestone:CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-OCCUPATION","relation":"contains"},{"source":"group:milestones","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-REGRET","relation":"contains"},{"source":"group:milestones","target":"milestone:MULTI-AGENT-STATIC-LEARNER-REGRET","relation":"contains"},{"source":"group:textbook-spine","target":"spine:13","relation":"contains"},{"source":"group:textbook-spine","target":"spine:14","relation":"contains"},{"source":"group:textbook-spine","target":"spine:15","relation":"contains"},{"source":"group:textbook-spine","target":"spine:16","relation":"contains"},{"source":"group:textbook-spine","target":"spine:17","relation":"contains"},{"source":"chapter:foundations","target":"chapter:probability","relation":"teaching order"},{"source":"chapter:probability","target":"chapter:etc","relation":"teaching order"},{"source":"chapter:etc","target":"chapter:ucb","relation":"teaching order"},{"source":"chapter:ucb","target":"chapter:oful","relation":"teaching order"},{"source":"chapter:oful","target":"chapter:thompson","relation":"teaching order"},{"source":"chapter:thompson","target":"chapter:exp3","relation":"teaching order"},{"source":"chapter:exp3","target":"chapter:tsallis","relation":"teaching order"},{"source":"chapter:tsallis","target":"chapter:finite-horizon-rl","relation":"teaching order"},{"source":"chapter:finite-horizon-rl","target":"chapter:frontier","relation":"teaching order"},{"source":"spine:13","target":"spine:14","relation":"source order"},{"source":"spine:14","target":"spine:15","relation":"source order"},{"source":"spine:15","target":"spine:16","relation":"source order"},{"source":"spine:16","target":"spine:17","relation":"source order"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ScalarENNReal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ScalarPseudoRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Regret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.CurvatureNoiseGapGeometry","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RegretDecomposition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RegretCountBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasureFoundation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.IntegrabilitySums","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasurableSums","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasurableRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RatMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HistoryFiltration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MartingaleDifference","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.BudgetStoppingTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RewardKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.PosteriorKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardFoundation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationFoundation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationSums","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.AffinityKL","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CodingEntropyBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HuffmanStep","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CrossEntropy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FixedLengthCoding","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.UniformCoding","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CommonDomination","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.SuccinctGeometryAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HighProbability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationWeightedPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCountBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationWeightedPullCountBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationFiniteBanditBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationVariance","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.UCBSummability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailClippedTransfer","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOOModel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOOCantorModel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOOLevels","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOOPacking","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOODimension","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOOPartition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOODepthOptimization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORewardFamily","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HOOCantorRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailPowerSum","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3Potential","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableHedge","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ExplorationBias","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticMaximal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticScheduled","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableRegretAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedRegretAllTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Exp3UniformRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.FTRLOneStep","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.KernelTrajectoryPrefix","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULEllipticalPotentialFoundation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianMixture","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCTrace","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.MeasureL2Indicator","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.ActionLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapHalfSet","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.MultiRegimeContract","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Literature","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Automation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.OpenProblems","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSConditionalReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConfidence","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretTail","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.FiniteGapCutoff","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.PowerTailIntegral","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.PowerCutoffNormalization","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalAllocationRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalParallelRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRealized","target":"module:BanditRLProof","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"module:BanditRLProof.Algorithms.ArmStreamPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRewardKernel","target":"module:BanditRLProof.Algorithms.CUCBActualReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBThreshold","target":"module:BanditRLProof.Algorithms.CUCBCharge","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"module:BanditRLProof.Algorithms.CUCBCharge","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConditional","target":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedMGF","target":"module:BanditRLProof.Algorithms.CUCBChargedConditional","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBCharge","target":"module:BanditRLProof.Algorithms.CUCBChargedMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","target":"module:BanditRLProof.Algorithms.CUCBChargedMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","target":"module:BanditRLProof.Algorithms.CUCBConcentration","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRoundMGF","target":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationConditionalMGF","target":"module:BanditRLProof.Algorithms.CUCBConditionalMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConcentration","target":"module:BanditRLProof.Algorithms.CUCBConfidence","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.Algorithms.CUCBConfidence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBChargedConcentration","target":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBDeterministicTrigger","target":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBNiceEvent","target":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"module:BanditRLProof.Algorithms.CUCBFiniteConcavity","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","target":"module:BanditRLProof.Algorithms.CUCBFiniteDeterministicExample","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"module:BanditRLProof.Algorithms.CUCBFiniteExample","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"module:BanditRLProof.Algorithms.CUCBFiniteExample","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFiniteExample","target":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","target":"module:BanditRLProof.Algorithms.CUCBFiniteSourceExample","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","target":"module:BanditRLProof.Algorithms.CUCBGapCutoff","relation":"imports"},{"source":"module:BanditRLProof.FiniteGapCutoff","target":"module:BanditRLProof.Algorithms.CUCBGapCutoff","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"module:BanditRLProof.Algorithms.CUCBGapInverse","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBConfidence","target":"module:BanditRLProof.Algorithms.CUCBNiceEvent","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBTrajectory","target":"module:BanditRLProof.Algorithms.CUCBObservationMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleMeasurable","target":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","target":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","relation":"imports"},{"source":"module:BanditRLProof.PowerTailIntegral","target":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBPolynomialIntegral","target":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","relation":"imports"},{"source":"module:BanditRLProof.PowerCutoffNormalization","target":"module:BanditRLProof.Algorithms.CUCBPolynomialRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBGapCutoff","target":"module:BanditRLProof.Algorithms.CUCBPolynomialThreshold","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","target":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretTail","target":"module:BanditRLProof.Algorithms.CUCBRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBActualReward","target":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","target":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"module:BanditRLProof.Algorithms.CUCBRegretTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBOracleSuccess","target":"module:BanditRLProof.Algorithms.CUCBRewardKernel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"module:BanditRLProof.Algorithms.CUCBRoundMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBFeedbackModel","target":"module:BanditRLProof.Algorithms.CUCBSourceModel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBThresholdTail","target":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBImpossibleCase","target":"module:BanditRLProof.Algorithms.CUCBSufficientSampling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBSourceModel","target":"module:BanditRLProof.Algorithms.CUCBThresholdTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBHistory","target":"module:BanditRLProof.Algorithms.CUCBTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBObservationMGF","target":"module:BanditRLProof.Algorithms.CUCBTriggerMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBGapInverse","target":"module:BanditRLProof.Algorithms.CUCBUnderCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBRegretDecomposition","target":"module:BanditRLProof.Algorithms.CUCBUnderCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CUCBUnderCount","target":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","relation":"imports"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"module:BanditRLProof.Algorithms.CUCBUnderCountIntegral","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalImportance","target":"module:BanditRLProof.Algorithms.CausalAllocation","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalExpectedRegret","target":"module:BanditRLProof.Algorithms.CausalAllocationRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalSampleMGF","target":"module:BanditRLProof.Algorithms.CausalConfidence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalTuning","target":"module:BanditRLProof.Algorithms.CausalConfidence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalRecommendation","target":"module:BanditRLProof.Algorithms.CausalExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"module:BanditRLProof.Algorithms.CausalHeterogeneous","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalHeterogeneous","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalAllocationRegret","target":"module:BanditRLProof.Algorithms.CausalHeterogeneousSampling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalMarginalLaw","target":"module:BanditRLProof.Algorithms.CausalImportance","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"module:BanditRLProof.Algorithms.CausalImportanceTransport","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalOrderedLaw","target":"module:BanditRLProof.Algorithms.CausalMarginalLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalAllocation","target":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"module:BanditRLProof.Algorithms.CausalParallelDesign","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalParallelDesign","target":"module:BanditRLProof.Algorithms.CausalParallelLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalParallelLaw","target":"module:BanditRLProof.Algorithms.CausalParallelRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalImportanceTransport","target":"module:BanditRLProof.Algorithms.CausalParallelRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalAllocationRegret","target":"module:BanditRLProof.Algorithms.CausalParallelRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalConfidence","target":"module:BanditRLProof.Algorithms.CausalRecommendation","relation":"imports"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"module:BanditRLProof.Algorithms.CausalRecommendation","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalSampling","target":"module:BanditRLProof.Algorithms.CausalSampleMGF","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"module:BanditRLProof.Algorithms.CausalSampleMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.CausalOptimalAllocation","target":"module:BanditRLProof.Algorithms.CausalSampling","relation":"imports"},{"source":"module:BanditRLProof.Regret","target":"module:BanditRLProof.Algorithms.ETC","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","relation":"imports"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","relation":"imports"},{"source":"module:BanditRLProof.MartingaleDifference","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.ETCBoundedRewardSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardIndependence","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","target":"module:BanditRLProof.Algorithms.ETCCenteredDiffSubGaussianWitnesses","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffRewardSubGaussian","target":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","target":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","relation":"imports"},{"source":"module:BanditRLProof.HistoryFiltration","target":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.Algorithms.ETCCountLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"module:BanditRLProof.Algorithms.ETCCountLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","relation":"imports"},{"source":"module:BanditRLProof.HistoryFiltration","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","relation":"imports"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","relation":"imports"},{"source":"module:BanditRLProof.RatMeasurability","target":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","target":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","target":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCondSubGaussianWitnesses","target":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.MeasurableRegret","target":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","target":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCMeasurability","target":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"module:BanditRLProof.Algorithms.ETCGeneratedHistoryPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedRegretAssembly","target":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCBoundedRewardInfinitePiSource","target":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMeanMeasurability","target":"module:BanditRLProof.Algorithms.ETCInfinitePiExpectedRegretAssembly","relation":"imports"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof.Algorithms.ETCMeasurability","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"module:BanditRLProof.Algorithms.ETCMeasurability","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","target":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","target":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof.Algorithms.ETCPairwiseCenteredSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","target":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.ETCPairwiseSubGaussianTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCArgmaxOracle","target":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"module:BanditRLProof.Algorithms.ETCPairwiseTailContract","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExactSubGaussianTail","target":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","relation":"imports"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"module:BanditRLProof.Algorithms.ETCRatArmLawRealKernel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","target":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","relation":"imports"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCExpectedPullCount","target":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealArgmaxTie","target":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","relation":"imports"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","relation":"imports"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"module:BanditRLProof.Algorithms.ETCRealLMLCompat","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealPrefixLawTransport","target":"module:BanditRLProof.Algorithms.ETCRealSourceAdapter","relation":"imports"},{"source":"module:BanditRLProof.RegretCountBounds","target":"module:BanditRLProof.Algorithms.ETCRegretLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","target":"module:BanditRLProof.Algorithms.ETCRegretLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCEmpiricalMean","target":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.Algorithms.ETCSumRewardsDiff","relation":"imports"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof.Algorithms.ETCTrace","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"module:BanditRLProof.Algorithms.ETCTrace","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCTrace","target":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"module:BanditRLProof.Algorithms.ETCTraceCountLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCenteredDiffCanonicalTail","target":"module:BanditRLProof.Algorithms.ETCWrongCommitCanonicalTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRegretLemmas","target":"module:BanditRLProof.Algorithms.ETCWrongCommitRegretAssembly","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORate","target":"module:BanditRLProof.Algorithms.HOOActualRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOConditionalMGF","target":"module:BanditRLProof.Algorithms.HOOConcentration","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"module:BanditRLProof.Algorithms.HOOConditionalMGF","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationConditionalMGF","target":"module:BanditRLProof.Algorithms.HOOConditionalMGF","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOConcentration","target":"module:BanditRLProof.Algorithms.HOOConfidence","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.Algorithms.HOOConfidence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORegretAlgebra","target":"module:BanditRLProof.Algorithms.HOODepthOptimization","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOORegretPartition","target":"module:BanditRLProof.Algorithms.HOOExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOSelectionTail","target":"module:BanditRLProof.Algorithms.HOOExpectedVisits","relation":"imports"},{"source":"module:BanditRLProof.HOOTailSum","target":"module:BanditRLProof.Algorithms.HOOExpectedVisits","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"module:BanditRLProof.Algorithms.HOOExpectedVisits","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"module:BanditRLProof.Algorithms.HOOHistory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOConfidence","target":"module:BanditRLProof.Algorithms.HOOIndexConfidence","relation":"imports"},{"source":"module:BanditRLProof.HOOModel","target":"module:BanditRLProof.Algorithms.HOOIndexConfidence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"module:BanditRLProof.Algorithms.HOOMeasurable","relation":"imports"},{"source":"module:BanditRLProof.HOOOptimalBranch","target":"module:BanditRLProof.Algorithms.HOOPathComparison","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOPrefix","target":"module:BanditRLProof.Algorithms.HOOPathComparison","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOHistory","target":"module:BanditRLProof.Algorithms.HOOPrefix","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOODepthOptimization","target":"module:BanditRLProof.Algorithms.HOORate","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedRegret","target":"module:BanditRLProof.Algorithms.HOORegretAlgebra","relation":"imports"},{"source":"module:BanditRLProof.HOOPartition","target":"module:BanditRLProof.Algorithms.HOORegretPartition","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"module:BanditRLProof.Algorithms.HOORegretPartition","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"module:BanditRLProof.Algorithms.HOORewardFamily","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOIndexConfidence","target":"module:BanditRLProof.Algorithms.HOOSelectionTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOPathComparison","target":"module:BanditRLProof.Algorithms.HOOSelectionTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOMeasurable","target":"module:BanditRLProof.Algorithms.HOOTrajectory","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailArmLaw","target":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailGapThreshold","target":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailAdaptive","target":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTailSum","target":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ArmStreamPolicy","target":"module:BanditRLProof.Algorithms.HeavyTailHistory","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"module:BanditRLProof.Algorithms.HeavyTailHistory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"module:BanditRLProof.Algorithms.HeavyTailRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"module:BanditRLProof.Algorithms.HeavyTailRegret","relation":"imports"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"module:BanditRLProof.Algorithms.HeavyTailRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegret","target":"module:BanditRLProof.Algorithms.HeavyTailRegretCap","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailSourceGap","target":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","target":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","relation":"imports"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"module:BanditRLProof.Algorithms.HeavyTailSourceCounterexample","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceAdaptive","target":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailExpectedCount","target":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailSourceSchedule","target":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailArmLaw","target":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourceExpectedCount","target":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailRegret","target":"module:BanditRLProof.Algorithms.HeavyTailSourceRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailHistory","target":"module:BanditRLProof.Algorithms.HeavyTailUCB","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","target":"module:BanditRLProof.Algorithms.KLUCBGeneratedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"module:BanditRLProof.Algorithms.MOSS","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","target":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","target":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"module:BanditRLProof.Algorithms.MOSSCanonicalReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSRewardBranch","target":"module:BanditRLProof.Algorithms.MOSSConditionalReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"module:BanditRLProof.Algorithms.MOSSConstants","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSOccupancy","target":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationIndexOccupancy","target":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSConstants","target":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","target":"module:BanditRLProof.Algorithms.MOSSExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"module:BanditRLProof.Algorithms.MOSSHistory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"module:BanditRLProof.Algorithms.MOSSHistory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"module:BanditRLProof.Algorithms.MOSSHistory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSConditionalReward","target":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryLaw","target":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSOptimism","target":"module:BanditRLProof.Algorithms.MOSSOccupancy","relation":"imports"},{"source":"module:BanditRLProof.PullCountReindex","target":"module:BanditRLProof.Algorithms.MOSSOccupancy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSPeeling","target":"module:BanditRLProof.Algorithms.MOSSOptimism","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationTailIntegration","target":"module:BanditRLProof.Algorithms.MOSSOptimism","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSS","target":"module:BanditRLProof.Algorithms.MOSSPeeling","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"module:BanditRLProof.Algorithms.MOSSPeeling","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationDyadicExponential","target":"module:BanditRLProof.Algorithms.MOSSPeeling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSStream","target":"module:BanditRLProof.Algorithms.MOSSRegret","relation":"imports"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"module:BanditRLProof.Algorithms.MOSSRegret","relation":"imports"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"module:BanditRLProof.Algorithms.MOSSRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","target":"module:BanditRLProof.Algorithms.MOSSRewardBranch","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSExpectedOccupancy","target":"module:BanditRLProof.Algorithms.MOSSStream","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSRegret","target":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSHistory","target":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","relation":"imports"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"module:BanditRLProof.Algorithms.MOSSStreamMeasurable","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSCanonicalHistory","target":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"module:BanditRLProof.Algorithms.MOSSUnusedCoordinate","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsCollision","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsReward","target":"module:BanditRLProof.Algorithms.MusicalChairsCollision","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordination","target":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsExploration","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"module:BanditRLProof.Algorithms.MusicalChairsExploration","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsRanking","target":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","target":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsHandoff","target":"module:BanditRLProof.Algorithms.MusicalChairsMarginal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCollision","target":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsRanking","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsPopulation","target":"module:BanditRLProof.Algorithms.MusicalChairsRanking","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsRealized","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationRegret","target":"module:BanditRLProof.Algorithms.MusicalChairsRealized","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsLearnerRegret","target":"module:BanditRLProof.Algorithms.MusicalChairsRealized","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsCoordinationTime","target":"module:BanditRLProof.Algorithms.MusicalChairsReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MusicalChairsExploration","target":"module:BanditRLProof.Algorithms.MusicalChairsReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditCorollaryOne","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditExponentialAudit","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremFourContractAudit","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoLatentReward","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","relation":"imports"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","relation":"imports"},{"source":"module:BanditRLProof.KernelTrajectoryPrefix","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativeTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCount","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNthPull","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoNativePrefix","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoSelectedIID","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","relation":"imports"},{"source":"module:BanditRLProof.IntegrabilitySums","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","relation":"imports"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTheoremTwoStarvation","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditAudit","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmInitialRecurrence","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmMeasurableRecurrence","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTrajectoryAudit","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRate","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditConditionalExponentialAudit","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmRecurrence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmFixedIID","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmTheoremOne","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmPathIntegrability","target":"module:BanditRLProof.Algorithms.StochasticGradientBanditTwoArmUnconditionalRecurrence","relation":"imports"},{"source":"module:BanditRLProof.Regret","target":"module:BanditRLProof.Algorithms.Thompson","relation":"imports"},{"source":"module:BanditRLProof.PosteriorKernel","target":"module:BanditRLProof.Algorithms.Thompson","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","target":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensity","target":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonAlgorithmDensityProcess","target":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonBayesRegretDecomposition","target":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","relation":"imports"},{"source":"module:BanditRLProof.PullCountDecomposition","target":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonMeasurableTrajectory","target":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalSampler","target":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","relation":"imports"},{"source":"module:BanditRLProof.HistoryFiltration","target":"module:BanditRLProof.Algorithms.ThompsonReferencePolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonClippedUCBScore","target":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"module:BanditRLProof.Algorithms.ThompsonStationaryReward","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationVariance","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.ExpectationSums","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCount","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.Regret","target":"module:BanditRLProof.Algorithms.UCB","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","target":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCFiniteArmRewardLaw","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealInfinitePiTail","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.RealKernelRegretPullCount","target":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamAsymptotics","target":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","relation":"imports"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","target":"module:BanditRLProof.Algorithms.UCBArmStreamSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"module:BanditRLProof.Algorithms.UCBArmStreamTail","relation":"imports"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"module:BanditRLProof.Algorithms.UCBArmStreamTail","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Algorithms.UCBArmStreamTail","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"module:BanditRLProof.Algorithms.UCBArmwiseBoundedFiniteArmSampledAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","target":"module:BanditRLProof.Algorithms.UCBBoundedFiniteArmSampledAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernel","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","relation":"imports"},{"source":"module:BanditRLProof.ExpectationRegretPullCount","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamProcess","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCCountLemmas","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","relation":"imports"},{"source":"module:BanditRLProof.ScalarPseudoRegret","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawTrajMeasure","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLaw","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawPolicy","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectoryReal","target":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledReal","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","relation":"imports"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"module:BanditRLProof.Algorithms.UCBContextDependentBoundedRewardKernel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","relation":"imports"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","relation":"imports"},{"source":"module:BanditRLProof.FiniteContextVarianceProxy","target":"module:BanditRLProof.Algorithms.UCBContextDependentSubGaussianRewardKernel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawCenteredKernelReal","target":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectorySampledAsymptotics","target":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"module:BanditRLProof.Algorithms.UCBFiniteArmSubGaussianSampledAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","target":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.Algorithms.UCBFixedCountPeeling","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardPairTrajectory","target":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBConditionalRewardLawRegret","target":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","target":"module:BanditRLProof.Algorithms.UCBFixedPolicyTelescopingAnytimeRegret","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealHistoryScore","target":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","relation":"imports"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"module:BanditRLProof.Algorithms.UCBRealHistoryIndex","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamExpectedPullCount","target":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryCanonicalKernelTrajectory","target":"module:BanditRLProof.Algorithms.UCBRealStationaryExplicitPolicy","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealLMLCompat","target":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamFiniteArmRewardLaws","target":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBRealStationaryFiniteArmRewardLaws","target":"module:BanditRLProof.Algorithms.UCBRealStationaryMeasurePreservingSource","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamConditionalReward","target":"module:BanditRLProof.Algorithms.UCBRealStationarySelectedRewardConsistency","relation":"imports"},{"source":"module:BanditRLProof.Literature","target":"module:BanditRLProof.Automation","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.BoundedRewardKernelLaw","relation":"imports"},{"source":"module:BanditRLProof.RewardKernel","target":"module:BanditRLProof.BoundedRewardKernelLaw","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"module:BanditRLProof.ConcentrationCappedOccupancy","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"module:BanditRLProof.ConcentrationConditionalMGF","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","target":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationGaussianOccupancy","target":"module:BanditRLProof.ConcentrationIndexOccupancy","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationMartingaleMaximal","target":"module:BanditRLProof.ConcentrationIndexOccupancy","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationCappedOccupancy","target":"module:BanditRLProof.ConcentrationIndexOccupancy","relation":"imports"},{"source":"module:BanditRLProof.MartingaleDifference","target":"module:BanditRLProof.ConcentrationMartingaleMaximal","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"module:BanditRLProof.ConcentrationQuadraticMaximal","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.ConcentrationQuadraticMaximal","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"module:BanditRLProof.ConcentrationQuadraticScheduled","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"module:BanditRLProof.ConcentrationSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"module:BanditRLProof.ConcentrationSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.ConcentrationSubGaussian","relation":"imports"},{"source":"module:BanditRLProof.RewardKernel","target":"module:BanditRLProof.ConditionalExpectationReward","relation":"imports"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"module:BanditRLProof.ConditionalExpectationReward","relation":"imports"},{"source":"module:BanditRLProof.Regret","target":"module:BanditRLProof.ConditionalExpectationReward","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"module:BanditRLProof.ConditionalRewardFoundation","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.ConditionalRewardLawSource","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.ConditionalRewardLawSource","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.ConditionalRewardLawSource","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFintypeGeometricAllTime","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryGeometricAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardLawSource","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryLaw","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","relation":"imports"},{"source":"module:BanditRLProof.ConditionalRewardPartialTrajectoryMaskedLaw","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFintypeTelescopingAllTime","target":"module:BanditRLProof.ConditionalRewardPartialTrajectoryTelescopingAllTime","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"module:BanditRLProof.DelayedFeedback.ActionLaw","relation":"imports"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"module:BanditRLProof.DelayedFeedback.ActionLaw","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.CausalView","target":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"module:BanditRLProof.DelayedFeedback.CausalView","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"module:BanditRLProof.DelayedFeedback.EliminatedArmInitialization","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.ActiveAllocation","target":"module:BanditRLProof.DelayedFeedback.Elimination","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","target":"module:BanditRLProof.DelayedFeedback.OrderedNoSwitchTrace","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Processing","target":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","target":"module:BanditRLProof.DelayedFeedback.OrderedProcessingTransition","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","target":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Accounting","target":"module:BanditRLProof.DelayedFeedback.Processing","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.ProcessedPrefixCounts","target":"module:BanditRLProof.DelayedFeedback.RecursiveProcessedState","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","target":"module:BanditRLProof.DelayedFeedback.StochasticGapOrderingAudit","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.Elimination","target":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","relation":"imports"},{"source":"module:BanditRLProof.DelayedFeedback.StochasticGoodEvent","target":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.DelayedFeedback.StochasticGoodEventAssembly","relation":"imports"},{"source":"module:BanditRLProof.Exp3ConditionalMoments","target":"module:BanditRLProof.Exp3ActionProcess","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"module:BanditRLProof.Exp3BernsteinAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"module:BanditRLProof.Exp3BernsteinExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3PureBernstein","target":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","target":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinRealizedHighProbabilityRegret","target":"module:BanditRLProof.Exp3BernsteinTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"module:BanditRLProof.Exp3BestArm","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"module:BanditRLProof.Exp3ComparatorBernstein","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"module:BanditRLProof.Exp3ComparatorBernstein","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"module:BanditRLProof.Exp3ComparatorConfidence","relation":"imports"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"module:BanditRLProof.Exp3ConditionalMoments","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","target":"module:BanditRLProof.Exp3DoubleVarianceSparseBestArmEventualRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"module:BanditRLProof.Exp3ExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableHedge","target":"module:BanditRLProof.Exp3ExplorationBias","relation":"imports"},{"source":"module:BanditRLProof.Exp3Potential","target":"module:BanditRLProof.Exp3HedgeRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"module:BanditRLProof.Exp3HighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3HedgeRegret","target":"module:BanditRLProof.Exp3ImportanceWeighted","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"module:BanditRLProof.Exp3MixedSquareBernstein","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"module:BanditRLProof.Exp3MixedSquareBernstein","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticFixedMGF","target":"module:BanditRLProof.Exp3MixedSquareBernstein","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BestArm","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedAllHorizon","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedBestArmAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquareBernsteinRealizedTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquareConfidence","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareConfidence","target":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernstein","target":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareBernsteinHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedDoublePredictableVarianceHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceRealizedMarkovHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceLossEnergyRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedDoublePredictableVarianceHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BestArm","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityBestArmAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsity","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedDoublePathwiseVarianceProbabilisticSparsityTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAESparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquareExponentialRealizedExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSmallLossRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsityTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovHighProbabilityRegret","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovProbabilisticSparsity","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BestArm","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityAllHorizon","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityBestArmAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedMarkovExplicitTuning","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsity","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceSparseLossRealizedPathwiseVarianceProbabilisticSparsityTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"module:BanditRLProof.Exp3MixedSquarePredictableVarianceTail","relation":"imports"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"module:BanditRLProof.Exp3PredictableAdversary","relation":"imports"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"module:BanditRLProof.Exp3PredictableHedge","relation":"imports"},{"source":"module:BanditRLProof.Exp3ExplorationBias","target":"module:BanditRLProof.Exp3PredictableIntegration","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"module:BanditRLProof.Exp3PredictableMoments","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.Exp3PredictableMoments","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"module:BanditRLProof.Exp3PredictableRegretAllTime","relation":"imports"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"module:BanditRLProof.Exp3PredictableRegretAllTime","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"module:BanditRLProof.Exp3PureBernstein","relation":"imports"},{"source":"module:BanditRLProof.Exp3PureConfidence","target":"module:BanditRLProof.Exp3PureBernstein","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorConfidence","target":"module:BanditRLProof.Exp3PureConfidence","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinAllHorizon","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedAllHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinExplicitTuning","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedExplicitTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedHighProbabilityRegret","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinTuning","target":"module:BanditRLProof.Exp3RandomSquareBernsteinRealizedTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3BernsteinHighProbabilityRegret","target":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableIntegration","target":"module:BanditRLProof.Exp3RandomSquareHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"module:BanditRLProof.Exp3RealizedConcentration","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.Exp3RealizedConcentration","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.Exp3RealizedConcentration","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedDeviationTail","target":"module:BanditRLProof.Exp3RealizedConfidence","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","target":"module:BanditRLProof.Exp3RealizedDeviationAllTime","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConcentration","target":"module:BanditRLProof.Exp3RealizedDeviationTail","relation":"imports"},{"source":"module:BanditRLProof.Exp3HighProbabilityRegret","target":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedConfidence","target":"module:BanditRLProof.Exp3RealizedHighProbabilityRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3MixedSquarePredictableVariance","target":"module:BanditRLProof.Exp3RealizedPredictableVariance","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticScheduled","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationConfidenceSchedule","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceAllTime","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationQuadraticMaximal","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceMaximal","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedPredictableVariance","target":"module:BanditRLProof.Exp3RealizedPredictableVarianceTail","relation":"imports"},{"source":"module:BanditRLProof.Exp3ExpectedRegret","target":"module:BanditRLProof.Exp3RealizedRegret","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableRegretAllTime","target":"module:BanditRLProof.Exp3RealizedRegretAllTime","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedDeviationAllTime","target":"module:BanditRLProof.Exp3RealizedRegretAllTime","relation":"imports"},{"source":"module:BanditRLProof.Exp3ScoreRegularity","target":"module:BanditRLProof.Exp3RecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonRecursiveSampler","target":"module:BanditRLProof.Exp3RecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableMoments","target":"module:BanditRLProof.Exp3SampledHedge","relation":"imports"},{"source":"module:BanditRLProof.Exp3RecursiveTrajectory","target":"module:BanditRLProof.Exp3SampledHistoryScore","relation":"imports"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"module:BanditRLProof.Exp3ScoreRegularity","relation":"imports"},{"source":"module:BanditRLProof.Exp3RealizedRegret","target":"module:BanditRLProof.Exp3UniformRegret","relation":"imports"},{"source":"module:BanditRLProof.IntegrabilitySums","target":"module:BanditRLProof.ExpectationBochnerSums","relation":"imports"},{"source":"module:BanditRLProof.ExpectationWeightedPullCountBounds","target":"module:BanditRLProof.ExpectationFiniteBanditBounds","relation":"imports"},{"source":"module:BanditRLProof.ExpectationFiniteBanditBounds","target":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","relation":"imports"},{"source":"module:BanditRLProof.MeasureFoundation","target":"module:BanditRLProof.ExpectationFoundation","relation":"imports"},{"source":"module:BanditRLProof.ExpectationFiniteBanditModelBounds","target":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","relation":"imports"},{"source":"module:BanditRLProof.ScalarPseudoRegret","target":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPseudoRegretOfRealBounds","target":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof.ExpectationPseudoRegretRatBounds","relation":"imports"},{"source":"module:BanditRLProof.ExpectationSums","target":"module:BanditRLProof.ExpectationPullCount","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof.ExpectationPullCount","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"module:BanditRLProof.ExpectationPullCountBounds","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.ExpectationRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"module:BanditRLProof.ExpectationRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.RegretDecomposition","target":"module:BanditRLProof.ExpectationRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.ExpectationFoundation","target":"module:BanditRLProof.ExpectationSums","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCount","target":"module:BanditRLProof.ExpectationWeightedPullCount","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCountCast","target":"module:BanditRLProof.ExpectationWeightedPullCount","relation":"imports"},{"source":"module:BanditRLProof.ExpectationWeightedPullCount","target":"module:BanditRLProof.ExpectationWeightedPullCountBounds","relation":"imports"},{"source":"module:BanditRLProof.ExpectationPullCountBounds","target":"module:BanditRLProof.ExpectationWeightedPullCountBounds","relation":"imports"},{"source":"module:BanditRLProof.BoundedRewardKernelLaw","target":"module:BanditRLProof.FiniteArmRewardKernelLaw","relation":"imports"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof.FiniteBanditModelInvariants","relation":"imports"},{"source":"module:BanditRLProof.FiniteArmRewardKernelLaw","target":"module:BanditRLProof.FiniteContextVarianceProxy","relation":"imports"},{"source":"module:BanditRLProof.FiniteGapLayerCake","target":"module:BanditRLProof.FiniteGapCutoff","relation":"imports"},{"source":"module:BanditRLProof.MeasurableLocalQuantities","target":"module:BanditRLProof.FiniteRealArgmax","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOExpectedVisits","target":"module:BanditRLProof.HOOCantorModel","relation":"imports"},{"source":"module:BanditRLProof.HOOCantorModel","target":"module:BanditRLProof.HOOCantorRate","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOActualRegret","target":"module:BanditRLProof.HOOCantorRate","relation":"imports"},{"source":"module:BanditRLProof.HOOPacking","target":"module:BanditRLProof.HOODimension","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOTree","target":"module:BanditRLProof.HOOGeometry","relation":"imports"},{"source":"module:BanditRLProof.HOOModel","target":"module:BanditRLProof.HOOLevels","relation":"imports"},{"source":"module:BanditRLProof.HOOGeometry","target":"module:BanditRLProof.HOOModel","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HOOTrajectory","target":"module:BanditRLProof.HOOModel","relation":"imports"},{"source":"module:BanditRLProof.HOOModel","target":"module:BanditRLProof.HOOOptimalBranch","relation":"imports"},{"source":"module:BanditRLProof.HOOLevels","target":"module:BanditRLProof.HOOPacking","relation":"imports"},{"source":"module:BanditRLProof.HOODimension","target":"module:BanditRLProof.HOOPartition","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTailSum","target":"module:BanditRLProof.HOOTailSum","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailScheduledConfidence","target":"module:BanditRLProof.HeavyTailArmLaw","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamTail","target":"module:BanditRLProof.HeavyTailArmLaw","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailClippedMoments","target":"module:BanditRLProof.HeavyTailClippedConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"module:BanditRLProof.HeavyTailClippedConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof.HeavyTailClippedConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailClipping","target":"module:BanditRLProof.HeavyTailClippedMoments","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"module:BanditRLProof.HeavyTailClippedMoments","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof.HeavyTailClippedScheduled","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailClippedConfidence","target":"module:BanditRLProof.HeavyTailClippedScheduled","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailClippedScheduled","target":"module:BanditRLProof.HeavyTailClippedTransfer","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailArmLaw","target":"module:BanditRLProof.HeavyTailClippedTransfer","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"module:BanditRLProof.HeavyTailClipping","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"module:BanditRLProof.HeavyTailConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTruncation","target":"module:BanditRLProof.HeavyTailFixedTilt","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"module:BanditRLProof.HeavyTailFixedTilt","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof.HeavyTailGapThreshold","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof.HeavyTailScheduledConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"module:BanditRLProof.HeavyTailScheduledConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailUnshiftedMGF","target":"module:BanditRLProof.HeavyTailSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailConfidence","target":"module:BanditRLProof.HeavyTailSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof.HeavyTailSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailSourcePolicy","target":"module:BanditRLProof.HeavyTailSourceGap","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailSourceConfidence","target":"module:BanditRLProof.HeavyTailSourceSchedule","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailTuning","target":"module:BanditRLProof.HeavyTailTailSum","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCBArmStreamSource","target":"module:BanditRLProof.HeavyTailTruncation","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailPowerSum","target":"module:BanditRLProof.HeavyTailTuning","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.HeavyTailUCB","target":"module:BanditRLProof.HeavyTailTuning","relation":"imports"},{"source":"module:BanditRLProof.HeavyTailFixedTilt","target":"module:BanditRLProof.HeavyTailUnshiftedMGF","relation":"imports"},{"source":"module:BanditRLProof.MeasureFoundation","target":"module:BanditRLProof.HistoryFiltration","relation":"imports"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof.IndependenceFoundation","relation":"imports"},{"source":"module:BanditRLProof.Regret","target":"module:BanditRLProof.LeafLemmas","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETC","target":"module:BanditRLProof.Literature","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.UCB","target":"module:BanditRLProof.Literature","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.Thompson","target":"module:BanditRLProof.Literature","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","target":"module:BanditRLProof.LowerBounds.AffinityKL","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","target":"module:BanditRLProof.LowerBounds.ArithmeticBlockCoding","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BlockEntropy","target":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticPrefixCode","target":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HuffmanConstruction","target":"module:BanditRLProof.LowerBounds.ArithmeticZeroExtension","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"module:BanditRLProof.LowerBounds.BanditHistoryDataProcessing","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ConditionalKernelKL","target":"module:BanditRLProof.LowerBounds.BanditHistoryKL","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"module:BanditRLProof.LowerBounds.BanditHistoryKL","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"module:BanditRLProof.LowerBounds.BlockEntropy","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"module:BanditRLProof.LowerBounds.CodingEntropyBound","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"module:BanditRLProof.LowerBounds.CommonDensityKL","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CommonDensityKL","target":"module:BanditRLProof.LowerBounds.CommonDensityOverlap","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"module:BanditRLProof.LowerBounds.CommonDomination","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"module:BanditRLProof.LowerBounds.CrossEntropy","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ArithmeticIntervals","target":"module:BanditRLProof.LowerBounds.DyadicAddresses","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FiniteDiscreteKL","target":"module:BanditRLProof.LowerBounds.FinitePartitionKL","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FinitePartitionKL","target":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","target":"module:BanditRLProof.LowerBounds.FinitePartitionKLRecovery","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.DyadicAddresses","target":"module:BanditRLProof.LowerBounds.FixedLengthCoding","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianMillsRatio","target":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"module:BanditRLProof.LowerBounds.GaussianMinimax","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BasicIdeas","target":"module:BanditRLProof.LowerBounds.GaussianMinimax","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"module:BanditRLProof.LowerBounds.GaussianMinimax","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"module:BanditRLProof.LowerBounds.GaussianTesting","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InstanceDependent","target":"module:BanditRLProof.LowerBounds.HighProbability","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HuffmanStep","target":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.HuffmanAlphabet","target":"module:BanditRLProof.LowerBounds.HuffmanConstruction","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","target":"module:BanditRLProof.LowerBounds.HuffmanStep","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.KLUCBBernoulli","target":"module:BanditRLProof.LowerBounds.InformationTheory","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.Minimax","target":"module:BanditRLProof.LowerBounds.InstanceDependent","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.BanditHistoryKL","target":"module:BanditRLProof.LowerBounds.InstanceDependent","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianMinimax","target":"module:BanditRLProof.LowerBounds.InstanceDependent","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"module:BanditRLProof.LowerBounds.Minimax","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.ShannonLengths","target":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeConstruction","target":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodePruning","target":"module:BanditRLProof.LowerBounds.PrefixCodeGreedy","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","target":"module:BanditRLProof.LowerBounds.PrefixCodePruning","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"module:BanditRLProof.LowerBounds.PrefixCodeSiblings","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.InformationTheory","target":"module:BanditRLProof.LowerBounds.RelativeEntropyFiltration","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianTesting","target":"module:BanditRLProof.LowerBounds.RelativeEntropyNonMetric","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.CodingEntropyBound","target":"module:BanditRLProof.LowerBounds.ShannonLengths","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.MOSSHistoryRegret","target":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.GaussianHypothesisTesting","target":"module:BanditRLProof.LowerBounds.SubgaussianMinimax","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.FixedLengthCoding","target":"module:BanditRLProof.LowerBounds.UniformCoding","relation":"imports"},{"source":"module:BanditRLProof.LowerBounds.PrefixCodeExchange","target":"module:BanditRLProof.LowerBounds.UniformCoding","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof.MathlibWrappers","relation":"imports"},{"source":"module:BanditRLProof.MeasurableSums","target":"module:BanditRLProof.MeasurableLocalQuantities","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.MeasurableLocalQuantities","relation":"imports"},{"source":"module:BanditRLProof.MeasureFoundation","target":"module:BanditRLProof.MeasurablePullCount","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof.MeasurablePullCount","relation":"imports"},{"source":"module:BanditRLProof.MeasurablePullCount","target":"module:BanditRLProof.MeasurablePullCountCast","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.MeasurableRegret","relation":"imports"},{"source":"module:BanditRLProof.MeasureFoundation","target":"module:BanditRLProof.MeasurableSums","relation":"imports"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof.MeasureFoundation","relation":"imports"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"module:BanditRLProof.OFULAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULMeasurableRecursiveSelection","target":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","relation":"imports"},{"source":"module:BanditRLProof.OFULSelfNormalizedMarkov","target":"module:BanditRLProof.OFULConfidenceEllipsoid","relation":"imports"},{"source":"module:BanditRLProof.OFULEllipticalPotential","target":"module:BanditRLProof.OFULEllipticalPotentialFoundation","relation":"imports"},{"source":"module:BanditRLProof.OFULInitialRoundGap","target":"module:BanditRLProof.OFULExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"module:BanditRLProof.OFULExpectedRegretAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"module:BanditRLProof.OFULExpectedRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"module:BanditRLProof.OFULExpectedRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"module:BanditRLProof.OFULFiniteActionOptimism","relation":"imports"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"module:BanditRLProof.OFULFiniteHorizonScoreGram","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianMixtureMeasurability","target":"module:BanditRLProof.OFULFiniteHorizonScoreGram","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianSpectralMixture","target":"module:BanditRLProof.OFULGaussianCovarianceMixture","relation":"imports"},{"source":"module:BanditRLProof.OFULFiniteHorizonScoreGram","target":"module:BanditRLProof.OFULGaussianEvaluatedMixture","relation":"imports"},{"source":"module:BanditRLProof.OFULSelfNormalizedConfidence","target":"module:BanditRLProof.OFULGaussianMixture","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianCovarianceMixture","target":"module:BanditRLProof.OFULGaussianMixtureMeasurability","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianMixture","target":"module:BanditRLProof.OFULGaussianSpectralMixture","relation":"imports"},{"source":"module:BanditRLProof.OFULConcreteHistoryRidgeSelection","target":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","relation":"imports"},{"source":"module:BanditRLProof.OFULUniformTimeConfidence","target":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","target":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","relation":"imports"},{"source":"module:BanditRLProof.HistoryFiltration","target":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","relation":"imports"},{"source":"module:BanditRLProof.OFULSelectedWidthSummation","target":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryConfidenceGap","target":"module:BanditRLProof.OFULGeneratedTrajectoryUniformConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"module:BanditRLProof.OFULHighProbabilityRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryPredictableConfidence","target":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"module:BanditRLProof.OFULInitialRoundGap","relation":"imports"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"module:BanditRLProof.OFULMeasurableRecursiveSelection","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ETCRealEmpiricalMean","target":"module:BanditRLProof.OFULMeasurableRecursiveSelection","relation":"imports"},{"source":"module:BanditRLProof.Algorithms.ThompsonCanonicalTrajectory","target":"module:BanditRLProof.OFULMeasurableRecursiveSelection","relation":"imports"},{"source":"module:BanditRLProof.OFULGeneratedTrajectoryRadiusWidth","target":"module:BanditRLProof.OFULNormalizedRadiusWidth","relation":"imports"},{"source":"module:BanditRLProof.OFULConfidenceEllipsoid","target":"module:BanditRLProof.OFULScalarRegularizationBias","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","target":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","relation":"imports"},{"source":"module:BanditRLProof.OFULHighProbabilityRegretRate","target":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllTimeConfidence","target":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","relation":"imports"},{"source":"module:BanditRLProof.OFULNormalizedRadiusWidth","target":"module:BanditRLProof.OFULScheduledAllHorizonCumulativeGap","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonAllRoundGap","target":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULAllTimeConfidence","target":"module:BanditRLProof.OFULScheduledAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULHistoryEnvironmentRewardLaw","target":"module:BanditRLProof.OFULScheduledAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","target":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","target":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonIndexedHighProbabilityRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedActionChargeBound","target":"module:BanditRLProof.OFULScheduledBlockStartForcedHorizonWindowFiniteHorizonTail","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAEAlignedWindowPositiveActionCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledBlockStartForcedPositiveActionCostBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"module:BanditRLProof.OFULScheduledBlockStartForcedPseudoRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegret","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretAsymptotics","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretRate","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledAllHorizonHighProbabilityRegretRate","target":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeHighProbabilityRegretRate","relation":"imports"},{"source":"module:BanditRLProof.BudgetStoppingTime","target":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","target":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledCumulativeAlignedWindowPositiveCostBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledCumulativePositiveCostBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledPositiveActionCostBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedAllTimeConfidence","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageConsistency","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULExpectedRegretConsistency","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityAverageRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegretAsymptotics","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHighProbabilityRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBlockStartForcedHistoryAlgorithm","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedIndexCount","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedHistoryAlgorithm","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedPseudoRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedAllTimeConfidence","target":"module:BanditRLProof.OFULScheduledPowerOfTwoForcedScalarChargeBound","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBoundedStoppingTimeExpectedRegret","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretExactMoment","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretClosed","target":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretSecondMoment","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledBudgetExhaustionExpectedRegret","target":"module:BanditRLProof.OFULScheduledUnitGrowthBudgetExhaustionExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.OFULFiniteActionOptimism","target":"module:BanditRLProof.OFULSelectedWidthSummation","relation":"imports"},{"source":"module:BanditRLProof.OFULEllipticalPotentialFoundation","target":"module:BanditRLProof.OFULSelectedWidthSummation","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.OFULSelfNormalizedConfidence","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.OFULSelfNormalizedConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULEllipticalPotentialFoundation","target":"module:BanditRLProof.OFULSelfNormalizedConfidence","relation":"imports"},{"source":"module:BanditRLProof.OFULGaussianEvaluatedMixture","target":"module:BanditRLProof.OFULSelfNormalizedMarkov","relation":"imports"},{"source":"module:BanditRLProof.OFULScalarRegularizationBias","target":"module:BanditRLProof.OFULUniformTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.OFULUniformTimeConfidence","relation":"imports"},{"source":"module:BanditRLProof.Automation","target":"module:BanditRLProof.OpenProblems","relation":"imports"},{"source":"module:BanditRLProof.HistoryFiltration","target":"module:BanditRLProof.PolicyMeasurability","relation":"imports"},{"source":"module:BanditRLProof.PowerTailIntegral","target":"module:BanditRLProof.PowerCutoffNormalization","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.PullCountDecomposition","relation":"imports"},{"source":"module:BanditRLProof.LeafLemmas","target":"module:BanditRLProof.PullCountReindex","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEmpiricalOptimisticRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceExpectedConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationRealizedBehaviorConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseRealizedBehaviorConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeCountMartingaleConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtCalibration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtAverageConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtHighProbabilityAverageConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtExplicitRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeInverseSqrtNormalizedRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeHoeffdingUCBVI","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBellmanInnovation","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","relation":"imports"},{"source":"module:BanditRLProof.Exp3ComparatorBernstein","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIChargeSummation","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAggregateTransition","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICounting","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIEpisodeRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVICoordinateAlignment","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVILocalBellman","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimalTailAlignment","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAlignment","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIProbabilityBudget","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIOptimism","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRecurrentOptimism","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIAdaptiveBellmanMartingale","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIRegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIClippedPlanner","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationFixedMGF","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISameSourceConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIConfidenceTuning","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVISimultaneousConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIMartingaleTuning","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITerminal","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVIBernsteinConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeUCBVITransitionValueConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticSource","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","relation":"imports"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeEpisodewiseCommonSpaceL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeDecayingExplorationBehaviorConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationRegularityClosedConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","relation":"imports"},{"source":"module:BanditRLProof.RewardTraceLaw","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticOccupancyEnvelope","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeRecommendedRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretExplicitIntegratedRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretInMeasureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalExplicitRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalModelConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticCumulativeExploratoryBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","relation":"imports"},{"source":"module:BanditRLProof.KernelTrajectoryPrefix","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalSource","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationCommonSpaceConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCommonSpaceConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentExplicitRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticExplicitBudgetRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardCumulativeDecayingExplorationConsistency","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardEmpiricalOptimisticProjection","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSource","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","target":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonCoordinateModelConfidence","target":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEpisodeBatchLaw","target":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonEpisodeBatchStandardBorel","relation":"imports"},{"source":"module:BanditRLProof.FiniteRealArgmax","target":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","target":"module:BanditRLProof.RL.FiniteHorizonEstimatedModelCertificate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportEpisodeThreshold","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportReachability","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveEmpiricalOptimisticConfidence","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","target":"module:BanditRLProof.RL.FiniteHorizonExploratoryReachabilityCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","target":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","target":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","target":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleVisitCountPositivity","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDEligibleEmpiricalTransitionConfidence","target":"module:BanditRLProof.RL.FiniteHorizonIIDGeneratedEmpiricalRewardExactness","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.RL.FiniteHorizonIIDMultiBatchCumulativeConfidenceRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.RL.FiniteHorizonIIDSimultaneousCountConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonEmpiricalModel","target":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceAlmostSureBehaviorExpectedRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalCommonSpaceBehaviorExpectedRegretFinitePrefixCumulativeAverageRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretLogRate","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.MeasureL2Indicator","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitExpectedAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitThreeQuarterGoodEventAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeHighProbabilityAverageRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeSingleModelEventHighProbabilityAverageRealizedBehaviorRegret","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterL1TruncationEquivalence","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1TruncationEquivalence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationExpectedTruncationReplacement","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedRealizedBehaviorRegretAndPolicyValueExpectedConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnL1Optimality","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnInMeasureAlmostSureOptimality","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnSimultaneousHighProbabilityOptimality","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedAverageSampledReturnExpectedOptimality","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledReturnAndSuccessorPolicyExpectedReturnConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterCauchySchwarzExpectedAbsoluteAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedDoubleLinearRawWindowFirstPassageSummableDelayAndEventualImmediateStoppingL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterDelayedEventExpectedContribution","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","relation":"imports"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableExpectedUpperBound","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterAEFiniteEventualImmediateStoppingAndInMeasureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalGrowingWindowGridStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExpectedPositivePartConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterL1Consistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteAsymptotics","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterExplicitTailStartExpectedAbsoluteBound","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterPolynomialSecondMomentExpectedAbsoluteBound","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterIntegrableFiniteStoppingTime","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterSquareIntegrableFiniteStoppingTime","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterExpectedRegretTruncationReplacement","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedStoppingTimeExplicitDeterministicMomentExpectedAverageRealizedBehaviorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterStoppedBehaviorExpectedRegretAndReturnDeviationL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformAbsoluteContinuity","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterLpConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdUnboundedHittingAfterUniformIntegrabilityExpectedConsistency","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBoundedWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalPolynomialBaseGrowingRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityBurninLogRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalBehaviorExpectedRegretHighProbabilityLogRate","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonAdaptiveStochasticRewardSampledEmpiricalOptimisticSelfConsistentCausalRealizedSuccessorRegret","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityLogRate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretInMeasureExplicitSchedule","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretHighProbabilityExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRealizedBehaviorRegretUpperTailInProbability","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalCappedDoubleLinearRawWindowFirstPassageStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalAverageRealizedBehaviorRegretAlmostSureExplicitSchedule","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalReciprocalThresholdCappedDoubleLinearRawWindowFirstPassageVanishingDelayProbabilityAndL1Consistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRandomPrefixAverageRealizedBehaviorRegretAlmostSureConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalStoppingTimeAverageRealizedBehaviorRegretAlmostSureConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalRateControlledRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","target":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalThresholdTriggeredDoubleLinearRawWindowStoppingTimeL1AverageRealizedBehaviorRegretConsistency","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonOptimality","target":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"module:BanditRLProof.RL.FiniteHorizonOptimality","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonOccupancyRegret","target":"module:BanditRLProof.RL.FiniteHorizonOptimisticCertificate","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonMDP","target":"module:BanditRLProof.RL.FiniteHorizonPolicy","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDCountConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStageTransitionJointFactorization","target":"module:BanditRLProof.RL.FiniteHorizonStageVisitFactorization","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonTrajectory","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","relation":"imports"},{"source":"module:BanditRLProof.RewardKernel","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","relation":"imports"},{"source":"module:BanditRLProof.ConcentrationSubGaussian","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","relation":"imports"},{"source":"module:BanditRLProof.ConditionalExpectationReward","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConditionalLaw","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardCumulativeConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDTrajectoryBatch","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardErasureLaw","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonIIDAllCoordinateFiniteBatchConfidence","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDEmpiricalRewardConfidence","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDAllCoordinateEmpiricalModelConfidence","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonExploratoryPathSupportExplicitCalibration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDExplicitCalibration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDSelfConsistentCalibration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardIIDTotalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardInitialLawTotalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardMarginal","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellmanInnovationConcentration","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTotalReturnConcentration","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardBellman","target":"module:BanditRLProof.RL.FiniteHorizonStochasticRewardTrajectory","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonPolicy","target":"module:BanditRLProof.RL.FiniteHorizonTrajectory","relation":"imports"},{"source":"module:BanditRLProof.RL.FiniteHorizonNaturalCausalInverseSqrtThresholdCappedUnboundedHittingAfterStoppedSampledAndSuccessorPolicyReturnDeterministicTailHighProbabilityOptimality","target":"module:BanditRLProof.RL.StoppedReturnJointErrorDeterministicTailHighProbability","relation":"imports"},{"source":"module:BanditRLProof.RealMeanRegretPullCount","target":"module:BanditRLProof.RealKernelRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.RealMeanRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.IntegrabilitySums","target":"module:BanditRLProof.RealMeanRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.RealMeanRegretPullCount","relation":"imports"},{"source":"module:BanditRLProof.Core","target":"module:BanditRLProof.Regret","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof.RegretCountBounds","relation":"imports"},{"source":"module:BanditRLProof.RegretDecomposition","target":"module:BanditRLProof.RegretCountBounds","relation":"imports"},{"source":"module:BanditRLProof.MathlibWrappers","target":"module:BanditRLProof.RegretDecomposition","relation":"imports"},{"source":"module:BanditRLProof.PolicyMeasurability","target":"module:BanditRLProof.RewardKernel","relation":"imports"},{"source":"module:BanditRLProof.RewardKernel","target":"module:BanditRLProof.RewardTraceLaw","relation":"imports"},{"source":"module:BanditRLProof.ScalarENNReal","target":"module:BanditRLProof.ScalarPseudoRegret","relation":"imports"},{"source":"module:BanditRLProof.RegretDecomposition","target":"module:BanditRLProof.ScalarPseudoRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","relation":"imports"},{"source":"module:BanditRLProof.Exp3Potential","target":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"module:BanditRLProof.TsallisConjugatePotentialStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","target":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","target":"module:BanditRLProof.TsallisFTRLConditionalStability","relation":"imports"},{"source":"module:BanditRLProof.Exp3ActionProcess","target":"module:BanditRLProof.TsallisFTRLConditionalStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLConditionalStability","target":"module:BanditRLProof.TsallisFTRLExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.ExpectationBochnerSums","target":"module:BanditRLProof.TsallisFTRLExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"module:BanditRLProof.TsallisFTRLFiniteHorizonSelection","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","target":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","target":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLStationarity","target":"module:BanditRLProof.TsallisFTRLInteriority","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLInteriority","target":"module:BanditRLProof.TsallisFTRLMinimizerExistence","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","target":"module:BanditRLProof.TsallisFTRLMinimizerMeasurability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLMinimizerExistence","target":"module:BanditRLProof.TsallisFTRLMinimizerUniqueness","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"module:BanditRLProof.TsallisFTRLOneStepStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"module:BanditRLProof.TsallisFTRLOneStepStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLExpectedStability","target":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Exp3SampledHistoryScore","target":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Exp3PredictableAdversary","target":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.Exp3SampledHedge","target":"module:BanditRLProof.TsallisFTRLRecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"module:BanditRLProof.TsallisFTRLRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLOneStepStability","target":"module:BanditRLProof.TsallisFTRLStationarity","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDArmDependentSuboptimalBoostRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveCorruptedRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","target":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDHorizonHistoryAdaptiveExpectedCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"module:BanditRLProof.TsallisFiniteArmIIDMeasurableHistoryArmGatedSuboptimalBoostRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","target":"module:BanditRLProof.TsallisFiniteArmIIDPreviousActionGatedSuboptimalBoostRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingCorruptedRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDTimeVaryingSuboptimalBoostRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDHistoryAdaptiveRefinedCorruptedRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIIDUniformSuboptimalBoostRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanAllRegimes","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRefinedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","target":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","target":"module:BanditRLProof.TsallisFiniteArmIndependentMeanSwitchCountDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"module:BanditRLProof.TsallisFiniteArmIndependentPathVariationDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDRewardLaw","target":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","target":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"module:BanditRLProof.TsallisFiniteArmIndependentRewardLaw","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountCompressedDynamicRegret","target":"module:BanditRLProof.TsallisFiniteArmIndependentSingleSwitchComparatorObstruction","relation":"imports"},{"source":"module:BanditRLProof.FiniteBanditModelInvariants","target":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"module:BanditRLProof.TsallisFiniteBanditMeanLoss","relation":"imports"},{"source":"module:BanditRLProof.Exp3ImportanceWeighted","target":"module:BanditRLProof.TsallisImportanceWeightedMoment","relation":"imports"},{"source":"module:BanditRLProof.TsallisRegularizer","target":"module:BanditRLProof.TsallisImportanceWeightedMoment","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentDriftingMeanDynamicRegret","target":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","target":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","target":"module:BanditRLProof.TsallisOracleRestartExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"module:BanditRLProof.TsallisOracleRestartExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","target":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedDynamicRegret","target":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIndependentGlobalMeanSwitchCountDynamicRegret","target":"module:BanditRLProof.TsallisOracleRestartGlobalMeanSwitchCount","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartGeneratedTrajectory","target":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartDynamicRegret","target":"module:BanditRLProof.TsallisOracleRestartPredictableRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedStability","target":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","relation":"imports"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"module:BanditRLProof.TsallisOracleRestartRefinedStabilityTuning","relation":"imports"},{"source":"module:BanditRLProof.TsallisOracleRestartExpectedRegret","target":"module:BanditRLProof.TsallisOracleRestartScoreAlignment","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","target":"module:BanditRLProof.TsallisRefinedAveragedStabilityObstruction","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedRegularity","target":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","relation":"imports"},{"source":"module:BanditRLProof.TsallisImportanceWeightedMoment","target":"module:BanditRLProof.TsallisRefinedImportanceWeightedMoment","relation":"imports"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"module:BanditRLProof.TsallisRefinedSuboptimalStability","relation":"imports"},{"source":"module:BanditRLProof.FTRLOneStep","target":"module:BanditRLProof.TsallisRegularizer","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","target":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","target":"module:BanditRLProof.TsallisScheduledAllTimesExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","target":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledAllRateExpectedStability","target":"module:BanditRLProof.TsallisScheduledExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"module:BanditRLProof.TsallisScheduledExpectedRegret","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledScoreAlignment","target":"module:BanditRLProof.TsallisScheduledExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisConjugatePotentialFiniteHorizon","target":"module:BanditRLProof.TsallisScheduledExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"module:BanditRLProof.TsallisScheduledFixedGapSelfBounding","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"module:BanditRLProof.TsallisScheduledIIDHistoryAdaptive","relation":"imports"},{"source":"module:BanditRLProof.IndependenceFoundation","target":"module:BanditRLProof.TsallisScheduledIIDMeanGap","relation":"imports"},{"source":"module:BanditRLProof.KernelIndependentExtension","target":"module:BanditRLProof.TsallisScheduledIIDMeanGap","relation":"imports"},{"source":"module:BanditRLProof.KernelTrajectoryPrefix","target":"module:BanditRLProof.TsallisScheduledIIDMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","target":"module:BanditRLProof.TsallisScheduledIIDMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledIIDMeanGap","target":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"module:BanditRLProof.TsallisScheduledIIDTimeVaryingMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"module:BanditRLProof.TsallisScheduledIndependentMeanGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedStability","target":"module:BanditRLProof.TsallisScheduledInitialExpectedStability","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLGeneratedMeasurability","target":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","target":"module:BanditRLProof.TsallisScheduledReferenceGapExpectedDeviationSelfBounding","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","target":"module:BanditRLProof.TsallisScheduledReferenceGapSelfBounding","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRefinedExpectedPenalty","target":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedGapSelfBounding","target":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRecursiveTrajectory","target":"module:BanditRLProof.TsallisScheduledScoreAlignment","relation":"imports"},{"source":"module:BanditRLProof.TsallisTimeVaryingPenalty","target":"module:BanditRLProof.TsallisScheduledScoreAlignment","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","target":"module:BanditRLProof.TsallisScheduledSelfBoundingInterpolation","relation":"imports"},{"source":"module:BanditRLProof.TsallisConstrainedQuadraticOptimization","target":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledExpectedRegret","target":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","relation":"imports"},{"source":"module:BanditRLProof.TsallisRefinedSuboptimalStability","target":"module:BanditRLProof.TsallisScheduledSuboptimalExpectedBound","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledConditionalMeanGap","target":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisFiniteArmIIDCorruptedRewardLaw","target":"module:BanditRLProof.TsallisScheduledTimeVaryingExpectedGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLEstimatedEnvironmentRegret","target":"module:BanditRLProof.TsallisSelfBounding","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","target":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledRefinedStabilityPenalty","target":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","relation":"imports"},{"source":"module:BanditRLProof.TsallisScheduledSelfBoundingOptimization","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleFixedGap","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","relation":"imports"},{"source":"module:BanditRLProof.TsallisSelfBoundingBetaRoot","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedTuning","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedScalar","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingRefinedWindow","relation":"imports"},{"source":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingOptimization","target":"module:BanditRLProof.TsallisSqrtScheduleSelfBoundingTuning","relation":"imports"},{"source":"module:BanditRLProof.Exp3Potential","target":"module:BanditRLProof.TsallisTimeVaryingPenalty","relation":"imports"},{"source":"module:BanditRLProof.TsallisConjugatePotentialStability","target":"module:BanditRLProof.TsallisTimeVaryingPenalty","relation":"imports"},{"source":"module:BanditRLProof.TsallisFTRLRegret","target":"module:BanditRLProof.TsallisTimeVaryingPenalty","relation":"imports"},{"source":"module:BanditRLProof.TsallisSelfBounding","target":"module:BanditRLProof.TsallisTimeVaryingPenalty","relation":"imports"},{"source":"module:BanditRLProof.ProbabilityUnionBound","target":"module:BanditRLProof.UCBSummability","relation":"imports"},{"source":"module:BanditRLProof.MeasureL2Indicator","target":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","relation":"imports"},{"source":"module:BanditRLProof.OFULScheduledUnboundedStoppingTimeExpectedRegretRate","target":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","relation":"imports"},{"source":"module:BanditRLProof.UnboundedStoppingTimeL2CoordinateIntegrability","target":"module:BanditRLProof.UnboundedStoppingTimeWeightedL2CoordinateIntegrability","relation":"imports"},{"source":"declaration:BanditRLProof.pullCount","target":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.pseudoRegret","target":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteBanditModel","target":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ProbabilityUnionBound.measure_biUnion_finset_le_of_uniform","target":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.tsum_ofReal_geometricConfidenceShare","target":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_abs_tail_random_pullCount_ennreal_delta_trajMeasure_on_horizon","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare","target":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.UCB.meanGap_le_two_radius_of_confidenceScore_max","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","target":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.KLUCB.bernoulliKLCore_le_sq_div","target":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.KLUCB.generatedIndexAt_le_selected_of_K_le","target":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","target":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","target":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Policy.MeasurablePolicy","target":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ETC.argmaxCommitOracle","target":"declaration:BanditRLProof.ETC.argmaxCommitOracle_encode_le_of_score_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","target":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","target":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le","target":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","target":"declaration:BanditRLProof.OFUL.fixedDirectionCompensatedScore_hasMGFUpperBoundAt","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.fixedDirectionCompensatedScore_hasMGFUpperBoundAt","target":"declaration:BanditRLProof.OFUL.measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","target":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.CanonicalLinearSubgaussianEnvironmentLaw","target":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_log","target":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","target":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq","target":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_bestAction","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_bestAction","target":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","target":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_historyScore","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_historyScore","target":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","target":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_mean_bestAction_sub_clippedUCB_le","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_sum_clippedUCB_action_sub_mean_le","target":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.hedge_regret_le_log_card_div_add_eta_mul_mixedSquaredLoss_of_nonneg","target":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_sqrt","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareRealizedRegret_tail","target":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.measure_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","target":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.tsum_ofReal_geometricConfidenceShare","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedDeviation_sum_tail_predictableVariance_fixedTilt","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.selectedLossCenteredSecondMoment_le_one","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictableRealizedVariance_sum_le_horizon","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictable_highProbabilityRegret_tail_total_delta","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Concentration.tsum_ofReal_geometricConfidenceShare","target":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","target":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisEstimatedRegret_pointMass_le_stability_add_penalty","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Tsallis.integral_sum_sampledScheduledHalfTsallisPotentialStabilityAtTime_le_allRateBound","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisEstimatedRegret_eq_environmentRegret","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_allRateBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisPredictableEnvironmentRegret_pointMass_le_sqrtSchedule_log_iidLossState","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Tsallis.iidLossStateMeanGap_finiteArmIIDRewardVectorLoss_eq_gap","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteBanditModel","target":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_log","target":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret_tendsto_zero","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Tsallis.sampledScheduledHalfTsallisPredictableMovingComparatorEnvironmentRegret_le_oracleRestartSwitchCountSqrt","target":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_dominates_and_is_attained","target":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_occupancyGapRemaining","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_occupancyGapRemaining","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.integrable_stoppedValue_of_uniform_secondMoment_of_memLp_two_rounds","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.TransitionCountSummary.sum_aggregateTransitionCount_eq_aggregateVisitCount","target":"declaration:BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.adaptiveCumulativeAggregateTransitionCountAt_eq_sum","target":"declaration:BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_pos","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.measurable_clippedPolicyTable","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_bernsteinCoordinateFailureEvent_le","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.trajectoryMeasure_optimalTailFailureEvent_le","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQTable_dominatesOptimal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_all_successorBatchAligned_ae","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_positive_batchedInvSqrt_le","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.sum_high_batchedRatio_le_log","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSuccessorBellmanInnovation_compensated_hasCondMGFUpperBoundAt","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.trajectoryMeasure_recurrentBellmanInnovation_sum_ge_le","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","target":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.worstCaseExpectedRegret","target":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret","target":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDSampleMeanLaw","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_zero_error_event","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_gap_error_event","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_id_gaussianReal_zero","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_gap_sub_id_gaussianReal","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_le_exp","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability_le_exp","target":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.exists_alternative_le_average","target":"declaration:BanditRLProof.LowerBounds.exists_leastExploredAlternative","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.alternativeExpectedPullBudget_le","target":"declaration:BanditRLProof.LowerBounds.exists_leastExploredAlternative","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.baseEnvironmentRegret","target":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.changedEnvironmentRegretLowerBound","target":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","target":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.uniquelyDecodable_range","target":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.kraft_inequality","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropy","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo","target":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy","target":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","target":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_restrict_add_compl","target":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliKLCore_event_le","target":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuberCore","target":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale","target":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_nonneg","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","target":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","target":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough","target":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_gaussianPDFReal_one","target":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","target":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_one","target":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_sq","target":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_historyKL_le_half","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianArm_zero_two_mul","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.sixteen_div_twentySeven_le_exp_neg_half","target":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.add","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_add_le_rpow","target":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy","target":"declaration:BanditRLProof.LowerBounds.divergenceInfimum","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_ge","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_le_perturbed","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","target":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.measurableSet_oneArmMajorityPullEvent","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","target":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_probability_charge_le_expectedPseudoRegret","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_of_exp_testing_bound","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMeanChange_produces_gap_contract","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","target":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.chapter16GaussianChangedEnvironment_armKL","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal_eq_sum_expectedPulls","target":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.correlationSum_le_one","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_basis_basis","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.exists_abs_eq_maxAbsCoefficient","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sumAbs_eq","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.abs_inner_strictBasis_supportSignCombination_eq_one","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.orthonormal","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.norm_sq_supportSignCombination_eq_size","target":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gradientCoordinate","target":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_eq_bestMean_sub_policyValue","target":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_historyParameter","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_contract","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_sum_eq_zero","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_terms_summable","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sq_div_two_mul_sourceC_abs_div_two","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.exp_mul_le_sourceEqEight","target":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.TwoArmBoundedFixedMeanEnvironmentContract","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","target":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardTrajectorySuccessorPotential","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqTelescope","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseInitialUnconditionalRecurrence","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmFailureMass_sq","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum_ge","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_le_maxGap_mul_failureMass","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.failureMass_eq_successFailure_add_sq","target":"declaration:BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_sourceTheoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_eq_generated","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_contract","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne_piecewise","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_piecewise_bound","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.isStoppingTime_twoArmNthOptimalPullTime","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_eq_top_iff","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_of_time_eq","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_eq_of_time_eq","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.UCB.armStreamMeasure_map_fixedArmFinitePrefix_eq_pi","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_latentCoordinate_ae","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.UCB.armStreamMeasure_map_frestrictLe_eq_pi","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.trajectoryMixture_map_history_action_eq_compProd","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.canonicalMeasurableEnvironmentTrajectoryKernel_map_history_action_eq_compProd","target":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Measure.map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.UCB.armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_joint_eq_compProd","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.UCB.armStreamSelectedRewardKernel","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_map_frestrictLe_eq_native","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.nativeStationaryTrajectoryMeasure","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextPair_eq_compProd","target":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_latentCoordinate_ae","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_eq_native","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_map_optimalPullTimeRewardBlock_eq_latentMasked","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmLatentMaskedOptimalPullBlock_preimage_appendixCObservedPhaseEvent","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmAppendixCObservedPhaseEvent","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCPureLatentRewardEvent_eq_union_phase_missing","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.disjoint_twoArmAppendixCLatentPhaseEvent_missing","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDTrajectoryMeasure_appendixCGeneratedPhaseEvent_eq_latent","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCRewardPhaseProbability_eq_generated_add_missing","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.mem_twoArmAppendixCMissingPullLatentPhaseEvent_iff","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmAppendixCMissingPullLatentPhaseEvent_subset_terminalCountBelow","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_visible_eq_generated","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_probability_le_countBelow","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDMissingPullLatentPhase_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_charge_mul_probability_le_integral","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneStarvationEvent","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_sampledPseudoRegret_eq","target":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound_pos","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_pos","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_ge","target":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_pos","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression_ge_quarter","target":"declaration:BanditRLProof.LowerBounds.randomRegret_ge_quarter_of_clippingDecomposition","relation":"teaching prerequisite"},{"source":"declaration:BanditRLProof.Causal.ParallelParameters.allocation_cost_le","target":"milestone:CAUSAL-PARALLEL-ACTUAL-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.ParallelParameters.graph_designCost","target":"milestone:CAUSAL-PARALLEL-ACTUAL-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel","target":"milestone:CAUSAL-PARALLEL-ACTUAL-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.ParallelParameters.expected_simpleRegret_parallel_optimal","target":"milestone:CAUSAL-PARALLEL-ACTUAL-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_source_bound","target":"milestone:CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_explicit_rate","target":"milestone:CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_uniform","target":"milestone:CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.GraphModel.expected_simpleRegret_optimal","target":"milestone:CAUSAL-COMMON-ALPHABET-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.HOO.RegularCovering.expected_pseudoRegret_rate_family","target":"milestone:HOO-REWARD-FAMILY-RATE","relation":"certifies"},{"source":"declaration:BanditRLProof.HOO.RegularCovering.expected_actual_eq_pseudoRegret_family","target":"milestone:HOO-REWARD-FAMILY-RATE","relation":"certifies"},{"source":"declaration:BanditRLProof.HOO.RegularCovering.expected_actualRegret_rate_family","target":"milestone:HOO-REWARD-FAMILY-RATE","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_ne_top","target":"milestone:HEAVY-TAIL-GENALTI-FINITE-SUPREMUM","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.process_value_mem","target":"milestone:HEAVY-TAIL-GENALTI-FINITE-SUPREMUM","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.GenaltiAudit.normalized_sSup_two_eq","target":"milestone:HEAVY-TAIL-GENALTI-FINITE-SUPREMUM","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.literal_printed_bound_false","target":"milestone:HEAVY-TAIL-PRINTED-REGRET-COUNTEREXAMPLE","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.SourceCounterexample.kernel_raw_moment","target":"milestone:HEAVY-TAIL-PRINTED-REGRET-COUNTEREXAMPLE","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.adaptive_corrupted_clipped_mean_tail","target":"milestone:HEAVY-TAIL-CLIPPED-CORRUPTION-TRANSFER","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.observed_corrupted_clipped_mean_tail","target":"milestone:HEAVY-TAIL-CLIPPED-CORRUPTION-TRANSFER","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.arm_corrupted_clipped_mean_tail","target":"milestone:HEAVY-TAIL-CLIPPED-CORRUPTION-TRANSFER","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_expected_regret","target":"milestone:HEAVY-TAIL-CORRECTED-SOURCE-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robust_integral_count_le_budget","target":"milestone:HEAVY-TAIL-CORRECTED-SOURCE-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_upper_tail_sum","target":"milestone:HEAVY-TAIL-SOURCE-SCHEDULE-CONFIDENCE","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.SourcePolicy.robustMean_lower_tail_sum","target":"milestone:HEAVY-TAIL-SOURCE-SCHEDULE-CONFIDENCE","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail","target":"milestone:HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_lower_tail","target":"milestone:HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_upper_tail_log_sharp","target":"milestone:HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","relation":"certifies"},{"source":"declaration:BanditRLProof.HeavyTail.source_truncated_mean_lower_tail_log_sharp","target":"milestone:HEAVY-TAIL-SOURCE-CONFIDENCE-FOUR","relation":"certifies"},{"source":"declaration:BanditRLProof.pseudoRegret_eq_finset_sum_gap_mul_pullCount","target":"milestone:FOUNDATION-REGRET-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","target":"milestone:PROBABILITY-GENERATED-COND-MGF","relation":"certifies"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_tsum_of_uniform","target":"milestone:PROBABILITY-FINTYPE-GEOMETRIC-ALL-TIME-UNION","relation":"certifies"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_geometricConfidenceShare","target":"milestone:PROBABILITY-FINTYPE-GEOMETRIC-ALL-TIME-UNION","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeGeometricAllTimeBadEvent","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_geometricAllTime_abs_tail_ennreal_delta_trajMeasure","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Concentration.tsum_ofReal_telescopingConfidenceShare","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_iUnion_fintype_le_delta_of_telescopingConfidenceShare","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.successorArmEmpiricalMeanFintypeTelescopingAllTimeBadEvent","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_fintype_telescopingAllTime_abs_tail_ennreal_delta_trajMeasure","target":"milestone:PROBABILITY-GENERATED-FINTYPE-EMPIRICAL-MEAN-TELESCOPING-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingHistoryPolicy","target":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescopingPairHistory_eq_finitePairHistoryOfTrace","target":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorTelescoping_allTimeBadEvent_le_trajMeasure","target":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorTelescoping_allHorizonPullCount_of_not_badEvent","target":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.lintegral_ofReal_pseudoRegret_selectedPolicySuccessorTelescoping_le_trajMeasure","target":"milestone:UCB-FIXED-POLICY-TELESCOPING-ANYTIME-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_generatedActionPartialTrajectoryPairLawSource_trajMeasure","target":"milestone:PROBABILITY-ETC-UCB-ROUTE-SURFACE","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.historyStepKernelFamily_centeredReward_succ_hasCondSubgaussianMGF_trajMeasure","target":"milestone:PROBABILITY-ETC-UCB-ROUTE-SURFACE","relation":"certifies"},{"source":"declaration:BanditRLProof.ConditionalExpectationReward.actionRewardHistoryStepKernelFamily_successorArmEmpiricalMean_simultaneous_finiteArmTime_abs_tail_ennreal_delta_trajMeasure","target":"milestone:PROBABILITY-ETC-UCB-ROUTE-SURFACE","relation":"certifies"},{"source":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","target":"milestone:ETC-CANONICAL-SUBGAUSSIAN-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.ETC.argmaxCommitOracle_argmax_finRange","target":"milestone:ETC-CANONICAL-RAT-LEAST-TIE","relation":"certifies"},{"source":"declaration:BanditRLProof.ETC.argmaxCommitOracle_encode_le_of_score_le","target":"milestone:ETC-CANONICAL-RAT-LEAST-TIE","relation":"certifies"},{"source":"declaration:BanditRLProof.ETC.integral_real_pseudoRegret_explorationArgmaxGeneratedAction_le_canonicalSubGaussianArmPerArmIntegralRegretBoundReal","target":"milestone:ETC-LML-PORT","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.integral_real_pseudoRegret_selectedPolicySuccessorGeneratedUCBRegretAction_le_textbookGapSum_finiteArmSubgaussianLaws_without_proxy_positivity","target":"milestone:UCB-FINITE-ARM-SUBGAUSSIAN","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","target":"milestone:UCB-EXPECTED-AVERAGE-CONSISTENCY","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.measure_selectedPolicySuccessorLargeGapEvent_generatedUCB_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","target":"milestone:UCB-HORIZON-INDEXED-CANONICAL-CHAIN","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.measure_successorArmPullCount_selectedPolicySuccessorGeneratedUCBAction_gt_explicitPullThreshold_le_ennreal_delta_actionRewardTrajMeasure_centeredKernel","target":"milestone:UCB-HORIZON-INDEXED-CANONICAL-CHAIN","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedPseudoRegret_nonneg_and_le","target":"milestone:UCB-HORIZON-INDEXED-CANONICAL-CHAIN","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.selectedPolicySuccessorFiniteArmSubgaussianExpectedAveragePseudoRegret_tendsto_zero","target":"milestone:UCB-HORIZON-INDEXED-CANONICAL-CHAIN","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.integral_realKernelRegret_armStreamAction_le_lml_sum","target":"milestone:UCB-LML-PORT","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.sum_range_min_prefix_update_le_two_trace_average_log","target":"milestone:OFUL-ELLIPTICAL-POTENTIAL","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.fixedDirectionCompensatedScore_hasMGFUpperBoundAt","target":"milestone:OFUL-SELF-NORMALIZED-RIDGE-CONFIDENCE","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.measure_finiteHorizonRidgeEstimate_error_matrixNorm_gt_confidenceRadius_le","target":"milestone:OFUL-SELF-NORMALIZED-RIDGE-CONFIDENCE","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.finiteHistoryTelescopingScalarRidgeOptimisticAlgorithm","target":"milestone:OFUL-MEASURABLE-GENERATED-POLICY","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","target":"milestone:OFUL-ALL-TIME-CONFIDENCE","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.telescopingCanonicalStandardHighProbabilityPseudoRegret_nonneg_and_allHorizon_tail_le_explicitBound_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","target":"milestone:OFUL-ALL-HORIZON-HIGH-PROBABILITY-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.canonicalStandardExpectedAveragePseudoRegret_tendsto_zero","target":"milestone:OFUL-EXPECTED-AVERAGE-CONSISTENCY","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_endpoint_add_envelope_mul_delta_of_linearSubgaussianEnvironment_of_featureBound_le_regularization","target":"milestone:OFUL-BOUNDED-STOPPING-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.OFUL.integral_stoppedValue_canonicalStandardHighProbabilityPseudoRegret_nonneg_and_le_quadraticCoefficient_mul_stoppingTimeRoundSecondMoment_add_initialGap_mul_sqrt_stoppingTimeRoundSecondMoment_mul_sqrt_delta_and_stoppedViolation_measure_le_of_squareIntegrableFiniteStoppingTime","target":"milestone:OFUL-UNBOUNDED-STOPPING-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.PosteriorKernel.canonicalPosterior_kernel_ae_eq_condDistrib_of_pair_map_eq","target":"milestone:THOMPSON-POSTERIOR-KERNEL","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.canonicalSampler_condDistrib_action_ae_eq_bestAction","target":"milestone:THOMPSON-CANONICAL-SAMPLER","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.uniformReferenceThompsonAlgorithm_trajectory_condDistrib_action_ae_eq_bestAction","target":"milestone:THOMPSON-RECURSIVE-PROBABILITY-MATCHING","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_historyScore","target":"milestone:THOMPSON-BAYES-CLIPPED-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.integral_trajectoryBayesMeanRegret_eq_add_clippedUCB","target":"milestone:THOMPSON-BAYES-CLIPPED-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.canonicalLatentArmStreamTrajectory_reward_eq_rewardFromArmStream_ae","target":"milestone:THOMPSON-LATENT-STREAM-SUPPORT","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.IsOptimalMeanSelector","target":"milestone:THOMPSON-STATIONARY-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","target":"milestone:THOMPSON-STATIONARY-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.stationaryLatentArmStreamCanonicalTrajectoryMeasure_integral_trajectoryBayesMeanRegret_le","target":"milestone:THOMPSON-GENERAL-PORT","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictable_expectedRegret_le_four_mul_sqrt","target":"milestone:EXP3-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonBernsteinSquareBestArmRealizedRegret_tail","target":"milestone:EXP3-BEST-ARM-REALIZED-HIGH-PROBABILITY","relation":"certifies"},{"source":"declaration:BanditRLProof.Concentration.measure_iUnion_scheduled_deviation_ge_inter_variance_le_delta_of_fixedTilt_quadratic_tail","target":"milestone:CONCENTRATION-COUNTABLE-SCHEDULED-QUADRATIC-TAIL","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRealizedDeviationAllTimeFailureSet_le","target":"milestone:EXP3-PREDICTABLE-VARIANCE-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledRealizedDeviationGeometricAllTimeFailureSet_le","target":"milestone:EXP3-REALIZED-DEVIATION-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledPredictableRegretGeometricAllTimeFailureSet_le","target":"milestone:EXP3-PREDICTABLE-REGRET-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.sampledTrajectoryRealizedRegret_eq_predictableRegret_add_realizedDeviation","target":"milestone:EXP3-REALIZED-REGRET-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","target":"milestone:EXP3-REALIZED-REGRET-GEOMETRIC-ALL-TIME","relation":"certifies"},{"source":"declaration:BanditRLProof.Exp3.sampledPredictable_allHorizonDoubleVarianceProbabilisticSparseLossBestArmRealizedRegret_tail_of_sparsityFailure_le","target":"milestone:EXP3-SPARSE-ALL-HORIZON","relation":"certifies"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","target":"milestone:TSALLIS-IID-LOG","relation":"certifies"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDHistoryAdaptiveExpectedCorruptedRewardLawRegret_le_allRegimes","target":"milestone:TSALLIS-HISTORY-ADAPTIVE-CORRUPTION","relation":"certifies"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIndependentDriftingMeanDynamicRegret_le_allRegimes","target":"milestone:TSALLIS-DYNAMIC-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledOracleRestartHalfTsallisFiniteArmIndependentGlobalMeanChangeDynamicRegret_le","target":"milestone:TSALLIS-ORACLE-RESTART-GENERATED","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.MDP","target":"milestone:RL-FINITE-MDP-BELLMAN","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.MDP.optimalValueAt_dominates_and_is_attained","target":"milestone:RL-FINITE-MDP-BELLMAN","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.MarkovPolicy.expectedRegret_eq_occupancyGapRemaining","target":"milestone:RL-OCCUPANCY-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.MDP.expectedRegret_eq_occupancyGap_nonneg_and_optimalPolicy_zero","target":"milestone:RL-OCCUPANCY-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeStochasticEmpiricalOptimisticSource.exploratorySource_trajectoryMeasure_projectedCumulativeInverseSqrtPathSupport_decayingExplorationAverageRealizedBehaviorConsistency_allWindows_of_standardBorel","target":"milestone:RL-ADAPTIVE-REALIZED-CONSISTENCY","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_squareIntegrableFiniteStoppingTime","target":"milestone:RL-UNBOUNDED-HITTINGAFTER-L2","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveStochasticSampledEmpiricalOptimisticSource.selfConsistentScheduledCausalSource_inverseSqrtThresholdUnboundedHittingAfter_stoppedAverageRealizedBehaviorRegret_integrable_and_integral_le_threshold","target":"milestone:RL-UNBOUNDED-HITTINGAFTER-EXPECTED-UPPER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.sum_adaptiveCumulativeAggregateTransitionCountAt_eq_visitCountAt","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.clippedQRemaining_of_aggregateVisitCount_pos","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_policyAt_succ","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_simultaneousTransitionFailureEvent_le_fifth","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentQTableOfTrajectory_dominatesOptimal","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.AdaptiveEpisodeBatchSource.recurrentSource_generatedSuccessorPseudoRegret_le_charge_add_innovation","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.totalGeneratedPairCharge_le_explicit","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_bellmanInnovation_sum_ge_threshold_le_fifth","target":"milestone:RL-UCBVI-HOEFFDING-GENERATED-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.recurrentSource_trajectoryMeasure_cumulativeEpisodePseudoRegret_gt_canonicalRegretBound_le","target":"milestone:RL-UCBVI-HOEFFDING-CANONICAL-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.FiniteHorizonRL.AdaptiveCumulativeHoeffdingUCBVI.integral_cumulativeEpisodePseudoRegret_recurrentSource_le_canonicalRegretBound_add_failure","target":"milestone:RL-UCBVI-HOEFFDING-CANONICAL-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.Budget.isStoppingTime_budgetExhaustionTime_of_adapted","target":"milestone:BWK-STOPPING-FOUNDATION","relation":"certifies"},{"source":"declaration:BanditRLProof.Budget.isStoppingTime_budgetExhaustionTime_of_adapted","target":"milestone:BWK-FINAL-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.bernoulliKL","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.index","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.historyPolicy","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.pairHistory_eq_finitePairHistoryOfTrace","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.generatedIndexAt_le_selected_of_K_le","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.measure_generatedKLAllTimeBadEvent_le_trajMeasure","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.allHorizonPullCount_of_not_badEvent","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.KLUCB.lintegral_ofReal_pseudoRegret_generatedKLUCBBounded_le_trajMeasure","target":"milestone:KL-UCB-BOUNDED-GENERATED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.worstCaseExpectedRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal.mem_policyClass","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal.eq_minimaxExpectedRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.expectedRegret_le_worstCaseExpectedRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret_le_worstCaseExpectedRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.le_minimaxExpectedRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_alternative_le_average","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.alternativeExpectedPullBudget_le","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_leastExploredAlternative","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.baseEnvironmentRegret","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.changedEnvironmentRegretLowerBound","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half","target":"milestone:TEXTBOOK-PART-IV-CH13-BASIC-IDEAS-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance_pos","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanLaw","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDObservationLaw","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianCoordinateAverage","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDSumLaw","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDSampleMeanLaw","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_zero_error_event","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_gap_error_event","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_id_gaussianReal_zero","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_gap_sub_id_gaussianReal","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianReal_zero_Ici_le_exp_neg_sq_div_two_variance","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianReal_gap_Iio_half_le_exp_neg_sq_div_two_variance","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_le_exp","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanGapErrorProbability_le_exp","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMills_lower_integral","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMills_upper_integral","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_source_bounds","target":"milestone:TEXTBOOK-PART-IV-CH13-GAUSSIAN-TESTING-COMPANION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.uniquelyDecodable_range","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.kraft_inequality","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropy_nonneg","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.expectedCodeLength","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.expectedCodeLength_nonneg","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.huffmanCode_optimal","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_prefixCode_of_uniquelyDecodable","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.length_antitone","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.huffmanCode_entropy_sandwich","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_rate_tendsto_entropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_payload_interval","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.sourceBlock_code_family_limit_ge_entropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_ceilingLogPrefixCode","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_crossEntropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.entropyTerm_tendsto_zero_right","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_asymmetry","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_triangle_counterexample","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_eq_if","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_le_half_commonDensityAffinity_sq","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.half_commonDensityAffinity_sq_le_overlap","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_same_variance","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussian_testing_max_error_three_twentieths","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_absolutelyContinuous_of_integrable","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_probability_absolutelyContinuous_of_integrable","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_not_absolutelyContinuous","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_zero_iff","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_le","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.absolutelyContinuous_iff_atom_support","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.rnDeriv_mul_atom","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.rnDeriv_atom_eq_div","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_atom_support_mismatch","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_klFun","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_sum_log","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_if","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_top_iff","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_iSup_densityApproximation_trim","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_fin_observation_densityApproximation","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.rnDeriv_restrict_restrict","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_restrict_add_compl","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliKLCore_event_le","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exp_neg_half_bernoulliKLCore_le_affinity","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.half_binaryAffinity_sq_le_eventError","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuberCore","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_nonneg","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","target":"milestone:TEXTBOOK-PART-IV-CH14-INFORMATION-THEORY-LEAN-SPINE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianArm","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianBandit","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_gaussianPDFReal_one","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_one","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianArm_zero_two_mul","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_sq","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_le_half","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-KL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure_succ","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finiteHistoryPullCountENNReal","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_eq_expectedPullCountThrough","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_canonicalBanditHistoryMeasure_eq_sum_realizedExpectedPullCount_mul_armKL","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","target":"milestone:TEXTBOOK-PART-IV-CH15-SAME-POLICY-HISTORY-KL-DECOMPOSITION","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_map_le","target":"milestone:TEXTBOOK-PART-IV-CH15-EX15-7-DATA-PROCESSING-LEAF","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_observedBanditHistory_le_expectedPulls_sum","target":"milestone:TEXTBOOK-PART-IV-CH15-EX15-7-DATA-PROCESSING-LEAF","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_leastExploredAlternative","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.base_event_probability_lower_bound","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.changed_complement_probability_lower_bound","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianMinimax_base_changed_history","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_historyKL_le_half","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.sixteen_div_twentySeven_le_exp_neg_half","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianWorstCaseExpectedPseudoRegret_ge","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge","target":"milestone:TEXTBOOK-PART-IV-CH15-GAUSSIAN-MINIMAX-LOWER-BOUND","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentPolicyOver","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.add","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_add_le_rpow","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.divergenceInfimum","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_le","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum_le","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_le_perturbed","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_ge","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajorityPullEvent","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.measurableSet_oneArmMajorityPullEvent","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_mul_eq_exp","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exp_testing_bound_of_majority_regret_bounds","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_of_exp_testing_bound","target":"milestone:TEXTBOOK-PART-IV-CH16-CONSISTENCY-DINF-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.measurable_finiteHistoryGapPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_ne_top","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret_toReal","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.sum_canonicalRealizedExpectedPullCountThrough_general","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough_ne_top","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_ne_top","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegretReal","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_forces_gapPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_forces_gapPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_probability_charge_le_expectedPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","target":"milestone:TEXTBOOK-PART-IV-CH16-EVENT-REGRET-PRODUCERS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedRegret_div_log_ge","target":"milestone:TEXTBOOK-PART-IV-CH16-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","target":"milestone:TEXTBOOK-PART-IV-CH16-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","target":"milestone:TEXTBOOK-PART-IV-CH16-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.tailAtLeast","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialHighProbabilityThreshold","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_tailMass_ge_of_integral_ge","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.exists_cdfTail_ge_of_integral_ge","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.measureReal_diff_ge_delta","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression_ge_quarter","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.randomRegret_ge_quarter_of_clippingDecomposition","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_corollary17_2","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.noUniformGaussianRandomPseudoRegretTail_corollary17_3","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialClippedGaussianReward","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialCenteredNoiseLaw","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_eq17_8","target":"milestone:TEXTBOOK-PART-IV-CH17-FIRST-MOMENT-AND-TAIL-DEPENDENCY-SLICE","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_corollary17_2","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.noUniformGaussianRandomPseudoRegretTail_corollary17_3","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_eq17_8","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_history_marginal","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_pull_le_half_claim17_6","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount_tail_claim17_7","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialFullRandomRegret_ge_boundary_eq17_8","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.integrable_adversarialTableRandomRegret","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialTable_strictTail_eq_one_sub_CDF","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_theorem17_4","target":"milestone:TEXTBOOK-PART-IV-CH17-SOURCE-TERMINALS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt","target":"milestone:TEXTBOOK-PART-IV-THEOREM-13-1-GAUSSIAN-MINIMAX","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.observedBefore","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.outstandingAt","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.observedBefore_disjoint_outstandingAt","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.observedBefore_union_outstandingAt","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.card_observedBefore_add_card_outstandingAt","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.outstandingCount","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.maxOutstandingBeforeThrough","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.outstandingCount_le_round","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.outstandingCount_le_maxOutstandingBeforeThrough","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.oneBasedDelayShift","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperMissingAtEnd","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperMissingAtEnd_eq_outstandingAt_oneBasedDelayShift","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_eq_outstandingCount_oneBasedDelayShift","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_le_round","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperSigmaMaxThrough","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.paperMissingCount_le_paperSigmaMaxThrough","target":"milestone:DELAYED-FEEDBACK-SOURCE-ACCOUNTING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.ActionTimeView","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.ActionTimeView.ext","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.CausalDecisionRule","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_lt","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_pastAction_of_not_lt","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_mem","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_observedLoss_of_not_mem","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_outstanding_loss_hidden","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.actionTimeViewAt_eq_of_observation_equivalent","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.causalDecision_eq_of_observation_equivalent","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.newlyObservedBefore","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.observedBefore_mono","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.processed_disjoint_newlyObservedBefore","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.processed_union_newlyObservedBefore","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.processAllNew","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.processAllNew_eq_observedBefore","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.previousObservedBefore_subset_current","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.processAllNew_from_previous_eq_current","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.outstandingAt_disjoint_newlyObservedBefore","target":"milestone:DELAYED-FEEDBACK-CAUSAL-PROCESSING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.inactiveArms","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.activeEqualShare","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_active","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_of_inactive","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.activeEqualShare_nonneg","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOProbability_nonneg","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sum_delayedSAPOProbability_eq_one","target":"milestone:DELAYED-SAPO-ACTIVE-ALLOCATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.eliminated","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_eliminated_iff","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.mem_remainingActive_iff","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.OptimalArmSurvivalCertificate","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.optimal_mem_remainingActive_of_certificate","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.remainingActive_nonempty_of_certificate","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminationSnapshot.sum_delayedSAPOProbability_after_elimination_eq_one","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.probability_nonnegative","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.sum_probability_eq_one","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.finiteActionDistribution","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOAllocation.actionMeasure_isProbabilityMeasure","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_isProbabilityMeasure","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.causalDelayedSAPOActionMeasureRule_eq_of_observation_equivalent","target":"milestone:DELAYED-SAPO-ELIMINATION-ACTION-LAW","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.sourceUcbStar","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.EliminationGoodEvent","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalMean_le_ucbStar_of_eliminationGoodEvent","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalArmSurvivalCertificate_of_eliminationGoodEvent","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimal_mem_remainingActive_of_eliminationGoodEvent","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.optimalSurvivalEventSet","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.eliminationGoodEventSet_subset_optimalSurvivalEventSet","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOSourceConfidenceSnapshot.measure_optimalSurvivalEventSet_compl_le_of_goodEvent","target":"milestone:DELAYED-SAPO-GOOD-EVENT-D9-PROJECTION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventComponent","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.componentFailure","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.failureSet","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceGoodEventSet_compl","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_sum","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.linearFailureBudget","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.doubleLinearFailureBudget","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sourceComponentFailureBudget","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.sum_sourceComponentFailureBudget_eq_nine_div","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_sourceGoodEventSet_compl_le_nine_div","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_eliminationGoodEventSet_compl_le_nine_div","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOGoodEventFailureFamily.measure_optimalSurvivalEventSet_compl_le_nine_div","target":"milestone:DELAYED-SAPO-D8-D9-ASSEMBLY","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_antitone","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_of_count_le_96_mul_scale","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.one_le_ten_mul_sourceEmpiricalWidthScale_two_log_of_small_count","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_one","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_one_four","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_one_le_four","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.not_sourceEmpiricalWidthScale_horizon_four_one_le_four","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.eight_mul_empiricalWidth_lt_gap_of_mem_eliminated","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.gap_le_sixteen_mul_empiricalWidth_of_mem_remainingActive_of_large_or_small_count","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_large_or_small_count","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.surrogateGap_le_gap","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_two_mul_surrogateGap","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.d12_gap_ordering_chain","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOD10D12GapOrderingContract.gap_le_twenty_mul_gap_of_eliminationPrefixIndex_le","target":"milestone:DELAYED-SAPO-D10-D12-GAP-ORDERING-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.processedPullCount","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefix.expectedPullMass_eq_of_active_throughout","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_nonneg","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_one","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.sourceEmpiricalWidthScale_le_three_of_count_le_eight_mul","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.expectedPullMass_eq_of_mem_active","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.quarter_count_sub_six_log_le_count_of_mem_active","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.eighth_count_le_count_of_large_count","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_three_of_large_count","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.empiricalWidth_le_ten_of_mem_active","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.ucbStar_le_empiricalMean_add_width","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.activeArmGapBranch","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedPrefixCountCertificate.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_countCertificate","target":"milestone:DELAYED-SAPO-D1-ACTIVE-COUNT-WIDTH-PRODUCER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefix","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalWidthAt","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.empiricalUpperAt","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toConfidenceSnapshot","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.D4CountClause","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.currentActive_subset_activeAt_sourceIndex","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.toProcessedPrefixCountCertificate","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOProcessedTraceSummary.gap_le_twenty_mul_gap_at_earlier_elimination_snapshot_of_traceSummary","target":"milestone:DELAYED-SAPO-PROCESSED-TRACE-SUMMARY-ADAPTER","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.source_le_roundStart","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOStructuralRoundState.currentActive_subset_activeAtSourceRound","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.sourceRound_not_mem","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_nodup","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedOrder_available","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.extendedSource_le_roundStart","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.currentActive_subset_extendedSourceActive","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.toPreEliminationSummary","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.line8RemainingActive_subset_currentActive","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_currentActive_subset_before","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchProcessOne.afterLine8_preserves_roundStart","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-PROCESS-ONE","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.processedOrder_toFinset_eq_observedBefore","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActionRound","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_processedOrder","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchRoundClose.nextRoundState_currentActive","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralStep","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPONoSwitchStructuralReachable","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralStep","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.currentActive_subset_of_structuralReachable","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.mem_earlierRemainingActive_of_laterEliminated","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.gap_le_twenty_mul_gap_of_ordered_no_switch_eliminations","target":"milestone:DELAYED-SAPO-ORDERED-NO-SWITCH-TRACE-ORDERING","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.finiteAverageGap","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.aboveTwiceAverageGap","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.two_mul_card_aboveTwiceAverageGap_le","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceStochasticLossGap","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceStochasticLossGap_nonneg","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.two_mul_card_sourceStochasticLossGap_aboveTwiceAverage_le","target":"milestone:DELAYED-SAPO-D11-NONNEGATIVE-GAP-HALF-SET","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOInitialEliminatedProbability","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.delayedSAPOInitialPhaseTarget","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.sourceEmpiricalWidthScale_pos","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmBank","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.ActiveArmsUninitialized","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_some_iff","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeIfEliminated_eq_none_iff","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_mem","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_of_not_mem","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.mem_eliminated_of_initializeNewlyEliminated_ne_prior","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_eq_prior_of_mem_remainingActive","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.prior_eq_none_of_mem_eliminated","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.remainingActive_uninitialized_after_initialize","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_arm","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationRound","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_eliminationProcessedOrder","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_errorCount","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseIndex","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_phaseSamples","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.ofProcessOne_processedAtProbabilityLevel","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_pos","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialProbability_le_one","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_nonneg","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.surrogateGap_pos","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_nonneg","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initialPhaseTarget_pos","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.DelayedSAPOEliminatedArmInitialization.initializeNewlyEliminated_spec_of_mem","target":"milestone:DELAYED-SAPO-LINE10-ELIMINATED-ARM-INITIALIZATION","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract","target":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim","target":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim","target":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.stochasticClaim_iff_shared_fields","target":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","relation":"certifies"},{"source":"declaration:BanditRLProof.DelayedFeedback.SameAlgorithmMultiRegimeContract.adversarialClaim_iff_shared_fields","target":"milestone:NEURIPS-2025-DELAYED-BOBW-CENTRAL-ENDPOINTS","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQSet_bddAbove","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.le_sourceQ_of_mem","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_le_norm","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_nonneg","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_zero","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.abs_inner_le_sourceQ_of_mem","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceQ_eq_zero_of_atom_orthogonal","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceR","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.sourceRSet_not_bddAbove_of_nonzero_atom_orthogonal","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.correlationSum_le_one","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_basis_basis","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_le_maxAbsCoefficient","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_nonneg","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.exists_abs_eq_maxAbsCoefficient","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportCombination","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.signedSupportAtom_mem","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_le","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_basis","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_signedSupportAtom","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportCombination_eq","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.abs_coefficientSign","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.supportSignCombination","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.maxAbsCoefficient_coefficientSign","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceQ_supportSignCombination","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_supportSignCombination","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_mul_sourceQ","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.inner_supportCombination_le_sumAbs_of_sourceQ_le_one","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceRSet_bddAbove","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.sourceR_supportCombination_eq","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.orthonormal","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_coefficientSign","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.coefficientSign_mul_self","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctSupport.norm_sq_supportSignCombination","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsSuccinctAt","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.IsStrictlySuccinctAt","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceR_eq_sumAbs","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.inner_supportSignCombination_eq_sumAbs","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sourceQ_supportSignCombination_eq_one","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.norm_sq_supportSignCombination_eq_size","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.sumAbs_eq","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.abs_inner_strictBasis_supportSignCombination_eq_one","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.SuccinctRepresentation.strictSize_le","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.StrictSuccinctRepresentation.abs_coefficient_pos","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.succinctSize_ge_strictSize","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.LowerBounds.Succinct.SuccinctUnitSystem.strictlySuccinctSize_unique","target":"milestone:NEURIPS-2025-SUCCINCT-LOWER-BOUND-GEOMETRY-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxDenominator","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxDenominator_pos","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_pos","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_nonneg","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_sum","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_le_one","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceIncrement_eq_indicator","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sum_sourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.policyValue","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gradientCoordinate","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_eq_bestMean_sub_policyValue","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapCoordinate","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.gapExpectedIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expectedSourceIncrement_eq_gapExpectedIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_ge_minGap_mul_failureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.gapExpectedIncrement_best_ge","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.bestParameterIncrementSum_ge","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceExpectedPseudoRegret","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.instantaneousGap_le_maxGap_mul_failureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.failureMass_eq_successFailure_add_sq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceRegretDecomposition_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_softmaxProbability","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_sourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_succ","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_historyParameter","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxFiniteActionDistribution","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historySoftmaxDistributionSource","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyAlgorithm","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyAlgorithm_policy","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_sourceIncrement_eq_expectedSourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_sourceIncrement_eq_expectedSourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_expectedSourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_sourceIncrement_eq_gapCoordinate","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryKernel","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action_zero_given_environment","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_action","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryMeasure_condDistrib_nextPair_given_environment_prefix","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_sum_eq_initial","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_zeroInitialization_sum","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_succ","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_sum_eq_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmParameterAt_one_eq_neg_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_one_eq_one_sub_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.finTwo_one_eq_neg_zero_of_sum_eq_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zero_div_one_sub_zero_eq_exp_two_mul","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.exp_two_mul_zero_mul_one_sub_softmaxProbability_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.exp_neg_two_mul_zero_mul_softmaxProbability_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_exp_two_mul_failure_eq_success","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmProbabilityAt_zero_div_failure_eq_exp_two_mul","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.historyParameter_exp_two_mul_zero_eq_odds","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.two_mul_abs_pow_div_factorial_add_two_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceC","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_terms_summable","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_nonneg","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_mono","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceC_le_exp_two_mul","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo_terms_summable","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.exp_eq_one_add_self_add_expTailTwo","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.expTailTwo_le_of_abs_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sq_div_two_mul_sourceC_abs_div_two","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.exp_mul_le_sourceEqEight","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_exp_mul_le_sourceEqEight_of_ae_abs_le_one","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentInitialPairKernel_exp_actionReward_le_sourceEqEight_of_mean","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_historyStepKernel_exp_actionReward_le_sourceEqEight_of_mean","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableEnvironmentHistoryStepKernel_exp_actionReward_le_sourceEqEight_of_mean","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardQ","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseQ","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardQ_mul_reward_eq_sourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseQ_mul_reward_eq_sourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardEqEightRemainder_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseEqEightRemainder_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_forwardSuccessor_le_add_success_sq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmHistoryStepKernel_exp_inverseSuccessor_le_sub_failure_sq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.softmaxProbability_zeroInitialization_finTwo","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardRecurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseRecurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardSuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseSuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardRecurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseRecurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.TwoArmBoundedFixedMeanEnvironmentContract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_forwardSuccessor_le_of_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_measurableTwoArmHistoryStepKernel_inverseSuccessor_le_of_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmEnvironmentPrefix","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNextPair","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmEnvironmentPrefix","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNextPair","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixSigma_mono","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixFiltration","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardTrajectorySuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInverseTrajectorySuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_forwardSuccessor_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.trajectoryPrefix_condDistrib_integral_inverseSuccessor_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_forwardIncrement_le_of_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialPairKernel_exp_inverseIncrement_le_of_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_zero_abs_le_one_ae","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_succ_abs_le_one_ae","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_reward_abs_le_one_ae","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_prefix_rewards_abs_le_one_ae","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_le_abs_reward_of_mem_Icc","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.abs_sourceIncrement_softmax_le_abs_reward","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmInitialPairKernel_sourceIncrement_of_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_measurableTwoArmHistoryStepKernel_sourceIncrement_of_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.abs_historyParameter_zeroInitialization_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessorPotential_eq_exp_historyParameter","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessorPotential_eq_exp_historyParameter","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardTrajectorySuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseTrajectorySuccessorPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_ae_eq_integral_condDistrib","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardTrajectorySuccessor_condExp_le_recurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseTrajectorySuccessor_condExp_le_recurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInversePotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSuccessProbability","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFailureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectoryParameterZero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmForwardPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInversePotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessProbability","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmFailureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInversePotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessProbability_sq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmFailureMass_sq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardSuccessor_eq_nextPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseSuccessor_eq_nextPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmForwardRecurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmInverseRecurrenceBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardUnconditionalRecurrence","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseUnconditionalRecurrence","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmScalarForwardIterate","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmScalarInverseTelescope","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqTelescope","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseFailureMassSqSum_le_initial_div","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialForwardPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialInversePotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialForwardPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialInversePotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardPotential_zero_eq_initial","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInversePotential_zero_eq_initial","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_zero_kernel_eq_initial","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInversePotential_zero_kernel_eq_initial","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardInitialUnconditionalRecurrence","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInverseInitialUnconditionalRecurrence","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmForwardFiniteIteration_from_source_initial","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFullFailureMassSqSum_le","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_apply","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDRewardKernel_isMarkov","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_initialFeedback_apply","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_feedback_apply","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDReward_aestronglyMeasurable","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDEnvironment_contract","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFixedIIDHistoryStepKernel_sourceIncrement_eq_gapCoordinate","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmTrajectorySourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectorySourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_integral_condDistrib","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectorySourceIncrement_condExp_ae_eq_successFailure","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryParameterZero_succ","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmTrajectoryParameterZero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectorySourceIncrement_eq_successFailure","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_succ","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInitialSourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmInitialSourceIncrement","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialSourceIncrement_eq_quarter_gap","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_zero","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_eq_successFailureSum","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSuccessFailureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSuccessFailureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmFailureMass_eq_successFailure_add_sq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_eq_parameter_add_failureSq","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_half_log_forwardPotential","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessProbability_sq_le_one","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmForwardPotential_le_source_bound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmTrajectoryParameterZero_le_source_log_bound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmActionGap","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmActionGap","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmInitialActionGap_eq_half","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSuccessorActionGap_eq_failureMass","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_eq_generated","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedExpectedPseudoRegret_le_sourceTheoremOne","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_sourceTheoremOne","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_theoremOne","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepOneMargin_pos","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFourSurvivalLowerBound_pos","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_ge","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourStepFour_survivalMass_pos","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourFiniteGeometricPhaseMass_le_inv","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.theoremFourFiniteTransientMass_le_inv","target":"milestone:NEURIPS-2025-STOCHASTIC-GRADIENT-BANDIT-MECHANISM-AUDIT","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmActionGap_le_gap","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_le_gap_mul_horizon","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmSampledPseudoRegret","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integral_twoArmSampledPseudoRegret_le_gap_mul_horizon","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceTheoremOne_margin_of_two_mul_eta_sourceC_le","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.sourceTheoremOne_constant_le_inv_eta","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_pos","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_sq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_le_one","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneRate","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneRate_nonneg","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_horizon_eq_rate","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneEta_mul_rate_eq_log","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_inv_eta_le_inv_log_two_mul_rate","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_log_argument_le_horizon_pow_four","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_log_term_le_two_mul_rate","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_gap_mul_horizon_le_exp_constant_mul_rate","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOneAbsoluteConstant","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.corollaryOne_piecewise_bound","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne_piecewise","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDDirac_corollaryOne","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmGeneratedAction","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneThreshold","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneTriggerEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmGeneratedAction","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmOptimalPullCount","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneTriggerEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmStepOneStarvationEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmTerminalOptimalPullCountEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurableSet_twoArmOptimalPullCountBelowEvent","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_eq_iUnion_terminalCount","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTerminalOptimalPullCountEvent_sampledPseudoRegret_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCountBelowEvent_charge_mul_probability_le_integral","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_nonneg","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_suboptimalPullCount","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_add_suboptimalPullCount_eq_horizon","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmSampledPseudoRegret_eq_gap_mul_horizon_sub_of_optimalPullCount_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.mem_twoArmStepOneStarvationEvent_of_lowProbability_noFurtherOptimalPull","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_sampledPseudoRegret_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.integrable_twoArmSampledPseudoRegret_of_finiteMeasure","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmStepOneStarvationEvent_charge_mul_probability_le_integral","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDStepOneStarvationEvent_charge_mul_probability_le_integral","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixGeneratedAction","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixGeneratedAction","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmPrefixOptimalPullCount","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmPrefixOptimalPullCount_environmentPrefix_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmInclusiveOptimalPullCountProcess","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmInclusiveOptimalPullCountProcess","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.isStoppingTime_twoArmNthOptimalPullTime","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullTime","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_eq_top_iff","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_succ_of_nthOptimalPullTime_eq_top","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmOptimalPullCount_lt_of_fin_nthOptimalPullTime_eq_top","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_succ_eq_of_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_action_eq_zero","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_count_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullTime_spec","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmGeneratedReward","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullReward","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_of_time_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.adapted_twoArmSuccessProbability","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.measurable_twoArmNthOptimalPullSuccessProbability","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullSuccessProbability_eq_of_time_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.armStreamMeasure_map_fixedArmFinitePrefix_eq_pi","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_fixedArmFinitePrefix_eq_pi","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmTrajectoryMeasure_dirac_eq_map_trajectoryKernel","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.stationaryRewardKernelAt_twoArmFixedIIDRewardKernel_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmFixedIIDLatentTrajectoryMeasure_map_optimalPrefix_eq_pi","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.StochasticGradientBandit.twoArmNthOptimalPullReward_eq_latentCoordinate_ae","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.armStreamMeasure_map_frestrictLe_eq_pi","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.extendArmStreamFinitePrefix","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.measurable_extendArmStreamFinitePrefix","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.extendArmStreamFinitePrefix_apply_of_le","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_of_streamPrefix_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixKernel","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_eq_prefixKernel_comap","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_stream_visiblePrefix_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.UCB.armStreamMeasure_map_output_coordinate_compProd_comap_without_eq_prod","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamVisiblePrefixNextAction","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamVisibleNextReward","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchKernel","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCap","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountCap","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.realHistoryPullCount_extendPairHistorySucc","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.LatentArmStreamVisiblePrefixNextActionBranchLocality","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod_of_locality","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryMeasure_map_visiblePrefix_nextAction_eq_compProd","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_eq_selectedCoordinate_ae","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_prefix_next_eq_compProd","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamFeedback_eq_of_withoutCoordinate_eq_of_selectedCoordinate_ne","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_eq_of_withoutCoordinate_eq_of_target_count_lt","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamNextActionNeSet","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamInitialSafeArmSet","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamNextActionNeSet","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_restrict_nextActionNe_eq_of_withoutCoordinate_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_zero","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.singletonPairHistory_preimage_latentArmStreamPrefixCountCap_zero","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCapLocality_zero","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountLt","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountLt","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountEq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamPrefixCountEq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.mem_latentArmStreamPrefixCountCap_extendPairHistorySucc_iff","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamPrefixCountCap_of_extendPairHistorySucc_mem","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.selectedCoordinate_ne_of_extendPairHistorySucc_mem_prefixCountCap","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamSuccessorCountCap_preimage","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamSuccessorCountCapSection","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurableSet_latentArmStreamSuccessorCountCapSection","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.historyStepKernel_apply_restrict_successorCountCap_eq_of_withoutCoordinate_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_succ","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamTrajectoryKernel_map_frestrictLe_restrict_countCap_eq_of_withoutCoordinate_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality_of_prefixCountCapLocality","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextActionBranchLocality","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_coordinate_branch_eq_prod","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.measurable_latentArmStreamSelectedCoordinate","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_branch_eq_prod","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_mixed_eq_compProd","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisiblePrefixNextAction_selectedCoordinate_eq_compProd","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_joint_eq_compProd","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleNextReward_condDistrib_ae_eq_nu","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_joint_eq_compProd","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Thompson.latentArmStreamVisibleTrajectoryMeasure_nextReward_condDistrib_ae_eq_nu","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Measure.compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.Measure.map_compProd_restrict_eq_of_base_restrict_eq_of_fiber_restrict_eq","target":"milestone:NEURIPS-2025-SGB-PHASE-TRANSITION-FOLLOWON","relation":"certifies"},{"source":"declaration:BanditRLProof.CUCB.SourceModel.theorem_one_refined_regret","target":"milestone:CUCB-FULL-REPAIRED-RATES","relation":"certifies"},{"source":"declaration:BanditRLProof.CUCB.SourceModel.theorem_two_deterministic","target":"milestone:CUCB-FULL-REPAIRED-RATES","relation":"certifies"},{"source":"declaration:BanditRLProof.CUCB.SourceModel.theorem_two_probabilistic","target":"milestone:CUCB-FULL-REPAIRED-RATES","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_source_bound","target":"milestone:CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_explicit_rate","target":"milestone:CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_uniform","target":"milestone:CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.Causal.NodeGraphModel.expected_simpleRegret_optimal","target":"milestone:CAUSAL-NATIVE-HETEROGENEOUS-EXPECTED-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.transition_unfixed_hazard_quarter","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-OCCUPATION","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.unfixed_survival","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-OCCUPATION","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.expected_unfixed_occupation_le","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-OCCUPATION","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.roundPseudoRegret_le_twice_unfixed","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.expectedCoordinationRegret_real_le","target":"milestone:MULTI-AGENT-STATIC-COORDINATION-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.source_expected_learnerRegret_coarse","target":"milestone:MULTI-AGENT-STATIC-LEARNER-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.realizedLearnerRegret_expected_eq_pseudo","target":"milestone:MULTI-AGENT-STATIC-LEARNER-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.MusicalChairs.source_expected_visibleRegret_le","target":"milestone:MULTI-AGENT-STATIC-LEARNER-REGRET","relation":"certifies"},{"source":"declaration:BanditRLProof.MOSS.canonicalGapExpectedRegret_le","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.subgaussianMinimax_sandwich","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.moss_nearMinimax","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.Concentration.measure_exists_le_independent_partialSum_ge_le_subgaussian","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.MOSS.historyAlgorithm","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.MOSS.selected_index_gt_mean_add_half_gap","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMills_lower_integral","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMills_upper_integral","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanZeroErrorProbability_source_bounds","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.worstCaseExpectedRegret","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsMinimaxOptimal","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanVariance_pos","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDObservationLaw","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDSumLaw","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianIIDSampleMeanLaw","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_zero_error_event","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.twoPointGaussianThresholdDecision_gap_error_event","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.hasSubgaussianMGF_id_gaussianReal_zero","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianSampleMeanThresholdRisk_le_exp","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.expectedRegret_le_worstCaseExpectedRegret","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.minimaxExpectedRegret_le_worstCaseExpectedRegret","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.le_minimaxExpectedRegret","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_alternative_le_average","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.alternativeExpectedPullBudget_le","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_leastExploredAlternative","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.baseEnvironmentRegret","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.changedEnvironmentRegretLowerBound","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half_sub_error","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.max_base_changed_regretLowerBound_ge_half","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge_one_div_fiftyFour_sqrt","target":"spine:13","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.BinaryPrefixCode.kraft_inequality","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.discreteEntropyBaseTwo_eq_div_log_two","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.expectedCodeLength","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.huffmanCode_optimal","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_prefixCode_of_uniquelyDecodable","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsOptimalPrefixCode.length_antitone","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.huffmanCode_entropy_sandwich","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_rate_tendsto_entropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.arithmeticBlockCode_payload_interval","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.sourceBlock_code_family_limit_ge_entropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_ceilingLogPrefixCode","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.fixedLength_uniformPowerTwo_optimal","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.uniform_three_fixedLength_not_optimal","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_crossEntropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.entropyTerm_tendsto_zero_right","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_absolutelyContinuous_of_integrable","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_of_probability_absolutelyContinuous_of_integrable","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_top_of_not_absolutelyContinuous","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_ne_top_iff","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_eq_zero_iff","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_trim_le","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_sum_log","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_if","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_eq_top_iff","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.finitePartitionRelativeEntropy_eq_relativeEntropy","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_commonDensity_eq_if","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_commonSigmaFiniteDominatingMeasure","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_finite_lt_top_iff_ac","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_asymmetry","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_triangle_counterexample","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_le_half_commonDensityAffinity_sq","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.half_commonDensityAffinity_sq_le_overlap","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.commonDensityOverlap_le_testingError","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_same_variance","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussian_testing_max_error_three_twentieths","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.rnDeriv_restrict_restrict","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.relativeEntropy_restrict_add_compl","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bernoulliRelativeEntropy_event_le","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.binaryBretagnolleHuber","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_antitone","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuber","target":"spine:14","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_compProd_same_left_eq_lintegral_klDiv_of_measurable","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_historyStep_samePolicy_eq_iterated_lintegral_armKL_general","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalBanditHistoryMeasure","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalRealizedExpectedPullCountThrough","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianArm","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianBandit","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.log_gaussianPDFReal_div_gaussianPDFReal_one","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.llr_gaussianReal_one_ae","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.integrable_llr_gaussianReal_one","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_gaussianReal_one","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_unitGaussianArm_zero_two_mul","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_sq","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_informationExponent_eq_half","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianMinimaxGap_le_half","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_sum","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_map_le","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.klDiv_observedBanditHistory_le_expectedPulls_sum","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.UnitGaussianBanditEnvironment","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianExpectedPseudoRegret","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_gaussianMinimax_historyKL_le_half","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.base_event_probability_lower_bound","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.changed_complement_probability_lower_bound","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.sixteen_div_twentySeven_le_exp_neg_half","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.finiteArmedGaussianMinimaxLowerBound","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianMinimaxExpectedPseudoRegret_ge","target":"spine:15","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentPolicyOver","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.add","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_add_le_rpow","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.IsConsistentRegret.eventually_log_add_div_log_le","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.divergenceInfimum","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.divergenceInfimum_le","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.parametricDivergenceInfimum_le","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_le_perturbed","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.unitGaussianDivergenceInfimum_eq","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.banditHistoryRelativeEntropy_eq_expectedPulls_mul_of_only_arm_changed","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajorityPullEvent","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.bretagnolleHuberScale_expectedPulls_mul_armKL_le_majorityErrors","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.finiteHistoryGapPseudoRegret","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.canonicalGapExpectedPseudoRegret_eq_sum_expectedPulls","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_probability_charge_le_expectedPseudoRegret","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMajority_compl_probability_charge_le_expectedPseudoRegret","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_gapPseudoRegret_of_only_arm_changed","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_of_exp_testing_bound","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.FiniteMeanBanditEnvironment","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.oneArmMeanChange_produces_gap_contract","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.expectedPullCount_ge_log_regret_changeOfMeasure","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.UnitVarianceGaussianBanditEnvironment","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianExpectedRegret_ge_finiteTimeInstanceDependent","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedRegret_div_log_ge","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.consistentPolicy_liminf_expectedPull_div_log_ge_inv_dInf","target":"spine:16","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.tailAtLeast","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.stochasticHighProbabilityThreshold","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.stochasticMinimaxHighProbabilityThreshold","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialHighProbabilityThreshold","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.exists_cdfTail_ge_of_integral_ge","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.measureReal_diff_ge_delta","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRegretLowerExpression_ge_quarter","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.randomRegret_ge_quarter_of_clippingDecomposition","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_theorem17_1","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.gaussianRandomPseudoRegret_ge_corollary17_2","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.noUniformGaussianRandomPseudoRegretTail_corollary17_3","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_eq17_8","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialNoiseHistoryJoint_pull_le_half_claim17_6","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialFullBoundaryCount_tail_claim17_7","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.integrable_adversarialTableRandomRegret","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.LowerBounds.adversarialRandomRegret_ge_theorem17_4","target":"spine:17","relation":"maps to source"},{"source":"declaration:BanditRLProof.Exp3.measure_sampledRealizedRegretGeometricAllTimeFailureSet_le","target":"group:proof-laboratory","relation":"benchmarked in"},{"source":"declaration:BanditRLProof.Tsallis.integral_sampledScheduledHalfTsallisFiniteArmIIDRewardLawRegret_le_log","target":"group:proof-laboratory","relation":"benchmarked in"},{"source":"declaration:BanditRLProof.OFUL.measure_telescopingCanonicalHistoryTrajectory_allTimeConfidenceFailureSet_le_of_linearSubgaussianEnvironment","target":"group:proof-laboratory","relation":"benchmarked in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.tangentPairing_add_const_of_isSimplexTangent","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.sum_weight_mul_sub_weightedCenter_eq_zero","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition_of_centered","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_decomposition","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_center_le","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_eq_center_iff","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_decomposition","target":"group:proof-laboratory","relation":"audited in"},{"source":"declaration:BanditRLProof.CurvatureNoiseGap.weightedShiftEnergy_add_le_two","target":"group:proof-laboratory","relation":"audited in"}]}